Reinforcement Learning · LLM Reasoning · AI Agents

SDPG vs SDAR vs AntiSD: Three Ways to Spend a Privileged Teacher in RLVR

RLVR gives one bit per trajectory. SDPG copies a hint-conditioned teacher on correct trajectories, SDAR gates the copy per token, AntiSD rewards disagreement: three ways to mine dense signal from one privileged context.

SDPG vs SDAR vs AntiSD: Three Ways to Spend a Privileged Teacher in RLVR

The privileged-teacher trick

Reinforcement learning from verifiable rewards has a well-known bandwidth problem: a long chain of thought earns one scalar at the end, and that scalar has to be smeared back across hundreds of tokens. See GRPO vs PPO for why the clipped group-relative baseline became the default anyway. The three papers in this guide attack the same bottleneck with the same raw material: the same model, run twice, once with extra context it will never see at deployment.

Call the deployable policy the student and the context-augmented copy the teacher. The teacher is never a bigger model. It is the student conditioned on information that is legitimate at training time but unavailable at inference: the ground-truth answer plus a worked solution, retrieved skills from past experience, or a partial solution used as a hint. The second forward pass costs almost nothing in weights, because it reuses the student’s own parameters, which is why this family of methods keeps appearing in 2026 RLVR papers.

What separates the three methods is what they do with the teacher’s output:

  • SDPG copies it: minimize the KL from student to teacher on trajectories the verifier already marked correct, and use a reference-policy KL term to stop the copy from collapsing the policy.
  • SDAR copies it, but only where the teacher is actually right: a per-token sigmoid gate, because in multi-turn agentic settings the teacher is worse than the student on more than half of all tokens.
  • AntiSD refuses to copy it at all: it treats the teacher’s disagreements with the student as a reward, computed as the pointwise mutual information between the next token and the privileged context.

Read together they form a spectrum: unconditional copying is the naive baseline that fails, the two middle papers are controlled copies with different gate designs, and AntiSD is the endpoint where “copy” is replaced by “mine the disagreement.” The DelTA vs AntiSD vs DVAO comparison covers AntiSD’s reward-shaping layer; this guide is about the privileged-teacher axis it shares with SDPG and SDAR.

SDPG: copy the teacher, but only on winning trajectories

SDPG (Self-Distilled Policy Gradient) is built for math RLVR where you already hold verified references for every training prompt. The privileged context is blunt: The correct answer is {answer}. A common way to solve this is: {solution}, where the solution is generated by Gemini 2.5 Pro. The same Qwen3-4B weights serve as both student (question only) and teacher (question plus hint), and the objective stacks three terms: a group-relative verifier advantage with a stabilized denominator A = (R - mean) / (std + eps_std), a full-vocabulary on-policy distillation loss, and a small reference-policy KL penalty (alpha = 1e-3).

Two controls are what make the copy survivable. Positive-advantage gating applies the distillation loss only on trajectories the verifier scored correct, so a sharp-but-imperfect teacher never reinforces plausible-looking tokens that sit on a globally wrong path. A beta schedule warms the distillation weight up, holds it, then decays it near the end of training, so the finished policy stops leaning on information it cannot access at test time.

The numbers, on Qwen3-4B at 400 steps (pass@1, mean@32): AIME2025 last-checkpoint 0.327 against 0.242 for GRPO-style training and 0.300 for the RLSD self-distillation baseline; AIME2024 0.380 against 0.280; AMC23 0.858 against 0.714. The telling internal comparison is against RLSD, which distills from the same privileged teacher without a brake and suffers entropy collapse around step 250. SDPG’s reference KL keeps actor entropy high through training; the gain is not “add self-distillation” but “add self-distillation and then keep it from eating the policy.” One honest caveat: the paper reports no wall-clock or token cost for the second teacher forward pass, and the AIME best-vs-last gap (0.408 vs 0.380) hints the runs are still noisy at this scale.

SDAR: the same trick in multi-turn agents, where the teacher is often wrong

SDAR (Self-Distilled Agentic Reinforcement Learning) moves the privileged-teacher idea from single-turn math to multi-turn agents (household tasks in ALFWorld, multi-hop Search-QA, web shopping in WebShop), trained with plain GRPO plus an auxiliary distillation loss L = L_GRPO + lambda * L_SDAR. The teacher is the same policy given retrieved skills it will not see at inference, which makes the target cheap and on-policy.

The reason SDAR exists is a measurement that should give anyone doing privileged-teacher RL pause: because skill retrieval is imperfect, more than 50% of tokens show a negative teacher-student log-prob gap, meaning the teacher is literally worse than the student on the majority of positions. An ungated copy would actively hurt on those tokens, and the paper’s evidence is that it does. So a sigmoid gate over the gap amplifies distillation where the teacher is genuinely better and softly attenuates it elsewhere. Even with random retrieval (pure noise as skills), the gate salvages a small gain (+1.9 on ALFWorld), which is the cleanest evidence that the gating, not the skills, does the work.

The results, SDAR over its own GRPO baseline: on Qwen2.5-3B-Instruct, ALFWorld 75.0% to 84.4% (+9.4), Search-QA 36.4% to 43.4% (+7.0), WebShop accuracy 63.3% to 68.0%; on Qwen2.5-7B, WebShop 72.6% to 82.8% (+10.2); the single largest jump is +20.3 on WebShop for Qwen3-1.7B, though that comes off a weak 38.3% baseline. Gains hold across retrieval tiers, which is the right way to read the method: it is a trust-management layer, not a retrieval improvement. The open question it leaves is the one SDPG never has to answer: whether a same-sized gated teacher beats simply distilling from a genuinely bigger model when one is available.

AntiSD: stop copying, start rewarding disagreement

AntiSD (Anti-Self-Distillation) begins from a diagnostic about why naive privileged-teacher copying stalls: standard self-distillation sharpens exactly the wrong tokens. It inflates confidence on structural boilerplate (line breaks, “Step”, equals signs) while reducing confidence on deliberation tokens where the model is actually weighing options. Minimizing teacher-student divergence spends gradient on tokens that were never the bottleneck.

So AntiSD inverts the sign. It computes the conditional pointwise mutual information between the next token and the privileged context and rewards the divergence, shaped with a smooth phi function bounded on the negative side so exploratory tokens cannot be punished too hard. An entropy gate calibrated over a 5-step warmup deactivates the reward once teacher entropy drops below 0.93 times its baseline (a confident teacher has no disagreement left to mine), which lets the signal be aggressive early and harmless late.

The results are sample-efficiency claims: AntiSD reaches GRPO’s accuracy in 2 to 10 times fewer training steps across five models from 4B to 30B, and finishes up to 11.5 points higher. The strongest single data point is Qwen3-8B on HMMT 2025: GRPO’s peak in roughly one fifth of the steps, ending about 15 points higher. Like the other two, it is validated where privileged contexts are easy to construct (competition math), and the constants (0.93, the warmup length) read as tuned.

Key numbers

MeasurementSDPGSDARAntiSDSetting and sourceSame harness?
Headline vs GRPOAIME25 0.327 vs 0.242 last-ckpt (+0.085); AMC23 0.858 vs 0.714 (+0.144)WebShop +10.2 (7B: 72.6 to 82.8); ALFWorld +9.4 (3B: 75.0 to 84.4)GRPO accuracy in 2 to 10x fewer steps; up to +11.5 final; HMMT25 about 1/5 steps, about +15 (Qwen3-8B)Each paper’s own tablesNo
Privileged contextGround-truth answer + Gemini 2.5 Pro solutionRetrieved skills from past experienceQuestion + partial solution (hint)Each paperYes (all “extra context at train, none at inference”)
Signal directionMinimize full-vocab forward KL to teacherGated minimize (sigmoid over teacher-student gap)Maximize PMI divergence as rewardEach paperYes
Gate designTrajectory-level: positive advantage only, plus beta annealToken-level: sigmoid gate; >50% tokens have negative gapEntropy gate at 0.93x warmup baseline, plus bounded phi shapingEach paperNo
Models testedQwen3-4B (+1.7B appendix)Qwen2.5-3B/7B, Qwen3-1.7BFive models 4B to 30B incl. Qwen3-8B, Qwen3-30B-A3B, OLMo-3-7B x2Each paperNo
DomainMath contests (AIME24/25, AMC23)Agentic: ALFWorld, Search-QA, WebShopMath (AIME, HMMT, MinervaMath), code check (HumanEval+/MBPP+)Each paperNo
Extra cost per stepSecond forward pass (unreported wall-clock)Retrieval + second forward passPrivileged-context forward passesEach paperYes
Collapse protectionReference KL (alpha = 1e-3); RLSD baseline collapses at step 250 without itGate attenuates the majority of tokensEntropy gate disables signal once teacher is confidentEach paperNo

Read the table for its structure, not as a ranking: no row except the design rows was measured under a shared harness, and the headline row bundles three different benchmarks, three model families, and three evaluation budgets. What is directly comparable is the design space: all three pay a second forward pass, all three must remove or gate the privileged signal, and each gates at a different granularity: trajectory (SDPG), token (SDAR), or time (AntiSD’s entropy schedule).

Limits and open questions

Three caveats apply to the whole family, on top of the per-paper ones above. First, the evidence is author-measured and narrow: every headline number comes from the paper’s own setup, no two papers share a harness, and none of the three has a third-party head-to-head replication. Second, the privileged context is assumed rather than explained: SDPG needs verified references per prompt, SDAR needs a skill store worth retrieving, AntiSD needs a hint that is genuinely informative; if you cannot construct one, the method does not degrade gracefully, it does not apply. Third, the training-time bill is understated in all three papers: everyone pays a second forward pass per step, only SDAR additionally pays retrieval, and only SDPG’s paper omits the wall-clock cost entirely. Treat the sample-efficiency and accuracy gains as real but bounded, and re-run the comparison on your own task distribution before committing.

When to use which

Start from what your privileged context is. If you hold verified references (answers, solutions, unit tests), SDPG is the most direct fit, and its controls (trajectory-level gating, reference KL, beta anneal) are the minimum viable kit for any privileged-teacher setup. Its AIME numbers are also the most conservative of the three: every gain is quoted against a same-budget GRPO run on the same model.

If your task is multi-turn, copying unconditionally is off the table. SDAR’s >50%-negative-gap measurement is the transferable result here: in any setting where the extra context is retrieved rather than verified, expect the teacher to be worse than the student on a large share of tokens, and gate accordingly. The random-retrieval ablation says the gate salvages value even from noise.

If rollout cost dominates and your domain admits a clean hint, AntiSD is the strongest claim: 2 to 10x fewer steps is a compute statement, not a leaderboard nudge, but it is also the most tuned: the 0.93 entropy threshold and 5-step warmup deserve a sensitivity check before you trust the transfer.

The common failure mode to budget for is collapse toward the teacher: SDPG watches entropy, SDAR watches per-token trust, AntiSD watches teacher confidence. Pick the guardrail that matches where your privileged signal is most likely to lie to you. And if none of your tasks admit a cheap privileged context, none of these methods apply; that is the actual entry fee, not the second forward pass. The broader lineage of dense-signal RLVR fixes is covered in DelTA vs AntiSD vs DVAO, and the distillation side of the story in on-policy distillation explained.

FAQ

Do SDPG, SDAR, and AntiSD need a bigger teacher model?

No; that is the point of the family. In all three, the teacher is the same weights conditioned on extra context: an answer-plus-solution string (SDPG), retrieved skills (SDAR), or a partial solution (AntiSD). The distillation target costs a second forward pass but zero extra parameters. The open question SDAR leaves is whether this same-sized teacher beats a genuinely bigger one when available.

Which method fixes the entropy collapse problem in self-distillation?

SDPG addresses it most directly. Its RLSD baseline (privileged-teacher distillation without a brake) collapses around step 250, while SDPG’s reference KL penalty (alpha = 1e-3) and decaying distillation weight keep actor entropy high. AntiSD sidesteps the issue structurally: it never copies the teacher, so there is no teacher to collapse onto, and its entropy gate disables the PMI reward once the teacher becomes confident.

Which of SDPG, SDAR, and AntiSD works without verified training answers?

Mostly none of them. SDPG explicitly needs ground-truth answers plus reference solutions per prompt. AntiSD needs a constructible partial solution, which is easy for math and hard elsewhere. SDAR is the least demanding (the privileged context is retrieved skills), but its gains assume retrieval exists and its best results still come from curated skill stores. All three are fine-tuning methods for settings where you already hold extra supervision, not ways to conjure supervision from a bare verifier reward.

How large are the gains, honestly?

SDPG: +0.085 to +0.144 last-checkpoint over GRPO on three math contests at Qwen3-4B, with the internal RLSD baseline between them. SDAR: +4.7 to +10.2 over GRPO on agentic benchmarks, with the largest single number (+20.3 WebShop) coming off the weakest baseline. AntiSD: 2 to 10x sample efficiency and up to +11.5 final points, with the +15 HMMT25 figure on one model. Each is a real but bounded gain, measured by its own authors on its own setup; none has been replicated head-to-head by a third party.

What do SDPG, SDAR, and AntiSD cost at inference time?

Nothing at deployment, because the privileged context is training-only in every case. What they cost is training-time: one extra forward pass per step to evaluate the teacher (SDPG, AntiSD) plus retrieval (SDAR). SDPG is the only paper that does not report the wall-clock price of that second pass, so benchmark it on your own setup before scaling.