One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation
On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost. The teacher is the model itself, conditioned on privileged information the student will not have at test time, such as a reference solution, a plan, or environment feedback. The teacher is no stronger than the student, only better informed. Early results were pro
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%tirth8205/code-review-graph →
“Fuzzy title match (0.73): “One Symptom, Three Levers: A Critical Review of On-Policy Se” ≈ “tirth8205/code-review-graph””
- LinkedLinked via arxiv author · 85%Justin Robert →
“One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation”
- LinkedLinked via arxiv author · 85%Raheel Qader →
“One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation”
