Read original ↗
paperarXivTrust 82 · PrimaryPublished 7d agoLive · 5d ago

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost. The teacher is the model itself, conditioned on privileged information the student will not have at test time, such as a reference solution, a plan, or environment feedback. The teacher is no stronger than the student, only better informed. Early results were pro

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%tirth8205/code-review-graph

    Fuzzy title match (0.73): “One Symptom, Three Levers: A Critical Review of On-Policy Se” ≈ “tirth8205/code-review-graph”

  • LinkedLinked via arxiv author · 85%Justin Robert

    One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

  • LinkedLinked via arxiv author · 85%Raheel Qader

    One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

Implements (incoming)

authored (incoming)

Related across the graph

Topics