Read original ↗
paperarXivTrust 82 · PrimaryPublished 2d agoLive · 16h ago

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). However, the teacher scores student-generated trajectories that are inherently off-policy for it, so the reliability of its supervision, and hence the source of the student's improvement, remains unclear. We quantitatively analyze teacher supervision during OPD training and find substantial noise whose prevalence increases with teacher scale. Surprisingly, the student policy is insensitive to such noise, converging

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Yi Ding

    Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

  • LinkedLinked via arxiv author · 85%Ruqi Zhang

    Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

authored (incoming)

Related across the graph

Topics