A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation
On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens remains poorly understood. We analyze the gradient of the per-token K2 estimator of reverse KL with respect to the student logits. The $\ell_1$ norm of this gradient factorizes into the absolute teacher--student log-probability gap and a student-side softmax factor that grows as the sampled token becomes less likely under the student. In our math-distillation runs, these per-token norms are highly non-uniform: low-stude
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%mudler/LocalAI →
“Shared author/contributor keys: guo”
- FuzzyOverlapping authors or contributors · 62%HKUDS/LightRAG →
“Shared author/contributor keys: jin”
- FuzzyOverlapping authors or contributors · 62%keras-team/keras →
“Shared author/contributor keys: jin”
- LinkedLinked via arxiv author · 85%Bing Shao →
“A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation”
- LinkedLinked via arxiv author · 85%Jiazheng Zhang →
“A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation”
- LinkedLinked via arxiv author · 85%Xiaolong Ma →
“A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation”
- LinkedLinked via arxiv author · 85%Yujiong Shen →
“A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation”
- LinkedLinked via arxiv author · 85%Senjie Jin →
“A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation”
