Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning
On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher support can favor locally plausible responses that omit evidence distributed across the input or violate global task constraints. Task-specific verifiers, in contrast, evaluate task completion at the response level and may return graded rewards that reflect partial success. We diagnose this mismatch on fixed responses from two representative long-context evidence-aggregation tasks. Across longer input ranges, trajectory
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%TauricResearch/TradingAgents →
“Shared author/contributor keys: xiao”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: zhou”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Lizhu Zhang →
“Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning”
- LinkedLinked via arxiv author · 85%Jixun Wang →
“Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning”
- LinkedLinked via arxiv author · 85%Xiaoang Xu →
“Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning”
- LinkedLinked via arxiv author · 85%Xiaorong Wang →
“Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning”
