Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, where a privileged teacher with access to reference solutions provides token-level supervision on the student's own generated trajectories. However, we find that OPSD consistently fails on long chain-of-thought (long-CoT) reasoning models, yielding at best marginal gains while destabilizing the reflective reasoning capability these models depend on. Through a novel decomposition of the teacher's supervision signal, we identify the root cause: the teacher's supervision is dominated by a reference

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 51%chrisliu298/awesome-on-policy-distillation
  • PossiblePossibly related (embedding) · 50%benjaminzwhite/reasoning-models
  • LinkedLinked via arxiv author · 85%Zhanming Shen

    Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

  • LinkedLinked via arxiv author · 85%Jintao Tong

    Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

  • LinkedLinked via arxiv author · 85%Shaotian Yan

    Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

  • LinkedLinked via arxiv author · 85%Chen Shen

    Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

  • LinkedLinked via arxiv author · 85%Hao Chen

    Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

  • LinkedLinked via arxiv author · 85%Wentao Ye

    Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

Implements

authored (incoming)

Implements (incoming)

Related across the graph

Topics