Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

On-Policy Delta Distillation

On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy distillation has been studied and applied across various settings, its fundamental design remains underexplored. In this paper, we introduce a new distillation reward, termed the delta signal, instead of directly imitating the teacher's output distribution. The delta signal is defined as the difference between the teacher model and its base model prior to instruction tuning

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%ultralytics/ultralytics

    Shared author/contributor keys: han

  • FuzzyOverlapping authors or contributors · 62%janhq/jan

    Shared author/contributor keys: han

  • LinkedLinked via arxiv author · 85%Byeongho Heo

    On-Policy Delta Distillation

  • LinkedLinked via arxiv author · 85%Jaehui Hwang

    On-Policy Delta Distillation

  • LinkedLinked via arxiv author · 85%Sangdoo Yun

    On-Policy Delta Distillation

  • LinkedLinked via arxiv author · 85%Dongyoon Han

    On-Policy Delta Distillation

Implements (incoming)

authored (incoming)

Related across the graph

Topics