On-Policy Delta Distillation
On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy distillation has been studied and applied across various settings, its fundamental design remains underexplored. In this paper, we introduce a new distillation reward, termed the delta signal, instead of directly imitating the teacher's output distribution. The delta signal is defined as the difference between the teacher model and its base model prior to instruction tuning
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%ultralytics/ultralytics →
“Shared author/contributor keys: han”
- FuzzyOverlapping authors or contributors · 62%janhq/jan →
“Shared author/contributor keys: han”
- LinkedLinked via arxiv author · 85%Byeongho Heo →
“On-Policy Delta Distillation”
- LinkedLinked via arxiv author · 85%Jaehui Hwang →
“On-Policy Delta Distillation”
- LinkedLinked via arxiv author · 85%Sangdoo Yun →
“On-Policy Delta Distillation”
- LinkedLinked via arxiv author · 85%Dongyoon Han →
“On-Policy Delta Distillation”
