Weak-to-Strong Generalization via Direct On-Policy Distillation
Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to repeat on every new strong model because the target model must generate many rollouts during training. As models scale, post-training itself becomes a bottleneck. We study a weak-to-strong alternative: run RL on a smaller model where rollouts are cheaper, then reuse what that RL run learned to improve a stronger target model. Directly distilling the post-RL weak teacher is not enough, because the teacher's final policy mixes useful RL gains with the limitati
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%nick7nlp/Awesome-LLM-On-Policy-Distillation →
- PossiblePossibly related (embedding) · 55%chrisliu298/awesome-on-policy-distillation →
- PossiblePossibly related (embedding) · 53%rllm-org/rllm →
- PossiblePossibly related (embedding) · 49%AgileRL/AgileRL →
- PossiblePossibly related (embedding) · 48%chrisliu298/awesome-llm-unlearning →
- LinkedLinked via arxiv author · 85%Shiyuan Feng →
“Weak-to-Strong Generalization via Direct On-Policy Distillation”
- LinkedLinked via arxiv author · 85%Huan-ang Gao →
“Weak-to-Strong Generalization via Direct On-Policy Distillation”
- LinkedLinked via arxiv author · 85%Haohan Chi →
“Weak-to-Strong Generalization via Direct On-Policy Distillation”
