repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 28d ago
nick7nlp/Awesome-LLM-On-Policy-Distillation
A curated collection of papers and resources on On-Policy Distillation for Large Language Models.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation →
- PossiblePossibly related (embedding) · 56%MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training →
- PossiblePossibly related (embedding) · 54%PHF: Privileged Hidden Flow for On-Policy Self-Distillation →
- PossiblePossibly related (embedding) · 53%Purified OPSD: On-Policy Self-Distillation Without Losing How to Think →
- PossiblePossibly related (embedding) · 50%Knowledge Distillation of Black-Box Large Language Models →
- PossiblePossibly related (embedding) · 51%Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation →
- PossiblePossibly related (embedding) · 62%DemoPSD: Disagreement-Modulated Policy Self-Distillation →
- PossiblePossibly related (embedding) · 47%TREK: Distill to Explore, Reinforce to Refine →
Implements
paperDistill to Detect: Exposing Stealth Biases in LLMs through Cartridge DistillationpaperMOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-TrainingpaperPHF: Privileged Hidden Flow for On-Policy Self-DistillationpaperPurified OPSD: On-Policy Self-Distillation Without Losing How to Think
Covers
Implements (incoming)
Covers (incoming)
Related across the graph
paperMOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-TrainingnewsKnowledge Distillation of Black-Box Large Language ModelspaperDemoPSD: Disagreement-Modulated Policy Self-DistillationpaperPurified OPSD: On-Policy Self-Distillation Without Losing How to ThinkpaperWeak-to-Strong Generalization via Direct On-Policy DistillationpaperPHF: Privileged Hidden Flow for On-Policy Self-DistillationnewsModel "distillation" accusations are getting way overblown at this pointpaperDistill to Detect: Exposing Stealth Biases in LLMs through Cartridge DistillationpaperNeuron-Aware Data Selection for Annotation-Free LLM Self-DistillationpaperTREK: Distill to Explore, Reinforce to RefinenewsThe LLM distillation process simplified for politicians:
