TREK: Distill to Explore, Reinforce to Refine
Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning trajectories, but it stalls on hard prompts whose correct solution modes lie outside the student's on-policy support. We propose TREK (Teacher-Routed Exploration via Forward KL), a simple staged procedure that uses distillation not for imitation but for exploration support expansion. A key advantage of TREK is its generality: because it only consumes verified output trajectories, it can use an external black-box teacher, a white-box teacher, or the same model given additional inferen
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%rllm-org/rllm →
- PossiblePossibly related (embedding) · 47%nick7nlp/Awesome-LLM-On-Policy-Distillation →
- PossiblePossibly related (embedding) · 46%AgileRL/AgileRL →
- PossiblePossibly related (embedding) · 45%chrisliu298/awesome-on-policy-distillation →
- PossiblePossibly related (embedding) · 49%ModelOriented/DALEX →
- LinkedLinked via arxiv author · 85%Yuanda Xu →
“TREK: Distill to Explore, Reinforce to Refine”
- LinkedLinked via arxiv author · 85%Zhengze Zhou →
“TREK: Distill to Explore, Reinforce to Refine”
- LinkedLinked via arxiv author · 85%Kayhan Behdin →
“TREK: Distill to Explore, Reinforce to Refine”
