Enhancing Rubric-based RL via Self-Distillation
Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout manages to satisfy (Unexplored Criteria, UC) receive no optimization signal. Recent methods address this by incorporating rubric information as external guidance during rollout, yet they introduce a train-inference mismatch: the policy is optimized on rollouts produced under external guidance while this guidance is absent at inference time, causing error accumulation through autoregressive decoding. Moreover, these
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Verifiers v1 Lets Agentic RL Training Exceed Model Context Windows via DAG Branching - Tech Times →
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Mingxuan Xia →
“Enhancing Rubric-based RL via Self-Distillation”
- LinkedLinked via arxiv author · 85%Yuhang Yang →
“Enhancing Rubric-based RL via Self-Distillation”
- LinkedLinked via arxiv author · 85%Chao Ye →
“Enhancing Rubric-based RL via Self-Distillation”
- LinkedLinked via arxiv author · 85%Shuai Zhu →
“Enhancing Rubric-based RL via Self-Distillation”
- LinkedLinked via arxiv author · 85%Shenzhi Yang →
“Enhancing Rubric-based RL via Self-Distillation”
