Read original ↗
paperarXivTrust 82 · PrimaryPublished 27d agoLive · 26d ago

Distilled Reinforcement Learning for LLM Post-training

Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and limited capability to acquire new knowledge. OPD, meanwhile, unconditionally matches teacher logits through KL divergence, which creates a dilemma: similar teachers provide little new knowledge, while substantially different teachers often yield ineffective guidance, largely restric

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning

    Fuzzy title match (0.73): “Distilled Reinforcement Learning for LLM Post-training” ≈ “aymericdamien/TopDeepLearning”

  • LinkedLinked via arxiv author · 85%Yuchen Wang

    Distilled Reinforcement Learning for LLM Post-training

  • LinkedLinked via arxiv author · 85%Zhaochun Li

    Distilled Reinforcement Learning for LLM Post-training

  • LinkedLinked via arxiv author · 85%Jionghao Bai

    Distilled Reinforcement Learning for LLM Post-training

  • LinkedLinked via arxiv author · 85%Yining Zhang

    Distilled Reinforcement Learning for LLM Post-training

  • LinkedLinked via arxiv author · 85%Hexuan Deng

    Distilled Reinforcement Learning for LLM Post-training

Implements (incoming)

authored (incoming)

Related across the graph

Topics