Read original ↗
paperarXivTrust 82 · PrimaryPublished 4d agoLive · 17h ago

Boosting LLM Exploration via Weak-Model Guidance in RLVR

Reinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and degraded pass@$k$ for large $k$. While existing methods mitigate this entropy collapse through algorithmic regularizations, cross-model non-parametric perturbation is also neglected. In this work, we propose a simple yet effective approach to preserve the generative diversity of LLMs during RLVR. Instead of relying solely on internal exploration, we force the target model to generate answers based on partial reasoning t

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • LinkedLinked via arxiv author · 85%Xingyu Shen

    Boosting LLM Exploration via Weak-Model Guidance in RLVR

  • LinkedLinked via arxiv author · 85%Huishuai Zhang

    Boosting LLM Exploration via Weak-Model Guidance in RLVR

  • LinkedLinked via arxiv author · 85%Pengpeng Liu

    Boosting LLM Exploration via Weak-Model Guidance in RLVR

  • LinkedLinked via arxiv author · 85%Yinchun Wang

    Boosting LLM Exploration via Weak-Model Guidance in RLVR

  • LinkedLinked via arxiv author · 85%Dongyan Zhao

    Boosting LLM Exploration via Weak-Model Guidance in RLVR

Implements (incoming)

authored (incoming)

Related across the graph

Topics