Read original ↗
paperarXivTrust 82 · PrimaryPublished 8d agoLive · 6d ago

How to Train a Critic Stably and Efficiently

Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. A reliable critic could instead estimate token-level advantages from one response, but standard critic-based training recipes are often unstable. We study this instability and develop \textbf{Best-Practice Critic Optimization (BPCO)}, a recipe that combines DPPO, value predictions bounded to the reward range, Monte Carlo value targets, unnormalized policy advantages, and length-adaptive generalized advantage estimation. Because the critic is

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%sgl-project/sglang

    Shared author/contributor keys: zhou

  • FuzzyOverlapping authors or contributors · 62%browser-use/browser-use

    Shared author/contributor keys: lee

  • LinkedLinked via arxiv author · 85%Penghui Qi

    How to Train a Critic Stably and Efficiently

  • LinkedLinked via arxiv author · 85%Xiangxin Zhou

    How to Train a Critic Stably and Efficiently

  • LinkedLinked via arxiv author · 85%Wee Sun Lee

    How to Train a Critic Stably and Efficiently

Implements (incoming)

authored (incoming)

Related across the graph

Topics