Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

High-throughput RLHF systems often decouple rollout generation from policy optimization, leading to the use of stale rollouts during learner updates. In this work, we study the effect of such staleness in asynchronous GRPO. We make the behavior policy explicit in the GRPO surrogate objective and distinguish between the surrogate-gradient mapping used by the learner and the true total derivative of a distribution-dependent population objective. Under assumptions of local boundedness, distributional smoothness, and behavior-policy smoothness, we show that stale rollouts introduce a per-step surr

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Jingwei Song

    Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

  • LinkedLinked via arxiv author · 85%Haofeng Xu

    Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

  • LinkedLinked via arxiv author · 85%Jie Xiao

    Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

  • LinkedLinked via arxiv author · 85%Chengke Bao

    Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

  • LinkedLinked via arxiv author · 85%Jingwei Shi

    Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

  • LinkedLinked via arxiv author · 85%Pengbin Feng

    Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

  • LinkedLinked via arxiv author · 85%Weixun Wang

    Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

  • LinkedLinked via arxiv author · 85%Yuhang Han

    Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

authored (incoming)

Related across the graph

Topics