Read original ↗
paperarXivTrust 82 · PrimaryPublished 27d agoLive · 26d ago

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives \textit{zero} learning signal. Providing privileged guidance during training, such as solution prefixes, can help overcome this learning cliff by steering the model towards {correct solutions with non-zero reward}. {We call these rollouts \textit{off-context}: they are generated from a training prompt that contains privileged guidance, while the target objective is defined by t

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%affaan-m/ECC

    Shared author/contributor keys: jiang

  • FuzzyOverlapping authors or contributors · 62%BerriAI/litellm

    Shared author/contributor keys: jiang

  • FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning

    Fuzzy title match (0.73): “Off-Context GRPO: Learning to Reason on Hard Problems using ” ≈ “aymericdamien/TopDeepLearning”

  • LinkedLinked via arxiv author · 85%Priyank Agrawal

    Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

  • LinkedLinked via arxiv author · 85%Ankur Samanta

    Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

  • LinkedLinked via arxiv author · 85%Shervin Ghasemlou

    Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

  • LinkedLinked via arxiv author · 85%Jalaj Bhandari

    Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

  • LinkedLinked via arxiv author · 85%Kavosh Asadi

    Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

Implements (incoming)

authored (incoming)

Related across the graph

Topics