Read original ↗
paperarXivTrust 82 · PrimaryPublished 26d agoLive · 24d ago

SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

Self-reflection is a powerful mechanism for credit assignment in human learning, converting sparse outcome feedback into actionable guidance. However, its potential for post-training Large Language Models (LLMs) remains underexplored. We propose Self-Reflective Policy Optimization (SRPO), a framework that internalizes this capability. SRPO enables LLMs to analyze their own completed trajectories, synthesize errors into concise "reflection patches," and use reflection-conditioned teacher scores on student on-policy rollouts as dense token-level training signals. This process effectively transfo

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Jialong Liu

    SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

  • LinkedLinked via arxiv author · 85%Yuling Shi

    SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

  • LinkedLinked via arxiv author · 85%Ning Yang

    SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

  • LinkedLinked via arxiv author · 85%Xiaodong Gu

    SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

  • LinkedLinked via arxiv author · 85%Zuchao Li

    SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

  • FuzzySimilar title/name (fuzzy) · 84%Thysrael/Horizon

    Fuzzy title match (0.92): “SRPO: Self-Reflective Policy Optimization for Long-Horizon R” ≈ “Thysrael/Horizon”

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

authored (incoming)

Implements (incoming)

Related across the graph

Topics