Read original ↗
paperarXivTrust 82 · PrimaryPublished 2d agoLive · 18h ago

Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-level advantage uniformly to all decisions, yielding coarse credit over long-horizon interactions. On-policy self-distillation offers finer supervision by re-evaluating sampled behavior with privileged information (PI) available only during training. However, fine-grained supervision is not necessarily fine-grained credit: PI-induced likelihood changes describe how additional information alters policy preference, but do not directly determine how an executable action should inherit

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 84%roboflow/supervision

    Fuzzy title match (0.92): “Reconciling Process Supervision with Outcome-Based Credit in” ≈ “roboflow/supervision”

  • FuzzyOverlapping authors or contributors · 62%langchain-ai/langchain

    Shared author/contributor keys: gan

  • FuzzySimilar title/name (fuzzy) · 59%WenyuChiou/awesome-agentic-ai-zh

    Fuzzy title match (0.73): “Reconciling Process Supervision with Outcome-Based Credit in” ≈ “WenyuChiou/awesome-agentic-ai-zh”

  • FuzzySimilar title/name (fuzzy) · 59%Fosowl/agenticSeek

    Fuzzy title match (0.73): “Reconciling Process Supervision with Outcome-Based Credit in” ≈ “Fosowl/agenticSeek”

  • LinkedLinked via arxiv author · 85%Jingxiao Yang

    Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

  • LinkedLinked via arxiv author · 85%Wangjie Gan

    Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

  • LinkedLinked via arxiv author · 85%Yingxuan Zhuang

    Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

  • LinkedLinked via arxiv author · 85%Wenqi Zhang

    Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

Implements (incoming)

authored (incoming)

Related across the graph

Topics