Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents
Reinforcement learning has become a promising paradigm for improving large language model (LLM) agents on long-horizon search tasks, where the agent must make a sequence of intermediate decisions before receiving a final outcome. However, existing methods still face a key limitation: the rollout budget is often allocated without explicitly assessing the utility of intermediate states. As a result, substantial computation may be spent on low-value states, even though different branches can vary drastically in their informativeness. In this paper, we propose Information Gain-based Rollout Policy
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%rllm-org/rllm →
- PossiblePossibly related (embedding) · 47%Best practices for multi-turn reinforcement learning in Amazon SageMaker AI →
- FuzzySimilar title/name (fuzzy) · 87%NirDiamant/GenAI_Agents →
“Fuzzy title match (0.94): “Information Gain-based Rollout Policy Optimization: An Adapt” ≈ “NirDiamant/GenAI_Agents””
- FuzzySimilar title/name (fuzzy) · 84%Unity-Technologies/ml-agents →
“Fuzzy title match (0.92): “Information Gain-based Rollout Policy Optimization: An Adapt” ≈ “Unity-Technologies/ml-agents””
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzySimilar title/name (fuzzy) · 59%datawhalechina/hello-agents →
“Fuzzy title match (0.73): “Information Gain-based Rollout Policy Optimization: An Adapt” ≈ “datawhalechina/hello-agents””
- FuzzySimilar title/name (fuzzy) · 59%Eigenwise/atomic-agents →
“Fuzzy title match (0.73): “Information Gain-based Rollout Policy Optimization: An Adapt” ≈ “Eigenwise/atomic-agents””
