The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Works
Dense per-step supervision is an appealing remedy for sparse-reward, long-horizon LLM agents: reward the agent for predicting its next observation, and memory should follow. We show that under group-normalized RL (GRPO), this recipe does not merely fail -- it destroys the policy. Across Qwen3-1.7B/4B/8B on ALFWorld, a potential-based prediction reward drives every run into a degenerate absorbing state (prediction accuracy -> 1.0, task success -> 0,episode length pinned at the horizon): the "dark room" pathology, built automatically by the optimizer. A single-factor ablation localizes the cause
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%RLHF →
- PossiblePossibly related (embedding) · 46%Reinforcement Learning With Metacognitive Feedback Is Offered As A Next-Gen Way To Shape AI LLMs - Forbes →
- FuzzySimilar title/name (fuzzy) · 87%NirDiamant/GenAI_Agents →
“Fuzzy title match (0.94): “The Dark Room in the Reward Channel: Dense Prediction Reward” ≈ “NirDiamant/GenAI_Agents””
- FuzzySimilar title/name (fuzzy) · 84%Unity-Technologies/ml-agents →
“Fuzzy title match (0.92): “The Dark Room in the Reward Channel: Dense Prediction Reward” ≈ “Unity-Technologies/ml-agents””
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzySimilar title/name (fuzzy) · 59%datawhalechina/hello-agents →
“Fuzzy title match (0.73): “The Dark Room in the Reward Channel: Dense Prediction Reward” ≈ “datawhalechina/hello-agents””
- LinkedLinked via arxiv author · 85%Xuanyu Wang →
“The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Work”
