ECHO: Prune to act, trace to learn with selective turn memory in agentic RL
Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Existing context-management methods make such rollouts feasible by truncating distant history, folding past turns into summaries, or selecting compact memory states. However, these breakthroughs introduce two coupled limitations. First, as the number of turns grows, historical observations are progressively removed or collapsed into compressed states, making it harder for the policy to reuse fine-grained evidence. Second, once the original turns are no longer
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownEvaluating long-term memory limits in stateless LLM chatbots — feedback needed [D] →
- LinkedLinked via unknownNoshkoto/Noshy →
- PossiblePossibly related (embedding) · 48%Why I built a proactive context curator instead of a compactor — and what I got wrong for three months [P] →
- PossiblePossibly related (embedding) · 49%zircote/rlm-rs →
