MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use
Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memories reshape model reasoning and affect performance on the current task. We identify memory-induced cognitive traps: even faithfully recorded and semantically relevant memories can distort model reasoning or beliefs and degrade current task performance. To systematically evaluate these failure modes,
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 30%MemTensor/MemOS →
“Possibly related via embedding similarity 0.67 (not asserted). Timestamp check: artifact slightly before paper (-49d).”
- PossiblePossibly related (embedding) · 61%I tried to give an LLM room to think. This is where it led. [P] →
- PossiblePossibly related (embedding) · 61%The biggest problem with AI memory isn't recall—it's stale facts [P] →
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Mengru Wang →
“MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use”
- LinkedLinked via arxiv author · 85%Haozhe Luo →
“MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use”
- LinkedLinked via arxiv author · 85%Zhenqian Xu →
“MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use”
- LinkedLinked via arxiv author · 85%Zhixiang Cui →
“MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use”
