MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations
Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, multi-session interactions. Existing benchmarks, however, evaluate such memory almost exclusively through downstream question answering, scoring only the correctness of a final answer. This black-box formulation conflates the heterogeneous causes of memory failure, such as missing the introduction of a relevant fact, binding an operation to the wrong target, or relying on stale values after a correction. As a result, it can credit correct answers despite their reliance on inconsiste
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 68%Evaluating long-term memory limits in stateless LLM chatbots — feedback needed [D] →
- PossiblePossibly related (embedding) · 61%NirDiamant/Agent_Memory_Techniques →
- PossiblePossibly related (embedding) · 59%basicmachines-co/basic-memory →
- PossiblePossibly related (embedding) · 59%las7/memharness →
- PossiblePossibly related (embedding) · 27%MemTensor/MemOS →
“Possibly related via embedding similarity 0.59 (not asserted). Timestamp check: artifact slightly before paper (-12d).”
- LinkedLinked via arxiv author · 85%Xixuan Hao →
“MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations”
- LinkedLinked via arxiv author · 85%Zeyu Zhang →
“MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations”
- LinkedLinked via arxiv author · 85%Zehao Lin →
“MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations”
