UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory
Long-term memory is increasingly important for conversational agents, yet existing benchmarks primarily measure memory through pointwise factual recall: whether a system can recover isolated facts or event-level details from prior interactions. Real-world memory use, however, often requires a more demanding capability: integrating distributed, implicit, and noisy evidence across extended interaction histories into coherent, task-oriented outputs. We call this capability memory utilization. Here, we introduce UtilMem, a diagnostic benchmark comprising 1,717 instances across five domains, design
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 63%TRACE: open-source hierarchical memory for LLM agents, 82.5% on MemoryAgentBench’s EventQA using gpt-oss-20B [P] →
- PossiblePossibly related (embedding) · 60%Evaluating long-term memory limits in stateless LLM chatbots — feedback needed [D] →
- PossiblePossibly related (embedding) · 61%Is designing a memory graph around known data structure “overfitting” if I never touch the questions? [D] →
- LinkedLinked via arxiv author · 85%Peijun Qing →
“UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory”
- LinkedLinked via arxiv author · 85%Fobo Shi →
“UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory”
- LinkedLinked via arxiv author · 85%Soroush Vosoughi →
“UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory”
