Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context
Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems. Existing long-context and memory benchmarks often expose session or topic boundaries, or probe direct personal-memory questions. These settings understate a harder assistant-memory regime: a flat mixed-topic thread where the system must infer which earlier episode makes a later task decision valid. We introduce SCALE-QA, a constraint-grounded task QA benchmark for flat unsegmented threads targeting episode integrity failure. The dataset contains 3,000 audited questions
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%Show HN: ContextVault – Shared memory layer for your AI and your team →
- PossiblePossibly related (embedding) · 54%Evaluating long-term memory limits in stateless LLM chatbots — feedback needed [D] →
- PossiblePossibly related (embedding) · 54%TRACE: open-source hierarchical memory for LLM agents, 82.5% on MemoryAgentBench’s EventQA using gpt-oss-20B [P] →
- PossiblePossibly related (embedding) · 53%We're building agents that can read millions of documents, but still forget a video they watched yesterday. →
- LinkedLinked via arxiv author · 85%Zhexi Feng →
“Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context”
- LinkedLinked via arxiv author · 85%Ruiyi Zhang →
“Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context”
- LinkedLinked via arxiv author · 85%Yongbo Yang →
“Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context”
- LinkedLinked via arxiv author · 85%Pengtao Xie →
“Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context”
