StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs
Humans effortlessly perceive the present while remembering the past, yet streaming VLMs often trade off real-time perception against long-term memory. Prior work shows that shortening the context can sharpen current-scene perception at the expense of long-range recall. To reconcile these abilities, we introduce StreamTTT, which writes long-range history into online-updated fast weights outside the attention context. This leaves a short sliding key-value cache dedicated to recent evidence, mitigating attention dilution. We train StreamTTT jointly on offline long-video QA and a newly constructed
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%Breakthrough in long-context efficiency announced →
- PossiblePossibly related (embedding) · 53%We're building agents that can read millions of documents, but still forget a video they watched yesterday. →
- LinkedLinked via arxiv author · 85%Joya Chen →
“StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs”
- LinkedLinked via arxiv author · 85%Zeyun Zhong →
“StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs”
- LinkedLinked via arxiv author · 85%Mike Zheng Shou →
“StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs”
