CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention
Recurrent models must forget in order to remember, yet the state of the art decides what to erase without consulting what is stored -- the gate sees only the arriving token, not the memory it is about to modify. This memory-blind gating is one of three coupled defects in the leading delta-rule architecture (GDN-2): the value-axis erase mask wastes parameters at the scale of the value projection, and -- as we prove -- mathematically prevents the WY-form triangular chunk solver that makes recurrent training competitive with Transformers. We introduce CARVE (Content-Aware Recurrent with Value E
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownBreakthrough in long-context efficiency announced →
- LinkedLinked via unknownNew Server Hopes to Break Through AI’s “Memory Wall” →
- LinkedLinked via unknownMatrix Orthogonalization Improves Memory in Recurrent Models →
- PossiblePossibly related (embedding) · 48%Hierarchos: Preliminary Findings From a 232M Recurrent Memory-Augmented Assistant Model [P] →
- PossiblePossibly related (embedding) · 45%Zefan-Cai/R-KV →
- PossiblePossibly related (embedding) · 53%A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets →
