Read original ↗
paperarXivTrust 82 · PrimaryPublished 4d agoLive · yesterday

Rethinking Expressivity and Efficiency in Test-Time Training

Test-Time Training (TTT) enables long-context processing via continuous weight updates during inference, but current methods struggle to balance the expressivity of per-token update dynamics with the hardware efficiency of chunk-wise approximations. We propose E$^2$-TTT (Expressive and Efficient TTT) to bridge this gap. Under the standard approximation of taking gradients at the chunk-start weights, we derive a closed-form state transition that exactly reproduces the chunk-end fast-weight and momentum states of the per-token recurrence. This enables fully parallelized chunk-level training whil

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 52%Breakthrough in long-context efficiency announced
  • FuzzyOverlapping authors or contributors · 62%ultralytics/yolov5

    Shared author/contributor keys: martin

  • LinkedLinked via arxiv author · 85%Zeyun Zhong

    Rethinking Expressivity and Efficiency in Test-Time Training

  • LinkedLinked via arxiv author · 85%Joya Chen

    Rethinking Expressivity and Efficiency in Test-Time Training

  • LinkedLinked via arxiv author · 85%Manuel Martin

    Rethinking Expressivity and Efficiency in Test-Time Training

  • LinkedLinked via arxiv author · 85%Frederik Diederichs

    Rethinking Expressivity and Efficiency in Test-Time Training

  • LinkedLinked via arxiv author · 85%Juergen Gall

    Rethinking Expressivity and Efficiency in Test-Time Training

  • LinkedLinked via arxiv author · 85%Juergen Beyerer

    Rethinking Expressivity and Efficiency in Test-Time Training

Covers

Implements (incoming)

authored (incoming)

Related across the graph

Topics