Training Continuous Chain of Thought Models: A Tale of Two Regimes
Continuous Chain-of-Thought methods replace verbose reasoning traces with a short sequence of dense latent representations. Earlier continuous CoT methods indirectly supervise the latent representations such that its final state match that of verbose reasoning traces, requiring autoregressive, slow generation during training. We introduce C-MTP, a simpler, faster direct supervision approach that models each latent as an average of the embeddings in the CoT traces to be compressed. Our approach outperforms a prior direct supervision method that approximates the distribution of compressed tokens
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning - Apple Machine Learning Research →
- PossiblePossibly related (embedding) · 51%Chain of Thought is a scaling trap. the next wave is latent reasoning (Coconut / HRM / RecrusiveMAS)... but then we hit the black box wall. Where does BDH fit? [D] →
- FuzzyOverlapping authors or contributors · 62%ultralytics/yolov5 →
“Shared author/contributor keys: choi”
- LinkedLinked via arxiv author · 85%Varun Yerram →
“Training Continuous Chain of Thought Models: A Tale of Two Regimes”
- LinkedLinked via arxiv author · 85%He He →
“Training Continuous Chain of Thought Models: A Tale of Two Regimes”
- LinkedLinked via arxiv author · 85%Eunsol Choi →
“Training Continuous Chain of Thought Models: A Tale of Two Regimes”
