How Much is Left? LLMs Linearly Encode Their Remaining Output Length
Large language models generate one token at a time, yet their responses show remarkably consistent length structure: step-by-step solutions converge in predictable token counts, retrievals stop after a few sentences, retractions extend responses by measurable amounts. We ask whether the model carries an internal estimate of how much response remains. Training minimal-capacity linear probes on frozen hidden states of three open-weight 7-8B models across seven completion-style datasets, we find three converging pieces of evidence. First, total response length is linearly decodable from the promp
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 58%chrisliu298/awesome-llm-unlearning →
- PossiblePossibly related (embedding) · 57%New Server Hopes to Break Through AI’s “Memory Wall” →
- PossiblePossibly related (embedding) · 56%thu-pacman/chitu →
- PossiblePossibly related (embedding) · 52%Looking for feedback on a small test SLM I built completely from scratch [P] →
- PossiblePossibly related (embedding) · 48%Atomic-man007/Awesome_Multimodel_LLM →
- LinkedLinked via arxiv author · 85%Mohamed Amine Merzouk →
“How Much is Left? LLMs Linearly Encode Their Remaining Output Length”
- LinkedLinked via arxiv author · 85%Dmitri Carpov →
“How Much is Left? LLMs Linearly Encode Their Remaining Output Length”
- LinkedLinked via arxiv author · 85%Mirko Bronzi →
“How Much is Left? LLMs Linearly Encode Their Remaining Output Length”
