A Formal Limitation on Learning Human Language From Textual Corpora
Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of text, including the hidden states of contemporary large language models. Modeling language use as a joint distribution over meanings, contexts, and utterances, we derive upper bounds on the probability that a decoder recovers a speaker's intended meaning from a representation of the utterance. The bounds are governed by the uncertainty that form leaves about meaning, which splits into an irreducible part and a part tha
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Understanding large language models demands distinguishing human projection from machine cognition - Nature →
- FuzzyOverlapping authors or contributors · 62%DietrichGebert/ponytail →
“Shared author/contributor keys: cheng”
- LinkedLinked via arxiv author · 85%Emily Cheng →
“A Formal Limitation on Learning Human Language From Textual Corpora”
- LinkedLinked via arxiv author · 85%Ryan Cotterell →
“A Formal Limitation on Learning Human Language From Textual Corpora”
