Neural Collapse Is Forbidden: Information Floors in Language Models
Within-class variance in language-model representations is commonly read as incomplete neural collapse. We argue it is allocated information storage, and that the allocation obeys a law. A one-line centering identity voids a family of simplex equiangular-tight-frame claims, including our own earlier ones; in dimensionless variance shares across 14 models, macro-category structure carries only 4-12% of representational variance and within-token context carries 79-91%, stable across a 100x parameter range. On the theory side, token-level weight decay penalizes a category in proportion to its typ
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%minimal-diffusion-lm →
- PossiblePossibly related (embedding) · 49%Transformer →
- PossiblePossibly related (embedding) · 48%A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models - Apple Machine Learning Research →
- PossiblePossibly related (embedding) · 47%New Server Hopes to Break Through AI’s “Memory Wall” →
- LinkedLinked via arxiv author · 85%Bruno Abrahao →
“Neural Collapse Is Forbidden: Information Floors in Language Models”
