Local and Global Regimes of Geometric Complexity in Language Model Representations
Intrinsic dimensionality (ID) is widely used to probe the representational complexity of language models, but it remains unclear whether ID differences reflect properties of language itself or artefacts of how the underlying dataset was constructed. In this paper, we focus specifically on how lexical diversity, the number of unique last-token items present in a dataset, affects ID estimates of that dataset. We find a scale-dependent transition between two regimes: at low lexical diversity, conditions with fewer unique final words produce higher ID, while at high lexical diversity, this orderin
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%Evaluating J-space entropy as an error predictor across 7 datasets on Qwen3-4B [R] →
- PossiblePossibly related (embedding) · 47%Understanding large language models demands distinguishing human projection from machine cognition - Nature →
- FuzzySimilar title/name (fuzzy) · 84%mudler/LocalAI →
“Fuzzy title match (0.92): “Local and Global Regimes of Geometric Complexity in Language” ≈ “mudler/LocalAI””
- LinkedLinked via arxiv author · 85%Arwa Osman →
“Local and Global Regimes of Geometric Complexity in Language Model Representations”
- LinkedLinked via arxiv author · 85%Marco Baroni →
“Local and Global Regimes of Geometric Complexity in Language Model Representations”
- LinkedLinked via arxiv author · 85%Iuri Macocco →
“Local and Global Regimes of Geometric Complexity in Language Model Representations”
