Language Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio Geometry
Language identification is commonly addressed using either neural architectures or statistical n-gram models. Neural approaches typically require substantial computational resources, whereas classical frequency-based methods offer efficient linear-time performance, but rely on distance metrics that are not always appropriate for compositional data. This work models character and bigram frequency distributions as compositional vectors constrained to the simplex and mapped via the centered log-ratio (CLR) transformation bijectively onto the $(D-1)$-dimensional zero-sum subspace of $\mathbb{R}^
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%Transformer →
- LinkedLinked via arxiv author · 85%Paul-Andrei Pogăcean →
“Language Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio Geometry”
- LinkedLinked via arxiv author · 85%Sanda-Maria Avram →
“Language Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio Geometry”
