Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content
Language models encode text as subword tokens, raw bytes, or rendered pixels, but these encodings are usually compared under modeling constraints that expose different amounts of linguistic content to models across different languages. We instead ask what each encoding preserves when both the content and the downstream capacity are controlled. Using verified parallel sentences across thirteen languages and five scripts, we compare tokens, bytes, and pixels through a shared bottleneck whose width is swept to trace rate-utility frontiers. This separates three quantities that are often conflated:
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%Large Language Models Are Still Getting Stronger, but Researchers Face New Bottlenecks in Data, Evaluation, and Safety | Newswise - Newswise →
- PossiblePossibly related (embedding) · 53%New Server Hopes to Break Through AI’s “Memory Wall” →
- LinkedLinked via arxiv author · 85%Ingo Ziegler →
“Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content”
- LinkedLinked via arxiv author · 85%Martin Krebs →
“Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content”
- LinkedLinked via arxiv author · 85%Desmond Elliott →
“Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content”
- PossiblePossibly related (embedding) · 54%GigaToken: ~1000x faster Language model tokenization →
- PossiblePossibly related (embedding) · 46%Can Large Language Models Capture Human Risk Preferences? - Mirage News →
