How Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple Mitigation
Large language models (LLMs) are increasingly critical to digital library workflows, yet their ability to process historical language remains poorly understood. Historical difficulty is typically treated as a monolithic barrier, conflating orthographic variation, linguistic distance, and pretraining exposure. In this paper, we propose a diagnostic framework that decomposes this difficulty into four distinct dimensions: tokenization cost, predictive uncertainty (surprisal), semantic robustness, and context sensitivity. We evaluate this framework on three datasets spanning three centuries: (1)
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownWhat exactly does word2vec learn? →
- LinkedLinked via unknownminimal-diffusion-lm →
- LinkedLinked via unknownTransformer →
- LinkedLinked via unknownBook Review: Domain-Specific Small Language Models by Guglielmo Iozzia →
- LinkedLinked via unknownKnowledge Distillation of Black-Box Large Language Models →
- LinkedLinked via unknownKnowledge Distillation of Black-Box Large Language Models (2024) →
- PossiblePossibly related (embedding) · 47%Apil-Shrestha/token-efficiency-lex →
- PossiblePossibly related (embedding) · 61%EuroEval/EuroEval →
