In-Place Tokenizer Expansion for Pre-trained LLMs
A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflecting the deployment priorities at that time. When those priorities shift, languages added later are split into many more tokens per word, which can raise latency, compute, and energy consumption for users of those languages. Cloud models can afford a broad vocabulary because the embedding and LM-head matrices are a small fraction of their parameters. On a compact model those matrices are a material share of per-token decode bandwidth, so on-device models ship small vocabularies a
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 58%New Server Hopes to Break Through AI’s “Memory Wall” →
- FuzzyOverlapping authors or contributors · 62%deepspeedai/DeepSpeed →
“Shared author/contributor keys: smith”
- FuzzyOverlapping authors or contributors · 62%browser-use/browser-use →
“Shared author/contributor keys: lee”
- LinkedLinked via arxiv author · 85%Jimmy T. H. Smith →
“In-Place Tokenizer Expansion for Pre-trained LLMs”
- LinkedLinked via arxiv author · 85%Tarek Dakhran →
“In-Place Tokenizer Expansion for Pre-trained LLMs”
- LinkedLinked via arxiv author · 85%Alberto Cabrera →
“In-Place Tokenizer Expansion for Pre-trained LLMs”
- LinkedLinked via arxiv author · 85%Simon S. Lee →
“In-Place Tokenizer Expansion for Pre-trained LLMs”
- LinkedLinked via arxiv author · 85%Paul Pak →
“In-Place Tokenizer Expansion for Pre-trained LLMs”
