One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography
Multilingual language models transfer knowledge across languages through shared subword vocabulary, a mechanism that breaks down when related languages use different writing systems. Prior work addresses this via script equalization (romanization or IPA transcription), but direct comparisons are rare; the focus has been on encoder-only models, with most work adapting existing pretrained models. We systematically compare different input representations in autoregressive multilingual pretraining, comparing orthographic text, IPA, and romanization in a controlled setup across three scales (467M,
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions - Apple Machine Learning Research →
- LinkedLinked via arxiv author · 85%Muge Zhang →
“One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography”
- LinkedLinked via arxiv author · 85%Aaron Jencks →
“One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography”
- LinkedLinked via arxiv author · 85%Krishna Badikela →
“One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography”
- LinkedLinked via arxiv author · 85%Yulia Tsvetkov →
“One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography”
- LinkedLinked via arxiv author · 85%Sachin Kumar →
“One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography”
- PossiblePossibly related (embedding) · 57%Study: Generative AI is making writing on Reddit and elsewhere boring →
