Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya
Multilingual pre-trained language models (PLMs) exhibit degraded performance on low-resource, non-Latin-script languages, driven by high out-of-vocabulary (OOV) rates and excessive subword fragmentation that result from Latin-script-centric tokenizer training. We introduce VEXMLM, a vocabulary-extended variant of XLM-R targeting the two highest-resource Ge'ez-script languages, Amharic and Tigrinya, and further evaluated on 17 additional low-resource African languages (19 total). We train a language-specific SentencePiece tokenizer on curated Amharic and Tigrinya monolingual corpora, extend XLM
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Hailay Kidu Teklehaymanot →
“Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya”
- LinkedLinked via arxiv author · 85%Debela Desalegn Yadeta →
“Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya”
- LinkedLinked via arxiv author · 85%Wolfgang Nejdl →
“Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya”
