Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya

Multilingual pre-trained language models (PLMs) exhibit degraded performance on low-resource, non-Latin-script languages, driven by high out-of-vocabulary (OOV) rates and excessive subword fragmentation that result from Latin-script-centric tokenizer training. We introduce VEXMLM, a vocabulary-extended variant of XLM-R targeting the two highest-resource Ge'ez-script languages, Amharic and Tigrinya, and further evaluated on 17 additional low-resource African languages (19 total). We train a language-specific SentencePiece tokenizer on curated Amharic and Tigrinya monolingual corpora, extend XLM

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Hailay Kidu Teklehaymanot

    Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya

  • LinkedLinked via arxiv author · 85%Debela Desalegn Yadeta

    Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya

  • LinkedLinked via arxiv author · 85%Wolfgang Nejdl

    Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya

authored (incoming)

Related across the graph

Topics