Read original ↗
paperarXivTrust 82 · PrimaryPublished 23d agoLive · 20d ago

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one monolingual models and five multilingual models spanning twenty-seven language varieties across Ghana, Sierra Leone, Nigeria, Senegal, Kenya and Zimbabwe. Models are fine-tuned primarily on read speech drawn from religious texts, which offer broad, license-clear and orthographically consistent coverage for languages that otherwise lack transcribed audio. We describe a two-step (and, f

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 87%google-bert/bert-base-uncased

    Fuzzy title match (0.94): “DONDO: Open w2v-BERT Speech-Recognition Base Models for Afri” ≈ “google-bert/bert-base-uncased”

  • FuzzySimilar title/name (fuzzy) · 87%huggingface/speech-to-speech

    Fuzzy title match (0.94): “DONDO: Open w2v-BERT Speech-Recognition Base Models for Afri” ≈ “huggingface/speech-to-speech”

  • LinkedLinked via arxiv author · 85%Paul Azunre

    DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

Has model

Implements (incoming)

authored (incoming)

Related across the graph

Topics