Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung
Vietnam's ethnic minority languages are almost absent from the field of Natural Language Processing (NLP), and the challenge goes beyond data scarcity: Cham, Khmer, and Tay-Nung differ sharply in script, Vietnamese contact, and standardization, conditions under which standard multilingual adaptation can learn the wrong signals. We introduce CKTN, the first corpus and benchmark for these languages (44,367 documents, 24M subword tokens), spanning continued pretraining, category classification, and summary-document retrieval. We show that existing multilingual encoders severely fragment these lan
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%PacificAI/langtest →
- PossiblePossibly related (embedding) · 49%chrisliu298/awesome-llm-unlearning →
- PossiblePossibly related (embedding) · 48%thu-pacman/chitu →
- PossiblePossibly related (embedding) · 48%EuroEval/EuroEval →
- PossiblePossibly related (embedding) · 47%ConlangCrafter Turns AI to Imagining Languages →
- LinkedLinked via arxiv author · 85%Anh Trac Duc Dinh →
“Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung”
- LinkedLinked via arxiv author · 85%Khang Nhat Hoang Vo →
“Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung”
- LinkedLinked via arxiv author · 85%Vinh Cong Doan →
“Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung”
