KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding
Evaluating large language models (LLMs) across languages remains challenging, as most multilingual benchmarks rely on translated English datasets, often obscuring linguistic and cultural specificity in the target language. This issue is particularly pronounced for less-resourced languages such as Kyrgyz, where reliable natively authored evaluation data are scarce. Building on previously introduced Kyrgyz-language evaluation datasets, this work reports the first systematic and large-scale evaluation of LLMs in Kyrgyz using the KyrgyzLLM-Bench benchmark suite. KyrgyzLLM-Bench comprises two nativ
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%Large Language Models Are Still Getting Stronger, but Researchers Face New Bottlenecks in Data, Evaluation, and Safety | Newswise - Newswise →
- LinkedLinked via arxiv author · 85%Timur Turatali →
“KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding”
- LinkedLinked via arxiv author · 85%Aida Turdubaeva →
“KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding”
- LinkedLinked via arxiv author · 85%Rustem Izmailov →
“KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding”
- LinkedLinked via arxiv author · 85%Anton M. Alekseev →
“KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding”
- LinkedLinked via arxiv author · 85%Sergey I. Nikolenko →
“KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding”
