PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional {underrepresented languages spanning 6 language families -- ranging
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%thu-pacman/chitu →
- PossiblePossibly related (embedding) · 55%chrisliu298/awesome-llm-unlearning →
- PossiblePossibly related (embedding) · 54%sileod/reasoning-core →
- PossiblePossibly related (embedding) · 53%edwardcapriolo/deliverance →
- PossiblePossibly related (embedding) · 53%New benchmark exposes reasoning gaps in top models →
- PossiblePossibly related (embedding) · 51%1.7B model leading strict-7 formal reasoning above Qwen3-8B and Gemma-4-26B - specialists eating generalist territory? →
- LinkedLinked via arxiv author · 85%Daryna Dementieva →
“PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages”
- LinkedLinked via arxiv author · 85%Nikolay Babakov →
“PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages”
