Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a f
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova →
- PossiblePossibly related (embedding) · 50%DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R] →
- PossiblePossibly related (embedding) · 49%New benchmark exposes reasoning gaps in top models →
- PossiblePossibly related (embedding) · 47%Evaluating J-space entropy as an error predictor across 7 datasets on Qwen3-4B [R] →
- LinkedLinked via arxiv author · 85%Ayoub Kirouane →
“Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See”
- LinkedLinked via arxiv author · 85%Christos Petrocheilos →
“Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See”
