Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs
Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across 22 languages, spanning high-, mid-, and low-resource settings. Using two human-curated Q\&A datasets, we compare open and closed box UE methods (nine in total) across different model sizes and architectures while eliciting long-form reasoning, avoiding LLM-as-a-judge and embedding-based scoring, which can introduce evaluation noise. We report three main actionable findings. First, we fi
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%chrisliu298/awesome-llm-unlearning →
- PossiblePossibly related (embedding) · 53%New benchmark exposes reasoning gaps in top models →
- PossiblePossibly related (embedding) · 51%thu-pacman/chitu →
- PossiblePossibly related (embedding) · 51%amitshekhariitbhu/llm-internals →
- PossiblePossibly related (embedding) · 51%wbopan/flashtrace →
- LinkedLinked via arxiv author · 85%Andrea Alfarano →
“Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs”
- LinkedLinked via arxiv author · 85%Andrea Bacciu →
“Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs”
- LinkedLinked via arxiv author · 85%Saab Mansour →
“Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs”
