Localize-Then-Decide Guarantees for LLM Judgments
Large language models (LLMs) are increasingly used as evaluators to assess output quality and preference alignment, yet providing reliable guarantees of agreement with human judgments remains challenging. Recent work introduces confidence-thresholding methods that provide such guarantees for pairwise comparisons, relying on the assumption that higher estimated confidence implies lower disagreement risk with humans. However, this assumption can break down when the number of candidate responses increases, since distributing probability mass across many alternatives can distort confidence estimat
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%keras-team/keras →
“Shared author/contributor keys: jin”
- FuzzyOverlapping authors or contributors · 62%HKUDS/LightRAG →
“Shared author/contributor keys: jin”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: zhou”
- LinkedLinked via arxiv author · 85%Xinyu Li →
“Localize-Then-Decide Guarantees for LLM Judgments”
- LinkedLinked via arxiv author · 85%Jiayi Zhou →
“Localize-Then-Decide Guarantees for LLM Judgments”
- LinkedLinked via arxiv author · 85%Guanqun Cao →
“Localize-Then-Decide Guarantees for LLM Judgments”
- LinkedLinked via arxiv author · 85%Zeyu Fu →
“Localize-Then-Decide Guarantees for LLM Judgments”
- LinkedLinked via arxiv author · 85%Tianjin Huang →
“Localize-Then-Decide Guarantees for LLM Judgments”
