Read original ↗
paperarXivTrust 82 · PrimaryPublished 5d agoLive · 5d ago

Localize-Then-Decide Guarantees for LLM Judgments

Large language models (LLMs) are increasingly used as evaluators to assess output quality and preference alignment, yet providing reliable guarantees of agreement with human judgments remains challenging. Recent work introduces confidence-thresholding methods that provide such guarantees for pairwise comparisons, relying on the assumption that higher estimated confidence implies lower disagreement risk with humans. However, this assumption can break down when the number of candidate responses increases, since distributing probability mass across many alternatives can distort confidence estimat

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%keras-team/keras

    Shared author/contributor keys: jin

  • FuzzyOverlapping authors or contributors · 62%HKUDS/LightRAG

    Shared author/contributor keys: jin

  • FuzzyOverlapping authors or contributors · 62%sgl-project/sglang

    Shared author/contributor keys: zhou

  • LinkedLinked via arxiv author · 85%Xinyu Li

    Localize-Then-Decide Guarantees for LLM Judgments

  • LinkedLinked via arxiv author · 85%Jiayi Zhou

    Localize-Then-Decide Guarantees for LLM Judgments

  • LinkedLinked via arxiv author · 85%Guanqun Cao

    Localize-Then-Decide Guarantees for LLM Judgments

  • LinkedLinked via arxiv author · 85%Zeyu Fu

    Localize-Then-Decide Guarantees for LLM Judgments

  • LinkedLinked via arxiv author · 85%Tianjin Huang

    Localize-Then-Decide Guarantees for LLM Judgments

Implements (incoming)

authored (incoming)

Related across the graph

Topics