Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees
Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer exists. However, objective tasks introduce a distinct reliability challenge for reference-free LLM judging. In the absence of a reference answer, the judge evaluates factual correctness either through its parametric knowledge or through tool augmentation. Although the former enables efficient evaluation, the judge may hallucinate or lack sufficient evidence for its verdi
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 60%Evaluate a model properly →
- PossiblePossibly related (embedding) · 54%Cutting RAG inference costs 6x starts with deciding what never reaches the LLM →
- PossiblePossibly related (embedding) · 45%Why Self-Correction Loops Can Degrade Reliability in LLM Pipelines (85% Down to 62%) →
- LinkedLinked via arxiv author · 85%Sher Badshah →
“Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees”
- LinkedLinked via arxiv author · 85%Ali Emami →
“Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees”
- LinkedLinked via arxiv author · 85%Hassan Sajjad →
“Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees”
