LLM Judges Can Be Too Generous When There Is No Reference Answer
LLM judges are increasingly being used to evaluate open-ended model responses, often in no-reference settings where a ground-truth answer is unavailable. However, can they reliably assess in such evaluation setups? We explore this question in this paper through a two stage pipeline with a) calibration experiments that assess the judge model's knowledge of the task it is evaluating, and b) sensitivity experiments that assess how the judge model's performance is impacted by the presence and positioning of the reference answer in the prompt. Across experiments covering three languages, we show th
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Evaluate a model properly →
- PossiblePossibly related (embedding) · 47%OpenDCAI/One-Eval →
- PossiblePossibly related (embedding) · 46%Competence Gate: gating tool-use on a small model's internal confidence signal instead of its verbalised one — Qwen3.5-4B, open weights [P] →
- LinkedLinked via arxiv author · 85%Chalamalasetti Kranti →
“LLM Judges Can Be Too Generous When There Is No Reference Answer”
- LinkedLinked via arxiv author · 85%Sowmya Vajjala →
“LLM Judges Can Be Too Generous When There Is No Reference Answer”
