Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs
Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved. Using a structured evidence-sufficiency prompt as a test case, we asked whether it reduces unsafe overconfident answers, how far that effect depends on the scoring judge, and what it costs in helpfulness. Methods: In a retrospective public-data benchmark (Real-POCQi, HealthBench, MedRBench), four models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3)
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Benchmarking large language models against practicing clinicians on psychopathological assessment - Nature →
- PossiblePossibly related (embedding) · 48%What does "Safe AI" look like? [D] →
- PossiblePossibly related (embedding) · 45%Evaluate a model properly →
- PossiblePossibly related (embedding) · 45%How many LLMs does it take to reason through a clinical decision? - Medical Xpress →
- LinkedLinked via arxiv author · 85%Koyar Afrasyab →
“Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs”
