Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking

Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated LLM-asa-Judge approaches. Whether such evaluators replicate clinical calibration and caution, however, remains untested. We introduce MedQADE, the first standardised open-response clinical benchmark for German, a major clinical language lacking native evaluation infrastructure, comprising 3,800 items annotated by ten practising physicians and nine Large Language Model (LLM) evaluators. The top-performing evaluator model, Gemini 3 Flash, reached

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Explains

authored (incoming)

Covers (incoming)

Implements (incoming)

Related across the graph

personWilliam PhilippnewsThe AI Validation Gap: Decision Support Tools Are Outrunning Their Own Evidence - The Clinical Trial VanguardpersonSebastian FudickarpersonMarkus HobertpersonTheresa PauluspersonFinn FassbenderpersonRebecca HerzogpersonJohanna ReimerpersonLukas GoedenewsBenchmarking large language models against practicing clinicians on psychopathological assessment - NaturenewsOpen-weight 4B models approach o3-level medical question answering in Swedish [P]personRonald BöcknewsTowards AI-augmented decision making in psychiatrynewsAddressing benchmarking gaps in large language models for health and medicine with dynamic red-teaming - NaturepersonThorsten LangernewsOpen-source Python library + no-code web dashboard for evaluating oncology AI models at clinical decision thresholds. [P]personMartje Paulyrepotoxy4ny/redteam-ai-benchmarknewsAI Analysis of Perioperative Data Improves Prediction of Postoperative Complications - OR Management NewspersonSebastian LönspersonIp Chi WangnewsThe AI Validation Gap: Decision Support Tools Are Outrunning Their Own Evidence - clinicaltrialvanguard.comtutorialEvaluate a model properlynewsArtificial Intelligence-Augmented Standardized Patient Models for AETCOM (Attitude, Ethics, and Communication) Competency Evaluation: A Pilot Study - CureuspersonAlexander BaumannnewsIntroducing GeneBench-PronewsCo-pilot, Not Autopilot: A Practical Method for Using Large Language Models in Interventional Cardiology - EMJrepocimeister/tokenizer-intrinsic-evals

Topics