RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching
As AI systems are increasingly used to draft radiology reports, reliably evaluating their clinical quality remains a critical challenge. Large language model (LLM)-based metrics are now the best-correlated with radiologist judgment, yet they output a single opaque score that neither a clinician nor a model builder can easily interpret or audit. We introduce RadMatch, a multi-stage, LLM-based metric that decomposes report comparison into a structured finding-level matching with significance-aware scoring and error characterization across seven clinical attribute dimensions (status, location, se
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%Using AI Could Help Patients Understand Radiology Reports - Radiological Society of North America | RSNA →
- PossiblePossibly related (embedding) · 50%Addressing benchmarking gaps in large language models for health and medicine with dynamic red-teaming - Nature →
- PossiblePossibly related (embedding) · 49%Towards AI-augmented decision making in psychiatry →
- PossiblePossibly related (embedding) · 49%Open-source Python library + no-code web dashboard for evaluating oncology AI models at clinical decision thresholds. [P] →
- LinkedLinked via arxiv author · 85%Charles Corbière →
“RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching”
- LinkedLinked via arxiv author · 85%Léo Machado →
“RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching”
- LinkedLinked via arxiv author · 85%Aubin Charley →
“RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching”
- LinkedLinked via arxiv author · 85%Baptiste Callard →
“RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching”
