Collapsibility of Performance Metrics in Clinical Predictive AI
Background: Population level assessments of predictive artificial intelligence (AI) can conceal performance disparities across subgroups. Fairness evaluations commonly rely on performance analyses across subgroups. However, some performance metrics are non-collapsible, meaning that the overall population performance value does not equal the weighted average of subgroup specific values. Objective: To examine the collapsibility properties of commonly reported performance metrics in predictive AI, with a focus on the area under the receiver operating characteristic curve (AUC, also known as c-s
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%Open-source Python library + no-code web dashboard for evaluating oncology AI models at clinical decision thresholds. [P] →
- PossiblePossibly related (embedding) · 49%Why Aren’t We Measuring How AI Affects Humans? →
- PossiblePossibly related (embedding) · 48%Artificial Intelligence-Augmented Standardized Patient Models for AETCOM (Attitude, Ethics, and Communication) Competency Evaluation: A Pilot Study - Cureus →
- PossiblePossibly related (embedding) · 47%Human-Governed Validation of Artificial Intelligence-Generated Medical Assessment Artifacts: A Technical Report - Cureus →
- LinkedLinked via arxiv author · 85%João Matos →
“Collapsibility of Performance Metrics in Clinical Predictive AI”
- LinkedLinked via arxiv author · 85%Ben Van Calster →
“Collapsibility of Performance Metrics in Clinical Predictive AI”
- LinkedLinked via arxiv author · 85%Richard D. Riley →
“Collapsibility of Performance Metrics in Clinical Predictive AI”
- LinkedLinked via arxiv author · 85%Paula Dhiman →
“Collapsibility of Performance Metrics in Clinical Predictive AI”
