Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection

Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a clinician-built synthetic benchmark and failure atlas that labels medical AI errors by severity (1--5) and safety gate type (missed urgent escalation, unsafe remote dosing, unsafe discharge reassurance, evidence fabrication, unsafe protocol execution, source support gap). The current public release (v0.2.1) contains 44 clinician-reviewed synthetic cases with severity annotations, a live HuggingFace leaderboard preview, a safety gate tax

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Implements (incoming)

Covers (incoming)

newsOpen-source Python library + no-code web dashboard for evaluating oncology AI models at clinical decision thresholds. [P]newsThe AI Validation Gap: Decision Support Tools Are Outrunning Their Own Evidence - The Clinical Trial VanguardnewsThe AI Validation Gap: Decision Support Tools Are Outrunning Their Own Evidence - clinicaltrialvanguard.comnewsCode-Free Classification of Pediatric Pneumonia on Chest Radiographs Using Google Cloud Vertex AI AutoML: A Proof-of-Concept Internal Validation Study - CureusnewsAI Tool Better than Current Screening for Hip Fracture Risk - Inside Precision MedicinenewsClinical AI Generators and Reviewers Must Be Evaluated Together - Bioengineer.orgnewsDevelopment of a machine-learning risk stratification tool for vasoactive medication need after two-bolus fluid resuscitation in pediatric suspected sepsis - NaturenewsLimited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison - NaturenewsTurning Medical AI Benchmark Scores into Trustworthy Clinical Readiness Claims - Bioengineer.orgnewsBeyond the Algorithm: A Critical and Evidence-Based Review of Artificial Intelligence in Chronic Pain Rehabilitation - CureusnewsClinical AI needs safeguards against hallucinations, data leaks and overreliance, review finds - Medical Xpressnews2MM: AI Roundup: Food and Drug Administration reviews liver injury prediction tool, Joint Commission launches healthcare artificial intelligence certification, governance playbooks aim to standardize adoption, and pediatric hospitals bring generative tools to t - 2 Minute MedicinenewsSafety and alignment in an era of long-horizon modelsnewsHuman-Governed Validation of Artificial Intelligence-Generated Medical Assessment Artifacts: A Technical Report - Cureus

authored (incoming)

Related across the graph

newsThe AI Validation Gap: Decision Support Tools Are Outrunning Their Own Evidence - The Clinical Trial VanguardpersonGoktug OzkannewsSafety and alignment in an era of long-horizon modelsnewsDevelopment of a machine-learning risk stratification tool for vasoactive medication need after two-bolus fluid resuscitation in pediatric suspected sepsis - NaturenewsHuman-Governed Validation of Artificial Intelligence-Generated Medical Assessment Artifacts: A Technical Report - CureusnewsChina works on AI safety benchmark as regulators target large model risks - South China Morning PostnewsOpen-source Python library + no-code web dashboard for evaluating oncology AI models at clinical decision thresholds. [P]news2MM: AI Roundup: Food and Drug Administration reviews liver injury prediction tool, Joint Commission launches healthcare artificial intelligence certification, governance playbooks aim to standardize adoption, and pediatric hospitals bring generative tools to t - 2 Minute MedicinenewsOIG Report: VA’s Aggressive Deployment of AI in the Clinical Field Considered ‘Risky’ - U.S. MedicinenewsAI Tool Better than Current Screening for Hip Fracture Risk - Inside Precision MedicinenewsBeyond the Algorithm: A Critical and Evidence-Based Review of Artificial Intelligence in Chronic Pain Rehabilitation - CureusnewsClinical AI needs safeguards against hallucinations, data leaks and overreliance, review finds - Medical XpressnewsLimited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison - NaturenewsClinical AI Generators and Reviewers Must Be Evaluated Together - Bioengineer.orgnewsThe AI Validation Gap: Decision Support Tools Are Outrunning Their Own Evidence - clinicaltrialvanguard.comnewsCode-Free Classification of Pediatric Pneumonia on Chest Radiographs Using Google Cloud Vertex AI AutoML: A Proof-of-Concept Internal Validation Study - Cureusrepojeinlee1991/chinese-llm-benchmarknewsNew framework improves clinical reasoning and decision making in AI systems - Tech XplorenewsIntroducing GeneBench-PronewsTurning Medical AI Benchmark Scores into Trustworthy Clinical Readiness Claims - Bioengineer.org

Topics