MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection
Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a clinician-built synthetic benchmark and failure atlas that labels medical AI errors by severity (1--5) and safety gate type (missed urgent escalation, unsafe remote dosing, unsafe discharge reassurance, evidence fabrication, unsafe protocol execution, source support gap). The current public release (v0.2.1) contains 44 clinician-reviewed synthetic cases with severity annotations, a live HuggingFace leaderboard preview, a safety gate tax
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%New framework improves clinical reasoning and decision making in AI systems - Tech Xplore →
- PossiblePossibly related (embedding) · 52%China works on AI safety benchmark as regulators target large model risks - South China Morning Post →
- PossiblePossibly related (embedding) · 52%OIG Report: VA’s Aggressive Deployment of AI in the Clinical Field Considered ‘Risky’ - U.S. Medicine →
- PossiblePossibly related (embedding) · 52%Introducing GeneBench-Pro →
- FuzzySimilar title/name (fuzzy) · 59%jeinlee1991/chinese-llm-benchmark →
“Fuzzy title match (0.73): “MedFailBench: A Clinician-Built Open-Source Benchmark for Me” ≈ “jeinlee1991/chinese-llm-benchmark””
- LinkedLinked via arxiv author · 85%Goktug Ozkan →
“MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection”
