HealMed: Multilingual Evaluation of Large Language Models in Medicine
We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions. Each translation was evaluated and revised by two experts fluent in English and the corresponding target language. On HealMed, performance declined most in low-resource languages, although the size of the gap varied ma
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 69%Addressing benchmarking gaps in large language models for health and medicine with dynamic red-teaming - Nature →
- PossiblePossibly related (embedding) · 63%Comparative Performance of Large Language Models in the Polish State Specialization Examination in Anesthesiology and Intensive Care Medicine - Cureus →
- PossiblePossibly related (embedding) · 60%Large Language Models Are Still Getting Stronger, but Researchers Face New Bottlenecks in Data, Evaluation, and Safety | Newswise - Newswise →
- PossiblePossibly related (embedding) · 59%Clinician Use of a General-Purpose Large Language Model in Hospital Medicine: A Mixed-Methods Pilot Study - Cureus →
- FuzzyOverlapping authors or contributors · 62%affaan-m/ECC →
“Shared author/contributor keys: jiang”
- FuzzyOverlapping authors or contributors · 62%keras-team/keras →
“Shared author/contributor keys: jin”
- FuzzyOverlapping authors or contributors · 62%browser-use/browser-use →
“Shared author/contributor keys: lee”
- FuzzyOverlapping authors or contributors · 62%HKUDS/LightRAG →
“Shared author/contributor keys: jin”
