Read original ↗
paperarXivTrust 82 · PrimaryPublished 6d agoLive · 3d ago

HealMed: Multilingual Evaluation of Large Language Models in Medicine

We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions. Each translation was evaluated and revised by two experts fluent in English and the corresponding target language. On HealMed, performance declined most in low-resource languages, although the size of the gap varied ma

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Implements (incoming)

authored (incoming)

Covers (incoming)

Related across the graph

personYingjian ChenpersonAosong FengpersonDhruvanewsClinician Use of a General-Purpose Large Language Model in Hospital Medicine: A Mixed-Methods Pilot Study - CureuspersonMichihiro YasunagapersonAbdul SamadpersonCesar CaraballopersonSantiago Gudiño-RosalespersonJihyo KwakpersonChanjun ParkpersonKanyakorn VeerakanjanapersonHugo Toshio ItikawapersonYusuke IwasawapersonIrene LipersonInsook ChonewsComparative Performance of Large Language Models in the Polish State Specialization Examination in Anesthesiology and Intensive Care Medicine - CureuspersonQingyu ChenpersonManish Guptarepokeras-team/keraspersonAkbar FaruqipersonJinghui LupersonJaewoo KangpersonHaoyu ZhangpersonFan GaopersonHyunjae Kimrepoaffaan-m/ECCpersonSherry T. TongnewsEvidence, use cases, and implementation safeguards of large language models in primary care - NaturepersonHeuiseok LimpersonEunji JeonpersonXiujie ChenpersonIsarar SiddiquepersonIsabelli MartinspersonCibele BrandãopersonHang JiangnewsAddressing benchmarking gaps in large language models for health and medicine with dynamic red-teaming - NaturepersonYutaka MatsuorepoHKUDS/LightRAGpersonKevin W. JinnewsLarge Language Models Are Still Getting Stronger, but Researchers Face New Bottlenecks in Data, Evaluation, and Safety | Newswise - NewswisepersonIsrar AhmedpersonZeo LapaluspersonRex YingpersonLuis Guilherme CardosopersonRenee DuapersonZixin XupersonKyeongeon LeepersonMinjin KimpersonGabriel Madera-SantiagopersonTianxing WupersonPiyalitt IttichaiwongpersonEthan GohpersonEdison Marrese-Taylorrepobrowser-use/browser-userepoBerriAI/litellm

Topics