MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clinical practice. Many rely on synthetic conversations or patient simulators, omit patient-uploaded medical images, or evaluate open-ended clinical responses using multiple-choice or lexical-overlap metrics that poorly reflect clinical quality. We introduce \textbf{MedRealMM}, a large-scale benchmark for multimodal online medical consultation built from de-identified patient-doctor interactions collected from a nationwide Chinese internet hospital. Med
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%tyang816/Awesome-TCM-LLM →
- PossiblePossibly related (embedding) · 52%Atomic-man007/Awesome_Multimodel_LLM →
- PossiblePossibly related (embedding) · 51%New research shows how AMIE, our medical AI, could help manage health conditions. →
- PossiblePossibly related (embedding) · 50%Clinician Use of a General-Purpose Large Language Model in Hospital Medicine: A Mixed-Methods Pilot Study - Cureus →
- PossiblePossibly related (embedding) · 50%Co-pilot, Not Autopilot: A Practical Method for Using Large Language Models in Interventional Cardiology - EMJ →
- LinkedLinked via arxiv author · 85%Runhan Shi →
“MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation”
- LinkedLinked via arxiv author · 85%Quan Zhou →
“MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation”
- LinkedLinked via arxiv author · 85%Yuqian Xu →
“MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation”
