Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations
Patients seeking medical information often ask questions that embed incorrect assumptions or misconceptions. In such cases, safe medical communication requires not only answering the question, but identifying and correcting the underlying false belief. These interactions naturally unfold over multiple turns, a pattern now mirrored in interactions with LLMs. Yet current evaluation frameworks do not capture model behavior in these settings, where misconceptions can emerge, persist, or evolve over the course of a conversation. Whether LLMs can reliably correct such misconceptions over time remain
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 58%Clinician Use of a General-Purpose Large Language Model in Hospital Medicine: A Mixed-Methods Pilot Study - Cureus →
- PossiblePossibly related (embedding) · 51%Large language models exhibit stigmatizing behaviour in contextual judgements of health conditions - Nature →
- PossiblePossibly related (embedding) · 51%KennispuntTwente/tidyprompt →
- PossiblePossibly related (embedding) · 49%yubol-bobo/Awesome-Multi-Turn-LLMs →
- PossiblePossibly related (embedding) · 48%Evaluating long-term memory limits in stateless LLM chatbots — feedback needed [D] →
- LinkedLinked via arxiv author · 85%Monica Munnangi →
“Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations”
- LinkedLinked via arxiv author · 85%Saiph Savage →
“Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations”
- PossiblePossibly related (embedding) · 46%When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure mode. →
