Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination Detection
Understanding the frequency of factual errors in chatbot-generated text and evaluating systems that detect these errors is critical for determining chatbot safety. Yet factual-error detection is often treated as a single-pass, single-annotator labeling problem. In long-form chatbot responses, factual errors can be subtle and embedded within mostly correct text. We develop a multi-perspective annotation study of medically relevant chatbot responses, combining first-pass annotation, LLM-as-a-Judge (LaJ) candidate discovery, and two forms of adjudication: medical-expert and evidence-based fact-
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%Suspecting court of using AI, man injected prompts in filings to try to win case →
- PossiblePossibly related (embedding) · 53%Why AI Literacy Isn’t Enough to Fix Chatbots: 6 Questions with Medical AI and Health Law Scholar Sofia Palmieri - Petrie-Flom Center →
- PossiblePossibly related (embedding) · 53%Clinical AI needs safeguards against hallucinations, data leaks and overreliance, review finds - Medical Xpress →
- PossiblePossibly related (embedding) · 51%AI Hallucinations: Why Artificial Intelligence Can Sound Convincing While Getting the Facts Wrong - USA Herald →
