The strength of clinical evidence is recoverable from language model representations but not from their stated grades
Large language models (LLMs) increasingly summarize clinical evidence, where a claim's weight depends on how strongly it is supported. Yet these models convey confidence poorly, and properties they never state, such as truth, are often readable from their activations. Whether a clinical model registers evidence strength, distinct from truth, and states it when asked is untested, and any such signal could be lexical. We compiled 45,134 clinical claims from six public sources, harmonized 20,611 into a four-level evidence grade under three independent frameworks, and tested 22 local, open-weight
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownTowards AI-augmented decision making in psychiatry →
- PossiblePossibly related (embedding) · 48%Large language models exhibit stigmatizing behaviour in contextual judgements of health conditions - Nature →
- PossiblePossibly related (embedding) · 55%Clinician Use of a General-Purpose Large Language Model in Hospital Medicine: A Mixed-Methods Pilot Study - Cureus →
- PossiblePossibly related (embedding) · 57%Benchmarking large language models against practicing clinicians on psychopathological assessment - Nature →
- PossiblePossibly related (embedding) · 51%Comparative Performance of Large Language Models in the Polish State Specialization Examination in Anesthesiology and Intensive Care Medicine - Cureus →
