newsGoogle News — LLMTrust 62 · AggregatorPublished 1mo agoLive · 1mo ago
Benchmarking large language models against practicing clinicians on psychopathological assessment - Nature
Benchmarking large language models against practicing clinicians on psychopathological assessment Nature
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%The strength of clinical evidence is recoverable from language model representations but not from their stated grades →
- PossiblePossibly related (embedding) · 54%Multi-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC) →
- PossiblePossibly related (embedding) · 53%Emo-gml/Awesome-Mental-Health-LLMs →
- PossiblePossibly related (embedding) · 53%PacificAI/langtest →
- PossiblePossibly related (embedding) · 52%Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking →
- PossiblePossibly related (embedding) · 50%Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs →
- PossiblePossibly related (embedding) · 48%PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image →
- PossiblePossibly related (embedding) · 50%Gotta Catch them all: the modes of Sycophancy →
Covers
paperThe strength of clinical evidence is recoverable from language model representations but not from their stated gradespaperMulti-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC)repoEmo-gml/Awesome-Mental-Health-LLMsrepoPacificAI/langtestpaperClinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking
Covers (incoming)
paperJudge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMspaperPathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology ImagepaperGotta Catch them all: the modes of SycophancypaperAnchorScore: A CLIP-Based Diagnostic of MLLM Annotation DifficultypaperToward Better Assessment of LLMs' Performance in Clinical Error DetectionpaperMove by Move: Measuring and Steering How LLMs Conduct PsychotherapypaperAugmenting Interviewer Judgments of Patient Experience with Automatic Language Analysis
Related across the graph
paperAnchorScore: A CLIP-Based Diagnostic of MLLM Annotation DifficultypaperMove by Move: Measuring and Steering How LLMs Conduct PsychotherapypaperPathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology ImagepaperJudge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMsrepoEmo-gml/Awesome-Mental-Health-LLMspaperMulti-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC)paperClinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI BenchmarkingpaperGotta Catch them all: the modes of SycophancypaperAugmenting Interviewer Judgments of Patient Experience with Automatic Language AnalysisrepoPacificAI/langtestpaperThe strength of clinical evidence is recoverable from language model representations but not from their stated gradespaperToward Better Assessment of LLMs' Performance in Clinical Error Detection
