newsReddit r/MachineLearningTrust 52 · CommunityPublished 22d agoLive · 22d ago
Open-weight 4B models approach o3-level medical question answering in Swedish [P]
I have been running some experiments with smaller open-weight LLMs on multiple-choice questions of Swedish medical licensing exams. On a dataset called MedQA-SWE, GPT-4 scored 84% accuracy in 2024 and o3 scored 88% in 2025 on a smaller, overlapping dataset. With post-training (SFT) on data from earlier years, I got MedGemma-1.5-4B to a passing score of 60% on the final year’s exam. Find the implementation here:
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%Clinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question Answering →
- PossiblePossibly related (embedding) · 52%Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking →
- PossiblePossibly related (embedding) · 46%From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b →
- PossiblePossibly related (embedding) · 46%MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams →
- PossiblePossibly related (embedding) · 46%CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning →
Covers
paperClinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question AnsweringpaperClinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI BenchmarkingpaperFrom Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14bpaperMIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical ExamspaperCLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning
Related across the graph
paperClinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question AnsweringpaperCLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical ReasoningpaperClinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI BenchmarkingpaperFrom Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14bpaperMIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams
