CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only by its final answer, making it hard to interpret and verify, especially in 3D radiology, where a diagnosis should be traceable to evidence in the scan. Existing chest CT question-answering datasets compound this by reducing expert radiology reports to answer-only pairs, dropping the reasoning that links findings to conclusions and omitting the patient history clinicians rely on. As a result, reasoning-capable 3D chest CT MLLMs remain o
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownUsing AI to help physicians diagnose rare genetic diseases affecting children →
- LinkedLinked via unknownRetrieval is underrated →
- LinkedLinked via unknownNorthwind AI →
- PossiblePossibly related (embedding) · 50%DIAGNijmegen/rse-grand-challenge →
- PossiblePossibly related (embedding) · 47%mlmed/torchxrayvision →
