Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA
Healthcare multimodal AI must combine visual and textual evidence while remaining reliable and interpretable. Using MediaEval Medico 2025 as a retrospective GI endoscopy case study, we analyze design choices across nine documented systems for question answering and explanation quality. Parameter-efficient adaptation of pretrained backbones provides strong challenge performance, but answer-level gains do not consistently translate into faithful and complete clinical reasoning. Methods enforcing structured reasoning and explicit grounding show more reliable behavior across heterogeneous question
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Sushant Gautam →
“Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA”
- LinkedLinked via arxiv author · 85%Vajira Thambawita →
“Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA”
- LinkedLinked via arxiv author · 85%Michael A. Riegler →
“Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA”
- LinkedLinked via arxiv author · 85%Pål Halvorsen →
“Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA”
- LinkedLinked via arxiv author · 85%Steven A. Hicks →
“Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA”
- PossiblePossibly related (embedding) · 60%Prompt-Based Knowledge Fusion Improves Faithfulness in Open-Domain Question Answering - Bioengineer.org →
