Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using dat
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%Large Language Models in Life Science Research: What Scientists Need to Know - Technology Networks →
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Patrick Phuoc Do →
“Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy”
- LinkedLinked via arxiv author · 85%Chau M. Ta →
“Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy”
- LinkedLinked via arxiv author · 85%Chaoli Wang →
“Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy”
