Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 28d ago

Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs

Understanding how vision-language models (VLMs) interpret data visualizations remains an open problem, and is increasingly important as these models are used for analytical tasks where reliable reasoning is essential. We introduce a lightweight, diagnostic saliency map method tailored for text generation over images using transformer models, the current state-of-the-art models in visualization interpretation. Our approach aggregates the language model's attention over the visual tokens across all heads and layers, then maps this attention back onto the vision encoder's patch grid to localise i

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Maeve Hutchinson

    Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs

  • LinkedLinked via arxiv author · 85%Abderrahmane Wassim Mehdaoui

    Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs

  • LinkedLinked via arxiv author · 85%Pranava Madhyastha

    Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs

authored (incoming)

Related across the graph

Topics