How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA
Compositional visual question answering requires Vision-Language Models (VLMs) to execute multiple reasoning operations like object selection, spatial relation resolution, and attribute verification. Despite strong aggregate performance, the mechanistic basis of VLM failures on this task remains underexplored. To address this gap, we analyze vision-operation misalignment in VLMs by examining how failures relate to specific reasoning operations and the internal computational pathways through which they arise and propagate. We introduce an Operation-centric mechanistic framework that decomposes
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “How Do VLMs Fail? Vision-Operation Misalignment in Compositi” ≈ “VioletVision-3B””
- LinkedLinked via arxiv author · 85%Navya Gupta →
“How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA”
- LinkedLinked via arxiv author · 85%Bingjie Xu →
“How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA”
- LinkedLinked via arxiv author · 85%Avinash Anand →
“How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA”
- LinkedLinked via arxiv author · 85%Timothy Liu →
“How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA”
- LinkedLinked via arxiv author · 85%Zhengchen Zhang →
“How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “How Do VLMs Fail? Vision-Operation Misalignment in Compositi” ≈ “pytorch/vision””
