Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models
Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descriptions as a benchmark, and have not focused on understanding the model's error types. Here, we introduce the Complex Social Behavior (CSB) dataset, containing 100 images depicting complex social interactions/behaviors. We analyze the progression of scene descriptions over a decade (2017-2025) of VLMs (four pre-Multimodal Large Language Mod
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%VioletVision-3B →
- PossiblePossibly related (embedding) · 53%vlm-starter →
- LinkedLinked via arxiv author · 85%Shravan Murlidaran →
“Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models”
- LinkedLinked via arxiv author · 85%Miguel P. Eckstein →
“Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models”
