Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descriptions as a benchmark, and have not focused on understanding the model's error types. Here, we introduce the Complex Social Behavior (CSB) dataset, containing 100 images depicting complex social interactions/behaviors. We analyze the progression of scene descriptions over a decade (2017-2025) of VLMs (four pre-Multimodal Large Language Mod

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 54%VioletVision-3B
  • PossiblePossibly related (embedding) · 53%vlm-starter
  • LinkedLinked via arxiv author · 85%Shravan Murlidaran

    Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

  • LinkedLinked via arxiv author · 85%Miguel P. Eckstein

    Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

Has model

Implements

authored (incoming)

Related across the graph

Topics