The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs
Despite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting. We investigate whether this reflects missing internal knowledge or a gap between internal representations and verbalized outputs. Training simple probes on activations from four VLMs across five counting datasets reveals that nonlinear probes can reliably detect counting errors, suggesting that VLMs often encode the correct count even when they output the wrong answer. SVCCA analysis shows that probes trained on ground-truth counts and probes trained on model outputs oc
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 63%vlm-starter →
- PossiblePossibly related (embedding) · 56%VioletVision-3B →
- LinkedLinked via arxiv author · 85%Ahmed Oumar El-Shangiti →
“The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs”
- LinkedLinked via arxiv author · 85%Abzal Nurgazy →
“The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs”
- LinkedLinked via arxiv author · 85%Hilal AlQuabeh →
“The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs”
- LinkedLinked via arxiv author · 85%Nikolai Rozanov →
“The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs”
- LinkedLinked via arxiv author · 85%Kentaro Inui →
“The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs”
