HoloCount: A Holistic Visual Counting Benchmark for MLLMs
Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning. While Multimodal Large Language Models (MLLMs) have achieved remarkable success in qualitative scene understanding, their quantitative precision remains a significant bottleneck, often characterized by persistent numerical hallucinations. Existing counting benchmarks primarily focus on basic perception in simplified contexts, failing to capture the complex failure modes that emerge under logical constraints or adversarial conditions. To address t
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%Atomic-man007/Awesome_Multimodel_LLM →
- PossiblePossibly related (embedding) · 53%voxel51/fiftyone →
- PossiblePossibly related (embedding) · 51%vlm-starter →
- LinkedLinked via arxiv author · 85%Jinhong Deng →
“HoloCount: A Holistic Visual Counting Benchmark for MLLMs”
- LinkedLinked via arxiv author · 85%Limeng Qiao →
“HoloCount: A Holistic Visual Counting Benchmark for MLLMs”
- LinkedLinked via arxiv author · 85%Guanglu Wan →
“HoloCount: A Holistic Visual Counting Benchmark for MLLMs”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: wan”
- FuzzySimilar title/name (fuzzy) · 59%jeinlee1991/chinese-llm-benchmark →
“Fuzzy title match (0.73): “HoloCount: A Holistic Visual Counting Benchmark for MLLMs” ≈ “jeinlee1991/chinese-llm-benchmark””
