Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models
Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering accuracy marks the response as wrong, while object-hallucination metrics may regard both the object and attribute as image-supported; neither reveals the transfer. This study formalizes this blind spot as Dense Same-Class Attribute Misbinding (DSCAM) and presents InstaBind-Lite, a controlled benchmark that makes it directly measurable. Its 524 images contain 529 curated groups of 3-6 same-class entities, 1773 boxed
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “Diagnosing Dense Same-Class Attribute Misbinding in Large Vi” ≈ “VioletVision-3B””
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “Diagnosing Dense Same-Class Attribute Misbinding in Large Vi” ≈ “pytorch/vision””
- FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%TauricResearch/TradingAgents →
“Shared author/contributor keys: xiao”
- FuzzyOverlapping authors or contributors · 62%hiyouga/LlamaFactory →
“Shared author/contributor keys: lin”
- LinkedLinked via arxiv author · 85%Yuanzhi Xu →
“Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models”
- LinkedLinked via arxiv author · 85%Qian Gao →
“Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models”
- LinkedLinked via arxiv author · 85%Xiangjun Fan →
“Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models”
