When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images
*Chulin Zhao and Ruoqi Hu contributed equally to this work. State-of-the-art text-to-image (T2I) models exhibit pronounced and systematic defects when prompts involve intricate compositional factors such as multiple entities and multiple attributes. In this paper, we investigate how humans identify such defects. Specifically, we manually select 651 reference images from the four categories of people, hand, object, and scene that exhibit complex compositional characteristics, from which prompts emphasizing compositional factors are derived by manually editing ChatGPT-generated prompts. We the
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%Understanding Generative AI: Beyond Chatbots and Prompts - themetropolitan.metrostate.edu →
- PossiblePossibly related (embedding) · 54%Understanding Generative AI: Beyond Chatbots and Prompts - Metro State University →
- PossiblePossibly related (embedding) · 51%Large Tabular Models Excel Where LLMs Fail →
- FuzzyOverlapping authors or contributors · 62%hiyouga/LlamaFactory →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos →
“Shared author/contributor keys: lin”
- PossiblePossibly related (embedding) · 49%A dataset with 52 Text to image model evaluation [P] →
- LinkedLinked via arxiv author · 85%Ruoqi Hu →
“When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images”
- LinkedLinked via arxiv author · 85%Chulin Zhao →
“When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images”
