Read original ↗
paperarXivTrust 82 · PrimaryPublished 26d agoLive · 25d ago

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple, free-form semantic aspects of similarity. Benchmarking a broad range of frontier vision-language models (VLMs) reveals a considerable performance gap compared to

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%Tongyi-MAI/Z-Image-Turbo

    Fuzzy title match (0.73): “The Many Senses of Visual Similarity: A Text-Prompted Image ” ≈ “Tongyi-MAI/Z-Image-Turbo”

  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • LinkedLinked via arxiv author · 85%Sheng-Yu Wang

    The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

  • LinkedLinked via arxiv author · 85%Yotam Nitzan

    The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

  • LinkedLinked via arxiv author · 85%Aaron Hertzmann

    The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

  • LinkedLinked via arxiv author · 85%Jun-Yan Zhu

    The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

  • LinkedLinked via arxiv author · 85%Eli Shechtman

    The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

Has model

Implements (incoming)

authored (incoming)

Covers (incoming)

Related across the graph

Topics