Read original ↗
paperarXivTrust 82 · PrimaryPublished 24d agoLive · 23d ago

Test-Time Training for Modality Order Consistency in Vision-Language Models

We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperforms question-first prompting, revealing a repeatable modality order failure. We use this gap to design an order-consistent test-time training method. Our method substantially closes the modality-order gap across all evaluated settings. Surprisingly, it also yields consistent improvements in the stronger image-first branch over the baseline, hence bootstrapping

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B

    Fuzzy title match (0.73): “Test-Time Training for Modality Order Consistency in Vision-” ≈ “VioletVision-3B”

  • FuzzySimilar title/name (fuzzy) · 84%pytorch/vision

    Fuzzy title match (0.92): “Test-Time Training for Modality Order Consistency in Vision-” ≈ “pytorch/vision”

  • FuzzyOverlapping authors or contributors · 62%microsoft/ML-For-Beginners

    Shared author/contributor keys: gupta

  • LinkedLinked via arxiv author · 85%Aditi Gupta

    Test-Time Training for Modality Order Consistency in Vision-Language Models

  • LinkedLinked via arxiv author · 85%Yossi Gandelsman

    Test-Time Training for Modality Order Consistency in Vision-Language Models

Has model

Implements (incoming)

authored (incoming)

Related across the graph

Topics