Read original ↗
paperarXivTrust 82 · PrimaryPublished 14d agoLive · 11d ago

Robustifying Vision-Language Models via Test-Time Prompt Adaptation

Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing test-time adaptation methods typically rely on sample-level confidence heuristics, overlooking the intrinsic distributional structure of the data. This sample-centric approach limits robustness, as it fails to distinguish confident adversarial mispredictions from true semantic consistency. In this work, we observe that adversarial distortion is structurally brittle: while holistic representations are corrupted, semantic

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Xingyu Zhu

    Robustifying Vision-Language Models via Test-Time Prompt Adaptation

  • LinkedLinked via arxiv author · 85%Huanshen Wu

    Robustifying Vision-Language Models via Test-Time Prompt Adaptation

  • LinkedLinked via arxiv author · 85%Shuo Wang

    Robustifying Vision-Language Models via Test-Time Prompt Adaptation

  • LinkedLinked via arxiv author · 85%Beier Zhu

    Robustifying Vision-Language Models via Test-Time Prompt Adaptation

  • LinkedLinked via arxiv author · 85%Jiannan Ge

    Robustifying Vision-Language Models via Test-Time Prompt Adaptation

  • LinkedLinked via arxiv author · 85%Jiaheng Zhang

    Robustifying Vision-Language Models via Test-Time Prompt Adaptation

  • LinkedLinked via arxiv author · 85%Tianlong Chen

    Robustifying Vision-Language Models via Test-Time Prompt Adaptation

authored (incoming)

Related across the graph

Topics