Robustifying Vision-Language Models via Test-Time Prompt Adaptation
Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing test-time adaptation methods typically rely on sample-level confidence heuristics, overlooking the intrinsic distributional structure of the data. This sample-centric approach limits robustness, as it fails to distinguish confident adversarial mispredictions from true semantic consistency. In this work, we observe that adversarial distortion is structurally brittle: while holistic representations are corrupted, semantic
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Xingyu Zhu →
“Robustifying Vision-Language Models via Test-Time Prompt Adaptation”
- LinkedLinked via arxiv author · 85%Huanshen Wu →
“Robustifying Vision-Language Models via Test-Time Prompt Adaptation”
- LinkedLinked via arxiv author · 85%Shuo Wang →
“Robustifying Vision-Language Models via Test-Time Prompt Adaptation”
- LinkedLinked via arxiv author · 85%Beier Zhu →
“Robustifying Vision-Language Models via Test-Time Prompt Adaptation”
- LinkedLinked via arxiv author · 85%Jiannan Ge →
“Robustifying Vision-Language Models via Test-Time Prompt Adaptation”
- LinkedLinked via arxiv author · 85%Jiaheng Zhang →
“Robustifying Vision-Language Models via Test-Time Prompt Adaptation”
- LinkedLinked via arxiv author · 85%Tianlong Chen →
“Robustifying Vision-Language Models via Test-Time Prompt Adaptation”
