On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline
Vision-Language Pre-training Models (VLPMs) are known to be vulnerable to adversarial attacks. Recent transferable attacks on VLPMs have followed a common pipeline with complicated loss functions or multi-stage text/image attacks. However, in this paper, we demonstrate that such a sophisticated attack pipeline can be simpler yet more successful. Specifically, we identify three previously overlooked issues caused by inappropriate cross-modal interactions and excessive operations. To address them, we propose the Simple Vision-Language Attack (SimVLA) pipeline, which observably improves transfera
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “On Success and Simplicity: A Second Look at Transferable Vis” ≈ “VioletVision-3B””
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “On Success and Simplicity: A Second Look at Transferable Vis” ≈ “pytorch/vision””
- FuzzyOverlapping authors or contributors · 62%hiyouga/LlamaFactory →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos →
“Shared author/contributor keys: lin”
- LinkedLinked via arxiv author · 85%Yuchen Ren →
“On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline”
- LinkedLinked via arxiv author · 85%Zhengyu Zhao →
“On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline”
- LinkedLinked via arxiv author · 85%Chenhao Lin →
“On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline”
- LinkedLinked via arxiv author · 85%Bo Yang →
“On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline”
