Towards Robustness against Typographic Attack with Training-free Concept Localization
Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Models (LVLMs). Despite their widespread adoption, CLIP models exhibit a critical yet underexplored failure mode: irrelevant text appearing within images confounds visual representations, biasing them toward lexical meaning rather than true visual semantics. This robustness issue, commonly described as a Typographic Attack (TA), exposes a vulnerability that poses a significant risk to safety-critical applications such as autonomous driving. To achieve
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%vlm-starter →
- PossiblePossibly related (embedding) · 52%VioletVision-3B →
- LinkedLinked via arxiv author · 85%Bohan Liu →
“Towards Robustness against Typographic Attack with Training-free Concept Localization”
- LinkedLinked via arxiv author · 85%Wenqian Ye →
“Towards Robustness against Typographic Attack with Training-free Concept Localization”
- LinkedLinked via arxiv author · 85%Guangzhi Xiong →
“Towards Robustness against Typographic Attack with Training-free Concept Localization”
- LinkedLinked via arxiv author · 85%Zhenghao He →
“Towards Robustness against Typographic Attack with Training-free Concept Localization”
- LinkedLinked via arxiv author · 85%Sanchit Sinha →
“Towards Robustness against Typographic Attack with Training-free Concept Localization”
- LinkedLinked via arxiv author · 85%Aidong Zhang →
“Towards Robustness against Typographic Attack with Training-free Concept Localization”
