Foveation-Guided Dynamic Token Selection for Robust and Efficient Vision Transformers
The human visual system (HVS) employs foveated sampling and eye movements to achieve efficient perception, conserving both metabolic energy and computational resources. Drawing inspiration from this robustness and adaptability, we introduce the Foveated Dynamic Transformer (FDT), a foveation-guided dynamic token-selection architecture that integrates these mechanisms into a vision transformer framework. The FDT exhibits strong resilience to various types of noise and adversarial attacks, despite not being explicitly trained for such challenges. This inherent robustness is achieved through the
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 46%VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization - Apple Machine Learning Research →
- PossiblePossibly related (embedding) · 45%How Outpost VFX Uses AWS to Accelerate AI Model Training for Visual Effects →
- PossiblePossibly related (embedding) · 45%autowarefoundation/vision_pilot →
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “Foveation-Guided Dynamic Token Selection for Robust and Effi” ≈ “VioletVision-3B””
- LinkedLinked via arxiv author · 85%Ibrahim Batuhan Akkaya →
“Foveation-Guided Dynamic Token Selection for Robust and Efficient Vision Transformers”
- LinkedLinked via arxiv author · 85%Kishaan Jeeveswaran →
“Foveation-Guided Dynamic Token Selection for Robust and Efficient Vision Transformers”
- LinkedLinked via arxiv author · 85%Bahram Zonooz →
“Foveation-Guided Dynamic Token Selection for Robust and Efficient Vision Transformers”
- LinkedLinked via arxiv author · 85%Elahe Arani →
“Foveation-Guided Dynamic Token Selection for Robust and Efficient Vision Transformers”
