TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models
Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated by the large number of visual tokens fed to the language model. Existing token reduction methods rely on attention-based scores or pairwise similarity, without an explicit semantic representation of each token. We introduce TORINO (TOken Reduction via Interpretable coNcept Overlap), a plug-and-play framework for adaptive visual token reduction in VLMs that requires no fine-tuning of the underlying model. TORINO leverages Sparse Autoencoders (SAEs) to project
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%vlm-starter →
- PossiblePossibly related (embedding) · 51%llmsresearch/llm-flashcards →
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “TORINO: Token Reduction via Interpretable Concept Overlap in” ≈ “VioletVision-3B””
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “TORINO: Token Reduction via Interpretable Concept Overlap in” ≈ “pytorch/vision””
- FuzzyOverlapping authors or contributors · 62%open-webui/open-webui →
“Shared author/contributor keys: nguyen”
- LinkedLinked via arxiv author · 85%Riccardo Renzulli →
“TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models”
- LinkedLinked via arxiv author · 85%Gabriele Spadaro →
“TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models”
- LinkedLinked via arxiv author · 85%Shruthi Gowda →
“TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models”
