Read original ↗
paperarXivTrust 82 · PrimaryPublished 27d agoLive · 26d ago

Searching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language Models

Vision-language models normally execute the same complete vision encoder for every question, even when OCR, counting, object, attribute, and spatial queries may not require identical computation. We study whether fixed-budget combinations of vision blocks can be skipped without fine-tuning. A shared K-block route skips one searched set of exactly K blocks for every question, while a capability-specific K-block policy selects one same-size route using a known capability label. We introduce a source-balanced evolutionary search and compare it with independent ranking, contiguous removal, and ran

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B

    Fuzzy title match (0.73): “Searching for Task-Specific Vision Paths: Evolutionary Block” ≈ “VioletVision-3B”

  • FuzzySimilar title/name (fuzzy) · 84%pytorch/vision

    Fuzzy title match (0.92): “Searching for Task-Specific Vision Paths: Evolutionary Block” ≈ “pytorch/vision”

  • LinkedLinked via arxiv author · 85%Tarun Tomar

    Searching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language Models

Has model

Implements (incoming)

authored (incoming)

Related across the graph

Topics