Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs
Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance across vision-language tasks. However, their high inference cost, arising from both the large number of input visual tokens and the heavy computation of the large language model (LLM), remains a key barrier to practical deployment. Recent work attempts to reduce the cost by adaptively optimizing individual dimensions, e.g., pruning redundant visual tokens or skipping LLM layers and heads. Nonetheless, prior approaches typically treat these dimensions independently and overlook a fundamental coupling: the ava
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 26%sgl-project/sglang →
“Possibly related via embedding similarity 0.57 (not asserted). Timestamp check: artifact slightly before paper (-20d).”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%browser-use/browser-use →
“Shared author/contributor keys: lee”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Pengcheng Wang →
“Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs”
- LinkedLinked via arxiv author · 85%Zhiquan Wang →
“Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs”
- LinkedLinked via arxiv author · 85%Jayoung Lee →
“Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs”
- LinkedLinked via arxiv author · 85%Zhuoyan Xu →
“Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs”
