Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding
Multimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-trained foundation model to further enhance the spatial awareness of MLLMs. In this paper, we first reveal that when integrating diverse foundation models into MLLMs, different models provide complementary spatial priors that benefit different tasks. Motivated by this, we propose $\textbf{ViPS}$, a novel multi-model prior framework designed to fully unleash the potential of incorporating multiple $\textbf{Vi}$sual $\t
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%hiyouga/LlamaFactory →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%ultralytics/ultralytics →
“Shared author/contributor keys: han”
- FuzzyOverlapping authors or contributors · 62%janhq/jan →
“Shared author/contributor keys: han”
- LinkedLinked via arxiv author · 85%Xiao Lin →
“Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding”
- LinkedLinked via arxiv author · 85%Xiaohu Huang →
“Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding”
- LinkedLinked via arxiv author · 85%Kai Han →
“Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding”
