Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 29d ago

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

Multimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-trained foundation model to further enhance the spatial awareness of MLLMs. In this paper, we first reveal that when integrating diverse foundation models into MLLMs, different models provide complementary spatial priors that benefit different tasks. Motivated by this, we propose $\textbf{ViPS}$, a novel multi-model prior framework designed to fully unleash the potential of incorporating multiple $\textbf{Vi}$sual $\t

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%hiyouga/LlamaFactory

    Shared author/contributor keys: lin

  • FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos

    Shared author/contributor keys: lin

  • FuzzyOverlapping authors or contributors · 62%ultralytics/ultralytics

    Shared author/contributor keys: han

  • FuzzyOverlapping authors or contributors · 62%janhq/jan

    Shared author/contributor keys: han

  • LinkedLinked via arxiv author · 85%Xiao Lin

    Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

  • LinkedLinked via arxiv author · 85%Xiaohu Huang

    Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

  • LinkedLinked via arxiv author · 85%Kai Han

    Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

Implements (incoming)

authored (incoming)

Related across the graph

Topics