Read original ↗
paperarXivTrust 82 · PrimaryPublished 25d agoLive · 22d ago

3D-Aware VLMs with Implicit and Explicit Geometries

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos. Our VLM-IE3D introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, as well as complementary Explicit Geometry Tokens (EGTs) that encode detailed geom

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • FuzzyOverlapping authors or contributors · 62%affaan-m/ECC

    Shared author/contributor keys: jiang

  • FuzzyOverlapping authors or contributors · 62%BerriAI/litellm

    Shared author/contributor keys: jiang

  • LinkedLinked via arxiv author · 85%Wenhao Liu

    3D-Aware VLMs with Implicit and Explicit Geometries

  • LinkedLinked via arxiv author · 85%Xueying Jiang

    3D-Aware VLMs with Implicit and Explicit Geometries

  • LinkedLinked via arxiv author · 85%Quanhao Qian

    3D-Aware VLMs with Implicit and Explicit Geometries

  • LinkedLinked via arxiv author · 85%Deli Zhao

    3D-Aware VLMs with Implicit and Explicit Geometries

  • LinkedLinked via arxiv author · 85%Ran Xu

    3D-Aware VLMs with Implicit and Explicit Geometries

Implements (incoming)

authored (incoming)

Related across the graph

Topics