3D-Aware VLMs with Implicit and Explicit Geometries
Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos. Our VLM-IE3D introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, as well as complementary Explicit Geometry Tokens (EGTs) that encode detailed geom
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%affaan-m/ECC →
“Shared author/contributor keys: jiang”
- FuzzyOverlapping authors or contributors · 62%BerriAI/litellm →
“Shared author/contributor keys: jiang”
- LinkedLinked via arxiv author · 85%Wenhao Liu →
“3D-Aware VLMs with Implicit and Explicit Geometries”
- LinkedLinked via arxiv author · 85%Xueying Jiang →
“3D-Aware VLMs with Implicit and Explicit Geometries”
- LinkedLinked via arxiv author · 85%Quanhao Qian →
“3D-Aware VLMs with Implicit and Explicit Geometries”
- LinkedLinked via arxiv author · 85%Deli Zhao →
“3D-Aware VLMs with Implicit and Explicit Geometries”
- LinkedLinked via arxiv author · 85%Ran Xu →
“3D-Aware VLMs with Implicit and Explicit Geometries”
