CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the ability to reason over real-world household environments containing multiple interconnected rooms and diverse object categories. We introduce CAIRN, a topology-aware 3D-LLM for multi-room 3D scene understanding. CAIRN aligns transformer attention with scene hierarchy, giving the model explicit awareness of object-level relations and room-level connectivity. It enriches object tokens with room-local relational context via a graph neural network, intro
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%ghbalf/freecad-ai →
- PossiblePossibly related (embedding) · 49%huggingface/transformers →
- LinkedLinked via arxiv author · 85%He Liang →
“CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models”
- LinkedLinked via arxiv author · 85%Chenyang Ma →
“CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models”
- LinkedLinked via arxiv author · 85%Yiming Zhang →
“CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models”
- LinkedLinked via arxiv author · 85%Sangyun Shin →
“CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models”
- LinkedLinked via arxiv author · 85%Andrew Markham →
“CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models”
- LinkedLinked via arxiv author · 85%Niki Trigoni →
“CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models”
