SceneBind: Binding What and Where Across Vision, Audio and Language
We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots. This representation explicitly captures object-level semantics, spatial attributes, and uncertainty. We further propose SceneB
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “SceneBind: Binding What and Where Across Vision, Audio and L” ≈ “VioletVision-3B””
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “SceneBind: Binding What and Where Across Vision, Audio and L” ≈ “pytorch/vision””
- LinkedLinked via arxiv author · 85%Mingfei Chen →
“SceneBind: Binding What and Where Across Vision, Audio and Language”
- LinkedLinked via arxiv author · 85%Zijun Cui →
“SceneBind: Binding What and Where Across Vision, Audio and Language”
- LinkedLinked via arxiv author · 85%Ruoke Zhang →
“SceneBind: Binding What and Where Across Vision, Audio and Language”
- LinkedLinked via arxiv author · 85%Hyeonggon Ryu →
“SceneBind: Binding What and Where Across Vision, Audio and Language”
- LinkedLinked via arxiv author · 85%Eli Shlizerman →
“SceneBind: Binding What and Where Across Vision, Audio and Language”
