Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 27d ago

SceneBind: Binding What and Where Across Vision, Audio and Language

We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots. This representation explicitly captures object-level semantics, spatial attributes, and uncertainty. We further propose SceneB

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B

    Fuzzy title match (0.73): “SceneBind: Binding What and Where Across Vision, Audio and L” ≈ “VioletVision-3B”

  • FuzzySimilar title/name (fuzzy) · 84%pytorch/vision

    Fuzzy title match (0.92): “SceneBind: Binding What and Where Across Vision, Audio and L” ≈ “pytorch/vision”

  • LinkedLinked via arxiv author · 85%Mingfei Chen

    SceneBind: Binding What and Where Across Vision, Audio and Language

  • LinkedLinked via arxiv author · 85%Zijun Cui

    SceneBind: Binding What and Where Across Vision, Audio and Language

  • LinkedLinked via arxiv author · 85%Ruoke Zhang

    SceneBind: Binding What and Where Across Vision, Audio and Language

  • LinkedLinked via arxiv author · 85%Hyeonggon Ryu

    SceneBind: Binding What and Where Across Vision, Audio and Language

  • LinkedLinked via arxiv author · 85%Eli Shlizerman

    SceneBind: Binding What and Where Across Vision, Audio and Language

Has model

Implements (incoming)

authored (incoming)

Related across the graph

Topics