Read original ↗
paperarXivTrust 82 · PrimaryPublished 25d agoLive · 24d ago

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding. Meanwhile, existing semantic reconstruction and 3D-aware vision-language methods largely rely on externally extracted 2D semantic cues or loosely coupled geometry inputs, limiting unified geometry-instance learning in long dynamic s

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%TauricResearch/TradingAgents

    Shared author/contributor keys: xiao

  • FuzzyOverlapping authors or contributors · 62%pytorch/pytorch

    Shared author/contributor keys: zou

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • FuzzyOverlapping authors or contributors · 62%sgl-project/sglang

    Shared author/contributor keys: zhou

  • LinkedLinked via arxiv author · 85%Xiaolin Zhou

    IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

  • LinkedLinked via arxiv author · 85%Fangzhou Hong

    IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

  • LinkedLinked via arxiv author · 85%Zhizhong Su

    IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

  • LinkedLinked via arxiv author · 85%Dingwen Zhang

    IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

Implements (incoming)

authored (incoming)

Related across the graph

Topics