IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer
Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding. Meanwhile, existing semantic reconstruction and 3D-aware vision-language methods largely rely on externally extracted 2D semantic cues or loosely coupled geometry inputs, limiting unified geometry-instance learning in long dynamic s
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%TauricResearch/TradingAgents →
“Shared author/contributor keys: xiao”
- FuzzyOverlapping authors or contributors · 62%pytorch/pytorch →
“Shared author/contributor keys: zou”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: zhou”
- LinkedLinked via arxiv author · 85%Xiaolin Zhou →
“IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer”
- LinkedLinked via arxiv author · 85%Fangzhou Hong →
“IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer”
- LinkedLinked via arxiv author · 85%Zhizhong Su →
“IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer”
- LinkedLinked via arxiv author · 85%Dingwen Zhang →
“IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer”
