4DAnyone: Create Anyone in 4D from a Casual Monocular Video
We present 4DAnyone, a framework for reconstructing 4D humans from an uncalibrated monocular video by generating reconstruction-grade multiview-consistent videos and lifting them into 4D Gaussian Splatting (4DGS). Existing camera-controlled video diffusion models synthesize plausible novel-view videos but fail to maintain consistency when scaled to the tens of target views required for 4DGS reconstruction. We identify this failure as a bounded-attention-context problem: when target views exceed the capacity of a single DiT forward pass, they must be split into groups, exposing two coupled bott
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Yudong Jin →
“4DAnyone: Create Anyone in 4D from a Casual Monocular Video”
- LinkedLinked via arxiv author · 85%Pengtao Xie →
“4DAnyone: Create Anyone in 4D from a Casual Monocular Video”
- LinkedLinked via arxiv author · 85%Qihang Zhang →
“4DAnyone: Create Anyone in 4D from a Casual Monocular Video”
- LinkedLinked via arxiv author · 85%Zehong Shen →
“4DAnyone: Create Anyone in 4D from a Casual Monocular Video”
- LinkedLinked via arxiv author · 85%Mingzhen Xu →
“4DAnyone: Create Anyone in 4D from a Casual Monocular Video”
- LinkedLinked via arxiv author · 85%Yujun Shen →
“4DAnyone: Create Anyone in 4D from a Casual Monocular Video”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: zhou”
- FuzzyOverlapping authors or contributors · 62%HKUDS/LightRAG →
“Shared author/contributor keys: jin”
