Read original ↗
paperarXivTrust 82 · PrimaryPublished 5d agoLive · 2d ago

4DAnyone: Create Anyone in 4D from a Casual Monocular Video

We present 4DAnyone, a framework for reconstructing 4D humans from an uncalibrated monocular video by generating reconstruction-grade multiview-consistent videos and lifting them into 4D Gaussian Splatting (4DGS). Existing camera-controlled video diffusion models synthesize plausible novel-view videos but fail to maintain consistency when scaled to the tens of target views required for 4DGS reconstruction. We identify this failure as a bounded-attention-context problem: when target views exceed the capacity of a single DiT forward pass, they must be split into groups, exposing two coupled bott

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Yudong Jin

    4DAnyone: Create Anyone in 4D from a Casual Monocular Video

  • LinkedLinked via arxiv author · 85%Pengtao Xie

    4DAnyone: Create Anyone in 4D from a Casual Monocular Video

  • LinkedLinked via arxiv author · 85%Qihang Zhang

    4DAnyone: Create Anyone in 4D from a Casual Monocular Video

  • LinkedLinked via arxiv author · 85%Zehong Shen

    4DAnyone: Create Anyone in 4D from a Casual Monocular Video

  • LinkedLinked via arxiv author · 85%Mingzhen Xu

    4DAnyone: Create Anyone in 4D from a Casual Monocular Video

  • LinkedLinked via arxiv author · 85%Yujun Shen

    4DAnyone: Create Anyone in 4D from a Casual Monocular Video

  • FuzzyOverlapping authors or contributors · 62%sgl-project/sglang

    Shared author/contributor keys: zhou

  • FuzzyOverlapping authors or contributors · 62%HKUDS/LightRAG

    Shared author/contributor keys: jin

authored (incoming)

Implements (incoming)

Related across the graph

Topics