Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, generating long, multi-view consistent videos of dynamic scenes remains unsolved. In this work, we present MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. Our key insight is that an autoregressive 3D reconstruction model naturally interfaces between autoregressively gen

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Gal Fiebelman

    MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

  • LinkedLinked via arxiv author · 85%Hadar Averbuch-Elor

    MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

  • LinkedLinked via arxiv author · 85%Sagie Benaim

    MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

  • FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses

    Fuzzy title match (0.73): “MV-Forcing: Long Multi-View Video Generation via 4D-Grounded” ≈ “Developer-Y/cs-video-courses”

authored (incoming)

Implements (incoming)

Related across the graph

Topics