Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

3D reconstruction and generation are commonly tackled by separate paradigms: pixel-based regression for reconstruction, and latent diffusion for generation. Recent works attempt to unify them in latent space, but with notable drawbacks: the diffusion objective is defined on latent features rather than the underlying 3D representation, and both branches suffer from information loss introduced by latent encoding, while requiring a pretrained Variational Autoencoder (VAE) or Representation Autoencoder (RAE). In this paper, we reformulate these two tasks under a unified pixel-space diffusion parad

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Sensen Gao

    PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

  • LinkedLinked via arxiv author · 85%Zhaoqing Wang

    PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

  • LinkedLinked via arxiv author · 85%Qihang Cao

    PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

  • LinkedLinked via arxiv author · 85%Dongdong Yu

    PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

  • LinkedLinked via arxiv author · 85%Changhu Wang

    PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

  • LinkedLinked via arxiv author · 85%Jia-Wang Bian

    PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

  • PossiblePossibly related (embedding) · 49%voxelmorph/voxelmorph
  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

authored (incoming)

Implements (incoming)

Related across the graph

Topics