Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions, or compress geometry into latent spaces in order to leverage pre-trained latent diffusion models. In this work, we show that such architectural overhead and intricate loss formulations are unnecessary. We introduce a minimalist pixel-space Diffusion Transformer, built on a plain ViT, that operates directly on raw 3D point map patches and is conditioned on image tokens from a pre-trained DINOv3. Unlike existing latent diffusion approaches, we train our diffusion backbone entire

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 46%Diffuse-XL
  • LinkedLinked via arxiv author · 85%Haofei Xu

    PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

  • LinkedLinked via arxiv author · 85%Rundi Wu

    PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

  • LinkedLinked via arxiv author · 85%Philipp Henzler

    PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

  • LinkedLinked via arxiv author · 85%Nikolai Kalischek

    PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

  • LinkedLinked via arxiv author · 85%Michael Oechsle

    PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

  • LinkedLinked via arxiv author · 85%Fabian Manhardt

    PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

  • LinkedLinked via arxiv author · 85%Marc Pollefeys

    PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

Has model

authored (incoming)

Implements (incoming)

Related across the graph

Topics