Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets. We argue this inherits more of the generative output interface than dense prediction requires: unlike RGB synthesis, dense prediction asks for pixel-correct, t

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%Tongyi-MAI/Z-Image-Turbo

    Fuzzy title match (0.73): “From RGB Generation to Dense Field Readout: Pixel-Space Dens” ≈ “Tongyi-MAI/Z-Image-Turbo”

  • LinkedLinked via arxiv author · 85%Zanyi Wang

    From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

  • LinkedLinked via arxiv author · 85%Xin Lin

    From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

  • LinkedLinked via arxiv author · 85%Haodong Li

    From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

  • LinkedLinked via arxiv author · 85%Dengyang Jiang

    From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

  • LinkedLinked via arxiv author · 85%Yijiang Li

    From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

  • LinkedLinked via arxiv author · 85%Pengtao Xie

    From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

  • FuzzyOverlapping authors or contributors · 62%affaan-m/ECC

    Shared author/contributor keys: jiang

Has model

authored (incoming)

Implements (incoming)

Related across the graph

Topics