Read original ↗
paperarXivTrust 82 · PrimaryPublished 28d agoLive · 25d ago

Unified Video Dense Prediction from Disjoint Data

Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified systems circumvent this by restricting training to fully co-annotated data, or by incurring the large computational cost of pseudo-labeling. To mitigate this, we introduce UniD, a unified video model that jointly predicts eight dense scene properties-depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials-all learned from disjoint, d

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Yihong Sun

    Unified Video Dense Prediction from Disjoint Data

  • LinkedLinked via arxiv author · 85%Seoung Wug Oh

    Unified Video Dense Prediction from Disjoint Data

  • LinkedLinked via arxiv author · 85%Jiahui Huang

    Unified Video Dense Prediction from Disjoint Data

  • LinkedLinked via arxiv author · 85%Bharath Hariharan

    Unified Video Dense Prediction from Disjoint Data

  • LinkedLinked via arxiv author · 85%Joon-Young Lee

    Unified Video Dense Prediction from Disjoint Data

authored (incoming)

Related across the graph

Topics