Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation

Tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video is a long-standing problem in computer vision. To tackle this problem, existing methods require inputs beyond the video itself-such as 3D models, depth maps, object masks, or task-specific learned features-and they struggle with textureless, transparent, reflective, or deformable surfaces. Here, we introduce ProxyPose, which recasts 6-DoF pose tracking as video-to-video translation. Given only a video and a single marked pixel in the first frame, a fine-tuned video diffusion model translates the input

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Ruihang Zhang

    ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation

  • LinkedLinked via arxiv author · 85%Felix Taubner

    ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation

  • LinkedLinked via arxiv author · 85%Pooja Ravi

    ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation

  • LinkedLinked via arxiv author · 85%Kiriakos N. Kutulakos

    ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation

  • LinkedLinked via arxiv author · 85%David B. Lindell

    ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation

authored (incoming)

Related across the graph

Topics