ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation
Tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video is a long-standing problem in computer vision. To tackle this problem, existing methods require inputs beyond the video itself-such as 3D models, depth maps, object masks, or task-specific learned features-and they struggle with textureless, transparent, reflective, or deformable surfaces. Here, we introduce ProxyPose, which recasts 6-DoF pose tracking as video-to-video translation. Given only a video and a single marked pixel in the first frame, a fine-tuned video diffusion model translates the input
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Ruihang Zhang →
“ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation”
- LinkedLinked via arxiv author · 85%Felix Taubner →
“ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation”
- LinkedLinked via arxiv author · 85%Pooja Ravi →
“ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation”
- LinkedLinked via arxiv author · 85%Kiriakos N. Kutulakos →
“ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation”
- LinkedLinked via arxiv author · 85%David B. Lindell →
“ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation”
