newsReddit r/MachineLearningTrust 52 · CommunityPublished 21d agoLive · 21d ago
Why first person video may matter for robot learning[D]
I can see why first-person video might help a robot model, but not because the robot can copy a human hand. The joints, reach, timing, and control space are all different. What may transfer is the sequence of visual attention: which object enters view, what changes before contact, and where the actor looks next. LingBot-VLA 2.0 (arXiv:2607.06403) uses first-person data alongside robot trajectories. A useful ablation separates visual prediction from robot
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 59%Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues? →
- PossiblePossibly related (embedding) · 58%Native Video-Action Pretraining for Generalizable Robot Control →
- PossiblePossibly related (embedding) · 57%Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation →
- PossiblePossibly related (embedding) · 55%The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction →
- PossiblePossibly related (embedding) · 54%Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots →
- PossiblePossibly related (embedding) · 48%ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models →
Covers
paperDo Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?paperNative Video-Action Pretraining for Generalizable Robot ControlpaperHuman-Centric Transferable Tactile Pre-Training for Dexterous Robotic ManipulationpaperThe Surprising Effectiveness of Video Diffusion Models for Hand Motion ReconstructionpaperTranslation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots
Covers (incoming)
Related across the graph
paperContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World ModelspaperNative Video-Action Pretraining for Generalizable Robot ControlpaperTranslation as a Bridging Action: Transferring Manipulation Skills from Humans to RobotspaperHuman-Centric Transferable Tactile Pre-Training for Dexterous Robotic ManipulationpaperDo Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?paperThe Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction
