Masked Visual Actions for Unified World Modeling
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model that pre
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Hadi Alzayer →
“Masked Visual Actions for Unified World Modeling”
- LinkedLinked via arxiv author · 85%Wenlong Huang →
“Masked Visual Actions for Unified World Modeling”
- LinkedLinked via arxiv author · 85%Haonan Chen →
“Masked Visual Actions for Unified World Modeling”
- LinkedLinked via arxiv author · 85%Christopher Luey →
“Masked Visual Actions for Unified World Modeling”
- LinkedLinked via arxiv author · 85%Lvmin Zhang →
“Masked Visual Actions for Unified World Modeling”
- LinkedLinked via arxiv author · 85%Maneesh Agrawala →
“Masked Visual Actions for Unified World Modeling”
- LinkedLinked via arxiv author · 85%Gordon Wetzstein →
“Masked Visual Actions for Unified World Modeling”
- LinkedLinked via arxiv author · 85%Li Fei-Fei →
“Masked Visual Actions for Unified World Modeling”
