CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents. CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor. However, cross-embodiment learning is no
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video →
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses →
“Fuzzy title match (0.73): “CLAP: Cross-Embodiment Video World Models are Zero-Shot Phys” ≈ “Developer-Y/cs-video-courses””
- LinkedLinked via arxiv author · 85%Kechen Liu →
“CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators”
- LinkedLinked via arxiv author · 85%Ola Shorinwa →
“CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators”
