DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm $\mathrm{SE}(3)$ transformations into attention via \textbf{PRoPE-style geometric encoding}, preserving a
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 84%liguodongiot/llm-action →
“Fuzzy title match (0.92): “DreamX-Phi 1.0: Action-Conditioned Video World Model for Rob” ≈ “liguodongiot/llm-action””
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%marimo-team/marimo →
“Shared author/contributor keys: team”
- FuzzyOverlapping authors or contributors · 62%recommenders-team/recommenders →
“Shared author/contributor keys: team”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%DreamX Team →
“DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation”
- LinkedLinked via arxiv author · 85%Kerui Chen →
“DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation”
- LinkedLinked via arxiv author · 85%Xiangxiang Chu →
“DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation”
