AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction
Predicting robot videos requires both precise motion reasoning and preservation of high-frequency appearance, yet monolithic pixel models entangle these objectives and often conceal their progress behind a strong last-frame baseline. We present AcrossVAM1.0, a lightweight, text-assisted video action model that factorizes future prediction into object-centric motion and dense appearance. A frozen SAM3-DLP codec decomposes four context frames into semantic particles for the robot, arm, and gripper, together with a background latent. A 0.28M-parameter spatio-temporal Transformer aligns particle i
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses →
“Fuzzy title match (0.73): “AcrossVAM1.0: Particle World Modeling for Text-Assisted Robo” ≈ “Developer-Y/cs-video-courses””
- LinkedLinked via arxiv author · 85%Yafei Zhang →
“AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction”
- LinkedLinked via arxiv author · 85%Nan Wu →
“AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction”
