Read original ↗
paperarXivTrust 82 · PrimaryPublished 5d agoLive · yesterday

AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction

Predicting robot videos requires both precise motion reasoning and preservation of high-frequency appearance, yet monolithic pixel models entangle these objectives and often conceal their progress behind a strong last-frame baseline. We present AcrossVAM1.0, a lightweight, text-assisted video action model that factorizes future prediction into object-centric motion and dense appearance. A frozen SAM3-DLP codec decomposes four context frames into semantic particles for the robot, arm, and gripper, together with a background latent. A 0.28M-parameter spatio-temporal Transformer aligns particle i

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses

    Fuzzy title match (0.73): “AcrossVAM1.0: Particle World Modeling for Text-Assisted Robo” ≈ “Developer-Y/cs-video-courses”

  • LinkedLinked via arxiv author · 85%Yafei Zhang

    AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction

  • LinkedLinked via arxiv author · 85%Nan Wu

    AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction

Implements (incoming)

authored (incoming)

Related across the graph

Topics