Native Video-Action Pretraining for Generalizable Robot Control
The advent of video-action models offers a promising path for robot control. Nevertheless, we argue that repurposing video generative models designed for digital content creation is inherently inadequate for physical environments. To bridge this gap, we present LingBot-VA 2.0, a video-action foundation model built from the ground up for embodiment. Four core design principles showcase its evolution from LingBot-VA. (1) Departing from traditional reconstruction-focused VAEs, we introduce a semantic visual-action tokenizer, which aligns visual representations with both semantics and actions, imp
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%lucidrains/mimic-video →
- FuzzySimilar title/name (fuzzy) · 87%lllyasviel/ControlNet →
“Fuzzy title match (0.94): “Native Video-Action Pretraining for Generalizable Robot Cont” ≈ “lllyasviel/ControlNet””
- FuzzySimilar title/name (fuzzy) · 87%lllyasviel/ControlNet-v1-1 →
“Fuzzy title match (0.94): “Native Video-Action Pretraining for Generalizable Robot Cont” ≈ “lllyasviel/ControlNet-v1-1””
- PossiblePossibly related (embedding) · 58%Flux 3 X Mimic: The Next Generation of Video-Action Models →
- PossiblePossibly related (embedding) · 58%Why first person video may matter for robot learning[D] →
- PossiblePossibly related (embedding) · 48%Lingbot-map: A 3D foundation model for reconstructing scenes from streaming data →
- FuzzySimilar title/name (fuzzy) · 87%google-deepmind/dm_control →
“Fuzzy title match (0.94): “Native Video-Action Pretraining for Generalizable Robot Cont” ≈ “google-deepmind/dm_control””
- FuzzyOverlapping authors or contributors · 78%sgl-project/sglang →
“Shared author/contributor keys: luo, zhou”
