Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs
Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions, and actions that are costly to collect at scale. We argue that this bottleneck stems from conflating two distinct learning objectives: acquiring physical competence (how to move) and acquiring semantic alignment (what to do). Crucially, only the latter requires language supervision. Building on this Decomposition Hypothesis, we propose Task-Agnostic Pretraining (TAP), a two-stage framework that first learns transferable motor priors from cheap,
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%vlm-starter →
- LinkedLinked via arxiv author · 85%Junhao Shi →
“Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs”
- LinkedLinked via arxiv author · 85%Siyin Wang →
“Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs”
- LinkedLinked via arxiv author · 85%Xiaopeng Yu →
“Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs”
- LinkedLinked via arxiv author · 85%Li Jin →
“Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs”
- LinkedLinked via arxiv author · 85%Jingjing Gong →
“Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs”
- LinkedLinked via arxiv author · 85%Xipeng Qiu →
“Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs”
