Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision

Cross-embodiment transfer in vision-language-action (VLA) models remains challenging because low-level state and action spaces differ fundamentally across robot platforms. We observe that the high-level cognitive process underlying manipulation, including scene perception, object identification, task planning, and sub-task decomposition, is largely shared across embodiments. Based on this observation, we present ZR-0, a 2.6 billion parameter end-to-end VLA model that uses dense Embodied Chain-of-Thought (ECoT) supervision to align cross-embodiment representations within the vision-language mod

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via unknownvlm-starter
  • FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B

    Fuzzy title match (0.73): “Training Vision-Language-Action Models with Dense Embodied C” ≈ “VioletVision-3B”

  • FuzzySimilar title/name (fuzzy) · 84%liguodongiot/llm-action

    Fuzzy title match (0.92): “Training Vision-Language-Action Models with Dense Embodied C” ≈ “liguodongiot/llm-action”

  • FuzzySimilar title/name (fuzzy) · 84%pytorch/vision

    Fuzzy title match (0.92): “Training Vision-Language-Action Models with Dense Embodied C” ≈ “pytorch/vision”

  • FuzzySimilar title/name (fuzzy) · 84%roboflow/supervision

    Fuzzy title match (0.92): “Training Vision-Language-Action Models with Dense Embodied C” ≈ “roboflow/supervision”

Implements

Has model

Covers (incoming)

Implements (incoming)

Related across the graph

Topics