Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection

Vision-Language-Action (VLA) models have shown remarkable promise in generalized robotic manipulation. However, their spatial generalization remains fragile. We argue that simply increasing the number of viewpoints is insufficient. Models often fall into the trap of Shortcut Learning, latching onto spurious correlations (e.g., fixed relative poses between objects or between the camera and robot base) rather than learning true spatial relationships. In this work, we propose a data-centric solution to enhance VLA spatial generalization. We utilize a dual-arm setup where one arm performs manipula

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 51%vlm-starter
  • LinkedLinked via arxiv author · 85%Jincheng Tang

    The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection

  • LinkedLinked via arxiv author · 85%Yilong Zhu

    The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection

  • LinkedLinked via arxiv author · 85%Zhengyuan Xie

    The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection

  • LinkedLinked via arxiv author · 85%Jiang-Jiang Liu

    The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection

  • LinkedLinked via arxiv author · 85%Jiaxing Zhang

    The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection

  • PossiblePossibly related (embedding) · 47%sou350121/VLA-Handbook
  • PossiblePossibly related (embedding) · 51%lucidrains/mimic-video

Implements

authored (incoming)

Implements (incoming)

Related across the graph

Topics