repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 4h ago
sou350121/VLA-Handbook
本项目旨在为致力于进入VLA(Vision-Language-Action)领域的算法工程师提供一份全中文、实战导向的学习/面试手册。 不同于通用的 CV/NLP 面试指南,本项目聚焦于 Robotics 特有的挑战
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%VioletVision-3B →
- PossiblePossibly related (embedding) · 49%CoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigation →
- PossiblePossibly related (embedding) · 49%SurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics →
- PossiblePossibly related (embedding) · 47%The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection →
- PossiblePossibly related (embedding) · 45%Visual Language Models Train Robots to Read Human Emotions →
- PossiblePossibly related (embedding) · 46%Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning →
- PossiblePossibly related (embedding) · 51%From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model →
- PossiblePossibly related (embedding) · 45%I trained a vision-language model to play Snake, and so can you. [P] →
Related to
Implements
Covers
Implements (incoming)
Covers (incoming)
Related across the graph
paperThe Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data CollectionnewsAI Creates Virtual Language Using Color and Gestures - 조선일보newsI trained a vision-language model to play Snake, and so can you. [P]paperDo Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied ReasoningpaperSurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical RoboticspaperCoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned NavigationpaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelnewsVisual Language Models Train Robots to Read Human EmotionsmodelVioletVision-3B
