repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · yesterday
lucidrains/mimic-video
Implementation of Mimic-Video, Video-Action Models for SOTA Generalizable Robot Control Beyond VLAs
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models →
- PossiblePossibly related (embedding) · 53%FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model →
- PossiblePossibly related (embedding) · 51%The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection →
- PossiblePossibly related (embedding) · 48%PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation →
- PossiblePossibly related (embedding) · 48%SurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics →
- PossiblePossibly related (embedding) · 54%Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning →
- PossiblePossibly related (embedding) · 51%From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model →
- PossiblePossibly related (embedding) · 53%Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement →
Implements
paperZ-1: Efficient Reinforcement Learning for Vision-Language-Action ModelspaperFurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action ModelpaperThe Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data CollectionpaperPhysisForcing: Physics Reinforced World Simulator for Robotic ManipulationpaperSurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics
Implements (incoming)
paperSimple-to-Complex Structured Demonstrations for Vision-Language-Action LearningpaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperTraining-Free Acceleration for Vision-Language-Action Models with Action Caching and RefinementpaperLift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware ManipulationpaperNative Video-Action Pretraining for Generalizable Robot Control
Covers (incoming)
Related across the graph
paperThe Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data CollectionnewsFlux 3 X Mimic: The Next Generation of Video-Action ModelspaperNative Video-Action Pretraining for Generalizable Robot ControlpaperSimple-to-Complex Structured Demonstrations for Vision-Language-Action LearningpaperLift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware ManipulationpaperPhysisForcing: Physics Reinforced World Simulator for Robotic ManipulationpaperSurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical RoboticspaperFrom Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelpaperTraining-Free Acceleration for Vision-Language-Action Models with Action Caching and RefinementpaperZ-1: Efficient Reinforcement Learning for Vision-Language-Action ModelspaperFurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model
