Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement
Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow matching-based VLA models have shown remarkable success due to their capability to generate precise and smooth action sequences and capture multimodal distributions. However, the iterative denoising process in the action head acts as a major computational bottleneck, posing a critical challenge for real-time deployment. To address this challenge, we propose ActionCache, a plug-and-play external cache that opportunistically reuses past intermediate actions to war
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%lucidrains/mimic-video →
- PossiblePossibly related (embedding) · 53%xlang-ai/OSWorld →
- LinkedLinked via arxiv author · 85%Ryuji Oi →
“Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement”
- LinkedLinked via arxiv author · 85%Hikari Otsuka →
“Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement”
- LinkedLinked via arxiv author · 85%Kosuke Matsushima →
“Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement”
- LinkedLinked via arxiv author · 85%Yuki Ichikawa →
“Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement”
- LinkedLinked via arxiv author · 85%Masato Motomura →
“Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement”
- LinkedLinked via arxiv author · 85%Tatsuya Kaneko →
“Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement”
