FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory
State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal memory. We introduce Fast and EffectIVE VLA (FIVE-VLA) to address these through two key contributions. First, we employ an efficient vision encoder that processes high-resolution ($448 \times 896$) images while generating only 98 tokens, over $5\times$ fewer than existing approaches, and bypass text generation entirely for single-pass trajectory prediction. Second, we propose Recurrent Action Memory
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Kemal Oksuz →
“FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory”
- LinkedLinked via arxiv author · 85%Alexandru Buburuzan →
“FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory”
- LinkedLinked via arxiv author · 85%Yuhan Yao →
“FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory”
- LinkedLinked via arxiv author · 85%Puneet K. Dokania →
“FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory”
- FuzzySimilar title/name (fuzzy) · 84%liguodongiot/llm-action →
“Fuzzy title match (0.92): “FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurre” ≈ “liguodongiot/llm-action””
