Read original ↗
paperarXivTrust 82 · PrimaryPublished 7d agoLive · 6d ago

FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal memory. We introduce Fast and EffectIVE VLA (FIVE-VLA) to address these through two key contributions. First, we employ an efficient vision encoder that processes high-resolution ($448 \times 896$) images while generating only 98 tokens, over $5\times$ fewer than existing approaches, and bypass text generation entirely for single-pass trajectory prediction. Second, we propose Recurrent Action Memory

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Kemal Oksuz

    FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

  • LinkedLinked via arxiv author · 85%Alexandru Buburuzan

    FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

  • LinkedLinked via arxiv author · 85%Yuhan Yao

    FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

  • LinkedLinked via arxiv author · 85%Puneet K. Dokania

    FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

  • FuzzySimilar title/name (fuzzy) · 84%liguodongiot/llm-action

    Fuzzy title match (0.92): “FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurre” ≈ “liguodongiot/llm-action”

authored (incoming)

Implements (incoming)

Related across the graph

Topics