Read original ↗
paperarXivTrust 82 · PrimaryPublished 6d agoLive · 4d ago

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting their ability to retain past observations and develop precise spatial perception. In this paper, we propose StreamPI, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters. One core design is instruction-anchored temporal modeling. It treats each (visual observation, language instruction) pair as an atomic temp

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B

    Fuzzy title match (0.73): “StreamPI: Streaming Multimodal Temporal Modeling for Vision-” ≈ “VioletVision-3B”

  • LinkedLinked via arxiv author · 85%Yizhe Liu

    StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

  • LinkedLinked via arxiv author · 85%Jinghua Hou

    StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

  • LinkedLinked via arxiv author · 85%Yuxiang Lu

    StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

  • LinkedLinked via arxiv author · 85%Zhenya Yang

    StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

  • LinkedLinked via arxiv author · 85%Xianzhe Fan

    StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

  • LinkedLinked via arxiv author · 85%Junwei Luo

    StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

  • LinkedLinked via arxiv author · 85%Junyi Li

    StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

Has model

authored (incoming)

Covers (incoming)

Implements (incoming)

Related across the graph

Topics