StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models
Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting their ability to retain past observations and develop precise spatial perception. In this paper, we propose StreamPI, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters. One core design is instruction-anchored temporal modeling. It treats each (visual observation, language instruction) pair as an atomic temp
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “StreamPI: Streaming Multimodal Temporal Modeling for Vision-” ≈ “VioletVision-3B””
- LinkedLinked via arxiv author · 85%Yizhe Liu →
“StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models”
- LinkedLinked via arxiv author · 85%Jinghua Hou →
“StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models”
- LinkedLinked via arxiv author · 85%Yuxiang Lu →
“StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models”
- LinkedLinked via arxiv author · 85%Zhenya Yang →
“StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models”
- LinkedLinked via arxiv author · 85%Xianzhe Fan →
“StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models”
- LinkedLinked via arxiv author · 85%Junwei Luo →
“StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models”
- LinkedLinked via arxiv author · 85%Junyi Li →
“StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models”
