Megadose AI progress, ranked daily.

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

· HF Daily Papers ·
StreamPI gives single-frame robot VLA models temporal memory without adding parameters.

The framework treats each visual observation and instruction pair as a temporal unit, allowing fusion inside the pair and causal streaming across pairs. Its training randomizes frame intervals to better match asynchronous real-robot deployment. The authors say it inherits pretrained single-frame weights and can run in single- or multi-frame modes. In real-robot tasks and LIBERO simulation, StreamPI outperformed pi0.5. HF Daily Papers' note

score 5

Categories: Research