StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models
StreamPI gives single-frame robot VLA models temporal memory without adding parameters.
The framework treats each visual observation and instruction pair as a temporal unit, allowing fusion inside the pair and causal streaming across pairs. Its training randomizes frame intervals to better match asynchronous real-robot deployment. The authors say it inherits pretrained single-frame weights and can run in single- or multi-frame modes. In real-robot tasks and LIBERO simulation, StreamPI outperformed pi0.5. HF Daily Papers' note
The framework treats each visual observation and instruction pair as a temporal unit, allowing fusion inside the pair and causal streaming across pairs. Its training randomizes frame intervals to better match asynchronous real-robot deployment. The authors say it inherits pretrained single-frame weights and can run in single- or multi-frame modes. In real-robot tasks and LIBERO simulation, StreamPI outperformed pi0.5. HF Daily Papers' note
score 5