Megadose Built for builders and researchers.

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

· HF Daily Papers ·
OneStreamer uses generated, time-grounded captions as memory so a streaming video model can answer later without rereading old frames.

The paper pairs that memory system with training for proactive responses, so the model can wait until enough evidence appears and then act. Its PHCM component records local details and event summaries; PSTL trims repeated “waiting” supervision while keeping state-change signals. The authors also introduce OneStreamer-1M, a streaming video interaction dataset with more than one million records. Their 4B model leads the compared methods on eight streaming video understanding benchmarks, according to the paper. HF Daily Papers' note

score 5

Categories: Research