MOSS-VL Technical Report
MOSS-VL is built to keep watching video while it is already speaking.
The report describes an open vision-language model family designed for real-time interaction, with gated cross-attention letting the decoder attend to incoming visual frames during generation. Its training includes a synthesized interaction corpus for when to speak, stay silent, or revise. The authors say MOSS-VL-Realtime leads three of four open-source streaming benchmark averages and scores 66.0 on OmniMMI Proactive Alerting versus 37.5 for the best baseline. They also release five checkpoints, the training curriculum, and real-time inference code. HF Daily Papers' note
The report describes an open vision-language model family designed for real-time interaction, with gated cross-attention letting the decoder attend to incoming visual frames during generation. Its training includes a synthesized interaction corpus for when to speak, stay silent, or revise. The authors say MOSS-VL-Realtime leads three of four open-source streaming benchmark averages and scores 66.0 on OmniMMI Proactive Alerting versus 37.5 for the best baseline. They also release five checkpoints, the training curriculum, and real-time inference code. HF Daily Papers' note
score 7