Megadose Built for builders and researchers.

MOSS-VL Technical Report

· HF Daily Papers ·
MOSS-VL is built to keep watching video while it is already speaking.

The report describes an open vision-language model family designed for real-time interaction, with gated cross-attention letting the decoder attend to incoming visual frames during generation. Its training includes a synthesized interaction corpus for when to speak, stay silent, or revise. The authors say MOSS-VL-Realtime leads three of four open-source streaming benchmark averages and scores 66.0 on OmniMMI Proactive Alerting versus 37.5 for the best baseline. They also release five checkpoints, the training curriculum, and real-time inference code. HF Daily Papers' note

score 7

Categories: Model Releases, OSS & Tools