Megadose AI progress, ranked and analyzed.

OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

· HF Daily Papers ·
The paper frames audio-video chat as a native model task, then supplies synthetic data, a benchmark, and RL training for it.

OmniVChat assumes the user’s question lives in audio and video together, with no separate text prompt, captioner, or speech recognizer. The authors say evaluation is hard because good answers may depend on surroundings, expressions, and objects, not just keywords. Their OmniVChat-Studio generates single- and multi-turn audio-visual dialogues for training and testing. Training Qwen3-Omni-Instruct with their OmniVChat-RL improves results on both the synthetic benchmark and a human-recorded version. HF Daily Papers' note

score 6

Categories: Research