OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs
OmniAssistBench tests whether omni-modal models can guide users through multi-turn video tasks, and current systems still miss key assistant behaviors.
The benchmark simulates continuous interactions by reverse-engineering internet videos into user goals and segmented turns. Models are given predefined priors so they can be judged against the same route through a task. Gemini-3-Pro scored 66.4/100, while Qwen3-Omni-Instruct scored 51.2. The paper says models often understand user input but give incomplete or wrong help, especially with gestures, conversation history, and waiting for the right event. HF Daily Papers' note
The benchmark simulates continuous interactions by reverse-engineering internet videos into user goals and segmented turns. Models are given predefined priors so they can be judged against the same route through a task. Gemini-3-Pro scored 66.4/100, while Qwen3-Omni-Instruct scored 51.2. The paper says models often understand user input but give incomplete or wrong help, especially with gestures, conversation history, and waiting for the right event. HF Daily Papers' note
score 5