Megadose Built for builders and researchers.

OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning

· HF Daily Papers ·
The paper introduces a benchmark and training pipeline for cases where audio and visual evidence both have to be used.

OmniReasoningBench has 1,150 multiple-choice and open-ended questions across video reasoning and reasoning beyond video. The authors also describe OmniQA, a data engine that builds evidence-grounded QA pairs with time-stamped clue chains, producing SFT and RL training sets. Their MFSD method assigns credit by testing responses against modality-specific clues and cross-modal interactions. The resulting OmniReasoning-30B-A3B reports 50.0% on OmniVideoBench and 42.5% on OmniReasoningBench, beating the Qwen3-Omni base model by 12.8 and 9.3 points. HF Daily Papers' note

score 6

Categories: Research