Megadose Built for builders and researchers.

BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

· ArXiv · AI/CL/LG ·
The paper introduces a 1,000-question English-and-Russian test for multimodal models on dense business and scientific documents.

BEAR-Bench is designed to evaluate reasoning over professional, text-heavy materials without requiring outside domain knowledge. The authors say existing benchmarks lean toward extraction tasks, broad settings, or English and Chinese coverage. They tested 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, and found room for improvement even among the strongest systems. They also used the model outputs to compare hallucination detection methods. ArXiv · AI/CL/LG's note

score 4

Categories: Research