HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
The benchmark is built to catch browsing agents that can search but cannot persist across languages, media, and weak clues.
HyperBrowseComp contains 423 human-validated questions in 13 languages, written by native or highly proficient speakers. Each question has a concise public answer, but finding it may require obscure evidence, multi-step clue chains, or sources like videos, scans, images, and maps. The authors filtered out easier items using models without internet access, aiming to reduce answers solvable from memorized knowledge alone. They evaluate models with native search and a shared retrieval setup, and compare against a human sample. Source: ArXiv · AI/CL/LG's note.
HyperBrowseComp contains 423 human-validated questions in 13 languages, written by native or highly proficient speakers. Each question has a concise public answer, but finding it may require obscure evidence, multi-step clue chains, or sources like videos, scans, images, and maps. The authors filtered out easier items using models without internet access, aiming to reduce answers solvable from memorized knowledge alone. They evaluate models with native search and a shared retrieval setup, and compare against a human sample. Source: ArXiv · AI/CL/LG's note.
score 6