HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
The benchmark is built to make browsing agents prove they can find obscure, verifiable evidence across languages and media.
HyperBrowseComp contains 423 manually written and human-validated questions in 13 languages. The questions require multi-step web investigation, including videos, scanned documents, images, and maps. The authors filtered easier items by testing against models without internet access, aiming to avoid answers recoverable from memorized knowledge alone. They also compare model performance under provider-native search and a shared retrieval setup, with a human evaluation for context. HF Daily Papers' note
HyperBrowseComp contains 423 manually written and human-validated questions in 13 languages. The questions require multi-step web investigation, including videos, scanned documents, images, and maps. The authors filtered easier items by testing against models without internet access, aiming to avoid answers recoverable from memorized knowledge alone. They also compare model performance under provider-native search and a shared retrieval setup, with a human evaluation for context. HF Daily Papers' note
score 5