SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
SnapBench tests how well mobile AI retrieval survives messy photos and imperfect questions.
The benchmark pairs 1,145 queries with 9,085 gallery items across 53 controlled corruption conditions, with human annotations. The authors evaluate 16 multimodal retrievers, including dual-tower encoders and embedding-based VLMs. Image noise hurts retrieval substantially; text noise mostly damages text-only retrieval and has less effect on joint retrieval. Clean image-only retrieval often beats joint retrieval, which the paper links to coarse text dragging results down and weak fallback across modalities. HF Daily Papers' note
The benchmark pairs 1,145 queries with 9,085 gallery items across 53 controlled corruption conditions, with human annotations. The authors evaluate 16 multimodal retrievers, including dual-tower encoders and embedding-based VLMs. Image noise hurts retrieval substantially; text noise mostly damages text-only retrieval and has less effect on joint retrieval. Clean image-only retrieval often beats joint retrieval, which the paper links to coarse text dragging results down and weak fallback across modalities. HF Daily Papers' note
score 5