CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
Six multimodal models scored low on CLBench-V, with the best overall result at 0.2847.
The benchmark tests whether models can use task-specific visual and textual context across grounding, applying new information, and learning new knowledge. It includes 3,443 instances across science, finance, long-document understanding, spatial reasoning, and web-based visual QA. InternVL3.5-30B-A3B led on grounding and new knowledge learning, while Qwen3.5-Plus led on new information application. HF Daily Papers' note
The benchmark tests whether models can use task-specific visual and textual context across grounding, applying new information, and learning new knowledge. It includes 3,443 instances across science, finance, long-document understanding, spatial reasoning, and web-based visual QA. InternVL3.5-30B-A3B led on grounding and new knowledge learning, while Qwen3.5-Plus led on new information application. HF Daily Papers' note
score 5