Megadose AI progress, ranked and analyzed.

MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

· HF Daily Papers ·
The benchmark tests whether MLLMs know when to answer and when to refuse under shifted image-question context.

MMOOC includes more than 41,000 image-question pairs across answerable shifted in-context cases and unanswerable out-of-context cases. The paper says current multimodal models still have trouble balancing refusal with robust answering when context shifts. The authors evaluate with accuracy, refusal rate, and an LLM-as-judge metric for reasoning correctness. HF Daily Papers' note

score 5

Categories: Research