MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
The benchmark tests whether MLLMs know when to answer and when to refuse under shifted image-question context.
MMOOC includes more than 41,000 image-question pairs across answerable shifted in-context cases and unanswerable out-of-context cases. The paper says current multimodal models still have trouble balancing refusal with robust answering when context shifts. The authors evaluate with accuracy, refusal rate, and an LLM-as-judge metric for reasoning correctness. HF Daily Papers' note
MMOOC includes more than 41,000 image-question pairs across answerable shifted in-context cases and unanswerable out-of-context cases. The paper says current multimodal models still have trouble balancing refusal with robust answering when context shifts. The authors evaluate with accuracy, refusal rate, and an LLM-as-judge metric for reasoning correctness. HF Daily Papers' note
score 5