Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs
The paper tests whether hybrid-thinking MLLMs keep clean final-answer behavior when switching out of deliberative mode.
The authors introduce PatternEval, a 2,415-prompt multimodal benchmark targeting failures beyond accuracy: chain-of-thought leakage, repetition, contradiction, and performative reasoning. They report that these response-pattern failures appear across providers and are substantially worse in non-thinking inference. Their proposed PatternRM and PatternRL add pattern-specific penalties during reinforcement learning, reducing cross-mode misalignment on Qwen3-VL-4B and 8B with only a marginal task-performance trade-off. HF Daily Papers' note
The authors introduce PatternEval, a 2,415-prompt multimodal benchmark targeting failures beyond accuracy: chain-of-thought leakage, repetition, contradiction, and performative reasoning. They report that these response-pattern failures appear across providers and are substantially worse in non-thinking inference. Their proposed PatternRM and PatternRL add pattern-specific penalties during reinforcement learning, reducing cross-mode misalignment on Qwen3-VL-4B and 8B with only a marginal task-performance trade-off. HF Daily Papers' note
score 4