Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
The new C4 benchmark finds even the strongest tested MLLMs recover creatively encoded meanings only about half the time.
The paper frames cross-concept understanding as the ability to infer intended meaning from non-obvious conceptual links, using Chinese idiom figures as the test bed. C4-Eval includes 184 synthetic items and 37 human-created figures, expanded into 884 answer-recovery cases across five task settings. Candidate constraints helped sharply, while bridge hints and explanation prompts produced only modest gains. HF Daily Papers' note
The paper frames cross-concept understanding as the ability to infer intended meaning from non-obvious conceptual links, using Chinese idiom figures as the test bed. C4-Eval includes 184 synthetic items and 37 human-created figures, expanded into 884 answer-recovery cases across five task settings. Candidate constraints helped sharply, while bridge hints and explanation prompts produced only modest gains. HF Daily Papers' note
score 4