Megadose Built for builders and researchers.

WOVEN: Weaving Visual World Modeling into Multimodal LLMs

· ArXiv · AI/CL/LG ·
WOVEN tests whether visual transition reasoning can be trained once and reused across multimodal tasks.

The paper introduces a 36,076-example training source and benchmark built around scenes, actions, and reasoning types. Its evaluation of 38 frontier MLLMs finds that even the strongest systems remain well below human performance on this kind of reasoning. Training on WOVEN improved 22 of 26 external benchmarks, with reported gains up to 27.3 percentage points. The authors argue the useful supervision is the reasoning operation being taught, not the specific scene or action shown. ArXiv · AI/CL/LG's note

score 5

Categories: Research