Megadose AI progress, ranked and analyzed.

MIRROR: Learning from the Other View for Multi-Modal Reasoning

· ArXiv · AI/CL/LG ·
MIRROR trains a model to learn from whichever modality solves the same geometry problem best.

The paper argues that VLMs can answer from text and fail on the diagram, or do the reverse, even when the problem is equivalent. The authors build ODA-Data to pair text-dominant, image-dominant, and combined views of the same geometry tasks. MIRROR evaluates all views, treats the best-performing one as the teacher, and trains the others toward it with a reverse-KL objective. They report better accuracy and more consistent multimodal behavior than standard RL. ArXiv · AI/CL/LG's note

score 5

Categories: Research