MindTopo: Can Foundation Models Reason in Topological Space?
MindTopo tests whether models can track spatial relations that survive deformation, and finds planning is the weak point.
The benchmark covers continuity, separation, order, enclosure, and knots across 11,030 procedurally generated instances. Fourteen multimodal large language models do better at identifying or inferring relations than at acting as closed-loop agents. The best model remains well below observed human performance. Fine-tuning and reinforcement learning helped Qwen3-VL-2B-Instruct more on reasoning than planning, while generated rollouts often failed to preserve topology across transitions. ArXiv · AI/CL/LG's note
The benchmark covers continuity, separation, order, enclosure, and knots across 11,030 procedurally generated instances. Fourteen multimodal large language models do better at identifying or inferring relations than at acting as closed-loop agents. The best model remains well below observed human performance. Fine-tuning and reinforcement learning helped Qwen3-VL-2B-Instruct more on reasoning than planning, while generated rollouts often failed to preserve topology across transitions. ArXiv · AI/CL/LG's note
score 5