Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models
The paper introduces RuleMaze, a benchmark for testing whether multimodal models can navigate visual mazes while obeying explicit natural-language rules.
The authors say the task is meant to isolate perception, rule interpretation, and constrained action planning. They also propose a rule-generation pipeline that turns natural-language rules into logic and executable validators. Their Disentangled Multimodal Planning method separates perception, execution, and rule checking, producing intermediate traces. In experiments, it improves rule compliance and planning success over end-to-end textual planning baselines. ArXiv · AI/CL/LG's note
The authors say the task is meant to isolate perception, rule interpretation, and constrained action planning. They also propose a rule-generation pipeline that turns natural-language rules into logic and executable validators. Their Disentangled Multimodal Planning method separates perception, execution, and rule checking, producing intermediate traces. In experiments, it improves rule compliance and planning success over end-to-end textual planning baselines. ArXiv · AI/CL/LG's note
score 5