Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution
Pattern-specific executors sharply reduced the gap between an agent’s stated plan and what it actually did.
The paper tests a “Planning-as-Routing” setup where an LLM declares one of four planning modes, then a deterministic router sends the task to a matching executor. Generic Plan+ReAct preserved the declared planning structure in only 22% to 45% of trajectories across three benchmarks. The routed executors improved task success from 0.48 to 0.92 on ALFWorld and from 0.36 to 0.44 on SWE-bench Verified. The authors still find that current LLMs do not reliably choose the best planning mode for a given task. ArXiv · AI/CL/LG's note
The paper tests a “Planning-as-Routing” setup where an LLM declares one of four planning modes, then a deterministic router sends the task to a matching executor. Generic Plan+ReAct preserved the declared planning structure in only 22% to 45% of trajectories across three benchmarks. The routed executors improved task success from 0.48 to 0.92 on ALFWorld and from 0.36 to 0.44 on SWE-bench Verified. The authors still find that current LLMs do not reliably choose the best planning mode for a given task. ArXiv · AI/CL/LG's note
score 4