Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
The paper claims agent runs improve when a separate controller spends inference budget deciding what work to do next.
The proposed harness splits execution between workers and a meta-reasoning controller that tracks progress, weighs options, and dispatches new work from compact memory. On ProgramBench, it reports 71.5% with GPT-5.5 versus 58.0% for Codex, and 67.2% with Opus 4.8 versus 65.5% for Claude Code. Across other long-horizon reasoning and proof tasks, it gains 3.6 to 4.2 points over direct control. The authors note the overhead can hurt at small budgets, but say the method keeps improving where direct control plateaus. ArXiv · AI/CL/LG's note
The proposed harness splits execution between workers and a meta-reasoning controller that tracks progress, weighs options, and dispatches new work from compact memory. On ProgramBench, it reports 71.5% with GPT-5.5 versus 58.0% for Codex, and 67.2% with Opus 4.8 versus 65.5% for Claude Code. Across other long-horizon reasoning and proof tasks, it gains 3.6 to 4.2 points over direct control. The authors note the overhead can hurt at small budgets, but say the method keeps improving where direct control plateaus. ArXiv · AI/CL/LG's note
score 5