What is Missing from AI Post-Training AI: An Empirical Analysis
The paper says post-training agents can execute a plan, but do not rethink the plan once evidence comes in.
The authors separate execution-level iteration from strategy-level revision. In their corpus of post-training runs, agents chose a training strategy at the start and then spent the remaining budget making local adjustments inside it. Scaffolds based on experience improved task scores, and human guidance could redirect the initial choice, but neither produced spontaneous mid-run strategy changes. Extra inference compute helped on easier tasks and barely moved the hardest one. ArXiv · AI/CL/LG's note
The authors separate execution-level iteration from strategy-level revision. In their corpus of post-training runs, agents chose a training strategy at the start and then spent the remaining budget making local adjustments inside it. Scaffolds based on experience improved task scores, and human guidance could redirect the initial choice, but neither produced spontaneous mid-run strategy changes. Extra inference compute helped on easier tasks and barely moved the hardest one. ArXiv · AI/CL/LG's note
score 5