Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models
The paper proposes three base-model screens that predict whether coding-agent post-training is worth the cost.
The method replays successful post-trained agent trajectories and finds the first code-changing step that turns tests from failing to passing. It then scores base checkpoints on probability assigned to that decisive action, multiple-choice selection of the right patch, and prefix-conditioned pass@K continuations. Across ten public base/post-trained model pairs, all three screens closely matched post-trained SWE-bench Verified pass@1 rankings. ArXiv · AI/CL/LG's note
The method replays successful post-trained agent trajectories and finds the first code-changing step that turns tests from failing to passing. It then scores base checkpoints on probability assigned to that decisive action, multiple-choice selection of the right patch, and prefix-conditioned pass@K continuations. Across ten public base/post-trained model pairs, all three screens closely matched post-trained SWE-bench Verified pass@1 rankings. ArXiv · AI/CL/LG's note
score 5