TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development
TraceML turns the ML-agent gap into a version-by-version planning record.
The paper compares 4,465 human Kaggle trajectories with paired agent runs on seven shared competitions. Each code version is labeled by action, intent, edit size, score effect, timestamp, and score. The authors find humans cycle among data work, validation, model changes, and ensembling, while the tested agents fall into narrower loops and rarely return to abandoned ideas. A human-derived planning prompt improves some behaviors and scores, but leaves the overall effort profile recognizably agent-like. ArXiv · AI/CL/LG's note
The paper compares 4,465 human Kaggle trajectories with paired agent runs on seven shared competitions. Each code version is labeled by action, intent, edit size, score effect, timestamp, and score. The authors find humans cycle among data work, validation, model changes, and ensembling, while the tested agents fall into narrower loops and rarely return to abandoned ideas. A human-derived planning prompt improves some behaviors and scores, but leaves the overall effort profile recognizably agent-like. ArXiv · AI/CL/LG's note
score 5