AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers
AgentHPOBench tests whether LLM agents can improve experiments step by step, not just produce code or answers.
The benchmark includes 30 executable machine-learning tasks across seven research categories. Agents start from a validated baseline, then make sequential hyperparameter changes after seeing prior configs, metrics, and logs. The authors evaluated 12 common agents alongside conventional HPO baselines, finding real optimization ability but weak sustained refinement and log diagnosis. ArXiv · AI/CL/LG's note
The benchmark includes 30 executable machine-learning tasks across seven research categories. Agents start from a validated baseline, then make sequential hyperparameter changes after seeing prior configs, metrics, and logs. The authors evaluated 12 common agents alongside conventional HPO baselines, finding real optimization ability but weak sustained refinement and log diagnosis. ArXiv · AI/CL/LG's note
score 5