A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
The paper introduces AgenticBBO-Bench to compare LLM agents on black-box optimization under one finite-budget protocol.
The benchmark spans synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design. In the authors’ experiments, agentic BBO beats direct LLM methods across all five domains and tops the best numerical optimizers in four. The study finds that extra numerical tools do not reliably help, task semantics usually do, and specific priors are less dependable. It also reports a five-task frontier challenge where GPT-6 Astra and DeepSeek-V4.1-Flash sit on the performance-cost Pareto frontier among seven evaluated LLMs. HF Daily Papers' note
The benchmark spans synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design. In the authors’ experiments, agentic BBO beats direct LLM methods across all five domains and tops the best numerical optimizers in four. The study finds that extra numerical tools do not reliably help, task semantics usually do, and specific priors are less dependable. It also reports a five-task frontier challenge where GPT-6 Astra and DeepSeek-V4.1-Flash sit on the performance-cost Pareto frontier among seven evaluated LLMs. HF Daily Papers' note
score 5