Megadose Built for builders and researchers.

Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?

· ArXiv · AI/CL/LG ·
The benchmark tests whether agents can turn one-off LLM skill into cheaper reusable systems.

BOTTLED gives agents an unlabeled workload plus fixed time, compute, and API budgets, then lets them choose tactics such as training a small model or writing a reusable program. The paper finds zero-shot strength is a poor predictor: 48 of 60 bottling runs fell below their model’s zero-shot confidence interval, and 31 lost to a small-model distillation baseline on the same token budget. One strong result still shows the upside: Opus 5 kept about 82% of its zero-shot macro-F1 on query-product relevance at roughly 657x lower reported cost. ArXiv · AI/CL/LG's note

score 6

Categories: Research