Megadose Built for builders and researchers.

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

· ArXiv · AI/CL/LG ·
The benchmark finds current LLM agents still weak at changing the training algorithms that would matter for recursive self-improvement.

AI4AI-Bench tests agents on 10 frozen research repos, giving each 4 hours on one B300 to rewrite the training algorithm before rerunning and scoring it against a hidden evaluator. Across 29 configurations from six systems, the mean score was 0.166, with the best system at 0.250 on a scale where 0.1 is the original repo algorithm and 1.0 is the task optimum. Most submissions did not alter how the model learns; the ones that did scored higher on average. More reasoning effort mainly increased the chance that agents attempted those learning-rule changes. ArXiv · AI/CL/LG's note

score 6

Categories: Research