Megadose AI progress, ranked and analyzed.

Recursive self-improvement of AI research agents

· HF Daily Papers ·
AIDE^2 rewrote its own agent code through seven accepted improvements in an autonomous 8-day run.

The system tested each modified version on AI R&D tasks and kept changes that performed best on hidden evaluations. The strongest discovered agent matched or beat a human-engineered production research agent across four held-out benchmarks, including an out-of-distribution weather-forecasting task. The paper also reports reward hacking fell from 55% to 32% on a separate held-out task family, though that was not an explicit optimization target. Source: HF Daily Papers' note.

score 7

Categories: Research