Megadose AI progress, ranked and analyzed.

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

· ArXiv · AI/CL/LG ·
The benchmark finds LLM agents can discover better training-data strategies, but often lose those gains before the run ends.

RSIBench-Data fixes the post-training stack so agents are judged on data-centric research decisions, not surrounding systems work. In tests of four frontier agents across six benchmarks, agents improved on their first valid attempt in 58.33% of settings. But when searches continued after the best observed score, 78.26% finished with a worse final attempt. The authors identify stronger runs as those with accurate hypotheses, validation-grounded supervision, behavior-aligned data, and preserved strong checkpoints. ArXiv · AI/CL/LG's note

score 6

Categories: Research