RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement
The benchmark finds LLM agents can discover better training-data strategies, but often lose those gains before the run ends.
RSIBench-Data fixes the post-training stack so agents are judged on data-centric research decisions, not surrounding systems work. In tests of four frontier agents across six benchmarks, agents improved on their first valid attempt in 58.33% of settings. But when searches continued after the best observed score, 78.26% finished with a worse final attempt. The authors identify stronger runs as those with accurate hypotheses, validation-grounded supervision, behavior-aligned data, and preserved strong checkpoints. ArXiv · AI/CL/LG's note
RSIBench-Data fixes the post-training stack so agents are judged on data-centric research decisions, not surrounding systems work. In tests of four frontier agents across six benchmarks, agents improved on their first valid attempt in 58.33% of settings. But when searches continued after the best observed score, 78.26% finished with a worse final attempt. The authors identify stronger runs as those with accurate hypotheses, validation-grounded supervision, behavior-aligned data, and preserved strong checkpoints. ArXiv · AI/CL/LG's note
score 6