ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning
ARISE-RL trains open-ended agents by having a task/rubric generator and a solver co-evolve around rubric-based rewards.
The paper says the setup targets RL settings where there are no gold answers and reward signals get brittle on long-horizon agent tasks. Its generator uses real tool observations to make valid, intermediate-difficulty tasks near the solver’s current capability. The solver learns from fine-grained rubric satisfaction signals while reasoning and using tools. The authors also introduce RG-SED and ECR-Bench, and report state-of-the-art results across their evaluated benchmarks. ArXiv · AI/CL/LG's note
The paper says the setup targets RL settings where there are no gold answers and reward signals get brittle on long-horizon agent tasks. Its generator uses real tool observations to make valid, intermediate-difficulty tasks near the solver’s current capability. The solver learns from fine-grained rubric satisfaction signals while reasoning and using tools. The authors also introduce RG-SED and ECR-Bench, and report state-of-the-art results across their evaluated benchmarks. ArXiv · AI/CL/LG's note
score 5