DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
DRACO turns one trajectory-level rubric score into step-level training credit without using verifiers.
The paper targets long-horizon agent training where ground-truth success checks are unavailable. DRACO generates rubrics during training, scores a completed trajectory once, then redistributes that judgment across the steps tied to the rubric items. The authors report gains on AppWorld over the base model and over GRPO trained with sparse ground-truth rewards. They also report out-of-domain improvement on Tau-Bench without a frontier judge. ArXiv · AI/CL/LG's note
The paper targets long-horizon agent training where ground-truth success checks are unavailable. DRACO generates rubrics during training, scores a completed trajectory once, then redistributes that judgment across the steps tied to the rubric items. The authors report gains on AppWorld over the base model and over GRPO trained with sparse ground-truth rewards. They also report out-of-domain improvement on Tau-Bench without a frontier judge. ArXiv · AI/CL/LG's note
score 5