Megadose AI progress, ranked and analyzed.

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

· HF Daily Papers ·
DRACO turns one trajectory-level rubric score into step-level training signal without using verifiers.

The paper targets long-horizon agent training where ground-truth success checks are unavailable. DRACO generates rubrics during training, scores completed trajectories once, then redistributes that judgment across the steps tied to those rubrics for GRPO. The authors report a 15.9-point gain over the base model on AppWorld and a 5.3-point gain over sparse ground-truth-reward GRPO. They also report a 5.3-point out-of-domain gain on Tau-Bench over the base model. HF Daily Papers' note

score 5

Categories: Research