Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR
DATPO trains reasoning models to explore more distinct solution paths, improving pass@k rather than only single-shot accuracy.
The paper argues that standard RLVR can raise single-sample performance while leaving reasoning coverage narrow. Its proposed method uses difficulty-adaptive tree rollouts, sentence-level entropy for branching, and a sibling-diversity advantage term. On mathematical reasoning benchmarks, the authors report stronger pass@k results than baselines, with better test-time scaling. HF Daily Papers' note
The paper argues that standard RLVR can raise single-sample performance while leaving reasoning coverage narrow. Its proposed method uses difficulty-adaptive tree rollouts, sentence-level entropy for branching, and a sibling-diversity advantage term. On mathematical reasoning benchmarks, the authors report stronger pass@k results than baselines, with better test-time scaling. HF Daily Papers' note
score 5