Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning
EAPO uses entropy differently for wins and losses, rewarding uncertain successful steps while penalizing confident failed ones.
The paper argues that treating all high-entropy positions as prime targets can over-punish places where a failed response still had room to recover. Its method, Entropic Advantage Policy Optimization, redistributes rollout-level advantage into token-level credit without auxiliary models or extra supervision. The authors report the best overall performance across reasoning tasks and say the method improves exploration, problem coverage, and answer diversity. HF Daily Papers' note
The paper argues that treating all high-entropy positions as prime targets can over-punish places where a failed response still had room to recover. Its method, Entropic Advantage Policy Optimization, redistributes rollout-level advantage into token-level credit without auxiliary models or extra supervision. The authors report the best overall performance across reasoning tasks and say the method improves exploration, problem coverage, and answer diversity. HF Daily Papers' note
score 5