Megadose AI progress, ranked and analyzed.

PACT: From Credit Assignment to Critic Alignment

· HF Daily Papers ·
PACT reports stronger post-training results by changing when the critic is trained.

The paper defines token-level credit through three conditions and uses that framework to compare OPD, RLOO, and GAE signals. Its proposed method, Policy Aligned Critic Training, updates the actor before the critic so critic training can be corrected toward the updated policy. In agentic math reasoning, it reports 72.87% average accuracy across four benchmarks, ahead of GRPO and PPO. On SWE-bench Verified, it reports a 67.4% pass rate, also above PPO, GRPO, and SAO. HF Daily Papers' note

score 5

Categories: Research