Megadose AI progress, ranked and analyzed.

When and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment

· ArXiv · AI/CL/LG ·
UECR-GRPO lets teacher scores affect both response ranking and token-level credit without changing the total verifier credit for a response.

The paper targets RLVR for math reasoning, where final-answer rewards give sparse feedback and teacher preferences can be dense but unreliable. Its method combines verifier reward with a teacher-to-anchor path log-ratio before group normalization, then redistributes token credit using teacher-policy gaps tempered by teacher entropy. On five math reasoning benchmarks, it reports Avg@12 gains over the strongest baseline of 0.89 points for Qwen3-1.7B and 0.56 points for Qwen3-4B. ArXiv · AI/CL/LG's note

score 4

Categories: Research