Megadose AI progress, ranked and analyzed.

Reinforcement Learning for Code Optimization

· ArXiv · AI/CL/LG ·
The paper says timed code rewards can work, but only after controlling the noise that usually breaks them.

The authors build DMC-Optim, a benchmark and sandbox meant to make execution-time feedback usable for RL. They report large pass@1 gains at stricter speed percentiles on Qwen 2.5 7B and CWM 32B while preserving correctness scores. Their setup also outperforms standard RLVR when the timing environment is degraded. On LCB, CWM 32B wins most median-sample speed comparisons, but still reaches only about half the human rate of complexity-class improvements. ArXiv · AI/CL/LG's note

score 5

Categories: Research