$λ$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource
The paper says Flow-GRPO’s denoising-step instability can be predicted and budgeted through a single “path variance” quantity.
The authors argue that drifting importance ratios, uneven clipping, and lost late-training samples are symptoms of the same per-step variance in the sampler’s Gaussian transition kernel. Their $\lambda$-Controlled GRPO uses that predicted law to calibrate updates and assign gradient effort across denoising steps. In tests on text-to-image reward tasks, it improves OCR-scored text rendering and preference-model reward over the strongest empirical stabilizer. ArXiv · AI/CL/LG's note
The authors argue that drifting importance ratios, uneven clipping, and lost late-training samples are symptoms of the same per-step variance in the sampler’s Gaussian transition kernel. Their $\lambda$-Controlled GRPO uses that predicted law to calibrate updates and assign gradient effort across denoising steps. In tests on text-to-image reward tasks, it improves OCR-scored text rendering and preference-model reward over the strongest empirical stabilizer. ArXiv · AI/CL/LG's note
score 5