Let It Go or Learn to Self-Correct: Continuous Diffusion for Constrained Discrete Tasks
A sampler that stops clinging to early noisy states lifted Sudoku validity from 31% to 95% without retraining.
The paper argues that standard DDPM sampling can lock in early discrete mistakes on globally constrained tasks such as Sudoku, graph connectivity, Latin squares, and N-queens. The authors compare that approach with sampling directly from the model’s clean prediction and report consistent gains across the tested tasks. They also introduce self-correction training, exposing the model to its own predictions so standard samplers become more robust at inference. ArXiv · AI/CL/LG's note
The paper argues that standard DDPM sampling can lock in early discrete mistakes on globally constrained tasks such as Sudoku, graph connectivity, Latin squares, and N-queens. The authors compare that approach with sampling directly from the model’s clean prediction and report consistent gains across the tested tasks. They also introduce self-correction training, exposing the model to its own predictions so standard samplers become more robust at inference. ArXiv · AI/CL/LG's note
score 5