Ambient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient Learning
RefineMix trains discrete diffusion models with out-of-distribution data only at diffusion times where it helps rather than shifts the sampler.
The paper targets severe data scarcity, especially in scientific settings. Its key claim is that masked discrete diffusion cannot use related data the same way continuous diffusion uses Gaussian noise, because surviving tokens still carry domain information. RefineMix instead uses out-of-distribution data at low noise levels, where disjoint domain supports can help learning without biasing generation. In protein sequence generation, the authors report that 197 in-domain examples nearly doubled generated proteins that were novel, foldable, and in-family versus standard finetuning. ArXiv · AI/CL/LG's note
The paper targets severe data scarcity, especially in scientific settings. Its key claim is that masked discrete diffusion cannot use related data the same way continuous diffusion uses Gaussian noise, because surviving tokens still carry domain information. RefineMix instead uses out-of-distribution data at low noise levels, where disjoint domain supports can help learning without biasing generation. In protein sequence generation, the authors report that 197 in-domain examples nearly doubled generated proteins that were novel, foldable, and in-family versus standard finetuning. ArXiv · AI/CL/LG's note
score 4