TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward
TILT steers diffusion sampling at test time using a reward drawn from the base model itself.
The paper frames compositional misses as cases where an image matches individual concepts but not their joint presence. Its reward favors samples where all requested concepts appear together, without an external reward model or added supervision. The authors derive a KL-constrained tilted target distribution and use it to guide diffusion sampling. On T2ICompBench prompts, they report better compositional alignment while maintaining image quality versus prior baselines. HF Daily Papers' note
The paper frames compositional misses as cases where an image matches individual concepts but not their joint presence. Its reward favors samples where all requested concepts appear together, without an external reward model or added supervision. The authors derive a KL-constrained tilted target distribution and use it to guide diffusion sampling. On T2ICompBench prompts, they report better compositional alignment while maintaining image quality versus prior baselines. HF Daily Papers' note
score 5