DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization
DiSCO is pitched as a black-box prompt-level safety layer for text-to-image models.
The paper targets “benign adversarial” prompts: text that appears safe but can still trigger NSFW outputs because of the model’s learned distribution. DiSCO adds optimized suffixes to prompts, using safe and unsafe image pools from the target model to steer generation. The authors report attack success rate reductions of 37.7% on undefended models and 25.13% on defended models, while preserving prompt meaning and improving coherence. HF Daily Papers' note
The paper targets “benign adversarial” prompts: text that appears safe but can still trigger NSFW outputs because of the model’s learned distribution. DiSCO adds optimized suffixes to prompts, using safe and unsafe image pools from the target model to steer generation. The authors report attack success rate reductions of 37.7% on undefended models and 25.13% on defended models, while preserving prompt meaning and improving coherence. HF Daily Papers' note
score 4