Megadose Built for builders and researchers.

Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering

· HF Daily Papers ·
The paper says diffusion language models can be steered during denoising to push targeted biased answers.

The authors describe an attack that watches a dLLM’s intermediate activations and adjusts a steering vector while the answer is still being revised. In tests, it increased LLaDA-8B-Instruct’s preference for a targeted demographic answer on ambiguous BBQ questions from 1.8 to 16.7 percentage points. On SocialStigmaQA, selection of stigmatizing answers rose from 17.6% to 58.1%. The paper argues that the denoising path itself becomes a control channel, so audits need to examine the serving setup as well as the frozen model. Source: HF Daily Papers' note.

score 4

Categories: Research