Megadose Built for builders and researchers.

Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation

· HF Daily Papers ·
NEEDLE removes identified LLM backdoors without needing the clean model or poisoned training data.

The method estimates a backdoor direction and a refusal subspace from activation vectors, then uses sequential weight orthogonalisation to suppress the trigger behavior while preserving refusal-related representations. The authors say existing defenses can shift outputs on benign prompts, hurting performance and safety. Across multiple model families and attack types, NEEDLE had the lowest mean attack success rate among tested defenses, including 0% on challenging code injection attacks. It also produced the lowest KL divergence and minimal changes in capability and safety. HF Daily Papers' note

score 4

Categories: Research