Megadose AI progress, ranked and analyzed.

Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons

· ArXiv · AI/CL/LG ·
The paper claims a training-free “tripwire” can push aligned models into refusal only when an attack is detected.

The method identifies safety-specific neurons with statistical tests and a utility filter, then clamps them to activations associated with harmful inputs. The authors describe two equivalent deployment modes: detector-gated inference intervention and an offline bias-patch weight edit. Across four aligned LLMs and four attack types, they report attack success rates of at most 2.0% with a 0.5% to 5.3% MT-Bench utility drop. Source: ArXiv · AI/CL/LG's note.

score 5

Categories: Research