Megadose AI progress, ranked and analyzed.

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning

· ArXiv · AI/CL/LG ·
The paper proposes a release-time safety gate meant to make open-weight models harder to jailbreak through later fine-tuning.

Its Unidirectional Safety Gate adds a Null Space Cubic Layer and an Inverse Adapter after the final Transformer layer. The authors say the layer suppresses gradients from harmful samples in a calibrated protected region, while the adapter preserves the base model’s normal forward behavior. In six evaluated model-dataset settings, the method kept post-fine-tuning attack success near the pre-release level under a fixed threshold, with trade-offs showing up on harder unsafe samples. ArXiv · AI/CL/LG's note

score 5

Categories: Research