Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning
The paper proposes a release-time safety gate meant to make open-weight models harder to jailbreak through later fine-tuning.
Its Unidirectional Safety Gate adds a Null Space Cubic Layer and an Inverse Adapter after the final Transformer layer. The authors say the layer suppresses gradients from harmful samples in a calibrated protected region, while the adapter preserves the base model’s normal forward behavior. In six evaluated model-dataset settings, the method kept post-fine-tuning attack success near the pre-release level under a fixed threshold, with trade-offs showing up on harder unsafe samples. ArXiv · AI/CL/LG's note
Its Unidirectional Safety Gate adds a Null Space Cubic Layer and an Inverse Adapter after the final Transformer layer. The authors say the layer suppresses gradients from harmful samples in a calibrated protected region, while the adapter preserves the base model’s normal forward behavior. In six evaluated model-dataset settings, the method kept post-fine-tuning attack success near the pre-release level under a fixed threshold, with trade-offs showing up on harder unsafe samples. ArXiv · AI/CL/LG's note
score 5