Local Sparsity Enables Unsupervised LLM Safety Detection
The paper argues unsafe prompts can be flagged by modeling only safe activation patterns, if the detector exploits sparse local structure.
The authors frame LLM safety detection as anomaly detection rather than supervised classification on known harms. Their method uses sparse autoencoder concept features and masks detection to the small active support shared by nearby points. They report validation across multiple architectures and datasets, including capability and safety benchmarks. With 1% out-of-distribution data for calibration, the locally sparse methods reach near-optimal performance while using only 1-2% of SAE neurons. ArXiv · AI/CL/LG's note
The authors frame LLM safety detection as anomaly detection rather than supervised classification on known harms. Their method uses sparse autoencoder concept features and masks detection to the small active support shared by nearby points. They report validation across multiple architectures and datasets, including capability and safety benchmarks. With 1% out-of-distribution data for calibration, the locally sparse methods reach near-optimal performance while using only 1-2% of SAE neurons. ArXiv · AI/CL/LG's note
score 5