Megadose AI progress, ranked daily.

When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs

· ArXiv · AI/CL/LG ·
The paper says activation-aware pruning preserves SAE behavior better than magnitude pruning because it better controls representation-space perturbations.

The authors frame pruning damage through “perturbation energy,” a covariance-weighted norm for a fixed sparse autoencoder. They argue magnitude pruning misses activation geometry and can distort the space SAEs learned to read. Wanda and SparseGPT are reported as more robust, while middle layers emerge as the most pruning-sensitive. The paper uses that finding to propose layer-wise sparsity allocation that lowers perplexity at the same average sparsity. ArXiv · AI/CL/LG's note

score 5

Categories: Research