Learning Functional Subspaces for Neural Network Compression
LSP learns which low-rank subspaces to discard across the whole model, instead of pruning each layer by local criteria.
The paper says local compression rules can let errors compound with depth, especially at high compression. Its Learnable Subspace Projections keep pretrained weights frozen while jointly optimizing projectors against output KL divergence or the original training loss. After training, those projectors become standard low-rank factors, with shared factors for tied layer groups. Reported results include stronger performance than baselines across several LLMs and ViT-B/16, plus up to 1.6x faster decoding at small batch sizes. HF Daily Papers' note
The paper says local compression rules can let errors compound with depth, especially at high compression. Its Learnable Subspace Projections keep pretrained weights frozen while jointly optimizing projectors against output KL divergence or the original training loss. After training, those projectors become standard low-rank factors, with shared factors for tied layer groups. Reported results include stronger performance than baselines across several LLMs and ViT-B/16, plus up to 1.6x faster decoding at small batch sizes. HF Daily Papers' note
score 4