Megadose AI progress, ranked and analyzed.

Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression

· ArXiv · AI/CL/LG ·
A block-level optimization step sharply improves low-rank LLM compression before full-model refinement.

The paper says independent SVD truncation can look optimal per matrix while still compounding errors inside a Transformer block. Its three-stage method moves from whitened SVD to joint block optimization to end-to-end language-modeling refinement, using 256 calibration sequences and no recovery data. On LLaMA-7B at 60% compression, WikiText-2 perplexity falls from 42.1 after L1 to 19.1 after L2 and 11.4 after L3. The authors limit the claim to perplexity and compression fidelity, noting downstream accuracy remains well below the dense model. ArXiv · AI/CL/LG's note

score 5

Categories: Research