LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates
LoRA-TSD reframes LoRA optimization as a manifold tangent-space update, claiming stronger results without full-matrix costs.
The paper introduces an optimizer that applies Muon-style spectral descent inside the fixed-rank tangent space and retracts the result back into LoRA factors. The authors say its retraction is up to 2.8x cheaper than truncated-SVD retractions used in earlier manifold approaches. They also give global convergence guarantees for LoRA-Pro and LoRA-TSD under a tangent-projected gradient stationarity measure. In tests on six commonsense and NLI benchmarks using Llama and Qwen models, LoRA-TSD beat the compared LoRA optimizers and stayed robust across adapter ranks. ArXiv · AI/CL/LG's note
The paper introduces an optimizer that applies Muon-style spectral descent inside the fixed-rank tangent space and retracts the result back into LoRA factors. The authors say its retraction is up to 2.8x cheaper than truncated-SVD retractions used in earlier manifold approaches. They also give global convergence guarantees for LoRA-Pro and LoRA-TSD under a tangent-projected gradient stationarity measure. In tests on six commonsense and NLI benchmarks using Llama and Qwen models, LoRA-TSD beat the compared LoRA optimizers and stayed robust across adapter ranks. ArXiv · AI/CL/LG's note
score 5