PaLoRA: Paced Low-Rank Adaptation for Continual Learning
PaLoRA argues that continual LoRA training needs rank-aware step scaling, not a fixed small learning rate.
The paper says finite-precision updates still leak into old-task subspaces even when gradients are projected away from prior knowledge. Its pacing law increases the restriction as the effective rank of past updates grows, aiming to balance retention with learning new tasks. PaLoRA combines adaptive SVD compression, nullspace projection, and that rank-aware pacing. The authors report consistent gains over prior methods, including about 4% accuracy on 50-task ImageNet-A and ImageNet-R benchmarks. HF Daily Papers' note
The paper says finite-precision updates still leak into old-task subspaces even when gradients are projected away from prior knowledge. Its pacing law increases the restriction as the effective rank of past updates grows, aiming to balance retention with learning new tasks. PaLoRA combines adaptive SVD compression, nullspace projection, and that rank-aware pacing. The authors report consistent gains over prior methods, including about 4% accuracy on 50-task ImageNet-A and ImageNet-R benchmarks. HF Daily Papers' note
score 4