Local Support Learning
The paper proposes gating fine-tuning adapters so they activate only on the distribution they were trained on.
Local Support Learning treats catastrophic forgetting as a geometry problem around each weight matrix. It pairs a normal trained adapter with a Gaussian-mixture gate meant to stay closed on inputs from prior phases, without using prior data. The authors report results on LLMs up to 7B parameters, retaining pretrained and finetuned capabilities across multiple phases with modest memory and compute. ArXiv · AI/CL/LG's note
Local Support Learning treats catastrophic forgetting as a geometry problem around each weight matrix. It pairs a normal trained adapter with a Gaussian-mixture gate meant to stay closed on inputs from prior phases, without using prior data. The authors report results on LLMs up to 7B parameters, retaining pretrained and finetuned capabilities across multiple phases with modest memory and compute. ArXiv · AI/CL/LG's note
score 5