Local Support Learning
LSL tries to stop forgetting by making new updates fire only on the kind of inputs they were trained for.
The paper frames catastrophic forgetting as a geometric problem inside each weight matrix’s input space. Its method pairs a normal trainable adapter with a gate, built from a Gaussian Mixture Model, that closes away from the current training distribution. The authors say this lets large language models keep pretrained and finetuned abilities across multiple phases without using prior data. They report results up to 7B-parameter LLMs, with memory and compute efficiency and robustness to hyperparameter choices. HF Daily Papers' note
The paper frames catastrophic forgetting as a geometric problem inside each weight matrix’s input space. Its method pairs a normal trainable adapter with a gate, built from a Gaussian Mixture Model, that closes away from the current training distribution. The authors say this lets large language models keep pretrained and finetuned abilities across multiple phases without using prior data. They report results up to 7B-parameter LLMs, with memory and compute efficiency and robustness to hyperparameter choices. HF Daily Papers' note
score 4