Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss
The paper argues that data selection, forgetting, and loss of plasticity come from the same update-level interaction inside a model.
The authors derive a token- and layer-wise way to estimate how learning from one token changes another prediction. They separate softmax force, shared readout geometry, and residual connections to expose two interaction channels. Positive interactions point to useful training experience, while negative ones explain interference through collision or erosion. Over longer training horizons, the shared geometry itself changes, weakening future learning signals. Source: HF Daily Papers' note
The authors derive a token- and layer-wise way to estimate how learning from one token changes another prediction. They separate softmax force, shared readout geometry, and residual connections to expose two interaction channels. Positive interactions point to useful training experience, while negative ones explain interference through collision or erosion. Over longer training horizons, the shared geometry itself changes, weakening future learning signals. Source: HF Daily Papers' note
score 5