Training Trajectories Determine Circuit Removability in Annealable Soft-Prior Transformers
The same small Transformer can keep or lose a retrieval circuit depending on how its soft positional prior is removed.
The paper tests an annealable soft-prior Transformer on associative recall and finds a sharp split: models trained with the prior active perform well, then collapse when the gate is zeroed. Smooth fade-to-zero training preserves most zero-gate accuracy, while forced-zero training, hard switching, and post hoc continuation do not. The same pattern appears on Markov induction, with linear regression ICL treated as a boundary case. Mechanistic traces suggest consolidation happens after the gate reaches zero, even though the heads involved differ by seed. ArXiv · AI/CL/LG's note
The paper tests an annealable soft-prior Transformer on associative recall and finds a sharp split: models trained with the prior active perform well, then collapse when the gate is zeroed. Smooth fade-to-zero training preserves most zero-gate accuracy, while forced-zero training, hard switching, and post hoc continuation do not. The same pattern appears on Markov induction, with linear regression ICL treated as a boundary case. Mechanistic traces suggest consolidation happens after the gate reaches zero, even though the heads involved differ by seed. ArXiv · AI/CL/LG's note
score 4