Feature Evolution and Migration during Vision Transformer Training
The paper tracks individual ViT features across both layers and epochs, finding that many shift location early in training.
The authors use sparse autoencoders on CLS-token representations to compare feature activations across epoch-layer pairs. They call this movement “feature migration,” meaning a feature becomes most detectable in a different layer as training progresses. In their experiments, migration is concentrated early, more often moves toward earlier layers, and fades as feature organization stabilizes. Deeper layers also stabilize earlier and more strongly than shallow ones. ArXiv · AI/CL/LG's note
The authors use sparse autoencoders on CLS-token representations to compare feature activations across epoch-layer pairs. They call this movement “feature migration,” meaning a feature becomes most detectable in a different layer as training progresses. In their experiments, migration is concentrated early, more often moves toward earlier layers, and fades as feature organization stabilizes. Deeper layers also stabilize earlier and more strongly than shallow ones. ArXiv · AI/CL/LG's note
score 4