Topographic Training Concentrates Causal Circuits Without Improving Neuron Monosemanticity
TopoLoss made ViT circuits more spatially concentrated, but did not make individual neurons cleaner to interpret.
The authors trained ViTs on ImageNet-100 with different TopoLoss weights and tested whether local topographic clusters carried more causal signal. At α=1.0, those clusters were 2.79× more causally sufficient than same-size random unit sets, with stronger effects at higher α. Sparse autoencoder measures worsened in some respects, including lower L0 sparsity and a 19-fold rise in dead features, while neuron-level monosemanticity scores stayed flat. The paper argues the gain is at the circuit level, not the single-neuron level. ArXiv · AI/CL/LG's note
The authors trained ViTs on ImageNet-100 with different TopoLoss weights and tested whether local topographic clusters carried more causal signal. At α=1.0, those clusters were 2.79× more causally sufficient than same-size random unit sets, with stronger effects at higher α. Sparse autoencoder measures worsened in some respects, including lower L0 sparsity and a 19-fold rise in dead features, while neuron-level monosemanticity scores stayed flat. The paper argues the gain is at the circuit level, not the single-neuron level. ArXiv · AI/CL/LG's note
score 4