Comparing Latent Concept Formation in State Space Models and Transformers via Sparse Autoencoders
The paper finds almost total feature-level alignment between Mamba-130m and Pythia-70m.
Using sparse autoencoders over a 10-million-token corpus, the authors report that 99.98% of Mamba features cluster near the upper alignment boundary with the Transformer model. The small divergent slice is described as syntax-heavy, not a broad semantic split. Their account is that Pythia can separate formatting edge cases, while Mamba compresses some rigid syntactic anomalies into polysemantic “junk drawer” neurons. ArXiv · AI/CL/LG's note
Using sparse autoencoders over a 10-million-token corpus, the authors report that 99.98% of Mamba features cluster near the upper alignment boundary with the Transformer model. The small divergent slice is described as syntax-heavy, not a broad semantic split. Their account is that Pythia can separate formatting edge cases, while Mamba compresses some rigid syntactic anomalies into polysemantic “junk drawer” neurons. ArXiv · AI/CL/LG's note
score 5