Megadose AI progress, ranked and analyzed.

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

· ArXiv · AI/CL/LG ·
SAE active latent sets diverged from human judgments of conceptual similarity.

The paper tests whether overlap among active sparse-autoencoder latents gives a more interpretable similarity signal than dense representation cosine similarity. It finds the set measure works in controlled toy settings and can form coherent text neighborhoods, but does not better recover human category boundaries or typicality. Under controlled semantic edits, changes in SAE active sets often failed to match human judgments of conceptual change. The authors read this as evidence against simple bag-of-features semantics for SAE features outside idealized cases. ArXiv · AI/CL/LG's note

score 4

Categories: Research