Megadose AI progress, ranked and analyzed.

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

· ArXiv · AI/CL/LG ·
SAE features can matter causally without acting like stable steering directions.

The paper introduces FEGA, which removes the same active SAE feature across contexts and studies the resulting changes in model logits. Across SAE variants, the authors find that clean one-dimensional downstream effects are uncommon. They separate value-like features, which more often have structured low-dimensional effects, from pointer-like features, whose effects are mostly diffuse and context-dependent. ArXiv · AI/CL/LG's note

score 5

Categories: Research