Megadose AI progress, ranked and analyzed.

SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

· ArXiv · AI/CL/LG ·
The paper proposes training an LLM to explain SAE features directly from their decoder directions.

SAEVerbalizer injects sparse-autoencoder decoder directions into an LLM’s internal representations, then fine-tunes downstream layers to verbalize what those directions mean. The authors say this avoids relying mainly on external behavioral evidence, which they describe as shallow and costly to collect at scale. In experiments, the method generalizes to unseen features, transfers across separately trained SAE dictionaries, and can extend to features from other LLMs with a lightweight adapter. ArXiv · AI/CL/LG's note

score 5

Categories: Research