SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
Frontier agents can find SAE features for target concepts, but still trail expert references, especially on causal steering.
SAEScientist-Bench tests agents on 20 mechanistic interpretability tasks using Gemma Scope’s 131K-plus feature dictionary for Gemma-2-9B-IT. The agents design contrastive probes, search for the best feature, and are judged against Neuronpedia-anchored expert features. The paper says they can separate target concepts from controls relatively well, but often misread their own measurements. The benchmark frames experimental model understanding as a measurable capability for autonomous AI R&D. ArXiv · AI/CL/LG's note
SAEScientist-Bench tests agents on 20 mechanistic interpretability tasks using Gemma Scope’s 131K-plus feature dictionary for Gemma-2-9B-IT. The agents design contrastive probes, search for the best feature, and are judged against Neuronpedia-anchored expert features. The paper says they can separate target concepts from controls relatively well, but often misread their own measurements. The benchmark frames experimental model understanding as a measurable capability for autonomous AI R&D. ArXiv · AI/CL/LG's note
score 5