SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
Frontier agents can find SAE features for target concepts, but still fall short of expert interpretability work.
The benchmark asks agents to design contrastive probes and search 131K-plus Gemma Scope features in Gemma-2-9B-IT. Across 10 agent setups and 20 tasks, agents approached expert performance on separating target concepts from controls. They lagged badly on causal steering, and the authors say they often misread experimental measurements. HF Daily Papers' note
The benchmark asks agents to design contrastive probes and search 131K-plus Gemma Scope features in Gemma-2-9B-IT. Across 10 agent setups and 20 tasks, agents approached expert performance on separating target concepts from controls. They lagged badly on causal steering, and the authors say they often misread experimental measurements. HF Daily Papers' note
score 5