Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?
The paper says circuit tests can look strong while failing to reproduce the model’s wrong answers.
The authors argue that circuit explanations should be checked on failures as well as successes. Across IOI, Docstring, and Mechanistic Interpretability Benchmark tasks, many circuits matched correct behavior but missed most errors. On GPT-2 small IOI under mean ablation, tested circuits agreed on 97.3–99.5% of correct prompts but only 11.4–41.7% of error prompts. Restoring omitted attention heads recovered many of those errors in a held-out IOI case study. ArXiv · AI/CL/LG's note
The authors argue that circuit explanations should be checked on failures as well as successes. Across IOI, Docstring, and Mechanistic Interpretability Benchmark tasks, many circuits matched correct behavior but missed most errors. On GPT-2 small IOI under mean ablation, tested circuits agreed on 97.3–99.5% of correct prompts but only 11.4–41.7% of error prompts. Restoring omitted attention heads recovered many of those errors in a held-out IOI case study. ArXiv · AI/CL/LG's note
score 5