Megadose AI progress, ranked and analyzed.

Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?

· ArXiv · AI/CL/LG ·
The paper says circuit tests can look strong while failing to reproduce the model’s wrong answers.

The authors argue that circuit explanations should be checked on failures as well as successes. Across IOI, Docstring, and Mechanistic Interpretability Benchmark tasks, many circuits matched correct behavior but missed most errors. On GPT-2 small IOI under mean ablation, tested circuits agreed on 97.3–99.5% of correct prompts but only 11.4–41.7% of error prompts. Restoring omitted attention heads recovered many of those errors in a held-out IOI case study. ArXiv · AI/CL/LG's note

score 5

Categories: Research