Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
Mechanist is presented as an agentic system for automating mechanistic interpretability work on AI models.
The paper says it combines an interpretability knowledge graph of about 13,000 papers with a broader 43-million-paper database and a library of 32 mechanism-analysis methods. Its authors report that Mechanist produces more valuable hypotheses and runs experiments more reliably than Claude Code and existing AI-scientist systems. The examples include finding a safety risk from cross-modal transfer through apparently safe training data, developing a theory of model “belief,” and using the results to steer model behavior, including DNA-sequence generation. HF Daily Papers' note
The paper says it combines an interpretability knowledge graph of about 13,000 papers with a broader 43-million-paper database and a library of 32 mechanism-analysis methods. Its authors report that Mechanist produces more valuable hypotheses and runs experiments more reliably than Claude Code and existing AI-scientist systems. The examples include finding a safety risk from cross-modal transfer through apparently safe training data, developing a theory of model “belief,” and using the results to steer model behavior, including DNA-sequence generation. HF Daily Papers' note
score 5