Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability
The paper says common faithfulness objectives can rank the wrong circuit above a better one.
The authors find “objective-level recovery gaps” where intervention-defined faithfulness prefers equal-size circuits that match the model’s behavior less well. They test this across four human-reference tasks and InterpBench, including outputs from EAP, EAP-IG, ACDC, and Edge-SP. Under resampling, KL misranks 9.4% to 41.2% of candidate pairs on the human-reference tasks. Restoring selected signals from the intact-model execution fixes 96 of 100 persistent KL misrankings without changing the circuits or their original behavioral scores. ArXiv · AI/CL/LG's note
The authors find “objective-level recovery gaps” where intervention-defined faithfulness prefers equal-size circuits that match the model’s behavior less well. They test this across four human-reference tasks and InterpBench, including outputs from EAP, EAP-IG, ACDC, and Edge-SP. Under resampling, KL misranks 9.4% to 41.2% of candidate pairs on the human-reference tasks. Restoring selected signals from the intact-model execution fixes 96 of 100 persistent KL misrankings without changing the circuits or their original behavioral scores. ArXiv · AI/CL/LG's note
score 5