Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability
Mechanistic interpretability tries to identify sparse circuit subgraphs that explain model behavior, typically evaluating candidates via intervention-defined faithfulness. This paper demonstrates that faithfulness objectives can prefer an equally sized circuit that reproduces intact-model behavior worse than an alternative, because ablating excluded components distorts the computational context of retained ones. The authors argue this creates an objective-level recovery gap: better discovery algorithms cannot recover true mechanisms if the evaluation metric itself systematically rewards the wrong candidate.
The paper builds a careful, multi-pronged case that faithfulness-based circuit selection is unreliable. By first testing controlled, equally sized edits to reference circuits and then examining outputs of EAP, EAP-IG, ACDC, and Edge-SP, the authors isolate the evaluation objective from the search procedure. A context-restoration intervention repairs 96 of 100 persistent misrankings, lending strong support to the proposed mechanism. The work is primarily diagnostic---it does not offer a replacement faithfulness metric---but its demonstration that human-reference tasks suffer more misranking than the semi-synthetic InterpBench benchmark makes it directly relevant to current practice in pretrained language models.
The experimental design is the paper's greatest strength. Testing controlled candidates before discovered outputs cleanly separates flaws in the objective from flaws in search (Section 3: 'If a score prefers the worse member of a pair, the failure cannot be attributed to the search procedure'). Equal-size controls rule out trivial size-based explanations, while the inclusion of both InterpBench and the Human suite shows the problem is not confined to small synthetic models---indeed, misranking is more frequent in pretrained language models. The context-restoration experiment is particularly elegant: by selectively replacing donor activations with intact-model signals while keeping circuits fixed, the authors show that repairing distorted inputs can correct rankings in 96% of cases (Section 5), supporting context distortion as a causal contributor.
The authors are appropriately cautious about their conclusions, noting that their behavioral criterion $Q$ measures output agreement under resampling rather than mechanism identity. Still, the reliance on pairwise reversals raises questions about whether the global optimization landscape is equally problematic; local misranking need not imply that the globally optimal circuit under $F^{\mathrm{KL}}$ is far from the best under $Q$. The context-restoration results are suggestive but incomplete: as the authors acknowledge, 'Restoration results concern a selected cohort of persistent KL failures; they neither identify minimal causal sets nor establish superiority over random restoration.' Additionally, the study is limited to GPT-2 small and one small code model on a handful of tasks, leaving open whether the effect scales to larger systems.
The evidence directly supports the misranking claims. Tables 2 and 4 report concrete failure rates---for instance, ACDC candidate--candidate KL misranking reaches 41.2% under resampling on human-reference tasks---and the magnitude of behavioral deficits is quantified alongside frequencies. The paper fairly situates its contribution within a growing literature on faithfulness unreliability, noting that prior work found 'existing methods are highly sensitive to seemingly insignificant changes in the ablation methodology' (Miller et al., 2024). The key advance over that prior work is showing that this sensitivity leads not merely to score variation but to systematic preference for behaviorally worse circuits, a stronger normative claim.
The authors document their pipeline meticulously. They use 'exactly the same frozen 100 validation prompts, 100 independent-test prompts, 100-prompt mean bank, and three donor assignments' across Sections 3 and 4 (Appendix A.1), and they rely on 'the authors' released implementations of EAP, EAP-IG, ACDC, and Edge-SP' (Section 4). What is missing is a public code release or permanent artifact link for the specific experimental harness used to generate candidates, form pairs, and compute scores. Without this, independent reproduction of the exact misranking rates and restoration experiments would require substantial reimplementation, even though the underlying discovery methods and model checkpoints are available.
Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circuit once it is found. We show that intervention-defined faithfulness can instead prefer an equally sized circuit that reproduces the model's behavior less well, creating an objective-level recovery gap. Across four human-reference tasks and InterpBench, we compare validation faithfulness with behavior on held-out prompts under fixed ordinary resampling. The behavioral criterion is agreement with the intact model, including its mistakes, except on Greater-Than, where we use semantic accuracy. Controlled reference edits reveal misranking without any discovery algorithm, and outputs of EAP, EAP-IG, ACDC, and Edge-SP exhibit the same failure. Under resampling, KL misranks 9.4%-41.2% of candidate pairs across these methods on the human-reference tasks. We investigate context distortion as an explanation: replacing excluded signals changes the inputs on which retained components operate. Restoring selected signals from the recipient's intact-model execution repairs 96 of 100 persistent KL misrankings from the discovery pool on both validation and held-out prompts. The circuits and their original behavioral scores remain unchanged. These findings show why better discovery alone is insufficient when its objective rewards the wrong candidate.
Pick a starting point or write your own. Challenges run in the background, so you can keep reading while the AI investigates.
No challenges yet. Disagree with the review? Ask the AI to revisit a specific claim.