Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability

cs.LG Chuqin Geng, Li Zhang, Haolin Ye, Mark Zhang, Luke Zhang, Xujie Si · Oct 1, 2026
Local to this browser
What it does
Mechanistic interpretability tries to identify sparse circuit subgraphs that explain model behavior, typically evaluating candidates via intervention-defined faithfulness. This paper demonstrates that faithfulness objectives can prefer an...
Why it matters
This paper demonstrates that faithfulness objectives can prefer an equally sized circuit that reproduces intact-model behavior worse than an alternative, because ablating excluded components distorts the computational context of retained...
Main concern
The paper builds a careful, multi-pronged case that faithfulness-based circuit selection is unreliable. By first testing controlled, equally sized edits to reference circuits and then examining outputs of EAP, EAP-IG, ACDC, and Edge-SP,...
Community signal
0
0 up · 0 down
Sign in to vote with arrows
AI Review AI reviewed
Plain-language introduction

Mechanistic interpretability tries to identify sparse circuit subgraphs that explain model behavior, typically evaluating candidates via intervention-defined faithfulness. This paper demonstrates that faithfulness objectives can prefer an equally sized circuit that reproduces intact-model behavior worse than an alternative, because ablating excluded components distorts the computational context of retained ones. The authors argue this creates an objective-level recovery gap: better discovery algorithms cannot recover true mechanisms if the evaluation metric itself systematically rewards the wrong candidate.

Critical review
Verdict
Bottom line

The paper builds a careful, multi-pronged case that faithfulness-based circuit selection is unreliable. By first testing controlled, equally sized edits to reference circuits and then examining outputs of EAP, EAP-IG, ACDC, and Edge-SP, the authors isolate the evaluation objective from the search procedure. A context-restoration intervention repairs 96 of 100 persistent misrankings, lending strong support to the proposed mechanism. The work is primarily diagnostic---it does not offer a replacement faithfulness metric---but its demonstration that human-reference tasks suffer more misranking than the semi-synthetic InterpBench benchmark makes it directly relevant to current practice in pretrained language models.

“We show that intervention-defined faithfulness can instead prefer an equally sized circuit that reproduces the model's behavior less well, creating an objective-level recovery gap.”
Geng et al. · Abstract
“These findings show why better discovery alone is insufficient when its objective rewards the wrong candidate.”
Geng et al. · Section 6
What holds up

The experimental design is the paper's greatest strength. Testing controlled candidates before discovered outputs cleanly separates flaws in the objective from flaws in search (Section 3: 'If a score prefers the worse member of a pair, the failure cannot be attributed to the search procedure'). Equal-size controls rule out trivial size-based explanations, while the inclusion of both InterpBench and the Human suite shows the problem is not confined to small synthetic models---indeed, misranking is more frequent in pretrained language models. The context-restoration experiment is particularly elegant: by selectively replacing donor activations with intact-model signals while keeping circuits fixed, the authors show that repairing distorted inputs can correct rankings in 96% of cases (Section 5), supporting context distortion as a causal contributor.

“We first remove discovery from the experiment. Starting from a reference circuit, we generate alternatives of the same size and ask how faithfulness ranks them. If a score prefers the worse member of a pair, the failure cannot be attributed to the search procedure”
Geng et al. · Section 3
“Restoring selected signals from the recipient's intact-model execution repairs 96 of 100 persistent KL misrankings from the discovery pool on both validation and held-out prompts.”
Geng et al. · Section 5
Main concerns

The authors are appropriately cautious about their conclusions, noting that their behavioral criterion $Q$ measures output agreement under resampling rather than mechanism identity. Still, the reliance on pairwise reversals raises questions about whether the global optimization landscape is equally problematic; local misranking need not imply that the globally optimal circuit under $F^{\mathrm{KL}}$ is far from the best under $Q$. The context-restoration results are suggestive but incomplete: as the authors acknowledge, 'Restoration results concern a selected cohort of persistent KL failures; they neither identify minimal causal sets nor establish superiority over random restoration.' Additionally, the study is limited to GPT-2 small and one small code model on a handful of tasks, leaving open whether the effect scales to larger systems.

“Our behavioral criterion Q depends on ordinary resampling and does not by itself establish mechanism identity.”
Geng et al. · Section 6
“Restoration results concern a selected cohort of persistent KL failures; they neither identify minimal causal sets nor establish superiority over random restoration.”
Geng et al. · Section 6
Evidence and comparison

The evidence directly supports the misranking claims. Tables 2 and 4 report concrete failure rates---for instance, ACDC candidate--candidate KL misranking reaches 41.2% under resampling on human-reference tasks---and the magnitude of behavioral deficits is quantified alongside frequencies. The paper fairly situates its contribution within a growing literature on faithfulness unreliability, noting that prior work found 'existing methods are highly sensitive to seemingly insignificant changes in the ablation methodology' (Miller et al., 2024). The key advance over that prior work is showing that this sensitivity leads not merely to score variation but to systematic preference for behaviorally worse circuits, a stronger normative claim.

“Faithfulness measurements can change substantially with ablation methodology (Miller et al., 2024), and activation-patching conclusions depend on corruption and scoring choices (Zhang and Nanda, 2024; Heimersheim and Nanda, 2024).”
Geng et al. · Section 2.1
“existing methods are highly sensitive to seemingly insignificant changes in the ablation methodology”
Reproducibility

The authors document their pipeline meticulously. They use 'exactly the same frozen 100 validation prompts, 100 independent-test prompts, 100-prompt mean bank, and three donor assignments' across Sections 3 and 4 (Appendix A.1), and they rely on 'the authors' released implementations of EAP, EAP-IG, ACDC, and Edge-SP' (Section 4). What is missing is a public code release or permanent artifact link for the specific experimental harness used to generate candidates, form pairs, and compute scores. Without this, independent reproduction of the exact misranking rates and restoration experiments would require substantial reimplementation, even though the underlying discovery methods and model checkpoints are available.

“Sections 3 and 4 share exactly the same frozen 100 validation prompts, 100 independent-test prompts, 100-prompt mean bank, and three donor assignments.”
Geng et al. · Appendix A.1
“We use the authors' released implementations of EAP, EAP-IG, ACDC, and Edge-SP”
Geng et al. · Section 4
Abstract

Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circuit once it is found. We show that intervention-defined faithfulness can instead prefer an equally sized circuit that reproduces the model's behavior less well, creating an objective-level recovery gap. Across four human-reference tasks and InterpBench, we compare validation faithfulness with behavior on held-out prompts under fixed ordinary resampling. The behavioral criterion is agreement with the intact model, including its mistakes, except on Greater-Than, where we use semantic accuracy. Controlled reference edits reveal misranking without any discovery algorithm, and outputs of EAP, EAP-IG, ACDC, and Edge-SP exhibit the same failure. Under resampling, KL misranks 9.4%-41.2% of candidate pairs across these methods on the human-reference tasks. We investigate context distortion as an explanation: replacing excluded signals changes the inputs on which retained components operate. Restoring selected signals from the recipient's intact-model execution repairs 96 of 100 persistent KL misrankings from the discovery pool on both validation and held-out prompts. The circuits and their original behavioral scores remain unchanged. These findings show why better discovery alone is insufficient when its objective rewards the wrong candidate.

Challenge the Review

Pick a starting point or write your own. Challenges run in the background, so you can keep reading while the AI investigates.

No challenges yet. Disagree with the review? Ask the AI to revisit a specific claim.