Source-linked AI summary

Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation

Ajay Pravin Mahale

arXiv:2608.13754v1cs.AI

TL;DR

Audit-relevant circuit evidence must support compatible conclusions across defensible analyses. Using a preregistered multiverse and deterministic regulatory claim mapping, the paper finds that current evidence does not meet that filing standard.

  • Problem

    Circuit-level evidence matters for filing only if independent analysts using the same system and tool reach compatible conclusions.

  • Method

    The study preregisters seven circuit-discovery axes and deterministically maps discovered circuits to structured Annex IV statements.

  • Results

    Circuit-level evidence does not meet the filing standard; standardising the metric removes 0.1377 of a 0.7316 flip rate.

  • Takeaways & Limitations

    Interpretability evidence’s evidentiary weight should be measured rather than assumed, and that measurement is relatively inexpensive.

  • Takeaways & Limitations

    The study covers one model and one task, so larger models may change the reported flip rate.

Abstract

from arXiv · show

The EU AI Act requires providers of high-risk systems to file technical documentation describing how the system reaches its decisions. Mechanistic interpretability is the obvious source of such evidence, and circuit discovery is its most developed instrument. We ask whether that evidence survives the condition under which it would be relied upon: two competent analysts, the same system, the same tool, different defensible settings. We pre-registered a crossed grid of seven analytic axes, every level taken from a published implementation, and mapped each discovered circuit through a deterministic claim map to a structured Annex IV statement. Across 15,840 pre-registered specifications on GPT-2 small and the indirect object identification task, of which 7,561 produced a claim, the derived statement flips across 73.2% of specification pairs (95% CI 0.725 to 0.738) and the modal claim commands 41.1% of the space. The evidence fails a filability criterion at every tolerance a conformity assessment body would plausibly accept. Standardising the single most influential choice, the evaluation metric, leaves the flip rate at 59.4%. Removing circuit size from the claim entirely and holding it fixed leaves 27.1% (95% CI 0.255 to 0.286), still above the pre-registered threshold. The circuits underlying these claims are structurally near-disjoint, median pairwise Jaccard overlap 4%, and functionally uncorrelated at Cohen's kappa 0.015, so the instability is not one mechanism described in different words. We give the filability criterion as a standalone protocol, and we report that one of the seven documented discovery objectives does not execute at all on the library's own canonical task. The study covers one model and one task, and whether the conclusion holds at scale is untested.

1 Introduction

The introduction frames audit value as requiring reproducible conclusions and asks whether circuit-level interpretability evidence remains compatible across defensible analytic choices. It presents a pre-registered measurement and filability criterion for assessing whether such evidence meets the standard that regulatory filing presumes.

  • An audit is valuable only when two auditors reach the same conclusion.
  • EU Regulation 2024/1689 requires high-risk AI providers to document system logic, interpretability measures, and meaningful explanations of automated decisions.Annex IV points 2(b) and 2(e), and Article 86(1), specify these documentation and explanation requirements.
  • Circuit discovery is presented as mechanistic interpretability’s mature tool for producing subgraphs that explain model behaviour.
  • A circuit-level filing has value only if another analyst using the same system and tool would produce a compatible account.
  • The study measures regulatory-claim movement across defensible specifications and introduces a standards-body decision rule without proposing a better method or rejecting interpretability.Its stated conclusion is that currently produced evidence does not meet the standard presupposed by filing it.

2 Related work

Prior work establishes that mechanistic-interpretability explanations may be non-unique, sensitive to methodological choices, and structurally distinct despite shared function. This study extends those concerns by crossing seven analytic axes, testing functional equivalence, and applying multiverse methods to regulatory evidence.

  • Non-identifiability: Méloux et al. show that a given behavior need not have a unique mechanistic-interpretability explanation; this study measures the cost when explanations are filed.The paper adopts their premise but examines downstream evidentiary consequences rather than publication-level explanation uniqueness.
  • Faithfulness metric sensitivity: Miller et al. find faithfulness scores highly sensitive to seemingly insignificant methodological changes, while this study crosses seven axes and propagates variation to a regulatory statement.Their released library is used as the instrument for this study’s stress test.
  • Structure versus function: Bayat Makou et al. find structurally distinct circuits implementing the same computation under fixed task conditions, a pattern termed phantom specialisation.This motivates testing function rather than inferring from circuit structure alone.
  • Auditability: Lan et al. document conflicting conclusions about the same behavior and call for audit standards, while Mueller et al. link comparison across methods to standardised evaluation.This study responds with a measurement and criterion aimed at standardised specification for certification.
  • Multiverse method: Steegen et al. advocate reporting results across defensible data-processing choices, while Simonsohn et al. provide the specification-curve visualization adapted here.Simmons et al. identify undisclosed analytic flexibility as enabling analysts to present almost anything as significant.
  • Explanation metrics under optimisation pressure: H​​sia et al. show that sufficiency and comprehensiveness can be inflated without changing predictions or explanations; two study metrics therefore receive separate discard-rate reporting.The paper treats metric-specific reporting as preferable to silently pooling results.

3 What this paper is not claiming

The paper does not claim that discovered circuits are random, that filings differ without mechanistic differences, or that its findings generalize beyond one model and task. It instead reports information-bearing circuits, substantial functional differences, evaluation-set variability, and limited scale claims.

  • 0.2746 against 0.4230 shows discovered circuits are more stable than a size-matched random null at fixed circuit size.The circuits carry information and are not claimed to be no better than random.
  • 0.3972 pooled and 0.4777 at the smallest circuits quantify functional instability, so differing filings do not arise without differing mechanisms.
  • Cohen’s kappa of 0.0146 indicates the circuits are functionally uncorrelated, not one mechanism described in different ways.
  • The study makes no claim about scale because it examines one model and one task.

4 Formalism

The formalism represents each analysis as a published, eight-axis specification, maps its discovered circuit to an Annex IV claim, and quantifies instability through pairwise disagreement. Filability is defined by requiring the modal claim share to meet a stated tolerance, accounting for the finite number of claim classes.

  • Specification space: Each specification is s = (o, a, d, m, τ, P, r, g), with every axis level taken from a published implementation.The axes are discovery objective, ablation operator, corruption distribution, evaluation metric, size threshold, prompt variant, seed, and granularity.
  • Circuit representation: Circuit discovery returns C(s) ⊆ E; for GPT-2 small’s factorised graph, |E| = 32,491.The edge count is confirmed both from the instrument and from an independent closed-form computation-order count.
  • Claim instability: The claim flip rate is the probability that two distinct specifications yield different structured Annex IV statements.The modal share is π∗ = maxc nc/N, where nc counts specifications producing claim class c.
  • Filability: Evidence is filable at tolerance α if and only if π∗ ≥ 1 −α, with nine observed claim classes giving a flip-rate ceiling of 0.8889.The observed flip rate 0.7316 is 82.3% of that ceiling.

5 The claim map

The claim map deterministically translates discovered circuits into nested statements for two regulatory addressees, using legally grounded ordering and a committed calibration rule for size bins. Its implementation was fixed in preregistration and published to isolate circuit variance from model-generated mapping variance.

  • Implementation: φ is fixed in preregistration, implemented as deterministic code, and published rather than generated by a language model.This avoids adding model variance to circuit variance, which would prevent separating the two sources.
  • Regulatory addressees: Two regulatory addressees receive nested claims: φoverseer follows Annex IV and Article 14(4)(c), while φaffected leads with the input segment under Article 86(1).The overseer ordering is dominant layer band, size class, then ranked input segment; the affected-person ordering begins with the input segment.
  • Claim structure: Three claim granularities are nested by construction, so finer-grained claims refine rather than contradict coarser claims.This nesting is part of the claim-map design.
  • Calibration: The size bins came from a committed calibration rule based on a measured node-count curve, not discretionary selection.Table 1 reports the resulting sensitivity.

6 Method

The study evaluates GPT-2 small on indirect object identification using a pre-registered, crossed discovery grid and reports uncertainty through specification-level bootstrap resampling. One library objective failed without repair, while the threshold rule discarded over half of specifications and produced an uneven surviving pool.

  • Model and task: GPT-2 small was evaluated on indirect object identification, enabling comparability with prior non-identifiability and faithfulness work and a meaningful size-matched random baseline.
  • Grid: 18,480 specifications were pre-registered across seven discovery objectives, seven ablation operators, four corruption distributions, four metrics, three thresholds, two prompt variants, and five seeds.Corruption distributions were nested because two ablation operators did not read them, avoiding duplicate specifications.
  • Statistics: 10,000 bootstrap replicates resampled specifications rather than pairs because flip rate and ¯J are pairwise U-statistics with dependent shared specifications.No p-value was computed across the designed specification grid.
  • Exclusions: 220 cells from one discovery objective failed at runtime because the library computed mean squared error against an integral target.The objective was not repaired because the library itself was the instrument under test and was reproduced verbatim.
  • Thresholding: 52.3% of 15,840 specifications produced no circuit under the metric-relative threshold rule, with discard rates distributed unevenly across metrics.The threshold rule admitted no circuit for 8,279 specifications, and the disclosure preceded the lock.
  • Thresholding: 52% comprehensiveness characterized the surviving pool, compared with 25% by design, and the authors reported this composition without smoothing.

7 Results

Across the pre-registered multiverse, circuit-derived claims are highly unstable, contradictory, and non-filable, with instability persisting after standardising the metric or removing circuit size from the claim. The circuits themselves show little functional agreement, indicating that the claims do not describe one mechanism in different words.

  • Claim instability: 73.16% was the claim flip rate (95% CI 72.47% to 73.80%), exceeding the 20% threshold and yielding a modal claim share of 41.09%, far below filability requirements.The pre-registered filability criterion required the interval lower bound to clear 0.20 and modal share to reach at least 0.80.
  • Claim instability: 41.1% and 28.5% of specifications respectively attributed the same model and task to early versus late layers, and the claims contradicted rather than differed in emphasis.The two claims were the most common outcomes in the claim space.
  • Claim instability: 39.5% of specification pairs still flipped under the coarsest two-class claim, while all six addressee-granularity combinations rejected filability.This shows the result was not caused by fine-grained claim language.
  • Robustness to standardisation: 59.39% of claims still flipped after completely standardising the metric, with the interval lower bound at 58.03%, remaining far above threshold.The metric was identified as the largest single lever for claim variance, but standardising it was insufficient.
  • Robustness to standardisation: 27.06% of claims still flipped after removing circuit size from the statement and holding size fixed, with a 95% CI of 25.50% to 28.59%.The interval lower bound remained above the pre-registered threshold, so attribution to model location remained unstable even without size information.
  • Functional agreement: 0.0146 was Cohen’s kappa for functional agreement, despite 60% raw agreement, showing that circuits were usually correct on the same fraction of examples rather than the same examples.Per-example functional instability was 0.3972, and the gap to claim flip rate was 0.3344 (95% CI 0.3297 to 0.3394).

8 The filability protocol

The filability protocol requires filers to declare the defensible specification space and a deterministic claim map, then report how much of that space agrees. It makes auditing implementable without requiring regulators to decide which analytic choice is correct.

  • Protocol requirements: The protocol begins by declaring the specification space S: defensible analytic axes and levels, each supported by a citation.This defines the alternatives a competent analyst could defend.
  • Protocol requirements: The claim map φ must be deterministic, published code rather than a language model.The protocol therefore fixes how circuits become filed claims.
  • Protocol requirements: Filers must report π∗, F, discard rates for each axis level, and within-group decompositions of pooled statistics.These outputs quantify agreement and attrition across the declared specification space.
  • Auditability: A filing that cannot state its own π∗ is asserted rather than audited, while the protocol requires only declaration of the space and reporting its agreement.Regulators need not adjudicate which analytic choice is correct for the protocol to be implementable.

9 Limitations

The study is limited to one model and one task, with replication on Pythia-160m pre-registered but not yet run. Its multiverse is constrained by library coverage, high and uneven discards, size-selection effects, and invalid bootstrap intervals for one reported family.

  • Scope: The study covers one model and one task, while a pre-registered 924-cell Pythia-160m replication has not yet been run.The replication uses a different architecture family and three seed levels instead of five.
  • Analytic limitations: Fixing circuit size drops F from 0.7316 to 0.2746, exceeding the reduction from standardising any pre-registered axis.Circuit size is selected by the threshold rule and distorts F, null-comparison direction, and functional-gap sign.
  • Analytic limitations: The measured F is a lower bound because the library omits optimal ablation, a defensible operator implemented in the field but not by the instrument.Thus, the multiverse spans the library’s choices rather than the literature’s full set of defensible choices.
  • Data composition: 52.3% of items are discarded, and the surviving pool is metric imbalanced despite a pre-fixed threshold rule and no repairs.The authors report this composition rather than adjust for it.
  • Uncertainty: Bootstrap intervals are unusable for one reported family because conditioning groups contain only two to five specifications and retained self-pairs dominate.Those quantities are reported as point estimates with the diagnostic attached.

10 Discussion

The discussion argues that standardising one analytic axis leaves substantial instability, while tooling limitations further constrain interpretability evidence. It proposes treating evidentiary weight as measurable rather than assumed.

  • What standardising one axis buys: 0.1377 of the 0.7316 flip rate is removed by standardising the metric, the biggest lever, so no single knob solves the problem.The passage warns that standardising one axis and declaring the problem solved would be mistaken by a wide margin.
  • On the instrument: One of seven documented, publicly exported discovery objectives does not run on the library’s own canonical task.The study reports this as a discard with its rate and as evidence about the maturity of tooling supporting regulatory evidence.
  • What a regulator can take from this: Interpretability evidence’s evidentiary weight should be measured rather than assumed, using a measurement cheap relative to the circuit-producing sweep.The passage explicitly distinguishes this conclusion from claiming that interpretability evidence is worthless.

11 Reproducibility

The confirmatory sweep was preregistered, reproducibly configured, and executed under a verified single environment, with the uploaded archive matching the cited code commit exactly.

  • 11 Reproducibility: 22 of 22 files matched between the uploaded archive and the recovered git tree, while every reported number traced to a config, seed, and environment fingerprint.The preregistration was timestamped before the confirmatory sweep, and the sweep covered all 1,320 cells under one verified environment.
Loading 2608.13754v1…