Source-linked AI summary
Causal Attribution for Agentic Decisions: Estimators, Coupling, and a Traceability Specification
Ajay Pravin Mahale
TL;DR
The paper addresses what records high-risk AI providers need for post-hoc causal attribution in agentic systems, where the required evidence is not established. It develops and checks alternative estimators, derives direct-effect and coupling machinery, and finds exact attribution failures while specifying the traceability record needed for future evaluation. The discrepancy experiment remains unrun because its live regulated pipeline was unavailable during the study window.
Problem
Agentic systems lack established requirements for records sufficient to determine post-hoc which system component caused a regulated decision.
Method
The paper separates marginal and common-random-numbers total effects, adds a pinned-downstream natural direct effect, derives coupling conditions, and pre-registers a discrepancy experiment and traceability specification.
Results
The marginal estimand makes an inert step identical to a decisive step on every planted-chain run, while the common-random-numbers effect is exactly zero for the decisive step on about one run in ten despite a direct effect of 0.25.
Takeaways & Limitations
An exact zero effect does not by itself establish that a step did nothing, so attribution requires accompanying execution records such as factual noise and change rates.
Takeaways & Limitations
The discrepancy experiment was specified but not run because the live regulated pipeline it required was unavailable during the study window.
Abstract
from arXiv · showhide
A provider of a high-risk AI system must keep records that make a decision traceable, and for agentic systems it has not been established what those records must contain for post-hoc causal attribution to be possible. We give the estimator framework and then the conditions under which it fails. We separate the marginal total effect that prior work measures from a common-random-numbers total effect that isolates a step's own contribution, add the natural direct effect under a pinned downstream, and check the estimators against hand derivations. Both estimands then fail, in the same direction. Under the marginal estimand a causally inert step has the identical total effect to the decisive one on every run of our planted chain, an algebraic identity and not a coincidence at one draw. Under common random numbers the decisive step returns exactly zero on the runs where the executing step flips, about one in ten, while its direct effect there is 0.25 and it demonstrably acts; an exact zero does not certify that a step did nothing, and we put that here rather than in the limitations. We derive the coupling that keeps the direct effect estimable once contexts diverge, with a closed form for its degradation, and show that the mediated share on which a natural ranking is built is not a share under suppression: where the direct and mediated paths oppose, it exceeds one and ranks a suppressed component above a pure mediator. We publish the discrepancy experiment's pre-registration rather than a result, because the live pipeline it requires was not available in the study window. We contribute the traceability specification such a filing would need, against a gap the Act's calendar opens: Article 86's right to an explanation has applied since 2 August 2026, while the Article 12 logging and Annex IV documentation that could evidence one were deferred to 2 December 2027 by Regulation (EU) 2026/1744.
1 Introduction
The paper asks what records would let post-hoc analysis establish which part of an agentic system caused a regulated decision. It separates competing causal estimands, identifies exact failure modes, and derives a traceability specification for evidence that remains unavailable or deferred.
- Agentic execution-provenance metrics remain proposed, without agreed definitions or adopted evaluation protocols.
- The paper separates marginal and common-random-numbers total effects, showing that the marginal estimand can make inert and decisive steps identical on every run.The common-random-numbers estimand separates those steps but can still return exactly zero for a decisive step.
- 52 of 60 draws receive different component names under the two estimands and the incumbent locus rule, with the direction characterized rather than scored.
- The paper adds a pinned-downstream natural direct effect, a coupling analysis with degradation formulas, and a mediation result showing that the proposed share is not a suppression share.Where direct and mediated paths oppose, the mediated share can exceed one and rank a suppressed component above a pure mediator.
- The discrepancy experiment was pre-registered and published but not run because the required live regulated pipeline was unavailable during the study window.The paper reports no rank correlation as a finding.
- The proposed traceability specification is mapped to filed documents and two Commission instruments that do not yet exist.A companion study concerns reproducibility across analysts, whereas this paper concerns whether evidence tracks causation.
2 Setting and estimands
The paper models an agentic decision as a trajectory of typed steps whose actions depend on prior actions and exogenous noise. Its estimands use interventions against the factual run, with resampling as the primary operator and direct-versus-mediated influence treated as distinct.
- A trajectory is a sequence of typed model, tool, retrieval, memory, or branch steps ending in a decision, with each action conditioned on the preceding prefix and exogenous noise.
- Every effect is defined as a contrast against the realized factual run, making the provider’s retained record part of what determines estimability.
- The framework adopts five intervention operators—resampling, forcing an action, forcing an observation, altering context, and altering policy—and uses resampling as primary.
- Removal is excluded because it can place the system off-distribution, paralleling instability from ablation choices in circuit-level interpretability.The cited comparison reports that filed interpretability claims flip across 73.2% of specification pairs.
- Agentic systems permit a step’s direct effect to be near zero while its total effect is large because it changes what a later step chooses to do.The paper treats this direct-versus-mediated distinction as unavailable in a single forward pass.
3 Two total effects, and why the distinction is the crux
The paper distinguishes marginal and common-random-numbers total effects because they answer different causal questions and fail in different ways. Marginal effects can make inert and decisive steps identical, while CRN effects can assign exact zero to an acting step.
- Estimands: Marginal and common-random-numbers total effects differ in whether downstream steps use fresh or factual noise.The former marginalizes downstream randomness; the latter holds it fixed after intervention.
- CRN failure mode: Under CRN, inert steps return exactly zero, but a decisive step also returns zero on runs of probability 1 −q.At q = 0.9 and w = 0.5, this occurs when the executing step flips, about one run in ten.
- Planted chain: Under the marginal estimand, inert and decisive steps have identical effects on every run of the planted chain.This is an algebraic identity, so no magnitude threshold can separate them.
- Interpretation: The CRN zero results from cancellation between opposing direct and mediated paths, so a zero alone does not establish inertness.The paper also notes that resampling cannot attribute influence once a policy has become deterministic.
- Empirical check: At q = 0.9, w = 0.5, the estimator found 18 exact-zero decisive-step runs, each with |TEcrn(1)| < 10^-12 and |DE(1)| = 0.254±0.014.The exact-zero rate was 0.133 at w = 1/2 and exactly 0 at every other swept value.
- Attribution target: The incumbent locus rule and the largest-CRN-effect rule selected different components on 52 of 60 draws because they answer different attribution questions.The locus rule identifies where the outcome became inevitable; CRN asks which action carried the difference.
4 The direct effect and the decomposition
The paper defines a natural direct effect by intervening at one step while pinning downstream actions, then treats the mediated effect as the residual from the common-random-numbers total effect. Pin plausibility must be reported separately by intervention arm because pooled plausibility is uninformative and the changed arm can make the direct-effect estimate a bound.
- Natural direct effect: The natural direct effect intervenes at step k while pinning every downstream step to its factual action.This is the paper’s operational definition of the direct-effect arm.
- Decomposition: The mediated effect is defined identically as ME = TEcrn − DE, not as an approximation.The paper checks the decomposition against hand derivations rather than treating floating-point subtraction checks as estimator validation.
- Decomposition: The mediated share for step 1 is 0.5, matching 1 − w because the mediated path carries weight 1 − w.Each of TEmarg, TEcrn, and DE separately matches its closed form within Monte Carlo error at every step.
- Pin plausibility: The pooled pin-plausibility value is 1/2 for every fidelity parameter, so it does not measure the pin’s fidelity.Measured pooled values are 0.483 at q = 0.9 and 0.482 at q = 0.99, while the arm-specific values differ sharply.
- Pin plausibility: At q = 0.9, pin plausibility is 0.899 when the intervention leaves the action unchanged and 0.097 when it changes it.At q = 0.99, the corresponding values are 0.990 and 0.010.
- Pin plausibility: When the pin changes the action, the continuation is one the policy would produce with probability 1 − q, so ME should be read as bounded rather than as a point estimate.Both intervention arms are therefore reported alongside every mediated-effect estimate.
5 Coupling across divergent contexts
The paper derives how shared-randomness coupling behaves when factual and counterfactual contexts diverge, distinguishing maximal coupling from shared-u inverse-transform sampling. It finds that agreement depends strongly on sampler design and perturbation details, while replay validity additionally requires reproducible, batch-invariant inference within a pinned stack.
- Shared-u coupling: Shared-u inverse-transform sampling keeps the same uniform draw across branches, but its agreement is determined by overlapping CDF intervals rather than guaranteed to attain maximal coupling.A first implementation incorrectly asserted that shared-u sampling reached the maximal-coupling bound; validation rejected that claim.
- Sampler order: At TV = 0.18, sorted-vocabulary sampling agrees 8.6% of the time, but sorting can instead raise agreement for peaked branches, so the robust requirement is shared, branch-independent order.The earlier general claim that sorting necessarily collapses agreement was withdrawn.
- Shared-u coupling: Mass moved between tokens shifts every subsequent CDF boundary, so an early index-order perturbation can decouple the entire tail.This explains why shared-u agreement can degrade when contexts diverge.
- Maximal coupling: Maximal coupling requires both branches’ distributions at draw time and achieves agreement 1 − TV, whereas the shared-u alternative has a closed-form gap from that optimum.The factual distribution must therefore be retained alongside the counterfactual one; this is available during factual replay at the cost of one extra forward pass per step.
- Agreement results: At TV = 0.179 and V = 32, maximal coupling agrees 0.821 of the time versus 0.412 for quantile coupling.Across twelve cells at TV = 0.18, maximal-coupling agreement remains 0.820 across vocabularies and source entropies, while quantile agreement varies.
- Scope of results: The reported agreement magnitudes are specific to vocabulary size and perturbation model, whereas only maximal coupling’s 1 − TV is model-independent.For example, at TV ≈ 0.18, changing the perturbation construction changes quantile agreement from 0.412 to 0.595 and sorted agreement from 0.086 to 0.321.
- Replay requirements: Replay measurements require batch-invariant kernels and a pinned inference stack because temperature zero with a fixed seed is insufficient for reproducibility.Where batch invariance is unavailable, effects smaller than the infrastructure’s null-replay floor cannot be reported as effects.
- Scope of results: The sampler result is established on synthetic categorical distributions and has not been run against a language model.Only maximal coupling’s 1 − TV is invariant across the modeled settings.
6 Inconsistent mediation, and a pre-registered statistic built on it
The paper shows that the mediated-share ratio |ME|/|TEcrn| is not a valid ranking quantity under inconsistent mediation. Opposing direct and mediated effects can make the ratio exceed one and rank suppression above pure mediation.
- The statistic: The pre-registered mediated-share statistic was built on a quantity mediation research has identified as unstable and inappropriate.The paper reports this defect rather than treating it as a footnote because the same construction can affect agent-attribution rankings.
- The ranking failure: A pure mediator has DE = 0 and mediated share 1, so suppression can outrank pure mediation under this ratio.The paper identifies this as a ranking failure that mixes opposing mechanisms into one score.
- The inequality: When direct and mediated effects have opposite signs, the mediated share can exceed one because the mediated effect is a difference rather than a non-negative part.The construction yields |TEcrn − DE| = |TEcrn| + |DE| under opposing signs.
- Operational interpretation: In an agent pipeline, inconsistent mediation can describe a retrieval that supports approval directly while causing a later verification step to raise a flag.This is presented as a realistic opposing-path configuration rather than a merely pathological case.
- Check against derivation: For the constructed chain, hand-derived values TEcrn(1) = +0.30, DE(1) = −0.10, ME(1) = +0.40, and ratio 4/3 matched rollout estimates within Monte Carlo error.The agreement validates the estimators while demonstrating that the quantity built on them is misleading.
7 The pre-registered experiment, published and not run
The discrepancy experiment is fully specified and pre-registered but was not run because the required live regulated pipeline was unavailable. The paper therefore publishes the procedures, guards, and analysis choices without presenting a rank-correlation finding.
- Status: The planned empirical question—whether filed observability recovers causally determining components—was specified but not answered because no live regulated pipeline was available.The authors explicitly publish the specification rather than substitute a synthetic proxy.
- Primary analysis: H2 is the pre-registered primary contrast: a joint two-degree-of-freedom Wald test of (βrec, βverb) = (0, 0) in a Plackett–Luce model.H1, H3, H4, and per-attributor analyses are exploratory.
- Models: The primary model uses decision-level clustering with a cluster-robust sandwich, while CR1 cluster-robust least squares is secondary and pre-specified for comparability.Any disagreement between primary and secondary analyses would be stated explicitly.
- Reporting: The reporting convention normalizes coefficients as g = β/∥β∥2, preserving scale invariance and bounding coefficients in [−1, 1].Decision-level bootstrap percentile intervals are used because the normalization constraint makes a symmetric delta-method interval unsuitable.
- Guards: The authors publish two tested separation guards and remove a |z| > 40 rule because it would flag strong, well-identified effects as separation as sample size grows.Detected separation makes the primary contrast indeterminate rather than prompting a refit reported as significant.
- Pre-registration: The pre-registration records operational details, including full threshold grids and unresolved rules whose values depend on measurements unavailable before deployment.This preserves later decision procedures without pretending to know unmeasured quantities.
8 What the Act requires, and the record it does not ask for
The Act requires system-level documentation and logging but does not specify the per-execution, component-level records needed for causal attribution in agentic systems. This gap remains while relevant obligations and supporting instruments are deferred or incomplete.
- Description versus record: Annex IV requires a general architecture description, whereas Article 12 concerns events over the system’s lifetime; neither supplies a per-instance causal account.The paper distinguishes design-time description from execution records.
- Logging scope: Article 12(2) bounds logging granularity by risk identification, post-market monitoring, and deployer monitoring rather than component-level attribution.The paper finds that none of these purposes imports a component-level requirement.
- Minimum content: For credit-scoring systems, Article 12(3) sets no minimum log content because its enumerated minimum applies only to remote biometric identification.The paper states that the same absence applies to every Annex III class except biometric identification.
- Standards: The standardisation request names no log content of its own and defers to Article 12, while the Regulation defers detailed specification to the future standard.The paper characterizes this as circular delegation rather than a concrete record specification.
- Access and granularity: Authorities have an access right to automatically generated logs, but the required granularity remains set by the provider because no minimum content is prescribed.Article 21(2) was deferred with the logging duty to 2 December 2027.
- Interpretation: The Regulation does not prohibit component-level attribution; it mandates logging for system-level purposes and leaves finer granularity to the provider.Appendix A therefore proposes the specification needed for causal traceability.
9 Limitations
The paper’s main empirical claim remains untested, and several estimator and specification results are established only on synthetic planted structures. The proposed traceability specification is also unevaluated in deployment.
- Empirical evidence: The discrepancy experiment was specified but not run because the required live regulated pipeline was unavailable during the study window.The authors identify this as the principal limitation.
- Estimator limitation: The common-random-numbers total effect is exactly zero for an acting step when direct and mediated paths balance and factual noise inverts the mediator.This occurs on one run in ten on the planted chain, while the real-pipeline rate is unknown.
- Validation scope: Estimator validation uses one family of planted generators, establishing correctness on known structures but not behavior on real trajectories.Real trajectories may have more steps, noisier outcomes, and non-Bernoulli policies.
- Coupling scope: The coupling results are exact for synthetic categorical distributions but have not been tested with a language model or serving stack.The paper notes progressively diverging contexts and untested batch invariance.
- Interactions: Single-step interventions omit interactions between steps, so per-step effects need not sum to a joint effect.The paper excludes a Shapley interaction arm rather than reproducing prior work.
- Specification status: Appendix A is derived from estimator requirements rather than filing experience, and its cost estimates are reasoned rather than measured.Requirement R6’s storage-versus-interval-width trade-off requires deployment to resolve.
10 Related work
The paper distinguishes post-hoc causal attribution from adjacent approaches in attribution, verification, provenance, mediation, and replay. It positions its specification against gaps in causal attribution evidence and observability protocols.
- Post-hoc attribution: Prior trajectory-attribution work uses oracle substitution or binary responsibility scores, rather than resampling under the same policy.These approaches primarily target failure repair or training supervision.
- Ex ante verification: Ex ante verification decides whether a proposed action should execute, whereas this paper explains decisions after a trajectory has run.The two functions do not substitute for one another.
- Provenance and observability: Evidence tracing and execution provenance lack agreed definitions and adopted protocols, motivating the paper’s regulatory observability specification.Existing semantic observability work provides requirements and a schema but addresses authorization non-identifiability as a scoping issue.
- Mediation methodology: The paper imports natural-direct-effect methodology and established cautions about mediated proportions, including inconsistent mediation when paths have opposite signs.It highlights the absence of this mediation literature from prior agent-attribution work.
- Determinism and replay: Replay research establishes forward-pass determinism limits and common-random-number methods, but prior rollout work does not address coupling agreement or attribution.These findings determine what a replay harness must control.
A A traceability specification for agentic high-risk systems
The specification defines cumulative record-keeping levels and requirements for reconstructing trajectories, replaying counterfactuals, and making attribution contestable. Its key threshold is L2, which requires counterfactual replayability and is not currently required by the Regulation.
- Levels: L0 records technical documentation and Article 12 logs, but is sufficient only for system-level risk and monitoring.It does not support any of levels L1 to L3.
- Levels: L1 makes a trajectory reconstructable by recovering and ordering the steps that produced a decision.Requirements R1 to R4 establish this level.
- Levels: L2 makes a named step counterfactually replayable while holding the rest of the trajectory at its factual randomness.This is the level at which TEcrn and DE become estimable.
- Levels: L3 supports attribution that remains contestable under adversarial reading, including its uncertainty.Requirements R10 to R12 define this level.
- Levels: The levels are cumulative, but the decisive jump is from L1 to L2: observability stacks commonly provide L1, while causal attribution requires L2.No current Regulation provision asks for L2.
- Core record requirements: R1–R4 require identifiers and ordering, step types, committed inputs and outputs, and a recorded non-model outcome function.Together they make the factual trajectory and its intervention targets recoverable.
- Replay requirements: R5–R9 preserve replay conditions through keyed randomness, draw-time distributions, serving-stack fingerprints, tool purity declarations, and external-state binding.Dropping R6 leaves shared-u coupling available but widens mediated-effect intervals; R5 cannot be added retrospectively.
- Accountability and filing: R10–R12 address retention parity, cross-actor completeness, and signed effect ratios with undefined-case rates.The specification also maps requirements to Annex IV and monitoring documentation, while identifying retention and accountability gaps.
A.5 What this specification does not claim
The specification establishes conditions for attribution to be possible from records, not evidence that attribution will be accurate or that the Regulation requires the full scheme.
- Scope boundary: The specification has not been evaluated against a live deployment and does not claim that conforming records make attribution accurate.Accuracy remains an empirical question addressed by the pre-registered Section 7 experiment.
B Measured coupling agreement
Measured coupling agreement tracks the closed-form results closely, while probability-sorted agreement is non-monotone in total variation divergence. The table shows the ordering mechanism behind the reported reversal.
- Measured coupling agreement: 0.0447 at TV = 0.3969 and 0.1655 at TV = 0.4245 show that probability-sorted agreement can fall and then rise as divergence increases.The reversal occurs because agreement depends on relative branch orderings, not divergence alone; all empirical columns are within 1.9 Monte Carlo standard errors of their closed forms.