Source-linked AI summary
Generalised Transportability via Causal Abstractions
Yorgos Felekis, Paris Giampouras, Fabio Massimo Zennaro, Theodoros Damoulas
TL;DR
Transportability traditionally asks whether individual causal queries can move from a source to a target, leaving non-transportable and target-agnostic settings unresolved. This paper reframes transportability as model-level causal abstraction, yielding certified intervals that bracket target queries across synthetic and ecological evaluations.
Problem
Existing transportability analysis treats queries individually, motivating a model-level relation between structured source and target causal models.
Method
The framework learns a single deterministic causal transport map and optimises it robustly over target-side mechanism and environment perturbations.
Results
Certified intervals bracketed ground-truth queries across benchmarks and an ecological dataset, while directional certificates were up to 17× tighter than general ones.
Takeaways & Limitations
Model-level transport provides practical query guarantees even for non-transportable effects and target-agnostic settings.
Takeaways & Limitations
Directional certificates can worsen when the declared mechanism-shift prior excludes the realised shift.
Abstract
from arXiv · showhide
Transporting a causal conclusion from a source study population to a target one is a fundamental problem in causal inference. The theory of transportability provides a criterion for when this is possible: given experimental data from the source and observational data from the target, it determines whether a target query is identifiable and does so completely; i.e. if the query can be transported, the criterion finds the exact formula. However, it works one query at a time and returns an expression rather than the value itself. It is also silent in two practically important regimes: when the query is not transportable and when no target data exist at all. To tackle both, we take a model-level perspective grounded in Causal Abstraction theory. Source and target share variables, graph, and interventions, differing only at a known set of mechanisms, which makes transportability a special case of same-level abstraction. Thus, instead of asking whether one query transports, we ask whether a single map aligns the source and target across their interventional behaviour. We characterise when such a map exists in both the Markovian and semi-Markovian settings; when it does, every target query transports at once. Our main contribution lies in the approximate case. When no exact map exists, the best approximate one still yields certified query intervals, recasting abstraction error as a quantitative notion of approximate transportability. We formulate model-level transport as distributionally robust optimisation over mechanism and environment perturbations of the unseen target and derive certificates for both challenging regimes: bounds for non-transportable queries, and guarantees under target-agnostic settings. We evaluate our framework on synthetic Markovian and semi-Markovian benchmarks and a real ecological dataset, and we show that the certified intervals bracket the true interventional query.
1 Introduction
The paper reframes transportability as a model-level causal-abstraction problem, replacing query-by-query verdicts with a map relating source and target interventional behaviour. It develops exact and approximate transport methods, including robust query certificates for non-transportable and target-agnostic regimes.
- Model-level perspective: Transportability is recast as same-resolution causal abstraction because source and target share variables, graphs, and intervention semantics while differing in known mechanisms.This model-level object replaces per-query transportability verdicts with a relation between domains.
- Model-level perspective: The framework defines a hierarchy of source–target consistency notions, culminating in a constructive map that transports the full interventional family.The hierarchy ranges from observational matching through full interventional transport and includes a query-restricted relaxation.
- Exact transport: Exact constructive abstractions are characterised for Markovian and semi-Markovian models, reducing to linear feasibility conditions with uniqueness criteria in finite state spaces.These characterisations specify when a single map exists across the two causal settings.
- Approximate transport: TraCA formulates approximate model-level transport as distributionally robust optimisation over target-side mechanism and environment perturbations around a trusted source SCM.Unlike standard DRO, uncertainty is placed on the unseen target rather than on the source data-generating distribution.
- Query-level certificates: Model-level transport error yields certified intervals for target queries, including bounds for non-identifiable queries and guarantees that hold uniformly over target ambiguity sets.The framework provides general certificates for any admissible target and sharper directional certificates when a prior is declared.
- Empirical validation: Across synthetic Markovian and semi-Markovian benchmarks and a real ecological dataset, certified intervals bracket ground-truth queries at the radius required by the true shift.The directional certificate is several times tighter than the general certificate.
2 Background
This section establishes the formal language for the framework: SCMs define causal domains and interventions generate post-interventional distributions. Query families separate relevant marginals from the functionals evaluated on them, while Markov kernels provide general transport operators between distributions.
- Structural causal models: An SCM consists of endogenous and exogenous variables, structural functions, and an environment distribution over exogenous variables.Acyclic structural assignments induce a causal DAG and a reduced-form mixing function from exogenous to endogenous states.
- Interventions and queries: Interventions fix selected variables, remove their incoming graph edges, and produce a post-interventional joint distribution.The framework considers a finite set of relevant interventions and associated post-intervention variables of interest.
- Interventions and queries: A query specification family identifies relevant interventional marginals, while a causal query is any measurable functional of one such marginal.The same interventional distribution can yield different queries, such as a mean, median, or variance, depending on the chosen functional.
- Transport operators: Markov kernels generalise deterministic transformations by assigning each input a probability distribution over outputs and mapping input measures to output measures.Deterministic measurable maps are recovered when the kernel is a Dirac measure.
- Transport operators: Together, SCMs, interventions, query families, and Markov kernels recast transportability as a structural relation between models rather than a collection of query-specific formulas.This language supports relations between source and target interventional distributions.
3 Transportability as a structural relation between models
Transportability assumes source and target share variables, graph, and intervention semantics while differing only in known mechanisms, but classical analysis answers one query at a time. This section reframes transportability as a hierarchy of model-level consistency relations, culminating in a common operator or map that can align interventional behaviour and imply simultaneous transport.
- Structural assumptions: Source and target SCMs share endogenous variables, a DAG, and intervention semantics, differing only in mechanisms or exogenous distributions at known shifted nodes K.Invariant nodes have identical mechanisms, while shifted nodes are represented by selection variables pointing into them.
- Query-level transportability: Classical transportability determines identifiability and derives a formula separately for each target causal query using source and target information plus selection-diagram assumptions.The criterion is complete for deciding whether a specified query is transportable, but it does not provide a single model-level relation across queries.
- Model-level relation: The model-level perspective asks whether one common structural map aligns the full source and target interventional families, a requirement stronger than transporting many individual queries.The paper develops a hierarchy of increasingly structured source-target relations, including an orthogonal query-family-restricted relaxation.
- Consistency hierarchy: Observational and per-intervention consistency are always achievable through Markov kernels, because a kernel can map any probability measure to any other.Per-intervention consistency allows the kernel to vary with each intervention, so it imposes no genuine compatibility constraint between the interventional families.
- Interventional consistency: Interventional consistency is the first nontrivial level, requiring one common operator across interventions; constructive consistency further requires a deterministic map that preserves full interventional semantics.Constructive and non-constructive interventional consistency remain distinct because a common kernel need not factorize.
- Implications for transportability: A constructively interventionally consistent map transports all queries simultaneously, whereas Q-restricted consistency guarantees transport only for queries in Q; transportability alone does not ensure a common map.The resulting mismatch between query-wise transportability and full-family alignment is called the compatibility gap.
4 Characterising Exact Model-level Transportability
Exact model-level transportability is characterised by recursively aligning shifted mechanisms on every parent context reached by the intervention family. In Markovian and semi-Markovian settings, stagewise assembly conditions are necessary and sufficient for a single componentwise transport kernel to recover all target interventional laws.
- Markovian case: Local alignment must cover every relevant parent configuration where a shifted mechanism is evaluated, including contexts created by interventions on its parents.Intervening directly on a shifted node bypasses its mechanism and imposes no condition on its local transport component.
- Markovian case: When shifted parents precede a shifted node, alignment is recursive and must preserve the relevant Markov structure of partially transformed laws.The later mechanism is matched against transported parent laws rather than original source conditionals.
- Markovian case: Theorem 8 characterises exact transport by stagewise matching: lifted local kernels successively align each prefix with the target interventional law, yielding the full target joint at the final stage.The componentwise kernel acts as the identity on invariant nodes, and the stagewise Markov-preserving condition is both sufficient and, within the componentwise class, without loss of genuine CICks.
- Semi-Markovian case: In the semi-Markovian setting, districts are assembled recursively with districtwise product kernels under stagewise district preservation.District contraction can create cycles, so the stated acyclic constructive route is a requirement of the construction rather than transportability itself.
- Semi-Markovian case: Theorem 12 gives a necessary-and-sufficient district-level criterion for a districtwise kernel to be a CICk, and genuine districtwise CICks automatically satisfy the required preservation conditions.The criterion checks partially transformed laws for every intervention and district in the graph.
5 Approximate Model-level Transportability
Section 5 relaxes exact model-level transportability to a discrepancy-based notion of transport error. It defines intervention-query and query-family errors that quantify how closely one common kernel transports source interventional distributions to target distributions.
- Motivation: Approximate transportability replaces exact equality between transported source and target interventional families with a quantitative discrepancy.This discrepancy defines model-level transport error, analogously to abstraction error.
- Intervention-query error: Intervention-query transport error measures, for a fixed intervention and output set, the discrepancy between transported source and target interventional distributions.Projecting both distributions onto the query coordinates yields the pointwise error.
- Query-family error: Aggregating pointwise errors over a query specification family defines the Q-restricted transport error of a common Markov kernel.The resulting quantity measures how far one transport operator is from transporting the relevant source interventional family to the target one.
- Connections and conditions: In the full post-interventional case, Q-restricted transport error reduces to model-level average abstraction error.Restricting kernels to deterministic or constructive classes recovers the corresponding approximate abstraction or constructive problems.
6 Modelling shifts in Linear Additive Noise models
This section models domain shifts in Linear Additive Noise models through changes in the environment and/or causal mechanism matrix. It derives finite perturbation expansions and stability bounds for interventional propagators, enabling robust optimisation.
- LAN representation: In LANs, X = AU with A = (I − W)^−1 = I + W + · · · + W^(d−1), where W is strictly lower triangular and nilpotent.W encodes direct causal mechanisms, while A encodes total causal effects along directed paths.
- Interventional propagators: Hard interventions remove incoming edges through a diagonal gating matrix Rι, yielding the interventional propagator Aι = (I − RιW)^−1.The same acyclic structure ensures a finite polynomial expansion for Aι.
- Sources of domain shift: Domain shifts are represented as changes in the environment U, the mechanism A, or both, with target uncertainty modeled by separate mechanism and environment perturbations.Mechanism perturbations preserve the fixed topological ordering and acyclic LAN class.
- Mechanism perturbation: Mechanism perturbations admit an exact finite expansion because A∆W is strictly lower triangular and nilpotent, so no convergence condition is required.Although ∆W may be sparse, the sandwiched form A∆WA can produce global propagator changes through directed-path propagation.
- Stability bounds: When γι is small, propagator sensitivity is locally linear in ∥∆W∥; for γι < 1, sharper control is governed by the geometric factor (1 − γι)^−1.These bounds convert structural-matrix perturbations into explicit post-interventional propagator bounds for target-agnostic robust optimisation.
7 Transportation via Causal Abstractions: TraCA
TraCA formulates target-agnostic transport as distributionally robust learning of a single map against structurally plausible shifts in an unseen target domain. Its robust objective and certificates improve as target ambiguity shrinks, approaching the full-information oracle under Lipschitz losses.
- Target-agnostic robust formulation: TraCA learns a single transport map T by minimizing worst-case transport loss over an ambiguity set containing plausible unseen target domains.The source SCM is treated as trusted; uncertainty concerns the target environment rather than source-data misspecification.
- Structured target shifts: The target-side adversary perturbs each intervention’s propagator, while mechanism perturbations are restricted to designated shifted nodes and may obey magnitude or pattern constraints.In the LAN setting, TraCA introduces a mechanism adversary ΔW and uses a shifted-node row mask to encode domain-varying mechanisms.
- Robustness guarantees: If A1 ⊆ A2, then R⋆(A1) ≤ R⋆(A2), so shrinking the ambiguity set lowers the robust objective and tightens the transport certificate.Whenever the true target perturbation lies in A1, the A1-robust map certifies that target.
- Robustness guarantees: Under Lipschitz losses, the gap between the robust objective and oracle value R⋆({ξ◦}) is controlled by ambiguity diameter and vanishes as ambiguity shrinks to ξ◦.These guarantees follow from the minimax structure and connect target-side information to certificate tightness.
8 Optimisation
Section 8 instantiates TraCA as robust min–max optimisation in Gaussian and empirical forms, sharing an alternating solver while differing in target-environment parameterisation. The resulting learned maps support explicit worst-case guarantees for model-level transport error and downstream queries.
- Optimisation formulations: TraCA uses Gaussian moment-based and empirical perturbation-based formulations that instantiate the same robust objective with different unseen-target parameterisations.Both formulations share the min–max backbone of alternating projected gradient descent-ascent; the empirical formulation additionally uses a Frobenius-budgeted noise perturbation.
- Transport-map constraints: Admissible transport maps are diagonal in the Markovian case and block-diagonal by district in the semi-Markovian case.Intervened coordinates are fixed to intervention values and excluded from Q-restricted losses, so their transport-map rows do not affect certified quantities.
- Unified alternating solver: Both formulations alternately update the transport map by projected descent and target-side adversarial variables by projected or proximal ascent with feasibility corrections.The common solver differs only in the parameterisation and projection of target-side environment perturbations.
- Mechanism ambiguity: Mechanism ambiguity sets constrain target structural perturbations while preserving the fixed acyclic support, ensuring the perturbed model remains a valid linear acyclic SCM.The framework supports global, row-wise, column-wise, entrywise, and directionally recentered perturbation geometries.
- Q-restricted objectives: The Q-restricted objective includes both full-joint and single-query settings as special cases.Choosing all non-intervened outputs recovers the full post-interventional objective, while a singleton intervention/output pair yields a single robust loss.
- Certified guarantees: The optimisation structure yields explicit worst-case guarantees for model-level transport error and downstream queries.These guarantees combine the learned transport map with the stability moduli αι.
9 Provable Robustness and Query-Level Transport
The section develops certified worst-case bounds for query-level transport by decomposing abstraction error into transport, mechanism, and environment terms. Lipschitz query functionals inherit these model-level guarantees as certified intervals, with directional certificates sharpening bounds for linear queries.
- Model-level certificates: Certificates decompose abstraction error into transport mismatch, mechanism instability, and environment ambiguity, and grow with each source of uncertainty.The mechanism term scales with α_ι, while the environment term scales with ε and the projected propagator norm.
- Model-level certificates: The query-restricted certificates specialise to full post-interventional output families and single-query settings, with corresponding bounds applying at objective minimisers.Full recovery uses Q_full, while the single-query case follows by taking |Q| = 1.
- Model-level certificates: For included queries, restricted certificates are never worse than corresponding full post-interventional certificates, whereas omitted queries receive no certificate from the restricted objective.The selector-free full-joint bounds are looser but useful when no query family is specified.
- Query-level transport: Lipschitz causal query functionals convert W2 or Frobenius abstraction-error bounds directly into certified intervals around the transported-source query value.The certificates apply to arbitrary designated queries, though they can be conservative for linear and mean queries.
- Query-level transport: Directional certificates sharpen linear-query bounds by replacing the global mechanism modulus with a readout-specific modulus that fixes the input and output directions.For benchmark single-outcome queries, the readout direction selects the outcome coordinate alone.
10 Experiments
Experiments evaluate transport-map learning and certified intervals across synthetic and ecological benchmarks, emphasizing radius selection, directional priors, and the trade-off between robustness and distortion. Results show that directional certificates and appropriately sized ambiguity sets can preserve informative coverage, while mismatched priors or over-robustification can undermine performance.
- Experimental protocol: Transport maps use diagonal rescaling, or block-diagonal transformations semi-Markovian settings, learned by minimizing worst-case transport loss with 5-fold cross-validation.Training and certification use the Q-restricted objective, with maps learned on source splits alone.
- ATE: On the ATE benchmark, observational transport-distance minimization selects an operating cell without interventional target data.With the offset prior, the selected cell has W2^2 = 0.040, versus 0.63 for the symmetric variant.
- ATE: Directional certificates remain informative beyond the radius where general certificates become vacuous.Under the symmetric variant, the general certificate is vacuous at rtrain ≥1.0, while the directional certificate remains informative across the grid for E[Y | do(X=0)] and until rtrain=2.0 for E[Y | do(X=1)].
- ATCE: ATCE exhibits a DRO crossover: the identity baseline wins at small test perturbations, whereas robustification becomes beneficial as perturbations increase.The crossover depends on training radius, so no single radius dominates across test magnitudes.
- LiLUCAS: At matched radius rtrain=0.5, the Gaussian directional certificate covers all four queries while the general certificate is already vacuous.The out-of-ball setting demonstrates the trade-off between interval tightness and coverage when the target lies outside the training ambiguity ball.
11 Guidelines for Practitioners
TraCA should be configured to match the available prior knowledge, with directional assumptions declared explicitly and ambiguity radii sized to genuine uncertainty. Practitioners should interpret general and directional certificates differently and treat observational radius calibration as a heuristic rather than a guarantee.
- Choosing the ambiguity geometry: Choose the ambiguity geometry to match prior knowledge: Frobenius, row- or column-wise, and entrywise sets correspond to increasingly specific mechanism-shift bounds.The entrywise box is described as the tightest and most informative choice when each coefficient has its own bound.
- Encoding a directional prior: Encode a directional prior through the offset δ: its sign specifies the expected shift direction, magnitude its size, and radius bjk the residual uncertainty.The guarantees are conditional on this declared domain knowledge; δ = 0 makes no directional commitment.
- What to expect from the map: Up to a 60% error reduction occurs when the realised shift is small and aligned with the declared direction, but oversized ambiguity radii can over-robustify the map.Under symmetric ambiguity, improvement over the identity appears only in specific regimes, including large shifts under the Gaussian objective.
- Calibrating the radius from observational data: Observational radius selection can work on ATE when observational distance has an interior optimum, but it fails as a guarantee on Portland with a monotonic distance increase.On ATE, the selected certified interval contains the true interventional query without using its value; on Portland, the selector chooses a radius whose directional certificate does not yet cover.
- Reading the certificates: TraCA provides general certificates valid across the ambiguity set and sharper directional certificates that exploit declared priors; directional intervals can be several times tighter.On some benchmarks, the general certificate is vacuous while the directional one remains informative, and coverage within the training ball preserves validity.
12 Conclusion … A.1 Markovian characterisation
The paper reframes transportability as model-level causal abstraction, yielding certified intervals for otherwise difficult regimes while identifying current linearity and graph-knowledge limitations. The appendices provide exact existence proofs, robustness analyses, benchmark specifications, and a recursive Markovian characterisation.
- 12 Conclusion: A single map can align the source and target interventional families, replacing query-by-query transportability with model-level causal abstraction.The framework defines consistency notions ranging from observational matching to constructive transport of the full interventional family.
- 12 Conclusion: 2.9× on ATCE and as much as 17× per query on LiLUCAS: directional certificates can substantially tighten certified intervals.Across non-transportable, target-agnostic, and ecological settings, the intervals bracket the ground-truth query at the radius required by the true shift.
- 12 Conclusion: TraCA returns a transported estimate with an interval containing the truth, including cases where general bounds are vacuous and target data are unavailable.On the real ecological data, the general bound remained the only usable interval at the radius required by the shift.
- 12 Conclusion: The framework is limited to linear abstraction maps and linear structural causal models, and it assumes access to the true causal DAG or an accurate causal-discovery estimate.Its stability bounds and certificate closed forms rely on the reduced form X = AU and the resolvent identity for A = (I −W)−1.
- 12 Conclusion: Future work includes doubly distributionally robust transport and extensions beyond acyclic district structures.The proposed robustness extension would account for uncertainty in both the source reference model and the target-side shift.
- Appendix overview: The appendices supply full proofs and finite-state refinements for exact Markovian and semi-Markovian existence results, transport-error attainment, and robustness certificates.They also include computable support-function bounds for directional linear-query certificates.
- Appendix overview: The appendices additionally document formulation-specific optimisation algorithms, benchmark structural equations and interventions, implementation details, and the relationship between TraCA and DiRoCA.These materials support the empirical and computational aspects of the unified presentation.
- Appendix A. Exact transportability existence results: Appendix A gives recursive proofs, finite-state reductions, and uniqueness refinements that make exact transportability existence results algorithmically checkable.A.1 formalises the stagewise prefix-matching intuition underlying the Markovian recursive assembly theorem.
A.2 Markovian finite-state stagewise compatibility, uniqueness, and deterministic specialisation · A.3 Semi-Markovian finite-state district compatibility, uniqueness, and related remarks
The finite-state theory reduces Markovian stagewise compatibility to linear feasibility, characterizes uniqueness by matrix rank, and identifies a stricter deterministic specialization. Semi-Markovian results carry these reductions to districts, where computational cost grows with district size but common kernels can enable universal transport.
- A.2 Markovian finite-state stagewise compatibility, uniqueness, and deterministic specialisation: At each Markovian stage, finite-state compatibility reduces to a linear feasibility problem, with stagewise Markov preservation checked separately.The formulation uses row-stochastic kernels, row-fixing constraints for intervention values, and compatibility equations for admissible contexts.
- A.2 Markovian finite-state stagewise compatibility, uniqueness, and deterministic specialisation: If rank(M_i) = n, the compatible Markovian stagewise kernel K_i is unique; if rank(M_i) < n, the unconstrained solution set has positive dimension.Stochasticity constraints can nevertheless reduce a nonunique affine solution set to a singleton.
- A.2 Markovian finite-state stagewise compatibility, uniqueness, and deterministic specialisation: Uniqueness requires at least n linearly independent rows of M_i, so admissible contexts and distinct intervention values together must number at least n.Richer intervention families may increase matrix rank and tighten or uniquely determine the compatible kernel.
- A.2 Markovian finite-state stagewise compatibility, uniqueness, and deterministic specialisation: A deterministic Markovian specialization exists exactly when one intervention- and context-independent state map τ_i transports every transformed source slice while preserving intervention values.Equivalently, K_i(· | x_i) = δ_{τ_i(x_i)}, which is stricter than stochastic compatibility.
- A.3 Semi-Markovian finite-state district compatibility, uniqueness, and related remarks: In the semi-Markovian setting, finite-state district compatibility and uniqueness are exact district-level analogues of the Markovian results.The local state space is dom[X_Dr], and compatible district kernels are obtained through the corresponding finite-dimensional constraints.
- A.3 Semi-Markovian finite-state district compatibility, uniqueness, and related remarks: Semi-Markovian compatible district-kernel existence is again a linear feasibility problem, with global existence requiring assembly into a district-preserving districtwise product kernel.The stagewise district kernels must be assembled compatibly across districts.
- A.3 Semi-Markovian finite-state district compatibility, uniqueness, and related remarks: The stagewise linear program at a shifted district has dimension N2^(sum_{i∈D_r} |dom[X_i]|), so computational cost grows rapidly with district size.This is identified as the main practical cost of moving from nodes to districts.
- A.3 Semi-Markovian finite-state district compatibility, uniqueness, and related remarks: A common constructive kernel can act as a universal semi-Markovian transport operator even when direct transportability fails.The query-level identity has the same formal shape as in the Markovian case, but its role is genuinely broader here.
Appendix B. Attainment of the transport error
The appendix characterises when transport error vanishes: this occurs exactly when an admissible common kernel realises exact Q-restricted consistency. Under compactness and lower semicontinuity, the minimum transport error is attained, preserving this equivalence.
- Exact attainment: EQ(K) = 0 exactly when the admissible common kernel K realises exact Q-restricted consistency.This is the pointwise equivalence for any admissible common kernel.
- Exact attainment: If exact Q-restricted consistency is realisable within K, a minimiser K⋆ has zero transport error and realises it.The minimiser argument uses the equivalence between zero error and exact consistency.
- Attainment conditions: Under compactness of K and lower semicontinuity of EQ, the infimum E⋆Q(K) is attained, so vanishing transport error is equivalent to exact realisability.The generalised Weierstrass theorem supplies attainment under these hypotheses.
Appendix C. Auxiliary structural results · C.1 Corollary 19
The auxiliary structural results show that changing a reduced-form mechanism cannot generally be replicated by changing only the exogenous law. Corollary 19 derives ambiguity-geometry-specific spectral-norm bounds for admissible perturbations.
- Appendix C. Auxiliary structural results: Two SCMs can share endogenous and exogenous spaces while differing in their reduced-form maps under a fixed environment.
- Appendix C. Auxiliary structural results: In general, no exogenous law can reproduce the distribution induced by replacing one reduced-form mechanism with another.
- Appendix C. Auxiliary structural results: A counterexample uses binary endogenous and exogenous spaces, a constant source map, and an identity target map under a uniform environment.
- Appendix C. Auxiliary structural results: Because the source map always outputs 0, every exogenous-law pushforward is δ0 and assigns zero mass to state 1, unlike the target distribution.
- C.1 Corollary 19: Corollary 19 bounds the perturbation term according to ambiguity geometry, converting geometry-specific matrix-norm controls into spectral-norm bounds.
- C.1 Corollary 19: Frobenius-ball bounds use submultiplicativity, while row budgets control the induced ∞-norm and column budgets control the induced 1-norm.
- C.1 Corollary 19: Since Rι is diagonal, it can zero out rows, and admissible perturbations vanish outside the allowed row set K.
- C.1 Corollary 19: Entrywise bounds yield |Aι|RιB control, while directional-box bounds replace bjk with |δjk| + bjk before taking suprema over admissible perturbations.
Appendix D. Information monotonicity of the robust value
The robust transport value is monotone in target-side information: shrinking the ambiguity set cannot worsen the robust objective or its certificate. Under smoothness, the gap to the oracle is controlled by ambiguity diameter and vanishes as the diameter goes to zero, with stronger convergence available under compactness and continuity assumptions.
- Quantitative tightening: When losses are Lipschitz in the target perturbation, the gap between robust and oracle values is bounded by ambiguity diameter.The oracle value is R⋆({ξ◦}), where ξ◦ is the true target-side perturbation.
- Information monotonicity: Shrinking the target ambiguity set can only decrease the robust transport value and improve the certificate on the true target loss.For A1 ⊆ A2, R⋆(A1) ≤ R⋆(A2).
- Oracle convergence: If a sequence of ambiguity sets has diamdA(Ak) → 0, the robust transport values converge to the oracle value.The claim follows from the stated two-sided bound and the diameter condition.
- Oracle convergence: Without Lipschitz continuity, nested compact ambiguity sets shrinking to the true perturbation yield stronger convergence under compactness of maps and joint continuity of losses.Berge’s maximum theorem and Dini’s theorem establish uniform convergence in the admissible map class.
Appendix E. Proofs of the robustness certificates … Appendix I. Note on relation to DiRoCA
The appendices prove robustness certificates by decomposing transport, mechanism, and environment discrepancies, then detail directional computation, optimisation, benchmarks, implementation, and the relation to DiRoCA. Together, they establish exact or conservative query certificates across Gaussian, empirical, Markovian, and semi-Markovian settings and document their evaluation domains.
- Appendix E. Proofs of the robustness certificates: The Gaussian and empirical proofs separately decompose discrepancy into transport, mechanism, and environment contributions and bound each term pointwise before aggregation.The empirical proof squares, averages over queries, and takes the supremum to obtain its certificate.
- E.1 Directional certificates for linear queries: Directional certificates split query error into transport, mechanism, and environment terms, with the mechanism term computed as a support-function budgeting problem over the admissible perturbations.For a single shifted node, nilpotency removes higher-order terms and yields a rank-one first-order calculation; multi-node budgets retain finite higher-order terms and may use the conservative bound γι/(1−γι) when γι < 1.
- E.1 Directional certificates for linear queries: The reported directional certificates are exact closed-form maxima rather than numerical-search outputs, and the argument applies to both diagonal Markovian and block-diagonal semi-Markovian transport maps.The rank-one identity was numerically verified across benchmarks, interventions, and admissible perturbations with residual at machine precision.
- Appendix F. Formulation-specific optimisation details: The optimisation appendices instantiate Algorithm 1 with Gaussian moment-based adversarial environments and smooth surrogates, versus empirical sample perturbations updated by projected ascent.For entrywise-box mechanism ambiguity, ATE, ATCE, and LiLUCAS enumerate at most 2^2 = 4 vertices and refine the worst case with three gradient-ascent steps; Portland is excluded because its root-only shift leaves the mechanism set empty.
- Appendix G. Benchmark datasets: The benchmark suite comprises three synthetic causal models and the Portland urban-streams domain, where canopy cover affects dissolved oxygen across source watersheds and target Fanno Creek under environmental shift.The appendix specifies structural matrices, intervention sets, and benchmark configurations; Portland contains 347 observations across four seasonal rounds and 87 unique sites, with N = 247 source and N = 100 target observations.
- G.1 ATE: The ATE benchmark is semi-Markovian with latent X↔Y confounding and shifts only the Y mechanism coefficient from βs = 1.0 to βt = 1.5.Its reported target intervention values are E[Y | do(X=0)] = 0.30 and E[Y | do(X=1)] = 1.80, compared with 1.30 on the source for the latter query.
- Appendix I. Note on relation to DiRoCA: TraCA shares DiRoCA’s min-max optimisation machinery but differs in scope: it transports between same-level source and target populations and converts abstraction error into certified causal-query intervals ℓ≤Qt≤L.DiRoCA instead robustifies a low/high-level abstraction within one system and bounds a scalar abstraction error.
Appendix J. Supplementary experimental results … J.4 Portland
Appendix J provides deferred benchmark results supporting Section 10.3’s claims, including certificate nesting, coverage transitions, query-scale effects, and radius-selection limitations. The Portland results show that observational fit can select a radius with zero coverage despite guarantees holding at declared radii.
- Appendix J. Supplementary experimental results: The appendix reports full profiles for the Section 10 query family under the Section 10.2 protocol, expanding claims referenced in Section 10.3.Results are organised by benchmark.
- J.1 ATE: At rtrain=0.2, both certificates are nested and widen with radius, while the general certificate widens faster; the offset band contains truth at every shown radius.Under the symmetric variant, truth falls outside the band at the smallest radius, with a coverage gap of 0.08.
- J.2 ATCE: Across three test magnitudes, directional coverage reaches one exactly when rtrain meets rtest, whereas the general certificate covers below the in-ball boundary.The general bound admits targets outside the ball for which it was constructed.
- J.3 LiLUCAS: For LiLUCAS, the empirical general certificate is absolutely narrower than the Gaussian certificate at every radius but remains vacuous throughout.Its vacuity threshold is less than half as large, showing that absolute width does not determine informativeness.
- J.4 Portland: On Portland, query coverage requires a radius containing the realised shift, while the directional certificate remains informative beyond that point.Radius selection therefore affects informativeness even when the guarantee itself remains valid.
- J.4 Portland: The observational transport distance increases monotonically with rtrain, so it has no interior optimum and cannot locate the radius needed for coverage.The observational selector returns the smallest ε, where coverage is zero.
- J.4 Portland: For Portland, the general certificate is vacuous from rtrain=0.5 onward, whereas the directional certificate stays informative throughout.The observational selector’s smallest-radius choice explains the zero-coverage outcome.