Source-linked AI summary

Foundations of Structural Causal Models with Cycles and Latent Variables

Stephan Bongers, Patrick Forré, Jonas Peters, Joris M. Mooij

arXiv:1611.06221v6stat.MEcs.AIcs.LG

TL;DR

Cyclic SCMs with latent variables lack several guarantees enjoyed by acyclic models, including solvability, unique distributions, valid marginalization, and graph-semantic consistency. The paper characterizes conditions restoring these properties and introduces simple SCMs, which preserve many acyclic SCM conveniences in the cyclic setting.

  • Problem

    Cyclic SCMs cannot reliably guarantee solutions, unique induced distributions, valid marginalization, Markov properties, or graph-semantic consistency.

  • Method

    The paper analyzes solvability conditions and introduces simple SCMs, defined by unique solvability with respect to every variable subset.

  • Results

    Simple SCMs have unique observational, interventional, and counterfactual distributions, remain closed under interventions and marginalization, and respect latent projections.

  • Takeaways & Limitations

    Simple SCMs extend acyclic SCMs to cyclic settings while preserving many convenient causal and probabilistic properties.

  • Takeaways & Limitations

    Unique solvability is sufficient but not necessary for a counterfactually equivalent marginal SCM, so the condition may sometimes be relaxed.

Abstract

from arXiv · show

Structural causal models (SCMs), also known as (nonparametric) structural equation models (SEMs), are widely used for causal modeling purposes. In particular, acyclic SCMs, also known as recursive SEMs, form a well-studied subclass of SCMs that generalize causal Bayesian networks to allow for latent confounders. In this paper, we investigate SCMs in a more general setting, allowing for the presence of both latent confounders and cycles. We show that in the presence of cycles, many of the convenient properties of acyclic SCMs do not hold in general: they do not always have a solution; they do not always induce unique observational, interventional and counterfactual distributions; a marginalization does not always exist, and if it exists the marginal model does not always respect the latent projection; they do not always satisfy a Markov property; and their graphs are not always consistent with their causal semantics. We prove that for SCMs in general each of these properties does hold under certain solvability conditions. Our work generalizes results for SCMs with cycles that were only known for certain special cases so far. We introduce the class of simple SCMs that extends the class of acyclic SCMs to the cyclic setting, while preserving many of the convenient properties of acyclic SCMs. With this paper we aim to provide the foundations for a general theory of statistical causal modeling with SCMs.

1. Introduction

The paper develops foundations for causal modeling with SCMs that allow cycles, latent variables, and nonlinear relationships. It characterizes when acyclic-model properties persist under solvability conditions and introduces simple SCMs to retain many of those properties in cyclic settings.

  • SCMs express causal relationships as deterministic functional relationships and underpin statistical methods for inferring causal structure from data.
  • Acyclic SCMs generalize causal Bayesian networks and have convenient properties including unique distributions, intervention and marginalization closure, latent-projection compatibility, and Markov properties.
  • Cycles arise in feedback systems and can cause SCMs to lack solutions or have multiple solutions, undermining unique observational, interventional, and counterfactual distributions.
  • The paper studies cycles, latent variables, and nonlinear functions, determining sufficient solvability conditions under which standard SCM properties hold.
  • Under ancestral unique solvability conditions, the paper restores compatibility between cyclic SCM graph structure and causal semantics under interventions.
  • Simple SCMs are uniquely solvable with respect to every variable subset and preserve convenient properties of acyclic SCMs, including latent-projection compatibility and directed global Markov properties.

2. Structural causal models

This section defines SCMs independently of their solving random variables, allowing latent variables and cyclic functional relations. It formalizes structural equations, solutions, observational distributions, equivalence, and key properties of interventions and graph representations.

  • SCM definition: An SCM consists of endogenous and exogenous variable domains, a measurable causal mechanism f, and a product distribution over exogenous variables.The exogenous components are mutually independent under the product measure.
  • Structural equations and interventions: Structural equations express each endogenous variable through a deterministic causal mechanism, while interventions modify mechanisms targeting selected variables.The framework imposes no acyclicity assumption and permits self-cycles.
  • Solutions: A solution is a pair of random variables whose exogenous distribution equals PE and whose endogenous variables satisfy X = f(X,E) almost surely.Solutions are defined up to almost sure equality.
  • Solutions and observational distributions: Cyclic SCMs may have multiple observational distributions or no solution at all, so observational behavior is not generally unique or guaranteed to exist.With f1(x,e) = x2 and f2(x,e) = x1, any distribution supported on {(x,x)} is observational; replacing f1 with x2 + 1 yields no solutions.
  • Equivalence and structural minimality: Equivalent SCMs have the same solutions and causal and counterfactual semantics, and every SCM has an equivalent structurally minimal representative.Structural minimality can therefore be achieved without changing the model's equivalence class.
  • Interventions: Perfect interventions on disjoint variable subsets commute, preserve acyclicity, and commute with the twin operation on SCMs and graphs.However, an intervention can produce an SCM without a solution.

3. Solvability

This section formalizes solvability and unique solvability through measurable solution functions for endogenous-variable subsets. It establishes equivalences and conditions that extend key acyclic-SCM properties to cyclic models, while identifying important limitations and graph-theoretic consequences.

  • Definitions: Solvability w.r.t. O means a measurable function maps external endogenous and exogenous inputs to values satisfying O's structural equations almost surely.The function is defined on X_pa(O)\O × E_pa(O) and outputs X_O.
  • Global solvability: An SCM has a solution if and only if its structural equations have a solution almost surely, if and only if it is solvable.The equivalence requires measurable-selection arguments in the cyclic case.
  • Scope and limitations: Acyclic SCMs are uniquely solvable w.r.t. every subset, and cyclic SCMs can share this property, but solvability for one subset generally does not transfer to strict subsets, supersets, unions, or intersections.The section gives a cyclic example uniquely solvable for every subset and separately notes the failure of general inheritance properties.
  • Unique solvability: For any subset O, unique solvability is equivalent to almost-sure uniqueness of the corresponding structural-equation solution; globally, it guarantees existence and a common observational distribution across all solutions.Unlike ordinary subset solvability, unique solvability requires the measurable solution function to be unique up to a null set.
  • Structural consequences: Solvability w.r.t. O is equivalent to ancestral solvability, and interventions preserve (unique) solvability when they include suitable external parents; for linear SCMs, unique solvability also implies ancestral unique solvability.The intervention condition requires pa(O) \ O ⊆ I ⊆ I \ O.

4. Equivalences

Section 4 defines observational, interventional, and counterfactual equivalence as progressively finer ways to compare SCMs through observational, interventional, and counterfactual distributions. Equivalence implies counterfactual equivalence, which implies interventional and observational equivalence, although these coarser notions need not preserve the SCM graph.

  • Observational equivalence: Observational equivalence means that two SCMs are indistinguishable by their observational distributions, but it does not imply full equivalence.Equivalent SCMs are observationally equivalent, while observationally equivalent SCMs can have different augmented graphs.
  • Interventional equivalence: Interventional equivalence requires observational equivalence after every perfect intervention on the variables under consideration.It is preserved when restricting the variable subset, but not necessarily when enlarging it; interventionally equivalent SCMs can also have different graphs.
  • Counterfactual equivalence: Counterfactual equivalence requires the twin SCMs to be interventionally equivalent on both the original variables and their copies.It is finer than interventional equivalence, yet counterfactually equivalent SCMs can still have different augmented graphs.
  • Hierarchy: The equivalence notions form a hierarchy: equivalence implies counterfactual equivalence, which implies interventional equivalence and observational equivalence.This hierarchy compares SCMs at different abstraction levels and formalizes the last two implications in the ladder of causation.

5. Marginalizations

Marginalizing an SCM over endogenous variables preserves observational, interventional, and counterfactual semantics when the model is uniquely solvable on the marginalized subset. Without suitable solvability, marginalization may fail, and the resulting graph need not equal the latent projection.

  • Semantic preservation: Unique solvability with respect to L is sufficient for constructing a marginal SCM on I \ L that preserves observational, interventional, and counterfactual semantics.The equivalence holds simultaneously for all three semantics on the retained variables.
  • Failure conditions: If an SCM is not solvable with respect to L, no SCM on I \ L can be interventionally equivalent to it with respect to the retained variables.A concrete example shows interventions on the margin for which the candidate marginal model has a solution but the original model does not.
  • Definition and construction: Marginalization is defined by substituting a measurable solution function for the marginalized variables into the remaining causal mechanisms.The resulting SCM is defined up to equivalence because solution functions are unique only up to probability-zero sets.
  • Composition and limitations: Marginalization is not defined for every subset, although sequential marginalization over disjoint subsets agrees with marginalization over their union when the required unique-solvability conditions hold.A variable with a self-cycle may be impossible to marginalize alone, while jointly marginalizing it with another variable can remain possible.
  • Counterfactual equivalence: Unique solvability is sufficient but not necessary for a counterfactually equivalent marginal SCM, so the condition can sometimes be relaxed.Interventional equivalence alone generally does not imply counterfactual equivalence, but this marginalization construction provides both under its stated condition.
  • Latent projection: A marginalized SCM need not respect the latent projection: its augmented graph can be a strict subgraph when substituted paths cancel, although acyclic SCMs remain closed under marginalization.A sufficient condition for respecting the latent projection is provided, while general marginalizations need not satisfy it.

6. Markov properties

For cyclic SCMs, the directed global Markov property based on d-separation can fail, even when the graph d-separates variables that remain conditionally dependent. For uniquely solvable SCMs, σ-separation provides a general directed global Markov property, while d-separation remains valid under specific solvability, discreteness, or linearity conditions.

  • Failure of d-separation: The directed global Markov property does not hold in general for cyclic SCMs.In Example 6.1, X1 and X2 are d-separated given {X3,X4}, yet are dependent under every solution.
  • General directed global Markov property: For uniquely solvable SCMs, the observational distribution exists uniquely and satisfies the general directed global Markov property when the model is componentwise uniquely solvable.The criterion is based on σ-separation, which extends d-separation through graph acyclification.
  • Conditions for d-separation: The ordinary directed global Markov property holds for acyclic SCMs, ancestrally uniquely solvable discrete SCMs, and certain linear SCMs with nontrivial exogenous dependence and exogenous density.These are the three conditions listed in Theorem 6.3.
  • Relationship between Markov properties: The σ-separation criterion is generally weaker than the d-separation criterion because σ-separation implies d-separation.The discrete result requires ancestral unique solvability, correcting an earlier erroneous theorem.

7. Causal interpretation of the graph of SCMs

This section restricts attention to cyclic SCMs with uniquely defined observational and interventional distributions, while showing that graph structure and causal behavior can still diverge. It provides sufficient intervention-based criteria for detecting directed and bidirected relations in latent projections, but not universal identification methods.

  • Assumptions: The analysis restricts attention to SCM graphs whose induced marginal observational and interventional distributions are uniquely defined.Without this restriction, cyclic SCMs may have none, one, or multiple induced distributions, and intervention can change this status.
  • Graph interpretation: Cyclic SCMs can yield causal behavior inconsistent with their graphs, including interventions on nonancestors changing a variable’s distribution.The example retains uniquely defined marginal distributions under intervention, yet intervening on variable 1 changes the marginal distribution of variable 4 even though 1 is not an ancestor of 4.
  • Directed relations: If interventions on i produce different unique marginal distributions for j under the stated solvability condition, the latent projection contains directed edge i →j.The criterion applies to detecting either a directed edge or path; setting O = I targets an edge, while O = {i,j} targets a directed path.
  • Identification limits: These criteria are sufficient but incomplete: some directed edges and bidirected edges cannot be identified from observational, interventional, or counterfactual distributions.Examples include a directed edge for which no intervention satisfies Proposition 7.1 and a bidirected edge where p(x2 |do(X1 = x1)) = p(x2 |X1 = x1) for every x1.
  • Bidirected relations: Under the proposition’s solvability and uniqueness assumptions, if j is not an ancestor of i, the latent projection contains a bidirected edge i ↔j.The condition compares observational and interventionally induced distributions while requiring unique marginal distributions for the relevant variables.

8. Simple SCMs

Simple SCMs are defined by unique solvability for every subset of variables, extending acyclic SCMs to cyclic models while preserving closure under marginalization, intervention, and counterfactual construction. Their distributions exist uniquely and satisfy global Markov properties, while graph-based causal relationships remain only partially identifiable from distributions.

  • Definition and scope: Simple SCMs are uniquely solvable with respect to every subset of variables, meaning each subset of equations has a unique solution for its associated variables.This class extends acyclic SCMs to cyclic models while preserving many convenient properties.
  • Closure properties: Simple SCMs are closed under arbitrary-order marginalization, perfect intervention, and the twin operation, with marginalizations respecting latent projection and remaining simple.These closure properties ensure that repeated marginalization and causal transformations stay within the class.
  • Graph and causal semantics: Simple SCMs contain acyclic SCMs as a subclass and have no self-cycles, because self-cycles prevent variables from being uniquely determined by their parents.They also satisfy solvability conditions that allow causal relationships to be defined from the graph.
  • Distributional properties: Observational, interventional, and counterfactual distributions for simple SCMs all exist uniquely and satisfy the corresponding general directed global Markov properties.Under any of Theorem 6.3's conditions (1a), (1b), or (1c), they also satisfy the directed global Markov property.
  • Causal relationships: Causal and confoundedness relationships can be characterized from graph structure under sufficient conditions, but distributions alone generally cannot identify all such relationships.This limitation already occurs in acyclic SCMs without further assumptions.
  • Counterfactuals and potential outcomes: All counterfactuals are defined for simple SCMs, including cyclic models, enabling potential outcomes to be defined through measurable solution functions after perfect interventions.Potential outcomes are represented as X_ξI := g_Mdo(I,ξI)(E_pa(I)).

9. Discussion … A.2. Markov properties

The paper shows that cycles and latent variables undermine several guarantees familiar from acyclic SCMs, while solvability conditions recover many of them. It introduces simple SCMs as a well-behaved cyclic class and explains how σ-separation supports Markov properties in general cyclic models.

  • 9. Discussion: Cyclic SCMs may lack solutions or unique observational, interventional, and counterfactual distributions, unlike the generally convenient acyclic setting.
  • 9. Discussion: Appropriate unique-solvability conditions extend acyclic operations and results to cyclic SCMs, including equivalence comparisons and marginalization.
  • 9. Discussion: Simple SCMs extend acyclic SCMs to cycles while preserving unique distributions, closure under interventions and marginalization, latent-projection respect, and Markov properties.
  • 9. Discussion: Simple SCM solutions satisfy conditional independencies implied by σ-separation, enabling direct extensions of acyclic adjustment, do-calculus, and identification methods.
  • 9. Discussion: Future work includes reducing exogenous variables while preserving semantics, extending identifiability results, representing selection bias, and proving completeness results.
  • A.1. Directed (mixed) graphs: The appendix introduces directed mixed graphs, cyclic graph terminology, and the DAG formed by strongly connected components.
  • A.2. Markov properties: Directed global Markov properties extend d-separation to acyclic mixed graphs but do not generally hold for cyclic SCMs.
  • A.2. Markov properties: Under unique solvability with respect to every strongly connected component, an SCM’s observational distribution satisfies the general directed global Markov property via σ-separation.

A.2.1. The directed global Markov property

The directed global Markov property links d-separations in a directed mixed graph to conditional independencies in an SCM’s observational distribution. It holds under specific solvability and structural conditions, but can fail even for uniquely solvable cyclic SCMs.

  • Definition and graphical criterion: D-separation associates graph-implied conditional independencies with the observational distribution through the directed global Markov property.The criterion generalizes d-separation for DAGs and m-separation for ADMGs and mDAGs.
  • Sufficient conditions: A uniquely solvable SCM satisfies the directed global Markov property if it is acyclic, ancestrally uniquely solvable with discrete endogenous spaces, or meets specified linearity and density conditions.The linear condition requires each causal mechanism to depend nontrivially on at least one exogenous variable, with PE having a Lebesgue density.
  • Sufficient conditions: Under these conditions, the observational distribution exists uniquely and satisfies the directed global Markov property relative to the SCM’s graph.This theorem applies to the graph G(M) associated with the SCM.
  • Counterexample: A uniquely solvable cyclic SCM can violate the property: X1 and X2 are dependent given {X3,X4} despite being d-separated by that set in the graph.The counterexample is even simple, showing that unique solvability alone does not guarantee the property in cyclic models.

A.2.2. The general directed global Markov property · A.3. Modular SCMs · A.3.1. Definition of a modular SCM

The general directed global Markov property extends d-separation to cyclic SCMs through σ-separation and holds for SCMs uniquely solvable within every strongly connected component. Modular SCMs instead specify a HEDG, compatible solution functions, and intervention semantics directly as modular graphical models.

  • A.2.2. The general directed global Markov property: Acyclification replaces each strongly connected component’s mechanisms with measurable solution functions, producing an acyclic SCM observationally equivalent to the original.The graph of the acyclified SCM is contained in the acyclification of the original graph, potentially as a strict subgraph.
  • A.2.2. The general directed global Markov property: Under unique solvability for every strongly connected component, an SCM has a unique observational distribution satisfying the general directed global Markov property relative to its graph.The proof acyclifies the SCM into an observationally equivalent acyclic SCM and applies the directed global Markov property there.
  • A.2.2. The general directed global Markov property: σ-separation extends d-separation by adding strongly connected-component conditions for blocking non-colliders, while reducing to d-separation in acyclic graphs.It is defined through σ-blocked paths and is equivalent for walks and paths.
  • A.2.2. The general directed global Markov property: The general directed global Markov property is weaker than the directed global Markov property, and d-separation need not imply σ-separation in cyclic graphs.Example A.18 exhibits d-separation without σ-separation for X1 and X2 conditional on {X3,X4}.
  • A.2.2. The general directed global Markov property: For simple SCMs, observational, interventional, and counterfactual distributions all exist uniquely and satisfy the general directed global Markov property under their corresponding graphs.With at least one condition from Theorem A.7, these distributions also satisfy the stronger directed global Markov property; σ-faithfulness remains an open completeness question.
  • A.3. Modular SCMs: Modular SCMs are causal graphical models defined on directed graphs with hyperedges, or HEDGs, rather than graphs derived from structural equations.A HEDG consists of a directed graph and a simplicial complex whose inclusion-maximal elements are maximal hyperedges.
  • A.3.1. Definition of a modular SCM: A modular SCM assigns compatible measurable solution functions to every loop, ensuring that solutions for nested loops agree when restricted to smaller loops.Its tuple contains a HEDG, product measurable spaces for variables and hyperedge noises, the compatible system, and a product noise measure.
  • A.3.1. Definition of a modular SCM: Solutions of modular SCMs are constructed inductively over strongly connected components, and perfect interventions are defined directly on the underlying HEDG using a mapping φ.The intervention changes hyperedges and associated noise spaces, so its semantics depend on the choice of φ.

A.3.2. Relation between SCMs and modular SCMs … B.2. (Unique) solvability w.r.t. strict super- and subsets

The paper relates modular SCMs to underlying SCMs that are loop-wisely solvable and have compatible solution functions, while characterizing simple SCMs through unique loop-wise solvability. It also summarizes graphical-model coverage and solvability conditions, including σ-compactness for subsets and preservation over ancestral subsets but not arbitrary strict supersets or subsets.

  • A.3.2. Relation between SCMs and modular SCMs: A modular SCM’s underlying SCM is loop-wisely solvable and has a compatible system of measurable solution functions.Every modular-SCM solution is also a solution of its underlying SCM.
  • A.3.2. Relation between SCMs and modular SCMs: A modular SCM is an SCM equipped with additional compatible solution-function structure, and its underlying graph can be sparser than the modular SCM’s HEDG.The induced graph may contain fewer bidirected edges and fewer loops than the HEDG.
  • A.3.2. Relation between SCMs and modular SCMs: Simple SCMs are exactly the SCMs that are loop-wisely uniquely solvable; for them, every family of measurable solution functions is compatible.This equivalence guarantees the existence of a compatible system for simple SCMs.
  • A.4. Overview of causal graphical models: The graphical-model overview distinguishes models representable by SCMs from the narrower class representable by acyclic SCMs.The SCM class is depicted by the gray area, while the acyclic-SCM class is depicted by the dark gray area.
  • APPENDIX B: (UNIQUE) SOLVABILITY PROPERTIES: Appendix B develops solvability and unique-solvability properties for subsets, strict super- and subsets, unions, and intersections, with proofs provided in Appendix E.The listed results organize the appendix’s treatment of solvability preservation and sufficient conditions.
  • B.1. Sufficient condition for solvability w.r.t. subsets: σ-compactness of the relevant solution space, together with nonemptiness, is a sufficient condition for solvability with respect to a subset.The condition applies for P_E-almost every e and every fixed x\O; σ-compactness includes countable discrete spaces, real intervals, and Euclidean spaces.
  • B.2. (Unique) solvability w.r.t. strict super- and subsets: Solvability with respect to O does not generally extend to arbitrary strict supersets or strict subsets, but it does extend to every ancestral subset in G(M)O.An example exhibits solvability for {1,2} and {2,3} but failure for {1}, {3}, and {1,2,3}.

B.3. (Unique) solvability w.r.t. unions and intersections

(Unique) solvability is not generally preserved under unions or intersections. However, unique solvability combines across ancestral subsets under suitable assumptions and can be checked nodewise on ancestral subsets.

  • General limitations: In general, (unique) solvability is not preserved under unions or intersections.This failure includes unions of disjoint subsets.
  • General limitations: An SCM can be uniquely solvable for {1,2} and {2,3} but not for their intersection.Example B.3 constructs such an SCM with three variables and mechanism functions f1(x)=0, f2(x)=x2 · (1 − 1{0}(x1 · x3)) + 1, and f3(x)=0.
  • Ancestral unions: If A and à are ancestral subsets and the SCM is uniquely solvable for A, Ã, and A ∩ Ã, it is uniquely solvable for A ∪ Ã.The subsets must lie within O and be ancestral in the induced graph G(M)O.
  • Nodewise checking: Ancestral unique solvability with respect to O is equivalent to unique solvability for the ancestral subset of O consisting of each node’s ancestors.Thus, checking these node-specific ancestral subsets suffices to establish ancestral unique solvability for O.

APPENDIX C: LINEAR SCMS · APPENDIX D: EXAMPLES

Appendix C characterizes solvability, unique solvability, and marginalization for linear SCMs through matrix conditions, while Appendix D supplies examples motivated by feedback systems and supporting the main text.

  • APPENDIX C: LINEAR SCMS: For linear SCMs, solvability with respect to L is characterized by a necessary-and-sufficient pseudoinverse matrix condition.The condition uses A_LL = I_L − B_LL and applies for P_E-almost every exogenous realization and every x_O.
  • APPENDIX C: LINEAR SCMS: When solvability holds, the resulting mapping g_v is a measurable solution function for the SCM.The construction applies for every vector v ∈ R^L.
  • APPENDIX C: LINEAR SCMS: Unique solvability with respect to L holds if and only if A_LL = I_L − B_LL is invertible, yielding a measurable solution function g_L.This specializes the general solvability condition to matrix invertibility for linear SCMs.
  • APPENDIX C: LINEAR SCMS: For linear SCMs, unique solvability with respect to L is equivalent to ancestral unique solvability and to unique solvability on every strongly connected component of G(M)_L.This is a positive result specific to the linear class.
  • APPENDIX C: LINEAR SCMS: If I_L − B_LL is invertible, marginalizing L produces a linear SCM by substitution, and the marginalization respects the latent projection.The marginal causal mechanism is defined on the remaining endogenous variables and exogenous variables.
  • APPENDIX C: LINEAR SCMS: Linear SCMs and their marginalizations are observationally, interventionally, and counterfactually equivalent with respect to the retained variables, and every marginalization respects the latent projection.Simple linear SCMs additionally remain closed under marginalization.
  • APPENDIX D: EXAMPLES: Appendix D presents examples of SCMs describing equilibrium states of feedback systems governed by random differential equations and additional examples supporting the main text.The feedback-system examples motivated the study of cyclic SCMs.

D.1. SCMs as equilibrium models · D.2. Additional examples

The section shows how cyclic SCMs represent equilibrium systems, including mechanical and market models, interventions, and counterfactuals. It also notes that additional examples support the paper’s main text.

  • D.1. SCMs as equilibrium models: Feedback systems such as economic markets and coupled masses can be modeled through SCMs representing the causal semantics of equilibrium states.These systems are naturally described using differential equations with feedback between observed variables.
  • D.1. SCMs as equilibrium models: A damped coupled harmonic oscillator has a unique equilibrium position when all friction coefficients satisfy b_i > 0.At equilibrium, velocities and accelerations vanish, and the forces on each mass sum to zero.
  • D.1. SCMs as equilibrium models: Treating spring lengths, spring constants, and endpoint distance as fixed parameters yields a linear SCM for the oscillator’s equilibrium positions.The construction follows by translating the equilibrium equations of the dynamical system into causal mechanisms.
  • D.1. SCMs as equilibrium models: The oscillator SCM supports perfect interventions, and forcing mass j to Q_j = ξ_j makes the equilibrium positions solve the intervened model M_do({j},ξ_j).The model is identified as a simple SCM using Proposition C.3.
  • D.1. SCMs as equilibrium models: The market equilibrium model represents price, supply, and demand with a self-cycle for price, thereby encoding the equilibrium condition X_D = X_S.The resulting linear SCM is uniquely solvable and illustrates how self-cycles enlarge the class of SCMs.
  • D.1. SCMs as equilibrium models: For the market model, a counterfactual intervention changing supply from s to s′ after observing price p yields the deterministic price p + (s′ − s)/β_D.The counterfactual is formulated through the intervened twin model and is represented by a Dirac measure at the resulting price.
  • D.2. Additional examples: Additional examples are provided to support the main text.This subsection serves as supplementary material rather than introducing a separate modeling result.

Section 2 … Section 7

The examples show that SCMs with latent variables and cycles can exhibit non-equivalent representations, failures of solvability and marginalization conditions, and counterfactuals that observational and interventional data cannot identify. They also demonstrate that graph structure may fail to reflect counterfactual equivalence or latent causal relations.

  • Section 2: Structural equations that agree almost surely can define the same solutions, while changing a mechanism on a probability-zero event leaves solution status unchanged.In Example D.4, the two mechanisms differ only when e = 0, which has probability zero.
  • Section 2: No interventionally equivalent SCM with the specified monotonicity and continuous-noise conditions can represent the latent confounder example.The contradiction follows because the transformed noise component would need both a standard-normal distribution and a nonzero mean.
  • Section 2: A counterfactual outcome can remain unidentified even with observational and randomized interventional data when it depends on an unobserved parameter.The query distribution is N(ρc,1 −ρ^2), while the observational and interventional densities do not depend on ρ.
  • Section 3: Perfect interventions can destroy solvability or uniqueness: one intervention yields no solvability, while another yields multiple induced distributions.The original SCM has the unique solution (0,1), but interventions do({1},ξ1) with ξ1 ≠ 0 and do({2},ξ2) with ξ2 > 1 have different failures.
  • Section 4: Counterfactually equivalent SCMs can have different graphs, so graph equality is not necessary for identical counterfactual behavior.Examples M and ˆ M are counterfactually equivalent although G(M) is not equal to G( ˆ M).
  • Section 5: A marginal SCM may exist even when the stated marginalization condition fails, and its graph can be a strict subgraph of the latent projection.Examples D.11 and D.12 show counterfactual equivalence after marginalization and the absence of a directed path present in the original augmented graph.
  • Section 7: Marginal interventional distributions can reveal bidirected-edge structure that differs between SCMs with otherwise similar observed mechanisms.Example D.13 compares SCMs whose second equations use independent E2 versus the shared E1, with their augmented graphs shown separately.

APPENDIX E: PROOFS … Section 8

The appendices provide proofs for the paper’s theoretical results, establishing solvability, uniqueness, measurable solution constructions, and consequences for interventions, marginalization, and simple SCMs. The supplied passages span appendices A–C and main-text Sections 2–5 and 8.

  • APPENDIX E: PROOFS: Appendix E contains proofs of the theoretical results in appendices A–C and the main text, using measure-theoretic results from Appendix F.This passage describes the appendix’s overall purpose rather than a specific theorem.
  • Appendix A: Under componentwise unique solvability, an SCM is uniquely solvable and all solutions share the same observational distribution; related constructions establish loop-wise solvability and compatibility.The appendix also proves path-reduction lemmas for C-d-open and C-σ-open walks and constructs compatible solution functions for induced SCMs.
  • Appendix B: Measurable selection proves solvability with respect to O, while combining solution functions for overlapping ancestral subsets yields unique solvability with respect to their union.The union result uses a measurable mapping assembled from solution functions for A and Ã.
  • Appendix C: For linear SCMs, unique solvability with respect to L is equivalent to invertibility of A_LL, and equivalently to unique solvability within every strongly connected component.The proof uses an upper triangular block form induced by the DAG of strongly connected components.
  • Section 2: Structural minimalization preserves equivalence, perfect interventions remove incoming functional dependencies and graph edges, and twin operations do not create directed cycles under the stated assumptions.These proofs connect the definitions of minimal SCMs, interventions, and twin constructions to their graphical effects.
  • Section 3: A measurable solution exists exactly when the structural equations have solutions for almost every exogenous value, and unique solvability makes the observational distribution the push-forward induced by the solution function.The section also proves that solvability transfers to ancestral subsets and extends appropriately under perfect interventions.
  • Section 4; Section 5: The proofs establish that successive marginalizations compose when the relevant unique-solvability conditions hold, and that observational and interventional equivalences follow under these conditions.The supplied passages show marginalization composition and the equivalence result using Lemma E.4 and Proposition 5.5.
  • Section 8: Simple SCMs are closed under the twin operation, and their resulting twin models are uniquely solvable; consequently, they induce unique observational, interventional, and counterfactual distributions.The supplied Section 8 passages construct measurable solution functions for the twin model and invoke Theorem 3.6 for distributional uniqueness.

APPENDIX F: MEASURABLE SELECTION THEOREMS

Appendix F develops the measure-theoretic foundations used in Appendix E, including standard spaces, analytic sets, and measurable mappings. It establishes two measurable selection theorems, with the second extending selection to product spaces under σ-compact fiber conditions.

  • Measure-theoretic foundations: Standard measurable spaces are spaces isomorphic to Borel spaces of Polish spaces, while standard probability spaces additionally carry a probability measure.Open and closed subsets of R^d and finite sets with their usual complete metric are examples.
  • Measure-theoretic foundations: Measurable sets in standard measurable spaces are analytic, and measurable images and preimages of analytic sets remain analytic.These closure properties support the measurable selection arguments that follow.
  • Measure-theoretic foundations: Analytic subsets of a standard probability space are measurable and can be sandwiched between measurable sets with equal completed-measure values.This result uses universal measurability of analytic sets.
  • Measurable selection theorems: The first measurable selection theorem provides a measurable g:E→X selecting from S for PE-almost every e when S is measurable and its projection misses only a PE-null set.The proof combines analytic uniformization with almost-everywhere replacement by a measurable mapping.
  • Measurable selection theorems: The product-space extension can fail without extra assumptions, but σ-compact fibers yield a measurable g:X×E→Y selecting from S for PE-almost every e and every x.The second theorem applies when the relevant fibers are nonempty and σ-compact outside a PE-null set.
Loading 1611.06221v6…