Source-linked AI summary

Causal Discovery from Heterogeneous/Nonstationary Data with Independent Changes

Biwei Huang, Kun Zhang, Jiji Zhang, Joseph Ramsey, Ruben Sanchez-Romero, Clark Glymour, Bernhard Schölkopf

arXiv:1903.01672v5cs.LGstat.ML

TL;DR

Heterogeneous and nonstationary data complicate causal discovery because causal mechanisms change across domains or time. CD-NOD detects changing modules, recovers skeletons, orients directions using independent distribution changes, and estimates low-dimensional driving forces. The paper reports extensions and experiments on synthetic, task-fMRI, and stock-market data.

  • Problem

    Causal discovery needs methods that identify structure, directions, and mechanism changes when causal models vary across domains or over time.

  • Method

    CD-NOD combines an enhanced constraint-based procedure, distribution-shift-based direction identification, and low-dimensional representation of changing mechanisms.

  • Results

    CD-NOD is reported to recover causal structure, exploit distribution shifts for direction identification beyond restricted functional classes, and perform successfully on synthetic and real-world datasets.

  • Takeaways & Limitations

    Distribution shifts can provide useful causal information, while CD-NOD supports nonparametric analysis without window segmentation and extends to dynamic relations and particular stationary confounders.

  • Takeaways & Limitations

    The framework allows related mechanism changes through particular confounders but assumes pseudo causal sufficiency rather than causal sufficiency for observed variables.

Abstract

from arXiv · show

It is commonplace to encounter heterogeneous or nonstationary data, of which the underlying generating process changes across domains or over time. Such a distribution shift feature presents both challenges and opportunities for causal discovery. In this paper, we develop a framework for causal discovery from such data, called Constraint-based causal Discovery from heterogeneous/NOnstationary Data (CD-NOD), to find causal skeleton and directions and estimate the properties of mechanism changes. First, we propose an enhanced constraint-based procedure to detect variables whose local mechanisms change and recover the skeleton of the causal structure over observed variables. Second, we present a method to determine causal orientations by making use of independent changes in the data distribution implied by the underlying causal model, benefiting from information carried by changing distributions. After learning the causal structure, next, we investigate how to efficiently estimate the "driving force" of the nonstationarity of a causal mechanism. That is, we aim to extract from data a low-dimensional representation of changes. The proposed methods are nonparametric, with no hard restrictions on data distributions and causal mechanisms, and do not rely on window segmentation. Furthermore, we find that data heterogeneity benefits causal structure identification even with particular types of confounders. Finally, we show the connection between heterogeneity/nonstationarity and soft intervention in causal discovery. Experimental results on various synthetic and real-world data sets (task-fMRI and stock market data) are presented to demonstrate the efficacy of the proposed methods.

1. Introduction

CD-NOD addresses causal discovery when mechanisms change across domains or time, using distribution shifts to recover structure, orient directions, and represent changing mechanisms. The framework extends to dynamic relations and particular stationary confounders while avoiding window segmentation.

  • Research questions: CD-NOD targets changing causal models by identifying changing local mechanisms, recovering observed causal skeletons, determining directions, and estimating low-dimensional driving forces.The paper formulates three questions covering skeleton recovery, direction identification, and representation of mechanism changes.
  • Phase II: Independent changes in causal modules provide additional information for causal direction identification, including for general causal mechanisms beyond restricted functional classes.Kernel embeddings of changing conditional distributions measure dependence between mechanism changes in candidate directions.
  • Phase I: The method introduces a surrogate domain or time variable C to unpack distribution shifts into causal representations and detect variables adjacent to C as mechanism-changing.Conditional independences given C match those implied by the true causal structure, under the stated faithfulness assumption.
  • Mechanism changes: The framework estimates an interpretable low-dimensional representation of changing mechanisms and avoids window segmentation for continuously changing causal processes.Kernel Nonstationary Visualization is introduced for representing the driving force of nonstationarity.
  • Extensions and evaluation: The approach extends to time-varying instantaneous and lagged relations, discusses stationary confounders, and combines CD-NOD with constrained functional causal-model methods.Experiments include synthetic data, task-fMRI, Hong Kong stock, and US stock datasets.

2. Problem Definition and Related Work

Changing causal models can make methods designed for fixed distributions produce spurious edges or wrong directions. CD-NOD instead models changing causal modules and uses distribution shifts to recover structure and identify directions.

  • Fixed causal models: Standard causal discovery methods generally assume a fixed causal model and recover structure from conditional independences or constrained functional relationships.Constraint-based methods target Markov equivalence classes, while functional causal models use independence between causes and noise.
  • Changing causal models: When related mechanism changes act like an unobserved quantity, standard constraint-based methods can infer spurious observed edges.In the illustrated setting, the estimated skeleton contains spurious edges V1−V4 and V2−V4.
  • Changing causal models: Fitting a fixed functional causal model to merged distribution-shifted data can destroy cause–noise independence and produce an incorrect causal direction.The example combines mechanisms V2 = 0.3V1 + E and V2 = 0.7V1 + E across two datasets.
  • Related work: Separate domain or sliding-window analyses are possible alternatives, but they require comparison and merging and are unsuitable for continuously changing mechanisms in some settings.Related approaches may also impose linearity or fail when nonstationarity comes from changing noise distributions.
  • CD-NOD: CD-NOD provides a nonparametric, computationally efficient approach that identifies changing causal modules, recovers causal structure, uses distribution shifts for directions, and estimates interpretable change representations.The framework is designed for heterogeneous and nonstationary data rather than fixed causal models.

3. CD-NOD Phase I: Changing Causal Module Detection and Causal Skeleton Estimation

CD-NOD Phase I uses the domain/time index as a surrogate to detect changing causal modules and recover the observed causal skeleton under pseudo causal sufficiency. Conditional-independence tests support asymptotic recovery, while the method allows changing or vanishing causal links and imposes no hard restrictions on functional forms or distributions.

  • Assumptions: Pseudo causal sufficiency permits confounders that are functions of the domain index or smooth functions of time, so their values are fixed within domains or time instances.The paper distinguishes pseudo, stationary, and nonstationary confounders.
  • Assumptions: CD-NOD Phase I allows causal mechanisms and parameters to change across domains or time, including causal links that vanish or appear.The framework focuses initially on contemporaneous relations and extends naturally to time-delayed relations.
  • Assumptions: The local causal process is represented by an SEM with domain/time-dependent parameters, independent disturbances, and optional pseudo confounders.The disturbance is independent of the domain/time index and parents, while changing parameters are specific to each variable.
  • Detection and skeleton recovery: The method treats C as a surrogate for unobserved changing factors and applies conditional-independence tests on V ∪ C to detect changing modules and estimate the causal skeleton.Algorithm 1 builds a complete undirected graph, tests variable–C independences, then tests variable-pair independences conditioned on observed variables and C.
  • Limitations: If a changing module becomes conditionally independent of C given an alternative subset despite dependence given its parents, Algorithm 1 may detect only some changing variables.The paper explicitly limits the detection claim under this failure of the additional assumption.
  • Detection and skeleton recovery: Under the stated assumptions, two observed variables are nonadjacent exactly when they are conditionally independent given some subset of the other variables and C.This criterion provides the asymptotic correctness basis for recovering the causal skeleton.
  • Confounding and guarantees: The independence tests can also identify pseudo confounders behind nonadjacent variables when conditioning on C reveals independence that conditioning only on observed variables does not.The criterion compares dependence under all observed-variable conditioning sets with independence after adding C.
  • Confounding and guarantees: Changing-module detection follows the principle of minimal changes, yielding as few edges involving C as possible while connecting C to variables with changing mechanisms.This graphical property follows from faithfulness on the augmented graph and edge minimality.

4. CD-NOD Phase II: Distribution Shifts Benefit Causal Direction Determination

CD-NOD uses distribution shifts to orient causal edges by testing independent changes in causal modules, extending invariance-based reasoning to general mechanisms and avoiding window segmentation.

  • Generalization of Invariance: Adding the surrogate variable C helps recover the observed causal skeleton and orient edges incident to C-specific variables.C-specific variables are those adjacent to C after skeleton discovery; standard unshielded-triple rules handle cases where the neighboring variable is not adjacent to C.
  • Generalization of Invariance: The invariance-based procedures include earlier methods as special cases, while independent changes apply more generally when both cause and conditional-effect distributions vary.If only one of the two modules changes, independence is immediate; the paper notes that this restriction is less generic.
  • Generalization of Invariance: Invariance tests are conditional-independence tests involving C: Vi ⊥⊥ C | S expresses that P(Vi | S,C) is unchanged across domains or times.When S is empty, the test reduces to marginal independence between Vi and C, or homogeneity of P(Vi).
  • Independent Changes: When both P(cause) and P(effect | cause) change, independent causal-module changes distinguish the causal direction because reverse-factorization modules are generally dependent.The method compares dependence between P(V1) and P(V2|V1) against dependence between P(V2) and P(V1|V2).
  • Independent Changes: CD-NOD extends HSIC with kernel embeddings of nonstationary conditional distributions to measure dependence between causal modules.The embedding and Gram matrices are estimated over the whole dataset, without sliding-window segmentation.
  • Multiple Variables: For multiple variables, Algorithm 3 uses minimal deconfounding sets and potential deconfounding sets before applying independent-change orientation procedures.These sets remove effects from common causes or account for possible confounding behind adjacent variables.

A. Let Z(1)

For each unoriented adjacent pair, CD-NOD defines minimal deconfounding sets consisting of sets adjacent to one endpoint that remove effects from common causes.

  • A. Let Z(1): A minimal deconfounding set Z(1)_lk is a set of variables adjacent to Vl used for the pair Vk and Vl.The section introduces Z(1)_lk for each adjacent pair as part of Algorithm 3's orientation procedure.

B. Let Z(2)

CD-NOD supplements minimal deconfounding sets with potential sets and provides identifiability conditions, algorithmic outputs, and a corresponding equivalence-class boundary.

  • B. Let Z(2): Algorithm 3 orients edges in a partially oriented graph by iterating over unoriented adjacent pairs and outputting directed edges when its conditions hold.The procedure operates on variables whose causal modules change and returns a graph with edges among those variables oriented.
  • B. Let Z(2): The worked example applies Algorithm 3 to output V3 →V1, V3 →V4, V1 →V2, and V4 →V2.These are the resulting directions reported for the example's successive orientation steps.
  • Identifiability: Under Assumptions 1–3, an adjacent pair is identifiable when it lies in a V-structure, changing modules are independent when both change, or an incident edge touches only one endpoint.If none of these conditions holds, the direction may remain unidentified even though the causal skeleton is identifiable.
  • Identifiability: The whole causal graph is identifiable when every adjacent pair satisfies at least one of the three direction-identifiability conditions.Otherwise, CD-NOD returns an equivalence class sharing the causal skeleton and all directions identifiable under those conditions.
  • Computational Complexity: Kernel-based independence and dependence measures in both CD-NOD phases have O(N^3) computational complexity.The PC search additionally depends on the number of observed variables m and maximal degree k.

5. CD-NOD Phase III: Nonstationary Driving Force Estimation

Phase III estimates a low-dimensional representation of changing causal modules, called the nonstationary driving force, using kernel embeddings and KPCA.

  • Kernel Representation: Kernel embedding represents changing conditional distributions without requiring prior knowledge of which causal-model parameters change.This provides a flexible nonparametric alternative to estimating known parameters across values of C.
  • Driving Force Definition: The nonstationary driving force λ_i(C) is a low-dimensional representation of changes in P(V_i | PA_i,C).A nonlinear mapping h_i summarizes the conditional distribution along C; λ_i(C) is constant when the module does not change.
  • Kernel Nonstationary Visualization: KNV applies KPCA to the estimated kernel embedding to capture variability across domains or time and obtain the driving-force components.The embedding derived from Proposition 1 is used as input, and principal components summarize its variation across C.
  • Kernel Nonstationary Visualization: The kernel trick lets KNV work directly with N × N kernel matrices rather than explicitly learning a high-dimensional embedding for each C.The driving force is estimated through eigenvalue decomposition, with the first few eigenvectors optionally retained to capture most variance.
  • Algorithm 4: Algorithm 4 takes observations of X and Y, computes a kernel Gram matrix, applies KPCA, and outputs an estimate of λ̂(C).Linear or Gaussian kernels can be used for the Gram matrix.

6. Extensions of CD-NOD

CD-NOD is extended to dynamic systems with time-varying instantaneous and lagged relations, and to settings with stationary confounders. Distribution shifts can help orient some causal relations, but independent changes in both adjacent causal modules may leave directions unresolved.

  • With Time-Varying Instantaneous and Lagged Causal Relationships: CD-NOD recovers both time-varying instantaneous and lagged causal relationships by reorganizing multivariate time series into a unit causal graph.The reorganization creates m*(P + 1) variables and constrains future variables from causing past variables.
  • With Time-Varying Instantaneous and Lagged Causal Relationships: Algorithm 5 detects changing modules and recovers lagged and instantaneous skeletons using independence tests involving the surrogate variable C.Lagged directions follow the rule that past causes future, while instantaneous directions use the direction-recovery procedures from Algorithms 2 and 3.
  • With Stationary Confounders: With stationary confounders, distribution shifts can distinguish V1 →V2 from V1 ←V2 when only V1’s causal module changes.The two directions induce different independence relations involving C, V1, and V2.
  • With Stationary Confounders: When both adjacent causal modules change independently under a stationary confounder, the two directions have the same observed and module-level independences and cannot be distinguished without further information.The dependence between P(V1) and P(V2|V1), and symmetrically between P(V2) and P(V1|V2), prevents orientation by this information alone.
  • With Stationary Confounders: If CD-NOD’s identifiability conditions fail, causal directions remain undetermined and may require constrained functional causal models as additional information.The paper discusses linear non-Gaussian and nonlinear additive-noise models as possible extensions.

7. Relations between Heterogeneity/Nonstationarity and Soft Intervention in Causal Discovery

The paper connects heterogeneous or nonstationary causal mechanisms with soft interventions, interpreting distribution shifts across domains or time as naturally occurring interventions. Its framework broadens standard soft-intervention assumptions by detecting changes automatically and allowing multiple affected variables and changing structures.

  • Soft Intervention: A soft intervention changes a variable’s causal module without making it independent of its causes or breaking incident edges.The intervention indicator is known, binary, exogenous, and directly causes only the intervened variable.
  • Connection to Heterogeneity/Nonstationarity: A changing conditional distribution P(Vi|PAi,C) across values c and c′ corresponds to a soft intervention on Vi.The paper explicitly identifies heterogeneity or nonstationarity with soft interventions carried out by nature.
  • Broader Framework: Unlike the compared intervention framework, CD-NOD can automatically detect changes that are discrete across domains or continuous over time.The earlier framework assumes a known two-state intervention indicator, whereas CD-NOD permits unknown change locations and forms.
  • Broader Framework: CD-NOD allows a changing quantity to affect several variables through pseudo confounders and permits causal edges to vanish or appear across domains or time periods.These extensions go beyond interventions restricted to a known direct cause of a single variable and fixed causal structure.

8. Experimental Results

Across synthetic and real-world experiments, CD-NOD generally achieved the strongest causal-structure recovery and identified changing mechanisms and driving forces. Applications to fMRI and stock returns produced interpretable structures, temporal changes, and sector-level patterns.

  • Setting 1: CD-NOD achieved the best F1 score and precision for causal-skeleton recovery on both heterogeneous and nonstationary data across all sample sizes.Its recall was similar to competing methods, while the original constraint-based, IB, and MC methods had lower precision in these settings.
  • Setting 1: CD-NOD’s detected changing causal modules performed well, with F1 score tending to increase as sample size grew in both heterogeneous and nonstationary scenarios.The evaluation used F1 score, precision, and recall.
  • Setting 1: CD-NOD produced the best F1 score for the whole causal graph and for inferred causal directions across the tested heterogeneous and nonstationary scenarios.The window-based method performed lower, while IB and MC were substantially weaker because they target linear systems.
  • Task-fMRI: KNV gave the best recovery of changing components in the reported comparisons, while recovered fMRI driving forces correlated with task states by 0.22, 0.49, 0.41, and 0.47 for LDLPFC, LIPL, LSGA, and RIPL.The fMRI driving-force changes matched intervals associated with the resting state without input stimuli.
  • Stock Returns: Stock-market graphs showed denser intra-sector connections, with energy, finance, public utilities, and basic industries more likely to cause stocks in other sectors.Change points in several stocks aligned with critical periods of the 2008 financial crisis, and highly central stocks often showed changes around both reported dates.
  • Stock Returns: In the stock experiments, 63 edges produced by the original PC algorithm were removed after accounting for nonstationarity, and 37 of 80 stocks had causal modules that changed over time.The removed edges included 19 within-sector and 44 between-sector edges; changing modules were concentrated in energy, finance, basic industry, and consumer service.

9. Discussion and Conclusions

CD-NOD uses conditional independence and independent changes to discover causal structure and characterize changing mechanisms in heterogeneous or nonstationary data. Experiments report improved synthetic-data performance and meaningful findings on task-fMRI and stock data, while several scope limitations remain open.

  • Contributions: CD-NOD locates changing causal modules, estimates the observed causal skeleton, determines directions from distribution shifts, and represents mechanism changes in low dimensions.The framework combines conditional independence with independent changes and also addresses stationary confounders.
  • Empirical findings: CD-NOD substantially improves synthetic-data F1 score and precision relative to other causal-discovery methods.On task-fMRI, it reveals information flows and changing causal influences; on stock data, learned relationships agree with background knowledge and driving forces align with the 2008 financial crisis.
  • Interpretation: Distribution shifts provide useful causal-discovery information because the causal model compactly describes how the joint distribution changes.The paper connects this perspective to domain adaptation and prediction in nonstationary environments.
  • Open questions: Future work must address changing causal directions, reduced conditional-independence-test power under distribution shift, and nonstationary confounders.The authors specifically mention direction flips, feedback loops, and broader confounder types.

Appendix A. Proof of Theorem 1

The proof establishes that conditional independences involving the surrogate variable C characterize nonadjacency in the augmented graph and therefore recover nonadjacency in the original graph. It relies on deterministic functions of C and faithfulness.

  • Implications of the SEM: Acyclic SEMs imply that every observed variable is a function of mechanism-change terms and C, making relevant conditional independences hold given C.The proof uses that the functions g_l(C) and parameters θ_m(C) are deterministic functions of C.
  • Reverse implication: Faithfulness converts conditional independence in the augmented graph into nonadjacency for the corresponding observed variables.This implication is used when the conditioning set excludes C.
  • Reverse implication: When C is included in the conditioning set, factorization and the proof’s conditional-independence identities likewise establish nonadjacency.The argument considers S = V_ij ∪ {C} and uses the displayed factorization together with subsequent substitutions.
  • Conclusion: Therefore, two observed variables are nonadjacent if and only if they are conditionally independent given some subset of the other observed variables together with C.This is the theorem-level conclusion of the appendix proof.

Proof

This proof section defines tensor-product feature representations and kernel mean embeddings used to estimate a conditional distribution quantity. The displayed transformation rewrites the estimator in feature-space matrix form.

  • Estimator transformation: The proof then rewrites equation (30) after applying tensor-product associativity and a feature-space inverse relation.The resulting expression is presented as the transformed estimator.
  • Feature representation: The proof uses tensor-product feature maps for joint representations of X and C.It defines φ⊗(X,C) as φ(X) ⊗ φ(C).
  • Feature representation: The notation Φ_y, Φ_x, Φ_x,c, and Φ_x,cn collects feature vectors across observations for the corresponding variables or paired inputs.These matrices support the subsequent estimator construction.

Appendix C. Proof of Theorem 2

Theorem 2’s proof gives conditional-independence criteria for orienting an adjacent pair of variables. It covers collider patterns, changes affecting one or both variables, and neighboring-edge configurations.

  • Collider-based orientation: Collider configurations orient an adjacent pair by testing which endpoint is conditionally independent of a third influencing variable.The two alternatives distinguish Vi → Vj ← Vk from Vj → Vi ← Vk.
  • Single changing influence: If only one endpoint is influenced by a change, conditional independence between that endpoint and C determines the orientation relative to the adjacent variable.The conditioning set either excludes or includes the opposite endpoint, yielding opposite directions.
  • Independent changes: Independent changes at both endpoints orient the edge by comparing independence between a marginal distribution and the other variable’s conditional distribution.The deconfounding set Z is used in both directional alternatives.
  • Neighboring-edge orientation: Additional conditional-independence rules apply when a neighboring edge is incident to one endpoint but not the other.These rules use separation sets that exclude or contain the opposite endpoint.
  • Identifiability: The direction between Vi and Vj is identifiable whenever at least one of the stated conditions is satisfied.This summarizes the proof’s orientation criterion.
Loading 1903.01672v5…