Source-linked AI summary

Causal Representation Learning from General Environments under Nonparametric Mixing

Ignavier Ng, Shaoan Xie, Xinshuai Dong, Peter Spirtes, Kun Zhang

arXiv:2604.23800v1cs.LGstat.ML

TL;DR

Causal representation learning remains difficult because latent variables and causal structure are generally unidentifiable without restrictive assumptions. This paper formalizes general environments and shows that nonparametric mixing with nonlinear latent causal models can still yield full latent-DAG recovery up to minor indeterminacies.

  • Problem

    Latent variables and causal structure remain difficult to recover from observations, while existing approaches often rely on restrictive assumptions about distribution changes or model structure.

  • Method

    The paper formalizes general environments and leverages sufficient changes in causal mechanisms up to third-order derivatives under nonparametric mixing and nonlinear latent causal models.

  • Results

    The method fully recovers the latent DAG and identifies latent variables up to permutation and mixing with surrounding parents for additive and heteroscedastic noise models.

  • Takeaways & Limitations

    General environments can support latent-DAG recovery under less restrictive distribution-change assumptions while narrowing latent-variable ambiguity to surrounding parents.

  • Takeaways & Limitations

    The identifiability conditions exclude latent causal mechanisms that are linear across all environments, and second-order approaches may recover only a Markov equivalence class.

Abstract

from arXiv · show

Causal representation learning aims to recover the latent causal variables and their causal relations, typically represented by directed acyclic graphs (DAGs), from low-level observations such as image pixels. A prevailing line of research exploits multiple environments, which assume how data distributions change, including single-node interventions, coupled interventions, or hard interventions, or parametric constraints on the mixing function or the latent causal model, such as linearity. Despite the novelty and elegance of the results, they are often violated in real problems. Accordingly, we formalize a set of desiderata for causal representation learning that applies to a broader class of environments, referred to as general environments. Interestingly, we show that one can fully recover the latent DAG and identify the latent variables up to minor indeterminacies under a nonparametric mixing function and nonlinear latent causal models, such as additive (Gaussian) noise models or heteroscedastic noise models, by properly leveraging sufficient change conditions on the causal mechanisms up to third-order derivatives. These represent, to our knowledge, the first results to fully recover the latent DAG from general environments under nonparametric mixing. Notably, our results match or improve upon many existing works, but require less restrictive assumptions about changing environments.

1 Introduction

Causal representation learning seeks latent causal variables and structure from indirect, high-dimensional observations, but remains difficult because nonlinear representations are generally unidentifiable and causal discovery is itself challenging. Existing approaches often rely on restrictive assumptions about distribution changes, motivating a general-environment setting that supports softer assumptions while enabling full latent-DAG recovery.

  • Motivation: Causal representation learning recovers latent causal variables and structure from observations that may be low-level, indirect, and high-dimensional.Examples include image pixels, linguistic tokens, and gene expressions.
  • Challenges: Nonlinear ICA is unidentifiable without additional assumptions, and causal representation learning additionally inherits the difficulty of causal discovery.Different latent representations can explain the same observed data without matching the underlying generating process.
  • Prior approaches: Existing multi-environment approaches commonly assume hard, single-node, coupled, or counterfactual distribution changes, alongside parametric or graphical constraints.These assumptions define major lines of prior causal representation learning research.
  • Contribution: General environments address realistic violations of these assumptions while permitting nonparametric mixing and nonlinear latent causal models.The paper claims full latent-DAG recovery and latent-variable identification up to minor indeterminacies.
  • Contribution: Third-order derivative conditions extract causal ordering information, enabling full latent-DAG recovery under additive Gaussian or heteroscedastic noise models.The paper validates its identifiability theory through simulation studies.

2 Problem Setting

The paper models observed variables as generated from latent variables through an unknown nonparametric mixing function, with latent variables sharing an unknown DAG across environments. It seeks recovery from multi-environment samples while allowing distribution changes represented by hard or soft interventions.

  • Generative model: Observed variables X are generated from latent variables Z through an unknown nonparametric mixing function g.The figure depicts X as observable and Z as latent.
  • Generative model: Across environments, latent variables follow structural equation models with the same unknown DAG G_Z.The environment index u specifies settings in which causal mechanisms may vary.
  • Estimation objective: The objective is to estimate latent variables Z and the latent DAG G_Z up to minor indeterminacies from samples of X across multiple environments.Environment changes may arise from heterogeneous data or nonstationary time series.
  • Environment changes: The framework represents environment changes through interventions that may be hard or soft.Soft interventions modify causal mechanisms without necessarily eliminating parent influence, whereas hard interventions eliminate targeted parent dependencies.
  • Identification: Observational equivalence compares estimated and true models when they generate matching observed-variable distributions across the given environments.The compared models include a mixing function, latent distribution, and latent DAG.

3 Desiderata for CRL from General Environments

The paper defines general environments through desiderata intended to make causal representation learning more realistic: interventions need not be hard, their targets need not be known, and multiple nodes may change simultaneously. These desiderata respond to the unidentifiability of nonlinear ICA and the restrictive intervention assumptions used in prior work.

  • Motivation: Nonlinear ICA is a special case of causal representation learning with an empty latent DAG, yet remains unidentifiable without assumptions.Causal representation learning is therefore more complex than this already difficult special case.
  • General environments: General environments broaden prior settings by formalizing realistic assumptions about changing distributions.The paper presents three desiderata defining this broader class.
  • Desideratum 1: Interventions may be soft or hard rather than requiring hard interventions.Soft interventions modify targeted causal mechanisms without fully removing parent influence and occur in examples such as RNA interference.
  • Desideratum 2: The setting assumes no prior knowledge of intervention targets or their coupling pattern across environments.Such knowledge is often infeasible when latent variables are unknown.
  • Desideratum 3: Interventions may affect single nodes or multiple nodes across environments.Simultaneous changes in multiple latent causal mechanisms motivate the multi-node allowance.
  • Intervention information: Knowing which single-node interventions share targets is equivalent to knowing intervention targets up to variable permutation.This result formalizes the relationship between intervention coupling information and target identification.

4 CRL with Latent Additive Noise Models

The paper develops identifiability theory for causal representation learning with latent additive noise models, using third-order derivative information to recover causal directions. Under sufficient change conditions, it recovers the latent DAG and identifies latent variables up to permutation and mixing with surrounding parents.

  • Second-order derivative conditions can recover the latent Markov network but generally cannot determine all causal directions, limiting identification to a Markov equivalence class.
  • Third-order derivatives can reveal causal directions in latent additive noise models, enabling recovery of the latent DAG under sufficient change conditions.
  • The approach uses sink-node properties to identify a sink from the latent Markov network, then infers its parents as neighboring nodes and iteratively repeats this process.
  • Compared with prior Markov-network identification, the result narrows latent-variable ambiguity from intimate neighbors to surrounding parents, a subset of those neighbors.
  • Theorem 1 establishes that the recovered DAG is identical to the true latent DAG up to permutation, while each recovered latent variable depends only on itself and its surrounding parents.
  • The sufficient-change assumptions support soft or hard, single-node or multi-node interventions without requiring prior knowledge of intervention targets or their coupling patterns.

5 CRL with Latent Heteroscedastic Noise Models

The section establishes identifiability for latent heteroscedastic noise models under general environments and nonparametric mixing. Third-order changes in causal mechanisms support recovery of the latent DAG and latent variables up to parent-related indeterminacies.

  • Model: Latent heteroscedastic noise models allow noise terms with nonconstant variances while retaining Gaussian noise and third-order differentiable mechanisms.The estimation method is nearly identical to that for additive noise models.
  • Identifiability theory: Third-order derivatives of sink-node mechanisms provide information for inferring the latent DAG.For a sink node, the relevant derivative expression vanishes for variables that are not its parents.
  • Identifiability theory: Sufficient changes across environments are imposed through linearly independent changes in the mechanism-derived vectors w(Z′,u).The condition requires L + 1 environment values for each relevant Z′.
  • Identifiability theory: Under the stated assumptions and faithfulness, Algorithm 1 recovers the latent DAG exactly after a permutation of estimated variables.The theorem applies to the latent HNM data-generating process.
  • Identifiability theory: Each permuted estimated latent variable is solely a function of the corresponding variable and a subset of its surrounding parents.This is the theorem’s permitted indeterminacy for latent-variable identification.

6 Simulation Studies

Simulation studies test the theory on three latent variables generated from predefined DAGs across environments with changing noise statistics and nonlinear mechanisms. The method recovers the correct structures and exhibits the predicted variable correspondences.

  • Simulation setup: The simulations use three latent variables with randomly varying noise means and variances across environments.Latent mechanisms and the injective mixing function are implemented with two-layer MLPs using LeakyReLU activations.
  • Simulation setup: The mixing function transforms latent variables into observations through an injective two-layer MLP with an orthogonal weight matrix.The simulation procedure satisfies the paper’s general-environment desiderata.
  • Results: The method successfully recovers the correct structure for both evaluated latent DAGs.Figure 2 compares estimated and true latent variables for a chain and a two-parent DAG.
  • Results: Z1 and Z2 show clear one-to-one correspondences with their estimated counterparts, indicating component-wise identification.The correspondence for Z3 is weaker because its estimate depends on Z3 and its surrounding parent Z2.

7 Conclusion

The paper formalizes general environments for causal representation learning and proves identifiability under nonparametric mixing with nonlinear latent causal models. Simulations validate the theory, while future work targets real-world applications and weaker noise assumptions.

  • Conclusion: General environments permit full latent-DAG recovery and latent-variable identification up to minor indeterminacies with nonparametric mixing.The result covers additive Gaussian and heteroscedastic noise models.
  • Conclusion: Sufficient changes in causal mechanisms up to third-order derivatives provide the information needed to recover causal ordering and the complete latent DAG.
  • Conclusion: Simulation studies validate the identifiability theory under the proposed setting.
  • Conclusion: Future work includes applying the method to real-world data and relaxing the Gaussian-noise assumption.

Checklist

The checklist records stated assumptions, proofs, reproducibility materials, training details, and asset or human-subject disclosures. The related-work material contrasts existing CRL identifiability results and intervention assumptions.

  • Theoretical reporting: The paper states that assumptions and complete proofs for theoretical results are provided, with explanations of the assumptions included.
  • Reproducibility and assets: Code, data, reproduction instructions, and training details are reported as included, while several infrastructure and asset fields are marked not applicable.
  • Human subjects: The paper reports no crowdsourcing or human-subject research, so participant instructions, risks, wages, and compensation are marked not applicable.
  • Related work: Existing multi-environment CRL results commonly rely on hard or soft interventions and may impose restrictive assumptions on environments or mixing functions.
  • Related work: The paper positions its result as allowing nonlinear structural equations and nonparametric mixing, unlike cited alternatives requiring linear mixing or yielding only the moral graph.
  • Related work: Prior ordering-based causal discovery methods search over orderings or use functional and noise assumptions to identify nodes sequentially.
  • Related work: Table 2 compares multi-environment CRL identifiability results under hard or soft interventions and summarizes their key assumptions and outcomes.

B Proof of Proposition 1

Proposition 1 shows that, for coupled single-node interventions, knowing which interventions share targets determines the intervention targets up to variable permutation.

  • Partitioning interventions by shared targets identifies groups with common intervention targets and separates groups targeting different variables.The proof constructs a partition whose blocks contain interventions sharing one target.
  • A suitable permutation assigns the intervention groups to latent variables, although the correct permutation is unknown.Therefore, target assignments are recoverable only up to variable permutation.

C Proofs of Lemmas 1 and 2

The lemmas analyze sink nodes in the latent DAG through derivative constraints on the latent log-density, using the Markov property and sink-node structure.

  • For a sink node, the proof applies the Markov property to simplify the latent distribution and obtain the required derivative relation.The displayed relation involves a third-order derivative of the log-density.
  • Lemma 2 similarly considers a sink node under the data-generating process in Eq. (5) and states a zero condition for variables outside its parent set.The proof again uses the Markov property together with the sink-node assumption.

D Proof of Theorem 1

Theorem 1 establishes that Algorithm 1 recovers the latent DAG exactly up to permutation and identifies each latent variable using only itself and surrounding parents.

  • After a permutation of the estimated variables, the recovered DAG Ĝ_Zπ is identical to the true latent DAG G_Z.This is the theorem’s first identifiability statement.
  • Each permuted estimated latent variable Ẑπ(i) is solely a function of a subset of the true variable Zi and its surrounding parents.The result narrows the remaining variable-level ambiguity to the variable and its DAG-neighborhood parents.
  • The proof uses a diffeomorphic transformation between estimated and true latent variables, then selects a permutation with nonzero diagonal Jacobian entries.This permutation supports the inductive comparison of the estimated and true causal structures.
  • The argument combines faithfulness, Markov-network structure, and progressively refined variable-function relationships to establish both DAG and latent-variable identifiability.Supporting lemmas relate Markov networks, moral graphs, and diffeomorphisms on variable subsets.

D.4 Proof of Proposition 2

Proposition 2 proves inductively that Algorithm 1 preserves causal edges and progressively restricts latent-variable dependencies while removing sink nodes.

  • Causal-edge preservation: At each iteration, estimated edges into the processed sink nodes match the corresponding true edges under a permutation α.The induction establishes this correspondence first for the newly processed sink and then for all previously processed sinks.
  • Sink identification: The proof identifies estimated sink nodes by taking third-order derivatives and exploiting linearly independent changes in causal-mechanism derivatives across environments.Differences across environment values isolate the relevant derivative coefficients.
  • Causal-edge preservation: Minimal Markov networks and faithfulness transfer adjacency relations between estimated and true variables, allowing sink-node edges to be oriented correctly.For a sink, moral-graph adjacency corresponds to directed parent-to-sink edges.
  • Latent-variable dependencies: The induction shows that each estimated latent variable is a function only of its corresponding true variable and a controlled set of surrounding variables.The dependency set is progressively narrowed through Markov-network neighborhoods and the induction hypothesis.
  • Inductive reduction: Removing the identified sink preserves a diffeomorphic transformation on the remaining variables, completing the inductive step.The argument shows that other variables cannot depend on the removed estimated sink and then applies the subset-diffeomorphism lemma.
  • Inductive reduction: The induction starts from a diffeomorphic transformation between all estimated and true latent variables and proceeds through the remaining n−1 iterations.The base case relies on both mixing functions being diffeomorphisms onto their images.

E Proof of Theorem 2

Theorem 2 establishes identifiability for latent heteroscedastic noise models: Algorithm 1 recovers the latent DAG exactly up to permutation and identifies each latent variable up to a function of itself and surrounding parents.

  • E Proof of Theorem 2: Under the theorem’s assumptions, Algorithm 1 outputs estimated latent variables and a graph satisfying the identifiability guarantees.The proof is stated to follow the same route as Theorem 1, using Propositions 2 and 3.
  • E Proof of Theorem 2: The recovered graph Ĝ_Zπ is identical to the true latent graph G_Z after a permutation π of the estimated latent variables.
  • E Proof of Theorem 2: Each permuted estimated variable Ẑπ(i) is solely a function of a subset of Z_i and its surrounding parents in the true graph.
  • E Proof of Theorem 2: The heteroscedastic-noise proof differs from Theorem 1 in the third-order derivative used in Eq. (7).
  • F.2 Selecting Hyperparameters: Model selection can choose the number of latent variables using the evidence lower bound (ELBO) loss.The empirical discussion recommends selecting the latent-variable count according to ELBO loss.
  • F.2 Selecting Hyperparameters: Choosing sparsity requires balancing reconstruction and graph density: overly sparse structures reconstruct poorly, whereas overly dense structures add unwanted edges and increase KL divergence.The supplied figure discussion gives λ1 = 0.01 as achieving an ELBO loss as good as λ1 = 0 while using fewer edges.
Loading 2604.23800v1…