Source-linked AI summary

On the Identifiability of the Post-Nonlinear Causal Model

Kun Zhang, Aapo Hyvarinen

arXiv:1205.2599v1stat.MLcs.LG

TL;DR

The paper addresses incomplete understanding of PNL model identifiability and the difficulty of extending it beyond two variables. It systematically analyzes the two-variable case and develops a Markov-equivalence-based procedure for larger systems, showing general identifiability in two variables and avoiding exhaustive causal-structure search in larger systems.

  • Problem

    The PNL model's identifiability and its application to causal discovery with more than two variables were not adequately addressed.

  • Method

    The paper analyzes two-variable PNL identifiability and applies the model to structures in the Markov equivalent class, testing disturbance independence from direct causes for each variable.

  • Results

    The PNL model is generally identifiable in the two-variable case, with non-identifiable situations characterized, while the multi-variable procedure avoids exhaustive search over all causal structures.

  • Takeaways & Limitations

    PNL can distinguish causes from effects in most two-variable cases and can recover larger causal structures without directly testing every possible DAG.

  • Takeaways & Limitations

    Nonparametric conditional-independence tests may be unreliable with many conditioning variables because of the curse of dimensionality.

Abstract

from arXiv · show

By taking into account the nonlinear effect of the cause, the inner noise effect, and the measurement distortion effect in the observed variables, the post-nonlinear (PNL) causal model has demonstrated its excellent performance in distinguishing the cause from effect. However, its identifiability has not been properly addressed, and how to apply it in the case of more than two variables is also a problem. In this paper, we conduct a systematic investigation on its identifiability in the two-variable case. We show that this model is identifiable in most cases; by enumerating all possible situations in which the model is not identifiable, we provide sufficient conditions for its identifiability. Simulations are given to support the theoretical results. Moreover, in the case of more than two variables, we show that the whole causal structure can be found by applying the PNL causal model to each structure in the Markov equivalent class and testing if the disturbance is independent of the direct causes for each variable. In this way the exhaustive search over all possible causal structures is avoided.

1 INTRODUCTION

The paper motivates the PNL causal model as a flexible but insufficiently characterized approach for distinguishing causes from effects. It investigates identifiability in two variables and proposes a reduced-search strategy for more than two variables.

  • 1 INTRODUCTION: Functional causal models can recover causal relations when well specified, but misspecification may produce misleading results.The paper therefore seeks models that are general enough to approximate data-generating processes while remaining identifiable.
  • 1 INTRODUCTION: The PNL model incorporates nonlinear causal effects, independent disturbances, and invertible measurement distortion in observed variables.It contains additive noise models as a special case when the post-nonlinear distortion is absent.
  • 1 INTRODUCTION: PNL achieved correct causal directions on all eight datasets in the Cause-effect pairs task.The paper attributes part of this performance to allowing measurement distortion, which is frequently encountered in practice.
  • 1 INTRODUCTION: The paper systematically studies two-variable identifiability under unbounded disturbance-density support and reports that the model is generally identifiable.It also enumerates non-identifiable cases and uses simulations to illustrate some of them.
  • 1 INTRODUCTION: For more than two variables, the proposed approach tests PNL models only across the Markov equivalent class rather than searching all causal structures.It tests whether each disturbance is independent of the associated direct causes, reducing the search space and avoiding mutual-independence tests involving more than two variables.

2 IDENTIFIABILITY

The paper investigates when the two-variable PNL causal model is identifiable by analyzing the consequences of assuming both causal directions. It shows that non-identifiability is confined to enumerated distributional and functional cases, yielding practical sufficient conditions for identifiability.

  • Two-variable setup: The two-variable analysis assumes the PNL model holds in both directions and derives strong constraints on the involved distributions and functions.The argument proceeds by contradiction under differentiability and support assumptions.
  • Non-identifiability cases: Special non-identifiable cases impose restrictive forms, including Gaussian disturbances with linear h or log-mix-lin-exp densities.When e2 is Gaussian, h must be linear and t1 must also be Gaussian; when h is linear, the paper lists Gaussian and log-mix-lin-exp alternatives.
  • Non-identifiability cases: Theorem 8 enumerates five situations in which the two-variable PNL causal model is not identifiable.Situation I is the linear Gaussian case; Situations II–V are additional cases identified by the paper.
  • Identifiability conditions: If the disturbance density is neither Gaussian, log-mix-lin-exp, nor a generalized mixture of two exponentials, the PNL model is identifiable under A1 and A2.The criterion depends on the disturbance and transformed-cause distributions rather than directly on the observed-variable distributions.
  • Identifiability conditions: If the nonlinear cause-effect function f1 is non-invertible, the PNL causal model is identifiable under A1 and A2.This is stated as Corollary 10 and is presented as an intuitively appealing result.

3 NONLINEAR ICA-BASED IDENTIFICATION METHOD

The paper identifies a candidate causal direction by estimating a disturbance representation and testing its independence from the proposed cause. Testing both directions distinguishes the causal relation when exactly one hypothesis is supported.

  • Direction testing: Under x1 →x2, the disturbance e2 can be estimated as ˆe2 = l2(x2) −l1(x1) and tested for independence from x1.The functions l1 and l2 are learned by minimizing mutual information between x1 and ˆe2.
  • Optimization: The estimation objective is equivalent to maximizing E log pˆe2(ˆe2) + E log |l′2(x2)|.Terms independent of l1 and l2 are omitted from the optimization.
  • Optimization: Multi-layer perceptrons represent l1 and l2, whose parameters are learned with gradient-based methods.A statistical independence test then evaluates whether the estimated disturbance is independent of x1.
  • Direction testing: Testing both x1 →x2 and x2 →x1 yields a directional decision when exactly one hypothesis holds.If neither holds, there is no PNL relation; if both hold, the model cannot distinguish cause from effect without additional information.

4 MORE THAN TWO VARIABLES

For more than two variables, the paper avoids exhaustive search by first finding the d-separation equivalent class and then testing PNL compatibility within that class. The approach is theoretically tied to independence of disturbances from their parents.

  • Motivation: Brute-force evaluation of all DAGs is computationally costly and does not scale as the number of variables increases.It also requires difficult mutual-independence tests across many estimated disturbances.
  • Two-step procedure: The proposed procedure first finds the d-separation equivalent class using conditional-independence methods.This restricts subsequent PNL testing to causal structures within the equivalent class.
  • Two-step procedure: For each structure in that class, the method estimates disturbances and checks whether each disturbance is independent of its variable’s parents.This avoids exhaustive search over all causal structures and high-dimensional mutual-independence tests.
  • Theoretical guarantee: The disturbances are mutually independent if and only if the causal Markov condition holds and each disturbance is independent of its parents.This theorem supplies the validity criterion for evaluating candidate DAGs with the PNL model.
  • Limitations: Nonparametric conditional-independence tests may be unreliable with many conditioning variables because of the curse of dimensionality.Handling nonlinear causal effects and non-Gaussian variables remains outside the paper’s scope and is left for future work.

5 SIMULATIONS

Simulations verify that specific non-identifiable situations predicted by the theory can occur, including cases where both causal directions explain the observed data.

  • 5.1 ON SITUATION II IN TABLE 1: In Situation II, independence tests accepted both candidate causal directions, so the PNL model could not distinguish cause from effect.The tests accepted independence for (t1, e2) and (z2, e1) at α = 0.05.
  • 5.1 ON SITUATION II IN TABLE 1: The Situation II simulation used h(t1) = −t1 and 2000 samples to construct variables t1 and e2 before testing the reverse representation.The variables t1 and e2 were generated independently, and z2 and e1 were then obtained through the reverse-direction transformation.
  • 5.1 ON SITUATION II IN TABLE 1: The nonlinear ICA-based identification method also accepted independence for both estimated directions in observed data generated under Situation II.Both (x1, ê2) and (x2, ê1) accepted the independence hypothesis, allowing both estimated PNL models to explain the data.
  • 5.2 ON SITUATION V IN TABLE 1: For Situation V, the study numerically solved the ODE system to obtain the density and transformation functions used to generate simulation data.The resulting curves and densities were plotted, and RBF networks learned the densities before random sampling.
  • 5.2 ON SITUATION V IN TABLE 1: In the Situation V simulation, independence was accepted under both hypotheses, so both causal directions explained the data and the PNL model was not identifiable.The result was reported for the independence tests summarized in Table 3.

6 CONCLUSION

The paper concludes that the PNL causal model is generally identifiable for two variables, while providing a Markov-equivalence-based procedure for larger systems.

  • 6 CONCLUSION: The two-variable PNL model is generally identifiable and can therefore distinguish cause from effect, with all particular non-identifiable situations reported in Theorem 8.Some of those situations were verified and illustrated by simulations.
  • 6 CONCLUSION: For more than two variables, the whole causal structure can be found by applying the PNL model to the Markov equivalent class and testing disturbance-parent independence.This avoids applying the model directly to every possible causal structure as the variable number increases.

APPENDIX: SOME PROOFS

The appendix proves identifiability results by converting independence assumptions into differential and functional equations, then classifies their valid solutions. It also establishes the multivariable independence result using entropy and Jacobian arguments.

  • Proof strategy: Theorem 1 uses diagonal Hessians of log joint densities to translate independence between observed variables and disturbances into constraints on the model transformation.Invertibility makes independence between x1 and e2 equivalent to independence between t1 and e2, and similarly for x2, z2, and e1.
  • Two-variable identifiability: If e2 is Gaussian, the appendix proves that h is linear and t1 is also Gaussian.The result identifies a Gaussian-linear configuration among the exceptional cases requiring special treatment.
  • Two-variable identifiability: When f1 is noninvertible, the PNL causal direction is unique because the composite h cannot satisfy the invertibility required by every nonidentifiable situation.If f1 has a flat domain, assuming both directions explain the data would instead make h nondifferentiable at the boundary.
  • Multivariable extension: Under the causal Markov condition and independence of each disturbance ei from its direct causes pai, all disturbances e1, ..., en are mutually independent.The proof uses conditional-density identities, a lower-triangular Jacobian, and the resulting mutual-information expression.
Loading 1205.2599v1…