Source-linked AI summary
Inferring deterministic causal relations
Povilas Daniusis, Dominik Janzing, Joris Mooij, Jakob Zscheischler, Bastian Steudel, Kun Zhang, Bernhard Schoelkopf
TL;DR
Distinguishing cause from effect is difficult for deterministic invertible relations because dependence-based methods cannot identify the direction from P(X, Y). This paper proposes using an independence postulate between the cause distribution and function, theoretically extends the method to low noise, and reports competitive empirical performance with substantially faster computation.
Problem
Dependence-based causal learning cannot distinguish X →Y from Y →X using the joint distribution P(X, Y ), and deterministic relations make this direction problem especially challenging.
Method
The method assumes a special independence relation between the cause distribution and an invertible function, interprets it through information geometry, and uses reference distribution families such as uniform distributions or Gaussians.
Results
The method works on simulated and real data, remains theoretically justified for small additive noise, achieves competitive accuracy, and runs in linear time in the number of data points.
Takeaways & Limitations
The approach extends causal-direction inference to deterministic relations and can handle cases where existing methods require noise, while uniform and Gaussian reference distributions both perform well empirically.
Takeaways & Limitations
The method may completely fail in the very large noise regime, whose behavior requires better understanding.
Abstract
from arXiv · showhide
We consider two variables that are related to each other by an invertible function. While it has previously been shown that the dependence structure of the noise can provide hints to determine which of the two variables is the cause, we presently show that even in the deterministic (noise-free) case, there are asymmetries that can be exploited for causal inference. Our method is based on the idea that if the function and the probability density of the cause are chosen independently, then the distribution of the effect will, in a certain sense, depend on the function. We provide a theoretical analysis of this method, showing that it also works in the low noise regime, and link it to information geometry. We report strong empirical results on various real-world data sets from different domains.
1 Introduction
The paper addresses the difficulty of distinguishing X →Y from Y →X when the variables are deterministically related by an invertible function. It proposes exploiting independence between the cause distribution and the function, whose consequences appear as structured irregularities in the effect distribution.
- Motivation: Conditional-independence methods cannot distinguish Markov-equivalent DAGs, including whether X →Y or Y →X from P(X, Y).This motivates methods based on other statistical asymmetries.
- Existing approaches: Additive noise models assume Y is a function of X plus noise E statistically independent of X, whereas the reverse model generically lacks this property.The statistical independence assumption is decisive because an uncorrelated-noise model can always be found.
- Deterministic setting: The deterministic case is harder, so the paper focuses on invertible relations Y = f(X), assuming without loss of generality that f is monotonically increasing.Noninvertible functions are treated as making the task solvable and are excluded from the paper’s focus.
- Core postulate: The central postulate is that, for X →Y, the distribution of X and the mapping f correspond to independent mechanisms of nature.The paper also frames this idea through separate descriptions of P(X) and f, while noting that practical dependence measures are needed because Kolmogorov complexity is uncomputable.
- Inference intuition: Under this postulate, peaks of the effect density correlate with points where the inverse function has steep slope, making the reverse causal hypothesis require mutual adjustments.For general input densities, the method relies on the shapes of pX and f being sufficiently uncorrelated.
2 Independence of f and pX in terms of information geometry
The paper formulates independence between the cause distribution and the invertible function through relative-entropy geometry. This yields an information-geometric causal inference rule based on the sign of an asymmetry measure.
- Independence condition: The paper interprets independence as the absence of correlation between peaks of pX and regions where f has large slope.This is expressed by treating p(x) and log f′(x) as random variables under the uniform distribution.
- Independence condition: In the backward direction, the corresponding quantities exhibit positive correlation between p(y) and log g′(y), where g = f^-1.This asymmetry distinguishes the causal direction from its inverse.
- Information geometry: The orthogonality condition is formulated as D(pX || vg) = D(pX || u) + D(u || vg), linking the independence assumption to information geometry.Here orthogonality refers to the additivity relation for relative-entropy distances.
- Information geometry: Relative-entropy additivity expresses effect irregularity as the sum of cause irregularity and function irregularity.For a uniform reference, D(pY || v) = D(pX || u) + D(uf || v).
- Exponential families: For noncompact-support distributions, the method replaces a single uniform reference with exponential families of smooth reference distributions.The irregularity measure becomes distance to an exponential family, whose projection is uniquely defined under the stated conditions.
- Inference rule: IGCI infers X → Y when CX→Y < 0 and Y → X when CX→Y > 0, except when the function is too simple.The rule follows from relative-entropy positivity and the nonzero function-irregularity term.
3 Special cases and estimators
The paper develops information-geometric causal criteria for deterministic invertible relations using reference distributions, examines uniform, Gaussian, and isotropic-Gaussian cases, and describes finite-sample estimators. It also identifies estimator limitations and connects the deterministic linear case to prior work.
- Information-geometric formulation: For a diffeomorphism f, the method compares the irregularities of cause and effect relative to reference-measure manifolds, using projections and entropy relationships.The analysis interprets causal asymmetry through information geometry and low-dimensional exponential manifolds.
- Uniform reference: With uniform references on [0, 1], the entropies of the projected cause and effect coincide, yielding a criterion based on the transformation and input density.This specializes the general relationship to arbitrary diffeomorphisms of the unit interval.
- Gaussian references: Gaussian references extend the framework to d-dimensional real vectors, where projections preserve the means and use isotropic covariances determined by the input variances.The choice of reference distributions can affect the inferred direction when the variances differ, although sign reversals were rare experimentally.
- Relation to prior linear methods: For deterministic linear models, the method includes the prior criterion of Janzing et al. as a special case, with a probabilistic justification based on symmetric randomization of the cause distribution.The equivalence holds when the reference manifolds are isotropic Gaussians and the transformation is linear and invertible.
- Reference-measure limitation: The nonlinear-reference method loses covariance information near linear relations, whereas isotropic-Gaussian references retain information except for overall joint scaling.This defines an important scope boundary for selecting reference measures.
- Finite-set extension: For finite probability spaces, the method remains applicable when the reference manifolds are not restricted to the uniform distribution alone.In the discrete-Gaussian example, transformed reference distributions are usually no longer discrete Gaussian, making the inference principle nontrivial.
- Empirical estimators: In finite samples, two natural estimators use either entropy estimates or a slope-based integral; they coincide deterministically, but the slope estimator diverges under noise as m →∞.The noisy slope estimator therefore requires subtraction of its reverse-direction analogue, while entropy estimators were used for the reported experiments.
4 Adding small noise
The paper analyzes robustness to small additive noise by bounding noise-generated entropy. When the deterministic effect has lower entropy than the cause, the causal criterion remains valid below an explicit noise threshold.
- Noise robustness: The analysis bounds the noise variance so that, under an additive noise model, the effect entropy remains lower than the cause entropy.This provides a partial theoretical analysis of the low-noise regime.
- Scope: The theorem establishes robustness only under the stated entropy ordering and noise condition.Its scope is therefore tied to the assumptions used in the additive-noise analysis.
- Entropy bound: The noise-generated entropy bound compares arbitrary unit-variance noise with Gaussian noise, with equality when both X and E are Gaussian.The Gaussian case makes the inequality tight.
- Robustness theorem: Theorem 1 states that if Y = f(X) and S(pY) < S(pX), the method remains robust for noise satisfying inequality (16) with σ < e^(2S(pX)−2S(pY)−1).The threshold depends on the entropy gap between cause and effect.
5 Experiments
Experiments evaluate the proposed method on artificial, CauseEffectPairs, and German Rhine data, covering deterministic, low-noise, and noisier real-world settings. Results show accurate inference under structured mechanisms and low noise, while very large noise remains problematic.
- 5.1 Simulated data: The simulations combine multiple input distributions, noise distributions, and monotonic mechanisms across repeated settings with sample size m = 1000.Each setting was repeated 100 times, and average correct-inference percentages were reported.
- 5.1 Simulated data: Deterministic inference can fail when high-derivative regions of f coincide with peaks in the input density pX.For structured mechanisms such as s5(x), the true direction was inferred quite accurately across the considered input distributions.
- 5.1 Simulated data: At λ = 0.03, performance remains similar across the tested noise distributions, whereas higher noise can strongly affect the inferred direction.The large-noise regime was not investigated in detail and was left for future work.
- Real-world data: The CauseEffectPairs evaluation contains 51 variable pairs from various domains, many with relatively high noise levels.The method was compared with two other causal-inference methods suitable for pairwise causal-direction inference.
- German Rhine data: On Rhine water-level pairs, the method achieved 82% accuracy with the uniform reference measure and 84% with the Gaussian reference measure.These correspond to 189 and 193 correct decisions, respectively, across city pairs; noise was larger for more distant stations.
6 Discussion
The discussion presents the method as an approach for inferring deterministic causal relations under an independence postulate involving the cause distribution and mechanism. It reports theoretical and empirical support, competitive accuracy, faster computation, deterministic applicability, and a limitation under very large noise.
- 6 Discussion: The method assumes a special independence relation between the cause distribution and function relative to selected reference-distribution families.Choosing the reference families determines which distributional information is treated as essential versus scaling and location information.
- 6 Discussion: The method works on simulated and real data, and theory explains why it can still work for functional relations with small noise.This extends the method's supported scope beyond strictly deterministic relations.
- 6 Discussion: The proposed method has competitive accuracy, computation time orders of magnitude faster than existing methods, and runtime linear in the number of data points.It also handles deterministic cases, whereas the cited existing methods require noise.
- 6 Discussion: In the very large noise regime, the method may completely fail, and the authors identify understanding this regime as future work.They also propose future work on estimating confidence in inferred causal directions.