Source-linked AI summary
How Many Samples Are Needed to Determine Causal Direction? Sharp Minimax Bounds for Bivariate LiNGAM
Jikai Jin
TL;DR
Determining causal direction in bivariate LiNGAM is difficult when the edge is weak, disturbances are nearly Gaussian, or error scales are uncertain. The paper proves a sharp local minimax sample-complexity law based on covariance and non-Gaussianity signals, identifying which signal governs recovery.
Problem
Existing LiNGAM theory establishes population identifiability but does not quantify sample-size divergence near indistinguishable directions.
Method
The paper combines exact covariance geometry, a quantitative rotation theorem, and adaptive covariance or independence testing with matching lower bounds.
Results
For sufficiently small β and ν, the required samples scale as log(1/δ)/(d_β^2+β^2ν^2), with d_β and βν representing covariance and non-Gaussianity signals.
Takeaways & Limitations
Recovery is governed by the stronger covariance or non-Gaussianity signal; when covariance classes overlap, identification relies entirely on βν.
Takeaways & Limitations
The sharpness claim is local in β and ν for fixed nuisance bounds, without explicit numerical constants uniform over varying source or rotation bounds.
Abstract
from arXiv · showhide
We study how many observations are needed to determine the causal direction between two linearly related variables. Classical LiNGAM theory shows that independent non-Gaussian disturbances identify the direction, but does not quantify the difficulty when the causal effect is weak or the disturbances are nearly Gaussian. Let $β$ bound the absolute structural coefficient from below, let $ν$ measure each standardized disturbance's distance from Gaussianity, and let the disturbance scales lie in $[\underlineσ,\overlineσ]$. We prove the sharp local minimax law \[ N_2^\star(β,ν,δ) \asymp \frac{\log(1/δ)} {d_β^2+β^2ν^2}, \qquad d_β= \left[β^2- \left(1-\frac{\underlineσ^2}{\overlineσ^2}\right)\right]_+. \] Previous theory established population identifiability or assumed a fixed separation between the two directions. By contrast, we establish the sharp sample complexity as a joint function of edge strength, distance from Gaussianity, and scale uncertainty, and characterize when identification comes from non-Gaussian dependence or from covariance alone. The proof was independently generated with GPT-5.6 Sol in Codex's Ultra mode during a two-hour session. The human author supplied the prompt and was responsible only forchecking the proof and revising and polishing the manuscript.
1 Introduction
The introduction frames causal-direction recovery as distinguishing two observational regressions through independence, with difficulty governed by non-Gaussianity, edge strength, and error-scale uncertainty. It presents a sharp sample-complexity law combining covariance and non-Gaussian dependence signals.
- Identification principle: LiNGAM identifies the causal direction because only the causal regressor is independent of its structural disturbance; reversing the regression mixes the original disturbances.OLS residuals are uncorrelated with regressors in both directions, so independence—not uncorrelatedness—creates the asymmetry.
- Identification principle: At Gaussianity, orthogonal rotations preserve independence, while error-scale restrictions can provide a separate identification route.Under equal error variances, covariance remains informative even as sources approach Gaussianity.
- Signal sources: The wrong-direction dependence signal is at least a constant multiple of βν, while scale uncertainty contributes covariance signal d_β.The rotation mixes independent sources by order β, and d_β vanishes exactly when forward and reverse covariance classes overlap.
- Sample complexity: N_2^⋆(β,ν,δ) is of order log(1/δ)/(d_β^2 + β^2ν^2) for sufficiently small β and ν under fixed tail, scale, and coefficient bounds.The dependence follows from a quantitative rotation theorem and a test combining covariance comparison with non-Gaussian dependence.
- Regimes: When covariance classes overlap, the rate is governed entirely by βν; under equal error variances, d_β = β^2, so covariance remains informative near Gaussianity.The stronger of the covariance signal d_β and non-Gaussian signal βν determines the effective difficulty.
2 Literature review and the precise open gap
Prior LiNGAM and ICA work established population identification, algorithms, or asymptotic results, but did not provide a uniform two-direction finite-sample minimax law over weak edges and near-Gaussian disturbances. This paper addresses that gap with a primitive-parameter complexity theory explaining the interaction of non-Gaussianity and error-scale uncertainty.
- Prior identification and algorithms: LiNGAM’s original and DirectLiNGAM results establish identification or algorithmic correctness, but not uniform finite-sample error bounds.DirectLiNGAM’s main correctness claim is explicitly for infinite sample size, while the original finite-sample discussion is algorithmic and empirical.
- The precise open gap: Stability characterizations and near-Gaussian ICA results do not identify the sharp testing exponent for the composite LiNGAM classes studied here.Existing near-Gaussian results are tied to specified contamination paths, asymptotic limits, or different probability metrics.
- The precise open gap: High-dimensional ICA and LiNGAM results impose stronger moment or kurtosis conditions and do not track deterioration as source-level non-Gaussianity tends to zero.Other recent kernel and Wasserstein approaches provide asymptotic, population, or empirical convergence results without a two-direction minimax lower bound or a uniform near-Gaussian rate.
- Our contribution: The contribution is a sharp local complexity theory parameterized by edge strength, source non-Gaussianity, and uncertainty in error scales.The theory supplies a quantitative modulus of LiNGAM identifiability over a nonparametric sub-Gaussian source class.
- Our contribution: Wrong-direction fitting creates observable dependence of order βν, while the resulting phase law distinguishes non-Gaussianity-based learning from covariance-based identification.Equal error variances explain why complexity need not diverge as ν approaches zero, because covariance can already contain directional information.
3 Problem formulation and main theorem
Section 3 defines the two directional LiNGAM experiments, their admissible source classes, and the minimax directional error. Theorem 3.1 gives the sharp local sample-complexity scaling, with distinct regimes depending on scale uncertainty and non-Gaussianity.
- Problem formulation: The forward and reverse classes impose bounded structural effects, disturbance scales, independent disturbances, and standardized source laws separated from Gaussianity.No density, symmetry, or nonvanishing-cumulant assumption is imposed.
- Problem formulation: The minimax criterion is the smallest sample size at which a decision rule’s worst directional error is at most δ.Decision rules may be measurable and randomized.
- Main theorem: Theorem 3.1 establishes constants and local radii such that the bivariate LiNGAM sample complexity obeys the sharp scaling in β, ν, and δ.The constants depend only on the fixed nuisance bounds.
- Regimes: When scale uncertainty makes d_β = 0 for sufficiently small β, the rate is ≍(β^2ν^2)^−1 log(1/δ); with known common scale, it is ≍ log(1/δ) / [β^2(β^2 + ν^2)].The common-scale case has d_β = β^2.
- Regimes: The rate saturates at β^−4 log(1/δ) when ν ≲ β, and the covariance-driven transition occurs at d_β ≍ βν.These regimes describe when non-Gaussianity or covariance separation controls identification.
- Scope of the theorem: The theorem is local in (β, ν), and its sharpness concerns scaling for fixed nuisance bounds rather than explicit constants uniform over varying source or rotation bounds.Away from the local region, confidence dependence is of order log(1/δ) when d_β^2 + β^2ν^2 is bounded below.
4 Proof outline
The proof combines source-class separation, covariance and non-Gaussianity-based separation, and matching lower-bound constructions to establish the directional-identification rate. Its analytic notation distinguishes the source class from the function spaces used in the argument.
- Proof steps: The source class is nonempty, and the two directional classes are disjoint whenever ν > 0.
- Proof steps: Proposition 6.1 computes exact covariance separation and provides an estimator based on it.
- Proof steps: A rotation by angle θ of admissible non-Gaussian sources has independence defect at least a∗|sin θ cos θ|ν in a third-order Gaussian Sobolev norm.
- Proof steps: The wrong OLS direction induces this rotation, yielding population score gap cβν; estimating the gap at n−1/2 scale gives exponent nβ2ν2.
- Proof steps: The upper bound selects the covariance or independence branch according to which known separation is larger.
- Notation map: The notation separates the source class Q(K, ν) from Sr, Qr, and Hr, which serve distinct roles in the analytic proof.
5 Nonemptiness and qualitative identifiability
The section establishes that the admissible non-Gaussian source class is nonempty and that forward and reverse bivariate LiNGAM models are qualitatively identifiable when β > 0 and ν > 0. Disjointness follows from the Darmois–Skitovich characterization, which forces simultaneously independent linear forms to have Gaussian sources.
- Qualitative identifiability: The bivariate Darmois–Skitovich theorem states that independent nondegenerate sources yielding two independent linear forms with all coefficients nonzero must both be Gaussian.The characterization requires neither densities nor additional regularity.
- Nonemptiness: νsrc > 0 ensures Q(K, ν) is nonempty for every admissible ν up to the constructed source-path range.The source path uses a bounded perturbation of the standard Gaussian density and sets νsrc := A◦ε◦.
- Qualitative identifiability: P→(β, ν) ∩ P←(β, ν) = ∅ whenever β > 0 and ν > 0.Thus no law in the model class admits both causal directions under positive edge strength and non-Gaussianity.
- Qualitative identifiability: In a putative reverse representation, the independent forms are Y = aV + E and X − bY = (1 − ab)V − bE, with all four coefficients nonzero.Applying the theorem forces V and E, hence the standardized disturbances, to be Gaussian, contradicting NGw(Zj) ≥ ν > 0.
6 Exact covariance geometry
Covariance geometry exactly quantifies the directional signal: it is d_β, with no uniform covariance-based direction signal when d_β = 0. When d_β > 0, the covariance classes are separated by a sharp gap, attained at boundary parameter values.
- Covariance signal: Covariance supplies exactly the signal d_β and no uniform directional signal when d_β = 0.The two model classes may be disjoint even when their covariance sets provide no uniform directional separation.
- Covariance separation: When d_β > 0, the two covariance half-spaces are separated by a gap 2Hd_β.This separation underlies covariance-based directional testing.
- Sharpness: At d_β > 0, equality in the covariance separation bound is attained at |a| = |b| = β, s = u = H, and t = v = L.These boundary values establish sharpness of the separation result.
- Covariance testing: Bernstein concentration converts the covariance gap into a finite-sample directional testing guarantee.The proof bounds covariance-estimation deviations and notes that the resulting event contains every sign error.
7 A quantitative Maxwell–Darmois–Skitovich inequality
The section proves a uniform quantitative rotation inequality: rotating independent standardized sub-Gaussian sources creates dependence proportional to mixing strength and their joint non-Gaussianity. A local Hermite estimate, nonlinear remainder control, and compactness argument establish the bound uniformly over the full source class.
- Proof strategy: The full proof combines local Hermite coercivity, nonlinear remainder estimates, Gaussian conjugation, and endpoint angular regularity.These ingredients provide the norm estimates and continuity needed for the uniform bound.
- Uniform quantitative inequality: Theorem 7.1 gives a uniform Sobolev Maxwell inequality for orthogonal rotations of independent centered, variance-one sub-Gaussian sources.The result assumes a finite sub-Gaussian bound K and nonzero rotation coefficient |c| ≥ c∗.
- Uniform quantitative inequality: The dependence defect is bounded below at scale |sc| times the joint weighted non-Gaussianity of the two sources.The mixing contributes |sc|, while departure from the rotation-invariant Gaussian pair contributes the non-Gaussianity factor.
- Proof strategy: The perturbative argument handles sources near Gaussianity, while compactness and Darmois–Skitovich exclude non-Gaussian zeros away from the Gaussian pair.The admissible sub-Gaussian characteristic functions are compact in the relevant Sobolev topology, yielding a positive minimum away from the local regime.
- Connection to sample complexity: Applied to the wrong OLS direction, the inequality yields a population gap of order βν and therefore a β^2ν^2 testing signal.Section 8 supplies |sc| ≳ β and η ≳ ν.
- Proof strategy: A linearized Hermite argument shows that standardization prevents cancellation, producing an S3 signal of size at least |sc|η.This leading term is compared with the nonlinear remainder.
8 Population score gap in the wrong direction
The section converts the abstract rotation result into a population score gap for the wrong OLS direction. The correct direction has zero defect, while the wrong direction is separated with strength controlled by β.
- Population score gap in the wrong direction: The wrong-direction standardized residuals have the abstract rotation form, with mixing strength bounded below by a constant multiple of β.This links Theorem 7.1 to the population score gap used by the test.
- Population score gap in the wrong direction: Proposition 8.1 states that the correct score is zero and the wrong score is separated for 0 < β ≤ β0.The separation constant κ depends only on the fixed constants.
- Population score gap in the wrong direction: The forward pair equals (Z1, Z2), so its defect vanishes.The reverse case follows by swapping X and Y, with the same constant.
- Population score gap in the wrong direction: The rotated source rows are exactly (c, s) and (−s, c), and the coefficient and scale bounds imply |c| ≥ c∗ > 0 and |sc| ≥ c1β.These bounds provide the uniform mixing control needed for wrong-direction separation.
9 A uniform estimator of the Sobolev score
Section 9 constructs a measurable, uniform estimator of both directional Sobolev scores by combining sample-split covariance estimation, Hilbert-space median-of-means estimation, and stability of the OLS transformation. This yields a uniform high-probability score estimator at the parametric deviation scale and a uniform directional upper bound.
- Construction: The estimator splits the sample to estimate covariance independently from Hilbert-valued characteristic-function feature means, enabling conditional robust mean estimation.The covariance estimate is obtained from (X^2, XY, Y^2), projected onto the admissible covariance set, while the remaining observations are OLS-transformed for both directions.
- Robust mean estimation: A measurable Hilbert-space median-of-means selector is built from block means, robust distance order statistics, and deterministic tie-breaking.Its construction uses only the data and the block parameter, not the unknown variance bound.
- Uniform stability: Uniform covariance conditioning and OLS-transform stability control how feature means and scores change when the estimated covariance replaces the population covariance.All model covariance matrices lie in a fixed compact spectral interval, and the stability lemma applies uniformly over both directions and admissible covariance matrices.
- Score representation: The Sobolev score is recovered by mapping joint and marginal characteristic-function derivative coordinates into their defect norm.The map subtracts products of marginal coordinates from joint coordinates, and the resulting norm equals the score defect norm component by component.
- Uniform guarantees: Proposition 9.3 provides measurable estimators of both directional scores with uniform high-probability guarantees at the parametric deviation scale.The result holds for 8 ≤ x ≤ c0n, using covariance estimation and conditional Hilbert-space median-of-means estimation with an e^-cx tail.
- Directional identification: For sufficiently small β and ν, Proposition 9.4 converts the score-estimation result into a uniform directional upper bound.The subsequent argument writes the bound in exponential form A exp(−anξ), with ξ determined by the section’s directional separation quantity.
10 Matching two-point lower bound
The matching lower bound constructs admissible forward and reverse LiNGAM laws whose squared Hellinger separation is O(d_β^2 + β^2ν^2). Tensorization then yields an exponential n-sample testing lower bound, showing both discrepancy mechanisms are necessary.
- Least-favorable pair: The discrepancy splits into a rotation of size βν and a nonorthogonal matrix perturbation of size d_β.These two perturbations correspond to the non-Gaussian and scale-uncertainty contributions in the separation bound.
- Least-favorable pair: The least-favorable forward–reverse pair has squared Hellinger separation at most C(d_β^2 + β^2ν^2).The construction applies for sufficiently small β and ν within the admissible model classes.
- Construction: The constructed laws satisfy the required disturbance-scale bounds, and Hellinger distance is preserved under their common invertible transformation.The proof controls the rotation and nonorthogonal perturbation separately before combining them with the Hellinger triangle inequality.
- Testing consequence: The testing consequence applies to every possibly randomized n-sample directional decision rule throughout a sufficiently small fixed (β, ν) range.The lower bound follows by tensorizing Hellinger affinity for the least-favorable pair.
- Testing consequence: The two testing errors have a nontrivial exponential lower bound in n(d_β^2 + β^2ν^2), with randomization handled by adjoining a shared independent seed.Affinity tensorization and the testing-error–total-variation relation produce the bound.
11 Proof of the main theorem and interpretation
The section completes the main theorem by combining matching exponential risk bounds and translating them into sample complexity. It also explains why uniform nonparametric coverage and the scale-dependent covariance phase are unavoidable.
- Proof of the main theorem: Matching exponential upper and lower risk bounds yield the theorem’s sample-complexity statement after common small parameter ranges are chosen.The proof selects β0 and ν0 small enough for the preceding propositions, then takes infima over sample sizes and decision rules.
- Proof of the main theorem: R_2,n(β, ν) ≥ 1/(4e^-Cn(d_β^2+β^2ν^2)) supplies the lower-risk half of the theorem.The bound follows by taking the infimum over all decision rules in Proposition 10.2; a minimizing rule need not exist.
- Interpretation: The upper bound requires no source density, while the smooth density is used only to construct a least-favorable lower pair.Thus the lower bound establishes non-improvability along one admissible source direction, and the upper bound supplies uniformity over the full nonparametric class.
- Interpretation: When d_β = 0, forward and reverse covariance classes intersect, so covariance provides no information for distinguishing the constructed pair.If the scale interval is a singleton, covariance remains directional even at the Gaussian limit.