Source-linked AI summary
The lesson of causal discovery algorithms for quantum correlations: Causal explanations of Bell-inequality violations require fine-tuning
Christopher J. Wood, Robert W. Spekkens
TL;DR
The paper examines whether causal discovery algorithms can distinguish Bell-inequality-violating from Bell-inequality-satisfying correlations. It shows that conditional-independence-based algorithms cannot, while causal explanations of nonsignalling Bell violations require fine-tuning that conflicts with their core principles.
Problem
Conditional-independence-based causal discovery algorithms may lack the information needed to distinguish correlations that violate Bell inequalities from those that satisfy them.
Method
The paper applies causal discovery tools to Bell correlations and analyzes causal models using observed independences, including marginal setting independence and no-signalling.
Results
Such algorithms cannot distinguish Bell-inequality-violating from Bell-inequality-satisfying correlations, and causal models reproducing Bell violations while respecting observed independences require fine-tuned causal parameters.
Takeaways & Limitations
Causal explanations of Bell-inequality violations must confront the failure of conditional-independence-only methods and the incompatibility between observed independences and the no-fine-tuning principle.
Takeaways & Limitations
Extending causal inference to quantum causal models requires justifying changes to both the physical framework and the rules of inference.
Abstract
from arXiv · showhide
An active area of research in the fields of machine learning and statistics is the development of causal discovery algorithms, the purpose of which is to infer the causal relations that hold among a set of variables from the correlations that these exhibit. We apply some of these algorithms to the correlations that arise for entangled quantum systems. We show that they cannot distinguish correlations that satisfy Bell inequalities from correlations that violate Bell inequalities, and consequently that they cannot do justice to the challenges of explaining certain quantum correlations causally. Nonetheless, by adapting the conceptual tools of causal inference, we can show that any attempt to provide a causal explanation of nonsignalling correlations that violate a Bell inequality must contradict a core principle of these algorithms, namely, that an observed statistical independence between variables should not be explained by fine-tuning of the causal parameters. In particular, we demonstrate the need for such fine-tuning for most of the causal mechanisms that have been proposed to underlie Bell correlations, including superluminal causal influences, superdeterminism (that is, a denial of freedom of choice of settings), and retrocausal influences which do not introduce causal cycles.
I. INTRODUCTION
The paper examines whether causal discovery algorithms can capture the causal implications of Bell correlations. It argues that conditional-independence-based algorithms cannot distinguish Bell-inequality violations and that causal explanations of such correlations require fine-tuning.
- Motivation: Causal discovery algorithms infer causal relations from observed correlations, with causal relations additionally supporting intervention and counterfactual reasoning.The paper situates its analysis within work associated with Pearl and Spirtes, Glymour, and Scheines.
- Bell correlations: Bell scenarios exhibit conditional independences shared by correlations that satisfy and violate Bell inequalities.Apart from special quantum-state degeneracies, these are the only conditional independences arising in Bell scenarios.
- Bell correlations: Algorithms using only conditional-independence relations therefore cannot distinguish Bell-inequality-violating correlations from Bell-inequality-satisfying correlations.Their input omits distributional features that reveal the Bell distinction.
- Approach: The paper applies standard algorithms in settings with and without hidden variables, warning that their outputs require careful interpretation.The analysis focuses on prominent algorithms whose input is the set of conditional independences among observed variables.
- Central result: Any causal model reproducing Bell-inequality violations while respecting observed independences must fine-tune its causal parameters.The relevant independences include marginal independence of measurement settings and no-signalling.
- Central result: The paper reframes Bell’s theorem as a constraint on causal explanations and places superluminal causes, superdeterminism, and non-cyclic retrocausality under a common fine-tuning criticism.This approach derives Bell’s local causality from no fine-tuning while also targeting the other two mechanisms.
U V Model parameters
This section introduces causal models as structures plus parameters that determine conditional probabilities, then explains how they generate joint distributions and conditional-independence predictions. It also presents the causal Markov condition as the main structural constraint.
- Model parameters: A causal model combines a causal structure, represented by a DAG, with causal-statistical parameters specifying each variable’s conditional probability given its parents.The model may be obtained from a deterministic model by eliminating some exogenous variables.
- Model parameters: Deterministic causal models specify functions fixing each variable from its parents, while statistical parameters specify distributions over exogenous variables.Deterministic models are a special case in which conditional probabilities correspond to deterministic functions.
- Joint distributions: A causal model predicts a joint distribution by multiplying the conditional probabilities of every variable given its parents.For a non-complete DAG, the distributions supported by the structure form only a subset of all possible distributions.
- Conditional independence: Conditional-independence relations are distributional features determined by causal structure rather than by causal-statistical parameters, and most causal discovery algorithms focus on them.The section defines conditional independence through equivalent conditional-probability and factorization conditions.
- Causal Markov condition: The causal Markov condition states that each variable is conditionally independent of its nondescendants given its parents.Additional conditional independences can be inferred using the semi-graphoid axioms or identified graphically by d-separation.
- Causal Markov condition: The example DAG yields conditional independences such as Y ⊥⊥ S | T through repeated applications of the Markov condition and semi-graphoid axioms.The derivation proceeds through decomposition, contraction, weak union, and symmetry.
III. CAUSAL DISCOVERY ALGORITHMS
Causal discovery algorithms infer causal structures from conditional independences, favoring explanations whose independences remain stable under parameter changes. In a simple three-variable example, natural and fine-tuned causal explanations generate the same distribution, but only the former robustly preserves the observed dependence pattern.
- Algorithms and causal explanations: Causal discovery algorithms solve an inverse problem by using observed correlations, especially conditional independences, to narrow the causal structures that could explain them.For direct applications to quantum-theoretic distributions, the finite-sample problem of inferring independences does not arise.
- Algorithms and causal explanations: Three distinct causal structures can support exactly the same conditional independence A ⊥⊥ B | C, so observations may leave the causal structure underdetermined.Such structures form an equivalence class, and additional information such as temporal ordering may be needed to identify one uniquely.
- Natural and unnatural models: For a three-variable distribution in which A and B are independent while each is dependent on C, the natural explanation makes A and B causally independent.The corresponding model is shown as the natural causal model for the relevant conditional independences.
- Natural and unnatural models: An alternative model can reproduce the same distribution by combining two causal mechanisms whose effects cancel, but this requires parameter values that cannot be chosen arbitrarily.This makes the explanation less natural because the observed independence depends on a special parameter choice.
- Natural and unnatural models: Changing parameters in the first model preserves the independence pattern, whereas changing parameters in the second model does not.Causal discovery algorithms therefore favor the first model because its statistical independences are robust to parameter variation.
- Faithfulness and minimality: Faithfulness, or no fine-tuning, requires every conditional independence to follow from causal structure rather than particular causal-statistical parameter values.Under a uniform parameter prior, parameter choices explaining structure-unimplied independences have measure zero.
- Faithfulness and minimality: Minimality prefers a causal model that another model can simulate but cannot itself simulate in return, treating lower expressive power as greater falsifiability.This simplicity criterion concerns the distributions a structure permits, not merely its number of variables or arrows.
- Faithfulness and minimality: Causal discovery is a fallible heuristic rather than a guaranteed identification procedure, and it can return an equivalence class instead of a unique causal structure.The faithfulness assumption is treated as an inference to the best explanation, illustrated by preferring one chair over two precisely aligned chairs.
A. Example of causal discovery assuming no latent variables
Under the no-latent-variable assumption, causal discovery tests candidate causal orderings and removes arrows or structures that conflict with observed conditional independences. The example shows that stability and temporal information can narrow the viable causal explanation.
- Candidate causal structures: The analysis assumes observed variables are the only causally relevant variables and considers every possible causal ordering.A causal ordering permits influences only from lower- to higher-ranked variables.
- Candidate causal structures: For S < T < C, the general factorization is P(S, T, C) = P(S)P(T | S)P(C | T, S).The conditional independence (S ⊥⊥C | T) permits replacing P(C | T, S) with P(C | T) and dropping S → C.
- Conditional independence: The resulting simplified structure generates distributions satisfying (S ⊥⊥C | T), making it a candidate rather than an arbitrary joint-distribution model.The same procedure is repeated for each causal ordering.
- Candidate causal structures: The six possible orderings yield five candidate structures, but two violate stability, leaving three viable structures.The orderings T < S < C and T < C < S produce the same structure.
- Additional ordering information: Adding the information that tar appears after smoking rules out structures with T < S and leaves the smoking → tar → cancer structure.Latent genetic explanations motivate extending causal discovery beyond the no-latent-variable setting.
- Algorithmic guarantee: The algorithm is correct when a set of causal structures exists that is both minimal and faithful to the observed correlations.More efficient related algorithms share this minimality-and-faithfulness correctness condition.
B. Example of causal discovery allowing for latent variables
Allowing latent variables requires richer causal representations and additional restrictions on candidate structures. The discussion shows that standard algorithms can identify minimal faithful structures from conditional independences, but those relations need not reproduce the full observed distribution.
- Latent-variable models: Latent-variable causal discovery is more complicated because unobserved but causally relevant variables must be incorporated into the analysis.The paper distinguishes models limited to pairwise common causes from unrestricted models with latent causes of more than two observed variables.
- Latent-variable models: Latent variables mediating an observed pair or acting as common effects do not enlarge the set of observable distributions beyond direct-cause models.This motivates focusing on latent variables that serve as common causes.
- Limitation: A model faithful to observed conditional independences need not reproduce the observed distribution, so CI-based minimality can select structures that fail at the distribution level.This distinction is identified as a significant shortcoming of prominent algorithms.
- Minimality and faithfulness: For a fixed set of conditional independences, any faithful model can be replaced by a faithful model limited to pairwise common causes, which minimality prefers.Consequently, standard algorithms need only search this restricted class when using conditional independences alone.
- Graphical representation: Patterns summarize sets of latent-and-observed-variable causal structures using graphs whose nodes represent only observed variables and whose edges have several possible forms.Bidirected, directed-with-circle, and undirected edges encode alternative underlying DAG connections.
- IC* algorithm: For the smoking example, IC* examines twenty-five combinations and eliminates those introducing new v-structures, leaving nine candidate causal structures.The algorithm is correct for minimal structures with pairwise common causes that are faithful to the observed conditional independences.
- Application to smoking: With temporal information that tar follows smoking, three options remain: smoking causes tar and cancer, a latent common cause explains smoking and tar, or both mechanisms operate.The latent-common-cause option can screen off smoking from cancer while leaving smoking noncausal for cancer.
IV. APPLYING CAUSAL DISCOVERY ALGORITHMS TO QUANTUM CORRELATIONS
The paper applies CI-based causal discovery to EPR and CHSH quantum correlations, comparing one Bell-compatible experiment with one Bell-inequality-violating experiment. Because both share the same conditional independences, these algorithms cannot distinguish their causal implications.
- Quantum scenarios: The analysis studies Bell experiments with two systems, two settings per measurement, and two outcomes per measurement, using S and T for settings and A and B for outcomes.Bell’s local-causality structure treats A and B as effects of their local settings and a shared hidden variable λ.
- EPR and CHSH correlations: The EPR correlations considered satisfy Bell inequalities and can therefore be explained by local causes.Measurements along the same axes yield perfect outcome correlation, while different axes yield no correlation.
- EPR and CHSH correlations: The CHSH setup produces correlations with probability approximately 0.85 for the specified correlation and anticorrelation cases, and violates a Bell inequality.Both experiments independently sample the settings.
- Conditional independences: EPR and CHSH have the same generating conditional independences: independent settings and the two no-signalling relations.The listed relations are (S ⊥⊥T), (A ⊥⊥T | S), and (B ⊥⊥S | T), after excluding pathological independences from maximal-entanglement degeneracy.
- Algorithmic consequence: Since the algorithms use only conditional independences, they draw the same causal conclusions for EPR and CHSH despite their different Bell-inequality status.EPR is locally explainable, whereas CHSH is not locally explainable.
- Algorithmic consequence: CI-based causal discovery does not capture Bell’s theorem because independences alone are insufficient; correlation strength must also be examined.The paper frames this as an inverse problem: infer possible causal structures from quantum correlations.
A. No latent variables
Under assumptions of no latent variables, setting-to-local-outcome causation, and acyclic structures, the candidate causal graphs for nonsignalling quantum correlations cannot represent all required independences faithfully. Explaining them therefore requires fine-tuning, and IC/SGS signals this failure by returning no valid pattern.
- Candidate structures: For the ordering S < T < A < B, the most general joint distribution factorizes as P(S)P(T|S)P(A|S, T)P(B|S, T, A).Other orderings consistent with S < A and T < B produce the causal structures shown in Fig. 23.
- Faithfulness failure: The Fig. 23a structure captures (S ⊥⊥T) and (A ⊥⊥T|S) but not (B ⊥⊥S|T) faithfully.The latter independence can hold only through fine-tuning of causal parameters, including dependencies between parameters for P(B|S,T,A) and P(A|S).
- Faithfulness failure: The analogous Fig. 23b structure has the same problem, so no no-latent-variable causal structure in this class faithfully captures the relevant no-signalling independences.This conclusion follows under the assumed causal ordering constraints and local setting-to-outcome influences.
- Algorithmic failure: Applying the IC algorithm, equivalently SGS, returns a graph that is not a valid pattern because no faithful causal structure exists in this setting.The result matches the algorithms’ correctness guarantee, which applies only when a faithful structure exists.
- Superluminal influences: Even allowing superluminal causal influences without hidden variables does not avoid fine-tuning if those influences must remain unable to transmit superluminal signals.The no-signalling requirement itself imposes the tuning constraint.
B. Latent variables allowed
Allowing latent variables, IC* returns causal patterns from conditional independences alone, but these patterns may fail to reproduce the observed quantum distribution. Consequently, such algorithms cannot determine whether correlations violate Bell inequalities or admit a locally causal explanation.
- IC* applied to nontrivial no-signalling correlations produces the pattern shown in Fig. 24.
- The algorithm’s output leaves ambiguity about whether settings are direct causes of outcomes or linked through freely chosen common causes.
- Among compatible pairwise-common-cause structures, the models align with Bell’s notion of local causality, and minimality favors the corresponding model.
- A causal structure can reproduce all relevant conditional independences while failing to reproduce the full distribution generated by a CHSH experiment.
- Conditional-independence-based minimality can therefore favor a model that cannot reproduce the observed distribution.
- Causal discovery algorithms must incorporate correlation strengths, not only conditional independences, to assess local causal explanations.
C. Some proposed causal explanations of quantum correlations
The paper applies causal-discovery ideas to proposed causal accounts of Bell-inequality-violating quantum correlations. It considers hidden-variable models generally and focuses on three mechanisms: superluminal causation, superdeterminism, and retrocausation.
- The analysis examines three proposed explanations of Bell-inequality-violating correlations: superluminal causation, superdeterminism, and retrocausation.
- The discussion begins with causal explanations that allow hidden variables, while treating structures without hidden variables as a special case.
1. Superluminal causation
Superluminal causal models can reproduce Bell correlations, but preserving the observed no-signalling independences requires fine-tuning. This issue applies across several proposed mechanisms, including direct cross-wing influences and deBroglie-Bohm-inspired models.
- Superluminal explanations posit causal influences from one wing’s setting or outcome to the other wing’s outcome.
- Fine-tuning is required to cancel or otherwise suppress the superluminal signals implied by these causal paths.In Fig. 25c, correlations along different paths can cancel; other structures use specially distributed hidden variables.
- The Toner-Bacon model illustrates a mechanism where signalling is prohibited only for a special distribution over shared random variables.
- The deBroglie-Bohm interpretation is a prominent superluminal-causation model, with variants differing over whether no-signalling is dynamically explained.
- Valentini’s equilibration-based account permits small deviations from equilibrium that could in principle enable superluminal signalling.
- Finite-speed superluminal models can deviate from quantum operational predictions when wings are space-like separated relative to the influence speed.
2. Superdeterminism
Superdeterministic models explain Bell correlations by allowing variables to influence measurement settings, thereby abandoning free choice. They can keep influences subluminal, but reproducing the observed independences still requires fine-tuning.
- 2. Superdeterminism: Superdeterminism allows hidden variables or other variables to causally influence one or both measurement settings.
- 2. Superdeterminism: The proposed causal influences can all be subluminal, but the models conflict with freely chosen settings.
- 2. Superdeterminism: Superdeterministic structures require fine-tuning to satisfy observed setting independence and the two no-signalling conditional independences.
- 2. Superdeterminism: When free will is abandoned, no-signalling becomes an observed statistical independence that the causal model must still reproduce.
- 2. Superdeterminism: Similar fine-tuning mechanisms enforce right-to-left no-signalling in the alternative structures shown in Figs. 26b and 26c.
3. Retrocausation
Retrocausal models can preserve relativistic spacetime structure by reversing causal directions, but acyclic versions still require fine-tuning to reproduce Bell correlations without signalling.
- Retrocausation: Retrocausation posits causal influences contrary to the standard arrow of time, potentially within or backward along the light cone.The section considers only retrocausal structures without causal cycles.
- Constructing acyclic models: Acyclic retrocausal models can be generated by reversing causal arrows entering measurement settings in superdeterministic models.One example reverses the µ → S arrow, producing the structure shown in Fig. 27c.
- Constructing acyclic models: The spacetime location assigned to µ can make retrocausal and superluminal explanations effectively indistinguishable as causal accounts.This follows especially if spatiotemporal relations are treated as supervening on causal relations.
- Fine-tuning: Retrocausal explanations require fine-tuning because generic parameters would correlate S and B, enabling signalling along the causal chain.Thus they face the same constraint as superluminal and superdeterministic explanations.
- Comparison with superluminal models: Superluminal structures without hidden variables are mostly excluded because some imply conditional independences absent from the observed correlations.Only diagram 28c remains a candidate, and it still requires fine-tuning to preserve no-signalling independence.
V. PROOF OF THE NECESSITY OF FINE-TUNING IN CAUSAL EXPLANATIONS OF BELL INEQUALITY VIOLATIONS
The proof examines all causal models compatible with the observed conditional independences and shows that any model avoiding fine-tuning yields a Bell-local decomposition, contradicting Bell-inequality violation.
- Proof strategy: The proof uses a brute-force search over causal structures for the four observed variables in Bell-type experiments.It asks whether any causal explanation can respect the core principles of CI-based causal discovery algorithms.
- Assumptions: The assumptions are quantum predictions, standard causal-model explanations, and faithfulness, which denies explaining observed independences through parameter fine-tuning.Together these assumptions are shown to be inconsistent for the target correlations.
- Eliminating causal structures: Observed independences exclude hidden common causes and direct influences for the pairs {S, T}, {S, B}, and {T, A}.The remaining potentially connected pairs are {A, B}, {S, A}, and {T, B}.
- Remaining structures: Without fine-tuning, A and B must share a common cause, while S-A and T-B may involve direct influences, common causes, or both.These are the only causal structures that can explain the observed conditional independences without fine-tuning.
- Contradiction: All remaining structures decompose P(A, B|S, T) into the form of Eq. (10), which implies Bell inequalities and therefore cannot explain their violation.The proof consequently exhausts all candidate causal structures.
VI. CONCLUSIONS
The paper concludes that conditional-independence-only causal discovery cannot distinguish Bell-violating from Bell-satisfying correlations, while causal explanations of the former require fine-tuning.
- Main conclusions: CI-only causal discovery algorithms cannot distinguish Bell-inequality-violating correlations from Bell-inequality-satisfying correlations.Algorithms that use correlation strength are needed to capture Bell’s theorem more fully.
- Main conclusions: Every standard causal model reproducing no-signalling Bell violations must explain observed independences through fine-tuned causal parameters.This applies to superluminal, superdeterministic, and acyclic retrocausal strategies.
- Research directions: The authors identify opportunities for quantum-foundations tools to improve causal discovery algorithms and for causal-model ideas to inform quantum-foundations research.They present these as directions for further research rather than established results.
- Quantum causal models: Quantum causal models replace classical probability with a noncommutative generalization while retaining analogous causal-structure assumptions.If given a sensible causal interpretation, they could explain Bell violations without fine-tuning.
Appendix A: d-separation
The appendix introduces d-separation as the DAG relation used to represent conditional independences, based on how chains, forks, and colliders block paths.
- Basic structures: DAGs represent conditional independence through d-separation, using chains, forks, and colliders as basic three-variable structures.These structures are illustrated in Fig. 29.
- Path blocking: A path between X and Y is blocked by Z when a conditioned chain or fork contains C, or when an unconditioned collider C has no conditioned descendant.These are the two path-blocking conditions used in the definition.
- Definition: X and Y are d-separated by Z exactly when Z blocks every path between them in the DAG.In a causal-network interpretation, d-separation represents a causal screening-off relation corresponding to statistical conditional independence.