Source-linked AI summary
Causality for Machine Learning
Bernhard Schölkopf
TL;DR
The article examines how causality connects to machine learning and AI, where these links have been limited or are still open. It synthesizes causal-modeling concepts and research directions, concluding that causal thinking is relevant to robust learning and unresolved AI problems.
Problem
Machine learning has only recently developed stronger connections to causality, while causal discovery, dynamics, and the role of objects remain difficult research problems.
Method
The article develops a conceptual account connecting causal models, interventions, function classes, independent mechanisms, invariance, robustness, and machine learning.
Results
The synthesis links causal modeling to more invariant or robust models and uses causality to frame representation learning, disentanglement, causal discovery, and other machine-learning problems.
Takeaways & Limitations
Causal thinking provides a framework for understanding machine-learning challenges and for developing data-driven causal methods and new learning methods.
Takeaways & Limitations
The account is personal and biased, may omit relevant work, and acknowledges that major machine-learning and AI problems remain unsolved.
Abstract
from arXiv · showhide
Graphical causal inference as pioneered by Judea Pearl arose from research on artificial intelligence (AI), and for a long time had little connection to the field of machine learning. This article discusses where links have been and should be established, introducing key concepts along the way. It argues that the hard open problems of machine learning and AI are intrinsically related to causality, and explains how the field is beginning to understand them.
1 Introduction
Machine learning has improved substantially but remains weak at transfer across problems, intervention-based reasoning, and acting in imagined spaces. The article argues that causal modeling can help address these conceptual limitations and support more invariant or robust models.
- Causal modeling can lead to more invariant or robust machine-learning models.
- Machine learning performs poorly at transfer across problems and other forms of generalization beyond IID samples.These capabilities differ from ordinary generalization between data points sampled from the same distribution.
- Machine learning often treats interventions, domain shifts, and temporal structure as nuisances rather than information.
- Causal modeling focuses on modeling and reasoning about interventions to help explain and resolve these limitations.
2 The Mechanization of Information Processing
The information revolution evolved from symbolic computing toward learning systems that extract information and infer rules from unstructured data. Its growing societal impact creates both opportunities and ethical responsibilities.
- Computers enabled industrial-scale information processing, while AI aims to perform that processing intelligently.
- Machine learning applications often convert user data into predictions about future behavior and money.
- The current information revolution has shifted from symbolic AI and programmed rules toward learning systems that infer rules from data.The learning phase extends information extraction to unstructured data.
- The information revolution is transforming IT companies into AI-first organizations and creating industries around data collection and clickwork.
- The analogy between energy and information is compelling, but the paper characterizes current understanding of information as incomplete.
- Digital goods can be copied at essentially zero cost, unlike many physical goods.
- Because information processing underlies major technological and societal changes, the current revolution may be especially significant.
- The ethical use of information technologies raises concerns involving privacy, clickwork, governance, and citizen incentives.
3 From Statistical to Causal Models
Statistical learning models associations in a joint distribution, whereas causal models represent mechanisms, interventions, and graph-structured factorization. Observational data alone generally cannot identify causal direction without additional assumptions.
- Machine-learning successes commonly rely on massive data, high-capacity models, high-performance computing, and IID problem settings.
- Observational dependence is symmetric, but interventions expose directional mechanisms such as a customer buying a rucksack after owning a laptop.
- The Common Cause Principle explains statistical dependence through a variable that causally influences both observables and renders them independent when conditioned on.
- Without additional assumptions, observational data cannot distinguish alternative causal structures because they generate the same class of distributions.
- Structural causal models assign each observable as a deterministic function of its graph parents and a stochastic unexplained variable.
- Recursive application of structural assignments yields the observational joint distribution and graph-implied conditional independences.
- Independent noise variables and graph structure imply a causal factorization into conditionals corresponding to structural assignments.
- Causal learning exploits causal conditionals and structural assumptions beyond the function-class assumptions of statistical learning.
4 Levels of Causal Modelling
Mechanistic models describe physical dynamics in detail, statistical models learn associations but generally do not predict interventions, and causal models occupy an intermediate level. Causal discovery seeks such models from data under weak assumptions.
- Mechanistic models: Differential equations can predict system evolution, intervention effects, statistical dependences, and causal structure.Their causal structure can be read from which state entries affect future entries.
- Statistical models: Statistical models usually omit time and predict variables only while experimental conditions remain unchanged.
- Statistical models: Statistical models can often be learned from data but do not predict the effects of interventions.
- Causal models: Causal discovery and learning aims to construct causal models from data using only weak assumptions.
- Causal models: Causal models lie between mechanistic and statistical models, abstracting from physical realism while retaining selected interventional and counterfactual capabilities.
5 Independent Causal Mechanisms
The Independent Causal Mechanisms principle treats a causal system as autonomous modules whose mechanisms remain independent, helping explain invariance across settings and supporting algorithmic accounts of causal structure.
- Independence principle: The Beuchet Chair illustrates how violating independence between an object and perceptual process can produce a misleading perceived structure.The illusion arises from a special vantage point that makes two separate objects appear as one chair.
- Independence principle: The ICM principle states that a system’s causal generative process consists of autonomous modules that do not inform or influence one another.In probabilistic models, each conditional mechanism is independent of the other mechanisms.
- Mechanism invariance: Under causal factorization, changing one mechanism does not change the others, while smaller distributional changes tend to affect only a sparse or local set of mechanisms.This supports invariance when moving between related settings or domains.
- Mechanism invariance: Independent mechanisms differ from statistical independence of variables, because variables in a causal graph may remain statistically dependent even when their mechanisms are independent.Mechanism dependence concerns conditional factors, not simply dependence between observed variables.
- Algorithmic independence: Algorithmic graphical models represent independent noise strings as programs run by fixed Turing machines, extending causal modeling beyond statistical distributions.The approach treats algorithmic independence as a relation between descriptions of mechanisms.
- Algorithmic independence: Applying algorithmic independence to an initial state and system dynamics yields non-decreasing Kolmogorov complexity, interpreted as entropy that stays constant or increases.This provides the thermodynamic arrow of time and implies the second law under the stated interpretation.
6 Cause-Effect Discovery
Cause-effect discovery from observational data is difficult because two-variable distributions can support multiple causal directions and finite-sample independence tests are hard. The paper explains how assumptions on function classes can break this symmetry and improve causal discovery methods.
- Limits of observational discovery: Observational data alone cannot distinguish the three two-variable causal cases because they generate the same class of observational distributions.A causal model therefore contains more information than a statistical model.
- Limits of observational discovery: Finite datasets make conditional independence testing difficult, especially with continuous, multidimensional conditioning sets; with two variables, the Markov condition has no nontrivial implications.These are separate obstacles for causal discovery based on conditional independences.
- Function-class assumptions: Function-class assumptions can make causal structure more learnable from finite data by restricting how unobserved variables select among possible functions.Smooth dependence on a selector variable reduces the effective function-class size.
- Function-class assumptions: For additive noise models, the observed distribution cannot generally be fit by an additive noise model in the reverse direction, breaking cause-effect symmetry.This result requires genericity assumptions and has exceptions, including linear functions with Gaussian noises.
- Function-class assumptions: Function-class assumptions have also supported progress in conditional independence testing through kernel function classes representing distributions in reproducing kernel Hilbert spaces.The cited methods extend independence testing with learned representations of probability distributions.
- Broader connection: Machine learning ideas can help address causality problems previously considered hard, while causality may also help improve machine learning.The paper presents this as a two-way connection between the fields.
7 Half-Sibling Regression and Exoplanet Detection
Causal models helped identify shared noise structure in stellar light curves, enabling exoplanet-transit detection despite severe instrumental errors. A binary example also shows that causal direction may remain unidentifiable from observational data.
- Exoplanet detection: The astronomy application builds on causal models inspired by additive noise models and the independent causal mechanisms assumption.The passage describes the resulting astronomy breakthrough as enabled by this modeling approach.
- Limits of causal identification: A causal edge X → Y can be observationally invisible when binary mechanisms and noise make Y uniform and independent of X.In that case, observational data cannot distinguish X → Y from alternatives.
- Exoplanet detection: Shared noise across stars motivated a method for removing systematic errors from Kepler light curves before searching for exoplanet transits.The transit signal was often orders of magnitude smaller than instrument errors.
- Exoplanet detection: Kepler’s reaction-wheel failure increased systematic error, creating a difficult setting that particularly suited the error-removal method.The augmented system combined error models with exoplanet-transit models and efficient light-curve search.
8 Invariance, Robustness, and Semi-Supervised Learning
This section connects causal direction and independent mechanisms to invariant transfer, semi-supervised learning, adversarial robustness, and broader challenges in machine learning. It also identifies dynamics and representation of complex environments as open problems.
- Invariance and transfer: Causal direction is presented as crucial for transfer and robustness under covariate shift, with effect-from-cause prediction expected to transfer more easily.The accompanying work also made a prediction for semi-supervised learning.
- Semi-supervised learning: For anticausal learning, unlabelled inputs can contain information about p(y|x), whereas using p(X) is predicted to be futile in the causal direction.This follows from independence of cause and mechanism in the causal direction and positive dependence in the reverse direction.
- Semi-supervised learning: A meta-analysis of published SSL benchmark studies corroborated the paper’s prediction, while later work added theoretical analyses and conditional SSL.The prediction concerned the usefulness of unlabelled data in the anticausal direction.
- Adversarial robustness: Adversarial examples violate IID assumptions because modified test inputs are interventions designed to expose non-robustness of anticausal p(y|x).The passage distinguishes this setting from ordinary IID prediction.
- Adversarial robustness: Causal classifiers are hypothesized to be harder to fool adversarially, and modeling the causal generative direction has been reported as a possible defense.The proposed defense is called analysis by synthesis in vision.
- Open problems: Animals’ ability to group pixels by common fate or intervention is offered as a causal perspective on why high-dimensional Atari reinforcement learning is difficult.The question contrasts original high-resolution games with downsampled versions that made DeepQ work.
- Open problems: Connecting causal learning to dynamics remains a large open area because existing causal models generally need not represent time.The altitude–temperature example illustrates an underlying temporal physical process beneath a static causal model.
9 Causal Representation Learning
Causal representation learning seeks variables and mechanisms that support causal modeling in high-dimensional, unstructured data. The proposed direction combines structural causal models with generative architectures to obtain transferable, manipulable representations for intervention and reasoning.
- Causal representation learning: Causal representation learning aims to discover causal variables from observations such as images rather than receiving suitable units in advance.This extends causal discovery beyond settings where variables are already defined.
- SCMs and generative models: Embedding an SCM inside a generative model can represent unexplained variables as latent noise while processing high-dimensional, unstructured inputs and outputs.Modern generative models similarly make desired randomness an exogenous model input.
- Causal representation learning: The representation-learning objective aligns causal-variable discovery with learning representations that are robust, transferable, interpretable, explainable, or fair.Identifying suitable causal units remains challenging for both human and machine intelligence.
- Transferable mechanisms: Independent causal mechanisms provide an inductive bias for transferring modules across substantially different domains.Competitive training is cited as one way to learn such models, including distorted handwritten characters.
- SCMs and generative models: An encoder–structural-assignment–decoder architecture maps observations into latent noise variables, applies causal mechanisms, and reconstructs the original data.With sufficiently large latent dimension, reconstruction training seeks p ◦ f ◦ q ≈ id on observed images.
- Interventional representations: The ICM assumption implies that interventions on latent representations corresponding to true causal variables can produce valid image data.The relevant interventions act on latent noises and the mechanisms driven by them.
- Interventional world models: Causality is proposed as a route beyond statistical representations toward models supporting intervention, planning, and reasoning.Such models are intended to represent alternative scenarios and actions in an imagined space.
10 Personal Notes and Conclusion
The author recounts how personal collaborations and workshops connected him to causality research, and describes contributions linking causality with machine learning. He concludes that these links remain an emerging area while machine learning’s hardest problems remain unsolved.
- Personal trajectory: The author’s engagement with causality developed through encounters with Judea Pearl, Dominik Janzing, and the causality community.A 2009 workshop helped establish membership in the causality community and led to meeting Peter Spirtes.
- Methodological context: The causal graph can become part of an unspecified model, while network topology can ensure each noise variable feeds only one subsequent unit and allow all DAGs to be learned.The passage also notes that interventions on the S_i can be performed, including decoders without encoders such as GANs.
- Research connections: The work links causality and machine learning in both directions: learning methods develop data-driven causal methods, while causal ideas inspire new learning methods.
- Research connections: Representation learning and disentanglement are presented as fields where causal questions and deep learning have increasingly converged.The author notes that research has begun combining both fields.
- Conclusion: The author cautions that this account is personal and biased, that the field remains in its infancy, and that hard machine learning and AI problems remain unsolved.