Source-linked AI summary

Granger Causality: A Review and Recent Advances

Ali Shojaie, Emily B. Fox

arXiv:2105.02675v2stat.MEcs.LGstat.ML

TL;DR

The paper reviews advances in Granger-causality methods for increasingly complex time-series settings. These developments address nonlinear, non-Gaussian, categorical, subsampled, and mixed-frequency observations while retaining important identification constraints.

  • Problem

    Modern recording modalities and large-scale data processing create challenges for analyzing time series, while causal interpretation requires the stringent assumption of no unmeasured confounders.

  • Method

    The paper synthesizes methods spanning high-dimensional VARs, nonlinear interaction models, categorical formulations, and subsampled or mixed-frequency time series.

  • Results

    Recent developments expand Granger causality to flexible methods for non-Gaussian and non-continuous observations and to differences between the true causal time scale and observed-data frequency.

  • Takeaways & Limitations

    These developments offer new opportunities for investigating causal effects among variables across broader observation types and time scales.

  • Takeaways & Limitations

    Granger-causality analyses require no unmeasured confounders, a stringent requirement, and multinomial logistic models are non-identifiable without further formulation.

Abstract

from arXiv · show

Introduced more than a half century ago, Granger causality has become a popular tool for analyzing time series data in many application domains, from economics and finance to genomics and neuroscience. Despite this popularity, the validity of this notion for inferring causal relationships among time series has remained the topic of continuous debate. Moreover, while the original definition was general, limitations in computational tools have primarily limited the applications of Granger causality to simple bivariate vector auto-regressive processes or pairwise relationships among a set of variables. Starting with a review of early developments and debates, this paper discusses recent advances that address various shortcomings of the earlier approaches, from models for high-dimensional time series to more recent developments that account for nonlinear and non-Gaussian observations and allow for sub-sampled and mixed frequency time series.

1 INTRODUCTION

Granger causality analyzes time-series interactions using temporal ordering and predictive relationships, but its traditional linear VAR formulation imposes restrictive assumptions. The review surveys advances that relax these restrictions for broader data types, dimensionalities, and sampling schemes.

  • Granger causality evaluates whether using one series’ history improves prediction of another, while restricting causal statements to past causing future.
  • Traditional analyses rely on linear VAR models and often test pairwise relationships, which can produce confounded inferences when many series interact.
  • The review covers network methods, high-dimensional analyses, discrete-valued and nonlinear series, and subsampled or mixed-frequency observations.
  • The standard framework assumes continuous-valued series, linear dynamics, known lag structure, regular discrete sampling, and sampling aligned with the causal time scale.
  • Real-world time series can violate these assumptions through nonlinear dynamics and irregular sampling, limiting where standard Granger causality applies.

2 THE HISTORY OF GRANGER CAUSALITY

The history of Granger causality centers on a predictability-based definition developed within identifiable linear models, alongside continuing debate about whether predictive relations establish causal effects. Early VAR-based methods were limited by assumptions about observability, dynamics, sampling, stationarity, lag structure, and dimensionality, motivating multivariate and broader approaches.

  • 2.1 Definition: Granger causality defines one series as causal of another when including its past improves the latter’s optimal prediction.
  • 2.1 Definition: The framework distinguishes Granger causality from formal causal definitions because it is based on predictability rather than directly establishing causal effects.
  • 2.1 Definition: Granger’s linear model is generally non-identifiable unless the contemporaneous matrix A0 is diagonal, yielding the simple VAR causal model.
  • 2.1 Definition: The VAR framework assumes continuous observations, linear dynamics, regular discrete sampling at the causal lag, known lag order, stationarity, perfect observation, and a complete system.
  • 2.1 Definition: These requirements are unlikely to hold simultaneously and are not verifiable, so causal effects may be unidentifiable or conclusions may be wrong.
  • 2.2 Early approaches and applications: Early methods focused on bivariate models, while network Granger causality sought to analyze larger variable sets and adjust for confounders.

3 NETWORK GRANGER CAUSALITY

Network Granger causality extends analysis beyond pairwise models to systems with many variables, representing causal relations through VAR lag-matrix structure. High-dimensional methods use sparsity, shrinkage, factor models, and lag-selection strategies, with consistency results supporting network recovery.

  • Motivation: Network Granger causality addresses settings where either many exogenous variables must be controlled or all variables are endogenous.The first setting seeks to prevent incorrect relations by incorporating relevant outside information; the second studies interactions across an entire system.
  • VAR formulation: In VAR models, xi is Granger-causal for xj when the corresponding lag-matrix entries contain a nonzero coefficient.This correspondence permits causal relations to be read from the zeros and nonzeros of autoregressive matrices.
  • VAR formulation: Network representations distinguish lagged Granger-causal edges from instantaneous dependencies encoded by the innovation precision matrix.Expanded graphs replicate variables across time, while compact graphs combine interactions across lags.
  • High-dimensional estimation: Factor-augmented VARs summarize many exogenous series with unobserved factors, estimated using principal components or maximum likelihood.This approach targets relationships among a small number of endogenous variables while incorporating broader information.
  • High-dimensional estimation: Sparse estimation replaces purely test-based approaches by directly selecting nonzero VAR coefficients through penalties such as lasso and group lasso.Lasso induces element-wise sparsity, whereas group penalties can select all lags for an interaction or larger related groups.
  • High-dimensional estimation: Consistency results connect high-dimensional VAR estimation to spectral eigen-structure and establish correct network selection for high-dimensional panel data.Basu and Michailidis relate the required sample size to the process spectrum, while Shojaie and Michailidis establish consistency for their selection algorithm.

4 MORE GENERAL NOTIONS OF GRANGER CAUSALITY

More general Granger-causality notions extend linear VAR analysis to nonlinear, non-Gaussian, discrete-valued, point-process, subsampled, and mixed-frequency time series. These extensions define causality through predictive irrelevance or conditional independence and address mismatches between observation and causal timescales.

  • Scope of classical definitions: Classical Granger causality is based on linear dynamics and real-valued Gaussian series, limiting analyses of nonlinear and discrete observations.The paper notes that many neuroscience and genomics interactions are inherently nonlinear and that count data motivate broader models.
  • Nonlinear notions: A component-wise nonlinear model extends Granger causality by treating xj as irrelevant to predicting xi when gi does not depend on xj’s history.This generalizes the VAR interpretation from zero coefficients to invariance of the relevant prediction function.
  • Nonlinear notions: Definition 1 characterizes Granger non-causality as invariance of the prediction function to the past of the candidate cause.The paper then notes that this formulation still assumes additive noise.
  • Nonlinear notions: Strong Granger causality further generalizes the framework through conditional independencies that can model arbitrary nonlinear relationships.A second definition expresses non-causality through equality of conditional distributions with and without the candidate series’ history.
  • Scope of advances: The review covers recent approaches for multivariate discrete-valued, nonlinear, point-process, subsampled, and mixed-frequency time series.These developments broaden the models and observation schemes to which Granger-causality analysis can be applied.
  • Sampling and frequency: Subsampling and mixed-frequency methods address data observed more slowly than the underlying causal scale or at different rates across series.The paper warns that slower observation can miss true interactions and add spurious ones, and reviews approaches for both settings.

4.1 Discrete-valued time series

Discrete-valued time series arise in applications involving counts, categories, and quantized continuous measurements, where traditional VAR-based Granger analysis is inappropriate. Recent models instead infer Granger causality through structured transition distributions, including MTD and mLTD formulations with tractable estimation and identifiability conditions.

  • 4.1 Discrete-valued time series: Discrete-valued series include count, binary, categorical, and quantized measurements from domains such as health, weather, finance, and sales.These data motivate models beyond traditional continuous-valued VAR representations.
  • 4.1 Discrete-valued time series: Traditional VAR-based Granger analysis is inappropriate for multivariate discrete-valued time series.The paper reviews models based on general Granger-causality definitions for these observations.
  • 4.1.1 Categorical time series: The MTD represents conditional probabilities as an intercept component plus weighted pairwise transition tables, with higher-order lags and interaction terms available.The weights form a probability distribution, and the intercept is critical for model identifiability.
  • 4.1.1 Categorical time series: A change-of-variables reparameterization makes MTD estimation convex and enables practical Granger-causality selection despite prior non-convexity and unknown identifiability conditions.The resulting optimization uses linear constraints, so it has no local optima separate from the globally optimal convex solution set.
  • 4.1.1 Categorical time series: For identifiable MTD models, xj is Granger non-causal for xi if and only if the columns of Zj are all equal; with the sparse special case, this is equivalent to Zj = 0.Equal columns mean the transition distribution does not depend on the lagged values of xj.
  • 4.1.1 Categorical time series: The reparameterized Zj entries quantify probability increases associated with lagged categorical states, while γj measures probability mass explained by xj.These parameters provide an interpretable dependence representation for categorical time series.
  • 4.1.2 Alternative formulation for categorical time series: The mLTD provides an alternative categorical transition model whose identifiability restrictions differ from MTD, while its non-causality interpretation is Zj = 0.The paper also notes that the reparameterized parameters support an interpretable notion of categorical dependence.

4.2 Methods for capturing interactions in non-linear time series

Nonlinear Granger-causality methods extend beyond linear VAR models, but differ in flexibility, interpretability, sample requirements, and handling of lag selection. Recent neural-network approaches use structured sparsity to identify nonlinear interactions and relevant lags.

  • Model-free methods: Model-free measures such as transfer entropy and directed information detect nonlinear dependencies with minimal predictive assumptions, but their estimators have high variance and require substantial data.These methods can therefore be inappropriate in high-dimensional settings because of the curse of dimensionality.
  • ODE-based methods: ODE-based approaches flexibly model nonlinear dynamics, but the additive formulations discussed here restrict the interaction mechanisms they can represent.System-identification variants use flexible functions f_i, and Granger non-causality corresponds to a zero component function f_ij.
  • Neural-network methods: Neural networks represent complex, nonlinear, and non-additive interactions, yet conventional autoregressive MLPs and RNNs are essentially black boxes with limited structural interpretability.Joint networks can also require many parameters and much data, perform poorly in high-dimensional settings, and impose common lag histories across series.
  • Structured neural models: Component-wise MLPs and LSTMs use group-based sparsity penalties on neural-network weights to identify Granger non-causal interactions.The MLP formulation can automatically detect nonlinear causality and each interaction’s lags, whereas the LSTM formulation avoids explicit lag selection by modeling long-range dependencies.
  • Structured neural models: The LSTM-based formulation sidesteps lag selection, while its structured penalties help address limited-data settings.The cMLP and cLSTM schematics encode non-causality through zero outgoing weights to the relevant output pathway.
  • Structured neural models: In component-wise MLPs, zeroing all first-layer weights associated with series x_j implies that x_j does not Granger-cause x_i.Group lasso penalizes outgoing weights across all lags, while hierarchical penalties emphasize higher lags and can produce sparse causal connections.

4.3 Subsampled and mixed-frequency time series

Subsampling and mixed-frequency observation can confound lagged and instantaneous Granger-causal effects, so analyses must model the observation structure explicitly. Under non-Gaussian independent shocks and stated assumptions, the underlying structural effects can remain identifiable.

  • Subsampling: Sampling below the process’s causal time scale can miss true interactions and introduce spurious lagged or instantaneous effects.Ignoring subsampling can make zero entries in the underlying lag matrix inconsistent with zero entries after aggregation through (A)^k.
  • Subsampling: A subsampled process has transition matrix (A)^k and aggregates shocks across unobserved time points, producing a higher-dimensional structural representation.The resulting shock vector has dimension kp and special structure in its structural matrix and error distributions.
  • Subsampling: Classical analysis that ignores subsampling can incorrectly estimate lagged and instantaneous interactions.One illustrated analysis finds no lagged effect between x1 and x2 but a relatively large instantaneous interaction.
  • Mixed-frequency sampling: Mixed-frequency data use series-specific sampling rates and indicator matrices to select observed time points, yielding a structural process analogous to the subsampled case.The mixed-frequency parameterization is written as (A, C, p_e; k), with k a vector of sampling rates.
  • Identifiability: Tank et al. show that accounting for subsampling and mixed-frequency structure can permit direct estimation of the underlying A and C matrices from observed data.The identifiability result assumes non-Gaussian independent shocks and additional stated assumptions.
  • Identifiability: Mixed-frequency sampling can resolve parameter ambiguities in non-Gaussian models, including identification of A_ij when series x_j and x_i differ by one sampling time step.For subsampled structural models with acyclic instantaneous effects, the causal structure can be identified without prior causal ordering.

5 CONCLUSION

The paper reviews how restrictive assumptions and historically simple methods have limited Granger-causality analyses, then surveys advances that broaden applicability. Despite progress, flexible approaches for complex, non-stationary, high-dimensional time series and richer causal evidence remain needed.

  • 5 CONCLUSION: Classical Granger-causality approaches were constrained by restrictive assumptions and simple methods for investigating causal relations.These limitations reflect assumptions needed to infer causal effects from time-series data and historically used analytical approaches.
  • 5 CONCLUSION: Recent work relaxes classical assumptions and generalizes Granger-causality analysis to larger and more diverse time-series settings.The reviewed advances include analyses with many variables, automatic lag selection, non-stationarity, and non-Gaussian or non-continuous observations.
  • 5 CONCLUSION: Recent developments also address differences between the true causal time scale and the frequency at which data are observed.This extends the framework toward settings involving sampling-frequency mismatches.
  • 5 CONCLUSION: These developments have expanded Granger-causality application domains and created opportunities to study interactions in complex systems from a systems perspective.The paper connects broader applicability with investigating interactions among components of complex systems.
  • 5 CONCLUSION: Important open needs include flexible nonparametric methods handling many observed series, unmeasured variables, and non-stationarity.The conclusion also points to interventions and perturbations as emerging data sources for discovering causal effects and restricting possible causal hypotheses.
Loading 2105.02675v2…