Source-linked AI summary
Discovering Graphical Granger Causality Using the Truncating Lasso Penalty
Ali Shojaie, George Michailidis
TL;DR
The paper addresses discovery of gene-regulatory causal relationships from time-course expression data, especially when the number of genes is large relative to the sample size. It introduces truncating lasso for graphical Granger models, showing consistent edge, sign, and VAR-order estimation while retaining regulatory time-lag information. The method improves on lasso-type estimates, particularly for small to moderate sample sizes, but relies on sufficiently fine-grained repeated observations.
Problem
Discovering causal gene-regulatory relationships from time-course expression data is important for understanding cellular mechanisms and improving biological estimation and inference.
Method
The paper introduces a truncating lasso penalty for graphical Granger models that determines effective VAR time lags, reduces covariates, and supports efficient estimation.
Results
The method consistently estimates variable selection, effect signs, and VAR order in high-dimensional sparse settings and improves estimates over lasso and adaptive lasso, especially for small to moderate sample sizes.
Takeaways & Limitations
The resulting estimates provide information about the time lags of regulatory effects while reducing unnecessary covariates and improving false positive and false negative rates.
Takeaways & Limitations
Reliable reverse engineering requires repeated observations on fine time grids because coarse intervals can make some causal effects indistinguishable.
Abstract
from arXiv · showhide
Components of biological systems interact with each other in order to carry out vital cell functions. Such information can be used to improve estimation and inference, and to obtain better insights into the underlying cellular mechanisms. Discovering regulatory interactions among genes is therefore an important problem in systems biology. Whole-genome expression data over time provides an opportunity to determine how the expression levels of genes are affected by changes in transcription levels of other genes, and can therefore be used to discover regulatory interactions among genes. In this paper, we propose a novel penalization method, called truncating lasso, for estimation of causal relationships from time-course gene expression data. The proposed penalty can correctly determine the order of the underlying time series, and improves the performance of the lasso-type estimators. Moreover, the resulting estimate provides information on the time lag between activation of transcription factors and their effects on regulated genes. We provide an efficient algorithm for estimation of model parameters, and show that the proposed method can consistently discover causal relationships in the large $p$, small $n$ setting. The performance of the proposed model is evaluated favorably in simulated, as well as real, data examples. The proposed truncating lasso method is implemented in the R-package grangerTlasso and is available at http://www.stat.lsa.umich.edu/~shojaie.
1 Introduction
Time-course gene expression data can support discovery of gene-regulatory causal relationships, but existing Granger-based methods face high-dimensionality and interpretability challenges. The paper addresses these issues with a truncating lasso penalty that determines effective time lags and simplifies graphical Granger models.
- Motivation: Time-course expression data can reveal regulatory relationships because changes in one gene’s expression may affect another gene’s expression over time.Graphical Granger causality formalizes this by testing whether past values of one gene improve prediction of another beyond that gene’s own history.
- Existing approaches: Dynamic Bayesian networks and Granger causality are established approaches for inferring causal relationships from time-series data.Dynamic Bayesian networks represent variables across time points, while Granger causality compares predictive autoregressive models.
- High-dimensional setting: When p ≫ n, sparse penalized models are needed because penalization can improve prediction accuracy with many genes and few observations.Prior work applied lasso and sparse VAR models to graphical Granger causality and gene-regulatory network discovery.
- Limitations of prior penalties: Lasso estimates may assign influence across multiple time lags, making models difficult to interpret and potentially harming model selection when covariates are added.Group lasso simplifies the model by averaging effects across lags, but loses lag, sign, and effect-magnitude information.
- Proposed approach: The proposed truncating lasso determines the VAR order, reduces covariates, and provides an efficient estimation algorithm with consistency guarantees.The method is evaluated on simulated and real data and is reported to provide better estimates than alternative penalization methods.
2 Model and Methods
The paper formulates graphical Granger causality through sparse, lagged regression models and introduces truncating lasso to select relevant variables and effective time lags. An iterative algorithm estimates the non-convex model, with consistency, false-positive control, and computational guarantees.
- 2.1 Graphical Models and Penalized Estimates of DAGs: Graphical models encode variables as nodes and directed causal relationships as edges, represented through an adjacency matrix.
- 2.2 Graphical Granger Causality: Granger causality declares X causal for Y when including both variables’ histories significantly improves prediction over using Y’s history alone.
- 2.3 Truncating Lasso for Graphical Granger Models: Weighted lasso estimates graphical Granger models through p separate ℓ1-regularized regressions, but may include many covariates across unknown time lags.The weighted lasso can use up to T − 1 lags, producing p(T − 1) covariates and potentially retaining edges from multiple time points.
- 2.3 Truncating Lasso for Graphical Granger Models: Truncating lasso automatically estimates the VAR order and simplifies the model by reducing covariates while retaining time-lag information.Its truncation factor forces later estimates to zero when the preceding estimate contains sufficiently few edges.
- 2.4 Choice of the Tuning Parameter: The estimator is consistent for variable selection, effect signs, and VAR-order recovery in high-dimensional sparse settings.The tuning parameter can also control a version of the false-positive rate at a specified level under scaled design columns.
- 2.5 Algorithm and Computational Complexity: Block-relaxation solves the non-convex problem through convex weighted-lasso subproblems, with each subproblem requiring O(np^2) operations.The algorithm converges to a stationary point and often converges in fewer than 10 iterations; for large T, it may be faster than lasso.
3 Results
Simulations and biological-network analyses evaluate truncating-lasso estimators against lasso-type and search-based alternatives, showing improved estimation and a time-lag advantage.
- Simulation Studies: 50 simulations found TAlasso to provide the best estimate across the evaluated criteria, with truncation advantages increasing for longer time series.The advantage was particularly significant for small sample sizes and diminished with large n.
- Simulation Studies: Truncation improved both false-positive and false-negative rates by reducing the d × p covariates considered in the graphical Granger model.In the p = 20 network, lasso and adaptive lasso included edges beyond the true VAR order and missed some true edges.
- E-coli Regulatory Network: TAlasso improved recall and achieved a higher F1 measure than Alasso on the E-coli regulatory-network example.The comparison used known and estimated regulatory networks evaluated with performance measures.
- HeLa BioGRID Network: CNET achieved the best performance on the HeLa BioGRID example, while TAlasso performed slightly better than group lasso.Penalization methods were less accurate than CNET but were described as more computationally efficient for large networks.
- HeLa BioGRID Network: The truncating-lasso estimate also reports effective time lags for regulatory effects, information overlooked by the other two methods.Table 1 provides details on the effective time lags of gene effects in the network.
4 Discussion
The paper frames truncating lasso as a method for estimating gene regulatory networks that can improve estimation while revealing regulatory time lags. Its use depends on sufficiently fine, repeated time-course observations.
- Gene regulatory networks can improve estimation and inference about pathways involved in environmental responses or disease progression.
- Truncating lasso correctly determines the underlying time-series order and uses it to reduce covariates, improving false positive and false negative rates.
- The method provides information on the time lags of regulatory effects between genes, including effective lags summarized for the BioGRID network.
- Repeated time-series observations on fine time grids are required because coarse intervals can make some underlying causal effects indistinguishable.
- The method offers significant improvements over lasso and adaptive lasso estimates, especially for small to moderate sample sizes.The improvement is attributed to excluding unnecessary covariates from the regression problem.
Appendix
The appendix establishes consistency properties for truncating adaptive lasso under high-dimensional assumptions. With asymptotically high probability, it excludes incorrect effects, estimates signs correctly, identifies true effects, and recovers VAR order.
- Under stated growth, variance, and partial-correlation assumptions, the consistency theorem analyzes truncating adaptive lasso in a high-dimensional setting.
- With probability converging to 1, no additional Granger-causal effects are included and the signs of the effects are correctly estimated.
- With probability asymptotically larger than 1−β, true Granger-causal effects and the VAR model order are correctly determined.
- The proof uses truncation after effective effects diminish over time, while adaptive lasso recovers true edges with exponentially large probability.
- False positives before the truncation threshold occur with exponentially small probability, supporting correct recovery of the VAR order.