Source-linked AI summary
Sparse and compositionally robust inference of microbial ecological networks
Zachary D. Kurtz, Christian L. Mueller, Emily R. Miraldi, Dan R. Littman, Martin J. Blaser, Richard A. Bonneau
TL;DR
Microbial network inference must handle compositional abundances and high-dimensional, under-powered data, while distinguishing direct from indirect associations. SPIEC-EASI combines compositional transformations with sparse graphical-model estimators and synthetic benchmarking, outperforming state-of-the-art methods in realistic simulations and producing reproducible networks on American Gut data.
Problem
Compositional OTU abundances and p > n dimensionality complicate reliable microbial interaction inference, while correlations can reflect indirect relationships.
Method
SPIEC-EASI applies compositionally robust transformations and uses neighborhood selection or sparse inverse covariance selection to infer conditional-dependence networks.
Results
SPIEC-EASI outperforms SparCC and CCREPE in interaction recovery and global network features across almost all tested realistic benchmark scenarios, while covariance selection is competitive in some settings.
Takeaways & Limitations
SPIEC-EASI provides a framework for reproducible microbial network inference and for incorporating additional biological information into future modeling.
Takeaways & Limitations
Inference relies on approximating the inverse covariance structure from transformed data and adding unit pseudocounts to handle zero counts.
Abstract
from arXiv · showhide
16S-ribosomal sequencing and other metagonomic techniques provide snapshots of microbial communities, revealing phylogeny and the abundances of microbial populations across diverse ecosystems. While changes in microbial community structure are demonstrably associated with certain environmental conditions, identification of underlying mechanisms requires new statistical tools, as these datasets present several technical challenges. First, the abundances of microbial operational taxonomic units (OTUs) from 16S datasets are compositional, and thus, microbial abundances are not independent. Secondly, microbial sequencing-based studies typically measure hundreds of OTUs on only tens to hundreds of samples; thus, inference of OTU-OTU interaction networks is severely under-powered, and additional assumptions are required for accurate inference. Here, we present SPIEC-EASI (SParse InversE Covariance Estimation for Ecological Association Inference), a statistical method for the inference of microbial ecological interactions from metagenomic datasets that addresses both of these issues. SPIEC-EASI combines data transformations developed for compositional data analysis with a graphical model inference framework that assumes the underlying ecological interaction network is sparse. To reconstruct the interaction network, SPIEC-EASI relies on algorithms for sparse neighborhood and inverse covariance selection. Because no large-scale microbial ecological networks have been experimentally validated, SPIEC-EASI comprises computational tools to generate realistic OTU count data from a set of diverse underlying network topologies. SPIEC-EASI outperforms state-of-the-art methods in terms of edge recovery and network properties on realistic synthetic data under a variety of scenarios. SPIEC-EASI also reproducibly predicts previously unknown microbial interactions using data from the American Gut project.
Introduction
Microbial interaction inference is difficult because sequencing data are compositional and high-dimensional, while correlations may reflect indirect relationships. SPIEC-EASI addresses these challenges by combining compositional transformations with sparse graphical-model inference and realistic synthetic benchmarking.
- Motivation: Microbiome studies seek to infer ecological interactions from OTU population data, where OTU counts proxy underlying microbial abundances.Interactions are typically inferred as significant, usually nondirectional associations between sampled populations.
- Challenges: Traditional correlation analysis can produce spurious results because normalized OTU abundances are compositional and therefore not independent.Correlation may also connect OTUs that are indirectly related in the ecological network.
- Challenges: With hundreds to thousands of OTUs but only tens to hundreds of samples, interaction inference operates in the underdetermined regime p > n and requires network assumptions.The dimensionality challenge motivates methods designed for high-dimensional data.
- Method: SPIEC-EASI transforms compositional OTU data and estimates interaction graphs using neighborhood selection or sparse inverse covariance selection.Its graphical-model formulation targets conditional independence rather than empirical correlation or covariance.
- Benchmarking: SPIEC-EASI includes a synthetic-data generator that produces realistic OTU counts from networks with diverse topologies for benchmarking inference methods.This addresses the absence of experimentally validated large-scale microbial interaction networks and limitations of earlier synthetic benchmarks.
- Results: SPIEC-EASI outperforms state-of-the-art methods on interaction recovery and network features across diverse realistic synthetic scenarios, while producing stable, reproducible networks on real data.American Gut analyses also identify family-associated clusters and composite network topologies.
Materials and Methods
SPIEC-EASI combines compositional-data transformation with sparse graphical-model inference to reconstruct microbial association networks, while also generating realistic synthetic OTU datasets for benchmarking.
- Pipeline overview: SPIEC-EASI contains separate inference and synthetic-data-generation modules for microbial ecological network analysis.The inference module analyzes real or synthetic OTU data, while the generation module creates benchmark datasets.
- Data processing and transformation: Count normalization creates unit-sum compositions whose components are dependent, making standard covariance and regression methods unsuitable.Closure effects can induce negatively biased covariance estimates.
- Data processing and transformation: The centered log-ratio transform maps compositions into Euclidean space and provides the basis for estimating dependence structure from transformed OTU data.A pseudo-count is added before transformation to avoid numerical problems caused by zero counts.
- Inference of microbial associations: Network inference uses neighborhood selection or penalized inverse covariance selection, both formulated as convex optimization procedures applicable when p > n under structural assumptions.The resulting graphical model targets conditional associations rather than correlations explainable through alternate network paths.
- Inference of microbial associations: Stability-based model selection chooses sparsity along the λ-dependent solution path, while inverse covariance selection also yields covariance estimates for downstream analyses.The tuning parameter λ controls the sparsity of the inferred model.
- Generation of synthetic microbial abundance datasets: Synthetic data combine fitted OTU marginal distributions with user-selected network topologies through the Normal to Anything approach.The pipeline fits marginal distributions to real OTU data and generates band-like, cluster, or scale-free graph structures.
Results
The study benchmarks SPIEC-EASI against existing methods on realistic synthetic data and American Gut data, evaluating edge recovery, global network properties, reproducibility, and sparsity. SPIEC-EASI generally performs best, especially S-E(MB), but recovery depends strongly on sample size and network topology.
- Synthetic benchmarking: Synthetic datasets modeled on American Gut data varied topology, association strength, and sample number to benchmark SPIEC-EASI against SparCC, CCREPE, and Pearson correlation.The evaluation used realistic OTU count data and precision-recall metrics across 960 independent synthetic datasets.
- Synthetic benchmarking: Performance improved with sample size and depended strongly on topology, with recovery best for band graphs, followed by cluster and scale-free graphs.The authors report that recovery performance nearly doubled across these topology classes at fixed sample size, taxa number, and condition number.
- Synthetic benchmarking: SPIEC-EASI methods performed as well as or significantly better than controls in most scenarios, with S-E(MB) superior to S-E(glasso).At n = 1360, S-E(MB) was the only method recovering a significant portion of edges under all tested scenarios.
- Interpretation: The authors caution that complete recovery is likely unrealistic with at most hundreds of samples, although high-confidence S-E(MB) edges achieved very high precision across network types.Final network construction still requires confidence-based edge selection, for which no optimal process exists.
- Global network properties: S-E(MB) recovered degree distributions better than all other methods, including at smaller sample sizes, while its advantage for betweenness centrality emerged mainly at n = 1360.At n = 1360, S-E(MB) was significantly better in five of six conditions; SparCC was best for scale-free networks with κ = 10.
- Global network properties: S-E methods had lower errors for connected-component recovery in band and scale-free networks, although all methods predicted too many components for cluster networks.The component-recovery comparison covered varying sample sizes and network topologies.
- American Gut network: On American Gut data, SPIEC-EASI produced more reproducible and sparser networks than SparCC and CCREPE, with S-E(MB) showing roughly 50 versus CCREPE’s 250 edge disagreements.All methods shared a 127-edge core network, and inferred interactions were largely assortative within taxonomic groups.
Discussion
SPIEC-EASI addresses compositionality and high dimensionality in microbial interaction inference through compositionally robust graphical models and synthetic-data benchmarking. It improves network inference in benchmarks and real American Gut data, while network topology remains an important boundary on recovery.
- SPIEC-EASI combines compositionally robust transformations with neighborhood and sparse inverse-covariance selection for 16S microbiome data.
- Synthetic benchmarking found S-E(MB) generally outperformed SparCC and CCREPE for recovering interactions and global network topology.
- Network topology and direct-interaction strength affect recovery beyond total sample size, informing sample-size planning and statistical-power assessment.
- American Gut analyses produced more consistent and sparser networks than SparCC and CCREPE, with greater interaction likelihood among phylogenetically related OTUs.
- Scale-free network structure can elude accurate inference despite global sparsity, although prior network or phylogenetic information can be incorporated into the framework.
- Regularized inverse-covariance estimates may support models of taxa responses to experimental design factors by representing correlations among taxa.
- Association networks can provide topology for dynamic gut-microbiome models used to formulate perturbation-response hypotheses.
- The authors conclude that SPIEC-EASI improves state-of-the-art microbial network inference and can support more sophisticated ecological and medical modeling.
A.1 Method comparison
The comparison table contrasts SPIEC-EASI with SparCC, CCREPE, and Pearson correlation across compositional correction, dependence concepts, and network assumptions.
- SPIEC-EASI, SparCC, and CCREPE apply compositional correction, whereas Pearson correlation does not.
- SPIEC-EASI uses conditional independence and assumes network sparsity, while SparCC assumes average correlation is zero.
A.2 Comparative marginal fits to American Gut Project count data
The synthetic-data pipeline fits marginal count distributions to American Gut Project OTU data after normalization and filtering.
- After filtering, p = 205 OTU count vectors from n = 558 samples were available for marginal fitting.
- Each OTU count vector was modeled independently using maximum-likelihood estimation.
A.2.1 Common count distributions
SPIEC-EASI’s generator considers common count distributions, including models that represent overdispersion, continuous abundance variation, and excess zeros.
- Five common distributions were considered for modeling univariate OTU count data.
- Log-normal: The log-normal model represents a continuous variable whose logarithm is normally distributed, with discrete counts obtained by rounding.
- Poisson: The Poisson model describes discrete event counts, with λ determining both the mean and variance.
- Zero-inflated Poisson: Zero-inflated Poisson models add an excess-zero component for count data with more zeros than a single-component model can represent.
- Negative Binomial: The negative binomial accommodates overdispersion through a hierarchical Poisson mixture, and approaches Poisson as r →∞.
- Zero-inflated Negative Binomial: Zero-inflated negative binomial augments the negative binomial with excess zeros; φ = 0 reduces it to the negative binomial.
A.2.2 Goodness-of-fit to the AGP data
The AGP OTU count data were compared with five candidate distributions using QQ plots. The zero-inflated Negative Binomial model was selected because it fit both tails accurately.
- A.2.2 Goodness-of-fit to the AGP data: The analysis compared AGP count data with five statistical models using log-log QQ plots.Synthetic data were generated from maximum-likelihood-fitted distributions, and the QQ plots compared synthetic with real data.
- A.2.2 Goodness-of-fit to the AGP data: r2 = 0.985 for the ziNB distribution, the only model accurately modeling both tails of the OTU data.The ziNB distribution was therefore used as SPIEC-EASI’s standard data-generation setting.
A.3 Heatmaps of AGP and synthetic data
Heatmaps compare American Gut Project data with synthetic datasets across band, graph, and scale-free network types. The synthetic datasets are consistent with the real data in dimensions and abundance patterns.
- A.3 Heatmaps of AGP and synthetic data: Synthetic datasets matched AGP data in OTU number, sample number, and OTU count abundances across samples.The heatmaps display log10(counts) for band, graph, and scale-free network types.
A.4 Relationship between target and empirical correlation after NorTA transformation
Figure 9 evaluates how well the NORTA process recovers target Pearson and inverse correlations from zero-inflated Negative Binomial data. The comparison spans untransformed and log-transformed counts.
- A.4 Relationship between target and empirical correlation after NorTA transformation: The figure compares recovered Pearson correlations with input multivariate Normal correlations for zero-inflated Negative Binomial data.Upper panels show empirical correlations, while lower panels show inverse correlations.
- A.4 Relationship between target and empirical correlation after NorTA transformation: The comparison covers untransformed and log-transformed counts using p = 205 OTUs, n = 20,000 samples, and 10 replicates per plot.These settings define the simulation shown in Figure 9.
A.5 Effect of Condition Number on Correlations Distribution
Figure 10 relates Precision-matrix condition number to the correlation distribution of synthetic networks. Higher condition numbers correspond to stronger network correlations.
- A.5 Effect of Condition Number on Correlations Distribution: The condition number is defined as the ratio of the largest to smallest eigenvalue or singular value of the Precision matrix.Figure 10 examines its relationship with the network’s correlation distribution.
- A.5 Effect of Condition Number on Correlations Distribution: Increasing condition number corresponds to increasing correlation strength in the synthetic network.The relationship links a matrix property to the distribution of network correlations.
A.6 Betweenness Centrality for Synthetic Networks
The figures assess whether SPIEC-EASI recovers structural properties of synthetic microbial networks, including betweenness centrality, geodesic distances, and cluster sizes. They also show comparative inference results and taxonomic assortativity across methods.
- Betweenness centrality: Betweenness centrality distributions predicted by S-E(MB) are compared with the distributions of each synthetic network type.The examples use κ = 100, n = 1360 samples, and p = 205 OTUs.
- Betweenness centrality: Betweenness-centrality recovery compares S-E(glasso), S-E(MB), SparCC, CCREPE, and Pearson using average KL-divergence across synthetic datasets.Bars aggregate three independent dataset sets, each containing seven datasets, with standard-error bars.
- Geodesic distance: Geodesic-distance distributions predicted by S-E(MB) are evaluated against distributions generated from each network type.The examples use κ = 100, n = 1360 samples, and p = 205 OTUs.
- Geodesic distance: Geodesic-distance recovery is compared across S-E(glasso), S-E(MB), SparCC, CCREPE, and Pearson using average KL-divergence.Asterisks denote significantly better recovery by an S-E method than each control method at p < 0.05.
- Cluster size: Cluster-size distributions predicted by S-E(MB) are compared with the distributions associated with each synthetic network type.The examples use κ = 100, n = 1360 samples, and p = 205 OTUs.
- Cluster size: Cluster-size recovery compares S-E(glasso), S-E(MB), SparCC, CCREPE, and Pearson using average KL-divergence across synthetic datasets.Asterisks indicate significantly better recovery by an S-E method than each control method at P < 0.05.