Source-linked AI summary
Estimating sample-specific regulatory networks
Marieke Lydia Kuijjer, Matthew Tung, GuoCheng Yuan, John Quackenbush, Kimberly Glass
TL;DR
Existing network methods usually produce aggregate models that do not capture regulatory heterogeneity among samples. This paper introduces LIONESS, which uses linear interpolation between aggregate models to estimate sample-specific networks, and demonstrates its accuracy and applicability across simulated and experimental datasets. The resulting networks support analyses of temporal topology and regulatory changes that may not be apparent in expression data.
Problem
Existing network methods generally produce one aggregate network from many samples, limiting representation of regulatory heterogeneity across individuals and biological states.
Method
LIONESS estimates each sample’s network by linearly interpolating predictions from aggregate models reconstructed with and without that sample.
Results
LIONESS demonstrated accuracy and applicability across simulated data, synchronized-yeast microarrays, and human lymphoblastoid-cell-line RNA-seq, including associations with chromatin state and biologically enriched differential targeting.
Takeaways & Limitations
Sample-specific regulatory networks provide a view of biological systems distinct from, but complementary to, other multi-omic data and can reveal temporal and regulatory changes not apparent in expression data.
Takeaways & Limitations
Phenotypic information was limited for the 65 individuals studied, and available phenotypes were not significantly associated with expression- or network-based clustering groups.
Abstract
from arXiv · showhide
Biological systems are driven by intricate interactions among the complex array of molecules that comprise the cell. Many methods have been developed to reconstruct network models of those interactions. These methods often draw on large numbers of samples with measured gene expression profiles to infer connections between genes (or gene products). The result is an aggregate network model representing a single estimate for the likelihood of each interaction, or "edge," in the network. While informative, aggregate models fail to capture the heterogeneity that is represented in any population. Here we propose a method to reverse engineer sample-specific networks from aggregate network models. We demonstrate the accuracy and applicability of our approach in several data sets, including simulated data, microarray expression data from synchronized yeast cells, and RNA-seq data collected from human lymphoblastoid cell lines. We show that these sample-specific networks can be used to study changes in network topology across time and to characterize shifts in gene regulation that may not be apparent in expression data. We believe the ability to generate sample-specific networks will greatly facilitate the application of network methods to the increasingly large, complex, and heterogeneous multi-omic data sets that are currently being generated, and ultimately support the emerging field of precision network medicine.
1. INTRODUCTION
Complex traits and diseases reflect heterogeneous, network-level regulatory processes rather than isolated molecular changes. Existing network methods typically produce one aggregate network, motivating LIONESS to estimate sample-specific networks that capture individual variation.
- 1. INTRODUCTION: Complex phenotypes often arise from altered biological networks whose structures change across cell states.Studying these changing connections may reveal mechanisms involved in disease.
- 1. INTRODUCTION: Large multi-omic and single-cell resources have highlighted regulatory diversity across cells, tissues, phenotypes, and environmental exposures.Individual-specific variations with relatively small effects may collectively contribute to disease.
- 1. INTRODUCTION: Network-level heterogeneity is important because biological processes, rather than individual molecules alone, help mediate phenotypic diversity.Capturing this heterogeneity supports increasingly individualized analyses of gene expression and regulation.
- 1. INTRODUCTION: Most existing network methods combine many samples but estimate a single aggregate network representing regulatory processes shared across the population.Such models do not capture sample-specific regulatory processes.
- 1. INTRODUCTION: LIONESS reverse engineers sample-specific networks by linearly interpolating predictions from existing aggregate network inference approaches.The method was evaluated on simulated, yeast microarray, and human lymphoblastoid-cell-line RNA-seq data.
2. METHODS
The methods frame aggregate networks as combinations of sample-level networks and use LIONESS to recover each sample’s contribution. The approach preserves complex within-network interactions while using a linear framework to compare networks across samples.
- 2. METHODS: Biological networks represent genes as nodes and inferred gene interactions as edges, while network heterogeneity reflects differences among cells and individuals.The methods distinguish relationships within a network from relationships between multiple networks.
- 2. METHODS: LIONESS applies linear interpolation to estimate a sample-specific network from aggregate models reconstructed with and without that sample.Figure 1 illustrates this leave-one-sample-out comparison.
- 2. METHODS: An aggregate network from N samples is modeled as the average of individual component networks, with each sample contributing to aggregate edge weights.This linear framework relates a set of networks representing different biological samples.
- 2. METHODS: LIONESS estimates an individual edge by taking the difference between all-sample and leave-one-out aggregate edge scores, scaling it, and adding it to the leave-one-out score.The derivation assigns equal sample weights q = 1/N in the analyses, although alternative weighting is possible.
- 2. METHODS: The LIONESS framework is independent of the aggregate inference method and can wrap approaches such as Pearson correlation, mutual information, PANDA, and CLR.The aggregate model can therefore retain complex or nonlinear biological relationships while LIONESS estimates sample-level differences.
3. RESULTS
LIONESS accurately reconstructs sample-specific networks across simulated and experimental data, revealing reproducible network variation and regulatory patterns that expression analyses alone may miss.
- 3.1. LIONESS accurately and reproducibly predicts networks using in silico data: LIONESS accurately predicts individual sample networks across heterogeneous in silico data, outperforming aggregate networks as heterogeneity increases.At low heterogeneity, LIONESS and aggregate-network accuracy are similar; sample-specific permuted edges are poorly captured by aggregate models.
- 3.1. LIONESS accurately and reproducibly predicts networks using in silico data: LIONESS performance remains independent of network size and retains predictive ability when expression data contain noise.
- 3.1. LIONESS accurately and reproducibly predicts networks using in silico data: As sample size increases, LIONESS accuracy remains constant, whereas aggregate models improve for baseline edges but poorly estimate sample-specific permuted edges.This analysis used subsets ranging from N = 20 to 5000 input samples and evaluated both overall and permuted-edge accuracy.
- 3.1. LIONESS accurately and reproducibly predicts networks using in silico data: LIONESS consistently predicts single-sample networks across Pearson correlation, PANDA, mutual information, and CLR aggregate-inference methods.
- 3.2. Estimating single-sample networks using experimental data from yeast: Across yeast technical replicates, LIONESS networks show strong reproducibility, although some methods, especially CLR, are sensitive to the background data.PANDA networks showed higher replicate similarity than the other reconstruction approaches and were used in subsequent analyses.
- 3.2. Estimating single-sample networks using experimental data from yeast: Yeast cell-cycle networks display oscillatory edge-weight patterns, with highly variable edges originating from MBP1, SWI4, SWI6, and STB1.Edge-weight oscillations occur at twice the frequency of target-gene expression oscillations, whose period approximately matches the yeast cell cycle.
- 3.5. Complex relationships between network targeting and gene expression: Expression-modified LIONESS targeting strongly correlates with promoter-DNase events, while transcription-factor targeting relationships can be nonlinear and activator- or repressor-like.The targeting-expression relationship is strongest at high or low target-gene expression, and additional regulatory mechanisms may contribute.
- 3.7. Differential-targeting of genes highlights important biological processes: Network targeting can distinguish cellular-process groups when expression-based analysis cannot, identifying enrichment for cell proliferation and immune function.Expression-based clustering found 2620 differentially expressed genes but no enriched biological functions, whereas differential targeting produced functional enrichment.
4. DISCUSSION
LIONESS estimates sample-specific regulatory networks by using how adding or removing one sample perturbs an aggregate network model. That perturbation is used to infer a sample’s contribution and corresponding network.
- LIONESS estimates sample-specific regulatory networks from aggregate network models.
- Adding or removing a single sample slightly perturbs the aggregate network model.
- The perturbation estimates each sample’s contribution to the aggregate network and therefore its network.
C Reactome Pathways
The discussion positions LIONESS as a general framework for estimating complete sample-specific networks and applying network-level analyses to heterogeneous biological data.
- The framework can be applied beyond gene-expression regulatory networks to other network inference methods and multi-sample omics data.
- LIONESS could support precision-medicine analyses by associating predicted network interactions and properties with patient phenotypes, genotypes, progression, survival, or drug response.
- LIONESS is presented as a method for estimating each sample’s complete network rather than repurposing differential-expression information.
8. SUPPLEMENTAL MATERIALS AND METHODS
Existing single-sample network analyses emphasize sample-specific expression or correlation changes, whereas LIONESS estimates complete networks while retaining relationships common across samples.
- Single-sample differential-expression and differential-correlation approaches cannot estimate network structures common across all samples.
- ssMARINa and VIPER estimate regulator or protein activity from sample-specific differential-expression signatures but do not estimate the networks producing that activity.
- DERA colors genes in a prior biological network using sample-specific differential-expression results, so sample-specific edges absent from the prior network cannot be identified.
- ssPCC assumes every edge has the same weight distribution across predicted single-sample models and filters edges using a known biological network.
- LIONESS differs by estimating each sample’s complete network instead of repurposing differential-expression information for network analysis.
- The compared approaches may overlook nonlinear dynamics and synergistic effects that characterize regulatory networks.
8.2. Aggregate network reconstruction approaches
The analyses apply LIONESS to four network reconstruction approaches spanning linear and nonlinear association measures, finding robust sample-level weighting and stronger accuracy potential for PANDA.
- The study evaluates Pearson correlation, PANDA, mutual information, and CLR as representative network reconstruction approaches.
- Pearson correlation measures linear relationships between transcription-factor and target-gene expression levels, whereas mutual information captures associations without assuming linearity or continuity.
- PANDA integrates a prior regulatory network with gene-expression and protein-interaction data through message passing.
- 0.9702 was the average LIONESS edge-weight for samples in the simulated strong-linear relationship.
- Mutual information identified and down-weighted samples deviating from the expected relationship more readily than Pearson correlation, while both approaches remained robust to sample number.
8.4. Data generation, normalization, and relevant pre-processing steps
The analyses generated heterogeneous sample-specific network models and matched expression data, then evaluated LIONESS across network size, edge permutation, expression noise, and sample count.
- 1000 Boolean initial states per randomized network were propagated to attractors, whose averages defined matched node-expression levels.This process was repeated across N randomized versions of the seed network.
- Networks with 100, 250, or 625 nodes were generated under six edge-permutation levels, with ρ ranging from 0.5 to 3.The permutation level controlled how different each sample network was from the seed network.
- LIONESS accurately estimated permuted sample-specific edges, whereas the aggregate network failed especially when heterogeneity was low.The LIONESS accuracy for permuted edges was similar to that for non-permuted edges.
- Expression-noise and sample-number analyses tested LIONESS under added Gaussian noise and subsets of up to 10,000 samples.The sample-number analysis used M = 100 nodes and ρ = 1.
Processing the yeast cell cycle expression data
The supplied processing passages describe preprocessing of yeast expression data and human RNA-seq and DNase datasets, including normalization, alignment, quality control, filtering, and motif construction.
- Yeast expression data were normalized separately by replicate, batch-corrected, and averaged across probe sets, yielding 5088 genes across 50 samples.The samples comprised 25 from each of two replicate datasets.
- A yeast motif prior was built from predicted binding sites for 204 transcription factors, covering 3551 genes and 105 targeting factors on the expression array.The prior used genes with tandem promoters and array coverage.
- RNA-seq processing aligned reads to hg19, counted and summarized gene reads, and removed poor-quality samples and samples lacking suitable DNase data.The retained dataset contained 153 samples from 65 cell lines.
- RNA-seq genes were retained when raw counts occurred in at least 50% of samples, with library-size and gene-length corrections applied.The count filter retained 21516 of 57820 genes.
- DNase peaks were called and mapped to nearby genes, producing promoter-DNase scores for 12424 genes with at least one promoter peak.Technical replicates were averaged before filtering to genes with RNA-seq and motif-prior data.
8.5. Analysis of the human single-sample networks
Human single-sample networks were analyzed by converting expression, LIONESS edge weights, and motif evidence into probabilities, clustering samples, and testing group differences in targeting and expression.
- The analysis defined gene degree in three ways when comparing network structure with DNase information.The supplied passage introduces the three-way degree calculation without specifying the three definitions.
- TF-expression probabilities came from inverse-CDF-transformed Z-scores relative to expression in other samples.This provided a sample-specific expression measure for each transcription factor.
- LIONESS edge probabilities were obtained by inverse-CDF-transforming predicted edge weights expressed in Z-score units.Motif-edge probabilities were defined separately from motif evidence.
- Motif-edge probabilities were binary, indicating whether a transcription-factor motif was found in a target promoter.The motif probability was 0 or 1 for each TF–gene pair.
- Hierarchical clustering used row-normalized edge weights or expression, Spearman distances, complete linkage, and a two-group dendrogram cut.Spearman correlation was selected because network edges are often non-normal.
- LIMMA and pre-ranked GSEA tested differential expression and differential targeting between the two sample groups.Targeting was defined as the sum of incoming edges, or gene in-degree.
8.7. Summary of data and analyses
The supplemental tables summarize the single-sample analysis approaches and the datasets used across the manuscript’s analyses.
- Supplemental Table S1 summarizes single-sample analysis approaches.
- Supplemental Table S2 summarizes the manuscript’s analyses and the data used for each.
8.8. Supplemental Figures
The supplemental figures illustrate LIONESS evaluation on simulated data, single-sample edge-weight estimation, and lymphoblastoid cell-line network and expression clustering.
- Supplemental Figures: LIONESS estimates single-sample edge weights from Pearson correlation for linear relationships and mutual information for nonlinear relationships.The examples use simulated expression values for two genes across samples.
- Supplemental Figures: A schematic depicts generation of in silico expression data from known underlying gene regulatory network models.
- Supplemental Figures: Simulated-data analyses compare aggregate and LIONESS networks across network sizes, permutation levels, and all versus permuted edges.The figure reports differences in AUC values for these evaluation settings.
- Supplemental Figures: Supplemental comparisons examine expression-profile similarity alongside similarity between networks reconstructed from corresponding replicates.The relative similarity panel converts Spearman correlations to Z-scores using the mean and standard deviation across correlation values.
- Supplemental Figures: For 153 lymphoblastoid cell-line samples, technical replicates cluster together in both predicted networks and gene-expression data.This matching occurs for 121/153 network samples and 121/153 expression samples.
9. LIONESS: LINEAR INTERPOLATION TO OBTAIN NETWORK ESTIMATES FOR SINGLE SAMPLES
LIONESS derives single-sample networks from aggregate network estimates using a linear interpolation framework that can operate with different reconstruction algorithms. The section examines its mathematical behavior for Pearson correlation and mutual information, including large-sample limits and convergence constraints.
- Derivation: LIONESS derives a single-sample network by comparing aggregate networks built with and without that sample.The derivation assumes that an aggregate edge is a weighted combination of sample-specific edge values; equal weighting gives q = 1/N.
- Generality: The framework is independent of the aggregate network inference method and applies to both linear and nonlinear reconstruction algorithms.The section applies it to Pearson correlation and mutual information, respectively.
- Limitations: The LIONESS networks do not converge in the limit of many samples, and the reported ND interquartile range is typically 0.01–1.The difference between all-sample and leave-one-sample-out aggregate estimates consistently exceeds 1/N.
- Pearson correlation: For Pearson correlation, averaging the derived single-sample networks recovers the aggregate correlation network in the large-sample limit.The derivation uses all-sample and leave-one-sample-out correlation estimates and explains why the relevant finite-sample terms retain a finite limit.
- Mutual information: For mutual information, the same linear edge relationship holds in the limit of a large number of samples despite mutual information measuring nonlinear associations.The derivation represents mutual information through binned counts and analyzes the effect of removing one data point.
- Mutual information: The mutual-information derivation reduces to the original data-point formulation of mutual information when N is large.This result depends on the limiting behavior of the binned-count formulation.