Source-linked AI summary

On Identifying Significant Edges in Graphical Models of Molecular Networks

Marco Scutari, Radhakrishnan Nagarajan

arXiv:1104.0896v5stat.MLstat.ME

TL;DR

Ad-hoc edge-confidence thresholds can distort molecular network topology and complicate the identification of statistically significant associations. The paper estimates a threshold by minimizing the L1 distance between observed and asymptotic confidence CDFs, and evaluates it across synthetic and experimental data. The method maintains specificity and accuracy near 1 while improving sensitivity relative to common ad-hoc thresholds.

  • Problem

    Unknown network distributions make confidence thresholds difficult to evaluate, leading researchers to use ad-hoc cutoffs that can affect significant-edge identification and network topology.

  • Method

    The method estimates the confidence threshold minimizing the L1 distance between the observed edge-confidence CDF and the asymptotic CDF.

  • Results

    Across networks and sample sizes, specificity and accuracy are close to 1, while the estimated threshold systematically improves sensitivity over ad-hoc thresholds with comparable specificity and accuracy.

  • Takeaways & Limitations

    Statistically motivated thresholding can reduce reliance on large ad-hoc cutoffs while retaining comparable specificity and accuracy and recovering more significant associations.

  • Takeaways & Limitations

    The estimated threshold depends jointly on all edge-confidence values, so individual edge decisions are not independent.

Abstract

from arXiv · show

Objective: Modelling the associations from high-throughput experimental molecular data has provided unprecedented insights into biological pathways and signalling mechanisms. Graphical models and networks have especially proven to be useful abstractions in this regard. Ad-hoc thresholds are often used in conjunction with structure learning algorithms to determine significant associations. The present study overcomes this limitation by proposing a statistically-motivated approach for identifying significant associations in a network. Methods and Materials: A new method that identifies significant associations in graphical models by estimating the threshold minimising the $L_{\mathrm{1}}$ norm between the cumulative distribution function (CDF) of the observed edge confidences and those of its asymptotic counterpart is proposed. The effectiveness of the proposed method is demonstrated on popular synthetic data sets as well as publicly available experimental molecular data corresponding to gene and protein expression profiles. Results: The improved performance of the proposed approach is demonstrated across the synthetic data sets using sensitivity, specificity and accuracy as performance metrics. The results are also demonstrated across varying sample sizes and three different structure learning algorithms with widely varying assumptions. In all cases, the proposed approach has specificity and accuracy close to 1, while sensitivity increases linearly in the logarithm of the sample size. The estimated threshold systematically outperforms common ad-hoc ones in terms of sensitivity while maintaining comparable levels of specificity and accuracy. Networks from experimental data sets are reconstructed accurately with respect to the results from the original papers.

1. Introduction and background

Graphical models represent dependencies among molecular variables, but assessing learned network structures is difficult when the true structure is unknown. The paper motivates statistically grounded edge-confidence thresholding to replace ad-hoc significance cutoffs.

  • Graphical models: Graphical models encode dependence relationships among random variables using nodes and edges.Structure learning seeks a graph representing conditional independencies in the data.
  • Model assessment: Structure-learning algorithms estimate network structures, but their statistical robustness is difficult to assess for real-world data with unknown true distributions.Reference-data comparisons are limited because real experimental networks lack known ground-truth structures.
  • Existing assessment approach: Bootstrap resampling and model averaging estimate edge probabilities by repeatedly learning graphical-model structures from resampled data.The procedure can use parametric or nonparametric bootstrap and is compatible with different structure-learning algorithms.
  • Edge confidence: Edge intensities estimate confidence that individual edges belong to the true network, but their threshold depends on the data and learning algorithm.The underlying distribution of learned networks is unknown, making the confidence threshold difficult to evaluate directly.
  • Proposed approach: The proposed estimator selects a confidence threshold by minimizing the L1 distance between observed and asymptotic edge-confidence CDFs.The approach is subsequently evaluated on experimental data sets from Nagarajan et al. and Sachs et al.

2. Selecting significant edges

The method estimates a statistically motivated threshold by approximating an ideal binary confidence configuration and minimizing its L1 CDF distance from observed edge confidences. The resulting threshold separates significant from non-significant edges while accounting for the full ordered confidence distribution.

  • Threshold construction: The ordered edge confidences are compared with an ideal configuration containing zeros for non-significant edges and ones for significant edges.This ideal configuration corresponds to the limit in which bootstrap networks share exactly the same structure, which may occur with consistent structure learning and large samples.
  • Distance minimization: The proposed estimator minimizes the L1 distance between the empirical CDF of observed confidences and the CDF of the ideal configuration.The empirical CDF is piecewise constant, enabling the distance to be simplified and computed in linear time; minimization can then be performed using linear programming.
  • Threshold construction: The parameter t represents the fraction of ideal confidence values equal to zero and supplies the separation threshold for significant and non-significant edges.Estimating t from data provides a statistically motivated threshold through the quantile function of the ideal configuration.
  • Distance minimization: The L1 norm is used because it places less weight on large deviations than small ones, making the procedure robust across varied confidence configurations.The edge-identification problem can therefore be viewed as least absolute deviations estimation or L1 approximation.
  • Illustrative example: Although edges are classified individually, their classifications are not independent because the estimated threshold depends on the complete ordered confidence vector.The threshold is a function of the whole confidence distribution rather than a separately estimated quantity for each edge.
  • Illustrative example: In the four-node example, the minimum occurs at t = 0.4999816, retaining only edges with confidence at least 0.7689.The selected edges are (A, D), (B, D) and (C, D).

3. Simulation results

Simulations across three Bayesian networks, three structure-learning algorithms, and varying normalized sample sizes show that the estimated threshold preserves high specificity and accuracy while improving sensitivity over ad-hoc thresholds.

  • Experimental setup: Three structure-learning algorithms— IAMB, HC, and MMHC—were evaluated on the ALARM, HAILFINDER, and INSURANCE networks using sensitivity, specificity, and accuracy.The simulations estimated confidence values from bootstrap samples, applied the estimated threshold, and compared the resulting networks with known true structures.
  • Sensitivity: Sensitivity increased with normalized sample size, with HC recovering about 50%–75% of true edges when n/p was at least 0.2.IAMB and MMHC recovered about 45%–50% of HAILFINDER and about 19%–40% of ALARM and INSURANCE in the corresponding comparisons.
  • Sensitivity: Sensitivity grew rapidly for n/p ≤ 1 and then converged asymptotically toward 1, with slower convergence for the denser INSURANCE network.INSURANCE has 1.92 edges per node, compared with 1.24 for ALARM and 1.17 for HAILFINDER.
  • Specificity and accuracy: Specificity and accuracy remained close to 1 across networks and sample sizes, even at very low n/p ratios.The high values reflect the relatively small numbers of true edges among all possible edges; INSURANCE showed lower values than the sparser networks.
  • Threshold behavior: The estimated threshold showed no apparent trend with n/p and could remain below 1 at high n/p values, while performance estimates were stable.The threshold varied substantially across samples, whereas the confidence intervals for sensitivity, specificity, and accuracy were small.
  • Comparison with ad-hoc thresholds: Compared with t = 0.70, 0.80, 0.90, and 0.95, the estimated threshold systematically improved sensitivity, especially at low n/p, while maintaining comparable specificity and accuracy.The difference in sensitivity diminished as n/p increased.

4. Applications to molecular expression profiles

The proposed thresholding approach was applied to gene and protein expression data, reproducing previously reported networks while selecting significant edges through an estimated confidence threshold.

  • 4.1. Differentiation potential of aged myogenic progenitors: The study reanalysed gene-expression data from Nagarajan et al. using GAPDH-normalised profiles, 500 bootstrap samples, and the IAMB algorithm.Significant edges were evaluated without regard to direction and then oriented according to the highest-frequency direction.
  • 4.1. Differentiation potential of aged myogenic progenitors: All edges identified in the earlier Nagarajan study across its learning and normalisation settings were also identified by the proposed approach.The earlier IAMB analysis using GAPDH alone detected considerably more additional edges.
  • 4.1. Differentiation potential of aged myogenic progenitors: The proposed approach may reduce false positives and spurious gene relationships while providing edge directionality absent from Nagarajan et al.'s undirected network.This directionality was obtained using the algorithm from Imoto et al.
  • 4.2. Cellular signalling networks: For Sachs et al.'s flow-cytometry data, the analysis used 854 non-intervention observations across 11 variables and compared the estimated threshold with 0.85.The combined perturbed and non-perturbed observations from the original study could not be analysed with this approach.
  • 4.2. Cellular signalling networks: Thresholds between 0.4 and 0.9 produced the same Sachs network, while the estimated threshold was p̂(i) >= 0.93.The confidence levels of significant and non-significant edges were widely separated in the empirical CDF.

5. Conclusions

The paper proposes a statistically motivated alternative to ad-hoc edge thresholds by matching observed confidence distributions to an asymptotic ideal configuration. It demonstrates the approach on synthetic networks and molecular expression data.

  • 5. Conclusions: Ad-hoc thresholds can reduce false positives but accentuate false negatives, affecting the topology of biological network abstractions.The proposed estimator instead minimises the L1 distance between observed and asymptotic confidence CDFs.
Loading 1104.0896v5…