Source-linked AI summary
Random-set methods identify distinct aspects of the enrichment signal in gene-set analysis
Michael A. Newton, Fernando A. Quintana, Johan A. den Boon, Srikumar Sengupta, Paul Ahlquist
TL;DR
Gene-set enrichment methods must detect over-representation of altered-expression genes while handling both selected lists and quantitative gene-level evidence. This paper develops conditional random-set scores, compares averaging with selection, and finds that each is superior in different enrichment regimes. The approach also supports calibrated multi-category inference without Monte Carlo.
Problem
Existing enrichment analyses are limited by selection dependence, unequal detection power across category sizes, and dependence among statistics from shared genes.
Method
The paper conditions on observed gene-level scores, treats categories as random gene sets, and analytically standardizes category statistics using their conditional mean and variance.
Results
Each method has a domain of superiority: averaging is more powerful for many modestly altered genes, whereas selection is preferred for slight enrichment with large effects.
Takeaways & Limitations
Random-set scoring measures distinct enrichment signals, extends Fisher-style analysis to quantitative scores, and enables multi-category inference without Monte Carlo.
Takeaways & Limitations
Random-set Z scores compare a category with hypothetical random sets conditional on differential-expression results, rather than with a hypothetical rerun of the whole expression study.
Abstract
from arXiv · showhide
A prespecified set of genes may be enriched, to varying degrees, for genes that have altered expression levels relative to two or more states of a cell. Knowing the enrichment of gene sets defined by functional categories, such as gene ontology (GO) annotations, is valuable for analyzing the biological signals in microarray expression data. A common approach to measuring enrichment is by cross-classifying genes according to membership in a functional category and membership on a selected list of significantly altered genes. A small Fisher's exact test $p$-value, for example, in this $2\times2$ table is indicative of enrichment. Other category analysis methods retain the quantitative gene-level scores and measure significance by referring a category-level statistic to a permutation distribution associated with the original differential expression problem. We describe a class of random-set scoring methods that measure distinct components of the enrichment signal. The class includes Fisher's test based on selected genes and also tests that average gene-level evidence across the category. Averaging and selection methods are compared empirically using Affymetrix data on expression in nasopharyngeal cancer tissue, and theoretically using a location model of differential expression. We find that each method has a domain of superiority in the state space of enrichment problems, and that both methods have benefits in practice. Our analysis also addresses two problems related to multiple-category inference, namely, that equally enriched categories are not detected with equal probability if they are of different sizes, and also that there is dependence among category statistics owing to shared genes. Random-set enrichment calculations do not require Monte Carlo for implementation. They are made available in the R package allez.
1. Introduction.
Gene-set enrichment analysis asks whether altered-expression genes are over-represented in functional categories. The paper contrasts selection-based and quantitative-score approaches, motivating random-set methods that address limitations of existing enrichment evaluations.
- Enrichment analysis asks whether genes with altered expression are over-represented in named functional categories.
- Selection methods test overlap between a category and a short list of significantly altered genes, but depend on selection stringency.
- Quantitative approaches retain gene-level scores and assess category statistics using permutation distributions tied to the differential-expression problem.
- Selected-gene analyses give equal weight to genes at both ends of the selected list.
- SAFE/GSEA adds computational burden, can be ineffective with few microarrays, and tests complete absence of differential expression rather than absence of enrichment.
2. Random-set enrichment scoring.
Random-set enrichment treats the category as a random gene set while conditioning on observed gene-level scores. This yields analytically standardized scores that extend Fisher-style testing to quantitative evidence and multiple score types.
- The unstandardized enrichment score X̄ averages gene-level scores over the genes in category C.
- Random-set scoring draws category C uniformly from m distinct genes among G genes, equivalent to shuffling scores across gene labels.
- The method analytically obtains the random-set distribution’s first two moments and uses an approximately Gaussian distribution instead of Monte Carlo.
- The standardized score Z has mean zero and unit variance under no enrichment, with large values favoring enrichment and approximate standard-normal behavior for moderate-to-large categories.
- Gene-level scores can include log fold changes, t-statistics, ranks, posterior probabilities, or transformed Spearman correlations.
- With binary scores, the random-set Z score approximately corresponds to Fisher’s or Pearson’s test; otherwise it generalizes those scores.
3. An analysis of host/virus associations in cancer.
The host-virus analysis compares selection- and averaging-based enrichment scores across GO categories using probe-set correlations with EBNA1. The methods identify different enrichment signals, with category-specific results affected by probe-set redundancy and calibration choices.
- Host-virus association: 65% of host probe sets were negatively correlated with EBNA1, indicating significant global negative association.The most extreme negative correlation was r = −0.75, with p-value = 0.04.
- Category scoring: The analysis scored 2761 GO categories with at least 10 annotated probe sets using Zave for averaged correlations and Zsel for selected genes.The annotated-probe-set universe contained G = 27,152 probe sets, and Zsel represented the normal-score version of Fisher’s exact test.
- Category scoring: Zave and Zsel capture different enrichment components: selection detects over-representation among the most altered genes, whereas averaging detects broadly shifted gene-level evidence.Categories can have high Zsel with unremarkable average correlation, or high Zave without over-representation on the selected list.
- Category-specific findings: The immune response category had strong Zave evidence but weaker Zsel evidence because many negatively correlated genes were not in the 5% FDR selected list.GO:0006955 contains m = 1494 probe sets, while its specific immune-response subsets generally had Zave > 0 and often Zsel < 0.
- Category-specific findings: GO:0019883 antigen presentation, endogenous antigen contained m = 48 probe sets, including x = 8 selected probes, and had zsel = 10.…; follow-up experiments confirmed the suggested negative correlations.The category scored highly by both methods, although the supplied passage does not resolve whether EBNA1 exploits disabled antigen presentation or changes host expression.
- Probe-set adjustment: Because genes map to multiple probe sets, the analysis applies an adjustment using gene and probe-set counts; for GO:0019883, zadjust;ave = 3.84 and zadjust;sel = 5.3.The adjustment approximates gene-level reduction and is described as generally conservative.
- Calibration comparison: Random-set and SAFE calibrations gave p < 10−10 (z = −7.1) versus p = 0.02 for GO:0019883, despite using the same average-rank statistic.The differing category rankings suggest that random-set calibration identifies a distinct enrichment signal from SAFE.
4. Averaging or selection? A theoretical comparison.
Theoretical comparison shows that selection and averaging have distinct power advantages depending on enrichment strength, gene-level effect size, and category composition; neither method is uniformly superior.
- Theoretical implications: The paper concludes that neither scoring method is always preferred because each has its own superiority region in the enrichment state space.This establishes sub-optimality of relying exclusively on either selection or averaging.
- Model and testing: The location model treats category enrichment as H0: πC = π versus H1: πC > π, with enrichment defined as πC −π.Category statistics X̄sel and X̄ave are compared against alternatives under a model where gene scores depend on latent differential-expression indicators.
- Model and testing: Averaging tests the mean quantitative score using a Normal(δπC,1/m) sampling distribution, so power increases with effect, enrichment, and category size.The illustrated case uses m = 20, π = 0.2, and α = 0.05.
- Power comparison: Both methods gain power as enrichment or effect increases, but their power surfaces create distinct domains of superiority.For m = 20 and π = 0.2, Figure 4 compares selection and averaging across enrichment πC −π and per-gene effect δ.
- Power comparison: Selection is better when enrichment is small but effects are large, whereas averaging is better when effects are small but enrichment is large.Averaging can combine substantial noise with signal when enrichment is weak, while selection benefits from a small number of highly altered genes.
- Theoretical implications: For sufficiently large categories, selection is more powerful under suitable conditions, while for any πC > π, averaging dominates when δ is sufficiently small.The theoretical claims require 0 < κ < 1; the selection-superiority interval is nonempty when κ is sufficiently small, such as below 0.133.
5. Simultaneous inference with multiple categories.
Multiple-category inference must address category-size effects and dependence from overlapping genes. The paper uses size-adjusted rankings and joint Gaussian simulation to illustrate simultaneous testing.
- Size imbalance: Because Zave has mean √m(πC −π)δ, larger categories can rank highly partly because of size rather than enrichment.The same size effect also occurs for Zsel, especially when the selected set is relatively large.
- Size imbalance: Ranking categories by Z/√m adjusts for the category-size effect and ranks them by estimated enrichment πC −π.This ranking remains useful for prioritizing categories across GO networks.
- Size imbalance: The Z/√m ranking is not itself error-controlled and does not account for its nonconstant variance.It therefore requires calibration for formal inference.
- Dependence and simultaneous testing: Joint category statistics can be modeled through intersections among categories, with dependence increasing as category overlap increases.The directed graphical structure of GO provides information about proper-subset relationships.
- Application: The paper illustrates a possible simultaneous-testing approach but leaves a full analysis of the multiple-category problem beyond its scope.The approach uses a multivariate normal approximation rather than fully developing all calculations.
- Dependence and simultaneous testing: Under the global null, the vector of category statistics is approximately Gaussian with zero means, unit variances, and covariances determined by category relationships.Cholesky-based simulation respects overlap-induced dependence and supports multiple-testing procedures such as maxT.
- Application: In the NPC example, 11 categories exceeded the maxT 95th-percentile threshold after categories were ranked using T = Z/√m.Figure 5 reports these categories as having enrichment scores exceeding 1.10.
6. Discussion.
The paper develops conditional random-set methods for detecting distinct enrichment signals and compares selection with averaging across enrichment problems. It also extends the framework to multiple categories while addressing category-size bias and dependence from shared genes.
- Random-set calibration: Random-set calibration extends Fisher’s selected-gene framework to quantitative gene-level scores using an easily computed approximation and standardized enrichment statistics.Conditioning on differential-expression results avoids revisiting raw data and supports diverse differential-expression outputs.
- Dependence: Among-gene dependence is conditioned away in random-set scoring, but the resulting Z scores compare a category statistic with a hypothetical random set rather than a rerun expression study.Dependence can itself be the biologically relevant enrichment signal, so significant findings still require care against spurious explanations.
- Selection versus averaging: Selection and averaging target different enrichment patterns: averaging is stronger for many modestly altered genes, whereas selection is preferred for slight enrichment with large effects.The comparison is made under a location-shift model in terms of power to detect enrichment.
- Multiple-category inference: For multiple-category inference, categories can be ranked by Z score divided by the square root of category size, while category-score dependence is mediated by category overlap.The joint Z-score distribution is approximately multivariate normal under the complete null, enabling procedures such as single-step maxT.
- Multiple-category inference: Standard multiple-testing adjustments cannot be transferred simply to random-set Z scores because category size and dependence complicate p-value-based procedures.The paper demonstrates single-step maxT and notes that stepwise family-wise-error and false-discovery-rate refinements remain possible.
APPENDIX: SELECTION VS AVERAGING (THEOREM 1)
The appendix characterizes when selection or averaging has greater power under a location-shift model. It defines the model quantities, identifies critical variance behavior, and gives sufficient conditions for selection superiority.
- Theorem 1: Averaging is better when Δ = τ_sel − τ_ave > 0, while selection is better when Δ < 0.Δ compares the Z-score critical values for level α tests.
- Model quantities: The model defines μ0 and μ1 as selection probabilities for nondifferentially expressed and differentially expressed genes, with δ representing differential-expression effect size.π_C − π measures enrichment, m is category size, and σ^2(·) is a variance function.
- Selection threshold: The selection threshold k can be chosen to target a specified level α̃ false discovery rate because the defining function h is monotone and invertible.The threshold uniquely satisfies h(k) = κ, with κ determined by α̃ and π.
- Selection superiority: For sufficiently large category size, the appendix derives conditions under which selection is superior, and the relevant interval is nonempty for 0 < κ < 0.133.The proof considers large effects and shows the category-size-dependent term can dominate.
- Small-effect regime: For small δ, Mills’ ratio yields the threshold approximation k ≈ (−log κ)/δ, and δ^2/μ0 diverges as μ0 → 0.This completes the appendix’s proof for the corresponding small-effect regime.
- Implementation: The R package allez implements random-set enrichment calculations, especially for GO categories.The reported calculations used R version 2.1 and Bioconductor hug133plus2 version 1.10.0.
Computing notes.
The paper acknowledges financial support from the National Cancer Institute and FONDECYT, along with assistance from named collaborators and referees.
- Acknowledgments: Support came from National Cancer Institute grants CA64364, CA22443, and CA97944, and FONDECYT grants 1020712 and 1060729.The acknowledgments also credit contributors to statistical analysis, molecular-study coordination, specimen collection, fieldwork, category analysis, R computing, and refereeing.