Source-linked AI summary
Entropy inference and the James-Stein estimator, with application to nonlinear gene association networks
Jean Hausser, Korbinian Strimmer
TL;DR
Small-sample, high-dimensional data make entropy estimation difficult, especially when maximum likelihood severely underestimates entropy. The paper develops a James-Stein shrinkage estimator and reports strong statistical performance with substantially lower computational cost than NSB.
Problem
Estimating entropy from small-sample, high-dimensional data is difficult, while maximum likelihood severely underestimates true entropy.
Method
The paper estimates entropy through James-Stein-type shrinkage applied at the level of cell frequencies.
Results
The NSB, Chao-Shen, and shrinkage estimators have small MSEs across four scenarios, while shrinkage is 1000 times faster than NSB.
Takeaways & Limitations
The shrinkage estimator provides an efficient procedure for entropy estimation and supports mutual-information-based gene association analysis.
Takeaways & Limitations
The estimator assumes p is fixed and known, although it appears robust when p is larger than necessary.
Abstract
from arXiv · showhide
We present a procedure for effective estimation of entropy and mutual information from small-sample data, and apply it to the problem of inferring high-dimensional gene association networks. Specifically, we develop a James-Stein-type shrinkage estimator, resulting in a procedure that is highly efficient statistically as well as computationally. Despite its simplicity, we show that it outperforms eight other entropy estimation procedures across a diverse range of sampling scenarios and data-generating models, even in cases of severe undersampling. We illustrate the approach by analyzing E. coli gene expression data and computing an entropy-based gene-association network from gene expression data. A computer program is available that implements the proposed shrinkage estimator.
5 October 2008; last revised 19 June 2009
The paper’s keywords identify entropy, shrinkage estimation, mutual information, and gene association networks as its central topics. The work received partial support from an Emmy Noether grant.
- The paper focuses on entropy, James-Stein shrinkage estimation, mutual information, and gene association networks.
- The work was partially supported by an Emmy Noether grant from the Deutsche Forschungsgemeinschaft.
1 Introduction
Entropy estimation is difficult in small-sample, high-dimensional settings because maximum likelihood severely underestimates entropy. The paper introduces an analytic James-Stein shrinkage estimator that is accurate, computationally inexpensive, and produces compatible cell-frequency estimates.
- Entropy estimation is challenging when n ≪ p, and maximum likelihood severely underestimates true entropy.When n ≫ p, ML estimation is reliable and optimal.
- Recent small-sample entropy benchmarks include the NSB and Chao-Shen estimators.
- The paper introduces a novel small-sample entropy estimator based on James-Stein shrinkage.
- The proposed procedure is fully analytic and computationally inexpensive.
- The estimator simultaneously estimates entropy and cell frequencies suitable for the Shannon entropy formula.
2 Conventional Methods for Estimating Entropy
Conventional entropy estimators include frequency plug-in, Bayesian, NSB, and Chao-Shen approaches. Their performance depends on the sampling regime and, for Bayesian methods, the choice of prior parameters.
- Entropy estimators either estimate cell frequencies for a plug-in calculation or infer entropy directly.
- The multinomial model connects observed counts y_k with cell probabilities θ_k.
- The ML estimator uses observed frequencies, but its plug-in entropy estimate is biased.
- The Miller-Madow estimator applies a first-order bias correction based on the number of observed cells.Here m>0 denotes the number of cells with y_k > 0.
- Bayesian estimators use Dirichlet pseudo-counts, with A interpreted as the prior sample size.
- There is no general agreement on the best noninformative Dirichlet parameters, and inappropriate choices can underperform ML.
- NSB uses an infinite Dirichlet mixture prior, whereas Chao-Shen combines Horvitz-Thompson estimation with Good-Turing correction.
- The Chao-Shen estimator has remarkably good statistical properties.
3 A James-Stein Shrinkage Estimator
The proposed method applies James-Stein shrinkage to cell-frequency estimation by combining a high-dimensional model with a lower-dimensional target. It is analytically connected to empirical Bayes estimation and can be extended to multiple targets.
- The contribution applies James-Stein-type shrinkage directly to cell frequencies for entropy estimation.
- James-Stein shrinkage regularizes high-dimensional inference and is suited to small samples, including the original n = 1 setting.
- The estimator averages a low-bias, high-variance model with a higher-bias, lower-variance model.
- The shrinkage intensity λ ranges from 0 for no shrinkage to 1 for full shrinkage, with the uniform distribution as a convenient target.
- The shrinkage estimator is equivalent to an empirical Bayes estimator with data-driven flattening constants.
- The method assumes fixed, known p but appears robust when p is specified larger than necessary.
- The approach can be modified to use multiple targets with different shrinkage intensities.
4 Comparative Evaluation of Statistical Properties
A simulation benchmark compared nine entropy estimators across four probability-generating scenarios and sample sizes from 10 to 10000 in dimension p = 1000. The NSB, Chao-Shen, and shrinkage estimators were most statistically efficient, while performance depended strongly on estimator and sampling scenario.
- Simulation design: The four scenarios covered sparse heterogeneous probabilities, homogeneous probabilities, homogeneous probabilities with structural zeros, and a Zipf-type power law.The first three scenarios used Dirichlet distributions with a = 0.0007, a = 1, and a = 1 with half the cells structurally zero.
- Results: Maximum likelihood and Miller-Madow performed worst except in scenario 1, and Miller-Madow’s bias correction was not particularly effective.The passage states that both estimators remain inappropriate even for moderately large sample sizes.
- Results: Bayesian estimators with pseudocounts 1/2 and 1 performed well in scenarios 2 and 3 but were less efficient in scenario 4 and failed completely in scenario 1.Thus, Bayesian performance depended on both prior choice and sampling scenario.
- Results: The NSB, Chao-Shen, and shrinkage estimators achieved small MSEs across all four scenarios regardless of sample size.The NSB and Chao-Shen estimators were nearly unbiased in scenario 3.
- Results: The three top-performing estimators were NSB, Chao-Shen, and shrinkage, while shrinkage was much faster than NSB in the simulations.The shrinkage estimator was reported to be faster than NSB by a factor of 1000.
5 Application to Statistical Learning of Nonlinear Gene Association Networks
The paper applies shrinkage-based entropy and mutual-information estimation to nonlinear gene association networks under strongly undersampled expression data. In the E. coli example, ARACNE reduced 5151 gene-pair estimates to 112 nonzero mutual-information links and produced a network with prominent hubs.
- Motivation: Mutual information provides a natural association measure for genes because it captures both linear and nonlinear relationships.It is non-negative, symmetric, and equals zero only when the variables are independent.
- Estimation workflow: Shrinkage regularization is needed because gene-expression mutual information is estimated from small samples, often with n much smaller than K2.The paper notes that simple maximum-likelihood approaches are not valid in this regime.
- Estimation workflow: The workflow discretizes gene measurements, estimates K2 contingency-table cell frequencies with shrinkage, and computes marginal, joint, and mutual-information entropies.The method requires discrete data and uses Freedman-Diaconis discretization when measurements are not already discrete.
- E. coli application: The E. coli analysis used 102 genes, 9 time points, 16 discretized expression levels, and 5151 gene pairs, with each pair represented by a 16 × 16 table.This gives p = 256 in a strongly undersampled setting.
- E. coli application: ARACNE retained 112 gene pairs with nonzero mutual information after selecting against indirect links using the information processing inequality.For each gene triplet, the pair with the smallest mutual information was discarded.
- E. coli application: The inferred network contained prominent hubs centered on hupB, sucA, and nuoL.hupB is a DNA-binding transcriptional regulator, while nuoL and sucA are key components of E. coli metabolism.
6 Discussion
The paper proposes a simple James-Stein-type shrinkage estimator for small-sample entropy and mutual information, showing statistical and computational efficiency and applying it to E. coli gene-association inference.
- The proposed James-Stein-type shrinkage estimator infers entropy and mutual information from small samples while remaining statistically and computationally efficient.The authors emphasize this efficiency despite the method’s simplicity.
- The estimator provides multinomial frequencies alongside entropy estimates, supporting mutual-information analyses of nonlinear pairwise dependencies.This frequency output can be plugged into the Shannon entropy formula.
- Unlike NSB, the proposed estimator is fully analytic, and unlike both NSB and Chao-Shen, it is described as simpler and more versatile.The paper identifies analytic computation and frequency estimates as its two versatility advantages.
- The method was applied to expression data from E. coli to estimate an entropy-based gene dependency network.The application targets genomics and systems biology.
- The authors position the estimator as a potential contribution to machine learning and statistics tools for high-dimensional data analysis.
Appendix A: Recipe For Constructing James-Stein-type Shrinkage Estimators
The appendix presents James-Stein shrinkage as regularized high-dimensional inference: it combines a variable high-dimensional estimate with a biased, lower-variance target using data-adaptive weighting.
- Recipe For Constructing James-Stein-type Shrinkage Estimators: James-Stein shrinkage is presented as a simple analytic device suited to small-sample settings, including the original n = 1 case.
- Recipe For Constructing James-Stein-type Shrinkage Estimators: For p ≥3, the James-Stein estimator dominates the maximum likelihood estimator in mean squared error.
- Recipe For Constructing James-Stein-type Shrinkage Estimators: The general recipe regularizes a high-dimensional estimator by linear combination and adaptively estimates the shrinkage parameter from data using quadratic risk.
- Recipe For Constructing James-Stein-type Shrinkage Estimators: James-Stein shrinkage combines a high-dimensional estimate with a lower-dimensional target that is less variable but more biased.The high-dimensional estimate has low bias but potentially large variance in small samples.
- Recipe For Constructing James-Stein-type Shrinkage Estimators: The James-Stein estimate is a data-driven weighted average of the high-dimensional estimate and the target.
- Recipe For Constructing James-Stein-type Shrinkage Estimators: The shrinkage estimate improves mean squared error relative to both component estimators.
- Recipe For Constructing James-Stein-type Shrinkage Estimators: The optimal shrinkage intensity can be calculated analytically without knowing the true parameter value.
- Recipe For Constructing James-Stein-type Shrinkage Estimators: The method has an empirical Bayes interpretation but requires only the first two moments of the component-estimate distributions.
Appendix B: Computer Implementation
The proposed entropy and mutual-information shrinkage estimators were implemented in R and distributed through the CRAN package “entropy.”
- The proposed shrinkage estimators and the other investigated entropy estimators were implemented in R.
- The corresponding R package “entropy” was deposited in CRAN under the GNU General Public License.