Source-linked AI summary
Bayesian inference for a covariance matrix
Ignacio Alvarez, Jarad Niemi, Matt Simpson
TL;DR
The paper asks how covariance-matrix prior choices affect posterior inference, given the inverse Wishart prior’s restrictive variance–correlation behavior and incomplete analytical understanding of alternatives. It compares several priors through simulation and bird-count data, finding especially strong inverse-Wishart bias for small variances and greater flexibility from separation-based modeling. These results support choosing priors with attention to bias patterns and computational framework.
Problem
Inverse Wishart priors can restrict variance information and impose prior dependence between variances and correlations, while analytical understanding of alternatives remains incomplete.
Method
The study compares inverse Wishart, scaled inverse Wishart, hierarchical inverse Wishart, and separation-strategy priors using simulations and Great Lakes bird-count data.
Results
Inverse Wishart shows extreme upward bias for small variances and corresponding correlation shrinkage toward zero, whereas other priors show only slight correlation bias that disappears with larger samples.
Takeaways & Limitations
The separation-based BMMmu prior provides the most flexibility because variances and correlations are independent by construction, while computational costs depend on the inference framework.
Takeaways & Limitations
The BMMmu prior is computationally more complex, although this restriction is less detrimental with HMC-based software than with Gibbs sampling.
Abstract
from arXiv · showhide
Covariance matrix estimation arises in multivariate problems including multivariate normal sampling models and regression models where random effects are jointly modeled, e.g. random-intercept, random-slope models. A Bayesian analysis of these problems requires a prior on the covariance matrix. Here we assess, through a simulation study and a real data set, the impact this prior choice has on posterior inference of the covariance matrix. Inverse Wishart distribution is the natural choice for a covariance matrix prior because its conjugacy on normal model and simplicity, is usually available in Bayesian statistical software. However inverse Wishart distribution presents some undesirable properties from a modeling point of view. It can be too restrictive because assume the same amount of prior information about every variance parameters and, more important, it shows a prior relationship between the variances and correlations. Some alternatives distributions has been proposed. The scaled inverse Wishart distribution, which give more flexibility on the variance priors conserving the conjugacy property but does not eliminate the prior relationship between variances and correlations. Secondly, it is possible to fit separate priors for individual correlations and standard deviations. This strategy eliminates any prior relationship within the covariance matrix parameters, but it is not conjugate and therefore computationally slow.
1 Introduction
The study examines how covariance-matrix prior choices affect posterior inference in multivariate and jointly modeled random-effects settings. It motivates this comparison by the inverse Wishart prior’s modeling limitations and evaluates alternatives through simulation and bird-count data.
- Bayesian covariance-matrix estimation requires a prior and arises in multivariate normal and jointly modeled random-effects problems.
- The inverse Wishart is widely used because it is conjugate and commonly implemented in Bayesian software.
- Inverse Wishart priors restrict variance uncertainty through one degree-of-freedom parameter and impose prior dependence between variances and correlations.
- Alternative proposals include scaled inverse Wishart, hierarchical inverse Wishart, and separation-strategy priors, but their properties are less fully understood analytically.
- The study uses simulation and a Great Lakes bird-count data set to assess how prior choices affect posterior covariance-matrix inference.
2 Statistical Models
The paper formulates covariance estimation under a multivariate normal model and compares several prior classes, emphasizing trade-offs between conjugacy, variance flexibility, parameter dependence, and computation.
- Model: The multivariate normal model assumes independent observations Yi ∈ R^d with mean µ and positive-definite covariance matrix Σ.
- Prior classes: The study considers inverse Wishart, scaled inverse Wishart, hierarchical inverse Wishart, and separation-strategy priors for Σ.
- Inverse Wishart prior: Inverse Wishart conjugacy yields inverse Wishart full conditional and marginal posterior distributions, facilitating MCMC implementation.
- Inverse Wishart prior: The inverse Wishart prior restricts variance-specific prior information, can bias posterior variances near zero, and links variances with correlations.
- Scaled inverse Wishart: Scaled inverse Wishart adds flexibility for standard-deviation priors while retaining inverse-Wishart structure, but preserves the inverse-Wishart correlation distribution.
- Hierarchical and separation priors: The hierarchical inverse Wishart implies half-t standard-deviation priors, while the separation strategy models standard deviations and correlations independently.
- Separation strategy: The separation strategy removes variance-correlation dependence by construction but is computationally disadvantaged, especially relative to HMC-based software such as Stan.
3 Simulation study results
The simulation study compares four covariance-matrix priors and finds that inverse Wishart dependence between variances and correlations can substantially distort posterior estimates when variances are small.
- Study design: The study simulates prior draws and multivariate-normal data to assess how covariance-matrix prior choice affects posterior inference.It examines two- and ten-dimensional settings, with factorial combinations of sample sizes, standard deviations, and correlations.
- Prior behavior: The inverse Wishart prior induces positive dependence among standard deviations and associates larger correlations with larger variances.The scaled and hierarchical inverse Wishart priors weaken this dependence, whereas BMMmu draws are independent by construction.
- Correlation inference: When the standard deviation is small, the inverse Wishart prior heavily shrinks posterior correlations toward zero, even when the true correlation is close to 1.The bias decreases with sample size but remains remarkably large with 250 observations when the standard deviation is around 0.01.
- Correlation inference: The other priors recover true correlations with only slight bias toward zero for small sample sizes.The reported ten-dimensional results are similar to the displayed bivariate results.
- Standard-deviation inference: The inverse Wishart prior overestimates small standard deviations, and the bias remains quite large at a true standard deviation of 0.01 with 250 observations.The other priors show no difficulty accurately estimating the standard deviation.
- Covariance inference: The inverse Wishart prior shows no bias for covariance estimates, while its small-variance behavior produces upward variance bias and correlation bias toward zero.The paper attributes covariance robustness to heavy tails in the marginal covariance prior and explains the correlation bias through covariance division by two standard deviations.
4 Bird counts in the Superior National Forest
The bird-count analysis compares covariance priors across total versus average counts and pairwise versus simultaneous species models. Correlation estimates are strongly affected by the response scale under the inverse Wishart prior, while total counts avoid the observed shrinkage.
- Data and models: The data comprise yearly counts for the 10 most abundant species in Superior National Forest from 1995 to 2013.Total count is the annual number of birds; mean count divides total count by approximately 500 yearly surveys.
- Data and models: The analyses compare four covariance priors plus an inverse Wishart prior with scaled data, using pairwise and simultaneous multivariate-normal models.The response is alternatively represented by total count or average count.
- Correlation results: Average-count correlations are shrunk toward zero under every model, with inverse-Wishart shrinkage described as extreme.The results match the pattern found in the simulation study.
- Correlation results: For total counts, the estimates show no shrinkage toward small correlations, and the inverse Wishart prior may best recover the Pearson correlation coefficient.This contrasts with the average-count analysis under the same general modeling framework.
- Correlation results: The SIW, HIWht, and BMMmu priors show similar correlation behavior regardless of whether total or average counts are modeled.Using any of these priors leads to basically the same conclusions about estimated correlations.
5 Summary
The study finds that covariance-prior choice affects posterior inference through prior dependence between variances and correlations. It recommends balancing modeling flexibility against computational resources and sampler capabilities.
- Prior properties: The inverse Wishart prior induces strong dependence between correlations and variances, while BMMmu makes them independent by construction.Scaled and hierarchical inverse Wishart priors retain similar dependence but appear more flexible than the inverse Wishart prior.
- Posterior behavior: The inverse Wishart posterior is biased toward larger values for small empirical variances and shrinks corresponding correlations toward zero.The other three priors show only slight bias toward zero correlations, which disappears with larger sample sizes.
- Modeling implications: BMMmu is appealing because correlations and variances can be modeled independently, allowing the data to define their relationship.A separation strategy can instead use distinct priors such as half-t standard deviations and uniform correlations.
- Computational implications: BMMmu is expected to be computationally most complex because the other priors retain inverse-Wishart conjugacy.SIW is reported as more computationally efficient than HIW or BMM even when using Stan.
- Practical guidance: With HMC samplers such as Stan, the separation strategy offers modeling flexibility and good inference properties; Gibbs samplers favor priors retaining conjugacy.When inverse Wishart is the only available option, the data can be pre-scaled for correlation estimation.