Source-linked AI summary

On weakly informative prior distributions for the heterogeneity parameter in Bayesian random-effects meta-analysis

Christian Röver, Ralf Bender, Sofia Dias, Christopher H. Schmid, Heinz Schmidli, Sibylle Sturtz, Sebastian Weber, Tim Friede

arXiv:2007.08352v3stat.ME

TL;DR

Few-study Bayesian NNHM analyses require careful heterogeneity-prior specification because weak data can leave results sensitive to prior assumptions. The paper develops structured guidance and examples for choosing such priors, concluding that their choice affects estimation, prediction, and shrinkage, while acknowledging that sparse data may leave the posterior largely prior-driven.

  • Problem

    Few studies make heterogeneity-prior assumptions influential, while generally accepted guidance for weakly informative priors is lacking.

  • Method

    The paper develops general guidance by relating heterogeneity priors to effect scales, anticipated magnitudes, unit information, empirical evidence, and prior or predictive quantiles.

  • Results

    Choice of heterogeneity priors affects estimation of the overall mean and especially prediction and shrinkage through borrowing-of-strength.

  • Takeaways & Limitations

    The paper provides guidance for designing, justifying, and prospectively pre-specifying weakly informative heterogeneity priors, including examples for common effect measures.

  • Takeaways & Limitations

    When data provide little information about heterogeneity, posterior results may be largely determined by the prior.

Abstract

from arXiv · show

The normal-normal hierarchical model (NNHM) constitutes a simple and widely used framework for meta-analysis. In the common case of only few studies contributing to the meta-analysis, standard approaches to inference tend to perform poorly, and Bayesian meta-analysis has been suggested as a potential solution. The Bayesian approach, however, requires the sensible specification of prior distributions. While non-informative priors are commonly used for the overall mean effect, the use of weakly informative priors has been suggested for the heterogeneity parameter, in particular in the setting of (very) few studies. To date, however, a consensus on how to generally specify a weakly informative heterogeneity prior is lacking. Here we investigate the problem more closely and provide some guidance on prior specification.

INTRODUCTION

The paper examines Bayesian random-effects meta-analysis with the NNHM, focusing on how to specify heterogeneity priors when few studies make prior assumptions influential. It aims to provide general guidance and consensus examples for weakly informative priors.

  • INTRODUCTION: The NNHM models study estimates and standard errors with normal distributions and represents between-study variability through a second normal level.
  • INTRODUCTION: Bayesian computation is straightforward, but few studies increase sensitivity to prior specifications, especially for variance parameters.
  • INTRODUCTION: The investigation develops general guidance for judging and deriving weakly informative heterogeneity priors and suggests consensus examples for common effect measures.
  • INTRODUCTION: The guidance may support prior design, justification, and prospective pre-specification while reducing suspicion of post-hoc prior tweaking.
  • INTRODUCTION: The paper considers proper weakly informative priors because commonly proposed noninformative priors can cause problems with few studies or marginal likelihoods.
  • INTRODUCTION: Prior specification is discussed through epistemic, regularisation, and pragmatic perspectives, without assuming a single universally correct prior.

Implications for interval estimation

The paper links heterogeneity priors to interval estimation, prediction, shrinkage, and model behavior, while emphasizing technical constraints and scale-specific interpretation. It presents several proper distributional families but notes that prior choice can dominate when data are sparse.

  • Implications for interval estimation: Heterogeneity priors must have integrable behavior near zero and an eventually decreasing upper tail to ensure posterior integrability.
  • Implications for interval estimation: When data add little information, the posterior can closely resemble the prior, making transparent and convincing prior specification especially important.
  • Implications for interval estimation: Larger heterogeneity produces wider prediction intervals and less shrinkage, while varying 휏 from zero to infinity spans pooled and separate analyses.
  • Implications for interval estimation: Half-normal, half-Student-t, half-Cauchy, half-logistic, exponential, and Lomax distributions can provide near-uniform behavior near zero with decaying tails.
  • Implications for interval estimation: Half-Student-t and Lomax distributions offer heavy-tailed variants, while mixtures can combine informative and heavy-tailed components for robustness.
  • Implications for interval estimation: Bounded uniform priors impose a sharp upper cutoff that may be difficult to justify, although a sufficiently large bound can be reasonable for log-ORs.
  • Implications for interval estimation: The effect scale is central to prior specification, and transformed effects should often be interpreted on the back-transformed scale.

Magnitudes of other effects

The paper proposes judging plausible heterogeneity magnitudes through implied effects, transformed predictive ranges, marginal predictive distributions, within-study variability, and shrinkage information.

  • Implied effect distributions: Plausible heterogeneity values can be assessed from the implied conditional distribution of true study effects around the overall mean.The distribution has standard deviation τ, with effects lying within μ ± 1.96τ with 95% probability.
  • Implied effect distributions: For a random pair of true study effects, the difference follows N(0, 2τ^2), while the absolute difference has median 0.95τ.These pairwise differences provide an interpretable scale for judging whether candidate τ values are reasonable or extreme.
  • Transformed effect measures: On logarithmic scales, predictive ranges should also be examined after back-transformation, such as exponentiation for odds ratios, relative risks, and hazard ratios.Tables compare 95% predictive intervals and median differences across τ values, including their exponentiated interpretations.
  • Marginal prior predictives: Different heterogeneity prior families produce different marginal predictive distributions, so family choice can reflect tail behavior, near-zero behavior, simplicity, or concrete prior information.Half-normal, half-Student-t, half-Cauchy, half-logistic, exponential, Lomax, log-normal, and uniform families are among those considered.
  • Within-study variability: The unit information standard deviation σ1 provides a landmark for comparing between-study heterogeneity τ with within-study variability.Usually τ ≪ σ1, so study means may differ while subject-level distributions remain largely overlapping; σ1 can therefore constrain plausible τ values.
  • Shrinkage information: Prior maximum sample size links τ and σ1 to the amount of information a meta-analysis can contribute to shrinkage.If ideal additional information should equal at most 16 subjects, the corresponding τ is at most one quarter of σ1.

Empirical information on 휏

Empirical studies can inform heterogeneity priors, but their usefulness depends on how representative the external evidence is for the meta-analysis at hand.

  • Empirical evidence: Prior heterogeneity information has been derived empirically from large collections of published meta-analyses for standardized mean differences and log-ORs.Examples include analyses of Cochrane Database of Systematic Reviews by Rhodes et al. and Turner et al.
  • Applicability of external evidence: For external empirical information, researchers must consider its representativeness, the relevant data subset, and what to do when no suitable sample exists.When applicability is uncertain, a robustified two-component mixture prior may reflect that uncertainty.

Guiding questions

Prior specification should begin by translating plausible heterogeneity into the effect measure’s scale and then using empirical evidence and predictive implications to choose a distribution. The paper emphasizes that endpoint scaling, near-zero behavior, tail behavior, and practical plausibility all matter.

  • Guiding questions: Plausible heterogeneity magnitudes should be determined in terms of τ or θ_i ranges before choosing a parametric prior family.The choice may then reflect near-zero behavior, heavy-tailedness, or simplicity.
  • Guiding questions: Absolute-scale endpoints lack universally applicable prior scales because changing units, such as hours versus minutes, changes the appropriate prior scale.For averages, the unit information standard deviation can provide orientation when standard errors scale with sample size.
  • Guiding questions: Standardized mean differences are dimensionless, but heterogeneity near zero may be especially unlikely because the scale is explicitly designed to compare endpoints across measurement systems.Common SMD benchmarks include 0.2, 0.5, and 0.8, which can help bound P(τ≤1).
  • Guiding questions: Empirical heterogeneity estimates provide scale-specific guidance, including SMD estimates with median 0.20 and 95% quantile 0.66.A separate healthcare dataset yielded a log-Student-t distribution for SMD heterogeneity with median 0.18 and 95% quantile 2.43.

Regression slopes

Regression-slope heterogeneity depends on the scale of the regressor as well as the outcome, so prior elicitation requires a common reference increment and appropriate rescaling. Correlation examples further illustrate how transformations and bounded domains affect plausible heterogeneity.

  • Regression slopes: For regression parameters, the anticipated heterogeneity depends on the units of the regressor, so a reference increment Δx must be specified across studies.Changing from per-week to per-day coefficients requires rescaling the prior by the corresponding factor.
  • Regression slopes: Fisher’s z transformation maps correlations to the real line and gives an approximate standard error depending only on study sample size.Values within −0.5<r_i<0.5 are little affected, while more extreme correlations change more substantially.
  • Regression slopes: A uniform correlation range over [−1,1] implies transformed-scale variance 0.912, making τ near 1.0 an extreme heterogeneity magnitude.The corresponding plain-correlation variance is 1/3.
  • Regression slopes: Moderate correlation ranges imply Var(z)=0.302, so τ=0.30 may already represent large heterogeneity.This follows for both r∼Uniform(−0.5,0.5) and r∼Uniform(0.0,0.8).
  • Regression slopes: Using untransformed correlations in the NNHM is problematic because the model does not reflect their bounded parameter space.The practice nevertheless remains common.
  • Regression slopes: Empirical correlation heterogeneity had median 0.12 and 95% quantile 0.29 across 539 analyses from 25 studies.These estimates provide empirical motivation for substantially smaller scales than τ near 1.0.

EXAMPLE APPLICATIONS

The examples apply weakly informative heterogeneity priors to mean differences and standardized mean differences, using substantive effect scales and empirical evidence to justify half-normal(0.5) choices. The resulting analyses illustrate shrinkage and uncertainty behavior with few studies.

  • Mean differences: For symptom duration measured in days, heterogeneity above 1 day appears implausible because it would exceed the expected antibiotic effect.A half-normal(0.5) prior gives P(τ≤1)≈95% and a roughly ±1-day 95% prior predictive interval.
  • Mean differences: The symptom-duration example uses a half-normal(0.5) prior, yielding an estimated reduction of about one symptom day with uncertainty of roughly a factor of two.The posterior heterogeneity remained close to its prior anticipation, and the prediction interval was only slightly longer than the overall-mean interval.
  • EXAMPLE APPLICATIONS: The example data are organized around mean differences in days and standardized mean differences derived from differing depression symptom scales.The corresponding tables provide study estimates and standard errors used in the NNHM.
  • Standardized mean differences: For standardized mean differences, τ=1.0 implies a median difference of approximately 0.95 between random pairs of true study means, which appears extreme.Values around τ=0.5 or below were judged more plausible, motivating half-normal(0.5).
  • Standardized mean differences: The SMD example adopts half-normal(0.5) as a slightly conservative prior because empirical evidence suggests potential heavy-tailedness and may have limited relevance to the example.The analysis included pronounced heterogeneity and produced a wide overall-effect interval.

Log-transformed effect scales

For log odds ratios and log incidence rate ratios, the paper uses half-normal(0.5) priors because τ values near 1 imply large multiplicative differences. With sparse data, these priors constrain implausibly large heterogeneity and influence predictive uncertainty.

  • Log-transformed effect scales: For log-ORs, τ=1.0 implies that a random pair of studies differs by a factor of 2.6, making τ=0.5 or below more plausible.A half-normal(0.5) prior mostly covers τ<1.0 and includes predictive effects within a factor of 3 around the overall mean.
  • Log-transformed effect scales: In the two-study log-OR example, the weakly informative prior down-weighted unreasonably large heterogeneity values and produced an estimated log-OR of −1.81.The original confidence intervals were compatible with heterogeneity from zero to very large values, illustrating the speculative nature of sparse-data inference.
  • Log-transformed effect scales: The log-OR analysis showed a substantial reduction in adverse events, while the overall-effect interval remained wide and the prediction interval was limited by prior and empirical heterogeneity information.The posterior heterogeneity was very similar to the prior, and shrinkage was observable.
  • Log-transformed effect scales: The log-IRR example concerns recurrent cardiovascular hospitalizations or cardiovascular death and analyzes logarithmic event-rate ratios.Negative log-rate ratios indicate reduced incidence rates.
  • Log-transformed effect scales: For log-IRRs, the analysis used half-normal(0.5) despite overlapping study intervals, because heterogeneous circumstances can still generate apparently homogeneous data.The pooled estimate was log-IRR −0.49, corresponding to an IRR of 61% with a confidence interval from 33% to 116%.

Log odds

The examples illustrate weakly informative heterogeneity priors across log-odds, regression, and correlation analyses, with posterior heterogeneity often lower than anticipated. The resulting analyses retain uncertainty in predictions and mean effects.

  • Log odds: The log-odds analysis produced a posterior predictive probability centered at 0.11, with a 95% posterior predictive interval from 0.03 to 0.34.Its posterior predictive standard error was 0.70, corresponding roughly to an effective sample size of 3.22 relative to the UISD.
  • Log odds: The log-odds example used a half-normal(1.0) prior for heterogeneity, with the prior’s 95% quantile approximately at τ=2.Values larger than τ=2 would imply an almost noninformative posterior predictive distribution.
  • Regression: The regression example used a half-normal(0.125) heterogeneity prior for a 5 percentage point decrease in LVEF, and the estimated effects were very homogeneous.The overall log-HR estimate was 0.19, corresponding to a 1.21-fold increased mortality hazard for a 5 percentage point decrease.
  • Correlations: For Fisher-z-transformed correlations, a half-normal(0.2) prior mostly covered heterogeneity values below 0.4, with a prior median at τ=0.13.The resulting posterior median was τ=0.12, with the 95% credible interval extending to 0.30.
  • Posterior heterogeneity: Figure 8 compares marginal prior and posterior densities for τ across seven examples, marking 95% credible intervals and posterior medians.Dashed lines represent priors, solid lines represent posteriors, dark-grey shading marks the credible interval, and vertical lines mark posterior medians.
  • Correlations: The correlation analysis estimated a mean effect of about 0.08, while its confidence interval ranged from −0.1 to +0.3.The result was described as neutral to slightly positive correlation between conscientiousness and medication adherence.

DISCUSSION

The discussion proposes a structured approach for specifying weakly informative heterogeneity priors, especially with few studies, while identifying practical consequences and scope boundaries.

  • Prior specification: A seven-question framework narrows heterogeneity-prior choices by considering effect scale, expected effects, UISD, empirical information, and prior quantiles or properties.The authors demonstrated this approach in seven applications involving few studies, obtaining sensible prior distributions and results.
  • Motivation: Careful prior specification is especially important for the heterogeneity parameter when only few studies provide limited information.The discussion also identifies prior specification as relevant when computing marginal likelihoods or Bayes factors, which require proper priors.
  • Limitations: Prior sensitivity should be assessed because results may be robust to prior variations, while exchangeability can be compromised by publication or reporting bias that is difficult to detect with few studies.If a suitable weakly informative heterogeneity prior cannot be specified, the authors suggest a more conservative uninformative-prior approach.
  • Implications: Heterogeneity priors affect estimation, prediction, and shrinkage because inferred heterogeneity determines the amount of borrowing-of-strength.Smaller heterogeneity produces stronger pooling, whereas larger heterogeneity leaves individual estimates more loosely connected through the model.
  • Applications: Standardizing heterogeneity priors may help avoid post hoc disputes when different priors produce conflicting interpretations in regulatory and HTA settings.The discussion notes an IQWiG effort to derive an empirical heterogeneity distribution for HTA applications.
  • Scope: The arguments may transfer to related pairwise models and hierarchical models, but more complex evidence-synthesis applications require additional model components and prior specifications.The authors report that, for a fixed prior median, distributional shape had little impact compared with prior scaling in sensitivity analyses.

APPENDIX

The appendix derives approximate standard errors and unit information standard deviations for several effect measures, using simplifying assumptions about group sizes, variances, and event counts.

  • Continuous outcomes: For a difference between two group means, assuming known common standard deviation and equal group sizes yields an approximate standard error and UISD.The derivation uses empirical group averages and neglects uncertainty in variance estimation.
  • Binary outcomes: The variance of a logarithmic odds estimate depends on the observed proportion, linking its uncertainty to the proportion-based standard-error expression.The passage notes similarity to the standard error of a logarithmic odds ratio.
  • Binary outcomes: For p = 1/2, the logarithmic odds-ratio UISD is twice as large, while its variance is four times as large, as the two logits’ variances add.Each logit also has twice the variance because it is based on half as many subjects per total sample size.
  • Rate outcomes: For incidence rate ratios, assuming equal treatment- and control-group event counts gives an approximate standard error and a per-event UISD.The derivation assumes a total event count m and equal event counts in the two groups.

B PRIOR DISTRIBUTION FAMILIES

The appendix characterizes prior families for nonnegative heterogeneity, emphasizing their moments, scale parameterizations, and construction through scale mixtures.

  • Distribution families: Table B1 compares half-normal, half-Student-t, half-Cauchy, half-logistic, exponential, Lomax, log-normal, and proper uniform families.The table lists densities, medians, 95% quantiles, means, variances, and coefficients of variation when defined.
  • Scale parameterization: For one-parameter scale families, quantiles, expectations, and standard deviations are proportional to scale, while the coefficient of variation is constant.The appendix gives half-normal, half-Student-t, and fixed-parameter Lomax families as examples.
  • Moment conditions: Some moments are undefined for particular families or parameter values, including half-Student-t expectations for ν ≤ 1 and variances for ν ≤ 2.The appendix also explains that exponential rate is the inverse of scale and exp(μ) acts as the log-normal scale.
  • Scale mixtures: Scale mixtures create heavier-tailed priors as mixing-scale variation increases, with the coefficient of variation controlling the departure from the conditional distribution.The appendix identifies Lomax and Student-t distributions as mixtures of exponential and normal distributions, respectively.
  • Lomax construction: An exponential prior with uncertain scale can be represented as a Lomax distribution by matching the scale’s expectation and coefficient of variation.The construction uses an inverse-gamma scale distribution, equivalently a gamma rate distribution, and derives the Lomax parameters.
  • Lomax construction: For an exponential scale mixture with expectation μ = 0.5 and coefficient of variation cv = 0.5, the corresponding Lomax parameters are derived from those target moments.The example illustrates prior specification through interpretable scale uncertainty rather than direct selection of Lomax parameters.

C.3 The (half-) Student-푡distribution as a normal scale mixture

This appendix section connects half-Student-t priors to normal scale mixtures and shows how degrees of freedom and scale can be chosen from target expectation and coefficient of variation.

  • Scale-mixture representation: The Student-t distribution can be represented as a normal distribution whose scale follows a scaled inverse-chi distribution with ν degrees of freedom.This representation makes the normal scale-mixture connection explicit and extends directly to half-Student-t distributions.
  • Scale-mixture representation: The inverse-chi distribution is the inverse square root of a chi-square deviate and is a special case of a square-root inverted-gamma distribution.Introducing an additional scale parameter yields the scaled inverse-chi distribution.
  • Parameter selection: Tables C2 and C3 relate inverse-chi degrees of freedom to coefficients of variation, allowing one quantity to be selected from the other.The relationship can be inverted numerically for chosen coefficient-of-variation targets.
  • Worked example: For target expectation μ = 0.5 and coefficient of variation cv = 0.5, the normal scale mixture implies ν = 4.2 degrees of freedom and a half-Student-t scale of 0.40.The unscaled Student-t construction would instead imply E[σ] = 1.24, so rescaling is required to achieve the intended expectation.
  • Worked example: A Student-t setting with ν = 4 and scale 0.5 implies coefficient of variation cv = 0.52 and expectation μ = 0.63.This contrasts direct parameter selection with moment-targeted scaling.

C.4 Scale mixture examples

The paper describes Lomax and half-Student-t priors arising as scale mixtures and illustrates how heterogeneity priors generate predictive distributions for treatment effects. R code demonstrates conditional and marginal prior predictive calculations and a Bayesian meta-analysis using a half-normal prior.

  • Scale mixture constructions: Fixing the expectation while increasing a scale mixture’s coefficient of variation produces a more skewed mixing distribution with a lower median.Simple rescaling of the exponential or half-normal scale distribution proportionally rescales heterogeneity and the predictive distribution.
  • Scale mixture constructions: Lomax and half-Student-t heterogeneity priors arise as scale mixtures of exponential and half-normal distributions, respectively.The corresponding tables also describe the induced distributions for scale parameters, heterogeneity, and predictions.
  • Conditional prior predictive distribution: A fixed τ=0.35 implies that approximately 95% of relative risks fall between 0.5 and 2.0 around the overall effect.The example uses a normal conditional distribution for log-relative risks and exponentiates the simulated effects.
  • Marginal prior predictive distribution: A half-normal(0.32) prior for τ implies that approximately 95% of odds ratios fall between 0.5 and 2.0.The marginal predictive distribution is obtained by drawing τ from the half-normal prior before generating log-odds ratios.
  • Bayesian meta-analysis workflow: The bayesmeta workflow fits the example data with a half-normal(0.5) prior for τ and provides printed results and a forest plot.The code loads the data, computes odds ratios for two randomized studies, performs the analysis, and visualizes it.

D.4 Sensitivity analyses

Sensitivity analyses vary the half-normal prior scale and the prior family while examining two four-study examples. Changing the prior family with a fixed median has little effect, whereas scale and noninformative-prior choices can materially affect uncertainty and extreme heterogeneity estimates.

  • Design: The sensitivity analyses use four-study MD and IRR examples, varying half-normal scales by factors of one-half or two and comparing prior families at a fixed median.Both examples originally use half-normal(0.5) priors.
  • Mean difference example: Larger plausible heterogeneity mainly lowers the effect confidence interval’s lower bound, while the median and upper bound change less.The pattern is attributed to greater weight for the most extreme and uncertain estimate under larger heterogeneity.
  • Prior-family sensitivity: Changing the prior family while holding its median fixed produces barely discernible differences in the overall effect estimates.The compared prior families have different shapes and properties, yet the resulting estimates remain remarkably similar.
  • Noninformative-prior comparison: Without a weakly informative prior, the empirical data may not sufficiently rule out implausible heterogeneity ranges in these few-study analyses.The paper notes that estimates differ more substantially under noninformative uniform or Jeffreys priors.
  • Scale sensitivity: Although confidence-interval width changes slightly across scale choices, the substantive conclusions do not change drastically.For the more optimistic half-normal(0.25) prior, the interval almost excludes zero.
  • Noninformative-prior comparison: Noninformative priors produce confidence intervals roughly 1.5 or 2.1 times wider than the proposed half-normal(0.5) prior.The comparison concerns the sensitivity analyses’ overall effect estimates and indicates a substantial precision difference.
Loading 2007.08352v3…