Source-linked AI summary
Penalising model component complexity: A principled, practical approach to constructing priors
Daniel P. Simpson, Håvard Rue, Thiago G. Martins, Andrea Riebler, Sigrunn H. Sørbye
TL;DR
Specifying priors becomes difficult as hierarchical models grow more complex, while existing general frameworks have important limitations. The paper develops Penalised Complexity priors that penalise departures from simpler base models and reports improved interpretability alongside a practical, tunable framework for prior construction.
Problem
As models become more complex, practical methods for specifying joint priors become scarce, while existing reference-prior frameworks are difficult to apply to moderately complex hierarchical models.
Method
The paper constructs Penalised Complexity priors by measuring deviation from a simpler base model and calibrating the resulting penalty with a user-defined scaling parameter, including multivariate extensions.
Results
The paper reports that decomposing models into favoured parametric and extra nonparametric components improves interpretability and indicates how far data require departure from the base model.
Takeaways & Limitations
PC priors provide a principled, widely applicable framework whose tuning supports vague, weakly informative, or strongly informative default priors.
Takeaways & Limitations
For sparsity, the basic PC prior can encode incorrect information when it treats each component as independently zero rather than modelling the vector's sparsity.
Abstract
from arXiv · showhide
In this paper, we introduce a new concept for constructing prior distributions. We exploit the natural nested structure inherent to many model components, which defines the model component to be a flexible extension of a base model. Proper priors are defined to penalise the complexity induced by deviating from the simpler base model and are formulated after the input of a user-defined scaling parameter for that model component, both in the univariate and the multivariate case. These priors are invariant to reparameterisations, have a natural connection to Jeffreys' priors, are designed to support Occam's razor and seem to have excellent robustness properties, all which are highly desirable and allow us to use this approach to define default prior distributions. Through examples and theoretical results, we demonstrate the appropriateness of this approach and how it can be applied in various situations.
1. INTRODUCTION
The paper addresses the difficulty of specifying sensible default priors for increasingly complex hierarchical Bayesian models by introducing interpretable Penalised Complexity priors. PC priors use user-controlled flexibility information and principled design goals, while remaining useful rather than uniquely optimal or universal.
- Motivation: Bayesian model quality depends on the quality of its prior and other inputs, making prior construction an unavoidable part of Bayesian data analysis.This challenge intensifies as models become more complex and expert information about structural parameters becomes harder to elicit.
- Contribution: The paper introduces Penalised Complexity priors as a broad framework for building informative priors for a large class of hierarchical models.The framework aims to provide sensible defaults, especially for general Bayesian computation software such as R-INLA.
- Contribution: PC priors encode information through four principles and a single user-set parameter controlling the flexibility allowed in the model.That parameter can be calibrated using weak information or another subjective criterion, while making the prior’s encoded information more interpretable and easier to elicit.
- Contribution: The proposed principles are designed to yield invariance to reparameterisations, a connection to Jeffreys’ prior, support for Occam’s razor, and empirical robustness.The authors present PC priors as unified, understandable, and straightforward enough for general practitioners.
- Scope and limitations: The authors do not claim that PC priors or their principles are optimal, unique, or universal, but argue that they are useful, understandable, conservative, and better than doing nothing.This is presented as a deliberately modest claim about the framework’s practical value.
- Scope and limitations: The paper restricts its main analysis to additive hierarchical models with latent structures whose components are controlled by a small number of flexibility parameters.It excludes settings where the number of flexibility parameters grows asymptotically, except for a sparse linear-model discussion involving a possible prior modification.
2. A GUIDED TOUR OF NON-SUBJECTIVE PRIORS FOR BAYESIAN HIERARCHICAL MODELS
This section reviews non-subjective priors for Bayesian hierarchical models, focusing on prediction and covering objective, weakly informative, risk-averse, and base-model approaches. It argues that hierarchical-model complexity and overfitting make prior suitability and stability central concerns.
- Scope: The review focuses on prior-specification methods applicable across the paper’s examples and on priors for prediction rather than testing.The authors note that testing may require different priors from prediction.
- Objective priors: Objective priors aim to inject minimal information but are design-dependent and difficult to derive for complex hierarchical models.Reference priors have a less successful history in hierarchical models because they depend on the model and can change when components are added or removed.
- Weakly informative priors: Weakly informative priors encode limited, plausible information and can produce better predictive inference than reference priors in some settings.The passage gives the half-Cauchy prior on a normal standard deviation as an example.
- Risk-averse priors: Risk-averse reuse of priors from similar applications is ad hoc, can import inappropriate choices, and includes known problematic defaults such as Γ(ϵ, ϵ) for precision parameters.The section emphasizes checking prior suitability for the specific application.
- Base models and overfitting: Independent priors on hyperparameters can fail to control flexibility in overspecified or partially identifiable models, allowing overfitting priors to overwhelm support for the base model.The section motivates shrinkage or sparsity and identifies stability as a primary concern for hierarchical inference.
D4: The prior should limit the flexibility of an over-specified model
The prior should shrink flexible, over-specified models toward simpler structures. Good shrinkage properties are presented as important for obtaining good inference in hierarchical models.
- D4: The prior should limit the flexibility of an over-specified model: Priors for over-specified models should have good shrinkage properties.This desideratum is linked to the discussion in Section 2.4.
- D4: The prior should limit the flexibility of an over-specified model: Without good shrinkage properties, priors are unlikely to support good inference for hierarchical models.The passage states this likelihood judgment directly.
D6: The prior should control what a parameter does, rather than its
A sensible prior should control a parameter’s role in the model rather than depend on its numerical representation. This motivates reparameterisation-invariant priors based on distances or divergences between models.
- D6: The prior should control what a parameter does, rather than its numerical value: Priors should be locally indifferent to parameterisation, so posterior inference does not depend on using a standard deviation, variance, or precision.The principle concerns equivalent representations of the same Gaussian random effect.
- D6: The prior should control what a parameter does, rather than its numerical value: Jeffreys’ work motivates using distance between models as a scale for constructing priors invariant to reparameterisation.Later work extends these ideas to objective priors for Bayesian hypothesis testing.
- D6: The prior should control what a parameter does, rather than its numerical value: Divergence-based priors derive proper priors from divergence measures between competing models.These priors are called divergence-based (DB) priors.
3. PENALISED COMPLEXITY PRIORS
This section defines penalised complexity priors by shrinking model components toward a simpler base model using principled complexity measurement, constant-rate penalisation, and user-defined scaling. It establishes reparameterisation invariance and derives the Gaussian random-effect precision prior, contrasting its behaviour with commonly used Gamma priors.
- Principles: PC priors prefer the simpler base model by penalising deviations according to the asymmetric Kullback–Leibler distance and a constant decay rate.The distance is d(f||g) = √(2KLD(f∥g)); constant-rate penalisation gives a prior whose relative change does not depend on d.
- Principles: User-defined tail probabilities determine the prior’s scale, allowing users to control informativeness through an interpretable bound U and tail weight α.The scaling parameter satisfies a condition based on an interpretable transformation Q(ξ), a sensible upper bound U, and tail-event weight α.
- Properties: PC priors are invariant to reparameterisation and align with desiderata including limited flexibility, effect-based control, informativeness, and computational feasibility.These properties follow from defining the prior on the distance scale and transforming it to the parameter of interest.
- Gaussian random-effect precision: For a Gaussian random effect, the PC prior for precision τ is independent of the structure matrix R, including when R is rank-deficient.The natural base model is the absence of the random effect, and the prior is derived as a type-2 Gumbel distribution.
- Gaussian random-effect precision: All commonly used Γ(a, b) priors with finite expected value a/b overfit, whereas the PC prior has πd(0) = 0.The PC prior corresponds to an exponential distribution on the standard deviation, with λ = −ln(α)/U controlling the penalty.
- Gaussian random-effect precision: The PC prior and matched-variance Γ(1, b) prior differ substantially at low precisions and in tail behaviour.The comparison uses U = 0.3/0.31 and α = 0.01, with b = 0.0076 selected to equalise marginal random-effect variances.
4. SOME PROPERTIES OF PC PRIORS
The section shows that PC priors’ behaviour near the base model drives their statistical properties, while tail behaviour is generally less important in moderate-dimensional models. It also establishes useful asymptotic and risk properties, but identifies failure of the basic construction for sparse high-dimensional shrinkage.
- Local behaviour: Near a regular base model, the PC prior is a tilted Jeffreys’ prior, with informativeness controlled by the scaled Riemannian distance parameter λ.The distance is defined using the Fisher information metric, and λ determines the degree of tilting.
- Asymptotic behaviour: At boundary base models, the prior’s behaviour near zero determines the posterior’s large-sample behaviour, and Principle 1 ensures correct asymptotic behaviour.For a finite prior density at the base model, the asymptotic behaviour matches that of the maximum likelihood estimator.
- Bayesian testing: Theorem 2 establishes consistent Bayes factors for testing ζ = 0 against ζ > 0 when the prior on ζ does not overfit.B01 →∞ under H0 and B01 →0 under H1.
- Shrinkage behaviour: For Gaussian random-effect precision, lighter-than-Student-t tails on the distance imply zero prior density at both zero and unit shrinkage, while λ controls exponential decay.Theorem 3 gives the conditions πκ(0) = 0 and πκ(1) = 0.
- Risk properties: In the normal means model, U = 5 yields risk results almost identical to the half-Cauchy prior, whereas U = 1 exceeds the minimax rate for large ∥x0∥.The difference from the half-Cauchy occurs only for really large ∥x0∥ and decreases as U increases.
- Sparse shrinkage limitation: The basic PC prior fails for high-dimensional sparse shrinkage because independent zero components are the wrong base model and its tails cannot support sparse, moderately sized signals.When true sparsity s0 = o(p), the required λ grows like O(…), forcing U = O(p−1 log(p)).
5. THE STUDENT-T CASE
The Student-t case treats degrees of freedom as a flexible extension of the Gaussian base model, with PC priors centered on the Gaussian limit. Compared with exponential priors, PC priors better preserve Occam’s razor and show more stable inference and coverage in simulations.
- Model setup: The Student-t distribution extends the Gaussian base model, attained when ν = ∞, while standardized unit precision preserves interpretability for finite ν > 2.The analysis focuses solely on the degrees-of-freedom parameter ν, with precision fixed.
- Prior behavior: For any absolutely continuous prior on ν > 2 with E(ν) < ∞, the induced prior has πd(0) = 0 and overfits.The zero density at the Gaussian base model prevents the mode from being centered at d = 0.
- Prior construction: Exponential and finite-interval uniform priors place zero density at the base model, whereas the PC prior is defined to have its mode at d = 0.The PC prior scale is elicited through Prob(ν < U) = α, with λ = −log(α)/d(U).
- Prior comparison: For exponential priors with means 5, 10 and 20, the distance-scale modes correspond to ν = 7.0, 17.0 and 37.0, respectively, so shrinkage targets depend on the prior.All six exponential and uniform priors shown have density zero as d → 0.
- Simulation results: PC-prior performance was barely affected by α and remained good across sample sizes, whereas exponential-prior inference remained strongly dependent on the prior mean even for n = 1 000 and n = 10 000.With ν = 100, PC priors widen credible intervals for few observations but learn the high degrees of freedom as data increase.
- Coverage: At the 95% level, PC-prior coverage probabilities were always at least 0.9, while exponential-prior coverage ranged from sensible to far too low, including zero in several settings.PC-prior coverage was generally slightly higher than the nominal level.
6. DISEASE MAPPING USING THE BYM MODEL
The BYM disease-mapping model is reparameterised into interpretable, nearly orthogonal components so that principled complexity-penalising priors can address scaling and confounding issues. In German larynx-cancer data, the posterior indicates that spatial structure explains nearly all marginal variance despite the prior favoring lower spatial contributions.
- BYM model: The BYM model combines an overall intercept, covariate effects, an unstructured Gaussian effect, and a spatial effect encoding similarity between nearby regions.The spatial component uses an intrinsic Gaussian Markov random field, while the unstructured effect has precision τvI.
- BYM model: The intrinsic spatial component has a 1-vector null space, so imposing 1T u = 0 prevents confounding with the intercept.Its density is invariant to adding a constant to u.
- Model limitations: The original BYM formulation has graph-dependent marginal variance and does not separate the structured and unstructured components independently.Consequently, priors cannot transfer directly across graphs, and the two components are difficult to interpret separately.
- Reparameterisation: Reparameterisation standardises the spatial component and introduces mixing parameter φ and precision τ, which have nearly orthogonal interpretations.The spatial and unstructured components explain fractions φ and 1 − φ of the marginal variance, while τ controls the marginal precision.
- PC priors: The PC-prior construction treats no effects from the spatial and unstructured components as the base model for τ, and no spatial dependency, φ = 0, as the fixed-precision base model.Increasing φ adds spatial dependency while keeping marginal precision constant.
- German larynx-cancer example: The posterior for φ concentrates around 1 although the PC prior assigns 2/3 probability to φ < 1/2, implying that only the spatial component contributes to marginal variance.This result comes from the larynx-cancer mortality data for 544 German districts, with 7283 deaths and an average of 13.4 per region.
7. MULTIVARIATE PROBIT MODELS
The PC prior methodology extends naturally to multivariate parameters by penalising distance from a base model on smooth manifolds. Multivariate probit models illustrate this construction for correlation matrices, including parameterisation-invariant priors and applications to exchangeable dependence structures.
- Multivariate probit models: Existing priors for multivariate probit correlation matrices are problematic because the joint uniform prior has informative marginals and Jeffreys’ prior concentrates near ±1 in high dimensions.A multivariate Gaussian prior restricted to positive-definite matrices is also described as unconvincing.
- Multivariate PC priors: PC priors extend to multivariate parameters on smooth manifolds while retaining the features of the univariate construction.The parameter space may be a manifold such as the symmetric positive definite matrices or correlation matrices.
- Multivariate PC priors: The multivariate construction uses distance level sets as coordinates, assigning an exponential distribution to distance from the base model.The level sets form disjoint embedded submanifolds, or leaves, that decompose the parameter space.
- Multivariate PC priors: Computational geometry can approximate multivariate PC prior densities in low dimensions, while exact expressions are available when distance level sets are simplexes or spheres.Linear and quadratic distance forms provide tractable cases for deriving and simulating these priors.
- Correlation matrices: For correlation matrices, a reparameterisation based on γij = −log(sin(θij)) places the PC prior in the linear multivariate case.The correlation matrix is represented as R = BBT, with p = q(q −1)/2 parameters in the γ-parameterisation.
- Application: Six Cities study: In the exchangeable model, the PC prior automatically adjusts densities for negative and positive ρ under the positive-definiteness constraint −1/(m −1) < ρ < 1.Reusing λ = 0.1 and 1.0 preserves comparable penalisation of distance from the base model across the general and reduced parameterisations.
8. DISTRIBUTING THE VARIANCE: HIERARCHICAL MODELS, AND ALTERNATIVE DISTANCES
The section extends PC priors to hierarchical models by using the global model structure to control overall variance and allocate contributions among components. It also describes recursive variance decompositions and practical multivariate approximations for constructing these priors.
- Hierarchical models: Hierarchical models require global rather than independent component priors because practitioners may not know each component’s relative effect or be able to control joint variance contributions.The framework addresses this by controlling the overall variance of the linear predictor and how each term contributes to it.
- Hierarchical models: A multivariate PC prior is placed on component variance fractions to account for the global model structure.The approach controls both overall variance and the allocation among linear and nonlinear effects.
- Computational construction: Within the PC-prior framework, priors can be built automatically for new models and datasets from graphical structure and model design, without depending on observations.The resulting priors can be integrated into software such as R-INLA or STAN.
- Computational construction: Taylor expansions provide practical multivariate PC-prior approximations: second order for an interior base model and linear for a boundary base model.The second-order implementation uses a numerical Hessian at the base model, and the joint prior for weights must be recomputed when φ changes.
- Variance decomposition: Variance is recursively distributed across covariates and nested effects, with component weights constrained to the 2-simplex and equal weights representing conditional exchangeability.For three covariates, the base model sets w1 = w2 = w3 = 1/3, while deeper decomposition separates linear from purely nonlinear effects.
- Variance decomposition: The recursive structure determines each term’s variance fraction by multiplying weights along its path from the top node.For example, the linear Age effect contributes w1(1−φ1), whereas g3(CR) contributes w3.
9. DISCUSSION
The discussion presents PC priors as a principled, component-wise response to weaknesses in practical prior specification, while identifying unresolved guidance and theoretical questions. It also emphasizes improved interpretability, robustness, and confidence against over-fitting through scale-based prior specification.
- Motivation: PC priors address arbitrary prior selection by providing a principled, widely applicable approach to specifying priors for difficult parameters.The paper frames prior selection as fundamental to Bayesian statistics and criticizes current practice for arbitrariness and inadequate sensitivity analysis.
- Component-wise specification: PC priors are defined component-wise, unlike reference priors, whose construction depends on the global model structure.This aligns PC priors with component-wise and often additive modelling workflows using directed acyclic graphs.
- Limitations and future work: PC priors still require case-by-case derivation, better scaling guidance informed by the global model, and implementation as defaults in packages such as R-INLA.The negative-binomial over-dispersion example illustrates that some components cannot be separated cleanly from other model parameters.
- Interpretability and robustness: The discussion links flexible extensions of simpler parametric base models to improved interpretability and added robustness in applications.The decomposition separates favoured parametric structure from extra nonparametric components, with a flexibility parameter indicating departure from the simpler model.
- Limitations and future work: Further theoretical work should examine prior-tail effects, since exponential tails appear sufficient in low dimensions while half-Cauchy tails differ in high-dimensional sparse settings.The paper notes that it remains unclear whether the observed difference is genuinely attributable to the tail.
- Practical impact: Scale-based PC priors simplify interpretation and communication while increasing confidence that prior choices do not force over-fitting.The discussion presents this as a route toward more sound Bayesian analysis and fewer copied prior choices from similar research articles.
APPENDIX A: DERIVATION OF PC PRIORS … A.4 The PC prior for the variance weights in additive models
Appendix A derives PC priors for singular models, multivariate-normal precisions, BYM mixing parameters, and additive-model variance weights. The derivations establish invariant distance constructions, scale-based calibration, and computational strategies for evaluating the resulting priors.
- A.1 The “distance” to a singular model: A.1 defines a renormalised distance from a singular base model using a finite Kullback–Leibler divergence limit.The construction yields a finite, parameterisation-invariant distance for defining PC priors at singular models.
- A.2 The PC prior for the precision in a multivariate Normal distribution: A.2 derives the PC prior for a multivariate-normal precision by approximating the singular base model with τ0 and taking τ0 →∞.For a full-rank R, the divergence scales with pτ0/τ; generalized inverses and determinants give the same result when R is rank-deficient.
- A.2 The PC prior for the precision in a multivariate Normal distribution: A.2 calibrates the prior’s scale through Prob(1/√τ > U) = α, giving θ = −ln(α)/U.An exponential prior on the distance uses rate λ = θ/√pτ0 and produces the type-2 Gumbel distribution.
- A.3 The PC prior for the mixing parameter in the BYM-model: A.3 represents the BYM flexible model as √(1 −φ)v + √φu, with v as the base model at φ = 0.The covariance is Σ1(φ) = (1 −φ)I + φR−1, and sparse-matrix traces, determinants, or eigenvalues support computation.
- A.3 The PC prior for the mixing parameter in the BYM-model: A.3 determines the penalisation parameter from Prob(φ < u) = α, requiring α > d(u)/d(1).For generalized hierarchical models, the matrix determinant lemma reduces the cost of evaluating det(Σ1(φ)) across φ values.
- A.4 The PC prior for the variance weights in additive models: A.4 computes the joint PC prior for additive-model variance weights from the covariance of a standardised linear predictor.Sparse extraction matrices Ai represent the required elements or linear combinations of scaled second-order random-walk components.
- A.4 The PC prior for the variance weights in additive models: A.4 reparameterises weights as ˜wi = log(wi/wn), placing the base model at ˜w = 0 before numerically approximating the KLD Hessian.The resulting PC prior follows from the Hessian approximation and Eq. (7.6).
APPENDIX B: PROOFS · B.1 Proofs of Theorems 1 and 5
The appendix proves Theorem 5 and notes that Theorem 1 follows by the same argument. The proof uses a large-ν expansion of the KLD between a unit-precision Student-t distribution and a standard Gaussian, together with a tail condition on πν(ν).
- B.1 Proofs of Theorems 1 and 5: Theorem 5 is proved in this appendix, while Theorem 1 follows along the same lines.
- B.1 Proofs of Theorems 1 and 5: The relevant divergence is the KLD between a unit-precision Student-t distribution with d.o.f. ν and a standard Gaussian.
- B.1 Proofs of Theorems 1 and 5: For large ν, the KLD expansion begins with 3 4ν−2.
- B.1 Proofs of Theorems 1 and 5: The next term in the large-ν expansion is 3 2ν−3.
- B.1 Proofs of Theorems 1 and 5: The expansion includes a remainder term O(ν−4).
- B.1 Proofs of Theorems 1 and 5: Because πν(ν) has a finite first moment, it is o(ν−2) as ν →∞.
B.2 Proof of Theorem 2 · B.3 Proof of Theorem 3
Theorem 2’s proof establishes Bayes-factor consistency under regular alternatives and derives rates under regular and irregular asymptotic behavior. It also extends the argument to priors with boundary behavior π(ζ) = O(ζ^k).
- B.2 Proof of Theorem 2: Under H1, Bayes-factor consistency follows because the alternative domain is open and contains only regular model points.The proof considers H0: ζ = 0 against H1: ζ ∼ π(ζ), with π(0) ∈ (0, ∞).
- B.2 Proof of Theorem 2: B01(y_n) = O_p(n^1/2) under regular asymptotic behavior.This follows because the first quotient is O_p(1) under Bochkina and Green’s Assumption M1.
- B.2 Proof of Theorem 2: Under H0, the posterior π(ζ|y_n) converges to a truncated normal distribution with mean parameter 0 and variance parameter n−1/2v.The proof evaluates this asymptotic density at ζ = n−1/2.
- B.2 Proof of Theorem 2: B01(y_n) = O_p(n) when the model has irregular asymptotic behavior.The proof replaces the truncated normal approximation with the appropriate Gamma distribution.
- B.2 Proof of Theorem 2: The Bochkina–Green results extend the argument to parameter-invariant extensions of the non-local priors considered by Johnson and Rossell.This extension is stated for the irregular-asymptotic case after replacing the truncated normal distribution by the appropriate Gamma distribution.
- B.2 Proof of Theorem 2: If π(ζ) = O(ζ^k) with k > −1 as ζ → 0, then B01(y_n) = O_p(n^1/2+k).The rate follows from a direct extension of the preceding argument.
B.4 Proof of Theorem 4
The proof establishes conditions under which a suitably rescaled prior on standard deviations induces a marginal prior for β concentrated on δp-sparse vectors. It further states that the prior on δp-dimensionality can be centred at the target’s true sparsity s.
- Theorem 6: For sufficiently large λ = λ(p) ↑∞, the rescaled prior πλ(σ) = λπ(λσ) gives the marginal prior for β mass on δp-sparse vectors.The setup assumes a prior π(σ) with π(0) = 1 and considers β ∼ N(0, D2) with Dii ∼ πλ(σ).
- Theorem 6: The prior on the δp-dimensionality is centred at the true sparsity s under the theorem’s stated sufficient condition.Here, δp = p−1 and s denotes the true sparsity of the target vector.
- Proof: The proof derives the result from bounds based on the definition of πλ(σ) and standard estimates for the exponential integral.It also uses πλ(σ) ≲1 and a Taylor expansion of the Lambert-W function for an upper bound.