Source-linked AI summary
A Widely Applicable Bayesian Information Criterion
Sumio Watanabe
TL;DR
The paper addresses the difficulty of estimating Bayes free energy for singular models when the RLCT depends on an unknown true distribution. It defines WBIC using posterior log-likelihood averaging at inverse temperature 1/log n and proves that WBIC has the same asymptotic expansion as Bayes free energy, including for singular and unrealizable models. WBIC therefore extends BIC to singular models while allowing RLCT estimation without knowing the true distribution.
Problem
Bayes free energy cannot generally be approximated by BIC for singular models, and its RLCT-based approximation is difficult to use because the true distribution is unknown.
Method
The paper defines WBIC as posterior log-likelihood averaging with an inverse temperature asymptotically governed by β∗log n →1.
Results
WBIC has the same asymptotic expansion as Bayes free energy even for singular and unrealizable models, and coincides with BIC for regular models.
Takeaways & Limitations
WBIC generalizes BIC to singular statistical models and enables RLCT estimation without information about the true distribution.
Takeaways & Limitations
Theoretical comparison of WAIC and WBIC in singular model selection remains an important problem for future study.
Abstract
from arXiv · showhide
A statistical model or a learning machine is called regular if the map taking a parameter to a probability distribution is one-to-one and if its Fisher information matrix is always positive definite. If otherwise, it is called singular. In regular statistical models, the Bayes free energy, which is defined by the minus logarithm of Bayes marginal likelihood, can be asymptotically approximated by the Schwarz Bayes information criterion (BIC), whereas in singular models such approximation does not hold. Recently, it was proved that the Bayes free energy of a singular model is asymptotically given by a generalized formula using a birational invariant, the real log canonical threshold (RLCT), instead of half the number of parameters in BIC. Theoretical values of RLCTs in several statistical models are now being discovered based on algebraic geometrical methodology. However, it has been difficult to estimate the Bayes free energy using only training samples, because an RLCT depends on an unknown true distribution. In the present paper, we define a widely applicable Bayesian information criterion (WBIC) by the average log likelihood function over the posterior distribution with the inverse temperature $1/\log n$, where $n$ is the number of training samples. We mathematically prove that WBIC has the same asymptotic expansion as the Bayes free energy, even if a statistical model is singular for and unrealizable by a statistical model. Since WBIC can be numerically calculated without any information about a true distribution, it is a generalized version of BIC onto singular statistical models.
1 Introduction
Singular models defeat the normal-approximation basis of BIC, while their Bayes free energy is governed by the RLCT, which depends on the unknown true distribution. The paper introduces WBIC and proves that it asymptotically matches Bayes free energy in singular settings while coinciding with BIC for regular models.
- Motivation: Singular models include neural networks, mixture models, reduced rank regressions, Bayesian networks, and hidden Markov models, and their likelihood functions cannot generally be approximated by normal distributions.Hierarchical layers, hidden variables, or grammatical rules commonly produce singularity.
- Motivation: For regular models, Bayes free energy is asymptotically approximated by BIC, whereas singular models generally require an RLCT-based expansion.The RLCT replaces half the parameter count in the singular-model asymptotics.
- Related work: RLCT theory has yielded results for several singular models, but practical evaluation cannot directly use them because the true distribution is unknown.The paper cites algebraic-geometrical procedures and studies of neural networks, mixtures, regressions, Bayesian networks, and hidden Markov models.
- Contribution: WBIC estimates Bayes free energy from training data by averaging the log likelihood over a posterior distribution, with inverse temperature β selected asymptotically near 1/log n.The paper defines WBIC to avoid requiring information about the true distribution.
- Results: WBIC has the same asymptotic behavior as Bayes free energy in singular models and coincides with BIC in regular models.The paper also reports lower numerical-calculation cost than direct Bayes free-energy computation and estimation of RLCTs without a known true distribution.
2 Statistical Models and Notations
This section defines the statistical-learning quantities used throughout the paper, including average log loss, KL divergence, optimal parameters, and the distinction between realizable and unrealizable models. It also formalizes singularity through the structure of the minimizing parameter set and Hessian.
- Statistical quantities: The model is represented by p(x|w), training samples follow the true distribution q(x), and L(w) denotes average log loss.The true distribution and model parameterization provide the basis for subsequent KL and asymptotic definitions.
- Statistical quantities: L(w) equals the true-distribution entropy plus D(q||p_w), so L(w) is minimized exactly when the model density equals q(x).The KL divergence is nonnegative, making S a lower bound for L(w).
- Assumptions: The paper assumes an interior parameter w0 minimizing L(w), while allowing multiple minimizing parameters that represent the same probability density.In singular models, the minimizing set can be an analytic or algebraic set with singularities.
- Notation: The paper defines the log density-ratio function and the functions K(w) and K_n(w) to analyze population and sample behavior.These quantities support the paper's formulation of the statistical-learning problem and its later asymptotic results.
- Definitions: A model is realizable when q(x)=p0(x); regularity additionally requires a unique minimizer and a strictly positive-definite Hessian at that point.Otherwise, the true distribution is unrealizable or singular for the model under the paper's definitions.
3 Singular Learning Theory
Singular learning theory addresses models whose optimal parameter sets may contain singularities, preventing quadratic approximations and standard BIC-style evaluation. Algebraic-geometrical resolution transforms these singularities into a standard form and yields the RLCT-based asymptotic expansion of Bayes free energy.
- Singular models: Singular models can have optimal parameter sets containing singularities, with non-positive-definite Hessians that prevent quadratic approximation of K(w).The paper studies conditions under which the asymptotic analysis remains tractable.
- Resolution of singularities: Theorem 1 uses a resolution map w = g(u) to transform arbitrary singularities into normal crossings on an analytic manifold.The representation introduces analytic and smooth functions whose local-coordinate forms support asymptotic analysis.
- RLCT: The real log canonical threshold λ and multiplicity m are birational invariants independent of the chosen resolution map.This invariance makes the RLCT an intrinsic quantity for the specified distribution, model, and prior.
- Birational invariance: The paper’s conjecture extends birational invariance of the parity quantity Q beyond realizable cases, while Lemma 2 establishes it when the true distribution is realizable.The paper notes that the general conjecture would follow from representing arbitrary nonnegative analytic functions as KL distances.
- Bayes free energy: Under the paper’s assumptions, Bayes free energy expands as F = nL_n(w0) + λ log n − (m − 1) log log n + R_n.Here R_n converges in law to a random variable as n approaches infinity.
- Motivation: The true distribution is unknown in applications, so λ and m cannot be directly used; the paper therefore develops a method to estimate F from training data.The stated purpose is to construct an estimate that does not require knowing the true distribution.
4 Main Results
The main results establish the optimal inverse temperature and show that WBIC reproduces the leading asymptotic terms of Bayes free energy. They also connect WBIC to BIC in regular settings, including unrealizable ones.
- Optimal inverse temperature: There is a unique optimal inverse temperature β* with 0 < β* < 1 when L_n(w) is nonconstant.The expected training log loss is decreasing in β, and β* is defined as the unique parameter satisfying the stated equation.
- WBIC asymptotics: Theorem 4 gives WBIC an asymptotic expansion whose first two main terms equal those of Bayes free energy.The theorem introduces a random remainder U_n with mean zero that converges in law to a Gaussian random variable.
- Corollaries: If the model’s parity is odd, the remainder term U_n is zero.This is stated as Corollary 1 under the condition Q(K(w), ϕ(w)) = 1.
- Corollaries: For β1 = β01/log n and β2 = β02/log n, the paper establishes convergence in probability for the corresponding WBIC quantities.The supplied passage states the scaling and convergence result but does not include the full displayed expression.
- Relation to BIC: For a true distribution regular for the model, WBIC and BIC differ by less than a constant-order term, even when the distribution is unrealizable by the model.Thus WBIC agrees asymptotically with the standard criterion in the regular setting while being formulated for broader settings.
- Singular-model evaluation: Replacing nL_n(w0) with nL_n(ŵ) is inappropriate in singular model evaluation because their difference can exceed the regular-model average d/2.In singular settings, the difference is asymptotically equal to the maximum value of a Gaussian process.
5 Proofs of Main Results
The proofs combine empirical-process arguments, regularity analysis, and blow-up maps. These steps establish the probabilistic and geometric properties needed for the main theorem and its corollaries.
- Empirical-process setup: The proof introduces an empirical process η_n(w) and uses its convergence to a random process and related random variables.The supplied proof passages state convergence in law and use these quantities in subsequent asymptotic bounds.
- Control away from the optimum: Outside the region where K(w) is small, the relevant contributions converge to zero in probability under β = β0/log n.The argument uses the condition 1 − r > 1/2 and boundedness of K on the compact parameter set.
- Regular case: When q(x) is regular for the model, positive definiteness of J(w*) near the optimum supplies uniform eigenvalue bounds for the proof.The minimum and maximum eigenvalues are taken over the region K(w) ≤ ε.
- Blow-up argument: A blow-up map resolves the local geometry and shows k_i = 1 in each coordinate chart, yielding Q(K(w), ϕ(w)) = 1.The Jacobian determinant of the blow-up is then used in the argument.
- Optimal temperature: The uniqueness of β* follows by combining the Cauchy–Schwarz inequality with the nonconstancy assumption and applying the mean value theorem.The proof derives existence of β* in the interval 0 < β* < 1.
5.4 First Preparation for Proof of Theorem 4
The first preparation for proving Theorem 4 represents the relevant random quantities as integrals on the resolved manifold. Existing convergence results for a Gaussian process then support their asymptotic evaluation.
- Proof setup: The proof of Theorem 4 reduces the problem to evaluating Eβ of the quantities introduced through the model’s likelihood and posterior expressions.The preparation defines A_n and B_n as the two values whose asymptotics are studied.
- Resolved-manifold integration: The resolution theorem converts integrals over {w ∈ W; K(w) < ε} into integrals over the manifold M.Local coordinate charts and smooth partition functions are used to represent the transformed integrals.
- Random-process asymptotics: The random process ξ_n(u) converges in law to a Gaussian random process under the Fundamental Conditions.When q(x) is realizable and u^2k = 0, an additional property is invoked in the proof.
- Local asymptotic analysis: The proof studies a Schwartz distribution and its local-coordinate measure to derive asymptotics involving δ(t − u^2k)|u^h| as t approaches zero.The measure is supported on the set where the transformed coordinates u_a vanish.
- Local lemma: Lemma 4 supplies the local asymptotic expansion for a real-valued C1 function G(u^2k, u_k, u) as t approaches zero.The lemma is obtained by applying the stated relation to the distribution Y(t).
5.6 Proof of Lemma 2
The proof analyzes auxiliary quantities using resolution maps and asymptotic expansions, establishing that the relevant coefficient is independent of the chosen resolution map.
- Assumption: For a realizable true distribution, the support of du* lies in u_2k = 0, allowing the preceding expansion to be applied.The realizability assumption supplies the support condition used in the proof.
- Case analysis: When Q(K(g(u)), ϕ(g(u))) = 1, an odd exponent makes σ_k^a take both signs, forcing the main-order coefficient to vanish.The sign variation is the key cancellation used in this case.
- Case analysis: When Q(K(g(u)), ϕ(g(u))) = 0, some function Φ(w) makes the main-order term nonzero.This distinguishes the second case from the cancellation case.
- Conclusion: Therefore, Q(K(g(u)), ϕ(g(u))) does not depend on the resolution map.This is the stated conclusion of the lemma.
- Proof strategy: The proof studies A_n and B_n through resolution-map coordinates and asymptotic expansions involving the gamma function.These quantities are then combined using identities to complete the argument.
5.8 Proof of Corollary 1
The proof establishes the corollary by defining a centered random variable and evaluating the posterior expectation at the optimal inverse temperature.
- Coordinate analysis: The support of du* is contained in {u = (0, u_b)}, and the proof uses local-coordinate sign behavior to analyze Θ.If Q(q, p, ϕ) = 1, σ_k^a takes both +1 and −1 in arbitrary local coordinates.
- Optimal temperature: Using the optimal inverse temperature β*, the proof defines T = 1/(β*log n) and expresses F as E_β*[nL_n(w)].Theorem 2 and Theorem 4 are then used to complete the derivation.
5.10 Proof of Corollary 3
The proof of Corollary 3 evaluates posterior expectations using matrix expansions under regularity, where the Fisher information matrix is positive definite.
- Regular-model setup: In a regular model, the maximum likelihood estimator converges to w_0 in probability, and J_n(w) is defined as a d × d matrix.The proof introduces w* to relate the sample matrix to the population matrix.
- Matrix convergence: J_n(w*) = J(w_0) + o_p(1), and J(w_0) is positive definite.Positive definiteness follows from regularity.
- Asymptotic factor: The resulting determinant term is (nβ)^(d/2) det(J(w_0) + o_p(1))^(1/2).This is the leading matrix factor used in the asymptotic evaluation.
- Remainder control: The proof also uses nK_n(ŵ) = O_p(1), because the true distribution is regular for the statistical model.Together with the preceding expansion, this completes Theorem 5.
6 A Method How to Use WBIC
This section illustrates WBIC for model selection and RLCT estimation through reduced rank regression experiments, emphasizing numerical evaluation rather than theorem proving.
- Model selection: WBIC is evaluated for model selection across six reduced rank regression models with H = 1, 2, 3, 4, 5, 6.The experiment compares WBIC1 and WBIC2 using averages and standard deviations.
- Experimental setup: The reduced rank regression experiment fixes σ = 0.1, M = N = 6, and generates 100 sets of n = 500 samples from a true rank-three model.Posterior sampling uses the Metropolis method with β = 1/log n.
- Model selection: With true rank H_0 = 3, the true model H = 3 was selected in all 100 independent training-sample sets.The experiment used M = N = 6 and n = 500 training samples per set.
- RLCT estimation: RLCT estimation uses β_1 = 1/log n and β_2 = 1.5/log n in the same experiment.The estimate is obtained from Corollary 3.
- RLCT estimation: Theoretical RLCTs were well estimated; differences arose from smaller-order terms than log n and, when m = 2, from a log log n term.For unrealizable distributions, the stated theoretical value is λ = H(M + N − H)/2.
7 Discussion
The discussion compares WBIC with alternative free-energy evaluation methods, clarifies its relation to WAIC, BIC, RLCTs, and model-selection goals, and connects WBIC to algebraic-geometric properties such as parity.
- 7.1 WAIC and WBIC: WBIC and WAIC generalize information criteria to singular statistical models, while coinciding with BIC and AIC, respectively, in realizable regular models.
- 7.1 WAIC and WBIC: WBIC has the asymptotic expansion WBIC = BIC + op(1) in regular models.
- 7.1 WAIC and WBIC: Singular models can yield WAIC and WBIC values smaller than AIC and BIC when the prior is positive at the optimal parameter set.
- 7.2 Other Methods How to Evaluate Free Energy: The all-temperatures method can estimate free energy without asymptotic theory, but accurate calculation requires sufficiently many temperatures and posterior expectations, causing high computational costs.
- 7.2 Other Methods How to Evaluate Free Energy: Importance sampling accuracy strongly depends on the choice of the approximating function H(w), while the two-step method assumes theoretical RLCT values for all relevant cases.
- 7.2 Other Methods How to Evaluate Free Energy: WBIC requires asymptotic theory but not theoretical RLCT values, and it can be used when the true distribution is unrealizable by the statistical model.
- 7.3 Algebraic geometry and Statistics: The paper defines parity for a statistical Q(K(w), ϕ(w)) and proves that it affects WBIC's asymptotic behavior and likelihood-fluctuation terms.
- 7.3 Algebraic geometry and Statistics: Parity is related to analytic continuation, restricted parameter sets, and the difference between K(w) and K_n(w).
8 Conclusion
The paper proposes WBIC for singular and unrealizable statistical models and proves that it has the same asymptotic expansion as the Bayes free energy.
- WBIC applies when a true distribution is unrealizable by and singular for a statistical model, while matching the Bayes free energy asymptotically.