Source-linked AI summary
A Generalization of Amari's Bayesian Duality
Mohammad Emtiyaz Khan, Thomas Möllenhoff
TL;DR
The paper addresses the limited scope and attention of Amari’s Bayesian duality. It connects Bayes’ rule to convex duality to derive a broader inverse mapping between posterior and likelihood, recovering Amari’s case while applying more generally. The resulting framework connects Bayesian duality with broader work in Bayesian inference and convex optimization, while retaining scope limitations in Amari’s original formulation.
Problem
Amari’s Bayesian duality received limited attention and was restricted to cases where likelihood and posterior share the same form, excluding many Bayesian models.
Method
The paper writes Bayes’ rule as an optimization problem and uses convex duality to find an inverse likelihood mapping in a dual function space.
Results
The convex-dual formulation recovers Amari’s posterior-likelihood bijection as a special case and can recover the likelihood from a given posterior up to normalization.
Takeaways & Limitations
The generalized viewpoint connects Amari’s Bayesian duality to broader literature on Bayesian inference and convex optimization and extends it beyond the original restricted setting.
Takeaways & Limitations
Amari’s original construction requires absorbing the prior into a valid exponential-family distribution, which is rarely possible and limits its general applicability.
Abstract
from arXiv · showhide
Amari's contributions to information geometry and machine learning are well known. Here, we revisit Amari's work on Bayesian duality which has not received as much attention. We connect Amari's Bayesian duality to a convex duality of Bayes' rule. Using this connection, we present a generalization of Amari's Bayesian duality and discuss its relevance for modern artificial intelligence.
1 Introduction
The paper revisits Amari’s underrecognized Bayesian duality, linking it to machine-learning ideas and generalizing its scope through convex duality of Bayes’ rule.
- Motivation: Amari’s Bayesian duality work received limited attention despite his broader foundational contributions to information geometry and machine learning.The paper situates this work alongside Amari’s contributions to stochastic gradient descent, recurrent networks, EM, and natural-gradient methods.
- Prior framework: The original framework connected likelihood and posterior manifolds through a bijection when both had the same exponential-family form.The roles of their dual coordinates were interchanged, but this setup did not cover most Bayesian cases.
- Generalization: The paper generalizes Amari’s theory by formulating Bayes’ rule through convex duality and connecting it to broader Bayesian-inference and convex-optimization literature.The authors also discuss the relevance of Bayesian duality for modern AI.
2 Amari’s Bayesian Duality Theory
Amari’s Bayesian duality connects likelihood and posterior manifolds when both have the same exponential-family form, swapping their dual coordinates through a bijection. This restriction fails in general Bayesian models, motivating a broader mapping based on posterior sufficient statistics and convex duality.
- 2.1 The Likelihood Function: Amari’s framework represents likelihood and posterior densities as exponential families with dual natural and expectation coordinates.These coordinates are related through the Legendre transform and form a dually-flat information-geometric structure.
- 2.2 The Posterior Distribution: Amari’s construction requires absorbing the prior into the likelihood base measure while preserving a valid exponential-family distribution with y as natural parameter.The required rearrangement is rarely possible, so the posterior generally does not have the same form as the likelihood.
- 2.2 The Posterior Distribution: When likelihood and posterior share the same exponential-family form, their manifolds are bijectively linked and their dual-coordinate roles are interchanged.Swapping y and θ in the likelihood representation yields the posterior representation, establishing Amari’s Bayesian duality.
- 2.3 An Example: In isotropic Gaussian examples, a uniform prior preserves the likelihood’s form and makes the posterior mean and natural parameter equal to y.The Gaussian likelihood and posterior therefore share the required representation in this special case.
- 2.4 How to Generalize Amari’s Framework?: Ridge regression violates Amari’s matching-form condition because its posterior variance differs from the likelihood variance.A mapping between the manifolds can still be found by expressing the likelihood using the posterior’s sufficient statistics.
- 2.4 How to Generalize Amari’s Framework?: The generalized construction uses Bayes’ rule as the forward map and finds t(θ) from the posterior as the reverse map, with the two operations related through convex duality.The likelihood is rewritten in terms of posterior sufficient statistics to support this dual structure.
3 Bayesian Duality via Convex Duality
The paper derives a convex-dual formulation of Bayes’ rule that generalizes Amari’s Bayesian duality beyond restrictive exponential-family cases. The dual formulation recovers likelihood mappings in Gaussian and ridge-regression examples, while allowing generic loss functions and non-unique solutions.
- Convex-dual formulation: Convex duality rewrites Bayes’ rule as an optimization problem whose dual searches for a likelihood mapping t(θ) in a dual function space.The primal optimum is the posterior, while the dual formulation connects posterior distributions to likelihood functions.
- Convex-dual formulation: For a posterior q(θ)=p(θ|y), the optimal t*(θ) recovers p(y|θ) up to a normalization constant and reproduces the posterior through Bayes’ rule.This establishes the inverse mapping that underlies the generalized Bayesian duality.
- Gaussian example: In the isotropic Gaussian example, Bayes’ rule reduces to quadratic optimization with m*=y, and the dual solution satisfies eλ*=m*=y.The resulting t*(θ) is proportional to N(y|θ,I), and substituting it into Bayes’ rule recovers the posterior.
- Generalization beyond Amari’s setting: The primal-dual Gibbs formulation realizes Amari’s bijection for the specified likelihood and posterior forms, but also applies when likelihoods are replaced by generic loss functions.The paper uses this broader applicability to motivate a ridge-regression example.
- Ridge-regression example: For ridge regression, the dual problem recovers the likelihood through optimization in transformed spaces, including the f-space representation with optimal parameters a*=y and B*=I.The formulation can recover the likelihood exactly, although the dual solution is not unique and therefore does not generally define a bijection.
- Ridge-regression example: Unlike Amari’s case, the generalized formulation may have multiple dual solutions, so posterior-to-likelihood mapping need not be bijective even though the interchanging operation is retained.The paper attributes this non-uniqueness to multiple ways of aggregating information.
4 Relevance for AI and Connections to Other Fields
The generalized Bayesian duality connects convex-duality formulations across Bayesian and non-Bayesian machine learning, including compact representations and ridge regression. The framework is presented as relevant to understanding relationships between models and data and as encompassing several existing dualities.
- Convex duality generalizes Amari’s Bayesian duality, with the dual problem providing an inverse mapping from posterior manifolds to likelihood manifolds that is sometimes bijective.
- The interaction between observations y and parameters θ motivates using Bayesian duality to study relationships between models and data.
- Bayesian duality connects to convex-duality methods for finding compact representations in data space, including representer theorems, kernel methods, support vector machines, and Gaussian processes.
- For ridge regression, the dual problem reformulates minimization over θ as maximization over an N-length real-valued dual vector α.
- Bayesian dualities with more expressive posterior forms are expected to contain dualities based on less flexible posteriors and non-Bayesian scenarios as special cases.
- Variational approximations use dual structures associated with simple posterior approximations, enabling practical algorithms based on Bayesian duality.
A.1 Derivation of the Dual Problem for Ridge Regression
The appendix evaluates the ridge-regression dual derivation by transforming the relevant integral, substituting the resulting expression, and solving stationarity conditions for the dual variables.
- A change of variables is used to evaluate the integral in Eq. 29 after bringing it into a suitable form.
- Substitution into Eq. 41 followed by taking the logarithm produces the objective used in the derivation.
- Differentiating the full objective in Eq. 28 with respect to eλ1 and eλ2 and setting the derivatives to zero yields the stationarity equations.
- Solving the first stationarity equation gives eλ1 = m∗(eλ2 + 1), which is then inserted into the other equation.
- Back-substitution gives eλ1 = m∗/σ2.
A.2 Connection to Convex Dual of Ridge Regression
The appendix connects primal and dual ridge-regression solutions through stationarity and the matrix inversion lemma. The resulting relation identifies the dual vector with the residual between observations and optimal predictions.
- The primal solution is the ridge solution.
- The dual solution is obtained from the stationarity condition of the dual problem.
- The matrix inversion lemma represents θ∗ in terms of α∗, linking the primal parameter vector to the dual solution.
- The relation α∗ = y − f∗ follows from the stationarity condition, where f∗ = xθ∗ is the optimal prediction.