Source-linked AI summary

On Sparse variational methods and the Kullback-Leibler divergence between stochastic processes

Alexander G. de G. Matthews, James Hensman, Richard E. Turner, Zoubin Ghahramani

arXiv:1504.07027v2stat.ML

TL;DR

The paper addresses the lack of a rigorous, general process-level interpretation of sparse variational Gaussian-process objectives and the limits of augmentation-based consistency arguments. It supplies a measure-theoretic derivation covering broader inducing variables and likelihoods, then characterizes when augmentation is consistent and applies the framework to interdomain approximations and Cox processes.

  • Problem

    Sparse variational Gaussian-process methods lack a fully established process-level KL interpretation, while marginal consistency of augmented models may not ensure consistency with the original model.

  • Method

    The paper develops finite-dimensional and measure-theoretic KL-divergence derivations for stochastic processes, including non-data inducing points, all-function likelihoods, and deterministic augmentation.

  • Results

    The framework establishes the sparse variational objective as a process-level KL interpretation, shows marginal augmentation consistency is insufficient in general, and identifies a sufficient deterministic-augmentation condition.

  • Takeaways & Limitations

    The resulting framework provides a precise foundation for sparse variational inference and clarifies interdomain sparse approximations and Cox-process inference.

  • Takeaways & Limitations

    The treatment relies on measurability conditions for stochastic processes and on Fubini’s theorem for the interdomain transformation.

Abstract

from arXiv · show

The variational framework for learning inducing variables (Titsias, 2009a) has had a large impact on the Gaussian process literature. The framework may be interpreted as minimizing a rigorously defined Kullback-Leibler divergence between the approximating and posterior processes. To our knowledge this connection has thus far gone unremarked in the literature. In this paper we give a substantial generalization of the literature on this topic. We give a new proof of the result for infinite index sets which allows inducing points that are not data points and likelihoods that depend on all function values. We then discuss augmented index sets and show that, contrary to previous works, marginal consistency of augmentation is not enough to guarantee consistency of variational inference with the original model. We then characterize an extra condition where such a guarantee is obtainable. Finally we show how our framework sheds light on interdomain sparse approximations and sparse approximations for Cox processes.

1 Introduction

The paper reframes sparse variational Gaussian-process methods through rigorous KL divergences between stochastic processes, generalizing the treatment beyond data-point inducing variables and finite likelihood dependence. It also examines augmentation consistency and applications to interdomain approximations and Cox processes.

  • Motivation: Titsias’ inducing-point framework is influential for scalable Gaussian-process approximations, with inducing locations treated as variational rather than model parameters.This treatment is associated with protection from overfitting, although the paper questions whether the usual justification is exact.
  • Existing framework: The sparse variational distribution factors as p(fD\Z|fZ)q(fZ), combining a prior conditional with a variational distribution over inducing points.For conjugate likelihoods, the optimal q(fZ) is analytically Gaussian; non-conjugate extensions were developed subsequently.
  • Contributions: Marginal consistency after augmenting the index set does not by itself make variational inference equivalent to inference in the original model.The paper characterizes a condition—deterministic augmentation conditioned on the whole latent function—under which the desired equivalence does hold.
  • Contributions: The paper gives a shorter, more general, and intuitive proof of the KL-divergence theorem for stochastic processes, including inducing points not selected from the data.It presents this connection as previously unremarked in the literature and as a firmer foundation for sparse variational methods.
  • Paper scope: The framework is developed from finite-dimensional intuition through measure-theoretic stochastic-process arguments and applied to interdomain sparse approximations and Cox-process inference.The infinite-index-set treatment also permits likelihoods depending on infinitely many function values.

2 Finite index set case

For finite index sets, extending the variational distribution to all function values lets the full-distribution KL divergence reduce exactly to Titsias’ objective. The remaining values marginalize, so the inducing variables on both sides are bookkeeping rather than a moving optimization target.

  • Finite index set setup: The finite-index-set analysis adds ∗ = X\(D ∪ Z), representing function values outside the data and inducing locations.These points may matter for predictions at held-out inputs.
  • Finite index set derivation: The variational distribution is extended to include the additional function values, and its KL divergence with the full posterior is then expanded using conditional-factorization identities.The derivation cancels shared conditional terms and uses marginalization of conditional densities.
  • Finite index set result: The final integral is exactly Titsias’ KL-divergence objective, establishing equivalence between the full-distribution and original finite-dimensional formulations.The result follows after integrating out the function values in ∗.
  • Finite index set result: The inducing values appearing on both sides of the objective are an accounting artifact because all other function values marginalize.Different inducing-point choices require tracking different subsets of function values.

3 Infinite index set case

For infinite index sets, ordinary vector integration fails because no suitable infinite-dimensional Lebesgue measure exists. The paper therefore uses measure-theoretic KL divergence and derives a sparse inducing-point objective that generalizes prior finite-dimensional results.

  • The measure-theoretic obstacle: Infinite index sets cannot be handled by integrating against an infinite-dimensional vector measure.The notation for integrating over an infinite-dimensional vector is not mathematically valid.
  • The measure-theoretic obstacle: No nonzero measure is both translation invariant and locally finite on the relevant infinite-dimensional space.The only measure satisfying both properties is the zero measure.
  • Measure-theoretic KL divergence: The rigorous KL definition uses Radon-Nikodym derivatives and integrates with respect to one probability measure instead of infinite-dimensional Lebesgue measure.When absolute continuity fails, the KL divergence is defined as infinity; in finite dimensions, the definition reduces to the familiar density-based form.
  • General sparse derivation: The paper generalizes sparse inducing-point derivations to inducing points outside the data and likelihoods depending on infinitely many function values.The derivation also avoids assuming finite-dimensional marginals have Lebesgue densities.
  • General sparse derivation: The process-level KL divergence equals the sparse objective KL[QZ||PZ]−EQD[log L(Y|fD)]+log L(Y).The marginal likelihood term is constant with respect to Q and can be ignored during optimization.

4 Augmented index sets

Augmenting the index set preserves marginal consistency but does not generally preserve the variational KL objective. Equality is recovered when the added variables are deterministic functions of the original function values, yielding matching conditional distributions.

  • Augmented index sets: Augmented models introduce a finite set I alongside the original index set X, with parameters θ governing the augmented prior.The construction follows the augmentation argument used in sparse variational frameworks and interdomain Gaussian processes.
  • Marginal consistency: Marginal consistency alone does not ensure KL[QX||P̂X]=KL[QX∪I||P̂X∪I].The chain rule exposes an additional conditional KL-divergence on the augmented variables.
  • Marginal consistency: Equality holds only when QI|X=PI|X, QX-almost surely.Otherwise, variational inference in the augmented family optimizes a two-sided objective that is not equivalent to inference in the original model.
  • Transformed augmentation: For the transformation (Ĩ,X̃)=(X\D,D), the KL divergence on the data set generally differs from that on the full index set, except when Z⊂D.This follows directly from applying the KL chain rule to the transformed problem.
  • Deterministic augmentation: Deterministic augmentation makes the augmented KL divergence equal to the unaugmented divergence when fI=h(fX).The conditional KL term vanishes because both measures use the same delta-function conditional distribution.
  • Deterministic augmentation: The deterministic condition covers copied inducing points from X and interdomain inducing variables.The induced marginal on I is the push-forward of the marginal on X under h.

5 Examples

The paper applies its extended variational-process framework to interdomain inducing variables and Cox processes, including likelihoods dependent on infinitely many function values.

  • 5.1 Variational interdomain approximations: The interdomain approximation extends sparse variational inference to inducing variables defined as random variables indexed by a separate set I.These variables are deterministic conditional on the whole function, bringing them within the paper’s general framework.
  • 5.1 Variational interdomain approximations: The framework justifies optimizing interdomain inducing-point parameters θ while retaining a well-defined KL-divergence objective and protection from overfitting.The paper identifies this as enabling potentially improved sparse approximations.
  • 5.2 Approximations to Cox process posteriors: For Cox processes, the Poisson likelihood depends on all points in the index set because unobserved regions also provide information about intensity.Consequently, the relevant data index set is D = X.
  • 5.2 Approximations to Cox process posteriors: The Cox-process treatment requires the likelihood’s integral to exist almost surely and avoids relying on a density with respect to infinite-dimensional Lebesgue measure.The paper instead applies its more general form of Bayes’ theorem.
  • 5.2 Approximations to Cox process posteriors: For the specific inverse link used by Lloyd et al. (2015), the derivation continues as in that work, and Cox-process approximations could be combined with interdomain methods.The paper presents this combination as a possible direction for further work.

6 Conclusion and acknowledgements

The paper connects sparse variational inducing-point methods to rigorous KL-divergence objectives and broadens their theoretical scope. It also identifies the condition needed for principled augmentation and applies the framework to interdomain and Cox-process approximations.

  • 6 Conclusion and acknowledgements: The paper formalizes the connection between Titsias’ variational inducing-point framework and KL-divergence between stochastic processes.It presents this as clarifying the framework’s measure-theoretic foundations.
  • 6 Conclusion and acknowledgements: The proof framework allows inducing points that are not data points and removes unnecessary dependence on Lebesgue measure.The authors emphasize the role of Radon-Nikodym derivatives in the proof.
  • 6 Conclusion and acknowledgements: Marginal consistency alone does not guarantee a principled optimization objective for augmented models, whereas deterministic conditional inducing points guarantee one and protect augmentation parameters from overfitting.This condition is used to justify principled variational optimization in the augmented setting.
  • 6 Conclusion and acknowledgements: The extended theory handles principled interdomain sparse approximations and Cox processes whose likelihood depends on an infinite set of function points.These applications demonstrate the framework’s reach beyond finite data-index settings.
  • 6 Conclusion and acknowledgements: The authors suggest that the measure-theoretic formulation may support broader generalization, including Hilbert-space approaches for interdomain inducing points and other Bayesian models.These are presented as future possibilities rather than established results.
Loading 1504.07027v2…