Source-linked AI summary
On a generalization of the Jensen-Shannon divergence
Frank Nielsen
TL;DR
The paper addresses how to generalize bounded, support-flexible Jensen-Shannon divergence beyond scalar skewing. It introduces vector-skew Jensen-Bregman and Jensen-Shannon divergences, characterizes their properties, constructs symmetric parameterized families, and develops iterative centroid computation for mixture families. The main outcome is a bounded vector-skew framework with applications to categorical distributions and normalized histograms.
Problem
The paper seeks a broader Jensen-Shannon construction that retains boundedness and support flexibility while replacing scalar skewing with vector-valued control.
Method
The paper vector-skews weighted Jensen-Bregman divergences, derives Jensen-Shannon-type divergences, constructs symmetric parameterized families, and applies a convex-concave procedure to mixture-family centroids.
Results
The proposed vector-skew Jensen-Shannon divergences preserve boundedness, admit symmetric parameterized families, and support Jensen-Shannon centroid computation for mixture-family densities.
Takeaways & Limitations
Vector parameters provide additional controls for tuning Jensen-Shannon-type dissimilarities, while the centroid procedure covers categorical distributions and normalized histograms within mixture families.
Takeaways & Limitations
For continuous mixture families with shared support, computing negentropy is generally intractable because of the differential-entropy log-sum term.
Abstract
from arXiv · showhide
The Jensen-Shannon divergence is a renown bounded symmetrization of the Kullback-Leibler divergence which does not require probability densities to have matching supports. In this paper, we introduce a vector-skew generalization of the scalar $α$-Jensen-Bregman divergences and derive thereof the vector-skew $α$-Jensen-Shannon divergences. We study the properties of these novel divergences and show how to build parametric families of symmetric Jensen-Shannon-type divergences. Finally, we report an iterative algorithm to numerically compute the Jensen-Shannon-type centroids for a set of probability densities belonging to a mixture family: This includes the case of the Jensen-Shannon centroid of a set of categorical distributions or normalized histograms.
1 Introduction
The paper motivates extending Jensen-Shannon divergence because it symmetrizes Kullback-Leibler divergence, remains bounded, and accommodates densities with arbitrary supports. It then introduces vector-skew generalizations and applies them to construct tunable symmetric divergences and compute Jensen-Shannon centroids.
- Motivation: The Kullback-Leibler divergence is asymmetric and may diverge, motivating bounded symmetrizations such as Jeffreys and Jensen-Shannon divergence.The Jensen-Shannon divergence can be interpreted as total KL divergence to the average distribution.
- Motivation: The Jensen-Shannon divergence applies to densities with arbitrary supports and is upper bounded by log 2.It saturates at log 2 when the supports are disjoint.
- Contributions: The paper generalizes scalar skewing to a k-dimensional vector α, preserving boundedness and applicability to densities with potentially different supports.The construction also derives vector-skew Jensen-Shannon divergence from the cross-entropy-minus-entropy decomposition of KLD.
- Contributions: Vector-controlled families of symmetric Jensen-Shannon-type divergences provide additional tuning parameters beyond scalar skewing.The parameters may be selected, for example, using cross-validation techniques.
- Contributions: The paper develops a concave-convex iterative procedure for Jensen-Shannon centroids of mixture-family densities, including categorical distributions and shared-component mixtures.Experiments graphically compare Jeffreys and Jensen-Shannon centroids for grey-valued image histograms.
2 Extending the Jensen-Shannon divergence
The paper defines vector-skew Jensen-Bregman and Jensen-Shannon divergences using weighted mixtures and studies their convexity, boundedness, symmetry, and recoverability of the ordinary Jensen-Shannon divergence. It also constructs infinitely many symmetric parameterized families.
- Vector-skew Jensen-Bregman divergences: A vector-skew Jensen-Bregman divergence uses a skewing vector α, weights w, and a scalar γ to combine Bregman divergences of weighted mixtures.Choosing γ equal to the weighted average skew produces a simplified Jensen diversity expression.
- Vector-skew Jensen-Bregman divergences: Jensen diversity generalizes cluster variance for Bregman divergences and is a k-point measure rather than a two-point divergence.It arises from the gap in Jensen's inequality for a strictly convex generator.
- Vector-skew Jensen-Shannon divergences: The weighted vector-skew Jensen-Shannon divergence is defined from weighted KL divergences between mixtures and recovers ordinary Jensen-Shannon divergence for k = 2, α1 = 0, α2 = 1, and equal weights.Its entropy representation uses the weighted average skew ᾱ.
- Properties: The resulting divergences are strictly separable convex, because they are convex combinations of weighted KL divergences with the corresponding strict convexity property.The component KLα,β divergence is strictly separable convex when α ≠ β.
- Properties: The vector-skew Jensen-Shannon divergence is bounded by log 1/[ᾱ(1 − ᾱ)] and is symmetric exactly when skew-weight pairs match under complementation.A symmetric construction can also be obtained by doubling the skew-vector dimension with complementary parameters.
- Symmetric families: The framework generates infinitely many vector-skew divergences and infinitely many symmetric families by increasing the skew-vector dimension or pairing complementary skew values.The two-dimensional construction imposes equal weights and complementary skew parameters for symmetry.
3 Jensen-Shannon centroids on mixture families
The paper formulates Jensen-Shannon centroids for mixture families as strictly convex optimization problems and solves them numerically with CCCP. The method applies to categorical distributions and normalized histograms, while continuous mixture families may require stochastic approximation.
- Mixture-family formulation: Mixture families preserve closure under weighted mixing, allowing Jensen-Shannon divergence to be represented through the Shannon negentropy generator.Categorical distributions are a principal example of such mixture families.
- Centroid optimization: Jensen-Shannon centroids minimize a weighted divergence objective and are unique when they exist because the divergence is strictly separable convex.Uniform weights yield centroids, while non-uniform weights yield barycenters.
- Numerical procedure: The centroid objective is a difference-of-convex program solved iteratively by CCCP, which requires no learning-rate or step-size selection.CCCP converges linearly to a local minimum or fixed point, with the practical challenge being inversion of the gradient of B.
- Categorical distributions: For categorical distributions, the paper provides an algorithm that converts natural and mixture parameters while iterating CCCP to approximate the Jensen-Shannon centroid.The procedure initializes the centroid from the input parameters and returns the converted distribution after a prescribed number of iterations.
- Examples and scope: The method computes centroids for normalized image histograms, including inputs with non-coinciding supports, where the centroid copies and scales regions supported by only one input.The experiments compare Jensen-Shannon and Jeffreys centroids for Lena, Barbara, and an inverted Barbara histogram.
- Examples and scope: Closed-form entropy and cross-entropy expressions support mixture-family calculations, but continuous densities with shared support can make negentropy intractable because of the differential-entropy log-sum term.For Gaussian mixture models with prescribed components, the negentropy may instead be approximated using Monte Carlo techniques.
4 Conclusion and discussion
The paper extends Jensen-Shannon divergence through vector skewing while preserving boundedness, support flexibility, and unnormalized-density formula expressions. It also develops tunable divergence families and a convex-concave procedure for Jensen-Shannon centroids.
- Motivation: The Jensen-Shannon divergence is a bounded symmetrization of Kullback-Leibler divergence with applications in machine learning and deep learning.Scalar skewing has experimentally improved performance in some tasks.
- Properties: The proposed Jensen-Shannon-type divergences retain support flexibility and extend to unnormalized densities with the same formula expression.These properties are listed among the essential properties preserved by the construction.
- Parametric families: The resulting families can be tuned by vector parameters, adding multiple controls for adapting divergence behavior in applications.The paper presents this as a generalization from scalar skewing to vector skewing.
- Centroids: For densities in mixture families, Jensen-Shannon centroids can be computed using the convex-concave procedure.The paper applies this setting to categorical distributions and compares Jeffreys and Jensen-Shannon centroids for grey-valued image histograms.
- Vector-skew divergences: Vector skewing generalizes scalar-skew Jensen-Shannon divergences while preserving boundedness and applicability to densities with potentially different supports.The construction uses vectors α and β to define weighted separable divergences.
- Vector-skew divergences: The vector-skew framework unifies Jeffreys divergence and Jensen-Shannon α-skew divergence through appropriate parameter settings.The bi-vector-skew construction recovers Jeffreys divergence and the Jensen-Shannon α-skew form as special cases.
- Vector-skew divergences: Correlating the skewing vectors yields an equivalent Jensen diversity and a vector-skew generalization of Jensen-Shannon divergence.For β=(¯α,…,¯α), the mean parameter is ¯α=Σ_i w_iα_i.