Source-linked AI summary

A Complete Characterization of Tensorizable $f$-divergences

Rodrigo Cruz, Flavio P. Calmon, Qian Yu

arXiv:2608.28556v1cs.ITmath.PRmath.ST

TL;DR

The paper addresses the lack of a complete characterization of tensorizable f-divergences and refines the formalism used to define tensorization. It shows that admissible tensorization formulas are multi-affine and identifies all tensorizable f-divergences under this notion.

  • Problem

    Tensorizable f-divergences are widely used but had not yet been formally characterized.

  • Method

    The paper refines tensorization and analyzes the induced functional equations, including multi-affine rules and local characteristics of symmetrized divergences.

  • Results

    All tensorizable f-divergences are nonnegative linear combinations of KL and reverse KL divergences or power divergences.

  • Takeaways & Limitations

    Tensorization itself provides a direct characterization of the f-divergences admitting compositional formulas under product measures.

Abstract

from arXiv · show

Csiszar's formulation of the $f$-divergence introduced a vast family of functionals for quantifying dissimilarity between probability distributions. However, many applications in statistics and information theory rely only on a few $f$-divergences, such as the Kullback-Leibler divergence, the $χ^2$-divergence, and the squared Hellinger distance. These divergences are especially useful because they admit simple compositional formulas under product measures, a property sometimes referred to as tensorization. In this work, we refine a formalism of tensorization previously introduced in the literature. Then, we show that any possible tensorization formula has a multi-affine form characterized by a single parameter, and identify all tensorizable $f$-divergences under our adopted notion of tensorization.

I. INTRODUCTION

Tensorizable f-divergences provide compositional discrepancy formulas for product measures, supporting tractable analysis in information theory, statistics, privacy, learning, and optimization.

  • Tensorizable f-divergences decompose discrepancies between product measures into expressions depending only on marginal discrepancies.
  • Their compositional formulas remain tractable for measures with many factors.
  • Tensorizable divergences support estimation of difficult discrepancies in hypothesis testing and cumulative privacy accounting.
  • They also underpin information-theoretic bounds in online learning and stochastic optimization.

A. Contributions and Methods

The paper addresses the missing formal characterization of tensorizable f-divergences and identifies both their admissible tensorization rules and the divergences that realize them.

  • The paper fills the gap that tensorizable f-divergences had not been formally characterized.
  • The only tensorizable f-divergences are nonnegative linear combinations of KL and reverse KL divergences or power divergences.
  • Every admissible tensorization formula belongs to a one-parameter family of multi-affine functions.
  • For symmetrized tensorizable divergences, the tensorization parameter and local characteristic provide two scalar identifying parameters.

B. Related Work

The work extends prior characterizations based on compositional behavior, while focusing directly on tensorization within the class of f-divergences.

  • Unlike prior work restricted to polynomial tensorization rules, this paper identifies every admissible rule without assumptions on its formula.
  • The analysis overlaps with axiomatic information-measure research because tensorization can be distilled into a functional equation.
  • The paper argues that tensorization alone uniquely identifies Tsallis α-relative entropy in the relevant regime.
  • Its setting studies a broader composition notion while restricting attention to the smaller class of f-divergences.

II. PRELIMINARIES

The preliminaries define f-divergences and tensorization over measurable spaces, establish their finite-value range, and state foundational structural properties.

  • An f-divergence is defined from a convex function f with f(1)=0 and Radon–Nikodym derivatives relative to a common dominating measure.
  • The finite-value range is R(D_f)={D_f(P∥Q):D_f(P∥Q)<∞}.
  • If f(0) and f′(∞) are finite, the range is [0,f(0)+f′(∞)]; otherwise it is [0,∞).
  • Tensorization requires a map τ_f that combines divergences of two factor pairs into the divergence of their product pair.
  • The analysis applies to arbitrary measurable spaces, and tensorization bases extend to n≥2 factors inductively.

3) Marginalization:

The range of every f-divergence includes zero, corresponding to identical distributions.

  • 0 ∈ R(D_f) for every f-divergence.

III. CHARACTERIZATION OF TENSORIZATION BASES

The paper represents f-divergences through log-likelihood-ratio distributions, where products become convolutions and tensorization forces affine dependence on marginal divergence values. Consequently, every admissible tensorization base has a one-parameter multi-affine form.

  • Df(P∥Q) is determined by the log-likelihood-ratio distribution of Q against P.
  • Under product measures, marginal log-likelihood ratios add, so the product pair corresponds to convolution of marginal LLR distributions.
  • Tensorization makes Df(µ1 ∗ µ2) depend only on the two scalar divergences Df(µ1) and Df(µ2), upgrading mixture affinity in distributions to affinity in divergence values.
  • Every section of the tensorization base is affine in each argument and symmetric under exchanging the arguments.
  • τ_f(d1,d2) = d1 + d2 + γ d1d2, where γ is the single tensorization parameter.

IV. SYMMETRIZED TENSORIZABLE f-DIVERGENCES

The symmetrized problem reduces tensorization to a functional equation on two-point symmetric LLR distributions. Its solutions are controlled by the tensorization parameter and a local characteristic, with uniqueness and an admissibility constraint.

  • The symmetrized divergence restricted to two-point symmetric LLR distributions becomes a real-valued function g satisfying a tensorization functional equation.
  • The local characteristic is the finite nonnegative quadratic rate at which g(r) vanishes as r approaches zero.
  • For fixed tensorization parameter γ and local characteristic ℓ, at most one continuous g solves the functional equation.
  • The restriction g uniquely determines the symmetrized f-divergence, so matching γ and ℓ yields uniqueness of the symmetrized divergence.
  • A symmetrized tensorizable f-divergence must satisfy γℓ ≥ −1/8; the excluded regime produces an oscillatory, nondecreasing-incompatible solution.

V. CHARACTERIZATION OF TENSORIZABLE f-DIVERGENCES

The final characterization solves the functional equation and imposes convexity. Every tensorizable f-divergence is therefore a nonnegative combination of KL and reverse KL divergences or a power divergence.

  • The proof evaluates tensorization on the convolution of an arbitrary two-point LLR distribution and a symmetric one, producing a linear functional equation for f.
  • The relevant solution space has dimension two, allowing the complete family of continuous solutions to be identified from kernel elements.
  • Convexity reduces all solutions to nonnegative combinations of KL and reverse KL divergences or to power divergences.
  • When ℓ > 0 and γ = 0, f(x) = A x log x − B log x + E(x − 1), with A,B ≥ 0 and A + B = 2ℓ.

VI. CONCLUSION

The work characterizes tensorizable f-divergences under its adopted definition and identifies the resulting divergence families. It also establishes scope boundaries and directions for extending the framework.

  • The characterization shows that tensorizable f-divergences are nonnegative combinations of KL and reverse KL divergences or power divergences.
  • The analysis connects the adopted tensorization formalism to classical functional-equation methods rather than defining tensorization directly through additivity.
  • The class of tensorizable f-divergences is not closed under addition, because the simplification rule for Df+g can depend on all marginal discrepancies.
  • Linear combinations of tensorizable f-divergences remain a proposed direction for studying broader compositional classes.
  • The measure-theoretic appendix establishes invariance of the log-likelihood-ratio distribution to the dominating measure and closure of realizable LLR distributions under mixtures and convolution.
  • The result applies to subclasses of measurable spaces closed under products because the proofs require tensorization for finite-support discrete distributions.

A. Existence of the Local Characteristic

The appendix establishes convergence of a normalized function along dyadic halving steps and uses this local behavior to identify solutions globally. The resulting characterization reduces the solution space to a small set of forms.

  • A. Existence of the Local Characteristic: The proof studies h(r)=g(r)/r^2 along dyadic halving steps, deriving recursive relations that yield convergence as r→0+.
  • A. Existence of the Local Characteristic: A uniform bound near the origin and continuity extend the common limit from a dense subset to an interval.
  • A. Existence of the Local Characteristic: Two continuous solutions with the same tensorization parameter and local characteristic agree near zero, and the recursive identity propagates equality to all positive arguments.
  • A. Existence of the Local Characteristic: The kernel analysis bounds its dimension by two and shows that values on exponential lattices are determined by two initial values.
  • A. Existence of the Local Characteristic: The resulting f-functions include Ax log x−B log x+E(x−1), with A+B=2ℓ and convexity requiring A,B≥0.
  • A. Existence of the Local Characteristic: The appendix verifies that the derived functions induce tensorizable f-divergences using logarithm and exponent identities, integral linearity, and Tonelli–Fubini.
Loading 2608.28556v1…