Source-linked AI summary
A Characterization of Entropy in Terms of Information Loss
John C. Baez, Tobias Fritz, Tom Leinster
TL;DR
The paper asks how information loss associated with measure-preserving functions can be characterized without focusing directly on individual probability measures. It imposes functoriality, convex linearity, and continuity, and concludes that information loss is a constant multiple of Shannon entropy difference, with an analogous Tsallis-entropy extension.
Problem
The paper addresses characterizations of entropy by shifting the focus from individual probability measures to information loss under measure-preserving functions.
Method
The paper models measure-preserving functions as processes and characterizes their information loss using functoriality, convex linearity, and continuity.
Results
Information loss is c(H(p) − H(q)) for Shannon entropy, while degree-α homogeneity yields a quantity proportional to Hα(p) − Hα(q) for Tsallis entropy.
Takeaways & Limitations
Shannon and Tsallis entropy differences provide characterized forms of information loss for measure-preserving functions under the stated axioms.
Abstract
from arXiv · showhide
There are numerous characterizations of Shannon entropy and Tsallis entropy as measures of information obeying certain properties. Using work by Faddeev and Furuichi, we derive a very simple characterization. Instead of focusing on the entropy of a probability measure on a finite set, this characterization focuses on the `information loss', or change in entropy, associated with a measure-preserving function. Information loss is a special case of conditional entropy: namely, it is the entropy of a random variable conditioned on some function of that variable. We show that Shannon entropy gives the only concept of information loss that is functorial, convex-linear and continuous. This characterization naturally generalizes to Tsallis entropy as well.
1. Introduction
The paper characterizes information loss associated with measure-preserving functions rather than entropy of individual probability measures. Under functoriality, convex linearity, and continuity, information loss is determined by Shannon entropy differences, with analogous results for general measures and Tsallis entropy.
- Novel perspective: The characterization focuses on entropy change under measure-preserving functions, recovering ordinary entropy from the unique map onto a one-point space.This shifts attention from a single probability measure to information loss associated with a process.
- Axioms: Information loss is required to be functorial, convex-linear, and continuous, representing additive loss across stages, probabilistic mixing, and small changes in processes.Functoriality gives F(f ◦g) = F(f) + F(g), while continuity requires F(f) to change only slightly when f changes slightly.
- Main result: Under these assumptions, there exists c ≥ 0 such that information loss for any measure-preserving function equals c(H(p) − H(q)).The conclusion identifies Shannon entropy differences as the unique form of information loss up to a multiplicative constant.
- General measures: For general finite measures, convex linearity is replaced by additivity and homogeneity, while the same information-loss formula remains valid up to scale.Any finite measure is written as λp, with H(λp) = λH(p).
- Tsallis extension: Replacing degree-one homogeneity by degree-α homogeneity yields information loss proportional to Hα(p) − Hα(q), where Hα is Tsallis entropy.The Tsallis characterization changes the homogeneity degree from 1 to α while retaining the same overall structure.
2. The main result
The paper characterizes information loss for measure-preserving morphisms between finite measured sets. Under functoriality, convex linearity, and continuity, the loss is exactly a nonnegative multiple of the decrease in Shannon entropy, with an analogous result for general finite measures.
- Definitions: A measure-preserving morphism in FinProb maps finite probability spaces while preserving singleton masses, and information loss is assigned to each such morphism.Continuity is defined through convergent sequences of morphisms with eventually fixed underlying sets and functions and pointwise-convergent measures.
- The main characterization: Theorem 2 states that any nonnegative information-loss map satisfying functoriality, convex linearity, and continuity has the form F(f)=c(H(p)−H(q)) for c≥0.Conversely, every nonnegative constant c produces a map satisfying the three axioms.
- General measures: For finite measured sets beyond probability measures, convex linearity is replaced by additivity and homogeneity over direct sums and nonnegative scalar multiplication.A finite measure is normalized to a probability measure when its total mass is nonzero, and its entropy is total mass times the normalized entropy.
- General measures: The generalized result again gives F(f)=c(H(p)−H(q)) for every measure-preserving morphism in FinMeas.The proof reduces nonzero finite measures to probability measures using homogeneity and handles zero-mass morphisms separately.
3. Why Shannon entropy works
The entropy difference associated with a measure-preserving function equals conditional entropy. This interpretation makes the functoriality of Shannon information loss transparent, while continuity and the remaining axioms follow directly or through the entropy construction.
- Axioms: For Shannon entropy, continuity is immediate, while additivity and homogeneity in the general-measure setting follow from the corresponding probability-measure result.The paper notes that the proof for FinMeas can also be checked directly.
- Conditional-entropy interpretation: For random variables x and y=f(x), the information loss F(f)=H(p)−H(q) is exactly the conditional entropy of x given y.Thus information loss is a special case of conditional entropy for deterministic processing.
- Functoriality: The conditional-entropy formulation makes the composition rule easy to verify by applying the entropy-loss identity to both sides.The paper identifies this composition rule with functoriality of information loss.
4. Faddeev’s theorem
Faddeev’s theorem provides the entropy characterization needed for the main proof. Its invariance, continuity, normalization, and grouping condition force any admissible functional to be a nonnegative multiple of Shannon entropy.
- The theorem: Faddeev’s theorem says that a nonnegative functional on finite probability measures satisfying invariance, continuity, normalization, and the grouping condition is a multiple of Shannon entropy.Conversely, every nonnegative multiple of Shannon entropy satisfies these conditions.
- The grouping rule: The grouping rule recursively decomposes a probability mass into subdistributions and adds the parent contribution weighted by that mass.The paper identifies this rule with recursivity and relates it to strong additivity.
- Proof idea: The theorem’s logarithmic form emerges by comparing values on uniform distributions and solving the resulting functional equations.Continuity extends the characterization from rational probability measures to arbitrary probability measures.
5. Proof of the main result
The proof recovers entropy as the information loss of the unique morphism collapsing a probability space to a point, then applies Faddeev’s theorem. Functoriality and convex linearity supply the conditions needed for that theorem.
- Recovering entropy: The unique morphism from a probability measure p to the one-point space is interpreted as collapsing p and losing all its information.The paper defines the entropy of p through the information loss of this collapse morphism.
- Applying Faddeev’s theorem: Functoriality makes information loss invariant under invertible morphisms and gives zero loss for the collapse of the one-point space.These facts provide Faddeev’s invariance and normalization conditions.
- Conclusion: Faddeev’s theorem then implies that the entropy functional is a constant multiple of Shannon entropy, completing the characterization.This establishes the hard direction of the main theorem.
- Applying Faddeev’s theorem: The proof derives Faddeev’s grouping condition by comparing information loss for a composite construction with its convex decomposition.Convex linearity and the definition of entropy supply the two expressions that are compared.
6. A characterization of Tsallis entropy
The paper characterizes Tsallis entropy through information loss on measure-preserving morphisms. The characterization matches the Shannon case except that homogeneity changes from degree 1 to degree α.
- 6. A characterization of Tsallis entropy: The Tsallis characterization is identical to the Shannon characterization except that convex linearity is replaced by homogeneity of degree α.For α = 1, the limiting entropy agrees with Shannon entropy.
- 6. A characterization of Tsallis entropy: For probability measures, composable morphisms satisfy information-loss composition, compatibility with convex combinations, and continuity.Under these axioms, the information loss is determined up to a nonnegative constant by Tsallis entropy.
- 6. A characterization of Tsallis entropy: The proof replaces Faddeev’s theorem with Furuichi’s theorem, whose modified condition yields the Tsallis characterization.The resulting proof follows the Shannon argument with this theorem substituted.
- 6. A characterization of Tsallis entropy: For arbitrary finite measures, the corresponding extension changes the homogeneity degree from 1 to α while retaining additivity.The paper defines Tsallis entropies for arbitrary measures to obtain this extension.
- 6. A characterization of Tsallis entropy: Tsallis entropy is characterized by information loss satisfying functoriality, additivity, degree-α homogeneity, and continuity.The resulting map is a nonnegative function on morphisms in FinMeas.