Source-linked AI summary
Persistent Cross Entropy
Sijin Yeom, Jae-Hun Jung
TL;DR
Persistent cross entropy addresses the lack of a natural cross-entropy comparison for persistence diagrams with different event spaces. It constructs an induced probability using similarity and persistence weighting, then evaluates PCE through theoretical stability results and three applications. The studies show that PCE distinguishes equal-entropy diagrams, separates causal directions without a joint diagram, and supports directional topology loss for knowledge distillation.
Problem
Shannon cross entropy cannot directly compare persistence-based probability measures defined on persistence diagrams with different event spaces.
Method
The paper combines a similarity function with persistence weighting, assigns unexplained mass to an added event, and defines persistent cross entropy on the resulting common space.
Results
PCE distinguishes diagrams with the same persistent entropy, separates causal directions without constructing a joint persistence diagram, and supports directional topology loss for knowledge distillation.
Takeaways & Limitations
PCE provides a directional comparison of persistence diagrams that preserves unexplained structure and extends entropy-based analysis beyond a single scalar summary.
Takeaways & Limitations
Preliminary tests on nonlinear and chaotic systems produced mixed results, so further study is needed to determine when PCE works well and how it can be improved.
Abstract
from arXiv · showhide
Persistent entropy is the Shannon entropy of a persistence-based probability measure defined on a persistence diagram. However, its cross-entropy version is not naturally defined because two persistence diagrams generally have different event spaces. To bridge these event spaces, we combine a similarity function with persistence weighting to define an induced probability. The induced probability reflects information from one diagram on the event space of the other diagram and assigns unexplained probability mass to the unexplained event. Using the induced probability, we extend cross entropy to persistence diagrams, called persistent cross entropy (PCE). We establish the main properties of both the induced probability and PCE and prove stability theorems for both. Through three numerical studies, we show that PCE distinguishes diagrams with the same persistent entropy, separates causal directions in dynamical systems without constructing a joint persistent diagram, and can be used as a directional topology loss for knowledge distillation.
1 Introduction
Persistence diagrams are difficult to compare with Shannon cross entropy because their probability measures live on different event spaces. The paper bridges those spaces with an induced probability and develops persistent cross entropy with stability results and applications.
- Persistence diagrams encode multiscale topological features but have unordered points and variable cardinality, motivating functional, image-based, and vector representations.
- Persistent entropy summarizes persistence-based probabilities with one Shannon-entropy scalar, so diagrams with different feature compositions can share the same entropy.
- Shannon cross entropy cannot directly compare pX and pY because they are defined on different persistence diagrams and therefore different event spaces.
- The induced probability places retained mass on DX according to how similarly DY explains each point and assigns the remainder to an unexplained event.
- The paper defines persistent cross entropy using this common event space and establishes properties and stability results for both the induced probability and PCE.
2 Background and Setup
Persistent homology represents features by birth–death pairs, while persistence probabilities weight their lifetimes. Persistent entropy then measures how concentrated or evenly distributed those weights are.
- Persistent homology builds a filtration of simplicial complexes and records each feature by its birth and death scales in a persistence diagram.
- Birth–persistence coordinates v = (b_v, ℓ_v) preserve the death coordinate through d = b + ℓ and map the diagonal to ℓ = 0.
- The finite off-diagonal multiset is used for probability construction, while the diagonal is adjoined only for Wasserstein or bottleneck matching.
- Persistence probabilities assign more mass to long-lived features and less to short-lived features, with repeated points treated as separate probability atoms.
- Persistent entropy is the Shannon entropy of the persistence probability, becoming small when persistence is concentrated and reaching log n for n equal weights.
3 Induced Probabilities
The induced probability transfers information from DY to DX by combining similarity and persistence weighting, while assigning unmatched mass to an unexplained event. It is normalized, direction-sensitive, and stable under diagram perturbations.
- Different event spaces prevent directly applying Shannon cross entropy to pX and pY, so the construction induces a probability on DX ∪{∂}.
- Each point’s explanatory score combines similarity with persistence weighting, and a fixed response scale converts explanatory differences into weights.Coordinate scales and the response scale remain fixed across comparisons rather than being recalibrated for each pair.
- The unexplained event ∂ receives the residual mass because explained masses over DX generally sum to less than one.The residual is obtained by summing the unexplained amount from each point of DX.
- Comparing DX with itself recovers pX, and the induced probability equals pX exactly when explanatory-power differences vanish on DX.
- The total variation distance between pX and the induced probability equals the unexplained mass.
- The unexplained mass vanishes as DY approaches DX, while the induced probability varies stably with DY under the 1-Wasserstein distance.A bound controls changes in the induced probability for fixed DX.
4 Persistent Cross Entropy
Persistent cross entropy extends Shannon cross entropy to persistence diagrams by using the induced probability on a shared event space. Its entropy excess is directional and locally stable as one or both diagrams vary.
- Persistent cross entropy is defined by applying Shannon cross entropy to pX and the probability induced by DY on DX ∪{∂}.
- The entropy excess measures how differently DX and DY explain DX under the selected similarity function and persistence weighting.
- PCE is directional because reversing the comparison changes the reference diagram, event space, and induced probability.
- The entropy excess vanishes exactly when the induced probability equals pX and vanishes as DY approaches DX in 1-Wasserstein distance.
- Entropy excess is locally stable when both diagrams vary, with changes bounded by the sum of their 1-Wasserstein perturbations.When DX is fixed, the bound reduces to dependence on the perturbation of DY alone.
5 Experiments
Three experiments show that PCE distinguishes diagrams with equal persistent entropy, identifies causal direction from separate diagrams, and provides a directional topology loss for knowledge distillation.
- 5.1 Two loops with equal persistent entropy: PCE separates four diagrams whose H1 persistent entropies all round to 1.500, with values increasing from 1.752 to 3.977.The corresponding unexplained masses are 0.194, 0.498, 0.751, and 0.770.
- 5.1 Two loops with equal persistent entropy: The unexplained event preserves the small-loop feature's strongest explanation while leaving 0.751 of the total mass unexplained in the Y3 comparison.Retained probability is not renormalized within DX, avoiding artificial inflation of the small-loop point.
- 5.2 Spring–mass benchmark: PCE directly produces directed comparisons from DA and DB without constructing the joint diagram DAB.The entropy excess measures the cost of using one diagram to explain the other, while the reverse comparison uses the opposite direction.
- 5.2 Spring–mass benchmark: All 16 one-way spring–mass cases fall in the expected directional half-plane, separating A →B from B →A.Under A →B, DB explains DA at lower cost; the interpretation reverses for B →A.
- 5.2 Spring–mass benchmark: Unexplained mass reaches R2 = 0.993 for balanced one-way separation, while persistence landscapes reach 0.876 and entropy excess reaches 0.740.Across four regimes, entropy excess gives the strongest separation, followed by unexplained mass; these PCE quantities avoid DAB.
- 5.3 Topology-aware knowledge distillation: EM-PCE gives the highest final accuracy and last-10-epoch mean in the seed-7 CIFAR-100 experiment, improving over TopKD by 0.50 and 0.606 percentage points.During PCE training, epoch-averaged unexplained mass decreases from 0.2494 in the first epoch to 0.00756 in the final epoch, treated as a topology-matching diagnostic.
A Proofs for Probabilities Induced on the Augmented Event Space
The stability framework assumes a bounded, monotonically increasing, globally Lipschitz weighting function g, with g(0)=0 ensuring diagonal matches contribute nothing.
- Regularity assumptions: g is assumed monotonically increasing and globally Lipschitz with finite bound Mg and Lipschitz constant Lg.These assumptions make stability constants independent of the diagrams and their maximum persistence.
- Regularity assumptions: The condition g(0) = 0 makes a point contribute zero when it is matched to the diagonal.
- Stability consequence: The assumptions support uniform stability estimates whose constants do not depend on diagram-specific maximum persistence.The construction combines g's regularity with the constant L0 from Definition 1.
- Experimental choice: The experimental function is g(t) = t/(1 + t), and it satisfies the stated assumptions.
A.2 Maximum-entropy derivation of the Gaussian similarity function
The appendix derives the Gaussian similarity and formalizes the augmented event-space probability construction, including repeated occurrences, unexplained mass, and stability prerequisites.
- Gaussian similarity: The Gaussian similarity is normalized by its maximum, removing the dimensional prefactor and yielding a coordinate-rescaled Euclidean construction.It has the form κ(u,v) = h(ρ(u,v)) with h(r) = e^(-r^2/2).
- Gaussian similarity: The function h satisfies h(0) = 1, decreases strictly for positive distances, and approaches zero as distance grows.
- Event-space construction: Persistence diagrams are represented through occurrence sets so repeated points remain distinct atoms while projection recovers their coordinates.
- Event-space construction: The augmented finite event space contains diagram occurrences and an unexplained event, with induced masses extending to probability measures.The unexplained-event mass is explicitly included rather than absorbed into the diagram-point probabilities.
- Probability properties: The construction assigns nonnegative occurrence-level masses summing to one and preserves positivity for every off-diagonal point.
- Diagram distance: The 1-Wasserstein distance matches occurrences, includes unmatched points through diagonal matches, and uses birth–persistence coordinates.
- Probability properties: The unexplained mass vanishes precisely when every point of the reference diagram is fully explained by the comparison diagram.
- Stability proof: Stability bounds are obtained by comparing matched points under admissible diagram matchings and then taking the infimum over matchings.The estimates use boundedness and Lipschitz properties of g and the similarity function.
A.7 Proof of Theorem 1
Theorem 1 bounds changes in the induced probability when persistence diagrams vary, using Wasserstein distance and total variation on a shared augmented event space.
- Stability bound: Theorem 1 controls changes in unexplained-event mass and induced probabilities through a Wasserstein-distance bound.The proof first establishes uniform pointwise bounds and then takes the supremum over diagram points.
- Proof mechanism: The proof applies the mean value theorem to control the change in the induced masses before summing pointwise bounds.
- Total variation: The total variation argument uses one half of the ℓ1 difference between probability measures defined on the same finite event space.
A.8 Proof of Proposition 2
Proposition 2 rewrites the KL divergence between the original and induced probabilities as persistent cross entropy minus persistent entropy, with a nonnegative squared-difference expression.
- Divergence identity: KL(pX∥pY_X) equals H(pX, pY_X) − H(pX), because the original measure assigns zero mass to the unexplained event.
- Nonnegativity: The final quantity is nonnegative because it is a positive constant times an expectation of a squared difference.Suppressing occurrence labels recovers the expectation over u ∼ pX.
A.9 Proof of Theorem 2
The proof establishes the theorem using constants fixed independently of the input diagrams, yielding a uniform constant across every diagram pair.
- The theorem holds with constants Lϕ and τ fixed independently of DX and DY.
A.10 Proof of Theorem 3
The proof develops bounds for changes in persistence-weighted quantities and similarity terms under diagram perturbations, then combines them into joint and one-diagram stability estimates.
- The proof compares matched diagram points, including diagonal matches, while assigning diagonal points zero persistence.This lets the matching argument cover both off-diagonal and diagonal cases.
- Changes in normalized persistence weights are bounded by analyzing total persistence and matched-point differences.The proof separately handles pairs with both points off the diagonal and pairs involving the diagonal.
- The proof constructs a finite uniform bound from the center diagrams, fixed functions, and perturbation scales.The resulting constant does not depend on DX′ or DY′ within the prescribed neighborhood.
- Substitution of the bounds proves the joint local stability claim for perturbations of both diagrams.
- When DX′ = DX, the joint estimate reduces to the final one-diagram stability claim.
B Additional Experimental Details
The experiments fix probability, similarity, persistence, dynamical-system, baseline, and training configurations explicitly, while verifying identities of the augmented probability construction.
- The induced probability combines similarity-weighted persistence mass on DX with remaining mass assigned to the unexplained event ∂.The reference probability remains pX(u) = ℓu/P, and no further normalization is required.
- The numerical construction satisfies unit total augmented mass, boundary-mass total variation, and the stated entropy-excess identity up to floating-point precision.
- The causal-direction experiment evaluates PCE directed quantities against symmetric baselines across the complete 9 × 9 coupling grid using fixed parameters.The PCE panels compare directed quantities from DA and DB, whereas symmetric baselines compare DAB separately with each individual diagram.
- The dynamical-systems study uses delay reconstructions with E = 4 and q = 1, retaining 997 common points across individual and joint clouds.The joint reconstruction uses (xt, xt+1, yt, yt+1), while PCE is computed directly from DA and DB without DAB.
- The knowledge-distillation study uses CIFAR-100, a pretrained ResNet56 teacher, a ResNet20 student, and fixed shared training settings.Teacher and student features are normalized, exact finite H0 Vietoris–Rips diagrams are computed, and the teacher diagram is detached.
- Validation selected PCE over TopKD, with best accuracies of 70.64% and 69.62%, respectively, before the configuration was fixed for full training.