Source-linked AI summary

Measuring Memory and Generalization as Separable Geometric Channels: The Topo^2 Framework

Zhanbo Zhang, Ming Liu, Qing Wang

arXiv:2608.30487v1cs.LGcs.AI

TL;DR

Noisy-label studies often use scalar measures that conflate memorization with overfitting and cannot identify or remove memorized representation content. Topo2 introduces geometric channels and causal interventions to separate these phenomena. Its FM0 prescription reaches the setting’s generalization ceiling with essentially no memory, while the framework establishes a deterministic within-channel trajectory and explicit scope boundaries.

  • Problem

    Scalar noisy-label memorization measures conflate memorization with overfitting and cannot distinguish where memorized information lives or whether it can be removed.

  • Method

    Topo2 combines a validated two-channel geometric measurement protocol with FM0 causal interventions, graded laws, and falsification tests.

  • Results

    FM0 reaches the setting’s generalization ceiling with memory approximately zero across 9/9 settings, while FM0-hat recovers 90–97% without the oracle.

  • Takeaways & Limitations

    The framework makes memorization measurable, separable, and causally manipulable in representation geometry, with a deterministic seed-conditional trajectory supporting predictability.

  • Takeaways & Limitations

    The framework does not claim independent prediction via Hcross, method competitiveness, or robustness of the break beyond single-seed decompositions.

Abstract

from arXiv · show

Deep networks trained on noisy labels simultaneously generalize on clean data and memorize flipped labels. These are usually conflated as pressures on one capacity. We present Topo^2, a measurement framework that makes them causally separable, measurable, and law-governed. Persistent-homology H1 structure of the representation space separates into a within-class manifold channel (a function of the training stopping point) and a cross-class channel (a monotone readout of memorized flipped samples). An intervention, the FM0 prescription (zero loss on flipped samples from epoch 0), reaches each setting's generalization ceiling while memorizing essentially nothing. Within the framework we establish a law set with graded evidence: (L2) FM0 separation prescription (9/9); (L1) the within-channel as a training-position function (mid-rise 6/6; convergence-back CIFAR 3/3, SVHN 2/3); (L3) a ring-construction identity (definitional, not a law); and TLS (memory-generalization topological layering): memory is causally additive, anchored (silencing clean collapses the representation), invertible (stripping memory restores near-ceiling generalization), and quantitatively billable (the memorization cost law, effective slope coefficient C ~ 0.38 at the reference capacity: CIFAR-10 0.3801 / SVHN 0.3806 / CIFAR-100 0.384 / VGG 0.3715, capacity-dependent in general and traced to clean-sample feature displacement). We also publish the framework's boundaries: a falsification ledger of nine dead ends, and an instrument-vindication section that excludes six families of global statistics as explanations of the within-channel. The framework turns "memorization" from an ill-defined capacity into a measurable, separable, invertible topological layer.

1 Introduction

Topo2 replaces scalar noisy-label memorization measures with a geometric and causal toolkit that separates representation channels. It validates the instrument, enables memory interventions, and states both graded laws and explicit boundaries.

  • Scalar noisy-label training loss conflates memorization with overfitting and cannot locate or remove memorized information in the representation.
  • The framework measures two channels, Hwithin and Hcross, with documented normalization and a dual-scope memory readout.
  • FM0 provides a causal intervention that creates a clean generalization substrate for overlay, stripping, and re-anchoring experiments.
  • The law set uses graded evidence and publishes falsified hypotheses rather than presenting only surviving findings.
  • Hwithin is analyzed as a deterministic function of the RNG sequence, while the framework explicitly limits claims about independent prediction and method competitiveness.

2 Measurement Methods

The measurement study vindicates Hwithin by excluding global-statistic explanations and calibrates its geometric semantics, while distinguishing Hcross as a local label-mixing readout. Additional controls establish robustness, self-corrections, and representation-layer alignment.

  • Instrument vindication: An FM0 pairing with val_acc differing by 0.3% and all 14 global statistics within 5.8% still shows Hwithin differing by 12.The checkpoints are eta40_FM0 (within=28) and eta50_FM0 (within=16).
  • Positive calibration: Hwithin measures local-density homogeneity along class-local interpolation paths, with homogeneous constructions high and hetero-scale constructions lowest.Values are 32.3 ± 2.6 versus 13.7 ± 2.4 across three seeds.
  • Positive calibration: Density variance strongly reduces Hwithin, whereas intrinsic dimension and anisotropy remain approximately flat.Density sensitivity changes Hwithin from 37.0 at s=0 to 3.0 at s=1.0; intrinsic-dimension and anisotropy values remain approximately 26–36.
  • Positive calibration: Hcross measures neighborhood label mixing and decreases monotonically with inter-class centroid distance, without responding to manifold-level interleaving.Cross falls from 157 at δ=0 to a floor of 34 at δ=24; interleaved crescent and orthogonal crossing are 31 and 34.
  • Robustness and controls: The Hcross slope remains positive and the Hwithin slope remains near zero across the 9-cell N_NEIGHBOR × PCA_DIM robustness core.Cross slopes range from +129 to +173, while within slopes range from −3 to +8.
  • Self-corrections: The Gaussian null and C2 self-corrections narrow interpretation: within≈30 is a pipeline artifact, and Hwithin measures homogeneity rather than topological richness.
  • Instrument vindication: The decisive exclusion chain falsifies six global-statistic families as explanations of Hwithin.
  • Third-party cross-checks: Neural-collapse metrics can organize monotonically while Hwithin stays flat, supporting a local rather than mean-level interpretation of the readout.

2.2 The 𝐻1 pipeline

The pipeline computes separate H1 quantities for within-class and cross-class complexes and requires fixed normalization for cross comparisons. Absolute cross values carry sampling noise, while the neighborhood and PCA settings are reported as robust.

  • Hwithin is H1 of same-class submanifolds, whereas Hcross is H1 of the cross-class complex.
  • The pipeline documents normalization obligations before comparing cross values.
  • Cross is normalized by N_PAIRS × N_INTERP, so comparisons must fix and report the pipeline.Varying both settings fourfold leaves cross_fraction at 0.86–0.87.
  • Absolute cross values have ±15–30% RNG sampling noise, making contrasts below 30% require error bars.Large FM0↔CONT↔full-overlay effects of at least 2× are unaffected.
  • N_NEIGHBOR and PCA_DIM are reported as robust pipeline settings.

2.3 Memory readout: the dual-scope mem_noisy

The dual-scope mem_noisy readout measures whether flipped training samples receive their noisy labels, paired with clean-sample agreement to distinguish memorization from clean performance. FM0 suppresses noisy-label memorization while preserving ceiling-level generalization, enabling causal add, remove, and re-anchor interventions.

  • Readout definition: mem_noisy is the fraction of flipped training samples predicted at their noisy labels and serves as the decisive memorization readout.Its dual-scope pairing uses mem_clean, the fraction of clean training samples predicted at their true labels.
  • Readout behavior: CONT reaches approximately 0.999 mem_noisy, FM0 approximately 0.004, and the K-half overlay approximately 0.50.These values establish the readout's separation across memorization settings.
  • FM0 intervention: Zero loss on flipped samples from epoch 0 reaches the setting's generalization ceiling while producing approximately zero memory.Reported ceilings are SVHN approximately 0.96 and CIFAR-10 approximately 0.91.
  • FM0 intervention: The FM0-hat data-driven mask recovers 90–97% without the noise distribution, making the deployable intervention distinct from the FM0 oracle.FM0-hat provides the substrate for the subsequent causal operations.
  • Causal interventions: Overlay adds memory by unfreezing K flipped samples, while strip removes memory during clean-only continuation and re-anchor tests memory without clean signal.Overlay interpolates FM0 to CONT as K increases; strip drains mem from 0.999 to 0.11, while re-anchoring collapses validation accuracy from 0.92 to 0.04.

2.6 Determinism and the eval cadence (a methodological core asset)

The training trajectory and Hwithin are deterministic functions of the RNG sequence, so intermediate evaluation cadence becomes a training variable. Hwithin varies substantially while validation accuracy changes little, making cadence disclosure essential for reproducibility and geometric analysis.

  • Cadence dependence: Hwithin takes distinct values under identical seeds, protocols, masks, and training code when only intermediate-evaluation cadence changes.Full-test 300-epoch values are 16, 26, and 25 for cadences 25, none, and 100; validation accuracy changes by only approximately 0.01.
  • Cadence dependence: Hwithin takes values 26 / 16 / 16 / 25 / 22 across cadences 0 / 25 / 50 / 100 / 200, while val_acc moves by only approximately 0.01.The geometric coordinate spans 16–26, a 38% spread, versus a 2% scalar spread.
  • Causal control: The waste control reproduces the evaluation condition exactly, with within equal and validation accuracy bit-identical to 1e-4.This pins the causal channel to the RNG footprint rather than evaluation itself.
  • Deterministic trajectory: The training trajectory is deterministic rather than chaotic: same-cadence repetitions are bit-identical, and 12 independent trainings show zero deviation.The waste control reproduces the full state, including weights, validation accuracy, and within.
  • Methodological rule: Because intermediate checkpointing with evaluation is itself a training variable, same-seed reproducibility claims must declare evaluation and save cadence.Comparisons mixing checkpoints from different cadences require qualification.

3 The Law Set

The law set separates training-position effects in H1_within from memorized-label effects in H1_cross, using FM0 and interventions to establish graded causal and quantitative structure.

  • FM0 separation prescription: The FM0 prescription reaches the setting’s generalization ceiling with approximately zero memory, with zero exceptions across 9/9 settings.Reported ceilings are SVHN 0.96+ and CIFAR 0.91+; FM0-hat recovers 90–97% without the oracle.
  • Ring identity: The ring identity within = single-pair rings + cross-pair same-class rings holds with zero residual across 6/6 checkpoints, but is definitional rather than a law.The proposed cross-pair isolation mechanism fails the multi-seed test: the decomposition reverses or mixes across seeds.
  • TLS: Memory is causally additive: overlaying flipped samples increases H1_cross monotonically across all 3 seeds and reaches CONT-level cross at saturation.The reported sequences are 90→221, 91→254, and 111→219.
  • TLS: Memory anchors to the clean representation and is invertible: silencing clean collapses validation accuracy, whereas stripping memory restores near-ceiling accuracy.CIFAR validation accuracy falls 0.92→0.04 when clean signal is silenced; stripping reduces memory 0.999→0.11 and restores accuracy 0.630→0.919.
  • TLS: The memorization cost law fits an effective coefficient near 0.38 at reference capacity, but C is capacity-dependent and traces to clean-sample displacement.Values are CIFAR-10 0.3801, SVHN 0.3806, CIFAR-100 0.384, and VGG 0.3715; a width sweep gives 0.471→0.342.

4 Empirical Panorama (key tables)

At the shared ∼11M reference capacity, the memorization cost coefficient is near 0.38 across four settings, while remaining capacity-dependent in general.

  • The memorization cost law uses mem and noise rate η, with C reported as an effective slope coefficient.Throughout, mem denotes mem_noisy and η denotes the noise rate.

5 What the Framework Does NOT Claim

The framework narrows its claims through falsification: several proposed mechanisms and predictors do not survive controlled tests, while validated claims retain explicit scope and evidence limits.

  • Hcross1 is a geometric signature of memory, not an independent generalization predictor.Partial correlations with memorization are weak and setting-dependent, including SVHN +0.08 with ΔR2 +0.000.
  • The reported evidence is uneven: several cells are single-seed, and VGG η50 has 7.7% cross-seed coefficient of variation.The paper states statistical strength per number and distinguishes stronger from weaker trajectory evidence.
  • The ledger records nine dead ends and treats intervention, matched-load comparisons, corrected readouts, and null baselines as necessary safeguards.Evidence is graded across cross-sectional, trajectory, and causal findings, with some results explicitly removed after multi-seed testing.
  • The controlled-overlay advantage over CONT was downgraded because oracle masking, memory amount, and co-training path were confounded.The supported claim is the tradeoff curve, not a single crossing-point comparison.
  • The cost coefficient remains capacity-dependent, while fixed-capacity class-count comparisons do not materially shift it.CIFAR-10 and CIFAR-100 give 0.3801 and 0.384, while superclass retraining gives 0.370 versus 0.369.
  • A corrected null claim replaces the dead within≈30 invariant: Gaussian point clouds return within=30.3 ± 8.7, matching the pipeline artifact.The framework therefore treats zero deviation from the random baseline—not the value 30—as the invariant.
  • All six global-statistic explanations were falsified: val_acc differed by 0.3%, statistics by ≤5.8%, yet within differed by 12.The decisive FM0 pairing supports instrument vindication by exclusion.

6 Related Work

Topo2 complements existing label-noise, memorization-theory, topological-data-analysis, and generalization-theory work by combining geometric measurement with causal manipulation and cost accounting.

  • Relative to method-oriented label-noise learning, FM0 supplies a mechanistic ceiling and common measurement language.The comparison names DivideMix, ELR, and Co-teaching as related methods.
  • Memorization theory is complementary: prior work addresses why, when, whether, and which samples are memorized, while Topo2 addresses where and cost.
  • Topo2 extends persistent-homology work from description toward causal manipulation through overlay and strip operators.Its measurement protocol is presented as law-governed rather than descriptive alone.
  • Generalization-theory capacity bounds are characterized as orthogonal to a billable memorization cost.

7 Discussion and Outlook

The framework converts memorization into a causally testable geometric quantity and supports deterministic trajectory analysis, while leaving mechanism, breadth, and low-noise behavior open.

  • Two channels, mem_noisy, FM0, and overlay/strip operators provide a common language for causal measurement of memorization.
  • Hwithin1 is a deterministic function of the RNG sequence, making trajectory predictability a necessary condition for elevating observation to law.
  • Open questions include the residual image term in C, co-evolution advantage, broader architectures and datasets, within-channel dynamics, and the unresolved chaos region.
  • The paper treats falsified hypotheses as boundary markers for the confirmed law set.

8 Conclusion

Topo2 makes memorization and generalization measurable, separable, and causally manipulable in representation geometry. It combines validated measurement, FM0 intervention, graded laws, falsification tests, and determinism analysis.

  • Topo2 makes memorization and generalization measurable, separable, and causally manipulable in representation geometry.
  • The framework includes an instrument vindicated against six global-statistic alternatives and two self-corrections, alongside a validated measurement protocol.
  • FM0 cleanly separates the memorization and generalization channels.
  • Topo2 establishes a law set with graded evidence and maintains a falsification ledger of nine dead ends.
  • A determinism analysis makes the trajectory law deterministic.
Loading 2608.30487v1…