Source-linked AI summary
Dataset Complexity Shapes Finite-Distance Loss Geometry in Neural Networks
Jaeyong Bae, Hawoong Jeong
TL;DR
Finite datasets can have similar basic statistics but different label organization, raising whether dataset complexity is reflected in local loss geometry. The paper measures multiscale label mixing and reference-conditioned local entropy, estimating the latter with adaptive sequential Monte Carlo. Greater complexity produces stronger inner contraction of effective solution volume, while the outer response weakens; real-image experiments show the same qualitative pattern, amplified by label randomization.
Problem
Finite datasets can share size and low-order statistics while differing in structural label organization, so the paper asks whether dataset complexity is reflected in local loss geometry.
Method
The paper pairs multiscale local label mixing with Franz–Parisi-inspired local entropy around trained references and estimates finite-network entropy using adaptive sequential Monte Carlo.
Results
Greater dataset complexity produces a deeper inner contraction of effective solution volume followed by a weaker outer response, with the qualitative trend also appearing in MNIST.
Takeaways & Limitations
Dataset structure changes where radial contraction is concentrated rather than simply making solutions uniformly sharper across distance.
Abstract
from arXiv · showhide
Finite datasets can share the same size and low-order statistics while differing strongly in structural complexity. We connect this dataset complexity to loss-landscape geometry by pairing local label mixing across neighborhood scales with local entropy around trained neural-network solutions. Adapted from the Franz--Parisi construction in spin-glass theory, local entropy measures the effective volume of low-loss, solution-like parameter configurations at each distance from a reference. We estimate it in finite networks using adaptive sequential Monte Carlo. In a controlled synthetic sweep, greater dataset complexity produces a larger decrease in local entropy near the reference. Farther away, its radial derivative becomes weak and nearly common across conditions. Dataset complexity therefore changes where the effective solution volume contracts, rather than making it decrease uniformly faster. Experiments on real image data show the same qualitative trend, with label randomization further amplifying the effect. These results show that dataset structure shapes how low-loss neighborhoods are organized across finite distances from trained solutions.
I. INTRODUCTION
Finite datasets with similar size and low-order statistics can differ in how labels are organized, motivating a complexity measure tied to local loss-landscape geometry. The study connects multiscale label mixing with local entropy around trained neural-network references.
- Motivation: Finite datasets may share sample count, class balance, and input means and covariances while differing in local label organization.Nearby labels may be interwoven or arranged in large, well-separated regions.
- Motivation: Dataset complexity denotes how strongly differently labeled examples are interwoven across local scales, independently of sample count, parameter count, and class balance.The operational measure increases with local label mixing and is normalized against random mixing.
- Approach: The study compares multiscale label mixing with local entropy centered on trained references to examine effective solution volume across finite distances.Local entropy complements path- or point-based probes by characterizing low-loss configurations within neighborhoods.
- Findings: Across controlled synthetic data and MNIST, stronger local mixing accompanies a stronger change in surrounding local entropy, concentrated over an inner distance range.The radial derivative becomes more negative before returning toward zero, indicating nonuniform contraction over distance.
- Experiments: The experiments span controlled two-dimensional label patterns and structured image data, including MNIST constructions with randomized labels.The synthetic family varies label organization systematically while complementary MNIST settings test persistence in higher-dimensional data.
B. Dataset complexity and local entropy
The paper operationalizes dataset complexity through multiscale nearest-neighbor label mixing and measures reference-conditioned effective solution volume with local entropy. Its finite-network estimator uses a continuous loss-weighted shell construction around zero-error trained references.
- Dataset complexity: Dataset complexity is measured through nearest-neighbor label mismatches across neighborhood scales and normalized against random label mixing.C_k increases as differently labeled samples become locally interwoven, while random mixing at fixed class counts gives C_MS = 1 in expectation.
- Dataset complexity: All datasets use n = 512 samples, with k_max = 22 possible neighbors, and larger C_MS indicates stronger local label mixing.The neighborhood scale is bounded because k_max grows sublinearly with n.
- Local entropy: Local entropy measures the effective volume of low-loss configurations at a fixed distance from a trained reference.This neighborhood-scale quantity supplements the single-point information provided by the loss at one solution.
- Local entropy: Retained references are zero-training-error interpolating solutions, and the reported entropy averages over the empirical ensemble of such references.The ensemble corresponds to retained trained references in the finite neural-network experiments.
- Local entropy: The finite-network construction continuously weights every direction by lower regularized loss using the shell inverse temperature β_sh = 100 and exp[−β_sh L_reg].The soft construction avoids the sharp separation produced by strict counting of exact solutions.
- Local entropy: The local-entropy decomposition separates a geometric term depending only on P and r from a dataset- and reference-dependent term.Thus, at fixed architecture, dataset differences enter through the loss-dependent contribution rather than the geometric shell factor.
C. Importance sampling for local entropy
The estimator targets low-loss configurations on a fixed-radius parameter shell by replacing inefficient uniform angular sampling with an importance distribution and adaptive sequential Monte Carlo. The importance distribution absorbs the analytically known regularization bias, leaving a loss-dependent factor for numerical estimation.
- Uniform angular sampling can miss relevant low-loss configurations when the weighted shell measure is concentrated in a narrow region.
- The method incorporates the quadratic regularization bias into the sampling measure rather than retaining it in a uniform Monte Carlo weight.The bias is largest in the direction minimizing the parameter norm on the shell.
- The resulting angular proposal is a von Mises–Fisher distribution centered at the norm-minimizing direction, with analytically available normalization.
- The full radial-shell local entropy is compared with an independently calculated thermodynamic-limit replica-symmetric result in the perceptron benchmark.The benchmark uses finite-size curves across N = 40, 80, 160, and 320 trainable weights.
- Adaptive sequential Monte Carlo estimates the remaining loss-partition factor through tempered particle distributions, resampling, and mutation.
III. RESULTS
The results section validates the finite-size shell estimator against a replica-symmetric perceptron calculation before applying it to neural-network loss landscapes. The benchmark shows convergence toward the analytic prediction as model size increases.
- The benchmark tests whether finite-size local-entropy estimates approach the independently calculated replica-symmetric prediction as trainable weights increase.
- The Gaussian perceptron uses signed Gaussian patterns and a soft logistic loss to weight shell configurations around exact reference classifiers.
- The replica-symmetric formulation separates local entropy into radial, directional, and energetic contributions whose competition determines retained solution-like volume.
- The finite-size SMC curve approaches the independently calculated replica-symmetric curve as the number of perceptron weights increases.The comparison uses N ∈ {40, 80, 160, 320}, with 215 particles at each radius.
B. Controlled synthetic label organization
The controlled synthetic family varies label organization while holding the sampled input setting fixed, producing a graded range of multiscale dataset complexity. Greater complexity is associated with earlier and stronger contraction of effective solution volume near trained references.
- Small βdata produces locally interwoven labels, whereas larger βdata favors extended same-label domains and smoother interfaces.
- The multiscale complexity CMS generally decreases with βdata, providing a graded measure of label organization across the synthetic sweep.
- Larger CMS makes centered local entropy decrease more rapidly, so effective solution volume narrows sooner with distance from the trained reference.
- The radial derivative develops a deeper negative inner minimum at high CMS but approaches a common weakly negative value at large radius.
- At r = 1, particle-weighted training accuracy is 0.598 at the highest-complexity endpoint versus about 0.93 in low-complexity conditions.The similar outer derivatives therefore coexist with markedly different shell accuracies, indicating that most volume differences accumulate inward.
- The total radial variation ATV follows the broad graded change across the synthetic complexity sweep.
C. Label organization and local entropy in MNIST
MNIST experiments extend the relationship between label organization and local entropy to fixed-input label randomization and naturally different digit-pair tasks. In both settings, greater complexity mainly changes inner radial contraction, with label randomization producing the strongest contrast.
- Label-noise sweep with fixed inputs: In the label-noise sweep, nested label flips change local organization while keeping each dataset’s input coordinates fixed, and CMS increases with noise.
- Label-noise sweep with fixed inputs: Increasing label noise makes local entropy fall more sharply just outside the trained reference, concentrating effective solution volume in a narrower neighborhood.
- Label-noise sweep with fixed inputs: The negative inner minimum of gE(r) deepens with CMS, while high-noise conditions can show gE(r) crossing zero at intermediate radii.
- Label-noise sweep with fixed inputs: The monotonic increase of ATV with label-noise probability summarizes the changing radial shape of local entropy.
- Label-noise sweep with fixed inputs: The clean–randomized contrast persists after symmetrizing the finite-distance loss response to cancel residual first-order tilt.The symmetrized response rises earlier and much more under random labels, supporting a geometric rather than tilt-only explanation.
- Naturally different digit-pair tasks: The digit-pair comparison has weaker experimental control because changing the pair also changes input distribution and geometry.
- Naturally different digit-pair tasks: Across naturally different digit-pair tasks, larger CMS is likewise associated with faster inner entropy decline and earlier narrowing around the trained reference.
- Naturally different digit-pair tasks: At r = 1, shell training accuracy remains high for both endpoints: 0.979 for task 0/1 and 0.893 for task 4/9.
IV. DISCUSSION
Dataset complexity changes the radial organization of low-loss neighborhoods around trained references: higher complexity sharpens and moves contraction inward rather than uniformly rescaling the landscape. The results extend flat–sharp descriptions across finite radial scales and show that random labels can produce qualitatively distinct outer responses.
- Finite-distance hardening: h(r) = −gE(r)/r shows that higher dataset complexity generally moves the strongest solution-volume contraction closer to the reference and makes it sharper.The rescaling compares radial shape and turning location, not absolute response magnitude.
- Finite-distance hardening: A fixed-Hessian Taylor baseline cannot generate the observed inner rise followed by a turning range, even if dataset complexity changes the reference Hessian.With a fixed quadratic form, h is constant for isotropic curvature and nonincreasing for anisotropic curvature.
- Finite-distance hardening: Across controlled synthetic and MNIST comparisons, greater label mixing produces deeper inner contraction followed by a weaker outer response.Natural digit-pair tasks show milder differences concentrated in the inner radial range.
- Random limits and natural data: Fixed-input label randomization strengthens the inner bottleneck, moves it toward the reference, and can make gE change from negative to positive and back to negative.This differs from the milder regime observed for selected natural digit-pair tasks.
- Finite-distance hardening: The radial response reflects weighted directional reorganization: loss responses strengthen faster than local entropy shifts weight toward easier directions.h is the weighted mean of κ among directions contributing to effective solution volume.
- Conclusion: The paper extends the flat–sharp description from one reference point to radial scales, while noting that broader generalization beyond selected tasks, architectures, representations, and parameter metrics remains open.Future paired interventions could separate effects of data organization from architecture.
Appendix A: Adaptive SMC and controls
The appendix describes an adaptive sequential Monte Carlo procedure for estimating local entropy and radial responses, using tempering, overlap-based controls, resampling, and mutation. Numerical estimates combine independently seeded particle pools and quantify uncertainty across dataset replicates.
- Adaptive SMC: The tempering schedule runs from t0 = 0 to tK = 1, estimating the loss-partition factor as a product of smaller changes.The initial and terminal partition factors satisfy ZL(0; r) = 1 and ZL(1; r) = ZL(r).
- Adaptive SMC: The sampler initializes particles from a known distribution, adaptively tempers them, and applies overlap thresholds to control increments and resampling.Metropolis–Hastings mutation follows each accepted transition to explore the tempered distribution and restore diversity after resampling.
- Adaptive SMC: Pool overlap compares accumulated particle weights with equal weights and triggers systematic resampling when the weighted representation becomes insufficient.Step overlap tests whether current locations represent the next target, whereas pool overlap tests whether the accumulated representation should be reconstructed.
- Numerical controls: The DNN calculations use two independently seeded pools of 512 particles, initialized from a vMF importance distribution and combined with shared normalizers.This reduces sensitivity to one atypical sampling path when estimating partition factors and terminal expectations.
- Numerical controls: The radial derivative gE(r) is evaluated directly from terminal particles rather than finite differences, with uncertainty computed across dataset replicates.The implementation checks that all evaluated directional derivatives are finite.
- Reference-sector inputs: The replica-symmetric calculation summarizes reference solutions with squared norm and typical directional overlap before passing these quantities to the shell calculation.The reference-sector saddle supplies the order parameters used in the reference-conditioned shell analysis.
Supplementary Material
The supplementary material defines the reference-conditioned local-entropy observable and develops its finite-network and replica-based formulation. The construction uses a hard exact-solution reference ensemble and a soft, loss-weighted shell at fixed distance.
- Perceptron formulation: The perceptron formulation absorbs labels into Gaussian patterns, so successful classification is represented by a positive field for each pattern.A logistic cross-entropy loss provides a soft measure of classification-constraint violations.
- Local-entropy construction: The reference configuration is sampled from an exact-classification sector with an L2 prior, while nearby configurations are weighted by soft classification loss and shell energy.The shell is hard in distance but soft in loss.
- Local-entropy construction: Local entropy measures the Gibbs mass of parameter vectors lying at a specified squared distance from a retained trained reference.The observable is conditioned on both the dataset and the selected reference.
- Replica formulation: Replica calculations separate reference normalization from the shell contribution and continue replica counts formally to zero.The shell calculation uses the selected reference representation and its associated normalization.
- Replica formulation: The reference-sector disorder dependence remains in the directional partition factor, while the radial factor is deterministic and disorder-independent.This separation preserves the nontrivial overlap structure in the directional term.
S3. Replica-symmetric reference saddle
The replica-symmetric reference saddle reduces the reference ensemble to macroscopic norm and overlap variables. Its stationary equations determine the reference-sector free entropy and the order parameters imported into the shell calculation.
- Replica-symmetric ansatz: The replica-symmetric ansatz represents each reference copy through its norm and pairwise directional overlap.The reference free-entropy density is expressed using these macroscopic quantities.
- Saddle evaluation: The RS prediction for the reference free-entropy density is used to select a stationary reference branch and pass its order parameters to the shell calculation.The scalar reference free entropy itself is not the final shell observable.
- Order parameters: The reference ensemble is summarized by squared norm eQ and typical overlap qref between two reference directions.qref is interpreted as the typical overlap of two solutions drawn from the same hard reference ensemble.
S4. Shell geometry and order parameters
The shell construction describes parameter configurations at fixed distance from a selected reference using radial norms, overlap matrices, and replica-symmetric reductions. The fixed-distance constraint concerns distance from the reference, while equal shell norms are an additional symmetry assumption.
- Shell constraints: Points on a sphere centered at the reference can have different norms, so setting all shell norms Qγ = Q is an additional replica-symmetry assumption.This assumption restores permutation symmetry among shell replicas.
- Scope of the ansatz: A more general radial-RSB construction would retain independent shell norms before symmetry reduction.The reported construction therefore adopts a narrower radial-symmetry treatment than the fully general alternative.
- Shell constraints: The shell replicas are constrained relative to the same selected reference, with reference–shell overlaps encoding the distance condition.The distance constraint is imposed separately for each shell replica.
- Overlap structure: The fully replicated shell object combines reference, shell, and mixed overlap blocks into a joint covariance structure.Independent disorder patterns contribute through repeated one-pattern energetic factors.
- Shell reduction: Resolving the distance constraints fixes the selected-reference row, leaving an integral over the remaining shell order parameters.The reduced shell target retains optimization over Q and the unfixed overlap variables.
- Shell reduction: The shell RS ansatz reduces the functional saddle to a scalar saddle over Q, p, and t.The reduction follows after resolving distance deltas and importing the reference saddle.
S6. Replica-symmetric shell functional
This section constructs a replica-symmetric shell description around a fixed selected reference branch. It specifies the reference and shell covariance structure, feasible domain, and resulting local-entropy functional.
- Reference and shell representation: The reference block is fixed to the selected replica-symmetric reference branch, with explicit definitions for selected and nonselected reference fields.The construction also introduces independent standard Gaussian variables and the shell-field representation.
- Functional construction: The shell construction combines entropic and energetic contributions into a prelimit one-pattern factor and a shell replica-symmetric functional.The shell entropic contribution, shell energetic term, and s-independent factors are assembled before defining the final functional.
- Reference and shell representation: Shell fields have unit variance and pairwise covariance p for distinct replicas.This covariance structure is used together with the replica-symmetric shell block.
- Feasible domain and entropy density: The feasible shell domain constrains Q, the correlation term c_d(Q, r), and p.The stated conditions include Q > 0, |c_d(Q, r)| ≤ 1, and p < 1.
- Feasible domain and entropy density: The RS shell local entropy density is obtained by evaluating the shell functional over the feasible shell variables.The definitions of the shell covariance, reference data, and functional determine this density.
S7. Saddle equations and constrained branches
This section derives interior saddle equations for the constrained shell functional and treats boundary conditions separately. The retained production branch remains interior across the screened grid.
- Interior saddle equations: The saddle analysis applies only in the interior domain Q > 0, |c_d(Q, r)| < 1, p < 1, and A(Q, p, t; r) > 0.These inequalities define the region where the unconstrained interior equations are valid.
- Interior saddle equations: Interior saddles are determined by equations for the Q, p, and t variables, including distance, entropic, and energetic derivatives.The energetic derivatives use inner and outer selected-reference averages.
- Interior saddle equations: The full-feasible interior saddle equations are identified as equations (S226)–(S228), under the assumption A > 0.These equations are the constrained-branch conditions for the shell functional.
- Constrained branches: Boundary parametrizations are required when A = 0 because the interior derivatives are not valid there.The section gives an explicit A = 0 boundary parametrization subject to |c_d| ≤ 1 and p < 1.
- Constrained branches: The retained production branch was interior throughout the 42-point grid, with A ≥ 7.04 × 10^-7.The ζ = 0 face was screened during branch selection but was not retained as the production branch.
- Constrained branches: The full shell local entropy density is obtained after inserting the reference saddle data into the constrained shell functional.The resulting expression is evaluated on the feasible domain, with interior saddles satisfying equations (S226), (S227), and (S228).