Source-linked AI summary

Learning curves of generic features maps for realistic datasets with a teacher-student model

Bruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mézard, Lenka Zdeborová

arXiv:2102.08127v3stat.MLcond-mat.dis-nncs.LGmath.PRmath.ST

TL;DR

Canonical teacher-student models assume Gaussian i.i.d. inputs, limiting their direct relevance to realistic datasets. This paper introduces a Gaussian covariate generalization with distinct teacher and student feature spaces, derives rigorous asymptotic error formulas, and shows that its predictions capture learning curves across several realistic-data settings while also identifying limitations.

  • Problem

    Canonical teacher-student models rely on Gaussian i.i.d. inputs, creating a need to test whether the framework can describe realistic datasets and generic feature maps.

  • Method

    The paper models teacher and student features as correlated Gaussian vectors from potentially different maps, then derives closed-form asymptotic errors for empirical-risk minimization.

  • Results

    The framework captures learning curves for realistic settings including neural-network features on generated data and ridge regression with real datasets.

  • Takeaways & Limitations

    The model provides a unified framework for analyzing learning curves across deterministic, random, and learned feature maps, including kernels, scattering transforms, and transfer learning.

  • Takeaways & Limitations

    On non-Gaussian real data, the strategy can fail beyond ridge regression and mean-squared test error because classification error depends on more than the first two moments.

Abstract

from arXiv · show

Teacher-student models provide a framework in which the typical-case performance of high-dimensional supervised learning can be described in closed form. The assumptions of Gaussian i.i.d. input data underlying the canonical teacher-student model may, however, be perceived as too restrictive to capture the behaviour of realistic data sets. In this paper, we introduce a Gaussian covariate generalisation of the model where the teacher and student can act on different spaces, generated with fixed, but generic feature maps. While still solvable in a closed form, this generalization is able to capture the learning curves for a broad range of realistic data sets, thus redeeming the potential of the teacher-student framework. Our contribution is then two-fold: First, we prove a rigorous formula for the asymptotic training loss and generalisation error. Second, we present a number of situations where the learning curve of the model captures the one of a realistic data set learned with kernel regression and classification, with out-of-the-box feature maps such as random projections or scattering transforms, or with pre-learned ones - such as the features learned by training multi-layer neural networks. We discuss both the power and the limitations of the framework.

1 Introduction

The paper generalizes teacher-student models beyond Gaussian i.i.d. inputs by allowing teacher and student feature maps to act on different correlated Gaussian spaces. It derives closed-form high-dimensional errors and argues that the framework captures learning curves for realistic datasets and diverse feature maps.

  • Motivation and model: The Gaussian covariate model lets teacher and student use different feature maps of data, represented as correlated Gaussian vectors with arbitrary covariance.The feature maps may be deterministic, random, or learned from data.
  • Realistic-data applications: The framework covers analyses involving kernel regression and classification, random projections, neural tangent kernels, scattering transforms, and transfer learning on generated data.The paper also discusses concrete limits of applicability.
  • Motivation and model: The model defines labels from teacher features while empirical-risk minimization learns from student features, supporting regression and classification losses with regularization.The teacher function can include randomness, while the student accesses only the student features.
  • Theoretical contributions: Theorems 1 and 2 provide rigorous closed-form characterizations of the estimator and its asymptotic training and generalisation errors using Gaussian comparison inequalities.The same expression can also be obtained through the replica method, placing earlier replica-based results on a rigorous basis.
  • Theoretical contributions: A Gaussian-equivalent model is conjectured to capture asymptotic training and generalisation errors for broad data distributions and generic feature maps.The claim extends beyond cases where the original data and feature maps preserve exact Gaussianity.
  • Realistic-data applications: The predictions capture learning curves in settings including generative-adversarial-network data with neural-network features and real datasets with ridge regression.The real-data illustration compares several feature maps on the even-versus-odd MNIST task.

2 Main technical results

The paper derives closed-form asymptotic training and generalisation errors for a Gaussian covariate teacher-student model under broad losses and regularizations. The results are expressed through fixed-point equations and apply to important regression and classification settings.

  • The main result gives closed-form asymptotic training and generalisation errors for the Gaussian covariate model with ℓ2 regularization.The theorem is stated in the high-dimensional limit under the model’s assumptions.
  • Assumptions: The model assumes n, p, and d grow jointly with fixed ratios, while covariance matrices satisfy positivity, spectral-convergence, and bounded-singular-value conditions.Additional regularity assumptions constrain the loss and regularization functions.
  • Fixed-point characterization: The asymptotic errors are obtained by solving self-consistent fixed-point equations for overlap parameters such as V*, q*, and m*.The equations can be iterated to find their fixed point, whose parameters enter the training and generalisation errors.
  • Fixed-point characterization: The parameters depend on the teacher projection, student covariance spectrum, and their limiting joint empirical distribution.The projected teacher weights and student eigenvalues are identified as relevant model statistics.
  • Applications: Special cases include ridge regression with mean-squared error and binary classification with sign outputs, classification error, and logistic training loss.These cases yield simplified asymptotic expressions used in the experiments.
  • Generality: The framework covers generic non-separable losses and regularizations, including finite-size distributional results and high-dimensional error observables.Theorem 2 provides a non-asymptotic characterization for optimal solutions under stated regularity conditions.

3 Applications of the Gaussian model

The paper applies its Gaussian covariate framework to kernels, random features, generative-network data, learned neural features, and real datasets. The theory often matches observed learning curves, while deviations emerge for some real-data settings and tasks.

  • Applications: The framework covers random kitchen sinks, kernels, generative neural networks, and feature maps that may be deterministic, random, or learned.Random-feature models use nonlinear maps of randomly projected inputs, while kernel methods are represented as teacher-student problems in feature space.
  • GAN-generated data and learned teachers: The GAN pipeline models realistic inputs as images generated from Gaussian latent vectors, then assigns labels using fitted teacher weights.The experiments use a dcGAN trained on CIFAR10 and a neural-network teacher trained for an odd-versus-even task.
  • GAN-generated data and learned teachers: The Gaussian predictions closely match logistic-regression learning curves on dcGAN-generated CIFAR10-like images across learned feature maps from different training epochs.The comparison includes generalisation classification error and unregularised training loss as functions of sample complexity.
  • GAN-generated data and learned teachers: At epoch 0, random features outperform early learned features in one experiment, while epoch-50 features become separable at lower α than epoch-200 features despite worse final performance.Here α = n/d, and the learned feature maps correspond to different stages of neural-network training.
  • Learning from real data sets: For real datasets, the method predicts ridge-regression mean-squared-error curves and accurately captures interpolation peaks for random and scattering features.The real-data procedure estimates covariance matrices and teacher quantities from the dataset before evaluating the theoretical equations.
  • Learning from real data sets: The real-data approximation deviates near the full dataset size and fails beyond ridge regression and mean-squared test error, including a classification-error mismatch.The limitation is attributed to classification error depending on more than the first two moments, while exact teacher recovery would predict zero test error in the asymptotic setup.

A Main result from the replica method

The appendix derives the Gaussian covariate model’s asymptotic performance using a replica calculation. It formulates a Gibbs measure, reduces the high-dimensional problem to overlap parameters, and obtains saddle-point equations in the high-dimensional limit.

  • Replica derivation: The replica method introduces a Gibbs measure and computes its free-energy density through replicated partition functions and overlap parameters.The overlap variables encode relations between teacher weights and student weights, reducing the estimation problem to scalar order parameters.
  • Result: The resulting equations characterize the estimator and its asymptotic training and generalisation errors, with an implementation provided for the losses discussed.The derivation is presented as a heuristic replica calculation, while the main manuscript also establishes the corresponding result using Gaussian comparison inequalities.
  • Model setup: The model draws jointly Gaussian teacher and student covariates with correlation matrices Ψ, Ω, and Φ, while labels depend on a teacher function of u.The student observes v and learns weights by regularised empirical risk minimisation.
  • Model setup: The analysis measures asymptotic training and generalisation errors using sample complexity α = n/d and aspect ratio γ = p/d.The estimator uses convex loss and regularisation functions, including logistic or square loss and ℓp regularisation.
  • Replica derivation: In the high-dimensional limit, the integral concentrates at extrema of a potential, and a replica-symmetric ansatz yields self-consistent saddle-point equations.The limit sends d to infinity while α and γ remain finite, with replica overlaps constrained to shared diagonal and off-diagonal values.
  • Loss dependence: The loss-dependent term is expressed through a Moreau envelope, whose derivative or proximal operator determines the scalar response function used in the fixed-point equations.For logistic loss the proximal equation lacks a closed form, whereas hinge-loss and some quadratic cases simplify.

A.5 Examples

The examples specialize the Gaussian covariate equations to ridge regression, binary classification, logistic loss, and hinge loss. They provide explicit or fixed-point characterizations of asymptotic errors and connect vanishing-regularization classifiers to max-margin solutions.

  • Ridge regression: For ridge regression with a linear teacher, the asymptotic training and generalisation errors are determined by fixed-point variables V⋆, q⋆, and m⋆.The task uses squared loss with matching linear teacher and predictor.
  • Ridge regression: The parameter V⋆ depends only on the population-covariance spectrum and describes the variance gap between generalisation and training error.Its interpretation links covariance spectral structure to the deformation of the Gaussian field governing the asymptotic output.
  • Binary classification: For binary classification with sign teachers, the asymptotic classification error is expressed through overlap variables satisfying self-consistent saddle-point equations.The framework recovers the special isotropic case with d = p and extends it to the broader Gaussian covariate setting.
  • Binary classification: As λ → 0, logistic and soft-margin solutions converge to the max-margin estimator.The classification equations generally require solving loss-specific fixed-point relations rather than direct integration.

A.6 Relation to previous models

The framework connects the Gaussian covariate model to kernel methods, random features, and models involving generative or pre-learned feature maps. These correspondences recover known kernel-regression equations and extend the model to broader supervised-learning settings.

  • Random features: Random-feature learning is represented by applying a random projection followed by a component-wise nonlinearity, with covariance parameters linked to the projection matrix.For Gaussian data, these relations hold asymptotically through a Gaussian equivalence theorem.
  • Generative and pre-learned features: The framework includes random-feature regression on data generated by pre-trained models and generalizes that setting to structured teachers and fixed feature maps.Examples include generative models, random features, scattering transforms, and pre-learned neural networks.
  • Kernel methods: Kernel regression can be rewritten as ridge regression in feature space, making it a special case of the Gaussian covariate model.The correspondence uses the kernel eigenvalues and eigenvectors to define the feature map and covariance matrices.
  • Kernel methods: The model recovers the self-consistent equations for kernel ridge regression and extends the analysis to kernel logistic regression and support vector machines.The recovery follows after rescaling n, ρ, m, q, and λ by d.

B Rigorous proof of the main result

The rigorous analysis formulates teacher and student data as correlated Gaussian blocks and studies convex empirical-risk minimization under high-dimensional assumptions. Its proof targets the estimator and associated training and generalisation errors.

  • Problem formulation: The proof analyzes matrices of teacher and student features, with the estimator defined through a convex objective involving the training function g.The formulation uses concatenated teacher vectors U and student vectors V.
  • Gaussian structure: The tuple governing teacher outputs, student features, and estimator quantities is bivariate Gaussian with a covariance determined by the model.The analysis introduces overlaps that characterize the estimator.
  • Main analytical result: The estimator distribution is computed in the weak sense from the unique solution of six scalar fixed-point equations.This result is stated for the fully generic formulation without introducing the spectral decomposition used in the ℓ2 case.
  • Assumptions: The generic theorem requires positive-definite covariance structure, convergent spectral distributions, bounded singular values, convex lower-semicontinuous functions, and coercivity.Additional conditions control scaling, independence, dimension ratios, and finite-sample concentration rates.
  • Assumptions: The assumptions ensure a non-vanishing teacher distribution and make the optimization and concentration arguments applicable to common machine-learning settings.The discussion specifically identifies ridge-regularized convex losses, LASSO, and elastic-net as examples covered by coercive regularization.

B.2 Main theorem

The main theorem derives asymptotic training loss and generalisation error through scalar potentials and their minimizers. Under the stated assumptions, the relevant quantities have finite high-dimensional limits and satisfy concentration results.

  • Scalar formulation: The analysis defines scalar quantities, Moreau-envelope terms, and a potential whose variables characterize the asymptotic behavior of the estimator.The potential uses Gaussian variables, the teacher function, and Moreau envelopes of the target functions.
  • Asymptotic limits: Under Assumption (B.1), the defined quantities admit finite limits as n, p, d →∞.These limits provide the asymptotic objects used in the theorem.
  • Optimization structure: The potential is jointly convex in selected variables and jointly concave in the remaining variables, with optimality conditions yielding self-consistent fixed-point equations.The fixed-point equations arise from the optimization problem for the potential.
  • Training and generalisation: Theorem 4 gives concentration bounds for the training loss and generalisation error of any optimal solution under Assumption (B.1).The theorem applies for sufficiently small positive ε and constants C, c, and c′.
  • Estimator observables: Theorem 5 extends the concentration statement to observables of the estimator, while broader classes of test functions lose exponential concentration rates.The rates depend on the regularity class of the functions used to evaluate the estimator and training variables.

B.3.1 A Gaussian comparison theorem

The proof uses Gaussian comparison and convex-analysis tools to replace a high-dimensional optimization problem with a simpler auxiliary problem. Concentration results then control Moreau-envelope and pseudo-Lipschitz observables.

  • Gaussian comparison: The Convex Gaussian Min-max Theorem compares a primary optimization problem with an auxiliary problem involving Gaussian vectors.The auxiliary problem is used to study the asymptotic properties of the original optimization.
  • Gaussian comparison: The primary and auxiliary optimization problems share a convex-concave structure that enables analysis of the simpler auxiliary formulation.The paper calls reformulations matching these forms acceptable primary and auxiliary problems.
  • Convex analysis: Moreau envelopes and proximal operators provide the convex-analytic representation used throughout the proof.The Moreau envelope is defined alongside its proximal operator, whose properties support subsequent bounds.
  • Convex analysis: The proof establishes monotonicity properties for several functions derived from proximal operators and Moreau envelopes.The stated result is that h1 is nondecreasing while h2, h3, and h4 are nonincreasing.
  • Concentration: Gaussian Poincaré inequalities yield concentration for Moreau envelopes and pseudo-Lipschitz functions under the required scaling assumptions.Exponential concentration is stated separately for separable pseudo-Lipschitz functions of order two.
  • Concentration: The resulting concentration lemmas support finite-dimensional-to-asymptotic control of the quantities used in the main theorem.The Moreau-envelope concentration result assumes proper convex functions satisfying the theorem’s scaling conditions.

B.4 Determining a candidate primary problem, auxiliary problem and its solution.

The proof reformulates the high-dimensional learning problem into primary and auxiliary optimization problems, then reduces the auxiliary problem to a scalar optimization over six parameters.

  • Primary and auxiliary problems: The original optimization is rewritten with auxiliary variables so that it can be analyzed through a primary problem and its corresponding auxiliary problem.The reformulation introduces variables such as z and λ while preserving the original minimization structure.
  • Feasibility and compactness: The feasibility sets are shown to be compact by combining coercivity, lower semicontinuity, bounded covariance operators, and scaling assumptions.The minimizers and auxiliary variables receive dimension-independent bounds.
  • Gaussian reformulation: The proof uses Gaussian decompositions and conditioning to isolate independent standard Gaussian matrices from the teacher and response variables.Orthogonal projections in Gaussian space produce equivalent formulations involving independent copies of Gaussian matrices.
  • Variable decomposition: Orthogonal decomposition separates the student weight into a component aligned with the projected teacher signal and an orthogonal component.A scalar Lagrange multiplier enforces orthogonality to the projected teacher direction.
  • Scalar reduction: The resulting auxiliary problem is reduced to a scalar optimization over six parameters involving Moreau envelopes of the loss and regularizer.The reduction introduces scalar variables including τ1, τ2, κ, η, ν, and m.

B.5 Study of the scalar equivalent problem : geometry and asymptotics.

The scalar equivalent problem has a controlled convex-concave geometry, admits an asymptotic limit, and yields fixed-point equations for its optimal solution.

  • Geometry: The finite-dimensional scalar objective is continuous, jointly convex in selected minimization variables, and jointly concave in selected maximization variables.The convex variables include (m, η, τ1), while the concave variables include (κ, ν, τ2).
  • Asymptotics: The scalar objective converges to an asymptotic potential whose optimal value characterizes the high-dimensional limit.The limiting potential retains the same convex-concave structure and is obtained through convergence of the finite-dimensional objective.
  • Asymptotic objective: The asymptotic potential is continuously differentiable and combines quadratic terms with limiting Moreau-envelope contributions from the loss and regularization.Its structure includes terms represented by Lg and Lr together with the scalar parameters of the optimization.
  • Fixed points: The zero-gradient conditions of the asymptotic optimization define fixed-point equations for every feasible solution.These equations can also be translated into the notation used by the replica method.
  • Uniqueness and stability: Strict concavity and strict convexity near minimizers support uniqueness and stability of the scalar optimization in its respective parameter blocks.The result applies to (τ2, κ, ν) under fixed minimization variables and to (η, m, τ1) on the specified set.

B.6 Back to the original problem : proof of Theorem 4 and 5

The proof transfers the scalar asymptotic characterization back to the original estimator, establishing finite-dimensional concentration and linking the result to replica predictions.

  • Assumptions: The analysis begins under assumptions ensuring bounded spectra, suitable teacher behavior, and regularity of the loss, regularizer, and test functionals.These assumptions support concentration and finite-size control for the original optimization problem.
  • Transfer to optimizers: Strong convexity of the auxiliary problem connects closeness of finite-dimensional objective values to closeness of the corresponding optimizer variables.This provides the bridge from scalar optimization errors to estimator-level errors.
  • Concentration and rates: Gaussian and concentration arguments establish convergence rates for the scalar objective and the original estimator under the stated regularity conditions.The proof combines concentration of random terms, Moreau envelopes, and an ε-net argument.
  • Rate limitation: Relaxing the regularity of f, g, φ1, or φ2 from the stated classes can replace exponential rates with linear rates.The loss of exponential rates is attributed to the weaker pseudo-Lipschitz assumptions.
  • Random teachers: Random teacher vectors can be handled by conditioning on the teacher and adding an expectation, provided they are independent of the Gaussian matrices and model randomness.The finite-size rates then depend on teacher assumptions and covariance-eigenvalue decay.
  • Replica equivalence: The rigorous result is used to prove the replica prediction for separable losses with ridge regularization in a Gaussian random-teacher setting.The appendix derives an exact analytical matching between the Gordon-theorem result and the replica prediction in this case.

D.1 Ridge regression on real data

The experiments evaluate theoretical learning curves for ridge and kernel regression on real image datasets using random, scattering, kernel, and learned neural-network features. Agreement is strong at smaller sample counts but degrades near the finite data-universe size; universality also breaks for classification error.

  • Experimental setup: The experiments compare ridge or kernel regression learning curves on MNIST and fashion-MNIST using random, scattering, kernel, and learned neural-network features.The feature maps include random projections, scattering transforms, kernel approximations, and snapshots of trained neural networks.
  • Ridge and kernel regression: Random features and scattering transforms show characteristic double descent, with theory accurately predicting the interpolation-transition peak.The random-feature dimension is matched to the scattering-transform dimension in the MNIST comparison.
  • Finite-universe effects: As n approaches ntot, theoretical learning curves begin deviating from simulations because population covariances are estimated from the finite universe.The NTK experiment with ntot = 7000 predicts perfect generalisation near ntot, while the simulated error approaches a constant.
  • Limits of universality: Changing from squared-error regression to binary classification produces a mismatch between theoretical and simulated curves, even when the estimator remains the ridge solution.The breakdown occurs for MNIST odd-versus-even classification with the square loss and predictor sign.

D.2 Binary classification on GAN generated data

The experiment constructs synthetic CIFAR10-looking images with a trained dcGAN, labels them using a teacher trained on real CIFAR10, and evaluates logistic regression on learned student features. The pipeline tests theoretical predictions in a realistic generated-data classification setting.

  • Student features: The student feature maps are obtained by removing the final layer of a trained three-layer neural network at multiple training epochs.Snapshots are taken at epochs 0, 5, 50, and 200.
  • Synthetic-data pipeline: A dcGAN maps i.i.d. Gaussian latent vectors to CIFAR10-looking images, which a teacher trained on real CIFAR10 labels for binary classification.The generated samples are then used as the data universe for the student task.
  • Classification experiment: Logistic regression is trained on the learned features using fresh dcGAN-generated samples and teacher-assigned labels.The reported points and error bars average over 10 independent runs.
  • Theory evaluation: The theoretical learning curves are computed from population covariances estimated by Monte Carlo sampling of the synthetic GAN distribution.The covariance estimates use 10^6 samples with precision of order 10^-5.

E Ridge regression with linear teachers

This appendix explains why Gaussian-covariate ridge-regression predictions can extend beyond Gaussian data: the relevant errors depend on covariance spectra and trace products, whose asymptotics can be universal. The argument remains heuristic and is developed explicitly only for linear teachers and restricted settings.

  • Scope: The appendix limits its universality discussion to linear teachers and notes that extending the argument to the full covariance model would require additional work.The cited random-matrix reasoning is therefore heuristic rather than a complete proof for every non-Gaussian setting.
  • Ridge-regression reduction: For ridge regression with a linear teacher, training and test errors can be expressed through empirical and population covariance matrices.The Gaussian calculation reduces the problem to evaluating traces involving these covariance matrices.
  • Random-matrix perspective: The Gaussian analysis uses random matrix theory to compute covariance-related traces, while replica methods can obtain the same result without explicit RMT calculations.The relevant quantities include covariance spectra and traces of products between empirical and population covariances.
  • Universality: Universality is expected when non-Gaussian data share the same population covariances and their empirical spectra and trace products converge to Gaussian-predicted values.These conditions are supported for broad classes of distributions by random-matrix results, though the argument is not fully general here.
  • Restricted example: In the restricted case u = v, Marchenko–Pastur spectral behavior helps explain why ridge-regression predictions can remain valid beyond Gaussian covariates.The discussion connects covariance-spectrum robustness to successful applications on real data.
Loading 2102.08127v3…