Source-linked AI summary

Landau theory of quenched criticality in linear in-context learning

Daesik Kim, Sumin Choi, Hyojae Jeon, Jung Hoon Han

arXiv:2608.28059v1cond-mat.dis-nncs.LG

TL;DR

Linear ICL exhibits a double-descent singularity near its interpolation threshold, raising the question of which quantity fluctuates critically. The paper compares annealed and quenched theories, integrates the cavity equation into a Landau potential, and finds that connected learning-parameter fluctuations correspond to a diverging susceptibility at τ_c = 1, with numerical support and a large-context pseudogap regime.

  • Problem

    Linear ICL has a double-descent singularity at the interpolation threshold, but the critically fluctuating quantity underlying this behavior remains an open physics question.

  • Method

    The paper compares annealed and quenched descriptions and integrates the cavity self-consistency equation for ξ into a Landau potential.

  • Results

    Connected sample-to-sample fluctuations of the learned parameters generate the singular prediction-error contribution, while the Landau susceptibility diverges at τ_c = 1.

  • Takeaways & Limitations

    The Landau order parameter has a geometric interpretation through zero modes, and the theory predicts a large-context pseudogap-like regime confirmed by numerical solutions.

Abstract

from arXiv · show

In-context learning (ICL) allows a pretrained model to infer a new task from examples supplied in its prompt without updating its parameters. In linear models of ICL, the prediction error develops a double-descent singularity when the number of pretraining samples becomes comparable to the number of learnable parameters. We formulate this interpolation singularity as a critical phenomenon of a quenched disordered system. By comparing annealed and quenched descriptions of the same linear ICL model, we identify the connected sample-to-sample fluctuations of the learned parameters as the microscopic origin of the singular error. A Landau potential is constructed by integrating the cavity self-consistency equation for the renormalized ridge parameter $ξ$. The role of (magnetization) order parameter is played by $ξ$, while the bare ridge parameter $λ$ becomes its conjugate magnetic field. The normalized sample complexity $τ$ acts as a temperature and the double-descent singularity occurs at the critical temperature $τ_c =1$. The Landau susceptibility is precisely the quantity that diverges in the fluctuation contribution to the prediction error. The order parameter is closely related to the fraction of zero eigenvalues of the empirical relaxation matrix in the ridgeless limit, which define flat directions in the learning dynamics. The Landau theory is generically cubic in the order parameter with critical exponents $(β_{\rm cr},δ_{\rm cr},γ_{\rm cr})=(1,2,1)$. In the large-context regime, there appears a pseudogap-like regime characterized by suppressed order parameter. Predictions of the Landau theory are independently confirmed from numerical solutions of the original learning problem with good quantitative agreement. Our results pave the way for solid statistical-physics understanding of the interpolation criticality in linear in-context learning.

I. INTRODUCTION

The paper frames linear in-context learning’s interpolation singularity as a critical phenomenon, identifying quenched learning-parameter fluctuations as its source and developing a Landau description validated numerically.

  • Linear ICL uses fixed pretrained parameters to infer task-dependent predictions from prompt demonstrations rather than updating parameters for each task.
  • The interpolation threshold occurs at τ = 1, where the number of pretraining constraints becomes comparable to the number of trainable directions and double descent appears in the ridgeless limit.
  • Comparing annealed and quenched calculations identifies divergent connected sample-to-sample fluctuations of the learned parameters as the source of the double-descent prediction-error singularity.
  • Integrating the cavity self-consistency equation defines a Landau potential in which ξ is the order parameter, τ is temperature-like, and λ is the conjugate field.
  • The order parameter is geometrically related to zero modes of the empirical relaxation matrix, whose flat directions produce no learning dynamics and explain the cubic Landau term.
  • For finite inverse context length, the critical point is τ_c = 1 with exponents (β_cr,δ_cr,γ_cr) = (1,2,1), while large contexts yield a suppressed-order-parameter pseudogap and a nonanalytic endpoint.
  • Numerical solutions of the original learning problem reproduce the predicted pseudogap suppression and crossover boundaries, supporting Landau theory as an effective description.

A. Calculating annealed averages and learning vectors

The annealed scheme replaces empirical matrices by their statistical averages in the joint asymptotic limit, yielding an analytically tractable learning vector and error description.

  • The annealing scheme solves the gradient-flow equation after replacing R and S with their trace-normalized expectations.
  • In the joint asymptotic limit, n scales as d^2 while k scales as d, so each training task is sampled O(d) times and empirical averages become statistical averages.
  • For quantities depending on task vectors, the empirical task average is replaced by the uniform average over the task pool and can be evaluated analytically for a specified training distribution.
  • For Gaussian training tasks, the relevant task-vector statistics can be computed explicitly.
  • The annealed result corresponds to the τ → ∞ limit of the quenching scheme, where τ acts as temperature and critical fluctuations vanish.
  • The training covariance is Wishart with Marchenko–Pastur eigenvalue statistics, and its zero-eigenvalue fraction is max(1 − κ,0).
  • In the joint asymptotic limit, subleading terms involving b_tr and b_te are O(1/d) and are omitted, leaving annealed error expressions based on C_tr and C_te.

C. ICL and IDG error estimation

The annealed and quenched analyses yield finite baseline errors, but quenched sample-to-sample fluctuations generate the singular ICL error at the interpolation threshold. In the large-sample limit, fluctuations vanish and quenched errors reduce to annealed errors.

  • ICL error develops a nonanalytic singularity, whereas the annealed error functions remain finite for all parameter values.
  • In the ridgeless limit, the zero-eigenvalue projector identifies the fraction of flat directions, and this structure contributes to the limiting error expressions.
  • The nonanalytic ICL contribution depends on task diversity κ and originates from distribution mismatch between training and test tasks.
  • The quenched error becomes singular at the critical sample complexity τ_c = 1, producing the double-descent behavior.
  • The quenched learning vector replaces the bare ridge parameter λ with a self-consistently determined renormalized parameter ξ.
  • At τ →∞, the fluctuation contribution vanishes, empirical averages converge to distribution averages, and quenched dynamics reduce to annealed dynamics.

V. LANDAU THEORY OF LICL

The Landau theory integrates the cavity self-consistency equation for ξ into a potential whose temperature-like control is τ and whose conjugate field is h = τλ. Its susceptibility diverges at τc = 1 and matches the singular fluctuation contribution to prediction error, while ξ is geometrically tied to flat directions.

  • Landau construction: The Landau potential is obtained by integrating the self-consistency equation for ξ, with τ as temperature and h = τλ as its conjugate field.At zero field, ξ vanishes at the critical temperature τc = 1.
  • Landau construction: The cubic term is allowed because ξ is non-negative and lacks ξ → −ξ symmetry, so ξ is not an ordinary spontaneous-symmetry-breaking order parameter.The paper instead gives ξ a geometric interpretation through the relaxation matrix.
  • Landau susceptibility: The fluctuation coefficient in quenched prediction error is proportional to the Landau susceptibility, which diverges as |t|^-1 at zero field.This identifies the double-descent divergence with a diverging Landau susceptibility.
  • Critical behavior: The generic critical exponents are (β_cr,δ_cr,γ_cr) = (1,2,1), satisfying the Widom scaling relation.These exponents describe the fixed-finite-γ critical point at τc = 1.
  • Geometric interpretation: In the ridgeless limit, ξ is related to the fraction of zero eigenvalues of the relaxation matrix, whose flat directions contain no learning dynamics.The trace relation connects ξ to the flat-direction density ρflat.

B. Crossover behavior for γ ≪ξ∗≪1

For large context length and κ < 1, the Landau solution exhibits a crossover at τ = κ before the true finite-γ critical point τc = 1. The residual order in the pseudogap region scales as O(γ) and vanishes in the strict infinite-context limit, where additional critical behavior appears.

  • Crossover behavior: For finite γ, the true critical temperature remains τc = 1, while the region κ < τ < 1 has residual order ξ∗ = O(γ).The residual order shrinks as γ → 0.
  • Geometric interpretation: For κ < 1, the infinite-context flat-direction density reduces to ρflat = 1 − min(τ,κ), recovering the crossover solution for ξ∗.This supports the geometric interpretation of the crossover order parameter.
  • Nonanalytic critical point: At κ = 1 and γ = 0, a nonanalytic ξ^(5/2) term replaces the cubic term and changes the critical exponent β_cr to 2.For small finite γ, this behavior appears only in an intermediate crossover regime.

VI. NUMERICAL TESTS OF LANDAU THEORY

Numerical solutions independently extract the Landau order parameter and reproduce its predicted phase structure, including a strongly suppressed pseudogap regime at large context lengths. The extracted scaling behavior and direct comparisons agree well with the Landau theory.

  • Numerical extraction: The empirical extraction scheme obtains ξ∗ independently of the cavity self-consistency equation.Formula (6.1) provides a separate route for extracting the order parameter from numerical solutions.
  • Phase diagram: For γ = 10^-2, the pseudogap suppresses ξ∗ prominently between the critical line τ = 1 and crossover line τ = κ.The suppression is especially visible in the region bounded by τ = 1 and τ = κ.
  • Theory comparison: At κ = 0.6, the empirically extracted ξ∗shows an excellent fit to curves obtained from the theoretical self-consistency equation.The comparison supports the reliability of the extraction scheme.
  • Scaling regimes: The fitted exponent p rises rapidly but smoothly from p ≈0 to p ≈1 across the crossover line τ = κ.This behavior matches the two scaling regimes predicted by the Landau theory.
  • Criticality: The double-descent peak is identified as a divergent connected fluctuation in the learning parameter through contrast between annealed and quenched averages.The Landau susceptibility, proportional to the inverse curvature V′′(ξ∗), controls this singular fluctuation contribution.
  • Geometric interpretation: In the ridgeless limit, ξ∗is related to the fraction of zero eigenvalues of the empirical relaxation matrix, corresponding to flat learning directions.The gradient flow preserves initial components along these directions.

Appendix A: Derivation of the formula for ˆyℓ+1 in linear in-context learning

The appendix derives the linear in-context prediction rule by expressing linear self-attention in terms of token components and vectorized learning variables. The resulting prediction is a weighted average of demonstrated labels, with weights determined by query-key overlaps.

  • Parameter dependence: The derivation depends on the product M = Q⊤K rather than on Q and K separately.This mirrors the dependence of the full Transformer attention mechanism on query-key interactions.
  • Linear self-attention: The token and matrix decompositions reduce the linear self-attention calculation to a compact expression for ŷℓ+1.The derivation sets irrelevant matrix components to zero and absorbs Vyy into redefined parameters.
  • Prediction rule: The prediction is an average of existing labels ym weighted by overlaps between the query qℓ+1 and keys km.Here qℓ+1 = Qxℓ+1 and km = Kxm.
  • Vector formulation: Vectorization rewrites the prediction as an inner product between two vectors, with spatial inputs rescaled by d for large-d scaling.The convention vec[uv⊤] = u⊗v is used in this reformulation.
  • Vector formulation: The learning vector P contains d(d +1) trainable parameters, while H contains the context and test-input information used for prediction.Together, P and H define the linear ICL scheme.

Appendix B: Computation of the annealed relaxation matrix and source vector

The annealed calculation decomposes context-dependent fields into averages and fluctuations, then computes the corresponding relaxation matrix and source vector. In the joint asymptotic limit, fluctuation corrections vanish under the stated assumptions.

  • Field decomposition: The annealing procedure decomposes each field Hµ into an average and a fluctuation.The normalized context length is α = ℓ/d.
  • Conditional averages: Conditional averaging over input vectors produces the mean and covariance structure associated with the training task vectors.The task vectors are averaged conditionally through their distribution and repeated-sample structure.
  • Annealed construction: The relaxation matrix and source vector are decomposed and then assembled from their conditional expectations.These operations complete the annealed construction.
  • Asymptotic limit: The joint asymptotic limit makes δE(wµ) vanish, simplifying the annealed field representation.The resulting expressions use the mean field without the vanishing fluctuation term.
  • Asymptotic assumption: When n far exceeds k, averages over n training samples can be replaced by averages over the uniform task-pool distribution.This replacement uses the assumption that the number of training samples greatly exceeds task diversity.
  • Annealed construction: The annealed relaxation matrix and source vector are completed after applying the conditional averaging steps.The construction supplies the quantities used in the annealed learning calculation.

Appendix C: Marchenko-Pastur distribution

The Marchenko-Pastur analysis characterizes the rank and zero-eigenvalue structure of the relevant Wishart matrix as a function of task diversity. Above unit task diversity the matrix is full rank, while below it a fraction of eigenvalues is zero.

  • Marchenko-Pastur law: In the simultaneous large-k,d limit, the Wishart eigenvalues follow the Marchenko-Pastur distribution at fixed κ = k/d.The analysis applies the distribution through its Stieltjes transform.
  • Full-rank regime: When k exceeds d, the Wishart matrix is full rank and its d eigenvalues are statistically nonzero.The continuous eigenvalue density carries unit total weight in this regime.
  • Low-rank regime: When κ < 1, the Wishart matrix is low-rank with approximately d − k zero eigenvalues and k nonzero eigenvalues.The delta-function contribution has weight 1 − κ = (d − k)/d.
  • Annealed error: The annealed error is derived from the relaxation-matrix and source-vector averages and simplifies in the asymptotic ICL and IDG limits.The resulting expression is extended to a general test distribution.
  • Annealed error: The general annealed error formula follows after combining the averaged terms and applying the stated simplifications and normalization assumptions.The expression uses eλ = γ + λ and neglects terms containing btr or bte within the paper’s scope.

Appendix E: Order estimation of the btr-dependent terms in the annealed test error

Appendix E shows that terms involving btr in the annealed test error vanish asymptotically for any small nonzero ridge parameter λ. The derivation specializes the expressions to the ICL setting after explicit inverse-matrix manipulations.

  • The calculation begins by explicitly computing elements of (Etr + λId+1)^−1.
  • The btr-dependent contributions to the annealed test error are summarized before specializing to the ICL setting.
  • The ICL specialization uses bte = 0 and Cte = Id to reduce the btr-dependent expression.
  • For any small but nonzero λ, all btr-containing terms in the test errors vanish in the joint asymptotic limit.This agrees with taking the ridgeless limit only after the joint asymptotic limit.

Appendix F: Derivation of the annealed ICL and IDG Errors

Appendix F derives the annealed ICL and IDG errors through spectral manipulations based on the Marchenko–Pastur distribution. It then establishes positivity of their difference and identifies a task-diversity-dependent zero-eigenvalue contribution.

  • The empirical eigenvalues of Ctr follow the MP distribution in the joint asymptotic limit.
  • The Stieltjes transform Iκ(z) of the MP distribution provides spectral identities used to derive the annealed errors.
  • The annealed ICL–IDG error difference is strictly positive for positive renormalized ridge parameter eλ.
  • For κ < 1, the ICL–IDG error gap remains 1−κ because of the zero-eigenvalue sector of Ctr, while it vanishes for κ ≥ 1.
  • The appendix derives the error-difference expression and lower bound from the annealed error formulas and spectral decomposition.
Loading 2608.28059v1…