Source-linked AI summary
Cleaning large correlation matrices: tools from random matrix theory
Joël Bun, Jean-Philippe Bouchaud, Marc Potters
TL;DR
High-dimensional covariance estimation becomes problematic when the number of observations is not large relative to the number of variables. The review develops RMT-based analytical tools and covariance estimators, showing that large-dimensional risk can substantially exceed optimal risk and that eigenvector information remains essential.
Problem
Estimating all entries or even the eigenvalues of a high-dimensional true correlation matrix becomes problematic when the observation count is not large relative to the number of variables.
Method
The review combines Coulomb-gas methods, free probability, Bayesian covariance estimation, and rotationally invariant priors to study large sample covariance matrices.
Results
For a wide class of processes, TrE^-1 = TrC^-1/(1−q), so out-of-sample risk can exceed optimal risk when q > 0 and diverge as q →1.
Takeaways & Limitations
Reliable high-dimensional covariance estimation requires attention to sample eigenvectors as well as eigenvalues, because the Marchenko-Pastur equation alone does not provide eigenvector information.
Takeaways & Limitations
The Marchenko-Pastur equation describes the eigenvalue spectrum of large sample covariance matrices but does not yield information about their eigenvectors.
Abstract
from arXiv · showhide
This review covers recent results concerning the estimation of large covariance matrices using tools from Random Matrix Theory (RMT). We introduce several RMT methods and analytical techniques, such as the Replica formalism and Free Probability, with an emphasis on the Marchenko-Pastur equation that provides information on the resolvent of multiplicatively corrupted noisy matrices. Special care is devoted to the statistics of the eigenvectors of the empirical correlation matrix, which turn out to be crucial for many applications. We show in particular how these results can be used to build consistent "Rotationally Invariant" estimators (RIE) for large correlation matrices when there is no prior on the structure of the underlying process. The last part of this review is dedicated to some real-world applications within financial markets as a case in point. We establish empirically the efficacy of the RIE framework, which is found to be superior in this case to all previously proposed methods. The case of additively (rather than multiplicatively) corrupted noisy matrices is also dealt with in a special Appendix. Several open problems and interesting technical developments are discussed throughout the paper.
1. Introduction
High-dimensional correlation matrices are difficult to estimate because sample matrices can differ substantially from population matrices when N/T remains non-negligible. The review develops RMT-based tools, including eigenvector-aware rotationally invariant estimators, and examines their implications for applications such as portfolio risk.
- 1.1. Motivations: Large datasets require reliable covariance estimation for applications including factor identification, PCA, generalized least squares, classification, and portfolio analysis.The relevant observables range from stock returns and biological indicators to physical and macroeconomic measurements.
- 1.1. Motivations: The sample correlation matrix E is computed from T realizations of N demeaned and standardized observables to estimate the population matrix C.When N is much smaller than T, E converges almost surely to C under classical multivariate statistics.
- 1.1. Motivations: When N and T grow with q = N/T non-negligible, estimating all entries or eigenvalues of C becomes problematic because E must be distinguished from the true matrix C.This regime is called the large dimension limit or Kolmogorov regime.
- 1.2–1.3. Historical survey and outline: The review surveys RMT techniques and develops consistent covariance estimators, culminating in an optimal rotationally invariant estimator that depends only on E in the large-dimensional limit.The methods include Coulomb-gas techniques, free probability, Bayesian approaches, eigenvector results, and resolvent-based analysis.
- 1.2. Historical survey: Eigenvalue-only inversion is insufficient because the Marčenko–Pastur equation does not describe empirical eigenvectors, which are not generally consistent estimators of population eigenvectors.This limitation motivates treating eigenvector statistics as essential to high-dimensional covariance estimation.
- 1.3. Outline: In portfolio applications, using E can overestimate realized out-of-sample risk by the factor (1−q)^-1 when q < 1.The review analyzes this effect in large-scale Markowitz optimization and studies the optimal RIE as an alternative.
2. Random Matrix Theory: overview and analytical tools
This section introduces Random Matrix Theory as a framework for analyzing large random matrices through limiting spectral behavior, rotational invariance, and analytical transforms. It develops the Coulomb-gas and Marčenko–Pastur perspectives while emphasizing that noisy covariance estimation also requires eigenvector information.
- 2.1.1. Large dimensional random matrices: Rotational invariance makes the matrix distribution unchanged under orthogonal conjugation and permits an eigenvalue–eigenvector representation involving the Vandermonde determinant.For real symmetric matrices, invariance is expressed as Pβ(M) = Pβ(ΩMΩ†) for Ω ∈ O(N).
- 2.1.1. Large dimensional random matrices: RMT studies large random matrices through universal limiting behavior, including deterministic limiting spectral densities that approximate finite matrices.The empirical spectral distribution can converge almost surely to a limiting spectral density, and self-averaging allows one large sample to represent that density.
- 2.1.2. Various RMT transforms: The resolvent is a continuous analytical object that contains complete information about a matrix’s eigenvalues and eigenvectors, while its pole residues project onto eigenspaces.This makes the resolvent useful for spectral analysis and for studying eigenvector statistics.
- 2.2.1. Stieltjes transform and potential function: The Coulomb-gas analogy models eigenvalues as logarithmically repelling particles in a confining potential, linking the potential function to the Stieltjes transform.A convex confining potential supports the one-cut assumption and prevents eigenvalue accumulation at the potential minimum.
- 2.2. Analytical examples: Characterizing the potential function can determine the limiting spectral density and reproduce diverse empirical spectral densities through suitable convex confining potentials.The Marčenko–Pastur law is obtained as a central example, with convergence behavior depending on the quality ratio q.
- 2.3. Free probability: Coulomb-gas methods describe uncorrupted spectra but are not directly useful for additively or multiplicatively noisy matrices when the corrupted-entry distribution is unknown.This limitation motivates additional RMT tools for separating true signals from noisy covariance observations.
N TrM, (2.61)
Free probability computes spectral laws and mixed moments for large random matrices, while replica methods extend resolvent analysis to eigenvector-sensitive, entrywise information.
- Asymptotic freeness applies when a fixed matrix is paired with a rotationally invariant random matrix in the large-N limit.This permits mixed-moment calculations and supports free-probability analysis of invariant matrix products.
- Free probability computes limiting spectral distributions of sums and products of invariant random matrices.Voiculescu’s free addition and multiplication formulas describe how noise modifies signal spectra.
- A rank-1 factor can create a spectral outlier when β > σ in the Wigner-noise example.The outlier criterion identifies when the largest eigenvalue lies outside the noise spectrum.
- For additive noise, eigenvalues outside the signal spectrum satisfy a criterion involving the signal resolvent, gA(z) = β^-1.This relation characterizes eigenvalues introduced by the noisy perturbation.
- Replica methods study resolvent entries and therefore eigenvector information that normalized-trace methods cannot capture.The replica identity rewrites the problem using multiple copies and yields entrywise asymptotic resolvent relations.
- The resulting resolvent relation is elementwise self-averaging, although the full resolvent matrix need not be deterministic.This distinction matters because entrywise behavior contains more spectral-decomposition information than the normalized trace alone.
3. Spectrum of large empirical covariance matrices
In the high-dimensional regime q = N/T of order unity, empirical covariance spectra systematically differ from population spectra, motivating RMT-based analysis and correction. The Marčenko–Pastur framework characterizes this distortion, including bulk spreading, outliers, and effects of heavy-tailed structure.
- High-dimensional setting: When q = N/T is of order unity, RMT provides precise statements about empirical covariance matrices beyond the classical q → 0 regime.The high-dimensional or Kolmogorov regime differs from fixed-N asymptotics, where the sample matrix converges to the population covariance.
- Marčenko–Pastur framework: The Marčenko–Pastur equation is a deterministic self-consistent relation between the Stieltjes transforms of the sample matrix E and population matrix C.Its right-hand side is deterministic when C is fixed, enabling analysis of the sample spectrum from the population spectrum.
- Spectral distortion: The sample spectrum is wider than the population spectrum: its variance equals the population variance plus q.This widening reflects measurement noise in the high-dimensional regime.
- Spectral distortion: Even when all population eigenvalues equal one, sample eigenvalues occupy [(1 −√q)^2, (1 + √q)^2], creating systematic estimation bias.The resulting deviations become significant when the sample size is comparable to the number of variables.
- Structured populations: Population structure can broaden the empirical spectrum further, with inverse-Wishart and power-law examples placing sample mass outside the population support.For inverse-Wishart covariance, the spectrum outside the population support is attributed to measurement noise; added structure can make the empirical distribution span nearly the positive real line.
- Outliers: A population spike separates from the Marčenko–Pastur bulk only when µ1 > 1 + √q; otherwise it merges at λ+ = (1 + √q)^2.Separated outliers converge to deterministic locations and fluctuate on the smaller T^-1/2 scale.
4. Statistics of the eigenvectors
The chapter studies how noise distorts eigenvectors of large sample covariance matrices and develops asymptotic overlap formulas, including comparisons between independently noisy estimates. Bulk eigenvectors carry little information about their population counterparts, whereas outlier eigenvectors retain structured alignment with population spikes.
- Sample eigenvectors raise two questions: their similarity to population eigenvectors and the information available from comparing noisy estimates.
- The analysis uses squared overlaps and the full resolvent to characterize eigenvector statistics in the high-dimensional limit.The resolvent is used instead of only its normalized trace.
- The bulk: Bulk sample eigenvectors are delocalized in the population basis, with rescaled overlaps of order O(1), while their direct projections onto corresponding true directions vanish as N grows.Consequently, non-outlier sample eigenvectors retain very little information about their corresponding population eigenvectors.
- Outliers: Outlier sample eigenvectors concentrate in cones around their parent population spikes and are delocalized in directions orthogonal to those spikes.Their overlap with non-parent population eigenvectors is only of order N^-1/2.
- Outliers: The population-spike/sample-spike overlap becomes progressively weaker as the population eigenvalue approaches the bulk threshold 1 + √q.
- Correlated sample covariance matrices: Overlap formulas for two noisy sample estimates self-average, can be evaluated without prior knowledge of the population spectrum, and quantify eigenvector stability when both samples share the same population matrix.Departures from the predicted self-overlap can indicate that the underlying population matrices differ.
5. Bayesian Random Matrix Theory
This section frames high-dimensional covariance estimation as a Bayesian decision problem and develops rotationally invariant estimators that use the observed sample spectrum while accounting for noisy eigenvectors.
- Motivation: In high dimensions, the sample estimator is inconsistent because its spectral density deviates from the true spectrum when q = O(1), while bulk eigenvectors are highly noisy.These effects make direct eigenvalue correction insufficient for accurately estimating C.
- Bayesian estimation: The proposed strategy minimizes a loss between an estimator Ξ(E) and the true matrix C, with Bayesian models encoding prior beliefs about C.The section specifically considers squared distance as a natural optimality criterion and expects mixed estimators to outperform classical ones in high dimensions.
- Bayesian estimation: Bayesian inference combines a prior distribution and likelihood into P(C,Y), tests the posterior P(C|Y), and yields the posterior mean as the MMSE estimator.The likelihood models the measurement process, while the prior represents beliefs about the structure of C.
- Bayesian estimation: For an isotropic prior, linear shrinkage takes the form Ξlin = αsE + (1 − αs)IN with αs = 1/(1 + 2qκ).Here αs lies in [0, 1] and κ > 0.
- Rotationally invariant estimators: Under a rotationally invariant prior, the estimator uses E's eigenvectors and replaces its eigenvalues with optimally chosen functions γi(Λ), generalizing linear shrinkage.The framework is motivated when no privileged directions are known, but the assumption is not optimal when E reveals non-trivial structure.
- Rotationally invariant estimators: In the large-N limit, the optimal shrinkage function depends on eigenvector-overlap statistics and the prior spectrum, yet can be estimated directly from E without explicitly choosing a prior.For N = 500, the reported agreement is excellent, and the framework reproduces linear shrinkage for an isotropic Inverse-Wishart matrix.
6. Optimal rotational invariant estimator for general covariance matrices
The optimal rotational invariant estimator constrains estimates to the sample eigenbasis, then converges in the large-dimensional limit to a deterministic nonlinear shrinkage function that no longer explicitly requires C. The same function applies to bulk eigenvalues and outliers, producing a narrower cleaned spectrum while accurately estimating diverse spectral configurations.
- 6.1. Oracle estimator: The oracle estimator is the best L2 estimate of C among positive-definite matrices diagonal in a predetermined eigenbasis U, typically the sample eigenbasis.
- 6.2. Explicit form of the optimal RIE: In the large-dimensional limit, the oracle estimator converges to a deterministic rotational invariant estimator that can be estimated directly from data without explicit knowledge of C.
- 6.2. Explicit form of the optimal RIE: The same universal nonlinear shrinkage function applies to both bulk eigenvalues and outliers as N →∞.
- 6.3. Some properties of the “cleaned” eigenvalues: The cleaning operation preserves the trace of the population matrix C and produces a spectrum narrower than C, which is itself narrower than the sample spectrum E.
- 6.3. Some properties of the “cleaned” eigenvalues: Small eigenvalues are enhanced, while large eigenvalues are reduced; asymptotically, shrinkage interpolates from λ/(1−q)^2 for small λ to λ−2q for large λ.
- 6.4. Numerical applications: The estimator accurately predicts bulk eigenvalues and outliers across isolated, deformed GOE, Toeplitz, and power-law spectral configurations.
7. Application: Markowitz portfolio theory and previous “cleaning” schemes
In the high-dimensional Markowitz setting, sample covariance estimation can severely underestimate realized risk, motivating covariance cleaning. The review compares rotationally invariant, clipping, shrinkage, and eigenvalue-substitution approaches, while noting limitations from temporal dependence, model assumptions, and eigenvector inconsistency.
- Markowitz portfolio risk: When N and T grow with finite q=N/T, sample-based portfolio weights can underestimate realized out-of-sample risk, becoming disastrous as q approaches 1.At q=1, in-sample risk is zero while out-of-sample risk diverges.
- Optimal RIE: The optimal RIE minimizes out-of-sample risk among rotationally invariant estimators under stated distributional assumptions and can maximize the Sharpe ratio.The oracle estimator is identified with the solution to the out-of-sample risk minimization problem.
- Previous cleaning schemes: Linear shrinkage and clipping regularize noisy spectra by shifting small eigenvalues upward, while clipping preserves the trace and replaces bulk eigenvalues with a common value.Linear shrinkage also pulls the largest eigenvalues downward; clipping uses a threshold based on the Marchenko-Pastur upper edge.
- Previous cleaning schemes: Estimating the inverse-Wishart parameter κ from Tr E^-1 is reliable only when κ is not too large; a two-sample test is offered as a more robust alternative.When C is close to the identity, the trace-based estimate can produce negative κ values.
- Previous cleaning schemes: Clipping can require an effective ratio qeff or tuning parameter αc because temporal autocorrelation and an inadequate identity null hypothesis distort the empirical spectral density.The corrected threshold is λ+ = (1 + √qeff)^2, and αc controls how many largest eigenvalues remain unchanged.
- Previous cleaning schemes: Eigenvalue substitution can approximately recover population eigenvalues by inverting the Marchenko-Pastur equation, but it is not optimal because sample eigenvectors are inconsistent.The recovered eigenvalues can instead be used to compute the optimal RIE.
8. Numerical Implementation and Empirical results
The section develops finite-sample implementations of RIE cleaning and evaluates them on synthetic and financial data. IW-regularization and QuEST correct small-eigenvalue bias and accurately approximate Oracle estimators, while RIE cleaning improves out-of-sample portfolio risk.
- Regularization: IW-regularization shifts small eigenvalues upward toward the Oracle estimator and corrects the downside bias that matters for covariance-matrix inversion.The method is motivated by an Inverse-Wishart model and can be further improved by sorting the regularized eigenvalues.
- Synthetic empirical studies: For N = 500 and q = 0.5, sorting improves IW-regularization accuracy significantly, while QuEST has the best large-N accuracy but no statistically significant advantage over IWs.The comparison uses 100 Wishart realizations and finds the performance improvement over the sample covariance matrix very substantial.
- QuEST: QuEST avoids systematic left-edge underestimation and handles outliers effectively in synthetic experiments, accurately estimating Oracle eigenvalues for finite N and q < 1.Its nonlinear, non-convex optimization yields little additional improvement over IWs as N increases, despite substantially higher computational cost.
- Financial applications: In financial data, IWs-regularized RIE and QuEST both accurately estimate the approximated Oracle estimator, with an effective ratio q_eff = 1.1q improving the European-stock estimate.The RIE tracks average realized risk closely for US stocks and shows similar conclusions for Japanese and European equities.
- Financial applications: Across portfolio tests, cleaned correlation matrices reduce out-of-sample risk, and regularized RIE generally gives the lowest risk except for one Japanese minimum-risk case.The RIE result is robust across dimensions N ≥ 300, apart from fluctuations at N = 100.
9. Conclusion and perspectives
The review closes by extending rotationally invariant covariance estimation toward temporal dependence, nonstandard distributions, rectangular matrices, and unresolved high-dimensional settings. It also summarizes empirical promise and identifies analytical tools and open problems for broader models.
- The review presents RMT techniques as useful for estimating large correlation matrices within a rotationally invariant framework, including real-world applications.
- Temporal correlations and temporal structure violate an important assumption of the sample covariance model in most real-life applications.The review proposes extending estimators to account for autocorrelations, including exponential autocorrelation models.
- EWMA covariance estimation makes older observations gradually less influential than more recent information through a constant α and time-series length T.
- Finite-fourth-moment assumptions behind the Marchenko–Pastur equation may fail in real datasets, especially finance, motivating more robust covariance estimates.
- Rectangular correlation matrices use singular values and associated vectors to identify strongly correlated linear combinations of input and output variables.The largest singular value identifies the strongest mutual correlation, while the corresponding vectors give predictor and predicted-variable weights.
- When N > T, the sample covariance matrix generically has N − T zero eigenvalues, creating an interpretation challenge for covariance estimation.
- Dyson’s Brownian Motion provides a physical interpretation of additive-noise eigenvalue and eigenvector dynamics and helps infer eigenvector overlaps.
A.1. Definitions and results.
This section defines the HCIZ integral over compact matrix groups and reviews exact and asymptotic results. The unitary case is explicit, while generalizations and soluble cases remain technically constrained.
- The HCIZ integral averages an exponential matrix expression over Haar measure on O(N), U(N), or Sp(N), with β equal to 1, 2, or 4.
- For the unitary case β = 2, the HCIZ integral has an exact finite-N determinant-ratio representation.
- The β = 1 and β = 4 expressions remain open problems in the reviewed treatment.
- Dyson’s Brownian motion yields the large-N asymptotic functional for β = 2, despite the difficulty of estimating alternating determinant terms.
- For arbitrary β, the asymptotic result satisfies Fβ(A, B) = βF2(A, B)/2.
- Explicit asymptotic solutions are available for Wigner matrices, a flat distribution, and one matrix of rank n ≪ N.The low-rank case is especially relevant to replica analyses with finitely many replicas.
A.2. Derivation of (A.5) in the Rank-1 case.
The rank-1 derivation evaluates the HCIZ integral by enforcing orthogonality, introducing a delta-function representation, and applying a saddle-point calculation. The argument extends to rank n ≪ N because orthogonality constraints become nearly automatic.
- The rank-1 case sets A = diag(a1, 0, …, 0) and rewrites the HCIZ integral with an explicit normalization.
- A delta-function representation enforces orthogonality constraints before the large-N integration is evaluated.
- The N →∞ integral over ζ is evaluated by a saddle-point method, producing the equation used in the derivation.
- The saddle-point solution is substituted back into the integral and checked by differentiating both sides, yielding the claimed result.
- For rank n ≪ N, the orthogonality constraint is nearly automatic because random unit vectors in N dimensions have scalar products of order 1/√N.Only the normalization constraint remains operative, allowing the integral to factorize into n independent rank-1 integrals.
- Schur-complement block identities provide determinant and inverse formulas, including Woodbury, determinant-lemma, Sylvester, and Sherman–Morrison identities.
B.3. Resolvent identities.
The resolvent identities derive recursive relations by expressing a block matrix through Schur complements. The construction applies to diagonal resolvent entries and extends to a two-dimensional block.
- The resolvent analysis rewrites H(z) as a block matrix and defines a Schur complement D = A − BC^-1B*.
- For K = 1, the scalar Schur complement gives an expression for a diagonal resolvent entry using a matrix minor.
- The diagonal-resolvent result holds for every other diagonal term of G.
- For K = 2, the same block representation produces identities involving a 2 × 2 Schur complement.
- The resulting relation applies to specified resolvent entries and supports a recursion relation for the resolvent.
C. Self-consistent relation for Green’s function and Central Limit Theorem
The review presents resolvent recursion as an RMT tool for deriving large-N spectral properties. Unlike Replica analysis, this approach permits non-identically distributed matrix entries and requires no ansatz.
- Resolvent recursion derives self-consistent equations for random-matrix spectral properties in the large-N limit.The approach uses recursion relations for a matrix resolvent together with a Central Limit Theorem argument.
- Unlike Replica analysis, the recursion technique allows non-identically distributed matrix entries.
- The recursion technique requires no ansatz to perform the calculations.
C.1. Wigner matrices.
For Wigner matrices, resolvent recursion and the Central Limit Theorem yield self-consistent relations whose normalized-trace form gives the semicircle law. The diagonal resolvent entries agree empirically with this prediction for large matrices.
- Wigner ensemble: The Wigner ensemble consists of symmetric matrices with iid entries and variance scaled as N^-1 so eigenvalues remain bounded as N grows.
- Resolvent calculation: Wick’s theorem and resolvent identities reduce the Wigner-matrix calculation to a large-N self-consistent relation.
- Resolvent calculation: The resulting off-diagonal resolvent entries scale as N^-1/2, while diagonal minors differ from diagonal entries by O(N^-1).
- Semicircle law: Taking the normalized trace yields the equation for the Stieltjes transform of the semicircle law.
- Empirical illustration: At N = 1000, the empirical imaginary resolvent estimate from one sample is compared with the theoretical curve and a confidence interval.The empirical estimate is shown in blue, the theoretical prediction in red, and the confidence interval with green dots.
- Empirical illustration: The GOE illustration shows excellent agreement between each diagonal resolvent entry and the semicircle law when N is large enough.
C.2. Sample covariance matrices.
The sample-covariance analysis uses a block resolvent to handle products of rectangular matrices and derives large-dimensional relations for both N × N and T × T formulations. The T × T formulation is often simpler because its resolvent can be approximated by a normalized trace.
- Block-resolvent setup: A block matrix of size (N + T) × (N + T) is introduced because the sample covariance matrix is a product of rectangular matrices.
- Block-resolvent calculation: The block-resolvent entries are analyzed using resolvent identities, Wick’s theorem, and the Central Limit Theorem.
- Large-system behavior: For distinct indices, off-diagonal blocks scale as T^-1/2 in the large-system analysis.
- Resolvent relations: The derivation connects the block-resolvent relations to the sample-covariance resolvent through gS(z) = qgE(z) + (1 − q)/z.
- Resolvent relations: The T × T sample covariance matrix S is often easier to analyze because its resolvent can be approximated by its normalized trace.
- Additive noise: The appendix also treats additively corrupted matrices, distinguishing the measured sample matrix M from the population matrix C.
D.1.2. Dyson Brownian Motion.
The review analyzes additive rotationally invariant noise through Dyson Brownian motion, resolvent evolution, free probability, and Replica methods. These approaches yield deterministic large-N resolvents in the basis of the population matrix and recover the free-addition relation.
- Dyson Brownian Motion: Dyson Brownian motion represents additive noise as a time-evolving symmetric matrix whose variance increases with fictitious time.
- Dyson Brownian Motion: The eigenvalues follow a Coulomb-gas dynamics, while the associated eigenvectors satisfy a conditional stochastic differential equation.
- Resolvent evolution: An alternative approach directly evolves the full resolvent using Itô calculus.
- Resolvent evolution: In the large-N limit, the resolvent evolution reduces to a matrix PDE, whose trace gives a Burgers equation for the Stieltjes transform.
- Resolvent evolution: For Gaussian orthogonal noise, the Stieltjes-transform solution is equivalent to the additive model with variance σ^2t.
- Free probability and Replica methods: Free probability and Replica calculations recover the free-addition formula and show that each resolvent entry converges to a deterministic quantity in the basis of C.The additive case is described as simpler than the multiplicative case.
D.3. Overlap and Optimal RIE formulas in the additive case.
The additive-noise analysis derives overlap formulas and an optimal rotationally invariant eigenvalue-cleaning rule. Numerical examples show agreement with theory and recover physically consistent spectra while preserving the trace.
- D.3. Overlap and Optimal RIE formulas in the additive case.: The resolvent of the additive model converges entrywise to a deterministic limit, which simplifies in the eigenbasis of C.This simplification enables explicit mean squared overlaps between sample and true eigenvectors.
- D.3.1. Mean squared overlaps.: For GOE noise with entry variance σ2/N, the transform reduces to Z(z) = z − σ2gM(z), yielding a simpler overlap formula.The resulting expression matches the earlier GOE-specific result.
- D.3.1. Mean squared overlaps.: The theoretical overlap prediction agrees remarkably with simulations generated from 200 noisy samples at N = 500 and T = 1000.The experiment uses a fixed Wishart signal and GOE additive noise with variance 1/N.
- D.3.2. Optimal RIE.: The asymptotic oracle estimator for bulk eigenvalues can be computed explicitly from the overlap formulas and trace identities.The derivation uses the resolvent relation involving Z(z) and gM(z).
- D.3.2. Optimal RIE.: The optimal additive-noise RIE keeps the eigenvectors of M and applies a nonlinear shrinkage function to its observable eigenvalues.The non-observable oracle estimator converges as N →∞ to a deterministic function of those eigenvalues.
- D.3.2. Optimal RIE.: In the GOE-signal and GOE-noise case, optimal cleaning rescales empirical eigenvalues by the signal-to-noise ratio, as in the Wiener filter.The cleaned empirical spectral distribution is narrower than the true one.
- D.3.2. Optimal RIE.: For a white Wishart signal, GOE noise can push eigenvalues negative, whereas optimal cleaning returns the cleaned spectrum to positive values.The corresponding Stieltjes transform is obtained from a cubic equation, using the solution with non-negative imaginary part.
- D.3.2. Optimal RIE.: Figure D.3 compares the optimal cleaning formula with naive eigenvalue substitution and shows that optimal cleaning narrows eigenvalue spacings.The comparison uses the same parameters as Figure D.2.