Source-linked AI summary
Sample eigenvalue based detection of high dimensional signals in white noise using relatively few samples
N. Raj Rao, Alan Edelman
TL;DR
The paper addresses signal-number estimation when high dimensionality and few samples make classical sample-eigenvalue methods unreliable. It develops a mathematically justified estimator based on large random matrix theory, identifies an effective number of signals, and uses simulations to assess consistency and finite-sample behavior.
Problem
Classical eigenvalue-based estimators lack reliable justification when the empirical covariance is singular, while weak or closely spaced signals may be asymptotically unidentifiable from sample eigenvalues.
Method
The paper develops a computationally simple sample-eigenvalue estimator using moments and distributional results from large random matrix theory.
Results
Simulations show consistency for the true signal count in the fixed-dimensional, large-sample limit and for the effective signal count in the large-system limit.
Takeaways & Limitations
The effective number of signals explains why underestimation is unavoidable when too few samples make some signals unidentifiable by sample-eigenvalue methods.
Abstract
from arXiv · showhide
We present a mathematically justifiable, computationally simple, sample eigenvalue based procedure for estimating the number of high-dimensional signals in white noise using relatively few samples. The main motivation for considering a sample eigenvalue based scheme is the computational simplicity and the robustness to eigenvector modelling errors which are can adversely impact the performance of estimators that exploit information in the sample eigenvectors. There is, however, a price we pay by discarding the information in the sample eigenvectors; we highlight a fundamental asymptotic limit of sample eigenvalue based detection of weak/closely spaced high-dimensional signals from a limited sample size. This motivates our heuristic definition of the effective number of identifiable signals which is equal to the number of "signal" eigenvalues of the population covariance matrix which exceed the noise variance by a factor strictly greater than 1+sqrt(Dimensionality of the system/Sample size). The fundamental asymptotic limit brings into sharp focus why, when there are too few samples available so that the effective number of signals is less than the actual number of signals, underestimation of the model order is unavoidable (in an asymptotic sense) when using any sample eigenvalue based detection scheme, including the one proposed herein. The analysis reveals why adding more sensors can only exacerbate the situation. Numerical simulations are used to demonstrate that the proposed estimator consistently estimates the true number of signals in the dimension fixed, large sample size limit and the effective number of identifiable signals in the large dimension, large sample size limit.
I. INTRODUCTION
The paper develops a mathematically justified, computationally simple sample-eigenvalue estimator for signal detection in high-dimensional, sample-starved settings. It also identifies an asymptotic limit on detecting weak or closely spaced signals and evaluates consistency in classical and large-system regimes.
- Motivation: High-dimensional, sample-starved settings can make classical eigenvalue-based estimators inapplicable because the empirical covariance matrix is singular.Ad-hoc modifications lack rigorous theoretical justification, making it unclear whether underestimation reflects a fundamental detection limit.
- Contribution: The paper develops a mathematically justified, computationally simple signal detector based on sample eigenvalues.The approach is designed to operate effectively when samples are limited.
- Identifiability: The effective number of identifiable signals captures a fundamental sample-size-dependent limit for detecting weak or closely spaced signals from sample eigenvalues.The concept is grounded in large random matrix theory and a threshold depending on noise variance, dimensionality, and sample size.
- Results: In the fixed-dimensional, large-sample limit, both the proposed and Wax-Kailath MDL estimators consistently estimate the true number of signals.In the large-system limit, simulations suggest consistency for the proposed estimator but not for the MDL estimator with respect to the effective number of signals.
- Results: The simulations also suggest that the proposed estimator remains applicable in moderate-dimensional settings.This evidence is numerical rather than a general consistency theorem.
II. PROBLEM FORMULATION
The paper formulates signal-number estimation from independent Gaussian snapshot vectors containing low-rank signals and unknown-variance white noise. It revisits classical information-criterion methods because they become degenerate or unjustified when samples are fewer than sensors, motivating a new sample-eigenvalue estimator with an explicit high-dimensional limitation.
- Model: The observation model combines a finite-dimensional Gaussian signal through an unknown mixing matrix with independent Gaussian white noise of unknown variance.The signal covariance is nonsingular and the mixing matrix is assumed to have full column rank.
- Problem: Estimating the number of signals is posed as model selection from m samples because the population covariance is unknown.If the covariance were known, k could be obtained from the multiplicity of its smallest eigenvalue.
- Prior methods: Wax-Kailath AIC is inconsistent, whereas its MDL form is consistent when dimensionality is fixed and sample size tends to infinity.The classical estimator relies on sample eigenvalues and assumes m>n.
- High-dimensional limitation: When m<n, the sample covariance matrix is singular and the classical estimators become degenerate.Restricting the candidate model order to k<min(n,m) is an ad-hoc modification without rigorous theoretical justification.
- Proposed method: The proposed estimator uses sample eigenvalues with computational complexity comparable to the modified Wax-Kailath estimator.Its derivation assumes the number of signals is much smaller than the system size, k≪n.
- Identifiability: Sample-eigenvalue methods cannot reliably detect sufficiently weak or closely spaced signals when too few samples are available.In that regime, the proposed estimator and other sample-eigenvalue detectors consistently underestimate the signal count.
III. PERTINENT RESULTS FROM RANDOM MATRIX THEORY
The paper uses large random matrix theory to characterize how sample eigenvalues spread around population eigenvalues when system dimension and sample size grow together. These results provide the mathematical basis for a sample-eigenvalue estimator in high-dimensional, sample-starved settings.
- Random-matrix basis: The joint distribution of sample eigenvalues is difficult to analyze directly because its exact form contains a multidimensional orthogonal or unitary-group integral.Classical fixed-dimensional asymptotics establish consistency, but do not describe high-dimensional eigenvalue spreading.
- Estimator rationale: The proposed estimator explicitly exploits analytical results describing the spreading of sample eigenvalues.This gives its use in high-dimensional, sample-starved settings a stronger mathematical basis than approaches relying on classical sample-eigenvalue consistency.
- Estimator rationale: Ad-hoc use of the equality between nonzero eigenvalues of bR and (1/m)X′X does not mathematically justify modifying classical estimators when m<n.The paper specifically challenges modifications based only on the nonzero sample eigenvalues.
- Asymptotic regime: Large random matrix theory describes eigenvalue distributions when n,m→∞ while n/m→c∈(0,∞).This regime differs from the classical n fixed, m→∞ asymptotic setting.
A. Eigenvalues of the signal-free SCM
For a signal-free Gaussian sample covariance matrix, the empirical eigenvalue distribution converges to the Marčenko-Pastur law in the proportional-dimensional limit. Its spread depends on the dimensionality-to-sample-size ratio, and the eigenvalue moments provide additional asymptotic information for estimator construction.
- Limiting distribution: The empirical eigenvalue distribution of a signal-free sample covariance matrix converges almost surely to the Marčenko-Pastur distribution when n,m→∞ with n/m→c.The result assumes independent Gaussian samples with variance λ=σ2.
- Limiting distribution: For noise variance 1, eigenvalues increasingly cluster around 1 as c→0, while spreading is significant for modest c.The density is illustrated for different values of c=n/m.
- Moments: Almost-sure convergence of the empirical distribution implies almost-sure convergence of eigenvalue moments to the corresponding Marčenko-Pastur moments.Finite-dimensional sample moments fluctuate around these limiting values.
- Estimator statistics: The paper uses the large-dimensional eigenvalue-moment results to construct a statistic whose distribution is independent of the unknown noise variance.The statistic is analyzed using the limiting moment behavior and the delta method.
B. Eigenvalues of the signal bearing SCM
In high-dimensional sample covariance matrices, signal eigenvalues exhibit a phase transition: only those above a threshold separate from the noise-only limit. The corresponding asymptotic behavior is supported by large random matrix results.
- As n,m→∞ with n/m→c, the noise eigenvalues retain the limiting empirical distribution because a fixed number of signal eigenvalues has vanishing weight.
- A signal eigenvalue changes the limiting largest sample eigenvalue only when it exceeds a threshold, producing a phase transition.
- Proposition 3.4 characterizes the almost-sure limits of the largest sample eigenvalues associated with population signal eigenvalues in the large-system regime.
IV. ESTIMATING THE NUMBER OF SIGNALS
The estimator selects the signal count through an information-theoretic criterion built from sample-eigenvalue moments, including singular covariance matrices. Its score is derived by approximating the noise-eigenvalue statistic with a normal distribution.
- The proposed estimator uses distributional properties of traces of powers, or eigenvalue moments, of large-dimensional Wishart sample covariance matrices.
- With unknown noise variance, the model for k signals has k+1 free parameters, and the noise variance is estimated from the ordered sample eigenvalues.
- The test statistic q_k is approximated as normally distributed because the n−k noise eigenvalues resemble those of a signal-free sample covariance matrix.
- Substituting the dimensionality-to-sample ratio n/m into the information criterion yields the estimator in (9), whose score function is illustrated in Figure 3.
V. EXTENSION TO FREQUENCY DOMAIN AND VECTOR SENSORS
The sample-eigenvalue estimator extends from time-domain covariance matrices to frequency-domain spectral estimates and quaternion-valued vector-sensor data. For wideband signals, criteria are combined across frequency bins under an independence assumption that may fail with severely limited snapshots.
- Frequency domain: For Fourier coefficient vectors, the sample covariance matrix becomes a periodogram estimate, and the estimator remains applicable to frequency-dependent eigenvalues.
- Frequency domain: Under the SPLOT assumption, Fourier coefficients at different frequencies are statistically independent when observation time greatly exceeds signal correlation times.
- Frequency domain: For wideband signals occupying M frequency bins, the detection criterion is obtained by summing the corresponding single-frequency criterion across those bins.
- Limitation: When snapshots are severely constrained, violation of the SPLOT assumption can make frequency coefficients dependent and likely degrade estimator performance.
- Vector sensors: For quaternion-valued narrowband measurement vectors, the estimator applies with β=4, including data representations from vector sensors.
VI. CONSISTENCY OF THE ESTIMATOR AND THE EFFECTIVE NUMBER OF IDENTIFIABLE SIGNALS
The proposed estimator is consistent for the true signal count in the fixed-dimension, large-sample regime and is conjectured to estimate only the effective identifiable count in the large-system regime. Signals below the random-matrix threshold are asymptotically indistinguishable from noise for sample-eigenvalue detection.
- Classical consistency: The proposed estimator is conjectured to be consistent for the true number of signals when n is fixed and m→∞.
- Effective identifiability: Signals with eigenvalues between the noise level and λ(1+√c)^2 have the same asymptotic detection behavior as omitted signals.
- Limitations: The asymptotic equivalence result provides no convergence-rate information, and the rate may depend on the covariance eigenvalue structure and be arbitrarily slow.
- Limitations: Consistency in the large-system regime remains unproved because it requires refined analysis of fluctuations among ordered noise eigenvalues; simulations provide non-definitive corroboration.
- Effective identifiability: The effective number of identifiable signals counts population covariance eigenvalues exceeding λ(1+√c)^2 in the large-system ratio c=n/m.
- Effective identifiability: In the large-system, large-sample limit, the estimator is conjectured to consistently estimate the effective rather than necessarily the actual signal count.
A. The asymptotic identifiability of two closely spaced signals
The paper analyzes when two closely spaced signals can be identified from sample eigenvalues alone, linking identifiability to system dimensionality, sample size, and signal-vector similarity. It also examines finite-sample eigenvalue fluctuations and uses ZSep to characterize detection reliability.
- The effective number of signals determines when both signals can be reliably detected asymptotically from sample eigenvalues alone.The analysis considers two uncorrelated signals and expresses their effective number through the population covariance eigenvalues.
- Equation (38) captures the tradeoff among identifying two closely spaced signals, system dimensionality, available snapshots, and the cosine of the angle between signal vectors.
- Finite dimensionality and sample size cause signal and noise eigenvalues to fluctuate around limiting positions, affecting discrimination from the largest noise eigenvalue.These fluctuations are illustrated in Figure 4.
- ZSep measures the theoretical separation of a signal eigenvalue from the largest noise eigenvalue in standard deviations of signal-eigenvalue fluctuations.
- Simulations suggest reliable detection of the effective number of signals when ZSep_j exceeds a range of approximately 5–15.The broad range indicates that finite-system performance depends on interactions between signal and noise eigenvalues.
VII. NUMERICAL SIMULATIONS
Monte-Carlo simulations compare the proposed estimator with Wax-Kailath MDL across system and sample sizes. The proposed method detects signals with fewer samples, tracks the effective identifiable count in large-system regimes, and can converge arbitrarily slowly.
- Simulation setup: Over 4000 Monte-Carlo simulations compare the new sample-eigenvalue estimator with the modified Wax-Kailath MDL estimator across n and m.The simulations use two signal eigenvalues, λ1 = 10 and λ2 = 3, with unit noise eigenvalues.
- Classical consistency: For n = 32 and n = 128, both estimators eventually detect two signals with high probability, but the new estimator requires significantly fewer samples.This evaluates the classical n fixed, m →∞ consistency regime.
- Effective signal count: The effective signal count matches regimes where the new estimator detects one signal with high probability, suggesting relevance beyond the asymptotic setting.The correspondence is observed using Figure 6 values of k_eff across the n and m settings considered.
- Severely sample-starved regime: With fewer than 10 samples, the new estimator often detects zero signals even when k_eff = 1, showing that the asymptotic prediction is unreliable when m ≪ n.The authors caution that signal identifiability predictions should not be expected to remain accurate in severely sample-starved settings.
- Large-system regimes: When m = 4n, the proposed estimator consistently detects two signals while Wax-Kailath MDL does not; when m = n/4, it consistently estimates k_eff = 1.For m = n/4, the second signal eigenvalue λ2 = 3 lies exactly on the asymptotic identifiability threshold.
- Convergence behavior: The rate of convergence can be arbitrarily slow and cannot be entirely explained by the separation metric ZSEP.Tables II-(c) and II-(d) provide additional evidence for this observation.
VIII. CONCLUDING REMARKS
The paper develops a sample-eigenvalue-only information-theoretic estimator for signal-count detection in white noise. It accounts for finite-size eigenvalue blurring, while several consistency and testing questions remain open.
- Contribution: The proposed information-theoretic estimator detects the number of signals from sample eigenvalues alone while explicitly accounting for finite-size blurring.The approach is intended for white-noise settings with relatively few samples.
- Open theory: The algorithm’s consistency conjecture remains to be proven in both the n fixed, m →∞ and n, m(n) →∞ with n/m(n) → c regimes.These are the two asymptotic settings discussed in the conclusion.
- Future work: Future work will extend signal-count estimation to arbitrary covariance when an independently estimated noise covariance is available from relatively few samples.The planned estimator will retain the form in (9) and use trace-of-powers results for structured Wishart matrices.
- Open testing problem: Neyman-Pearson analysis of the most powerful test under a false-detection constraint remains open, especially near the threshold for low-level signals.The authors note that nested hypothesis tests might improve performance in this regime.