Source-linked AI summary

Compact and Infinite-Order Error Analysis for Null-Space SVD Estimation

Xin Li, Jonathan Cohen, Rami Puzis

arXiv:2608.30374v1math.STcs.LG

TL;DR

The paper asks how noisy SVD estimates of null spaces behave across perturbation orders, risks, and retained subspaces. It derives compact and infinite-order SVD and risk formulas, extends them to invariant subspaces, and computes convergence from exceptional points. The results establish distinct empirical and generalization rankings, spectral convergence boundaries, and separate finite-sample diagnostics from the asymptotic BBP threshold.

  • Problem

    Null-space estimation requires understanding exact estimator errors, all-order behavior, risk rankings, and convergence limits across simple and multiple nullities.

  • Method

    The paper derives an exact compact simple-nullity solution, Taylor recursions for vectors, projectors, and risks, block invariant-subspace recursions, and exceptional-point convergence analysis.

  • Results

    Strict expected second-order empirical ranking follows from the Wishart splitting matrix; isotropic subspaces have strict expected generalization ranking, while unequal spikes are addressed by an overlap criterion and a simultaneous 99% confidence certificate.

  • Takeaways & Limitations

    The equal–ranked–equal phenomenon is a finite-sample spectral-mixing diagnostic, and tolerance crossings, exceptional-point radii, and the BBP threshold must be treated as distinct quantities.

  • Takeaways & Limitations

    For unequal-spike matrices, the overlap condition is sufficient, and endpoint equalization alone does not imply its hypotheses, unimodality, or a single connected visibility interval.

Abstract

from arXiv · show

We study null-space estimation from a noisy matrix. For a simple left null space, we first derive an exact compact expression for the error of the smallest left singular vector. We then give an all-order series for the SVD vector and projector, followed by compact and consistently truncated series forms for the fixed-realization empirical risk and conditional population generalization risk. The recursion extends to a multiple-dimensional null space by following the complete invariant subspace. The convergence radius is not inferred from an error plot: it is computed independently from the nearest complex exceptional point that joins a retained eigenvalue branch to its complement. A reduced-nullity experiment shows that moving this spectral boundary can increase the radius, although the improvement is not monotone in the retained nullity. For individually ordered null directions under Gaussian training with \(τ\geq m\), we prove that the Wishart splitting matrix \(W\) gives a strict second-order empirical ranking. Gaussian averaging equalizes the leading generalization risks at both small and very large noise, while a column-swap theorem proves strict expected generalization ranking for an isotropic signal subspace. For unequal spikes, an exact population-overlap criterion and a simultaneous \(99\%\) Monte Carlo confidence certificate explain the observed intermediate ranking. A sixth-order risk correction improves the lower-crossover estimate in the reported experiment. This equal--ranked--equal phenomenon is a finite-sample diagnostic related to spectral mixing, but its tolerance crossings, the exceptional-point radius, and the asymptotic BBP threshold are three distinct quantities.

1 Contributions

The paper develops exact and all-order tools for noisy null-space SVD estimation, then analyzes risk rankings, convergence boundaries, and numerical validation. Its results distinguish empirical and generalization behavior across noise regimes and signal structures.

  • An exact compact reconstruction of the smallest left singular vector leads to arbitrary-order vector and projector recursions and consistently truncated empirical and population-risk expansions.The framework covers simple nullity and extends to block and basis-free formulations.
  • Empirical branch risk equals the corresponding ordered sample eigenvalue divided by τ, while the Wishart splitting matrix W gives strict expected second-order ranking under Gaussian training with τ ≥ m.For generalization risk, Gaussian training yields common small- and large-noise endpoint laws, exact strict ranking for isotropic signal subspaces, and an overlap criterion for unequal spikes.
  • The projector-series convergence disk is set by the nearest complex exceptional point joining a retained eigenvalue sheet to its complement.Changing retained nullity moves the spectral boundary and can expose a different collision, so the radius may increase or decrease.
  • Numerical experiments verify the compact identities, arbitrary-order convergence, exceptional-point radii, and the nonmonotonic effect of reducing nullity.For an unequal-spike example, a simultaneous 99% Monte Carlo certificate verifies intermediate-noise overlap inequalities and expected ranking, while fourth-, sixth-, and inverse-noise corrections predict crossings.

2 Problem setting

The problem estimates a clean left null space from a noisy data matrix with a fixed noise realization. The complete null space is represented by its orthogonal projector, while reduced retained nullity selects only part of it.

  • The clean data matrix is Z = HX, and the noisy observation is eZ(σ) = Z + σE with E a fixed standardized unit-variance noise realization.Here H, X, and E have dimensions and roles specified by the model.
  • Under full row rank of X and rank(H) = n, the clean rank is r = n and the measurement-based and system-based left null spaces coincide.The clean nullity is q0 = m − n.
  • The q smallest left singular vectors of eZ define the estimated subspace, whose invariant object is the orthogonal projector because bases may rotate internally.This formulation avoids treating an arbitrary basis as the primary estimator.
  • Choosing q = q0 estimates the complete clean null space, whereas q < q0 is called reducing the retained nullity.
  • The paper distinguishes conditional scalar risks for one finite training realization from their expectations over training realizations.These are denoted Remp,q and Rgen,q conditionally, and εemp,q and εgen,q after averaging.

3 Exact eigenvalue problem and compact solution

The noisy SVD estimator is recast as a symmetric eigenvalue problem. For simple nullity, an exact compact system and scalar closure characterize the eigenvalue branch descending from the clean zero eigenvalue and its oriented eigenvector.

  • The estimator is equivalent to the symmetric eigenvalue problem for A = (Z + E)(Z + E)^T, with A0 = ZZT, A1 = ZET + EZT, and A2 = EET.
  • For q0 = 1, Proposition 1 gives an exact compact representation using Q and an equivalent scalar closure whose roots are eigenvalues.The compact construction applies where Q is invertible.
  • The compact system also yields the normalized eigenvector and error, with the preceding equivalence proving that the scalar closure and oriented eigenpair describe the same eigenspace.
  • The oriented eigenpair is characterized by Qbη = βη, λβ = ηTE(Z + E)T bη, bηT bη = 1, and β > 0.The orientation convention selects the plus branch.
  • The scalar closure has a unique analytic root near (σ, λ) = (0, 0), corresponding to the eigenvalue branch descending from the clean zero eigenvalue.Analyticity follows from the implicit-function theorem because Q(0, 0) = Im and ∂λF(0, 0) = 1.

4 Infinite-order recursion

The paper expands the exact compact solution into sequential all-order Taylor recursions for eigenvalues, vectors, errors, and projectors. For multiple nullity, it propagates the complete invariant subspace and uses a basis-free projector recursion.

  • 4.1 Simple-nullity recursion: The arbitrary-order coefficients are obtained by coefficientwise Taylor expansion of the exact closed compact system, rather than by an independent construction.At each order, previously computed coefficients determine the next eigenvalue, perpendicular vector component, normalization coefficient, and full vector.
  • 4.1 Simple-nullity recursion: The order-K Taylor approximation is formed by truncating the eigenvector error series after computing its coefficients sequentially.
  • 4.1 Simple-nullity recursion: The first-order error is the least-squares expression, and the second-order expansion adds a normalization correction.
  • 4.2 Block recursion for multiple nullity: For q0 > 1, the recursion follows a moving orthonormal frame spanning the noisy invariant subspace descending from the clean null space.The restriction matrix need not be diagonal in the chosen gauge, and the gauge fixes arbitrary internal rotations.
  • 4.2 Block recursion for multiple nullity: The first nonzero splitting inside the clean null space occurs at second order, selecting the small-noise branches used by reduced SVD targets.
  • 4.2 Block recursion for multiple nullity: The basis-free projector recursion starts with P1 = −GA1P0 − P0A1G and is the natural coordinate-free comparison with the exact SVD projector.It uses projector identities and commutators to generate higher-order terms.

5 Finite-sample empirical risk

The section develops fixed-realization empirical-risk identities from SVD eigenvalues and shows that Gaussian training produces a strict ordered-branch ranking through the Wishart splitting matrix W.

  • For simple nullity, empirical risk equals the smallest eigenvalue branch divided by τ, making the compact eigenvalue reconstruction exact.
  • The empirical-risk series is a fixed-realization expansion; averaging it term by term requires a separate domination or uniform-convergence argument.
  • For every retained rank, ordered branch risks decompose the aggregate empirical risk before and after training expectation.
  • The first nonzero small-noise separation satisfies λ^(i)(σ)/σ2 → µ_i, where the µ_i are ordered eigenvalues of W.When W is simple, these branches identify the locally analytic sheets.
  • Under Gaussian training with τ ≥ m, W has a positive simple Wishart spectrum almost surely, so expected empirical branch risks are strictly ordered.The strict ordering follows because ordered eigenvalue gaps are positive almost surely and integrable.

6 Finite-sample generalization risk

The section derives compact and consistently truncated conditional generalization-risk forms, then analyzes ordered expected risks across noise regimes and signal spectra.

  • 6.1 Conditional compact and series forms: The compact scalar generalization-risk identity follows by substituting the compact singular-vector formula and is exact whenever that eigenvector formula is exact.
  • 6.1 Conditional compact and series forms: An order-K risk approximation retains only coefficient products of total degree at most K, preventing inconsistent comparison with exact SVD risk.
  • 6.1 Conditional compact and series forms: Termwise averaging of fixed-realization risk series requires a common convergence disk and integrable coefficient bounds; unbounded Gaussian noise does not automatically satisfy this.
  • 6.2 Expected ordered-branch generalization: At small noise, every ordered branch has the same leading coefficient, while rank dependence first appears in the O(σ4) term.
  • 6.2 Expected ordered-branch generalization: At large noise, Gaussian averaging again equalizes the leading law, with rank dependence first appearing in the t2 coefficient and normalized gaps O(ζ−4).
  • 6.2.1 Why the intermediate generalization risks are ranked: For an isotropic signal subspace and τ ≥ m, all ordered expected generalization risks are strictly increasing for every finite σ > 0.
  • 6.2.1 Why the intermediate generalization risks are ranked: For unequal spikes, cumulative population-overlap inequalities are sufficient for risk ordering, but endpoint equalization alone does not guarantee them or a single visibility interval.The reported experiment supplies a simultaneous 99% Monte Carlo certificate at one representative middle-noise point.

7 Convergence radius and exceptional points

The convergence radius is determined algebraically by the nearest exceptional point separating retained and complementary eigenvalue sheets, independently of numerical error curves.

  • A norm bound is a sufficient small-noise check, not the exact convergence radius.
  • Analytic continuation in complex noise scale identifies exceptional-point candidates where eigenvalue sheets meet.
  • Only collisions joining a retained sheet to its complement limit the grouped spectral projector; internal cluster collisions do not.
  • The radius ρ_q is computed independently from the projector-error curve, which serves only as numerical validation.
  • The projector series converges for |s| < ρ_q, has truncation error O(|s|K+1), and cannot be made convergent beyond ρ_q by increasing K.
  • The scalar generalization-risk radius satisfies ρ_gen,q ≥ ρ_q, with equality generically unless the risk functional cancels the projector singularity.
  • In the reported experiment, ρ_gen,5 = 1.24437996, and the Taylor series diverges at σ = 2 because 2 > ρ_gen,5.

8 Reduced nullity

Reducing the retained nullity changes which eigenvalue branches form the spectral boundary, so the convergence radius can increase or decrease without a general monotonic pattern. The reduced target is a different lower-dimensional subspace, not a repair of the full-nullity projector.

  • Reduced-nullity construction: W is diagonalized to label the zero-origin branches and select the bottom-q SVD subspace when the boundary eigenvalues are separated.For 1 ≤ q < q0, the condition µq−1 < µq gives a unique small-noise limit.
  • Numerical experiment: The full-nullity experiment plots conditional generalization-risk series convergence at σ = 10−3 and the order-120 error δ[120].
  • Numerical experiment: The order-120 conditional generalization-risk error is shown for each retained nullity alongside independently computed projector radii ρq.Observed growth at the same boundaries provides numerical evidence that no plotted scalar risk cancels the limiting projector singularity for this realization.
  • Spectral boundary: Reducing nullity moves one branch between the retained cluster and its complement, potentially changing which collision limits the projector series.An old limiting collision can become internal, while a previously internal collision can become a new boundary collision.
  • Interpretation: The resulting convergence range belongs to a different lower-dimensional target and does not repair the original full-nullity projector.

9 Numerical verification

Numerical experiments verify the compact identities, arbitrary-order projector recursion, risk-ordering claims, and independently computed exceptional-point convergence radii. Reduced-nullity series can converge where the full-nullity series fails, but the radius improvement is nonmonotone.

  • 9.1 Compact formula: The compact reconstruction matches the direct SVD calculation, with the comparison confirming the formula and the need for M = GEET.Table 2 reports zero normalization error to machine precision and cond(Q) = 1.31077.
  • 9.2 Infinite-order convergence: At σ = 10−3, projector-series errors fall from 3.7381×10−7 at order 1 to 1.4536×10−15 at order 120, with four orders reaching floating-point accuracy.The block recursion converges to the exact full-nullity SVD projector.
  • 9.3 Ordered-branch generalization regimes: Expected empirical risks are exactly ordered, while generalization risks are nearly equal at both noise endpoints and visibly separated at intermediate noise.The empirical second-order separation is governed by W; the unequal-spike intermediate ranking is supported by overlap-tail estimates and a simultaneous 99% confidence certificate.
  • 9.3 Ordered-branch generalization regimes: Including the fourth-order correction moves the lower crossing prediction from 0.3201 to 0.3101, while the reported endpoint errors are 0.04% and 1.87%.The calibration also uses a high-noise inverse-power law and compares measured with predicted crossings of δvis = 10−3.
  • 9.4 Exceptional-point and reduced-nullity results: The full-nullity radius is ρ5 = 1.24437996, increasing to 1.27907567 for q = 4 and reaching ρ2 = 1.40126671, but the improvement is nonmonotone.At σ = 1.26, the q = 4 series converges whereas the original q = 5 Taylor series does not; the radius is determined by boundary-crossing exceptional points.

10 Relation to the BBP transition

The paper relates finite-sample spectral separation to BBP-type rank transitions while distinguishing the corresponding thresholds by definition and asymptotic setting.

  • Spectral-branch connection: The reduced-nullity experiment and BBP theory both track a selected spectral branch while it remains separated from its complement.Changing the selected rank changes which branch crossing is relevant, potentially producing sequential rank transitions.
  • Distinct thresholds: The exceptional-point radius ρq is a finite-dimensional Taylor radius along a fixed noise direction, whereas BBP is a real-axis phase transition after dimensions grow.Visibility endpoints instead depend on an ensemble-averaged statistic and a chosen tolerance.
  • Distinct thresholds: The visibility endpoints depend on tolerance and are not phase transitions, despite being consistent with finite-sample signal–noise mixing.BBP theory also depends on the boundary of the selected rank.
  • Scope of the connection: An asymptotic connection would require a suitable statistic and a proof over increasing dimensions and many realizations.The present experiments provide finite-sample mechanisms related to BBP-type transitions but do not prove an asymptotic identity.
  • Scope of the connection: The reported finite-realization radius differs from the asymptotic BBP threshold because the realization lies outside a dimension-growing experiment.The paper explicitly contrasts ρ5 ≈ 1.24438 with the BBP notion.

11 Conclusion

The conclusion combines exact fixed-realization SVD expansions with risk-ordering results, convergence analysis, and finite-sample qualifications. It finds strict empirical ordering, endpoint-equalized expected generalization, and distinct spectral thresholds.

  • Conclusion: The compact formula is exact for simple nullity, while block propagation and the matrix W handle multiple nullity and reduced SVD targets.The expected-risk coefficient series requires interchangeability of expectation and Taylor expansion.
  • Conclusion: Gaussian ordered branches share the leading low-noise generalization law, with branch dependence first entering through fourth-order projector coefficients.The normalized expected generalization profile is asymptotically uniform at both noise endpoints.
  • Conclusion: Empirical branch risks are exactly ordered at every noise level, while expected generalization ordering is strict for isotropic signals and certified experimentally for unequal spikes.The unequal-spike certificate establishes three positive adjacent expected gaps at σ = ∥H∥2 with simultaneous 99% Monte Carlo confidence.
  • Conclusion: The sixth-order low-noise correction improves the lower-crossover estimate, while the high-noise spread is O((h⋆/σ)^4).These corrections refine the observed intermediate-noise visibility regime.
  • Conclusion: The convergence radius is the nearest exceptional point crossing the retained-complement boundary, and reducing nullity changes it nonmonotonically.In the reproducible experiment, q = 2 is optimal among five tested nullities with ρ2 = 1.40126671.
  • Conclusion: Visibility endpoints, exceptional-point radii, and the asymptotic BBP threshold are related spectrally but remain distinct quantities.The divergence of the plotted series reflects leaving its Taylor disk, not failure of the exact SVD.
Loading 2608.30374v1…