Source-linked AI summary

When Is the Sharp Covariance Envelope Tight? Feature-Only Geometry for Volume-Sampled Least Squares

Kihun Rhee

arXiv:2608.26877v1cs.LGstat.ML

TL;DR

The paper addresses whether a globally sharp covariance envelope is tight on a particular fixed feature pool, beyond prior arbitrary-response results at s=d. It combines a universal Loewner envelope with residual augmentation and feature-only geometry, showing an exact strictness-versus-attainment phase under stated conditions. The scope is conditional coefficient covariance for ordinary volume-sampled selected OLS, not population generalization.

  • Problem

    Prior exact arbitrary-fixed-response loss and prediction-covariance formulas for ordinary volume sampling were established at the rank-size endpoint s=d, leaving the fixed-design attainability question open.

  • Method

    The paper proves a globally sharp Loewner envelope, uses residual augmentation for a response-aware slack mechanism, and defines a feature-only margin to classify fixed-pool attainability.

  • Results

    νA>0 iff every compatible residual is strictly below the normalized spectral envelope, while νA=0 iff some compatible residual attains it at every strict-interior budget.

  • Takeaways & Limitations

    Sound lower certificates convert verified positive margins into conservative cardinality decisions without changing the sampler or least-squares fit.

  • Takeaways & Limitations

    The claims concern conditional centered, full-Gram-whitened coefficient covariance for one indexed fixed pool and do not establish population generalization or exact general-design computation of νA.

Abstract

from arXiv · show

Prior analyses by Derezinski and Warmuth established all-size sampling identities, selected-OLS unbiasedness, and inverse moments for ordinary volume sampling, while their exact arbitrary-fixed-response loss and prediction-covariance formulas are at the rank-size endpoint s=d. We establish a Loewner envelope for centered coefficient covariance for every full-rank fixed pool, response, and legal budget d <= s <= m under ordinary indexed fixed-size volume sampling followed by selected unweighted least squares; its coefficient is globally sharp over the full-rank class. Global sharpness does not determine attainability on the pool in hand. Under positive loss, strict-interior budgets, and no coloops, a feature-only margin nu_A gives the exact fixed-design spectral phase: nu_A > 0 if and only if the normalized spectral envelope is strict for every compatible residual, whereas nu_A = 0 if and only if some compatible residual is spectrally tight; the same zero-margin residual is tight at every strict-interior budget. A residual-augmented change of measure supplies the response-aware mechanism and a one-sided quantitative slack bound, while support saturation proves the attainment direction. Critical equal-leverage geometry interprets the boundary, and sound lower certificates yield conservative same-primitive cardinality decisions. Frozen-feature examples show that the certificate is nonvacuous and measure the fixed-pool cost of its authorized reduction. The claims concern conditional centered, full-Gram-whitened coefficient covariance, not population generalization.

1 Introduction

The paper asks when a globally sharp covariance envelope is actually tight on a particular feature pool. It answers with a feature-only margin that separates uniform strictness from existential attainment under stated interior conditions.

  • Motivation: Frozen-feature least-squares refitting can vary along concentrated coefficient directions that scalar same-pool loss does not reveal.The analysis isolates the linear readout and studies its most variable refitted-coefficient direction, not classification loss or representation quality.
  • Problem: The fixed-pool problem maximizes conditional centered, full-Gram-whitened coefficient covariance over compatible residuals, with subset sampling as the only randomness.The resulting worst-residual margin depends on features rather than the observed response.
  • Gap: Prior exact arbitrary-response loss and prediction-covariance formulas reached the rank-size endpoint s=d, leaving expected-loss bounds and covariance extensions for s>d open.The paper extends the analysis to legal budgets beyond the rank-size boundary.
  • Main result: The feature-only margin νA exactly separates uniform strictness for every compatible residual from attainment by some compatible residual.When νA=0, one compatible positive-loss residual attains the normalized spectral envelope at every strict-interior subset size; when νA>0, all compatible residuals remain below it.
  • Supporting results: Theorems 1–2 provide a globally sharp all-budget benchmark and a response-aware residual-augmentation mechanism behind the fixed-design phase.Residual augmentation yields an exact interior representation and a boundary-valid one-sided resolvent.
  • Implications: Critical equal-leverage geometry identifies the boundary, while sound lower certificates convert only verified positive margins into conservative cardinality actions.Frozen-feature examples demonstrate nonvacuous certificate decisions and measure their fixed-pool cost.

2 Setting, metric, and scope

The setting is a deterministic full-rank regression pool with ordinary indexed fixed-size volume sampling and selected unweighted least squares. The metric is conditional centered coefficient covariance in full-Gram-whitened coordinates, with explicit endpoint and zero-support conventions.

  • Setting: The paper fixes a full-column-rank design X and deterministic response y, so randomness comes only from the indexed subset draw.All expectations are conditional on the fixed pair (X,y).
  • Sampling: Ordinary unrescaled fixed-size volume sampling selects indexed subsets S of size d≤s≤m using determinant-weighted support.Repeated rows at different indices remain distinct observations.
  • Support conventions: Zero-volume subsets have zero probability and receive no pseudoinverse, ridge term, rescaling, weighting, replacement rule, or selected fit.Only positive-volume subsets support the estimator.
  • Metric: The centered second moment is expressed in an invariant form for the selected coefficient estimator.The analysis later normalizes its largest eigenvalue by αL* only when αL*>0.
  • Metric: The normalized directional quantity qs is λmax(Ms)/(αL*) rather than a raw Euclidean eigenvalue.No directional quotient is formed at a zero denominator.
  • Endpoints: At m=d, s=m, or L*=0, the full fit or supported selected fits have zero loss and covariance, so residual directions and singular-denominator expressions are excluded.At s=d<m, supported square systems interpolate and residual augmentation is not invoked.

3 Universal budget envelope

Theorem 1 establishes a design-independent Loewner covariance envelope for every full-rank pool, response, and legal budget. Its coefficient is globally sharp, while regular positive-loss interiors are strictly below the envelope.

  • Universal budget envelope: The envelope coefficient is α=(m−s)/(m−d), obtained using indexed Cauchy–Binet normalization, rank-size moments, padding, and selected-residual moments.The construction is design-independent.
  • Universal budget envelope: Theorem 1 bounds centered whitened coefficient covariance in Loewner order for every full-column-rank X, fixed y, and d≤s≤m.The bound controls all coefficient contrasts simultaneously before taking a trace.
  • Sharpness: The coefficient is attained by a full-column-rank core-plus-zero family, so global sharpness is a coefficient-attainment statement rather than an equality classification.The same normalized coefficient is a perturbative supremum for regular interior designs.
  • Interior behavior: For positive-loss row-general-position designs with d≥2 and d<s<m, the matrix inequality is strict.The interior proof couples a size-s sample to a rank-size basis with uniform padding and applies total covariance.
  • Boundary: At arbitrary full-column-rank boundary designs, a determinant-weighted directional limit preserves the one-sided envelope without assigning inverses to singular subsets.The boundary argument does not continue the interior decomposition.

4 Residual-augmented mechanism

Theorem 2 augments the whitened design with the normalized residual to expose the response-dependent mechanism behind the universal covariance ceiling. It yields an exact interior change of measure and a one-sided boundary resolvent, plus a second-moment tightening.

  • Augmentation: Residual augmentation appends the normalized residual to the whitened design only for analysis; sampling and fitting remain ordinary volume sampling and selected OLS.The construction applies when L*>0 and d<s<m.
  • Change of measure: The response-dependent auxiliary law PB is related to the ordinary law PX by an exact loss-weighted change of measure on row-general-position interiors.The auxiliary support consists of distinct rank-(d+1) subsets and can be smaller than the ordinary support.
  • Boundary limitation: The exact inverse-moment covariance identity does not extend through a rank-changing boundary; only a congruent one-sided inequality remains there.This boundary limitation preserves an envelope but not the interior identity.
  • Resolvent: The augmented first moment and operator Jensen produce an interior resolvent bound and its linear consequence for the covariance ceiling.The minus sign in the middle identity reverses the matrix order.
  • Refinement: Every supported augmented sub-Gram matrix is a positive contraction, enabling an exact resolvent identity and a computable second-centered-Gram-moment tightening.Proposition 16 extends the hierarchy one-sidedly to arbitrary full-column-rank boundaries.

5 Robust pre-response geometry

Theorem 3 uses a feature-only margin to determine whether the universal spectral covariance envelope is uniformly strict or attainable by a compatible residual. A residual-augmented contraction supplies the positive branch, while support saturation proves tightness in the zero branch.

  • Robust pre-response geometry: The feature-only margin νA answers whether the globally sharp envelope is strict for every compatible residual or attained by some residual on the fixed pool.This resolves the fixed-design attainability question left open by class-level sharpness.
  • Robust pre-response geometry: Under d < s < m, m ≥ d + 2, positive loss, and no coloops, νA is defined over compatible normalized residual directions after whitening A⊤A = Id.The theorem uses ordinary unrescaled indexed size-s volume sampling and selected unweighted least squares.
  • Robust pre-response geometry: νA > 0 certifies uniform strictness, whereas νA = 0 yields a compatible residual that attains the envelope at every strict-interior subset size.The same zero-margin witness works simultaneously for s′ ∈ {d + 1, …, m − 1}.
  • Robust pre-response geometry: The residual-augmented matrix B = [A z] has orthonormal columns, enabling a response-aware contraction and a lower bound on spectral slack.The positive slack bound is one-sided and does not identify the exact slack magnitude.
  • Robust pre-response geometry: Zero margin forces augmented-leverage support saturation, whose mutually exclusive omission events under volume sampling reproduce the universal envelope.Conversely, equality forces a zero Rayleigh quotient and therefore zero margin.

6 Critical equal-leverage geometry

At critical redundancy with equal leverage, the zero-margin boundary has repeated-pair geometry inherited from binary erasure theory, while residual-coupled weights provide a continuous formulation. The resulting witness is budget-independent across strict-interior sizes.

  • Critical frame and erasure geometry: The zero boundary is the projectively repeated-pair boundary with an orthogonal remainder.This gives the critical equal-leverage interpretation of νA = 0.
  • Critical frame and erasure geometry: The continuous margin uses one fixed Naimark complement, one removed residual direction, and the induced lower frame bound rather than a binary erasure mask.Coherence bounds νA from two sides but generally does not compute it exactly.
  • Critical residual witness: Saturated residual coordinates force a repeated pair through unit-norm and Parseval constraints, linking the zero Rayleigh quotient to critical frame structure.The signed difference of a coherence-maximizing pair supplies the complementary quantitative bound.
  • Critical residual witness: At νA = 0, one compatible residual fixed before sampling attains the universal spectral envelope for every s ∈ {d + 1, …, 2d − 1}.The witness is independent of the budget throughout the strict-interior range.
  • Critical frame and erasure geometry: At m = 2d with equal leverage, Figure 2 contrasts classical binary two-erasure and Naimark-complement geometry with residual-coupled continuous weights.Its second panel separates the existential zero branch from the uniform positive branch of the νA phase.

7 Sound certificates and same-primitive action

Sound feature-only certificates can certify the positive-margin branch without computing νA exactly. Their positive passes authorize conservative cardinality reductions while failures and abstentions remain inconclusive.

  • Sound pre-response certificates: A verified feature-only scalar t(A) with 0 ≤ t(A) ≤ νA certifies positive margin whenever t(A) > 0.This avoids computing the exact minimum over the residual sphere.
  • Sound pre-response certificates: Failure, zero output, or abstention does not imply νA = 0, envelope attainment, or the absence of another sound certificate.The rule is deliberately one-sided and conservative.
  • Same-primitive action: For the predeclared (m, d, ξ) = (512, 64, 1/3) profile, a verified pass authorizes s = 338; otherwise the rule falls back to s = 363.Both branches retain the same primitive and the same one-sided 1/3L∗Id covariance ceiling.

8 Pre-response decisions and their measured cost

Frozen feature pools demonstrate that the sound certificate can make nonvacuous pre-response cardinality decisions while sometimes abstaining. The measured fixed-pool cost is descriptive and does not establish population-level performance or envelope-phase membership.

  • Measured fixed-pool decisions: Across 72 frozen-encoding cells, the sufficient rule produced action/fallback counts of 20/4, 4/20, and 5/19 across three released blocks.The rule both authorized the smaller cardinality and abstained on other cells.
  • Measured fixed-pool decisions: The frozen-feature action layer determines whether to authorize the smaller subset or invoke the universal-envelope fallback before observing a response.This preserves the same sampling and estimation primitive.
  • Scope of the measurements: These descriptive cells do not estimate TRU_spec, test Theorem 3, or establish population performance or selector dominance.A fallback is inconclusive rather than evidence of the zero phase.

9 Related work

The related-work comparison distinguishes prior results by sampling law, estimator, budget-specific object, response model, covariance target, and equality quantifier. The cited neighboring works overlap only partially because they study different primitives or targets.

  • Ordinary volume sampling: Prior ordinary-volume results include all-size determinant laws, selected-OLS unbiasedness, and inverse-Gram moments.The arbitrary-fixed-response loss and prediction-covariance formulas remain identified at the rank-size endpoint s = d.
  • Different primitives: Other audited neighbors change the sampling law, estimator, response model, or covariance target relative to this work.Examples include volume-rescaled and random-design regression, generic-sketch covariance, with-replacement debiased fits, and active or dependent-leverage designs.
  • Pre-response selection: Pre-response selection studies include weighted sparsification, weighted empirical risk minimization, residual/leverage deletion, scalar full-pool loss, and dependent-leverage sampling.These results address different sampling primitives or target statistics.
  • Inverse-Gram comparisons: Prior inverse or pseudoinverse results average matched selected-Gram objects under the ordinary feature law, unlike the residual-augmented correction here.The comparison is therefore to a different response-aware calculation.
  • Comparison framework: Table 1 compares the present fixed-pool coefficient-covariance envelope and strict/tight phase against audited theorem-level results.Its relations are not claims about volume sampling as a whole.
  • Residual augmentation: The residual-augmented calculation uses a response-induced law–target mismatch only in the supplementary contraction correction, while operational sampling remains under the ordinary feature law.Its first-order step is classical operator Jensen; the ordered remainder is a specialized local calculation.

10 Limitations and conclusion

The paper’s scope is conditional coefficient covariance for one fixed feature pool under ordinary volume sampling and selected unweighted OLS. Under stated conditions, feature geometry decides strictness versus tightness, while certificates support conservative cardinality decisions without claiming population generalization.

  • Limitations: The analysis concerns conditional centered, full-Gram-whitened coefficient covariance for one indexed fixed feature pool, with subset selection as the only randomness.Sampling is ordinary unrescaled fixed-size volume sampling followed by selected unweighted OLS.
  • Limitations: The exact phase theorem requires positive full-fit loss, a strict-interior budget, and no coloops; the residual-augmented identity additionally requires row general position.Its boundary form is one-sided.
  • Limitations: Feature-only screens are sufficient but not complete: passing certifies positive margin, whereas failure or abstention is inconclusive.The paper does not establish the computational complexity of evaluating ν_A exactly for a general design.
  • Conclusion: The cardinality rule and finite-pool evidence do not claim population generalization, selector dominance, a new sampler, or predictive utility.These are explicit scope boundaries rather than conclusions about downstream performance.
  • Conclusion: Under the stated conditions, ν_A > 0 means uniform strictness, while ν_A = 0 means one compatible residual is tight at every strict-interior budget.Critical geometry interprets the boundary, and sound lower certificates yield conservative cardinality decisions without changing the sampler or fit.
  • Disclosure: AI-assisted drafting and verification workflows are not evidence for mathematical correctness, novelty, empirical claims, or citation validity.The author retains responsibility for the manuscript’s mathematical and empirical content.

Proof support and supplementary material

The supplementary material follows the dependency order from ordinary volume sampling and covariance envelopes through residual augmentation, robust geometry, critical geometry, and feature certificates. It also provides a fixed-threshold instantiation with certified cardinality branches.

  • Appendix organization: The appendix organizes derivations by result dependencies: ordinary law and covariance envelope, residual augmentation, robust pre-response geometry, critical geometry, and the response-uniform feature certificate.It also retains balanced and finite diagnostic material as supplementary content.
  • Certificate construction: The fixed-threshold route defines its predicate in the original feature coordinates.The supplied passage identifies the route but does not state the full predicate.
  • Certificate construction: The strict guards 0 < ℓ_i < 1 for every i, together with Φ_13(X) ⪰ 0, certify 13/60 ≤ ν_A.This certificate follows from the derivation in Appendix G.
  • Certified instantiation: For the fixed profile (m, d, ξ) = (512, 64, 1/3), the residual dimension is n = 448.This profile is used for the certified instantiation.
  • Certified instantiation: When the strict guards and fixed-threshold predicate pass, the certified count is s_13 = 338 for every fixed positive-loss response.The passage states the count and response condition but does not reproduce the associated inequality.
  • Fallback route: If the sufficient route abstains, the universal-envelope fallback gives q_H = 149 and s_H = 363 with the same one-sided tolerance certificate.Both branches retain ordinary indexed fixed-size volume sampling and selected unweighted OLS; only cardinality differs.

Theorem-to-proof map and scope

The paper maps its theorem chain from universal covariance envelopes and residual-augmented transforms to fixed-design phase results, critical geometry, certificates, and supplementary fixed-pool diagnostics, while recording their domains and quantifiers.

  • Universal envelope: Theorem 8, Proposition 10, Lemma 11, and Proposition 13 establish universal unbiasedness and the covariance envelope.
  • Universal envelope: Equation (67) derives the full-pool loss consequence using the full-pool Pythagoras identity.
  • Attainment and slack: Propositions 14 and 15 address attainment, strict-interior slack, and the supremum, while Proposition 17 distinguishes an exact interior identity from a positive-slack boundary inequality.
  • Response-aware mechanism: Equations (94), (98), and (103)–(121) provide the residual-augmented transform, resolvent, and boundary inequality within their stated domains.
  • Fixed-design phase: Theorem 3 gives the robust zero/positive phase classification for no-coloop whitened designs, while Theorem 4 gives critical equal-leverage geometry and a sharp witness.
  • Certificate and scope: Corollary 5 supplies a sufficient feature-only lower certificate, but it is silent for d < m ≤2d and positivity is not guaranteed when m > 2d.

C.1 Fixed-cardinality normalizer and theorem statement

The section derives a fixed-cardinality normalizer and establishes a universal Loewner covariance envelope for volume-sampled least squares. It then separates global sharpness from fixed-design attainment, identifying strict interior slack and boundary attaining families.

  • Theorem statement: Theorem 8 bounds centered, full-Gram-whitened coefficient covariance for every full-column-rank X, fixed response y, and legal budget d ≤ s ≤ m.The bound is matrix-valued, preserving directional information rather than only controlling the trace.
  • Mean and covariance: The selected coefficient is unbiased for the full-pool least-squares coefficient, with the identity applying without requiring L∗ > 0.The selected residual mean identity supports the covariance and loss analysis, while singular selected Gram matrices require separate boundary arguments.
  • Boundary extension: The covariance inequalities and corresponding loss bound extend from row-general-position designs to every full-column-rank X by one-sided boundary passage.The continuation retains limiting positive-volume terms and drops only newly supported nonnegative terms; it does not continue the exact covariance decomposition to singular selected-Gram boundaries.
  • Attainment: The coefficient is attained by a core-plus-zero full-column-rank family, establishing global sharpness over the full-rank class.This is a coefficient-attainment statement, not an equality classification for all designs.
  • Interior phase: For d ≥ 2 and d < s < m, positive-loss row-general-position designs are strictly below the envelope, although the same coefficient remains their operator-norm supremum.At s = d, equality holds for every positive-loss row-general-position design; at s = m, no normalized ratio is formed.

D.6 One-sided extension to arbitrary full-column-rank boundary designs

The boundary extension preserves one-sided covariance-envelope inequalities for arbitrary full-column-rank designs, while exact covariance identities remain restricted. Balanced geometry and finite diagnostics clarify the resulting gaps and their limits.

  • One-sided boundary extension: The augmented exclusion calculation yields an exact first moment without requiring row general position.This supports the response-aware upper inequality on boundary designs.
  • One-sided boundary extension: Boundary passage from perturbed designs preserves one-sided comparisons after singular-support terms are dropped.Zero-volume sets contribute no boundary estimator, while newly supported perturbed terms are nonnegative.
  • One-sided boundary extension: The boundary resolvent inequality can remain strict, with slack (209/126)I2 ≻0.This demonstrates strictness for a concrete boundary calculation.
  • Phase and attainment: Theorem 2 gives νA > 0 ⇒ spec(A, s) < αs, while zero-margin support saturation supplies a witness attaining the envelope across strict-interior budgets.The same feature-selected witness and direction work for every s′ ∈ {d + 1, ..., m − 1}.
  • Balanced geometry: At m = 2d, the scalar certificate cX may vanish even though response-aware analysis still certifies a covariance gap.The Parseval core-plus-zero design shows that redundancy alone is insufficient for cX positivity.
  • Balanced geometry: For d = 4n and 1 ≤ r < d/2, one fixed design-response pair supplies a strict upper witness while uniform lower bounds keep the gap of order r/d.The matching-order squeeze applies only on the stated subsequence and budget range.
  • Scope limits: The balanced near-attainment statement is not asserted for positive limiting budget fractions, all budgets, trace covariance, or augmented-leverage margins.The explicit witness has an augmented coloop, so no augmented-margin claim follows.
  • Finite diagnostics: Finite residual-localization diagnostics show monotone ceiling contraction, but their exact and Monte-Carlo outputs are mechanism illustrations rather than theorem evidence.Monte-Carlo diagnostics used seeded draws and jackknife error bars, with sampled endpoints below the deterministic ceiling at reported resolution.
Loading 2608.26877v1…