Source-linked AI summary

Spectral Rank Certification for Foundation Model Adapters

Mohammed Ahnouch, Lotfi Elaachak

arXiv:2608.15351v1cs.LGstat.ME

TL;DR

LoRA’s nominal rank does not reveal how many components are statistically distinguishable from null structure. This paper develops finite-sample spectral calibration and an empirical-null workflow, finding that calibrated effective rank is usually far below nominal rank and can disagree with energy retention.

  • Problem

    LoRA rank is typically an engineering hyperparameter, while the number of components distinguishable from null structure remains an inferential question.

  • Method

    The paper combines exact Gaussian/Haar finite-sample theory with empirical-null testing, Monte Carlo p-values, deflation, block tests, and BH reporting for adapter spectra.

  • Results

    Calibrated effective rank is usually much smaller than nominal rank, with module-wise BH retaining a median of one component versus three under 95% energy retention.

  • Takeaways & Limitations

    Energy retention and statistical surprise answer different questions, so nominal rank should not be treated as calibrated effective rank.

  • Takeaways & Limitations

    The Gaussian/Haar reference model does not claim trained adapter residuals are isotropic Gaussians, so deployment relies on empirical nulls for the full testing pipeline.

Abstract

from arXiv · show

Nominal LoRA rank is a design parameter; calibrated spectral evidence is a separate inferential quantity. This article develops a finite-sample framework for inferring effective rank structure in public foundation-model adapters. The theoretical core is an exact chi-square divergence for the fixed-dimensional Gaussian rank-one reference experiment, with an unknown signal direction integrated under a rotation-invariant reference prior. The resulting series yields a computable finite-sample Le Cam bound at concrete layer sizes, an explicit remainder bound for numerical truncation, and the rectangular Baik-Ben Arous-Peche (BBP) limit. A compact-manifold Laplace expansion shows that finite-sample likelihood evidence also depends on leading spectral gaps through the factor $s_1^{|m-n|}\prod_{i\ge2}(s_1^2-s_i^2)$, motivating joint calibration of clustered singular values. Building on these results, we introduce an empirical-null workflow for PEFT LoRA adapters: factor reconstruction, Monte Carlo $p$-values, stagewise and block testing, and module-wise and corpus-level BH reporting. In an audit of 26 public adapters, 684 modules, six architecture families, and 31,770 public-checkpoint spectra rows, calibrated effective rank is typically much smaller than nominal rank and differs systematically from 95\% energy retention. A measured RoBERTa-RTE slice on $n=24$ examples illustrates the measurement path from calibrated ranks to task evaluation, without treating the slice as a utility study. The main empirical finding is that calibrated effective rank is usually far below nominal rank, and that energy retention and statistical surprise answer different questions.

1 Introduction · 2 Related Work

The paper separates nominal LoRA rank and energy-based truncation from inferential evidence about components distinguishable from null structure. It develops finite-sample spectral calibration and an empirical-null workflow, positioning these contributions relative to rank-allocation methods and asymptotic random-matrix theory.

  • 1 Introduction: Nominal LoRA rank controls trainable parameters and memory, whereas calibrated evidence asks which adapter components are distinguishable from null structure.Adapters with the same nominal rank can have different spectra, and similar reconstruction energy can yield different false-positive interpretations.
  • 1 Introduction: The target output is a reproducible label identifying whether component j is surprising under a specified null construction N_ℓ at level q.This calibrated decision complements pruning and rank-allocation methods by separating statistical evidence from downstream engineering tradeoffs.
  • 1 Introduction: The reference experiment tests H0: Y = σG against H1: Y = θuv^⊤ + σG, with independent Gaussian noise and Haar-uniform signal directions.The Haar law provides a rotation-invariant reference prior for unknown signal directions; singular-value thresholds require calibration when average-direction detection has limited power.
  • 1 Introduction: The theoretical contributions include an exact fixed-(m, n) second moment, finite-sample Le Cam and series-remainder bounds, the rectangular BBP limit, and a Laplace expansion involving leading spectral gaps.The integrated likelihood depends on leading gaps through the factor s_1^{|m-n|}∏_{i≥2}(s_1^2-s_i^2), motivating joint calibration of clustered singular values.
  • 1 Introduction: The deployable workflow combines PEFT factor reconstruction, Monte Carlo p-values, stagewise calibration, block tests, and module-wise and corpus-level BH reporting.These procedures are designed for empirical-null analysis of public adapter factors.
  • 1 Introduction: The public-checkpoint audit finds calibrated effective rank typically far below nominal rank, while energy retention and statistical surprise answer different questions.A small RTE illustration records retained-rank tradeoffs under the same truncation rules without turning the slice into a utility study.
  • 2 Related Work: LoRA and related PEFT methods constrain updates to low-rank subspaces, while AdaLoRA reallocates rank budgets and DoRA, VB-LoRA, and LoRA-Mini alter update parameterization or decomposition.These methods address how to allocate or reparameterize rank during learning, rather than how to infer calibrated evidence after training.
  • 2 Related Work: Spiked random-matrix models and singular-value shrinkage provide asymptotic detection, outlier, and MSE-oriented spectral rules but do not by themselves provide finite-sample false-positive bounds at concrete layer dimensions.This gap motivates the paper’s fixed-dimensional calibration framework.

3 Exact Rank-One Bound

Section 3 derives an exact chi-square divergence for the Haar-mixture rank-one Gaussian reference experiment and establishes explicit finite-sample control. It also gives the BBP-scale limit and converts divergence into a Le Cam power bound at a chosen false-positive rate.

  • Exact chi-square divergence: Theorem 1 gives the exact chi-square divergence for independent Haar directions (u,v) and (u′,v′), with signal-to-noise parameter λ = θ^2/σ^2.The derivation uses Gaussian moment-generating functions, symmetry, independence, and rotational invariance to evaluate the required even moments.
  • Finite-sample control: At fixed m and n, the divergence series is entire and has an explicit truncation remainder whenever q_K+1 < 1.The bound uses q_j = λ^2/[(m + 2j)(n + 2j)]; a crude global alternative is R_K ≤ e^|λ| P{Poisson(|λ|) ≥ 2K + 2}.
  • BBP-scale limit: When m,n →∞ and λ^2/(mn) → β ∈ [0, 1), the exact divergence converges to the stated BBP-scale limit.The proof applies fixed-j moment asymptotics, geometric domination, and the central-binomial generating function.
  • Testing consequence: The finite-sample testing consequence translates the divergence calculation into a bound on attainable power at a chosen false-positive rate for the reference experiment.The final inequality follows from Cauchy-Schwarz on L−1.

4 Full-Spectrum Evidence and Higher-Rank Signals

The section contrasts asymptotic spectral scales with finite-sample calibration via null quantiles. It also shows that integrated likelihood depends on leading spectral gaps, while clustered components and higher-rank signals require joint rather than purely marginal evidence.

  • Finite-sample calibration: The unnormalized null edge is approximately σ(√m+√n), while the additive rectangular outlier scale is σ(mn)1/4.These asymptotic quantities provide scales, not finite-sample false-positive rates.
  • Finite-sample calibration: A calibrated spectral rule should use the null quantile rather than treating a fixed asymptotic edge as a finite-sample error threshold.The null threshold is defined through the distribution of the top singular value under P0.
  • Full-spectrum likelihood evidence: Near s1 = s2, a block statistic is more appropriate than the nondegenerate single-component expansion.This follows because the expansion requires s1 > s2 and depends on the leading spectral gaps.
  • Higher-rank signals: Componentwise BH tests each singular value marginally and therefore does not test the coupling induced by higher-rank signals.Clustered near-critical components should be tested jointly or handled by a global rank budget.

5 Empirical-Null Algorithms for Adapters

This section defines an empirical-null calibration workflow for adapter spectra, emphasizing that null construction is part of the scientific claim. The deployable procedure reconstructs module updates, performs stagewise Monte Carlo testing, and reports retained ranks, diagnostics, and abstentions.

  • Empirical-null construction: Null construction must be reported because it constitutes part of the scientific claim, with options including seeds, residual-tail fitting, sign flips, shuffles, permutations, and task-preserving bootstraps.
  • Reporting outputs: The reporting algorithm outputs retained ranks, diagnostics, and abstentions when the selected null construction falls outside calibration criteria.
  • Algorithm inputs: The procedure accepts an adapter checkpoint or factors, null constructor N, target level q, Monte Carlo count B, random seed, and module list.
  • Module processing: For each module ℓ, it reconstructs the dense update and computes singular values s_ℓ,1 ≥ · · · ≥ s_ℓ,rℓ.Public PEFT-style LoRA checkpoints generally store factors, while the reconstructed update has rank at most nominal rank r up to numerical tolerance.
  • Stagewise testing: Calibration null samples are generated independently of evaluation samples, and each component uses the same deflation or block operation applied to the observed adapter.

4. Use the add-one Monte Carlo value

The workflow calibrates the full data-dependent testing pipeline by rerunning deflation, p-value computation, and selection on null adapters. When singular-value gaps are small, it switches from componentwise inference to conservative block-level testing.

  • BH outputs should be described as empirically calibrated unless dependence conditions are proved for the testing pipeline.
  • Null calibration reruns deflation, computes all p-values, applies the same selection rule, and reports realized false-positive and discovery summaries with Wilson or bootstrap intervals.
  • When relative gaps fall below a threshold, componentwise tests can overstate distinctions between adjacent singular directions, so block inference becomes the fallback.
  • The default block-reporting threshold is δ = 0.10, with adjacent components grouped when gj < δ and joint block statistics calibrated using the same null constructor and deflation history.
  • If block and componentwise decisions disagree, report both and treat the block result as the conservative block-level decision.

6 Synthetic Validation

Synthetic validation shows that finite-sample calibration materially changes operating characteristics: the asymptotic edge has false-positive rate 0.129 versus nominal 0.05, near-critical spikes have subunit power, and strong spikes are detected almost surely. Sequential rank testing uses stage-specific residual dimensions and BH retention at q = 0.10, while trained-adapter validation requires the empirical-null program in Section 7.

  • Sequential-null diagnostic: Sequential testing uses nominal residual dimensions (m − j + 1, n − j + 1) at each stage and reports BH retention at q = 0.10.These retention frequencies characterize the stated sequential-null simulation.
  • Scope and deployment analogue: Validation on trained adapters is deferred to an empirical-null program using held-out seeds, task-preserving randomizations, or another declared construction.The synthetic experiment alone does not establish analogous effects on trained adapters.
  • Synthetic operating characteristics: 0.129 versus nominal 0.05: the finite-sample asymptotic edge is anti-conservative, while strong spikes remain detectable almost surely.Figure 3 likewise emphasizes that calibration changes the reported false-positive rate while strong spikes remain detectable.
  • Synthetic operating characteristics: Near-critical spikes have calibrated power well below one, demonstrating that detectability is not uniform across spike strengths.The synthetic results isolate this behavior as an effect exposed by the reference model.

7 Public Adapter Checkpoints

The study audits public LoRA adapter checkpoints across six architecture families using reconstructed spectra and empirical-null calibration. Calibrated effective rank is generally much lower than nominal rank and 95% energy retention, while the RTE slice serves only as a reproducible measurement path.

  • Corpus and implementation: The weights-only corpus covers RoBERTa, RoBERTa-large, XLM-RoBERTa, FLAN-T5, Whisper, and LLaMA adapters, with spectra reconstructed from pinned repositories and revisions.Derived spectra, metadata, and checksums are archived while third-party weights remain with their original sources.
  • Null calibration: Each module compares a primary factor-channel shuffle null with factor-sign and matched-Gaussian sensitivity nulls.The primary null uses 999 samples per module, giving a smallest Monte Carlo p-value of 0.001 and supporting module-wise BH at q = 0.10 for ranks up to 64.
  • Calibrated effective rank: Most trained modules contain at least one empirically surprising component, yet module-wise BH retains a median of one component versus three under the 95% energy rule.The MRPC rank sweep reports median module-wise BH ranks of 0 for r = 1 and 1 for r = 8 and r = 16.
  • Measurement-path illustration: The RoBERTa RTE illustration evaluates full, energy-95, corpus-level calibrated BH, and module-wise calibrated BH adapters on the first n = 24 validation examples.It uses CPU inference, verifies the saved classifier head, and applies best rank-k approximations to dense adapter updates.
  • Measurement-path illustration: A two-example gap on the RTE slice is not a significance claim; the experiment is a reproducible measurement path rather than a utility study.The rank-budget result matches the full-adapter count, while energy-95 retains about 26%.

8 Limitations

The Gaussian/Haar model provides finite-sample calibration for a reference rank-one experiment, not a claim about the distribution of trained adapter residuals. Deployment therefore relies on empirical nulls that reproduce the full testing pipeline.

  • Model scope: The Gaussian/Haar experiment is a calibration reference, not an isotropic-Gaussian model of trained adapter residuals.Real residuals may exhibit anisotropy, heavy tails, row/column variance heterogeneity, tensor structure, or optimizer-induced correlations.
  • Deployment calibration: Equation (5) bounds attainable power only for the stated rank-one reference experiment.Deployment extends the bound through empirical nulls reproducing the full deflation and block-testing pipeline.

9 Reproducibility Statement

The reproducibility artifact regenerates synthetic results, extracts spectra from pinned public adapter revisions, and validates numerical-rank, p-value, retained-rank, and null-calibration checks. It archives derived spectra, metadata, checksums, and measured RTE evaluation rows while leaving third-party weights at their original sources.

  • Artifact and data provenance: The artifact regenerates synthetic tables and writes public-checkpoint adapter spectra from pinned repositories and revisions.It also checks numerical rank rank(∆W) ≤r.
  • Validation checks: The pipeline validates p-values in [0, 1], confirms retained counts do not exceed nominal rank, and runs null-only calibration.These checks provide explicit reproducibility safeguards for the empirical-null workflow.
  • Archived outputs: Derived spectra, metadata, checksums, and measured RTE evaluation rows are archived, while third-party weights remain with their original sources.Public-adapter summary rows use real weights-only spectra from pinned PEFT adapters.

10 Conclusion

The conclusion separates nominal adapter rank from calibrated effective rank, combining finite-sample spectral theory with empirical-null testing for real adapters. It emphasizes full-spectrum and spectral-gap effects while treating training dynamics and task-specific sensitivity as possible explanations rather than established mechanisms.

  • Theoretical and empirical contributions: Calibrated effective rank is distinct from nominal rank, and the exact chi-square series supplies a finite-sample Le Cam bound with an explicit numerical-truncation remainder.The result concerns the rank-one reference experiment.
  • Theoretical and empirical contributions: The Laplace expansion shows that likelihood evidence can depend on the full spectrum, not only the top singular value, including leading spectral-gap effects.This motivates considering full-spectrum information in calibration.
  • Theoretical and empirical contributions: For real adapters, empirical nulls extend the reference calculation to full deflation and block-testing pipelines.The conclusion links misspecification and adaptive-deflation results to structured empirical nulls and full-pipeline resampling.
  • Theoretical and empirical contributions: Multiplicity control is part of the proposed empirical workflow, alongside structured empirical nulls and full-pipeline resampling.The conclusion cites Benjamini and Hochberg (1995), Efron (2010), and Westfall and Young (1993) in this context.
  • Interpretation and limitations: Training dynamics and task-specific sensitivity remain possible explanations for rank differences, not established mechanisms.The conclusion explicitly limits the interpretation of these explanations.
Loading 2608.15351v1…