Source-linked AI summary

Representation Measurements Under Function-Preserving Reparameterizations

Abdullah Karasan

arXiv:2608.27020v1stat.MLcs.LG

TL;DR

Representation-derived measurements should not change when hidden coordinates are reparameterized without changing a model’s function, but the paper finds that column-permutation parallel analysis violates this requirement. Using theoretical analysis, controlled experiments, and invariant comparators, it shows that PA-derived counts and decisions can vary with coordinate choice despite an unchanged observed covariance spectrum.

  • Problem

    Hidden coordinates are not uniquely determined by a language model’s input–output function, raising whether representation-derived measurements are invariant under function-preserving changes of basis.

  • Method

    The paper analyzes column-permutation parallel analysis theoretically and empirically, including structural reference-procedure results, controlled reparameterizations, seed and centering controls, and orthogonally invariant comparators.

  • Results

    Across the tested settings, median fixed-threshold decision disagreement is 0.26 under reparameterization versus 0.00 under the independent-seed control, while invariant comparators retain similar observed discrimination.

  • Takeaways & Limitations

    PA-derived component counts and downstream decisions can reflect hidden-coordinate choice rather than a well-defined property of the model function.

  • Takeaways & Limitations

    The empirical scope covers five small open-weight models, three question-answering domains, and tested architecture-compatible transformations; the theoretical marginal-preservation limitation requires the full orthogonal group O(d).

Abstract

from arXiv · show

Hidden coordinates are not uniquely determined by a language model's input--output function, so representation-derived measurements should be invariant to function-preserving changes of basis. This study shows that column-permutation parallel analysis violates function-preserving reparameterization invariance because its reference distribution and selected component count can change while the model function and observed covariance spectrum remain fixed. More generally, a data-internal reference procedure cannot simultaneously preserve every coordinate marginal, remain orthogonally equivariant, and remove cross-coordinate covariance. Empirically, across five models, three retrieval domains, and 75 transformations, median component-count disagreement is 0.79 and median fixed-threshold decision disagreement is 0.26. A centering-only control isolates the reference-driven effect, with 1,141 of 1,200 component counts changing despite an unchanged observed spectrum, whereas independent parallel analysis seeds change none of the corresponding decisions. By contrast, orthogonally invariant comparator scores remain numerically stable with similar held-out discrimination. Together, these results show that parallel analysis-derived component counts and decisions can reflect hidden-coordinate choice rather than a well-defined property of the model.

1 Introduction

The paper tests whether representation measurements remain invariant when hidden coordinates are changed without changing the model function. It finds that column-permutation parallel analysis can make component counts and downstream decisions depend on coordinate choice.

  • Motivation: Function-preserving reparameterizations should leave statistics interpreted as properties of the model function unchanged.The same input–output function can have multiple hidden-coordinate parameterizations.
  • Problem: Column-permutation parallel analysis preserves coordinate marginals but uses a reference distribution that changes with the chosen basis.Orthogonal reparameterization leaves the observed covariance spectrum unchanged while changing covariance diagonals used by the reference.
  • Study design: The study evaluates coordinate dependence in raw PA, implemented scoring rules, and fixed-threshold classifications of evidence-containing versus evidence-ablated contexts.The retrieval pipeline selects a context subspace with PA and scores answers by projection onto that subspace.
  • Structural result: Data-internal reference procedures cannot simultaneously preserve every coordinate marginal, remain right-orthogonally equivariant, and remove cross-coordinate covariance.Column-permutation PA is presented as one instance of this broader incompatibility.
  • Theoretical result: 15.2% of 3,600 context–layer spectra receive a nonzero orbit-wide range certificate for PA-selected component counts.The bound is sufficient rather than a sharp characterization and is inconclusive elsewhere.
  • Empirical study: The empirical study separates reference-driven instability from randomness and preprocessing alternatives across models, domains, and function-preserving transformations.The supplied passage introduces the controlled comparison, while the detailed outcomes are reported later in the paper.

2 Related Work

Related work grounds the paper’s invariance criterion in neural-network symmetries, representation comparison, identifiability, permutation methods, dimensionality estimation, and retrieval evaluation. These strands motivate testing whether representation-derived measurements are well-defined under function-preserving coordinate changes.

  • Representation invariance: Neural-network symmetries show that parameters and hidden representations can vary while the input–output function remains fixed.Examples include hidden-unit permutations, compensating rescalings, and orthogonal transformations absorbed into adjacent weights.
  • Representation comparison: RSA and CKA compare relational or Gram-matrix structure rather than individual activation coordinates, motivating orthogonally invariant comparators.CKA is described as invariant to orthogonal transformations and isotropic scaling.
  • Identifiability and interpretability: Identifiability studies show that representations may be determined only up to invertible linear transformations, while probes and neuron analyses can depend on analysis choices.This motivates explicit estimands and control tasks when interpreting representation measurements.
  • Permutation methods: Permutation-reference validity depends on the invariance or exchangeability assumptions of the tested null and the transformation-group construction.Studentization can restore asymptotic validity in some settings, but finite random-permutation exactness also depends on group construction.
  • Dimensionality estimation: Dimensionality estimators quantify different representation properties, so whether each estimand is invariant to hidden-basis changes is a separate question.The literature distinguishes intrinsic manifold dimension, covariance rank, effective rank, and factor count.
  • Retrieval evaluation: The retrieval case study treats evidence-containing contexts as sufficient and matched evidence-ablated contexts as insufficient for testing representation-score invariance.Held-out AUROC measures discrimination, while invariance under function-preserving transformations tests measurement validity.

3 Setup and Theoretical Analysis

The analysis defines function-preserving reparameterization invariance and shows that column-permutation parallel analysis can change with hidden-coordinate choice even when the model function and observed spectrum remain fixed. It proves a broader incompatibility for data-internal references and introduces orthogonally invariant comparator scores.

  • Setup: Function-preserving reparameterization requires statistics interpreted as properties of a model function to satisfy T(XR, AR) = T(X, A) and induces invariant threshold decisions.The admissible transformation class is architecture-specific: full O(d) for RMSNorm models and a mean-preserving subgroup for LayerNorm.
  • Coordinate Dependence of Column-Permutation Parallel Analysis: Column-permutation parallel analysis compares an orthogonally invariant observed spectrum with a reference covariance that depends on the hidden-coordinate system.Orthogonal transformations preserve singular values but can change the covariance diagonal used to construct the permutation reference.
  • Coordinate Dependence of Column-Permutation Parallel Analysis: Theorem 1 gives a spectrum-dependent lower bound on the range of selected component counts across the full orthogonal orbit.The lower endpoint is attained by a constant covariance diagonal, while aligning a coordinate with a leading eigenvector yields zero selected components.
  • Coordinate Dependence of Column-Permutation Parallel Analysis: Theorem 2 establishes different parallel-analysis component counts within the LayerNorm-compatible subgroup, despite identical singular values under compatible transformations.Together, the theorems show fixed observed spectra can coexist with basis-dependent permutation references and component counts.
  • Limitations: The theorems establish coordinate dependence of the PA estimator, not ambiguity in algebraic or covariance rank; more permutations improve reference estimation without equalizing coordinate-specific distributions.The sufficient bound can be zero even when actual counts vary, and its finite-sample guarantee is conservative rather than a sharp characterization.
  • Limits of Marginal-Preserving Reference Distributions: No data-internal reference constructed solely from X can simultaneously preserve every coordinate marginal, remain right-orthogonally equivariant, and alter X⊤X.The impossibility applies directly to the full orthogonal class used for RMSNorm, while the LayerNorm subgroup is handled by a separate construction.
  • Invariant Context–Answer Scores: Comparator scores built from jointly right-orthogonally invariant Gram-matrix quantities remain invariant under the specified transformations.Eigenvalue ties require explicit tie-safe treatment for the cumulative-variance principal-subspace score; ridge-shrinkage remains invariant with repeated eigenvalues.

4 Empirical Analysis

The empirical study tests whether function-preserving reparameterizations destabilize PA-based measurements, whether randomness or preprocessing explains the effect, and whether invariant alternatives retain discrimination. Across models and retrieval domains, component counts and downstream decisions change substantially, while centering-only controls isolate the reference-distribution mechanism and invariant scores remain stable.

  • Study design: The study evaluates PA-based measurements across five models, four architecture families, three retrieval domains, and five reparameterizations per model.The retrieval domains are FinanceBench, QASPER, and HotpotQA; paired conditions differ by annotated evidence while holding the question and answer fixed.
  • Reparameterization-induced instability: 0.79 is the median context-level selected-component-count disagreement across 75 model–domain–transformation comparisons.The median absolute component-count change is 2 (IQR 1–4; range 0–12), with original and reparameterized median counts of 39 and 37, respectively.
  • Reparameterization-induced instability: 0.26 is the median fixed-threshold decision disagreement rate, while within-question ordering reversal has a median rate of 0.10.The decisions concern evidence-containing versus evidence-ablated contexts using the threshold selected on the original calibration split.
  • Randomness and preprocessing checks: 0 of 1,200 matched decisions change under an independent permutation seed without reparameterization, compared with a median reparameterization-induced decision-disagreement rate of 0.29.The median no-reparameterization seed-change rate is 0.06, versus 0.76 for reparameterization; reparameterization is larger in all 12 model–domain combinations.
  • Context discrimination: Right-orthogonally invariant comparator scores remain numerically stable while retaining discrimination comparable to the PA-selected score.Across 15 model–domain combinations, all four scores have median held-out AUROC within a 0.023 band; the ridge-shrinkage score has median AUROC 0.712 versus 0.703 for the PA-selected score.
  • Robustness analysis: 0.87 is the median disagreement rate at the original robustness setting, and the median remains between 0.86 and 0.99 across the complete 72-setting grid.Increasing the permutation budget or changing the quantile, selection convention, or aggregation level does not materially reduce the instability.

5 Discussion and Limitations

The study interprets column-permutation PA instability as a measurement-validity problem rather than evidence that PA lacks descriptive value. Invariant comparators remain stable with similar discrimination, but the empirical evidence is limited in scope.

  • Discussion: Column-permutation PA can distinguish evidence-containing from evidence-ablated contexts, but its counts and decisions depend on hidden coordinates.The comparators remain numerically stable and serve as methodological baselines rather than proposed optimal retrieval scores.
  • Discussion: All four comparator scores lie within 0.023 median AUROC, and representation-based scores do not consistently outperform answer log-probability.
  • Limitations: The experiments cover five small open-weight models, three question-answering domains, and tested architecture-compatible transformations.The seed control uses one independent rerun, the standardization ablation one fixed transformation, and the robustness grid one fixed transformation per representative combination.
  • Limitations: These controlled choices do not exhaust larger models, other tasks, or all function-preserving transformations.The general marginal-preservation limitation is established for O(d), not extended to the LayerNorm mean-preserving subgroup.

6 Conclusion

The conclusion is that column-permutation PA is not invariant under function-preserving reparameterization. Although model behavior, total variance, and the observed covariance spectrum remain unchanged, coordinate-dependent references can change downstream measurements and decisions.

  • Conclusion: Column-permutation PA changes selected component counts, projection scores, and fixed-threshold decisions under function-preserving reparameterization.
  • Conclusion: The theoretical conflict is established for the full orthogonal class used with RMSNorm and a separate LayerNorm-compatible subgroup.
  • Conclusion: 0.26 median fixed-threshold decision disagreement occurs under reparameterization, compared with 0.00 in the independent-seed control.
  • Conclusion: 1,141 of 1,200 component counts change under centering alone despite an unchanged observed spectrum.
  • Conclusion: Right-orthogonally invariant comparators remove the tested sensitivity while retaining similar observed discrimination.All four scores lie within 0.023 median AUROC, and representation-based scores do not consistently outperform answer log-probability.

7 Code and Data Availability

The study makes its protocols, analysis code, intermediate results, and figure-generation scripts publicly available. The appendix specifies the implemented scorer, including standardization, component selection, projection scoring, comparator definitions, and numerical safeguards.

  • Code and Data Availability: Protocols, analysis code, intermediate results, and figure-generation scripts are publicly available.Pre-specified, post-hoc control, and robustness analyses are documented separately.
  • Implemented Pipeline: The implemented scorer standardizes coordinates, enforces a minimum component count of one, and converts the selected subspace into a context–answer score.
  • Implemented Pipeline: Coordinatewise sample standard deviations are floored at 10^-6 before standardization.
  • Implemented Pipeline: Reparameterization can change both the standardized observed spectrum and its permutation reference, so the experiments test the complete scoring pipeline.
  • Implemented Pipeline: The PA-selected subspace score projects target-answer tokens onto leading right singular vectors after selecting components at each layer.
  • Numerical Definitions: The implementation floors score denominators at 10^-12 and assigns zero to a zero answer vector.

Appendix B. Proofs of the Main Technical Results

The appendix collects the longer proofs omitted from the main presentation, providing the technical details behind the paper’s principal theoretical results.

  • Appendix B. Proofs of the Main Technical Results: The appendix records the longer proofs omitted from the main presentation.
  • Appendix B. Proofs of the Main Technical Results: These proofs supplement the main theoretical presentation with complete technical derivations.
  • Appendix B. Proofs of the Main Technical Results: Appendix B serves as the repository for omitted proof details supporting the main technical results.

B.1 Proof of Lemma 2

The proof characterizes the covariance moments of column-permutation references and uses them to derive spectral probability and threshold bounds.

  • For off-diagonal entries, the permuted covariance has the distribution of x_j^⊤Px_k/n under a uniform permutation.The argument uses sampling without replacement and column centering.
  • The diagonal of the permuted covariance equals the original covariance diagonal deterministically.
  • The proof sums off-diagonal second moments, applies Markov’s inequality to the squared Frobenius norm, and uses Weyl’s inequality for eigenvalues.
  • The quantile choice δ = α yields the stated threshold bound.

B.2 Proof of Theorem 1

The theorem constructs two orthogonal orientations with identical observed spectra but sharply different parallel-analysis outcomes: one forces zero retained components, while the other guarantees a spectrum-dependent lower bound.

  • Theorem 1: Choosing the leading eigenvector as the first coordinate makes every reference realization have largest eigenvalue at least λ1(S), so brPA(XRloc) = 0.Column permutation preserves each covariance diagonal entry, causing the strict first comparison to fail.
  • Theorem 1: A Schur–Horn construction produces an orientation whose diagonal is sufficiently flat for every componentwise threshold.
  • Theorem 1: The flat orientation retains at least mα(S, n) components while the observed eigenvalues remain λk(S).Lemma 2 supplies simultaneous threshold bounds for all components.

B.3 Proof of Theorem 2

The proof gives a LayerNorm-compatible pair of function-preserving orientations with the same rank-one observed spectrum but different first-component decisions.

  • Theorem 2: Both observed covariance matrices have spectrum (λ, 0, …, 0), and a mean-preserving orthogonal transformation maps the local orientation to the mixed one.
  • Theorem 2: For Xloc, every reference realization has spectrum unchanged, so t1 = λ and the strict first comparison fails.
  • Theorem 2: For Xmix, the largest covariance diagonal entry is 4λ/9, and the sample-size condition makes its reference bound strictly smaller than λ.
  • Theorem 2: The mixed orientation retains exactly one component because the first eigenvalue exceeds its threshold while the second observed eigenvalue is zero.

B.4 Proof of Theorem 3

The proof shows that exact marginal preservation and right-orthogonal equivariance force any data-internal reference output to be a row permutation of the input.

  • Theorem 3: For any unit vector v, exact marginal preservation fixes the first projected empirical measure of a reference draw from K(XR, ·).Right-orthogonal equivariance transfers the resulting degenerate distributional statement across orientations.
  • Theorem 3: Applying the argument on a dense subset of the sphere and using continuity extends the identity to every direction.
  • Theorem 3: The Cramér–Wold theorem then implies equal empirical measures, so equal atom weights force the reference output to satisfy Z = PX.
Loading 2608.27020v1…