Source-linked AI summary

Geometric Stability: The Missing Axis of Representations

Prashant C. Raju

arXiv:2601.09173v5cs.LGcs.CLq-bio.QMstat.ML

TL;DR

Similarity metrics compare representational alignment but do not measure whether a single representation’s geometry is reliably recoverable across feature subsets. The paper introduces Shesha, which measures this geometric stability, and finds it dissociates from similarity and transferability across controlled manipulations and vision-model benchmarks.

  • Problem

    Existing representation-similarity methods measure whether two systems encode similar structure, but not whether one system’s geometry remains reliably recoverable across feature subsets.

  • Method

    Shesha correlates dissimilarity matrices from complementary random halves of feature dimensions to measure whether a representation’s pairwise geometry is redundantly encoded across its basis.

  • Results

    Across benchmark regimes, geometric stability dissociates from similarity and transferability; under compression, the metrics anti-correlate at ρ = −0.47.

  • Takeaways & Limitations

    Geometric stability is a distinct diagnostic of representational redundancy and should not be treated as a general proxy for representation quality or corruption robustness.

  • Takeaways & Limitations

    Shesha is a global scalar and does not resolve localized instabilities affecting rare categories, low-frequency stimuli, or specific representation submanifolds.

Abstract

from arXiv · show

Representational similarity analysis and related methods compare the internal geometries of neural networks, but they measure only alignment between spaces, leaving a blind spot -- whether a representation's structure is reliably recoverable, not merely similar. We introduce geometric stability, a distinct axis, and \textit{Shesha}, a metric that quantifies it from a single representation by correlating dissimilarity matrices built from complementary random halves of the feature dimensions. Unlike CKA and Procrustes distance, Shesha is provably non-invariant to orthogonal rotations of the feature basis. This is by design: the basis is privileged for learned models, since probes, patching, and steering act on coordinates, and a rotation-invariant metric cannot see whether the targeted structure survives them. A double dissociation isolates the mechanism -- removing the top principal component collapses CKA while Shesha holds, whereas rotating a representation into its eigenbasis, which preserves the spectrum and CKA exactly, collapses Shesha. Across 2,463 encoder configurations in seven domains, the metrics are redundant under geometry-preserving transforms and anti-correlate under compression ($ρ=-0.47$). Across 170 vision models spanning 6 clean and 38 corruption-shifted datasets, DINOv2 ranks first or second in transferability on three of six clean datasets yet bottom-quartile in stability on five, an isolated dissociation rather than a trade-off.

1 Introduction

The paper introduces geometric stability as a distinct axis from representational similarity, measuring whether a single representation’s geometry is redundantly recoverable across feature coordinates. Shesha captures this property and reveals dissociations that similarity and transferability measures can miss.

  • Problem and distinction: Similarity metrics overlook stability because CKA, RSA, and Procrustes are invariant to orthogonal feature rotations, whereas Shesha is not.This basis sensitivity is intentional because interpretability interventions operate on feature coordinates.
  • Contribution and method: Shesha measures geometric stability by correlating RDMs from complementary random feature subsets, quantifying whether one representation’s pairwise geometry is redundantly recoverable across its basis.It averages Spearman rank correlations over K independent splits and is intentionally sensitive to orthogonal rotations.
  • Empirical validation: Across 2,463 encoder configurations in seven domains, similarity and stability are redundant under geometry-preserving transformations but anti-correlate under compression (ρ = −0.47).The pooled correlation is near zero (Spearman ρ = −0.01, 95% CI [−0.06, +0.03]), because the transformation regime determines the relationship.
  • Mechanism: Removing the top principal component collapses CKA to 0.27 while Shesha remains 0.95, demonstrating that similarity can fail while geometric stability is preserved.The complementary manipulation—retaining only the top component—forms the opposing half of the reported double dissociation.
  • Interpretability implications: DINOv2 has among the least recoverable geometry in the benchmark despite strong similarity and transfer signals, making coordinate-level probes, patching, and steering potentially unreliable.The warning is specific rather than universal: contrastively aligned models such as CLIP are described as both transferable and geometrically stable.

2 The Geometric Stability Framework

Geometric stability measures whether a representation’s pairwise geometry is recoverable from complementary feature subsets, revealing how information is distributed across its coordinate basis. Shesha is invariant to scaling and permutations but intentionally non-invariant to orthogonal rotations, unlike CKA and Procrustes.

  • Basis dependence: Shesha depends on both the eigenspectrum and coordinate orientation, measuring how geometric information is distributed across the feature basis rather than the eigenspectrum alone.Redundant variance across coordinates produces high stability, whereas asymmetric concentration in a few coordinates produces low stability; therefore broad spectra can still yield low S.
  • Invariance structure: Shesha is invariant to global scaling, monotonic distance transformations, and coordinate permutations, while its orthogonal non-invariance distinguishes it from basis-invariant similarity metrics.Spearman correlation supplies rank invariance, cosine distance supplies global-scaling invariance, and random partition exchangeability supplies permutation invariance.
  • Invariance structure: Under orthogonal rotation, CKA remains exactly 1 because XX⊤ is preserved, whereas Shesha generally changes because rotation redistributes geometric information across coordinate axes.This formal dissociation makes Shesha sensitive to basis-level structure that CKA and Procrustes cannot detect.
  • Relation to RSA: Unlike RSA noise ceilings, which assess measurement reliability across observations, Shesha assesses whether a representation’s geometry is reliably recoverable across feature subsets.The two diagnostics are complementary: noise ceilings audit data reliability, whereas Shesha audits geometric reliability.

3 Distinctness of Stability and Similarity

This section establishes that Shesha captures geometric stability as a property distinct from representational similarity, formally through basis-rotation sensitivity and empirically through controlled dissociations. Across synthetic and large-scale encoder studies, similarity and stability separate, especially when variance is concentrated across coordinates.

  • Synthetic validation: Shesha recovers controlled stability nearly perfectly (ρ = 0.997, p < 10−86), while cases with CKA > 0.97 and near-zero S show that similarity does not imply stability.Balanced sampling across the stability–similarity space also confirms the converse: high stability need not imply high similarity.
  • Spectral sensitivity: Removing one leading principal component drops CKA from 1.0 to 0.27 while Shesha remains 0.95; after 26 removals, Shesha retains roughly 92× CKA’s signal.With a power-law eigenspectrum, Shesha decays gradually as leading components are stripped, whereas CKA depends strongly on the spectrum’s head.
  • Spectral compression: Retaining 300 PCA components restores vision CKA to 0.972 but changes Shesha from +0.732 to −0.136; language CKA likewise recovers to 0.993.As retained dimensionality increases, CKA rises toward its full-rank value while Shesha moves oppositely into negative values.
  • Transformation regimes: Geometry-preserving transformations couple CKA and Shesha positively (ρ = +0.75, N = 1,278), whereas PCA compression is the sole negative regime (ρ = −0.47, N = 948).The negative coupling occurs when variance concentrates in a low-dimensional, axis-aligned subspace, keeping CKA high while Shesha collapses.
  • Large-scale validation: Across 2,463 encoder configurations in seven domains, the dissociation reproduces domain by domain, with 15 random seeds per configuration and linear CKA comparisons.A mixed-effects model controlling for base-model identity attributes under 10% of stability variance to encoder identity (ICC = 0.10).
  • Substrate dependence: The dissociation extends across substrates: protein encoders show ρ = −0.36, while molecular profiles show negligible correlation (ρ = +0.06).In proteins, the negative relationship is driven by PCA compression of low-dimensional encoders; the molecular-profile result is consistent with the natural encoding regime described in the passage.

4 Geometric Stability in Pretrained Vision Models

Across 170 pretrained vision models, geometric stability reveals dissociations from transferability, effective dimensionality, accuracy, and CKA. SheshaFS measures whether geometry is recoverable from random coordinate subsets, exposing basis-dependent structure relevant to feature-subset methods.

  • Reproducibility: SheshaFS is highly reproducible across three CIFAR-10 feature-partition seeds, with Spearman ρ ≥0.993 and median per-model CV 0.75% across all 170 models.DINOv2 also retains the lowest family-mean SheshaFS under all three seeds.
  • Transferability versus stability: DINOv2 ranks first or second in transferability on three of six datasets but bottom-quartile in geometric stability on five, revealing an isolated dissociation.It ranks 36/36, 35/36, and 36/36 in stability on Flowers-102, CIFAR-10, and CIFAR-100, respectively, and 33/36 on Oxford Pets and 29/36 on DTD.
  • Architecture and training objective: Contrastive alignment predicts higher stability: CLIP-family models outperform self-supervised models on all six datasets, while EVA-02 ranks among the most stable models.The alignment target, rather than the training mechanism, determines geometric stability.
  • Stability versus effective dimensionality: SheshaFS and effective dimensionality diverge because participation ratio is rotation-invariant, whereas SheshaFS depends on variance distribution across learned coordinates.DINOv2 has the lowest SheshaFS yet the highest participation ratio among 36 families on CIFAR-10: 98.98 versus a benchmark mean of 51.64.

5 Discussion · Appendix A. Shesha Variants · A.1 Feature-Split Shesha (SheshaFS)

The discussion frames geometric stability as a second, basis-sensitive axis alongside similarity, revealing dissociations relevant to model selection and interpretability. It also establishes SheshaFS as the primary feature-split variant and identifies its scope, limitations, and cross-domain applicability.

  • 5 Discussion · 5.1 Two Axes of Representational Geometry: Geometric stability is a distinct axis from similarity because it measures whether a single representation’s geometry is reliably encoded across its coordinate basis.CKA and related metrics are rotation-invariant, whereas stability depends on how variance is distributed across coordinates.
  • 5.1 Two Axes of Representational Geometry: Under geometry-preserving transformations the axes are redundant, but under compression similarity remains high while stability falls, showing that stability adds information when variance is concentrated.This supports assessing stability before similarity, because low S means pairwise geometry may not be recoverable from independent feature subsets.
  • 5.1 Two Axes of Representational Geometry: Rotation sensitivity is intentional for learned representations because probes, patching, and steering operate in the privileged model coordinate basis rather than an arbitrary basis.Whitening is a boundary condition: by equalizing the eigenspectrum, it makes stability and similarity partially redundant for whitened representations.
  • 5.2 Geometric Stability as a Distinct Selection Axis: DINOv2 uniquely combines top-tier transfer with bottom-quartile stability, while transferability and stability show no general trade-off across architectural families (Theil-Sen ρ = +0.21, not significant).Contrastive families such as CLIP and SigLIP achieve high transfer and high stability together.
  • 5.2 Geometric Stability as a Distinct Selection Axis: Because stability is independent of transferability, it should be included as a separate model-selection criterion alongside transfer metrics.DINOv3’s Gram anchoring recovers stability at little or no transfer cost on four of six datasets, while CLIP-family models are more stable than single-modality self-supervised models on all six.
  • 5.3 Relation to Mechanistic Interpretability: A low stability score warns that probe, patching, or steering targets may depend on sampled features rather than robust global geometry.The discussion links DINOv2’s lower stability to non-redundant coordinate-basis encoding, while noting that this association does not establish the training objective’s causal role.
  • 5.4 Geometric Stability Across Substrates: Geometric stability generalizes across protein, molecular, and neural representations, distinguishing fragile representational geometry from noisy RDM estimation.Protein encoders reproduce the negative stability–similarity correlation under PCA compression, while neural recordings show the negligible-correlation natural-encoder regime.
  • 5.5 Limitations: Shesha is global rather than localized, and its broader empirical scope is limited by single-seed vision estimates and incomplete testing of realistic shifts and interventions.The method may miss fragile rare-category submanifolds; CIFAR-10 rankings were stable across three seeds, but the other five vision datasets remain point estimates.

A.2 Sample-Split Shesha (SheshaSS)

Sample-Split Shesha (SheshaSS) partitions data points into disjoint subsets and correlates their within-subset representational dissimilarity matrices to assess robustness to input variation. The variant is included for completeness but is not used in the present paper.

  • Method: SheshaSS partitions data points into two disjoint subsets and computes a representational dissimilarity matrix within each subset.
  • Method: It evaluates correlation on overlapping sample pairs or through anchor-based approaches, thereby measuring robustness to input variation across subsets.
  • Interpretation and scope: A low SheshaSS value may indicate sensitivity to sampling noise or reliance on spurious input-specific information, but this variant is not used here.

Appendix B. Invariance Proofs and Counterexample … B.6.1 Numerical Confirmation at Fixed Spectrum

Appendix B proves Shesha’s invariance to scaling, feature permutations, and monotone distance transformations, while showing that PCA compression and orthogonal rotations can reduce it despite preserved representational geometry. A constructive counterexample and a fixed-spectrum numerical setup establish the intended basis sensitivity.

  • B.1 Global Scaling Invariance; B.2 Isotropic Scaling Invariance: Shesha remains unchanged under global and isotropic positive scaling because cosine distances are preserved.
  • B.3 Feature Permutation Invariance: Shesha is invariant to feature permutations because uniformly random partitions relabel coordinates without changing the partition distribution.
  • B.4 Monotonic Distance Invariance: Shesha is invariant to strictly increasing transformations of distances because Spearman correlation depends only on pairwise-distance rankings.
  • B.5 Non-Invariance to PCA Compression: PCA compression can lower Shesha even when CKA remains approximately 1, because random halves may lose all informative coordinates and become degenerate.
  • B.6 Non-Invariance to Orthogonal Transformations: Constructive Counterexample: An explicit orthogonal rotation yields S(XQ) ≠ S(X) while preserving linear CKA exactly at 1.
  • B.6 Non-Invariance to Orthogonal Transformations: Constructive Counterexample: The counterexample generalizes: rotations that concentrate column energy into fewer coordinates reduce Shesha according to how unevenly geometry is distributed across axes.
  • B.6.1 Numerical Confirmation at Fixed Spectrum: A numerical fixed-spectrum experiment constructs a 200×64 representation from a five-dimensional Gaussian latent through a dense random map, obtaining S(X) = 0.903 before rotation into its eigenbasis.

Appendix C. Connection to RSA Noise Ceiling … E.5.1 Protocol

Across its appendices, Shesha is shown to measure feature-axis geometric stability distinctly from similarity, with reliable convergence, strong ground-truth recovery, and sensitivity to dimensionality and spectral structure. The protocol adapts RSA split-half machinery to feature partitions and requires only a single representation matrix.

  • Appendix C. Connection to RSA Noise Ceiling: Shesha adapts RSA noise-ceiling split-half correlations from observations to feature dimensions, assessing whether geometry is redundantly encoded using only one matrix X ∈ Rn×d.RSA evaluates measurement reliability across repeated observations, whereas Shesha evaluates recoverability across complementary feature subsets.
  • Appendix D. Shesha Computation: All SheshaFS estimates used complementary equal-size feature halves, cosine-distance RDMs, Spearman correlation of upper triangles, and K=30 random partitions averaged together.For odd feature counts, one half received (d+1)/2 features.
  • Appendix E. Ground Truth Validation: ρ = 0.997 for ground-truth stability recovery, while balanced designs reduce SheshaFS–debiased CKA correlation to ρ = 0.204, supporting sensitivity and discriminant validity.These results establish that SheshaFS tracks known stability while remaining distinct from representational similarity.
  • E.1 Convergence Over K and Subsampling: |¯∆| = 0.0115 across 30 model–dataset combinations, and all 15 models achieved stable estimates at n = 400 versus n = 1,600.Per-model mean drifts ranged from 0.0002 for ResNet-50 to 0.0176 for ViT-Tiny, with mean 0.0077.
  • E.2 Dimensionality Sensitivity: −0.112 versus +0.620 after reduction to 64 PCA components, showing that compression prevents pairwise geometry from being redundantly recovered across arbitrary coordinate halves.The range after reduction was [−0.204, −0.055] across 30 conditions.
  • E.3 Sensitivity to Known Stability Levels: ρ = 0.997 between SheshaFS and parametrically controlled stability across 21 signal-to-noise levels, confirming near-perfect monotonic recovery of the ground truth.The synthetic representations mixed low-rank signal with isotropic noise while varying α from 0 to 1.
  • E.4 Spectral Deletion: At k = 1 removed principal component, CKA, PWCKA, and Procrustes all fell below 0.5, whereas SheshaFS retained sensitivity to the remaining eigenspectrum.The experiment used representations with a power-law eigenspectrum and progressively removed leading components.
  • E.5 Tail Noise Ablation: The tail-noise ablation injected isotropic Gaussian noise into 478 spectral-tail components after the top 34 components explained 90% of variance, testing whether SheshaFS produces false alarms.Noise was varied across 14 scales from σ ∈ [0.001, 50], with each condition repeated using three independent draws.

E.5.2 Results · E.5.3 Interpretation

SheshaFS remains stable under amplified non-functional tail noise while CKA and Shesha RDM similarity degrade in parallel, indicating no false alarms. This contrasts with its detection of genuine tail-structure removal, explained by Spearman rank-order correlation preserving distance rankings under monotone noise changes.

  • E.5.2 Results: SheshaFS holds at its baseline value of 0.971 when noise is low (σ ≤0.01), while CKA and Shesha RDM similarity exceed 0.999.
  • E.5.2 Results: At moderate noise (σ = 0.05), CKA and Shesha RDM similarity remain above 0.99, showing limited sensitivity to non-functional tail-noise amplification.
  • E.5.2 Results: CKA falls below 0.95 at σ = 0.20, whereas Shesha RDM similarity crosses that threshold slightly earlier at σ = 0.10.
  • E.5.2 Results: Across the full noise range, CKA and Shesha degradation profiles track in parallel, with no regime where Shesha detects a change that CKA misses.
  • E.5.3 Interpretation: SheshaFS detects spectral-tail deletion but not tail-noise amplification because deletion alters pairwise distance rankings whereas amplification preserves them.
  • E.5.3 Interpretation: The underlying mechanism is SheshaFS’s Spearman rank-order correlation, which distinguishes genuine geometric changes from energetic perturbations that preserve rankings.

E.6 Preprocessing Ablation · E.6.1 Mechanistic Interpretation of Whitening

The Shesha–CKA divergence is robust to raw, centered, and L2-normalized preprocessing but breaks under whitening, which equalizes the spectrum and amplifies noise. Whitening therefore provides a mechanistic explanation for the divergence’s dependence on spectral structure.

  • E.6 Preprocessing Ablation: The preprocessing ablation tested raw, centered, centered with L2 normalization, and whitened representations.Whitening used ZCA with shrinkage λ = 0.1, following Walther et al. (2016).
  • E.6 Preprocessing Ablation: The Shesha–CKA divergence persists under raw, centered, and L2-normalized preprocessing, but whitening equalizes the spectrum and removes the divergence.This robustness pattern is reported across the preprocessing conditions in Table S1.
  • E.6 Preprocessing Ablation: The divergence is therefore robust across preprocessing choices that preserve spectral structure.The reported exception is whitening, which equalizes the spectrum.
  • E.6.1 Mechanistic Interpretation of Whitening: At k = 30, whitened CKA remains negative at −0.054, compared with −0.076 under raw preprocessing.Whitening weakens but does not eliminate the negative CKA value at this ablation level.
  • E.6.1 Mechanistic Interpretation of Whitening: Whitening reduces the Shesha baseline from 0.98 to 0.50 at k = 0.The drop reflects noise amplification caused by spectral equalization.
  • E.6.1 Mechanistic Interpretation of Whitening: Whitening’s mechanistic effect is consistent with spectral equalization amplifying noise in Shesha.The whitened Shesha baseline falls at k = 0 while CKA remains negative at k = 30.

E.7 Comparison with RSA Reliability Methods … F.3.2 Procrustes Similarity

Across reliability checks, geometric interventions, domains, and alternative similarity metrics, Shesha remains reproducible and distinct from CKA-like similarity measures. Its divergence is driven by basis-dependent spectral structure: geometry-preserving transforms preserve both metrics, whereas PCA-style variance concentration preserves similarity while collapsing Shesha.

  • E.7 Comparison with RSA Reliability Methods; B. Shesha vs Whitened Variant; C. Shesha Across Preprocessing: ρ = 1.000 between standard and whitened Shesha, while raw-space Shesha avoids whitening-related numerical instability and noise amplification.This connects Shesha to established RSA reliability practices without requiring whitening.
  • A. Stability vs Similarity Metrics; D. CKA Across Preprocessing: Removing top principal components collapses similarity metrics immediately but causes Shesha to degrade gracefully, preserving sensitivity to spectral-tail structure.An orthogonal rotation that concentrates variance into a coordinate subset lowers SheshaFS while leaving CKA unchanged, establishing basis dependence as the cause.
  • E.8 Seed Stability: 0.0047 mean seed sensitivity across 30 model–dataset combinations, with a maximum of 0.0142, confirms reproducible Shesha estimates from K = 30 random splits.All combinations were below the 0.05 stability threshold, and 25/30 were below 0.01.
  • E.9 Dissociation with Balanced Quadrant Sampling: ρ = 0.204 between Shesha and debiased CKA under balanced quadrant sampling, demonstrating that the metrics assess largely different attributes.The four quadrants separately realize high/high, high/low, low/low, and low/high stability–similarity combinations.
  • F.1.4 Video (N=128): ρ = −0.24 versus ρ = −0.27 in a preliminary alternate-source video analysis, suggesting that the stability–similarity relationship is robust to video-source diversity.The analysis used 100 clips from a single-source Jellyfish video with the same base models and preprocessing as the main video analysis.
  • Appendix F. Encoders; F.1 Cross-Domain Validation: Data Sources and Preprocessing; F.2 Encoder Transformations; F.2.1 PCA; F.2.2 Random Projection; F.2.3 Top-Variance Feature Selection; F.2.4 Random Feature Subsets; F.2.5 Gaussian Noise Injection; F.2.6 Normalization: 2,463 configurations across seven domains show that distance-preserving transforms keep CKA and Shesha high, whereas PCA compression keeps CKA high but collapses Shesha.The negative coupling arises only in the variance-concentration regime, while pooled correlations are near zero because the regimes differ.
  • F.3 Similarity Metrics; F.3.1 Effective-Rank Projection-Weighted CKA (PWCKA); F.3.2 Procrustes Similarity: Alternative similarity metrics in the language domain all maintain |ρ| < 0.30 with Shesha, extending their distinctness beyond CKA.The comparison includes effective-rank projection-weighted CKA and Procrustes similarity, with Procrustes aligning representations by an optimal orthogonal transformation after unit-Frobenius normalization.

F.4 Statistical Methods … G.3 Transferability Metrics

The appendix establishes statistically distinct stability and similarity across 2,463 encoder configurations, then details the 170-model vision benchmark and transferability measurements. Its analyses use bootstrap inference, mixed-effects controls, nonparametric comparisons, and dataset-wise testing procedures.

  • F.4.1 Bootstrap Inference: Distinctness was assessed with Spearman correlations and 10,000 bootstrap replicates resampling encoder configurations within each domain.Confidence intervals are bootstrap percentile intervals at 95%.
  • F.4.2 Mixed-Effects Models: The mixed-effects model attributes less than 10% of stability variance to base-model identity, with CKA’s slope on Shesha equal to −0.03 (95% CI [−0.08, +0.02]).The model uses Shesha as outcome, debiased CKA as fixed effect, and base model as a random intercept.
  • F.4.3 Mann-Whitney U Tests: Architectural comparisons used two-sided Mann-Whitney U tests on SheshaFS scores with exact p-values.The comparisons covered contrastive versus self-supervised and hierarchical versus columnar architectures.
  • F.4 Statistical Methods; F.4.4 Multiple Comparisons: Per-dataset vision tests used no multiplicity correction because each dataset was treated as an independent evaluation domain.The distinctness analysis also reports correlations by encoder type and domain.
  • F.4.4 Multiple Comparisons: ρ = −0.01 (95% CI [−0.06, +0.03]) across 2,463 encoder configurations and seven domains, confirming negligible net correlation between Shesha and CKA.Robustness checks maintain |ρ| < 0.10, while aProtein’s moderate negative correlation is driven by PCA on low-dimensional sequence encoders.
  • Appendix G. Vision Benchmark: Extended Results; G.1 Model Selection: The vision benchmark selected 170 pretrained timm models across training objectives, architectural families, and model scales.The models included supervised, self-supervised, contrastive, generative, columnar, hierarchical, hybrid, and convolutional families.
  • G.2 Feature Extraction: Penultimate-layer features were extracted from fixed random dataset subsets using seed 320 and each model’s standard preprocessing transform.Sample sizes ranged from 1,500 to 5,000 images across the six datasets, with replacement for smaller Flowers-102.
  • G.3 Transferability Metrics: LogME was computed on the extracted features, while LEEP was computed for models with classification heads.Both transferability metrics used the same extracted features; LogME followed the authors’ implementation.

G.4 Extended Results

Extended results show a DINOv2 paradox: strong transferability coexists with low geometric stability across most datasets, except EuroSAT. They also find higher stability for contrastive than self-supervised models across all six datasets.

  • DINOv2 paradox: Bottom-quartile SheshaFS stability for DINOv2-giant occurs on every dataset except EuroSAT, despite top-six transferability on CIFAR-10, CIFAR-100, and Flowers-102.On EuroSAT, DINOv2-giant is among the most stable models, ranking 4/170.
  • Model-family stability: Contrastive models are more stable than self-supervised models on all six datasets.The comparison includes 29 contrastive models and 41 self-supervised models, evaluated with Mann–Whitney U tests.
  • Family-level confirmation: Across five datasets, DINOv2 occupies the high-transfer, low-stability corner, whereas on EuroSAT it ranks among the most stable families.These per-dataset family rankings provide within-family confirmation of the concentration-stability relationship described in Section 4.4.

G.5 Seed Stability of Vision Models · G.6 SAM vs. SGD Ablation: Full Protocol and Extended Results · G.6.1 Protocol

SheshaFS rankings are highly reproducible across feature-partition seeds, preserving the DINOv2 stability paradox despite negligible sub-percent variation. The SAM ablation tests whether changing optimization geometry affects SheshaFS using controlled ResNet-18 experiments, with representations and metrics extracted under a fully specified protocol.

  • G.5 Seed Stability of Vision Models: The seed-stability sweep recomputed CIFAR-10 scores for all 170 models using three independent seeds, each generating K = 30 random feature partitions.Per-model scores were summarized by their mean, standard deviation, and coefficient of variation.
  • G.5 Seed Stability of Vision Models: Rankings are nearly seed-invariant: SheshaFS achieves Kendall τ ∈[0.944, 0.945] and Spearman ρ ∈[0.993, 0.995] across seeds, with median CV 0.75%.Only 5 of 170 models exceed a 5% SheshaFS CV, supporting practical reproducibility despite a small Friedman-test seed effect.
  • G.5 Seed Stability of Vision Models: The DINOv2 paradox is seed-invariant: the register-augmented large variant ranks 170 of 170 under all three seeds, and DINOv2 has the lowest family-mean SheshaFS each time.Its SheshaFS values are 0.291, 0.287, and 0.297 across the three seeds.
  • G.5 Seed Stability of Vision Models: A Friedman test detects a small SheshaFS seed effect (χ2 = 8.48, p = 0.014), but negligible magnitude and near-unchanged rankings motivate reporting seed 320 in the main text.The reported practical-stability indicators are median CV 0.75%, Spearman ρ ≥0.993, and Kendall τ ≥0.944.
  • G.6 SAM vs. SGD Ablation: Full Protocol and Extended Results: The SAM ablation varies only the perturbation radius ρ ∈{0, 0.01, 0.02, 0.05, 0.1, 0.2}, with ρ = 0 recovering standard SGD, across 15 random training seeds.All configurations use identical ResNet-18 training hyperparameters; SAM tests whether SheshaFS responds to optimization geometry when accuracy and learned features remain approximately constant.
  • G.6.1 Protocol: The protocol trains CIFAR-10 and CIFAR-100 ResNet-18 models for 100 epochs with learning rate 0.05, momentum 0.9, weight decay 5 × 10−4, and batch size 128.Cosine annealing is used, and the CIFAR adaptation replaces the usual initial setup with a 3×3 convolution and identity max-pooling.
  • G.6.1 Protocol: After training, the study extracts 512-dimensional penultimate-layer representations from 2,000 test images and computes accuracy, debiased linear CKA, SheshaFS with K = 30 splits, and three additional metrics.The representations are taken after average pooling and CKA is measured against the SGD baseline.

G.6.2 Results … G.7.3 Stratified analysis

SAM separates SheshaFS from CKA while producing an interior, dataset-dependent optimum and reducing label-aware geometric concentration. Across 170 vision models, SheshaFS predicts probe instability, with the relationship strongest at intermediate accuracy and attenuated at low accuracy.

  • G.6.2 Results: SAM dissociates SheshaFS from CKA: CKA declines as ρ increases, while SheshaFS exceeds the SGD baseline at the peak radius on every seed in both datasets.CKA falls from 1.000 to 0.925 on CIFAR-10 and from 1.000 to 0.772 on CIFAR-100; SheshaFS gains +0.067 on CIFAR-10 and +0.018 on CIFAR-100 at their peak radii.
  • G.6.2 Results: The SheshaFS optimum is interior and dataset-dependent, peaking at ρ = 0.05–0.1 on CIFAR-10 before declining at ρ = 0.2.The CIFAR-10 peak is 0.872, while the value at ρ = 0.2 is 0.851; the variance ratio reverses correspondingly, bottoming at 0.786 before rising to 0.790.
  • G.6.2 Results: Label-aware variants move opposite to SheshaFS, with variance and class-separation ratios declining while supervised alignment remains approximately constant.On CIFAR-10, the variance ratio falls from 0.849 to 0.790 and class-separation ratio from 3.08 to 2.41, while supervised alignment stays near 0.51; CIFAR-100 follows the same pattern.
  • G.7.1 Protocol: The probe-sensitivity protocol fixes a stratified CIFAR-10 sample split and varies only random feature halves across 170 models, isolating feature-subset effects.Representations contained 512–1536 dimensions; each model used 20 random halves with subset fraction 0.5.
  • G.7.2 Headline Result: SheshaFS predicts probe-accuracy variability across 170 vision models, including after controlling for mean accuracy, supporting geometric stability as a practical diagnostic.The correlations are ρ = −0.302 for standard deviation, ρ = −0.260 for range, and ρ_partial = −0.382 after controlling for mean probe accuracy.
  • G.7.3 Stratified analysis: The SheshaFS–probe-variability relationship is strongest in the middle accuracy tercile and weakens at high and low accuracy.Correlations are ρ = −0.473 for middle accuracy, ρ = −0.261 for high accuracy, and ρ = −0.072 for low accuracy; DINOv2 models have SheshaFS ≈0.29 to 0.37 but consistently mediocre probes.
Loading 2601.09173v5…