Source-linked AI summary
More Data Cannot Break a Symmetry: Identifiability by Design
Jing Xu, Christopher Kanan
TL;DR
The paper asks how stimulus-geometry symmetry limits unsupervised correspondence recovery before data exist. It formalizes the symmetry as a design-time diagnostic and tests an intervention in colour. Symmetric designs remain ambiguous despite more optimisation, whereas changing the sampled colours cuts catastrophic failures from 75% to 2%.
Problem
Stimulus-geometry automorphisms can make unsupervised alignment non-identifiable, but their consequences for experimental design require a diagnostic available before data collection.
Method
The paper formalizes automorphisms and catastrophic relabelling cost, then evaluates symmetry-breaking colour designs across candidate geometries and representations.
Results
75% to 2%: choosing nine colours by the design diagnostic cuts catastrophic alignment failures across 93 model representations.
Takeaways & Limitations
Identifiability should be checked during experimental design by testing candidate geometries and breaking invariant sampling structures when catastrophic cost is zero.
Takeaways & Limitations
Real observers’ hue asymmetry costs 0.26 for a rotation versus 0.003 under the HSV cylinder, so the nominal design degeneracy is not exercised by those data.
Abstract
from arXiv · showhide
Unsupervised representational alignment recovers a stimulus-by-stimulus correspondence from geometry alone, but the automorphism group of the stimulus geometry bounds what any such alignment can identify, before data exist. The obvious diagnostic for this degeneracy, the cheapest non-identity relabelling, ranks two published designs in the wrong order, because dense sampling creates near-duplicates whose transposition is nearly free. We turn this known invariance (Demetci et al., 2024) into a design-time diagnostic and intervention. In colour, where candidate geometries have closed form, we show that the failure is structural: sixty-four times the restart budget leaves a symmetric design unmoved while an asymmetric set at the same N recovers every time. Discriminating representational models and recovering a correspondence are essentially uncorrelated objectives (r = -0.02 over 3,000 subsets). Choosing nine colours by this diagnostic alone, without consulting any learned representation, moves all 93 model representations away from the degenerate point and cuts catastrophic alignment failures from 75% to 2% with the models, the layers, N and the solver all held fixed. The same risk arises wherever a regular design meets its candidate geometry's isometry group, including evenly spaced orientations, tones, or motion directions, and the check costs one function call before data collection.
1. Introduction
Unsupervised alignment can recover stimulus correspondences from geometry, but stimulus-geometry symmetries limit identifiability before data collection. The paper makes this limitation an experimental-design variable and reports four colour-based consequences.
- Unsupervised alignment recovers stimulus correspondences by minimising a Gromov–Wasserstein objective over permutations.
- Stimulus-geometry automorphisms can make the recovered correspondence non-unique, so identifiability is partly controlled by experimental design.
- The paper shows that the cheapest non-identity relabelling can rank published colour designs incorrectly.
- The paper finds structural failure rather than computational failure, weak association between model discrimination and correspondence recovery, and a design intervention reducing catastrophic failures from 75% to 2%.
2. Identifiability of an unlabelled correspondence
The paper formalizes identifiability through the automorphism group of a stimulus geometry and shows why the cheapest relabelling is a misleading diagnostic. A catastrophic-cost measure instead captures how damaging an ambiguity can be.
- The automorphism group consists of relabellings that leave every pairwise dissimilarity unchanged.
- For any alignment permutation, composing with a stimulus-geometry automorphism leaves the Gromov–Wasserstein objective unchanged.
- Nine equally spaced hues under the HSV cylinder have rotational symmetry, making the objective flat across all nine rotations.
- The cheapest non-identity relabelling gives 0.0794 for the nine-colour set under CIELAB and 0.0025 for the 93-colour set, reversing the published-design ranking.
- The catastrophic cost is the cheapest relabelling that moves at least half the stimuli beyond one quarter of the set diameter.
- Under this measure, the nine-colour set scores exactly 0 while the 93-colour set scores 0.33.
3. Structural failure, not optimisation failure
Restarting the optimiser cannot resolve ambiguity caused by a symmetric stimulus geometry. In colour, asymmetric sampling succeeds under the same protocol, while real observer judgements avoid the nominal degeneracy through hue asymmetry.
- Increasing the restart budget from 5 to 320 raises real nine-colour judgements from 62% to 100% exact recovery.
- Under the HSV cylinder, rotations of the nominal nine-colour set remain within 0.003 cost of the identity.
- With sixty-four times the search, the synthetic ring changes from 0.638 to 0.623, while the 93-colour displacement falls from 0.155 to 0.052.
- Nine irregular points with the same N, noise, and protocol recover at every budget and never catastrophically.
- Across 300 repetitions, every ring solution is an exact dihedral element at cost 0.0000, whereas real judgements return the identity every time.
- Exact matching accuracy is 0% on the 93-colour data at every budget, even as displacement reaches 0.052 and no run is catastrophic.
- Real observers’ hue asymmetry costs 0.26 for a rotation versus 0.003 under the HSV cylinder, so the design’s nominal degeneracy is not exercised.
4. Designing a set that can answer the question
Breaking stimulus-geometry symmetries improves correspondence identifiability without sacrificing model discrimination. A design-time choice of nine colours sharply reduces catastrophic alignment failures.
- 0.000 against 0.188 and 0.231: varying saturation and value lifts catastrophic cost at N = 9, whereas adding hues along the invariant submanifold does not.More hues at S = V = 1 leave cost exactly zero from N = 6 to N = 72.
- r = −0.02: discriminability and catastrophic cost are essentially uncorrelated across 3,000 random nine-colour subsets.None of the 3,000 subsets is degenerate; the published set is the only zero.
- 75% to 2%: changing only which nine colours are shown cuts catastrophic alignment failures across 93 model representations.The prescription holds the models, layers, N and solver fixed while selecting the design before data collection.
Appendix A. Robustness of the diagnostic
The diagnostic is robust to threshold choices but depends on the candidate geometry and on supplying its symmetry group. It generalizes from regular colour designs to other symmetric and semantic stimulus sets, with nonzero searches remaining non-certifying.
- Robustness of the diagnostic: Zero is robust across all 15 threshold combinations for the nine-colour hue ring, while the 93-colour set is never below 0.31.The published nine-colour set scores zero in 14 of 15 cells.
- Robustness of the diagnostic: The worst-case cost changes with the candidate geometry: removing the HSV cylinder raises it to 0.040, versus 0.159 under CIELAB alone.The HSV cylinder is therefore decisive for the reported zero.
- Robustness of the diagnostic: Regular designs sampled invariantly under isometries can be exactly degenerate, whereas 200 THINGS objects score 0.224 rather than zero.The diagnostic can screen semantic sets, but a nonzero search result does not prove that no degeneracy exists.
- Robustness of the diagnostic: A zero is a proof, but a nonzero minimum is only an upper bound unless the relevant symmetry group is included in the relabelling pool.Random relabellings can miss structured symmetries such as an icosahedron’s antipodal map.
- Robustness of the diagnostic: At matched N = 9, the hue-circle design remains at zero while departures from the invariant submanifold lift catastrophic cost immediately.The published 93-colour set follows the same curve at its own N, so more data alone does not explain the improvement.
- Robustness of the diagnostic: Before data collection, the prescription is to compute the worst-case catastrophic cost across candidate geometries and vary sampling if it is zero.This design-time check targets the invariant structure rather than adding stimuli along it.
Appendix B. Model representations
Across 93 model representations, stimulus design substantially affects catastrophic alignment, while model-discrimination performance does not predict correspondence recovery.
- The published nine-colour set has median catastrophic cost 0.123 and minimum 0.029 across 93 representations; no representation falls below 0.01.Five of twelve models have at least one layer closer to zero than human observers’ cost of 0.085.
- Matched-size colour-solid draws have median cost 0.163 versus 0.123 for the published ring, with no draw below 0.05 versus 3.2% of ring representations.Representations, layers, and stimulus count were held fixed while only the stimulus set differed.
- Changing only the nine displayed colours raises exact recovery from 8.6% to 52.7%, lowers mean displacement from 0.415 to 0.034, and cuts catastrophic failures from 75.3% to 2.2%.The designs were selected from candidate geometries before consulting the 93 unseen representations, with models, layers, N, and solver fixed.
- The geometry-discrimination design matches a random draw for correspondence, with mean displacement 0.163 versus 0.164 (p = 0.99), whereas the correspondence-optimised design performs better (p < 10^-10).Discriminating candidate geometries therefore does not itself identify a stimulus correspondence.
- Leaving the degenerate point is insufficient: random colour-solid draws score 0.175 yet still leave 15.1% of representations with catastrophic correspondences.Reaching 2.2% requires a design chosen specifically for correspondence recovery.
- The correspondence-optimised design reduces median inter-representation distance from 0.555 to 0.101, making models agree rather than appear more separable.The published ring’s apparent model separation is accompanied by split-half instability, with single representations failing to agree three times in four.
Appendix C. Orientation
Evenly spaced orientation stimuli form a highly symmetric regular twelve-gon, producing severe correspondence degeneracy in learned representations. Uneven spacing helps, whereas small jitter raises design cost without improving alignment.
- Twelve evenly spaced orientations form a regular twelve-gon carrying D12 exactly under the π-periodic embedding (cos 2θ, sin 2θ).The experiment compares evenly spaced, jittered, and unevenly spaced gratings with other image properties held identical.
- The closest representation comes within 0.0001 of the degenerate point, and half of the 93 representations score below 0.05 under the evenly spaced design.Both uneven alternatives lift nearly every representation away from degeneracy.
- Uneven spacing reduces catastrophic split-half alignment from 51.6% to 19.4%, improving 61 of 93 representations (p < 10^-6).The median catastrophic-cost increase is +0.103, with 92 of 93 representations improved in cost (p < 10^-16).
- Jitter raises design-time cost for 92 of 93 representations (p < 10^-16) but leaves alignment unchanged (p = 0.51), with 16 fixed against 12 broken.The result motivates sampling away from the invariant submanifold rather than merely perturbing a regular design.
Appendix D. fMRI
An atlas-based fMRI analysis found significant cross-subject structure in hV4 under one condition, but several acquisition and localisation limitations prevent distinguishing absent structure from structure too noisy to align.
- The analysis includes all 35 recruited subjects after independent fMRIPrep-based preprocessing, without a separate motion-exclusion step.The original study excluded four subjects for head motion, so their influence on estimates cannot be ruled out.
- In hV4 under the no-report condition, the noise-ceiling lower bound is +0.127 (p = 0.004), with six analysis variants ranging from 0.104 to 0.127.All variants survive Benjamini–Hochberg correction at q = 0.05 across 17 tests.
- Interpretation is limited by atlas-based rather than individually localised ROIs, unavailable fieldmaps, 300 ms rapid event-related presentation, and no functional localiser.The noise ceiling cannot resolve whether absent shared structure reflects no structure or structure too noisy to align.
Appendix E. Four measurements that cannot separate the candidates
Four measurements on the nine-colour set fail to distinguish the HSV cylinder, CIELAB, and cone-opponent candidate geometries. Their high agreement partly reflects a rank-deficient HSV representation.
- Four independent measurements at S = V = 1 fail to distinguish the three candidate colour geometries.This motivates testing identifiability separately from representational-model discrimination.
- Pairwise canonical correlations range from 0.972 to 0.984 across the three geometry pairs.The HSV representation is rank-deficient because constant saturation and value collapse its cylinder to a one-dimensional circle.
- Human similarity judgements correlate with the hue-circle prediction at ρ = 0.800 against a noise ceiling of 0.890.Linear decoding from twelve vision models reaches R2 = 0.98–0.99 at best for all three coordinate systems.
Appendix F. Probe-set identifiability beyond perception
Probe-set identifiability depends on the chosen probes and candidate geometry: exact or approximate symmetries can leave correspondences ambiguous even with more data. The appendix distinguishes this ambiguity from parameterisation symmetries and notes that its empirical importance beyond controlled settings remains untested.
- Non-trivial automorphisms of a probe-induced dissimilarity matrix bound the correspondence that geometry-only alignment can uniquely recover.The bound applies whenever systems are compared using geometry induced by a shared probe set.
- Additional probes from the same symmetry-preserving support cannot identify features with identical response profiles over that support.Breaking the ambiguity requires probes that vary along directions capable of separating the features, rather than merely increasing sampling density on one manifold.
- Probe-induced ambiguity is conditional on the probe set and candidate geometry, unlike hidden-unit parameterisation symmetries, which exist independently of the probes.Parameterisation symmetries may instead be handled by quotienting out or accounting for equivalent model parameterisations.
- Whether restricted natural-image or language probes create practically important approximate ambiguities in cross-model feature matching remains untested.The colour and orientation experiments establish the principle only where the geometry, symmetry group, and ground-truth correspondence are known exactly.