Source-linked AI summary
When is Test-Time Adaptation Identifiable From Unlabeled Evidence?
Kartik Jhawar, Lipo Wang
TL;DR
The paper asks whether unlabeled evidence contains enough information to select the best TTA action, rather than assuming selector capacity is the limiting issue. It formalizes identifiability, analyzes a finite-batch Gaussian boundary, and tests the information limit on public benchmarks. The results show that deployment structures can reverse oracle actions under unchanged coarse evidence, while richer evidence helps only when it resolves the relevant ambiguity.
Problem
TTA methods can help or harm, creating a need to determine whether unlabeled evidence is sufficient to choose the best action.
Method
The paper formalizes oracle-action identifiability for fixed TTA actions and observation channels, then studies finite-batch theory and benchmark deployments.
Results
The paper finds that indistinguishable deployments can have different oracle actions; a Gaussian KEEP-versus-RECENTER boundary scales as 1/√n, and benchmarks show reversals under unchanged global evidence.
Takeaways & Limitations
TTA selection requires matching the observation channel to deployment variations that can change action rankings, not only increasing selector capacity.
Takeaways & Limitations
The theory uses a finite action menu and simple one-dimensional Gaussian model, while experiments use one main source-model family and a limited set of deployments.
Abstract
from arXiv · showhide
Test-time adaptation (TTA) offers many ways to update a deployed model without labels, but choosing the wrong update can make a strong source model worse. Recent methods therefore try to predict which adaptation will work from unlabeled test data. We ask a prior question: does the evidence given to the selector contain enough information to determine the best action at all? We show that this is not guaranteed, even with a perfect selector. If an observation channel makes two deployments look the same while their TTA rankings differ, reliable selection is impossible from that channel; richer evidence can restore the decision only when it resolves the relevant ambiguity. We make this boundary exact in a finite-batch Gaussian TTA model, where doing nothing beats mean recentering for small shifts, recentering wins beyond a unique critical shift, and the boundary shrinks as $1/\sqrt n$. Public benchmark studies on CIFAR-100-C and DomainNet-126 show the same failure mode with modern TTA methods: changing only deployment structure can reverse the oracle action while global order-blind evidence remains unchanged. The result is a practical way to separate two failure modes that are usually mixed together: a weak selector versus an information channel that cannot support the desired decision in the first place.
1 INTRODUCTION
The paper asks whether unlabeled evidence can determine the best fixed TTA action, showing that this depends on whether the observation channel preserves deployment properties that affect action rankings.
- Motivation: TTA selection can fail because identical evidence may correspond to deployments with opposite best actions, making reliable selection impossible from that channel.The issue is informational, not merely a weakness of the selector model.
- Problem formulation: The paper formalizes TTA selection as an information problem over a fixed action menu and specified unlabeled evidence channel.It asks whether the evidence is sufficient before asking how to build a stronger selector.
- Theory: A finite-batch TTA analysis gives a unique KEEP-versus-RECENTER transition, with the critical shift scaling as 1/√n.The analysis also constructs shift mechanisms with identical coarse evidence but opposite preferred actions.
- Empirical evidence: Benchmark evidence on CIFAR-100-C and DomainNet-126 shows that changing deployment structure can reverse the preferred TTA action while global order-blind evidence remains unchanged.Stream-aware evidence can capture these reversals, although that advantage depends on the deployment family.
- Implication: The practical distinction is between a weak selector and an information channel that cannot support the desired decision.When relevant deployment information is missing, richer evidence may matter more than greater selector capacity.
2 PROBLEM SETUP
The setup defines TTA actions, deployment worlds, prospective risks, and oracle actions, then separates information-level identifiability from learning a selector from finite data.
- Actions: The action menu contains KEEP, which uses the frozen source model, and data-dependent TTA procedures applied to an unlabeled adaptation batch.Benchmark actions include SOURCE/KEEP, DeYO, Tent, ROID, and test-time normalization; RECENTER is introduced as a tractable theoretical prototype.
- Deployment worlds: A deployment world specifies both the target distribution and the batch protocol, including ordering or temporal structure that can change stateful adaptation.The same image multiset can therefore produce different adapted predictors when images arrive differently.
- Risk and oracle: Prospective risk evaluates an action on an independent fresh target point, averaging over both the adaptation batch and evaluation point.Target labels define oracle risks only after experiments; selectors do not see them before acting.
- Observation channel: The evidence channel maps the source model and unlabeled batch to observed features such as entropy, logits, global moments, or order statistics.The induced evidence law determines which deployment worlds are indistinguishable to the selector.
- Identifiability: Oracle-action identifiability means that all worlds sharing the same allowed evidence have at least one common oracle-best action.This is an information question distinct from learning the decision from limited samples.
- Decision difficulty: The framework measures selector regret as excess risk over the oracle and asks when zero regret is possible or positive error is unavoidable.It separates knowing the winning action from knowing every risk or the full target mechanism.
3 WHEN IS THE ORACLE TTA ACTION IDENTIFIABLE?
TTA action identifiability depends on whether the allowed unlabeled evidence distinguishes deployments with different oracle actions. Exact impossibility results, compatibility conditions, and a finite-batch Gaussian model show how shift mechanism, batch size, and evidence choice determine selection.
- 3.1 IF TWO WORLDS LOOK THE SAME BUT NEED DIFFERENT ACTIONS, SELECTION IS IMPOSSIBLE: If two deployments have identical evidence laws but different unique oracle actions, every selector incurs at least 1/2 average action error and at least Δ/2 average regret when non-oracle actions cost Δ.This impossibility persists even with unlimited unlabeled data when the full allowed observation remains identical.
- 3.2 THE POSITIVE SIDE: WHEN THE ACTION IS IDENTIFIABLE: An oracle action is identifiable exactly when all worlds compatible with the observed evidence share at least one oracle-best action.A positive worst-case margin strengthens this condition by requiring the common action to beat every alternative by a positive amount.
- 3.3 THE PERMUTATION OBSTRUCTION: Order-blind evidence cannot identify the oracle when permuting a deployment leaves the evidence unchanged but changes the best action of a stateful TTA method.Matched benchmark experiments test this obstruction using the same images under different deployment orderings.
- 3.4 A TTA-SPECIFIC FINITE-BATCH KEEP-VERSUS-RECENTER BOUNDARY: KEEP is optimal near zero translation, whereas RECENTER wins beyond a unique critical shift δc(n), because KEEP worsens with shift while RECENTER retains an estimation-noise floor.The critical boundary scales as Θ(n^-1/2), reflecting the sample-mean standard error.
- 3.5 THE SAME TARGET MEAN CAN COME FROM TWO DIFFERENT SHIFTS: Mean-only evidence cannot identify the oracle under either translation or pure class-prior shifts when both worlds share the same target mean but favor opposite actions.Under pure prior shift, mean recentering moves the threshold in the wrong direction; richer statistics may separate the worlds.
- 3.6 IMPLICATIONS: The best TTA action depends on the shift, batch size, deployment structure, and information retained by the selector.The experiments are intended to test this information boundary rather than propose a universal TTA algorithm.
4 EXPERIMENTAL DESIGN
The experiments evaluate pre-action selection across controlled corruption and natural domain shifts using fixed TTA candidates and progressively richer unlabeled evidence. Matched deployments isolate whether order and stream structure provide information unavailable to global summaries.
- Experimental protocol: The study evaluates five fixed actions—SOURCE/KEEP, DeYO, Tent, ROID, and test-time normalization—from the same frozen source state, using labels only to define oracle outcomes.Selection evidence is label-free and computed before executing the chosen adaptation.
- 4.2 CIFAR-100-C: CONTROLLED DEPLOYMENTS: CIFAR-100-C contains 675 worlds spanning 15 corruptions, three severities, five deployment regimes, and three seeds, including matched IID and correlated worlds with identical images but different order.This isolates deployment structure while holding the selected image set fixed.
- Evidence hierarchy: The evidence hierarchy progresses from global uncertainty summaries (Z1) to permutation-invariant statistics (Z2), successive-window behavior (Z3), and ordered source-logit projections (Z4).Evaluation uses oracle regret, action accuracy, negative transfer, benefit capture, bootstrap intervals, and paired tests.
- 4.3 DOMAINNET-126: NATURAL DOMAIN SHIFTS: DomainNet-126 tests six directional transfers across 90 worlds, three seeds, and five deployment regimes, with 2,000 target images per world and execution batches of 32.Matched IID/correlated worlds again preserve the selected image multiset and class quota while changing only order.
5 RESULTS
Matched benchmark experiments show that deployment structure can reverse the oracle TTA action while global evidence remains unchanged, whereas stream-aware evidence can recover the relevant signal. This information advantage is conditional: it disappears under a different stochastic deployment family.
- Matched deployment experiments: 157 of 270 CIFAR-100-C matched pairs changed oracle action with identical selected images, including 82 strong flips; DomainNet-126 changed in 33 of 36 pairs, including 27 strong flips.Aligned SOURCE predictions and SOURCE error remained unchanged across the matched pairs.
- Interpretation: The benchmark results instantiate an information boundary: an observation channel can discard a deployment variable that changes the ranking of TTA actions.Selector quality therefore depends on whether its evidence retains the information needed to distinguish such deployments.
- CIFAR-100-C: Z3 changed action in 92.7% of CIFAR-100-C strong pairs with 0.093 pp mean pair regret, while global MORPHEUS selectors changed action in 0%.The global selectors had 6.189, 2.873, and 2.921 pp pair regret for entropy, NC-text, and NC-appendix variants, respectively.
- DomainNet-126: 0.034 pp regret and 93.3% exact action accuracy were achieved by the 2,990-dimensional DomainNet Z3 RF, while global Z2 reached 1.933 pp regret and 45.6% accuracy.At execution scale W = 32, the compact 11-dimensional ORDER11 Ridge model still achieved 0.061 pp regret and 92.2% accuracy.
- DomainNet-126: On 27 strong DomainNet flips, Z3 and ORDER11-RF changed action in all 27 with zero pair regret, while global Z2 changed none and had 1.942 pp mean pair regret.Ridge changed action in 26/27 pairs and had 0.044 pp pair regret.
- Deployment-family dependence: Under randomized stochastic CIFAR deployments, the controlled Z3 advantage disappeared: MORPHEUS NC-text had 0.210 pp regret versus 0.338 pp for Z3.The result shows that richer stream evidence is not automatically superior; its value depends on the deployment family.
6 RELATED WORK AND POSITIONING
Prior TTA work develops reliable adaptation, unlabeled performance estimation, and method selection, while this paper studies whether pre-action evidence can identify the best action. Its positioning draws on adjacent limits in TTA, domain adaptation, and decision theory without claiming those neighboring results as its own.
- TTA selection: Existing TTA methods address adaptation reliability, while AETTA, TTALine, and MORPHEUS estimate performance or select methods using unlabeled evidence.MORPHEUS is the closest operational neighbor because it uses pre-adaptation entropy and Neural-Collapse geometry.
- TTA selection: Cygert et al. evaluate unsupervised selection after candidate adaptations run, giving a richer information and computation budget than this paper’s pre-action setting.The paper treats this as an adjacent comparison rather than a same-budget method comparison.
- TTA limits: Recent TTA-limit studies examine recovery complexity and underspecified objectives, whereas this paper asks whether the best action is identifiable from pre-action evidence.These works are close in spirit but study different objects.
- Theoretical neighbors: Domain-adaptation and transfer studies provide impossibility, unlabeled-estimation, mapping, target-structure, and transfer-selection precedents, but not this paper’s specific TTA criterion or matched-stream result.The neighboring theories motivate the setting without directly supplying the finite-batch KEEP/RECENTER boundary.
- Decision theory: The paper uses established partial-identification and decision-theory machinery as a foundation rather than presenting Action-ID, the two-point lower bound, or the TV inequality as new general decision theory.Its contribution is concentrated on TTA-specific information limits and positive conditions.
7 DISCUSSION AND LIMITATIONS
The discussion separates weak selector models from insufficient evidence channels and shows that richer evidence helps only when it captures deployment variation that changes action rankings. The conclusions are bounded by the model, experiments, and coarse-evidence construction.
- Discussion: A selector can fail because its regression model is weak or because its evidence makes differently ranked deployments indistinguishable; only the former is addressed by more capacity.This distinction is the paper’s central practical diagnostic.
- Discussion: On CIFAR-100-C and DomainNet-126, changing stream organization alone can alter the required action, while global evidence may remain insufficient and small stream-aware summaries may recover the signal.The relevant evidence is determined by the deployment variation that changes action rankings.
- Discussion: Order-aware evidence is not universally better: its controlled advantage disappears under a stochastic deployment family, so observation channels must match the deployment variations under study.The paper’s conclusion is conditional rather than a blanket preference for richer or order-aware evidence.
- Limitations: The theory uses a finite action menu and one-dimensional Gaussian model, while experiments use one main source-model family and a simplified DomainNet action boundary.Broader architectures and action menus remain future tests.
- Limitations: The translation/prior-shift construction is deliberately coarse: full target distributions differ even though the specified population-mean channel is shared, and richer observations may resolve the ambiguity.Identifiability therefore belongs jointly to the deployment family and observation channel.
REPRODUCIBILITY STATEMENT
The experiments freeze target worlds, actions, seeds, evidence definitions, selectors, and statistical tests, with proofs in Appendix A and the experimental audit in Appendix B.
- Reproducibility: The study freezes target worlds, actions, random seeds, evidence definitions, selectors, and statistical tests, with proofs in Appendix A and an experimental audit in Appendix B.The submission also states that code and frozen result artifacts will be included for reproduction.
AI USE STATEMENT
The authors used generative AI tools during the research workflow for debugging, organization, mathematical checking, and language editing, while retaining responsibility for the work.
- Generative AI tools supported code debugging, organization, mathematical checking, and language editing.
- The authors remained responsible for theorem assumptions, proofs, experiments, numerical claims, citations, and manuscript statements.
A.1 PROOF OF THEOREM 1
The proofs establish that indistinguishable evidence can make reliable TTA selection impossible, while identifiable actions require a shared optimum across compatible worlds. In the Gaussian model, KEEP is optimal below a unique shift threshold and RECENTER above it.
- Identical evidence laws with opposite oracle actions force at least 1/2 average action error and at least Δ/2 expected regret.The discrete construction obtains Δ/2 = 0.15875.
- The discrete two-world construction keeps P(X) identical while changing target posteriors, reversing the oracle action from ADAPT to KEEP.World A prefers ADAPT, whereas World B prefers KEEP.
- When evidence laws differ, total variation bounds the remaining selection ambiguity and yields a regret lower bound proportional to (1 − TV)/2.
- An oracle action is identifiable exactly when all worlds compatible with the evidence share at least one optimal action.
- Invariant observation channels preserve indistinguishability under transformations, so differing optimal actions imply non-identifiability.
- In the Gaussian KEEP-versus-RECENTER model, KEEP has lower risk below a unique critical shift δc(n), while RECENTER wins above it.The proof establishes existence, uniqueness, and ordering of the transition.
A.7 PROOF OF COROLLARY 2
The same population mean can favor opposite actions because translation and class-proportion changes affect the mean differently. This ambiguity persists for finite-batch RECENTER for sufficiently large n.
- A balanced translation makes population RECENTER optimal because it restores source risk, while KEEP incurs the larger risk g(m).
- A class-proportion shift can make KEEP better than population RECENTER despite producing the same target mean.
- For sufficiently large n, the opposite-action conclusion persists for finite-batch RECENTER because the critical boundary tends to zero.
B.1 CIFAR-100-C: COMPLETE CONTROLLED STUDY
The controlled CIFAR-100-C study shows that deployment structure can reverse oracle actions for identical selected images, while stream-aware evidence substantially improves selection. The advantage is not universal across stochastic deployment families.
- Across the controlled 675-world study, Z3 achieved 0.312 pp mean oracle regret, versus 4.135 pp for entropy and 1.170 pp for the text Neural-Collapse baseline.
- 157/270 matched CIFAR-100-C pairs changed oracle action, including 82 strong flips, despite using exactly the same selected images.
- On 82 strong flips, Z3 changed its action in 92.7% of pairs with 0.093 pp mean pair regret, while MORPHEUS selectors changed action in 0%.
- The Z3 advantage disappears under randomized stochastic deployments, where MORPHEUS selectors outperform Z3 on mean regret.Mean regret was 0.210 pp for MORPHEUS NC-text versus 0.338 pp for Z3.
- DomainNet matched worlds changed oracle action in 33/36 pairs, with 27 strong flips, confirming the deployment-structure mechanism.
- On DomainNet, global Z2 achieved 1.933 pp regret and 45.6% accuracy, whereas Z3 achieved 0.034 pp regret and 93.3% accuracy.