Source-linked AI summary
Game-theoretic statistics and safe anytime-valid inference
Aaditya Ramdas, Peter Grünwald, Vladimir Vovk, Glenn Shafer
TL;DR
Optional stopping and continuous monitoring can undermine conventional statistical practice, motivating evidence measures that remain valid at arbitrary stopping times. The paper develops SAVI through betting-based martingales and e-processes, surveys composite and nonparametric applications, and identifies scope limitations in several constructions.
Problem
Sampling and peeking decisions can produce misleading significance results, creating a need for inference valid under arbitrary stopping and monitoring.
Method
The paper develops safe anytime-valid inference using test martingales, e-processes, and game-theoretic betting, with constructions for composite and nonparametric settings.
Results
The paper reports advances including coarsened-filtration likelihood-ratio test martingales, improved symmetry bets, exchangeability results, and near-optimal sequential change detection.
Takeaways & Limitations
Game-theoretic methods provide a framework for extending likelihood-ratio intuition to composite and nonparametric testing and estimation problems.
Takeaways & Limitations
Some constructions are setting-dependent: reverse information projection may fail to produce an e-process, and admissibility questions remain open.
Abstract
from arXiv · showhide
Safe anytime-valid inference (SAVI) provides measures of statistical evidence and certainty -- e-processes for testing and confidence sequences for estimation -- that remain valid at all stopping times, accommodating continuous monitoring and analysis of accumulating data and optional stopping or continuation for any reason. These measures crucially rely on test martingales, which are nonnegative martingales starting at one. Since a test martingale is the wealth process of a player in a betting game, SAVI centrally employs game-theoretic intuition, language and mathematics. We summarize the SAVI goals and philosophy, and report recent advances in testing composite hypotheses and estimating functionals in nonparametric settings.
1. INTRODUCTION
Safe anytime-valid inference addresses irreproducibility from optional stopping by providing evidence measures that remain valid while data collection and analysis evolve. Its game-theoretic foundation uses nonnegative martingale wealth processes to support testing and estimation, including composite and nonparametric problems.
- Motivation: Sampling until significance appears can produce misleading findings, and such peeking remains prevalent despite longstanding statistical criticism.Examples include the power-posing study and a survey in which 56% of psychologists admitted deciding whether to collect more data after checking significance.
- Motivation: SAVI provides tools that accommodate continuous monitoring, optional stopping, continuation, and analysis of accumulating data without relying on prespecified stopping times.Its measures include power-one tests, confidence sequences, anytime-valid p-values, and e-processes.
- Game-theoretic foundation: Test martingales are wealth processes for Skeptic in a game where Forecaster assigns probabilities and Reality reveals successive observations.Skeptic’s cumulative betting score, or e-value, quantifies evidence against the Forecaster’s probabilities.
- Technical scope: Generalized game protocols address composite and extremely rich nulls when ordinary test martingales are unavailable or only trivial processes exist.They either restrict Skeptic’s information through an Intermediary or run parallel games whose net wealth is bounded across subsets of the null.
- Technical scope: Nonnegative martingales and e-processes act as nonparametric, composite generalizations of likelihood ratios, revealing hidden games behind many testing and estimation problems.The paper presents this as a source of new methodology and theoretical insight, while noting that a full understanding remains open.
2. CENTRAL CONCEPTS
The paper develops core SAVI concepts through e-values, test martingales, e-processes, and sequential tests, interpreting them as betting wealth processes. It explains how these tools extend likelihood-ratio ideas to composite settings and support anytime-valid evidence and confidence sequences.
- Composite hypotheses: Universal inference always yields an e-process, whereas reverse information projection yields e-values that may sometimes form an e-process directly.For alternatives without a common reference measure, the paper instead centers the design of nonnegative martingales or e-processes.
- E-values and martingales: An e-variable is a nonnegative random variable with expectation at most one under every distribution in the null, while its observed value is an e-value.When the expectation equals one, it is called a unit bet against the null.
- E-values and martingales: A nonnegative martingale starting at one is a test martingale, and optional stopping makes its stopped value an e-value.Game-theoretically, it is the wealth process of a gambler betting against the null.
- Composite hypotheses: E-processes generalize test martingale reasoning to composite nulls, where nontrivial test martingales may not exist because the null is too large.Composite test martingales can be viewed as simultaneous likelihood ratios, but e-processes are needed beyond that setting.
- Sequential testing: Sequential tests convert accumulating evidence into increasing binary rejection decisions, and Ville’s inequality supports rejection when a process reaches 1/α.The running maximum of an e-process is not itself an e-process, so additional calibration may be needed for nondecreasing evidence measures.
- Construction principles: The paper uses mixtures to combine e-values, provided the component e-values and weights are selected without looking at the data.This remains valid even when the component e-values are dependent, including when computed from the same data.
3. GENERAL PRINCIPLES AND METHODOLOGY
The methodology constructs e-processes for increasingly difficult composite testing problems, using growth-rate criteria, plug-in or mixture alternatives, and reverse information projection. It also identifies conditions under which these constructions are valid or optimal.
- Growth-rate optimality selects test martingales or e-processes that accumulate evidence quickly under plausible alternatives.The expected logarithmic growth criterion also controls evidence accumulation at stopping times.
- Risking all current wealth can make an e-process permanently zero, preventing it from recognizing later evidence against the null.This occurs when a betting strategy can go bankrupt with positive probability under an alternative.
- The likelihood ratio is growth-rate optimal for simple-vs.-simple testing, including at every stopping time.This result assumes the alternative is absolutely continuous with respect to the null.
- For composite alternatives, progressively learned plug-in distributions estimate the best-fitting alternative from data observed before each round.This extends the simple-vs.-simple likelihood-ratio strategy to alternatives containing unknown distributions.
- Reverse information projection yields an e-variable that is GRO relative to a stopping time, although it may be difficult to calculate and may fail to form an e-process.The construction is universal because it avoids regularity assumptions and asymptotics, while its validity depends on the chosen alternative construction.
- When learned plug-in or mixture alternatives are used, RIPr need not be an e-process, so composite alternatives may require mixing over e-processes instead.A fixed alternative can yield an e-process in settings where a learning mixture cannot.
4. PARAMETRIC EXAMPLES
Parametric examples reduce composite testing through likelihood ratios, invariance, conditioning, and mixture constructions. These approaches cover classical tests and produce e-processes with optimality or computational guarantees in selected settings.
- Likelihood ratios provide test martingales for simple alternatives and remain central to Bayesian and sequential testing constructions.Bayes factors for simple nulls are test martingales, while moment-generating-function constructions provide another class.
- Coarsening or conditioning on sufficient statistics can turn a composite parametric null into a simple null with a valid likelihood-ratio e-process.The resulting conditional distribution is common across the null family.
- Monotone likelihood ratios make the boundary alternative optimal for one-sided exponential-family tests and support plug-in or mixture approaches for composite alternatives.The same property can extend beyond exponential families, including fixed-degree-of-freedom noncentral t-distributions.
- Scale-free coarsening reduces the unknown-variance t-test to a simple-vs.-simple likelihood-ratio test.The resulting e-process is invariant to common rescaling and has been connected to GROW and REGROW optimality.
- Invariance methods also cover location, rotation, and Gaussian linear-regression problems with nuisance parameters.The affine-group construction extends the same reduction principle beyond scale invariance.
- For dependence and two-sample settings, e-variables can be built from conditional outcome models and flattened across multiple data streams.These constructions extend to conditional independence under model-X assumptions and sequential two-sample testing.
5. NONPARAMETRIC EXAMPLES
Nonparametric examples extend testing-by-betting to composite distributions, settings without ordinary martingales, and confidence sequences for functionals. The constructions include robust and adaptive mean-estimation procedures.
- The case studies construct test martingales for composite nonparametric nulls, supermartingales when martingales do not exist, and confidence sequences using reversed submartingales.They also include e-processes in settings lacking supermartingales and implementations in confseq.
- Mixture supermartingales for sub-Gaussian means grow exponentially under alternatives and grow faster as the alternative mean moves farther from the null mean.This makes the evidence adapt automatically to the difficulty of the testing problem.
- Mixture choices determine the iterated-logarithm behavior of confidence-sequence widths, trading a log t term for log log t through different behavior near the origin.Continuous mixtures around the origin yield the log t behavior, while unbounded mixtures at the origin can achieve log log t at the cost of constants.
- Inverting test supermartingales produces confidence sequences for a mean and, with time-varying conditional means, for the running mean.The same inversion framework translates betting evidence into anytime-valid estimation.
5.2 Heavy-Tailed, Robust Mean Estimation (Case B)
Heavy-tailed mean estimation relaxes sub-Gaussian assumptions through bounded-variance and moment-based constructions, while variance-adaptive methods exploit empirical variance for bounded data. Robust extensions address adversarial corruption.
- Conditional-variance-bounded models extend sub-Gaussian mean confidence sequences beyond Gaussian tails.The resulting confidence sequences appear visually almost identical to sub-Gaussian ones, with extensions when the p-th moment is finite for p > 1.
- Huber-robust test supermartingales and confidence sequences handle adversarial corruptions in heavy-tailed data.This extends the mean-estimation framework to contaminated observations.
- Sub-Gaussian or variance-bound parameters must be supplied in advance, because the data cannot generally identify the variance parameter itself.This restriction distinguishes these constructions from variance-adaptive methods for bounded variables.
- For bounded means, subexponential supermartingales use predictable estimates and an empirical variance term, then yield closed-form confidence sequences after gamma mixing.The predictable estimate can reduce the empirical variance contribution.
- Plug-in betting strategies adapt to the underlying distribution, particularly its mean and variance, and connect to Chernoff, empirical-likelihood, and dual-likelihood methods.These strategies are presented as statistically powerful for bounded mean estimation.
5.4 Testing Symmetry (Case A)
Testing symmetry illustrates that nontrivial martingales may be unavailable or inadmissible for rich nonparametric nulls, while universal inference and filtration choices can recover useful e-processes or martingales. The section also develops e-detectors for sequential change detection, with error control and, in some settings, near-optimal delay.
- Testing Symmetry (Case A): A test supermartingale exists for the symmetric-distribution class, using an empirical variance term and allowing mixtures or plug-in methods.
- Testing Symmetry (Case A): Ramdas et al. show that R_t is inadmissible because a test martingale R̃_t is always at least as large and typically larger.
- Testing Exchangeability and Log-Concavity (Case C): For binary exchangeability, no nontrivial test supermartingale exists in the original filtration, but universal inference yields a nontrivial and powerful e-process.
- Testing Exchangeability and Log-Concavity (Case C): Shrinking the filtration to conformal p-values enables nontrivial test martingales, although experiments find them less powerful than the binary exchangeability e-process.
- Testing Exchangeability and Log-Concavity (Case C): For log-concave distributions, no test supermartingale exists, whereas universal inference provides a powerful e-process; some nulls admit no nontrivial e-process at all.
- Sequential Change Detection: E-detectors sum e-processes begun at consecutive times, providing ARL at least 1/α and, in parametric settings, detection delay scaling like log(1/α)/D(P||Q).
- Sequential Change Detection: Backward confidence sequences complement e-detectors when pre- and post-change classes are unavailable or intersecting, with nonasymptotic ARL control and near-optimal parametric delay.
5.8 Time-Uniform Central Limit Theory and Asymptotic Confidence Sequences
This section develops time-uniform and asymptotic confidence sequences for sequential estimation, including doubly robust average-treatment-effect inference and nonparametric functionals. These methods extend familiar CLT- and offline-test-based procedures to settings with continuous monitoring, non-i.i.d. data, and sampling without replacement.
- Time-Uniform Central Limit Theory and Asymptotic Confidence Sequences: The same framework constructs doubly robust asymptotic confidence sequences for average treatment effects, producing anytime versions of corresponding confidence intervals.
- Time-Uniform Central Limit Theory and Asymptotic Confidence Sequences: Asymptotic confidence sequences approximate unknown nonasymptotic confidence sequences, with symmetric-difference error vanishing faster than a loglogt/t rate.
- Time-Uniform Central Limit Theory and Asymptotic Confidence Sequences: A universal asymptotic confidence sequence replaces σ with empirical variance when data have more than two moments, yielding a time-uniform analogue of the CLT.
- Time-Uniform Central Limit Theory and Asymptotic Confidence Sequences: Asymptotic confidence sequences extend to M-estimation and other semiparametric and nonparametric functional-estimation problems where CLT-based confidence intervals are standard offline.
- Time-Uniform Central Limit Theory and Asymptotic Confidence Sequences: Confidence sequences cover prespecified quantiles and entire univariate cumulative or quantile functions, extending the Dvoretzky-Kiefer-Wolfowitz inequality uniformly over time.
- Time-Uniform Central Limit Theory and Asymptotic Confidence Sequences: Plug-in test supermartingales for sampling without replacement can be inverted into confidence sequences, with later test martingales yielding the tightest known sequences in that setting.
6. MULTIPLE HYPOTHESIS TESTING
Game-theoretic methods extend anytime-valid inference to meta-analysis, false-discovery-rate control, and post-selection confidence intervals. The resulting procedures support continuous updating, adaptive data collection, arbitrary dependence in several settings, and decisions made after monitoring.
- Global Null Testing and Meta-Analysis: ALL-IN meta-analysis can update after every observation while retaining type-I error guarantees and informing whether studies should start, stop early, or expand.
- Global Null Testing and Meta-Analysis: Study-level e-values can be multiplied into a meta-level test martingale while studies are initiated, changed, or stopped adaptively.
- False Discovery Rate: The e-BH procedure controls FDR at α under arbitrary dependence among e-values, unlike the stated p-value analogue.
- False Discovery Rate: In bandit multiple testing, e-BH controls FDR at any data-dependent stopping time despite dependence induced by adaptive treatment assignment.
- False Discovery Rate: E-values can serve as unnormalized weights in weighted FDR methods, providing a stated power advantage over normalized weights.
- False Discovery Rate: The e-BY procedure controls FCR at α by reporting (1 − α|S|/K)-e-CIs for any dependence structure and data-dependent selection rule.
- False Discovery Rate: Stopped confidence sequences can be combined with e-BH and e-BY so adaptive monitoring, selection, and corrected interval reporting remain aligned.
- False Discovery Rate: E-values legitimize peeking but do not prevent other abuses: selecting the best of many betting strategies constitutes e-hacking.
7. OTHER APPLICATIONS
Game-theoretic statistics is rapidly evolving and has additional application areas beyond those discussed in detail. The section signals a broader research landscape rather than presenting a specific new application.
- OTHER APPLICATIONS: Game-theoretic statistics is described as rapidly evolving.
- OTHER APPLICATIONS: The paper points to additional topics where game-theoretic statistics is relevant.
- OTHER APPLICATIONS: This section serves as an entry point to applications beyond the preceding discussion.
Comparing/Evaluating Forecasters.
The paper asks how probabilistic forecasters can be evaluated for calibration and compared, using game-theoretic tests based on test supermartingales, e-processes, and confidence sequences. It also describes Jeffreys’s law: sufficiently different reliable forecasters cannot both remain undiscredited.
- Probabilistic forecasters can be evaluated for qualities such as calibration and compared with one another.
- Recent approaches address these questions using test supermartingales, e-processes, and confidence sequences.
- Jeffreys’s law states that two reliable forecasters must agree in the long run.
- If two forecasters differ too much, a Skeptic observing both can discredit at least one of them.
Multi-Armed Bandits and Reinforcement Learning.
In contextual multi-armed bandits and reinforcement learning, sequential decisions use contexts, actions, and observed rewards. The paper frames a central question about evaluating data collected under an exploratory policy before using it for decision making.
- Contextual bandits and reinforcement learning present contexts, actions, and observed rewards in sequential decision problems.
- A policy maps contexts to actions, while an exploratory policy is used to learn the unknown reward function.
- A central question is whether data collected under an exploratory policy can support subsequent decision-making analysis.
8. DISCUSSION
The discussion places SAVI methods alongside Bayesian, MDL, online-learning, and sequential-analysis ideas while identifying open questions about filtrations, admissibility, and universal inference. It emphasizes that filtration choices trade safety against power and that several characterizations remain unresolved.
- Connections to Bayesian and coding methods: Bayesian tools, MDL universal codes, and online sequential prediction strategies are closely connected to SAVI betting and prediction procedures.
- Choice of filtration: The choice of filtration affects both safety, through allowable stopping times, and power, through the rate at which wealth grows under alternatives.
- Choice of filtration: For exchangeability testing, conformal p-values enable a nontrivial test martingale only in a coarsened filtration, restricting optional stopping to times that cannot see the original data.
- Choice of filtration: For known discrete support, an e-process in the original filtration avoids that sacrifice and appears at least as powerful as the conformal test martingale in experiments.
- Admissibility: A necessary condition for admissibility is not sufficient: universal inference has the required infimum form but is known to be inadmissible in some examples.
APPENDIX A: SAVI AS A FREQUENTIST – EVIDENTIAL – BAYESIAN MIDDLE GROUND?
The appendix presents SAVI as a middle ground combining Bayesian prior information, evidential interpretation, and frequentist error control and coverage. It contrasts e-confidence intervals and Bayesian credible intervals, and extends e-posteriors toward minimax decision guarantees.
- E-processes quantify evidence against null hypotheses and remain meaningful even outside sequential testing, including batch multiple-testing settings.
- For simple nulls, admissible e-processes and Bayes factors coincide; for composite and nonparametric tests, they may differ substantially.
- An e-confidence interval includes parameter values whose e-posterior support exceeds α, whereas a Bayesian credible interval requires posterior probability 1−α on average over the prior.
- Because e-confidence intervals require pointwise support, Bayesian credible intervals are narrower in practice.
- The e-posterior can support minimax-optimal decisions for arbitrary loss functions, with guarantees independent of the chosen prior but weaker for atypical data.
- SAVI unifies prior information, evidentially meaningful numbers, and frequentist error control and coverage without disqualifying the Bayesian, evidential, or Neyman–Pearsonian paradigms.
APPENDIX B: DOES SAVI COME AT A PRICE?
SAVI does entail tradeoffs relative to classical methods, but whether these count as a “price” depends on the performance criterion. Its flexibility under optional stopping, robustness, and evidence-combination benefits can outweigh fixed-sample disadvantages in some settings.
- Compared with Bayesian credible intervals, anytime-valid confidence sets are wider, although not wider than the support interval.This comparison is presented as indicative rather than as a direct ranking of methods.
- At fixed sample sizes, uniformly most powerful Neyman–Pearson tests have more power than e-variables with GRO status for the same hypotheses.Anytime-valid confidence sequences also tend to be 1.5 to 2 times wider than standard confidence intervals.
- With optional stopping upon rejection, expected minimal stopping times under the alternative can match or be smaller than the fixed n required by Neyman–Pearson tests.This illustrates that the apparent cost depends on how performance quality is measured.
- SAVI imposes stronger type-I error control under optional stopping or continuation while accepting a slightly higher type-II error.The framework’s error definitions also differ from classical statistics because it optimizes expected evidence rather than simply minimizing error probabilities.
- SAVI optimizes evidence growth and reproducibility-related benefits rather than classical power alone, making direct “price” comparisons incomplete.The paper highlights robustness to dependence, evidence combination, and post-hoc loss functions as additional advantages.