Source-linked AI summary
Safe Testing
Peter Grünwald, Rianne de Heide, Wouter Koolen
TL;DR
The paper addresses how to conduct and combine hypothesis tests when decisions to continue testing depend on earlier outcomes. It develops e-value-based methods, including GRO constructions for composite hypotheses, and shows that they preserve Type-I error guarantees under optional continuation while supporting several statistical interpretations.
Problem
Standard p-values do not provide the same optional-continuation behavior as e-values when additional studies are selected based on previous outcomes.
Method
The paper develops e-variable constructions and GRO criteria for composite null and alternative hypotheses, including plug-in methods for optional continuation.
Results
E-values preserve Type-I error guarantees under optional continuation, and GRO constructions apply to broad testing problems with composite hypotheses.
Takeaways & Limitations
E-values provide a framework for combining evidence across adaptively continued studies while retaining interpretations based on gambling, p-values, and Bayesian Bayes factors.
Takeaways & Limitations
Multiplicative combination requires separate data batches, and GRO justification is strongest under independent study outcomes with moderate log-e-variable variances.
Abstract
from arXiv · showhide
We develop the theory of hypothesis testing based on the e-value, a notion of evidence that, unlike the p-value, allows for effortlessly combining results from several studies in the common scenario where the decision to perform a new study may depend on previous outcomes. Tests based on e-values are safe, i.e. they preserve Type-I error guarantees, under such optional continuation. We define growth-rate optimality (GRO) as an analogue of power in an optional continuation context, and we show how to construct GRO e-variables for general testing problems with composite null and alternative, emphasizing models with nuisance parameters. GRO e-values take the form of Bayes factors with special priors. We illustrate the theory using several classic examples including a one-sample safe t-test and the 2 x 2 contingency table. Sharing Fisherian, Neymanian and Jeffreys-Bayesian interpretations, e-values may provide a methodology acceptable to adherents of all three schools.
1 Introduction and Overview
The paper develops hypothesis testing with e-values, emphasizing their validity under optional continuation and their construction for broad composite testing problems. It presents growth-rate optimality and situates e-values among gambling, frequentist, and Bayesian interpretations.
- E-variables are nonnegative evidence measures proposed as an alternative to classical p-values.
- E-value threshold tests preserve Type-I error guarantees under optional continuation.The decision to conduct additional testing may depend on earlier outcomes.
- The paper develops general optimality criteria and constructions for e-variables with composite null and alternative hypotheses.The initial construction uses a prior on the alternative and yields GRO e-variables with frequentist Type-I error control.
- GRO e-variables have a Bayes-factor form with a special prior on the null.
- The paper examines optional continuation, stopping, competitiveness with classical methods, historical context, and connections among major statistical paradigms.
1. First Interpretation: Gambling
The gambling interpretation treats an e-variable as a ticket purchased for one dollar whose realized payout is E dollars. Under the null, the expected payout does not exceed the purchase price.
- An e-variable can be interpreted as a one-dollar ticket that pays E dollars after the data are realized.Buying r tickets yields expected final wealth at most r dollars under the null.
- Under the null hypothesis, purchasing such tickets is not expected to generate profit.
- The interpretation permits buying several tickets and positive fractional amounts.
2. Second Interpretation: Conservative p-Value, Type I Error Probability
E-values provide conservative p-value connections and retain Type-I error control when tests are combined through optional continuation. Their validity under multiplication depends on using separate data batches.
- For any e-variable E, the reciprocal 1/E is a conservative p-value.
- Threshold tests reject the null when E ≥ 1/α and preserve Type-I error probability α.
- Reciprocals of standard p-values are generally not e-variables, and e-values may require more extreme data for rejection on the same dataset.
- Products of e-variables from independent studies remain valid even when later studies are chosen based on earlier outcomes.
3. Third Interpretation: Bayes Factors
For composite nulls, e-variables can be constructed as Bayes factors with specially chosen null priors, including when nuisance parameters are present. Their conditional products remain valid under optional continuation, stopping, and changing priors, provided study batches are separate.
- Bayes factors: For a simple null, the Bayes factor is a sharp e-variable with expectation exactly 1.
- Bayes factors: For composite nulls, a suitable prior on H0 can pair with any prior on H1 to make the Bayes factor an e-variable.The resulting priors may be unusual, including degenerate priors.
- Bayes factors: GRO e-variables target rapid growth of evidence under H1 and extend to settings without a prior, including nuisance parameters.The paper uses maximin and relative maximin criteria for increasingly general cases.
- Optional continuation: Conditional e-variables have conditional expectation at most 1 under every null distribution, and their running product is a test super-martingale.This construction supports evidence accumulation across sequential studies.
- Optional continuation: Stopping a product of conditional e-variables preserves e-value validity and Type-I error guarantees, even when continuation or stopping depends on prior outcomes.Each e-variable may be constructed just in time from available information, and different later priors may depend arbitrarily on the past.
- Optional continuation: Multiplicative combination requires separate data batches; reusing data across studies is not allowed.E-values based on overlapping data can instead be combined by averaging, not multiplication.
2 The GRO e-Variable
The paper defines GRO e-variables for a given alternative by maximizing expected log-capital growth, with the optimal construction characterized as a Bayes factor involving a specially chosen null distribution.
- Theorem 1: Theorem 1 constructs an e-variable E∗ = q(Y)/p∗0(Y) under a full-support alternative Q when the minimum KL divergence to null Bayes marginals is finite.The denominator distribution P∗0 may be a sub-probability distribution.
- Theorem 1: The optimal expected log-growth equals the infimum KL divergence from Q to null Bayes marginals, D(Q∥P∗0).When the minimizing prior exists, P∗0 is the reverse information projection of Q onto the null marginal family.
- Optimal denominator: With a composite null, the optimal Bayes-factor denominator is essentially unique: any other null prior generally fails to produce an e-variable with the same numerator.If P∗0 = PW∗0, then W∗0 essentially uniquely minimizes the KL divergence.
- Simple alternatives: For a simple alternative, the θ1-GRO e-variable maximizes expected log-capital growth under Pθ1 among e-variables valid for H0.GRO measures evidence quality through expected capital growth rather than Neyman–Pearson power.
- Sequential growth: The logarithm is preferred for sequential growth because products can be ruined by an early zero, whereas expected log-growth is asymptotically optimal for independent samples.When expected log-growth is positive, the product grows exponentially, with maximal growth attained by maximizing expected log E.
- Example: In the illustrated normal-model case, the GRO construction uses a degenerate null prior and can require more data for rejection than the standard two-sided or one-sided 0.05 tests.The comparison is made using the threshold E ≥ 20, which gives a Type-I error guarantee of 0.05.
3 The GROW e-Variable
For composite alternatives, the paper defines GROW as worst-case expected log-growth and shows that the resulting e-variable is a Bayes factor based on priors selected through joint information projection.
- The GROW criterion: GROW selects, among e-variables valid for H0, the one maximizing the worst-case expected capital growth across all distributions in H1.This is a maximin criterion for composite alternatives when no prior on Θ1 is available.
- General construction: Under the theorem’s finite-divergence, full-support conditions, the GROW e-variable achieves worst-case growth D(PW∗1∥P∗0), essentially uniquely.The achieved value equals the supremum over valid e-variables of their minimum expected log-growth across Θ1.
- Information projection: The GROW e-variable is a Bayes factor between the components of the joint information projection, with the null component equal to a Bayes marginal when the relevant minimum is attained.This extends the simple-alternative construction to composite alternatives.
- One-parameter models: For separated one-parameter hypotheses with a monotone likelihood ratio, the GROW e-variable is the Bayes factor between null and alternative Bayes marginals minimizing KL divergence.The construction applies to Θ0 = {θ ≤ δ−} and Θ1 = {θ ≥ δ+}.
- Exponential families: In the one-parameter exponential-family example, the optimal priors W∗1 and W∗0 are degenerate at δ+ and δ−, respectively.The result relies on the sufficient statistic having a monotone likelihood ratio.
4 The REGROW e-variable: general composite H1 case
REGROW extends growth-rate optimization to composite alternatives by offsetting the unattainable oracle growth rate, including nested models and nuisance parameters. The resulting e-variables can be characterized through specially chosen Bayes factors, including right-Haar constructions for the t-test.
- REGROW defines an e-variable relative to an offset f, generalizing GROW beyond constant offsets and accommodating nested alternatives or nuisance parameters.The offset represents the growth benchmark used for relative worst-case optimization.
- 4.1 Composite H1, no effect size known: For alternatives with no pre-stated effect size, raw GROW collapses to the uninformative constant E*=1, motivating REGROW instead.This occurs when the alternative contains parameters arbitrarily close to the null in KL divergence.
- 4.1 Composite H1, no effect size known: In simple-null problems, the REGROW construction reduces to a redundancy-capacity problem with coding and Jeffreys-prior interpretations.For suitable parameter spaces, the redundancy penalty is (d/2) log n+O(1), linking the construction to MDL and objective Bayes methods.
- Discussion: The paper's examples typically eliminate nuisance parameters with REGROW before applying GROW to parameters of interest, while noting that this ordering is not always best.The discussion reports good practical performance but explicitly leaves open alternative REGROW choices for some settings.
- 4.3 Theorem 1 in Full: Application to Bayesian and Sequential t-test: For nuisance-parameter models such as the t-test, the theory is extended to coarsened data and convex sets of Bayes marginal distributions because the minimum KL divergence may not be attained.The coarsening framework allows the construction to model a scale-free marginal while handling nuisance variation.
- 4.3 Theorem 1 in Full: Application to Bayesian and Sequential t-test: Despite using the improper scale-invariant prior wH(σ) ∝ 1/σ, the resulting right-Haar Bayes factor is a valid e-variable and has a GROW property under compatible priors.The stated GROW result applies among e-variables for the full data under the paper's weak prior-moment condition.
5 (RE)GRO(W), Optional Continuation and Stopping
This section distinguishes optional continuation across studies from optional stopping within a data stream, and develops seqdec e-variable specifications that safely support both. It also examines coherence, practical constructions, and conditions for GRO behavior.
- Optional continuation: Optional continuation allows later study decisions, stopping times, and conditional e-variables to depend on previously observed studies.The study-level filtration records information available after completed studies and governs these choices.
- Applying e-specifications: The plug-in method constructs the m-th study’s e-variable by evaluating the specification function s[n_m] on that study’s data.For seqdec specifications, the resulting study-level products form a test martingale.
- Seqdec and stopping: Seqdec specifications decompose into sequential factors, so the same construction can support optional stopping on the concatenated data stream.The resulting factors are conditional e-variables and preserve Type-I error safety at the data level.
- Caveats: The t-test illustrates a filtration constraint: study-level safety can fail if stopping times use information finer than the filtration required by the construction.Such improperly specified stopping rules can create fake conditional e-variables with expectation above one under the null.
- Seqdec and stopping: Sequential application is coherent: applying a seqdec specification separately by study or globally to concatenated data gives the same result.This applies to the listed Bayesian W-GRO, GROW, and t-test specifications where the seqdec property holds.
- GRO constructions: For composite alternatives, batch size and model structure can determine whether a seqdec specification is optimally or almost optimally GRO.The 2 × 2 example motivates conditions under which REGROW-like behavior can be achieved.
6 Competitiveness: GRO and Power
The section compares GRO e-value tests with fixed-sample Neyman–Pearson tests in evidence and power planning. Optional stopping can reduce expected data requirements while increasing worst-case planning requirements.
- Evidence: GRO e-variables are designed to maximize expected log-evidence against composite null hypotheses while remaining valid under optional continuation.For simple nulls, they coincide with likelihood ratios and Bayes factors.
- Evidence planning: nGROW = ⌈L/D(Pδ∥P0)⌉ gives the smallest sample size expected to reach target growth L in the one-dimensional exponential-family setting.In the Gaussian location model, this becomes n = ⌈2L/δ2⌉.
- Power: For α = 0.05 and β = 0.2, the fixed-sample Neyman–Pearson benchmark has Cnp ≈6.180 in nnp = Cnp/δ2.This provides the reference scale for comparing e-value-based sample sizes.
- Power: A GROW test can require up to about twice the classical data amount, while its induced confidence interval has the same order of width.The authors characterize this as qualitatively closer to a standard Neyman–Pearson test than a standard Bayes factor approach.
- Power: A standard Bayesian prior avoids specifying δ in advance, but the resulting W1-GRO test needs a logarithmic factor more samples to reach power 1 −β.The evidential target-growth and maximal-power planning approaches can nevertheless be matched through a suitable choice of L.
- Optional stopping: Optional stopping requires less data on average for a desired power but more data in the worst case.The paper identifies this mismatch as a practical obstacle because individual studies must plan for the larger worst-case requirement.
7 Earlier and Related Work
The paper situates e-values and test martingales within earlier work on safe testing, sequential testing, confidence sequences, and p-value calibration. It presents the main novelty as the paper’s four versions of Theorem 1 rather than the safe-testing setting itself.
- Safe and anytime-valid inference: Test martingales underlie anytime-valid p-values, anytime-valid tests, and anytime-valid confidence sequences.The related literature traces these ideas through work by Shafer, Vovk, Robbins, Ramdas, and others.
- Safe and anytime-valid inference: Anytime-valid confidence sequences arise by collecting test martingales over null hypotheses and varying or inverting the null.The paper emphasizes batch-level analysis, while related anytime-valid work often focuses on individual data points.
- Novelty: The paper’s central novelty is the four versions of Theorem 1, while the safe or anytime-valid setting itself is not claimed as novel.The authors connect the theorems to standard, reverse, and joint information projections.
- Sequential testing: Sequential testing uses mathematically similar processes but is conceptually distinct from e-variable testing based on test martingales.The comparison concerns the roles of conditional e-variables and reciprocal processes under the null and alternative.
- P-value relations: Calibrators map p-values to e-variables through decreasing functions such as f(p) = 1/√p −1.The choice of calibrator is described as essentially arbitrary within the validity condition.
8 GRO: Discussion and Open Problems
The discussion qualifies when GRO is strongly justified and identifies open problems for high-variance and composite-alternative settings. It also notes a necessary compatibility condition for a proposed e-variable to be valid and GRO.
- GRO justification: In optional continuation, GRO is most strongly justified when study outcomes are independent and log e-variable variances are not too large.High variance can produce prolonged draw-downs before asymptotic growth becomes favorable.
- Validity conditions: A proposed ratio based on an information projection is an e-variable only when the denominator projection also minimizes KL divergence; then it must be GRO.This links validity to the specific projection used in the denominator.
- Open problems: For composite alternatives, GROW and REGROW may not always be preferable, especially when the goal is the fastest-shrinking always-valid confidence interval.The paper notes that REGROW e-variables do not generally achieve the usual O((log log n)/n) shrinkage rate.
9 Could Fisher, Jeffreys and Neyman Have Agreed on a Currency for Testing?
The paper argues that e-variable-based testing can connect Fisherian, Neyman-Pearson, and Bayesian perspectives while supporting safety under optional continuation.
- Synthesis: The paper presents e-variable testing as a possible common methodology for adherents of the three testing traditions amid continuing disputes over p-values and alternatives.
- Neyman-Pearson perspective: E-variable tests preserve Type-I error guarantees at fixed significance levels and extend this safety across non-pre-specified sequences of studies.Optional continuation replaces the usual single-study setting and motivates growth-rate optimality as an analogue of power.
- Fisherian perspective: Like p-values, e-values provide evidence against the null without requiring a specific alternative, although alternatives close to the data-generating process make them grow faster.
- Bayesian perspective: E-variables can also be represented as Bayes factors, with priors accommodating subjective knowledge or objective Jeffreys and right-Haar constructions.
A.1 Proof of Theorem 1, Simplest Version, and Corollary 2
The simplest proof constructs an e-variable from a reverse information projection and establishes its optimality and essential uniqueness through KL divergence and convexity arguments.
- E-variable property: Differentiating KL divergence along mixtures with null distributions shows that the candidate ratio satisfies the e-variable constraint.
- Construction: The proof identifies a minimizing measure P*0 for the KL divergence from Q over the convex null class and forms the candidate ratio q/p*0.
- Optimality: The candidate maximizes expected log evidence among null-valid e-variables, using the properness of the log scoring rule and the minimizing KL property.
- Uniqueness: Strict convexity of KL divergence and Jensen’s inequality establish that any e-variable attaining the same expected log evidence is essentially equal to the candidate.
A.2 Proof of full version of Theorem 1
The full theorem extends the reverse-information-projection argument to composite alternatives by combining a simple theorem with a minimax saddle-point result under regularity conditions.
- Reduction: The proof first establishes the e-variable property and essential uniqueness for the general construction by reducing relevant marginal distributions to the simplest theorem.
- Optimality: The saddle-point theorem converts the minimization and maximization relations into the desired growth-rate optimality inequality for the constructed e-variable.
- Minimax construction: A loss function based on KL divergence and the alternative-side adjustment f is introduced to apply a nonstandard minimax or saddle-point theorem.
- Regularity: The argument verifies well-definedness and lower semicontinuity using boundedness, finite KL divergence, weak convergence, and lower semicontinuity properties of KL divergence.
A.3 Remarks on and Checking of Conditions for Theorem 1
The appendix checks the theorem’s support, KL-finiteness, and minimizing-prior conditions, showing how these requirements apply in the paper’s examples.
- General conditions: Full support ensures the optimal e-variable is almost surely well-defined, while finite KL divergence ensures that expectations in the proof are well-defined.
- General conditions: For standard parametric models, excluding boundary points typically guarantees the required support and finite-divergence conditions; the 2×2 example uses Θ1=(0,1)2.
- Minimizing priors: The existence of a minimizing alternative prior is a strong additional condition, but the appendix verifies it for the composite examples and the safe t-test setting.
- Examples: In Examples 6 and 7, compactness, lower semicontinuity, and convexity establish existence of minimizing priors, while boundary-mass arguments preserve admissibility.
- Examples: The appendix proves that the minimizing alternative prior has full support and that the minimizing null prior assigns mass only to the interior parameter space.
- Examples: For fixed sample size, Carathéodory’s theorem bounds the number of support points needed to represent the relevant Bayes marginal distribution.
B Additional Clarifications and Proofs
The clarifications extend the sequential framework to conditional distributions and external side information, establish safety under optional continuation, and supply proofs and examples for GRO e-variables, t-tests, sample-size comparisons, and Brownian-motion approximations.
- Filtration extensions: The filtration extends to 2 × 2 data by modeling paired streams and conditioning Y on X.The construction uses conditional distributions pθ1(yi | xi) and augments the filtration with the observed covariates.
- Filtration extensions: External side information can guide whether to start another study and determine its sample size under a conditional-independence assumption.The resulting filtration may include observed data, side information, covariates, or a coarsening required by the testing setting.
- Filtration extensions: The construction ensures conditional expected e-values remain at most 1, preserving safety under optional continuation.This applies to both plug-in e-variable specifications and the sequential specification discussed for the sequential setting.
- Proofs of optimality: Monotone likelihood ratios imply stochastic dominance, which establishes the GRO property of the density-ratio e-variable pδ+(Y)/pδ−(Y).The argument uses the fact that the denominator corresponds to a prior concentrated on δ− and invokes the GRO characterization.
- Proofs of optimality: For the one-sample t-test, the relevant densities form a monotone likelihood-ratio family in the t-statistic, allowing Proposition 3 to apply.The noncentral t-distribution has ν := n − 1 degrees of freedom and noncentrality parameter µ = √nδ.