Source-linked AI summary
Hypothesis testing with e-values
Aaditya Ramdas, Ruodu Wang
TL;DR
The book addresses the need for a unified treatment of e-values in hypothesis testing and develops methods and theory spanning their construction, optimization, merging, and use in multiple testing. It presents universal inference, log-optimality, and related results, including strong duality and asymptotic log-optimality, while focusing primarily on hypothesis testing rather than estimation.
Problem
The book addresses the need for a unified treatment of e-values, a recently named and under-studied tool fundamental to hypothesis testing and estimation.
Method
The book synthesizes foundational and advanced results while developing methods for constructing e-values, including universal inference, merging, admissibility, and strong duality.
Results
The book establishes that an empirically adaptive martingale merging function for iid e-values is asymptotically log-optimal.
Takeaways & Limitations
The book provides a coherent treatment of e-values intended for education, research, and practical use in statistics and related fields.
Takeaways & Limitations
The book focuses on hypothesis testing rather than estimation and only briefly treats the closely connected literature on confidence sequences.
Abstract
from arXiv · showhide
This book is written to offer a humble, but unified, treatment of e-values in hypothesis testing. It is organized into three parts: Fundamental Concepts, Core Ideas, and Advanced Topics. The first part includes four chapters that introduce the basic concepts. The second part includes five chapters of core ideas such as universal inference, log-optimality, e-processes, operations on e-values, and e-values in multiple testing. The third part contains seven chapters of advanced topics. The book collates important results from a variety of modern papers on e-values and related concepts, and also contains many results not published elsewhere. It offers a coherent and comprehensive picture on a fast-growing research area, and is ready to use as the basis of a graduate course in statistics and related fields.
Lists of notation, conventions, and examples
The book presents e-values as a unified foundation for inference, connecting them to tests and p-values while enabling post-hoc, sequential, multiple-testing, and log-optimal procedures.
- E-values serve as technical tools in multiple testing, dependent p-value merging, martingale methods, and related mathematical theories.Compound e-values underpin e-BH representations of FDR-controlling procedures, while e-values also arise in p-value merging and nonnegative martingales.
- E-values provide a unified foundation linking hypothesis tests, p-values, and broader statistical decision procedures.The book connects their existence and thresholding properties to tests and p-values across composite hypotheses.
- E-values support inference when testing problems, decisions, losses, or data collection depend on the observed data.The book highlights post-hoc decision making, data-dependent losses, optional continuation, and unknown sampling schemes.
- E-values preserve the expressive power of level-α testing because any such test can be represented by thresholding an appropriate e-variable at 1/α.The all-or-nothing construction recovers the original rejection decision, although general p-to-e calibrators need not preserve decisions.
- Log-optimal e-values maximize expected logarithmic evidence and exist for composite nulls, while avoiding zero values that destroy multiplicative evidence.The logarithm also governs the exponential growth rate of products of sequential e-values.
Validity: E-values under the null
Under the null, e-values support valid testing through expectation control, but their calibration and admissibility depend on the construction and testing problem.
- An e-to-p calibrator is essentially unique, whereas p-to-e calibrators form a much richer class.The canonical e-to-p map is supplied by Markov’s inequality, but converting in both directions can lose substantial evidence.
- E-values can be converted into conservative p-values by e 7→min(1, 1/e), while randomized calibration can yield strictly smaller p-values.The randomized transform e 7→(U/e) ∧1 dominates the non-randomized transform in all but degenerate cases.
- Admissible e-variables are characterized by nondominance, and for simple nulls every admissible e-variable can be written as a likelihood ratio.The admissibility characterization also yields exactness conditions and extends to broader composite settings.
- For finite product spaces, multiplying admissible e-variables produces an admissible e-variable, and the same conclusion extends to several one-parameter testing problems.The product result is stated for finite hypothesis classes and product spaces, with related extensions for ordered parameter families.
- In exponential families with monotone likelihood ratios, likelihood-ratio e-variables can be log-optimal against specified alternatives.The normal example supplies an e-variable for testing nonpositive means against positive alternatives, and the exponential-family construction is log-optimal for a fixed alternative.
- E-values remain valid under broad data conditions: a normalized sum tests a nonnegative mean without distributional or independence assumptions, while squared statistics require stronger variance control.For the sub-Gaussian class, Z^2 is valid under the restricted null P0 but not for the broader class because heavy left tails may make the variance unbounded.
Bibliographical note
This chapter develops a unified treatment of e-values, connecting powered tests, e-power, and log-optimality while collecting results on their construction and limitations.
- Bibliographical note: This chapter combines results from the e-value literature with additional derivations to form a coherent treatment.Its coverage includes testing equivalences, likelihood-ratio optimality, calibrators, and asymptotic e-power.
- 3.2 A unified existence result: The chapter establishes equivalences between powered tests, bounded powered e-variables, and, under condition (U), powered p-variables.For any α, [0,1/α]-valued e-variables correspond one-to-one with level-α tests.
- 3.2 A unified existence result: Condition (U) is necessary for constructing powered p-variables in general; finite sample spaces can permit powered e-variables without such p-variables.The finite-space obstruction arises because p-variables take finitely many positive values, preventing power at sufficiently small thresholds.
- 3.3 E-power and powered e-values: In the iid normal setting, likelihood-ratio e-values have e-power nµ^2/2, growing linearly with sample size, while their positive-e-power region depends on the alternative threshold.For Qθ=N(θ,1), positive e-power occurs when θ exceeds the relevant boundary, uniformly when ν>µ/2.
- 3.3 E-power and powered e-values: E-power is defined by EQ[log E], requiring a geometric mean above one and implying poweredness, whereas poweredness alone need not imply positive e-power.For simple alternatives, a small affine transformation can convert a powered e-variable into one with positive e-power.
- 3.3 E-power and powered e-values: The chapter also records counterexamples showing that poweredness, boundedness, positive e-power, and asymptotic behavior need not coincide.Examples include uniformly powered e-variables whose affine transforms have negative e-power and e-variables with infinite versus negative e-power.
- 3.5 Log-optimality and likelihood ratios: Log-optimal e-values maximize e-power, avoid zero almost surely, and are uniquely represented by likelihood ratios dQ/dP under the stated conditions.In the normal model, choosing θ=µ maximizes the e-power n(θµ−θ^2/2).
- 3.8 Axiomatic justification of the e-power: The chapter characterizes e-power axiomatically as EQ[log E] up to strictly increasing transformation, especially for product e-values with asymptotic consistency.The geometric mean exp EQ[log E] is an equivalent transformed measure.
Bibliographical note
The chapter largely synthesizes results and concepts from important e-value papers while also presenting results derived specifically for this book.
- Bibliographical note: The chapter combines results from important papers on e-values with results derived for the book’s coherent presentation.
Post-hoc testing and decision making with e-values
E-values support valid inference when error levels or decisions are chosen after seeing data, and they remain valid under optional continuation and stopping. The book also develops universal inference and subsampling methods for constructing and improving e-values in irregular and composite problems.
- Testing at data-dependent error levels: Every post-hoc valid family of tests is dominated by a family generated from one e-variable, with binary tests obtained by thresholding it at 1/α.The corresponding p-variable identifies the smallest level at which an admissible post-hoc binary family rejects.
- Post-hoc decisions using p-values and e-values: E-values provide risk-controlled decisions in non-binary or data-dependent decision problems, whereas p-value rules may fail to provide the desired guarantee.The e-value rule preserves risk control, including when the decision space is not binary.
- Optional continuation of experiments: E-values permit data-dependent continuation because multiplying conditional e-values preserves validity, allowing repeated monitoring and rejection when evidence exceeds 1/α.Ville’s inequality controls type-I error under continuous monitoring, while optional stopping preserves validity at a data-dependent stopping time.
- Universal inference: Universal inference constructs e-variables for composite nulls and alternatives without regularity conditions when the null maximum likelihood can be efficiently bounded.This provides tests for many irregular problems for which no other inferential methods are known.
- Likelihood-ratio methods: The subsampled likelihood-ratio e-variable has e-power at least as large as the split likelihood-ratio e-variable, while cross-fit and subsampling methods can yield smaller confidence regions than split LRT.In the reported power ordering, classical LRT is highest, followed by subsampling, cross-fit, and split LRT.
Bibliographical note
The chapter develops a duality theory linking log-optimal e-variables, numeraires, and reverse information projections, then relates these objects to universal inference, Bayes factors, and effective null hypotheses.
- The numeraire is log-optimal, exists without assumptions on P or Q, and is unique up to Q-nullsets.In the simple-null case, it equals the likelihood ratio dQ/dP.
- The RIPr P∗ lies in the bipolar P◦◦, which represents the effective null hypothesis associated with the e-variables for P.The inclusion may be strict relative to Conv(P).
- Strong duality identifies the numeraire's log-e-power with the minimum KL divergence from Q to the effective null, attained at P∗.Under Q ≪ P, the relation is EP[log E∗] = sup EQ[log E] = inf P∈P◦◦ KL(Q,P) = KL(Q,P∗), with quantities possibly infinite.
- The numeraire is the unique e-variable that can be represented as a likelihood ratio between Q and an equivalent member of P◦◦.This characterization provides a way to identify log-optimal e-variables in examples.
- For a point alternative, universal inference constructs EUI as the likelihood ratio of q(Z) to the maximum likelihood under the null.The numeraire is at least as large as EUI up to L-nullsets in the stated setting.
- Universal inference is always logically coherent, whereas numeraire families can be incoherent, creating a tradeoff between e-power and interpretability.For composite nulls, the numeraire generally dominates universal inference and may do so strictly.
- For nonparametric settings, the numeraire remains an e-variable and dominates Un; when the RIPr lies in the convex hull, it is the unique Bayes factor that is also an e-variable.When both hypotheses are simple, the main methods coincide; with composite hypotheses, they generally differ.
- For point alternatives, likelihood-ratio e-variables are numeraires, extending the monotone-likelihood-ratio example.The conclusion follows because the likelihood ratio is an e-variable and the RIPr characterization applies.
Bibliographical note
The chapter draws on foundational and recent work establishing reverse information projections, optimal e-variables, and related optimality criteria, while noting unresolved links and coherence issues.
- The chapter is largely shaped by Larsson et al. [2025a], while the RIPr traces back to Csiszár and Tusnády [1984] and foundational work by Li [1999].Grünwald et al. [2024a] later highlighted the RIPr's role in constructing optimal e-variables.
- Grünwald et al. [2024a] study GROW, which maximizes worst-case e-power, but the chapter describes this criterion as pessimistic because it may not adapt to simpler alternatives.Their work motivates the REGROW criterion.
- Related work derives numeraires for exchangeability, group invariance, and one-parameter exponential families using different techniques.The cited results span Koning, Larsson et al., Pérez-Ortiz et al., and Grünwald et al.
- The minimum-KL result is related to but distinct from the KLinf metric, and the exact relationship remains unresolved.Strong duality also extends to some divergences beyond KL, including an example based on Rényi divergence.
- Numeraire e-variables are generally logically incoherent, unlike universal inference, a distinction first observed by Bickel [2024].This difference affects how evidence behaves when the null is enlarged.
Sequential anytime-valid inference using e-processes
Sequential e-processes provide anytime-valid evidence for testing, allowing continuous monitoring and optional stopping while supporting both decisions and evidence assessment. The likelihood-ratio process connects this framework to Wald’s SPRT and is optimal in several simple testing settings.
- Wald’s sequential probability ratio test: The SPRT reduces expected sample sizes to 137.7 under the null and 139.3 under the alternative, about half the fixed-sample LRT requirement.The comparison concerns testing p = 0.5 against p = 0.6 with target type-I and type-II errors of 0.05.
- SAVI and Wald’s SPRT: Unlike Wald’s framework, SAVI methods prioritize continuously updated evidence and guarantees across all stopping times rather than one prespecified decision time.Binary decisions remain possible by thresholding the resulting evidence process.
- Optional stopping and anytime validity: Sequential e-processes support valid inference at any stopping time, addressing the type-I error failure caused by repeatedly peeking at ordinary p-values.Ville’s inequality turns an e-process threshold crossing into a valid sequential test.
- E-processes as a general framework: Every sequential test, including tests for composite and possibly nonparametric nulls, can be represented using e-processes.This makes e-process construction a general route to sequential testing rather than merely a way to analyze one stopping rule.
- Evidence interpretation: E-processes remain interpretable as evidence because larger values indicate more evidence against the null, even when no thresholded decision is made.They can be thresholded when a binary decision is needed, but thresholding is not required for interpretation.
- Log-optimality: For simple nulls, likelihood-ratio processes are log-optimal and achieve the largest asymptotic growth rate, equal to KL(Q, P) in the iid setting.The same likelihood-ratio martingale can also optimize Wald’s error–stopping-time tradeoff in the simple setting, though the criteria generally differ.
Bibliographical note
The book consolidates recent results on e-processes and e-value merging into a unified framework, including characterizations of admissible merging functions. Its results also expose trade-offs between optimal expected growth and undesirable distributional behavior.
- Bibliographical context: The book situates these results within a literature spanning modern definitions of e-processes, their equivalence, universal inference, Ville’s inequality, and testing by betting.The cited developments include work by Wald, Ville, Doob, Howard, Ramdas, Vovk, Wang, and others.
- Merging under arbitrary dependence: The arithmetic mean essentially dominates all symmetric e-merging functions, while every admissible e-merging function belongs to the class Mλ or is dominated by one of its members.This characterization covers arbitrary dependence and positively homogeneous symmetric merging functions.
- Independent e-values: Products and U-statistic functions yield valid admissible merging methods for independent e-variables, including the product, arithmetic average, and constant-one cases.Convex combinations of these U-statistic functions are also admissible independent e-merging functions.
- Sequential and martingale merging: For sequential e-values, convex mixtures of U-statistic functions are admissible, and validity requires sequentiality in only one of the K! possible orders.Martingale merging functions are characterized by admissibility, with sequential monotonicity replacing ordinary coordinatewise monotonicity.
- Adaptive merging: The empirically adaptive martingale merging function is a default choice without prior information and achieves asymptotic log-optimality for iid e-values.Its asymptotic growth rate matches that of the corresponding empirically adaptive e-process.
- Trade-offs of product merging: The product function maximizes expected value under powered alternatives and variance under the exact independent null, but its all-in strategy can converge to zero almost surely.Its large expected value therefore comes with a distributional risk that the book describes as highly undesirable.
Bibliographical note
The book develops e-BH and related procedures with finite-sample FDR guarantees under arbitrary dependence, while showing that broader FDR procedures can be represented or improved through compound, adaptive, and randomized e-values.
- FDR control for self-consistent e-testing procedures, including e-BH, holds at level α for arbitrary dependence among compound e-variables.
- The e-BH procedure controls FDR without dependence assumptions, unlike ordinary BH applied to reciprocal p-values, which requires the Benjamini–Yekutieli correction under arbitrary dependence.
- Every FDR-controlling procedure can be reproduced by e-BH applied to suitable compound e-values, with tight choices available for admissible procedures.
- Compound e-values admit simple separable constructions linked to Bayes estimators, optimal densities, and compound-decision theory.
- Closed and minimally adaptive constructions dominate the base e-BH procedure while retaining FDR control at level α.
- Randomized rounding can preserve FDR control while weakly expanding e-BH discoveries, and De-BH applies two successive rounding steps for further power improvement.
Bibliographical notes
The book situates its e-value results within prior work on FDR control, compound decision theory, randomization, and applications such as knockoffs and meta-analysis.
- The e-BH procedure and its theoretical guarantees build on prior work by Wang and Ramdas, including related boosting methods.
- Compound e-values trace to earlier work without that name, while related literature uses terms such as generalized e-values and relaxed e-values.
- The book connects compound e-values to Robbins’ compound decision theory and empirical Bayes methods for sequence models.
- The combination and derandomization recipe has been applied to model-X knockoffs and meta-analysis with limited reported information.
- The book identifies prior sources for minimally adaptive e-BH, stochastic rounding, intersection methods, and closed e-BH.
Advanced Topics
Advanced topics extend e-values to approximate and asymptotic validity, compound multiple testing, and procedures whose guarantees persist under relaxed conditions.
- Approximate p-values and e-values allow multiplicative error ε and probability-level error δ, with δ providing the more lenient form of approximation.
- The atomless assumption is needed for certain equivalences between approximation formulations, although external randomization makes it automatic.
- Controlling δ is stronger than controlling ε because positive ε can be converted to a strictly smaller additional probability error δ′.
- The chapter develops asymptotic e-variables and p-variables for growing datasets, including sequences that become strongly valid under specified constructions.
- Under alternatives with positive mean and finite variance, the constructed asymptotic e-values grow to infinity with probability one.
- Approximate per-hypothesis objects yield approximate compound e-values and asymptotic FDR control when used with e-BH.
Bibliographical note
The book uses optimal transport to characterize merging functions and studies how e-values combine with p-values, including quotient and Fisher rules for multiple testing and meta-analysis.
- Optimal transport duality supplies the main technical framework because merging under arbitrary dependence is an optimization problem over joint distributions with fixed marginals.
- The chapter proves a characterization of e-merging functions and shows that valid one-dimensional transformations are dominated by affine forms (1 − λ) + λE.
- The quotient combiner outputs a capped p-value quotient and typically offers greater power, motivating e-values as unnormalized weights in multiple testing.
- Fisher combination is preferable for exchangeable datasets, whereas quotient combination can outperform it when the p-value comes from a substantially more powerful primary dataset.
- The e-weighted BH procedure uses evidence-dependent, non-normalized weights and requires no dependence assumption within the e-values.
Bibliographical note
The bibliographical note develops relationships among p-variables, average p-variables, e-values, and p-merging functions, including calibration and admissibility results.
- Twice an average p-variable is a p-variable, and the factor 2 cannot be improved.Average p-variables form a convex, distributionally closed class, unlike p-variables; they are the convex hull of p-variables.
- For an average p-variable P, convex calibrators produce e-variables, while (2E)^-1 ∧1 converts any e-variable E into an average p-variable.
- Average p-values mediate e-to-p calibration: composing the factor-2 conversions yields the unique admissible calibration e 7→e^-1 ∧1.
- The naive procedure of merging p-values through e-values is generally inadmissible, although varying the underlying function recovers all admissible homogeneous p-merging functions.
- Admissible p-merging functions can be represented through admissible calibrators, and homogeneous symmetric cases permit identical calibrators and weights.
- Randomization and exchangeability improve p-value merging, while online rules avoid fixing the number of p-values but require a calibrator independent of K.
Bibliographical note
The note covers e-confidence intervals, e-based multiple testing, and majority-vote aggregation of dependent uncertainty sets, including randomized and exchangeable improvements.
- Every e-confidence interval is a confidence interval, and level-free e-CI families support e-BY false coverage-rate control under arbitrary dependence and selection.
- Majority vote combines arbitrarily dependent uncertainty sets with near-input coverage, but intersections can have inadequate coverage, as low as 1−Kα.
- Worst-case majority-vote error is αK/⌈K/2⌉, while practical coverage may approach 1−α or exceed it under favorable dependence.
- Exchangeable or randomized ordering yields sets no larger than majority vote while retaining a 1−2α coverage guarantee.
- The randomized set CU has coverage at least 1−α and is no smaller than CR, while CR has coverage at least 1−2α and is no larger than CM.
- For equal-width intervals, the median-of-midpoints construction contains majority vote, preserves 1−2α coverage, and has at most the input width.
Bibliographical note
The note surveys duality and representation results connecting e-values and p-values, plus power, calibration, comonotonicity, and distribution-informed threshold improvements.
- Under monotonicity, 1/S(X;P) is simultaneously the supremum of e-variables and the reciprocal of the infimum of p-variables.
- E-variables built from data are conditional expectations of likelihood ratios, with optimality for simple-null versus alternative testing.
- Powered exact p-values and e-values exist under equivalent structural conditions, including pivotal and bounded variants.
- Conditional e-to-p calibrators based on probability bounds can improve on the standard reciprocal calibration and support procedures such as Benjamini–Hochberg.
- For comonotonic e-variables, the supremum remains valid and can have higher power than a mixture of likelihood ratios.
- Distributional shape can raise e-value thresholds by nearly a factor of 2 for decreasing or unimodal densities, whereas standard Markov bounds cannot improve for several log-transformed classes.
Bibliographical note
The note introduces risk measures and their domains, then connects coherent risk constraints to e-values and backtesting procedures.
- Risk measures map random variables or distributions to real-valued risk assessments, with this chapter focusing on real-valued distributions and omitting broader variants.
- VaR and ES are central risk measures; ES is coherent, whereas VaR lacks subadditivity.
- Coherent risk measures satisfy monotonicity, cash invariance, subadditivity, and positive homogeneity.
- Constraints from ESβ characterize e-variables for null distributions whose likelihood ratios relative to P0 are bounded by 1/(1−β).
- With multiple observations, the associated testing problem can learn the true data-generating distribution, unlike the corresponding single-observation formulations.
- Backtest e-statistics turn risk comparisons into e-valid tests; for nonnegative variables, e(x,r)=x/r is valid when r bounds the mean and has expectation above one when the mean exceeds r.
Appendix
The appendix establishes general foundations for e-merging functions, rich and atomless statistical models, calibrators, and quantile-based constructions. It also records the conditions and limitations under which these abstractions support the book’s results.
- Statistical models: A statistical model is rich when a uniformly distributed [0,1] random variable exists under every probability measure, and any model can be made rich by adjoining uniform randomness.For a single probability measure, richness is equivalent to atomlessness.
- E-merging functions: An e-merging function transforms K e-values into one e-value through an increasing Borel function satisfying the required validity condition.The definition is formulated on a statistical model and can be reduced to a fixed atomless probability space.
- E-merging functions: For e-merging functions, validity on one rich statistical model is equivalent to validity across statistical models, using products with a uniform probability space.The equivalence relies on transferring distributions between rich models and then applying symmetry.
- Limitations: The richness assumption is essential: under a model of point masses, the maximum of e-variables remains an e-variable, but the maximum function is not a valid e-merging function.Thus, properties of e-merging functions cannot be extended indiscriminately to non-rich models.
- Quantiles: Quantile constructions provide measurable threshold events and translation rules, while left and right quantile functions agree almost everywhere.On atomless spaces, an event with any prescribed probability can be chosen between strict and weak upper level sets of a random variable.