Source-linked AI summary
Confidence Intervals for Policy Evaluation in Adaptive Experiments
Vitor Hadad, David A. Hirshberg, Ruohan Zhan, Stefan Wager, Susan Athey
TL;DR
Adaptive experiments can leave non-targeted treatment values difficult to estimate because undersampling and unstable assignment probabilities undermine conventional inference. The paper adaptively reweights augmented inverse-propensity estimates to obtain approximately normal, lower-variance inference, with favorable empirical RMSE and coverage comparisons. The approach supports frequentist confidence intervals under stated probability conditions, while requiring sufficient sampling of each arm.
Problem
Extreme undersampling or non-convergence of assignment probabilities makes it challenging to reuse adaptively collected data for estimating parameters not targeted by the experiment.
Method
The paper constructs averaging estimators with adaptive weights that control term contributions to variance and support asymptotic normality under severe adaptivity.
Results
The proposed methods compare favorably with existing alternatives in numerical experiments, while adaptively weighted estimators achieve asymptotically nominal coverage at a pre-specified horizon.
Takeaways & Limitations
The estimator enables approximately normal frequentist confidence intervals for policy values, including questions not targeted by the adaptive design.
Takeaways & Limitations
Consistent estimation requires infinite sampling; aggressive bandit procedures that sample some arms only finitely many times are excluded.
Abstract
from arXiv · showhide
Adaptive experiment designs can dramatically improve statistical efficiency in randomized trials, but they also complicate statistical inference. For example, it is now well known that the sample mean is biased in adaptive trials. Inferential challenges are exacerbated when our parameter of interest differs from the parameter the trial was designed to target, such as when we are interested in estimating the value of a sub-optimal treatment after running a trial to determine the optimal treatment using a stochastic bandit design. In this context, typical estimators that use inverse propensity weighting to eliminate sampling bias can be problematic: their distributions become skewed and heavy-tailed as the propensity scores decay to zero. In this paper, we present a class of estimators that overcome these issues. Our approach is to adaptively reweight the terms of an augmented inverse propensity weighting estimator to control the contribution of each term to the estimator's variance. This adaptive weighting scheme prevents estimates from becoming heavy-tailed, ensuring asymptotically correct coverage. It also reduces variance, allowing us to test hypotheses with greater power - especially hypotheses that were not targeted by the experimental design. We validate the accuracy of the resulting estimates and their confidence intervals in numerical experiments and show our methods compare favorably to existing alternatives in terms of RMSE and coverage.
1 Introduction
Adaptive designs improve efficiency for targeted objectives but can make post-experiment inference difficult, especially when assignments undersample some treatments. The paper develops approximately normal confidence intervals and adaptive estimators to address these challenges.
- Adaptive experiments can improve efficiency for objectives such as welfare maximization, best-arm identification, or hypothesis testing, but may sacrifice information about other questions.
- The paper proposes confidence intervals based on approximate normality even when assignment probabilities converge to zero or fail to converge, assuming known probabilities satisfying certain conditions.
- In a two-stage trial, adaptive sampling can bias the sample mean downward because initially low outcomes reduce an arm’s later sampling.The bias arises directly from adaptive data collection, even without selection effects.
- Inverse-probability weighting corrects this bias by up-weighting observations from rarely assigned arms, but decaying assignment probabilities create heavier-tailed, non-normal distributions.
- Naive estimators become reasonable only when adaptivity vanishes quickly enough for assignment probabilities to stabilize, a condition difficult to ensure under limited experimental budgets.
2 Policy Evaluation with Adaptively Collected Data
The paper develops policy-value estimators for adaptively collected data that remain asymptotically normal under severe propensity-score behavior. History-adapted evaluation weights stabilize variance and support confidence intervals, while specific allocation schemes address erratic or vanishing assignment probabilities.
- Setup: Treatment assignments are history-dependent propensity scores, and the target is the arm value Q(w) even when data collection did not target it.The framework assumes i.i.d. potential outcomes but allows observed outcomes to be dependent through adaptive assignments.
- Unbiased scoring rules: AIPW scoring rules preserve unbiasedness while potentially reducing variance through a control variate based on an outcome-mean estimator.When the outcome-mean estimator is zero, AIPW reduces to IPW.
- Asymptotically normal test statistics: Simple averages can be heavy-tailed and non-normal because inverse assignment probabilities may diverge or fail to converge, making conditional variance unstable.The resulting IPW distribution can mix tightly concentrated and highly dispersed behavior across adaptive trajectories.
- Asymptotically normal test statistics: Adaptively weighting unbiased scores stabilizes variance and yields consistent estimators with asymptotically normal studentized statistics.The weights may introduce small finite-sample bias when their random denominator is used, but they control tails and variance asymptotically.
- Constructing adaptive weights: Variance-stabilizing weights require allocation-rate conditions that limit how quickly propensities decay; these conditions still permit sublinear regret.The method cannot consistently estimate arms sampled only finitely many times.
- Constructing adaptive weights: The constant allocation scheme guarantees variance convergence and asymptotic normality, while the two-point scheme adapts between persistently high and rapidly decaying assignment probabilities.The two-point scheme averages these scenarios using the posterior probability that an arm is optimal under Thompson sampling.
3 Estimating Treatment Effects
The framework extends beyond single-arm values to treatment effects and more general causal targets. Treatment-effect inference can use either directly weighted difference scores or differences of separately estimated arm values, with the latter studied for higher power in the experiments.
- Treatment effects: Treatment effects are defined as differences in arm values, Δ(w1, w2) = E[Yt(w1) − Yt(w2)].The framework considers both direct scoring rules for the difference and differences of separately estimated arm values.
- Direct difference estimation: Directly weighting differences of AIPW scores yields asymptotically normal estimates of treatment effects.This approach is obtained by applying the adaptively weighted aggregation framework to unbiased difference scores.
- Difference of value estimates: Separately estimating each arm value and taking their difference also supports asymptotically normal inference under additional variance conditions.Theorem 4 gives joint asymptotic normality for the studentized arm statistics and consistency for the difference estimator.
- Comparison of approaches: The experiments use the difference of separately weighted arm estimates because separate evaluation weights provide more variance control and were found to have higher power.The paper notes that directly targeted difference estimators may still be useful in some applications.
- General targets: The results also cover general targets admitting doubly robust estimators whose Riesz representer depends on the treatment assignment mechanism.A dose-shift estimand in an adaptive clinical trial is given as an example.
4 Numerical Experiments
The experiments compare several estimators and confidence intervals for arm values and treatment-effect differences under three modified-Thompson-sampling settings. Adaptively weighted AIPW estimators generally provide approximately normal statistics and roughly correct coverage, while alternatives face variance, bias, or validity trade-offs.
- Experimental design: The study compares sample mean, AIPW, and constant- and two-point adaptively weighted AIPW estimators for arm values, treatment-effect differences, and their confidence intervals.AIPW-based intervals use plug-in reward means and a variance estimate; sample-mean intervals use a usual variance estimate, while time-uniform confidence sequences are also evaluated.
- Experimental design: Three settings use K = 3 arms with uniform[−1, 1] reward noise and no, low, or high gaps between arm values under modified Thompson sampling.Arm values are Q(w) = 1 in the no-signal case, Q(w) = 0.9 + 0.1w in the low-signal case, and Q(w) = 0.5 + 0.5w in the high-signal case.
- Estimator behavior: The unweighted AIPW estimator is unbiased but has poor RMSE and confidence-interval width, with far-from-normal studentized statistics in the no-signal case.In low- and high-signal settings, its variance is high because it does not account for undersampling of the bad arm.
- Estimator behavior: Adaptively weighted AIPW estimators perform relatively well, with roughly correct normal-interval coverage and approximately normal studentized statistics; two-point allocation gives smaller RMSE and tighter intervals.Even in the longest high-signal experiments, the bad arm receives only around 50 observations.
- Estimator behavior: Naive normal confidence intervals for the sample mean are invalid, showing severe under-coverage with little or no signal, whereas Howard et al.’s confidence sequences are conservative but often wide.The sample mean’s bias for undersampled arms can be non-monotonic in signal strength.
- Trade-offs: The simulations favor two-point adaptively weighted AIPW for asymptotically nominal fixed-horizon coverage, while sample-mean confidence sequences remain valid under arbitrary stopping but are often wider.The adaptively weighted estimator requires known propensity scores and sufficiently slow decay; confidence sequences require no assignment-process restrictions.
5 Related literature
Related work studies adaptive policy evaluation through optimal-policy learning, sequential randomization, debiasing, and stabilized weighting. The cited methods establish normality under varying assumptions, while the paper’s comparisons include their implications for arm-value estimation.
- Prior approaches: Much prior work focuses on learning or estimating the value of an optimal policy, including treatment allocation strategies designed to optimize clinical-trial criteria.The literature also addresses welfare maximization, best-arm identification, and hypothesis-testing power.
- Prior approaches: For sequentially randomized treatment, prior estimators reduce to augmented inverse propensity weighting or adaptively weighted estimators with constant allocation rates when specialized to arm-value estimation.The cited asymptotic-normality results rely on assumptions implying that a non-negligible treatment proportion remains assigned throughout the study.
- Debiasing: W-decorrelation provides consistent and asymptotically normal linear-regression estimates under strong serial correlation, but its arm-value estimates have high variance in the paper’s numerical evaluation.Its multi-armed-bandit construction uses arm indicators as covariates, along with arm counts, sample averages, and a tuning parameter.
6 Discussion
The paper develops adaptively weighted estimators for normal confidence intervals in adaptive experiments and reports favorable empirical performance, while identifying open questions about estimator optimality and broader sampling designs.
- Discussion: The proposed estimators produce normal confidence intervals with low variance and outperform existing alternatives in mean squared error and coverage.The approach is presented as a step toward policy learning and evaluation with adaptively collected data.
- Open alternatives: The construction is not the only route to normal confidence intervals, and a simpler estimator has essentially indistinguishable numerical performance in the single-arm setting.
- Limitations and open questions: The paper provides no optimality guarantees for confidence-interval width and leaves extension to contextual, non-stationary, or randomly stopped designs open.
- A.1 Behavior of arm value estimators over time: The adaptively weighted estimator estimates the good arm value Q(3) with negligible bias, small root mean-squared error, and roughly correct coverage for large T.
- A.3 Comparison to W-decorrelation: Both the two-point allocation method and W-decorrelation attain correct coverage, but W-decorrelation typically has much higher mean squared error.
- A.4 Two-point allocation over time: The two-point allocation scheme makes earlier observations receive more weight as λt decays over time.
A.5 Introduction example, revisited.
In the two-stage example, the adaptively weighted estimator is less biased than the sample mean and more thin-tailed than inverse-propensity weighting, yielding an approximately normal studentized distribution.
- Estimator distributions: The adaptively weighted estimator is more thin-tailed than inverse-propensity weighting and less biased than the sample average.
- Studentized statistics: The adaptively weighted estimator has a centered and approximately normal studentized distribution, unlike the sample mean and inverse-propensity weighted estimator.The latter two estimators do not have normal asymptotic distributions, although their deviations from normality are relatively small.
- Studentized statistics: Studentization alone is not sufficient to establish asymptotic normality because the evaluation weights also play a role.
B Limit theorems for adaptively weighted unbiased scores
The appendix proves more general versions of Theorems 2–4 for adaptively weighted unbiased scores.
- General limit theorems: The appendix develops general versions of the paper’s main limit theorems.
- General limit theorems: These results concern Theorems 2–4 from the body of the paper.
- General limit theorems: The appendix’s role is to provide broader theorem formulations beyond the main-body statements.
B.1 Setting
The general framework uses adaptively weighted unbiased scores for functionals of potential outcomes, with conditions ensuring consistency and asymptotic normality of estimates and differences.
- Setting: The setting has iid potential outcomes, observes only the outcome for the assigned treatment, and uses known history-dependent assignment probabilities.
- Scoring rules and weighting: Scoring rules transform observed outcomes into unbiased arm evaluations, which are averaged with data-adaptive weights to control variance before studentization.
- Functional setting: The framework defines L2(Pt−1) through conditional second moments and imposes continuity of the target functional with respect to that norm.
- Limit theorem: Theorem 5 establishes consistency and asymptotic normality for the studentized adaptively weighted estimator under boundedness, convergence, and weight conditions.
- Variance-stabilizing weights: Theorem 6 shows that recursively constructed variance-stabilizing weights satisfy the required sampling, variance-convergence, and Lyapunov conditions under specified allocation rates.
- Differences of functionals: Theorem 7 extends the result to differences of two linear functionals, giving jointly normal studentized statistics and a consistent estimator of their difference.
B.2 Specializing general results
The appendix specializes the general scoring-rule results to arm-specific values, treatment-effect differences, and average derivatives, establishing consistency and asymptotic normality under the corresponding representers and conditions.
- Arm-specific value: Arm-specific value estimation uses the conditional mean at treatment level w, with Riesz representer I{· = w}/e_t(w).Substituting this functional and representer into the general scoring rule yields the AIPW score used for the arm-value estimator.
- Arm-specific value: The adaptively weighted arm-value estimator is covered by Theorem 2, which establishes sufficient conditions for consistency and asymptotic normality.Theorem 2 is presented as a special case of the general Theorem 5.
- Treatment effects: Treatment-effect estimation can use the difference of two unbiased arm-value scores, with representer I{· = w1}/e_t(w1) − I{· = w2}/e_t(w2).Theorem 5 supplies conditions for consistency and asymptotic normality of this difference estimator.
- Treatment effects: An alternative treatment-effect estimator assigns each unbiased scoring rule its own sequence of evaluation weights, covered by Theorem 4 as a special case of Theorem 7.For the two arm values, the general conditions simplify to the counterparts stated in Theorem 4, with one condition always satisfied.
- Average derivative: The average derivative functional ψ(m) = ∫m′(w)f(w)dw is covered by Theorem 5 using representer γ_t(w) = −f′(w)/f_t(w), assuming f(w) vanishes at ±∞.The representer is verified through integration by parts.
C.1 Proof of Theorem 5
The proof of Theorem 5 establishes a martingale central limit theorem by controlling higher moments, conditional variances, and negligible remainder terms for the general adaptively weighted estimator.
- Proof strategy: The proof targets the Lyapunov and variance-convergence conditions required for the auxiliary martingale sequence.These conditions are used to obtain consistency and asymptotic normality of the studentized estimator.
- Lyapunov condition: The Lyapunov argument shows that the normalized contribution of large score terms converges to zero under the stated moment condition.Negligible-weight lemmas and inequalities such as Jensen’s and Hölder’s control the relevant higher-moment terms.
- Variance control: Lemma 12 characterizes the conditional variance of the unbiased score and bounds it above and below by positive constants.The lower bound follows from the assumed variance bound, while the upper bound uses uniform boundedness of the variance and outcome regressions.
- Variance convergence: The conditional-variance sum is decomposed into a dominant term and a remainder, with the remainder shown to be op(AT).The ratio is a weighted average of a bounded sequence converging almost surely to zero, with individually negligible weights.
- Variance convergence: The dominant variance term converges to its expectation, yielding ZT/EZT = 1 + op(1).This establishes the required variance convergence after the remainder terms are controlled.
C.2 General CLT for adaptively-weighted unbiased score differences
The general CLT proof establishes joint asymptotic normality for adaptively weighted score differences and then transfers it to feasible studentized statistics and linear combinations.
- Joint normality: The proof establishes asymptotic normality of every linear combination of the auxiliary martingales, then invokes the Cramér–Wold theorem for joint normality.The resulting vector converges to a jointly normal random variable Z ∼ N(0, I).
- Covariance control: Cross terms in the score covariance are shown to converge to zero in probability, allowing the covariance structure to simplify.The argument expands the covariance, cancels terms using iterated expectations and the Riesz representation, and bounds the remaining terms.
- Covariance control: Variance convergence and the Lyapunov condition make the remaining covariance terms negligible under the stated consistency and boundedness assumptions.The proof combines weighted-array arguments, Cauchy–Schwarz, and the variance-convergence condition.
- Studentization: The feasible studentized statistics converge to the same jointly normal limit as the population-normalized statistics.This follows because the ratio of empirical to population variances converges to one.
- Random weighting: Premultiplying by normalized random vectors preserves a standard normal limit when the vectors converge to a unit-norm constant.Slutsky’s theorem removes the vanishing terms, leaving a normally distributed projection with unit variance.
C.3 Proof of Theorem 6
Theorem 6 verifies that variance-stabilizing weights satisfy the general estimator conditions by maintaining controlled variance and higher moments under allocation-rate restrictions.
- Proof structure: The proof verifies infinite-sampling, variance-convergence, and Lyapunov conditions for the variance-stabilizing weights.These are the conditions required by the general theorem for asymptotic normality.
- Variance convergence: The variance-convergence numerator is constant under the variance-stabilizing recursion and the terminal condition λ_T = 1.The proof then establishes the required behavior of the denominator.
- Lyapunov condition: The Lyapunov numerator converges to zero when the allocation-rate exponent satisfies α < δ/(2+δ).The denominator is constant because the variance-stabilizing weights ensure E[∑T_t=1 ...] = 1.
- Lyapunov condition: The auxiliary function f_α is uniformly bounded for α < 1, which controls the final factor in the Lyapunov bound.Its boundedness follows from monotonicity and the limit as x approaches one.
- Infinite sampling: The infinite-sampling condition holds because the relevant allocation-weight sums diverge for α < 1 in both cases determined by T0.The proof treats separately T0 ≥ T/2 and T0 ≤ T/2, obtaining divergence in each case.
C.4 Asymptotic normality of alternative estimators
The proof establishes asymptotic normality for an alternative estimator by comparing its studentized statistic with that of an oracle adaptively weighted estimator. It also shows the associated variance estimates are asymptotically equivalent under the stated sampling conditions.
- Conclusion: Slutsky’s theorem transfers the central limit theorem from the adaptively weighted estimator to the alternative estimator.The transfer holds under the same conditions as the adaptively weighted estimator.
- Asymptotic comparison: The alternative estimator’s studentized statistic is shown to asymptotically behave like that of an adaptively weighted estimator using the oracle plugin ˆm_t(w) = Q(w).The oracle plugin is consistent under the conditions of Theorem 2.
- Asymptotic comparison: The numerators of the two studentized statistics are asymptotically equivalent under the infinite sampling assumption.The relevant remainder term is op(1).
- Variance equivalence: Only the first term in the variance expansion is asymptotically non-negligible; the second and third terms are op(1).The argument uses consistency of eQ_T, nonnegative weights, bounded absolute moments, Markov’s inequality, and the infinite sampling condition.
- Variance equivalence: The ratio of weights multiplying the leading variance term converges to one in probability, yielding ˜V_T = bV_h-avg_T(1 + op(1)).This establishes asymptotic equivalence between the alternative variance estimate and the leading variance term.