Source-linked AI summary
Time series experiments and causal estimands: exact randomization tests and trading
Iavor Bojinov, Neil Shephard
TL;DR
The paper addresses causal inference for temporal experiments when only one outcome path is observed and conventional population estimands are difficult to estimate without strong assumptions. It develops potential-outcome-path estimands and exact randomization inference under non-anticipating treatment assignment, then applies them to simulations and trading experiments. In the trading application, one trading method has a lower slippage rate than the alternative.
Problem
For time series experiments, population-averaged estimands are poorly suited to personalized settings and are difficult to estimate without strong, often unrealistic assumptions.
Method
The paper defines causal estimands for single-unit treatment paths and derives randomization-based inference, including an exact test, under non-anticipating assignment.
Results
One trading method has a lower slippage rate than the alternative in the empirical application.
Takeaways & Limitations
The framework supports causal statements about temporal treatments using the treatment randomization, including in the trading experiments analyzed.
Takeaways & Limitations
Because only one outcome path is observed, estimating the causal effects requires strong assumptions; increasing q also increases the estimator’s variance.
Abstract
from arXiv · showhide
We define causal estimands for experiments on single time series, extending the potential outcome framework to dealing with temporal data. Our approach allows the estimation of some of these estimands and exact randomization based p-values for testing causal effects, without imposing stringent assumptions. We test our methodology on simulated "potential autoregressions,"which have a causal interpretation. Our methodology is partially inspired by data from a large number of experiments carried out by a financial company who compared the impact of two different ways of trading equity futures contracts. We use our methodology to make causal statements about their trading methods.
1 Introduction
The paper develops causal estimands and exact randomization inference for experiments that apply treatment paths over time to a single unit, addressing limits of conventional population-averaged estimands. It applies the framework to simulated potential autoregressions and trading experiments comparing human traders with computer algorithms.
- Motivation: Time series experiments follow one unit over time, where conventional population estimands are difficult to estimate without strong, often unrealistic assumptions.The paper emphasizes that the length of the experiment can exceed the number of available units.
- Contribution: The framework generalizes one-period experimental treatments on multiple units to multiple-period treatment paths on a single unit.It represents causal effects through potential outcome paths and also extends the results to multiple units.
- Method: Under a relatively weak non-anticipating treatment-assignment assumption, the paper estimates several causal effects and derives two non-parametric randomization-based inference strategies.One strategy provides an exact randomization test of no causality for time series.
- Empirical application: The methodology is motivated in part by experiments that randomly assigned trading jobs to human traders or computer algorithms.The data cover one year of experiments across 10 equity-index futures markets and target causal effects on relative trading costs.
- Evaluation and application: The paper evaluates its procedures with simulated potential autoregressions and uses the trading database to make causal statements about alternative trading methods.The empirical illustration measures the causal effect of trading methods.
2 The treatment path and potential outcome paths
The paper represents treatment and outcome histories as paths, so each possible treatment sequence has corresponding potential outcomes while only one path is observed. It imposes non-anticipation so treatment assignment cannot depend on future potential outcomes, enabling causal analysis of temporal treatments.
- Treatment paths: At each time step, a single unit receives binary treatment 0 or 1, creating a treatment path whose potential outcomes depend on current and past treatments.The framework can be generalized to multiple treatments.
- Potential outcome paths: A length-t treatment history generates 2^t potential outcomes at time t and 2(2^t − 1) potential outcomes across all times through t.For T = 3, the framework contains 14 potential outcomes across eight potential paths.
- Potential outcome paths: The potential outcome path for treatment sequence w1:t is Y1:t(w1:t) = {Y1(w1), Y2(w1:2), ..., Yt(w1:t)}.The notation collects the outcome at each time under the corresponding treatment history.
- Observed path: An experiment observes only the potential path associated with the administered treatment sequence, leaving the other paths missing.For the example treatment path (1, 1, 0), the observed outcomes are Y1(1), Y2(1, 1), and Y3(1, 1, 0).
- Extensions: The setup links the observed outcome path to the administered treatment path under a single-version-of-treatment assumption and can condition assignment on observed covariates through an information set.The framework also extends to multiple units and allows treatment or outcomes in one series to affect future treatment probabilities in another.
- Treatment assignment: Non-anticipation requires treatment assignment at time t to depend on past treatments and past potential outcomes, not future potential outcomes.This assumption is also described as the condition that future potential outcomes do not Granger-cause the current treatment.
3 Causal effects
The paper defines time-series causal effects by comparing potential outcomes at fixed times and emphasizes temporal averages over treatment paths. It develops estimands that can be estimated from one experimental unit by conditioning on observed treatment history, while using stepping to reduce that dependence.
- 3.1 General causal effects: Time-series causal effects compare potential outcomes at a fixed time, with the temporal average treatment effect as the primary object of interest.The formulation avoids super-population or model-based averaging; randomization of treatment is the only randomness invoked.
- 3.1 General causal effects: The framework defines a broad class of causal estimands by comparing potential outcomes under different treatment paths and averaging them with non-stochastic weights.Weights can emphasize probable paths or assign zero weight to impossible paths, incorporating all relevant potential outcomes in the average.
- 3.2 p ≥0 lag causal effect of treatment on outcome: With one observed outcome path, general path comparisons are not estimable without strong assumptions, so the paper defines estimands conditional on observed past treatment.This yields a class estimable from one experimental unit, at the cost of making the estimand depend on the observed treatment path.
- 3.2 p ≥0 lag causal effect of treatment on outcome: The p-lag causal effect compares treatment and control at time t−p, averaging over treatment assignments at later times; p = 0 gives the contemporaneous effect.Uniform weights are used for most theoretical results, and the temporal average p-lag effect averages these time-specific effects.
- 3.3 Causal effects and treatment path: Conditioning on observed treatment paths supports estimability without further assumptions, while the resulting effects still satisfy a central limit theorem for inference.The authors relate this conditional perspective to the average effect on the treated.
- 3.3 Causal effects and treatment path: The q-step p-lag effect averages over treatment paths preceding the focal lag, reducing dependence on the observed treatment path as q increases.At q = t−p−1 it recovers the general path-based effect, while q = 0 leaves no such averaging.
4 Experiments and estimation
The paper develops Horvitz–Thompson estimators for temporal causal effects under randomized treatment paths, establishing unbiasedness and variance results without further assumptions for some estimands. Proxy outcomes can improve precision when they predict future potential outcomes, including under non-stationarity.
- Estimators and variance: The proposed estimators are unbiased over the randomization distribution, with variances depending on potential outcomes and unbiased estimators available for variance upper bounds.The upper bound is distinct from the usual Neyman-style bound, and its connection to that bound is not established.
- Randomization framework: The framework uses adapted propensity scores and conditions moments on the treatment randomization while holding all potential outcomes fixed.Probabilistic assignment requires path probabilities to lie strictly between zero and one, enabling randomization-based moment calculations.
- Estimators and variance: Conditional on all potential outcomes, estimation errors are martingale differences, making the estimators conditionally unbiased and uncorrelated through time.Theorem 1 gives conditional mean-zero errors and their conditional variance; the martingale property also yields unconditional zero covariance across distinct times.
- Average p-lag effects: The average p-lag causal effect averages time-specific effects, but its estimand depends on the observed treatment path, so different paths can produce different estimates.The paper derives unbiasedness and variance properties for this average estimator from the corresponding time-specific results.
- Proxy outcomes: Proxy outcomes can reduce estimator variance, and a good predictor of future potential outcomes can make the proxy-based estimator much more efficient, especially for non-stationary outcomes.The proxy is a function of the available past information and is used in a decomposition of the causal effect estimator.
5 Experiments and randomization inference
The paper proposes non-parametric inference based on randomized treatment paths, combining exact tests for sharp nulls with asymptotic tests and confidence intervals for average temporal effects. The conservative average-effect test is reported to have high power in practice.
- Asymptotic inference: The martingale-array central limit theorem yields asymptotically Gaussian scaled estimation errors when p is finite, regularity conditions hold, and T tends to infinity.The resulting asymptotic framework supports hypothesis tests and confidence intervals beyond the sharp-null setting.
- Exact randomization tests: The sharp null asserts no temporal causal effect for every treatment path, whereas the lag-p null asserts zero p-lag effects at every time point.The latter implies a zero average p-lag effect and can be tested against a portmanteau alternative.
- Exact randomization tests: Exact tests simulate treatment paths from their conditional randomization distribution, which determines the exact distribution of causal estimands under the sharp null.Under the sharp null, all potential outcomes are known from the observed outcomes, allowing causal statistics to be recomputed for each treatment path.
- Asymptotic inference: The conservative test for no average temporal causal effects uses an upper variance bound and is reported to have high power in practice.The unobserved-potential-outcome variance term is replaced by the bound derived earlier.
6 Connection to other work
The paper situates its model-free temporal causal effects within longitudinal, impulse-response, potential-outcome, and macroeconomic literatures. Its effects differ from model-based impulse responses and from related potential-outcome constructions in how treatment paths and expectations are defined.
- Longitudinal causal inference: The paper’s p-lag causal effects generalize several longitudinal-study estimands, including the blip, contemporaneous, and lagged effects under particular weighting choices.The related estimands are special cases obtained when weights equal reciprocal adapted propensities or when the lag is specialized.
- Longitudinal causal inference: Unlike marginal structural models that restrict the potential-outcome structure, the paper’s approach is framed from a finite-population perspective using randomization for inference.The paper contrasts its model-free treatment with approaches that impose functional restrictions to reduce the number of potential outcomes.
- Impulse responses: The paper’s p-lag effects connect to Sims-style impulse responses, but impulse responses are conditional expectations defined relative to a time-series model.In the potential autoregression example, the impulse response has the form IRF_t,s = φ^sσκ.
- Potential-outcome approaches: The paper distinguishes its potential-outcome effects from Angrist and Kuersteiner’s effects because their construction holds future innovations fixed without guaranteeing identical future treatment paths.Thus the two approaches are related in spirit but define different causal effects.
- Other connections: The paper notes that its causal effects have no direct connection to Granger causality, although Granger causality appears in one assumption.This separates predictive notions of causality from the paper’s intervention-based estimands.
7 Multiple units
The framework extends from one time series to multiple units with staggered observation times and potentially dependent series. It combines unit-level information for more accurate population-average effects while retaining randomization-based inference under stated assignment conditions.
- Setup and assumptions: The setup assumes temporal SUTVA, so each unit’s potential outcome depends only on that unit’s treatment path.Treatment assignment can still depend on other series’ treatment or outcome histories, allowing cross-series dependence.
- Setup and assumptions: The multiple-unit extension combines information across units to obtain more accurate treatment-effect estimates without adding assumptions.Units may begin and end experiments at different times, and the framework accommodates these staggered observation schedules.
- Population effects: The population p-lag effect is a weighted average of unit-level average p-lag effects, interpretable as an effect averaged across time and units.The paper focuses on equal unit weights, with stepped effects defined analogously.
- Multiple-unit inference: Randomization-based tests extend by sampling a new treatment path for each unit and computing a pooled statistic, while unit-level estimators remain conditionally unbiased regardless of cross-unit dependence.The sharp-null and conservative average-effect procedures are adapted from the single-unit case.
- Multiple-unit inference: When treatment paths are independent across units, unit-level p-values are independent and can be combined with Fisher’s method, though unequal estimator variances can reduce its power.The paper notes alternative methods for settings where unit-level variances differ.
8 Simulation study
The simulation study evaluates estimator sampling behavior, conservative-test calibration, power, and pooled estimation under potential autoregressions. Results indicate approximate normality and near-exact testing performance under the studied settings, with heavy-tailed noise requiring longer series.
- 8.1 Study design: For a univariate impulse potential autoregression with T = 100, μ1 = 0.5, μ0 = 0, φ = 0.5, and σ = 1, the running estimator tracks the true estimand near 0.8.The figure displays point estimates over time, a running average, and a 95% confidence interval against the true value.
- 8.2.1 Fixing the potential outcomes: Different treatment paths produce estimator distributions that quickly approach normality, whereas heavy-tailed noise requires longer experiments for approximate normality.The study fixes potential outcomes when examining variation over treatment paths; Cauchy noise is used as a heavy-tailed case.
- 8.2.2 Replicating over potential outcomes: The conservative test has slightly lower α level because it overestimates the variance, while its power is only slightly below the randomization-based test.The comparison studies power as treatment effects and the φ parameter vary.
- 8.3 Pooled estimation: averaging over multiple units: Pooling two independent experiments produces an estimator whose distribution is well approximated by the CLT and whose variance is lower than either individual experiment's variance.The pooled experiments use T1 = T2 = 100.
9 Empirical example from finance
The finance application uses randomized assignments between two trading methods to evaluate slippage across 10 equity futures markets. Exact randomization-based analyses find method A generally performs better contemporaneously, with little evidence of lagged effects.
- Trading futures contracts: The experiments assign human and algorithmic traders randomly, with i.i.d. Bernoulli treatment assignments that ignore past data.This setup satisfies the non-anticipating treatments assumption.
- Trading futures contracts: Method A typically trades more often and takes roughly three times longer than method B to fill an order.The frequency difference varies across markets.
- Definition of financial slippage: Slippage is measured using signed VWAP minus mid-price, scaled by mid-price, and lower values represent lower trading costs.The observed slippage series supplies the primary outcome for the analysis.
- Pooled estimation and lagged estimation: Only 1 of 20 lagged statistics was statistically significant, indicating little evidence of lagged causal dependence.The lone significant result occurred at lag 2 for market 7.
- Inference on b̄τ_p: The randomization-based test uses the experiment's treatment assignments rather than a model of financial-return dynamics.This avoids modeling thick-tailed returns and long-range volatility clustering.
- Pooled estimation and lagged estimation: Across 10 markets, method A performed better than method B in contemporaneous slippage, with a pooled p-value close to zero.The pooled contemporaneous result is strongly in favor of method A in a causal sense.
10 Conclusion
The paper develops a potential-outcome framework for causal estimands in single time series and extends it to multiple units. Its exact randomization tests applied to trading experiments identify lower slippage for one method, while the framework relies on probabilistic, non-anticipating treatment assignment.
- Conclusion: The framework defines a broad class of single-time-series causal estimands that can be estimated without assumptions on potential outcomes.Estimators are unbiased over the randomization distribution.
- Conclusion: The inferential framework requires treatment assignments to be non-anticipating and probabilistic.These conditions support the randomization-based inference.
- Conclusion: The framework provides three strategies for generalizing the analysis from one time series to multiple units.The paper applies the methods to a database of experiments from a quantitative hedge fund.
- Conclusion: Exact randomization tests can be conducted for the proposed causal estimands, alongside a CLT-based procedure for average temporal effects.The second procedure estimates an upper variance bound.
- Conclusion: Applied to trading experiments, the methods show that one trading method has a lower slippage rate than the alternative.The conclusion reports the direction of the comparison without identifying the methods here.
A.1 Proof of Theorem 1
The appendix develops randomization-based estimators and inference for temporal causal effects, using adapted propensity scores and variance bounds under non-anticipating probabilistic treatment assignment.
- Temporal causal estimands: The framework uses adapted propensity scores to construct estimators for current, lagged, and stepped temporal effects.The treatment-history window is encoded through conditional assignment probabilities.
- Variance and temporal effects: Under non-anticipating probabilistic assignment, the proposed estimators have martingale-difference errors and variance that can be bounded.The appendix derives estimable upper bounds for these variances.
- Temporal causal estimands: An m-period causal-impact condition states that outcomes at time t depend on treatment history through m prior periods.The m = 0 case restricts influence to the current outcome value.
- Exact randomization inference: Randomization tests generate treatment paths under the assignment mechanism to approximate the null distribution and estimate a p-value.As the number of simulated assignments increases, the simulated average estimates the null p-value more closely.
- Inference: The standardized estimator has a known variance under the sharp null, but simulation evidence elsewhere shows its test has lower power than exact and conservative alternatives.The standardized construction is based on a martingale-difference sequence.
B.4.1 Simulation evidence
The simulation evidence compares standardized testing with exact and conservative procedures. The standardized test has worse power than both alternatives.
- Simulation evidence: The standardized test has worse power than the exact and conservative tests in the simulation experiments.The comparison is reported for the simulation results shown in Figure 8.
B.4.2 Empirical results
The empirical analysis compares standardized and unstandardized randomization-based results, examines test power, and reports trading-market outcomes using pooled and market-specific analyses.
- The standardized statistics are broadly in line with the unstandardized ones in the empirical analysis.
- The conservative test performs only slightly worse than the exact randomization test across treatment-effect and autoregressive-parameter comparisons.The comparisons use fixed φ = 0.5 for treatment-effect changes and fixed τ = 0.5 for φ changes, with Gaussian noise of variance 1.
- The simulation histograms show that all four estimators have similar variance and are centered at the true data-generating value.
- The pooled estimator has lower variance than the corresponding unpooled version in simulations with n = 2 experiments and T = 100.
- The market analysis records slippage relative to the asset mid-price and presents randomization distributions and pooled hypothesis-test results for the markets.Figure 11 displays slippage over time for eight markets, while Figure 12 displays randomization distributions for Markets 2 through 9.