Source-linked AI summary

Program Evaluation and Causal Inference with High-Dimensional Data

Alexandre Belloni, Victor Chernozhukov, Ivan Fernández-Val, Christian Hansen

arXiv:1311.2645v8math.STecon.EMstat.MEstat.ML

TL;DR

The paper asks how to conduct reliable treatment-effect inference with many controls, endogenous treatment, heterogeneous effects, and functional outcomes. It combines approximate sparsity, regularization, and orthogonal or doubly robust moments to construct efficient estimators and honest confidence bands. The framework covers local and exogenous-treatment effects and is illustrated using 401(k) participation and eligibility.

  • Problem

    Reliable causal inference is difficult when treatment is endogenous or heterogeneous and controls are high-dimensional, especially because model selection can omit moderate confounders and bias inference.

  • Method

    The paper uses approximately sparse reduced-form relationships, regularization or selection methods, and orthogonal moment conditions for uniformly valid post-selection inference.

  • Results

    The framework provides efficient estimators and honest inference for LATE, LQTE, ATE, QTE, distributional effects, and function-valued parameters, with an application finding larger 401(k) participation effects at high asset quantiles.

  • Takeaways & Limitations

    Orthogonal moments allow valid treatment-effect inference after imperfect high-dimensional model selection, including inference over continua of quantiles and other functional parameters.

  • Takeaways & Limitations

    The approach requires approximately sparse predictive relationships and nuisance estimators converging at o(n^-1/4); non-orthogonal moments generally fail to deliver uniform root-n inference.

Abstract

from arXiv · show

In this paper, we provide efficient estimators and honest confidence bands for a variety of treatment effects including local average (LATE) and local quantile treatment effects (LQTE) in data-rich environments. We can handle very many control variables, endogenous receipt of treatment, heterogeneous treatment effects, and function-valued outcomes. Our framework covers the special case of exogenous receipt of treatment, either conditional on controls or unconditionally as in randomized control trials. In the latter case, our approach produces efficient estimators and honest bands for (functional) average treatment effects (ATE) and quantile treatment effects (QTE). To make informative inference possible, we assume that key reduced form predictive relationships are approximately sparse. This assumption allows the use of regularization and selection methods to estimate those relations, and we provide methods for post-regularization and post-selection inference that are uniformly valid (honest) across a wide-range of models. We show that a key ingredient enabling honest inference is the use of orthogonal or doubly robust moment conditions in estimating certain reduced form functional parameters. We illustrate the use of the proposed methods with an application to estimating the effect of 401(k) eligibility and participation on accumulated assets.

1. Introduction

The paper addresses causal treatment-effect estimation when treatment assignment, heterogeneous effects, and the choice among many controls complicate inference. It develops uniformly valid methods based on orthogonal moments, approximate sparsity, and high-dimensional regularization for scalar and function-valued parameters.

  • Motivation: Treatment-effect analysis becomes difficult when assignment is non-random, effects are heterogeneous, and researchers must select controls from many variables and transformations.The framework considers endogenous binary treatment with a binary instrument and accommodates many potential controls.
  • Motivation: Post-selection inference can be invalid because moderate nonzero controls may be omitted, creating substantial omitted-variables bias.The paper avoids relying on restrictive beta-min conditions that make perfect model selection possible.
  • Contributions: The paper provides theoretically valid program-evaluation inference in approximately sparse models even when model selection is imperfect.Its procedures target key treatment-effect parameters and remain valid across models where perfect selection may be impossible.
  • Contributions: Orthogonal moment conditions make estimation first-order insensitive to errors in high-dimensional nuisance parameters, supporting regular root-n inference after regularization.The nuisance estimator is required to achieve o(n^-1/4) error and avoid excessive complexity or overfitting.
  • Contributions: The framework covers LATE, LQTE, distributional effects, and functional response data, while also extending to exogenous-treatment ATE and QTE settings.Functional response methods support inference over a range of quantile indices rather than only one quantile.
  • Contributions: The paper develops general uniformly valid inference for continua of moment-condition parameters, including functional central limit theorems, multiplier bootstrap validity, and a functional delta method.The results apply beyond treatment effects to structural econometric and other data-science problems.

2. The Treatment Effects Setting and Target Parameters

The paper studies treatment effects with binary treatment and instruments, allowing endogenous receipt, functional outcomes, and many controls. Its target parameters include local effects for compliers and treated compliers, with ATE and QTE as exogenous-treatment special cases.

  • 2.1. Observables and Reduced Form Parameters: The observed data include functional outcomes indexed by u, covariates X, binary instrument Z, and binary treatment D.Functional outcomes can represent threshold indicators, age-indexed height, or dosage-indexed health outcomes.
  • 2.1. Observables and Reduced Form Parameters: The framework treats D as potentially endogenous while assuming Z is randomly assigned conditional on observable covariates X.The 401(k) application treats eligibility as exogenous only after conditioning on income and other individual characteristics.
  • 2.1. Observables and Reduced Form Parameters: Reduced-form parameters average conditional expectations gV(z,x)=EP[V|Z=z,X=x] over X and underpin the structural treatment-effect parameters.The target variable V varies by context, including treatment-state outcome and treatment indicators.
  • 2.2. Target Structural Parameters – Local Treatment Effects: LATE measures the treatment-effect difference for compliers, the subgroup whose treatment status can be influenced by the instrument, rather than the entire population.When D≡Z, the local structural function and LATE become the population ASF and ATE.
  • 2.2. Target Structural Parameters – Local Treatment Effects: LQTE compares potential-outcome quantiles between treated and non-treated states for compliers, extending the framework from scalar to distributional effects.Threshold-indexed outcomes generate local distribution treatment effects, whose quantile inverse yields LQTE.
  • 2.3. Target Structural Parameters – Local Treatment Effects on the Treated: LATE-T measures the average effect for treated compliers, while corresponding treated-group distributional and quantile effects are also defined.Under conditional exogeneity, LQTE and LQTE-T reduce to QTE and QTE-T.
  • 2.3. Target Structural Parameters – Local Treatment Effects on the Treated: The proposed results provide estimation and inference for these treatment-effect-on-the-treated quantities and their exogenous-treatment special cases.Random assignment conditional on controls yields ASF-T and ATE-T; conditional exogeneity yields QTE and QTE-T.

3. Estimation of Reduced-Form and Structural Parameters in a Data-Rich Environment

The estimation strategy models reduced-form functions with high-dimensional methods, estimates their functionals using orthogonal equations, and obtains structural effects by plug-in. Approximate sparsity and orthogonality support uniformly valid inference, including functional confidence bands.

  • 3. Estimation of Reduced-Form and Structural Parameters in a Data-Rich Environment: The key reduced-form quantities are αV(z)=EP[gV(z,X)] and γV=EP[V], where gV(z,X)=EP[V|Z=z,X].The variable V changes with the treatment-effect functional under study.
  • 3. Estimation of Reduced-Form and Structural Parameters in a Data-Rich Environment: The procedure estimates predictive relationships, estimates reduced-form parameters with orthogonal equations, and then obtains structural effects through a plug-in rule.Orthogonal equations immunize reduced-form estimation against imperfect first-step model selection.
  • 3.1. First Step: Modeling and Estimating gV and mZ: The functions gV and mZ are approximated using generalized linear combinations of rich control dictionaries and known links such as linear, logistic, or probit.Approximation errors are represented by rV and rZ.
  • 3.1. First Step: Modeling and Estimating gV and mZ: Approximate sparsity permits many potential controls when only a small number of coefficients achieve sufficiently small approximation errors.The dictionary dimension may exceed sample size, subject to conditions including log p=o(n^1/3).
  • 3.1. First Step: Modeling and Estimating gV and mZ: Lasso and Post-Lasso estimators of gV and mZ achieve near-oracle convergence rates while allowing a continuum of response variables.The paper’s results extend high-dimensional sparse modeling beyond settings with known relevant controls.
  • 3.2. Second Step: Robust Estimation of the Reduced-Form Parameters αV (z) and γV: Orthogonal moment functions, tied to efficient influence functions, reduce sensitivity to the non-regularity of penalized and post-selection estimators.The approach targets reduced-form parameters such as αV(z) and γV.
  • 3.2. Second Step: Robust Estimation of the Reduced-Form Parameters αV (z) and γV: The reduced-form process is asymptotically Gaussian uniformly over rich data-generating-process classes that include imperfect model selection.The multiplier bootstrap consistently approximates its large-sample law.
  • 3.3. Third Step: Robust Estimation of the Structural Parameters: Structural estimators are asymptotically Gaussian, and the resulting theory supports simultaneous confidence bands and tests of functional hypotheses.Structural results follow from reduced-form theory and a functional delta method extended to uniformity in P.

4. Theory: Estimation and Inference on Local Treatment Effects Functionals

The section establishes uniform Gaussian limit theory and bootstrap validity for reduced-form processes and smooth structural functionals under approximate sparsity and regularity conditions. These results support uniformly valid inference despite possible model-selection mistakes.

  • Functional data: The framework accommodates functional response data through measurability, continuity, entropy, instrument-strength, and boundedness conditions.These conditions cover the continuum of indexed outcomes and structural parameters used in the theory.
  • Assumptions: Approximate sparsity and boundedness conditions control nuisance-function approximations and sparse estimators uniformly over data-generating processes.The assumptions also require equivalence of empirical and population norms on sparse subsets.
  • Reduced-form inference: The reduced-form empirical process admits a uniform linearization, √n(bρ −ρ) = Zn,P + oP(1), and is asymptotically Gaussian.The Gaussian process has bounded, uniformly continuous paths under the stated conditions.
  • Reduced-form inference: The multiplier bootstrap consistently approximates the large-sample law of the reduced-form process uniformly over P ∈ Pn.This provides a basis for inference on reduced-form parameters when perfect model selection is impossible.
  • Structural functionals: For smooth structural functionals, uniform Hadamard differentiability transfers the limit theory and bootstrap validity from reduced forms to normalized structural estimators.The resulting limit is a zero-mean tight Gaussian process, and the conditional bootstrap law approaches the estimator’s large-sample law.

5. General Theory: Honest Inference in General Moment Condition Problems with Nuisance Functions Estimated by Machine Learning Methods

The paper develops uniform inference for a continuum of moment-condition parameters with high-dimensional nuisance functions estimated by machine learning. Orthogonal moments, functional limit theory, bootstrap validity, and data-splitting support honest inference under approximate sparsity.

  • Framework: The framework targets a continuum of function-valued parameters identified by moment conditions with potentially infinite-dimensional nuisance functions.It treats nuisance functions as objects that may vary with the indexing parameter u.
  • Novelty: The continuum result is presented as new because prior approaches relied on Donsker conditions that are precluded by the estimated nuisance-function classes considered here.The paper therefore replaces that route with uniform functional limit theory compatible with high-dimensional machine-learning nuisance estimation.
  • Nuisance estimation: Approximately sparse nuisance functions can be estimated using Lasso, Post-Lasso, and other modern statistical or machine-learning methods.The framework also allows over-identified cases through pointwise optimal combinations of moment conditions, although these are not analyzed explicitly.
  • Orthogonality: Neyman orthogonality makes the moment condition insensitive to small perturbations in nuisance estimates and supports regular estimation of the target parameter.In the linear example, the resulting score is semiparametrically efficient, and the estimator is uniformly root-n consistent and asymptotically normal under approximately sparse models.
  • Uniform inference: Theorem 5.1 establishes a functional central limit theorem for the continuum of estimators uniformly over a broad class of data-generating processes.The limiting process has almost surely uniformly continuous paths under the stated regularity conditions.
  • Bootstrap inference: The multiplier bootstrap consistently approximates the large-sample law of the estimator process and of smooth functionals obtained through the functional delta method.The validity result requires a uniformly accurate estimator of the relevant derivative or Jacobian object.

6. Theory: Lasso and Post-Lasso for Functional Response Data

This section develops uniform Lasso and Post-Lasso theory for functional responses under linear and logistic links. The results control selection errors over an index set and establish uniformly valid performance under approximate sparsity.

  • Model and estimators: The section develops Lasso and Post-Lasso estimators for functional responses under linear or logistic link functions.The results are presented autonomously but also support nuisance estimation and treatment-effect inference.
  • Model and estimators: Approximate sparsity represents each indexed response relationship using many transformations of covariates plus a small approximation error.The dictionary may contain powers, splines, and interactions, while the link is linear or logistic.
  • Uniform estimation: Under logistic-link conditions, the functional-response Lasso and Post-Lasso likewise satisfy uniform performance bounds and uniform sparsity.These results cover binary functional responses through the logistic link.
  • Uniform estimation: Penalty levels and loading matrices are chosen to control selection errors uniformly over the response index u.The implementation iteratively updates penalty loadings using Lasso and Post-Lasso residuals.
  • Uniform estimation: Under sufficient regularity conditions, the linear-link Lasso is uniformly sparse and satisfies uniform performance bounds, with corresponding results for Post-Lasso.The rates apply uniformly over probability laws and the index set.

7. Application: the Effect of 401(k) Participation on Asset Holdings

The application estimates 401(k) eligibility and participation effects on accumulated assets using increasingly rich controls and high-dimensional selection. Selection-based estimates are stable, while unselected flexible specifications can be erratic or substantially less precise.

  • Data and design: The design addresses endogenous 401(k) participation by treating eligibility as an instrument after conditioning on controls, especially income-related confounders.The high-dimensional approach permits broad control dictionaries while assuming a relatively low-dimensional confounding structure.
  • Data and design: The application uses 9,915 SIPP households, net financial assets as the outcome, participation as treatment, and eligibility as the instrument.It evaluates multiple treatment effects under several control specifications.
  • Average effects: Selection-based average-effect estimates are stable across control specifications and broadly consistent with low-dimensional estimates.Bootstrap and analytic standard errors are also quite similar.
  • Average effects: In the quadratic-spline-plus-interactions specification, unselected ATE and LATE estimates are substantially larger and their standard errors are roughly three times larger than selected estimates.The selected estimates remain similar to those from the low-dimensional setting.
  • Quantile effects: Selected LQTE estimates indicate a small participation effect at low quantiles and a larger effect at high quantiles, with uniform intervals rejecting zero and constant effects.The result is statistically significant at the 5% level.
  • Variable selection: Across specifications, selected models use between two and 22 variables for several average-effect results, while quantile selections range from none to 237 variables across u.Selected variables mostly capture income-related effects.

B.6. Auxiliary Result: Conditional Multiplier CLT in Rd uniformly in P ∈P.

This auxiliary section establishes conditional multiplier central limit results uniformly over probability laws and supports empirical-process limits for classes that change with sample size. The arguments use subsequence convergence and moment, entropy, and measurability conditions.

  • Conditional multiplier CLT: The conditional multiplier CLT assumes independent random vectors and multipliers with zero mean, unit variance, uniformly bounded q-th moments, and q > 2.The result is formulated uniformly over the indexed probability laws.
  • Conditional multiplier CLT: Along suitable subsequences, the distribution of the random vectors converges in Mallow’s metric and the associated covariance matrices converge to a bounded limit.The limiting distribution and covariance may depend on the subsequence.
  • Conditional multiplier CLT: The multiplier distribution converges weakly to the corresponding Gaussian limit in bounded-Lipschitz distance.The proof combines Lindeberg’s central limit theorem with an extended continuous mapping argument.
  • Functional extensions: The auxiliary proofs extend these convergence arguments through subsequence decompositions, continuous mappings, and functional derivatives.The framework also handles bootstrap processes and smooth functionals under appropriate continuity conditions.
  • Changing function classes: For empirical-process classes that change with n, entropy and envelope conditions yield stochastic equicontinuity and Gaussian convergence when covariance functions converge pointwise.The function class itself may depend on the underlying probability law through n.

E.1. Proof of Theorem 5.1.

The proof establishes uniform convergence and bootstrap validity for a continuum of target parameters by combining preliminary rates, linearization, orthogonality, and empirical-process arguments.

  • Step 1. A Preliminary Rate Result: The preliminary estimator achieves a uniform rate sup_u∈U ∥bθ_u−θ_u∥≲τ_n with probability 1−o(1).
  • Step 3. Linearization: Linearization decomposes estimation error into parameter, nuisance, and empirical-process terms, which are controlled through successive bounds.
  • Orthogonality: Orthogonality makes the nuisance-direction contribution vanish or become negligible, preventing first-order bias from nuisance estimation.
  • Empirical-process control: Entropy, envelope, moment, and growth conditions yield uniformly controlled empirical processes for the function classes used in the proof.
  • Limit theory: The argument concludes with uniform stochastic convergence and a functional central limit theorem, including multiplier-bootstrap validity.

F.1. Causal Interpretations for Structural Parameters.

The paper gives causal interpretations for structural treatment parameters under exogeneity, first-stage relevance, instrument non-degeneracy, and monotonicity.

  • Causal parameters: The framework distinguishes causal ASF, ASF-T, LASF, LASF-T, ATE, ATE-T, LATE, and LATE-T parameters.
  • Assumption F.1: Identification requires conditional instrument exogeneity, a nonzero first stage, instrument overlap, and monotonicity of potential participation.
  • Identification: Under these assumptions, complier potential-outcome means equal the identified parameters θ_Yu(d) and ϑ_Yu(d).
  • Identification: The identification results connect effects among compliers and treated compliers to the corresponding structural causal quantities.

Appendix G. Additional Results for Section 3

The appendix analyzes an alternative high-dimensional strategy that models conditional outcome and treatment probabilities directly, preserving approximate sparsity and near-oracle estimation rates.

  • Constructing g_V: The relation e_V(d,z,x)l_D(d,z,x) is used to construct the resulting estimator for g_V.
  • Regression functions: The functions e_V and l_D represent conditional means and conditional treatment probabilities given D, Z, and X.
  • Alternative modeling strategy: The alternative strategy models e_V(d,z,x) and l_D(d,z,x) as generalized linear approximations with sparse coefficients and approximation errors.
  • Estimation results: The estimator of e_V achieves the near-oracle rate p(s log p)/n under the stated sparse approximation framework.
  • Estimation results: The remaining estimation steps follow the strategy developed in the main text.

Appendix H. Additional Results for Section 4

The appendix supplies additional conditions and finite-sample results supporting functional inference and bootstrap procedures for high-dimensional Lasso and Post-Lasso estimators.

  • Assumption H.1: Assumption H.1 imposes approximate sparsity, bounded approximation errors, controlled dimensionality, estimator accuracy, and sparse-eigenvalue conditions.
  • Functional limit theory: Under the stated assumptions, the alternative reduced-form process satisfies a functional central limit theorem and a multiplier-bootstrap functional central limit theorem.
  • Lasso and Post-Lasso: Finite-sample Lasso and Post-Lasso analysis covers penalty selection, convergence rates, sparsity bounds, and restricted-eigenvalue conditions.
  • Sufficient conditions: The appendix develops sufficient conditions for functional-response processes, including VC structure, conditional smoothness, and boundedness cases.

I.3. Finite Sample Results:

The finite-sample analysis establishes uniform bounds for weighted logistic estimators, their sparsity, and post-selection convergence under restricted-eigenvalue and approximation conditions.

  • Weighted logistic analysis: The analysis uses both the design matrix and a weighted counterpart based on conditional outcome variances.The weights satisfy 0 ≤ w_ui ≤ 1.
  • Regularity conditions: A logistic restricted eigenvalue condition underpins finite-sample control of penalized estimators.The relevant sparse support is T_u = supp(θ_u), with sparsity bounded by s.
  • Penalized estimator: Uniform finite-sample bounds are derived for the ℓ1-penalized logistic estimator when the stated side conditions hold.Lemma I.6 provides rates of convergence under the displayed assumptions.
  • Sparsity: The ℓ1-penalized logistic estimator has uniformly controlled support size under a penalty and score condition.Lemma I.7 bounds the number of non-zero coefficients in bθ_u uniformly over u.
  • Post-selection estimation: Post-ℓ1-logistic estimation delivers uniform convergence bounds after selecting the support by penalized logistic regression.The post-selection estimator is defined using the selected support eT_u and requires conditions on the combined sparse set.
  • Norm consequences: Sparse-vector norm relations convert prediction or restricted-eigenvalue results into ℓ1-norm convergence bounds.For a sparse vector with k non-zero entries, ℓ1 control follows through φ_min(k).

Appendix J. Additional Results for Section 7

The appendix supplements the main application with additional wealth outcomes, control specifications, treatment definitions, and plots of quantile treatment effects.

  • Outcomes and controls: Additional results use both total wealth and net total financial assets as outcome variables.The appendix reports detailed results for four different sets of controls.
  • Treatment effects: The appendix reports intention-to-treat effects using 401(k) eligibility and instrumental-variable results using participation instrumented by eligibility.It also plots QTE and QTE-T for eligibility treatment and LQTE and LQTE-T for instrumented participation.

Appendix K. Auxiliary Results: Algebra of Covering Entropies

The auxiliary results establish entropy bounds for function classes and provide empirical-process ingredients used to control the paper’s high-dimensional estimators and processes.

  • Entropy algebra: Lemma K.1 develops covering-entropy algebra for VC-subgraph classes and related function-class operations.The results cover products, unions, and Lipschitz-type transformations of function classes.
  • Conditional expectations: Lemma K.2 bounds the entropy of conditional-expectation classes when the integration measure varies with the conditioning variable.The result extends fixed-measure integral-class arguments to conditional distributions μ_w.
  • Proof tools: The auxiliary proof framework uses Jensen’s inequality and finitely discrete probability measures to establish the entropy comparison.The resulting bound is weakly increasing in the integrability index s.
  • Process setup: The theorem-proof setup defines influence functions, low-bias moments, nuisance-function classes, and stacked estimator processes.These objects organize the subsequent linearization and empirical-process arguments.
  • Linear representation: Under the stated sparsity and approximation assumptions, the estimator process admits a linear representation in ℓ∞(U)^dρ.The displayed result is √n(bρ − ρ) = Z_n,P + o_P(1).
  • Empirical-process control: The auxiliary conditions establish bounded envelopes and uniform covering-entropy controls for the function classes used in the process approximation.These controls rely on VC-subgraph structure and Lemma K.2.

L.2. Proof of Theorem 4.2.

The proof of Theorem 4.2 linearizes the estimator and establishes weak convergence and bootstrap validity through empirical-process and influence-function arguments.

  • Setup: The proof uses empirical-process notation for independent observations and the corresponding centered process G_n.The preparation step fixes the asymptotic sequence of probability measures and expectations.
  • Linearization: The estimator’s linearization error arises entirely from estimating the influence function and would vanish if that function were known.The error is denoted ζ_n,P(D_n, B_n).
  • Influence-function estimation: The proof controls the plug-in influence-function representation uniformly over variables and treatment indices.The displayed bounds use the nuisance estimators and empirical-process terms.
  • Remainder control: Entropy and envelope bounds control the relevant function class, yielding a negligible remainder in the uniform representation.The argument applies covering-number bounds and exponential-tail control for the multiplier variables.
  • Weak convergence: The normalized estimator process converges weakly to a Gaussian process under sequences of probability laws in the model class.The limit is expressed through the empirical process of the influence functions.
  • Changing classes: A Donsker theorem for classes changing with sample size supplies asymptotic tightness and Gaussian convergence when covariance functions converge pointwise.The proof permits the probability law and indexed process to depend on n.

Appendix N. Proofs for Section 6 and Appendix I

The appendix proves uniform convergence and sparsity results for post-Lasso estimators under approximate-sparsity, moment, and sparse-eigenvalue conditions. These results support valid high-dimensional inference across a continuum of target parameters.

  • Proof strategy: The proofs establish uniform results by applying lemmas under sequences of probability measures P = P_n within the model class.The arguments invoke Lemmas I.3–I.8 and verify their required events and conditions.
  • Penalty-loadings control: Penalty-loadings events hold with probability 1 − o(1), including bounds relating estimated loadings to ideal loadings uniformly over parameters and covariates.The events E1, E2, and E3 control approximation errors, empirical score processes, and penalty loadings.
  • Assumptions: Approximate sparsity and covering conditions imply the required high-dimensional concentration condition for the estimating equations.The proofs use bounded second and third moments together with tuning-parameter choices to verify Condition WL.
  • Rates and sparsity: Sparse-eigenvalue conditions keep the relevant restricted eigenvalue quantities bounded away from zero, enabling convergence and sparsity bounds.This conclusion is obtained for both the general and binary-outcome settings under their respective assumptions.
  • Rates and sparsity: The Post-Lasso estimators satisfy the stated uniform convergence rates and sparsity bounds with probability approaching one.The proof concludes by combining the verified events with Lemmas I.5 and I.8.

Appendix O. Simulation Experiment

The simulation compares orthogonal and naive post-selection inference in a high-dimensional exogenous-treatment design. It also reports quantile and local quantile treatment-effect estimates across increasingly rich control specifications.

  • Inference comparison: The simulation evaluates a proposed estimator using orthogonal estimating equations and variable selection against a naive variable-selection estimator.Both procedures use post-model-selection conditional-expectation estimates, while differing in their estimating equations.
  • Simulation design: The design uses exogenous treatment conditional on controls, 250 covariates, and a sample size of 200.Covariates are Gaussian with AR(1)-type covariance, and coefficient strength varies across data-generating designs.
  • Inference comparison: The proposed procedure is evaluated through 5% average-treatment-effect t-tests using model selection, orthogonal estimating equations, and plug-in standard errors.The proposed tests are displayed in the right panel of Figure 11, while naive tests appear in the left panel.
  • Inference comparison: Naive 5% tests exhibit substantial size distortions for many coefficient designs, with near-nominal size occurring in only a handful of cases.The underlying data-generating process is unknown in practice, making the naive procedure's favorable cases difficult to identify.
Loading 1311.2645v8…