Source-linked AI summary
An Exact and Robust Conformal Inference Method for Counterfactual and Synthetic Controls
Victor Chernozhukov, Kaspar Wüthrich, Yinchu Zhu
TL;DR
Policy-effect inference with counterfactual and synthetic-control methods requires robust procedures for assessing causal effects. This paper develops a general conformal permutation framework that is exact under exchangeability and finds a significant decrease in female gonorrhea incidence in its application.
Problem
The paper addresses how to make inferences on the causal effect of policy interventions.
Method
It develops a generic counterfactual modeling framework and robust inference procedures that encompass diverse counterfactual-analysis methods.
Results
The procedures have exact finite-sample validity under iid or exchangeable data, and the application finds that criminalizing indoor prostitution significantly decreased female gonorrhea incidence.
Takeaways & Limitations
The framework provides generic procedures for making inferences on policy effects estimated by counterfactual and synthetic-control methods.
Takeaways & Limitations
The difference-in-differences results should be interpreted with caution.
Abstract
from arXiv · showhide
We introduce new inference procedures for counterfactual and synthetic control methods for policy evaluation. We recast the causal inference problem as a counterfactual prediction and a structural breaks testing problem. This allows us to exploit insights from conformal prediction and structural breaks testing to develop permutation inference procedures that accommodate modern high-dimensional estimators, are valid under weak and easy-to-verify conditions, and are provably robust against misspecification. Our methods work in conjunction with many different approaches for predicting counterfactual mean outcomes in the absence of the policy intervention. Examples include synthetic controls, difference-in-differences, factor and matrix completion models, and (fused) time series panel data models. Our approach demonstrates an excellent small-sample performance in simulations and is taken to a data application where we re-evaluate the consequences of decriminalizing indoor prostitution. Open-source software for implementing our conformal inference methods is available.
1 Introduction
The paper develops generic, robust inference procedures for policy effects estimated by counterfactual and synthetic control methods. It reframes inference as counterfactual prediction and structural-break testing, supporting high-dimensional estimators under correct specification and misspecification.
- Problem: The setting is policy-effect inference for one treated unit observed before and after intervention, potentially alongside many untreated control units.The paper targets aggregate time-series applications using counterfactual and synthetic control methods.
- Framework: The framework nests and generalizes difference-in-differences, synthetic control, factor, matrix-completion, interactive fixed-effects, and time-series approaches.It focuses on methods generating mean-unbiased proxies for untreated counterfactual outcomes.
- Inference method: The procedures recast inference as counterfactual prediction and structural-break testing, obtaining p-values by permuting blocks of estimated residuals across time.They target hypotheses about effect trajectories and per-period effects, with pointwise confidence intervals constructed by test inversion.
- Validity: The methods are approximately valid when counterfactual estimators are consistent and shocks are stationary and weakly dependent.They are also valid under misspecification when the data satisfy stationarity and weak dependence and the estimated counterfactual model is stable under perturbations.
- Theoretical results: Theoretical results provide finite-sample non-asymptotic size bounds, implying exactness as T0 →∞ and exact finite-sample validity for iid or exchangeable data.The bounds describe how factors affect finite-sample accuracy and show that imposing the null can substantially improve small-sample size accuracy when exchangeability fails.
- Additional contributions: Additional contributions include a tuning-free constrained Lasso that nests synthetic control and difference-in-differences, consistency results with many controls, and extensions to averages, multiple treated units, and placebo tests.The paper also establishes validity with time-series misspecification and stability and provides verifiable sufficient conditions for many CSC methods.
2 A Conformal Inference Method
The method models untreated potential outcomes with mean-unbiased counterfactual predictors and tests post-treatment policy-effect trajectories using residual-based permutations. Under iid data and a null-imposed counterfactual estimate, it has exact finite-sample validity, while permutation choices affect precision and asymptotic requirements.
- Counterfactual Model: The counterfactual model assumes mean-unbiased predictors or proxies for untreated potential outcomes and permits fixed or random predictors without restricting their dependence on shocks.The shock sequence is centered and stationary, with additional assumptions later allowing iid or weakly dependent processes.
- Counterfactual Model: The fundamental identifying assumption is that the shock sequence is invariant under the intervention, requiring intervention timing to be independent of factors changing its distribution.If policy changes shock distributions, the method can instead be interpreted as a structural-break test or provide valid prediction sets for random policy effects.
- Hypotheses and Residuals: The procedure tests a post-treatment policy-effect trajectory against a sharp null, which fully determines post-treatment counterfactual outcomes.It obtains a counterfactual proxy estimate under the null and computes residuals for inference.
- Hypotheses and Residuals: Under iid data, imposing the null guarantees model-free exact finite-sample validity, whereas estimating counterfactuals without the null prevents exact validity even with iid data.Null-imposed estimation is essential for good small-sample performance.
- Permutation Inference: The method uses iid or overlapping moving-block permutations to compute p-values; with exchangeable residuals, either set preserves exact finite-sample validity.The larger all-permutations set yields more precise p-values and permits testing at lower significance levels, while asymptotic validity depends on shock-process assumptions.
3 Theory
The theory establishes finite-sample and asymptotic validity for conformal inference under exchangeability, time-series dependence, and weak estimation conditions. It also accommodates misspecified or inconsistent estimators when they are stable, while identifying limitations under unconditional shock heteroscedasticity.
- 3 Theory: Under iid or exchangeable data, the procedure is exactly valid in finite samples.The time-series results are non-asymptotic, with finite-sample size bounds that imply exactness as T0 →∞.
- 3.1 Consistent Estimation: For time-series data, validity requires only weak, easy-to-verify estimator error conditions alongside stationarity and weak dependence of the stochastic shocks.The conditions include pointwise consistency and consistency in the prediction norm, and the framework accommodates non-stationary data.
- 3.2 Misspecification and Estimator Stability: Under stationarity and weak dependence, estimator stability—not consistency or correct specification—is sufficient for approximate validity under misspecification.Stable estimators are approximately independent of individual observations, enabling moving block permutations and accommodating high-dimensional CSC methods.
- 3.2 Misspecification and Estimator Stability: Theorem 2’s validity bound tends to zero under exponentially decaying mixing coefficients, while iid shocks permit iid permutations and dependent shocks require moving block permutations.Conditional on the estimator, the p-value is approximately uniform; with iid or exchangeable data, it is exactly uniform.
- 3.1 Consistent Estimation: Theorem 1 gives approximately unbiased p-values under consistent estimation, with a non-asymptotic bound supporting uniform validity across diverse data-generating processes.The bound depends on T∗, T0, log T, and constants related to the shock process, mixing, and shock-density bound.
4 Sufficient Conditions for Consistent Estimation
This section gives sufficient conditions ensuring counterfactual mean proxies are accurate enough for asymptotically valid inference. It establishes consistency results for constrained least squares and factor-based estimators under weak, standard regularity conditions.
- General conditions: The proposed inference procedure is asymptotically valid when counterfactual mean proxies satisfy sufficient estimation conditions that verify Assumption 3.The section assumes T0 → ∞ and, when applicable, J → ∞; regularity conditions include bounded moments and weak serial dependence.
- Constrained least squares: Constrained least squares estimators are consistent when the parameter space lies in a bounded ℓ1-ball, the true weights are feasible, and the data are exponentially β-mixing under identification.The result provides performance bounds and does not require sparsity, allowing dense weight vectors.
- Constrained least squares: The constrained least squares result permits large control dimensions, requiring only log J = o(T^c), where c > 0 depends on the β-mixing coefficients.This accommodates settings where the number of potential controls and time periods have similar orders of magnitude.
- Factor models: 1/N + 1/T is the pure factor model’s proxy mean-squared-error bound under standard Bai (2003) regularity conditions as N → ∞ and T → ∞.No restriction is imposed on the relative growth of N and T.
- Factor models: 1/T + 1/N is the interactive fixed-effects model’s proxy mean-squared-error bound under Bai (2009) conditions and the factor identification condition.Analogous validity results can be established for common correlated effects estimators, while the stated bound is intended to hold more generally than N and T having the same order.
5 Empirical Application
The application revisits Rhode Island’s 2003 decriminalization of indoor prostitution, examining female gonorrhea incidence with three counterfactual prediction methods. Placebo tests support the inference for synthetic control and constrained Lasso, while difference-in-differences warrants caution; all methods find later significant reductions that are robust to control-state exclusions.
- Data and design: The study examines log female gonorrhea incidence per 100,000 after Rhode Island decriminalized indoor prostitution in July 2003.The data come from the CDC’s Gonorrhea Surveillance Program and cover 1985 onward.
- Data and design: The analysis applies difference-in-differences, canonical synthetic control, and constrained Lasso with K = 1 using all other U.S. states and Washington, D.C. as potential controls.Constrained Lasso nests both difference-in-differences and synthetic control.
- Main results: The zero-effect null is rejected at the 10% level under both moving-block and iid permutations for all three methods.Pointwise 90% confidence intervals show similar results across methods.
- Placebo and specification checks: Placebo tests and pre-treatment residual plots support the inference procedure with synthetic control and especially constrained Lasso, but caution against interpreting difference-in-differences results uncritically.The placebo tests assess whether pre-treatment effects are absent, whose rejection would undermine the procedure’s assumptions.
- Main results: After effects were not or only marginally significant during the first three years, legalizing indoor prostitution significantly decreased female gonorrhea incidence thereafter.This pattern corroborates Cunningham and Shah (2018).
- Robustness: Leave-one-out analyses show that the results are not driven by a single control state, with all but one specification significant at the 10%-level.The check iteratively excludes states receiving non-zero synthetic-control or constrained-Lasso weights.
Figures Main Text
The main-text figures present small-sample properties, permutation mechanics, raw data, placebo checks, pointwise confidence intervals, and leave-one-out robustness checks. The application figures use state-level log female gonorrhea data from Cunningham and Shah (2018).
- Figures Main Text: Figure 1 reports small-sample-size properties at a nominal level of 10%.The experiment tests H0 : θT0+1 = 0 using T0 = 19 and J = 50, with canonical synthetic-control weights estimated as (1/3, 1/3, 1/3, 0, …, 0).
- Figures Main Text: Figure 2 provides a graphical illustration of the permutations.
- Figures Main Text: Figure 3 displays the raw state-level data on log female gonorrhea cases per 100,000.The data are from Cunningham and Shah (2018).
- Figures Main Text: Figures 4–6 present graphical placebo checks, pointwise confidence intervals, and leave-one-out robustness checks, respectively.The leave-one-out figure shows p-values from testing null hypothesis (17) after omitting one non-zero-weight control state; labels include DID, SC, and CL.
Tables Main Text … B Interpretation as a Structural Breaks Test
The paper extends its conformal inference procedures to placebo tests, average effects, and multiple treated units, while interpreting failures of distributional invariance as structural breaks. Placebo non-rejections support but do not prove validity, whereas the structural-break formulation has power against both policy-effect shifts and distributional changes.
- Tables Main Text: The main-text tables report placebo specification and zero-effect tests using p-values computed from pre-treatment and post-treatment hypotheses.The tables use data from Cunningham and Shah (2018), with placebo horizons τ ∈ {1, 2, 3} and a post-treatment null covering 2004–2009.
- Online Supplemental Appendix to “An Exact and Robust Conformal Inference: The supplemental appendix develops extensions for average effects over time, multiple treated units, and placebo testing.These extensions are presented in the Online Supplemental Appendix to the paper’s conformal inference method.
- A.1 Testing Hypotheses about Average Effects over Time: Average-effect hypotheses can be tested by collapsing observations into non-overlapping time blocks and permuting residuals from an estimated average proxy.The procedure requires consistent identification and estimation of the average proxy; its effective sample size is T/T∗, so T must be substantially larger than T∗.
- A.1 Testing Hypotheses about Average Effects over Time: For average-effect testing, identification may fail for nonlinear and dynamic models, unlike SC and other regression-based estimators under stated orthogonality and consistency conditions.The formal test properties rely on stationarity and weak dependence of the aggregated disturbances when the average proxy is consistently estimated.
- A.2 Multiple Treated Units: With multiple treated units, the method separately tests unit-specific effects and can test average treatment effects by averaging outcomes across treated units.The aggregated residuals are permuted, and the resulting test inherits the formal properties established for the core procedure.
- A.3 Placebo Tests: The authors recommend easy-to-implement placebo tests that impose a pre-intervention placebo and apply the inference method with an earlier cutoff.Plots of pre-treatment residuals can complement the formal testing results.
- A.3 Placebo Tests: Placebo rejections undermine the procedure’s assumptions, while non-rejections support but do not prove validity and cannot assess assumptions such as distributional invariance.In the empirical application, placebo tests favor SC and constrained Lasso but call for caution when interpreting difference-in-differences results.
- B Interpretation as a Structural Breaks Test: When disturbance-distribution invariance fails, the procedure becomes a structural-break test with power against nonzero policy effects and distributional changes that increase S(u) quantiles.Examples include scale shifts, under the stationary and weakly dependent null framework.
C Prediction Sets for Random Policy Effects
This section extends the procedure from fixed to random policy effects, showing that it yields valid prediction sets under conformal-prediction-style conditions. The resulting sets target policy effects defined as differences between potential outcomes, while allowing policy-induced changes in distributional features.
- Random policy effects: The procedure generates valid prediction sets for random policy effects θ_t, paralleling classical conformal prediction for future random target quantities.The fixed-effect assumption is replaced by a random-effects framework, including settings such as Cattaneo et al. (2021).
- Validity conditions: Theorem C.1 establishes validity when T* is fixed, the counterfactual predictor satisfies Assumption 3, the relevant Assumption 2 condition holds, and S(u) has density bounded by D.The theorem applies Assumption 2.1 for Π = Π_all and Assumption 2.2 for Π = Π_→.
- Validity conditions: The theorem’s approximation term is δ̃_T = (T*/T_0)^(1/4)(log T), with constant C depending on T*, M, and D but not T.This term appears in the characterization involving C_1−α(t) defined by Algorithm 1.
- Interpretation: The sets are called prediction sets rather than conventional confidence sets because they predict random policy effects defined by differences between the two potential outcomes.This terminology follows the conformal prediction literature and reflects the distinction discussed for synthetic-control confidence and prediction intervals.
- Interpretation: With random effects, θ_t captures policy impacts on mean proxies and other distributional features, while only the no-policy residual process {u_t^N} is required to be stationary.Policy-induced changes in the mean, variance, or other distributional features are absorbed into {θ_t}.
D Model-free Exact Validity under Exchangeability · E Sufficient Conditions for Estimator Stability · E.1 Generic Sufficient Condition for Low-dimensional Models
Under exchangeability, the conformal procedure has exact finite-sample size control without requiring a correct or consistent counterfactual estimator. The paper then gives sufficient conditions under which low-dimensional estimators remain stable and yield asymptotic size control.
- D Model-free Exact Validity under Exchangeability: Under iid or exchangeable data, the conformal inference procedure achieves exact finite-sample size control without requiring a correct or consistent counterfactual mean estimator.The result is model-free and relies on permutation symmetry rather than estimator consistency.
- D Model-free Exact Validity under Exchangeability: Exact validity follows when the estimator is permutation-invariant, which holds for regression-based synthetic controls, constrained Lasso, and penalized regression when the null is imposed during estimation.The required invariance may fail for dynamic autoregressive models.
- D Model-free Exact Validity under Exchangeability: The procedure requires exchangeable residuals, which can occur even with non-exchangeable data when transformations such as differencing remove a common trend.The difference-in-differences example illustrates how residual exchangeability can support finite-sample validity.
- D Model-free Exact Validity under Exchangeability: Imposing the null during estimation improves size accuracy in small-pre-treatment CSC settings and remains beneficial even with substantially larger pre-treatment samples.The passages specifically compare T0 = 19 and T0 = 99 and report better performance when the null is imposed.
- E Sufficient Conditions for Estimator Stability: For estimator stability, the paper presents generic sufficient conditions for low-dimensional models and notes that high-dimensional models require case-by-case analysis.Constrained Lasso is verified separately, while Ridge regression is treated in Appendix F.
- E.1 Generic Sufficient Condition for Low-dimensional Models: Uniform convergence of the full-sample and subset losses, continuity at a unique minimizer, and compactness of the parameter space imply maxH∈H ∥β̂(Z) − β̂(ZH)∥2 = oP(1).The condition applies to a class of subsets H, although Assumption 4 itself only requires a singleton.
- E.1 Generic Sufficient Condition for Low-dimensional Models: Estimator consistency to the pseudo-true value under weak conditions translates stability into asymptotic size control, yielding |P (p̂ ≤ α) − α| = oP(1).The argument uses a uniform law of large numbers and allows the stability tolerance to be a constant δ > 0.
E.2 Constrained Lasso
This section develops constrained Lasso perturbation-stability results under possible misspecification, requiring sparse solutions and stability of empirical moments rather than estimator convergence. The resulting residual stability can hold in high dimensions, while stability alone does not guarantee closeness to pseudo-true residuals.
- E.2 Constrained Lasso: Constrained Lasso stability is studied under misspecification, including nonlinear relationships or constraint sets that exclude the true parameter.The framework allows EX_t(Y_t − X_t′β) ≠ 0 for every β ∈ W.
- E.2 Constrained Lasso: Lemma E.2 guarantees perturbation stability when empirical covariance and moment estimates are stable, sparse solutions satisfy support bounds, and covariates are bounded.The stated conditions include sup-norm perturbation bounds, a sparse eigenvalue condition, support sizes at most s/2, and max_t∥X_t∥∞ ≤ κ2.
- E.2 Constrained Lasso: The stability argument does not require the constrained Lasso estimator to converge to any target, unlike prior one-observation perturbation results relying on correct specification and consistent variable selection.The authors identify this as the first result of this kind to their knowledge.
- E.2 Constrained Lasso: s = o(T0/ log(T0)) yields stable estimated residuals, and stability should readily hold when J ≪ T0/ log(T0) because solution supports are bounded by J.For |H| ≲ log(T0), the required sparsity condition ensures the perturbation effect vanishes asymptotically.
- E.2 Constrained Lasso: Stability does not imply that constrained Lasso residuals are close to pseudo-true residuals, because high-dimensional covariates can prevent prediction-error terms from vanishing.The section contrasts its weaker stability requirement with stronger conditions needed to establish closeness to pseudo-true residuals.
F Consistency and Estimator Stability
The Ridge example shows that estimator stability can hold under correct specification even when consistency fails. Under simple sub-Gaussian or boundedness conditions, this stability supports robust inference, while extensions to more general processes remain open.
- Motivation: Ridge regression may be inconsistent under correct specification while still satisfying estimator stability.The section uses Ridge regression as a simple, analytically tractable illustration of the distinction between stability and consistency.
- Robustness guarantee: This perturbation stability can preserve inference validity even when a badly chosen tuning parameter makes consistency fail.The procedure requires estimator stability under perturbations, whereas consistency would be expected under an ideal tuning-parameter choice.
- Sufficient conditions: Independent data across t with sub-Gaussian rows of X provide simple sufficient conditions for verifying Lemma F.1’s assumptions.The verification uses random matrix theory for X′X/T and the Hanson-Wright inequality for terms involving X′u.
- Limitations: The analysis is limited to simple sub-Gaussian or boundedness assumptions, with analogous results for more complicated processes left for future research.The authors state that Ridge can be stable without consistency in this simple setting and expect similar results more generally.
G Simulation Study … H.1.5 Proof of Lemma H.5
Simulations show that the conformal inference procedure has exact or near-correct size under exchangeability and stationary dependence, while nonstationary misspecification can distort size. The accompanying proofs establish approximate validity by combining permutation ergodicity and estimation-error bounds under moving-block and iid permutations.
- G Simulation Study: The study evaluates difference-in-differences, canonical SC, and constrained Lasso with K = 1 across varied data-generating processes.The simulations vary ρu, ρϵ, T0, J, and F2t, with DGP differences arising from the weight specification w.
- G Simulation Study: Exact size control holds for iid data, irrespective of whether the prediction model is correctly specified.The iid setting has ρu = ρϵ = 0, implying exchangeable residuals.
- G Simulation Study: Under stationary dependence with ρu = ρϵ = 0.6, size remains close to correct under both correct specification and misspecification.This supports robustness under estimator stability and stationarity.
- G Simulation Study: With trending factors F2t ∼ N(t, 1), correct specification yields excellent size, whereas misspecification can cause size distortions.The deterioration under misspecification is specific to the nonstationary setting considered.
- G Simulation Study: Under correct specification, the method has excellent small-sample power and approaches the oracle power bound based on the true marginal distribution of ut.The power analysis uses T0 = 19, J = 50, and ρu = ρϵ = 0.6.
- Additional Notation; H.1 Proof of Theorem 1: The proofs introduce notation and establish approximate validity through high-level conditions, permutation ergodicity, and estimation-error bounds.The arguments cover both moving block and iid permutations, with the theorem proof combining the corresponding lemmas.
- H.1.1 Proof of Lemma H.1: Lemma H.1 bounds the difference between feasible and oracle p-values using estimation error, approximate ergodicity, and the bounded density of S(u).The proof proceeds in two steps: controlling the p-value difference and deriving the desired test result.
H.2 Proof of Theorem 2 · H.2.1 Proof of Lemma H.6
The proof of Theorem 2 bounds the uniform discrepancy between the empirical distribution and Ψ through four steps, then derives the desired result. The proof of Lemma H.6 uses Berbee coupling and the Dvoretzky-Kiefer-Wolfowitz inequality to control the empirical process under β-mixing dependence.
- H.2 Proof of Theorem 2: The proof constructs coupled random elements that are independent of separated data while preserving the original distribution, enabling the comparison arguments.The construction is made on an enlarged probability space and uses β-mixing coupling.
- H.2 Proof of Theorem 2: Theorem 2’s proof proceeds in 4 steps: the first three bound sup_x∈R |F̂(x) − Ψ(x; β)|, while the fourth derives the desired result.The target distribution is Ψ(x; β) = P(φ(Z_t, ..., Z_t+T*−1; β) ≤ x).
- H.2 Proof of Theorem 2: The argument controls deviations through mixing and approximation terms, including βmixing(k − T* + 1), γ1,T, and γ2,T.The proof also selects m1 so that m1 ≍ (log m)1/D3 and m1/2βmixing(m1 − T* + 1) ≲ m−1.
- H.2 Proof of Theorem 2: The final calibration step establishes that the relevant transformed variable has uniform distribution on (0, 1), and identifies the events ˆp ≥ 1 − α and F̂(G(Z)) < α.The proof then concludes after combining the displayed bounds and choosing m1.
- H.2.1 Proof of Lemma H.6: Lemma H.6 begins by bounding the block-aggregation remainder using |1{W_t ≤ x} − G(x)| ≤ 1, yielding sup_x∈R |Δ(x)| ≤ T − mK ≤ m−1.This provides the initial deterministic control before coupling.
- H.2.1 Proof of Lemma H.6: Berbee’s coupling constructs variables with the same distribution as the original sequence whose corresponding block coordinates are independent across blocks.The construction satisfies P(∪_{t=1}^{mK}{W̄_t ≠ W_t}) ≤ mKβmixing(m) ≤ Tβmixing(m).
- H.2.1 Proof of Lemma H.6: Applying the Dvoretzky-Kiefer-Wolfowitz inequality to the independent coupled blocks controls the resulting empirical distribution uniformly.The coupled empirical distribution equals the original empirical distribution with probability at least 1 − Tβmixing(m).
H.3 Proof of Theorem C.1 · H.4 Proof of Theorem D.1
Theorem C.1 is established by relating the conformal event to a p-value and controlling the approximation through a bound involving ˜δ_T. Theorem D.1 follows from exchangeability of data and residuals, yielding finite-sample p-value validity and a no-ties lower bound.
- H.3 Proof of Theorem C.1: The proof defines the residual-based conformal p-value through the permutation distribution of S(ˆu_π(Z∗)), with ˆu(Z∗) = Y^N − ˆP^N.The predictor ˆP^N is computed using Z∗.
- H.3 Proof of Theorem C.1: Theorem C.1’s approximation argument uses ˜δ_T = (T∗/T_0)1/4(log T), with constant C depending on T∗, M, and D but not on T.This is the rate-control quantity introduced in the proof.
- H.3 Proof of Theorem C.1: The target coverage event θ_t ∈ C_1−α(t) is equivalent to the p-value event ˆp_Z∗ > α because Z(θ_0) = Z∗.The equivalence completes the proof of the desired result.
- H.4 Proof of Theorem D.1: Theorem D.1’s proof proceeds by showing that exchangeable data produce exchangeable residuals, which implies P(ˆp ≤ α) ≤ α.The proof is organized into three steps and invokes standard randomization-inference arguments.
- H.4 Proof of Theorem D.1: The randomization quantiles are invariant because Π_all and Π_→ form groups satisfying Π_π = Π for every π ∈ Π.This invariance supports the exchangeability argument for the p-value.
- H.4 Proof of Theorem D.1: When the distribution of {S(ˆu_π)}_π∈Π is continuous, ties occur with probability one zero, enabling the lower bound α−1/n ≤ P(ˆp ≤ α).The final step follows by arguments analogous to Step 2.
H.5 Proof of Lemma 1 … H.10 Proof of Lemma 6
The appendix establishes the lemmas underpinning the inference procedure under mixing, factor-model, matrix-completion, and time-series regularity conditions. The proofs derive estimator consistency and approximation bounds using coupling, concentration, and existing factor-model results.
- H.5 Proof of Lemma 1: Under moment, β-mixing, growth, and tuning conditions, Lemma 1 obtains high-probability bounds for the ℓ1-constrained estimator.The proof uses Berbee’s coupling to replace dependent observations with independent counterparts and shows the approximation terms vanish under log J = o(T 4τ/(3τ+4)).
- H.6 Proof of Lemma 2: Lemma 2 applies Bai (2003)’s regularity conditions to establish the factor-estimation relation ˆF_t − λ′_1F_t = OP(1/δNT).The argument relies on convergence of the rotation matrix H to a nonsingular limit and moving-block permutation size n = |Π| = T.
- H.7 Proof of Lemma 3: Lemma 3 uses Bai (2009)’s assumptions to control factor, loading, and residual-estimation errors, including pointwise bounds for the estimated factors.The proof combines bounds for ∆β and ∆F with N ≍ T and establishes ∥ˆF_t∥2 = OP(1).
- H.8 Proof of Lemma 4: Lemma 4 proves matrix-completion estimation accuracy under conditional cross-sectional independence, moment bounds, nuclear-norm constraints, and a vanishing rate condition.The proof bounds the error through trace duality and controls the noise norm using Lemma H.10, whose argument applies Liapunov’s inequality and concentration results.
- H.10 Proof of Lemma 6: Lemma 6 follows directly when the tuning sequence satisfies ℓT∥ˆρ − ρ∥ = oP(1).This condition transfers consistency of the autoregressive parameter estimate into the desired result.
H.11 Proof of Lemma 7 … I Tables and Figures Appendix
The appendix establishes technical lemmas supporting estimator behavior under stationarity, compactness, uniqueness, and matrix conditions, then documents simulation tables and figures for size and power. The proofs deliver consistency, pointwise convergence, eigenvalue control, and perturbation bounds under their stated assumptions.
- H.11 Proof of Lemma 7: H.11 derives bounded regressor moments and a positive lower eigenvalue with probability approaching one under stationarity and roots bounded away from the unit circle.These results support the proof of Lemma 7.
- H.11 Proof of Lemma 7: H.11 proves ˆρ = ρ + oP(1) and pointwise convergence when estimated residuals converge pointwise and the stationary regressors remain OP(1).The proof controls the average product of estimated regressors and perturbed errors before establishing the pointwise result.
- H.12 Proof of Lemma E.1: H.12 proves ∥ˆβ(Z) − β∗∥2 = oP(1) from compactness, continuity, and uniqueness of the minimum of L(·).The argument uses an event with probability 1 − o(1) and an arbitrary η > 0.
- H.12 Proof of Lemma E.1: H.12 extends the consistency result uniformly across H by showing maxH∈H ∥ˆβ(ZH)−β∗∥2 = oP(1).The bound holds on the same high-probability event used for the single-estimator result.
- H.13 Proof of Lemma E.2: H.13 bounds estimator differences uniformly over H, yielding ∆′ ˆΣ∆ ≥ κ1∥∆∥2 2 ≥ κ1s−1∥∆∥2 1 with probability at least 1 − γ1,T − γ2,T − γ3,T.The proof compares quadratic objectives along the segment between ˆβ and ˆβH.
- H.14 Proof of Lemma F.1: H.14 analyzes inconsistency and stability of a perturbed estimator under eigenvalue and perturbation-event conditions.The displayed rate establishes a lower bound proportional to λ2T/(T + λ)2 for the relevant expression.
- I Tables and Figures Appendix: The tables and figures appendix reports finite-sample size properties, power curves, and empirical rejection rates under the stated simulation designs.The simulations use nominal level α = 0.1 and 5000 repetitions; additional designs vary F2t and error autocorrelation.