Source-linked AI summary
Inference in High Dimensional Panel Models with an Application to Gun Control
Alexandre Belloni, Victor Chernozhukov, Christian Hansen, Damian Kozbur
TL;DR
The paper addresses estimation and inference in high-dimensional panel models with additive fixed effects, where many time-varying variables and within-individual dependence complicate variable selection. It develops Cluster-Lasso-based procedures with valid post-selection inference and supports them theoretically, through simulations, and in a gun-prevalence application. The application’s interpretation relies on using firearm-suicide shares as a useful gun-ownership proxy and on the included controls capturing confounding variation.
Problem
High-dimensional panel inference must handle many time-varying regressors while allowing unrestricted, potentially dense individual fixed effects.
Method
The paper demeans observations to remove fixed effects, applies Cluster-Lasso with dependence-robust penalty loadings, and builds inference procedures for fixed-effects and IV models.
Results
Cluster-Lasso retains favorable model-selection and prediction properties under clustered dependence, with simulations showing markedly better performance than procedures that ignore clustering.
Takeaways & Limitations
The methods support variable selection and inference with unrestricted fixed effects and within-individual dependence, and the gun application broadly agrees with earlier findings using broader controls.
Takeaways & Limitations
The empirical interpretation assumes firearm-suicide share is a useful gun-ownership measure and that included controls capture variables jointly associated with crime and that proxy.
Abstract
from arXiv · showhide
We consider estimation and inference in panel data models with additive unobserved individual specific heterogeneity in a high dimensional setting. The setting allows the number of time varying regressors to be larger than the sample size. To make informative estimation and inference feasible, we require that the overall contribution of the time varying variables after eliminating the individual specific heterogeneity can be captured by a relatively small number of the available variables whose identities are unknown. This restriction allows the problem of estimation to proceed as a variable selection problem. Importantly, we treat the individual specific heterogeneity as fixed effects which allows this heterogeneity to be related to the observed time varying variables in an unspecified way and allows that this heterogeneity may be non-zero for all individuals. Within this framework, we provide procedures that give uniformly valid inference over a fixed subset of parameters in the canonical linear fixed effects model and over coefficients on a fixed vector of endogenous variables in panel data instrumental variables models with fixed effects and many instruments. An input to developing the properties of our proposed procedures is the use of a variant of the Lasso estimator that allows for a grouped data structure where data across groups are independent and dependence within groups is unrestricted. We provide formal conditions within this structure under which the proposed Lasso variant selects a sparse model with good approximation properties. We present simulation results in support of the theoretical developments and illustrate the use of the methods in an application aimed at estimating the effect of gun prevalence on crime rates.
1. Introduction
Panel data accommodates unrestricted individual heterogeneity, but high-dimensional covariates and within-unit dependence complicate selection and inference. The paper develops clustered Lasso-based procedures and illustrates them in simulations and a gun-prevalence application.
- High-dimensional panel datasets contain many time-varying variables from measured characteristics, transformations, interactions, and flexible trends.
- Approximate sparsity makes high-dimensional estimation feasible by assuming many covariates are available but only a small number are important for prediction.
- Individual heterogeneity may be relevant for every individual, related to observed variables unrestrictedly, and different across individuals, making naive sparse-model methods unreliable.
- Within-individual dependence violates independent-observation assumptions and can cause substantial inference distortions when ignored.
- Cluster-Lasso accommodates clustered covariance while partialing out unrestricted additive fixed effects before variable selection.
- Simulations show Cluster-Lasso procedures perform markedly better than selection procedures that ignore clustering, while the gun application broadly agrees with earlier results using broader controls.
2. Dimension Reduction and Regularization via Lasso Estimation in Panels
The paper removes fixed effects by within-individual demeaning and applies Lasso with clustered penalty loadings to select variables under within-unit dependence. Post-selection estimation and inference rely on the selected model and valid regularization conditions.
- The estimation strategy eliminates fixed effects through within-individual demeaning, producing a within model for variable selection.
- Cluster-Lasso estimates coefficients by penalized minimization on the within model, using a main penalty level and covariate-specific penalty loadings.
- 2.2. Clustered Penalty Loadings.: Clustered penalty loadings accommodate dependence within individuals, heteroscedasticity, and non-Gaussian data.
- Post-Cluster-Lasso refits least squares using variables selected by Cluster-Lasso, whose good selected-model properties transfer to the post-selection estimator.
- The regularization event requires penalty parameters to dominate the score vector, setting coefficients insufficiently large relative to sampling noise to zero.
- Under valid feasible loadings, the regularization event holds with probability tending to one.
3. Regularity Conditions and Performance Results for Cluster-Lasso Under Grouped Dependence
The theory allows independent groups with unrestricted within-group dependence and establishes Cluster-Lasso performance under approximate sparsity, sparse-eigenvalue, and regularity conditions. These results support model selection and inference after removing fixed effects.
- The model permits individual effects to depend unrestrictedly on time-varying covariates, with observations independent across individuals but unrestrictedly dependent within individuals.
- The analysis covers both n →∞, T fixed and n →∞, T →∞ joint asymptotics.
- Within-group information ranges from ı_T = 1 under perfect dependence to ı_T = T under perfect independence.
- Approximate sparsity requires the outcome function to be well approximated by a linear combination of a dictionary with p ≫ n allowed.
- Sparse-eigenvalue conditions require only suitably small submatrices of the high-dimensional Gram matrix to be well behaved.
- Theorem 1 shows that feasible Cluster-Lasso selects a model of size at most Ks with probability 1 −o(1) and retains favorable selection and prediction properties under clustered dependence.
4. Applications of Cluster-Lasso
The applications use Cluster-Lasso to select instruments or controls in high-dimensional fixed-effects models, enabling uniformly valid inference after selection. The framework covers instrumental variables and partially linear treatment models while allowing unrestricted within-individual dependence.
- 4.1. Selection of Instruments: Cluster-Lasso selects a sparse set of instruments for high-dimensional fixed-effects IV models, providing good approximations to optimal instruments.The first-stage model assumes approximate sparsity after eliminating fixed effects.
- Applications: The analysis extends to a fixed small number of included exogenous variables and to vector-valued endogenous treatments, but the latter extension is omitted from the presentation.The IV setup also relies on conditions excluding weak identification through strong identification of the structural parameter.
- 4.1. Selection of Instruments: The selected-instrument IV estimator is consistent and asymptotically normal, with valid inference based on the usual clustered standard error estimator.The result holds uniformly over a large class of data-generating processes under the stated SMIV conditions.
- 4.2. Selection of Control Variables: The partially linear fixed-effects treatment model allows the treatment effect to be inferred while conditioning on high-dimensional confounders and unrestricted individual-specific heterogeneity.The model specifies separate outcome and treatment equations with additive fixed effects and allows dependence over time within individuals.
- 4.2. Selection of Control Variables: Post-double-selection OLS using variables selected by Cluster-Lasso is consistent and asymptotically normal, with conventional clustered standard errors supporting uniformly valid inference.This remains valid even when perfect variable selection is impossible.
5. Simulation Examples
The simulations evaluate Cluster-Lasso for fixed-effects IV and linear models with many variables, instruments, and dense individual heterogeneity. Clustered penalty loadings generally deliver the strongest feasible inference, while difficult or misspecified designs reduce performance.
- Simulation setup: The simulations assess Cluster-Lasso inference in fixed-effects IV models with many instruments and linear fixed-effects models.The designs vary sample size, instrument count, coefficient structure, and within-individual dependence.
- Simulation setup: The fixed effects are dense and may correlate with instruments, so selecting over them can invalidate instruments and inference.The design therefore eliminates fixed effects rather than treating them as a sparse set for selection.
- IV results: Design 3 is difficult because diffuse signal can yield no selected instruments and weak identification at small sample sizes.The problem diminishes with larger samples, but performance remains different from the oracle benchmark.
- IV results: Ignoring within-individual dependence in penalty loadings produces substantial bias and large size distortions, even with clustered standard errors.Smaller penalties can spuriously include instruments, creating post-selection endogeneity bias.
- IV results: Cluster-Lasso clearly dominates other feasible IV procedures, with approximately correct test size in Designs 1 and 2.In Design 1, its Bias, RMSE, and test size are similar to the infeasible Oracle benchmark.
- Linear-model results: In linear fixed-effects simulations, post-double-selection with clustered penalty loadings produces RMSE comparable to oracle procedures across designs.Competing procedures tend to have poor bias and coverage, while selecting over dense fixed effects performs poorly.
6. Empirical Example: The Social Cost of Gun Ownership, Cook and Ludwig (2006)
The empirical application extends the gun-prevalence and crime analysis to a much larger control set using fixed effects and Cluster-Lasso selection. The selected-control estimates support a positive association with firearm homicide, provide some evidence for overall homicide, and remain inconclusive for non-gun homicide.
- Design: The application reexamines Cook and Ludwig’s gun-prevalence analysis using Cluster-Lasso after partialing out county and time fixed effects.The extension allows a substantially larger set of potential controls than the original analysis.
- Design: The expanded controls include county-level demographics, income, crime, spending, housing, education, voting, employment, and migration measures.Interactions of initial control values with linear, quadratic, and cubic time terms allow time-varying confounding patterns.
- Results: Using all 978 controls yields imprecise estimates for overall, gun, and non-gun homicide, with potentially inaccurate standard errors.The reported estimates are -.010, .00004, and -.033 for the three outcomes, respectively.
- Results: For overall homicide, Cluster-Lasso estimates .079 (.043), compared with .070 (.035) using baseline controls.For gun homicide, the corresponding estimates are .171 (.047) and .178 (.046), despite substantively different selected controls.
- Results: For non-gun homicide, the Cluster-Lasso estimate is -.019 (.040), compared with -.071 (.038) using baseline controls, and remains statistically inconclusive.The results cannot rule out moderate positive or negative effects at conventional levels.
- Conclusion: Overall, the Cluster-Lasso results show a strong positive effect for firearm homicide, some evidence for overall homicide, and an imprecise non-gun homicide effect.The authors describe these findings as broadly consistent with Cook and Ludwig (2006) despite the richer confounder set.
7. Conclusion
The paper develops clustered Lasso methods for variable selection and inference in high-dimensional panel models with fixed effects. Simulations and a gun-prevalence application support the methods, while causal interpretation remains dependent on the exogeneity of the firearm-suicide-rate measure.
- 7. Conclusion: Clustered Lasso accommodates strong within-individual dependence and fixed effects, supporting variable selection and inference in high-dimensional IV and partially linear panel models.The method uses clustered penalty loadings and partials out individual fixed effects before selection.
- 7. Conclusion: Causal interpretation requires believing that lagged firearm-suicide rates provide an exogenous measure of gun prevalence after controlling for county-level controls and county and time fixed effects.Without that belief, the estimates are not valid causal estimates.
- 7. Conclusion: The proposed methods perform well in a simulation study and produce gun-prevalence estimates broadly consistent with Cook and Ludwig (2006) despite using a broader control set.The empirical application estimates the effect of gun prevalence on crime.
Appendix A. Cluster-Lasso Penalty Loadings Implementation
The appendix specifies feasible penalty-loading choices for Cluster-Lasso and describes the tuning constants used in the empirical and simulation examples.
- Appendix A. Cluster-Lasso Penalty Loadings Implementation: The implementation provides feasible options for setting Cluster-Lasso penalty levels and loadings for j = 1, . . . , p.These options are used to establish asymptotic validity of the proposed algorithm.
- Appendix A. Cluster-Lasso Penalty Loadings Implementation: The examples use c = 1.1, γ = 0.1/ log(p ∨nT), and K = 15 iterations, with Post-Lasso preferred at each step.The notation Lasso/Post-Lasso allows either estimator, although the preferred approach uses Post-Lasso.
Algorithm of Cluster-Lasso penalty loadings
The penalty-loading algorithm initializes loadings, repeatedly updates them using residuals, and yields asymptotically valid loadings under the stated regularity conditions.
- Algorithm of Cluster-Lasso penalty loadings: The algorithm first specifies initial penalty loadings, computes a Lasso/Post-Lasso estimate, and obtains residuals from that estimate.The initial option is based on equations (2.1), (2.2), and (A.22).
- Algorithm of Cluster-Lasso penalty loadings: If K > 1, the procedure refines penalty loadings, updates the Lasso/Post-Lasso estimator, and recomputes residuals.Residuals are recomputed using the updated coefficients.
- Algorithm of Cluster-Lasso penalty loadings: If K > 2, the algorithm repeats the update step K −2 times.This completes the iterative loading procedure.
- Algorithm of Cluster-Lasso penalty loadings: The resulting loadings satisfy ℓφj ⩽ bφj ⩽uφj for every j with probability 1−o(1), with ℓP→1 and u ⩽C < ∞.The bounds establish asymptotic validity of the constructed loadings under the extended regularity condition.
- Algorithm of Cluster-Lasso penalty loadings: The proof accounts for clustering while adapting an argument used for the basic penalty-loading option.The proposition relies on the conditions of Theorem 1 and Condition R’.
A.1. Comments on Condition R’.
Condition R’ governs when the feasible penalty-loading procedure is valid, with requirements depending on the time-series dependence and the relative growth of n and T. Clustered variance estimation can impose stricter growth conditions and reduce variable-selection effectiveness.
- A.1. Comments on Condition R’.: Condition R’ consists of high-level conditions that can be verified using lower-level primitive assumptions.Condition R’(i) generally requires the most stringent conditions.
- A.1. Comments on Condition R’.: When T is fixed or dependence is strong, rates are governed by the cross-sectional dimension, and Condition R’(i) may hold when n grows sufficiently quickly relative to log(p).The result uses moment and boundedness conditions similar to those in related high-dimensional work.
- A.1. Comments on Condition R’.: When T grows with weak dependence, Condition R’(i) is established under strictly stationary strongly mixing data, bounded observed variables, and cross-sectional independence of heterogeneity sequences.The example assumes mixing coefficients satisfying θ(j) ⩽exp{−2cj}.
- A.1. Comments on Condition R’.: The clustered variance estimator requires a stronger condition on T relative to n because its numerator can behave like the variance of a strongly dependent process.The denominator depends only on a weakly dependent process in the discussed setting.
- A.1. Comments on Condition R’.: Without the growth condition, penalty loadings may diverge and prevent selection of variables even when strong predictors exist.This creates a potential variable-selection cost relative to a covariance estimator tailored to weak dependence.
B.1. Proof of Theorem 1.
The proof establishes finite-sample and asymptotic performance bounds for Cluster-Lasso and Post-Cluster-Lasso under the stated penalty, sparsity, and eigenvalue conditions. It proceeds from score and loading control through restricted-eigenvalue arguments, sparsity bounds, and final probability statements.
- Proof structure: The proof first derives Cluster-Lasso bounds and then analyzes Post-Cluster-Lasso using the selected support.The argument is organized into eight steps, with the first four addressing Cluster-Lasso and the next three focusing on Post-Cluster-Lasso before final condition bounds.
- Penalty control: The prescribed penalty level makes the regularization event occur with probability 1 − o(1).This follows from controlling the maximum score and verifying the penalty construction under the stated conditions.
- Cluster-Lasso bounds: The proof obtains prediction and ℓ1-rate bounds using penalty-loadings control and restricted-eigenvalue inequalities.The relevant bounds use the loading restrictions and relate restricted eigenvalues to sparse eigenvalues.
- Post-Cluster-Lasso bounds: Post-Cluster-Lasso bounds follow from projection arguments, sparse-eigenvalue control, and finite-sample results for the selected support.The proof explicitly defines the projection onto selected regressors and derives bounds for prediction, estimation, and sparsity quantities.
- Sparsity control: The proof establishes a sparsity bound for Lasso that supports the subsequent rate results and Post-Lasso analysis.It combines empirical pre-sparsity, maximal sparse-eigenvalue sub-linearity, and a data-driven-penalty sparsity result.
- Final conditions: Under the stated conditions, the auxiliary constants satisfy µ = O(1), 1/κ̄_C = O(1), and µϕ_min(k̄ + s)/κ̄_C = O(1) with probability 1 − o(1).These bounds complete the transition from the intermediate inequalities to the theorem’s stated results.
B.2. Proof of Theorem 2.
The proof of Theorem 2 reduces the estimator to a normalized sum, verifies the required moment and dependence conditions, and applies a central limit theorem. It also establishes the consistency of the associated variance estimator.
- Proof setup: The proof begins from the Post-Lasso rate bounds supplied by Theorem 1 and decomposes the estimator into terms needed for asymptotic analysis.The decomposition is used to isolate the leading stochastic component from remainder terms.
- Asymptotic normality: The normalized leading term is shown to converge in distribution to N(0, 1) as n, T → ∞ jointly and when n → ∞ with T fixed.The argument verifies the conditions of the cited central-limit result for both asymptotic sequences.
- Remainder control: The proof uses moment, conditional-mean, and moderate-deviation conditions to control the score and remainder terms.These controls support the bounds required for the central-limit argument.
- Variance estimation: The final step establishes consistency of the variance estimator by reducing it to several separately controlled terms.The proof applies Cauchy–Schwarz and triangle inequalities and verifies the required component bounds.
Step 4. Bounds
This bounds subsection controls the remaining terms in the proof by combining previously established probability bounds with the model’s SMIV conditions.
- Remaining bounds: The proof shows that A3.1 = A4 = o_P(1) and controls A3.2 using the stated SMIV condition.These bounds address the remaining components needed in the variance and asymptotic arguments.
B.3. Proof of Theorem 3.
The proof of Theorem 3 establishes asymptotic normality and variance-estimator consistency for the partially linear model. It organizes the argument into a leading-term decomposition and bounds for the resulting remainder components.
- Notation and projections: The proof defines stacked residualized data, projection operators, and restricted least-squares coefficients for subsets of selected regressors.These objects provide the notation for analyzing the selected-support estimator and its projection terms.
- Proof structure: The proof proceeds in seven steps: the first establishes asymptotic normality, steps two through six provide supporting bounds, and step seven proves variance-estimator consistency.This structure separates the central-limit argument from the auxiliary rate and remainder controls.
- Asymptotic normality: The desired central limit theorems follow for joint growth of n and T and for n → ∞ with T fixed.The proof obtains these results by applying reasoning analogous to the earlier central-limit argument.
- Support control: The proof uses the Post-Lasso support size and sparse-eigenvalue conditions to control projection and estimation terms.In particular, the selected support size is bounded through Theorem 1 and the relevant sparse eigenvalue is bounded away from zero.
- Variance estimation: Variance-estimator consistency is established by decomposing the estimator into B- and C-terms and bounding each component.The proof uses substitutions, expansions, and inequalities to show the required component orders.