Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
101,235 characters · 13 sections · 2 citation commands
Inference in High Dimensional Panel Models with an Application to Gun Control
\sloppy
The use of panel data is extremely common in empirical economics. Panel data is appealing because it allows researchers to estimate the effects of variables of interest while accounting for time invariant individual specific heterogeneity in a flexible manner. For example, the most widespread model employed in empirical analyses using panel data in economics is the linear fixed effects model which treats individual specific heterogeneity as a set of additive fixed effects to be estimated jointly with other model parameters. This approach is attractive because it allows the researcher to estimate the common slope parameters of the model without imposing any structure over the individual specific heterogeneity.
Many panel data sets also have a large number of time varying variables available for each observation; i.e. they are “high dimensional” data. The large number of available variables may arise because the number of measured characteristics is large. For example, many panel data analyses in economics make use of county, state, or country level panels where there is a large set of measured characteristics and aggregates such as output, employment, demographic characteristics, etc. available for each observation. A large number of time varying variables may also be present due to a researcher wishing to allow for flexible dependence of an outcome variable on a small set of observed time varying covariates and thus considering a variety of transformations and interactions of the underlying set of variables. Identification of effects of interest in panel data contexts is also often achieved through a strategy where identification becomes more plausible as one allows for flexible trends that may differ across treatment states. Allowing for flexible trends that may differ based on observable characteristics may then be desirable but potentially introduces a large number of control variables.
A difficulty in high dimensional settings is that useful predictive models and informative inference about model parameters is complicated by the presence of the large number of explanatory variables. The problem with building a model for prediction can easily be seen when one considers the example of forecasting using a linear regression model in which there are exactly as many linearly independent explanatory variables as there are observations. In this case, the ordinary least squares estimator will fit the data perfectly, returning an $R^2$ of one. However, the estimated model is likely to provide very poor out-of-sample prediction properties because the model estimated by unrestricted least squares is overfit. The least squares fit captures not just the signal about how the predictor variables may be used to forecast the outcome but also perfectly captures the noise in the given sample which is not useful for generating out-of-sample predictions. Constraining the estimated model to avoid perfectly fitting the sample data, or “regularization,” is necessary for building a useful predictive model. Similarly, informative inference about parameters in a linear regression model is clearly impossible if the number of explanatory variables is larger than the sample size if one is unwilling to impose additional model structure.
A useful structure which has been employed in the recent econometrics literature focusing on inference in high dimensional settings is approximate sparsity; see, for example, \citeasnoun{BellChernHans:Gauss}, \citeasnoun{BellChenChernHans:nonGauss}, and \citeasnoun{BelloniChernozhukovHansen2011}. A leading example is the approximately sparse linear regression model which is characterized by having many covariates of which only a small number are important for predicting the outcome.\footnote{There are many statistical methods designed for doing prediction in exactly sparse models in which the relationship between an outcome and a large dimensional set of covariates is perfectly captured by a model using only a small number of non-zero parameters; see \citeasnoun{elements:book} for a review. Approximately sparse models generalize exactly sparse models by allowing for approximation errors from using a low-dimensional approximation.} Approximately sparse models nest conventional parametric regression models as well as standard sieve and series based nonparametric regression models. In addition to nesting standard econometric models, the framework is appealing as it reduces the problem of finding a good predictive model to a variable selection problem. Sensible estimation methods appropriate for this framework also yield models with a relatively small set of variables which aids interpretability of the results and corresponds to the usual approach taken in empirical economics where models are typically estimated using a small number of control variables.
There are a variety of sensible variable selection estimators that are appropriate for estimating approximately sparse models. For example, $\ell_1$-penalized methods such as the Lasso estimator of \citeasnoun{FF:1993} and \citeasnoun{T1996} have been proposed for model selection problems in high dimensional least squares problems in part because they are computationally efficient. Many $\ell_1$-penalized methods and related methods have been shown to have good estimation properties with i.i.d. data even when perfect variable selection is not feasible; see, e.g., \citeasnoun{CandesTao2007}, \citeasnoun{MY2007}, \citeasnoun{BickelRitovTsybakov2009}, \citeasnoun{horowitz:lasso}, \citeasnoun{BC-PostLASSO} and the references therein. Such methods have also been shown to extend to nonparametric and non-Gaussian cases as in \citeasnoun{BickelRitovTsybakov2009} and \citeasnoun{BellChenChernHans:nonGauss}, the latter of which also allows for conditional heteroscedasticity.
While the models and methods mentioned above are useful in a variety of contexts, they do not immediately apply to standard panel data models. There are two key points of departure between conventional approximately sparse high dimensional models and conventional panel data models used in empirical economics. The first is that the approximately sparse framework seems highly inappropriate for usual beliefs about individual specific heterogeneity in fixed effects models.\footnote{There are other natural alternatives to dimension reduction over individual specific heterogeneity. For example, one could assume that individual specific heterogeneity is drawn from some common distribution as in conventional random effects approaches. This distribution could also be allowed to be flexibly specified as in \citeasnoun{altonji:matzkin} or \citeasnoun{bester:hansen:index}. Alternatively, one could consider a grouped structure over heterogeneity as in \citeasnoun{bester:hansen:group} or \citeasnoun{bonhomme:manresa}.} Specifically, the approximately sparse structure would imply that individual specific heterogeneity differs from some constant level for only a small number of individuals and may be completely ignored for the vast majority of individuals. A seemingly more plausible model for individual specific heterogeneity is one in which it is allowed to be relevant for every individual, related to observed variables in an unrestricted manner, and different for each individual. Under these beliefs, naively applying a method designed for approximately sparse models may result in a procedure with poor estimation and inference properties.
The second key difference is that the assumption of independent observations is inappropriate for many panel data sets used in economics. Many economic panels appear to exhibit substantial correlation between observations within the same cross-sectional unit of observation. It is well-known that failing to account for this correlation when doing inference about model parameters in panel data with a small number of covariates may lead to tests with substantial size distortions. This concern has led to the routine use of “clustered standard errors” which are robust to within-individual correlation and heterogeneity across individuals in empirical research.\footnote{See \citeasnoun{arellano:feinf}, \citeasnoun{bdm:cluster}, and \citeasnoun{hansen:cluster} among others.} In the context of variable selection in high dimensional models, failing to account for this correlation may result in substantial understatement of sampling variability. This understatement of sampling variability may then lead to a variable selection device selecting too many variables, many of which have no true association to the outcome of interest. The presence of these spuriously selected variables may have a substantial negative impact on the resulting estimator as the spuriously selected variables are, by construction, the most strongly correlated to the noise within the sample.
A key contribution of this paper is offering a variant of the Lasso estimator that accommodates a clustered covariance structure (Cluster-Lasso). We provide formal conditions under which the estimator performs well in the sense of returning a sparse estimate and having good forecasting and rate of convergence properties. By providing results allowing for a clustered error structure, we are also able to allow for the presence of unrestricted additive individual specific heterogeneity which are treated as fixed effects that are partialed out of the model before variable selection occurs. Accommodating this structure requires partialing out a number of covariates that is proportional to the sample size under some asymptotic sequences we consider. In general, partialing out a number of variables proportional to the sample size will induce a non-standard, potentially highly dependent covariance structure in the partialed-out data. The structure of the fixed effects model is such that partialing out the fixed effects cannot induce correlation across individuals, though it may induce strong correlation within the observations for each individual. Because this structure is already allowed for in the clustered covariance structure, partialing out the fixed effects poses no additional burden after allowing for clustering.
The second contribution of this paper is taking the derived performance bounds for the proposed Lasso variant and using them to provide methods for doing valid inference following variable selection in two canonical models with high dimensional components: the linear instrumental variables (IV) regression with high dimensional instruments and additive fixed effects and the partially linear treatment model with high dimensional controls and additive fixed effects. Inference in these settings is complicated due to the fact that variable selection procedures inevitably make model selection mistakes which may result in invalid inference following model selection; see \citeasnoun{potscher} and \citeasnoun{leeb:potscher:pms} for examples. It is thus important to offer procedures that are robust to such model selection mistakes. To address this concern, we follow the approach of \citeasnoun{BellChenChernHans:nonGauss} in the IV model and \citeasnoun{BelloniChernozhukovHansen2011} in the partially linear model making use of the Cluster-Lasso to accommodate within-individual dependence and partialing out of fixed effects. We show that standard inference following these procedures results in inference about model parameters of interest that is uniformly valid within a large class of approximately sparse regression models as long as a clustered covariance estimator is used in estimating the parameters' asymptotic variance. The results of this paper thus allow valid inference about a prespecified, fixed set of model parameters of interest in canonical panel data models with additive fixed effects in the realistic scenario where a researcher is unsure about the exact identities of the relevant set of variables to be included in addition to the variables of interest and the fixed effects.
In addition to theoretical guarantees, we illustrate the performance of the proposed methods through simulation examples. In the simulations, we consider a fixed effects IV model and a conventional linear fixed effects model. The IV model has a single endogenous variable whose coefficient we would like to infer and a large number of instruments which satisfy the IV exclusion restriction only after eliminating the fixed effects. Given the large number of instruments, we use the Cluster-Lasso to select a small set of instruments to use in an IV estimator as in \citeasnoun{BellChenChernHans:nonGauss}. In the linear model, we have a single treatment variable of interest that is related to a set of fixed effects and additionally have a large number of potential confounding variables. We estimate the effect of the variable of interest using the double selection procedure of \citeasnoun{BelloniChernozhukovHansen2011} with the Cluster-Lasso used as the variable selector. In both cases, we find that the Cluster-Lasso-based procedures perform well in terms of both estimation risk and inference properties as measured by size of tests. The most interesting feature of the simulation results is that the Cluster-Lasso-based procedures perform markedly better than variable selection procedures that do not allow for clustering. This difference in performance suggests that additional modifications of Lasso-type procedures to account for other dependence structures may be worthwhile.
We also provide results from an empirical example that looks at estimating the effect of guns on crime using data from a panel of U.S. counties following the analysis of \citeasnoun{cook:ludwig:guns}. In their analysis, \citeasnoun{cook:ludwig:guns} use a conventional linear fixed effects model with a small number of time varying county level control variables and find a positive and statistically significant effect of their proxy of gun prevalence on the overall homicide rate and the gun homicide rate and a negative but insignificant effect on the non-gun homicide rate. We extend this analysis by considering a broad set of county-level demographic characteristics and a set of flexible trends that are interacted with baseline county-level characteristics. We then use the methods developed in this paper to select a small, data-dependent set of variables that it is important to control for if one wants to hold fixed important sources of confounding variation. Interestingly, our findings are largely consistent with those of \citeasnoun{cook:ludwig:guns} despite allowing for a much richer set of conditioning information. Specifically, we find a strong positive relationship between gun prevalence and gun homicides, a small and statistically insignificant relationship between gun prevalence and non-gun homicides, and a borderline significant positive effect of gun prevalence on the overall homicide rate.
There are many regularization or dimension reduction techniques available in the statistics and econometrics literature. An appealing method for estimating sparse high dimensional linear models is the Lasso. Lasso estimates regression coefficients by minimizing a least squares objective plus an $\ell_1$ penalty term. We begin with an informal discussion of Lasso in linear models with fixed effects before proceeding with more precise specifications and modeling assumptions. In particular, our goal for this section is to outline a Lasso procedure for estimating the model $$y_{it} = x_{it}'\beta + \alpha_i + \epsilon_{it}, \quad i =1,..., n, \quad t = 1,..., T,$$ where $y_{it}$ is an outcome of interest, $x_{it}$ are covariates, $\alpha_i$ are individual specific effects, and $\epsilon_{it}$ is an idiosyncratic disturbance term which is mean zero conditional on covariates but may have dependence within an individual.\footnote{We abstract from issues arising from unbalanced panels for notational convenience but note that the arguments go through immediately provided that the missing observations are missing at random.}
The first step in our estimation strategy is to eliminate the fixed effect parameters. For simplicity, we will always consider removing the fixed effects by within individual demeaning but note that removing the fixed effects using other differencing methods could be accommodated using similar arguments. To this end, we define$$\ddot x_{it} = x_{it} - \frac{1}{T}\sum_{t=1}^T x_{it}.$$ We define the quantities $\ddot y_{it}$ and $\ddot \epsilon_{it}$ similarly and note that the double dot notation will signify deviations from within individual means throughout the paper. Eliminating the fixed effects by substracting individual specific means leads to the “within model”:
$$\ddot y_{it} = \ddot x_{it}' \beta + \ddot \epsilon_{it}.$$
The Cluster-Lasso coefficient estimate $\widehat \beta_L$ is defined by the solution to the following penalized minimization problem on the within model:
Solving the problem ((ref)) requires two user-specified tuning parameters: the main penalty level, $\lambda$, and covariate specific penalty loadings, $\{\widehat \phi_j\}_{j=1}^{p}$. The main penalty parameter dictates the amount of regularization in the Lasso procedure and serves to balance overfitting and bias concerns. The covariate specific penalty loadings $\{\widehat \phi_j\}_{j=1}^{p}$ are introduced to allow us to handle data which may be dependent within individual, heteroscedastic, and non-Gaussian. We provide further discussion of the specific choices of penalty parameters in the next subsection. Correct choice of these penalty parameters is particularly important when Lasso is used and the ultimate goal is inference about parameters of a “structural” model.
We will also make use of a post model selection estimator; see for example \citeasnoun{BC-PostLASSO}. The Post-Cluster-Lasso estimator is defined with respect to the variables selected by Cluster-Lasso: $\widehat I = \{ j : \widehat \beta_{Lj} \neq 0\}$. The Post-Cluster-Lasso estimator is simply the least squares estimator subject to the constraint that covariates not selected in the initial Cluster-Lasso regression must have zero coefficients:
As discussed in Section (ref), the selected model $\widehat I$ has good properties under regularity conditions and approximate sparsity of the coefficient $\beta$. Just as in \citeasnoun{BC-PostLASSO}, good properties of the selected set of variables will then translate into good properties for the Post-Cluster-Lasso estimator.
An important condition used in proving favorable performance of Cluster-Lasso and inference following Cluster-Lasso-based model selection is the use of penalty loadings and penalty parameters that dominate the score vector in the sense that
for some constant slack parameter $c>1$. \citeasnoun{BelloniChernozhukovHansen2011} refer to condition ((ref)) as the “regularization event”. Note that the term $\frac{1}{nT}\sum_{i=1}^n \sum_{t=1}^T \ddot x_{itj} \ddot \epsilon_{it}$ intuitively captures the sampling variability in learning about coefficient $\beta_j$. The regularization event thus corresponds to selecting penalty parameters large enough to dominate the noise in estimating model coefficients. Looking at the structure of the Lasso optimization problem, ((ref)), we can see that ((ref)) leads to settings all coefficients whose magnitude is not big enough relative to sampling noise exactly to zero in the Lasso solution. This property makes Lasso-based methods appealing for forecasting and variable selection in sparse models where many of the model parameters can be taken to be zero and it is desirable to exclude any variables from the model that cannot reliably be determined to have strong predictive power.
Given the importance of event ((ref)) in verifying desirable properties of Lasso-type estimators, it is key that penalty loadings and the penalty level are chosen so that ((ref)) occurs with high probability. The intuition for suitable choices can be seen by considering $\widehat \phi_j = \phi_j$ where $$\phi_j^2 = \frac{1}{nT} \sum_{i=1}^n \left ( \sum_{t=1}^T \ddot x_{itj} { \ddot \epsilon}_{it} \right)^2 = \frac{1}{nT} \sum_{i=1}^n \sum_{t=1}^T \sum_{t' = 1}^T \ddot x_{itj} \ddot x_{it'j} { \ddot \epsilon}_{it} { \ddot \epsilon}_{it'}.$$ Note that the quantity $\phi_j^2$ is a natural measure for the noise in estimating $\beta_j$ that allows for arbitrary within-individual dependence. With these loadings, we can apply the moderate deviation theorems for self-normalized sums due to \citeasnoun{jing:etal} to conclude that $$\frac{P( \phi_j^{-1} \frac{1}{\sqrt{nT}} \sum_{i=1}^n \sum_{t=1}^T \ddot x_{itj} \ddot \epsilon_{it}>m)}{P(N(0,1)>m)} = o(1), \quad \text{uniformly in } |m| = o(n^{1/6}), \ j \in {1,...,p}. $$ It follows from this result and the union bound that setting $\lambda$ large enough to dominate $p$ standard Gaussian random variables with high-probability, specifically as in ((ref)), will implement condition ((ref)) using $\widehat \phi_j = \phi_j$.
This form of loadings is an extension of the loadings considered in \citeasnoun{BellChenChernHans:nonGauss} which apply in settings with non-Gaussian and heteroscedastic but independent data to the present setting where we need to accommodate strong within-individual dependence. Formally verifying the validity of this approach requires control of the tail behavior of each of the sums $\phi_j^{-1} \frac{1}{{nT}} \sum_{i=1}^n \sum_{t=1}^T \ddot x_{itj} \ddot \epsilon_{it}$ uniformly over $j \leqslant p$. \citeasnoun{BellChenChernHans:nonGauss} verify the appropriate uniform tail control under independence across observations using loadings suitable for independent observations and $\lambda = 2 c \sqrt{nT} \Phi^{-1}(1-o(1)/2p)$ by using results from the theory of moderate deviations of self-normalized sums from \citeasnoun{jing:etal}. We extend these results to allow for a clustered dependence structure in the observations which is appropriate for panel data and well-suited to fixed effects applications.
In practice, the values $\{\phi_j\}_{j=1}^{p}$ are infeasible since they depend on the unobservable $\ddot \epsilon_{it}$. To make estimation feasible, we use preliminary estimates of $\ddot \epsilon_{it}$, denoted $\widehat {{\epsilon}}_{it}$, in forming feasible loadings:
The $\widehat{ \epsilon}_{it}$ can be calculated through an iterative algorithm given in Appendix (ref) which follows the algorithm given in \citeasnoun{BellChenChernHans:nonGauss} and \citeasnoun{BelloniChernozhukovHansen2011}. We define the Feasible Cluster-Lasso and Feasible Post-Cluster-Lasso estimates as the Cluster-Lasso and Post-Cluster-Lasso estimates using the feasible penalty loadings. A key property of the feasible penalty loadings needed for validity of the approach is that
Under this condition and setting
with $\gamma = o(1)$, the regularization event ((ref)) holds with probability tending to one.
It is worth noting that failure to use the clustered penalty loadings defined in ((ref)) (or their infeasible version) can lead to an inflated probability of failure of the regularization event. When this event fails to hold, covariates which are only spuriously related with the outcome have a non-negligable chance of entering the selected model. In the simulation experiments provided in Section (ref), we demonstrate how inclusion of such variables can be problematic for post-model-selection inference essentially due to their introducing a type of endogeneity bias.\footnote{A variable which is spuriously selected must have non-negligible correlation to the errors within sample which results in similar behavior as when an endogenous variable is included in a regression.}
This section gives conditions under which Cluster-Lasso and Cluster-Post-Lasso attains favorable performance bounds. These bounds are useful in their own right and are important elements in establishing the properties of inference following Lasso variable selection discussed in Section (ref). In establishing our formal results, we consider the additive fixed effects model
where $e_i$ represents time invariant individual specific heterogeneity that is allowed to depend on $w_{i} = \{w_{it}\}_{t=1}^{T}$ in an unrestricted manner. Throughout, we will assume that $\{y_{it},w_{it}\}_{t=1}^{T}$ are i.i.d. across $i$ but do not restrict the within individual dependence.\footnote{Note that we impose that data are i.i.d. across $i$ for notational convenience. We could allow for data that are i.n.i.d. across $i$ at the cost of complicating the notation and statement of the regularity conditions.} Our results will hold under $n \rightarrow \infty$, $T$ fixed asymptotics and $n \rightarrow \infty$, $T \rightarrow \infty$ joint asymptotics.\footnote{Because we want to accommodate both $n \rightarrow \infty$, $T$ fixed asymptotics and $n \rightarrow \infty$, $T \rightarrow \infty$ joint asymptotics in a simple and unified manner, we maintain strict exogeneity of $w_i$ throughout and do not consider time effects. We note that time effects may be included easily under $n \rightarrow \infty$, $T$ fixed asymptotics but that some modification of the formal results would be needed under $T \rightarrow \infty$ sequences.}
A key distinction between the analysis in this paper and previous work on Lasso is allowing for within-individual dependence. To aid discussion of this feature, we let $$ \imath_T := T \min_{1 \leqslant j \leqslant p }\frac{ {\mathrm{E}} [ \frac{1}{T} \sum_{t=1}^T \ddot x^2_{itj} \ddot \epsilon^2_{it}]}{ {\mathrm{E}}[ \frac{1}{T} (\sum_{t=1}^T \ddot x_{itj} \ddot \epsilon_{it})^2] } = T \min_{1 \leqslant j \leqslant p} \frac{{\mathrm{E}} [ \frac{1}{T} \sum_{t=1}^T \ddot x^2_{itj} \ddot \epsilon^2_{it}] }{{\mathrm{E}} [\phi^2_j]} $$ be the index of information induced by the “time" or “within-group" dimension. This time information index, $\imath_T$, is inversely related to the strength of within-individual dependence and can vary between two extreme cases:
There are many interesting cases between these extremes. A leading case is where $\imath_T \propto T$ which occurs when there is weak dependence within clusters and results in clustering only affecting the constants in the Lasso performance bounds. The case where $\imath_T \propto T^{a}$ for some $0 \leqslant a < 1$ corresponds to stronger forms of dependence within clusters which could be generated, for example, by fractionally integrated data. Our results will allow for the two extreme cases as well as those falling between these two extremes.
We begin the presentation of formal conditions by defining approximately sparse models. Note that the model $f$ as well as the set of covariates $w_{it}$ may depend on the sample size, but we suppress this dependence for notational convenience.
Condition ASM. (Approximately Sparse Model). The function $f(w_{it})$ is well-approximated by a linear combination of a dictionary of transformations, $x_{it} = X_{nT}(w_{it})$, where $x_{it}$ is a $p \times 1$ vector with $p \gg n$ allowed, and $X_{nT}$ is a measurable map. That is, for each $i$ and $t$ $$ f(w_{it}) = x_{it}' \beta + r(w_{it}),$$ where the coefficient $\beta$ and the remainder term $r(w_{it})$ satisfy $$\|\beta\|_0 \leqslant s = o(n\imath_T) \ \ \ and \ \ \ \left [\frac{1}{nT}\sum_{i=1}^n \sum_{t=1}^T r(w_{it})^2 \right ]^{1/2} \leqslant A_{s} = O_{\mathrm{P}}( \sqrt{s/n\imath_T}).$$
We note that the approximation error $r(w_{it})$ is restricted to be of the same order as or smaller than sampling uncertainty in $\beta$ provided that the true model were known. Because we will mainly be concerned with the within model, we note that it is straightforward to show that $\ddot y_{it} = \ddot f(w_{it}) + \ddot \epsilon_{it}$ satisfies Condition ASM when the original model does.
The next assumption controls the behavior to the empirical Gram matrix. Let $\ddot M$ be the $p \times p$ matrix of the sample covariances between the variables $\ddot x_{itj}$. Thus, $$ \ddot M = \{M_{jk}\}_{j,k=1}^{p}, \quad M_{jk} = \frac{1}{nT} \sum_{i=1}^n \sum_{t=1}^T \ddot x_{itj} \ddot x_{itk}. $$ In standard regression analysis where the number of covariates is small relative to the sample size, a conventional assumption used in establishing desirable properties of conventional estimators of $\beta$ is that $\ddot M$ has full rank. In the high dimensional setting, $\ddot M$ will be singular if $p \geqslant n$ and may have an ill-behaved inverse even when $p < n$. However, good performance of the Lasso estimator only requires good behavior of certain moduli of continuity of $\ddot M$. There are multiple formalizations and moduli of continuity that can be considered in establishing the good performance of Lasso; see \citeasnoun{BickelRitovTsybakov2009}. We focus our analysis on a simple eigenvalue condition that is suitable for most econometric applications. It controls the minimal and maximal $m$-sparse eigenvalues of $\ddot M$ defined as
where $$ \Delta(m) = \{ \delta \in \mathbb{R}^p: \|\delta\|_0\leqslant m, \|\delta\|_2 = 1\}, $$ is the $m$-sparse subset of a unit sphere.
In our formal development, we will make use of the following simple sufficient condition:
Condition SE. (Sparse Eigenvalues) For any $C>0$, there exist constants $0< \kappa' < \kappa'' < \infty$, which do not depend on $n$ but may depend on $C$, such that with probability approaching one, as $n \to \infty$, $\kappa' \leqslant \varphi_{{\rm min}}(Cs)(\ddot M) \leqslant \varphi_{{\rm max}}(Cs)(\ddot M) \leqslant \kappa''$.
Condition SE requires only that certain “small" $Cs \times Cs$ submatrices of the large $p \times p$ empirical Gram matrix are well-behaved. This condition seems reasonable and will be sufficient for the results that follow. Note that we prefer to write the eigenvalue conditions in terms of the demeaned covariates as it is straightforward to show that the conditions continue to hold under data generating processes where the covariates have nonzero within-individual variation. Condition SE could be shown to hold under more primitive conditions by adapting arguments found in \citeasnoun{BC-PostLASSO} which build upon results in \citeasnoun{ZhangHuang2006} and \citeasnoun{RudelsonVershynin2008}; see also \citeasnoun{RudelsonZhou2011}.
The final condition collects various rate and moment restrictions. Again, the conditions are expressed in terms of demeaned quantities for convenience. Also consider the following third moment: $$ \varpi_j = \left({\mathrm{E}}\left [ \left | \frac{1}{\sqrt{T}} \sum_{t=1}^T \ddot x_{itj} \ddot \epsilon_{it}\right |^3 \right]\right)^{1/3}. $$
Condition R. (Regularity Conditions) Assume that for data $\{y_{it},w_{it}\}$ that are i.i.d. across $i$, the following conditions hold with $x_{it}$ defined as in Condition ASM with probability $1-o(1)$:
(i) $\left(\frac{1}{T} \sum_{t=1}^ T {\mathrm{E}}[\ddot x_{itj}^2 \ddot \epsilon_{it}^2]\right) + \left(\frac{1}{T} \sum_{t=1}^ T {\mathrm{E}}[\ddot x_{itj}^2 \ddot \epsilon_{it}^2]\right)^{-1} = O(1)$,
(ii) $1 \leqslant \max_{1 \leqslant j \leqslant p } \phi_j/\min_{1 \leqslant j \leqslant p} \phi_j = O(1)$,
(iii) $1 \leqslant \max_{1 \leqslant j \leqslant p} \varpi_j/\sqrt{{\mathrm{E}} \phi_j^2} = O(1),$
(iv) $\log^3 (p) = o(nT) $ and $\ s \log (p\vee nT) = o(n\imath_T)$,
(v) $\max_{1 \leqslant j \leqslant p}|\phi_j - \sqrt{{\mathrm{E}} \phi_j^2}|/\sqrt{{\mathrm{E}} \phi_j^2} = o(1)$.
This condition is sufficient under the high-level assumption ((ref)) on the availability of the valid feasible data loadings. In the appendix we provide additional conditions under which we exhibit validity of data-dependent loadings constructed using an iterative algorithm of the type proposed in \citeasnoun{BellChenChernHans:nonGauss}.
With the above conditions in place, we can state the asymptotic performance bounds of Cluster-Lasso.
Theorem (ref) shows that Cluster-Lasso and Post-Cluster-Lasso continue to have good model selection and prediction properties allowing for a clustered dependence structure when penalty loadings that account for this dependence are used. Establishing these results is important as applied researchers in economics typically assume the data have a clustered dependence structure and because it allows us to accommodate partialing out a large number of fixed effects which in general will induce a clustered error structure in the demeaned data. The simulation results in Section (ref) also illustrate the importance of allowing for clustering not just in calculating standard errors but also in forming penalty loadings when selecting variables using Lasso, showing that inference about coefficients of interest following variable selection using Lasso may be poor when this dependence is ignored when forming penalty loadings.
The bounds derived in Section (ref) allow us to derive the properties of inference methods following variable selection with the Cluster-Lasso. In this section, we use these results to provide two different applications of using Cluster-Lasso to select variables for use in causal inference.
Instrumental variables techniques are widely used in applied economic research. While these methods give an important tool for calculating structural effects, they are often imprecise. One way to improve the precision of instrumental variables estimators is to use many instruments or to try to approximate the optimal instruments as in \citeasnoun{amemiya:optimalIV}, \citeasnoun{chamberlain}, and \citeasnoun{newey:optimaliv}.
In this section, we follow \citeasnoun{BellChenChernHans:nonGauss} who consider using Post-Lasso to estimate optimal instruments. Using Lasso-based methods to form first-stage predictions in IV estimation provides a practical approach to obtaining the efficiency gains from using optimal instruments while dampening the problems associated with many instruments. We prove that Cluster-Lasso-based procedures produce first-stage predictions that provide good approximations to the optimal instruments when controlling for individual heterogeneity through fixed effects. We consider the following model:
where ${\mathrm{E}}[\epsilon_{it} u_{it}] \ne 0$ but ${\mathrm{E}}[\epsilon_{it}|w_{i1},...,w_{iT}] = {\mathrm{E}}[u_{it}|w_{i1},...,w_{iT}] = 0$.\footnote{The extension to $d_{it}$ an $r \times 1$ vector with $r \ll nT$ fixed is straightforward and omitted for convenience. The results also carry over immediately to the case with a small number of included exogenous variables $y_{it} = \alpha d_{it} + x_{it}'\beta + e_i + \epsilon_{it}$ where $x_{it}$ is a $k \times 1$ vector with $k \ll nT$ fixed that will be partialed out with the fixed effects.}
We consider estimation of the parameter of interest $\alpha$, the coefficient on the endogenous regressor, using Cluster-Lasso to select instruments. We assume that the first-stage follows an approximately sparse model with $h(w_{it}) = z_{it}'\pi + r(w_{it})$ where we let $z_{it} = z(w_{it})$ denote a dictionary of transformations of underlying instrument $w_{it}$ and $\pi$ be a sparse coefficient as in Condition ASM. After eliminating the fixed effect terms through demeaning, the model reduces to
where we set $D_{it}=h(w_{it})$ for notational convenience. By Theorem 1, the Cluster-Lasso estimate of the coefficients on $\ddot z_{it}$ when we use $\ddot z_{it}$ to predict $\ddot d_{it}$, $\widehat \pi$, will be sparse with high probability. Letting $\widehat I_\pi = \{ j : \widehat \pi_j \neq 0 \}$, the Cluster-Lasso-based estimator of $\alpha$ may be calculated by standard two stage least squares using only the instruments selected by Cluster-Lasso: $\ddot z_{it \widehat I_\pi} := \left(\ddot z_{itj} \right)_{ j \in \widehat I_\pi}$. That is, we define the Post-Cluster-Lasso IV estimator for $\alpha$ as
$\widehat {D_{it}}$ is the fitted value from the regression of $\ddot d_{it}$ on $\left(\ddot z_{itj} \right)_{ j \in \widehat I_\pi}$,\footnote{That is, $\widehat D_{it}$ is the Post-Cluster-Lasso forecast of $\ddot d_{it}$ using $\ddot z_{it}$ as predictors.} and $\widehat \epsilon_{it} = \ddot y_{it} - \widehat \alpha \ddot d_{it}$. We then define an estimator of the asymptotic variance of $\widehat \alpha$, which will be used to perform inference for the parameter $\alpha$ after proper rescaling, as
Scaled appropriately, the estimate $\widehat V$ will be close to the quantity $$V = \frac{\imath_T^D}{T} Q^{-1} \Omega Q^{-1}, \ \ \text{with probability $1-o(1)$} $$ where $$Q={\mathrm{E}}[ \frac{1}{T} \sum_{t=1}^T \ddot D_{it}^2], \ \ \ \Omega = \frac{1}{nT} \sum_{i=1}^n \sum_{t=1}^T \sum_{t'=1}^T {\mathrm{E}}[ \ddot D_{it} \ddot D_{it'} \ddot \epsilon_{it} \ddot \epsilon_{it'}].$$
Finally, it is convenient to define the following quantities that are useful in discussing formal conditions for our estimation procedure. We define appropriate moments and information indices analogous to those used to derive properties of Cluster-Lasso and Post-Cluster-Lasso. For any arbitrary random variables, $A=\{A_{it}\}_{i\leqslant n,t\leqslant T}$, define
$$ \phi^2(A) = \frac{1}{n}\sum_{i=1}^n \left( \frac{1}{\sqrt T} \sum_{t=1}^T A_{it} \right)^2, \ \ \varpi (A)= {\mathrm{E}}\left [ \left | \frac{1}{\sqrt{T}} \sum_{t=1}^T A_{it}\right |^3 \right]^{1/3}, \ \ \imath_T(A) = T \frac{{\mathrm{E}} \left[\frac{1}{T}\sum_{t=1}^T A_{it}^2 \right] } {{\mathrm{E}} \left[ \phi^2(A) \right]}. $$ For use in the instrumental variables estimation, we let
To derive asymptotic properties of these estimators, we will make use of the following condition in addition to those assumed in Section (ref).
Condition SMIV
(i) Sufficient conditions for Post-Cluster-Lasso: ASM, SE, R hold for model (ref).
(ii) Sufficient conditions for asymptotic normality of $\widehat \alpha$ and consistency of $\frac{\imath_T^D}{T}\widehat V$:
(a) ${\mathrm{E}}\left[ \frac{1}{T} \sum_{t=1}^T \ddot D_{it}^2\right],$ $ {\mathrm{E}}\left[\frac{1}{T} \sum_{t=1}^T \ddot \epsilon_{it}^2 \ddot D_{it}^2\right]$, ${\mathrm{E}} \left[ \left( \frac{1}{T} \sum_{t=1}^T \ddot d_{it}^2 \right)^2 \right] $ are bounded uniformly from above and away from zero, uniformly in $n, T.$ Additionally, ${\mathrm{E}}\left[ \left( \frac{1}{T}\sum_{t=1}^T \ddot \epsilon_{it}^2 \right)^q\right] = O(1)$ for some $q>4$.
(b) $ \varpi_D / \sqrt{ {\mathrm{E}} \phi_D^2 } = O(1)$, $\max_{1 \leqslant j \leqslant p} \varpi_{z_j\epsilon}/\sqrt{{\mathrm{E}} \phi_{z_j \epsilon}^2} = O(1),$
(c) $\max_j \frac{\imath_T^{z_j \epsilon}}{T} \phi_{z_j \epsilon}^2 = O_P(1) $, $\frac{1}{T} \phi_{dD}^2 = O_{\mathrm{P}}(1),$ $\max_j \frac{1}{T} \phi_{z_j d}^2 = O_P(1)$
(d) $\frac{s^2 \log^2 (p\vee nT) }{n\imath_T} \max \{ 1, \max_{1 \leqslant j \leqslant p} \frac{\imath_T^D}{\imath_T^{z_j \epsilon}} \} =o(1)$ and $ \frac{\imath_T^D}{\imath_T}n^{2/q}\frac{s \log(p \vee nT)}{n}=o(1)$
The conditions assumed in Condition SMIV are fairly standard. Outside of moment conditions, the main restriction in Condition SMIV is condition (ii)(a) that guarantees that the parameter $\alpha$ would be strongly identified if $\ddot D_{it}$ could be observed. Coupled with the approximately sparse model, this condition implies that using a small number of the variables in $z_{it}$ is sufficient to strongly identify $\alpha$ which rules out the case of weak-instruments as in \citeasnoun{ss:weakiv} and many-weak-instruments as in \citeasnoun{NeweyEtAl-JIVE}.\footnote{See also \citeasnoun{RJIVE} who consider many-weak-instruments in a $p > n$ setting.}
With the model and conditions in place, we provide the following results which can be used to perform inference about the structural parameter $\alpha$.
This theorem verifies that the IV estimator formed with instruments selected by Cluster-Lasso in a linear IV model with fixed effects is consistent and asymptotically normal. In addition, one can use the result with $\widehat V$ defined in ((ref)), which is simply the usual clustered standard error estimator arellano:feinf, to perform valid inference for $\alpha$ following instrument selection. Note that this inference will be valid uniformly over a large class of data generating processes which includes cases where perfect instrument selection is impossible.
A second strategy for identifying structural effects in economic research is based on assuming that variables of interest are as good as randomly assigned conditional on time varying observables and time invariant fixed effects. Since this approach relies on including the right set of time varying observables, a practical problem researchers face is the choice of which control variables to include in the model. The high dimensional framework provides a convenient setting for exploring data-dependent selection of control variables. In this section, we consider the problem of selecting a set of variables to include in a linear model from a large set of possible control variables in the presence of unrestricted individual specific heterogeneity.
The structure of the Lasso optimization problem ensures that any estimated coefficient that is not set to zero can be reliably differentiated from zero relative to estimation noise when ((ref)) holds while any coefficient that can not be distinguished reliably from zero will be estimated to be exactly zero. This property makes Lasso-based methods appealing for variable selection in sparse models. However, this property also complicates inference after model selection in approximately sparse models which may have a set of variables with small but non-zero coefficients in addition to strong predictors. In this case, satisfaction of condition ((ref)) will result in excluding variables with small but non-zero coefficients which may lead to non-negligible omitted variables bias and irregular sampling behavior of estimates of parameters of interest. This intuition is formally developed in \citeasnoun{potscher} and \citeasnoun{leeb:potscher:pms}. Offering solutions to this problem with fully independent data is the focus of a number of recent papers; see, for example, \citeasnoun{BellChernHans:Gauss}; \citeasnoun{BellChenChernHans:nonGauss}; \citeasnoun{ZhangZhang:CI}; \citeasnoun{BCH2011:InferenceGauss}; \citeasnoun{BelloniChernozhukovHansen2011}; \citeasnoun{vdGBRD:AsymptoticConfidenceSets}; \citeasnoun{JM:ConfidenceIntervals}; and \citeasnoun{BCFH:Policy}.\footnote{These citations are ordered by date of first appearance on arXiv.} In this section, we focus on extending the approach of \citeasnoun{BelloniChernozhukovHansen2011} to the panel setting with dependence within individuals.
To be precise, we consider estimation of the parameter $\alpha$ in the partially linear additive fixed effects panel model:
where $y_{it}$ is the outcome variable, $d_{it}$ is the policy/treatment variable whose impact $\alpha$ we would like to infer,\footnote{The analysis extends easily to the case where $d_{it}$ is an $r \times 1$ vector where $r$ is fixed and is omitted for convenience.} $z_{it}$ represents confounding factors on which we need to condition, $e_i$ and $f_i$ are fixed effects which are invariant across time, and $\zeta_{it}$ and $u_{it}$ are disturbances that are independent of each other. Data are assumed independent across $i$, and dependence over time within individual is largely unrestricted.
The confounding factors $z_{it}$ affect the policy variable via the function $m(z_{it})$ and the outcome variable via the function $g(z_{it})$. Both of these functions are unknown and potentially complicated. We use linear combinations of control terms $x_{it} = P(z_{it})$ to approximate $g(z_{it})$ and $m(z_{it})$ with $ x_{it}'\beta_{g}$ and $x_{it}'\beta_{m}$. In order to allow for a flexible specification and incorporation of pertinent confounding factors, we allow the dimension, $p$, of the vector of controls, $x_{it} = P(z_{it})$, to be large relative to the sample size.\footnote{High dimensional $x_{it}$ typically occurs in either of two ways. First, the baseline set of conditioning variables itself may be large so $x_{it} = z_{it}$. Second, $z_{it}$ may be low-dimensional, but one may wish to entertain many nonlinear transformations of $z_{it}$ in forming $x_{it}$. In the second case, one might prefer to refer to $z_{it}$ as the controls and $x_{it}$ as something else, such as technical regressors. For simplicity of exposition and as the formal development in the paper is agnostic about the source of high dimensionality, we call the variables in $x_{it}$ controls or control variables in either case.} Upon substituting these approximation into ((ref)) and ((ref)), we are essentially left with a conventional linear fixed effects model with a high dimensional set of potential confounding variables:
where $r_g(z_{it})$ and $r_m(z_{it})$ are approximation errors. The fixed effects can again be eliminated by subtracting within group means yielding
Informative inference about $\alpha$ is not possible in this model without imposing further structure since we allow for $p > n$ elements in $x_{it}$. The additional structure is added by assuming that condition ASM applies to both $g(z_{it})$ and $m(z_{it})$ which implies that exogeneity of $d_{it}$ may be taken as given once one controls linearly for a relatively small number, $s < n$, of the variables in $x_{it}$ whose identities are a priori unknown.\footnote{Note that this condition is stronger than necessary but convenient; see, e.g., \citeasnoun{farrell:JMP}. We briefly explore this in the simulation example in Section (ref) where we consider a design where Condition ASM holds only in one equation and show that our procedure still yields good results in that setting.} Under this condition, estimation of $\alpha$ may then proceed by using variable selection methods to choose a set of relevant control variables from among the set $\ddot x_{it}$ to use in estimating ((ref)).
To estimate $\alpha$ in this environment, we adopt the post-double-selection method of \citeasnoun{BelloniChernozhukovHansen2011}. This method proceeds by first substituting ((ref)) into ((ref)) to obtain predictive relationships for the outcome $\ddot y_{it}$ and the treatment $\ddot d_{it}$ in terms of only control variables:
We then use two variable selection steps. Cluster-Lasso is applied to equation ((ref)) to select a set of variables that are useful for predicting $\ddot y_{it}$; we collect the controls $x_{itj}$ for which $\widehat \pi_{j} \neq 0$ in the set $\widehat I_{RF}$. Cluster-Lasso is then applied to equation ((ref)) to select a set of variables that are useful for predicting $\ddot d_{it}$; we again collect the controls $x_{itj}$ for which $\widehat \beta_{m,j} \neq 0$ in the set $\widehat I_{FS}$. The set of controls that will be used is then defined by the union $\widehat I = \widehat I_{FS} \cup \widehat I_{RF}$. Estimation and inference for $\alpha$ may then proceed by ordinary least squares estimation of $\ddot y_{it}$ on $\ddot d_{it}$ and the set of controls in $\widehat I$ using conventional clustered standard errors arellano:feinf.
\citeasnoun{BelloniChernozhukovHansen2011} develop and discuss the post-double-selection method in detail. They note that including the union of the variables selected in each variable selection step helps address the issue that model selection is inherently prone to errors unless stringent assumptions are made. As noted by \citeasnoun{leeb:potscher:pms}, the possibility of model selection mistakes precludes the possibility of valid post-model-selection inference based on a single Lasso regression within a large class of interesting models. The chief difficulty arises with covariates whose effects in ((ref)) are small enough that the variables are likely to be missed if only ((ref)) is considered but have large effects in ((ref)). The exclusion of such variables may lead to substantial omitted variables bias if they are excluded which is likely if variables are selected using only ((ref)).\footnote{The argument is identical if only ((ref)) is used for variable selection exchanging the roles of ((ref)) and ((ref)). The argument also holds if considering only one of ((ref)) or ((ref)) for variable selection.} Using both model selection steps guards against such model selection mistakes and guarantees that the variables excluded in both model selection steps have a neglible contribution to omitted variables bias under Condition ASM.
We present additional moment and rate conditions before stating a result which can be used for performing inference about $\alpha$. Results will be valid uniformly over the large class of models that satisfy the following conditions as well as appropriate conditions from Section (ref). We again define several moments using the same notation introduced before condition SMIV.
Condition SMPLM
(i) Sufficient conditions for Post-Cluster-Lasso: ASM, SE, R hold for models (ref) and (ref).
(ii) Sufficient conditions for asymptotic normality of $\widehat \alpha$ and consistency of $\frac{\imath_T^D}{T}\widehat V$:
(a) $Q={\mathrm{E}}\left[ \frac{1}{T} \sum_{t=1}^T \ddot u_{it}^2\right]$, $ {\mathrm{E}}\left[\frac{1}{T} \sum_{t=1}^T \ddot u_{it}^2 \ddot \zeta_{it}^2\right] $, $ {\mathrm{E}}\left[\left( \frac{1}{T} \sum_{t=1}^T \ddot u_{it}^2 \right)^2\right] $ are bounded uniformly from above and away from zero, uniformly in $n,T$. Additionally, ${\mathrm{E}}\left[ \left(\frac{1}{T} \sum_{t=1}^T \ddot \zeta_{it}^2\right)^q \right] = O(1), \ {\mathrm{E}}\left[ \left(\frac{1}{T} \sum_{t=1}^T \ddot u_{it}^2\right)^q \right] = O(1)$ and ${\mathrm{E}}\left[ \left(\frac{1}{T} \sum_{t=1}^T \ddot d_{it}^2\right)^q \right] = O(1)$ for some $q>4$. $|\alpha| \leqslant B < \infty$.
(b) $ \varpi_{u\zeta} / \sqrt{ {\mathrm{E}} \phi_{u\zeta}^2} = O(1)$, $\max \limits_{1 \leqslant j \leqslant p} \varpi_{x_j\zeta}/\sqrt{{\mathrm{E}} \phi_{x_j \zeta}^2} = O(1),$ $\max \limits_{1 \leqslant j \leqslant p} \varpi_{x_ju}/\sqrt{{\mathrm{E}} \phi_{x_j u}^2} = O(1)$.
(c) $ \max_j \frac{\imath_T^{x_j \zeta}}{T} \phi_{x_j \zeta}^2 = O_P(1) $, $\frac{\imath_T^{u\zeta}}{T} \phi_{u\zeta}^2 = O_{\mathrm{P}}(1),$ $\max_j \frac{\imath_T^{x_j u}}{T} \phi_{x_j u}^2 = O_P(1)$, $\frac{1}{T} \phi_{ud}^2 = O_{\mathrm{P}}(1)$.
(d) $\frac{\imath_T^{u\zeta}}{\min \{\imath_T^{RF},\imath_T^{FS},\min_j \{\imath_T^{x_j \zeta}\}\}}\left(s + n^{2/q}\right)\left(\max_{i,t,j} \ddot x_{itj}^2\right) \frac{s \log^2(p \vee nT)}{n} =o_{\mathrm{P}}(1)$.
Finally, we define the following variance estimators for the post double selection procedure:
$$\widehat V_n = \widehat Q^{-1} \widehat \Omega \widehat Q^{-1}$$ $$\widehat Q = {\frac{1}{nT}\sum_{i=1}^n \sum_{t=1}^T} \widehat { u}_{it}^2, \ \ \ \widehat \Omega = {\frac{1}{nT}\sum_{i=1}^n \sum_{t=1}^T}\sum_{t'=1}^T \widehat {u}_{it} \widehat{ u}_{it'} \widehat{ \zeta}_{it} \widehat{ \zeta}_{it'},$$ where
Given the above conditions, we have the following central limit theorem for the Post-Double Cluster-Lasso estimator $\widehat \alpha$.
This theorem verifies that the OLS estimator which regresses $\ddot y_{it}$ on $\ddot d_{it}$ and the union of variables selected by Cluster-Lasso from ((ref)) and ((ref)) is consistent and asymptotically normal with asymptotic variance that can be estimated with the conventional clustered standard error estimator. Inference based on this result will be valid uniformly over a large class of data generating processes which includes cases where perfect variable selection is impossible.
The results in the previous sections suggest that Cluster-Lasso based estimates should have good estimation and inference properties in panel models with individual specific heterogeneity provided the sample size $n$ is large. In this section, we provide simulation evidence about the performance of our asymptotic approximation for inference about structural parameters in IV models with fixed effects and many instruments and linear fixed effects models when Cluster-Lasso is used for variable selection. We also provide a comparison with several other standard estimators. The Cluster-Lasso based procedures compare favorably to all other feasible approaches considered.
The first simulation illustrates the performance of the Cluster-Lasso based IV estimator in a simple instrumental variables model with fixed effects and many instruments. In our simulation experiments, we generate data from the linear IV model
We generate disturbances according to
with initial conditions for $\epsilon_{it}$ and $u_{it}$ drawn from their stationary distribution. We generate the individual heterogeneity $e_i$ for $i = 1,...,n$ as correlated normal random variables with ${\mathrm{E}}[e_i] = 0$, Var$(e_i) = \frac{4}{T}$, and Corr$(e_i,e_j) = .5^{|i-j|}$ for all $i$ and $j$. We set $f_i = e_i$. We draw the instruments conditional on the fixed effects from
where $\varphi_{itj}$ are normal random variables with ${\mathrm{E}}[\varphi_{itj}] = 0$, Var$(\varphi_{itj}) = 1$, and Corr$(\varphi_{itj},\varphi_{itk}) = .5^{|j-k|}$ that are independent across $i$ and $t$. In all of our simulations, we set $\rho_{\epsilon} = \rho_{u} = \rho_z = .8$, and we set $\rho_{\nu} = .5$. We also set $\alpha = .5$. We redraw the disturbances $\epsilon$ and $u$ at each simulation replication but condition on one realization of the fixed effects and instruments. We consider different sample sizes set to $n=50,100,150,200$ all with $T=10$.
Note that the instruments are not valid without conditioning on the fixed effects within this structure. The fixed effects are also dense in the sense that most of the generated effects will be small but non-zero. This feature would lead to a failure of variable selection methods that included the set of fixed effects in the variables over which selection will occur. Such methods would fail to include the majority of the effects which would then result in invalidity of the instruments and substantial bias in the resulting estimates of $\alpha$. A simple and widely-used way to bypass this problem is removing the entire set of fixed effects as considered in this paper.
The final features of the design are the number of instruments and the structure of the coefficients on the instruments, $\pi$. We consider three different coefficient vectors $\pi_1, \pi_2,$ and $\pi_3$ defined as
for $1 \leqslant j \leqslant p$ where $\lfloor a \rfloor$ returns the integer part of $a$. We refer to designs using $\pi_{1}$, $\pi_{2}$, and $\pi_3$ as Design 1, Design 2, and Design 3 respectively. Design 1 is approximately sparse and should be the most favorable setting as the majority of the signal concentrates in the smallest number of variables among the designs considered. Design 2 does not satisfy Condition ASM, but a substantial amount of the signal is captured by the first few variables. Cluster-Lasso should be able to reliably identify these variables which should result in reasonable properties of the Cluster-Lasso based IV estimator, though there should be a substantial loss of efficiency relative to the infeasible setting where the exact values of $\pi_2$ are known. Design 3 is exactly sparse but should be the most difficult design because the signal is diffused equally over twice as many variables as in Design 1. This spreading of the signal will make it harder to reliably detect any of the instruments. Finally, we consider two different numbers of instruments, $p=n \times (T-2)$ and $p = n \times (T+2)$, for each sample size and design of first-stage coefficients.
For each setting, we report results from five different estimators. We consider IV estimates based on variables selected using the clustered penalty loadings developed in this paper (Clustered Loadings). As a comparison, we also consider IV estimates based on variables selected using the loadings that are valid with heteroscedastic and independent data from \citeasnoun{BellChenChernHans:nonGauss} (Heteroscedastic Loadings). In cases with $p<nT$, we report estimates using 2SLS on the full set of instruments (All). Finally, we consider two different infeasible oracle estimators. The first oracle knows the value of the coefficients $\pi$ (Oracle) while the second also knows the exact values of the fixed effects (FE Oracle). Thus, both oracle estimators use a single instrument that uses the true values of the first stage coefficients, $z_{it}'\pi$. The difference between the two is that the fixed effects are removed by taking differences of all variables from within-individual means in the Oracle results while the true values of the FE are directly subtracted from $y_{it}$ and $d_{it}$ in the FE Oracle results.
The results are based on 1000 simulations for each setting described above. For results based on All, Heteroscedastic Loadings, Clustered Loadings, and Oracle, the fixed effects are treated as unknown parameters and eliminated by taking deviations from within-individual means. For each estimator, we report mean bias, root mean squared error, and rejection rates for a 5%-level test of $H_0:\alpha=.5$ using both clustered standard errors and heteroscedastic standard errors.\footnote{Since moments of IV estimators may not exist, we calculate truncated bias and truncated RMSE, truncating at $\pm 10,000$.} In some of the simulation replications, the IV estimator using variables selected by Lasso is undefined as Lasso sets all coefficients to zero. In such a case, we record a failure to reject the null which is a conservative alternative to applying the Sup-score statistic described in \citeasnoun{BellChenChernHans:nonGauss}. Mean bias and root-mean-square-error for Lasso-based estimates are calculated conditional on Lasso selecting at least one instrument.
The results for estimation of $\alpha$ with first stage coefficients $\pi_1$, $\pi_2$, and $\pi_3$ are reported respectively in Tables (ref), (ref), and (ref) when $p = n \times (T-2)$ and Tables (ref), (ref), and (ref) when $p = n \times (T+2)$. The two oracle estimators provide infeasible benchmarks. Looking at these results, we do see that IV based on the infeasible instruments formed using the true values of the first-stage coefficients perform well in the designs considered. As expected given the well-known properties of 2SLS, the 2SLS estimates using the full set of instruments when $p < nT$ exhibit large bias relative to standard error, large RMSE, and produce tests that suffer from large size distortions.
The Lasso-based results where we do variable selection using loadings that are appropriate under independence but ignore within-individual dependence are quite interesting. This approach performs relatively well compared to naive 2SLS using all of the instruments. However, using instruments selected by Lasso with loadings that ignore the dependence produces an IV estimator of $\alpha$ that has a substantial bias and results in tests that have large size distortions even when clustered standard errors are applied. The presence of this bias illustrates the point that care must be taken when selecting instruments for a post model selection analysis. In general, ${\mathrm{E}}[z_{itj}\epsilon_{it}| \ \ j \ \ \text{selected}] \neq 0$ though the difference from zero is ignorable when ((ref)) occurs. However, in the absence of the regularization event ((ref)), this conditional expectation can be large which introduces a type of “endogeneity” bias as the selected instruments are effectively invalid. We see this behavior when using the heteroscedastic loadings in the designs we consider as these loadings produce smaller penalty levels than the appropriate clustered loadings which results in the spurious inclusion of instruments.
Finally, we see that IV based on instruments selected by Cluster-Lasso clearly dominates the other feasible procedures in the simulation designs considered. In Designs 1 and 2, using this procedure produces tests that have approximately correct size that is comparable to size of tests based on both oracle models considered. We also see that the performance for Bias, RMSE, and size of tests is similar to the infeasible Oracle benchmark in Design 1. Designs 2 and 3 were both designed to be difficult. As expected, we see a substantial loss in RMSE relative to the infeasible oracles in Design 2 as the Lasso-based variable selection is unable to consistently identify and exploit the signal available in the variables with small, non-zero coefficients. It is reassuring that performance is still reasonable in this setting. We also see that the more diffuse signal in Design 3 poses challenges when $n$ is small as the contribution each variable makes to the overall signal is too weak to be reliably detected. Thus, there are many replications where no instruments are selected and replications where only a subset of the relevant variables are selected, effectively resulting in weak identification. For the larger sample sizes, this problem is diminished though performance still deviates from the oracle benchmark. Overall, these results are favorable in that using Cluster-Lasso to select instruments outperforms the other methods explored here and performs reasonably well even in somewhat adversarial conditions.
In this simulation, we consider estimation of a coefficient on a variable of interest in a standard linear fixed effects model. Specifically, we generate data according to the model
We generate disturbances according to
with initial conditions for $\epsilon_{it}$ and $u_{it}$ drawn from their stationary distribution. We generate $e_i$, $f_i$, and $z_{it}$ exactly as in Section (ref), so we omit the details for brevity. We again set $\rho_{\epsilon} = \rho_{u} = .8$ and set $\alpha = .5$. We redraw the disturbances $\epsilon$ and $u$ at each simulation replication but condition on one realization of the fixed effects and controls. We use sample sizes set to $n=50,100,150,200$ with $T=10$. As in Section (ref), failure to condition on the full set of fixed effects may result in substantial biases in estimation of $\alpha$. The structure of the fixed effects will also invalidate methods that include the set of fixed effects in the variables over which selection will occur as illustrated in the results below.
As in Section (ref), we consider three different specifications for the coefficient vectors $\beta$ and $\gamma$:
for $1 \leqslant j \leqslant p$ where $\lfloor a \rfloor$ returns the integer part of $a$. Again, in each design, we consider $p=n \times (T-2)$ and $p = n \times (T+2)$. Designs 1 and 3 clearly fall within the set of models covered by our theoretical development, and we expect the estimation and inference about $\alpha$ following selection of controls using Cluster-Lasso to perform well in either case, though Design 3 is more difficult than Design 1. The relationship between $d_{it}$ and $z_{it}$ in Design 2 does not satisfy Condtion ASM, but the relationship between $y_{it}$ and $z_{it}$ satisfies this condition.
For this simulation, we consider six estimators of $\alpha$. When $p \leqslant nT$, we use the conventional fixed effects estimator including all the variables in $z_{it}$ (All). We use the post-double-selection method with penalty loadings appropriate for independent, heteroscedastic data in each Lasso stage (Heteroscedastic Loadings) and with our clustered loadings in each Lasso stage (Clustered Loadings). We also consider a post-double-selection estimator which includes the fixed effects in the set of variables over which selection occurs (Select over FE) using the approach of \citeasnoun{kock:hdpanel}. We also consider two oracle estimators. The first oracle knows the values of the coefficients $\beta$ and $\gamma$ (Oracle) while the second also knows the exact values of the fixed effects (FE Oracle). The Oracle estimate of $\alpha$ is thus obtained by regressing $\ddot y_{it} - \ddot z_{it}'\beta$ onto $\ddot d_{it} - \ddot z_{it}'\gamma$ while the FE Oracle estimate of $\alpha$ is obtained by regressing $y_{it} - z_{it}'\beta - e_i$ onto $d_{it} - z_{it}'\gamma - f_i$. As before, the results are based on 1000 simulation replications for each setting; and we report mean bias, root mean squared error, and rejection rates for a 5%-level test of $H_0:\alpha=.5$ using both clustered standard errors and heteroscedastic standard errors for each estimator.
The results for the three partially linear model designs are reported in Tables (ref), (ref), and (ref) for $p=n \times (T-2)$ and Tables (ref), (ref), and (ref) for $p = n \times (T+2)$. The two oracle estimators provide infeasible benchmarks and unsurprisingly produce estimators with small bias and RMSE and tests with reasonable size as long as clustered standard errors are used as is conventional in the literature, e.g. \citeasnoun{bdm:cluster}. In all simulations, estimates using the full set of controls when feasible have small bias but large variability leading to large RMSE relative to oracle estimates. Tests based on estimators using the full set of controls are also badly size distorted regardless of whether heteroscedastic or clustered standard errors are used. This distortion results from the difficulty in robustly estimating standard errors when many variables are included which is an unresolved topic of current research; see, e.g., \citeasnoun{CJN:PLMStandardError}. This feature suggests that one may not wish to simply including many controls without regularization even when possible.
Estimates based on the double selection method using Lasso with penalty loading appropriate under heteroscedasticity and independence or using Lasso to also select over the fixed effects perform better than simply including all controls in the $p < nT$ case but tend to perform poorly in terms of bias and coverage probabilities. The bias and poor coverage properties of the estimator that attempts to select over the fixed effects is due to the difficulties in performing selection over the dense part of the model and shows the importance of eliminating fixed effect parameters via demeaning or differencing. These difficulties arise because sparsity provides a poor approximation to the true fixed effects structure. Note that a dense model over unobserved heterogeneity where heterogeneity matters differentially for each individual seems quite reasonable in many economic applications and suggests that attempting to select over fixed effects may result in undesirable features at least when inference about model parameters is the goal of the empirical analysis.
We find it more surprising that using heteroscedastic penalty loadings also leads to noticeable bias and a distortion in statistical size. The heteroscedastic loadings lead to less penalization in our designs which result in inclusion of a few spurious variables. Usual intuition for linear models suggests that including a few extra variables has little impact on say ordinary least squares estimates of parameters of interest. The difficulty arises because the spuriously included variables are not included at random but are exactly those variables with little to no impact that are most highly correlated to the noise and are not properly screened out because the penalty is too low for ((ref)) to be a reliable guide. Choosing the variables most highly correlated to the noise then yields that ${\mathrm{E}}[ x_{itj} \epsilon_{it} | \ \ j \ \ \text{selected} ]$ is not negligible due to the use of incorrect penalty loadings leading to biased estimation just as in the instrumental variables case.
Finally, we again see that basing estimation and inference for $\alpha$ on the post-double-selection method using clustered penalty loadings clearly dominates the other feasible procedures in the simulation designs considered. This procedure yields an estimator with RMSE comparable to the oracles across all designs considered. We also see that feasible inference based on this procedure does a relatively good job controlling size across all designs considered. Overall, these results are favorable to Lasso-based variable selection using clustered penalty loadings after partialing out fixed effects and suggests that these methods may offer useful tools to empirical researchers faced with high dimensional panel data.
In the earlier sections, we provided results on the performance of Lasso as a model selection device for panel data models with fixed effects and discussed how to apply Lasso to problems of economic interest in such settings. In this section, we demonstrate the use of Cluster-Lasso by reexamining the \citeasnoun{cook:ludwig:guns} study of the impact of gun ownership on crime. We briefly review \citeasnoun{cook:ludwig:guns} before presenting the results using the methods described in this paper.
\citeasnoun{cook:ludwig:guns} give several arguments suggesting that gun ownership levels may impose externalities on a community. On the one hand, widespread prevelance of guns can act a deterrent to criminal activity. On the other hand, higher gun prevelance in the general population can lead to higher gun ownership among dangerous people, perhaps through theft or illegal sales, which may lead to an increase in crime. Thus, it is unclear whether the net effect of guns is positive or negative. To investigate the impact of guns, \citeasnoun{cook:ludwig:guns} estimate the effect of gun prevelance on several measures of crime rates. In this example, we revisit their estimation of the effect of gun prevelance on homicide rates.
A major contribution of \citeasnoun{cook:ludwig:guns} is to provide an improved measure of gun ownership in order to get more accurate estimates of the social costs of gun prevalence.\footnote{Previously, several authors had obtained conflicting estimates for the effect of interest; see, e.g., \citeasnoun{lottbook} and \citeasnoun{duggan2001}.} Because exact gun-ownership numbers in the U.S. are difficult to obtain, \citeasnoun{cook:ludwig:guns} instead use the fraction of suicides committed with a firearm (abbreviated FSS) within a county as a proxy for county-level gun ownership rates. \citeasnoun{cook:ludwig:guns} argue that if guns are prevalent within a county, then they should be more accessible for the purpose of suicide. They show that their proxy for gun prevelance, FSS, matches up with survey data directly measuring gun ownership from the General Social Survey better than previously used measures. In our analysis, we take it as given that FSS provides a useful measure of gun ownership and that learning the causal effect of FSS is an interesting goal. We thus abstract from any further measurement issues in order to give a clear illustration of our methods.
The main strategy employed by \citeasnoun{cook:ludwig:guns} to estimate causal effects of gun prevalence is to exploit differences in gun ownership across counties and over time. \citeasnoun{cook:ludwig:guns} construct a panel of 195 large United States counties between the years 1980 through 1999 and use this data to estimate linear fixed effects models of the form
where $\alpha_i$ and $\delta_t$ are respectively unobserved county and year effects that will be treated as parameters to be estimated, $X_{it}$ are additional covariates meant to control for any factors related to both gun ownership rates and crime rates that vary across counties and over time, and $Y_{it}$ is one of three dependent variables: the overall homicide rate within county $i$ in year $t$, the firearm homicide rate within county $i$ in year $t$, or the non-firearm homicide rate within county $i$ in year $t$. \citeasnoun{cook:ludwig:guns} consider controls, $X_{it}$, for percent African American, percent of households with female head, nonviolent crime rates, and percent of the population that lived in the same house five years earlier.\footnote{Details about the controls can be found in \citeasnoun{cook:ludwig:guns}.}
Interpreting the estimated effect of gun prevelance as measured by FSS as causal relies on the belief that there are no variables associated both to crime rates and FSS that are not included in ((ref)). The inclusion of county and time fixed effects accounts for any aggregate macroeconomic conditions that affect all counties uniformly and any county-level characteristics that do not vary over time. The additional variables used by \citeasnoun{cook:ludwig:guns} in $X_{it}$ are then meant to capture all other sources of variation that are correlated to both FSS and the log of of the homicide rate. Of course, one might worry that the set of controls included in $X_{it}$ does not adequately capture remaining confounds after controlling for time and county effects.
We extend the analysis performed in \citeasnoun{cook:ludwig:guns} by allowing for a much larger set of potential control variables which may strengthen the plausibility of the claim that all sources of confounding variation have been captured. Specifically, we consider an essentially identical model $$\log Y_{it} = \beta_0 + \beta_1 \text{log FSS}_{it-1} + W_{it}'\beta_W + \alpha_i + \delta_t + \epsilon_{it}$$ which differs from ((ref)) by our consideration of a large set of variables in $W_{it}$. We form $W_{it}$ by taking variables compiled by the US Census Bureau. Basic variables include county-level measures of demographics, the age distribution, the income distribution, crime rates, federal spending, home ownership rates, house prices, educational attainment, voting paterns, employment statistics, and migration rates.\footnote{The exact identities of the variables are available upon request. The entire dataset is taken from the U.S. Census Bureau USA Counties Database and can be downloaded at http://www.census.gov/support/USACdataDownloads.html.} We note that $W_{it}$ includes measures meant to capture all the variables controlled for in \citeasnoun{cook:ludwig:guns} in their $X_{it}$, though with our data and construction we do not reproduce their results exactly. However, we show below that we obtain similar results with both sets of variables. A key concern with the fixed effects model is that there is some feature of the counties that is correlated not just to the level of crime rates and gun ownership but also to the evolution of these variables. To flexibly allow for this possibility, we also include interactions of the initial (1980) values of all control variables with a linear, quadratic, and cubic term in time. With the main effects and interactions of initial conditions with a cubic trend, we end up with 978 total control variables.
While controlling for a large set of variables may make the assumption that all relevant confounds have been included in the model more plausible, including too many covariates may lower estimation precision and also complicates estimation of the variance of estimators as illustrated in Section (ref). Thus, a researcher faces a tradeoff between making sure that relevant confounds are included in the model and being able to draw meaningful conclusions from the data. Using variable selection as outlined in this paper offers one potential resolution to this tension by allowing consideration of a large set of controls while maintaining parsimony and producing valid inferential statements under the assumption that the set of confounds that needs to be included after accounting for the full set of fixed effects is small relative to the sample size.
We present estimation results in Table (ref) with results for each dependent variable presented across the columns and each row corresponding to a different specification. As a baseline, we report numbers taken directly from the first row of Table 3 in \citeasnoun{cook:ludwig:guns} in the first row of Table 1 (“Cook and Ludwig (2006) Baseline”). \citeasnoun{cook:ludwig:guns} obtained these results by regressing log homicide rates on lagged log FSS, county and time fixed effects, and the baseline set of controls mentioned above and use these numbers as their baseline results.
We report results obtained from our data in Rows 2-4 (labeled “FSS $+$ Census Baseline”, “ Full Set of Controls”, and “Cluster Post-Double Selection”).\footnote{All results are based on weighted regression where we weight by the within-county average population over 1980-1999.} In Row 2 of Table (ref) (“FSS $+$ Census Baseline”), we attempt to replicate the result from Row 1 using control variables gathered from the census that correspond to the variables indicated as being used in \citeasnoun{cook:ludwig:guns} Table 3, Row 1. Despite using slightly different data, we produce results that are fairly similar to those reported in \citeasnoun{cook:ludwig:guns}. Specifically, \citeasnoun{cook:ludwig:guns} give point estimates (standard errors) of the coefficient on lagged log FSS of .086 (.038) for overall homicide rates and .173 (.049) for gun homicide rates; and we obtain estimated effects (standard errors) of .070 (.035) for overall homicide rates and of .178 (.046) for gun homicide rates. The discrepancy between the results is somewhat larger for non-gun homicide rates, though the results are still broadly consistent with each other. \citeasnoun{cook:ludwig:guns} report an estimated effect (standard error) of -.033 (.040) while we estimate the effect to be -.071 with a standard error of .038.
We provide the results based on the large set of controls in Rows 3 and 4 of Table (ref). In Row 3, we present the results based on using all 978 potential controls in addition to the full set of county and time effects. Using all of the controls, the estimated effect of lagged suicide rates is small for each dependent variable. The estimated coefficients (standard errors) are only -.010 (.033) for overall homicide rates, .00004 (.044) for gun homicide rates, and -.033 (.042) for non-gun homicide rates. These results are relatively imprecise, and one could not rule out moderate sized positive or negative effects for any of the dependent variables. In addition, the simulation results illustrate that the estimated standard errors with a large number of controls may be inaccurate, suggesting that one should be hesitant in trusting these results as accurate standard errors may be even larger. Of course, it is not obvious that one would believe that all 978 controls are necessary though one may not be sure of the exact identities of the variables that should be included. If this is the case, the methods for variable selection developed in this paper offer one avenue for finding a relevant set of controls that should be included.
In Row 4, we present estimates of the effect of gun prevalence on homicide rates based on the post-double-selection method using Cluster-Lasso to select controls after partialing out the fixed effects. We also provide the identities of the selected controls in Table (ref). For both overall homicide rates and gun homicide rates, the estimates based on Cluster-Lasso selected controls are very similar to those obtained with the baseline set of controls in our data though standard errors are slightly larger. For overall homicide results, the Cluster-Lasso estimate (standard error) is .079 (.043) compared to .070 (.035) with the baseline controls; and the Cluster-Lasso estimate (standard error) is .171 (.047) compared to .178 (.046) with the baseline controls when gun homicide is the dependent variable. This similarity is interesting given that the set of variables selected by Lasso differs substantively from the set of baseline controls. For the overall homicide rate, we would fail to reject the null hypothesis that gun prevalence as measured by suicide rates is not associated to homicide rates after controlling for a broad set of variables at the 5% level in the Cluster-Lasso results, though we would reject the hypothesis of no effect of gun prevalence on overall homicide rates at the 10% level. The result is stronger when the gun homicide rate is the dependent variable. In this case, one would draw the conclusion that more guns, as measured by the firearm suicide rate, is strongly positively associated with more homicides committed with firearms. Under the assumption that the set of controls considered is sufficient to account for relevant confounds, one could also take these estimated effects as causal. This assumption seems more plausible in the Cluster-Lasso results which allow for consideration of a richer set of controls than the baseline results.
Finally, we turn to the results with non-gun homicide as the dependent variable. In this case, there is a larger discrepancy between the baseline results and the results using controls selected by Cluster-Lasso, though one would draw the same qualitative conclusion in either case. With the baseline intuitively selected set of controls, the estimated effect of gun prevalence is -.071 with an estimated standard error of .038; and the estimated effect is smaller in magntitude, at -.019, with an estimated standard error of .040 using the Cluster-Lasso selected controls. In both cases, we would fail to reject the null hypothesis that gun prevalence as measured by suicide rates is not associated to non-gun homicide rates after controlling for a broad set of variables at conventional levels, and one could not rule out moderate positive or negative effects of gun prevalence on non-gun homicide at conventional levels using the Cluster-Lasso based results.
Overall, our Cluster-Lasso based results are broadly consistent with the claims of \citeasnoun{cook:ludwig:guns}. We find a strong positive effect of gun prevalence on the firearm homicide rate after allowing for a large set of confounds and including a full set of county and time effects. We also find some evidence of a positive effect of gun prevalence on overall homicide rates, but produce an imprecise estimate of the effect on non-gun homicides which could be consistent with moderate positive or negative effects. The similarity to the \citeasnoun{cook:ludwig:guns} results adds further credibility to their claims as we allow for a richer set of confounding variables.\footnote{Of course, the estimates are still not valid causal estimates if one does not believe the claim that the lagged log of the firearm suicide rate provides an exogenous measure of gun prevalence after controlling for our large set of county level controls, county, and time fixed effects. Also, note that the similarity between results following selection and results based on an intuitively selected initital set of controls is not mechanical as evidenced, for example, in the empirical example in \citeasnoun{BelloniChernozhukovHansen2011}.}
In this paper, we have considered variable selection using Lasso in panel data allowing for a clustered error structure. This structure allows for strong dependence among observations within the same individual and is commonly assumed in panel data applications in economics. Handling this structure is also important for allowing us to verify good properties of variable selection after partialing out individual specific fixed effects by taking deviations from within individual means which, in general, will induce a clustering structure. We show that Lasso continues to have good selection and estimation properties allowing for this structure when clustered penalty loadings are used, and we use this good performance to establish results for doing inference following variable selection in IV models with additive fixed effects and partially linear models with additive fixed effects. We show that these methods perform well in a simulation study and illustrate their use in estimating the effect of gun prevalence on crime as in \citeasnoun{cook:ludwig:guns}. In the empirical example, we find that our results are broadly consistent with results from \citeasnoun{cook:ludwig:guns} despite allowing for a broader set of controls.