Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
78,455 characters · 19 sections · 105 citation commands
Omitted variable bias of Lasso-based inference methods: A finite sample analysis
Since their introduction, post double Lasso belloni2014inference and debiased Lasso javanmard2014confidence,vandergeer2014asymptotically,Zhang_Zhang have quickly become the most popular inference methods for problems with many control variables. Given the rapidly growing theoretical\footnote{
\singlespacing} and applied\footnote{
} literature on these methods, it is crucial to take a step back and examine the performance of these procedures in empirically relevant settings and to better understand their merits and limitations relative to other alternatives.
In this paper, we study the performance of post double Lasso and debiased Lasso from two different angles. First, we develop a theory of the finite sample behavior of post double Lasso and debiased Lasso. We show that under-selection of the Lasso can lead to substantial omitted variable biases (OVBs) in finite samples. Our theoretical results on the OVBs are non-asymptotic; they complement but do not contradict the existing asymptotic theory. Second, we conduct extensive simulations and two empirical studies to investigate the performance of Lasso-based inference methods in practically relevant settings. We find that the OVBs can render the existing asymptotic approximations inaccurate and lead to size distortions.
Following the literature, we consider the following standard linear model
Here $Y_{i}$ is the outcome, $D_{i}$ is the scalar treatment variable of interest, and $X_{i}$ is a $(1\times p)$-dimensional vector of control variables. In the main text, we focus on the performance of post double Lasso for estimating and making inferences (e.g., constructing confidence intervals) on the treatment effect $\alpha^{\ast}$ in settings where $p$ can be larger than or comparable to $n$. We present results for debiased Lasso in Appendix (ref).
Post double Lasso consists of two Lasso selection steps: a Lasso regression of $Y_{i}$ on $X_{i}$ and a Lasso regression of $D_{i}$ on $X_{i}$. In the third step, the estimator of $\alpha^{\ast}$, $\tilde{\alpha}$, is the OLS regression of $Y_{i}$ on $D_{i}$ and the union of controls selected in the two Lasso steps. OVB arises in post double Lasso whenever the relevant controls (i.e., the controls with non-zero coefficients) are selected in neither Lasso steps, a situation we refer to as double under-selection. Results that can explain when and why double under-selection occurs are scarce in the existing literature. This paper shows theoretically that this phenomenon can even occur in simple examples with classical assumptions (e.g., normal homoscedastic errors, orthogonal designs for the relevant controls), which are often viewed as favorable to the performance of the Lasso. We prove that if the products of the absolute values of the non-zero coefficients and the variances of the controls are no greater than half the regularization parameters derived based on standard Lasso theory\footnote{
\singlespacing}, Lasso fails to select these controls in both steps with high probability.\footnote{
\singlespacing} This result allows us to derive the first non-asymptotic lower bound formula in the literature for the OVB of the post double Lasso estimator $\tilde{\alpha}$. Our lower bound provides explicit universal constants, which are essential for understanding the finite sample behavior of post double Lasso and its limitations.
The OVB lower bound is characterized by the interplay between the probability of double under-selection and the magnitude of the coefficients corresponding to the relevant controls in (ref)--(ref). In particular, there are three regimes: (i) the omitted controls have coefficients whose magnitudes are large enough to cause substantial bias; (ii) the omitted controls have coefficients whose magnitudes are too small to cause substantial bias; (iii) the magnitude of the relevant coefficients is large enough such that the corresponding controls are selected with high-probability. We illustrate these three regimes in Figure (ref), which plots the bias of post double Lasso and the number of selected control variables as a function of the magnitude of the non-zero coefficients. [FIGURE (ref) HERE.]
Our theoretical analysis of the OVB has important implications for inference procedures based on the post double Lasso. belloni2014inference show that $\sqrt{n}(\tilde\alpha-\alpha^\ast)$ is asymptotically normal with zero mean. We show that in finite samples, the OVB lower bound can be more than twice as large as the standard deviation obtained from the asymptotic distribution in belloni2014inference. This is true even when $n$ is much larger than $p$ and $\beta^\ast$ and $\gamma^\ast$ are sparse. To illustrate, assume that (ref) and (ref) share the same set of $k$ non-zero coefficients and set $(n,\,p)=(14238,\,384)$ as in angrist2019machine, who use post double Lasso to estimate the effect of elite colleges. The ratio of our OVB lower bound to the standard deviation in belloni2014inference is $0.27$ if $k=1$ and $2.4$ if $k=10$. This example shows that the requirement on the sparsity parameter $k$ for the OVBs to be negligible is quite stringent. We emphasize that our findings do not contradict the existing results on the asymptotic distribution of post double Lasso in belloni2014inference. Rather, our results suggest that the OVBs can make the asymptotic zero-mean approximation of $\sqrt{n}(\tilde\alpha-\alpha^\ast)$ inaccurate in finite samples.
To better understand the practical implications of the OVB of post double Lasso, we perform extensive simulations. Our simulation results can be summarized as follows. (i) Large OVBs are persistent across a range of empirically relevant settings and can occur even when $n$ is large and larger than $p$, and the sparsity parameter $k$ is small. (ii) The OVBs can lead to invalid inferences and under-coverage of confidence intervals. (iii) The performance of post double Lasso varies substantially across different popular choices of regularization parameters, and no single choice outperforms the others across all designs. While it may be tempting to choose a smaller regularization parameter than the standard recommendation in the literature to mitigate under-selection, we find that this idea does not work in general and can lead to rather poor performance.
In addition to the simulations, we consider two empirical applications: the analysis of the effect of 401(k) plans on savings by belloni2017program and chernozhukov2018double and the study of the racial test score gap by fryerlevitt2013. We draw samples of different sizes from the large original data and compare the subsample estimates to the estimates based on the original data.\footnote{
} In both applications, we find substantial biases even when $n$ is considerably larger than $p$, and we document that the magnitude of the biases varies substantially depending on the regularization choice.
Given our theoretical results, simulations, and empirical evidence, a natural question is how to make statistical inferences in a reliable manner if one is concerned about OVBs. In many economic applications, $p$ is comparable to but still smaller than $n$. This motivates the recent development of high-dimensional OLS-based inference procedures cattaneo2018inference,dadamo2018cluster,jochmans2020heteroscedasticity,kline2020leave. These methods are based on OLS regressions with all controls and rely on novel variance estimators that are robust to the inclusion of many controls (unlike conventional variance estimators). Based on extensive simulations, we find that OLS with standard errors proposed by cattaneo2018inference demonstrates excellent coverage accuracy across all our simulation designs. Another advantage of OLS-based methods over Lasso-based inference methods is that the former do not rely on any sparsity assumptions. This is important because sparsity assumptions may not be satisfied in applications and, as this paper shows, the OVBs of Lasso-based inference procedures can be substantial even when $k$ is small and $n$ is large and larger than $p$. However, OLS yields somewhat wider confidence intervals than the Lasso-based inference methods, suggesting a trade-off between coverage accuracy and the length of the confidence intervals.
Our analyses suggest two main recommendations concerning the use of post double Lasso in empirical studies. First, if the estimates of $\alpha^{*}$ are robust to increasing the recommended regularization parameters in both Lasso steps, this suggests that either the OVBs are negligible (Regime (ii)) or under-selection is unlikely (Regime (iii)). In either case, post double Lasso is a reliable and efficient method. Otherwise, modern high-dimensional OLS-based inference methods constitute a possible alternative when $p$ is smaller than $n$. Second, our findings highlight the importance of augmenting the final OLS regression in post double Lasso with control variables motivated by economic theory and prior knowledge, as suggested by belloni2014inference.
Consider the following linear regression model
where $\left\{ Y_{i}\right\} _{i=1}^{n}=Y$ is an $n$-dimensional response vector, $\left\{ X_{i}\right\} _{i=1}^{n}=X$ is an $n\times p$ matrix of covariates with $X_{i}$ denoting the $i$th row of $X$, $\left\{ \varepsilon_{i}\right\} _{i=1}^{n}=\varepsilon$ is a zero-mean error vector, and $\theta^{*}$ is a $p$-dimensional vector of unknown coefficients.
The Lasso estimator of $\theta^{*}$, which was first proposed by tibsharini1996regression, is given by
where $\lambda$ is the regularization parameter. Let $\varepsilon\sim\mathcal{N}\left(0_{n},\,\sigma^{2}I_{n}\right)$ and $X$ be a fixed design matrix with normalized columns (i.e., $n^{-1}\sum_{i=1}^{n}X_{ij}^{2}=1$ for all $j=1,\dots,p$). In this example, bickel2009simultaneous set $\lambda=2\sigma\sqrt{2n^{-1}\left(1+\tau\right)\log p}$ (where $\tau>0$) to establish upper bounds on $\sqrt{\sum_{j=1}^{p}(\hat{\theta}_{j}-\theta_{j}^{*})^{2}}$ with a high probability guarantee. To establish perfect selection, wainwright2009sharp sets $\lambda$ proportional to $\sigma\phi^{-1}\sqrt{\left(\log p\right)/n}$, where $\phi\in(0,\,1]$ is a measure of correlation between the covariates with nonzero coefficients and those with zero coefficients.
Besides the classical choices in bickel2009simultaneous and wainwright2009sharp, other choices of $\lambda$ are available in the literature. For instance, belloni2012sparse and belloni2016cluster propose choices that accommodate heteroscedastic and clustered errors. The regularization choice of belloni2012sparse, which is recommended by belloni2014inference for post double Lasso, is based on the following Lasso program:
where $(\hat{l}_1,\dots,\hat{l}_p)$ are penalty loadings obtained using the iterative algorithm developed in belloni2012sparse.
Finally, a very popular practical approach for choosing $\lambda$ is cross-validation; see, for example, homrighausen2013thelasso,homrighausen2014leaveoneout and chetverikov2020cross for theoretical results on cross-validated Lasso.
The model (ref)\textendash (ref) implies the following reduced form model for $Y_{i}$:
where $\pi^{*}=\gamma^{*}\alpha^{*}+\beta^{*}$ and $u_{i}=\eta_{i}+\alpha^{*}v_{i}$.
The post double Lasso, introduced by belloni2014inference, essentially exploits the Frisch-Waugh theorem, where the regressions of $Y$ on $X$ and $D$ on $X$ are implemented with the Lasso:
The final estimator $\tilde{\alpha}$ of $\alpha^{*}$ is then obtained from an OLS regression of $Y$ on $D$ and the union of selected controls
where $\hat{I}_{1}=\textrm{supp}\left(\hat{\pi}\right)=\left\{ j:\,\hat{\pi}_{j}\neq0\right\} $ and $\hat{I}_{2}=\textrm{supp}\left(\hat{\gamma}\right)=\left\{ j:\,\hat{\gamma}_{j}\neq0\right\} $.
This section presents a simple numerical example illustrating the OVB of post double Lasso. All computations were performed in Matlab MATLAB2020. The Lasso is implemented using the built-in function lasso. We consider a simple but classical setting that is often considered favorable to the performance of the Lasso. The data are simulated according to the structural model (ref)--(ref), where $X_{i}\overset{iid}\sim \mathcal{N}\left(0_p,I_{p}\right)$, $\eta_{i}\overset{iid}\sim\mathcal{N}(0,1)$, and $v_{i}\overset{iid}\sim \mathcal{N}(0,1)$ are independent of each other. Our object of interest is $\alpha^{*}$. We set $n=500$, $p=200$, $\alpha^{*}=0$, and consider a sparse setting where $\beta_j^{*}=\gamma_j^{*}=c \cdot 1\left\{j\le k\right\}$ for $j=1,\dots,p$ and $k=5$. Following the simulation exercise in belloni2014inference, we vary the population $R^2$s in (ref) and (ref) by varying the magnitude of the non-zero coefficients $c$. We employ the regularization parameter choice by bickel2009simultaneous.
Figure (ref) displays the finite sample distribution of post double Lasso for different values of $R^2$. For comparison, we plot the distribution of the “oracle estimator” of $\alpha^\ast$, a regression of $Y_i-X_i\pi^\ast$ on $D_i-X_i\gamma^\ast$. The finite sample behavior of post double Lasso depends on how many of the $k=5$ relevant controls get selected in both Lasso steps. Figure (ref) shows histograms of the number of selected relevant controls. [FIGURES (ref) AND (ref) HERE.]
When $R^2=0.5$, post double Lasso exhibits an excellent performance. The finite sample distribution is well-approximated by the normal distribution of the oracle estimator and centered at $\alpha^\ast=0$. The reason for the excellent performance is that all $k=5$ controls are selected with high probability such that post double Lasso essentially coincides with an OLS regression of $Y_i$ on $D_i$ and the five relevant controls.
Let us now consider what happens if we decrease the magnitude of the coefficients and the implied $R^2$. For $R^2=0.3$, post double Lasso exhibits a large finite sample bias. Moreover, the distribution of post double Lasso differs substantially from the distribution of the oracle estimator: it has a larger standard deviation and is slightly skewed. This distribution is a mixture of the distributions of OLS conditional on the two Lasso steps selecting different combinations of relevant controls (Panel (b) in Figure (ref)). For $R^2=0.1$, post double Lasso again exhibits a significant bias. At the same time, the shape of the distribution is similar to that of the oracle estimator. This is because, with high probability, none of the controls gets selected, while the coefficients are large enough to cause an OVB. Decreasing the $R^2$ to $0.01$ reduces the bias. However, it does not change the shape of the finite sample distribution because the selection performance remains unchanged.
The simple numerical example in this section shows that post double Lasso can suffer from OVBs when the two Lasso steps do not select all relevant controls. The magnitude of the coefficients corresponding to the omitted controls can be large enough such that the OVB shifts the location of the finite sample distribution far away from the true value $\alpha^\ast=0$. The issue documented here is not a “small sample” phenomenon but persists even in large sample settings; see Appendix (ref).
This section provides a theoretical analysis of the OVB of post double Lasso. Our goal here is to demonstrate that, even in simple examples with classical assumptions (e.g., normal homoscedastic errors, orthogonal designs for the relevant controls), which are often viewed favorable to the performance of Lasso, the finite sample OVBs of post double Lasso can be substantial relative to the standard deviation provided in the existing literature. We first establish a new necessary result for the Lasso's inclusion and then derive lower and upper bounds on the OVBs of post double Lasso. These results are derived for fixed $\left(n,\,p,\,k\right)$ and are also valid when $\left(k\log p\right)/n\rightarrow0$ or $\left(k\log p\right)/n\rightarrow\infty$. As it will become clear in the following, $p$ needs to be large enough for our results to be informative. Without loss of generality, we normalize the matrix $X$ such that $\left(X_{j}^{T}X_{j}\right)/n=1$ for all $j=1,\dots,p$. We focus on fixed designs (of $X$) to highlight the essence of the problem; see Appendix (ref) for an extension to random designs.
For the convenience of the reader, here we collect the notation to be used in the theoretical analysis. Let $1_{m}$ denote the $m-$dimensional (column) vector of “1”s and $0_{m}$ is defined similarly. The $\ell_{1}-$norm of a vector $v\in\mathbb{R}^{m}$ is denoted by $\left|v\right|_{1}:=\sum_{i=1}^{m}\left|v_{i}\right|$ and the $\ell_{\infty}-$norm of a vector $v\in\mathbb{R}^{m}$ is denoted by $\left|v\right|_{\infty}:=\max_{i=1,\dots,m}\left|v_{i}\right|$. The $\ell_{\infty}$ matrix norm (maximum absolute row sum) of a matrix $A$ is denoted by $\left\Vert A\right\Vert _{\infty}:=\max_{i}\sum_{j}\left|a_{ij}\right|$. For a vector $v\in\mathbb{R}^{m}$ and a set of indices $T\subseteq\left\{ 1,\dots,m\right\} $, let $v_{T}$ denote the sub-vector (with indices in $T$) of $v$. For a matrix $A\in\mathbb{R}^{n\times m}$, let $A_{T}$ denote the submatrix consisting of the columns with indices in $T$. For a vector $v\in\mathbb{R}^{m}$, let $\textrm{sgn}(v):=\left\{ \textrm{sgn}(v_{j})\right\} _{j=1,\dots,m}$ denote the sign vector such that $\textrm{sgn}(v_{j})=1$ if $v_{j}>0$, $\textrm{sgn}(v_{j})=-1$ if $v_{j}<0$, and $\textrm{sgn}(v_{j})=0$ if $v_{j}=0$. Given a set $K$, let $\textrm{card}\left(K\right)$ denote the cardinality of $K$. We denote $\max\left\{ a,\,b\right\} $ by $a\vee b$ and $\min\left\{ a,\,b\right\} $ by $a\wedge b$.
Post double Lasso exhibits OVBs whenever the relevant controls are selected in neither (ref) nor (ref). To the best of our knowledge, there are no formal results strong enough to show that, with high probability, Lasso can fail to select the relevant controls in both steps. Therefore, we first establish a new necessary result for the single Lasso's inclusion in Lemma (ref). To derive this result, we consider the following classical assumptions, which are often viewed favorable to the performance of the Lasso.
Known as the incoherence condition due to wainwright2009sharp, part (iii) in Assumption (ref) is needed for the exclusion of the irrelevant controls. Note that if the columns in $X_{K^{c}}$ are orthogonal to the columns in $X_{K}$ (but within $X_{K^{c}}$, the columns need not be orthogonal to each other), then $\phi=1$. Obviously a special case of this is when the entire $X$ consists of mutually orthogonal columns (which is possible if $n\geq p$). To provide some intuition for Assumption (ref)(iii), let us consider the simple case where $k=1$ and $K=\left\{ 1\right\} $, $X$ is centered (such that $\left\{n^{-1} \sum_{i=1}^{n}X_{ij}\right\} _{j=1}^{p}=0_{p}$), and the columns in $X$ are normalized such that the standard deviations of $X_{1}$ and $X_{j}$ (for any $j\in\left\{ 2,3,\dots,p\right\} $) are identical. Then, $1-\phi$ is simply the maximum of the absolute (sample) correlations between $X_{1}$ and each of the $X_{j}$s with $j\in\left\{ 2,3,\dots,p\right\} $.
Lemma (ref) shows that for large enough $p$, Lasso fails to select any of the relevant covariates with high probability if ((ref)) holds for all $l\in K$. If such conditions hold with respect to both (ref) and (ref), then Lemma (ref) implies that the relevant controls are selected in neither (ref) nor (ref) with probability at least $1-2/p^{\tau}$; see Panels (c) and (d) of Figure (ref) for an illustration.
Let us rewrite ((ref)) as $\left|\theta_{l}^{*}\right|=2^{-1}\left|a\right|\lambda$ with $\left|a\right|\in(0,\,1]$. Assume that $\left|a\right|\phi^{-1}\sigma$ is bounded from above and away from zero; moreover, $\lambda$ satisfies (ref) and scales as $\sqrt{\left(\log p\right)/n}$. These conditions imply that $\left|\theta_{l}^{*}\right|\asymp\sqrt{1/n}$ in the classical asymptotic framework where $n\rightarrow \infty$ and $p$ is fixed. This regime of $\theta_{l}^{*}$ is exactly where classical model selection procedures struggle to distinguish a coefficient from zero in low-dimensional settings.
In the introduction, we have discussed the relationship of our result to that in lahiri2021necessary. It is also interesting to compare Lemma (ref) with the results in wainwright2009sharp. Note that ((ref)) implies $\mathbb{P}\left(\hat{\theta}_{l}\neq0\right)\leq1/p^{\tau}$ for any $l\in K$ subject to ((ref)). In comparison, wainwright2009sharp shows that whenever $\theta_{l}^{*}\in\left(\lambda\textrm{sgn}\left(\theta_{l}^{*}\right),\,0\right)$ or $\theta_{l}^{*}\in\left(0,\,\lambda\textrm{sgn}\left(\theta_{l}^{*}\right)\right)$ for some $l\in K$,
Constant bounds in the form of ((ref)) cannot explain that, with high probability, Lasso fails to select the relevant covariates in both (ref) and (ref) when $p$ is sufficiently large.
In this section, we apply Lemma (ref) to derive lower bounds on the OVB of post double Lasso. We consider the structural model (ref)--(ref), which can be written in matrix notation as
In matrix notation, the reduced form (ref) becomes
where $\pi^{*}=\gamma^{*}\alpha^{*}+\beta^{*}$ and $u=\eta+\alpha^{*}v$. We make the following assumptions about model ((ref))--((ref)).
Proposition (ref) derives a lower bound formula for the OVB of post double Lasso concerning the case where $\alpha^{*}=0$.
Let us compare the non-asymptotic lower bound in Proposition (ref) to the implications of the existing asymptotic results for the bias of post double Lasso. If $\sigma_{v}$ is bounded away from zero and $\sigma_{\eta}$ is bounded from above, the existing theory would imply that the biases of post double Lasso are bounded from above by $\texttt{constant}\cdot\left(k\log p\right)/n$, irrespective of whether Lasso fails to select the relevant controls or not, and how small $\left|a\right|$ and $\left|b\right|$ are. The (positive) $\texttt{constant}$ does not depend on $\left(n,\,p,\,k,\,\beta_{K}^{*},\,\gamma_{K}^{*},\,\alpha^{*}\right)$, and bears little meaning in the asymptotic framework which simply assumes $\left(k\log p\right)/\sqrt{n}\rightarrow0$ among other sufficient conditions. [The existing theoretical framework makes it difficult to derive an informative $\texttt{constant}$, and to our knowledge, the literature provides no such derivation.] The asymptotic upper bound $\texttt{constant}\cdot\left(k\log p\right)/n$ does not distinguish cases that vary in $\left(\beta_{K}^{*},\,\gamma_{K}^{*},\,\alpha^{*}\right)$. By contrast, our lower bound analyses are informative about whether the upper bound can be attained by the magnitude of the OVBs and provide explicit constants. These features of our analysis are crucial for understanding the finite sample limitations of post double Lasso. In view of Proposition (ref), $\underline{OVB}$ is not a simple linear function of $\left(k\log p\right)/n$ in general, but roughly linear in $\left(k\log p\right)/n$ when $\phi^{-2}n^{-1}k\log p=o\left(1\right)$ and $k\exp\left(-\left(4\phi^{2}\right)^{-1}b^{2}\left(1+\tau\right)\log p\right)+2/p^{\tau}=o\left(1\right)$.
The finite sample behavior of post double Lasso can be characterized by three regimes: (i) non-negligible OVBs, (ii) negligible OVBs, and (iii) absence of OVBs.
Regime (i) (non-negligible OVB). When double under-selection occurs with high probability and $\left(k\log p\right)/\sqrt{n}$ is not small enough, according to Proposition (ref), the OVB lower bound can be substantial compared to the standard deviation obtained from the asymptotic distribution in belloni2014inference. To gauge the magnitude of the OVB and explain why the confidence intervals proposed in the literature can exhibit under-coverage, it is instructive to compare $\underline{OVB}$ with $\ensuremath{\sigma_{\tilde{\alpha}}=n^{-1/2}\left(\sigma_{\eta}/\sigma_{v}\right)}$, the standard deviation (of $\tilde{\alpha}$) obtained from the asymptotic distribution in belloni2014inference.\footnote{
} Let us consider an example with $n=14238$, $p=384$ angrist2019machine, $a=b=1$, $\sigma_{\eta}=\sigma_{v}=1$, $\tau=0.5$, and $\phi=0.5$. If $\left|\beta_{j}^{*}\right|=0.07$ and $\left|\gamma_{j}^{*}\right|=0.07$ for all $j\in K$, then $\underline{OVB}/\sigma_{\tilde{\alpha}}\approx0.27$ when $k=1$ and $\underline{OVB}/\sigma_{\tilde{\alpha}}\approx2.4$ when $k=10$; see Panel (c) of Figures (ref) and (ref) for an illustration of Regime (i). The OVBs can also be non-negligible when some but not all relevant controls are selected; see Panel (b) of Figures (ref) and (ref). These results suggest that, non-asymptotically, post double Lasso cannot avoid the “post-selection inference issues” raised in a series of papers by Leeb and P\"otscher leeb2005model,leeb2008can,leeb2017testing.
Regime (ii) (negligible OVB). By Lemma (ref) and ((ref)), $\left|\hat{\beta}-\beta^{*}\right|_{1}=2^{-1}\left|a\right|k\lambda_{1}$ and $\left|\hat{\gamma}-\gamma^{*}\right|_{1}=2^{-1}\left|b\right|k\lambda_{2}$ with probability at least $1-2/p^{\tau}$. If $\sigma_{v}$ is bounded away from zero, $\sigma_{\eta}$ is bounded from above, and $\left|a\right|=\left|b\right|=o\left(1\right)$, by a similar argument as in belloni2014inference, we can show that $\sqrt{n}\left(\tilde{\alpha}-\alpha^{*}\right)$ is approximately normal and centered at zero, even if $\left(k\log p\right)/\sqrt{n}$ is bounded away from zero and scales as a constant. Holding other factors constant, the magnitude of OVBs decreases as $\left|\beta_{K}^{*}\right|$ and $\left|\gamma_{K}^{*}\right|$ decrease (i.e., as $\left|a\right|$ and $\left|b\right|$ decrease). As $\left|a\right|$ and $\left|b\right|$ become very small, the relevant controls become essentially irrelevant. Panel (d) of Figures (ref) and (ref) provides an illustration of Regime (ii).
Regime (iii) (absence of OVB). All relevant controls will be selected when the magnitude of their coefficients is large enough. Specifically, for all $j\in K,\;a>3,\,b>3$, if
then $\mathbb{P}\left[\textrm{supp}\left(\hat{\pi}\right)=\textrm{supp}\left(\pi^{*}\right)\right]\geq1-1/p^{\tau}$ or $\mathbb{P}\left[\textrm{supp}\left(\hat{\gamma}\right)=\textrm{supp}\left(\gamma^{*}\right)\right]\geq1-1/p^{\tau}$ by standard arguments wainwright_2019. As a result, $\mathbb{P}\left(\left\{ \hat{I}_{1}\cup\hat{I}_{2}\right\} =K\right)\geq1-1/p^{\tau}$ (where $\hat{I}_{1}$ and $\hat{I}_{2}$ are defined in Section (ref)); i.e., the final OLS step ((ref)) includes all the relevant controls with high probability. By similar argument as in Appendix (ref), on the high probability event $\left\{ \left\{ \hat{I}_{1}\cup\hat{I}_{2}\right\} =K\right\} $, the OVB of $\tilde{\alpha}$ is zero. Panel (a) of Figures (ref) and (ref) provides an illustration of Regime (iii).
This “triple-regime” characterization suggests that one may assess the robustness of post double Lasso by increasing the penalty level. If increasing $\lambda_{1}$ and $\lambda_{2}$ yields similar estimates $\tilde{\alpha}$, then the underlying model could be in the regime where either the OVBs are negligible (Regime (ii)) or under-selection in both Lasso steps is unlikely (Regime (iii)). By contrast, under Regime (i), the performance of post double Lasso can be quite sensitive to an increase of $\lambda_{1}$ and $\lambda_{2}$. The rationale behind this heuristic lies in that the final step of post double Lasso, ((ref)), is simply an OLS regression of $Y$ on $D$ and the union of selected controls from ((ref))--((ref)). A natural question is by how much $\lambda_{1}$ and $\lambda_{2}$ should be increased for the robustness checks. For the regularization choice proposed in belloni2014inference, we will show in the simulations of Section (ref) that an increase by $50\%$ works well in practice.
Finally, our theoretical results have interesting implications for the comparison between post double Lasso and post single Lasso, where the latter only relies on one Lasso step to select the relevant controls. Note that the magnitude of the OVB of the post single Lasso estimator of $\alpha^{*}$ also falls into three regimes: (i) non-negligible OVB, (ii) negligible OVB, and (iii) absence of OVB. Thus, qualitatively, the OVBs of post single Lasso and post double Lasso have a similar behavior in finite samples. However, quantitatively, the magnitude of the OVBs can be much larger for the post single Lasso than for the post double Lasso, as illustrated by the following example. If $\beta_{j}^{*}=a\phi^{-1}\sigma_{\eta}\sqrt{2n^{-1}\left(1+\tau\right)\log p}$ with $\left|a\right|\in(0,\,1]$, and $\left|\gamma_{j}^{*}\right|\geq1$, then the argument for showing Proposition (ref) in Appendix (ref) implies that the OVB lower bound scales roughly as $\left|a\right|k\sqrt{\left(\log p\right)/n}$ for the post single Lasso estimator of $\alpha^{*}$ (and note that $\left(\left|a\right|k\sqrt{\left(\log p\right)/n}\right)/\left(1/\sqrt{n}\right)=\left|a\right|k\sqrt{\log p}$).
In this section, we briefly summarize the additional theoretical results that are provided in the appendix. First, we also consider cases where $\alpha^{*}\neq0$. The conditions required to derive the explicit formula are difficult to interpret when $\alpha^{*}\ne0$. However, it is possible to provide easy-to-interpret scaling results (without explicit constants) for cases where $\alpha^{*}\ne0$. Roughly, the scaling of our OVB lower bound can be as large as
and
These results reveal an interesting feature of the post double Lasso. The scaling of the OVB lower bound depends on $\left|\alpha^{*}\right|$ when the relevant controls are not selected. This is because the error $u$ in the reduced form equation (ref) involves $\alpha^{*}$ such that the choice of $\lambda_{1}$ in (ref) depends on $\left|\alpha^{*}\right|$ via the variance of $u$. By contrast, it is well-known that the OVB of OLS does not depend on $\left|\alpha^{*}\right|$ when relevant controls are omitted. Interested readers are referred to Propositions (ref) and (ref) in Appendix (ref) for details.
Second, we also provide upper bounds on the OVB. We have seen in equation (ref) that the lower bound on the OVBs scales as $ \left(\sigma_{\eta}/\sigma_{v}\right)\vee\left|\alpha^{*}\right|$ when $\left(k\log p\right)/n\rightarrow\infty$. Interestingly enough, we can also show that the upper bound on the OVB scales as in equation (ref) despite $\left(k\log p\right)/n\rightarrow\infty$ and the Lasso being inconsistent in the sense $\sqrt{n^{-1}\sum_{i=1}^{n}\left(X_{i}\hat{\pi}-X_{i}\pi^{*}\right)^{2}}\rightarrow\infty$, $\sqrt{n^{-1}\sum_{i=1}^{n}\left(X_{i}\hat{\gamma}-X_{i}\gamma^{*}\right)^{2}}\rightarrow\infty$ with high probability. Interested readers are referred to Propositions (ref) and (ref) in Appendix (ref) for details.
To better understand the practical implications of the OVB of post double Lasso, in this section, we present the results from simulations and two empirical applications with widely-used regularization choices available in standard software packages. The analyses were carried out using Matlab MATLAB2020, R R2021, and Stata stata2021.
In this section, we present simulation evidence on the performance of post double Lasso with three choices of the regularization parameter: (i) the heteroscedasticity-robust proposal of belloni2012sparse,belloni2014inference ($\lambda_{\text{BCCH}}$) implemented using the R-package hdm with the double selection option hdm2016, (ii) the regularization parameter with the minimum cross-validated error ($\lambda_{\text{min}}$) implemented using the R-package glmnet glmnet, and (iii) the regularization parameter corresponding to the minimum plus one standard deviation cross-validated error ($\lambda_{\text{1se}}$) also implemented using glmnet. We use the same type of regularization parameter choice in both Lasso steps.
We simulate data according to the DGP of Section (ref). To illustrate the role of the sample size $n$ and the sparsity parameter $k$, we consider (i) $(n,p,k)=(500,200,5)$, (ii) $(n,p,k)=(1000,200,5)$, and (iii) $(n,p,k)=(500,200,10)$. We show results for $R^2\in \{0.01,0.05,0.1,0.2,0.3,0.4,0.5\}$ based on 1,000 simulation repetitions. Appendix (ref) presents additional simulation evidence, where we vary $n$, the distribution of $(X_i,\eta_i,v_i)$, the true value $\alpha^\ast$, and also consider a heteroscedastic DGP.
We start by investigating the selection performance of the two Lasso steps of post double Lasso. Panel (a) of Figure (ref) displays the average number of selected controls (i.e., the cardinality of $\hat{I}_{1}\cup\hat{I}_{2}$ in (ref)) as a function of $R^2$. Lasso with $\lambda_{\text{BCCH}}$ selects the lowest number of controls. Choosing $\lambda_{\text{1se}}$ leads to a somewhat higher number of selected controls and results in moderate over-selection for larger values of $R^2$. Lasso with $\lambda_{\text{min}}$ selects the highest number of controls and exhibits substantial over-selection. Panel (b) shows the corresponding average numbers of selected relevant controls. [FIGURE (ref) HERE.]
Figure (ref) presents evidence on the bias of post double Lasso. To make the results easier to interpret, we report the ratio of the bias to the empirical standard deviation. Post double Lasso with $\lambda_{\text{BCCH}}$ can exhibit biases that are more than two times larger than the standard deviation when $(n,p,k)=(500,200,10)$. The bias can still be comparable to the standard deviation when $(n,p,k)=(1000,200,5)$. Consistent with our theoretical discussions in Sections (ref) and (ref), the relationship between $R^2$ and the ratio of bias to standard deviation is non-monotonic: it is increasing for small $R^2$ and decreasing for larger $R^2$. The bias is somewhat smaller for $\lambda=\lambda_{\text{1se}}$. Setting $\lambda=\lambda_{\text{min}}$ yields the smallest ratio of bias to standard deviation. Finally, we note that when $R^2$ is large enough such that there is no under-selection, post double Lasso performs well and is approximately unbiased for all regularization parameters. [FIGURE (ref) HERE.]
The additional simulation evidence reported in Appendix (ref) confirms these results but further shows that $\alpha^\ast$ is an important determinant of the performance of post double Lasso because of its direct effect on the magnitude of the coefficients and the error variance in the reduced form equation (ref). Moreover, we show that, while choosing $\lambda=\lambda_{\text{min}}$ works well when $\alpha^\ast=0$, this choice can yield poor performances when $\alpha^\ast\ne 0$ (see Figure (ref)). A similar phenomenon arises when using $0.5\lambda_{\text{BCCH}}$ instead of $\lambda_{\text{BCCH}}$: this choice works well when $\alpha^\ast=0$ (see Figure (ref) below), but yields biases when $\alpha^\ast\ne 0$. We found that, under our DGPs, this is related to the fact that when $\alpha^\ast\ne 0$, (ref) and (ref) differ with respect to the underlying coefficients and noise variances, which leads to differences in the (over-)selection behavior of the Lasso. Thus, there is no simple recommendation for how to choose the regularization parameters in practice.
The substantive performance differences between the three regularization choices suggest that post double Lasso is sensitive to the penalty levels in the intermediate case where $R^2$ is small enough so that under-selection occurs but large enough to cause substantial OVBs. To further investigate this issue, we compare the results for $\lambda_{\text{BCCH}}$, $0.5 \lambda_{\text{BCCH}}$, and $1.5 \lambda_{\text{BCCH}}$. Figure (ref) displays the average numbers of all selected controls (relevant or not) and selected relevant controls in both Lasso steps. The differences in the selection performance are substantial. Lasso with $0.5 \lambda_{\text{BCCH}}$ over-selects for all $n$ and $R^2$, while Lasso with $1.5 \lambda_{\text{BCCH}}$ under-selects unless $R^2$ and $n$ are large and $k=5$. The differences get smaller as $n$ increases and larger as $k$ increases. [FIGURE (ref) HERE.]
Figure (ref) displays the ratio of bias to standard deviation. Choosing $0.5 \lambda_{\text{BCCH}}$ yields small biases relative to the standard deviations for all $R^2$. By contrast, choosing $1.5 \lambda_{\text{BCCH}}$ yields biases that can be more than six times larger than the standard deviations when $(n,p,k)=(500,200,10)$ and still substantial when $(n,p,k)=(1000,200,5)$. For very small and large values of $R^2$, post double Lasso is less sensitive to the penalty level. In Section (ref), we discuss how to interpret and use robustness checks with respect to the regularization parameters in empirical applications. [FIGURE (ref) HERE.]
In sum, our simulation evidence shows that (i) under-selection can lead to large biases relative to the standard deviations, (ii) the performance of post double Lasso can be very sensitive to the choice of regularization parameters, and (iii) there is no simple recommendation for how to choose the regularization parameters in practice.
We revisit the analysis of the causal effect of eligibility for 401(k) plans ($D$) on total wealth ($Y$).\footnote{
} We use the data on $n=9915$ households from the 1991 SIPP belloni2017data analyzed by belloni2017program and chernozhukov2018double with high-dimensional methods. We consider two different specifications of the control variables ($X$).
Table (ref) presents post double Lasso estimates based on the whole sample with $\lambda_{\text{BCCH}}$, $0.5\lambda_{\text{BCCH}}$, and $1.5\lambda_{\text{BCCH}}$. For comparison, we also report OLS estimates with and without controls. For both specifications, the results are qualitatively similar across the different regularization choices and similar to OLS with all controls. This is possible as $n$ is much larger than $p$. Nevertheless, there are some non-negligible quantitative differences between the point estimates. A comparison to OLS without control variables shows that omitting controls can yield substantial OVBs in this application.[TABLE (ref) HERE.]
To investigate the impact of under-selection, we perform the following exercise. We draw random subsamples of size $n_s\in \{200,400,800,1600\}$ with replacement from the original dataset. Based on each subsample, we estimate $\alpha^\ast$ using post double Lasso with $\lambda_{\text{BCCH}}$, $0.5\lambda_{\text{BCCH}}$, and $1.5\lambda_{\text{BCCH}}$ and compute the bias as the difference between the average subsample estimate and the point estimate based on the original data with the same type of regularization choice in Table (ref). The results are based on 1,000 simulation repetitions.
Figures (ref) displays the bias and the ratio of bias to standard deviation for both specifications. We find that post double Lasso can exhibit large finite sample biases. The biases under the QSI specification tend to be smaller (in absolute value) than the biases under the TWI specification. Interestingly, the ratio of bias to standard deviation may not be monotonically decreasing in $n_s$ (in absolute value) due to the standard deviation decaying faster than the bias. Finally, we find that post double Lasso can be very sensitive to the penalty level. [FIGURE (ref) HERE.]
We revisit fryerlevitt2013's analysis of the racial differences in the mental ability of young children based on data from the US Collaborative Perinatal Project fryer2013data. As in the reanalysis of chernozhukov2020generic, we restrict the sample to Black and White children so that our final sample includes $n=30002$ observations. We focus on the standardized test score in the Wechsler Intelligence Test at the age of seven as our outcome variable ($Y$). The variable of interest ($D$) is an indicator for Black children. We use the same specification as in fryerlevitt2013, excluding interviewer fixed effects. The control variables ($X$) include extensive information on socio-demographic characteristics, the home environment, and the prenatal environment; see their Table 1B for descriptive statistics. After removing collinear terms there are $p=78$ controls.
Table (ref) shows the results for post double Lasso with $\lambda_{\text{BCCH}}$, $0.5\lambda_{\text{BCCH}}$, and $1.5\lambda_{\text{BCCH}}$, as well as OLS with and without controls based on the whole sample. Since $n=30002$ is much larger than $p=78$, all methods except for OLS without controls yield similar results. [TABLE (ref) HERE.]
To investigate the impact of under-selection, we draw random subsamples of size $n_s\in \{200,400,800,1600\}$ with replacement from the original dataset. In each sample, we estimate $\alpha^\ast$ using post double Lasso with $\lambda_{\text{BCCH}}$, $0.5\lambda_{\text{BCCH}}$, and $1.5\lambda_{\text{BCCH}}$ and compute the bias as the difference between the average estimate based on the subsamples and the estimate based on the original data with the same type of regularization choice. The results are based on 1,000 simulation repetitions.
Figure (ref) displays the bias and the ratio of bias to standard deviation. While the magnitude of the bias is decreasing in $n_s$, it can be substantial and larger than the standard deviation when $n_s$ is small. Moreover, the performance of post double Lasso is very sensitive to the choice of the regularization parameters. With $0.5\lambda_{\text{BCCH}}$, post double Lasso is approximately unbiased for all $n_s$, whereas, with $1.5\lambda_{\text{BCCH}}$, the bias is comparable to the standard deviation even when $n_s=1600$. [FIGURE (ref) HERE.]
The OVBs have important consequences for making inferences based on post double Lasso. Figure (ref) displays the coverage rates of 90% confidence intervals based on the DGPs in Section (ref) and shows that the OVB of post double Lasso can cause substantial under-coverage even when $n=1000$ and $k=5$. The most important determinant of the under-coverage is the sparsity parameter $k$. Our results show that the requirement on $k$ for guaranteeing a good finite sample coverage accuracy for all $R^2$ can be quite stringent. [FIGURE (ref) HERE.]
These results prompt the question of how to make inference in a reliable manner when one is concerned about OVBs. In many economic applications, $p$ is comparable to but still smaller than $n$. In such settings, OLS-based inference procedures provide a natural alternative to Lasso-based methods. Under classical conditions, OLS is the best linear unbiased estimator and admits exact finite sample inference as long as $p+1\leq n$ (recalling that the number of regression coefficients is $p+1$ in (ref)). Unlike the Lasso-based inference methods, OLS does not rely on any sparsity assumptions. This is important because sparsity assumptions may not be satisfied in applications and, as we show in this paper, the OVBs of Lasso-based inference procedures can be substantial even when $k$ is small and $n$ is large and larger than $p$. Indeed, OLS-based inference exhibits desirable optimality properties absent sparsity (or other restrictions) on $\beta^\ast$.\footnote{
}
While OLS is unbiased, constructing standard errors is challenging when $p$ is large. For instance, cattaneo2018inference show that conventional Eicker-White robust standard errors are inconsistent under asymptotics where $p$ grows as fast as $n$. This result motivates a recent literature to develop high-dimensional OLS-based inference procedures that are valid in settings with many controls cattaneo2018inference,jochmans2020heteroscedasticity,kline2020leave.
Figures (ref) compares the finite sample performance of post double Lasso and OLS with the heteroscedasticity robust HCK standard errors proposed by cattaneo2018inference. Panel (a) shows that OLS exhibits close-to-exact empirical coverage rates irrespective of the magnitude of the coefficients and the implied $R^2$. The additional simulation evidence in Appendix (ref) confirms the excellent performance of OLS with HCK standard errors. Panel (b) displays the average length of 90% confidence intervals and shows that OLS yields somewhat wider confidence intervals than post double Lasso. [FIGURE (ref) HERE.]
In sum, our simulation results suggest that modern OLS-based inference methods that accommodate many controls may constitute a viable alternative to Lasso-based inference methods. These methods are unbiased and demonstrate an excellent size accuracy, irrespective of the magnitude of the coefficients corresponding to the relevant controls. However, there is a trade-off because OLS yields somewhat wider confidence intervals than post double Lasso.
Finally, it is worth noting that we consider settings where one can easily invert $X^T X$ and the OLS and HCK variance estimators are numerically stable. In the case of singular or nearly singular $X^T X$, regularization is often unavoidable; see Section (ref) for a discussion of alternatives to OLS and Lasso-based inference methods.
Here we summarize the practical implications of our results and provide guidance for empirical researchers.
First, the simulation evidence in Section (ref) and Appendix (ref) along with the theoretical results (see Section (ref)) suggest the following heuristic: if the estimates of $\alpha^{*}$ are robust to increasing the theoretically recommended regularization parameters in the two Lasso steps, post double Lasso could be a reliable and efficient method. Therefore, we recommend to always check whether empirical results are robust to increasing the regularization parameters. Based on our simulations, a simple rule of thumb is to increase by $50\%$ the regularization parameters proposed in belloni2014inference. Robustness checks are standard in other contexts (e.g., bandwidth choices in regression discontinuity designs), and our results highlight the importance of such checks in the context of Lasso-based inference methods.
Second, following belloni2014inference, we recommend to always augment the union of selected controls with an “amelioration” set of controls motivated by economic theory and prior knowledge to mitigate the OVBs.
Third, our simulations show that in moderately high-dimensional settings where $p$ is comparable to but smaller than $n$, recently developed OLS-based inference methods that are robust to the inclusion of many controls exhibit better size properties. These simulation results suggest that high-dimensional OLS-based procedures constitute a possible alternative to Lasso-based inference methods.
Forth, OLS-based methods are not applicable when $p>n$, and the OLS and variance estimators can be numerically unstable under severe multi-collinearity even if $p<n$. In such cases, regularization is often needed. Ridge regressions, which impose restrictions on the Euclidean norm of $\beta^\ast$, avoid variable selection and may be a useful alternative to the Lasso; see also armstrong2020bias for related restrictions on $\beta^\ast$.
Finally, in many economic applications, researchers start with a small number of raw controls and want to use a flexible non-parametric model to capture the dependence of outcomes on controls while maintaining a simple parametric form for modeling the variables of interest. Such a specification leads to the classical partially linear models. In fact, belloni2014inference motivate post double Lasso with these models. If one is concerned about OVBs, inference methods that do not rely on variable selection are natural alternatives to post double Lasso. Under suitable smoothness restrictions on the non-parametric component, inference on the parameter of interest in partially linear models is a well-studied problem robinson1988root,newey1994asymptotic. The frameworks proposed in these papers can be built upon procedures such as sieves chen2007chapter16, local non-parametric methods fan1996local, and kernel ridge regressions scholkopf2002learning.
Given the rapidly increasing popularity of Lasso and Lasso-based inference methods in empirical economic research, it is crucial to better understand the merits and limitations of these new tools, and how they compare to other alternatives such as the high-dimensional OLS-based procedures.
This paper presents theoretical results as well as simulation and empirical evidence on the finite sample behavior of post double Lasso and the debiased Lasso (in the appendix). Specifically, we analyze the finite sample OVBs arising from the Lasso not selecting all the relevant control variables. Our results have important practical implications, and we provide guidance for empirical researchers.
We focus on the implications of under-selection for post double Lasso and the debiased Lasso in linear regression models. However, our results on the under-selection of the Lasso also have important implications for other inference methods that rely on Lasso as a first-step estimator. Towards this end, an interesting avenue for future research would be to investigate the impact of under-selection on the performance of the Lasso-based approaches proposed by belloni2014inference, farrell2015robust, belloni2017program, and chernozhukov2018double for non-linear models. In moderately high-dimensional settings where $p$ is smaller than but comparable to $n$, it would also be interesting to compare the treatment effects estimators in belloni2014inference to the robust finite sample methods proposed by rothe2017robust.
Finally, this paper motivates further examinations of the practical usefulness of Lasso-based inference procedures and other modern high-dimensional methods. For example, angrist2019machine present interesting simulation evidence on the finite sample behavior of Lasso-based IV methods belloni2012sparse. It would be interesting to explore the implications of our theoretical results on the under-selection of the Lasso in problems with weak instruments.