EconBase
← Back to paper

Omitted variable bias of Lasso-based inference methods: A finite sample analysis

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

78,455 characters · 19 sections · 105 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Omitted variable bias of Lasso-based inference methods: A finite sample analysis

abstractWe study the finite sample behavior of Lasso-based inference methods such as post double Lasso and debiased Lasso. We show that these methods can exhibit substantial omitted variable biases (OVBs) due to Lasso not selecting relevant controls. This phenomenon can occur even when the coefficients are sparse and the sample size is large and larger than the number of controls. Therefore, relying on the existing asymptotic inference theory can be problematic in empirical applications. We compare the Lasso-based inference methods to modern high-dimensional OLS-based methods and provide practical guidance. \noindentKeywords: Lasso, post double Lasso, debiased Lasso, OLS, omitted variable bias, size distortions, finite sample analysis \noindentJEL codes: C21, C52, C55

Introduction

Since their introduction, post double Lasso belloni2014inference and debiased Lasso javanmard2014confidence,vandergeer2014asymptotically,Zhang_Zhang have quickly become the most popular inference methods for problems with many control variables. Given the rapidly growing theoretical\footnote{

doublespaceSee, for example, farrell2015robust, belloni2017program, zhang2017simultaneous, chernozhukov2018double, caner2018asymptotically among others.

\singlespacing} and applied\footnote{

doublespaceSee, for example, chen2015can, decker2016health, schmitz2017informal, breza2019social, jones2019what, cole2020mobilizing, and enke2020moral among others.

} literature on these methods, it is crucial to take a step back and examine the performance of these procedures in empirically relevant settings and to better understand their merits and limitations relative to other alternatives.

In this paper, we study the performance of post double Lasso and debiased Lasso from two different angles. First, we develop a theory of the finite sample behavior of post double Lasso and debiased Lasso. We show that under-selection of the Lasso can lead to substantial omitted variable biases (OVBs) in finite samples. Our theoretical results on the OVBs are non-asymptotic; they complement but do not contradict the existing asymptotic theory. Second, we conduct extensive simulations and two empirical studies to investigate the performance of Lasso-based inference methods in practically relevant settings. We find that the OVBs can render the existing asymptotic approximations inaccurate and lead to size distortions.

Following the literature, we consider the following standard linear model

eqnarray[eqnarray omitted — 175 chars of source]

Here $Y_{i}$ is the outcome, $D_{i}$ is the scalar treatment variable of interest, and $X_{i}$ is a $(1\times p)$-dimensional vector of control variables. In the main text, we focus on the performance of post double Lasso for estimating and making inferences (e.g., constructing confidence intervals) on the treatment effect $\alpha^{\ast}$ in settings where $p$ can be larger than or comparable to $n$. We present results for debiased Lasso in Appendix (ref).

Post double Lasso consists of two Lasso selection steps: a Lasso regression of $Y_{i}$ on $X_{i}$ and a Lasso regression of $D_{i}$ on $X_{i}$. In the third step, the estimator of $\alpha^{\ast}$, $\tilde{\alpha}$, is the OLS regression of $Y_{i}$ on $D_{i}$ and the union of controls selected in the two Lasso steps. OVB arises in post double Lasso whenever the relevant controls (i.e., the controls with non-zero coefficients) are selected in neither Lasso steps, a situation we refer to as double under-selection. Results that can explain when and why double under-selection occurs are scarce in the existing literature. This paper shows theoretically that this phenomenon can even occur in simple examples with classical assumptions (e.g., normal homoscedastic errors, orthogonal designs for the relevant controls), which are often viewed as favorable to the performance of the Lasso. We prove that if the products of the absolute values of the non-zero coefficients and the variances of the controls are no greater than half the regularization parameters derived based on standard Lasso theory\footnote{

doublespaceNote that the existing Lasso theory requires the regularization parameter to exceed a certain threshold, which depends on the standard deviations of the noise and the covariates.

\singlespacing}, Lasso fails to select these controls in both steps with high probability.\footnote{

doublespaceThe “half the regularization parameters" type of condition on the magnitude of non-zero coefficients was independently discovered in lahiri2021necessary. We are grateful to an anonymous referee for making us aware of this paper. Our proof strategies differ from the asymptotic ones in lahiri2021necessary and allow us to derive an explicit lower bound with meaningful constants for the probability of under-selection for fixed $n$, which is needed for deriving an explicit formula for the OVB lower bound. While some of the arguments in lahiri2021necessary can be made non-asymptotic, one of their core arguments for showing necessary conditions for variable selection consistency of the Lasso relies on $n$ tending to infinity. It is not clear that such an argument can lead to an explicit lower bound with meaningful constants for the probability of under-selection. On the other hand, our non-asympototic argument can easily lead to asymptotic conclusions.

\singlespacing} This result allows us to derive the first non-asymptotic lower bound formula in the literature for the OVB of the post double Lasso estimator $\tilde{\alpha}$. Our lower bound provides explicit universal constants, which are essential for understanding the finite sample behavior of post double Lasso and its limitations.

The OVB lower bound is characterized by the interplay between the probability of double under-selection and the magnitude of the coefficients corresponding to the relevant controls in (ref)--(ref). In particular, there are three regimes: (i) the omitted controls have coefficients whose magnitudes are large enough to cause substantial bias; (ii) the omitted controls have coefficients whose magnitudes are too small to cause substantial bias; (iii) the magnitude of the relevant coefficients is large enough such that the corresponding controls are selected with high-probability. We illustrate these three regimes in Figure (ref), which plots the bias of post double Lasso and the number of selected control variables as a function of the magnitude of the non-zero coefficients. [FIGURE (ref) HERE.]

Our theoretical analysis of the OVB has important implications for inference procedures based on the post double Lasso. belloni2014inference show that $\sqrt{n}(\tilde\alpha-\alpha^\ast)$ is asymptotically normal with zero mean. We show that in finite samples, the OVB lower bound can be more than twice as large as the standard deviation obtained from the asymptotic distribution in belloni2014inference. This is true even when $n$ is much larger than $p$ and $\beta^\ast$ and $\gamma^\ast$ are sparse. To illustrate, assume that (ref) and (ref) share the same set of $k$ non-zero coefficients and set $(n,\,p)=(14238,\,384)$ as in angrist2019machine, who use post double Lasso to estimate the effect of elite colleges. The ratio of our OVB lower bound to the standard deviation in belloni2014inference is $0.27$ if $k=1$ and $2.4$ if $k=10$. This example shows that the requirement on the sparsity parameter $k$ for the OVBs to be negligible is quite stringent. We emphasize that our findings do not contradict the existing results on the asymptotic distribution of post double Lasso in belloni2014inference. Rather, our results suggest that the OVBs can make the asymptotic zero-mean approximation of $\sqrt{n}(\tilde\alpha-\alpha^\ast)$ inaccurate in finite samples.

To better understand the practical implications of the OVB of post double Lasso, we perform extensive simulations. Our simulation results can be summarized as follows. (i) Large OVBs are persistent across a range of empirically relevant settings and can occur even when $n$ is large and larger than $p$, and the sparsity parameter $k$ is small. (ii) The OVBs can lead to invalid inferences and under-coverage of confidence intervals. (iii) The performance of post double Lasso varies substantially across different popular choices of regularization parameters, and no single choice outperforms the others across all designs. While it may be tempting to choose a smaller regularization parameter than the standard recommendation in the literature to mitigate under-selection, we find that this idea does not work in general and can lead to rather poor performance.

In addition to the simulations, we consider two empirical applications: the analysis of the effect of 401(k) plans on savings by belloni2017program and chernozhukov2018double and the study of the racial test score gap by fryerlevitt2013. We draw samples of different sizes from the large original data and compare the subsample estimates to the estimates based on the original data.\footnote{

doublespaceFor example, kolesar2018inference use a similar of exercise to illustrate the issues with discrete running variables in regression discontinuity designs.

} In both applications, we find substantial biases even when $n$ is considerably larger than $p$, and we document that the magnitude of the biases varies substantially depending on the regularization choice.

Given our theoretical results, simulations, and empirical evidence, a natural question is how to make statistical inferences in a reliable manner if one is concerned about OVBs. In many economic applications, $p$ is comparable to but still smaller than $n$. This motivates the recent development of high-dimensional OLS-based inference procedures cattaneo2018inference,dadamo2018cluster,jochmans2020heteroscedasticity,kline2020leave. These methods are based on OLS regressions with all controls and rely on novel variance estimators that are robust to the inclusion of many controls (unlike conventional variance estimators). Based on extensive simulations, we find that OLS with standard errors proposed by cattaneo2018inference demonstrates excellent coverage accuracy across all our simulation designs. Another advantage of OLS-based methods over Lasso-based inference methods is that the former do not rely on any sparsity assumptions. This is important because sparsity assumptions may not be satisfied in applications and, as this paper shows, the OVBs of Lasso-based inference procedures can be substantial even when $k$ is small and $n$ is large and larger than $p$. However, OLS yields somewhat wider confidence intervals than the Lasso-based inference methods, suggesting a trade-off between coverage accuracy and the length of the confidence intervals.

Our analyses suggest two main recommendations concerning the use of post double Lasso in empirical studies. First, if the estimates of $\alpha^{*}$ are robust to increasing the recommended regularization parameters in both Lasso steps, this suggests that either the OVBs are negligible (Regime (ii)) or under-selection is unlikely (Regime (iii)). In either case, post double Lasso is a reliable and efficient method. Otherwise, modern high-dimensional OLS-based inference methods constitute a possible alternative when $p$ is smaller than $n$. Second, our findings highlight the importance of augmenting the final OLS regression in post double Lasso with control variables motivated by economic theory and prior knowledge, as suggested by belloni2014inference.

Lasso and post double Lasso

The Lasso

Consider the following linear regression model

equation[equation omitted — 85 chars of source]

where $\left\{ Y_{i}\right\} _{i=1}^{n}=Y$ is an $n$-dimensional response vector, $\left\{ X_{i}\right\} _{i=1}^{n}=X$ is an $n\times p$ matrix of covariates with $X_{i}$ denoting the $i$th row of $X$, $\left\{ \varepsilon_{i}\right\} _{i=1}^{n}=\varepsilon$ is a zero-mean error vector, and $\theta^{*}$ is a $p$-dimensional vector of unknown coefficients.

The Lasso estimator of $\theta^{*}$, which was first proposed by tibsharini1996regression, is given by

equation[equation omitted — 185 chars of source]

where $\lambda$ is the regularization parameter. Let $\varepsilon\sim\mathcal{N}\left(0_{n},\,\sigma^{2}I_{n}\right)$ and $X$ be a fixed design matrix with normalized columns (i.e., $n^{-1}\sum_{i=1}^{n}X_{ij}^{2}=1$ for all $j=1,\dots,p$). In this example, bickel2009simultaneous set $\lambda=2\sigma\sqrt{2n^{-1}\left(1+\tau\right)\log p}$ (where $\tau>0$) to establish upper bounds on $\sqrt{\sum_{j=1}^{p}(\hat{\theta}_{j}-\theta_{j}^{*})^{2}}$ with a high probability guarantee. To establish perfect selection, wainwright2009sharp sets $\lambda$ proportional to $\sigma\phi^{-1}\sqrt{\left(\log p\right)/n}$, where $\phi\in(0,\,1]$ is a measure of correlation between the covariates with nonzero coefficients and those with zero coefficients.

Besides the classical choices in bickel2009simultaneous and wainwright2009sharp, other choices of $\lambda$ are available in the literature. For instance, belloni2012sparse and belloni2016cluster propose choices that accommodate heteroscedastic and clustered errors. The regularization choice of belloni2012sparse, which is recommended by belloni2014inference for post double Lasso, is based on the following Lasso program:

equation[equation omitted — 193 chars of source]

where $(\hat{l}_1,\dots,\hat{l}_p)$ are penalty loadings obtained using the iterative algorithm developed in belloni2012sparse.

Finally, a very popular practical approach for choosing $\lambda$ is cross-validation; see, for example, homrighausen2013thelasso,homrighausen2014leaveoneout and chetverikov2020cross for theoretical results on cross-validated Lasso.

Post double Lasso

The model (ref)\textendash (ref) implies the following reduced form model for $Y_{i}$:

eqnarray[eqnarray omitted — 66 chars of source]

where $\pi^{*}=\gamma^{*}\alpha^{*}+\beta^{*}$ and $u_{i}=\eta_{i}+\alpha^{*}v_{i}$.

The post double Lasso, introduced by belloni2014inference, essentially exploits the Frisch-Waugh theorem, where the regressions of $Y$ on $X$ and $D$ on $X$ are implemented with the Lasso:

eqnarray[eqnarray omitted — 385 chars of source]

The final estimator $\tilde{\alpha}$ of $\alpha^{*}$ is then obtained from an OLS regression of $Y$ on $D$ and the union of selected controls

equation[equation omitted — 302 chars of source]

where $\hat{I}_{1}=\textrm{supp}\left(\hat{\pi}\right)=\left\{ j:\,\hat{\pi}_{j}\neq0\right\} $ and $\hat{I}_{2}=\textrm{supp}\left(\hat{\gamma}\right)=\left\{ j:\,\hat{\gamma}_{j}\neq0\right\} $.

Numerical example

This section presents a simple numerical example illustrating the OVB of post double Lasso. All computations were performed in Matlab MATLAB2020. The Lasso is implemented using the built-in function lasso. We consider a simple but classical setting that is often considered favorable to the performance of the Lasso. The data are simulated according to the structural model (ref)--(ref), where $X_{i}\overset{iid}\sim \mathcal{N}\left(0_p,I_{p}\right)$, $\eta_{i}\overset{iid}\sim\mathcal{N}(0,1)$, and $v_{i}\overset{iid}\sim \mathcal{N}(0,1)$ are independent of each other. Our object of interest is $\alpha^{*}$. We set $n=500$, $p=200$, $\alpha^{*}=0$, and consider a sparse setting where $\beta_j^{*}=\gamma_j^{*}=c \cdot 1\left\{j\le k\right\}$ for $j=1,\dots,p$ and $k=5$. Following the simulation exercise in belloni2014inference, we vary the population $R^2$s in (ref) and (ref) by varying the magnitude of the non-zero coefficients $c$. We employ the regularization parameter choice by bickel2009simultaneous.

Figure (ref) displays the finite sample distribution of post double Lasso for different values of $R^2$. For comparison, we plot the distribution of the “oracle estimator” of $\alpha^\ast$, a regression of $Y_i-X_i\pi^\ast$ on $D_i-X_i\gamma^\ast$. The finite sample behavior of post double Lasso depends on how many of the $k=5$ relevant controls get selected in both Lasso steps. Figure (ref) shows histograms of the number of selected relevant controls. [FIGURES (ref) AND (ref) HERE.]

When $R^2=0.5$, post double Lasso exhibits an excellent performance. The finite sample distribution is well-approximated by the normal distribution of the oracle estimator and centered at $\alpha^\ast=0$. The reason for the excellent performance is that all $k=5$ controls are selected with high probability such that post double Lasso essentially coincides with an OLS regression of $Y_i$ on $D_i$ and the five relevant controls.

Let us now consider what happens if we decrease the magnitude of the coefficients and the implied $R^2$. For $R^2=0.3$, post double Lasso exhibits a large finite sample bias. Moreover, the distribution of post double Lasso differs substantially from the distribution of the oracle estimator: it has a larger standard deviation and is slightly skewed. This distribution is a mixture of the distributions of OLS conditional on the two Lasso steps selecting different combinations of relevant controls (Panel (b) in Figure (ref)). For $R^2=0.1$, post double Lasso again exhibits a significant bias. At the same time, the shape of the distribution is similar to that of the oracle estimator. This is because, with high probability, none of the controls gets selected, while the coefficients are large enough to cause an OVB. Decreasing the $R^2$ to $0.01$ reduces the bias. However, it does not change the shape of the finite sample distribution because the selection performance remains unchanged.

The simple numerical example in this section shows that post double Lasso can suffer from OVBs when the two Lasso steps do not select all relevant controls. The magnitude of the coefficients corresponding to the omitted controls can be large enough such that the OVB shifts the location of the finite sample distribution far away from the true value $\alpha^\ast=0$. The issue documented here is not a “small sample” phenomenon but persists even in large sample settings; see Appendix (ref).

Theoretical analysis

This section provides a theoretical analysis of the OVB of post double Lasso. Our goal here is to demonstrate that, even in simple examples with classical assumptions (e.g., normal homoscedastic errors, orthogonal designs for the relevant controls), which are often viewed favorable to the performance of Lasso, the finite sample OVBs of post double Lasso can be substantial relative to the standard deviation provided in the existing literature. We first establish a new necessary result for the Lasso's inclusion and then derive lower and upper bounds on the OVBs of post double Lasso. These results are derived for fixed $\left(n,\,p,\,k\right)$ and are also valid when $\left(k\log p\right)/n\rightarrow0$ or $\left(k\log p\right)/n\rightarrow\infty$. As it will become clear in the following, $p$ needs to be large enough for our results to be informative. Without loss of generality, we normalize the matrix $X$ such that $\left(X_{j}^{T}X_{j}\right)/n=1$ for all $j=1,\dots,p$. We focus on fixed designs (of $X$) to highlight the essence of the problem; see Appendix (ref) for an extension to random designs.

For the convenience of the reader, here we collect the notation to be used in the theoretical analysis. Let $1_{m}$ denote the $m-$dimensional (column) vector of “1”s and $0_{m}$ is defined similarly. The $\ell_{1}-$norm of a vector $v\in\mathbb{R}^{m}$ is denoted by $\left|v\right|_{1}:=\sum_{i=1}^{m}\left|v_{i}\right|$ and the $\ell_{\infty}-$norm of a vector $v\in\mathbb{R}^{m}$ is denoted by $\left|v\right|_{\infty}:=\max_{i=1,\dots,m}\left|v_{i}\right|$. The $\ell_{\infty}$ matrix norm (maximum absolute row sum) of a matrix $A$ is denoted by $\left\Vert A\right\Vert _{\infty}:=\max_{i}\sum_{j}\left|a_{ij}\right|$. For a vector $v\in\mathbb{R}^{m}$ and a set of indices $T\subseteq\left\{ 1,\dots,m\right\} $, let $v_{T}$ denote the sub-vector (with indices in $T$) of $v$. For a matrix $A\in\mathbb{R}^{n\times m}$, let $A_{T}$ denote the submatrix consisting of the columns with indices in $T$. For a vector $v\in\mathbb{R}^{m}$, let $\textrm{sgn}(v):=\left\{ \textrm{sgn}(v_{j})\right\} _{j=1,\dots,m}$ denote the sign vector such that $\textrm{sgn}(v_{j})=1$ if $v_{j}>0$, $\textrm{sgn}(v_{j})=-1$ if $v_{j}<0$, and $\textrm{sgn}(v_{j})=0$ if $v_{j}=0$. Given a set $K$, let $\textrm{card}\left(K\right)$ denote the cardinality of $K$. We denote $\max\left\{ a,\,b\right\} $ by $a\vee b$ and $\min\left\{ a,\,b\right\} $ by $a\wedge b$.

Stronger necessary results on the Lasso's inclusion

Post double Lasso exhibits OVBs whenever the relevant controls are selected in neither (ref) nor (ref). To the best of our knowledge, there are no formal results strong enough to show that, with high probability, Lasso can fail to select the relevant controls in both steps. Therefore, we first establish a new necessary result for the single Lasso's inclusion in Lemma (ref). To derive this result, we consider the following classical assumptions, which are often viewed favorable to the performance of the Lasso.

assumptionIn terms of model (ref), suppose: (i) $K=\left\{ j:\,\theta_{j}^{*}\neq0\right\} \neq\emptyset$ and $\textrm{card}\left(K\right)= k\leq\left(n\wedge p\right)$; (ii) $X_{K}^{T}X_{K}$ is a diagonal matrix; (iii) $\left\Vert \left(X_{K^{c}}^{T}X_{K}\right)\left(X_{K}^{T}X_{K}\right)^{-1}\right\Vert _{\infty}=1-\phi$ for some $\phi\in(0,\,1]$, where $K^{c}$ is the complement of $K$.

Known as the incoherence condition due to wainwright2009sharp, part (iii) in Assumption (ref) is needed for the exclusion of the irrelevant controls. Note that if the columns in $X_{K^{c}}$ are orthogonal to the columns in $X_{K}$ (but within $X_{K^{c}}$, the columns need not be orthogonal to each other), then $\phi=1$. Obviously a special case of this is when the entire $X$ consists of mutually orthogonal columns (which is possible if $n\geq p$). To provide some intuition for Assumption (ref)(iii), let us consider the simple case where $k=1$ and $K=\left\{ 1\right\} $, $X$ is centered (such that $\left\{n^{-1} \sum_{i=1}^{n}X_{ij}\right\} _{j=1}^{p}=0_{p}$), and the columns in $X$ are normalized such that the standard deviations of $X_{1}$ and $X_{j}$ (for any $j\in\left\{ 2,3,\dots,p\right\} $) are identical. Then, $1-\phi$ is simply the maximum of the absolute (sample) correlations between $X_{1}$ and each of the $X_{j}$s with $j\in\left\{ 2,3,\dots,p\right\} $.

lemma[Necessary result on the Lasso's inclusion] In model (ref), suppose the $\varepsilon_{i}$s are independent over $i=1,\dots,n$ and $\varepsilon_{i}\sim\mathcal{N}\left(0,\sigma^{2}\right)$, where $\sigma\in\left(0,\,\infty\right)$. Let Assumption (ref) hold. We solve the Lasso ((ref)) with \begin{equation} \lambda\geq\frac{2\sigma}{\phi}\sqrt{\frac{2\left(1+\tau\right)\log p}{n}} \end{equation} where $\tau>0$. Let $E_{1}$ denote the event that $\textrm{sgn}\left(\hat{\theta}_{j}\right)=-\textrm{sgn}\left(\theta_{j}^{*}\right)$ for at least one $j\in K$, and $E_{2}$ denote the event that $\textrm{sgn}\left(\hat{\theta}_{l}\right)=\textrm{sgn}\left(\theta_{l}^{*}\right)$ for at least one $l\in K$ with \begin{equation} \left|\theta_{l}^{*}\right|\leq\frac{\lambda}{2}. \end{equation} Then, we have \begin{equation} \mathbb{P}\left(E_{1}\cap\mathcal{\mathcal{E}}\right)=\mathbb{P}\left(E_{2}\cap\mathcal{\mathcal{E}}\right)=0 \end{equation} where $\mathcal{E}=\left\{ \left|n^{-1}X^{T}\varepsilon\right|_{\infty}\leq\sigma\phi^{-1}\sqrt{2n^{-1}\left(1+\tau\right)\log p}\right\} $ and $\mathbb{P}\left(\mathcal{E}\right)\geq1-1/p^{\tau}$. If ((ref)) holds for all $l\in K$, we have \begin{equation} \mathbb{P}\left(\hat{\theta}=0_{p}\right)\geq1-\frac{1}{p^{\tau}}. \end{equation}

Lemma (ref) shows that for large enough $p$, Lasso fails to select any of the relevant covariates with high probability if ((ref)) holds for all $l\in K$. If such conditions hold with respect to both (ref) and (ref), then Lemma (ref) implies that the relevant controls are selected in neither (ref) nor (ref) with probability at least $1-2/p^{\tau}$; see Panels (c) and (d) of Figure (ref) for an illustration.

Let us rewrite ((ref)) as $\left|\theta_{l}^{*}\right|=2^{-1}\left|a\right|\lambda$ with $\left|a\right|\in(0,\,1]$. Assume that $\left|a\right|\phi^{-1}\sigma$ is bounded from above and away from zero; moreover, $\lambda$ satisfies (ref) and scales as $\sqrt{\left(\log p\right)/n}$. These conditions imply that $\left|\theta_{l}^{*}\right|\asymp\sqrt{1/n}$ in the classical asymptotic framework where $n\rightarrow \infty$ and $p$ is fixed. This regime of $\theta_{l}^{*}$ is exactly where classical model selection procedures struggle to distinguish a coefficient from zero in low-dimensional settings.

In the introduction, we have discussed the relationship of our result to that in lahiri2021necessary. It is also interesting to compare Lemma (ref) with the results in wainwright2009sharp. Note that ((ref)) implies $\mathbb{P}\left(\hat{\theta}_{l}\neq0\right)\leq1/p^{\tau}$ for any $l\in K$ subject to ((ref)). In comparison, wainwright2009sharp shows that whenever $\theta_{l}^{*}\in\left(\lambda\textrm{sgn}\left(\theta_{l}^{*}\right),\,0\right)$ or $\theta_{l}^{*}\in\left(0,\,\lambda\textrm{sgn}\left(\theta_{l}^{*}\right)\right)$ for some $l\in K$,

equation[equation omitted — 157 chars of source]

Constant bounds in the form of ((ref)) cannot explain that, with high probability, Lasso fails to select the relevant covariates in both (ref) and (ref) when $p$ is sufficiently large.

remarkAs the choices of regularization parameters used in the vast majority of literature bickel2009simultaneous,wainwright2009sharp,belloni2012sparse,belloni2013least,belloni2014inference, the choice of $\lambda$ in Lemma (ref) is derived from the principle that $\lambda$ should be no smaller than $2\max_{j=1,\dots,p}\left|n^{-1}X_{j}^{T}\varepsilon\right|$ with high probability. In particular, our choice for $\lambda$ takes the form of that in wainwright2009sharp, but ours involves a sharper universal constant. Choosing regularization parameters in this form ensures the exclusion of irrelevant controls. In addition, our choice for $\lambda$ has a scaling that can be achieved by the regularization parameters in belloni2012sparse, belloni2013least, belloni2014inference, and coincides with that in bickel2009simultaneous when the columns in $X_{K^{c}}$ are orthogonal to the columns in $X_{K}$ (but within $X_{K^{c}}$, the columns need not be orthogonal to each other).

Lower bounds on the OVBs

In this section, we apply Lemma (ref) to derive lower bounds on the OVB of post double Lasso. We consider the structural model (ref)--(ref), which can be written in matrix notation as

eqnarray[eqnarray omitted — 105 chars of source]

In matrix notation, the reduced form (ref) becomes

eqnarray[eqnarray omitted — 52 chars of source]

where $\pi^{*}=\gamma^{*}\alpha^{*}+\beta^{*}$ and $u=\eta+\alpha^{*}v$. We make the following assumptions about model ((ref))--((ref)).

assumption(i) The error terms $\eta$ and $v$ consist of independent entries drawn from $\mathcal{N}\left(0,\,\sigma_{\eta}^{2}\right)$ and $\mathcal{N}\left(0,\,\sigma_{v}^{2}\right)$, respectively, where $\eta$ and $v$ are independent of each other; (ii) the data are centered: $\bar{D}=n^{-1}\sum_{i=1}^{n}D_{i}=0$, $\bar{X}=\left\{n^{-1}\sum_{i=1}^{n}X_{ij}\right\} _{j=1}^{p}=0_{p}$, and $\bar{Y}=n^{-1}\sum_{i=1}^{n}Y_{i}=0$; (iii) $K=\left\{ j:\,\beta_{j}^{*}\neq0\right\} =\left\{ j:\,\gamma_{j}^{*}\neq0\right\} \neq\emptyset$ and $\textrm{card}\left(K\right)= k\leq\left(n\wedge p\right)$.

Proposition (ref) derives a lower bound formula for the OVB of post double Lasso concerning the case where $\alpha^{*}=0$.

proposition[OVB lower bound] Suppose $\alpha^{*}=0$. Let Assumption (ref)(ii)-(iii) and Assumption (ref) hold; $\lambda_{1}=2\phi^{-1}\sigma_{\eta}\sqrt{2n^{-1}\left(1+\tau\right)\log p}$ and $\lambda_{2}=2\phi^{-1}\sigma_{v}\sqrt{2n^{-1}\left(1+\tau\right)\log p}$; for all $j\in K$ and $\left|a\right|,\left|b\right|\in(0,\,1]$, \begin{equation} both\quad\beta_{j}^{*}=a\phi^{-1}\sigma_{\eta}\sqrt{\frac{2\left(1+\tau\right)\log p}{n}}\quad and\quad\gamma_{j}^{*}=b\phi^{-1}\sigma_{v}\sqrt{\frac{2\left(1+\tau\right)\log p}{n}}. \end{equation} In terms of $\tilde{\alpha}$ obtained from ((ref)), we have \[ \left|\mathbb{E}\left(\tilde{\alpha}-\alpha^{*}\vert\mathcal{M}\right)\right|\geq\underset{:=\underline{\text{OVB}}}{\underbrace{\max_{r\in(0,1]}T_{1}\left(r\right)T_{2}\left(r\right)}} \] where \begin{eqnarray*} T_{1}\left(r\right) & = & \frac{\left(1+\tau\right)\left|ab\right|\phi^{-2}\sigma_{\eta}\frac{k\log p}{n}}{4\left(1+\tau\right)\phi^{-2}b^{2}\sigma_{v}\frac{k\log p}{n}+\left(1+r\right)\sigma_{v}},\\ T_{2}\left(r\right) & = & 1-k\exp\left(\frac{-b^{2}\left(1+\tau\right)\log p}{4\phi^{2}}\right)-\frac{1}{p^{\tau}}-\exp\left(\frac{-nr^{2}}{8}\right), \end{eqnarray*} for any $r\in(0,\,1]$, and $\mathcal{M}$ is an event with $\mathbb{P}\left(\mathcal{M}\right)\geq1-k\exp\left(-\left(4\phi^{2}\right)^{-1}b^{2}\left(1+\tau\right)\log p\right)-2/p^{\tau}$.
remarkIn our theoretical results, we implicitly assume $p$ is sufficiently large such that $1-k\exp\left(-\left(4\phi^{2}\right)^{-1}b^{2}\left(1+\tau\right)\log p\right)-2/p^{\tau}>0$. Indeed, probabilities in such a form are often referred to as the “high-probability” guarantees in the literature of (non-asymptotic) high-dimensional statistics concerning large $p$ and small enough $k$. Recalling the definitions of $\hat{I}_{1}$ and $\hat{I}_{2}$ in Section (ref), the event $\mathcal{M}$ is the intersection of $\left\{ \hat{I}_{1}=\hat{I}_{2}=\emptyset\right\} $ and an additional event $\mathcal{E}_{t^{*}}=\left\{ \left|n^{-1}X_{K}^{T}v\right|_{\infty}\leq t^{*},\,t^{*}=4^{-1}\left|b\right|\lambda_{2}\right\} $. The event $\left\{ \hat{I}_{1}=\hat{I}_{2}=\emptyset\right\} $ occurs with probability at least $1-2/p^{\tau}$, and the event $\mathcal{E}_{t^{*}}$ occurs with probability at least $1-k\exp\left(-\left(4\phi^{2}\right)^{-1}b^{2}\left(1+\tau\right)\log p\right)$. The additional event $\mathcal{E}_{t^{*}}$ is needed for us to derive a non-trivial lower bound. In particular, with probability at most $k\exp\left(-\left(4\phi^{2}\right)^{-1}b^{2}\left(1+\tau\right)\log p\right)$, we have $\left|n^{-1}X_{K}^{T}v\right|_{\infty}\geq t^{*}$, and on this event, the lower bound in Proposition (ref) can be negative, which is uninformative for the absolute value of OVBs.

Let us compare the non-asymptotic lower bound in Proposition (ref) to the implications of the existing asymptotic results for the bias of post double Lasso. If $\sigma_{v}$ is bounded away from zero and $\sigma_{\eta}$ is bounded from above, the existing theory would imply that the biases of post double Lasso are bounded from above by $\texttt{constant}\cdot\left(k\log p\right)/n$, irrespective of whether Lasso fails to select the relevant controls or not, and how small $\left|a\right|$ and $\left|b\right|$ are. The (positive) $\texttt{constant}$ does not depend on $\left(n,\,p,\,k,\,\beta_{K}^{*},\,\gamma_{K}^{*},\,\alpha^{*}\right)$, and bears little meaning in the asymptotic framework which simply assumes $\left(k\log p\right)/\sqrt{n}\rightarrow0$ among other sufficient conditions. [The existing theoretical framework makes it difficult to derive an informative $\texttt{constant}$, and to our knowledge, the literature provides no such derivation.] The asymptotic upper bound $\texttt{constant}\cdot\left(k\log p\right)/n$ does not distinguish cases that vary in $\left(\beta_{K}^{*},\,\gamma_{K}^{*},\,\alpha^{*}\right)$. By contrast, our lower bound analyses are informative about whether the upper bound can be attained by the magnitude of the OVBs and provide explicit constants. These features of our analysis are crucial for understanding the finite sample limitations of post double Lasso. In view of Proposition (ref), $\underline{OVB}$ is not a simple linear function of $\left(k\log p\right)/n$ in general, but roughly linear in $\left(k\log p\right)/n$ when $\phi^{-2}n^{-1}k\log p=o\left(1\right)$ and $k\exp\left(-\left(4\phi^{2}\right)^{-1}b^{2}\left(1+\tau\right)\log p\right)+2/p^{\tau}=o\left(1\right)$.

Key takeaways of our theoretical results

The finite sample behavior of post double Lasso can be characterized by three regimes: (i) non-negligible OVBs, (ii) negligible OVBs, and (iii) absence of OVBs.

Regime (i) (non-negligible OVB). When double under-selection occurs with high probability and $\left(k\log p\right)/\sqrt{n}$ is not small enough, according to Proposition (ref), the OVB lower bound can be substantial compared to the standard deviation obtained from the asymptotic distribution in belloni2014inference. To gauge the magnitude of the OVB and explain why the confidence intervals proposed in the literature can exhibit under-coverage, it is instructive to compare $\underline{OVB}$ with $\ensuremath{\sigma_{\tilde{\alpha}}=n^{-1/2}\left(\sigma_{\eta}/\sigma_{v}\right)}$, the standard deviation (of $\tilde{\alpha}$) obtained from the asymptotic distribution in belloni2014inference.\footnote{

doublespaceWe thank Ulrich M\"uller for suggesting this comparison.

} Let us consider an example with $n=14238$, $p=384$ angrist2019machine, $a=b=1$, $\sigma_{\eta}=\sigma_{v}=1$, $\tau=0.5$, and $\phi=0.5$. If $\left|\beta_{j}^{*}\right|=0.07$ and $\left|\gamma_{j}^{*}\right|=0.07$ for all $j\in K$, then $\underline{OVB}/\sigma_{\tilde{\alpha}}\approx0.27$ when $k=1$ and $\underline{OVB}/\sigma_{\tilde{\alpha}}\approx2.4$ when $k=10$; see Panel (c) of Figures (ref) and (ref) for an illustration of Regime (i). The OVBs can also be non-negligible when some but not all relevant controls are selected; see Panel (b) of Figures (ref) and (ref). These results suggest that, non-asymptotically, post double Lasso cannot avoid the “post-selection inference issues” raised in a series of papers by Leeb and P\"otscher leeb2005model,leeb2008can,leeb2017testing.

Regime (ii) (negligible OVB). By Lemma (ref) and ((ref)), $\left|\hat{\beta}-\beta^{*}\right|_{1}=2^{-1}\left|a\right|k\lambda_{1}$ and $\left|\hat{\gamma}-\gamma^{*}\right|_{1}=2^{-1}\left|b\right|k\lambda_{2}$ with probability at least $1-2/p^{\tau}$. If $\sigma_{v}$ is bounded away from zero, $\sigma_{\eta}$ is bounded from above, and $\left|a\right|=\left|b\right|=o\left(1\right)$, by a similar argument as in belloni2014inference, we can show that $\sqrt{n}\left(\tilde{\alpha}-\alpha^{*}\right)$ is approximately normal and centered at zero, even if $\left(k\log p\right)/\sqrt{n}$ is bounded away from zero and scales as a constant. Holding other factors constant, the magnitude of OVBs decreases as $\left|\beta_{K}^{*}\right|$ and $\left|\gamma_{K}^{*}\right|$ decrease (i.e., as $\left|a\right|$ and $\left|b\right|$ decrease). As $\left|a\right|$ and $\left|b\right|$ become very small, the relevant controls become essentially irrelevant. Panel (d) of Figures (ref) and (ref) provides an illustration of Regime (ii).

Regime (iii) (absence of OVB). All relevant controls will be selected when the magnitude of their coefficients is large enough. Specifically, for all $j\in K,\;a>3,\,b>3$, if

equation[equation omitted — 251 chars of source]

then $\mathbb{P}\left[\textrm{supp}\left(\hat{\pi}\right)=\textrm{supp}\left(\pi^{*}\right)\right]\geq1-1/p^{\tau}$ or $\mathbb{P}\left[\textrm{supp}\left(\hat{\gamma}\right)=\textrm{supp}\left(\gamma^{*}\right)\right]\geq1-1/p^{\tau}$ by standard arguments wainwright_2019. As a result, $\mathbb{P}\left(\left\{ \hat{I}_{1}\cup\hat{I}_{2}\right\} =K\right)\geq1-1/p^{\tau}$ (where $\hat{I}_{1}$ and $\hat{I}_{2}$ are defined in Section (ref)); i.e., the final OLS step ((ref)) includes all the relevant controls with high probability. By similar argument as in Appendix (ref), on the high probability event $\left\{ \left\{ \hat{I}_{1}\cup\hat{I}_{2}\right\} =K\right\} $, the OVB of $\tilde{\alpha}$ is zero. Panel (a) of Figures (ref) and (ref) provides an illustration of Regime (iii).

This “triple-regime” characterization suggests that one may assess the robustness of post double Lasso by increasing the penalty level. If increasing $\lambda_{1}$ and $\lambda_{2}$ yields similar estimates $\tilde{\alpha}$, then the underlying model could be in the regime where either the OVBs are negligible (Regime (ii)) or under-selection in both Lasso steps is unlikely (Regime (iii)). By contrast, under Regime (i), the performance of post double Lasso can be quite sensitive to an increase of $\lambda_{1}$ and $\lambda_{2}$. The rationale behind this heuristic lies in that the final step of post double Lasso, ((ref)), is simply an OLS regression of $Y$ on $D$ and the union of selected controls from ((ref))--((ref)). A natural question is by how much $\lambda_{1}$ and $\lambda_{2}$ should be increased for the robustness checks. For the regularization choice proposed in belloni2014inference, we will show in the simulations of Section (ref) that an increase by $50\%$ works well in practice.

Finally, our theoretical results have interesting implications for the comparison between post double Lasso and post single Lasso, where the latter only relies on one Lasso step to select the relevant controls. Note that the magnitude of the OVB of the post single Lasso estimator of $\alpha^{*}$ also falls into three regimes: (i) non-negligible OVB, (ii) negligible OVB, and (iii) absence of OVB. Thus, qualitatively, the OVBs of post single Lasso and post double Lasso have a similar behavior in finite samples. However, quantitatively, the magnitude of the OVBs can be much larger for the post single Lasso than for the post double Lasso, as illustrated by the following example. If $\beta_{j}^{*}=a\phi^{-1}\sigma_{\eta}\sqrt{2n^{-1}\left(1+\tau\right)\log p}$ with $\left|a\right|\in(0,\,1]$, and $\left|\gamma_{j}^{*}\right|\geq1$, then the argument for showing Proposition (ref) in Appendix (ref) implies that the OVB lower bound scales roughly as $\left|a\right|k\sqrt{\left(\log p\right)/n}$ for the post single Lasso estimator of $\alpha^{*}$ (and note that $\left(\left|a\right|k\sqrt{\left(\log p\right)/n}\right)/\left(1/\sqrt{n}\right)=\left|a\right|k\sqrt{\log p}$).

Additional theoretical results

In this section, we briefly summarize the additional theoretical results that are provided in the appendix. First, we also consider cases where $\alpha^{*}\neq0$. The conditions required to derive the explicit formula are difficult to interpret when $\alpha^{*}\ne0$. However, it is possible to provide easy-to-interpret scaling results (without explicit constants) for cases where $\alpha^{*}\ne0$. Roughly, the scaling of our OVB lower bound can be as large as

equation*[equation* omitted — 160 chars of source]

and

equation[equation omitted — 159 chars of source]

These results reveal an interesting feature of the post double Lasso. The scaling of the OVB lower bound depends on $\left|\alpha^{*}\right|$ when the relevant controls are not selected. This is because the error $u$ in the reduced form equation (ref) involves $\alpha^{*}$ such that the choice of $\lambda_{1}$ in (ref) depends on $\left|\alpha^{*}\right|$ via the variance of $u$. By contrast, it is well-known that the OVB of OLS does not depend on $\left|\alpha^{*}\right|$ when relevant controls are omitted. Interested readers are referred to Propositions (ref) and (ref) in Appendix (ref) for details.

Second, we also provide upper bounds on the OVB. We have seen in equation (ref) that the lower bound on the OVBs scales as $ \left(\sigma_{\eta}/\sigma_{v}\right)\vee\left|\alpha^{*}\right|$ when $\left(k\log p\right)/n\rightarrow\infty$. Interestingly enough, we can also show that the upper bound on the OVB scales as in equation (ref) despite $\left(k\log p\right)/n\rightarrow\infty$ and the Lasso being inconsistent in the sense $\sqrt{n^{-1}\sum_{i=1}^{n}\left(X_{i}\hat{\pi}-X_{i}\pi^{*}\right)^{2}}\rightarrow\infty$, $\sqrt{n^{-1}\sum_{i=1}^{n}\left(X_{i}\hat{\gamma}-X_{i}\gamma^{*}\right)^{2}}\rightarrow\infty$ with high probability. Interested readers are referred to Propositions (ref) and (ref) in Appendix (ref) for details.

Simulations and empirical evidence

To better understand the practical implications of the OVB of post double Lasso, in this section, we present the results from simulations and two empirical applications with widely-used regularization choices available in standard software packages. The analyses were carried out using Matlab MATLAB2020, R R2021, and Stata stata2021.

Simulations

In this section, we present simulation evidence on the performance of post double Lasso with three choices of the regularization parameter: (i) the heteroscedasticity-robust proposal of belloni2012sparse,belloni2014inference ($\lambda_{\text{BCCH}}$) implemented using the R-package hdm with the double selection option hdm2016, (ii) the regularization parameter with the minimum cross-validated error ($\lambda_{\text{min}}$) implemented using the R-package glmnet glmnet, and (iii) the regularization parameter corresponding to the minimum plus one standard deviation cross-validated error ($\lambda_{\text{1se}}$) also implemented using glmnet. We use the same type of regularization parameter choice in both Lasso steps.

We simulate data according to the DGP of Section (ref). To illustrate the role of the sample size $n$ and the sparsity parameter $k$, we consider (i) $(n,p,k)=(500,200,5)$, (ii) $(n,p,k)=(1000,200,5)$, and (iii) $(n,p,k)=(500,200,10)$. We show results for $R^2\in \{0.01,0.05,0.1,0.2,0.3,0.4,0.5\}$ based on 1,000 simulation repetitions. Appendix (ref) presents additional simulation evidence, where we vary $n$, the distribution of $(X_i,\eta_i,v_i)$, the true value $\alpha^\ast$, and also consider a heteroscedastic DGP.

We start by investigating the selection performance of the two Lasso steps of post double Lasso. Panel (a) of Figure (ref) displays the average number of selected controls (i.e., the cardinality of $\hat{I}_{1}\cup\hat{I}_{2}$ in (ref)) as a function of $R^2$. Lasso with $\lambda_{\text{BCCH}}$ selects the lowest number of controls. Choosing $\lambda_{\text{1se}}$ leads to a somewhat higher number of selected controls and results in moderate over-selection for larger values of $R^2$. Lasso with $\lambda_{\text{min}}$ selects the highest number of controls and exhibits substantial over-selection. Panel (b) shows the corresponding average numbers of selected relevant controls. [FIGURE (ref) HERE.]

Figure (ref) presents evidence on the bias of post double Lasso. To make the results easier to interpret, we report the ratio of the bias to the empirical standard deviation. Post double Lasso with $\lambda_{\text{BCCH}}$ can exhibit biases that are more than two times larger than the standard deviation when $(n,p,k)=(500,200,10)$. The bias can still be comparable to the standard deviation when $(n,p,k)=(1000,200,5)$. Consistent with our theoretical discussions in Sections (ref) and (ref), the relationship between $R^2$ and the ratio of bias to standard deviation is non-monotonic: it is increasing for small $R^2$ and decreasing for larger $R^2$. The bias is somewhat smaller for $\lambda=\lambda_{\text{1se}}$. Setting $\lambda=\lambda_{\text{min}}$ yields the smallest ratio of bias to standard deviation. Finally, we note that when $R^2$ is large enough such that there is no under-selection, post double Lasso performs well and is approximately unbiased for all regularization parameters. [FIGURE (ref) HERE.]

The additional simulation evidence reported in Appendix (ref) confirms these results but further shows that $\alpha^\ast$ is an important determinant of the performance of post double Lasso because of its direct effect on the magnitude of the coefficients and the error variance in the reduced form equation (ref). Moreover, we show that, while choosing $\lambda=\lambda_{\text{min}}$ works well when $\alpha^\ast=0$, this choice can yield poor performances when $\alpha^\ast\ne 0$ (see Figure (ref)). A similar phenomenon arises when using $0.5\lambda_{\text{BCCH}}$ instead of $\lambda_{\text{BCCH}}$: this choice works well when $\alpha^\ast=0$ (see Figure (ref) below), but yields biases when $\alpha^\ast\ne 0$. We found that, under our DGPs, this is related to the fact that when $\alpha^\ast\ne 0$, (ref) and (ref) differ with respect to the underlying coefficients and noise variances, which leads to differences in the (over-)selection behavior of the Lasso. Thus, there is no simple recommendation for how to choose the regularization parameters in practice.

The substantive performance differences between the three regularization choices suggest that post double Lasso is sensitive to the penalty levels in the intermediate case where $R^2$ is small enough so that under-selection occurs but large enough to cause substantial OVBs. To further investigate this issue, we compare the results for $\lambda_{\text{BCCH}}$, $0.5 \lambda_{\text{BCCH}}$, and $1.5 \lambda_{\text{BCCH}}$. Figure (ref) displays the average numbers of all selected controls (relevant or not) and selected relevant controls in both Lasso steps. The differences in the selection performance are substantial. Lasso with $0.5 \lambda_{\text{BCCH}}$ over-selects for all $n$ and $R^2$, while Lasso with $1.5 \lambda_{\text{BCCH}}$ under-selects unless $R^2$ and $n$ are large and $k=5$. The differences get smaller as $n$ increases and larger as $k$ increases. [FIGURE (ref) HERE.]

Figure (ref) displays the ratio of bias to standard deviation. Choosing $0.5 \lambda_{\text{BCCH}}$ yields small biases relative to the standard deviations for all $R^2$. By contrast, choosing $1.5 \lambda_{\text{BCCH}}$ yields biases that can be more than six times larger than the standard deviations when $(n,p,k)=(500,200,10)$ and still substantial when $(n,p,k)=(1000,200,5)$. For very small and large values of $R^2$, post double Lasso is less sensitive to the penalty level. In Section (ref), we discuss how to interpret and use robustness checks with respect to the regularization parameters in empirical applications. [FIGURE (ref) HERE.]

In sum, our simulation evidence shows that (i) under-selection can lead to large biases relative to the standard deviations, (ii) the performance of post double Lasso can be very sensitive to the choice of regularization parameters, and (iii) there is no simple recommendation for how to choose the regularization parameters in practice.

Empirical evidence

The effect of 401k plans on total wealth

We revisit the analysis of the causal effect of eligibility for 401(k) plans ($D$) on total wealth ($Y$).\footnote{

doublespaceThe effect of 401(k) plans is well-studied. We estimate intention to treat effect of 401(k) eligibility on assets as, e.g., in Poterbaetal1994,Poterbaetal1995,Poterbaetal1998 and Benjamin2003. Other studies have used 401(k) eligibility to instrument for the actual 401(k) participation status abadie2003semiparametric,CH2004,belloni2017program,wuthrich2019closed.

} We use the data on $n=9915$ households from the 1991 SIPP belloni2017data analyzed by belloni2017program and chernozhukov2018double with high-dimensional methods. We consider two different specifications of the control variables ($X$).

enumerate[itemsep=0pt] • Two-way interactions (TWI) specification. We use the same set of low-dimensional control variables as in Benjamin2003 and CH2004: seven income dummies, five age dummies, family size, four education dummies, and dummies for marital status, two-earner status, defined benefit pension status, individual retirement account (IRA) participation status, and homeownership. Following common empirical practice, we augment this baseline specification with all two-way interactions. After removing collinear columns there are $p=167$ control variables. • Quadratic spline & interactions (QSI) specification. This is the “Quadratic Spline Plus Interactions specification” of belloni2017program. It contains dummies for marital status, two-earner status, defined benefit pension status, IRA participation status, and homeownership, second-order polynomials in family size and education, a third-order polynomial in age, a quadratic spline in income with six breakpoints, as well as interactions of all the non-income variables with each term in the income spline. After removing collinear columns there are $p=272$ control variables.

Table (ref) presents post double Lasso estimates based on the whole sample with $\lambda_{\text{BCCH}}$, $0.5\lambda_{\text{BCCH}}$, and $1.5\lambda_{\text{BCCH}}$. For comparison, we also report OLS estimates with and without controls. For both specifications, the results are qualitatively similar across the different regularization choices and similar to OLS with all controls. This is possible as $n$ is much larger than $p$. Nevertheless, there are some non-negligible quantitative differences between the point estimates. A comparison to OLS without control variables shows that omitting controls can yield substantial OVBs in this application.[TABLE (ref) HERE.]

To investigate the impact of under-selection, we perform the following exercise. We draw random subsamples of size $n_s\in \{200,400,800,1600\}$ with replacement from the original dataset. Based on each subsample, we estimate $\alpha^\ast$ using post double Lasso with $\lambda_{\text{BCCH}}$, $0.5\lambda_{\text{BCCH}}$, and $1.5\lambda_{\text{BCCH}}$ and compute the bias as the difference between the average subsample estimate and the point estimate based on the original data with the same type of regularization choice in Table (ref). The results are based on 1,000 simulation repetitions.

Figures (ref) displays the bias and the ratio of bias to standard deviation for both specifications. We find that post double Lasso can exhibit large finite sample biases. The biases under the QSI specification tend to be smaller (in absolute value) than the biases under the TWI specification. Interestingly, the ratio of bias to standard deviation may not be monotonically decreasing in $n_s$ (in absolute value) due to the standard deviation decaying faster than the bias. Finally, we find that post double Lasso can be very sensitive to the penalty level. [FIGURE (ref) HERE.]

Racial differences in the mental ability of children

We revisit fryerlevitt2013's analysis of the racial differences in the mental ability of young children based on data from the US Collaborative Perinatal Project fryer2013data. As in the reanalysis of chernozhukov2020generic, we restrict the sample to Black and White children so that our final sample includes $n=30002$ observations. We focus on the standardized test score in the Wechsler Intelligence Test at the age of seven as our outcome variable ($Y$). The variable of interest ($D$) is an indicator for Black children. We use the same specification as in fryerlevitt2013, excluding interviewer fixed effects. The control variables ($X$) include extensive information on socio-demographic characteristics, the home environment, and the prenatal environment; see their Table 1B for descriptive statistics. After removing collinear terms there are $p=78$ controls.

Table (ref) shows the results for post double Lasso with $\lambda_{\text{BCCH}}$, $0.5\lambda_{\text{BCCH}}$, and $1.5\lambda_{\text{BCCH}}$, as well as OLS with and without controls based on the whole sample. Since $n=30002$ is much larger than $p=78$, all methods except for OLS without controls yield similar results. [TABLE (ref) HERE.]

To investigate the impact of under-selection, we draw random subsamples of size $n_s\in \{200,400,800,1600\}$ with replacement from the original dataset. In each sample, we estimate $\alpha^\ast$ using post double Lasso with $\lambda_{\text{BCCH}}$, $0.5\lambda_{\text{BCCH}}$, and $1.5\lambda_{\text{BCCH}}$ and compute the bias as the difference between the average estimate based on the subsamples and the estimate based on the original data with the same type of regularization choice. The results are based on 1,000 simulation repetitions.

Figure (ref) displays the bias and the ratio of bias to standard deviation. While the magnitude of the bias is decreasing in $n_s$, it can be substantial and larger than the standard deviation when $n_s$ is small. Moreover, the performance of post double Lasso is very sensitive to the choice of the regularization parameters. With $0.5\lambda_{\text{BCCH}}$, post double Lasso is approximately unbiased for all $n_s$, whereas, with $1.5\lambda_{\text{BCCH}}$, the bias is comparable to the standard deviation even when $n_s=1600$. [FIGURE (ref) HERE.]

Implications for inference and comparison to high-dimensional OLS-based methods

The OVBs have important consequences for making inferences based on post double Lasso. Figure (ref) displays the coverage rates of 90% confidence intervals based on the DGPs in Section (ref) and shows that the OVB of post double Lasso can cause substantial under-coverage even when $n=1000$ and $k=5$. The most important determinant of the under-coverage is the sparsity parameter $k$. Our results show that the requirement on $k$ for guaranteeing a good finite sample coverage accuracy for all $R^2$ can be quite stringent. [FIGURE (ref) HERE.]

These results prompt the question of how to make inference in a reliable manner when one is concerned about OVBs. In many economic applications, $p$ is comparable to but still smaller than $n$. In such settings, OLS-based inference procedures provide a natural alternative to Lasso-based methods. Under classical conditions, OLS is the best linear unbiased estimator and admits exact finite sample inference as long as $p+1\leq n$ (recalling that the number of regression coefficients is $p+1$ in (ref)). Unlike the Lasso-based inference methods, OLS does not rely on any sparsity assumptions. This is important because sparsity assumptions may not be satisfied in applications and, as we show in this paper, the OVBs of Lasso-based inference procedures can be substantial even when $k$ is small and $n$ is large and larger than $p$. Indeed, OLS-based inference exhibits desirable optimality properties absent sparsity (or other restrictions) on $\beta^\ast$.\footnote{

doublespaceFor one-sided testing problems, the one-sided $t$-test based on OLS with all controls is the uniformly most powerful test; for two-sided problems, the two-sided $t$-test is the uniformly most powerful unbiased test vandervaart1998book. We refer to Section 4 in armstrong2016, Section 5.5 in elliott2015nearly, and Section 2.1 in li2021linear for further discussions.

}

While OLS is unbiased, constructing standard errors is challenging when $p$ is large. For instance, cattaneo2018inference show that conventional Eicker-White robust standard errors are inconsistent under asymptotics where $p$ grows as fast as $n$. This result motivates a recent literature to develop high-dimensional OLS-based inference procedures that are valid in settings with many controls cattaneo2018inference,jochmans2020heteroscedasticity,kline2020leave.

Figures (ref) compares the finite sample performance of post double Lasso and OLS with the heteroscedasticity robust HCK standard errors proposed by cattaneo2018inference. Panel (a) shows that OLS exhibits close-to-exact empirical coverage rates irrespective of the magnitude of the coefficients and the implied $R^2$. The additional simulation evidence in Appendix (ref) confirms the excellent performance of OLS with HCK standard errors. Panel (b) displays the average length of 90% confidence intervals and shows that OLS yields somewhat wider confidence intervals than post double Lasso. [FIGURE (ref) HERE.]

In sum, our simulation results suggest that modern OLS-based inference methods that accommodate many controls may constitute a viable alternative to Lasso-based inference methods. These methods are unbiased and demonstrate an excellent size accuracy, irrespective of the magnitude of the coefficients corresponding to the relevant controls. However, there is a trade-off because OLS yields somewhat wider confidence intervals than post double Lasso.

Finally, it is worth noting that we consider settings where one can easily invert $X^T X$ and the OLS and HCK variance estimators are numerically stable. In the case of singular or nearly singular $X^T X$, regularization is often unavoidable; see Section (ref) for a discussion of alternatives to OLS and Lasso-based inference methods.

Recommendations for empirical practice

Here we summarize the practical implications of our results and provide guidance for empirical researchers.

First, the simulation evidence in Section (ref) and Appendix (ref) along with the theoretical results (see Section (ref)) suggest the following heuristic: if the estimates of $\alpha^{*}$ are robust to increasing the theoretically recommended regularization parameters in the two Lasso steps, post double Lasso could be a reliable and efficient method. Therefore, we recommend to always check whether empirical results are robust to increasing the regularization parameters. Based on our simulations, a simple rule of thumb is to increase by $50\%$ the regularization parameters proposed in belloni2014inference. Robustness checks are standard in other contexts (e.g., bandwidth choices in regression discontinuity designs), and our results highlight the importance of such checks in the context of Lasso-based inference methods.

Second, following belloni2014inference, we recommend to always augment the union of selected controls with an “amelioration” set of controls motivated by economic theory and prior knowledge to mitigate the OVBs.

Third, our simulations show that in moderately high-dimensional settings where $p$ is comparable to but smaller than $n$, recently developed OLS-based inference methods that are robust to the inclusion of many controls exhibit better size properties. These simulation results suggest that high-dimensional OLS-based procedures constitute a possible alternative to Lasso-based inference methods.

Forth, OLS-based methods are not applicable when $p>n$, and the OLS and variance estimators can be numerically unstable under severe multi-collinearity even if $p<n$. In such cases, regularization is often needed. Ridge regressions, which impose restrictions on the Euclidean norm of $\beta^\ast$, avoid variable selection and may be a useful alternative to the Lasso; see also armstrong2020bias for related restrictions on $\beta^\ast$.

Finally, in many economic applications, researchers start with a small number of raw controls and want to use a flexible non-parametric model to capture the dependence of outcomes on controls while maintaining a simple parametric form for modeling the variables of interest. Such a specification leads to the classical partially linear models. In fact, belloni2014inference motivate post double Lasso with these models. If one is concerned about OVBs, inference methods that do not rely on variable selection are natural alternatives to post double Lasso. Under suitable smoothness restrictions on the non-parametric component, inference on the parameter of interest in partially linear models is a well-studied problem robinson1988root,newey1994asymptotic. The frameworks proposed in these papers can be built upon procedures such as sieves chen2007chapter16, local non-parametric methods fan1996local, and kernel ridge regressions scholkopf2002learning.

Conclusion

Given the rapidly increasing popularity of Lasso and Lasso-based inference methods in empirical economic research, it is crucial to better understand the merits and limitations of these new tools, and how they compare to other alternatives such as the high-dimensional OLS-based procedures.

This paper presents theoretical results as well as simulation and empirical evidence on the finite sample behavior of post double Lasso and the debiased Lasso (in the appendix). Specifically, we analyze the finite sample OVBs arising from the Lasso not selecting all the relevant control variables. Our results have important practical implications, and we provide guidance for empirical researchers.

We focus on the implications of under-selection for post double Lasso and the debiased Lasso in linear regression models. However, our results on the under-selection of the Lasso also have important implications for other inference methods that rely on Lasso as a first-step estimator. Towards this end, an interesting avenue for future research would be to investigate the impact of under-selection on the performance of the Lasso-based approaches proposed by belloni2014inference, farrell2015robust, belloni2017program, and chernozhukov2018double for non-linear models. In moderately high-dimensional settings where $p$ is smaller than but comparable to $n$, it would also be interesting to compare the treatment effects estimators in belloni2014inference to the robust finite sample methods proposed by rothe2017robust.

Finally, this paper motivates further examinations of the practical usefulness of Lasso-based inference procedures and other modern high-dimensional methods. For example, angrist2019machine present interesting simulation evidence on the finite sample behavior of Lasso-based IV methods belloni2012sparse. It would be interesting to explore the implications of our theoretical results on the under-selection of the Lasso in problems with weak instruments.

Figures and tables

figure[figure omitted — 1,049 chars of source]
figure[figure omitted — 340 chars of source]
figure[figure omitted — 207 chars of source]
figure[figure omitted — 843 chars of source]
figure[figure omitted — 385 chars of source]
figure[figure omitted — 915 chars of source]
figure[figure omitted — 435 chars of source]
figure[figure omitted — 713 chars of source]
figure[figure omitted — 313 chars of source]
figure[figure omitted — 385 chars of source]
figure[figure omitted — 891 chars of source]
table[table omitted — 1,120 chars of source]
table[table omitted — 599 chars of source]