Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
85,026 characters · 18 sections · 61 citation commands
Ill-posed Estimation in High-Dimensional Models with Instrumental Variables
In econometric applications, we may want to include a large number of regressors to account for heterogeneity of individuals or simply because economic theory is not explicit about which regressors to include in the model. These settings often lead to high-dimensional models where the number of parameters to be estimated is close to the sample size or even larger.
In this paper, we consider an instrumental variable (IV) model where the vector of parameters $\beta^0$ is identified through
for a scalar dependent variable $Y$, a possibly endogenous vector of covariates $X$, and a vector of instrumental variables and exogenous covariates $Z$. Our setup is high-dimensional in the sense that the dimension of $\beta^0$ may be larger than the sample size $n$.
This paper is concerned with inference on inner products of $\beta^0$ of the type $a^T\beta^0$ for some vector $a$. In this sense, our model has a semi-parametric interpretation. When a low-dimensional subvector of $\beta^0$ is the parameter of interest and the remaining components of $\beta^0$ are considered as nuisance parameters, then inference on $a^T\beta^0$ implies inference on this low-dimensional subvector of $\beta^0$ for an appropriate choice of the vector $a$. We also allow the subvector of $\beta^0$ of interest to increase slowly with the sample size and provide inference for it. Our main example is when the low-dimensional subvector of $\beta^0$ is associated with endogenous regressors.
As the number of regressors in $X$ may increase with the sample size $n$, also the singular values of the matrix $M$ defined as
depend on $n$. In particular, including additional control variables in the model might affect the dependence between endogenous regressors and instruments and hence the cross second moment. This leads to situations where the singular values of $M$ decrease with $n$ and the vector $\beta^0$ is thus not strongly identified, following the terminology in andrews2012estimation. Also, when the number of endogenous regressors increases with $n$ it is well known that the singular values of $M$ converge to zero in general and might even have an exponential decay. In the high-dimensional case, we then require some form of sparsity of the matrix $M$, i.e., that many entries of $M$ are zero or sufficiently small.
A crucial insight of this paper is to show how the mapping properties of the matrix $M$ affect the asymptotic behavior of our estimator. For instance, we see that the minimal eigenvalue of $M$ slows down the rate of convergence and enlarges the asymptotic variance of our estimator. In addition, the sparsity pattern of $M$ and the sparsity pattern of the parameter vector $\beta^0$ are shown to be related: less sparsity of $M$ requires a higher degree of sparsity of $\beta^0$ and vice versa. This can be interpreted as an $\ell_1$ analog of the so-called source condition used in the inverse problems literature which links the smoothness of the unknown function to the smoothing properties of the operator that characterizes the inverse problem.
This paper proposes a novel estimation procedure based on a Lasso type estimator, suitably modified to have a tractable limiting distribution for inner products of $\beta^0$. While the Lasso estimator makes use of the underlying sparsity constraints, it is well known that it does not have a tractable limiting distribution. In this paper, we use the methodology of desparsification to make up for this drawback. Our desparsified IV Lasso estimator for $\beta^0$ corrects the high-dimensional two stage least squares (2SLS) estimator by subtracting a regularization bias. In the case of low dimensions, i.e. under a known sparsity structure, the resulting estimator coincides with the ordinary 2SLS estimator.
We establish the rate of convergence of inner products of our estimator, and show that the rate is affected by the minimum singular value of $M$ (opportunely normalized). In particular, we can show an analog to the nonparametric IV case, where slow rates of convergence are common. Moreover, inner products of our estimator for $\beta^0$ are shown to be asymptotically normal. The normalization factor for the estimator is shown to be driven by the minimal singular value of $M$. We derive confidence intervals and hypothesis testing procedures for inner products of $\beta^0$. As discussed above, inference results on inner products of $\beta^0$ imply inference results on low-dimensional subvectors of $\beta^0$ or even on subvectors of $\beta^0$ slowly increasing with the sample size. In Monte Carlo simulations, we show that the proposed confidence intervals have accurate size.
It is interesting to note that having the rate of our estimator affected by the minimum singular value of $M$ is similar to what happens for sieve estimation in the nonparametric IV (NPIV) literature. In NPIV literature the rate of convergence is derived under smoothness assumptions of the underlying IV regression functions instead of under sparsity constraints of the IV regression coefficients as in this paper. In particular, model (ref) can be also seen as an approximation of the true relationship between $Y$ and a vector of endogenous covariates based on a dictionary $X$ of transformations of the endogenous covariates. Hence, the two types of assumptions (smoothness and sparsity) provide two alternative frameworks to deal with high-dimension in nonparametric IV regression models.
\paragraph{Related Literature.} Our paper contributes to the growing literature on inference for structural parameters in sparse high-dimensional IV settings. Much work in this setting focuses on the case where the dimension of the endogenous variable is small but where there is a large number of available instruments, see ng2009selecting, belloni2012, and BelloniChernozhukovHansen2011. In this context, ChaoSwanson2005 and HansenHausmanNewey2008, among others, propose methods to account for weak identification when the number of instruments is allowed to be large but not larger than $n$. HANSEN2014 and CarrascoDoukali2017 extend this literature by considering cases where the number of weak instruments is larger than $n$ and propose a Ridge regularized jackknife instrumental variable estimation. When the number of endogenous regressors in model (ref) is fixed and there are high-dimensional control variables, ChernozhukovHansenSpindler2015 propose a three step estimator where high-dimensional sparse linear models with only exogenous variables are fitted. In particular, Lasso is only used for the fit of nuisance parameters and the use of the Lasso estimates follows standard lines. For the fit of the parameters of the endogenous covariates, the criterion function is orthogonalized such that errors in the estimation of the other parameters (i.e. of the nuisance parameters) enter into the model only quadratically. For this reason the classical bounds for the errors of the Lasso estimates of the nuisance parameters suffice. In particular, no debiasing/desparsification of Lasso estimates is needed at any point of the procedure. BelloniChernozhukovFVHansen consider estimation of treatment effects in IV models with binary instrument and endogenous variable in the presence of a high-dimensional set of control variables.
Also relevant to this paper is the literature concerning choice of valid instruments. In the context of a scalar endogenous variable, guo2016 propose a method to select valid instruments based on hard thresholding in setups where the number of instruments and of exogenous variables may tend to infinity. Their proposal is related to LASSO approaches for the selection of valid instruments in finite dimensional setups. kang2016instrumental use Lasso to instrumental variable selection in the context of invalid instruments. Based on an initial median estimator, windmeijer2017 use adaptive Lasso for instrument selection and establish consistency of their procedure.
In model (ref) which allows for increasing dimension of endogenous regressors, gautier2011high establish a novel estimation procedure based on novel sensitivity characteristics of the empirical counterpart of $M$ to obtain confidence sets with length depending on the strength of instruments. gautier2011high also establish confidence bands after bias correction. belloni2017simultaneous use such sensitivity to construct simultaneously valid confidence regions and have proposed a multiplier bootstrap procedure to compute critical values and establish its validity. Their approach is based on orthogonality restrictions when considering linear combinations of the original instruments.\\ Our approach is essentially different from the previous ones as our sparsity condition is based on the population matrix $M$ and not on its empirical counterpart. This allows us to provide a novel link between high-dimensional and NPIV estimation where the first is based on assuming sparsity while the latter is based on assuming smoothness in the underlying model, see e.g. AiChen2003, NP03econometrica, DFR02, Chen08, and references therein for NPIV estimation. fan2014 propose a modified Lasso approach for estimation in high-dimensional instrumental variables models. Our paper is also related to guo2016 and gold2018 that, as we propose in our paper, use two-step estimators using a threshold procedure and Lasso estimation, respectively, in the first step and desparsification in the second step. However, gold2018 make assumptions about sparsity that differ from ours and their settings exclude cases where the estimator of components of $\beta$ does not achieve a parametric $\sqrt{n}$-rate. On the other hand, we do allow for singular values of $M$ to tend to zero which yields slower rates and provide novel inference results for inner products of the estimator of $\beta$ of increasing dimension. This is an important feature of our paper as we are thus able to provide an interpretation that is close to the nonparametric IV estimation. Recently, neykov2018unified considered univariate confidence set estimation in a high-dimensional setting but require that the number of instruments coincides with the number of endogenous variables. In contrast, our approach is efficient under sparsity constraints and moreover, it is robust against violations of strong identification, which, as far as we know, has not been addressed in the related literature.
Our paper is also related to the rich statistical literature on high-dimensional statistical models that contain only exogenous variables and where endogeneity and instrumental variables are not considered, see, zhang2014, javanmard2014, javanmard2014G and vandegeer2014. An alternative approach to our desparsified Lasso estimator is ridge regression where an $\ell_2$ penalty is used and the asymptotic distribution results can be readily obtained. This approach in high-dimensional Gaussian regression is considered by buhlmann2013. In an extensive simulation study, however, javanmard2014G show that the ridge regression approach is overly conservative, which is in line with the theoretical results. This is why we also pursue to desparsify the Lasso estimator rather than using the ridge regression.
The remainder of the paper is organized as follows. In Section (ref), we describe the model, motivate the desparsification procedure and discuss sparsity requirements. Section (ref) contains the rates of convergence and the asymptotic normality results of our estimator. Section (ref) is concerned with the finite sample performance of our estimator. Simulations and numerical implementation of our inference procedure are in Section (ref). In Section (ref) we present an application to demand estimation for automobiles using market share data. All proofs can be found in the appendix.
\paragraph{Notation.} The $\ell_p$ norm of a vector $a$ is denoted by $\|a\|_p$, $1\leq p\leq \infty$. For a set $S$, the cardinality of $S$ is denoted by $|S|$. For a vector $a$, and $S$ a set of indices, $a_S$ denotes the restriction of $a$ to indices in $S$: $a_S:=\{a_j; j\in S\}$. Further, for a matrix $A$ we use the notation
for the element-wise sup-norm,
for the operator norm, and
for the $\ell_1$ norm. For vectors $a$ we have $\|a\|_{op,\infty} = \| a\|_{\infty}$ and for a matrix $A$ it holds $\|A\|_{op,\infty} = \|A^T\|_{1}.$ The smallest and largest eigenvalue of $A$ are denoted by $\lambda_{\min}(A)$ and $\lambda_{\max}(A)$, respectively. We denote by $A_{j}$ the $j$-th column of the matrix $A$ and by $A_{-j}$ the matrix $A$ without the $j$-th column. We denote by $e_j$ the $j$-th unit column vector. For two positive sequences $a_n$, $b_n$ we use the notation $a_n\sim b_n$ to mean that there are two positive and universal constants $C_1, C_2$ such that $C_1\leq a_n/b_n\leq C_2$. We abbreviate “with probability approaching one” to “wpa1”, and say that a sequence of events ${\left\lbrace B_n\right\rbrace }$ holds wpa1 if $\mathbb{P}(B_n^c) = o(1)$ as $n\rightarrow \infty$.
Consider again model (ref), the high-dimensional instrumental variable model is given by
where $\beta^0$ is the $ p$--dimensional, unknown parameter of interest. Some of the covariates in $X$ are possibly endogenous in the sense that they are related to the unobservables $U$, i.e., ${\mathbb E}[ UX]$ does not vanish. Here, $Y$ is a scalar dependent variable, $ X$ is a $ p$--dimensional vector of endogenous and exogenous covariates, $Z$ is a $q$--dimensional vector of instrumental variables and exogenous covariates. So, the vectors $Z$ and $X$ may have elements in common if $X$ contains exogenous covariates.\\ To ensure identification of the parameter $\beta^0$, we assume throughout the paper that $q\geq p$. This condition can be met in at least three situations: (1) the case where one has low dimensional endogenous variables and high-dimensional exogenous controls and the interest is in inference on the coefficients of the low dimensional endogenous variables (examples are: [a] the case where one includes many exogenous controls to account for complex heterogeneity, or [b] the case where one includes many exogenous variables to account for nonlinearities to approximate a partial linear structure); (2) more generally, the case where a high-dimensional linear sieve approach is used to approximate nonlinear and nonparametric instrumental variables models; (3) models that have a rich information in exogenous variation, corresponding to high-dimensional instrumental variables, as for instance in AngristKrueger1991.
We also assume throughout the paper that the matrices $M := {\mathbb E}\left(ZX^T\right)$ and $\Sigma := {\mathbb E}\left(ZZ^T\right)$ are of full column rank. Thus, the parameter vector $\beta^0$ is identified through
Estimating $\beta^0$ by simply replacing the matrices on the right hand side by their empirical counterparts fails for two reasons. First, the empirical counterparts of $M$ and $\Sigma$ are in general not of full rank in the high-dimensional case. Second, it is well known that, for large matrices, estimators simply based on the sample mean do not provide satisfactory performance. In this paper, we address these challenges by using regularization procedures.
A common assumption to obtain consistent estimation results in the high-dimensional setting is a sparsity restriction: most of the parameters of $\beta^0$ are zero (exact sparsity) or sufficiently small (approximate sparsity) which implies that a relatively small number of regressors in $X$ is sufficient in describing the dependent variable $Y$.
Note that the model is not identified if the minimal eigenvalue of $M^T\Sigma^{-1} M$ is zero, which we rule out throughout the paper (see Assumption (ref)). We thus introduce
which satisfies $\omega <\infty$ for each $n\geq 1$ under Assumption (ref). Below, we also assume that the maximal eigenvalue of $M^T\Sigma^{-1} M$ is bounded from above uniformly in $n\geq 1$ and hence, $\omega $ is strictly positive for all $n\geq 1$. On the other hand, in many cases we expect that $\omega $ might increase with the sample size $n$ either because the model requires a large number of functions to account for nonlinearity in the endogenous covariates or because the instruments are weak and thus the model is not strongly identified. In the first case, $X^T\beta$ is an approximation of the true nonlinear instrumental regression through approximating functions stored in $X$ whose number increases with $n$. In the second case, weakness of the instruments is captured by close to zero elements in the matrices $M$ and $\Sigma^{-1/2}M$.
Similar to andrews2012estimation, we consider the strongly identified case where $\omega $ is uniformly bounded above, and the semi-strongly identified case where $\omega $ is unbounded but satisfies $n/\omega \to\infty$. We show below that $\omega $ slows down the rate of convergence of our estimator. In the semi-strong case, the size of the confidence sets increases relative to $\omega $. There is also a third case which is the weak identified case where $n/\omega =O(1)$ but the results of our paper do not apply to it. In this paper, we show that under appropriate assumptions the rate of convergence of our estimator for each component of $\beta^0$ is $\sqrt{\omega /n}$.
Throughout the paper, we assume that a sample $(Y_i,X_i,Z_i)$, $1\leq i\leq n$ of independent and identically distributed copies of $(Y,X,Z)$ is available. We write the vector and matrices of observations as $\mathbf Y = (Y_1,\dots,Y_n)^T$, $\mathbf X=(\mathbf X_1,\dots,\mathbf X_p)$ with $\mathbf X_j=( X_{1,j},\dots, X_{n,j})^T$ for $1\leq j\leq p$ , and $\mathbf Z=(\mathbf Z_1,\dots,\mathbf Z_q)$ with $\mathbf Z_j=(Z_{1,j},\dots, Z_{n,j})^T$ for $1\leq j\leq q$. Moreover, the $n$-vector of unobservables is denoted by $\mathbf U = (U_1,\ldots, U_n)^T$. The $(d\times d)$--dimensional identity matrix is denoted by $I_d$ and its $j$--th column by $e_j$.
In this section, we introduce our estimation procedure which is based on desparsifying a Lasso estimator. The methodology is based on regularized estimators of the matrices $\Theta:=\Sigma^{-1}$, $M := {\mathbb E}[Z X^T]$, and $\Theta^M := (M^T\Theta M)^{-1}$ denoted by $\widehat \Theta$, $\widehat M$, and $\widehat\Theta^M$, respectively, which are defined in Subsection (ref). We propose the following desparsified IV Lasso estimator of $\beta^0$ given by
where $\widetilde \beta$ is a consistent estimator of $\beta^0$ that makes use of the underlying sparsity assumption. The first summand on the right hand side of (ref) corresponds to a regularized empirical analog of $\beta^0$ as in (ref). The second summand on the right hand side of (ref) accounts for the regularization bias of our matrix estimators.
The proposed estimator naturally extends the 2SLS estimator to the high-dimensional case. Consider the situation of a known sparsity structure where regularization is not required and so $\widehat \Theta$, $\widehat M$, and $\widehat\Theta^M$ are the usual empirical counterparts of $\Theta$, $M$ and $\Theta^M$. In this case it holds $\widehat\Theta^M\widehat M^T \widehat\Theta\mathbf Z^T\mathbf X/n=I_p$ and moreover, $\widehat \beta$ coincides with the 2SLS estimator.
The choice of $\widetilde\beta$ is motivated by our sparsity assumption given below and by the asymptotic properties for $\widehat\beta$ that we want to obtain. To derive the asymptotic results of our desparsified IV Lasso estimator we make use of the following key decomposition
for a remainder term $\Delta$ which is given by
Then, we have to show that $\|\Delta\|_\infty$ is asymptotically negligible under regularity assumptions. In particular, to show this we require that $\|\widetilde\beta - \beta^0\|_1$ is sufficiently small. This property is satisfied by the Lasso estimator and thus we choose $\widetilde\beta$ in equation (ref) to be the Lasso estimator which makes use of the underlying sparsity structure imposed on $\beta^0$. Therefore, our estimation procedure is based on the IV Lasso estimator of $\beta^0$ given by
for some tuning parameter $\lambda>0$. Then, we replace $\widetilde\beta$ in equation (ref) to obtain $\widehat\beta$.
In this section we introduce some notations and assumptions about sparsity that we tacitly maintain all along the paper. Let $s_0$ denote the cardinality of the set $S_0$, i.e., $s_0 := |S_0|$, where $S_0$ is a set such that $\|\beta_{S_0^c}^0\|_1$ is sufficiently small. That is, we assume that the set $S_0$ is rich enough such that the parameter vector $\beta^0$ satisfies
for some constant $C>0$. Inequality (ref) imposes approximate sparsity on $\beta^0$: The absolute value of the parameters outside the sparsity set $S_0$ is bounded by some value which tends to zero as the sample size tends to infinity.
Hereafter, we assume that $\Theta$ and $\Theta^M $ exist and assume sparsity with respect to rows of $\Theta:= \Sigma^{-1} $. To this purpose we define
The sparsity restriction on $\Theta$ has the following interpretation: if the $(jk)$--th component of $\Theta$ is zero, then the variables $Z_j$ and $Z_k$ are partially uncorrelated, given the other variables. In particular if $Z$ is jointly normal then we have that the variables $Z_j$ and $Z_k$ are conditionally independent, given the other variables. This also motivates the use of an $\ell_1$--penalty for the estimation of $\Sigma^{-1}$, as we do in Section (ref) and which was proposed by meinshausen2006. Note that it is possible to relax the sparsity constraints but this would lead to a less efficient estimator.
We need to assume some sparsity pattern on $M$, that is, most of the elements in each row or column of $M$ are zero. We conjecture that it would suffice to assume only approximate sparsity for $M$ and $\Theta$ but at the cost of much more technical proofs and notation. For the sparsity of $M$ we introduce the notation
Hence, $\|M\|_1 \leq s_M\|M\|_{\infty}$. For $j=1,\dots,p$, we denote $\gamma_j := \mathop{\textrm{argmin}}_{\gamma\in\mathbb{R}^{p-1}}\|(\Theta^{1/2}M)_j - (\Theta^{1/2} M)_{-j}\gamma\|_2^2$ and impose the following approximate sparsity condition on $\gamma_j$: we assume there exists a set $S_j$ such that
for some constant $C>0$. Below we also denote $s_j^M := |S_j|$ and for convenience we use the notation $s_{\max}^M := \max_{1\leq j\leq q}s_j^M$.
In this section, we provide the regularization schemes to construct the approximate inverses $\widehat\Theta$ and $\widehat\Theta^M$ as well as the regularized estimator $\widehat M$. Asymptotic properties of these estimator will be studied in Section (ref).
Here we construct a regularized estimator of the inverse of $\Sigma$ denoted by $\widehat\Theta$. The basic idea to construct such an estimator is to relate the inversion of a $q\times q$ matrix to $q$ regression problems of $\mathbf Z_j$ over $\mathbf Z_{-j}$, where for $1\leq j\leq q$, $\mathbf Z_j=(Z_{1,j},\dots, Z_{n,j})^T$ is the $j$-th column vector of the matrix $\mathbf{Z}$ and $\mathbf Z_{-j}=(\mathbf Z_{1},\dots, \mathbf Z_{j-1},\mathbf Z_{j+1},\ldots, \mathbf Z_{q})$. This approach was introduced by meinshausen2006 and consists in using the Lasso estimator for nodewise regression. For every $j=1,\ldots, q$ we consider the Lasso estimator
for some tuning parameter $\lambda_j^{\Theta}>0$ that will be let to tend to zero as the sample size increases to get asymptotic results. We introduce the $q$-column vector $\widehat \Gamma_j=(\widehat \Gamma_{kj})_{k=1}^q$ such that
with $\widehat\xi_{j} = (\widehat\xi_{jk})_{k\in\{1,\ldots,q\}\setminus \{j\}}$. By the definition of $\widehat \Gamma_j$, it holds $\mathbf Z_j - \mathbf Z_{-j}\widehat\xi_j = \mathbf Z \widehat \Gamma_j$. Then, the matrix $\widehat\Theta = \big(\widehat\Theta_{1},\dots, \widehat\Theta_{q}\big)^T$ is constructed as
Note that while the population counterpart $\Theta$ is symmetric, its estimator $\widehat\Theta$ does not need to be so. For more details on this procedure, we refer to meinshausen2006.
A standard sample matrix estimator for the matrix $M$ does not have good performance in the high-dimensional case and regularization is needed. Hence, we propose a thresholding estimator of $M$. Intuitively, we want to eliminate those values of the empirical matrix $\widetilde{M} := \mathbf Z^T\mathbf X/n$ that lie below some specified threshold. More precisely, we propose to use the thresholding estimator $\widehat{M} = (\widehat{M}_{jk})$ where
and $C_0 > 0$ is a constant defined as in Proposition (ref) and is related to the constant appearing in the large deviation inequality for the components of $\widetilde{M}$. For symmetric matrices such a regularization scheme has been considered in BickelLevina2008 and CaiZhou2012 among others, and we refer to these papers for a discussion on this constant. For practical implementation, we choose $c_n := C_0\sqrt{\log(q)/n}$ by cross-validation, see Section (ref) for more details.
In this section, we construct the estimator $\widehat\Theta^M$ which is an approximate inverse of $\widehat M^T \widehat\Theta \widehat M$. This estimator involves the regularized estimators $\widehat \Theta$ and $\widehat M$ obtained in Sections (ref) and (ref) and the square root of $\widehat\Theta$, denoted by $\widehat \Theta^{1/2}$. Note that we can make use of the Schur decomposition of $\widehat\Theta$ to compute its square root given that $\widehat\Theta$ is not necessarily symmetric.
Let $(\widehat \Theta^{1/2}\widehat M)_j$ denote the $j$-th column vector of the matrix $\widehat \Theta^{1/2}\widehat M$ and
Remark that $\widehat \Theta^{1/2}\widehat M$ is the empirical cross moment of $X$ and the (approximately) orthonormalized $Z$. The approximate orthonormalization of $\mathbf Z$ is performed by premultiplication by $\widehat\Theta^{1/2}$. As for the construction of $\widehat\Theta$, we relate the regularized inversion of a $p\times p$ matrix to $p$ regression problems of $(\widehat \Theta^{1/2}\widehat M)_j$ on $(\widehat \Theta^{1/2}\widehat M)_{-j}$. To do that, for every $j=1,\ldots,p$ we consider the Lasso estimator:
for some tuning parameter $\lambda_j^M>0$ that will be let to tend to zero as the sample size increases to get asymptotic results. Let $\widetilde \Gamma_j=(\widetilde \Gamma_{kj})_{k=1}^p$ be the $p$-column vector determined by
with $\widetilde\gamma_{j} = (\widetilde\gamma_{jk})_{k\in\{1,\ldots,p\}\setminus \{j\}}$. The matrix $\widehat\Theta^M$ is then set equal to $\widehat\Theta^M = (\widehat\Theta_{1}^M,\dots,\widehat\Theta_p^M)^T$ where
As already stressed in vandegeer2014, other regularization methods to obtain the approximate inverses of $\widehat\Sigma$ and $(\widehat M^T \widehat \Theta \widehat M)$ that do not deliver a bound for $\big\|\widehat M^T \widehat \Theta\widehat M\,\widehat\Theta_{j}^M-e_j\big\|_\infty$, like the ridge regularization, may not be optimal because without this bound we cannot directly obtain asymptotic distribution results for components of $\beta^0$. The regularization methods that we use to construct $\widehat\Theta$ and $\widehat\Theta^M$ automatically include this bound in the optimization problem.
In this section, we derive the asymptotic distribution of the desparsified IV Lasso estimator $\widehat \beta$ given in (ref). To obtain asymptotic results on which our inference will be based we have to show that the remainder term $\Delta$ in the key decomposition (ref) is asymptotically negligible. We start by providing all the assumptions that we need to obtain our asymptotic results. After that, we first provide results about rates of convergence for the estimated matrices and for $\widehat\beta$, and then asymptotic normality will be established.
In this section we gather assumptions which we require to establish our inference results. Below, a random vector $W\in\mathbb R^d$ is called sub-Gaussian if ${\mathbb E}\exp\big(|v^T W|^2/C\big)=O(1)$ for all $v\in\mathbb R^d$ such that $\|v\|_2\leq 1$ and some sufficiently large constant $C>0$.
Sub-Gaussianity, as imposed in Assumption (ref) $(ii)$, is satisfied, for instance, if the random vectors have bounded support. Assumption (ref) $(iii)$ implies that $\Sigma_{jj} = O(1)$ uniformly in $j$ since $\Sigma_{jj} \leq \lambda_{\max}(\Sigma)$. Similarly, it also implies that $\|\Theta_j\|_2\leq \lambda_{\min}(\Sigma)=O(1)$ uniformly in $j$ and consequently, $\|\Theta\|_1=O\big( \sqrt{s_{\max}}\big)$ which we use below.We also make use of the notation ${\cal B}:=\{\beta:\, \|\beta_{S_0^c}\|_1\leq 3\|\beta_{S_0}\|_1\}$.
The previous result shows that a modified version of the so called compatibility condition, see e.g. buhlmann2011, is satisfied with high probability. Note that such conditions are required in the high-dimensional estimation context in order to relax the requirement of non-zero eigenvalues of associated estimated matrices. The following assumption provides more details about the choice of regularization parameters and imposes conditions on the underlying sparsity.
Assumption (ref) $(i)$ specifies the rate of the tuning parameters $\lambda$ used for the plug-in Lasso and $\lambda_j^{\Theta}$, $\lambda_j^M$ used for the nodewise Lasso estimators. The rate of the regularization parameters $\lambda$ and $\lambda_j^M$ is larger by $\sqrt{\log(q)}$ than the common choices of it, which is due to the additional estimation step that is involved for our initial IV Lasso estimator. Assumption (ref) $(ii)$ imposes upper bounds on the maximal element (in absolute value) of $M$ and $M^T\Sigma^{-1} M$, and the conditional variance of $U$ given $Z$, which is standard in the literature and is a mild restriction on the heteroscedasticity of the model. Assumption (ref) $(iii)$ imposes sparsity restrictions which we require in order to obtain our inference results. Specifically, this assumption restricts the sparsity of $\beta_0$ (captured by $s_0$) in relation to the sparsity of $M$ (captured by $s_M$). Finally, condition (ref) implies $\log(p)/\sqrt{n} = o(1)$.
For the next assumption, recall that $ \gamma_j:= \mathop{\textrm{argmin}}_{\gamma\in\mathbb{R}^{p-1}}\|(\Theta^{1/2}M)_j - (\Theta^{1/2} M)_{-j}\gamma\|_2^2$ for $1\leq j\leq p$. Introduce a vector $\Gamma_j:=(\Gamma_{kj})_{k=1}^p$ with $\Gamma_{kj}=-\gamma_{kj}$ for $k\neq j$ and $1$ otherwise, where $\gamma_{kj}$ is the $k$--th entry of $\gamma_j$.
Assumption (ref) $(i)$ imposes upper bounds on moments associated to $ZX^T$ while Assumption (ref) $(ii)$ imposes mild rate conditions on moments of $ZZ^T$. Note that the logarithmic rates in Assumption (ref) can be replaced by other powers of logarithms to allow for more heavy tailed variables. This would require slight changes in our constraints on the growth of dimension parameters $p$ and $q$ and somewhat more restrictive sparsity constraints.
In this section we provide rates of convergence for the regularized matrices used to construct our estimator $\widehat \beta$. These results are then used to establish asymptotic normality results in the next section.
In the following result, we derive a rate of convergence for $\widehat M$ in the $\ell_1$ norm. The first part of the theorem provides a large deviation inequality for the components of $\widetilde M$ and it is derived by exploiting sub-Gaussianity of the rows of $\mathbf X$ and $\mathbf Z$ and a Bernstein-type inequality for sub-exponential random variables.
The constant $c$ in inequality (ref) depends on the second order moments and cross moments of the elements in $Z$ and $X$ as well as on their sub-Gaussian norms. Its expression can be deduced from the proof of the proposition given in the appendix.\\ The next result gives a key upper bound for the approximation error of the relaxed inverses $\widehat\Theta_j$ and $\widehat\Theta_{j}^M$. These upper bounds depend on the regularization parameters and the values $\widehat\tau_j$ or $\widetilde\tau_j$. For the inference on the structural parameter, we thus have to control the asymptotic behavior of $\widehat\tau_j$ and $\widetilde\tau_j$.
We now establish the rate of convergence of the regularized estimators $\widehat\Theta$ and $\widehat\Theta^M$. The first result in the next proposition was established by vandegeer2014, and hence the proof is omitted.
In this subsection, we derive the rate of convergence of the desparsified IV Lasso estimator $\widehat \beta$. The next theorem provides an asymptotic upper bound of the bias term $\Delta$, which is key to derive further inference results.
From Theorem (ref) we see that the rate of convergence of the desparsified Lasso estimator $\widehat \beta$ is affected by the possibly increasing parameter $\omega$. In the next result, we show that the bias term $\Delta$ is indeed asymptotically negligible under additional rate requirements. We also see below that the rate of convergence of our estimator is given by $\sqrt{\omega/ n}$ under a mild assumption.
We will see in the next section that $V$ after standardization converges to the standard normal distribution and, in particular, that $\sqrt\omega\, V$ is stochastically bounded. We also see that $\omega$ enters the sparsity condition in equation (ref). In the strong identified case, the components of $\widehat\beta$ are $\sqrt n$ consistent. In the semi-strongly identified case, this rate of convergence may slow down depending on the asymptotic behavior of $\omega$.
Also the next result is an immediate consequence of Corollary (ref) and provides a bound for linear functionals of $\widehat \beta-\beta^0$ uniformly over representers $a\in\mathbb R^p$ with $\ell_1$ norm which might increase at a rate $K:=K(n)$. For some constant $C>0$, we define $\mathcal A_K={\left\lbrace a\in\mathbb R^p: \,\|a\|_1^2/K\leq C\right\rbrace }$.
The sparsity restriction (ref) becomes more restrictive for large values of $K$. Two examples of linear functionals for which Corollary (ref) holds are given by vectors $a$ selecting one component of $\beta$ and vectors $a$ selecting linear combinations of a finite number of components of $\beta$, for which $K=1$ and $K$ is bounded, respectively.
In this subsection, we establish asymptotic normality of inner products of the desparsified Lasso estimator $\widehat\beta$. We also see that asymptotic normality of components of $\widehat\beta$ immediately follows.
To achieve the asymptotic distribution of our estimator $\widehat\beta$ we consider a normalization factor to standardize the estimator $\widehat\beta$. This normalization factor involves the empirical counterpart of the covariance matrix of the 2SLS estimator which is given by
We then require the following assumption on this covariance matrix $\Omega$. We introduce the set $\mathcal A={\left\lbrace a\in\mathbb R^p: \,a\in\ell_2\text{ and }\|a\|_1\leq C\|a\|_2\right\rbrace }$ for some constant $C>0$.
Assumption (ref) can be easily verified under mild regularity assumptions, such as, the lower bound $\sqrt{{\mathbb E}[U^2|Z]}\geq \underline{\sigma}$, which is a common condition to derive asymptotic distribution results. Indeed, the condition $\sqrt{{\mathbb E}[U^2|Z]}\geq \underline{\sigma}$ implies $a^T\Omega\,a\geq \underline{\sigma}^2a^T\Theta^M\,a$ and hence, Assumption (ref) holds, for instance, if the eigenvalues of $\Theta^M$ have a polynomial or exponential decay.
We now propose a heteroscedasticity robust covariance estimator. To obtain the empirical counterpart of $\Omega$, denoted by $\widehat\Omega$, we replace the matrices $\Theta^M$, $M$, and $\Theta$ by their regularized empirical counterparts defined in Section (ref):
for the vector of Lasso residuals $\widehat{\mathbf U}= \big(Y_1-X_1^T\widetilde\beta,\dots, Y_n-X_n^T\widetilde\beta\big)$ and $\widetilde\beta$ is the IV Lasso estimator given in (ref). We now establish asymptotic normality of linear combinations of the components of $\widehat \beta$.
Below, we provide some implications of Theorem (ref). An immediate consequence of Theorem (ref) is componentwise asymptotic normality, in which case $a=e_j$ for some $1\leq j\leq p$ where $e_j$ is a $p$-vector of zeros but for the $j$-th component that is equal to $1$. Another consequence of Theorem (ref) is asymptotic normality of linear combinations of a finite number of components of $\widehat \beta$. In both cases, the restriction imposed in $\mathcal A$ is satisfied. But even if the dimension of the low-dimensional subvector of interest increases, the condition $\|a\|_1/\|a\|_2\leq const.$ can be justified as the following example illustrates.
The next theorem establishes asymptotically valid confidence intervals and testing procedures for inner products of $\beta^0$. The following two corollaries are direct implications of Theorem (ref) and hence, their proofs are omitted. Below, $\Phi$ denotes the cumulative distribution function of the standard normal distribution.
The following examples illustrate the previous theorem for the componentwise case where $a=e_j$.
Another direct implication of Theorem (ref) concerns hypothesis testing. For some $a\in\mathbb{R}^p$ (satisfying condition (ref)) consider the null hypothesis $H_{a,0}:\, a^T\beta^0=a^T\beta^H$ for a given vector $\beta^H\in\mathbb{R}^p$.
This section presents Monte Carlo experiments to analyze the finite sample properties of our estimator. We consider the situation where we have a linear reduced form equation but allow for approximate sparsity. We consider three cases: the case where the true structural relationship is linear and we have homoscedasticity (Section (ref)), the case where the true structural relationship is linear and we have heteroscedasticity (Section (ref)), and finally the homoscedastic case where the true structural relationship is nonlinear in the endogenous variable and we use a series approximation (Section (ref)). All experiments are based on 1000 Monte Carlo iterations. The choice of tuning parameters is based on $10$-fold cross-validation where we make use of the R function \verb"cv.glmnet" of the \verb"glmnet" package (see Appendix (ref) for a description of the cross-validation procedure).
We generate i.i.d. data from the following model
with
where $\Sigma = \left((0.5)^{|j-k|}\right)_{jk}$ is a $q\times q$ matrix. The parameter $\rho$ captures the degree of endogeneity and is varied in the experiments below. The parameters are set in the following way: $\beta_1 = 2$, $\beta_{-1,j} = 1 + (j-1)*c$ for $1\leq j \leq 50$, where $c$ is a constant such that the parameters $\beta_{-1,j}$ are equispaced between $1$ and $3$, $\beta_{-1,j} = 0$ for $51 \leq j \leq (p-1)$, and $\alpha_{-1,j} = 1/(2j^{3})$ for $1\leq j\leq (q-1)$. The parameter $\alpha_1$ accounts for the strength of the instrument $Z_1$ and is varied in the experiments, i.e., we consider $\alpha_1\in{\left\lbrace 1,0.75, 0.5, 0.25\right\rbrace }$. Note that we multiply the error term in the second equation by $\sqrt{1-\alpha_1^2}$, to ensure that the variance of $X_1$ does not depend on the value $\alpha_1$. Since ${\mathbb E}[U^2|Z]=1$ we are in the homoscedastic case where the covariance matrix simplifies to $\Omega= {\mathbb E}[U^2]\Theta^M$.\\ The desparsified IV Lasso estimator $\widehat \beta$ is computed as in Subsection (ref). It is based on the initial IV Lasso $\widetilde\beta$ given in (ref) where the tuning parameter $\lambda$ is chosen via $10$-fold cross-validation. The regularized estimators $\widehat \Theta$, $\widehat M$, and $\widehat\Theta^M$ are implemented as described in Subsection (ref) with the tuning parameters $\lambda_j^{\Theta}$, $j=1,\ldots,q$, and $\lambda_j^M$, $j=1,\ldots,p$, and $c_n= C_0\sqrt{\log(q)/n}$, chosen by 10-fold cross-validation. We emphasize that our implementation of the estimators for high dimensional matrices follows standard procedures in the related literature see, for instance, meinshausen2006. Alternatively, one could use the procedure proposed in vandegeer2014 and choose the same tuning parameter, say $\lambda_j^{\Theta} = \lambda_{\Theta}$ (resp. $\lambda_j^M = \lambda_M$), by $10$-fold cross-validation among all the $q$ (resp. $p$) nodewise regressions. We examined this procedure but it slows down the computational time and the results were not better. For large choices of $q$ and $p$ we made use of parallel computing (which is straightforward in R given the \verb"parallel" package).
To estimate the covariance matrix $\Omega$ we adapt to the instrumental variable setting the idea proposed by sun2012scaled, which consists in replacing the variance of $U$ by the error variance estimator obtained with the initial IV Lasso $\widetilde\beta$, $\widetilde\sigma^2 := \sum_{i=1}^n(Y_i - \widetilde\beta_1 X_1 - \widetilde\beta_{-1}^T X_{-1})^2$. Then, given the estimator $\widehat\Omega= \widetilde\sigma^2(\hat\Theta^M)^T$ we compute the confidence interval for the structural parameter $\beta_1$ by following Example (ref).
We first study the effect of $\rho$ and $\alpha$ on the results of our inference procedure. Here, we take $p = 100$ with one endogenous variable and $q = 100$ exogenous variables (included and excluded covariates). Then, we look at the effect of $\alpha$ when $p=q=200$. The sample size is fixed to $n=100$. The results are in Table (ref). Here, we report the absolute values of the mean bias for the desparsified IV Lasso estimator $\widehat\beta_1$ and for the IV Lasso estimator $\widetilde \beta_1$, for different values of the parameters $\rho$ and $\alpha_1$. The absolute mean is computed over the $1000$ Monte Carlo replications. We also report the coverage of our confidence interval for $\beta_1$ at the nominal level $95\%$. Table (ref) also reports the average coverage of the intervals for individual coefficients corresponding to variables in either $S_0$ or $S_0^c$ computed as follows: $AvgCov_{\alpha}(S_0) = s_0^{-1}\sum_{j\in S_0} \widehat{\mathbb{P}}(\beta_j^0 \in CI_j(\alpha))$ and $AvgCov_{\alpha}(S_0^c) = (p - s_0)^{-1}\sum_{j\in S_0^c} \widehat{\mathbb{P}}(\beta_j^0 \in CI_j(\alpha))$, where $CI_j(\alpha) =[\widehat\beta_j \pm \Phi^{-1}(1-\alpha/2)\, \widehat\Omega_{jj}^{1/2}/n^{1/2}]$ according to Corollary (ref) and $\widehat{\mathbb{P}}$ is obtained as an average over $1000$ Monte Carlo iterations.
From Table (ref) we see that the absolute mean bias of the desparsified IV Lasso estimator $\widehat \beta_1$ is considerably smaller than the absolute mean bias of the IV Lasso estimator $\widetilde\beta_1$ for each value of $\rho$ and $\alpha_1$, and also as $p=q$ increases. As $\alpha_1$ decreases, i.e., the strength of instruments declines, we see that the values of the absolute mean bias of both the desparsified IV Lasso and of the initial lasso estimator $\widetilde\beta_1$ become larger when $p=q=100$. When $p=q=200$ we see that the effect on the instrument strength on mean the bias of the IV Lasso estimator $\widetilde\beta_1$ and our desparsified estimator $\widehat \beta_1$ is mixed. From the third column of Table (ref) we see that the empirical coverage for $\beta_1$ is close to the nominal level of $95\%$. Concerning the coefficients corresponding to variables in $S_0$, we have some undercoverage (see the fourth column of Table (ref)), which is yet less severe when $p=q=200$. Undercoverage for coefficients in $S_0$ has been shown for the desparsified Lasso in reduced from regression by vandegeer2014 in different simulation designs. On the other hand, the coverage for the coefficients corresponding to variables in $S_0^c$ is accurate and even somewhat larger than the nominal coverage probability when $p=q=100$.
Figure (ref) shows the histograms approximating the sampling distribution of our estimator $\widehat\beta_1$ for different values of $\alpha_1$ when $\rho=0.5$. From this figure we see that there is a perfect fit and that for $\alpha_1$ small the distribution is slightly right skewed. We have superposed the probability density function of a standard normal, which corresponds to the asymptotic distribution of the estimator.
Random support of $\beta^0$. As a further exercise, we have analyzed the situation where the support of $\beta^0$ is randomly selected. We fix the cardinality of the active set of $\beta^0$ equal to $s_0 = 15$ and then the support $S_0$ of $\beta^0$ is obtained as $S_0 = \{u_1,\ldots, u_{15}\}$ where $u_1,\ldots, u_{15}$ is a realization of $15$ draws without replacement from $\{1,\ldots,p\}$. The simulation design is as in (ref) but where we present here only the result for $\alpha_1=0.05$. In Table (ref) we report the absolute value of the estimated bias (again computed as the difference between the average over the Monte Carlo iterations and the true value of $\beta_1$) for our desparsified IV Lasso estimator $\widehat\beta_1$ and for the IV Lasso estimator $\widetilde\beta_1$. We see that the coverage of our confidence interval for $\beta_1$ is close to the nominal coverage of $95\%$. Table (ref) also reports the average coverage of the intervals for individual coefficients corresponding to variables in either $S_0$ or $S_0^c$. Again there is some undercoverage for the $S_0$ coefficients.
We generate i.i.d. data from the model (ref) where
where $\Sigma$ is a $q\times q$ matrix that can be set in two different ways. Denote $p_z = q - p_w$, $q = \dim(Z)$ and $p_w = \dim(X_{-1})$, then we have the two following designs for $\Sigma$.
In the rest of this section, we fix the degree of endogeneity and the strength of the instruments by setting $\rho = 0.5$ and $\alpha_1=1$. The other parameters are set as in the previous simulation with homoscedastic errors. The covariance estimator under heteroscedasticity is implemented as the matrix $\widehat\Omega$ in (ref).\\ In Tables (ref) and (ref) we report the absolute value of the estimated bias -- computed as the difference between the average over $1000$ Monte Carlo iterations and the true value of $\beta_1$ given by $2$ -- for our desparsified IV Lasso estimator $\widehat\beta_1$ and for the IV Lasso estimator $\widetilde\beta_1$. Table (ref) refers to Design 1 while Table (ref) refers to Design 2. We see that the bias of $\widehat\beta_1$ is again considerably smaller than the one of the initial IV Lasso estimator $\widetilde\beta_1$ in absolute value. Compared to the bias reported in Table (ref), in presence of heteroskedasticity the bias is larger in absolute value. However the bias of our estimator $\widehat\beta_1$ is less affected by heteroscedasticity than the bias of the initial IV Lasso estimator $\widetilde\beta_1$. In addition, we report the average coverage of our confidence interval for $\beta_1$ at the confidence level of $95\%$. We see that the coverage increases with $q$. We also report the average coverage of the intervals for individual coefficients corresponding to variables in either $S_0$ or $S_0^c$. For the Design 2 we also outline in Table (ref) the effect of augmenting $p$. Figure (ref) repots the histograms relative to Design 2 which show that the distribution of $\widehat\beta_1$ is more and more concentrated around the true value of $\beta_1$ as $n$ and $q$ increase.
This simulation corresponds to Example (ref) about series approximation. Let $\phi^J(X_1):= (\phi_1(X_1), \ldots, \phi_J(X_1))^T$ be a $J$-vector of basis functions. We generate i.i.d. data from the following model
with $(U,V,Z)$ is generated from (ref) again with $\Sigma = \left((0.5)^{|j-k|}\right)_{jk}$. Moreover, $\varphi(X_1) = X_1^2/4$ so that if $(\phi_j)_j$ are polynomials we have that $\phi^J(X_1) = (1, X_1, (X_1/2)^2, \ldots, (X_1/J)^J)$ and $\beta_1 = (0,0,1,0,\ldots,0)^T$. We take $J=10$, $\dim(X_{-1}) = 100 - J$ and $q = 100$ exogenous variables (included and excluded covariates). The parameters are set in the following way: $\beta_{-1,j} = 1 + (j-1)*c$ for $1\leq j \leq 50$, where $c$ is a constant such that the parameters $\beta_{-1,j}$ are equispaced between $1$ and $3$, $\alpha_{1,j}$ has components equispaced between 1 and 0.5, $\beta_{-1,j} = 0$ for $51 \leq j \leq (p-J)$and $\alpha_{-1,j} = j^{-3}/2$ for $1\leq j\leq 50$. The results of this simulation are reported in Table (ref). We report the absolute mean of the bias for the desparsified IV estimator and the initial IV Lasso estimator of the parameter $\beta_{1,3} = 1$, that is, the coefficient of the second order polynomial. We see that we obtain some undercoverage for the coefficient $\beta_{1,3}$ but the coverages for $S_0$-coefficients and $S_0^c$-coefficients are close or beyond the $95\%$ nominal level.
In this section we apply our method to estimate the price coefficient in a logit model of demand for automobiles using market share data. This application follows the empirical illustration in ChernozhukovHansenSpindler2015 and the aim is to estimate the price effect on the market share of a particular car. We consider the following system of equations:
where $s_{it}$ is the market share of product $i$ in market $t$, $s_{0t}$ denotes the outside option, $p_{it}$ is the price which is endogenous, $x_{it}$ are observed product characteristics which are exogenous, and $z_{it}$ is a set of instrumental variables.\\ In our application we use the same product characteristics as in ChernozhukovHansenSpindler2015 and BerryLevinsohnPakes_1995, that is, $x_{it}$ contains an air conditioning dummy, horsepower divided by weight, miles per dollar, vehicle size and a time trend. We center all the variables in order to eliminate the constant. The instruments for price are formed by using the idea developed in BerryLevinsohnPakes_1995 that characteristics of other products satisfy an exclusion restriction of the type $ \mathbb{E}[u_{it}|x_{jt'}] = 0$ for any $t'$ and any $j\neq i$. Therefore, any function of characteristics of other products may be used as an instrument for price. We follow ChernozhukovHansenSpindler2015 and form instruments as
where $x_{k,it}$ and $z_{k,it}$ denote the $k$-th element of $x_{it}$ and $z_{it}$, respectively, and $\mathcal{I}_f$ denotes the set of products produced by firm $f$. In this way we have a set of $10$ excluded instruments.\\ In addition, because economic theory does not specify the functional form in which the elements of $x_{it}$ enter the regression model, we also consider first-order interaction terms of the variables in $x_{it}$, and quadratic and cubic transformations of the continuous variables in $x_{it}$ for a total of $18$ new variables. In this way, we have a vector of augmented controls, denoted by $x_{it}^a$ that contains $x_{it}$ and these new variables. The corresponding vector of augmented excluded instruments is then given by $z_{it}^a$ where $z_{it}^{a}$ is constructed as in (ref) but with $x_{k,jt}$ replaced by $x_{k,jt}^a$. By using the generic notation in the paper: $X = (p_{it},x_{it}^{aT})^T$, $Z = (z_{it}^{aT},x_{it}^{aT})^T$, and $\beta^0 = (\beta_{0}^{0},\beta_1^{0T})^T$.\\ In our data set, we have a total of $n = 2217$ observations, $23$ augmented controls, and $71$ augmented instruments $Z$. According to our theory, the strength of identification is measured through the parameter $\omega$. In the non-augmented framework with $10$ excluded instruments, the estimated $\omega$ is equal to $6.44\cdot10^{-06}$ and hence, relatively small. When we augment the number of controls and instruments the estimate of $\omega$ increases. This means that adding polynomial transformations and interactions, if on the one hand makes the model more flexible, on the other hand reduces the strength of identification, which is not surprising. This is not a problem since our approach is robust to semi strongly identified models.
In our application we estimate the covariance matrix to construct the confidence intervals by using our heteroscedastic robust estimator. In Table (ref) we show the results obtained with different estimators. Together with the point estimate, we also report the lower and upper bound of the $95\%$-confidence interval. We first compute the OLS and 2SLS estimators obtained without augmenting the controls and the instruments. Then, we show the results obtained with augmented controls and instruments with three estimators: the OLS, the 2SLS and our desparsified IV Lasso estimator. To incorporate uncertainty induced by sample splitting for the selection of the tuning parameters, we use the finite-sample adjustments proposed by Chernozhukov2018EJ. Specifically, we present estimation results as the median of desparsified IV Lasso estimate for 200 different sample splits. Here, we retain estimates of our desparsified IV Lasso estimate gives a low number of products with inelastic demand. As we explained below, inelastic demand is unrealistic in a setup of firms maximizing their profit. The standard deviation is computed from the median, over the same $200$ seeds, of the adjusted variances (following the variance adjustment in equation (3.14) in Chernozhukov2018EJ).
We see that the estimated price coefficient becomes larger in absolute value when we move from the Baseline OLS ($-0.0886$ with a standard deviation of $0.0043$) to the Baseline 2SLS ($-0.1419$ with a standard deviation of $0.0119$), which can be interpreted as the fact that the OLS estimator is biased because of endogeneity of price. The magnitude of the OLS estimated price coefficient increases when we use augmented controls (the Augmented OLS estimate is $-0.0991$ with a standard deviation of $0.0046$) and, when we use augmented controls and instruments, the Augmented 2SLS price coefficient estimate is -0.1273 with a standard deviation of $0.0076$. The largest value in absolute value is obtained with our desparsified IV Lasso estimator which gives a price coefficient estimate equal to $-0.2104$ with a (finite-sample adjusted) standard deviation of $0.0306$. The estimated price coefficient that we obtain with our estimator is similar to the one in ChernozhukovHansenSpindler2015 based on the double-selection approach, which is equal to $-0.221$. The latter is contained in our $95\%$-confidence interval for $\beta_0^0$. The slight difference is mainly due to the randomness in choosing the tuning parameters.\\ Notice that as we move from the baseline results to the results based on augmented controls and instruments, the estimates become more plausible from an economic theory point of view. Indeed, in our setup firms maximizing their profit should face elastic demand for all products. In line with this insight, whereas the baseline OLS (resp. 2SLS) point estimates imply inelastic demand for 1502 (resp. $670$) products, our desparsified IV Lasso estimate inelastic demand for only $32$ products using augmented controls and variables. The number of products with inelastic demand are reported on the last column of Table (ref). For our desparsified IV estimator, the number of inelastic is larger than the one based on the double-selection procedure of ChernozhukovHansenSpindler2015 which is equal to $12$. This difference is due to the fact that our estimate of the price coefficient is slightly lower than the double-selection based estimate, as discussed above.\% and have been calculated as the number of products for which $\widehat\beta_{0}^0 p_{it}(1 - s_{it}) > -1$.
Overall, we see that our desparsified IV estimator and inference procedure perform well in empirical applications and give plausible results. In addition, because our procedure is robust to heteroscedasticity, our inference remains valid when the regression error term is heteroscedastic.