EconBase
← Back to paper

Complete Subset Averaging with Many Instruments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

83,290 characters · 11 sections · 78 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Complete Subset Averaging with Many Instruments

abstractWe propose a two-stage least squares (2SLS) estimator whose first stage is the equal-weighted average over a complete subset with $k$ instruments among $K$ available, which we call the complete subset averaging (CSA) 2SLS. The approximate mean squared error (MSE) is derived as a function of the subset size $k$ by the nagar1959bias expansion. The subset size is chosen by minimizing the sample counterpart of the approximate MSE. We show that this method achieves the asymptotic optimality among the class of estimators with different subset sizes. To deal with averaging over a growing set of irrelevant instruments, we generalize the approximate MSE to find that the optimal $k$ is larger than otherwise. An extensive simulation experiment shows that the CSA-2SLS estimator outperforms the alternative estimators when instruments are correlated. As an empirical illustration, we estimate the logistic demand function in \citet*{berry1995automobile} and find the CSA-2SLS estimate is better supported by economic theory than the alternative estimates.\\\\ Keywords: two-stage least squares, many instruments, endogeneity, model averaging, equal-weight.

Introduction

Instrumental variables (IV) estimators are commonly used to estimate the parameters associated with endogenous variables in economic models. The two-stage least squares (2SLS) is the most popular IV estimator for linear regression models with endogenous regressors. While the 2SLS estimator is usually applied to just-identified models where the number of instruments is the same as the number of endogenous variables, there are many applications where more instruments are used than the number of endogenous variables (over-identified). In particular, when the instrument set is a large set of dummy variables or is constructed by interacting the original instruments with exogenous variables, the total number of instruments can be quite large. For example, angrist1991does use as many as 180 instruments for one endogenous variable in one of their specifications to get a tighter confidence interval for the structural parameter than using a smaller set of instruments.

Although using a large set of instruments can improve efficiency, there is a trade-off in terms of increased bias in the point estimate, e.g. kunitomo1980asymptotic, morimune1983approximate, and bekker1994alternative. Motivated by the trade-off relationship, donald2001choosing propose to select the number of instruments by minimizing the approximate mean squared error (MSE) of IV estimators. kuersteiner2010constructing propose an IV estimator that applies the model averaging approach of hansen2007least in the first stage and show that the selected weights attain optimality in the sense of li1987asymptotic. okui2011instrumental proposes to average the first stage using shrinkage to obtain the shrinkage IV estimators. When the number of instruments can be much larger than the sample size, \citet*{belloni2012sparse} propose to select instruments using Lasso in the first stage under approximate sparsity and carrasco2012regularization proposes a regularized IV estimator using all the instruments.

In this paper, we propose a 2SLS estimator whose first stage is the equal-weighted average over a complete subset with $k$ instruments among the total of $K$ instruments. We call this estimator the complete subset averaging (CSA) 2SLS estimator. Our approach differs from the important existing work in the many instruments literature: Unlike donald2001choosing, the CSA-2SLS is based on model averaging in the first-stage rather than model selection and does not require ordering of the instruments; Unlike kuersteiner2010constructing, it does not require weight estimation; Unlike okui2011instrumental, it does not require to specify the main set of instruments a priori.

The main theoretical contribution of this paper is three-fold. First, we derive the approximate MSE for our CSA-2SLS estimator by the nagar1959bias expansion. It is technically challenging because the average of non-nested projection matrices is not idempotent. In contrast, the existing literature usually assumes nested models. The derived formula shows the bias-variance trade-off and some interesting features which will be discussed in detail in the following sections. Second, we generalize the approximate MSE formula when irrelevant instruments exist whose number grows as the sample size increases. A penalty term decreasing with the subset size $k$ appears in the generalized formula, which suggests that the choice of a larger $k$ than otherwise would be desirable in the presence of irrelevant instruments. Third, we prove that the CSA-2SLS estimator with the subset size minimizing the sample approximate MSE is asymptotically optimal in the sense that it attains the lowest possible MSE among the class of the CSA-2SLS estimators with different subset sizes. Our optimality proof is based on li1987asymptotic and whittle1960bounds.

Our approach is motivated by the following observations. First, a set of instruments in economic applications can be correlated with each other, often by construction. In this case, the model averaging approach in the first-stage is more appropriate. We borrow the intuition from the high-dimensional prediction literature. The model selection method, such as Lasso breaks down when predictors are highly correlated. Alternatively, ridge regression works better for highly correlated predictors and \citet*{elliott2013complete} show the connection between the complete subset averaging and ridge regression. Therefore, we believe that the averaging approach can work better with correlated instruments. Figure (ref) shows the correlation structure of instruments in three important studies: 180 instruments (interactions of dummy variables) in \citet*{angrist1991does}, 10 instruments (average characteristics of similar products) in \citet*{berry1995automobile}, and 46 instruments (lagged dependent variables) in \citet*{acemoglu2008income}. The instruments of the first study do not exhibit high correlation, but those of the other two show high correlation. Therefore, it would be empirically relevant to consider the averaging approach suggested in this paper.

figure[figure omitted — 511 chars of source]

Our second motivation is that, in model averaging, estimating weights can cause finite sample efficiency loss, especially when the weight vector's dimension is large. In the forecasting literature it is not surprising to see that equal-weighted averaging outperforms other sophisticated optimal weighting schemes, e.g., see clemen1989combining, stock2004combination, and smith2009simple. Bootstrap aggregating (known as bagging; breiman1996bagging) is another example of equal-weight model averaging and is a popular method in the machine learning literature.

It is worth emphasizing the important work by \citet*{elliott2013complete, elliott2015complete} who propose the equal-weight complete subset regressions in the forecasting context. They demonstrate the complete subset regression's excellent performance relative to competing methods such as ridge regression, Lasso, Elastic Net, bagging, and Bayesian model averaging. We build on their idea of the complete subset regression to provide a formal theoretical justification of the CSA-2SLS estimator with extensive Monte Carlo simulations. Indeed, we find that the CSA-2SLS exhibits potentially huge gains in terms of the bias and the MSE relative to existing methods especially when the instruments are correlated and there is large endogeneity.

There are two limitations to be noted. First, conditionally homoskedastic errors are assumed in the derivation of the approximate MSE, similarly in donald2001choosing, kuersteiner2010constructing, hansen2007least, okui2011instrumental, and \citet*{carrasco2012regularization}. This is required to obtain the explicit order of the higher-order terms in the expansion of the MSE, which allows one-to-one comparison with the existing literature. \citet*{donald2009choosing} derive the approximate MSE under heteroskedasticity for the efficient generalized method of moment (GMM) and the generalized empirical likelihood (GEL) estimators at the expense of more complicated expressions. Also, note that the 2SLS is no longer efficient under heteroskedasticity.

Second, we focus on the 2SLS estimator because of its utmost popularity among applied researchers. The limited information maximum likelihood (LIML) estimator has gained considerable attention recently because of its theoretical advantage over the 2SLS having a smaller bias under the many instruments asymptotics (donald2001choosing).

commentOur simulation result shows that the CSA-2SLS estimator exhibits a very small bias across different specifications of the data-generating processes. Thus, although our analysis can be extended to the $k$-class estimators, including the LIML and the bias-corrected 2SLS, we expect that improvements would be smaller compared to the 2SLS.

To maintain our focus on the new averaging method, however, we defer any extension to the $k$-class estimators such as LIML and bias-corrected 2SLS to future research.

Finally, we summarize related literature. The model averaging approach becomes prevalent in the econometrics literature. hansen2007least shows that the weight choice based on the Mallows criterion achieves optimality. hansen2012jackknife propose the jackknife model averaging, which allows heteroskedasticity. \citet*{ando2014model} present a model averaging for high-dimensional regression. ando2017weight and \citet*{zhang2016optimal} consider the class of generalized linear models to show the optimality under the Kullback-Leibler loss function. zhang2018spatial propose a model averaging in the spatial autoregressive models. kitagawa2016model propose the propensity score model averaging estimator for the average treatment effects for treated. lee2015averaged propose a model averaging approach over complete subsets in the second stage IV regression.

Our approach is different from the many weak instruments asymptotics in chao2005consistent, stock2005asymptotic, han2006gmm, andrews2007testing, and \citet*{hansen2008estimation}, where the concentration parameter can grow slower than the sample size. Alternative estimators under heteroskedasticity and many instruments are proposed by \citet*{hausman2011properties} and \citet*{hausman2012instrumental}. kuersteiner2012kernel extends the instrument selection criteria of donald2001choosing to the time series setting and proposes a GMM estimator using lags as instruments. kang2018higher derives the approximate MSE of IV estimators with locally invalid instruments. \citet*{antoine2014conditional} and escanciano2017simple propose estimators free from choice variables by adopting the continuum of unconditional moment condition in the first stage.

The remainder of the paper is organized as follows. Section (ref) describes the model and proposes the CSA-2SLS estimator. Section (ref) derives the approximate MSE formula of the estimator and investigates its properties. We also show the asymptotic optimality result as well as the implementation procedure. Section (ref) extends the result by allowing an increasing number of irrelevant instruments. Section (ref) investigates the relationship between our approximate MSE and that of \citet*{donald2001choosing} when the instruments are orthogonal. Section (ref) studies the finite sample properties via Monte Carlo simulations. Section (ref) provides an empirical illustration. The online supplement contains proofs and additional simulation results.

Model and Estimator

We follow the setup of donald2001choosing and kuersteiner2010constructing. The model is

eqnarray[eqnarray omitted — 379 chars of source]

where $y_{i}$ is a scalar outcome variable, $Y_{i}$ is a $d_{1}\times 1$ vector of endogenous variables, $x_{1i}$ is a $d_{2}\times 1$ vector of included exogenous variables, $z_{i}$ is a vector of exogenous variables (including $x_{1i}$), $\ensuremath{\varepsilon}_{i}$ and $u_{i}$ are unobserved random variables with finite second moments which do not depend on $z_{i}$, and $f(\cdot)$ is an unknown function of $z$. Let $f_{i} = f(z_{i})$ and $d=d_{1}+d_{2}$. The second equation represents a nonparametric reduced form relationship between $Y_{i}$ and the exogenous variables $z_{i}$, with $E[\eta_{i}|z_{i}]=0$ by construction. Define the $N\times1$ vectors $y = (y_{1},\ldots,y_{N})'$, $\ensuremath{\varepsilon} = (\ensuremath{\varepsilon}_{1},\ldots,\ensuremath{\varepsilon}_{N})'$, and the $N\times d$ matrices $X=(X_{1},\ldots,X_{N})'$, $f=(f_{1},\ldots,f_{N})'$, and $u=(u_{1},\ldots,u_{N})'$.

The set of instruments has the form $Z_{K,i}\equiv (\psi_{1}(z_{i}),\ldots,\psi_{K}(z_{i}),x_{1i})'$, where $\psi_{k}$'s are functions of $z_{i}$ such that $Z_{K,i}$ is a $(K(N)+d_{2})\times1$ vector of instruments. The total number of instruments $K(N)$ increases as $N\rightarrow\infty$ but we suppress the dependency on $N$ and write $K$ unless we need to express the dependence of $K$ on $N$ explicitly. Define its matrix version as $Z_{K}=(Z_{K,1},\ldots,Z_{K,N})'$. To define the complete subset averaging estimator, consider a subset of $k$ excluded instruments. Note that the included exogenous variables $x_{1i}$ are always included in the instruments set. To avoid confusion, we will use “instruments” to refer to excluded instruments throughout the paper. The number of subsets with $k$ instruments is \[\binom{K}{k}=\frac{K!}{k!(K-k)!}.\] A complete subset with size $k$ is the collection of all these subsets. Let $M(K,k) = \binom{K}{k}$, which is the number of models given $K$ and $k$. For brevity, the dependence of $M$ on $K$ and $k$ will be suppressed unless it is necessary. For any model $m$ with $k$ instruments, let $Z_{m,i}^{k}$ be an $(k+d_{2})\times 1$ vector of subset instruments including $x_{1i}$ where $d_{1}\leq k\leq K$ and $Z_{m}^{k}=(Z_{m,1}^{k},\ldots,Z_{m,N}^{k})'$ be an $N\times (k+d_{2})$ matrix for $m=1,\ldots, M$. For each $m$, the first stage equation can be rewritten as

equation[equation omitted — 79 chars of source]

or equivalently,

equation[equation omitted — 55 chars of source]

where $\Pi_{m}^{k}$ is the $(k+d_{2})\times d$ dimensional projection coefficient matrix for model $m$ with $k$ instruments, $u_{m,i}^{k}$ is the projection error, and $u_{m}^{k}=(u_{m,1}^{k},\ldots,u_{m,N}^{k})'$. The projection coefficient matrix is estimated by

equation[equation omitted — 98 chars of source]

and the average fitted value of $X$ over complete subset $k$ becomes

equation[equation omitted — 220 chars of source]

where $P_{m}^{k} = Z_{m}^{k}(Z_{m}^{k'}Z_{m}^{k})^{-1}Z_{m}^{k'}$ and $P^{k} = \frac{1}{M}\sum_{m=1}^{M}P_{m}^{k}$. The complete subset averaging (CSA) 2SLS estimator with a given $k$ is now defined as

equation[equation omitted — 109 chars of source]

The matrix $P^{k}$ is the average of projection matrices. We call this matrix as the complete subset averaging $P$ (CSA-$P$) matrix. Note that the CSA-$P$ matrix is symmetric but not idempotent in general.

Therefore, we can estimate the model by the CSA-2SLS estimator for a given $k$. We will deliberate on the choice of $k$ in the next section and close this section by characterizing the CSA-2SLS estimator and the CSA-$P$ matrix in a broader context.

The CSA-2SLS estimator can be interpreted as the minimizer of the average of 2SLS criterion functions. For each subset of instruments $Z_{m,i}^{k}$, the corresponding moment condition is

equation[equation omitted — 121 chars of source]

The standard 2SLS estimator given the moment condition (ref) minimizes

equation[equation omitted — 111 chars of source]

This equation is the GMM criterion with the weight matrix $(Z_{m}^{k'}Z_{m}^{k})^{-1}$. Conventional model average estimators are based on a weighted average of $\widehat{\beta}^{k}_{m}$, the minimizer of (ref), over different models. For this type of model averaging estimators, see hansen2007least for the OLS estimator, lee2015averaged for the 2SLS estimator, and \citet*{chen2016averaging} for the general approach based on moment conditions.

In contrast, the CSA-2SLS estimator minimizes the average of (ref) over different models directly:

equation[equation omitted — 192 chars of source]

The optimal model averaging 2SLS estimator of kuersteiner2010constructing can be interpreted similarly as the minimizer of the average of 2SLS criterion functions with data-dependent weights.

Subset Size Choice and Optimality

In this section, we discuss how to choose the subset size $k$. First, we derive the approximate MSE of the CSA-2SLS estimator by the nagar1959bias expansion. This high-order expansion shows the bias-variance trade-off that helps improve the finite sample performance of the estimator (e.g. donald2001choosing, kuersteiner2010constructing, and carrasco2012regularization). The subset size $k$ is chosen to minimize the sample counterpart of the approximate MSE whose formula is provided. We prove the optimality of the chosen subset size in the sense of li1987asymptotic.

Approximate MSE

In our analysis, the CSA-$P$ matrix $P^{k}$ plays an important role. Since $P^{k}$ is not idempotent unless $k=K(N)$, we cannot directly apply the existing technique in donald2001choosing or kuersteiner2010constructing to our analysis. We first list regularity conditions. Let $\|A\|=\sqrt{\text{tr}(A'A)}$ denote the Frobenius norm for a matrix $A$.

assumption\ \begin{enumerate} • $\{y_{i},X_{i},z_{i}\}$ are i.i.d. with finite fourth moment and $E[\ensuremath{\varepsilon}_{i}^{2}]=\sigma_{\varepsilon}^{2}>0$. • $E[\ensuremath{\varepsilon}_{i}|z_{i}]=0$ and $E[u_{i}|z_{i}]=0$. • Let $u_{ia}$ be the $a$th element of $u_{i}$. Then $E[\ensuremath{\varepsilon}_{i}^{r}u_{ia}^{s}|z_{i}]$ are constant and bounded for all $a$ and all $r,s\geq0$ and $r+s\leq4$. • $f_{i}$ is bounded. • $\overline{H} = Ef_{i}f_{i}'$ exists and is nonsingular. • For each $k\geq d$ and all $m=1,\ldots,M$, $Z_{m}^{k'}Z_{m}^{k}$ is nonsingular with probability approaching one. • Let $P_{ii}^{k}$ denote the $i$th diagonal element of $P^{k}$. Then $\max_{i\leq N}P_{ii}^{k}\xrightarrow{p}0$ as $N\rightarrow\infty$. \end{enumerate}
assumptionLet $c>0$ be given. There exists a sequence $\{\underline{k}(N)\}$ such that for all $k(N)\geq \underline{k}(N)$ there exists $\Pi_m^{k(N)}$ that satisfies \[\frac{1}{M}\sum_{m=1}^{M}E\|f(z_{i})-\Pi_{m}^{k(N)'}Z_{m,i}^{k(N)}\|^{2} < c\] for large enough $N$.

Assumption (ref) collects standard moment and identification conditions similar to those of donald2001choosing and kuersteiner2010constructing.

Assumption (ref) controls the speed of the lower bound $\underline{k}(N)$ such that $f(z_i)$ is arbitrarily well-approximated by the size-$k(N)$ complete subset projections. This assumption is not particularly stronger than Assumption 2(ii) of donald2001choosing since they are equivalent when $\underline{k}(N)=K(N)$. The approximation error can be small even with $\underline{k}(N) < K(N)$ when instruments are correlated. For example, consider that $f(z_{i})=z_i'\pi_0+o_p(1)$ and all elements of $z_i$ are the linear transformation of $z_{1i}$. Then, Assumption 2 holds with $k=1$ and $M=K$. Below we provide an example where $k(N)$ can grow at arbitrarily slow rate when the instruments are equally correlated at the rate of $\rho$. It is also possible that Assumption (ref) holds only if ${k}(N)/K(N)\rightarrow 1$. A leading case is when the instruments are orthogonal. We provide a detailed analysis for orthogonal instruments in Section (ref). In sum, Assumption 2 can hold for a fixed $k$, a growing $k(N)$ slower than $K(N)$, or a $k(N)$ growing as fast as $K(N)$, depending on the correlation structure of the instruments.

We require Assumption (ref) only to derive a valid approximate MSE expression, which gives an intuition on how CSA-2SLS can reduce MSE. It also provides a data adaptive method to choose the subset size $k$ balancing between the bias and the variance in a finite sample. When the bias is relatively large, $k<K$ can be chosen as we confirm both in the simulation studies in Section (ref) and in the empirical application in Section (ref). We also provide an alternative method of choosing $k$ other than the approximate MSE in Section (ref) and investigate the approximate MSE under different circumstances in Section (ref).

example[Equicorrelated Instruments] We provide a special example to further discuss Assumption (ref). Consider the function \[f(z_{i}) = \gamma_{K}\sum_{j=1}^{K}z_{ij}\] for $\gamma_{K}\neq0$, $K>1$, and $dim(f)=1$, where $z_{ij}$'s have mean zero, the unit variance, and the constant correlation $\rho$. Since \begin{equation*} Var(f(z_{i})) = \gamma_{K}^{2}K(1 + (K-1)\rho), \end{equation*} where $1+(K-1)\rho>0$, which is the smallest eigenvalue of the covariance matrix of $Z$, we set \begin{equation*} \gamma_{K} = \frac{1}{\sqrt{K(1 + (K-1)\rho)}} \end{equation*} so that the variance of $f(z_i)$ is fixed at 1. Setting $\gamma_{K}$ as a decreasing sequence in $K$ does not necessarily imply weak instruments. Since $K^2/N\rightarrow 0$, $\gamma_K$ converges to zero at a slower rate than $1/\sqrt{N}$ so that $\sqrt{N}$-consistency of the CSA-2SLS estimator is still maintained. This sequence of instruments corresponds to “nearly weak” instruments of \citet*{hahn2002discontinuities} and \citet*{antoine2009efficient}. Note that the weak instruments sequence of \citet*{staiger1997instrumental} assumes that the first-stage coefficients are $O(1/\sqrt{N})$ and 2SLS estimators are not $\sqrt{N}$-consistent. The minimized mean squared approximation error of the CSA with $k$ instruments\footnote{Since the minimization problem is symmetric in $\pi$, we set $\pi_{1}=\pi_{2}=\cdots=\pi_{k}=\pi$ and solve the first-order condition to obtain the minimizer and (ref).} is \begin{align} &\min_{{\Pi_{m}^{k}}}\frac{1}{M}\sum_{m=1}^{M}E\|f(z_{i})-\Pi_{m}^{k'}Z_{m,i}^{k}\|^{2} \notag \\ &=\min_{\pi_{1},\cdots,\pi_{k}}E[(f(z_{i})-\pi_{1} z_{i1} - \pi_{2} z_{i2} - \cdots -\pi_{k} z_{ik})^{2}] \notag \\ &= \frac{1-\rho}{1+(k-1)\rho}\left(1-\frac{k}{K}\right). \end{align} Some interesting observations can be made for the expression (ref). First, the approximation error (ref) converges to zero so long as $k = O(K^{\alpha})$ for $0<\alpha\leq1$ as $K\rightarrow\infty$. Since $k$ can grow at an arbitrarily slow rate if $\rho\neq0$, at least for this example, Assumption 2 imposes a very mild condition on $k$. In contrast, if $\rho=0$ or $\rho=1/K$, then (ref) converges to zero if and only if $k/K\rightarrow1$, which would be the least favourable scenario for the CSA-2SLS estimator. Second, for any given $(k,K)$, the approximation error gets smaller as $\rho$ increases and it is equal to zero when $\rho=1$. Figure (ref) shows the approximation errors when $K=20$ and $k=1,3,5$. Third, the approximation error is decreasing with $k$ uniformly in $\rho$, which is shown in Figure (ref). Evaluating (ref) at $k$ and $k+1$, respectively, the difference is $(1-\rho)(1+(K-1)\rho)/(1+\rho)\geq 0$. \begin{figure}[t] \caption{Approximation error when $K=20$: $k=1$ (blue), $k=3$ (red), $k=5$ (yellow)} \end{figure}

We now provide the approximate MSE formula of CSA-2SLS. Define $H = f'f/N$ and $P_{f} = f(f'f)^{-1}f'$.

theoremSuppose that Assumptions (ref)-(ref) are satisfied, $k^{2}/N\rightarrow0$, and $E[u_{i}\ensuremath{\varepsilon}_{i}|z_{i}] = \sigma_{u\varepsilon}\neq0$. Then, we have the following results for the CSA-2SLS estimator: \begin{eqnarray} N(\widehat{\beta}-\beta)(\widehat{\beta}-\beta)' &=& \widehat{Q}(k) + \widehat{r}(k),\\ E\left[\widehat{Q}(k)|Z\right] &=& \sigma_{\ensuremath{\varepsilon}}^{2}H^{-1} + S(k) + T(k), \\ \frac{\widehat{r}(k)+T(k)}{tr(S(k))} & = & o_p(1), as k \rightarrow \infty, N \rightarrow \infty, \end{eqnarray} with \begin{equation} S(k) = H^{-1}\left[\sigma_{u\ensuremath{\varepsilon}}\sigma_{u\ensuremath{\varepsilon}}'\frac{k^{2}}{N} + \sigma_{\ensuremath{\varepsilon}}^{2}\frac{f'(I-P^{k})(I-P_{f})(I-P^{k})f}{N}\right]H^{-1}. \end{equation}

In the expansion, $\widehat{r}(k)$ and $T(k)$ are terms of smaller order in probability than those in $S(k)$ and that the term $\sigma_{\ensuremath{\varepsilon}}^{2}H^{-1}$ is the first-order asymptotic variance under homoskedasticity. The first term of $S(k)$ (ignoring pre- and post-multiplied $H^{-1}$), $\sigma_{u\ensuremath{\varepsilon}}\sigma_{u\ensuremath{\varepsilon}}'k^{2}/N$, corresponds to the bias and is similar to that of donald2001choosing. The second term is different from the usual higher-order variance in similar expansions in the literature. Let $V(k) = f'(I-P^{k})(I-P_{f})(I-P^{k})f/N$. It can be decomposed as

equation[equation omitted — 150 chars of source]

In donald2001choosing, the higher-order variance term, which takes the form of $f'(I-P^{K})f/N$ where $P^{K}$ is the projection matrix consisting of $K$ instruments, decreases with $K$. This gives the bias-variance trade-off in their expression. In contrast, $V(k)$ is the sum of two components and the monotone decrease with respect to $k$ is not always guaranteed.

comment\todo[inline]{ CSA-2SLS does not require to determine the exact growth rate of $\underline{k}(N)$ since it selects a higher $k$ automatically when the approximation error is quite large for a fixed sample size $N$. For example, if the higher order variance caused by the approximation error dominates the approximate MSE, then CSA-2SLS would select $k=K$. However, it is more common in practice to choose $k<K$ to exploit the benefit of correlated $z_i$, which is confirmed in the simulation studies and the empirical applications below. }

We further elaborate on the behavior of the higher-order variance term $V(k)$ using shrinkage. Since $I-P_{f}$ is idempotent, we can write

equation[equation omitted — 60 chars of source]

where $u^k\equiv(I-P^k)f$. There exist two factors that force to move $V(k)$ into opposite directions as $k$ varies. First, the norm of $u^k$ decreases as $k$ increases since $P^{k^{*}}$ for $k^{*}>k$ is the average of the projection matrices on the space spanned by a larger set of instruments. Second, $P^k$ can cause less shrinkage for $P^{k}f$ for a larger $k$, which makes the norm of $(I-P_f)u^k$ larger. Note that

equation[equation omitted — 112 chars of source]

and that $P^k f$ with less shrinkage lies farther away from the space spanned by $f$. Therefore, if the second factor dominates the first then $V(k)$ increases along with $k$.

comment\begin{example} Let $N=3$ and $K=2$. Suppose that $f=(2,2,1)'$, $Z_1=(1,0,0)'$, $Z_2=(0,1,0)'$, and $Z = [Z_1, Z_2]$. Since $P^{1} = (Z_{1}(Z_{1}'Z_{1})^{-1}Z_{1}' +Z_{2}(Z_{2}'Z_{2})^{-1}Z_{2}' )/2$ and $P^{2} = Z(Z'Z)^{-1}Z'$, we calculate $u^1=(I-P^1)f=(1,1,1)$ and $u^2=(I-P^2)f=(0,0,1)'$. Thus, $\Vert u^{2} \Vert < \Vert u^{1}\Vert$. However, it is clear from the figure below that $u^1$ is closer to $f$ than $u^2$. Indeed, we have $\Vert (I-P_f)u^1\Vert \approx 0.47$ and $\Vert (I-P_f)u^2\Vert \approx 0.94$. \end{example}

If there are more than one endogenous variable, then we choose $k$ that minimizes a linear combination of the MSE, $S_{\lambda}(k)\equiv \lambda' S(k)\lambda$ for a user-specified $\lambda$ similarly in donald2001choosing.

commentHowever, for a model containing only one endogenous variable, the choice of $\lambda$ is irrelevant. To see this, first observe that $\sigma_{u \ensuremath{\varepsilon}} = \sigma_{\eta\ensuremath{\varepsilon}}e_{1}$ where $e_{1}$ is the first unit vector. In addition, the components of $f$ that are included in the main equation lie on the column space spanned by the instruments (including the included exogenous variables) so that \begin{align*} f'(I-P^{k})(I-P^{k})f &= \overline{Y}'(I-P^{k})(I-P^{k})\overline{Y}e_{1}e_{1}',\\ f'(I-P^{k})P_{f}(I-P^{k})f &= \left(\overline{Y}'(I-P^{k})\overline{Y}\right)^{2}\left(e_{1}'H^{-1}e_{1}\right)\frac{1}{N}e_{1}e_{1}', \end{align*} where $\overline{Y} = (E[Y_{1}|z_{1}],\ldots,E[Y_{N}|z_{N}])'$. Thus we can write \begin{equation*} \lambda'S(k)\lambda = \left(\lambda'H^{-1}e_{1}\right)^{2}\left[\sigma_{\eta\ensuremath{\varepsilon}}^{2}\frac{k^{2}}{N} + \sigma_{\ensuremath{\varepsilon}}^{2}\left(\frac{\overline{Y}'(I-P^{k})(I-P^{k})\overline{Y}}{N}-\left(\frac{\overline{Y}'(I-P^{k})\overline{Y}}{N}\right)^{2}e_{1}'H^{-1}e_{1}\right)\right], \end{equation*} which shows that the minimization problem with a single endogenous variable does not depend on $\lambda$.

Implementation and Optimality of Approximate MSE

The approximate MSE $S(k)$ in (ref) gives useful guidance on the data-adaptive choice of $k$. Both terms in $S(k)$ converge to zero asymptotically, but there is a trade-off between the higher-order bias and variance terms in finite samples. Similarly, in donald2001choosing and kuersteiner2010constructing, the approximate MSE sheds light on the choice of $k$ only when there exists a non-negligible size of bias for a given sample size $N$ with many instruments.

We first introduce notation to construct the sample counterpart of the approximate MSE. Let $\widetilde{\beta}$ be a preliminary estimator that is fixed across different values of $k$ and $\widetilde{\varepsilon}=y-X\widetilde{\beta}$. Let $\widetilde{f}$ be an estimate of $f$. The residual matrix is denoted by $\widetilde{u}=X-\widetilde{f}$ and $\widetilde{u} = (u_1,u_2,\ldots,u_N)'$ where $\widetilde{u}_{i}$ is a $d\times1$ vector. Define $\widetilde{H} = \widetilde{f}'\widetilde{f}/N$, $\widetilde{\sigma}_{\ensuremath{\varepsilon}}^{2} = \widetilde{\ensuremath{\varepsilon}}'\widetilde{\ensuremath{\varepsilon}}/N$, $\widetilde{\sigma}_{u\ensuremath{\varepsilon}} = \widetilde{u}'\widetilde{\ensuremath{\varepsilon}}/N$, $\widetilde{\sigma}_{\lambda\ensuremath{\varepsilon}} = \widetilde{\lambda}'\widetilde{H}^{-1}\widetilde{\sigma}_{u\ensuremath{\varepsilon}}$, and $\widetilde{\Sigma}_u=\widetilde{u}'\widetilde{u}/N$. The feasible criterion function for $S_{\lambda}(k)$ is defined as below:

align*[align* omitted — 426 chars of source]

where

align*[align* omitted — 460 chars of source]

Then, we choose $\widehat{k}$ as a minimizer of the feasible approximate MSE function, $\widehat{S}_{\lambda}(k)$. Given $\widehat{k}$, we can compute the CSA-2SLS estimator defined in (ref).

We remark on two practical issues in implementation. First, our theory requires the preliminary estimator $\widetilde{\beta}$ be consistent, and it is simply achieved by 2SLS estimation with any valid IVs. However, the performance of the estimator in a finite sample might rely on the choice. Following donald2001choosing and kuersteiner2010constructing, we propose choosing the set of IVs based on the first stage Mallows criterion, which brings satisfactory performance results in the extensive simulations experiments reported in Section (ref). Second, it could be infeasible to estimate $P^k$ when $K$ is too large. Note that the number of complete subsets (all the subsets with different $k$'s) grows exponentially, $2^K-1$. To deal with this computational issue, we propose to draw a smaller number of submodels randomly as follows. Let $\mathcal{R}^k$ be a class of $R$ subsets with $k$ elements that are randomly selected from $\{1,\ldots,M\}$. Then, the CSA projection matrix with randomly drawn submodels is defined as $\breve{P} = R^{-1} \sum_{r\in\mathcal{R}^k} P^k_r$, where $P^k_r$ is a projection matrix using the set of IVs indexed by $r$. Table (ref) in Section (ref) of the Appendix shows simulation results such that randomly drawn submodels work well with a feasible size of $R$.

We finalize this subsection by proving that the CSA-2SLS estimator achieves the asymptotic optimality in the sense that it minimizes the MSE among the class of CSA-2SLS estimators with different complete subset sizes. We collect additional regularity conditions below.

assumption\ \begin{enumerate} • $\widetilde{\sigma}^2_{\ensuremath{\varepsilon}}\xrightarrow{p} \sigma^2_{\ensuremath{\varepsilon}}$, $\widetilde{\sigma}^2_{u\ensuremath{\varepsilon}}\xrightarrow{p} \sigma^2_{u\ensuremath{\varepsilon}}$, $\widetilde{\lambda} \xrightarrow{p} \lambda$, $\widetilde{H} \xrightarrow{p} \overline{H}$, and $\lambda'\overline{H}^{-1}\sigma_{u\ensuremath{\varepsilon}} \neq0$. • $\lim_{N\rightarrow \infty} \sum_{k=1}^{K(N)} \left(N S_{\lambda}(k)\right)^{-1} =0$ almost surely in $Z$. • There exist a constant $\phi \in (0,1/2)$ such that $\|\widetilde{\Sigma}_u- \Sigma_u\| = O_p(N^{-1/2+\phi}S_{\lambda}(k)^{\phi})$ where $\Sigma_{u} = E[u_{i}u_{i}'|z_{i}]$. \end{enumerate}

Assumptions (ref)(i) and (ii) are similar to Assumptions 4--5 in donald2001choosing. Assumption (ref)(i) is a high-level assumption on the consistency of the preliminary estimators. Assumption (ref)(ii) is a standard assumption in the model selection or model averaging literature (see, e.g.\ Assumption (A.3) in li1987asymptotic). This condition excludes the case that $f$ is perfectly explained by a finite number of instruments. Assumption (ref)(iii) is a mild requirement on the convergence rate of $\widetilde{f}$ that the standard series estimator easily satisfies. For example, consider a series estimator with $\widetilde{k}$ b-spline bases as in newey1997convergence, where $\sup_{z} \Vert f(z) - \Pi^{\widetilde{k} \prime}_0 Z_{\widetilde{k}}(z) \Vert =O_p(\widetilde{k}^{-\alpha})$ where $\alpha=1$ and $Z_{\widetilde{k}}(z)\equiv(\psi_{1}(z),\ldots,\psi_{K}(z),x_{1})'$. By choosing the optimal rate for $\widetilde{k}$, we get the rate for the error variance, $\Vert \widetilde{\Sigma}_u - \Sigma_u \Vert =O_p(N^{-1/3})$. Since $S_{\lambda}(k)=o_p(1)$, we can write $S_{\lambda}(k)=O_p(N^{-c})$ for some $c>0$. Then, Assumption (ref)(iii) requires $\Vert \widetilde{\Sigma}_u - \Sigma_u \Vert =O_p(N^{-1/2+\phi-c})$. Comparing this rate with $O_p(N^{-1/3})$ above, we can find $\phi >1/6+c$.

The following result shows that the proposed estimator is optimal among the class of the CSA-2SLS estimators.

theoremUnder Assumptions (ref), (ref), and (ref), as $N\rightarrow\infty$, \begin{align} \frac{S_{\lambda}(\widehat{k})}{\min_{k } S_{\lambda}(k)} \xrightarrow{p} 1. \end{align}

Cross-validation Method

In this subsection we consider an alternative method of choosing the subset size $k$. The approximate MSE provides useful guidance in choosing $k$. More importantly, it gives an intuition on how CSA-2SLS could reduce MSE further by providing the expression for the higher-order bias and variance. However, the theoretical illustration in the previous subsections depends on a few simplifying assumptions: homoskedasticity and the asymptotic approximation of $f(z_i)$ by the complete subset model averaging (Assumption (ref)). When they are violated, the subset size choice based on the approximate MSE improves the finite sample property of CSA-2SLS less than expected. Therefore, it is worthwhile to propose some alternative methods for the subset size choice.

We propose to choose $k$ by the cross-validation (CV) method that minimizes the first-stage mean squared prediction error. We explain it by the leave-one-out cross-validation. This can be immediately extended to any $b$-fold cross-validation, where $b$ is the number of subsamples. For $k=1,\ldots, K$, we define the following cross-validation objective function:

align*[align* omitted — 150 chars of source]

where $\widehat{\Pi}_{m,-i}^{k \prime}$ is the jackknife estimator for $\widehat{\Pi}_{m}^{k \prime}$, which is estimated by (ref) without using the $i$-th observation $(X_i, Z_i)$. Then, we choose $k$ that minimizes $CV_n(k)$. lee2020complete apply a similar cross-validation method in the context of quantile prediction and show that $\widehat{k}$ is asymptotically equivalent to the infeasible best subset size. However, the optimality result is not for the second-stage MSE but for the prediction error of the first-stage regression. Therefore, the CV method is robust to the requirement of homoskedasticity and Assumption (ref) but it might perform worse than the approximate MSE when those conditions are satisfied.

commentThird, we can adopt information-based criteria such as Akaike information criterion (AIC) or Bayesian information criterion (BIC) in the first-stage regression and choose $k$. The same remarks for the CV method apply: it will be more robust to the requirement of some regularity conditions but it does not guarantee the optimal choice of $k$ that minimizes the second stage MSE. We note that all these alternative methods relax some stringent assumptions required for the approximate MSE analysis but they are based on heuristics and cast many open questions for their theoretical properties. To tackle on those questions is beyond the scope of this paper and we leave them for future research.

Approximate MSE under Different Circumstances

Irrelevant Instruments

The higher-order MSE expansion in the previous section assumes an increasing sequence of instruments where the instruments are strong enough for $\sqrt{N}$-consistency and the existence of higher-order terms. This implies that the concentration parameter, which measures the strength of the instruments, increases at the rate of $N$. An important feature of the CSA-2SLS is that no instrument is excluded either in finite samples or asymptotically due to equal-weighted averaging. Assume that the econometrician does not know whether irrelevant instruments are included or not. Although irrelevant instruments can be excluded when a pre-screening method is available, it is important to investigate whether the CSA-2SLS estimator and the MSE expansion are still valid if irrelevant instruments happen to be included in the averaging. In contrast, irrelevant instruments can be excluded by zero or negative weights in kuersteiner2010constructing.

In this section, we derive the MSE under a growing set of irrelevant instruments as $N$ and $K$ increase. This makes the concentration parameter smaller in level for given $N$ and $K$ than the case where all the instruments are relevant. By comparing the MSEs with and without irrelevant instruments we can analyze the effect of irrelevant instruments on the MSE. Under the many weak instruments asymptotics of chao2005consistent, the growth rate of the concentration parameter can be slower than $N$. In this case, the 2SLS estimator is not $\sqrt{N}$-consistent anymore and it would be difficult to compare the higher-order expansion with the other case. Thus, we leave the analysis with many weak instruments for future research.

For the sake of simplicity, we assume that there is no exogenous variables, $x_{1i}$, so the model simplifies to

eqnarray[eqnarray omitted — 123 chars of source]

where $X_{i}$ is a $d\times 1$ vector of endogenous variables and $z_{i}$ is a vector of exogenous variables. We set $f(z_{i}) = E[X_{i}|z_{i}]$. We divide $z_{i}$ into the relevant and irrelevant ones, $z_{1i}$ and $z_{2i}$ and make the following definition:

definition[Definition 4.1] A vector of instruments $z_{2i}$ is irrelevant if $E[X_{i}|z_{1i},z_{2i}] = E[X_{i}|z_{1i}]$.

Under Definition (ref), the set of instruments $Z_{K,i}$ can be divided into two sets, relevant and irrelevant instruments. Let $K_{1}$ and $K_{2}$ be the number of relevant and irrelevant instruments, respectively.

We introduce additional definitions. Let $\mathcal{M}(K,k) =\left\{1,2,\ldots,M(K,k)\right\}$ be the index set for all subsets with $k$ instruments and let $\mathcal{M}_{1}(K,K_{1},k)$ and $\mathcal{M}_{2}(K,K_{2},k)$ be the index sets for subsets with at least $d$ relevant instruments and those with less than $d$, respectively. By construction, $\mathcal{M}_{1}(K,K_{1},k)\cup\mathcal{M}_{2}(K,K_{2},k) =\mathcal{M}(K,k)$. Let $M_{1}(K,K_{1},k)$ and $M_{2}(K,K_{2},k)$ be the cardinality of $\mathcal{M}_{1}(K,K_{1},k)$ and $\mathcal{M}_{2}(K,K_{2},k)$, respectively. For brevity, we suppress the dependence on $K$, $K_{1}$, $K_{2}$, and $k$.

Each element of the index sets $\mathcal{M}_{1}$ and $\mathcal{M}_{2}$ corresponds to a relevant and an irrelevant first stage model (a set of $k$ instruments). Since there are $d$ endogenous variables, those models in $\mathcal{M}_{1}$ satisfy the order condition for identification and (completely) approximate $f(z_{i})$ as $k$ increases. In contrast, those irrelevant models in $\mathcal{M}_{2}$ do not approximate $f(z_{i})$ as $k$ increases.

We make assumptions that will replace Assumption (ref).

assumption\ \begin{enumerate} • Let $c>0$ be given. There exists a sequence $\{\underline{k}(N)\}$ such that for all $k(N)\geq \underline{k}(N)$ there exists $\Pi_m^{k(N)}$ that satisfies \[\frac{1}{M_1}\sum_{m \in \mathcal{M}_1}E\|f(z_{i})-\Pi_{m}^{k(N)'}Z_{m,i}^{k(N)}\|^{2} < c\] for large enough $N$. • For each $k\geq d$ and some positive constant $C$, $0 <C^{-1}\leq M_{1}/M$. • $M_{2}/M=o(1/\sqrt{k})$ as $N\rightarrow\infty$. \end{enumerate}

Assumption (ref)(i) is a version of Assumption (ref) with relevant instruments. Assumptions (ref)(ii)-(iii) are restrictions on the proportion of relevant and irrelevant first-stage models. Assumption (ref)(ii) sets a non-zero lower bound for the proportion of relevant first-stage models but still allows an increasing number of irrelevant instruments. Assumption (ref)(iii) states that the proportion of irrelevant models decreases to zero at the $1/\sqrt{k}$ rate. Since $M_{2}(K,K_{2},k)=0$ for $k\geq K_{2}+d-1$ by construction (e.g. if $d=1$, $K=10$, and $K_{2}=5$, one cannot choose a subset of $6$ or more irrelevant instruments), Assumption (ref)(iii) controls the growth rate of the sequence $k(N)$ as $N$, $K$, and $K_{2}$ grow.

To fix ideas, suppose that $d=1$ (one endogenous variable). $M_{1}$ is the number of subsets with $k$ instruments including at least 1 relevant instrument and $M_{2}$ is the number of subsets with $k$ irrelevant instruments. Since $M=M_{1}+M_{2}$ by definition and

equation*[equation* omitted — 81 chars of source]

for $k=1,...,K_{2}$ and $M_{2}=0$ for $k=K_{2}+1,...,K$, we have

equation*[equation* omitted — 218 chars of source]

Thus, $M_{1}/M$ is monotonically increasing in $k$ for given $K$ and $K_{1}$ with the minimum $K_{1}/K$ at $k=1$. Thus, Assumption (ref)(ii) holds if there is a positive constant $C$ such that $0< C^{-1} \leq K_{1}/K$.

We can also find a sufficient condition for Assumption (ref)(iii) when $d=1$. By using the bounds for binomial coefficients and Stirling's formula for a large $k$,

equation[equation omitted — 264 chars of source]

Thus, Assumption (ref)(iii) holds if $K_{2}/K<e^{-1}$. Since this is a sufficient condition, Assumption (ref)(iii) can hold even when the proportion of irrelevant instruments is much higher than $e^{-1}$. Figure (ref) shows that $(M_{2}/M)\sqrt{k}\rightarrow0$ as long as $k$ grows sufficiently fast as $(K,K_{2})$ increases when $K_{2}/K=0.8$.

figure[figure omitted — 194 chars of source]

Now we can generalize the main theorem allowing for an increasing number of irrelevant instruments.

theoremIf Assumptions (ref) and (ref) are satisfied, $k^{2}/N\rightarrow0$, and $\sigma_{u\varepsilon}\neq0$, then for the CSA-2SLS estimator the equations (ref)-(ref) are satisfied with \begin{equation} S(k) = \left(\frac{M}{M_{1}}\right)^{2}H^{-1}\left[\sigma_{u\ensuremath{\varepsilon}}\sigma_{u\ensuremath{\varepsilon}}'\frac{k^{2}}{N} + \sigma_{\ensuremath{\varepsilon}}^{2}\frac{f'(I-P^{k})(I-P_{f})(I-P^{k})f}{N}\right]H^{-1}. \end{equation}

Theorem (ref) shows that the approximate MSE formula in the presence of irrelevant instruments is the same as the formula in Theorem (ref) except that $(M/M_{1})^{2}$ is multiplied. Since $(M/M_{1})^{2}$ depends on $k$, the optimal $k$ that minimizes (ref) may differ from the minimizer of (ref).

Specifically, by observing that $(M/M_{1})^{2}$ places larger penalties for smaller $k$, we can conclude that the optimal $k$ will be (weakly) larger than those without the term $(M/M_{1})^{2}$. A practical recommendation is to try larger values of $k$ than the minimizer of the sample approximate MSE if the presence of irrelevant (or very weak) instruments is suspected.

Orthogonal Instruments

There are special cases when the CSA-2SLS estimator coincides exactly with the 2SLS estimator using all the instruments. The two estimators are the same if (a) $k=K$, which is trivial, or (b) the instruments are mutually orthogonal. We elaborate on the second case.

The projection matrix of the orthogonal instruments is equal to the sum of the projection matrices of each of the instruments. Thus, the projection matrix of the 2SLS estimator, $P$, becomes

align*[align* omitted — 86 chars of source]

where $\widetilde{P}^{1}_{j}$ for $j=1,...,K$ is the projection matrix based on the orthogonal instrument $j$ with the subset size 1. Now consider the CSA-$P$ matrix with subset size $k$ constructed from the orthogonal instruments:

align*[align* omitted — 183 chars of source]

Since the term $k/K$ is cancelled out when we plug-in $(k/K)P$ into Equation (ref), the CSA-2SLS estimator becomes identical to the 2SLS estimator for any $k$.

The identity result does not hold in general when instruments are correlated although they can be always orthogonalized without affecting the column space they span. We illustrate this point by a simple example with $K=2$ and $k=1$. Let $(Z_1,Z_2)$ be the vector of instruments and $(\widetilde{Z}_1,\widetilde{Z}_2)$ be the corresponding vector of orthogonalized instruments. Then,

align*[align* omitted — 218 chars of source]

where $\widetilde{P}$ denotes a generic projection matrix of the orthogonalized instruments. Therefore, the CSA-2SLS estimator using $P^1$ does not give the same estimate as that using $\frac{1}{2} P^2$, the 2SLS estimator.

The approximate MSE in (ref) is simplified further when instruments are orthogonal. Let $\Delta_{K} = \text{tr}\left(\frac{f'(I-P)f}{N}\right)$. Because of $P^k=(k/K)P$, $(I-P_{f})f=0$, and the idempotent property of $I-P$, we have

align[align omitted — 361 chars of source]

Thus, (ref) can be written as

equation[equation omitted — 110 chars of source]

where

equation*[equation* omitted — 206 chars of source]

Note that $\widetilde{S}(K)$ is the approximate MSE for 2SLS using $K$ instruments as shown in \citet*{donald2001choosing} and is independent of $k$. We show in the following Corollary that the approximate MSE in (ref) is asymptotically equivalent to $\widetilde{S}(K)$ when instruments are orthogonal.

corollarySuppose that Assumptions (ref)--(ref) are satisfied with orthogonal instruments. Then, for any non-zero $d\times 1$ vector $\lambda$, as $N \rightarrow \infty$, \begin{align} \frac{S_{\lambda}(k)}{\widetilde{S}_{\lambda}(K)} \ensuremath{\stackrel{p}{\rightarrow}} 1. \end{align}

Corollary (ref) holds because Assumption (ref)--(ref) imply $k/K\rightarrow1$ and $\Delta_{K}\xrightarrow{p}0$ in (ref) when instruments are orthogonal. The full proof is given in Appendix A.

Recall that CSA-2SLS does not depend on $k$ when instruments are orthogonal. Corollary (ref) establishes the logical consistency by showing that the approximate MSE becomes independent of $k$ as well when instruments are orthogonal.

Simulation

We investigate the finite sample properties of the CSA-2SLS estimator by conducting Monte Carlo simulation studies\footnote{The replication R codes for both the Monte Carlo experiments and empirical applications are available at \url{https://github.com/yshin12/ls-csa}.}. We consider the following simulation design:

align[align omitted — 124 chars of source]

where $Y_{i}$ is a scalar, $(\beta_0,\beta_1)$ is set to be (0, 0.1), $\beta_1$ is the parameter of interest, and $Z_{i}\sim \text{i.i.d. } N(0,\Sigma_{z})$. The diagonal terms of $\Sigma_{z}$ are ones and off-diagonal terms are $\rho_z$'s. Thus, $\rho_{z}$ denotes the correlation between instruments and is set to $\rho_z=0$ or $0.5$. The error terms $(\ensuremath{\varepsilon}_{i},u_{i})$ are i.i.d.\ over $i$, bivariate normal with variances 1 and covariance $\sigma_{u \ensuremath{\varepsilon}}$. Note that $\sigma_{u\ensuremath{\varepsilon}}$ denotes the severity of endogeneity and is set to $\sigma_{u\ensuremath{\varepsilon}}=0.1$ or $0.9$. We consider three designs for $\pi$: flat, decreasing, half-zero signals depending on the element values of the vector $\pi$. In each design, the signal-noise ratio is controlled by $R_f^2$, which is either 0.01 or 0.1. The exact formula of $\pi$ and $R_f^2$ are provided in the online supplement. The number of observations and the number of instruments are set to $(N,K)=(100,20)$ and $(1000,30)$. Note that the flat and decreasing designs of $\pi$ with $\rho_z=0$ coincide with those used in donald2001choosing and kuersteiner2010constructing. The simulation results are based on 400 replications.

In addition to the CSA-2SLS (denoted by CSA hereafter) estimator based on the approximate MSE, we also estimate the model by OLS, 2SLS with full instruments, the optimal 2SLS by donald2001choosing denoted by DN, and the model average estimator by kuersteiner2010constructing denoted by KO and compare their performance.\footnote{Since we build up the idea of CSA based on 2SLS, we focus on the comparison with similar 2SLS type estimators in various situations. We leave it for future research to develop a CSA estimator in different classes (e.g.\ LIML or JIVE) and compare the performance with other types of estimators.}

Tables (ref)--(ref) report the simulation results for designs with $\sigma_{u \ensuremath{\varepsilon}}=0.9$. The complete simulation results including $\sigma_{u \ensuremath{\varepsilon}}=0.1$ are collected in the Appendix. In the tables, we report the mean squared errors (MSE), bias (Bias), median absolute deviation (MAD), median bias (Median Bias), interdecile range (Range), and the coverage probability of the 95% confidence interval (Coverage). The mean and median of $\widehat{k}$ are also reported for DN and CSA.

As expected from the theory, the simulation results are quite promising for CSA when the instruments are correlated ($\rho_z=0.5$). Table (ref) reports the results of $(N,K)=(100,20)$. It confirms that CSA chooses smaller $k$'s and shows smaller variances. DN also shows a similar bias level but MSE is much larger. KO performs worse than CSA in terms of both bias and MSE in all designs in the table. However, MSE of KO is still comparable to CSA and is much smaller than that of DN. This indicates that the estimators based on model averaging mitigate the well-known `moments problem' caused by the large outliers with a small number of instruments, e.g. mariano1972existence.\footnote{We thank the referee for pointing this out.}

Next, we compare some robust statistics such as MAD and Median Bias. CSA shows better performance in both measures than DN except two designs (weak IV signal/decreasing $\pi_0$ and strong IV signal/flat $\pi_0$), where the losses are much smaller than the gains in other designs. Overall, KO shows smaller MAD but larger Median Bias than CSA. The coverage of CSA is close to the nominal level across different specifications except weak IV signal/half-zero $\pi_0$ but that of KO is far from the nominal level especially when the instrument signal is weak.\footnote{We construct the confidence intervals by assuming the instrument/weight/subset selection in each method is correct. See Section (ref) of the Appendix for details.}

Finally, the performance of CSA.CV, which chooses $\hat{k}$ by the 10-fold cross-validation, is slightly worse than CSA in these simulation designs. In Table (ref), we observe similar patterns when $(N,K)=(1000,30)$ and $R_f^2=0.01$. However, the performance of DN, KO, and CSA is quite similar in the lower panel when the signal from the instruments is strong.

Tables (ref)--(ref) report the simulation results when the instruments are uncorrelated. As expected from the theory, the performance of CSA is quite similar to 2SLS. We partially confirm the results in donald2001choosing with the same designs. The coverage of DN is better than other estimators that show quite similar performance but MAD of DN is much larger.\footnote{We could duplicate the MAD statistics in donald2001choosing only by the formula $med\left\vert \widehat{\beta}_{1} -\beta_{1,0} \right\vert$, i.e. re-centered at the true parameter value. However, we use the formula $med\left\vert \widehat{\beta}_{1} -med (\widehat{\beta}_{1}) \right\vert$ for MAD in this paper. See, for example, rousseeuw1993alternatives for the definition.}

table[table omitted — 4,021 chars of source]
table[table omitted — 3,667 chars of source]
table[table omitted — 3,681 chars of source]
table[table omitted — 3,694 chars of source]

We finish this section by a remark on the computation. When we compute the CSA-P matrix, we use 100 random draws from complete subsets when the complete subset size is bigger than 100. Table (ref) in the Appendix reports simulation results with different sizes of random draws in one design, and it shows that CSA performs satisfactorily with a relatively small number of draws.

Empirical Illustration

In this section we illustrate the usefulness of the CSA-2SLS estimator by estimating a logistic demand function for automobiles in \citet*{berry1995automobile}. The model specification is

align*[align* omitted — 169 chars of source]

where $S_{it}$ is the market share of the product $i$ in market $t$ with product 0 denoting the outside option, $P_{it}$ is the endogenous price variable, $X_{it}$ is a vector of included exogenous variables, and $Z_{it}$ is a vector of instruments. The parameter of interest is $\alpha_0$ from which we can calculate the price elasticity of demand.

table[table omitted — 1,653 chars of source]

We first estimate the model by using the same set of regressors and instruments used by \citet*{berry1995automobile}. The vector of covariates $X_{it}$ includes 5 variables: a constant, an air conditioning dummy, horsepower divided by weight, miles per dollar, and vehicle size. The vector of original instruments $Z_{it}$ includes 10 variables and is constructed by the characteristics of other car models. We also consider an extended design by adopting 48 instruments and 24 regressors constructed by \citet*{chernozhukov2015post}. We presume that all instruments are valid and relevant. Based on the previous simulation results, we adopt two methods for choosing $k$ for CSA: (a) the approximate MSE (AMSE) method and (b) the cross-validation(CV) method. We use the Mallows criterion for selecting IVs in AMSE and apply the 10-fold cross-validation for CV.\ For both estimators, we set $R=100$ when calculating the CSA-P matrix.

Table (ref) summarizes the estimation results. In addition to two CSA estimates, we also estimate the model using OLS, 2SLS, DN, and KO for comparison. For each design we report the choice of $k$ in DN and CSAs, the estimate of $\alpha$, and the heteroskedasticity and cluster robust standard errors of $\alpha$ given that we have chosen the correct model for $k$ or the optimal weight of KO. Finally, we report the number of products whose price elasticity of demand is inelastic, which is computed by the following formula:

align*[align* omitted — 99 chars of source]

where $1\{A\}$ is an indicator function equal to 1 if $A$ holds.

It is interesting to note that the estimation results of CSA contrast sharply with those of other estimation methods in the extended design. Recall that the economic theory predicts elastic demand in this market. Similar to the result in \citet*{berry1995automobile}, the OLS estimate is biased towards zero and makes 1,405 out of 2,217 products (63%) have inelastic own price demand. The 2SLS estimator mitigates the bias but not enough. Still, 874 (41%) products show inelastic own price elasticity. It is interesting to note that DN and KO are not particularly better than 2SLS in this empirical example although they are supposed to correct the bias caused by many instruments. The estimation result of DN comes close to 2SLS by choosing most instruments, 47 out of 48, and that of KO coincides exactly with 2SLS by putting the whole weight to the largest set of instruments. On the contrary, only 7 products (0.3%) have inelastic demand according to both estimation results of CSA. The $\alpha$ estimates by CSA is about twice as large as those by other estimators in absolute term. Since the bias caused by many instruments is towards the OLS estimates, this result can be viewed as a correction for the many instrument bias. However, the standard error of CSA is larger than others and there is a potential trade-off between the bias and the variance. Finally, the original design has fewer instruments and the bias correction by CSAs is not as large as in the extended design. However, CSA.CV shows a larger bias correction and its robustness to heteroskedasticity could be a potential reason.

In sum, 2SLS with all the available instruments suffers from many instruments bias in this application. DN and KO do not correct the bias enough to make the estimation results consistent with the prediction by the economic theory. In contrast, the CSA point estimates reduce the bias substantially in the extended design. Therefore, it is worthwhile to estimate a model with CSA and compare it with other existing methods when the model contains many instruments.