EconBase
← Back to paper

A Conditional Linear Combination Test with Many Weak Instruments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

103,974 characters · 8 sections · 171 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Conditional Linear Combination Test with Many Weak Instruments

abstractWe consider a linear combination of jackknife Anderson-Rubin (AR), jackknife Lagrangian multiplier (LM), and orthogonalized jackknife LM tests for inference in IV regressions with many weak instruments and heteroskedasticity. Following I.Andrews(2016), we choose the weights in the linear combination based on a decision-theoretic rule that is adaptive to the identification strength. Under both weak and strong identifications, the proposed test controls asymptotic size and is admissible among certain class of tests. Under strong identification, our linear combination test has optimal power against local alternatives among the class of invariant or unbiased tests which are constructed based on jackknife AR and LM tests. Simulations and an empirical application to Angrist-Krueger(1991)'s (Angrist-Krueger(1991)) dataset confirm the good power properties of our test. Keywords: Many instruments, power, size, weak identification JEL codes: C12, C36, C55

Introduction

Various recent surveys in leading economics journals suggest that weak instruments remain important concerns for empirical practice. For instance, I.Andrews-Stock-Sun(2019) survey 230 instrumental variable (IV) regressions from 17 papers published in the American Economic Review (AER). They find that many of the first-stage F-statistics (and non-homoskedastic generalizations) are in a range that raises such concerns, and virtually all of these papers report at least one first-stage F with a value smaller than 10. Similarly, in lee2021's (lee2021) survey of 123 AER articles involving IV regressions, 105 out of 847 specifications have first-stage Fs smaller than 10. Moreover, many IV applications involve a large number of instruments. For example, in their seminal paper, Angrist-Krueger(1991) study the effect of schooling on wages by interacting three base instruments (dummies for the quarter of birth) with state and year of birth, resulting in 180 instruments. Hansen-Hausman-Newey(2008) show that using the 180 instruments gives tighter confidence intervals than using the base instruments even after adjusting for the effect of many instruments. In addition, as pointed out by MS22, in empirical papers that employ the “judge design" (e.g., see maestas2013, sampat2019, and dobbie2018), the number of instruments (the number of judges) is typically proportional to the sample size, and the famous Fama-MacBeth two-pass regression in empirical asset pricing (e.g., see fama1973, shanken1992, and anatolyev2022) is equivalent to IV estimation with the number of instruments proportional to the number of assets. Similarly, belloni2012 consider an IV application involving more than one hundred instruments for the study of the effect of judicial eminent domain decisions on economic outcomes. carrasco2015 used many instruments in the estimation of the elasticity of intertemporal substitution in consumption. Furthermore, as pointed out by Goldsmith(2020), the shift-share or Bartik instrument (e.g., see bartik1991 and blanchard1992), which has been widely applied in many fields such as labor, public, development, macroeconomics, international trade, and finance, can be considered as a particular way of combining many instruments. For example, in the canonical setting of estimating the labor supply elasticity, the corresponding number of instruments is equal to the number of industries, which is also typically proportional to the sample size.

In this paper, following the seminal study by I.Andrews(2016), we propose a jackknife conditional linear combination (CLC) test that is robust to weak identification, many instruments, and heteroskedasticity. The proposed test also achieves efficiency under strong identification against local alternatives. The starting point of our analysis is the observation that, under strong identification, an orthogonalized jackknife Lagrangian multiplier (LM) test is the uniformly most powerful (UMP) test against local alternatives among the class of tests that are constructed based on jackknife LM and Anderson-Rubin (AR) tests and are either unbiased or invariant to sign changes. However, the orthogonalized LM test may not have good power under weak identification or against certain fixed alternatives. Therefore, we consider a linear combination of jackknife AR, jackknife LM, and orthogonalized LM tests. Specifically, we follow I.Andrews(2016) and determine the linear combination weights by minimizing the maximum power loss, which can be viewed as a maximum regret and is further calibrated based on the limit experiment of interest and a sufficient statistic for the identification strength under many instruments. Then, similar to I.Andrews(2016), we show such a jackknife CLC test is adaptive to the identification strength in the sense that (1) it achieves correct asymptotic size, (2) it is asymptotically and conditionally admissible under weak identification among some class of tests, (3) it converges to the UMP test mentioned above under strong identification against local alternatives,\footnote{We emphasize that the UMP property of our CLC test under strong identification holds within the class of sign-invariant or unbiased tests that are constructed based on jackknife AR and LM tests only. It may be possible to construct more efficient tests using test statistics besides the jackknife AR and LM. How to construct a globally optimal test under strong identification with many IVs and heteroskedastic errors is a topic that remains to be explored in future research. } and (4) it has asymptotic power equal to 1 under strong identification against fixed alternatives. The properties of jackknife AR, jackknife LM, orthogonalized LM, and our CLC tests are summarized in Table (ref). Simulations based on the limit experiment as well as calibrated data confirm the good power properties of our test. Then, we apply the new jackknife CLC test to Angrist-Krueger(1991)'s (Angrist-Krueger(1991)) dataset with the specifications of 180 and 1,530 instruments. We find that, in both specifications, our confidence intervals (CIs) are the shortest among those constructed by weak identification robust tests, namely, the jackknife AR, LM, and CLC tests, and the two-step procedure. Furthermore, our CIs are found to be even shorter than the non-robust Wald test CIs based on the jackknife IV estimator (JIVE) proposed by Angrist(1999), which is in line with the theoretical result that the jackknife CLC test is adaptive to the identification strength and is efficient under strong identification.

table[table omitted — 499 chars of source]

Relation to the literature. The contributions in the present paper relate to two strands of literature. First, it is related to the literature on many instruments; see, for example, Kunitomo1980, morimune1983, Bekker(1994), donald2001, chamberlain2004, Chao-Swanson(2005), stock2005, han2006, D.Andrews-Stock(2007), Hansen-Hausman-Newey(2008), Newey-Windmeijer(2009), anderson2010, kuersteiner2010, anatolyev2011, belloni2011, okui2011, belloni2012, carrasco2012, Chao(2012), Haus2012, hansen2014, carrasco2015, Wang_Kaffo_2016, kolesar2018, Matsushita2020, solvsten2020, crudu2021, and MS22, among others. In the context of many instruments and heteroskedasticity, Chao(2012) and Haus2012 provide standard errors for Wald-type inferences that are based on JIVE and jackknifed versions of the limited information maximum likelihood (LIML) and Fuller(1977)'s (Fuller(1977)) estimators (HLIM and HFUL). These estimators are more robust to many instruments than the commonly used two-stage least squares (TSLS) estimator because they can correct the bias caused by the high dimension of IVs.\footnote{Specifically, the rate of growth of the concentration parameter, which measure the overal instrument strength, is denoted as $\mu_n^2$. JIVE, HLIM, and HFUL remain consistent with heteroskedastic errors even when instrument weakness is such that $\mu_n^2$ is slower than the number of instruments $K$, provided that $\mu_n^2/\sqrt{K} \rightarrow \infty$ as the number of observations $n \rightarrow \infty$ (Chao(2012), Haus2012). In contrast, TSLS is less robust to instrument weakness as it is shown to be consistent only under homoskedasticity if $\mu_n^2/K \rightarrow \infty$ (Chao-Swanson(2005), Chao-Swanson(2005)).} In simulations derived from the data in Angrist-Krueger(1991), which is representative of empirical labor studies with many instrument concerns, Angrist-Frandsen2022 show that such bias-corrected estimators outperform the TSLS that is based on the instruments selected by the least absolute shrinkage and selection operator (LASSO) introduced in belloni2012 or the random forest-fitted first stage introduced in athey2019. Furthermore, under many weak moment asymptotics, Newey-Windmeijer(2009) provide new variance estimators for the jackknife GMM and the class of generalized empirical likelihood (GEL) estimators, which includes the continuous updating estimator (CUE) and EL estimator as special cases. In the linear heteroskedastic IV model, consistency and asymptotic normality of CUE require $m^2/n \rightarrow 0$ and $m^3/n \rightarrow 0$, respectively, where $m$ and $n$ denote the number of moment conditions and the sample size (e.g., see p.689 of Newey-Windmeijer(2009)). Such conditions are needed to simultaneously control the estimation error for all the elements of the heteroskedasticity consistent weighting matrix. Somewhat stronger rate conditions are required for other GEL estimators.

However, the Wald-type inference methods are invalid under weak identification, which occurs when the ratio of the concentration parameter over the square root of the number of instruments remains bounded as the sample size increases to infinity. In this case, all the estimators mentioned earlier become inconsistent, and there is no consistent test for the structural parameter of interest (see Section 3 of MS22). For weak identification robust inference under many instruments, D.Andrews-Stock(2007) consider the AR test, the score test introduced in Kleibergen(2002), and the conditional likelihood ratio test introduced in Moreira(2003). Their IV model is homoskedastic and requires the number of instruments to diverge slower than the cube root of the sample size ($K^3/n \rightarrow 0$, where $K$ denotes the number of instruments). anatolyev2011 propose a modified AR test that allows for the number of instruments to be proportional to the sample size but still require homoskedastic errors. Recently, crudu2021 and MS22 propose jackknifed versions of the AR test in a model with many instruments and heteroskedasticity. Both tests are robust to weak identification, but MS22's (MS22) jackknife AR test has better power properties due to the use of a cross-fit variance estimator. However, the jackknife AR tests may be inefficient under strong identification. To address this issue, MS22 also propose a new pre-test for weak identification under many instruments and apply it to form a two-stage testing procedure with a Wald test based on the JIVE introduced in Angrist(1999). The JIVE-Wald test is more efficient than the jackknife AR under strong identification. Therefore, an empirical researcher can employ the jackknife AR if the pre-test suggests weak identification and the JIVE-Wald if the pre-test suggests strong identification. In addition to the jackknife AR, Matsushita2020 propose a jackknife LM test, which is also robust to weak identification, many instruments, and heteroskedastic errors. However, the jackknife CLC test introduced in our paper is more efficient than the jackknife AR, the jackknife LM, and the two-step test under strong identification and local alternatives, while still being robust to weak identification.

Second, our paper is related to the literature on weak identification under the framework of a fixed number of instruments or moment conditions, in which various robust inference methods are available for non-homoskedastic errors; see, for example, Stock-Wright(2000), Kleibergen(2005), D.Andrews-Cheng(2012), I.Andrews(2016), I.Andrews-Mikusheva(2016), I.Andrews(2018), Moreira-Moreira(2019), D.Andrews-Guggenberger(2019), and lee2021. In particular, our jackknife CLC test extends the work of I.Andrews(2016) to the framework with many weak instruments. I.Andrews(2016) considers the convex combination between the generalized AR statistic (S statistic) introduced by Stock-Wright(2000) and the score statistic (K statistic) introduced by Kleibergen(2005). We find that under many weak instruments, the orthogonalized jackknife LM statistic plays a role similar to the K statistic. However, the trade-off between the jackknife AR and orthogonalized LM statistics turns out to be rather different from that between the S and K statistics. As pointed out by I.Andrews(2016), in the case with a fixed number of weak instruments (or moment conditions), the K statistic picks out a particular (random) direction corresponding to the span of a conditioning statistic that measures the identification strength and restricts attention to deviations from the null along this specific direction. In contrast to the K statistic, the S statistic treats all deviations from the null equally. Therefore, the trade-off between the K and S statistics is mainly from the difference in attention to deviation directions. We find that with many weak instruments, the jackknife AR and orthogonalized LM tests do not have such difference in deviation directions. Instead, their trade-off is mostly between local and non-local alternatives. Furthermore, although the standard LM test (without orthogonalization) is not weak identification robust under I.Andrews(2016)'s framework, the jackknife LM test is under many instruments. Therefore, we consider a linear combination of jackknife AR, jackknife LM, and orthogonalized jackknife LM tests and find that the resulting CLC test has good power properties in a variety of scenarios.

Notation. We denote $\mathcal{Z}(\mu)$ as the normal random variable with unit variance and expectation $\mu$ and $[n] = \{1,2,\cdots, n\}.$ We further simplify $\mathcal{Z}(0)$ as $\mathcal{Z}$, which is just a standard normal random variable. We denote $z_{\alpha}$ as the $(1-\alpha)$ quantile of a standard normal random variable and $\mathbb{C}_{\alpha}(a_1,a_2;\rho)$ as the $(1-\alpha)$ quantile of random variable $a_1 \mathcal{Z}_1^2 + a_2(\rho \mathcal{Z}_1 + (1-\rho^2)^{1/2} \mathcal{Z}_2)^2 + (1-a_1 - a_2)\mathcal{Z}_2^2$ where $\mathcal{Z}_1$ and $\mathcal{Z}_2$ are two independent standard normal random variables, $\alpha$ is the significance level, $\rho$ is a constant in $(-1,1)$, and $a_{1}$ and $a_{2}$ are the weights of the first and second components in the random variable. We further simplify $\mathcal{C}_{0,0;\rho}$ as $\mathbb{C}_{\alpha}$, which is just the $1-\alpha$ quantile of $\mathcal{Z}^2$. We let $\mathbb{C}_{\alpha,\max}(\rho) = \sup_{(a_1,a_2) \in \mathbb{A}_0} \mathbb{C}_{\alpha}(a_1,a_2;\rho)$, where $\mathbb{A}_0 = \{(a_1,a_2) \in [0,1] \times [0,1], a_1+a_2\leq \overline{a}\}$ for some $\overline{a}<1$. We suppress the dependence of $\mathbb{C}_{\alpha,\max}(\rho)$ on $\overline{a}$ for simplicity of notation. The operators $\mathbb{E}^*$ and $\mathbb{P}^*$ are expectation and probability taken conditionally on data, respectively. For example, $\mathbb{E}^*1\{\mathcal{Z}^2(\hat{\mu}) \geq \mathbb{C}_{\alpha}\}$, in which $\hat{\mu}$ is some estimator of the expectation $\mu$ based on data, means the expectation is taken over the normal random variable by treating $\hat{\mu}$ as deterministic. We use $\rightsquigarrow$ to denote convergence in distribution, $U \stackrel{d}{=} V$ to denote that $U$ and $V$ share the same distribution, and $\text{maxeig}(\mathcal{V})$ and $\text{mineig}(\mathcal{V})$ to denote maximum and minimum eigenvalues of a positive semidefinite matrix $\mathcal{V}$. For two sequences of random variables $U_n$ and $V_n$, we write $U_n \stackrel{d}{=}V_n +o_P(1)$ if there exist $\tilde U_n \stackrel{d}{=}U_n$ and $\tilde V_n \stackrel{d}{=}V_n$ such that $\tilde U_n - \tilde V_n = o_P(1)$.

Setup and Limit Problems

We consider the linear IV regression with a scalar outcome $Y_i$, a scalar endogenous variable $X_i$, and a $K \times 1$ vector of instruments $Z_i$ such that

align[align omitted — 101 chars of source]

where $\Pi_i = \mathbb{E}X_i$ and $\{Z_i\}_{i\in [n]}$ is treated as fixed, following the many-instrument literature. We let $K$ diverge with sample size $n$, allowing for the case that $K$ is of the same order of magnitude as $n$. We further have $\mathbb{E}V_i = 0$ by construction, and $\mathbb{E}e_i = 0$ by IV exogeneity. We allow $(e_i, V_i)$ to be heteroskedastic across $i$. Also, following the literature on many instruments (e.g., MS22), we assume that there are no controls included in our model as they can be partialled out from $(Y_i,X_i,Z_i)$. We provide more discussions about the effect of partialling out the covariates after Assumption (ref) below.

We are interested in testing $\beta = \beta_0$. Let $e_i(\beta_0) = Y_i - X_i \beta_0 = e_i + X_i \Delta$, where $\Delta = \beta- \beta_0$. We collect the transpose of $Z_i$ in each row of $Z$, an $n \times K$ matrix of instruments, and denote $P = Z (Z^\top Z)^{-1}Z^\top$. In addition, Let $Q_{a,b} = \frac{\sum_{i \in [n]}\sum_{j \neq i}a_i P_{ij}b_j}{\sqrt{K}}$ and $\mathcal{C} =Q_{\Pi,\Pi}$. Then, as pointed out by MS22, the rescaled $\mathcal{C}$ is the concentration parameter that measures the strength of identification in the heteroskedastic IV model with many instruments. Specifically, the parameter $\beta$ is weakly identified if $\mathcal{C}$ is bounded and strongly identified if $|\mathcal{C}|\rightarrow \infty$. We consider drifting sequence asymptotics so that all quantities are implicitly indexed by the sample size $n$ except specified otherwise. We omit such dependence for notation simplicity.

Throughout the paper, we consider three scenarios: (1) weak identification and fixed alternatives in which $\mathcal{C} \rightarrow \widetilde{\mathcal{C}}$ for some fixed constant $\widetilde{\mathcal{C}} \in \Re$ and $\Delta$ is fixed and bounded, (2) strong identification and local alternatives in which $\mathcal{C} = \widetilde{\mathcal{C}}/d_n $, $\Delta = \widetilde{\Delta}d_n $, $\widetilde{\mathcal{C}}$ and $\widetilde{\Delta}$ are bounded constants independent of $n$, and $d_n \rightarrow 0$ is a deterministic sequence, and (3) strong identification and fixed alternatives in which $\mathcal{C} = \widetilde{\mathcal{C}}/d_n $ for the same $\widetilde{\mathcal{C}}$ and $d_n$ defined in case (2) and $\Delta$ is fixed and bounded.\footnote{If we follow the setup in Chao(2012) and Haus2012 and assume $\Pi_i = \mu_n \pi_i/\sqrt{n}$ so that $\infty>C\geq \sum_{i \in [n]}\sum_{j \neq i}\pi_iP_{ij}\pi_j/n \geq c>0$ for some constants $c,C$, then $ \mathcal{C} = \frac{\mu_n^2}{\sqrt{K}}\frac{\sum_{i \in [n]}\sum_{j \neq i}\pi_iP_{ij}\pi_j}{n}$, implying that $d_n = \sqrt{K}/\mu_n^2$. Then, our definition of strong identification ($d_n \rightarrow 0$) is equivalent to that defined in Chao(2012) and Haus2012 ($\mu_n^2/\sqrt{K} \rightarrow \infty$).} Many weak identification robust tests proposed in the literature (namely, the jackknife AR tests proposed by crudu2021 and MS22 and the jackknife LM test proposed by Matsushita2020) depend on a subset of the following three quantities: $(Q_{e(\beta_0),e(\beta_0)},Q_{X,e(\beta_0)},Q_{X,X})$. Throughout the paper, we maintain the following high-level assumption.

assUnder both weak and strong identification, the following weak convergence holds: \begin{align} \begin{pmatrix} & Q_{e,e} \\ & Q_{X,e} \\ & Q_{X,X} - \mathcal{C} \end{pmatrix} \rightsquigarrow \mathcal{N}\left(\begin{pmatrix} 0 \\ 0 \\ 0 \end{pmatrix},\begin{pmatrix} \Phi_1 & \Phi_{12} & \Phi_{13} \\ \Phi_{12} & \Psi & \tau \\ \Phi_{13} & \tau & \Upsilon \end{pmatrix}\right), \end{align} for some $(\Phi_1, \Phi_{12}, \Phi_{13}, \Psi, \tau, \Upsilon)$.

Although there are no controls in the model ((ref)), we further verify Assumption (ref) in Section (ref) of the Online Supplement for a proper linear IV regression that includes a fixed dimension of exogenous control variables, which are then partialled out from the original outcome variable, endogenous variable, and instruments.\footnote{ Here, we focus on the case where the number of exogenous control variables is treated as fixed. In the case where the dimension of the exogenous variables is also large and assumed to diverge to infinity with the sample size, CNT23 propose new versions of various jackknife IV estimators and show they are consistent and asymptotically normal under strong identification. We conjecture that it is possible to replace our jackknife construct (i.e. $Q_{a,b}$) by the new version and consider weak identification robust tests and their linear combinations in the same manner as studied in this paper. This is left as a topic for future research.}

Assumption (ref) implies that,\footnote{Note that $

pmatrix[pmatrix omitted — 78 chars of source]

=

pmatrix[pmatrix omitted — 77 chars of source]
pmatrix[pmatrix omitted — 51 chars of source]

.$} under both strong and weak identification,

align[align omitted — 485 chars of source]

where

align[align omitted — 466 chars of source]

In particular, under strong identification, we have $Q_{X,X}d_n \stackrel{p}{\longrightarrow} \widetilde{\mathcal{C}}$, which has a degenerate distribution. Also, under local alternatives, we have $\Delta = o(1)$ so that

align*[align* omitted — 152 chars of source]

To describe a feasible version of the test, we assume we have consistent estimates for all the variance components.

assLet $\rho(\beta_0) = \frac{\Phi_{12}(\beta_0)}{\sqrt{\Phi_1(\beta_0)\Psi(\beta_0)}}$, $\widehat{\gamma}(\beta_0) = (\widehat{\Phi}_1(\beta_0), \widehat{\Phi}_{12}(\beta_0), \widehat{\Phi}_{13}(\beta_0), \widehat{\Psi}(\beta_0), \widehat{\tau}(\beta_0), \widehat{\Upsilon}, \widehat{\rho}(\beta_0))$ be an estimator, and $\mathcal{B} \in \Re$ be a compact parameter space. Then, we have $\inf_{\beta_0 \in \mathcal{B}}\Phi_1(\beta_0)>0$, $\inf_{\beta_0 \in \mathcal{B}}\Psi(\beta_0)>0$, $\Upsilon >0$, and for $\beta_0 \in \mathcal{B}$, \begin{align*} ||\widehat{\gamma}(\beta_0) - \gamma(\beta_0)||_2 = o_p(1), \end{align*} where $\gamma(\beta_0) \equiv (\Phi_1(\beta_0), \Phi_{12}(\beta_0),\Phi_{13}(\beta_0),\Psi(\beta_0), \tau(\beta_0),\Upsilon,\rho(\beta_0))$.

Several remarks on Assumption (ref) are in order. First, Chao(2012) propose a consistent estimator for $\Psi$ where there is strong identification and many instruments. It is possible to compute $\widehat{\gamma}(\beta_0)$ based on Chao(2012)'s (Chao(2012)) estimator with their JIVE-based residuals $\hat{e}_i$ from the structural equation replaced by $e_i(\beta_0)$. Under weak identification and $\beta_0 = \beta$, crudu2021 and Matsushita-Otsu2021 establish the consistency of such estimators for $\Phi_1(\beta_0)$ and $\Psi(\beta_0)$, respectively. Similar arguments can be used to show the consistency of the rest of the elements in $\widehat{\gamma}(\beta_0)$ under both weak and strong identification. In addition, the consistency can be established under both local and fixed alternatives. We provide more details in Section (ref) in the Online Supplement. Second, motivated by KSS2020, MS22 propose cross-fit estimators $\widehat{\Phi}_1(\beta_0)$ and $\widehat{\Upsilon}$, which are consistent under both weak and strong identification and lead to better power properties. Following their lead, one can write down the cross-fit estimators for the rest of the elements in $\gamma(\beta_0)$ and show they are consistent.\footnote{For example, MS22 establish the limit of their cross-fit estimator $\widehat{\Psi}$ under weak identification and many instruments when the residual $\hat{e}_i$ from the structural equation is computed based on the JIVE estimator. We can construct $\widehat{\Psi}(\beta_0)$ by replacing $\hat{e}_i$ by $e_i(\beta_0)$. Then, the argument, as theirs with $Q_{X,e}/Q_{X,X}$ replaced by $\Delta$, establishes that $\widehat{\Psi}(\beta_0) \stackrel{p}{\longrightarrow} \Psi(\beta_0)$.} We provide more details in Section (ref) in the Online Supplement. Note that both crudu2021's (crudu2021) and MS22's (MS22) estimators are consistent under heteroskedasticity and allow for $K$ to be of the same order of $n$. Third, the consistency of $\widehat \gamma(\beta_0)$ over the entire parameter space under both strong and weak identifications is more than necessary and maintained mainly for simplicity of presentation. In fact, in order for our jackknife CLC test proposed below to control size under both weak and strong identification, it suffices to require $\widehat{\gamma}(\beta_0)$ to be consistent under the null only. The power analyses in Lemmas (ref) and (ref) below, and subsequently, Theorems (ref) and (ref), only require the consistency of $\widehat{\gamma}(\beta_0)$ under strong identification with local alternatives and weak identification with fixed alternatives, respectively.

Under this framework, crudu2021 and MS22 consider the jackknife AR test

align[align omitted — 150 chars of source]

and Matsushita2020 consider the jackknife LM test

align[align omitted — 150 chars of source]

Both tests are robust to weak identification, many instruments, and heteroskedasticity. Lemma (ref) below characterizes the joint limit distribution of $(AR(\beta_0),LM(\beta_0))^\top$ under strong identification and local alternatives.

lemSuppose Assumptions (ref) and (ref) hold and we are under strong identification with local alternatives, that is, there exists a deterministic sequence $d_n \rightarrow 0$ such that $\mathcal{C} = \widetilde{\mathcal{C}}/d_n $ and $\Delta = \widetilde{\Delta}d_n $, where $\widetilde{\mathcal{C}}$ and $\widetilde{\Delta}$ are bounded constants independent of $n$. Then, we have \begin{align*} \begin{pmatrix} AR(\beta_0) \\ LM(\beta_0) \end{pmatrix} \rightsquigarrow \begin{pmatrix} \mathcal{N}_1 \\ \mathcal{N}_2 \end{pmatrix} \stackrel{d}{=} \mathcal{N}\left(\begin{pmatrix} 0 \\ \frac{\widetilde{\Delta} \widetilde{\mathcal{C}}}{\Psi^{1/2}} \end{pmatrix},\begin{pmatrix} 1 & \rho \\ \rho & 1 \end{pmatrix}\right) \end{align*} where $\rho = \Phi_{12}/\sqrt{\Phi_1\Psi}$.

Two remarks are in order. First, under strong identification, we consider local alternatives so that $ \beta-\beta_0 \rightarrow 0$. This is why we have $(\Psi(\beta_0),\Phi_1(\beta_0),\Phi_{12}(\beta_0))$ converge to $(\Psi,\Phi_1,\Phi_{12})$, which are just the counterparts of $(\Psi(\beta_0),\Phi_1(\beta_0),\Phi_{12}(\beta_0))$ when $\beta_0$ is replaced by $\beta$. Second, although $AR(\beta_0)$ has zero mean, and hence, no power in this case, it is correlated with $LM(\beta_0)$. It is therefore possible to use $AR(\beta_0)$ to reduce the variance of $LM(\beta_0)$ and obtain a test that is more powerful than the LM test.

lemConsider the limit experiment in which researchers observe $(\mathcal{N}_1,\mathcal{N}_2)$ with \begin{align*} \begin{pmatrix} \mathcal{N}_1 \\ \mathcal{N}_2 \end{pmatrix} \stackrel{d}{=} \mathcal{N}\left(\begin{pmatrix} 0 \\ \theta \end{pmatrix},\begin{pmatrix} 1 & \rho \\ \rho & 1 \end{pmatrix}\right), \end{align*} know the value of $\rho$ and that $\mathbb{E}\mathcal{N}_1 = 0$, and want to test for $\theta=0$ versus the two-sided alternative. In this case, $1\{\mathcal{N}_2^{*2} \geq \mathbb{C}_{\alpha}\}$ is UMP among level-$\alpha$ tests that are either invariant to sign changes or unbiased, where \begin{align*} \mathcal{N}_2^* = (1-\rho^2)^{-1/2}(\mathcal{N}_2 - \rho \mathcal{N}_1) \end{align*} is the normalized residual from the projection of $\mathcal{N}_2$ on $\mathcal{N}_1$.

Let the orthogonalized jackknife LM statistic be $LM^*(\beta_0) = (1-\widehat{\rho}(\beta_0)^2)^{-1/2} (LM(\beta_0)- \widehat{\rho}(\beta_0) AR(\beta_0))$. Then, Lemma (ref) implies, under strong identification and local alternatives,

align[align omitted — 390 chars of source]

Lemma (ref) with $\theta = \widetilde{\Delta} \widetilde{\mathcal{C}}\Psi^{-1/2}$ implies, in this case, that the test $1\{LM^{*2}(\beta_0) \geq \mathbb{C}_{\alpha}\}$ is asymptotically strictly more powerful than the jackknife AR and LM tests based on $AR(\beta_0)$ and $LM(\beta_0)$ against local alternatives as long as $\rho \neq 0$. In addition, under strong identification and local alternatives, MS22's (MS22) two-step test statistic is asymptotically equivalent to $LM(\beta_0)$, and thus, is less powerful than $LM^*(\beta_0)$ too.

Next, we compare the behaviors of $AR(\beta_0)$, $LM(\beta_0)$, and $LM^*(\beta_0)$ under strong identification and fixed alternatives.

lemSuppose Assumption (ref) holds, $( Q_{e(\beta_0),e(\beta_0)} - \Delta^2 \mathcal{C}, Q_{X,e(\beta_0)} - \Delta \mathcal{C},Q_{X,X} - \mathcal{C})^\top = O_p(1)$, and we are under strong identification so that $d_n \mathcal{C} \rightarrow \widetilde{\mathcal{C}}$ for some $d_n \rightarrow 0$. Then, we have, for any fixed $\Delta \neq 0$, \begin{align*} d_n^2 \begin{pmatrix} AR^2(\beta_0) \\ LM^2(\beta_0) \\ LM^{*2}(\beta_0) \end{pmatrix} \stackrel{p}{\longrightarrow} \begin{pmatrix} \Phi_1^{-1}(\beta_0) \Delta^4 \widetilde{\mathcal{C}}^2 \\ \Psi^{-1}(\beta_0) \Delta^2 \widetilde{\mathcal{C}}^2\\ (1-\rho^2(\beta_0))^{-1}(\Psi^{-1/2}(\beta_0) - \rho(\beta_0) \Phi_1^{-1/2}(\beta_0)\Delta)^2 \Delta^2 \widetilde{\mathcal{C}}^2 \end{pmatrix}. \end{align*}

Given $d_n \rightarrow 0$ and both $\Phi_1^{-1}(\beta_0) \Delta^4 \widetilde{\mathcal{C}}^2>0$ and $\Phi_1^{-1}(\beta_0) \Delta^2 \widetilde{\mathcal{C}}^2>0$, $AR^2(\beta_0)$ and $LM^2(\beta_0)$ have power $1$ against fixed alternatives asymptotically. By contrast, $LM^{*2}(\beta_0)$ may not have power if $\Delta = \Delta_*(\beta_0) \equiv \Phi_1^{1/2}(\beta_0)\Psi^{-1/2}(\beta_0)\rho^{-1}(\beta_0)$.

Next, we compare the performance of $AR(\beta_0)$ and $LM^*(\beta_0)$ under weak identification and fixed alternatives.

lemSuppose Assumptions (ref) and (ref) hold and we are under weak identification so that $\mathcal{C} \rightarrow \widetilde{\mathcal{C}} \in \Re$. Then, we have, for any fixed $\Delta \neq 0$, \begin{align} \begin{pmatrix} AR(\beta_0) \\ LM^*(\beta_0) \end{pmatrix} \rightsquigarrow \begin{pmatrix} \mathcal{N}_1 \\ \mathcal{N}_2^* \end{pmatrix} \stackrel{d}{=} \mathcal{N}\left(\begin{pmatrix} m_1(\Delta) \\ m_2(\Delta) \end{pmatrix},\begin{pmatrix} 1 & 0 \\ 0 & 1 \end{pmatrix}\right), \end{align} where $\rho(\beta_0) = \frac{\Phi_{12}(\beta_0)}{\sqrt{\Psi(\beta_0)\Phi_1(\beta_0) }}$ and \begin{align*} \begin{pmatrix} m_1(\Delta) \\ m_2(\Delta) \end{pmatrix} & = \begin{pmatrix} \Phi_1^{-1/2}(\beta_0)\Delta^2\widetilde{\mathcal{C}}\\ (1-\rho^2(\beta_0))^{-1/2}\Psi^{-1/2}(\beta_0)\Delta \widetilde{\mathcal{C}}-\rho(\beta_0) (1-\rho^2(\beta_0))^{-1/2}\Phi_1^{-1/2}(\beta_0)\Delta^2\widetilde{\mathcal{C}} \end{pmatrix}. \end{align*} In particular, as $\Delta \rightarrow \infty$, we have \begin{align*} m_1(\Delta) & \rightarrow \frac{\widetilde{\mathcal{C}}}{\Upsilon^{1/2}} \quad and \quad m_2(\Delta) \rightarrow \frac{\widetilde{\mathcal{C}}}{\Upsilon^{1/2}} \frac{\rho_{23}}{(1-\rho_{23}^2)^{1/2}}, \end{align*} where $\rho_{23} = \frac{\tau}{(\Psi \Upsilon)^{1/2}}$ is the correlation between $Q_{X,e}$ and $Q_{X,X}$.\footnote{We suppress the dependence of $m_1(\Delta)$ and $m_2(\Delta)$ on $\gamma(\beta_0)$ and $\widetilde{\mathcal{C}}$ for notation simplicity. }

By comparing the means of the normal limit distribution in ((ref)), we notice that under weak identification and fixed alternatives, neither $LM^*(\beta_0)$ dominates $AR(\beta_0)$ or vice versa. We also notice from Lemma (ref) that for testing distant alternatives, the power of $LM^*(\beta_0)$ is different from $AR(\beta_0)$ by a factor of $\rho_{23}/\sqrt{1-\rho^2_{23}}$, so that it will be lower when $|\rho_{23}| \leq 1/\sqrt{2}$. Under weak identification and homoskedasticity,\footnote{Specifically, we say the data are homoskedastic if the covariance matrices of $(e_i,V_i)$ are constant across $i$.} we have $\rho_{23} = \rho = \Phi_{12}/\sqrt{\Psi \Phi_1}$. Therefore, although the test $1\{LM^{*2}(\beta_0) \geq \mathbb{C}_{\alpha}\}$ has a power advantage under strong identification against local alternatives, it may lack power under weak identification against distant alternatives if the degree of endogeneity is low. Furthermore, $LM^*(\beta_0)$ may not have power if $\Delta = \Delta_*(\beta_0)$.

In the current setting with many instruments, $AR(\beta_0)$ and $LM^*(\beta_0)$ play roles similar to that of Stock-Wright(2000)'s (Stock-Wright(2000)) S and Kleibergen(2005)'s (Kleibergen(2005)) K statistics in I.Andrews(2016)'s (Andrews(2016)) setting, respectively. In the fixed number of IVs case, the power trade-off between S and K statistics is based on the direction of deviations from the null. However, as shown in Lemma (ref) (the case with weak identification and fixed alternatives), the deviations of $AR(\beta_0)$ and $LM^*(\beta_0)$ from the null do not have such a difference in direction under the many-instrument setting because $\widetilde {\mathcal{C}}$ is just a scalar. Instead, their power trade-off is between local and non-local alternatives. This is in stark contrast to the setting in I.Andrews(2016).

To achieve the advantages of $AR(\beta_0)$, $LM(\beta_0)$, and $LM^*(\beta_0)$ in all three scenarios above, we need to combine them in a way that is adaptive to the identification strength. Following I.Andrews(2016), we consider the linear combination of $AR^2(\beta_0)$, $LM^2(\beta_0)$, and $LM^{*2}(\beta_0)$. Recall that $(\mathcal{N}_1,\mathcal{N}_2^*)$ are the limits of $(AR(\beta_0),LM^{*}(\beta_0))$ in either strong or weak identification. See (ref) and (ref) for their expressions in these two cases. Then, in the limit experiment, the linear combination test can be written as

align[align omitted — 238 chars of source]

where $(a_1,a_2) \in \mathbb{A}_0$ are the combination weights, $\mathcal{N}_1 \sim \mathcal{Z}(\theta_1)$, and $\mathcal{N}_2^{*} \sim \mathcal{Z}(\theta_2)$; the mean parameters $\theta_1$ and $\theta_2$ are defined in Lemmas (ref) and (ref) for strong and weak identification, respectively; and $\tilde \rho$ is the limit of $\widehat{\rho}(\beta_0)$.\footnote{Under fixed alternatives, $\tilde \rho = \rho(\beta_0)$; under local alternatives, $\tilde \rho = \rho$.} Let the eigenvalue decomposition of the matrix $

pmatrix[pmatrix omitted — 151 chars of source]

$ be

align[align omitted — 310 chars of source]

where, by construction, $\nu_1(a_1,a_2)\geq \nu_2(a_1,a_2) \geq 0$ and $\mathcal{U}$ is a $2 \times 2$ unitary matrix. We highlight the dependence of eigenvalues $(\nu_1,\nu_2)$ on the weights $(a_1,a_2)$. The dependence of $\mathcal{U}$ on $(a_1,a_2)$ is suppressed for notation simplicity. Then, we have

align*[align* omitted — 234 chars of source]

and $\phi_{a_1,a_2,\infty} = 1\{\nu_1(a_1,a_2) \widetilde{\mathcal{N}}_1^2 + \nu_2(a_1,a_2) \widetilde{\mathcal{N}}_2^2 \geq \mathbb{C}_{\alpha}(a_1,a_2; \tilde \rho)\}$, where

align[align omitted — 211 chars of source]

and $\widetilde{\mathcal{N}}_1$ and $\widetilde{\mathcal{N}}_2$ are independent normal random variables with unit variance. This implies that $\phi_{a_1,a_2,\infty}$ can be viewed as a linear combination test of two independent chi-squared random variables with one degree of freedom, and those two chi-squared random variables are obtained by properly rotating $\mathcal{N}_1$ and $\mathcal{N}_2^*$ (i.e., the limits of $AR(\beta_0)$ and $LM^*(\beta_0)$).

Theorem (ref) states the key properties of $\phi_{a_1, a_2, \infty}$ under the limit experiment.

thm\begin{enumerate}[label=(\roman*)] • Suppose we are under weak identification and fixed alternatives and let $\mathcal{N}_1 \sim \mathcal{Z}(\theta_1)$, $\mathcal{N}_2^{*} \sim \mathcal{Z}(\theta_2)$, and they are independent, where $\theta_1 = m_1(\Delta)$ and $\theta_2 = m_2(\Delta)$ as in (ref). We consider the test of $H_0: \theta_1 = \theta_2 = 0$ against $H_1: \theta_1 \neq 0$ or $\theta_2 \neq 0$. Let $\Phi_{\alpha}$ denote the class of size-$\alpha$ tests for $H_0: \theta_1 = \theta_2 = 0$ constructed based on $(\widetilde{\mathcal{N}}_1^2,\widetilde{\mathcal{N}}_2^2)$ defined in (ref). Then, for any $(a_1,a_2) \in \mathbb{A}_0$, $\phi_{a_1,a_2,\infty}$ defined in (ref) is an admissible test within $\Phi_{\alpha}$. In addition, let $(\widetilde \theta_1, \widetilde \theta_2) = ( \theta_1, \theta_2) \mathcal{U}$. If $(\widetilde \theta_1^2, \widetilde \theta_2^2) = b \cdot (\nu_1(a_1,a_2),\nu_2(a_1,a_2))$ for some positive constant $b$, then for any test $\phi \in \Phi_\alpha$, there exists some $\overline{b}>0$ such that for any $0 < b < \overline{b}$, we have $\mathbb{E}\phi \leq \mathbb{E}\phi_{a_1,a_2,\infty}$. • Suppose we are under strong identification and local alternatives and \begin{align*} \begin{pmatrix} \mathcal{N}_1 \\ \mathcal{N}_2 \end{pmatrix} \stackrel{d}{=}\mathcal{N} \left( \begin{pmatrix} 0 \\ \theta \end{pmatrix}, \begin{pmatrix} 1 & \rho \\ \rho & 1 \end{pmatrix} \right), \end{align*} where $\theta =\frac{\widetilde{\Delta}\widetilde{\mathcal{C}}}{\Psi^{1/2}} $. We consider the test of $H_0: \theta = 0$ against $H_1: \theta \neq 0$. Then, $\phi_{a_1,a_2,\infty}$ defined in (ref) is UMP among the class of level-$\alpha$ tests that are constructed based on $(\mathcal{N}_1,\mathcal{N}_2)$ and invariant to the sign change if and only if $a_1 = 0$ and $a_2\rho = 0$. In this case, this test is also UMP among the class of unbiased level-$\alpha$ tests that are constructed based on $(\mathcal{N}_1,\mathcal{N}_2)$. • Suppose Assumption (ref) holds, $( Q_{e(\beta_0),e(\beta_0)} - \Delta^2 \mathcal{C}, Q_{X,e(\beta_0)} - \Delta \mathcal{C},Q_{X,X} - \mathcal{C})^\top = O_p(1)$, and we are under strong identification with fixed alternatives. If $1 \geq a_{1,n} \geq \frac{\tilde{q}\Phi_1(\beta_0)}{\mathcal{C}^2\Delta_*^4(\beta_0)}$ for some constant $\tilde{q} > \mathbb{C}_{\alpha,max}(\rho(\beta_0))$ and $(a_{1,n},a_{2,n}) \in \mathbb{A}_0$, where $\Delta_*(\beta_0) = \Phi_1^{1/2}(\beta_0)\Psi^{-1/2}(\beta_0)\rho^{-1}(\beta_0)$, then \begin{align*} 1\{a_{1,n} AR^2(\beta_0) + a_{2,n}LM^2(\beta_0)+ (1-a_{1,n} - a_{2,n})LM^{*2}(\beta_0) \geq \mathbb{C}_{\alpha}(a_{1,n},a_{2,n};\widehat{\rho}(\beta_0))\} \stackrel{p}{\longrightarrow} 1. \end{align*} \end{enumerate}

Several remarks are in order. First, unlike the one-sided jackknife AR test proposed by MS22, we construct the jackknife CLC test based on $AR^2(\beta_0)$ for several reasons. First, under weak identification, when the concentration parameter $\mathcal{C}$, and thus, $m_1(\Delta)$ defined in Lemma (ref) is nonnegative, the one-sided test has good power. However, even in this case, the power curves simulation in Section (ref) shows that our jackknife CLC test is more powerful than the one-sided AR test in most scenarios. Second, our jackknife CLC test will have good power even when $\mathcal{C}$ is negative.\footnote{We note that $\mathcal{C} = \frac{\sum_{i \in [n]} \sum_{ j \neq i} \Pi_i P_{ij}\Pi_j}{\sqrt{K}} = \frac{ \sum_{i \in [n]}(1-P_{ii})\Pi_i^2 - \Pi^\top M \Pi}{\sqrt{K}}$, where $M = I-P$. If $\Pi^\top M \Pi$ and $\sum_{i \in [n]}P_{ii}\Pi_i^2$ are sufficiently large, $\mathcal{C}$ can be negative. MS22 further assume that $\Pi^\top M \Pi \leq \frac{C \Pi^\top \Pi}{K}$ for some constant $C>0$, which implies that $\mathcal{C}>0$.} Third, we show below that under strong identification and local alternatives, our jackknife CLC test converges to the UMP test $1\{\mathcal{N}^{*2}_2>\mathbb{C}_{\alpha}\}$ whereas both the one- and two-sided tests based on $AR(\beta_0)$ have no power, as shown in Lemma (ref). Fourth, under strong identification and fixed alternatives, our jackknife CLC test has asymptotic power equal to 1, as shown in Lemma (ref) and Theorem (ref) below. In this case, using the one-sided jackknife AR test cannot further improve the power. Fifth, combining $LM^{*2}(\beta_0)$ with $AR^2(\beta_0)$ (and $LM^2(\beta_0)$), rather than $AR(\beta_0)$, can substantially mitigate the impact of power loss of $LM^*(\beta_0)$ at $\Delta_*(\beta_0)$, as shown in the numerical investigation in Section (ref).

Second, Theorem (ref)(i) implies that $\phi_{a_1,a_2,\infty}$ is admissible among tests that are also quadratic functions of $\mathcal{N}_1$ and $\mathcal{N}_2^*$ with the same rotation $\mathcal{U}$ but different eigenvalues $(\tilde \nu_1,\tilde \nu_2)$; that is,

align*[align* omitted — 216 chars of source]

Specifically, in the special case with $a_2 = 0$ (i.e., we put zero weight on $LM^2(\beta_0)$), the rotation matrix $\mathcal{U} = I_2$ and $\phi_{a_1,0,\infty}$ is admissible among level-$\alpha$ tests based on the test statistics of the form $a_1\mathcal{N}_1^2 + (1-a_1)\mathcal{N}_2^{*2}$ for $a_1 \in [0,1]$, which is similar to the result for the linear combination of S and K statistics in I.Andrews(2016).

Third, similar to I.Andrews(2016), Theorem (ref)(i) also shows that our linear combination test is optimal against certain alternatives under weak identification. Additionally, in the case with $a_2 = 0$, the power optimality result in (ref)(i) also carries over to $\phi_{a_1, 0, \infty}$ among level-$\alpha$ tests of the form $a_1 \mathcal{N}_1^2 + (1-a_1)\mathcal{N}^{*2}_2$ for $a_1 \in [0,1]$.

Fourth, when $a_1 = 0$ and $a_2\rho = 0$ and under strong identification and local alternatives, we have $\phi_{a_1,a_2,\infty} = 1\{\mathcal{N}_2^{*2} \geq \mathbb{C}_{\alpha}\}$, which is both the UMP invariant and unbiased test. When $\rho=0$ and under local alternatives, $a_2\mathcal{N}_2^{*2}$ in the second and third terms of $\phi_{a_1,a_2,\infty}$ cancels out, implying that $\phi_{a_1,a_2,\infty} = 1\{\mathcal{N}_2^{*2} \geq \mathbb{C}_{\alpha}\}$ as long as $a_1 = 0$.

Fifth, we note that both the rotation matrix $\mathcal{U}$ and the eigenvalues $\nu_1$ and $\nu_2$ in (ref) are functions of $(a_1,a_2)$. We choose this specific parametrization so that $\phi_{a_1,a_2,\infty}$ can be written as a linear combination of $AR^2(\beta_0)$, $LM^2(\beta_0)$, and $LM^{*2}(\beta_0)$. It is possible to use other parametrizations to combine $AR(\beta_0)$ and $LM^*(\beta_0)$. For example, let $$\mathcal{O}(\zeta) =

pmatrix[pmatrix omitted — 72 chars of source]

$$ be a rotation matrix with angle $\zeta$ and $

pmatrix[pmatrix omitted — 71 chars of source]

= \mathcal{O}(\zeta)

pmatrix[pmatrix omitted — 45 chars of source]

$. Then, in the limit experiment, the linear combination test statistic can be written as

align[align omitted — 95 chars of source]

where $(\mathcal{N}^{\dagger}_1,\mathcal{N}^{\dagger}_2)$ are the limits of $(AR^\dagger(\beta_0,\zeta) ,LM^\dagger(\beta_0,\zeta) )$ under either weak or strong identification. In the following, we will use a minimax procedure to determine the optimal weights $(a_1,a_2)$ for our jackknife CLC test $\phi_{a_1,a_2,\infty}$. Similarly, we can use this procedure to select the value of $a$ and $\zeta$ for the new parametrization in ((ref)). Under strong identification and local alternatives, Lemma (ref) shows that the test $1\{LM^{*2}(\beta_0) \geq \mathbb{C}_\alpha\}$ is the most powerful test against local alternatives. This is achieved by our jackknife CLC test $\phi_{a_1,a_2,\infty}$ with $a_1=0$ and $a_2\rho =0$. In this case, the new parametrization does not bring any additional power.

A Conditional Linear Combination Test

In this section, we determine the weights $(a_1,a_2)$ in the jackknife CLC test via a minimax procedure. Under weak identification, the limit test statistic of the jackknife CLC test with weights $(a_1,a_2)$ is

align[align omitted — 331 chars of source]

where $m_1(\Delta)$ and $m_2(\Delta)$ are defined in Lemma (ref), and $\mathcal{Z}_1(\cdot)$ and $\mathcal{Z}_2(\cdot)$ are independent. In this case, we can be explicit and write $\phi_{a_1,a_2,\infty} = \phi_{a_1,a_2,\infty}(\Delta)$. However, the limit power of the jackknife CLC test will typically remain unknown as the true parameter $\beta$ (and hence $\Delta$) is unknown. To overcome this issue, we follow I.Andrews(2016) and calibrate the power, i.e, $\mathbb{E}\phi_{a_1,a_2,\infty}(\delta)$, where $\delta$ ranges over all possible values that $\Delta$ can potentially take; we define $\phi_{a_1,a_2,\infty}(\delta)$ as well as the range of potential values of $\Delta$ below.

Let $\widehat{D} = Q_{X,X} - (Q_{e(\beta_0),e(\beta_0)},Q_{X,e(\beta_0)})

pmatrix[pmatrix omitted — 131 chars of source]

^{-1}

pmatrix[pmatrix omitted — 71 chars of source]

$ be the residual from the projection of $Q_{X,X}$ on $(Q_{e(\beta_0),e(\beta_0)},Q_{X,e(\beta_0)})$. By (ref), under weak identification,

align*[align* omitted — 95 chars of source]

where

align*[align* omitted — 563 chars of source]

We note that $\widehat{D}$ is a sufficient statistic for $\mu_D$, which contains information about the concentration parameter $\mathcal{C}$ and is asymptotically independent of $AR(\beta_0)$, $LM(\beta_0)$, and hence $LM^*(\beta_0)$.

Under weak identification, we observe that $m_1(\Delta)$ and $m_2(\Delta)$ in Lemma (ref) can be written as

align[align omitted — 168 chars of source]

where

align[align omitted — 537 chars of source]

By (ref), we see that $ \phi_{a_1,a_2,\infty}=\phi_{a_1,a_2,\infty}(\Delta)$ defined in (ref) can be written as

align*[align* omitted — 303 chars of source]

This motivates the definition that

align[align omitted — 352 chars of source]

To emphasize the dependence of $\phi_{a_1,a_2,\infty}(\delta)$ on $\mu_D$ and $\gamma(\beta_0)$, we further write $\phi_{a_1,a_2,\infty}(\delta)$ as $\phi_{a_1,a_2,\infty}(\delta,\mu_D,\gamma(\beta_0))$.

The range of values that $\Delta$ can take is defined as $\mathcal{D}(\beta_0) = \{\delta: \delta+\beta_0 \in \mathcal{B}\}$, where $\mathcal{B}$ is the parameter space. For instance, in their empirical application of returns to education, MS22 assume that $\beta$ (i.e., the return to education) ranges from -0.5 to 0.5, with $\mathcal{B} = [-0.5,0.5]$. We adopt the same practice in our simulations based on calibrated data in Section (ref) and empirical application in Section (ref). Specifying the parameter space is almost inevitable for any weak-identification-robust inference method, but additional simulation results in Section (ref) of the Online Supplement show that our method is insensitive to the choice of parameter space.

Following the lead of I.Andrews(2016), we define the highest attainable power for each $\delta \in \mathcal{D}(\beta_0)$ as $\mathcal{P}_{\delta,\mu_D} = \sup_{(a_1,a_2) \in \mathbb{A}(\mu_D,\gamma(\beta_0))} \mathbb{E}\phi_{a_1,a_2,\infty}(\delta,\mu_D,\gamma(\beta_0))$, which means that

align*[align* omitted — 105 chars of source]

is the power loss when the weights are set as $(a_1,a_2)$. Here we denote the domain of $(a_1,a_2)$ as $\mathbb{A}(\mu_D,\gamma(\beta_0))$ and define it as $\mathbb{A}(\mu_D,\gamma(\beta_0)) = \{(a_1,a_2) \in \mathbb{A}_0, a_1 \in [\underline{a}(\mu_D,\gamma(\beta_0)),1] \}$ where $\mathbb{A}_0 = \{(a_1,a_2) \in [0,1] \times [0,1], a_1+a_2\leq \overline{a}\}$ for some $\overline{a}<1$,

align[align omitted — 221 chars of source]

the two tuning parameters $(p_1,p_2) = (0.01,1.1)$, $\Delta_*(\beta_0) = \Phi_1^{1/2}(\beta_0)\Psi^{-1/2}(\beta_0)\rho^{-1}(\beta_0)$ as defined after Lemma (ref), and

align*[align* omitted — 320 chars of source]

The maximum power loss over $\delta \in \mathcal{D}(\beta_0)$ can be viewed as a maximum regret. Then, we choose $(a_1,a_2)$ that minimizes the maximum regret; that is,

align[align omitted — 303 chars of source]

Four remarks on the domain of $(a_1,a_2)$ (i.e., $\mathbb{A}(\mu_D,\gamma(\beta_0))$) are in order. First, the lower bound $\underline{a}(\mu_D,\gamma(\beta_0))$ is motivated by Theorem (ref)(iii). Specifically, we require $p_1 \in (0,1)$ and close to 0 and $p_2>1$. In the Online Supplement, we provide a detailed report on the finite sample performance of our CLC test for both simulation designs analyzed in Section (ref) and the empirical application in Section (ref), where we consider different values of $p_1$ and $p_2$. The results indicate that our test's finite sample performance is not affected by the specific values chosen for $(p_1,p_2)$, as all the results are very close to those reported in the main paper. Second, under weak identification, $\mu_D$ is bounded, and $\frac{1.1 \mathbb{C}_{\alpha,\max}(\rho(\beta_0)) \Phi_1(\beta_0) c_{\mathcal{B}}(\beta_0) }{ \Delta_*^4(\beta_0) \mu_D^2}$ may be larger than $0.01$. In this case, we have $\mathbb{A}(\mu_D,\gamma(\beta_0)) = \{(a_1,a_2) \in \mathbb{A}_0, a_1 \in [0.01,1]\}$. Third, under strong identification and local alternatives, $\frac{1.1 \mathbb{C}_{\alpha,\max}(\rho(\beta_0)) \Phi_1(\beta_0) c_{\mathcal{B}}(\beta_0) }{ \Delta_*^4(\beta_0) \mu_D^2}$ will converge to zero so that

align*[align* omitted — 240 chars of source]

We show in Theorem (ref) below that in this case, the minimax jackknife CLC test converges to $1\{\mathcal{N}_2^{*2} \geq \mathbb{C}_{\alpha}\}$ defined in Lemma (ref), which is the UMP invariant and unbiased test. Furthermore, the minimax $a_1$ satisfies the requirement in Theorem (ref)(iii) with $\tilde{q} = 1.1\mathbb{C}_{\alpha,\max}(\rho(\beta_0))$ so that under strong identification, our CLC test has asymptotic power 1 against fixed alternatives, as shown in Theorem (ref). Fourth, we require $\overline{a}<1$ for some technical reason. In our simulations, we have not observed the minimax $a_1+a_2$ reaching the upper bound. Therefore, setting the upper bound to $\overline{a}$ or $1$ does not have any numerical impact.

Since we cannot observe the values of $\mu_D$ and $\gamma(\beta_0)$ in practice, we adopt the plug-in method described in Section 6 of I.Andrews(2016). Specifically, we replace $\gamma(\beta_0)$ with its consistent estimator $\widehat{\gamma}(\beta_0)$ as specified in Assumption (ref). To obtain a proxy of $\mu_D$,\footnote{In fact, as $\phi_{a_1,a_2,\infty}(\delta,\mu_D,\gamma(\beta_0))$ only depends on $\mu_D^2$, we aim to find a good estimator for $\mu_D^2$.} we define

align*[align* omitted — 380 chars of source]

which is a function of $\widehat{\gamma}(\beta_0)$ and a consistent estimator of $\sigma_D$ by Assumption (ref). Then, under weak identification, we have $\widehat{D}^2/\widehat{\sigma}_D^2 = D^2 /\sigma_D^2 + o_p(1) \stackrel{d}{=} \mathcal{Z}^2(\mu_D/\sigma_D) + o_p(1)$ and $D^2 /\sigma_D^2$ is a sufficient statistic for $\mu_D^2$. Let $\widehat{r} = \widehat{D}^2/\widehat{\sigma}_D^2$. We consider two estimators for $\mu_D$ as functions of $\widehat{D}$ and $\widehat{\sigma}_D$, namely, $f_{pp}(\widehat{D},\widehat{\gamma}(\beta_0)) = \widehat{\sigma}_D \sqrt{\widehat{r}_{pp}}$ and $f_{krs}(\widehat{D},\widehat{\gamma}(\beta_0)) = \widehat{\sigma}_D \sqrt{\widehat{r}_{krs}}$, where $\widehat{r}_{pp} = \max(\widehat{r}-1,0)$ and

align*[align* omitted — 186 chars of source]

Specifically, krs93 show that $\widehat{r}_{krs}$ is positive as long as $\widehat{r}>0$ and $\widehat{r} \geq \widehat{r}_{krs} \geq \widehat{r}-1$. It is also possible to consider the MLE based on a single observation $\widehat{D}^2/\widehat{\sigma}_D^2$. However, such an estimator is harder to use because it does not have a closed-form expression.

In practice, we estimate $\mathbb{E}\phi_{a_1,a_2,\infty}(\delta,\mu_D,\gamma(\beta_0))$ by $ \mathbb{E}^*\phi_{a_1,a_2,s}(\delta,\widehat{D},\widehat{\gamma}(\beta_0))$ for $s \in \{pp,krs\}$, where

align[align omitted — 632 chars of source]

and $(\widehat{C}_{1}(\delta),\widehat{C}_{2}(\delta))$ are similarly defined as $(C_1(\delta),C_2(\delta))$ in (ref) with $\gamma(\beta_0)$ replaced by $\widehat{\gamma}(\beta_0)$; that is,

align*[align* omitted — 649 chars of source]

Let $\mathcal{P}_{\delta,s}(\widehat{D},\widehat{\gamma}(\beta_0)) = \sup_{(a_1,a_2) \in \mathbb{A}(f_{s}(\widehat{D},\widehat{\gamma}(\beta_0)),\widehat{\gamma}(\beta_0)) } \mathbb{E}^*\phi_{a_1,a_2,s}(\delta,\widehat{D},\widehat{\gamma}(\beta_0))$. Then, for $s \in \{pp,krs\}$, we can estimate $a(\mu_D,\gamma(\beta_0))$ in (ref) by $\mathcal{A}_s(\widehat{D},\widehat{\gamma}(\beta_0)) = (\mathcal{A}_{1,s}(\widehat{D},\widehat{\gamma}(\beta_0)),\mathcal{A}_{2,s}(\widehat{D},\widehat{\gamma}(\beta_0)))$ defined as

align[align omitted — 402 chars of source]

where $\phi_{a_1,a_2,s}(\delta,\widehat{D},\widehat{\gamma}(\beta_0))$ is defined in (ref),

align*[align* omitted — 241 chars of source]
align*[align* omitted — 329 chars of source]
align*[align* omitted — 390 chars of source]

and $\widehat{\Delta}_*(\beta_0) = \widehat{\Phi}_1^{1/2}(\beta_0)\widehat{\Psi}^{-1/2}(\beta_0) \widehat{\rho}^{-1}(\beta_0)$. Then, the feasible jackknife CLC test is, for $s \in \{pp,krs\}$,

align[align omitted — 532 chars of source]

Asymptotic Properties

We first consider the asymptotic properties of the jackknife CLC test under weak identification and fixed alternatives, in which $\mathcal{C}\rightarrow \widetilde{\mathcal{C}}$ and $\Delta$ is treated as fixed so that we have

align*[align* omitted — 92 chars of source]

We see from (ref) and (ref) that $\mathcal{A}_s(d,r) = (a_1(f_s(d,r),r),a_2(f_s(d,r),r))$ is a function of $(d,r) \in \Re \times \Gamma$, where $\Gamma$ is the parameter space for $\gamma(\beta_0)$ and $s \in \{pp,krs\}$. We make the following assumption on $\mathcal{A}_s(\cdot)$.

assLet $\mathcal{S}_s$ be the set of discontinuities of $\mathcal{A}_s(\cdot,\gamma(\beta_0)): \Re \mapsto [0,1] \times [0,1]$. Then, we assume $\mathcal{A}_s(d,r)$ is continuous in $r$ for any $d \in \Re/\mathcal{S}_s$, and the Lebesgue measure of $\mathcal{S}_s$ is zero for $s \in \{pp,krs\}$.

Assumption (ref) is a technical condition that allows us to apply the continuous mapping theorem. It is mild because $\mathcal{A}_s(\cdot)$ is allowed to be discontinuous in its first argument. In practice, we can approximate $\mathcal{A}_s(\cdot)$ by a step function defined over a grid of $d$ so that there is a finite number of discontinuities. The continuity of $\mathcal{A}_s(\cdot)$ in its second argument is due to the smoothness of the bivariate normal PDF with respect to the covariance matrix. Therefore, in this case, Assumption (ref) holds automatically.

thmSuppose we are under weak identification and fixed alternatives and that Assumptions (ref)--(ref) hold. Then, for $s\in \{pp,krs\}$, \begin{align*} \mathcal{A}_s(\widehat{D},\widehat{\gamma}(\beta_0)) & \rightsquigarrow \mathcal{A}_s(D,\gamma(\beta_0))= (a_1(f_s(D,\gamma(\beta_0)),\gamma(\beta_0)),a_2(f_s(D,\gamma(\beta_0)),\gamma(\beta_0))) \end{align*} and\footnote{We assume that $\frac{C}{0} = +\infty$ if $C>0$ and $\min(C,+\infty) = C.$} \begin{align*} \mathbb{E}\widehat{\phi}_{\mathcal{A}_s(\widehat{D},\widehat{\gamma}(\beta_0))} \rightarrow \mathbb{E}\phi_{a_1(f_s(D,\gamma(\beta_0)),\gamma(\beta_0)),a_2(f_s(D,\gamma(\beta_0)),\gamma(\beta_0)),\infty}(\Delta,\mu_D,\gamma(\beta_0)), \end{align*} where $\phi_{a_1,a_2,\infty}(\delta)$ is defined in (ref) and $a_l(f_s(D,\gamma(\beta_0)),\gamma(\beta_0))$ is interpreted as $a_l(\mu_D,\gamma(\beta_0))$ defined in (ref) with $\mu_D$ replaced by $f_s(D,\gamma(\beta_0))$ for $l = 1,2$. In addition, let $BL_1$ be the class of functions $h(\cdot)$ of $D$ that is bounded and Lipschitz with Lipschitz constant 1. Then, if the null hypothesis holds such that $\Delta = 0$, we have \begin{align*} \mathbb{E}(\widehat{\phi}_{\mathcal{A}_s(\widehat{D},\widehat{\gamma}(\beta_0))} - \alpha)h(\widehat{D}) \rightarrow 0, \quad \forall h \in BL_1. \end{align*}

Several remarks on Theorem (ref) are in order. First, we see that the power of our jackknife CLC test is $\mathbb{E}\phi_{\mathcal{A}_s(D,\gamma(\beta_0)),\infty}(\Delta,\mu_D,\gamma(\beta_0))$, which does not exactly match the minimax power $$\mathbb{E}\phi_{a_1(\mu_D,\gamma(\beta_0)),a_2(\mu_D,\gamma(\beta_0)),\infty}(\Delta,\mu_D,\gamma(\beta_0))$$ in the limit problem. This is because under weak identification, it is impossible to consistently estimate $\mu_D$, or equivalently, the concentration parameter. A similar result holds under weak identification with a fixed number of moment conditions in I.Andrews(2016). The best we can do is to approximate $\mu_D$ by reasonable estimators based on $D$ such as $f_{pp}(D,\gamma(\beta_0))$ and $f_{krs}(D,\gamma(\beta_0))$. Second, Theorem (ref) implies that our jackknife CLC test controls size asymptotically conditionally on $\widehat{D}$, and thus, unconditionally. Last, according to Theorem (ref), the CLC test's asymptotic power, with weights $(a_1,a_2)$ chosen through the minimax procedure, is equivalent to the limit experiment's asymptotic power when the weights are $\mathcal{A}_s(D,\gamma(\beta_0))$, which is a function of $D$. As $D$ is independent of the normal random variables in $\phi_{a_1,a_2,\infty}(\delta)$ in (ref), the two optimality results stated in Theorem (ref)(i) also hold asymptotically, conditional on $\widehat{D}$. To make this statement precise, we define the eigenvalue decomposition

align[align omitted — 838 chars of source]

Define a class of tests

align*[align* omitted — 471 chars of source]

where $(\mathcal{Z}_1,\mathcal{Z}_2)$ are two independent standard normal random variables. Further define, for $s \in \{pp,krs\}$,

align*[align* omitted — 240 chars of source]
assSuppose $\mathcal{U}_s(d,r)$ is continuous in $r$ and the set of discontinuities of $\mathcal{U}_s(\cdot)$ w.r.t. its first argument has zero Lebesgue measure.
corSuppose we are under weak identification and fixed alternatives and that Assumptions (ref)--(ref) hold. Let $\tilde \phi(\cdot) \in \Phi_\alpha$ and for any $d \in \Re$, denote $(\theta_1,\theta_2) = (m_1(\Delta),m_2(\Delta))\mathcal{U}_s(d,\gamma(\beta_0))$. Then, the following two optimality results hold. \begin{enumerate}[label=(\roman*)] • If for some $d \in \Re$ and $s \in \{pp,krs\}$, we have \begin{align*} \lim_{\varepsilon \rightarrow 0} \lim_{n\rightarrow \infty} & \frac{\mathbb{E} \tilde \phi(\widetilde {AR}_s^2(\beta_0), \widetilde {LM}^{*2}_s(\beta_0),\widehat{D},\widehat{\gamma}(\beta_0)) 1\{ |\widehat D -d| \leq \varepsilon \} }{\mathbb{E} 1\{ |\widehat D -d| \leq \varepsilon \} } \\ & \geq \lim_{\varepsilon \rightarrow 0} \lim_{n\rightarrow \infty} \frac{\mathbb{E} \hat \phi_{\mathcal{A}_s(\widehat D, \widehat \gamma(\beta_0))}1\{ |\widehat D -d| \leq \varepsilon \} }{\mathbb{E} 1\{ |\widehat D -d| \leq \varepsilon \} }, \end{align*} for all $(\theta_1,\theta_2) \in \Re^2$, then \begin{align*} \lim_{\varepsilon \rightarrow 0} \lim_{n\rightarrow \infty} & \frac{\mathbb{E} \tilde \phi(\widetilde {AR}_s^2(\beta_0), \widetilde {LM}^{*2}_s(\beta_0),\widehat{D},\widehat{\gamma}(\beta_0)) 1\{ |\widehat D -d| \leq \varepsilon \} }{\mathbb{E} 1\{ |\widehat D -d| \leq \varepsilon \} } \\ & = \lim_{\varepsilon \rightarrow 0} \lim_{n\rightarrow \infty} \frac{\mathbb{E} \hat \phi_{\mathcal{A}_s(\widehat D, \widehat \gamma(\beta_0))}1\{ |\widehat D -d| \leq \varepsilon \} }{\mathbb{E} 1\{ |\widehat D -d| \leq \varepsilon \} }, \end{align*} for all $(\theta_1,\theta_2) \in \Re^2$. • If $( \theta_1^2, \theta_2^2) = b \cdot (\nu_1(d,\gamma(\beta_0)),\nu_2(d,\gamma(\beta_0)))$ for some positive constant $b$, then there exists $\overline{b}>0$ such that if $0<b<\overline{b}$, we have \begin{align*} \lim_{\varepsilon \rightarrow 0} \lim_{n\rightarrow \infty} & \frac{\mathbb{E} \tilde \phi(\widetilde {AR}_s^2(\beta_0), \widetilde {LM}^{*2}_s(\beta_0),\widehat{D},\widehat{\gamma}(\beta_0)) 1\{ |\widehat D -d| \leq \varepsilon \} }{\mathbb{E} 1\{ |\widehat D -d| \leq \varepsilon \} } \\ & \leq \lim_{\varepsilon \rightarrow 0} \lim_{n\rightarrow \infty} \frac{\mathbb{E} \hat \phi_{\mathcal{A}_s(\widehat D, \widehat \gamma(\beta_0))}1\{ |\widehat D -d| \leq \varepsilon \} }{\mathbb{E} 1\{ |\widehat D -d| \leq \varepsilon \} }, \end{align*} \end{enumerate}

Corollary (ref) shows that under weak identification and fixed alternatives, our jackknife CLC test is asymptotically admissible and optimal against certain alternatives conditional on $\widehat D$.

Next, we consider the performance of $\widehat{\phi}_{\mathcal{A}_s(\widehat{D},\widehat{\gamma}(\beta_0))} $ defined in (ref) under strong identification and local alternatives. To precisely state the optimality result, we further consider the class of level-$\alpha$ tests against $\theta =0$ v.s. the two-sided alternative that are constructed based on one observation of $(\mathcal{N}_1,\mathcal{N}_2)$, where $\theta = \widetilde \Delta \widetilde {\mathcal{C}} \Psi^{-1/2}$ and

align*[align* omitted — 231 chars of source]

Specifically, denote

align*[align* omitted — 315 chars of source]

and

align*[align* omitted — 335 chars of source]

as the classes of sign-invariant and unbiased tests, respectively.

thmSuppose that Assumptions (ref) and (ref) hold. Further suppose that we are under strong identification and local alternatives as described in Lemma (ref). Then, for $s \in \{pp,krs\}$, we have \begin{align*} \mathcal{A}_{1,s}(\widehat{D},\widehat{\gamma}(\beta_0)) \stackrel{p}{\longrightarrow} 0, \quad \mathcal{A}_{2,s}(\widehat{D},\widehat{\gamma}(\beta_0))\rho \stackrel{p}{\longrightarrow} 0, \quad and \quad \widehat{\phi}_{\mathcal{A}_s(\widehat{D},\widehat{\gamma}(\beta_0))} \rightsquigarrow 1\{\mathcal{N}_2^{*2} \geq \mathbb{C}_{\alpha}\}, \end{align*} where $\mathcal{N}_2^{*} \stackrel{d}{=} \mathcal{N}\left(\frac{\widetilde{\Delta} \widetilde{\mathcal{C}}}{[(1-\rho^2)\Psi]^{1/2}},1\right)$. In addition, suppose $\breve{\phi}_n = \phi(AR(\beta_0),LM(\beta_0)) +o_P(1)$ for some $\phi \in \Phi_{\alpha}^I \cup \Phi_{\alpha}^U$ and the sequence $\{\breve{\phi}_n \}_{n \geq 1}$ is uniformly integrable. Then, we have \begin{align*} \lim_{n \rightarrow \infty} \mathbb{E}\widehat{\phi}_{\mathcal{A}_s(\widehat{D},\widehat{\gamma}(\beta_0))} = \sup_{\phi \in \Phi_{\alpha}^I \cup \Phi_{\alpha}^U } \lim_{n \rightarrow \infty} \mathbb{E}\phi(AR(\beta_0),LM(\beta_0)) \geq \lim_{n \rightarrow \infty} \mathbb{E}\breve{\phi}_n. \end{align*}

Five remarks are in order. First, under strong identification, $\mu_D$, and thus, $D$ approaches infinity, and so does our estimator $\widehat D$. This is how our estimator $\widehat D$ can detect the identification strength. In addition, we show in the proof of Theorem (ref) that under strong identification, the calibrated power gap $\mathcal{P}_{\delta,s}(\widehat{D},\widehat{\gamma}(\beta_0)) - \mathbb{E}^*\phi_{a_1,a_2,s}(\delta,\widehat{D},\widehat{\gamma}(\beta_0))$ is maximized when $\delta$ is in the region of local alternatives. However, in this region, as shown by Lemma (ref), the maximum power gap can achieve zero if all the weights are put on $LM^*(\beta_0)$, which leads to the first result in Theorem (ref). Second, our jackknife CLC test is adaptive to identification strength. In practice, econometricians do not know whether the true value $\beta$ is close to the null $\beta_0$. Therefore, our jackknife CLC test calibrates power across all possible values of $\delta$ (i.e., $\delta \in \mathcal{D}(\beta_0)$), which include both local and fixed alternatives. Yet, Theorem (ref) shows that the minimax procedure can produce the most powerful test as if it is known that $\beta$ belongs to the region of local alternatives. Third, Theorem (ref) shows that under strong identification and local alternatives, our jackknife CLC test converges to the UMP level-$\alpha$ test that is either invariant to the sign change or unbiased and constructed based on $AR(\beta_0)$ and $LM(\beta_0)$. Therefore, it is more powerful than the jackknife AR and LM tests. Fourth, under strong identification and local alternatives, the JIVE-based Wald test proposed by Chao(2012) is asymptotically equivalent to the jackknife LM test, which implies that the jackknife AR and JIVE-Wald-based two-step test in MS22 is also dominated by the jackknife CLC test. Fifth, consider the HLIM based Wald test statistic proposed by Haus2012, which is denoted as $W_h(\beta_0)$. In Section (ref) in the Online Supplement, we show that, under local alternative and strong identification,

align*[align* omitted — 142 chars of source]

where $\tilde{\rho} = plim_{n\rightarrow \infty} X^\top e(\beta_0)/(e(\beta_0)^\top e(\beta_0))$ and $\Psi_h = \Psi - 2 \tilde \rho \Phi_{12} + \tilde \rho^2 \Phi_1$ is the corresponding asymptotic variance. Then, by letting $\breve{\phi}_n = 1\{W_h^2(\beta_0) \geq \mathbb{C}_{\alpha}\}$ and

align*[align* omitted — 205 chars of source]

Theorem (ref) implies our jackknife CLC test is more powerful than the HLIM based Wald test under strong identification against local alternatives. In fact, by direct calculation, we can see that, for $\theta = \widetilde \Delta \widetilde{\mathcal{C}}\Psi^{-1/2}$,

align*[align* omitted — 324 chars of source]

The noncentrality parameter for the HLIM based Wald test is weakly smaller than that of the CLC test, which explains the power comparison. The equality holds if $\tilde \rho \Phi_1^{1/2}\Psi^{-1/2} = \rho$, which further holds in the special case of many weak IVs and homoskedasticity in the sense that $\Pi^\top \Pi/K = o(1)$ and $\mathbb{E}(V_i,e_i)^\top (V_i,e_i)$ does not vary across $i$.

Combining Theorems (ref) and (ref), we can show the uniform size control of our jackknife CLC test no matter the identification is strong or weak. Let $\lambda_n \in \Lambda_n$ be the data generating process of $n$ observations of $(e,V,Z)$. Under $\lambda_n$, the covariance matrix of $(Q_{e,e},Q_{X,e},Q_{X,X})$ is denoted as $\mathbb V_n$. We impose the following restriction on the sequence of classes of DGPs ($\{\Lambda_n \}_{n \geq 1}$):\footnote{In (ref), we focus on the model without exogenous control variables. The independence and moment conditions for $(e_i,V_i)$ are sufficient for Assumption (ref). We further verify in Section (ref) of the Online Supplement that the joint asymptotic normality (Assumption 1) holds in the case with exogenous controls. }

align[align omitted — 633 chars of source]

In Sections (ref) and (ref) of the Online Supplement, we further verify that Assumption (ref) holds, respectively, for the standard variance estimators, which follow the construction in crudu2021, and the cross-fit variance estimators, which follow MS22. Theorem (ref) shows that our jackknife CLC test has correct asymptotic size, under similar arguments as those in ACG2020 and I.Andrews(2016).

thmSuppose Assumption (ref) holds, $\{\Lambda_n \}_{n \geq 1}$ satisfies (ref), and we are under the null hypothesis that $\beta_0 = \beta$. Then, we have \begin{align*} \liminf_{n \rightarrow \infty} \inf_{\lambda_n \in \Lambda_n} \mathbb{E}_{\lambda_n}(\widehat{\phi}_{\mathcal{A}_s(\widehat{D},\widehat{\gamma}(\beta_0))}) = \limsup_{n \rightarrow \infty} \sup_{\lambda_n \in \Lambda_n} \mathbb{E}_{\lambda_n}(\widehat{\phi}_{\mathcal{A}_s(\widehat{D},\widehat{\gamma}(\beta_0))}) = \alpha. \end{align*}

Last, we show that, under strong identification, the jackknife CLC test $\widehat{\phi}_{\mathcal{A}_s(\widehat{D},\widehat{\gamma}(\beta_0))}$ defined in (ref) has asymptotic power 1 against fixed alternatives.

thmSuppose Assumption (ref) holds, and $( Q_{e(\beta_0),e(\beta_0)} - \Delta^2 \mathcal{C}, Q_{X,e(\beta_0)} - \Delta \mathcal{C},Q_{X,X} - \mathcal{C})^\top = O_p(1)$. Further suppose that we are under strong identification with fixed alternatives so that $\Delta = \beta - \beta_0$ is nonzero and fixed. Then, we have \begin{align*} \widehat{\phi}_{\mathcal{A}_s(\widehat{D},\widehat{\gamma}(\beta_0))} \stackrel{p}{\longrightarrow} 1. \end{align*}

Simulation

Power Curve Simulation for the Limit Problem

In this section, we present simulation results to compare the power performance of various tests under the limit problem described in Section (ref). We consider the following tests with a nominal rate of $5\%$: (i) our jackknife CLC test, where $\mu_D$ is estimated using either $pp$ or $krs$ method, (ii) the one-sided jackknife AR test defined in ((ref)), (iii) the jackknife LM test defined in ((ref)), and (iv) the test that is based on the orthogonalized jackknife LM statistic $LM^{*2}(\beta_0)$ defined in this paper. We conduct 5,000 simulation replications to obtain stable simulation results.

We set the parameter space for $\beta$ as $\mathcal{B} = [-6/\mathcal{C},6/\mathcal{C}]$, where $\mathcal{C} = 3$ and $6$. The choice of parameter space follows that in I.Andrews(2016). We set $\beta_0 = 0$, and the values of the covariance matrix in ((ref)) are set as follows: $\Phi_1 = \Psi = \Upsilon =1$, and $\Phi_{12} = \Phi_{13} = \tau = \rho$, where $\rho \in \{0.2, 0.4, 0.7, 0.9\}$. We then compute $\gamma(\beta_0)$ based on (ref) as $\beta$ ranges over $\mathcal{B}$ and generate $AR(\beta_0)$ and $LM(\beta_0)$ based on (ref). Last, we implement our CLC test purely based on $AR(\beta_0)$, $LM(\beta_0)$, $\gamma(\beta_0)$, and $\mathcal{B}$ without assuming the knowledge of $(\mathcal{C},\beta)$. We have tried to simulate under alternative settings of the covariance matrix, and the obtained patterns of the power behavior are very similar.

Figures (ref)--(ref) plot the power curves for $\rho=0.2, 0.4, 0.7,$ and $0.9$. In each figure, we report the results under both $\mathcal{C}=3$ and $6$. We observe that overall, the two jackknife CLC tests have the best power properties in terms of minimizing the maximum regret. Especially when the identification is relatively strong ($\mathcal{C}=6$) and/or the degree of endogeneity is not very low ($\rho =0.4, 0.7$, or $0.9$), the jackknife CLC tests outperform their AR and LM counterparts by a large margin. In addition, we notice that when $\mathcal{C}=3$, for some parameter values $LM^*(\beta_0)$ can suffer from substantial declines in power relative to the other tests, which is in line with our theoretical predictions. By contrast, our jackknife CLC tests are able to guard against such substantial power loss because of the adaptive nature of their minimax procedure. In Section (ref) of the Online Supplement, we further report power curves for alternative values of the tuning parameters $(p_1,p_2)$ in ((ref)) and of $\mathcal{C}$, and find that the overall patterns remain very similar.

figure[figure omitted — 159 chars of source]
figure[figure omitted — 160 chars of source]
figure[figure omitted — 160 chars of source]
figure[figure omitted — 160 chars of source]

Simulation Based on Calibrated Data

We follow the approach of Angrist-Frandsen2022 and MS22 and use a data generating process (DGP) calibrated based on the 1980 census dataset from Angrist-Krueger(1991). We define the instruments as $$\tilde Z_i = \big( (1 \{Q_i = q, C_i = c\})_{q \in \{2,3,4\}, c \in \{31,\cdots,39\} }, (1 \{Q_i = q, P_i = p \})_{q \in \{2,3,4\}, p \in \{\text{51 states}\} } \big),$$ where $Q_i, C_i, P_i$ are individual $i$'s quarter of birth (QOB), year of birth (YOB) and place of birth (POB), respectively, so that there are 180 instruments. Note that the dummy with $q = 1$ and $c = 30$ is omitted in $Z_i$. We denote $\tilde Y_i$ as income, $\tilde X_i$ as the highest grade completed, and $\tilde W_i$ as the full set of YOB-POB interactions; that is, $$\tilde W_i = \big( 1 \{C_i = c, P_i = p\}_{c \in \{30,...,39\}, p \in \{\text{51 states}\} } \big),$$ which is a $510 \times 1$ matrix.

As in Angrist-Frandsen2022, using the full 1980 sample (consisting of 329,509 individuals), we first obtain the average $\tilde X_i$ for each QOB-YOB-POB cell; we call this $\bar{s}(q,c,p)$. Next we use LIML to estimate the structural parameters in the following linear IV regression: $$\tilde Y_i = \tilde X_i \beta_X + \tilde W_i^\top \beta_W +e_i,$$ $$\tilde X_i = \tilde Z_i^\top \Gamma_Z + \tilde W_i^\top \Gamma_W + V_i,$$ where $\tilde X$ is endogenous and instrumented by $\tilde Z_i$ and $\tilde W_i$ is the exogenous control variable. Denote the LIML estimate for $\beta_{X,W} \equiv (\beta_X^\top, \beta_W^\top)^\top$ as $\widehat\beta_{LIML}^\top = (\widehat\beta_{LIML, X}^\top, \widehat\beta_{LIML,W}^\top)$. We let $\widehat{y}(C_i,P_i) = \tilde W_i^\top \widehat{\beta}_{LIML,W}$ and $$\omega(Q_i,C_i,P_i) = \tilde Y_i - \tilde X_i \widehat{\beta}_{LIML,X} - \tilde W_i^\top \widehat{\beta}_{LIML,W}.$$

Based on the LIML estimate and the calibrated $\omega(Q_i,C_i,P_i)$, we simulate the following two DGPs:

enumerate• DGP 1: \begin{align} \widetilde{y}_i = \bar{y} + \beta \widetilde{s}_i + \omega(Q_i,C_i,P_i)(\nu_i +\kappa_2 \xi_i) \end{align} $$\widetilde{s}_i \sim Poisson(\mu_i),$$ where $\beta$ is the parameter of interest, $\nu_i$ and $\xi_i$ are independent standard normal, $\bar{y} = \frac{1}{n}\sum_{i=1}^n \widehat{y}(C_i,P_i)$, $\mu_i \equiv max\{1,\gamma_0+ \gamma_Z^\top \tilde Z_i + \kappa_1 \nu_i \}$, and $\gamma_0 +\gamma^\top_Z \tilde Z_i$ is the projection of $\bar{s}_i(q,c,p)$ onto a constant and $\tilde Z_i$. We set $\kappa_1 = 1.7$ and $\kappa_2 = 0.1$ as in MS22. • DGP 2: Same as DGP 1 except that $\kappa_1 = 2.7$ and $$\widetilde{s}_i \sim \lfloor Poisson(2\mu_i)/2 \rfloor.$$

We consider sample sizes of 0.5%, 1%, and 1.5% of the full sample size. Upon obtaining $n$ observations, we exclude instruments with $\sum_{i=1}^n \tilde Z_{ij} <5$. This results in three different sample sizes: small, medium, and large, with 1,648, 3,296, and 4,943 observations, respectively. The number of instruments also varies across sample sizes, with 119, 142, and 150 instruments for small, medium, and large samples, respectively. Our DGP 1 is exactly the same as that in MS22, with the correlation parameter of $\rho =0.41$. DGP 2 has a higher correlation parameter of $\rho =0.7$. The identification strength increases with the sample size. For DGP 1, the concentration parameters $\mathcal{C}/\Upsilon^{1/2}$ for small, medium, and large samples are 2.15, 3.62, and 4.85, respectively. For DGP 2, they are 2.38, 3.97, 5.28, respectively.

We emphasize that following Angrist-Frandsen2022 and MS22, we only use $\tilde W_i$ to compute the LIML estimator and calibrate $\omega(Q_i,C_i,P_i)$, but do not use it to generate new data. Therefore, for the simulated data, the outcome variable is $\tilde{y}_i$, the endogenous variable is $\tilde s_i$, the IV $\tilde Z_i$ is viewed to be fixed, and the exogenous control variable is just an intercept. We then denote the demeaned versions of $\tilde{y}_i$, $\tilde{s}_i$, and $\tilde Z_i$ as $Y_i$, $X_i$, and $Z_i$, respectively, in (ref) and implement various inference methods described below. Following MS22, we test the null hypothesis that $\beta = \beta_0$ for $\beta_0 =0.1$ while varying the true value $\beta \in \mathcal{B}$. The parameter space is set as $\mathcal{B} = [-0.5,0.5]$, which is consistent with the choice of parameter space for the empirical application below. The results below are based on 1,000 simulation repetitions. We provide more details about the implementation in Section (ref) in the Online Supplement. We set $(p_1,p_2) = (0.01,1.1)$ in (ref). Additional simulation results using other choices of $(p_1,p_2)$ and $\mathcal{B}$ are reported in Section (ref) in the Online Supplement. All of them are very close to what we report here.

We compare the following tests with a nominal rate of $5\%$:

enumerate• pp: our jackknife CLC test when $\mu_D$ is estimated by the method $pp$. • krs: our jackknife CLC test when $\mu_D$ is estimated by the method $krs$. • AR: the one-sided jackknife AR test with the cross-fit variance estimator proposed by MS22. • LM_CF: Matsushita-Otsu2021's (Matsushita-Otsu2021) jackknife LM test, but with a cross-fit variance estimator (details are given in Section (ref) in the Online Supplement). • 2-step: MS22's (MS22) two-step estimator in which the overall size is set at $5\%$. • LM$^*$: LM$^*$ test defined in this paper. • LM_MO: Matsushita-Otsu2021's (Matsushita-Otsu2021) original jackknife LM test.
figure[figure omitted — 135 chars of source]
figure[figure omitted — 135 chars of source]

Figures (ref) and (ref) plot the power curves of the aforementioned tests. We can make four observations. First, all methods control size well because they are all weak identification robust. Second, the performance of the jackknife CLC test with $krs$ is slightly better than that with $pp$, which is consistent with the power curve simulation in Section (ref). Third, in DGP 1 with a small sample size, the power of the jackknife AR test is at most about 9.2% higher than that of the $krs$ test when $\beta$ is around -0.3. However, for alternatives close to the null (e.g., when $\beta$ is around 0), the power of the $krs$ test is 24% higher, which implies that the power of the $krs$ test is still better than that for the jackknife AR test in the minimax sense. The power of the jackknife LM tests is similar to that of the $krs$ test in DGP 1 with a small sample size. Fourth, for the rest of the scenarios, the power of the $krs$ test is the highest in most regions of the parameter space. The power of the jackknife AR and LM is at most 0.7% higher than that of the $krs$ test at some point. For DGP 1 with medium and large sample sizes, the maximum power gaps between our $krs$ test and the jackknife LM are about 8.6% and 5.6%, and about 43.2% and 50% compared with the jackknife AR. Furthermore, they are 23.3%, 19.5%, and 18.5% compared with the jackknife LM for DGP 2 with small, medium, and large sample sizes, respectively, and about 41.5%, 55.3%, and 55.85% compared with the jackknife AR.

Figures (ref) and (ref) show the average values of $(a_1,a_2)$, which represents the weights assigned to $AR(\beta_0)$ and $LM(\beta_0)$ in our CLC tests, under DGPs 1 and 2, respectively. The weight assigned to $LM^*(\beta_0)$ is simply $1-a_1-a_2$. As shown in Table (ref), under weak identification and fixed alternatives, there is no clear winner among $AR(\beta_0)$, $LM(\beta_0)$, and $LM^*(\beta_0)$, and thus, our CLC test assigns weights to all the three tests. However, under strong identification and local alternative, $LM^*(\beta_0)$ is the UMP test and should carry all the weights, which means $a_1+a_2$ should be minimum. On the other hand, under strong identification and for some fixed alternatives, $LM^*(\beta_0)$ may lack power while both $AR(\beta_0)$ and $LM(\beta_0)$ have power 1. In this case, as long as we do not assign all weights on $LM^*(\beta_0)$, our CLC test should also have power 1. We observe that our simulation results are consistent with these theoretical predictions. First, when $\beta_0$ is close to the null 0.1, both $a_1$ and $a_2$ are small, indicating that most of the weights are put on $LM^*(\beta_0)$. Second, we observe from Figures (ref) and (ref) that the power of $LM^*(\beta_0)$ drops rapidly when $\beta$ is smaller than around zero. Therefore, our CLC test assigns more weights on $AR(\beta_0)$ and $LM(\beta_0)$. Third, for distant alternatives, significant weights are assigned to $AR(\beta_0)$ and $LM(\beta_0)$, which ensures the good power of our CLC test. Additionally, we note that the weights assigned to $AR(\beta_0)$ ($a_1$) are higher on the left side of the parameter space relative to the right, since $AR(\beta_0)$ is more powerful on the left.

figure[figure omitted — 146 chars of source]
figure[figure omitted — 146 chars of source]

Empirical Application

In this section, we consider the linear IV regressions with the specification underlying Angrist-Krueger(1991), using the full original dataset.\footnote{The dataset can be downloaded from MIT Economics, Angrist Data Archive, https://economics.mit.edu/faculty/angrist/data1/data/angkru1991.} The outcome variable $Y$ and endogenous variable $X$ are log weekly wages and schooling, respectively. We follow Angrist-Krueger(1991) and focus on two specifications with 180 and 1,530 instruments. The 180 instruments consist of 30 quarter and year of birth interactions (QOB-YOB) and 150 quarter and place of birth interactions (QOB-POB). The second specification includes full interactions among QOB-YOB-POB, resulting in 1,530 instruments. The exogenous control variables have been partialled out from the outcome, endogenous variables, and IVs. Further details on the empirical application can be found in Section (ref) in the Online Supplement. The considered tests are similar to those in the previous section. The jackknife AR test is defined in (ref) with $\widehat{\Phi}_1$ being the cross-fit estimator in MS22. The jackknife LM test is defined in ((ref)) with the cross-fit estimator for $\Psi(\beta_0)$. The $pp$ and $krs$ tests are our jackknife CLC tests. The two-step procedure is given by MS22. Specifically, the researcher accepts the null if $\widetilde{F}>9.98$ and $Wald(\beta_0) < \mathbb{C}_{0.02}$\footnote{$\widetilde{F} = Q_{X,X}/\widehat{\Upsilon}$, where $\widehat{\Upsilon}$ is the cross-fit estimator. $Wald(\beta_0)$ is defined as $\left(\frac{\hat{\beta} - \beta_0}{\hat{V}}\right)^2$, where $\hat{\beta}$ is the JIVE estimator and $\hat{V}$ is a cross-fit estimator of the asymptotic variance of $\hat{\beta}$. We refer interested readers to MS22 for more details.} or if $\widetilde{F}\leq 9.98$ and $AR(\beta_0) < \emph{z}_{0.02}$. In the case of 180 instruments, because $\widetilde{F} = 13.42 > 9.98$, the lower and upper bounds of the 95% confidence interval (CI) for the two-step procedure correspond respectively to the minimum and maximum of the set $\{\beta_0 \in \Re: Wald(\beta_0)<\mathbb{C}_{0.02}\}$; similarly, for the 1,530 instruments, as $\widetilde{F} = 6.32 \leq 9.98$, the lower and upper bounds of the CI for the two-step procedure correspond respectively to the minimum and maximum of the set $\{\beta_0 \in \Re: AR(\beta_0)<\emph{z}_{0.02} \}$. We also report the 95% Wald test CI based on the JIVE estimator, denoted as JIVE-t. Table (ref) reports the 95% CIs by inverting the corresponding 5% tests mentioned above for the parameter space $\mathcal{B} = [-0.5,0.5]$. Note all CIs except JIVE-t are robust to weak identification. As $\widetilde{F}$'s are higher than $4.14$ in both cases, the JIVE-t (5%) has the Stock-Yogo(2005b) (Stock-Yogo(2005b))-type guarantee with at most a 5% size distortion (i.e., the overall size is less than 10%). We set $(p_1,p_2)$ in (ref) as $(0.01,1.1)$. The empirical results with other choices of $(p_1,p_2)$ and $\mathcal{B}$ are reported in Section (ref) in the Online Supplement. All of them are very close to what we report here.

table[table omitted — 1,079 chars of source]

Table (ref) highlights that the CIs generated by our jackknife CLC tests are the shortest among all the weak identification robust CIs (i.e., pp, krs, jackknife AR, jackknife LM, and two-step). Furthermore, the jackknife CLC CIs are $7.6\%$ and $2.0\%$ shorter than the non-robust JIVE-t CIs with 180 and 1,530 instruments, respectively, which is in line with our theoretical result that the CLC tests are adaptive to the identification strength and efficient under strong identification.