Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,743 characters · 11 sections · 0 citation commands
\centerline{{\bf Inference with Many Weak Instruments }}
\centerline{Anna Mikusheva\footnote{ Department of Economics, M.I.T. Address: 77 Massachusetts Avenue, E52-526, Cambridge, MA, 02139. Email: [email removed]. National Science Foundation support under grant number 1757199 is gratefully acknowledged. We are grateful to Josh Angrist, Kirill Evdokimov, Whitney Newey and Mikkel S{\o}lvsten for advice, to Brigham Frandsen for sharing code for simulations, and to Ben Deaner and Sylvia Klosin for research assistance.} and Liyang Sun\footnote{UC Berkeley and CEMFI. Email: [email removed]. Support from the Jerry A. Hausman Graduate Dissertation Fellowship and Institute of Education Sciences, U.S. Department of Education, through Grant R305D200010, is gratefully acknowledged.}}
\centerline{\bf Abstract}
{{We develop a concept of weak identification in linear IV models in which the number of instruments can grow at the same rate or slower than the sample size. We propose a jackknifed version of the classical weak identification-robust Anderson-Rubin (AR) test statistic. Large-sample inference based on the jackknifed AR is valid under heteroscedasticity and weak identification. The feasible version of this statistic uses a novel variance estimator. The test has uniformly correct size and good power properties. We also develop a pre-test for weak identification that is related to the size property of a Wald test based on the Jackknife Instrumental Variable Estimator. This new pre-test is valid under heteroscedasticity and with many instruments. }
\noindentKey words: instrumental variables, weak identification, dimensionality asymptotics.
\noindentJEL classification codes: C12, C36, C55.}
Recent empirical applications of instrumental variables (IV) estimation often involve many instruments that together may or may not be strongly relevant. For example, in a prominent paper by Angrist and Krueger (1991) that started the weak IV literature, the authors construct 180 instruments by interacting dummies for the quarter of birth with state and year of birth, and use these instruments to study the effect of schooling on wage. Other examples include papers that employ an empirical strategy known as “judge design” (Maestas et al., 2013; Sampat and Williams, 2015; Dobbie et al., 2018). Fueled by rich administrative data, these papers use the exogenous assignment of cases to judges as instruments for treatment. Since each judge can only process a certain number of cases out of the total court cases, the number of judges (the number of instruments) is usually proportional to the sample size. Another example is the famous Fama-MacBeth procedure in Asset Pricing (Fama and MacBeth, 1973; Shanken, 1992), which is equivalent to IV estimation procedure with the number of instruments proportional to the number of assets.
This paper answers three questions in an environment with many instruments: how to define weak identification, what to do if identification is weak, and how to pre-test for weak instruments. We model many-instrument asymptotics by allowing the number of instruments to grow at most proportionally with the sample size. Firstly, we define weak identification for linear IV models with many instruments by providing necessary and sufficient conditions for the existence of a consistent test. Secondly, we introduce a test that works when there are many instruments, but is also robust to weak identification and heteroscedasticity. Finally, we propose a pre-test for weak identification. This pre-test forms the basis for a two-step procedure that is analogous to that of Stock and Yogo (2005). The two-step test controls size distortion under many-instrument asymptotics, regardless of the strength of identification or the presence of heteroscedasticity.
We define weak identification as a situation where an analog of the concentration parameter divided by the square root of the number of instruments stays bounded in large samples. We prove that even in a homoscedastic model with known covariance, an asymptotically consistent test does not exist if the ratio of the concentration parameter over the square root of the number of instruments stays bounded in large samples. Thus, a necessary condition for a consistent test to exist is that the concentration parameter grows faster than the square root of the number of instruments. Later, we show that this is also a sufficient condition by constructing a robust test that becomes consistent when this condition is satisfied.
We propose a new jackknifed version of the Anderson-Rubin (AR) test which is robust to both weak identification and heteroscedasticity in a model with many instruments. The new test uses an asymptotic approximation based on a Central Limit Theorem (CLT) for quadratic forms. The new AR test has the correct size regardless of identification strength and becomes consistent as soon as the concentration parameter grows faster than the square root of the number of instruments.
As an important technical contribution, we introduce a novel variance estimator for the quadratic form CLT in the absence of a consistent estimator for the structural parameter. The target variance is a quadratic form of the individual (heteroscedastic) variances of errors. We apply cross-fitting (Newey and Robins, 2018; Kline et al., 2020) to produce unbiased proxies for the individual variances of errors. We adjust the quadratic form to remove the bias due to correlations between proxies. We prove the consistency of the new estimator under the null and local alternatives under a wide range of identification scenarios.
Finally, we propose a new pre-test for weak identification which is easy to use and is consistent with our definition of weak identification. An empirical researcher can use our pre-test to decide between employing our jackknife AR test if the pre-test suggests that the identification is weak or a Wald test based on the Jackknife Instrumental Variable Estimator (JIVE, Angrist et al., 1999) if the pre-test suggests that the identification is strong. We guarantee the size of this two-step procedure. Chao et al. (2012) prove that JIVE is consistent in a heteroscedastic model when the concentration parameter grows faster than the square root of the number of instruments. Chao et al. (2012) also derive a consistent estimator of the JIVE standard error. The two-step procedure is appealing because when identification is strong, the JIVE-Wald is more efficient and easy to implement and report.
Our pre-test is in the spirit of Stock and Yogo (2005), but it differs from theirs in two important ways. Firstly, our pre-test allows for a general form of heteroscedasticity, while the pre-test proposed in Stock and Yogo (2005) works only under conditionally homoscedastic errors. Secondly, the Stock and Yogo (2005) pre-test is designed for a small number of instruments and is based on the Two-Stage Least Squares (TSLS) estimator. With many instruments TSLS is consistent only when the concentration parameter grows faster than the number of instruments, which makes the Stock and Yogo (2005) pre-test not very informative.
We apply our pre-test to Angrist and Krueger (1991) and find that their identification is strong. Consequently the JIVE confidence set is reliable (has coverage within 5% tolerance level of the declared coverage). Our weak identification-robust jackknife AR confidence set is somewhat wider than the JIVE confidence set but is still informative.
{\bf{Relation to the Literature}}. Our paper contributes to both the literature on weak IV and the literature on many instruments. The weak IV literature relates identification strength to the size of the concentration parameter and proposes robust tests that work only when there are a small number of instruments. Generalizations to many weak instruments either strongly restrict the number of instruments (Andrews and Stock, 2007) or work only under homoscedasticity (Anatolyev and Gospodinov, 2011). Crudu et al. (2020) recently proposed a jackknife-type AR test that is robust towards weak identification and heteroscedasticity. While employing a similar statistic, their test uses a variance estimator different from ours, that can be shown to lead to a power loss at distant alternatives and inconsistency of the test in some settings where our test is consistent.
The many weak instruments literature started with a prominent paper by Bekker (1994). It mostly establishes conditions for consistency and asymptotic gaussianity for particular estimators. For example, Chao and Swanson (2005) show that in a homoscedastic model limited information maximum likelihood (LIML) and bias-corrected TSLS (BTSLS) are consistent when the concentration parameter grows faster than the square root of the number of instruments. In a heteroscedastic model, consistency of LIML and BTSLS requires that the concentration parameter grows faster than the number of instruments. By contrast, JIVE remains consistent when the concentration parameter grows faster than the square root of the number of instruments (Chao et al., 2012). Our paper shows that the condition in Chao et al. (2012) is necessary for consistency and if it is violated it is impossible to consistently distinguish between any two values of the structural parameter.
The remainder of this paper is organized as follows. Section (ref) summarizes our proposal for empirical researchers. In Section (ref) we introduce our definition of weak identification in an environment with many instruments. In Section (ref) we construct the jackknife AR test and establish its power properties. In Section (ref) we present the pre-test and prove that it controls size. Section (ref) conducts a simulation exercise inspired by Angrist and Frandsen (2019), and Section (ref) concludes. Some proofs and additional results may be found in the Supplementary Appendix.
In empirical applications using instrumental variables, concerns about weak identification are widespread. The current consensus practice is to report the first stage $F$ statistic and as long as it is above 10, researchers are allowed to rely on standard $t$-statistics inferences. This practice has foundations in Stock and Yogo (2005) which showed that the concentration parameter fully characterizes the size distortion of the TSLS-Wald test, and empirically the concentration parameter can be judged based on the first stage $F$ statistics. This result has been obtained under the assumptions of homoscedasticity and for a fixed number of instruments.
While the first stage $F$ pre-test provides reasonable classification for homoscedastic IV models with a small number of instruments, it is inadequate for settings with many instruments. Hansen et al (2008) argue that the TSLS estimator should not be used in applications with many instruments as it becomes very biased. They also argue that a low first stage $F$ statistic is not always indicative of a weak identification issue and $t$-statistics inferences based on more appropriate estimators other than TSLS, along with corrected standard errors, may still be reliable. Estimators with known good properties in heteroscedastic settings with many instruments include JIVE (Chao and Swanson, 2005) and heteroscedasticity-robust Fuller (Hausman et al, 2012).
While theoretical econometrics literature provides recommendations on the choice of estimator, it is largely silent on how to determine whether the concerns of weak identification are valid in a given data set with many instruments. This prompted empirical researchers to formulate econometric arguments and perform simulation studies to support their usage of $t$-statistics. For example, Bhuller et al (2020), recently published in the Journal of Political Economy, used judges design instruments to study the effects of incarceration on recidivism and employment. Concerned about potentially having many weak instruments, Bhuller et al (2020) included a 10-page-long Appendix D with a simulation study to support their usage of the JIVE $t$-statistic.
Our paper proposes a new recipe for empirical researchers to gauge weak identification in applications with many instruments. Specifically, we argue that theoretically, the strength of identification is measured by the concentration parameter divided by the square root of the number of instruments. This is in contrast to the first stage $F$ statistics from Stock and Yogo (2005), which implicitly divide the concentration parameter by the number of instruments. We suggest applied researchers calculate a new pre-test $\widetilde{F}$ (see equation ((ref))) and compare it to a cutoff of 4.14. If $\widetilde{F}$ is above the cutoff then the researcher can rely on the JIVE $t$-statistic with the caveats analogous to Stock and Yogo (2005): Namely, the size distortions of the JIVE $t$-statistic are within 5% tolerance level of the nominal size. If $\widetilde{F}$ is below the cutoff we suggest researchers report a confidence set obtained by inversion of our newly proposed weak-identification robust jackknife AR test (see Equation ((ref))). As discussed in Section (ref), applied researchers may also choose other cutoffs depending on their tolerance level of size distortions. Here we illustrate this recipe with an example from Angrist and Krueger (1991) (hereafter referred to as AK91).
AK91 provided a motivating example for the weak identification literature, starting with the seminal work by Bound et al. (1995). Staiger and Stock (1997) suggested that the relatively low value of the first stage $F$ statistic can be seen as a sign of potentially weak instruments in the AK91 application. Hansen et al. (2008) argued that many instruments may be a more relevant description of the identification issue encountered in AK91. They suggested that estimators other than the TSLS may restore the reliability of standard inferences. We resolve the controversy of whether the instruments are weak in this example utilizing a formal pre-test.
The original AK91 application estimated the effect of schooling ($X_{i}$) on log weekly wage ($Y_{i}$) using quarter of birth as instruments in a sample from the 1980 census of 329,509 men born in 1930-39. There are multiple specifications in the original AK91 study. We focus on the specification with 180 instruments and also on an extension of this specification using 1,530 instruments. The 180 instruments include 30 quarter and year of birth interactions (QOB-YOB) and 150 quarter and state of birth interactions (QOB-POB). For the second specification with 1,530 instruments, we also include full interactions among QOB-YOB-POB. Table (ref) reports the first stage $F$ statistics (FF), our proposed pre-test statistics $\widetilde{F}$, 5% and 2% confidence sets based on the JIVE $t$-statistic and the jackknife AR statistic proposed in this paper.
While the first stage $F$ statistic is below 10 and the current empirical practice would point towards weak identification for both specifications, the instruments turn out to be strong in both specifications based on our pre-test. According to the results of this paper discussed in Section (ref), both of the reported confidence sets based on a nominal 5% JIVE $t$-test are reliable with the same caveats as in Stock and Yogo (2005), namely, the actual rejection rate under the null hypothesis does not exceed 10%. This pre-test is based on the statistic $\widetilde{F}$ and rejects whenever $\widetilde{F} > 4.14$. Based on the pre-test, the empirical researcher may report the JIVE confidence set only, and not the identification-robust AR confidence set.
This two-step procedure is similar to that popularized by Stock and Yogo's (2005). Choosing between JIVE and AR confidence set to report based on the pretest in the first step guarantees that the reported confident set from this two-step procedure has an overall size of 15%. If the applied researcher prefers that the two-step procedure has an overall size of 5%, results in Section (ref) suggest using a higher cutoff of 9.98 for $\widetilde{F}$ and smaller nominal sizes to construct the confidence sets. Specifically, the applied researcher should choose between a nominal 98%-level JIVE confidence set and a nominal 98%-level AR confidence set. In this case, for the specification with 180 instruments, the applied researcher can still report the JIVE confidence set as the corresponding $\widetilde{F}$ is greater than 9.98. However, for the specification with 1530 instruments, the applied researcher needs to report the identification-robust AR confidence set instead as the corresponding $\widetilde{F}$ is less than 9.98.
An alternative to pre-test is to always report a robust confidence set, which would be the 5% jackknife AR confidence set in this case. We do see that the AR confidence sets are wider, yet still informative.
We study the linear IV regression with a scalar outcome $Y_i$, a potentially endogenous scalar regressor $X_i$ and a $K\times 1$ vector of instrumental variables $Z_i$:
for $ i=1,..., N.$ We denote $Y$ to be the $N\times 1$ vector of outcome and $X$ to be the $N\times 1$ vector of endogenous regressors. We collect the transpose of $Z_i$ in each row of $Z$, a $N\times K$ matrix of instruments. We denote $\Pi_i=\mathbb{E}[X_i|Z_i]$ and allow the instruments to affect the endogenous regressor in a non-linear way. All results in this paper hold conditionally on a realization of the instruments. Thus, we treat the instruments as fixed (non-random) and $\Pi_i$ as some constants. We collect $\Pi_i$ in $\Pi$, a $N\times 1$ vector. The mean-zero errors $(e_i,v_i)$ are independent across $i$ but not identically distributed and may be heteroscedastic. We assume without loss of generality that there are no controls included in our model as they may be partialled out.
Weak identification under small $K$ is studied extensively in the weak IV literature. For Gaussian homoscedastic errors $(e_i,v_i)$ and linear first stage ($\Pi_i=\pi'Z_i$), the strength of the instruments corresponds directly to the concentration parameter, $\frac{\pi^\prime Z'Z\pi}{\sigma_v^2}$ where $\sigma_v^2=Var(v_i)$. The concentration parameter equals the signal-to-noise ratio in the first-stage regression and is related to the bias of the TSLS estimator and the quality of Gaussian approximation for the TSLS $t$-statistic. For the general case with homoscedastic errors, Staiger and Stock (1997) introduced weak instrument-asymptotics in which one considers a sequence of models so that the concentration parameter converges to a constant as $N\to\infty$. Under this asymptotic embedding, neither a consistent estimator of $\beta$ nor a consistent test of the null hypothesis that $\beta$ equals some scalar exists, and the test based on the TSLS $t$-statistic severely over-rejects.
The magnitude of the concentration parameter is not a good indicator of identification strength when the number of instruments is large. Inspired by Bekker (1994), we model large $K$ by considering $K\to\infty$ as $N\to\infty$, with the only restriction that $K$ is at most a fraction of $N$. Under this many instrument-asymptotics, Theorem (ref) below shows that the re-scaled concentration parameter $\frac{\pi^\prime Z'Z\pi}{\sigma_v^2\sqrt{K}}$ provides a characterization of weak identification in terms of the consistency of tests.
The setting considered in Theorem (ref) is quite favorable: the first stage is linear, errors are Gaussian and homoscedastic with known covariance matrix. So the only unknown parameters are $\beta$ and $\pi$. Theorem (ref) states that even in this favorable setting there exists no test that consistently differentiates any $\beta^*$ from $\beta_0$ if the ratio of the concentration parameter to the square root of the number of instruments is bounded. Indeed, for any test $\psi$ we can find its guaranteed power $\mathbb{E}_{\beta^*,\pi}\psi$ by minimizing over the alternatives $(\beta^*,\pi)$ with bounded ratio of the concentration parameter over $\sqrt{K}$. We show that even in this favorable setting the test that achieves the maximum guaranteed power has guaranteed power strictly less than one asymptotically. With heteroscedasticity of unknown form, sufficient statistics of low dimensions are not known, making the setting even less favorable. Later we show that in a more general heteroscedastic model we can construct a robust test that becomes consistent when $\frac{\Pi^\prime \Pi}{\sqrt{K}}\to\infty$.
Theorem (ref) can also be used to characterize weak identification in terms of consistent estimation since it implies there exists no consistent estimator for $\beta$ when the ratio of the concentration parameter to $\sqrt{K}$ is bounded. Our result complements the literature on estimation with many instruments. Chao and Swanson (2005) show that with homoscedastic errors, when $K$ grows proportionally to the sample size the TSLS estimator is consistent only if the concentration parameter grows faster than the number of instruments $K$, while LIML and BTSLS estimators are consistent when the concentration parameter grows faster than $\sqrt{K}$. However, under heteroscedasticity, even when $\frac{\pi^\prime Z'Z\pi}{\sqrt{K}}\to\infty$, LIML and BTSLS become inconsistent, but JIVE is still consistent, according to Chao et al. (2012).
The proof of Theorem (ref) builds on several classical papers. Following the approach of Andrews et al. (2006), we first reduce the class of tests to those based on a sufficient statistic. Among these tests, the minimal power is achieved by a test invariant to rotations of the instruments. This observation allows us to further reduce our attention to invariant tests, which depend on the data only through its maximal invariant under rotations. Then we derive a limit experiment for $K\to\infty$ similar to that derived in Andrews and Stock (2007). In this limit experiment the minimax power is less than one. Finally we use the argument of M{\"u}eller (2011) to bound the desired asymptotic minimax power using the minimax power obtained in the limit experiment.
The goal of this section is to introduce a test robust to weak identification in the heteroscedastic IV model when the number of instruments, $K$, is large.
The existing weak IV literature proposes several weak identification-robust tests of the null hypothesis $H_0:\beta=\beta_0$, when $K$ is small. These tests have correct size when the identification is weak and become consistent when the identification is strong. One example is the AR test. Specifically, the IV model ((ref)) implies that under a given null hypothesis $H_0:\beta=\beta_0$, the exogeneity assumption holds $\mathbb{E}[Z'e(\beta_0)]=0$ for the implied error $e(\beta_0)=Y-\beta_0X$. Then under mild assumptions, the scaled sample analog $\frac{1}{\sqrt{N}}Z'e(\beta_0)\Rightarrow N(0,\Sigma)$ satisfies a $K$-dimensional CLT. The AR statistic is defined as $ \frac{1}{N}e(\beta_0)'Z\widehat\Sigma^{-1}Z'e(\beta_0) $, where $\widehat\Sigma$ is a consistent estimator of $Var\left(\frac{1}{\sqrt{N}}Z'e\right)$. The AR test rejects the null hypothesis when the AR statistic exceeds the $(1-\alpha)$ quantile of the $\chi^2_K$ distribution. The AR test has asymptotically correct size regardless of the value of the first stage coefficients $\Pi_i$ and is asymptotically consistent when an analog of the concentration parameter grows to infinity.
Generalizing the AR statistic to the large-$K$ setting is challenging for multiple reasons. Firstly, the covariance matrix $\Sigma$ has dimension $K\times K$. Its consistent estimation is problematic if not impossible under general heteroscedasticity. Secondly, the AR statistic under the null has an improperly centered limit distribution because $\chi_K^2$ has a very large mean. Thirdly, the $K$-dimensional CLT provides a poor approximation to the AR statistic when $K$ is large.
We propose an analog of the AR test that is heteroscedasticity-robust and weak identification-robust in the presence of a large number of instruments. Denote the projection matrix $P=Z(Z'Z)^{-1}Z'$. Our test rejects the null of $H_0:\beta=\beta_0$ when the jackknife AR statistic
exceeds the $(1-\alpha)$ quantile of the standard normal distribution. We defer the discussion of the estimator of the variance $\widehat\Phi$ to the next subsection.
To address the challenges with the existing AR statistic, the AR statistic we propose uses the default homoscedasticity-inspired weighting $(Z'Z)^{-1}$ in place of $\widehat\Sigma^{-1}$. With the $(Z'Z)^{-1}$ weighting, the existing AR statistic has a quadratic form $e(\beta_0)'Pe(\beta_0)$. However, this quadratic form is not centered at zero as it contains the term $\sum_{i=1}^NP_{ii}e_i^2$, and each summand has positive mean. We thus remove this term from the quadratic form. This re-centering can be referred to as leave-one-out or jackknife. In the context of consistent estimation under many instruments, this leave-one-out idea was introduced by Angrist et al. (1999) and fruitfully exploited in a number of papers including Hausman et al. (2012) and Chao et al. (2012). Recently, this idea has been used in Chao et al. (2014) and Crudu et al. (2020). In order to create a test of correct size based on our AR statistic, we use a CLT for quadratic forms proved in Chao et al. (2012) that is restated below.
The assumption $P_{ii}\leq \delta<1$ implies that $\frac{K}{N}=\frac{1}{N}\sum_{i=1}^NP_{ii}\leq \delta<1$. This assumption is often referred to as a balanced design assumption. In the case of group-dummies instruments, $P_{ii}$ is equal to the ratio of the size of the group that observation $i$ belongs to over $N$. Assumption (ref) can be checked for any specific design.
While Lemma (ref) requires $K\to\infty$, the Gaussian approximation may work well for smaller $K$ as well. For example, if $K$ is fixed and errors are homoscedastic, then $$ \frac{1}{\sqrt{K}\sqrt{\Phi}}\sum_{i=1}^N\sum_{j\neq i}P_{ij}\eta_i\eta_j\Rightarrow \frac{\chi^2_K-K}{\sqrt{2K}} ~~~\mbox{ as }~~~ N\to\infty. $$ We prove this statement in the Supplementary Appendix S4. While the limit here is not Gaussian it is very well approximated by a standard normal distribution even for relatively small $K$. The random variable $\frac{\chi^2_K-K}{\sqrt{2K}}$ exceeds the 95% quantile of the standard normal distribution at most 7% of the time for all $K$, and at most 6% of the time for $K>40$.
In order to conduct asymptotically valid inference based on the normal approximation in Lemma (ref), we need an estimator for the scale parameter $\Phi$, which is consistent under the null. One `naive' estimator that achieves this is $ \widehat\Phi_1=\frac{2}{K}\sum_{i=1}^N\sum_{j\neq i}P_{ij}^2e_i^2(\beta_0)e_j^2(\beta_0), $ which uses the square of the implied error as an estimator for the $i$-th error variance. Under the null when $e_i(\beta_0)=e_i$, the estimator $\widehat\Phi_1$ is consistent under relatively mild conditions. However, using $\widehat\Phi_1$ in a test would result in poor power. To see this, note that under an alternative value of the parameter $\beta=\beta_0+\Delta$, we can plug in the first stage and write the implied error $e_i(\beta_0)=Y_i-\beta_0X_i$ as the sum of a non-trivial mean $\Delta \Pi_i$ and a mean-zero random term $\eta_i = e_i+\Delta v_i$:
The AR statistics of a form similar to ((ref)) with $\widehat\Phi_1$ has been recently and independently proposed by Crudu et al. (2020). The aforementioned paper establishes robustness of their proposed test towards weak identification and heteroscedasticity in terms of size. While squaring $e_i(\beta_0)$ makes it an unbiased estimator for $Var(e_i)$ under the null, it is biased under the alternative when $\Delta \neq 0$. The bias in $\widehat\Phi_1$ grows at the same order as the fourth power of $\Delta$, which brings down the power of the test against distant alternatives. In Section (ref), we discuss the power implications of the `naive' estimator in more detail.
In order to remove the bias in $e_i^2(\beta_0)$ under the alternatives, one may residualize the implied error before squaring. However, this introduces a bias under the null. Denote $M=I-P$ and let $M_i$ be the $i$th row of $M$. Even under the null, the squared residualized error is biased $\mathbb{E}(M_ie)^2\neq Var(e_i)$. This is because the squared residual contains not only the squared error $e_i$ but also the square of regression estimation mistake. The latter can be large when the number of regressors $K$ is large.
This bias can be removed successfully using the cross-fit variance estimator suggested in Kline et al. (2020) and Newey and Robins (2018). Namely, they show that a product of the implied error and residual achieves both goals: it removes the linearly predictable part of the implied error and remains an unbiased estimator of the variance $$ \mathbb{E}\left[\frac{e_iM_ie}{M_{ii}}\right]=Var(e_i). $$
Our challenge is that the scale parameter $\Phi$ defined in Lemma (ref) is a quadratic form with a double summation. Residuals $M_ie(\beta_0)$ and $M_je(\beta_0)$ are correlated since they contain the same estimation mistake. One can show that $$ \mathbb{E}\left[e_iM_iee_jM_je\right]=(M_{ii}M_{jj}+M_{ij}^2)Var(e_i)Var(e_j). $$ Our proposed estimator of the scale parameter $\Phi$ re-weights each term in the summation to remove the bias described above:
We establish the consistency of $\widehat\Phi$ under the null and extend this result to local alternatives.
Theorem (ref) combined with Lemma (ref) implies that under the null $H_0:\beta=\beta_0$ our proposed AR statistic has an asymptotically standard normal distribution. Since no assumption about identification is made, the resulting AR test has asymptotically correct size regardless of the strength of identification.
Theorem (ref) establishes the consistency of the variance estimator when the null hypothesis does not hold. We use Theorem (ref) to derive local power curves of the AR test discussed in the next section. The variance estimator ((ref)) residualizes the implied errors $M_ie(\beta_0)$ to remove non-trivial mean of $e(\beta_0)$ under the alternative. The residualization is complete if the first stage is linear $\Pi_i=\pi'Z_i$. We do not impose such an assumption in Theorem (ref). Instead we require that the approximation of $\Pi_i$ by a linear combination of instruments improves with the number of instruments as measured by the norm of the approximation mistake, $\Pi'M\Pi$. In their Assumption 4, Chao et al. (2012) impose that $\frac{\Pi'M\Pi}{N}\to 0$, which may be weaker or stronger than our assumption $\Pi'M\Pi\leq\frac{C}{K}\Pi'\Pi$ depending on the identification strength. The variance estimation in Chao et al. (2012) is valid only under strong identification as it relies on the consistency of the JIVE estimator. The residuals from structural equation, with the JIVE estimate for $\beta$ plugged in, approximate the structural errors well. In contrast, our variance estimator remains valid under weak identification when no consistent estimator for $\beta$ exists. This is why we need stricter assumptions on the linear approximation to produce reliable residuals under weak identification.
Let us introduce a jackknife measure of the information contained in the instruments: $$ \mu^2=\sum_{i=1}^N\sum_{j\neq i}P_{ij}\Pi_i\Pi_j. $$ For the linear first stage $\Pi_i=\pi'Z_i$, we have $\mu^2=\pi'Z'Z\pi-\sum_{i=1}^NP_{ii}(\pi'Z_i)^2$. Assumption (ref) guarantees that $(1-\delta)\pi'Z'Z\pi\leq\mu^2\leq \pi'Z'Z\pi.$ Thus, the two measures $\frac{\mu^2}{\sqrt{K}}$ and $\frac{\pi'Z'Z\pi}{\sqrt{K}}$ are of the same order and increase to infinity or not simultaneously. In the general case where the instruments may affect the endogenous regressor in an arbitrarily non-linear way, the linear IV regression only uses the projection of $\Pi$ onto the linear space of the instruments. Thus the projection matrix appears naturally in our measure of identification strength. The parameter $\mu^2$ can be considered as a jackknife generalization of the parameter $\pi'Z'Z\pi$ to non-linear case.
Equation ((ref)) of Theorem (ref) characterizes the local power curves of the jackknife AR test. The power under the alternative $\beta=\beta_0+\Delta$ is a function of the distance $\Delta$ between the alternative $\beta$ and the null $\beta_0$, the number of instruments $K$, a measure of identification strength $\mu^2$ and the degree of uncertainty $\sqrt{\Phi}$. Our jackknife AR statistic can be negative, unlike the AR statistic from the small-$K$ case which is always non-negative. We reject the null when $AR(\beta_0)$ exceeds the $(1-\alpha)$ quantile of the standard normal distribution. Under the alternative $\beta=\beta_0+\Delta$, the AR statistics has a positive drift and produces non-trivial power for both positive and negative $\Delta$. The second statement of Theorem (ref) shows that the AR test consistently distinguishes $\beta$ from $\beta_0$ as long as $\frac{\mu^2}{\sqrt{K}\sqrt{\Phi}}\to\infty$.
Theorem (ref) implies that $\frac{\mu^2}{\sqrt{K}}\to\infty$ is a sufficient condition for the consistency of the jackknife AR test in a model with a linear first stage . This complements Theorem (ref) which implies that $\frac{\pi'Z'Z\pi}{\sqrt{K}}\to\infty$ is necessary for the consistency of any test. This condition has appeared before in Chao et al. (2012) as a sufficient condition for the consistency of the JIVE estimator and asymptotic validity and consistency of the JIVE $t$-test. The important difference between the proposed jackknife AR test and the JIVE $t$-test is that even under weak identification ($\frac{\pi'Z'Z\pi}{\sqrt{K}}\not\to\infty$), the former maintains asymptotically valid size, while the latter does not. It is worth noticing that the condition $\frac{\Pi'\Pi}{K}\to0$ imposed by Theorem (ref) is quite weak as it covers both weakly and strongly identified cases.
\paragraph{ Power implications of variance estimation.} While the leave-one-out AR test with our proposed cross-fit variance estimator is consistent against fixed alternatives when identification is strong, the same test with a `naive' variance estimator $\widehat{\Phi}_1$ is in general not consistent. The difference between the implied error $e_i(\beta_0)$ and $\eta_i$ as defined in equation ((ref)) results in that the difference between $\widehat{\Phi}_1$ and $\Phi$ is a fourth degree polynomial of $\Delta$. This makes the stochastic shift for the AR statistic with the naive variance estimator to stabilize at the finite level when $\Delta\to\pm\infty$: $$ \frac{\Delta^2\mu^2}{\sqrt{K}\sqrt{\widehat{\Phi}_1}}\approx \frac{\Delta^2\mu^2}{\sqrt{K}\sqrt{\Phi}}\sqrt{\frac{\Phi}{c\Delta^4+\Phi}}\to C_{\pm} \mbox{ as }\Delta\to\pm\infty, $$ while it increases unboundedly for the statistic with the cross-fit variance estimator. Here $c=\frac{2}{K}\sum_{i=1}^N\sum_{j\neq i}P_{ij}^2\Pi_i^2\Pi_j^2$.
Theoretical inconsistency of a test may or may not result in power differences of empirical relevance for commonly used significance levels. This depends partially on whether the level at which the stochastic shift stabilizes is above the typically used critical values. In very strongly identified cases where $\frac{\mu^2}{\sqrt{K}}$ is large, we may detect no significant difference between two statistics for alternatives with relatively small $\Delta$, implying a small power difference. While there is an increasingly large difference in realized values of statistics for large $\Delta$, such a difference might not translate to a power difference either since both tests would reject. For example, in the AK91 example discussed in Section (ref), using the `naive' variance estimator $\widehat{\Phi}_1$ would yield nearly identical jackknife AR confidence sets. For the specification that uses 1530 instruments, jackknife AR confidence sets based on the `naive' variance estimator are $[-0.048, 0.202]$ (5%) and $[-0.662, 0.224]$ (2%), which are very close to the ones based on $\widehat\Phi$ as reported in Table (ref).
We find larger power differences for moderately weak instruments under a sparse first stage. The divergence between two statistics depends positively on parameter $c$. While large values of the first stage coefficients $\Pi_i$ tend to produce large values of both $\frac{\mu^2}{\sqrt{K}}$ and $c$, the relation between the last two is not proportional. A more sparse first stage tends to produce higher values of $c$ (and larger power differences) for the same level of the identification strength $\frac{\mu^2}{\sqrt{K}}$, and therefore more stark power loss from using the `naive' variance estimator. Based on a simple simulation design, Figure (ref) plots the power curves for the leave-one-out AR test with different variance estimators under a sparse first stage (a) and a dense first stage (b). We include additional power comparisons in Section (ref) and in the Supplementary Appendix.
In a prominent paper, Stock and Yogo (2005) introduced a pre-test for weak identification that has gained enormous popularity in applied work. In homoscedastic IV models with small $K$, the concentration parameter fully characterizes the worst bias of the TSLS as a fraction of the OLS bias and the worst rejection rate of TSLS-Wald test. Stock and Yogo (2005) suggest a set of cut-offs for the first stage $F$ statistic, above which a researcher can guarantee with high (prespecified) probability that the bias of TSLS is not larger than 10% of the OLS bias, or that the TSLS-Wald statistic does not over-reject by more than 5%. The cut-offs depend on the goal (bias or size) and the number of instruments. However, these details seem to be mostly disregarded in empirical practice that uses a cut-off of 10, regardless of the goal or the number of instruments.
As with any procedure of such generality, the Stock-Yogo pre-test suffers from multiple drawbacks. First, the pre-test is valid only if the model is homoscedastic. Andrews (2018) shows that in models calibrated to commonly-used data sets with heteroscedasticity one may find cases with the first stage $F$ statistics exceeding 1000, that have large over-rejections of the TSLS-Wald test. Second, the TSLS estimator is less robust to weak identification when $K$ is large. In a homoscedastic model when $K$ is growing proportionally to the sample size, the TSLS estimator is consistent only if $\frac{\pi^\prime Z'Z\pi}{K}\to\infty$, while LIML and BTSLS estimators are consistent when $\frac{\pi^\prime Z'Z\pi}{\sqrt{K}}\to\infty$ as shown in Chao and Swanson (2005). In this case, the pre-test becomes too conservative. Indeed, if $\frac{\pi^\prime Z'Z\pi}{\sqrt{K}}\to\infty$ but $\frac{\pi^\prime Z'Z\pi}{K}\nrightarrow\infty$, then the pre-test most likely declares weak identification as the expectation of the first stage $F$ equals to $\frac{\pi^\prime Z'Z\pi}{K\sigma_v^2}+1$, even though there exist consistent estimators and a reasonable Wald-test can be constructed.
We propose a new pre-test for weak identification, that allows us to assess the reliability of the JIVE-Wald test. Our pre-test uses statistic
here $\widehat{\Upsilon}=\frac{2}{K}\sum_{i}\sum_{j\neq i}\frac{P_{ij}^2}{M_{ii}M_{jj}+M_{ij}^2}X_{i}M_{i}XX_{j}M_{j}X$ is an estimate of the variance $\Upsilon$ defined in ((ref)). The JIVE-Wald test uses the JIV2 estimator introduced in Angrist et al. (1999): $$ \widehat{\beta}_{JIVE}=\frac{\sum_{i=1}^{N}\sum_{j\neq i}P_{ij}Y_{i}X_{j}}{\sum_{i=1}^{N}\sum_{j\neq i}P_{ij}X_{i}X_{j}}. $$ We use the following estimator of the JIVE variance, that is a cross-fit version of the estimator derived in Chao et al. (2012): $$ \widehat{V}=\frac{\sum_{i=1}^{N}\left(\sum_{j\neq i}P_{ij}X_{j}\right)^2 \frac{\widehat{e}_{i}M_i\widehat{e}}{M_{ii}}+\sum_{i=1}^{N}\sum_{j\neq i}\widetilde{ P}_{ij}^{2}M_iX\widehat{e}_{i}M_jX\widehat{e}_{j}}{\left(\sum_{i=1}^{N}\sum_{j\neq i}P_{ij}X_{i}X_{j}\right)^{2}}, $$ where $\widehat{e}_{i}=Y_{i}-X_{i}\widehat{\beta}_{JIVE}$ and $\widetilde{ P}_{ij}^{2}= \frac{P_{ij}^2}{M_{ii}M_{jj}+M_{ij}^2}.$ The Wald statistic is defined as $Wald(\beta_{0})=\frac{\left(\widehat{\beta}_{JIVE}-\beta_{0}\right)^{2}}{\widehat{V}}.$ Our choice of JIVE is based on two considerations. First, according to Hausman et al. (2012), in a heteroscedastic IV model, when $\frac{\pi^\prime Z'Z\pi}{\sqrt{K}}\to\infty$, LIML and BTSLS become inconsistent, but JIVE is consistent. Second, the JIVE estimator is a ratio of two quadratic forms similar to the jackknife AR statistic, which motivates the following characterization.
Theorem (ref) shows that the distribution of the JIVE-Wald statistics can be quite different from its conventional $\chi^2_1$ limit when $\frac{\mu^{2}}{\sqrt{K}\sqrt{\Upsilon}}$ is small. If $\frac{\mu^{2}}{\sqrt{K}\sqrt{\Upsilon}}$ is large, then most realizations of the random variable $\nu$ are large as well and the limit of the JIVE-Wald is close to the distribution of $\xi^2$, which is $\chi^2_1$. This suggests that $\frac{\mu^{2}}{\sqrt{K}\sqrt{\Upsilon}}$ is a good measure for identification strength. The assumption $\frac{\Pi^\prime\Pi}{K^{2/3}} \to 0$ is somewhat restrictive but covers both weakly and strongly identified cases.
Using Theorem (ref) we create a pre-test for one definition of weak identification following Stock and Yogo (2005), which stipulates whether the actual size of the conventional 5% JIVE Wald test could exceed 10%. First, we calculate the worst asymptotic rejection rate of the JIVE-Wald test for a given theoretical strength of identification $S=\frac{\mu^2}{\sqrt{K}\sqrt{\Upsilon}}$: \[ R^{\max}_{\alpha}\left(S\right)=\max_{\varrho\in[-1,1]}\mathbb{P}_{S, \varrho}\left\{ \frac{\xi^{2}}{1-2\varrho\frac{\xi}{\nu}+\frac{\xi^{2}}{\nu^{2}}}\geq\chi_{1,1-\alpha}^{2}\right\}, \] where $\mathbb{P}_{S, \varrho}$ is the probability distribution of $(\xi,\nu)$ as described in Theorem (ref). The quantity $R^{\max}_{\alpha}$ can be straightforwardly obtained from simulations (the maximum rejection occurs at $\varrho=1$). Specifically $S=\frac{\mu^2}{\sqrt{K}\sqrt{\Upsilon}}>2.5 $ implies $R^{\max}_{5\%}\left(S\right)<10\%$.
The strength of identification parameter as measured by $S=\frac{\mu^2}{\sqrt{K}\sqrt{\Upsilon}}$ is unknown in practice. Theorem (ref) also allows us to construct a 5%-test for the null hypothesis that the unknown strength of identification parameter $S=\frac{\mu^2}{\sqrt{K}\sqrt{\Upsilon}} $ is lower than 2.5. This test is based on the statistic $\widetilde{F}$ and rejects whenever $\widetilde{F} > 4.14$. This test is therefore the analog to Stock and Yogo (2005) first stage $F$ pre-test, which tests whether the actual size of the conventional 5% JIVE Wald test could exceed 10%.
An advantage of the new pre-test based on $\widetilde{F}$ for weak identification is that when it is combined with any weak identification robust test, such as our jackknife AR test, to be used when $\widetilde{F}$ is below the cut-off, we can guarantee that the size of such two-step procedure is within a tolerance bound of 10% from the declared nominal size.
The attraction of the two-step procedure is that confidence sets based on the JIVE-Wald test are relatively easy to construct and are well understood by the practitioners. As we illustrate in simulations, the jackknife AR confidence sets tend to be wider than the JIVE-Wald confidence sets when identification is strong. Simulations also suggest the Bonferroni bounds derived in Corollary (ref) tend to be conservative, as the actual size of the two-step test does not exceed 7%.
The 5% Wald confidence set with 10% tolerance described in Corollary (ref) is the leading case considered by Stock and Yogo (2005). However, Theorem (ref) also allows us to create a two-step procedure with the overall size of 5% or 10% by adjusting the cut-off for $\widetilde{F}$ and using Wald and the jackknife AR confidence sets with smaller nominal sizes (and correspondingly larger critical values). Table (ref) tabulates a few combinations of valid cut-offs and critical values. As an example of a 5% two-step procedure, the researcher may compare the $\widetilde{F}$ statistic with 9.98. If $\widetilde{F}$ exceeds the cut-off, the researcher reports a JIVE-Wald confidence set that uses the 98% quantile of the $\chi^2_1$ as the critical value. Otherwise, the researcher reports a jackknife AR confidence set that uses the 98% quantile of the standard normal distribution as the critical value. We apply this procedure to the AK91 example discussed in Section (ref) and report the results in Table (ref) in italic.
In this section we conduct Monte Carlo simulations to show that the jackknife AR and the pre-test we develop are robust to many weak instruments unlike canonical IV estimators. To maintain the practical relevance, we attempt to preserve the structure of AK91 as described in Section (ref). Specifically, we adopt the simulation design by Angrist and Frandsen (2019). There is very little endogeneity in the original AK91, which makes it hard to study the biases of different estimators. Thus, we follow Angrist and Frandsen (2019) to introduce additional omitted variable bias to the simulated data. The simulated data has a nonlinear first stage and is heteroscedastic. We deviate from Angrist and Frandsen (2019) in two respects. First, we vary the sample size $N$ of the simulated data to be 1.5%, 1% and 0.5% of the original sample size. This is to vary the identification strength. We report the identification strength by $\frac{\mu^2}{\sqrt{K}\sqrt{\Upsilon}}$ as well as the average $\widetilde{F}$ across simulations. Simulations with sample size equal to 1.5% of the original sample size produce strong identification in our definition, 1% still produce strong identification but close to the weak identification region, while 0.5% produce weak identification.When we reduce the sample size we also need to exclude the instruments of the groups that are no longer populated. Second, both in data simulation and in estimation we do not include controls in order to isolate the implications of many instruments. The Appendix provides more details on our simulation design.
We evaluate the performance of common estimators and tests based on 1000 simulation draws. In Table (ref), we report the bias and size of Wald tests based on OLS, 2SLS, LIML and JIVE estimators. For the Wald test based on the LIML estimator, we calculate the standard errors as in Hansen et al. (2008). While Hansen et al. (2008) correct the canonical standard error estimator to be robust to many instruments, this test is not robust to heteroscedasticity as LIML itself is inconsistent under heteroscedasticity. For the Wald test based on the JIVE estimator, we calculate the heteroscedasticity-robust standard errors as described in Section (ref).
We find that due to many instruments 2SLS has large bias even under strong identification. While Hausman et al. (2012) show LIML is inconsistent under many instruments and heteroscedasticity, LIML is not too biased in our simulated data, as long as identification is not weak. We find that JIVE has low bias when identification is strong, but its bias increases when identification is weak. The Wald test based on either LIML or JIVE is not robust to many weak instruments, and we find substantial size distortion for LIML under weak identification. Surprisingly we do not find large size distortion for JIVE.
In Table (ref) we report the rejection frequency of the robust test we developed in this paper based on the jackknife AR test statistic. We find that the jackknife AR controls size even under weak identification. Our proposed pre-test also controls size and is able to switch to the JIVE-Wald test when identification is strong. In contrast, the first stage F statistics of Stock and Yogo (2005) (FF) are very small even under strong identification, which makes it not very informative.
In Table (ref) we compare the length of confidence intervals formed by inverting various tests. In particular, when identification is strong, jackknife AR confidence sets are longer (less efficient) but are not unreasonably long compared to the Wald tests based on LIML and JIVE. In this case, a pre-test can improve the efficiency by switching to the Wald test based on JIVE. As with the canonical AR test, the jackknife AR test can result in confidence intervals with infinite length. We report the probability of infinite length in the last column of Table (ref), and note that such probability increases as identification gets weaker.
To complement the discussion in Section (ref), we compare the performance of the jackknife AR test based on our proposed “cross-fit” variance estimator with that based on the “naive” variance estimator. Since power loss does not show up with strong identification, we further reduce the sample size to be 0.25% of the original size. In Table (ref) we confirm that the size is not affected by the choice of variance estimator. Figure (ref) demonstrates the difference in power for the jackknife AR tests with the cross-fit and the naive variance estimators. The “cross-fit” variance estimator performs slightly better in terms of power when identification is weak. As shown in the last two columns of Table (ref), the power difference is also reflected in fewer unbounded confidence intervals based on the jackknife AR test, and shorter confidence intervals when the bounded using the “cross-fit" variance estimator.
In this paper, we focus on identification for linear IV models with many instruments. In this environment, we characterize weak identification as a situation where an analog of the concentration parameter stays bounded relative to the square root of the number of instruments in large samples. We introduce a jackknifed version of the AR test that is robust to our definition of weak identification and heteroscedasticity. We also propose a pre-test for weak identification and correspondingly a two-step testing procedure in the spirit of Stock and Yogo (2005). Unlike the pre-test proposed by Stock and Yogo (2005), our two-step test controls size distortion even under heteroscedasticity and with many instruments. As an empirical example, our pre-test rejects weak identification in Angrist and Krueger (1991) where up to 1,530 instruments are used.
The data underlying this article are available in “Replication package for Inference with Many Weak Instruments", at \url{https://doi.org/10.5281/zenodo.5546157}.
Anatolyev, S., and Gospodinov, N. (2011). “Specification Testing in Models with Many Instruments." Econometric Theory 27, 427\textendash 441. \\ Andrews, I. (2018). “Valid Two-Step Identification-Robust Confidence Sets for GMM." The Review of Economics and Statistics, 100, 337\textendash348. Supplementary Appendix \\ Andrews, D.W.K., and Stock, J.H. (2007). “Testing with many weak instruments.” Journal of Econometrics 138, 24\textendash 46. \\ Andrews D, Moreira M, Stock J. (2006). “Optimal two-sided invariant similar tests of instrumental variables regression." Econometrica 74:715\textendash 752 \\ Angrist, J.D., and Frandsen B. (2019). “Machine Labor" NBER working paper 26584.\\ Angrist, J.D., Imbens, G.W., and Krueger, A.B. (1999). “Jackknife instrumental variables estimation.” \emph{Journal of Applied Econometrics} 14, 57\textendash67. \\ Angrist, J.D., and Krueger, A.B. (1991). “Does Compulsory School Attendance Affect Schooling and Earnings?" \emph{The Quarterly Journal of Economics} 106, 979\textendash1014.\\ Bekker, P.A. (1994). “Alternative Approximations to the Distributions of Instrumental Variable Estimators.” \emph{Econometrica} 62, 657\textendash681.\\ Bhuller, M., Dahl, G.B., Loken, K.V. and M. Mogstad (2020): “Incarceration, Recidivism, and Employment,” \emph{Journal of Political Economy}, 128(4), 1269\textendash1324.\\ Bound, J., Jaeger, D.A. and Baker, R.M. (1995) “Problems with Instrumental Variables Estimation when the Correlation between the Instruments and the Endogenous Explanatory Variable is Weak,” \emph{Journal of the American Statistical Association}, 90:430, 443\textendash450.\\ Chao, J.C., Hausman, J.A., Newey, W.K., Swanson, N.R., and Woutersen, T. (2014). “Testing overidentifying restrictions with many instruments and heteroskedasticity.” \emph{Journal of Econometrics} 178, 15\textendash21.\\ Chao, J.C., and Swanson, N.R. (2005). “Consistent Estimation with a Large Number of Weak Instruments." \emph{Econometrica} 73, 1673\textendash1692.\\ Chao, J.C., Swanson, N.R., Hausman, J.A., Newey, W.K., and Woutersen, T. (2012). “Asymptotic Distribution of JIV in a heteroscedastic IV Regression with Many Instruments.” \emph{Econometric Theory} 28, 42\textendash 86. \\ Crudu, F., Mellace, G., and Sandor, Z. (2020). “Inference in Instrumental Variables Models with Heteroskedasticity and Many Instruments.” \emph{Econometric Theory}, forthcoming.\\ Dobbie, W., Goldin, J., and Yang, C.S. (2018). “The Effects of Pretrial Detention on Conviction, Future Crime, and Employment: Evidence from Randomly Assigned Judges." \emph{American Economic Review} 108, 201\textendash240.\\ Fama, E.F., and MacBeth, J.D. (1973). “Risk, Return, and Equilibrium: Empirical Tests." \emph{Journal of Political Economy} 81, 607\textendash636.\\ Hansen, C., Hausman, J., and Newey, W. (2008). “Estimation With Many Instrumental Variables." \emph{Journal of Business & Economic Statistics} 26, 398\textendash 422.\\ Hausman, J.A., Newey, W.K., Woutersen, T., Chao, J.C., and Swanson, N.R. (2012). “Instrumental variable estimation with heteroscedasticity and many instruments.” \emph{Quantitative Economics} 3, 211\textendash 255. \\ Kleibergen F. (2002). “Pivotal statistics for testing structural parameters in instrumental variables regression.” \emph{Econometrica} 70:1781\textendash 1803. \\ Kline, P., Saggio, R., and S{\o}lvsten, M. (2020). “Leave-out estimation of variance components.” \emph{Econometrica} 88, 1859\textendash1898.\\ Maestas, N., Mullen, K.J., and Strand, A. (2013). “Does Disability Insurance Receipt Discourage Work? Using Examiner Assignment to Estimate Causal Effects of SSDI Receipt.” \emph{American Economic Review} 103, 1797\textendash1829.\\ M{\"u}ller, U. K. (2011). “Efficient Tests Under a Weak Convergence Assumption.” \emph{Econometrica} 79 (2): 395\textendash435. \\ Newey, W. (2004). “Many Instrument Asymptotics.” \\ Newey, W.K., and Robins, J.R. (2018). “Cross-Fitting and Fast Remainder Rates for Semiparametric Estimation.” \\ Newey, W.K., and Windmeijer, F. (2009). “Generalized Method of Moments With Many Weak Moment Conditions.” \emph{Econometrica} 77, 687\textendash 719. \\ Sampat, B., and Williams, H.L. (2019). “How Do Patents Affect Follow-On Innovation? Evidence from the Human Genome.” \emph{American Economic Review} 109, 203\textendash236.\\ Shanken, J. (1992). “On the Estimation of Beta-Pricing Models." \emph{The Review of Financial Studies} 5, 1\textendash33.\\ Staiger, D., and Stock, J.H. (1997). “Instrumental Variables Regression with Weak Instruments.” \emph{Econometrica} 65 (3): 557\textendash86. \\ Stock, J.H., and Yogo, M. (2005). “Testing for weak instruments in Linear IV regression. In Identification and Inference for Econometric Models: Essays in Honor of Thomas Rothenberg," pp. 80\textendash108.\\