Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
38,245 characters · 11 sections · 30 citation commands
The First-stage F Test with Many Weak Instruments
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 {
\affil{{Department of Statistics and Actuarial Science, The University of Hong Kong} \\ {(e-mail: [email removed])}}
\affil{{Department of Statistics and Actuarial Science, The University of Hong Kong} \\ {(e-mail: [email removed])}}
\affil{{School of Data Science, Chinese University of Hong Kong (Shenzhen) \\ {(e-mail: [email removed])}} }
} \fi
\if10 {
} \fi
Keywords: weak instruments, many instruments, $F$ test, size distortions\\ JEL Classification numbers: C12, C26\\ Word Count: 5740
\spacingset{1.8}
The first-stage $F$ statistic introduced by stock2002testing is commonly used to detect weak instruments in empirical research. Evidence of its popularity can be found in American Economic Review, where 15 of 17 papers published between 2014 and 2018 reported at least one first-stage $F$ statistic andrews2019weak. However, this approach was originally developed for a fixed number of instrumental variables (IVs), and does not address the case of a large number of instruments, which is commonly encountered in practice angrist1991does, dobbie2018effects,bhuller2020incarceration.
Several studies have pointed out limitations of applying SY2005's $F$ test with many instruments. For example, hansen2008estimation demonstrated through empirical examples and simulations that a low $F$ statistic does not necessarily indicate weak instruments. More recently, mikusheva2021inference described that the classical $F$ test can mistakenly identify weak instruments mainly due to the insufficiency of the conventional measure for instrument strength, known as the concentration parameter. However, these studies only narratively discussed the unreliability of the $F$ test. The theoretical basis for not recommending the $F$ test in practice has yet to be established.
In this paper, we study the asymptotic behavior of the first-stage $F$ statistic within the many-instrument framework, where the number of instruments and the sample size go to infinity simultaneously and proportionally. We show that the more appropriate distribution of the $F$ statistic shifts to the normal distribution, instead of the conventional noncentral Chi-squared distribution. The inadequacy of the noncentral Chi-squared distribution provides poor finite sample approximations to the $F$ statistic with many instruments, leading to size distortion of the classical $F$ test. These size distortions occur regardless of the pretested IV estimator or Wald test and become increasingly severe as the number of instruments approaches the sample size.
Our second goal is to correct SY2005's two-step procedure to enhance the usability of the $F$ test with many instruments. Apart from the inadequacy of the noncentral Chi-squared distribution, SY2005's two-step procedure suffers from the insufficiency of the concentration parameter when measuring instrument strength. In the case of many instruments, chao2005consistent and MS2022 show that the appropriate measure is the re-scaled concentration parameter, which is the ratio of the concentration parameter over the square root of the number of instruments. In our asymptotic result, the re-scaled concentration parameter appears in the centering term of the $F$ statistic. Building on this, we propose a two-step procedure based on the $F$ statistic to detect many weak instruments that is analogous to that of MS2022. Our proposed statistic is directly derived from the classical $F$ statistic and follows the standard normal distribution, making it both conceptually familiar and straightforward to apply.
By identifying the deficiencies of the first-stage $F$ statistic with many instruments, this study contributes to the literature on discussing its limitations and implications for empirical analysis. In the case of a fixed number of instruments, lee2022valid and keane2023instrument focus on the performance of the IV t-test and show that using the rule-of-thumb $F > 10$ as a diagnostic cannot guarantee its well-controlled size and power. They further suggest that a higher threshold should be adopted in practice. Our study provides theoretical justification for the unreliability of the $F$ test to gauge instrument strength in many-instrument settings.
Additionally, this study contributes to the literature on measuring the strength of many instruments. hahn2002new proposed a test to examine the adequacy of the standard asymptotic result in IV regression models. They argue that if the test rejects their null, then weakness in instruments may arise. However, lee2012hahn proved that it is indeed a test for the exogeneity of the instruments. MS2022 and carrasco2022testing considered heteroscedastic models and proposed novel $F$-type tests for many weak instruments. Our study focuses on the original $F$ test statistic and makes corrections for the effects of many instruments.
The paper is organized as follows. In Section (ref), we introduce the model, followed by a discussion of the concentration parameter. In Section (ref), we show the unreliability of the first-stage $F$ test by proving its size distortions through the noncentral Chi-squared approximation. In Section (ref), we propose a two-step procedure using the first-stage $F$ statistic with many instruments. Section (ref) presents an analysis of the returns to education data in angrist1991does. Section (ref) concludes with some further discussions.
We consider the following model:
for $i=1,\dots,n$, where $y_i$ is a scalar outcome, $Y_i$ is a scalar endogenous variable, $\mathbf{Z}_i$ is a $K_n\times 1$ vector of instrument variables. Errors $(u_i,v_i)$ have zero mean, covariance $\sigma_{vu}$ and variances $\sigma_{uu}^2$ and $\sigma_{vv}^2$, respectively. We denote by $\mathbf{y}$, $\mathbf{Y}$, $\mathbf{u}$, and $\mathbf{v}$ the $n\times 1$ vectors that collect the corresponding scalars, $\mathbf{Z}=(\mathbf{Z}_1',\dots,\mathbf{Z}_n')'$ the $n\times K_n$ matrix of observations on the $K_n$ instrumental variables. Moreover, $\mathbf{P}_Z=\mathbf{Z}(\mathbf{Z}'\mathbf{Z})^{-1}\mathbf{Z}'$, and $\mathbf{M}_Z=\mathbf{I}_n-\mathbf{P}_Z$ are two projection matrices, and $\mathbf{D}_Z=\mathrm{diag}(P_{11},\dots,P_{nn})$ is the diagonal matrix containing the diagonal terms of $\mathbf{P}_Z$.
The behaviors of IV estimation and inference methods crucially depend on the magnitude of the concentration parameter,
which characterizes the strength of instruments. When $K_n\equiv K$, SY2005 demonstrated that a small value of $\mu_n^2$ indicates weak instruments. When $K_n \rightarrow \infty$, a more appropriate measure of the strength of instruments is ${\mu_n^2}/{\sqrt{K_n}}$ that leverages the effect of many instruments. chao2005consistent showed that the bias-corrected 2SLS (B2SLS) estimator nagar1959bias estimator, the limited information maximum likelihood (LIML) estimator anderson1949estimation and the jackknife instrumental variable estimator (JIVE) angrist1995split are consistent only when $\mu_n^2$ grows faster than $\sqrt{K_n}$. Wald-tests based on the above estimators therefore over-reject when $\mu_n^2/ \sqrt{K_n}$ is bounded. Furthermore, MS2022 showed that there exists no consistent test for testing $\beta=\beta_0$ when $\mu_n^2/\sqrt{K_n}$ stays bounded; for this reason, they defined the instruments to be weak if $\mu_n^2/ \sqrt{K_n}$ stays bounded. We therefore focus on the measure $\mu_n^2/ \sqrt{K_n}$, which characterizes instrument strength within the many instruments framework.
In this section, we first review SY2005's influential $F$ test for detecting weak instruments, and show that it has distorted sizes when detecting many weak instruments\footnote{Stock-Yogo also showed that the $F$ test remains valid when $K_n^4/n \rightarrow 0$. However, this condition in fact requires very small $K_n$. For example, when the sample size is large enough to reach 10000, the number of instruments should be much smaller than 10 to satisfy the asymptotic scheme. Therefore, this setting cannot cover practical situations where $K_n$ is in hundreds.}.
SY2005 defines instruments to be weak if the bias of IV estimators (e.g., the 2SLS-OLS relative bias) or rejection rate of IV-Wald tests (e.g., the 2SLS-Wald test) exceeds a predetermined tolerance level (e.g., 10%). They further showed that $\mu_n^2$ can fully determine both the level of the estimation bias and rejection rate. Therefore, in SY2005's first step, a theoretical value of $\mu_n^2=\mu_0^2$ that indicates weak instruments is obtained. In the second step, SY2005 proposed to use the first-stage $F$ statistic to test $H_0^{SY}:\mu_n^{2}\leq \mu_0^2$ and showed that:
when $K_n\equiv K$, where $\chi^2_{K}(\mu_0^2)$ denotes the non-central Chi-squared distribution with $K$ degrees of freedom and noncentrality parameter $\mu_0^2$. To summarize, SY2005's two-step testing procedure for weak instruments is formulated as follows:
We first examine the empirical sizes of the first-stage $F$ statistic using simulations. Let $n=1000$ and $K_n=5,300,500$ and 800. Consider $\beta=1$, $\boldsymbol{Z}_i \stackrel{i.i.d.}{\sim} N_{K_n}(\mathbf{0},\mathbf{I}_{K_n})$, $\{v_i\}_{i=1}^n$ are i.i.d. normal with $\sigma_{vv}^2=1$, and $\mu^2_0=5$ and 500.
The first two rows of Table (ref) report that the conventional $F$ test has correct sizes with a fixed number of instrument, but over-rejects $H_0^{SY}$ when the number of instruments becomes large, regardless of the magnitude of $\mu_0^2$. Moreover, the over-rejection phenomenon gets increasingly severe when $K_n$ gets close to $n$. For example, when $\mu_0^2=5$ and $K_n$ increases from 500 to 800, the empirical sizes increases from 12.6% to 23%, which both far exceed the nominal level 5%.
It is natural to expect that the distribution in ((ref)) can explain the size distortion phenomenon in Table (ref) after letting $K \rightarrow \infty$. However, the expectation for this sequential limit scheme (SEQ-L: $n\rightarrow\infty$, followed by $K_n\rightarrow\infty$) turns out to be incorrect. Specifically, after renormalizing the noncentral Chi-squared distribution, the SEQ-L will provide the CLT: $ \sqrt{K_n}\left(F-1-{\mu_n^2}/{K_n }\right)\stackrel{d}{\rightarrow}N(0,2)$ that leads to the following result:
Proposition (ref) shows that the SEQ-L predicts the classical $F$ test to have correct sizes with many instruments. Therefore, it fails to characterize the size distortion phenomena observed in Table (ref). Such inadequacy of the SEQ-L motivates us to study the asymptotic behaviour of the $F$ statistic under the simultaneous limit scheme (SIM-L), where $n$ and $K_n$ go to infinity simultaneously and proportionally. The following assumptions are used in the sequel.
Assumption (ref) is standard in the many IV literature which was initially introduced in bekker1994alternative. Assumption (ref) assumes the homoscedastic first-stage errors. Our results are established under homoscedasticity as we focus on the behaviour of the original $F$ statistic, which was developed in such context. Investigating the performance of the $F$ statistic under heteroscedastiticty is beyond the scope of this paper. We establish the limiting distribution of the first-stage $F$ statistic for a large $K_n$ in the following theorem.
The condition $\mu_n^2 =o(K_n)$ is relatively weak as it covers both the weakly identified and strongly identified cases. This condition is also made in MS2022. The asymptotic normality of the $F$ statistic, as shown in ((ref)), stands in stark contrast to the conventional Chi-squared distribution, which only holds for a fixed number of instruments. Applying Theorem (ref), the following corollary confirmed the size distortions of the classical $F$ test observed in Table (ref).
Corollary (ref) theoretically identifies the limitation of the first-stage $F$ test with many instruments due to the poor approximation using the noncentral Chi-squared distribution. Particularly, when dealing with asymptotically balanced instruments ($\omega=\alpha^2$) or mesokurtic first-stage errors, the classical $F$ test would be oversized. Moreover, the size distortions become more severe as $\alpha$ approaches 1.
Corollary (ref) shows that our result under the SIM-L successfully recognizes the size distortion phenomena. The ratio $\alpha$ plays a crucial role as it depicts the effect of the magnitude of $K_n$ that is invisible under the SEQ-L. This difference between the asymptotic behaviours of $F$ under the SIM-L and SEQ-L allows us to explain from a theoretical perspective the over-rejection phenomenon of the classical $F$ test when the number of instruments is relatively large. The last row in Table (ref) reports the predicted sizes from Theorem (ref), which aligns perfectly with the empirical counterpart in the first two rows.
Corollary (ref) also serves as a warning to researchers using the classical $F$ test to detect many weak instruments of the size distortion problem, no matter which IV estimator or IV-Wald test is pre-tested. For example, relying on the popular rule-of-thumb that compares $F$ and the cutoff of 10 can still fail to control the rejection rate of B2SLS-Wald test within 10%. Therefore, empirical researchers are warned not to use the classical $F$ test to detect many weak instruments.
To enhance the usability of the first-stage $F$ statistic, we present a new two-step procedure for many weak instruments. In the first step, we consider controlling the worst rejection rate of the B2SLS-Wald test as B2SLS is consistent in the homoscedasticity setting when $\mu_n^2/\sqrt{K_n}\rightarrow \infty$. In the second step, we propose a corrected $F$ test to assess the reliability of the B2SLS-Wald test.
We re-consider the behaviour of the B2SLS-Wald test statistic in Section 3.4 of SY2005:
where
with $\mathbf{P}_b=\mathbf{P}_Z-K_n/n\mathbf{I}_n$, $\hat{\mathbf{u}}=\mathbf{y}-\mathbf{Y}\hat{\boldsymbol{\beta}}_{B2SLS}$ and $\hat{\sigma}_{uu}^2$ being the B2SLS-residuals estimator. To test for $H_0:{\mu_n^2}/{\sqrt{K_n}}\leq C$, we propose a corrected $F$ test using statistic
where $C$ is a constant obtained in our first step that is formulated later. We establish the behaviour of the B2SLS-Wald statistic as follows:
Assumption (i) $n^{-1}\sum_{i=1}^n(P_{ii}-\alpha)^2 \rightarrow 0$ or (ii) i.i.d normal $\{(u_i,v_i)\}_{i=1,\dots,n}$ imposes conditions on instrument designs or errors, respectively. The former is known as the asymptotically balanced instruments design that is often imposed in the many IV literature, see, for example, hausman2012instrumental, anatolyev2011specification and wang2016bootstrap. We refer to anatolyev2017asymptotics on the detailed discussions on this assumption. Under either Assumption (i) which implies $\omega=\alpha^2$ or Assumption (ii) which provides analytic error moments, $W$ converges in distribution to a mixture of two normal random variables. This result is largely different from the standard Chi-squared distribution that holds under a fixed number of instruments. It indicates that $W$ will behave close to the Chi-squared distribution only when ${\mu_n^2}/{\sqrt{K_n}}$ is unbounded. However, if ${\mu_n^2}/{\sqrt{K_n}}$ is bounded, the Chi-squared distribution will produce poor finite sample approximations and lead to size distortions. It further confirms that ${\mu_n^2}/{\sqrt{K_n}}$ is an adequate indicator for the strength of many instruments. Therefore, based on ((ref)), we can control the worst rejection rate of the B2SLS-Wald test for a given tolerance level $T$:
Using simulations, a theoretical value of $\mu_n^2/\sqrt{K_n}$ that corresponds to the tolerance level $T$, denoted by $C$, can be determined. Consequently, the null hypothesis of many weak IVs can be formulated by $H_0:{\mu_n^2}/{\sqrt{K_n}}\leq C$, that can be tested using the $F_c$ statistic as follows:
Finally, we propose the following two-step procedure to detect many weak instruments based on the first-stage $F$ statistic:
One notable advantage of this two-step procedure is that the implementation of the first step is identical to that of Section 5 in MS2022, except for the different measures for the instrument strength, see discussions in Appendix. Therefore, one can directly obtain the upper bound $C$ without simulating the first-step using the relationship:
where $C_0$ is proposed to be 2.5 in MS2022. For example, if $K_n=100$ and $n=1000$ in practice, then $C=3.7$ and researcher can use the $F_c$ test to give a fast and reliable assessment of the instrument strength.
In this section, we re-analyse the returns to education data of angrist1991does (henceforth referred to as AK1991) using quarter of birth as an instrument for educational attainment, and construct confidence intervals for the strength of instruments. One of the specifications in the original AK1991 uses up to 180 instruments that include 30 quarter and year of birth interactions and 150 quarter and state of birth interactions. At the time of publication, the issue of weak instruments had received little attention. Later it has been widely suggested that the setup suffers from a weak instrument problem (angrist1995split; bound1995problems). MS2022 applied their proposed pre-test and argued the instrument set is strong with the original full data.
As the original sample size (329,509) is larger than usual for empirical research, we consider the sample size to be 0.1% ($n=330$), 0.2% ($n=660$), 0.5% ($n=1650$) and 1% ($n=3300$) of the original data, more in line with the typical empirical application. We examine the specification with 180 instruments and 1530 instruments that extend the model by including the interactions among quarter and year and state of birth. We evaluate the performance of the first-stage $F$ statistic, the $\widetilde{F}$ statistic and the $F_c$ statistic based on 1000 randomly chosen subsamples and report the results in Table (ref)\footnote{A normality check with the Shapiro-Wilk test shows that the first-stage errors are plausibly normal ($p=0.68$) so that our proposed method is applicable.}.
For the 0.1% subsample with 180 instruments, the average $F$ statistics is 1.53, which is far below the conventional cut-off of 10. However, the average $\widetilde{F}$ statistic is 4.55, which exceeds MS2022's cutoff of 2.5. It provides an evidence that 0.1%-scheme produces strong instruments subsamples. Our proposed $F_c$ turns out to be 2.65 ($C=5.2$ according to ((ref))), which also claims that the instrument set is strong. When the sample size increases to 660, the first-stage $F$ statistic is uninformative. While both our proposed $F_c$ test and the $\widetilde{F}$ test determine the instruments to be weak ($C=4.1$). The findings for the case of 1530 instruments are similar. In conclusion, our proposed method is informative to identify the strength of many instruments.
Empirical researchers often use a large number of instruments in practice. In this paper, we investigate the behaviour of the first-stage $F$ statistic with many instruments. We establish that the first-stage $F$ statistic is asymptotically standard normal after appropriate normalization and recentering, which contrasts with the conventional noncentral Chi-squared distribution. We show that SY2005’s $F$ test will lead to size distortions for detecting many weak instruments, no matter which IV estimator or IV-Wald test is pretested.
As a byproduct of our main theory, we propose a two-step procedure for many weak instruments based on the $F$-statistic. The proposed method is conceptually familiar and computationally simple. This suggests that researchers can still assess the strength of many instruments relying on the $F$ statistic after proper corrections.
For future directions, it would be interesting to study the asymptotic behaviour of olea2013robust's effective $F$ statistic under the many-instrument setting as it is robust to heteroscedasticity, autocorrelation, and clustering. We conjecture that, after proper recentering and renormalizations, it would be asymptotically normal, indicating that the effective $F$ test would also have size distortions with many instruments. To establish such theoretical justifications, new tools such as the joint CLT for several sesquilinear forms under non-i.i.d. settings are needed.
We first prove Theorem (ref), then prove Corollary (ref) and Proposition (ref) applying Theorem (ref). Finally, we prove Theorem (ref).
Denote $\boldsymbol{\Upsilon}=\mathbf{Z}\pi$, we first establish the following lemma:
To prove Theorem (ref), firstly note that
Besides,
as
We then can apply Delta method with Lemma (ref) and the function $f:\mathbb{R}^2\rightarrow\mathbb{R}$ satisfying that $$f\left(\frac{\mathbf{v}'\mathbf{v}}{n}, \frac{\mathbf{v}'\mathbf{P}_Z\mathbf{v}}{n}\right)=F.$$ We have $$\nabla f\left(\sigma_{vv}^2,\frac{K_n}{n}\sigma_{vv}^2\right)=\left(-\frac{n\sum_{i=1}\Upsilon_i^2+nK_n\sigma_{vv}^2}{K_n(n-K_n)\sigma_{vv}^4},\frac{n\sum_{i=1}\Upsilon_i^2+n^2\sigma_{vv}^2}{K_n(n-K_n)\sigma_{vv}^4}\right),$$ which yields with Lemma (ref) that $$\sqrt{n}(F-1-\frac{\mu_n^2}{K_n })\stackrel{d}{\rightarrow}N\left(0,\frac{(\omega-\alpha^2)\mathrm{E}(v_1^4)+(2\alpha-3\omega+\alpha^2)\sigma_{vv}^4}{\alpha^2(1-\alpha)^2\sigma_{vv}^4}\right).$$
By the noncentral Chi-squared distribution's normal approximation for large $K_n$, we verify that
so
Denote $$\sigma_F^2=\frac{(\omega-\alpha^2)\mathrm{E}(v_1^4)+(2\alpha-3\omega+\alpha^2)\sigma_{vv}^4}{\alpha^2(1-\alpha)^2\sigma_{vv}^4}.$$Then, using Theorem (ref),
Especially, when $\omega=\alpha^2$ or $v_i$ has zero excess kurtosis, then $\sigma_F^2=2/(\alpha(1-\alpha))$, so that
Similar to the proofs of Theorem (ref), we have
where the last convergence holds by the CLT under the SEQ-L: $\sqrt{K_n}(F-1-\mu_0^2/K_n)\stackrel{d}{\rightarrow}N(0,2)$.
Define
We first introduce two lemmas:
The B2SLS-Wald test statistic can be rewritten to:
where the denominator expands to
From Lemma (ref), this expansion further converges to $\sigma_1^2-2\sigma_{12}{Q_{Yu}}/{Q_{YY}}+\sigma_2^2Q_{Yu}^2/Q_{YY}^2$, so that