EconBase
← Back to paper

A specification test for the strength of instrumental variables

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

46,389 characters · 12 sections · 42 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A specification test for the strength of instrumental variables

\def\spacingset#1{ {#1}} \spacingset{1}

\if00 \fi

\if10 {

center[center omitted — 35 chars of source]

} \fi

abstractThis paper develops a new specification test for the instrument weakness when the number of instruments $K_n$ is large with a magnitude comparable to the sample size $n$. The test relies on the fact that the difference between the two-stage least squares (2SLS) estimator and the ordinary least squares (OLS) estimator asymptotically disappears when there are many weak instruments, but otherwise converges to a non-zero limit. We establish the limiting distribution of the difference within the above two specifications, and introduce a delete-$d$ Jackknife procedure to consistently estimate the asymptotic variance/covariance of the difference. Monte Carlo experiments demonstrate the good performance of the test procedure for both cases of single and multiple endogenous variables. Additionally, we re-examine the analysis of returns to education data in angrist1991does using our proposed test. Both the simulation results and empirical analysis indicate the reliability of the test.

{\it Keywords:} weak instruments, many instruments, two stage least squares estimator, specification test

\spacingset{1.8}

Introduction

In some instrumental variable (IV) models, empirical researchers often face situations where the number of instruments is large with a magnitude comparable to the sample size arellano1991some, dagenais1997higher, bhuller2020incarceration. Instead of gaining efficiency as the conventional asymptotic theory indicates, the increasing number of instruments would deteriorate various IV estimators and test statistics bekker1994alternative, han2006gmm, anatolyev2011specification. This is the first strand of theoretical econometric literature on the IV estimator, the so-called “many instruments problem".

The second strand focuses on the statistical properties of the related estimation and inference methods, the so-called “weak instrument problem", where the instruments are weakly correlated with the endogenous variables. It is well understood that weakness in instruments will lead to bias and inconsistency of the IV estimators and test statistics bound1995problems,staiger1994instrumental. The key quantity here is the concentration parameter\footnote{When there are multiple endogenous variables, the key quantity is the concentration matrix, whose eigenvalues contain the strength of instruments. For simplicity of illustration, we use the concentration parameter as a unified term. } rothenberg1984approximating, which depicts the strength of the instrument set. The conventional asymptotic setting assumes that the concentration parameter grows at the same rate as the sample size. However, in the weak instrument asymptotics, this parameter is typically of smaller order than the sample size. For example, staiger1994instrumental proposed the local-to-zero framework to keep the concentration parameter roughly constant against the sample size and show that the classical IV estimators are not consistent. More generally, chao2005consistent introduced a unified framework to amalgamate the aforementioned two strands by modeling the concentration parameter with an arbitrary and divergent sequence. They show that the consistent estimation is still feasible when the concentration parameter is of higher order than $\sqrt{K_n}$. From the perspective of inference, mikusheva2021inference established that no asymptotically consistent test for the interested coefficients exists if the ratio of the concentration parameter over $\sqrt{K_n}$ stays bounded. Therefore, in applications involving many instruments, it raises a natural question of how to determine which range the concentration parameter is in.

Compared with developing estimation and inference methods that are robust to the many weak instruments problem, the literature on assessing the strength of instruments is exceptionally scarce. When the number of instruments is fixed, stock2002testing proposed to use a first-stage $F$ test for the single endogenous variable and the minimum eigenvalue of the Cragg-Donald statistics for multiple endogenous variables, which depict the magnitude of the concentration parameter. They have tabulated the non-standard critical values by simulation experiments, and a “rule of thumb" becomes a commonplace pre-test: reject the null of weak instruments when this $F$ statistic exceeds 10. However, the validity of the test remains unclear when the number of instruments is large, even though the authors have shown its reliability when the size of the instrument set grows at a much slower rate than the sample size ($K_n^4/n\rightarrow0$). Also, for the case of multiple endogenous variables, the tabulated critical values are known to be conservative as pointed out in stock2002testing. Building on their work, sanderson2016weak considered tests for the purposes of estimation and inference on one of multiple endogenous variables, but still keeping the number of instruments fixed. For the case of many instruments, hahn2002new proposed a test to examine the adequacy of the standard asymptotic result in IV regression models. They argue that if the test rejects the null, then weakness in instruments may arise. However, the test lacks power in detecting weak instruments hausman2005asymptotic. lee2012hahn proved that it is indeed a test for the exogeneity of the instruments. For the case of single endogenous variable and many instruments, mikusheva2021inference defined the instrument set to be weak when the concentration parameter in bounded in $\sqrt{K_n}$, and developed a pre-test for the weakness of the instruments in a spirit of the first-stage $F$ test.

In this paper, our contribution is to propose a specification test to distinguish between two different cases where the concentration parameter is at the same rate or a smaller rate of $K_n$ within the many instruments setup. Determining the order of the concentration parameter relative to $K_n$ is essential as the asymptotic behavior of the IV estimators and related inference methods heavily rely on it chao2006asymptotic,anatolyev2011specification. Our proposed test can also be viewed as a follow-up procedure to the test proposed in mikusheva2021inference: if their test concludes that the ratio of the concentration parameter over $\sqrt{K_n}$ is large, then one can apply our test to gain further insight into the employed instruments set by testing whether the concentration parameter is smaller than $K_n$ in order.

Unlike the tests in stock2002testing and mikusheva2021inference that can depict the magnitude of the concentration parameter, our test is in the same manner of the test proposed in hahn2002new, where they proposed a procedure to test for strong instruments by means of comparing forward and reverse 2SLS. Our test is based on the finding that the difference between the 2SLS and OLS estimators disappears under many weak instruments asymptotics but deviates from zero under many strong instruments asymptotics. As an important technical contribution, we establish the joint distribution of seven quadratic/bilinear forms that make up the 2SLS and OLS estimators. This complex joint distribution has its own interest: it can be used for the study of other related estimator such as the B2SLS estimator nagar1959bias in this context of many weak instruments. As the asymptotic covariance matrix has no explicit form, to implement our test we introduce a subsampling technique and propose a delete-$d$ Jackknife covariance matrix estimator. Both the theory and Monte Carlo experiments show that the test is robust to the number of endogenous variables and the type of error distributions.

The paper is organised as follows. In Section (ref) we introduce the model and assumptions, followed by a discussion of the order of the concentration parameter. In Section (ref), we first give the convergence results of the 2SLS and OLS estimators under many strong instruments and many weak instruments asymptotics in Section (ref). Then in Section (ref), we derive the limiting distributions of the difference of the 2SLS and OLS estimators under the above two specifications. In Section (ref) and (ref), we show that the delete-$d$ Jackknife is a valid estimation method for the asymptotic covariance matrix under the null and formalize the new test. Section (ref) reports the results of Monte Carlo simulations, followed by an empirical analysis in Section (ref) where we re-exaimine the results in angrist1991does. Section (ref) provides concluding remarks and further discussion.

Notations. Throughout the paper, for a vector $\mathbf{a}$, $\mathbf{a}'$ denotes its transpose. For an $a\times b$ matrix $\mathbf{A}$, we denote by $\mathbf{A}_i$ its $i$-th row, by $\mathbf{A}(j)$ its $j$-th column, and by $A_{ij}$ its $(i,j)$ entry, so that $\mathbf{A}=[\mathbf{A}_1',\mathbf{A}_2',\dots,\mathbf{A}_{a}']'=[\mathbf{A}(1),\mathbf{A}(2),\dots,\mathbf{A}(b)]$. In addition, $\|\cdot\|$ represents the Euclidean norm for a vector and the induced operator norm for a matrix. $\mathbf{I}_{a}$ and $\mathbf{0}_{a}$ are $a \times a$ identity and null matrices, with $\mathbf{0}_{a \times b}$ being $a \times b$ matrix of zeros; $\boldsymbol{J}^{ij}$ is the single-entry matrix with 1 at $(i, j)$ and zero elsewhere and $``\boldsymbol{j}^{i}$ is the single-entry vector with 1 at $i$-th entry and zero elsewhere. Finally, $``\stackrel{p}{\rightarrow} "$ denotes the convergence in probability and $``\stackrel{d}{\rightarrow}$ " the convergence in distribution.

Model and assumptions

Consider the following model:

equation[equation omitted — 71 chars of source]
equation[equation omitted — 69 chars of source]

where $\mathbf{y}$ and $\mathbf{Y}$ are, respectively, an $n\times 1$ vector and an $n \times p$ matrix of observations on the endogenous variables of the system, $\mathbf{Z}$ is an $n\times K_n$ matrix of observations on the $K_n$ instrumental variables, and $\mathbf{u}$ and $\mathbf{V}$ are, respectively, an $n\times 1$ vector and an $n \times p$ matrix of random disturbances. Within this setup, we specify $\boldsymbol{\Pi} \left(K_{n} \times p\right)$ to depend on $n$ in order to model the effects of having many weak instruments. In particular, the effect of many instruments can be examined by letting $K_{n} \rightarrow \infty$ as $n \rightarrow \infty$, while the effect of weak instruments can be accounted for by shrinking $\boldsymbol{\Pi}$ toward a zero matrix as $n$ grows. We also drop other potential exogenous variables without loss of generality. The following assumptions are used in the sequel.

assumption(a) As $n\rightarrow\infty$, $K_n/n \rightarrow\alpha \in (0,1)$; and (b) there exists a non-decreasing sequence of positive real numbers $\{s_n\}$ such that $s_n/n\rightarrow\kappa\in [0,\infty)$, and $ \boldsymbol{\Theta}_n=\boldsymbol{\Pi}'\mathbf{Z}'\mathbf{Z}\boldsymbol{\Pi}/s_n {\rightarrow} \boldsymbol{\Theta}$ almost surely, for some $p\times p$ nonrandom positive definite matrix $\boldsymbol{\Theta}$.
assumption(a) $\mathbf{Z}$ is independent of $\mathbf{u}$ and $\mathbf{V}$, and the $\mathbf{Z}_i$'s are independently and identically distributed (i.i.d.); (b) $(\mathbf{V}_i,u_i)'$ is i.i.d. with zero mean and the covariance matrix of $\boldsymbol{\Sigma}$ satisfying that $\mathrm{Cov}(\mathbf{V}_i)=\boldsymbol{\Sigma}_{VV}$, $\mathrm{Var}(u_i)=\sigma_u^2$, and $\mathrm{Cov}(\mathbf{V}_i,u_i)=\boldsymbol{\Sigma}_{Vu}$; and (c) there exist some positive constants $C_1<\infty$ and $C_2<\infty$ such that $E(u_i^4)<C_1$ and $E(V_{ij}^4)<C_2$.

Assumption (ref)(a) adopts the many instruments asymptotic framework. We are particularly interested in the case where $K_n$ and $n$ are of the same order. Given that the concentration matrix, $\boldsymbol{\Sigma}_{VV}^{-1/2}\boldsymbol{\Pi}'\mathbf{Z}'\mathbf{Z}\boldsymbol{\Pi}\boldsymbol{\Sigma}_{VV}^{-1/2}$, is a natural measure of the instrument strength, Assumption (ref)(b) models the strength of instruments by the order of the magnitude of $s_n$. We focus on the cases where $s_n$ grows no faster than $n$. In general, the slower the divergence of $s_n$ is, the weaker the instruments are. Assumption (ref)(a) assumes the instruments are valid and i.i.d., and Assumption (ref)(b) and (c) require the error terms to be homoscedastic and have the finite fourth moment.

A key parameter in the discussion is $s_n$, which measures the strength of the instrument set. The parameter plays an essential role in IV regressions: For example, the consistent estimation and the reliable inference of $\boldsymbol{\beta} $ crucially depend on this parameter. To further clarify the discussion, we propose the following taxonomy of instruments in terms of their strength:

itemize• The instrument set is “completely weak" if $\frac{s_n}{\sqrt{n}} \rightarrow \kappa_1 \in [0, \infty),$ as $n\rightarrow \infty$. • The instrument set is “moderately weak" if $\frac{s_n}{\sqrt{n}} \rightarrow \infty$, but $\frac{s_n}{n}\rightarrow 0$ as $n\rightarrow \infty$. • The instrument set is “strong" if $\frac{s_n}{n}\rightarrow \kappa_0 \in (0, \infty)$ as $n\rightarrow \infty$.

In case (a), no findings have been reported on the consistent estimation of $\boldsymbol{\beta}$ to date, and mikusheva2021inference (hereafter referred to as MS2022) show that a consistent test is absent as well. In case (b), chao2005consistent show that the 2SLS estimator is inconsistent but consistent estimation is still feasible by using estimators such as the B2SLS and LIML anderson1949estimation estimators. However, they have a slow convergence rate $\frac{s_n}{\sqrt{n}}$ within this specification chao2006asymptotic, hansen2008estimation. In this paper we call the set of instruments “weak" if $s_n=o(n)$ by integrating the cases (a) and (b): instruments are weak as a group and it refers to the many-weak-instrument asymptotics hausman2012instrumental, bekker2015jackknife. In case (c), $s_n$ has the order of $n$ and the concentration parameter grows as fast as the sample size. This case corresponds to the standard many-strong-instrument asymptotics: instruments are many and they are strong as a group donald2001choosing,anderson2010asymptotic,anatolyev2013instrumental. In this case, the 2SLS estimator is still inconsistent, and the B2SLS and the LIML estimators remain consistent with the convergence rate $\sqrt{n}$.

In empirical applications, it is of practical importance to distinguish between the case of $\frac{s_n}{n}\rightarrow 0$ and the case of $\frac{s_n}{n}\rightarrow \kappa_0>0$. Firstly, as mentioned above, the B2SLS, LIML and JIVE angrist1995split estimators have the normal convergence rate $\sqrt{n}$ in the latter case. However, in the former case, they have the slower convergence rate $\frac{s_n}{\sqrt{n}}$ with moderately weak instruments and become inconsistent with completely weak instruments. For example, when $s_nn^{\gamma+1/2}$ for a constant $0<\gamma<1/2$, the convergence rate is then $n^\gamma$, which can be slow (e.g. $\gamma=0.1$) even though consistency can still be achieved. Such slow convergence rate of the aforementioned estimators can lead to biased results in finite sample. Besides, some robust inference methods with many instruments require the instruments to be strong, i.e., they are developed under many-strong-instrument asymptotics. For example, lee2012hahn modified the overidentification $J$ test for testing the hypothesis of valid instruments and established the asymptotic normality the modified test statistic when $\frac{s_n}{n}\rightarrow \kappa_0$. Similar results can also be found in anatolyev2011specification. If the instruments are not strong, then one should be cautious about the usage of their proposed tests. Therefore, determining the choice of many-instrument aymptotics is crucial when applying such inference methods. We therefore consider testing the hypothesis

equation[equation omitted — 152 chars of source]

The new specification test

In this section, we introduce our testing procedure to test for the strength of the available instrument set.

Limits of 2SLS and OLS

To motivate the new specification test, we firstly study the limits of $\hat{\boldsymbol{\beta}}^{2SLS}=(\mathbf{Y}'\mathbf{P}_{Z}\mathbf{Y})^{-1}$\\$\mathbf{Y}'\mathbf{P}_{Z}\mathbf{y}$ and $\hat{\boldsymbol{\beta}}^{OLS}=(\mathbf{Y}'\mathbf{Y})^{-1}\mathbf{Y}'\mathbf{y}$ under many instruments asymptotics in the following.

theoremAssume that Assumptions (ref) and (ref) hold. Then, as $n\rightarrow\infty$, (a) under $H_1$, $\hat{\boldsymbol{\beta}}^{2SLS} \stackrel{p}{\rightarrow} \boldsymbol{\beta}+ (\kappa_0/\alpha\boldsymbol{\Theta}+\boldsymbol{\Sigma}_{VV})^{-1}\boldsymbol{\Sigma}_{Vu}$, $\hat{\boldsymbol{\beta}}^{OLS} \stackrel{p}{\rightarrow} \boldsymbol{\beta}+ (\kappa_0\boldsymbol{\Theta}+\boldsymbol{\Sigma}_{VV})^{-1}\boldsymbol{\Sigma}_{Vu}$; and (b) under $H_0$, both $\hat{\boldsymbol{\beta}}^{2SLS}$ and $\hat{\boldsymbol{\beta}}^{OLS}$ converge in probability to $ \boldsymbol{\beta}+ \boldsymbol{\Sigma}_{VV}^{-1}\boldsymbol{\Sigma}_{Vu}$.

It has been well documented that 2SLS suffers from the many instruments problem. Theorem (ref) provides a more quantitatively asymptotic analysis of the phenomenon. Part (a) shows that when the instruments are strong as a group, the asymptotic bias of the 2SLS and OLS estimators differs, and the bias of 2SLS is strictly smaller than the bias of OLS. As the number of instruments increases, the asymptotic bias of 2SLS gets more severe, which is also observed by bekker1994alternative. When $\alpha$ is close to zero (a fixed number of instruments), 2SLS will apply with little bias, and when $\alpha$ approaches one, 2SLS is biased towards OLS buse1992bias. Part (b) indicates that this difference disappears in the presence of many weak instruments. We also find that the asymptotic bias of two estimators in (b) becomes greater compared with those in (a), suggesting that weakness in instruments exacerbates the estimation bias in many instruments asymptotics.

Limiting distribution of the difference between 2SLS and OLS

Inspired by Theorem (ref), we adopt the ideas of the specification test approach of hausman1978specification involving the 2SLS and OLS estimators and see how far apart they are. If the difference between the two estimators is small, one should not reject the asymptotics adopted in the mode. If the difference is large, one should come to the opposite conclusion. We firstly derive the limiting distribution of the difference $\hat{\boldsymbol{\beta}}^{2SLS}-\hat{\boldsymbol{\beta}}^{OLS}$ under both many weak and many strong asymptotics in the following theorem,

theorem(a) Under $H_0$, Assumptions (ref) and (ref), as $n\rightarrow\infty$, \begin{equation} \sqrt{n}\left( \hat{\boldsymbol{\beta}}^{2SLS}- \hat{\boldsymbol{\beta}}^{OLS} \right)\stackrel{d}{\rightarrow}N_p(\boldsymbol{0}, \boldsymbol{\Sigma}_{0}), \end{equation} where $\boldsymbol{\Sigma}_{0}=-\boldsymbol{g}_3\boldsymbol{\Sigma}_{3}\boldsymbol{g}_3'+\frac{1}{\alpha^2}\boldsymbol{g}_3\boldsymbol{\Sigma}_{4}\boldsymbol{g}_3'$ with $\boldsymbol{g}_3$, $\boldsymbol{g}_4$, $ \boldsymbol{\Sigma}_{3}$ and $\boldsymbol{\Sigma}_{4}$ defined in ((ref)), ((ref)), ((ref)) and ((ref)), respectively. (b) Under $H_1$, Assumptions (ref) and (ref), as $n\rightarrow\infty$, \begin{equation} \sqrt{n}\left( \hat{\boldsymbol{\beta}}^{2SLS}- \hat{\boldsymbol{\beta}}^{OLS}-\boldsymbol{\Delta} \right)\stackrel{d}{\rightarrow}N_p(\boldsymbol{0},\boldsymbol{\Sigma}_{A}), \end{equation} where $\boldsymbol{\Delta}=(\kappa_0/\alpha\boldsymbol{\Theta}+\boldsymbol{\Sigma}_{VV})^{-1}\boldsymbol{\Sigma}_{Vu}-(\kappa_0\boldsymbol{\Theta}+\boldsymbol{\Sigma}_{VV})^{-1}\boldsymbol{\Sigma}_{Vu}+O_p(n^{-1/2})$ and $\boldsymbol{\Sigma}_{A}=\boldsymbol{h}_1\boldsymbol{\Sigma}_{1}\boldsymbol{h}_1'+\boldsymbol{h}_2\boldsymbol{\Sigma}_{2}\boldsymbol{h}_2'+\boldsymbol{h}_3\boldsymbol{\Sigma}_{3}\boldsymbol{h}_3'+\alpha\boldsymbol{h}_3\boldsymbol{\Sigma}_{3}\boldsymbol{h}_4'+\alpha\boldsymbol{h}_4\boldsymbol{\Sigma}_{3}\boldsymbol{h}_3'+\boldsymbol{h}_4\boldsymbol{\Sigma}_{4}\boldsymbol{h}_4'$ with $\boldsymbol{h}_1$, $\boldsymbol{h}_2$, $\boldsymbol{h}_3$, $\boldsymbol{h}_4$, $\boldsymbol{\Sigma}_{1}$, $\boldsymbol{\Sigma}_{2}$ defined in ((ref)), ((ref)), ((ref)), ((ref)), ((ref)) and ((ref)), respectively.

An important case in the empirical literature is $p=1$, that is a single endogenous variable, where we have the following result.

corollaryWhen $p=1$, under $H_0$, Assumptions (ref) and (ref), as $n\rightarrow\infty$, \begin{equation} \sqrt{n}\left( \hat{\beta}^{2SLS}- \hat{\beta}^{OLS} \right)\stackrel{d}{\rightarrow}N(0,\sigma^2), \end{equation} where $ \sigma^2=\frac{1-\alpha}{\alpha}(\frac{\sigma_{u}^2}{\sigma_{vv}^2}-\frac{\sigma_{vu}^2}{\sigma_{vv}^4})$ with $\sigma_{u}^2, \sigma_{vu}$ and $\sigma_{vv}^2$ the corresponding variance and covariance of the errors.

Note that in the case of a single endogenous variable and under the null, the asymptotic variance of the difference has an explicit form. However, consistently estimating $\sigma_{u}^2$ and $\sigma_{vu}$ is infeasible with completely weak instruments. Besides, the asymptotic covariance matrix in ((ref)) has no explicit form for multiple endogenous variables. To implement a test procedure, we therefore propose a consistent variance/covariance estimator based on the delete-$d$ jackknife variance/covariance estimation following a theory proposed in shao1989general.

Jackknife covariance estimation

Let $d=\lambda n$ be an integer for $0<\lambda<1$ and $r=n-d$. Define $\mathbf{S}_{r}$ to be the collection of subsets of $\{1, \cdots, n\}$ with size $r$. For $s=\left\{i_{1}, \ldots, i_{r}\right\} \in \mathbf{S}_{r}$, let $\hat{\boldsymbol{\theta}}_s=\hat{\boldsymbol{\beta}}_s^{2SLS}-\hat{\boldsymbol{\beta}}_s^{OLS}$ be the subsample estimation, where $\hat{\boldsymbol{\beta}}_s^{2SLS}$ and $\hat{\boldsymbol{\beta}}_s^{OLS}$ are the 2SLS and OLS estimates based on the corresponding subsample. The delete-$d$ jackknife estimator of $\boldsymbol{\Sigma}_0$ is then

equation[equation omitted — 320 chars of source]

where $N={n \choose d} $. To establish the consistency and asymptotic unbiasedness of $\widehat{\boldsymbol{\Sigma}}_0$, we impose the following regularity conditions:

assumption$ \mathrm{E}(\mathbf{Y}'\mathbf{P}_Z\mathbf{Y})^{-1}=O(n^{-1}).$

This assumption is natural as $\mathbf{Y}'\mathbf{P}_Z\mathbf{Y}$ has the stochastic order of $n$ under Assumptions (ref) and (ref).

theoremUnder $H_0$, Assumptions (ref), (ref) and (ref), as $n\rightarrow\infty$, $\widehat{\boldsymbol{\Sigma}}_0$ is a consistent and asymptotically unbiased estimator of $\boldsymbol{\Sigma}_0$.

By using ((ref)) with large $d$, we obtain a consistent and efficient variance estimator. This is achieved at the expense of a large number of computations. As $N$ increases extremely fast as $n$ increases, we further consider the jackknife-sampling variance estimator (JSVE) proposed in shao1989efficiency to reduce the computation burden. Define a simple random sample (without replacement) $S_{m}$ of size $m$ from $\mathbf{S}_{r}$. We then compute $\hat{\theta}_{s}$ for $s \in S_{m}$ and apply

equation[equation omitted — 293 chars of source]

as the variance/covariance estimator. As stated in shao1989efficiency, JSVE is still asymptotically unbiased and consistent and we illustrate the result in the following corollary:

corollaryUnder $H_0$, Assumptions (ref), (ref) and (ref), as $n\rightarrow\infty$, $n/m\rightarrow0$, $\widehat{\boldsymbol{\Sigma}}_0^S$ is a consistent and asymptotically unbiased estimator of $\boldsymbol{\Sigma}_0$.
remarkNote that the classical bootstrap cannot be a solution to the weak instrument problem since it cannot replicate the correlation between instruments and structural errors in bootstrap samples anderson2010asymptotic,wang2016bootstrap.

Test procedure

We now introduce our method to detect many weak instruments. The new specification testing procedure can be implemented with the following steps:

description\setlength\itemsep{0.2em} • Obtain $N$ subsamples by randomly deleting $d$ observations from the sample without replacement for $N$ times; • Compute $ \widehat{\boldsymbol{\Sigma}}_0^S$ based on the $N$ subsamples taken in Step 1; • Given $r=n-d$, compute $\boldsymbol{\theta}_s$ based on a random subsample of size $r$; and • Realize the test statistics $\boldsymbol{T}_n=\boldsymbol{\theta}_s'( \widehat{\boldsymbol{\Sigma}}_0^S)^{-1}\boldsymbol{\theta}_s$. Reject $H_0$ if $\boldsymbol{T}_n>\chi^2_c(p)$ at the significant level c.
remarkStep 3 ensures that the JSVE matches with the theoretical asymptotic variance. As for the choice of $\lambda$, wu1990asymptotic recommended the interval $[0.25,0.75]$. In our simulation results, we have taken $\lambda=0.45$ which is about the middle of this interval, regardless of the number of endogenous variables considered in various settings.

Monte Carlo simulations

Monte Carlo designs

The goal of this section is to evaluate the finite sample performance of the test procedure introduced in the previous section. We consider the following setup with

equation[equation omitted — 60 chars of source]
equation[equation omitted — 76 chars of source]

where $\boldsymbol{\beta}=\mathbf{1}$, $\boldsymbol{Z}_i \stackrel{i.i.d.}{\sim} N_{K_n}(\mathbf{0},\mathbf{I}_{K_n})$, for $i=1,\dots, n$. The number of instruments $K_n$ varies in $\{50,100, 200\}$ with corresponding sample size $n$ satisfying three different ratios $K_n/n=\frac{1-\lambda}{3},\frac{1-\lambda}{2}$ and $\frac{2(1-\lambda)}{3}$. $\lambda$ is taken to be 0.45 as mentioned in the previous section. To investigate the effects of the number of endogenous variables, we choose $p=1,2,$ and 3. For the error terms, we generate them following two distributions: (i) multivariate normal, $N_{p+1}(\mathbf{0},\boldsymbol{\Sigma})$; and (ii) multivariate t, $t_5(\mathbf{0},\boldsymbol{\Sigma})$, where we set

equation*[equation* omitted — 346 chars of source]

for $p=1$, $p=2$ and $p=3$, respectively. The degree of endogeneity, $\rho$, is $0.9$ (highly endogenous) or $0.5$ (mildly endogenous).

In this setup, the parameter $\boldsymbol{\Pi}$ measures the strength of the instruments. To generate data under the null, we adopt the local to zero asymptotics as follows:

equation[equation omitted — 90 chars of source]

with $\boldsymbol{C}=\{c_{ij}\}_{1\leq i\leq K_n, 1\leq j\leq p }\stackrel{i.i.d.}{\sim}N(0,1)$. It can be verified that under ((ref)), $s_n$ is at the same order of the magnitude of $pc_n^2K_n$. We conclude that

itemize$H_0\Leftrightarrow c_n=o(1) $; • $H_1 \Leftrightarrow c_n=O(1)$,

since $K_n$ and $n$ are of the same order. Thus, we set $c_n=0.1$ and $c_n=1$ to examine the sizes and powers of the proposed test, respectively. For the JSVE, we choose $m=n^{3/2}$. The Monte Carlo experiments are conducted using 1000 replications.

Simulation results

Table (ref) reports the empirical biases, and the root mean squared errors (RMSE) of the JSVE for each combination of $(K_n,n)$. We can see that the JSVE performs very well in these experiments.

Table (ref) reports the empirical sizes with 5% and 10% nominal sizes when $c_n=0.1$. Though for $p=3$ with small ratio $K_n/n$, the proposed test is slightly oversized due to the resampling error. For example, the empirical sizes are 9.3% and 15.8% when $K_n=200$ and errors are Gaussian. The test still successfully controls the sizes under almost all settings, irrespective of endogenous degrees, error distributions, combinations of $(K_n,n)$ and the number of endogenous variables. When there is a single endogenous variable (the leading empirical case), the test performs stably.

Table (ref) demonstrates the empirical powers of the proposed test when $c_n=1$. It has generally satisfactory power performance across the board, especially for normal errors. For a fixed ratio $K_n/n$, the empirical power increases with $K_n$ as expected. For a fixed $p$, the empirical power decreases as the ratio $K_n/n$ increases. For example, the test has powers of 85.8, 68.7 and 41.2 for three $K_n/n$ ratios when $\rho=0.9$, $p=1$, $K_n=50$ under normal errors at 95% significant level. This is because that $\boldsymbol{\Delta}$ in ((ref)) is generally negative in our simulation settings and increases in $\alpha$. When keeping all the settings fixed but the endogenous degree $\rho$, we observe that the test has less power with mildly endogenous variables ($\rho=0.5)$ compared to the same setting with highly endogenous variables ($\rho=0.9)$. For instance, the powers are 20% and 96.9% when $p=2$, $K_n=50$ and $K_n/n=(1-\lambda)/3$ under normal errors at 95% significant level, for $\rho=0.5$ and $\rho=0.9$, respectively. The reason lies in the fact that $\boldsymbol{\Delta}$ is generally decreasing in $\rho$, according to ((ref)). When the errors follow the multivariate-t distribution, the test has less power. Especially, the test has quite low power when $\rho=0.5$ and $p=1$ under student-t errors. However, with increasing concentration parameter, say $c_n=2$, the test shows greater power as expected even though the errors follow multivariate-t distribution.

table[table omitted — 768 chars of source]
table[table omitted — 3,971 chars of source]
table[table omitted — 4,105 chars of source]

An empirical illustration: Return to education

In this section, we re-analyse the returns to education data of angrist1991does (henceforth referred to as AK1991) using quarter of birth as an instrument for educational attainment. One of the setups in the original AK1991 application use up to 180 instruments that include 30 quarter and year of birth interactions and 150 quarter and state of birth interactions. Later it has been widely suggested that the setup suffers from a weak instrument problem (angrist1995split; bound1995problems). MS2022 appled their proposed $\widetilde{F}$ test and argued that the instrument set is not completely weak with the original full data.

As the original sample size (329,509) of the data is larger than usual for empirical research, we consider subsamples with 0.1% ($n=330$) and 0.5% ($n=1650$) of the original sample size, more in keeping with the typical empirical application. We examine the specification of 180 instruments and 1350 instruments, respectively, that extend the model by including the interactions among quarter and year and state of birth. We evaluate the performance of the OLS estimator, the 2SLS estimator, the $\widetilde{F}$ test and our proposed test based on 1000 randomly chosen subsamples and report the results in Table (ref). Within our specification with 0.1% subsamples ($n=330,K_n=180)$, we find that the OLS and 2SLS estimates are close in magnitude: the average of OLS is 0.0483 and the average of 2SLS is 0.0514. Result in column (d) is in favor of this finding where our proposed test shows that the two estimates are close significantly. The average $\widetilde{F}$ is 4.84, which is above the cut-off 4.14 suggested by MS2022. It provides evidence that 1%-scheme produces subsamples which are not completely weak. Therefore, our test helps provide more information and explain why the OLS and 2SLS estimators behave similarly: The instruments in 1% subsample are moderately weak, that is, $s_n/\sqrt{n} \rightarrow \infty$ but $s_n/n \rightarrow 0$.

For the 0.5% subsamples with $K_n=1350$ instruments and $n=1650$, the difference between OLS and 2SLS estimators is again negligible as shown in columns (a), (b) and (d). It therefore indicates the presence of weak instruments. The $\widetilde{F}$ test then can provide detailed categorization: it average value is 4.65, evidencing that the instruments are moderately weak. Though not as precise as the result obtained from the $\widetilde{F}$ test, our test is still informative to identify the strength of instruments.

table[table omitted — 797 chars of source]

Conclusion

We find that the limits of the 2SLS and OLS estimators coincide under the many weak instruments asymptotics but differ under the many strong instruments asymptotics. Building on this, we propose a test statistic to distinguish between these two specifications. The proposed test allows for multiple endogenous variables and shows robustness to general forms of non-normality in the error distribution.

Our proposed test can be applied to determine which asymptotic scheme should be used. We suggest that the many strong asymptotics are trustworthy if the 2SLS and OLS estimates are not significantly close. Estimation methods can be applied while achieving the optimal convergence rate, and so as inferential tools leaning on the assumption of many strong instruments.

Besides, our proposed test can be viewed as a follow-up test for the $\widetilde{F}$ test in MS2022. If empirical researchers draw conclusion that the large instrument set is not completely weak, then one can apply our proposed test to differentiate the case with strong instruments from the case with moderately weak instruments.