Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
480,170 characters · 33 sections · 193 citation commands
Wild Bootstrap Inference for Instrumental Variables Regressions with Weak and Few Clusters
{18pt} {10pt} \belowdisplayskip\abovedisplayskip {5pt} \abovedisplayshortskip \belowdisplayshortskip {8pt} \belowdisplayskip\abovedisplayskip {4pt} \linespread{1.6}
The instrument variable (IV) regression is one of the five most commonly used causal inference methods identified by Angrist-Pischke(2008), and it is often applied with clustered data. For example, Young(2021) analyzes 1,359 IV regressions in 31 papers published by the American Economic Association (AEA), out of which 24 papers account for clustering of observations. Three issues arise when running IV regressions with clustered data. First, the strength of IVs may be heterogeneous across clusters with a few clusters providing the main identification power. Indeed, Young(2021) finds that in the average paper of his AEA samples, with the removal of just one cluster or observation, the first-stage $F$ can decrease by 28%, and 38% of reported 0.05 significant two-stage least squares (TSLS) results can be rendered insignificant at that level. Second, the number of clusters is small in many IV applications. For instance, Acemoglu2011 cluster the standard errors at the country/polity level, resulting in 12-19 clusters, Glitz2020 cluster at the sectoral level with 16 sectors, and Rogall2021 clusters at the province (district) level with 11 provinces (30 districts), respectively. When the number of clusters is small, conventional cluster-robust inference procedures may be unreliable. Third, it is also possible that IVs are weak in all clusters, in which case researchers need to use weak-identification-robust inference methods Andrews-Stock-Sun(2019).
Motivated by these issues, in this paper we study the inference for IV regressions with a small, and thus, fixed number of clusters and weak within-cluster dependence. We denote clusters in which the parameters of the endogenous variables are strongly identified as strong IV clusters. First, we show that a wild bootstrap Wald test, with or without the cluster-robust variance estimator (CRVE), controls size asymptotically up to a small error, as long as there exists at least one strong IV cluster. Second, we show that the wild bootstrap tests have power against local alternatives at 10% and 5% significance levels when there are at least five and six strong IV clusters, respectively. Third, in the common case of testing a single restriction, we show the bootstrap Wald test with CRVE is more powerful than that without CRVE for distant local alternatives. Fourth, we develop the full-vector inference based on a wild bootstrap Anderson-Rubin(1949) test, which controls size asymptotically up to a small error regardless of instrument strength. Fifth, in the Online Supplement we show that for the common case with a single endogenous variable and a single IV, a wild bootstrap test based on the unstudentized Wald statistic (i.e., the one without CRVE) is asymptotically equivalent to a certain wild bootstrap AR test under both null and alternative, implying that in such a case it is fully robust to weak IV. Sixth, we show in the Online Supplement that bootstrapping weak-IV-robust tests other than the AR test (e.g., the Lagrange multiplier test or the conditional quasi-likelihood ratio test) controls asymptotic size when there is at least one strong IV cluster.
Our inference procedure is empirically relevant. First, we notice that besides the aforementioned examples, the numbers of clusters may also be rather small in studies that estimate the region-wise effects of certain intervention if the partition of clusters is at the state level. We illustrate the usefulness of our methods in Section (ref) by applying them to the well-known dataset of ADH2013 in the estimation of the effects of Chinese imports on local labor markets in three US Census Bureau-designated regions (South, Midwest, and West) with 11-16 clusters at the state level. Second, our bootstrap inference is flexible with respect to IV strength: the bootstrap Wald test allows for cluster-level heterogeneity in the first stage, while its AR counterpart is fully robust to weak IVs. Figure 1 reports the estimated first-stage coefficients for each cluster (state) in ADH2013's (ADH2013) dataset, which suggests that there exists substantial variation in the IV strength among states. Specifically, the first-stage coefficients of some states are quite large compared with the rest in the region, while some other states have coefficients that are rather close to zero, and thus, potentially subject to weak identification. Some states even have opposite signs for their first-stage coefficients. However, there is no existing proven-valid inference method for IV regressions with few clusters, where IVs may be weak in some clusters. Therefore, we believe our bootstrap methods enrich practitioners' toolbox by providing a reliable inference in this context. Third, different from the analytical inference based on the widely used heteroskedasticity and autocorrelation consistent (HAC) estimators, our approach is agnostic about the within-cluster (weak) dependence structure and thus avoids the use of tuning parameters to estimate the covariance matrix for dependent data.
The contributions in the present paper relate to several strands of literature. First, it is related to the literature on the cluster-robust inference.\footnote{See Cameron(2008), Conley2011inference, Imbens-Kolesar(2016), abadie2022should, Hagemann(2017), Hagemann(2019), Hagemann2019placebo, Hagemann2020inference, Mackinnon-Webb(2017), Djogbenou-Mackinnon-Nielsen(2019), Mackinnon-Nielsen-Webb(2019), Ferman2019inference, Hansen-Lee(2019), Menzel2021bootstrap, Mackinnon2021, among others, and mackinnon2022cluster for a recent survey.} Djogbenou-Mackinnon-Nielsen(2019), Mackinnon-Nielsen-Webb(2019), and Menzel2021bootstrap show the bootstrap validity under the asymptotic framework in which the number of clusters diverges to infinity.\footnote{We refer interested readers to mackinnon2022cluster for detailed discussions on this asymptotic framework and the alternative asymptotic framework that treats the number of clusters as fixed.} Ibragimov-Muller(2010), Ibragimov-Muller(2016), Bester-Conley-Hansen(2011), Canay-Romano-Shaikh(2017), Hagemann(2019), Hagemann2019placebo, Hagemann2020inference, Canay-Santos-Shaikh(2020), and Hwang(2020) consider an alternative asymptotic framework in which the number of clusters is treated as fixed, while the number of observations in each cluster is relatively large and the within cluster dependence is sufficiently weak. However, the inference methods proposed by BCH and Hwang(2020) require an (asymptotically) equal cluster-level sample size,\footnote{See Bester-Conley-Hansen(2011) and Hwang(2020) for details.} while those proposed by IM and CRS require strong identification in all clusters. In contrast, our bootstrap Wald tests are more flexible as it does not require an equal cluster size and only needs one strong IV cluster for size control and five to six for local power, thus allowing for substantial heterogeneity in IV strength. To our knowledge, there is no alternative method in the literature that remains valid in such a context. For full-vector inference, we further provide the bootstrap AR tests, which are fully robust to weak identification.
Second, Canay-Santos-Shaikh(2020) prove the validity of wild bootstrap with a few large clusters by innovatively connecting it with a randomization test with sign changes. Our results rely heavily on this technical breakthrough, but also complement Canay-Santos-Shaikh(2020) in the following aspects. First, Canay-Santos-Shaikh(2020) focus on the linear regression with exogenous regressors and then extend their analysis to a score bootstrap for the GMM estimator. Instead, we focus on extending the wild restricted efficient (WRE) cluster bootstrap, which is popular among empirical researchers, for general $k$-class IV estimators. Therefore, our procedure cannot be formulated as a score bootstrap in the GMM setting. Second, we study the local power for the Wald test both with and without CRVE. The power analysis of the Wald test with CRVE is new to the literature and technically involved because with a fixed number of clusters, the CRVE has a random limit. We rely on the Sherman–Morrison–Woodbury formula for the matrix inverse to analyze the behaviors of the test statistic and its wild bootstrap counterparts under local alternatives. Based on such a detailed asymptotic analysis, we establish the local power of the bootstrap Wald test studentized by CRVE, and show it is higher than that for the Wald test without CRVE studentization for distant local alternatives in the empirically prevalent case of testing a single restriction. Indeed, as the linear regression is a special case of the linear IV regression studied in this paper, we believe the same result holds for the local power of the Wald tests in Canay-Santos-Shaikh(2020). Third, different from its unstudentized counterpart, the local power of the Wald test with CRVE is established without the assumption that the cluster-level first-stage coefficients have the same sign, which may not hold in some empirical studies (e.g., Figure (ref)).
Third, our paper is related to the literature of WRE bootstrap for IV regressions.\footnote{See, for example, Davidson-Mackinnon(2008), Davidson-Mackinnon(2010), finlay2014, Finlay-Magnusson(2019), Roodman-Nielsen-MacKinnon-Webb(2019), and Mackinnon2021.} In particular, Davidson-Mackinnon(2008), Davidson-Mackinnon(2010) first proposed the restricted efficient (RE) and WRE bootstrap procedures for IV regression with homoskedastic and heteroskedastic errors, respectively. finlay2014, Finlay-Magnusson(2019) extended it to clustered data, and Roodman-Nielsen-MacKinnon-Webb(2019) incorporated it into the Stata package “boottest". Our key contribution to this strand of literature is to theoretically prove the validity of the (modified) WRE cluster bootstrap under the alternative asymptotic framework in which IVs may be weak for some clusters and the number of clusters is treated as fixed. In particular, we carefully re-design the first stage of the WRE procedure to allow for cluster-level heterogeneity of IV strength. Our theoretical analysis also covers the general $k$-class IV estimators, including TSLS, the limited information maximum likelihood (LIML) estimator or the modified LIML estimator proposed by Fuller(1977). In addition, we theoretically prove the power advantages of the bootstrap Wald test with CRVE compared with the one without. This result is new and justifies the proposal in the literature of using the bootstrap Wald test studentized by CRVE from the power perspective.
Furthermore, our paper is related to the literature on weak identification, in which various normal approximation-based inference approaches are available for nonhomoskedastic cases, among them Stock-Wright(2000), Kleibergen(2005), Andrews-Cheng(2012), Andrews(2016), Andrews-Mikusheva(2016), Andrews(2018), Moreira-Moreira(2019), and Andrews-Guggenberger(2019). As Andrews-Stock-Sun(2019) remark, an important question concerns the quality of the normal approximations with nonhomoskedastic observations. On the other hand, when implemented appropriately, alternative approaches such as bootstrap may substantially improve the inference for IV regressions.\footnote{See, for example, Davidson-Mackinnon(2008), Davidson-Mackinnon(2010), Moreira-Porter-Suarez(2009), Wang-Kaffo(2016), Kaffo-Wang(2017), Finlay-Magnusson(2019), and Young(2021), among others. In addition, we notice that tuvaandorj2021robust develops permutation versions of weak-IV-robust tests with (non-clustered) heteroskedastic errors.} We contribute to this literature by establishing formally that wild bootstrap-based AR tests control asymptotic size under both weak identification and a small number of clusters. However, contrary to the Wald tests, we find in simulations that the bootstrap AR test without CRVE has better power properties than the one studentized by CRVE. We formally establish the local power of the bootstrap AR test without CRVE. Furthermore, we show that the bootstrap validity for other weak-IV-robust tests requires at least one strong IV clusters (Sections (ref)-(ref) in the Online Supplement), which provides theoretical support to the simulation results in finlay2014, Finlay-Magnusson(2019) documenting that with clustered observations, bootstrapping the AR tests results in better finite sample size control than bootstrapping alternative robust tests.
Finally, we note that although empirical applications often involve settings with substantial first-stage heterogeneity such as those shown in Figure (ref), related econometric literature remains rather sparse. Abadie-Gu-Shen(2019) exploit such heterogeneity to improve the asymptotic mean squared error of IV estimators with independent and homoskedastic observations. Instead, we focus on developing inference methods that are robust to the first-stage heterogeneity for data with a small number of clusters, while allowing for (weak) within-cluster dependence and heteroskedasticity.
The remainder of this paper is organized as follows. Section (ref) presents the setup and the bootstrap procedures. Section (ref) presents main assumptions, while Sections (ref)-(ref) provide asymptotic results. Simulations in Section (ref) suggest that our procedures have outstanding finite sample size control and, in line with our theoretical analysis, the bootstrap Wald test with CRVE has power advantages compared with the other bootstrap tests. The empirical application is presented in Section (ref). We conclude in Section (ref).
Throughout the paper, we consider the setup of a linear IV regression with clustered data,
where the clusters are indexed by $j \in [J] = \{1, ..., J\}$ and units in the $j$-th cluster are indexed by $i \in I_{n,j} = \{1, ..., n_j \}$. In (ref), we denote $y_{i,j} \in \textbf{R}$, $X_{i,j} \in \textbf{R}^{d_x}$, and $W_{i,j} \in \textbf{R}^{d_w}$ as an outcome of interest, endogenous regressors, and exogenous regressors, respectively. Furthermore, we let $Z_{i,j} \in \textbf{R}^{d_z}$ denote the IVs for $X_{i,j}$. $\beta \in \textbf{R}^{d_x}$ and $\gamma \in \textbf{R}^{d_w}$ are unknown structural parameters.
We allow the parameter of interest $\beta$ to shift with respect to (w.r.t.) the sample size, which incorporates the analyses of size and local power in a concise manner: $\beta_n = \beta_0 + \mu_{\beta}/r_n$, where $\mu_{\beta} \in \textbf{R}^{d_x}$ is the local parameter and $r_n$ is the convergence rate of the score defined later. We let $\lambda^{\top}_{\beta} \beta_0 = \lambda_0$, $\lambda_{\beta} \in \textbf{R}^{d_x \times d_r}$, $\lambda_0 \in \textbf{R}^{d_r}$, and $d_r$ denotes the number of restrictions under the null hypothesis. Define $\mu = \lambda_{\beta}^{\top} \mu_{\beta}$. Then, the null and local alternative hypotheses studied in this paper can be written as
Throughout the paper, we consider the $k$-class IV estimators of the form:
where $\vec{Z} = [Z:W]$, $\vec{X} = [X:W]$, $Y$, $X$, $Z$, and $W$ are $n \times 1$, $n \times d_x$, $n \times d_z$, and $n \times d_w$-dimensional vectors and matrices formed by $y_{i,j}$, $X^{\top}_{i,j}$, $Z^{\top}_{i,j}$, and $W^{\top}_{i,j}$, respectively. We also denote $P_A = A(A^{\top} A)^{-1}A^{\top}$ and $M_A = I_n - P_A$, where $A$ is an $n$-dimensional matrix and $I_n$ is an $n$-dimensional identity matrix. Specifically, we focus on four $k$-class estimators: (1) the TSLS estimator, where $\hat{\kappa} = \hat{\kappa}_{tsls} = 1$, (2) the LIML estimator, where
(3) the FULL estimator, where $\hat{\kappa} = \hat{\kappa}_{full} = \hat{\kappa}_{liml} - C/(n-d_z-d_w)$ with some constant $C$, and (4) the bias-adjusted TSLS (BA) estimator proposed by Nagar(1959) and Rothenberg(1984), where $\hat{\kappa} = \hat{\kappa}_{ba} = n/(n-d_z+2)$. Theoretically, we show that these four estimators are asymptotically equivalent.\footnote{More specifically, we show in Lemma S.B.1 in the Online Supplement that under our asymptotic framework, for $L \in \{\text{liml},\text{full},\text{ba}\}$, $\hat{\beta}_{L} = \hat{\beta}_{tsls} + o_P(r_n^{-1})$ and $\hat{\beta}_{tsls} - \beta_n = O_P(r_n^{-1})$, where $r_n = O(n^{1/2})$.} However, in simulations, we find that FULL has the best finite sample performance in the overidentified case.
For the rest of the paper, we define $\widetilde{Z}_{i,j}$ as the residual of regressing $Z_{i,j}$ on $W_{i,j}$ using the entire sample, and that for any random vectors $U_{i,j}$ and $V_{i,j}$, $\widehat{Q}_{UV,j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}U_{i,j}V_{i,j}^\top$ and $\widehat{Q}_{UV} = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}U_{i,j}V_{i,j}^\top$.
For inference, we construct Wald statistics based on the $k$-class estimator $\hat{\beta}_L$ defined in ((ref)) with $\hat k = \hat k_L$ for $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$. When the $d_r \times d_r$ weighting matrix $\hat{A}_{r}$ is asymptotically deterministic in the sense of Assumption (ref) below (such as $\hat{A}_r = I_{d_r}$, the $d_r \times d_r$ identity matrix), we denote $T_n$ as the Wald statistic without CRVE and define it as
where $||u||_A = \sqrt{u^{\top}Au}$ for a generic vector $u$ and a weighting matrix $A$.
Further define $\hat{A}_{r,CR}$ as the inverse of the CRVE:
$\widehat{\Omega}_{CR} = n^{-1} \sum_{j \in [J]} \sum_{i \in I_{n,j}} \sum_{k \in I_{n,j} } \widetilde{Z}_{i,j} \widetilde{Z}_{k,j}^{\top} \hat{\varepsilon}_{i,j} \hat{\varepsilon}_{k,j}$, and $\hat{\varepsilon}_{i,j}$ is the unrestricted residual defined in ((ref)). The corresponding Wald statistic studentized by CRVE is denoted as
In the case of testing the value of a single endogenous variable $\beta = \beta_0$, the test statistics reduce to $T_n = |\hat{\beta}_L - \beta_0|$, and $T_{CR,n} = |(\hat{\beta}_L - \beta_0)/\widehat{V}^{1/2}|$, respectively, where $\widehat{V}$ denotes the usual cluster-robust variance estimator.
We reject the null hypothesis in ((ref)) at $\alpha$ significance level if the test statistics are greater than their corresponding critical values ($\hat c_n (1-\alpha)$ and $\hat c_{CR,n}(1-\alpha)$ defined below for $T_n$ and $T_{CR,n}$, respectively). We compute the critical values by a wild bootstrap procedure described below.
Several remarks are in order.
In this section, we describe a full-vector wild bootstrap procedure for Anderson-Rubin (AR) type weak-IV-robust statistics with or without CRVE. Recall that $\beta_n = \beta_0 + \mu_{\beta}/r_n$. Under the null, we have $\mu_\beta = 0$, or equivalently, $\beta_n =\beta_0$. First, define the AR statistic without CRVE as
where $\hat{A}_z$ is a $d_z \times d_z$ weighting matrix with an asymptotically deterministic limit as specified in Assumption (ref) below, $f_{i,j} = \widetilde{Z}_{i,j}\bar{\varepsilon}^r_{i,j}$, $\bar{\varepsilon}^r_{i,j} = y_{i,j} - X^{\top}_{i,j} \beta_0 - W_{i,j}^{\top} \bar{\gamma}^r $, and $\bar{\gamma}^r$ is the null-restricted ordinary least squares (OLS) estimator of $\gamma$: $$\bar{\gamma}^r = \left( \sum_{j \in [J]} \sum_{i \in I_{n,j}} W_{i,j} W_{i,j}^{\top} \right)^{-1}\sum_{j \in [J]} \sum_{i \in I_{n,j}} W_{i,j} (y_{i,j} - X^{\top}_{i,j} \beta_0).$$ Second, we also define the AR statistic with the (null-imposed) CRVE as
Our wild bootstrap procedure for the AR statistics is defined as follows.
We note that unlike the $T_{CR,n}$-based Wald test in Section (ref), there is no need to bootstrap the CRVE for the $AR_{CR,n}$ test even though $\hat{A}_{CR}$ also admits a random limit. This is because $\hat{A}_{CR}$, unlike $\hat{A}_{r,CR}$ defined in (ref), is invariant to the sign changes. Two remarks are in order.
In this section, we introduce the assumptions that will be used in our analysis of the wild bootstrap tests under a small number of clusters in Sections (ref)-(ref). Define $\widehat{Q} = \widehat{Q}_{\widetilde{Z}X}^{\top} \widehat{Q}_{\widetilde{Z}\widetilde{Z}}^{-1} \widehat{Q}_{\widetilde{Z}X}$ and $Q$ as the probability limits of $\widehat{Q}$.
We conclude this section with more concrete examples. Specifically, Assumptions (ref) and (ref) hold in Examples (ref) and (ref), and we explain why they may not hold in Examples (ref) and (ref). For Example (ref), Assumption (ref) holds while Assumption (ref) does not. Consequently, the wild bootstrap AR tests are still applicable while the Wald tests are not.
For the Wald test, we further assume the following assumption.
Several remarks are in order.
Assumption (ref) requires that the weighting matrix $\hat{A}_r$ in ((ref)) has a deterministic limit. It rules out the case that $\hat{A}_r$ equals the inverse of CRVE, which has a random limit under a small number of clusters. We will discuss the bootstrap Wald test with CRVE in Section (ref).
Two remarks are in order.
We next examine the power of the wild bootstrap test against local alternatives.
Several remarks are in order.
Now we consider a wild bootstrap test for the Wald statistic studentized by CRVE.
We require $J>d_r$ because otherwise CRVE and its bootstrap counterpart are not invertible. Theorem (ref) states that with at least one strong IV cluster, the $T_{CR,n}$-based wild bootstrap test controls size asymptotically up to a small error. Next, we turn to the local power.
The size control of the bootstrap Wald tests with or without CRVE relies on Assumption (ref)(i), which rules out overall weak identification in which all clusters are weak. In the case that the parameter of interest may be weakly identified in all clusters, we can consider the bootstrap AR tests for the full vector of $\beta_n$ as defined in Section (ref).
Theorem (ref) below shows that, in the general case with multiple IVs, the limiting null rejection probability of the $AR_{n}$-based bootstrap test does not exceed the nominal level $\alpha$, and that of the $AR_{CR,n}$ test does not exceed $\alpha$ by more than $1/2^{J-1}$ when $J>d_z$, irrespective of IV strength.
Next, to study the power of the $AR_n$-based bootstrap test against the local alternative, we let $\lambda_{\beta} = I_{d_x}$ for $\mathcal{H}_{1,n}$ in ((ref)) so that $\mu = \mu_{\beta}$ and impose the following condition.
If $d_x = d_z = 1$, Assumption (ref)(i) implies strong identification. When $d_z>1$ or $d_x >1$, Assumption (ref)(i) only rules out the case that $Q_{\widetilde{Z}X}$ is a zero matrix, while allowing it to be nonzero but not of full column rank. As noted below, this means the AR test can still have local power in some direction even without strong identification. Assumption (ref)(ii) is similar to Assumption (ref)(ii). In particular, it holds automatically if Assumption (ref)(i) holds and $d_x= d_z=1$ (i.e., single endogenous regressor and single IV).
In this section, we investigate the finite sample performance of the wild bootstrap tests and alternative inference methods. We consider four different data generating processes (DGPs) with four dependence structures, namely cluster fixed effect (DGP 1), network dependence (DGP 2), serial dependence (DGP 3), and spatial dependence (DGP 4). For all four DGPs, the number of bootstrap replications equals 399. For DGPs 1, 3, and 4, the number of Monte Carlo replications is 50,000. The second DGP requires implementing spectral clustering for each Monte Carlo replication, which is computationally heavy. For this reason, we set the number of replications as 10,000. The nominal level $\alpha$ is set at 10%. In addition, following the recommendation in Djogbenou-Mackinnon-Nielsen(2019) and mackinnon2022cluster, we conduct cluster-level demean for all DGPs. For the conciseness of the paper, we only report the results for DGPs 1 and 2 in the paper and relegate the results for DGPs 3 and 4 to Section (ref) in the Online Supplement. The overall patterns observed from DGPs 3 and 4 are very similar to those from DGPs 1 and 2.
DGP 1. The first DGP is similar to that in Section IV of Canay-Santos-Shaikh(2020) and extend theirs to the IV model. The data are generated as
for $i=1, ..., n$ and $j=1, ..., J$. We set the number of clusters $J \in \{6,9,12,20\}$. The total number of observations $n$ is equal to 500. To allow for unbalanced clusters, we follow Djogbenou-Mackinnon-Nielsen(2019) and Mackinnon2021 and set the cluster sizes as
and $n_J = n - \sum_{j \in [j]} n_j$. We let $r=4$ to generate substantial heterogeneity in cluster sizes.\footnote{$r=4$ corresponds to the highest value of the heterogeneity parameter considered in the simulations of Djogbenou-Mackinnon-Nielsen(2019) and mackinnon2022cluster.} For example, when $J=6$, the cluster sizes are $8,17,33,65,127,$ and $250$, respectively.\footnote{When $J=20$, the cluster sizes are $2,2,3,3,4,5,6,8,10,12,15,18,22,27,33,41,50,61,75$, and $103$, respectively. We use this setting to investigate the robustness of the inference methods under a scenario that is quite different from the fixed-$J$ asymptotic framework. In addition, when $J=20$ and $d_z=3$, we combine the two clusters with $n_j=2$ and combine the two clusters with $n_j=3$, so that IM and CRS can also be implemented.} In addition, $(\varepsilon_{i,j}, v_{i,j})$, $(a_{\varepsilon,j}, a_{v,j})$, $Z_{i,j}$, and $\sigma(Z_{i,j})$ are specified as follows:
where $I_{d_z}$ is a $d_z \times d_z$ identity matrix and $Z_{i,j,k}$ denotes the $k$-th element of $Z_{i,j}$.\footnote{As mentioned above, we project out the cluster fixed effects. Specifically, we can define $\widetilde{Z}_{i,j} = Z_{i,j} - \frac{1}{n_j}\sum_{i \in I_{n,j}}Z_{i,j}$. Note that $\frac{1}{\sqrt{n_j}}\sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}\sigma(Z_{i,j})(a_{\varepsilon,j}+\varepsilon_{i,j})$ is asymptotically normal conditional on $a_{\varepsilon,j}$, as $n_j$ diverges to infinity. This is because $\left(\frac{1}{\sqrt{n_j}}\sum_{i \in I_{n,j}}Z_{i,j}\left(\sum_{k=1}^{d_z}Z_{i,j,k}^2\right), \left[\frac{1}{\sqrt{n_j}}\sum_{i \in I_{n,j}}Z_{i,j}\right] \left[\frac{1}{n_j} \sum_{i \in I_{n,j}} \left(\sum_{k=1}^{d_z}Z_{i,j,k}^2\right)\right], \frac{1}{\sqrt{n_j}}\sum_{i \in I_{n,j}}Z_{i,j}\left(\sum_{k=1}^{d_z}Z_{i,j,k}^2\right)\varepsilon_{i,j}\right)$ are jointly asymptotically normal given $(a_{\varepsilon,j})_{j \in [J]}$. Then, we have
which is asymptotically normal given $(a_{\varepsilon,j})_{j \in [J]}$ as the cluster size diverges to infinity, so that Assumption (ref)(ii) is satisfied.} The number of IVs is set to be $d_z \in \{1,3\}$, and $\rho_{\varepsilon v} \in \{0.3, 0.5, 0.7\}$ corresponds to the degree of endogeneity. In addition, we introduce cluster heterogeneity to the $d_z \times 1$ first-stage coefficients $\Pi_{z,j}$ by letting $\Pi_{z,j} = (\Pi/2, ..., \Pi/2)^{\top}$ for $1\leq j \leq J/3$, $\Pi_{z,j} = (\Pi, ..., \Pi)^{\top}$ for $J/3 < j \leq 2J/3$, and $\Pi_{z,j} = (2\Pi, ..., 2\Pi)^{\top}$ for $2J/3 < j \leq J$, where $\Pi \in \{0.25, 0.375, 0.5\}$. Such a DGP satisfies our assumptions and allows for heterogeneity in both cluster sizes and IV strengths. The values of $\beta$ and $\gamma$ are both set to $1$. In Section (ref) in the Online Supplement, we further report the simulation results for the case where $\Pi$ remains the same across clusters, so that the cluster-level heterogeneity in identification strength originates solely from the heterogeneity in cluster size. We find that the overall patterns remain very similar to those reported in Section (ref) below.
DGP 2. We consider the linear-in-mean social interaction model detailed in Example (ref). Specifically, following L22, we generate the network $\mathcal A$ with $n=500$ nodes as $\mathcal A_{i,j} = 1\{||\eta_{i} - \eta_{j}||_2 \geq (7/(\pi n))^{1/2}\}$, where $\eta_i \stackrel{i.i.d.}{\sim} \text{Uniform}[0,1]^2$. Following BDF09, we generate data according to (ref) in which $\varepsilon_i \stackrel{i.i.d.}{\sim} \mathcal{N}(0,0.1)$,
where $\varepsilon_i' \stackrel{i.i.d.}{\sim}\mathcal{N}(0,1)$ and $(\varepsilon_i,\varepsilon_i',\eta_i)$ are independent. The observables $(y_i,X_i,W_i,Z_i)$ are as defined in Example (ref) with $d_z=2$. According the calibration study by BDF09, we set the parameters $(\alpha,\beta,\delta) = (0.7683,0.4666,0.1507)$. Because the IV strength scales with $\gamma \beta + \delta$, we study the performance of our method under various identification strength by setting this value to be $(0.11, 0.1498, 0.1896)$, which correspond to values $(-0.0872, -0.0019, 0.0834)$ for $\gamma$. For comparison, we note that BDF09 set $\gamma = 0.0834$. We follow the classification procedure proposed by L22 to obtain the clusters.
We set $n=500$ and let $L$ be $[5,10,20,30]$. The number of clusters ($J$) depends on $L'$ which varies across simulation replications. We report the average $J$ below.
We investigate the finite sample performance of the following inference methods.
Except for the methods AR-ASY, AR-B, and AR-B-S, all other methods require the estimation of parameter $\beta$. All four k-class estimators mentioned in the paper can be used. Here, we report the results regarding the TSLS estimator for simulations with one IV and report those regarding the FULL estimator for simulations with multiple IVs, as it is known that FULL has reduced finite sample bias than TSLS in the over-identified case.\footnote{The performance of LIML is very similar to FULL in our simulations. We therefore omit the LIML results but they are available upon request.}
DGP 1. Tables (ref)-(ref) report the null empirical rejection frequencies of the ten inference methods described above with $J=6$ or $12$. For succinctness, we report the results for $J=9$ or $20$ in Section (ref) of the Online Supplement. Also as mentioned above, the results of the Wald tests are based on TSLS when $d_z=1$ and based on FULL when $d_z=3$, respectively. Following the recommendation in the literature, we set the tuning parameter of FULL to be 1.\footnote{In this case FULL is best unbiased to a second order among $k$-class estimators under normal errors Rothenberg(1984).} Several observations are in order.
Furthermore, for power comparisons we focus on the inference methods that control size throughout the above simulations (namely, WRE, W-B, W-B-S, AR-ASY, AR-B, and AR-B-S). The true value of $\beta$ is set equal to 1 and we vary the value of $\beta_0-\beta$ from $-3$ to $3$. $\rho_{\varepsilon v}$ is set equal to $0.5$. Figures (ref) and (ref) show the power curves of the bootstrap Wald tests and the AR tests, respectively, with $d_z=1$ and $J=6$ or $12$. Figures (ref) and (ref) in the Online Supplement show the power curves with $d_z=1$ and $J=9$ or $20$. In addition, Figures (ref)--(ref) report those for $d_z=3$. We highlight several observations below.
DGP 2. Table (ref) reports the size for the network data. Note the identification strength increases with $\gamma$. Similar to DGP 1, we find that ASY and BCH does not control size when the number of clusters is small, while IM and CRS have large size distortions when the number of clusters is large so that within each cluster, the number of observations is not sufficiently large to render asymptotic normality. Instead, the wild bootstrap-based inference methods (WRE, W-B, W-B-S, AR-B, and AR-B-S) have good size control regardless of the number of clusters. In Section (ref), we further report the power results for DGP 2. Consistent with DGP 1 and other simulation designs, we find that our W-B-S has the best power, especially against distant alternatives.
In an influential study, ADH2013 analyze the effect of rising Chinese import competition on US local labor markets between 1990 and 2007, when the share of total US spending on Chinese goods increased substantially from 0.6% to 4.6%. The dataset of ADH2013 includes 722 commuting zones (CZs) that cover the entire mainland US. In this section, we further analyze the region-wise effects of such import exposure by applying IV regression with the proposed wild bootstrap procedures to three Census Bureau-designated regions: South, Midwest and West, with 16, 12, and 11 states, respectively, in each region.\footnote{The Northeast region is not included in the study because of the relatively small number of states ($9$) and small number of CZs in each state (e.g., Connecticut and Rhode Island have only 2 CZs).}
Following ADH2013, we consider the linear IV model
where the outcome variable $y_{i,j}$ denote the decadal change in average individual log weekly wage in a given CZ and the endogenous variable $X_{i,j}$ is the change in Chinese import exposure per worker in a CZ, where imports are apportioned to the CZ according to its share of national industry employment. The Bartik-type instrument $Z_{i,j}$ (e.g., see Goldsmith(2020)) proposed by ADH2013 is Chinese import growth in other high-income countries, where imports are apportioned to the CZs in the same way as that for $X_{i.j}$.\footnote{See Sections I.B and III.A in ADH2013 for a detailed definition of these variables.} As can be seen from Figure (ref), the predictive power of the IV varies considerably among states. The exogenous variables $W_{i,j}$ include the characteristic variables of CZs and decade specified in ADH2013 as well as state dummies. Our regressions are based on the CZ samples in each region, and the samples are clustered at the state level, following ADH2013. Besides the results for the full sample, we also report those for female and male samples separately. Tables (ref)--(ref) in the Online Supplement show the values of $\widehat{Q}_{\tilde{Z}W,j}$ for each cluster in the three regions, and all of them are close to zero. ((ref)) in Assumption (ref)(i) is thus plausible in this application.
The main result of the IV regression for the three regions is given in Table (ref), with the number of observations ($n$) and clusters ($J$) for each region. As we have only one endogenous variable and one IV, the TSLS estimator is used throughout this section. We further report the conventional 90% confidence intervals based on CRVE with the standard normal critical value (ASY) and the 90% bootstrap confidence sets (CSs) by inverting the corresponding $T_n$ (W-B), $T_{CR,n}$ (W-B-S), and $AR_n$ (AR-B) bootstrap tests with a 10% nominal level.\footnote{AR-B and AR-B-S are numerically equivalent with one instrument.} The computation of the bootstrap CSs is conducted over the parameter space $[-10,10]$ with a step size of 0.01, and the number of bootstrap draws is set at 2,000 for each step. We highlight the main findings below.
In many empirical applications of IV regressions with clustered data, the IVs may be weak for some clusters and the number of clusters may be small. In this paper, we provide valid inference methods in this context. For the Wald tests with and without CRVE, we extend the WRE cluster bootstrap procedure to allow for the setting of few clusters and cluster-level heterogeneity in IV strength. For the full-vector inference, we develop wild bootstrap AR tests that control size asymptotically irrespective of IV strength, and also show that the bootstrap validity of other weak-IV-robust tests requires at least one strong IV cluster. Furthermore, we establish the power properties of the Wald and AR tests. Finally, recent studies by hansen2022jackknife and mackinnon2022leverage suggest that asymptotic or bootstrap inference based on jackknife variance estimators can be more reliable than that based on the conventional CRVE for OLS models. Therefore, it may be interesting to consider extending jackknife inference to the current setting. We leave this line of investigation for future research.
\singlespacing
This section provides the pseudo-code for our recommended wild bootstrap inference procedure W-B-S. We emphasize that when constructing the bootstrapped sample $X^*_{i,j}(g)$ and $y^*_{i,j}(g)$ in Step 4.2 of Algorithm 1 below, we use the interacted instruments $\overline{Z}_{i,j}$ defined in Step 2 of the algorithm. However, we still use the original instruments $Z_{i,j}$ to estimate the estimator and construct the test statistic $T_{CR,n}$.
\IncMargin{-1em}
\DecMargin{-1em}
Tables (ref)--(ref) report the values of $\widehat{Q}_{\tilde{Z}W,j}$ for each cluster in the South, Midwest, and West regions, and all of them are close to zero. For example, in Table (ref), the column $``W_1"$ reports the values of $\widehat{Q}_{\tilde{Z}W_1,j}$, where $j=1, ..., 16$ for the South region. Therefore, ((ref)) in Assumption (ref)(i) is plausible in this application.
We define $\hat{\beta}_L$ as the $k$-class estimator with $\hat{\kappa}_L$ for $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$. Their null-restricted and bootstrap counterparts are denoted as $\hat{\beta}_L^r$ and $\hat{\beta}_{L,g}^*$, respectively, and $\hat{\gamma}_L$, $\hat{\gamma}_L^r$, and $\hat{\gamma}_{L,g}^*$ are similarly defined. In the following, we show that $\hat{\beta}_{tsls}$, $\hat{\beta}_{liml}$, $\hat{\beta}_{full}$, $\hat{\beta}_{ba}$ are asymptotically equivalent, and so be their null-restricted and bootstrap counterparts.
Proof. First, $L \in \{\text{liml},\text{full},\text{ba}\}$, we have $ \hat{\mu}_L = \hat{\kappa}_L-1$ and
Then, by applying the Frisch–Waugh–Lovell Theorem we obtain that
By construction, $\hat{\beta}_{tsls}$ corresponds to $\hat{\mu}_{tsls} = 0$. For the LIML estimator, we have
which implies that
We note that
In addition, let $\hat{\varepsilon}_{i,j}$ be the residual from the full sample projection of $\varepsilon_{i,j}$ on $W_{i,j}$. Then, we have
where $c$ is a positive constant, the first inequality is by the definition that $\dot{\varepsilon}_{i,j}$ is the residual from the cluster-level projection of $\varepsilon_{i,j}$ on $W_{i,j}$, and the second inequality is by Assumption (ref)(iii). Combining (ref)--(ref), we have $\hat{\mu}_{liml} = O_P(r_n^{-2})$. In addition, we have $\frac{1}{n} X^{\top} M_{\vec{Z}} X = O_P(1)$, $\frac{1}{n} X^{\top} M_{\vec{Z}} Y = O_P(1)$, and
where the last equality holds by Assumptions (ref)(ii) and (ref)(i). This means
In addition, we note $\hat{\mu}_{full} = \hat{\mu}_{liml} - \frac{C}{n-d_z-d_w} = O_P(r_n^{-2})$ and $\hat{\mu}_{ba} = O(n^{-1})$, respectively. Therefore, we have the same results for $(\hat{\beta}_{full},\hat{\beta}_{ba})$.
For the second statement in the lemma, we have
where the last equality holds because $\widehat{Q}_{\widetilde{Z}\varepsilon} = \frac{1}{n}\sum_{j \in [J]} \sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}\varepsilon_{i,j} = O_P(r_n^{-1})$.
Next, we turn to the third statement in the lemma. We note that, for $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$ and $\lambda_{\beta}^{\top}\hat{\beta}^r_L = \lambda_0$,
As $\hat{\mu}_{L} = O_P(r_n^{-2})$ and $\hat{\beta}_L = \hat{\beta}_{tsls}+o_P(r_n^{-1})$ for $L \in \{\text{liml},\text{full},\text{ba}\}$, we have
For the last statement in the lemma, we note that
where the last equality holds because $\lambda_{\beta}^{\top}\beta_n - \lambda_0= \mu r_n^{-1}$ by construction and $\hat{\beta}_{tsls} - \beta_n = O_P(r_n^{-1})$. $\blacksquare$
Proof. By the same argument in the proof of Lemma (ref), for $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$ and $g \in \textbf{G}$, we have
such that $\hat{\mu}_{tsls,g}^* = 0$, $\hat{\mu}_{full,g}^* = \hat{\mu}_{liml,g}^* - \frac{C}{n-d_z-d_x}$, $\hat{\mu}_{ba,g}^* = \hat{\mu}_{ba}$, and $$\hat{\mu}^*_{liml,g}=\min_{r} r^{\top} \vec{Y}^{*\top}(g) M_W Z(Z^{\top} M_W Z)^{-1} Z^{\top} M_W \vec{Y}^*(g) r / (r^{\top} \vec{Y}^{*\top}(g) M_{\vec{Z}}\vec{Y}^*(g)r),$$ where $\vec{Y}^{*}(g) = [Y^*(g) : X^*(g)]$, $Y^*(g)$ is an $n \times 1$ vector formed by $Y_{i,j}^*(g)$, and $r=(1,-\beta^{\top} )^{\top}$.
Following the same argument previously, we have
where $\varepsilon_g^{*r}$ is an $n\times 1$ vector formed by $g_j\hat{\varepsilon}_{i,j}^r$. We first note that
where the last line is by Assumptions (ref)(ii), Lemma (ref), and the assumption that $\widehat{Q}_{\widetilde{Z}W,j}(\hat{\gamma}_L^r-\gamma) = o_P(r_n^{-1})$. Second, we have
where the last equality is by $\frac{1}{n_j} \sum_{i \in I_{n,j}}\hat{\varepsilon}^{r}_{i,j} \widetilde{Z}_{i,j} = O_P(r_n^{-1})$, as in ((ref)). We further note that
where $\hat{\beta}_L^r - \beta_n = O_P(r_n^{-1})$ and
Therefore, we have
where $\mathscr{E}_{i,j} = \varepsilon_{i,j} - W_{i,j}^\top\widehat{Q}_{WW}^{-1} \widehat{Q}_{W\varepsilon}.$ Similarly, we can show that
which, combined with (ref), implies
where $\hat{\theta}_g = \widehat{Q}_{WW}^{-1}\left[\frac{1}{n} \sum_{j \in [J]}\sum_{i \in I_{n,j}}(g_j\mathscr{E}_{i,j} W_{i,j})\right]$ and the second equality holds because
Recall that $\dot{\varepsilon}_{i,j}$ is the residual from the cluster-level projection of $\varepsilon_{i,j}$ on $W_{i,j}$, which means there exists a vector $\hat{\theta}_j $ such that $\varepsilon_{i,j} = \dot{\varepsilon}_{i,j} + W_{i,j}^\top \hat{\theta}_j$ and $\sum_{i \in I_{n,j}}\dot{\varepsilon}_{i,j}W_{i,j} = 0$. Then, we have
where the last inequality is by Assumption (ref)(iii). This implies
Combining (ref), (ref), and (ref), we obtain that $\hat{\mu}_{liml,g}^* = O_P(r_n^{-2})$, and thus, $\hat{\mu}^*_{full,g} = O_P(r_n^{-2})$. It is also obvious that $\hat{\mu}^*_{ba,g} = O_P(n^{-1})$. Given that $\hat{\mu}^*_{L,g} = O_P(r_n^{-2} + n^{-1})$, to establish $\hat{\beta}_{L,g}^* = \hat{\beta}_{tsls,g}^* + o_P(r_n^{-1})$, it suffices to show $\frac{1}{n}X^{*\top}(g) M_{\vec{Z}} X^*(g) = O_P(1)$, $\frac{1}{n}X^{*\top}(g) M_{\vec{Z}} Y^*(g) = O_P(1)$, and $(\frac{1}{n}X^{*\top}(g) P_{\widetilde{Z}} X^*(g))^{-1} = O_P(1)$.
For $\frac{1}{n}X^{*\top}(g) M_{\vec{Z}} Y^*(g) = \frac{1}{n}X^{*\top}(g) Y^*(g) - \frac{1}{n}X^{*\top}(g) P_{\vec{Z}} Y^*(g)$, we first show
where $\widehat{Q}^*_{XX}(g) = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}X^*_{i,j}(g)X^{*\top}_{i,j}(g)$, $\widehat{Q}^*_{XW}(g) = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}X^*_{i,j}(g)W^{\top}_{i,j}$, and $\widehat{Q}^*_{X\varepsilon}(g) = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}g_j X^*_{i,j}(g) \hat{\varepsilon}^r_{i,j}$. Notice that
because $\widehat{Q}_{XX,j}=O_P(1)$, $\widehat{Q}_{X\tilde{v},j}=O_P(1)$, and $\widehat{Q}_{\tilde{v}\tilde{v},j}=O_P(1)$. To see these three relations, we note that
by $\widehat{Q}_{XX,j}=O_P(1)$, $\widehat{Q}_{X\overline{Z},j} = O_P(1)$, $\widehat{Q}_{XW,j}=O_P(1)$, $\tilde{\Pi}_{\overline{Z}}=O_P(1)$, and $\widetilde{\Pi}_w = O_P(1)$, where $\widehat{Q}_{X\overline{Z},j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}X_{i,j}\overline{Z}_{i,j}^{\top}$, and $\widehat{Q}_{XW,j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}X_{i,j}W_{i,j}^{\top}$. Similar arguments hold for $\widehat{Q}_{\tilde{v}\tilde{v},j}$. We can also show that
by $\widehat{Q}_{XW,j}=O_P(1)$ and $\widehat{Q}_{\tilde{v}W,j}=O_P(1)$, where $\widehat{Q}_{XW,j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}X_{i,j}W_{i,j}^{\top}$ and $\widehat{Q}_{\tilde{v}W,j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}\tilde{v}_{i,j}W_{i,j}^{\top}$. In addition, we have
by $\widehat{Q}_{X\hat{\varepsilon},j} = O_P(1)$ and $\widehat{Q}_{\tilde{v}\hat{\varepsilon},j}=O_P(1)$ under similar arguments as those in ((ref)), where
Combining ((ref)), ((ref)), ((ref)), $\hat{\beta}^r_L=O_P(1)$, and $\hat{\gamma}^r_L= \widehat{Q}_{WW}^{-1} \widehat{Q}_{w\varepsilon}+o_P(1)=O_P(1)$, we obtain ((ref)). Next, by the fact that $\frac{1}{n}\sum_{j \in [J]} \sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}W_{i,j}^{\top} = 0$, we have
where
Following the same lines of reasoning, we can show that, for all $g \in \textbf{G}$, $\widehat{Q}_{\widetilde{Z}X}^*(g) = O_P(1)$, $\widehat{Q}_{\widetilde{Z}Y}^*(g)=O_P(1)$, and $\widehat{Q}_{WY}^*(g)=O_P(1)$. In addition, by Assumptions (ref)(iv) and (ref)(iii), $\widehat{Q}_{\widetilde{Z}\widetilde{Z}}^{-1} = O_P(1)$ and $\widehat{Q}_{WW}^{-1} = O_P(1)$, which further implies that $\frac{1}{n}X^{*\top}(g) M_{\vec{Z}} Y^*(g) = O_P(1)$.
By a similar argument, we can show $\frac{1}{n}X^{*\top}(g) M_{\vec{Z}} X^*(g) = O_P(1)$. Last, we have
where we use the fact that $\widehat{Q}_{\widetilde{Z}X}^{*}(g) = \widehat{Q}_{\widetilde{Z}X} + o_P(1)$, which is established in Step 1 in the proof of Theorem (ref). In addition, $Q_{\widetilde{Z}X}^{\top}Q_{\widetilde{Z}\widetilde{Z}}^{-1}Q_{\widetilde{Z}X}$ is invertible by Assumptions (ref)(ii) and (ref)(i). This implies $(\frac{1}{n}X^{*\top}(g) P_{\widetilde{Z}} X^*(g))^{-1} = O_P(1)$, which further implies $\hat{\beta}_{L,g}^* = \hat{\beta}_{tsls,g}^* + o_P(r_n^{-1})$.
For the second result, we note that
where the second equality holds because $\sum_{j \in [J]}\sum_{i \in I_{n,j}}g_j\widetilde{Z}_{i,j}W_{i,j}^\top = 0.$ In addition, note that
Combining this with the fact that $\hat{\beta}_{tsls}^r - \beta_n = O_P(r_n^{-1})$ as shown in Lemma (ref), we have $\hat{\beta}_{tsls,g}^* - \beta_n = O_P(r_n^{-1})$. $\blacksquare$
Proof. Note that we assume $\widehat{Q}_{\widetilde{Z}W,j}=o_P(1)$ and $\frac{1}{n}\sum_{j \in [J]} \sum_{i \in I_{n,j}}W_{i,j}\varepsilon_{i,j} = O_P(r_n^{-1})$. The first statement holds because
Next, if Assumption (ref)(i) also holds, then
where the last equality holds by Lemma (ref). This implies $\widehat{Q}_{\widetilde{Z}W,j}(\hat{\gamma}_L-\gamma) = o_P(r_n^{-1})$. In the same manner, we can show $\widehat{Q}_{\widetilde{Z}W,j}(\hat{\gamma}_L^r-\gamma) = o_P(r_n^{-1})$. Last, given Assumptions (ref), (ref), (ref)(i), and the fact that $\widehat{Q}_{\widetilde{Z}W,j}(\hat{\gamma}_L^r-\gamma) = o_P(r_n^{-1})$, Lemma (ref) shows $\hat{\beta}_{L,g}^* - \beta_n = (\hat{\beta}_{L,g}^* - \hat{\beta}_{tsls,g}^*) + (\hat{\beta}_{tsls,g}^* - \beta_n) = O_P(r_n^{-1})$. Then, following the same argument in (ref), we can show $\hat{\gamma}_{L,g}^* - \gamma = O_P(r_n^{-1})$, which leads to the desired result. $\blacksquare$
When $d_x =1$, we define $a_j = Q^{-1}Q_{\widetilde{Z}X}^\top Q_{\widetilde{Z}\widetilde{Z}}^{-1} Q_{\widetilde{Z}X,j}$. For $d_x>2$, $a_j$ is defined in Assumption (ref). Let $\mathbb{S} \equiv \textbf{R}^{d_z \times d_x} \times \textbf{R}^{d_z \times d_z} \times \otimes_{j \in [J]} \textbf{R}^{d_z} \times \textbf{R}^{d_r \times d_r}$ and write an element $s \in \mathbb{S}$ by $s = \left(s_1, s_2, \{s_{3,j}: j \in [J]\}, s_4 \right)$ where $s_{3,j} \in \textbf{R}^{d_z}$ for any $j \in [J]$. Define the function $T$: $\mathbb{S} \rightarrow \textbf{R}$ to be given by
for any $s \in \mathbb{S}$ such that $s_2$ and $s_1^{\top}s_2^{-1}s_1$ are invertible and let $T(s)=0$ otherwise. We also identify any $(g_1, ..., g_q) = g \in \textbf{G} = \{-1, 1\}^J$ with an action on $s \in \mathbb{S}$ given by $gs = \left( s_1, s_2, \{ g_j s_{3,j} : j \in [J] \}, s_4 \right)$. For any $s \in \mathbb{S}$ and $\textbf{G'} \subseteq \textbf{G}$, denote the ordered values of $\{ T(gs) : g \in \textbf{G'} \}$ by $T^{(1)}(s | \textbf{G'}) \leq \ldots \leq T^{(|\textbf{G'}|)} (s | \textbf{G'})$. In addition, for any $\textbf{G'} \subseteq \textbf{G}$, denote the ordered values of $\{ T^*_n(g) : g \in \textbf{G'} \}$ by $T^{*(1)}_n(\textbf{G'}) \leq \ldots \leq T_n^{*(|\textbf{G'}|)}(\textbf{G'})$.
Given this notation we can define the statistics $S_n, \widehat{S}_n \in \mathbb{S}$ as
Let $E_n$ denote the event $ E_n = I\left\{\text{$\widehat{Q}_{\widetilde{Z}X}$ is of full rank value and $\widehat{Q}_{\widetilde{Z}\widetilde{Z}}$ is invertible}\right\}, $ and Assumptions (ref)-(ref) imply that $\liminf_{n \rightarrow \infty} \mathbb{P} \{ E_n=1 \} =1.$ Also let $T^*_n(g)=0$ if $E_n=0$.
We first give the proof for the Wald statistic based on TSLS. Note that whenever $E_n=1$ and $\mathcal{H}_0$ is true, the Frisch-Waugh-Lovell theorem implies that
where $\widehat{Q} = \widehat{Q}^{\top}_{\widetilde{Z}X} \widehat{Q}_{\widetilde{Z}\widetilde{Z}}^{-1} \widehat{Q}_{\widetilde{Z}X}$ and $\iota_J \in \textbf{G}$ is a $J \times 1$ vector of ones.
In the following, we divide the proof into three steps. In the first step, we show
In the second step, we show
In the last step, we prove the desired result.
Step 1. By the continuous mapping theorem, it suffices to show $\widehat{Q}^*_{\widetilde{Z}X}(g) = \widehat{Q}_{\widetilde{Z}X} + o_P(1)$. Note that
Therefore, it suffices to show $\frac{1}{n_j} \sum_{i \in I_{n,j}} \widetilde{Z}_{i,j} \tilde{v}_{i,j} = o_P(1)$ for all $j \in [J]$. Recall $\overline{Z}_{i,j}$ is just $\widetilde{Z}_{i,j}$ interacted with all the cluster dummies and $\widetilde{\Pi}_{\overline{Z}}$ is the OLS coefficient of $\overline{Z}_{i,j}$ defined in ((ref)). Denote $\widetilde{\Pi}_{\overline{Z},j}$ as the $j$-th block of $\widetilde{\Pi}_{\overline{Z}}$, which corresponds to the OLS coefficient of $\widetilde{\Pi}_{\overline{Z},j}$. We have
where the second equality in (ref) holds by $\frac{1}{n_j}\sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}W_{i,j}^{\top}=o_P(1)$ and $\widetilde{\Pi}_{w} = O_P(1)$. In particular,
where $\widehat{Q}_{\widetilde{W}\widetilde{W}} = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\widetilde{W}_{i,j}\widetilde{W}_{i,j}^{\top}$, $\widehat{Q}_{\widetilde{W}\hat{\varepsilon}} = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\widetilde{W}_{i,j}\hat{\varepsilon}_{i,j}$, $\widehat{Q}_{\hat{\varepsilon}\hat{\varepsilon}} = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\hat{\varepsilon}_{i,j}^2$, $\widetilde{W}_{i,j} = W_{i,j} - \widehat{\Gamma}_{w,j}^{\top} \widetilde{Z}_{i,j}$, and $\widehat{\Gamma}_{w,j} = \widehat{Q}^{-1}_{\widetilde{Z}\widetilde{Z},j}\widehat{Q}_{\widetilde{Z}W,j}$. Notice that by $\widehat{Q}_{\widetilde{Z}W,j} = o_P(1)$, $\widehat{\Gamma}_{w,j} = o_P(1)$ so that
Similarly, we have
where the second equality follows from $\widehat{Q}_{W\hat{\varepsilon}}=0$ by the first-order condition of the $k$-class estimators. Furthermore, $\widehat{Q}_{\hat{\varepsilon}\hat{\varepsilon}} \geq c > 0$ by Assumption (ref)(iii) and $\widehat{Q}_{\hat{\varepsilon}\hat{\varepsilon}} \geq \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\dot{\varepsilon}^2_{i,j}$, where $\dot{\varepsilon}_{i,j}$ is the residual from the cluster-level projection of $\varepsilon_{i,j}$ on $W_{i,j}$. Therefore, $\widehat{Q}_{\widetilde{W}\widetilde{W}} - \widehat{Q}_{\widetilde{W}\hat{\varepsilon}} \widehat{Q}^{-1}_{\hat{\varepsilon}\hat{\varepsilon}} \widehat{Q}_{\hat{\varepsilon}\widetilde{W}} = \widehat{Q}_{WW} + o_P(1)$ by combining ((ref)) and ((ref)), and further by Assumption (ref)(iv),
Next, we define $\widehat{Q}_{\widetilde{Z}\dot{X},j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}\dot{X}_{i,j}^{\top}$ and recall $\widehat{Q}_{\widetilde{Z}\widetilde{Z},j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}\widetilde{Z}_{i,j}^{\top}$, where $\dot{X}_{i,j} = X_{i,j} - \widetilde{\Pi}_{w}^{\top}W_{i,j} - \widetilde{\Pi}_{\hat{\varepsilon}}^{\top}\hat{\varepsilon}_{i,j}$, and $\widetilde{\Pi}_{w}$ and $\widetilde{\Pi}_{\hat{\varepsilon}}$ are the OLS coefficients of $W_{i,j}$ and $\hat{\varepsilon}_{i,j}$, respectively, from regressing $X_{i,j}$ on $(\overline{Z}^{\top}_{i,j}, W^{\top}_{i,j}, \hat{\varepsilon}_{i,j})^{\top}$ using the entire sample. Then, we have
where we use the facts that $\widehat{Q}_{\widetilde{Z}W,j} = o_P(1)$, $\widetilde{\Pi}_w = O_P(1)$, $\widetilde{\Pi}_{\hat{\varepsilon}} = O_P(1)$, and $\widehat{Q}_{\widetilde{Z}\hat{\varepsilon},j} = o_P(1)$. In particular,
where $\widehat{Q}_{\tilde{\varepsilon}\tilde{\varepsilon}} = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\tilde{\varepsilon}^2_{i,j}$, $\widehat{Q}_{\tilde{\varepsilon}W} = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\tilde{\varepsilon}_{i,j}W^{\top}_{i,j}$, $\tilde{\varepsilon}_{i,j} = \hat{\varepsilon}_{i,j} - \widehat{\Gamma}_{\hat{\varepsilon},j}^{\top} \widetilde{Z}_{i,j}$, and $\widehat{\Gamma}_{\hat{\varepsilon},j} = \widehat{Q}_{\widetilde{Z}\widetilde{Z},j}^{-1}\widehat{Q}_{\widetilde{Z}\hat{\varepsilon},j}$. Then, by using similar arguments as those for ((ref)), ((ref)), and ((ref)), we obtain
To see the last equality in (ref)}, we note that
where the last equality holds by Assumption (ref)(ii), Lemma (ref), and Lemma (ref). Plugging (ref)} into (ref), we obtain the desired result that $\frac{1}{n_j} \sum_{i \in I_{n,j}} \widetilde{Z}_{i,j} \tilde{v}_{i,j} = o_P(1)$.
Step 2. We note that whenever $E_n=1$, for every $g \in \textbf{G}$,
By Lemma (ref), we have
Note under both cases in Assumption (ref)(ii), we have $\sum_{j \in [J]} \xi_j a_j = 1$ and
Then, for any $\varepsilon >0$, we have
where the last equality holds because $\lambda_{\beta}^{\top}\hat{\beta}^r_{tsls}=\lambda_0$ under $\mathcal{H}_0$.
Note that $T(g \widehat{S}_n) = T(g S_n)$ whenever $E_n=0$ as we have defined $T(s)=0$ for any $s=\left(s_1, s_2, \{s_{3,j}: j \in [J]\}, s_4 \right)$ whenever $s_2$ or $s_1^{\top}s_2^{-1}s_1$ is not invertible. Therefore, results in ((ref)), ((ref)) and ((ref)) imply (ref).
Step 3. Note that by \hyperlink{assumption: 1}{Assumptions (ref)}, \hyperlink{assumption: 2}{(ref)}, (ref), and the continuous mapping theorem, we have
where $\xi_j > 0$ for all $j \in [J]$ by Assumption (ref)(iii). Therefore, we obtain from ((ref)), ((ref)), ((ref)), and the continuous mapping theorem that
For any $x \in \textbf{R}$ letting $\lceil x \rceil$ denote the smallest integer larger than $x$ and $k^* \equiv \lceil |\textbf{G}| (1-\alpha) \rceil$, we obtain from ((ref)) that
Since $\liminf_{n \rightarrow \infty} P\{ E_n =1\} =1$, we have
where the second inequality is due to the Portmanteau's theorem. To see the last inequality, we note that for all $g \in \textbf{G}$, $\mathbb{P}\{ T(gS) = T(-gS)\} = 1$, $\mathbb{P}\{ T(gS) = T(\tilde{g}S)\} = 0$ for $\tilde{g} \notin \{ g, -g \}$, and under the null, $T(gS)$ has the same distribution across $g \in \textbf{G}$. Let $|\textbf{G}| = 2^q$. Then, we have
which implies
For the lower bound, first note that $k^* > |\textbf{G}| - 2$ implies that $\alpha - \frac{1}{2^{J-1}} \leq 0$, in which case the result trivially follows. Now assume $k^* \leq |\textbf{G}| - 2$, then
where the first equality follows from ((ref)), the first inequality follows from Portmanteau's theorem, the second inequality holds because $\mathbb{P} \{ T^{(\textbf{z}+2)} (S | \textbf{G}) > T^{(\textbf{z})}(S | \textbf{G}) \} =1$ for any integer $\textbf{z} \leq |\textbf{G}| -2$ by ((ref)) and Assumption (ref), and the last inequality follows from noticing that $k^*+2 = \lceil |\textbf{G}|((1-\alpha)+2/|\textbf{G}|) \rceil = \lceil |\textbf{G}| (1-\alpha') \rceil$ with $\alpha' = \alpha - \frac{1}{2^{J-1}}$ and the properties of randomization tests.
Lemma (ref) and (ref) in the Supplement further show the other $k$-class estimators and their null-restricted and bootstrap counterparts are asymptotically equivalent to those of the TSLS estimator. Therefore, the results for LIML, FULL, and BA estimators can be derived in the same manner. $\blacksquare$
For the power of the $T_n$-based wild bootstrap test, we focus on the TSLS estimator. As shown in the end of the proof of Theorem (ref), the test statistics constructed using LIML, FULL, and BA estimators and their bootstrap counterparts are asymptotically equivalent to those constructed based on TSLS estimator, which leads to the desired result. Note that
Notice that Assumptions (ref), (ref), (ref)(i), and Lemma (ref) imply that $r_n(\hat{\beta}^r_{tsls} - \beta_n)$ is bounded in probability. This implies
Recall $\widehat{Q}$ and $\widehat{Q}^{*}_g$ defined in (ref) and (ref), respectively. We have $\widehat{Q}^{*-1}_g = \widehat{Q}^{-1}+o_P(1)$ and
where the last equality follows from Lemma (ref) and $\widehat{Q}^*_{\widetilde{Z}X}(g) = \widehat{Q}_{\widetilde{Z}X} + o_P(1)$. Furthermore, we notice that
Therefore, employing ((ref)) with $r_n (\lambda_{\beta}^{\top} \beta_n - \lambda_0) = \lambda_{\beta}^{\top} \mu_{\beta}$, we conclude that whenever $E_n=1$,
Together with $(\ref{eq: boot-unstud-power-2})$, this implies that
This implies
In addition, let $\textbf{G}_s = \textbf{G} \backslash \textbf{G}_w$, where $\textbf{G}_w = \{g \in \textbf{G}: g_{j} = g_{j'}, \forall j,j' \in \mathcal{J}_s\}$. To establish the power result for the bootstrap test, we request that $\hat c_n(1-\alpha)$ does not take values of $T^*_n(g)$ for $g \in \textbf{G}_w$ because otherwise the test statistic and the critical value are asymptotically equivalent even under the alternative. We note that the following condition is imposed in the theorem: $|\textbf{G}_s| = |\textbf{G}| - 2^{J-J_s+1} \geq k^*$. Therefore, based on (ref) and (ref), to establish the desired result, it suffices to show that as $||\mu||_2 \rightarrow \infty$,
which follows under similar arguments as those employed in the proof of Theorem 3.2 in Canay-Santos-Shaikh(2020). $\blacksquare$
Following the same argument in the proof of Theorem (ref), we can show that
where $T_{CR,n}(g)$ is defined in the proof of Theorem (ref) with $\mu=0$ as we are under the null. Then, the rest of the proof is similar to Step 3 in the proof of Theorem (ref). We omit the detail for brevity. $\blacksquare$
For the power analysis, we focus on the TSLS estimator. The results for other $k$-class estimators can be derived in the same manner given Lemma (ref). Recall $a_j$ defined in Assumption (ref). We further define
where $\tilde{Q} = Q^{-1} Q_{\widetilde{Z}X}^{\top} Q_{\widetilde{Z}\widetilde{Z}}^{-1}$, $Q = Q_{\widetilde{Z}X}^\top Q_{\widetilde{Z}\widetilde{Z}}Q_{\widetilde{Z}X}$, $c_{0,g} = \sum_{j \in [J]} \xi_j g_j a_j$, and
We order $\{T_{CR,\infty}(g)\}_{g \in \textbf{G}}$ in ascending order: $(T_{CR,\infty})^{(1)} \leq \cdots \leq (T_{CR,\infty})^{|\textbf{G}|}$. In the proof of Theorem (ref), we have already shown that, under $\mathcal{H}_{1,n}$,
Next, we derive the limit of $\hat{A}_{r,CR}$ and $\hat{A}^*_{r,CR,g}$. We first note that
where the first equality holds by Lemma (ref). This implies
where $\iota_J$ is a $J \times 1$ vector of ones. Similarly, we have
where $\widehat{Q}_{\widetilde{Z}X,j}^*(g) = \frac{1}{n_j}\sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}X_{i,j}^{*\top}(g)$, the second equality is by Lemma (ref), and the last equality holds because $\widehat{Q}_{\widetilde{Z}X,j} = Q_{\widetilde{Z}X,j} + o_P(1)$, $\widehat{Q}_{\widetilde{Z}X,j}^*(g) = Q_{\widetilde{Z}X,j} + o_P(1)$ proved in Step 1 of the proof of Theorem (ref), $r_n(\hat{\beta}^r_{tsls}-\beta_n) = O_P(1)$, and $r_n(\hat{\beta}^*_{tsls,g} - \beta_n) = O_P(1)$. In addition, following the same arguments that lead to (ref), we have
Note $\widehat{Q}^{*-1}_g\widehat{Q}_{\widetilde{Z}X}^{*\top}(g)\widehat{Q}_{\widetilde{Z}\widetilde{Z}}^{-1} \stackrel{p}{\longrightarrow} \tilde{Q}$. Therefore, combining (ref) and (ref), we have
where the second equality is by the fact that $\tilde{Q}Q_{\widetilde{Z}X,j} = a_j I_{d_x}$ and the last convergence is by the fact that $\lambda_\beta^\top r_n(\hat{\beta}^r_{tsls} - \beta_n) = -\mu$. This implies
By the Portmanteau theorem, we have
We aim to show that, as $||\mu ||_2\rightarrow \infty$, we have
where $\textbf{G}_s = \textbf{G} \backslash \textbf{G}_w$, and $\textbf{G}_w = \{g \in \textbf{G}: g_{j} = g_{j'}, \forall j,j' \in \mathcal{J}_s\}$. Then, given $|\textbf{G}_s| = |\textbf{G}|-2^{J-J_s+1}$ and $k^* = \lceil |\textbf{G}|(1-\alpha) \rceil \leq |\textbf{G}|-2^{J-J_s+1}$, (ref) implies that as $ ||\mu||_2 \rightarrow \infty$,
Therefore, it suffices to establish (ref).
By ((ref)), we see that $$T_{CR,\infty}(\iota_J) = \left\Vert \lambda_\beta^\top \tilde{Q} \left[\sum_{j \in [J]} \mathcal{Z}_j \right]+ \mu \right\Vert_{A_{r,CR,\iota_J}},$$ and $A_{r,CR,\iota_J}$ is independent of $\mu$ as $c_{0, \iota_J}=1$. In addition, we have $\lambda_{\min}(\tilde{Q}^\top \lambda_\beta^\top A_{r,CR,\iota_J} \lambda_\beta \tilde{Q})>0$ with probability one. Therefore, for any $\delta>0$, we can find a sufficiently small constant $c>0$ and a sufficiently large constant $M>0$ such that with probability greater than $1-\delta$, for any $\mu$,
On the other hand, for $g \in \textbf{G}_s$, we can write $T_{CR,\infty}(g)$ as
where for $j \in [J]$, $N_{0,g} = \lambda_\beta^\top \tilde{Q} \left[\sum_{j \in [J]} g_j \mathcal{Z}_j \right]$, $N_{j,g} = \lambda_\beta^\top \tilde{Q} \left[g_j\mathcal{Z}_j - a_j \xi_j \sum_{\tilde{j} \in [J]} g_{\tilde{j}}\mathcal{Z}_{\tilde{j}} \right]$, and $c_{j,g} = \xi_j (g_j - c_{0,g})a_j$.
We claim that for $g \in \textbf{G}_s$, $c_{j,g} \neq 0$ for some $j \in \mathcal J_s$. Suppose it does not hold, then it implies that $g_j = c_{0,g}$ for all $j \in \mathcal J_s$, i.e., for all $j \in \mathcal J_s$, $g_j$ shares the same sign, and thus, contradicts the definition of $\textbf{G}_s$. Therefore, combining the claim with the assumption that $\min_{j \in \mathcal J_s}|a_j| > 0$, we have $\min_{g \in \textbf{G}_s} \sum_{j \in [J]} c_{j,g}^2 >0.$
In addition, we note that
where we denote $M_1 = \sum_{j \in [J]} N_{j,g} N_{j,g}^\top$, $M_2 = \sum_{j \in [J]} c_{j,g} N_{j,g}$, and $\overline{c}^2 = \sum_{j \in [J]} c_{j,g}^2$. For notation ease, we suppress the dependence of $(M_1,M_2,\overline{c})$ on $g$. Then, we have
Note for any $d_r \times 1$ vector $u$, by the Cauchy–Schwarz inequality,
where the equal sign holds if and only if there exist $(u,g,c) \in \Re^{d_r} \times \textbf{G}_s \times \Re$ such that
or equivalently,
This occurs with probability zero if $J - 1> d_r$ as $\{g_j \mathcal Z_j\}_{j \in [J]}$ are independent and non-degenerate normal vectors and the matrix $\left(I_J + (-a_1\xi_1,\cdots,-a_J\xi_J)^\top\iota_J^\top \right)$ has rank $J-1$. Therefore, the matrix $\mathbb{M} \equiv M_1 - \frac{M_2 M_2^\top}{\overline{c}^2}$ is invertible with probability one. Specifically, denote $\mathbb{M}$ as $\mathbb{M}(g)$ to highlight its dependence on $g$. We have $\max_{g \in \textbf{G}_s }(\lambda_{\min}(\mathbb{M}(g)))^{-1} = O_P(1)$. In addition, denote $\frac{M_2}{\overline{c}} + \overline{c} \mu $ as $\mathbb{V}$, which is a $d_r \times 1$ vector. Then, we have
where the second equality is due to the Sherman–Morrison–Woodbury formula.
Next, we note that
where $ \mathbb{M}_0 = N_{0,g} - \frac{c_{0,g} M_2}{\overline{c}^2} = N_{0,g} - \frac{c_{0,g} (\sum_{j \in [J]} c_{j,g} N_{j,g})}{\sum_{j \in [J]} c_{j,g}^2}.$ With these notations, we have
where the first inequality holds because $(u+v)^\top A (u+v) \leq 2 (u^\top A u + v^\top A v)$ for some $d_r\times d_r$ positive semidefinite matrix $A$ and $u,v \in \Re^{d_r}$, the second inequality holds due to the fact that $\mathbb{M}^{-1} \mathbb{V} (1+ \mathbb{V}^\top \mathbb{M}^{-1}\mathbb{V})^{-1} \mathbb{V}^\top \mathbb{M}^{-1}$ is positive semidefinite, the third inequality holds because $\mathbb{V}^\top \mathbb{M}^{-1}\mathbb{V}$ is a nonnegative scalar, and the last holds by substituting in the expressions for $\mathbb{M}_0$ and $\overline{c}$.
Then, we have
Combining (ref) and (ref), we have, as $||\mu||_2 \rightarrow \infty$,
where the second inequality is by the fact that $k^* \leq |\textbf{G}_s|$, and thus, $(T_{CR,\infty})^{(k^*)} \leq \max_{g \in \textbf{G}_s} T_{CR,\infty}(g)$ and the last convergence holds because $\max_{g \in \textbf{G}_s} C(g) = O_P(1)$ and does not depend on $\mu$. As $\delta$ is arbitrary, we can let $\delta \rightarrow 0$ and obtain the desired result.
For the proof of part (ii), let $\tilde{c}_{CR,n}(1-\alpha)$ denote the $(1-\alpha)$ quantile of
i.e., the bootstrap statistic $T_n^*(g)$ studentized by the original CRVE instead of the bootstrap CRVE. Because $d_r = 1$, we have
Therefore, we have
where we use the fact that
Further note that $\max_{g \in \textbf{G}}|r_n\lambda_{\beta}^{\top}(\hat{\beta}^*_{tsls,g} - \beta_n)| = O_P(1)$ and does not depend on $\mu$, and
and $A_{r,CR,\iota_J}$ does not depend on $\mu$ either.
Last, although $\hat{c}_{CR,n}(1-\alpha)$ depends on $\mu$, we have
for some $C(g)$ that does not depend on $\mu$ as has been proved in Part (i). Therefore, for any $\delta>0$, there exists a constant $c_{\mu}>0$, such that when $|\mu| > c_\mu$,
where $C$ in the first inequality is a constant such that for $n$ being sufficiently large,
the second inequality is by Portmanteau theorem, and the last inequality holds if $c_\mu$ is sufficiently large. This concludes the proof. $\blacksquare$
The proof for the $AR_n$-based wild bootstrap test follows similar arguments as those in \hyperlink{theo: boot-t}{Theorem (ref)}, and thus we keep exposition more concise. Let $\mathbb{S} \equiv \otimes_{j \in [J]} \textbf{R}^{d_z} \times \textbf{R}^{d_z \times d_z}$ and write an element $s \in \mathbb{S}$ by $s = ( \{s_{1j} : j \in [J]\}, s_2)$ where $s_{1j} \in \textbf{R}^{d_z}$ for any $j \in [J]$. Define the function $T_{AR}$: $\mathbb{S} \rightarrow \textbf{R}$ to be given by
Given this notation we can define the statistics $S_n, \widehat{S}_n \in \mathbb{S}$ as
where $\bar{\varepsilon}^r_{i,j} = y_{i,j}-X_{i,j}^{\top}\beta_0-W^{\top}_{i,j}\bar{\gamma}^r$. Note that by the Frisch-Waugh-Lovell theorem,
Therefore, letting $k^* \equiv \lceil |\textbf{G}| (1-\alpha) \rceil$, we obtain from ((ref))-((ref)) that
Furthermore, note that we have
where the third equality follows from $\sum_{j \in [J]} \sum_{i \in I_{n,j}} \widetilde{Z}_{i,j}W_{i,j}^{\top} = 0$. ((ref)) implies that if $k^* > |\textbf{G}| -2$, then $1 \{ T_{AR}(S_n) > T_{AR}^{(k^*)} (\widehat{S}_n | \textbf{G}) \} = 0$, and this gives the upper bound in Theorem \hyperlink{theo: boot-AR}{(ref)}. We therefore assume that $k^* \leq |\textbf{G}| -2$, in which case
Then, to examine the right hand side of ((ref)), first note that by Assumptions (ref) and (ref), and the continuous mapping theorem we have
where $\xi_j > 0$ for all $j \in [J]$. Furthermore, by Assumptions (ref)(i), Lemma (ref), and $\beta_n=\beta_0$, we have
and thus, for every $g \in \textbf{G}$,
We thus obtain from results in ((ref))-((ref)) and the continuous mapping theorem that
Then, by the Portmanteau’s theorem and the properties of randomization tests, we have
The lower bounds follow by applying similar arguments as those for \hyperlink{theo: boot-t}{Theorem (ref)}.
For the $AR_{CR,n}$-based wild bootstrap test, define the statistics $S_{CR,n}, \widehat{S}_{CR,n} \in \mathbb{S}$ as
Then, we have for any action $g \in \textbf{G}$ that
We set $E_n \in \textbf{R}$ to equal $E_n \equiv 1 \left\{ \frac{r_n^2}{n^2} \sum_{j\in J} \sum_{i \in I_{n,j}} \sum_{k \in I_{n,j}} \widetilde{Z}_{i,j} \widetilde{Z}_{k,j}^{\top} \bar{\varepsilon}^r_{i,j} \bar{\varepsilon}^r_{k,j} \; \text{is invertible} \right\},$ and have
In addition, similar to the case with $AR_n$ and $AR^*_n(g)$, we have under $\beta_n = \beta_0$,
where $A_{CR} = \sum_{j \in [J]} \mathcal{Z}_j \mathcal{Z}_j^{\top}$, and
Therefore, we have
which follows from ((ref)), ((ref)), the continuous mapping theorem and Portmanteau's theorem. The claim of the upper bound in the theorem then follows from similar arguments as those in Theorem (ref). $\blacksquare$
Define $AR_{\infty}(g) = || \sum_{j \in [J]}g_j \mathcal{Z}_j + \sum_{j \in [J]}\xi_jg_ja_jQ_{\widetilde{Z}X}\mu ||_{A_z}$. In particular, notice that $AR_{\infty}(\iota_J) = || \sum_{j \in [J]} \mathcal Z_j + Q_{\widetilde{Z}X} \mu ||_{A_z}$ since $\sum_{j \in [J]} \xi_j a_j =1$. Following same arguments in the proof of Theorem (ref), we can show that, under $\mathcal{H}_{1,n}$ with $\lambda_{\beta} = I_{d_x}$,
Similar to the proofs of Theorems (ref) and (ref), in order to establish Theorem (ref), it suffices to show that as $||Q_{\widetilde{Z}X}\mu||_2 \rightarrow \infty$,
By Triangular inequality, we have $AR_{\infty}(\iota_J) \geq ||Q_{\widetilde{Z}X}\mu||_{A_z} - O_P(1)$, and $$\max_{g \in \textbf{G}_s}AR_{\infty}(g) \leq \max_{g \in \textbf{G}_s}|\sum_{j \in [J]}g_j \xi_j a_j| ||Q_{\widetilde{Z}X} \mu||_{A_z} + O_P(1).$$ In addition, $\max_{g \in \textbf{G}_s} |\sum_{j \in [J]} g_j \xi_j a_j|<1$ so that as $||Q_{\widetilde{Z}X}\mu||_2 \rightarrow \infty$, $||Q_{\widetilde{Z}X}\mu||_{A_z} \rightarrow \infty$ and
This concludes the proof. $\blacksquare$
First, we show that in the specific case with one endogenous regressor and one IV, if the wild bootstrap procedure for the $AR_n$ statistic is applied to the unstudentized Wald statistic $T_n$, then the resulting test will be asymptotically equivalent to the $AR_n$-based bootstrap test, both under the null and the alternative. More precisely, for $T_n = | \hat{\beta} - \beta_0 |$, where $\hat{\beta}$ is the TSLS estimator, the wild bootstrap generates
Notice that in this case, the restricted TSLS estimator $\hat{\gamma}^r$ and the restricted OLS estimator $\bar{\gamma}^r$ are the same, which implies the residuals $\hat{\varepsilon}^r_{i,j}$ and $\Bar{\varepsilon}^r_{i,j}$ defined in Sections (ref) and (ref), respectively, are the same. Let $\hat{c}^s_n(1-\alpha)$ denote the $(1-\alpha)$-th quantile of $\{ T^{s*}_n(g) \}_{g \in \textbf{G}}$. We show this equivalence below.
Several remarks are in order. First, with one endogenous variable and one IV, the Jacobian matrix $\widehat{Q}_{\widetilde{Z}X}$ is just a scalar, which shows up in both the Wald statistic $T_n$ and the bootstrap critical value. After the cancellation of $\widehat{Q}_{\widetilde{Z}X}$, $T_n$ and its critical value are numerically the same as their AR counterparts, which leads to Theorem (ref). Second, for $\widehat{Q}_{\widetilde{Z}X}$ to be cancelled, we only need $\liminf_{n \rightarrow \infty}\mathbb{P} \{ \widehat{Q}_{\widetilde{Z}X} \neq 0 \} = 1$, which is very mild. It holds when $\widehat{Q}_{\widetilde ZX}$ is continuously distributed. Even when both $\widetilde{Z}_{i,j}$ and $X_{i,j}$ are discrete, it still holds if there exists at least one strong IV cluster, i.e., $Q_{\widetilde{Z}X} \neq 0$, where $Q_{\widetilde{Z}X}$ is the probability limit of $\widehat{Q}_{\widetilde{Z}X}$. When both $\widetilde{Z}_{i,j}$ and $X_{i,j}$ are discrete and $Q_{\widetilde{Z}X} = 0$, this condition still holds if some type of CLT holds such that $r_n \widehat{Q}_{\widetilde{Z}X} \xrightarrow{\enskip d \enskip} N(c,\sigma^2)$, as $N(c,\sigma^2)$ is continuous. Third, the robustness of the $T_n$-based bootstrap test in ((ref)) does not carry over to the general case with multiple IVs as $T_n$ and its bootstrap statistic can no longer be reduced to their AR counterparts. Fourth, the robustness to weak IV cannot be extended to the $T_{CR,n}$-based bootstrap test, for which we have to further bootstrap the CRVE.
Second, we discuss wild bootstrap inference with other weak-IV-robust statistics. To introduce the test statistics, we define the sample Jacobian as
and define the orthogonalized sample Jacobian as
where $\widehat{\Omega} = n^{-1}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\sum_{k \in I_{n,j}}f_{i,j}f_{k,j}^{\top} $, and $\widehat{\Gamma}_{l} = n^{-1} \sum_{j \in [J]} \sum_{i \in I_{n,j}} \sum_{k \in I_{n,j}} \left( \widetilde{Z}_{i,j} X_{i,j,l} \right) f_{k,j}^\top,$ for $l=1, ..., d_x$. Therefore, under the null $\beta_n = \beta_0$ and the framework where the number of clusters tends to infinity, $\widehat{D}$ equals the sample Jacobian matrix $\widehat{G}$ adjusted to be asymptotically independent of $\widehat{f}$.
Then, the cluster-robust version of Kleibergen(2002),Kleibergen(2005)'s LM statistic is defined as
In addition, the conditional quasi-likelihood ratio (CQLR) statistic in Kleibergen(2005), Newey-Windmeijer(2009), and Guggenberger-Ramalho-Smith(2012) are adapted from Moreira(2003)'s conditional likelihood ratio (CLR) test, and its cluster-robust version takes the form
where $rk_n$ is a conditioning statistic and we let $rk_n = n \widehat{D}^\top \widehat{\Omega}^{-1} \widehat{D}$.\footnote{This choice follows Newey-Windmeijer(2009). Kleibergen(2005) uses alternative formula for $rk_n$, and Andrews-Guggenberger(2019) introduce alternative CQLR test statistic.}
The wild bootstrap procedure for the LM and CQLR tests is as follows. We compute
for any $g = (g_1, ..., g_q) \in \textbf{G}$, where the definition of $\widehat{f}_g^*$ and $f^*_{k,j}(g_j)$ is the same as that in Section (ref). Then, we compute the bootstrap analogues of the test statistics as
Let $\hat{c}_{LM,n}(1-\alpha)$ and $\hat{c}_{LR,n}(1-\alpha)$ denote the $(1-\alpha)$-th quantile of $\{LM_{n}^*(g)\}_{g \in \textbf{G}}$ and $\{LR^*_{n}(g)\}_{g \in \textbf{G}}$, respectively. We notice that with at least one strong cluster,
where $\widetilde{D} = \left( \widetilde{D}_1, ..., \widetilde{D}_{d_x}\right)$, and for $l=1, ..., d_x$, $$ \widetilde{D}_l = Q_{\widetilde{Z}X} - \left\{ \sum_{j \in [J]} \left(\xi_j Q_{\widetilde{Z}X,j,l} \right) \left( \sqrt{\xi_j} \mathcal{Z}_{\varepsilon,j} \right) \right\} \left\{ \sum_{j \in [J]} \xi_j \mathcal{Z}_{\varepsilon,j} \mathcal{Z}_{\varepsilon,j}^{\top} \right\}^{-1} \sum_{j \in [J]} \sqrt{\xi_j} \mathcal{Z}_{\varepsilon,j}.$$ Although the limiting distribution is nonstandard, we are able to establish the validity results by connecting the bootstrap LM test with the randomization test and by showing the asymptotic equivalence of the bootstrap LM and CQLR tests in this case. We conjecture that similar results can also be established for other weak-IV-robust statistics proposed in the literature.
Notice that when $d_x = d_z = 1$, $\widehat{Q}_{\widetilde{Z}X}$ is a scalar, and the restricted TSLS estimator $\hat{\gamma}^r$ is equivalent to the restricted OLS estimator $\bar{\gamma}^r$, which is well defined by Assumption (ref)(iv), so that $\hat{\varepsilon}_{i,j}^r = \bar{\varepsilon}_{i,j}^r$. Therefore,
and whenever $\widehat{Q}_{\widetilde{Z}X} \neq 0$,
by $\beta_0 = \hat{\beta}^r_{tsls}$, $\sum_{j \in [J]}\sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}W_{i,j}^{\top}=0$, and $\hat{\varepsilon}^r_{i,j} = \varepsilon_{i,j} - X_{i,j}^{\top}( \hat{\beta}^r_{tsls} - \beta_n ) - W_{i,j}^{\top}(\hat{\gamma}^r_{tsls} - \gamma)$.
In addition, for the bootstrap statistics we have
and whenever $\widehat{Q}_{\tilde{Z}X} \neq 0$,
Therefore, $1\{ T_n > \hat{c}^s_n(1-\alpha) \}$ is equal to $1\{ AR_n > \hat{c}_{AR,n}(1-\alpha) \}$ whenever $\widehat{Q}_{\widetilde{Z}X} \neq 0$. We conclude that $\liminf_{n \rightarrow \infty} \mathbb{P} \left\{ \phi_n^s = \phi_n^{ar} \right\} =1$ because $\liminf_{n \rightarrow \infty} \mathbb{P} \left\{ \widehat{Q}_{\widetilde{Z}X} \neq 0 \right\}=1$. $\blacksquare$
The proof for the bootstrap LM test follows similar arguments as those for the studentized version of the bootstrap AR test. Let $\mathbb{S} \equiv \textbf{R}^{d_z \times d_x} \times \otimes_{j \in [J]} \textbf{R}^{d_z}$, and write an element $s \in \mathbb{S}$ by $s = \left( \{ s_{1,j} : j \in [J] \}, \{ s_{2,j} : j \in [J] \} \right)$. We identify any $(g_1, ..., g_q) = g \in \textbf{G} = \{-1, 1\}^J$ with an action on $s \in \mathbb{S}$ given by $gs = \left( \{ s_{1,j} : j \in [J] \}, \{ g_j s_{2,j} : j \in [J] \} \right)$. We define the function $T_{LM}: \mathbb{S} \rightarrow \textbf{R}$ to be given by
for any $s \in \mathbb{S}$ such that $\sum_{j \in [J]} s_{2,j} s_{2,j}^{\top}$ and $D(s)^{\top} \left( \sum_{j \in [J]} s_{2,j} s^{\top}_{2,j} \right)^{-1} D(s)$ are invertible and set $T_{LM}(s)=0$ whenever one of the two is not invertible, where
for $s_{1,j} = (s_{1,j,1}, ..., s_{1,j,d_x})$ and $l=1, ..., d_x$.
Furthermore, define the statistic $S_n$ as
Note that for $l = 1, ..., d_x$ and $j \in [J]$, by Assumptions (ref)(iii) and (ref)(i) we have
where $Q_{\tilde{Z}X,j,l}$ denotes the $l$-th column of the $d_z \times d_x$-dimensional matrix $Q_{\tilde{Z}X,j}$. Then, by Assumptions (ref)(ii) and (ref)(iii), (ref)(i) and (ref)(ii), and the continuous mapping theorem we have
where $\xi_j > 0$ for all $j \in [J]$. Also notice that for $l=1, ..., d_x$,
and $\frac{r_n}{n}\sum_{i \in I_{n,j}} \tilde{Z}_{i,j}\bar{\varepsilon}_{i,j}^r = \frac{r_n}{n}\sum_{i \in I_{n,j}} \tilde{Z}_{i,j}\varepsilon_{i,j} + o_P(1)$ by $\beta_n = \beta_0$ and Lemma (ref). In addition, we set $A_n \in \textbf{R}$ to equal
and we have
which holds because $\left\{ \mathcal{Z}_j : j \in [J] \right\}$ are independent and continuously distributed with covariance matrices that are of full rank, and $Q_{\tilde{Z}X,j}$ are of full column rank for all $j \in \mathcal{J}_s$, by Assumptions (ref)(ii), (ref)(i), and (ref)(ii).
It follows that whenever $A_n=1$,
In what follows, we denote the ordered values of $\{ T_{LM}(gs): g \in \textbf{G} \}$ by
Next, we have
where the final inequality follows from ((ref)), ((ref)), ((ref)), ((ref)), the continuous mapping theorem and Portmanteau's theorem. Therefore, setting $k^* \equiv \lceil |\textbf{G}|(1-\alpha)\rceil$, we can obtain from ((ref)) that
where the final inequality follows by $gS \overset{d}{=} S$ for all $g \in \textbf{G}$ and the properties of randomization tests. Then, we notice that for all $g \in \textbf{G}$, $T_{LM}(gS)=T_{LM}(-gS)$ with probability 1, and $P\{ T_{LM}(gS) = T_{LM}(\tilde{g}S) \}=0$ for $\tilde{g} \notin \{g, -g\}$. Therefore,
The claim of the upper bound in the theorem then follows from ((ref)) and ((ref)}). The proof for the lower bound is similar to that for the bootstrap AR test, and thus is omitted.
To prove the result for the CQLR test, we note that
where the third equality follows from the mean value expansion $\sqrt{1+x} = 1+ (1/2)(x+o(1))$, the fourth and last equalities follow from $AR_{CR,n} - rk_n <0$ w.p.a.1 since $AR_{CR,n} = O_P(1)$ while $rk_n \rightarrow \infty$ w.p.a.1 under \hyperlink{assumption: 3}{Assumption (ref)(i)}. Using arguments similar to those in ((ref)), we obtain that for each $g \in \textbf{G}$,
by $AR^*_{CR,n}(g) - rk_n <0$ w.p.a.1 since $AR^*_{CR,n}(g) = O_P(1)$ for each $g \in \textbf{G}$. Then, it follows that whenever $A_n=1$,
Then, we obtain that
where the second inequality follows from ((ref)), ((ref)), ((ref)), ((ref)), the continuous mapping theorem and Portmanteau's theorem. Finally, the upper and lower bounds for the studentized bootstrap CQLR test follows from the previous arguments for the bootstrap LM test. $\blacksquare$
In this section, we report further simulation results for DGP 1. We notice that the overall pattern is very similar to that observed in Section (ref).
First, we report the results for DGP 1 with the number of clusters $J$ equal to $9$ or $20$. The size results with $d_z=1$ and $d_z=3$ are reported in Tables (ref) and (ref), respectively. In addition, we report the further power results in Figures (ref) - (ref) for $d_z=1$ and $J=9$ or $20$, Figures (ref) - (ref) for $d_z=3$ and $J=6$ or $12$, and Figures (ref) - (ref) for $d_z=3$ and $J=9$ or $20$, respectively. For the power results, we vary the value of $\beta_0-\beta$ from $-3$ to $3$ with the true value of $\beta$ equal to 1 throughout the simulations. $\rho_{\varepsilon v}$ is set equal to $0.5$.
Second, we report the simulation results under DGP 1 for the case where $\Pi \in \{0.25, 0.375, 0.5\}$ remains the same across clusters (i.e., homoskedastic $\Pi$), so that the cluster-level heterogeneity in identification strength originates solely from the heterogeneity in cluster size. Other settings remain the same as those in DGP 1 in Section (ref). Tables (ref)-(ref) report the size properties of the ten inference methods with $d_z=1$ or $3$. In addition, Figures (ref)-(ref) report the power properties of the Wald and AR tests, respectively, with $d_z=1$. Figures (ref)-(ref) report those for the Wald and AR tests, respectively, with $d_z=3$.
Furthermore, we investigate the relationship between the performance of alternative procedures and the number of clusters $J$. Specifically, we plot the size results as a function of $J$ (from $J=6$ to $30$) for $d_z=1$ in Figure (ref) and $d_z=3$ in Figure (ref), respectively. In order to handle the cases with relatively large values of $J$ (e.g., $J=30$), we let the cluster heterogeneity parameter $r$ equal to 3 for $d_z=1$ and 2.4 for $d_z=3$, respectively, so that inference methods that are based on cluster-level estimators (IM and CRS) can also be implemented even with $J=30$.\footnote{As the total number of observations $n$ is equal to 500, the smallest and largest cluster sizes are 2 and 63 for $r=3$ (4 and 56 for $r=2.4$), respectively, under $J=30$.} We focus on DGP 1 in (ref) of the main text with $\Pi=0.25$ and $\rho_{\epsilon v}=0.5$ throughout these simulations.
We observe that similar to the findings in previous simulations, IM and CRS tests can have considerable over-rejections when the number of clusters is relatively large, while ASY and BCH tests typically have large size distortions when the number of clusters is relatively small. Among the seven Wald-based inference methods, only the three wild bootstrap tests (WRE, W-B, and W-B-S) have reasonable size control across different values of $J$. Although slightly lower than the nominal level when $J$ is small, their null rejection frequencies tend to become rather close to $10\%$ as the number of clusters becomes sufficiently large (e.g., when $J$ is larger than 20). We also notice that all the three AR-based inference methods (AR-ASY, AR-B, and AR-B-S) tend to control size. The two bootstrap AR tests over-reject slightly when $J$ is small, while the asymptotic AR test (AR-ASY) can be rather conservative in the over-identified case (Figure (ref)).
Finally, we investigate the power properties of alternative procedures as a function of $J$, from $J=6$ to $30$. Similar to previous simulations, we focus on the power performance of the inference methods that have good size control (namely, WRE, W-B, W-B-S, AR-B, AR-B-S, and AR-ASY). The true value of $\beta$ is set equal to 1 throughout the simulations. Figures (ref) and (ref) report the results for $d_z=1$ with $\beta_0 = 3$ and $-1$, respectively. In addition, Figures (ref) and (ref) report those for $d_z=3$ with $\beta_0 = 3$ and $-1$, respectively. We highlight several observations below. First, the power of all the six inference methods tends to increase with the number of clusters. Second, our recommended procedure W-B-S has the highest power among the six methods across different values of $J$. Third, among the AR tests, the unstudentized bootstrap AR test (AR-B) has the best power performance while AR-ASY has the lowest power. The three observations are in line with our theoretical results in the main text.
In this section, we present the power results for DGP 2. Specifically, Figures (ref) and (ref) present the results for the Wald and AR tests, respectively, with average $J$ equal to 5.437 or 20.189. In addition, Figures (ref) and (ref) present those with average $J$ equal to 10.437 or 29.276.
In this section, we introduce two more simulation designs: DGPs 3 and 4.
DGP 3. We consider a time series IV model.
where $\iota_{d_z}$ is a $d_z$-dimensional vector of ones. We vary the IV strength $\Pi \in \{0.25, 0.375, 0.5\}$, the number of IVs $d_z \in \{1,5\}$, and the AR(1) parameter $\rho_{_{TS}} \in \{0.3, 0.5, 0.7\}$. We let $\beta=1$ and $\gamma =0$. The number of observations is set at 160. The simulation design is similar to the one for time series data considered by Ibragimov-Muller(2010), Bester-Conley-Hansen(2011), and Canay-Romano-Shaikh(2017). We divide the data into $J$ consecutive blocks (clusters) of equal size, where $J \in \{8, 10, 16\}$.
DGP 4. We consider a spatial IV model.
where $\beta = 1$ and $\gamma = 0$. The number of observations is 160 and the observations lie within a square of dimension $40 \times 40$. The location for each observation $s_i = (s_{i_1}, s_{i_2})$ are drawn uniformly within the square, i.e., we let $s_{i_1} \sim U[0, 40]$ independently of $s_{i_2} \sim U[0, 40]$. The distance between observations at locations $s_i$ and $s_j$ is Euclidean: $d_{ij} = \sqrt{(s_{i_1} - s_{j_1})^2 + (s_{i_2} - s_{j_2})^2}$. The simulation design is similar to those considered in Bester-Conley-Hansen(2011), sun2015, and lee2016. The errors $(u_i, v_i)^{\top}$ are drawn from $N\left(
,
\right)$. In addition, $u_i$ and $u_j$ have a correlation equal to $\rho_{_{SP}}^{d_{ij}}$ between them. Similarly, $v_i$ and $v_j$ have a correlation equal to $\rho_{_{SP}}^{d_{ij}}$. Furthermore, the IVs $Z_i$ are drawn from $N(0, I_{d_z})$, with a correlation $\rho_{_{SP}}^{d_{ij}}$ between each element of $Z_i$ and its corresponding element of $Z_j$. Therefore, $\rho_{SP}$ controls the degree of spatial dependence and we let $\rho_{_{SP}} \in \{0.3, 0.5, 0.7\}$. We vary the IV strength $\Pi \in \{0.25, 0.375, 0.5\}$ and the number of IVs $d_z \in \{1,5\}$. To implement inference, we split the data into $J$ groups (clusters) with equal area, where $J \in \{6, 9, 12\}$.\footnote{Specifically, for $J=9$, we split the data into 9 groups formed by splitting into three equal parts both vertically and horizontally. For $J=6$, we split into two equal parts vertically and three horizontally, while for $J=12$, we split into three equal parts vertically and four horizontally.}
Results for DGP 3. Tables (ref)-(ref) report the size results for the time series IV model with $d_z=1$ and $d_z=5$, respectively. We observe that the standard Wald test (ASY) does not control size when the number of clusters is small, especially in the over-identified case. For example, when $d_z=5$, $J=8$, and $\Pi=0.25$, its null rejection frequencies are $0.183, 0.196$, and $0.251$ for $\rho_{TS} = 0.3, 0.5,$ and $0.7$, respectively. BCH improves upon ASY but still has null rejection frequencies equal to $0.135$, $0.143$, and $0.186$, respectively, under such settings. On the other hand, IM and CRS typically have large size distortion when the number of clusters is large. For example, when $d_z=5$, $J=16$, and $\Pi=0.25$, the null rejection frequencies for IM are $0.259, 0.259,$ and $0.249$, for $\rho_{TS} = 0.3, 0.5,$ and $0.7$, respectively, while those for CRS are $0.267, 0.266,$ and $0.256$. By contrast, the wild bootstrap-based Wald inference methods (WRE, W-B, W-B-S) have good size control regardless of the number of clusters and the number of IVs. Therefore, we focus on the bootstrap methods when comparing power. Figures (ref) and (ref) report the power properties of the Wald and AR tests, respectively, with $d_z=1$. We find that the power curves are very similar to those in DGPs 1 and 2. Overall, our W-B-S has the best power, especially against distant alternatives.
Results for DGP 4. Tables (ref)-(ref) report the size for the spatial IV model with $d_z=1$ and $d_z=5$, respectively. The patterns are similar to those observed in other DGPs. ASY and BCH tend to have relatively large size distortions when the number of clusters $J$ is small, while IM and BCH tend to have large distortions when $J$ is large. In addition, the size distortions of these four inference methods increase when the number of IVs becomes large ($d_z=5$). By contrast, WRE, W-B, and W-B-S have good size control across different settings. Figures (ref) and (ref) report the power for the Wald and AR tests, respectively. Again, we see that overall W-B-S has the best power among these tests.