EconBase
← Back to paper

Wild Bootstrap for Instrumental Variables Regressions with Weak and Few Clusters

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

480,170 characters · 33 sections · 193 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
This text was truncated for display. The citation measures were computed over the complete text.

Wild Bootstrap Inference for Instrumental Variables Regressions with Weak and Few Clusters

abstractWe study the wild bootstrap inference for instrumental variable regressions under an alternative asymptotic framework that the number of independent clusters is fixed, the size of each cluster diverges to infinity, and the within cluster dependence is sufficiently weak. We first show that the wild bootstrap Wald test controls size asymptotically up to a small error as long as the parameters of endogenous variables are strongly identified in at least one of the clusters. Second, we establish the conditions for the bootstrap tests to have power against local alternatives. We further develop a wild bootstrap Anderson-Rubin test for the full-vector inference and show that it controls size asymptotically even under weak identification in all clusters. We illustrate their good performance using simulations and provide an empirical application to a well-known dataset about US local labor markets. \\ Keywords: Wild Bootstrap, Weak Instrument, Clustered Data, Randomization Test. JEL codes: C12, C26, C31

{18pt} {10pt} \belowdisplayskip\abovedisplayskip {5pt} \abovedisplayshortskip \belowdisplayshortskip {8pt} \belowdisplayskip\abovedisplayskip {4pt} \linespread{1.6}

Introduction

The instrument variable (IV) regression is one of the five most commonly used causal inference methods identified by Angrist-Pischke(2008), and it is often applied with clustered data. For example, Young(2021) analyzes 1,359 IV regressions in 31 papers published by the American Economic Association (AEA), out of which 24 papers account for clustering of observations. Three issues arise when running IV regressions with clustered data. First, the strength of IVs may be heterogeneous across clusters with a few clusters providing the main identification power. Indeed, Young(2021) finds that in the average paper of his AEA samples, with the removal of just one cluster or observation, the first-stage $F$ can decrease by 28%, and 38% of reported 0.05 significant two-stage least squares (TSLS) results can be rendered insignificant at that level. Second, the number of clusters is small in many IV applications. For instance, Acemoglu2011 cluster the standard errors at the country/polity level, resulting in 12-19 clusters, Glitz2020 cluster at the sectoral level with 16 sectors, and Rogall2021 clusters at the province (district) level with 11 provinces (30 districts), respectively. When the number of clusters is small, conventional cluster-robust inference procedures may be unreliable. Third, it is also possible that IVs are weak in all clusters, in which case researchers need to use weak-identification-robust inference methods Andrews-Stock-Sun(2019).

Motivated by these issues, in this paper we study the inference for IV regressions with a small, and thus, fixed number of clusters and weak within-cluster dependence. We denote clusters in which the parameters of the endogenous variables are strongly identified as strong IV clusters. First, we show that a wild bootstrap Wald test, with or without the cluster-robust variance estimator (CRVE), controls size asymptotically up to a small error, as long as there exists at least one strong IV cluster. Second, we show that the wild bootstrap tests have power against local alternatives at 10% and 5% significance levels when there are at least five and six strong IV clusters, respectively. Third, in the common case of testing a single restriction, we show the bootstrap Wald test with CRVE is more powerful than that without CRVE for distant local alternatives. Fourth, we develop the full-vector inference based on a wild bootstrap Anderson-Rubin(1949) test, which controls size asymptotically up to a small error regardless of instrument strength. Fifth, in the Online Supplement we show that for the common case with a single endogenous variable and a single IV, a wild bootstrap test based on the unstudentized Wald statistic (i.e., the one without CRVE) is asymptotically equivalent to a certain wild bootstrap AR test under both null and alternative, implying that in such a case it is fully robust to weak IV. Sixth, we show in the Online Supplement that bootstrapping weak-IV-robust tests other than the AR test (e.g., the Lagrange multiplier test or the conditional quasi-likelihood ratio test) controls asymptotic size when there is at least one strong IV cluster.

Our inference procedure is empirically relevant. First, we notice that besides the aforementioned examples, the numbers of clusters may also be rather small in studies that estimate the region-wise effects of certain intervention if the partition of clusters is at the state level. We illustrate the usefulness of our methods in Section (ref) by applying them to the well-known dataset of ADH2013 in the estimation of the effects of Chinese imports on local labor markets in three US Census Bureau-designated regions (South, Midwest, and West) with 11-16 clusters at the state level. Second, our bootstrap inference is flexible with respect to IV strength: the bootstrap Wald test allows for cluster-level heterogeneity in the first stage, while its AR counterpart is fully robust to weak IVs. Figure 1 reports the estimated first-stage coefficients for each cluster (state) in ADH2013's (ADH2013) dataset, which suggests that there exists substantial variation in the IV strength among states. Specifically, the first-stage coefficients of some states are quite large compared with the rest in the region, while some other states have coefficients that are rather close to zero, and thus, potentially subject to weak identification. Some states even have opposite signs for their first-stage coefficients. However, there is no existing proven-valid inference method for IV regressions with few clusters, where IVs may be weak in some clusters. Therefore, we believe our bootstrap methods enrich practitioners' toolbox by providing a reliable inference in this context. Third, different from the analytical inference based on the widely used heteroskedasticity and autocorrelation consistent (HAC) estimators, our approach is agnostic about the within-cluster (weak) dependence structure and thus avoids the use of tuning parameters to estimate the covariance matrix for dependent data.

figure[figure omitted — 5,039 chars of source]

The contributions in the present paper relate to several strands of literature. First, it is related to the literature on the cluster-robust inference.\footnote{See Cameron(2008), Conley2011inference, Imbens-Kolesar(2016), abadie2022should, Hagemann(2017), Hagemann(2019), Hagemann2019placebo, Hagemann2020inference, Mackinnon-Webb(2017), Djogbenou-Mackinnon-Nielsen(2019), Mackinnon-Nielsen-Webb(2019), Ferman2019inference, Hansen-Lee(2019), Menzel2021bootstrap, Mackinnon2021, among others, and mackinnon2022cluster for a recent survey.} Djogbenou-Mackinnon-Nielsen(2019), Mackinnon-Nielsen-Webb(2019), and Menzel2021bootstrap show the bootstrap validity under the asymptotic framework in which the number of clusters diverges to infinity.\footnote{We refer interested readers to mackinnon2022cluster for detailed discussions on this asymptotic framework and the alternative asymptotic framework that treats the number of clusters as fixed.} Ibragimov-Muller(2010), Ibragimov-Muller(2016), Bester-Conley-Hansen(2011), Canay-Romano-Shaikh(2017), Hagemann(2019), Hagemann2019placebo, Hagemann2020inference, Canay-Santos-Shaikh(2020), and Hwang(2020) consider an alternative asymptotic framework in which the number of clusters is treated as fixed, while the number of observations in each cluster is relatively large and the within cluster dependence is sufficiently weak. However, the inference methods proposed by BCH and Hwang(2020) require an (asymptotically) equal cluster-level sample size,\footnote{See Bester-Conley-Hansen(2011) and Hwang(2020) for details.} while those proposed by IM and CRS require strong identification in all clusters. In contrast, our bootstrap Wald tests are more flexible as it does not require an equal cluster size and only needs one strong IV cluster for size control and five to six for local power, thus allowing for substantial heterogeneity in IV strength. To our knowledge, there is no alternative method in the literature that remains valid in such a context. For full-vector inference, we further provide the bootstrap AR tests, which are fully robust to weak identification.

Second, Canay-Santos-Shaikh(2020) prove the validity of wild bootstrap with a few large clusters by innovatively connecting it with a randomization test with sign changes. Our results rely heavily on this technical breakthrough, but also complement Canay-Santos-Shaikh(2020) in the following aspects. First, Canay-Santos-Shaikh(2020) focus on the linear regression with exogenous regressors and then extend their analysis to a score bootstrap for the GMM estimator. Instead, we focus on extending the wild restricted efficient (WRE) cluster bootstrap, which is popular among empirical researchers, for general $k$-class IV estimators. Therefore, our procedure cannot be formulated as a score bootstrap in the GMM setting. Second, we study the local power for the Wald test both with and without CRVE. The power analysis of the Wald test with CRVE is new to the literature and technically involved because with a fixed number of clusters, the CRVE has a random limit. We rely on the Sherman–Morrison–Woodbury formula for the matrix inverse to analyze the behaviors of the test statistic and its wild bootstrap counterparts under local alternatives. Based on such a detailed asymptotic analysis, we establish the local power of the bootstrap Wald test studentized by CRVE, and show it is higher than that for the Wald test without CRVE studentization for distant local alternatives in the empirically prevalent case of testing a single restriction. Indeed, as the linear regression is a special case of the linear IV regression studied in this paper, we believe the same result holds for the local power of the Wald tests in Canay-Santos-Shaikh(2020). Third, different from its unstudentized counterpart, the local power of the Wald test with CRVE is established without the assumption that the cluster-level first-stage coefficients have the same sign, which may not hold in some empirical studies (e.g., Figure (ref)).

Third, our paper is related to the literature of WRE bootstrap for IV regressions.\footnote{See, for example, Davidson-Mackinnon(2008), Davidson-Mackinnon(2010), finlay2014, Finlay-Magnusson(2019), Roodman-Nielsen-MacKinnon-Webb(2019), and Mackinnon2021.} In particular, Davidson-Mackinnon(2008), Davidson-Mackinnon(2010) first proposed the restricted efficient (RE) and WRE bootstrap procedures for IV regression with homoskedastic and heteroskedastic errors, respectively. finlay2014, Finlay-Magnusson(2019) extended it to clustered data, and Roodman-Nielsen-MacKinnon-Webb(2019) incorporated it into the Stata package “boottest". Our key contribution to this strand of literature is to theoretically prove the validity of the (modified) WRE cluster bootstrap under the alternative asymptotic framework in which IVs may be weak for some clusters and the number of clusters is treated as fixed. In particular, we carefully re-design the first stage of the WRE procedure to allow for cluster-level heterogeneity of IV strength. Our theoretical analysis also covers the general $k$-class IV estimators, including TSLS, the limited information maximum likelihood (LIML) estimator or the modified LIML estimator proposed by Fuller(1977). In addition, we theoretically prove the power advantages of the bootstrap Wald test with CRVE compared with the one without. This result is new and justifies the proposal in the literature of using the bootstrap Wald test studentized by CRVE from the power perspective.

Furthermore, our paper is related to the literature on weak identification, in which various normal approximation-based inference approaches are available for nonhomoskedastic cases, among them Stock-Wright(2000), Kleibergen(2005), Andrews-Cheng(2012), Andrews(2016), Andrews-Mikusheva(2016), Andrews(2018), Moreira-Moreira(2019), and Andrews-Guggenberger(2019). As Andrews-Stock-Sun(2019) remark, an important question concerns the quality of the normal approximations with nonhomoskedastic observations. On the other hand, when implemented appropriately, alternative approaches such as bootstrap may substantially improve the inference for IV regressions.\footnote{See, for example, Davidson-Mackinnon(2008), Davidson-Mackinnon(2010), Moreira-Porter-Suarez(2009), Wang-Kaffo(2016), Kaffo-Wang(2017), Finlay-Magnusson(2019), and Young(2021), among others. In addition, we notice that tuvaandorj2021robust develops permutation versions of weak-IV-robust tests with (non-clustered) heteroskedastic errors.} We contribute to this literature by establishing formally that wild bootstrap-based AR tests control asymptotic size under both weak identification and a small number of clusters. However, contrary to the Wald tests, we find in simulations that the bootstrap AR test without CRVE has better power properties than the one studentized by CRVE. We formally establish the local power of the bootstrap AR test without CRVE. Furthermore, we show that the bootstrap validity for other weak-IV-robust tests requires at least one strong IV clusters (Sections (ref)-(ref) in the Online Supplement), which provides theoretical support to the simulation results in finlay2014, Finlay-Magnusson(2019) documenting that with clustered observations, bootstrapping the AR tests results in better finite sample size control than bootstrapping alternative robust tests.

Finally, we note that although empirical applications often involve settings with substantial first-stage heterogeneity such as those shown in Figure (ref), related econometric literature remains rather sparse. Abadie-Gu-Shen(2019) exploit such heterogeneity to improve the asymptotic mean squared error of IV estimators with independent and homoskedastic observations. Instead, we focus on developing inference methods that are robust to the first-stage heterogeneity for data with a small number of clusters, while allowing for (weak) within-cluster dependence and heteroskedasticity.

The remainder of this paper is organized as follows. Section (ref) presents the setup and the bootstrap procedures. Section (ref) presents main assumptions, while Sections (ref)-(ref) provide asymptotic results. Simulations in Section (ref) suggest that our procedures have outstanding finite sample size control and, in line with our theoretical analysis, the bootstrap Wald test with CRVE has power advantages compared with the other bootstrap tests. The empirical application is presented in Section (ref). We conclude in Section (ref).

Setup, Estimation, and Inference

Setup

Throughout the paper, we consider the setup of a linear IV regression with clustered data,

eqnarray[eqnarray omitted — 113 chars of source]

where the clusters are indexed by $j \in [J] = \{1, ..., J\}$ and units in the $j$-th cluster are indexed by $i \in I_{n,j} = \{1, ..., n_j \}$. In (ref), we denote $y_{i,j} \in \textbf{R}$, $X_{i,j} \in \textbf{R}^{d_x}$, and $W_{i,j} \in \textbf{R}^{d_w}$ as an outcome of interest, endogenous regressors, and exogenous regressors, respectively. Furthermore, we let $Z_{i,j} \in \textbf{R}^{d_z}$ denote the IVs for $X_{i,j}$. $\beta \in \textbf{R}^{d_x}$ and $\gamma \in \textbf{R}^{d_w}$ are unknown structural parameters.

We allow the parameter of interest $\beta$ to shift with respect to (w.r.t.) the sample size, which incorporates the analyses of size and local power in a concise manner: $\beta_n = \beta_0 + \mu_{\beta}/r_n$, where $\mu_{\beta} \in \textbf{R}^{d_x}$ is the local parameter and $r_n$ is the convergence rate of the score defined later. We let $\lambda^{\top}_{\beta} \beta_0 = \lambda_0$, $\lambda_{\beta} \in \textbf{R}^{d_x \times d_r}$, $\lambda_0 \in \textbf{R}^{d_r}$, and $d_r$ denotes the number of restrictions under the null hypothesis. Define $\mu = \lambda_{\beta}^{\top} \mu_{\beta}$. Then, the null and local alternative hypotheses studied in this paper can be written as

align[align omitted — 109 chars of source]

K-Class IV Estimators

Throughout the paper, we consider the $k$-class IV estimators of the form:

align[align omitted — 261 chars of source]

where $\vec{Z} = [Z:W]$, $\vec{X} = [X:W]$, $Y$, $X$, $Z$, and $W$ are $n \times 1$, $n \times d_x$, $n \times d_z$, and $n \times d_w$-dimensional vectors and matrices formed by $y_{i,j}$, $X^{\top}_{i,j}$, $Z^{\top}_{i,j}$, and $W^{\top}_{i,j}$, respectively. We also denote $P_A = A(A^{\top} A)^{-1}A^{\top}$ and $M_A = I_n - P_A$, where $A$ is an $n$-dimensional matrix and $I_n$ is an $n$-dimensional identity matrix. Specifically, we focus on four $k$-class estimators: (1) the TSLS estimator, where $\hat{\kappa} = \hat{\kappa}_{tsls} = 1$, (2) the LIML estimator, where

align*[align* omitted — 223 chars of source]

(3) the FULL estimator, where $\hat{\kappa} = \hat{\kappa}_{full} = \hat{\kappa}_{liml} - C/(n-d_z-d_w)$ with some constant $C$, and (4) the bias-adjusted TSLS (BA) estimator proposed by Nagar(1959) and Rothenberg(1984), where $\hat{\kappa} = \hat{\kappa}_{ba} = n/(n-d_z+2)$. Theoretically, we show that these four estimators are asymptotically equivalent.\footnote{More specifically, we show in Lemma S.B.1 in the Online Supplement that under our asymptotic framework, for $L \in \{\text{liml},\text{full},\text{ba}\}$, $\hat{\beta}_{L} = \hat{\beta}_{tsls} + o_P(r_n^{-1})$ and $\hat{\beta}_{tsls} - \beta_n = O_P(r_n^{-1})$, where $r_n = O(n^{1/2})$.} However, in simulations, we find that FULL has the best finite sample performance in the overidentified case.

Wild Bootstrap Inference

Inference Procedure for Wald Statistics

For the rest of the paper, we define $\widetilde{Z}_{i,j}$ as the residual of regressing $Z_{i,j}$ on $W_{i,j}$ using the entire sample, and that for any random vectors $U_{i,j}$ and $V_{i,j}$, $\widehat{Q}_{UV,j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}U_{i,j}V_{i,j}^\top$ and $\widehat{Q}_{UV} = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}U_{i,j}V_{i,j}^\top$.

For inference, we construct Wald statistics based on the $k$-class estimator $\hat{\beta}_L$ defined in ((ref)) with $\hat k = \hat k_L$ for $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$. When the $d_r \times d_r$ weighting matrix $\hat{A}_{r}$ is asymptotically deterministic in the sense of Assumption (ref) below (such as $\hat{A}_r = I_{d_r}$, the $d_r \times d_r$ identity matrix), we denote $T_n$ as the Wald statistic without CRVE and define it as

eqnarray[eqnarray omitted — 121 chars of source]

where $||u||_A = \sqrt{u^{\top}Au}$ for a generic vector $u$ and a weighting matrix $A$.

Further define $\hat{A}_{r,CR}$ as the inverse of the CRVE:

align[align omitted — 498 chars of source]

$\widehat{\Omega}_{CR} = n^{-1} \sum_{j \in [J]} \sum_{i \in I_{n,j}} \sum_{k \in I_{n,j} } \widetilde{Z}_{i,j} \widetilde{Z}_{k,j}^{\top} \hat{\varepsilon}_{i,j} \hat{\varepsilon}_{k,j}$, and $\hat{\varepsilon}_{i,j}$ is the unrestricted residual defined in ((ref)). The corresponding Wald statistic studentized by CRVE is denoted as

align[align omitted — 118 chars of source]

In the case of testing the value of a single endogenous variable $\beta = \beta_0$, the test statistics reduce to $T_n = |\hat{\beta}_L - \beta_0|$, and $T_{CR,n} = |(\hat{\beta}_L - \beta_0)/\widehat{V}^{1/2}|$, respectively, where $\widehat{V}$ denotes the usual cluster-robust variance estimator.

We reject the null hypothesis in ((ref)) at $\alpha$ significance level if the test statistics are greater than their corresponding critical values ($\hat c_n (1-\alpha)$ and $\hat c_{CR,n}(1-\alpha)$ defined below for $T_n$ and $T_{CR,n}$, respectively). We compute the critical values by a wild bootstrap procedure described below.

enumerate• For $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$, compute the null-restricted residual $$\hat{\varepsilon}^r_{i,j} = y_{i,j} - X_{i,j}^{\top} \hat{\beta}^r_L - W_{i,j}^{\top} \hat{\gamma}^r_L,$$ where $\hat{\beta}^r_L$ and $\hat{\gamma}^r_L$ are null-restricted $k$-class IV estimators of $\beta$ and $\gamma$ from $(y_{i,j}, X^{\top}_{i,j}, W^{\top}_{i,j}, Z^{\top}_{i,j})^{\top}$,\footnote{The null-restricted $k$-class estimator is defined as \begin{align*} & \hat{\beta}^r_L = \hat{\beta}_L - \left(X^{\top}P_{\widetilde{Z}}X -\hat{\mu}_L X^{\top}M_{\vec{Z}}X\right)^{-1} \lambda_{\beta} \left(\lambda_{\beta}^{\top}(X^{\top}P_{\widetilde{Z}}X - \hat{\mu}_L X^{\top}M_{\vec{Z}}X)^{-1} \lambda_{\beta}\right)^{-1} (\lambda_{\beta}^{\top}\hat{\beta}_L - \lambda_0), \\ & \hat{\gamma}^r_L = (W^\top W)^{-1}W^\top(Y-X\hat{\beta}_L^r), \;\; where \;\; \hat{\mu}_L = \hat{\kappa}_L-1 \;\; and \;\; \tilde{Z} = M_W Z, \end{align*} for $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$; e.g., see Appendix B of Roodman-Nielsen-MacKinnon-Webb(2019) for a general formula.} and the unrestricted residual is defined as \begin{align} \hat{\varepsilon}_{i,j} = y_{i,j} - X_{i,j}^{\top} \hat{\beta}_L - W_{i,j}^{\top} \hat{\gamma}_L, \end{align} where $\hat{\beta}_L$ and $\hat{\gamma}_L$ are defined in ((ref)) with $\hat{\kappa}_L$. • Construct $\overline{Z}_{i,j}$ as \begin{align*} \overline{Z}_{i,j} = \left( \widetilde{Z}^{\top}_{i,j} 1\{j=1\}, ......, \widetilde{Z}^{\top}_{i,j} 1\{j = J\} \right)^{\top}, \end{align*} where $\widetilde{Z}_{i,j}$ is the residual of regressing $Z_{i,j}$ on $W_{i,j}$ using the entire sample. • Compute the first-stage residual \begin{align} \tilde{v}_{i,j} = X_{i,j} - \widetilde{\Pi}_{\overline{Z}}^\top \overline{Z}_{i,j} - \widetilde{\Pi}_{w}^\top W_{i,j}, \end{align} where $\widetilde{\Pi}_{\overline{Z}}$ and $\widetilde{\Pi}_{w}$ are the OLS coefficients of $\overline{Z}_{i,j}$ and $W_{i,j}$ from regressing $X_{i,j}$ on $(\overline{Z}^{\top}_{i,j}, W^{\top}_{i,j}, \hat{\varepsilon}_{i,j})^{\top}$ using the entire sample. • Let $\textbf{G} = \{-1, 1\}^J$ and for any $g = (g_1, ..., g_J) \in \textbf{G}$ generate \begin{eqnarray} X_{i,j}^*(g) = \widetilde{\Pi}_{\overline{Z}}^\top\overline{Z}_{i,j} + \widetilde{\Pi}_{w}^\top W_{i,j} + g_j \tilde{v}_{i,j}, \quad y_{i,j}^*(g) = X_{i,j}^{*\top}(g) \hat{\beta}^r_L + W_{i,j}^{\top} \hat{\gamma}^r_L + g_j \hat{\varepsilon}^r_{i,j}. \end{eqnarray} For each $g=(g_1, ..., g_J) \in \textbf{G}$, compute $\hat{\beta}^*_{L,g}$ and $\hat{\gamma}^*_{L,g}$, the analogues of the estimators $\hat{\beta}_L$ and $\hat{\gamma}_L$ using $\left( y_{i,j}^{*}(g), X_{i,j}^{*\top}(g) \right)^{\top}$ in place of $\left( y_{i,j}, X_{i,j}^{\top} \right)^{\top}$ and the same $(Z^{\top}_{i,j}, W^{\top}_{i,j})^{\top}$. Compute the bootstrap analogue of the Wald statistics:\footnote{Let $X^*(g)$ and $Y^*(g)$ be the $n \times d_x$ matrix constructed using $X_{i,j}^*(g)$ and the $n \times 1$ vector constructed using $Y_{i,j}^*(g)$, respectively. For $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$ and $g \in \textbf{G}$, \begin{align*} \hat{\beta}^*_{L,g}= \left(X^{*\top}(g) P_{\widetilde{Z}} X^*(g) - \hat{\mu}^*_{L,g} X^{*\top}(g) M_{\vec{Z}} X^*(g) \right)^{-1}\left(X^{*\top}(g) P_{\widetilde{Z}} Y^*(g) - \hat{\mu}^*_{L,g} X^{*\top}(g) M_{\vec{Z}} Y^*(g)\right), \;\; \text{where} \;\; \hat{\mu}^*_{L,g} = \hat{\kappa}^*_{L,g}-1. \end{align*}} \begin{align} T^*_{n}(g) = || \lambda_{\beta}^{\top} \hat{\beta}^*_{L,g} - \lambda_0 ||_{\hat{A}_{r}}, \quad T^*_{CR,n}(g) = ||\lambda_{\beta}^{\top} \hat{\beta}^*_{L,g} - \lambda_0||_{\hat{A}_{r,CR,g}^*}. \end{align} $\hat{A}^*_{r,CR,g}$ is the bootstrap counterpart of $\hat A_{r,CR}$ defined as \begin{align*} \hat{A}^*_{r,CR,g} = \left( \lambda^{\top}_{\beta} \widehat{V}^*_g \lambda_{\beta} \right)^{-1}, \quad \widehat{V}^*_g = \widehat{Q}_g^{*-1} \widehat{Q}^{*\top}_{\widetilde{Z}X}(g) \widehat{Q}^{-1}_{\widetilde{Z}\widetilde{Z}} \widehat{\Omega}^*_{CR,g} \widehat{Q}^{-1}_{\widetilde{Z}\widetilde{Z}}\widehat{Q}^*_{\widetilde{Z}X}(g) \widehat{Q}_g^{*-1}, \end{align*} where for $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$, \begin{align} & \widehat{\Omega}^*_{CR,g} = \frac{1}{n} \sum_{j \in [J]} \sum_{i \in I_{n,j}} \sum_{k \in I_{n,j} } \widetilde{Z}_{i,j} \widetilde{Z}_{k,j}^{\top} \hat{\varepsilon}^*_{i,j}(g) \hat{\varepsilon}^*_{k,j}(g), \quad \widehat{Q}^*_{\widetilde{Z}X}(g) = \frac{1}{n} \sum_{j \in [J]} \sum_{i \in I_{n,j}} \widetilde{Z}_{i,j} X_{i,j}^{*\top}(g), \notag \\ & \hat{\varepsilon}^*_{i,j}(g) = y^*_{i,j}(g) - X_{i,j}^{*\top}(g) \hat{\beta}^*_{L,g} - W_{i,j}^{\top} \hat{\gamma}^*_{L,g}, \quad \text{and} \quad \widehat{Q}^*_g = \widehat{Q}_{\widetilde{Z}X}^{*\top}(g) \widehat{Q}_{\widetilde{Z}\widetilde{Z}}^{-1} \widehat{Q}^*_{\widetilde{Z}X}(g). \end{align} • \hypertarget{Step 5}{We} compute the $1-\alpha$ quantiles of $\{ T^*_{n}(g): g \in \textbf{G} \}$ and $\{ T^*_{CR,n}(g): g \in \textbf{G} \}$: \begin{align*} & \hat{c}_{n}(1-\alpha) = \inf \left\{ x \in \textbf{R}: \frac{1}{|\textbf{G}|} \sum_{g \in \textbf{G}}1 \{ T^*_{n}(g) \leq x \} \geq 1-\alpha \right\},\\ & \hat{c}_{CR,n}(1-\alpha) = \inf \left\{ x \in \textbf{R}: \frac{1}{|\textbf{G}|} \sum_{g \in \textbf{G}}1 \{ T^*_{CR,n}(g) \leq x \} \geq 1-\alpha \right\}, \end{align*} where $1\{E\}$ equals one whenever the event $E$ is true and equals zero otherwise and $|\textbf{G}| = 2^J$. The bootstrap test for $\mathcal{H}_0$ rejects whenever $T_{CR,n}$ exceeds the critical value $\hat{c}_{CR,n}(1-\alpha)$ and $T_{n}$ exceeds $\hat{c}_{n}(1-\alpha)$ for Wald statistics with and without CRVE, respectively.

Several remarks are in order.

remStep 1 imposes null when computing the residuals in the structural equation ((ref)), which is advocated by Cameron(2008), Davidson-Mackinnon(2010), mackinnon2022cluster, and Canay-Santos-Shaikh(2020), among others. The estimators $\widetilde{\Pi}_{\overline{Z}}$ and $\widetilde{\Pi}_w$ in Step 3 are similar to the efficient reduced-form estimators in the WRE bootstrap procedures which have superior finite sample performance for IV regressions, even when the instruments are rather weak. In this paper, we focus on extending the WRE cluster bootstrap procedure\footnote{See, e.g., Mackinnon2021 for a detailed discussion of the WRE cluster bootstrap procedure and an efficient computation algorithm.} because (1) we find the resulting bootstrap also has excellent finite sample performance for IV regressions with a small number of clusters, and (2) we want to be consistent with the suggestions in the literature.
remTo adapt to the new framework, we carefully modify the original WRE cluster bootstrap procedure and use the fully interacted IVs $\overline{Z}_{i,j}$ in Step 3, which effectively estimate the cluster-level first-stage coefficient for $\widetilde Z_{i,j}$. Such a modification is necessary because we allow for the IV strength, and thus, the first stage coefficient of $\widetilde Z_{i,j}$ to be different across clusters.\footnote{Finlay-Magnusson(2009), Finlay-Magnusson(2019), Roodman-Nielsen-MacKinnon-Webb(2019), and Mackinnon2021 considered the IV model with clustered data in which the first-stage coefficient is homogeneous across clusters. } We then need to preserve such a heterogeneity in the bootstrap sample by estimating the first-stage coefficient within each cluster. If we further assume the identification strength is homogeneous across clusters, then we can just use $\widetilde Z_{i,j}$ to generate $X_{i,j}^*(g)$ in Step 3. We emphasize that in the whole procedure, the fully interacted IVs $\overline{Z}_{i,j}$ are only needed to construct $X_{i,j}^*(g)$, and we still use the uninteracted IVs $\widetilde Z_{i,j}$ when computing $(\hat{\beta}_L^{\top}, \hat{\gamma}_L^{\top})^{\top}$ in ((ref)), where $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$, and their null-restricted and bootstrap counterparts (i.e., $(\hat{\beta}^{r\top}_L, \hat{\gamma}_L^{r\top})^{\top}$ in Step 1 and $(\hat{\beta}^{*\top}_{L,g}, \hat{\gamma}^{*\top}_{L,g})^{\top}$ in Step 4).
remOur second modification of the original WRE cluster bootstrap procedure is that when regressing $X_{i,j}$ on $(\overline{Z}^{\top}_{i,j}, W^{\top}_{i,j}, \hat{\varepsilon}_{i,j})^{\top}$ in Step 3, we use the unrestricted residuals $\hat{\varepsilon}_{i,j}$ instead of the null-restricted residuals $\hat{\varepsilon}_{i,j}^r$. This modification is required to establish the asymptotic power results (Theorems (ref) and (ref)).

Inference Procedure for Weak-instrument-robust Statistics

In this section, we describe a full-vector wild bootstrap procedure for Anderson-Rubin (AR) type weak-IV-robust statistics with or without CRVE. Recall that $\beta_n = \beta_0 + \mu_{\beta}/r_n$. Under the null, we have $\mu_\beta = 0$, or equivalently, $\beta_n =\beta_0$. First, define the AR statistic without CRVE as

eqnarray[eqnarray omitted — 181 chars of source]

where $\hat{A}_z$ is a $d_z \times d_z$ weighting matrix with an asymptotically deterministic limit as specified in Assumption (ref) below, $f_{i,j} = \widetilde{Z}_{i,j}\bar{\varepsilon}^r_{i,j}$, $\bar{\varepsilon}^r_{i,j} = y_{i,j} - X^{\top}_{i,j} \beta_0 - W_{i,j}^{\top} \bar{\gamma}^r $, and $\bar{\gamma}^r$ is the null-restricted ordinary least squares (OLS) estimator of $\gamma$: $$\bar{\gamma}^r = \left( \sum_{j \in [J]} \sum_{i \in I_{n,j}} W_{i,j} W_{i,j}^{\top} \right)^{-1}\sum_{j \in [J]} \sum_{i \in I_{n,j}} W_{i,j} (y_{i,j} - X^{\top}_{i,j} \beta_0).$$ Second, we also define the AR statistic with the (null-imposed) CRVE as

eqnarray*[eqnarray* omitted — 216 chars of source]

Our wild bootstrap procedure for the AR statistics is defined as follows.

enumerate• Compute the null-restricted residual $\bar{\varepsilon}^r_{i,j} = y_{i,j} - X_{i,j}^{\top} \beta_0 - W_{i,j}^{\top} \bar{\gamma}^r.$ • Let $\textbf{G} = \{-1, 1\}^J$ and, for any $g = (g_1, ..., g_J) \in \textbf{G}$, define \begin{eqnarray*} \widehat{f}^*_g = n^{-1} \sum_{j \in [J]} \sum_{i \in I_{n,j}} f^*_{i,j}(g_j),\quad and \quad f^*_{i,j}(g_j) = \widetilde{Z}_{i,j}\varepsilon^*_{i,j}(g_j), \end{eqnarray*} where $\varepsilon^*_{i,j}(g_j) = g_j \bar{\varepsilon}^r_{i,j}$. Compute the bootstrap statistics: \begin{align*} AR^*_{n}(g) = \big\Vert \widehat{f}^*_g \big\Vert_{\hat{A}_{z}} \quad and \quad AR^*_{CR,n}(g) = \big\Vert \widehat{f}^*_g \big\Vert_{\hat{A}_{CR}}. \quad \end{align*} • Let $\hat{c}_{AR,n}(1-\alpha)$ and $\hat{c}_{AR,CR,n}(1-\alpha)$ denote the $(1-\alpha)$-th quantile of $\{AR_{n}^*(g)\}_{g \in \textbf{G}}$ and $\{AR^*_{CR,n}(g)\}_{g \in \textbf{G}}$, respectively.

We note that unlike the $T_{CR,n}$-based Wald test in Section (ref), there is no need to bootstrap the CRVE for the $AR_{CR,n}$ test even though $\hat{A}_{CR}$ also admits a random limit. This is because $\hat{A}_{CR}$, unlike $\hat{A}_{r,CR}$ defined in (ref), is invariant to the sign changes. Two remarks are in order.

remIf $2^J$ is too large, the researcher can replace $\textbf{G} = \{-1,1\}^J$ by $\hat{\textbf{G}}$ where $\hat{\textbf{G}} = \{g_1,\cdots,g_B\}$, $g_1 = \iota_J$, $\iota_J$ is a $J$-dimension vector of ones, and $g_2,\cdots,g_B$ are i.i.d. Uniform $\textbf{G}$, which is equivalent to generating each element of $g$ by a Rademacher random variable. This is known as the stochastic approximation and recommended by Canay-Romano-Shaikh(2017) and Canay-Santos-Shaikh(2020) for the inference under the asymptotic framework with a fixed number of clusters. All the theoretical results in this paper remains true if $\textbf{G}$ is replaced by $\hat{\textbf{G}}$ by letting $B \rightarrow \infty$. See LR06 for more discussion on the validity of the stochastic approximation.
remIn the following, we study the size and power properties of the proposed wild bootstrap procedure applied to the $k$-class estimators. There are three main takeaways for applied researchers from our theoretical investigation. First, for inference based on Wald statistics, we recommend the bootstrap Wald test studentized by CRVE because of its superior power properties. This method is denoted as W-B-S, and we provide its pseudo code in Section (ref) in the Online Supplement. Second, for the over-identified case, we recommend using the bootstrap Wald tests with Fuller's modified LIML estimator over TSLS because the former has a smaller finite sample bias. Third, for the full-vector inference, if researchers are concerned that all clusters are weak, and thus, would like to implement a weak-identification-robust procedure, then we recommend the bootstrap AR test without CRVE studentization because of its power advantage over alternative asymptotic and bootstrap AR tests that are studentized by CRVE.

Main Assumptions and Several Examples

In this section, we introduce the assumptions that will be used in our analysis of the wild bootstrap tests under a small number of clusters in Sections (ref)-(ref). Define $\widehat{Q} = \widehat{Q}_{\widetilde{Z}X}^{\top} \widehat{Q}_{\widetilde{Z}\widetilde{Z}}^{-1} \widehat{Q}_{\widetilde{Z}X}$ and $Q$ as the probability limits of $\widehat{Q}$.

assumption(i) For each $j \in [J]$, \begin{align} \widehat Q_{\widetilde ZW,j} = o_P(1) \end{align} and \begin{align} \frac{r_n}{n} \sum_{i \in I_{n,j}} W_{i,j}\varepsilon_{i,j} = O_P(1), \end{align} where the convergence rate $r_n$ satisfies $r_n \rightarrow \infty$ and $r_n = O(\sqrt{n})$. (ii) There exists a collection of independent random variables $\{ \mathcal{Z}_j : j \in [J] \}$, where $\mathcal{Z}_j \sim N(0, \Sigma_j)$ with positive definite $\Sigma_j$ for all $j \in [J]$ such that \begin{align} \left\{ \frac{r_n}{n} \sum_{i \in I_{n,j}} \widetilde{Z}_{i,j}\varepsilon_{i,j} : j \in [J] \right\} \xrightarrow{\enskip d \enskip} \left\{ \mathcal{Z}_{j} : j \in [J] \right\}. \end{align} (iii) For each $j \in [J]$, $n_j/n \rightarrow \xi_j >0$. (iv) $\frac{1}{n}\sum_{j \in [J]} \sum_{i \in I_{n,j}}W_{i,j}W_{i,j}^\top$ is invertible.
remWe have (ref) if \begin{eqnarray} \frac{1}{n_j} \sum_{i \in I_{n,j}} \Big\Vert W_{i,j}^{\top} \left( \widehat{\Gamma}_n - \widehat{\Gamma}_{n,j} \right) \Big\Vert^2 \stackrel{p}{\longrightarrow} 0, \end{eqnarray} where $\widehat{\Gamma}_n$ and $\widehat{\Gamma}_{n,j}$ are the $d_w \times d_z$ matrices that satisfy the following orthogonality conditions: $\sum_{j \in [J]} \sum_{i \in I_{n,j}}W_{i,j}(Z_{i,j} - \widehat{\Gamma}_n^\top W_{i,j})^\top =0$, and $\sum_{i \in I_{n,j}}W_{i,j}(Z_{i,j} - \widehat{\Gamma}_{n,j}^\top W_{i,j})^\top =0.$ Canay-Santos-Shaikh(2020) impose the same condition as ((ref)) and point out that one sufficient but not necessary condition for ((ref)) is that the distributions of $(Z^{\top}_{i,j}, W^{\top}_{i,j})_{i \in I_{n,j}}$ are the same across clusters. In our empirical application in Section (ref), we compute $\widehat{Q}_{\widetilde{Z}W,j}$ for each cluster and find that they are all close to zero. We refer to Tables (ref)--(ref) in the Online Supplement for more details.{\footnote{It is also possible to rigorously test whether $Q_{\widetilde ZW,j} $ is close to zero for $j \in [J]$ if we can consistently estimate the covariance matrix of $\widehat Q_{\widetilde ZW,j}$ under a more detailed within-cluster dependence structure.}} Assumption (ref)(iii) gives the restriction on cluster sizes, and Assumption (ref)(iv) ensures $\widetilde Z_{i,j}$ is uniquely defined.
remWe have (ref) and (ref) hold if the within cluster dependence is sufficiently weak for some type of CLT to hold. We provide three examples of data structure below that satisfy our requirements. We focus on (ref) and let $U_{i,j} = \widetilde Z_{i,j}\varepsilon_{i,j}$, but similar results apply to ((ref)) by letting $U_{i,j} = W_{i,j}\varepsilon_{i,j}$.
example[Serial Dependence] We use $i$ and $j$ to index time period and cluster, respectively, so that observations have serial dependence over time and are asymptotically independent across clusters. Such settings were considered in BCH (Lemma 1 and Section 4.1), IM (Section 3.1), and CRS (Section S.1) for time series data and IM (Section 3.2) for panel data,\footnote{Specifically, for time series data, they propose to divide the full sample into $J$ (approximately) equal sized consecutive blocks (clusters). For panel data, assuming independence across individuals, then one may treat the observations for each individual as a cluster.} among others. In this setup, we can verify (ref) under different levels of serial dependence. \begin{enumerate} • ($L_q$-Mixingale) Let $U_{i,j}^{(k)}$ denote the $k$-th element of $U_{i,j}$. Suppose there exists a filtration $\mathcal{F}_{i,j}$ that satisfies the following conditions: for some $q\geq 3$ and any $l\geq 0$ and $j \in [J]$ \begin{align*} \left\Vert \mathbb E (U_{i,j}^{(k)} \mid \mathcal{F}_{i-l,j}) \right\Vert_q \leq c_{n_j,i}\psi_l, \quad \left\Vert U_{i,j}^{(k)}- \mathbb E (U_{i,j}^{(k)} \mid \mathcal{F}_{i+l,j}) \right\Vert_q \leq c_{n_j,i}\psi_{l+1}, \end{align*} and $\max_{j \in [J]}\max_{i \in I_{n,j}}c_{n_j,i} = o(n^{1/2})$. Then, LL20 implies (ref) holds with $r_n = \sqrt{n}$. In fact, they show that the partial sum process of $\{U_{i,j}\}_{i \in I_{n,j}}$ can be approximated by a martingale, and thus, is called a mixingale. It forms a very general class of models, including martingale differences, linear processes and various types of mixing and near-epoch dependence processes as special cases. • (Long Memory) Suppose $U_{i,j} = \sum_{l=0}^{\infty}\Psi_l a_{i-l,j}$ where the innovations $a_{i,j} = (a_{i,j}^{(1)},\cdots,a_{i,j}^{(K)})^\top$ are $K$-dimensional martingale difference with respect to an filtration $\mathcal{F}_{i,j}$ such that \begin{align*} \max_{i,j,k}\mathbb E (|a_{i,j}^{(k)}|^{2+d}\mid \mathcal{F}_{i-1,j}) < \infty, a.s. \quad and \quad \mathbb E (a_{i,j}a_{i,j}^\top \mid \mathcal{F}_{i-1,j}) = \Sigma_a, a.s. \end{align*} The $K \times K$ matrix coefficient $\Psi_l$ can be approximated by \begin{align*} \Psi_l \sim \frac{l^{d-1}}{\Gamma(d)} \Pi, \quad as \quad l \rightarrow \infty, \end{align*} where $\Gamma(\cdot)$ is the gamma function, $\Pi$ is a non-singular $K \times K$ matrix of constants that are independent of $l$, and $d \in (0,0.5)$ is the memory parameter. Then, C02 implies ((ref)) holds with $r_n = n^{1/2-d}$. \end{enumerate}
example[Spatial Dependence] This example is proposed by Bester-Conley-Hansen(2011). Suppose we have $n$ individuals indexed by $l$. For $l$-th individual, its location is denoted as $s_l$, which is an $m$-dimensional integer. The distance between individual $l_1$ and $l_2$ is measured by the maximum coordinatewise metric $\text{dist}(l_1,l_2) = ||s_{l_1}-s_{l_2}||_\infty$. Both $\widetilde Z$ and $\varepsilon$ are indexed by the location so that $(\widetilde Z_l,\varepsilon_l) = (\widetilde Z_{s_l},\varepsilon_{s_l})$ for $l \in [n]$. The clusters $I_{n,j}$ for $j \in [J]$ are defined as disjoint regions ($\Lambda_1,\cdots,\Lambda_J$). Let $\mathcal{F}_{\Lambda}$ be the $\sigma$-field generated by a given random field $(\widetilde Z_s,\varepsilon_s)$, $s \in \Lambda$ with $\Lambda$ compact and let $|\Lambda|$ be the number of $s \in \Lambda$. Let $\Upsilon_{\Lambda_1,\Lambda_2}$ denote the minimum distance from an element of $\Lambda_1$ to an element of $\Lambda_2$ where the distance is measured by the maximum coordinatewise metric. The mixing coefficient is then \begin{align*} \alpha_{k_1,k_2}(l) & = \sup\{ \mathbb P(A \cap B) - \mathbb P(A)\mathbb P( B) \}, \\ & s.t. \quad A \in \mathcal{F}_{\Lambda_1}, B \in \mathcal{F}_{\Lambda_2}, |\Lambda_1| \leq k_1, \quad |\Lambda_2| \leq k_2, \Upsilon(\Lambda_1,\Lambda_2) \geq l. \end{align*} Bester-Conley-Hansen(2011) assume the mixing coefficients satisfy (1) $\sum_{l = 1}^\infty l^{m-1} \alpha_{1,1}(l)^{\delta_1/(2+\delta_1)}<\infty$, (2) $\sum_{l = 1}^\infty l^{m-1} \alpha_{k_1,k_2}(l)<\infty$ for $k_1+k_2\leq 4$, and (3) $\alpha_{1,\infty}(l) = O(l^{-m-\delta_2})$ for some $\delta_1>0$ and $\delta_2>0$. Under this assumption and other regularity conditions in their Assumptions 1 and 2, Bester-Conley-Hansen(2011) verifies (ref) with a finite number of clusters ($J$ fixed) and $r_n = \sqrt{n}$.
example[Network Dependence] Suppose we observe $n$ units indexed by $\ell \in [n]$ and an adjacency matrix $\mathcal{A} = \{A_{\ell,\ell'}\}$ where $A_{\ell,\ell'} = 1$ means units $\ell$ and $\ell'$ are linked and $A_{\ell,\ell'} = 0$ means otherwise. We consider a linear-in-means social interaction model studied by BDF09. Specifically, we have \begin{align} y_\ell = \alpha + \beta \frac{\sum_{\ell': A_{\ell,\ell'}=1}y_{\ell'}}{n_\ell} + \gamma B_\ell + \delta \frac{\sum_{\ell': A_{\ell,\ell'}=1}B_{\ell'}}{n_\ell} + \varepsilon_\ell, \end{align} where $n_\ell = |\ell': A_{\ell,\ell'}=1|$ denote node $\ell$th's number of friends and $B_\ell$ represents node $\ell$th's background characteristics, and we assume $\mathbb E(\varepsilon_\ell|B_\ell) = 0$. In this setup, we have $X_\ell = \frac{\sum_{\ell': A_{\ell,\ell'}=1}y_{\ell'}}{n_\ell}$ and $W_\ell = (1, B_\ell, \frac{\sum_{\ell': A_{\ell,\ell'}=1}B_{\ell'}}{n_\ell})^\top$. Following the literature, we assume the adjacency matrix is independent of $\{B_\ell,\varepsilon_\ell\}_{\ell \in [n]}$. Then, BDF09 show that one can use $Z = (\tilde A^2 B, \tilde A^3 B)$ as IVs where $\tilde A$ is the $n \times n$ normalized adjacency matrix with a typical entry $\tilde A_{\ell,\ell'} = A_{\ell,\ell'}/n_{\ell}$ and $B$ is a $n \times 1$ vector of $\{B_\ell\}_{\ell \in [n]}$. The clusters $I_{n,j}$, $j \in [J]$ form a partition of $[n]$. For a subset of indexes $S \subset [n]$, define the conductance of $S$ as $\phi_{\mathcal{A}}(S) = \frac{|\partial_{\mathcal{A}}(S)|}{vol_A(S)}$, where $|\partial_{\mathcal{A}}(S)| = \sum_{\ell \in S}\sum_{\ell' \in [n]/S}{\mathcal{A}}_{\ell,\ell‘}$ is the number of links involving a unit in $S$ and a unit not in $S$ and $vol_{\mathcal{A}}(S) = \sum_{\ell \in S}\sum_{\ell' \in [n]}\mathcal{A}_{\ell,\ell‘}$ is the sum of degrees $\sum_{\ell' \in [n]}{\mathcal{A}}_{\ell,\ell'}$ of units $\ell \in S$. Then, L22 shows (ref) holds with a finite number of clusters ($J$ fixed) and $r_n = \sqrt{n}$ when $\max_{j \in [J]}\phi_\mathcal{A}(I_{n,j}) (\frac{1}{n}\sum_{\ell \in [n]}\sum_{\ell' \in [n]}\mathcal{A}_{\ell,\ell'})\rightarrow 0$ as $n \rightarrow \infty$ and the observations exhibit weak network dependence in the sense of L22. If the clusters $\{I_{n,j}\}_{j \in [J]}$ are latent, L22 further shows it is possible to recover the clusters by spectral clustering, a method that clusters the leading $J$ eigenvectors of network graph Laplacian by the k-means algorithm. We provide more details in Section (ref).
remAs mentioned in the Introduction, our asymptotic framework follows previous studies such as Ibragimov-Muller(2010), Bester-Conley-Hansen(2011), and Canay-Romano-Shaikh(2017), which proposed valid inference methods under the setting that treats the number of clusters as fixed. However, we find in simulations (e.g., see Section (ref)) that for IV regressions, these methods may not control size under data generating processes with heterogeneous cluster sizes and IV strengths, weak IV clusters, or a small number of clusters. The motivation of our study is to propose wild bootstrap-based inference methods that perform well even in such cases and provide the corresponding regularity conditions. We emphasize that the main restriction of the asymptotic framework that treats the number of clusters as fixed is it requires the within-cluster dependence to be sufficiently weak for some type of CLT to hold within each cluster (as illustrated in Examples (ref)-(ref)). For example, as pointed out by mackinnon2022cluster, this requirement rules out the case that $\varepsilon_{i,j}$ follows a factor structure, i.e., \begin{align} \varepsilon_{i,j} = \lambda_{i,j} \eta_{j} + u_{i,j}, \end{align} where $u_{i,j}$ denotes the idiosyncratic error, $\eta_j$ denotes the cluster-wide shock, and the effect of $\eta_j$ on $\varepsilon_{i,j}$ is given by the factor loading $\lambda_{i,j}$. By contrast, such a dependence structure can be handled under the asymptotic framework that lets the number of clusters $J$ to diverge to infinity. Under this framework, a CLT is applied to (appropriately normalized) $\sum_{j \in [J]}\sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}\varepsilon_{i,j}$ and it is thus possible to allow for arbitrary within-cluster dependence under suitable conditions on the cluster heterogeneity; e.g., see Djogbenou-Mackinnon-Nielsen(2019) and Hansen-Lee(2019) for the regularity conditions. In particular, the factor structure in ((ref)) is allowed as long as $n^{-1}\sup_{j \in [J]}n_j \rightarrow 0$. We also refer interested readers to mackinnon2022cluster for an excellent discussion on the two alternative asymptotic frameworks.
remIf there exist cluster fixed effects in the IV model, we can transform our data by projecting out the fixed effects so that our $(y,X,W,Z)$ are expressed as deviations from cluster means and assume that our conditions hold for the model involving the transformed data. Such a practice is also recommended by Djogbenou-Mackinnon-Nielsen(2019) and mackinnon2022cluster. We refer to Example (ref) below for more details. Panel data is a special case of Example (ref) in which the clusters are individual units.
assumption(i) The quantities $\widehat{Q}_{\widetilde{Z}X,j}$, $\widehat{Q}_{\widetilde{Z}\widetilde{Z},j}$, $\widehat{Q}_{\widetilde{Z}X}$, and $\widehat{Q}_{\widetilde{Z}\widetilde{Z}}$ converge in probability to deterministic matrices, which are denoted as $Q_{\widetilde{Z}X,j}$, $Q_{\widetilde{Z}\widetilde{Z},j}$, $Q_{\widetilde{Z}X}$, and $Q_{\widetilde{Z}\widetilde{Z}}$, respectively. (ii) The matrices $Q_{\widetilde{Z}\widetilde{Z},j}$ is invertible for $j \in [J]$. (iii) For all $j \in [J]$, $\widehat{Q}_{XX,j} = O_P(1)$, $\widehat{Q}_{X\varepsilon,j} = O_P(1)$, and $\widehat{Q}_{\dot{\varepsilon}\dot{\varepsilon},j} \geq c>0$ for constant $c$ with probability approaching one, where $\dot{\varepsilon}_{i,j}$ is the residual from the cluster-level projection of $\varepsilon_{i,j}$ on $W_{i,j}$.
remAssumption (ref) is required for our analysis of the Wald test, but not for the AR test. Assumption (ref)(i) holds if the dependence of units within clusters is weak enough to render some type of LLN to hold. Assumption (ref)(ii) is standard in the literature and holds regardless of the IV strength.

We conclude this section with more concrete examples. Specifically, Assumptions (ref) and (ref) hold in Examples (ref) and (ref), and we explain why they may not hold in Examples (ref) and (ref). For Example (ref), Assumption (ref) holds while Assumption (ref) does not. Consequently, the wild bootstrap AR tests are still applicable while the Wald tests are not.

example[Cluster Fixed Effects] Suppose \begin{align} \mathcal Y_{i,j} = \mathcal X_{i,j}^\top \beta + \mathcal W_{i,j}^\top \theta + \eta_{j} + \mathcal E_{i,j}, \end{align} where $\mathcal Y_{i,j}$ is the outcome variable, $\mathcal X_{i,j}$ contains the endogenous variables, $\eta_{j}$ denotes cluster fixed effects, $\mathcal W_{i,j}$ contains the individual-level exogenous variables, $\mathcal E_{i,j}$ is the individual-level idiosyncratic error, and $\mathcal Z_{i,j}$ denotes the IVs. As recommended in Djogbenou-Mackinnon-Nielsen(2019) and mackinnon2022cluster, we can transform the data by projecting out the fixed effects so that our $(y,X,W,Z)$ are expressed as deviations from cluster means. Specifically, let $y_{i,j} = \mathcal Y_{i,j} - \frac{1}{n_j} \sum_{i \in I_{n,j}}\mathcal Y_{i,j}$, $X_{i,j} = \mathcal X_{i,j} - \frac{1}{n_j} \sum_{i \in I_{n,j}}\mathcal X_{i,j}$, $W_{i,j} = \mathcal W_{i,j} - \frac{1}{n_j} \sum_{i \in I_{n,j}}\mathcal W_{i,j}$, $e_{i,j} = \mathcal E_{i,j} - \frac{1}{n_j} \sum_{i \in I_{n,j}}\mathcal E_{i,j}$, and $Z_{i,j} = \mathcal Z_{i,j} - \frac{1}{n_j} \sum_{i \in I_{n,j}}\mathcal Z_{i,j}$ be the cluster-level demeaned versions. Then, we can rewrite (ref) as (ref) and assume our conditions hold for the model involving the transformed data.
example[Heterogeneous IV Strength across Clusters] We allow for cluster-level heterogeneity with regard to IV strength. Consider the following first stage regression: \begin{align*} X_{i,j} = Z_{i,j}^\top \Pi_{z,j,n} + W_{i,j}^\top \Pi_{w,j,n} + V_{i,j}, \end{align*} where $\Pi_{z,j,n}$ and $\Pi_{w,j,n}$ are the coefficients of $Z_{i,j}$ and $W_{i,j}$, respectively, via the cluster-level population projection of $X_{i,j}$ on $Z_{i,j}$ and $W_{i,j}$, for each $j \in [J]$.\footnote{We note $\Pi_{z,j,n}$ and $\Pi_{w,j,n}$ depend on the sample size because the underlying distribution is indexed by $n$.} Then, our model (ref) allows for both $\Pi_{z,j,n}$ and $\Pi_{w,j,n}$ to vary across clusters. In particular, we allow for some of $\Pi_{z,j,n}$ to decay to or be zero for the bootstrap Wald tests and all $\Pi_{z,j,n}$ to decay to or be zero for the bootstrap AR tests. We will come back to this point in Sections (ref)--(ref) with more details.
example[Heterogeneous Slope for the Endogenous Variable] Similar to Canay-Santos-Shaikh(2020), we cannot allow for $\beta$ to be heterogeneous across clusters. As a stylized example, let \begin{align} y_{i,j} = X_{i,j}^\top \beta_j + W_{i,j}^\top \gamma_j +e_{i,j}, \end{align} where $W_{i,j}$ is just the cluster dummies. For $\beta$ equal to a suitable weighted average of $\beta_j$'s, we may rewrite (ref) as (ref) with $\varepsilon_{i,j} = X_{i,j}^\top(\beta_j - \beta) + e_{i,j}$, which implies \begin{align*} \frac{r_n}{n} \sum_{i \in I_{n,j}} \widetilde{Z}_{i,j}\varepsilon_{i,j} = \frac{r_n}{n} \sum_{i \in I_{n,j}} (Z_{i,j} - \bar Z_j) (X_{i,j}^\top(\beta_j - \beta) + e_{i,j} ). \end{align*} We then see that Assumption (ref)(ii) is violated unless $\beta_j = \beta$.
example[Difference-in-Difference] We do not allow for the regression with cluster-level IVs, which usually occurs in difference-in-difference analysis with endogenous treatments. In this setting, the treatment status and group assignment are interpreted as our $X_{i,j}$ and $Z_{i,j}$, respectively, and they are different due to imperfect compliance. In addition, when the assignment is at the cluster level (i.e., some clusters are assigned to the treatment group while the others are assigned to the control group), the value of $Z_{i,j}$ is invariant within each cluster. If there exists cluster fixed effects which needs to be partialled out as in Example (ref), we have $\tilde Z_{i,j} = 0$, which means our Assumption (ref)(ii) is violated because $\Sigma_j$ is zero, and thus, degenerate. Following the suggestion by Canay-Santos-Shaikh(2020), it may be possible to avoid such an issue by clustering more coarsely (e.g., pairing one treated cluster with one control cluster).
example[Cluster-level Endogenous Variables] If $X_{i,j}$ is a cluster-level variable (say, $X_j$), then the resulting within-cluster limiting Jacobian matrix $Q_{\widetilde{Z}X,j}$ may be random and potentially correlated with the within-cluster score component $\mathcal{Z}_j$ as $X_{j}$ is endogenous, which violates Assumption (ref)(i). We notice that similar issues can arise for the approaches by Bester-Conley-Hansen(2011), Hwang(2020), IM, and CRS. On the other hand, our wild bootstrap AR tests ($AR_n$ and $AR_{CR,n}$) only requires Assumption (ref) but not Assumption (ref), and thus, remain valid in this case.

Asymptotic Results for the Wald Tests without CRVE

For the Wald test, we further assume the following assumption.

assumption(i) $Q_{\widetilde{Z}X}$ is of full column rank. (ii) One of the following two conditions holds: (1) $d_x = 1$, or (2) there exists a scalar $a_j$ for each $j \in [J]$ such that $Q_{\widetilde{Z}X,j} = a_j Q_{\widetilde{Z}X}$.

Several remarks are in order.

remAssumption (ref)(i) requires (overall) strong identification for $\beta_n$. Assumption (ref)(ii)(1) states that if there is only one endogenous variable, no further restrictions are required. A single endogenous variable is the leading case in empirical applications involving IV regressions. In fact, 101 out of 230 specifications in Andrews-Stock-Sun(2019)'s (Andrews-Stock-Sun(2019)) sample and 1,087 out of 1,359 in Young(2021)'s (Young(2021)) sample has one endogenous regressor and one IV. lee2021 find that among 123 papers published in AER between 2013 and 2019 that include IV regressions, 61 employ single-IV regressions.\footnote{They point out that the single-IV case “includes applications such as randomized trials with imperfect compliance (estimation of LATE, Imbens_Angrist_1994), fuzzy regression discontinuity designs (see discussion in Lee_Lemieux_2010), and fuzzy regression kink designs (see discussion in Card_et_al_2015)."} Angrist2021one also point out that “most studies using IV (including Angrist(1990) and Angrist-Krueger(1991)) report just-identified IV estimates computed with a single instrument."\footnote{In our empirical application, we revisit the influential study by ADH2013, which also has only one IV.} Assumption (ref)(ii)(1) further allows for the case of single endogenous regressor and multiple IVs.
remAssumption (ref)(ii)(2) is only needed if we have multiple endogenous variables. The condition is similar to that in Canay-Santos-Shaikh(2020), which restricts the type of heterogeneity of the within-cluster Jacobian matrices. However, it is still weaker than restrictions assumed in the literature for cluster-robust Wald tests with few clusters. For example, Bester-Conley-Hansen(2011) and Hwang(2020) provide asymptotic approximations that are based on $t$ and $F$ distributions for the Wald statistics with CRVE. Their conditions require the within-cluster Jacobian matrices to have the same limit for all clusters (i.e., Assumption (ref)(ii)(2) to hold with $a_j=1$ for all $j \in [J]$).\footnote{For example, see Bester-Conley-Hansen(2011) and Hwang(2020) for details.} They also impose that the cluster sizes are approximately equal for all clusters and the cluster-level scores in Assumption (ref)(ii) have the same normal limiting distribution for all clusters, which are not necessary for the wild bootstrap. Finally, Assumption (ref) will not be needed for the bootstrap AR tests in Section (ref), which require neither strong identification nor homogeneity Jacobian matrices. To further clarify our setting, we can relate the Jacobian matrices with the first-stage projection coefficient. Specifically, recall $\Pi_{z,j,n}$ is the coefficient of $Z_{i,j}$ via the cluster-level population projection of $X_{i,j}$ on $Z_{i,j}$ and $W_{i,j}$ as defined in Example (ref). Then, we have $\lim_{n \rightarrow \infty}\Pi_{z,j,n} = \Pi_{z,j} := Q_{\widetilde{Z}\widetilde{Z},j}^{-1} Q_{\widetilde{Z}X,j}$ under our framework. Also define $\Pi_z = Q_{\widetilde{Z}\widetilde{Z}}^{-1} Q_{\widetilde{Z}X}$. Assumption (ref)(i) ensures that overall we have strong identification as $Q_{\widetilde{Z}X}$ (and $\Pi_z$) is of full column rank. Furthermore, we call the clusters in which $Q_{\widetilde{Z}X,j}$ (and $\Pi_{z,j}$) are of full column rank the strong IV clusters, i.e., $\beta_n$ is strongly identified in these clusters. On the other hand, strong identification for $\beta_n$ is not ensured in the rest of the clusters. We can draw three observations from Assumption (ref). First, given the number of clusters is fixed, only one strong IV cluster is needed for Assumption (ref)(i) to hold. Second, Assumption (ref)(ii)(2) implies that when $a_j \neq 0$, $Q_{\widetilde{Z}X,j}$ (and $\Pi_{z,j}$) is of full column rank, so that the $j$-th cluster is a strong IV cluster. Third, Assumption (ref)(ii)(2) excludes the case that $Q_{\widetilde{Z}X,j}$ is of a reduced rank but is not a zero matrix when we have multiple endogenous variables. It is possible to select out the clusters with Jacobian matrices of reduced rank RR00,KP06,CF19. To formally extend these rank tests to the case with a fixed number of large clusters would merit a separate paper and is thus left as a topic for future research.
assumptionSuppose $|| \hat{A}_r - A_r ||_{op} = o_P(1)$, where $A_r$ is a $d_r \times d_r$ symmetric deterministic weighting matrix such that $0 < c \leq \lambda_{min}(A_r) \leq \lambda_{max}(A_r) \leq C < \infty$ for some constants $c$ and $C$, $\lambda_{\min}(A)$ and $\lambda_{\max}(A)$ are the minimum and maximum eigenvalues of the generic matrix $A$, and $||A||_{op}$ denotes the operator norm of the matrix $A$.

Assumption (ref) requires that the weighting matrix $\hat{A}_r$ in ((ref)) has a deterministic limit. It rules out the case that $\hat{A}_r$ equals the inverse of CRVE, which has a random limit under a small number of clusters. We will discuss the bootstrap Wald test with CRVE in Section (ref).

theoremSuppose that Assumptions (ref)-(ref) hold. Then under $\mathcal{H}_0$, for all four estimation methods (namely, TSLS, LIML, FULL, and BA), \begin{eqnarray*} \alpha - \frac{1}{2^{J-1}} \leq \liminf_{n \rightarrow \infty} \mathbb{P} \{ T_{n} > \hat{c}_{n}(1-\alpha) \} \leq \limsup_{n \rightarrow \infty} \mathbb{P} \{ T_{n} > \hat{c}_{n}(1-\alpha) \} \leq \alpha + \frac{1}{2^{J-1}}. \end{eqnarray*}

Two remarks are in order.

remFirst, Theorem (ref) states that as long as there exists at least one strong IV cluster, the $T_n$-based wild bootstrap test has limiting null rejection probability no greater than $\alpha+1/2^{J-1}$ and no smaller than $\alpha - 1/2^{J-1}$. For the test to have power against local alternatives, we later show that at least 5 or 6 strong IV clusters are needed. The error $1/2^{J-1}$ can be viewed as the upper bound for the asymptotic size distortion, which vanishes exponentially with the total number of clusters rather than the number of strong IV clusters. Intuitively, although the weak IV clusters do not contribute to the identification of $\beta_n$, the scores of such clusters still contribute to the limiting distributions of the IV estimators, which in turn determines the total number of sign changes in the bootstrap Wald statistics. We note that $1/2^{J-1}$ equals $1.56\%$ and $0.2\%$ when $J=7$ and $10$, respectively. If such an distortion is still of concern, researchers can replace $\alpha$ in our context by $\alpha - 1/2^{J-1}$ to ensure null rejection rate.
remIt is well known that estimators such as LIML and FULL have reduced finite sample bias relative to TSLS in the over-identified case, especially when the IVs are not strong. Since the validity of the randomization with sign changes requires a distributional symmetry around zero, the LIML and FULL-based bootstrap Wald tests may therefore achieve better finite sample size control than that based on TSLS. This is also confirmed by our simulation experiments.\footnote{To theoretically document the asymptotic bias due to the dimensionality of IVs, one needs to consider an alternative framework in which the number of clusters is fixed but the number of IVs tends to infinity, following the literature on many/many weak instruments Bekker(1994), Chao-Swanson(2005), MS22. We leave this direction of investigation for future research.}

We next examine the power of the wild bootstrap test against local alternatives.

theoremSuppose that Assumptions (ref)-(ref) hold. Further suppose that there exists a subset $\mathcal{J}_s$ of $[J]$ such that $a_j>0$ for each $j \in \mathcal{J}_s$, $a_j = 0$ for $j \in [J] \backslash \mathcal{J}_s$, and $\lceil |\textbf{G}|(1-\alpha) \rceil \leq |\textbf{G}|-2^{J-J_s+1}$, where $|\textbf{G}| = 2^J$, $J_s = |\mathcal{J}_s|$, and $a_j$ is defined in Assumption (ref).\footnote{When $d_x= 1$, we define $a_j = Q^{-1}Q_{\widetilde{Z}X}^\top Q_{\widetilde{Z}\widetilde{Z}}^{-1} Q_{\widetilde{Z}X,j}$.} Then under $\mathcal{H}_{1,n}$ in ((ref)), for all four estimation methods (namely, TSLS, LIML, FULL, and BA), \begin{eqnarray*} \lim_{|| \mu ||_2 \rightarrow \infty} \liminf_{n \rightarrow \infty} \mathbb{P} \{ T_{n} > \hat{c}_{n}(1-\alpha) \} =1. \end{eqnarray*}

Several remarks are in order.

remTo establish the power of the $T_n$-based wild bootstrap test against $r_n$-local alternatives, we need homogeneity of the signs of Jacobians for the strong IV clusters (i.e., $a_j>0$ for each $j \in \mathcal J_s$). For example, in the case with a single IV, it requires $\Pi_{z,j}$ to have the same sign across all the strong IV clusters. We notice that this condition is not needed for the bootstrap Wald test studentized by CRVE described in Section (ref).
remOur test compares the test statistic $T_n$ with the critical value $\hat c_n(1-\alpha)$, where $T_n$ is asymptotically equivalent to $T^*_n (\iota_J)$, i.e., the bootstrap test statistic with $g$ equal to a $J \times 1$ vector of ones, and the critical value $\hat c_n(1-\alpha)$ is just the $\lceil |\textbf{G}|(1-\alpha) \rceil$th order statistic of $\{T^*_n (g)\}_{ g \in \textbf{G}}$. We show that, when $|| \mu ||_2 \rightarrow \infty$ and the signs of $g_j$ for all strong IV clusters are the same (only weak IV clusters have different signs), $T^*_n (g)$ is equivalent to $T^*_n (\iota_J)$, and thus, $T_n$, even under the alternative. Intuitively, as $|| \mu ||_2 \rightarrow \infty$, the asymptotic behaviour of $T_n^*(g)$ can be affected by the sign changes only through the strong IV clusters (the effect of sign changes from the weak IV clusters becomes negligible). Let us denote the set of such $g$'s that only flip the sign of weak IV clusters as $\textbf{G}_w = \{g: g_j = g_{j'}, \forall j, j' \in \mathcal J_s\}$. Given there are $J_s$ strong IV clusters, the cardinality of $\textbf{G}_w$ is $2^{J-J_s+1}$. To establish the power against $\mathcal{H}_{1,n}$ in Theorem (ref), we request that our bootstrap critical value $\hat c_n(1-\alpha)$ does not take values of $T^*_n(g)$ for $g \in \textbf{G}_w$ because otherwise the test statistic $T_n$ and the critical value are asymptotically equivalent even under the alternative. This implies the inequality that \begin{align*} \lceil |G|(1-\alpha) \rceil \leq |G|-2^{J-J_s+1}. \end{align*} Therefore, we need a sufficient number of strong IV clusters to establish the power result. For instance, the condition $\lceil |\textbf{G}|(1-\alpha) \rceil \leq |\textbf{G}|-2^{J-J_s+1}$ requires that $J_s \geq 5$ and $J_s \geq 6$ for $\alpha=10\%$ and $5\%$, respectively. Theorem (ref) suggests that although the size of the wild bootstrap test is well controlled even with only one strong IV cluster, its power depends on the number of strong IV clusters.
remThe wild bootstrap test has resemblance to the group-based $t$-test in IM and the randomization test with sign changes in CRS. IM and CRS approaches separately estimate the parameters using the samples in each cluster (say, $\hat{\beta}_{1}, ..., \hat{\beta}_J$), and therefore requires $\beta_n$ to be strongly identified in all clusters. In contrast, our method allows for some but not all clusters to have weak IVs in the sense of Staiger-Stock(1997), where $\Pi_{z,j,n}$ has the same order of magnitude as $n^{-1/2}$.\footnote{The cluster-level IV estimators of such weak clusters would become inconsistent and have highly nonstandard limiting distributions.} Also, if there exist both strong and “semi-strong" IV clusters, in which the (unknown) convergence rates of IV estimators can vary among clusters Andrews-Cheng(2012), then the estimators with the slowest convergence rate\footnote{The clusters corresponding to these estimators will become effective (dominant) clusters among the $J$ clusters.} will dominate in the test statistics that are based on the cluster-level estimators. In such cases, IM and CRS can become invalid while our bootstrap test remains valid. On the other hand, if $\beta_n$ is strongly identified in all clusters and the cluster-level IV estimators have minimal finite sample bias, IM and CRS have an advantage over the wild bootstrap when there are multiple endogenous variables as they do not require Assumption (ref)(ii). Therefore, the two types of approaches could be considered as complements, and practitioners may choose between them according to the characteristics of their data and models.

Asymptotic Results for the Wald Tests with CRVE

Now we consider a wild bootstrap test for the Wald statistic studentized by CRVE.

theoremSuppose that Assumptions (ref)-(ref) hold, and $J> d_r$. Then under $\mathcal{H}_0$, for all four estimation methods (namely, TSLS, LIML, FULL, and BA), \begin{align*} \alpha - \frac{1}{2^{J-1}} & \leq \liminf_{n \rightarrow \infty} \mathbb{P} \{ T_{CR,n} > \hat{c}_{CR,n}(1-\alpha) \} \\ & \leq \limsup_{n \rightarrow \infty} \mathbb{P} \{ T_{CR,n} > \hat{c}_{CR,n}(1-\alpha) \} \leq \alpha + \frac{1}{2^{J-1}}. \end{align*}

We require $J>d_r$ because otherwise CRVE and its bootstrap counterpart are not invertible. Theorem (ref) states that with at least one strong IV cluster, the $T_{CR,n}$-based wild bootstrap test controls size asymptotically up to a small error. Next, we turn to the local power.

theorem(i) Suppose that Assumptions (ref)-(ref) hold, $J-1> d_r$, and that there exists a subset $\mathcal{J}_s$ of $[J]$ such that $\min_{j \in \mathcal J_s} |a_j| > 0$, $a_j=0$ for each $j \in [J] \backslash \mathcal{J}_s$, and $\lceil |\textbf{G}|(1-\alpha) \rceil \leq |\textbf{G}|-2^{J-J_s+1}$, where $|\textbf{G}| = 2^J$, $J_s = |\mathcal{J}_s|$, and $a_j$ is defined in Assumption (ref).\footnote{When $d_x= 1$, we define $a_j = Q^{-1}Q_{\widetilde{Z}X}^\top Q_{\widetilde{Z}\widetilde{Z}}^{-1} Q_{\widetilde{Z}X,j}$.} Then under $\mathcal{H}_{1,n}$ in ((ref)), for all four estimation methods (namely, TSLS, LIML, FULL, and BA), \begin{eqnarray*} \lim_{|| \mu ||_2 \rightarrow \infty} \liminf_{n \rightarrow \infty} \mathbb{P} \{ T_{CR,n} > \hat{c}_{CR,n}(1-\alpha) \} =1. \end{eqnarray*} (ii) Further suppose that $d_r=1$. Then under $\mathcal{H}_{1,n}$ in ((ref)), for any $\delta>0$, there exists a constant $c_{\mu}>0$ such that when $|\mu| >c_\mu$, \begin{align*} \liminf_{n \rightarrow \infty}\mathbb{P}(\phi^{cr}_n \geq \phi_n) \geq 1-\delta, \end{align*} where $\phi^{cr}_n = 1 \{ T_{CR,n} > \hat{c}_{CR,n}(1-\alpha)\}$ and $\phi_n = 1 \{ T_{n} > \hat{c}_{n}(1-\alpha) \}$.
remDifferent from Theorem (ref), the power result in Theorem (ref) does not require the homogeneity condition on the sign of first-stage coefficients for the strong IV clusters (i.e., it only requires $\min_{j \in \mathcal J_s} |a_j|>0$). Therefore, the bootstrap Wald test studentized by CRVE has an advantage over its unstudentized counterpart, given that the homogeneous sign condition may not hold in some empirical studies (e.g., Figure (ref)). To study the local power without the sign condition, we rely on arguments very different from those in Canay-Santos-Shaikh(2020).
remWe further establish in Theorem (ref)(ii) that in the case of a $t$-test (i.e., when the null hypothesis involves one restriction), the rejection of the $T_{CR,n}$-based bootstrap test dominates that based on $T_n$ with a large probability when the local parameter $\mu$ is sufficiently different from zero. To see this, we note that, when there is just one restriction, $ 1\{T_n > \hat{c}_{n}(1-\alpha)\} = 1\{T_{CR,n} > \tilde{c}_{CR,n}(1-\alpha)\},$ where $\tilde{c}_{CR,n}(1-\alpha)$ denotes the $(1-\alpha)$ quantile of $\left\{|(\lambda_{\beta}^{\top}\hat \beta_{L,g}^* - \lambda_0)|/\sqrt{\lambda_{\beta}^{\top}\widehat{V}\lambda_{\beta}}: g \in \textbf{G}\right\}.$ Then, Theorem (ref)(ii) follows because $\tilde{c}_{CR,n}(1-\alpha) > \hat{c}_{CR,n}(1-\alpha)$ with large probability as $||\mu||_2$ becomes sufficiently large.

Asymptotic Results for Weak-instrument-robust Tests

The size control of the bootstrap Wald tests with or without CRVE relies on Assumption (ref)(i), which rules out overall weak identification in which all clusters are weak. In the case that the parameter of interest may be weakly identified in all clusters, we can consider the bootstrap AR tests for the full vector of $\beta_n$ as defined in Section (ref).

assumption$|| \hat{A}_z - A_z ||_{op} = o_P(1)$, where $A_z$ is a $d_z \times d_z$ symmetric deterministic weighting matrix such that $0 < c \leq \lambda_{min}(A_z) \leq \lambda_{max}(A_z) \leq C < \infty$ for some constants $c$ and $C$.

Theorem (ref) below shows that, in the general case with multiple IVs, the limiting null rejection probability of the $AR_{n}$-based bootstrap test does not exceed the nominal level $\alpha$, and that of the $AR_{CR,n}$ test does not exceed $\alpha$ by more than $1/2^{J-1}$ when $J>d_z$, irrespective of IV strength.

theoremSuppose Assumption (ref) holds and $\beta_n=\beta_0$. For $AR_n$, further suppose Assumption (ref) holds. For $AR_{CR,n}$, further suppose $J> d_z$. Then, \begin{align*} \alpha - \frac{1}{2^{J-1}} & \leq \liminf_{n \rightarrow \infty} \mathbb{P} \{ AR_{n} > \hat{c}_{AR,n}(1-\alpha) \} \notag \\ & \leq \limsup_{n \rightarrow \infty} \mathbb{P} \{ AR_{n} > \hat{c}_{AR,n}(1-\alpha) \} \leq \alpha, \; and \notag \\ \alpha - \frac{1}{2^{J-1}} & \leq \liminf_{n \rightarrow \infty} \mathbb{P} \{ AR_{CR,n} > \hat{c}_{AR,CR,n}(1-\alpha) \} \\ & \leq \limsup_{n \rightarrow \infty} \mathbb{P} \{ AR_{CR,n} > \hat{c}_{AR,CR,n}(1-\alpha) \} \leq \alpha + \frac{1}{2^{J-1}}. \end{align*}
remFirst, for the bootstrap AR test studentized by the CRVE, we require the number of IVs to be smaller than the number of clusters because otherwise, the CRVE is not invertible. Second, the behavior of wild bootstrap for other weak-IV-robust statistics proposed in the literature is more complicated as they depend on an adjusted sample Jacobian matrix (e.g., see Kleibergen(2005), Andrews(2016), Andrews-Mikusheva(2016), and Andrews-Guggenberger(2019), among others). We study the asymptotic properties of wild bootstrap for these statistics in Section (ref) in the Online Supplement and show that their validity would require at least one strong IV cluster. This provides theoretical support to the findings in the literature that the bootstrap AR test has better finite sample size properties than other bootstrap weak-IV-robust tests for clustered data (e.g., see finlay2014, Finlay-Magnusson(2019)). Third, for the weak-IV-robust subvector inference, one may use a projection approach Dufour-Taamouti(2005) after implementing the wild bootstrap AR tests for $\beta_n$, but the result may be conservative.\footnote{ Alternative subvector inference methods (e.g., see Section 5.3 in Andrews-Stock-Sun(2019) and the references therein) provide a power improvement over the projection approach with a large number of observations/clusters. However, it is unclear whether they can be applied to the current setting. It is unknown whether the asymptotic critical values given by these approaches will still be valid with a small number of clusters. Also, Wang-Doko(2018) show that bootstrap tests based on the subvector statistics therein may not be robust to weak IVs even under conditional homoskedasticity.}

Next, to study the power of the $AR_n$-based bootstrap test against the local alternative, we let $\lambda_{\beta} = I_{d_x}$ for $\mathcal{H}_{1,n}$ in ((ref)) so that $\mu = \mu_{\beta}$ and impose the following condition.

assumption(i) $Q_{\widetilde{Z}X} \neq 0$. (ii) There exists a scalar $a_j$ for each $j \in [J]$ such that $Q_{\widetilde{Z}X,j} = a_jQ_{\widetilde{Z}X}$.

If $d_x = d_z = 1$, Assumption (ref)(i) implies strong identification. When $d_z>1$ or $d_x >1$, Assumption (ref)(i) only rules out the case that $Q_{\widetilde{Z}X}$ is a zero matrix, while allowing it to be nonzero but not of full column rank. As noted below, this means the AR test can still have local power in some direction even without strong identification. Assumption (ref)(ii) is similar to Assumption (ref)(ii). In particular, it holds automatically if Assumption (ref)(i) holds and $d_x= d_z=1$ (i.e., single endogenous regressor and single IV).

theoremSuppose Assumptions (ref), (ref), and (ref) hold. Further suppose that there exists a subset $\mathcal{J}_s$ of $[J]$ such that $\min_{j \in \mathcal J_s} a_j > 0$, $a_j=0$ for each $j \in [J] \backslash \mathcal{J}_s$, and $\lceil |\textbf{G}|(1-\alpha) \rceil \leq |\textbf{G}|-2^{J-J_s+1}$, where $J_s = |\mathcal{J}_s|$ and $a_j$ is defined in Assumption (ref). Then, under $\mathcal{H}_{1,n}$ with $\lambda_{\beta} = I_{d_x}$, \begin{eqnarray*} \lim_{|| Q_{\widetilde{Z}X}\mu ||_2 \rightarrow \infty} \liminf_{n \rightarrow \infty} \mathbb{P} \{ AR_{n} > \hat{c}_{AR,n}(1-\alpha) \} =1. \end{eqnarray*}
remNotice that the bootstrap AR test not studentized by CRVE has power against local alternatives as long as $||Q_{\widetilde{Z}X}\mu||_2 \rightarrow \infty$, which may hold even when $\beta_n$ is not strongly identified and $Q_{\widetilde{Z}X}$ is not of full column rank.
remWhen $d_z = 1$, we have $1\{AR_n > \hat{c}_{AR,n}(1-\alpha) \} = 1\{AR_{CR,n} > \hat{c}_{AR,CR,n}(1-\alpha) \}$, which implies the $AR_n$ and $AR_{CR,n}$-based bootstrap tests have the same power against local alternatives. However, such a power equivalence does not hold when $d_z>1$. Unlike the Wald test with CRVE in Section (ref), we cannot establish the power of the AR statistics with CRVE for the general case due to the fact that the CRVE for the Wald test is computed using the estimated residual $\hat{\varepsilon}_{i,j}$ whereas the CRVE for the AR test is computed by imposing the null (see Section (ref)). Consequently, unlike the Wald test, the AR test with CRVE is not theoretically guaranteed to have better power than the AR test without CRVE. Indeed, we observe in Section (ref) that the $AR_{CR,n}$-based bootstrap test has inferior finite sample power properties compared with its $AR_n$-based counterpart.
remIn Section (ref) in the Online Supplement, we further show that for the common case with a single endogenous variable and a single IV, a wild bootstrap test based on the unstudentized Wald statistic is asymptotically equivalent to a certain wild bootstrap AR test under both null and alternative. Second, we show that bootstrapping the Lagrange multiplier test or conditional quasi-likelihood ratio test controls asymptotic size when there is at least one strong IV cluster.

Monte Carlo Simulations

In this section, we investigate the finite sample performance of the wild bootstrap tests and alternative inference methods. We consider four different data generating processes (DGPs) with four dependence structures, namely cluster fixed effect (DGP 1), network dependence (DGP 2), serial dependence (DGP 3), and spatial dependence (DGP 4). For all four DGPs, the number of bootstrap replications equals 399. For DGPs 1, 3, and 4, the number of Monte Carlo replications is 50,000. The second DGP requires implementing spectral clustering for each Monte Carlo replication, which is computationally heavy. For this reason, we set the number of replications as 10,000. The nominal level $\alpha$ is set at 10%. In addition, following the recommendation in Djogbenou-Mackinnon-Nielsen(2019) and mackinnon2022cluster, we conduct cluster-level demean for all DGPs. For the conciseness of the paper, we only report the results for DGPs 1 and 2 in the paper and relegate the results for DGPs 3 and 4 to Section (ref) in the Online Supplement. The overall patterns observed from DGPs 3 and 4 are very similar to those from DGPs 1 and 2.

Simulation Designs

DGP 1. The first DGP is similar to that in Section IV of Canay-Santos-Shaikh(2020) and extend theirs to the IV model. The data are generated as

eqnarray*[eqnarray* omitted — 228 chars of source]

for $i=1, ..., n$ and $j=1, ..., J$. We set the number of clusters $J \in \{6,9,12,20\}$. The total number of observations $n$ is equal to 500. To allow for unbalanced clusters, we follow Djogbenou-Mackinnon-Nielsen(2019) and Mackinnon2021 and set the cluster sizes as

align*[align* omitted — 143 chars of source]

and $n_J = n - \sum_{j \in [j]} n_j$. We let $r=4$ to generate substantial heterogeneity in cluster sizes.\footnote{$r=4$ corresponds to the highest value of the heterogeneity parameter considered in the simulations of Djogbenou-Mackinnon-Nielsen(2019) and mackinnon2022cluster.} For example, when $J=6$, the cluster sizes are $8,17,33,65,127,$ and $250$, respectively.\footnote{When $J=20$, the cluster sizes are $2,2,3,3,4,5,6,8,10,12,15,18,22,27,33,41,50,61,75$, and $103$, respectively. We use this setting to investigate the robustness of the inference methods under a scenario that is quite different from the fixed-$J$ asymptotic framework. In addition, when $J=20$ and $d_z=3$, we combine the two clusters with $n_j=2$ and combine the two clusters with $n_j=3$, so that IM and CRS can also be implemented.} In addition, $(\varepsilon_{i,j}, v_{i,j})$, $(a_{\varepsilon,j}, a_{v,j})$, $Z_{i,j}$, and $\sigma(Z_{i,j})$ are specified as follows:

align*[align* omitted — 464 chars of source]

where $I_{d_z}$ is a $d_z \times d_z$ identity matrix and $Z_{i,j,k}$ denotes the $k$-th element of $Z_{i,j}$.\footnote{As mentioned above, we project out the cluster fixed effects. Specifically, we can define $\widetilde{Z}_{i,j} = Z_{i,j} - \frac{1}{n_j}\sum_{i \in I_{n,j}}Z_{i,j}$. Note that $\frac{1}{\sqrt{n_j}}\sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}\sigma(Z_{i,j})(a_{\varepsilon,j}+\varepsilon_{i,j})$ is asymptotically normal conditional on $a_{\varepsilon,j}$, as $n_j$ diverges to infinity. This is because $\left(\frac{1}{\sqrt{n_j}}\sum_{i \in I_{n,j}}Z_{i,j}\left(\sum_{k=1}^{d_z}Z_{i,j,k}^2\right), \left[\frac{1}{\sqrt{n_j}}\sum_{i \in I_{n,j}}Z_{i,j}\right] \left[\frac{1}{n_j} \sum_{i \in I_{n,j}} \left(\sum_{k=1}^{d_z}Z_{i,j,k}^2\right)\right], \frac{1}{\sqrt{n_j}}\sum_{i \in I_{n,j}}Z_{i,j}\left(\sum_{k=1}^{d_z}Z_{i,j,k}^2\right)\varepsilon_{i,j}\right)$ are jointly asymptotically normal given $(a_{\varepsilon,j})_{j \in [J]}$. Then, we have

align*[align* omitted — 528 chars of source]

which is asymptotically normal given $(a_{\varepsilon,j})_{j \in [J]}$ as the cluster size diverges to infinity, so that Assumption (ref)(ii) is satisfied.} The number of IVs is set to be $d_z \in \{1,3\}$, and $\rho_{\varepsilon v} \in \{0.3, 0.5, 0.7\}$ corresponds to the degree of endogeneity. In addition, we introduce cluster heterogeneity to the $d_z \times 1$ first-stage coefficients $\Pi_{z,j}$ by letting $\Pi_{z,j} = (\Pi/2, ..., \Pi/2)^{\top}$ for $1\leq j \leq J/3$, $\Pi_{z,j} = (\Pi, ..., \Pi)^{\top}$ for $J/3 < j \leq 2J/3$, and $\Pi_{z,j} = (2\Pi, ..., 2\Pi)^{\top}$ for $2J/3 < j \leq J$, where $\Pi \in \{0.25, 0.375, 0.5\}$. Such a DGP satisfies our assumptions and allows for heterogeneity in both cluster sizes and IV strengths. The values of $\beta$ and $\gamma$ are both set to $1$. In Section (ref) in the Online Supplement, we further report the simulation results for the case where $\Pi$ remains the same across clusters, so that the cluster-level heterogeneity in identification strength originates solely from the heterogeneity in cluster size. We find that the overall patterns remain very similar to those reported in Section (ref) below.

DGP 2. We consider the linear-in-mean social interaction model detailed in Example (ref). Specifically, following L22, we generate the network $\mathcal A$ with $n=500$ nodes as $\mathcal A_{i,j} = 1\{||\eta_{i} - \eta_{j}||_2 \geq (7/(\pi n))^{1/2}\}$, where $\eta_i \stackrel{i.i.d.}{\sim} \text{Uniform}[0,1]^2$. Following BDF09, we generate data according to (ref) in which $\varepsilon_i \stackrel{i.i.d.}{\sim} \mathcal{N}(0,0.1)$,

align*[align* omitted — 171 chars of source]

where $\varepsilon_i' \stackrel{i.i.d.}{\sim}\mathcal{N}(0,1)$ and $(\varepsilon_i,\varepsilon_i',\eta_i)$ are independent. The observables $(y_i,X_i,W_i,Z_i)$ are as defined in Example (ref) with $d_z=2$. According the calibration study by BDF09, we set the parameters $(\alpha,\beta,\delta) = (0.7683,0.4666,0.1507)$. Because the IV strength scales with $\gamma \beta + \delta$, we study the performance of our method under various identification strength by setting this value to be $(0.11, 0.1498, 0.1896)$, which correspond to values $(-0.0872, -0.0019, 0.0834)$ for $\gamma$. For comparison, we note that BDF09 set $\gamma = 0.0834$. We follow the classification procedure proposed by L22 to obtain the clusters.

enumerate• Input: a positive integer $L$ and network $\mathcal A$. • Compute all separated components (no links between two components) of the network denoted as $\{\mathcal V_{h}\}_{h \in \{0\} \cup [H]}$, where $\{\mathcal V_{h}\}_{h \in \{0\} \cup [H]}$ is a partition of $[n]$ (n vertexes) and they are sorted in ascending order according to their sizes. We keep all the components whose sizes are greater than 5. Denote the number of components left as $L'+1$ for some $L' \geq 0$. • Suppose the biggest component $\mathcal V_0$ has size $\tilde n_0$. By permuting labels, we suppose $\mathcal V_0 = [\tilde n_0]$ and denote its adjacency matrix as $\mathcal A_0$. Then, we compute the graph Laplacian as \begin{align*} \mathcal L_0 = I_{\tilde n_0} - D_0^{-1/2} \mathcal A_0 D_0^{-1/2}, \end{align*} where $D_0 = \text{diag}(\sum_{i \in [\tilde n_0]} A_{0,1,i},\cdots,\sum_{i \in [\tilde n_0]} A_{0,\tilde n_0,i})$ is an $\tilde n_0 \times \tilde n_0$ diagonal matrix of degrees. • Obtain the top $L$ eigenvector matrix of $\mathcal L_0$ corresponding to its $L$ largest eigenvalues and denote it as $ V = [V_1^\top,\cdots,V_{\tilde n_0}^\top]^\top,$ where $V_{i} \in \Re^{L}$ for $i \in [\tilde n_0]$. • Apply K-means algorithm to $V$ and divide $\mathcal V_0 = [\tilde n_0]$ into $L$ groups, denoted as $\mathcal V_{0,1},\cdots,\mathcal V_{0,L}$. • Output: we obtain $J = L+L'$ clusters $(\mathcal V_{0,1},\cdots,\mathcal V_{0,L},\mathcal V_{1},\cdots,\mathcal V_{L'})$.

We set $n=500$ and let $L$ be $[5,10,20,30]$. The number of clusters ($J$) depends on $L'$ which varies across simulation replications. We report the average $J$ below.

Inference Methods

We investigate the finite sample performance of the following inference methods.

enumerate• IM: This inference method is proposed by Ibragimov-Muller(2010), which compares their group-based $t$-test statistic with the critical value of a $t$-distribution with $J-1$ degrees of freedom. The test statistic is constructed by separately running the IV regression using the samples in each cluster. • CRS: The approximate randomization test proposed by Canay-Romano-Shaikh(2017), which compares IM's test statistic with the critical value of the sign changes-based randomization distribution of the statistic. • ASY: Compare $T_{CR,n}$, the Wald statistic studentized by CRVE as described in Section (ref), with a standard normal critical value (as $r=1$ here). This is the standard Wald test. • BCH: This inference method is proposed by Bester-Conley-Hansen(2011), which compares $T_{CR,n}$ with $\sqrt{\frac{J}{J-1}}$ times the critical value of a $t$-distribution with $J-1$ degrees of freedom. • WRE: Compare the studentized Wald statistic $T_{CR,n}$ with the critical value generated by the WRE cluster (WREC) bootstrap procedure, as described in Davidson-Mackinnon(2010), Finlay-Magnusson(2019), Roodman-Nielsen-MacKinnon-Webb(2019), and Mackinnon2021. • W-B: Compare the unstudentized Wald statistic $T_n$, which sets the weighting matrix $\hat A_r = 1$ in ((ref)), with the wild bootstrap critical value $\hat c_n$ as described in Section (ref). Therefore, this bootstrap procedure does not involve the CRVE-based studentization. • W-B-S: Compare the studentized Wald statistic $T_{CR,n}$ with the wild bootstrap critical value $\hat c_{CR,n}$ as described in Section (ref). Its pseudo code is provided in Section (ref) in the Online Supplement. • AR-ASY: Compare $AR_{CR,n}$ with $\chi^2(d_z)$ critical value, where $AR_{CR,n}$ is the AR statistic studentized by CRVE and $d_z$ is the number of instruments. • AR-B: Compare the unstudentized AR statistic $AR_n$, which sets the weighting matrix $\hat A_z = I_{d_z}$ in ((ref)), with the wild bootstrap critical value $\hat c_{AR,n}$ as described in Section (ref). • AR-B-S: Compare $AR_{CR,n}$ with the wild bootstrap critical value $\hat c_{AR,CR,n}$ as described in Section (ref).

Except for the methods AR-ASY, AR-B, and AR-B-S, all other methods require the estimation of parameter $\beta$. All four k-class estimators mentioned in the paper can be used. Here, we report the results regarding the TSLS estimator for simulations with one IV and report those regarding the FULL estimator for simulations with multiple IVs, as it is known that FULL has reduced finite sample bias than TSLS in the over-identified case.\footnote{The performance of LIML is very similar to FULL in our simulations. We therefore omit the LIML results but they are available upon request.}

Simulation Results

DGP 1. Tables (ref)-(ref) report the null empirical rejection frequencies of the ten inference methods described above with $J=6$ or $12$. For succinctness, we report the results for $J=9$ or $20$ in Section (ref) of the Online Supplement. Also as mentioned above, the results of the Wald tests are based on TSLS when $d_z=1$ and based on FULL when $d_z=3$, respectively. Following the recommendation in the literature, we set the tuning parameter of FULL to be 1.\footnote{In this case FULL is best unbiased to a second order among $k$-class estimators under normal errors Rothenberg(1984).} Several observations are in order.

enumerate• In general, size distortions increase when the first-stage coefficient $\Pi$ become small, the degree of endogeneity $\rho_{\varepsilon v}$ becomes high, or the IV model becomes over-identified ($d_z=3$). • IM and CRS tests can have considerable over-rejections when the number of clusters is relatively large. For example, when $J=12$, the maximum null empirical rejection frequencies for IM are 0.185 in Table (ref) and 0.209 in Table (ref), while the maximum rejection frequencies for CRS are 0.137 and 0.242, respectively. Similarly, when $J=20$, the maximum null rejection frequencies for IM are 0.209 in Table (ref) and 0.286 in Table (ref), while those for CRS are 0.142 and 0.302, respectively, in Section (ref) of the Online Supplement. • ASY and BCH tests typically have large size distortions when the number of clusters is relatively small. For example, when $J=6$ and $\Pi = 0.25$, for $\rho_{\varepsilon v}=0.3, 0.5,$ and $0.7$, the null empirical rejection frequencies of ASY are 0.211, 0.219, and 0.228, respectively, in Table (ref) and 0.254, 0.258, and 0.265, respectively, in Table (ref). BCH improves upon ASY, but still has corresponding rejection frequencies equal to 0.147, 0.154, and 0.168, respectively, in Table (ref), and 0.180, 0.185, and 0.193, respectively, in Table (ref). • Among the Wald-based inference methods (IM, CRS, ASY, BCH, WRE, W-B, and W-B-S), only the three wild bootstrap tests (WRE, W-B, and W-B-S) have good size control across different settings of $J$, $\Pi$, $\rho_{\varepsilon v}$, and $d_z$. • All the three AR-based inference methods control the size across different settings, but AR-ASY, which is based on the chi-squared critical values, can be rather conservative in the over-identified case. In particular, in Table (ref), it does not reject the null when $J=6$ and has very low rejection frequencies when $J=9$ in Table (ref) in the Online Supplement.\footnote{We notice that the null rejection probabilities of the $AR_{CR,n}$-based asymptotic test decrease toward zero when $d_z$ approaches $J$. If $d_z$ is equal to $J$, the value of $AR_{CR,n}$ will be exactly equal to $d_z$ (or $J$), and thus has no variation (for $\bar{f} = \left(\bar{f}_{1}, ..., \bar{f}_J \right)^{\top}$ and $\bar{f}_{j} = n^{-1} \sum_{i \in I_{n,j}} f_{i,j},$ $ AR_{CR,n} = \iota_J^{\top} \bar{f} \left( \bar{f}^{\top} \bar{f} \right)^{-1} \bar{f}^{\top} \iota_J = \iota_J^{\top} \iota_J = d_z $ as long as $\bar{f}$ is invertible, where $\iota_J$ denotes a $J$-dimensional vector of ones). By contrast, the $AR_{n}$-based bootstrap test works well even when $d_z$ is larger than $J$.}

Furthermore, for power comparisons we focus on the inference methods that control size throughout the above simulations (namely, WRE, W-B, W-B-S, AR-ASY, AR-B, and AR-B-S). The true value of $\beta$ is set equal to 1 and we vary the value of $\beta_0-\beta$ from $-3$ to $3$. $\rho_{\varepsilon v}$ is set equal to $0.5$. Figures (ref) and (ref) show the power curves of the bootstrap Wald tests and the AR tests, respectively, with $d_z=1$ and $J=6$ or $12$. Figures (ref) and (ref) in the Online Supplement show the power curves with $d_z=1$ and $J=9$ or $20$. In addition, Figures (ref)--(ref) report those for $d_z=3$. We highlight several observations below.

enumerate• In general, all the tests become more powerful when the IV strength becomes stronger and/or the number of clusters becomes larger. • According to Figures (ref), (ref), (ref), and (ref), when the alternative is sufficiently distant, W-B-S (the bootstrap Wald test studentized by CRVE) is more powerful than W-B, especially when the identification is relatively weak and/or the number of clusters is small. This is in line with our theoretical result in Theorem (ref)(ii) of Section (ref). • In Figures (ref), (ref), (ref), and (ref), the power curves of WRE are between those of W-B-S and W-B in most cases and highest for certain alternatives, but they may decrease when the alternative becomes more distant, especially in the cases with relatively weak IVs and few clusters. • Figures (ref), (ref), (ref), and (ref) show that among the AR tests, AR-B (the wild bootstrap AR tests without CRVE studentization) has the highest power, followed by AR-B-S.\footnote{In the case with one IV ($d_z=1$), AR-B and AR-B-S are equivalent so that their power curves coincide with each other.} AR-ASY has the lowest power among the three, which is in line with the under-rejections observed in the size results. • Comparing Figures (ref) and (ref), we find that overall W-B-S has the best power properties among these tests. Similar observation can be made by comparing the corresponding figures of bootstrap Wald and AR tests in the Online Supplement.

DGP 2. Table (ref) reports the size for the network data. Note the identification strength increases with $\gamma$. Similar to DGP 1, we find that ASY and BCH does not control size when the number of clusters is small, while IM and CRS have large size distortions when the number of clusters is large so that within each cluster, the number of observations is not sufficiently large to render asymptotic normality. Instead, the wild bootstrap-based inference methods (WRE, W-B, W-B-S, AR-B, and AR-B-S) have good size control regardless of the number of clusters. In Section (ref), we further report the power results for DGP 2. Consistent with DGP 1 and other simulation designs, we find that our W-B-S has the best power, especially against distant alternatives.

table[table omitted — 6,970 chars of source]
table[table omitted — 6,937 chars of source]
figure[figure omitted — 9,655 chars of source]
table[table omitted — 6,505 chars of source]

Empirical Application

In an influential study, ADH2013 analyze the effect of rising Chinese import competition on US local labor markets between 1990 and 2007, when the share of total US spending on Chinese goods increased substantially from 0.6% to 4.6%. The dataset of ADH2013 includes 722 commuting zones (CZs) that cover the entire mainland US. In this section, we further analyze the region-wise effects of such import exposure by applying IV regression with the proposed wild bootstrap procedures to three Census Bureau-designated regions: South, Midwest and West, with 16, 12, and 11 states, respectively, in each region.\footnote{The Northeast region is not included in the study because of the relatively small number of states ($9$) and small number of CZs in each state (e.g., Connecticut and Rhode Island have only 2 CZs).}

Following ADH2013, we consider the linear IV model

align[align omitted — 81 chars of source]

where the outcome variable $y_{i,j}$ denote the decadal change in average individual log weekly wage in a given CZ and the endogenous variable $X_{i,j}$ is the change in Chinese import exposure per worker in a CZ, where imports are apportioned to the CZ according to its share of national industry employment. The Bartik-type instrument $Z_{i,j}$ (e.g., see Goldsmith(2020)) proposed by ADH2013 is Chinese import growth in other high-income countries, where imports are apportioned to the CZs in the same way as that for $X_{i.j}$.\footnote{See Sections I.B and III.A in ADH2013 for a detailed definition of these variables.} As can be seen from Figure (ref), the predictive power of the IV varies considerably among states. The exogenous variables $W_{i,j}$ include the characteristic variables of CZs and decade specified in ADH2013 as well as state dummies. Our regressions are based on the CZ samples in each region, and the samples are clustered at the state level, following ADH2013. Besides the results for the full sample, we also report those for female and male samples separately. Tables (ref)--(ref) in the Online Supplement show the values of $\widehat{Q}_{\tilde{Z}W,j}$ for each cluster in the three regions, and all of them are close to zero. ((ref)) in Assumption (ref)(i) is thus plausible in this application.

The main result of the IV regression for the three regions is given in Table (ref), with the number of observations ($n$) and clusters ($J$) for each region. As we have only one endogenous variable and one IV, the TSLS estimator is used throughout this section. We further report the conventional 90% confidence intervals based on CRVE with the standard normal critical value (ASY) and the 90% bootstrap confidence sets (CSs) by inverting the corresponding $T_n$ (W-B), $T_{CR,n}$ (W-B-S), and $AR_n$ (AR-B) bootstrap tests with a 10% nominal level.\footnote{AR-B and AR-B-S are numerically equivalent with one instrument.} The computation of the bootstrap CSs is conducted over the parameter space $[-10,10]$ with a step size of 0.01, and the number of bootstrap draws is set at 2,000 for each step. We highlight the main findings below.

enumerate• Table (ref) reports that the TSLS estimates of the average effect of Chinese imports on wages equal $-0.97$ and $-1.05$ for the South and West regions, respectively, while equals $-0.025$ for the Midwest region.\footnote{That is, a \$$1,000$ per worker increase in a CZ's exposure to Chinese imports is estimated to reduce average weekly earnings by $0.97$, $1.05$, and $0.025$ log points, respectively, for the three regions (the corresponding TSLS estimate in ADH2013 for the entire mainland US is $-0.76$).} In addition, we find that the values of the effective first-stage F statistics Olea-Pflueger(2013) are 85.26, 8.69, and 63.45 for South, Midwest, and West, respectively, suggesting the identification is relatively weak for Midwest compared with the other two regions.\footnote{We used CRVE as the variance estimator of the effective first-stage F statistic. We also note that the asymptotic results and critical values for the effective F statistic in Olea-Pflueger(2013) are based on the consistency of the variance estimator. In the case of CRVE, this would require the number of clusters to diverge to infinity. Therefore, when the number of available clusters is small, the critical values provided by Olea-Pflueger(2013) might not have satisfactory performance. Still, we can learn from the values of the effective F statistics that the identification for Midwest may be relatively weak.} • All the three types of bootstrap CSs (W-B, W-B-S, and AR-B) are longer than the conventional asymptotic CSs (ASY) for all cases in Table (ref). For example, the studentized Wald bootstrap CSs (W-B-S) are $22.3\%$, $26.2\%$, and $26.4\%$ longer than the ASY CSs for South, Midwest, and West, respectively, in the all-sample case. Similar differences in the CS length can also be observed for female and male samples, respectively. This is in line with the simulation results in Section (ref), suggesting that conventional CSs can be too short in such cases and result in empirically relevant under-coverage.\footnote{We also computed CSs with BCH's critical values, which are based on $t$ distributions instead of the standard normal distribution, and find that the three types of bootstrap CSs are longer than the BCH CSs in all cases.} • Only the effect in the South region is significantly different from zero at the 10% level under all three bootstrap CSs, while the effects in the other two regions are not. This is in line with the result of effective F statistics, which suggests that the identification is strongest for South. • We note that the effect on West is only significant under the studentized bootstrap Wald CS (W-B-S, $[-1.22, -0.47]$). In addition, W-B-S CSs have the shortest length among the three types of bootstrap CSs in most cases in Table (ref), which is in line with our power results in Sections (ref) and (ref). • Comparing the results for female and male samples, we find that across all the regions, the effects are more substantial for the male samples. Furthermore, the effects for both female and male samples in the South are significantly different from zero.
table[table omitted — 1,562 chars of source]

Conclusion

In many empirical applications of IV regressions with clustered data, the IVs may be weak for some clusters and the number of clusters may be small. In this paper, we provide valid inference methods in this context. For the Wald tests with and without CRVE, we extend the WRE cluster bootstrap procedure to allow for the setting of few clusters and cluster-level heterogeneity in IV strength. For the full-vector inference, we develop wild bootstrap AR tests that control size asymptotically irrespective of IV strength, and also show that the bootstrap validity of other weak-IV-robust tests requires at least one strong IV cluster. Furthermore, we establish the power properties of the Wald and AR tests. Finally, recent studies by hansen2022jackknife and mackinnon2022leverage suggest that asymptotic or bootstrap inference based on jackknife variance estimators can be more reliable than that based on the conventional CRVE for OLS models. Therefore, it may be interesting to consider extending jackknife inference to the current setting. We leave this line of investigation for future research.

\singlespacing

Pseudo-Code for the W-B-S Procedure

This section provides the pseudo-code for our recommended wild bootstrap inference procedure W-B-S. We emphasize that when constructing the bootstrapped sample $X^*_{i,j}(g)$ and $y^*_{i,j}(g)$ in Step 4.2 of Algorithm 1 below, we use the interacted instruments $\overline{Z}_{i,j}$ defined in Step 2 of the algorithm. However, we still use the original instruments $Z_{i,j}$ to estimate the estimator and construct the test statistic $T_{CR,n}$.

\IncMargin{-1em}

algorithm[algorithm omitted — 2,253 chars of source]

\DecMargin{-1em}

$\widehat{Q}_{\tilde{Z}W,j}$ in the Empirical Application

Tables (ref)--(ref) report the values of $\widehat{Q}_{\tilde{Z}W,j}$ for each cluster in the South, Midwest, and West regions, and all of them are close to zero. For example, in Table (ref), the column $``W_1"$ reports the values of $\widehat{Q}_{\tilde{Z}W_1,j}$, where $j=1, ..., 16$ for the South region. Therefore, ((ref)) in Assumption (ref)(i) is plausible in this application.

table[table omitted — 1,600 chars of source]
table[table omitted — 1,281 chars of source]
table[table omitted — 1,190 chars of source]

Equivalence Among $k$-Class Estimators

We define $\hat{\beta}_L$ as the $k$-class estimator with $\hat{\kappa}_L$ for $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$. Their null-restricted and bootstrap counterparts are denoted as $\hat{\beta}_L^r$ and $\hat{\beta}_{L,g}^*$, respectively, and $\hat{\gamma}_L$, $\hat{\gamma}_L^r$, and $\hat{\gamma}_{L,g}^*$ are similarly defined. In the following, we show that $\hat{\beta}_{tsls}$, $\hat{\beta}_{liml}$, $\hat{\beta}_{full}$, $\hat{\beta}_{ba}$ are asymptotically equivalent, and so be their null-restricted and bootstrap counterparts.

lemmaSuppose Assumptions (ref), (ref), and (ref)(i) hold. Then, for $L \in \{\text{liml},\text{full},\text{ba}\}$, we have \begin{align*} & \hat{\beta}_{L} = \hat{\beta}_{tsls} + o_P(r_n^{-1}), \quad \hat{\beta}_{tsls} - \beta_n = O_P(r_n^{-1}), \\ & \hat{\beta}_{L}^r = \hat{\beta}_{tsls}^r + o_P(r_n^{-1}), \quad and \quad \hat{\beta}_{tsls}^r - \beta_n = O_P(r_n^{-1}). \end{align*}

Proof. First, $L \in \{\text{liml},\text{full},\text{ba}\}$, we have $ \hat{\mu}_L = \hat{\kappa}_L-1$ and

align*[align* omitted — 440 chars of source]

Then, by applying the Frisch–Waugh–Lovell Theorem we obtain that

align*[align* omitted — 231 chars of source]

By construction, $\hat{\beta}_{tsls}$ corresponds to $\hat{\mu}_{tsls} = 0$. For the LIML estimator, we have

align*[align* omitted — 213 chars of source]

which implies that

align[align omitted — 282 chars of source]

We note that

align[align omitted — 392 chars of source]

In addition, let $\hat{\varepsilon}_{i,j}$ be the residual from the full sample projection of $\varepsilon_{i,j}$ on $W_{i,j}$. Then, we have

align[align omitted — 1,009 chars of source]

where $c$ is a positive constant, the first inequality is by the definition that $\dot{\varepsilon}_{i,j}$ is the residual from the cluster-level projection of $\varepsilon_{i,j}$ on $W_{i,j}$, and the second inequality is by Assumption (ref)(iii). Combining (ref)--(ref), we have $\hat{\mu}_{liml} = O_P(r_n^{-2})$. In addition, we have $\frac{1}{n} X^{\top} M_{\vec{Z}} X = O_P(1)$, $\frac{1}{n} X^{\top} M_{\vec{Z}} Y = O_P(1)$, and

align[align omitted — 207 chars of source]

where the last equality holds by Assumptions (ref)(ii) and (ref)(i). This means

align*[align* omitted — 74 chars of source]

In addition, we note $\hat{\mu}_{full} = \hat{\mu}_{liml} - \frac{C}{n-d_z-d_w} = O_P(r_n^{-2})$ and $\hat{\mu}_{ba} = O(n^{-1})$, respectively. Therefore, we have the same results for $(\hat{\beta}_{full},\hat{\beta}_{ba})$.

For the second statement in the lemma, we have

align*[align* omitted — 394 chars of source]

where the last equality holds because $\widehat{Q}_{\widetilde{Z}\varepsilon} = \frac{1}{n}\sum_{j \in [J]} \sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}\varepsilon_{i,j} = O_P(r_n^{-1})$.

Next, we turn to the third statement in the lemma. We note that, for $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$ and $\lambda_{\beta}^{\top}\hat{\beta}^r_L = \lambda_0$,

align*[align* omitted — 316 chars of source]

As $\hat{\mu}_{L} = O_P(r_n^{-2})$ and $\hat{\beta}_L = \hat{\beta}_{tsls}+o_P(r_n^{-1})$ for $L \in \{\text{liml},\text{full},\text{ba}\}$, we have

align*[align* omitted — 316 chars of source]

For the last statement in the lemma, we note that

align*[align* omitted — 511 chars of source]

where the last equality holds because $\lambda_{\beta}^{\top}\beta_n - \lambda_0= \mu r_n^{-1}$ by construction and $\hat{\beta}_{tsls} - \beta_n = O_P(r_n^{-1})$. $\blacksquare$

lemmaSuppose Assumptions (ref), (ref), and (ref)(i) hold and $\widehat{Q}_{\widetilde{Z}W,j}(\hat{\gamma}_L^r-\gamma) = o_P(r_n^{-1})$. Then, for $L \in \{\text{liml},\text{full},\text{ba}\}$ and $g \in \textbf{G}$, we have \begin{align} \hat{\beta}_{L,g}^* = \hat{\beta}_{tsls,g}^* + o_P(r_n^{-1}) \quad and \quad \hat{\beta}_{tsls,g}^*- \beta_n = O_P(r_n^{-1}). \end{align}

Proof. By the same argument in the proof of Lemma (ref), for $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$ and $g \in \textbf{G}$, we have

align[align omitted — 265 chars of source]

such that $\hat{\mu}_{tsls,g}^* = 0$, $\hat{\mu}_{full,g}^* = \hat{\mu}_{liml,g}^* - \frac{C}{n-d_z-d_x}$, $\hat{\mu}_{ba,g}^* = \hat{\mu}_{ba}$, and $$\hat{\mu}^*_{liml,g}=\min_{r} r^{\top} \vec{Y}^{*\top}(g) M_W Z(Z^{\top} M_W Z)^{-1} Z^{\top} M_W \vec{Y}^*(g) r / (r^{\top} \vec{Y}^{*\top}(g) M_{\vec{Z}}\vec{Y}^*(g)r),$$ where $\vec{Y}^{*}(g) = [Y^*(g) : X^*(g)]$, $Y^*(g)$ is an $n \times 1$ vector formed by $Y_{i,j}^*(g)$, and $r=(1,-\beta^{\top} )^{\top}$.

Following the same argument previously, we have

align[align omitted — 310 chars of source]

where $\varepsilon_g^{*r}$ is an $n\times 1$ vector formed by $g_j\hat{\varepsilon}_{i,j}^r$. We first note that

align[align omitted — 515 chars of source]

where the last line is by Assumptions (ref)(ii), Lemma (ref), and the assumption that $\widehat{Q}_{\widetilde{Z}W,j}(\hat{\gamma}_L^r-\gamma) = o_P(r_n^{-1})$. Second, we have

align[align omitted — 1,232 chars of source]

where the last equality is by $\frac{1}{n_j} \sum_{i \in I_{n,j}}\hat{\varepsilon}^{r}_{i,j} \widetilde{Z}_{i,j} = O_P(r_n^{-1})$, as in ((ref)). We further note that

align*[align* omitted — 147 chars of source]

where $\hat{\beta}_L^r - \beta_n = O_P(r_n^{-1})$ and

align*[align* omitted — 263 chars of source]

Therefore, we have

align*[align* omitted — 173 chars of source]

where $\mathscr{E}_{i,j} = \varepsilon_{i,j} - W_{i,j}^\top\widehat{Q}_{WW}^{-1} \widehat{Q}_{W\varepsilon}.$ Similarly, we can show that

align*[align* omitted — 473 chars of source]

which, combined with (ref), implies

align*[align* omitted — 499 chars of source]

where $\hat{\theta}_g = \widehat{Q}_{WW}^{-1}\left[\frac{1}{n} \sum_{j \in [J]}\sum_{i \in I_{n,j}}(g_j\mathscr{E}_{i,j} W_{i,j})\right]$ and the second equality holds because

align*[align* omitted — 183 chars of source]

Recall that $\dot{\varepsilon}_{i,j}$ is the residual from the cluster-level projection of $\varepsilon_{i,j}$ on $W_{i,j}$, which means there exists a vector $\hat{\theta}_j $ such that $\varepsilon_{i,j} = \dot{\varepsilon}_{i,j} + W_{i,j}^\top \hat{\theta}_j$ and $\sum_{i \in I_{n,j}}\dot{\varepsilon}_{i,j}W_{i,j} = 0$. Then, we have

align*[align* omitted — 394 chars of source]

where the last inequality is by Assumption (ref)(iii). This implies

align[align omitted — 112 chars of source]

Combining (ref), (ref), and (ref), we obtain that $\hat{\mu}_{liml,g}^* = O_P(r_n^{-2})$, and thus, $\hat{\mu}^*_{full,g} = O_P(r_n^{-2})$. It is also obvious that $\hat{\mu}^*_{ba,g} = O_P(n^{-1})$. Given that $\hat{\mu}^*_{L,g} = O_P(r_n^{-2} + n^{-1})$, to establish $\hat{\beta}_{L,g}^* = \hat{\beta}_{tsls,g}^* + o_P(r_n^{-1})$, it suffices to show $\frac{1}{n}X^{*\top}(g) M_{\vec{Z}} X^*(g) = O_P(1)$, $\frac{1}{n}X^{*\top}(g) M_{\vec{Z}} Y^*(g) = O_P(1)$, and $(\frac{1}{n}X^{*\top}(g) P_{\widetilde{Z}} X^*(g))^{-1} = O_P(1)$.

For $\frac{1}{n}X^{*\top}(g) M_{\vec{Z}} Y^*(g) = \frac{1}{n}X^{*\top}(g) Y^*(g) - \frac{1}{n}X^{*\top}(g) P_{\vec{Z}} Y^*(g)$, we first show

align[align omitted — 373 chars of source]

where $\widehat{Q}^*_{XX}(g) = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}X^*_{i,j}(g)X^{*\top}_{i,j}(g)$, $\widehat{Q}^*_{XW}(g) = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}X^*_{i,j}(g)W^{\top}_{i,j}$, and $\widehat{Q}^*_{X\varepsilon}(g) = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}g_j X^*_{i,j}(g) \hat{\varepsilon}^r_{i,j}$. Notice that

align[align omitted — 235 chars of source]

because $\widehat{Q}_{XX,j}=O_P(1)$, $\widehat{Q}_{X\tilde{v},j}=O_P(1)$, and $\widehat{Q}_{\tilde{v}\tilde{v},j}=O_P(1)$. To see these three relations, we note that

align*[align* omitted — 172 chars of source]

by $\widehat{Q}_{XX,j}=O_P(1)$, $\widehat{Q}_{X\overline{Z},j} = O_P(1)$, $\widehat{Q}_{XW,j}=O_P(1)$, $\tilde{\Pi}_{\overline{Z}}=O_P(1)$, and $\widetilde{\Pi}_w = O_P(1)$, where $\widehat{Q}_{X\overline{Z},j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}X_{i,j}\overline{Z}_{i,j}^{\top}$, and $\widehat{Q}_{XW,j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}X_{i,j}W_{i,j}^{\top}$. Similar arguments hold for $\widehat{Q}_{\tilde{v}\tilde{v},j}$. We can also show that

align[align omitted — 192 chars of source]

by $\widehat{Q}_{XW,j}=O_P(1)$ and $\widehat{Q}_{\tilde{v}W,j}=O_P(1)$, where $\widehat{Q}_{XW,j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}X_{i,j}W_{i,j}^{\top}$ and $\widehat{Q}_{\tilde{v}W,j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}\tilde{v}_{i,j}W_{i,j}^{\top}$. In addition, we have

align[align omitted — 215 chars of source]

by $\widehat{Q}_{X\hat{\varepsilon},j} = O_P(1)$ and $\widehat{Q}_{\tilde{v}\hat{\varepsilon},j}=O_P(1)$ under similar arguments as those in ((ref)), where

align*[align* omitted — 259 chars of source]

Combining ((ref)), ((ref)), ((ref)), $\hat{\beta}^r_L=O_P(1)$, and $\hat{\gamma}^r_L= \widehat{Q}_{WW}^{-1} \widehat{Q}_{w\varepsilon}+o_P(1)=O_P(1)$, we obtain ((ref)). Next, by the fact that $\frac{1}{n}\sum_{j \in [J]} \sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}W_{i,j}^{\top} = 0$, we have

align*[align* omitted — 352 chars of source]

where

align*[align* omitted — 380 chars of source]

Following the same lines of reasoning, we can show that, for all $g \in \textbf{G}$, $\widehat{Q}_{\widetilde{Z}X}^*(g) = O_P(1)$, $\widehat{Q}_{\widetilde{Z}Y}^*(g)=O_P(1)$, and $\widehat{Q}_{WY}^*(g)=O_P(1)$. In addition, by Assumptions (ref)(iv) and (ref)(iii), $\widehat{Q}_{\widetilde{Z}\widetilde{Z}}^{-1} = O_P(1)$ and $\widehat{Q}_{WW}^{-1} = O_P(1)$, which further implies that $\frac{1}{n}X^{*\top}(g) M_{\vec{Z}} Y^*(g) = O_P(1)$.

By a similar argument, we can show $\frac{1}{n}X^{*\top}(g) M_{\vec{Z}} X^*(g) = O_P(1)$. Last, we have

align*[align* omitted — 417 chars of source]

where we use the fact that $\widehat{Q}_{\widetilde{Z}X}^{*}(g) = \widehat{Q}_{\widetilde{Z}X} + o_P(1)$, which is established in Step 1 in the proof of Theorem (ref). In addition, $Q_{\widetilde{Z}X}^{\top}Q_{\widetilde{Z}\widetilde{Z}}^{-1}Q_{\widetilde{Z}X}$ is invertible by Assumptions (ref)(ii) and (ref)(i). This implies $(\frac{1}{n}X^{*\top}(g) P_{\widetilde{Z}} X^*(g))^{-1} = O_P(1)$, which further implies $\hat{\beta}_{L,g}^* = \hat{\beta}_{tsls,g}^* + o_P(r_n^{-1})$.

For the second result, we note that

align*[align* omitted — 754 chars of source]

where the second equality holds because $\sum_{j \in [J]}\sum_{i \in I_{n,j}}g_j\widetilde{Z}_{i,j}W_{i,j}^\top = 0.$ In addition, note that

align*[align* omitted — 309 chars of source]

Combining this with the fact that $\hat{\beta}_{tsls}^r - \beta_n = O_P(r_n^{-1})$ as shown in Lemma (ref), we have $\hat{\beta}_{tsls,g}^* - \beta_n = O_P(r_n^{-1})$. $\blacksquare$

lemmaSuppose Assumptions (ref) and (ref) hold. Then, we have, for $j \in [J]$, $$\widehat{Q}_{\widetilde{Z}W,j}(\bar{\gamma}^r - \gamma) = o_P(r_n^{-1}).$$ If, in addition, Assumption (ref)(i) holds, then we have, For $j \in [J]$, $L \in \{\text{tsls},\text{liml},\text{full},\text{ba}\}$, and $g \in \textbf{G}$, \begin{align*} \widehat{Q}_{\widetilde{Z}W,j} (\hat{\gamma}_L - \gamma) = o_P(r_n^{-1}), \quad \widehat{Q}_{\widetilde{Z}W,j} (\hat{\gamma}^r_L - \gamma) = o_P(r_n^{-1}), \quad and \quad \widehat{Q}_{\widetilde{Z}W,j} (\hat{\gamma}^*_{L,g} - \gamma) = o_P(r_n^{-1}). \end{align*}

Proof. Note that we assume $\widehat{Q}_{\widetilde{Z}W,j}=o_P(1)$ and $\frac{1}{n}\sum_{j \in [J]} \sum_{i \in I_{n,j}}W_{i,j}\varepsilon_{i,j} = O_P(r_n^{-1})$. The first statement holds because

align[align omitted — 392 chars of source]

Next, if Assumption (ref)(i) also holds, then

align[align omitted — 401 chars of source]

where the last equality holds by Lemma (ref). This implies $\widehat{Q}_{\widetilde{Z}W,j}(\hat{\gamma}_L-\gamma) = o_P(r_n^{-1})$. In the same manner, we can show $\widehat{Q}_{\widetilde{Z}W,j}(\hat{\gamma}_L^r-\gamma) = o_P(r_n^{-1})$. Last, given Assumptions (ref), (ref), (ref)(i), and the fact that $\widehat{Q}_{\widetilde{Z}W,j}(\hat{\gamma}_L^r-\gamma) = o_P(r_n^{-1})$, Lemma (ref) shows $\hat{\beta}_{L,g}^* - \beta_n = (\hat{\beta}_{L,g}^* - \hat{\beta}_{tsls,g}^*) + (\hat{\beta}_{tsls,g}^* - \beta_n) = O_P(r_n^{-1})$. Then, following the same argument in (ref), we can show $\hat{\gamma}_{L,g}^* - \gamma = O_P(r_n^{-1})$, which leads to the desired result. $\blacksquare$

Proof of Theorem (ref)

When $d_x =1$, we define $a_j = Q^{-1}Q_{\widetilde{Z}X}^\top Q_{\widetilde{Z}\widetilde{Z}}^{-1} Q_{\widetilde{Z}X,j}$. For $d_x>2$, $a_j$ is defined in Assumption (ref). Let $\mathbb{S} \equiv \textbf{R}^{d_z \times d_x} \times \textbf{R}^{d_z \times d_z} \times \otimes_{j \in [J]} \textbf{R}^{d_z} \times \textbf{R}^{d_r \times d_r}$ and write an element $s \in \mathbb{S}$ by $s = \left(s_1, s_2, \{s_{3,j}: j \in [J]\}, s_4 \right)$ where $s_{3,j} \in \textbf{R}^{d_z}$ for any $j \in [J]$. Define the function $T$: $\mathbb{S} \rightarrow \textbf{R}$ to be given by

align[align omitted — 189 chars of source]

for any $s \in \mathbb{S}$ such that $s_2$ and $s_1^{\top}s_2^{-1}s_1$ are invertible and let $T(s)=0$ otherwise. We also identify any $(g_1, ..., g_q) = g \in \textbf{G} = \{-1, 1\}^J$ with an action on $s \in \mathbb{S}$ given by $gs = \left( s_1, s_2, \{ g_j s_{3,j} : j \in [J] \}, s_4 \right)$. For any $s \in \mathbb{S}$ and $\textbf{G'} \subseteq \textbf{G}$, denote the ordered values of $\{ T(gs) : g \in \textbf{G'} \}$ by $T^{(1)}(s | \textbf{G'}) \leq \ldots \leq T^{(|\textbf{G'}|)} (s | \textbf{G'})$. In addition, for any $\textbf{G'} \subseteq \textbf{G}$, denote the ordered values of $\{ T^*_n(g) : g \in \textbf{G'} \}$ by $T^{*(1)}_n(\textbf{G'}) \leq \ldots \leq T_n^{*(|\textbf{G'}|)}(\textbf{G'})$.

Given this notation we can define the statistics $S_n, \widehat{S}_n \in \mathbb{S}$ as

align*[align* omitted — 452 chars of source]

Let $E_n$ denote the event $ E_n = I\left\{\text{$\widehat{Q}_{\widetilde{Z}X}$ is of full rank value and $\widehat{Q}_{\widetilde{Z}\widetilde{Z}}$ is invertible}\right\}, $ and Assumptions (ref)-(ref) imply that $\liminf_{n \rightarrow \infty} \mathbb{P} \{ E_n=1 \} =1.$ Also let $T^*_n(g)=0$ if $E_n=0$.

We first give the proof for the Wald statistic based on TSLS. Note that whenever $E_n=1$ and $\mathcal{H}_0$ is true, the Frisch-Waugh-Lovell theorem implies that

align[align omitted — 540 chars of source]

where $\widehat{Q} = \widehat{Q}^{\top}_{\widetilde{Z}X} \widehat{Q}_{\widetilde{Z}\widetilde{Z}}^{-1} \widehat{Q}_{\widetilde{Z}X}$ and $\iota_J \in \textbf{G}$ is a $J \times 1$ vector of ones.

In the following, we divide the proof into three steps. In the first step, we show

eqnarray[eqnarray omitted — 122 chars of source]

In the second step, we show

align[align omitted — 122 chars of source]

In the last step, we prove the desired result.

Step 1. By the continuous mapping theorem, it suffices to show $\widehat{Q}^*_{\widetilde{Z}X}(g) = \widehat{Q}_{\widetilde{Z}X} + o_P(1)$. Note that

eqnarray*[eqnarray* omitted — 267 chars of source]

Therefore, it suffices to show $\frac{1}{n_j} \sum_{i \in I_{n,j}} \widetilde{Z}_{i,j} \tilde{v}_{i,j} = o_P(1)$ for all $j \in [J]$. Recall $\overline{Z}_{i,j}$ is just $\widetilde{Z}_{i,j}$ interacted with all the cluster dummies and $\widetilde{\Pi}_{\overline{Z}}$ is the OLS coefficient of $\overline{Z}_{i,j}$ defined in ((ref)). Denote $\widetilde{\Pi}_{\overline{Z},j}$ as the $j$-th block of $\widetilde{\Pi}_{\overline{Z}}$, which corresponds to the OLS coefficient of $\widetilde{\Pi}_{\overline{Z},j}$. We have

align[align omitted — 439 chars of source]

where the second equality in (ref) holds by $\frac{1}{n_j}\sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}W_{i,j}^{\top}=o_P(1)$ and $\widetilde{\Pi}_{w} = O_P(1)$. In particular,

align*[align* omitted — 416 chars of source]

where $\widehat{Q}_{\widetilde{W}\widetilde{W}} = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\widetilde{W}_{i,j}\widetilde{W}_{i,j}^{\top}$, $\widehat{Q}_{\widetilde{W}\hat{\varepsilon}} = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\widetilde{W}_{i,j}\hat{\varepsilon}_{i,j}$, $\widehat{Q}_{\hat{\varepsilon}\hat{\varepsilon}} = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\hat{\varepsilon}_{i,j}^2$, $\widetilde{W}_{i,j} = W_{i,j} - \widehat{\Gamma}_{w,j}^{\top} \widetilde{Z}_{i,j}$, and $\widehat{\Gamma}_{w,j} = \widehat{Q}^{-1}_{\widetilde{Z}\widetilde{Z},j}\widehat{Q}_{\widetilde{Z}W,j}$. Notice that by $\widehat{Q}_{\widetilde{Z}W,j} = o_P(1)$, $\widehat{\Gamma}_{w,j} = o_P(1)$ so that

align[align omitted — 100 chars of source]

Similarly, we have

align[align omitted — 129 chars of source]

where the second equality follows from $\widehat{Q}_{W\hat{\varepsilon}}=0$ by the first-order condition of the $k$-class estimators. Furthermore, $\widehat{Q}_{\hat{\varepsilon}\hat{\varepsilon}} \geq c > 0$ by Assumption (ref)(iii) and $\widehat{Q}_{\hat{\varepsilon}\hat{\varepsilon}} \geq \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\dot{\varepsilon}^2_{i,j}$, where $\dot{\varepsilon}_{i,j}$ is the residual from the cluster-level projection of $\varepsilon_{i,j}$ on $W_{i,j}$. Therefore, $\widehat{Q}_{\widetilde{W}\widetilde{W}} - \widehat{Q}_{\widetilde{W}\hat{\varepsilon}} \widehat{Q}^{-1}_{\hat{\varepsilon}\hat{\varepsilon}} \widehat{Q}_{\hat{\varepsilon}\widetilde{W}} = \widehat{Q}_{WW} + o_P(1)$ by combining ((ref)) and ((ref)), and further by Assumption (ref)(iv),

align[align omitted — 250 chars of source]

Next, we define $\widehat{Q}_{\widetilde{Z}\dot{X},j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}\dot{X}_{i,j}^{\top}$ and recall $\widehat{Q}_{\widetilde{Z}\widetilde{Z},j} = \frac{1}{n_j}\sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}\widetilde{Z}_{i,j}^{\top}$, where $\dot{X}_{i,j} = X_{i,j} - \widetilde{\Pi}_{w}^{\top}W_{i,j} - \widetilde{\Pi}_{\hat{\varepsilon}}^{\top}\hat{\varepsilon}_{i,j}$, and $\widetilde{\Pi}_{w}$ and $\widetilde{\Pi}_{\hat{\varepsilon}}$ are the OLS coefficients of $W_{i,j}$ and $\hat{\varepsilon}_{i,j}$, respectively, from regressing $X_{i,j}$ on $(\overline{Z}^{\top}_{i,j}, W^{\top}_{i,j}, \hat{\varepsilon}_{i,j})^{\top}$ using the entire sample. Then, we have

align[align omitted — 479 chars of source]

where we use the facts that $\widehat{Q}_{\widetilde{Z}W,j} = o_P(1)$, $\widetilde{\Pi}_w = O_P(1)$, $\widetilde{\Pi}_{\hat{\varepsilon}} = O_P(1)$, and $\widehat{Q}_{\widetilde{Z}\hat{\varepsilon},j} = o_P(1)$. In particular,

align*[align* omitted — 342 chars of source]

where $\widehat{Q}_{\tilde{\varepsilon}\tilde{\varepsilon}} = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\tilde{\varepsilon}^2_{i,j}$, $\widehat{Q}_{\tilde{\varepsilon}W} = \frac{1}{n}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\tilde{\varepsilon}_{i,j}W^{\top}_{i,j}$, $\tilde{\varepsilon}_{i,j} = \hat{\varepsilon}_{i,j} - \widehat{\Gamma}_{\hat{\varepsilon},j}^{\top} \widetilde{Z}_{i,j}$, and $\widehat{\Gamma}_{\hat{\varepsilon},j} = \widehat{Q}_{\widetilde{Z}\widetilde{Z},j}^{-1}\widehat{Q}_{\widetilde{Z}\hat{\varepsilon},j}$. Then, by using similar arguments as those for ((ref)), ((ref)), and ((ref)), we obtain

align*[align* omitted — 410 chars of source]

To see the last equality in (ref)}, we note that

align*[align* omitted — 431 chars of source]

where the last equality holds by Assumption (ref)(ii), Lemma (ref), and Lemma (ref). Plugging (ref)} into (ref), we obtain the desired result that $\frac{1}{n_j} \sum_{i \in I_{n,j}} \widetilde{Z}_{i,j} \tilde{v}_{i,j} = o_P(1)$.

Step 2. We note that whenever $E_n=1$, for every $g \in \textbf{G}$,

align[align omitted — 698 chars of source]

By Lemma (ref), we have

align[align omitted — 424 chars of source]

Note under both cases in Assumption (ref)(ii), we have $\sum_{j \in [J]} \xi_j a_j = 1$ and

align[align omitted — 161 chars of source]

Then, for any $\varepsilon >0$, we have

align[align omitted — 663 chars of source]

where the last equality holds because $\lambda_{\beta}^{\top}\hat{\beta}^r_{tsls}=\lambda_0$ under $\mathcal{H}_0$.

Note that $T(g \widehat{S}_n) = T(g S_n)$ whenever $E_n=0$ as we have defined $T(s)=0$ for any $s=\left(s_1, s_2, \{s_{3,j}: j \in [J]\}, s_4 \right)$ whenever $s_2$ or $s_1^{\top}s_2^{-1}s_1$ is not invertible. Therefore, results in ((ref)), ((ref)) and ((ref)) imply (ref).

Step 3. Note that by \hyperlink{assumption: 1}{Assumptions (ref)}, \hyperlink{assumption: 2}{(ref)}, (ref), and the continuous mapping theorem, we have

align[align omitted — 390 chars of source]

where $\xi_j > 0$ for all $j \in [J]$ by Assumption (ref)(iii). Therefore, we obtain from ((ref)), ((ref)), ((ref)), and the continuous mapping theorem that

align*[align* omitted — 164 chars of source]

For any $x \in \textbf{R}$ letting $\lceil x \rceil$ denote the smallest integer larger than $x$ and $k^* \equiv \lceil |\textbf{G}| (1-\alpha) \rceil$, we obtain from ((ref)) that

align[align omitted — 137 chars of source]

Since $\liminf_{n \rightarrow \infty} P\{ E_n =1\} =1$, we have

align*[align* omitted — 369 chars of source]

where the second inequality is due to the Portmanteau's theorem. To see the last inequality, we note that for all $g \in \textbf{G}$, $\mathbb{P}\{ T(gS) = T(-gS)\} = 1$, $\mathbb{P}\{ T(gS) = T(\tilde{g}S)\} = 0$ for $\tilde{g} \notin \{ g, -g \}$, and under the null, $T(gS)$ has the same distribution across $g \in \textbf{G}$. Let $|\textbf{G}| = 2^q$. Then, we have

align*[align* omitted — 187 chars of source]

which implies

align*[align* omitted — 162 chars of source]

For the lower bound, first note that $k^* > |\textbf{G}| - 2$ implies that $\alpha - \frac{1}{2^{J-1}} \leq 0$, in which case the result trivially follows. Now assume $k^* \leq |\textbf{G}| - 2$, then

align*[align* omitted — 390 chars of source]

where the first equality follows from ((ref)), the first inequality follows from Portmanteau's theorem, the second inequality holds because $\mathbb{P} \{ T^{(\textbf{z}+2)} (S | \textbf{G}) > T^{(\textbf{z})}(S | \textbf{G}) \} =1$ for any integer $\textbf{z} \leq |\textbf{G}| -2$ by ((ref)) and Assumption (ref), and the last inequality follows from noticing that $k^*+2 = \lceil |\textbf{G}|((1-\alpha)+2/|\textbf{G}|) \rceil = \lceil |\textbf{G}| (1-\alpha') \rceil$ with $\alpha' = \alpha - \frac{1}{2^{J-1}}$ and the properties of randomization tests.

Lemma (ref) and (ref) in the Supplement further show the other $k$-class estimators and their null-restricted and bootstrap counterparts are asymptotically equivalent to those of the TSLS estimator. Therefore, the results for LIML, FULL, and BA estimators can be derived in the same manner. $\blacksquare$

Proof of Theorem (ref)

For the power of the $T_n$-based wild bootstrap test, we focus on the TSLS estimator. As shown in the end of the proof of Theorem (ref), the test statistics constructed using LIML, FULL, and BA estimators and their bootstrap counterparts are asymptotically equivalent to those constructed based on TSLS estimator, which leads to the desired result. Note that

align*[align* omitted — 922 chars of source]

Notice that Assumptions (ref), (ref), (ref)(i), and Lemma (ref) imply that $r_n(\hat{\beta}^r_{tsls} - \beta_n)$ is bounded in probability. This implies

align[align omitted — 320 chars of source]

Recall $\widehat{Q}$ and $\widehat{Q}^{*}_g$ defined in (ref) and (ref), respectively. We have $\widehat{Q}^{*-1}_g = \widehat{Q}^{-1}+o_P(1)$ and

align[align omitted — 980 chars of source]

where the last equality follows from Lemma (ref) and $\widehat{Q}^*_{\widetilde{Z}X}(g) = \widehat{Q}_{\widetilde{Z}X} + o_P(1)$. Furthermore, we notice that

align[align omitted — 609 chars of source]

Therefore, employing ((ref)) with $r_n (\lambda_{\beta}^{\top} \beta_n - \lambda_0) = \lambda_{\beta}^{\top} \mu_{\beta}$, we conclude that whenever $E_n=1$,

align*[align* omitted — 571 chars of source]

Together with $(\ref{eq: boot-unstud-power-2})$, this implies that

align*[align* omitted — 1,514 chars of source]

This implies

align[align omitted — 378 chars of source]

In addition, let $\textbf{G}_s = \textbf{G} \backslash \textbf{G}_w$, where $\textbf{G}_w = \{g \in \textbf{G}: g_{j} = g_{j'}, \forall j,j' \in \mathcal{J}_s\}$. To establish the power result for the bootstrap test, we request that $\hat c_n(1-\alpha)$ does not take values of $T^*_n(g)$ for $g \in \textbf{G}_w$ because otherwise the test statistic and the critical value are asymptotically equivalent even under the alternative. We note that the following condition is imposed in the theorem: $|\textbf{G}_s| = |\textbf{G}| - 2^{J-J_s+1} \geq k^*$. Therefore, based on (ref) and (ref), to establish the desired result, it suffices to show that as $||\mu||_2 \rightarrow \infty$,

align*[align* omitted — 117 chars of source]

which follows under similar arguments as those employed in the proof of Theorem 3.2 in Canay-Santos-Shaikh(2020). $\blacksquare$

Proof of Theorem (ref)

Following the same argument in the proof of Theorem (ref), we can show that

align*[align* omitted — 175 chars of source]

where $T_{CR,n}(g)$ is defined in the proof of Theorem (ref) with $\mu=0$ as we are under the null. Then, the rest of the proof is similar to Step 3 in the proof of Theorem (ref). We omit the detail for brevity. $\blacksquare$

Proof of Theorem (ref)

For the power analysis, we focus on the TSLS estimator. The results for other $k$-class estimators can be derived in the same manner given Lemma (ref). Recall $a_j$ defined in Assumption (ref). We further define

align[align omitted — 186 chars of source]

where $\tilde{Q} = Q^{-1} Q_{\widetilde{Z}X}^{\top} Q_{\widetilde{Z}\widetilde{Z}}^{-1}$, $Q = Q_{\widetilde{Z}X}^\top Q_{\widetilde{Z}\widetilde{Z}}Q_{\widetilde{Z}X}$, $c_{0,g} = \sum_{j \in [J]} \xi_j g_j a_j$, and

align*[align* omitted — 451 chars of source]

We order $\{T_{CR,\infty}(g)\}_{g \in \textbf{G}}$ in ascending order: $(T_{CR,\infty})^{(1)} \leq \cdots \leq (T_{CR,\infty})^{|\textbf{G}|}$. In the proof of Theorem (ref), we have already shown that, under $\mathcal{H}_{1,n}$,

align*[align* omitted — 372 chars of source]

Next, we derive the limit of $\hat{A}_{r,CR}$ and $\hat{A}^*_{r,CR,g}$. We first note that

align*[align* omitted — 969 chars of source]

where the first equality holds by Lemma (ref). This implies

align*[align* omitted — 507 chars of source]

where $\iota_J$ is a $J \times 1$ vector of ones. Similarly, we have

align[align omitted — 1,064 chars of source]

where $\widehat{Q}_{\widetilde{Z}X,j}^*(g) = \frac{1}{n_j}\sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}X_{i,j}^{*\top}(g)$, the second equality is by Lemma (ref), and the last equality holds because $\widehat{Q}_{\widetilde{Z}X,j} = Q_{\widetilde{Z}X,j} + o_P(1)$, $\widehat{Q}_{\widetilde{Z}X,j}^*(g) = Q_{\widetilde{Z}X,j} + o_P(1)$ proved in Step 1 of the proof of Theorem (ref), $r_n(\hat{\beta}^r_{tsls}-\beta_n) = O_P(1)$, and $r_n(\hat{\beta}^*_{tsls,g} - \beta_n) = O_P(1)$. In addition, following the same arguments that lead to (ref), we have

align[align omitted — 647 chars of source]

Note $\widehat{Q}^{*-1}_g\widehat{Q}_{\widetilde{Z}X}^{*\top}(g)\widehat{Q}_{\widetilde{Z}\widetilde{Z}}^{-1} \stackrel{p}{\longrightarrow} \tilde{Q}$. Therefore, combining (ref) and (ref), we have

align*[align* omitted — 1,018 chars of source]

where the second equality is by the fact that $\tilde{Q}Q_{\widetilde{Z}X,j} = a_j I_{d_x}$ and the last convergence is by the fact that $\lambda_\beta^\top r_n(\hat{\beta}^r_{tsls} - \beta_n) = -\mu$. This implies

align*[align* omitted — 284 chars of source]

By the Portmanteau theorem, we have

align*[align* omitted — 182 chars of source]

We aim to show that, as $||\mu ||_2\rightarrow \infty$, we have

align[align omitted — 144 chars of source]

where $\textbf{G}_s = \textbf{G} \backslash \textbf{G}_w$, and $\textbf{G}_w = \{g \in \textbf{G}: g_{j} = g_{j'}, \forall j,j' \in \mathcal{J}_s\}$. Then, given $|\textbf{G}_s| = |\textbf{G}|-2^{J-J_s+1}$ and $k^* = \lceil |\textbf{G}|(1-\alpha) \rceil \leq |\textbf{G}|-2^{J-J_s+1}$, (ref) implies that as $ ||\mu||_2 \rightarrow \infty$,

align*[align* omitted — 204 chars of source]

Therefore, it suffices to establish (ref).

By ((ref)), we see that $$T_{CR,\infty}(\iota_J) = \left\Vert \lambda_\beta^\top \tilde{Q} \left[\sum_{j \in [J]} \mathcal{Z}_j \right]+ \mu \right\Vert_{A_{r,CR,\iota_J}},$$ and $A_{r,CR,\iota_J}$ is independent of $\mu$ as $c_{0, \iota_J}=1$. In addition, we have $\lambda_{\min}(\tilde{Q}^\top \lambda_\beta^\top A_{r,CR,\iota_J} \lambda_\beta \tilde{Q})>0$ with probability one. Therefore, for any $\delta>0$, we can find a sufficiently small constant $c>0$ and a sufficiently large constant $M>0$ such that with probability greater than $1-\delta$, for any $\mu$,

align[align omitted — 80 chars of source]

On the other hand, for $g \in \textbf{G}_s$, we can write $T_{CR,\infty}(g)$ as

align*[align* omitted — 193 chars of source]

where for $j \in [J]$, $N_{0,g} = \lambda_\beta^\top \tilde{Q} \left[\sum_{j \in [J]} g_j \mathcal{Z}_j \right]$, $N_{j,g} = \lambda_\beta^\top \tilde{Q} \left[g_j\mathcal{Z}_j - a_j \xi_j \sum_{\tilde{j} \in [J]} g_{\tilde{j}}\mathcal{Z}_{\tilde{j}} \right]$, and $c_{j,g} = \xi_j (g_j - c_{0,g})a_j$.

We claim that for $g \in \textbf{G}_s$, $c_{j,g} \neq 0$ for some $j \in \mathcal J_s$. Suppose it does not hold, then it implies that $g_j = c_{0,g}$ for all $j \in \mathcal J_s$, i.e., for all $j \in \mathcal J_s$, $g_j$ shares the same sign, and thus, contradicts the definition of $\textbf{G}_s$. Therefore, combining the claim with the assumption that $\min_{j \in \mathcal J_s}|a_j| > 0$, we have $\min_{g \in \textbf{G}_s} \sum_{j \in [J]} c_{j,g}^2 >0.$

In addition, we note that

eqnarray*[eqnarray* omitted — 352 chars of source]

where we denote $M_1 = \sum_{j \in [J]} N_{j,g} N_{j,g}^\top$, $M_2 = \sum_{j \in [J]} c_{j,g} N_{j,g}$, and $\overline{c}^2 = \sum_{j \in [J]} c_{j,g}^2$. For notation ease, we suppress the dependence of $(M_1,M_2,\overline{c})$ on $g$. Then, we have

align*[align* omitted — 243 chars of source]

Note for any $d_r \times 1$ vector $u$, by the Cauchy–Schwarz inequality,

align*[align* omitted — 218 chars of source]

where the equal sign holds if and only if there exist $(u,g,c) \in \Re^{d_r} \times \textbf{G}_s \times \Re$ such that

align*[align* omitted — 52 chars of source]

or equivalently,

align*[align* omitted — 195 chars of source]

This occurs with probability zero if $J - 1> d_r$ as $\{g_j \mathcal Z_j\}_{j \in [J]}$ are independent and non-degenerate normal vectors and the matrix $\left(I_J + (-a_1\xi_1,\cdots,-a_J\xi_J)^\top\iota_J^\top \right)$ has rank $J-1$. Therefore, the matrix $\mathbb{M} \equiv M_1 - \frac{M_2 M_2^\top}{\overline{c}^2}$ is invertible with probability one. Specifically, denote $\mathbb{M}$ as $\mathbb{M}(g)$ to highlight its dependence on $g$. We have $\max_{g \in \textbf{G}_s }(\lambda_{\min}(\mathbb{M}(g)))^{-1} = O_P(1)$. In addition, denote $\frac{M_2}{\overline{c}} + \overline{c} \mu $ as $\mathbb{V}$, which is a $d_r \times 1$ vector. Then, we have

align*[align* omitted — 284 chars of source]

where the second equality is due to the Sherman–Morrison–Woodbury formula.

Next, we note that

align*[align* omitted — 192 chars of source]

where $ \mathbb{M}_0 = N_{0,g} - \frac{c_{0,g} M_2}{\overline{c}^2} = N_{0,g} - \frac{c_{0,g} (\sum_{j \in [J]} c_{j,g} N_{j,g})}{\sum_{j \in [J]} c_{j,g}^2}.$ With these notations, we have

align*[align* omitted — 1,306 chars of source]

where the first inequality holds because $(u+v)^\top A (u+v) \leq 2 (u^\top A u + v^\top A v)$ for some $d_r\times d_r$ positive semidefinite matrix $A$ and $u,v \in \Re^{d_r}$, the second inequality holds due to the fact that $\mathbb{M}^{-1} \mathbb{V} (1+ \mathbb{V}^\top \mathbb{M}^{-1}\mathbb{V})^{-1} \mathbb{V}^\top \mathbb{M}^{-1}$ is positive semidefinite, the third inequality holds because $\mathbb{V}^\top \mathbb{M}^{-1}\mathbb{V}$ is a nonnegative scalar, and the last holds by substituting in the expressions for $\mathbb{M}_0$ and $\overline{c}$.

Then, we have

align[align omitted — 116 chars of source]

Combining (ref) and (ref), we have, as $||\mu||_2 \rightarrow \infty$,

align*[align* omitted — 509 chars of source]

where the second inequality is by the fact that $k^* \leq |\textbf{G}_s|$, and thus, $(T_{CR,\infty})^{(k^*)} \leq \max_{g \in \textbf{G}_s} T_{CR,\infty}(g)$ and the last convergence holds because $\max_{g \in \textbf{G}_s} C(g) = O_P(1)$ and does not depend on $\mu$. As $\delta$ is arbitrary, we can let $\delta \rightarrow 0$ and obtain the desired result.

For the proof of part (ii), let $\tilde{c}_{CR,n}(1-\alpha)$ denote the $(1-\alpha)$ quantile of

align*[align* omitted — 164 chars of source]

i.e., the bootstrap statistic $T_n^*(g)$ studentized by the original CRVE instead of the bootstrap CRVE. Because $d_r = 1$, we have

align*[align* omitted — 93 chars of source]

Therefore, we have

align*[align* omitted — 772 chars of source]

where we use the fact that

align*[align* omitted — 251 chars of source]

Further note that $\max_{g \in \textbf{G}}|r_n\lambda_{\beta}^{\top}(\hat{\beta}^*_{tsls,g} - \beta_n)| = O_P(1)$ and does not depend on $\mu$, and

align*[align* omitted — 151 chars of source]

and $A_{r,CR,\iota_J}$ does not depend on $\mu$ either.

Last, although $\hat{c}_{CR,n}(1-\alpha)$ depends on $\mu$, we have

align*[align* omitted — 219 chars of source]

for some $C(g)$ that does not depend on $\mu$ as has been proved in Part (i). Therefore, for any $\delta>0$, there exists a constant $c_{\mu}>0$, such that when $|\mu| > c_\mu$,

align*[align* omitted — 763 chars of source]

where $C$ in the first inequality is a constant such that for $n$ being sufficiently large,

align*[align* omitted — 145 chars of source]

the second inequality is by Portmanteau theorem, and the last inequality holds if $c_\mu$ is sufficiently large. This concludes the proof. $\blacksquare$

Proof of Theorem (ref)

The proof for the $AR_n$-based wild bootstrap test follows similar arguments as those in \hyperlink{theo: boot-t}{Theorem (ref)}, and thus we keep exposition more concise. Let $\mathbb{S} \equiv \otimes_{j \in [J]} \textbf{R}^{d_z} \times \textbf{R}^{d_z \times d_z}$ and write an element $s \in \mathbb{S}$ by $s = ( \{s_{1j} : j \in [J]\}, s_2)$ where $s_{1j} \in \textbf{R}^{d_z}$ for any $j \in [J]$. Define the function $T_{AR}$: $\mathbb{S} \rightarrow \textbf{R}$ to be given by

align[align omitted — 97 chars of source]

Given this notation we can define the statistics $S_n, \widehat{S}_n \in \mathbb{S}$ as

align*[align* omitted — 277 chars of source]

where $\bar{\varepsilon}^r_{i,j} = y_{i,j}-X_{i,j}^{\top}\beta_0-W^{\top}_{i,j}\bar{\gamma}^r$. Note that by the Frisch-Waugh-Lovell theorem,

align[align omitted — 190 chars of source]

Therefore, letting $k^* \equiv \lceil |\textbf{G}| (1-\alpha) \rceil$, we obtain from ((ref))-((ref)) that

align*[align* omitted — 146 chars of source]

Furthermore, note that we have

align[align omitted — 548 chars of source]

where the third equality follows from $\sum_{j \in [J]} \sum_{i \in I_{n,j}} \widetilde{Z}_{i,j}W_{i,j}^{\top} = 0$. ((ref)) implies that if $k^* > |\textbf{G}| -2$, then $1 \{ T_{AR}(S_n) > T_{AR}^{(k^*)} (\widehat{S}_n | \textbf{G}) \} = 0$, and this gives the upper bound in Theorem \hyperlink{theo: boot-AR}{(ref)}. We therefore assume that $k^* \leq |\textbf{G}| -2$, in which case

align[align omitted — 437 chars of source]

Then, to examine the right hand side of ((ref)), first note that by Assumptions (ref) and (ref), and the continuous mapping theorem we have

align[align omitted — 235 chars of source]

where $\xi_j > 0$ for all $j \in [J]$. Furthermore, by Assumptions (ref)(i), Lemma (ref), and $\beta_n=\beta_0$, we have

align*[align* omitted — 342 chars of source]

and thus, for every $g \in \textbf{G}$,

align[align omitted — 86 chars of source]

We thus obtain from results in ((ref))-((ref)) and the continuous mapping theorem that

align*[align* omitted — 195 chars of source]

Then, by the Portmanteau’s theorem and the properties of randomization tests, we have

align*[align* omitted — 285 chars of source]

The lower bounds follow by applying similar arguments as those for \hyperlink{theo: boot-t}{Theorem (ref)}.

For the $AR_{CR,n}$-based wild bootstrap test, define the statistics $S_{CR,n}, \widehat{S}_{CR,n} \in \mathbb{S}$ as

align*[align* omitted — 317 chars of source]

Then, we have for any action $g \in \textbf{G}$ that

align[align omitted — 246 chars of source]

We set $E_n \in \textbf{R}$ to equal $E_n \equiv 1 \left\{ \frac{r_n^2}{n^2} \sum_{j\in J} \sum_{i \in I_{n,j}} \sum_{k \in I_{n,j}} \widetilde{Z}_{i,j} \widetilde{Z}_{k,j}^{\top} \bar{\varepsilon}^r_{i,j} \bar{\varepsilon}^r_{k,j} \; \text{is invertible} \right\},$ and have

align[align omitted — 93 chars of source]

In addition, similar to the case with $AR_n$ and $AR^*_n(g)$, we have under $\beta_n = \beta_0$,

align*[align* omitted — 141 chars of source]
align*[align* omitted — 119 chars of source]

where $A_{CR} = \sum_{j \in [J]} \mathcal{Z}_j \mathcal{Z}_j^{\top}$, and

align[align omitted — 243 chars of source]

Therefore, we have

align*[align* omitted — 360 chars of source]

which follows from ((ref)), ((ref)), the continuous mapping theorem and Portmanteau's theorem. The claim of the upper bound in the theorem then follows from similar arguments as those in Theorem (ref). $\blacksquare$

Proof of Theorem (ref)

Define $AR_{\infty}(g) = || \sum_{j \in [J]}g_j \mathcal{Z}_j + \sum_{j \in [J]}\xi_jg_ja_jQ_{\widetilde{Z}X}\mu ||_{A_z}$. In particular, notice that $AR_{\infty}(\iota_J) = || \sum_{j \in [J]} \mathcal Z_j + Q_{\widetilde{Z}X} \mu ||_{A_z}$ since $\sum_{j \in [J]} \xi_j a_j =1$. Following same arguments in the proof of Theorem (ref), we can show that, under $\mathcal{H}_{1,n}$ with $\lambda_{\beta} = I_{d_x}$,

align*[align* omitted — 161 chars of source]

Similar to the proofs of Theorems (ref) and (ref), in order to establish Theorem (ref), it suffices to show that as $||Q_{\widetilde{Z}X}\mu||_2 \rightarrow \infty$,

align*[align* omitted — 109 chars of source]

By Triangular inequality, we have $AR_{\infty}(\iota_J) \geq ||Q_{\widetilde{Z}X}\mu||_{A_z} - O_P(1)$, and $$\max_{g \in \textbf{G}_s}AR_{\infty}(g) \leq \max_{g \in \textbf{G}_s}|\sum_{j \in [J]}g_j \xi_j a_j| ||Q_{\widetilde{Z}X} \mu||_{A_z} + O_P(1).$$ In addition, $\max_{g \in \textbf{G}_s} |\sum_{j \in [J]} g_j \xi_j a_j|<1$ so that as $||Q_{\widetilde{Z}X}\mu||_2 \rightarrow \infty$, $||Q_{\widetilde{Z}X}\mu||_{A_z} \rightarrow \infty$ and

align*[align* omitted — 158 chars of source]

This concludes the proof. $\blacksquare$

Further Results for the Weak-IV-Robust Statistics

First, we show that in the specific case with one endogenous regressor and one IV, if the wild bootstrap procedure for the $AR_n$ statistic is applied to the unstudentized Wald statistic $T_n$, then the resulting test will be asymptotically equivalent to the $AR_n$-based bootstrap test, both under the null and the alternative. More precisely, for $T_n = | \hat{\beta} - \beta_0 |$, where $\hat{\beta}$ is the TSLS estimator, the wild bootstrap generates

equation[equation omitted — 203 chars of source]

Notice that in this case, the restricted TSLS estimator $\hat{\gamma}^r$ and the restricted OLS estimator $\bar{\gamma}^r$ are the same, which implies the residuals $\hat{\varepsilon}^r_{i,j}$ and $\Bar{\varepsilon}^r_{i,j}$ defined in Sections (ref) and (ref), respectively, are the same. Let $\hat{c}^s_n(1-\alpha)$ denote the $(1-\alpha)$-th quantile of $\{ T^{s*}_n(g) \}_{g \in \textbf{G}}$. We show this equivalence below.

theoremSuppose that $d_x=d_z=1$, $\liminf_{n \rightarrow \infty}\mathbb{P}(\widehat{Q}_{\widetilde{Z}X} \neq 0) = 1$, $\hat{\beta}$ in the Wald test is computed via TSLS, and Assumption (ref)(iv) holds. Then, \begin{eqnarray*} \liminf_{n \rightarrow \infty} \mathbb{P} \{ \phi_n^s = \phi_n^{ar} \} =1, \end{eqnarray*} where $\phi_n^s = 1 \{ T_n > \hat{c}^s_n(1-\alpha) \}$ and $\phi_n^{ar} = 1 \{ AR_n > \hat{c}^{ar}_n(1-\alpha) \}$.

Several remarks are in order. First, with one endogenous variable and one IV, the Jacobian matrix $\widehat{Q}_{\widetilde{Z}X}$ is just a scalar, which shows up in both the Wald statistic $T_n$ and the bootstrap critical value. After the cancellation of $\widehat{Q}_{\widetilde{Z}X}$, $T_n$ and its critical value are numerically the same as their AR counterparts, which leads to Theorem (ref). Second, for $\widehat{Q}_{\widetilde{Z}X}$ to be cancelled, we only need $\liminf_{n \rightarrow \infty}\mathbb{P} \{ \widehat{Q}_{\widetilde{Z}X} \neq 0 \} = 1$, which is very mild. It holds when $\widehat{Q}_{\widetilde ZX}$ is continuously distributed. Even when both $\widetilde{Z}_{i,j}$ and $X_{i,j}$ are discrete, it still holds if there exists at least one strong IV cluster, i.e., $Q_{\widetilde{Z}X} \neq 0$, where $Q_{\widetilde{Z}X}$ is the probability limit of $\widehat{Q}_{\widetilde{Z}X}$. When both $\widetilde{Z}_{i,j}$ and $X_{i,j}$ are discrete and $Q_{\widetilde{Z}X} = 0$, this condition still holds if some type of CLT holds such that $r_n \widehat{Q}_{\widetilde{Z}X} \xrightarrow{\enskip d \enskip} N(c,\sigma^2)$, as $N(c,\sigma^2)$ is continuous. Third, the robustness of the $T_n$-based bootstrap test in ((ref)) does not carry over to the general case with multiple IVs as $T_n$ and its bootstrap statistic can no longer be reduced to their AR counterparts. Fourth, the robustness to weak IV cannot be extended to the $T_{CR,n}$-based bootstrap test, for which we have to further bootstrap the CRVE.

Second, we discuss wild bootstrap inference with other weak-IV-robust statistics. To introduce the test statistics, we define the sample Jacobian as

align*[align* omitted — 251 chars of source]

and define the orthogonalized sample Jacobian as

align*[align* omitted — 241 chars of source]

where $\widehat{\Omega} = n^{-1}\sum_{j \in [J]}\sum_{i \in I_{n,j}}\sum_{k \in I_{n,j}}f_{i,j}f_{k,j}^{\top} $, and $\widehat{\Gamma}_{l} = n^{-1} \sum_{j \in [J]} \sum_{i \in I_{n,j}} \sum_{k \in I_{n,j}} \left( \widetilde{Z}_{i,j} X_{i,j,l} \right) f_{k,j}^\top,$ for $l=1, ..., d_x$. Therefore, under the null $\beta_n = \beta_0$ and the framework where the number of clusters tends to infinity, $\widehat{D}$ equals the sample Jacobian matrix $\widehat{G}$ adjusted to be asymptotically independent of $\widehat{f}$.

Then, the cluster-robust version of Kleibergen(2002),Kleibergen(2005)'s LM statistic is defined as

align*[align* omitted — 144 chars of source]

In addition, the conditional quasi-likelihood ratio (CQLR) statistic in Kleibergen(2005), Newey-Windmeijer(2009), and Guggenberger-Ramalho-Smith(2012) are adapted from Moreira(2003)'s conditional likelihood ratio (CLR) test, and its cluster-robust version takes the form

align*[align* omitted — 132 chars of source]

where $rk_n$ is a conditioning statistic and we let $rk_n = n \widehat{D}^\top \widehat{\Omega}^{-1} \widehat{D}$.\footnote{This choice follows Newey-Windmeijer(2009). Kleibergen(2005) uses alternative formula for $rk_n$, and Andrews-Guggenberger(2019) introduce alternative CQLR test statistic.}

The wild bootstrap procedure for the LM and CQLR tests is as follows. We compute

align*[align* omitted — 404 chars of source]

for any $g = (g_1, ..., g_q) \in \textbf{G}$, where the definition of $\widehat{f}_g^*$ and $f^*_{k,j}(g_j)$ is the same as that in Section (ref). Then, we compute the bootstrap analogues of the test statistics as

align*[align* omitted — 307 chars of source]

Let $\hat{c}_{LM,n}(1-\alpha)$ and $\hat{c}_{LR,n}(1-\alpha)$ denote the $(1-\alpha)$-th quantile of $\{LM_{n}^*(g)\}_{g \in \textbf{G}}$ and $\{LR^*_{n}(g)\}_{g \in \textbf{G}}$, respectively. We notice that with at least one strong cluster,

align*[align* omitted — 430 chars of source]

where $\widetilde{D} = \left( \widetilde{D}_1, ..., \widetilde{D}_{d_x}\right)$, and for $l=1, ..., d_x$, $$ \widetilde{D}_l = Q_{\widetilde{Z}X} - \left\{ \sum_{j \in [J]} \left(\xi_j Q_{\widetilde{Z}X,j,l} \right) \left( \sqrt{\xi_j} \mathcal{Z}_{\varepsilon,j} \right) \right\} \left\{ \sum_{j \in [J]} \xi_j \mathcal{Z}_{\varepsilon,j} \mathcal{Z}_{\varepsilon,j}^{\top} \right\}^{-1} \sum_{j \in [J]} \sqrt{\xi_j} \mathcal{Z}_{\varepsilon,j}.$$ Although the limiting distribution is nonstandard, we are able to establish the validity results by connecting the bootstrap LM test with the randomization test and by showing the asymptotic equivalence of the bootstrap LM and CQLR tests in this case. We conjecture that similar results can also be established for other weak-IV-robust statistics proposed in the literature.

theoremSuppose Assumptions (ref), (ref)(i), and (ref) hold, $\beta_n=\beta_0$, and $q > d_z$, then \begin{eqnarray*} \alpha - \frac{1}{2^{J-1}} \leq \liminf_{n \rightarrow \infty} \mathbb{P} \{LM_n > \hat{c}_{LM,n}(1-\alpha) \} \leq \limsup_{n \rightarrow \infty} \mathbb{P} \{LM_n > \hat{c}_{LM,n}(1-\alpha) \} \leq \alpha + \frac{1}{2^{J-1}}; \\ \alpha - \frac{1}{2^{J-1}} \leq \liminf_{n \rightarrow \infty} \mathbb{P} \{ LR_n > \hat{c}_{LR,n}(1-\alpha) \} \leq \limsup_{n \rightarrow \infty} \mathbb{P} \{ LR_n > \hat{c}_{LR,n}(1-\alpha) \} \leq \alpha + \frac{1}{2^{J-1}}. \end{eqnarray*}

Proof of Theorem (ref)

Notice that when $d_x = d_z = 1$, $\widehat{Q}_{\widetilde{Z}X}$ is a scalar, and the restricted TSLS estimator $\hat{\gamma}^r$ is equivalent to the restricted OLS estimator $\bar{\gamma}^r$, which is well defined by Assumption (ref)(iv), so that $\hat{\varepsilon}_{i,j}^r = \bar{\varepsilon}_{i,j}^r$. Therefore,

align*[align* omitted — 243 chars of source]

and whenever $\widehat{Q}_{\widetilde{Z}X} \neq 0$,

align*[align* omitted — 791 chars of source]

by $\beta_0 = \hat{\beta}^r_{tsls}$, $\sum_{j \in [J]}\sum_{i \in I_{n,j}}\widetilde{Z}_{i,j}W_{i,j}^{\top}=0$, and $\hat{\varepsilon}^r_{i,j} = \varepsilon_{i,j} - X_{i,j}^{\top}( \hat{\beta}^r_{tsls} - \beta_n ) - W_{i,j}^{\top}(\hat{\gamma}^r_{tsls} - \gamma)$.

In addition, for the bootstrap statistics we have

align*[align* omitted — 255 chars of source]

and whenever $\widehat{Q}_{\tilde{Z}X} \neq 0$,

align*[align* omitted — 235 chars of source]

Therefore, $1\{ T_n > \hat{c}^s_n(1-\alpha) \}$ is equal to $1\{ AR_n > \hat{c}_{AR,n}(1-\alpha) \}$ whenever $\widehat{Q}_{\widetilde{Z}X} \neq 0$. We conclude that $\liminf_{n \rightarrow \infty} \mathbb{P} \left\{ \phi_n^s = \phi_n^{ar} \right\} =1$ because $\liminf_{n \rightarrow \infty} \mathbb{P} \left\{ \widehat{Q}_{\widetilde{Z}X} \neq 0 \right\}=1$. $\blacksquare$

Proof of Theorem (ref)

The proof for the bootstrap LM test follows similar arguments as those for the studentized version of the bootstrap AR test. Let $\mathbb{S} \equiv \textbf{R}^{d_z \times d_x} \times \otimes_{j \in [J]} \textbf{R}^{d_z}$, and write an element $s \in \mathbb{S}$ by $s = \left( \{ s_{1,j} : j \in [J] \}, \{ s_{2,j} : j \in [J] \} \right)$. We identify any $(g_1, ..., g_q) = g \in \textbf{G} = \{-1, 1\}^J$ with an action on $s \in \mathbb{S}$ given by $gs = \left( \{ s_{1,j} : j \in [J] \}, \{ g_j s_{2,j} : j \in [J] \} \right)$. We define the function $T_{LM}: \mathbb{S} \rightarrow \textbf{R}$ to be given by

align[align omitted — 254 chars of source]

for any $s \in \mathbb{S}$ such that $\sum_{j \in [J]} s_{2,j} s_{2,j}^{\top}$ and $D(s)^{\top} \left( \sum_{j \in [J]} s_{2,j} s^{\top}_{2,j} \right)^{-1} D(s)$ are invertible and set $T_{LM}(s)=0$ whenever one of the two is not invertible, where

align[align omitted — 260 chars of source]

for $s_{1,j} = (s_{1,j,1}, ..., s_{1,j,d_x})$ and $l=1, ..., d_x$.

Furthermore, define the statistic $S_n$ as

align[align omitted — 246 chars of source]

Note that for $l = 1, ..., d_x$ and $j \in [J]$, by Assumptions (ref)(iii) and (ref)(i) we have

align[align omitted — 131 chars of source]

where $Q_{\tilde{Z}X,j,l}$ denotes the $l$-th column of the $d_z \times d_x$-dimensional matrix $Q_{\tilde{Z}X,j}$. Then, by Assumptions (ref)(ii) and (ref)(iii), (ref)(i) and (ref)(ii), and the continuous mapping theorem we have

align[align omitted — 188 chars of source]

where $\xi_j > 0$ for all $j \in [J]$. Also notice that for $l=1, ..., d_x$,

align[align omitted — 636 chars of source]

and $\frac{r_n}{n}\sum_{i \in I_{n,j}} \tilde{Z}_{i,j}\bar{\varepsilon}_{i,j}^r = \frac{r_n}{n}\sum_{i \in I_{n,j}} \tilde{Z}_{i,j}\varepsilon_{i,j} + o_P(1)$ by $\beta_n = \beta_0$ and Lemma (ref). In addition, we set $A_n \in \textbf{R}$ to equal

align[align omitted — 122 chars of source]

and we have

align[align omitted — 90 chars of source]

which holds because $\left\{ \mathcal{Z}_j : j \in [J] \right\}$ are independent and continuously distributed with covariance matrices that are of full rank, and $Q_{\tilde{Z}X,j}$ are of full column rank for all $j \in \mathcal{J}_s$, by Assumptions (ref)(ii), (ref)(i), and (ref)(ii).

It follows that whenever $A_n=1$,

align[align omitted — 184 chars of source]

In what follows, we denote the ordered values of $\{ T_{LM}(gs): g \in \textbf{G} \}$ by

align[align omitted — 95 chars of source]

Next, we have

align[align omitted — 442 chars of source]

where the final inequality follows from ((ref)), ((ref)), ((ref)), ((ref)), the continuous mapping theorem and Portmanteau's theorem. Therefore, setting $k^* \equiv \lceil |\textbf{G}|(1-\alpha)\rceil$, we can obtain from ((ref)) that

align[align omitted — 406 chars of source]

where the final inequality follows by $gS \overset{d}{=} S$ for all $g \in \textbf{G}$ and the properties of randomization tests. Then, we notice that for all $g \in \textbf{G}$, $T_{LM}(gS)=T_{LM}(-gS)$ with probability 1, and $P\{ T_{LM}(gS) = T_{LM}(\tilde{g}S) \}=0$ for $\tilde{g} \notin \{g, -g\}$. Therefore,

align[align omitted — 129 chars of source]

The claim of the upper bound in the theorem then follows from ((ref)) and ((ref)}). The proof for the lower bound is similar to that for the bootstrap AR test, and thus is omitted.

To prove the result for the CQLR test, we note that

align[align omitted — 573 chars of source]

where the third equality follows from the mean value expansion $\sqrt{1+x} = 1+ (1/2)(x+o(1))$, the fourth and last equalities follow from $AR_{CR,n} - rk_n <0$ w.p.a.1 since $AR_{CR,n} = O_P(1)$ while $rk_n \rightarrow \infty$ w.p.a.1 under \hyperlink{assumption: 3}{Assumption (ref)(i)}. Using arguments similar to those in ((ref)), we obtain that for each $g \in \textbf{G}$,

align[align omitted — 103 chars of source]

by $AR^*_{CR,n}(g) - rk_n <0$ w.p.a.1 since $AR^*_{CR,n}(g) = O_P(1)$ for each $g \in \textbf{G}$. Then, it follows that whenever $A_n=1$,

align[align omitted — 184 chars of source]

Then, we obtain that

align[align omitted — 442 chars of source]

where the second inequality follows from ((ref)), ((ref)), ((ref)), ((ref)), the continuous mapping theorem and Portmanteau's theorem. Finally, the upper and lower bounds for the studentized bootstrap CQLR test follows from the previous arguments for the bootstrap LM test. $\blacksquare$

Further Monte Carlo Simulation Results

Further Simulation Results for DGP 1

In this section, we report further simulation results for DGP 1. We notice that the overall pattern is very similar to that observed in Section (ref).

First, we report the results for DGP 1 with the number of clusters $J$ equal to $9$ or $20$. The size results with $d_z=1$ and $d_z=3$ are reported in Tables (ref) and (ref), respectively. In addition, we report the further power results in Figures (ref) - (ref) for $d_z=1$ and $J=9$ or $20$, Figures (ref) - (ref) for $d_z=3$ and $J=6$ or $12$, and Figures (ref) - (ref) for $d_z=3$ and $J=9$ or $20$, respectively. For the power results, we vary the value of $\beta_0-\beta$ from $-3$ to $3$ with the true value of $\beta$ equal to 1 throughout the simulations. $\rho_{\varepsilon v}$ is set equal to $0.5$.

Second, we report the simulation results under DGP 1 for the case where $\Pi \in \{0.25, 0.375, 0.5\}$ remains the same across clusters (i.e., homoskedastic $\Pi$), so that the cluster-level heterogeneity in identification strength originates solely from the heterogeneity in cluster size. Other settings remain the same as those in DGP 1 in Section (ref). Tables (ref)-(ref) report the size properties of the ten inference methods with $d_z=1$ or $3$. In addition, Figures (ref)-(ref) report the power properties of the Wald and AR tests, respectively, with $d_z=1$. Figures (ref)-(ref) report those for the Wald and AR tests, respectively, with $d_z=3$.

Furthermore, we investigate the relationship between the performance of alternative procedures and the number of clusters $J$. Specifically, we plot the size results as a function of $J$ (from $J=6$ to $30$) for $d_z=1$ in Figure (ref) and $d_z=3$ in Figure (ref), respectively. In order to handle the cases with relatively large values of $J$ (e.g., $J=30$), we let the cluster heterogeneity parameter $r$ equal to 3 for $d_z=1$ and 2.4 for $d_z=3$, respectively, so that inference methods that are based on cluster-level estimators (IM and CRS) can also be implemented even with $J=30$.\footnote{As the total number of observations $n$ is equal to 500, the smallest and largest cluster sizes are 2 and 63 for $r=3$ (4 and 56 for $r=2.4$), respectively, under $J=30$.} We focus on DGP 1 in (ref) of the main text with $\Pi=0.25$ and $\rho_{\epsilon v}=0.5$ throughout these simulations.

We observe that similar to the findings in previous simulations, IM and CRS tests can have considerable over-rejections when the number of clusters is relatively large, while ASY and BCH tests typically have large size distortions when the number of clusters is relatively small. Among the seven Wald-based inference methods, only the three wild bootstrap tests (WRE, W-B, and W-B-S) have reasonable size control across different values of $J$. Although slightly lower than the nominal level when $J$ is small, their null rejection frequencies tend to become rather close to $10\%$ as the number of clusters becomes sufficiently large (e.g., when $J$ is larger than 20). We also notice that all the three AR-based inference methods (AR-ASY, AR-B, and AR-B-S) tend to control size. The two bootstrap AR tests over-reject slightly when $J$ is small, while the asymptotic AR test (AR-ASY) can be rather conservative in the over-identified case (Figure (ref)).

Finally, we investigate the power properties of alternative procedures as a function of $J$, from $J=6$ to $30$. Similar to previous simulations, we focus on the power performance of the inference methods that have good size control (namely, WRE, W-B, W-B-S, AR-B, AR-B-S, and AR-ASY). The true value of $\beta$ is set equal to 1 throughout the simulations. Figures (ref) and (ref) report the results for $d_z=1$ with $\beta_0 = 3$ and $-1$, respectively. In addition, Figures (ref) and (ref) report those for $d_z=3$ with $\beta_0 = 3$ and $-1$, respectively. We highlight several observations below. First, the power of all the six inference methods tends to increase with the number of clusters. Second, our recommended procedure W-B-S has the highest power among the six methods across different values of $J$. Third, among the AR tests, the unstudentized bootstrap AR test (AR-B) has the best power performance while AR-ASY has the lowest power. The three observations are in line with our theoretical results in the main text.

table[table omitted — 6,940 chars of source]
table[table omitted — 6,907 chars of source]

Further Simulation Results for DGP 2

In this section, we present the power results for DGP 2. Specifically, Figures (ref) and (ref) present the results for the Wald and AR tests, respectively, with average $J$ equal to 5.437 or 20.189. In addition, Figures (ref) and (ref) present those with average $J$ equal to 10.437 or 29.276.

Simulation Results for DGPs 3 and 4

In this section, we introduce two more simulation designs: DGPs 3 and 4.

DGP 3. We consider a time series IV model.

align*[align* omitted — 401 chars of source]

where $\iota_{d_z}$ is a $d_z$-dimensional vector of ones. We vary the IV strength $\Pi \in \{0.25, 0.375, 0.5\}$, the number of IVs $d_z \in \{1,5\}$, and the AR(1) parameter $\rho_{_{TS}} \in \{0.3, 0.5, 0.7\}$. We let $\beta=1$ and $\gamma =0$. The number of observations is set at 160. The simulation design is similar to the one for time series data considered by Ibragimov-Muller(2010), Bester-Conley-Hansen(2011), and Canay-Romano-Shaikh(2017). We divide the data into $J$ consecutive blocks (clusters) of equal size, where $J \in \{8, 10, 16\}$.

DGP 4. We consider a spatial IV model.

align*[align* omitted — 89 chars of source]

where $\beta = 1$ and $\gamma = 0$. The number of observations is 160 and the observations lie within a square of dimension $40 \times 40$. The location for each observation $s_i = (s_{i_1}, s_{i_2})$ are drawn uniformly within the square, i.e., we let $s_{i_1} \sim U[0, 40]$ independently of $s_{i_2} \sim U[0, 40]$. The distance between observations at locations $s_i$ and $s_j$ is Euclidean: $d_{ij} = \sqrt{(s_{i_1} - s_{j_1})^2 + (s_{i_2} - s_{j_2})^2}$. The simulation design is similar to those considered in Bester-Conley-Hansen(2011), sun2015, and lee2016. The errors $(u_i, v_i)^{\top}$ are drawn from $N\left(

pmatrix[pmatrix omitted — 21 chars of source]

,

pmatrix[pmatrix omitted — 33 chars of source]

\right)$. In addition, $u_i$ and $u_j$ have a correlation equal to $\rho_{_{SP}}^{d_{ij}}$ between them. Similarly, $v_i$ and $v_j$ have a correlation equal to $\rho_{_{SP}}^{d_{ij}}$. Furthermore, the IVs $Z_i$ are drawn from $N(0, I_{d_z})$, with a correlation $\rho_{_{SP}}^{d_{ij}}$ between each element of $Z_i$ and its corresponding element of $Z_j$. Therefore, $\rho_{SP}$ controls the degree of spatial dependence and we let $\rho_{_{SP}} \in \{0.3, 0.5, 0.7\}$. We vary the IV strength $\Pi \in \{0.25, 0.375, 0.5\}$ and the number of IVs $d_z \in \{1,5\}$. To implement inference, we split the data into $J$ groups (clusters) with equal area, where $J \in \{6, 9, 12\}$.\footnote{Specifically, for $J=9$, we split the data into 9 groups formed by splitting into three equal parts both vertically and horizontally. For $J=6$, we split into two equal parts vertically and three horizontally, while for $J=12$, we split into three equal parts vertically and four horizontally.}

table[table omitted — 6,947 chars of source]
table[table omitted — 6,944 chars of source]
table[table omitted — 6,947 chars of source]
table[table omitted — 6,944 chars of source]

Results for DGP 3. Tables (ref)-(ref) report the size results for the time series IV model with $d_z=1$ and $d_z=5$, respectively. We observe that the standard Wald test (ASY) does not control size when the number of clusters is small, especially in the over-identified case. For example, when $d_z=5$, $J=8$, and $\Pi=0.25$, its null rejection frequencies are $0.183, 0.196$, and $0.251$ for $\rho_{TS} = 0.3, 0.5,$ and $0.7$, respectively. BCH improves upon ASY but still has null rejection frequencies equal to $0.135$, $0.143$, and $0.186$, respectively, under such settings. On the other hand, IM and CRS typically have large size distortion when the number of clusters is large. For example, when $d_z=5$, $J=16$, and $\Pi=0.25$, the null rejection frequencies for IM are $0.259, 0.259,$ and $0.249$, for $\rho_{TS} = 0.3, 0.5,$ and $0.7$, respectively, while those for CRS are $0.267, 0.266,$ and $0.256$. By contrast, the wild bootstrap-based Wald inference methods (WRE, W-B, W-B-S) have good size control regardless of the number of clusters and the number of IVs. Therefore, we focus on the bootstrap methods when comparing power. Figures (ref) and (ref) report the power properties of the Wald and AR tests, respectively, with $d_z=1$. We find that the power curves are very similar to those in DGPs 1 and 2. Overall, our W-B-S has the best power, especially against distant alternatives.

table[table omitted — 7,690 chars of source]
table[table omitted — 7,749 chars of source]

Results for DGP 4. Tables (ref)-(ref) report the size for the spatial IV model with $d_z=1$ and $d_z=5$, respectively. The patterns are similar to those observed in other DGPs. ASY and BCH tend to have relatively large size distortions when the number of clusters $J$ is small, while IM and BCH tend to have large distortions when $J$ is large. In addition, the size distortions of these four inference methods increase when the number of IVs becomes large ($d_z=5$). By contrast, WRE, W-B, and W-B-S have good size control across different settings. Figures (ref) and (ref) report the power for the Wald and AR tests, respectively. Again, we see that overall W-B-S has the best power among these tests.

table[table omitted — 7,751 chars of source]
table[table omitted — 7,751 chars of source]
figure[figure omitted — 9,652 chars of source]
figure[figure omitted — 9,652 chars of source]
figure[figure omitted — 9,652 chars of source]
figure[figure omitted — 9,698 chars of source]
figure[figure omitted — 9,698 chars of source]
figure[figure omitted — 8,311 chars of source]