EconBase
← Back to paper

A Dimension-Agnostic Bootstrap Anderson-Rubin Test For Instrumental Variable Regressions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

97,031 characters · 15 sections · 183 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Dimension-Agnostic Bootstrap Anderson-Rubin Test For Instrumental Variable Regressions

abstract{Weak-identification-robust tests for instrumental variable (IV) regressions are typically developed separately depending on whether the number of IVs is treated as fixed or increasing with the sample size, forcing researchers to make a stance on the asymptotic behavior, which is often ambiguous in practice. This paper proposes a bootstrap-based, dimension-agnostic Anderson�Rubin (AR) test that achieves correct asymptotic size regardless of whether the number of IVs is fixed or diverging, and even accommodates cases where the number of IVs exceeds the sample size. By incorporating ridge regularization, our approach reduces the effective rank of the projection matrix and yields regimes where the limiting distribution of the AR statistic can be a weighted chi-squared, a normal, or a mixture of the two. Strong approximation results ensure that the bootstrap procedure remains uniformly valid across all regimes, while also delivering substantial power gains over existing methods by exploiting rank reduction.}

Keywords: Anderson-Rubin Test, Weak Identification, Bootstrap, High Dimension, Quadratic Form, Ridge Regularization

JEL Classification: C12, C36, C55

\setcounter{oldtocdepth}{\value{tocdepth}} \addtocontents{toc}{\setcounter{tocdepth}{-10}}

Introduction

Weak and numerous instruments remain persistent concerns in instrumental variable (IV) regressions across various fields. Surveys by Andrews-Stock-Sun(2019) and lee2021 (lee2021) find that a considerable number of IV regressions in the American Economic Review report first-stage F-statistics below 10. In addition, empirical studies often involve many instruments, such as the 180 IVs used by Angrist-Krueger(1991) to examine the effect of schooling on wages. In “judge design” studies, the number of instruments (number of judges) is typically proportional to the sample size (MS22, MS22).\footnote{E.g., see kling2006, doyle2007, dahl2014, dobbie2018, sampat2019, agan2023, frandsen2023, chyn2024 and the references therein.} Similar patterns of many IVs occur in Fama-MacBeth regressions (fama1973, fama1973; shanken1992, shanken1992), shift-share IVs (Goldsmith(2020), Goldsmith(2020)), wind-direction IVs (deryugina2019mortality, deryugina2019mortality; bondy2020crime, bondy2020crime), granular IVs (gabaix2024, gabaix2024), local average treatment effect estimation (blandhol2022, blandhol2022; boot2024, boot2024; sloczynski2024should, sloczynski2024should), and Mendelian randomization (davey2003, davey2003; davies2015, davies2015).

However, existing weak-identification-robust inference methods for IV regressions are either based on an asymptotic framework in which the number of instruments $K$ is treated as fixed\footnote{E.g., see Staiger-Stock(1997), Stock-Wright(2000), Kleibergen(2002), Kleibergen(2005), Moreira(2003), Andrews-Cheng(2012), Andrews-Mikusheva(2016), Andrews(2018), Andrews-Guggenberger(2019), Moreira-Moreira(2019), among others.} or on an alternative one that allows $K_n$ to diverge to infinity with the sample size $n$.\footnote{E.g., see Andrews-Stock(2007), Newey-Windmeijer(2009), anatolyev2011, crudu2021, MS22, matsushita2024, LWZ(2023), DKM24, among others.} These methods compare distinct test statistics with distinct critical values so that procedures formulated under the fixed-$K$ asymptotics generally do not have correct size control under the diverging-$K$ asymptotics and vice versa. An empirical researcher is, therefore, forced to take a stance on the asymptotic regime of the number of instruments to implement them, which can be ambiguous in many empirical applications. For example, when $K$ is moderate compared with $n$ (e.g., $K=10$ and $n=200$), it is unclear which test the researcher should use. Furthermore, as we will see below, a third asymptotic regime may arise with the use of regularization, making dimension-robust inference even more challenging.

Motivated by this issue, we propose a bootstrap-based, dimension-agnostic AR test. First, by deriving strong approximations for the proposed test statistic and its bootstrap counterpart (under both the null and alternative hypotheses), we show that the new bootstrap test has a correct asymptotic size, regardless of whether the number of IVs $K$ is fixed or diverging. Our proof, which relies on the Lindeberg swapping strategy, contributes a general result on the strong approximation for quadratic forms with independent and heteroskedastic errors. Additionally, our (conditional) strong approximation derivation for bootstrap statistics involving quadratic forms is novel and may be of independent interest. Second, by employing a ridge-regularized projection matrix, our AR test remains valid in high-dimensional cases where $K$ exceeds the sample size $n$. Third, the characterization of the errors in strong approximation offers a theoretically sound basis for selecting the ridge regularizer without taking a stance on the specific regime of $K$. Our choice of the regularizer also helps to reduce the rank of the projection matrix, which can potentially improve the power performance of the test. Fourth, we show that depending on the asymptotic behavior of both $K$ and $K_\lambda$ (the effective rank of the regularized projection matrix), the limit distribution of the test statistic can be (1) normal, (2) weighted chi-squared, or (3) a mixture of weighted chi-squared and normal distributions. Given the strong approximation result, our bootstrap inference remains uniformly valid regardless of the asymptotic regimes, and we further provide its power properties under each scenario. Fifth, the strong approximation and uniform inference results are all established when the number of control variables is allowed to diverge at the same rate or even faster than $\sqrt{n}$. Sixth, simulation experiments and an empirical application to the dataset of card(2009) confirm the excellent size and power properties of our bootstrap test compared to alternative methods.

Relation to the literature: For weak-identification-robust inference based on the classical AR test, Andrews-Stock(2007) showed its validity under many instruments, but requires the number of instruments to diverge more slowly than the cube root of the sample size $n$ ($K^3/n \rightarrow 0$). Newey-Windmeijer(2009) proposed a GMM-AR test under many (weak) moment conditions but imposed the same rate condition on $K$. anatolyev2011 constructed a modified AR test that allows $K$ to be proportional to $n$ but requires homoskedastic errors, and Kaffo-Wang(2017) proposed a bootstrap version of their test. For estimation with many instruments, carrasco2012, carrasco2015, carrasco2016efficient, hansen2014, and carrasco2017 proposed regularization approaches for two-stage least squares, limited information maximum likelihood, and jackknife IV (Angrist(1999), Angrist(1999)) estimators. Furthermore, carrasco2016 first proposed a ridge-regularized AR test that allows for $K$ being larger than $n$ with homoskedastic errors. Bun-Farbmacher-Poldermans(2020) compared the centered and uncentered GMM-AR test and identified a missing degrees-of-freedom correction when $K/n \rightarrow 0$. Recently, crudu2021 and MS22 proposed jackknifed versions of the AR test under many instruments and general heteroskedasticity. DKM24 developed a ridge-regularized version of the jackknife AR test, which is further robust to the scenario where $K$ diverges faster than the sample size. However, the jackknife AR tests are based on standard normal critical values that require $K$ to diverge; thus, they may not have the correct size under fixed $K$. tuvaandorj2024 established the validity of a permutation AR test under heteroskedasticity and diverging $K$, requiring $K^3/n \rightarrow 0$. In contrast to the above methods, our test remains valid with heteroskedastic errors uniformly across a broad asymptotic regime for $K$, spanning from fixed to diverging faster than the sample size.

Furthermore, belloni2012 proposed a Lasso-based method for selecting optimal instruments, valid under high-dimensional IVs and heteroskedasticity, but requiring strong identification and sparse first-stage regressions. However, WZ23 showed that both Lasso and debiased Lasso linear regressions can suffer from significant omitted variable bias, even when the coefficient vector is sparse and the sample size exceeds the number of controls. In such cases, the “long regression,” which includes all regressors, often outperforms the Lasso-based methods. kolesar-muller-roelsgaard(2025) similarly recommended using the “long regression" unless the number of regressors is comparable to or exceeds the sample size. belloni2012 also proposed a weak-identification-robust sup-score test that is dimension-agnostic and does not rely on sparsity. Similar to DKM24, our simulation study shows that the power of our ridge-regularized bootstrap AR test matches the sup-score test when IVs have strong but sparse signals while offering substantially more power when the signal is weak but dense. N23 introduced a jackknife version of the Kleibergen(2002)'s K test and combined it with the sup-score test, but his method relies on a sparse $\ell_1$-regularized estimation of $\rho(Z_i)$, the conditional correlation between the endogenous variable and the outcome error. Without the sparsity assumption, the estimation of $\rho(Z_i)$ may be inconsistent when the dimension of $Z_i$ is large. boot-ligtenberg(2023) developed a dimension-robust AR test based on continuous updating, but relied on an invariance assumption. In contrast to the aforementioned approaches, our bootstrap inference procedure accommodates many instruments and heteroskedastic errors, yet does not rely on invariance or sparsity assumptions.

Our paper also relates to the literature on bootstrap inference for IV regressions. It is found in this literature that when implemented appropriately, bootstrap approaches may substantially improve the inference accuracy for IV models, including the cases where IVs may be rather weak.\footnote{E.g., see Davidson-Mackinnon(2008), Davidson-Mackinnon(2010), Davidson-Mackinnon(2014b), Moreira-Porter-Suarez(2009), Wang-Kaffo(2016), Finlay-Magnusson(2019), roodman2019, young2022, and Wang-Zhang2024, among others.} However, no existing study has uniformly established the bootstrap validity with regard to the number of IVs. We fill this gap by deriving strong approximation results for both the test statistic and its bootstrap counterpart. The strong approximation for the AR statistic is related to the analysis of quadratic forms by HS01. Additionally, our results of (conditional) strong approximation for bootstrap statistics with a quadratic form are, based on our best knowledge, new to the literature.

Our test also remains valid even when the number of control variables diverges at a rate of $\sqrt{n}$ or faster, provided it remains of a smaller order than $n$, regardless of whether $K$ is fixed or diverging. As pointed out by Chao(2023) and mikusheva2024weak, the presence of many controls can introduce additional bias in jackknife IV estimators and AR tests. This phenomenon, often referred to as the quadratic barrier (see CJM18; LSMPW24), poses a major challenge for inference. To address this, we design a debiasing procedure for the AR statistic following the construction in CJN18. Furthermore, to achieve valid bootstrap inference under many controls, we explicitly account for the impact of debiasing on the dispersion of the AR statistic by appropriately adjusting the bootstrap statistic.

Lastly, Anatolyev-Solvsten(2023) proposed an analytical dimension-agnostic $F$ test for linear regressions by analyzing the asymptotic behavior of quadratic forms under two distinct regimes: (1) a fixed number of restrictions, resulting in a weighted chi-squared limiting distribution, and (2) a growing number of restrictions, yielding a normal limiting distribution. Their $F$-test is, in principle, applicable to our setting by testing zero restrictions on the IV coefficients in a linear regression under the null, and it is more general in two respects: (1) it accommodates control variables whose dimension can be of the same order as the sample size $n$, and (2) it allows for testing general linear restrictions. However, our bootstrap inference offers several key advantages. First, although we require the number of controls to be of a smaller order than $n$, we allow the number of instruments $K$ to exceed $n$, a case not covered by their framework. Second, our use of ridge regularization reduces the rank of the projection matrix and gives rise to a third asymptotic regime, where $K$ diverges but a Lindeberg-type condition for asymptotic normality fails, resulting in a limiting distribution that is a mixture of weighted chi-squared and normal variables, akin to the regime analyzed in KSS2020 and YGZ24. This regime does not arise in Anatolyev-Solvsten(2023) due to the absence of regularization. Analytical inference in this setting requires knowledge of the number of dominant eigenvalues, which can be a challenging task. In contrast, our bootstrap approach circumvents such difficulty by directly employing the strong approximation and remains uniformly valid across all three regimes. In our simulations, when the number of instruments is proportional to the sample size, the use of ridge regularization places the test statistic in the third asymptotic regime. Our bootstrap inference procedure has excellent size control even in this challenging setting, and further provides substantial power gains compared to alternative methods because of the rank reduction.

Structure of the paper: Section (ref) makes precise the model setup and provides the testing procedure for our dimension-robust AR test statistic. Sections (ref) and (ref) provide the strong approximation results under both null and alternative for our test statistic and its bootstrap counterpart, respectively. We derive the power properties of our test under the fixed-$K$ and diverging-$K$ asymptotics, respectively, in Section (ref). Section (ref) presents the results of Monte Carlo simulations and Section (ref) applies our test to an empirical application. Proofs of the theorems are given in the Supplemental Appendix, along with additional lemmas and simulation results.

Notations: We denote by $[n]$ the set $\{1,\cdots,n\}$, and use $||A||_{op}$ and $||A||_F$ to refer to the operator and Frobenius norms of a matrix $A$, respectively.

Setup and Testing Procedure

Setup

Consider the linear instrumental variable regression

align[align omitted — 180 chars of source]

where $\widetilde X_i$ denotes a scalar endogenous variable and $W_i \in \mathbb{R}^{d_w}$ denotes the exogenous control variables. In addition, we have $K$-dimensional instrumental variables (IVs) denoted as $\widetilde Z_i$, and $ \widetilde{\Pi}_i \equiv \mathbb{E}( \widetilde{X}_i | \widetilde{Z}_i, W_i)$. We stack $\widetilde Z_i^\top$ up and denote the resulting $n \times K_n$ matrix $\widetilde Z$. We define $\widetilde Y \in \mathbb R^n$, $\widetilde X \in \mathbb R^n$, $\widetilde \Pi \in \mathbb R^n$, $\widetilde v \in \mathbb R^n$, and $ W \in \mathbb R^{n \times d_w}$ in the same manner. Throughout the paper, we also allow $d_w$ to diverge to infinity but at rate that is slower than the sample size $n$, i.e., $d_w = o(n)$. We further require $W$ to be of full rank so that its projection matrix $P_W = W(W^\top W)^{-1} W^\top$ is well defined. We allow, but do not require, $K_n$ to increase with $n$. Specifically, the dimension of $Z$ can be fixed, grow proportional to, or even faster than $n$.

We focus on the model with a scalar endogenous variable for two reasons. First, in many empirical applications of IV regressions, there is only one endogenous variable (as can be seen from the surveys by Andrews-Stock-Sun(2019) and lee2021). Second, the strong approximation results derived in Sections (ref) and (ref) extend directly to the general case of full-vector inference with multiple endogenous variables. Additionally, for the dimension-robust subvector inference, one may use a projection approach (Dufour-Taamouti(2005), Dufour-Taamouti(2005)) after implementing our test on the whole vector of endogenous variables.\footnote{Alternative subvector inference methods for IV regressions (e.g., see GKMC(2012), Andrews(2017), and GKM(2019), GKM(2021)) provide a power improvement over the projection approach under fixed $K$. However, whether they can be applied to the current setting is unclear. Also, Wang-Doko(2018) and Wang(2020) show that bootstrap tests based on the standard subvector AR statistic may not be robust to weak identification even under fixed $K$ and conditional homoskedasticity.}

To proceed, we first partial out the exogenous control variables $W$ from our IV regressions. Specifically, we stack up $(Y_i,X_i,e_i,\Pi_i,v_i)$ to $(Y,X,e,\Pi,v)$, which are defined as $Y = M_W \widetilde{Y}$, $X = M_W \widetilde{X}$, $\Pi = M_W \widetilde{\Pi}$, $e = M_W \widetilde{e}$, and $v = M_W \widetilde{v}$, where $M_W = I_n - P_W$ and $I_n$ is an $n \times n$ identity matrix. In addition, we define $Z = M_W \widetilde{Z}$. Then, (ref) can be rewritten as

align[align omitted — 85 chars of source]

Throughout our analysis, we treat $(Z,W)$ as fixed, which is equivalent to taking all expectations and probability measures conditionally on $(Z,W)$.

Test Statistic

Given that we allow $K_n$ to be greater than $n$, the matrix $Z^\top Z$ is not necessarily invertible. Therefore, we define $P_\lambda = Z (Z^\top Z + \lambda I_{K_n})^{-1}Z^\top$ as the ridge-regularized projection matrix of $Z$ with some ridge penalty $\lambda$ that will be chosen based on $Z$ only. As we treat the instruments and control variables as fixed, so are the ridge-regularizer $\lambda$ and matrix $P_\lambda$. The $(i,j)$ element of $P_{\lambda}$ is denoted as $P_{\lambda,ij}$. Further denote $e_i(\beta_0) = Y_i - X_i \beta_0$. Then, our dimension-agnostic AR test statistic is written as

align[align omitted — 278 chars of source]

where $\kappa = (M_W \circ M_W)^{-1}$,\footnote{Here $\circ$ denotes the Hadamard product and $M_W \circ M_W$ is invertible as long as $d_w < n/2$ as shown by CJN18.} $A_{\lambda,ii} = 2 P_{\lambda,ii} P_{W,ii} - B_{\lambda,ii}$, $B_{\lambda,jk} = \sum_{i \in [n]} P_{W,ik} P_{W,ij} P_{\lambda,ii} = [P_W D_{\lambda} P_W]_{jk}$,

align*[align* omitted — 108 chars of source]
align[align omitted — 108 chars of source]

and

align*[align* omitted — 179 chars of source]

In particular, we can regard $K_\lambda$ as the effective rank under the ridge regularization.

We note that under the null (i.e., $\beta = \beta_0$), the first quadratic term of $\widehat Q(\beta_0)$ in (ref) does not have an exact zero mean due to partialling out controls. This bias is not asymptotically negligible when the dimension of the controls ($d_w$) is of the order $\sqrt{n}$ or greater. The second term of $\widehat Q(\beta_0)$ in (ref), inspired by the variance estimator proposed by CJN18, is used to correct such a bias.

The regularizer $\lambda$ is chosen as

align[align omitted — 330 chars of source]

where $P_{\theta} = Z (Z^\top Z + \theta I_{K_n})^{-1}Z^\top$, $P_{\theta, ij}$ is the $(i,j)$ entry of $P_{\theta}$, $\overline \theta = ||Z^{\top}Z||_{op}$, $K_{\theta}$ is defined in (ref) with $\lambda$ replaced by $\theta$, while $c_1$ and $c_2$ are two positive constants chosen by the researcher such that $c_1$ is sufficiently small. We view $\frac{c}{0} = + \infty$ for any $c>0$. In practice, we use $c_1 = 0.1$ and $c_2 = 1$.\footnote{We have done extensive simulations and find that the results of our test are not sensitive to the specific choice of $c_1$ and $c_2$. The simulation results with alternative choices of $c_1$ and $c_2$ are reported in the Supplemental Appendix.} If there is no $\lambda$ that satisfies both inequalities in (ref), then we choose

align*[align* omitted — 186 chars of source]
remSeveral remarks regarding the choice of the regularizer are in order. First, we select the regularizer $\lambda$ as the largest value over the interval $[0,\overline \theta]$ that both \begin{align*} \frac{\max_{i \in [n]}P_{\lambda,ii}^2}{K_\lambda} \left(1+\sum_{i \in [n]} P_{W,ii}^2 \right) \quad and \quad \frac{\max_{i \in [n]}\sum_{j \in [n], j \neq i}\Xi_{\lambda,ij}^2}{K_\lambda} \end{align*} remain small. These are two critical conditions for ensuring the validity of our strong approximation results for both the test statistic and the bootstrap critical value. We will discuss the theoretical properties of these two terms in detail below. Moreover, since the choice of $\lambda$ depends solely on the instruments, which are treated as fixed (i.e., non-random, or conditioned upon), it does not introduce any model selection bias. Second, given that the conditions for strong approximation are satisfied, we choose the regularizer $\lambda$ as large as possible over $[0, \bar\theta]$. Such a choice is inspired by carrasco2012, carrasco2015, carrasco2016efficient, and carrasco2017, who showed that their proposed regularized IV estimators can be more efficient than those without regularization by employing a sufficiently large value of the regularizer relative to the overall instrument strength (concentration parameter).\footnote{For example, see Proposition 1 of carrasco2012, Proposition 2 of carrasco2015, and Proposition 2 of carrasco2016efficient, in which regularized IV estimators are shown to achieve the semiparametric efficiency bound under homoskedastic errors, given a sufficiently large value of the regularizer relative to the concentration parameter.} Third, our choice of the upper bound for $\lambda$ as $\overline \theta = ||Z^{\top}Z||_{op}$ is motivated by the fact that the ridge regularization transforms the eigenvalues of $Z^\top Z$. Specifically, consider the case where $K_n \leq n$ and the singular value decomposition of $Z$ as $Z = \mathcal U \mathcal S \mathcal V^\top,$ where $\mathcal U \in \Re^{n \times K_n}$ with $\mathcal U^\top \mathcal U = I_{K_n}$, $S = \text{diag}(s_1,\cdots,s_{K_n})$ is a diagonal matrix of non-zero singular values in descending order, $\mathcal V \in \Re^{K_n \times K_n}$, and $\mathcal V^\top \mathcal V = I_{K_n}$. Then, the regularized projection matrix is given by \begin{align*} P_\lambda = \mathcal U diag\left(\frac{s_1^2}{s_1^2 + \lambda},\cdots,\frac{s_{K_n}^2}{s_{K_n}^2 + \lambda} \right) \mathcal U^\top. \end{align*} If for some $k \in [K_n]$, the ratio $s_k/s_1$ is close to zero, then choosing $\lambda$ on the order of $s_1^2 = ||Z^{\top}Z||_{op}$ will cause the $k$-th singular value of $P_\lambda$ (i.e., $\frac{s_k^2}{s_k^2 + \lambda}$) to be close to zero. Intuitively, a large $\lambda$ attenuates the contributions of directions associated with small singular values, effectively reducing the rank of $P_\lambda$ and helping to improve the power performance of our test. We will give more details on this point in Section (ref) (e.g., see Remark (ref)).

Bootstrap Critical Value

To implement the dimension-agnostic test, we propose to use bootstrap critical values. Specifically, let $\{\eta_i\}_{i \in [n]}$ be an independent sequence of random variables with zero mean and unit variance that are generated independently from the samples. Our bootstrap AR test statistic is denoted as $\widehat{Q}^*(\beta_0)$ and defined as

align[align omitted — 190 chars of source]

Then, the bootstrap critical value is denoted as $\widehat{\mathcal C}^*_{\alpha}(\beta_0)$ and defined as the $(1-\alpha)$-th percentile of $\widehat{Q}^*(\beta_0)$ conditional on data, where $\alpha$ is the nominal level of rejection under the null. We reject the null hypothesis of $\beta = \beta_0$ if $\widehat{Q}(\beta_0) > \widehat{\mathcal C}^*_{\alpha}(\beta_0).$

remUnlike the first term of $\widehat Q(\beta_0)$ defined in (ref), we use $\Xi_\lambda$ instead of $P_\lambda$ to define the bootstrap AR statistic. Note that under the null, $e(\beta_0) = M_W \tilde e$, whose elements are not independent from each other. When the dimension of controls $d_w$ diverges at a rate $\sqrt{n}$ or higher, such a cross-sectional dependence is not asymptotically negligible. However, the bootstrap multipliers $\{\eta_i\}_{i \in [n]}$ are independent and, thus, unable to mimic the dependence. Instead, we explicitly account for this difference by adjusting the middle matrix $P_\lambda$ in the original statistic to $\Xi_\lambda$, so that valid bootstrap inference can be achieved under many controls. Additionally, we impose the null on the bootstrap data generating process, following the recommendations in the literature of bootstrap for IV regressions or non-homoskedastic errors, such as Cameron(2008), Davidson-Mackinnon(2010), Roodman-Nielsen-MacKinnon-Webb(2019), and mackinnon2023fast, among others.
remAs pointed out by anatolyev2011 and MS22, when $K$ is fixed, no regularization is used, and the errors are homoskedastic, the test statistic admits the usual re-centered chi-squared approximation: \begin{align*} \frac{\widehat{Q}(\beta_0)}{c_n} \rightsquigarrow \frac{\chi^2_K - K}{\sqrt{2K}} \end{align*} for some normalization scalar $c_n$ computed under homoskedasticity. Furthermore, MS22 noted that this re-centered chi-squared distribution converges quickly to the standard normal distribution as $K$ increases. This suggests that critical values based on $\frac{\chi^2_K - K}{\sqrt{2K}}$ remain valid whether $K$ is fixed or diverging, making it a dimension-agnostic strong approximation for (the re-scaled) $\widehat{Q}(\beta_0)$ under homoskedasticity. In this paper, we extend this idea to the heteroskedastic setting by deriving a weighted re-centered chi-squared approximation for $\widehat{Q}(\beta_0)$ and establishing conditions under which a bootstrap critical value yields valid inference uniformly across different asymptotic regimes. In doing so, we also accommodate a diverging number of controls and ridge regularization, which allows the number of instruments $K$ to exceed the sample size and provides power improvement as well.
remOur proposed bootstrap test is AR-based. It is possible to extend our dimension-agnostic inference procedure to score-based Lagrangian Multiplier (LM) tests provided that the first-stage residual $\tilde v$ is consistently estimable. Given the consistency of residuals, we conjecture that our bootstrap inference remains valid for score-based statistics, including the cases where the effect of the endogenous variable $\tilde X$ may be heterogeneous and the structural equation (ref) is thus misspecified.\footnote{In such settings of heterogeneous treatment effects, especially when the number of instruments diverges with the sample size, researchers typically assume that the reduced-form models for both endogenous variables $\tilde Y$ and $\tilde X$ are linear; see, for example, kolesar2018, EK18, boot2024, and yap2024inference. In such cases, we require the consistency of reduced-form residuals for the bootstrap validity.} Specifically, this may require restricting the dimension of $(W,\tilde Z)$ to be of a smaller order of $n$, imposing some sparsity conditions, and/or assuming that the reduced form regressions for $(\tilde Y, \tilde X)$ are approximately linear. One advantage of our AR-based inference procedure is that it imposes minimal assumptions on the first stage. For instance, we do not have any restriction on $\tilde\Pi$, aligned closely with the setting in MS22.

Strong Approximation of the Test Statistic

This section is concerned with the conditions under which the null distribution of the test statistic defined in (ref) can be approximated by its bootstrap counterpart, no matter whether the dimension $K_n$ of the IVs is fixed or diverging with the sample size. We make the following assumptions on the data-generating process (DGP) to establish this result.

ass\begin{enumerate} • Suppose (ref) holds in which $W$ and $Z$ are treated as fixed, $\{\tilde e_i,\tilde v_i\}_{i \in [n]}$ are independent, mean zero, but potentially heteroskedastic. • There exist constants $C \in (0,\infty)$ and $q > 6$ such that $\max_{i \in [n]} \mathbb E (\tilde e_{i}^{2q} + X_i^{2q}) \leq C$. • Let $\tilde \sigma_i^2 = \mathbb E \tilde e_i^2$. Then, there exist constants $\infty > \bar c > \underline c>0$ such that $$\bar c \geq \max_{i \in [n]} \tilde \sigma_i^2 \geq \min_{i \in [n]} \tilde \sigma_i^2 \geq \underline c.$$ • The matrix $W^\top W$ is invertible and $\max_{i \in [n]} P_{W,ii} = o(1)$, where $P_{W,ii}$ denotes the i-th diagonal element of the projection matrix $P_W$. • Suppose $p_n = \max_{i \in [n]} \frac{ \sum_{j \in [n], j \neq i} \Xi_{\lambda,ij}^2}{K_\lambda}$ and ${p_n'} = \max_{i \in [n]} \frac{ P_{\lambda,ii}^2}{K_\lambda}$. Then, we have $p_n n^{3/q} = o(1)$ and ${p_n'} (1+\sum_{i \in [n]} P_{W,ii}^2) = o(1)$. • Suppose that $\{\eta_i\}_{i \in [n]}$ are i.i.d. and independent of data, have mean zero, unit variance, and sub-Gaussian tail in the sense that $\inf\left\{u>0: \mathbb{E}\exp\left(\frac{|\eta|}{u}\right)^2\leq 2\right\} \leq C <\infty$ for some fixed constant $C \in (0,\infty)$. \end{enumerate}

Assumptions (ref).1--(ref).3 are standard regularity conditions. Assumption (ref).4 allows the dimension of control variables to diverge at a rate that is slower than the sample size, i.e., $d_w = o(n)$. The impact of partialling out $W$ from both $Y$ and $X$ becomes asymptotically negligible only when $d_w = o(\sqrt{n})$, reflecting a broader phenomenon commonly referred to as the quadratic barrier. See, for example, CJM18 and LSMPW24 for further discussions. We overcome this barrier and establish bootstrap validity by carefully debiasing the AR statistic and further adjusting the middle matrix of the bootstrap quadratic form (as noted in Remark (ref)). To the best of our knowledge, this is the mildest rate condition regarding the number of controls established for bootstrap inference with high-dimensional IVs (without imposing a sparsity assumption). We note that analytical inference remains feasible even when $d_w$ is proportional to $n$, as demonstrated in Anatolyev-Solvsten(2023). However, in such a high-dimensional control setting, our current bootstrap inference procedure may fail to control size. At present, it is unclear whether any valid resampling-based inference method exists in this regime, let alone one that remains valid uniformly over the dimensions of both $Z$ and $W$. We leave this important question for future research. In the following, we provide further comparisons between analytical and bootstrap inference approaches in Remarks (ref) and (ref).

Assumption (ref).5 requires that $p_n$ and ${p_n'}$ vanish sufficiently fast. Consider the case without ridge regularization (i.e., $\lambda=0$) and where the projection matrix is well-defined (i.e., $K_n < n$). If the diagonal elements of $P_\lambda$ (with $\lambda = 0$) are well-balanced in the sense that $P_{\lambda,ii} = K_n/n$, then we have

align*[align* omitted — 140 chars of source]

This implies $p_n = O(n^{-1})$ and ${p_n'} = O(n^{-1})$. Importantly, we note that these results hold regardless of whether $K_n$ is fixed or increasing with $n$. If $P_W$ is also well-balanced such that $\max_{i \in [n]} P_{W,ii} \leq Cd_w/n$, then

align*[align* omitted — 86 chars of source]

as long as $d_w = o(n)$. In the minimum, even we only have ${p_n'} =o(1)$, if $d_w = O(\sqrt{n})$, then $\sum_{i \in [n]}P_{W,ii}^2 = O(1)$, which still guarantees that ${p_n'} (1 + \sum_{i \in [n]}P_{W,ii}^2)=o(1)$.

These calculations imply that our inference procedure remains valid even when the number of control variables diverges at the rate $\sqrt{n}$ or faster. Moreover, in high-dimensional settings where $K_n > n$, our ridge-regularized approach with the choice of $\lambda$ in (ref) ensures that Assumption (ref).5 holds provided $q>6$.

Finally, Assumption (ref).6 requires the bootstrap weights $\eta_i$ to have sub-Gaussian tails. In practice, we recommend using standard normal or Rademacher random variables, both of which satisfy this condition.

To proceed, we need to introduce some more notation. Define $\Delta = \beta - \beta_0$, $\tilde \tau_i = \mathbb E(\tilde e_i \tilde v_i)$, $\tilde \varsigma_i^2 = \mathbb E \tilde v_i^2$, $\breve e_i(\beta_0) = \tilde e_i(\beta_0) + \Pi_i \Delta$, and $\tilde e_i(\beta_0) = \tilde e_i + \tilde v_i \Delta $. Then, we denote

align*[align* omitted — 405 chars of source]

In addition, let

align[align omitted — 217 chars of source]

and

align[align omitted — 199 chars of source]

where $\{g_{i}\}_{i \in [n]}$ are i.i.d. standard normal random variables that are generated independent of data. The following theorem shows that our proposed AR test statistic $\widehat Q(\beta_0)$ can be strongly approximated by $Q(\beta_0) + C(\Delta)$ in Kolmogorov distance, where

align*[align* omitted — 172 chars of source]

Furthermore, Theorem (ref) in the next section shows the bootstrap statistic $\widehat Q^*(\beta_0)$ can be strongly approximated by $Q^*(\beta_0)$ in Kolmogorov distance conditionally on data. Note that $Q^*(\beta_0)$ is equal to $Q(\beta_0)$ under the null hypothesis.\footnote{Under the null, we have $\Delta=0$ and $\tilde \sigma_i(\beta_0) = \breve \sigma_i(\beta_0)$.}

thmSuppose Assumption (ref) holds, and $||\Pi||_2^2 \Delta^2 / \min\left(K_\lambda^{1/2},K_\lambda^{2/3}\right)$ is bounded. Then, we have \begin{align*} \sup_{y\in \Re}\left|\mathbb P(\widehat Q(\beta_0) \leq y)- \mathbb P(Q(\beta_0) + C(\Delta)\leq y)\right| = o(1). \end{align*}
remWe note that $Q(\beta_0)$ is implicitly indexed by the sample size $n$, which explains why we call it a strong approximation rather than a limit of our AR statistic $\widehat Q(\beta_0)$. Second, as noted in Remark (ref), the cross-sectional dependence between the elements of $e(\beta_0)$ is not asymptotically negligible when $d_w$ diverges at a rate $\sqrt{n}$ or higher. On the other hand, $\{g_{i}\}_{i \in [n]}$ in $Q(\beta_0)$ and $Q^*(\beta_0)$ are i.i.d. standard normal random variables. We account for this by adjusting $P_\lambda$ in the original statistic to $\Xi_\lambda$ to (ref)-(ref). Third, we can see that \begin{align} \mathbb E Q(\beta_0) + C(\Delta) = \frac{\sum_{i \in [n]} \sum_{j \in [n], j \neq i} \Pi_i P_{\lambda,ij} \Pi_j \Delta^2 }{\sqrt{K_\lambda}}, \end{align} which is the non-centrality parameter for the AR statistic under the alternative.
remThe strong approximation $Q(\beta_0) + C(\Delta)$ encompasses three asymptotic regimes in a unified framework: (1) when both $K$ and $K_\lambda$ are bounded, $Q(\beta_0) + C(\Delta)$ asymptotically follows a weighted non-central chi-squared distribution; (2) when both $K$ and $K_\lambda$ diverge so that a Lindeberg-type condition holds, it converges in distribution to a normal random variable; and (3) when $K$ diverges but $K_\lambda$ is bounded, it converges to a mixture of a weighted sum of non-central chi-squared distributions and a normal distribution. These three regimes are discussed separately by KSS2020 in the setting of estimation of variance components. For testing linear restrictions, Anatolyev-Solvsten(2023) proposed an analytical inference procedure that is valid under regimes (1) and (2). However, in their setting, where ridge regularization is not employed, the third regime does not arise. A key advantage of our bootstrap procedure and the associated strong approximation results is that they are valid irrespective of the asymptotic regime, including the challenging case with a mixture of distributions. We will provide further details on the regimes in Section (ref).
remTheorem (ref) is valid under both the null and alternative hypotheses. The nature of the alternatives depends on the magnitude of $||\Pi||_2^2 \Delta^2 / \min\left(K_\lambda^{1/2},K_\lambda^{2/3}\right)$. When $K_\lambda$ is bounded, we have weak (strong) identification when the concentration parameter $||\Pi||_2^2$ is bounded (diverging). When $K_\lambda$ is diverging, as shown by MS22, weak (strong) identification arises when the concentration parameter $||\Pi||_2^2/\sqrt{K_\lambda}$ is bounded (diverging). Under either regime with regard to $K_{\lambda}$, Theorem (ref) accommodates (i) fixed alternatives under weak identification and (ii) local alternatives scaled by the square root of the concentration parameter under strong identification.

Strong Approximation of the Bootstrap Statistic

This section concerns the strong approximation of the bootstrap statistic defined in Section (ref) in Kolmogorov distance conditionally on data. The approximation in Theorem (ref) is the same as that for the original statistic under the null hypothesis, as established in Theorem (ref), which directly implies that the proposed test with bootstrap critical values achieves a correct asymptotic size. Such a result holds no matter whether the dimension of IVs $K_n$ is fixed or diverging to infinity.

thmLet $\mathcal D$ denote all observations in our sample. Suppose Assumption (ref) holds and $\Delta$ is bounded. Then, we have \begin{align*} \sup_{y \in \Re} |\mathbb P( \widehat Q^*(\beta_0) \leq y | \mathcal D) - \mathbb P( Q^*(\beta_0) \leq y) | = o_P(1). \end{align*}
remTheorem (ref) remains valid under both the null and alternative hypotheses. In contrast to Theorem (ref), it accommodates fixed alternatives even in the presence of strong identification (without requiring $||\Pi||_2^2 \Delta^2 / \min\left(K_\lambda^{1/2},K_\lambda^{2/3}\right)$ to be bounded). This distinction originates from the fact that the alternative hypothesis $\Delta$ affects $Q(\beta_0)$ and $Q^*(\beta_0)$ differently -- introducing non-centrality bias in the former and variance in the latter (e.g., see the non-centrality bias in (ref) and the definition of $Q^*(\beta_0)$ in (ref), respectively). The distinction also underpins the power of our dimension-agnostic AR test, which will be analyzed in detail in Section (ref).\footnote{Similar phenomenon of an inflated variance under the alternative hypothesis also occurs with the jackknife AR tests using analytical variance estimators that impose the null hypothesis (e.g., crudu2021, MS22, and DKM24) so that the resulting tests can be robust to weak identification.}
remTheorem (ref) can be viewed as a general result of strong approximation for the bootstrap version of the quadratic forms. The proof extends the Lindeberg swapping strategy mentioned in the previous section. Indeed, compared with Theorem (ref), it is substantially more involved to establish Theorem (ref) because the second moment of the bootstrap statistic $\widehat Q^*(\beta_0)$ conditional on data is random and does not exactly match that of its strong approximation (i.e., $Q^*(\beta_0)$). We rely on the concentration inequalities for quadratic forms (i.e., Hanson-Wright inequality) and linear forms of martingale difference sequence to bound the approximation error in Kolmogorov distance due to the mismatch of the second moments. This technique seems new to the literature and may be of independent interest.
remIn addition, we observe from (ref)-(ref) that $Q(\beta_0)$ and $Q^*(\beta_0)$ have the same marginal distribution under the null hypothesis ($\Delta=0$). This means that, under the null, the bootstrap statistic closely approximates the test statistic in Kolmogorov distance when conditioned on the data, whether $K_n$ is fixed or diverging. This equivalence forms the basis for our bootstrap test to achieve the correct size. To rigorously validate this assertion, the following regularity condition is required.
assDenote $\mathcal C_{\alpha}(\beta_0) = \inf\{y \in \Re: 1-\alpha \leq F_{\beta_0}(y)\}$, where $F_{\beta_0}(y) = \mathbb P(Q(\beta_0) \leq y)$. Let the $\varepsilon$-neighborhood around $\mathcal C_{\alpha}(\beta_0)$ be $\mathbb B(\mathcal C_{\alpha}(\beta_0), \varepsilon)$. Then, under the null, the density of $Q(\beta_0)$ exists in the neighborhood $\mathbb B(\mathcal C_{\alpha}(\beta_0), \varepsilon)$ and is denoted as $f_{n}(\cdot)$. In addition, there exits an $\varepsilon>0$ such that $\liminf_{n \rightarrow \infty} \inf_{y \in \mathbb B(\mathcal C_{\alpha}(\beta_0), \varepsilon)}f_{n}(y) \geq \underline c >0,$ where $\underline c>0$ is a fixed constant.
remWe discuss three asymptotic regimes for power analysis in Section (ref). In each regime, the limiting distribution of $Q(\beta_0)$ is either normal, weighted chi-squared, or a mixture of the two. This ensures that Assumption (ref) holds automatically in all three regimes.
thmSuppose we are under the null hypothesis $\beta = \beta_0$ and Assumptions (ref) and (ref) hold. Then, we have \begin{align*} \mathbb P(\widehat{Q}(\beta_0) > \widehat{\mathcal C}^*_{\alpha}(\beta_0)) \rightarrow \alpha. \end{align*}

Asymptotic Power

In this section, we discuss the power of the bootstrap inference by focusing on three separate cases: (I) both $K$ and $K_\lambda$ diverge, (II) $K$ diverges but $K_\lambda $ is bounded, and (III) both $K$ and $K_\lambda$ are bounded.

The Case with Diverging $K$ and $K_\lambda$

To proceed, we let $\Psi(\beta_0) = 2 \sum_{i \in [n]} \sum_{j \in [n], j \neq i} \tilde \sigma^2_i(\beta_0) \Xi_{\lambda,ij}^2 \tilde \sigma_j^2(\beta_0)/K_\lambda$, where $\tilde \sigma_i^2(\beta_0) = Var(\tilde e_i (\beta_0) ) = \tilde \sigma_i^2 + 2 \Delta \tilde \tau_i + \Delta^2 \tilde \varsigma_i^2$.

ass\begin{enumerate} • $K \rightarrow \infty$, $K_{\lambda} \rightarrow \infty$, and $||\Pi||_2^2 \Delta^2/\sqrt{K_{\lambda}}$ is bounded. • $\Delta$ and $\max_{i \in [n]}|\Pi_i|$ are bounded. • $\Psi^{-1/2}(\beta_0) \frac{ \sum_{i \in [n]} \sum_{j \in [n], j \neq i} \Pi_i P_{\lambda,ij} \Pi_j \Delta^2 }{\sqrt{K_\lambda} } \rightarrow \mu(\beta_0)$. \end{enumerate}
thmSuppose Assumptions (ref) and (ref) hold. Then, we have \begin{align*} \mathbb P(\widehat Q(\beta_0) > \widehat C^*_\alpha(\beta_0)) \rightarrow \mathbb P \left( \mathcal{N}(\mu(\beta_0),1) > z_\alpha \right), \end{align*} where $\mathcal{N}(\mu,1)$ is a normal random variable with mean $\mu$ and unit variance and $z_\alpha$ is the $(1-\alpha)$ quantile of a standard normal random variable.
remLet us denote $(\varpi_1,\cdots,\varpi_n)$ as the eigenvalues of the matrix \begin{align*} diag(\tilde \sigma_1(\beta_0),\cdots,\tilde \sigma_n(\beta_0)) \Xi_{\lambda} diag(\tilde \sigma_1(\beta_0),\cdots,\tilde \sigma_n(\beta_0)), \end{align*} ordered such that $|\varpi_1| \geq |\varpi_2| \geq \cdots \geq |\varpi_n|$. From the proof of Theorem (ref), we note that under the null, \begin{align*} \widehat Q(\beta_0) = \frac{\sum_{i=1}^n (g_i^2 -1)\varpi_i}{\sqrt{K_{\lambda}}} + o_P(1), \end{align*} where $\{g_i\}_{i \in [n]}$ is an i.i.d. sequence of standard normal random variables. Furthermore, we have \begin{align*} \varpi_1 = O(1) \quad and \quad \sum_{i \in [n]} \varpi_i^2 \geq \underline c K_\lambda, \end{align*} for some constant $\underline c > 0$. This implies when $K_\lambda \to \infty$, \begin{align*} \frac{\varpi_1^2}{\sum_{i \in [n]} \varpi_i^2 } = o(1), \end{align*} which is a Lindeberg-type condition that guarantees asymptotic normality of the test statistic, as established in Theorem (ref).
remWhen $d_w = o(\sqrt{n})$, Theorem (ref) holds if we replace $\Xi_{\lambda}$ by $P_{\lambda}$ in the definition of $\Psi(\beta_0)$ as the effect of partialling out controls is asymptotically negligible. If $K_n < n$ and we set $\lambda = 0$ so that $P_\lambda = P$ (i.e., without ridge regularization), the local power of our test is asymptotically equivalent to that of the jackknife AR tests proposed by crudu2021 and MS22.\footnote{crudu2021 and MS22 proposed different variance estimators for the jackknife AR statistic. Under local alternatives characterized under Assumption (ref), the two variance estimators are asymptotically equivalent.} In general, the regularizer $\lambda$ can affect the power through $\mu(\beta_0)$, which depends on $P_{\lambda}$ and $K_\lambda$. Specifically, following Remark (ref), we note that the alternative $\Delta$ affects the limiting distribution through (1) the non-centrality bias of the test statistic $\widehat Q(\beta_0)$, given by \begin{align} \frac{ \sum_{i \in [n]} \sum_{j \in [n], j \neq i} \Pi_i P_{\lambda,ij} \Pi_j \Delta^2 }{\sqrt{K_\lambda} } \equiv \mathbb{C}_\lambda \Delta^2, \end{align} where $\mathbb{C}_\lambda$ denotes the concentration parameter under the ridge regularization, and (2) the variance of the statistic, captured by $\Psi(\beta_0) = \tilde \sigma_i^2 + 2 \Delta \tilde \tau_i + \Delta^2 \tilde \varsigma_i^2$. Both components contribute to the mean $\mu(\beta_0)$ of the limiting distribution. Furthermore, we note that (ref) also motivates our choice of the regularizer $\lambda$. In particular, as we restrict the upper bound $\bar\theta$ for the regularizer to be $||Z^\top Z||_{op}$, the ridge regularization $\lambda I_{K_n}$ will not dominate $Z^\top Z$. Then it is plausible that the numerator of the concentration parameter, i.e., $\sum_{i,j \in [n]^2, i \neq j}\Pi_i P_{\lambda, ij} \Pi_j$ does not change order for the range of $\lambda$ we consider. On the other hand, as $\lambda$ increases, the effective rank $K_\lambda$ decreases, which typically causes the non-centrality in (ref) to increase and thus lead to power improvement. Notice that it is possible for $\mathbb{C}_\lambda$ in (ref) to achieve a higher order of magnitude than the concentration parameter without ridge regularization \begin{align} \frac{ \sum_{i \in [n]} \sum_{j \in [n], j \neq i} \Pi_i P_{ij} \Pi_j }{\sqrt{K}}, \end{align} as long as $K_\lambda = o(K)$ (e.g., $K$ diverges but $K_\lambda$ is fixed). Such an advantage of regularization has been pointed out in previous studies, such as carrasco2015, carrasco2016efficient and carrasco2017. In the next section, we further study in detail the case where $K_\lambda$ is bounded while $K$ diverges.

The Case with Diverging $K$ but Bounded $K_\lambda$

Following Remark (ref), we now consider the case where $K_\lambda$ remains bounded, resulting in the failure of the Lindeberg-type condition for asymptotic normality.

ass\begin{enumerate} • Suppose there exists a fixed positive integer $R$ such that \begin{align*} \frac{\varpi_i}{ (\sum_{j \in [n]} \varpi_i^2)^{1/2} } \rightarrow r_i \neq 0, \quad \forall i = 1,\cdots,R, \quad and \quad \frac{\varpi_{R+1}^2}{ \sum_{i \in [n]} \varpi_i^2 } = o(1). \end{align*} • Denote $(\varpi_1^*,\cdots,\varpi_n^*)$ as the eigenvalues of the matrix \begin{align*} diag(\breve \sigma_1(\beta_0),\cdots,\breve \sigma_n(\beta_0)) \Xi_{\lambda} diag(\breve \sigma_1(\beta_0),\cdots,\breve \sigma_n(\beta_0)), \end{align*} ordered such that $|\varpi_1^*| \geq |\varpi_2^*| \geq \cdots \geq |\varpi_n^*|$. Suppose there exists a fixed positive integer $R^*$ such that \begin{align*} \frac{\varpi_{i}^*}{ \left(\sum_{j \in [n]} \varpi_i^{*2} \right)^{1/2} } \rightarrow r_i^* \neq 0, \quad \forall i = 1,\cdots,R^*, \quad and \quad \frac{\varpi_{R^*+1}^{*2}}{ \sum_{i \in [n]} \varpi_i^{*2} } = o(1). \end{align*} • Suppose $||\Pi||_2^2 \Delta^2/\sqrt{K_{\lambda}}$, $\Delta$, and $\max_{i \in [n]}|\Pi_i|$ are bounded. • $\Psi^{-1/2}(\beta_0) \frac{ \sum_{i \in [n]} \sum_{j \in [n], j \neq i} \Pi_i P_{\lambda,ij} \Pi_j \Delta^2 }{\sqrt{K_\lambda} } \rightarrow \mu(\beta_0)$ and $\frac{\breve \Psi(\beta_0)}{\Psi(\beta_0)} \rightarrow \psi(\beta_0)>0$, where $\breve \Psi(\beta_0) = 2 \sum_{i \in [n]} \sum_{j \in [n], j \neq i} \breve \sigma^2_i(\beta_0) \Xi_{\lambda,ij}^2 \breve \sigma_j^2(\beta_0)/K_\lambda$. \end{enumerate}
thmSuppose Assumptions (ref) and (ref) hold. Then, we have \begin{align*} \mathbb P(\widehat Q(\beta_0) > \widehat C^*_\alpha(\beta_0)) \rightarrow \mathbb P \left(\chi(\{r_i\}_{i \in [R]}) + \mu(\beta_0) > \psi^{1/2}(\beta_0) \mathcal C_\alpha (\{r_i^*\}_{i \in [R^*]}) \right), \end{align*} where the random variable $\chi(\{r_i\}_{i \in [R]})$ has the distribution $$\frac{\sum_{ i \in [R]} (g_i^2 -1) r_i}{\sqrt{2}} + \left(1-\sum_{i \in [R]}r_i^2 \right)^{1/2} g_{R+1},$$ with $\{g_i\}_{i \in [R+1]}$ being i.i.d. standard normal random variables, and $\mathcal C_\alpha (\{r_i\}_{i \in [R]})$ is the $(1-\alpha)$-th quantile of $\chi(\{r_i\}_{i \in [R]})$.
remWhen the Lindeberg-type condition fails due to $K_\lambda$ being bounded, the limiting distribution of our test statistic becomes a mixture of a weighted sum of chi-squared random variables and a standard normal random variable. This is similar to the scenario described in KSS2020 and YGZ24. Analytical inference in this regime is difficult, as it requires estimating the number of dominant eigenvalues $R$ driving the asymptotic distribution or reporting (the union of) confidence intervals corresponding to consecutive values of $R$ (see, e.g., Section 7.2 of KSS2020). A key advantage of our bootstrap inference procedure is that it does not require prior knowledge of the number of dominant eigenvalues $R$ or associated weights $\{r_i\}_{i \in [R]}$, since it is valid regardless of the asymptotic regime. In our simulations in Section (ref) with $K = 160$, with our choice of $\lambda$, we observe one dominant eigenvalue ($R= 1$) and $r_1 = \sqrt{0.948}$. Furthermore, in Section (ref), we observe that $K_\lambda$ is equal to 2.015 and 1.550, respectively, for the specification with 38 and 342 IVs, suggesting that this regime applies to our empirical application of card(2009)'s dataset as well.
remAs mentioned earlier, the alternative $\Delta$ affects the location and scale of the test statistic and the bootstrap critical value, represented by $\mu(\beta_0)$ and $\psi(\beta_0)$, respectively. When $K_\lambda$ diverges, the scale effect becomes asymptotically negligible, as indicated by $\psi(\beta_0) = 1$ in Theorem (ref). However, when $K_\lambda$ is bounded, $\psi(\beta_0)$ may differ from one, and the scale effect remains relevant in the limiting distribution.
remFurthermore, we note that the ridge-regularized concentration parameter $\mathbb{C}_\lambda$ in (ref) can achieve a higher divergence rate than that without regularization, given that $K_\lambda$ is bounded while $K$ diverges. In particular, as established by MS22, $\sum_{i \in [n]}\sum_{j \in [n], j\neq i}\Pi_i P_{ij}\Pi_j /\sqrt{K} \rightarrow \infty$ is required for the jackknife AR test (without regularization) to be consistent. By contrast, with the help of regularization, our test only requires $\sum_{i \in [n]}\sum_{j \in [n], j\neq i}\Pi_i P_{\lambda,ij}\Pi_j \rightarrow \infty$ to be consistent if $K_\lambda$ is bounded.

The Case with Bounded $K$ and $K_\lambda$

In this section, we consider the power property of our bootstrap AR test in the asymptotic framework that the dimension of $Z$ (i.e., $K_n = K$) is fixed. To rigorously state the regularity conditions, we recall the singular value decomposition of $Z$ as $Z = \mathcal U \mathcal S \mathcal V^\top,$ where $\mathcal U \in \Re^{n \times n}$, $\mathcal U^\top \mathcal U = I_n$, $\mathcal S = [ S_0, 0_{K, n-K}]^\top$, $S_0$ is a diagonal matrix of non-zero singular values, $0_{K, n-K} \in \Re^{K \times (n-K)}$ is a matrix of zeros, $\mathcal V \in \Re^{K \times K}$, and $\mathcal V^\top \mathcal V = I_K$. Denote $\mathcal U = [\mathcal U_1,\mathcal U_2]$ such that $\mathcal U_1 \in \Re^{n \times K}$, $\mathcal U_2 \in \Re^{n \times (n-K)}$, $\mathcal U_1^\top \mathcal U_1 = I_K$, $\mathcal U_1^\top \mathcal U_2 = 0_{K,n-K}$, and $\mathcal U_2^\top \mathcal U_2 = I_{n-K}$. Further denote $\Omega(\beta_0) \equiv \mathcal U_1^\top \text{diag}(\tilde \sigma_1^2(\beta_0),\cdots,\tilde \sigma_n^2(\beta_0) ) \mathcal U_1$ and the eigenvalue decomposition

align*[align* omitted — 226 chars of source]

Last, denote $\nu(\beta_0) = \lim_{n \rightarrow \infty} \mathbb U^\top \Omega^{-1/2}(\beta_0) \Delta \mathcal U_1^\top \Pi$.

ass\begin{enumerate} • Suppose the IVs $Z$ have a fixed dimension $K$, $\left(\max_{i \in [n]}P_{\lambda,ii}\right) d_w = o(1)$, and $K_\lambda \geq c$ for some constant $c>0$. • $\max_{i \in [n]}||\mathcal{U}_{1,i}||_2 = o(1)$, where $\mathcal{U}_{1,i}^\top \in \Re^{1 \times K}$ is the $i$-th row of $\mathcal{U}_1$. • $||\Pi||_2^2 \Delta^2 /\sqrt{K_\lambda}$ is bounded. \end{enumerate}

The following theorem establishes our AR test's power property in the fixed $K_\lambda$ scenario.

thmSuppose Assumptions (ref) and (ref) hold. Then, we have \begin{align*} \mathbb P(\widehat Q(\beta_0) > \widehat C^*_\alpha(\beta_0)) \rightarrow \mathbb P \left(\sum_{k\in [K]} \omega_k \chi_k^2(\nu_k^2(\beta_0)) > \mathcal C_{\omega}(1-\alpha) \right), \end{align*} where $\omega = (\omega_1,\cdots,\omega_K)$, $\{ \chi_k^2(\nu_k^2(\beta_0))\}_{k \in [K]}$ is a sequence of independent non-central chi-squared random variables with one degree of freedom and noncentrality parameter $\nu_k^2(\beta_0)$, $\nu_k(\beta_0)$ is the $k$-th element of $\nu(\beta_0)$, $\mathcal C_{\omega}(1-\alpha)$ is the $(1-\alpha)$ quantile of a weighted chi-squared random variable $\sum_{k \in [K]} \omega_k \chi^2_k$, and $\{\chi^2_k\}_{k \in [K]}$ is a sequence of i.i.d. centered chi-squared random variables with one degree of freedom.

Next, we demonstrate that when $K$ is fixed, our dimension-agnostic AR test is (asymptotically) admissible within a specific class of tests, which includes the standard (heteroskedasticity-robust) AR test designed for fixed $K$. Let

align[align omitted — 149 chars of source]

and $\widehat {\mathcal G}_k (\beta_0)$ be the $k$-th element of $\widehat {\mathcal G} (\beta_0)$, where $\hat \Omega(\beta_0)$ is a consistent estimator of $\Omega(\beta_0)$. We observe that, in the scenario where $K$ is fixed, the standard AR test rejects if

align[align omitted — 196 chars of source]

where $\iota_K$ is a $K$-dimensional vector of ones and $\mathcal C_{\iota_K}(1-\alpha)$ is just the $(1-\alpha)$ quantile of the centered chi-squared random variable with $K$ degrees of freedom. On the other hand, our bootstrap AR test is asymptotically equivalent to a test that rejects if

align*[align* omitted — 119 chars of source]

In addition, the proof of Theorem (ref) shows

align*[align* omitted — 212 chars of source]

We consider the class $\Phi_\alpha$ of tests $\phi(\cdot)$ defined as

align*[align* omitted — 362 chars of source]

Both the standard and bootstrap AR tests control size, and thus, belong to this class. The power of any test $\phi(\cdot) \in \Phi_\alpha$ is determined by $\nu(\beta_0) \in \Re^K$.

thmSuppose Assumptions (ref) and (ref) hold. In addition, let $\widehat {\mathcal G}(\beta_0)$ be defined in (ref), $\hat \Omega(\beta_0) \stackrel{p}{\longrightarrow} \Omega(\beta_0)$, and $0<c \leq \lambda_{\min}\left(\Omega(\beta_0) \right) \leq \lambda_{\max}\left(\Omega(\beta_0) \right) \leq C < \infty.$ Then, our bootstrap test $\phi_0 = 1\{ \widehat{Q}(\beta_0) > \widehat{\mathcal C}^*_{\alpha}(\beta_0)\}$ is asymptotically admissible w.r.t. $\Phi_\alpha$ in the sense that if there exists a test $\phi^* \in \Phi_\alpha$ such that for all values of $\nu(\beta_0) \in \Re^K$, \begin{align*} \lim_{n \rightarrow \infty} \mathbb E \phi^*(\widehat {\mathcal G}_1^2 (\beta_0),\cdots,\widehat {\mathcal G}_K^2 (\beta_0)) \geq \lim_{n \rightarrow \infty} \mathbb E \phi_0, \end{align*} then we must have \begin{align*} \lim_{n \rightarrow \infty} \mathbb E \phi^*(\widehat {\mathcal G}_1^2 (\beta_0),\cdots,\widehat {\mathcal G}_K^2 (\beta_0)) = \lim_{n \rightarrow \infty} \mathbb E \phi_0, \end{align*} for all $\nu(\beta_0) \in \Re^K$.
remIt is reasonable to assume there exists a consistent estimator $\hat \Omega(\beta_0)$ for $\Omega(\beta_0)$, which is a $K \times K$ matrix with $K$ fixed.
remBecause the standard AR test defined in (ref) belongs to $\Phi_\alpha$, Theorem (ref) implies our bootstrap test $\phi_0$ is not dominated by the standard AR test for all alternatives. In fact, the standard AR test is also admissible among the tests in $\Phi_\alpha$ so that it is not dominated by $\phi_0$ either. However, our bootstrap test is dimension-robust, while the standard AR test does not have the correct size under the regimes in Sections (ref) and (ref).
remUnder strong identification against local alternatives, the K test proposed by Kleibergen(2002) is the uniformly most powerful unbiased test when the number of IVs is treated as fixed and, thus, dominates both the standard AR and our test. However, the K test is not dimension-robust, similar to the standard AR test. In fact, LWZ(2023) proposed a counterpart of the K test in the setting of many weak instruments with heteroskedastic errors (but it may be invalid under a fixed number of IVs). Furthermore, both the K test and its many-weak-IV counterpart have power ditches, and thus, no power against certain fixed alternatives, even under strong identification (e.g., see Section 3.1 of A16 and Lemma 2.3 of LWZ(2023)).
remN23 proposed a dimension-robust version of the K test, which de-correlates the endogenous variable $X_i$ and outcome error $e_i$ conditionally on $Z_i$. This approach requires consistently estimating the conditional correlation $\rho(Z_i) = \mathbb E(X_i e_i|Z_i)$. However, when the dimension of $Z_i$ is large, in general, $\rho(Z_i)$ cannot be consistently estimated. Instead, N23 imposes a sparsity condition and estimates $\rho(Z_i)$ by an $\ell_1$-regularized regression. According to his simulations, the dimension-robust K test can also suffer from the power ditch issue due to the (null-imposed) decorrelation. Unlike N23's (N23) procedure, our test achieves robustness against the dimension of IVs without imposing any additional structure. Furthermore, if one is comfortable with imposing the sparsity assumption on $\rho(Z_i)$, then it is possible to combine our test and N23's (N23) K test (e.g., by constructing a dimension-robust version of the conditional linear combination test in LWZ(2023), which is efficient under strong identification and also solves the power ditch issue).

Monte Carlo Simulations

This section investigates the finite sample size and power performance of existing tests and our proposed test. To begin, we explicitly define these tests and their corresponding critical values. In addition, following belloni2012 and DKM24, upon obtaining a given instrument set $Z$, we standardize it by $\frac{1}{n} \sum_{i=1}^n Z_{ij}^2 = 1$, for $j=1,...,K.$ Note that the tests described in section (ref) below are based on the standardized $Z$. Throughout the simulations, we set the number of Monte Carlo and bootstrap replications equal to $5,000$ and $10,000$ respectively, and set the nominal level $\alpha=0.05$.

Description of Tests

Specifically, we consider the following eleven tests:

enumerate• BS: Our bootstrap test based on (ref) and (ref), which rejects $H_0$ whenever $\widehat{Q}(\beta_0) > \widehat{\mathcal C}^*_{\alpha}(\beta_0),$ and we let the upper bound defined in (ref) be $\overline{\theta} \equiv || Z^{\top}Z||_{op}$; • JAR$_{\rm std}$: The jackknife AR test based on crudu2021's standard variance estimator for diverging $K$, which rejects $H_0$ whenever \begin{align*} \frac{1}{\sqrt{\widehat{\Phi}^{std}(\beta_0)} \sqrt{K}} \sum_{i \in [n]} \sum_{j \in [n], j \neq i} P_{ij} e_i(\beta_0) e_j(\beta_0) > q_{1-\alpha} \left(\mathcal{N}(0,1)\right), \end{align*} where $\widehat{\Phi}^{std}(\beta_0) := \frac{2}{K}\sum_{i \in [n]}\sum_{j \neq i}P_{ij}^2e_i^2(\beta_0)e_j^2(\beta_0)$ and $P_{ij}$ denotes the $(i,j)$ element of $P:= Z(Z^{\top}Z)^{-1}Z^{\top}$;\footnote{Note that this statistic is slightly different from the one proposed by crudu2021, in that they replace $P_{ij}$ by $C_{ij}$, where $C$ is defined in Section 3.2 of their paper.} • JAR$_{\rm cf}$: MS22's jackknife AR test, which is based on a cross-fit variance estimator for diverging $K$ and rejects $H_0$ whenever \begin{align*} \frac{1}{\sqrt{\widehat{\Phi}^{cf}(\beta_0)} \sqrt{K}} \sum_{i \in [n]} \sum_{j \in [n], j \neq i} P_{ij} e_i(\beta_0) e_j(\beta_0) > q_{1-\alpha} \left(\mathcal{N}(0,1)\right), \end{align*} where $\widehat{\Phi}^{cf}(\beta_0) := \frac{2}{K}\sum_{i \in [n]}\sum_{j \neq i}\frac{P_{ij}^2}{M_{ii}M_{jj} + M_{ij}^2}[e_i(\beta_0)M_ie(\beta_0)][e_j(\beta_0)M_je(\beta_0)]$, $M=I_n-P$, and $M_i$ denotes the $i$th row of $M$;\footnote{In the simulations, the cross-fit variance estimator $\widehat{\Phi}^{cf}(\beta_0)$ can be negative at times. To ensure the JAR$_{\rm cf}$ test is well-defined, we set the variance estimator to be $\max\left(\widehat{\Phi}^{cf}(\beta_0),\frac{1}{\sqrt{n \log(n)}}\right)$.} • AR: The classical heteroskedasticity-robust AR test for fixed $K$, rejecting $H_0$ whenever \begin{align*} J_n^{\top}(\beta_0) \widehat{\Omega}_n(\beta_0)^{-1} J_n(\beta_0) > q_{1-\alpha}(\chi^2_K), \;\; \end{align*} where $J_n(\beta_0) := n^{-1/2} Z^{\top} e(\beta_0)$ and $\widehat{\Omega}_n(\beta_0):= n^{-1} Z^{\top} \{ diag(e_1^2(\beta_0),...,e_n^2(\beta_0) ) \} Z$; • RJAR: The ridge-regularized jackknife AR test for diverging $K$ proposed by DKM24, which rejects $H_0$ whenever \begin{align*} \frac{1}{\sqrt{\widehat{\Phi}_{\gamma_n^*}(\beta_0)} \sqrt{r_n}} \sum_{i \in [n]} \sum_{j \in [n], j \neq i} P_{\gamma_n^*,ij} e_i(\beta_0) e_j(\beta_0) > q_{1-\alpha} \left( \mathcal{N}(0,1) \right), \end{align*} where $P_{\gamma_n^*,ij}$ denotes the $(i,j)$ element of $P_{\gamma_n^*}:= Z(Z^{\top}Z + \gamma_n^* I_K)^{-1}Z^{\top}$, $r_n := rank(Z)$, $\widehat{\Phi}_{\gamma_n^*}(\beta_0) := \frac{2}{r_n} \sum_{i \in [n]}\sum_{j \neq i}(P_{\gamma_n^*,ij})^2e_i^2(\beta_0)e_j^2(\beta_0)$, $\gamma_n^* := \max \operatorname*{arg\,max}_{\theta \in \Gamma_n} \sum_{i \in [n]} \sum_{j \neq i} P_{\theta,ij}^2$, and $ \Gamma_n := \{ \gamma_n \in \mathbb{R}: \gamma_n \geq 0 \text{ if } r_n = K, \text{ and } \gamma_n \geq 1 \text{ if } r_n < K \} $; • BCCH: belloni2012's sup-score test, which rejects $H_0$ whenever \begin{align*} \max_{1 \leq j \leq K} \frac{ \left| \sum_{i \in [n]} e_i(\beta_0) Z_{ij} \right| }{ \sqrt{\sum_{i \in [n]} e_i^2(\beta_0) Z_{ij}^2 }} > c_{_{BCCH}} q_{1-\alpha/(2K)}(\mathcal{N}(0,1)), \end{align*} where we let $c_{_{BCCH}} = 1.1$, following belloni2012's recommendation; • CT: The ridge-regularized AR test proposed by carrasco2016, which rejects $H_0$ whenever \begin{align*} \frac{n e(\beta_0)^{\top}P_{0.05} e(\beta_0)}{e(\beta_0)^{\top}(I_n - P_{0.05})e(\beta_0)} > \widehat{\mathcal C}^*_{\alpha, CT}(\beta_0), \end{align*} where $\widehat{\mathcal C}^*_{\alpha, CT}(\beta_0)$ denotes the bootstrap critical value discussed in Section 3 of their paper, and the choice of the fixed scalar 0.05 for the regularizer (which does not depend on $n$) follows that used in the simulations of DKM24.\footnote{carrasco2016 show that under homoskedastic errors, their test statistic converges to an infinite sum of weighted $\chi^2_1$ distributions. For inference, they proposed a residual bootstrap procedure, which is based on the empirical distribution of residuals.} • LM: The jackknife LM test for diverging $K$ proposed by matsushita2024, which rejects $H_0$ whenever \begin{align*} \frac{1}{\sqrt{\widehat\Psi(\beta_0)} \sqrt{K}} \sum_{i \in [n]} \sum_{j \neq i} P_{ij}X_i e_j(\beta_0) > q_{1-\alpha}\left( \mathcal{N}(0,1) \right), \end{align*} where $\widehat\Psi(\beta_0) := \frac{1}{K} \left( \sum_{i \in [n],j \neq i} P_{ij} X_j e_i^2(\beta_0)+\sum_{i \in [n],j \neq i} P_{ij}^2 X_i X_j e_i(\beta_0)e_j(\beta_0) \right);$ • AS: The dimension-robust $F$ test proposed by Anatolyev-Solvsten(2023), which rejects $H_0$ whenever \begin{align*} F > \widehat{C}_{\alpha, AS}, \end{align*} where $F$ and $\widehat C_{\alpha,AS}$ denote the $F$-test statistic and the critical value described in Sections 2.1 and 2.3, respectively, in Anatolyev-Solvsten(2023);\footnote{The code for their test can be found at \url{https://github.com/mikkelsoelvsten/manyRegressors/blob/master/R/LOFtest.R}. Translating our model to that of Anatolyev-Solvsten(2023), our structural equation of (ref) can be given by $\tilde Y - \tilde X \beta = W \Gamma + \tilde Z \Theta + \tilde e,$ where the AS test corresponds to testing $\Theta = 0_{K_n}$ under the null hypothesis $\beta = \beta_0$. In terms of the notation in Section 2 of their paper, $y_i = \widetilde{Y}_i - \widetilde X_i \beta_0$ and $x_i = (W_i^\top, Z_i^\top)^\top$ with $H^{AS}_0: \textbf{R} \pmb{\beta}^{AS} = \textbf{q},$ where $\pmb{\beta}^{AS} =(\Gamma^\top,\Theta^\top)^\top$, $\textbf{R} = \left[ \begin{array}{cc} \mathbf{0}_{K \times d_w} & I_K \end{array} \right]$, and $\textbf{q} = \mathbf{0}_{K \times 1}$. } • Empirical: The bootstrap test based on our test statistic in (ref) but with its critical value generated by the empirical distribution of $e_i(\beta_0)$ instead, which rejects $H_0$ whenever $$\widehat{Q}(\beta_0) > \widetilde{\mathcal C}^*_{\alpha}(\beta_0),$$ where $\widetilde{\mathcal C}^*_{\alpha}(\beta_0)$ is the $(1-\alpha)$-th percentile of $\widetilde{Q}^*(\beta_0)$ conditional on data, $ \widetilde{Q}^*(\beta_0):= \frac{1}{\sqrt{K_\lambda}} \sum_{i \in [n]} \sum_{j \in [n], j \neq i} e^*_{i}(\beta_0)\Xi_{\lambda,ij} e^*_{j}(\beta_0)$, and $\{e^*_{i}(\beta_0)\}_{i \in [n]}$ is drawn from the empirical distribution of $\{e_i(\beta_0)\}_{i \in [n]}$. We use the same regularizer as that for the BS test in (1). • JK: The jackknife K test proposed by N23, which rejects $H_0$ whenever \begin{align*} JK(\beta_0) > q_{1-\alpha}(\chi^2_{1}), \end{align*} with $JK(\beta_0)$ defined in (2.5) of his paper.

Simulations Based on Haus2012

In this section, we consider a model based on the DGP given by Haus2012, with a sample size $n=200$ and a heteroskedastic error structure.

align*[align* omitted — 1,048 chars of source]

We assume that the errors are independent across $i$. We vary the number of instruments $K \in \{2,10,40,160,300\}$ and $\beta \in [-2,2]$ to investigate the size and power properties of the eleven tests under both fixed and diverging $K$ settings. Specifically, the $i$th instrument observation for $K \geq 10$ is given by $$Z_i^{\top} = (z_{i1},z_{i1}^2,z_{i1} \mathbb{1}(z_{i1}<q_{25}), z_{i1} \mathbb{1}(q_{25} \leq z_{i1} < q_{50}), z_{i1} \mathbb{1}(q_{50} \leq z_{i1} < q_{75}) ,z_{i1}D_{i1},...,z_{i1}D_{i,K-5}),$$ where $q_{\alpha}$ is the $\alpha$-percentile of $\{ z_{i1}\}_{i \in [n]}$, $D_{ik} \in \{0,1\}$ is a dummy variable that is independent across $(i,k)$ with $\mathbb{P}(D_{ik} = 1) = 1/2$, so that $Z_i \in \mathbb{R}^K$. Furthermore, for the case with $K=2$, we let $$Z_i^\top := (z_{i1},z_{i1}^2).$$ We define $\mu^2 := n \pi^{\top}\pi,$ and consider $\mu^2 = 72$ for $K=2$, while $\mu^2 = 8$ for $K \geq 10$, following Haus2012.\footnote{Specifically, for $K=2$, we let $\pi = \frac{0.6}{\sqrt{K}} \iota_K$; for $K \geq 10$ we let $\pi = \frac{0.2}{\sqrt{K}} \iota_K$. We allow $\mu^2$ to be larger for $K=2$ to demonstrate a non-negligible power; otherwise, all the tests would have a trivial power.}

table[table omitted — 1,555 chars of source]

Size Properties: Table (ref) reports the null rejection probabilities of the eleven tests across different $K$. We make several observations below. First, the $\rm RJAR$, $\rm JAR_{std}$, $\rm JAR_{\rm cf}$, $\rm CT$, $\rm Empirical$, $\rm LM$ and $\rm JK$ tests suffer from remarkable over-rejections under some or all values of $K$. Second, the classical $\rm AR$ test for fixed $K$ and the $\rm AS$ test control size for all values of $K$, but become conservative when $K$ is large. Similarly, we observe that the $\rm BCCH$ test is relatively conservative across different numbers of IVs. Indeed, we will see from Figures (ref)-(ref) that these tests tend to suffer from power decline when $K$ becomes large. By contrast, our proposed dimension-robust $\rm BS$ test largely resolves the size-distortion issues for all values of $K$ considered. Overall, our $\rm BS$ test has the best size properties among the eleven tests.

Power Properties: Figures (ref) and (ref) report the power curves for $10,40,160,$ and $300$ IVs, respectively. The power curve for $2$ IVs is reported in Figure (ref) in the Supplemental Appendix. Several remarks are in order. First, $\rm JAR_{\rm std}$ and $\rm RJAR$ have the same power curves for $K \in \{2, 10, 40, 160\}$, because RJAR's chosen regularizer $\gamma_n^*$ equals zero under the current DGP. Additionally, the power of $\rm JAR_{\rm std}$ and $\rm RJAR$ become relatively low as the number of IVs becomes large (e.g., $K=40,160$, and also $K=300$ for $\rm RJAR$). Second, the power curves of $\rm JAR_{\rm cf}$ are similar to those of $\rm JAR_{\rm std}$, but with higher rejections under the null. Third, for the cases with a larger $K$ (e.g., $K=40,160$), the power of the classical $\rm AR$, $\rm CT$, $\rm LM$, $\rm AS$ and $\rm JK$ tests is relatively low (some also suffer from serious size distortions). Fourth, the $\rm JK$ test has relatively low power with $10$ and $40$ IVs, but relatively good power with 160 and 300 IVs. Its size distortions are also small when the number of IV is large. Fifth, for the current DGP, $\rm BCCH$ typically has good power performance for $\beta <0$ but its power can be relatively low for $\beta > 0$. Overall, our BS test has the best power properties, with its power curves much higher than the other test in many cases.

Regularizers: Recall $\overline{\theta} = ||Z^{\top}Z||_{op}$. When $K = 2, 10, 40$ and $160$, we have $\lambda = \overline{\theta} = 200$; $\lambda = 256$, $\overline{\theta} = 1233.927$; $\lambda = \overline{\theta} = 3665.477$; $\lambda = \overline{\theta}= 13790.551$, respectively. When $K = 300$, $\lambda = \overline{\theta}= 125589.052$, while $\gamma_n^* = 41$. For $K \in \{2,10,40,160\}$, we have $\gamma_n^* = 0$ under this DGP.

figure[figure omitted — 967 chars of source]
figure[figure omitted — 978 chars of source]

Empirical Application

In this section, we consider an empirical application of IV regressions with underlying specifications based on card(2009), Goldsmith(2020), and DKM24. Specifically, we consider a single cross-section of data in year $2000$ across $124$ locations (cities) by using the following model:

align*[align* omitted — 82 chars of source]

where $\beta_s$ is the coefficient of interest and can be interpreted as the (negative) inverse elasticity of substitution between immigrants and natives in the relevant skill group $s$. In addition, $Y_{ls}$ denotes the difference between the residual log wages\footnote{As discussed in card(2009), residual log wages are log wages after controlling for education, age, gender, race, and ethnicity of the U.S. workforce.} for immigrant and native men in skill group $s \in \{h,c\}$ (high school or college equivalent) and location (city) $l = 1,...,124$. The vector of location-level controls is denoted as $W_l$; in this application we include the following controls: $(1)$ log of city size, $(2)$ college shares, $(3)$ manufacturing shares in both $(i)$ $1980$ and $(ii)$ $1990$, $(4)$ mean wage residuals for $(i)$ all natives and $(ii)$ all immigrants in $1980$, together with $(5)$ a constant (so that there are $9$ controls available for each city, i.e., $W_l \in \mathbb{R}^9$).\footnote{See Table 6 in card(2009) for more details on the controls.} $X_{ls}$ denotes the log ratio of immigrant to native hours worked in skill group $s$ of both men and women in the city $l$.

The ratio $X_{ls}$ is potentially endogenous because unobserved city-specific factors may shift the relative demand for immigrant workers, leading to higher relative wages and higher relative employment shares, thereby confounding the estimation of the inverse elasticity of substitution. To overcome this issue, card(2009) suggests using the ratio of the total number of immigrants from foreign country $m$ in city $l$ to the total number of immigrants from country $m$ as an instrument. The rationale for such a choice is that existing immigrant enclaves are likely to attract additional immigrant labor through social and cultural channels unrelated to labor market outcomes. To define the instruments, we can exploit settlement patterns at some initial period (possibly together with the arrival rate of immigrants from specific countries in subsequent periods) to determine the inflow of immigrants in each location. Specifically, we let $N_{lm,1980}$ be the number of immigrants from country $m = 1...,38$ settling into location $l$ in $1980$ and let $N_{l,1980}$ be the total number of immigrants in location $l$ in $1980$, respectively. In addition, we let $P_{l,2000}$ denote the population size of location $l$ in $2000$, including both immigrants and natives.

To proceed, we consider four sets of potential instruments for $X_{ls}$. The definition of the first two sets of instruments follows from DKM24. Specifically, we let the instruments for each location $l$ be given by $z_{l,1980} := \left\{ z_{lm,1980} \right\}_{m=1}^{38} \equiv \left\{ \frac{N_{lm,1980}}{N_{l,1980}} \times \frac{1}{P_{l,2000}} \right\}_{m=1}^{38} \in \mathbb{R}^{38 \times 1}$, so that our first set of instruments can be written as $Z_{38} := (z_{1,1980},...,z_{124,1980} )^{\top} \in \mathbb{R}^{124 \times 38}$. For the second set of instruments, we let each of the 38 IVs interact with the $9$ controls described above, so that $z_{l,1980} \in \mathbb{R}^{342 \times 1}$ (i.e., each $z_{lm,1980}$ is interacted with $9$ controls). Then, our second set of instruments is defined as $Z_{342} \in \mathbb{R}^{124 \times 342}$. Furthermore, given that our proposed bootstrap test is dimensional-robust, we consider a case with $K=1$ for our third and fourth sets of instruments for the high school and college skill groups, respectively. Specifically, following Goldsmith(2020), we construct Bartik-type instruments, which are given by $Z_{Bartik,s} = \{B_{ls}\}_{l=1}^{124} \in \mathbb{R}^{124 \times 1}$ for $s \in \{h,c\}$, where $B_{ls}:= \sum_{m=1}^{38} z_{l,1980} \times g_{ms}$ and $g_{ms}$ is the number of immigrants from country $m$ in skill group $s$ arriving in US from $1990$ to $2000$. Note that while the Bartik IVs depend on the skill group $s$ (i.e., $Z_{Bartik,h}$ and $Z_{Bartik,c}$), the first and second sets of instruments (i.e., $Z_{38}$ and $Z_{342}$) do not depend on $s$.

The empirical results of our bootstrap test and those in DKM24 are given in Tables (ref)--(ref). For the Bartik instruments (i.e., $K=1$), the result for $\rm JAR_{\rm cf}$ is not reported because the cross-fit variance estimator is negative. Table (ref) shows the $95\%$ confidence intervals (CIs) with the Bartik IVs. $Z_{Bartik,h}$ and $Z_{Bartik,c}$ are applied separately to their respective skill groups. The regularizers for methods $\rm RJAR$ and $\rm BS$ are $\gamma_n^* = 0$ and $\lambda = 0$, respectively, with ${p_n'} = 0.077$ and $p_n = 0.216$.\footnote{When $K=1$, ${p_n'}(\lambda)$ and $p_n(\lambda)$ are independent of $\lambda$, and we set $\lambda = 0$.} $K_\lambda$ for BS is equal to $0.457$ and $0.253$ for high-school and college workers, respectively. In addition, the number beneath each CI represents its relative length compared to the $\rm BS$ CI. For $K=1$, all CIs have similar lengths. Methods $\rm RJAR$ and $\rm JAR_{std}$ have shorter CIs, but this is because these methods may not control size when $K$ is fixed and tend to over-rejects under the null, as observed in our simulation studies. Among the CIs that are theoretically valid for small $K$ (i.e., $\rm AR$, $\rm BS$, and $\rm BCCH$), the $\rm BS$ CIs are the shortest across both skill groups.

Table (ref) reports the $95\%$ CIs for high-school and college workers, respectively, with $K = 38$; the set of instruments used for both skill groups is $Z_{38}$. The regularizers for methods $\rm RJAR$ and $\rm BS$ are $\gamma_n^* = 0$ and $\lambda = 13.8$, respectively, with $p_n = 0.022$ and ${p_n'} = 0.089$. $K_\lambda$ for BS is equal to $2.015$ for 38 IVs. We find that $\rm BS$ has the shortest CI for college workers, while $\rm JAR_{\rm cf}$ has the shortest CI for high-school workers. But based on our simulation studies, $\rm JAR_{\rm cf}$ may over-reject under the null, which can result in shorter CIs. Furthermore, the $\rm BS$ CIs are shorter than their $\rm BCCH$ counterparts for both high-school and college workers.

Table (ref) shows the $95\%$ CIs with $K = 342$; the set of instruments used for both skill groups is $Z_{342}$. The regularizers for methods $\rm RJAR$ and $\rm BS$ are $\gamma_n^* = 5.3$ and $\lambda = 67.4$, respectively, with $p_n = 0.016$ and ${p_n'} = 0.089$. $K_\lambda$ for BS is equal to $1.550$ for 342 IVs. For both high-school and college workers, $\rm CT$ rejects all null hypotheses and thus results in empty confidence intervals, potentially due to heteroskedastic errors. $\rm BS$ again has the shortest confidence interval for college workers, and is of similar length with $\rm RJAR$ for high-school workers. Finally, $\rm BCCH$ has relatively wide CIs compared with $\rm BS$ and $\rm RJAR$.

table[table omitted — 1,595 chars of source]
table[table omitted — 1,753 chars of source]
table[table omitted — 1,285 chars of source]

\singlespacing