Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
97,031 characters · 15 sections · 183 citation commands
A Dimension-Agnostic Bootstrap Anderson-Rubin Test For Instrumental Variable Regressions
Keywords: Anderson-Rubin Test, Weak Identification, Bootstrap, High Dimension, Quadratic Form, Ridge Regularization
JEL Classification: C12, C36, C55
\setcounter{oldtocdepth}{\value{tocdepth}} \addtocontents{toc}{\setcounter{tocdepth}{-10}}
Weak and numerous instruments remain persistent concerns in instrumental variable (IV) regressions across various fields. Surveys by Andrews-Stock-Sun(2019) and lee2021 (lee2021) find that a considerable number of IV regressions in the American Economic Review report first-stage F-statistics below 10. In addition, empirical studies often involve many instruments, such as the 180 IVs used by Angrist-Krueger(1991) to examine the effect of schooling on wages. In “judge design” studies, the number of instruments (number of judges) is typically proportional to the sample size (MS22, MS22).\footnote{E.g., see kling2006, doyle2007, dahl2014, dobbie2018, sampat2019, agan2023, frandsen2023, chyn2024 and the references therein.} Similar patterns of many IVs occur in Fama-MacBeth regressions (fama1973, fama1973; shanken1992, shanken1992), shift-share IVs (Goldsmith(2020), Goldsmith(2020)), wind-direction IVs (deryugina2019mortality, deryugina2019mortality; bondy2020crime, bondy2020crime), granular IVs (gabaix2024, gabaix2024), local average treatment effect estimation (blandhol2022, blandhol2022; boot2024, boot2024; sloczynski2024should, sloczynski2024should), and Mendelian randomization (davey2003, davey2003; davies2015, davies2015).
However, existing weak-identification-robust inference methods for IV regressions are either based on an asymptotic framework in which the number of instruments $K$ is treated as fixed\footnote{E.g., see Staiger-Stock(1997), Stock-Wright(2000), Kleibergen(2002), Kleibergen(2005), Moreira(2003), Andrews-Cheng(2012), Andrews-Mikusheva(2016), Andrews(2018), Andrews-Guggenberger(2019), Moreira-Moreira(2019), among others.} or on an alternative one that allows $K_n$ to diverge to infinity with the sample size $n$.\footnote{E.g., see Andrews-Stock(2007), Newey-Windmeijer(2009), anatolyev2011, crudu2021, MS22, matsushita2024, LWZ(2023), DKM24, among others.} These methods compare distinct test statistics with distinct critical values so that procedures formulated under the fixed-$K$ asymptotics generally do not have correct size control under the diverging-$K$ asymptotics and vice versa. An empirical researcher is, therefore, forced to take a stance on the asymptotic regime of the number of instruments to implement them, which can be ambiguous in many empirical applications. For example, when $K$ is moderate compared with $n$ (e.g., $K=10$ and $n=200$), it is unclear which test the researcher should use. Furthermore, as we will see below, a third asymptotic regime may arise with the use of regularization, making dimension-robust inference even more challenging.
Motivated by this issue, we propose a bootstrap-based, dimension-agnostic AR test. First, by deriving strong approximations for the proposed test statistic and its bootstrap counterpart (under both the null and alternative hypotheses), we show that the new bootstrap test has a correct asymptotic size, regardless of whether the number of IVs $K$ is fixed or diverging. Our proof, which relies on the Lindeberg swapping strategy, contributes a general result on the strong approximation for quadratic forms with independent and heteroskedastic errors. Additionally, our (conditional) strong approximation derivation for bootstrap statistics involving quadratic forms is novel and may be of independent interest. Second, by employing a ridge-regularized projection matrix, our AR test remains valid in high-dimensional cases where $K$ exceeds the sample size $n$. Third, the characterization of the errors in strong approximation offers a theoretically sound basis for selecting the ridge regularizer without taking a stance on the specific regime of $K$. Our choice of the regularizer also helps to reduce the rank of the projection matrix, which can potentially improve the power performance of the test. Fourth, we show that depending on the asymptotic behavior of both $K$ and $K_\lambda$ (the effective rank of the regularized projection matrix), the limit distribution of the test statistic can be (1) normal, (2) weighted chi-squared, or (3) a mixture of weighted chi-squared and normal distributions. Given the strong approximation result, our bootstrap inference remains uniformly valid regardless of the asymptotic regimes, and we further provide its power properties under each scenario. Fifth, the strong approximation and uniform inference results are all established when the number of control variables is allowed to diverge at the same rate or even faster than $\sqrt{n}$. Sixth, simulation experiments and an empirical application to the dataset of card(2009) confirm the excellent size and power properties of our bootstrap test compared to alternative methods.
Relation to the literature: For weak-identification-robust inference based on the classical AR test, Andrews-Stock(2007) showed its validity under many instruments, but requires the number of instruments to diverge more slowly than the cube root of the sample size $n$ ($K^3/n \rightarrow 0$). Newey-Windmeijer(2009) proposed a GMM-AR test under many (weak) moment conditions but imposed the same rate condition on $K$. anatolyev2011 constructed a modified AR test that allows $K$ to be proportional to $n$ but requires homoskedastic errors, and Kaffo-Wang(2017) proposed a bootstrap version of their test. For estimation with many instruments, carrasco2012, carrasco2015, carrasco2016efficient, hansen2014, and carrasco2017 proposed regularization approaches for two-stage least squares, limited information maximum likelihood, and jackknife IV (Angrist(1999), Angrist(1999)) estimators. Furthermore, carrasco2016 first proposed a ridge-regularized AR test that allows for $K$ being larger than $n$ with homoskedastic errors. Bun-Farbmacher-Poldermans(2020) compared the centered and uncentered GMM-AR test and identified a missing degrees-of-freedom correction when $K/n \rightarrow 0$. Recently, crudu2021 and MS22 proposed jackknifed versions of the AR test under many instruments and general heteroskedasticity. DKM24 developed a ridge-regularized version of the jackknife AR test, which is further robust to the scenario where $K$ diverges faster than the sample size. However, the jackknife AR tests are based on standard normal critical values that require $K$ to diverge; thus, they may not have the correct size under fixed $K$. tuvaandorj2024 established the validity of a permutation AR test under heteroskedasticity and diverging $K$, requiring $K^3/n \rightarrow 0$. In contrast to the above methods, our test remains valid with heteroskedastic errors uniformly across a broad asymptotic regime for $K$, spanning from fixed to diverging faster than the sample size.
Furthermore, belloni2012 proposed a Lasso-based method for selecting optimal instruments, valid under high-dimensional IVs and heteroskedasticity, but requiring strong identification and sparse first-stage regressions. However, WZ23 showed that both Lasso and debiased Lasso linear regressions can suffer from significant omitted variable bias, even when the coefficient vector is sparse and the sample size exceeds the number of controls. In such cases, the “long regression,” which includes all regressors, often outperforms the Lasso-based methods. kolesar-muller-roelsgaard(2025) similarly recommended using the “long regression" unless the number of regressors is comparable to or exceeds the sample size. belloni2012 also proposed a weak-identification-robust sup-score test that is dimension-agnostic and does not rely on sparsity. Similar to DKM24, our simulation study shows that the power of our ridge-regularized bootstrap AR test matches the sup-score test when IVs have strong but sparse signals while offering substantially more power when the signal is weak but dense. N23 introduced a jackknife version of the Kleibergen(2002)'s K test and combined it with the sup-score test, but his method relies on a sparse $\ell_1$-regularized estimation of $\rho(Z_i)$, the conditional correlation between the endogenous variable and the outcome error. Without the sparsity assumption, the estimation of $\rho(Z_i)$ may be inconsistent when the dimension of $Z_i$ is large. boot-ligtenberg(2023) developed a dimension-robust AR test based on continuous updating, but relied on an invariance assumption. In contrast to the aforementioned approaches, our bootstrap inference procedure accommodates many instruments and heteroskedastic errors, yet does not rely on invariance or sparsity assumptions.
Our paper also relates to the literature on bootstrap inference for IV regressions. It is found in this literature that when implemented appropriately, bootstrap approaches may substantially improve the inference accuracy for IV models, including the cases where IVs may be rather weak.\footnote{E.g., see Davidson-Mackinnon(2008), Davidson-Mackinnon(2010), Davidson-Mackinnon(2014b), Moreira-Porter-Suarez(2009), Wang-Kaffo(2016), Finlay-Magnusson(2019), roodman2019, young2022, and Wang-Zhang2024, among others.} However, no existing study has uniformly established the bootstrap validity with regard to the number of IVs. We fill this gap by deriving strong approximation results for both the test statistic and its bootstrap counterpart. The strong approximation for the AR statistic is related to the analysis of quadratic forms by HS01. Additionally, our results of (conditional) strong approximation for bootstrap statistics with a quadratic form are, based on our best knowledge, new to the literature.
Our test also remains valid even when the number of control variables diverges at a rate of $\sqrt{n}$ or faster, provided it remains of a smaller order than $n$, regardless of whether $K$ is fixed or diverging. As pointed out by Chao(2023) and mikusheva2024weak, the presence of many controls can introduce additional bias in jackknife IV estimators and AR tests. This phenomenon, often referred to as the quadratic barrier (see CJM18; LSMPW24), poses a major challenge for inference. To address this, we design a debiasing procedure for the AR statistic following the construction in CJN18. Furthermore, to achieve valid bootstrap inference under many controls, we explicitly account for the impact of debiasing on the dispersion of the AR statistic by appropriately adjusting the bootstrap statistic.
Lastly, Anatolyev-Solvsten(2023) proposed an analytical dimension-agnostic $F$ test for linear regressions by analyzing the asymptotic behavior of quadratic forms under two distinct regimes: (1) a fixed number of restrictions, resulting in a weighted chi-squared limiting distribution, and (2) a growing number of restrictions, yielding a normal limiting distribution. Their $F$-test is, in principle, applicable to our setting by testing zero restrictions on the IV coefficients in a linear regression under the null, and it is more general in two respects: (1) it accommodates control variables whose dimension can be of the same order as the sample size $n$, and (2) it allows for testing general linear restrictions. However, our bootstrap inference offers several key advantages. First, although we require the number of controls to be of a smaller order than $n$, we allow the number of instruments $K$ to exceed $n$, a case not covered by their framework. Second, our use of ridge regularization reduces the rank of the projection matrix and gives rise to a third asymptotic regime, where $K$ diverges but a Lindeberg-type condition for asymptotic normality fails, resulting in a limiting distribution that is a mixture of weighted chi-squared and normal variables, akin to the regime analyzed in KSS2020 and YGZ24. This regime does not arise in Anatolyev-Solvsten(2023) due to the absence of regularization. Analytical inference in this setting requires knowledge of the number of dominant eigenvalues, which can be a challenging task. In contrast, our bootstrap approach circumvents such difficulty by directly employing the strong approximation and remains uniformly valid across all three regimes. In our simulations, when the number of instruments is proportional to the sample size, the use of ridge regularization places the test statistic in the third asymptotic regime. Our bootstrap inference procedure has excellent size control even in this challenging setting, and further provides substantial power gains compared to alternative methods because of the rank reduction.
Structure of the paper: Section (ref) makes precise the model setup and provides the testing procedure for our dimension-robust AR test statistic. Sections (ref) and (ref) provide the strong approximation results under both null and alternative for our test statistic and its bootstrap counterpart, respectively. We derive the power properties of our test under the fixed-$K$ and diverging-$K$ asymptotics, respectively, in Section (ref). Section (ref) presents the results of Monte Carlo simulations and Section (ref) applies our test to an empirical application. Proofs of the theorems are given in the Supplemental Appendix, along with additional lemmas and simulation results.
Notations: We denote by $[n]$ the set $\{1,\cdots,n\}$, and use $||A||_{op}$ and $||A||_F$ to refer to the operator and Frobenius norms of a matrix $A$, respectively.
Consider the linear instrumental variable regression
where $\widetilde X_i$ denotes a scalar endogenous variable and $W_i \in \mathbb{R}^{d_w}$ denotes the exogenous control variables. In addition, we have $K$-dimensional instrumental variables (IVs) denoted as $\widetilde Z_i$, and $ \widetilde{\Pi}_i \equiv \mathbb{E}( \widetilde{X}_i | \widetilde{Z}_i, W_i)$. We stack $\widetilde Z_i^\top$ up and denote the resulting $n \times K_n$ matrix $\widetilde Z$. We define $\widetilde Y \in \mathbb R^n$, $\widetilde X \in \mathbb R^n$, $\widetilde \Pi \in \mathbb R^n$, $\widetilde v \in \mathbb R^n$, and $ W \in \mathbb R^{n \times d_w}$ in the same manner. Throughout the paper, we also allow $d_w$ to diverge to infinity but at rate that is slower than the sample size $n$, i.e., $d_w = o(n)$. We further require $W$ to be of full rank so that its projection matrix $P_W = W(W^\top W)^{-1} W^\top$ is well defined. We allow, but do not require, $K_n$ to increase with $n$. Specifically, the dimension of $Z$ can be fixed, grow proportional to, or even faster than $n$.
We focus on the model with a scalar endogenous variable for two reasons. First, in many empirical applications of IV regressions, there is only one endogenous variable (as can be seen from the surveys by Andrews-Stock-Sun(2019) and lee2021). Second, the strong approximation results derived in Sections (ref) and (ref) extend directly to the general case of full-vector inference with multiple endogenous variables. Additionally, for the dimension-robust subvector inference, one may use a projection approach (Dufour-Taamouti(2005), Dufour-Taamouti(2005)) after implementing our test on the whole vector of endogenous variables.\footnote{Alternative subvector inference methods for IV regressions (e.g., see GKMC(2012), Andrews(2017), and GKM(2019), GKM(2021)) provide a power improvement over the projection approach under fixed $K$. However, whether they can be applied to the current setting is unclear. Also, Wang-Doko(2018) and Wang(2020) show that bootstrap tests based on the standard subvector AR statistic may not be robust to weak identification even under fixed $K$ and conditional homoskedasticity.}
To proceed, we first partial out the exogenous control variables $W$ from our IV regressions. Specifically, we stack up $(Y_i,X_i,e_i,\Pi_i,v_i)$ to $(Y,X,e,\Pi,v)$, which are defined as $Y = M_W \widetilde{Y}$, $X = M_W \widetilde{X}$, $\Pi = M_W \widetilde{\Pi}$, $e = M_W \widetilde{e}$, and $v = M_W \widetilde{v}$, where $M_W = I_n - P_W$ and $I_n$ is an $n \times n$ identity matrix. In addition, we define $Z = M_W \widetilde{Z}$. Then, (ref) can be rewritten as
Throughout our analysis, we treat $(Z,W)$ as fixed, which is equivalent to taking all expectations and probability measures conditionally on $(Z,W)$.
Given that we allow $K_n$ to be greater than $n$, the matrix $Z^\top Z$ is not necessarily invertible. Therefore, we define $P_\lambda = Z (Z^\top Z + \lambda I_{K_n})^{-1}Z^\top$ as the ridge-regularized projection matrix of $Z$ with some ridge penalty $\lambda$ that will be chosen based on $Z$ only. As we treat the instruments and control variables as fixed, so are the ridge-regularizer $\lambda$ and matrix $P_\lambda$. The $(i,j)$ element of $P_{\lambda}$ is denoted as $P_{\lambda,ij}$. Further denote $e_i(\beta_0) = Y_i - X_i \beta_0$. Then, our dimension-agnostic AR test statistic is written as
where $\kappa = (M_W \circ M_W)^{-1}$,\footnote{Here $\circ$ denotes the Hadamard product and $M_W \circ M_W$ is invertible as long as $d_w < n/2$ as shown by CJN18.} $A_{\lambda,ii} = 2 P_{\lambda,ii} P_{W,ii} - B_{\lambda,ii}$, $B_{\lambda,jk} = \sum_{i \in [n]} P_{W,ik} P_{W,ij} P_{\lambda,ii} = [P_W D_{\lambda} P_W]_{jk}$,
and
In particular, we can regard $K_\lambda$ as the effective rank under the ridge regularization.
We note that under the null (i.e., $\beta = \beta_0$), the first quadratic term of $\widehat Q(\beta_0)$ in (ref) does not have an exact zero mean due to partialling out controls. This bias is not asymptotically negligible when the dimension of the controls ($d_w$) is of the order $\sqrt{n}$ or greater. The second term of $\widehat Q(\beta_0)$ in (ref), inspired by the variance estimator proposed by CJN18, is used to correct such a bias.
The regularizer $\lambda$ is chosen as
where $P_{\theta} = Z (Z^\top Z + \theta I_{K_n})^{-1}Z^\top$, $P_{\theta, ij}$ is the $(i,j)$ entry of $P_{\theta}$, $\overline \theta = ||Z^{\top}Z||_{op}$, $K_{\theta}$ is defined in (ref) with $\lambda$ replaced by $\theta$, while $c_1$ and $c_2$ are two positive constants chosen by the researcher such that $c_1$ is sufficiently small. We view $\frac{c}{0} = + \infty$ for any $c>0$. In practice, we use $c_1 = 0.1$ and $c_2 = 1$.\footnote{We have done extensive simulations and find that the results of our test are not sensitive to the specific choice of $c_1$ and $c_2$. The simulation results with alternative choices of $c_1$ and $c_2$ are reported in the Supplemental Appendix.} If there is no $\lambda$ that satisfies both inequalities in (ref), then we choose
To implement the dimension-agnostic test, we propose to use bootstrap critical values. Specifically, let $\{\eta_i\}_{i \in [n]}$ be an independent sequence of random variables with zero mean and unit variance that are generated independently from the samples. Our bootstrap AR test statistic is denoted as $\widehat{Q}^*(\beta_0)$ and defined as
Then, the bootstrap critical value is denoted as $\widehat{\mathcal C}^*_{\alpha}(\beta_0)$ and defined as the $(1-\alpha)$-th percentile of $\widehat{Q}^*(\beta_0)$ conditional on data, where $\alpha$ is the nominal level of rejection under the null. We reject the null hypothesis of $\beta = \beta_0$ if $\widehat{Q}(\beta_0) > \widehat{\mathcal C}^*_{\alpha}(\beta_0).$
This section is concerned with the conditions under which the null distribution of the test statistic defined in (ref) can be approximated by its bootstrap counterpart, no matter whether the dimension $K_n$ of the IVs is fixed or diverging with the sample size. We make the following assumptions on the data-generating process (DGP) to establish this result.
Assumptions (ref).1--(ref).3 are standard regularity conditions. Assumption (ref).4 allows the dimension of control variables to diverge at a rate that is slower than the sample size, i.e., $d_w = o(n)$. The impact of partialling out $W$ from both $Y$ and $X$ becomes asymptotically negligible only when $d_w = o(\sqrt{n})$, reflecting a broader phenomenon commonly referred to as the quadratic barrier. See, for example, CJM18 and LSMPW24 for further discussions. We overcome this barrier and establish bootstrap validity by carefully debiasing the AR statistic and further adjusting the middle matrix of the bootstrap quadratic form (as noted in Remark (ref)). To the best of our knowledge, this is the mildest rate condition regarding the number of controls established for bootstrap inference with high-dimensional IVs (without imposing a sparsity assumption). We note that analytical inference remains feasible even when $d_w$ is proportional to $n$, as demonstrated in Anatolyev-Solvsten(2023). However, in such a high-dimensional control setting, our current bootstrap inference procedure may fail to control size. At present, it is unclear whether any valid resampling-based inference method exists in this regime, let alone one that remains valid uniformly over the dimensions of both $Z$ and $W$. We leave this important question for future research. In the following, we provide further comparisons between analytical and bootstrap inference approaches in Remarks (ref) and (ref).
Assumption (ref).5 requires that $p_n$ and ${p_n'}$ vanish sufficiently fast. Consider the case without ridge regularization (i.e., $\lambda=0$) and where the projection matrix is well-defined (i.e., $K_n < n$). If the diagonal elements of $P_\lambda$ (with $\lambda = 0$) are well-balanced in the sense that $P_{\lambda,ii} = K_n/n$, then we have
This implies $p_n = O(n^{-1})$ and ${p_n'} = O(n^{-1})$. Importantly, we note that these results hold regardless of whether $K_n$ is fixed or increasing with $n$. If $P_W$ is also well-balanced such that $\max_{i \in [n]} P_{W,ii} \leq Cd_w/n$, then
as long as $d_w = o(n)$. In the minimum, even we only have ${p_n'} =o(1)$, if $d_w = O(\sqrt{n})$, then $\sum_{i \in [n]}P_{W,ii}^2 = O(1)$, which still guarantees that ${p_n'} (1 + \sum_{i \in [n]}P_{W,ii}^2)=o(1)$.
These calculations imply that our inference procedure remains valid even when the number of control variables diverges at the rate $\sqrt{n}$ or faster. Moreover, in high-dimensional settings where $K_n > n$, our ridge-regularized approach with the choice of $\lambda$ in (ref) ensures that Assumption (ref).5 holds provided $q>6$.
Finally, Assumption (ref).6 requires the bootstrap weights $\eta_i$ to have sub-Gaussian tails. In practice, we recommend using standard normal or Rademacher random variables, both of which satisfy this condition.
To proceed, we need to introduce some more notation. Define $\Delta = \beta - \beta_0$, $\tilde \tau_i = \mathbb E(\tilde e_i \tilde v_i)$, $\tilde \varsigma_i^2 = \mathbb E \tilde v_i^2$, $\breve e_i(\beta_0) = \tilde e_i(\beta_0) + \Pi_i \Delta$, and $\tilde e_i(\beta_0) = \tilde e_i + \tilde v_i \Delta $. Then, we denote
In addition, let
and
where $\{g_{i}\}_{i \in [n]}$ are i.i.d. standard normal random variables that are generated independent of data. The following theorem shows that our proposed AR test statistic $\widehat Q(\beta_0)$ can be strongly approximated by $Q(\beta_0) + C(\Delta)$ in Kolmogorov distance, where
Furthermore, Theorem (ref) in the next section shows the bootstrap statistic $\widehat Q^*(\beta_0)$ can be strongly approximated by $Q^*(\beta_0)$ in Kolmogorov distance conditionally on data. Note that $Q^*(\beta_0)$ is equal to $Q(\beta_0)$ under the null hypothesis.\footnote{Under the null, we have $\Delta=0$ and $\tilde \sigma_i(\beta_0) = \breve \sigma_i(\beta_0)$.}
This section concerns the strong approximation of the bootstrap statistic defined in Section (ref) in Kolmogorov distance conditionally on data. The approximation in Theorem (ref) is the same as that for the original statistic under the null hypothesis, as established in Theorem (ref), which directly implies that the proposed test with bootstrap critical values achieves a correct asymptotic size. Such a result holds no matter whether the dimension of IVs $K_n$ is fixed or diverging to infinity.
In this section, we discuss the power of the bootstrap inference by focusing on three separate cases: (I) both $K$ and $K_\lambda$ diverge, (II) $K$ diverges but $K_\lambda $ is bounded, and (III) both $K$ and $K_\lambda$ are bounded.
To proceed, we let $\Psi(\beta_0) = 2 \sum_{i \in [n]} \sum_{j \in [n], j \neq i} \tilde \sigma^2_i(\beta_0) \Xi_{\lambda,ij}^2 \tilde \sigma_j^2(\beta_0)/K_\lambda$, where $\tilde \sigma_i^2(\beta_0) = Var(\tilde e_i (\beta_0) ) = \tilde \sigma_i^2 + 2 \Delta \tilde \tau_i + \Delta^2 \tilde \varsigma_i^2$.
Following Remark (ref), we now consider the case where $K_\lambda$ remains bounded, resulting in the failure of the Lindeberg-type condition for asymptotic normality.
In this section, we consider the power property of our bootstrap AR test in the asymptotic framework that the dimension of $Z$ (i.e., $K_n = K$) is fixed. To rigorously state the regularity conditions, we recall the singular value decomposition of $Z$ as $Z = \mathcal U \mathcal S \mathcal V^\top,$ where $\mathcal U \in \Re^{n \times n}$, $\mathcal U^\top \mathcal U = I_n$, $\mathcal S = [ S_0, 0_{K, n-K}]^\top$, $S_0$ is a diagonal matrix of non-zero singular values, $0_{K, n-K} \in \Re^{K \times (n-K)}$ is a matrix of zeros, $\mathcal V \in \Re^{K \times K}$, and $\mathcal V^\top \mathcal V = I_K$. Denote $\mathcal U = [\mathcal U_1,\mathcal U_2]$ such that $\mathcal U_1 \in \Re^{n \times K}$, $\mathcal U_2 \in \Re^{n \times (n-K)}$, $\mathcal U_1^\top \mathcal U_1 = I_K$, $\mathcal U_1^\top \mathcal U_2 = 0_{K,n-K}$, and $\mathcal U_2^\top \mathcal U_2 = I_{n-K}$. Further denote $\Omega(\beta_0) \equiv \mathcal U_1^\top \text{diag}(\tilde \sigma_1^2(\beta_0),\cdots,\tilde \sigma_n^2(\beta_0) ) \mathcal U_1$ and the eigenvalue decomposition
Last, denote $\nu(\beta_0) = \lim_{n \rightarrow \infty} \mathbb U^\top \Omega^{-1/2}(\beta_0) \Delta \mathcal U_1^\top \Pi$.
The following theorem establishes our AR test's power property in the fixed $K_\lambda$ scenario.
Next, we demonstrate that when $K$ is fixed, our dimension-agnostic AR test is (asymptotically) admissible within a specific class of tests, which includes the standard (heteroskedasticity-robust) AR test designed for fixed $K$. Let
and $\widehat {\mathcal G}_k (\beta_0)$ be the $k$-th element of $\widehat {\mathcal G} (\beta_0)$, where $\hat \Omega(\beta_0)$ is a consistent estimator of $\Omega(\beta_0)$. We observe that, in the scenario where $K$ is fixed, the standard AR test rejects if
where $\iota_K$ is a $K$-dimensional vector of ones and $\mathcal C_{\iota_K}(1-\alpha)$ is just the $(1-\alpha)$ quantile of the centered chi-squared random variable with $K$ degrees of freedom. On the other hand, our bootstrap AR test is asymptotically equivalent to a test that rejects if
In addition, the proof of Theorem (ref) shows
We consider the class $\Phi_\alpha$ of tests $\phi(\cdot)$ defined as
Both the standard and bootstrap AR tests control size, and thus, belong to this class. The power of any test $\phi(\cdot) \in \Phi_\alpha$ is determined by $\nu(\beta_0) \in \Re^K$.
This section investigates the finite sample size and power performance of existing tests and our proposed test. To begin, we explicitly define these tests and their corresponding critical values. In addition, following belloni2012 and DKM24, upon obtaining a given instrument set $Z$, we standardize it by $\frac{1}{n} \sum_{i=1}^n Z_{ij}^2 = 1$, for $j=1,...,K.$ Note that the tests described in section (ref) below are based on the standardized $Z$. Throughout the simulations, we set the number of Monte Carlo and bootstrap replications equal to $5,000$ and $10,000$ respectively, and set the nominal level $\alpha=0.05$.
Specifically, we consider the following eleven tests:
In this section, we consider a model based on the DGP given by Haus2012, with a sample size $n=200$ and a heteroskedastic error structure.
We assume that the errors are independent across $i$. We vary the number of instruments $K \in \{2,10,40,160,300\}$ and $\beta \in [-2,2]$ to investigate the size and power properties of the eleven tests under both fixed and diverging $K$ settings. Specifically, the $i$th instrument observation for $K \geq 10$ is given by $$Z_i^{\top} = (z_{i1},z_{i1}^2,z_{i1} \mathbb{1}(z_{i1}<q_{25}), z_{i1} \mathbb{1}(q_{25} \leq z_{i1} < q_{50}), z_{i1} \mathbb{1}(q_{50} \leq z_{i1} < q_{75}) ,z_{i1}D_{i1},...,z_{i1}D_{i,K-5}),$$ where $q_{\alpha}$ is the $\alpha$-percentile of $\{ z_{i1}\}_{i \in [n]}$, $D_{ik} \in \{0,1\}$ is a dummy variable that is independent across $(i,k)$ with $\mathbb{P}(D_{ik} = 1) = 1/2$, so that $Z_i \in \mathbb{R}^K$. Furthermore, for the case with $K=2$, we let $$Z_i^\top := (z_{i1},z_{i1}^2).$$ We define $\mu^2 := n \pi^{\top}\pi,$ and consider $\mu^2 = 72$ for $K=2$, while $\mu^2 = 8$ for $K \geq 10$, following Haus2012.\footnote{Specifically, for $K=2$, we let $\pi = \frac{0.6}{\sqrt{K}} \iota_K$; for $K \geq 10$ we let $\pi = \frac{0.2}{\sqrt{K}} \iota_K$. We allow $\mu^2$ to be larger for $K=2$ to demonstrate a non-negligible power; otherwise, all the tests would have a trivial power.}
Size Properties: Table (ref) reports the null rejection probabilities of the eleven tests across different $K$. We make several observations below. First, the $\rm RJAR$, $\rm JAR_{std}$, $\rm JAR_{\rm cf}$, $\rm CT$, $\rm Empirical$, $\rm LM$ and $\rm JK$ tests suffer from remarkable over-rejections under some or all values of $K$. Second, the classical $\rm AR$ test for fixed $K$ and the $\rm AS$ test control size for all values of $K$, but become conservative when $K$ is large. Similarly, we observe that the $\rm BCCH$ test is relatively conservative across different numbers of IVs. Indeed, we will see from Figures (ref)-(ref) that these tests tend to suffer from power decline when $K$ becomes large. By contrast, our proposed dimension-robust $\rm BS$ test largely resolves the size-distortion issues for all values of $K$ considered. Overall, our $\rm BS$ test has the best size properties among the eleven tests.
Power Properties: Figures (ref) and (ref) report the power curves for $10,40,160,$ and $300$ IVs, respectively. The power curve for $2$ IVs is reported in Figure (ref) in the Supplemental Appendix. Several remarks are in order. First, $\rm JAR_{\rm std}$ and $\rm RJAR$ have the same power curves for $K \in \{2, 10, 40, 160\}$, because RJAR's chosen regularizer $\gamma_n^*$ equals zero under the current DGP. Additionally, the power of $\rm JAR_{\rm std}$ and $\rm RJAR$ become relatively low as the number of IVs becomes large (e.g., $K=40,160$, and also $K=300$ for $\rm RJAR$). Second, the power curves of $\rm JAR_{\rm cf}$ are similar to those of $\rm JAR_{\rm std}$, but with higher rejections under the null. Third, for the cases with a larger $K$ (e.g., $K=40,160$), the power of the classical $\rm AR$, $\rm CT$, $\rm LM$, $\rm AS$ and $\rm JK$ tests is relatively low (some also suffer from serious size distortions). Fourth, the $\rm JK$ test has relatively low power with $10$ and $40$ IVs, but relatively good power with 160 and 300 IVs. Its size distortions are also small when the number of IV is large. Fifth, for the current DGP, $\rm BCCH$ typically has good power performance for $\beta <0$ but its power can be relatively low for $\beta > 0$. Overall, our BS test has the best power properties, with its power curves much higher than the other test in many cases.
Regularizers: Recall $\overline{\theta} = ||Z^{\top}Z||_{op}$. When $K = 2, 10, 40$ and $160$, we have $\lambda = \overline{\theta} = 200$; $\lambda = 256$, $\overline{\theta} = 1233.927$; $\lambda = \overline{\theta} = 3665.477$; $\lambda = \overline{\theta}= 13790.551$, respectively. When $K = 300$, $\lambda = \overline{\theta}= 125589.052$, while $\gamma_n^* = 41$. For $K \in \{2,10,40,160\}$, we have $\gamma_n^* = 0$ under this DGP.
In this section, we consider an empirical application of IV regressions with underlying specifications based on card(2009), Goldsmith(2020), and DKM24. Specifically, we consider a single cross-section of data in year $2000$ across $124$ locations (cities) by using the following model:
where $\beta_s$ is the coefficient of interest and can be interpreted as the (negative) inverse elasticity of substitution between immigrants and natives in the relevant skill group $s$. In addition, $Y_{ls}$ denotes the difference between the residual log wages\footnote{As discussed in card(2009), residual log wages are log wages after controlling for education, age, gender, race, and ethnicity of the U.S. workforce.} for immigrant and native men in skill group $s \in \{h,c\}$ (high school or college equivalent) and location (city) $l = 1,...,124$. The vector of location-level controls is denoted as $W_l$; in this application we include the following controls: $(1)$ log of city size, $(2)$ college shares, $(3)$ manufacturing shares in both $(i)$ $1980$ and $(ii)$ $1990$, $(4)$ mean wage residuals for $(i)$ all natives and $(ii)$ all immigrants in $1980$, together with $(5)$ a constant (so that there are $9$ controls available for each city, i.e., $W_l \in \mathbb{R}^9$).\footnote{See Table 6 in card(2009) for more details on the controls.} $X_{ls}$ denotes the log ratio of immigrant to native hours worked in skill group $s$ of both men and women in the city $l$.
The ratio $X_{ls}$ is potentially endogenous because unobserved city-specific factors may shift the relative demand for immigrant workers, leading to higher relative wages and higher relative employment shares, thereby confounding the estimation of the inverse elasticity of substitution. To overcome this issue, card(2009) suggests using the ratio of the total number of immigrants from foreign country $m$ in city $l$ to the total number of immigrants from country $m$ as an instrument. The rationale for such a choice is that existing immigrant enclaves are likely to attract additional immigrant labor through social and cultural channels unrelated to labor market outcomes. To define the instruments, we can exploit settlement patterns at some initial period (possibly together with the arrival rate of immigrants from specific countries in subsequent periods) to determine the inflow of immigrants in each location. Specifically, we let $N_{lm,1980}$ be the number of immigrants from country $m = 1...,38$ settling into location $l$ in $1980$ and let $N_{l,1980}$ be the total number of immigrants in location $l$ in $1980$, respectively. In addition, we let $P_{l,2000}$ denote the population size of location $l$ in $2000$, including both immigrants and natives.
To proceed, we consider four sets of potential instruments for $X_{ls}$. The definition of the first two sets of instruments follows from DKM24. Specifically, we let the instruments for each location $l$ be given by $z_{l,1980} := \left\{ z_{lm,1980} \right\}_{m=1}^{38} \equiv \left\{ \frac{N_{lm,1980}}{N_{l,1980}} \times \frac{1}{P_{l,2000}} \right\}_{m=1}^{38} \in \mathbb{R}^{38 \times 1}$, so that our first set of instruments can be written as $Z_{38} := (z_{1,1980},...,z_{124,1980} )^{\top} \in \mathbb{R}^{124 \times 38}$. For the second set of instruments, we let each of the 38 IVs interact with the $9$ controls described above, so that $z_{l,1980} \in \mathbb{R}^{342 \times 1}$ (i.e., each $z_{lm,1980}$ is interacted with $9$ controls). Then, our second set of instruments is defined as $Z_{342} \in \mathbb{R}^{124 \times 342}$. Furthermore, given that our proposed bootstrap test is dimensional-robust, we consider a case with $K=1$ for our third and fourth sets of instruments for the high school and college skill groups, respectively. Specifically, following Goldsmith(2020), we construct Bartik-type instruments, which are given by $Z_{Bartik,s} = \{B_{ls}\}_{l=1}^{124} \in \mathbb{R}^{124 \times 1}$ for $s \in \{h,c\}$, where $B_{ls}:= \sum_{m=1}^{38} z_{l,1980} \times g_{ms}$ and $g_{ms}$ is the number of immigrants from country $m$ in skill group $s$ arriving in US from $1990$ to $2000$. Note that while the Bartik IVs depend on the skill group $s$ (i.e., $Z_{Bartik,h}$ and $Z_{Bartik,c}$), the first and second sets of instruments (i.e., $Z_{38}$ and $Z_{342}$) do not depend on $s$.
The empirical results of our bootstrap test and those in DKM24 are given in Tables (ref)--(ref). For the Bartik instruments (i.e., $K=1$), the result for $\rm JAR_{\rm cf}$ is not reported because the cross-fit variance estimator is negative. Table (ref) shows the $95\%$ confidence intervals (CIs) with the Bartik IVs. $Z_{Bartik,h}$ and $Z_{Bartik,c}$ are applied separately to their respective skill groups. The regularizers for methods $\rm RJAR$ and $\rm BS$ are $\gamma_n^* = 0$ and $\lambda = 0$, respectively, with ${p_n'} = 0.077$ and $p_n = 0.216$.\footnote{When $K=1$, ${p_n'}(\lambda)$ and $p_n(\lambda)$ are independent of $\lambda$, and we set $\lambda = 0$.} $K_\lambda$ for BS is equal to $0.457$ and $0.253$ for high-school and college workers, respectively. In addition, the number beneath each CI represents its relative length compared to the $\rm BS$ CI. For $K=1$, all CIs have similar lengths. Methods $\rm RJAR$ and $\rm JAR_{std}$ have shorter CIs, but this is because these methods may not control size when $K$ is fixed and tend to over-rejects under the null, as observed in our simulation studies. Among the CIs that are theoretically valid for small $K$ (i.e., $\rm AR$, $\rm BS$, and $\rm BCCH$), the $\rm BS$ CIs are the shortest across both skill groups.
Table (ref) reports the $95\%$ CIs for high-school and college workers, respectively, with $K = 38$; the set of instruments used for both skill groups is $Z_{38}$. The regularizers for methods $\rm RJAR$ and $\rm BS$ are $\gamma_n^* = 0$ and $\lambda = 13.8$, respectively, with $p_n = 0.022$ and ${p_n'} = 0.089$. $K_\lambda$ for BS is equal to $2.015$ for 38 IVs. We find that $\rm BS$ has the shortest CI for college workers, while $\rm JAR_{\rm cf}$ has the shortest CI for high-school workers. But based on our simulation studies, $\rm JAR_{\rm cf}$ may over-reject under the null, which can result in shorter CIs. Furthermore, the $\rm BS$ CIs are shorter than their $\rm BCCH$ counterparts for both high-school and college workers.
Table (ref) shows the $95\%$ CIs with $K = 342$; the set of instruments used for both skill groups is $Z_{342}$. The regularizers for methods $\rm RJAR$ and $\rm BS$ are $\gamma_n^* = 5.3$ and $\lambda = 67.4$, respectively, with $p_n = 0.016$ and ${p_n'} = 0.089$. $K_\lambda$ for BS is equal to $1.550$ for 342 IVs. For both high-school and college workers, $\rm CT$ rejects all null hypotheses and thus results in empty confidence intervals, potentially due to heteroskedastic errors. $\rm BS$ again has the shortest confidence interval for college workers, and is of similar length with $\rm RJAR$ for high-school workers. Finally, $\rm BCCH$ has relatively wide CIs compared with $\rm BS$ and $\rm RJAR$.
\singlespacing