Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Robust Inference with High-Dimensional Instruments
abstractWe propose a weak-identification-robust test for linear instrumental variable (IV) regressions with high-dimensional instruments, whose number is allowed to exceed the sample size.
In addition, our test is robust to general error dependence, such as network dependence and spatial dependence. The test statistic takes a self-normalized form and the asymptotic validity of the test is established by using random matrix theory. Simulation studies are conducted to assess the numerical performance of the test, confirming good size control and satisfactory testing power across a range of various error dependence structures.
Keywords: High-Dimensional Instrumental Variables, Weak Identification, Robust Inference, Self Normalization, Random Matrix Theory.
JEL Classification: C12; C22; C55
Introduction
Various recent surveys in leading economics journals suggest that weak instruments remain important concerns for empirical practice. For instance, I.Andrews-Stock-Sun(2019) survey 230 instrumental variable (IV) regressions from 17 papers published in the American Economic Review (AER). They find that many of the first-stage F-statistics (and non-homoskedastic generalizations) are in a range that raises such concerns, and virtually all of these papers report at least one first-stage F with a value smaller than 10.
Similarly, in lee2021's (lee2021) survey of 123 AER articles involving IV regressions, 105 out of 847 specifications have first-stage Fs smaller than 10.
Moreover, many IV applications involve a large number of instruments.
For example, in their seminal paper, Angrist-Krueger(1991) study the effect of schooling on wages by interacting three base instruments (dummies for the quarter of birth) with state and year of birth, resulting in 180 instruments. Hansen-Hausman-Newey(2008) show that using the 180 instruments gives tighter confidence intervals than using the base instruments even after adjusting for the effect of many instruments.
In addition, as pointed out by MS22, in empirical papers that employ the “judge design" (e.g., see maestas2013, sampat2019, and dobbie2018), the number of instruments (the number of judges) is typically proportional to the sample size, and the famous Fama-MacBeth two-pass regression in empirical asset pricing (e.g., see fama1973, shanken1992, and anatolyev2022) is equivalent to IV estimation with the number of instruments proportional to the number of assets.
Similarly, belloni2012 consider an IV application involving more than one hundred instruments for the study of the effect of judicial eminent domain decisions on economic outcomes.
carrasco2015 used many instruments in the estimation of the elasticity of intertemporal substitution in consumption.
Furthermore, as pointed out by Goldsmith(2020), the shift-share or Bartik
instrument (e.g., see bartik1991 and blanchard1992), which has been widely applied in many fields such as labor, public, development, macroeconomics, international trade, and finance, can be considered as a particular way of combining many instruments. For example, in the canonical setting of estimating the labor supply elasticity, the corresponding number of instruments is equal to the number of industries, which is also typically proportional to the sample size.
Similar patterns of many IVs occur in wind-direction IVs, granular IVs, local average treatment effect estimation, and Mendelian randomization.\footnote{E.g., see deryugina2019mortality, bondy2020crime, gabaix2024, blandhol2022, boot2024, sloczynski2024should, davey2003, and davies2015.}
The problem of inference under many instruments has been well studied in the literature (e.g., see the literature review below). However, to our knowledge, there still does not exist a valid inference method that is robust to the high dimensionality of instruments and general error dependence, such as network and spatial dependence, simultaneously. To fill this gap, we propose in this paper a novel procedure for inference under high-dimensional IVs. Our test statistic takes a self-normalized form, and we establish the asymptotic validity of our test under general error dependence structure, by using random matrix theory (RMT). Additionally, we derive the power property of our test under many instruments. Monte Carlo simulations are conducted to assess the numerical performance of the test, confirming good size control and satisfactory testing power across a range of various error dependence structures.
Literature Review. The contributions in the present paper is related to the large literature on many (weak) instruments.\footnote{See, for example, Kunitomo1980, morimune1983, Bekker(1994), donald2001, chamberlain2004, Chao-Swanson(2005), stock2005, han2006, Andrews-Stock(2007), Hansen-Hausman-Newey(2008), Newey-Windmeijer(2009), anderson2010, kuersteiner2010, anatolyev2011, belloni2011, okui2011, belloni2012, carrasco2012,
Chao(2012), Haus2012,
K13,
hansen2014, carrasco2015, Wang_Kaffo_2016, kolesar2018, EK18, solvsten2020, CNT23, boot2024,
among others.}
In the context of many instruments and heteroskedasticity, Chao(2012) and Haus2012 provide standard errors for Wald-type inferences that are based on JIVE and jackknifed versions of the limited information maximum likelihood (LIML) and Fuller(1977)'s (Fuller(1977)) estimators (HLIM and HFUL). These estimators are more robust to many instruments than the commonly used two-stage least squares (TSLS) estimator because they can correct the bias caused by the high dimension of IVs.\footnote{Specifically, the rate of growth of the concentration parameter, which measure the overall instrument strength, is denoted as $\mu_n^2$. JIVE, HLIM, and HFUL remain consistent with heteroskedastic errors even when instrument weakness is such that $\mu_n^2$ is slower than the number of instruments $K$, provided that $\mu_n^2/\sqrt{K} \rightarrow \infty$ as the number of observations $n \rightarrow \infty$ (Chao(2012), Haus2012). In contrast, TSLS is less robust to instrument weakness as it is shown to be consistent only under homoskedasticity if $\mu_n^2/K \rightarrow \infty$ Chao-Swanson(2005). In simulations derived from the data in Angrist-Krueger(1991), which is representative of empirical labor studies with many instrument concerns, Angrist-Frandsen2022 show that such bias-corrected estimators outperform the TSLS that is based on the instruments selected by the least absolute shrinkage and selection operator (LASSO) introduced in belloni2012 or the random forest-fitted first stage introduced in athey2019.}
Furthermore, carrasco2012, carrasco2015, carrasco2016efficient, hansen2014, and carrasco2017 proposed regularization approaches for two-stage least squares, limited information maximum likelihood, and jackknife IV (Angrist(1999)) estimators. Under many weak moment asymptotics, Newey-Windmeijer(2009) provide new variance estimators for the jackknife GMM and the class of generalized empirical likelihood (GEL) estimators, which includes the continuous updating estimator (CUE) and EL estimator as special cases.\footnote{In the linear heteroskedastic IV model, consistency and asymptotic normality of CUE require $m^2/n \rightarrow 0$ and $m^3/n \rightarrow 0$, respectively, where $m$ and $n$ denote the number of moment conditions and the sample size (e.g., see p.689 of Newey-Windmeijer(2009)). Such conditions are needed to simultaneously control the estimation error for all the elements of the heteroskedasticity consistent weighting matrix. Somewhat stronger rate conditions are required for other GEL estimators.}
However, the Wald-type inference methods are invalid under weak identification, which occurs when the concentration parameter remains bounded as the sample size increases to infinity. In this case, all the estimators mentioned earlier become inconsistent, and there is no consistent test for the structural parameter of interest.\footnote{E.g., see Section 3 of MS22.}
For weak-identification-robust inference based on the classical AR test, Andrews-Stock(2007) show its validity under many instruments, but their IV model is homoskedastic and requires the number of instruments to diverge slower than the cube root of the sample size ($K^3/n \rightarrow 0$). Newey-Windmeijer(2009) proposed a GMM-AR test under many (weak) moment conditions but imposed the same rate condition on $K$. anatolyev2011 constructed a modified AR test that allows the number of instruments to be proportional to the sample size but requires homoskedastic errors, and Kaffo-Wang(2017) proposed a bootstrap version of anatolyev2011's test. Furthermore, carrasco2016 first proposed a ridge-regularized AR test that allows for $K$ being larger than $n$ with homoskedastic errors. Bun-Farbmacher-Poldermans(2020) compared the centered and uncentered GMM-AR test and identified a missing degrees-of-freedom correction when $K/n \rightarrow 0$.
Recently, crudu2021 and MS22 proposed jackknifed versions of the AR test in a model with many instruments and general heteroskedasticity.
DKM24 developed a ridge-regularized version of the jackknife AR test.
tuvaandorj2024 established the validity of a permutation AR test under heteroskedasticity and diverging $K$, requiring $K^3/n \rightarrow 0$. boot-ligtenberg(2023) developed a dimension-robust AR test based on continuous updating but relied on an invariance assumption.
Furthermore, LWZ2024dimension proposed a dimension-agnostic bootstrap AR test by deriving strong approximation results for both test statistic and its bootstrap version. Their bootstrap AR test is valid with heteroskedastic errors uniformly across a broad asymptotic regime for $K$, spanning from fixed to diverging faster than the sample size.
In addition to the AR tests, Matsushita2020 propose a jackknife LM test. Lim2024 consider a linear combination of jackknife AR, jackknife LM, and orthogonalized jackknife LM tests and find that the resulting conditional linear combination (CLC) test has good power properties in a variety of scenarios.
N23 proposed a dimension-robust version of Kleibergen(2002)'s K test, and his method relies on a sparse $\ell_1$-regularized estimation of $\rho(Z_i)$, the conditional correlation between the endogenous variable and the outcome error.
Yap24 proposed a leave-three-out variance estimator, in the same spirit as that proposed by AS23, for inference under heterogeneous treatment effects.
The above methods are robust to weak identification, many instruments, and heteroskedastic errors.
However, the literature on inference procedures that are robust to more general error dependence structure remains sparse. ligtenberg2023 proposed a leave-cluster-out AR test that is robust to cluster dependence structure under many weak instruments. \nocite{Wang-Zhang2024}
The rest of this paper is organized as follows. In Section (ref), we formally define the observed intermediary, which is used to formulate our proposed statistic. Then, we establish their limiting distributions with power theories.
Section (ref) reports the finite sample performance of our proposed statistics across multiple scenarios through Monte Carlo simulations.
The Appendices are provided separately as the supplementary document, including
proof of theorems and relevant lemmas.
Model and Test Statistic
We consider the linear IV regression with a scalar outcome $y_i$, a scalar endogenous variable $x_i$, and a $K \times 1$ vector of instruments $z_i$ such that
align[align omitted — 115 chars of source]
Following the literature on many instruments, we assume that there are no (low-dimensional) controls included in our model as they can be partialled out from $(Y_i,X_i,Z_i)$.
The source of endogeneity is caused by a correlation between $\varepsilon_{i}$ and $v_{i}$, which represent the errors associated with the structural and the first-stage equation, respectively. The two sequences of regression errors, $\left\{ \varepsilon_{i}\right\} _{i=1}^{N}$ and $\left\{ v_{i}\right\} _{i=1}^{N}$, follow a mean-zero random process whose functional forms are unknown and potentially heteroskedastic. The unknown parameters are $\beta\in\mathbb{R}$ and $\pi\in\mathbb{R}^{K\times1}$, where $K$ is comparable or exceeds its sample size.
We are interested in testing the null hypothesis
equation[equation omitted — 104 chars of source]
even when the identification for the structural parameter of interest $\beta$ may be weak.
We focus on the model with a scalar endogenous variable for two reasons. First, in many empirical applications of IV regressions, there is only one endogenous variable (as can be seen from the surveys by Andrews-Stock-Sun(2019) and lee2021). Second, the asymptotic results derived in the paper extend directly to the general case of full-vector inference with multiple endogenous variables.
Additionally, for the subvector inference, one may use a projection approach Dufour-Taamouti(2005) after implementing our test on the whole vector of endogenous variables.\footnote{Alternative subvector inference methods for IV regressions (e.g., see GKMC(2012), Andrews(2017), and GKM(2019), GKM(2021)) provide a power improvement over the projection approach under fixed $K$. However, whether they can be applied to the current setting is unclear.
Also, Wang-Doko(2018) and Wang(2020) show that bootstrap tests based on the standard subvector AR statistic may not be robust to weak identification even under fixed $K$ and conditional homoskedasticity.}
To proceed, let us denote $y^*_{i}=y_{i}-x_{i}\beta_{0}.$
Let a singular value decomposition of the normalized instrument matrix be expressed as
align[align omitted — 115 chars of source]
where $Z:=\left(z_{1},\ldots,z_{N}\right)^{'}\in\mathbb{R}^{N\times K},$ $\lambda_{\ell}$ denotes the $\ell$-th largest non-zero eigenvalue of the matrix $S_{N}:=\frac{1}{N}Z^{'}Z,$ $q_{\ell}$ and $w_{\ell}$ denotes the associated orthonormal eigenvector of the $\ell$-th largest eigenvalue of $\underline{S}_{N}:=\frac{1}{N}ZZ^{'}$ and $S_{N}$, respectively.
To conduct the hypothesis testing in (ref), the associated test statistic proposed in this paper can be obtained by equating the sample analogue of
equation[equation omitted — 166 chars of source]
where $Y^{*}=Y-X\beta_{0}$, $Y=\left(y_{1},\ldots,y_{N}\right)^{'}$,
$X=\left(x_{1},\ldots,x_{N}\right)^{'}$,
and $\gamma_{j}^{2}:=\mathbb{E}\varepsilon_{j}^{2}$.
Similar (high-dimensional) moment condition was studied in Feng2024 to overcome the curse of dimensionality of coefficients in linear regression models with a presence of (unconditional) heteroskedasticity and autocorrelation of an unknown form.
After substituting $\gamma_{j}^{2}$ with its plug-in estimator under the null, namely $\frac{Y^{*'}Y^{*}}{N}$, the statistic of interest is obtained by evaluating its associated sample counterpart:
align[align omitted — 321 chars of source]
where $\underline{Y}:=Y^{*}/\left(Y^{*'}Y^{*}\right)^{1/2}\in\left(0,1\right)$. Then, with a proper convergence rate (shown in Theorem (ref)), our (oracle) statistic of interest is given by
align[align omitted — 435 chars of source]
Unless stated otherwise, we denote $\underline{a}:a/\left(a^{'}a\right)^{1/2}$ a self-normalized version of any column vector $a.$
To build intuition, suppose for the moment that the errors are homoskedastic, i.e., $\mathbb{E}(\varepsilon \varepsilon' |Z) = \gamma^2I_N$, where $\varepsilon = (\varepsilon_1, ..., \varepsilon_N)'$.
By the law of iterated expectation,
align[align omitted — 438 chars of source]
under the null hypothesis, where $\ell$ is randomly chosen from the set $\{1, ..., \min\{K,N\}\}$.
We note that $q'_{\ell}q_{s}$ and $w'_{\ell}w_s$ are equal to one when $\ell = s$ for all $\ell, s = 1, ..., \min\{K,N\}$ but zero otherwise. On the other hand, under the alternative $\beta \neq \beta_0$,
align[align omitted — 220 chars of source]
where $\Delta = \beta - \beta_0$.
For general error dependence, i.e., $\mathbb{E}(\varepsilon \varepsilon' | Z) = \Psi$, it is expected that under the null hypothesis,
align[align omitted — 265 chars of source]
where $\Psi^2_{ii}$ denotes the $i$-th diagonal entry of $\Psi$. The proposed test becomes robust against general error dependence with high-dimensional IVs if there exists, for some $i \in \{1, ..., N\}$, the sequence $\{ (q_{\ell}' \varepsilon \cdot \Psi_{ii}^{-1})^2\}_{\ell=1}^{\min\{K, N\}}$
converges to one in probability as $K$ and $N$ grow,
where the quantity $q_{\ell}'\varepsilon \cdot \Psi_{ii}^{-1}$ is the inner product of an eigen basis generated by the Gram matrix $\underline{S}_N$ and the normalized regression errors.
We establish the asymptotic validity of our inference procedure under the following regularity conditions.
assumption[High-dimensional asymptotic regime]
The dimension of instruments is considered high-dimension such that $K\rightarrow\infty,$ $N\rightarrow\infty$, and $K/N \rightarrow c\in\left(0,\infty\right)$.
With rank deficiency generally observed in sample covariance matrices, this regime is widely regarded as a standard high-dimensional paradigm in the literature Ledoit2002, Bai2007, Onatski2013. This paper introduces a more adaptable framework at a general level of intensity beyond that of Khan2024, where the number of covariates may exceed its sample size.
assumption[Instruments]
The data generating process of $Z=\left(z_{1},\ldots,z_{N}\right)^{'}\in\mathbb{R}^{N\times K}$ is given by $z_{i}=\Sigma^{1/2}f_{i},$ where $f_{i}=\left(f_{1i},\ldots,f_{Ki}\right)^{'}$ for all $i=1,\ldots,N$, $\Sigma^{1/2}$ is the square root of a deterministic non-negative definite Hermitian matrix $\Sigma,$ and $\lambda_{max}\left(\Sigma\right)=O\left(1\right)$. Moreover, $f_{i}$ consists of $K$ independent and identically distributed (i.i.d.) random variables with mean zero, unit variance, and finite fourth moment.
Assumption (ref) is recognized as a typical design in RMT literature Karoui2008, dobriban2018high, hastie2022surprises, Li2024, farbmacher2024revisiting, Zhang2025,which we employ to characterize correlations within instruments. The instruments, therefore, possess a weak form of cross-sectional dependence defined in Chudik2011. Qualitative instruments can also be included, but require standardization or variable transformation to ensure zero mean, unit variance, and a finite fourth moment (see, Theorem (ref)).
Assumption (ref) allows us to apply RMT, which helps to generalize the error dependence structure, in the current setting with high-dimensional instruments, as can be seen in Assumption (ref) below.
assumptionDenote $a_{-i}$ a whole sequence of the random variable $a$ with the $i$-th entry removed. The two unobserved errors in the IV model are given by
\[
\varepsilon_{i}=\Xi_{\varepsilon}\left(i,\varepsilon_{-i}^{'},v^{'},\mathcal{M}^{'}\right), \;\; \text{and} \;\;
v_{i}=\Xi_{v}\left(i,v_{-i}^{'},\varepsilon^{'},\mathcal{M}^{'}\right),
\]
where $\Xi_{\varepsilon}$ and $\Xi_{v}$ are unknown functions and $\mathcal{M}=\left(m_{1}^{'},\ldots,m_{N}^{'}\right)^{'}$ where $m_{i}$ denotes a vector of exogenous variables that jointly explains the variation of the error at individual $i$ for all $i=0, \ldots, N$.
The data generating process of $z$ is independent of $\varepsilon$ and $v$.
The permitted structure of errors are general and flexible, accommodating complicated network or spatial dependence structure into account. Assumption (ref) ensures the mean independence of $\varepsilon$ and $v$ against instruments $z$. Specifically, $\mathbb{E}\left(\left.f_{i}^{\delta}\right|\varepsilon_{i},v_{i}\right)=\mathbb{E}\left(f_{i}^{\delta}\right)$, for $0<\delta\leq4$.
Define the self-normalized vector $u:=\left(u_{1},\ldots,u_{N}\right)$ where $u_{i}=\varepsilon_{i}/\left(\varepsilon^{'}\varepsilon\right)^{1/2}.$ The following theorem provides the null limiting distribution of $Q_{N}$, the oracle statistic, under high-dimensional IVs with a presence of (unconditional) heteroskedasticity in error processes.
thm[Oracle statistic]
Suppose that Assumptions (ref)-(ref) are satisfied. Further suppose that (i) $\mathbb{E}\left(f_{11}^{4}\right)=\kappa<\infty$,and (ii) $\sum_{i=1}^{N}\left|u_{i}^{3}\right|=o_{\mathbb{P}}\left(1\right).$
Then, under the null hypothesis in (ref),
\[
Q_{N}=\sqrt{\frac{N^{2}}{2\text{tr}\left(\Sigma^{2}\right)}}\left(\underline{Y}{}^{'}\underline{S}_{N}\underline{Y}-\frac{1}{N}\text{tr}\left(\underline{S}_{N}\right)\right)\overset{d}{\rightarrow}\mathcal{N}\left(0,1\right).
\]
Replacing $\text{tr}\left(\Sigma^{2}\right)$ with the consistent estimator proposed by Li2012,
we have the following asymptotic result for the feasible statistic.
cor[Feasible statistic]
Under Assumptions (ref)-(ref), and all conditions in Theorem (ref), we further assume that $\sum_{i=1}^{N}u_{i} = O_{\mathbb{P}}\left(1\right).$ Then,
\[
\widehat{Q}_{N} := \sqrt{\frac{N^{2}}{2\widehat{\text{tr}\left(\Sigma^{2}\right)}}}\left(\underline{Y}^{'}\underline{S}_{N}\underline{Y}-\frac{1}{N}\text{tr}\left(\underline{S}_{N}\right)\right)\overset{d}{\rightarrow}\mathcal{N}\left(0,1\right),
\]
where $\widehat{\text{tr}\left(\Sigma^{2}\right)}:=\frac{1}{N\left(N-1\right)}\sum_{i\neq j}^{N}\left(z_{i}^{'}z_{j}\right)^{2}$ is a ratio-consistent estimator of $\text{tr}\left(\Sigma^{2}\right)$.
Let $\Delta = \beta - \beta_0$. The alternative with local departure is defined as
equation[equation omitted — 145 chars of source]
where $h = \Delta \cdot N^{1/2}/ (2 tr(\Sigma^2))^{1/5}$ denotes a (deterministic) level of departure.
By construction, the quantity $\left(\frac{\left(2\text{tr}\left(\Sigma^{2}\right)\right)^{2/5}}{N}\right)^{1/2}$ implies a rate of signal strength in the coefficient of interest that the test can detect.
thm[Power theory]
Let Assumptions (ref)-(ref) hold. In addition, suppose that: (i) $\frac{1}{N}\varepsilon^{'}\varepsilon\overset{p}{\rightarrow}\alpha_{1}$ and $\frac{1}{N}v^{'}v\overset{p}{\rightarrow}\alpha_{2}$ where $\alpha_{1},\alpha_{2}\in\left(0,\infty\right)$, (ii) $\mathbb{E}\left(f_{11}^{8}\right)<\infty$,
and (iii) $\pi'\pi/\sqrt{K}$ is bounded.
Then, $Q_{N}-\varpi_{N}\overset{d}{\rightarrow}\mathcal{N}\left(0,1\right),$ where
\begin{align*}
\varpi_{N} & =\frac{N}{\left\Vert \varepsilon\right\Vert ^{2}\left(2tr\left(\Sigma^{2}\right)\right)^{1/10}}h^{2} \pi'\left(S_{N}^{2}
-\frac{1}{N}tr\left(S_{N}\right)I_{N}\right)\pi \\
& + \frac{N}{\left\Vert \varepsilon\right\Vert ^{2}\left(2tr\left(\Sigma^{2}\right)\right)^{1/10}}h^{2}
v'\left(\frac{1}{N} S_N - \frac{1}{N^2} tr(S_N) I_N \right)v \\
& + \frac{2 N^{1/2}}{\left\Vert \varepsilon\right\Vert ^{2}\left(2tr\left(\Sigma^{2}\right)\right)^{3/10}}hv^{'}\left(S_{N} - \frac{1}{N}\text{tr}\left(S_{N}\right)I_{N}\right)\varepsilon.
\end{align*}
Several comments are in order.
First, the power of the feasible statistic is analogous to that of the oracle one, since one merely replaces the unknown quantity $\text{tr}\left(\Sigma^{2}\right)$ with its consistent estimator. Namely $\left(\widehat{Q}_{N}-Q_{N}\right)+Q_{N}-\varpi_{N}=Q_{N}-\varpi_{N}+o_{\mathbb{P}}\left(1\right).$
Second, Condition (ii) can be relaxed to a finite fourth moment with a cost of technical complication.
Monte Carlo Simulations
In this part, we present the power performance of our proposed statistics across a range of data generating processes. The statistical software adopted for numerical results is MATLAB 2023a with the default seed. We further note that all parameters and functional forms postulated in the error processes are arbitrarily chosen. To design regimes of high dimensionality, we fix the sample size to $400$, then vary the number of instruments so that the data intensities ($K/n$) are $1/4,1/2,1,2$ and $3$. The numerical size (power) of statistics is calculated as the probability of null rejections over 1,000 replications, given the data is generated under the null (alternative) hypothesis. Throughout the analysis, an intercept is set to $2$ and a level of significance is set to 0.05.
table[table omitted — 2,783 chars of source]
The errors in the structural equation, i.e., $\varepsilon$, are generated by one of the three following processes: (i) the network dependent process, (ii) the spatial dependent process, and (iii) the multiplicative heteroskedasticity of a mixture of two i.i.d. processes.
Specifically, for the network dependence process, we follow the simulation design of the network formation model in Section 5 of Kojevnikov2021, and set $\gamma=0.5$ in (5.1) of their model. The processes (ii) and (iii) were also implemented in Feng2024. For the case with spatial dependence, the errors follow the spatially autoregressive error process of order one studied by kim2011spatial, that is,
align*[align* omitted — 45 chars of source]
where $\rho_s=0.8$ and the innovation $e$ is independently generated by a standardized $t$-distribution with 5 degrees of freedom. In addition, $W=[c_iw_{ij}]$ is a $N \times N$ spatial weighting matrix. To specify such a matrix, we set all diagonal entries to zero and quantify the off-diagonal entry $w_{ij}$ to one if and only if $w_{ij}=a>0.5$, where $a$ is drawn from a uniform distribution over an interval $[0,1]$, otherwise zero. Denote $c_i$ a scalar that standardizes each row of $W$ to sum to unity.
For the multiplicative heteroskedasticity and autocorrelation of a mixture of two i.i.d. processes, we let
align*[align* omitted — 82 chars of source]
where $\zeta_i = 1 + \left( a i/N \right)$
and $\omega_i = \sqrt{1-0.7^2} \eta_{1,i} + 0.7 \eta_{2,i}$,
$\eta_{1,i}$ is generated by i.i.d. standardized $t$-distribution with 5 degrees of freedom, and $\eta_{2,i}$ is generated by i.i.d. standardized chi-square distribution with 6 degrees of freedom. As a result, the underlying process $\omega_i$ is considered both heavy-tailed and highly-skewed.
The scalar $s_1:= 289a^2/300 + (1+ a/2)^2$ and $s_2:= 1.2(a+2)$ normalizes the random variable $\varepsilon_i$ to mean zero and variance one. Under this design, the errors are equi-correlated with a correlation approximately equal to $0.52$ for $i, j=1, ..., N$. The impact of the skedastic component is set to $a=10$.
For brevity, the three processes are abbreviated as {[}NET-E{]}, {[}SPA-E{]}, and {[}MUL-E{]}, respectively. The errors in the first-stage equation are correlated with the other given by $v_{i}=\rho\varepsilon_{i}+\sqrt{1-\rho^{2}}\eta_{3,i}$, where $\eta_3$ is independently generated by a standardized $t$-distribution with $5$ degrees of freedom (abbreviated as $t_{5}$ hereafter). The two processes are strongly correlated, taking values of $\rho\in\left\{ -.9,.5,.9\right\} .$
Being consistent with our asymptotic theory, the instruments are generated according to a latent factor model studied in Section 2.1 of bhattacharya2011sparse. Namely,
align*[align* omitted — 60 chars of source]
where $\eta_{4,i}\in\mathbb{R}^{3}$ and $f_{i}\in\mathbb{R}^{K}$ are independently drawn from $t_{5}$ and the standard normal distribution, respectively. $\Lambda$ is a $K\times3$ loading matrix such that $\Lambda^{'}\Lambda=\text{diag}\left(6,5,3\right)$. The population covariance matrix $\Sigma$ is the Toeplitz matrix with the $\left(i,j\right)$-th elements equaling to $0.7^{\left|i-j\right|}$. The strength of IVs is set to be weak such that $\pi^{'}\pi=1$ throughout the simulations.
The finite sample performance of the feasible statistic is analogous to that of the oracle one. To save space, we thus only report the results for the feasible test. We highlight several findings below.
First, the simulation results illustrate that the empirical size of our test is close to the nominal level across different settings of the number of instruments ($K/N$), the degree of endogeneity ($\rho$), and the error dependence structure.
Second, we note that the power of our test improves as the departure $h$ increases and attains unity, ensuring complete distributional separation between hypotheses, when $h$ exceeds 5.
Third, in line with our asymptotic theory, the correlation between the two error processes influences the power of the test. A negative contemporary dependence enhances the power overall, as well as its magnitude.
Conclusion
In this paper, we propose a weak-identification-robust test for linear instrumental variable (IV) regressions with high-dimensional instruments, whose number is allowed to exceed the sample size.
In addition, our test is robust to general error dependence, such as network dependence and spatial dependence. The test statistic takes a self-normalized form and the asymptotic validity of the test is established by using RMT. Simulation studies are conducted to assess the numerical performance of the test, confirming good size control and satisfactory testing power across a range of various error dependence structures. For future research agenda, we note that although our new test is robust to high-dimensional instruments under general error dependence, it is not robust to the case with a fixed number of instruments.
On the other hand, LWZ2024dimension recently proposed a dimension-agnostic bootstrap AR test under many weak instruments and (independent) heteroskedastic errors by using strong approximation. For dimension-robust inference under the current setting, it may be interesting to explore a bootstrap-based inference procedure and combine it with RMT.
thebibliography\bibitem[\citeauthoryear{Anatolyev and Gospodinov}{Anatolyev and Gospodinov}{2011}]{anatolyev2011}
Anatolyev, S. and N. Gospodinov (2011).
\newblock Specification testing in models with many instruments.
\newblock {\em Econometric Theory\/} {\em 27\/}(2), 427--441.
\bibitem[\citeauthoryear{Anatolyev and Mikusheva}{Anatolyev and Mikusheva}{2022}]{anatolyev2022}
Anatolyev, S. and A. Mikusheva (2022).
\newblock Factor models with many assets: strong factors, weak factors, and the two-pass procedure.
\newblock {\em Journal of Econometrics\/} {\em 229\/}(1), 103--126.
\bibitem[\citeauthoryear{Anatolyev and S{\o}lvsten}{Anatolyev and S{\o}lvsten}{2023}]{AS23}
Anatolyev, S. and M. S{\o}lvsten (2023).
\newblock Testing many restrictions under heteroskedasticity.
\newblock {\em Journal of Econometrics\/} {\em 236\/}(1), 105473.
\bibitem[\citeauthoryear{Anderson, Kunitomo, and Matsushita}{Anderson et al.}{2010}]{anderson2010}
Anderson, T., N. Kunitomo, and Y. Matsushita (2010).
\newblock On the asymptotic optimality of the liml estimator with possibly many instruments.
\newblock {\em Journal of Econometrics\/} {\em 157\/}(2), 191--204.
\bibitem[\citeauthoryear{Andrews}{Andrews}{2017}]{Andrews(2017)}
Andrews, D. W. (2017).
\newblock Identification-robust subvector inference.
\newblock Technical report, Cowles Foundation Discussion Paper 2105.
\bibitem[\citeauthoryear{Andrews and Stock}{Andrews and Stock}{2007}]{Andrews-Stock(2007)}
Andrews, D. W. K. and J. H. Stock (2007).
\newblock Testing with many weak instruments.
\newblock {\em 138\/}(1), 24--46.
\bibitem[\citeauthoryear{Andrews, Stock, and Sun}{Andrews et al.}{2019}]{Andrews-Stock-Sun(2019)}
Andrews, I., J. H. Stock, and L. Sun (2019).
\newblock \hypertarget{Andrews-Stock-Sun(2019)}{Weak instruments in instrumental variables regression: Theory and practice}.
\newblock {\em Annual Review of Economics\/} {\em 11\/}(1), 727--753.
\bibitem[\citeauthoryear{Angrist and Frandsen}{Angrist and Frandsen}{2022}]{Angrist-Frandsen2022}
Angrist, J. and B. Frandsen (2022).
\newblock Machine labor.
\newblock {\em Journal of Labor Economics\/} {\em 40\/}(S1), S97--S140.
\bibitem[\citeauthoryear{Angrist, Imbens, and Krueger}{Angrist et al.}{1999}]{Angrist(1999)}
Angrist, J., G. Imbens, and A. Krueger (1999).
\newblock \hypertarget{Angrist(1999)}{Jackknife instrumental variables estimates}.
\newblock {\em 14\/}(1), 57--67.
\bibitem[\citeauthoryear{Angrist and Krueger}{Angrist and Krueger}{1991}]{Angrist-Krueger(1991)}
Angrist, J. D. and A. B. Krueger (1991).
\newblock Does compulsory school attendance affect schooling and earning?
\newblock {\em Quarterly Journal of Economics\/} {\em 106\/}(4), 979--1014.
\bibitem[\citeauthoryear{Athey, Tibshirani, and Wager}{Athey et al.}{2019}]{athey2019}
Athey, S., J. Tibshirani, and S. Wager (2019).
\newblock Generalized random forests.
\newblock {\em The Annals of Statistics\/} {\em 47\/}(2), 1148--1178.
\bibitem[\citeauthoryear{Bai, Miao, and Pan}{Bai et al.}{2007}]{Bai2007}
Bai, Z. D., B. Q. Miao, and G. M. Pan (2007).
\newblock {On asymptotics of eigenvectors of large sample covariance matrix}.
\newblock {\em The Annals of Probability\/} {\em 35}, 1532 -- 1572.
\bibitem[\citeauthoryear{Bai and Silverstein}{Bai and Silverstein}{1998}]{bai1998apx}
Bai, Z. D. and J. W. Silverstein (1998).
\newblock No eigenvalues outside the support of the limiting spectral distribution of large-dimensional sample covariance matrices.
\newblock {\em The Annals of Probability\/} {\em 26}, 316--345.
\bibitem[\citeauthoryear{Bartik}{Bartik}{1991}]{bartik1991}
Bartik, T. J. (1991).
\newblock {\em Who benefits from state and local economic development policies? Kalamazoo, MI: WE Upjohn Institute for Employment Research}.
\bibitem[\citeauthoryear{Bekker}{Bekker}{1994}]{Bekker(1994)}
Bekker, P. (1994).
\newblock \hypertarget{Bekker(1994)}{Alternative approximations to the distributions of instrumental variable estimators}.
\newblock {\em Econometrica\/} {\em 62\/}(3), 657--681.
\bibitem[\citeauthoryear{Belloni, Chen, Chernozhukov, and Hansen}{Belloni et al.}{2012}]{belloni2012}
Belloni, A., D. Chen, V. Chernozhukov, and C. Hansen (2012).
\newblock Sparse models and methods for optimal instruments with an application to eminent domain.
\newblock {\em Econometrica\/} {\em 80}, 2369--2429.
\bibitem[\citeauthoryear{Belloni, Chernozhukov, and Hansen}{Belloni et al.}{2011}]{belloni2011}
Belloni, A., V. Chernozhukov, and C. Hansen (2011).
\newblock Inference for high-dimensional sparse econometric models.
\newblock {\em arXiv preprint arXiv:1201.0220\/}.
\bibitem[\citeauthoryear{Bhattacharya and Dunson}{Bhattacharya and Dunson}{2011}]{bhattacharya2011sparse}
Bhattacharya, A. and D. B. Dunson (2011).
\newblock Sparse bayesian infinite factor models.
\newblock {\em Biometrika\/} {\em 98}, 291--306.
\bibitem[\citeauthoryear{Blanchard, Katz, Hall, and Eichengreen}{Blanchard et al.}{1992}]{blanchard1992}
Blanchard, O. J., L. F. Katz, R. E. Hall, and B. Eichengreen (1992).
\newblock Regional evolutions.
\newblock {\em Brookings Papers on Economic Activity\/} {\em 1992\/}(1), 1--75.
\bibitem[\citeauthoryear{Blandhol, Bonney, Mogstad, and Torgovitsky}{Blandhol et al.}{2022}]{blandhol2022}
Blandhol, C., J. Bonney, M. Mogstad, and A. Torgovitsky (2022).
\newblock When is tsls actually late?
\newblock Technical report, National Bureau of Economic Research Cambridge, MA.
\bibitem[\citeauthoryear{Bondy, Roth, and Sager}{Bondy et al.}{2020}]{bondy2020crime}
Bondy, M., S. Roth, and L. Sager (2020).
\newblock Crime is in the air: The contemporaneous relationship between air pollution and crime.
\newblock {\em Journal of the Association of Environmental and Resource Economists\/} {\em 7\/}(3), 555--585.
\bibitem[\citeauthoryear{Boot and Ligtenberg}{Boot and Ligtenberg}{2023}]{boot-ligtenberg(2023)}
Boot, T. and J. W. Ligtenberg (2023).
\newblock Identification- and many instrument-robust inference via invariant moment conditions.
\newblock {\em arXiv:2303.07822\/}.
\bibitem[\citeauthoryear{Boot and Nibbering}{Boot and Nibbering}{2024}]{boot2024}
Boot, T. and D. Nibbering (2024).
\newblock Inference on lates with covariates.
\newblock {\em arXiv preprint arXiv:2402.12607\/}.
\bibitem[\citeauthoryear{Carrasco}{Carrasco}{2012}]{carrasco2012}
Carrasco, M. (2012).
\newblock A regularization approach to the many instruments problem.
\newblock {\em Journal of Econometrics\/} {\em 170\/}(2), 383--398.
\bibitem[\citeauthoryear{Carrasco and Doukali}{Carrasco and Doukali}{2017}]{carrasco2017}
Carrasco, M. and M. Doukali (2017).
\newblock Efficient estimation using regularized jackknife iv estimator.
\newblock {\em Annals of Economics and Statistics\/} (128), 109--149.
\bibitem[\citeauthoryear{Carrasco and Tchuente}{Carrasco and Tchuente}{2015}]{carrasco2015}
Carrasco, M. and G. Tchuente (2015).
\newblock Regularized liml for many instruments.
\newblock {\em Journal of Econometrics\/} {\em 186\/}(2), 427--442.
\bibitem[\citeauthoryear{Carrasco and Tchuente}{Carrasco and Tchuente}{2016a}]{carrasco2016efficient}
Carrasco, M. and G. Tchuente (2016a).
\newblock Efficient estimation with many weak instruments using regularization techniques.
\newblock {\em Econometric Reviews\/} {\em 35\/}(8-10), 1609--1637.
\bibitem[\citeauthoryear{Carrasco and Tchuente}{Carrasco and Tchuente}{2016b}]{carrasco2016}
Carrasco, M. and G. Tchuente (2016b).
\newblock Regularization based anderson rubin tests for many instruments.
\newblock Technical report, School of Economics Discussion Papers, University of Kent.
\bibitem[\citeauthoryear{Chamberlain and Imbens}{Chamberlain and Imbens}{2004}]{chamberlain2004}
Chamberlain, G. and G. Imbens (2004).
\newblock Random effects estimators with many instrumental variables.
\newblock {\em Econometrica\/} {\em 72\/}(1), 295--306.
\bibitem[\citeauthoryear{Chao and Swanson}{Chao and Swanson}{2005}]{Chao-Swanson(2005)}
Chao, J. C. and N. R. Swanson (2005).
\newblock \hypertarget{Chao-Swanson(2005)}{Consistent Estimation with a Large Number of Weak Instruments}.
\newblock {\em Econometrica\/} {\em 73\/}(5), 1673--1692.
\bibitem[\citeauthoryear{Chao, Swanson, Hausman, Newey, and Woutersen}{Chao et al.}{2012}]{Chao(2012)}
Chao, J. C., N. R. Swanson, J. A. Hausman, W. K. Newey, and T. Woutersen (2012).
\newblock Asymptotic distribution of jive in a heteroskedastic iv regression with many instruments.
\newblock {\em Econometric Theory\/} {\em 28\/}(1), 42--86.
\bibitem[\citeauthoryear{Chao, Swanson, and Woutersen}{Chao et al.}{2023}]{CNT23}
Chao, J. C., N. R. Swanson, and T. Woutersen (2023).
\newblock Jackknife estimation of a cluster-sample iv regression model with many weak instruments.
\newblock {\em Journal of Econometrics\/}.
\bibitem[\citeauthoryear{Chudik, Pesaran, and Tosetti}{Chudik et al.}{2011}]{Chudik2011}
Chudik, A., M. H. Pesaran, and E. Tosetti (2011).
\newblock Weak and strong cross-section dependence and estimation of large panels.
\bibitem[\citeauthoryear{Crudu, Mellace, and S{\'a}ndor}{Crudu et al.}{2021}]{crudu2021}
Crudu, F., G. Mellace, and Z. S{\'a}ndor (2021).
\newblock Inference in instrumental variable models with heteroskedasticity and many instruments.
\newblock {\em Econometric Theory\/} {\em 37\/}(2), 281--310.
\bibitem[\citeauthoryear{Davey Smith and Ebrahim}{Davey Smith and Ebrahim}{2003}]{davey2003}
Davey Smith, G. and S. Ebrahim (2003).
\newblock Mendelian randomization: Can genetic epidemiology contribute to understanding environmental determinants of disease?
\newblock {\em International journal of epidemiology\/} {\em 32\/}(1), 1--22.
\bibitem[\citeauthoryear{Davies, von Hinke Kessler Scholder, Farbmacher, Burgess, Windmeijer, and Smith}{Davies et al.}{2015}]{davies2015}
Davies, N. M., S. von Hinke Kessler Scholder, H. Farbmacher, S. Burgess, F. Windmeijer, and G. D. Smith (2015).
\newblock The many weak instruments problem and mendelian randomization.
\newblock {\em Statistics in medicine\/} {\em 34\/}(3), 454--468.
\bibitem[\citeauthoryear{Deryugina, Heutel, Miller, Molitor, and Reif}{Deryugina et al.}{2019}]{deryugina2019mortality}
Deryugina, T., G. Heutel, N. H. Miller, D. Molitor, and J. Reif (2019).
\newblock The mortality and medical costs of air pollution: Evidence from changes in wind direction.
\newblock {\em American Economic Review\/} {\em 109\/}(12), 4178--4219.
\bibitem[\citeauthoryear{Dobbie, Goldin, and Yang}{Dobbie et al.}{2018}]{dobbie2018}
Dobbie, W., J. Goldin, and C. S. Yang (2018).
\newblock The effects of pretrial detention on conviction, future crime, and employment: Evidence from randomly assigned judges.
\newblock {\em American Economic Review\/} {\em 108\/}(2), 201--40.
\bibitem[\citeauthoryear{Dobriban and Wager}{Dobriban and Wager}{2018}]{dobriban2018high}
Dobriban, E. and S. Wager (2018).
\newblock High-dimensional asymptotics of prediction: Ridge regression and classification.
\newblock {\em The Annals of Statistics\/} {\em 46\/}(1), 247--279.
\bibitem[\citeauthoryear{Donald and Newey}{Donald and Newey}{2001}]{donald2001}
Donald, S. G. and W. K. Newey (2001).
\newblock Choosing the number of instruments.
\newblock {\em Econometrica\/} {\em 69\/}(5), 1161--1191.
\bibitem[\citeauthoryear{Dov{\`\i}, Kock, and Mavroeidis}{Dov{\`\i} et al.}{2024}]{DKM24}
Dov{\`\i}, M.-S., A. B. Kock, and S. Mavroeidis (2024).
\newblock A ridge-regularized jackknifed anderson-rubin test.
\newblock {\em Journal of Business & Economic Statistics\/}, 1--12.
\bibitem[\citeauthoryear{Dufour and Taamouti}{Dufour and Taamouti}{2005}]{Dufour-Taamouti(2005)}
Dufour, J.-M. and M. Taamouti (2005).
\newblock {Projection}-based statistical inference in linear structural models with possibly weak instruments.
\newblock {\em Econometrica\/} {\em 73\/}(4), 1351--1365.
\bibitem[\citeauthoryear{Evdokimov and Koles{\'a}r}{Evdokimov and Koles{\'a}r}{2018}]{EK18}
Evdokimov, K. S. and M. Koles{\'a}r (2018).
\newblock Inference in instrumental variables analysis with heterogeneous treatment e ects.
\bibitem[\citeauthoryear{Fama and MacBeth}{Fama and MacBeth}{1973}]{fama1973}
Fama, E. F. and J. D. MacBeth (1973).
\newblock Risk, return, and equilibrium: Empirical tests.
\newblock {\em Journal of Political Economy\/} {\em 81\/}(3), 607--636.
\bibitem[\citeauthoryear{Farbmacher, Groh, M{\"u}hlegger, and Vollert}{Farbmacher et al.}{2024}]{farbmacher2024revisiting}
Farbmacher, H., R. Groh, M. M{\"u}hlegger, and G. Vollert (2024).
\newblock Revisiting the many instruments problem using random matrix theory.
\newblock {\em arXiv preprint arXiv:2408.08580\/}.
\bibitem[\citeauthoryear{Feng, Jaidee, Pan, and Zhu}{Feng et al.}{2024}]{Feng2024}
Feng, Q., S. Jaidee, G. Pan, and W. Zhu (2024).
\newblock Robust testing in high dimensional linear models.
\newblock {\em archive\/}.
\bibitem[\citeauthoryear{Fuller}{Fuller}{1977}]{Fuller(1977)}
Fuller, W. A. (1977).
\newblock Some properties of a modification of the limited information estimator.
\newblock {\em Econometrica\/} {\em 45\/}(4), 939--953.
\bibitem[\citeauthoryear{Gabaix and Koijen}{Gabaix and Koijen}{2024}]{gabaix2024}
Gabaix, X. and R. S. Koijen (2024).
\newblock Granular instrumental variables.
\newblock {\em Journal of Political Economy\/} {\em 132\/}(7), 000--000.
\bibitem[\citeauthoryear{Goldsmith-Pinkham, Sorkin, and Swift}{Goldsmith-Pinkham et al.}{2020}]{Goldsmith(2020)}
Goldsmith-Pinkham, P., I. Sorkin, and H. Swift (2020).
\newblock Bartik instruments: What, when, why, and how.
\newblock {\em American Economic Review\/} {\em 110\/}(8), 2586--2624.
\bibitem[\citeauthoryear{Guggenberger, Kleibergen, and Mavroeidis}{Guggenberger et al.}{2019}]{GKM(2019)}
Guggenberger, P., F. Kleibergen, and S. Mavroeidis (2019).
\newblock A more powerful subvector anderson rubin test in linear instrumental variables regression.
\newblock {\em Quantitative Economics\/} {\em 10\/}(2), 487--526.
\bibitem[\citeauthoryear{Guggenberger, Kleibergen, and Mavroeidis}{Guggenberger et al.}{2021}]{GKM(2021)}
Guggenberger, P., F. Kleibergen, and S. Mavroeidis (2021).
\newblock A powerful subvector anderson rubin test in linear instrumental variables regression with conditional heteroskedasticity.
\newblock {\em arXiv preprint arXiv:2103.11371\/}.
\bibitem[\citeauthoryear{Guggenberger, Kleibergen, Mavroeidis, and Chen}{Guggenberger et al.}{2012}]{GKMC(2012)}
Guggenberger, P., F. Kleibergen, S. Mavroeidis, and L. Chen (2012).
\newblock \hypertarget{GKMC(2012)}{On the asymptotic sizes of subset Anderson--Rubin and Lagrange multiplier tests in linear instrumental variables regression}.
\newblock {\em Econometrica\/} {\em 80\/}(6), 2649--2666.
\bibitem[\citeauthoryear{Han and Phillips}{Han and Phillips}{2006}]{han2006}
Han, C. and P. C. Phillips (2006).
\newblock Gmm with many moment conditions.
\newblock {\em Econometrica\/} {\em 74\/}(1), 147--192.
\bibitem[\citeauthoryear{Hansen, Hausman, and Newey}{Hansen et al.}{2008}]{Hansen-Hausman-Newey(2008)}
Hansen, C., J. Hausman, and W. Newey (2008).
\newblock Estimation with many instrumental variables.
\newblock {\em 26\/}(4), 398--422.
\bibitem[\citeauthoryear{Hansen and Kozbur}{Hansen and Kozbur}{2014}]{hansen2014}
Hansen, C. and D. Kozbur (2014).
\newblock Instrumental variables estimation with many weak instruments using regularized jive.
\newblock {\em Journal of Econometrics\/} {\em 182\/}(2), 290--308.
\bibitem[\citeauthoryear{Hastie, Montanari, Rosset, and Tibshirani}{Hastie et al.}{2022}]{hastie2022surprises}
Hastie, T., A. Montanari, S. Rosset, and R. J. Tibshirani (2022).
\newblock Surprises in high-dimensional ridgeless least squares interpolation.
\newblock {\em Annals of statistics\/} {\em 50\/}(2), 949.
\bibitem[\citeauthoryear{Hausman, Newey, Woutersen, Chao, and Swanson}{Hausman et al.}{2012}]{Haus2012}
Hausman, J. A., W. K. Newey, T. Woutersen, J. C. Chao, and N. R. Swanson (2012).
\newblock Instrumental variable estimation with heteroskedasticity and many instruments.
\newblock {\em Quantitative Economics\/} {\em 3\/}(2), 211--255.
\bibitem[\citeauthoryear{Kaffo and Wang}{Kaffo and Wang}{2017}]{Kaffo-Wang(2017)}
Kaffo, M. and W. Wang (2017).
\newblock \hypertarget{Kaffo-Wang(2017)}{On bootstrap validity for specification testing with many weak instruments}.
\newblock {\em Economics Letters\/} {\em 157}, 107--111.
\bibitem[\citeauthoryear{Karoui}{Karoui}{2008}]{Karoui2008}
Karoui, N. E. (2008).
\newblock {Spectrum estimation for large dimensional covariance matrices using random matrix theory}.
\newblock {\em The Annals of Statistics\/} {\em 36}, 2757--2790.
\bibitem[\citeauthoryear{Khan, Lan, Tamer, and Yao}{Khan et al.}{2024}]{Khan2024}
Khan, S., X. Lan, E. Tamer, and Q. Yao (2024).
\newblock Estimating high dimensional monotone index models by iterative convex optimization.
\newblock {\em Journal of Econometrics\/}, 105901.
\bibitem[\citeauthoryear{Kim and Sun}{Kim and Sun}{2011}]{kim2011spatial}
Kim, M. S. and Y. Sun (2011).
\newblock Spatial heteroskedasticity and autocorrelation consistent estimation of covariance matrix.
\newblock {\em Journal of Econometrics\/} {\em 160\/}(2), 349--371.
\bibitem[\citeauthoryear{Kleibergen}{Kleibergen}{2002}]{Kleibergen(2002)}
Kleibergen, F. (2002).
\newblock \hypertarget{Kleibergen(2002)}{Pivotal statistics for testing structural parameters in instrumental variables regression}.
\newblock {\em Econometrica\/} {\em 70\/}(5), 1781--1803.
\bibitem[\citeauthoryear{Kojevnikov, Marmer, and Song}{Kojevnikov et al.}{2021}]{Kojevnikov2021}
Kojevnikov, D., V. Marmer, and K. Song (2021).
\newblock Limit theorems for network dependent random variables.
\newblock {\em Journal of Econometrics\/} {\em 222}, 882--908.
\bibitem[\citeauthoryear{Koles{\'a}r}{Koles{\'a}r}{2013}]{K13}
Koles{\'a}r, M. (2013).
\newblock Estimation in instrumental variables models with heterogeneous treatment effects.
\newblock {\em Working Paper\/}.
\bibitem[\citeauthoryear{Koles{\'a}r}{Koles{\'a}r}{2018}]{kolesar2018}
Koles{\'a}r, M. (2018).
\newblock Minimum distance approach to inference with many instruments.
\newblock {\em Journal of Econometrics\/} {\em 204\/}(1), 86--100.
\bibitem[\citeauthoryear{Kuersteiner and Okui}{Kuersteiner and Okui}{2010}]{kuersteiner2010}
Kuersteiner, G. and R. Okui (2010).
\newblock Constructing optimal instruments by first-stage prediction averaging.
\newblock {\em Econometrica\/} {\em 78\/}(2), 697--718.
\bibitem[\citeauthoryear{Kunitomo}{Kunitomo}{1980}]{Kunitomo1980}
Kunitomo, N. (1980).
\newblock Asymptotic expansions of the distributions of estimators in a linear functional relationship and simultaneous equations.
\newblock {\em Journal of the American Statistical Association\/} {\em 75\/}(371), 693--700.
\bibitem[\citeauthoryear{Ledoit and Wolf}{Ledoit and Wolf}{2002}]{Ledoit2002}
Ledoit, O. and M. Wolf (2002).
\newblock Some hypothesis tests for the covariance matrix when the dimension is large compared to the sample size.
\newblock {\em The Annals of Statistics\/} {\em 30}, 1081--1102.
\bibitem[\citeauthoryear{Lee, McCrary, Moreira, and Porter}{Lee et al.}{2022}]{lee2021}
Lee, D. S., J. McCrary, M. J. Moreira, and J. R. Porter (2022).
\newblock Valid t-ratio inference for iv.
\newblock {\em American Economic Review\/} {\em 112\/}(10), 3260--90.
\bibitem[\citeauthoryear{Li and Chen}{Li and Chen}{2012}]{Li2012}
Li, J. and S. X. Chen (2012).
\newblock {Two sample tests for high-dimensional covariance matrices}.
\newblock {\em The Annals of Statistics\/} {\em 40}, 908--940.
\bibitem[\citeauthoryear{Li, Zhang, Cai, and Li}{Li et al.}{2024}]{Li2024}
Li, S., L. Zhang, T. T. Cai, and H. Li (2024).
\newblock Estimation and inference for high-dimensional generalized linear models with knowledge transfer.
\newblock {\em Journal of the American Statistical Association\/} {\em 119}, 1274--1285.
\bibitem[\citeauthoryear{Ligtenberg}{Ligtenberg}{2023}]{ligtenberg2023}
Ligtenberg, J. W. (2023).
\newblock Inference in iv models with clustered dependence, many instruments and weak identification.
\newblock {\em arXiv preprint arXiv:2306.08559\/}.
\bibitem[\citeauthoryear{Lim, Wang, and Zhang}{Lim et al.}{2024a}]{Lim2024}
Lim, D., W. Wang, and Y. Zhang (2024a).
\newblock A conditional linear combination test with many weak instruments.
\newblock {\em Journal of Econometrics\/} {\em 238\/}(2), 105602.
\bibitem[\citeauthoryear{Lim, Wang, and Zhang}{Lim et al.}{2024b}]{LWZ2024dimension}
Lim, D., W. Wang, and Y. Zhang (2024b).
\newblock A dimension-agnostic bootstrap anderson-rubin test for instrumental variable regressions.
\newblock {\em arXiv preprint arXiv:2412.01603\/}.
\bibitem[\citeauthoryear{Maestas, Mullen, and Strand}{Maestas et al.}{2013}]{maestas2013}
Maestas, N., K. J. Mullen, and A. Strand (2013).
\newblock Does disability insurance receipt discourage work? using examiner assignment to estimate causal effects of ssdi receipt.
\newblock {\em American Economic Review\/} {\em 103\/}(5), 1797--1829.
\bibitem[\citeauthoryear{Matsushita and Otsu}{Matsushita and Otsu}{2022}]{Matsushita2020}
Matsushita, Y. and T. Otsu (2022).
\newblock Jackknife lagrange multiplier test with many weak instruments.
\newblock {\em Econometric Theory\/} {\em forthcoming}.
\bibitem[\citeauthoryear{Maurice J. G. Bun and Poldermans}{Maurice J. G. Bun and Poldermans}{2020}]{Bun-Farbmacher-Poldermans(2020)}
Maurice J. G. Bun, H. F. and R. W. Poldermans (2020).
\newblock Finite sample properties of the gmm anderson-rubin test.
\newblock {\em Econometric Reviews\/} {\em 39\/}(10), 1042--1056.
\bibitem[\citeauthoryear{Mikusheva and Sun}{Mikusheva and Sun}{2022}]{MS22}
Mikusheva, A. and L. Sun (2022).
\newblock Inference with many weak instruments.
\newblock {\em Review of Economic Studies\/} {\em 89\/}(5), 2663--2686.
\bibitem[\citeauthoryear{Morimune}{Morimune}{1983}]{morimune1983}
Morimune, K. (1983).
\newblock Approximate distributions of k-class estimators when the degree of overidentifiability is large compared with the sample size.
\newblock {\em Econometrica\/} {\em 51\/}(3), 821--841.
\bibitem[\citeauthoryear{Navjeevan}{Navjeevan}{2023}]{N23}
Navjeevan, M. (2023).
\newblock An identification and dimensionality robust test for instrumental variables models.
\newblock {\em arXiv preprint arXiv:2311.14892\/}.
\bibitem[\citeauthoryear{Newey and Windmeijer}{Newey and Windmeijer}{2009}]{Newey-Windmeijer(2009)}
Newey, W. K. and F. Windmeijer (2009).
\newblock \hypertarget{Newey-Windmeijer(2009)}{Generalized method of moments with many weak moment conditions}.
\newblock {\em Econometrica\/} {\em 77\/}(3), 687--719.
\bibitem[\citeauthoryear{Okui}{Okui}{2011}]{okui2011}
Okui, R. (2011).
\newblock Instrumental variable estimation in the presence of many moment conditions.
\newblock {\em Journal of Econometrics\/} {\em 165\/}(1), 70--86.
\bibitem[\citeauthoryear{Onatski, Moreira, and Hallin}{Onatski et al.}{2013}]{Onatski2013}
Onatski, A., M. J. Moreira, and M. Hallin (2013).
\newblock Asymptotic power of sphericity tests for high-dimensional data.
\newblock {\em The Annals of Statistics\/} {\em 41}, 1204--1231.
\bibitem[\citeauthoryear{Sampat and Williams}{Sampat and Williams}{2019}]{sampat2019}
Sampat, B. and H. L. Williams (2019).
\newblock How do patents affect follow-on innovation? evidence from the human genome.
\newblock {\em American Economic Review\/} {\em 109\/}(1), 203--36.
\bibitem[\citeauthoryear{Shanken}{Shanken}{1992}]{shanken1992}
Shanken, J. (1992).
\newblock On the estimation of beta-pricing models.
\newblock {\em The Review of Financial Studies\/} {\em 5\/}(1), 1--33.
\bibitem[\citeauthoryear{S{\l}oczy{\'n}ski}{S{\l}oczy{\'n}ski}{2024}]{sloczynski2024should}
S{\l}oczy{\'n}ski, T. (2024).
\newblock When should we (not) interpret linear iv estimands as late?
\newblock {\em arXiv preprint arXiv:2011.06695\/}.
\bibitem[\citeauthoryear{S{\o}lvsten}{S{\o}lvsten}{2020}]{solvsten2020}
S{\o}lvsten, M. (2020).
\newblock Robust estimation with many instruments.
\newblock {\em Journal of Econometrics\/} {\em 214\/}(2), 495--512.
\bibitem[\citeauthoryear{Stock and Yogo}{Stock and Yogo}{2005}]{stock2005}
Stock, J. and M. Yogo (2005).
\newblock Asymptotic distributions of instrumental variables statistics with many instruments.
\newblock {\em Identification and inference for econometric models: Essays in honor of Thomas Rothenberg\/} {\em 6}, 109--120.
\bibitem[\citeauthoryear{Tuvaandorj}{Tuvaandorj}{2024}]{tuvaandorj2024}
Tuvaandorj, P. (2024).
\newblock Robust permutation tests in linear instrumental variables regression.
\newblock {\em Journal of the American Statistical Association\/} (forthcoming), 1--24.
\bibitem[\citeauthoryear{Wang}{Wang}{2020}]{Wang(2020)}
Wang, W. (2020).
\newblock \hypertarget{Wang(2020)}{On the inconsistency of nonparametric bootstraps for the subvector Anderson-Rubin test}.
\newblock {\em Economics Letters\/}, 109157.
\bibitem[\citeauthoryear{Wang and Doko Tchatoka}{Wang and Doko Tchatoka}{2018}]{Wang-Doko(2018)}
Wang, W. and F. Doko Tchatoka (2018).
\newblock \hypertarget{Wang-Doko(2018)}{On Bootstrap inconsistency and Bonferroni-based size-correction for the subset Anderson--Rubin test under conditional homoskedasticity}.
\newblock {\em Journal of Econometrics\/} {\em 207\/}(1), 188--211.
\bibitem[\citeauthoryear{Wang and Kaffo}{Wang and Kaffo}{2016}]{Wang_Kaffo_2016}
Wang, W. and M. Kaffo (2016).
\newblock Bootstrap inference for instrumental variable models with many weak instruments.
\newblock {\em Journal of Econometrics\/} {\em 192\/}(1), 231--268.
\bibitem[\citeauthoryear{Wang and Zhang}{Wang and Zhang}{2024}]{Wang-Zhang2024}
Wang, W. and Y. Zhang (2024).
\newblock Wild bootstrap inference for instrumental variables regressions with weak and few clusters.
\newblock {\em Journal of Econometrics\/} {\em 241\/}(1), 105727.
\bibitem[\citeauthoryear{Yap}{Yap}{2024}]{Yap24}
Yap, L. (2024).
\newblock Inference with many weak instruments and heterogeneity.
\newblock {\em arXiv preprint arXiv:2408.11193\/}.
\bibitem[\citeauthoryear{Zhang, Ding, Zhou, and Wang}{Zhang et al.}{2025}]{Zhang2025}
Zhang, Z., P. Ding, W. Zhou, and H. Wang (2025).
\newblock With random regressors, least squares inference is robust to correlated errors with unknown correlation structure.
\newblock {\em Biometrika\/} {\em 112}, asae054.