Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
89,412 characters · 14 sections · 106 citation commands
An Improved Inference for IV Regressions
Keywords: Many Weak Instruments, Shift-Share Instruments, Combination Test.
JEL Classification: C12, C36, C55
\setcounter{oldtocdepth}{\value{tocdepth}} \addtocontents{toc}{\setcounter{tocdepth}{-10}}
Empirical applications of instrumental variables (IV) regressions in economics often involve multiple sets of candidate instruments, some with dimensionality proportional to the sample size and others more parsimonious. A canonical example is the influential study by Angrist-Krueger(1991), which estimates the causal effect of schooling on wages using three quarter-of-birth (QoB) dummies as instruments, as well as their interactions with state- and year-of-birth dummies, yielding a total of 1,530 instruments. A different but related class of examples arises in the context of shift-share IVs,\footnote{See, for instance, bartik1991, blanchard1992, adao2019shift, Goldsmith(2020), borusyak2022quasi, borusyak2025practical, and references therein.} which, as noted by Goldsmith(2020), can be interpreted as a particular way of aggregating many underlying base instruments. It is common practice to report results based on a one-dimensional shift-share IV and then supplement them with results obtained using the full set of base IVs.
This naturally raises the question of whether one can take the low-dimensional IV regression as the benchmark specification and combine it with its many-IV counterpart in a way that enhances the power of IV inference, for free, in the sense that no additional identification restrictions are imposed on the many-IV specification. In this paper, we provide a constructive solution.
Specifically, we consider an optimal combined inference procedure based on three commonly reported test statistics for IV regressions in a clustered setting, where the data consist of many clusters of bounded size. The three component statistics are: (i) a standard cluster-robust Wald statistic constructed from the low-dimensional IVs; (ii) a leave-one-cluster-out jackknife Lagrange multiplier (LM) statistic; and (iii) a leave-one-cluster-out jackknife Anderson--Rubin (AR) statistic. The LM and AR statistics are constructed using the many-IV specification, and the leave-one-cluster-out design removes the many-IV bias in the presence of within-cluster error dependence.
We first show that, under the null hypothesis and local alternatives, the Wald, LM, and AR statistics are jointly asymptotically normal, provided that the low-dimensional IVs strongly identify the parameter of interest, while allowing the many-IV specification to be weakly identified. Standard optimal testing theory then implies that, in the corresponding Gaussian limit experiment, the uniformly most powerful unbiased (UMPU) test rejects for large absolute values of an appropriate linear combination of the three limiting Gaussian statistics.\footnote{See, for instance, Section 4.2 of Lehmann-Romano(2006) and Lemma 2.2 of LWZ24 for formal arguments.} Our proposed test, as a function of the three statistics, is asymptotically equivalent to this UMPU test and is therefore weakly more powerful than each of the Wald, LM, and AR tests.
Importantly, the combination test adapts automatically to the identification strength of the many-IV specification. When the parameter of interest is weakly identified under many IVs in the sense of MS22, the optimal linear combination places asymptotically negligible weight on the LM and AR components, so that the resulting test reduces to the Wald test. In this case, the procedure remains valid and retains the same asymptotic power as the standalone Wald test.
The notion of optimality in our study is defined relative to the class of tests based on the Wald, LM, and AR statistics and serves primarily as a device for improving inference within this class. We do not attempt to derive a globally optimal test by searching over all possible combinations of a given set of base IVs, nor do we aggregate results across alternative specifications that employ different sets of base IVs or different low-dimensional IVs. Rather, our objective is to take the low-dimensional specification as the benchmark and combine its associated test statistic with those constructed from a given set of many IVs, thereby strengthening inference while remaining agnostic about the identification strength of the many-IV specification. Nonetheless, we attribute a notion of asymptotic optimality to the resulting combination test in the sense of mueller2011.
The confidence interval implied by the combination test has the usual “estimator plus and minus a standard error times a critical value" form. Its center is an estimator that linearly combines a standard GMM estimator based on the low-dimensional IVs with a leave-one-cluster-out jackknife IV estimator based on many IVs and the AR statistic, using weights that capture both the relative identification strength from the two IV sets and the UMPU weights. The efficiency gain manifests itself as an almost surely shorter confidence interval. We measure this gain by the percentage reduction in the length of the resulting confidence interval relative to that of the Wald test based on the low-dimensional IVs. As an illustration, in the immigrant enclave application of card(2009) (card(2009)), the combination procedure reduces the length by between 5% and 17%. This reduction depends mainly on the identification strengths of the low-dimensional and many IVs, together with the limiting correlations between the Wald and LM statistics, and between the LM and AR statistics. In Section (ref), we translate these relationships into a practical rule of thumb based solely on the ratio of the standard errors of the low-dimensional and many-IV estimators, which can be directly read from reported regression tables prior to implementing our combination procedure. We illustrate this rule using the application in card(2009).
Relation to the literature. This paper contributes to the large literature on many (weak) instruments,\footnote{See, for instance, Kunitomo1980, morimune1983, Bekker(1994), donald2001, Chao-Swanson(2005), stock2005, han2006, Andrews-Stock(2007), Hansen-Hausman-Newey(2008), Newey-Windmeijer(2009), anderson2010, kuersteiner2010, anatolyev2011, okui2011, belloni2012, carrasco2012, Chao(2012), Haus2012, K13, hansen2014, carrasco2015, Wang_Kaffo_2016, kolesar2018, EK18, solvsten2020, CNT23, LWZ24, BN24, Yap24, lim2024dimension, among others.} and is particularly related to LWZ24. Building on Andrews(2016), they propose a jackknife conditional linear combination (CLC) test for independent data that is robust to weak identification, many instruments, and heteroskedasticity, combining jackknife AR, LM, and orthogonalized LM statistics to achieve good power across identification regimes. Their analysis, however, is confined to many-IV settings and does not directly apply to our framework. In contrast, we explicitly incorporate low-dimensional IVs within a clustered environment and develop a unified procedure that efficiently combines estimation and inference from both low- and high-dimensional IV specifications. Our primary objective is to construct a theoretically grounded combination test that improves upon the conventional Wald test based solely on low-dimensional IVs. We also note that, in a different context, jiang2025adjustments proposes an optimal linear combination of adjusted and unadjusted estimators for the average treatment effect under covariate-adaptive randomization.
Our paper is also related to the literature on inference with many IVs and clustered data. FLM23 study leave-one-cluster-out jackknife versions of the IV estimator, and ligtenberg2023inference develop weak-identification-robust procedures in clustered environments. We depart from this literature by not focusing solely on estimation or inference within the many-IV specification itself. Instead, we leverage the many-IV specification to enhance the power of low-dimensional Wald inference through an optimal combination of test statistics. CNT23 and kolesar2026cluster consider settings with both many IVs and many control variables, which introduce additional challenges for cluster-robust inference. In this paper, we restrict attention to models with a fixed number of control variables and leave the case with many IVs and many controls for future research.
Structure of the paper: Section (ref) details the rationale behind the proposed rule of thumb and demonstrates its use in an empirical application; practitioners mainly interested in applications may focus on this section directly. Section (ref) introduces the model and key preliminaries, while Section (ref) develops the large-sample theory for the combination test and formalizes its theoretical properties. Section (ref) presents the practical implementation of the combination test in an empirically relevant setting and employs simulations to evaluate its power properties, and Section (ref) concludes. An additional case with weak low-dimensional IVs alongside strong many IVs, as well as all proofs, is presented in the Online Appendices.
Notation. We write $[n] \equiv \{1,\ldots,n\}$ and $[G] \equiv \{1,\ldots,G\}$. Let $A$ be an $n\times m$ matrix and let $\{n_g\}_{g\in [G]}$ be positive integers with $\sum_{g=1}^G n_g = n$. We denote by $A_{[g]}$ the $g$-th row-wise block of $A$, of dimension $n_g \times m$. When $A$ is an $n \times n$ square matrix, we denote by $A_{[g,h]}$ the $(g,h)$-th block of $A$. For a positive semi-definite square matrix $A$, denote its largest and smallest eigenvalues by $\lambda_{\max}(A)$ and $\lambda_{\min}(A)$, respectively. Let $C$ be a generic positive constant independent of $n$, whose value may change from line to line. For brevity, we write $\sum_{g,h \in [G]^2, g\neq h} := \sum_{g \in [G]}\sum_{h \in [G], h \neq g}$.
In this section, we develop a practical rule of thumb that can be directly applied to reported estimates and standard errors from regressions using low-dimensional IVs and from regressions employing many IVs. We illustrate its empirical relevance using the application in Goldsmith(2020), which builds on card(2009).
We quantify the efficiency gain from the combination test by the percentage reduction in the length of its confidence interval, relative to that of the Wald test, in large samples. Section (ref) formally derives an analytical expression for this reduction as a function of the identification strengths of both the low-dimensional and many instrumental variables (IVs), as well as the limiting correlations between the Wald and LM statistics and between the LM and AR statistics. Because the percentage reduction is monotonically increasing in the absolute correlation between the LM and AR statistics, we further obtain a lower bound on this reduction by setting this correlation to zero. This yields a lower bound on the efficiency gain that depends only on the correlation between the Wald and LM statistics (denoted as $\rho_1$) and on the ratio of the standard deviations of the GMM estimator based on low-dimensional IVs to the leave-one-cluster-out jackknife IV estimator (JIVE) based on many IVs (or, equivalently, on the relative strength of the many IVs to the low-dimensional IVs).
Figure (ref) plots the lower bound as a function of the standard deviation ratio for different values of $\rho_1$, the limiting correlation between the Wald and LM statistics. Two observations emerge. First, for any fixed $\rho_1$, we show theoretically that whenever the standard deviation ratio exceeds $\rho_1$, the lower bound on efficiency gains increases with the ratio, reflecting the fact that the combination test exploits the additional precision provided by the LM statistic. Second, once the standard deviation ratio exceeds one, the lower bound decreases as $\rho_1$ increases, because the LM statistic then contributes relatively little additional information beyond the highly correlated Wald statistic. As a simple rule of thumb that only requires a back of envelope calculation based on the reported standard errors, we propose: for empirically plausible values of $\rho_1$ (between $-0.7$ and $0.7$), whenever the standard error from the regression with low-dimensional IVs divided by that from the regression with many IVs is greater than $1.05$, the corresponding confidence interval is reduced by at least $10\%$.\footnote{In additional (unreported) plots for $\rho_1 \in [-0.99, 0.99]$, the confidence interval is at least $10\%$ shorter whenever the standard deviation ratio exceeds $1.11$ (with thresholds $1.05$ for $5\%$ and $1.25$ for $20\%$).} We further argue in Section (ref) that the same rule-of-thumb can be applied to other many-IV estimators such as HLIML and HFUL developed by Haus2012.
To illustrate the empirical relevance of efficiency gains and the associated rule of thumb, we implement the combination test in an empirical application that estimates the (negative) inverse elasticity of substitution between immigrants and natives, following card(2009). As in Goldsmith(2020), we examine two separate sets of results by skill group: high school equivalent workers and college equivalent workers. The analysis is based on cross-sectional regressions for each skill group in 124 cities in 2000. The dependent variable is the residual log wage gap between immigrant and native men, and the main regressor of interest is the log ratio of immigrant-to-native hours for both men and women within the same skill group. Because a positive labor demand shock to immigrants can simultaneously increase their earnings and labor supply relative to natives, this introduces potential endogeneity. To construct the one-dimensional Bartik instrument (i.e., shift-share instrument), immigration shares from 38 origin countries (groups) in $1980$ are used as the base instruments, and the final instrument is formed as a weighted average of these country-specific shares, where the weights are given by the number of arrivals to the United States between 1990 and 2000 by origin country group and skill group (see Goldsmith(2020) for further details). As argued in card(2009), the rationale for the IV is that existing immigrant enclaves are likely to attract additional immigrant labor through social and cultural channels unrelated to labor market outcomes.
Table (ref) reports the point estimates obtained from the two-stage least squares (TSLS) estimator using the Bartik instrument, $\hat \beta_1$, along with those from the leave-one-cluster-out estimator, $\hat \beta_2$, which relies on all the 38 base IVs. We present results separately for specifications that include and exclude city-level controls. As expected and consistent with the findings in Goldsmith(2020), the estimates are broadly similar within each skill group. However, in every specification, the confidence intervals constructed from $\hat \beta_1$ (Wald CI) differ to some extent from those based on $\hat \beta_2$ (LM CI). This discrepancy is partly due to the different standard errors of the two estimators, which we exploit in our combination test to obtain strictly shorter confidence intervals across all specifications. For example, for workers with college equivalent skills, our confidence intervals are roughly $7\%$ and $17\%$ shorter, respectively, than the Wald CIs in the specifications with and without controls.
Figure (ref) displays the realized percentage reduction in confidence interval length achieved by our combination test, together with the asymptotic lower bounds for that reduction, analogous to Figure (ref), but calculated using the specification-specific estimate $\hat \rho_1$ of the limiting correlation between the Wald and LM statistics. Two observations are worth noting. First, up to finite-sample estimation error and across different specifications, the actual percentage reductions generally fall near or above the corresponding lower bounds, in line with our theoretical predictions. Note also that these bounds are derived solely from the Wald and LM statistics, indicating that, in this empirical setting, incorporating the AR statistic yields little additional efficiency gain. This aligns with our earlier theoretical discussion, which posits that efficiency gains increase monotonically with the absolute correlation between LM and AR statistics, and, in fact, the corresponding consistent estimates $\hat\rho_2$ are not particularly large in almost every specification.
Second, Figure (ref) clearly illustrates our proposed rule of thumb. All estimated $\hat \rho_1$ values fall within $[-0.7, 0.7]$. When the standard error ratio exceeds $1.05$, the actual reduction in the length of the confidence interval is $17\%$ (specification without controls for college equivalent workers), much above the rule-of-thumb benchmark $10\%$ for its lower bound. In contrast, when the standard error ratio remains at or below $1.05$, the improvements are modest, reflecting the converse of our rule of thumb. Nevertheless, even in these cases, our combination test can still deliver confidence intervals that are about $7\%$ (specification with controls for college equivalent workers), $6\%$ (specification without controls for high school equivalent workers), and $5\%$ (specification with controls for high school equivalent workers) shorter. In Supplementary Appendix F, we further consider the return to education application of Angrist-Krueger(1991), which illustrates even more notable efficiency gains provided by our approach.
We consider a clustered dataset with $G$ clusters and denote the size of the $g$-th cluster as $n_g$ for $g \in [G]$. We index observations by units followed by clusters. Denote $I_g = [N_{g-1}+1,\cdots,N_{g}]$, where $N_0 = 0$, $N_g = \sum_{g'=0}^g n_{g'}$, and $N_G =n$. Then, $\{I_g\}_{g \in [G]}$ forms a partition of $[n]$, and if $i \in I_g$, this means that the $i$-th observation belongs to the $g$-th cluster. We then consider a linear IV regression with clustered data:
where we denote $\tilde Y_{i,g} \in \Re$, $\tilde X_{i,g} \in \Re$, and $W_{i,g} \in \Re^{d_w}$ as an outcome variable, an endogenous regressor, and exogenous regressors, respectively. Further denote $\tilde Z_{i,g} \in \Re^{K}$ as the IVs for $\tilde X_{i,g}$. The first-stage equation can be written as
where $\tilde \Pi_{i,g} = \mathbb E (\tilde X_{i,g}|\{\tilde Z_{j,g},W_{j,g}\}_{j \in I_g})$ is not assumed to be linear in $\tilde Z_{i,g}$ and $W_{i,g}$. We assume that $\mathbb E \tilde e_{i,g} =0$ and $\mathbb E \tilde V_{i,g} = 0$, and $\left\{ \tilde e_{i,g}, \tilde V_{i,g} \right\}_{i \in I_g, g \in [G]}$ are independent between clusters, but allow them to have a general dependence structure within each cluster. Throughout the paper, the dimension $d_w$ of $W_{i,g}$ is assumed to be fixed. If researchers want to include cluster fixed effects in the model, they can obtain (ref) by first demeaning the data (outcome, endogenous regressor, controls, and instruments) at the cluster level. We assume that $K$, the dimension of $\tilde Z_{i,g}$, diverges to infinity with the sample size.
Let $\tilde Y$, $\tilde X$, $\tilde \Pi$, $W$, $\tilde Z$ be $n \times 1$, $n \times 1$, $n\times 1$, $n \times d_w$, and $n \times K$-dimensional vectors and matrices formed by $\tilde Y_{i,g}$, $\tilde X_{i,g}$, $\tilde \Pi_{i,g}$, $W_{i,g}$, and $\tilde Z_{i,g}$, respectively. More specifically, $\tilde Y$ is constructed by stacking up $\tilde Y_{i,g}$ across $i \in I_g$ followed by $g \in [G]$, and similarly for $\tilde X$, $\tilde \Pi$, $W$ and $\tilde Z$. We then partial out $W$ from $\tilde Y$, $\tilde X$, and $\tilde Z$, so that the model in (ref)-(ref) can be written in a vector form as
where $Y = M_W \tilde Y$, $X = M_W \tilde X$, $\Pi = M_W \tilde \Pi$, $e = M_W \tilde e$, $V = M_W \tilde V$, $M_W = I_n - P_W$, $P_W = W(W^\top W)^{-1}W^\top$, and $I_n$ denotes an $n \times n$ identity matrix. We further denote $Z = M_W \tilde Z$.
In addition, besides the $K$-dimensional many IVs $\tilde Z_{i,g} \in \Re^K$, we assume that there is another set of low-dimensional IVs
where $\{f_{i,g}(\cdot)\}_{i \in I_g, g\in [G]}$ is a list of known nonstochastic functions of $d_z$ dimension. Specifically, as illustrated by the example of Angrist-Krueger(1991) in the Introduction, researchers may begin with certain low-dimensional base IVs $\tilde{z}_{i,g}$, such as the three QoB dummies, and construct a large number of new IVs by taking the interaction between $\tilde z_{i,g}$ and control variables $W_{i,g}$ (e.g., state- and year-of-birth dummies in Angrist-Krueger(1991)). Then, $\tilde{z}_{i,g}$ is a subset of the $K$-dimensional many IVs $\tilde Z_{i,g}$ for the model in (ref)-(ref), which include both the low-dimensional base IVs and interacted IVs. The second example of $\tilde z_{i,g}$ is the widely used shift-share IV. As pointed out by Goldsmith(2020), under their identification strategy that treats the shares as exogenous, the one-dimensional shift-share IV can be regarded as a weighted average of many base IVs. For instance, in the canonical setting of estimating the inverse elasticity of labor supply (e.g., see Sections I and VI of Goldsmith(2020)), the observations are typically clustered at the location level, such as the US commuting zone level, with a short panel dataset of $T$ time periods. The structural equation of interest can thus be written as
where $g$ indexes a location, $t$ a time period, $\tilde Y_{g,t}$ is wage growth, $\tilde X_{g,t}$ is employment growth, and $W_{g,t}$ is a vector of controls which could include location and time fixed effects. Then, according to our notation, $G$ is equal to the number of locations, and $n_g = T$ for all $g \in [G]$. The shift-share IV is an inner product of the initial industry-location shares and the industry-period growth rates, i.e., $\tilde z_{g,t} = \sum_{k=1}^K \tilde s_{0,g,k} h_{k,t}$, where the employment share of industry $k$ in the location $g$ in a certain initial period corresponds to the base IV $\tilde s_{0,g,k}$, the growth rate of industry $k$ in the period $t$ corresponds to the weight $h_{k,t}$, and $K$ is the number of industries.
Additionally, let $\tilde z$ be the $n \times d_z$-dimensional matrix formed by $\tilde z_{i,g}$, and denote $z = M_W \tilde z$. In many empirical applications, the dimension $d_z$ is just one (e.g., the shift-share IV), but our setup also allows for $d_z>1$, while maintaining the requirement that $d_z$ is fixed with respect to the sample size $n$.
The null and alternative hypotheses studied are $\mathcal{H}_0: \beta = \beta_0 \ \text{against}\ \mathcal{H}_{1}: \beta \neq \beta_0$. We focus on the model with a scalar endogenous variable for two reasons. First, in many empirical applications of IV regressions, there is only one endogenous variable.\footnote{For example, 101 out of 230 specifications in andrews2019weak's sample and 1,087 out of 1,359 in young2022consistency's sample feature one endogenous regressor and one IV. Similarly, lee2022valid find that 61 out of 123 IV papers published in AER between 2013 and 2019 use a single IV. Our setting further allows for the overidentified case with one endogenous regressor and multiple IVs. In general, empirical researchers can generate many IVs by using polynomials or interactions based on their low-dimensional base IVs and control variables, in the same spirit as Angrist-Krueger(1991). Then, it is possible to achieve efficiency improvement using our combination procedure.} Second, if we assume that at least the low-dimensional IVs provide strong identification, our results can be extended to testing of scalar restrictions with multiple endogenous variables by applying standard subvector inference methods, without appealing to a projection-based weak-identification-robust inference approach.\footnote{For weak-identification-robust subvector inference, in general, one may use a projection approach (Dufour-Taamouti(2005), Dufour-Taamouti(2005)) after implementing inference on the whole vector of endogenous variables. However, the projection approach typically leads to conservative inference. Alternative subvector inference methods for IV regressions (e.g., see GKMC(2012), Andrews(2017), GKM(2019), GKM(2021), and Wang-Doko(2018)) provide a power improvement over under a fixed number of instruments (some of these methods further require conditional homoskedasticity). However, whether they can be applied to the setting of many weak instruments is unclear.} In addition, our method could potentially be extended to the case in which the identification strength provided by many IVs is mixed for the endogenous variables, by following the approach of Chao(2012). We leave these extensions for future research.
Under our setting, it is possible to conduct inference directly based on the low-dimensional IVs. Specifically, given a $d_z \times d_z$ positive definite weighting matrix $\hat A_n$, the generalized method of moments (GMM) estimator can be written as
It is also possible to construct test statistics using the $K$-dimensional many IVs. Denote $P = Z(Z^\top Z)^{-1}Z^\top$ as the projection matrix of $Z$. Then, the leave-one-cluster-out jackknife IV estimator of $\beta$ is denoted as $\hat \beta_2$ and defined as
where $\bar P$ is the block diagonal matrix corresponding to $P$ such that the $g$-th block on its diagonal is $P_{[g,g]}$. Note that under independent data, $\hat \beta_2$ reduces to the JIVE estimator in angrist1999, Chao(2012), and MS22.
Given $\hat \beta_1$ and $\hat \beta_2$, we define the estimator $\hat \beta$ as
where $\dot \Phi_1$ and $\ddot \Phi_2$ are the variance estimators for $\hat \beta_1$ and $\hat \beta_2$, respectively, to be defined later in Section (ref). We show that \(\hat{\beta}\) is consistent whenever either the low-dimensional IV estimator \(\hat{\beta}_1\) or the many-IV estimator \(\hat{\beta}_2\) is consistent (i.e., either the low-dimensional IVs or the many IVs provide strong identification for $\beta$). Because researchers do not need to know which of the two estimators is consistent when constructing \(\hat{\beta}\), the estimator is doubly robust.
We then use the doubly robust estimator \(\hat{\beta}\) to re-estimate the variances associated with \(\hat{\beta}_1\) and \(X^{\top}(P-\bar P)e\), denoted by \(\widehat{\Phi}_1\) and \(\hat{\Sigma}\), respectively, and defined in Section (ref). These variance estimates are used to construct the Wald statistic (with low-dimensional IVs) and the leave-one-cluster-out jackknife LM statistic (with many IVs):
where \(e(\beta_0) = Y - X\beta_0\).
Lastly, as pointed out by Haus2012, LWZ24, and MS24, it is possible to use the jackknife AR statistic to further improve the efficiency of the jackknife LM statistic. In the current setting with clustered data, we define the leave-one-cluster-out jackknife AR statistic as
where $\hat \Upsilon$ is a consistent variance estimator for the numerator defined later, and $\hat e = Y -X \hat \beta$.
We use the consistent estimator $\hat \beta$ to construct our tests for two reasons. First, using the null value $\beta_0$ to form the AR statistic and then combine it with the LM statistic can produce a non-monotonic power curve (i.e., a power ditch against certain alternatives; see Andrews(2016) and LWZ24 for related discussions). Constructing the AR statistic with $\hat \beta$ avoids this issue. Consequently, the AR statistic in (ref) does not depend on $\beta_0$ but serves as a normalized estimator of zero, used solely to improve the efficiency of our procedure. Second, the combination test relies on correlations between $T(\beta_0)$ and $LM(\beta_0)$ and between $LM(\beta_0)$ and $AR$, whose consistent estimation requires a consistent $\beta$ estimator. This consistency is crucial to ensure proper size control, especially when the many-IV specification is weakly identified.
In the next section, we illustrate how to optimally combine the three test statistics \(T(\beta_0)\), \(LM(\beta_0)\), and \(AR\). The corresponding variance estimators \((\dot{\Phi}_1, \ddot{\Phi}_2, \widehat{\Phi}_1, \hat{\Sigma}, \hat{\Upsilon})\) are introduced later in Section (ref).
Given the three test statistics $(T(\beta_0), LM(\beta_0), AR)$, we seek to combine them in a theoretically justified way that can improve on the Wald test based only on the low-dimensional IVs. The key insight of our paper is that under certain local alternative $\beta - \beta_0 = \delta d_n$ with some deterministic sequence $d_n \downarrow 0$, we have the following joint limiting distribution:
for some $a_1$, $a_2$, $\rho_1$ and $\rho_2$ to be defined later. In this limiting problem, the UMPU level-$\alpha$ test for the default null hypothesis $\delta=0$ against two-sided alternatives, which are solely based on the limiting three-dimensional normal random vector, can be obtained by invoking standard hypothesis-testing results (see, for example, Section 4.2 of Lehmann-Romano(2006)), and is stated in Proposition (ref) below.
In light of this optimal testing result in the limiting problem, one may wish to propose implementing the following test:
and then investigate its asymptotic justification.
However, the parameters $a_1,a_2,\rho_1,\rho_2$ are usually unknown and need to be estimated. In addition, it turns out that the weights $(\omega_1,\omega_2,\omega_3)$ are invariant to the scale normalization of $(b_1,b_2,b_3)$, and thus, $(a_1,a_2)$. Therefore, to construct the UMPU test, it suffices to consistently estimate $\alpha_1 = a_1/\sqrt{a_1^2+a_2^2}$ and $\alpha_2 = a_2/\sqrt{a_1^2+a_2^2}$ along with $\rho_1$ and $\rho_2$.
Given the consistent estimators $(\hat \alpha_1,\hat \alpha_2,\hat \rho_1,\hat \rho_2)$ for $( \alpha_1, \alpha_2, \rho_1, \rho_2)$ specified in Section (ref), we then implement the feasible version of the combination test:
In this section, we investigate the asymptotic behavior of our combination test. We begin by stating and discussing general assumptions about the data-generating process and the identification strength of both the low-dimensional and many IVs. We then establish the asymptotic efficiency properties of the combination test and, finally, compare its efficiency to that of the conventional Wald test based solely on low-dimensional IVs via the limiting length ratio of their confidence intervals.
As in Chao(2012), we treat $\tilde Z$ and $W$ as fixed. This is equivalent to treating them as random and repeating all the analyses in the paper by conditioning on them. For the data-generating process, we impose the following assumptions.
For the low-dimensional IVs, we focus on the case of strong identification strength in the main text. We further investigate the case in which the low-dimensional IVs have weak identification strength in Supplementary Appendix A, and show that, under certain conditions, when the many IVs provide strong identification, our optimal combination test still asymptotically controls size.
With strong identification, the asymptotic variance of $\hat \beta_1$ is
where $\Omega = \mathbb E \sum_{g \in [G]} \left(z_{[g]}^\top \tilde e_{[g]}\right) \left(z_{[g]}^\top \tilde e_{[g]}\right)^\top$, $\acute \Pi = z A_n z^\top \Pi$ and $\Psi=\Pi^\top z A_n \Omega A_n z^\top \Pi$. The cluster-robust variance estimator $\widehat \Phi_1$ for the Wald statistic with low-dimensional IVs is
where $\hat \Omega = \sum_{g \in [G]} \left(z_{[g]}^\top \hat e_{[g]}\right) \left(z_{[g]}^\top \hat e_{[g]}\right)^\top$, $\acute X = z \hat A_n z^\top X$ and $\hat \Psi = X^\top z \hat A_n \hat \Omega \hat A_n z^\top X$. The initial estimator $\dot{\Phi}_1$ for $\Phi_1$ used in the computation of $\hat{\beta}$ in (ref) is defined in the same way as $\widehat{\Phi}_1$, except that $\hat e_{[g]} = Y_{[g]} - X_{[g]} \hat{\beta}$ is replaced by $\dot e_{[g]} = Y_{[g]} - X_{[g]} \hat{\beta}_1$.
We make the following assumptions regarding the inference with low-dimensional IVs.
If the many-IV-based identification is strong, similar to Chao(2012), we can show that the asymptotic variance of $\hat \beta_2$ is
where
is the asymptotic variance of $X^\top (P - \bar P) e$. A natural estimator for $\Phi_2$ is thus
where $\hat \Sigma$ is a consistent estimator of $\Sigma$. Such an estimator is proposed in Chao(2012) for the case with independent data. Here, in addition to extending to clustered data, we need to account for the fact that $W$ has already been partialled out, whereas in Chao(2012) the coefficients for $W$ are also estimated. Therefore, some adjustments are required, as in matsushita2024. For that purpose, define $Q = M_W (P - \bar P) M_W$, and let $\bar Q$ be the block diagonal matrix corresponding to $Q$ such that the $g$-th block on its diagonal is $Q_{[g,g]}$. Our variance estimator is similar to the one in Chao(2012) but with $P-\bar P$ replaced by $Q- \bar Q$, i.e.,
The initial estimator $\ddot{\Phi}_2$ for $\Phi_2$ used in the computation of $\hat{\beta}$ is defined in the same way as $\widehat{\Phi}_2$, except that $\hat e_{[g]} = Y_{[g]} - X_{[g]} \hat{\beta}$ is replaced by $\ddot e_{[g]} = Y_{[g]} - X_{[g]} \hat{\beta}_2$.
Last, for the jackknife AR statistic, the variance estimator is given by
which is consistent for the asymptotic variance of $\hat e^\top (P - \bar P) \hat e$, given by
We make the following assumptions regarding the inference with many IVs.
We now investigate the asymptotical properties of $\phi^*_n$ when the low-dimensional IVs are strong, in the sense that Assumption (ref) holds. This allows us to define the local alternative according to the asymptotic variance of $\hat \beta_1$ and $\hat \beta_2$ and the limiting covariance structure of the component statistics, from which the joint limiting distribution of $(T(\beta_0), LM(\beta_0), AR)$ can be derived. The results with weak low-dimensional IVs are given in Supplementary Appendix A. The formal regularity condition is stated as follows.
The following theorem establishes the joint distribution of the three test statistics above under the local alternative.
To implement the optimal test $\phi_n^*$ defined in (ref), we still need to estimate $\alpha_1$, $\alpha_2$, $\rho_1$, and $\rho_2$. For $\alpha_1$ and $\alpha_2$, we propose the following estimators:
In addition, let $\hat X = M_W (P - \bar P) X$. For $\rho_1$ and $\rho_2$, we propose the following estimators:
By combining Proposition (ref) and Theorem (ref), and invoking the approach developed by mueller2011, we obtain a precise sense of asymptotic optimality for our proposed test $\phi^*_n$, which is formalized in the following Theorem (ref).
Finally, for any fixed alternative, both $T(\beta_0)$ and $LM(\beta_0)$ are consistent and, by construction, avoid the issue of non-monotonic power (noted at the end of Section (ref)) when their corresponding set of IVs is strong. Hence, it is reasonable to anticipate that our combined test will retain these desirable properties, a result that we formalize in the theorem below. However, we emphasize once more that these results remain valid regardless of the strength of the many IVs.
We measure the efficiency improvement of the combination test over the conventional Wald test that uses only low-dimensional IVs by the percentage reduction in the asymptotic length of the resulting confidence interval. Recall from Remark (ref) that, under weak many instruments, the combination test is asymptotically equivalent to the Wald test. Consequently, the associated confidence interval takes the usual “estimator plus and minus a standard error times a critical value" form: $\bigl[\hat\beta_1-\sqrt{\widehat\Phi_1}\times\sqrt{\mathbb C_{\alpha}},\;\hat\beta_1+\sqrt{\widehat\Phi_1}\times\sqrt{\mathbb C_{\alpha}}\bigr]$, where $\hat\beta_1$ and $\hat\Phi_1$ are as defined above, and $\sqrt{\mathbb C_{\alpha}}$ is the standard normal critical value. In this case, the asymptotic efficiency gain over the conventional Wald test is zero.
When the many IVs provide strong identification, the confidence interval associated with our combination test can also be expressed in the familiar “estimator plus and minus a standard error times a critical value" form. To see this, observe that under strong identification by many IVs, the component LM statistic can be represented as follows:
where the $o_P(1)$ term comes from the fact that under strong identification,
Inserting this into ((ref)) gives the following form of our combination test:
The resulting confidence interval is asymptotically equivalent to
where $\hat\beta^*$ is a combined estimator of $\beta$,
The intuition behind the combined estimator is fundamentally efficiency-driven. First, $\widehat \Phi_1$ and $\widehat \Phi_2$ estimate the asymptotic variances of $\hat \beta_1$ and $\hat \beta_2$, respectively, so for given weights $(\hat\omega_1,\hat\omega_2)$ the combined estimator $\hat\beta^*$ assigns greater weight to the estimator with the smaller asymptotic variance. Second, although the AR statistic is asymptotically centered at zero and therefore does not affect the location of the combined estimator, its correlation with the many-IV-based estimator $\hat \beta_2$ allows it to reduce the variance of the combined estimator. This mechanism parallels the construction of the HLIML and HFUL estimators in Haus2012, which can be more efficient than the JIVE estimator. Third, the estimated weights $(\hat \omega_1,\hat \omega_2,\hat \omega_3)$ are consistent for the population weights $(\omega_1,\omega_2,\omega_3)$ in (ref) associated with the UMPU test. Finally, the optimal weights defined in (ref) solve the following problem:
The quadratic constraint ensures that $CI^*$ attains the correct asymptotic coverage, while the objective corresponds to the asymptotic variance of the combined estimator $\hat\beta^*$, since
Thus, the optimal weights in (ref), originally motivated by the UMPU testing problem, also yield an optimally efficient combined estimator of $(\hat \beta_1, \hat \beta_2, AR)$ that achieves the minimal asymptotic variance.
A direct implication of the confidence interval in ((ref)) is that, asymptotically, the percentage reduction in its length relative to the confidence interval based on the conventional Wald test can be derived analytically as follows,
As equation ((ref)) shows, the efficiency gain is primarily driven by the relative identification strength of the low-dimensional and many IVs, summarized by \[ \frac{a_2}{a_1} = \lim_{n \to \infty} \sqrt{\frac{\Phi_1}{\Phi_2}}, \] as well as by the limiting correlations between the Wald and LM statistics, denoted by $\rho_1$, and between the LM and AR statistics, denoted by $\rho_2$. In particular, consistent with Theorem (ref), when $a_2/a_1 = \rho_1$, the combination test $\phi_n^*$ yields no efficiency improvement, and its confidence interval is asymptotically of the same length as that based on $\tilde \phi_n$. This case includes the weak many-IV scenario, where $\rho_1 = a_2 = 0$.
By contrast, whenever $a_2/a_1 \neq \rho_1$, the resulting confidence interval is strictly shorter, implying improved efficiency. Moreover, the efficiency gain in ((ref)) is monotonically increasing in $|\rho_2|$, which leads to the lower bound reported in ((ref)) by setting $\rho_2 = 0$. This lower bound can also be interpreted as the CI length reduction achieved by optimally combine $\hat \beta_1$ and $\hat \beta_2$ only.\footnote{The optimal reduction in this case is $\lim_{n \rightarrow \infty} 1- \frac{1/\left(\tilde \omega_1/\sqrt{ \Phi_1} + \tilde \omega_2 / \sqrt{ \Phi_2} \right)}{\sqrt{ \Phi_1}} = 1- \frac{1}{\left(\tilde \omega_1 + \tilde \omega_2 a_2/ a_1 \right)}$, where the optimal weights are computed by
By direct calculation, we can show that
}
Figure (ref) plots this bound as a function of $a_2/a_1$, the ratio of the standard deviations of $\hat\beta_1$ and $\hat\beta_2$, for various values of $\rho_1$.
In practice, $a_2/a_1$ can be estimated by the ratio of any consistent estimator of $\Phi_1$ and $\Phi_2$ under strong identification.\footnote{If the many-IV specification is weakly identified, then $\Phi_1 = o(1)$ and $1/\Phi_2 = O(1)$, so that $a_2/a_1 = 0$. Suppose we estimate $\Phi_1$ and $\Phi_2$ by $\dot \Phi_1$ and $\ddot \Phi_2$, respectively. Although $\ddot \Phi_2$ is not consistent under weak many IVs, Section C.8 shows that $\ddot \Phi_2^{-1/2} = O_P(1)$ and $\dot \Phi_1^{1/2} = o_P(1)$, implying $\sqrt{\dot \Phi_1 / \ddot \Phi_2} \stackrel{p}{\longrightarrow} 0 = a_2/a_1$. Hence, the plug-in estimator of $a_2/a_1$ remains consistent even under weak identification of the many-IV specification.} This yields a simple rule of thumb, discussed in Section (ref): for empirically plausible values of $\rho_1$ between $-0.7$ and $0.7$, if the ratio of the reported standard errors $\sqrt{\dot \Phi_1}$ and $\sqrt{\ddot \Phi_2}$ exceeds $1.05$, then the associated confidence interval shortens by at least $10\%$.
In some applications (e.g., the replication of card(2009) by Goldsmith(2020)), researchers report alternative many-IV estimators (denoted $\tilde \beta_2$ with variance $\tilde \Phi_2$), such as HFUL, instead of the JIVE estimator ($\hat \beta_2$) considered above. Let $\tilde \rho_1$ denote the correlation between the GMM estimator $\hat \beta_1$ and $\tilde \beta_2$. By the same argument, if $\tilde \rho_1 \in [-0.7,0.7]$ (or $\tilde \rho_1 \in [-0.99,0.99]$), the confidence interval $\widetilde{CI}$ based on the optimal combination of $\hat \beta_1$ and $\tilde \beta_2$ is at least $10\%$ shorter than the conventional Wald interval whenever the ratio $\sqrt{\Phi_1/\tilde \Phi_2}$ exceeds $1.05$ (or $1.1$). Moreover, if $\tilde \beta_2$ is asymptotically equivalent to a combination of the JIVE estimator $\hat \beta_2$ and the $AR$ statistic, as is the case for HLIML and HFUL, then our optimal confidence interval ($CI^*$) is weakly shorter than $\widetilde{CI}$ and therefore at least $10\%$ shorter than the conventional Wald interval. Hence, the rule of thumb is not specific to the JIVE estimator $\hat \beta_2$, but applies more broadly to many-IV estimators asymptotically equivalent to a combination of JIVE and the $AR$ statistic.
In this section, we first synthesize the preceding discussions to describe how to practically implement our combination test in an empirically relevant model and then apply it in simulations to assess its finite-sample power performance. In particular, we consider the following model with clustered data,
where $\alpha_g$ and $\xi_g$ denote cluster-specific fixed effects, and $\bar X_{i,g}$, $\bar W_{i,g}$, and $\bar Z_{i,g}$ represent, respectively, the endogenous regressor, exogenous regressors, and (potentially many) base instruments. By demeaning at the cluster level, we partial out the fixed effects and obtain $\tilde Y_{i,g}$, $\tilde X_{i,g}$, $W_{i,g}$, $\tilde Z_{i,g}$, $\tilde V_{i,g}$, and $\tilde e_{i,g}$, echoing the notation in our initial model setup in (ref)-(ref). The setting we consider in this paper is one in which practitioners are often interested in testing $\mathcal{H}_0: \beta = \beta_0 \ \text{against}\ \mathcal{H}_{1}: \beta \neq \beta_0$ using an alternative set of low-dimensional IVs, $\tilde{z}_{i,g} = f_{i,g}(\tilde Z,W)$. These instruments can be formed, for example, by selecting a subset of the original base IVs $\tilde Z_{i,g}$, or by taking a weighted or simply an unweighted average of them.
Our implementation procedure starts by partialling out the covariates $\tilde W$ from $\tilde Y$, $\tilde X$, $\tilde Z$, and $\tilde z$. This produces $Y$, $X$, $Z$, and $z$, which are consistent with the notation used in Section (ref). These transformed variables serve as the effective observations in all subsequent estimation and inference procedures. We outline these subsequent steps below in sequential order.
All of our simulations are based on (ref) and (ref). The fixed effects are generated by $\alpha_g = u_{1g} + g/G$ and $\xi_g = u_{2g} + g/G$, $g=1, \dots, G$, where $u_{1g}$ and $u_{2g}$ are independent standard normal random variables. The control variables in $\bar W_{i,g}$ are generated by the standard normal distribution, and the dimension of $\bar W$ is fixed at $d_w = 10$. The instruments in $\bar Z_{i,g}$ are normally distributed with mean $0$ and cluster-level dependence: within each cluster $g$, the covariance matrix is given by $$\Omega_{1g}=
_{n_g \times n_g}, \quad g=1, \dots, G,$$ and between clusters these instruments are independent of each other; we set $\theta_1 = 0.5$ in our simulations. Finally, to obtain an arguably complex error structure, we first generate $\acute e_{i,g} = \rho \varepsilon_{i,g} + \sqrt{1-\rho^2} \sigma_{i,g} v_{g}$ and $\acute V_{i,g} = \rho \eta_{i,g} + \sqrt{1-\rho^2} \sigma_{i,g} v_{g}$, where $\sigma_{i,g} = \sqrt{\left(0.2+(\bar W_{i,g}^\top\tau)^2\right)/2.4}$. Here, $\varepsilon_{i,g}$, $\eta_{i,g}$, and $v_{g}$ are mutually independent standard normal random variables, $\rho$ governs the degree of endogeneity, $\tau$ is specified below, and we fix $\rho = 0.5$ in all simulations. Next, within each cluster, we premultiply the vectors $\acute e_{i,g}$ and $\acute V_{i,g}$ by $$ \Omega_{2g}=
_{n_g \times n_g}, \quad g=1, \dots, G, $$ thereby generating $\bar e_{i,g}$ and $\bar V_{i,g}$. In our simulations, we set $\theta_2 = 0.7$.
For the parameters, we set $\beta = 0.3$, $\gamma = \tau = (1/\sqrt{d_w}) \times \iota_{d_w}$, where $\iota_{d_w}$ is a $d_w \times 1$ vector of ones. We specify geometrically decaying coefficients of IVs as $\pi = \left( \phi^0, \phi^1, \ldots, \phi^{K-1} \right)$, where $K$ denotes the number of (many) base IVs and $\phi$ controls the relative weight assigned to each instrument. It is noted that $\phi = 0$ represents the case in which only the first instrument has identification strength, while $\phi = 1$ corresponds to the case in which each instrument has the same identification strength. We further normalize $\pi$ to have $\left\Vert \pi \right\Vert_2 = \sqrt{\psi \sqrt{K}/n}$, where $\psi$ controls the identification strength of the many IVs. The one-dimensional IV is constructed by taking the average of the many IVs; as $\phi$ approaches one, the identification strength of the low-dimensional IV becomes stronger since it is closer to the optimal instrument. We set the sample size at $n = 2{,}000$ and the number of clusters at $G = 500$, and then generate heterogeneous cluster sizes following a procedure similar to that in Djogbenou-Mackinnon-Nielsen-2019. Specifically, for $g=1, \cdots,G-1$, we set $n_g = \max\left\{1, n \exp(2 g / G) / (1+\sum_{g=1}^{G-1} \exp(2 g / G)) \right\}$ and then the size of the last cluster as $n_{G} = \max\left\{1, n - \sum_{g=1}^{G-1}n_g \right\}$. For the dimension of $\bar Z$, we consider $K = 100$ and $K = 500$, respectively. All the results below are based on $5,000$ simulations.
Figure (ref) displays the power curves for our combination test $\phi_n^*$ along with those for the component Wald and jackknife LM tests, at different values of $K$ (the dimension of the many IVs), $\psi$ (which governs the identification strength of the many IVs), and $\phi$ (which controls the identification strength of the one-dimensional IV relative to the many IVs). We identified three main observations, each of which aligns with our large-sample theory. First, in every scenario, the combination test $\phi^*_n$ attains the correct size and is more powerful than each of the other two tests. In particular, as shown in Panel C, the power curve $\phi^*_n$ is never dominated by that of the Wald test, regardless of the strength of the many IVs, thereby underscoring the “free lunch" efficiency gains delivered by our combination test. Second, within Panels A and B of Figure (ref), we observe that for fixed $K$ and $\psi$, the power improvement of $\phi_n^*$ over the Wald test becomes more substantial as the identification strength of the one-dimensional IV weakens relative to that of the many IVs (i.e., as $\phi$ decreases). This is reflected in the widening gap between the power curves of $\phi_n^*$ and the Wald test. It emphasizes how many-IV-based LM and AR statistics contribute to the power enhancement. Third, in contrast to the cases in which the one-dimensional IV dominates many IVs in strength (first figure in Panel A or B) and the power curve of $\phi_n^*$ coincides with that of the Wald test, noticeable gaps persist between the power curves of $\phi_n^*$ and the LM test in the flipped cases (second and third figures in Panel A or B). These gaps highlight how the AR component contributes to power enhancement through its correlation with the LM statistic.
This paper proposes an inference approach that improves conventional estimation and inference in instrumental variables regressions using low-dimensional instruments (e.g., aggregated shift-share IVs) and their underlying high-dimensional base instruments. The procedure requires strong identification only for the low-dimensional IV regression while allowing the many-instrument specification to be weakly identified. We also provide a practical rule of thumb for when inference improves by at least $10\%$, based solely on the ratio of variances of the low- and high-dimensional IV estimators commonly reported in applications. Extensions to settings with many controls or diverging cluster sizes, alternative bootstrap procedures, and combinations with tests such as the sup-score test are left for future work.