EconBase
← Back to paper

Inference with Many Weak Instruments and Heterogeneity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

90,736 characters · 17 sections · 81 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference with Many Weak Instruments and Heterogeneity

abstractThis paper considers inference in a linear instrumental variable regression model with many potentially weak instruments, in the presence of heterogeneous treatment effects. I first show that existing test procedures, including those that are robust to either weak instruments or heterogeneous treatment effects, can be arbitrarily oversized. I propose a novel and valid test based on a score statistic and a “leave-three-out" variance estimator. In the presence of heterogeneity and within the class of tests that are functions of the leave-one-out analog of a maximal invariant, this test is asymptotically the uniformly most powerful unbiased test. In two applications to judge and quarter-of-birth instruments, the proposed inference procedure also yields a bounded confidence set while some existing methods yield unbounded or empty confidence sets.

\vskip 0.2in Keywords: Many Weak Instruments, Heterogeneous Treatment Effects. \vskip 0.2in

Introduction

Many empirical studies in economics involve instrumental variable (IV) models with many instruments. A prominent example is the examiner design: several studies argue that judges or case workers are as good as randomly assigned and can affect the treatment status, so they are used as instruments to study the effects of foster care doyle2007child, incarceration kling2006incarceration, detention dobbie2018effects, disability benefits autor2019disability, and misdemeanor prosecution agan2023misdemeanor, among others. When the IV is a vector of indicators for judges, the number of instruments can be large relative to the sample size. Another example of many IV is a single instrument interacted with discrete covariates. When angrist1991does used the quarter of birth as an instrument to study the returns to education, interacting the quarter of birth with the state of birth generates 150 instruments.

commentRecent econometric research also suggests that many instruments should be used. With covariates, blandhol2022tsls and sloczynski2020should show that the standard two-stage least squares (TSLS) estimator cannot be interpreted as a non-negatively weighted average of causal effects in general, as the TSLS estimator can put negative weights on local average treatment effects (LATE). The standard interpretation of TSLS as a non-negatively weighted average of LATE's is retained only with some parametric assumptions or when the regression is fully saturated (i.e., where instruments are fully interacted with covariates). With a saturated regression, several jackknife estimators recover a positively weighted average of LATE's. evdokimov2018inference, chao2023jackknife, boot2024inference Unless there are a few (or no) discrete covariates, fully interacting the instrument with covariates creates many instruments, further motivating the many IV setting.

Despite the pervasiveness and importance of this setting, there does not yet exist an inference procedure that is robust to both heterogeneous treatment effects and weak instruments, which is a gap this paper aims to fill. Weak IV refers to a setting where the first-stage coefficients converge to zero at a rate such that no consistent estimator for the object of interest exists (following from the definition in mikusheva2022inference rather than chao2012asymptotic); and heterogeneous treatment effects refers to a setting where different subsets of the many IV may estimate different local average treatment effects (LATE). There are several recent proposals crudu2021inference, mikusheva2022inference, matsushita2022jackknife that are robust to weak IV, but they assume constant treatment effects. A separate literature evdokimov2018inference proposed variance estimators for the Jackknife IV Estimator (JIVE) that are robust to heterogeneous treatment effects, but their $t$-statistic test is still not robust to many weak IV. While it is clear that weak IV can lead to substantial distortions in inference (e.g., dufour1997some, StaigerStock97), it is less obvious if procedures developed under constant treatment effects that are robust to weak IV are still valid with heterogeneous treatment effects.

In this paper, I first show that neglecting either heterogeneity or weak instruments can result in substantial distortions in inference. (ref) presents a simple simulation that has both weak instruments and heterogeneous treatment effects. For a nominal 5% test, using the procedure from mikusheva2022inference (MS22), which is robust to weak instruments but not heterogeneity, can result in 100% rejection under the null, because their test statistic is not centered correctly when there is heterogeneity. This result is attributed to how their test is a joint test of both the parameter value and the null of no heterogeneity. Similarly, the procedure from evdokimov2018inference (EK18), which is robust to heterogeneity but not weak instruments, can be severely oversized, as expected due to dufour1997some. Additionally, this section documents how an empirically common practice of constructing a “leniency measure" that combines the many instruments and then using weak IV robust procedures from the just-identified IV literature is invalid.

Given the evidence on how existing methods have substantial distortions in inference, (ref) proposes a procedure for valid inference. Following the many instruments literature, the JIVE estimand is the object of interest --- this estimand can be interpreted as a weighted average of treatment effects when there is heterogeneity (e.g., EK18). Using weak identification asymptotics, I show that the Lagrange Multiplier (LM) (i.e., score) statistic, earlier proposed by matsushita2022jackknife under constant treatment effects, is mean zero and asymptotically normal even with treatment effect heterogeneity. In fact, I prove a stronger normality result that a set of jackknife statistics that includes the LM is jointly normal, which is the first technical challenge of this paper. This normality result uses an asymptotic environment that nests the asymptotic environments of EK18 and MS22 in that normality holds if either the number of instruments is large or the instruments are strong. This normality implies that, as long as the variance of LM is consistently estimable, a $t$-statistic can be calculated and critical values from the standard normal distribution are valid for inference. Obtaining a consistent variance estimator is the second technical challenge of the paper, since reduced-form coefficients are not consistently estimable when there are few observations per instrument. Motivated by anatolyev2023testing who proposed a method to jointly test the significance of many covariates in OLS, I construct a leave-three-out (L3O) variance estimator for the LM variance and show that it is consistent, even when reduced-form coefficients are not consistently estimable. Due to the generality of the setting considered, beyond its robustness to weak IV and heterogeneity, the procedure proposed in this paper is also robust to heteroskedasticity, and potentially many covariates, so it retains the advantages of existing procedures in the literature.

(ref) argues that the proposed LM procedure is powerful. In the over-identified IV environment with normal homoskedastic errors, moreira2009maximum showed that, if we are willing to restrict our attention to tests that are invariant to rotations of the instruments, it suffices to consider tests that are functions of three statistics. These three statistics are known as a “maximal invariant". To be robust to non-normality and heteroskedasticity in the many IV environment, I focus on the leave-one-out (L1O) analog of this maximal invariant. The proposed LM statistic is one of the three statistics in the L1O analog, and I show that the two-sided LM test is asymptotically uniformly most powerful unbiased (UMPU) within the class of tests that are functions of this L1O analog, for the interior of the alternative space (i.e., where heterogeneity is imposed).

commentTo investigate the power properties of one-sided tests beyond the set, I construct a power envelope (i.e., the highest achievable power for any feasible test that is valid across the entire null space for a given alternative) using the algorithm in elliott2015nearly (henceforth EMW). Using an adversarial covariance matrix, the power of the LM test can be substantially lower than the power envelope. Calibrating the variance matrix based on the empirical application of angrist1991does, LM is essentially indistinguishable from the EMW power envelope. This numerical evidence suggests that LM is not necessarily uniformly most powerful, but it is nonetheless simple and informative, and the empirical gains from a more powerful procedure are likely small in the given asymptotic problem.

Simulation results in (ref) show how the procedure is robust even with a small number of instruments, and it is reasonably powerful even with constant treatment effects. (ref) contains two empirical applications that show how being robust to many weak IV and heterogeneity can change conclusions.\footnote{ Implementation code can be found at: \url{https://github.com/lutheryap/mwivhet}. } In the angrist1991does quarter of birth application, the matsushita2022jackknife procedure that are robust to many weak IV but not heterogeneity have unbounded confidence sets while L3O has a bounded confidence set. In the agan2023misdemeanor judge application, MS22 has an empty confidence set, and the length of the L3O confidence interval is more than twice that of EK18, which is not robust to many weak IV.

This paper contributes to the following strands of literature. First, this paper contributes to a growing literature on many weak instruments. There is a strand of literature dealing with many instruments (e.g., chao2012asymptotic) and another separate strand dealing with weak instruments (e.g., StaigerStock97, lmmpy23). While recent procedures accommodate both simultaneously (e.g., crudu2021inference, mikusheva2022inference, matsushita2022jackknife, lim2024conditional), their focus has been on the linear IV model with constant treatment effects. This paper augments their setup by allowing for heterogeneity in treatment effects, and contributes new results on the limitations of their procedures under heterogeneity. Further, I show how heterogeneity can be understood in a framework analogous to weak instruments.

Second, this paper contributes to the literature on heterogeneous treatment effects (e.g., kolesar2013estimation, evdokimov2018inference, blandhol2022tsls). The previous papers exploit consistent estimation of the object of interest to conduct inference. In contrast, this paper uses the (more general) weak IV environment where the object of interest may not be consistently estimated. Two recent papers allow weak IV and heterogeneity. boot2024inference study a single discrete instrument interacted and saturated with many covariates. Their setup is a special case of the environment considered in this paper, so it is unclear if their procedure generalizes to many instruments without covariates (e.g., judges). kleibergen2025double target a continuous updating (CU) GMM estimator with a fixed number of instruments rather than many instruments.\footnote{ They show that their CU-GMM estimator corresponds to the limited information maximum likelihood (LIML) estimator. However, it is also known that the LIML estimand may not be interpretable as a weighted average of LATE's. kolesar2013estimation I am unaware of any paper that allows both weak IV and heterogeneity with a fixed number of instruments and targets a parameter that is a weighted average of LATE's. }

Third, this paper contributes to a literature on inference when coefficients cannot be consistently estimated. The difficulty in having such a general robust inference procedure lies in consistent variance estimation when the number of coefficients is large. Recent literature that has made substantial progress in a different context. In doing inference in OLS with many covariates, cattaneo2018inference and anatolyev2023testing proposed consistent variance estimators that are robust to heteroskedasticity, which involve inverting a large ($n$ by $n$, where $n$ is the sample size) matrix and a L3O approach respectively. boot2024inference adapt the cattaneo2018inference variance estimator for inference. In contrast, this paper adapts the approach from anatolyev2023testing that does not require an inversion of an $n$ by $n$ matrix, and whose L3O implementation is fast when using matrix operations.

Fourth, this paper contributes to a literature on optimal tests. While the UMPU test for just-identified IV has been established since moreira2009tests, obtaining a UMPU test in the over-identified IV environment has thus far been more challenging. In the over-identified IV environment with constant treatment effects, several statistics are informative of the object of interest. Consequently, there is a large literature that numerically compares various valid tests and characterizes various forms of optimality (e.g., andrews2016conditional, andrews2019optimal, van2023power, lim2024conditional). By imposing heterogeneity in the environment, the problem is (somewhat surprisingly) simplified. Since only one statistic in the asymptotic distribution is directly informative of the object of interest, I obtain a UMPU result.

Challenges in Conventional Practice

This section explains the challenges faced in conventional practice by considering a simple potential outcomes model without covariates that exhibits weak instruments and heterogeneity in treatment effects. This model is a special case of the general model in Section (ref). A simulation from the model shows how weak instruments and heterogeneity can lead to substantial distortions in inference for procedures recently proposed in the econometric literature. A common empirical practice of constructing a leave-one-out instrument and then applying inference methods for the instrument as if it is not constructed also has high rejection rates. In contrast, the method proposed in this paper has a rejection rate that is close to the nominal rate.

Setting for Simple Example

The simple example uses the canonical latent variable framework of heckman2005structural. We are interested in the effect of $X_i \in \{ 0,1 \}$ (e.g., incarceration) on some outcome $Y_i$, for $i = 1, \cdots n$ that indexes individuals. To instrument for $X_i$, we use a vector of judges indicators: $Z_i$ is a $(K+1)$-dimensional vector of indicators for judges, indexed $k=1,\cdots, K+1$, each with $c=5$ individual cases, so the vector takes value 1 for the $k$th component when individual $i$ is matched to judge $k$, and 0 elsewhere. Then, $n = (K+1)c$. The problem of many instruments arises when $c$ is fixed while $K$ increases. Let $Y_i(0)$ and $Y_i(1)$ denote the untreated and treated potential outcomes respectively, and we observe $Y_i = Y_i (X_i)$. The treatment status given some instrument value $z$ is $X_i(z)$, and we observe $X_i (Z_i)$. The model is:

equation[equation omitted — 102 chars of source]

where $1\{ \cdot \}$ is an indicator function that takes the value 1 if the argument is true and 0 otherwise. Here, $Z_i^\prime \lambda = \lambda_{k(i)}$, where $k(i)$ is the judge that individual $i$ is matched to. With individual unobservable $v_i \sim U[0,1]$, the probability of treatment (i.e., $X_i=1$) given judge $k$ is $\lambda_k$. I set $\lambda_k = 1/2$ for the base judge, and evenly split all other $K$ judges to take 4 different values of $\lambda_k$. Potential outcomes are $Y_i(0) = \varepsilon_i$ and $Y_i(1) = f(v_i) + \varepsilon_i$ so $Y_i(1) - Y_i(0) = f(v_i)$ is the treatment effect. The individual-specific residuals $v_i$ and $\varepsilon_i$ are allowed to be arbitrarily correlated. Let $\beta_k$ denote the local average treatment effect (LATE) when comparing judge $k$ to the base judge: for instance, when $\lambda_k > 1/2$, $\beta_k = \frac{1}{\lambda_k-1/2} \int_{1/2}^{\lambda_k} f(v) dv$. The values of $(\lambda_k, \beta_k)$ for the 4 groups of judges are $(1/2 -s, \beta - h/s), (1/2(1-s), \beta + 2h/s), (1/2(1+s), \beta - 2h/s)$, and $(1/2+s, \beta + h/s)$. The function $f(v)$ that delivers these parameters and further details of this example are in (ref).

The $\lambda_k$ and $\beta_k$ values are parameterized by objects $s$ and $h$, which control the IV strength and heterogeneity in the model respectively. The impact of these parameters are illustrated in (ref) that plots the point masses for the four groups of judges in reduced-form. Parameter $s$ controls how far $E[X\mid Z]$ are spread across judges, which then affects the instrument strength. Parameter $h$ controls the distance between the mass points and a line with slope $\beta$ --- this slope is the object of interest. If the impact of $X$ on $Y$ is homogeneous, then $h=0$, and all mass points must lie on a line --- this implication is falsifiable by the data.

comment\begin{table} \caption{Parameter Values for Simple Example} \begin{tabular}{lrrrrr} \toprule $\textcolor{blue}{\lambda_k}$ & $\frac{1}{2}-\textcolor{blue}s$ & $\frac{1}{2}-\frac{1}{2}\textcolor{blue}s$ & $\frac{1}{2}$ & $\frac{1}{2}+\frac{1}{2}\textcolor{blue}s$ & $\frac{1}{2}+\textcolor{blue}s$\\ \midrule \textcolor{red}{$\beta_k$} & $\beta- \frac{\textcolor{red}h}{s}$ & $\beta+ 2\frac{\textcolor{red}h}{s}$ & NA & $\beta-2\frac{\textcolor{red}h}{s}$ & $\beta+\frac{\textcolor{red}h}{s}$ \\ \bottomrule \end{tabular} \end{table}
figure[figure omitted — 1,111 chars of source]
table[table omitted — 2,195 chars of source]

The simulation designs vary the values of $s$ and $h$ through the parameters:

equation[equation omitted — 138 chars of source]

The statistic $T_{FS}$ is the leave-one-out (L1O) analog of the first-stage “F" statistic in this model, and $T_{AR}$ is similarly the L1O analog of the Anderson-Rubin statistic under the null. These objects are explained in detail in (ref), but it suffices to mention here that, for the given $c$ and $K$, there is a one-to-one mapping between $(E[T_{FS}],E[T_{AR}])$ and $(s,h)$. Using StaigerStock97 asymptotics, $E[T_{FS}]$ is the parameter that determines whether there is strong or weak identification. Where $C$ is some positive arbitrary constant, $E[T_{FS}] \rightarrow \infty$ is an environment with strong identification where the object of interest can be estimated consistently, and $E[T_{FS}] \rightarrow C < \infty$ is an environment with weak identification where no consistent estimator exists.

For every design, I generate data under the null and calculate the frequency that each inference procedure rejects the null of $\beta_{0}=0$. These procedures include the standard TSLS $t$-test, procedures that are robust to either weak instruments (MO, MS) or heterogeneity (EK), and procedures that use a constructed instrument ($\tilde{X}$). The results are presented in (ref), which I will refer to in the remainder of this section as I explain them.

Issue with Many Weak Instruments

Many IV and weak IV are different but related issues. The many IV problem arises when the number of cases per judge $c$ does not diverge to infinity, so that $K$ is large relative to $n$. When $c$ is small, the judge-specific $\lambda_k$ and $\beta_k$ cannot be consistently estimated and hence inference procedures like the TSLS $t$-test can be oversized. In (ref), a small $c$ is attributed to the sample uncertainty surrounding each black circle. The weak IV problem arises from $E[T_{FS}]$ not diverging: since $E[T_{FS}]$ is a function of $K,c$, and $s$, the weak IV issue is related to the many IV issue.

If we simply run the TSLS $t$-test for an over-identified model, then the estimator can be asymptotically biased and inference is invalid, a fact already known in the literature. This fact is also evident in (ref), where TSLS has 100% rejection in many designs. In TSLS, the first stage regresses $X$ on $Z$ to get a predicted $\hat{X} = Z \hat{\pi}$, where $\hat{\pi}$ is the estimated coefficient; the second stage regresses $Y$ on $\hat{X}$. With constant treatment effects, the asymptotic bias of the TSLS estimator depends on $\sum_i \varepsilon_i \hat{X}_i / \sum_i \hat{X}_i^2 $. When every judge only has $c=5$ cases, the influence of $v_i$ on $\hat{\pi}_{k(i)}$ and hence $\hat{X}_i$ is non-negligible. Since $\varepsilon_i$ and $v_i$ can be arbitrarily correlated, the numerator is biased. If the instruments are weak such that the denominator $\sum_i \hat{X}_i^2$ does not diverge sufficiently quickly, then the asymptotic bias can be large. Due to the asymptotic bias, the $t$-statistic is not centered around $\beta_0$ when data is generated under the null, so we observe over-rejection in (ref).

Since the bias in the TSLS estimator arises from using $X_i$ to estimate $\hat{\pi}$, a natural solution to address that bias is to use the JIVE to estimate $\beta$. Instead of using $\hat{X}_i = Z_i^\prime \hat{\pi}$ in the second stage, we instead use $\tilde{X}_i = Z_i^\prime \hat{\pi}_{-i}$, where $\hat{\pi}_{-i}$ is the coefficient from the first-stage regression that leaves out observation $i$. I call $\hat{\pi}_{-i}$ the leave-one-out (L1O) coefficient. With $P=Z\left(Z^{\prime}Z\right)^{-1}Z^{\prime}$ denoting the projection matrix, $\tilde{X}_i = Z_i^\prime \hat{\pi}_{-i}$ can be written as $\tilde{X}_i = \sum_{j\ne i}P_{ij}X_{j}$. Then, the JIVE is:

equation[equation omitted — 138 chars of source]

In the many IV context with constant treatment effects, the asymptotic distribution of the $t$-statistic of the JIVE is the same as the distribution of the $t$-statistic of the TSLS estimator in the just-identified environment mikusheva2022inference --- it is a ratio of two normally distributed random variables. It is well-known that, in the just-identified IV context with weak IV, the rejection rate of the standard $t$-statistic can be up to 100% for a nominal 5% test (e.g., dufour1997some). Hence, like the just-identified IV context, by using a structural model that has sufficiently weak instruments and high covariance, the simulation can deliver high rejection rates.

EK18 have a procedure that is robust to heterogeneity, but not weak instruments, so even if we use their variance estimator for the $t$-statistic, this problem is not alleviated. This fact is evident in the EK column of (ref), where, with a sufficiently large correlation in the individual unobservables, rejection rates can be large.\footnote{ The rejection rate of EK can be 100% under the null in some simulations: one example is given in (ref) in Online (ref). } Hence, ignoring the issue of weak instruments can lead to substantial distortions in inference. In fact, even with strong instruments, there is no guarantee that EK18 achieves the nominal rate, because their variance estimation method requires consistent estimation of the first-stage coefficients $\hat{\pi}$. A condition for consistent variance estimation is that the number of cases per judge is large, which is not $c=5$.

remarkIn the literature, there have been several definitions of weak IV, which I clarify in this remark. Using (ref), there are three asymptotic regimes, ordered from the strongest to the weakest: (i) $\frac{1}{\sqrt{K}}E[T_{FS}] \rightarrow \infty$, (ii) $E[T_{FS}] \rightarrow \infty$, and (iii) $E[T_{FS}] \rightarrow C < \infty$. Regime (i) is a necessary condition for the TSLS estimator to be consistent, so $\frac{1}{\sqrt{K}}E[T_{FS}] \rightarrow C < \infty$ is what StockYogo05 would refer to as weak instruments. Regime (ii) is a necessary condition for the JIVE to be consistent (e.g., chao2012asymptotic, EK18). Regime (iii) is where no estimator is consistent (e.g., MS22). If $K$ is fixed, then (i) and (ii) are the same asymptotically, and (iii) is the relevant weak identification asymptotic regime. If $K \rightarrow \infty$, then there is more ambiguity in what weakness means: chao2012asymptotic and EK18 who assume (ii) are robust to weak instruments when defined in the StockYogo05 sense, because $s$ can converge to 0, albeit at a slower rate than $\sqrt{K}$. In this paper, I follow the StaigerStock97 standard of weak identification where no consistent estimator exists, which corresponds to (iii) that EK18 is not robust to.

Issue with Heterogeneity

Next, consider proposals for inference that are developed for contexts with many weak IV. MS22 (and crudu2021inference) propose using an Anderson-Rubin (AR) statistic $T_{AR}=\frac{1}{\sqrt{K}}\sum_{i}\sum_{j\ne i}P_{ij}e_{i}\left( \beta_0 \right)e_{j}\left( \beta_0 \right)$, for $e_{i}\left( \beta_0 \right):=Y_{i}-X_{i}\beta_{0}$ where $\beta_{0}$ is the hypothesized null value. With constant treatment effects, $e_i := Y_i - X_i \beta$ is the residual. Hence, if the instrument is orthogonal to the residual, then $E[Z_i e_i]=0$.\footnote{ An equivalent way to see how heterogeneity affects inference is through the framework of hall2003large and lee2018consistent: $E[Z_i e_i]=0$ is a special case of a misspecified over-identified GMM problem. The instruments are individually valid, but every component of the $K$ moments in $E[Z_i e_i]=0$ identifies a different treatment effect, so there is no parameter that satisfies all moments simultaneously under heterogeneity. Then, while the estimand is still interpretable as a combination of these treatment effects due to how GMM weights these moments, there are additional components in the variance that affect inference.} Then, $T_{AR}$ is the L1O analog for the quadratic form that tests the moment $E[Z_i e_i]=0$. Since observations are independent, the critical value for the test is obtained from a mean-zero normal distribution. In this model, $E\left[T_{AR}\right]=\sqrt{K} (c-1) h^2$ under the null.\footnote{This result can be obtained as a special case of (ref) in (ref) and using the fact that $\sum_{i}\sum_{j\ne i}P_{ij}^{2}=\sum_{i}\sum_{j\ne i}\left(1/c^{2}\right)=\sum_{i}\frac{c-1}{c^{2}}=\sum_{k}\frac{c-1}{c}$.} Hence, when there are constant treatment effects such that $h=0$ for all $k$, the statistic is unbiased. However, in the setup with heterogeneity, the $T_{AR}$ can be biased: in fact, when $h$ does not converge to zero, $E[T_{AR}]$ diverges, resulting in a 100% rejection rate under the null, even if the oracle variance were used. Further, there does not exist any estimand $\beta$ such that $E\left[T_{AR}\right]=0$, as stated in (ref) of (ref).

A further problem with the feasible MS procedure is that when there is strong heterogeneity ($E[T_{AR}]=2\sqrt{K}$) in this simulation, their cross-fit variance estimate is negative for all simulation draws, as the negative heterogeneity terms are larger in magnitude than the positive variances of the residuals. The formal analysis requires more notation from (ref), so details are deferred to (ref).

Another proposal in the literature that is robust to many weak instruments is matsushita2022jackknife (MO22) who use the statistic $T_{LM}=\frac{1}{\sqrt{K}}\sum_{i}\sum_{j\ne i}P_{ij}e_{i} \left( \beta_0 \right) X_{j} = \frac{1}{\sqrt{K}}\sum_{i}e_{i}\left( \beta_0 \right)\tilde{X}_{i}$. This statistic can be interpreted as the LM (or score) statistic that uses the moment $E[e \tilde{X}] =0$. They propose the following variance estimator $\hat{\Psi}_{MO}$:

equation[equation omitted — 255 chars of source]

While $T_{LM}$ has zero mean under the null even with heterogeneity, a result shown later in (ref), the MO22 variance estimator was constructed under constant treatment effects, so the variance estimand differs from the true variance. It can be shown that $E\left[\hat{\Psi}_{MO}\right]\ne \textup{Var}\left(T_{LM}\right)$, and $\hat{\Psi}_{MO}$ is inconsistent in general, so when it is used to construct the $t$-statistic of $T_{LM}$, the normalized statistic is not distributed $N(0,1)$ asymptotically. Consequently, by constructing a DGP where $\hat{\Psi}_{MO}$ underestimates the variance, it is possible to get over-rejection of the MO22 procedure, as in the cases of (ref) where $E[T_{AR}]$ diverges. As expected, when there is no heterogeneity such that $h=0$, the rejection rate of MO22 and MS22 are close to the nominal rate.

Issue with a Constructed Instrument

In light of problems with weak identification and heterogeneity, there is a large applied literature that transforms a many IV environment into a just-identified single-IV environment. With a single IV, the anderson1949estimation (AR) procedure (among others) is robust to both weak identification and heterogeneity. However, this subsection will argue that such an approach is invalid.

Due to how the JIVE is written, there are several empirical papers that treat $\tilde{X}_{i}=\sum_{j\ne i}P_{ij}X_{j}$ as the “instrument” so that $\hat{\beta}=\sum_{i}Y_{i}\tilde{X}_{i}/\sum_{i}X_{i}\tilde{X_{i}}$, and proceed with inference as if $\tilde{X}_{i}$ is not constructed, but is an observed scalar instrument, usually referred to as a leniency measure. While the resulting estimator is numerically identical to JIVE, there are distortions in inference because the variance estimators do not account for the variability in constructing $\tilde{X}_{i}$.

If the TSLS $t$-statistic inference is used as if $\tilde{X}_{i}$ is the instrument, then its rejection rates in designs with heterogeneity are usually higher than rejection rates of EK18 that accounts for the variance accurately, by comparing the $\tilde{X}$-t and EK columns in (ref). Consequently, in the cases where EK under-rejects, $\tilde{X}$-t can have close to nominal rejection rates by coincidence.

Even if the weak IV robust AR procedure for just-identified IV were used, there can still be distortion in inference (see $\tilde{X}$-AR in (ref)). The AR $t$-statistic is $t_{\tilde{X}AR} := \sum_{i}e_{i}\left( \beta_0 \right)\tilde{X}_{i}/ \sqrt{\hat{V}}$, where $\hat{V} = \sum_{i}\tilde{X}_{i}^{2}\hat{\varepsilon}_{i}^{2}/ \left(\sum_{i}\tilde{X}_{i}^{2}\right)^{2}$ and $\hat{\varepsilon}_{i}=e_{i}\left( \beta_0 \right)-\tilde{X}_{i}\left( \sum_{i}e_{i}\left( \beta_0 \right)\tilde{X}_{i}\right) / \left( \sum_{i}\tilde{X}_{i}^{2} \right)$. Even though $t_{\tilde{X}AR}$ is mean zero and asymptotically normal, the variance estimand is inaccurate, much like MO22. In particular, when $\beta_0=\beta=0$, the leading term of the variance estimand is $E\left[\sum_{i}\tilde{X}_{i}^{2}e_{i}^{2}\right]$, and it does not converge to the true variance derived in (ref) in general. Hence, using the just-identified AR procedure with a constructed instrument results in over-rejection. There are several papers that cluster standard errors by judges, but this approach faces a similar issue.\footnote{ Details of this discussion are relegated to Online (ref). }

As a preview, the L3O procedure proposed in this paper has rejection rates close to the nominal rate while the other procedures can over-reject.

Valid Inference

In light of how existing procedures are invalid in an environment with many weak instruments and heterogeneity as documented in the previous section, this section describes a novel inference procedure and shows that it is valid. I set up a general model, then show that an LM statistic is asymptotically normal and a feasible variance estimator is consistent, which suffices for inference.

Setting: Model and Asymptotic Distribution

The general setup mimics evdokimov2018inference. With an independently drawn sample of individuals $i=1,\dots,n$, we observe each individual's scalar outcome $Y_{i}$, scalar endogenous variable $X_{i}$, instrument $Z_{i}$, and covariates $W_{i}$, with $\dim\left(Z_{i}\right)=K$.\footnote{ The endogenous variable $X_i$ can be extended to a vector with some technical modifications and without conceptual complications. } For every instrument value $z$, there is an associated potential treatment $X_{i}\left(z\right)$, and we observe $X_{i}=X_{i}\left(Z_{i}\right)$. Similarly, potential outcomes are denoted $Y_{i}\left(x\right)$, with $Y_{i}=Y_{i}\left(X_{i}\right)$. Let $R_{i}:=E\left[X_{i}\mid Z_{i},W_{i}\right]$ and $R_{Yi}:=E\left[Y_{i}\mid Z_{i},W_{i}\right]$ be linear in $Z_i$ and $W_i$. The model, written in the reduced-form and first-stage equations, is:

align*[align* omitted — 325 chars of source]

The setup implicitly conditions on $Z_{i},W_{i}$, so $R_{i},R_{Yi}$ are nonrandom.\footnote{If we are interested in a superpopulation where $Z$ is random, then the estimands would be defined as the probability limit of the conditional objects. Then, it suffices to have regularity conditions to ensure that the conditional object converges to the unconditional object.} Linearity in $Z$ and $W$ is not necessarily restrictive when there is full saturation or when $K$ is large.\footnote{Any nonlinear function of the instruments can be arbitrarily well-approximated by a spline with a large number of pieces or a high-order polynomial. Moreover, the arguments in this paper could presumably be extended to a linear approximation of nonlinear functions as long as there are regularity conditions to ensure that higher-order terms are asymptotically negligible. }

Define $e_{i}:=Y_{i}-X_{i}\beta$, where $\beta$ is some estimand of interest, and $e_{i}$ is a linear transformation. Let $e_i \left( \beta_0 \right) := Y_i - X_i \beta_0$ denote the feasible null-imposed linear transformation. Let $R_{\Delta i}:=R_{Yi}-R_{i}\beta$ and $\nu_{i}:=\zeta_{i}-\eta_{i}\beta$. These definitions imply $e_{i}=R_{\Delta i}+\nu_{i}$ and $R_{\Delta i} = Z_i^\prime (\pi_Y - \pi \beta) + W_i^\prime (\gamma_Y - \gamma \beta)$. Since $E\left[\nu_{i}|Z_{i},W_i\right]=0$ from the model, $E\left[e_{i}|Z_{i},W_i\right]=R_{\Delta i}$, which need not be zero. For data matrix $A$, let $H_{A}=A\left(A^{\prime}A\right)^{-1}A^{\prime}$ denote the hat (i.e., projection) matrix and $M_{A}=I-H_{A}$ its corresponding annihilator matrix. With $Z,W$ denoting the corresponding data matrices of the instrument and covariates, let $Q=(Z,W)$, $P=H_{Q}$, and $M=I-P$. $C$ denotes arbitrary constants.

remarkWhile $E\left[e_{i}|Z_{i}, W_i\right]=R_{\Delta i}$ need not be zero under heterogeneous treatment effects, $E\left[e_{i}|Z_{i}, W_i\right]=R_{\Delta i}=0$ under constant treatment effects. Since $R_{\Delta i} = Z_i^\prime (\pi_Y - \pi \beta) + W_i^\prime (\gamma_Y - \gamma \beta)$ for all $i$, constant treatment effects with $E[Y_i - X_i \beta \mid Z_{i}, W_i] = 0$ also implies $\pi_Y = \pi \beta$ and $\gamma_Y = \gamma \beta$ outside of edge cases (e.g., when $Z_i,W_i$ are always 0). These $R_\Delta$ objects hence capture the impact of having heterogeneous treatment effects in the many instruments model.

The (conditional) object of interest and its corresponding estimator are: \[ \beta_{JIVE}:=\frac{\sum_{i}\sum_{j\ne i}G_{ij}R_{Yi}R_{j}}{\sum_{i}\sum_{j\ne i}G_{ij}R_{i}R_{j}}, \text{ and }\quad\hat{\beta}_{JIVE}=\frac{\sum_{i}\sum_{j\ne i}G_{ij}Y_{i}X_{j}}{\sum_{i}\sum_{j\ne i}G_{ij}X_{i}X_{j}}, \] where $G$ is an $n\times n$ matrix that can take several forms. If there are no covariates, using the projection matrix $G=H_Z=P$ is the standard JIVE, and when there are covariates, I use the unbiased JIVE “UJIVE" kolesar2013estimation with $G=\left(I-\textup{diag}\left(H_{Q}\right)\right)^{-1}H_{Q}-\left(I-\textup{diag}\left(H_{W}\right)\right)^{-1}H_{W}$. In an environment with a binary instrument and many covariates interacted with the instrument, the saturated estimand “SIVE" chao2023jackknife, boot2024inference uses $G=P_{BN}-M_QD_{BN}M_Q$, where $P_{BN}=M_{W}Z\left(Z^{\prime}M_{W}Z\right)^{-1}Z^{\prime}M_{W}$ and $D_{BN}$ is defined as a diagonal matrix with elements such that $P_{BN,ii}=\left[M_QD_{BN}M_Q\right]_{ii}$. With constant treatment effects, the estimand is the same for all the estimators: $R_{Yi}=R_{i}\beta$ so $\beta_{JIVE}=\beta$. Depending on the application, the estimand is usually interpretable as some weighted average of treatment effects when using JIVE without covariates or UJIVE with covariates with a saturated regression.\footnote{ In the judge example without covariates above, we have $G=P$ and $\pi_{Yk} = \beta_k \pi_k$ where $\beta_k$ is the local average treatment effect (LATE) between judge $k$ and the base judge, so $\beta_{JIVE} = \frac{\sum_k \pi_{Yk} \pi_k}{\sum_k \pi_k^2 } = \frac{\sum_k \pi_k^2 \beta_k}{\sum_k \pi_k^2}$ is a weighted average of LATE's. } (EK18) The focus of this paper is on inference, so I will not discuss the estimand in detail. The results for valid inference are established for any $G$ that satisfies properties that will be formally stated in the theorem.

This paper restricts its attention to the following statistics:

equation[equation omitted — 226 chars of source]

It suffices to focus on $\left(T_{AR},T_{LM},T_{FS}\right)$ for inference as they correspond to a linear transformation of the leave-one-out analog of a maximal invariant --- details are in (ref). $T_{AR}$ is the (unnormalized) AR statistic used by MS22 for inference, and $T_{LM}$ is the LM (score) statistic used by MO22. $T_{FS}$ corresponds to a first-stage F statistic that can be used as a diagnostic for weak instruments.

Asymptotic theory in this paper uses $r_n/ \sqrt{K} \rightarrow \infty$, where

equation[equation omitted — 185 chars of source]

which nests the environments of EK18, MS22, and MO22: as long as one of the three objects in (ref) diverges at a rate above $\sqrt{K}$, we obtain $r_n/ \sqrt{K} \rightarrow \infty$. EK18 assume strong identification that translates to $\sum_i \left( \sum_{j \ne i} G_{ij} R_j \right)^2/ \sqrt{K} \rightarrow \infty$ in this setting, but $r_n/ \sqrt{K} \rightarrow \infty$ can also be achieved if either of the latter terms in $r_n$ diverges. MS22 and MO22 assume $K \rightarrow \infty$. Without covariates, $G=P$, so $\sum_i \sum_{j \ne i} G_{ij}^2 = O(K)$, and hence $r_n/ \sqrt{K} \rightarrow \infty$. Hence, to apply the asymptotic theory in this paper, it suffices to have either strong identification, or $K \rightarrow \infty$. The only case ruled out is where $K$ is fixed, and there is weak identifcation in that $\sum_i \left( \sum_{j \ne i} G_{ij} R_j \right)^2/ \sqrt{K}$ does not diverge.

assumption\begin{enumerate} [topsep=0pt,label=(\alph*)] • There exists $C < \infty$ such that $E[\eta_i^4] + E[\nu_i^4] \leq C$ for all $i$. • $E\left[\nu_{i}^{2}\right]$ and $E\left[\eta_{i}^{2}\right]$ are bounded away from 0 and $|\textup{corr}\left(\nu_{i},\eta_{i}\right)|$ is bounded away from 1. • There exists $\underline{c}>0$ such that for any $c_1, c_2, c_3$ that are not all 0, \\ $\frac{1}{r_n} \sum_i \left( c_{3}\sum_{j\ne i}\left(G_{ij}+G_{ji}\right)R_{j}+c_{2}\sum_{j\ne i}G_{ji}R_{\Delta j} \right)^2$ \\$+ \frac{1}{r_n} \sum_i \left( c_{1}\sum_{j\ne i}\left(G_{ij}+G_{ji}\right)R_{\Delta j}+c_{2}\sum_{j\ne i}G_{ij}R_{j} \right)^2 $ \\ $+ \frac{1}{r_n} Var\left( \sum_{i}\sum_{j\ne i} G_{ij}\left( c_1 \nu_i \nu_j + c_2 \nu_i \eta_j + c_3 \eta_i \eta_j \right) \right)$ $ \geq c$. \item $\frac{1}{r_n^2} \sum_{i} \left( \left(\sum_{j\ne i}G_{ij}R_{j}\right)^{4} + \left(\sum_{j\ne i}G_{ij}R_{\Delta j}\right)^{4} + \left(\sum_{j\ne i}G_{ji}R_{j}\right)^{4} + \left(\sum_{j\ne i}G_{ji}R_{\Delta j}\right)^{4} \right) \rightarrow0$. \item $|| \frac{1}{r_n} G_{L}G_{L}^{\prime}||_{F}+|| \frac{1}{r_n} G_{U}G_{U}^{\prime}||_{F}\rightarrow 0$, where $G_{L}$ is a lower-triangular matrix with elements $G_{L,ij}=G_{ij}1\left\{ i>j\right\} $ and $G_{U}$ is an upper-triangular matrix with elements $G_{U,ij}=G_{ij}1\left\{ i<j\right\} $. \end{enumerate}

(ref) states high-level conditions that mimic EK18 so that a central limit theorem (CLT) can be applied. These conditions hence accommodate the $G$ that EK18 consider with covariates. Having bounded moments in (a) is standard. Conditions (b) and (c) are sufficient to ensure that the variance is non-zero asymptotically. In particular, (b) rules out perfect correlation: in the simulation, $\textup{corr}(\eta_i,\nu_i) = -1$ is the pathological case that makes the variance zero, but $\textup{corr}(\eta_i,\nu_i) =1$ still allows non-zero variance. Conditions (d) and (e) ensure that the weights placed on the individual stochastic terms are not too large. The condition that $r_n/\sqrt{K} \rightarrow \infty$ is implied by (e) when $G=P$: due to Lemma B3 of chao2012asymptotic, under weak IV asymptotics where $P_{ii}\leq C<1$, we obtain $||G_{L}G_{L}^{\prime}||_{F} \leq C \sqrt{K}$.

Mechanically, if there is weak IV and fixed K, then $|| \frac{1}{r_n} G_{L}G_{L}^{\prime}||_{F} = \frac{1}{K} O(\sqrt{K}) \ne o(1)$, so (e) fails when $r_n/\sqrt{K}$ does not diverge. Notably, the conditions do not require $P_{ii} \rightarrow 0$ so the $\pi,\pi_{Y}$ coefficients need not be consistently estimated.

theoremIf (ref) holds and $\beta = \beta_0$, then $\hat{\beta}_{JIVE}-\beta_{JIVE} = T_{LM}/ T_{FS}$, and for $V = Var(T_{AR},T_{LM},T_{FS})$, \begin{equation} V^{-1/2} \left(\begin{array}{c} T_{AR} - \frac{1}{\sqrt{K}}\sum_{i}\sum_{j\ne i}G_{ij}R_{\Delta i}R_{\Delta j}\\ T_{LM}\\ T_{FS} -\frac{1}{\sqrt{K}}\sum_{i}\sum_{j\ne i}G_{ij}R_{i}R_{j} \end{array}\right)\xrightarrow{d}N\left(\left(\begin{array}{c} 0\\ 0\\ 0 \end{array}\right),I_3\right). \end{equation}

In (ref), $E[T_{FS}] = \frac{1}{\sqrt{K}}\sum_{i}\sum_{j\ne i}G_{ij}R_{i}R_{j}$ is the concentration parameter corresponding to the instrument strength. In the model of (ref), the mapping to the reduced-form $\pi$ can be found in (ref), so the concentration parameter is given by $E[T_{FS}] = \frac{1}{\sqrt{K}} \sum_k (c-1) \pi_k^2 = \frac{5}{8} \sqrt{K} (c-1) s^2$.\footnote{ This concentration parameter is comparable to the concentration parameter in just-identified IV. With slight abuse of notation, suppose the just-identified IV model has a first stage equation with $X= Z \pi + v$ where $\pi = s$. Then, omitting variance normalizations, using the notation from lmmpy23, the concentration parameter is $f_0 = \sqrt{n} s$, which determines if the TSLS estimator is consistent. In the L1O asymptotics, $\sqrt{K} (c-1)s^2 \approx ns^2/\sqrt{K}$ by using $n=(K+1)c$ and approximations $\sqrt{K/(K+1)} \approx 1$ and $(c-1)/\sqrt{c} \approx \sqrt{c}$. By comparing the L1O concentration parameter $ns^2/\sqrt{K}$ with the just-identified IV concentration parameter $f_0 = \sqrt{n} s$, I obtain the notions of weak identification in (ref). } If the instruments are strong, then $E[T_{FS}]\rightarrow\infty$, so $\hat{\beta}_{JIVE}-\beta_{JIVE}\xrightarrow{d}0$. With weak IV, $E[T_{FS}]$ converges to some constant $C<\infty$, so comparing the JIVE $t$-statistic with the standard normal distribution leads to invalid inference even in large samples.

The asymptotic distribution follows from establishing a quadratic CLT that may be of independent interest: it is proven by rewriting the leave-one-out sums as a martingale difference array, and then applying the martingale CLT. While there are existing quadratic CLT available, they do not fit the context exactly. chao2012asymptotic Lemma A2 requires $G$ to be symmetric, which works for $G=P$, but $G$ for UJIVE is not symmetric in general. EK18 Lemma D2 is established for scalar random variables, so I extend it to random vectors.

commentThe theorem states that the asymptotic distribution of $\hat{\beta}$ is a ratio of two normals, which is identical to the distribution of the just-identified TSLS $t$-statistic. While MS22 have observed this result with many weak IV, their results are restricted to the case with constant treatment effects. Here, I show that the distribution holds even with heterogeneous treatment effects. Consequently, under weak identification, comparing the JIVE $t$-statistic with the standard normal distribution leads to invalid inference even in large samples.

(ref) states that $T_{LM}$ is mean zero and asymptotically normal. Hence, if we have access to the oracle variance of $T_{LM}$, we can simply use the statistic $T_{LM}/\sqrt{\textup{Var}(T_{LM})}$ for testing because it has a standard normal distribution under the null. Obtaining a consistent estimator is an issue addressed in the next subsection.

Variance Estimation

To test the null that $H_{0}:\beta=\beta_{0}$, we can calculate $T_{LM}$ using the null-imposed $\beta_{0}$ and an estimator for the variance of $\sqrt{K} T_{LM}$, $\hat{V}_{LM}$, defined later in this section. Then, reject if $KT_{LM}^{2}/\hat{V}_{LM}\geq\Phi\left(1-\alpha/2\right)^{2}$ for a size $\alpha$ test where $\Phi(.)$ is the standard normal CDF. This procedure is valid when $T_{LM}$ is asymptotically normal with mean zero as we have established in the previous section, and when $\hat{V}_{LM}$ is consistent.

Before stating the variance estimator, I first decompose the variance expression in the equation below, which follows from substituting $e_i=R_{\Delta i} + \nu_i$ and $X_i = R_i + \eta_i$ into the variance. For $V_{LM}:=\textup{Var}\left(\sum_{i}\sum_{j\ne i}G_{ij}e_{i}X_{j}\right)$,

equation[equation omitted — 532 chars of source]

With constant treatment effects, only the first line appears in the variance as $R_{\Delta} =0$. With $G=P$, the expression for $\textup{Var}\left(\sum_{i}\sum_{j\ne i}P_{ij}e_{i}X_{j}\right)$ matches the expression in EK18 Theorem 5.3, but their variance estimator cannot be used directly as they required consistent estimation of reduced-form coefficients. By adapting the leave-three-out (L3O) approach of anatolyev2023testing (AS23), an unbiased and consistent variance estimator can be obtained. Intuitively, just as the own-observation bias in TSLS that involves a single sum can be addressed with L1O, an unbiased estimator for the variance expression that involves a triple sum can be obtained with L3O. Let $\tau:=(\pi^\prime, \gamma^\prime)^\prime$ and $\tau_\Delta := ((\pi_Y-\pi \beta)^\prime, (\gamma_Y - \gamma \beta)^\prime)^\prime$ denote the coefficients on $Q$ when running the regression of $X$ and $e$ respectively. The variance estimator is:

equation[equation omitted — 75 chars of source]

with

align*[align* omitted — 886 chars of source]

where

align*[align* omitted — 443 chars of source]

Following AS23, (ref) ensures that the L3O estimator is well-defined.\footnote{ If these conditions are not satisfied, then we can follow the modification in AS23 so that the variance estimator is conservative.}

assumption\begin{enumerate} [topsep=0pt,label=(\alph*)] • $\sum_{l\ne i,j,k}Q_{l}Q_{l}^{\prime}$ is invertible for every $i,j,k\in\left\{ 1,\cdots,n\right\} $. • $\max_{i\ne j\ne k\ne i}D_{ijk}^{-1}=O_{P}(1)$, where $D_{ijk}:=M_{ii}D_{jk}-\left(M_{jj}M_{ik}^{2}+M_{kk}M_{ij}^{2}-2M_{jk}M_{ij}M_{ik}\right)$. \end{enumerate}

(ref)(a) corresponds to AS23 Assumption 1 and (ref)(b) corresponds to AS23 Assumption 4. For consistent variance estimation, we additionally require regularity conditions that are stated in (ref) of (ref). These conditions are satisfied when $G$ is a projection matrix. With these conditions, (ref) below claims that the variance estimator is consistent.

theoremIf $\beta = \beta_0$, Assumptions (ref)-(ref) hold, and (ref) in (ref) holds, then $E\left[\hat{V}_{LM}\right]=V_{LM}$ and $\hat{V}_{LM}/V_{LM}\xrightarrow{p}1$.

With many instruments and potentially many covariates, the reduced-form coefficients $\pi,\pi_{Y},\gamma,\gamma_Y$ are not consistently estimable. The usual approach to constructing variance estimators calculates residuals by using the estimated coefficients, but this approach no longer works when these estimated coefficients are inconsistent. To be precise, applying Chebyshev's inequality for any $\epsilon>0$ yields:

equation[equation omitted — 295 chars of source]

Without an unbiased estimator and when reduced-form coefficients cannot be consistently estimated, the second term in ((ref)) is not necessarily asymptotically negligible. To overcome this problem, I use an unbiased variance estimator so that the second term is exactly zero. Then, it suffices to show that the variance of individual components of the variance are asymptotically small compared to $V_{LM}^2$, so that the first term in ((ref)) is $o(1)$ by applying the Cauchy-Schwarz inequality.

To obtain an unbiased estimator, I use estimators for the reduced-form coefficients $\pi,\pi_{Y},\gamma,\gamma_Y$ that are unbiased and independent of objects that they are multiplied with. The leave-three-out (L3O) approach has this unbiasedness property for linear regressions: when leaving three observations out in the inner-most sum of the $A$ expressions, the estimated coefficient $\hat{\tau}_{-ijk}$ is independent of $i,j,k$ and is unbiased for $\tau$. Then, when taking the expectation through a product of random variables of $i,j,k$ and $\hat{\tau}_{-ijk}$, $\tau$ can be used in place of the $\hat{\tau}_{-ijk}$ component, and the expectations of individual components can be isolated. For instance,

equation[equation omitted — 433 chars of source]

which recovers the triple sums in the $V_{LM}$ expression of ((ref)). Without leaving out observations $j$ and $k$, we would not be able to isolate $E[X_j]$ and $E[X_k]$ in the first equality. Without leaving out observation $i$, we would not be able to isolate $\tau_{\Delta}$ on expectation to obtain $E[\nu_i^2]$ in the second equality. An analogous argument applies to other components of $V$ in ((ref)). Assuming that the residuals have zero mean conditional on $Q$ is crucial: if we merely have $E[Q \zeta]=0$, this argument can no longer be applied.

remarkWhile the proposed $\hat{V}_{LM}$ is motivated by AS23, the contexts and estimators are different. First, the statistic that we are estimating the variance for is different: AS23 demeaned their $\mathcal{F}$ statistic using $\hat{E}_\mathcal{F}$, where $\hat{E}_\mathcal{F}$ is estimated using L1O, so they are interested in the variance of $\mathcal{F} - \hat{E}_\mathcal{F}$ that is mean zero; I use a mean-zero L1O statistic directly in $T_{LM}$. Second, the expectation of their variance estimator takes the form of their (9), which is analogous to the sum of $A_1$ and $A_4$ using the notation above, so repeated applications of their estimator is insufficient to recover all five terms. Hence, to adjust for the $A_4$ and $A_5$ terms here, I additionally require another estimator, and its form is similarly motivated by a L3O reasoning.

Inverting the test to obtain a confidence set is straightforward, as the test statistic $T_{LM}^2$ and variance estimator $\hat{V}_{LM}=B_{0}+B_{1}\beta_{0}+B_{2}\beta_{0}^{2}$ are quadratic in $\beta_0$ (for some $B_0,B_1,B_2$ that are functions of the data), so the confidence set is obtained by solving a quadratic inequality.\footnote{Details are relegated to the online appendix.}

Power Properties

This section characterizes power properties of the valid LM procedure. I first argue that we can restrict our attention to three statistics that are jointly normal by extending the argument from moreira2009maximum. Since the covariance matrix can be consistently estimated, the remainder of the section focuses on the 3-variable normal distribution with a known covariance matrix. With this asymptotic distribution, I show that the two-sided LM test is the uniformly most powerful unbiased test within the interior of the parameter space.

Sufficient Statistics and Maximal Invariant

As is standard in the literature, I consider the canonical model without covariates where the reduced-form errors are normal and homoskedastic (e.g., andrews2006optimal, moreira2009maximum). Suppose $\left(\eta,\zeta\right)$ in the model of (ref) are jointly normal with known variance:

equation[equation omitted — 272 chars of source]

Define: \[ \left(

array[array omitted — 29 chars of source]

\right):=\left(

array[array omitted — 103 chars of source]

\right). \] I restrict attention to tests that are invariant to rotations of $Z$, i.e., transformations of the form $Z \rightarrow ZF^\prime$ where $F$ is a $K \times K$ orthogonal matrix. In particular, an invariant test $\phi(s_1,s_2)$ is one for which $\phi(Fs_1, Fs_2) =\phi(s_1,s_2)$ for all $K \times K$ orthogonal matrices $F$. If we focus on invariant tests, then the maximal invariant contains all relevant information from the data for inference.

Due to moreira2009maximum Proposition 4.1, $\left(s_{1}^{\prime},s_{2}^{\prime}\right)^{\prime}$ are sufficient statistics for $\left(\pi_{Y}^{\prime},\pi^{\prime}\right)^{\prime}$. Further, $\left(s_{1}^{\prime}s_{1},s_{1}^{\prime}s_{2},s_{2}^{\prime}s_{2}\right)$ is a maximal invariant, and \[ \left(

array[array omitted — 29 chars of source]

\right)\sim N\left(\left(

array[array omitted — 89 chars of source]

\right),\Omega\otimes I_{K}\right). \] The maximal invariant $(s_{1}^{\prime}s_{1},s_{1}^{\prime}s_{2},s_{2}^{\prime}s_{2})$ is jointly normal with a mean that depends on $\Omega$ when $K\rightarrow \infty$.\footnote{ This result is stated in Online (ref). } Extending the argument to allow for heterogeneous treatment effects, the object of interest is $\beta = \frac{\pi^\prime Z^\prime Z \pi_Y}{\pi^\prime Z^\prime Z \pi}$ (following EK18), which is invariant to rotations of the instrument.\footnote{ With rotation matrix $F^\prime$ such that $F^\prime F = I$, observe that $X= ZF^\prime F\pi + \eta$, so if we were to run the regression on $ZF^\prime$ instead of $Z$, we would obtain coefficients $F \pi$ instead of $\pi$. Then, the estimand is $ \frac{\pi^\prime F^\prime F Z^\prime Z F^\prime F \pi_Y}{\pi^\prime F^\prime F Z^\prime Z F^\prime F \pi} = \frac{\pi^\prime Z^\prime Z \pi_Y}{\pi^\prime Z^\prime Z \pi}$ as before. }

To be robust to many instruments, heteroskedasticity, and non-normality, I use the leave-one-out (L1O) analog of the maximal invariant (following MS22; lim2024conditional).\footnote{ With heteroskedasticity, the variances are not consistently estimable, so we cannot correct for the variances directly. These variances no longer feature in the L1O analog of the maximal invariant. } Without covariates such that $G=P$, the L1O analog $\frac{1}{\sqrt{K}} \sum_i \sum_{j\ne i} P_{ij} (Y_i Y_j, Y_i X_j, X_i X_j)$ is a linear transformation of $(T_{AR}, T_{LM}, T_{FS})$.\footnote{ To see that $\frac{1}{\sqrt{K}} \sum_i \sum_{j\ne i} P_{ij} (Y_i Y_j, Y_i X_j, X_i X_j)$ is a linear transformation, use the fact that $e= Y+ X\beta$. Then, $\frac{1}{\sqrt{K}} \sum_i \sum_{j\ne i} P_{ij} ((e_i+X_i\beta) (e_j+X_j\beta), (e_i+X_i\beta) X_j, X_i X_j) = (T_{AR} + 2 T_{LM} \beta + T_{FS}\beta^2, T_{LM} - T_{FS}\beta, T_{FS})$.} In the remainder of this section, I focus on testing the null that $\beta_0 =0$ so $e\left( \beta_0 \right)= Y$ and the L1O of the maximal invariant is exactly $(T_{AR}, T_{LM}, T_{FS})$. The results are generalized in the appendix.

commentConsidering the leave-one-out (L1O) analog of the maximal invariant is attractive in this context because it removes the need to subtract the variance objects on the left-hand side of (ref). Without covariates such that $G=P$, I define $(T_{YY}, T_{YX}, T_{XX}) := \frac{1}{\sqrt{K}} \sum_i \sum_{j\ne i} P_{ij} (Y_i Y_j, Y_i X_j, X_i X_j)$. This $(T_{YY}, T_{YX}, T_{XX})$ is the L1O analog of the maximal invariant $(s_{1}^{\prime}s_{1},s_{1}^{\prime}s_{2},s_{2}^{\prime}s_{2})$, which relates to JIVE directly because $\hat{\beta}_{JIVE}= T_{YX}/ T_{XX}$.\footnote{ To see this analogy, $s_1^\prime s_2 = Y^\prime Z (Z^\prime Z)^{-1} Z^\prime X = Y^\prime P X = \sum_i \sum_j P_{ij} Y_i X_j$.} As a corollary of (ref), since $(T_{YY}, T_{YX}, T_{XX})$ is a linear transformation of $(T_{ee}, T_{eX}, T_{XX})$ that is jointly normal, $(T_{YY}, T_{YX}, T_{XX})$ is also jointly normal.\footnote{ To see that $(T_{YY}, T_{YX}, T_{XX})$ is a linear transformation, use the fact that $e= Y+ X\beta$. Then, $(T_{YY}, T_{YX}, T_{XX}) := \frac{1}{\sqrt{K}} \sum_i \sum_{j\ne i} P_{ij} ((e_i+X_i\beta) (e_j+X_j\beta), (e_i+X_i\beta) X_j, X_i X_j) = (T_{ee} + 2 T_{eX} \beta + T_{XX}\beta^2, T_{eX} - T_{XX}\beta, T_{XX})$.} Since $(T_{YY}, T_{YX}, T_{XX})$ is the L1O analog and has the same distribution as the maximal invariant, I restrict our attention to tests that are functions of $(T_{YY}, T_{YX}, T_{XX})$. While validity results in (ref) apply even when $K$ is small, the optimality results here do not apply. Based on (ref), the distribution of the maximal invariant is approximately normal when $K$ is large. When $K$ is fixed, the distribution of the maximal invariant is different from the distribution of L1O statistics, and focusing on the L1O statistics is not justified.

The asymptotic problem involving $(T_{AR}, T_{LM}, T_{FS})$ is:

equation[equation omitted — 533 chars of source]

While $\mu_2 =0$ under the null, $\mu_2$ may not be zero under the alternative. There are several restrictions in the $\mu$ vector, which is assumed to be finite. Since $P$ is a projection matrix, $\sum_i \sum_{j \ne i} P_{ij} R_{i} R_j = \sum_i R_i (\sum_{j } P_{ij} R_j - P_{ii} R_i) = \sum_i M_{ii} R_i^2$. Since the annihilator matrix $M$ has positive entries on its diagonal, we obtain $\mu_3 \geq 0$ and a similar argument yields $\mu_1 \geq 0$. With $\mu_2 = \sum_i \sum_{j \ne i} P_{ij} R_{\Delta i} R_j = \sum_i M_{ii} R_{\Delta i} R_i$, the Cauchy-Schwarz inequality implies $\mu_2^2 \leq \mu_1 \mu_3$. Constant treatment effects implies $\mu_2^2 = \mu_1 \mu_3$, which is a special case of the environment here. Even with covariates, if the regression is fully saturated with $G$ given by UJIVE, the same inequality restrictions hold.\footnote{ See (ref) in Online (ref). } These properties do not contradict the joint normality: even though $\mu_3\geq 0$, $T_{FS}$ can still be negative when using the L1O statistic. The inequalities $\mu_1, \mu_3 \geq0$ and $\mu_2^2 \leq \mu_1 \mu_3$ are also the only restrictions on $\mu$, as it can be shown that there exists a structural model where there are no further restrictions.\footnote{ Online (ref) establishes that there exists a structural model where $\Sigma$ is uninformative about $\mu$, and $\mu_1, \mu_3 \geq0$. Since the model in (ref) is binary, it is insufficient for such a general result, and a continuous $X$ is required. While the result establishes that there exists a structural model where there are no further restrictions, for any given structural model, there can still be further restrictions. }

commentBeyond the necessary restrictions that $\mu_1, \mu_3 \geq0$ and $\mu_2^2 \leq \mu_1 \mu_3$, there is also a question of whether $\Sigma$ places further restrictions on $\mu$, which can give more information about $\mu_2$. While $\Sigma$ is uninformative when we have normal homoskedastic reduced-form errors, it is less obvious if there exists any structural model where this result still holds when $\beta$ features in $\Sigma$. With more structure, there can be more restrictions on $\mu$, but if there is no structural model where $\Sigma$ is uninformative, then any necessary restriction should be accounted for in the asymptotic problem. In particular, the model of (ref) with weak identification and weak heterogeneity where $\sqrt{K}s^{2}\rightarrow C_{S}<\infty$ and $\sqrt{K}h^{2}\rightarrow C_{H}<\infty$ is one such example. Consequently, it suffices to focus on the 3-variable normal model with the non-negativity and Cauchy-Schwarz restrictions on $\mu$. The following lemma expresses the reduced-form objects in terms of structural parameters, and the proposition formally states the result.

Optimality Result

With a size $\alpha$ test, the two-sided LM test against the alternative that $\mu_2 \ne 0$ rejects when $T_{LM}^2/\textup{Var}(T_{LM}) > \Phi(1-\alpha/2)^2$. I consider the benchmark of a uniformly most powerful unbiased test (e.g., lehmann2005testing, moreira2009tests).

propositionConsider a restriction of the alternative $\mu$ space to the interior i.e., $\mu_1, \mu_3 >0$ and $\mu_2^2 < \mu_1 \mu_3$. Then, within the class of tests that are functions of $(T_{AR}, T_{LM}, T_{FS})$, the two-sided LM test is the uniformly most powerful unbiased test for testing $H_0: \mu_2 =0$ against $H_1: \mu_2 \ne 0$ in the asymptotic problem of ((ref)).

The argument for optimality applies a standard optimality result from lehmann2005testing on the exponential family, which includes the normal distribution. To apply the lehmann2005testing result, we require a convex parameter space and the the existence of alternative values above and below the null value. It can be verified that the restricted parameter space is still convex, and the restriction to the interior ensures the latter condition is satisfied. The proposition claims optimality within the class of unbiased tests, and makes no statement about tests that are biased (i.e., where the power somewhere in the alternative space can be lower than the size).

comment\begin{remark} With the characterized asymptotic distribution, there are several other tests that are valid. (1) We can implement a Bonferroni-type correction that constructs a 99% confidence set for both $\mu_{1}$ and $\mu_{3}$, then a 97% test for LM. (2) VtF from lmmpy23 can be adapted, because the asymptotic distribution does not rely on homogeneous treatment effects and the JIVE $t$ statistic has the same distribution as the just-identified TSLS $t$ statistic. (3) With a given structural model, the algorithm from elliott2015nearly can also be applied by using a grid on structural parameters. \end{remark}
remarkBeyond the two-sided UMPU result, we may also consider other power properties. The one-sided LM test is shown to be the most powerful test against a particular subset of the alternative space. Numerically, using a covariance matrix calibrated from an empirical application, the power of the two-sided LM test is also close to that of the nearly optimal test against a weighted average over a grid of alternative values, constructed using the algorithm from elliott2015nearly. Details are in Online (ref).

Studying optimality in the over-identified IV environment has thus far been complicated. With constant treatment effects, both $s_1^\prime s_1$ and $s_1^\prime s_2$ are informative of the object of interest $\beta$, because constant treatment effects implies $\mu_1 = \beta^2 \mu_3$ in addition to $\mu_2 = \beta \mu_3$. However, once we impose $\mu_1>0$ under the null that $\beta=0$, we rule out constant treatment effects by focusing on the interior of the alternative space. Then, the statistic associated with $\mu_1$ is no longer directly informative of $\beta$. Imposing heterogeneity is hence the key to obtaining this UMPU result.

Simulations

This section focuses on the simple example from (ref). I report two sets of simulations that assess the size and one that assesses power. One set of size simulations uses a large $K$ while the other a small $K$. Robustness checks that involve different data generating processes are relegated to the online appendix.\footnote{There are more simulation results using several different structural models in (ref), including settings with continuous treatment $X$, and with covariates. The results are qualitatively similar in those simulations, suggesting that the numerical findings are not unique to the data-generating process chosen.}

(ref) in (ref) reports rejection rates under the null for a relatively large number of judges with $K=400$, each with a small number of cases at $c=5$. L3O performs well across various designs, while existing procedures can substantially over-reject in at least one design. The LMorc column is included as an infeasible theoretical benchmark that uses an oracle variance. The difference between LMorc and L3O is attributed to the variance estimation procedure.

(ref) reports rejection rates under the null for a small number of judges with $K=4$ and a large number of cases at $c=200$. Based on the theory in (ref), L3O should be valid when the instrument is strong, i.e., in the cases with $E[T_{FS}] = .5c$, which is what we observe. Notably, even when $E[T_{FS}] = 2$ or $E[T_{FS}]=0$, the over-rejection for L3O is not too severe. EK performs very well in the cases with $E[T_{FS}]=.5c$ as expected in their theory. In contrast, MS and MO can over-reject severely with strong heterogeneity, even when instruments are strong.

(ref) reports rejection rates under the alternative. When $E[T_{FS}]=0$, the instrument should be completely uninformative about the true parameter, so we should have 0.05 rejection rate for a valid test, which is what we observe for L3O. When $E[T_{FS}]=2\sqrt{K}$, all procedures, including L3O, are very informative. Considering the designs with $E[T_{AR}]=0$ is most interesting, because this is an environment where MS and MO are valid, and the theoretical optimality result excludes this case. Looking at the case with $E[T_{AR}]=0, E[T_{FS}]=2$, L3O is less powerful than MS and MO in small samples, but the loss is less than 7 percentage points.

table[table omitted — 1,220 chars of source]
table[table omitted — 1,278 chars of source]

Empirical Applications

Returns to Education

angrist1991does were interested in the impact of years of education (X) on log weekly wages (Y). They instrument for education using the quarter of birth (QOB). I implement UJIVE using full interaction of QOB with the state of birth and year of birth (resulting in 1530 instruments) without other controls, which is similar to Table VII(2) of angrist1991does that uses the same set of controls but without full saturation. The implementation here differs from the implementation of MS and MO in that I do not linearly partial out other covariates, but merely saturate on state and year of birth. This implementation is motivated by recent econometric research (e.g., blandhol2022tsls, sloczynski2020should) that argue that the standard interpretation of estimands as a weighted average of LATE's is only retained with some parametric assumptions or when the specification controls for covariates richly, which can be achieved with full saturation.\footnote{ A further advantage of this implementation is that the code is fast: when $G$ is block-diagonal, it suffices to loop over blocks. } To ensure that the different procedures are directly comparable, I adapt the MS and MO inference procedures to target UJIVE, so that the estimand is the same across all procedures and differing results can be attributed purely to inference. TSLS is consequently not a meaningful comparison as the estimand is different from the others.

The results are reported in (ref). In addition to the aforementioned procedures, I include results from implementing the procedure in crudu2021inference (CMS) that uses the $T_{AR}$ statistic like MS22, but uses a plug-in variance estimator like MO22.\footnote{ MS22 use a cross-fit variance estimator, while they refer to the CMS variance estimator as the “naive" variance estimator. MS22 argue that their cross-fit variance is more powerful, which corroborates how MS has a bounded confidence set while CMS does not. } Being robust to weak and many IV results in L3O having a longer confidence interval than EK. With full saturation, CMS and MO yield unbounded confidence sets, while L3O yields a bounded confidence set, showing how robustness to heterogeneity changes the shape of the confidence set in this context. The shape of the confidence set depends on the coefficient on $\beta_{0}^2$. In particular, for $\Psi_{2}:= \frac{1}{K}\sum_{i}\left(\sum_{j\ne i}G_{ij}X_{j}\right)^{2}X_{i}^{2}+\frac{1}{K}\sum_{i \ne j}^n G_{ij}^{2}X_{i}^{2}X_{j}^{2}$, MO is unbounded when $T_{FS}^{2}-q\Psi_{2}<0$ and L3O is unbounded when $T_{FS}^{2}-qB_{2}<0$, where $q$ is 3.84 for a 5% test and $B_2$ is the coefficient on $\beta_0^2$ in the expression of $\hat{V}_{LM}$. Consequently, in this application, we can think of $T_{FS}^2/ \Psi_2 = 0.102$ and $T_{FS}^2/ B_2 = 11.8$ as first-stage statistics for MO and L3O respectively that determine whether the confidence sets are bounded.\footnote{ These statistics are “F" statistics with different variance estimators, suggesting that the instruments are meaningfully weak. The MS and MO variance estimators converge to the same object under weak identification such that $\frac{1}{K}\sum_{i} \sum_{j \ne i} G_{ij} R_i R_j \rightarrow 0$, which is not imposed by the asymptotic regime in this paper. } Analogously, when solving a quartic equation in CMS, an unbounded set occurs as $T_{FS}^2/ \left( \frac{2}{K} \sum_{i \ne j}^n G_{ij}^{2}X_{i}^{2}X_{j}^{2} \right) = 0.0545$, where the denominator is their coefficient on $\beta_0^4$ in their variance estimator. In contrast, the MS confidence set is bounded with $T_{FS}^2/ \left( \frac{2}{K} \sum_{i \ne j}^n \frac{G_{ij}^{2}}{M_{ii} M_{jj} + M_{ij}^2} X_{i}^{2}X_{j}^{2} \right) = 23.9$.

Due to the $\check{M}$ terms in the L3O expression, it is difficult to compare the estimates directly. However, it is possible to compare the estimands of these coefficients in the judge example without covariates:

equation[equation omitted — 158 chars of source]

If $R_{i}^{2}>3\left(1-2P_{ii}\right)E\left[\eta_{i}^{2}\right]$, then MO is more likely unbounded. Intuitively, the MO estimator contains additional products of $R$ that are not present in the true variance, and there are products of $R$ and the error present in the true variance that MO does not account for, motivating the aforementioned difference. We can interpret this condition as MO being more likely unbounded when the signal-to-noise ratio $R_i^2/E\left[\eta_{i}^{2}\right]$ is sufficiently large. In this application, by observing that the first-stage statistic of L3O is an order of magnitude larger than that of MO (i.e., $B_2$ is an order of magnitude smaller), and by comparing the CMS and MO first-stage statistics, there is evidence that $R_i^2/E\left[\eta_{i}^{2}\right]$ is large.\footnote{ Comparing the statistics between MO and MS implies $\frac{1}{K}\sum_{i}\left(\sum_{j\ne i}G_{ij}X_{j}\right)^{2}X_{i}^{2} < \frac{1}{K} \sum_{i} \sum_{j \ne i} G_{ij}^{2}X_{i}^{2}X_{j}^{2}$, which can equivalently be written as $\sum_i \sum_{j \ne i} \sum_{k \ne i,j} G_{ij} G_{ik} X_j X_k X_i^2 <0$. With full saturation, observations $i,j$ have $G_{ij} <0 $ when they are in the same covariate group but have different instrument values. Under the MS and MO asymptotic regimes where $\frac{1}{K} \sum_i \sum_{j \ne i} G_{ij} R_i R_j \rightarrow 0$ so $R_i^2/ E[\eta_i^2]$ is negligible, we obtain $\sum_{i} \sum_{j\ne i} G_{ij}^2 E[X_i^2] E[X_j^2]= \sum_{i} \sum_{j\ne i} G_{ij}^2 \left( R_i^2 + E[\eta_i^2] \right) \left( R_j^2 + E[\eta_j^2] \right) = \sum_{i} \sum_{j\ne i} G_{ij}^2 E[\eta_i^2] E[\eta_j^2] + o(1)$, and $\frac{1}{K}\sum_{i} \sum_{j\ne i} \sum_{k \ne i,j} G_{ij} G_{ik} E\left[ X_j X_k X_i^2 \right] = \frac{1}{K}\sum_{i} \sum_{j\ne i} \sum_{k \ne i,j} G_{ij} G_{ik} R_j R_k \left( R_i^2 + E[\eta_i^2] \right) = o(1)$ is asymptotically negligible. Since the difference between MS and MO is the same magnitude as the MS statistic, $\frac{1}{K}\sum_{i} \sum_{j\ne i} \sum_{k \ne i,j} G_{ij} G_{ik} E\left[ X_j X_k X_i^2 \right]$ is of similar order as $\sum_{i} \sum_{j\ne i} G_{ij}^2 E[X_i^2] E[X_j^2]$, so $R_i^2 / E[\eta_i^2]$ is non-negligible in this application. } This result does not depend on heterogeneity, because the coefficient of $\beta_0^2$ depends only on how $X$ is combined in the variance estimator.

table[table omitted — 735 chars of source]
table[table omitted — 624 chars of source]

Misdemeanor Prosecution

agan2023misdemeanor were interested in the effect of misdemeanor prosecution (X) on criminal complaint in two years (Y). They instrument for misdemeanor prosecution using the assistant district attorneys (ADAs) who decide if a case should be prosecuted in the Suffolk County District Attorney's Office in Massachusetts. As agan2023misdemeanor argued that as-if randomization holds conditional on court-by-time controls and that individual covariates are not required for relevance or exogeneity to hold in this context, the confidence set is constructed using full saturation of court-by-year and court-by-day-of-week fixed effects with no other controls for individual covariates.

As reported in (ref), with full saturation, the UJIVE is $-0.11$, so not prosecuting decreases the probability of criminal involvement by 11 percentage points.\footnote{This result is smaller than $-0.36$ reported in their Table III(3) that uses TSLS with a leniency measure. The result is more similar to the UJIVE robustness check in their Table A.1(5) of $-0.15$ with full saturation of the instrument, but their specification includes case/ defendant covariates, which results in a different estimator. } The L3O confidence interval (CI) is more than twice that of EK: unlike (ref) where $n/K = 221$, we have $n/K =11.9$ here, so a variance estimator that is robust to many IV has a larger impact on CI. MS has an empty confidence set while L3O has a bounded set, showing how being robust to heterogeneity can change conclusions.\footnote{ An empty confidence set using the AR procedure also suggests that the model with constant treatment effects is rejected, so there is meaningful heterogeneity. Since the variances of MO and L3O converge to the same object under homogeneity, the difference between MO and L3O confidence sets also suggests that there is heterogeneity. } Mechanically, the confidence set for MS solves a quartic equation, so an empty set can occur, but it is difficult to characterize when this phenomenon occurs in general.

The L3O CI is also shorter than MO, so being robust to heterogeneity decreases the length of the CI. Considering how the length of the MO confidence set is longer than L3O while being oversized in simulations, there is a question of when MO is conservative. While it is difficult to compare the confidence intervals or variance estimators directly, it is possible to compare the null-imposed variance estimands in the judge example without covariates. It can be shown that:

align*[align* omitted — 282 chars of source]

Then, MO is conservative when: (i) $R_{i}^{2}> (1- 2P_{ii})E\left[\eta_{i}^{2}\right]$, and (ii) $E\left[\eta_{i}\nu_{i}\right]$ is negatively correlated with $R_{i}R_{\Delta i}$, when $P_{ii} < 1/2$. In (i), $R_{\Delta i}^2$ only affects the magnitude of the difference, and not the sign, so this condition can be interpreted as a condition on the signal-to-noise ratio as before. Condition (ii) results from the $\sum_{i} M_{ii} (1- 2P_{ii}) E\left[\eta_{i}\nu_{i}\right]R_{i}R_{\Delta i}$ term that MO does not account for, and covariances can be positive or negative in general.

Conclusion

This paper has documented how weak instruments and heterogeneity can interact to invalidate existing procedures in the environment of many instruments. Addressing both problems simultaneously, this paper contributes a feasible and robust method for valid inference. The procedure is shown to be valid as the limiting distribution of commonly used statistics, including the LM statistic, in an environment with many weak instruments and heterogeneity, is normal, and a leave-three-out variance estimator is consistent for obtaining the variance of the LM statistic. Beyond its validity, the LM test is also optimal, as it is the uniformly most powerful unbiased test in the asymptotic distribution for the interior of the alternative space. In light of the broader econometric literature on the value of saturated regressions and how many instruments can arise from them, this paper presents a highly applicable, robust, and powerful inference procedure for IV.