EconBase
← Back to paper

When Should We (Not) Interpret Linear IV Estimands as LATE?

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

106,446 characters · 15 sections · 167 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

When Should We (Not) Interpret Linear IV Estimands as LATE?

\singlespacing

titlepage\begin{abstract} In this paper I revisit the interpretation of the linear instrumental variables (IV) estimand as a weighted average of conditional local average treatment effects (LATEs). I focus on a situation in which additional covariates are required for identification while the reduced-form and first-stage regressions may be misspecified due to an implicit homogeneity restriction on the effects of the instrument. I show that the weights on some conditional LATEs are negative and the IV estimand is no longer interpretable as a causal effect under a weaker version of monotonicity, i.e. when there are compliers but no defiers at some covariate values and defiers but no compliers elsewhere. The problem of negative weights disappears in the interacted specification of AI1995, which avoids misspecification and seems to be underused in applied work. I illustrate my findings in an application to the causal effects of pretrial detention on case outcomes. In this setting, I reject the stronger version of monotonicity, demonstrate that the interacted instruments are sufficiently strong for consistent estimation using the jackknife methodology, and present several estimates that are economically and statistically different, depending on whether the interacted instruments are used. \end{abstract} Keywords: instrumental variables, local average treatment effects, model misspecification, monotonicity, negative weights, two-stage least squares \\ JEL classification: C21, C26, C52, K42 \thispagestyle{empty}

\setcounter{page}{2}

\onehalfspacing

Introduction

Many instrumental variables are only valid after conditioning on additional covariates. The draft eligibility instrument in Angrist1990 requires controlling for the year of birth. The college proximity instrument in Card1995 is invalid without conditioning on several individual characteristics of workers Kitagawa2015. Even in the case of randomized experiments with noncompliance, it is often necessary to control for covariates correlated with treatment probability, such as household size and survey wave in Finkelsteinetal2012.

When conditioning on additional covariates is necessary for instrument validity, interpreting the linear instrumental variables (IV) and two-stage least squares (2SLS) estimands becomes complicated. AI1995 (hereafter, \citetalias{AI1995}) provide an influential interpretation of the 2SLS estimand in this context as a convex combination of conditional local average treatment effects (LATEs), i.e. average effects of treatment for individuals whose treatment status is affected by the instrument. However, this result is restricted to saturated models with discrete covariates and first-stage regressions that include a complete set of interactions between these covariates and the instrument. Such specifications are rare in empirical work, as is evident from several recent surveys of applications of IV methods.\footnote{BBMT2022,BBMT2025 consider a sample of 99 papers and find a single application of \citetalias{AI1995}'s specification. MTW2021 consider a sample of 122 papers and identify seven with specifications that include some covariate interactions with a single instrument.} This makes \citetalias{AI1995}'s result inappropriate for interpreting the vast majority of IV estimates encountered in economic applications Abadie2003.

In this paper I revisit the question of the causal interpretability of standard instrumental variables estimands. In particular, I focus on whether these estimands can be written as weighted averages of conditional LATEs with positive weights and, if so, whether these weights have an intuitive interpretation. To do so, I consider two variants of the usual monotonicity assumption: “weak monotonicity,” which postulates that at every covariate value, the instrument either does not discourage or does not encourage anyone to take treatment, and “strong monotonicity,” which additionally requires that the direction of this effect is uniform across covariate values.

My first contribution is to demonstrate that under weak monotonicity, the weights on some conditional LATEs may be negative in the usual application of IV, which restricts the first-stage effects of the instrument to be homogeneous. This finding implies that the resulting estimand is not a useful summary measure of average treatment effects; this parameter could be negative (positive) even if treatment effects are positive (negative) for everyone in the population. Under the same assumptions, all weights are necessarily positive in \citetalias{AI1995}'s interacted specification.

My second contribution is to explicitly compare the weights in both specifications with the “desired” weights, which recover the unconditional LATE parameter. Under strong monotonicity, when the weights in the usual application of IV and \citetalias{AI1995}'s specification are positive, both specifications overweight the effects in groups with large variances of the instrument, while the latter also overweights the effects in groups with strong first stages. It follows that the usual application of IV might be preferable when violations of strong monotonicity are not an issue.

However, if weak monotonicity is plausible but strong monotonicity is not, my theoretical results suggest that \citetalias{AI1995}'s interacted specification is preferable to the usual application of IV\@. Unfortunately, \citetalias{AI1995}'s specification is also difficult to estimate without bias; when the researcher divides the sample into many groups and subsequently creates an interacted instrument for each, 2SLS will be subject to the “many instrument” bias (see, e.g., Bekker1994). An alternative approach to estimating \citetalias{AI1995}'s specification, such as the fixed effect jackknife IV (FEJIV) estimator of CSW2023, should be used instead. Another concern about specifications with many instruments is whether they are jointly strong enough to enable consistent estimation. In this context, I consider a recent pretest for weak identification developed by MS2022. As an illustration, I perform an extensive simulation study. In these simulations, MS2022's pretest does a great job differentiating between cases where the best estimators of \citetalias{AI1995}'s specification, such as FEJIV, perform well and cases where all estimators perform badly.

To corroborate the concern about violations of strong monotonicity, I also replicate a sample of 988 instrumental variables regressions from 25 papers published in journals of the American Economic Association between 2006 and 2015. Every specification in my sample is based on a linear first-stage regression that restricts the effects of the instrument to be homogeneous. If strong monotonicity is violated but weak monotonicity is not, the homogeneous first stage will be misspecified and the conditional first stage will be positive for some covariate values but negative for others. First, I present strong suggestive evidence of the latter phenomenon, which directly translates to the incidence of negative weights in the usual application of IV\@. Then, I formally reject the null hypothesis of first-stage homogeneity in more than 70% of specifications in an average paper, despite accounting for multiple hypothesis testing.

In addition, I illustrate my findings in an application to the causal effects of pretrial detention on case outcomes Stevenson2018. Here, I consider several saturated specifications, which allows me to compare the estimates of \citetalias{AI1995}'s specification with the usual application of IV\@. I can also formally test whether the conditional first stage is positive for some covariate values and negative for others, and I conclusively reject the null hypothesis of sign homogeneity. The estimates based on \citetalias{AI1995}'s specification are smaller than in the usual application of IV, and the difference is often statistically significant. MS2022's pretest rejects in every case I consider, which supports the notion that the estimates based on \citetalias{AI1995}'s specification are preferable.

Finally, I supplement this paper with companion MATLAB, R, and Stata packages, fejiv, available at the MATLAB Central File Exchange, the Comprehensive R Archive Network (CRAN), and the Statistical Software Components (SSC) Archive, respectively.\footnote{To download the MATLAB package from the File Exchange, search for fejiv in the Add-On Explorer. To download the R package from CRAN, type \,install.packages("fejiv")\, in the R/RStudio console. To download the Stata package from SSC, type \,ssc install fejiv\, in the Command window.} These packages, based on the MATLAB code of CSW2023, can be used to implement the FEJIV estimator in practice. Of these, the MATLAB package is most appropriate when working with large datasets. A Stata package to implement MS2022's pretest is also available from Sun2023.

Two papers closely related to this are Kolesar2013 and BBMT2022,BBMT2025. Like this paper, Kolesar2013 studies the interpretation of 2SLS estimands under weak monotonicity while also considering the probability limits of several jackknife-type estimators, limited information maximum likelihood (LIML), and other alternatives to 2SLS\@. Kolesar2013's main result on 2SLS is not particular to any specification but instead represents a generic two-step IV estimand as a weighted average of conditional LATEs. The resulting weights are positive, subject to an additional condition that needs to be verified on a case-by-case basis.\footnote{This condition essentially requires that the first stage postulated by the researcher provides a sufficiently good approximation to the true first stage HV2005.} In contrast, this paper focuses specifically on the usual application of IV and \citetalias{AI1995}'s interacted specification. The benefit is that this allows me to considerably simplify the representation and obtain results that are more transparent and easy to interpret. This includes the novel result that in the usual application of IV, the weights on some conditional LATEs may be negative under weak monotonicity. In another contribution, released after this paper first circulated, BBMT2022,BBMT2025 focus on the consequences of misspecification of the model for the instrument propensity score that is implicit in IV and 2SLS estimation. In this paper I focus instead on violations of strong monotonicity and their implications.

The remainder of the paper is organized as follows. Section (ref) introduces my framework. Section (ref) provides my theoretical contributions, a review of the literature on many instruments, and a simulation study. Section (ref) studies negative first stages and first-stage heterogeneity in a sample of recent applications of IV methods and illustrates my findings in an analysis of the causal effects of pretrial detention on case outcomes. Section (ref) concludes. The appendix contains my proofs as well as additional simulation and estimation results.

Framework

In this section I formally define the objects of interest, i.e. the conditional and unconditional IV and 2SLS estimands. I reserve the term “2SLS” for the appropriate estimand in a model with interacted instruments; see equation ((ref)) below. When a single instrument is used instead, I use the term “IV” or “linear IV”; see equation ((ref)). In what follows, I also review identification in the LATE framework with covariates Abadie2003. Throughout the paper I assume that the appropriate moments exist whenever necessary.

Notation and Estimands

Suppose we are interested in the causal effects of a treatment, $D \in \{ 0,1 \}$, on an outcome, $Y = Y(D)$, where $Y(1)$ and $Y(0)$ are potential outcomes. An instrument, $Z \in \{ 0,1 \}$, is also available, and it determines which of the potential treatment states, $D(1)$ and $D(0)$, is observed, $D = D(Z)$. In principle, we could let $Y = Y(Z,D)$, but we will rule out direct effects of $Z$ on $Y$ below. Finally, let $X = \left( 1, X_1, \ldots, X_J \right)$ denote a row vector of covariates. In some cases I will allow for the possibility that additional instruments have been created by interacting $Z$ with all elements of $X$; then, $Z_{\mathrm{C}} = \left( Z, Z X_1, \ldots, Z X_J \right)$ will be used to denote the resulting row vector of instruments.

To provide motivation for what follows, let us consider the standard single-equation linear model for $Y$:

equation[equation omitted — 70 chars of source]

where $X$ and the instrument(s) are assumed to be uncorrelated with the error term $\upsilon$. Also, $\beta$ is the coefficient of interest. In this paper I do not assume that equation ((ref)) is correctly specified; in particular, I allow the effect of $D$ on $Y$ to be correlated with both observables and unobservables.

In practice, however, many researchers act as if this model is correctly specified and use linear IV or 2SLS for estimation. In what follows, I will focus on the interpretation of the probability limits of the IV and 2SLS estimators of $\beta$ when equation ((ref)) is possibly misspecified. With a single instrument, the probability limit of linear IV or, simply, the (linear) IV estimand is

equation[equation omitted — 162 chars of source]

where $W = \left( D, X \right)$, $Q = \left( Z, X \right)$, and $\left[ \cdot \right] _k$ denotes the $k$th element of the corresponding vector. Clearly, when a single instrument is available, equation ((ref)) characterizes the target of estimation in most empirical studies, which I also call the “usual” or “standard” estimand. This specification corresponds to reduced-form and first-stage regressions that project $Y$ and $D$ on $X$ and $Z$, excluding any interactions between $X$ and $Z$\@. Hence, I also refer to this specification as “noninteracted.”

On the other hand, if a vector of interacted instruments, $Z_{\mathrm{C}}$, is used in 2SLS estimation of equation ((ref)), the relevant probability limit or, simply, the 2SLS estimand is

equation[equation omitted — 436 chars of source]

where $Q_{\mathrm{C}} = \left( Z_{\mathrm{C}}, X \right)$. In this specification, the corresponding reduced-form and first-stage regressions project $Y$ and $D$ on $X$ and $Z_{\mathrm{C}}$, which implies that the effects of $Z$ on $Y$ and $D$ are allowed to vary with $X$ due to the interactions between $X$ and $Z$\@. Thus, I also refer to this specification as “interacted” or “fully interacted.”

Regardless of the implicit restrictions on the effects of the instrument, the true first stage can be written as

equation[equation omitted — 90 chars of source]

where

equation[equation omitted — 112 chars of source]

is the conditional first-stage slope coefficient or, equivalently, the coefficient on $Z$ in the regression of $D$ on $1$ and $Z$ in the subpopulation with $X=x$. Similarly, the conditional IV (or Wald) estimand can be written as

equation[equation omitted — 191 chars of source]

This parameter is equivalent to the coefficient on $D$ in the IV regression of $Y$ on $1$ and $D$ in the subpopulation with $X=x$, with $Z$ as the instrument for $D$\@.

Local Average Treatment Effects

In the LATE framework of IA1994 and AIR1996, the population consists of four latent groups: always-takers, for whom $D(1) = D(0) = 1$; never-takers, for whom $D(1) = D(0) = 0$; compliers, for whom $D(1) = 1$ and $D(0) = 0$; and defiers, for whom $D(1) = 0$ and $D(0) = 1$. As demonstrated by IA1994, if, among other things, we rule out the existence of defiers and assume that $X$ is orthogonal to $Z$, the unconditional IV estimand, $\beta_{\mathrm{IV}} = \frac{\e \left[ Y \mid Z=1 \right] - \e \left[ Y \mid Z=0 \right]}{\e \left[ D \mid Z=1 \right] - \e \left[ D \mid Z=0 \right]}$, recovers the average treatment effect for compliers, also referred to as the local average treatment effect (LATE)\@.

Some of my results will allow for the existence of both compliers and defiers, and hence throughout this paper I instead follow Kolesar2013 in defining the LATE as

equation[equation omitted — 106 chars of source]

i.e. the average treatment effect for individuals whose treatment status is affected by the instrument. This group includes both compliers and defiers; it will be restricted to compliers whenever the existence of defiers is ruled out. It is useful to note that this unconditional LATE parameter can also be written as

equation[equation omitted — 130 chars of source]

where

equation[equation omitted — 80 chars of source]

is the conditional LATE and

equation[equation omitted — 67 chars of source]

is the conditional proportion of compliers and defiers. The following assumption, together with additional assumptions below, will be used to identify $\tau(x)$ and $\pi(x)$, and thereby also $\tau_{\mathrm{LATE}}$\@. Recall that $Y = Y(D)$ when direct effects of $Z$ on $Y$ are ruled out and $Y = Y(Z,D)$ otherwise.

uassumption{IV} \begin{enumerate} • • (Conditional independence) \; $\big( Y(0,0), Y(0,1), Y(1,0), Y(1,1), D(0), D(1) \big) \perp Z \mid X$; • (Exclusion restriction) \; $\pr \left[ Y(1,d)=Y(0,d) \mid X \right] = 1$ for $d \in \{ 0,1 \}$ a.s.; • (Relevance) \; $0 < \pr \left[ Z=1 \mid X \right] < 1$ and $\pr \left[ D(1)=1 \mid X \right] \neq \pr \left[ D(0)=1 \mid X \right]$ a.s. \end{enumerate}

Assumption (ref) is standard but not sufficient to identify $\tau(x)$ and $\pi(x)$. It is also necessary to restrict the existence of defiers IA1994. The following assumption, due to Abadie2003, rules out the existence of defiers at any value of covariates.

uassumption{SM}[Strong monotonicity] $\pr \left[ D(1) \geq D(0) \mid X \right] = 1$ a.s.

In many applications, Assumption (ref) may be too restrictive deChaisemartin2017,DHM2023. A testable implication of Assumption (ref) is that $\omega(x)$, the conditional first-stage slope coefficient, is always non-negative. If this is formally rejected or otherwise implausible, an alternative assumption is necessary to obtain point identification. One possibility is to restrict treatment effect heterogeneity, as discussed by HV2005 and MT2018, in which case we will be able to identify the average treatment effect rather than the unconditional LATE parameter. Another possibility is to replace Assumption (ref) with a weaker assumption that postulates the existence of compliers but no defiers at some covariate values and the existence of defiers but no compliers elsewhere. While the relative appeal of these two assumptions is context dependent, I will focus on the latter in what follows.

uassumption{WM}[Weak monotonicity] There exists a subset of the support of $X$ such that $\pr \left[ D(1) \geq D(0) \mid X \right] = 1$ on it and $\pr \left[ D(1) \leq D(0) \mid X \right] = 1$ on its complement.

To understand the difference between Assumptions (ref) and (ref), consider a recent paper by DHMMR2019, who estimate the health effects of air pollution using an instrument based on changes in local wind direction. Imagine a pollution source located to the east of a particular city. When the wind also blows from the east, the city will experience relatively high levels of pollution; the opposite is true when the wind blows from the west. Assumption (ref) would require that every city reacts to a specific wind direction (say, east) in the same way (say, high pollution). This, however, is known not to be true. DHMMR2019 explain, for example, that air pollution is relatively high in San Francisco when the wind blows from the southeast, while the same is true in Boston when the wind blows from the southwest. Indeed, Assumption (ref) would allow for the possibility that different locations react to a specific wind direction in different ways.\footnote{If we knew the pollution-inducing wind direction for every location, as we do in the case of Boston and San Francisco, Assumption (ref) might remain plausible for an appropriately redefined instrument. However, if this direction needs to be estimated, as is likely the case in practice, Assumption (ref) will be more appropriate.}

Importantly, Assumption (ref), together with Assumption (ref), is sufficient to identify $\tau(x)$ and $\pi(x)$. Before stating the relevant lemma, it is useful to define an auxiliary function

equation[equation omitted — 128 chars of source]

where $\sgn(\cdot)$ is the sign function. Clearly, $c(x)$ equals 1 if there are only compliers at $X=x$ and $-1$ if there are only defiers at $X=x$.

The following lemma summarizes identification of the conditional LATE parameter and the conditional proportion of individuals whose treatment status is affected by the instrument.

lemma\begin{enumerate} • • Under Assumptions (ref) and (ref), $\tau(x) = \beta(x)$ and $\pi(x) = \omega(x)$. • Under Assumptions (ref) and (ref), $\tau(x) = \beta(x)$ and $\pi(x) = \left\lvert \omega(x) \right\rvert = c(x) \cdot \omega(x)$. \end{enumerate}

Lemma (ref) consists of well-known results and straightforward extensions of these results, and as such it is stated without proof AIR1996,AP2009. Note that strong monotonicity implies weak monotonicity, which means that every statement that is true under weak monotonicity is also true under strong monotonicity as a special case. I will follow this logic in the statement of the theoretical results below.

Negative Weights in Linear IV

AI1995, Revisited

Let us begin by revisiting \citetalias{AI1995}'s representation of the 2SLS estimand. Recall that \citetalias{AI1995} study a special case of the model in equation ((ref)) where all covariates are binary and represent membership in disjoint groups or strata. In this case, each of the original covariates needs to be discrete or discretized, which means that the population can be divided into $K$ groups, where $K$ denotes the number of possible combinations of values of these variables. (For example, with six binary variables, we have $K = 2^6 = 64$.) Let $G \in \{ 1, \ldots, K \}$ denote group membership and $G_k = 1[G = k]$ denote the resulting group indicators. \citetalias{AI1995} consider a model where original covariates are replaced with these group indicators, $X = \left( 1, G_1, \ldots, G_{K-1} \right)$, while reduced-form and first-stage regressions include a full set of interactions between $X$ and $Z$; that is, $Z_{\mathrm{C}} = \left( Z, Z G_1, \ldots, Z G_{K-1} \right)$. The following lemma restates \citetalias{AI1995}'s and Kolesar2013's interpretation of the 2SLS estimand in this context.

lemma[AI1995,Kolesar2013] Suppose that $X = \left( 1, G_1, \ldots, G_{K-1} \right)$ and $Z_{\mathrm{C}} = \left( Z, Z G_1, \ldots, Z G_{K-1} \right)$. Suppose further that Assumptions (ref) and (ref) hold. Then \begin{equation*} \beta_{\mathrm{2SLS}} = \frac{\e \left[ \sigma^2(X) \cdot \tau(X) \right]}{\e \left[ \sigma^2(X) \right]}, \end{equation*} where $\sigma^2(X) = \var \left[ \e \left[ D \mid X, Z \right] \mid X \right] = \e \left[ \left( \e \left[ D \mid X, Z \right] - \e \left[ D \mid X \right] \right) ^2 \mid X \right]$.

Lemma (ref) establishes that the 2SLS estimand in \citetalias{AI1995}'s interacted specification is a convex combination of conditional LATEs, with weights equal to the conditional variance of the first stage. This result is due to \citetalias{AI1995} and has usually been interpreted as requiring that the existence of defiers is completely ruled out AP2009. Kolesar2013 demonstrates that it also holds under weak monotonicity.

It may not be immediately obvious how the 2SLS weights in Lemma (ref) differ from the “desired” weights in equation ((ref)). The following result facilitates this comparison.

theoremSuppose that $X = \left( 1, G_1, \ldots, G_{K-1} \right)$ and $Z_{\mathrm{C}} = \left( Z, Z G_1, \ldots, Z G_{K-1} \right)$. Suppose further that Assumptions (ref) and (ref) hold. Then \begin{equation*} \beta_{\mathrm{2SLS}} = \frac{\e \left[ \left[ \pi(X) \right] ^2 \cdot \var \left[ Z \mid X \right] \cdot \tau(X) \right]}{\e \left[ \left[ \pi(X) \right] ^2 \cdot \var \left[ Z \mid X \right] \right]}. \end{equation*}

Theorem (ref) shows that the 2SLS estimand in \citetalias{AI1995}'s interacted specification is a convex combination of conditional LATEs, with weights equal to the product of the squared conditional proportion of compliers or defiers and the conditional variance of $Z$\@.\footnote{See also Walters2018 for a related remark that focuses on “descriptive” estimands and does not use the LATE framework for interpretation.} Since the “desired” weights in equation ((ref)) consist only of the conditional proportion of compliers or defiers, \citetalias{AI1995}'s specification overweights the effects in groups with strong first stages and with large variances of $Z$\@. Importantly, this result does not require strong monotonicity; weak monotonicity is sufficient.

remarkAlthough Lemma (ref) and Theorem (ref) show that \citetalias{AI1995}'s specification can avoid negative weights, practitioners rarely use multiple interacted instruments. In a survey of recent applications of IV methods, BBMT2022,BBMT2025 determine that only 1 out of 99 applicable papers has used \citetalias{AI1995}'s specification. Specifications with many interactions between the instrument(s) and covariates were more common in earlier work using IV methods Angrist1990,AK1991 but have since become rare, likely out of concern for the many instrument bias.\footnote{Indeed, BJB1995 write that their results “indicate that the common practice of adding interaction terms as excluded instruments may exacerbate the problem” (emphasis mine). On the other hand, some recent applications of the wind instrument DHMMR2019,BRS2020 and the “judges design” AD2015,MS2015,Stevenson2018 interact the instrument with selected covariates, which is similar in spirit to \citetalias{AI1995}'s specification. However, quantitatively speaking, this is still very rare in practice: in a sample of 122 papers considered by MTW2021, only seven include specifications with some covariate interactions with a baseline instrument.}

Usual Application of IV

Remark (ref) suggests that Theorem (ref) cannot be used directly to interpret most empirical studies because modern applications of IV methods avoid using many interacted instruments. A similar point is made by AP2009, who maintain, however, that an indirect argument in Abadie2003 implies that “some kind of covariate-averaged LATE” is estimated in noninteracted specifications as well. In what follows, I show that AP2009's assertion would be false under weak monotonicity. The claim is true under strong monotonicity, which I will be able to demonstrate directly, deriving the exact form of “covariate-averaged LATE” that linear IV estimates. I also revisit Abadie2003's indirect argument later on.

To save space, I combine two extensions of \citetalias{AI1995}'s analysis in what follows. On the one hand, I am interested in the interpretation of the IV estimand when we retain \citetalias{AI1995}'s restriction that the model for covariates is saturated but no longer use the interacted instruments. This analysis does not require any additional assumptions. On the other hand, I am also interested in the interpretation of the IV estimand in nonsaturated specifications. This analysis proceeds under the assumption that the instrument propensity score, defined as

equation[equation omitted — 50 chars of source]

is linear in $X$\@. This assumption is standard and has been used by Kolesar2013, LM2015, EK2019, and Ishimaru2024, among others.

uassumption{PS}[Instrument propensity score] $e(X) = X \alpha$.

Assumption (ref) holds automatically when $Z$ is randomized, and also when all covariates are discrete and the model for covariates is saturated. (This is why the statement of the theoretical results below only invokes Assumption (ref) and does not separately mention saturated specifications.) Assumption (ref) may also provide a good approximation to $e(X)$ in other situations, especially when $X$ includes powers and cross-products of the original covariates. This assumption is critical. BBMT2022,BBMT2025 determine that Assumption (ref) is necessary for the IV and 2SLS estimands to maintain their interpretation as a convex combination of conditional LATEs.

Let us first consider the case of weak monotonicity. The following result shows that the interpretation of the linear IV estimand is very unappealing in this context.

theoremSuppose that Assumptions (ref), (ref), and (ref) hold. Then \begin{equation*} \beta_{\mathrm{IV}} = \frac{\e \left[ c(X) \cdot \pi(X) \cdot \var \left[ Z \mid X \right] \cdot \tau(X) \right]}{\e \left[ c(X) \cdot \pi(X) \cdot \var \left[ Z \mid X \right] \right]}. \end{equation*}

Theorem (ref) provides a new representation of the IV estimand in the standard specification, i.e. one that, perhaps incorrectly, restricts the effects of the instrument in the reduced-form and first-stage regressions to be homogeneous across covariate values. Unlike in \citetalias{AI1995}'s specification, the estimand in the standard specification is not necessarily a convex combination of conditional LATEs. This is because $c(x)$ takes the value $-1$ for every value of covariates where there exist defiers but no compliers, and hence the corresponding weights in Theorem (ref) are negative as well. It follows that, when IV is applied in the usual way, the estimand may no longer be interpretable as a causal effect. It is even possible that this parameter may be negative (positive) when treatment effects are positive (negative) for everyone in the population.

The following result demonstrates that this problem disappears when we impose the strong version of monotonicity.

corollarySuppose that Assumptions (ref), (ref), and (ref) hold. Then \begin{equation*} \beta_{\mathrm{IV}} = \frac{\e \left[ \pi(X) \cdot \var \left[ Z \mid X \right] \cdot \tau(X) \right]}{\e \left[ \pi(X) \cdot \var \left[ Z \mid X \right] \right]}. \end{equation*}

Corollary (ref) provides a direct argument for AP2009's assertion that the standard specification of IV recovers a convex combination of conditional LATEs. As noted previously, however, this statement is no longer true under weak monotonicity. If strong monotonicity holds, then the weights in Corollary (ref) may be more desirable than those in \citetalias{AI1995}'s specification. Indeed, a comparison of Corollary (ref) and equation ((ref)) shows that the standard specification, like \citetalias{AI1995}'s specification, overweights the effects in groups with large variances of $Z$ but not, unlike the latter, in groups with strong first stages.\footnote{To be clear, both specifications attach a greater weight to conditional LATEs in groups with strong first stages, as required by equation ((ref)). But \citetalias{AI1995}'s specification places even more weight on such conditional LATEs than is necessary to recover the unconditional LATE parameter.}

remarkAbadie2003 shows that, under Assumptions (ref), (ref), and (ref), the IV estimand is equivalent to the coefficient on $D$ in the linear projection of $Y$ on $D$ and $X$ among compliers. In other words, IV is analogous to ordinary least squares (OLS), with the exception of its ability to implicitly condition the analysis on the (latent) subpopulation of compliers. Corollary (ref) provides another argument that “IV is like OLS\@.” Indeed, as shown by Angrist1998, the only difference between the OLS estimand and the ATE is in the dependence of the OLS weights on $\var \left[ D \mid X \right]$\@. Similarly, Corollary (ref) shows that, under strong monotonicity, the only difference between the IV estimand and the LATE is in the dependence of the IV weights on $\var \left[ Z \mid X \right]$\@. However, this analogy between OLS and IV may be problematic for IV given the undesirable properties of the OLS estimand under treatment effect heterogeneity Sloczynski2022.
remarkBWW2007 discuss the interpretation of interacted and noninteracted specifications in randomized experiments with noncompliance in which the existence of defiers is completely ruled out. In this case, the standard specification of IV recovers the unconditional LATE parameter but the interacted specification does not.\footnote{Instead, the interacted specification recovers a convex combination of conditional LATEs, which is generally different from the unconditional LATE parameter. A similar point about models with fully independent instruments is made by HK2020, who also revisits the link between the existence of defiers and negative weights in this context IA1994,deChaisemartin2017,DHM2023 and recommends interacted specifications.} This is a special case of the difference between Theorem (ref) and Corollary (ref) where $\var \left[ Z \mid X \right]$ is constant. However, Theorem (ref) makes it clear that under weak monotonicity the standard specification no longer recovers the unconditional LATE parameter or even a convex combination of conditional LATEs.
remarkTheorem (ref) and Corollary (ref) are also related to Theorem 1 in Kolesar2013, which provides a common representation of any two-step instrumental variables estimand in the case of a binary $D$, a discrete $Z$, and under conditions similar to Assumptions (ref), (ref), and (ref)\@. To present this result, it is necessary to introduce some additional notation. Let $P = \e \left[ D \mid Z,X \right]$, $P^L = \mathrm{L} \left[ D \mid Z_{\mathrm{G}},X \right]$, and $\tilde{P}^L = P^L - \mathrm{L} \left[ D \mid X \right]$, where $\mathrm{L}[\cdot]$ is the linear projection and $Z_{\mathrm{G}} = z_{\mathrm{G}}(X,Z)$ is the vector of constructed instruments, which may include (some) interactions between $X$ and $Z$\@. Also, let $\mathcal{P}_x$ denote the support of $P$ conditional on $X=x$ and $J_x$ denote the number of support points, with $\mathcal{P}_x = \big\{ p_{1,x} < \ldots < p_{J_x,x} \big\}$. Then, Kolesar2013 shows that \begin{equation} \beta_{\mathrm{TSIV}} = \int \sum_{j=1}^{J_x-1} \frac{\theta_j(x)}{\int \sum_{j=1}^{J_x-1} \theta_j(x) \, \mathrm{d} F^X(x)} \, \tau(p_{j,x};x) \, \mathrm{d} F^X(x), \end{equation} where $\beta_{\mathrm{TSIV}}$ is any two-step instrumental variables estimand (e.g., 2SLS) which uses $Z_{\mathrm{G}}$ as instruments, $\theta_j(x) = \left( p_{j+1,x} - p_{j,x} \right) \cdot \pr \left[ P > p_{j,x} \mid X=x \right] \cdot \e \left[ \tilde{P}^L \mid X=x, P > p_{j,x} \right]$, and $\tau(p_{j,x};x) = \frac{\e \left[ Y \, \mid \, P=p_{j+1,x}, \, X=x \right] \; - \; \e \left[ Y \, \mid \, P=p_{j,x}, \, X=x \right]}{p_{j+1,x} \; - \; p_{j,x}}$ is the conditional LATE based on two adjacent elements of $\mathcal{P}_x$. Kolesar2013's result is generic in the sense that it applies to any given vector of instruments $Z_{\mathrm{G}} = z_{\mathrm{G}}(X,Z)$\@. At the same time, Theorem (ref) is specific to the IV estimand. However, its focus on that particular specification simplifies the result, making it more transparent and easier to interpret than equation ((ref)).\footnote{Using equation ((ref)) to determine whether a given specification rules out the incidence of negative weights requires verifying the condition $\pr \left[ \theta_j(X) \geq 0 \right] = 1$ on a case-by-case basis.} In Appendix (ref), I also present an alternative proof of Theorem (ref), which uses Kolesar2013's representation of $\beta_{\mathrm{TSIV}}$.
remarkA testable implication of strong monotonicity is that $\omega(x)$, the conditional first-stage slope coefficient, is always non-negative. In a saturated specification with $X = \left( 1, G_1, \ldots, G_{K-1} \right)$, it is straightforward to construct a formal test based on this observation.\footnote{See also Semenova2025 for an analogous test in the context of endogenous sample selection.} If we define \begin{equation} \omega = \Big( \e \left[ D \mid Z=1, G=k \right] - \e \left[ D \mid Z=0, G=k \right] \Big)_{k=1}^{K}, \end{equation} then the null hypothesis can be written as \begin{equation} H_0 : \quad \left( -1 \right) \cdot \omega \leq 0 \end{equation} and the test statistic as \begin{equation} T = \max_{1 \leq k \leq K} \frac{\left( -1 \right) \cdot \hat{\omega}_k}{\hat{\sigma}_{\hat{\omega}_k}}. \end{equation} One possible choice of critical values for this test statistic are the one-step self-normalized critical values of CCK2019. Another is based on the Bonferroni procedure, which requires, however, that $K$ is much smaller than the sample size.
remarkSuppose we are interested in the estimand of Corollary (ref), but we are only willing to assume weak monotonicity. If $\omega(x)$ were known, we could define a new, “reordered” instrument as $Z_{\mathrm{R}} = 1 [ \omega(X) > 0 ] \cdot Z + 1 [ \omega(X) < 0 ] \cdot \left( 1-Z \right)$ and subsequently use it in a noninteracted specification. In Appendix (ref), I show that this procedure would recover the estimand of interest. In practice, however, $\omega(x)$ is unknown and would need to be estimated. I leave the study of the properties of the resulting reordered IV estimator to future work.

Finite Sample Considerations

Given the theoretical results in Sections (ref) and (ref), it seems reasonable to consider \citetalias{AI1995}'s interacted specification whenever weak monotonicity is plausible but strong monotonicity is not. However, this approach has some limitations in finite samples: it requires dividing the sample into $K$ groups, and when $K$ is sufficiently large relative to the sample size, some groups will be small. With many groups and instruments, this situation leads to bias, which results from overfitting the first stage. In other words, the first-stage fitted values pick up the noise, not just the signal, and a large amount of noise, particularly likely with many small groups, translates to poor estimates of the first stage and bias in the second stage.

This phenomenon, known as the “many instrument” bias, has been extensively studied in the econometrics literature. Recent surveys include Anatolyev2019 and MS2024.\footnote{The classic literature on many instruments has focused on the homogeneous effects model, but I interpret its results through the lens of the framework in Section (ref).} In the remainder of this section, I first review several solutions to this problem, which offer finite sample improvements over 2SLS when estimating specifications with many instruments (e.g., \citetalias{AI1995}'s specification). Then, I review a recent pretest designed to evaluate whether, in a given dataset, the instruments are jointly strong enough to ensure consistency. I conclude with a simulation study.

Estimation with Many Instruments

The problem of the many instrument bias is usually studied using the asymptotic sequence of Kunitomo1980, Morimune1983, and Bekker1994, which allows the number of instruments, $K$, to increase in proportion with the sample size, $N$\@. In the context of \citetalias{AI1995}'s interacted specification, fixing the ratio of $K$ to $N$ does not allow the group sizes to grow when the sample size grows, which reproduces the practical problem of small groups.

Under this asymptotic sequence, 2SLS is inconsistent unless the concentration parameter, a measure of instrument strength, grows faster than the number of instruments. The classic alternatives include the limited information maximum likelihood (LIML) estimator of AR1949 and the bias-corrected two-stage least squares (B2SLS) estimator of Nagar1959, both of which are consistent under homoskedasticity when the concentration parameter grows faster than the square root of the number of instruments CS2005. However, homoskedasticity of first-stage errors is impossible when the treatment is binary. Under heteroskedasticity, LIML and B2SLS require the same (stronger) condition as 2SLS CSHNW2012.

Under heteroskedasticity, the weaker condition that the concentration parameter grows faster than the square root of the number of instruments is sufficient for the consistency of the jackknife IV estimator (JIVE) of AIK1999, as also shown by CSHNW2012. The basic idea underlying jackknife-type estimators is that using a “leave-one-out” predictor of the treatment---effectively a separate first stage for each unit---will reduce the noise and bias.

At the same time, however, most of the estimators discussed so far are inconsistent under the asymptotic sequence that allows the number of covariates, alongside the number of instruments, to increase in proportion with the sample size. This is potentially a major limitation because, in \citetalias{AI1995}'s specification, the number of covariates and the number of instruments are the same and equal to the number of groups. Still, several modifications to JIVE and B2SLS are robust to many instruments and many covariates, including the improved jackknife IV estimator (IJIVE) of AD2009, the modified bias-corrected two-stage least squares (MB2SLS) estimator of Anatolyev2013, the unbiased jackknife IV estimator (UJIVE) of Kolesar2013, and three jackknife-type estimators of CSW2023, referred to as the fixed effect jackknife IV (FEJIV) estimator, the fixed effect limited information maximum likelihood (FELIM) estimator, and the fixed effect Fuller1977 (FEFUL) estimator. Although the performance of LIML is not additionally affected by many covariates Anatolyev2013, both LIML and MB2SLS rely on the homoskedasticity assumption. Furthermore, LIML does not even share the estimand with two-step IV estimators, such as 2SLS, MB2SLS, JIVE, IJIVE, UJIVE, and FEJIV, making it inappropriate in settings with treatment effect heterogeneity Kolesar2013. FELIM and FEFUL do not belong to the class of two-step IV estimators either. Finally, CSW2023 discuss the limitations of IJIVE and, to a lesser extent, UJIVE, making FEJIV the likely estimator of choice.

While the framework of CSW2023 does not explicitly allow for treatment effect heterogeneity, the suitability of the FEJIV estimator in my framework follows from Kolesar2013, who shows that any member of a broad class of two-step IV estimators has a common weighted average representation under treatment effect heterogeneity (cf. Remark (ref)). Because both 2SLS and FEJIV fall into this class, their estimands have the same interpretation under treatment effect heterogeneity and standard asymptotics. The difference is that under many instrument asymptotics, 2SLS becomes inconsistent for this estimand, whereas FEJIV remains consistent.

Weak Identification

Specifications with many instruments require that they are sufficiently strong as a group, although they can be individually weak or even irrelevant Anatolyev2019. In the context of \citetalias{AI1995}'s specification, the original instrument can be weak in some groups as long as it is sufficiently strong in others. But how strong is strong enough?

MS2022 study weak identification in linear models with many instruments, which is a situation where the concentration parameter divided by the square root of the number of instruments remains bounded as the sample size grows. They also develop a pretest for this phenomenon to evaluate whether identification is strong in a given dataset. (Their test statistic $\widetilde{F}$ should be compared to a cutoff of 4.14.) Under the null of weak identification, no consistent estimator exists, and inference can instead be based on a jackknifed version of the AR test statistic. When the pretest rejects, MS2022 recommend the jackknife IV estimator, which is consistent under the alternative CSHNW2012.

Simulations

In what follows, I study the finite sample performance of several two-step IV estimators of \citetalias{AI1995}'s specification, with a focus on settings with many small groups, treatment effect heterogeneity, and violations of Assumption (ref)\@. I adapt the data-generating process from BBMT2022, which was designed to mimic the college proximity study in Card1995. The simulation design also originally assumed homogeneous treatment effects and no monotonicity violations. As we will see, these restrictions are responsible for BBMT2022's conclusion that the usual application of IV is easier to estimate without bias than the interacted specification.

In the baseline data-generating process, as in BBMT2022, I draw $X$ uniformly from a Halton sequence $\mathcal{X}$ on $[0,1]$, subsequently drawing $Z$, $D(Z)$, and $Y(D)$ as \begingroup \allowdisplaybreaks

eqnarray[eqnarray omitted — 213 chars of source]

\endgroup where $(U,V)$ are standard multivariate normal with correlation 0.527, drawn independently of $(X,Z)$\@. I also set $|\mathcal{X}| = 250$, $p(0) = \pr \left[ D=1 \mid Z=0 \right] = 0.22$, and $p(1) = \pr \left[ D=1 \mid Z=1 \right] = 0.29$. In this setting, treatment effects are homogeneous and equal to 1.2. Strong monotonicity is satisfied even though the instrument is relatively weak, with the proportion of compliers independent of $X$ and equal to $p(1) - p(0) = 0.07$. Again, these parameters are calibrated to the data in Card1995.

In subsequent modifications of this data-generating process, I introduce treatment effect heterogeneity by specifying $Y(1)$ and $Y(0)$ as \begingroup \allowdisplaybreaks

eqnarray[eqnarray omitted — 206 chars of source]

\endgroup while also allowing for violations of strong (but not weak) monotonicity. This is accomplished by switching the values of $p(0)$ and $p(1)$ for some groups. Specifically, to generate what I refer to as “moderate” monotonicity violations, I reverse the values of $p(0)$ and $p(1)$ if $X>0.75$. For “large” monotonicity violations, the threshold value of $X$ is 0.5. I also consider a setting with “weak cells,” that is, values of $X$ where the proportion of compliers and defiers is zero. Here, I set $p(0) = p(1) = 0.22$ if $1/3 < X < 2/3$ and reverse the original values of $p(0)$ and $p(1)$ if $X>2/3$.

Two final modifications involve the number and relative sizes of groups and the instrument strength. So far, the groups were equal sized. To reproduce the likely scenario that some groups are large while others are small, I also consider a setting with $|\mathcal{X}| = 20$, but where $X$ is not drawn uniformly. Specifically, I set $\pr \left[ G=k \right]$ to be proportional to $1.3^k$, making the largest group $1.3^{19}$ times larger than the smallest. As in BBMT2022, I also consider a scenario where the instrument is stronger than the “weak” case above, with 0.52 replacing 0.29 as the larger value of $p(Z)$ whenever $p(0) \neq p(1)$ conditional on $X$\@. This sets the conditional proportion of compliers or defiers equal to 0.3, except in the “weak cells” design, where it is either 0.3 or 0.

table[table omitted — 3,520 chars of source]

The total number of simulation designs is sixteen, with $|\mathcal{X}| = 20$ or $|\mathcal{X}| = 250$, two levels of instrument strength (“weak” or “strong”), and four scenarios of violations of strong monotonicity, referred to as no violations, moderate violations, large violations, and violations with weak cells. Treatment effects are homogeneous when strong monotonicity holds and heterogeneous otherwise. The target parameter is the estimand in Theorem (ref), which is, except in the “weak cells” design, equal to that in Corollary (ref), making monotonicity violations the only reason why the estimands of the interacted and noninteracted specifications may be different. I consider two sample sizes, $N=3{,}000$ and $N=10{,}000$, when $|\mathcal{X}| = 20$, and additionally $N=50{,}000$ when $|\mathcal{X}| = 250$. The smallest sample size, $N=3{,}000$, is similar to the sample size in Card1995.

table[table omitted — 3,516 chars of source]

Table (ref) reports simulation results for a number of estimators in the “weak” IV case with 250 groups and no monotonicity violations. The first three columns, setting $N=3{,}000$, correspond to the baseline results in BBMT2022. Even though I consider a larger number of estimators than BBMT2022, I reach the same conclusion: all estimators are severely biased, with the only exception of IV in the noninteracted specification, whose bias is less than 10% and median bias is practically zero. However, panel B of Table (ref) reveals that this conclusion is predictable: the average value of MS2022's test statistic, $\widetilde{F}$, is 1.83, well below the cutoff of 4.14, which means that consistent estimation of the interacted specification is impossible. The remaining columns report simulation results for $N=10{,}000$ and $N=50{,}000$. Here, the strength of identification gradually increases, with the average value of $\widetilde{F}$ exceeding 10 when $N=50{,}000$. Indeed, when this is the case, the best-performing estimators of \citetalias{AI1995}'s specification---IJIVE, UJIVE, and FEJIV---are practically unbiased, in line with the results in MS2022.

Table (ref) introduces moderate monotonicity violations. With $N=3{,}000$, the average value of $\widetilde{F}$ is again below 2. Now, however, every estimator is severely biased, including IV in the noninteracted specification. (Estimation of this specification is biased because of monotonicity violations. Estimation of the interacted specification is biased because of insufficient instrument strength.) With larger sample sizes, $N=10{,}000$ and $N=50{,}000$, identification gets stronger. Specifically, when $N=50{,}000$, the average value of $\widetilde{F}$ again exceeds 10, and IJIVE, UJIVE, and FEJIV perform very well. IV estimation of the noninteracted specification remains biased; however, it is competitive with the best-performing estimators in terms of MSE\@.

table[table omitted — 3,510 chars of source]
table[table omitted — 3,533 chars of source]

Tables (ref) and (ref) consider large monotonicity violations and “weak cells.” It remains the case that IJIVE, UJIVE, and FEJIV are nearly unbiased whenever the average value of $\widetilde{F}$ is large enough. This includes the “weak cells” design in Table (ref), which underscores the notion that the instrument can be weak in some groups as long as it is sufficiently strong in others.\footnote{Intuitively, if $\pi(x)=0$ when $X=x$, $\tau(x)$ is not identified. However, because $\tau_{\mathrm{LATE}} = \frac{\e \left[ \pi(X) \cdot \tau(X) \right]}{\e \left[ \pi(X) \right]}$, the weight on $\tau(x)$ in $\tau_{\mathrm{LATE}}$ would have been zero anyway, and analogously for the estimands in Theorem (ref), Theorem (ref), and Corollary (ref). That is, as long as the overall instrument strength is sufficient MS2022, it does not matter that some conditional LATEs cannot be well estimated due to a conditional-on-$X$ weak IV problem, because those conditional LATEs are irrelevant for the target estimand.} On the other hand, unlike in Table (ref), IV estimation of the noninteracted specification is not only biased in Tables (ref) and (ref), but also noisy, which leads to very high values of MSE\@.

The remaining simulation results, for “weak” IV with $|\mathcal{X}| = 20$ and for “strong” IV with both values of $|\mathcal{X}|$, are reported in Tables (ref)--(ref) in Appendix (ref)\@. The bottom line is still that MS2022's pretest does a great job differentiating between cases where IJIVE, UJIVE, and FEJIV perform well or very well, and cases where all estimators of the interacted specification perform badly. Roughly speaking, values of $\widetilde{F}$ exceeding the recommended cutoff of 4.14 are associated with low bias, even when, with $|\mathcal{X}| = 250$ and $N=3{,}000$, there are only 12 units in each group; values of $\widetilde{F}$ exceeding 10--15 are associated with negligible or no bias, at least in the data-generating process under consideration.

Other estimators are clearly not competitive with IJIVE, UJIVE, and FEJIV\@. When there are violations of monotonicity, the usual application of IV is biased and often unstable. 2SLS estimation of the interacted specification is generally biased, as expected. MB2SLS is usually dominated by IJIVE, UJIVE, and FEJIV, especially on bias. JIVE is generally biased and unstable.

To be clear, the purpose of this simulation study is not to claim that \citetalias{AI1995}'s specification can be estimated without bias in a typical application of IV methods. Instead, the simulations show that IJIVE, UJIVE, and FEJIV estimation of \citetalias{AI1995}'s specification is reliable if the instruments are jointly strong enough, which can be verified using MS2022's pretest. Future research should examine whether the number of instruments in \citetalias{AI1995}'s specification could be reduced using appropriate regularization techniques, perhaps a modification of the existing approaches in CHS2015a,CHS2015b and Wiemann2024.

Empirical Applications

The results so far underscore the importance of using the interacted specification when weak monotonicity is plausible but strong monotonicity is not. In this section I present evidence of violations of strong monotonicity and first-stage homogeneity in a sample of recent applications of IV methods. Then, I revisit a study of the effects of pretrial detention on case outcomes in Philadelphia, where violations of strong monotonicity are particularly evident Stevenson2018.

Review of Applications of Instrumental Variables

In what follows, I use a sample of 1,309 instrumental variables regressions previously analyzed by Young2022, which corresponds to the universe of IV estimates reported in the main text of 30 papers published in journals of the American Economic Association between 2006 and 2015.\footnote{Young2022's goal was to cover the universe of replicable IV applications in this period subject to a small number of additional inclusion criteria reported in his paper.} After dropping specifications with multiple instruments, without additional covariates, or based on panel data, I obtain my final sample of 988 regressions in 25 papers.\footnote{Because Young2022 only considered papers with replication code in Stata, I define “specifications based on panel data” as those using Stata's xtivreg or xtivreg2 commands in the original replication package. The number of applicable regressions in several papers would decrease substantially if we eliminated not only duplicate IV regressions---which Young2022 already did---but also duplicate first stages. However, my preliminary attempt to do so did not meaningfully change any of the results reported in this section.} The number of regressions per paper is highly uneven in this sample, with the mean equal to $988/25 = 39.52$ and the quartiles equal to 8, 14, and 40. The list of papers under consideration is provided in Appendix (ref)\@.

Given the inclusion criteria above, every specification in my sample is based on a linear first-stage regression of $D$ on $Z$ and $X$, without any interactions between $Z$ and $X$\@. In my first exercise, I implicitly include these interactions by means of separate regressions of $D$ on $X$ given $Z=1$ and $Z=0$.\footnote{If the original treatment or instrument are not binary, I replace them with indicators for whether these variables are above their medians. I demonstrate robustness to other binarizations in Tables (ref) and (ref) in Appendix (ref)\@.} This simple approach allows me to estimate the conditional first stage at every value of $X$ as the difference in conditional means, $\hat{\omega}(x) = \hat{\e} \left[ D \mid Z=1, X=x \right] - \hat{\e} \left[ D \mid Z=0, X=x \right]$. Subsequently, I report the fraction of these estimates that are opposite in sign (“negative”) to the estimate in the original first stage, which is equivalent to the fraction of observations with negative weights in the usual application of IV (cf. Theorem (ref)). This is analogous to the recommendation of dCDH2020 to report the fraction of units with negative weights in two-way fixed effects regressions. Similarly, Semenova2025 reports the fraction of observations with negative predictions in a sample selection context related to mine.

Panel A of Table (ref) indicates that negative first stages are a common occurrence in recent applications of IV\@. The average fraction of observations with a negative first stage is 21.8% when using the linear probability model (LPM) to estimate the conditional means in $\omega(x)$ and 17.6% when using the probit model. After weighting by the inverse of the number of applicable regressions associated with a given paper, these averages increase to 28.5% and 28.0%, respectively, giving the average of the within-paper averages.

It may be the case that a portion of the estimated negative first stages is due to noise. However, the regressions in my sample are usually not saturated, which means that the formal test of violations of monotonicity in Remark (ref) is not appropriate. Instead, in my second exercise, I explicitly add interaction terms to each original (linear) first stage and test whether the corresponding coefficients are jointly equal to zero. Under the alternative, the true first stage is heterogeneous, which is a necessary condition for strong monotonicity being false but weak monotonicity being true.

table[table omitted — 2,122 chars of source]

Panel B of Table (ref) reports the results of this exercise. Using the Bonferroni procedure to account for multiple hypothesis testing separately for each paper, I conclude that 22 of 25 papers have at least one first stage that is heterogeneous. Using the Holm correction, I reject an average of 71.5% of homogeneous first stages per paper. The last column demonstrates that these conclusions are robust to using the probit instead of the linear probability model (LPM)\@.\footnote{The smaller number of papers under consideration when using the probit model reflects convergence and other estimation problems in the missing specifications.}

Reanalysis of Stevenson2018

Now, I turn to a reanalysis of Stevenson2018's study of the effects of pretrial detention on case outcomes. In this application, recently reanalyzed by Cunningham2021, CHMW2024, and MT2024, violations of strong monotonicity are evident, which I will be able to formally demonstrate.

table[table omitted — 1,017 chars of source]

The data are based on the Philadelphia court records and cover 331,971 arrests between 2006 and 2013. The “treatment” of interest is pretrial detention or, in other words, whether the defendant was incarcerated in the period between their arrest and disposition; the purpose of such detention is that they appear in court and do not commit another crime. The empirical question is whether pretrial detention has a causal effect on case outcomes, such as conviction and incarceration length. Naturally, pretrial detention is endogenous, and Stevenson2018's identification strategy is based on random assignment of bail magistrates (judges) to cases. These judges have broad authority to set bail---the amount required for pretrial release---at a level they choose. Thus, being assigned a strict judge makes the defendant less likely to be able to pay bail and more likely to be detained.

In Philadelphia, bail hearings usually last one or two minutes, which made it possible for only eight judges to hear all the cases in Stevenson2018's data. Table (ref) reports the number of cases and the detention rate for each judge. The magistrate I refer to as “Judge C” is the most lenient, with a relatively low detention rate of 0.395. In what follows---unlike Stevenson2018, who uses the full set of judge indicators as instruments---I focus on a single instrument defined as whether a given case was heard by Judge C\@. A simple regression of pretrial detention on the “Judge C” dummy reveals a first stage of --0.0195 with a standard error of 0.0023.

In the present context, strong monotonicity requires that every defendant detained by Judge C would also have been detained by other judges. However, this condition seems implausible, with the likely dimensions of monotonicity violations including the offense type Stevenson2018 and the defendant's race ABM2012. As in Stevenson2018, I focus on the seventeen most common offense types.\footnote{These include drug possession, drug sale, aggravated assault, robbery, first offense DUI, simple assault, drug purchase, burglary, shoplifting, theft, marijuana possession, murder, motor vehicle theft, prostitution, third-degree felony firearm possession, second-degree felony firearm possession, and vandalism.} I also consider three racial categories: Black, White, and other. The offense types are not mutually exclusive, which means that, in principle, the sample could be divided into $3 \cdot 2^{17}$ groups based on the defendant's race and the offense type. However, most of these groups are empty, and I also drop nonempty groups with fewer than three cases heard by Judge C or not heard by Judge C\@. As a result, for this specification, the final sample consists of 431 groups and 327,560 cases.

With such a saturated specification, a formal test of violations of strong monotonicity is straightforward, as discussed in Remark (ref). Given that Judge C is more lenient than others, the overall first stage is negative---not positive, as assumed previously---and the null hypothesis requires that all the group-specific first stages are also non-positive. In other words, the null can be written as

equation[equation omitted — 38 chars of source]

while the test statistic equals

equation[equation omitted — 113 chars of source]

When I implement this test, I obtain a test statistic of 5.637 and a $p$-value of 3.7e-06, despite accounting for multiple hypothesis testing.\footnote{In this application, the approach of CCK2019 produces almost identical critical values as the Bonferroni procedure.} In this application, strong monotonicity is clearly rejected. Further details on the group-specific impact of Judge C on pretrial detention are provided in Table (ref). Because presenting estimates for 431 groups is impractical, I restrict my attention to twenty groups with the largest (most positive) and ten groups with the smallest (most negative) $z$ statistics. For each group, I report the number of cases, the conditional first stage and its standard error, and the corresponding Holm $p$-value. At any conventional significance level, we can reject that the first stage is non-positive in two groups: defendants charged with burglary and vandalism who are neither Black nor White and Black defendants charged with robbery, simple assault, and theft. In general, a common feature of many of the groups with the largest $z$ statistics is a combination of being charged with a property crime (e.g., theft or burglary) and a violent crime (e.g., simple assault or aggravated assault). In fact, seven of these groups are charged with robbery, which is simultaneously a violent crime and a crime against property.\footnote{This is consistent with Stevenson2018's account that “[t]he magistrate that is most lenient overall is actually strictest when it comes to robbery.” However, the test discussed in Remark (ref) and the results in Table (ref) are otherwise different from the analysis in Stevenson2018.} Many of these groups comprise of defendants who are neither Black nor White. On the other hand, the groups with the smallest $z$ statistics are universally Black or White, and charged with nonviolent crimes.

table[table omitted — 4,304 chars of source]

Because strong monotonicity is rejected, the noninteracted specification cannot be used to estimate a convex combination of conditional LATEs (cf. Theorem (ref)). However, if weak monotonicity is plausible, the interacted specification will be appropriate, at least as long as the interacted instruments are sufficiently strong. In the present context, weak monotonicity seems quite sensible. Given that bail hearings in Philadelphia are extremely short, it is unlikely that more than a handful of factors---such as the offense type and the demographic characteristics of the defendant---could determine the amount of bail and the resulting likelihood of detention.

To incorporate additional factors into the analysis, I also consider two alternative specifications. First, I define the groups based on the offense type, the defendant's race, and their gender (male or female). In theory, the number of groups could be as large as $3 \cdot 2^{18}$ in this specification, but only 563 groups remain after I drop those that are empty or otherwise too small---requiring, as above, that there are at least three observations for every $(G,Z)$ combination. Second, I define the groups based on the offense type, the defendant's race and gender, and three time periods considered by Stevenson2018. The relevance of these specific time periods---divided by February 23, 2009 and February 23, 2011---results from concurring changes in the composition of magistrates other than Judge C\@. This sets the maximum number of groups in this specification at $3^2 \cdot 2^{18}$; in practice, the number of groups that are nonempty and sufficiently large is 981.

Table (ref) reports the main results of my analysis. In panels C and D, for each of the three specifications described above, I report the Bonferroni/CCK2019 $p$-value for the test of violations of strong monotonicity as well as MS2022's test statistic, $\widetilde{F}$, for the test of weak identification. The test results leave little doubt that strong monotonicity is violated while identification is strong. The $p$-values for the former test never exceed 0.0015, despite accounting for simultaneous testing of up to 981 hypotheses.\footnote{In the second and third specifications, the largest $z$ statistic is obtained in very small groups, which makes the normal approximation questionable. However, the rejection of strong monotonicity remains solid. The smallest Holm $p$-values in groups with at least 100 cases are 0.0109 and 0.0105 in the second and third specifications, respectively. The corresponding smallest Holm $p$-values in groups with at least 500 cases are 0.0408 and 0.0105. Note that these $p$-values are conservative, because they implicitly penalize hypothesis testing in groups smaller than 100 or 500 cases, even though such groups are ignored in this context.} The values of $\widetilde{F}$ range between 19.32 and 21.56. If the simulations in Section (ref) are any guide, we should expect negligible bias when estimating the interacted specification, at least when using the jackknife-type estimators such as IJIVE, UJIVE, and FEJIV\@.

Panels A and B of Table (ref) report OLS, IV, 2SLS, IJIVE, UJIVE, and FEJIV estimates of the effects of pretrial detention on conviction and incarceration length. The noninteracted specification, marked as “IV,” suggests that pretrial detention leads to a 17--19 p.p. increase in the likelihood of being convicted and an increase in incarceration length of 670--720 days. Such effects would have been substantial, but the validity of these estimates is questionable given the clear rejection of strong monotonicity in this application. When we turn to the interacted specification, the estimates become smaller. The effects on conviction are closer to zero---in the range of 4 to 15 p.p.---but often remain significant.\footnote{In the case of FEJIV, I report the standard errors derived by CSW2023. Practitioners should also consider a recent alternative proposed by BN2024, which explicitly accounts for treatment effect heterogeneity.} The effects on incarceration length are much smaller than in the noninteracted specification and suggestive of an effect of 50--160 days. These estimates are also usually not significantly different from zero, except for 2SLS\@. For both outcomes and each specification, the IJIVE, UJIVE, and FEJIV estimates are practically indistinguishable from each other but also clearly different from the 2SLS estimate.

table[table omitted — 5,035 chars of source]

My conclusions are generally in line with Stevenson2018, whose paper includes a relatively rare recent example of using specifications with interacted instruments (cf. Remark (ref)), although not of \citetalias{AI1995}'s interacted specification, which is implemented here. I also provide additional results in Appendix (ref)\@. Table (ref) reports MB2SLS and JIVE estimates of the effects of pretrial detention. While the MB2SLS estimates are largely similar to the results in Table (ref), the JIVE estimates are noisy and appear unreliable. Table (ref) reports the results of a bootstrap test for comparisons between the IV estimates in the noninteracted specification and various estimates of the interacted specification.\footnote{I perform this test for 2SLS, MB2SLS, JIVE, and UJIVE, but not IJIVE and FEJIV, because the latter estimators are very computationally demanding in the specifications that I consider, at least in my implementation.} At the 5% level, I nearly always reject the null that the estimands are the same in the case of incarceration length but not conviction. Finally, Table (ref) reports the results of a similar bootstrap test for comparisons between 2SLS and other estimators of the interacted specification. These differences are often highly statistically significant, which reaffirms the importance of correcting for the many instrument bias.

Conclusion

In this paper I studied the interpretation of linear IV and 2SLS estimands when both the treatment and the instrument are binary, and when additional covariates are required for identification. Using the LATE framework of IA1994 and AIR1996, I argued that the common practice of interpreting standard IV estimands as a convex combination of conditional LATEs, or even as the (unconditional) local average treatment effect, is substantially more problematic than previously thought. I showed that the interpretation of the usual application of IV, which limits the effects of the instrument in the reduced-form and first-stage regressions to be homogeneous, hinges critically on the specific variant of the monotonicity assumption that the researcher is willing to entertain. Under “weak monotonicity,” some of the IV weights may be negative and the IV estimand may no longer be interpretable as a causal effect.

What should applied researchers do in practice? In this paper I argued that it might be worthwhile to revisit the interacted specification of AI1995, which is guaranteed to eliminate negative weights under the same assumptions that are problematic for the usual application of IV\@. Specifications with many interacted instruments were used in influential papers by Angrist1990 and AK1991 but appear to have been largely abandoned in subsequent work out of concern for the many instrument bias. Unsurprisingly, however, the modern tools to estimate such specifications are substantially better than in the 1990s, as I also demonstrate in an extensive simulation study. A pretest for weak identification developed by MS2022 can be used to determine whether consistent estimation of the interacted specification is possible. When the pretest rejects, several jackknife-type estimators can be used, including the FEJIV estimator of CSW2023, which I also implement in the companion MATLAB, R, and Stata packages, fejiv.

There are at least two important situations when this recommendation will not be satisfactory. First, in some applications in which strong monotonicity is rejected, weak monotonicity will be implausible, too. If this is the case, it may be worthwhile to instead consider the partial identification approach of Noack2021, which evaluates the sensitivity of what can be learned about the local average treatment effect under violations of (weak) monotonicity. Second, a convex combination of conditional LATEs, which AI1995's specification is guaranteed to produce under weak monotonicity (and the usual application of IV under strong monotonicity), may be considered an imperfect substitute for the (unconditional) local average treatment effect.\footnote{A similar argument is made by CGBS2024 in the context of difference-in-differences designs.} If this is the case, there are many existing estimators that are consistent for the LATE under strong monotonicity (see, e.g., SUW_kappa, and the references therein). An important avenue for future research is to develop estimators of the LATE that are also robust to weak monotonicity.

\singlespacing

\setlength\bibsep{0pt}