EconBase
← Back to paper

A Practical Guide to Instrumental Variables Methods with Heterogeneous Treatment Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,787 characters · 10 sections · 148 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Practical Guide to Instrumental Variables Methods with Heterogeneous Treatment Effects

\onehalfspacing

Instrumental variables (IV) methods represent one of the most popular approaches to identifying and estimating causal effects in economics. CKZ2020 establish that the popularity of IV methods in applied microeconomics has grown continuously since the 1980s, as indicated by the content of the National Bureau of Economic Research (NBER) working papers. Likewise, GP2026 documents that roughly 30% of NBER working papers in applied microeconomics mention instrumental variables.

Standard textbook treatments of IV methods assume a linear model with constant effects. At the same time, the influential local average treatment effect (LATE) framework of IA1994, AI1995, and AIR1996 allows for a very general form of treatment effect heterogeneity, which the textbook model rules out. In this paper, we will offer a nontechnical, practical guide to the literature on IV methods with heterogeneous treatment effects, which originates from the work of AI1995. Instead of aiming to provide a comprehensive survey of the literature, our focus will be on highlighting several areas in which existing work in applied microeconomics deviates from what we regard as best practices motivated by the recent theoretical literature.

The first area we explore is the choice of the target parameter. Recent research highlights possible differences between the local average treatment effect (LATE), as defined by IA1994, and the probability limits of usual IV and two-stage least squares (2SLS) estimators, especially in cases where the instrument is only valid conditional on covariates. As an alternative, we will discuss an intuitive, general strategy to estimate the LATE when covariates matter. The second area we consider is the flexibility of parametric specifications. Recent research shows that a causal interpretation of IV and 2SLS estimands requires that the conditional mean of the instrument given covariates be linear. Parametric misspecification is also possible when estimating the LATE with covariates. As a solution to such problems, we will discuss flexible estimation strategies based on machine learning. The third area we explore is possible violations of the assumptions underlying the LATE framework. These assumptions have testable implications that have spurred a sizable theoretical literature, yet have had limited influence on applied work. We will discuss the intuition behind the resulting tests and their implementation. We will also discuss estimation approaches that are more robust to violations of monotonicity, which is often a controversial assumption.

Throughout the paper, we focus on the standard LATE framework with a binary treatment and a binary instrument as the leading case. Interested readers should also consider several other papers that offer a comprehensive survey of the existing literature, or focus instead on a subset of possible applications of IV methods. MT2018 use the marginal treatment effect (MTE) framework of HV2005 to discuss IV identification and extrapolation under treatment effect heterogeneity. MT2024 review the literature on IV methods with heterogeneous treatment effects, including extensions to multivalued treatments, multivalued instruments, and multiple instruments. BHJ2025 provide a guide to the literature on shift-share instruments. CFL2025 and GPHK2025 survey the literature on “judge leniency” designs, which is another important application of IV methods.

Target Parameters

In this section we discuss the differences between three leading estimands (target parameters) in applications of IV methods: the probability limit in the usual linear IV regression, the 2SLS estimand in a saturated specification with multiple interacted instruments, and the “true” local average treatment effect. We argue that this last parameter should be more commonly estimated in relevant applied work---ideally as the main parameter of interest or at least as a robustness check. We also illustrate the possible differences between these parameters in two empirical applications.

Motivating Example: Stratified RCTs with Imperfect Compliance

We start by considering a leading application of instrumental variables methods: a randomized controlled trial (RCT) with imperfect compliance. Concretely, let $Y_i$ denote the outcome, $D_i$ denote the binary treatment, and $Z_i$ denote the randomized assignment to treatment. We will refer to $Z_i$ as the instrument, following standard econometrics terminology. We want to leverage the randomness in $Z_i$ to estimate the causal effect of $D_i$ on $Y_i$. Let $D_i(1)$ and $D_i(0)$ denote the two potential treatments, i.e., the counterfactual treatment statuses that would be observed if a unit were assigned ($Z_i=1$) or not assigned ($Z_i=0$) to be treated. (Observed and assigned treatment statuses may differ due to imperfect compliance.) Following the usual terminology, we refer to units with $D_i(1)>D_i(0)$ as compliers. Under the standard assumptions underlying the local average treatment effect framework---namely, independence of the instrument, exclusion restriction, relevance, and monotonicity---the usual linear IV regression identifies the LATE, i.e., the average treatment effect for compliers IA1994. This is the so-called Wald estimand:

displaymath\beta \; = \; \frac{ \ensuremath{\mathbb{E}}(Y_i \mid Z_i=1) - \ensuremath{\mathbb{E}}(Y_i \mid Z_i=0) }{ \ensuremath{\mathbb{E}}(D_i \mid Z_i=1) - \ensuremath{\mathbb{E}}(D_i \mid Z_i=0) } \; = \; \ensuremath{\mathbb{E}}[Y_i(1)-Y_i(0) \mid D_i(1)>D_i(0)],

where $Y_i(1)$ and $Y_i(0)$ are potential outcomes.

Treatment assignment, however, is rarely completely random. Consider a stratified RCT where the assignment is only random once we condition on an appropriate set of covariates, $X_i$. For example, in the Oregon Health Insurance Experiment (OHIE), a randomized lottery that assigned low-income adults the opportunity to apply for Medicaid, it is necessary to control for household size and survey wave Finkelsteinetal2012. The independence assumption, $(Y_i(1),Y_i(0), D_i(1), D_i(0)) \perp Z_i$, no longer holds unconditionally because these covariates may be related to potential outcomes and treatments as well as the probability of treatment assignment. Specifically, household size had a direct effect on the chance of winning this particular lottery, while, at the same time, it is plausibly related to health expenditures, health care utilization, and baseline health status. Similarly, survey wave captures timing effects that may have influenced who was reached by the survey, while it also picks up time-varying shocks to outcomes. Consequently, $Z_i$ is no longer independent of potential outcomes and treatments unless we control for these particular covariates.

There is not, however, a unique method to control for covariates, and this makes the LATE framework with covariates substantially more complicated than the baseline model without covariates. In what follows, we discuss three approaches to control for covariates in stratified RCTs with imperfect compliance. First, consider a linear IV regression, a population regression specification where covariates, $X_{i}$, are controlled in an additively separable fashion in the outcome equation,

align[align omitted — 99 chars of source]

as well as in the first-stage regression,

align[align omitted — 81 chars of source]

In the case of the OHIE, the elements of $X_{i}$ are discrete and consist of mutually exclusive indicators for different combinations of values of household size and survey wave. The method of analysis in Finkelsteinetal2012 is equivalent to estimating $\beta_{\text{IV}}$ using these covariates. To simplify the comparison with other approaches, we assume, for now, that $X_{i}=(X_{i1},...,X_{iJ})$ does indeed saturate the model. That is, $X_{i}$ is a vector of dimension $J$, each entry $X_{ij}$ is a dummy variable, and for every observation $i$, $\sum_j X_{ij} = 1$.

Second, consider a 2SLS regression, a population regression specification where the outcome equation is still specified as additively separable but covariates are interacted with the instrument in the first-stage regression:

align[align omitted — 171 chars of source]

AP2009 refer to this specification as the “saturate and weight” approach. Sloczynski2025 calls it “AI's specification,” after AI1995. With saturated covariates, the first-stage regression in equation (ref) is correctly specified by design, and each $\pi_j$ may be interpreted as the causal effect of $Z_i$ on $D_i$ for the subgroup with covariate $X_{ij}=1$, which is the same as the proportion of compliers in that subgroup.\footnote{We return to the issue of misspecification in the next section.}

Although equation (ref) is correctly specified, equation (ref) still imposes a constant effect of $D_{i}$ on $Y_{i}$. Thus, the third approach we consider is grounded in a nonparametric identification result rather than a specific estimator. As shown by Frolich2007, the average treatment effect for compliers is identified as

align[align omitted — 337 chars of source]

which means that the LATE is the ratio of the average treatment effect of the instrument on the outcome and the average treatment effect of the instrument on the treatment. (As we will see below, $\beta_{\text{IV}}$ and $\beta_{\text{AI}}$ are not generally equal to the LATE\@.) In practice, this ratio can be estimated by a variety of approaches, which consist of using one's favorite estimator of the average treatment effect twice---once to compute the numerator and again to compute the denominator of the expression in (ref). With saturated covariates, this is particularly straightforward, because many standard estimators will be numerically identical and equivalent to a procedure that consists of computing the effect of the instrument on the outcome and the treatment in each covariate cell (using simple differences in means), and aggregating them together using the proportions of observations in those cells.

Under the LATE assumptions generalized to the case with covariates, as in Abadie2003, we can derive a weighted average representation of the three estimands in terms of covariate-specific LATEs. Denote by $\tau_j = \ensuremath{\mathbb{E}}[ Y_{i}(1)-Y_{i}(0) \mid D_{i}(1)>D_{i}(0), X_{ij} = 1 ]$ the average treatment effect for compliers with covariate level $j$ and by $p_{j} = \ensuremath{\mathbb{P}}( X_{ij} = 1 )$ the share of individuals with that covariate level. From Sloczynski2025, we have

equation[equation omitted — 244 chars of source]

and

equation[equation omitted — 245 chars of source]

where the representation in (ref) is also implied by Theorem 3 in AI1995. Also, from Frolich2007, we have

equation[equation omitted — 205 chars of source]

As is clear from these expressions, the average treatment effect for compliers, $\beta_{\text{LATE}}$, may be substantially different from $\beta_{\mathrm{IV}}$, $\beta_{\mathrm{AI}}$, or both---as long as $\pi_j$, the conditional proportion of compliers, or $\text{Var} \left[ Z_{i} \mid X_{ij} = 1 \right]$, the conditional variance of the instrument, vary across covariate values.\footnote{The interpretation of $\beta_{\text{LATE}}$, $\beta_{\mathrm{IV}}$, and $\beta_{\mathrm{AI}}$ will change when the monotonicity assumption is violated. We return to this issue below.} Relative to the LATE, $\beta_{\mathrm{IV}}$ overweights the covariate cells with large values of $\text{Var} \left[ Z_{i} \mid X_{ij} = 1 \right]$, while $\beta_{\mathrm{AI}}$ overweights those with large values of the product of $\pi_j$ and $\text{Var} \left[ Z_{i} \mid X_{ij} = 1 \right]$ Sloczynski2025. In a model with homogeneous treatment effects, such weighting may be desirable in practice due to improved estimation efficiency, but with heterogeneous treatment effects, it would be difficult to substantiate an explicit interest in $\beta_{\mathrm{IV}}$ or $\beta_{\mathrm{AI}}$. That being said, the weighted average representations in (ref)--(ref) only imply that the three estimands may be substantially different from each other, but it is an empirical question whether they will actually differ in a particular application.

To illustrate this empirical question, we use data from Finkelsteinetal2012. In this setting, as mentioned above, the instrument (the ability to apply for Medicaid) is randomized conditional on household size and survey wave, and the identifying assumptions underlying the LATE framework seem plausible. In our subsequent analysis, we focus on the impact of Medicaid coverage on the number of outpatient visits in the last six months. To simplify the illustration, unlike Finkelsteinetal2012, we do not use survey weights.

Panel A of Table (ref) reports the estimates of $\beta_{\text{LATE}}$, $\beta_{\mathrm{IV}}$, and $\beta_{\mathrm{AI}}$ using the full sample from the OHIE\@. Although there is substantial heterogeneity in covariate-specific LATEs, the different estimands yield nearly identical point estimates and standard errors, suggestive of a precisely estimated increase in the number of outpatient visits in response to Medicaid enrollment by about 1 in six months on average. Apparently, the correlation between $\widehat{\text{Var}} \left[ Z_{i} \mid X_{ij} = 1 \right]$ and $\hat{\tau}_j$ as well as $\hat{\pi}_j \cdot \widehat{\text{Var}} \left[ Z_{i} \mid X_{ij} = 1 \right]$ and $\hat{\tau}_j$ is sufficiently weak among compliers to render the differences between $\hat{\beta}_{\mathrm{LATE}}$, $\hat{\beta}_{\mathrm{IV}}$, and $\hat{\beta}_{\mathrm{AI}}$ meaningless.

table[table omitted — 1,162 chars of source]

The picture is very different, however, when we restrict our attention to the subsample of three-person households in Panel B of Table (ref). This restriction does not alter the identifying assumptions, i.e., randomization continues to hold conditional on survey wave, but it changes the practical importance of the different components in the weighted average representations in (ref)--(ref). Indeed, the estimated LATE, 4.022, is now substantially larger than $\hat{\beta}_{\mathrm{IV}}$ and $\hat{\beta}_{\mathrm{AI}}$, which are equal to 2.790 and 2.268, respectively.

The differences between $\hat{\beta}_{\mathrm{LATE}}$, $\hat{\beta}_{\mathrm{IV}}$, and $\hat{\beta}_{\mathrm{AI}}$ follow directly from their distinct weighting schemes. The estimated LATE aggregates covariate-specific treatment effects in proportion to each cell's estimated share among the compliers, $\hat{p}_j \hat{\pi}_j / \sum_k \hat{p}_k \hat{\pi}_k$. By contrast, the IV and 2SLS estimates additionally weight each cell by the conditional variance of the instrument, $\widehat{\text{Var}} \left[ Z_{i} \mid X_{ij} = 1 \right]$, and the latter again by $\hat{\pi}_j$. As a result, these estimates place relatively more weight on covariate cells in which treatment effects can be estimated more precisely, which need not coincide with those that are most representative of the complier population.

Table (ref) illustrates this mechanism. For each of the three covariate cells in the subsample of three-person households, we report its sample proportion, $\hat{p}_j$, the conditional variance of $Z_{i}$, as well as the estimated proportion of compliers, $\hat{\pi}_j$, and conditional LATE, $\hat{\tau}_j$. We also report the estimated weights underlying each of the three estimands, which are simply the sample counterparts of the expressions in (ref)--(ref). For example, $\hat{w}_{\text{LATE},j} = \hat{p}_j \hat{\pi}_j / \sum_k \hat{p}_k \hat{\pi}_k$ and $\hat{w}_{\text{IV},j} = \hat{p}_j \hat{\pi}_j \cdot \widehat{\text{Var}} \left[ Z_{i} \mid X_{ij} = 1 \right] / \sum_k \hat{p}_k \hat{\pi}_k \cdot \widehat{\text{Var}} \left[ Z_{i} \mid X_{ik} = 1 \right]$. In the first covariate cell, $\hat{w}_{\text{LATE},1} = \left( 0.241 \cdot 0.455 \right) / \left( 0.241 \cdot 0.455 + 0.448 \cdot 0.280 + 0.310 \cdot 0.250 \right) \approx 0.351$ and $\hat{w}_{\text{IV},1} = \left( 0.241 \cdot 0.168 \cdot 0.455 \right) \; / \; \left( 0.241 \cdot 0.168 \cdot 0.455 + 0.448 \cdot 0.037 \cdot 0.280 + 0.310 \cdot 0.099 \cdot 0.250 \right) \approx 0.600$. (All values are rounded for illustration.) Each of the estimates reported in Panel B of Table (ref) can be obtained as the dot product of $\hat{\tau}_j$ and the respective weights. For example, $\hat{\beta}_{\mathrm{LATE}} = 0.351 \cdot 1.067 + 0.401 \cdot 6 + 0.248 \cdot 5 \approx 4.022$ and $\hat{\beta}_{\mathrm{IV}} = 0.600 \cdot 1.067 + 0.151 \cdot 6 + 0.249 \cdot 5 \approx 2.790$. In this subsample, because $\widehat{\text{Var}} \left[ Z_{i} \mid X_{ij} = 1 \right]$ and $\hat{\pi}_j$ are large when $\hat{\tau}_j$ is small, $\hat{\beta}_{\mathrm{IV}}$ and $\hat{\beta}_{\mathrm{AI}}$ underweight the covariate cells with large estimated treatment effects, producing a downward bias relative to the LATE\@. As a result, the three estimands answer different causal questions despite relying on the same experimental variation.\footnote{One way to see this is through the lens of the framework developed by PS2025 to quantify the internal validity and representativeness of various estimands. If we treat the population of compliers---the largest subpopulation for which the average treatment effect is nonparametrically identified in an IV setting under standard assumptions---as our target, then $\beta_{\text{LATE}}$ is associated with the largest possible value of the measure of internal validity, equal to 1, while the measures associated with $\beta_{\mathrm{IV}}$ and $\beta_{\mathrm{AI}}$ will generally be less than 1. This means that these estimands correspond to average treatment effects for some (possibly small) subsets of the complier subpopulation, whereas $\beta_{\text{LATE}}$ is the average over all compliers.} A cautious researcher should investigate whether a similar pattern arises in their specific setting.

table[table omitted — 1,958 chars of source]

Beyond Stratified RCTs

While strata indicators necessarily saturate the model in stratified RCTs, saturated specifications may be infeasible in other applications of instrumental variables methods, especially when there are multiple covariates or some covariates are continuously distributed. If saturation is infeasible, so is the specification in equations (ref)--(ref). The linear IV regression in equations (ref)--(ref) remains feasible, but a weighted average representation of $\beta_{\mathrm{IV}}$, analogous to equation (ref), additionally requires the parametric restriction that the conditional mean of $Z_{i}$ given $X_{i}$ is linear in $X_{i}$. BBMT2025 refer to this assumption as “rich covariates” and prove that it is necessary for a causal interpretation of $\beta_{\mathrm{IV}}$.

As in the case of stratified RCTs, we recommend explicitly estimating $\beta_{\mathrm{LATE}}$ in every relevant non-RCT application, even if only as a robustness check alongside $\hat{\beta}_{\mathrm{IV}}$. The two parameters, $\beta_{\mathrm{IV}}$ and $\beta_{\mathrm{LATE}}$, may often be similar, but when they are different, $\beta_{\mathrm{LATE}}$ is arguably of much greater interest. As discussed above, the expression in (ref) implies that the LATE can be estimated as the ratio of estimates of the average treatment effect of the instrument on the outcome and treatment. In the absence of saturation, there are many distinct estimators of the average treatment effect with good properties, which implies that there are also many suitable estimators of the LATE\@.\footnote{Some of these estimators include those in Tan2006, Frolich2007, Uysal2011, Heiler2022, SSX2022, SUW_DR,SUW_kappa, and MSSU2026. Related machine learning approaches have been developed by BCFH2017, Chernozhukovetal2018, ST2022, SS2024, and others.}

One such estimator is based on inverse probability weighting (IPW), which requires first-step estimation of the conditional mean of $Z_{i}$ given $X_{i}$, often referred to as the instrument propensity score. This is the same conditional mean that is central to the “rich covariates” assumption. Here, however, we focus on estimating $\beta_{\mathrm{LATE}}$ and will likely use a logit or probit model, not the linear probability model implicit in $\beta_{\mathrm{IV}}$. In any case, the IPW estimator uses the first-step estimates of the instrument propensity score to reweight the units with $Z_{i}=1$ and $Z_{i}=0$ so that both subsamples become comparable in terms of their covariate values. This, in turn, enables straightforward estimation of the average treatment effects of the instrument on the outcome and treatment using simple differences in means of reweighted data. The ratio of these differences is the IPW estimator of the LATE, implied by equation (ref) and proposed by Frolich2007. SUW_kappa emphasize the importance of using normalized weights in this context, explore the connections to “kappa weighting” of Abadie2003, and develop the Stata package kappalate. The package allows both binary and non-binary (discrete or continuous) treatments; the instrument, however, must be binary.

To illustrate the possible differences between $\beta_{\mathrm{IV}}$ and $\beta_{\mathrm{LATE}}$ in a nonsaturated setting, we estimate both parameters using Abadie2003's (Abadie2003) sample from the 1991 Survey of Income and Program Participation (SIPP)\@. In this application, based on PVW1994, one of the causal effects of interest is that of participation in a 401(k) retirement plan on participation in an individual retirement account (IRA)\@. While 401(k) participation is likely endogenous, PVW1994 and Abadie2003 argue that 401(k) eligibility, which is determined by the employer, can be used as an instrument for 401(k) participation. Because 401(k) eligibility is not randomized or “as good as randomly assigned,” Abadie2003 controls for family income, age, age squared, marital status, and family size to estimate the causal effects of 401(k) participation using this instrument.

table[table omitted — 1,342 chars of source]

Table (ref) replicates the linear IV estimate of the effect of 401(k) on IRA participation, as reported in column (2) of Table 3 in Abadie2003. This estimate, statistically significant at the 5% level, suggests that participation in a 401(k) retirement plan increases the probability of IRA participation by about 2.74 percentage points. However, Table (ref) also reports two IPW estimates of the LATE\@. These estimates, based on logit and probit instrument propensity scores, are much smaller than the linear IV estimate, with $p$-values equal to 0.221 and 0.197, respectively. This calls into question the initial conclusion about the positive impact of 401(k) on IRA participation. If we ignore estimation uncertainty and the possibility of misspecification, the difference between $\hat{\beta}_{\mathrm{IV}}$ and $\hat{\beta}_{\mathrm{LATE}}$ must be driven by the different weights underlying the corresponding estimands. In our view, the possibility that such differences may materialize necessitates explicit estimation of $\beta_{\mathrm{LATE}}$ in relevant applied work---either as the main parameter of interest or at least as a robustness check.

Panel A of Appendix Table (ref) lists several R and Stata packages, including kappalate, which can be used to estimate the LATE when covariates matter.

Parametric Misspecification

In this section we discuss the possibility that estimates of $\beta_{\mathrm{IV}}$ and $\beta_{\mathrm{LATE}}$ may be biased due to parametric misspecification. If all the requisite covariates are observed and controlled for, the leading case of misspecification occurs when important interaction and higher-order terms are omitted. First, we discuss how this concern applies to estimation of both $\beta_{\mathrm{IV}}$ and $\beta_{\mathrm{LATE}}$, and how it is possible to test for parametric misspecification. Second, we recommend modern estimation approaches based on machine learning as a general solution to this problem.

Rich Covariates

A large literature on weighted average representations of $\beta_{\mathrm{IV}}$, $\beta_{\mathrm{AI}}$, and related estimands has used the assumption that the conditional mean of $Z_{i}$ given $X_{i}$ is linear in $X_{i}$. As mentioned in the previous section, a recent paper by BBMT2025 shows that this assumption, termed “rich covariates,” is not only sufficient but also necessary for the resulting estimand to represent a non-negatively weighted average of complier causal effects. If covariates are not rich, the estimand is not “weakly causal,” which means that when all conditional average treatment effects have the same sign, the estimand may assume the opposite sign. PS2025 show that this is equivalent to the absence of any subpopulation of compliers whose average treatment effect is recovered by the estimand of interest regardless of the pattern of treatment effect heterogeneity.

An immediate implication is that applied researchers targeting simple instrumental variables estimands should only use covariate specifications that are plausibly rich. To examine whether a given specification is rich, BBMT2025 recommend the regression specification error test (RESET) of Ramsey1969. In its canonical form, RESET, as applied to this problem, consists of estimating a regression of $Z_{i}$ on $X_{i}$ using OLS, obtaining fitted values, $\hat{Z}_{i}$, and, finally, estimating a regression of $Z_{i}$ on $X_{i}$ and several higher-order terms of $\hat{Z}_{i}$, such as $\hat{Z}_{i}^2$, $\hat{Z}_{i}^3$, and $\hat{Z}_{i}^4$ (again using OLS)\@. In this setup, RESET is simply the $F$ test of joint significance of the higher-order terms in the second-step regression, and should be interpreted as a test for neglected nonlinearity Wooldridge2010. If we reject the null hypothesis, we conclude that the original covariate specification is not rich.

It would be inaccurate to suggest that estimators of $\beta_{\mathrm{LATE}}$ cannot suffer from similar specification problems. For example, the IPW estimator discussed in the previous section relies on the assumption that the conditional mean of $Z_{i}$ given $X_{i}$ is correctly specified, however it is specified. This offers more flexibility than the result in BBMT2025, because we are no longer tied to the linear probability model and can instead use, say, the logit or probit model. On the other hand, the logit or probit specification will not be correctly specified if we fail to include relevant interaction and higher-order terms, as in the case of “rich covariates.” To test for parametric misspecification of the logit or probit model for the instrument propensity score, one may use the extension of RESET proposed by PW1996. Here, we need to augment the original specification with several higher-order terms of the fitted linear index and, again, test for their joint significance.

Table (ref) illustrates the problem of parametric misspecification. Panel A revisits the estimates of $\beta_{\mathrm{IV}}$ and $\beta_{\mathrm{LATE}}$ in Table (ref), and reports the RESET $p$-values associated with each. With the original covariate specification, we decidedly reject the “rich covariates” assumption but also the correct specification of the logit and probit models for the instrument propensity score. Can a more flexible specification salvage a causal interpretation of our estimates? Panel B of Table (ref) uses a quintic in family income, a quintic in age, an indicator for marital status, and indicators for each value of family size. With this set of covariates, RESET does not reject any of the null hypotheses of correct specification of the linear probability, logit, and probit models. $\hat{\beta}_{\mathrm{IV}}$ and $\hat{\beta}_{\mathrm{LATE}}$ are now very similar to each other, although $\hat{\beta}_{\mathrm{LATE}}$ has barely changed relative to Panel A while $\hat{\beta}_{\mathrm{IV}}$ has decreased by about 40%.

table[table omitted — 2,130 chars of source]

An interesting question is whether this empirical illustration correctly indicates that $\hat{\beta}_{\mathrm{IV}}$ may be less robust to parametric misspecification than estimators of $\beta_{\mathrm{LATE}}$. While we are not aware of any such results for generic IPW estimators, there are other parametric approaches to estimate $\beta_{\mathrm{LATE}}$ with improved robustness properties, such as the covariate balancing estimators of Heiler2022, SSX2022, and SUW_kappa, and the “doubly robust” estimators of Tan2006, Uysal2011, SUW_DR, and MSSU2026. Several of these estimators are implemented in the Stata package drlate SUW_DR. We recommend these estimators for wider use, together with their machine learning counterparts discussed below.

Double Machine Learning

The previous subsection explained that a “weakly causal” interpretation of $\beta_{\mathrm{IV}}$ and $\beta_{\mathrm{LATE}}$ may fail when the conditional mean of the instrument given covariates is misspecified, perhaps due to omitted interaction or higher-order terms. One practical response is to manually enrich the covariate specification and then apply RESET\@. In many empirical applications, however, the number of potential interactions and nonlinear transformations is large, making manual specification and specification testing cumbersome.

A complementary approach is to use double/debiased machine learning (DML) to flexibly approximate the relevant conditional means. DML methods for estimating $\beta_{\mathrm{LATE}}$ have been developed by BCFH2017, Chernozhukovetal2018, ST2022, and SS2024, whereas Chernozhukovetal2018 also applied the DML methodology to estimating $\beta_{\mathrm{IV}}$\@. An application of double/debiased machine learning in an IV context proceeds as follows. The researcher supplies a rich dictionary of covariates to a machine learning algorithm (e.g., lasso, ridge, or regression tree). This method then approximates the relevant nonlinear relationships---the conditional means of potential outcomes and potential treatments, as well as that of the instrument---in a data-driven way, with tuning parameters typically chosen by cross-validation. Finally, DML combines the resulting flexible predictions with orthogonal scores and cross-fitting to decrease the underlying regularization bias. In this sense, DML can be viewed as an automated approach to constructing “rich” covariate specifications, reducing the need for ad hoc choices of interaction and higher-order terms.

Relative to manually specified parametric models, restrictions of DML are less explicit and arise implicitly from the choice of the machine learning algorithm. For example, lasso assumes that the target functions can be well approximated by a sparse subset of the dictionary, while tree-based methods rely on recursive partitioning and local averaging. Thus, even in the DML framework, the validity of the researcher's conclusions continues to hinge on the chosen approximation. Accordingly, it is still recommended to compare multiple machine learning algorithms to assess robustness and to reduce reliance on the approximation properties, and thus the implicit parametric restrictions, of any single method.

Table (ref) illustrates these considerations by reporting DML estimates of $\beta_{\mathrm{IV}}$ and $\beta_{\mathrm{LATE}}$ in the 401(k) application considered above. In the case of both parameters, we estimate the relevant conditional means using six different machine learning algorithms: lasso, elastic net, ridge, random forest, regression tree, and XGBoost. The resulting estimates are generally smaller than the initial linear IV estimate in Table (ref) but larger than the corresponding LATE estimates. At the same time, the DML estimates remain somewhat sensitive to the choice of the machine learning algorithm. Thus, while DML relaxes explicit functional form assumptions, it is still advisable to verify robustness across alternative approximation methods (as we do in Table (ref)). We recommend researchers routinely report estimates using at least two or three different algorithms to assess the sensitivity of their conclusions.

table[table omitted — 2,113 chars of source]

Panel B of Appendix Table (ref) lists the leading R and Stata packages that can be used to implement DML methods.

Assumption Violations

In this section we advise practitioners to systematically assess whether the assumptions underlying the LATE framework may be violated. We begin by reviewing several testable implications of these assumptions, and explain how the resulting statistical tests can be implemented in practice. We argue that using such tests can give more credence to IV applications. We also discuss an intuitive estimation approach that is more robust to violations of monotonicity, which is a controversial assumption in many applications. The validity of this assumption may also be rejected by one of the statistical tests that we review, in which case such alternative estimation approaches are particularly useful.

Testing Instrument Validity

As first observed by BP1997, the standard LATE assumptions---specifically, independence of the instrument, exclusion restriction, and monotonicity---have implications that can be used to test instrument validity. If the LATE assumptions hold, then it must be the case that

equation[equation omitted — 171 chars of source]

and

equation[equation omitted — 171 chars of source]

for any subset $A$ of outcome values (e.g., any interval or range).\footnote{Formally, the inequalities in (ref) and (ref) must hold for any Borel set $A$\@.} Because the inequalities in (ref) and (ref) are defined in terms of moments of observed variables, they can be examined with empirical data. Kitagawa2015 derived a formal statistical test based on these inequalities, and proved that they are sharp (i.e., cannot be improved upon) as well as necessary but not sufficient for instrument validity. In other words, we may be able to reject invalid instruments, but it is fundamentally impossible to confirm instrument validity, even with infinite data. In a subsequent contribution, HM2015 developed a sharp test of weaker assumptions that still suffice to identify the LATE\@. MW2017 reformulated the inequalities in (ref) and (ref) as two conditional moment inequalities, with $Y_i$ entering as a conditioning variable, which facilitates implementation. Sun2023 and KR2024 extended this framework to the case of multivalued treatments.

A careful reader might ask: Why is it the case that the inequalities in (ref) and (ref) must hold whenever an instrument is valid? One answer to this question follows from IR1997, who proved that the difference in (ref) identifies the joint distribution of the treated outcome and complier status, $\ensuremath{\mathbb{P}}( Y_i(1) \in A, D_i(1) > D_i(0) )$, while the difference in (ref) identifies $\ensuremath{\mathbb{P}}( Y_i(0) \in A, D_i(1) > D_i(0) )$, the joint distribution of the untreated outcome and being a complier. Because these probabilities, like any other probabilities, must be nonnegative, the differences in (ref) and (ref) must be nonnegative, too.

Useful intuition for these testable implications is also provided by KR2024. Note that the difference in (ref) is simply the coefficient on $Z_i$ in the population regression of the compound outcome $1[Y_i \in A, D_i = 1]$ on $Z_i$. If the instrument $Z_i$ is valid and there are no defiers, any difference can only be driven by always-takers, never-takers, and compliers. In the case of always-takers and never-takers, $Z_i$ has no effect on $D_i$. If so, then the exclusion restriction implies that it also has no effect on $Y_i$, ensuring that the effect of $Z_i$ on the compound outcome is zero for both groups. It follows that the entire difference in (ref) is driven by compliers. However, in the case of compliers, $Z_i = 0$ implies that $D_i = 0$, which means that $\ensuremath{\mathbb{P}}( Y_i \in A, D_i = 1 \mid Z_i = 1 ) - \ensuremath{\mathbb{P}}( Y_i \in A, D_i = 1 \mid Z_i = 0 ) = \ensuremath{\mathbb{P}}( Y_i \in A, D_i = 1 \mid Z_i = 1 )$. This probability, too, cannot be negative, which is exactly what the inequality in (ref) requires. A similar intuition justifies the inequality in (ref).

Statistical tests based on the BP1997 conditions may seem impractical in some applications, especially when the instrument is only valid conditional on covariates.\footnote{As an alternative, MS2020 and CK2023 show how to incorporate covariates to test the validity of the identifying assumptions in the MTE framework of CHV2011.} In such cases, the covariates enter the inequalities in (ref) and (ref) as conditioning variables, which increases the computational burden. To facilitate implementation, FGK2022 proposed testing the BP1997 conditions using machine learning. Specifically, their procedure uses causal forests to estimate the conditional average treatment effects of $Z_i$ on the compound outcomes $1[Y_i \in A, D_i = 1]$ and $-1[Y_i \in A, D_i = 0]$, and subsequently selects groups with probable assumption violations using regression trees. With sample splitting, group selection in one subsample still enables “honest” testing in another.

One limitation of the procedure in FGK2022 is that it only considers the BP1997 conditions, which require discretizing the outcome variable to construct the compound outcomes. This would not be necessary, for example, when testing the MW2017 conditions. It also seems reasonable to use a similar machine-learning-based approach to construct a formal test of the implication of independence and monotonicity that the conditional first stage, $\ensuremath{\mathbb{E}}( D_i \mid Z_i = 1, X_i ) - \ensuremath{\mathbb{E}}( D_i \mid Z_i = 0, X_i )$, must be nonnegative at all covariate values.\footnote{This implication is sometimes tested by practitioners, especially in applications of the “judge leniency” design, but the choice of groups to examine is typically ad hoc. See, for example, MMS2013, DGY2018, and AKMS2019. An appropriate procedure based on machine learning can systematically search for groups where violations of monotonicity are most likely to be found.} These testable implications of the LATE assumptions, as well as other features, are implemented in the R package montest Andresen2026.

table[table omitted — 1,889 chars of source]

Table (ref) reports the resulting $p$-values for three IV applications: the Oregon Health Insurance Experiment (OHIE), as in Tables (ref) and (ref), the 401(k) application of Abadie2003, as in Tables (ref)--(ref), and the study of causal effects of pretrial detention on case outcomes in Stevenson2018. This last application uses data on more than 300,000 Philadelphia arrests between 2006 and 2013, and instruments for pretrial detention with indicators for randomly assigned judges. Following the replication of Stevenson2018 in Sloczynski2025, we use incarceration length as the outcome variable as well as a saturated covariate specification with indicators for each combination of offense type, race and gender of the defendant, and three time periods considered by Stevenson2018. We also focus on a binary instrument that indicates whether a given case was assigned to the most lenient judge or not.

The $p$-values in Table (ref) clearly indicate that there is fundamentally no evidence against instrument validity in the first two applications, regardless of whether we consider the BP1997 (“BP”) conditions, the MW2017 (“MW”) conditions, or the nonnegativity of the conditional first stage (“FS”)\@. In fact, this last condition is trivially satisfied in the 401(k) application of Abadie2003, which is characterized by one-sided noncompliance. At the same time, we clearly reject the null hypothesis of instrument validity using the data from Stevenson2018. The conclusion is essentially the same whether we consider the testable implications of independence, exclusion, and monotonicity (BP and MW) or the testable implications of independence and monotonicity (FS)\@. Assuming that independence is satisfied, which is plausible, this implies that either monotonicity is violated but exclusion is satisfied, or both monotonicity and exclusion are violated. Our discussion in the next subsection will implicitly assume the former possibility.

Panel C of Appendix Table (ref) lists several R packages, including montest, which can be used to implement various tests of the LATE assumptions.

Robustness to Monotonicity Violations

If the monotonicity assumption is rejected or otherwise implausible, as in Stevenson2018, estimation of $\beta_{\text{LATE}}$ will be difficult, and the previously discussed estimators of $\beta_{\text{LATE}}$ and $\beta_{\text{IV}}$ will not consistently estimate any “weakly causal” target parameter. At the same time, appropriate estimators of $\beta_{\text{AI}}$ will not be subject to such concerns, at least when a weaker assumption, termed “weak monotonicity” in Sloczynski2025, is assumed to hold. Weak monotonicity, unlike strong monotonicity (i.e., the usual assumption), allows both compliers and defiers to exist, but they must not coexist at any value of covariates; that is, weak monotonicity requires that there be no defiers at some covariate values and no compliers elsewhere. This assumption will be plausible in some applications, but not in others. In Stevenson2018, it seems reasonable that only a small number of observed characteristics of cases and defendants would determine the relative leniency of any particular judge.

To see why standard estimators of $\beta_{\text{LATE}}$ and $\beta_{\text{IV}}$ will not work under weak monotonicity, note that the term $\pi_j$ in the weighted average representations in (ref)--(ref) only retains its interpretation as the conditional proportion of compliers under strong monotonicity. In the absence of any assumptions, it is simply equal to $\ensuremath{\mathbb{E}}( D_i \mid Z_i = 1, X_i ) - \ensuremath{\mathbb{E}}( D_i \mid Z_i = 0, X_i )$, that is, the conditional first stage. (Under strong monotonicity, the conditional first stage identifies the conditional proportion of compliers.) However, under weak monotonicity, the conditional first stage will be positive at some covariate values and negative at others, which directly translates to the incidence of “negative weights” in (ref) and (ref). On the other hand, $\pi_j$ enters the expression in (ref) as a quadratic rather than linearly, which ensures that the resulting weights are positive even if some conditional first stages are not. Sloczynski2025 argues that this makes focusing on $\beta_{\text{AI}}$ a useful compromise under weak monotonicity.

However, focusing on $\beta_{\text{AI}}$ is not without drawbacks. Because, by construction, $\beta_{\text{AI}}$ uses multiple instruments, its estimation will be subject to many-instrument bias, which results from overfitting the first stage. Recent surveys of the literature on many instruments include Anatolyev2019 and MS2024. The standard response to this problem is to estimate the specification in equations (ref)--(ref) with an alternative estimator, such as the fixed effect jackknife IV (FEJIV) estimator of CSW2023 or the unbiased jackknife IV estimator (UJIVE) of Kolesar2013 and GPHK2025. Both of these estimators, unlike 2SLS, are consistent under the asymptotic sequence that allows the number of instruments and the number of covariates to increase in proportion with the sample size. In addition, MS2022 propose a pretest for weak identification that can be used to determine whether multiple instruments are jointly strong enough for consistent estimation. While the pretest is not formally designed for settings with many covariates, the simulation results in Sloczynski2025 suggest that it effectively distinguishes between settings in which FEJIV and UJIVE perform well and those in which all estimators of the specification in equations (ref)--(ref) perform poorly.

Table (ref) replicates several of the estimates of causal effects of pretrial detention on incarceration length in Table 8 in Sloczynski2025. The specification is the same as in Table (ref) above. The linear IV estimate suggests that pretrial detention leads to a large, statistically significant increase in incarceration length of about 666 days. We know, however, that the underlying estimand is not “weakly causal” because of violations of (strong) monotonicity. Indeed, the 2SLS estimate of $\beta_{\text{AI}}$, which eliminates the negative weights, is 80% smaller than the linear IV estimate; it also remains statistically significant. While this estimate might be more believable, it suffers from many-instrument bias, unlike the associated FEJIV and UJIVE estimates. These, however, only suggest a small increase in incarceration length of less than 2 months, and are not statistically different from zero. This conclusion contrasts sharply with the initial IV and 2SLS estimates.\footnote{To be clear, Stevenson2018 does not use either of the problematic estimation approaches as her method of choice. Instead, she uses a nonsaturated specification with a large number of interacted instruments, and estimates the model using the jackknife IV estimator (JIVE) of AIK1999. Unlike FEJIV and UJIVE, this estimator is not formally appropriate for settings with many covariates.} It is also supported by the fact that MS2022's pretest rejects weak identification, as reported by Sloczynski2025, which indicates that consistent estimation is possible.

table[table omitted — 1,318 chars of source]

Panel D of Appendix Table (ref) lists several MATLAB, R, and Stata packages that can be used to implement FEJIV, UJIVE, and other relevant estimators, as well as the pretest in MS2022.

Conclusion

The local average treatment effect framework of AI1995 has fundamentally reshaped how applied microeconomists think about instrumental variables. It is now common to acknowledge that IV models only identify average treatment effects for compliers, examine the characteristics of that subpopulation, and develop careful arguments in favor of instrument validity. In this paper, we have explored three areas in which we believe empirical practice could further benefit from closer engagement with the recent theoretical literature. As we outline above, we recommend that practitioners explicitly target the “true” LATE rather than the usual IV estimand, consider flexible functional forms to avoid parametric misspecification, and use formal tools in response to possible assumption violations, including statistical tests of instrument validity and specifications with many interacted instruments.

\singlespacing

\setlength\bibsep{0pt}

\onehalfspacing

\setcounter{table}{0}

table[table omitted — 2,428 chars of source]