EconBase
← Back to paper

Bridging Methodologies: Angrist and Imbens' Contributions to Causal Identification

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

178,324 characters · 22 sections · 186 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Bridging Methodologies: Angrist and Imbens' Contributions to Causal Identification

modification\begin{abstract} In the 1990s, Joshua Angrist and Guido Imbens studied the causal interpretation of Instrumental Variable estimates (a widespread methodology in economics) through the lens of potential outcomes (a classical framework to formalize causality in statistics). Bridging a gap between those two strands of literature, they stress the importance of treatment effect heterogeneity and show that, under defendable assumptions in various applications, this method recovers an average causal effect for a specific subpopulation of individuals whose treatment is affected by the instrument. They were awarded the Nobel Prize primarily for this Local Average Treatment Effect (LATE). The first part of this article presents that methodological contribution in-depth: the origination in earlier applied articles, the different identification results and extensions, and related debates on the relevance of LATEs for public policy decisions. The second part reviews the main contributions of the authors beyond the LATE. J. Angrist has pursued the search for informative and varied empirical research designs in several fields, particularly in education. G. Imbens has complemented the toolbox for treatment effect estimation in many ways, notably through propensity score reweighting, matching, and, more recently, adapting machine learning procedures. \end{abstract}
commentRESUME (abstract en francais) Dans les années 1990, Joshua Angrist et Guido Imbens se demandèrent comment interpréter causalement les estimations obtenues au moyen de variables instrumentales (une méthode courante en économie) en s’appuyant sur la notion de variables potentielles (un cadre classique pour formaliser les relations causales en statistique). Ils comblèrent un fossé entre ces deux disciplines en mettant en évidence l’importance de considérer l’hétérogénéité des effets d’un traitement et en montrant que, sous des hypothèses raisonnables dans de nombreuses situations pratiques, cette méthode permet d’estimer un effet causal moyen sur une sous-population spécifique d’individus, ceux dont le traitement est affecté par l’instrument. Ils reçurent le prix Nobel d’économie essentiellement pour cette notion de « local average treatment effect » (LATE). La première partie de cet article présente en détail cet apport méthodologique : ses racines visibles dans des articles appliqués antérieurs, les différents résultats d’identification et leurs extensions, ainsi que les débats portant sur l’intérêt du LATE pour éclairer des décisions de politique publique. La seconde examine les principales contributions de ces deux auteurs en plus du LATE. J. Angrist a poursuivi ses travaux empiriques dans plusieurs champs, en particulier celui de l’éducation, toujours avec une attention singulière accordée à la stratégie d’identification en recherchant et en utilisant des expériences naturelles informatives et variées. G. Imbens a continué à enrichir la boite à outils permettant d’estimer les effets causaux d’un traitement ou d’une politique publique avec de nombreuses avancées méthodologiques, notamment le matching sur le score de propension et, plus récemment, l’adaptation des techniques d’apprentissage statistique (« machine learning ») aux problématiques économétriques.
smallKeywords: instrumental variables (IV); Neyman-Rubin causal model; local average treatment effect (LATE); natural experiments; returns to schooling; US educational system; propensity score matching; causal machine learning.
commenten français : variables instrumentales (IV) ; modèle causal de Neyman-Rubin ; effet local moyen du traitement (LATE) ; expériences naturelles ; rendements de l’éducation ; système éducatif états-unien ; appariement sur le score de propension ; apprentissage statistique causal. titre en français : A la croisée de A la frontière de l’économétrie / économie et de la statistique économie plus qu'économétrie pour l'aspect interprétation, attention au sens économique A la croisée de l'économie et de la statistique : les contributions d’Angrist et d’Imbens à l’identification causale Entre économie et statistique : les contributions d’Angrist et d’Imbens à l’identification causale Unifier des méthodologies : les contributions d’Angrist et d’Imbens à l’identification causale Réunir économie et statistique : Rapprocher économie et statistique : Faire dialoguer économistes et statisticiens : les contributions d’Angrist et d’Imbens à l’identification causale

Introduction

In the mid-1990s, Professors Joshua Angrist and Guido Imbens contributed to a series of foundational papers imbens_angrist_1994, angrist_imbens_1995, angrist_imbens_rubin_1996 which brought together two powerful yet seemingly distant concepts, namely Neyman-Rubin's potential outcome framework rubin1974estimating, rubin1990 and instrumental variables wright_1928, Reiersl1945ConfluenceAB. At the time, the former was mainly popular in medical studies with random allocation of treatment to patients. In that field, the idea that each patient is faced with two potential health statuses, one with treatment and one without, was already widely accepted. Thanks to a random treatment allocation, comparisons across treatment groups allow for uncovering the average change in health status due to treatment, also termed the average causal effect of the treatment. This approach does not require imposing strong {a priori} assumptions on the nature of the link between treatment and health status, which is appealing.

Unlike medical researchers, economists were (and still are) mostly faced with observational data, namely data collected on agents who make choices in uncontrolled environments. In this context, it is the rule rather than the exception to witness agents who make interdependent and sometimes simultaneous choices. Researchers try to identify the impact of one choice on another in a ceteris paribus way, that is, holding the rest of the environment fixed. Unfortunately, co-occurring choices are often based on factors unobserved by the researcher, which makes it hard to disentangle causes and consequences leading to the well-known endogeneity issue (see Wooldridge2010 for an introduction). As an example, let us consider the question: “What is the impact of education years on potential wage?” Intrinsic motivation, which is hardly captured in collected data, can affect agents' decisions both regarding education and labor market outcomes. Without more information, it is thus impossible to interpret correlations between observed educational choices and wages as evidence of the ceteris paribus influence of education on wages. A long-established solution is to resort to instrumental variables (IV) plus the assumption that choices affect one another linearly, leading to the linear IV model. In the education-wage example, IVs are, loosely speaking, variables that capture variations in education unrelated to motivation\textcolor{couleurModification}{; for instance, the distance between home and the nearest college has been used in the literature card1993using.} Such variables enable researchers to identify the effect of education on wages holding motivation fixed.

The popularity of the linear IV model among applied economists largely comes from its ease of implementation and the simplicity of the model connecting the endogenous variable (the number of years of education in the previous example) and the outcome (wages). Having a simple-to-use estimator is an unquestionable advantage from a practical perspective. The simplicity of the underlying model has a major downside, however. As an illustration, a linear model of wages in terms of education implies that the ceteris paribus effect of education on wages is assumed to be the same for every individual in the population. In other words, basic linear models impose individuals be perfectly homogeneous in their response to a given (economic) stimulus. This is the starting point of J. Angrist and G. Imbens's common research agenda of the 1990s. Roughly speaking, these authors ask: “Assume one runs a linear IV regression while the true underlying model is not linear. Does the estimated coefficient on the endogenous variable still capture an economically meaningful quantity?” Addressing this ambitious question is not straightforward and requires, in particular, defining what is meant by economically meaningful. To do so, the authors propose to extend the canonical Neyman-Rubin potential outcome framework to make it compatible with the non-experimental designs in which IV models are typically applied. The economic quantities they are primarily interested in are average treatment effects, as in the classical Neyman-Rubin setup. Their idea is fairly simple and turns out to be extremely powerful in analyzing the properties of linear IV estimators in populations of heterogeneous agents. As we will see in Section (ref), the authors give a partially positive answer to the question presented above: linear IV can recover average treatment effects on some specific subgroups of individuals provided instruments are exogenous, relevant, and restrictions are placed on the direction in which instruments affect the endogenous variables; the so-called monotonicity assumption. On the other hand, treatment effects can remain fully heterogeneous at the individual level. Since only average effects for certain subpopulations can be identified, the authors called those local average treatment effects (LATE), a term now ubiquitous in economics.

Several aspects of J. Angrist and G. Imbens's message about their LATE-type identification results are often overlooked. Even though they claim that LATEs are, in some sense, the best one can hope to identify using IV methods when treatment effects' heterogeneity is not restricted, they unambiguously urge empiricists to reflect upon the relevance of the instruments used in any particular application. They emphasize the need for researchers to first focus on obtaining data with clean-enough sources of exogenous variation in the instruments, and they give detailed accounts of situations where the monotonicity of the instrument may or may not be satisfied. They also acknowledge that LATE may or may not be informative about average treatment effects in the entire population. To get round this intrinsic limitation, they encourage applied researchers to replicate IV estimations in varied scenarios to gain insight into the potential to extrapolate from LATE to more general average treatment effects. In J. Angrist and G. Imbens's eyes, the LATE paradigm only makes sense when it is combined with cautious empirical practices and rich enough sources of variation in the data at hand imbens2010, angrist_pischke_2010.

As made explicit by the Nobel committee, Joshua Angrist and Guido Imbens have been awarded the prize primarily for their breakthrough findings around the LATE KVA_Nobel_2021. However, their contributions extend way beyond this series of papers. J. Angrist is a recognized labor and education economist. He is renowned in particular for his work on the link between family composition and labor market outcomes angrist_evans_1998, the labor market prospects of war veterans angrist_1990, the impact of class size on educational achievement angrist_lavy_1999, or the effect of additional education years on earnings angrist_krueger_1991. G. Imbens is primarily a theoretical econometrician. He has worked in many different fields, including calibration methods inspired by survey sampling imbens1992, efficient estimation in models identified through moment conditions imbens_johnson_spady_1998, and causal inference at large. His contributions to the latter literature are numerous. He has worked on many methodological issues related to the estimation of average treatment effects, including propensity score reweighting hirano_et_al_2003, matching abadie_imbens_2006, regression discontinuity imbens_lemieux_2008, and synthetic control arkhangelsky_et_al_2021. He has also worked on the identification of policy-relevant parameters in instrumental variable models imbens_newey_2009 and on quantifying treatment effects' heterogeneity athey_imbens_2015. The academic influence of these two authors is also to be seen in the two reference econometrics textbooks angrist_pischke_2008, imbens_rubin_2015 they have co-authored with Jörn-Steffen Pischke and Donald Rubin, respectively.

The rest of this article is organized as follows. Section (ref) is devoted to a detailed presentation of the “LATE revolution” that J. Angrist and G. Imbens initiated in the 1990s. It aims at giving a comprehensive account of the rich set of results put forward by the authors in the series of articles that gave birth to the LATE paradigm. The theory framed indeed extends way beyond the canonical LATE theorem with one binary treatment and one binary instrument traditionally presented in econometrics courses. In Section (ref), we discuss the numerous other fields to which these authors have contributed throughout their careers. \textcolor{couleurModification}{Section (ref) concludes highlighting the similarity between the questions (and answers) brought by J. Angrist and G. Imbens on linear IV models and current issues about the interpretation of the parameters identified in two-way fixed effects models.}

A LAsTing rEvolution

This section focuses on the LATE methodological contribution, by which we refer to the framework, assumptions, and identification results introduced by J. Angrist and G. Imbens along with D. Rubin, mainly in the three following articles (the “LATE trilogy”):

itemize[noitemsep,nolistsep] • “Identification and Estimation of Local Average Treatment Effects”, Econometrica, imbens_angrist_1994 (henceforth IA94), • “Two-Stage Least Squares Estimation of Average Causal Effects in Models With Variable Treatment Intensity”, Journal of the American Statistical Association, angrist_imbens_1995 (henceforth AI95), • “Identification of Causal Effects Using Instrumental Variables”, Journal of the American Statistical Association, angrist_imbens_rubin_1996 (henceforth AIR96),

and direct extensions.

The fundamental ideas of the LATE contribution are now standard in many econometrics courses and textbooks, and various references are available. In the context of this review, a good reference, for instance, is the article of the Royal Swedish Academy of Sciences on the scientific background of the Nobel 2021 Prize KVA_Nobel_2021. For completeness, Appendix (ref) recalls basic notions of the Neyman-Rubin potential outcomes framework and their relations to Ordinary Least Square (OLS) and Instrumental Variables (IV) or Two-Stage Least Squares (TSLS) in econometrics. Then, it presents the textbook LATE theorem, that is, the identification of the average effect on compliers in the setting of a binary treatment, a binary instrument, and no covariates. That result may be familiar to many readers. Hence our choice to report this more standard material as annexes and devote this section to (i) locate the different settings and results of the three seminal articles (IA94, AI95, AIR96), which extend the textbook LATE theorem (Section (ref)); (ii) investigate the origination of these ideas that can be found in earlier applied articles (Section (ref)); (iii) discuss some specific points more in-depth, notably those related to controversies about the interest of the LATE and the opposition between “structural” and “causal” approaches (Section (ref)). Before doing so, we briefly recall the setting, main assumptions, and insights of the LATE contribution. Appendix (ref) presents the framework and notations with more details for curious readers or less familiar with this formalization.

\paragraph{Setting and notation}

We are interested in the causal effect of a treatment \(D\) on an outcome variable \(Y\). Following Neyman-Rubin's model, we introduce the following random variables.

itemize[noitemsep,nolistsep] • \(Z\) is the instrument; it is a scalar real random variable that will be binary, namely \(\textrm{Support}(Z) = \{0,1\}\), or discrete, multi-valued with finite support; in this case, the support of \(Z\) is \(\{z_0, z_1, \ldots, z_K\}\), also encoded as \(\{0, 1, \ldots, K\}\). • For any \(z \in \textrm{Support}(Z)\), \(D(z)\) denotes the potential treatment variable: what would be the treatment status of the unit if its instrument took value/modality \(z\). • \(D := D(Z)\) is the observed (as opposed to potential) treatment variable; like \(D(z)\), it is a scalar real random variable that will be binary (\(\textrm{Support}(D) = \{0,1\}\)), or discrete with a quantitative ordered meaning and taking values in a finite set \(\{0, 1, \ldots, J\}\), or a continuous random variable with \(\textrm{Support}(D) = [0, +\infty)\). • For any \(z \in \textrm{Support}(Z)\) and \(d \in \textrm{Support}(D)\), \(Y(z,d)\) is the potential outcome variable: what would be the outcome of the unit had it got \(Z = z\) for the instrument and \(D = d\) for the treatment; under the exclusion restriction (see Assumption (ref) below), it does not depend on the instrument so that \(Y(z,d)\) does not depend on \(z\) and is simply denoted \(Y(d)\). • \(Y := Y(Z, D(Z))\) or, under the exclusion restriction, \(Y := Y(D)\), is the observed outcome variable; like \(Y(d,z)\), \(Y\) is a real random variable that can be binary (qualitative outcome with two modalities) or continuous (quantitative outcome). • If present, \(X\) is a vector of observed covariates: each covariate only takes a finite number of distinct values; that is, we restrict to discrete covariates.

We additionally define individual causal effects as differences in potential variables. For instance, under the exclusion restriction and with a binary treatment, \(\Delta := Y(1) - Y(0)\) is the individual causal/treatment effect (used as synonyms) of \(D\) on \(Y\). \(\Delta\) is a real random variable and, a priori, with an unrestricted distribution (heterogeneous effects as opposed to homogeneous causal effects that assume \(\Delta = \delta_0 \in \mathbb{R}\) almost surely).

We observe an independent and identically distributed (i.i.d) sample \((Y_i, D_i, Z_i, X_i)_{i=1, \ldots, n}\) drawn from the population of interest. All the previous variables are defined at the level of an individual/agent/(statistical) unit indexed by \(i \in \{1, \ldots, n\}\). Thanks to i.i.d-ness and to lighten notation, we often omit the index \(i\): a random variable not indexed by \(i\) denotes a generic instance with the same distribution. For instance, \((Y, D, Z, X)\) has the same joint distribution, denoted \(\mathrm{P}_{(Y, D, Z, X)}\), as any observation \((Y_i, D_i, Z_i, X_i)\) for \(i \in \{1, \ldots, n\}\). Potential variables are also defined at the individual level. For instance, \(D(1)_i\) is unit \(i\)'s treatment had it received \(Z_i = 1\) for the instrument; \(D(1)\) is the same for an arbitrary individual of the population of interest. Likewise, \(\Delta_i\) is the individual causal effect for unit \(i\).

Miscellaneous notations. For any random variables \(A\) and \(B\), \(\mathrm{P}_{A}\) denotes the distribution of \(A\), and \(\mathrm{P}_{A \,|\, B}\) the conditional distribution of \(A\) knowing \(B\); \(\textrm{Support}(A)\) denotes the support of \(\mathrm{P}_{A}\). \(\perp\mkern-9.5mu\perp\) denotes independence between random variables; \(A \perp\mkern-9.5mu\perp B\) means that \(\mathrm{P}_{A \,|\, B} = \mathrm{P}_{A}\). \(\mathbb{P}\), \(\operatorname{\mathbb{E}}\), etc. denote the probability, expectation, etc. operator. The symbol \(:=\), as opposed to \(=\), denotes equality by definition (the left-hand term is defined by the right-hand term; or the reverse with \(=:\)). \(f\) or \(g\) denote a function (not necessarily explicitly defined) used locally for reasoning.

\paragraph{Main assumptions}

The following assumptions will prove central. We refer to them as the “LATE assumptions”.

itemize[noitemsep,nolistsep] • Exclusion restriction (the instrument affects the outcome only through the treatment): \begin{equation*} \tag{E} \forall z, z' \in Support(Z), \, \forall d \in Support(D), \, Y(z,d) = Y(z',d). \end{equation*} • Independence (of the instrument; it can be considered “as-if” randomized): \begin{equation*} \tag{I} Z \perp\mkern-9.5mu\perp \big\{ (D(z))_{z \in Support(Z)}, (Y(z, d))_{z \in Support(Z), d \in \textrm{Support}(D)} \big\}. \end{equation*} • Relevance (of the instrument; it has a causal effect on the treatment): \textcolor{couleurModification}{when \(Z\) is binary, it can be written as} \begin{equation*} \tag{\textnormal{R}} \operatorname{\mathbb{E}}[D(1) - D(0)] \neq 0. \end{equation*} \textcolor{couleurModification}{With more general instruments, this assumption can be formulated in several ways. We introduce such extensions on a case-by-case basis in the rest of the article.} • Monotonicity (of the effect of the instrument on the treatment): \begin{equation*} \tag{\textnormal{M}} \forall z, w \in \textrm{Support}(Z), \, \text{ either } \, D(z) \geq D(w) \text{ almost surely (a.s)} \, \text{ or } \, D(z) \leq D(w) \text{ a.s}, \end{equation*} \textcolor{couleurModification}{which, with a binary instrument (and a conventional choice for the ranking), amounts to} \begin{equation*} \tag{\textnormal{M'}} \textcolor{couleurModification}{D(1) \geq D(0) \text{ almost surely.}} \end{equation*}
modification\paragraph{Definition of relevant subpopulations} According to their potential treatment variables, individuals or units can be partitioned into several subpopulations which play a crucial role in the definition and interpretation of LATE parameters. It is useful to focus on the case of a binary instrument and binary treatment to describe these populations in an intuitive fashion. In this canonical setting, individuals or units can be partitioned into four types depending on their potential treatment variables (Table (ref)). \begin{table}[H] \caption{Partition of the population with a binary instrument and a binary treatment.} \begin{tabular}{l | c | c } & \({D(1)} = 0\) & \({D(1)} = 1\) \\ \hline \({D(0)} = 0\) & never-taker (NT) & complier (C) \\ \hline \({D(0)} = 1\) & defier (D) & always-taker (AT) \end{tabular} \end{table} Never- and always-takers are individuals whose treatment status is not affected by the instrument. To resume the education-wage example of the introduction, the treatment \(D\) could be defined as the indicator of going to college and the instrument \(Z\) as the indicator of living less than, say, 10 miles away to a college. An individual is an always-taker (respectively never-taker) if, whatever the distance between the nearest college and her home, she goes (does not go) to college. In contrast, the instrument does affect the treatment of compliers and defiers. For instance, in our example, a complier is defined by the fact that she would go to college had she lived close enough to a college but, all other things equal, would not go had she lived further. In the simple case of \(\textrm{Support}(Z) = \textrm{Support}(D) = \{0, 1\}\), (ref) is equivalent to the absence of defiers: if the instrument affects the treatment, it must do so in the same direction for each individual. In our example, living close to a college cannot induce anyone not going to college. \paragraph{Comments on the assumptions and insights} As just explained, for compliers, the instrument has an effect on the treatment (see the relevance assumption (ref) above). In addition, the second condition for a valid instrument relates to its exogeneity: Can it be considered as-if randomly assigned like in a controlled experiment (see independence assumption (ref) above)? Does it affect the outcome only through the treatment, without “direct” effect (see exclusion assumption (ref))? In a nutshell, the LATE methodological contribution presented at length below can be formulated as follows. Without restricting the heterogeneity of treatment effects, (i) the identification of a sensible causal parameter requires that the instrument (weakly) affects all individuals in the same direction (see the monotonicity assumption (ref) above), (ii) information learned on causal effects only concerns individuals whose treatment is affected by the instrument; in that sense, the average causal effect recovered is only local, it is the average on the subpopulation of compliers. That is why the complier subpopulation is central in the interpretation of LATE-type results.

Early LATEs?

The trilogy IA94, AI95, and AIR96 forms the core of the LATE methodological contribution pioneered by J. Angrist (J.A) and G. Imbens (G.I) along with D. Rubin. Interestingly, some of the key ingredients of the LATE contribution can be found in earlier papers. This subsection reviews a number of articles and ideas that were forerunners of the LATE revolution: origins of the potential outcomes paradigm and its early use in economics (Section (ref)), concerns about possibly heterogeneous causal effects (Section (ref)), and the attention devoted to the exclusion restriction (Section (ref)).

The (non-)use of potential outcomes in economics

\paragraph{The tradition of potential outcomes in statistics}

As outlined above, one of the main ways to present the LATE contribution is that it builds a bridge between statistics' potential outcome framework and instrumental variables used by economists in a setting of simultaneous (structural) equation models. The authors themselves claim that connection.\footnote{ In particular, it is noteworthy that they conclude their rejoinder to AIR96's comments (following the article, the issue presents comments by James Robins and Sander Greenland, James Heckman, Robert Moffitt, and Paul Rosenbaum) on this theme: “{The comments on our article cover a wide range of opinions, partly reflecting the gap between competing paradigms for evaluation research in statistics and econometrics. We believe that the gap between the two approaches can be narrowed. We hope that our article will make statisticians more appreciative of the insights offered by the IV framework invented by econometricians, while making economists more aware of the benefits of causal inference conducted in the potential outcomes framework developed by statisticians}” (Journal of the American Statistical Association June 1996, Vol. 91, No. 434, page 472). } In this respect, it is interesting to remark that the first paper of the trilogy was published in a leading economics journal; in contrast, the two others were published in a leading statistics journal.

Donald Rubin's work rubin1974estimating, rubin1978 is central in the development of the potential outcome framework since, although the notion can be found in the “potential yields” of neyman1923, D. Rubin generalizes it beyond randomized experiments to include observational studies that do not involve randomization. Causal inference in statistics (and related fields such as epidemiology) has been inherently connected to the notion of potential outcomes a la Neyman-Rubin at least since Rubin's articles were published. In contrast, potential outcomes were not very common in economics at the time of the redaction of the LATE trilogy, with a few notable exceptions maddala1983, bjorklund1987, heckman1990. We argue below this gap was meant to be filled due to the strong natural connections between potential outcomes and economic reasoning and the joint interest of G. Imbens in statistics and economics.

\paragraph{The connection with demand and supply curves}

The notion of potential outcomes naturally connects with economics. Perhaps the most straightforward objects to see this are demand and supply functions, which are ubiquitous in that field. For instance, a demand function outputs the quantity asked by customers given a prescribed vector of prices. An interest in the elasticity of demand with respect to price can be interpreted, in the framework of potential outcomes, as an interest in the causal effect of price on demand. A demand function, by the fact that it is a function, shares with potential outcomes the critical feature of being potential or counterfactual. It formalizes the concept of conceiving what would have been the quantity asked (the outcome variable) had the price (the treatment variable) been set to some specified value. This example shows that, although latent, the notion of potential outcomes was already present in economists' thinking from the start. J.A, G.I and D. Rubin quote tinbergen1930 about this example of supply and demand curves in their answer to comments to AIR96 (see Journal of the American Statistical Association, June 1996, Vol. 91, No. 434, page 469). They also quote haavelmo1944 for early inception of potential outcomes.

This example is not uniquely illustrative since the analysis of competitive markets and the estimation of demand (or supply) curves have been a case in point of the use of instrumental variables. J. Angrist, Kathryn Graddy, and G. Imbens notably extend the LATE formalism to this type of analysis with their article “The Interpretation of Instrumental Variables Estimators in Simultaneous Equations Models with an Application to the Demand for Fish” angrist_graddy_imbens_2000. This article is an important extension of the LATE trilogy, and we return to it in Section (ref).

\paragraph{An influence of G. Imbens?}

Second, it is interesting to notice that, in two influential applied articles,

itemize[noitemsep,nolistsep] • “Lifetime Earnings and the Vietnam Era Draft Lottery: Evidence from Social Security Administrative Records”, The American Economic Review, angrist_1990 (henceforth A90) • “Does Compulsory School Attendance Affect Schooling and Earnings?” The Quarterly Journal of Economics, angrist_krueger_1991 (henceforth AK91)

published just before the LATE trilogy, J. Angrist and his co-author A. Krueger do not resort to the potential outcome framework. Instead, causal parameters are introduced through simultaneous structural equation models. It may suggest that G. Imbens imported the notion of potential outcomes, even if his early works are not directly connected with causal inference issues but with another statistical theme, namely efficiency issues (see Section (ref)).

A concern about heterogeneity

A second central aspect of the LATE contribution relates to the attention devoted to heterogeneity: causal effects are a priori specific to each agent. Consequently, average causal effects over different subpopulations are expected to differ. It is interesting to compare the two above-mentioned applied articles in that dimension. In particular, although contemporary, A90 and AK91 substantially differ in their study of heterogeneity.

\paragraph{Heterogeneity in A90: already LATE}

A90 uses the Vietnam-era lottery draft as an instrumental variable for the endogenous veteran status and thus proposes TSLS estimates of the effect of military service on earnings. A90 appears remarkable as a precursor of the formalization of LATE regarding its attention to heterogeneity. The fifth section of this article, entitled “Caveats,” examines potential deviations from the main model. It encompasses treatment effect heterogeneity that “{merely result[s] in a reinterpretation of the estimates}” (A90, page 329), from an average treatment effect (ATE) to a LATE interpretation. It anticipates the distinction between what will soon be called always-takers (“true volunteers” in A90) and compliers. Indeed, J. Angrist writes “{Suppose that the impact of military service on the earnings of true volunteers differs from the impact on the earnings of draftees and men who enlisted because of the draft. Then using functions of the draft lottery as instruments will only identify the effect of military service for the latter group.}” (A90, pages 329-330, we italicize). Except for the potential outcome formalization, this is already exactly the LATE interpretation.

\paragraph{AK91: heterogeneity analysis based on observed covariates versus LATE-type heterogeneity}

AK91 uses the combination of quarters of birth and mandatory school attendance laws to study the causal effect of education (measured as the accumulated time spent in studies) on earnings. Quarters of birth - interacted with state fixed effects - are used as instruments of the number of years of education. In line with A90, AK91 is concerned with possible sources of treatment effect heterogeneity since they run the analysis both on the overall male population and on the sub-sample of black men (see section II.C of AK91). A major difference between the two articles is however that AK91 is concerned with heterogeneity stemming from an observed covariate, namely ethnicity, while A90 focuses on heterogeneity across unobserved populations, namely compliers and always-takers.

commentBoth types of heterogeneity can be relevant. Yet, it is interesting that AK91 does not consider the “modern” LATE-type heterogeneity. Indeed, one of the contributions underscored in AK91 is the comparison between OLS and TSLS estimates. “Using season of birth as an instrument for education in an earnings equation, we find a remarkable similarity between the OLS and the TSLS estimates of the monetary return to education. [\dots] This evidence casts doubt on the importance of omitted variables bias in OLS estimates of the return to education, at least for years of schooling around the compulsory schooling level.” (AK91, pages 1009-1010). To make sense, such a comparison requires that those two sets of estimators recover the same parameter.\footnote{ This remark provides the opportunity to mention another early article by J. Angrist, angrist1991. Among other results, this article shows that a TSLS over-identification test statistic is a test statistic for the equality of alternative Wald estimates of the same parameter. Wald or IV estimates rely on a single instrument, while TSLS estimates combine instruments (the end of Section (ref) in appendices discusses more that terminology). With the LATE result in mind and allowing for heterogeneous treatment effects, there is no reason for IV estimators based on two distinct instruments to estimate the same target parameter since the definition of the subpopulation of compliers is specific to the instrument. } Assuming a causal interpretation is feasible (no endogeneity issue for OLS, valid instrumental strategy for TSLS), this would be the case under the hypothesis of homogeneous causal effects. Without, OLS identify a weighted average over the entire population, whereas TSLS identify the famous LATE, an average causal effect over the subpopulation of compliers. Despite this stressed contribution, AK91 alludes to this issue with the precision: “{at least for years of schooling around the compulsory schooling level}”. Besides, the next sentence is: “Our results provide support for the view that students who are compelled to attend school longer by compulsory schooling laws earn higher wages as a result of their extra schooling.” (AK91, p1010). We italicize as this fragment closely connects to the characterization of compliers. Overall, AK91 appears mixed about the issue of treatment effects heterogeneity, notably compared to A90, which, except for the potential outcome formalization, is already remarkably close to the LATE.

Beware the exclusion restriction!

The exclusion restriction (ref) is one of the three requirements for a valid instrument. In this series of applied works, it is interesting to see that both A90 and AK91 are very careful about this condition, anticipating in that sense one of the assumptions of the LATE theorem and, especially, the attention devoted to this assumption compared to the simultaneous equation modeling (see Section (ref) for further details).

A90 devotes one subsection of the fifth “Caveats” section to it. AK91 deals with this issue in the third section of the paper. The following remark from AK91 is particularly explicit: “if season of birth influences earnings for reasons other than compulsory schooling, our approach is called into question” (AK91, p1007).

Although formalization still relies on traditional simultaneous equations instead of potential outcomes, the space and attention devoted to this issue suggest J. Angrist and A. Krueger were already well aware of the importance of this hypothesis.

The LATE trilogy

Section (ref) presents identification of the LATE in the simple setting of a binary instrument, a binary treatment, and no covariates. It corresponds to a standard econometrics textbook presentation, whose advantage is to focus on the novel (and now Nobel) ideas of the LATE. However, the three seminal papers, IA94, AI95, and AIR96, are far richer and address more general settings, multi-valued instruments, multi-valued treatments, presence of covariates.

This subsection presents and contrasts the more general results of these three papers:\footnote{ Contrary to Appendix (ref), for the sake of brevity, we do not enter into the complete formalization and proof of the results. We can but encourage interested readers to read the three original papers for more details. } imbens_angrist_1994 (Section (ref)), angrist_imbens_1995 (Section (ref)), \citet*{angrist_imbens_rubin_1996} (Section (ref)). We also evoke an essential, closely connected extension (to continuous treatment) of the LATE framework, namely \citet*{angrist_graddy_imbens_2000} (Section (ref)). For easier comparisons, Table (ref), displayed at the end of Section (ref) on page (ref), proposes a summary of the main results of IA94, AI95, AIR96, and extensions contrasting the different settings in terms of treatments and instruments.

Identification and Estimation of Local Average Treatment Effects (IA94)

This 9-page article is the canonical reference that introduces the monotonicity assumption and identification of the LATE, including the name itself.\footnote{ “We call this a local average treatment effect (LATE).” (IA94, p467). } It considers a binary treatment, \(\textrm{Support}(D) = \{0, 1\}\). In contrast, the instrument is more general: it is a discrete random variable (possibly more than binary), either scalar or multi-dimensional (although the two are equivalent in the sense that we can return to a scalar enumerating the different possible modalities of the vector).

The potential outcomes framework is used. The exclusion restriction (ref) is less formally explicit in so far as the only potential outcome variables introduced are \(Y(0)\) and \(Y(1)\), thus with only the treatment as an argument, not the instrument. That being said, a comment of their Condition 1 (“Existence of instruments”) refers to the exclusion restriction more directly, suggesting that a “random” determination of the instrument is not enough to guarantee what corresponds to hypotheses (ref) and (ref).\footnote{ “Note that random assignment of \(Z_i\) does not guarantee part (i) is satisfied because although random assignment implies that \(Z_i\) is independent of \(D_i(w)\), it does not imply that \(Z_i\) is independent of \(Y_i(0)\), \(Y_i(1)\).” (IA94, p468). Here \(w\) is any possible value taken by the instrument (as \(z \in \textrm{Support}(Z)\) above). }

\paragraph{The first LATE theorem}

Since the instrument is not restricted to be binary, Theorem 1 of IA94 is already more general than the textbook LATE theorem. Under the existence of valid instruments (their Condition 1 corresponding to (ref), (ref), and (ref)) and the monotonicity assumption (Condition 2 of IA94, which is exactly (ref)), for any pair \(z\) and \(w\) of distinct values of \(Z\) such that the instrument influences the treatment for these values (\(P(z) := \operatorname{\mathbb{E}}[D \,|\, Z = z] \neq \operatorname{\mathbb{E}}[D \,|\, Z = w]\)), the average causal parameter

align[align omitted — 155 chars of source]

is identified from the joint distribution \(\mathrm{P}_{(Y, D, Z)}\).

The causal parameter \(\delta^\textnormal{C}_{z,w}\) is defined as a function of the values \(z\) and \(w\) considered for the instrument. Of course, similar to the usual binary-instrument, binary-treatment LATE (denoted \(\delta^\textnormal{C}\) in Appendix (ref)), it is a local average treatment effect since the average is over the subpopulation of \textcolor{couleurModification}{compliers,} “those who can be induced to change participation status [the treatment] by a change in the instrument [from \(z\) to \(w\)]” (IA94, p470).

The textbook LATE theorem is a corollary of this result when the instrument is binary, \(\textrm{Support}(Z) = \{0, 1\}\). Indeed, the only two distinct values of the instrument are then \(z = 1\) and \(w = 0\) (or the reverse, this is symmetric) and, thanks to monotonicity (with the conventional order of (ref)), \(\{D(1) \neq D(0)\}\) is equivalent to \(\{D(1) > D(0)\}\), which defines compliers \textcolor{couleurModification}{(see Table (ref) in page (ref))}. Thus, \(\delta^\textnormal{C} := \operatorname{\mathbb{E}}[Y(1) - Y(0) \,|\, D(1) > D(0)] = \delta^\textnormal{C}_{1,0}\).

\paragraph{LATE interpretation of the IV estimand}

In a sense, Theorem 1 still focuses on a binary instrument since it only considers two distinct values of the instrument. IA94 then moves on to study the case of a more general instrument: “One way to exploit a multi-valued instrument is to estimate the ratio of the covariance of \(Y\) and some scalar function \(g(Z)\), and the covariance of \(D\) and \(g(Z)\). If \(Z\) is a scalar random variable, then the choice \(g(z) = z\) leads to the standard IV estimator. If \(Z\) is a vector, \(g(z)\) is often an estimate of \(P(z)\) [the probability of being treated conditional on \(Z = z\)]”. (IA94, p470).

To do so, IA94 requires Condition 3 (IA94, p470). In essence, it ensures a coherent choice of the function \(g(\cdot)\) in terms of ranking: for any \(z\), \(w \in \textrm{Support}(Z)\), \(P(z) \leq P(w) \implies g(z) \leq g(w)\), plus the relevance condition \textcolor{couleurModification}{\(\mathbb{C}\mathrm{ov}(D, g(Z)) \neq 0\)}. Importantly, in frequent situations (binary \(Z\) or \(g(\cdot) = P(\cdot)\)), the first part of this condition is automatically satisfied.

Under Conditions 1, 2 and 3 of IA94, Theorem 2 of this article shows that the IV estimand (that is, the limit in probability of the IV estimator instrumenting \(D\) by \(g(Z)\) in the regression of \(Y\) on \(D\)) is a convex combination of LATE for pairs of adjacent values of the instrument (where the values are ranked according to the conditional probability of being treated).\footnote{ Appendix (ref) (paragraph “The monotonicity assumption”) provides more details regarding this ranking. Note that, under the LATE assumptions, it is only a re-labeling without loss of generality. } If \(\textrm{Support}(Z) = \{z_0, z_1, \ldots, z_K\}\) (ordered such that \(k < m\) implies \(P(z_k) \leq P(z_m)\)),

equation[equation omitted — 413 chars of source]

with weights

equation*[equation* omitted — 192 chars of source]

that are non-negative and add up to one.

In particular, the stronger the effect of the instrument from \(z_{k-1}\) to \(z_k\) on the treatment (the higher \(P(z_k) - P(z_{k-1})\)), the more weight received by the LATE \(\delta^\textnormal{C}_{z_k, z_{k-1}}\) (which is \(\delta^\textnormal{C}_{z,w}\) of Theorem 1 with \(z = z_k\) and \(w = z_{k-1}\)) in the convex combination.

\paragraph{Estimation and inference}

We have only discussed identification issues so far. In practice, we also need estimation and inference results. That is why the third section of IA94 concludes with Theorem 3 which shows the asymptotic normality of the IV estimator and provides the expression of its asymptotic variance (when \(g(\cdot)\) is a known function; the Appendix of IA94 derives the asymptotic distribution when \(g(\cdot)\) depends on an unknown parameter that is jointly estimated).

\paragraph{Examples and discussion} The fourth and final section of IA94 (“Examples”) discusses the applicability of the conditions with three examples that “exploit the manner in which a particular program or treatment is implemented to create instruments that are exogenous. Evaluations of this type are sometimes referred to as natural experiments.” (IA94, p471, we italicize):

itemize[noitemsep,nolistsep] • “Draft Lottery,” that underlies notably angrist_1990; • “Administrative Screening,” with a situation where monotonicity can be doubtful; • “Randomization of Intention-to-Treat,” the typical textbook presentation of a randomized experiment with a binary instrument \(Z\) which is the randomized assignment to treatment.

Two-Stage Least Squares Estimation of Average Causal Effects in Models With Variable Treatment Intensity (AI95)

This 12-page article “generalizes [their] earlier result (imbens_angrist_1994) to models with variable treatment intensity” (AI95, p435). In IA94, the treatment \(D\) is binary, while in AI95, the treatment \(D\) (denoted \(S\) in the article for the number of years of schooling, which corresponds to their application) is a multi-valued quantitative ordered variable with finite support \(\{0, 1, \ldots, J\}\).\footnote{ Note that considering integer values for the support is without loss of generality: “\(S\) is assumed to take on only integer values between \(0\) and \(J\). It is enough, however, that \(S\) be bounded and take on a finite number of rational values. Then one can always use a linear transformation to ensure that \(S\) takes on integer values only between \(0\) and \(J\).” (AI95, p435). } This covers many applications where the treatment can be received in different doses, levels, or intensities (used as synonyms below).

In the introduction of AI95, J.A and G.I recall the widespread use of IV and TSLS estimators and the distinction between econometrics' structural equations versus statistics' potential outcomes to formalize causality. The second section presents the underlying application based on AK91. It provides a typical example of a treatment with various intensities, namely the number of years of schooling.

\paragraph{Average Causal Response (ACR) with a binary instrument}

The third section presents the main theoretical result of the article. At this stage, the authors consider a binary instrument \(Z\) and no control variables. \(D(z)\) (denoted \(S_z\) in the original) still represents the potential treatment variable.

As in IA94, Assumption 1 combines conditions (ref) and (ref) regarding the instrument \(Z\). Assumption 2 is the monotonicity assumption\textcolor{couleurModification}{ (ref)}. A novelty of AI95, specific to the multi-valued treatment case, is to discuss some testable implications of monotonicity. It cannot be tested as such since it involves unobserved counterfactual variables. However, J.A and G.I note that (ref), \(D(1) \geq D(0)\), implies that \(\mathbb{P}\!\left(D(1) \geq j\right) \geq \mathbb{P}\!\left(D(0) \geq j\right)\) for any level \(j \in \{0, \ldots, J\}\) of the treatment. If (ref) is maintained, for \(z \in \{0, 1\}\), \(\mathbb{P}\!\left(D(z) \geq j\right) = \mathbb{P}\!\left(D \geq j \,|\, Z = z \right)\), which is identified. Hence, a testable implication of (ref) (combined with (ref)) is that the cumulative distribution functions of \(D\) conditional on \(Z = 0\) and on \(Z = 1\) do not cross.

Theorem 1 of AI95 shows that the estimand of the IV estimator (instrumenting \(D\) by \(Z\) in the regression of \(Y\) on \(D\) and a constant) is equal to a convex combination of several LATEs across distinct levels of treatment. Formally, under (ref), (ref), (ref), and a form of the relevance condition, namely \(\mathbb{P}\!\left(D(1) \geq j > D(0)\right) > 0\) for at least one \(j \in \{0, \ldots, J\}\), the theorem writes

equation[equation omitted — 508 chars of source]

where, for any \(j \in \{1, \ldots, J\}\),

equation*[equation* omitted — 224 chars of source]

so that \(0 \leq \omega_j \leq 1\) and the weights add up to one.

comment\begin{equation} \underbrace{\frac{\operatorname{\mathbb{E}}[Y \,|\, Z = 1] - \operatorname{\mathbb{E}}[Y \,|\, Z = 0]}{\operatorname{\mathbb{E}}[D \,|\, Z = 1] - \operatorname{\mathbb{E}}[D \,|\, Z = 0]}}_{the identified IV estimand (\(\beta_{\textnormal{Wald}}\))} \;\, = \;\;\, \sum_{j = 1}^J \omega_j \, \operatorname{\mathbb{E}}[Y(j) - Y(j-1) \,|\, D(1) \geq j > D(0)] \, =: \, \delta^C_{ACR; 1, 0} \end{equation} where, for any \(j \in \{1, \ldots, J\}\), \begin{equation*} \omega_j := \frac{\mathbb{P}\!\left(D(1) \geq j > D(0)\right)}{\sum_{\ell = 1}^n \mathbb{P}\!\left(D(1) \geq \ell > D(0)\right)} \, \propto \, \mathbb{P}\!\left(D(1) \geq j > D(0)\right), \end{equation*} so that \(0 \leq \omega_j \leq 1\) and the weights add up to one.

The causal parameter \(\delta^\textnormal{C}_{\textnormal{ACR}; 1, 0}\) is obviously more complex than the basic LATE presented in (ref) which illustrates the additional layer introduced by a multi-valued \textcolor{couleurModification}{treatment}. The quantities \(\operatorname{\mathbb{E}}[Y(j) - Y(j-1) \,|\, D(1) \geq j > D(0)]\), \(j = 1, \ldots, J\) are actually LATEs as the population of individuals who switch treatment from less than $j$ to more than $j$ when the instrument jumps from 0 to 1 cannot be observed.

We remark that these LATEs \(\operatorname{\mathbb{E}}[Y(j) - Y(j-1) \,|\, D(1) \geq j > D(0)]\), \(j = 1, \ldots, J\), that are averaged (with weights \(\omega_j\)) can differ from one another in two dimensions:

enumerate[noitemsep,nolistsep] • the subpopulations of compliers, because a priori \(\{D(1) \geq j > D(0)\}\) is not equivalent to \(\{D(1) \geq \ell > D(0)\}\) for \(\ell \neq j\); • linearity of the causal effect is not assumed here; hence, the individual or average causal effect of a one-unit change from \(j-1\) to \(j\) is a priori different than the one from \(\ell-1\) to \(\ell\).\footnote{If linear causal effects were assumed, (ref) (see Appendix (ref) for details), the right-hand side of \(\delta^\textnormal{C}_{\textnormal{ACR}; 1, 0}\) would simplify to \(\sum_{j = 1}^J \omega_j \, \operatorname{\mathbb{E}}[\Delta \,|\, D(1) \geq j > D(0)]\), and only the variability of type (i) would arise.}

We further note that \(\delta^\textnormal{C}_{\textnormal{ACR}; 1, 0}\) can be reinterpreted as an average of individual causal effects from one level of the treatment to the next one, \(Y(j) - Y(j-1)\), which explains the characterization of \(\delta^\textnormal{C}_{\textnormal{ACR}; 1, 0}\) as “a weighted average of per-unit treatment effects along the length of a causal response function” (IA95, p431). J.A and G.I also refer to the \(\delta^\textnormal{C}_{\textnormal{ACR}; 1, 0}\) as “the average causal response (ACR)” (AI95, p435).

Finally, following the point made by a referee, J.A and G.I remark that, although the ACR is a weighted average, “it averages together components that are potentially overlapping” (IA95, p435). In general, nothing indeed prevents an individual from satisfying both \(\{D(1) \geq j > D(0)\}\) and \(\{D(1) \geq \ell > D(0)\}\) for \(\ell \neq j\). \textcolor{couleurModification}{For example, when the multi-valued treatment is the number of years of education and the binary instrument is living close enough to a college, an individual induced to go and graduate from college (thus reaching 16 years of education), but who would have completed only, say, 12 years of education (not going to college) had she lived far from a college, belongs to several complier populations: \(\{D(1) \geq j > D(0)\}\) for \(j = 13, 14, 15,\) and \(16\). In such a situation}, the individual contributes to several LATEs in \(\delta^\textnormal{C}_{\textnormal{ACR}; 1, 0}\). IA95 adds “In the schooling and other examples, however, most individuals would probably not be involved in an overlap of this sort, because the instrument would typically be expected to cause no more than a one-unit increment in treatment intensity for any particular individual.” (IA95, p436). There is no further comment from the authors regarding the issue. In all cases, like the monotonicity assumption, the point should induce researchers to be careful about the choice of the instrument and its effects on the treatment.

IA95 then presents a corollary of Theorem 1 for a causal interpretation of the IV estimand when a “variable [meaning multi-valued] treatment is incorrectly parameterized as a binary treatment” (IA95, p436). The result is interesting because it is a common situation in various applications; for instance, looking at the effect of high school graduation or college graduation as a summary of the number of years of education. In such cases, quantitatively, the estimation is upward biased; qualitatively, the sign remains, nonetheless, correctly estimated.

\paragraph{Extension to a multi-valued instrument and covariates}

The fourth section of IA95, “Multiple instruments and models with covariates”, presents two other identification results that generalize Theorem 1 when there can be multiple non-binary instruments and covariates. These additional results have the same flavor as those presented above in the sense that TSLS still identify positively weighted combinations of LATEs. Consequently, a concise summary of the different theorems of AI95 is the following. Under the LATE assumptions, TSLS applied to a causal model with variable treatment intensity identifies a weighted average of per-unit causal responses in a wide variety of models. We present those results below.

Theorem 2 of AI95 considers a non-binary instrument: \(K\) mutually exclusive binary instruments or, equivalently, a scalar instrument \(Z\) taking values in \(\{0, 1, \ldots, K\}\).\footnote{ Following the article, the encoding in \(\{0, 1, \ldots, K\}\) is used. Note that, equivalently, we could denote \(\textrm{Support}(Z) = \{z_0, z_1, \ldots, z_K\}\) as in IA94, and use a generic modality/value \(z_k\) instead of \(k\). Behind those choices, the setting is that of a multi-valued, finitely discrete instrument \(Z\). } As in IA94, the points of support of \(Z\) are ordered such that, for any \(k, m \in \{0, \ldots, K\}\), \(k < m\) implies \(\operatorname{\mathbb{E}}[D \,|\, Z = k] < \operatorname{\mathbb{E}}[D \,|\, Z = m]\), which combines the relevance and the monotonicity assumptions. We can then define the equivalent of \(\delta^\textnormal{C}_{\textnormal{ACR}; 1, 0}\), which is identified in the binary instrument case, for any pair \((k, k-1)\) of adjacent values of \(Z\): for any \(k \in \{1, \ldots, K\}\),

equation*[equation* omitted — 198 chars of source]

with \( \displaystyle \omega_{j,k} := \frac{\mathbb{P}\!\left(D(k) \geq j > D(k-1)\right)}{\sum_{\ell = 1}^n \mathbb{P}\!\left(D(k) \geq \ell > D(k-1)\right)}\), for any \(j \in \{1, \ldots, J\}\).

When instrumenting the variable treatment \(D\) by the multi-valued discrete instrument \(Z\), Theorem 2 shows that the estimand corresponding to the IV estimator is equal to a convex combination of ACRs across the different adjacent values of the instrument:\footnote{ The first-stage of the TSLS estimation procedure is the saturated regression of \(D\) on \(Z\) so that the theoretical predicted value used in the second-stage is the conditional expectation \(\operatorname{\mathbb{E}}[D \,|\, Z]\); hence the expression of \(\beta_{\textnormal{IV}}\), the estimand of the IV estimator in this case (see the last paragraph of Appendix (ref) for further details). }

equation[equation omitted — 433 chars of source]

where the weights satisfy \(0 \leq \mu_k \leq 1\) and \(\sum_{k = 1}^K \mu_k = 1\).

Moreover, \(\mu_k \propto (\operatorname{\mathbb{E}}[D \,|\, Z = k] - \operatorname{\mathbb{E}}[D \,|\, Z = k-1])\). Therefore, the stronger the instrument when varying between values \(k\) and \(k-1\), the more weight the corresponding ACR, \(\delta^\textnormal{C}_{\textnormal{ACR}; k, k-1}\), receives in the convex combination \(\delta^\textnormal{C}_{\textnormal{ACR}; Z}\).

Finally, Theorem 3 of AI95 adds discrete covariates \(X\), with a finite number of distinct values (so that, formally, the control covariates are included in the regression as the mutually exclusive indicators of each possible value \(x \in \textrm{Support}(X)\)). \textcolor{couleurModification}{Thanks to that restriction, for the first-stage of TSLS estimation, J.A and G.I can consider a saturated regression of \(D\) on \(Z\) and \(X\), that is, with as many parameters as possible values for the set of explanatory variables, \((Z, X)\) here.} By construction, the theoretical prediction of the endogenous treatment is then \(\operatorname{\mathbb{E}}[D \,|\, Z, X]\). \textcolor{couleurModification}{A saturated model for the covariates \(X\) happens to be crucial for the interpretation of the resulting TSLS estimand as an average causal effect on compliers. The next paragraph warns against possible misuses of the LATE in models with covariates that do not rely on a saturated specification. } In this context, their Theorem 3 shows that the resulting TSLS estimator has a probability limit, denoted \(\beta_{\textnormal{TSLS}}\), equal to a weighted average of the \(\delta^\textnormal{C}_{\textnormal{ACR}; Z}\) of Theorem 2 when they are defined conditionally on a value of the covariates:\footnote{ In the original paper, \(\delta^\textnormal{C}_{\textnormal{ACR}; Z}\) is denoted \(\beta_Z\), and the right-hand-side term of Theorem 3 (Equation (ref) here) is written \(\operatorname{\mathbb{E}}[\beta(X) \Theta(X)] \,/\, \operatorname{\mathbb{E}}[\Theta(X)]\) where “\(\beta(X)\) is the TSLS estimate, \(\beta_Z\), constructed using \(Z\) as an instrument in a population where \(X\) is fixed” (AI95, p437). }

equation[equation omitted — 260 chars of source]

with random weights

equation*[equation* omitted — 345 chars of source]

and where, for any \(x \in \textrm{Support}(X)\), \(\delta^\textnormal{C}_{\textnormal{ACR}; Z}(x)\) is defined as the equivalent of \(\delta^\textnormal{C}_{\textnormal{ACR}; Z}\) conditional on \(X = x\). From Theorem 2, \(\delta^\textnormal{C}_{\textnormal{ACR}; Z}(x)\) is thus identified by

equation*[equation* omitted — 324 chars of source]

Again, the (random) weights are connected to the strength of the instrument; namely, conditional on \(X\), the extent to which \(Z\) affects \(D\).

modification\paragraph{Warning: identification of LATEs in models with covariates} The articles of the LATE trilogy have opened the study of causal interpretation of IV/TSLS estimators. In particular, the last theorem of AI95 studies that problem in models with covariates. Causal interpretation of TSLS estimands with covariates as LATEs has not been entirely addressed in AI95 though and remains an active strand of the literature. A recent contribution, blandhol2022tsls, “When is TSLS Actually LATE?”, notably shows that in the absence of additional parametric assumptions on the effect of covariates a causal interpretation of a TSLS estimand as a LATE requires “saturated” specifications, which control for covariates nonparametrically. Without such additional assumptions, in specifications that are not saturated, the TSLS estimand mixes average causal effects for both compliers and noncompliers (always-takers or never-takers) and weighs some of the noncompliers average treatment effects negatively. In such cases, the TSLS estimand cannot be interpreted as an average causal effect on the subpopulation of compliers. That result is rather negative, all the more so as including covariates is frequent in practice and often important conceptually to support the validity of the instrument. In particular, the independence assumption (ref) might be doubtful as such, unconditionally, but more plausible when considered conditionally on covariates. It is therefore important to keep in mind the results of blandhol2022tsls: either additional parametric assumptions or a nonparametric control through saturated regressions are required for the usual LATE interpretation in models with covariates. It is worth noting that the only result of the LATE trilogy involving covariates (Theorem 3 of AI95) does use a saturated specification restricting to discrete covariates with a finite support.

\paragraph{Estimation of the weights and interpretation}

In addition to the LATE-type identification results of AI95, the authors show that the weights \(\omega_j / \omega_{j,k}\) are identified under the LATE assumptions. Moreover, they “can be consistently estimated from the difference between the empirical cumulative distribution functions of \(D\) given \(Z\)” (AI95, p436).

In the fifth section, “IV estimates of the returns to schooling: for whom?”, the authors illustrate the interest of such an estimation. As discussed above regarding Theorem 1, the weights add another layer to the interpretation of an IV estimate. In addition to the fact that this is a LATE (the average is over the subpopulation of compliers), the estimated weights reveal which compliers contribute most to the weighted average. Estimating and analyzing the weights thus permits a more comprehensive interpretation and understanding of the results.

For instance, J. Angrist and G. Imbens show that the weights associated with 12 to 16 years of schooling are lower for the 1920-1929 sample compared to the 1930-1939 sample. “Therefore, men who ended up completing some college because they were forced to graduate high school contribute more to the estimates for men born in 1930-1939 than to the estimates for men born in 1920-1929. This difference may explain the higher Wald and TSLS estimates for men born in 1930-1939 [\dots] because the returns to the last year of college tend to be substantially higher than those for any single year of high school (card_krueger_1992).” (AI95, p440).

Identification of Causal Effects Using Instrumental Variables (AIR96)

The last article of the LATE trilogy is co-authored with Donald Rubin. Its angle somewhat differs from IA94 and AI95.\footnote{ Also, in terms of mathematical modeling, contrary to IA94 and AI95, AIR96 uses a finite population setting: the potential variables are described as “fixed but unknown values” (AIR96, p446); “\(E[g]\) denotes the average over the population of \(N\) units of any function \(g(\cdot)\) [\dots] We emphasize that this notation simply reflects averages and frequencies in a finite population or subpopulation.” (AIR96, p447); the independence assumption (ref) is formulated in a survey-type random assignment of the instrument (AI96, Assumption 2, p446). Interestingly, more than two decades later, G. Imbens gets back to related questions with abadie2020sampling, where the authors discuss sampling-based as opposed to design-based uncertainty. } It considers the simplest setting (binary instrument, binary treatment, no covariates), and the first identification result (Proposition 1 of AIR96 in the third section “Causal estimands with instrumental variables”) is a particular case of Theorem 1 of IA94. On the other hand, AIR96 appears more precise and comprehensive as regards the underlying assumptions:

itemize[noitemsep,nolistsep] • it introduces the Stability and Unit Treatment Value Assumption (SUTVA, Assumption 1 of AIR96), which was only implicit in IA94 and AI95; • the exclusion restriction (Assumption 3) and the independence assumption (Assumption 2) are explicitly separated with the exclusion restriction expressed as a functional relationship (like in (ref)); • the relevance condition is interpreted in terms of the causal effect of \(Z\) on \(D\).

Besides, AIR96 proposes a more detailed comparison with the econometrics framework of simultaneous structural equations, notably in the second section, “Structural equation models in economics,” and the fourth, “Comparing the structural equation and potential outcomes frameworks” (see Section (ref) for more details on this thread). \paragraph{The textbook LATE theorem}

Although the least general result among those of the three LATE papers, Proposition 1 of AIR96 is probably the most well-known. Indeed, it is the textbook presentation of the LATE theorem: binary instrument, binary treatment, no covariates. This result is introduced and proved in Appendix (ref), Equation (ref). As explained above after Equation (ref), it is a direct corollary of Theorem 1 of IA94 when \(\textrm{Support}(Z) = \{0,1\}\):

equation[equation omitted — 202 chars of source]

\paragraph{Sensitivity to exclusion and monotonicity restrictions}

In contrast, Propositions 2 and 3 of AIR96 are two identification results that propose new results. Indeed, the fifth section, “Sensitivity of the IV estimand to critical assumptions,” derives the asymptotic bias of the IV estimand (compared to the targeted LATE, \(\delta^\textnormal{C}\)) when either exclusion (ref) or monotonicity (ref) is violated, while the other LATE assumptions are maintained.

These theoretical identification results have a practical interest as they give qualitative and quantitative insights for sensitivity analyses. As an illustration, the sixth section of AIR96, “An application: the effect of military service on civilian mortality,” “show[s] how the sensitivity of the estimated average treatment effect to violations of the exclusion restriction and the monotonicity assumption can be explored using the results from the previous section” (AIR96, p452). For instance, the authors discuss the required deviations from monotonicity to reverse the sign of the LATE. This exercise anticipates sign-reversal concerns that are currently at the heart of the literature studying difference-in-differences-type estimators.

Proposition 2 relaxes the exclusion restriction. Even before identification, this raises fundamental questions regarding the definition of causal parameters. For never-takers and always-takers, the instrument does not affect the treatment, \(D(0) = D(1)\), and AIR96 (Equation (13) of the paper) can define the individual causal effect of the instrument \(Z\) on \(Y\) for these individuals as \(H := Y(1, d) - Y(0, d)\), where \(d = 0\) (respectively \(d = 1\)) for a never-taker (resp. an always-taker). This presentation provides another interesting interpretation of the exclusion restriction: it ensures that the causal effect of \(Z\) on \(Y\) is null for always-takers and never-takers.\footnote{ Remark that, with a binary instrument and a binary treatment, conditional on \(\{D(0) = 0, D(1) = 1\}\), \(D = Z\); in that sense, the causal effect of \(D\) on \(Y\) coincides with the causal effect of \(Z\) on \(Y\) for a complier under (ref). }

The situation is more complicated for compliers. The authors explore that point through the assumption that assignment \(Z\) and treatment \(D\) have additive effects on the outcome \(Y\) for all compliers.\footnote{ As far as we understood, Proposition 2 of AIR96 relaxes the exclusion restriction (ref) for noncompliers (always-takers and never-takers) but maintains this restriction for compliers so that the LATE parameter, \(\delta^\textnormal{C}\), remains well-defined. However, “When there is a direct effect of assignment on the outcome for noncompliers, it is plausible that there is also a direct effect of assignment on outcome for complier” (AIR96, p451). The situation of Proposition 2 does not appear very credible; hence the relaxation of (ref) for anyone, including compliers, and the modeling using additively separable effects of \(Z\) and of \(D\) on the outcome to investigate the asymptotic bias. } They show that the bias relative to the average causal effect of \(D\) on \(Y\) for compliers positively depends on (i) the average size of the direct effect of \(Z\) on \(Y\) for noncompliers, \(\operatorname{\mathbb{E}}[H \,|\, D(1) = D(0)]\), and (ii) the odds of noncompliance, \(\mathbb{P}\{D(1) = D(0)\} / \, \mathbb{P}\{D(1) > D(0)\}\); remember that (ref) is still assumed to hold so there are no defiers.

On the other hand, Proposition 3 of AIR96 relaxes the monotonicity assumption. It shows that, in this case, the IV estimand is equal to

equation*[equation* omitted — 292 chars of source]

with \( \displaystyle \lambda \, := \, \frac{\mathbb{P}\{D(1) < D(0)\}}{ \mathbb{P}\{D(1) > D(0)\} - \mathbb{P}\{D(1) < D(0)\}} \).

In the setting of a binary instrument and a binary treatment, under the LATE assumptions except (ref), the IV estimand is thus a linear combination of the average causal effect on the compliers, \(\delta^\textnormal{C}\), and of the average causal effect on the defiers, \(\delta^\textnormal{D}\). However, the two weights, \(1 + \lambda\) and \(- \lambda\), are outside the unit interval \([0,1]\) because \(\lambda > 0\) whenever there are defiers. As a result, the IV estimand can be negative ({resp.} positive) even when \(\delta^\textnormal{C}\) and \(\delta^\textnormal{D}\) are both positive ({resp.} negative). \textcolor{couleurModification}{The monotonicity assumption is thus crucial to ensure the no-sign-reversal property of IVs.}

The asymptotic bias of IVs is proportional to the share of defiers in the population, \(\mathbb{P}\{D(1) < D(0)\}\). Another determinant is the extent of the heterogeneity of treatment effects: “The less variation there is in the causal effect of \(D\) on \(Y\), the smaller the bias from violations of the monotonicity assumption” (AIR96, p451); more specifically, in terms of heterogeneity between the average effect on compliers and on defiers. In particular, \(\delta^\textnormal{C} = \delta^\textnormal{D}\) is a straightforward sufficient condition to cancel the bias.

The Interpretation of Instrumental Variables Estimators in Simultaneous Equations Models with an Application to the Demand for Fish (AGI00)

commentThe literature on causal interpretations of TSLS is still very active. For instance, in addition to the whole literature devoted to TWFE mentioned in the introduction, recently, blandhol2022tsls, “When is TSLS Actually LATE?”, examines specifications including covariates. Absent parametric specification of the effect of covariates, it shows that the validity of the LATE interpretation crucially depends on using “saturated” specifications that control for covariates nonparametrically. Otherwise, in general, the TSLS estimand mixes average causal effects for both compliers and noncompliers (always-takers or never-takers) and, furthermore, with negative weights for some of the noncompliers average treatment effects. This is a rather negative result. It is worth noting that the only result of the LATE trilogy involving covariates (Theorem 3 of AI95) does use a saturated specification and restricts to discrete covariates.

In angrist_graddy_imbens_2000, “The Interpretation of Instrumental Variables Estimators in Simultaneous Equations Models with an Application to the Demand for Fish”, Review of Economic Studies (henceforth AGI00), the authors extend the LATE identification results to the setting of a continuous treatment, where the causal response function (that is, the potential outcome) \(d \mapsto Y(d)\) can be assumed to have a derivative. This section presents some of the prominent results from this paper. Overall, these can be seen as the continuous counterparts of Theorems 1 and 2 of AI95 (see Equations (ref) and (ref) above).

AGI00 develops its arguments and results in the setting of a demand-supply framework, considering, as a conventional choice, the demand function: what is the causal effect of a continuous treatment (price) on the outcome variable (quantity)? Price is endogenous due to classical simultaneity issues. The results are more general, and we adapt the notation for easier comparison.\footnote{ AIG00 considers demand (superscript \(d\)) and supply (\(s\)) functions/potential variables: \(q_t^d(p, z)\) and \(q_t^s(p, z)\), where \(p\) is any price (the equivalent of our \(d\)) and \(z\) any value of the instrument. The markets are indexed by \(t\), which plays the role of the unit/individual index \(i\) in our notation, and that we omit given i.i.d.-ness. } Also, for simplicity, we do not include covariates.\footnote{ Actually, before section 3.5 of AGI00, “Estimation with covariates,” covariates are formally present, but the results are only conditional on a given value \(x\) of the covariate \(X\) (which is equivalent to performing “unconditional” analyses on separate subsamples). Covariates are more explicitly introduced in section 3.5 with a parametric specification through AGI00's Assumption 5: “The average equilibrium price and quantity are linear and additive in covariates” (AGI00, p512); the corresponding result is Lemma 2 of AGI00 that we do not cover here. } Remember that, for any values \(d \in \textrm{Support}(D)\) and \(z \in \textrm{Support}(Z)\), \(Y(z,d)\) denotes a potential outcome variable. Under the exclusion restriction (ref), it does not depend on \(z\): \(Y(z,d) = Y(d)\), and the partial derivative with respect to \(d\) is the derivative of \(d \mapsto Y(d)\); we denote \(d \mapsto Y'(d)\) that derivative function. Following the application of AGI00 (non-negative prices for treatment), we consider \(\textrm{Support}(D) = [0, +\infty)\) but it could as well be the entire real line.

Theorem 1 of AGI00 shows that, in the case of a binary instrument, \(\textrm{Support}(Z) = \{0, 1\}\), the IV estimand is equal to

equation[equation omitted — 237 chars of source]

where the weighting function

equation*[equation* omitted — 205 chars of source]

is non-negative and integrates to one. This result is the continuous counterpart of AI95's Theorem 1.\footnote{ In Equation (ref), it was important that some inequalities were large (\(\geq\) some intensity \(j\) of the treatment) and other strict (\(> j\)). Here, the distinction does not matter since \(D\) and \(D(z)\), \(z \in \textrm{Support}(Z)\), are assumed to be continuous random variables, that is, admitting a density with respect to Lebesgue's measure. } The derivative \(Y'(d)\) can be interpreted as the individual marginal causal effect of \(D\) on \(Y\) at \(D = d\) (instead of the one-unit causal effect \(Y(j) - Y(j-1)\) of IA95 with a discrete treatment, \(\textrm{Support}(D) = \{0, 1, \ldots, J\}\)).

Then, AGI00 generalizes the previous result to the case of a discrete instrument with finite support \(\{z_0, z_1, \ldots, z_K\}\). As in IA94 and AI95, the distinct values of the instrument are ranked such that \(\operatorname{\mathbb{E}}[D \,|\, Z = z_k] < \operatorname{\mathbb{E}}[D \,|\, Z = z_m]\) for \(k < m\). Theorem 2 of AGI00 shows that, in this setting and under the LATE assumptions, the IV estimand, using some scalar function \(g(Z)\) as the instrument (like in IA94's second theorem), is a weighted average of the previous causal parameters across adjacent (in terms of conditional expectations of the treatment) values of the instrument:

equation[equation omitted — 499 chars of source]

For any \(k \in \{1, \ldots, K\}\) and \(d \in \textrm{Support}(D)\), the weights $\omega_k(d)$ satisfy

equation*[equation* omitted — 179 chars of source]

and the weights \((\alpha_k)_{k = 1, \ldots, K}\) add up to one and are proportional to

equation*[equation* omitted — 285 chars of source]

Moreover, the latter are non-negative for reasonable choices of \(g(\cdot)\), in particular for the choice \(g(z) = \operatorname{\mathbb{E}}[D \,|\, Z = z]\) (see AGI00, p511 for details). This is the choice made in AI95's Theorem 2 (a saturated regression for the first-stage). Thus, this result is, again, the continuous counterpart of AI95's Theorem 2.\footnote{ On top of the discrete/continuous treatment difference, AI95 uses the simpler re-encoding \(\textrm{Support}(Z) = \{0, 1, \ldots, K\}\) (and avoids therefore the introduction of the function \(g(\cdot)\)). }

Discussions, criticisms, and answers about the LATE

Following the presentation of the main identification results in Section (ref), this subsection focuses on the relevance of LATE-type parameters (Section (ref)), notably in contrast with more “structural” simultaneous equation models (Section (ref)). It also investigates the ability of LATEs to inform public policy decisions beyond the complier subpopulation. This question has been subject to controversies, as illustrated by the article “Better LATE Than Nothing: Some Comments on Deaton (2009) and Heckman and Urzua (2009)”, Journal of Economic Literature, imbens2010 (henceforth I10), which are evoked below.

What relevance of the LATE to inform public policy decisions?

The most debated point likely relates to the interest of the LATE regarding the information it provides to decide whether to implement a treatment (some public policy; for instance, a training program).\footnote{ For instance, James Heckman writes in his comment to AIR96: “LATE is a controversial parameter because it is defined for an unobservable subpopulation. Its use as an evaluation parameter thus is of questionable value.” (comments to AIR96, Journal of the American Statistical Association June 1996, Vol. 91, No. 434, page p459). } Related to this issue is the fact that compliers are not identified in the data. We briefly review some arguments connected to these debates in the setting of a binary treatment.

\paragraph{Relevant causal parameters}

The starting point is: what is the relevant causal parameter to inform a public policy decision? One answer appears indisputable: case-by-case analysis depending on the policy that is considered.

In some settings, the LATE might correspond to the exact parameter of interest, while it is the ATE (Average Treatment Effect, \(\delta := \operatorname{\mathbb{E}}[Y(1) - Y(0)]\)) or the ATT (Average Treatment effect on the Treated, \(\delta^{\textnormal{T}} := \operatorname{\mathbb{E}}[Y(1) - Y(0) \,|\, D = 1]\)) in other cases. For instance, in the context of A90 (effects of Vietnam military service on earnings), “One could imagine that the policy interest is in compensating those who were involuntarily taxed by the draft, in which case the compliers are exactly the population of interest. If, on the other hand, the question concerns future drafts that may be more universal than the Vietnam era one, the overall population may be closer to the population of interest” (I10, p414).\footnote{ In the example of this quote, however, we can wonder how such a compensation, whatever its amount, could be implemented since it would require identifying the compliers to know to whom the compensation should be sent. Remember that, as we never observe at the same time both \(D(0)\) and \(D(1)\), the compliers are not individually identified. Nonetheless, the point remains that the LATE could be the causal parameter of interest for the intention of specific policies. }

comment\footnote{ Those settings are perhaps not the most common but could be sensible, at least in the intention/theory of the public policy compared to its practice. Indeed, the following quotation is interesting but, concretely, we can wonder how such a compensation could be implemented since it would require to identify the compliers in order to know to whom the compensation should be sent. Yet, we cannot identify compliers individually without further assumptions. (It does not entirely prevent such a policy nonetheless as we could construct something like a probability of being a complier and use that as a threshold or a determinant of the compensation; but, for sure, it is not as simple as identifying the compliers.) That first answer is therefore not really an acceptable response of the LATE about being an interesting parameter. That being said, it is not a problem. The aim of the discussion here is essentially to convince that the contribution regarding identification of the LATE, by itself, does not have to answer to the interest of the LATE; it is another dimension. }

Even when the ATE or the ATT are of greater policy interest than the LATE,\footnote{ J.A and G.I ackowledge in various instances that the LATE has substantial limitations as a policy parameter; for instance, “Note that this group [compliers] need not be representative of the population, and that the members of this group cannot be identified from the data because membership involves unobserved counterfactual treatment status.” (AI95, p435). } the latter may still be the best one can hope for: identifying the ATE or the ATT relies on stronger assumptions than LATE identification and these assumptions may be too strong in some situations. In that case, recovering the LATE can be a reasonable alternative from which one can try to extrapolate and infer something on the ATE or the ATT. To sum up, as G. Imbens says, “better LATE than nothing”.

Another parameter often considered is the Intention-To-Treat (ITT). In the setting of a Randomized Controlled Trial (RCT) with imperfect compliance, where the binary instrument \(Z\) is the randomized assignment to the treatment, the ITT corresponds to the average causal effect of \(Z\) on \(Y\) and writes

equation*[equation* omitted — 101 chars of source]

Under independence (ref) only, the ITT is identified by \(\operatorname{\mathbb{E}}[Y \,|\, Z = 1] - \operatorname{\mathbb{E}}[Y \,|\, Z = 0]\). However, it is not always a relevant causal parameter for policymakers since it captures a fairly weak notion of average treatment effects: it measures how individuals react on average to being assigned into the treatment group. Because of imperfect compliance, this is distinct from the impact of getting treated. As explained above when comparing the LATE with the ATT and ATE, one could thus favor the LATE (or the ATT/ATE) over the ITT for policy-related reasons at the cost of more stringent identifying assumptions.

Overall, J. Angrist, G. Imbens, and D. Rubin underscore in their works that a trade-off exists between policy relevance of a parameter and ease of identification. They also prompt researchers to report several sets of estimators corresponding to the different parameters of interest to better inform public policy decisions.\footnote{ “It should be stressed, however, that the assumptions needed for a causal interpretation of the instrumental variables estimand (Assumptions 1 and 3-5) are substantially stronger than those needed for the causal interpretation of the intention-to-treat estimand (Assumption 1). The plausibility of the additional assumptions (i.e., the exclusion restriction and the monotonicity assumption) must be taken into account when facing the choice to report estimates of the intention-to-treat estimands, of the IV estimands, or both.” (AIR96, p450). }

commentA second answer is that, arguably, the ATE or the ATT is often more relevant than the LATE. Yet, be it ATE, ATT, or LATE, they are unknown parameters that need to be identified and estimated from the data. The main defense of the LATE could then resemble the following. In some situations, assumptions underlying the identification of the ATE or the ATT appear too far from reality to allow us to learn, say, directly something on the ATE or the ATT, whereas more plausible assumptions enable us to identify/learn the LATE. Granted, the LATE may not be the most relevant parameter.\footnote{ J.A and G.I make this point clear in various instances; for instance, “Note that this group [compliers] need not be representative of the population, and that the members of this group cannot be identified from the data because membership involves unobserved counterfactual treatment status.” (AI95, p435). } But, then, in a {second} step, up to you to try to learn indirectly something on the ATE or the ATT (or any parameter relevant to the contemplated public policy) by extrapolating one way or another from the LATE. To sum up, as G. Imbens says, “better LATE than nothing”.

\paragraph{Structural and causal approaches}

Under which assumptions and how to perform such an extrapolation is the next question. G. Imbens advocates that the two sets of assumptions regarding identification of the LATE on the one hand and extrapolation to other causal parameters more related to a given public policy decision on the other, should be separated as they are of different nature: “I would prefer to keep those assumptions separate and report both the local average treatment effect, with its high degree of internal but possibly limited external validity, and possibly add a set of estimates for the overall average effect with the corresponding additional assumptions, with lower internal, but higher external, validity” (I10, p415). In the previous quote, G.I distinguishes between two types of models:

itemize[noitemsep,nolistsep] • LATE-type approaches (labeled in I10 as “causal approaches”, including further extensions like regression discontinuity designs or difference-in-differences) which target one specific treatment parameter (some average of the individual causal effects) and, consequently, do not recover primitives of an underlying structural model, which is absent. • structural models which seek to identify the primitives behind the generic economic behavior of any individual. Provided primitives can be recovered, structural models are therefore able to perform richer counterfactual analyses than causal ones.

In a nutshell, causal approaches put more focus on internal validity (a credible estimation of a specific causal parameter on a given population from which the data is sampled) while structural approaches target, somewhat directly, external validity (a credible generalization of a causal effect to other populations).\footnote{ We extend the comparison between “causal” and “structural” approaches in Section (ref) below, which contrasts the way of expressing assumptions between the LATE framework and simultaneous equation models. }

commentG.I relates this discussion to the distinction between causal and structural models (see notably the sixth section, “Internal versus External Validity,” of I10): causal approaches put more focus on internal validity (a credible estimation of a specific causal parameter on a given population from which the data is sampled) while structural approaches target, somewhat directly, external validity (a credible generalization of a causal effect to other populations). We refer to the article for further details and simply recall what we think are interesting points to keep in mind about this debate:\footnote{ On another dimension of the comparison between “causal” and “structural” approaches, Section (ref) below contrasts the way of expressing assumptions between the LATE framework and simultaneous equation models. } \begin{itemize}[noitemsep,nolistsep] • By construction, LATE-type approaches (labeled in I10 as “causal approaches”, including further extensions like RDD, DID, etc.) only target one specific treatment parameter (some average of the individual causal effects) and, consequently, do not recover primitives of an underlying structural model, which is absent. • On the contrary (at least schematically), structural models describe, given unknown primitives, the behavior of the different units. Consequently, provided those primitives can be recovered, they are able to perform richer counterfactual analyses. \end{itemize}

\paragraph{Extrapolating from the LATE} We discuss here intuitively some general thoughts for extrapolating from the LATE.\footnote{ We come back to this question in Section (ref) when discussing angrist_fernandezval_2010 and angrist_rokkanen_2015. }

commentEARLIER VERSION WITH SHORT PRESENTATION OF angrist_fernandezval_2010

In general, the LATE, \(\delta^\textnormal{C} := \operatorname{\mathbb{E}}[Y(1) - Y(0) \,|\, D(1) > D(0)]\), is not equal to the ATE, \(\delta := \operatorname{\mathbb{E}}[Y(1) - Y(0)]\), nor to the ATT, \(\delta^{\textnormal{T}} := \operatorname{\mathbb{E}}[Y(1) - Y(0) \,|\, D = 1]\). Two separate dimensions contribute to this difference: (i) individual causal effects \(\Delta := Y(1) - Y(0)\) are in general heterogeneous; (ii) the different subpopulations of interest are, a priori, not the same.

Obviously, if effects are homogeneous (\(Y(1) - Y(0) = \delta_0 \in \mathbb{R}\) almost surely), then \(\delta^\textnormal{C} = \delta = \delta^{\textnormal{T}} = \delta_0\). Beyond that polar case, a limited heterogeneity supports the extrapolation of the LATE to other causal parameters. Provided several instruments, and thus several subpopulations of compliers, comparing LATE estimates across instruments is a way to assess the magnitude of heterogeneity. More generally, J.A and G.I insist on the importance of multiplying experiments or causal-approach studies to inform a public policy decision (see, for instance, the example in imbens2010, section 6, that illustrates this logic).

Another type of justification to extrapolate relates to (ii). Firstly, in the binary treatment binary instrument setting, extrapolation from the compliers to another population is possible in the particular case of one-sided noncompliance (in the terminology of RCT), where individuals not assigned to the treatment cannot be effectively treated. Formally, \(D(0) = 0\) for any individual, which implies (ref), but also the absence of always-takers. As a consequence, the subpopulation of treated individuals and the subpopulation of compliers coincide and \(\delta^\textnormal{C} = \delta^{\textnormal{T}}\): the LATE approach identifies the ATT.\footnote{ More precisely, the subpopulation of compliers assigned to the treatment. Yet, given that the assignment is (as-if) random (assumption (ref)), compliers assigned to the treatment (thus treated) and compliers not assigned (hence not treated) are comparable, on average, in terms of potential outcomes. } This case is important because of that theoretical result combined with the fact that \(D(0) = 0\) can be sensible in various applications.

More generally, it is sufficient to extrapolate that the compliers be “representative” of the group of interest, ideally, in terms of individual causal effects. If \(\mathrm{P}_{\Delta \,|\, D(1) > D(0)} = \mathrm{P}_{\Delta}\), the conditional distribution of treatment effects among compliers is the same as the marginal distribution (that is, among the entire population), then \(\delta^\textnormal{C} = \delta\) (and similar reasoning applies to the ATT). As \(\Delta\) cannot be observed for anyone, it is impossible to assess that directly. However, the idea is to compare, in terms of some observed covariates \(X\), compliers to the group of interest, say, the entire population, if we target the ATE. If those covariates are thought to be informative about the causal effect of \(D\) on \(Y\) and compliers and the other group are similar in terms of \(X\),\footnote{ Formally, when \(\mathrm{P}_{X \,|\, D(1) > D(0)} = \mathrm{P}_{X}\); compliers are said to be representative of the population as regards \(X\). } this supports using knowledge of the average individual causal effects among compliers to learn (partially) about the ATE. This second justification raises another question. \(X\) being observed, we can learn its distribution. However, since we cannot identify compliers individually, is it possible to learn about \(\mathrm{P}_{X \,|\, D(1) > D(0)}\)?

\paragraph{Learning about the compliers}

Although we cannot identify compliers individually, learning about the subpopulation is possible. First of all, under (ref) and (ref), the population proportion of compliers, \(\mathbb{P}\{D(1) > D(0)\}\), is identified and can be consistently estimated. Evrything else equal, the more numerous the compliers, the more likely they are representative of the whole population, the more grounded extrapolating from the LATE to the ATE. The proportion of the treated individuals who are compliers is also identified, which can be interesting, notably when trying to link the LATE to the ATT.

Furthermore, it is possible to identify the distribution of compliers' characteristics. For discrete covariates, Bayes's formula shows that the relative probability that compliers have a given characteristic \(X = x\), for any \(x \in \textrm{Support}(X)\), (relative to the entire population)

equation*[equation* omitted — 229 chars of source]

Under the LATE assumptions with (ref) extended to hold conditional on \(X\), the right-hand side is identified. More generally, abadie2003 is an important brick to the LATE contribution. Among other results, it shows that, for any real function \(g(\cdot)\) of \((Y, D, X)\), its expectation among compliers is identified since

equation[equation omitted — 207 chars of source]

where

equation*[equation* omitted — 177 chars of source]

The conditions for this result are the LATE assumptions holding conditional on the covariates (see Assumption 2.1 and Theorem 3.1 of abadie2003). In the formal sense of Equation (ref), kappa-weighting “detects compliers” and thus allows to identify virtually any characteristic of the complier subpopulation.

comment

\paragraph{Beyond average effects}

The previous result concerns any function of the observed variables \((Y, D, X)\). We could wonder what happens regarding the potential outcomes \(Y(0)\) and \(Y(1)\) (still in the setting of a binary instrument). In deaton2010, the author notes that in the case of an RCT with perfect compliance, identification of the ATE additionally relies on linearity of expectations. In fact, linearity of expectations implies that if we are interested only in the expectation of \(Y(1) - Y(0)\), knowledge of the two marginal distributions \(\mathrm{P}_{Y(0)}\) and \(\mathrm{P}_{Y(1)}\) is sufficient for identification. However, this is not the case if we are interested in features beyond the average, like distributional effects of the treatment, which require in general knowledge of \(\mathrm{P}_{Y(1) - Y(0)}\).

G. Imbens does not contest that reasoning but argues that a policymaker's typical interest lies in the two marginal distributions instead of the distribution of the difference.\footnote{ “In many cases, average effects of (functions of) outcomes are indeed what is of interest to policymakers, not quantiles of differences in potential outcomes. The key insight is an economic one -- a social planner, maximizing a welfare function that depends on the distribution of outcomes in each state of the world, would only care about the two marginal distributions, not about the distribution of the difference.” (I10, p409). } Besides, in the LATE framework, the marginal distribution of potential outcomes among compliers, \(\mathrm{P}_{Y(0) \,|\, D(1) > D(0)}\) and \(\mathrm{P}_{Y(1) \,|\, D(1) > D(0)}\), are identified. The result is shown in “Estimating Outcome Distributions for Compliers in Instrumental Variables Models”, The Review of Economic Studies, imbens_rubin_1997 (henceforth IR97). IR97 is an important extension to the LATE trilogy as it strengthens the identification result of the LATE theorem in the setting of a binary instrument and a binary treatment.

Another extension beyond average effects is quantile effects, developed in “Instrumental Variables Estimates of the Effect of Subsidized Training on the Quantiles of Trainee Earning”, Econometrica, abadie_angrist_imbens_2002 (henceforth AAI02). AAI02 focuses on the binary instrument binary treatment case with covariates and is one of the leading extensions of the LATE trilogy. Their Assumption 2.1 (AAI02, p91) is the classical LATE assumptions ((ref), (ref), (ref), (ref)) conditional on covariates \(X\) (plus non-trivial assignment: \(\mathbb{P}\!\left(Z=1 \,|\, X \right) \in (0,1)\)). Their Lemma 2.1 is the basis of all the LATE identification results derived in the paper:

equation[equation omitted — 104 chars of source]

The decisive consequence of (ref) is that “in the population of compliers, comparisons by \(D\) conditional on \(X\) have a causal interpretation.” (AAI02, p94). It was the case for averages; it is also the case for quantiles. AAI02 focuses on a linear model for conditional quantiles, which allows considering a single treatment effect (at any given quantile \(\tau \in (0, 1)\)), namely the difference between the \(\tau\)-quantiles of \(Y(1)\) and of \(Y(0)\) for compliers conditional on \(X\):

equation[equation omitted — 226 chars of source]

where \(\operatorname{\mathbb{Q}}_\tau\) denotes the \(\tau\)-quantile operator.\footnote{ To connect to the previous discussion, compared to the expectation \(\operatorname{\mathbb{E}}[\cdot]\), \(\operatorname{\mathbb{Q}}_\tau[\cdot]\) is not linear. In particular, \(\gamma_\tau^{\textnormal{C}}\) is not the \(\tau\)-quantile, conditional on \(X\), of \(\Delta := Y(1) - Y(0)\). } Relying on abadie2003's kappa function, AAI02 shows that the causal parameter \(\gamma_\tau^{\textnormal{C}}\) is identified in the LATE setting and proposes the quantile treatment effects (QTE) estimator to estimate it.

LATE and simultaneous equation models frameworks

For now, this review has focused on the identification results of the LATE trilogy and some of its main extensions. Another important aspect of these articles, also claimed as such by the authors, is to compare to the set-up of simultaneous (structural) equation models (SEM).\footnote{ It is in particular very present in AIR96, whose Sections 2 and 4 are entitled respectively “Structural Equation Models in Economics” and “Comparing the Structural Equation and Potential Outcomes Framework”. AGI00 also underscores this comparison focusing on a system of supply and demand equations: “A second contribution is the formulation of critical assumptions underlying identification of heterogeneous, time-varying demand functions in terms of potential demand and supply at different values of prices and instruments. This contrasts with most of the literature on simultaneous equations models, which casts critical assumptions in terms of unobservable functional-form-specific residuals.” (AGI00, p500). } J.A, G.I, and their co-authors claim that a critical contribution of the LATE framework is to make the identifying assumptions more transparent than in SEMs. This section briefly presents these arguments, following notably the exposition in AIR96. In the setting of a binary treatment, AIR96 considers a basic dummy endogenous variable model:\footnote{ We note that, in a comment to this article, J. Heckman proposes more advanced structural models. However, the target of this Section (ref), which is to illustrate the different types of assumptions as expressed in the LATE framework compared to the SEM formalization, remains in this simple setting. }

align[align omitted — 271 chars of source]
commentOTHER WRITING (taking more space) \begin{equation} D = \begin{cases} 1 & if D^* \geq 0, \\ 0 & if D^* < 0. \end{cases} \end{equation}

The first major difference is the formalization of causal effects. The LATE framework, following Neyman-Rubin's, defines them jointly with potential outcomes. In contrast, the parameter \(\beta_1\) represents the causal effect of \(D\) on \(Y\) in the SEM defined by (ref)-(ref)-(ref). It is possible to extend the model to introduce individual-specific causal effects (through random coefficients), hence heterogeneous causal effects. Nonetheless, we could say that, by default, LATE modeling posits heterogeneous causal effects by introducing \(\Delta := Y(1) - Y(0)\) as the building block, whereas SEMs more naturally consider a homogeneous causal effect through the parameter \(\beta_1\).

In SEMs, the relevance condition corresponds to \(\gamma_1 \neq 0\). There is not much debate about the formulation of this assumption as it is simpler conceptually and can be easily tested in the first-stage regression. We now consider the other LATE assumptions.

\paragraph{Exogeneity as opposed to independence and exclusion}

Much more debated is the standard SEM assumption of the “exogeneity” of the instrument, namely that \(Z\) is uncorrelated/orthogonal to the error terms or disturbances of the two equations (ref) and (ref):\footnote{ Without loss of generality due to the intercepts \(\beta_0\) and \(\gamma_0\), \(\varepsilon\) and \(\nu\) are centered; hence \(\operatorname{\mathbb{E}}[Z \varepsilon] = \mathbb{C}\mathrm{ov}(Z, \varepsilon)\) and \(\operatorname{\mathbb{E}}[Z \nu] = \mathbb{C}\mathrm{ov}(Z, \nu)\). }

equation[equation omitted — 135 chars of source]

The fact that \(Z\) does not intervene in Equation (ref) and the first part of (ref) (\(Z\) and \(\varepsilon\) are uncorrelated) capture the idea that \(Z\) affects \(Y\) only through \(D\). As a comparison, the LATE-type exclusion restriction (ref) is expressed through restricting the potential outcomes: \(Y(z,d) = Y(z',d)\). Condition (ref) also embeds the independence assumption (ref).

Conceptually, (ref) and (ref) are quite different assumptions. To assess (ref), “the researcher must consider, at the unit level, the effect of changing the value of the instrument while holding the value of the treatment fixed.” (AIR96, p449). In ordinary language, (ref) relates to the individual effect of \(Z\) on \(Y\). On the contrary, (ref) relates to the assignment mechanism of \(Z\), that is, how the instrument is determined and, most importantly, can it be considered “as-if” randomly determined. As a consequence, J. Angrist, G. Imbens, and D. Rubin consider that “pooling these assumptions into the single assumption of zero correlation between instruments and disturbances has led to confusion about the essence of the identifying assumptions and hinders assessment and communication of the plausibility of the underlying model” (AIR96, p450). It is noteworthy that deaton2010, while globally criticizing the LATE approach, also underscores this distinction by proposing and adopting different words: “external” for instruments that satisfy (ref) only, “exogenous” for instruments satisfying (ref) (and (ref)).\footnote{ “Whether any of these instruments is exogenous (or satisfies the exclusion restrictions) depends on the specification of the equation of interest, and is not guaranteed by its externality.” (deaton2010, p431). }

An additional point made in AIR96 is that assumptions such as (ref), cast in terms of moment conditions based on error terms, are hard to interpret from the start as those error terms are not straightforward to apprehend (without reference to potential outcomes).

\paragraph{Monotonicity}

Likewise, AIR96 defends that the LATE framework is more explicit about the monotonicity assumption. The argument may be less convincing here since, under homogeneous causal effects, the monotonicity is irrelevant. It thus appears severe to blame SEMs both for homogeneous causal effects and keeping implicit the monotonicity assumption. That being said, it remains interesting to notice that the use of a model with a constant parameter \(\gamma_1\) for the relation between \(Z\) and \(D\) (Equations (ref) and (ref)) implies that (ref) is automatically satisfied.

\paragraph{Equivalence and conciliation}

Let us conclude this Section (ref) with one remark and one result that may (or may not) reconcile the two positions. First, as noted by Robert Moffitt in his comments to AIR96, “Of course, one should not expect economists and statisticians, or even different individuals within each discipline, to find their intuition in the same way, and there is no reason not to have the model translated into multiple frameworks.” (Journal of the American Statistical Association June 1996, Vol. 91, No. 434, page 463). Second, vytlacil2002 shows an equivalence between the two frameworks in the sense that, given the classical LATE assumptions, it is possible to construct a latent-index model (as Equations (ref) and (ref)) that generates \(D(0)\) and \(D(1)\).

\newgeometry{left=5mm, right=5mm, top=5mm, bottom=5mm} \afterpage{ \thispagestyle{empty}

landscape\begin{tabular}{|p{4.2cm}||p{4.4cm}|p{5cm}|p{2.6cm}|P{9.6cm}|} \hline {Results and settings} & {Treatment} \(D\) & {Instrument} \(Z\) & {Covariates} \(X\) & {Identified causal parameter} \\ \hline \hline AIR 1996 Proposition 1\newline { (see (ref) and (ref))} & binary \newline \(\textrm{Support}(D) = \{0,1\}\) & binary \newline \(\textrm{Support}(Z) = \{0,1\}\) & none & \(\underbrace{\operatorname{\mathbb{E}}[Y(1) - Y(0) \,|\, D(1) > D(0)]}_{=: \, \delta^\textnormal{C} \,=\, \delta^\textnormal{C}_{1,0}, \text{ ``the'' Local Average Treatment Effect (LATE)}}\) \\ \hline IR 1997 Section 3 \newline { (see (ref))} & binary \newline \(\textrm{Support}(D) = \{0,1\}\) & binary \newline \(\textrm{Support}(Z) = \{0,1\}\) & none & \( \underbrace{\mathrm{P}_{Y(0 \,|\, D(1) > D(0)} \text{ and } \mathrm{P}_{Y(1 \,|\, D(1) > D(0)}}_{\text{the marginal distributions of potential outcomes among compliers}} \) \\ \hline AAI 2002 Section 3.1 \newline { (see (ref))} & binary \newline \(\textrm{Support}(D) = \{0,1\}\) & binary \newline \(\textrm{Support}(Z) = \{0,1\}\) & conditionally \newline on \(X\) & \( \underbrace{\operatorname{\mathbb{Q}}_{\tau}[Y(1) \,|\, D(1) > D(0), X] - \operatorname{\mathbb{Q}}_{\tau}[Y(0) \,|\, D(1) > D(0), X]}_{=: \, \gamma_\tau^{\textnormal{C}}, \text{ local \(\tau\)-quantile treatment effect on compliers}} \) \\ \hline IA 1994 Theorem 1 \newline { (see (ref))} & binary \newline \(\textrm{Support}(D) = \{0,1\}\) & two modalities \(z\) and \(w\) \newline of a discrete multi-valued \(Z\) & none & \( \underbrace{\operatorname{\mathbb{E}}[Y(1) - Y(0) \,|\, D(z) \neq D(w)]}_{=: \, \delta^\textnormal{C}_{z,w}} \) \\ \hline IA 1994 Theorem 2 \newline { (see (ref))} & binary \newline \(\textrm{Support}(D) = \{0,1\}\) & discrete multi-valued, finite \newline \(\textrm{Support}(Z) = \{z_0, z_1, \ldots, z_K\}\) \newline ranked by increasing conditional \newline on \(Z\) probability of being treated & none & a convex combination over adjacent \(Z\)-values of LATE \newline \( \sum_{k = 1}^K \lambda_k \underbrace{\operatorname{\mathbb{E}}[Y(1) - Y(0) \,|\, D(z_k) = 1, D(z_{k-1}) = 0]}_{= \, \delta^\textnormal{C}_{z_k, z_{k-1}}} \) \\ \hline AI 1995 Theorem 1 \newline { (see (ref))} & multi-valued, ordered finite \newline \(\textrm{Support}(D) = \{0, 1, \ldots, J\}\) & binary \newline \(\textrm{Support}(Z) = \{0,1\}\) & none & a convex combination over \(D\)-values of one-level-change LATE \newline \( \underbrace{\sum_{j = 1}^J \omega_j \, \operatorname{\mathbb{E}}[Y(j) - Y(j-1) \,|\, D(1) \geq j > D(0)]}_{=: \, \delta^\textnormal{C}_{\textnormal{ACR}; 1, 0}, \text{ Average Causal Response (ACR)}} \) \\ \hline \textbf{IA 1995 Theorem 2} \newline { (see (ref))} & multi-valued, ordered finite \newline \(\textrm{Support}(D) = \{0, 1, \ldots, J\}\) & discrete multi-valued, finite \newline \(\textrm{Support}(Z) = \{0, 1, \ldots, K\}\) \newline ranked by increasing conditional \newline expectation of \(D\), \(\operatorname{\mathbb{E}}[D \,|\, Z = \cdot]\) & none & \(\delta^\textnormal{C}_{\textnormal{ACR}; Z}\): a convex combination over adjacent \(Z\)-values of ACR \newline \( \sum_{k=1}^K \mu_k \big( \underbrace{\sum_{j = 1}^J \omega_{j,k} \operatorname{\mathbb{E}}[Y(j) - Y(j-1) \,|\, D(k) \geq j > D(k-1)]}_{=: \, \delta^\textnormal{C}_{\textnormal{ACR}; k, k-1}} \big) \) \\ \hline \textbf{IA 1995 Theorem 3} \newline { (see (ref))} & multi-valued, ordered finite \newline \(\textrm{Support}(D) = \{0, 1, \ldots, J\}\) & discrete multi-valued, finite \newline \(\textrm{Support}(Z) = \{0, 1, \ldots, K\}\) \newline ranked by increasing \(\operatorname{\mathbb{E}}[D \,|\, Z = \cdot]\) & yes, \textcolor{couleurModification}{with a} \newline \textcolor{couleurModification}{saturated model} \newline { (discrete \(X\) with finite support)} & \( \operatorname{\mathbb{E}}[\delta^\textnormal{C}_{\textnormal{ACR}; Z}(X) \, \Theta(X)] \, / \, \operatorname{\mathbb{E}}[\Theta(X)] \): a weighted average of \newline \(\delta^\textnormal{C}_{\textnormal{ACR}; Z}(x)\) (the equivalent of \(\delta^\textnormal{C}_{\textnormal{ACR}; Z}\) conditional on \(X = x\)) \newline with random weight \(\Theta(X) = \mathbb{V}\{ \operatorname{\mathbb{E}}[D \,|\, Z, X] \,|\, X \}\) \\ \hline \textbf{AGI 2000 Theorem 1} \newline { (see (ref))} & continuous \newline \(\textrm{Support}(D) = [0, +\infty)\) & binary \newline \(\textrm{Support}(Z) = \{0,1\}\) & none \newline { (conditionally on \(X\) in original)} & a (continuous) convex combination of LA marginal TE \newline \( \underbrace{\int_0^{+\infty} \operatorname{\mathbb{E}}[Y'(d) \,|\, D(1) \geq d \geq D(0)] \, \omega(d) \, \textrm{d} d}_{=: \, \delta^\textnormal{C}_{\textnormal{AmarginalCR}; 1, 0}, \text{ Average Marginal Causal Response (AMCR)}} \) \\ \hline \textbf{AGI 2000 Theorem 2} \newline { (see (ref))} & continuous \newline \(\textrm{Support}(D) = [0, +\infty)\) & discrete multi-valued, finite \newline \(\textrm{Support}(Z) = \{z_0, z_1, \ldots, z_K\}\) \newline ranked by increasing conditional \newline expectation of \(D\), \(\operatorname{\mathbb{E}}[D \,|\, Z = \cdot]\) & none & a convex combination over adjacent \(Z\)-values of AMCR \newline \( \sum_{k = 1}^K \alpha_k \underbrace{ \int_0^{+\infty} \operatorname{\mathbb{E}}[Y'(d) \,|\, D(z_k) \geq d \geq D(z_{k-1})] \, \omega_k(d) \, \textrm{d} d}_{=: \, \delta^\textnormal{C}_{\textnormal{AmarginalCR}; z_k, z_{k-1}}} \) \\ \hline \end{tabular} \captionof{table}{{Summary of the LATE trilogy results and of some extensions}. The results are ordered, firstly, by the richness of the treatment \(D\): binary and qualitative; quantitative in an ordered finite set; quantitative and continuous. In these two latter cases, remark that the linearity of causal effects (ref) is not assumed (for binary treatment, linearity is not applicable/automatic). Secondly, by the richness of the instrument \(Z\). We follow the original presentation of IA95 with integer values for \(Z\); it could be extended to a general finite set \(\textrm{Support}(Z) = \{z_0, z_1, \ldots, z_K\}\) at the cost of considering a scalar instrument \(g(Z)\). All results are stated under the LATE assumptions: (ref), (ref), (ref), (ref) (or (ref)), possibly conditional on \(X\).}

} \restoregeometry

The authors' influence beyond LATE

Joshua Angrist and Guido Imbens have been awarded the Nobel Memorial Prize in Economic Sciences mostly for the articles about LATE identification they coauthored in the 1990s. This series of works, however, represents only a small fraction of these authors' rich input to economics research. Perhaps surprisingly, J. Angrist and G. Imbens have hardly ever coauthored beyond their LATE series of articles. In this section, we give a detailed account of the numerous fundamental contributions these two economists have made in their respective fields.

Joshua Angrist

Zooming out on J. Angrist's career, the early 2000s can be seen as a turning point. In 2000 and 2002, angrist_graddy_imbens_2000 and abadie_angrist_imbens_2002 are the last two collaborations between J. Angrist (J.A) and G. Imbens (G.I), both on topics related to the identification of local average treatment effects in instrumental variable frameworks. These articles ended a decade of joint works that laid the foundations of the LATE framework. In this subsection, we present J.A's contributions from 2000 onwards not coauthored with G.I, a strand of works we refer to as post-LATE.

2000-2010: teachers' quality, labor market issues, and much more

\paragraph{Looking for informative and varied empirical research designs}

The 2000s decade is probably the most prolific period for J.A in terms of applied topics covered by his research: between 2000 and 2010, he made impactful contributions in fields as diverse as education angrist_et_al_2002, labor acemoglu_angrist_2001 or criminology angrist_2006.

All these articles have in common to exploit randomized or natural experiments that display large variations in treatment exposure across populations under study. These experiments are, therefore, adequate to measure the causal impact of economic policies by resorting to the potential outcome framework presented at length in Section (ref). Looking for good natural or randomized experiments is indeed at the heart of J. Angrist's research agenda. Quoting J.A and his coauthor Jörn-Steffen Pischke, in their celebrated review article “The Credibility Revolution in Empirical Economics: How Better Research Design is Taking the Con out of Econometrics”, Journal of Economic Perspectives, angrist_pischke_2010: “{Empirical microeconomics has experienced a credibility revolution, with a consequent increase in policy relevance and scientific impact. Sensitivity analysis played a role in this, but as we see it, the primary engine driving improvement has been a focus on the quality of empirical research designs.}” (page 4).

J.A not only looks for good experiments but also for varied ones which allow building knowledge on specific questions across distinct environments. In J. Angrist's view, this is the most efficient tool to understand how much external validity specific estimates contain. Another illuminating quote from angrist_pischke_2010 illustrates this well: “{A constructive response to the specificity of a given research design is to look for more evidence, so that a more general picture begins to emerge.}” (page 23).

To measure possible effects present in the experimental data he uses, J.A favors simple linear models, controlling for endogeneity of the policy if needed and for fixed effects in possibly multiple dimensions. In our views, such choices are made (i) to promote ease of interpretation of estimated effects and (ii) to stress that it is the experiment rather than convoluted modeling restrictions that are the driving force of policy effect identification. As we discuss below, the simple modeling approaches promoted by J.A however sometimes call for a cautious causal interpretation of the estimated effects.

\paragraph{Education and lotteries}

One applied topic, already present in J. Angrist's works of the 1990s, is still dominant in his research effort of the 2000s: the issue of returns to education. J.A's interest focuses mainly on two more specific sub-questions: the impact of teachers' quality and financial incentives or aids on educational attainments.

The interplay between teachers' quality and educational outcomes is addressed in two articles, angrist_lavy_2001 and angrist_guryan_2004. Those two articles exploit natural experiments, in Jerusalem and in the United States, respectively. In angrist_lavy_2001, the authors take advantage of a sudden increase in funds for on-the-job training of teachers that affected only a fraction of Jerusalem's schools in 1995. This shock is then used to study the effect of teachers' training on pupils' achievement at school. The authors conclude to a positive impact of teachers' training on pupils' results for some types of schools only (the non-religious ones). In angrist_guryan_2004, state-level variations in requirements to become a teacher in the US are exploited to analyze whether enforcing standards for teachers has consequences on the quality of teachers actually hired. The conclusion of this article is rather negative on the usefulness of testing teachers prior to hiring: these tests tend to increase teachers' wages but have little effect on teachers' quality. These results have been followed by numerous other analyses trying to assess how robust these findings were (e.g., boyd_2007 and goldhaber2007). In angrist_lavy_2001 and angrist_guryan_2004, endogeneity concerns are limited so that no instrumental variable approach is applied, which is the exception rather than the rule in J. Angrist's works.

The question of how relevant financial incentives are to shape pupils' or students' educational trajectories is tackled in four articles published by J.A between 2000 and 2010: angrist_et_al_2002, angrist_et_al_2006, angrist_lavy_2009, and angrist_lang_oreopoulos_2009. Among those, angrist_et_al_2002 and angrist_et_al_2006 are of particular interest because of the nature of the experiment at play: in the 1990s, a large-scale educational program involving 125,000 pupils from low-income families took place in Colombia. Targeted pupils were offered very generous vouchers that were renewable over the years conditional on having high enough grades. In certain areas, the demand for vouchers was larger than the supply by local authorities, which pushed the latter to introduce a lottery system to select among eligible pupils. As J. Angrist and his coauthors note, some lottery winners chose not to accept the voucher so that the lottery is an instance of a randomized experiment with imperfect compliance. These researchers take advantage of this fairly challenging research design thanks to a very clever idea: they remark that lottery winners have a much higher probability of resorting to financial aid of any type (vouchers, scholarships, etc.) than lottery losers. They thus interpret lottery winning as a credible instrument for financial incentives to study, which they use to analyze the effect of financial incentives on educational attainments. In other words, the analysis is recast in the canonical LATE framework with a binary instrument (winning the lottery or not) and a binary treatment (resorting to financial aid or not). The result primarily put forward by the authors is a positive and significant impact of financial aid on the propensity to complete secondary school.

Using lotteries as an informative source of exogenous variation to identify causal relationships is not new to J. Angrist: this idea already featured in his early works on the civil life consequences of GI conscription angrist_1990. angrist_et_al_2002 and angrist_et_al_2006 are however the first works by J.A that take advantage of lotteries to investigate school-related questions. As we will see below, lottery-based experiments will be even more at the heart of J.A's research agenda on school-related topics in the 2010s. Among J. Angrist's other interesting contributions to school topics, we can cite angrist_lang_2004. This article pioneers J.A's research interests of the 2010s: it is indeed J. Angrist's first work that deals with an experiment aimed at improving the school prospects of children living in impoverished inner-city districts in the US (here Boston).

\paragraph{Labor market issues}

J. Angrist also investigates several hot labor market questions in the 2000s decade, notably in these two articles: acemoglu_angrist_2001 and angrist_kugler_2003.

In acemoglu_angrist_2001, the authors investigate the consequences of the Americans with Disabilities Acts (ADA), a US federal ruling active since 1992 that seeks to promote employment of people with disabilities in the US. Given that the act targets all disabled people, this group is interpreted as a treatment group, and a difference-in-difference (DID) analysis is conducted.\footnote{ DIDs are an identification and estimation strategy to recover average treatment effects in natural experiments where two groups (one treatment and one control group) are observed at two periods, and one group is affected by a treatment between the two periods. The idea of DIDs can be traced back to the 19th century snow1855. It has been very popular among applied economists since its use by David Card in the early 1990s. David Card was awarded the Nobel Memorial Prize in Economic Sciences in 2021, along with Joshua Angrist and Guido Imbens (we refer to the dedicated article about D. Card's contributions in the same issue of this journal). } The authors report a significantly negative effect of the act on the employment of disabled men.

angrist_kugler_2003 studies the consequences of immigration on the labor market prospects of natives in different European countries in the 1990s. This question is difficult: immigration to and natives' employment in a country are both partly determined by the overall economic environment in that specific country, making immigration an endogenous predictor of natives' employment. J. Angrist and Adriana Kugler solve that puzzle by exploiting a massive and unexpected immigration shock that occurred in Europe in the 1990s. Between 1991 and 2001, the Yugoslav Wars put hundreds of thousands of civilians on the road who fled massively to nearby Western Europe. Instrumenting immigration with these Yugoslav Wars, the authors point to a negative impact of immigration on natives' employment, especially in countries with rigid economic institutions (high degree of protection of workers, for instance). This article is quite similar in spirit to card1990. In this path-breaking project, David Card uses an unexpected and large immigration wave from Cuba to Florida in 1980 to unveil the link between immigration and labor market outcomes in the US.

\paragraph{Methodological concerns and contributions}

As highlighted above, J.A's research agenda of the 2000s places much emphasis on looking for good experiments in order to make credible causal claims. Are there other significant methodological outputs for policy analysis that have emerged from this research agenda? The answer is ambiguous. In most of his articles, J. Angrist deals with individual or group-level panel data where treatment can be suspected to be endogenous, typically because of noncompliance issues. He centers his analysis on linear IV strategies to correct for possible endogeneity of the treatment variable of interest, and he typically adds control variables and fixed effects in multiple dimensions (time, individual/group, etc.) to partial out systematic heterogeneity, see for instance angrist_kugler_2003.

Following the seminal LATE articles of the 1990s, we would be tempted to interpret the second-stage coefficient on the endogenous treatment as some (local) average treatment effect. This is, however, more complex. The data designs used by J. Angrist in his applied works are much richer than those considered in imbens_angrist_1994, angrist_imbens_1995 or angrist_imbens_rubin_1996. It is a priori unclear whether the findings of these articles extend to much more complex designs with covariates and multiple fixed effects. This puzzle has, in fact, been solved very recently.

The ongoing two-way fixed effects (TWFE) literature addresses exactly that question and shows that running a linear regression with a treatment variable and time and individual fixed effects has some causal interpretation only if the heterogeneity of treatment effects across individuals and over time is severely restricted.\footnote{ We elaborate on the connection between LATE-type identification results and the TWFE literature in the conclusion (see Section (ref)). } J. Angrist was unaware of these results in the 2000s, and the difficulty of extending the LATE formalism to richer data designs is palpable in his projects of that period. In angrist_lavy_2001, the authors remark it is natural to impose a constant treatment effect assumption to rationalize the use of a standard linear DID regression with covariates. In angrist_et_al_2002, the experiment under study involves one binary treatment plagued with noncompliance and one binary instrument, which is similar to the seminal LATE framework. However, the authors hardly mention the connection with LATE and the potential outcome formalism. The discussion about the causal interpretation of the IV parameter in their application remains quite minimal as well. The limitations of linear IV models with covariates and fixed effects are, however, counterbalanced in J.A's works by frequent use of separate analyses on subpopulations defined by age or gender (e.g., angrist_et_al_2002, angrist_kugler_2003 or angrist_lavy_2009). Intuitively, running separate regressions for distinct gender-age groups restores some robustness of the IV specification against heterogeneity of treatment effects along the gender-age dimension.

As discussed in this subsection, the topics of interest to J. Angrist in the 2000s are more applied than purely methodological. Still, he proposes a few methodological contributions that complement his foundational works of the 1990s. In angrist_et_al_2006, for instance, nonparametric bounds on the distribution of treatment effects in the presence of noncompliance are proposed and applied to the school lottery experiment rolled out in Colombia in the early 1990s. angrist_cherno_fernandezval_2006 is another interesting econometric contribution. This article tackles the following simple question: what is the interpretation of running a linear quantile regression of an outcome on a set of covariates if the true quantile regression is not linear? Even though this project is not concerned with causality, there is an obvious connection with the LATE articles of the 1990s that seek to give a meaningful (causal) interpretation of the linear IV model even when the data is not generated according to such a model. angrist_cherno_fernandezval_2006 is quite thought-provoking from an econometrician's point of view: the authors show that linear quantile regression can always be interpreted as a weighted least square estimator, with weights that depend on the true quantile regression function.

From 2010 onwards: some more methodology, re-inspection of the 1990s' applied works plus an in-depth analysis of the US educational system

The research produced by J. Angrist since 2010 has three distinctive features in our view: (i) more emphasis on the causal interpretation of methods used in applications than in the previous decade, (ii) a reanalysis of some applied questions addressed by the author in the 1990s, and, most importantly, (iii) an active contribution to the public debate over education in the US. We present each of these features in detail below.

\paragraph{Back to the LATE framework and its external validity}

In the 2000s, the connection between the IV methodologies used in practice by J. Angrist and his coauthors and the causal framework laid down in the 1990s seemed less present. In the 2010s, this link was unquestionably reasserted. The articles angrist_fernandezval_2010 and angrist_rokkanen_2015 illustrate this phenomenon well. These articles give simple conditions under which LATE parameters, identified using IV methods, can be aggregated to recover global parameters such as average treatment effects on the treated or in the general population.\footnote{ angrist_rokkanen_2015 is not solely concerned with IV setups. This article focuses on regression discontinuity designs (RDD) which exploit thresholds (on wages, class sizes, etc.) that split otherwise similar populations into two groups, one that benefits from some policy (say above the threshold) and the other that remains unaffected. As explained in imbens_lemieux_2008, RDD allows researchers to identify average treatment effects at the threshold when treatment compliance is perfect and average treatment effects for compliers at the threshold otherwise. angrist_rokkanen_2015 proposes a framework to recover (local) treatment effects away from the threshold in RDD. } The conditions are necessarily restrictive: the present articles impose that all the differences between compliers and the rest of the population can be captured through observable covariates. Still, these papers simultaneously allow for a certain form of treatment effect heterogeneity and to recover both local and “global” average treatment effects in situations where those quantities need not be equal. angrist_fernandezval_2010 and angrist_rokkanen_2015 therefore contribute to improving the LATE framework on its arguably most controversial aspect, namely what external validity does a local treatment effect contain angrist_pischke_2010, imbens2010.

\paragraph{Re-inspection of 1990s' applied topics}

Angrist also reinspects topics he worked on in the 1990s: the civilian consequences of veteran status in the US angrist_chen_frandsen_2010, angrist_chen_2011, the socio-economic impact of having extra children on parents or already-born children angrist_lavy_schlosser_2010, and the interplay between class size and students' achievement angrist_lavy_etal_2019. In angrist_chen_frandsen_2010, angrist_chen_2011, and angrist_lavy_etal_2019, the authors use augmented versions of the datasets used in the original papers. In angrist_lavy_schlosser_2010, the analysis is conducted using Israeli data while the seminal work angrist_evans_1998 deals with American data.

These “replication” exercises are necessary to assess the long-term stability of previously estimated effects. J.A and his coauthors indeed conclude that some significant effects observed in the 1990s' papers cannot be recovered with more recent data, for instance, the positive impact of being a US army veteran on civilian earnings angrist_chen_2011 or the negative link between class size and educational attainments angrist_lavy_etal_2019.

\paragraph{Achievement gap in the US educational system}

Educational issues were at the center of J. Angrist's attention in the 2000s and still are in the 2010s decade. For more than a decade, he has been especially interested in the US secondary school system. The latter is a much-commented topic in American public life (see angrist_etal_boston_2012 for some contextual elements thereon). A reason for this can be found in the huge disparities in secondary school achievement in American cities, with poor and mostly nonwhite inner cities underperforming considerably compared to more affluent and mostly white suburbs. This phenomenon is called the achievement gap. In many US cities, the standard public school system competes with various alternatives: fully-private schools, exam schools (which are elite public schools with selection based on grades), and independent public schools that target children from impoverished backgrounds (known as charter schools). There is a heated debate over the relative merits of those alternatives in closing the achievement gap. In a still ongoing series of works, J.A, along with coauthors, actively takes part in this debate by providing evidence on the achievement trajectories of students attending different types of schools (see angrist_2022 for instance for an exhaustive review of these works). This research program is quite unique in its scope, with 13 articles coauthored by J.A on the said subject since 2010. It is conducted for the most part by researchers at Blueprint Labs\footnote{ Blueprint Labs is a recently created policy-evaluation lab hosted by the MIT and directed by Joshua Angrist. } and consists in investigating the performance of school systems in many big cities in the US: Boston, New York City, Chicago, Denver, New Orleans, \dots, and the list is still growing. To do so, researchers at Blueprint Labs have access to very rich data both on students and schools in the cities under study. As explained below, those researchers seek to develop a coherent set of econometric tools that ensure comparability across evaluations. Having that in mind, we can say that this series of works epitomize J. Angrist and G. Imbens' repeated calls for replicating similar experiments across environments to gain some insights into the external validity of effects measured locally angrist_pischke_2010, imbens2010.

The starting point of the research conducted at Blueprint Labs is the following: local school systems in the US offer a number of (quasi-)experimental variations in the allocation of students to schools that can be exploited to reconstruct potential achievement trajectories in alternative schools for the same student. These variations stem from the massive over-subscription faced by the most appealing schools, in particular exam and charter schools. Exam schools offer seats based on test scores. We can expect that marginally rejected students and those who are offered the last seats are quite indistinguishable, a situation amenable to a regression discontinuity analysis. Charter schools do not select their students based on school achievement, but because of over-subscription, they are forced to organize lotteries to offer seats. Lotteries are random by construction, so students who are offered a seat should be similar to those who are not. In both cases, students who are offered seats do not always accept them, which creates possible endogeneity issues that can be solved by using typical IV techniques yielding LATE-type identified treatment effects. In some studies, the analysis is made significantly more challenging by the design of certain school systems: students and schools in Chicago, New York, or Denver are allocated to one another through a centralized Deferred Acceptance algorithm gale_shapley_1962 so that it may be difficult to trace back the origins of observed choices ultimately made by students and schools. J. Angrist and his coauthors thus propose extensions to traditional potential outcome models that allow to identify and estimate clearly-defined local treatment effects in these complex designs abdulkadiroglu_etal_elite_2014, abdulkadiroglu_etal_denver_2017. The methods they come up with are fairly easy to use and nicely trade-off the richness of the analysis and practical considerations.

The research conducted on US secondary schools by J. Angrist and his coauthors has produced unexpected and strikingly consistent results across a wide range of environments. Charter schools have been shown to improve significantly student achievement (compared to standard public schools, for instance) in Boston adbulkadiroglu_etal_boston_2011, angrist_etal_boston_2012, angrist_etal_boston_2016, Denver abdulkadiroglu_etal_denver_2017, or New Orleans abdulkadiroglu_etal_nola_boston_2016. The ability of charter schools to close the achievement gap between black and white students has also been demonstrated angrist_etal_boston_2013. On the contrary, students who enroll in elite exam schools do not experience an achievement premium, be it in Boston or New York abdulkadiroglu_etal_elite_2014, and can even see their achievement deteriorates slightly in Chicago abdulkadiroglu_etal_chicago_2017. The mechanisms behind these results are nicely summarized in angrist_2022: exam schools are mostly selected by students to join already high-achieving peers, but these very same students would have performed better in a charter school.

Guido Imbens

Unlike Joshua Angrist, Guido Imbens (G.I) is primarily a theoretical econometrician. His methodological contributions outside the LATE are countless and span an impressive array of different topics. He is especially recognized for his works in survey sampling theory, efficient estimation of moment-restricted models, development of a rich set of tools to estimate treatment effect parameters, and identification in very general econometric models. G. Imbens' desire to draw deep connections across seemingly distant questions is also visible in his articles: as early as the 1990s, he underlined many parallels between survey sampling, treatment effect estimation, and efficient estimation of moment-restricted models that are still largely overlooked in the econometric community today. G.I's taste for state-of-the-art methods developed by the statistics community is another striking feature of his career: he was, for instance, an early promoter of the use of empirical likelihood and machine learning techniques in econometrics. The rest of this subsection is devoted to presenting in detail G. Imbens' rich and multifaceted career beyond the LATE.

The 1990s and early 2000s: the relentless search for statistical efficiency

Outside of his collaboration with Joshua Angrist and Don Rubin, Guido Imbens' early career focused on topics largely motivated by survey sampling concerns. In the early 1990s, G.I is particularly interested in semiparametric models\footnote{ As a reminder, a model is called semiparametric when it is indexed by a finite-dimensional parameter that does not fully characterize the data distribution. Conditional maximum likelihood is a prototypical example thereof: the distribution of the outcome given covariates is fully characterized by a finite-dimensional parameter, but the distribution of covariates is left unrestricted. } estimated with choice-based sampled data. Choice-based sampling arises when the distribution of a discrete outcome of interest (in a regression model, say) in the sample at hand differs from the true distribution of that outcome in the general population. This is the rule rather than the exception for researchers who work with survey data: it is indeed widespread to oversample or downsample certain subpopulations in a survey for cost reasons or because some subpopulations are of larger interest for that specific survey. Choice-based sampling may also happen in an uncontrolled fashion when there is systematic nonresponse to a survey, for example. The question addressed by G. Imbens in the first chapter of his dissertation and in several of his early published papers imbens1992, imbens_lancaster_1996 is more precisely the following: how to estimate efficiently the parameter of a semiparametric model when data is drawn according to a choice-based sampling design?\footnote{ In imbens_lancaster_1996, the authors consider a more general stratified sampling mechanism. The derived results are very close to those presented for choice-based sampling. We thus stick to the latter framework in this discussion. }

To tackle this question, G. Imbens and Tony Lancaster propose a novel estimator which is both efficient in the sense of chamberlain_1987 and computationally more attractive than existing alternatives such as cosslett_1981. Their contribution is quite innovative: they derive their novel estimator by solving explicitly for the conditions given in chamberlain_1987 to characterize an efficient estimator. This may seem to be only a far-fetched theoretical refinement, but it is actually quite the contrary: this finding can be seen as an early instance of a generic approach to designing well-behaved estimators that has come back to the forefront in the recent treatment effect literature employing machine learning tools cherno2018; a literature to which G.I actively participates as we review below.

imbens1992 and imbens_lancaster_1996 put forward another puzzling finding: if some extra moments of the data are available to the researcher, for instance, the probability that the discrete outcome of interest be equal to some value in the entire population, then it may still be useful to re-estimate these moments from the sampled data to build more efficient estimators of the parameters of the model. Counter-intuitive results of that form have long existed in survey sampling (see hajek_1971 for example). Again, this theoretical finding has surprisingly far-reaching consequences: as discussed further down, the very influential article hirano_et_al_2003 coauthored by G. Imbens shows that it is more efficient to estimate the propensity score even when the latter is known in order to recover average treatment effects.

imbens_lancaster_1994 is another nice article pointing in the same direction as imbens1992 and imbens_lancaster_1994. In this article, the authors show that when some moments of the data distribution are exactly known, and interest lies in estimating the parameter of a conditional likelihood model, then transforming the problem into a Generalized Method of Moment one where one seeks to match both the likelihood score and the exactly known moments yields an estimator more efficient than the traditional maximum likelihood one.\footnote{ The Generalized Method of Moments (GMM) was introduced in hansen_1982 as a generic approach to estimate parameters in semiparametric models where identification stems from a set of (conditional) moment restrictions. }

From the mid-1990s to the mid-2000s, G. Imbens explores the efficiency topic beyond his initial findings in survey sampling models. He devotes a large share of his research effort to investigating alternatives to generalized methods of moments (GMMs) to estimate parameters in moment-restricted models. GMMs have indeed one major drawback: to build efficient GMM estimators, one needs to resort to a cumbersome two-step approach. Parallel to the GMM literature, a novel family of promising estimators blossoms in the statistics community: these are called generalized empirical likelihood (GEL) estimators. This literature originates in owen_1990. Unlike GMMs, GELs automatically yield efficient estimators in one step.

Following in the steps of Art Owen, G.I is one of the early proponents of these methods in econometrics. He studies the properties of GEL techniques in a series of articles imbens1997, hellerstein_imbens_1999, imbens2002, donald_imbens_newey_2003. In these works, he draws interesting links between GEL methods and older techniques not commonly viewed through this lens, such as the raking estimator in survey sampling imbens1997. Despite its theoretical elegance, the GEL framework has not developed among applied researchers due to the erratic behavior of some GEL estimators in practice guggenberger_2008. This may explain why G. Imbens does not pursue this research effort beyond the mid-2000s. Still, the GEL framework provides a sound theoretical benchmark that G.I puts to good use to derive some fundamental properties of treatment effect estimators presented hereafter.

Enlarging the econometrician's toolbox for treatment effect estimation

G. Imbens has contributed to the treatment effect literature way beyond the articles about LATEs he coauthored in the 1990s. Since the early 2000s, he has been one of the most active econometricians in the field of treatment effect estimation. He has considerably enriched the practitioner's toolbox to estimate treatment effects through his works on propensity score weighting, matching, doubly robust estimation, or, more recently, machine learning-assisted treatment effect estimation. He has also broadened our understanding of the workings of the methods previously mentioned with notable theoretical breakthroughs in matching or regression discontinuity, for instance.

Since the early 2000s, two topics have been particularly central to G. Imbens' research program on the estimation of average treatment: propensity score reweighting and matching. To explain the rationale behind propensity score reweighting and matching, let us briefly recall the canonical Neyman-Rubin potential outcome framework described in Section (ref). We consider a binary treatment variable \(D\), a set of covariates \(X\) unaffected by the treatment, and two potential outcomes \(Y(0)\) and \(Y(1)\). For every individual, we only have access to \(Y = DY(1) + (1-D)Y(0)\). Our goal is to identify and estimate the average treatment effect \(\operatorname{\mathbb{E}}[Y(1) - Y(0)]\).

Under a number of conditions, in particular independence between $D$ and $(Y(0), Y(1))$ conditional on $X$, we can show that

equation[equation omitted — 290 chars of source]

In words, Equation (ref) implies that suitably reweighting observed data allows us to recover the ATE. The map \(x \mapsto \mathbb{P}\!\left(D = 1 \,|\, X = x \right)\) is called the propensity score function, hence the concept of propensity score reweighting.

Reweighting data to construct consistent estimators is, in fact, a very old idea that has roots in survey sampling, for instance horvitz_thompson_1952. G. Imbens has contributed to the literature on propensity score weighting mainly in two directions. In imbens2000 and hirano_imbens_2004, he extends the notion of propensity scores to more complex experiments that involve a non-binary, possibly continuous, treatment. In hirano_et_al_2003, the authors prove the counter-intuitive result that estimating the propensity score leads to a more efficient estimator of the ATE even when the propensity score is known. This finding has direct implications for practitioners working with randomized experiments in which the propensity score is exactly known. To prove their result, the authors elegantly recast their estimator as an instance of the GEL family of estimators mentioned earlier. This finding is also reminiscent of the results presented in imbens1992 and imbens_lancaster_1996.

Matching is a very intuitive idea. For every treated, \(D = 1\) ({respectively} control, \(D = 0\)), individual, we only observe their potential outcome with ({resp.} without) treatment. To “recreate” the other potential outcome, a simple idea consists in identifying, for each individual $i$, a small set of individuals in the other group with similar covariate profiles. This group is called the matching group of individual $i$. The average observed outcome in the matching group serves as an imputed value of the unobserved potential outcome for $i$.

The conceptual simplicity of matching is arguably a reason for its popularity among applied economists. However, it is not intuitive at all why matching should work from a theoretical viewpoint As recalled in abadie_imbens_2006, matching with a fixed number of matches can be seen as an extreme form of nonparametric regression, which has unappealing theoretical properties even when the number of individuals in the sample is large. In abadie_imbens_2006, the authors succeed in showing that, despite the limitations of matching with a fixed number of matches, this method can still yield consistent and asymptotically normal estimates of the ATE when there is at most one continuous covariate. This result is not completely favorable to matching but is still quite unexpected. Alberto Abadie and G.I have continued exploring the intriguing theoretical properties of matching for ATE estimation in a series of articles abadie_imbens_2008, abadie_imbens_2011, abadie_imbens_2012.

A widely used extension to propensity score reweighting and matching is propensity score matching rosenbaum_rubin_1983, which combines the two original methods as its name suggests. More precisely, individuals from one group are matched to individuals from the other group based on how close their propensity scores are. In a recent article abadie_imbens_2016, A. Abadie and G. Imbens derive the asymptotic properties of this method. The obtained results are more favorable than for standard matching. This can be understood on the following grounds: in propensity score matching, matching is performed based on a unique variable, the propensity score, instead of possibly many covariates.

The list of topics to which Guido Imbens has contributed is still long. Among those, regression discontinuity designs (RDDs) and the use of machine learning methods for causal analysis (also termed “causal ML”) are currently two of G. Imbens' most active research areas. G.I's most influential input on RDDs is his review paper coauthored with Thomas Lemieux imbens_lemieux_2008, which provides applied researchers with simple guidelines to implement RDD methods rigorously. Even without covariates, implementing RDD techniques requires some care: standard RDD methods involve nonparametric regressions on each side of the treatment allocation threshold. As recalled in imbens_kalyanaraman_2011, nonparametric regressions depend on hyper-parameters that are generally difficult to choose and have a strong influence on the obtained results. In that same article, the authors propose a practical choice of the hyper-parameter associated with local linear regression (a form of nonparametric regression) that is also theoretically justified. Through this article, G.I has played a crucial role in the massive adoption of local linear regression techniques for RDD estimation. Those initial theoretical findings have been recently complemented in gelman_imbens_2019.

Since the mid-2010s, G. Imbens has been a strong advocate of the use of machine learning algorithms to address policy questions in economics. In the 2010s decade, researchers working at the frontier between economics and statistics brilliantly demonstrated that modern predictive algorithms developed in the machine learning community could be adapted to improve economists' ability to solve complicated policy questions (see athey_imbens_2017 and athey_imbens_2019 for comprehensive reviews and cherno2018 for a recent theoretical contribution). In athey_imbens_2015, G.I along with his coauthor Susan Athey proposes to adapt the well-known Classification And Regression Tree algorithm (CART) introduced in breiman1984 to the problem of estimating and conducting inference on conditional average treatment effects \(\operatorname{\mathbb{E}}[Y(1) - Y(0) \,|\, X]\). This work illustrates the conceptual challenge associated with adapting machine learning tools to causal estimation and inference: the original algorithms are designed to predict an observed outcome accurately given observed covariates. Thus, many adaptations need to be made to translate these algorithms into the potential outcome framework and ensure they deliver sensible inference claims.

Identifying increasingly complex economic models

The vast majority of the projects described in the previous subsection are concerned with estimating average treatment effects. These average effects are just the tip of the iceberg: economists are keen on looking at an array of distributional phenomena not visible at the mean.

Recovering more than average effects is challenging, though: it requires placing more structure on the data. As a simple illustration, in a toy experiment with a binary treatment \(D\), potential outcomes \(Y(0)\) and \(Y(0)\) and no covariates, the ATE is identified as soon as \(\operatorname{\mathbb{E}}[Y(1) \,|\, D = 1] = \operatorname{\mathbb{E}}[Y(1)]\) and \(\operatorname{\mathbb{E}}[Y(0) \,|\, D = 0] = \operatorname{\mathbb{E}}[Y(0)]\). On the other hand, the marginal distributions of \(Y(0)\) and \(Y(1)\) are identified only under the much more stringent condition that \(D\) be independent of \((Y(0),Y(1))\). Furthermore, identifying these marginal distributions is still not enough to recover the distribution of treatment effects, that is, the distribution of \(Y(1) -Y(0)\).\footnote{ The paragraph “Beyond average effects” in Section (ref) elaborates on this distinction. }

G. Imbens worked on these difficult questions in the 2000s. One article of that period has particularly profound implications in our view, athey_imbens_2006. In this article, S. Athey and G. Imbens search for reasonable conditions under which the marginal distribution of potential outcomes \(Y(0)\) and \(Y(1)\) can be identified in a DID setup. This is much more challenging than in a randomized experiment: treatment is not randomized, even conditional on covariates, so that both differences across groups and over time need to be exploited. The authors come up with a very elegant solution that is also very easy to implement (when there are no covariates in the analysis, at least). Their result can be interpreted as a nonlinear extension of the classical DID identification of average treatment effects.

G. Imbens also considers identification in complex nonlinear models plagued with endogeneity in cherno_et_al_2007 and imbens_newey_2009. In the last paper, a variety of policy-relevant parameters are shown to be identified under conditions carefully discussed by the authors on specific examples derived from economic theory.

modification\section{Conclusion} In the series of works that laid the foundations of the LATE revolution imbens_angrist_1994, angrist_imbens_1995, angrist_imbens_rubin_1996, J. Angrist, G. Imbens, and D. Rubin brought together several fundamental ideas that had never been truly connected before in economics: (i) the use of potential outcomes to formalize causal effects, (ii) the attention that should be devoted to credible sources of identification, and (iii) the acknowledgment that heterogeneity of treatment effects matters when giving a causal meaning to OLS or IV estimands. These have now become the dominant paradigm for discussing the estimation of causal effects in econometrics. The currently rapidly growing two-way fixed effects (TWFE) literature perfectly illustrates the lasting impact of the LATE revolution in applied economics. Linear TWFE models are a popular tool among applied economists to measure average treatment effects purged from unobserved heterogeneity in two dimensions, typically at the time and individual levels in a panel environment. They can be viewed as a generalization of the basic difference-in-differences model to designs with multiple groups and periods. Nonetheless, until recently, the precise nature of the average effects recovered by TWFE models had not been studied. Recent articles falling under the TWFE banner have typically tried to fill in this gap (see dechaisemartin2021twoway and billinski_et_al_2022 for two exhaustive reviews of this literature). The results that have emerged from this research effort have had a lot of impact on applied researchers: linear TWFE models are unable to capture meaningful average treatment effects in general as soon as treatment effects are not perfectly homogeneous over time and across individuals. Even though these results are more negative than the findings of the early LATE contributions, the ingredients underpinning the TWFE literature are fundamentally the same as those behind the LATE papers. Indeed, the authors focus on a popular estimation technique, adapt the canonical potential outcome framework (here to accommodate multiple time periods), and conclude that, without restrictions on treatment effect heterogeneity, standard tools should be used with care.
comment\begin{modification} This subsection investigates a few of the numerous and various developments following the LATE contribution. Importantly, it appears that, beyond instrumental variables where it was initially introduced, (i) the use of potential outcomes to formalize causal effects, the attention devoted (ii) to credible sources of identification and (iii) to the interpretation of causal parameters when treatment effects can be heterogeneous have nowadays become the dominant paradigm to discuss the estimation of causal effects in econometrics. This is probably the deepest lasting impact of the LATE revolution. This subsection could thus continue presenting various methods such as regression discontinuity designs (RDD, see lee_lemieux_2010 for a review), differences-in-differences (DID, see dechaisemartin2021twoway and billinski_et_al_2022 for two reviews of this literature), or synthetic control methods (see, for instance, abadie_2021 for a recent review). Instead, we decided to focus on a few points more directly related to the LATE and instrumental variables. Some of them have been the subject of controversies, as illustrated by “Better LATE Than Nothing: Some Comments on Deaton (2009) and Heckman and Urzua (2009)”, Journal of Economic Literature, imbens2010 (henceforth I10). Our intention here is not to enter those controversies. That being said, they are useful to shed light on some aspects of the LATE, and it is also the opportunity to cover extensions related to the LATE trilogy. XXXX To have a better sense of the influence of the articles written by Joshua Angrist and Guido Imbens (along with Donald Rubin) in the 1990s on research in economics, let us first note that, according to Google Scholar, the three articles considered as the basis of the LATE theory imbens_angrist_1994, angrist_imbens_1995, angrist_imbens_rubin_1996 total up to more than 14,000 citations as of February 2023. Perhaps more interestingly, these LATE articles have very close connections with a currently rapidly growing literature in applied economics, namely the two-way fixed effects (TWFE) literature. Linear TWFE models are a popular tool among applied economists to measure average treatment effects purged from unobserved heterogeneity in two dimensions, typically at the time and individual levels in a panel environment. They can be viewed as a generalization of the basic difference-in-differences model to designs with multiple groups and periods.\footnote{ The idea of difference-in-differences (DID) can be traced back to the 19th century snow1855. It has been very popular among applied economists since its use by David Card in the early 1990s. David Card was awarded the Nobel Memorial Prize in Economic Sciences in 2021, along with Joshua Angrist and Guido Imbens (we refer to the dedicated article about D. Card's contributions in the same issue of this journal). } Nonetheless, until recently, the precise nature of the average effects recovered by TWFE models had not been studied. Recent articles falling under the TWFE banner have typically tried to fill in this gap (see dechaisemartin2021twoway and billinski_et_al_2022 for two exhaustive reviews of this literature). The results that have emerged from this research effort have had a lot of impact on applied researchers: linear TWFE models are unable to capture meaningful average treatment effects in general as soon as treatment effects are not perfectly homogeneous over time and across individuals. Even though these results are more negative than the findings of the early LATE contributions, the ingredients underpinning the TWFE literature are fundamentally the same as those behind the LATE papers: the authors focus on a popular estimation technique, adapt the canonical potential outcome framework (here to accommodate multiple time periods), and conclude that without restrictions on treatment effect heterogeneity standard tools should be used with care. \end{modification} \begin{modification} XXXX Paragraphe de l'introduction sur les TWFE à insérer Cf. commentaire 1 To have a better sense of the influence of the articles written by Joshua Angrist and Guido Imbens (along with Donald Rubin) in the 1990s on research in economics, let us first note that, according to Google Scholar, the three articles considered as the basis of the LATE theory imbens_angrist_1994, angrist_imbens_1995, angrist_imbens_rubin_1996 total up to more than 14,000 citations as of February 2023. Perhaps more interestingly, these LATE articles have very close connections with a currently rapidly growing literature in applied economics, namely the two-way fixed effects (TWFE) literature. Linear TWFE models are a popular tool among applied economists to measure average treatment effects purged from unobserved heterogeneity in two dimensions, typically at the time and individual levels in a panel environment. For the past decade, these methods have been considered a reasonable compromise as they are very easy to use, even with rich panel datasets. In dechaisemartin2022differenceindifferences, the authors note that 26 out of the 100 most cited articles published in The American Economic Review between 2015 and 2019 run a TWFE specification. TWFE methods can be viewed as an intuitive generalization of the basic difference-in-differences model to designs with multiple groups and periods.\footnote{ The idea of difference-in-differences (DID) can be traced back to the 19th century snow1855. It has been very popular among applied economists since its use by David Card in the early 1990s. David Card was awarded the Nobel Memorial Prize in Economic Sciences in 2021, along with Joshua Angrist and Guido Imbens (we refer to the dedicated article about D. Card's contributions in the same issue of this journal). } Nonetheless, until recently, the precise nature of the average effects recovered by TWFE models had not been studied. Recent articles falling under the TWFE banner have typically tried to fill in this gap (see dechaisemartin2021twoway and billinski_et_al_2022 for two exhaustive reviews of this literature). The results that have emerged from this research effort have had a lot of impact on applied researchers: linear TWFE models are unable to capture meaningful average treatment effects in general as soon as treatment effects are not perfectly homogeneous over time and across individuals. Even though these results are more negative than the findings of the early LATE contributions, the ingredients underpinning the TWFE literature are fundamentally the same as those behind the LATE papers: the authors focus on a popular estimation technique, adapt the canonical potential outcome framework (here to accommodate multiple time periods), and conclude that without restrictions on treatment effect heterogeneity standard tools should be used with care. \end{modification}