EconBase
← Back to paper

The State of Applied Econometrics - Causality and Policy Evaluation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

79,438 characters · 24 sections · 124 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The State of Applied Econometrics - Causality and Policy Evaluation

abstractIn this paper we discuss recent developments in econometrics that we view as important for empirical researchers working on policy evaluation questions. We focus on three main areas, where in each case we highlight recommendations for applied work. First, we discuss new research on identification strategies in program evaluation, with particular focus on synthetic control methods, regression discontinuity, external validity, and the causal interpretation of regression methods. Second, we discuss various forms of supplementary analyses to make the identification strategies more credible. These include placebo analyses as well as sensitivity and robustness analyses. Third, we discuss recent advances in machine learning methods for causal effects. These advances include methods to adjust for differences between treated and control units in high-dimensional settings, and methods for identifying and estimating heterogenous treatment effects.
center[center omitted — 13 chars of source]

JEL Classification: C14, C21, C52

Keywords:\ Causality, Supplementary Analyses, Machine Learning, Treatment Effects, Placebo Analyses, Experiments \pagestyle{plain} \baselineskip=20pt \setcounter{page}{1}

Introduction

This article synthesizes recent developments in econometrics that may be useful for researchers interested in estimating the effect of policies on outcomes. For example, what is the effect of the minimum wage on employment? Does improving educational outcomes for some students spill over onto other students? Can we credibly estimate the effect of labor market interventions with observational studies? Who benefits from job training programs? We focus on the case where the policies of interest had been implemented for at least some units in an available dataset, and the outcome of interest is also observed in that dataset. We do not consider here questions about outcomes that cannot be directly measured in a given dataset, such as consumer welfare or worker well-being, and we do not consider questions about policies that have never been implemented. The latter type of question is considered in a branch of applied work referred to as “structural” analysis; the type of analysis considered in this review is sometimes referred to as “reduced-form,” or “design-based,” or “causal” methods.

The gold standard for drawing inferences about the effect of a policy is the randomized controlled experiment; with data from a randomized experiment, by construction those units who were exposed to the policy are the same, in expectation, as those who were not, and it becomes relatively straightforward to draw inferences about the causal effect of a policy. The difference between the sample average outcome for treated units and control units is an unbiased estimate of the average causal effect. Although digitization has lowered the costs of conducting randomized experiments in many settings, it remains the case that many policies are expensive to test experimentally. In other cases, large-scale experiments may not be politically feasible. For example, it would be challenging to randomly allocate the level of minimum wages to different states or metropolitan areas in the United States. Despite the lack of such randomized controlled experiments, policy makers still need to make decisions about the minimum wage. A large share of the empirical work in economics about policy questions relies on observational data--that is, data where policies were determined in a way other than random assignment. But drawing inferences about the causal effect of a policy from observational data is quite challenging.

To understand the challenges, consider the example of the minimum wage. It might be the case that states with higher costs of living, as well as more price-insensitive consumers, select higher levels of the minimum wage. Such states might also see employers pass on higher wage costs to consumers without losing much business. In contrast, states with lower cost of living and more price-sensitive consumers might choose a lower level of the minimum wage. A naive analysis of the effect of a higher minimum wage on employment might compare the average employment level of states with a high minimum wage to that of states with a low minimum wage. This difference is {\it not} a credible estimate of the causal effect of a higher minimum wage: it is not a good estimate of the change in employment that would occur if the low-wage state raised its minimum wage. The naive estimate would confuse correlation with causality. In contrast, if the minimum wages had been assigned randomly, the average difference between low-minimum-wage states and high-minimum-wage states would have a causal interpretation.

Most of the attention in the econometrics literature on reduced-form policy evaluation focuses on issues surrounding separating correlation from causality in observational studies, that is, with non-experimental data. There are several distinct strategies for estimating causal effects with observational data. These strategies are often referred to as “identification strategies,” or “empirical strategies” (angristkrueger_strategies) because they are strategies for “identifying” the causal effect. We say that a causal effect is “identified” if it can be learned when the data set is sufficiently large. Issues of identification are distinct from issues that arise because of limited data. In Section (ref), we review recent developments corresponding to several different identification strategies. An example of an identification strategy is one based on “regression discontinuity.” This type of strategy can be used in a setting when allocation to a treatment is based on a “forcing” variable, such as location, time, or birthdate being above or below a threshold. For example, a birthdate cutoff may be used for school entrance or for the decision of whether a child can legally drop out of school in a given academic year; and there may be geographic boundaries for assigning students to schools or patients to hospitals. The identifying assumption is that there is no discrete change in the characteristics of individuals who fall on one side or the other of the threshold for treatment assignment. Under that assumption, the relationship between outcomes and the forcing variable can be modeled, and deviations from the predicted relationship at the treatment assignment boundary can be attributed to the treatment. Section (ref) also considers other strategies such as synthetic control methods, methods designed for networks settings, and methods that combine experimental and observational data.

In Section (ref) we discuss what we refer to in general as {\it supplementary analyses}. By supplementary analyses we mean analyses where the focus is on providing support for the identification strategy underlying the primary analyses, on establishing that the modeling decisions are adequate to capture the critical features of the identification strategy, or on establishing robustness of estimates to modeling assumptions. Thus the results of the supplementary analyses are intended to convince the reader of the credibility of the primary analyses. Although these analyses often involve statistical tests, the focus is not on goodness of fit measures. Supplementary analyses can take on a variety of forms, and we discuss some of the most interesting ones that have been proposed thus far. In our view these supplementary analyses will be of growing importance for empirical researchers. In this review, our goal is to organize these analyses, which may appear to be applied unsystematically in the empirical literature, or may have not received a lot of formal attention in the econometrics literature.

In Section (ref) we discuss briefly new developments coming from what is referred to as the machine learning literature. Recently there has been much interesting work combining these predictive methods with causal analyses, and this is the part of the literature that we put special emphasis on in our discussion. We show how machine learning methods can be used to deal with datasets with many covariates, and how they can be used to enable the researcher to build more flexible models. Because many common identification strategies rest on assumptions such as the ability of the researcher to observe and control for confounding variables (e.g. the factors that affect treatment assignment as well as outcomes), or to flexibly model the factors that affect outcomes in the absence of the treatment, machine learning methods hold great promise in terms of improving the credibility of policy evaluation, and they can also be used to approach supplementary analyses more systematically.

As the title indicates, this review is limited to methods relevant for policy analysis, that is, methods for causal effects. Because there is another review in this issue focusing on structural methods, as well as one on theoretical econometrics, we largely refrain from discussing those areas, focusing more narrowly on what is sometimes referred to as reduced-form methods, although we prefer the terms causal or design-based methods, with an emphasis on recommendations for applied work. The choices for topics within this area is based on our reading of recent research, including ongoing work, and we point out areas where we feel there are interesting open research questions. This is of course a subjective perspective.

New Developments in Program Evaluation

The econometric literature on estimating causal effects has been a very active one for over three decades now. Since the early 1990s the potential outcome, or Neyman-Rubin Causal Model, approach to these problems has gained substantial acceptance as a framework for analyzing causal problems. (We should note, however, that there is a complementary approach based on graphical models (e.g., pearl) that is widely used in other disciplines, though less so in economics.) In the potential outcome approach, there is for each unit $i$, and each level of the treatment $w$, a potential outcome $Y_i(w)$, that describes the level of the outcome under treatment level $w$ for that unit. In this perspective, causal effects are comparisons of pairs of potential outcomes for the same unit, e.g., the difference $Y_i(w')-Y_i(w)$. Because a given unit can only receive one level of the treatment, say $W_i$, and only the corresponding level of the outcome, $Y_i^{\rm obs}=Y_i(W_i)$ can be observed, we can never directly observe the causal effects, which is what holland calls the “fundamental problem of causal inference.” Estimates of causal effects are ultimately based on comparisons of different units with different levels of the treatment.

A large part of the causal or treatment effect literature has focused on estimating average treatment effects in a binary treatment setting under the unconfoundedness assumption (e.g., rosenbaum1983central), \[ W_i\ \perp\!\!\!\perp\ \Bigl(Y_i(0),Y_i(1)\Bigr)\ \Big|\ X_i.\] Under this assumption, associational or correlational relations such as $\mathbb{E}[Y^{\rm obs}_i|W_i=1,X_i=x]-\mathbb{E}[Y^{\rm obs}_i|W_i=0,X_i=x]$ can be given a causal interpretation as the average treatment effect $\mathbb{E}[Y_i(1)-Y_i(0)|X_i=x]$. The literature on estimating average treatment effects under unconfoundedness is by now a very mature literature, with a number of competing estimators and many applications. Some estimators use matching methods, some rely on weighting, and some involve the propensity score, the conditional probability of receiving the treatment given the covariates, $e(x)={\rm pr}(W_i=1|X_i=x)$. There are a number of recent reviews of the general literature (imbens2004, imbens2015causal, and for a different perspective heckman, heckman2007), and we do not review it in its entirety in this review. However, one area with continuing developments concerns settings with many covariates, possibly more than there are units. For this setting connections have been made with the machine learning and big data literatures. We review these new developments in Section (ref). In the context of many covariates there has also been interesting developments in estimating heterogenous treatment effects; we cover this literature in Section (ref). We also discuss, in Section (ref), settings with unconfoundedness and multiple levels for the treatment.

Beyond settings with unconfoundedness we discuss issues related to a number of other identification strategies and settings. In Section (ref), we discuss regression discontinuity designs. Next, we discuss synthetic control methods as developed in the abadie2010, which we believe is one the most important development in program evaluation in the last decade. In Section (ref) we discuss causal methods in network settings. In Section (ref) we draw attention to some recent work on the causal interpretation of regression methods. We also discuss external validity in Section (ref), and finally, in Section (ref) we discuss how randomized experiments can provide leverage for observational studies.

In this review we do not discuss the recent literature on instrumental variables. There are two major strands of that by now fairly mature literature. One focuses on heterogenous treatment effects, with a key development the notion of the local average treatment effect (imbens1994, angrist1996). This literature has recently been reviewed in imbens2014. There is also a separate literature on weak instruments, focusing on settings with a possibly large number of instruments and weak correlation between the instruments and the endogenous regressor. See bekker1994, stock1997, chamberlain for specific contributions, and andrews2007 for a survey. We also do not discuss in detail bounds and partial identification analyses. Since the work by Manski (e.g., manski_bounds) these have received a lot of interest, with an excellent recent review in tamer.

Regression Discontinuity Designs

A regression discontinuity design is a research design that exploits discontinuities in incentives to participate in a treatment to evaluate the effect of these treatment.

Set Up

In regression discontinuity designs, we are interested in the causal effect of a binary treatment or program, denoted by $W_i$. The key feature of the design is the presence of an exogenous variable, the forcing variable, denoted by $X_i$, such that at a particular value of this forcing variable, the threshold $c$, the probability of participating in the program or being exposed to the treatment changes discontinuously: \[ \lim_{x\uparrow c} {\rm pr}(W_i=1|X_i=x)\neq \lim_{x\downarrow c} {\rm pr}(W_i=1|X_i=x).\] If the jump in the conditional probability is from zero to one, we have a {\it sharp} regression discontinuity (SRD) design; if the magnitude of the jump is less than one, we have a {\it fuzzy} regression discontinuity (FRD) design. The estimand is the discontinuity in the conditional expectation of the outcome at the threshold, scaled by the discontinuity in the probability of receiving the treatment: \[ \tau^{\rm rd}=\frac{\lim_{x\downarrow c}\mathbb{E}[Y_i|X_i=x]- \lim_{x\uparrow c}\mathbb{E}[Y_i|X_i=x]}{\lim_{x\downarrow c}\mathbb{E}[W_i|X_i=x]-\lim_{x\uparrow c}\mathbb{E}[W_i|X_i=x]}.\] In the SRD case the denominator is equal to one, and we just focus on the discontinuity of the conditional expectation of the outcome given the forcing variable at the threshold. In that case, under the assumption that the individuals just to the right and just to the left of the threshold are comparable, the estimand has an interpretation as the average effect of the treatment for individuals close to the threshold. In the FRD case, the interpretation of the estimand is the average effect for compliers at the threshold (i.e., individuals at the threshold whose treatment status would have changed had they been on the other side of the threshold) hahntodd.

Estimation and Inference

In the general FRD case, the estimand $\tau^{\rm rd}$ has four components, each of them the limit of the conditional expectation of a variable at a particular value of the forcing variable. We can think of this, after splitting the sample by whether the value of the forcing variable exceeds the threshold or not, as estimating the conditional expectation at a boundary point. Researchers typically wish to use flexible (e.g., semiparametric or nonparametric) methods for estimating these conditional expectations. Because the target in each case is the conditional expectation at a boundary point, simply differencing average outcomes close to the threshold on the right and on the left leads to an estimator with poor properties, as stressed by porter. As an alternative porter suggested “local linear regression,” which involves estimating linear regressions of outcomes on the forcing variable separately on the left and the right of the threshhold, weighting most heavily observations close to the threshold, and then taking the difference between the predicted values at the threshold. This local linear estimator has substantially better finite sample properties than nonparametric methods that do not account for threshold effects, and it has become the standard. There are some suggestions that using local quadratic methods may work well given the current technology for choosing bandwidths (e.g., calonico). Some applications use global high order polynomial approximations to the regression function, but there has been some criticism of this practice. gelmanimbens argue that in practice it is difficult to choose the order of the polynomials in a satisfactory way, and that confidence intervals based on such methods have poor properties.

Given a local linear estimation method, a key issue is the choice of the bandwidth, that is, how close observations need to be to the threshold. Conventional methods for choosing optimal bandwidths in nonparametric estimation, e.g., based on cross-validation, look for bandwidths that are optimal for estimating the entire regression function, whereas here the interest is solely in the value of the regression function at a particular point. The current state of the literature suggests choosing the bandwidth for the local linear regression using asymptotic expansions of the estimators around small values for the bandwidth. See imbenskalyanaraman and cattaneo2010 for further discussion.

In some cases, the discontinuity involves multiple exogenous variables. For example, in jacob and matsudaira, the focus is on the causal effect of attending summer school. The formal rule is that students who score below a threshold on either a language or a mathematics test are required to attend summer school. Although not all the students who are required to attend summer school do so (so that this a fuzzy regression discontinuity design), the fact that the forcing variable is a known function of two observed exogenous variables makes it possible to estimate the effect of summer school at different margins. For example, one can estimate of the effect of summer school for individuals who are required to attend summer school because of failure to pass the language test, and compare this with the estimate for those who are required because of failure to pass the mathematics test. Even more than the presence of other exogenous variables, the dependence of the threshold on multiple exogenous variables improves the ability to detect and analyze heterogeneity in the causal effects.

An Illustration

Let us illustrate the regression discontinuity design with data from jacob. jacob use administrative data from the Chicago Public Schools which instituted in 1996 an accountability policy that tied summer school attendance and promotional decisions to performance on standardized tests. We use the data for 70,831 third graders in years 1997-99. The rule was that individuals score below the threshold (2.75 in this case) on either a reading or mathematics score before the summer were required to attend summer school. It should be noted that the initial scores range from 0 to 6.8, with increments equal to 0.1. The outcome variable $Y_i^{\rm obs}$ is the math score after the summer school, normalized to have variance one. Out of the 70,831 third graders, 15,846 score below the threshold on the mathematics test, 26,833 scored below the threshold on the reading test, 12,779 score below the threshold on both tests, and 29,900 scored below the threshold on at least one test.

Table (ref) presents some of the results. The first row presents an estimate of the effect on the mathematics test, using for the forcing variable the minimum of the initial mathematics score and the initial reading score. We find that the program has a substantial effect. Figure 1 shows which students contribute to this estimate. The figure shows a scatterplot of 1.5% of the students, with uniform noise added to their actual scores to show the distribution more clearly. The solid line shows the set of values for the mathematics and reading scores that would require the students to participate in the summer program. The area enclosed by the dashed line contains all the students within the bandwidth from the threshold.

We can partition the sample into students with relatively high reading scores (above the threshold plus the Imbens-Kalyanaraman bandwidth), who could only be in the summer program because of their mathematics score, students with relatively high mathematics scores (above the threshold plus the bandwidth) who could only be in the summer program because of their reading score, and students with low mathematics and reading scores (below the threshold plus the bandwidth). Rows 2-4 present estimates for these separate subsamples. We find that there is relatively little evidence of heterogeneity in the estimates of the program.

The last row demonstrates the importance of using local linear rather than standard kernel (local constant) regressions. Using the same bandwidth, but using a weighted average of the outcomes rather than a weighted linear regression, leads to an estimate equal to -0.15: rather than benefiting from the summer school, this estimate counterintuitively suggests that the summer program hurts the students in terms of subsequent performance. This bias that leads to these negative estimates is not surprising: the students who participate in the program are on average worse in terms of prior performance than the students who do not participate in the program, even if we only use information for students close to the threshold.

table[table omitted — 562 chars of source]

Regression Kink Designs

One of the most interesting recent developments in the area of regression discontinuity designs is the generalization to discontinuities in derivatives, rather than levels, of conditional expectations. The first discussions of these regression kink designs are in nielsen,cardkink,dong2. The basic idea is that at a threshold for the forcing variable, the slope of the outcome function (as a function of the forcing variable) changes, and the goal is to estimate this change in slope.

To make this clearer, let us discuss the example in cardkink. The forcing variable is a lagged earnings variable that determines unemployment benefits. A simple rule would be that unemployment benefits are a fixed percentage of last year's earnings, up to a maximum. Thus the unemployment benefit, as a function of the forcing variable, is a continuous, piecewise linear function. Now suppose we are interested in the causal effect of an increase in the unemployment benefits on the duration of unenmployment spells. Because the benefits are a deterministic function of lagged earnings, direct comparisons of individuals with different levels of benefits are confounded by differences in lagged earnings. However, at the threshold, the relation between benefits and lagged earnings changes. Specifically, the derivative of the benefits with respect to lagged earnings changes. If we are willing to assume that in the absence of the kink in the benefit system, the derivative of the expected duration would be smooth in lagged earnings, then the change in the derivative of the expected duration with respect to lagged earnings is informative about the relation between the expected duration and the benefit schedule, similar to the identification in a regular regression discontinuity design.

To be more precise, suppose the benefits as a function of lagged earnings satisfy \[ B_i=b(X_i),\] with $b(x)$ known and continuous, with a discontinuity in the first derivative at $x=c$. Let $b'(v)$ denote the derivative, letting $b'(c+)$ and $b'(c-)$ denote the derivatives from the right and the left at $x=c$. If the benefit schedule is piecewise linear, we would have \[ B_i= \beta_{0}+\beta_{1-}\cdot (X_i-c),\ \ X_i<c,\] \[ B_i=\beta_{0}+\beta_{1+}\cdot (X_i-c),\ \ X_i\geq c.\] This relationship is deterministic, making this a sharp regression kink design. Here, as before, $c$ is the threshold. The forcing variable $X_i$ is lagged earnings, $B_i$ is the unemployment benefit that an individual would receive. As a function of the benefits $b$, the logarithm of the unemployment duration, denoted by $Y_i$, is assumed to satisfy \[Y_i(b)=\alpha+\tau\cdot \ln (b)+\varepsilon_i.\] Let $g(x)=\mathbb{E}[Y_i|X_i=x]$ be the conditional expectation of $Y_i$ given $X_i=x$, with derivative $g'(x)$. The derivative is assumed to exist everywhere other than at $x=c$, where the limits from the right and the left exist. The idea is to characterize $\tau$ as \[ \tau=\frac{\lim_{x\downarrow c} g'(x)-\lim_{x\uparrow c} g'(x)}{\lim_{x\downarrow c} b'(x)-\lim_{x\uparrow c} b'(x)}.\] cardkink propose estimating $\tau$ by first estimating $g(x)$ by local linear or local quadratic regression around the threshold. We then divide the difference in the estimated derivative from the right and the left by the difference in the derivatives of $b(x)$ from the right and the left at the threshold.

In some cases, the relationship between $B_i$ and $X_i$ is not deterministic, making it a fuzzy regression kink design. In the fuzzy version of the regression kink design, the conditional expectation of $B_i$ given $X_i$ is estimated using the same approach to get an estimate of the change in the derivative at the threshold.

Summary of Recommendations

There are some specific choices to be made in regression discontinuity analyses, and here we provide our recommendations for these choices. We recommend using local linear or local quadratic methods (see for details on the implementation hahntodd, porter, calonico) rather than global polynomial methods. gelmanimbens present a detailed discussion on the concerns with global polynomial methods. These local linear methods require a bandwidth choice. We recommend the optimal bandwidth algorithms based on asymptotic arguments involving local expansions discssed in imbenskalyanaraman, calonico. We also recommend carrying out supplementary analyses to assess the credibility of the design, and in particular to test for evidence of manipulation of the forcing variable. Most important here is the McCrary test for discontinuities in the density of the forcing variable (mccrary), as well as tests for discontinuities in average covariate values at the threshold. We discuss examples of these in the section on supplementary analyses (Section (ref)). We also recommend researchers to investigate external validity of the regression discontinuity estimates by assessing the credibility of extrapolations to other subpopulations (bertanhaimbens, angristrokkanen, angristfernandez, donglewbel). See Section (ref) for more details.

The Literature

Regression Discontinuity Designs have a long history, going back to work in psychology in the fifties by campbell, but the methods did not become part of the mainstream economics literature until the early 2000s (with goldberger1, goldberger2 an exception). Early applications in economics include black angrist1999, vanderklaauw, lee. Recent reviews include imbenslemieux, leelemieux, vanderklaauw2, titiunik. More recently there have been many applications (e.g., greenstone) and a substantial amount of new theoretical work which has led to substantial improvements in our understanding of these methods.

Synthetic Control Methods and Difference-In-Differences

Difference-In-Differences (DID) methods have become an important tool for empirical researchers. In the basic setting there are two or more groups, at least one treated and one control, and we observe (possibly different) units from all groups in two or more time periods, some prior to the treatment and some after the treatment. The difference between the treatment and control groups post treatment is adjusted for the difference between the two groups prior to the treatment. In the simple DID case these adjustments are linear: they take the form of estimating the average treatment effect as the difference in average outcomes post treatment minus the difference in average outcomes pre treatment. Here we discuss two important recent developments, the synthetic control approach and the nonlinear changes-in-changes method.

Synthetic Control Methods

Arguably the most important innovation in the evalulation literature in the last fifteen years is the synthetic control approach developed by abadie2010,abadie2014 and abadie2003. This method builds on difference-in-differences estimation, but uses arguably more attractive comparisons to get causal effects. We discuss the basic abadie2010 approach, and highlight alternative choices and restrictions that may be imposed to further improve the performance of the methods relative to difference-in-differences estimation methods.

We observe outcomes for a number of units, indexed by $i=0,\ldots,N$, for a number of periods indexed by $t=1,\ldots,T$. There is a single unit, say unit $0$, who was exposed to the control treatment during periods $1,\ldots,T_0$ and who received the active treatment, starting in period $T_0+1$. For ease of exposition let us focus on the case with $T=T_0+1$ so there is only a single post-treatment period. All other units are exposed to the control treatment for all periods. The number of control units $N$ can be as small as 1, and the number of periods $T$ can be as small as 2. We may also observe exogenous fixed covariates for each of the units. The units are often aggregates of individuals, say states, or cities, or countries. We are interested in the causal effect of the treatment for this unit, $Y_{0T}(1)-Y_{0T}(0)$.

The traditional DID approach would compare the change for the treated unit (unit 0) between periods $t$ and $T$, for some $t<T$, to the corresponding change for some other unit. For example, consider the classic difference-in-differences study by cardmariel. Card is interested in the effect of the Mariel boatlift, which brought Cubans to Miami, on the Miami labor market, and specifically on the wages of low-skilled workers. He compares the change in the outcome of interest, for Miami, to the corresponding change in a control city. He considers various possible control cities, including Houston, Petersburg, Atlanta.

The synthetic control idea is to move away from using a single control unit or a simple average of control units, and instead use a weighted average of the set of controls, with the weights chosen so that the weighted average is similar to the treated unit in terms of lagged outcomes and covariates. In other words, instead of choosing between Houston, Petersburg or Atlanta, or taking a simple average of outcomes in those cities, the synthetic control approach chooses weights $\lambda_{\rm h}$, $\lambda_{\rm p}$, and $\lambda_{\rm a}$ for Houston, Petersburg and Atlanta respectively, so that $\lambda_{\rm h}\cdot Y_{{\rm h}t}+ \lambda_{\rm p}\cdot Y_{{\rm p}t}+ \lambda_{\rm a}\cdot Y_{{\rm a}t}$ is close to $Y_{{\rm m}t}$ (for Miami) for the pre-treatment periods $t=1,\ldots,T_0$, as well as for the other pretreatment variables (e.g., peri2015). This is a very simple, but very useful idea. Of course, if pre-boatlift wages are higher in Houston than in Miami, and higher in Miami than in Atlanta, it would make sense to compare Miami to the average of Houston and Atlanta rather than to Houston or Atlanta. The simplicity of the idea, and the obvious improvement over the standard methods, have made this a widely used method in the short period of time since its inception.

The implementation of the synthetic control method requires a particular choice for estimating the weights. The original paper abadie2010 restricts the weights to be non-negative and requires them to add up to one. Let $K$ be the dimension of the covariates $X_i$, and let $\Omega$ be an arbitrary positive definite $K\times K$ matrix. Then let $\lambda(\Omega)$ be the weights that solve \[ \lambda(\Omega)=\arg\min_\lambda \left(X_0-\sum_{i=1}^N \lambda_i\cdot X_{i}\right)' \Omega \left(X_0-\sum_{i=1}^N \lambda_i\cdot X_{i}\right).\] abadie2010 choose the weight matrix $\Omega$ that minimizes \[ \sum_{t=1}^{T_0} \left(Y_{0t}-\sum_{i=1}^N \lambda_i(\Omega)\cdot Y_{it}\right)^2.\] If the covariates $X_i$ consist of the vector of lagged outcomes, this estimate amounts to minimizing \[ \sum_{t=1}^{T_0} \left(Y_{0t}-\sum_{i=1}^N \lambda_i\cdot Y_{it}\right)^2,\] subject to the restrictions that the $\lambda_i$ are non-negative and summ up to one.

doudchenko point out that one can view the question of estimating the weights in the Abadie-Diamond-Hainmueller synthetic control method differently. Starting with the case without covariates and only lagged outcomes, one can consider the regression function \[ Y_{0t}=\sum_{i=1}^N \lambda_i\cdot Y_{it}+\varepsilon_t,\] with $T_0$ units and $N$ regressors. The absence of the covariates is rarely important, as the fit typically is driven by matching up the lagged outcomes rather than matching the covariates. Estimating this regression by least squares is typically not possible because the number of regressors $N$ (the number of control units) is often larger than, or the same order of magnitude as, the number of observations (the number of time periods $T_0$). We therefore need to regularize the estimates in some fashion or another. There are a couple of natural ways to do this. abadie2010 impose the restriction that the weights $\lambda_i$ are non-negative and add up to one. That often leads to a unique set of weights. However, there are alternative ways to regularize the estimates. In fact, both the restrictions that abadie2010 impose may hurt performance of the model. If the unit is on the extreme end of the distribution of units, allowing for weights that sum up to a number different from one, or allowing for negative weights may improve the fit. We can do so by using alternative regularization methods such as best subset regression, or LASSO (see Section (ref) for a description of LASSO) where we add a penalty proportional to the sum of the weights. doudchenko explore such approaches.

Nonlinear Difference-in-Difference Models

A commonly noted concern with difference-in-difference methods is that functional form assumptions play an important role. For example, in the extreme case with only two groups and two periods, it is not clear whether the change over time should be modeled as the same for the two groups in terms of levels of outcomes, or in terms of percentage changes in outcomes. If the initial period mean outcome is different across the two groups, the two different assumptions can give different answers in terms of both sign and magnitude. In general, a treatment might affect both the mean and the variance of outcomes, and the impact of the treatment might vary across individuals.

For the case where the data includes repeated cross-sections of individuals (that is, the data include individual observations about many units within each group in two different time periods, but the individuals can not be linked across time periods or may come from a distinct sample), atheyimbens_cic propose a non-linear difference-in-difference model which they refer to as the changes-in-changes model that does not rely on functional form assumptions.

Modifying the notation from the last subsection, we now imagine that there are two groups, $g \in \{A,B\}$, where group A is the control group and group B is the treatment group. There are many individuals in each group with potential outcomes denoted $Y_{gti}(w)$. We observe $Y_{gti}(0)$ for a sample of units in both groups when $t=1$, and for group A when $t=2$; we observe $Y_{gti}(1)$ for group B when $t=2.$ Denote the distribution of the observed outcomes in group $g$ at time $t$ by $F_{gt}(\cdot)$. We are interested in the distribution of treatment effects for the treatment group in the second period, $Y_{B2i}(1)-Y_{B2i}(0)$. Note that the distribution of $Y_{B2i}(1)$ is directly estimable, while the counterfactual distribution of $Y_{B2i}(0)$ is not, so the problem boils down to learning the distribution of $Y_{B2i}(0)$, based on the distributions of $Y_{B1i}(0)$, $Y_{A2i}(0)$, and $Y_{A1i}(0)$. Several assumptions are required to accomplish this. First is that the potential outcome in the absence of the treatment can be written as a monotone function of an unobservable $U_i$ and time: $Y_{gti}(0)=h(U_i,t)$. Note that the function does not depend directly on $g$, so that differences across groups are attributed to differences in the distribution of $U_i$ across groups. Second, the function $h$ is strictly increasing. This is not a restrictive assumption for a single time period, but it is restrictive when we require it to hold over time, in conjunction with a third assumption, namely that the distribution of $U_i$ is stable over time within each group. The final assumption is that the support of $U_i$ for the treatment group is contained in the support of $U_i$ for the control group. Under these assumptions, the distribution of $Y_{B2i}(0)$ is identified, with the formula for the distribution given as follows: \[ Pr(Y_{B2i}(0) \le y) = F_{B1}(F^{(-1)}_{A1}(F_{A2}(y))). \] atheyimbens_cic show that an estimator based on the empirical distributions of the observed outcomes is efficient and discuss extensions to discrete outcome settings.

The nonlinear difference-in-difference model can be used for two distinct purposes. First, the distribution is of direct interest for policy, beyond the average treatment effect. Further, a number of authors have used this approach as a robustness check, i.e., a supplementary analysis in the terminology of Section (ref), for the results from a linear model.

Estimating Average Treatment Effects under Unconfoundedness in Settings with Multivalued Treatments

Much of the earlier econometric literature on treatment effects focused on the case with binary treatments. For a textbook discussion, see imbens2015causal. Here we discuss the results of the more recent multi-valued treatment effect literature. In the binary treatment case, many methods have been proposed for estimating the average treatment effect. Here we focus on two of these methods, subclassification with regression and and matching with regression, that have been found to be effective in the binary treatment case (imbens2015causal). We discuss how these can be extended to the multi-valued treatment setting without increasing the complexity of the estimators. In particular, the dimension reducing properties of a generalized version of the propensity score can be maintained in the multi-valued treatment setting.

Set Up

To set the stage, it is useful to start with the binary treatment case. The standard set up postulates the existence of two potential outcomes, $Y_i(0)$ and $Y_i(1)$. With the binary treatment denoted by $W_i\in\{0,1\}$, the realized and observed outcome is \[Y_i^{\rm obs}=Y_i(W_i)=\left\{

array[array omitted — 87 chars of source]

\right. \] In addition to the treatment indicator and the outcome we may observe a set of pretreatment variables denoted by $X_i$. Following rosenbaum1983central a large literature focused on estimation of the population average treatment effect $\tau=\mathbb{E}[Y_i(1)-Y_i(0)],$ under the unconfoundedness assumption that \[ W_i\ \perp\!\!\!\perp\ \Bigl(Y_i(0),Y_i(1)\Bigr)\ \Bigl|\ X_i.\] In combination with overlap, requiring that the propensity score $e(x)={\rm pr}(W_i=1|X_i=x),$ is strictly between zero and one, the researcher can estimate the population average treatment effect by adjusting the differences in outcomes by treatment status for differences in the pretreatment variables: \[ \tau=\mathbb{E}\Bigl[ \mathbb{E}[Y_i^{\rm obs}|X_i,W_i=1]-\mathbb{E}[Y_i^{\rm obs}|X_i,W_i=0] \Bigr].\] In that case many estimation strategies have been developed, relying on regression hahn1998role, matching abadie2006, inverse propensity weighting hirr, subclassification rosenbaum1983central, as well as doubly robust methods robins1, robins2. rosenbaum1983central established a key result that underlies a number of these estimation strategies: unconfoundedness implies that conditional on the propensity score, the assignment is independent of the potential outcomes: \[ W_i\ \perp\!\!\!\perp\ \Bigl(Y_i(0),Y_i(1)\Bigr)\ \Bigl|\ e(X_i).\] In practice the most effective estimation methods appear to be those that combine some covariance adjustment through regression with a covariate balancing method such as subclassification, matching, or weighting based on the propensity score (imbens2015causal).

Substantially less attention has been paid to the case where the treatment takes on multiple values. Exceptions include imbens2000, lechner2001, imai2004, cattaneo2010, hirano2004 and yang2016. Let $\mathbb{W}=\{0,1,\ldots,T\}$ be the set of values for the treatment. In the multivalued treatment case, one needs to be careful in defining estimands, and the role of the propensity score is subtly different. One natural set of estimands is the average treatment effect if all units were switched from treatment level $w_1$ to treatment level $w_2$:

equation[equation omitted — 72 chars of source]

To estimate estimands corresponding to uniform policies such as ((ref)), it is not sufficient to take all the units with treatment levels $w_1$ or $w_2$ and use methods for estimating treatment effects in a binary setting. The latter strategy would lead to an estimate of $\tau'_{w_1,w_2}=\mathbb{E}[Y_i(w_2)-Y_i(w_1)|W_i\in\{w_1,w_2\}]$, which differs in general from $\tau_{w_1,w_2}$ because of the conditioning. Focusing on unconditional average treatment effects like $\tau_{w_1,w_2}$ maintains transitivity: $\tau_{w_1,w_2}+\tau_{w_2,w_3}=\tau_{w_1,w_3}$, which would not necessarily be the case for $\tau'_{w_1,w_2}$. There are other possible estimands, but we do not discuss alternatives here.

A key first step is to note that this estimand can be written as the difference in two marginal expectations: $\tau_{w_1,w_2}=\mathbb{E}[Y_{i}(w_2)]-\mathbb{E}[Y_{i}(w_1)]$, and that therefore identification of marginal expectations such as $\mathbb{E}[Y_{i}(w)]$ is sufficient for identification of average treatment effects.

Now suppose that a generalized version of unconfoundedness holds: \[ W_i\ \perp\!\!\!\perp\ \Bigl(Y_i(0),Y_i(1),\ldots, Y_i(T)\Bigr)\ \Bigl|\ X_i.\] There is no scalar function of the covariates that maintains this conditional independence relation. In fact, with $T$ treatment levels one would need to condition on $T-1$ functions of the covariates to make this conditional independence hold. However, unconfoundedness is in fact not required to enjoy the benefits of the dimension-reducing property of the propensity score. imbens2000 introduces a concept, called weak unconfoundedness, which requires only that the indicator for receiving a particular level of the treatment and the potential outcome for that treatment level are conditionally independent: \[ \boldsymbol{1}_{W_i=w}\ \perp\!\!\!\perp\ Y_i(w)\ \Bigl|\ X_i, \ \ {\rm for\ all}\ w\in\{0,1,\ldots,T\}.\] imbens2000 shows that weak uncnfoundedness implies similar dimension reduction properties as are available in the binary treatment case. He further introduced the concept of the generalized propensity score: \[ r(w,x)={\rm pr}(W_i=w|X_i=x).\] Weak unconfoundedness implies that, for all $w$, it is sufficient for the removal of systematic biases to condition on the generalized propensity score for that particular treatment level: \[ \boldsymbol{1}_{W_i=w}\ \perp\!\!\!\perp\ Y_i(w)\ \Bigl|\ r(w,X_i).\] This in turn can be used to develop matching or propensity score subclassification strategies as outlined in yang2016. This approach relies on the equality $ \mathbb{E}[Y_i(w)] =\mathbb{E}\Bigl[ \mathbb{E}[Y_i^{\rm obs}|X_i,W_i=w] \Bigr]$. As shown in yang2016, it follows from weak unconfoundedness that \[ \mathbb{E}[Y_i(w)] =\mathbb{E}\Bigl[ \mathbb{E}[Y_i^{\rm obs}|r(w,X_i),W_i=w] \Bigr].\] To estimate $\mathbb{E}[Y_i(w)]$, divide the sample into $J$ sublasses based on the value of $r(w,X_i)$, with $B_i\in\{1,\ldots,J\}$ denoting the subclass. We estimate $\mu_j(w)=\mathbb{E}[Y_i(w)|B_i=j]$ as the average of the outcomes for units with $W_i=w$ and $B_i=j$. Given those estimates, we estimate $\mu(w)=\mathbb{E}[Y_i(w)]$ as a weighted average of the $\hat\mu_j(w)$, with weights equal to the fraction of units in subclass $j$. The idea is not to find subsets of the covariate space where we can interpret the difference in average outcomes by all treatment levels as estimates of causal effects. Instead we find subsets where we can estimate the marginal average outcome for a particular treatment level as the conditional average for units with that treatment level, one treatment level at a time. This opens up the way for using matching and other propensity score methods developed for the case with binary treatments in settings with multivalued treatments, irrespective of the number of treatment levels.

A separate literature has gone beyond the multi-valued treatment setting to look at dynamic treatment regimes. With few exceptions most of these studies appear in the biostatistical literature: see hernanrobins for a general discussion.

Causal Effects in Networks and Social Interactions

An important area that has seen much novel work in recent years is that on peer effects and causal effects in networks. Compared to the literature on estimating average causal effects unconfoundedness without interference, the literature has not focused on a single setting; rather, there are many problems and settings with interesting questions. Here, we will discuss some of the settings and some of the progress that has been made. However, this review will be brief, and incomplete, because this continues to be a very active area, with work ranging from econometrics (manski1993) to economic theory (jackson).

In general, the questions in this literature focus on causal effects in settings where units, often individuals, interact in a way that makes the no-interference or sutva (rosenbaum1983central, imbens2015causal) assumptions that are routinely made in the treatment effect literature implausible. Settings of interest include those where the possible interference is simply a nuisance, and the interest continuous to be in causal effects of treatments assigned to a particular unit on the outcomes for that unit. There are also settings where the interest is in the magnitude of the interactions, or peer effects, that is, in the effects of changing treatments for one unit on the outcomes of other units. There are settings where the network (that is, the set of links connecting the individuals) is fixed exogenously, and some where the network itself is the result of a possibly complex set of choices by individuals, possibly dynamic and possibly affected by treatments. There are settings where the population can be partitioned into subpopulations with all units within a subpopulation connected, as, for example, in classroom settings (e.g., manski1993,carrell), workers in a labor market (crepon) or roommates in college (sacerdote), or with general networks, where friends of friends are not necessarily friends themselves (christakis). Sometimes it is more reasonable to think of many disconnected networks, where distributional approximations rely on the number of networks getting large, versus a single connected network such as Facebook. It maybe reasonable in some cases to think of the links as undirected (symmetric), and in others as directed. These links can be binary, with links either present or not, or contain links of different strengths. This large set of scenarios has led to the literature becoming somewhat fractured and unwieldy. We will only touch on a subset of these problems in this review.

Models for Peer Effects

Before considering estimation strategies, it is useful to begin by considering models of the outcomes in a setting with peer effects. Such models have been proposed in the literature. A seminal paper in the econometric literature is Manski's linear-in-means model (manski1993, bramoulle, goldsmith). Manski's original paper focuses on the setting where the population is partioned into groups (e.g., classrooms), and peer effects are constant within the groups. The basic model specification is \[ Y_{i}=\beta_0+ \beta_{\overline{Y}}\cdot {\overline{Y}}_i+\beta_X'X_i+\beta_{\overline{X}}'{\overline{X}}_i+\beta_Z'Z_i+\varepsilon_i,\] where $i$ indexes the individual. Here $Y_{i}$ is the outcome for individual $i$, ${\overline{Y}}_i$ is the average outcome for individuals in the peer group for individual $i$, $X_i$ is a set of exogenous characteristics of individual $i$, ${\overline{X}}_i$ is the average value of the characteristics in individual $i$'s peer group, and $Z_i$ are group characteristics that are constant for all individuals in the same peer group. Manski considers three types of peer effects. Outcomes for individuals in the same group may be correlated because of a shared environment. These effects are called correlated peer effects, and captured by the coefficient on $Z_i$. Next are the exogenous peer effects, captured by the coefficient on the group average ${\overline{X}}_i$ of the exogenous variables. The third type is the endogenous peer effect, captured by the coefficient on the group average outcomes $ {\overline{Y}}_i$. Manski concludes that identification of these effects, even in the linear model setting, relies on very strong assumptions and is unrealistic in many settings. In subsequent empirical work, researchers have often ruled out some of these effects in order to identify others.

graham_ectrica focuses on a setting very similar to that of Manski's linear-in-means model. He considers restrictions on the covariance matrix within peer groups implied by the model assuming homoskedasticity at the individual level. bramoulle allows for a more general network configuration than Manski, and investigate the benefits of such configurations for identification in the Manski-style linear-in-means model. hudgens start closer to the Rubin Causal Model or potential outcome setup. Like Manski they focus on a setting with a partitioned network. Following the treatment effect literature they focus primarily on the case with a binary treatment. Let $W_i$ denote the treatment for individual $i$, and let ${\boldsymbol{W}}_i$ denote the vector of treatments for the peer group for individual $i$. The starting point in the hudgens set up is the potential outcome $Y_i({\boldsymbol{w}})$, with restrictions placed on the dependence of the potential outcomes on the full treatment vector ${\boldsymbol{w}}$. aronowsamii allow for general networks and peer effects, investigating the identifying power from randomization.

Models for Network Formation

Another part of the literature has focused on developing models for network formation. Such models are of interest in their own right, but they are also important for deriving asymptotic approximations based on large samples. Such approximations require the researcher to specify in what way the expanding sample would be similar to or different from the current sample. For example, it would require the researcher to be specific in the way the additional units would be linked to current units or other new units.

There is a wide range of models considered, with some models relying more heavily on optimizing behavior of individuals, and others using more statistical models. See goldsmith,christakisimbens,mele,jackson,jacksonwolinsky for such network models in economics, and hollandleinhardt for statistical models. arun2015network develops a model for network formation and develops a corresponding central limit theorem in the presence of correlation induced by network links. arun2015econometrics surveys the econometrics of network formation.

Exact Tests for Interactions

One challenge in testing hypotheses about peer effects using methods based on standard asymptotic theory is that when individuals interact (e.g., in a network), it is not clear how interactions among individuals would change as the network grows. Such a theory would require a model for network formation, as discussed in the last subsection. This motivates an approach that allows us to test hypotheses without invoking large sample properties of test statistics (such as asymptotic normality). Instead, the distributions of the test statistics are based on the random assignments of the treatment, that is, the properties of the tests are based on randomization inference. In randomization inference, we approximate the distribution of the test statistic under the null hypothesis by re-calculating the test statistic under a large number of alternative (hypothetical) treatment assignment vectors, where the alternative treatment assignment vectors are drawn from the randomization distribution. For example, if units were independently assigned to treatment status with probability $p$, we re-draw hypothetical assignment vectors with each unit assigned to treatment with probability $p$. Of course, re-calculating the test statistic requires knowing the values of units' outcomes. The randomization inference approach is easily applied if the null hypothesis of interest is “sharp”: that is, the null hypothesis specifies what outcomes would be under all possible treatment assignment vectors. If the null hypothesis is that the treatment has no effect on any units, this null is sharp: we can infer what outcomes would have been under alternative treatment assignment vectors, in in particular, outcomes would be the same as the realized outcomes under the realized treatment vector.

More generally, however, randomization inference for tests for peer effects is more complicated than in settings without peer effects because the null hypotheses are often {\it not} sharp. aronow2012,atheyeckles develop methods for calculating exact p-values for general null hypotheses on causal effects in a single connected network, allowing for peer effects. The basic case aronow2012,atheyeckles consider is that where the null hypothesis rules out peer effects but allows for direct (own) effects of a binary treatment assigned randomly at the individual level. Given that direct effects are not specified under the null, individual outcomes are not known under alternative treatment assignment vectors, and so the null is not sharp. To address this problem, atheyeckles introduce the notion of an artificial experiment that differs from the actual experiment. In the artificial experiment, some units have their treatment assignments held fixed, and we randomize over the remaining units. Thus, the randomization distribution is replaced by a conditional randomization distribution, where treatment assignments of some units are re-randomized conditional on the assignment of other units. By focusing on the conditional assignment given a subset of the overall space of assignments, and by focusing on outcomes for a subset of the units in the original experiment, they create an artificial experiment where the original null hypothesis that was not sharp in the original experiment is now sharp. To be specific, the artificial experiments starts by designating an arbitrary set of units to be focal. The test statistics considered depend only on outcomes for these focal units. Given the focal units, the set of assignments that, under the null hypothesis of interest, does not change the outcomes for the focal units is derived. The exact distribution of the test statistic can then be inferred for such test statistics under that conditional randomization distribution under the null hypothesis considered.

atheyeckles extend this idea to a large class of null hypotheses. This class includes hypotheses restricting higher order peer effects (peer effects from friends-of-friends) while allowing for the presence of peer effects from friends. It also includes hypotheses about the validity of sparsification of a dense network, where the question concerns peer effects of friends according to the pre-sparsified network while allowing for peer effects of the sparsified network. Finally, the class also includes null hypotheses concerning the exchangeability of peers. In many models peer effects are restricted so that all peers have equal effects on an individual's outcome. It may be more realistic to allow effects of highly connected individuals, or closer friends, to be be different from those of less connected or more distant friends. Such hypotheses can be tested in this framework.

Randomization Inference and Causal Regressions

In recent empirical work, data from randomized experiments are often analyzed using conventional regression methods. Some researchers have raised concerns with the regression approach in small samples (freedman1, freedman2, young, atheyimbens2016, imbens2015causal), but generally such analyses are justified at least in large samples, even in settings with many covariates (binyu, tibs). There is an alternative approach to estimation and inference, however, that does not rely on large sample approximations, using approximations for the distribution of estimators induced by randomization. Such methods, which go back to fisher1925, fisher1935, neyman1923, neyman1935, clarify how the act of randomization allows for the testing for the presence of treatment effects and the unbiased estimation of average treatment effects. Traditionally these methods have not been used much in economics. However, recently there has been some renewed interest in such methods. See for example imbensrosenbaum, young, atheyimbens2016). In completely randomized experiments these methods are often straightforward, although even there analyses involving covariates can be more complicated.

However, the value of the randomization perspective extends well beyond the analysis of actual experiments. It can shed light on the interpretation of observational studies and the complications arising from finite population inference and clustering. Here we discuss some of these issues and more generally provide an explicitly causal perspective on linear regression. Most textbook discussions of regression specify the regression function in terms of a dependent variable, a number of explanatory variables, and an unobserved component, the latter often referred to as the error term: \[ Y_i=\beta_0+\sum_{k=1}^K \beta_k\cdot X_{ik}+\varepsilon_i.\] Often the assumption is made that in the population the units are randomly sampled from, the unobserved component $\varepsilon_i$ is independent of, or uncorrelated with, the regressors $X_{ik}$. The regression coefficients are then estimated by least squares, with the uncertainty in the estimates interpreted as sampling uncertainty induced by random sampling from the large population.

This approach works well in many cases. In analyses using data from the public use surveys such as the Current Population Survey or the Panel Study of Income Dynamics it is natural to view the sample at hand as a random sample from a large population. In other cases this perspective is not so natural, with the sample not drawn from a well-defined population. This includes convenience samples, as well as settings where we observe all units in the population. In those cases it is helpful to take an explictly causal perspective. This perspective also clarifies how the assumptions underlying identification of causal effects relate to the assumptions often made in least squares approaches to estimation.

Let us separate the covariates $X_i$ into a subset of causal variables $W_i$ and the remainder, viewed as fixed characteristics of the units. For example, in a wage regression the causal variable may be years of education and the characteristics may include sex, age, and parental background. Using the potential outcomes perspective we can interpret $Y_i(w)$ as the outcome corresponding to a level of the treatment $w$ for unit or individual $i$. Now suppose that for all units $i$ the function $Y_i(\cdot)$ is linear in in its argument, with a common slope coefficient, but a variable intercept, $Y_i(w)=Y_i(0)+\beta_W\cdot w$. Now write $Y_i(0)$, the outcome for unit $i$ given treatment level $0$ as \[ Y_i(0)=\beta_0+\beta_Z'Z_i+\varepsilon_i,\] where $\beta_0$ and $\beta_Z$ are the population best linear predictor coefficients. This representation of $Y_i(0)$ is purely definitional and does not require assumptions on the population. Then we can write the model as \[ Y_i(w)=\beta_0+\beta_W\cdot w+\beta_Z'Z_i+\varepsilon_i,\] and the realized outcome as \[ Y_i=\beta_0+\beta_W\cdot W_i+\beta_Z'Z_i+\varepsilon_i.\] Now we can investigate the properties of the least squares estimator $\hat \beta_W$ for $\beta_W$, where the distribution of $\hat\beta_W $ is generated by the assignment mechanism for the $W_i$. In the simple case where there are no characteristics $Z_i$ and the cause $W_i$ is a binary indicator, the assumption that the cause is completely randomly assigned leads to the conventional Eicker-Huber-White standard errors (eicker, huber, white1980robust). Thus, in that case viewing the randomness as arising from the assignment of the causes rather than as sampling uncertainty provides a coherent way of interpreting the uncertainty.

This extends very easily to the case where $W_i$ is binary and completely randomly assigned but there are other regressors included in the regression function. As lin and imbens2015causal show there is no need for assumptions about the relation of those regressors to the outcome, as long as the cause $W_i$ is randomly assigned. abadieathey extend this to the case where the cause is multivalued, possibly continuous, and the characteristics $Z_i$ are allowed to be generally correlated with the cause $W_i$. aronowsamii discuss the interpretation of the regression estimates in a causal framework. abadieathey2 discuss extensions to settings with clustering where the need for clustering adjustments in standard errors arises from the clustered assignment of the treatment rather than through clustered sampling.

External Validity

One concern that has been raised in many studies of causal effects is that of external validity. Even if a causal study is done carefully, either in analysis or by design, so that the internal validity of such a study is high, there is often little guarantee that the causal effects are valid for populations or settings other than those studied. This concern has been raised particularly forcefully in experimental studies where the internal validity is guaranteed by design. See for example the discussion in deaton, imbens2010 and manski2013public. Traditionally, there has been much emphasis on internal validity in studies of causal effects, with some arguing for the primacy of internal validity. Some have argued that without internal validity, little can be learned from a study (shadishcookcampbell, imbensmanski). Recently, however, deaton, manski2013public, banerjee have argued that external validity should receive more emphasis.

Some recent work has taken concerns with external validity more seriously, proposing a variety of approaches that directly allow researchers to assess the external validity of estimators for causal effects. A leading example concerns settings with instrumental variables with heterogenous treatment effects (e.g., angrist2004, angristfernandez, donglewbel, angristrokkanen, bertanhaimbens, kowalski, mogstad). In the modern literature with heterogenous treatment effects the instrumental variables estimator is interpreted as an estimator of the local average treatment effect, the average effect of the treatment for the compliers, that is, individuals whose treatment status is affected by the instrument. In this setting, the focus has been on whether the instrumental variables estimates are relevant for the entire sample, that is, have external validity, or only have local validity for the complier subpopulation.

In that context, angrist2004 suggests testing whether the difference in average outcomes for always-takers and never-takers is equal to the average effect for compliers. In this context, a Hausman test hausman for equality of the ordinary least squares estimate and an instrumental variables estimate can be interpreted as testing whether the average treatment effect is equal to the local average treatment effect; of course, the ordinary least squares estimate only has that interpretation if unconfoundedness holds. bertanhaimbens suggest testing a combination of two equalities, first that the average outcome for untreated compliers is equal to the average outcome for never-takers, and second, that the average outcome for treated compliers is equal to the average outcome for always-takers. This turns out to be equivalent to testing both the null hypothesis suggested by angrist2004 and the Hausman null. angristfernandez consider extrapolating local average treatment effects by exploiting the presence of other exogenous covariates. The key assumption in the angristfernandez approach, “conditional effect ignorability,” is that conditional on these additional covariates the average effect for compliers is identical to the average effect for never-takers and always-takers.

In the context of regression discontinuity designs, and especially in the fuzzy regression discontinuity setting, the concerns about external validity are especially salient. In that setting the estimates are in principle valid only for individuals with values of the forcing variable equal to, or close to, the threshold at which the probability of receipt of the treatment changes discontinuously. There have been a number of approaches to assess the plausibility of generalizing those local estimates to other parts of the population. The focus and the applicability of the various methods to assess external validity varies. Some of them apply to both sharp and fuzzy regression discontinuity designs, and some apply only to fuzzy designs. Some require the presence of additional exogenous covariates, and others rely only on the presence of the forcing variable. donglewbel observe that in general, in regression discontinuity designs with a continuous forcing variable, one can estimate the magnitude of the discontinuity as well as the magnitude of the change in the first derivative of the regression function, or even higher order derivatives. Under assumptions about the smoothness of the two conditional mean functions, knowing the higher order derivatives allows one to extrapolate away from values of the forcing variable close to the threshold. This method apply both in the sharp and in the fuzzy regression discontinuity design. It does not require the presence of additional covariates. In another approach, angristrokkanen do require the presence of additional exogenous covariates. They suggest testing whether whether conditional on these covariates, the correlation between the forcing variable and the outcome vanishes. This would imply that the assignment can be thought of as unconfounded conditional on the additional covariates. Thus it would allow for extrapolation away from the threshold. Like the Dong-Lewbel approach, the Angrist-Rokkanen methods apply both in the case of sharp and fuzzy regression discontinuity designs. Finally, bertanhaimbens propose an approach requiring a fuzzy regression discontinuity design. They suggest testing for continuity of the conditional expectation of the outcome conditional on the treatment and the forcing variable, at the threshold, adjusted for differences in the covariates.

Leveraging Experiments

Randomized experiments are the most credible design to learn about causal effects. However, in practice there are often reasons that researchers cannot conduct randomized experiments to answer the causal questions of interest. They may be expensive, or they may take too long to give the researcher the answers that are needed now to make decisions, or there may be ethical objections to experimentation. As a result, we often rely on a combination of experimental results and observational studies to make inferences and decisions about a wide range of questions. In those cases we wish to exploit the benefits of the experimental results, in particular the high degree of internal validity, in combination with the external validity and precision from large scale representative observational studies. At an abstract level, the observational data are used to estimate rich models that allow one to answer many questions, but the model is forced to accommodate the answers from the experimental data for the limited set of questions the latter can address. Doing so will improve the answers from the observational data without compromising their ability to answer more questions.

Here we discuss two specific settings where experimental studies can be leveraged in combination with observational studies to provide richer answers than either of the designs could provide on their own. In both cases, the interest is in the average causal effect of a binary treatment on a primary outcome. However, in the experiment the primary outcome was not observed and so one cannot directly estimate the average effect of interest. Instead an intermediate outcome was observed. In a second study, both the intermediate outcome and the primary outcome were observed. In both studies there may be additional pretreatment variables observed and possibly the treatment indicator.

These two examples do not exhaust the set of possible settings where researchers can leverage experimental data more effectively, and this is likely to be an area where more research is fruitful.

Surrogate Variables

In the first setting, studied in atheychetty, in the second sample the treatment indicator is not observed. In this case researchers may wish to use the intermediate variable, denoted $S_i$, as a surrogate. Following prentice, begg2000, frangakisrubin, the key condition for an intermediate variable to be a surrogate is that in the experimental sample, conditional the surrogate and observed covariates, the (primary) outcomes and the treatment are independent: $Y_i \perp\!\!\!\perp W_i | (S_i, X_i)$. There is a long history of attempts to use intermediate health measures in medical trials as surrogates (prentice). The results are mixed, with the condition often not satisfied in settings where it could be tested. However, many of these studies use low-dimensional surrogates. In modern settings there is often a large number of intermediate variables recorded in administrative data bases that lie on or close to the causal path between the treatment and the primary outcome. In such cases it may be more plausible that the full set of surrogate variables satisfies at least approximately the surrogacy condition.

For example, suppose an internet company is considering a change to the user experience on the company's website. They are interested in the effect of that change on the user's engagement with the website over a year long period. They carry out a randomized experiment over a month, where they measure details about the user's engagement, including the number of visits, webpages visited, and the length of time spent on the various webpages. In addition, they may have historical records on user characteristics including past engagement, for a large number of users. The combination of the pretreatment variables and the surrogates may be sufficiently rich so that conditional on the combination the primary outcome is independent of the treatment.

Given surrogacy, and given comparability of the observational and experimental sample (which requires that the conditional distribution of the primary outcome given surrogates and pretreatment variables is the same in the experimental and observational sample), atheychetty develop two methods for estimating the average effect. The first corresponds to estimating the relation between the outcome and the surrogates in the observational data and using that to impute the missing outcomes in the experimental sample. The second corresponds to estimating the relation between the treatment and the surrogates in the experimental sample and use that to impute the treatment indicator in the observational sample. They also derive the biases from violations of the surrogacy assumption.

Experiments and Observational Studies

In the second setting, studied in atheychettyimbens, the researcher again has data from a randomized experiment containing information on the treatment and the intermediate variables, as well as pretreatment variables. In the observational study the researcher now observes the same variables plus the primary outcome. If in the observational study unconfoundedness (selection-on-observables) were to hold, the researcher would not need the experimental sample, and could simply estimate the average effect of the treatment on the primary outcome by adjusting for differences between treated and control units in pretreatment variables. However, one can compare the estimates of the average effect on the intermediate outcomes based on the observational sample, after adjusting for pretreatment variables, with those from the experimental sample. The latter are known to be consistent, and so if one finds substantial and statistically significant differences, unconfoundedness need not hold. For that case atheychettyimbens develop methods for adjusting for selection on unobservables exploiting the observations on the intermediate variables.

Multiple Experiments

An issue that has not received as much attention, but provides fertile ground for future work concerns the use of multiple experiments. Consider a setting where a number of experiments were conducted. The experiments may vary in terms of the population that the sample is drawn from, or in the exact nature of the treatments included. The researcher may be interested in combining these experiments to obtain more efficient estimates, predicting the effect of a treatment in another population, or estimating the effect of a treatment with different characteristics. Such inferences are not validated by the design of the experiments, but the experiments are important in making such inferences more credible. These issues are related to external validity concerns, but include more general efforts to decompose experimentally estimated effects into components that can inform decisions on related treatments. In the treatment effect literature aspects of these problems have been studied in hotzimbensmortimer,imbens2010,alcott. They have also received some attention in the literature on structural modeling, where the experimental data are used to anchor aspects of the structural model, e.g., toddwolpin.