EconBase
← Back to paper

Interpreting OLS Estimands When Treatment Effects Are Heterogeneous: Smaller Groups Get Larger Weights

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

60,537 characters · 11 sections · 77 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Interpreting OLS Estimands When Treatment Effects Are Heterogeneous: Smaller Groups Get Larger Weights

\doublespacing

titlepage\begin{abstract} \begin{small} Applied work often studies the effect of a binary variable (“treatment”) using linear models with additive effects. I study the interpretation of the OLS estimands in such models when treatment effects are heterogeneous. I show that the treatment coefficient is a convex combination of two parameters, which under certain conditions can be interpreted as the average treatment effects on the treated and untreated. The weights on these parameters are inversely related to the proportion of observations in each group. Reliance on these implicit weights can have serious consequences for applied work, as I illustrate with two well-known applications. I develop simple diagnostic tools that empirical researchers can use to avoid potential biases. Software for implementing these methods is available in R and Stata. In an important special case, my diagnostics only require the knowledge of the proportion of treated units. \end{small} \end{abstract} \thispagestyle{empty}

\setcounter{page}{2}

\setlength\abovedisplayskip{5pt} \setlength\belowdisplayskip{5pt} \setlength\abovedisplayshortskip{0pt} \setlength\belowdisplayshortskip{5pt}

Introduction

Many applied researchers study the effect of a binary variable (“treatment”) on the expected value of an outcome of interest, holding fixed a vector of control variables. As noted by Imbens2015, despite the availability of a large number of semi- and nonparametric estimators for average treatment effects, applied researchers often continue to use conventional regression methods. In particular, numerous studies use ordinary least squares (OLS) to estimate

equation[equation omitted — 62 chars of source]

where $y$ denotes the outcome, $d$ denotes the treatment, and $X$ denotes the row vector of control variables, $\left( x_1, \ldots, x_K \right)$. Usually, $\tau$ is interpreted as the average treatment effect (ATE)\@. This estimation strategy is used in many influential papers in economics VV2012, AGN2013, AEFLM2016, as well as in other disciplines.

The great appeal of the model in ((ref)) comes from its simplicity AP2009. At the same time, however, a large body of evidence demonstrates the importance of heterogeneity in effects Heckman2001, BGH2006, which is explicitly ruled out by this same model. In this paper I contribute to the recent literature on interpreting $\tau$, the OLS estimand, when treatment effects are heterogeneous Angrist1998, Humphreys2009, AS2016. I demonstrate that $\tau$ is a convex combination of two parameters, which under certain conditions can be interpreted as the average treatment effects on the treated (ATT) and untreated (ATU)\@. Surprisingly, the weight that is placed by OLS on the average effect for each group is inversely related to the proportion of observations in this group. The more units are treated, the less weight is placed on ATT\@. One interpretation of this result is that OLS estimation of the model in ((ref)) is generally inappropriate when treatment effects are heterogeneous.

It is also possible, however, to present a more pragmatic view of my main result. I derive a number of corollaries of this result which suggest several diagnostic methods that I recommend to applied researchers. These diagnostics are applicable whenever the researcher is: (i) studying the effects of a binary treatment, (ii) using OLS, and (iii) unwilling to maintain that ATT is exactly equal to ATU\@. Typically, such a homogeneity assumption would be undesirably strong, because those choosing or chosen for treatment may have unusually high or low returns from that treatment, which would directly contradict the equality of ATT and ATU\@.

In deriving my diagnostics, I assume that the researcher is ultimately interested in ATE, ATT, or both, and that she wishes to estimate the model in ((ref)) using OLS but is concerned about treatment effect heterogeneity. In this case, my diagnostics are able to detect deviations of the OLS weights from the pattern that would be necessary to consistently estimate a given parameter. These diagnostics are easy to implement and interpret; they are bounded between zero and one in absolute value and they give the proportion of the difference between ATU and ATT (or between ATT and ATU) that contributes to bias. Thus, if a given diagnostic is close to zero, OLS is likely a reasonable choice; but if a diagnostic is far from zero, other methods should be used.

In an important special case, these diagnostics become particularly simple and immediate to report. If we wish to estimate ATT, this “rule of thumb” variant of my diagnostic is equal to the proportion of treated units, $\pr \left( d=1 \right)$; if our goal is to estimate ATE, the diagnostic is equal to $2 \cdot \pr \left( d=1 \right) - 1$, twice the deviation of $\pr \left( d=1 \right)$ from 50%. In short, OLS is expected to provide a reasonable approximation to ATE if both groups, treated and untreated, are of similar size. If we wish to estimate ATT, it is necessary that the proportion of treated units is very small.

It follows that OLS might often be substantially biased for ATE, ATT, or both. How common are these biases in practice? In a subset of 37 estimates from CKW2018, a recent survey of evaluations of active labor market programs, the mean proportion of treated units is 17.7%.\footnote{This sample is restricted to studies that CKW2018 coded as “selection on observables” and “regression.”} Using the “rule of thumb” variants of my diagnostics, I establish that on average the difference between the OLS estimand and ATE is expected to correspond to 64.6% of the difference between ATT and ATU\@. Similarly, the expected difference between OLS and ATT is on average equal to 17.7% of the difference between ATU and ATT\@. In other words, these biases might often be large.

The remainder of the paper is organized as follows. Section (ref) presents a leading example and the main theoretical results. Section (ref) discusses two empirical applications. In a study of the effects of a training program LaLonde1986, OLS estimates are very similar to $\widehat{\mathrm{ATT}}$\@. On the other hand, in a study of the effects of cash transfers AEFLM2016, OLS estimates are similar to $\widehat{\mathrm{ATU}}$\@. Section (ref) concludes. Proofs and several extensions are provided in the online appendices. The main results are implemented in newly developed R and Stata packages, {hettreatreg}.

A Weighted Average Interpretation of OLS

Leading Example

To illustrate the problem with OLS weights, consider the classic example of the National Supported Work (NSW) program. Because this program originally involved a social experiment, the difference in mean outcomes between the treated and control units provides an unbiased estimate of the effect of treatment. LaLonde1986 studies the performance of various estimators at reproducing this experimental benchmark when the experimental controls are replaced by an artificial comparison group from the Current Population Survey (CPS) or the Panel Study of Income Dynamics (PSID)\@. AP2009 reanalyze the NSW--CPS data and conclude that OLS estimates of the effect of NSW program on earnings in 1978 are similar to the experimental benchmark of \$1,794.\footnote{Subsequently to LaLonde1986, these data were studied by DW1999, ST2005, and many others. AP2009 analyze the subsample of the experimental treated units constructed by DW1999, combined with “CPS-1” or “CPS-3,” i.e. two of the nonexperimental comparison groups from CPS, constructed by LaLonde1986. In this replication, I focus on “CPS-1.”} In particular, their richest specification delivers an estimate of \$794. As I will show, this conclusion is driven by the small proportion of treated units in these data.

In this example, ATT and ATU are likely to be substantially different. This is because the treated group, unlike the CPS comparison (untreated) group, was highly economically disadvantaged. It is plausible that ATU might be zero or, due to the opportunity cost of program participation, even negative. Also, only 1.1% of the sample was treated, so ATE and ATU will be similar.

To demonstrate this, I modify the model in ((ref)) to include all interactions between $d$ and $X$\@. Estimation of this expanded model, again using OLS, allows us to separately compute $\widehat{\mathrm{ATE}}$, $\widehat{\mathrm{ATT}}$, and $\widehat{\mathrm{ATU}}$. This method is usually referred to as “regression adjustment” Wooldridge2010 or “Oaxaca--Blinder” Kline2011, GP2018. Using the control variables that deliver the estimate of \$794, we obtain $\widehat{\mathrm{ATE}} = -\$4 \mathrm{,} 930$, $\widehat{\mathrm{ATT}} = \$796$, and $\widehat{\mathrm{ATU}} = -\$4 \mathrm{,} 996$. It turns out that, since $\widehat{\mathrm{ATE}}$ and $\widehat{\mathrm{ATU}}$ are indeed negative, the OLS estimate and $\widehat{\mathrm{ATE}}$ have different signs. Moreover, if we represent the OLS estimate as a weighted average of $\widehat{\mathrm{ATT}}$ and $\widehat{\mathrm{ATU}}$ with weights that sum to unity, we can write $\$794 = \hat{w}_{ATT} \cdot \$796 + \left( 1-\hat{w}_{ATT} \right) \cdot \left( -\$4 \mathrm{,} 996 \right)$, where $\hat{w}_{ATT}$ is the weight on $\widehat{\mathrm{ATT}}$. Solving for $\hat{w}_{ATT}$ yields $\hat{w}_{ATT} = 99.96\%$. In other words, the hypothetical OLS weight on the effect on the treated is similar to the proportion of untreated units, 98.9%.

This “weight reversal” is not a coincidence. As I demonstrate below, the intuition from this example holds more generally, even though the OLS estimand is not necessarily a convex combination of two parameters from a procedure that controls for the full vector $X$\@.

Main Result

This section presents my main result, which focuses on the algebra of OLS and “descriptive” estimands that I define below. A causal interpretation of OLS also requires introducing the notion of potential outcomes as well as certain conditions that I discuss in section (ref)(ref), including an ignorability assumption. However, this is not needed for my main result.

If $\lp \left( \cdot \mid \cdot \right)$ denotes the linear projection, we are interested in the interpretation of $\tau$ in the linear projection of $y$ on $d$ and $X$,

equation[equation omitted — 91 chars of source]

when this linear projection does not correspond to the (structural) conditional mean. Let

equation[equation omitted — 45 chars of source]

be the unconditional probability of treatment and let

equation[equation omitted — 103 chars of source]

be the “propensity score” from the linear probability model or, equivalently, the best linear approximation to the true propensity score. Generally, the specification in ((ref)) and ((ref)) can be arbitrarily flexible, so this approximation can be made very accurate; in fact, we can think of equation ((ref)) as partially linear, where we may include powers and cross-products of original control variables.

After defining $p \left( X \right)$, it is helpful to introduce two linear projections of $y$ on $p \left( X \right)$, separately for $d=1$ and $d=0$, namely

equation[equation omitted — 129 chars of source]

and also

equation[equation omitted — 130 chars of source]

Note that equations ((ref)), ((ref)), and ((ref)) are definitional. It is sufficient for my main result that the linear projections introduced so far exist and are unique.

assumption(i) $\e ( y^2 )$ and $\e ( \| X \| ^2 )$ are finite. (ii) The covariance matrix of $\left( d,X \right)$ is nonsingular.
assumption$\var \left[ p \left( X \right) \mid d=1 \right]$ and $\var \left[ p \left( X \right) \mid d=0 \right]$ are nonzero, where $\var \left( \cdot \mid \cdot \right)$ denotes the conditional variance (with respect to $\e \left[ p \left( X \right) \mid d=j \right]$, $j=0,1$).

Assumption (ref) guarantees the existence and uniqueness of the linear projections in ((ref)) and ((ref)). Similarly, Assumption (ref) ensures that the linear projections in ((ref)) and ((ref)) exist and are unique.\footnote{Both assumptions are generally innocuous, although Assumption (ref) rules out a small number of interesting applications, such as regression adjustments in Bernoulli trials and completely randomized experiments. In these cases, however, OLS is consistent for the average treatment effect under general conditions IR2015.}

The next step is to use the linear projections in ((ref)) and ((ref)) to define the average partial linear effect of $d$ as

equation[equation omitted — 160 chars of source]

as well as the average partial linear effect of $d$ on group $j$ ($j=0,1$) as

equation[equation omitted — 174 chars of source]

These estimands are well defined under Assumptions (ref) and (ref), and have a causal interpretation under additional assumptions, as discussed in section (ref)(ref) below.\footnote{Moreover, $\tau_{APLE}$ is similar to the “average regression coefficient” or “average slope coefficient” in GP2018, which is also a descriptive estimand in the sense of AAIW2020.} When the linear projections in equations ((ref)) and ((ref)) represent the conditional mean of $y$, the average partial linear effects of $d$ overlap with its average partial effects. It should be stressed, however, that Theorem (ref), the main result of this paper, is more general and only requires Assumptions (ref) and (ref).

theorem[Weighted Average Interpretation of OLS] Under Assumptions (ref) and (ref), \begin{eqnarray} \tau &=& w_1 \cdot \tau_{APLE, 1} + w_0 \cdot \tau_{APLE, 0}, \nonumber \end{eqnarray} where $w_1 = \frac{\left( 1 - \rho \right) \cdot \var \left[ p \left( X \right) \mid d=0 \right]}{\rho \cdot \var \left[ p \left( X \right) \mid d=1 \right] + \left( 1 - \rho \right) \cdot \var \left[ p \left( X \right) \mid d=0 \right]}$ and $w_0 = 1 - w_1 = \frac{\rho \cdot \var \left[ p \left( X \right) \mid d=1 \right]}{\rho \cdot \var \left[ p \left( X \right) \mid d=1 \right] + \left( 1 - \rho \right) \cdot \var \left[ p \left( X \right) \mid d=0 \right]}$.
proofSee online appendix (ref)\@.

Theorem (ref) shows that $\tau$, the OLS estimand, is a convex combination of $\tau_{APLE, 1}$ and $\tau_{APLE, 0}$. The definition of $\tau_{APLE, j}$ makes it clear that $\tau$ is equivalent to the outcome of a particular three-step procedure. In the first step, we obtain $p \left( X \right)$, i.e. the “propensity score.” Next, in the second step, we obtain $\tau_{APLE, 1}$ and $\tau_{APLE, 0}$, as in ((ref)), from two linear projections of $y$ on $p \left( X \right)$, separately for $d=1$ and $d=0$. This is analogous to the “regression adjustment” procedure in section (ref)(ref), although now we control for $p \left( X \right)$ rather than the full vector $X$\@. Finally, in the third step, we calculate a weighted average of $\tau_{APLE, 1}$ and $\tau_{APLE, 0}$. The weight on $\tau_{APLE, 1}$, $w_1$, is decreasing in $\frac{\var \left[ p \left( X \right) \mid d=1 \right]}{\var \left[ p \left( X \right) \mid d=0 \right]}$ and $\rho$ and the weight on $\tau_{APLE, 0}$, $w_0$, is increasing in $\frac{\var \left[ p \left( X \right) \mid d=1 \right]}{\var \left[ p \left( X \right) \mid d=0 \right]}$ and $\rho$.\footnote{A formal proof that the relationship between $\rho$ and $w_1$ ($w_0$) is indeed always negative (positive) is provided in online appendix (ref)\@. This proof additionally assumes that the conditional mean of $d$ is linear in $X$\@.} This is clearly undesirable, since $\tau_{APLE} = \rho \cdot \tau_{APLE, 1} + \left( 1 - \rho \right) \cdot \tau_{APLE, 0}$.

This weighting scheme is also surprising: the more units belong to group $j$, the less weight is placed on $\tau_{APLE, j}$, i.e. the effect for this group. There are several ways to provide intuition for this result. One is provided in the next section. Another intuition follows from an alternative proof of Theorem (ref), which is provided with discussion in online appendix (ref)\@. It parallels the intuition in Angrist1998 and AP2009 that OLS gives more weight to treatment effects that are better estimated in finite samples.\footnote{This proof uses a result from Deaton1997 and SHW2015 as a lemma. The main proof of Theorem (ref) uses a result on decomposition methods from EGH2010. See online appendix (ref) for more details.}

Causal Interpretation

The fact that Theorem (ref) only requires the existence and uniqueness of several linear projections makes this result very general. On the other hand, one concern about this result might be that $\tau_{APLE, 1}$ and $\tau_{APLE, 0}$ do not necessarily correspond to the usual (causal) objects of interest. To define these objects, we need two potential outcomes, $y(1)$ and $y(0)$, only one of which is observed for each unit, $y = y(d) = y(1) \cdot d + y(0) \cdot \left( 1-d \right)$. The parameters of interest, ATE, ATT, and ATU, are defined as $\tau_{ATE} = \e \left[ y(1) - y(0) \right]$, $\tau_{ATT} = \e \left[ y(1) - y(0) \mid d=1 \right]$, and $\tau_{ATU} = \e \left[ y(1) - y(0) \mid d=0 \right]$. A causal interpretation of OLS also entails the following assumptions.

assumption[Ignorability in Mean] (i) $\e \left[ y(1) \mid X,d \right] = \e \left[ y(1) \mid X \right]$; and (ii) $\e \left[ y(0) \mid X,d \right] = \e \left[ y(0) \mid X \right]$.
assumption(i) $\e \left[ y(1) \mid X \right] = \alpha_1 + \gamma_1 \cdot p \left( X \right)$; and (ii) $\e \left[ y(0) \mid X \right] = \alpha_0 + \gamma_0 \cdot p \left( X \right)$.

Assumptions (ref) and (ref) ensure that $\tau$ admits a causal interpretation. Assumption (ref) is standard in the program evaluation literature Wooldridge2010. Assumption (ref) is not commonly used. Sufficient for this assumption, but not necessary, is that the conditional mean of $d$ is linear in $X$ and the conditional means of $y(1)$ and $y(0)$ are linear in the true propensity score, which is now equal to $p \left( X \right)$. Linearity of $\e \left( d \mid X \right)$ is assumed in AS2016 and AAIW2020. This assumption is not necessarily strong, since $X$ might include powers and cross-products of original control variables. It is also satisfied automatically in saturated models, as in Angrist1998 and Humphreys2009. The linearity assumption for $\e \left[ y(1) \mid p \left( X \right) \right]$ and $\e \left[ y(0) \mid p \left( X \right) \right]$ dates back to RR1983 but is restrictive. See also IW2009 and Wooldridge2010 for a discussion.

corollary[Causal Interpretation of OLS] Under Assumptions (ref), (ref), (ref), and (ref), \begin{eqnarray} \tau & = & w_1 \cdot \tau_{ATT} + w_0 \cdot \tau_{ATU}. \nonumber \end{eqnarray}
proofAssumption (ref) implies that $\e \left[ y(1)-y(0) \mid X \right] = \e \left( y \mid X,~d=1 \right) - \e \left( y \mid X,~d=0 \right)$. Then, Assumption (ref) implies that $\e \left[ y(1)-y(0) \mid X \right] = \left( \alpha_1 - \alpha_0 \right) + \left( \gamma_1 - \gamma_0 \right) \cdot p \left( X \right)$, which in turn implies that $\tau_{ATT} = \tau_{APLE, 1}$ and $\tau_{ATU} = \tau_{APLE, 0}$. This, together with Theorem (ref), completes the proof.

Corollary (ref) states that, under Assumptions (ref), (ref), (ref), and (ref), the OLS weights from Theorem (ref) apply to the causal objects of interest, $\tau_{ATT}$ and $\tau_{ATU}$. Hence, $\tau$ has a causal interpretation. The greater the proportion of treated units, the smaller is the OLS weight on $\tau_{ATT}$. Again, this is undesirable, since $\tau_{ATE} = \rho \cdot \tau_{ATT} + \left( 1 - \rho \right) \cdot \tau_{ATU}$.

To aid intuition for this surprising result, recall that an important motivation for using the model in ((ref)) and OLS is that the linear projection of $y$ on $d$ and $X$ provides the best linear predictor of $y$ given $d$ and $X$ AP2009. However, if our goal is to conduct causal inference, then this is not, in fact, a good reason to use this method. Ordinary least squares is “best” in predicting actual outcomes but causal inference is about predicting missing outcomes, defined as $y_m = y(1) \cdot \left( 1-d \right) + y(0) \cdot d$. In other words, the OLS weights are optimal for predicting “what is.” Instead, we are interested in predicting “what would be” if treatment were assigned differently.

Intuition suggests that if our goal were to predict “what is” and, without loss of generality, group one were substantially larger than group zero, we would like to place a large weight on the linear projection coefficients of group one ($\alpha_1$ and $\gamma_1$), because these coefficients can be used to predict actual outcomes of this group. As noted by Deaton1997 and SHW2015, the OLS weights are consistent with this idea. Indeed, Theorem (ref) also implies that

equation[equation omitted — 289 chars of source]

Namely, the OLS estimand is equal to the simple difference in means of $y$ plus an adjustment term that depends on the difference in means of $p \left( X \right)$ and a weighted average of $\gamma_1$ and $\gamma_0$. When group one is “large,” $w_0$, the weight on $\gamma_1$, is large as well.

Conversely, if group one is “large” but our goal is to predict missing outcomes, we need to place a large weight on $\alpha_0$ and $\gamma_0$, because these coefficients can be used to predict counterfactual outcomes of group one. To see this point, note that it follows from the discussion in IW2009 that when the conditional means of $y(1)$ and $y(0)$ are linear in $X$, we can write

equation[equation omitted — 268 chars of source]

where $\beta_1$ and $\beta_0$ are the coefficients on $X$ in the conditional means of $y(1)$ and $y(0)$, respectively. Equations ((ref)) and ((ref)) reiterate the point of Corollary (ref) that $\tau$ and $\tau_{ATE}$ have a very similar structure but they differ substantially in how they assign weights. Indeed, in the case of $\tau_{ATE}$, when group one is “large,” the weight on $\beta_1$ is small, the opposite of what we have seen for OLS\@.\footnote{Note that the (infeasible) linear projection of the missing outcome, $y_m$, on $d$ and $X$ would solve our problem of “weight reversal.” The weights on $\tau_{ATT}$ and $\tau_{ATU}$ would still be different than $\rho$ and $1 - \rho$ if $\var \left[ p \left( X \right) \mid d=1 \right]$ and $\var \left[ p \left( X \right) \mid d=0 \right]$ were different; but, at least, the weight on $\tau_{ATT}$ ($\tau_{ATU}$) would be increasing (decreasing) in $\rho$.}

Implications of Theorem (ref)

There are several practical implications of my main result. Throughout this section, I assume that the researcher is interested in estimating $\tau_{ATE}$, $\tau_{ATT}$, or both, and that she wishes to use OLS to estimate the model in ((ref)) but is concerned about the implications of Theorem (ref) and Corollary (ref). In Corollaries (ref) and (ref), I show how to decompose the difference between $\tau$ and $\tau_{ATE}$ or $\tau$ and $\tau_{ATT}$ into components attributable to (i) the difference between $\tau_{APLE, 1}$ and $\tau_{ATT}$, (ii) the difference between $\tau_{APLE, 0}$ and $\tau_{ATU}$ (jointly referred to as “bias from nonlinearity”), and (iii) the OLS weights on $\tau_{ATT}$ and $\tau_{ATU}$ (“bias from heterogeneity”).\footnote{Because “bias from nonlinearity” arises when Assumptions (ref) and/or (ref) are violated, it might be more accurate to refer to this component as “bias from endogeneity and nonlinearity.” Yet, I use the former term for brevity.} Because this paper generally focuses on what I now term “bias from heterogeneity,” my discussion below is restricted to this source of bias, which is equivalent to implicitly making Assumptions (ref) and (ref).

corollaryUnder Assumptions (ref) and (ref), \begin{eqnarray} \tau - \tau_{ATE} & = & \underbrace{w_0 \cdot \left( \tau_{APLE, 0} - \tau_{ATU} \right) + w_1 \cdot \left( \tau_{APLE, 1} - \tau_{ATT} \right)}_{bias from nonlinearity} \; + \; \underbrace{\delta \cdot \left( \tau_{ATU} - \tau_{ATT} \right)}_{bias from heterogeneity}, \nonumber \end{eqnarray} where $\delta = \rho - w_1 = \frac{\rho ^2 \cdot \var \left[ p \left( X \right) \mid d=1 \right] - \left( 1 - \rho \right) ^2 \cdot \var \left[ p \left( X \right) \mid d=0 \right]}{\rho \cdot \var \left[ p \left( X \right) \mid d=1 \right] + \left( 1 - \rho \right) \cdot \var \left[ p \left( X \right) \mid d=0 \right]}$. Also, under Assumptions (ref), (ref), (ref), and (ref), \begin{eqnarray} \tau - \tau_{ATE} & = & \delta \cdot \left( \tau_{ATU} - \tau_{ATT} \right). \nonumber \end{eqnarray}
corollaryUnder Assumptions (ref) and (ref), \begin{eqnarray} \tau - \tau_{ATT} & = & \underbrace{w_0 \cdot \left( \tau_{APLE, 0} - \tau_{ATU} \right) + w_1 \cdot \left( \tau_{APLE, 1} - \tau_{ATT} \right)}_{bias from nonlinearity} \; + \; \underbrace{w_0 \cdot \left( \tau_{ATU} - \tau_{ATT} \right)}_{bias from heterogeneity}. \nonumber \end{eqnarray} Also, under Assumptions (ref), (ref), (ref), and (ref), \begin{eqnarray} \tau - \tau_{ATT} & = & w_0 \cdot \left( \tau_{ATU} - \tau_{ATT} \right). \nonumber \end{eqnarray}

The proofs of Corollaries (ref) and (ref) follow from simple algebra and are omitted. These results show that, regardless of whether we focus on $\tau_{ATE}$ or $\tau_{ATT}$, the bias from heterogeneity is equal to the product of a particular measure of heterogeneity, namely the difference between $\tau_{ATU}$ and $\tau_{ATT}$, and an additional parameter that is easy to estimate, $\delta$ for $\tau_{ATE}$ and $w_0$ for $\tau_{ATT}$. While $w_0$ is guaranteed to be positive under Assumptions (ref) and (ref), $\delta$ may be positive or negative. Both $w_0$ and $\delta$, however, are bounded between zero and one in absolute value. Thus, $w_0$ and $\vert \delta \vert$ can be interpreted as the percentage of our measure of heterogeneity, $\tau_{ATU} - \tau_{ATT}$, which contributes to bias.\footnote{To be precise, $\vert \delta \vert$ can be interpreted as the percentage of $\sgn( \delta ) \cdot \left( \tau_{ATU} - \tau_{ATT} \right)$ that contributes to bias when focusing on $\tau_{ATE}$. Both $\delta$ and $w_0$ also have an intuitive interpretation as the difference between (i) the weight that we should place on $\tau_{ATT}$ when focusing on $\tau_{ATE}$ or $\tau_{ATT}$ and (ii) the weight that OLS actually places on this parameter. Indeed, $\delta$ is equal to the difference between $\rho$ and $w_1$. Similarly, $w_0 = 1-w_1$.} It might be useful to report estimates of $w_0$ and $\delta$ in studies that use OLS to estimate the model in ((ref)).

As an example, consider the empirical application in section (ref)(ref)\@. In this case, $\hat{w}_0 = 0.017$ and $\hat{\delta} = -0.971$. The interpretation of these estimates is as follows: if our goal is to estimate $\tau_{ATT}$, using the model in ((ref)) and OLS is expected to bias our estimates by only 1.7% of the difference between $\tau_{ATU}$ and $\tau_{ATT}$. If instead we wanted to interpret $\tau$ as $\tau_{ATE}$, our estimates would be biased by an estimated 97.1% of the difference between $\tau_{ATT}$ and $\tau_{ATU}$. Thus, in this application, it might perhaps be acceptable to interpret $\tau$ as $\tau_{ATT}$ but clearly not as $\tau_{ATE}$.

assumption$\var \left[ p \left( X \right) \mid d=1 \right] = \var \left[ p \left( X \right) \mid d=0 \right]$.

The calculation of $\delta$ and $w_0$ is further simplified under Assumption (ref). If we use $\delta^*$ and $w_0^*$ to denote the values of $\delta$ and $w_0$ in this special case, we can write $\delta^* = 2 \rho - 1$ and $w_0^* = \rho$. In this setting, the knowledge of $\delta$ and $w_0$ only requires information on $\rho$, the proportion of units with $d=1$. Of course, the special case where $\var \left[ p \left( X \right) \mid d=1 \right] = \var \left[ p \left( X \right) \mid d=0 \right]$ is hardly to be expected in practice. Still, $\delta^* = 2 \rho - 1$ and $w_0^* = \rho$ can potentially serve as a rule of thumb.

The practical implications of Assumption (ref) are particularly clear when $\rho$ is close to 0%, 50%, or 100%. When few units are treated, $\tau \simeq \tau_{ATT}$. When most of the units are treated, $\tau \simeq \tau_{ATU}$. Finally, when both groups are of similar size, $\tau \simeq \tau_{ATE}$. This can also be seen from Corollary (ref).

corollaryUnder Assumptions (ref), (ref), and (ref), \begin{eqnarray} \tau &=& \left( 1 - \rho \right) \cdot \tau_{APLE, 1} + \rho \cdot \tau_{APLE, 0}. \nonumber \end{eqnarray} Also, under Assumptions (ref), (ref), (ref), (ref), and (ref), \begin{eqnarray} \tau &=& \left( 1 - \rho \right) \cdot \tau_{ATT} + \rho \cdot \tau_{ATU}. \nonumber \end{eqnarray}

The proof follows immediately from simple algebra. Corollary (ref) provides conditions under which OLS reverses the “natural” weights on $\tau_{APLE, 1}$ and $\tau_{APLE, 0}$ (or $\tau_{ATT}$ and $\tau_{ATU}$). Indeed, under Assumption (ref), $\tau$ is a convex combination of group-specific average effects, with “reversed” weights attached to these parameters. Namely, the proportion of units with $d=1$ is used to weight the average effect of $d$ on group zero, and vice versa.

The results in this section allow empirical researchers to interpret the OLS estimand when treatment effects are heterogeneous. Alternatively, it might be sensible to use any of the standard estimators for average treatment effects under ignorability, such as regression adjustment (see section (ref)(ref)), weighting, matching, and various combinations of these approaches.\footnote{For recent reviews, see IW2009, Wooldridge2010, and AC2018.} It might also help to estimate a model with homogeneous effects using weighted least squares (WLS)\@. Indeed, in online appendix (ref), I demonstrate that when we regress $y$ on $d$ and $p \left( X \right)$, with weights of $\frac{1 - \rho}{w_0}$ for units with $d=1$ and $\frac{\rho}{w_1}$ for units with $d=0$, the WLS estimand is equal to $\tau_{APLE}$. In practice, of course, $\tau_{APLE}$ can also be obtained directly from equation ((ref)).

Related Work

This section discusses the relationship between my main result and those in Angrist1998 and Humphreys2009. These papers focus on saturated models with discrete covariates, in which the estimating equation includes an indicator for each combination of covariate values (“stratum”). In particular, Angrist1998 provides a representation of $\tau_n$ in

equation[equation omitted — 128 chars of source]

where $x_1, \ldots, x_S$ are stratum indicators. More precisely, Angrist1998 demonstrates that

equation[equation omitted — 300 chars of source]

where $\tau_s = \e \left( y \mid d=1, x_s=1 \right) - \e \left( y \mid d=0, x_s=1 \right)$. In online appendix (ref), I demonstrate that this result follows from Corollary (ref) when the model for $y$ is saturated.\footnote{Also, note that AS2016 show that this result in Angrist1998 is not specific to saturated models; instead, it is sufficient to assume that the model for $d$ is linear in $X$\@. My analysis in online appendix (ref) covers the results in both Angrist1998 and AS2016.} At the same time, the interpretation of OLS in Angrist1998 is different from Theorem (ref) and Corollary (ref). On the one hand, unlike Corollary (ref) and Humphreys2009, Angrist1998 does not restrict the relationship between $\tau_s$ and $\pr \left( d=1 \mid x_s=1 \right)$ in any way. On the other hand, Theorem (ref) and Corollary (ref) make it arguably easier to identify whether in a given application the OLS estimand will be close to any of the parameters of interest (cf. Corollaries (ref) to (ref)). In particular, Angrist1998 does not recover a pattern of “weight reversal,” which is discussed in detail in this paper.

Unlike Angrist1998, Humphreys2009 does not derive a new representation of $\tau_n$, but instead presents further analysis of the result in equation ((ref)). In particular, Humphreys2009 notes that $\tau_n$ can take any value between $\min ( \tau_s )$ and $\max ( \tau_s )$. Then, he demonstrates that $\tau_n$ is also bounded by $\tau_{ATT}$ and $\tau_{ATU}$ if we restrict the relationship between $\tau_s$ and $\pr \left( d=1 \mid x_s=1 \right)$ to be monotonic. According to Corollary (ref), $\tau$ is a convex combination of $\tau_{ATT}$ and $\tau_{ATU}$ if, among other things, both potential outcomes are linear in $p \left( X \right)$, which also implies a linear relationship between $\tau_s$ and $\pr \left( d=1 \mid x_s=1 \right)$ when the model for $y$ is saturated. Of course, this linearity assumption is stronger than the monotonicity assumption in Humphreys2009. However, in return, we are able to derive a closed-form expression for $\tau$ in terms of $\tau_{ATT}$ and $\tau_{ATU}$, which is a major advantage over the earlier literature, such as Angrist1998 and Humphreys2009.\footnote{Humphreys2009 also provides a brief informal remark that the OLS estimand, as represented in Angrist1998, is similar to $\tau_{ATT}$ ($\tau_{ATU}$) if propensity scores are “small” (“large”) in every stratum. This is a special case of the rule of thumb derived from Corollaries (ref) and (ref). My rule of thumb does not impose any such restrictions on the propensity score other than the requirement that the unconditional probability of treatment is close to zero or one.}

Empirical Applications

This section discusses two empirical illustrations of Theorem (ref) and its corollaries.\footnote{In a follow-up paper, I apply these results in the study of racial gaps in test scores and wages Sloczynski_ILRR.} In online appendices (ref) and (ref), I discuss the implementation of these results in Stata and R\@. Throughout the current section $\tau_{APLE}$, $\tau_{APLE, 1}$, and $\tau_{APLE, 0}$ are implicitly treated as equivalent to $\tau_{ATE}$, $\tau_{ATT}$, and $\tau_{ATU}$, respectively. Although this might be restrictive, I also demonstrate that in both applications sample analogues of $\tau_{APLE}$, $\tau_{APLE, 1}$, and $\tau_{APLE, 0}$, reported in the body of the paper, are similar to other estimates of $\tau_{ATE}$, $\tau_{ATT}$, and $\tau_{ATU}$, reported in online appendix (ref)\@.

The Effects of a Training Program on Earnings

I first consider the example from section (ref)(ref) in more detail. This replication of the study of the effects of NSW program in AP2009 constitutes an optimistic scenario for OLS\@. In this application, as I explained in section (ref)(ref), the effect for the treated group (ATT) is likely to be substantially larger than the effect for the CPS comparison group (ATU)\@. Moreover, since the experimental benchmark of \$1,794 corresponds to $\widehat{\mathrm{ATT}}$ and not to $\widehat{\mathrm{ATU}}$, the researcher should also focus on ATT\@. It turns out that my diagnostic for estimating ATT, $\hat{w}_0$, indicates that this parameter should approximately be recovered by OLS, even if treatment effects are heterogeneous.\footnote{It is well known that, in the NSW--CPS data, there is limited overlap in terms of covariate values between the treated and untreated units DW1999, ST2005. Thus, it is important to note that my theoretical results in section (ref) do not impose the overlap assumption.}

table[table omitted — 3,793 chars of source]

The top and middle panels of Table (ref) reproduce the estimates from AP2009 and report my diagnostics. The specification in column 4 was discussed in section (ref)(ref)\@. It turns out that $\hat{w}_0$ is between 0.1% and 1.9% for all specifications; similarly, the “rule of thumb” value of this diagnostic, $\hat{w}_0^*$, is, as always, equal to the proportion of treated units (only 1.1% in this sample). These results are very simple to interpret. Namely, as in section (ref)(ref), we estimate that the difference between the OLS estimand and ATT is less than 2% of the difference between ATU and ATT\@. In this case, it might indeed be sensible to rely on the OLS estimates of the effect of treatment.

The bottom panel of Table (ref) provides an application of Corollary (ref) to these results. In other words, the estimates from AP2009 are now decomposed into two components, $\widehat{\mathrm{ATT}}$ and $\widehat{\mathrm{ATU}}$\@. The difference between these estimates is substantial. In column 4, while the estimate of ATT is \$928, ATU is estimated to be --\$6,840. In other words, the OLS estimate of \$794, reported in AP2009 and discussed in section (ref)(ref), is actually a weighted average of these two estimates. The fact that it is close to \$928, and not to --\$6,840, is a consequence of the small proportion of treated units in this sample, 1.1%. The weight on \$928, $\hat{w}_1$, is 98.3% and the weight on --\$6,840, $\hat{w}_0$, is only 1.7%.

We might expect that if the proportion of treated units was larger, the weight on $\widehat{\mathrm{ATT}}$ would be smaller and the “performance” of OLS in replicating the experimental benchmark would deteriorate. I confirm this conjecture in online appendix (ref) by quasi-discarding “random” subsamples of untreated units over a range of sample sizes. In particular, I reestimate the model in ((ref)) using WLS, with weights of 1 for treated and $\frac{1}{k}$ for untreated units. Figures (ref) to (ref) show that in this application WLS estimates become more negative as $k$ increases. This is because larger values of $k$ correspond to greater proportions of untreated units being “discarded,” and hence larger weights on $\widehat{\mathrm{ATU}}$, which is substantially more negative than $\widehat{\mathrm{ATT}}$\@.

Additional extensions of my analysis are also presented in online appendix (ref)\@. For each specification in Table (ref), I provide both a linear and a nonparametric estimate of the conditional mean of the outcome given $p \left( X \right)$, separately for treated and untreated units (Figures (ref) to (ref))\@. A visual comparison of both estimates provides an informal test of Assumption (ref), which is necessary for a causal interpretation of $\tau_{APLE}$, $\tau_{APLE, 1}$, and $\tau_{APLE, 0}$. The linearity assumption appears to be approximately satisfied for the treated but usually not for the untreated units.

Thus, as a robustness check, I also report a number of alternative estimates of the effects of NSW program in Table (ref)\@. I consider regression adjustment, as in section (ref)(ref), as well as matching on $p \left( X \right)$ and on the logit propensity score.\footnote{In particular, the estimates discussed in section (ref)(ref) are reported in column 4 of the bottom panel of Table (ref)\@.} In each case, I separately estimate ATE, ATT, and ATU\@. These estimates are consistent with the claim that the general pattern of results in Table (ref) is driven by the OLS weights. The estimates of ATE and ATU are always negative and large in magnitude; the estimates of ATT are much closer to the experimental benchmark.

Finally, I repeat the following exercise from section (ref)(ref)\@. When we match the OLS estimates in Table (ref) with the corresponding estimates of ATT and ATU in Table (ref), we can write $\hat{\tau} = \hat{w}_{ATT} \cdot \hat{\tau}_{ATT} + \left( 1-\hat{w}_{ATT} \right) \cdot \hat{\tau}_{ATU}$. Unless $\hat{\tau}_{ATT}$ and $\hat{\tau}_{ATU}$ are sample analogues of $\tau_{APLE, 1}$ and $\tau_{APLE, 0}$, $\hat{w}_{ATT}$ does not need to be bounded between zero and one. Yet, we can solve for $\hat{w}_{ATT}$ for each set of estimates. The mean of $\hat{w}_{ATT}$ across all sets of estimates in Table (ref) is 98.3%, which is nearly identical to the sample proportion of untreated units, 98.9%. This is reassuring for my claims.

The Effects of Cash Transfers on Longevity

In my second application, I replicate a recent paper by AEFLM2016 and study the effects of cash transfers on longevity of the children of their beneficiaries, as measured by their log age at death. In particular, AEFLM2016 analyze the administrative records of applicants to the Mothers' Pension (MP) program, which supported poor mothers with dependent children in pre-WWII United States. In this study, the untreated group consists only of children of mothers who applied for a transfer, were initially deemed eligible, but were ultimately rejected. This strategy is used to ensure that treated and untreated individuals are broadly comparable, and hence an ignorability assumption might be plausible. Nevertheless, rejected mothers were slightly older and came from slightly smaller and richer families than accepted mothers. Thus, as before, there is no reason to believe that ATT and ATU are equal, although it is perhaps less clear a priori which is larger. Unlike in section (ref)(ref), it seems plausible that the researcher might be interested either in the average effect of cash transfers, ATE, or in their average effect for accepted applicants, ATT\@.

table[table omitted — 3,988 chars of source]

The top and middle panels of Table (ref) reproduce the baseline estimates from AEFLM2016 and report my diagnostics. While the OLS estimates are positive and statistically significant, my diagnostics indicate that these results should be approached with caution. Namely, treated units constitute the vast majority (or 87.5%) of the sample. It follows that OLS is expected to place a disproportionately large weight on $\widehat{\mathrm{ATU}}$, in which case the OLS estimates might be very biased for both ATE and ATT (cf. Corollaries (ref) and (ref)). Indeed, my estimates of $\delta$ suggest that the difference between the OLS estimand and ATE is equal to 65.9--74.5% of the difference between ATU and ATT\@. Also, the estimates of $w_0$ suggest that the difference between OLS and ATT corresponds to 78.4--87.0% of this measure of heterogeneity. The estimates of $\delta^*$ and $w_0^*$ are similar. It turns out that in this application the OLS estimates might be substantially biased for both of our parameters of interest. This would be a pessimistic scenario for OLS\@.

The results in the bottom panel of Table (ref) suggest that these biases are indeed substantial. In this panel, following Corollary (ref), each OLS estimate from AEFLM2016 is represented as a weighted average of estimates of two effects, on accepted (ATT) and rejected (ATU) applicants. The estimates of ATU are consistently larger than those of ATT\@. Thus, OLS overestimates both ATE (since $\hat{\delta} > 0$) and ATT\@. While the implicit OLS estimates of these parameters remain statistically significant in columns 1 and 2, this is no longer the case in columns 3 and 4, following the inclusion of county fixed effects. Perhaps more importantly, these estimates of ATT are half smaller than the corresponding OLS estimates. Clearly, this difference is economically quite meaningful.

To assess the robustness of these findings, I present several extensions of my analysis in online appendix (ref)\@. The informal test of Assumption (ref), as discussed in section (ref)(ref), appears to suggest that the conditional mean of the outcome given $p \left( X \right)$ is approximately linear for both the treated and untreated units (see Figures (ref) to (ref))\@. I also report a number of alternative estimates of the effects of cash transfers in Table (ref)\@. These additional results support my conclusion. Only one in twelve estimates of ATT is statistically different from zero, and four of the insignificant estimates are negative. While it is possible that cash transfers increase longevity, the OLS estimates reported in AEFLM2016 are almost certainly too large. Interestingly, this bias appears to be driven by the implicit OLS weights on ATT and ATU, which were the focus of this paper.\footnote{I also repeat two further exercises from section (ref)(ref)\@. First, after I reestimate the model in ((ref)) using WLS, with weights of 1 for treated and $\frac{1}{k}$ for untreated units, I demonstrate in Figures (ref) to (ref) that these estimates become more positive as $k$ increases. As before, larger values of $k$ translate into larger weights on $\widehat{\mathrm{ATU}}$, which is now greater than $\widehat{\mathrm{ATT}}$\@. Second, when I use the estimates of ATT and ATU in Table (ref) to recover the hypothetical OLS weights, I obtain 22.8% as the mean of $\hat{w}_{ATT}$. This is reasonably similar to the proportion of untreated units, 12.5%.}

Conclusion

This paper proposed a new interpretation of the OLS estimand for the effect of a binary treatment in the standard linear model with additive effects. According to the main result of this paper, the OLS estimand is a convex combination of two parameters, which under certain conditions are equivalent to the average treatment effects on the treated (ATT) and untreated (ATU)\@. Surprisingly, the weights on these parameters are inversely related to the proportion of observations in each group, which can lead to substantial biases when interpreting the OLS estimand as ATE or ATT\@.

One lesson from this result is that it might be preferable, as suggested by a body of work in econometrics, to use any of the standard estimators of average treatment effects under ignorability, such as regression adjustment, weighting, matching, and various combinations of these approaches. Empirical researchers with a preference for OLS might instead want to use the diagnostic tools that this paper also provided. These diagnostics, which are implemented in the {hettreatreg} package in R and Stata, are applicable whenever the researcher is: (i) studying the effects of a binary treatment, (ii) using OLS, and (iii) unwilling to maintain that ATT is exactly equal to ATU\@. In an important special case, these diagnostics only require the knowledge of the proportion of treated units.