EconBase
← Back to paper

Two-Way Fixed Effects and Differences-in-Differences with Heterogeneous Treatment Effects: A Survey

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

103,480 characters · 19 sections · 255 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Two-Way Fixed Effects and Differences-in-Differences with Heterogeneous Treatment Effects: A Survey

abstractLinear regressions with period and group fixed effects are widely used to estimate policies' effects: 26 of the 100 most cited papers published by the American Economic Review from 2015 to 2019 estimate such regressions. It has recently been shown that those regressions may produce misleading estimates, if the policy's effect is heterogeneous between groups or over time, as is often the case. This survey reviews a fast-growing literature that documents this issue, and that proposes alternative estimators robust to heterogeneous effects. We use those alternative estimators to revisit wolfers2006did. Keywords: two-way fixed effects regressions, differences-in-differences, parallel trends, heterogeneous treatment effects, panel data, repeated-cross section data, policy evaluation.

Introduction

A popular method to estimate the effect of a policy, or treatment, on an outcome is to compare over time groups experiencing different evolutions of their exposure to treatment. In practice, this idea is implemented by regressing $Y_{g,t}$, the outcome in group $g$ and at period $t$, on group fixed effects, period fixed effects, and $D_{g,t}$, the treatment of group $g$ at period $t$. For instance, to measure the effect of the minimum wage on employment in the US, researchers have often regressed employment in county $g$ and year $t$ on county fixed effects, year fixed effects, and the minimum wage in county $g$ and year $t$.

Such two-way fixed effects (TWFE) regressions are probably the most-commonly used technique in economics to measure the effect of a treatment on an outcome. de2020difference conducted a survey of the 20 papers with the most Google Scholar citations published by the American Economic Review in 2015, and of the similarly selected papers in 2016, 2017, 2018, and 2019. Of those 100 papers, 26 have estimated at least one TWFE regression to estimate the effect of a treatment on an outcome. TWFE regressions are also very commonly used in political science, sociology, and environmental sciences.

Researchers have long thought that TWFE estimators are equivalent to differences-in-differences (DID) estimators. With two groups and two periods, a DID estimator compares the outcome evolution from period $1$ to $2$ between a treatment group $s$ that switches from untreated to treated, and a control group $n$ that is untreated at both dates:

equation[equation omitted — 92 chars of source]

$\text{DID}$ relies on a parallel trends assumption: in the absence of the treatment, both groups would have experienced the same outcome evolution. Specifically, for every $g\in \{s,n\}$ and $t\in \{1,2\}$, let $Y_{g,t}(0)$ and $Y_{g,t}(1)$ denote the potential outcomes in group $g$ at period $t$ without and with the treatment, respectively.\footnote{Implicitly, this notation rules out dynamic treatment effects, and assumes that groups' potential outcomes only depend on their current treatment, not on their past treatments. This restriction is not of essence to derive Equation (ref) below, but it is of essence for some of the other results we cover, as noted later in the paper. We relax it in Section 3.2.} Parallel trends requires that the expected evolution of the untreated outcome be the same in both groups: $$E\left[Y_{s,2}(0)-Y_{s,1}(0)\right]=E\left[Y_{n,2}(0)-Y_{n,1}(0)\right].$$ Under that assumption, $\text{DID}$ is unbiased for the average treatment effect (ATE) in group $s$ at period $2$ (see, e.g., Abadie05):

align[align omitted — 359 chars of source]

where the last equality follows from the parallel trends assumption. Parallel trends is partly testable, by comparing the outcome trends of groups $s$ and $n$, before group $s$ received the treatment. In practice, such pre-trends tests sometimes fail, but other times they indicate that the two groups were indeed on parallel paths before $s$ got treated.\footnote{Pre-trends tests come with caveats unveiled by a recent literature, see kahn2020promise, bilinski2018nothing, and roth2019pre. Similarly, recent papers have proposed relaxations of the parallel trends assumption manski2018right,rambachan2019honest,freyaldenhoven2019pre. Though we allude to it in Section (ref), this literature is mostly beyond the scope of this survey. See Roth2022 for a review.}

Motivated by the fact that in the two-groups and two-periods design described above, $\text{DID}$ is equal to the treatment coefficient in a TWFE regression, researchers have also estimated TWFE regressions in more complicated designs with many groups and periods, variation in treatment timing, treatments switching on and off, and/or non-binary treatments. Recent research has shown that in those more complicated designs, TWFE estimators are unbiased for an ATE if parallel trends holds, and if another assumption is satisfied: the treatment effect should be constant, between groups and over time. Unlike parallel trends, this assumption is unlikely to hold, even approximately, in most of the applications where TWFE regressions have been used. For instance, the effect of the minimum wage on employment is likely to differ in counties with highly educated workers, and in counties with less educated workers.

The realization that one of the most commonly used empirical methods in social science relies on an often-implausible assumption has spurred a flurry of methodological papers diagnosing the seriousness of the issue, and proposing alternative estimators. This review aims to provide an overview of this recent literature, which has developed in such a quick and dynamic manner that some practitioners may have gotten lost in the whirlwind of new working papers. We start by giving an overview of the papers that have identified TWFE's regressions lack of robustness to heterogeneous treatment effects, and that have proposed diagnostic tools practitioners may use to assess the seriousness of this issue. We then give an overview of the papers that have proposed alternative estimators robust to heterogeneous treatment effects. Finally, we revisit wolfers2006did, a famous TWFE application, in light of the recent literature discussed in this survey. As a word of caution, note that this literature is very recent, so several of the papers we review are still working papers, which have not been through the peer-review process yet.

Table (ref) in the conclusion summarizes the heterogeneity-robust estimators available to applied researchers, depending on their research design. When available, the Stata and R commands implementing the diagnostics tools and alternative estimators discussed in this review are referenced, and the basic syntax of the Stata command is provided. We refer the reader to the commands' help files for further details on their syntax. Finally, the Stata code for our re-analysis of wolfers2006did, where several of the estimators discussed in this survey are computed, is available at:

\url{https://drive.google.com/file/d/156Fu73avBvvV_H64wePm7eW04V0jEG3K/view?usp=sharing}.

TWFE regressions with heterogeneous treatment effects

\setcounter{equation}{0}

TWFE regressions may not identify a convex combination of treatment effects

We consider a panel of $G$ groups observed at $T$ periods, respectively indexed by the placeholders $g$ and $t$, which can refer to any group or time period. Typically, groups are geographical entities gathering many observations, but a group could also just be a single individual or firm. Let $\widehat{\beta}_{fe}$ denote the coefficient of $D_{g,t}$, the treatment in group $g$ at period $t$, in an OLS regression of $Y_{g,t}$, the outcome of group $g$ at period $t$, on group fixed effects, period fixed effects, and $D_{g,t}$:

equation[equation omitted — 121 chars of source]

where $\epsilon_{g,t}$ denotes the regression residual. We assume that the regression is unweighted, but it is sometimes weighted by $N_{g,t}$, the population of group $g$ at period $t$. The results discussed below also apply to this weighted regression, see dcDH2020.\footnote{ The regression could also be estimated using more disaggregated outcome data. For instance, groups may be US counties, and one may estimate the regression using individual-level outcome measures, assigning group membership based on county of residence. This disaggregated regression is equivalent to the aggregated regression in (ref), provided $Y_{g,t}$ is defined as the average outcome of individuals in cell $(g,t)$, and the aggregated regression is weighted by the number of individuals in cell $(g,t)$. Accordingly, the results below also apply to disaggregated regressions, see dcDH2020.}

dcDH2020 show that under a parallel trends assumption on the potential outcome without treatment $Y_{g,t}(0)$,

equation[equation omitted — 130 chars of source]

If the treatment is binary, $TE_{g,t}=Y_{g,t}(1)-Y_{g,t}(0)$, the ATE in group $g$ at time $t$. If the treatment is discrete or continuous, $TE_{g,t}=(Y_{g,t}(D_{g,t})-Y_{g,t}(0))/D_{g,t}$, the effect of moving the treatment from $0$ to $D_{g,t}$ scaled by $D_{g,t}$.\footnote{dcDH2020 derive Equation (ref) assuming that groups' potential outcomes only depend on their current treatment, not on their past treatments. With dynamic effects, Equation (ref) still holds if the treatment is binary and staggered, except that some of the $TE_{g,t}$s become effects of having been treated for more than one period.} The $W_{g,t}$ are weights summing to 1, that are proportional to and of the same sign as

equation[equation omitted — 69 chars of source]

where $D_{g,.}$ is the average treatment of group $g$ across periods, $D_{.,t}$ is the average treatment at period $t$ across groups, and $D_{.,.}$ is the average treatment across groups and periods.

Equations (ref) and (ref) have two important consequences. First, $W_{g,t}$ is in general not equal to one divided by the number of treated $(g,t)$ cells, so $\widehat{\beta}_{fe}$ may be biased for the average treatment effect across those cells, the ATT. A special case where $W_{g,t}$ is equal to one divided by the number of treated $(g,t)$ cells, and where $\widehat{\beta}_{fe}$ is therefore unbiased for the ATT is when (i) the design is staggered, meaning that groups' treatment can only increase over time and can change at most once;\footnote{Together, (i) and (ii) imply that groups can only switch from untreated to treated, and may do so at different points in time. This is probably the definition of a staggered design many people have in mind. (i) extends the definition of a staggered design to non-binary treatments.} (ii) the treatment is binary; and (iii) there is no variation in treatment timing: all treated groups start receiving the treatment at the same date. However, conditions (i)-(iii) are seldom met in practice. $\widehat{\beta}_{fe}$ can also be unbiased for the ATT if one is ready to make more assumptions than just parallel trends. For instance, if one is also ready to assume that $D_{g,t}-D_{g,.}-D_{.,t}+D_{.,.}$ is uncorrelated with $TE_{g,t}$, the treatment effects that are up- and down-weighted by $\widehat{\beta}_{fe}$ do not systematically differ, and one can then show that $\widehat{\beta}_{fe}$ is unbiased for the ATT dcDH2020.\footnote{A special case of this “no-correlation” condition is if the treatment effect is constant, i.e. $TE_{g,t}=\delta$ for all $(g,t)$. Then, it directly follows from Equation (ref) that $E\left[\widehat{\beta}_{fe}\right]=\delta$. However, constant effect is most often an implausible assumption.} Unfortunately, this no-correlation condition is often implausible. To see this, note that $D_{g,t}-D_{g,.}-D_{.,t}+D_{.,.}$ is decreasing in $D_{g,.}$, meaning that $\widehat{\beta}_{fe}$ downweights the treatment effect of groups with the highest average treatment from period $1$ to $T$. However, groups with the largest and lowest average treatment may have systematically different treatment effects. Similarly, $D_{g,t}-D_{g,.}-D_{.,t}+D_{.,.}$ is decreasing in $D_{.,t}$, and the treatment effects at time periods with the highest average treatment may also systematically differ from the treatment effects at time periods where the average treatment is lower. In staggered adoption designs, $D_{.,t}$ is increasing in $t$ so the weights are decreasing in $t$. If the treatment effect is also monotonically increasing or decreasing in $t$, this no-correlation condition will fail. This no-correlation condition is partly testable, if one observes a proxy variable $P_{g,t}$ that is likely to be correlated with $TE_{g,t}$. Then, one can just test if $D_{g,t}-D_{g,.}-D_{.,t}+D_{.,.}$ is correlated with $P_{g,t}$.

Second, and perhaps more worryingly, Equation (ref) implies that some of the weights $W_{g,t}$ may be negative. This means that in the minimum wage example, $\widehat{\beta}_{fe}$ could be estimating something like $3$ times the effect of the minimum wage on employment in Santa Clara county, minus $2$ times the effect in Wayne county. Then, if raising the minimum wage by one dollar decreases employment by 5% in Santa Clara county and by 20% in Wayne county, one would have $E\left[\widehat{\beta}_{fe}\right]=3\times -0.05-(2\times -0.2)=0.25$. $E\left[\widehat{\beta}_{fe}\right]$ would be positive, while the minimum wage's effect on employment is negative both in Santa Clara and in Wayne county. This example shows that $\widehat{\beta}_{fe}$ may not satisfy the “no-sign reversal property”: $E\left[\widehat{\beta}_{fe}\right]$ could for instance be positive, even if the treatment effect is strictly negative in every $(g,t)$. This phenomenon can only arise when some of the weights $W_{g,t}$ are negative: when all those weights are positive, $\widehat{\beta}_{fe}$ does satisfy the no-sign reversal property. Note that despite its intuitive appeal and its popularity among applied researchers, the no-sign reversal property is not grounded in statistical decision theory, unlike other commonly-used criteria to discriminate estimators such as the mean-squared error. Still, it is connected to the economic concept of Pareto efficiency. If an estimator satisfies “no-sign-reversal”, the estimand attached to it can only be positive if the treatment is not Pareto-dominated by the absence of treatment, meaning that not everybody is hurt by the treatment. Conversely, the estimand can only be negative if the treatment does not Pareto-dominate the absence of treatment. On the other hand, if an estimator does not satisfy “no-sign-reversal”, the estimand attached to it could for instance be positive, even if the treatment is Pareto-dominated.

Inasmuch as “no-sign-reversal” is a desirable property, it becomes interesting to understand when $\widehat{\beta}_{fe}$ may satisfy it. Equation (ref) shows that with a binary treatment, the weights attached to $\widehat{\beta}_{fe}$ could all be positive. With a binary treatment, all the $(g,t)$s entering the summation in (ref) must have $D_{g,t}=1$, so for a weight $W_{g,t}$ to be strictly negative, one must have $1+D_{.,.}< D_{g,.}+D_{.,t}.$ This cannot happen if $D_{g,.}+D_{.,t}\leq 1$ for every $(g,t)$. Accordingly, all the weights are likely to be positive when there is no group that is treated most of the time, and no time periods where most groups are treated. In staggered designs, this has led jakiela2021simple to propose to drop the last periods of the data, those when $D_{.,t}$ is the highest, to mitigate or eliminate the negative weights. One could also drop the always-treated groups, if there are any.

On the other hand, Equation (ref) shows that with a non-binary treatment, it becomes more likely that some of the weights $W_{g,t}$ are negative. gentzkow2011 study the effect of the number of newspapers in county $g$ and year $t$ on turnout in presidential elections. Assume that in year $t$, county $g$ has $1$ newspaper ($D_{g,t}=1$), which is below its average number of newspapers across years, equal, say, to $2$ ($D_{g,.}=2$). At the same time, the average number of newspapers across counties in year $t$ is equal to 2 ($D_{.,t}=2$), which is above the average number of newspapers across all counties and years, equal, say, to $1$ ($D_{.,.}=1$). Then, it follows from (ref) that the weight assigned to the effect of newspapers in county $g$ and year $t$ is strictly negative. More generally, a necessary condition to have that all weights are positive is that in every period where the population's treatment is higher than its average across periods ($D_{.,t}\geq D_{.,.}$), the treatment of each treated group must also be larger than its average across periods ($D_{g,t}\geq D_{g,.}$ for all $g$s such that $D_{g,t}\ne 0$). This condition is likely to often fail.

The twowayfeweights Stata twowayfeweightsStata and R twowayfeweightsR commands compute the weights $W_{g,t}$ in (ref). The basic syntax of the Stata command is:

twowayfeweights outcome groupid timeid treatment, type(feTR)

A decomposition similar to (ref) can be obtained for TWFE regressions with control variables, and for $\widehat{\beta}_{fd}$, the treatment's coefficient in a regression of the outcome's first difference on the treatment's first difference and period fixed effects. dcDH2020 also derive decompositions similar to (ref), for $\widehat{\beta}_{fe}$ and $\widehat{\beta}_{fd}$, under common trends and under the assumption that the treatment effect does not change over time. The weights in all those decompositions are also computed by the twowayfeweights Stata and R commands.

dcDH2020 use the twowayfeweights Stata command to revisit gentzkow2011. The authors regress the change in turnout in county $g$ between two elections on the change of the county's number of newspapers and state-year fixed effects. They find that $\widehat{\beta}_{fd}=0.0026$ (s.e. $=0.0009$): one more newspaper increases turnout by 0.26 percentage points. Using the twowayfeweights Stata package, dcDH2020 find that under parallel trends, $\widehat{\beta}_{fd}$ estimates a weighted sum of the effects of newspapers on turnout in 10,077 county$\times$election cells, where 5,472 effects are weighted positively while 4,605 are weighted negatively, and where negative weights sum to -1.43. Accordingly, $\widehat{\beta}_{fd}$ is far from estimating a convex combination of effects. The weights are negatively correlated with the election year: $\widehat{\beta}_{fd}$ is more likely to upweight newspapers' effects in early elections, and to downweight or weight negatively newspapers' effects in late elections. This may lead $\widehat{\beta}_{fd}$ to be biased if newspapers' effects change over time. Similar results apply to $\widehat{\beta}_{fe}$: more than half of the weights attached to that coefficient are negative, and negative weights sum to -0.53.

The decomposition in (ref) is the main result in dcDH2020. Related results have appeared earlier in Theorems S1 and S2 of the Supplementary Material of deChaisemartin15c. borusyak2016 consider the case with a binary and staggered treatment. In their Lemma 1 and Proposition 1, they assume that the treatment effect varies with the duration elapsed since one has started receiving the treatment but does not vary across groups and over time. Then, they show that $\widehat{\beta}_{fe}$ estimates a weighted sum of effects, that may assign negative weights to long-run treatment effects. Their Appendix C also contains another result related to that in Equation (ref).\footnote{Prior to that, chernozhukov2013average had shown that one-way FE regressions may be biased for the average treatment effect, though unlike TWFE regressions they always estimate a convex combination of effects.}

The origin of the problem: “forbidden comparisons”

Forbidden comparisons when the treatment is binary and the design is staggered

goodman2021difference shows that when the treatment is binary and the design is staggered, meaning that groups can switch in but not out of treatment, we have

equation[equation omitted — 105 chars of source]

where $DID_{g,g',t,t'}$ is a DID comparing the outcome evolution of two groups $g$ and $g'$ from a pre period $t$ to a post period $t'$, and where $v_{g,g',t,t'}$ are non-negative weights summing to one, with $v_{g,g',t,t'}>0$ if and only if $g$ switches treatment between $t$ and $t'$ while $g'$ does not.\footnote{goodman2021difference actually decomposes $\widehat{\beta}_{fe}$ as a weighted average of DIDs between cohorts of groups becoming treated at the same date, and between periods of time where their treatment remains constant. One can then further decompose his decomposition, as we do here.} Some of the $DID_{g,g',t,t'}$s in Equation (ref) compare a group switching treatment from $t$ to $t'$ to a group untreated at both dates, while other $DID_{g,g',t,t'}$s compare a switching group to a group treated at both dates. The negative weights in (ref) originate from this second type of DIDs.

To see that, let us consider a simple example, first introduced by borusyak2016,\footnote{borusyak2016 have also coined the “forbidden comparisons” expression we borrow here.} with two groups and three periods. Group $e$, the early-treated group, is untreated at period 1 and treated at periods 2 and 3. Group $\ell$, the late-treated group, is untreated at periods 1 and 2 and treated at period 3. In this example, Equation (ref) reduces to

equation[equation omitted — 113 chars of source]

with

align*[align* omitted — 173 chars of source]

$\text{DID}_{e,\ell,1,2}$ compares the period-1-to-2 outcome evolution of group $e$, that switches from untreated to treated from period $1$ to $2$, to the outcome evolution of group $\ell$ that is untreated at both periods. $\text{DID}_{e,\ell,1,2}$ is similar to the $\text{DID}$ estimator in Equation (ref), and under parallel trends it is unbiased for the treatment effect in group $e$ at period $2$:

align[align omitted — 93 chars of source]

$\text{DID}_{\ell,e,2,3}$, on the other hand, compares the period-2-to-3 outcome evolution of group $\ell$, that switches from untreated to treated from period $2$ to $3$, to the outcome evolution of group $e$ that is treated at both dates. At both periods, $e$'s outcome is its treated potential outcome, which is equal to the sum of its untreated outcome and its treatment effect. Accordingly, $$Y_{e,3}-Y_{e,2}=Y_{e,3}(0)+TE_{e,3}-(Y_{e,2}(0)+TE_{e,2}).$$ On the other hand, group $\ell$ is only treated at period $3$, so $$Y_{\ell,3}-Y_{\ell,2}=Y_{\ell,3}(0)+TE_{\ell,3}-Y_{\ell,2}(0).$$ Taking the expectation of the difference between the two previous equations,

align[align omitted — 114 chars of source]

where $E\left[Y_{e,3}(0)-Y_{e,2}(0)\right]$ and $E\left[Y_{\ell,3}(0)-Y_{\ell,2}(0)\right]$ cancel out under the parallel trends assumption. Finally, it follows from Equations (ref), (ref), and (ref) that

align[align omitted — 125 chars of source]

In this simple example, Equation (ref) reduces to (ref). The right-hand side of Equation (ref) is a weighted sum of three ATEs where one ATE receives a negative weight. As the previous derivation shows, this negative weight comes from the fact $\widehat{\beta}_{fe}$ leverages $\text{DID}_{\ell,e,2,3}$, a DID comparing a group switching from untreated to treated to a group treated at both periods.

To make things more concrete, Figure (ref) below shows the actual and counterfactual outcome evolution, in a numerical example with three periods and an early and a late treated group. All treatment effects are positive: the actual outcomes, on the solid lines, are always above the counterfactual outcomes on the dashed lines. However, $\widehat{\beta}_{fe}$ is negative. $\widehat{\beta}_{fe}$ is the simple average of the DID comparing the early- to the late-treated group from period one to two, which is positive, and of the DID comparing the late- to the early-treated group from period two to three, which is negative, and larger in absolute value than the first DID. The reason why the second DID is negative is that the treatment effect of the early-treated group increases substantially from period two to three, so this group's outcome increases more than that of the late-treated group.

figure[figure omitted — 204 chars of source]

If one is ready to assume that the treatment effect does not change over time, $TE_{e,3}=TE_{e,2}$, and (ref) simplifies to

align[align omitted — 97 chars of source]

Then, the negative weight in (ref) disappears, and $\widehat{\beta}_{fe}$ estimates a weighted average of treatment effects. This extends beyond this simple example: Theorem S2 of the Web Appendix of dcDH2020 and Equation (16) of goodman2021difference show that in staggered adoption designs with a binary treatment, $\widehat{\beta}_{fe}$ estimates a convex combination of effects, if the treatment effect does not change over time but may still vary across groups. This conclusion, however, no longer holds if the treatment is not binary or the design is not staggered. Moreover, assuming constant treatment effects over time is often implausible as this rules out both dynamic treatment effects and calendar time effects.

The decomposition in Equation (ref) is key to understand why $\widehat{\beta}_{fe}$ may not identify a convex combination of treatment effects. On the other hand, it cannot be used to assess if $\widehat{\beta}_{fe}$ does indeed estimate a convex combination of effects in a given application. Consider an example similar to that above, but with a third group $n$ that remains untreated from period $1$ to $3$. In this second example, the decomposition in (ref) now indicates that $\widehat{\beta}_{fe}$ assigns a weight equal to 1/6 to DIDs comparing a switcher to a group treated at both periods. On the other hand, all the weights in (ref) are positive in this second example. This phenomenon can also arise in real data sets. In the data of stevenson2006bargaining used by goodman2021difference in his empirical application, if one restricts the sample to states that are not always treated and to the first ten years of the panel, all the weights in (ref) are positive, but the sum of the weights in (ref) on DIDs comparing a switcher to a group treated at both periods is equal to $0.06$. Beyond these examples, one can show that having DIDs comparing a switcher to a group treated at both periods in (ref) is necessary but not sufficient to have negative weights in (ref). Similarly, the sum of the weights on DIDs comparing a switcher to a group treated at both periods in (ref) is always larger than the absolute value of the sum of the negative weights in (ref). The reason why Equation (ref) “overestimates” the negative weights in (ref) is that as soon as there are three distinct treatment dates, there is not a unique way of decomposing $\widehat{\beta}_{fe}$ as a weighted average of DIDs, and there exists other decompositions than Equation (ref) putting less weight on DIDs using a group treated at both periods as the control group.\footnote{To see that, let $t_0<t_1<t_2$ be three dates, let $e$ be an early-treated group becoming treated at $t_1$, let $\ell$ be a late-treated group becoming treated at $t_2$, and let $n$ be a group untreated yet at $t_2$. Let $\underline{v}=\min(v_{\ell,e,t_1,t_2},v_{e,n,t_0,t_2})>0$. One has

equation[equation omitted — 151 chars of source]

Then, it follows from Equation (ref) that

align[align omitted — 336 chars of source]

Plugging Equation (ref) into Equation (ref) will yield a different decomposition of $\widehat{\beta}_{fe}$ as a weighted average of DIDs. But the weight on DIDs using a group treated at both periods as the control group is equal to $v_{\ell,e,t_1,t_2}$ in the left-hand-side of Equation (ref), and to $(v_{\ell,e,t_1,t_2}-\underline{v})$ in its right-hand side. Accordingly, this new decomposition puts strictly less weight than Equation (ref) on DIDs using a group treated at both periods as the control group.}

The bacondecomp Stata bacondecompStata and R bacondecompR commands compute the $\text{DID}_{g,g',t,t'}$s entering in (ref), the weights assigned to them, as well as the sum of the weights on $\text{DID}_{g,g',t,t'}$s using a group treated at both periods as the control group. The basic syntax of the bacondecomp Stata command is:

bacondecomp outcome treatment, ddetail

“Forbidden comparisons” when the design is not staggered or treatment is not binary

When the treatment is not staggered or when it is not binary, $\widehat{\beta}_{fe}$ may leverage another type of comparison: it may compare the outcome evolution of a group $m$ whose treatment increases more to the outcome evolution of a group $\ell$ whose treatment increases less. In fact, with two groups $m$ and $\ell$ and two periods, one can show that

equation[equation omitted — 166 chars of source]

where the right hand side of the previous display is the Wald-DID estimator studied by deChaisemartin15b. The Wald-DID compares the outcome evolution of groups $m$ and $\ell$, and scales that comparison by the differential evolution of $m$'s and $\ell$'s treatments. deChaisemartin15b show that the Wald-DID may not estimate a convex combination of effects, unless the treatment effect is constant over time and is the same in groups $m$ and $\ell$. This second requirement was not present in the binary and staggered case. In that case, we have seen before that if the treatment effect is constant over time, $\widehat{\beta}_{fe}$ estimates a convex combination of effects, even if the treatment effect varies between groups.

To see that with a non-binary or non-staggered treatment $\widehat{\beta}_{fe}$ may not estimate a convex combination of effects even if the treatment effect is constant over time, let us consider a simple example. Assume that group $m$ goes from $0$ to $2$ units of treatment from period $1$ to $2$, while group $\ell$ goes from $0$ to $1$ unit. Then, the denominator of the Wald-DID is equal to $2-0-(1-0)=1$, so

equation*[equation* omitted — 89 chars of source]

To simplify, let us also assume that in both groups, potential outcomes are linear in the number of treatment units, with slopes that are constant over time but may differ for groups $m$ and $\ell$:

align*[align* omitted — 90 chars of source]

Then, under parallel trends,

align*[align* omitted — 360 chars of source]

a weighted sum of $m$ and $\ell$'s treatment effects, where group $\ell$'s effect is weighted negatively. Intuitively, group $\ell$ is also treated at period two, and $\widehat{\beta}_{fe}$, which uses $\ell$ as a control group, subtracts its treatment effect out. This example also shows that $\widehat{\beta}_{fe}$ may fail to identify a convex combination of effects, even without variation in treatment timing: here, both $m$ and $\ell$ start getting treated at period 2.

To make things more concrete, Figure (ref) below shows the actual and counterfactual outcome evolution, in a numerical example with two periods, a group whose treatment increases more, from $0$ to $2$ units, and a group whose treatment increases less, from $0$ to $1$ unit. All treatment effects are positive: the actual outcomes, on the solid lines, are always above the counterfactual outcomes on the dashed lines. However, $\widehat{\beta}_{fe}$, which is equal to the DID comparing the more- and the less-treated groups from period one to two, is negative. The reason why this DID is negative is that the treatment effect, per treatment unit, of the less-treated group is more than twice larger than the treatment effect of the more-treated group. Accordingly, the outcome of the less-treated group increases more, despite the fact that this group receives a twice smaller treatment dose in period 2.

figure[figure omitted — 200 chars of source]

Decomposition results for other TWFE regression coefficients

Dynamic TWFE regressions

In staggered designs with a binary treatment, abraham2018 consider event-study regressions:

align[align omitted — 175 chars of source]

where $F_g$ is the first period at which group $g$ is treated. In words, the outcome is regressed on group and period fixed effects, and relative-time indicators $1\{F_g=t-\ell\}$ equal to 1 if group $g$ started receiving the treatment $\ell$ periods ago. For $\ell\geq 0$, $\widehat{\beta}_\ell$ is supposed to estimate the cumulative effect of $\ell+1$ treatment periods. For $\ell\leq -2$, $\widehat{\beta}_\ell$ is supposed to be a placebo coefficient testing the parallel trends assumption, by comparing the outcome trends of groups that will and will not start receiving the treatment in $|\ell|$ periods. Researchers have sometimes estimated a variant of this regression, where the first and last indicators $1\{F_g=t+K\}$ and $1\{F_g=t-L\}$ are respectively replaced by an indicator for being at least $K$ periods away from adoption ($1\{F_g\geq t+K\}$) and an indicator for having adopted at least $L$ periods ago ($1\{F_g\leq t-L\}$). Such endpoint binning is for instance recommended by Schmidheiny20: without it, the regression implicitly assumes that the treatment no longer has any effect after $L$ periods. Instead, with endpoint binning the regression assumes that that the treatment effect is constant after $L$ periods, a more plausible assumption.

abraham2018 show that under parallel trends, for $\ell\geq 0$,

equation[equation omitted — 166 chars of source]

where $TE_g(\ell)$ is the cumulative effect of $\ell+1$ treatment periods in group $g$, and $w_{g,\ell}$ and $w_{g,\ell'}$ are weights such that $\sum_{g}w_{g,\ell}=1$ and $\sum_{g}w_{g,\ell'}=0$ for every $\ell'$.\footnote{Equation (ref) follows from Proposition 3 in abraham2018, assuming no binning and that the treatment does not have an effect after $L+1$ periods of exposure. A slight difference is that the decomposition in abraham2018 gathers groups that started receiving the treatment at the same period into cohorts. Their decomposition can then be further decomposed, as we do here.} The first summation in the right-hand side of Equation (ref) is a weighted sum across groups of the cumulative effect of $\ell+1$ treatment periods, with weights summing to 1 but that may be negative. This first summation resembles that in the decomposition of the “static” TWFE coefficient in (ref), and it implies that $\widehat{\beta}_\ell$ may be biased if the cumulative effect of $\ell+1$ treatment periods varies across groups. The second summation is a weighted sum, across $\ell'\ne \ell$ and groups, of the cumulative effect of $\ell'+1$ treatment periods in group $g$, with weights summing to 0. This second summation was not present in the decomposition of the static TWFE coefficient. Importantly, its presence implies that $\widehat{\beta}_\ell$, which is supposed to estimate the cumulative effect of $\ell+1$ treatment periods, may in fact be contaminated by the effects of $\ell'+1$ treatment periods. As $\sum_{g}w_{g,\ell'}=0$ for every $\ell'$, this second summation disappears if $TE_g(\ell')$ does not vary across groups, but it is often implausible that the treatment effect does not vary across groups.

For $\ell\leq -2$, and without assuming parallel trends, abraham2018 show that $\widehat{\beta}_\ell$ estimates the sum of two terms. As intended, the first term measures deviations from parallel trends between groups that will and will not start receiving the treatment in $|\ell|$ periods. But the second term is similar to the second summation in the right-hand side of Equation (ref): a weighted sum, across $\ell'\geq 0$ and groups, of the cumulative effect of $\ell'+1$ treatment periods in group $g$, with weights summing to zero. Due to the presence of this second term, the expectation of $\widehat{\beta}_\ell$ may differ from zero even if parallel trends holds, and it may be equal to zero even if parallel trends fails. Thus, an important consequence of the results in abraham2018 is that in the presence of heterogeneous treatment effects, (ref) cannot be used to test for parallel trends.

The eventstudyweights Stata command EVENTSTUDYWEIGHTS computes the weights attached to event-study regressions. Its basic syntax is:

eventstudyweights \{rel_time_list\}, absorb(i.groupid i.timeid) \\ cohort(first_treatment) rel_time(ry),

where rel_time_list is the list of relative-time indicators $1\{F_g=t-\ell\}$ included in (ref), first_treatment is a variable equal to the period when group $g$ got treated for the first time, and ry is a variable equal to timeid minus first_treatment, the number of periods elapsed since group $g$ started receiving the treatment.

Event-study regressions can only be used in staggered designs with a binary treatment. In more complicated designs where the treatment is not binary or a group's treatment can increase or decrease multiple times, some researchers have estimated TWFE regressions of the outcome on the treatment and its first $K$ lags, the so-called distributed-lag regression. Other researchers have estimated a panel-data version of the local-projection method proposed by jorda2005estimation for time-series data: $Y_{g,t+\ell}$ is regressed on group and period FEs and $D_{g,t}$, for $\ell \in \{0,...,K\}$. de2020difference show that those regressions suffer from similar issues as the event-study regression: under parallel trends, the distributed-lag and local-projection regressions may produce biased estimates of the treatment's instantaneous and dynamic effects, if effects are heterogeneous across groups and over time. In particular, they do not satisfy the no-sign reversal property: one could have that the treatment's instantaneous and dynamic effects are positive in every $(g,t)$ cell, but the expectations of those regression coefficients are negative. de2020difference also show that the panel-data version of the local-projection method may yield biased estimates even if effects are homogeneous.

TWFE regressions with more than one treatment

Another case of interest is TWFE regressions with several treatments. For instance, to estimate separately the effect of medical and recreational marijuana laws on consumption, one may regress marijuana consumption in state $g$ and year $t$ on state and year fixed effects, on whether state $g$ has a recreational marijuana law in year $t$, and on whether state $g$ has a medical law in year $t$. de2020two show that in those regressions, the coefficient on a given treatment identifies a weighted sum of that treatment's effect across $(g,t)$s, with weights summing to $1$ but that may be negative, plus weighted sums of the effects of the other treatments in the regression, with weights summing to 0. In the example above, the coefficient on recreational laws may be contaminated by the effect of medical laws. The weights attached to TWFE regressions with several treatments are also computed by the twowayfeweights Stata and R commands.

Alternative heterogeneity-robust DID estimators

\setcounter{equation}{0}

In this section, we review several recently-proposed alternatives to TWFE regressions. We restrict our attention to estimators relying on parallel trends assumptions, like TWFE regressions, but that do not restrict treatment effect heterogeneity between groups and over time, unlike TWFE regressions. This excludes papers that have assumed randomized treatment timing athey2021design,roth2021efficient or sequential treatment randomization bojinov2020panel, rather than parallel trends. Intuitively, all the estimators below carefully choose valid control groups, to avoid making the “forbidden comparisons” that render TWFE estimators non-robust to heterogeneous treatment effects. We start by reviewing estimators ruling out dynamic effects, i.e. that assume that a group's current outcome only depends on its current treatment, before reviewing estimators that allow dynamic effects. In complicated designs, say with a continuous treatment that changes often, allowing for dynamic effects comes with a number of costs: it may result in imprecise estimators, and may complicate the interpretation of the estimated effects. Then, one may want to carefully evaluate if past treatments are indeed likely to affect the current outcome.

Estimators ruling out dynamic effects

With a binary treatment, dcDH2020 propose to use the $\text{DID}_{\text{M}}$ estimator. With two time periods, $\text{DID}_{\text{M}}$ is merely a weighted average of

align*[align* omitted — 153 chars of source]

and of

align*[align* omitted — 153 chars of source]

where for all $(d_1,d_2)\in \{0,1\}^2$, $N_{d_1,d_2}$ denotes the number of groups such that $D_{g,1}=d_1$ and $D_{g,2}=d_2$.\footnote{Implicitly, this definition of $\text{DID}_+$ and $\text{DID}_-$ assumes that all groups have the same sizes. The $\text{DID}_{\text{M}}$ estimator can easily be extended to instances where groups have heterogeneous sizes, see dcDH2020.} $\text{DID}_+$ is a DID comparing the period-one-to-two outcome evolution of groups going from untreated to treated, the “switchers in”, and of groups untreated at both dates. It is similar to the $\text{DID}$ estimator in Equation (ref), and it is unbiased for the treatment effect of the switching-in groups at period $2$, under a parallel trends assumption on the untreated outcome $Y_{g,t}(0)$. $\text{DID}_-$ is a DID comparing the period-one-to-two outcome evolution of groups treated at both dates, and of groups going from treated to untreated, the “switchers out”. $\text{DID}_-$ is also similar to the $\text{DID}$ estimator in Equation (ref), switching “treatment” and “non-treatment”. Then, one can show that $\text{DID}_-$ is unbiased for the treatment effect of the switching-out groups at period $2$, under a parallel trends assumption on the treated outcome $Y_{g,t}(1)$.

The $\text{DID}_{\text{M}}$ estimator can easily be extended to applications with more than two time periods. For each pair of consecutive time periods, one can compute a $\text{DID}_{+,t}$ estimator comparing groups going from untreated to treated from $t-1$ to $t$ to groups untreated at both dates, and a $\text{DID}_{-,t}$ estimator comparing groups treated at $t-1$ and $t$ to groups going from treated to untreated from $t-1$ to $t$. Then, one averages the $\text{DID}_{+,t}$ and $\text{DID}_{-,t}$ estimators across $t$. dcDH2020 show that the resulting estimator is unbiased for the average treatment effect across all switching $(g,t)$ cells, namely cells such that $D_{g,t}\ne D_{g,t-1}$. They also propose placebo estimators to test the parallel trends assumptions underlying $\text{DID}_{\text{M}}$. The placebos compare the outcome trends of switchers and non-switchers, before the switchers switch.

With more than two time periods, the $\text{DID}_{\text{M}}$ estimator may be biased if the treatment has dynamic effects. For instance, to infer the counterfactual trend that groups going from untreated to treated from $t-1$ to $t$ would have experienced without that switch, $\text{DID}_{+,t}$ uses as controls all groups untreated at $t-1$ and $t$. However, some of those groups may have been treated, say, at $t-2$. If the treatment has dynamic effects, this past treatment may affect their period $t-1$-to-$t$ outcome evolution, thus making them potentially invalid controls. Note that if the treatment is binary and staggered, such situations cannot arise: groups untreated at $t-1$ and $t$ have been untreated all along. Accordingly, $\text{DID}_{\text{M}}$ is robust to dynamic effects in binary and staggered designs.

The $\text{DID}_{\text{M}}$ estimator can easily be extended to non-binary treatments taking a finite number of values. Then, it is a weighted average, across $d$ and $t$, of DIDs comparing the $t-1$ to $t$ outcome evolution of groups whose treatment goes from $d$ to some other value from $t-1$ to $t$, and of groups with a treatment equal to $d$ at both dates, normalized by the intensity of the treatment change experienced by the switchers. For instance, in gentzkow2011, a county going from $2$ to $4$ newspapers is compared to a county with $2$ newspapers at both dates. The multi-period DID estimator in imai2021use is related to the $\text{DID}_{\text{M}}$ estimator. It can be used with a binary treatment, to estimate the switchers-in's treatment effect.

The $\text{DID}_{\text{M}}$ estimator is computed by the did_multiplegt Stata did_multiplegtStata and R did_multiplegtR commands. The basic syntax of the Stata command is:

did_multiplegt outcome groupid timeid treatment

dcDH2020 compute the $\text{DID}_{\text{M}}$ estimator in the gentzkow2011 example mentioned above, that studies the effect of newspapers on turnout in US presidential elections. dcDH2020 find that $\text{DID}_{\text{M}}=0.0043$ (s.e. $=0.0014$), meaning that one more newspaper increases turnout by 0.43 percentage point. $\text{DID}_{\text{M}}$ is 66% larger than, and significantly different from, $\widehat{\beta}_{fd}$, the estimator reported by gentzkow2011.

chaisemartin2022continuous extend the $\text{DID}_{\text{M}}$ estimator to continuous treatments. To simplify, we present their estimators in the case with two time periods, though they readily extend to the case with more periods. chaisemartin2022continuous assume that from period one to two, the treatment of some units, hereafter referred to as the movers, changes. They also assume that the treatment of other units, hereafter referred to as the stayers, does not change. This assumption is likely to be met when the treatment is say, trade tariffs: tariffs' reforms rarely apply to all products, so it is likely that tariffs of at least some products stay constant over time. On the other hand, this assumption is unlikely to be met when the treatment is say, precipitations: geographical units never experience the exact same precipitations over two consecutive years.

Under the assumption that there are some stayers, the estimator proposed by chaisemartin2022continuous compares the outcome evolution of movers and stayers, with the same period-one treatment. With a continuous treatment, such comparisons can either be achieved by reweigthing stayers by propensity score weights, or by adjusting movers' outcome change using a nonparametric regression of the outcome change on the period-one treatment among the stayers. Under parallel trends assumptions, the corresponding estimands identify a weighted average of the effect, across all movers, of moving their treatment from its period-one to its period-two value, scaled by the difference between these two values. This effect is a weighted average of the slopes of movers' potential outcome function, between their period-one and period-two treatments.

The estimators in chaisemartin2022continuous can be extended to the case where there are no stayers, provided there are quasi-stayers, meaning units whose treatment barely changes from period one to two. Alternatively, one could also use the estimator proposed by graham2012identification, which compares the outcome evolution of movers and quasi stayers, but without conditioning on units' period-one treatment. Their estimator relies on a linear treatment effect assumption, unlike those in chaisemartin2022continuous. When there are no true stayers, both estimators require choosing a bandwidth, namely the lowest treatment change below which a unit can be considered as a quasi-stayer. Neither chaisemartin2022continuous or graham2012identification derive an “optimal” bandwidth, so for now bandwidth choice is left to the discretion of the researcher. If the data has at least three periods, one could also use the correlated-random-coefficient estimator proposed by chamberlain1992efficiency. While it allows for some treatment effect heterogeneity, that estimator relies on a linear treatment effect assumption, like the estimator in graham2012identification.

chaisemartin2022continuous show that after some relabelling, some of their estimators are equivalent or nearly equivalent to estimators that had been previously proposed by deChaisemartin15b, Abadie05, and callaway2018. This implies that their estimators can be computed, up to small tweaks, by the companion software for those papers. We refer the reader to chaisemartin2022continuous for a precise description of how their estimators can be computed using existing software.

Estimators allowing for dynamic effects when the treatment is binary and the design is staggered.

For any $t\in \{1,...,T\}$, let $\bm{0}_t$ (resp. $\bm{1}_t$) denote a vectors of $t$ zeros (resp. ones). With dynamic effects, group $g$'s outcome at time $t$ is allowed to depend on her past treatments. For any $(d_1,...,d_t)$, let $Y_{g,t}(d_1,...,d_t)$ denote group $g$'s potential outcome at period $t$ with treatments $(d_1,...,d_t)$ from period $1$ to $t$.\footnote{This notation implicitly rules out anticipation effects: the outcome cannot depend on a group's future treatment.} In particular, $Y_{g,t}(\bm{0}_t)$ is group $g$'s outcome without ever being treated from period $1$ to $t$. With dynamic effects, callaway2018 and abraham2018 have proposed to replace the parallel trends assumption on $Y_{g,t}(0)$ by a parallel trends assumption on $Y_{g,t}(\bm{0}_t)$: for all $g\ne g'$ and $t\geq 2$,

equation[equation omitted — 161 chars of source]

We now review the estimators proposed by callaway2018, abraham2018, and borusyak2020revisiting for binary and staggered treatments, under the parallel trends assumption in Equation (ref).

The estimators proposed by callaway2018

In a staggered adoption design, groups can be aggregated into cohorts that start receiving the treatment at the same period. For all $c$ and $t$, and for all $\ell \in \{0,...,t\}$ let $\overline{Y}_{c,t}$ denote the average outcome at period $t$ across groups belonging to cohort $c$, and let $\overline{Y}_{n,t}$ denote the average outcome at period $t$ across groups that remain untreated from period $1$ to $T$, hereafter referred to as the never-treated groups, assuming for now that such groups exist. callaway2018 define their parameters of interest as $$TE_{c,c+\ell}=E\left[\overline{Y}_{c,c+\ell}(\bm{0}_{c-1},\bm{1}_{\ell+1})-\overline{Y}_{c,c+\ell}(\bm{0}_{c+\ell})\right],$$ the average effect of having been treated for $\ell+1$ periods in the cohort that started receiving the treatment at period $c$, for every $c\in \{2,...,T\}$ and $\ell \geq 0$ such that $\ell+c\leq T$. To estimate, say, $TE_{c,c}$, callaway2018 propose $$\overline{\text{DID}}_{c,0}=\overline{Y}_{c,c}-\overline{Y}_{c,c-1}-\left(\overline{Y}_{n,c}-\overline{Y}_{n,c-1}\right),$$ a DID estimator comparing the period $c-1$-to-$c$ outcome evolution in cohort $c$ and in the never-treated groups $n$. $\overline{\text{DID}}_{c,0}$ is unbiased for $TE_{c,c}$:

align*[align* omitted — 608 chars of source]

where the last equality follows from Equation (ref). More generally, to estimate $TE_{c,c+\ell}$, callaway2018 propose $$\overline{\text{DID}}_{c,\ell}=\overline{Y}_{c,c+\ell}-\overline{Y}_{c,c-1}-\left(\overline{Y}_{n,c+\ell}-\overline{Y}_{n,c-1}\right),$$ a DID estimator comparing the period-$c-1$-to-$c+\ell$ outcome evolution in cohort $c$ and in the never-treated groups $n$.

callaway2018 extend those baseline estimators in various directions. First, they propose more aggregated estimators, such as $\text{DID}_{\ell}$, a weighted average of the $\overline{\text{DID}}_{c,\ell}$ estimators across all cohorts reaching $\ell$ periods after their first treatment before the end of the panel. Second, they propose estimators similar to those above, but that use the not-yet-treated instead of the never-treated as controls. For instance, all groups not yet treated at period $c$ can be used as control groups in the definition of $\overline{\text{DID}}_{c,0}$. This is very useful when there is no never-treated group: in that case, the effects $TE_{c,c+\ell}$ can still be estimated, for every $c\geq 2$ and $\ell \geq 0$ such that $\ell+c\leq U$, where $U$ is the last period when at least one group is still untreated. Even when there are never-treated groups, one may worry that such groups are less comparable to groups that get treated at some point, and researchers sometimes prefer to discard them and only leverage variation in treatment timing. Finally, even when one is fine with keeping the never-treated groups, the not-yet-treated is a larger control group, and may lead to more precise estimators. Note that in staggered adoption designs with a binary treatment, the $\text{DID}_{\text{M}}$ estimator proposed by dcDH2020 also uses the not-yet-treated as controls, and is identical to the $\text{DID}_{0}$ estimator of the instantaneous treatment effect using the not-yet-treated as controls in callaway2018. Third, callaway2018 also propose estimators relying on a conditional parallel trends assumption. Fourth, they suggest placebo estimators to test the parallel trends assumptions underlying their estimators. These placebos are robust to heterogeneous effects, unlike the coefficients $\widehat{\beta}_\ell$ for $\ell\leq -2$ from the event-study regression in (ref).

The estimators proposed by callaway2018 are computed by the csdid Stata command csdidStata, and by the did R command didR. The basic syntax of the Stata command is

csdid outcome, time(timeid) gvar(cohort)

where cohort is equal to the period when a group starts receiving the treatment.

The estimators proposed by abraham2018

abraham2018 also propose DID estimators of the cohort-and-period specific effects $TE_{c,c+\ell}$ that only rely on the parallel trends assumption in Equation (ref), and that are robust to heterogeneous treatment effects. Their estimators either use the never-treated groups as controls, or the last-treated groups if there are no never-treated. With the former control group, their estimators of the $TE_{c,c+\ell}$ parameters are identical to those proposed by callaway2018 with the same control group. Operationally, they show that their estimators can be computed via a simple linear regression, which may reduce computing time. Unlike callaway2018, they do not propose estimators relying on a conditional parallel trends assumption, and they also do not propose estimators using the not-yet-treated as controls.

Their estimators are computed by the eventstudyinteract Stata command eventstudyinteractStata. Its basic syntax is

eventstudyinteract outcome \{rel_time_list\}, absorb(i.groupid i.timeid) \\ cohort(first_treatment) control_cohort(controlgroup)

where rel_time_list is the list of relative-time indicators $1\{F_g=t-\ell\}$ one would include in the event-study regression in (ref), first_treatment is a variable equal to the period when group $g$ got treated for the first time, and controlgroup is an indicator for the control group observations (e.g.: the never treated).

The estimators proposed by borusyak2020revisiting, gardnertwo, and liu2021practical

borusyak2020revisiting, gardnertwo, and liu2021practical have proposed estimators that may be more efficient than those in callaway2018 and abraham2018, under some assumptions. We start by reviewing borusyak2020revisiting, before discussing the connection between their results and those in gardnertwo and liu2021practical. The estimators in borusyak2020revisiting can be obtained by running a TWFE regression of the outcome on group and time fixed effects, and fixed effects for every treated $(g,t)$ cell. To be concrete, if the data has $50$ groups, $10$ time periods, and $100$ treated $(g,t)$ cells, the regression has a constant and 158 fixed effects ($49$ for groups, $9$ for time periods, and $100$ for the treated $(g,t)$ cells). Under the assumptions of the Gauss-Markov theorem, the coefficients from this regression are the linear estimators of the population coefficients with the lowest variance. But under parallel trends, the population coefficient on the fixed effect for treated cell $(g,t)$ is actually equal to $TE_{g,t}$, the ATE in cell $(g,t)$, so the estimators in borusyak2020revisiting are the linear estimators of those ATEs with the lowest variance. With estimators of $TE_{g,t}$ in hand, one can estimate $TE_{c,c+\ell}$ as the average of all the $TE_{g,t}$s such that group $g$ started receiving the treatment at period $c$ and $t=c+\ell$. Again, Gauss-Markov ensures that this estimator is the best linear estimator of $TE_{c,c+\ell}$. As the estimators in callaway2018 and abraham2018 are also linear estimators, those in borusyak2020revisiting have a lower variance.

A second, numerically equivalent way of computing the estimators in borusyak2020revisiting amounts to fitting a regression of the outcome on group and time fixed effects in the sample of untreated observations, and using that regression to predict the counterfactual outcome of treated observations. Estimates of the treatment effect of those observations are then merely obtained by substracting their counterfactual to their actual outcome. This imputation method is computationally faster than the first. It also readily generalizes to more complicated specifications, such as triple-differences, or models allowing for group-specific linear trends. Using this representation of their estimator, borusyak2020revisiting show that it can also be used to estimate the effect of a binary and non-staggered treatment, if that treatment does not have dynamic effects. This imputation method is the one used by the did_imputation Stata command did_imputationStata and by the didimputation R command did_imputationR to compute the estimators proposed by borusyak2020revisiting. The basic syntax of the Stata command is:

did_imputation outcome groupid timeid first_treatment,

where first_treatment is a variable equal to the period when group $g$ first got treated.

Before borusyak2020revisiting, liu2021practical and gardnertwo have proposed the same imputation method as in borusyak2020revisiting,\footnote{Even before that, gobillon2016regional have proposed a similar strategy to estimate treatment effects under a factor model.} but the result showing that the resulting estimators are efficient under the assumptions of the Gauss-Markov theorem only appears in borusyak2020revisiting. Note that wooldridge2021two has also proposed an estimation strategy connected, and in some cases numerically equivalent, to that of borusyak2020revisiting.

Understanding the differences between those estimators

Under parallel trends, the estimators in borusyak2020revisiting may offer precision gains with respect to those in callaway2018 or abraham2018, under the assumptions of the Gauss-Markov theorem. Those require, among other things, that the never treated potential outcomes $Y_{g,t}(\bm{0}_t)$ be independent of each other, both across groups and over time. It is, of course, often implausible that the potential outcomes of the same group are uncorrelated over time. With serial correlation, it is no longer guaranteed that the estimators in borusyak2020revisiting will always be more efficient than those in callaway2018 and abraham2018, but simulations in borusyak2020revisiting suggest that one can still expect efficiency gains with moderate serial correlation.

If trends are not exactly parallel, the estimators in borusyak2020revisiting may be more or less biased than those in callaway2018 or abraham2018 depending on the nature of the violation of parallel trends. borusyak2020revisiting do not provide a closed-form of their estimators, but one can show that with only one treated group $s$, which starts to receive the treatment at period $t_s$, their estimator of that group's treatment effect at $t_s+\ell$ is

equation[equation omitted — 184 chars of source]

while the estimator in callaway2018 and abraham2018 is

equation[equation omitted — 125 chars of source]

Equation (ref) shows that the estimator in callaway2018 and abraham2018 use groups' $t_s-1$ outcome, the last period before $s$ gets treated, as the baseline outcome, while Equation (ref) shows that the estimator in borusyak2020revisiting instead uses the average outcome from period $1$ to $t_s-1$ as the baseline. This is why the latter estimator is often more precise. However, it is also more biased, when parallel trends does not exactly hold and the discrepancy between groups' trends gets larger over longer horizons, as would for instance happen when there are group-specific linear trends. In such instances, roth2019pre notes that leveraging earlier pre-treatment periods increases the bias of a DID estimator, since one makes comparisons from earlier periods. If, on the other hand, parallel trends fails due to anticipation effects arising a few periods before $t_s$, Equations (ref) and (ref) imply that the estimator in borusyak2020revisiting is less biased than that in callaway2018 and abraham2018. However, these two types of violations of parallel trends may not be equally problematic. Often times, both estimators can be immunized against anticipation effects, by redefining $t_s$ as the date when the treatment was announced. On the other hand, it is often harder to immunize them against differential trends widening over time de2020difference. Beyond the simple example we consider here, deriving a closed-form expression of the estimators in borusyak2020revisiting is not straightforward. Whether the conclusions we derive in this simple example carry through to more complicated designs is thus an open question.

If one views parallel trends as a reasonable first-order approximation rather than an assumption that holds exactly, it may make sense to investigate how sensitive one's findings are to violations of parallel trends. To do so, one may for instance implement the partial identification approach in manski2018right or rambachan2019honest. The latter approach assumes that parallel trends do not hold exactly, and that the magnitude of placebo estimators is informative as to the magnitude of the bias in the actual estimators caused by differential trends. The estimators proposed by callaway2018 and abraham2018 may be more amenable to the approach in rambachan2019honest than the estimators proposed by borusyak2020revisiting. Consider again the same simple example as above. For any $\ell\leq t_s-2$, one can construct the following placebo estimator:

equation[equation omitted — 137 chars of source]

This placebo compares the treated and control groups' outcome evolution, from period $t_s-\ell-2$ to $t_s-1$, namely over $\ell+1$ periods before group $s$ got treated. It exactly mimicks the estimator of group $s$'s treatment effect at period $t_s+\ell$ proposed by callaway2018 and abraham2018, which compares the same groups, over the same number of periods. Accordingly, the magnitude of that placebo may indeed be informative as to the magnitude of the bias of the estimator in Equation (ref), as requested by rambachan2019honest. Building a placebo that would similarly mimick the estimator proposed by borusyak2020revisiting is not feasible, precisely because that estimator leverages all pre-treatment periods to construct its baseline. See de2020difference for more discussion of the advantages of having placebos that mimick actual estimators.

Another difference between these approaches is that borusyak2020revisiting impose parallel trends for every group and between every pair of consecutive time periods.\footnote{dcDH2020 and abraham2018 also impose that assumption.} callaway2018, on the other hand, impose a weaker parallel trends assumption: from period $c$ onwards, cohort $c$ must be on the same trend as the never-treated groups, but before that cohort $c$ may have been on a different trend. The assumption in callaway2018 is the minimal assumption ensuring that all the $TE_{c,c+\ell}$ can be unbiasedly estimated, but it is conditional on the design: which groups are required to be on parallel trends at which dates depends on groups' realized treatments. It is also not testable. We refer the reader to Marcus2020 and borusyak2020revisiting for further discussion on the differences between parallel trends assumptions.

Overall, whether the estimators in borusyak2020revisiting should be preferred to those in callaway2018 and abraham2018 may depend on one's degree of confidence in the parallel trends assumption, on the type of violations of this assumption that seems more likely to arise in the application at hand, on whether it is possible to immunize the estimators against anticipation effects by redefining the treatment date as the announcement date, and on one's willingness to undertake a sensitivity analysis such as the one proposed by rambachan2019honest. Note also that if the estimators proposed by borusyak2020revisiting, callaway2018, and abraham2018 are significantly different, this implies that the parallel trends assumption, at least the “strong version” of this assumption imposed by borusyak2020revisiting and abraham2018, must be violated.

Estimators allowing for dynamic effects when the treatment is not binary or the design is not staggered.

de2020difference propose treatment effect estimators robust to heterogeneous and dynamic treatment effects and that can be used even if the treatment is not binary or the design is not staggered. In their survey of 26 highly cited 2015-2019 AER papers using a TWFE regression, they find that 4 have a binary treatment and a staggered design, so being able to accommodate more general designs is important. The paper's main idea is to propose a generalization of the event-study approach to such designs, by defining the event as the period where a group's treatment changes for the first time. With a binary-and-staggered treatment, the event per this definition is the period where a group gets treated, so this definition extends the standard one to general designs.

More specifically, de2020difference start by showing that for any group $g$ whose treatment changed for the first time at period $F_g$, the instantaneous and dynamic effects of that change can be unbiasedly estimated. Let $$\delta_{g,\ell}=E(Y_{g,F_g+\ell}-Y_{g,F_g+\ell}(D_{g,1},...,D_{g,1}))$$ be the expected difference between group $g$'s actual outcome at $F_g+\ell$ and the counterfactual “status quo” outcome it would have obtained if its treatment had remained equal to its period-one value from period one to $F_g+\ell$. Let $N^c_{g,\ell}$ denote the number of groups whose treatment has not changed yet at $F_g+\ell$, and with the same treatment as $g$ at period one. de2020difference show that

equation*[equation* omitted — 167 chars of source]

a DID estimator comparing the $F_g-1$-to-$F_g+\ell$ outcome evolution between group $g$ and groups whose treatment has not changed yet at $F_g+\ell$ and with the same treatment as $g$ at period one, is unbiased for $\delta_{g,\ell}$ under parallel trends assumptions. To test those parallel trends assumptions, they propose placebo estimators comparing the outcome trends of switchers and non-switchers before the switchers switch.

Then, de2020difference aggregate the $\text{DID}_{g,\ell}$ estimators into an estimator of the effect of having experienced a weakly higher amount of treatment for $\ell$ periods. For any real number $x$ and $t\in \{1,...,T\}$, let $\bm{x}_t$ denote a $1\times t$ vector with coordinates equal to $x$. When the treatment is binary, for groups untreated at period one, $D_{g,1}=0$, so $$\delta_{g,\ell}=E(Y_{g,F_g+\ell}(\bm{0}_{F_g-1},1,D_{g,F_g+1},...,D_{g,F_g+\ell})-Y_{g,F_g+\ell}(\bm{0}_{F_g+\ell})).$$ For groups treated at period one, $D_{g,1}=1$, so $$-\delta_{g,\ell}=E(Y_{g,F_g+\ell}(\bm{1}_{F_g+\ell})-Y_{g,F_g+\ell}(\bm{1}_{F_g-1},0,D_{g,F_g+1},...,D_{g,F_g+\ell})).$$ The right-hand side of the two equations above are effects of having experienced a weakly higher amount of treatment for $\ell+1$ periods. Accordingly, the $\text{DID}_{g,\ell}$ estimators are aggregated into a $\text{DID}_{\ell}$ estimator, multiplying by minus one the $\text{DID}_{g,\ell}$ of groups treated at period one. With a non-binary treatment, one can also aggregate the $\text{DID}_{g,\ell}$ to estimate the effect of having experienced a weakly higher amount of treatment for $\ell+1$ periods.

Ultimately, this approach leads to an event-study graph, with the distance to the first treatment change on the $x$-axis, the $\text{DID}_{\ell}$ estimators on the $y$-axis to the right of zero, and placebo estimators on the $y$-axis to the left of zero. This event-study graph is useful to test the parallel trends assumption, and to provide reduced-form evidence of whether weakly increasing the treatment for $\ell+1$ periods increases or decreases the outcome on average. However, interpreting the magnitude of the $\text{DID}_{\ell}$ estimators might be complicated. For instance, with three periods and three groups such that $(D_{1,1}=0,D_{1,2}=4,D_{1,3}=1)$, $(D_{2,1}=0,D_{2,2}=2,D_{2,3}=3)$, and $(D_{2,1}=0,D_{2,2}=0,D_{2,3}=0)$, $\text{DID}_{1}$ estimates the average of $E(Y_{1,3}(0,4,1)-Y_{1,3}(0,0,0))$ and $E(Y_{2,3}(0,2,3)-Y_{2,3}(0,0,0))$. Accordingly, $\text{DID}_{1}$ does not estimate by how much the outcome increases on average when the treatment increases by a given amount for a given number of periods.

To circumvent this important limitation, two strategies can be implemented. First, the reduced-form event-study graph described above can be complemented with a first-stage event-study graph, where the outcome is replaced by the treatment. The estimators on the first-stage graph show the average value of $|D_{g,F_g+\ell}-D_{g,1}|$ across all groups entering in $\text{DID}_{\ell}$. In the example above, the first two estimates on the first-stage graph are equal to $1/2(D_{1,2}-D_{1,1}+D_{2,2}-D_{2,1})=3$ and $1/2(D_{1,3}-D_{1,1}+D_{2,3}-D_{2,1})=2$. This reflects the fact that in this example, $\text{DID}_{1}$ is an effect produced by increasing the previous and current treatment by $3$ and $2$ units on average. Second, a weighted average across $\ell$ of the reduced-form estimators divided by a weighted average across $\ell$ of the first-stage estimators is unbiased for a parameter with a clear economic interpretation. That parameter may be used to conduct a cost-benefit analysis comparing groups' actual treatments to the status quo scenario where they would have kept all along the same treatment as in period one. In other words, that parameter can be used to determine if the policy changes that took place over the duration of the panel led to a better situation than the one that would have prevailed if no policy change had been undertaken, a natural policy question. Importantly, that parameter can also be interpreted as an average total effect per unit of treatment, where “total effect” refers to the sum of the instantaneous and dynamic effects of a treatment.

The estimators proposed by de2020difference are computed by the did_multiplegt Stata and R commands. To compute those estimators rather than those proposed in dcDH2020, the Stata command's basic syntax is:

did_multiplegt outcome groupid timeid treatment, robust_dynamic dynamic(\#) \\ average_effect placebo(\#) longdiff_placebo breps(\#) cluster(groupid),

where dynamic(\#) specifies the horizon over which effects of a first treatment switch have to be estimated, and placebo(\#) specifies the number of placebos to be estimated.

The estimators in de2020difference can be used with a binary treatment switching on and off, with a discrete treatment, or with a continuous and staggered treatment (groups start getting treated at different dates, with differing intensities, but once a group gets treated its treatment intensity never changes). The estimators proposed by callaway2021difference can also accommodate continuous and staggered treatments. For continuous and non-staggered treatments, in their Section 4.3 chaisemartin2022continuous extend their baseline estimators to allow for dynamic effects. With respect to their baseline estimators, the main difference is that when allowing for dynamic effects, fewer units can be used as controls. Without dynamic effects, at period $t$, any unit whose treatment has not changed between $t-1$ and $t$ can be used as a valid control. With dynamic effects, only units whose treatments have not changed from period 1 to $t$ can be used as valid controls. Therefore, the need for “stayers” becomes even stronger when allowing for dynamic effects: many units need to keep the same value of the treatment for a large number of time periods. Developing estimators robust to dynamic effects that can be used with a continuous treatment and no stayers has not been done yet and is a promising area for future research.

The estimators in de2020difference can, of course, also be used with a binary and staggered treatment. Without covariates in the estimation, they are then equivalent to the estimators proposed by callaway2018 using the not-yet-treated as controls. With covariates, the estimators in callaway2018 and de2020difference differ. callaway2018 consider time-invariant covariates, and assume that trends are parallel once we condition on them. de2020difference instead consider time-varying covariates and assume that trends are parallel once the linear effect of those time-varying covariates on the outcome is accounted for. This for instance allows them to include group-specific linear trends in the estimation. With covariates, the parallel trends conditions in callaway2018 and de2020difference are not nested, and in principle one could combine both.

Finally, it is worth noting that de2020two propose estimators for the case with several treatments. They propose both estimators that generalize the $\text{DID}_{\text{M}}$ estimator in dcDH2020 and rule out dynamic effects, and estimators that generalize those in callaway2018 and allow for dynamic effects.

Application

\setcounter{equation}{0}

In this section, we revisit an application with a binary and staggered treatment, thus allowing us to compute several of the heterogeneity-robust DID estimators reviewed above. Between 1968 and 1988, 29 US states adopted a unilateral divorce law (UDL). wolfers2006did, building upon friedberg1998did, studies the effects of those laws on divorce rates, using a version of the event-study regression in (ref). We use his data (Wolfersdata, Wolfersdata) to revisit this question. In what follows, estimates are weighted by states' populations and standard errors are clustered at the state level, as in wolfers2006did. As the author estimates UDLs' dynamic effects up to 15 years after adoption, in our replication we focus on heterogeneity-robust DID estimators allowing for dynamic effects, and present the estimated effects over the same horizon. We use Stata for this replication exercise, and the versions of the twowayfeweights, eventstudyinteract, csdid, did_imputation, and did_multiplegt commands available from the SSC repository at the end of April 2022.

Figure (ref) below shows the instantaneous and dynamic effects of passing a UDL, according to six estimation methods. In the top-left panel, we show the estimates from the event-study regression in (ref), with $L=15$, $K=10$, and endpoint binning. According to this regression, UDLs increase the divorce rate on the year when the law is passed and for seven years thereafter. 11 years after those laws are passed, their effect becomes significantly negative. Those effects are consistent with those in Column (1) of Table 2 of wolfers2006did. Our event-study regression and that in wolfers2006did differ on two dimensions: wolfers2006did does not include any placebo indicator for pre-adoption periods, and he includes post-adoption indicators for bins of two years (one indicator for the year when the law is passed and the year after that, one indicator for the second and third years after the law is passed, etc.). Results seem fairly robust to those specification choices. The placebo estimates are small, and individually and jointly insignificant (F-test p-value=$0.863$).

We follow abraham2018, and compute the weights attached to UDLs' instantaneous effect in this event-study regression.\footnote{In practice, we use the twowayfeweights Stata command, which has an option to compute the correlation between the weights and other variables that we use below.} As shown in Equation (ref), this coefficient can be decomposed as the sum of two terms. The first term is a weighted sum of UDL's effects in the year when they are passed, across 27 states, where all effects receive a positive weight. The weights are negatively correlated with the year variable (correlation=$-0.232$), so this first term upweights UDLs' instantaneous effects in states passing a law early, and downweights UDLs' instantaneous effects in states passing a law late. Accordingly, this first term may differ from the average instantaneous effects of UDLs if those effects vary between early- and late-adopting states, but it at least estimates a convex combination of effects. The second term is a weighted sum of UDLs' effects in the years after they are passed. 29 effects of having passed a UDL a year ago enter in that second term. 16 enter with a positive weight, and 13 enter with a negative weight. The positive and negative weights respectively sum to $0.012$ and $-0.012$. 28 effects of having passed a UDL two years ago enter in that second term. 10 effects enter with a positive weight, and 18 enter with a negative weight. The positive and negative weights respectively sum to $0.010$ and $-0.010$. Effects of having passed a UDL three, four, ..., 14, and more than 15 years ago also enter in that second term. In total, the positive and negative weights in that second term respectively sum to around $0.064$ and $-0.064$. If UDLs' dynamic effects vary across states, that second term may not be equal to zero, thus further biasing the estimated instantaneous effect in the event-study regression. However, those contamination weights are not very large, so this bias is likely to be small. Overall, this event-study regression seems fairly robust to heterogeneous treatment effects.

In the top-centre panel of Figure (ref), we use the eventstudyinteract command to compute the estimators proposed by abraham2018. The estimated effects are very similar to those in the top-left panel. This could either be due to the fact that UDLs effects are not very heterogeneous, or to the fact that the event-study regression is fairly robust to heterogeneous treatment effects, as suggested above. Interestingly, the confidence intervals are, if anything, slightly wider in the top-left than in the top-centre panel of Figure (ref), thus showing that heterogeneity-robust DID estimators are not always less precise than TWFE estimators. The placebos are individually insignificant. They are also substantially smaller than the estimated effects of UDLs: it does not seem that violations of parallel trends can fully account for those estimated effects.

In the top-right panel of Figure (ref), we use the csdid command to compute the estimators proposed by callaway2018, using the “not-yet-treated” states as the control group. The estimated effects are very similar to those in the top-centre panel. 19 states never adopt a UDL over the period under consideration, so the group of “never-treated” states used as controls by eventstudyinteract is quite large, and accounts for a relatively large fraction of the group of “not-yet-treated” states used as controls by csdid. This may explain why in this application, the two commands yield very similar estimates. Using the larger control group of “not-yet-treated” states also does not lead to markedly more precise estimates: the widths of the confidence intervals are similar in the two panels. The placebos produced by csdid are small and individually insignificant. The placebos are much smaller in the top-right than in the top-centre panel. This is because csdid computes first-difference placebos, comparing the outcome evolution of treated and not-yet treated states, before the treated start receiving the treatment, and between pairs of consecutive periods.\footnote{csdid has an option to compute long-difference placebos, but it returned an error when we used it.} On the other hand, \texttt{eventstudyinteract} computes long-difference placebos. For instance, the second placebo, shown at $t=-3$ on the graph, compares the outcome evolution of treated and never-treated states, from $F_g-1$, the period before the treated start getting treated, to $F_g-3$. See de2020difference for a discussion of the respective advantages of long- and first-difference placebos.

In the bottom-left panel of Figure (ref), we use the did_imputation command to compute the estimators proposed by borusyak2020revisiting. The effects are very similar to those found by the previous two estimators. The confidence interval of the instantaneous effect is much tighter in the bottom-left panel than in all other panels: for that treatment effect, the estimator proposed by borusyak2020revisiting does lead to a large precision gain. However, the opposite can hold when one considers dynamic effects. For instance, the confidence interval of the effect two years after passing a UDL is more than 50% larger per did_imputation than per csdid. Accordingly, the estimators proposed by borusyak2020revisiting do not always lead to precision gains, relative to those proposed by abraham2018 or callaway2018. The placebos produced by did_imputation are small, individually insignificant, and jointly insignificant (F-test p-value = 0.541).\footnote{We did not report a joint test that all placebos are equal to 0 based on eventstudyinteract: this command does not readily allow to compute this test, as it does not return the covariances between the estimators. Similarly, csdid does not allow to jointly test if the placebos in Figure (ref) are significant: it computes a joint nullity test, but for more disaggregated placebos.} Note that the placebos computed by \texttt{did_imputation} are different from those computed by the other commands. Essentially, the command estimates a TWFE regression among all the untreated $(g,t)$, with $K$ leads of the treatment. To be consistent with the other estimations, we run the command with 9 leads. Then, everything is relative to 10 periods prior to treatment, which is why the placebo estimate is set to 0 at $t=-10$ in the bottom-left panel, instead of at $t=-1$ in the other panels.

In the bottom-centre panel of Figure (ref), we use the did_multiplegt command to compute the estimators proposed by de2020difference. The resulting estimates are extremely close to those produced by the csdid command. The only reason why the two sets of estimates are not identical is that the estimation is weighted by states' population, and the two commands seem to handle weights slightly differently. Without weighting, the two sets of estimates are identical, as expected given that there are no covariates in the estimation and we used csdid with the not-yet-treated as controls. The placebos computed by did_multiplegt are long-difference placebos, similar to those computed by eventstudyinteract, except that did_multiplegt uses the not-yet-treated as controls. They are small, and individually and jointly insignificant (F-test p-value = 0.427).

figure[figure omitted — 1,125 chars of source]

The estimates discussed so far do not control for state-specific linear trends. Whether such trends should or should not be included to estimate the effect of UDLs has been a debated issue in this literature, with friedberg1998did arguing in their favor, and wolfers2006did arguing that they may conflate dynamic effects. The results presented so far already suggest that including state-specific linear trends is unnecessary, as placebos are small and insignificant without them. To confirm that, we run the did_multiplegt command again, controlling for state-specific linear trends.\footnote{csdid does not allow for group-specific trends. did_imputation allows in principle for such trends but returned an error when such trends were added. eventstudyinteract allows for such trends.} The results, displayed in the bottom-right panel of Figure (ref), show that results are fairly insensitive to the inclusion of state-linear trends. If anything, adding them makes the estimated long-run effects more noisy. The only argument in favor of state-specific trends is that the placebos are slightly smaller with them, though the difference is most likely insignificant.

Finally, to synthetize our results and obtain a point estimate that can be compared to the results in wolfers2006did, we average UDL's effects from the year the law is passed to seven years thereafter. The results are displayed in Table (ref). We do not include therein the estimates from the eventstudyinteract and csdid commands, as one cannot readily obtain the standard error of this average effect from these commands. The results show that according to all estimation methods, UDLs positively affect the divorce rate from the year the law is passed to seven years thereafter. All estimates are fairly similar to each other and point towards an increase of 20%. The estimated standard error is substantially lower using the author's original specification, which is not surprising as it is less flexible than the other estimation methods. The estimated standard error is slightly larger using borusyak2020revisiting than the flexible event-study regression or the estimators proposed by de2020difference.

table[table omitted — 1,114 chars of source]

Conclusion, and avenues for future research

\setcounter{equation}{0}

The literature reviewed in this survey has shown that TWFE regressions may not always estimate a convex combination of treatment effects. In such cases, it may be hard to give them a causal interpretation, as TWFE coefficients could for instance be of a different sign than every unit's treatment effect. Table (ref) below summarizes the alternative estimators available to applied researchers, depending on their research design and on whether they are ready or not to rule out dynamic effects. The table shows that the literature so far has mostly focused on providing alternative estimators for the case with a binary treatment and staggered adoption. Heterogeneity-robust DID estimators that can be used in more complicated designs are scarce, while many applications where TWFE regressions have been used either do not have a staggered design, or do not have a binary treatment. Developing more estimators that can be used in such designs is a promising avenue for future research. This can often be done by building upon the insights gained from studying the binary-and-staggered case. For instance, the estimators proposed by de2020difference build upon those proposed by callaway2018 for the binary-and-staggered case. We hope that the whirlwind of DID working papers shall continue, till heterogeneity-robust DID estimators are as widely applicable as TWFE regressions.

It is also important to stress that at this stage, it is still unclear whether researchers should systematically abandon TWFE estimators. Those estimators sometimes estimate a convex combination of effects under the parallel trends assumption, they may estimate the ATT if the weights attached to them are uncorrelated with the treatment effects $TE_{g,t}$, and they often have a lower variance than the heterogeneity-robust estimators reviewed in the previous section. While there are examples where TWFE and heterogeneity-robust DID estimators are economically and statistically different dcDH2020,de2020difference,de2020two,baker2022much, the previous section also shows a data set where TWFE and heterogeneity-robust DID estimators lead to very similar conclusions. Understanding the circumstances where TWFE and heterogeneity-robust DID estimators are more likely to differ is an important question. We conjecture that differences are likely to be larger in complicated designs (e.g.: a non-binary treatment that can turn on and off multiple times, or several treatments) than in simple designs (e.g.: a single binary and staggered treatment). This conjecture is based on our discussion of Equation (ref) in Section 3. This is also a pattern we found when computing TWFE and heterogeneity-robust DID estimators in four different data sets, in the empirical examples of this survey and of dcDH2020 (dcDH2020,de2020difference,de2020two). But those examples are not enough to draw general conclusions: a systematic comparison of TWFE and heterogeneity-robust DID estimators in a broad set of applications is in order.

Analyzing estimators' robustness to heterogeneous treatment effects is important, as the assumption that all units are affected in the same way by a treatment is seldom credible. In this survey, we have focused on estimators relying on parallel trends assumptions, but this question is also relevant for other estimators. See for instance sloczynski2020should and blandhol2022tsls for instrumental variables estimators with covariates. More closely related to our set-up, the impact of heterogeneous treatment effects in the “group fixed-effects” model of bonhomme2015grouped remains to be studied.

table[table omitted — 2,618 chars of source]