Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
80,645 characters · 0 sections · 48 citation commands
Two-way fixed effects estimators with heterogeneous treatment effects
Keywords: linear regressions, fixed effects, heterogeneous treatment effects, difference-in-differences.
JEL Codes: C21, C23
\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Introduction}
A popular method to estimate the effect of a treatment on an outcome is to compare over time groups experiencing different evolutions of their exposure to treatment. In practice, this idea is implemented by estimating regressions that control for group and time fixed effects. Hereafter, we refer to those as two-way fixed effects (FE) regressions. We conducted a survey, and found that 20% of all empirical articles published by the American Economic Review (AER) between 2010 and 2012 have used a two-way FE regression to estimate the effect of a treatment on an outcome. When the treatment effect is constant across groups and over time, such regressions estimate that effect under the standard “common trends” assumption. However, it is often implausible that the treatment effect is constant. For instance, the minimum wage's effect on employment may vary across US counties, and may change over time. This paper examines the properties of two-way FE regressions when the constant effect assumption is violated.
We start by assuming that all observations in the same $(g,t)$ cell have the same treatment and that the treatment is binary, as is for instance the case when the treatment is a county-level law. We consider the regression of $Y_{i,g,t}$, the outcome of unit $i$ in group $g$ at period $t$ on group fixed effects, period fixed effects, and $D_{g,t}$, the treatment in group $g$ at period $t$. Let $\widehat{\beta}_{fe}$ denote the coefficient of $D_{g,t}$, and let $\beta_{fe}$ denote its expectation. Under the common trends assumption, we show that $\beta_{fe}$ is equal to a weighted sum of the treatment effect in each treated $(g,t)$ cell:
$\Delta_{g,t}$ is the average treatment effect (ATE) in group $g$ and period $t$ and the weights $W_{g,t}$s sum to one but may be negative. Negative weights arise because $\widehat{\beta}_{fe}$ is a weighted sum of several difference-in-differences (DID), which compare the evolution of the outcome between consecutive time periods across pairs of groups. However, the “control group” in some of those comparisons may be treated at both periods. Then, its treatment effect at the second period gets differenced out by the DID, hence the negative weights.
The negative weights are an issue when the ATEs are heterogeneous across groups or periods. Then, one could have that $\beta_{fe}$ is negative while all the ATEs are positive. For instance, $1.5\times 1-0.5\times 4$, a weighted sum of $1$ and $4$, is strictly negative. Using the data set of gentzkow2011, we find that 40% of the weights attached to $\beta_{fe}$ are negative, so $\beta_{fe}$ is not robust to heterogeneous effects.\footnote{ gentzkow2011 do not estimate $\beta_{fe}$, but $\beta_{fd}$, the treatment coefficient in the first-difference regression defined below. 46% of the weights attached to $\beta_{fd}$ are strictly negative.}
Researchers may want to know how serious that issue is in the application they consider. We show that conditional on all treatments, the absolute value of the expectation of $\widehat{\beta}_{fe}$ divided by the standard deviation of the weights is equal to the minimal value of the standard deviation of the ATEs across the treated $(g,t)$ cells under which the average treatment on the treated (ATT) may actually have the opposite sign than that coefficient. One can estimate that ratio to assess the robustness of the two-way FE coefficient. If that ratio is close to 0, that coefficient and the ATT can be of opposite signs even under a small and plausible amount of treatment effect heterogeneity. In that case, treatment effect heterogeneity would be a serious concern for the validity of that coefficient. On the contrary, if that ratio is very large, that coefficient and the ATT can only be of opposite signs under a very large and implausible amount of treatment effect heterogeneity.
Finally, we propose a new estimator, $\text{DID}_{\text{M}}$, that is valid even if the treatment effect is heterogeneous over time or across groups. It estimates the average treatment effect across all the $(g,t)$ cells whose treatment changes from $t-1$ to $t$. It relies on common trends assumptions on both potential outcomes. Those conditions are partly testable, and we propose a test that amounts to looking at pre-trends. This test differs from the standard event study pre-trends test autor2003, which has been shown to be invalid when treatment effects are heterogeneous abraham2018. We show that our estimator is asymptotically normal. We compute it in the data sets of gentzkow2011 and vella1998whose, and in both cases we find that it is significantly different from $\widehat{\beta}_{fe}$.\footnote{In both cases, our estimator is also significantly different from $\widehat{\beta}_{fd}$.} Our estimator can be used in applications where, for each pair of consecutive dates, there are groups whose treatment does not change. We estimate that this condition is satisfied for around 80% of the papers using two-way fixed effects regressions found in our survey of the AER.
Overall, our paper has implications for applied researchers estimating two-way fixed effects regressions. First, we recommend that they compute the weights attached to their regression and the ratio of $|\widehat{\beta}_{fe}|$ divided by the standard deviation of the weights. To do so, they can use the twowayfeweights Stata package that is available from the SSC repository. If many weights are negative, and if the ratio is not very large, we recommend that they compute our new estimator, using the fuzzydid and did_multiplegt Stata packages, also available from the SSC repository deChaisemartin18.
We extend our results in several important directions. First, another commonly-used regression is the first-difference regression of $Y_{g,t}-Y_{g,t-1}$, the change in the mean outcome in group $g$, on period fixed effects and on $D_{g,t}-D_{g,t-1}$, the change in the treatment. We let $\beta_{fd}$ denote the expectation of the coefficient of $D_{g,t}-D_{g,t-1}$. We show that under common trends, $\beta_{fd}$ also identifies a weighted sum of treatment effects, with potentially some negative weights. Second, in our Web Appendix we show that our results extend to fuzzy designs, where the treatment varies within $(g,t)$ cells, and to two-way fixed effects regressions with a non-binary treatment and with covariates.
Our paper is related to the DID literature. Our main result generalizes Theorem 1 in deChaisemartin15b. When the data has two groups and two periods, the Wald-DID estimand considered therein is equal to $\beta_{fe}$ and $\beta_{fd}$. Our results on $\beta_{fe}$ and $\beta_{fd}$ are thus extensions of that theorem to the case with multiple periods and groups.\footnote{In fact, a preliminary version of our main result appeared in a working paper version of deChaisemartin15b deChaisemartin15c.} Moreover, our $\text{DID}_{\text{M}}$ estimator is related to the Wald-TC estimator with many groups and periods proposed in deChaisemartin15b, and to the multi-period DID estimator proposed by imai2018. In Section (ref), we explain the differences between those three estimators.
More recently, borusyak2016, abraham2018, athey2018, callaway2018, and goodman2018 study the special case of staggered adoption designs, where the treatment of a group is weakly increasing over time. Those papers derive some important results specific to that design that we do not consider here. Still, some of the results in those papers are related to ours, and we describe precisely those connections later in the paper. The most important dimension on which our paper differs from those is that our results apply to any two-way fixed effects regressions, not only to those with staggered adoption. In our survey of the AER papers estimating two-way fixed effects regressions, less than 10% have a staggered adoption design. This suggests that while staggered adoptions are an important research design, they may account for a relatively small minority of the applications where two-way fixed effects regressions have been used.
The paper is organized as follows. Section (ref) introduces the set-up. Section (ref) presents our decomposition results. Section (ref) introduces our alternative estimator. Section (ref) briefly describes some of the extensions covered in our Web Appendix. Section (ref) presents our survey of the articles published in the AER, and our two empirical applications.
\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Set up}
One considers observations that can be divided into $G$ groups and $T$ periods. For every $(g,t)\in \{1,...,G\}\times \{1,...,T\}$, let $N_{g,t}$ denote the number of observations in group $g$ at period $t$, and let $N=\sum_{g,t}N_{g,t}$ be the total number of observations. The data may be an individual-level panel or repeated cross-section data set where groups are, say, individuals' county of birth. The data could also be a cross-section where cohort of birth plays the role of time. For instance, Duflo01 compares the schooling of different cohorts in Indonesia, some of which were exposed to a school construction program. It is also possible that for all $(g,t)$, $N_{g,t}=1$, e.g. a group is one individual or firm. All of the above are special cases of the data structure we consider.
One is interested in measuring the effect of a treatment on some outcome. Throughout the paper we assume that treatment is binary, but our results apply to any ordered treatment, as we show in Section (ref) of the Web Appendix. Then, for every $(i,g,t)\in \{1,...,N_{g,t}\}\times \{1,...,G\}\times \{1,...,T\}$, let $D_{i,g,t}$ and $(Y_{i,g,t}(0),Y_{i,g,t}(1))$ respectively denote the treatment status and the potential outcomes without and with treatment of observation $i$ in group $g$ at period $t$.
The outcome of observation $i$ in group $g$ and period $t$ is $Y_{i,g,t}=Y_{i,g,t}(D_{i,g,t})$. For all $(g,t)$, let
$D_{g,t}$ denotes the average treatment in group $g$ at period $t$, while $Y_{g,t}(0)$, $Y_{g,t}(1)$, and $Y_{g,t}$ respectively denote the average potential outcomes without and with treatment and the average observed outcome in group $g$ at period $t$.
Throughout the paper, we maintain the following assumptions.
Assumption (ref) requires that no group appears or disappears over time. This assumption is often satisfied. Without it, our results still hold but the notation becomes more complicated as the denominators of some of the fractions below may then be equal to zero.
Assumption (ref) requires that units' treatments do not vary within each $(g,t)$ cell, a situation we refer to as a sharp design. This is for instance satisfied when the treatment is a group-level variable, for instance a county- or a state-law. This is also mechanically satisfied when $N_{g,t}=1$. In our survey in Section (ref), we find that almost 80% of the papers using two-way fixed effects regressions and published in the AER between 2010 and 2012 consider sharp designs. We focus on sharp designs because of their prevalence, but in Section (ref) of the Web Appendix, we show that all the results in Sections (ref)-(ref) below can be extended to fuzzy designs.
We consider $D_{g,t}$, $Y_{g,t}(0)$, $Y_{g,t}(1)$ as random variables. For instance, aggregate random shocks may affect the average potential outcomes of group $g$ at period $t$. The treatment status of group $g$ at period $t$ may also be random. The expectations below are taken with respect to the distribution of those random variables. Assumption (ref) allows for the possibility that the treatments and potential outcomes of a group may be correlated over time, but it requires that the potential outcomes and treatments of different groups be independent.
Assumption (ref) requires that the shocks affecting a group's $Y_{g,t}(0)$ be mean independent of that group's treatment sequence. This rules out the possibility that a group gets treated because it experiences negative shocks, the so-called Ashenfelter's dip ashenfelter1978estimating. Assumption (ref) is related to the strong exogeneity condition in panel data models, which, as is well-known, is necessary to obtain the consistency of the fixed effects estimator wooldridge2002.
We now define the FE regression described in the introduction.\footnote{ Throughout the paper, we assume that $D_{g,t}$ in Regression (ref) and $D_{g,t}-D_{g,t-1}$ in Regression (ref) below are not collinear with the other independent variables in those regressions, so $\widehat{\beta}_{fe}$ and $\widehat{\beta}_{fd}$ are well-defined.}
For all $g$ and $t$, let $N_{g,.}=\sum_{t=1}^{T}N_{g,t}$ and $N_{.,t}=\sum_{g=1}^{G}N_{g,t}$ respectively denote the total number of observations in group $g$ and in period $t$. For any variable $X_{g,t}$ defined in each $(g,t)$ cell, let $X_{g,.}=\sum_{t=1}^{T}(N_{g,t}/N_{g,.})X_{g,t}$ denote the average value of $X_{g,t}$ in group $g$, let $X_{.,t}=\sum_{g=1}^{G}(N_{g,t}/N_{.,t})X_{g,t}$ denote the average value of $X_{g,t}$ in period $t$, and let $X_{.,.}=\sum_{g,t}(N_{g,t}/N)X_{g,t}$ denote the average value of $X_{g,t}$. For instance, $D_{3,.}$ and $D_{.,2}$ respectively denote the average treatment in group 3 across time and in period 2 across groups, whereas $Y_{.,.}$ denotes the average value of the outcome across groups and time. Finally, for any variable $X_{g,t}$, we let $\boldsymbol{X}$ denote the vector $(X_{g,t})_{(g,t)\in \{1,...,G\}\times \{1,...,T\}}$ collecting the values of that variable in each $(g,t)$ cell. For instance, $\boldsymbol{D}$ is the vector $(D_{g,t})_{(g,t)\in \{1,...,G\}\times \{1,...,T\}}$ collecting the treatments of all the $(g,t)$ cells.
\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Two-way fixed effects regressions}
\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{A decomposition result}
We study the FE regression under the following common trends assumption.
Assumption (ref) requires that the expectation of the outcome without treatment follow the same evolution over time in every group. When $t$ represents birth cohorts, Assumption (ref) requires that the outcome difference between consecutive cohorts be the same across groups.
Let $N_1=\sum_{i,g,t}D_{i,g,t}$ denote the number of treated units, let $$\Delta^{TR} = \frac{1}{N_1} \sum_{(i, g,t):D_{g,t}=1}\left[Y_{i,g,t}(1)-Y_{i,g,t}(0)\right]$$ denote the average treatment effect across all treated units, and let $\delta^{TR}=E\left[\Delta^{TR}\right]$ denote the expectation of that parameter, hereafter referred to as the ATT. For any $(g,t)\in \{1,...,G\}\times \{1,...,T\}$, let $$\Delta_{g,t}=\frac{1}{N_{g,t}}\sum_{i=1}^{N_{g,t}} \left[Y_{i,g,t}(1)-Y_{i,g,t}(0)\right]$$ denote the ATE in cell $(g,t)$. $\delta^{TR}$ is equal to the expectation of a weighted average of the treated cells' $\Delta_{g,t}$s:
Under the common trends assumption, we show that $\beta_{fe}$ is also equal to the expectation of a weighted sum of the $\Delta_{g,t}$s, with potentially some negative weights.
Let $\varepsilon_{g,t}$ denote the residual of observations in cell $(g,t)$ in the regression of $D_{g,t}$ on group and period fixed effects:\footnote{ $\varepsilon_{g,t}$ arises from a unit-level regression, where the dependent and independent variables only vary at the $(g,t)$ level. Therefore, all the units in the same $(g,t)$ cell have the same value of $\varepsilon_{g,t}$.} $$D_{g,t}=\alpha+\gamma_g+\lambda_t+\varepsilon_{g,t}.$$ One can show that if the regressors in Regression (ref) are not collinear, the average value of $\varepsilon_{g,t}$ across all treated $(g,t)$ cells differs from 0: $\sum_{(g,t):D_{g,t}=1}(N_{g,t}/N_1)\varepsilon_{g,t}\neq 0$. Then we let $w_{g,t}$ denote $\varepsilon_{g,t}$ divided by that average: $$w_{g,t}=\frac{\varepsilon_{g,t}}{\sum_{(g,t):D_{g,t}=1}\frac{N_{g,t}}{N_1}\varepsilon_{g,t}}.$$
This result implies that in general, $\beta_{fe}\neq \delta^{TR}$, so $\widehat{\beta}_{fe}$ is a biased estimator of the ATT. To illustrate this, we consider a simple example of a staggered adoption design with two groups and three periods, and where the treatments are non-stochastic: group 1 is untreated at periods 1 and 2 and treated at period 3, while group 2 is untreated at period 1 and treated both at periods 2 and 3.\footnote{ A similar example appears in borusyak2016.} We also assume that $N_{g,t}/N_{g,t-1}$ does not vary across $g$: all groups experience the same growth of their number of observations from $t-1$ to $t$, a requirement that is for instance satisfied when the data is a balanced panel. Then, one can show that
thus implying that
The residual is negative in group 2 and period 3, because the regression predicts a treatment probability larger than one in that cell, a classic extrapolation problem with linear regressions. Then, under the common trends assumption, it follows from Theorem (ref) and the fact that the treatments are non-stochastic that
$\beta_{fe}$ is equal to a weighted sum of the ATEs in group 1 at period 3, group 2 at period 2, and group 2 at period 3, the three treated $(g,t)$ cells. However, the weight assigned to each ATE differs from $1/3$, the proportion that each cell accounts for in the population of treated observations. Therefore, $\beta_{fe}$ is not equal to $\delta^{TR}$. Perhaps more worryingly, not all the weights are positive: the weight assigned to the ATE in group 2 period 3 is strictly negative. Consequently, $\beta_{fe}$ may be a very misleading measure of the treatment effect. Assume for instance that $E\left[\Delta_{1,3}\right]=E\left[\Delta_{2,2}\right]=1$ and $E\left[\Delta_{2,3}\right]=4$. At the period when they start receiving the treatment, both groups experience a modest positive ATE. But this effect builds over time and in period 3, one period after it has started receiving the treatment, group 2 now experiences a large ATE. Then,
$\beta_{fe}$ is strictly negative, while $E\left[\Delta_{1,3}\right]$, $E\left[\Delta_{2,2}\right]$, and $E\left[\Delta_{2,3}\right]$ are all positive. More generally, the negative weights are an issue if the $E\left[\Delta_{g,t}\right]$s are heterogeneous, across groups or over time.\footnote{ On the other hand, $\beta_{fe}$ does not rule out heterogeneous treatment effects within $(g,t)$ cells, as it is identified by variations across $(g,t)$ cells, and does not leverage any within-cell variation.} If $E\left[\Delta_{1,3}\right]=E\left[\Delta_{2,2}\right]=E\left[\Delta_{2,3}\right]=1$, then $\beta_{fe}=1=\delta^{TR}$.
Here is some intuition as to why one weight is negative in this example. It follows from Equation (ref) in the proof of Theorem (ref) goodman2018 that in this simple example, $\beta_{fe}=(\text{DID}_1 +\text{DID}_2)/2$, with
The first DID compares the evolution of the mean outcome from period 1 to 2 in group 2 and in group 1. The second one compares the evolution of the mean outcome from period 2 to 3 in group 1 and in group 2. The control group in the second DID, group 2, is treated both in the pre and in the post period. Therefore, under the common trends assumption, it follows from Lemma (ref) in Appendix (ref) (a similar result appears in Lemma 1 of Chaisemartin2011fuzzy and in Equation (13) of goodman2018) that $\text{DID}_1=E\left[\Delta_{2,2}\right]$, but $$\text{DID}_2 = E\left[\Delta_{1,3}\right] - (E\left[\Delta_{2,3}\right] - E\left[\Delta_{2,2}\right]).$$ $\text{DID}_2$ is equal to the ATE in group 1 period 3, minus the change in group 2's ATE between periods 2 and 3. Intuitively, the mean outcome of groups 1 and 2 may follow different trends from period 2 to 3 either because group 1 becomes treated, or because group 2's ATE changes. The intuition that negative weights arise because $\widehat{\beta}_{fe}$ uses treated observations as controls also appears in borusyak2016.
We now generalize the previous illustration by characterizing the $(g,t)$ cells whose ATEs are weighted negatively by $\beta_{fe}$.
Proposition (ref) shows that $\beta_{fe}$ is more likely to assign a negative weight to periods where a large fraction of groups are treated, and to groups treated for many periods. Then, negative weights are a concern when treatment effects differ between periods with many versus few treated groups, or between groups treated for many versus few periods.
Proposition (ref) has interesting implications in staggered adoption designs, a special case of sharp designs defined as follows.
Assumption (ref) is satisfied in applications where groups adopt a treatment at heterogeneous dates athey2002. In that design, borusyak2016 show that $\beta_{fe}$ is more likely to assign a negative weight to treatment effects at the last periods of the panel. This result is a special case of Proposition (ref): in staggered adoption designs, $D_{.,t}$ is increasing in $t$, so Proposition (ref) implies that $w_{g,t}$ is decreasing in $t$.\footnote{borusyak2016 assume that the treatment effect of cell $(g,t)$ only depends on the number of periods since group $g$ has started receiving the treatment, whereas Proposition (ref) does not rely on that assumption.} Proposition (ref) also implies that in that design, groups that adopt the treatment earlier are more likely to receive some negative weights.
Finally, in staggered adoption designs, athey2018 derive a decomposition of $\beta_{fe}$ that resembles to, but differs from, that in Theorem (ref). They derive their decomposition under the assumption that the dates at which each group starts receiving the treatment are randomly assigned, while we derive ours under a common trends assumption.
\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Robustness to heterogeneous treatment effects}
Theorem (ref) shows that in sharp designs with many groups and periods, $\widehat{\beta}_{fe}$ may be a misleading measure of the treatment effect under the standard common trends assumption, if the treatment effect is heterogeneous across groups and time periods. In the corollary below, we propose two robustness measures that can be used to assess how serious that concern is.
Those robustness measures are defined conditional on $\bm{D}$, the vector stacking together the treatments of all the $(g,t)$ cells. Specifically, for all $(g,t)\in\{1,...,G\}\times\{1,...,T\}$, let $\widetilde{\Delta}_{g,t}=E\left(\Delta_{g,t}\middle|\bm{D}\right)$ denote the ATE in cell $(g,t)$ conditional on $\bm{D}$,\footnote{ $\widetilde{\Delta}_{g,t}$ may differ from $E(\Delta_{g,t})$. To see this, let us consider a simple example where $T=2$. Then, under Assumption (ref), one has $\widetilde{\Delta}_{g,t}=E\left(\Delta_{g,t}\middle|D_{g,1},D_{g,2}\right)$. One may for instance have $E\left(\Delta_{g,1}\middle|D_{g,1}=0,D_{g,2}=0\right)< E\left(\Delta_{g,1}\middle|D_{g,1}=1,D_{g,2}=1\right)$, if a group is more likely to be treated if her treatment effect is initially high.} let $\widetilde{\Delta}^{TR}=E\left(\Delta^{TR}\middle|\bm{D}\right)$ denote the ATT conditional on $\bm{D}$, and let $\widetilde{\beta}_{fe}=E\left(\widehat{\beta}_{fe}\middle|\bm{D}\right)$. The first measure we consider is the minimal value of the standard deviation of the $\widetilde{\Delta}_{g,t}$s under which one could have that $\widetilde{\beta}_{fe}$ is of a different sign than $\widetilde{\Delta}^{TR}$. Therefore, this summary measure applies to $\widetilde{\beta}_{fe}$ and $\widetilde{\Delta}^{TR}$, rather than $\beta_{fe}$ and $\delta^{TR}$, the unconditional expectations of $\widehat{\beta}_{fe}$ and $\Delta^{TR}$ on which we have focused so far. However, one can show that when $G$, the number of groups, goes to infinity, $\widetilde{\beta}_{fe}-\beta_{fe}$ and $\widetilde{\Delta}^{TR}-\delta^{TR}$ both converge to 0. So if the number of groups is large, $\widetilde{\beta}_{fe}$ and $\widetilde{\Delta}^{TR}$ should not differ much from $\beta_{fe}$ and $\delta^{TR}$, and our robustness measure “almost” applies to $\beta_{fe}$ and $\delta^{TR}$.
Let
$\sigma(\boldsymbol{\widetilde{\Delta}})$ is the standard deviation of the conditional ATEs, and $\sigma(\boldsymbol{w})$ is the standard deviation of the $\boldsymbol{w}$-weights,\footnote{ One can show that $\sum_{(g,t):D_{g,t}=1}(N_{g,t}/N_1)w_{g,t}=1$.} across the treated $(g,t)$ cells. Let $n=\#\{(g,t):D_{g,t}=1\}$ denote the number of treated cells. For every $i\in\{1,...,n\}$, let $w_{(i)}$ denote the $i$th largest of the weights of the treated cells: $w_{(1)}\geq w_{(2)} \geq ... \geq w_{(n)}$, and let $N_{(i)}$ and $\widetilde{\Delta}_{(i)}$ be the number of observations and the conditional ATE of the corresponding cell. Then, for any $k\in\{1,...,n\}$, let $P_k=\sum_{i\geq k}N_{(i)}/N_1$, $S_k=\sum_{i\geq k} (N_{(i)}/N_1) w_{(i)}$ and $T_k=\sum_{i\geq k}(N_{(i)}/N_1) w^2_{(i)}$.
$\underline{\sigma}_{fe}$ and $\underline{\underline{\sigma}}_{fe}$ can be estimated simply by replacing $\widetilde{\beta}_{fe}$ by $\widehat{\beta}_{fe}$. An estimator of $\underline{\sigma}_{fe}$ can be used to assess the robustness of $\widehat{\beta}_{fe}$ to treatment effect heterogeneity across groups and periods. If $\underline{\sigma}_{fe}$ is close to 0, $\widetilde{\beta}_{fe}$ and $\widetilde{\Delta}^{TR}$ can be of opposite signs even under a small and plausible amount of treatment effect heterogeneity. In that case, treatment effect heterogeneity would be a serious concern for the validity of $\widehat{\beta}_{fe}$. On the contrary, if $\underline{\sigma}_{fe}$ is very large, $\widetilde{\beta}_{fe}$ and $\widetilde{\Delta}^{TR}$ can only be of opposite signs under a very large and implausible amount of treatment effect heterogeneity. Then, treatment effect heterogeneity is less of a concern.
Similarly, if $\underline{\underline{\sigma}}_{fe}$ is close to 0, one may have, say, $\widetilde{\beta}_{fe}>0$, while $\widetilde{\Delta}_{g,t}\leq 0$ for all $(g,t)$, even if the dispersion of the $\widetilde{\Delta}_{g,t}$s across $(g,t)$ cells is relatively small. Notice that $\underline{\underline{\sigma}}_{fe}$ is only defined if at least one of the weights is strictly negative: if all the weights are positive, then one cannot have that $\widetilde{\beta}_{fe}$ is of a different sign than all the $\widetilde{\Delta}_{g,t}$s.
When some of the weights $w_{g,t}$ are negative, $\widehat{\beta}_{fe}$ may still be robust to heterogeneous treatment effects across groups and periods, provided the assumption below is satisfied.
Assumption (ref) requires that the weights attached to the fixed effects estimator be uncorrelated with the conditional ATEs in the treated $(g,t)$ cells. This is often implausible. For instance, groups treated the most are also those with the lowest value of $w_{g,t}$, as shown in Proposition (ref). But those groups could also be those with the largest treatment effect. This would then induce a negative correlation between $\boldsymbol{w}$ and $\boldsymbol{\widetilde{\Delta}}$. The plausibility of Assumption (ref) can be assessed, by looking at whether $\boldsymbol{w}$ is correlated with a predictor of the treatment effect in each $(g,t)$ cell. In the two applications we revisit in Section (ref), this test is rejected.
\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Extension to the first-difference regression}
Instead of Regression (ref), many articles have estimated the first-difference regression defined below:
When $T=2$ and $N_{g,2}/N_{g,1}$ does not vary across $g$, meaning that all groups experience the same growth of their number of units from period 1 to 2, one can show that $\widehat{\beta}_{fe}=\widehat{\beta}_{fd}$. $\widehat{\beta}_{fe}$ differs from $\widehat{\beta}_{fd}$ if $T>2$ or $N_{g,2}/N_{g,1}$ varies across $g$.
We start by showing that a result similar to Theorem (ref) also applies to $\widehat{\beta}_{fd}$. For any $(g,t)\in \{1,...,G\}\times \{2,...,T\}$, let $\varepsilon_{fd,g,t}$ denote the residual of observations in group $g$ and at period $t$ in the regression of $D_{g,t}-D_{g,t-1}$ on period fixed effects, among observations for which $t\geq 2$. For any $g\in \{1,...,G\}$, let $\varepsilon_{fd,g,1}=\varepsilon_{fd,g,T+1}=0$. One can show that if the regressors in Regression (ref) are not perfectly collinear, $$\sum_{(g,t):D_{g,t}=1}\frac{N_{g,t}}{N_1}\left(\varepsilon_{fd,g,t}- \frac{N_{g,t+1}}{N_{g,t}}\varepsilon_{fd,g,t+1}\right)\neq 0.$$ Then we define $$w_{fd,g,t}=\frac{\varepsilon_{fd,g,t}- \frac{N_{g,t+1}}{N_{g,t}}\varepsilon_{fd,g,t+1}}{\sum_{(g,t):D_{g,t}=1}\frac{N_{g,t}}{N_1}\left(\varepsilon_{fd,g,t}- \frac{N_{g,t+1}}{N_{g,t}}\varepsilon_{fd,g,t+1}\right)}.$$
Theorem (ref) shows that under Assumption (ref), $\beta_{fd}$ is equal to a weighted sum of the ATEs in each treated $(g,t)$ cell with potentially some strictly negative weights, just as $\beta_{fe}$. We now characterize the $(g,t)$ cells whose ATEs are weighted negatively by $\beta_{fd}$. To do so, we focus on staggered adoption designs, as outside of this case it is more difficult to characterize those cells. Our characterization relies on the fact that for every $t\in \{2,...,T\}$, $\varepsilon_{fd,g,t}=D_{g,t}-D_{g,t-1}-\left(D_{.,t}-D_{.,t-1}\right)$. $\varepsilon_{fd,g,t}$ is the difference between the change of the treatment in group $g$ between $t-1$ and $t$, and the average change of the treatment across all groups.
Proposition (ref) shows that for all $t\in \{2,...,T-1\}$ such that the increase in the proportion of treated units is larger from $t-1$ to $t$ than from $t$ to $t+1$, the period-$t$ ATE of groups already treated in $t-1$ receives a negative weight. Moreover, if, at period $T$, at least one group becomes treated, the ATE of groups already treated in $T-1$ also receives a negative weight. Therefore, the treatment effect arising at the date when a group starts receiving the treatment does not receive a negative weight, only long-run treatment effects do. Then, negative weights are a concern when instantaneous and long-run treatment effects may differ. Proposition (ref) also shows that the prevalence of negative weights depends on how the number of groups that start receiving the treatment at date $t$ evolves with $t$. Assume for instance that this number decreases with $t$: many groups start receiving the treatment at date 1, a bit less start at date 2, etc., a case hereafter referred to as the “more early adopters” case. Then, if $N_{g,t}$ is constant across $(g,t)$, $D_{.,t}-D_{.,t-1}$ is decreasing in $t$, and all the long-run treatment effects receive negative weights, except maybe those of period $T$ if $D_{.,T}=D_{.,T-1}$. Conversely, assume that the number of groups that start receiving the treatment at date $t$ increases with $t$: few groups start receiving the treatment at date 1, a bit more start at date 2, etc., a case hereafter referred to as the “more late adopters” case. Then, if $N_{g,t}$ is constant across $(g,t)$, $D_{.,t}-D_{.,t-1}$ is increasing in $t$, and only the period-$T$ long-run treatment effects receive negative weights. Overall, negative weights are much more prevalent in the “more early adopters” than in the “more late adopters” case.
We now come back to general sharp designs where the treatment may not follow a staggered adoption. Let $\widetilde{\beta}_{fd}=E\left(\widehat{\beta}_{fd}\middle|\bm{D}\right)$ denote the expectation of $\widehat{\beta}_{fd}$ conditional on the vector of treatment assignments $\bm{D}$. Just as for $\widetilde{\beta}_{fe}$, one can show that the minimal value of $\sigma(\boldsymbol{\widetilde{\Delta}})$ compatible with $\widetilde{\beta}_{fd}$ and $\widetilde{\Delta}^{TR}=0$ is $\underline{\sigma}_{fd}=|\widetilde{\beta}_{fd}|/\sigma(\boldsymbol{w_{fd}}),$ where
is the standard deviation of the $\boldsymbol{w_{fd}}$-weights. One can also show that $\underline{\underline{\sigma}}_{fd}$, the minimal value of $\sigma(\boldsymbol{\widetilde{\Delta}})$ compatible with $\widetilde{\beta}_{fd}$ and $\widetilde{\Delta}_{g,t}$ of a different sign than $\widetilde{\beta}_{fd}$ for all $(g,t)$, has the same expression as $\underline{\underline{\sigma}}_{fe}$, except that one needs to replace the weights $w_{g,t}$ by the weights $w_{fd,g,t}$ in its definition. Estimators of $\underline{\sigma}_{fe}$ and $\underline{\sigma}_{fd}$ (or $\underline{\underline{\sigma}}_{fe}$ and $\underline{\underline{\sigma}}_{fd}$) can then be used to determine which of $\widehat{\beta}_{fe}$ or $\widehat{\beta}_{fd}$ is more robust to heterogeneous treatment effects.
Finally, and similarly to the result shown in Corollary (ref) for $\beta_{fe}$, $\beta_{fd}$ is equal to $\delta^{TR}$ under common trends and the following assumption:
Note that under the common trends assumption, one can jointly test Assumption (ref) and Assumption (ref), the assumption that the weights attached to $\beta_{fe}$ are uncorrelated with the $\Delta_{g,t}$s: if $\widehat{\beta}_{fe}$ and $\widehat{\beta}_{fd}$ are significantly different, at least one of these two assumptions must fail. In the second application we revisit in Section (ref), $\widehat{\beta}_{fe}$ and $\widehat{\beta}_{fd}$ are significantly different.
\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{An alternative estimator}
In this section, we show that it is possible to estimate a well-defined causal effect even if treatment effects are heterogeneous across groups or over time. Let
with $N_S=\sum_{(g,t):t\geq 2,D_{g,t}\ne D_{g,t-1}}N_{g,t}$. $\delta^S$ is the ATE of all switching cells. In staggered adoption designs, $\delta^S$ is the average of the treatment effect at the time when a group starts receiving the treatment, across all groups that become treated at some point.
We now show that $\delta^S$ can be unbiasedly estimated by a weighted average of DID estimators. This result holds under the following supplementary assumptions.
Assumption (ref) is the equivalent of Assumption (ref), for the potential outcome with treatment. It requires that the shocks affecting a group's $Y_{g,t}(1)$ be mean independent of that group's treatment sequence.
Again, Assumption (ref) is the equivalent of Assumption (ref), for the potential outcome with treatment. It requires that between each pair of consecutive periods, the expectation of the outcome with treatment follow the same evolution over time in every group. Assumptions (ref) and (ref) ensure that one can reconstruct the potential outcome that groups leaving the treament between $t-1$ and $t$ would have experienced if they had remained treated. In staggered adoption designs, Assumption (ref) and (ref) are not necessary for identification, because no group leaves the treatment. Together, Assumptions (ref) and (ref) imply that the ATE follows the same evolution over time in every group: $E\left(\Delta_{g,t}\right)=\eta_t+\theta_g$.\footnote{ It should be possible to weaken Assumptions (ref)-(ref), in particular to account for dynamic effects where $\Delta_{g,t}$ may depend on $(D_{g,1},...,D_{g,t-1})$. This would however introduce complications that are beyond the scope of this paper.} This still allows for heterogeneous treatment effects across groups and over time.\footnote{Imposing Assumptions (ref) and (ref) does not change the decompositions obtained in Theorems (ref) and (ref). $Y_{g,t}(1)$ is observed for all the treated $(g,t)$ cells entering these decompositions, so those assumptions do not bring identifying information for those cells.}
The first point of the stable groups assumption requires that between each pair of consecutive time periods, if there is a “joiner” (i.e., a group switching from being untreated to treated), then there should be another group that is untreated at both dates. The second point requires that between each pair of consecutive time periods, if there is a “leaver” (i.e., a group switching from being treated to untreated), then there should be another group that is treated at both dates.
Notice that under Assumption (ref), groups' treatments are not independent, so Assumption (ref) cannot hold. Accordingly, we replace Assumption (ref) by Assumption (ref) below. Assumption (ref) requires that conditional on its own treatments, a group's outcomes be mean independent of the other groups' treatments. It is weaker than Assumption (ref). Assumption (ref) is necessary to show that our estimator is unbiased, but it is not necessary to show that it is consistent. Accordingly, in Section (ref) of the Web Appendix, we show that our estimator is consistent under Assumption (ref). For every $g\in \{1,...,G\}$, let $\bm{D}_g=(D_{1,g},...,D_{T,g})$.
We can now define our estimator. For all $t\in\{2,...,T\}$ and for all $(d,d')\in \{0,1\}^2$, let
denote the number of observations with treatment $d'$ at period $t-1$ and $d$ at period $t$. Let
Note that $\text{DID}_{+,t}$ is not defined when there is no group such that $D_{g,t}=1,D_{g,t-1}=0$, or no group such that $D_{g,t}=0,D_{g,t-1}=0$. In such instances, we let $\text{DID}_{+,t}=0$. Similarly, let $\text{DID}_{-,t}=0$ when there is no group such that $D_{g,t}=1,D_{g,t-1}=1$ or no group such that $D_{g,t}=0,D_{g,t-1}=1$. Finally, let $$\text{DID}_{\text{M}} = \sum_{t=2}^{T}\left(\frac{N_{1,0,t}}{N_S}\text{DID}_{+,t}+\frac{N_{0,1,t}}{N_S}\text{DID}_{-,t}\right).$$
In Section (ref) of the Web Appendix, we also show that when $G$ goes to infinity, $\text{DID}_{\text{M}}$ is a consistent and asymptotically normal estimator of $\delta^S$. The $\text{DID}_{\text{M}}$ estimator is computed by the fuzzydid and did_multiplegt Stata packages.
Here is the intuition underlying Theorem (ref). $\text{DID}_{+,t}$ compares the evolution of the mean outcome between $t-1$ and $t$ in two sets of groups: the joiners, and those remaining untreated. Under Assumptions (ref) and (ref), $\text{DID}_{+,t}$ estimates the joiners' treatment effect. Similarly, $\text{DID}_{-,t}$ compares the evolution of the outcome between $t-1$ and $t$ in two sets of groups: those remaining treated, and the leavers. Under Assumptions (ref) and (ref), it estimates the leavers' treatment effect. Finally, $\text{DID}_{\text{M}}$ is a weighted average of those DIDs. Note that in staggered designs, there are no groups whose treatment decreases over time, so $\text{DID}_{\text{M}}$ is only a weighted average of the $\text{DID}_{+,t}$ estimators. Note also that one can separately estimate the joiners' and the leavers' treatment effect, by computing separately weighted averages of the $\text{DID}_{+,t}$ and $\text{DID}_{-,t}$ estimators. The former estimator only relies on Assumptions (ref) and (ref), while the latter only relies on Assumptions (ref) and (ref).
$\text{DID}_{\text{M}}$ is related to two other estimators. First, it is related to the Wald-TC estimator in point 2 of Theorem S1 in the Web Appendix of deChaisemartin15b, but the weighting of $\text{DID}_{+,t}$ and $\text{DID}_{-,t}$ therein differs. As a result, $\text{DID}_{\text{M}}$ estimates $\Delta^S$ under weaker assumptions. $\text{DID}_{\text{M}}$ is also related to the multi-period DID estimator in imai2018. However, the multi-period DID estimator is a weighted average of the $\text{DID}_{+,t}$, so it does not estimate the leavers' treatment effect, and applies to a smaller population. Besides, imai2018 do not establish the properties of their estimator. Finally, they do not generalize it to non-binary treatments, something we do in Section (ref) of the Web Appendix.
There may be a bias-variance trade-off between $\text{DID}_{\text{M}}$ and the two-way fixed effects regression estimators. For instance, assume that Regression (ref) is correctly specified:
Then, if the errors $\varepsilon_{g,t}$ are homoskedastic and uncorrelated, it follows from the Gauss-Markov theorem that $\widehat{\beta}_{fe}$ is the linear estimator of $\delta$, the constant treatment effect parameter, with the lowest variance. As $\text{DID}_{\text{M}}$ is also an unbiased linear estimator of $\delta$, the variance of $\widehat{\beta}_{fe}$ must be lower than that of $\text{DID}_{\text{M}}$. With heteroskedastic or correlated errors, one can construct examples where the variance of $\widehat{\beta}_{fe}$ is higher than that of $\text{DID}_{\text{M}}$, but this still suggests that $\text{DID}_{\text{M}}$ may often have a larger variance than that of $\widehat{\beta}_{fe}$, as we find in our applications in Section (ref).
$\text{DID}_{\text{M}}$ uses groups whose treatment is stable to infer the trends that would have affected switchers if their treatment had not changed. This strategy could fail, if switchers experience different trends than groups whose treatment is stable. To assess if this is a serious concern, we propose to use the following placebo estimator, that essentially compares the outcome's evolution from $t-2$ to $t-1$, in groups that switch and do not switch treatment between $t-1$ and $t$. This placebo estimator is defined under a modified version of Assumption (ref).
For all $t\in\{2,...,T\}$ and for all $(d,d',d'')\in \{0,1\}^3$, let
denote the number of observations with treatment status $d''$ at period $t-2$, $d'$ at period $t-1$, and $d$ at period $t$. Let
When there is no group such that $D_{g,t}=1,D_{g,t-1}=D_{g,t-2}=0$ or no group such that $D_{g,t}=D_{g,t-1}=D_{g,t-2}=0$, we let $\text{DID}^{\text{pl}}_{+,t}=0$, and we adopt the same convention for $\text{DID}^{\text{pl}}_{-,t}=0$. Let
$\text{DID}^{\text{pl}}_{+,t}$ compares the evolution of the mean outcome from $t-2$ to $t-1$ in two sets of groups: those untreated at $t-2$ and $t-1$ but treated at $t$, and those untreated at $t-2$, $t-1$, and $t$. If Assumptions (ref) and (ref) hold, then $E\left[\text{DID}^{\text{pl}}_{+,t}\right]=0$. Similarly, if Assumptions (ref) and (ref) hold, $E\left[\text{DID}^{\text{pl}}_{-,t}\right]=0$. Then, $E\left[\text{DID}_{\text{M}}^{\text{pl}}\right]=0$ is a testable implication of Assumptions (ref), (ref), (ref), and (ref), so finding $\text{DID}_{\text{M}}^{\text{pl}}$ significantly different from 0 would imply that those assumptions are violated: groups that switch treatment experience different trends before that switch than the groups used to reconstruct their counterfactual trends when they switch.\footnote{See also callaway2018, who propose another placebo test in staggered adoption designs.} Note that $\text{DID}_{\text{M}}^{\text{pl}}$ compares the trends of switching and stable groups one period before the switch. One can define other placebo estimators comparing those trends, say, two or three periods before the switch. $\text{DID}_{\text{M}}^{\text{pl}}$ and all those other placebo estimators are computed by the did_multiplegt Stata package.
\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Extensions}
In this section, we briefly review some of the extensions in our Web Appendix. First, we show that the decomposition of $\beta_{fe}$ in Theorem (ref) can be extended to fuzzy designs where the treatment varies within $(g,t)$ cells, to applications with a non binary treatment, and to two-way fixed effects regressions with control variables.\footnote{The decomposition of $\beta_{fd}$ in Theorem (ref) can also be extended to all of those cases.} In fuzzy designs or with a non-binary treatment, the weights in Theorem (ref) remain essentially unchanged.
We also consider two-way fixed effects regressions with covariates. Specifically, we study the coefficient of $D_{g,t}$ in a regression of $Y_{i,g,t}$ on group and period fixed effects, $D_{g,t}$, and a vector of covariates $X_{g,t}$. We show that a result very similar to Theorem (ref) applies to that coefficient, up to two differences. First, including covariates allows for different trends across groups, provided those differential trends are fully accounted for by a linear model in $X_{g,t}-X_{g,t-1}$, the change in a group's covariates. Specifically, instead of Assumptions (ref) and (ref), one needs to assume that $$E\left(Y_{g,t}(0)\middle|\bm{D}_g, \bm{X}_g\right) - E\left(Y_{g,t-1}(0)\middle|\bm{D}_g, \bm{X}_g\right)=(X_{g,t}-X_{g,t-1})'\gamma+\lambda_t,$$ for some vector $\gamma$ and constant $\lambda_t$, and where $\bm{X}_g=(X_{g,1},...,X_{g,T})$. Importantly, when the covariates are group-specific linear trends, the equation above is equivalent to $$E\left(Y_{g,t}(0)\middle|\bm{D}_g, \bm{X}_g\right) - E\left(Y_{g,t-1}(0)\middle|\bm{D}_g, \bm{X}_g\right)=\gamma_g+\lambda_t,$$ meaning that from $t-1$ to $t$, the evolution of $Y(0)$ in group $g$ should deviate from its group-specific linear trend $\gamma_g$ by an amount $\lambda_t$ common to all groups. Second, the residual $\varepsilon_{g,t}$ in the weights in Theorem (ref) has to be replaced by $\varepsilon^X_{g,t}$, the residual of observations in cell $(g,t)$ in the regression of $D_{g,t}$ on group and period fixed effects and $X_{g,t}$. Some of the corresponding weights may still be negative, as in Theorem (ref). Overall, two-way fixed effects regressions with covariates may rely on a more plausible common trends assumptions than those without covariates, but they still require that the treatment effect be homogeneous, across time and between groups.
Third, we show that under the common trends assumption and the assumption that the ATE of a $(g,t)$ cell does not change over time, $\beta_{fe}$ and $\beta_{fd}$ identify weighted sums of the ATEs of the $(g,t)$ cells whose treatment changes between $t-1$ and $t$. In sharp designs, the weights attached to $\beta_{fd}$ are all positive, while for $\beta_{fe}$, the same only holds in staggered adoption designs.
Fourth, we show that our $\text{DID}_{\text{M}}$ estimator can easily be extended to non-binary, discrete treatments. Then, we define it as a weighted average of DIDs comparing the evolution of the outcome in groups whose treatment went from $d$ to $d'$ between $t-1$ and $t$ and in groups with a treatment of $d$ at both dates, across all possible values of $d$, $d'$, and $t$.
Finally, our twowayfeweights, fuzzydid, and did_multiplegt Stata packages can handle all of those extensions.
\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Applicability, and applications}
\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Applicability}
We conducted a review of all papers published in the American Economic Review (AER) between 2010 and 2012 to assess the importance of two-way fixed effects regressions in economics. Over these three years, the AER published 337 papers. Out of these 337 papers, 33 or 9.8% of them estimate the FE or FD Regression, or other regressions resembling closely those regressions. When one withdraws from the denominator theory papers and lab experiments, the proportion of papers using these regressions raises to 19.1%.
Table (ref) shows descriptive statistics about the 33 2010-2012 AER papers estimating two-way fixed effects regressions. Panel A shows that 13 use the FE regression; six use the FD regression; six use regressions the FE or FD regression with several treatment variables; three use the FE or FD 2SLS regression discussed in Section (ref) of the Web Appendix; five use other regressions that we deemed sufficiently close to the FE or FD regression to include them in our count.\footnote{For instance, two papers use regressions with three-way fixed-effects instead of two-way fixed effects.} Panel B shows that more than three fourths of those papers consider sharp designs, while less than one fourth consider fuzzy designs. Finally, Panel C assesses whether, in those applications, there are groups whose exposure to the treatment remains stable between each pair of consecutive time periods, the condition that has to be met to be able to compute the $\text{DID}_{\text{M}}$ estimator. For about a half of the papers, reading the paper was not enough to assess this with certainty. We then assessed whether they presumably have stable groups or not. Overall, 12 papers have stable groups, 14 presumably have stable groups, five presumably do not have stable groups, and two do not have stable groups.
In Section (ref) of the Web Appendix, we review each of the 33 papers. We explain where two-way fixed effects regressions are used in the paper, and we detail our assessment of whether the design is a sharp or a fuzzy design, and of whether the stable groups assumption holds or not.
\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Application to gentzkow2011}
gentzkow2011 study the effect of newspapers on voters' turnout in US presidential elections between 1868 and 1928. They regress the first-difference of the turnout rate in county $g$ between election years $t-1$ and $t$ on state-year fixed effects and on the first difference of the number of newspapers available in that county. This corresponds to Regression (ref), with state-year fixed effects as controls. As reproduced in Table (ref) below, gentzkow2011 find that $\widehat{\beta}_{fd}=0.0026$ (s.e.= $9\times 10^{-4}$). According to this regression, one more newspaper increased voters' turnout by 0.26 percentage points. On the other hand, $\widehat{\beta}_{fe}=-0.0011$ (s.e.= $0.0011$). $\widehat{\beta}_{fe}$ and $\widehat{\beta}_{fd}$ are significantly different (t-stat=2.86).
We use the twowayfeweights Stata package, downloadable with its help file from the SSC repository, to estimate the weights attached to $\widehat{\beta}_{fe}$. 6,212 are strictly positive, 4,161 are strictly negative. The negative weights sum to -0.53. $\widehat{\underline{\sigma}}_{fe}=3\times 10^{-4}$, meaning that $\beta_{fe}$ and the ATT may be of opposite signs if the standard deviation of the ATEs across all the treated $(g,t)$ cells is equal to $0.0003$.\footnote{The number of newspapers is not binary, so strictly speaking, in this application the parameter of interest is the average causal response parameter introduced in Section (ref) of our Web Appendix, rather than the ATT.} $\widehat{\underline{\underline{\sigma}}}_{fe}=7\times 10^{-4}$, meaning that $\beta_{fe}$ may be of a different sign than the ATEs of all the treated $(g,t)$ cells if the standard deviation of those ATEs is equal to $0.0007$. We also estimate the weights attached to $\widehat{\beta}_{fd}$. 5,472 are strictly positive, and 4,605 are strictly negative. The negative weights sum to -1.43. $\widehat{\underline{\sigma}}_{fd}=4\times 10^{-4}$, and $\widehat{\underline{\underline{\sigma}}}_{fd}=6\times 10^{-4}$.
Therefore, $\beta_{fe}$ and $\beta_{fd}$ can only receive a causal interpretation if the weights attached to them are uncorrelated with the intensity of the treatment effect in each county$\times$election-year cell (Assumptions (ref) and (ref), respectively). This is not warranted. First, as $\widehat{\beta}_{fe}$ and $\widehat{\beta}_{fd}$ significantly differ, Assumptions (ref) and (ref) cannot jointly hold. Moreover, the weights attached to $\widehat{\beta}_{fe}$ and $\widehat{\beta}_{fd}$ are correlated with variables that are likely to be themselves associated with the intensity of the treatment effect in each cell. For instance, the correlation between the weights attached to $\widehat{\beta}_{fd}$ and $t$, the year variable, is equal to $-0.06$ (t-stat=-3.28). The effect of newspapers may be different in the last than in the first years of the panel. For instance, new means of communication, like the radio, appear in the end of the period under consideration, and may diminish the effect of newspapers.\footnote{ In fact, gentzkow2011 analyze the 1868 to 1928 period separately from later periods, because the growth of the radio may have changed newspapers' effects.} This would lead to a violation of Assumption (ref).
The stable groups assumption holds: between each pair of consecutive elections, there are counties where the number of newspapers does not change. We use the fuzzydid Stata package, downloadable with its help file from the SSC repository, to estimate a modified version of our $\text{DID}_{\text{M}}$ estimator, that accounts for the fact that the number of newspapers is not binary (see section (ref) of our Web Appendix, where we define this modified estimator). We include state-year fixed effects as controls in our estimation. We find that $\text{DID}_{\text{M}}=0.0043$, with a standard error of $0.0015$. $\text{DID}_{\text{M}}$ is 66% larger than $\widehat{\beta}_{fd}$, and the two estimators are significantly different at the 10% level (t-stat=1.69). $\text{DID}_{\text{M}}$ is also of a different sign than $\widehat{\beta}_{fe}$.
Our $\text{DID}_{\text{M}}$ estimator only relies on a common trends assumption. To assess its plausibility, we compute $\text{DID}_{\text{M}}^{\text{pl}}$, the placebo estimator introduced in Section (ref).\footnote{Again, we need to slightly modify $\text{DID}_{\text{M}}^{\text{pl}}$ to account for the fact that the number of newspapers is not binary.} As shown in Table (ref) below, our placebo estimator is small and not significantly different from 0, meaning that counties where the number of newspapers increased or decreased between $t-1$ and $t$ did not experience significantly different trends in turnout from $t-2$ to $t-1$ than counties where that number was stable. Our placebo estimator is estimated on a subset of the data: for each pair of consecutive time periods $t-1$ and $t$, we only keep counties where the number of newspapers did not change between $t-2$ and $t-1$. Still, almost 80% of the county $\times$ election-year observations are used in the computation of the placebo estimator. Moreover, when reestimated on this subsample, the $\text{DID}_{\text{M}}$ estimator is very close to the $\text{DID}_{\text{M}}$ estimator in the full sample.
\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{The effect of union membership on wages}
A number of articles have estimated the effect of union membership on wages using panel data and controlling for workers' fixed effects. For instance, jakubson1991 has found a 8.3% union membership premium using that strategy, in a sample of American males from the PSID followed from 1976 to 1980. vella1998whose estimate a similar regression and find similar results, in a sample of young American males from the NLSY followed from 1980 to 1987.\footnote{The fixed effects regression is not the main specification in vella1998whose. The authors favor instead a dynamic selection model.}
We use the data in vella1998whose to compute various estimators of the union wage premium. As union status is often measured with error freeman1984,card1996effect, we discard changes in union status happening twice in three consecutive years. Specifically, for individuals with $D_{i,t-1}=0$, $D_{i,t}=1$, and $D_{i,t+1}=0$, we replace $D_{i,t}$ by 0. Similarly, for individuals with $D_{i,t-1}=1$, $D_{i,t}=0$, and $D_{i,t+1}=1$, we replace $D_{i,t}$ by 1. Doing so, we discard half of the union status changes in the initial data.\footnote{ Keeping the original data does not change much the results presented below, except that the placebo estimator $\text{DID}_{\text{M}}^{\text{pl},2}$ becomes significant.}
We start by estimating a two-way fixed effects regression of wages on union membership with worker and year fixed effects. Table (ref) below shows that $\widehat{\beta}_{fe}=0.107$ (s.e.= $0.030$), a result close to that of the worker fixed effects regressions in jakubson1991 and vella1998whose.
Then, we estimate the weights attached to $\widehat{\beta}_{fe}$. 820 are strictly positive, 196 are strictly negative, but the negative weights only sum to -0.01. Still, $\widehat{\underline{\sigma}}_{fe}=0.097$, meaning that $\beta_{fe}$ and the ATT may be of opposite signs if the standard deviation of the treatment effect across the unionized worker $\times$ year observations is equal to $0.097$, a substantial but still possible amount of heterogeneity. The weights are negatively correlated with workers' years of schooling (correlation =$-0.12$, t-stat =$-1.88$). The union premium may be lower for more educated workers freeman1984unions, as they may be less substitutable than less educated ones. Then, $\widehat{\beta}_{fe}$ may overestimate $\delta^{TR}$, the average union premium across all unionized worker $\times$ year observations. We also find that $\widehat{\beta}_{fd}=0.060$ (s.e.= $0.032$) and that $\widehat{\beta}_{fe}$ and $\widehat{\beta}_{fd}$ significantly differ (t-stat=1.91),\footnote{The standard error of $\widehat{\beta}_{fe}-\widehat{\beta}_{fd}$ is computed with a worker-level clustered bootstrap.} thus casting further doubt on Assumptions (ref) and (ref).
The stable groups assumption holds: between each pair of consecutive years, there are workers whose union membership status does not change. We therefore compute our $\text{DID}_{\text{M}}$ estimator. Table (ref) shows that it is equal to $0.041$ (s.e.$=0.034$). $\text{DID}_{\text{M}}$ is significantly different from $\widehat{\beta}_{fe}$ (t-stat=2.60) and $\widehat{\beta}_{fd}$ (t-stat=2.36).\footnote{The standard errors of $\widehat{\beta}_{fe}-\text{DID}_{\text{M}}$ and $\widehat{\beta}_{fd}-\text{DID}_{\text{M}}$ are computed with a worker-level clustered bootstrap.} As discussed in Section (ref), we can also estimate separately the union premium for workers joining and leaving a union, something that was previously done by freeman1984. The joiners' effect estimate is equal to $0.059$ (s.e.$=0.053$), the leavers' effect is equal to $0.021$ (s.e.$=0.044$), and the two estimates do not significantly differ (t-stat$=0.55$).
$\text{DID}_{\text{M}}$ relies on a common trends assumption. To assess its plausibility, we compute $\text{DID}_{\text{M}}^{\text{pl}}$, the placebo estimator introduced in Section (ref). $\text{DID}_{\text{M}}^{\text{pl}}$ compares the wage growth of workers changing and not changing their union status one period before that change. We also compute $\text{DID}_{\text{M}}^{\text{pl},2}$ and $\text{DID}_{\text{M}}^{\text{pl},3}$, two other placebo estimators performing the same comparison two and three periods before the change. As shown in Table (ref) below, $\text{DID}_{\text{M}}^{\text{pl}}$ is large, positive, and significant (t-stat=2.49). On the other hand $\text{DID}_{\text{M}}^{\text{pl},2}$ and $\text{DID}_{\text{M}}^{\text{pl},3}$ are smaller and insignificant. Workers that become unionized start experiencing a differential positive pre-trend one year before becoming unionized. This differential pre-trend mostly comes from union joiners: for them, the placebo estimator is equal to $0.119$ (s.e.$=0.051$), while for union leavers the placebo is smaller ($0.061$) and insignificant (s.e.$=0.057$). Therefore, the placebos suggest that even the already small and insignificant $\text{DID}_{\text{M}}$ estimator may overestimate the union premium, due to a positive pre-trend. In fact, the estimate of leavers' effect, for which there is no evidence of a pre-trend, is very close to 0. Overall, our results indicate that there may not be a significant union wage premium.
\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Conclusion}
Almost 20% of empirical articles published in the AER between 2010 and 2012 use regressions with groups and period fixed effects to estimate treatment effects. In this paper, we show that under a common trends assumption, those regressions estimate weighted sums of the treatment effect in each group and period. The weights may be negative: in one application, we find that almost 50% of the weights are negative. The negative weights are an issue when the treatment effect is heterogeneous, between groups or over time. Then, one could have that the treatment's coefficient in those regressions is negative while the treatment effect is positive in every group and time period. We therefore propose a new estimator to address this problem. This estimator estimates the treatment effect in the groups that switch treatment, at the time when they switch. It does not rely on any treatment effect homogeneity condition. It is computed by the fuzzydid and did_multiplegt Stata packages. In the two applications we revisit, this estimator is significantly and economically different from the two-way fixed effects estimators.