Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
99,195 characters · 11 sections · 50 citation commands
Treatment Effects in Staggered Adoption Designs with Non-Parallel Trends
\abstract{ This paper considers identifying and estimating causal effect parameters in a staggered treatment adoption setting --- that is, where a researcher has access to panel data and treatment timing varies across units. We consider the case where untreated potential outcomes may follow non-parallel trends over time across groups. This implies that the identifying assumptions of leading approaches such as difference-in-differences do not hold. We mainly focus on the case where untreated potential outcomes are generated by an interactive fixed effects model and show that variation in treatment timing provides additional moment conditions that can be used to recover a large class of target causal effect parameters. Our approach exploits the variation in treatment timing without requiring either (i) a large number of time periods or (ii) requiring any extra exclusion restrictions. This is in contrast to essentially all of the literature on interactive fixed effects models which requires at least one of these extra conditions. Rather, our approach directly applies in settings where there is variation in treatment timing. Although our main focus is on a model with interactive fixed effects, our idea of using variation in treatment timing to recover causal effect parameters is quite general and could be adapted to other settings with non-parallel trends across groups such as dynamic panel data models.}
JEL Codes: C14, C21, C23
Keywords: Treatment Effects, Panel Data, Interactive Fixed Effects, Difference-in-Differences, Treatment Effect Heterogeneity
\onehalfspacing
Exploiting access to panel data is one of the most common, if not the most common, approach for researchers hoping to learn about the causal effect of economic policies (or some other treatment) on some outcome of interest. In a setting with panel data, economists have traditionally been most interested in models that include time-invariant unobserved heterogeneity that may be correlated with the treatment variable rather than, say, models that are primarily driven by lagged outcomes. A main reason for this is that economic theory often suggests models that include variables that may not be observed by the researcher. Following a large, recent literature on panel data approaches to causal inference that are robust to treatment effect heterogeneity, our starting point is a model for untreated potential outcomes such as
where $Y_{it}(0)$ is unit $i$'s untreated potential outcome in time period $t$ (i.e., the outcome that unit $i$ would experience in period $t$ if it did not participate in the treatment), $h_t$ is some (unknown) non-parametric function that can change over time, $\xi_i$ is unit-specific time-invariant unobserved heterogeneity (that may be distributed differently for the treated group relative to the untreated group and is not necessarily restricted to be scalar), and $e_{it}$ are time-varying unobservables; for simplicity, we abstract from observed covariates. This is obviously a challenging model to make much progress with (even under additional conditions on the time-varying unobservables such as being independent of the treatment and $\xi_i$). Moreover, in a large fraction of applications in economics, researchers only have access to a few periods of data with which to estimate a model or recover some target parameters --- this leads to substantial drawbacks for estimation strategies that rely on estimating $\xi_i$ for each unit due to the incidental parameters problem. Thus, it is very common to significantly simplify this model to the following
In this setting, if the distribution of $e_{it}$ is the same across groups, then it is straightforward to difference out the unit fixed effects, $\xi_i$, recover $\theta_t$ (given a large number of cross-sectional units), and hence to recover average treatment effect parameters. In fact, this is exactly the sort of model that leads to difference-in-differences identification strategies which are the dominant approach to causal inference with short panels in economics.\footnote{To be clear, this is not the only viable approach here. See, for example, athey-imbens-2006,chernozhukov-val-hahn-newey-2013 for alternative approaches that can work with short panels though, like difference-in-differences, these approaches require additional auxiliary assumptions relative to the model in (ref).}
However, we emphasize that, even if many theoretical arguments lead to models such as the one in (ref), the extra linearity condition that leads to (ref) is often not implied by economic theory --- despite it being a key requirement for the identification strategy to work. This is an important distinction relative to using a linear projection to estimate a possibly nonlinear conditional expectation where the conditioning variables are observed. In this case, the linear projection model has certain good properties (such as being the best linear approximation to conditional expectation) but, unlike (ref), linearity does not serve an important role as an identification assumption.
In the current paper, we instead consider the following model for untreated potential outcomes
where we have split the time-invariant unobserved heterogeneity, $\xi_i$, into two components, $\eta_i$ and $\lambda_i$, and allow for the effect of some of the components of the unobserved heterogeneity to vary over time. This is an interactive fixed-effects model for untreated potential outcomes. Viewed together, $\lambda_i'F_t$ can vary across units and periods, thus notably weakening the parallel trends assumption that results from the simpler model in (ref). Moreover, there are a number of features of this model that are attractive in the context of policy evaluation. For example, callaway-karami-2023 argue that this model (rather than, say, a unit-specific linear trends model) arises naturally in a setting where a researcher believes that parallel trends holds after conditioning on some other variables, but those variables are not observed by the researcher.\footnote{To give a more specific example, suppose that a researcher was interested in the effect of some treatment on a person's income. Further, suppose that the researcher thinks that parallel trends holds after conditioning on a person's ability. If that researcher were using data from the National Longitudinal Survey, then there are measures of a person's ability (such as the person's score on the Armed Forces Qualification Test (AFQT)) that the researcher could use. On the other hand, if the researcher were using data from the Current Population Survey, there are no obvious measures of ability that are available. Thus, in the first case, the researcher could just directly include AFQT score in the first case in place of $\lambda_i$ in (ref) and rely on a version of conditional parallel trends similar to the case considered in heckman-ichimura-smith-todd-1998; in the second case, however, $\lambda_i$ would be unobserved and that would lead to the type of interactive fixed effects model that we consider in the current paper. Notice that, by contrast, this discussion does not lead to linear trends models (where $\lambda_i'F_t$ is replaced by $\lambda_i t$) as, if ability were observed, it seems very unlikely that any researcher would include it in the model attached to the linear trend term $t$.}
In the current paper, we propose an identification strategy to recover causal effect parameters when (i) untreated potential outcomes are generated by an interactive fixed effects model as in (ref) and (ii) in a setting where there is variation in treatment timing across units (often referred to as staggered treatment adoption). Several other papers (e.g., gobillon-magnac-2016,xu-2017,callaway-karami-2023,imbens-kallus-mao-2021,brown-butts-2022; see below for additional related discussion) have proposed approaches for recovering treatment effect parameters under an interactive fixed effects model for untreated potential outcomes. These papers typically either require (i) the number of time periods to be large (in the sense of growing with the sample size rather than being fixed) or (ii) require additional auxiliary conditions such as exclusion restrictions or assumptions about the time-varying error terms being serially uncorrelated. These types of extra conditions are, however, implausible in many applications in economics, and, moreover, are not required by more common approaches such as difference-in-differences or including unit-specific linear trends. As discussed earlier, long panels are simply not available in a large fraction of applications in economics (and, even when a long panel is available, it may not be credible that a particular model is stable over a long time horizon; see arkhangelsky-etal-2021 for more discussion of this point). Likewise, assumptions that rule out serial correlation in time-varying error terms are often seen to be implausible in applications in economics (for example, this is a main issue discussed in bertrand-duflo-mullainathan-2004). Arguably, these issues at least partially explain the relative lack of popularity of interactive fixed effects models in empirical work in microeconomics.
Regarding staggered treatment adoption, a number of recent papers have considered difference-in-differences approaches to identifying causal effect parameters in the presence of treatment effect heterogeneity and staggered treatment adoption. This literature has pointed out a number of weaknesses of two-way fixed effects regressions (which have been the dominant approach to implementing difference-in-differences identification strategies for many years) in this context (chaisemartin-dhaultfoeuille-2020,borusyak-jaravel-spiess-2022,goodman-2021,sun-abraham-2021) and proposed alternative estimators that circumvent these issues (callaway-santanna-2021,gardner-2022, among others). These papers all treat variation in treatment timing as a nuisance and propose approaches that sidestep issues that are caused by staggered treatment adoption but do not show up in cases where the treatment timing is common across all units.
In this paper, we consider a similar setting --- we consider a case with the same data availability and are interested in the same types of causal effect parameters. However, rather than viewing staggered treatment adoption as a nuisance, we show how staggered treatment adoption can be exploited to recover causal effect parameters under substantially more sophisticated models than are typically used in the difference-in-differences literature. And, in particular, we show that staggered treatment adoption provides the opportunity to recover causal effect parameters under substantial/complex violations of parallel trends. Our approach does not require a large number of time periods nor does our approach require any of the auxiliary assumptions mentioned above. Conceptually, our insight is that, rather than relying on exclusion restrictions or assumptions ruling out serially correlation in the time-varying error terms to introduce additional moment conditions to identify the relevant parameters in a model with a fixed number of time periods, the groups (defined by the timing of treatment) provide an additional source of moment conditions that can be used to identify the relevant parameters in the model. We also emphasize that, although we mainly focus on an interactive fixed effects model for untreated potential outcomes, this same insight could be quite useful in a number of other applications that require extra moment conditions to identify the model.
Our paper is related to a large literature on interactive fixed effects. Foundational work in this literature includes pesaran-2006,bai-2009,ahn-lee-schmidt-2013, among many others. This literature is vast, so we will mostly confine this section to papers at the intersection of the interactive fixed effects literature and the causal inference literature. That being said, we note that, of the papers mentioned above, our approach is most closely related to the strand of the literature that builds on ahn-lee-schmidt-2013 as our approach does not require a large number of time periods and relies on GMM types of identification/estimation arguments.
There are a few recent papers that consider a treatment effects setting in the context of an interactive fixed effects model for untreated potential outcomes. Some of these papers require access to a large number of periods. This includes gobillon-magnac-2016,xu-2017,chan-kwok-2021. Work on the synthetic control method also broadly falls in this category; for example, abadie-diamond-hainmueller-2010 emphasize the connection between the synthetic control method and interactive fixed effects models, and, like the papers mentioned above, the synthetic control method is often rationalized in a setting where the number of time periods is large. Other related work includes athey-bayati-doudchenko-imbens-khosravi-2021,bai-ng-2021; see liu-wang-xu-2021 for more details on these types of approaches. Several recent papers have considered interactive fixed effects models for untreated potential outcomes without relying on arguments where the number of time periods needs to be large (in particular, callaway-karami-2023,imbens-kallus-mao-2021,brown-butts-2022,brown-butts-westerlund-2023). This strand of the literature is most similar to what we consider in the current paper. Unlike these papers, however, we do not require assumptions about time invariance of parameters in the model for untreated potential outcomes (which can allow for particular covariates to be used as instruments) or independence assumptions on the time-varying unobservables (which can allow for outcomes in other time periods to be used as instruments) or other auxiliary conditions. Instead, we are able to generate moment conditions to identify the parameters in the model for untreated potential outcomes from the staggered nature of the treatment adoption.
Our work also builds on recent papers in the difference-in-differences literature (chaisemartin-dhaultfoeuille-2020,goodman-2021,callaway-santanna-2021,sun-abraham-2021,marcus-santanna-2021,wooldridge-2021,gardner-2022,borusyak-jaravel-spiess-2022,dube-girardi-jorda-taylor-2023, among others). Staggered treatment adoption has been a main case that has been considered in this literature. However, that literature has seen staggered treatment adoption as a nuisance and proposed ways to circumvent issues with traditional panel data estimation strategies that arise due to staggered treatment adoption. Notable among these papers, marcus-santanna-2021 point out that, in a staggered treatment adoption setting, identification strategies that are based on parallel trends assumptions can be highly over-identified. Given over-identification, they propose more efficient estimators and tests for parallel trends. Our arguments are related to the over-identification case they consider but we instead allow for the extra moment conditions to be used to identify causal effect parameters while substantially relaxing the parallel trends assumption.
More generally, violations of the parallel trends assumption for untreated potential outcomes is a primary concern in empirical work that utilizes panel data in order to identify causal effect parameters. For example, the vast majority of empirical papers using DID-type identification strategies report event studies that usually include pre-treatment estimates of pseudo-treatment effects with the goal of assessing the credibility of the parallel trends assumption (see freyaldenhoven-hansen-shapiro-2019,roth-2022 for related discussion). When parallel trends seems likely to be violated (or as a robustness check), another common strategy is to introduce individual-specific linear trends in untreated potential outcomes (heckman-hotz-1989,wooldridge-2005,mora-reggio-2019) which allow for certain (though relatively limited) violations of parallel trends. The approach proposed in the current paper generalizes these linear trend models. Another approach is to allow for violations of parallel trends that result in partial identification of treatment effect parameters of interest (e.g., these ideas include things like allowing for (i) “not-too-big” violations of linear trends or (ii) that violations of parallel trends in post-treatment periods are “not-too-different” from violations of parallel trends in pre-treatment periods (manski-pepper-2018,rambachan-roth-2020).
\paragraph{Notation}
We consider a case where there are $\mathcal{T}$ time periods of panel data with $n$ units. We denote a particular time period by $t \in \{1, 2 \ldots, \mathcal{T}\}$. We focus on the case where $\mathcal{T}$ is fixed. We consider the case with a binary treatment $D_{it}$ that is equal to 1 if unit $i$ is treated in time period $t$ and is equal to 0 otherwise. To formalize the idea of staggered treatment adoption, we make the following assumption.
(ref) implies that, once a unit becomes treated, it remains treated in subsequent periods. Staggered treatment adoption is common in many applications in economics. For example, many location-specific (e.g., state-level) policies follow a staggered adoption pattern where policies are implemented in different locations at different points in time while remaining in place in subsequent periods (at least over the relatively short time horizons that are the relevant case for the setting that we consider). Other treatments in economics can be “scarring” in the sense that once a unit becomes treated, the unit “changes state” and is considered to be treated in subsequent periods as well. To give a specific example, sun-abraham-2021 (in the context of discussing dobkin-finkelstein-kluender-notowidigdo-2018) consider the effect of hospitalization (the treatment) on various individual-level economic variables. In this setting hospitalization is scarring in the sense that treated units are not hospitalized year after year, but rather once an individual becomes hospitalized, that individual permanently moves over to being in the treated group. See also chaisemartin-dhaultfoeuille-2022a,callaway-2023 for additional discussion about staggered treatment adoption.
One noteworthy implication of staggered treatment adoption is that we can fully account for a unit's entire treatment history by that unit's “group”. We define a unit's group by the time period when the unit becomes treated, and use the notation $G_i$ for this variable. For units that do not participate in the treatment, we (somewhat arbitrarily) set $G_i = \infty$. If there are units that are already treated in the first period, we drop those units.\footnote{It is without loss of generality that we have access to a never-treated group. In applications where all units eventually become treated, it is not possible to use an identification strategy that relies on having a comparison group (such as difference-in-differences or the strategy that we propose in the current paper) to recover treatment effect parameters in periods after all units become treated. Therefore, our approach in this case (and the approach taken in most other papers with the same sort of data) would be to drop all periods after all units have become treated. In this case, in the last remaining period, there would be one not-yet-treated group which would be the never-treated group that we refer to here.} This is standard in the literature on policy evaluation with panel data; for example, difference-in-differences types of identification arguments drop units treated in the first period because (i) they are not useful for learning about the path of untreated potential outcomes (except under strong additional assumptions) and (ii) it is not possible to recover treatment effect parameters for this group as we never observe their untreated potential outcomes --- the same issues apply to the setting that we consider below. We denote the full set of groups by $\mathcal{G} \subseteq \{2,\ldots,\mathcal{T},\infty\}$. We additionally denote the set of groups that ever participate in the treatment by $\bar{\mathcal{G}} = \mathcal{G} \setminus \infty$. To deal with the interactive fixed effects, we also have to drop some early-treated groups, but this depends on the number of interactive fixed effects in the model (in particular, we need to have access to at least $R+1$ pre-treatment periods where $R$ is the number of interactive fixed effects in the model). We return to this issue below.
Next, we define potential outcomes. Let $Y_{it}(g)$ denote the outcome that unit $i$ would experience in time period $t$ if it were in group $g$, and, for notational convenience, we define $Y_{it}(0)$ as the outcome that unit $i$ would experience in time period $t$ if it did not participate in the treatment in any time period --- we refer to this as a unit's untreated potential outcome. In each time period, the observed outcome is given by the potential outcome corresponding to a unit's actual group. That is, we observe $Y_{it} = Y_{it}(G_i)$. We make the following assumption.
(ref) says that outcomes in pre-treatment periods are not affected by participating in the treatment in subsequent periods. This assumption is common in the literature though we note that it is straightforward to weaken this assumption to “limited anticipation” where outcomes in periods that are “far enough” away from the treatment period are not affected by eventually participating in the treatment. In the current paper, in order to focus on main ideas, we do not consider this extension, but the related ideas in callaway-santanna-2021,sun-abraham-2021,callaway-karami-2023 would immediately apply to our setting (in particular, these arguments would essentially suggest just to “back up” the identification strategy into earlier periods). Finally, for this section, we make an assumption about the sampling process.
(ref) says that we have access to an iid sample across units. In practice, this allows, for example, for the outcomes to be serially correlated. Our identification arguments below apply immediately in settings with clustering, and it is straightforward to extend our inference results to cases with clustering (we provide more details below).
Next, we introduce the parameters that our approach will target below. Our immediate target parameter of interest is the group-time average treatment effect, which, for $t \geq g$ (post-treatment time periods), is defined as
This is the mean difference between treated potential outcomes and untreated outcomes for group $g$ in time period $t$. Group-time average treatment effects have been emphasized in recent work on difference-in-differences and show up as building blocks both for understanding limitations of TWFE regressions (chaisemartin-dhaultfoeuille-2020) and for proposing alternative estimation strategies that circumvent the limitations of TWFE regressions (callaway-santanna-2021). It is important to note that, conditional on $G=g$, $Y_t(g)$ is an observed outcome. However, $Y_t(0)$ is not an observed outcome. Thus, the challenge for identifying $ATT(g,t)$ is in recovering $\mathbb{E}[Y_t(0) | G=g]$.
In the treatment effects literature with panel data, it is common to aggregate $ATT(g,t)$'s into lower-dimensional treatment effect parameters. We briefly discuss the two most popular aggregations, but note that others are also possible (see callaway-santanna-2021 for alternative aggregations). We start with an event-study type of aggregation. First, define $e:=t-g$ which denotes the length of exposure to the treatment; for example, $e=0$ when $t=g$ which is the period that units in group $g$ become treated. Also, define $\mathcal{G}_e := \{ g \in \mathcal{G} | g+e \leq \mathcal{T} \}$, which is the set of groups that are observed to have participated in the treatment for $e$ periods. Then, consider the parameter
which is the average effect of participating in the treatment across units that have been exposed to the treatment for exactly $e$ time periods (note that the time period where exposure to the treatment is equal to $e$ can vary across units).
Next, we consider an aggregation into an overall treatment effect parameter. As a step in this direction, first define
which is the average effect of participating in the treatment that units in group $g$ experienced across all their post-treatment time periods. Then, a natural overall treatment effect parameter is
which is the average effect of participating in the treatment across all units that participated in the treatment in any time period.
Both the event study and overall average treatment effect are commonly reported in applications. For the identification results below, it is important to note that both of these parameters are weighted averages of $ATT(g,t)$'s (with weights that are straightforward to estimate). This implies that, if we can identify each $ATT(g,t)$, then it will be possible to translate $ATT(g,t)$'s into event study parameters or an overall treatment effect parameter if these are the ultimate target parameter(s) for a particular application. Thus, our arguments below focus on identifying group-time average treatment effects. Finally, some of our identification arguments result in the identification of group-time average treatment effects for a subset of groups or time periods; in those cases, the aggregated parameters that we discuss here may only be identified for a restricted set of groups and time periods as well. We defer these sorts of issues to later in the paper.
Given the framework discussed above, we now introduce an interactive fixed effects model for untreated potential outcomes and, subsequently, our approach to identifying group-time average treatment effects in this context. We make the following assumptions:
(ref) says that untreated potential outcomes are generated by an interactive fixed effects model. Because we consider a setting with a fixed number of time periods, we treat $\theta_t$ and $F_t$ as being fixed parameters (or, alternatively, our approach can be seen as conditional on the realizations of $\theta_t$ and $F_t$). In general, because the number of cross-sectional units is large, our strategy will be to consistently estimate (functions of) these parameters. On the other hand, we treat $\eta_i$ and $\lambda_i$ as being random. We also interpret $\eta_i$ and $\lambda_i$ as unobserved heterogeneity that can be distributed differently across groups. These differences in distribution can lead to different levels and trends of untreated potential outcomes for different groups. The literature on interactive fixed effects models often refers to $F_t$ as factors and $\lambda_i$ as factor loadings. For convenience, we sometimes use this terminology below, but for the most relevant applications to our approach (ones with a binary treatment and fixed-$\mathcal{T}$), interpreting $\lambda_i$ as unobserved heterogeneity and $F_t$ as a time-varying effect of unobserved heterogeneity is probably most natural.
The model in (ref) reduces to the sort of TWFE model for untreated potential outcomes that leads to DID identification strategies (see, e.g., blundell-dias-2009) when either $F_t$ is constant across $t$ (in this case the interactive fixed effects term is absorbed into the individual fixed effect $\eta_i$) or if $\lambda$ has the same mean across groups.\footnote{Related to this discussion, we also explicitly include unit and time fixed effects. Earlier work typically noted that two-way fixed effects models were special cases of interactive fixed effects models, but it is common in more recent work to explicitly (and separately) include the two-way structure (see, for example, callaway-karami-2023,brown-butts-2022 for related discussion).} A related side-effect of including explicit unit and time fixed effects is that the dimension of the interactive fixed effect term, is governed by the number of factors that vary over time and by the number of factor loadings whose means vary across groups (see the discussion below for more details).
(ref) says that, if one could observe/condition on the unobserved heterogeneity terms $\eta$ and $\lambda$ then the average untreated potential outcome would be the same across groups. Another way to think about this assumption is that, in terms of generating untreated potential outcomes, the important differences between groups are due to their distribution of $\eta$ and $\lambda$. This sort of assumption is very common in the literature on treatment effects with panel data; see, for example, gobillon-magnac-2016,xu-2017,gardner-2020,callaway-karami-2023. Importantly, these assumptions do not put any structure on how treated potential outcomes are generated. They also allow for units to select into participating in the treatment on the basis of their treated potential outcomes and their unobserved heterogeneity ($\eta$ and $\lambda$).
An implication of (ref) is that
which we use below as a source of moment conditions to identify parameters from the interactive fixed effects model.
(ref) are the main assumptions that we make in the paper (up to some rank conditions discussed in the next sections). Before continuing, it is worth emphasizing what we have not assumed. First, notice from the conditional exogeneity condition in (ref) that correlation in $e_{it}$ across $t$ is not restricted. This rules out strategies that rely on using outcomes in other periods as an additional source of identifying information as in some of the arguments in callaway-karami-2023 and imbens-kallus-mao-2021. Serial correlation in $e_{it}$ is generally thought to be quite prevalent in most of the fixed-$\mathcal{T}$ policy evaluation settings that our approach is relevant to (see, in particular, bertrand-duflo-mullainathan-2004). Second, we do not require any extra conditions on covariates that enter the model as in callaway-karami-2023,brown-butts-2022,brown-butts-westerlund-2023 that allow them to be used as excluded instruments.\footnote{In our case, there are no covariates that are even included in the model. To be clear, it would be straightforward to include covariates in the approach that we propose below. These could be handled in analogous ways to including covariates in typical panel data applications (and are, therefore, in some sense, not very interesting from an econometrics standpoint). This is in contrast to the other approaches mentioned above that exploit restrictions on the covariates (such as restrictions on how the covariates effects can vary over time or by imposing auxiliary models for the covariates) to achieve identification. See (ref) below for additional discussion along these lines.} In contrast to these approaches, below we will exploit the staggered treatment adoption in order to achieve identification.
To start with, we consider a small case that is helpful to understand the identification strategy in the current paper. In the next section, we consider generalizations allowing for (i) more time periods, (ii) more interactive fixed effects, and (iii) more groups. For now, suppose that $R=1$, so that the model in (ref) becomes
In addition, suppose that $\mathcal{T}=4$ and $\mathcal{G} = \{3,4,\infty\}$ so that there is a group that is treated in periods 3, 4, and an untreated group. Here, we focus on identifying $ATT(3,3)$ (the average effect of participating in the treatment for group 3 in period 3). First, notice that
where we use the notation $\Delta Y_t := Y_t - Y_{t-1}$ and where the first equality is just the definition of $ATT(3,3)$, the second equality adds and subtracts $\mathbb{E}[Y_2(0)|G=3]$, which is the mean untreated potential outcome for group 3 in time period 2, and the last equality holds because $Y_3(3)$ and $Y_2(0)$ are observed outcomes for group 3. (ref) highlights that, in order to recover $ATT(3,3)$, the key identification challenge is to recover $\mathbb{E}[\Delta Y_3(0)|G=3]$, that is, how untreated potential outcomes would have changed over time for group 3 had it not become treated in period 3. For this section, we make the following additional assumptions.
(ref) says that the factors change between the first two periods. (ref) says that the mean of $\lambda$ is different between group 4 and the never-treated group (which are the two relevant comparison groups in this section). We discuss both of these assumptions and their practical significance in more detail at the end of this section.
As a first step towards recovering $\mathbb{E}[\Delta Y_3(0) | G=3]$ from (ref), notice that
Similarly,
and this second equation implies that
where this expression holds just by rearranging terms from (ref). Notice that this step uses (ref) to avoid dividing by 0. Thus, from plugging the expression for $\lambda_i$ in (ref) back into (ref), we have that
where we define $\theta_3^* := \Delta \theta_3 - \frac{\Delta F_3}{\Delta F_2}\Delta \theta_2$, $F_3^* := \frac{\Delta F_3}{\Delta F_2}$, and $v_{i3} := \Delta e_{i3} - \frac{\Delta F_3}{\Delta F_2}\Delta e_{i2}$. Taking the expectation of (ref) conditional on being in group 3, we have that
(ref) holds from (ref) because (i) $\Delta Y_{i2}(0)$ is observed for units in group 3 (thus, the untreated potential outcomes on the right hand side of (ref) are observed outcomes for group 3), and (ii) $\mathbb{E}[v_{3} | G=3] = \mathbb{E}\Big[ \mathbb{E}[v_3|\eta, \lambda, G=3] \Big| G=3\Big] = 0$ (which holds by the law of iterated expectations and then by (ref)).
(ref) therefore suggests that identifying $ATT(3,3)$ hinges on identifying the two-dimensional parameter vector $(\theta_{3}^*, F_{3}^*)'$. Towards this end, notice that for groups 4 and $\infty$, which are not-yet-treated in period 3, $\Delta Y_{i3}(0)$ is observed. This suggests the possibility of recovering $(\theta_3^*,F_3^*)'$ using data coming from those groups through (ref). This is the gist of the strategy that we use below. An important complication arises, however, because, by construction, $\Delta Y_{i2}(0)$ is correlated with $v_{i3}$ in (ref) --- this is because $v_{i3}$ contains $\Delta e_{i2}$. This rules out recovering $(\theta_3^*,F_3^*)'$ directly from the regression of $\Delta Y_{i3}$ on $\Delta Y_{i2}$ using units that are not-yet-treated by period 3. This is the same type of issue that shows up in all of the literature on treatment effects in interactive fixed effects in fixed-$\mathcal{T}$ settings; it is the reason that the existing approaches discussed above impose extra conditions to generate additional moment conditions to be able to recover parameters from the interactive fixed effects model and, subsequently, treatment effect parameters.
The key insight of our identification strategy is that, if we have two distinct groups that are not-yet-treated in period 3, those two groups provide two moment conditions that can potentially be used to recover the two parameters $\theta_3^*$ and $F_3^*$. In particular, (ref) implies that, for any group $g$ and for any time period $t$
where the first equality holds by the law of iterated expectations and the second equality holds by (ref). In the setting considered here, this further implies that, for any group $g$
This implies the following moment conditions
$\Delta Y_{i3}(0)$ is not observed for units in group 3, so this moment condition is infeasible to estimate for group 3 (i.e., as expected, group 3 is not useful for recovering $(\theta_3^*,F_3^*)'$); however, $\Delta Y_{i3}(0)$ is observed for units in group 4 and the never-treated group. Thus, there are two available moment conditions and two parameters to estimate which implies that the necessary order condition is satisfied here. From these two moment equations separately for group 4 and the never-treated group (and after some straightforward algebra), it can be shown that
which implies that $F_3^*$ and $\theta_3^*$ are identified.
Before continuing, it is worth briefly commenting on the role that (ref) play in the discussion above. Towards this end, notice that for $t=2,3$,
which follows from the model for untreated potential outcomes in (ref) and similar arguments as are used throughout this section. Notice that these are the expressions that show up in the numerator and in the denominator for $F_3^*$ above, and further notice that these expressions depend on differences in the mean of $\lambda$ across groups. Next, it is clear that $F_3^*$ will not be identified if $\mathbb{E}[\Delta Y_2 | G=\infty] = \mathbb{E}[\Delta Y_2 | G=4]$ (this is the term that shows up in the denominator of the expression for $F_3^*$ above). From (ref), we can see that both assumptions are required for this difference to be non-zero. More intuitively, by construction our setup rules out (i) $F_1=F_2=F_3=F_4$ (otherwise, the interactive fixed effect term would be absorbed into the unit fixed effect $\eta_i$) and also rules out (ii) $\mathbb{E}[\lambda | G=3] = \mathbb{E}[\lambda|G=4] = \mathbb{E}[\lambda | G=\infty]$ (otherwise, the interactive fixed effect term would be absorbed into the time fixed effect $\theta_t$); in either case, given our setup, it would imply that $R=0$ rather than $R=1$. (ref) impose additional requirements relative to this baseline. First, if $F_1=F_2$, then all three groups will have the same trends in outcomes between the first two periods --- and these are the only two pre-treatment periods for group 3. In this case, we would not be able to distinguish between effects of the treatment in group 3 relative to changes in $F_3$ (i.e., differences in mean outcomes across groups in period 3 could arise for either of those reasons). (ref) rules out this case. Second, if $\mathbb{E}[\lambda|G=4] = \mathbb{E}[\lambda|G=\infty]$, then group 4 and the never-treated group will have the same trends in outcomes over time, regardless of changes in $F_t$ over time, and, therefore, would not be useful for learning about changes in $F_t$ in periods after group 3 becomes treated. (ref) rules out this case.
Given that $F_3^*$ and $\theta_3^*$ are identified, it immediately follows that $ATT(3,3)$ is identified and is given by
To conclude this section, notice that the main cost of our approach is that we are not able to identify as many group-time average treatment effects as would be possible using other identification strategies. For example, callaway-karami-2023 suppose that the researcher has access to a time-invariant covariate that does not affect the path of untreated potential outcomes and show that this sort of covariate can be used to generate additional moment conditions to identify the parameters in the model for untreated potential outcomes. Our approach does not require this sort of covariate (which is a key advantage of our approach), but it comes at the cost of only being able to identify $ATT(3,3)$ without recovering $ATT(3,4)$ or $ATT(4,4)$ which would be feasible using the approach in callaway-karami-2023.
In this section, we extend the results above to a setting with more periods, more groups, and allow for more interactive fixed effects. The arguments in this section target recovering $ATT(g,t)$ for a particular group $g \in \bar{\mathcal{G}}$. Given the model in (ref), after taking first differences, we have that
which eliminates the unit fixed effect $\eta_i$. Next, define
where $\Delta Y_i(0)$, $\Delta \theta$, and $\Delta e_i$ are all $(\mathcal{T}-1)$ dimensional vectors, and $\bm{\Delta} \mathbf{F}$ is $(\mathcal{T}-1)\times R$ matrix. Similarly, define
where $\Delta Y_i^{pre(g)}(0)$, $\Delta \theta^{pre(g)}$, and $\Delta e_i^{pre(g)}$ are all $(g-2) \times 1$ vectors and $\bm{\Delta} \mathbf{F}^{pre(g)}$ is a $(g-2)\times R$ matrix. Next, define $p_g := \textrm{P}(G=g)$ and
where the notation $\Big[\ \cdot \ \Big]_{g \in \mathcal{G}}$ indicates that there is one row in the matrix for each group in $\mathcal{G}$. Thus, $\mathbf{\Lambda}$ is a $|\mathcal{G}| \times (R+1)$ matrix where $|\mathcal{G}|$ denotes the cardinality of the set $\mathcal{G}$ (which is the number of groups). By construction, we have that $\textrm{Rank}(\bm{\Delta} \mathbf{F}) = R$ and that $\textrm{Rank}(\mathbf{\Lambda}) = R+1$. For example, if it were the case that $\textrm{Rank}(\bm{\Delta} \mathbf{F}) = (R-1)$, then the unit fixed effect would effectively absorb one of the interactive fixed effects terms and the model could equivalently be re-written with $R-1$ factors. Similarly, if $\textrm{Rank}(\mathbf{\Lambda}) = R$, then the time fixed effect would effectively absorb one of the interactive fixed effects terms and the model could equivalently be re-written with $R-1$ factors (see (ref) for a more detailed explanation). These conditions also impose that $(\mathcal{T}-1) \geq R$ (a restriction on having enough time periods) and that $|\mathcal{G}| \geq (R+1)$ (a restriction on having enough groups); however, we strengthen both of these conditions in the discussion below.
Below, we focus on identifying treatment effects for a particular group $g$ in a particular post-treatment period $t \geq g$. Define $\mathcal{G}^{comp}(g,t) := \{ g' \in \mathcal{G} : t < g' \}$. (i.e., this is the set of groups that have not yet been treated by period $t$). Additionally, define
which is an $|\mathcal{G}^{comp}(g,t)|$ dimensional vector, and also define
which is a $|\mathcal{G}^{comp}(g,t)| \times (R+1)$ matrix. We make the following two assumptions
(ref) generalize (ref) from the case with a single factor in the previous section to the case with $R$ factors considered here. The intuition is similar too. (ref) requires that there be enough variation in the factors themselves in pre-treatment periods (and, like the earlier case, rules out cases where there is limited variation in the factors in early periods but where there is more variation in the factors in post-treatment periods). (ref) requires enough variation in the means of $\lambda$ across the available comparison groups. Besides these similarities, in this case, these conditions also imply some additional restrictions on the groups and time periods for which our approach can recover $ATT(g,t)$. First, (ref) immediately requires that $(g-2) \geq R$. This means that we have enough pre-treatment periods given the number of factors $R$. And, in particular, this means that we can only recover $ATT(g,t)$ for groups for which we observe at least $R+1$ pre-treatment periods (e.g., if $R=2$, then we can only identify treatment effects for groups that become treated in period 4 or later). Second, (ref) implies that we must have that $|\mathcal{G}^{comp}(g,t)| \geq (R+1)$. This limits the periods for which we can recover $ATT(g,t)$. In particular, define $t^{max}(g)$ to be the largest value of $t$ such that $|\mathcal{G}^{comp}(g,t)| \geq R+1$. Then, this condition rules out recovering $ATT(g,t)$ in periods $t > t^{max}(g)$ (i.e., later periods where there is not a large enough set of groups that have not been treated). By extension, for any group that becomes treated in these late periods, we are not able to recover any group-time average treatment effects for that group. Given some number of factors $R$, we denote the set of groups that satisfy both criteria mentioned above (i.e., that both have enough pre-treatment periods and that are treated early enough to have a large enough comparison group) as $\mathcal{G}^\dagger$; that is, $\mathcal{G}^\dagger := \{g \in \mathcal{G} : (R+2) \leq g \leq t^{max}(g) \}$. This is the set of groups for which we will recover $ATT(g,t)$ using our approach.
To give an example, suppose that $R=2$ and that $\mathcal{G} = \{2,\ldots,\mathcal{T},\infty\}$ (relative to the notation above, here we are supposing that we have groups that become treated in every period as well as a never-treated group). In this case, $|\mathcal{G}^{comp}(g,t)| = \mathcal{T} - t + 1$; if $R=2$, then $\mathcal{G}^\dagger = \{4, \ldots , \mathcal{T}-2\}$, and we can recover $ATT(g,t)$ for any group $g \in \mathcal{G}^\dagger$ in periods $g \leq t \leq (\mathcal{T}-2 )$ (i.e., all post-treatment periods for group $g$ except the last two periods: $(\mathcal{T}-1)$ and $\mathcal{T}$).
Next, we provide our main identification arguments. Given the interactive fixed effects model in (ref), notice that
which implies that
where $\mathbf{H}^{pre(g)} := (\mathbf{\Omega}' \bm{\Delta} \mathbf{F}^{pre(g)})^{-1} \mathbf{\Omega}'$ and $\mathbf{\Omega}$ is a known (or consistently estimable) $(g-2) \times R$ matrix with rank $R$ that satisfies $\textrm{Rank}(\mathbf{\Omega}'\bm{\Delta} \mathbf{F}^{pre(g)}) = \textrm{Rank}(\bm{\Delta} \mathbf{F}^{pre(g)})$.\footnote{In practice, there are several leading candidates for $\mathbf{\Omega}$. First, the approach suggested in callaway-karami-2023 effectively sets $\mathbf{\Omega} = [\mathbf{0}_{R\times (g-2)-R}, \mathbf{I}_R ]'$. A second candidate is to use the first $R$ principal components of $\Delta Y_i^{pre(g)}$, so that $\mathbf{\Omega}$ is the probability limit of the first $R$ eigenvectors of $\bm{\Delta} {\mathbf{Y}^{pre(g)}}^{'} \bm{\Delta} \mathbf{Y}$ (here $\bm{\Delta} \mathbf{Y}^{pre(g)}$ is the $n^{comp(g,t)} \times (g-2)$ data matrix of pre-treatment outcomes where $n^{comp(g,t)}$ is the number of units in the comparison group for group $g$ in period $t$). There are tradeoffs to different choices here. For example, the first choice mentioned above for $\mathbf{\Omega}$ results in only using the most recent $R$ periods of $\Delta Y_{it}(0)$ on the right-hand side of (ref) rather than all pre-treatment periods. callaway-karami-2023 argue that using fewer pre-treatment periods has an advantage of being more robust to the model for untreated potential outcomes not holding in periods further away from the treatment. But it comes at the cost of requiring that $
'$ has rank $R$ (which strengthens \Cref{ass:rank-delta-F}); note that this approach could be tweaked to ``select'' $\Delta F_t$ in periods where the researcher is confident that the rank condition holds for that subset of periods. On the other hand, using the principal components uses data from all periods, but $\mathbf{\Omega}$ must be estimated and there may be other auxiliary conditions needed for $Rank\big(\mathbf{\Omega}'\bm{\Delta} \mathbf{F}^{pre(g)}\big) = R$. The most appropriate choice for $\mathbf{\Omega}$ may vary across different applications.} \Cref{eqn:lambda_i} holds from \Cref{eqn:DelY_pre} by multiplying by $\mathbf{H}^{pre(g)}$ and then cancelling and re-arranging terms. Moreover, for some particular post-treatment period $t$, we have that
where
Notice that $\theta^*(g,t)$ is scalar, $F^*(g,t)$ is a $R \times 1$ vector and $v_i(g,t)$ is scalar. Similar to the simpler case discussed above, we will target estimating $\theta^*(g,t)$ and $F^*(g,t)$ below. Because $\mathbf{\Omega}$ is known, $\widetilde{\Delta Y}_i^{pre(g)}$ amounts to being an $R$ dimensional vector of regressors in (ref). Furthermore, for any group $g' \in \mathcal{G}$, $\mathbb{E}[v_i(g,t)|G=g'] = 0$ (which holds under (ref)). Similar to the simpler case discussed above, these will be the source of moment conditions that we use below. In order to show that $ATT(g,t)$ is identified, we proceed in two steps. First, we show that $ATT(g,t)$ is identified if we can identify $\theta^*(g,t)$ and $F^*(g,t)$. Second, we show that (under certain conditions), $\theta^*(g,t)$ and $F^*(g,t)$ are indeed identified. Towards showing the first part, notice that
where the first equality comes from the definition of $ATT(g,t)$, the second equality adds and subtracts $\mathbb{E}[Y_{g-1}(0)|G=g]$, and the third equality holds by (i) plugging in (ref), (ii) because $\mathbb{E}[v_i(g,t) | G=g] = 0$, and (iii) by replacing potential outcomes with their observed counterparts. Given that $\theta^*(g,t)$ and $F^*(g,t)$ are known, this implies that $ATT(g,t)$ is identified and completes the first part of the argument described above.
Next, we move to showing that $\theta^*(g,t)$ and $F^*(g,t)$ are indeed identified. For a group $g' \in \mathcal{G}^{comp}(g,t)$, we have that
This gives one moment condition for each group in $\mathcal{G}^{comp}(g,t)$. Therefore, the order condition is that $|\mathcal{G}^{comp}(g,t)| \geq R+1$. For example, in the case where $R=1$, there are two parameters to identify: $\theta^*(g,t)$ and $F^*(g,t)$ (which is a scalar in this case). Thus, we need to have at least two groups that are not-yet-treated by period $t$. Stacking the above moment conditions, we have that
Then, identification hinges on the matrix
$\mathbf{\Gamma}(g,t)$ is a $|\mathcal{G}^{comp}(g,t)| \times (R+1)$ matrix, and, for relevance to hold, we need that $\textrm{Rank}\big(\mathbf{\Gamma}(g,t)\big) = R+1$. In the next proposition, we show that relevance holds under the combination of (ref).
Next, let $\mathbf{W}(g,t)$ denote a $|\mathcal{G}^{comp}(g,t)| \times |\mathcal{G}^{comp}(g,t)|$ positive definite weighting matrix. The following result shows that $ATT(g,t)$ is indeed identified under the conditions considered in this section.
(ref) is our main identification result in the paper. It formalizes the conditions under which group-time average treatment effects can be identified in settings with staggered treatment adoption when untreated potential outcomes are generated by an interactive fixed effects model.
To conclude this section, we discuss what sort of aggregated parameters can be recovered in our setting. To start with, we consider a version of an event study parameter. As earlier, we denote event time by $e=t-g$. Then, define $\mathcal{G}^\dagger_e := \{ g \in \mathcal{G}^\dagger : g+e \leq \mathcal{T} \}$. Then, consider the parameter
This is an event study parameter that measures the average effect of the treatment at different lengths of exposure to the treatment. However, because $\mathcal{G}^\dagger_e \subseteq \mathcal{G}_e$ (i.e., the set of available groups for which our approach can recover group-time average treatment effects may be smaller than the full set of groups), $ATT^{ES^{\dagger}}(e)$ is, in general, not the same as $ATT^{ES}(e)$ defined earlier; nor is it possible to recover $ATT^{ES}(e)$ using our approach (except in the special case where $R=0$ --- though, as noted earlier, this case essentially reduces to a version of difference-in-differences). That being said, if one were using alternative identification/estimation strategies where $ATT(g,t)$ for a larger set of groups and time periods and wished to compare the resulting event studies, it would be possible to apply the weights here to the full set of $ATT(g,g+e)$ (where $g \in \mathcal{G}_e$ rather than $\mathcal{G}^\dagger_e$). This would essentially zero-out the contribution of $ATT(g,g+e)$ for groups that are not in $\mathcal{G}^\dagger_e$.
Next, consider an aggregation into a single overall average treatment parameter. Towards this end, for groups $g \in \mathcal{G}^\dagger$, define
where $t^{max}(g)$ is the last time period such that $|\mathcal{G}^{comp}(g,t)| \geq R+1$ (i.e., that there are enough groups that are still untreated that the identification strategy can be implemented). This is the average treatment effect for group $g$ in all its post-treatment periods for which we are able to recover $ATT(g,t)$. We can aggregate this into an overall treatment effect parameter by
$ATT^{O^{\dagger}}$ is the average treatment effects across groups among all groups for which our approach is able to recover any group-time average treatment effects. Like $ATT^{ES^{\dagger}}(e)$, $ATT^{O^{\dagger}}$ is not directly comparable to $ATT^O$ defined earlier when there is treatment effect heterogeneity because it does not include as many time periods or groups. However, it does summarize the average effect of the treatment across the periods and groups for which we are able to learn about the average effect of the treatment. Moreover, it can be compared to overall average treatment effects using alternative identification strategies by applying these weights to those group-time average treatment effects.
We conclude this section with several additional remarks.
It is straightforward to estimate $\big(\theta^*(g,t),F^*(g,t)'\big)'$ using the sample analogue of (ref). To conserve on notation in this section, define
Applying the analogue principle to (ref) given a positive definite matrix $\widehat{\mathbf{W}}$, the estimator of $\delta^*(g,t)$ is
where $\mathbb{E}_n[\cdot]$ denotes the sample average operator across $i=1,\dots,n$ and
Let $\hat{p}_g:=\mathbb{E}_n[\mathbf{1}\{G_i=g\}]$. By the definition of the set of all groups $\mathcal{G}$, we have that $p_g > 0$ for all $g \in \mathcal{G}$. From (ref), $ATT(g,t)$ can be re-written as
which, by the analogue and plugin principles, suggests the following estimator:
We impose the following standard assumption.
As each element in $\ell^{comp}$ is bounded, a fourth-moment bound on $\ell^{comp}$ is thus satisfied automatically, i.e, $||\ell_i^{comp}||^4 < \infty$ by construction. In (ref) in (ref), we show that
where
To proceed with the asymptotic distribution of $\widehat{ATT}(g,t)$, we introduce more notation. Define
Next, we move to establishing the joint asymptotic normality of $ATT(g,t)$ across all the periods and groups for which it is identified and to provide a uniform inference procedure as well as the limiting distributions of aggregated treatment effect parameters such as the event study and overall average treatment effect discussed above. Recall that all of the aggregated treatment effect parameters discussed above can be written as and estimated by
respectively, where $\theta$ is generic notation for an aggregated treatment effect parameter, $ATT$ is a vector that stacks $ATT(g,t)$ for $(g,t) \in \mathcal{G}^\dagger \times \{g, \ldots, t^{max}(g)\}$, $w$ is a vector that stacks weights on each $ATT(g,t)$ to aggregate them into $\theta$; and $\hat{\theta}$, $\hat{w}$, and $\widehat{ATT}$ are the corresponding estimators. Also, notice that, in practice, the $w(g,t)$'s (the weights on particular group-time average treatment effects) need to be estimated, and, therefore, we should take into account their estimation effect. Given the conditions discussed above, for all of the aggregated parameters that we consider, the weights are asymptotically linear, which we write generically as
with $\mathbb{E}[\mathcal{W}_i]=0$. As a last piece of additional notation, define $\Psi_i$ as the vector comprising the set $\{\psi_{igt}\}_{(g,t)\in \mathcal{G}^\dagger \times \{g, \ldots, t^{max}(g)\}}$. The following result provides an asymptotic linear representation and an asymptotic normality result.
(ref) and (ref) establish the joint limiting distribution of all of the identified group-time average treatment effects as well as the aggregated treatment effect parameters. These results can be used as the basis for conducting inference by estimating $\mathbb{E}[\Psi \Psi']$ and/or $\sigma^2_w$. Instead of taking this approach, we opt for conducting inference using a multiplier bootstrap procedure that builds on the previous result. For brevity, we focus on inference for the group-time average treatment effects, but analogous results hold for the aggregated treatment effect parameters given the results above. The multiplier bootstrap involves perturbing the influence function, and offers a number of advantages relative to the nonparametric bootstrap in terms of computational speed. To fix ideas, consider some $iid$ mean-zero $\zeta_i$ with unit variance and a finite third moment that is drawn independently of the data, e.g., the standard normal or a binary $(-1,1)$ each with probability 1/2. The bootstrap estimate is obtained using
where $\widehat{\Psi}_i$ is an estimate of $\Psi_i$ that replaces population parameters with estimates. See kline-santos-2012 for more discussion of the multiplier bootstrap.
The next corollary establishes the (asymptotic) validity of the multiplier bootstrap procedure discussed above.
(ref) says that, under standard conditions and conditional on the observed data, $\sqrt{n}(\widehat{ATT}^\star - \widehat{ATT})$ has the same asymptotic distribution as $\sqrt{n}(\widehat{ATT} - ATT)$ in (ref). For each $(g,t)\in (g,t) \in \mathcal{G}^\dagger \times \{g, \ldots, t^{max}(g)\}$, one can construct standard errors for $\widehat{ATT}(g,t)$ using $\sigma_{gt}^\star = (q_{0.75}(g,t) - q_{0.25}(g,t))/(z_{0.75} - z_{0.25}) $ where $q_\tau(g,t)$ denotes the $\tau$'th quantile from the empirical distribution of $\sqrt{n}(\widehat{ATT}^\star(g,t) - \widehat{ATT}(g,t))$ and $z_\tau$ denotes the $\tau$'th quantile of the standard normal distribution.
In this section, we provide Monte Carlo simulations to illustrate the finite sample properties of our proposed estimation strategy. And, in particular, we compare our approach to the one in callaway-karami-2023 and to difference-in-differences and unit specific linear-trends approaches. We generate untreated potential outcomes by
This corresponds to the model in (ref) except for the term $Z_i'\beta$. In our case, because this term is time-invariant, it is absorbed into the unit-fixed effect $\eta_i$, but this is the key term in the estimation strategy proposed in callaway-karami-2023. Following the simulations in callaway-karami-2023, we consider the case where $\lambda_i$ and $Z_i$ are both vectors containing three elements. We also consider the case where $\mathcal{T}=8$ and $\mathcal{G} = \{5,6,7,8,\infty\}$ (so that we have units that become treated in periods 5, 6, 7, 8, and some units that remain untreated in all periods). We assign units to each group with equal probability. For $j=1,2,3$, we take $Z_{ij} \sim_{iid} \mathcal{N}(0,1)$. We also set $\beta_j = 0$ (which implies that it does not actually directly affect untreated potential outcomes though it can still be an additional source of moment conditions for the approach in callaway-karami-2023). Next, for $h \in \{0,1\}$, we set $\eta_i \sim h G_i + \varepsilon_{\eta,i}$ where $\varepsilon_{\eta,i} \sim \mathcal{N}(0,0.1)$. We vary $h$ across simulations --- when $h=0$, the distribution of $\eta_i$ is the same across groups, but otherwise it is different. We set $\lambda_{i1} = 1 + 2G_i + \rho_1 Z_{i1} + \varepsilon_{i1}$, $\lambda_{i2} = 1 - 5G_i + \rho_2 Z_{i2} + \varepsilon_{i2}$, and $\lambda_{i3} = 5 - 10G_i + \rho_3Z_{i3} + \varepsilon_{i3}$ where $\varepsilon_j \sim_{iid} \mathcal{N}(0,1)$ for $j \in \{1,2,3\}$. For all the simulations reported below, we set $\rho_j = 0.2$ for $j=1,2,3,$ which corresponds to the “medium strength instrument” case considered in callaway-karami-2023. Next, we set $\tilde{F}_{1t} = t$, $\tilde{F}_{2t} = (-1)^t t \log(t)$ and $\tilde{F}_{3t} = (-1)^{\mathbf{1}\{t > 5\}} (5-|5-t|)^2$. We consider four types of designs below. In designs labeled “no unobs.\,het.”, we set $F_t=(0,0,0)'$ and $h=0$; in designs that labeled “0 IFE”, we set $F_t=(0,0,0)'$ and $h=1$; for “1 IFE”, we set $F_t=(\tilde{F}_{1t}, 0, 0)'$ and $h=1$; for “2 IFE”, we set $F_t=(\tilde{F}_{1t}, \tilde{F}_{2t}, 0)'$ and $h=1$. We (with a few exceptions) provide results for the overall average treatment effect. For all the simulations below, $n=1000$, the results are based on 1000 Monte Carlo simulations. Finally, in all simulations, we set $Y_{it}(1) = Y_{it}(0)$, so the true value of all treatment effect parameters is equal to 0.
The first set of results is provided in (ref). To start with, notice that difference-in-differences and linear trends approaches seem to perform well in the cases where they are expected to perform well and perform poorly in expected cases --- for difference-in-differences these cases are with no unobserved heterogeneity and when there are 0 interactive fixed effects; for linear trends, these cases are when there is no unobserved heterogeneity, when there are 0 interactive fixed effects, or with 1 interactive fixed effect (this case holds because the first factor is linear in the DGP that we consider here). Similarly, callaway-karami-2023 performs roughly as expected. The estimation strategy generally performs well when the correct number of interactive fixed effects are included in the model. Including too few interactive fixed effects results in very poor performance of the estimation strategy. On the other hand, including too many interactive fixed effects tends to result in somewhat higher RMSE and MAD (as well as conservative inference).
As for our Staggered IFE approach, it performs somewhat better than callaway-karami-2023 --- when the correct number of interactive fixed effects are included, our approach performs well. It has lower mean squared error than CK-2023 in all specifications. In the case with 2 interactive fixed effects, the root mean squared error is notably large though this seems to be driven by a small number of simulations where the estimator performs poorly as the bias is low, the median absolute deviation is small, and the estimator appears to control size well. In the case with one interactive fixed effect, our estimator over-rejects, but in the other cases, when the number of interactive fixed effects is correct, our inference procedure appears to work well.
It is also interesting to consider settings where the researcher includes the wrong number of interactive fixed effects. First, including too few interactive fixed effects can result in very poor performance (one notable exception to this in the results in the table occurs in the case where we included only 1 interactive fixed effect but the true number of interactive fixed effects was 2). On the other hand, including too many tended to mainly result in conservative inference rather than biased estimates (in some cases the bias and RMSE are large, but the estimator typically performs well in terms of median absolute deviation). We leave a formal analysis of the implications of including too many interactive fixed effects to future work, but these simulations suggest an asymmetry between including too many interactive fixed effects relative to too few interactive fixed effects similar to what is noted in moon-weidner-2015 in a somewhat different context.
Next, we provide additional simulation results that vary the differences between groups (in terms of the mean of their interactive fixed effects). We use largely the same data-generating process as above except that we only consider the case where the number of interactive fixed effects is exactly equal to 1. We also set $\lambda_{i1} = 1 + G_i l + \epsilon_{i1}$. We vary $l$ among $\{0.5, 0.1, 0.01, 0.001\}$. Smaller values of $l$ indicate that the groups are more similar to each other in terms of their interactive fixed effects. For these simulations, we only provide results using the Staggered IFE approach considered in the current paper. We vary the number of interactive fixed effects included in the model and report results for the overall average treatment effect as well as event study parameters for $e=0,1,2$.
These results are provided in (ref). The results are most interesting for $l=0.01$ and $l=0.001$. These cases essentially correspond to very small violations of parallel trends, and this is an empirically relevant Monte Carlo simulation where a researcher might be unsure about whether or not the parallel trends assumption holds. It is especially interesting to compare the rows labeled $IFE=0$ (this estimate is essentially a GMM version of difference-in-differences where no interactive fixed effects are included in the model) and $IFE=1$ (which is the correctly specified model). Here, the RMSE and MAD are sometimes smaller when no interactive fixed effects are included in the model. This is especially true for $e=0$ and somewhat less true (particularly for MAD) for $e=1$ or $e=2$. On the other hand, including too few interactive fixed effects tends to result in over-rejection. This suggests an interesting trade-off in terms of including more or less interactive fixed effects, and it is also interesting to note that, in cases where the violation of parallel trends is small, inappropriately imposing parallel trends performs better closer to the time period when the treatment is implemented relative to periods further away from the time period when the treatment is implemented. This makes some sense too as the violations of parallel trends here accumulate over more and more periods as units move further away from the time period in which they were treated.
In this paper, we have introduced a new approach to recovering causal effect parameters in a setting where untreated potential outcomes are generated by an interactive fixed effects model, where the researcher only has access to a few periods of data, and where there is staggered treatment adoption. In our view, interactive fixed effects models are attractive in a large number of applications in economics where it is not clear whether or not it is reasonable to think that the effect of time-invariant unobservables is constant over time (this is evidenced by the ubiquity of pre-testing in difference-in-differences applications in economics). However, their adoption has, at least arguably, been hampered due to either requiring a large number of time periods or relatively strong auxiliary assumptions in order to estimate. The approach that we proposed in the current paper works without requiring a large number of periods or extra assumptions. For the vast majority of empirical applications in economics, the only additional requirement for using our approach is that there is variation in treatment timing across different units. This is common in many applications, and, in fact, has often been framed as a “complication” in many applications; instead, we argue that variation in treatment timing may present a useful opportunity to explore more complicated models in treatment effect applications with panel data.
\printbibliography