EconBase
← Back to paper

Estimating Dynamic Treatment Effects in Event Studies with Heterogeneous Treatment Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

98,978 characters · 24 sections · 94 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimating Dynamic Treatment Effects in Event Studies with Heterogeneous Treatment Effects

\thispagestyle{empty}

abstractTo estimate the dynamic effects of an absorbing treatment, researchers often use two-way fixed effects regressions that include leads and lags of the treatment. We show that in settings with variation in treatment timing across units, the coefficient on a given lead or lag can be contaminated by effects from other periods, and apparent pretrends can arise solely from treatment effects heterogeneity. We propose an alternative estimator that is free of contamination, and illustrate the relative shortcomings of two-way fixed effects regressions with leads and lags through an empirical application.

Keywords: difference-in-differences, two-way fixed effects, pretrend test

Introduction

Rich panel data has fueled a growing literature estimating treatment effects with two-way fixed effects regressions. This body of applied work has prompted a corresponding econometrics literature investigating the assumptions required for these regressions to yield causally interpretable estimates. For example, athey_imbens_staggered_adoption, kirill, callaway_santanna_ssrn2018, double_fe and goodman-bacon_difference--differences_2018 interpret the coefficient on the treatment status when there is treatment effects heterogeneity and variation in treatment timing. Researchers are often also interested in dynamic treatment effects, which they estimate by the coefficients $\mu_{\ell}$ associated with indicators for being $\ell$ periods relative to the treatment, in a specification that resembles the following:

equation[equation omitted — 131 chars of source]

Here $Y_{i,t}$ is the outcome of interest for unit $i$ at time $t$, $E_{i}$ is the time when unit $i$ initially receives the binary absorbing treatment, and $\alpha_{i}$ and $\lambda_{t}$ are the unit and time fixed effects. Units are categorized into different cohorts based on their initial treatment timing. The relative times $\ell=t-E_{i}$ included in ((ref)) cover most of the possible relative periods, but may still exclude some periods.

The first goal of this paper is to uncover potential pitfalls associated with using the estimates of the relative period coefficients $\mu_{\ell}$ as “reasonable” measures of dynamic treatment effects. We decompose $\mu_{\ell}$ to show it can be expressed as a linear combination of cohort-specific effects from both its own relative period $\ell$ and other relative periods; unless strong assumptions regarding treatment effects homogeneity hold, the terms that include treatment effects from other relative periods will not cancel out and will contaminate the estimate of $\mu_{\ell}.$ Importantly, this demonstrates that the widespread practice of using estimates of treatment leads in ((ref)) as a way of testing for parallel pretrends is problematic. roth_pretest_2019, in his survey of the applied literature, notes that checking whether $\mu_{\ell}=0$ for $\ell$ leads of treatment is a common test for pretrends. Our decomposition result implies that such a test would be invalid because the estimate of $\mu_{\ell}$ is affected by both pretrends and treatment effects heterogeneity, thus any test of $\mu_{\ell}=0$ cannot accept or reject the existence of pretrends without further assumptions on treatment effects.

We show how to calculate the weights underlying the linear combination of treatment effects in $\mu_{\ell}$ using an auxiliary regression. This auxiliary regression depends only on the distribution of cohorts and the relative time indicators included in ((ref)). Examining the weights allows researchers to gauge how large the amount of treatment effects heterogeneity needs to be for $\mu_{\ell}$ to be contaminated by treatment effects from other relative periods. Our publicly-available Stata package {eventstudyweights} automates the estimation of these weights using the panel dataset underlying any given specification of ((ref)).

The second goal of this paper is to propose an alternative regression-based method that is more robust to treatment effects heterogeneity than regression ((ref)). For dynamic treatment effects, researchers are usually interested in estimating some average of treatment effects from $\ell$ periods relative to the treatment. Our alternative method estimates the shares of cohort as weights. These weights are more interpretable than the weights underlying regression ((ref)) in the presence of treatment effects heterogeneity, and the resulting weighted average of treatment effects extends beyond a convex combination of treatment effects sloczynski_general_2018. As discussed in Section (ref), using the procedures of callaway_santanna_ssrn2018, our alternative method can also accommodate covariates.

We illustrate both our decomposition results and our alternative method via an empirical application, estimating the dynamic effects of a hospitalization. We follow dobkin_finkelstein_kluender_notowidigdo_aer2018 in using the publicly-available dataset, Health and Retirement Study (HRS), to first estimate two-way fixed effects regressions. We then illustrate our alternative estimation method with this example. Among the outcomes studied by dobkin_finkelstein_kluender_notowidigdo_aer2018, we focus on out-of-pocket medical spending and labor earnings. Our alternative method yields similar big-picture findings as the original paper that uses two-way fixed effects regressions: the earnings decline due to hospitalization is substantial compared to the transitory out-of-pocket spending increase. However, the two-way fixed effects estimates sometimes fall outside the convex hull of the underlying effects. In contrast, estimates using our alternative method, by construction, are guaranteed to be easy-to-interpret because they are weighted averages of the underlying effects, with weights corresponding to cohort shares.

The rest of the paper is organized as follows. In the next subsection, we review the theoretical literature. Section (ref) formally introduces the event study design and discusses our definition in relation to the applied literature. Section (ref) derives the estimands of two-way fixed effects regression, and introduces sufficient assumptions for them to be causally interpretable. Section (ref) develops our alternative estimator. Section (ref) illustrates our results using an empirical example and Section (ref) concludes. All proofs are contained in the Online Appendix.

Related Literature

This paper makes two main contributions within an active literature on the causal interpretations of two-way fixed effects models in settings with staggered treatment adoption (athey_imbens_staggered_adoption; kirill; callaway_santanna_ssrn2018; double_fe; goodman-bacon_difference--differences_2018). Our paper is also related to the traditional literature analyzing non-separable panel and treatment effects models e.g. heckman_ichimura_smith_todd_ecma1998,heckman_ichimura_todd_restud1997,blundell_jeea2004,semiparametric_did,victor_panel.

The first main contribution of our paper is to interpret estimates from two-way fixed effects specifications when researchers include \textquotedblleft dynamic\textquotedblright indicators for time relative to treatment and when treatment effects are heterogeneous across adoption cohorts. We derive our results for a general class of two-way fixed effects specifications where \textquotedblleft dynamic\textquotedblright indicators can be flexibly specified as single relative periods $\ell$ or sets of relative periods $g$ (thus also capturing any \textquotedblleft static\textquotedblright specification where all post-treatment indicators are collected in a single set). This class of specifications encompasses all specifications addressed by athey_imbens_staggered_adoption, kirill, callaway_santanna_ssrn2018, double_fe and goodman-bacon_difference--differences_2018.

As a building block for the causal interpretation of estimates, we define $CATT_{e,\ell}$, the cohort average treatment effects on the treated as the cohort-specific average difference in outcomes relative to never being treated. Our choice of a \textquotedblleft building block\textquotedblright is governed by the counterfactual and the type of heterogeneity of interest. This object coincides with the \textquotedblleft group-time average treatment effect\textquotedblright studied by callaway_santanna_ssrn2018 and is more granular than the building block used by goodman-bacon_difference--differences_2018 that is an average of $CATT_{e,\ell}$ over some relative period range. athey_imbens_staggered_adoption consider an alternate counterfactual to never being treated: being treated at a different time. kirill implicitly assume away heterogeneity across cohorts within a relative period, so their building block reduces to $ATT_{\ell}$. double_fe allow for heterogeneous treatment paths within a cohort over time across \textquotedblleft groups\textquotedblright , thus their building block is at the group level. We defer a discussion of assumptions underlying the causal interpretation of these building blocks to Section (ref).

The second main contribution of our paper is to propose a simple regression-based alternative estimation strategy that produces a more sensible estimand than conventional two-way fixed effects models under heterogeneous treatment effects. Our procedure is most similar to callaway_santanna_ssrn2018, but has the following differences. First, in the setting where there is no never-treated group, our method uses the last cohort to be treated as a control group, whereas callaway_santanna_ssrn2018 use the set of not-yet-treated cohorts. Our method and theirs thus rely on different, but non-nested parallel trends assumptions. Second, our estimation method can be cast as a regression specification and thus may be more familiar to applied researchers. However, a third difference is that the procedure of callaway_santanna_ssrn2018 allows for conditioning on time-varying covariates. double_fe and goodman-bacon_difference--differences_2018 respectively propose alternative estimators and diagnostic tools for estimation of causal effects in staggered settings, but do not consider the estimation of the dynamic path of treatment effects as we do.

Event studies design

In this section we first formalize the “event studies design”. As discussed in Section (ref), based on how this term is deployed in the empirical literature, an event study design is a staggered adoption design where units are treated at different times, and there may or may not be never treated units. It also nests a difference-in-differences design, where units are either first treated at time $t_{0}$ or never treated.

Specifically, we consider a setting with a random sample of $N$ units observed over $T+1$ time periods, where $T$ is fixed. For each $i\in\{0,\dots,N\}$ and $t\in\{0,...,T\}$, we observe the outcome $Y_{i,t}$ and treatment status $D_{i,t}\in\left\{ 0,1\right\} $: $D_{i,t}=1$ if $i$ is treated in period $t$ and $D_{i,t}=0$ if $i$ is not treated in period $t$. Throughout we assume that the observations $\{Y_{i,t},D_{i,t}\}_{t=0}^{T}$ are independent and identically distributed (i.i.d.).

In the general case of event studies we focus on an absorbing treatment such that the treatment status over time is a non-decreasing sequence of zeros and then ones, i.e. $D_{i,s}\leq D_{i,t}$ for $s<t$. We can thus uniquely characterize a treatment path by the time period of the initial treatment, denoted with $E_{i}=\min\left\{ t:D_{i,t}=1\right\} $. If unit $i$ is never treated i.e. $D_{i,t}=0$ for all $t$, we set $E_{i}=\infty$. Based on when they first receive the treatment, we can also uniquely categorize units into disjoint cohorts $e$ for $e\in\{0,\dots,T,\infty\}$, where units in cohort $e$ are first treated at the same time $\{i:E_{i}=e\}$.

We define $Y_{i,t}^{e}$ to be the potential outcome in period $t$ when unit $i$ is first treated in time period $e$. We define $Y_{i,t}^{\infty}$ to be the potential outcome if unit $i$ never receives the treatment, which we call the “baseline outcome”. Since the timing of the initial treatment uniquely characterizes one's treatment path, we can represent the observed outcome for unit $i$ as

align[align omitted — 147 chars of source]

For any treatment that is not absorbing, if we replace the treatment status $D_{i,t}$ with an indicator for ever having received the treatment, the new treatment is absorbing by construction.��Oftentimes the effect of having ever received the treatment is of interest, as it captures the path of treatment effects even though the treatment itself may be transient.�For example, deryugina_aej2017 is interested in the fiscal cost for a county that has been hit by a hurricane. While a hurricane itself may be transient, the impact of having had a hurricane may not be transient, hence why deryugina_aej2017 codes the year of the first hurricane experienced in a county as $E_{i}$.

In the next section, we use the notation developed above to define the treatment effect of an event study design.

Defining treatment effect of an event study design

In an event study design, we define the unit-level treatment effect as the difference between the observed outcome relative to the never-treated counterfactual outcome: $Y_{i,t}-Y_{i,t}^{\infty}$. Recall that $Y_{i,t}^{\infty}$ denotes the potential outcome if unit $i$ never receives the treatment. This particular counterfactual outcome $Y_{i,t}^{\infty}$ is a reasonable “baseline outcome”, though other counterfactual outcomes may be of interest as well. For example, athey_imbens_staggered_adoption also consider the treatment effect relative to the always-treated counterfactual outcome: $Y_{i,t}-Y_{i,t}^{0}$. sianesi_restat2004 defines the unit-level treatment effect to be relative to the not-yet-treated counterfactual outcome: $Y_{i,t}-Y_{i,t}^{e}$ for $e>t$.

When dynamic treatment effects are of interest, empirical researchers commonly report the coefficient estimate $\widehat{\mu}_{\ell}$ associated with indicators for being $\ell$ periods relative to the treatment in regression ((ref)) as an estimate for the average lagged effect. To assess the causal interpretation of $\mu_{\ell}$, we need “building blocks” for its decomposition, which are the average of unit-level treatment effects at a given relative period across units first treated at time $E_{i}=e$, i.e. units in the same cohort $e$. We call this average the cohort-specific average treatment effects on the treated, formally defined below. Later in Section (ref) we use them as building blocks for the interpretation of the relative period coefficients $\mu_{\ell}$ from two-way fixed effects regressions.

defnThe cohort-specific average treatment effect on the treated (CATT) $\ell$ periods from initial treatment is \begin{equation} CATT_{e,\ell}=E[Y_{i,e+\ell}-Y_{i,e+\ell}^{\infty}\mid E_{i}=e]. \end{equation}

Each $CATT_{e,\ell}$ represents the average treatment effect $\ell$ periods from the initial treatment for the cohort of units first treated at time $e$. We shift from calendar time index $t$ to relative period index $\ell$ which denotes the periods since treatment; for cohort $e$, $\ell$ ranges from $-e$ to $T-e$ because we observe at most $e$ periods before the initial treatment and $T-e$ periods after the initial treatment. Relative periods allow us to compare cohorts while holding their exposure to the treatment constant.

Identifying assumptions

With the above definitions, we formalize three potential identifying assumptions for outcomes of interest in our event study design. The first assumption is a generalized form of a parallel trends assumption. The second assumption requires no anticipation of the treatment. The third assumption imposes no variation across cohorts. For each assumption, we first discuss its meaning and then compare it with similar assumptions made in the literature interpreting two-way fixed effects regressions. Later in Section (ref) we interpret the relative period coefficients $\mu_{\ell}$ from two-way fixed effects regressions under different combinations of these assumptions.

assumption(Parallel trends in baseline outcomes.) For all $s\neq t$, the $E[Y_{i,t}^{\infty}-Y_{i,s}^{\infty}\vert E_{i}=e]$ is the same for all $e\in supp(E_{i})$.

If an application includes never-treated units so that $\infty\in supp(E_{i})$, we need to especially consider whether these never-treated units satisfy the parallel trends assumption. Never-treated units are likely to differ from ever-treated units in many ways, and may not share the same evolution of baseline outcomes. If the never-treated units are unlikely to satisfy the parallel trends assumption, then we should exclude them from the estimation to avoid violation of this assumption.

While common in the applied literature, the parallel trends assumption is strong and oftentimes violated. For example, ashenfelter_1978 documented that participants in job training programs experience a decline in earnings prior to the training period (Ashenfelter\textquoteright s dip). The timing of job training is dependent on the evolution of individual's baseline earnings, and this scenario therefore does not satisfy the parallel trends assumption. Proposition (ref) is the only result in this paper derived without this assumption, but there is an active literature studying inference under violations of the parallel trends assumption e.g. roth_rambachan_2020.

Our parallel trends assumption coincides with that of double_fe. One could substitute this assumption with a different identifying assumption that baseline outcomes are mean independent of $E_{i}$ i.e. at each $t$, $E[Y_{i,t}^{\infty}\vert E_{i}=e]$ is the same for all $e\in supp\left(E_{i}\right)$ and in particular is equal to $E[Y_{i,t}^{\infty}]$. This stronger assumption is plausible when the timing of treatment is indeed randomized, which is the assumption used by athey_imbens_staggered_adoption. By taking the “fully dynamic” specification as their DGP, kirill implicitly assume this version of a parallel trends assumption. callaway_santanna_ssrn2018 propose a weaker version that is conditional on covariates. Finally, for a particular estimand, goodman-bacon_difference--differences_2018 identifies a weaker version that only requires a weighted average of $E[Y_{i,t}^{\infty}-Y_{i,s}^{\infty}\vert E_{i}=e]$ (averaged across cohorts) to be zero.

assumption(No anticipatory behavior prior to treatment.) There is no treatment effect in pre-treatment periods i.e. $E[Y_{i,e+\ell}^{e}-Y_{i,e+\ell}^{\infty}\mid E_{i}=e]=0$ for all $e\in supp(E_{i})$ and all $\ell<0$.

Assumption (ref) requires potential outcomes in any $\ell$ periods before treatment to be equal to the baseline outcome on average as in malani_reif_anti_jpube2015 and botosaru_gutierrez_jae2018. This is most plausible if the full treatment paths are not known to units. If they have private knowledge of the future treatment path they may change their behavior in anticipation and thus the potential outcome prior to treatment may not represent baseline outcomes. For example, hendren_consumption shows that knowledge of future job loss leads to decreases in consumption. If the periods with anticipation behavior are known, then we may consider an alternative version of Assumption (ref), which holds for pre-periods in a subset of pre-treatment periods. Depending on the application, it may still be plausible to assume no anticipation until $K$ periods before the treatment.

The no anticipation assumption proposed by athey_imbens_staggered_adoption is a deterministic condition which stipulates that $Y_{i,e+\ell}^{e}=Y_{i,e+\ell}^{\infty}$ for all units $i$ and $e$ and $\ell<0$. By taking the “fully dynamic” specification as their DGP, kirill allow anticipation by including pre-trends indicators in the DGP. callaway_santanna_ssrn2018 and goodman-bacon_difference--differences_2018 implicitly assume no anticipation by using observed outcomes in time periods before the initial treatment as the untreated potential outcomes.

assumption(Treatment effect homogeneity.) For each relative period $\ell$, $CATT_{e,\ell}$ does not depend on cohort $e$ and is equal to $ATT_{\ell}$.

Assumption (ref) requires that each cohort experiences the same path of treatment effects. Treatment effects need to be the same across cohorts in every relative period for homogeneity to hold, whereas for heterogeneity to occur, treatment effects only need to differ across cohorts in one relative period. The assumption of treatment effect homogeneity is therefore strong, and in Section (ref), we describe how it can be violated in applied settings.

Our notion of treatment effect homogeneity does not preclude dynamic treatment effects; it only imposes that cohorts share the same path of treatment effects. The related literature sometimes formulates restrictions on the dynamics of treatment effects as another notion of treatment effect homogeneity. athey_imbens_staggered_adoption propose an assumption that “restricts the heterogeneity of the treatment effects over time,” which implies $CATT_{e,\ell}$ can vary over $e$ but not over $\ell$. kirill refer to one type of treatment effects heterogeneity as “only across the time horizon,” which implies $CATT_{e,\ell}$ can vary over $\ell$ but not over $e$. callaway_santanna_ssrn2018 allow for “arbitrary treatment effect heterogeneity” when $CATT_{e,\ell}$ varies across cohorts and over time. Similarly, double_fe describe treatment effects that may be “heterogeneous across groups and over time periods.” goodman-bacon_difference--differences_2018 allow heterogenous effects to either “vary across units but not over time” or “vary over time but not across units.” The literature has not converged on a single notion of treatment effects heterogeneity with time-varying treatment. Since researchers are interested in dynamic treatment effects when using a “dynamic” specification, we do not restrict the path of treatment effects but rather use “heterogeneity” to describe variation across cohorts only.

Relevance in the applied literature

To gauge the empirical relevance of our results, we survey the estimation methods used by the twelve papers collected by roth_pretest_2019 from three leading economics journals that contain the phrase “event study” in their main text.\footnote{We follow the selection criteria in roth_pretest_2019: the original sample consists of 70 total papers, but is further constrained to these twelve papers with publicly available data and code. The data and code are used to determine exactly the specification estimated in these papers.} From this sample of applied papers, we learn what specifications empirical researchers are actually using when estimating two-way fixed effects regressions. Four papers in this sample consider the simple setting where units either receive their first treatment at the same time or never receive the treatment. The other eight papers in this sample consider the more complex setting where treated units receive their treatment at various times, and there may or may not be never treated units. This observation suggests an event study in the applied literature nests two popular research designs: difference-in-differences design, but also the design where units receive their first treatments at various times, which is our focus.\footnote{This setup is the same as the staggered adoption design proposed by athey_imbens_staggered_adoption, but we keep the term event study because it is common in the applied literature. }

We summarize the main specifications in this sample of twelve applied papers in Table (ref).\footnote{We focus on the first specification underlying the event study estimates in each paper, which we view as a reasonable proxy for the main specification in the paper.} The columns of Table (ref) collect key properties of these specifications. In Section (ref), we introduce a general class of specifications that encompasses all of these specification. The estimates for relative period coefficients $\mu_{\ell}$ from all these papers therefore fall under our analysis in the next section.

These papers demonstrated that event studies are used to address a broad range of research questions. As an example of this literature, bailey_goodman_bacon_aer2015 use the rollout of the first Community Health Centers (CHCs) to study the longer-term health effects of increasing access to primary care. As another example, tewari_aej2014 uses variation in the timing of deregulation across states to estimate the impact of financial development on homeownership.

Estimators from linear two-way fixed effects regression

We consider a two-way fixed effects (FE) regression of the following form, estimated on a panel of $i=1,\dots,N$ units for $t=0,1,\dots,T$ calendar time periods:

equation[equation omitted — 135 chars of source]

Here $Y_{i,t}$ is the outcome of interest for unit $i$ at time $t$, $E_{i}$ is the time for unit $i$ to initially receive a binary absorbing treatment, and $\alpha_{i}$ and $\lambda_{t}$ are the unit and time fixed effects. The set $\mathcal{G}$ collects disjoint sets $g$ of relative periods $\ell\in[-T,T]$. We allow some relative periods to be excluded from the specification and denote the excluded set with $g^{excl}=\{\ell:\ell\not\in\underset{g\in\mathcal{G}}{\bigcup}g\}$. We denote by $\mu_{g}$ the relative period coefficients from regression ((ref)), i.e. the population regression coefficients. Their corresponding OLS estimators are denoted by $\widehat{\mu}_{g}$ respectively.

We are interested in the properties of $\mu_{g}$ when there are variations in the initial treatment timing, and there may or may not be never-treated units. Below in Section (ref) we illustrate how the choice of $\mathcal{G}$ coincides with a large number of specifications encountered in practice such as the “fully dynamic” specification. We next decompose $\mu_{g}$ in terms of $CATT_{e,\ell}$ when various combinations of the three identifying assumptions fail. For Propositions (ref)-(ref), we state the results in terms of the general specification ((ref)). To specialize these results to the “fully dynamic” specification ((ref)), we note the corresponding decomposition would replace bins with $g=\{\ell\}$ for each relative period $\ell$ included in the specification and $g^{excl}$ would contain all excluded relative periods. The decomposition remains unchanged though the summation over $\ell\in g$ simplifies since each $g$ is a singleton. For Proposition (ref), the decomposition further simplifies for the “fully dynamic” specification as discussed below.

Researchers may assume that $\mu_{g}$ can be interpreted as a convex average of $CATT_{e,\ell}$ for periods $\ell\in g$ from its corresponding set $g$; they may further assume the underlying weights have policy-relevant interpretation, e.g. weights depending on proportions of cohorts. For example, bailey_goodman_bacon_aer2015 interpret them as “intention-to-treat effects” of the treatment in a given relative year. However, our results show that $\mu_{g}$ may not represent the parameter of interest without strong assumptions such as treatment effect homogeneity. Section (ref) provides intuition for these negative results, and demonstrate the weights are actually non-linear functions of proportions of cohorts. Section (ref) illustrates how treatment effects heterogeneity invalidates the pretrends test for a simple three-period setting.

Common specifications

Common specifications can be broadly categorized as either “static” or “dynamic”. Static specifications estimate a single treatment effect that is time invariant. In contrast, dynamic specifications allow for non-parametric changes in the treatment effects over time. Within dynamic specifications, researchers also need to address issues of multi-collinearity, and may bin or trim distant relative periods. All of these choices can be written as instances of ((ref)) with the correct specification of $\mathcal{G}$, meaning that our results are applicable for a wide range of specifications employed in the empirical literature.

To clarify how to specify $\mathcal{G}$ in regression ((ref)), we define $D_{i,t}^{\ell}\coloneqq\mathbf{1}\{t-E_{i}=\ell\}$ to be an indicator for unit $i$ being $\ell$ periods away from initial treatment at calendar time $t$. For never-treated units $E_{i}=\infty$, we set $D_{i,t}^{\ell}=0$ for all $\ell$ and all $t$. We can represent the relative period bin indicator as

equation[equation omitted — 117 chars of source]

Static specification. For a “static” specification $\mathcal{G}$ contains a single element equal to $g=[0,T]$. The indicator $\mathbf{1}\{t-E_{i}\in g\}$ is equivalent to an indicator for whether unit $i$ has received its initial treatment by $t$: $\mathbf{1}\{E_{i}\leq t\}$. The “static” specification thus takes the following form

equation[equation omitted — 99 chars of source]

and the corresponding set of excluded relative periods is $g^{excl}=[-T,-1]$.

Dynamic specification. “Dynamic” specifications encompass any specifications of $\mathcal{G}$ where $\mathcal{G}$ contains more than one element, thus treatment effects are allowed to vary over time non-parametrically. In its most flexible form, a “fully dynamic” specification takes the following form

equation[equation omitted — 165 chars of source]

and the corresponding set of excluded relative periods is $g^{excl}=\{-T,\dots,-K-1,-1,L+1,\dots,T\}$.

Excluding some relative periods from the “fully dynamic” specification is necessary to avoid multi-collinearity, either among the relative period indicators $D_{i,t}^{\ell}$, or with the unit and time fixed effects. For example, when there are no never-treated units i.e. $\infty\not\in supp(E_{i})$ but with a panel balanced in calendar time, we need to exclude at least two relative period indicators in $\mathcal{G}$. These collinearities are discussed by kirill: one multi-collinearity comes from the relative period indicators summing to one for every unit $\sum_{\ell\in[-T,T]}D_{i,t}^{\ell}=1$, and the other multi-collinearity comes from the linear relationship between two-way fixed effects and the relative period indicators, namely $t-E_{i}=\ell$.

Excluding relative periods close to the initial treatment is common in practice. Normalizing relative to the period prior to treatment is the most common - six out of the eight papers we survey do so, as reflected in the above specification where we drop $D_{i,t}^{-1}$. The remaining two papers exclude $D_{i,t}^{0}$.

Excluding distant relative periods is however less common (only one of the eight papers we survey does so). Instead researchers “bin” or “trim” distant relative periods. For “binning”, researchers bin distant relative periods into $[-T,-K)$ and $(L,T]$ and estimate a “binned” specification

equation[equation omitted — 247 chars of source]

without excluding any distant relative periods so that $g^{excl}=\{-1\}$. For “trimming”, researchers trim their panel to be balanced in relative periods.

Neither \textquotedblleft binning\textquotedblright nor \textquotedblleft trimming\textquotedblright resolves the issues of contamination discussed below (i.e. the possibility that treatment effects from other periods affect the estimate for a given $\mu_{g}$). We show the contamination issue for the general specification ((ref)) encompasses both practices. For a given coefficient in the dynamic specification, \textquotedblleft trimming\textquotedblright does however mechanically remove any treatment effects from the relative periods \textquotedblleft trimmed\textquotedblright from the specification. For the static specification put forth in kirill, they noted that \textquotedblleft trimming\textquotedblright also does not resolve the contamination issue they identified with the static specification.

Interpreting the coefficients under no assumptions

First we show that without any assumptions, we can write $\mu_{g}$ as a linear combination of differences in trends.

propThe population regression coefficient on relative period bin $g$ is a linear combination of differences in trends from its own relative period $\ell\in g$, from relative periods $\ell\in g'$ belonging to other bins $g'\neq g$ but included in the specification, and from relative periods excluded from the specification $\ell\in g^{excl}$: \begin{align} \mu_{g}= & \sum_{\ell\in g}\sum_{e}\omega_{e,\ell}^{g}\left(E[Y_{i,e+\ell}-Y_{i,0}^{\infty}\vert E_{i}=e]-E[Y_{i,e+\ell}^{\infty}-Y_{i,0}^{\infty}]\right)\\ & +\sum_{g'\neq g,g'\in\mathcal{G}}\sum_{\ell\in g'}\sum_{e}\omega_{e,\ell}^{g}\left(E[Y_{i,e+\ell}-Y_{i,0}^{\infty}\vert E_{i}=e]-E[Y_{i,e+\ell}^{\infty}-Y_{i,0}^{\infty}]\right)\\ & +\sum_{\ell\in g^{excl}}\sum_{e}\omega_{e,\ell}^{g}\left(E[Y_{i,e+\ell}-Y_{i,0}^{\infty}\vert E_{i}=e]-E[Y_{i,e+\ell}^{\infty}-Y_{i,0}^{\infty}]\right). \end{align} We use the superscript $g$ to associate the weight $\omega_{e,\ell}^{g}$ with the coefficient $\mu_{g}$. The weight $\omega_{e,\ell}^{g}$ is equal to the population regression coefficient on $\mathbf{1}\{t-E_{i}\in g\}$ from regressing $D_{i,t}^{\ell}\cdot\mathbf{1}\left\{ E_{i}=e\right\} $ on all bin indicators $\{\mathbf{1}\{t-E_{i}\in g\}\}_{g\in\mathcal{G}}$ included in the specification ((ref)) and two-way fixed effects.

The above proposition is a direct result of regression mechanics. We provide an intuitive derivation for the closed-form expressions for the weights using the classical “omitted variables bias formula” in Section (ref). We defer the formal derivation to the Appendix. Here we mention the following four properties of the weights $\omega_{e,\ell}^{g}$.

itemize• For relative periods of $\mu_{g}$'s own bin i.e. $\ell\in g$, their associated weights as displayed in ((ref)) sum to one $\sum_{\ell\in g}\sum_{e}\omega_{e,\ell}^{g}=1$. • For relative periods belonging to some other bin included in ((ref)) i.e. $\ell\in g'$ for $g'\neq g$ and $g'\in\mathcal{G}$, their associated weights as displayed in ((ref)) sum to zero $\sum_{\ell\in g'}\sum_{e}\omega_{e,\ell}^{g}=0$ for each bin $g'$. • For relative periods not included in $\mathcal{G}$, their associated weights as displayed in ((ref)) sum to negative one $\sum_{\ell\in g^{excl}}\sum_{e}\omega_{e,\ell}^{g}=-1$. • If there are never-treated units i.e. $\infty\in supp(E_{i})$, we have $\omega_{\infty,\ell}^{g}=0$ for all $g$ and $\ell$.

We can easily estimate the weights $\omega_{e,\ell}^{g}$ for any given specification of $\mathcal{G}$ using the following auxiliary regression:

equation[equation omitted — 201 chars of source]

which regresses $D_{i,t}^{\ell}\cdot\mathbf{1}\left\{ E_{i}=e\right\} $ on all bin indicators included in regression ((ref)) and two-way fixed effects.

All of the above properties can be extended to a case where covariates are added to regression ((ref)) by partialling out the covariates before proceeding. In other words, the terms in parentheses in ((ref)), ((ref)) and ((ref)) would be replaced by terms for which the covariates are partialled out. The weights can be estimated by controlling for covariates in regression ((ref)) the same way as they are controlled for in the original regression.

However, covariates complicate the interpretations of $\mu_{g}$ in terms of $CATT_{e,\ell}$ as we describe below in Proposition 2-4. Depending on how covariates are controlled for in regression ((ref)), we may need an additional assumption that the counterfactual trends are linear in the time-varying covariates $X_{i,t}$. We leave a full investigation of the introduction of covariates for future work.

Interpreting the coefficients under parallel trends assumption only

propUnder Assumption (ref) (parallel trends) only, the population regression coefficient on the indicator for relative period bin $g$ is a linear combination of $CATT_{e,\ell\in g}$ as well as $CATT_{e,\ell'}$ from other relative periods $\ell'\not\in g$, with the same weights stated in Proposition (ref): \begin{equation} \mu_{g}=\sum_{\ell\in g}\sum_{e}\omega_{e,\ell}^{g}CATT_{e,\ell}+\sum_{g'\neq g,g'\in\mathcal{G}}\sum_{\ell'\in g'}\sum_{e}\omega_{e,\ell'}^{g}CATT_{e,\ell'}+\sum_{\ell'\in g^{excl}}\sum_{e}\omega_{e,\ell'}^{g}CATT_{e,\ell'}. \end{equation}

Under Assumption (ref), the terms in Proposition (ref) reduce to a linear combination of the causally interpretable building blocks $CATT_{e,\ell}$ as follows:

equation[equation omitted — 218 chars of source]

for $t=e+\ell$. However, two issues for interpretability remain. First, the coefficient $\mu_{g}$ can be written as an average of not only $CATT_{e,\ell}$ from own periods $\ell\in g$, but also $CATT_{e,\ell^{\prime}}$ from other periods. Second, the weights are still non-linear functions of the distribution of cohorts, same as those in ((ref)), and they are not restricted to lie in $[0,1]$.

The properties of these weights as described following Proposition (ref) imply that contamination from other periods wanes once we impose restrictions on treatment effects. In the next two subsections we illustrate how that can happen.

Interpreting the coefficients under parallel trends and no anticipation assumptions

propIf Assumption (ref) (parallel trends) holds and Assumption (ref) (no anticipatory behavior in all periods before the initial treatment) holds, the population regression coefficient $\mu_{g}$ is a linear combination of post-treatment $CATT_{e,\ell'}$ for all $\ell'\geq0$, with the same weights stated in Proposition (ref): \begin{equation} \mu_{g}=\sum_{\ell'\in g,\ell'>0}\sum_{e}\omega_{e,\ell}^{g}CATT_{e,\ell}+\sum_{g'\neq g,g'\in\mathcal{G}}\sum_{\ell'\in g',\ell'>0}\sum_{e}\omega_{e,\ell'}^{g}CATT_{e,\ell'}+\sum_{\ell'\in g^{excl},\ell'>0}\sum_{e}\omega_{e,\ell'}^{g}CATT_{e,\ell'}. \end{equation}

Once we restrict pre-treatment $CATT_{e,\ell\leq0}$ to be zero under the no anticipatory behavior assumption, the expression for $\mu_{g}$ simplifies as terms involving $CATT_{e,\ell\leq0}$ drop out. However, the second term in the expression for $\mu_{g}$ remains unless we further impose treatment effect homogeneity for its summands to cancel out each other. Thus, $\mu_{g}$ may be non-zero for pre-treatment periods even if parallel trends holds.

This result immediately implies a shortcoming of using pre-treatment coefficients (i.e. $\mu_{g}$ where $g$ contains only leads to the treatment $\ell<0$) to test for pretrends. Under the no anticipatory behavior assumption, cohort-specific treatment effects prior to treatment are all zero: $CATT_{e,\ell}=0$ for all $\ell<0$. Therefore, any linear combination of these $CATT_{e,\ell}$ is also zero. However, $\mu_{g}$ is a function of post-treatment $CATT_{e,\ell'\geq0}$ as well, even when $g$ only contains elements with $\ell<0$. We revisit this implication in greater depth in Section (ref). callaway_santanna_ssrn2018 provides alternative tests for pretrends that do not suffer from this drawback.

Sources of treatment effect heterogeneity

Since treatment effects heterogeneity violates Assumption (ref) and can alter how we interpret $\mu_{g}$, it is important to think through when different cohorts likely experience different paths of treatment effect. Such heterogeneity could arise for many reasons. For example, cohorts may differ in their covariates, which affect how they respond to treatment. We will explore a concrete example in our application: if treatment effects differ with age, and there is variation in age across units first treated at different times, we will have heterogeneous effects (see Section (ref) for details). After controlling for covariates, cohorts may still vary in their responses to the treatment if units select their initial treatment timing based on treatment effects. This source of heterogeneity is still compatible with our parallel trends assumption, which only rules out selection in the initial treatment timing based on the evolution of the baseline outcome. In addition to these two sources of heterogeneity, treatment effects may vary across cohorts due to calendar time-varying effects (e.g. macroeconomic conditions could govern the effects on labor market outcomes across cohorts).

Interpreting the coefficients under parallel trends and treatment effect homogeneity

propIf Assumption (ref) (parallel trends) holds and Assumption (ref) (treatment effect homogeneity) holds, then $CATT_{e,\ell}=ATT_{\ell}$ is constant across $e$ for a given $\ell$, and the population regression coefficient $\mu_{g}$ is equal to a linear combination of $ATT_{\ell\in g}$, as well as $ATT_{\ell'\not\in g}$ from other relative periods: \begin{equation} \mu_{g}=\sum_{\ell\in g}\omega_{\ell}^{g}ATT_{\ell}+\sum_{g'\neq g}\sum_{\ell'\in g'}\omega_{\ell'}^{g}ATT_{\ell'}+\sum_{\ell'\in g^{excl}}\omega_{\ell'}^{g}ATT_{\ell'} \end{equation} The weight $\omega_{\ell}^{g}=\sum_{e}\omega_{e,\ell}^{g}$ sums over the weights $\omega_{e,\ell}^{g}$ from Proposition (ref), and is equal to the population regression coefficient from the following auxiliary regression: \begin{equation} D_{i,t}^{\ell}=\alpha_{i}+\lambda_{t}+\sum_{g\in\mathcal{G}}\omega_{e,\ell}^{g}\mathbf{1}\{t-E_{i}\in g\}+\upsilon_{i,t} \end{equation} which regresses $D_{i,t}^{\ell}$ on all bin indicators included in regression ((ref)) and two-way fixed effects.

We note that even under treatment effect homogeneity $\mu_{\ell}$ can still be contaminated by treatment effects from the excluded periods. This contamination, however, can be avoided by adjusting the specification to only exclude periods with zero treatment effect.

For specifications with relative time bins, we note that there can still be contamination from other bins as suggested by the second term of expression ((ref)). A sufficient condition to avoid such contamination would be to group relative periods $\ell'$ into a bin only when their effects are the same since their weights $\omega_{\ell'}^{g}$ would sum to zero.

For the “fully dynamic” specification ((ref)) where all $g$'s are singletons of relative time periods, the weight $\omega_{\ell'}^{\ell}$ is zero for each relative period $\ell'\ne\ell$ that is included in the specification. The decomposition therefore simplifies to

equation[equation omitted — 118 chars of source]

Intuition for contamination

Proposition (ref) demonstrates that even under the assumptions of parallel trends and no anticipation, estimates $\mu_{g}$ can still be contaminated by treatment effects from other periods. In this section we explain the intuition behind why this contamination occurs for the “fully dynamic” specification ((ref)). A decomposition of $\mu_{\ell}$ into a weighted average of $CATT_{e,\ell'}$ demonstrates that contamination is driven by the interaction of two elements: the weights and $CATT_{e,\ell'}$. The weights underlying the contamination are non-linear functions of the distribution of the cohorts. We do not attempt to provide a heuristic for determining the magnitude of the weights, but instead describe how to estimate the weights and later on in Section (ref) how to estimate each $CATT_{e,\ell}$. This allows researchers to directly determine the degree of contamination in their application. Our publicly-available Stata package {eventstudyweights} automates the estimation of these weights using the panel dataset underlying any given specification of ((ref)).

We apply the familiar omitted variable bias (OVB) formula to arrive at our decomposition. In an event study where individuals receive the treatment at different times, the panel can never be balanced in both calendar time and time relative to the initial treatment. As a result, the relative time indicators are still correlated even after controlling for unit and time fixed effects in a two-way fixed effects regression. We use the saturated regression and the OVB formula to illustrate how this correlation leads to contamination. We defer its formal derivation to Appendix (ref).

Under the parallel trends assumption only, the saturated regression is

align[align omitted — 513 chars of source]

where the regressors are cohort fixed effects, time fixed effects, and cohort-specific relative time indicators. Furthermore, let $g^{incl}$ collect the relative time included in ((ref)). The coefficient associated with the cohort-specific relative time indicator $D_{i,t}^{\ell}\cdot\mathbf{1}\left\{ E_{i}=e\right\} $ is the cohort-specific average treatment effects $CATT_{e,\ell}$. To decompose the coefficient $\mu_{\ell}$ from ((ref)) in terms of this saturated regression ((ref)), the OVB formula multiplies the coefficients in the saturated regression ((ref)), $CATT_{e,\ell'}$, with the regression coefficients from ((ref)), $\omega_{e,\ell'}^{\ell}$, which leads to the following decomposition for $\mu_{\ell}$ as a linear combination of $CATT_{e,\ell'}$:

equation[equation omitted — 109 chars of source]

Since $\omega_{e,\ell'}^{\ell}$ is equal to a regression coefficient from ((ref)), we can write it as

equation[equation omitted — 146 chars of source]

Below we briefly comment on each of the three elements in the above expression to highlight how they depend on the distribution of the cohorts. We defer their detailed definitions and derivations to Appendix (ref).

itemize$\sigma_{e,\cdot}$ is a vector of the covariance between cohort $e$ and the other cohorts, namely $Cov\left(\mathbf{1}\{E_{i}=e\},\mathbf{1}\{E_{i}=e'\}\right)$. This term thus scales quadratically in the share of cohort $e$, and is small for small cohorts. • $\mathbf{\Delta}_{t}$ is a matrix of demeaned relative time indicators. The entry that corresponds to cohort $e'$ and relative time indicator $D_{i,t}^{\ell}$ is $E[D_{i,t}^{\ell}\mid E_{i}=e']-\frac{1}{T+1}\mathbf{1}\left\{ e'\in\mathcal{I}_{\ell}\right\} $. When $T$ is large, i.e. the panel is long, the second term is small and this entry is therefore approximately equal to the relative time indicator. • $A_{\ell}^{-1}$ is the row of $A^{-1}$ that corresponds to the relative time indicator $D_{i,t}^{\ell}$ for $A$ the covariance matrix of demeaned relative time indicators. Specifically, the entry in $A$ that corresponds to the covariance between demeaned $D_{i,t}^{\ell}$ and $D_{i,t}^{\ell'}$ is \begin{equation} \sum_{t}Cov\left(D_{i,t}^{\ell},D_{i,t}^{\ell'}\right)-\frac{1}{T+1}Cov\left(\mathbf{1}\left\{ E_{i}\in\mathcal{I}_{\ell}\right\} ,\mathbf{1}\left\{ E_{i}\in\mathcal{I}_{\ell'}\right\} \right). \end{equation} Within any time period $D_{i,t}^{\ell}$ and $D_{i,t}^{\ell'}$ are negatively correlated because no cohorts can be in these two relative times at the same time. The second covariance term is in general also non-zero because being in one cohort predicts (not) being in another cohort. Therefore $A$ is in general not a diagonal matrix and $A^{-1}$ would depend on the distribution of the cohorts non-linearly.

The three elements of ((ref)) demonstrate the weights are non-linear functions of the distribution of the cohorts, and they are in general non-zero. Nonetheless these weights can be estimated easily by the auxiliary regression ((ref)).

Invalidity of pretrend tests based on pre-period coefficients.

Contamination undermines the practice of testing for pretrends using pre-period coefficients. Proposition (ref) implies that when effects are not homogenous across cohorts, it is problematic to interpret non-zero estimates for $\mu_{g}$ as evidence for pretrends, where the set $g$ contains some leads $\ell<0$. Proposition (ref) implies that even with homogeneous treatment effect, if the effects associated with the excluded periods are not zero, then contamination may still occur. Therefore without strong assumptions, pre-period coefficients should not be used to test for pretrends because contamination can lead to estimates that are non-zero in the absence of pretrends or zero in the presence of pre-trends.

Testing for pretrends using pre-period coefficients is commonly used in practice. As an example, he_wang_aej2017 mention \textquotedblleft the estimated coefficients of the leads of treatments, i.e. $\delta_{k}$ for all $k\leq-2$ are statistically indifferent from zero\textquotedblright as evidence for lack of pretrends. As another example, Chetty_qje2014 assert “there is no trend toward higher individual pension contributions prior to year 0 ... as one would expect if individuals\textquoteright tastes for saving were changing around the job switch” based on pre-period coefficient estimates. These tests are only appropriate when the authors are willing to make strong assumptions.

To provide further intuition for why this test is not meaningful without additional assumptions we walk through a simple example of the fully dynamic specification. Consider a balanced panel with $T=2$ and cohorts $E_{i}\in\{1,2\}$. There are at least two multi-colinearities from including all four relative time indicators. To form the fully dynamic specification we include $g^{incl}=\{-2,0\}$ and exclude $g^{excl}=\{-1,1\}$:

equation[equation omitted — 142 chars of source]

The choice of $g^{excl}$ is based on the common practice of normalizing relative to the $-1$ period and distant lags.

When there are no never treated units, we can express the pre-trend coefficient $\mu_{-2}$ in terms of $CATT$s:

align[align omitted — 248 chars of source]

It is apparent the weights maintain the structure described in Proposition (ref). Without any anticipation effect, the effects $CATT_{e,\ell<0}$ are zero and thus we expect $\mu_{-2}$ to be zero regardless of the cohort shares. With homogeneous treatment effect, cohorts 1 and 2 experience the same treatment effect at relative time 0 so that $CATT_{1,0}$ and $CATT_{2,0}$ cancel. But even with homogeneous treatment effect, the last term reflects the role of excluded periods as $CATT_{1,-1}$, $CATT_{2,-1}$ and $CATT_{1,1}$ receive non-zero weights. If there is any lagged effect and $CATT_{1,1}$ is non-zero, the coefficient $\mu_{-2}$ would be non-zero even without any anticipation effect. Note such behavior is independent of the distribution of the two cohorts.

We can further introduce never treated units to our example to demonstrate how the weight $\omega_{e,\ell'}^{-2}$ can be non-linear in the distribution of cohorts while maintaining the structure described in Proposition (ref). In Figure (ref) we plot the weights $\omega_{e,\ell'}^{-2}$ as we vary the distribution of cohorts. Specifically, we vary the total share of treated cohorts (shown on the $x$-axis), holding the shares of cohort 1 and 2 equal to each other and setting the remaining to be the share of never treated units. For any distribution of cohorts, we have $\omega_{2,-2}^{-2}=1$ (not pictured). Panel (a) shows the weights for the included period $\ell'=0$ while panel (b) shows the weights for the excluded periods $\ell'=-1$ or 1.

The example shown in Figure (ref) provides a visualization of three takeaways regarding the contamination in the pre-trend coefficient $\mu_{-2}$. First, both panels show that weights are a non-linear function of cohort shares. Second, panel (a) confirms the structure for weights associated with $\ell'\neq\ell$ but $\ell'\in g^{incl}$ as described in Proposition (ref), namely $\sum_{e}\omega_{e,\ell'}^{-2}=0$ for $\ell'\neq-2$. However, these weights have non-zero magnitude. When the effect is homogenous across cohorts, the contaminations are equal to $\sum_{e}\omega_{e,\ell'}^{-2}ATT_{\ell'}$ and cancel each other out. In contrast, when effects are heterogeneous the different $CATT_{e,\ell'}$ will not necessarily cancel and will contaminate the estimate for $\mu_{-2}$. Third, panel (b) confirms the structure for weights associated with excluded periods as described in Proposition (ref), namely $\sum_{e}\omega_{e,\ell'\in g^{excl}}^{-2}=-1$. In other words, contaminations from excluded periods can be thought of as a type of \textquotedblleft normalization\textquotedblright : a weighted average of excluded $CATT_{e,\ell'}$ is subtracted off the estimated treatment effect. However because these weights are not contained in $[0,-1]$, this average may lie outside of the convex hull of $CATT_{e,\ell'}$ for excluded periods. The latter issue can be alleviated by an assumption that $CATT_{e,\ell'}$ are the same in all excluded periods; we can fully avoid contamination from excluded period treatment effects by assuming all associated $CATT_{e,\ell'}$ are equal to zero.

Alternative estimation method

We propose a new estimation method that is robust to treatment effects heterogeneity. The goal of our method is to estimate a weighted average of $CATT_{e,\ell}$ for $\ell\in g$ with reasonable weights, namely weights that sum to one and are non-negative. In particular, we focus on the following weighted average of $CATT_{e,\ell}$, where the weights are shares of cohorts that experience at least $\ell$ periods relative to treatment, normalized by the size of $g$:

equation[equation omitted — 126 chars of source]

One can aggregate $CATT_{e,\ell}$ to form other parameters of interest, such as those proposed by callaway_santanna_ssrn2018. We focus on the above aggregation $\nu_{g}$ since our goal is to improve the non-convex and non-zero weighting in $\mu_{g}$. The weights in $\nu_{g}$ are guaranteed to be convex and have an interpretation as the representative shares corresponding to each $CATT_{e,\ell}$. Thus, our alternative estimator $\widehat{\nu}_{g}$ improves upon the two-way fixed effects estimator $\widehat{\mu}_{g}$ by estimating an interpretable weighted average of $CATT_{e,\ell\in g}$.

Our method proceeds by replacing each component in $\nu_{g}$ with its consistent estimator. We first estimate each $CATT_{e,\ell}$ using an interacted two-way fixed effects regression, then estimate the weight $Pr\{E_{i}=e\mid E_{i}\in[-\ell,T-\ell]\}$ using their sample analogs. In the final step, we average over the cohort-specific estimates associated with relative period $\ell$. This method has a similar flavor as the method proposed by gibbons_serrato_urbancic_jem2018. They first use an interacted model to estimate the treatment effect for each fixed effect group; the resulting group-specific estimates are averaged to provide the ATE. Their method improves fixed effects regressions in a cross-sectional setting, and our method builds on theirs by improving two-way fixed effects regressions in a panel setting. We therefore follow their terminology in calling our alternative estimator an “interaction-weighted” estimator.

Interaction-weighted estimator

We describe the estimation procedure in three steps (with more detailed definitions stated in Definition (ref) of Online Appendix (ref)).

Step 1. We estimate $CATT_{e,\ell}$ using a linear two-way fixed effects specification that interacts relative period indicators with cohort indicators, excluding indicators for cohorts from some set $C$:

align[align omitted — 196 chars of source]

The exact specification depends on the cohort shares for a given application. If there is a never-treated cohort, i.e. $\infty\in supp\{E_{i}\}$, then we may set $C=\{\infty\}$ and estimate regression ((ref)) on all observations. If there are no never-treated units, i.e. $\infty\not\in supp\{E_{i}\}$, then we may set $C=\{\max\{E_{i}\}\}$, i.e. the latest-treated cohort and estimate regression ((ref)) on observations from $t=0,\dots,\max\{E_{i}\}-1$. Lastly, if there is a cohort that is always treated, i.e. $0\in supp\{E_{i}\}$, then we need to exclude this cohort from estimation.

The coefficient estimator $\widehat{\delta}_{e,\ell}$ from regression ((ref)) is a DID estimator for $CATT_{e,\ell}$ with particular choices of pre-periods and control cohorts. As DID is likely a familiar estimator for applied researchers, we separate the more in-depth discussion of its definition, choices of pre-periods, and choices of control cohorts in Section (ref).

Step 2. We estimate the weights $Pr\{E_{i}=e\mid E_{i}\in[-\ell,T-\ell]\}$ by sample shares of each cohort in the relevant period(s) $\ell\in g$.

Step 3. To form our IW estimator, we take a weighted average of estimates for $CATT_{e,\ell}$ from Step 1 with weight estimates from step 2. More formally, the IW estimator is

equation[equation omitted — 157 chars of source]

where $\widehat{\delta}_{e,\ell}$ is returned from step 1 and $\widehat{Pr}\{E_{i}=e\mid E_{i}\in[-\ell,T-\ell]\}$ is the estimated weight returned from step 2. We normalize the weights further by the size of $g$. If $g$ is a singleton, then its size is $\left|g\right|=1$.

Validity of the IW estimator. Under the parallel trends and no anticipation assumptions the coefficient estimator $\widehat{\delta}_{e,\ell}$ from regression ((ref)) is a consistent estimator for $CATT_{e,\ell}$. The sample shares of each cohort are also consistent estimators for the population shares. Thus, the IW estimator is consistent for a weighted average of $CATT_{e,\ell}$ with weights equal to the share of each cohort in the relevant period(s).

With a few standard assumptions (which we present as Assumption (ref) in Online Appendix (ref)) on regression ((ref)), we can show that each IW estimator is asymptotically normal and derive its asymptotic variance. The large sample approximation allows us to estimate the variance of IW estimators directly without relying on bootstrapping as in callaway_santanna_ssrn2018. However, we only construct pointwise confidence interval valid for a given IW estimator $\widehat{\nu}_{g}$. The bootstrap-based inference by callaway_santanna_ssrn2018 constructs simultaneous confidence intervals that are valid for the entire path of $\widehat{\nu}_{g}$.

Difference-in-differences estimator for $CATT_{e,\ell}$

defnAssume cohort $e$ is non-empty i.e. $\sum_{i=1}^{N}\text{\ensuremath{\mathbf{1}}}\{E_{i}=e\}>0$. Assume there exists some pre-period $s<e$ and some set of control cohorts $C\subseteq\left\{ c:e+\ell<c\leq T\right\} $ that are non-empty i.e. $\sum_{i=1}^{N}\text{\ensuremath{\mathbf{1}}}\{E_{i}\in C\}>0$. Using the notion $\mathbb{E}_{N}$ to abbreviate the symbol $\frac{1}{N}\sum_{i=1}^{N}$, the DID estimator with pre-period $s$ and control cohorts $C$ estimates $CATT_{e,\ell}$ as \begin{align} \hat{\delta}_{e,\ell}=\frac{\mathbb{E}_{N}[\left(Y_{i,e+\ell}-Y_{i,s}\right)\cdot\ensuremath{\mathbf{1}}\left\{ E_{i}=e\right\} ]}{\mathbb{E}_{N}[\ensuremath{\mathbf{1}}\left\{ E_{i}=e\right\} ]}-\frac{\mathbb{E}_{N}[\left(Y_{i,e+\ell}-Y_{i,s}\right)\cdot\ensuremath{\mathbf{1}}\left\{ E_{i}\in C\right\} ]}{\mathbb{E}_{N}[\ensuremath{\mathbf{1}}\left\{ E_{i}\in C\right\} ]}. \end{align} The assumption of non-empty cohort $e$, existence of pre-period and non-empty control cohorts makes the DID estimator well-defined. For example, DID estimators for cohort 0 are not well-defined because a pre-period does not exist for this cohort, which is why we exclude them from estimating regression ((ref)) in Step 1. Using the above definition, in regression ((ref)) from Step 1 of our proposed method, the coefficient estimator $\widehat{\delta}_{e,\ell}$ is a DID estimator for $CATT_{e,\ell}$ with pre-period $s=e-1$ (because we exclude relative period $\ell=-1$) and some choice of control cohorts $C$. If there is a never-treated cohort i.e. $\infty\in supp\{E_{i}\}$, then we set the control cohort to be never-treated units $C=\{\infty\}$. If there are no never-treated units, i.e. $\infty\not\in supp\{E_{i}\}$, then we set the control cohort $C=\{\max\{E_{i}\}\}$, i.e. the latest-treated cohort. Among all possible pre-periods, for regression ((ref)) from step 1 of our proposed method we choose $s=e-1$ and $C=\{\infty\}$ with never-treated units (or $\{\max\{E_{i}\}\}$ without never-treated units) because the resulting specification is a natural extension of the common specifications of two-way fixed effects regression. Note that without never-treated units we need to drop time periods $t\geq\max\{E_{i}\}$ from estimating regression ((ref)) because every unit will be treated in these periods. The DID estimators for $CATT_{e,\ell}$ for $e+\ell\geq\max\{E_{i}\}$ are thus not well-defined as the control cohort is empty $C=\emptyset$. For example, when there are just two cohorts with $E_{i}\in\{1,T\}$, one treated at $t=1$ and the other treated in the last period, we then need to omit interaction terms involving the latest-treated cohorts as well as dropping observations from the last period $t=T$ from estimation.

Under some assumptions, this DID estimator $\widehat{\delta}_{e,\ell}$ is an unbiased and consistent estimator for $CATT_{e,\ell}$, a fact that we build on in deriving the probability limit of the IW estimator. We state this in the following proposition.

propIf Assumptions (ref) and (ref) hold, then the DID estimator using any pre-period $s<e$ and non-empty control cohorts $C$ is an unbiased and consistent estimator for $CATT_{e,\ell}$.

It is possible to relax the parallel trends assumption to allow the timing of treatment to depend on covariates. One can estimate $CATT_{e,\ell}$ consistently based on the inverse propensity score reweighted estimator proposed by semiparametric_did and callaway_santanna_ssrn2018, the outcome regression approach by heckman_ichimura_todd_restud1997, and the doubly robust estimator recently proposed by santanna_zhao_dr_did2018. The resulting estimates $\hat{\delta}_{e,l}$ can then be plugged into step 3 to form our IW estimator. In particular, without covariates and for the case with never treated units, our approach coincides with callaway_santanna_ssrn2018. Therefore one can use the did R package developed by callaway_R to form the IW estimator.

The choice of the pre-period $s$ and the control cohorts $C$ depends on the trade-off between relaxing Assumption (ref) or Assumption (ref). If Assumption (ref) is likely to hold, we can choose any pre-period $s<e$ and include not-yet-treated units in control cohorts $C=\{c:c>e+\ell\}$ as in callaway_santanna_ssrn2018. This choice allows us to relax the parallel trends assumption to be just $E[Y_{i,e+\ell}^{\infty}-Y_{i,0}^{\infty}\vert E_{i}=e]=E[Y_{i,e+\ell}^{\infty}-Y_{i,0}^{\infty}\vert E_{i}>e+\ell]$. If instead we want to relax Assumption (ref) to allow some anticipation, we then need the parallel trends assumption to hold. For example, suppose we are only willing to assume away anticipation 2 periods before a unit is treated, we may still be able to recover several $CATT_{e,\ell}$ by appropriately selecting a pre-period and a smaller set of control cohorts. We may choose any $s<e-2$ for the pre-period, and for control cohorts, we may choose cohorts treated only after at least 2 periods from now $C\subseteq\left\{ c:e+\ell+2<c\leq T\right\} $, or choose the latest-treated cohort $C=\{\max\{E_{i}\}\}$ as specified in regression ((ref)).

Empirical Illustration

We illustrate our findings in the setting of dobkin_finkelstein_kluender_notowidigdo_aer2018. dobkin_finkelstein_kluender_notowidigdo_aer2018 study the economic consequences of hospitalization, which is a large source of economic risk for adults in the United States. To quantify these economic risks, in the first part of their analysis, dobkin_finkelstein_kluender_notowidigdo_aer2018 leverage variation in the timing of hospitalization observed in the publicly-available dataset, Health and Retirement Study (HRS), which we describe in more detail in Section (ref). Their estimation of the dynamic effects of hospitalization using two-way fixed effects regressions provides a good context for demonstrating our results. As we argue below in Section (ref), parallel trends and no anticipation assumptions are plausible in this setting. However, the effect of hospitalization is potentially heterogenous across individuals hospitalized in different years. Our findings of non-convex and non-zero weighting would therefore apply to their two-way fixed effects regression estimates, and our alternative estimator could lead to different estimates. Furthermore, this dataset is publicly available, which allows us to provide replication files.

Data

Our sample selection closely follows dobkin_finkelstein_kluender_notowidigdo_aer2018 but we include a cursory explanation here for completeness with an emphasis on how our final sample differs from their main analysis sample. Our primary source of data is the biennial Health and Retirement Study (HRS). We identify the sample of individuals who appear in two sequential waves of surveys and newly report having a hospital admission over the last two years (the “index” or initial admission) at the second survey. To focus on health “shocks”, we restrict attention to non-pregnancy-related hospital admissions as in dobkin_finkelstein_kluender_notowidigdo_aer2018. We also follow dobkin_finkelstein_kluender_notowidigdo_aer2018 by focusing on adults who are hospitalized at ages 50-59.

Unlike dobkin_finkelstein_kluender_notowidigdo_aer2018, we restrict our analysis to a subsample of these individuals who appear throughout waves 7-11 (roughly 2004-2012). Our sample of analysis therefore includes HRS respondents with index hospitalization during waves 8-11. The purpose of this sample restriction is to maintain a balanced panel with a reasonable sample size.

Here $i$ indexes an individual, and $t$ indexes survey wave ($T=4$) and is normalized to zero for wave 7, the first wave in our sample. Among the outcomes $Y_{i,t}$ studied by dobkin_finkelstein_kluender_notowidigdo_aer2018, we focus on two: out-of-pocket medical spending and labor earnings. They are derived from self-reports, adjusted to 2005 dollars and censored at the 99.95th percentile.

Summary statistics. Table (ref) presents basic summary statistics for our analysis sample before hospitalization. We have a slightly lower fraction of white in our sample, but otherwise have a similar sample to dobkin_finkelstein_kluender_notowidigdo_aer2018.

In Panel D, we compare means of the cross-sectional distributions of outcomes for individuals who have not been hospitalized by each wave. The size of the sample conditional on not having been hospitalized strictly decreases with each subsequent wave. There are apparent time trends in our outcomes of interest prior to hospitalization as we observe distributional changes across waves. Out-of-pocket medical spendings fluctuate and earnings decrease with each wave on average as more individuals are retired in each subsequent wave.

Setting

We illustrate how variations in the timing of hospitalization fit the event studies design proposed in Section (ref). We define treatment $D_{i,t}$ to be ever having been hospitalized. In our terminology, we categorize individuals into cohorts based on $E_{i}$, which is defined as the survey wave of their initial hospitalization. Since we restrict the sample to individuals who were ever hospitalized in waves 8-11, there are four cohorts $E_{i}\in\left\{ 1,2,3,4\right\} $. Although hospitalization itself may not be an absorbing state, we are trying to model the impact of having had any hospitalization. Thus, the cohort-specific average treatment effects $CATT_{e,\ell}$ trace out the path of treatment effects for cohort $e$ following a negative health shock (even though the shock itself may be transient), as opposed to never having been hospitalized. Next we discuss whether each of the three identifying assumptions proposed is likely to hold in the context of unexpected hospitalizations.

Parallel trends (Assumption (ref)). Hospitalization is likely to be earlier among sicker individuals with high out-of-pocket medical spending and low labor earnings. Thus, it is not plausible that the baseline outcome $Y_{i,t}^{\infty}$ is mean independent of the timing of hospitalization. The parallel trends assumption is more plausible as it allows the timing to depend on unobserved time-invariant characteristics such as chronic disease. Furthermore, hospitalized individuals might be on a downward trend for labor earnings already prior to hospitalization compared to individuals who are never hospitalized. To reduce such confounding, we follow dobkin_finkelstein_kluender_notowidigdo_aer2018 to restrict the parallel trends assumption to individuals who were ever hospitalized.

No anticipatory behavior (Assumption (ref)). It is plausible that there is no anticipatory behavior prior to the hospitalization, given that the treatment is restricted to conditions that are likely unexpected hospitalizations. This assumption may be violated if individuals have private information about the probability of these hospitalizations over time and thus adjust their behavior prior to hospitalization.

Treatment effect heterogeneity (Assumptions (ref)). For out-of-pocket medical spending, the effect of hospitalization is potentially heterogenous across individuals hospitalized in different waves. Individuals hospitalized in later waves are mechanically older at the time of hospitalization than individuals hospitalized in earlier waves. The effect on out-of-pocket medical spending is largely determined by generosity of health insurance, which may decrease as individuals age into Medicare. The effect on labor earnings is also likely heterogenous as it depends on the labor market condition at the time of hospitalization: for example, individuals hospitalized during the financial crisis may find it more difficult to return to the labor force, and suffer a more grave decrease in earnings.

Illustrating weights in two-way fixed effects regression

We illustrate our results on two-way fixed effects regression by estimating the following specification with indicators for up to three leads and lags following equation (3) of dobkin_finkelstein_kluender_notowidigdo_aer2018

equation[equation omitted — 200 chars of source]

For their estimation, dobkin_finkelstein_kluender_notowidigdo_aer2018 trim their sample, keeping only observations up to three waves prior to the hospitalization and three waves after the hospitalization, and weight their regression with survey weights. To fully illustrate issues with this specification, we do not trim, but rather use a sample balanced in calendar time for $t\in\{0,\dots,4\}$, and do not apply survey weights. With a sample balanced in calendar time, we need to exclude at least two relative period indicators due to multicollinearity. Following dobkin_finkelstein_kluender_notowidigdo_aer2018, we exclude the period right before hospitalization ($\ell=-1$). We also exclude $\ell=-4$. Since we do not trim, note that our results are not directly comparable to dobkin_finkelstein_kluender_notowidigdo_aer2018 even though they are quite similar.

We focus on a single coefficient $\mu_{-2}$ that is supposed to test for any pre-trend of hospitalizations. As in Proposition (ref), we can decompose $\mu_{-2}$ as

equation[equation omitted — 224 chars of source]

As discussed in Proposition (ref) we can estimate the underlying weights $\omega_{e,\ell}^{-2}$ by regressing $\mathbf{1}\left\{ E_{i}=e\right\} \cdot D_{i,t}^{\ell}$ on the relative wave indicators included in specification ((ref)) i.e. $\{D_{i,t}^{\ell}\}_{\ell=-3,\neq-1}^{3}$ and two-way fixed effects. The coefficient estimator of $D_{i,t}^{-2}$ in such regression, $\widehat{\omega}_{e,\ell}^{-2}$, consistently estimates $\omega_{e,\ell}^{-2}$.

Figure (ref) plots these estimated weights. As described in our decomposition results, these weights have the following properties: (a) the four weights from relative wave $\ell=-2$ sum to one; (b) the weights for other included relative waves $\ell\in\{-3,0,1,2,3\}$ sum to zero for each included relative wave; and (c) the weights from excluded relative waves $\ell\in\{-4,-1\}$ sum to negative one across these excluded relative waves.

The weights are non-negative for lags of treatments, which suggest that the FE estimate $\widehat{\mu}_{-2}$ is particularly sensitive to estimates of the dynamic effects of hospitalizations and does not isolate the pre-trends. Specifically, treatment effects heterogeneity in $\ell=-3,0,1,2$ can affect the FE estimate $\widehat{\mu}_{-2}$. Applied researchers can make similar plots to visualize the role of weights in their settings with our publicly-available Stata package {eventstudyweights} by sun_stata.

Comparing FE and IW estimates

We illustrate our alternative method (IW estimator) following the three steps outline in Section (ref). First, we estimate the interacted specification ((ref)) as

equation[equation omitted — 207 chars of source]

for $t=0,\dots,3$. Specifically, we estimate $CATT_{e,\ell}$ using a DID estimator $\widehat{\delta}_{e,\ell}$ with pre-period $s=-1$ and control cohort $C=4$, the cohort hospitalized in the last period. This means we need to drop $t=4$ from estimation because everyone has been hospitalized by $t=4$, and a control cohort for estimating $CATT_{e,\ell}$ in $t=4$ does not exist. Second, we estimate the sample share of each cohort $e$ across cohorts that experience at least $l$ periods relative to hospitalization by its sample analog. Third, we form IW estimates $\widehat{\nu}_{\ell}$ by taking weighted averages of $\widehat{\delta}_{e,\ell}$ (returned from step one) with sample cohort share (returned from step two) as weights.

In Table (ref), we report the FE estimates $\widehat{\mu}_{\ell}$ and the IW estimates $\widehat{\nu}_{\ell}$, as well as the underlying $CATT_{e,\ell}$ estimates $\widehat{\delta}_{e,\ell}$. In this application, the FE estimates and IW estimates happen to be very similar in their magnitude. The conclusion from dobkin_finkelstein_kluender_notowidigdo_aer2018 based on FE estimates similar to ours still holds: we find a substantial and persistent decline in the earnings due to hospitalization, and the increase in out-of-pocket spending is transitory and small in comparison.

However, the FE estimates still suffer from non-convex and non-zero weighting as we show in Section (ref), which can lead to interpretability issues depending on the amount of treatment effect heterogeneity. Even though for both outcomes $\widehat{\mu}_{0}$ falls in the convex hull of its underlying $CATT_{e,0}$ estimates, $\widehat{\mu}_{-2}$ turns out to be outside the convex hull of its underlying $CATT_{e,-2}$ estimates. In contrast, by construction, the IW estimates $\widehat{\nu}_{\ell}$ fall within the convex hull of its underlying $CATT_{e,\ell}$ estimates and are unaffected by $CATT_{e,\ell'}$ estimates from other periods $\ell'\neq\ell$. Thus, they have an interpretation as an average effect of the treatment on the treated $\ell$ periods after initial treatment.

Conclusions

This paper analyzes the behavior of relative period coefficients $\mu_{\ell}$ on the indicator for being $\ell$ periods away from the treatment from two-way fixed effects regressions in settings with variation in treatment timing and treatment effects heterogeneity. For dynamic treatment effects, researchers are usually interested in estimating some average of treatment effects from $\ell$ periods relative to the treatment, and it is common to report the coefficient estimate $\widehat{\mu}_{\ell}$ assuming that interpretation is valid. However, we show that in the presence of heterogenous treatment effects, the coefficient $\mu_{\ell}$ does not necessarily capture the dynamic treatment effect as it can fall outside the convex hull of $CATT_{e,\ell}$ from its corresponding period $\ell$ and may pick up spurious terms consisting of treatment effects from periods other than $\ell$. This negative result is based on the decomposition of the coefficients $\mu_{\ell}$ as a linear combination of cohort-specific average treatment effects on the treated $CATT_{e,\ell}$. The weights in the linear combination can be non-convex, and non-zero on relative periods $\ell'\neq\ell$.

Given these negative results on two-way fixed effects regression estimators, we propose “interaction-weighted” (IW) estimators for estimating dynamic treatment effects. The IW estimators are formed by first estimating $CATT_{e,\ell}$ with a regression saturated in cohort and relative period indicators, and then averaging estimates of $CATT_{e,\ell}$ across $e$ at a given $\ell$. These $CATT_{e,\ell}$ are identified under parallel trends and no anticipation assumptions. These estimators are easy to implement and robust to heterogenous treatment effects across cohorts; the IW estimator associated with relative period $\ell$ is guaranteed to estimate a convex average of $CATT_{e,\ell}$ using weights that are sample share of each cohort $e$.

Finally, we illustrate the empirical relevance of our results by estimating the dynamic effect of hospitalization on the out-of-pocket medical spending and labor earnings using a setup similar to dobkin_finkelstein_kluender_notowidigdo_aer2018. We find non-convex and non-zero weights for two-way fixed effects regression in this example, and show that the resulting estimates indeed sometimes fall outside the convex hull of the underlying CATT estimates due to contamination by treatment effects from other relative periods. IW estimates, on the other hand, are weighted averages of the underlying CATT estimates with weights representative of cohort share.

More broadly, our paper suggests that researchers can do more in this context to assess the validity of underlying assumptions and how violations may impact their estimates. We demonstrate the sensitivity of two-way fixed effects regressions to underlying assumptions and provide researchers with tools to assess and address these issues. Specifically, for average treatment effects estimated using two-way fixed effects regressions we recommend that empirical researchers directly estimate the underlying weights on cohort-specific average treatment effects using our proposed auxiliary regressions. This exercise allows empirical researchers to assess the degree of potential contamination under any assumed structure of treatment effects heterogeneity. We also recommend empirical researchers consider estimation methods more robust to treatment effects heterogeneity such as the IW estimator we propose. Developing additional tools for applied researchers to identify and estimate interpretable parameters of interest is a promising avenue for future research.

figure[figure omitted — 945 chars of source]
figure[figure omitted — 847 chars of source]
figure[figure omitted — 636 chars of source]
sidewaystable\begin{centering} \caption{Survey of Applied Papers} \begin{tabular}{>p{0.2\textwidth}>p{0.1\textwidth}>p{0.1\textwidth}>p{0.1\textwidth}>p{0.1\textwidth}>p{0.1\textwidth}>p{0.1\textwidth}>p{0.1\textwidth}} \hline Paper & Binary Absorbing Treatment & Variation in Treatment Timing & Pre-Treatment Relative Period Excluded & Exclude Distant Relative Periods & Bin Distant Relative Periods & Includes Never Treated Units & Panel Balanced in Relative Time\tabularnewline \hline Bosch and Campos-Vazquez (2014) & X & & & & & & \tabularnewline \hline Fitzpatrick and Lovenheim (2014) & & & & & & & \tabularnewline \hline Gallagher (2014) & X & X & -1 & & X & X & \tabularnewline \hline Tewari (2014) & X & X & 0 & & & & X\tabularnewline \hline Ujhelyi (2014) & X & X & -1 & & X & X & \tabularnewline \hline Bailey and Goodman-Bacon (2015) & X & X & -1 & & X & X & X\tabularnewline \hline Deryugina (2017) & X & X & -1 & & X & X & X\tabularnewline \hline Deschenes et al. (2017) & X & & & & & & \tabularnewline \hline He and Wang (2017) & X & X & -1 & & X & & \tabularnewline \hline Lafortune et al. (2017) & X & X & 0 & X & & X & \tabularnewline \hline Kuziemko et al. (2018 & X & X & -1 & & X & & \tabularnewline \hline Markevich and Zhuravskaya (2018) & & & & & & & \tabularnewline \hline \end{tabular} \end{centering} Notes: This table consolidates key properties of main event study specifications across a sample of applied papers. We follow the selection criteria in roth_pretest_2019: the original sample consists of 70 total papers, but is further constrained to these twelve papers with publicly available data and code. The data and code are used to determine exactly the specification estimated in these papers. We focus on the first specification underlying the event study estimates in each paper, which we view as a reasonable proxy for the main specification in the paper. Note that two papers (Fitzpatrick and Lovenheim (2014); Markevich and Zhuravskaya (2018)) have none of the attributes listed in the columns.
table[table omitted — 2,097 chars of source]
table[table omitted — 3,354 chars of source]