Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
82,085 characters · 9 sections · 63 citation commands
Difference in Differences with Time-Varying Covariates
\abstract{ This paper considers identification and estimation of causal effect parameters from participating in a binary treatment in a difference in differences (DID) setup when the parallel trends assumption holds after conditioning on observed covariates. Relative to existing work in the econometrics literature, we consider the case where the value of covariates can change over time and, potentially, where participating in the treatment can affect the covariates themselves. We propose new empirical strategies in both cases. We also consider two-way fixed effects (TWFE) regressions that include time-varying regressors, which is the most common way that DID identification strategies are implemented under conditional parallel trends. We show that, even in the case with only two time periods, these TWFE regressions are not generally robust to (i) time-varying covariates being affected by the treatment, (ii) treatment effects and/or paths of untreated potential outcomes depending on the level of time-varying covariates in addition to only the change in the covariates over time, (iii) treatment effects and/or paths of untreated potential outcomes depending on time-invariant covariates, (iv) treatment effect heterogeneity with respect to observed covariates, and (v) violations of strong functional form assumptions, both for outcomes over time and the propensity score, that are unlikely to be plausible in most DID applications. Thus, TWFE regressions can deliver misleading estimates of causal effect parameters in a number of empirically relevant cases. We propose both doubly robust estimands and regression adjustment/imputation strategies that are robust to these issues while not being substantially more challenging to implement.}
{ {JEL Codes:}} C14, C21, C23
{ {Keywords:}} Difference-in-Differences, Time Varying Covariates, Two-way Fixed Effects Regression, Doubly Robust, Conditional Parallel Trends, Treatment Effect Heterogeneity
\onehalfspacing
In this paper, we study difference in differences identification strategies where (i) the parallel trends assumption holds only after conditioning on covariates, (ii) some or all of these covariates vary over time, and (iii) some of the time varying covariates could themselves be affected by the treatment.
A number of papers (e.g., heckman-ichimura-smith-todd-1998,abadie-2005,santanna-zhao-2020) show that certain causal effect parameters, typically the average treatment effect on the treated (ATT), are identified under conditional parallel trends assumptions. These types of conditional parallel trends assumptions are attractive in applications where the path of untreated potential outcomes may differ among units with different characteristics. However, work in the econometrics literature typically considers the case where covariates involved in the parallel trends assumption either do not vary over time or are “pre-treatment” (that is, the value of a time-varying covariate is set to its value in the pre-treatment period; see bonhomme-sauder-2011,lechner-2011 for some discussions on using pre-treatment values of time-varying covariates). In contrast, empirical work in economics often only includes covariates that vary over time. In this case, identification must implicitly assume that the treatment does not have an effect on the covariates themselves, which is implausible in some applications.
Covariates that could have been affected by participating in the treatment are often referred to as “post-treatment” or as “bad controls.” The received wisdom seems to be that this type of covariate should not be included in empirical research.\footnote{For example, angrist-pischke-2008 discuss “bad controls” in the context of deciding whether or not to control for occupation when studying causal effects of graduating from college on earnings. In that case, occupation is likely to be affected by attending college and, therefore, can make comparisons in earnings among those with the same occupation who graduated or did not graduate from college hard to interpret (even if college were randomly assigned). angrist-pischke-2008 note that “...we would do better to control only for variables that are not themselves caused by education.” We return to a related example later in this section on the effect of job displacement on earnings where occupation is potentially affected by job displacement. } However, we provide several examples below where it seems important to condition on the value of the covariate that would have occurred in the absence of the treatment; in these cases, it would not generally be sufficient to just “not include” this sort of covariate. We propose several different strategies for dealing with time-varying covariates that show up in the parallel trends assumption while also potentially being affected by the treatment.
Difference in differences identification strategies are most often implemented using two-way fixed effects (TWFE) regressions. The most common version of a TWFE regression that includes covariates is the following
where $\theta_t$ is a time fixed effect, $\eta_i$ is individual-level unobserved heterogeneity (i.e., an individual fixed effect), $D_{it}$ is the treatment indicator, and $X_{it}$ are time varying covariates. In the TWFE regression in (ref), $\alpha$ is the parameter of interest and it is often interpreted as “the causal effect of the treatment” or at least would be hoped to be a weighted average of underlying heterogeneous treatment effects. Being able to include covariates is one of the main attractions of using a TWFE regression to implement a DID design. For example, angrist-pischke-2008 write: “A second advantage of regression-DD is that it facilitates empirical work with regressors other than switched-on/switched off dummy variables.”\footnote{ angrist-pischke-2008 also briefly mention “bad controls” in the context of difference in differences (Section 5.2.1).} TWFE regressions have come under much scrutiny in recent work in terms of how well they perform for implementing DID identification strategies. In particular, TWFE regressions can perform very poorly in the presence of more than two time periods, variation in treatment timing across units, and treatment effect heterogeneity (particularly, treatment effect dynamics); see goodman-2021,chaisemartin-dhaultfoeuille-2020. Although with only two time periods, TWFE regressions are known to be reliable under unconditional parallel trends, here we point out a number of problems with TWFE regressions for implementing DID identification strategies that rely on conditional parallel trends assumptions even in the case with only two time periods.
In particular, we show that TWFE regressions can deliver poor estimates of the average treatment effect on the treated (which is the natural target parameter for DID identification strategies) for any of four reasons: (1) time-varying covariates that are themselves affected by the treatment, (2) ATTs and/or parallel trends assumptions that depend on the pre-treatment level of time varying covariates in addition to (or instead of) only the change in the covariates over time, (3) ATTs and/or paths of untreated potential outcomes that depend on time-invariant covariates, and (4) violations of strong functional form assumptions both for outcomes over time and for the propensity score. All four of these issues are common in applications in economics.
In applications where none of the four issues mentioned above occur, TWFE regressions deliver a weighted average of conditional ATTs where all the weights are positive. However, even in this best-case scenario, TWFE regressions still suffer from a “weight-reversal” property similar to the one pointed out in sloczynski-2020 under unconfoundedness with cross-sectional data. In our case, conditional ATTs for relatively uncommon values of the covariates among the treated group (relative to the untreated group) are given large weights while conditional ATTs for common values of the covariates among the treated group are given small weights. In order to get around this weight reversal issue, one needs to additionally rule out heterogeneous treatment effects across different values of the covariates. Adding this condition to the previous four implies that TWFE regressions deliver the ATT; however, we stress that these are a very stringent set of requirements for TWFE regressions to perform well for estimating the ATT when the parallel trends assumption depends on time-varying covariates.
We propose several new strategies for dealing with time varying covariates that are required for the parallel trends assumption to hold. When the researcher is confident that the covariates evolve exogenously with respect to the treatment, we provide a doubly robust estimand for the ATT (these arguments are similar to the ones in santanna-zhao-2020 for the case with time invariant covariates). Doubly robust estimators have the property that they deliver consistent estimates of the ATT if either an outcome regression model or a propensity score model is correctly specified, thus giving researchers an extra chance to correctly specify a model relative to regression adjustment or propensity score weighting strategies. Besides this, our doubly robust estimands can also be used in the context of the double/debiased machine learning literature where the propensity score and outcome regression model can be estimated using a wide variety of modern machine learning techniques (see chernozhukov-etal-2018 for the general case and chang-2020 in the context of DID).\footnote{Using machine learning in this context may be particularly useful because the expressions for the ATT involve conditioning on time-varying covariates across different time periods. In many applications, time-varying covariates may be highly serially correlated, and it may be challenging to specify simple parametric models involving these covariates in this context. However, machine learning estimators may perform much better in this context.} When the time-varying covariates can be affected by the treatment, we provide sufficient (and easy-to-interpret) conditions under which the strategy of conditioning on “pre-treatment” covariates, which is common in the econometrics literature, is justified. We also discuss other cases where this strategy is not reasonable. In these cases, we propose regression adjustment-type and doubly robust-type expressions for the ATT. Finally, when a researcher is willing to make an additional function form assumption for untreated potential outcomes, we propose some even simpler approaches based on regression adjustment (these approaches are also broadly similar to recent “imputation estimators” proposed in liu-wang-xu-2021,gardner-2021,borusyak-jaravel-spiess-2021). We also show that stronger functional form assumptions for the model for untreated potential outcomes can allow for parallel trends-type assumptions for the covariates to be sufficient for identification of the ATT.
Before moving into our main arguments, we provide three examples to illustrate the types of questions that we address in the current paper. We revisit these applications at relevant parts of the paper.
The examples above are broadly representative of applications that invoke DID identification assumptions with time varying covariates. The first example involves time-varying covariates that can reasonably be thought of as evolving exogenously with respect to the treatment. The following two examples both involve covariates that are potentially affected by the treatment. Later in the paper, we point out some further conceptual differences between these latter two examples.
\paragraph{Related Literature} \
Our paper shares a similar motivation to zeldow-hatfield-2021 which considers different possible sources of bias due to controlling for time-varying covariates that are possibly affected by the treatment. That paper mainly considers how sensitive existing strategies are (e.g., controlling for only pre-treatment covariates or additionally including lagged outcomes) to covariates that can be affected by the treatment. Relative to that paper, we make explicit assumptions on how the treatment can affect the covariates and, under these extra conditions, are able to propose estimation strategies that are guaranteed to perform well (up to regularity conditions) in those cases.
Our paper is also related to the literature on causal inference with panel data using structural nested mean models (robins-1997) and marginal structural models (robins-hernan-brumback-2000); see blackwell-glynn-2018 for a recent review. These approaches, however, are based on “sequential ignorability” assumptions rather than allowing for time-invariant unobserved heterogeneity. Sequential ignorability implies that treated and untreated potential outcomes are independent of treatment status conditional on pre-treatment values of covariates (and possibly pre-treatment outcomes).\footnote{Another difference between the current paper and much of the sequential ignorability literature is that these papers are typically primarily interested in recovering causal effects of different treatment paths (e.g., where each unit can move into or out of the treatment in each period). The arguments in our paper could likely be extended in this direction but our main results apply to the case where there are only two time periods and treatment can only take place in the second time period.} Unlike the bulk of this literature, the current paper focuses on the case where a researcher would like to invoke a parallel trends assumptions -- rather than sequential ignorability -- for identification. However, the current paper also invokes an additional assumption on how treated and untreated potential covariates are generated; this type of assumption is not made in this literature. The reason for this is that the timing that we consider differs from what is typically considered in the literature on sequential ignorability; in our case, units potentially become treated, then their covariate realizes (and may itself be affected by treatment) and this covariate needs to be controlled for identification. By contrast, the sequential ignorability literature typically has the covariate realized first, then the treatment, then the outcome, and controlling for, effectively, the covariate in the previous period is sufficient for identifying parameters of interest. That said, like the current paper, that literature does take seriously how covariates evolve over time and how participating in the treatment can affect covariates themselves. Of papers broadly in this literature, the most similar to the current paper is imai-kim-wang-2018 which focuses on a conditional parallel trends assumption that can hold after conditioning on past values of the covariates as well as past values of the outcome.
Our paper is also related to the literature on mediation analysis. Like a mediator, our covariates can be affected by treatment participation. However, the mediation literature is typically interested in decomposing treatment effects into direct effects of the treatment and indirect effects due to the effect of the treatment on the mediator (see huber-2020 for a recent review of this literature). Our paper is less ambitious in that we only seek to identify the overall effect of the treatment on outcomes; the tradeoff is that we are able to generally make weaker assumptions than would be required to separately recover direct and indirect effects of participating in the treatment. That said, it would be interesting to extend our arguments to additionally identifying direct and indirect effects of participating in the treatment, and it seems likely that existing arguments from the mediation analysis literature could be applied in this case. Our paper is relatively more similar to rosenbaum-1984,lechner-2008,flores-lagunes-2009; these papers consider identification of treatment effect parameters under unconfoundedness (and with cross-sectional data) where the covariates that are required for the unconfoundedness assumption to hold could have been affected by the treatment. Besides this, our paper is related to a large literature in econometrics on strict exogeneity and pre-determinedness in panel data models (see, for example, arellano-honore-2001).
Finally, our results on interpreting TWFE regressions build on work on interpreting cross-sectional regressions under the assumption of unconfoundedness and in the presence of treatment effect heterogeneity; this literature includes angrist-1998,aronow-samii-2016,sloczynski-2020,goldsmith-hull-kolesar-2021,ishimaru-2021. goodman-2021,chaisemartin-dhaultfoeuille-2020,ishimaru-2022 all provide decompositions of the TWFE regression in (ref). In some ways, the decompositions in these papers are more general than our decomposition as they all consider the case with more than two time periods and with variation in treatment timing. On the other hand, our results zoom in on the “textbook” case with exactly two periods and where no one is treated in the first period; our decomposition emphasizes a number of possible limitations of TWFE regressions even in the case with exactly two periods. Indeed, moving to more complicated cases with more periods and variation in treatment timing would make the case for using TWFE regressions even weaker, as it would introduce additional issues particularly related to using already treated units as comparison units (which can lead to negative weights on underlying treatment effect parameters), as all three papers mentioned above imply. See (ref) below for a more detailed comparison.
For this section, we focus on a baseline case where the researcher has access to two time periods of panel data. We label the second time period $t^*$ and the first time period $t^*-1$, and use $t$ to indicate a generic time period. In each time period, we observe outcomes $Y_t$, a time-varying covariate $X_t$, and time invariant covariates $Z$. As is standard in the DID literature, we suppose that no one is treated in the first time period. We use the binary variable $D$ to indicate whether or not a unit participates in the treatment. Importantly for our setup, we allow for the possibility that the time varying covariate can itself be affected by the treatment; in order to do this, we define $X_{t}(1)$ to be the value that the covariate would take if a unit participated in the treatment and $X_{t}(0)$ to be the value that the covariate would take if a unit did not participate in the treatment; for simplicity, we often refer to these as “treated potential covariates” and “untreated potential covariates.” Next, we define treated potential outcomes as $Y_t(1,X_t(1))$ (this is the outcome that a unit would experience in time period $t$ if they participated in the treatment and their covariate took on its value under the treatment) and untreated potential outcomes as $Y_t(0,X_t(0))$ (this is the outcome that a unit would experience in time period $t$ if they did not participate in the treatment and their covariate took its value in the absence in the treatment). For most of the arguments in the current paper, it is sufficient to use the shorter notation $Y_t(1) := Y_t(1,X_t(1))$ and $Y_t(0) := Y_t(0,X_t(0))$. In this setup, the observed covariates in each time period are: $X_{t^*} = D X_{t^*}(1) + (1-D)X_{t^*}(0)$ and $X_{t^*-1} = X_{t^*-1}(0)$. In other words, in the second time period, we observe treated potential covariates for units that participate in the treatment, and we observe untreated potential covariates for units that do not participate in the treatment. In the first time period, since no units are treated yet, we observe untreated potential covariates for all units. Likewise, observed outcomes are given by $Y_{t^*} = DY_{t^*}(1) + (1-D)Y_{t^*}(0)$, and $Y_{t^*-1} = Y_{t^*-1}(0)$.
Following the vast majority of the DID literature, we target identifying the average treatment effect on the treated (ATT). It is given by
which is the average difference between treated and untreated potential outcomes among the treated group.
Throughout the paper, we make the following assumptions
(ref) says that we observe iid panel data. (ref) says that, on average, the path of untreated potential outcomes is the same for the treated group as for the untreated group after conditioning on untreated potential covariates in time period $t^*,$ pre-treatment covariates $X_{t^*-1}$, and time-invariant covariates $Z$. Relative to standard conditional parallel trends assumptions (heckman-ichimura-todd-1997,abadie-2005,callaway-santanna-2021), the set of covariates being conditioned on includes untreated potential covariates which are unobserved for the treated group and therefore can complicate existing identification strategies. (ref) is an overlap assumption, and this typo of assumption is standard in the treatment effects literature. Part (a) implies that, for any values of $X_{t^*}$, $X_{t^*-1}$, and $Z$, there will be some untreated units with those values of the covariates in the population. Part (b) is similar but holds for any values of $X_{t^*-1}$, $W_{t^*-1}$, and $Z$.
Next, we provide two distinct assumptions for dealing with covariates that vary over time.
We call the first assumption covariate exogeneity because it implies that participating in the treatment does not change the distribution of covariates for the treated group. This assumption is technically weaker than assumptions like, for all units $X_{it^*}(1) = X_{it^*}(0) = X_{it^*}$ though this would certainly be a leading case where this sort of condition might hold. (ref) allows for covariates to change values over time, but it imposes that (in distribution) they are not affected by participating in the treatment. This sort of condition may be reasonable in some applications (e.g., Example 1 above). In other cases, this assumption may be less reasonable (e.g., Examples 2 and 3 above).
(ref) is an unconfoundedness assumption for untreated potential covariates. It allows for the treatment to effect the time varying covariates, but it says that the distribution of untreated potential covariates is the same for the treated group and the untreated group after conditioning on the vector of pre-treatment covariates $(X_{t^*-1}, W_{t^*-1}, Z)$. This assumption allows us to recover the conditional distribution of untreated potential covariates for the treated group. This distribution is a key ingredient for identifying the ATT below.
In (ref), we allow for the possibility that $W_{t^*-1}$ is empty; in fact, this is a leading case. In this case, unconfoundedness for untreated potential covariates holds after conditioning on the lag of the time-varying covariates $X_{t^*-1}$ and other time invariant covariates $Z$. Below, we connect this specific condition to the common practice in the econometrics literature on DID of conditioning on pre-treatment values of time-varying covariates. With a slight abuse of notation, we also allow for the possibility that $W_{t^*-1}$ includes the lagged outcome $Y_{t^*-1}$. For example, another interesting case is when $W_{t^*-1} = Y_{t^*-1}$, so that covariate unconfoundedness holds after conditioning on pre-treatment covariates, time invariant covariates, and the pre-treatment outcome. Interestingly, we show below that, under this condition, both the path of outcomes over time and the lag of the outcome show up in the expression for $ATT$ which is unusual in DID applications (see, chabe-2017 for related discussion). In the results below, we provide separate results that invoke either (ref) or (ref).
Next, we state our main identification result.
The intuition for part (1) of (ref) is relatively straightforward. Under the conditional parallel trends assumption and when covariates evolve exogenously, one can recover the ATT by (i) taking the path of outcomes experienced by the treated group and adjusting it by the path of outcomes experienced by the untreated group (conditional on $X_{t^*}$, $X_{t^*-1}$, and $Z$) and then (ii) accounting for differences in the distribution of $X_{t^*}$, $X_{t^*-1}$, and $Z$ across groups. This result is very similar to existing results with time invariant covariates (e.g., heckman-ichimura-todd-1997) as well as lechner-2011).
The intuition for part (2) is somewhat more complicated. The term $\mathbbm{E}[\Delta Y_{t^*}|X_{t^*},X_{t^*-1},Z,D=0]$ is the average change in outcomes over time conditional on $X_{t^*}$, $X_{t^*-1}$, and $Z$ among the untreated group. Under (ref), this is the path of outcomes that, conditional on $X_{t^*}(0), X_{t^*-1}$, and $Z$, the treated group would have experienced if they had not participated in the treatment. The next expectation is over the distribution of $X_{t^*}(0)$ (conditional on $X_{t^*-1}, W_{t^*-1}$, and $Z$) for the untreated group. Under (ref), this is the same conditional distribution that $X_{t^*}(0)$ follows for the treated group. Finally, the outside expectation is over the distribution of $X_{t^*-1}$, $W_{t^*-1}$, and $Z$ for the treated group and, therefore, allows for these variables to be distributed differently in the treated group relative to the untreated group.
(ref) provides two important special cases for the results in part (2) of (ref). The first part provides a formal justification for the common practice in the econometrics literature on DID with time varying covariates of including only “pre-treatment” covariates. In particular, this result says that, when unconfoundedness holds for the time varying covariate conditional on time-invariant covariates and other pre-treatment covariates, then it is sufficient for the researcher to only “account for” pre-treatment and time-invariant covariates in order to recover the ATT.\footnote{We provide the proof of this result in (ref). The proof is relatively straightforward, but it appears to be a new contribution in the literature.} The second part of (ref) is also interesting in that it relates the ATT to an expression that includes the lagged outcome. There are a number of papers that explore the idea of including lagged outcomes in a DID framework (e.g., chabe-2017,imai-kim-wang-2018,zeldow-hatfield-2021) though it is challenging to provide a justification for including lagged outcomes in DID settings --- our approach justifies the inclusion of lagged outcomes (in the manner specified in the corollary) in cases where unconfoundedness for the time-varying covariate holds after conditioning on the lag of the outcome variable.
Next, we provide alternative expressions for $ATT$ that are useful for estimation.
Both of the expressions in (ref) involve both an outcome regression (the conditional expectation terms in each expression) and a propensity score. They are both doubly robust in the sense that a researcher can consistently estimate the ATT if either the model for the propensity score or the outcome regression model is correctly specified (references on double robustness include robins-rotnitzky-zhao-1994,scharfstein-rotnitzky-robins-1999,sloczynski-wooldridge-2018,santanna-zhao-2020). Besides this, they also provide a connection to the DID literature on estimating the ATT under conditional parallel trends using double/debiased machine learning; see, in particular, chang-2020. This may be particularly useful in the first case where the propensity score and outcome regression depends on time-varying covariates in both periods. These can be practically difficult to estimate because, in many cases, $X_{t^*}$ and $X_{t^*-1}$ may be highly collinear. Conventional methods typically invoke functional form assumptions that impose, for example, that these functionals only depending on $\Delta X_{t^*}.$ As noted below, these sorts of restrictions may be implausible in many applications.
To conclude this section, we revisit the three examples from the introduction.
\paragraph{Example 1 (Stand-your-ground, cont'd)} Our example on stand-your-ground laws involved conditioning on time-varying covariates that evolved exogenously with respect to the treatment. This suggests that this example is most related to our results in part (1) of (ref) and part (1) of (ref). In particular, machine learning estimators of the propensity score and outcome regression functions in (ref) are particularly attractive as they do not require strong functional form assumptions on these nuisance functions.\footnote{This particular application uses state-level data, so, in practice, it may be difficult to use machine learning approaches with only 50 or so observations. See (ref) for some more parametric approaches that may be more suitable for applications with limited data. That said, the more general point here though is that, in cases where covariates evolve exogenously, machine learning estimators, given enough data, are likely to be attractive in many applications.}
\paragraph{Example 2 (Shelter-in-place, cont'd)} In our example of shelter-in-place orders on various economic outcomes, the parallel trends assumption held after conditioning on the number of Covid-19 cases that would have occurred if the policy had not been implemented. That is, “untreated potential Covid-19 cases” plays the role of $X_{t^*}(0)$ in this case. callaway-li-2021b show that, under a SIRD model --- which is the leading pandemic model in the epidemiology literature --- controlling for the pre-treatment “state” of the pandemic is sufficient for unconfoundedness to hold. That is, the conditions in part (1) of (ref) and part (2) of (ref) hold when one wants to control for the number of untreated potential Covid-19 cases.
\paragraph{Example 3 (Job displacement, cont'd)} Finally, recall our example on the effect of job displacement on earnings where the parallel trends assumption holds only after conditioning on, for example, “untreated potential occupation” --- that is, the occupation that a worker would have had if they had not been displaced from their job. In this case, an unconfoundedness assumption for occupation may be more likely to hold if it conditions on (i) pre-treatment time-varying covariates (including pre-treatment occupation), (ii) time invariant covariates (such as demographics and education), and (iii) pre-treatment earnings. In particular, conditioning on pre-treatment earnings could be important if there are occupation specific wage premiums and high-earning workers are more likely to (in the absence of job displacement) stay in the same occupation over time relative to low-earning workers. This application would then be covered by the results from part (2) of (ref).
In this section, we consider how to interpret $\alpha$ in the TWFE regression in (ref). We continue to consider the “textbook” case with two time periods where no one is treated in the first time period and where some (but not all) units become treated in the second time period. This is a best-case for TWFE regressions as it does not introduce well-known problems related to using already-treated units as comparison units that show up when using TWFE regressions with multiple periods, variation in treatment timing, and treatment effect heterogeneity (goodman-2021,chaisemartin-dhaultfoeuille-2020). In the case with exactly two periods, it is helpful to equivalently re-write (ref) as
where we define $\Delta X_{t^*} := (1,X_{t^*}-X_{t^*-1})'$ which is the change in the covariate over time and is augmented with an intercept term for the time fixed effect. We also slightly abuse notation by taking $\beta$ to include an extra parameter in its first position corresponding to the intercept. Our interest in this section is in determining what kind of conditions are required to interpret $\alpha$ as the ATT or at least as a weighted average of some underlying treatment effect parameters.
Denote the linear projection of $\Delta Y_{t^*}$ on $\Delta X_{t^*}$ by $\textrm{L}(\Delta Y_{t^*} | \Delta X_{t^*}) := \Delta X_{t^*}'\mathbbm{E}[\Delta X_{t^*} \Delta X_{t^*}']^{-1} \mathbbm{E}[\Delta X_{t^*} \Delta Y_{t^*}]$, and define the corresponding projection error $e := \Delta Y_{t^*} - \textrm{L}(\Delta Y_{t^*} | \Delta X_{t^*})$. Similarly, define the linear projection of $D$ on $\Delta X_{t^*}$ as $\textrm{L}(D|\Delta X_{t^*}) := \Delta X_{t^*}'\mathbbm{E}[\Delta X_{t^*} \Delta X_{t^*}']^{-1} \mathbbm{E}[\Delta X_{t^*} D]$ and the corresponding projection error $u:= D - \textrm{L}(D|\Delta X_{t^*})$. Standard Frisch-Waugh type arguments imply that
Below, to keep the notation concise, it is useful to define $X^{all}(d) := (X_{t^*}(d), X_{t^*-1}, Z)$. We also define $ATT_{X^{all}(0)}(X^{all}(0)) := \mathbbm{E}[Y_{t^*}(1) - Y_{t^*}(0) | X^{all}(0), D=1]$ which is the ATT conditional on $X_{t^*}(0)$, $X_{t^*-1}$, and $Z$. And we further define $p(X^{all}(0)) = \textrm{P}(D=1|X^{all}(0))$. Next, we state a main result decomposing $\alpha$ from the TWFE regression.
The result in (ref) indicates that $\alpha$ is equal to a weighted average of underlying conditional ATTs (we discuss the nature of the weights in more detail below) plus a number of undesirable “bias” terms. We provide formal conditions under which each of these extra terms are equal to zero below. But, informally, term (A) contains bias from the treatment potentially affecting the covariates in time period $t^*$. Term (B) comes from ignoring time invariant covariates. Term (C) comes up when paths of outcomes depend on the levels of time-varying covariates instead of only on the change in covariates over time. Term (D) arises when the conditional expectation is nonlinear in the change in covariates over time. Term (E) is conceptually different from terms (A)-(D) and is non-zero when the propensity score is not equal to a linear projection of the treatment on the change in covariates over time.\footnote{Without further assumptions, some of the expressions that involve $X_{t^*}(0)$ are not necessarily identified (this includes $ATT_{X^{all}(0)}(X^{all}(0)), \mathbbm{E}[\Delta Y_{t^*}|X_{t^*}(0), X_{t^*-1},Z,D=1]$ and all of the weights as they depend on $p(X^{all}(0))$. However, if we additionally invoke (ref), then all of these terms are identified and Term (A) is equal to zero (see the discussion below for more details along these lines). The reason we do not invoke this assumption in (ref) is to point out that time-varying covariates being affected by the treatment can itself be an additional complication for TWFE regressions.}
Next, we introduce several additional assumptions that are useful for eliminating the bias terms in (ref). We also use the additional notation: $ATT_{X_{t^*}(0),X_{t^*-1}}(x_{t^*}(0), x_{t^*-1}) := \mathbbm{E}[Y_{t^*}(1) - Y_{t^*}(0) | X_{t^*}(0)=x_{t^*}(0), X_{t^*-1}=x_{t^*-1}]$ and $ATT_{\Delta X_{t^*}(0)}(\Delta x_{t^*}(0)) := \mathbbm{E}[Y_{t^*}(1) - Y_{t^*}(0) | \Delta X_{t^*}(0) = \Delta x_{t^*}(0)]$ --- these define different types of conditional ATTs.
The first part of (ref) says that, conditional on $X_{t^*}(0)$ and $X_{t^*-1}$, conditional ATTs do not depend on time invariant covariates $Z$. The second part says that, conditional on $X_{t^*}(0)$ and $X_{t^*-1}$, the path of untreated potential outcomes does not depend on time invariant covariates $Z$. This implies that the conditional parallel trends assumption holds without conditioning on time invariant covariates (and thus strengthens (ref)). (ref) is similar; the first part says that conditional ATTs further only depend on changes in time-varying covariates over time, and the second part says that the conditional parallel trends assumption only depends on the change in time-varying covariates over time rather than their level.
(ref) says that conditional ATTs and paths of untreated potential outcomes are linear in changes in untreated potential covariates over time. Jointly, (ref) imply that (i) time varying covariates are not affected by the treatment, (ii) that conditional ATTs (conditional on $X_{t^*}(0), X_{t^*-1},$ and $Z$) only depend on the change in time-varying covariates (and not on their levels or time invariant covariates) and are linear in time-varying covariates, and (iii) that the conditional parallel trends assumption in (ref) only depends on the change in time-varying covariates over time (and not on their levels or time invariant covariates) and is linear in time-varying covariates over time.
(ref) says that the propensity score (conditional on $X_{t^*}(0)$, $X_{t^*-1}$, and $Z$) is linear in $\Delta X_{t^*}(0)$. This type of assumption is very common in the literature on interpreting regressions under unconfoundedness with cross-sectional data (e.g., angrist-1998,aronow-samii-2016,sloczynski-2020,ishimaru-2021). In those cases, it sometimes holds by construction (e.g., when the covariates are all discrete and a full set of interactions is included in the model). In our case, though, it seems particularly implausible as (i) it requires the propensity score to only depend on changes in covariates over time, and (ii) even with fully interacted discrete regressors, the propensity score is unlikely to be linear in changes in the regressors over time.\footnote{For example, suppose that the only covariate is binary. In the cross-sectional case considered by other papers mentioned above, the propensity score would be linear by construction. However, the change in the covariate over time would be a single variable that can take the values -1, 0, or 1; moreover, the change in a binary covariate over time is equal to 0 in cases when the covariate is equal to 1 in both periods or when the covariate is equal to 0 in both periods. This suggests that the propensity score would not be linear in the change in covariates over time even in this very simple case. }
The proof of (ref) is provided in (ref). In the proof, we provide more specific results on which conditions are required for each term in terms (A)-(D) in (ref) to be equal to 0.
The result in (ref) suggests a number of potential issues with the TWFE regressions as in (ref). First, even if one is willing to maintain the additional assumptions in (ref) (which are likely to be very strong in most applications), $\alpha$ from the TWFE regression is still hard to interpret. Maintaining these additional assumptions (particularly, (ref)) implies that all of the weights, conditional ATTs, and linear projections in the first part of (ref) are identified and directly estimable. That said, the weights on conditional ATTs, $\omega_{ATT}$, do not have the property that they have mean one and the nuisance expression in term (E) may be non-negligible.
The second part of (ref) says that, if we are willing to assume that the propensity score is equal to the linear projection of the treatment on the change in time-varying covariates over time, then the weights on conditional ATTs will have mean one and the nuisance expression in term (E) will be equal to zero. Even in this case, the weights have a “weight-reversal” property analogous to the one pointed out in sloczynski-2020 in the context of unconfoundedness and cross-sectional data. What this means is that conditional ATTs are given more weight for values of the covariates that are relatively uncommon among the treated group relative to the untreated group; and that conditional ATTs are given less weight for values of the covariates that are relatively common among the treated group relative to the untreated group.
Finally, if in addition to all the previous conditions, conditional ATTs are constant across different values of the covariates, then $\alpha$ will be equal to the $ATT$. This is a treatment effect homogeneity condition with respect to the covariates. It is somewhat weaker than individual-level treatment effect homogeneity and it allows for treatment effects to still be systematically different for treated units relative to untreated units; instead, for the treated group, treatment effects cannot be systematically different across different values of the covariates.
These results are much different from our earlier results in (ref). Those results did not require any of the additional assumptions in (ref). In fact, when covariates evolve exogenously with respect to the treatment (as under (ref)), then the doubly robust expressions for the ATT in part (1) of (ref) only require that either the propensity score or the outcome regression model be correctly specified; in cases where these are estimated using machine learning, even these parametric assumptions can be substantially relaxed. Moreover, in contrast with the TWFE regressions considered in this section, our earlier additional results can accommodate cases where the time-varying covariates are affected by the treatment.
In this section, we provide several alternative strategies that involve stronger parametric assumptions on the path of untreated potential outcomes than we made in (ref). The approaches discussed in this section are generally simpler to estimate than would be the case for the expressions coming from (ref) and, in some cases, can allow for weaker (or at least alternative) assumptions on how the treatment affects time-varying covariates. The strategies that we propose in this section are also able to avoid the issues with TWFE regressions pointed out in (ref), and (when desired) can allow for the possibility that the treatment has an effect on the covariates.
The ideas in this section are broadly similar to regression adjustment strategies in the treatment effects literature (see, for example, imbens-wooldridge-2009) and the imputation estimators proposed in liu-wang-xu-2021,gardner-2021,borusyak-jaravel-spiess-2021 though they allow for (i) time-varying effects of time varying covariates and (ii) the possibility that the treatment directly affects the time-varying covariates.
To start with, it is well known (e.g. blundell-dias-2009,gardner-2021,borusyak-jaravel-spiess-2021) that there is a close connection between unconditional parallel trends assumptions and the following model for untreated potential outcomes
where $\theta_t$ is a time fixed effect, $\eta_i$ is time invariant unobserved heterogeneity (i.e., an individual fixed effect), and $v_{it}$ are idiosyncratic, time varying unobservables. An unconditional version of parallel trends holds in this model for untreated potential outcomes under the condition that $\mathbbm{E}[\Delta v_{t} | D=1] = \mathbbm{E}[\Delta v_t | D=0]$ for all time periods (this would hold by construction if $v_t$ is independent of treatment status in all time periods), but allows for $\eta$ to be distributed differently across groups and does not impose any modeling assumptions on treated potential outcomes.
As discussed above, the econometrics literature on difference in differences often considers the case where the covariates in the parallel trends assumption are time invariant. In that case, the analogous model for untreated potential outcomes is given by
where the distribution of $\eta$ can vary across groups (as well as vary with $Z$) and the key condition for the conditional parallel trends assumption to hold is that $\mathbbm{E}[\Delta v_t | Z, D=1] = \mathbbm{E}[\Delta v_t | Z, D=0]$ (see, for example, heckman-ichimura-todd-1997 for a discussion of this kind of model).\footnote{To see this, notice that $\mathbbm{E}[\Delta Y_t(0) | Z, D=1] = g_t(Z) - g_{t-1}(Z) = \mathbbm{E}[\Delta Y_t(0) | Z, D=0]$ which implies that conditional parallel trends holds.}
In this setup, the main challenge is estimating $g_t(z)$ (though note that this is a practical, estimation challenge rather than an identification challenge). The natural way to parameterize this model is
where we now take $Z$ to include an intercept (and, therefore, $\delta_t$ absorbs the time fixed effect). Given this framework, $ATT=\mathbbm{E}[\Delta Y_{t^*} | D=1] - \mathbbm{E}[Z|D=1]'\delta_{t^*}^*$ where $\delta_{t^*}^* := (\delta_{t^*} - \delta_{t^*-1})$ which can be consistently estimated from the regression of $\Delta Y_{t^*}$ on $Z$ using only observations from the untreated group.\footnote{This is closely related to regression adjustment estimators (see, for example, heckman-ichimura-smith-todd-1998,imbens-wooldridge-2009,santanna-zhao-2020 for related discussion). An alternative strategy would be to estimate the ATT by $n_1^{-1} \sum_{i=1,D_i=1}^n (Y_{it^*} - \hat{Y}_{it^*}(0))$ where $n_1$ is the number of treated observations and $\hat{Y}_{it^*}(0)$ is an imputed untreated potential outcome given by $\hat{Y}_{it^*}(0) = Y_{it^*-1} + Z_i'\hat{\delta}_{t^*}^*$. This imputation estimator is numerically equal to the regression adjustment estimator, but the imputation formulation is convenient particularly in the case with multiple periods and variation in treatment timing; see (ref) below for more details.}
The same sort of arguments imply that, when there are some covariates that vary over time (as above, we consider the case of a single time-varying covariate but note that it is straightforward to extend these arguments to cases with more time-varying covariates), a natural motivating model is
which implies that
Moreover, the same sorts of arguments as above imply that (ref) holds in this model. Similar to the previous case, the main practical challenge is that $g_t(z,x_t(0))$ is likely to be challenging to estimate. Like the previous case, the natural way to parameterize this model is
which implies that
where $\beta_{t^*}^* := (\beta_{t^*} - \beta_{t^*-1})$. Notice that, because untreated potential outcomes and untreated potential covariates are observed for the untreated group, the parameters in the model above can be recovered from a regression of the change in outcomes over time on time invariant covariates, the change in time varying covariates, and the level of the time varying covariates in the pre-treatment period. The model in (ref) is conceptually appealing as (up to the parametric assumptions) it compares units with both the same initial level of the time-varying covariate and that have the same change in time-varying covariates over time.
Although it is straightforward to recover the parameters in (ref), recall that,
Given that the parameters are identified, every term is identified in this expression except for $\mathbbm{E}[\Delta X_{t^*}(0) | D=1]$ (because $X_{t^*}(0)$ is not observed for the treated group). We briefly consider six settings for recovering $\mathbbm{E}[\Delta X_{t^*}(0)|D=1]$ --- three of these come from the assumptions we have already considered for untreated potential covariates and three involve parallel trends assumptions for untreated potential covariates. Several of these cases involve averaging over conditional expectations of $\Delta X_t(0)$. In this section we additionally impose linear models for these conditional expectations; under this extra condition, researchers are able to estimate ATT while potentially allowing for the treatment to affect time-varying covariates using only regressions and averaging.
\paragraph{Case 1: (ref) holds} \
Under (ref), $\mathbbm{E}[\Delta X_{t^*}(0) | D=1] = \mathbbm{E}[\Delta X_{t^*}(1) | D=1] = \mathbbm{E}[\Delta X_{t^*} | D=1]$. That is, when covariates evolve exogenously with respect to the treatment, we can replace the average change in untreated potential covariates for the treated group with the average change in covariates actually experienced by the treated group.
\paragraph{Case 2: (ref) holds conditional on $\bm{(Z,X_{t^*-1})}$} \
In this case, if we are willing to assume the following linear model for untreated potential covariates
where it follows by the conditions in this case that $\mathbbm{E}[u_{t^*}|Z,X_{t^*-1},D=d]=0$ for $d \in \{0,1\}$. Plugging this expression into (ref) implies that
where $\delta^*_{2,t^*} := \delta^*_{t^*} + \gamma_{t^*} \beta_{t^*}$, $\beta^*_{2,t^*} := \beta^*_{t^*} + \lambda_{t^*} \beta_{t^*}$ and $v_{2,it^*} := \beta_{t^*} u_{it^*} + \Delta v_{it^*}$. Further, notice that $\mathbbm{E}[v_{2,t^*} | Z, X_{t^*-1}, D=d] = 0$ for $d \in \{0,1\}$. Thus, in this case, one can estimate $\delta^*_{2,t^*}$ and $\beta^*_{2,t^*}$ from a regression of the change in outcomes over time using the untreated group, and then estimate the ATT from the sample analogue of
Thus, this particular case bypasses the need for actually estimating a separate model for the change in the time-varying covariate over time. This is perhaps not surprising as these are the same conditions as in (ref) where it was sufficient for the researcher to condition on the pre-treatment value of the covariates to recover the ATT.
\paragraph{Case 3: (ref) holds conditional on $\bm{X_{t^*-1},W_{t^*-1},Z}$} \
In this case,
where the first equality holds by the law of iterated expectations, and the second equality holds by (ref) and by assuming a linear model for the change in untreated covariates over time. This suggests estimating $\mathbbm{E}[\Delta X_{t^*}(0)|D=1]$ by running a regression of $\Delta X_{t^*}$ on $Z$, $X_{t^*-1}$, and $W_{t^*-1}$ using only untreated observations in order to estimate the parameters $\gamma_{t^*}$, $\lambda_{t^*}$, and $\xi_{t^*}$, and then to estimate $\mathbbm{E}[\Delta X_{t^*}(0) | D=1]$ by using the sample analogue of the expression in (ref).\footnote{In the special case (and perhaps leading case) considered in (ref) where $W_{t^*-1}$ includes the lagged outcome, $Y_{t^*-1}$ (in addition to all time-invariant covariates and the pre-treatment version of the covariates), one can follow this same procedure with $Y_{t^*-1}$ substituting for $W_{t^*-1}$.}
\paragraph{Case 4: Unconditional Parallel Trends holds for time-varying covariates} \
For this case, we assume that $\mathbbm{E}[\Delta X_{t^*}(0) | D=1] = \mathbbm{E}[\Delta X_{t^*}(0) | D=0]$. It immediately follows that
This expression is very similar to the one in Case 1, except that one should use the change in untreated potential covariates for the untreated group.
\paragraph{Case 5: Conditional Parallel Trends holds for time-varying covariates} \
For this case, we assume that $\mathbbm{E}[\Delta X_{t^*}(0) | Z, D=1] = \mathbbm{E}[\Delta X_{t^*}(0) | Z, D=0]$. In this case,
where the first equality holds by the law of iterated expectations, the second equality holds under conditional parallel trends for time-varying covariates, and the last equality holds under a linearity assumption. This suggests estimating $\gamma_{t^*}$ by running a regression of $\Delta X_{t^*}$ on $Z$ using only untreated observations and then to estimate $\mathbbm{E}[\Delta X_{t^*}(0)|D=1]$ from the sample analogue of (ref).
\paragraph{Case 6: Conditional Parallel Trends Holds under Generic Parallel Trends Assumption} \
For this case, we assume that $\mathbbm{E}[\Delta X_{t^*}(0) | Z, W_{t^*-1}, D=1] = \mathbbm{E}[\Delta X_{t^*}(0) | Z, W_{t^*-1}, D=0]$. In this case,
where the first equality holds by the law of iterated expectations, and the second equality holds by the conditional parallel trends assumption used in this case and a linearity assumption. Similarly to above, this suggests running a regression of $\Delta X_{t^*}$ on $Z$ and $W_{t^*-1}$ using only untreated observations to estimate $\gamma_{t^*}$ and $\xi_{t^*}$ and then to estimate $\mathbbm{E}[\Delta X_{t^*}(0)|D=1]$ from the sample analogue of (ref).
All of the approaches discussed in this section are substantially more robust than the TWFE regressions discussed in (ref). In particular, unlike TWFE regressions, they allow for the path of untreated potential outcomes to depend on (i) time-invariant covariates, (ii) the pre-treatment level of time-varying covariates, and (iii) the change in time-varying covariates over time that would have occurred if the treatment had not taken place. Given any of a number of assumptions on the path of time-varying covariates in the absence of the treatment (as in Cases 1-6 above), they allow for the treatment to have an effect on time-varying covariates. They allow for general forms of treatment effect heterogeneity; for example, they do not require conditional ATTs to be linear in covariates (as in (ref)(a)) nor do they require any of the extra treatment effect homogeneity conditions for TWFE regressions in (ref). Finally, they do not require any linearity conditions for the propensity score as in (ref). Relative to the approach discussed in (ref), the approaches considered in this section require linearity assumptions on the model for untreated potential outcomes and, in some cases, on a model for the change in untreated potential covariates over time. The two main advantages of this approaches in this section are (i) parallel trends assumptions for time varying covariates can be strong enough to identify the ATT, and (ii) the approaches in this section are also easy to implement --- they only requiring running regressions and computing averages.
\begin{additional-material}
Many applications in economics have more than exactly two time periods available. In this section, we extend previous results in two directions. First, following callaway-santanna-2021, we show how the previous arguments can be naturally extended to identifying group-time average treatment effects and then be aggregated into, for example, overall treatment effect parameters or event studies.
Second, I consider pre-testing the main identifying assumptions in the case where a researcher has access to more than one pre-treatment period.
* comment on common practice of just conditioning on covariates in the first period
* can you do joint pre-test??
* may be more efficient estimators using GMM, but we ignore this now
The main challenge with (ref) is a practical one --- $\mathbbm{E}[\Delta Y_t|X_t,X_{t-1},D=0]$ is challenging to estimate. For one reason, (although we abstract from this issue) in most realistic applications, the dimension of $X$ may be fairly large. This means that nonparametric estimation may be challenging in practice.\footnote{This is a primary concern even in the case with time invariant covariates and is primary motivation for recent work on doubly robust estimators in the context of difference in differences (santanna-zhao-2020).} A second practical challenge is that, in many applications, $X_t$ and $X_{t-1}$ may be highly correlated with each other. This suggests that approaches such as trying to estimate that conditional expectation using a linear model may perform poorly in practice.\footnote{That being said, one potentially promising approach here would be to use some kind of machine learning approach here. chang-2020 studies difference in differences under a conditional parallel trends assumption using machine learning estimation strategies. That paper does not explicitly consider time varying covariates, but the expression in (ref) fits into that framework simply by including both $X_t$ and $X_{t-1}$ as covariates.} Indeed, trying to estimate this sort of expression either parametrically, nonparametrically, or using machine learning is not common in applied work.
Good candidates for applications:
\end{additional-material}
In the current paper, we have considered DID identification strategies where the parallel trends assumption holds only after conditioning on time varying covariates that may themselves be affected by the treatment. This setting is common in empirical applications in economics, and we have provided several approaches that offer a number of advantages relative to more commonly used TWFE regressions that include covariates (even in the case where there are only two time periods). In addition, the new approaches that we have proposed are generally not much more complicated to implement than TWFE regressions.
\onehalfspace
\printbibliography