Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
128,230 characters · 35 sections · 137 citation commands
Robust difference-in-differences models
\nonstopmode
\address{Rochester Institute of Technology and UNC-Chapel Hill}
{ Keywords: Differences-in-differences, baseline information, selection bias, robust bounds, ATT.\\ JEL subject classification: C14, C31, C33, C35.}
The difference-in-differences (DID) technique is one of the most popular methods in the social sciences when an experimental research design cannot be used. The DID method requires at least observational data consisting of two different groups (a treatment group and a control group) and two time periods of pre-treatment and post-treatment. Under some assumptions, the method identifies the average treatment effects on the treated (ATT) as the DID estimand. The key identifying assumption of interest is the so-called parallel trends (PT) assumption. This assumption states that the untreated potential outcome variable for the treatment group would have followed on average the same trend as that for the control group had they not been treated. However, it is difficult to empirically verify the PT assumption because it restricts a hypothetical quantity that is not identifiable. Accordingly, convincing readers to approve the PT assumption has been the most vital and controversial part of the DID literature. For instance, kearney2015mtv's (kearney2015mtv) ambitious identification strategy on discovering the effects of an MTV reality show on teen childbearing provided insightful findings that would not have been discovered without the study, but there have been heated debates on the validity of its main PT assumption as well jaeger2018, kahnlang2019. The most common and widely understood approach for justifying the PT assumption is the pre-treatment period examination. If a null hypothesis of the same trend in the untreated potential outcome mean for both treatment and control groups in the pre-treatment periods cannot be rejected, then the researcher will believe that the PT assumption is likely to hold for the post-treatment period as well. Still, rigorously speaking, the evidence of pre-treatment PT is different from the PT assumption in the post-treatment period that is of interest, and thus additional arguments should be established for the PT assumption separately. CallawaySantAnna2022 discussed this nuance well in their vignette on pre-testing in a DID framework:
“Importantly, this is just a pre-test; it is different from an actual test. Whether or not the parallel trends assumption holds in pre-treatment periods does not actually tell you if it holds in the current period (and this is when you need it to hold!). It is certainly possible for the identifying assumptions to hold in previous periods but not hold in current periods; it is also possible for identifying assumptions to be violated in previous periods but for them to hold in current periods. That being said, we view the pre-test as a piece of evidence on the credibility of the DiD design in a particular application.”
See also Freyaldenhoven_al2019, kahnlang2019, Roth2022, etc. Hence, this paper develops a generalized DID framework that can utilize all the information available not only from the pre-treatment periods but also from multiple baseline covariates or data sources. In doing so, we develop a DID method that is robust to violations of PT that can be captured in the pre-treatment periods.
First, our approach is unique in that we follow Heckman_al1998 to interpret the PT assumption in a different way using a notion of selection bias (also known as confounding bias in statistics),\footnote{Some recent papers also use similar interpretations of the PT assumption Henderson2023, Sofer2016, Parkal2023.} which enables us to generalize the standard DID estimand by defining an information set that can represent a set of multiple pre-treatment periods or other baseline covariates. We define selection bias as the mean-difference of the untreated potential outcome between the treatment and control groups. We introduce the concept of generalized difference-in-differences (GDID) estimand defined as the difference between the ordinary least squares estimand in the post-treatment period and a selection bias in that period, which is a correspondence of the available information set and the baseline period selection bias. Under the assumption that the treatment has no anticipatory effects, we identify the baseline period selection bias as the difference-in-means of that period's observed outcome between the treatment and control groups.
Second, we consider assumptions under which the above correspondence is known. Our main assumption states that the selection bias in the post-treatment period lies within the convex hull of all selection biases in the pre-treatment periods. We provide a sufficient condition for this assumption to hold. For example, we discuss and illustrate that this assumption may be plausible in economic settings where Ashenfelter1978's (Ashenfelter1978) dip is present. It is well documented that in such contexts, the PT assumption is usually not plausible (e.g., see AshenfelterCard1985, HeckmanSmith1999, Heckman_etal1999). Based on the baseline information set we construct, we provide an identified set for the ATT that always contains the true ATT under our identifying assumption, and also the standard DID estimand, given that we are using a weaker assumption than the PT assumption. If in fact PT holds in the pre-treatment periods, our bounds naturally collapse to the standard DID estimand. Importantly, we show how the baseline covariates can help define the correspondence and therefore help partially identify the ATT of interest. Unlike the standard DID framework where covariates are required to be time-invariant, our method allows for exogenous time-varying covariates. To the best of our knowledge, only few papers in the DID literature allow for time-varying covariates Caetano_etal2022, Shahnal2022. Our paper contributes to this literature. We provide multiple illustrative examples where the standard DID estimand does not identify the ATT while our bounds cover it.
Third, we discuss alternative ways of defining the correspondence. For example, when the pre-treatment periods selection biases show some clear pattern in terms of trends, we discuss how the researcher can model such a pattern and use that model to forecast the post-treatment period selection bias. In the same direction, we propose a class of criteria on the selection biases from the perspective of a policymaker that can achieve a point identification of ATT. We call this point estimand a policy-oriented GDID, as it may not necessarily have a causal interpretation.
Fourth, we propose an implementation procedure for our bounds. We provide a doubly-robust estimand in the presence of covariates in the post-treatment period. Our proposed confidence bounds are valid in the sense that they will cover the true identified set with a pre-specified probability. However, they may be too conservative. We believe that this inference procedure can be improved and leave this improvement for future research.
Fifth, we show how our framework extends to the DID with multiple treatment periods, and the synthetic control (SC) settings. On the one hand, we derive bounds on each post-treatment period ATT. This approach can help reveal the heterogeneity in the treatment effects over time. As before, if PT holds for the treatment status in each treatment period, our bounds collapse to a DID estimand. We can therefore identify the ATT in each treatment period. We also extend the method to the identification of more causal parameters, including those considered in CallawaySantAnna2021. The causal parameters we consider can also help reveal the dynamic effect of the treatment. On the other hand, instead of finding the optimal weights for elements in the donor pool to create a counterfactual synthetic control for the treated unit, we propose bounds on the ATT by considering each donor as a potential control unit. Here again, we use the convex hull of all pre-treatment periods selection biases as the identified set for the selection bias in the post-treatment period.
Finally, we illustrate the empirical relevance of our methodology by revisiting kresch2020, cawley2021SSB, and cai2016. We apply our method to investigate the causal effect of a 2007 reform in Brazil, that gave municipalities the ultimate authority to provide some services, on different types of investment. We find that the effect was less/not significant as initially found by the author. The main issue was that the pre-treatment periods selection biases were not stable over time. This makes the PT assumption less reliable in this application. On the other hand, cawley2021SSB examine the pass-through of a tax of two cents per ounce on sugar-sweetened beverages (SSB tax) enacted in Boulder, Colorado, using the standard DID framework. Both the DID method and our GDID bounds lead to the conclusion that the policy effect on the post prices is statistically significant. Furthermore, we also revisit cai2016 who investigates the impact of insurance provision on tobacco production using a household-level panel dataset provided by the Rural Credit Cooperative (RCC), the main rural bank in China. Using the DID approach and the robust GDID bounds, we conclude as the author that the effect of the insurance is positive on both the area and share of tobacco.
Our paper is closely related to two papers in the literature: ManskiPepper2018, and RambachanRoth2020. While these papers mainly focus on a trends-based / space-based relaxation of the PT assumption, our paper relies on a selection-based relaxation approach.\footnote{Our selection-based relaxation can use information over time (e.g., when the baseline information set is the set of pre-treatment periods) or across space (e.g., when the baseline information is the set of baseline covariates, which can include geographic region).} On the one hand, ManskiPepper2018 introduce bounded-variation assumptions that relax the PT assumption. Instead of requiring that the untreated potential outcomes for the treatment and control groups follow on average the same trends between the baseline and treatment periods, the authors assume that the absolute difference in trends is bounded by a known sensitivity parameter. When this parameter is equal to zero, their assumption reduces to the PT assumption. ManskiPepper2018's (ManskiPepper2018) approach is robust to violations of the PT assumption when the sensitivity parameter is big enough. Our approach is robust to violations of PT that can be captured in the pre-treatment periods but is not necessarily robust to post-treatment violations that cannot be captured in the pre-treatment periods. For example, when PT holds in the pre-treatment periods while it is actually violated in the post-treatment period, our set identification method coincides with the standard DID approach, which will not identify the ATT as the needed PT assumption does not hold. However, the choice of the sensitivity parameter in the ManskiPepper2018 bounding approach remains unclear. On the other hand, RambachanRoth2020 generalizes ManskiPepper2018's (ManskiPepper2018) bounding method by considering a large class of restrictions that impose that the post-treatment violations of parallel trends cannot be “too different”\footnote{In the terminology of RambachanRoth2020.} from the pre-trends. Our bounding strategy falls into this class of restrictions that the authors consider and can therefore be viewed as a special case of their approach. However, we tackle the problem with a different perspective and our identifying assumptions have not been considered in RambachanRoth2020. Like in ManskiPepper2018, the specific restrictions they consider require the knowledge of a sensitivity parameter, whose choice still remains unclear in their case. Moreover, our approach does not require an explicit choice of a sensitivity parameter. The sensitivity parameter is implicitly embedded in the baseline information set.
Our paper also contributes to the growing literature on sensitivity analysis in the DID framework. Freyaldenhoven_al2019 propose a method that estimates a policy effect using a two-stage least squares approach in a linear panel event-study design where unobserved confounds may be related both to the outcome and the the policy variable of interest. Their identification strategy relies on the existence of covariates related to the policy only through the confounds. Keele_al2019 develop a method of sensitivity analysis that allows researchers to quantify the amount of bias from time-varying confounders necessary to change a study’s conclusions in the DID model, relying on baseline covariates. In the same direction, Ye_al2022 propose a partial identification strategy that relaxes the PT assumption to a monotone trends assumption relying on two groups of control units whose outcomes relative to the treated units exhibit a negative correlation. Our approach does not require an existence of two control groups and our identifying assumption may still hold even their monotone trends assumption fails to hold. Similarly to our bounds, their identified set is of a union bounds form that involves the minimum and maximum operators. While our confidence bounds may be too conservative, they propose a novel bootstrap method to construct uniformly valid confidence bounds for the identified set and parameter of interest. It may be possible to implement their proposed method in our framework. We find our inference method attractive as it is easy to implement, especially when the baseline information set is discrete. Leavitt2020 develops an empirical Bayes' procedure that allows for other trend assumptions in the DID framework. On the other hand, BilinskiHatfield2020 and DetteSchumann2020 propose more reliable inference methods to detect meaningful violations of the PT assumption in the pre-treatment periods. Finally, by extending our proposed method to the multiple treatment periods setting, we contribute to the growing literature on the causal interpretation of event-study coefficients in two-way fixed effects models in the presence of staggered treatment timing and heterogeneous treatment effects (as in Borusyak_al2022; AtheyImbens2022; goodmanbacon2021ddtiming; CallawaySantAnna2021; deChaisemartinDHaultfoeuille2020; SunAbraham2021) when the PT assumption fails to hold in the pre-treatment periods. Building on Wooldridge2021's (Wooldridge2021) idea, we show how two-way fixed effects regression methods can help compute our confidence set when the baseline information set is the set of pre-treatment periods. It would be interesting to extend this developed approach to the changes-in-changes model considered in AtheyImbens2006, since the identifying assumptions may be sensitive to functional forms as is the case for the PT assumption RothSantanna2021. Some recent papers have also proposed alternative point-identification results in this setting Parkal2023, Wooldridge2022. Other papers propose point-identification and estimation results in DID settings where the standard PT assumption may be questionable, while relying on some additional assumptions Henderson2023, Richardson2023, Dukesal2022, Brown2023.
The remainder of the paper is organized as follows. Section (ref) presents the model and a preview of our approach and results, and it introduces the generalized DID concept. Section (ref) formally discusses the assumptions and the main identification results, while Section (ref) introduces the policy-oriented generalized DID concept. Section (ref) briefly discusses the implementation of the proposed bounds. Section (ref) presents two extensions of our approach. Section (ref) shows the practical relevance of the method through three empirical illustrations, and Section (ref) concludes. Proofs of the main results are relegated to the appendix.
Consider the following two-period model:
where the vector $(Y_0, Y_1,D, I_0, X_0, X_1)$ represents the observed data, while the vector $(Y_1(0), Y_1(1))$ is latent. In this model, the variables $Y_0, Y_1 \in \mathcal Y$ are respectively the observed outcomes in the baseline period 0 and the follow-up period 1, while $D\in \left\{0,1\right\}$ is the observed treatment that occurred between periods 0 and 1, $Y_1(0)$ and $Y_1(1)$ are the potential outcomes that would have been observed in period 0 had the treatment $D$ been externally set to 0 and 1, respectively. The variable $Y_0(0)$ is the potential outcome that is realized in the baseline period when no individual/unit was treated. As is common in the DID literature, model ((ref)) assumes that there is no anticipatory effect of the treatment, so that $Y_0(1)=Y_0(0)$. The set $I_0 \in \mathcal I_0$ contains information on baseline data, while $X_0 \in \mathcal X_0$ and $X_1 \in \mathcal X_1$ denote the vector of covariates in periods 0 and 1, respectively. The baseline information $I_0$ could be a subset of $X_0$ but does not have to be.
In this paper, we are interested in identifying the average treatment effect on the treated (ATT) defined as $$ATT\equiv \mathbb E[Y_1(1)-Y_1(0)\vert D=1].$$ We first focus on the case without covariates.
We start by defining the standard ordinary least squares (OLS) estimand, which is the same as the difference in means estimand, as $$\theta_{OLS} \equiv \mathbb E[Y_1\vert D=1] - \mathbb E[Y_1 \vert D=0].$$ We can rewrite this OLS estimand as the ATT plus a bias term. Indeed, we have
where $SB_t\equiv \mathbb E[Y_t(0) \vert D=1]-\mathbb E[Y_t(0) \vert D=0]$.
Equation ((ref)) shows that the standard OLS estimand in period 1 can be decomposed as equal to the ATT of interest plus a bias term that we call selection bias. Therefore, in order to identify the ATT with the help of the OLS estimand, we need to identify this selection bias. The main question we are asking at this point is how to obtain the selection bias $SB_1$. The literature provides at least two solutions to this problem. One can randomize the treatment and then get rid of the selection bias when there is full compliance, or one can rely on the PT assumption. While a successful randomized experiment yields zero selection bias ($SB_1=0$), it is often difficult and costly to implement (e.g., because of some ethical concerns, feasibility). On the other hand, the PT assumption could be too restrictive in some cases. For this reason, our approach aims at relaxing the PT assumption and provides credible bounds on the ATT instead of point-identifying this parameter.
Following Heckman_al1998, we reinterpret the PT assumption as a bias equality assumption: the selection bias in period 1 is equal to the selection bias in period 0, i.e., $SB_1=SB_0$, which is identified as the difference in the baseline outcome means between the treatment and control groups, under the no-anticipatory effects assumption. Indeed, we have
The terminology parallel trends is more appropriate when the untreated potential outcome mean has linear trends in the treatment and control groups. However, when the trends are nonlinear, they may not be parallel even if the mathematical definition of the parallel trends assumption holds. Indeed, equality of the selection biases in periods 0 and 1 is sufficient for the mathematical definition of parallel trends to hold, regardless of what the untreated potential outcome mean trends are for the treatment and control groups. To illustrate this, consider a simple version of model ((ref)) where
and the information set $\mathcal I_0$ is the set of two pre-treatment periods $\mathcal T_0$, i.e., $\mathcal I_0=\mathcal T_0=\{-1,0\}$. In this model, the selection bias $SB_t=(1-2 \vert t \vert + 2 t^2)(\alpha_1-\alpha_0)$ where $\alpha_1=\frac{\phi(1)}{1-\Phi(1)}\approx 1.53$ and $\alpha_0=-\frac{\phi(1)}{\Phi(1)} \approx -0.29$. Therefore, we have $SB_0= \alpha_1-\alpha_0 =SB_{-1}=SB_1$. So, the standard parallel trends assumption holds. Yet, the trends of $\mathbb E[Y_t(0)\vert D=1]$ and $\mathbb E[Y_t(0)\vert D=0]$ are not parallel. We refer to these trends as spurious parallel trends. Figure (ref) displays those trends for $\theta=2$. Because of the existence of spurious trends, when a researcher is doubtful about the validity of the PT assumption, a selection-based relaxation approach could be more appropriate than a trends-based approach in some circumstances. In this sense, we view our approach as a complement to the existing trends-based relaxation approaches. For instance, suppose that the time unit on the $x$-axis is the year, and a semestral dataset is available. If a researcher is interested in identifying the treatment effect at period $t=1/2$ (first semester), the standard DID estimand will fail to identify the causal effect, as $SB_{\frac{1}{2}} \neq SB_0$. Our proposed approach will be robust to this kind of spurious parallel trends as long as our information set includes the period $t=-1/2$.
Before we present the formal results, we heuristically show the intuition behind our main identification strategy.
Suppose the information $I_0$ contains two pre-treatment periods such that $\mathcal I_0=\{-1,0\}$. In general, when $SB_{-1} \neq SB_0$, it is difficult to believe that $SB_0=SB_1$. Note that none of the conditions implies the other. Yet, researchers often rely on this pre-test to check the plausibility of PT. Our approach is to assume that $SB_1$ lies within the convex hull of $\{SB_{-1},SB_0\}$, that is, $SB_1 \in [\min\{SB_{-1}, SB_0\},\max\{SB_{-1}, SB_0\}]$. Under our assumption, we obtain the following bounds on the ATT: $$ATT \in \left[\theta_{OLS}-\max\{SB_{-1}, SB_0\}, \theta_{OLS}-\min\{SB_{-1}, SB_0\}\right].$$ Hence, our bounding approach is robust to violations of parallel trends that can be captured in the pre-treatment periods. However, this does not ensure that our identifying assumption is valid. One would have to justify why the selection bias in period 1 would lie within the convex hull of the selection biases in pre-treatment periods. As can be seen, the standard DID estimand ($\theta_{OLS}-SB_0$) lies within our bounds.
Suppose now that the baseline period 0 is the only pre-treatment period for which a data is available, and we observe a baseline covariate $X_0$. For simplicity, assume $I_0=X_0$. Unlike the standard approach which requires $X_0$ to be equal to $X_1$ i.e., $X_0=X_1=X$ abadie2005, we allow $X_0$ to be different from $X_1$ in our framework. Yet, to illustrate our contribution over the existing approaches, we consider the simple case where $X_0=X_1=X \in \{x_0,x_1\}$. Define $SB_t(x)\equiv \mathbb E[Y_t(0)\vert D=1, X=x]-\mathbb E[Y_t(0)\vert D=0, X=x]$. Existing methods assume $SB_0(x)=SB_1(x)$, while ours assumes $SB_1(x)\in [\min\{SB_0(x_0), SB_0(x_1)\},\max\{SB_0(x_0), SB_0(x_1)\}]$. As we can see, we allow for $SB_0(x) = SB_1(x)$ for some $x$, $SB_0(x)\neq SB_1(x)$ for all $x$, $SB_0(x) = SB_1(x')$ for some $(x,x')$ or $SB_0(x)\neq SB_1(x')$ for all $(x,x')$.
Given our above assumption, we partially identify $ATT(x)\equiv \mathbb E[Y_1(1)-Y_1(0)\vert D=1, X=x]$ as follows: $$ATT(x) \in \left[\theta_{OLS}(x)-\max\{SB_0(x_0), SB_0(x_1)\}, \theta_{OLS}(x)-\min\{SB_0(x_0), SB_0(x_1)\}\right],$$ where $\theta_{OLS}(x) \equiv \mathbb E[Y_1\vert D=1, X=x] - \mathbb E[Y_1 \vert D=0, X=x]$. Hence, we can integrate the bounds on $ATT(x)$ over the conditional distribution of $X$ given $D=1$ to obtain bounds on $ATT$.
The following example shows a data generating process where our identifying assumption holds in a situation where we only have two periods and a time-invariant covariate is available.
Interpretation of our assumption. Let $X$ be the variable gender. The standard assumption $SB_0(x)=SB_1(x)$ states that the selection bias for females in period 0 is equal to selection bias for females in period 1, and similarly for males. Our assumption states that the selection bias for females (resp. males) in period 1 lies between those for males and females in period 0. Our assumption allows the selection bias for females in period 1 to be equal to that of males in period 0, and vice versa. Furthermore, we allow for the possibility that the selection bias for females (resp. males) in period 1 be different from those for males and females in period 0.
As we explain above, the standard DID estimand is defined as the difference between the OLS estimand in period 1 and the selection bias in period 0: $\theta_{DID} \equiv \theta_{OLS}-SB_0$. We introduce a generalized version of this estimand.
In the above definition, if $SB_1(SB_0,\mathcal I_0)=SB_0$, the generalized DID estimand is the same as the standard DID estimand. Note however that $SB_1(SB_0,\mathcal I_0)$ is allowed to be a set of values. In such a case, the generalized DID estimand will be a set instead of a single value.
In the next section, we formally discuss our assumptions and the main results.
In this section, we state our identifying assumptions and present our main results.
Let us first consider the simple case with no covariates in the model. We now state our main assumption.
Assumption (ref) is weaker than the standard “parallel/common trends” assumption. Indeed, if $\mathcal{I}_0$ is the singleton of a single baseline information $I_0=\{\iota_0\}$, then Assumption (ref) is equivalent to $SB_1= SB_0(\iota_0),$ which is equivalent to the parallel trends assumption, as discussed earlier. For example, suppose that the information set contains two pre-treatment periods such that $\mathcal I_0=\{-1,0\}$. The PT assumption $SB_1=SB_0$ implies $$SB_1 \in \left[\min\{SB_{-1},SB_0\}, \max\{SB_{-1},SB_0\}\right],$$ which is equivalent to Assumption (ref) in this example.
Instead of assuming parallel trends or bias equality, we assume that the convex hull of the set of selection biases in the pre-treatment periods is stable over time. This assumption could be violated in many situations. For example, when the information set is ordered (e.g., time, ordered covariates) and the baseline selection biases change monotonically with the elements in the information set, then the selection bias in the follow-up period will likely be outside the set of pre-treatment periods selection biases. In such a case, Assumption (ref) may not hold. We propose a solution for this context in Section (ref). Furthermore, when the potential outcome in the treatment period is linear (but not a random walk) in the baseline period (e.g., $Y_1(0)=\alpha Y_0(0) + \varepsilon$, where $\alpha \neq 1$, and $\varepsilon$ is exogenous), then neither parallel trends nor bias set stability holds.
Note that the set $\mathcal{I}_0$ could contain all pre-treatment periods, observed baseline characteristics, or information from other data sources. For example, suppose that $\mathcal I_0$ contains gender. The standard parallel trends assumption conditional on gender states that the selection bias for males in period 0 would be the same for males in period 1, and similarly for females. As discussed above, our assumption (ref) allows the selection bias for females in period 1 to be equal to that for males in period 0, and vice versa. We believe that this latter assumption is more flexible. Suppose that we collect information from multiple data sources that may not be representative of the population of interest. In this situation, $I_0$ could denote a categorical random variable for each data set. For example, for a study on the US population, a researcher can combine data from multiple states. Each of these data can be considered as a piece of information $\iota_0$.
Assumption (ref) defines a particular correspondence $SB_1(SB_0,\mathcal I_0)$ for the selection bias $SB_1$. Plugging this in the generalized DID estimand yields the following bounds on the ATT.
The bounds in Proposition (ref) are never empty, as they always contain the standard DID estimand under the parallel trends assumption. However, they may not contain the OLS estimand in period 1, $\theta_{OLS}$, as 0 may not lie within the set $\Delta_{SB_0}$. If all pre-treatment periods selection biases are equal, i.e., $SB_0(\iota_0)=SB_0$ for all $\iota_0$, then our bounds collapse to a point, the standard DID estimand. In case the information set $\mathcal I_0$ is the set of pre-treatment periods, the above bounds are robust to violations of PT that can be captured in the pre-treatment periods. In this sense, our method can be seen as a way of salvaging (using the language of Masten_al2021) the standard DID model from violations of PT in the pre-treatment periods. However, our identification strategy does not rely on finding the falsification frontier.
An economic setting where our Assumption (ref) may hold would be the evaluation of a job training program in which the so-called Ashenfelter1978 dip occurs. As pointed out by Ashenfelter1978, individuals who participate in a job training program are usually those who have experienced a decline in employment and earnings prior to their enrollment in the program. If the decline were transitory, such individuals would normally experience a rebound in employment and earnings, even if they did not participate in the program. This phenomenon would likely make the PT assumption violated. We provide Example (ref) below that shows this pattern as depicted in Figure (ref). This Ashenfelter1978 dip has also been documented in the evaluation of incarceration on subsequent earnings and employment (e.g., see Lalonde_Cho2008, Jung2011).
The example below presents a data generating process (DGP) where the PT assumption fails, while Assumption (ref) holds. The DGP is inspired from a random-growth model with a factor structure discussed in HeckmanHotz1989. It shows how informative the bounds can be.
One question that comes to people's mind when they think about Assumption (ref) is: what are the conditions under which this assumption will hold? To this question, we provide a sufficient condition on the data generating process for the counterfactual untreated potential outcome under which this assumption holds.
Assumption (ref).((ref)) postulates a factor (or interactive fixed effects) structure for the untreated potential outcome, which is commonly used in applied research. Assumption (ref).((ref)) imposes some symmetry condition on the factor function $g_t$. If the function $g_t(.)$ is even in $t$, then the selection bias is symmetric in $t$, and the set of biases before the baseline period will be identical to the set of biases after the baseline period. In such a case, our Assumption (ref) will hold. This symmetry condition is similar to the intuition behind the symmetric DID discussed in AshenfelterCard1985. This assumption can be relaxed. See Example (ref) in the appendix where $g_t=t$ is odd, but our Assumption (ref) holds. Assumption (ref).((ref)) postulates that the selection into treatment is function of the time-invariant unobservables in the model.
Our main assumption (Assumption (ref)) will generally hold in settings where there exist some common life-cycle factors that affect the untreated potential outcome. These life-cycle factors could translate into the symmetry condition or some periodicity in the potential outcome. Although we provide a sufficient condition for our main assumption, we believe that a deeper understanding of it through its connection to structural economic choice models, as discussed in Ghanem_al2022 and Marx_al2022 in the context of the PT assumption, would be an interesting direction for future research. A more general sufficient condition is provided below. The factor structure considered in Assumption (ref) is a special case of Assumption (ref) below.
For example, Assumption (ref) holds in the following DGPs: $Y_t=\sqrt{t^2+U}+\varepsilon+\theta D*t\mathbbm{1}\{t\geq 0\}$, $D=\mathbbm{1}\{U \geq 1\}$, $U \sim N(0,\sigma^2)$, or $Y_t=\sqrt{(t+2)(t-1)+U}+\varepsilon+\theta D*t\mathbbm{1}\{t\geq 0\}$.
In this subsection, we include covariates in the analysis. We allow the baseline characteristics $X_0$ to be different from those in the follow-up period $X_1$. We denote the baseline information by $I(X_0)$ to explicitly show that it depends on the baseline covariates $X_0$. For the sake of clarity of the exposition, assume $I(X_0)=X_0$.
Define
The main idea behind Assumption (ref) is to use the baseline characteristics to help identify the set of possible values for the selection bias in the treatment period. The intuition is that observing different realizations of the baseline selection bias $SB_0$ can inform us about the range of possible values that the treatment period selection bias $SB_1$ can take. Assumption (ref) implies that the convex hull of all selection biases in the baseline period 0 is the same as that of all possible selection biases in period 1. Note that Assumption (ref) is weaker than the standard conditional PT in abadie2005, heckman1997, SantAnna_etal2020, etc. It allows for time-varying covariates as in Caetano_etal2022 but is different from their conditional PT assumption.
The next assumption is a common support assumption that requires that conditional on each period covariates, there exits at least a nonnegligible set of individuals in both treatment and control groups that share these characteristics. This assumption is standard when the covariates are time-invariant.
The identification results are summarized in Proposition (ref) below.
Using the results in Proposition (ref), we can then obtain sharp bounds on $ATT$ by integrating the bounds over the conditional distribution of $X_1$ in the treatment group ($D=1$):
The following Proposition (ref) provides a doubly robust estimand for $\int\theta_{OLS}(x_1)d F_{X_1\vert D=1}(x_1)$. This result is probably achieved in the literature, but since we could not find a closed-form expression of a doubly robust estimand for this quantity, we provide an estimand along with its proof for completeness.
Note that the proposed estimand $\tau^{DR}$ is equal to the desired quantity $\int\theta_{OLS}(x_1)d F_{X_1\vert D=1}(x_1)$ even if either the propensity score function or the conditional outcome mean function is misspecified. However, if both functions are misspecified, $\tau^{DR}$ is generally different from $\int\theta_{OLS}(x_1)d F_{X_1\vert D=1}(x_1)$.
As in the previous section, we are going to provide a sufficient condition under which Assumption (ref) holds. We slightly modify Assumption (ref) to the following.
Assumption (ref) is a modified version of Assumption (ref) to allow the factor to depend on some time-varying covariate $X_t$. Assumption (ref).((ref)) postulates that in the interactive fixed effects structure for the untreated potential outcome, the time-varying factor is determined by some potentially time-varying covariate $X_t$. Assumption (ref).((ref)) imposes some monotonicity condition on the factor function $g$. It also imposes a support condition on the time-varying covariate $X_t$ over time, which holds if the covariate is not changing over time. Assumption (ref).((ref)) is the same as Assumption (ref).((ref)) and postulates that the selection into treatment is function of the time-invariant unobservables in the model.
We can broaden conditions (ref).((ref)) and (ref).((ref)) in Assumption (ref) by replacing them by $Y_t(0)=g_t(X_t) \lambda(U)+\gamma(V)+\varepsilon_t$ along with the other restrictions, and $Supp(g_1(X_1)) \subseteq Supp(g_0(X_0))$, respectively. The function $g$ in the potential outcome model now has a subscript $t$, which allows $X_t$ to be the same random variable across time periods ($X_0=X_1$).
Before we move on, let us elaborate on our contribution to the literature. As can be seen from Propositions (ref) and (ref), our approach does not require the support of the outcome variable to be bounded as it is customary in the literature on partial identification. Furthermore, the approach does not rely on a sensitivity parameter as in ManskiPepper2018, and RambachanRoth2020. Our bounds can still be informative in situations where there are only two periods, 0 (baseline) and 1 (follow-up), as long as there exists other information available from observed baseline characteristics, or multiple data sources that may not be representative of the target population. Below, we provide a deeper comparison of our method to that of RambachanRoth2020 when our information set contains only pre-treatment periods.
First, for the sake of simplicity suppose the information set $I_0$ contains two pre-treatment periods $-1$ and $0$, such that $\mathcal I_0=\{-1,0\}$. Define $\delta\equiv(\delta_{-1},\delta_1)'$, where
Observe that $\delta_1=SB_1-SB_0$, and $\delta_{-1}=SB_{-1}-SB_0$.\footnote{In the terminology of RambachanRoth2020, $\delta_{-1}=\delta_{pre}$ and $\delta_{1}=\delta_{post}$.} Note that $\delta_{-1}$ is identified. Lemma 2.1 in RambachanRoth2020 provides a general characterization of the ATT if a researcher is willing to make a restriction that $\delta_1$ belongs to a closed and convex set. In our setting, we assume that $SB_1 \in [\min\{SB_{-1}, SB_0\},\max\{SB_{-1}, SB_0\}]$, which implies $\delta_{1} \in \left[\min\{SB_{-1}-SB_0,0\}, \max\{SB_{-1}-SB_0,0\}\right]$. Therefore, we can recast our framework in theirs where $\delta \in \{SB_{-1}-SB_0\}\times \left[\min\{SB_{-1}-SB_0,0\}, \max\{SB_{-1}-SB_0,0\}\right],$ which is a closed and convex set. Hence, our approach can be viewed as a special case of their method. However, they do not consider the type of restrictions we consider in this paper. We now compare our assumptions to the restrictions considered in RambachanRoth2020.
The differential trends evolve smoothly over time with slope changing by no more than $M$ between consecutive periods:
where $\delta_0$ is normalized to be equal to zero. We then have $\Delta^{SD}(M)\equiv \left\{\delta: \vert \delta_1 + \delta_{-1} \vert \leq M \right\}$. The parameter $M\geq 0$ is like a sensitivity parameter and governs the amount by which the slope of the differential trends can change between consecutive periods.
Under the smoothness restriction, we obtain the following bounds on the selection bias $SB_1$:
Our bounding approach yields the following bounds on $SB_1$:
In Appendix (ref), we show that if $SB_{-1} \neq SB_0$, there exists no value of $M$ such that the above two sets of bounds on $SB_1$ coincide. Furthermore, we show that there exist no values of $M$ for which RambachanRoth2020's (RambachanRoth2020) bounds are tighter than ours, while there exist values of $M$ for which our bounds are tighter than theirs $(M > 2\vert SB_0-SB_{-1}\vert)$.
This approach bounds the worst-case post-treatment violation of parallel trends in terms of the worst-case violation in the pre-treatment period:
where $\bar{M}\geq 0$ behaves as a sensitivity parameter. This implies the following bounds on $SB_1$:
In Appendix (ref), we show that if $SB_{-1} \neq SB_0$, there exists no value of $\bar{M}$ such that the above bounds on $SB_1$ coincide with ours. When $\bar{M }>1$, our bounds are tighter than RambachanRoth2020's (RambachanRoth2020), and there exist no positive values of $\bar{M}$ for which their bounds are tighter than ours.
Second, our approach covers the case where no pre-treatment trend exists, but there are multiple elements (baseline characteristics/data sets) available in the information set in period 0. Their methodology is silent about such a case.
Third, our approach does not require the knowledge of a sensitivity parameter, while the two restrictions they consider do. How to choose the values of the sensitivity parameters $M$ and $\bar{M}$ remains unclear in their approach.
In this subsection, we study the existence of possible discordancy between the restrictions we consider in the paper and those considered in RambachanRoth2020. We find that under the smoothness restrictions, when $SB_{-1} \neq SB_0$, the two bounds are discordant if $M < \vert SB_0-SB_{-1}\vert$, i.e., their intersection is empty if $M < \vert SB_0-SB_{-1}\vert$. Kedagni_etal2020 pointed out that when a full model is rejected, researchers should be cautious about the way they relax the model to avoid this kind of situations. We then recommend researchers not to use the smoothness restrictions with values of $M$ less than $\vert SB_0-SB_{-1}\vert$. This scenario never happens under the bounding relative magnitudes restriction, i.e., there exists no possible discordancy between our bounds and those obtained under this latter restriction. The two bounds always overlap.
The main question we are trying to answer is how to obtain the selection bias $SB_1$. Given the baseline information $I_0$, we are going to assume that the decision maker will choose the selection bias $SB_1$ in such a way that a loss function is minimized. By plugging such an optimal selection bias $SB_1$ into the definition of the generalized DID, we obtain what we call a policy-oriented generalized difference-in-differences (PO-GDID) estimand. This estimand may not have a causal interpretation, but it may help the policy-maker in her decision making process.
In this paper, we consider the class of $p$-norm losses defined as: $$\mathcal L_p (SB_1,SB_0, I_0)=\left(\mathbb E_{I_0}\left[\vert SB_1 - SB_0(I_0)\vert ^p \right]\right)^{1/p},$$ where $1 \leq p \leq \infty$. We are going to derive the optimal selection bias $SB_1$ for $p\in \left\{1,2,\infty\right\}$. We consider those special loss functions because the solutions to the optimization problem have closed-form expressions. Other loss functions can also be considered.
$\mathcal L_1 (SB_1,SB_0, I_0)=\mathbb E_{I_0}\left[\vert SB_1 - SB_0(I_0)\vert \right]$.
Given this $L1$ loss function, under Assumption (ref), the decision maker solves the following optimization problem:
The optimal decision is to set the selection $SB_1$ to be equal to the median selection bias in the baseline period, i.e., $SB_1=Med_{I_0}(SB_0(I_0))$. In such a case, the policy-oriented generalized DID estimand is given by $$\theta_{PO-GDID}=\theta_{OLS}-Med_{I_0}(SB_0(I_0)).$$
$\mathcal L_2 (SB_1,SB_0, I_0)=\left(\mathbb E_{I_0}\left[\vert SB_1 - SB_0(I_0)\vert ^2 \right]\right)^{1/2}$.
Minimizing the RMSE is equivalent to minimizing the mean square error (MSE). Therefore, under Assumption (ref), the decision maker solves the following optimization problem:
This yields an optimal decision for the selection $SB_1$ to be set equal to the average selection bias in the baseline period, i.e., $SB_1=\mathbb E_{I_0}[SB_0(I_0)]$. Hence, we have $$\theta_{PO-GDID}=\theta_{OLS}-\mathbb E_{I_0}[SB_0(I_0)].$$
$\mathcal L_{\infty} (SB_1,SB_0, I_0)=\text{ess}\sup_{\mathcal I_0}\vert SB_1 - SB_0(I_0)\vert$, where $\text{ess}\sup$ denotes essential supremum and is defined as follows: $$\text{ess}\sup_{\mathcal I_0} f=\inf\left\{M: \mathbb P(\iota_0 \in \mathcal I_0: f(\iota_0) \leq M)=1\right\}.$$
For simplicity, assume $\text{ess}\sup_{\mathcal I_0}\vert SB_1 - SB_0(I_0)\vert =\sup_{\iota_0 \in \mathcal I_0}\vert SB_1 - SB_0(\iota_0)\vert$. Then
Therefore, the minimum of $\mathcal L_{\infty} (SB_1, SB_0, I_0)$ is obtained when the two arguments of the max function are equal, i.e., $SB_1 - \inf_{\iota_0 \in \mathcal I_0}SB_0(\iota_0) = \sup_{\iota_0 \in \mathcal I_0}SB_0(\iota_0)-SB_1$. This implies $SB_1=\frac{1}{2}(\inf_{\iota_0 \in \mathcal I_0}SB_0(\iota_0)+\sup_{\iota_0 \in \mathcal I_0}SB_0(\iota_0))$, and $\mathcal L_{\infty} (SB_1)=\frac{1}{2}(\inf_{\iota_0 \in \mathcal I_0}SB_0(\iota_0)+\sup_{\iota_0 \in \mathcal I_0}SB_0(\iota_0))$.
This optimization problem with the $L\infty$ loss is equivalent to a minimax criterion, and yields the mid-point of the bounds on $SB_1$ stated in Assumption (ref). Hence, the PO-GDID estimand is given by $$\theta_{PO-GDID}=\theta_{OLS}-\frac{1}{2}(\inf_{\iota_0 \in \mathcal I_0}SB_0(\iota_0)+\sup_{\iota_0 \in \mathcal I_0}SB_0(\iota_0)).$$
Note that in all cases, if the information in the baseline period is a singleton, then the optimal $SB_1$ is the selection bias in the baseline period $SB_0$, which is equivalent to the parallel trends assumption. Unlike the PO-GDID estimand obtained from $L1$ and $L2$ loss functions, that obtained from the $L\infty$ loss function does not require the knowledge of the distribution of the information $I_0$ but only its support and is easy to compute. However, when the distribution of $SB_0(I_0)$ is uniform over $\left[ \inf_{\iota_0\in \mathcal I_0} SB_0(\iota_0), \sup_{\iota_0 \in \mathcal I_0} SB_0(\iota_0)\right]$, then the optimal selection bias $SB_1$ is the same in all three cases.
Let $\Lambda$ denote the set of possible distributions for $SB_0(I_0)$, and $SB_1(SB_0,\lambda)$ denote the optimal selection bias in period 1 given the distribution $\lambda \in \Lambda$ for $SB_0(I_0)$. Define $ATT_{\lambda} \equiv \theta_{OLS}-SB_1(SB_0,\lambda)$.
The following lemma holds.
A sufficient condition for the policy-oriented generalized DID estimand to be equal to the ATT is that the potential outcome $Y_1(0)$ satisfies: $$Y_1(0)\sim \mathbb E[Y_1\vert D=0]+SB_1D +\varepsilon,$$ where $\mathbb E[\varepsilon \vert D]=0.$
Suppose that the baseline information set $\mathcal I_0$ is ordered (e.g., a set of multiple pre-treatment periods $\mathcal I_0=\{-T_0, -T_0+1, \ldots, -1, 0\}$ or a continuous baseline covariate $X_0$ like age). We can regress $SB(I_0)$ on $\{I_0, I_0^2, \ldots\}$ and use this regression to predict $SB_1$.
For instance, if $\mathcal I_0=\{-T_0, -T_0+1, \ldots, -1, 0\}$ and the selection bias $SB(I_0)$ is increasing over time, Assumption (ref) may not hold. The researcher could instead use this increasing trend information about the selection bias to forcast the next period selection bias $\widehat{SB}_1$.
We briefly describe our estimation and inferential method. We assume that the information set $\mathcal I_0$ is finite. We can write the robust DID bounds $\Theta_I$ as the convex hull of the doubly-robust DID estimands as follows:
We can then take the convex hull of the confidence intervals of all DID estimands $\tau^{DR} - SB_0(\iota_0)$ to obtain valid confidence bounds for $\Theta$. More precisely, the confidence bounds can written as
where $CI^{1-\alpha}_{LB}(\tau^{DR}-SB_0(\iota_0))$ (resp. $CI^{1-\alpha}_{UB}(\tau^{DR}-SB_0(\iota_0))$) denotes the lower (resp. upper) bound of the $(1-\alpha)$-confidence interval of the parameter $\tau^{DR}-SB_0(\iota_0)$. But, these confidence bounds could be too conservative.
The proof of validity of this procedure is provided in Appendix (ref). The argument is similar to that of Berger1996 for union bounds. Indeed, Berger1996 showed that the union of the confidence regions has at least the same coverage rate as each confidence region. The confidence bounds in (ref) are similar to those in KolesarRothe2018 derived in a regression discontinuity design setting.
To implement these confidence bounds, we first estimate the propensity score function $P(X_1)$ (e.g., logit specification) and the outcome regression function $\mu_0(X_1)$ (e.g., linear or quadratic specification). Second, in order to obtain correct standard errors for each estimator $\hat{\tau}^{DR}-\widehat{SB}_0(\iota_0),$ we use a bootstrap method.\footnote{Note that we do not bootstrap $\min_{\iota_0 \in \mathcal I_0}\{\hat{\tau}^{DR}-\widehat{SB}_0(\iota_0)\}$ nor $\max_{\iota_0 \in \mathcal I_0}\{\hat{\tau}^{DR}-\widehat{SB}_0(\iota_0)\}$. As pointed out by Fang_al2019, the standard bootstrap is inconsistent in this case since the limiting distributions of these estimators are not Gaussian.}
In this subsection, we generalize our analysis to a setting where the treatment receipt occurs at multiple periods. We consider the following multiple treatment periods model:
where $Y_t$ denotes the observed outcome in period $t$, $D_t$ is the observed treatment status in period $t$ with $D_0=0$ by definition, while $Y_t(0,d_1, \ldots, d_T)$ is the potential outcome when the treatment path $(D_0,D_1, \ldots, D_T)$ is externally set to $(0,d_1,\ldots, d_T)$.\footnote{See Robins1986, Robins1987, and Han2021 for a similar definition of the potential outcome model.} Under the no-anticipation assumption, we have $Y_0(0,d_1,\ldots, d_T) =Y_0(0)$ for all $(d_1,\ldots,d_T) \in \{0,1\}^T.$ We assume that individuals do not anticipate any effects of the treatment before it occurs for the first time. However, we allow the individuals to anticipate the effects of the treatment for the rest of the period. This assumption is less restrictive than the commonly used no-anticipatory effects assumption.
In the above framework, the parameter of interest is the average treatment effect on the treated group following the path $(0,d_1',\ldots, d_T')$ to $(0,d_1,\ldots, d_T)$ in period $t$, which is defined as:
This parameter may help reveal some dynamic effect of the treatment. For example, in a non-staggered design framework, the parameter $ATT_t[(0,\ldots,0,d_s=0,0,\ldots, 0)\rightarrow (0,\ldots,0,d_s=1,0,\ldots, 0)]$ measures the dynamic effect of the treatment in period $t$ on people who were only treated in period $s$ compared to the status where they would have never been treated. Note that this setting requires a panel structure in the data. In a staggered design setting, the average treatment effect in period $t$ on units who are treated for the first time in period $g$ could be an interesting parameter, as considered in CallawaySantAnna2021: $$ATT_t[(0,\ldots,0,d_g=0,0,\ldots, 0)\rightarrow (0,\ldots,0,d_g=1,1,\ldots, 1)].$$
Similarly to what we have in the one post-treatment setting, we can write the difference-in-means estimand $(\theta_{DIM}^t)$ between the two groups $(0,d_1',\ldots, d_T')$ and $(0,d_1,\ldots, d_T)$ in period $t$ as:
where $SB_t[(0,d_1',\ldots, d_T')\rightarrow (0,d_1,\ldots, d_T)]\equiv \mathbb E\left[Y_t(0,d_1',\ldots, d_T') \vert (D_0,D_1, \ldots, D_T)=(0,d_1,\ldots, d_T)\right]-\mathbb E\left[Y_t(0,d_1',\ldots, d_T') \vert (D_0,D_1, \ldots, D_T)=(0,d_1',\ldots, d_T')\right]\equiv SB_t.$ We extend Assumption (ref) to the current setting.
Assumption (ref) is a generalization of Assumption (ref). In the appendix, we provide a sufficient condition for it to hold (Assumption (ref)).
One could be interested in a weighted average of all time periods treatment effects, i.e., $ATT[(0,d_1',\ldots, d_T')\rightarrow (0,d_1,\ldots, d_T)]=\sum_{t=1}^T \omega_t ATT_t[(0,d_1',\ldots, d_T')\rightarrow (0,d_1,\ldots, d_T)]$, with pre-specified weights $\omega_t \in [0,1]$. Bounds on this weighted average $ATT$ can then be obtained as follows:
For example, one can set $\omega_t=\frac{n_t}{\sum_{t=1}^T n_t}$, where $n_t$ denotes the cardinality of the treatment group in period $t$.
Another parameter that could be of interest is average treatment effect on people who are ever treated over the treatment period:
When the outcome variable is only a function of the current period treatment status, we denote by $Y_t(d_t)$ the potential outcome in period $t$ as $Y_t(0,d_1, \ldots, d_t, \ldots, d_T)=Y_t(0,d_1', \ldots, d_t, \ldots, d_T')$ for all $(0,d_1, \ldots, d_t, \ldots d_T)$ and $(0,d_1', \ldots, d_t, \ldots d_T')$. In such a case, without ambiguity, we denote $ATT_t=\mathbb E[Y_t(1)-Y_t(0) \vert D_t=1]$, $\theta^t_{OLS}=\mathbb E[Y_t \vert D_t=1]-\mathbb E[Y_t \vert D_t=0]$, $SB_t=\mathbb E[Y_t(0) \vert D_t=1]-\mathbb E[Y_t(0) \vert D_t=0]$, and $SB^t_{\iota_0}=\mathbb E[Y_{\iota_0}(0) \vert D_t=1]-\mathbb E[Y_{\iota_0}(0) \vert D_t=0]$.
Below, we propose a DGP in which parallel trends holds for each period.
In the next example, we propose a DGP in which PT does not holds, but Assumption (ref) does.
It is important to note that the above framework applies to both staggered and non-staggered designs. In a non-staggered design DID framework, we have $2^T$ possible treatment paths, while in the staggered design case there are $T+1$ possible treatment paths.
Without covariates, our identified set for $ATT_t[(0,d_1',\ldots, d_T')\rightarrow (0,d_1,\ldots, d_T)]$ can be computed using a two-way fixed effects (TWFE) regression approach. Suppose that we observe $T$ treatment periods. Define $D^g\equiv \mathbbm{1}\{(D_0,D_1,\ldots,D_g,\ldots,D_T)=(0,0,\ldots,d_g=1,\ldots,1)\}$, and $D^0\equiv \mathbbm{1}\{(D_0,D_1,\ldots,D_T)=(0,0,\ldots,0)\}$. Consider the following regression for $t\in \{0,1,\ldots,T\}$
where the subscript $i$ refers to individual $i$, and $i=1\ldots,N$.
We have
Then,
Therefore under PT, $\mathbb E[\varepsilon_{is} \vert D^g_i=1]-\mathbb E[\varepsilon_{is} \vert D^0_i=1]=\mathbb E[\varepsilon_{i0} \vert D^g_i=1]-\mathbb E[\varepsilon_{i0} \vert D^0_i=1],$ and we have
That is, $\theta_{DIM}^s(D^g=1)-SB_0^s(D^0=1)=\theta^g_s$. For illustration, see Example (ref) in the appendix. This result is similar to the idea developed in Wooldridge2021 when PT holds.
Now, let us consider the case where PT may not hold. Suppose that the information set $\mathcal I_0$ is the set of pre-treatment periods. For each $\iota_0 \in \mathcal I_0$, we can run the TWFE regression for $t\in \{\iota_0,1,2, \ldots, T\}$. We then obtain a 95% confidence interval for $CI^{\iota_0}(\hat{\theta}^g_s)$ for $\theta^{g,\iota_0}_s$ from the TWFE regression. Therefore, we can obtain a 95% CI for $ATT_s(D^g=1 \rightarrow D^0=1)$ as $$\left[\min_{\iota_0 \in \mathcal I_0}CI^{\iota_0}_{LB}(\hat{\theta}^g_s), \max_{\iota_0 \in \mathcal I_0}CI^{\iota_0}_{UB}(\hat{\theta}^g_s)\right],$$ where $CI^{\iota_0}_{LB}(\hat{\theta}^g_s)$ and $CI^{\iota_0}_{UB}(\hat{\theta}^g_s)$ denote the lower and upper bounds on the confidence interval for $\theta^{g,\iota_0}_s$, respectively.
Let
where $X = (X_1, \cdots, X_T)$. Then, the result in Proposition (ref) holds, except that we replace $\theta_{OLS}(x_1)$ by $\theta_{DIM}^t(g, x),$ where $x=(x_1, \cdots, x_T)$.
The following proposition holds.
Suppose we observe $J+1$ units, and without loss of generality only the first unit is exposed to the intervention. Let $\mathcal J=\{2, \ldots, J+1\}$ denote the donor pool. For simplicity, suppose first that we only have two periods, such that $\mathcal I_0=\{0\}$. We write the model as:\footnote{One can alternatively assume that $Y_1(0) \equiv \sum_{j\in \mathcal J} \lambda_j Y_1^j(0) + \varepsilon_1(0)$. As long as $\varepsilon_1(0)$ is exogenous, i.e., $\varepsilon_1(0)\ \rotatebox[origin=c]{90}{$\models$}\ D$, our approach would work, since we only need $\mathbb E[Y_1(0)] = \sum_{j\in \mathcal J} \lambda_j \mathbb E[Y_1^j(0)].$}
Define $ATT^j=\mathbb E[Y_1(1)-Y_1^j(0) \vert D=1]$. We can check that $ATT=\sum_{j\in \mathcal J} \lambda_j ATT^j$.
Before explaining how our approach can be extended to this synthetic control (SC) framework, we briefly discuss the SC method. Abadie_etal2003 and Abadie_etal2010 propose to choose $\lambda_2$, \ldots, $\lambda_{J+1}$ so that the resulting synthetic control best resembles the pre-intervention values for the treated unit of predictors of the outcome variable, subject to the restriction that the weights are nonnegative and sum to one. There are at least two issues with their approach. First, the weights obtained using their approach may not be the same weights that we are looking for in the intervention period. The approach implicitly relies on the assumption that the weights are stable across covariates and also between baseline and treatment periods. Second, a solution to their problem may not exist Shial2023. Our approach does not suffer from these above issues.
Each donor $j \in \mathcal J$ is a potential control group for the treatment group: $\lambda_j \geq 0$ for all $j\in \mathcal J$, and $\sum_{j \in \mathcal J} \lambda_j=1$. However, we do not know the weights $\lambda_j \geq 0$ for any donor $j$. Assuming that the selection bias when considering each donor $j$ as a control in period 1 lies within the convex hull of all selection biases in period 0, we obtain the worst-case bounds for the ATT as: $\left[\min_j \underline{\theta}_{ATT}^j, \max_j \overline{\theta}_{ATT}^j\right]$, where $\underline{\theta}_{ATT}^j=\theta^j_{OLS}-\max_j SB^j_0$, $\overline{\theta}_{ATT}^j=\theta^j_{OLS}-\min_j SB^j_0$, and $SB^j_t=\mathbb E[Y_t(0)\vert D=1]-\mathbb E[Y^j_t(0)]$. Indeed, we have:
where the third equality holds from the definition of $Y_1(0)$, the fourth holds from the definition of the potential outcome model, the fifth holds because $Y_1^j \equiv Y_1 \vert J=j$, and $\sum_{j \in \mathcal J} \lambda_j=1$. We are abusing the notation by considering $J$ as a random variable.
Similarly, we can show that $\theta_{OLS}=ATT + \sum_{j \in \mathcal J} \lambda_j SB^j_1$. Therefore, $$ATT=\sum_{j \in \mathcal J} \lambda_j (\theta_{OLS}^j - SB^j_1).$$ Hence, under the assumption $SB^j_1 \in [\min_j SB^j_0, \max_j SB^j_0]$, the above bounds for the ATT are valid.
When the information set $\mathcal I_0$ has more than one element, the bounds on the selection bias $SB^j_1$ become: $SB^j_1 \in \left[\min_{\iota_0 \in \mathcal I_0}\min_j SB^j_0(\iota_0), \max_{\iota_0 \in \mathcal I_0} \max_j SB^j_0(\iota_0)\right]$.
In this section, we illustrate our framework using some empirical examples. First, using the dataset from kresch2020, we show our robust GDID bounds and our PO-GDID estimands with the various loss functions we discussed as well as the point-estimand obtained by forcasting the treatment period selection bias using a linear projection method. There are five pre-treatment periods that represent the information set we consider. We use our doubly-robust estimand throughout this illustration. Secondly, we consider the application of cawley2021SSB where we do not have covariates to control for. We construct the information set using the two pre-treatment periods available in the data, and some of the qualitative results can be shown to be robust to the relaxation of the standard PT assumption. Lastly, cai2016's cai2016 analysis is chosen to demonstrate the multiple treatment period extension as well as the robust bounds or the PO-GDID estimands.
kresch2020 analyzed the effect of the legal reform in Brazil that clarified the relationship between municipal (local) and state governments in the water and sanitation sector. The reform was designed to eliminate the takeover threat by state companies toward municipal companies. Accordingly, the author tried to examine if the risk before the reform caused sub-optimal investment by the municipal providers by investigating if the reform led to increased investment in self-run municipal systems. The original estimation equation is the following standard two-way fixed effects (TWFE) specification with covariates $X$ entering the model linearly:
where $Y_{mt}$ is the investment level of municipality $m$ in year $t$, and the data includes 12 years from 2001 to 2012. The reform $Reform_{mt}$ is equal to 1 for self-run municipalities after the legislation was proposed ($t > 2005$),\footnote{DentehKedagni2022 pointed out that using the proposed bill date instead of its passage date could introduce misclassification in the DID framework. Investigating the consequences of misclassification in our framework is beyond the scope of the current paper.} and there are 5 pre-treatment periods. kresch2020 considered 7 outcome variables of $Y_{mt}$, and we follow him by considering the same outcomes as described in Table (ref).
The covariates $X_{mt}$ includes municipality $m$’s population, gross domestic product (GDP) and taxes (including national and state shares), water-intensive industry variables (e.g., agriculture, livestock production), and annual temperature and rainfall measures. Since each of the covariates is continuous, for the sake of tractability, we define our information set $\mathcal I_0$ using only the 5 pre-treatment periods. In particular, we assume a slightly different version of Assumption (ref) as follows:\footnote{Note that this is different from the simple bias set stability assumption (Assumption (ref)), and the results from this assumptions are also presented below.}
for all $x_1 \in \mathcal X_1$, where $\mathcal I_0 \equiv \{2001, \cdots, 2005 \}$.
For the doubly robust estimand $\tau^{DR}$, we primarily consider a logit model for the true propensity score $P(X_1)\equiv \mathbb E[D \vert X_1]$ and a linear model for the conditional outcome mean $\mu_0(X_1)\equiv \mathbb E[Y_1 \vert D=0, X_1]$. We also present results from a quadratic specification for the conditional outcome mean.
Figure (ref) shows the scatter plot for the unconditional selection bias in the pre-treatment periods for the 7 outcome variables. Recall that for each outcome variable, we consider the convex hull of the pre-treatment periods selection biases to be stable before and after the reform proposal in order to (partially) identify the ATT. For the last outcome variable (other investments) in particular, we demonstrate the projection-based identification method in the end of this subsection as there seems to be a trend in the selection biases in the pre-treatment periods.
Table (ref) summarizes the results for the robust GDID bounds where each row represents the results of each outcome variable. The first two columns show point estimates of the GDID bounds, and the third and fourth columns are the confidence bounds obtained using a bootstrap method. The last four columns contain the $\tau^{DR}$ estimates, the constructed selection bias sets from which the GDID bounds are estimated, and the results from kresch2020.
On the other hand , Table (ref) shows a slightly different results (though qualitatively the same as those in Table (ref)) from the quadratic specification of the conditional outcome mean $\mathbb E[Y_1 \vert D=0, X_1]$. Note that the slight difference in the results is driven by the different estimates for $\tau^{DR}$ in the fifth column, since the constructed selection bias sets remain the same (the last two columns).
The findings from Tables (ref)) and (ref) suggest that the increase in investment after the reform bill was introduced is less significant than what the results in kresch2020 suggest. Our point estimate bounds show that the magnitude of the increase in investment is much smaller for all seven outcomes. The confidence bounds for all outcomes contain 0, suggesting that there is no significant change in investment after the reform. One could argue that our confidence bounds are too conservative and this may be driving the results. This does not seem to be the main reason since the point estimate bounds which are usually tighter than any confidence bounds lead to a similar conclusion. Note that in contrast to what the theory suggests kresch2020's kresch2020 point estimates lie outside our point estimate bounds. As pointed out by drdid2020 in their Remark 1, the TWFE specification ((ref)) considered in kresch2020 imposes some additional restrictions on the data generating process when assuming conditional PT. More precisely, the treatment effect is homogeneous in $X_{mt}$, and it rules out X-specific trends in both treatment and control groups $(\mathbb E[Y_1-Y_0 \vert X, D=d]=\mathbb E[Y_1-Y_0 \vert D=d])$. When these restrictions do not hold (which is likely the case here), the parameter $\delta$ will not identify the ATT and may not have any causal interpretation. This might be the reason why the point estimates in kresch2020 do not lie within our point estimate bounds.
We also compute our bounds without the use of covariates. Table (ref) displays the GDID bounds estimates under Assumption (ref) to demonstrate how the results change when ignoring the covariates. In this estimation, $\theta_{OLS}$ estimate is used in lieu of $\tau^{DR}$. We notice that total investment as well as self-financed investment and investment from loans and debt have significantly increased, while the other types of investment remained statistically stable after the reform.
We now present the PO-GDID estimands considering the three loss functions discussed in Section (ref): $L1$, $L2$, and $L\infty$. We weight each element in the information set by the ratio of the number of observations in it to the total number of observations in the information set. Since each element in the information set has the same number of observations (balanced panel), the probability weight for each of the five elements in the information set is equal to $1/5$. However, the distribution of the selection bias $SB_0(I_0)$ is not uniform over the interval $\left[ \inf_{\iota_0\in \mathcal I_0} SB_0(\iota_0), \sup_{\iota_0 \in \mathcal I_0} SB_0(\iota_0)\right]$, and the PO-GDID estimates are different across the three loss functions.
Table (ref) shows the estimation results for the three PO-GDID estimands. Each set of three columns contains the point estimates and 95% confidence intervals for each of the loss functions. The results are consistent with the findings above. There is not a statistically significant increase in any of the types of investment considered.
Finally, we demonstrate another type of GDID estimand results where the selection bias $SB_1$ is forcast from a simple regression of pre-treatment periods selection biases on time. More precisely, we first regress the five available $SB_0(\iota_0), \iota \in \mathcal{I}_0$ on the time variable $t$ as follows
Visually, the regression line obtained from the last outcome variable “Other Investments” is illustrated in Figure (ref), and we predict $SB_1(t)$ for the middle point of the post-treatment periods $t=2009$. Hence, using the estimate for $\tau^{DR}$ and $\widehat{SB_1}$, we obtain the GDID estimate. For instance, using “Other Investments,” we have $\widehat{\theta_{GDID}} = \widehat{\tau^{DR}} - \widehat{SB_1} = 495 - ( - 138) = 633$. The results for all other types of investment are summarized in Table (ref).
We can still find that the reform effect was not statistically significant for any type of investment at 5%, but for “Other Investments,” it was statistically significant at 10% as the lower bound of the 90% confidence interval is greater than 0.
The authors examine the pass-through of a tax of two cents per ounce on sugar-sweetened beverages (SSB tax) enacted in Boulder, Colorado, using the standard DID framework. They considered both store and restaurant prices and collected two different datasets for each of them: hand-collected data and Nielsen retail scanner data for the store prices, and hand-collected data and web-scrapped (OrderUp.com) data for the restaurant prices. Hence, this exercise could have been the best example for us to explore the information set consisting of the multiple datasets, but we focus on utilizing multiple pre-treatment periods of the hand-collected datasets in this subsection due to the data limitation.\footnote{Nielsen retail scanner data are proprietary, and the unit price information is not available in the web-scrapped (OrderUp.com) data.}
Each dataset is bimonthly-collected and has four periods April, June, August, and October, where the tax was imposed on July 1st of the same year. Thus, our information set has two elements April and June. Moreover, we can implement the event-study type DID analysis (static heterogeneous treatment effects in multiple treatment periods model as in Equation ((ref))) to capture the non-parametric evolution of the treatment effects over the post-treatment periods. For the first dataset of the store prices, we consider three different prices ($post\_tax$, $reg\_tax$, and $untax$) as our outcome variables, and $fount$ is selected from the hand-collected restaurant dataset as another outcome variable of interest. In particular, $post\_tax$ uses post prices on the shelves, $reg\_tax$ uses prices at the register,\footnote{cawley2021SSB found that not all retailers included the tax in the posted (or shelf) prices; i.e., some retailers added the tax at the register making it less salient.} and $untax$ uses prices of products irrelevant to SSB tax (e.g., diet soda, products in which milk is the primary ingredient, alcoholic mixers, or coffee drinks) for the blind test. On the other hand, $fount$ represents restaurant fountain drink prices.
The control community is Fort Collins, Colorado, which is geographically close to Boulder and similar in demographic characteristics as well. Hence, the standard PT assumption states that the average equilibrium beverage price differences between Boulder and Fort Collins in the post-treatment periods would have been the same as the average equilibrium price differences in the pre-treatment periods if there had not been the SSB tax in Boulder. On the other hand, our GDID model assumes that the average equilibrium price differences without the tax in the post-treatment periods would have lain between the average equilibrium price differences in April and June between the two cities. Our approach can be seen as a robustness check of the findings in cawley2021SSB.
Before we proceed to the estimation results, we show the scatter plot of the selection biases in the pre-treatment periods. Figure (ref) shows the selection biases in April and June for each of the outcome variables $post\_tax$, $reg\_tax$, $untax$, and $fount$. Accordingly, we construct a set that contains both of the selection biases in April and June for each panel or the outcome variable and assume that the set will be stable in the following post-treatment periods and contain the unobserved post-period selection bias.
Tables (ref) shows the standard DID results and our GDID bounds results. The first column presents the standard DID estimate for the ATT, and the second and third columns are corresponding 95% confidence intervals. The fourth and fifth columns show lower and upper bounds of our identified set for the ATT as presented in Proposition (ref), and the corresponding 95% confidence intervals are given in the sixth and seventh columns. Note that we are still able to reject the null hypothesis that the effect on the post prices is not different from zero under a significance level of 5% from our GDID model for $post\_tax$, $reg\_tax$, and $fount$, implying that the same qualitative conclusion can be drawn from the GDID model where we do not have to maintain the standard PT assumption. However, our results suggest that a pass-through rate higher than 100% cannot be rejected from the GDID model for $post\_tax$ whereas the standard DID estimates rule out that case; the market could be imperfectly competitive.
Figure (ref) shows the multiple treatment periods GDID bounds estimates over the post-treatment periods. The red and dark blue dashed lines are the upper and lower bound of the ATT over the periods 1 and 2, and their 95% confidence regions are depicted as gray areas with dotted lines. Although we have only two post-treatment periods, we observe the following patterns. First, the pass-through rates of SSB tax on store prices seem relatively stable over time compared to the restaurant fountain drink prices. Second, given the increasing pass-through rates on restaurant drinks, especially with the 100% pass-through rate within the bound estimates in Oct (the second post-treatment period), it would be interesting to examine further whether or not there is any excessive market power exercised through the restaurant drink prices in later periods. Finally, the figure for untaxed product prices shows that the impact of SSB tax seems to be transmitted to the other drinks in Boulder city over time, but it is not statistically significant.
Table (ref) summarizes the estimation results for the three PO-GDID estimands where each set of three columns contains point estimates and 95% confidence intervals for each of the loss functions. Here, we obtain more or less the same results as the standard DID estimates, but it is important to point out that those PO-GDID estimands are derived under stronger assumptions than the robust bounds.
Lastly, Table (ref) shows the results from applying the linear projection method that we discussed in the previous example kresch2020. We can see that every positive effect that we observed has disappeared because of the increasing trends shown in Figure (ref). Hence, even the GDID bounds results (Table (ref)) are to be taken with cautiousness if we cannot rule out the existence of trends in the selection biases.
cai2016 investigates the impact of insurance provision on tobacco production using a household-level panel dataset provided by the Rural Credit Cooperative (RCC), the main rural bank in China. The regression equation used in cai2016 is as follows:
where $i, r, t$ are household, region, and year indices, respectively, and $Y$ is the outcome variable ($area\_tob$: area of tobacco production measured in mu,\footnote{1 mu corresponds to 1/15 ha.} $tobshare$: share of tobacco production in total area of agricultural production). The covariates $X$ linearly enters the equation to be controlled for and consist of the household size, education level, and age of the household head. Note that as is common in the applied research literature, the author interprets $\alpha_3$ in Equation ((ref)) as the ATT under the standard parallel trend assumption. As we previously discussed, this model specification can be too restrictive, especially when the treatment effect is heterogeneous in the covariates $X$.
Figure (ref) and (ref) show the selection biases in the pre-treatment periods for each of the outcome variables $area\_tob$, and $tobshare$. We do not see any clear pattern for the pre-treatment periods selection biases.
Tables (ref) and (ref) show the GDID bounds estimation results where each table uses either the linear or quadratic specifications for the outcome regression model $\mu_0(X_1)$. The first four columns represent the estimated GDID bounds and their 95% confidence intervals, and the following columns show the doubly-robust estimate, the standard OLS estimate, the constructed selection bias set $\Delta_{SB_0}$, and the original point estimates from cai2016. Note that the results are not significantly different across the specifications, and we still conclude that the effect of the insurance is positive on both the area and share of tobacco at 5% significance level. Different from kresch2020, on the other hand, $\tau^{DR}$ and $\theta_{OLS}$ are close to each other, and most of the estimated GDID bounds contain the original estimates from cai2016 except for $tobshare$ under the quadratic specification. These findings suggest that the specification in Equation ((ref)) is supported by the data.
Figure (ref) shows the event-study type DID analysis (static treatment effects in multiple treatment periods model as in Equation (ref)) to capture the non-parametric evolution of the treatment effects over the post-treatment periods. The red and dark blue dashed lines are respectively the upper and lower bounds of the time-specific treatment effects, and their 95% confidence regions are depicted as gray areas with dotted lines. From this analysis, we observe that the initial impact of the insurance provision on the tobacco production area is relatively small that the null hypothesis cannot be rejected at 10% significance levels. But, the effect becomes substantially significant over time.
Table (ref) shows the GDID estimates obtained from the three types of the loss functions. As can be seen, the results are not significantly different across the three loss functions and appear to be qualitatively the same.
Finally, Table (ref) summarizes the GDID estimation results from the linear projection. Note that the observed trend of selection bias in pre-treatment periods (Figure (ref)) has caused higher predicted selection bias in the post-treatment period ($\widehat{SB_1} = -0.062$), and the effect on $tobshare$ is now no longer statistically not significant at 5% level.
Following CallawaySantAnna2021, we applied our framework to investigate the impact of minimum wage increases on teen employment. Specifically, we used the Quarterly Workforce Indicators (QWI) data used in dube2016 to collect the first quarter teen employment as our outcome variable.
CallawaySantAnna2021 considered 7 years of periods between 2001 and 2007 where the federal minimum wage did not change over time, and 3 different control groups of $g=2004, 2006,$ and $2007$ with states that raised their minimum wage in or right before the beginning of years 2004, 2006, and 2007, respectively. The specific timing of the raise can be found in CallawaySantAnna2021, and it should be noted that there is some heterogeneity in the size of the minimum wage increase within each group. The control group consists of states that did not raise their minimum wage during this period, and the complete classification can be found in Table (ref).
For illustration purposes, we implement our bounding approach under Assumption (ref), where we defined the information set using the pre-treatment periods before the first treatment in 2004 (i.e., $\mathcal I_0 = { 2001, 2002, 2003 }$). Hence, by estimating (ref) three times and taking the convex hull of each estimate / confidence interval, we were able to estimate the bounds in Proposition (ref) with the corresponding confidence intervals. The results are summarized in Figure (ref) for each treatment group.
The vertical line in each panel of Figure (ref) represents the treatment timing, and the black dots shows upper/lower bound estimates of $ATT_t[(0, \ldots,0, d_g = 0, 0, \ldots, 0)\rightarrow (0, \ldots,0, d_g = 1, 1, \ldots, 1)]$ for each $g=2004, 2006, 2007$ and $t=2004, \cdots, 2007$ as well as the corresponding 95% confidence intervals. We confirm the similar and statistically significant treatment effect trends for $g=2004$ as those found in CallawaySantAnna2021. However, our estimates are not able to reject the null hypotheses of zero treatment effects after the treatment for $g=2006$ or $g=2007$ due to more dispersed and unstable selection biases in the pre-treatment periods, resulting in larger $\Delta_{SB}$. Note however that we do not include any covariates in this illustration.
In this paper, we propose a new DID method that is robust to violations of parallel trends that can be captured in the pre-treatment periods. Under a weaker assumption than the standard (conditional) parallel trends assumption, we derive novel bounds for the ATT, which we call the robust generalized DID bounds. These bounds always cover the standard DID estimand. If the PT assumption holds in the pre-treatment periods, our robust generalized DID bounds collapse to a point, the standard DID estimand. To construct the bounds, we define an information set in the baseline period where no individual was treated yet. This information set helps define the set of all pre-treatment periods selection biases. We therefore assume that the post-treatment period selection bias lies within the convex hull of all pre-treatment periods selection biases. We provide a sufficient condition for this assumption. We also show how baseline covariates can help in the identification strategy.
As the information set grows, our bounds become wider and may become less relevant for the policymaker. We therefore discuss different ways to select the post-treatment period selection bias optimally by minimizing a loss function chosen by the policymaker. Doing so will yield a point estimand that may not necessarily have a clear causal interpretation but could be relevant for the policymaker's decision making process. We call this parameter a policy-oriented generalized DID.
We show how our method can be extended to the multiple treatment periods DID designs and the synthetic control method. We illustrate our proposed method through some numerical and empirical examples. In the multiple treatment periods DID framework, our approach partially identifies various causal parameters that can help reveal some dynamic effects of the treatment. In this setup, we propose a two-way fixed effects regression inference method. Currently, the information set is static in our proposed approach as it does not change over time. Making the information set dynamic in order to allow past outcomes to influence current and future outcomes could be an interesting area for future research. For example, investment which is the outcome variable in one of our applications follows a dynamic process.