EconBase
← Back to paper

Robust Difference-in-differences Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

128,230 characters · 35 sections · 137 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Robust difference-in-differences models

\nonstopmode

\address{Rochester Institute of Technology and UNC-Chapel Hill}

abstractThe difference-in-differences (DID) method identifies the average treatment effects on the treated (ATT) under mainly the so-called parallel trends (PT) assumption. The most common and widely used approach to justify the PT assumption is the pre-treatment period examination. If a null hypothesis of the same trend in the outcome means for both treatment and control groups in the pre-treatment periods is rejected, researchers believe less in PT and the DID results. This paper develops a robust generalized DID method that utilizes all the information available not only from the pre-treatment periods but also from multiple data sources. Our approach interprets PT in a different way using a notion of selection bias, which enables us to generalize the standard DID estimand by defining an information set that may contain multiple pre-treatment periods or other baseline covariates. Our main assumption states that the selection bias in the post-treatment period lies within the convex hull of all selection biases in the pre-treatment periods. We provide a sufficient condition for this assumption to hold. Based on the baseline information set we construct, we provide an identified set for the ATT that always contains the true ATT under our identifying assumption, and also the standard DID estimand. We extend our proposed approach to multiple treatment periods DID settings. We propose a flexible and easy way to implement the method. Finally, we illustrate our methodology through some numerical and empirical examples.

{ Keywords: Differences-in-differences, baseline information, selection bias, robust bounds, ATT.\\ JEL subject classification: C14, C31, C33, C35.}

Introduction

The difference-in-differences (DID) technique is one of the most popular methods in the social sciences when an experimental research design cannot be used. The DID method requires at least observational data consisting of two different groups (a treatment group and a control group) and two time periods of pre-treatment and post-treatment. Under some assumptions, the method identifies the average treatment effects on the treated (ATT) as the DID estimand. The key identifying assumption of interest is the so-called parallel trends (PT) assumption. This assumption states that the untreated potential outcome variable for the treatment group would have followed on average the same trend as that for the control group had they not been treated. However, it is difficult to empirically verify the PT assumption because it restricts a hypothetical quantity that is not identifiable. Accordingly, convincing readers to approve the PT assumption has been the most vital and controversial part of the DID literature. For instance, kearney2015mtv's (kearney2015mtv) ambitious identification strategy on discovering the effects of an MTV reality show on teen childbearing provided insightful findings that would not have been discovered without the study, but there have been heated debates on the validity of its main PT assumption as well jaeger2018, kahnlang2019. The most common and widely understood approach for justifying the PT assumption is the pre-treatment period examination. If a null hypothesis of the same trend in the untreated potential outcome mean for both treatment and control groups in the pre-treatment periods cannot be rejected, then the researcher will believe that the PT assumption is likely to hold for the post-treatment period as well. Still, rigorously speaking, the evidence of pre-treatment PT is different from the PT assumption in the post-treatment period that is of interest, and thus additional arguments should be established for the PT assumption separately. CallawaySantAnna2022 discussed this nuance well in their vignette on pre-testing in a DID framework:

“Importantly, this is just a pre-test; it is different from an actual test. Whether or not the parallel trends assumption holds in pre-treatment periods does not actually tell you if it holds in the current period (and this is when you need it to hold!). It is certainly possible for the identifying assumptions to hold in previous periods but not hold in current periods; it is also possible for identifying assumptions to be violated in previous periods but for them to hold in current periods. That being said, we view the pre-test as a piece of evidence on the credibility of the DiD design in a particular application.”

See also Freyaldenhoven_al2019, kahnlang2019, Roth2022, etc. Hence, this paper develops a generalized DID framework that can utilize all the information available not only from the pre-treatment periods but also from multiple baseline covariates or data sources. In doing so, we develop a DID method that is robust to violations of PT that can be captured in the pre-treatment periods.

First, our approach is unique in that we follow Heckman_al1998 to interpret the PT assumption in a different way using a notion of selection bias (also known as confounding bias in statistics),\footnote{Some recent papers also use similar interpretations of the PT assumption Henderson2023, Sofer2016, Parkal2023.} which enables us to generalize the standard DID estimand by defining an information set that can represent a set of multiple pre-treatment periods or other baseline covariates. We define selection bias as the mean-difference of the untreated potential outcome between the treatment and control groups. We introduce the concept of generalized difference-in-differences (GDID) estimand defined as the difference between the ordinary least squares estimand in the post-treatment period and a selection bias in that period, which is a correspondence of the available information set and the baseline period selection bias. Under the assumption that the treatment has no anticipatory effects, we identify the baseline period selection bias as the difference-in-means of that period's observed outcome between the treatment and control groups.

Second, we consider assumptions under which the above correspondence is known. Our main assumption states that the selection bias in the post-treatment period lies within the convex hull of all selection biases in the pre-treatment periods. We provide a sufficient condition for this assumption to hold. For example, we discuss and illustrate that this assumption may be plausible in economic settings where Ashenfelter1978's (Ashenfelter1978) dip is present. It is well documented that in such contexts, the PT assumption is usually not plausible (e.g., see AshenfelterCard1985, HeckmanSmith1999, Heckman_etal1999). Based on the baseline information set we construct, we provide an identified set for the ATT that always contains the true ATT under our identifying assumption, and also the standard DID estimand, given that we are using a weaker assumption than the PT assumption. If in fact PT holds in the pre-treatment periods, our bounds naturally collapse to the standard DID estimand. Importantly, we show how the baseline covariates can help define the correspondence and therefore help partially identify the ATT of interest. Unlike the standard DID framework where covariates are required to be time-invariant, our method allows for exogenous time-varying covariates. To the best of our knowledge, only few papers in the DID literature allow for time-varying covariates Caetano_etal2022, Shahnal2022. Our paper contributes to this literature. We provide multiple illustrative examples where the standard DID estimand does not identify the ATT while our bounds cover it.

Third, we discuss alternative ways of defining the correspondence. For example, when the pre-treatment periods selection biases show some clear pattern in terms of trends, we discuss how the researcher can model such a pattern and use that model to forecast the post-treatment period selection bias. In the same direction, we propose a class of criteria on the selection biases from the perspective of a policymaker that can achieve a point identification of ATT. We call this point estimand a policy-oriented GDID, as it may not necessarily have a causal interpretation.

Fourth, we propose an implementation procedure for our bounds. We provide a doubly-robust estimand in the presence of covariates in the post-treatment period. Our proposed confidence bounds are valid in the sense that they will cover the true identified set with a pre-specified probability. However, they may be too conservative. We believe that this inference procedure can be improved and leave this improvement for future research.

Fifth, we show how our framework extends to the DID with multiple treatment periods, and the synthetic control (SC) settings. On the one hand, we derive bounds on each post-treatment period ATT. This approach can help reveal the heterogeneity in the treatment effects over time. As before, if PT holds for the treatment status in each treatment period, our bounds collapse to a DID estimand. We can therefore identify the ATT in each treatment period. We also extend the method to the identification of more causal parameters, including those considered in CallawaySantAnna2021. The causal parameters we consider can also help reveal the dynamic effect of the treatment. On the other hand, instead of finding the optimal weights for elements in the donor pool to create a counterfactual synthetic control for the treated unit, we propose bounds on the ATT by considering each donor as a potential control unit. Here again, we use the convex hull of all pre-treatment periods selection biases as the identified set for the selection bias in the post-treatment period.

Finally, we illustrate the empirical relevance of our methodology by revisiting kresch2020, cawley2021SSB, and cai2016. We apply our method to investigate the causal effect of a 2007 reform in Brazil, that gave municipalities the ultimate authority to provide some services, on different types of investment. We find that the effect was less/not significant as initially found by the author. The main issue was that the pre-treatment periods selection biases were not stable over time. This makes the PT assumption less reliable in this application. On the other hand, cawley2021SSB examine the pass-through of a tax of two cents per ounce on sugar-sweetened beverages (SSB tax) enacted in Boulder, Colorado, using the standard DID framework. Both the DID method and our GDID bounds lead to the conclusion that the policy effect on the post prices is statistically significant. Furthermore, we also revisit cai2016 who investigates the impact of insurance provision on tobacco production using a household-level panel dataset provided by the Rural Credit Cooperative (RCC), the main rural bank in China. Using the DID approach and the robust GDID bounds, we conclude as the author that the effect of the insurance is positive on both the area and share of tobacco.

Our paper is closely related to two papers in the literature: ManskiPepper2018, and RambachanRoth2020. While these papers mainly focus on a trends-based / space-based relaxation of the PT assumption, our paper relies on a selection-based relaxation approach.\footnote{Our selection-based relaxation can use information over time (e.g., when the baseline information set is the set of pre-treatment periods) or across space (e.g., when the baseline information is the set of baseline covariates, which can include geographic region).} On the one hand, ManskiPepper2018 introduce bounded-variation assumptions that relax the PT assumption. Instead of requiring that the untreated potential outcomes for the treatment and control groups follow on average the same trends between the baseline and treatment periods, the authors assume that the absolute difference in trends is bounded by a known sensitivity parameter. When this parameter is equal to zero, their assumption reduces to the PT assumption. ManskiPepper2018's (ManskiPepper2018) approach is robust to violations of the PT assumption when the sensitivity parameter is big enough. Our approach is robust to violations of PT that can be captured in the pre-treatment periods but is not necessarily robust to post-treatment violations that cannot be captured in the pre-treatment periods. For example, when PT holds in the pre-treatment periods while it is actually violated in the post-treatment period, our set identification method coincides with the standard DID approach, which will not identify the ATT as the needed PT assumption does not hold. However, the choice of the sensitivity parameter in the ManskiPepper2018 bounding approach remains unclear. On the other hand, RambachanRoth2020 generalizes ManskiPepper2018's (ManskiPepper2018) bounding method by considering a large class of restrictions that impose that the post-treatment violations of parallel trends cannot be “too different”\footnote{In the terminology of RambachanRoth2020.} from the pre-trends. Our bounding strategy falls into this class of restrictions that the authors consider and can therefore be viewed as a special case of their approach. However, we tackle the problem with a different perspective and our identifying assumptions have not been considered in RambachanRoth2020. Like in ManskiPepper2018, the specific restrictions they consider require the knowledge of a sensitivity parameter, whose choice still remains unclear in their case. Moreover, our approach does not require an explicit choice of a sensitivity parameter. The sensitivity parameter is implicitly embedded in the baseline information set.

Our paper also contributes to the growing literature on sensitivity analysis in the DID framework. Freyaldenhoven_al2019 propose a method that estimates a policy effect using a two-stage least squares approach in a linear panel event-study design where unobserved confounds may be related both to the outcome and the the policy variable of interest. Their identification strategy relies on the existence of covariates related to the policy only through the confounds. Keele_al2019 develop a method of sensitivity analysis that allows researchers to quantify the amount of bias from time-varying confounders necessary to change a study’s conclusions in the DID model, relying on baseline covariates. In the same direction, Ye_al2022 propose a partial identification strategy that relaxes the PT assumption to a monotone trends assumption relying on two groups of control units whose outcomes relative to the treated units exhibit a negative correlation. Our approach does not require an existence of two control groups and our identifying assumption may still hold even their monotone trends assumption fails to hold. Similarly to our bounds, their identified set is of a union bounds form that involves the minimum and maximum operators. While our confidence bounds may be too conservative, they propose a novel bootstrap method to construct uniformly valid confidence bounds for the identified set and parameter of interest. It may be possible to implement their proposed method in our framework. We find our inference method attractive as it is easy to implement, especially when the baseline information set is discrete. Leavitt2020 develops an empirical Bayes' procedure that allows for other trend assumptions in the DID framework. On the other hand, BilinskiHatfield2020 and DetteSchumann2020 propose more reliable inference methods to detect meaningful violations of the PT assumption in the pre-treatment periods. Finally, by extending our proposed method to the multiple treatment periods setting, we contribute to the growing literature on the causal interpretation of event-study coefficients in two-way fixed effects models in the presence of staggered treatment timing and heterogeneous treatment effects (as in Borusyak_al2022; AtheyImbens2022; goodmanbacon2021ddtiming; CallawaySantAnna2021; deChaisemartinDHaultfoeuille2020; SunAbraham2021) when the PT assumption fails to hold in the pre-treatment periods. Building on Wooldridge2021's (Wooldridge2021) idea, we show how two-way fixed effects regression methods can help compute our confidence set when the baseline information set is the set of pre-treatment periods. It would be interesting to extend this developed approach to the changes-in-changes model considered in AtheyImbens2006, since the identifying assumptions may be sensitive to functional forms as is the case for the PT assumption RothSantanna2021. Some recent papers have also proposed alternative point-identification results in this setting Parkal2023, Wooldridge2022. Other papers propose point-identification and estimation results in DID settings where the standard PT assumption may be questionable, while relying on some additional assumptions Henderson2023, Richardson2023, Dukesal2022, Brown2023.

The remainder of the paper is organized as follows. Section (ref) presents the model and a preview of our approach and results, and it introduces the generalized DID concept. Section (ref) formally discusses the assumptions and the main identification results, while Section (ref) introduces the policy-oriented generalized DID concept. Section (ref) briefly discusses the implementation of the proposed bounds. Section (ref) presents two extensions of our approach. Section (ref) shows the practical relevance of the method through three empirical illustrations, and Section (ref) concludes. Proofs of the main results are relegated to the appendix.

Analytical framework and overview of the results

The baseline model

Consider the following two-period model:

eqnarray[eqnarray omitted — 136 chars of source]

where the vector $(Y_0, Y_1,D, I_0, X_0, X_1)$ represents the observed data, while the vector $(Y_1(0), Y_1(1))$ is latent. In this model, the variables $Y_0, Y_1 \in \mathcal Y$ are respectively the observed outcomes in the baseline period 0 and the follow-up period 1, while $D\in \left\{0,1\right\}$ is the observed treatment that occurred between periods 0 and 1, $Y_1(0)$ and $Y_1(1)$ are the potential outcomes that would have been observed in period 0 had the treatment $D$ been externally set to 0 and 1, respectively. The variable $Y_0(0)$ is the potential outcome that is realized in the baseline period when no individual/unit was treated. As is common in the DID literature, model ((ref)) assumes that there is no anticipatory effect of the treatment, so that $Y_0(1)=Y_0(0)$. The set $I_0 \in \mathcal I_0$ contains information on baseline data, while $X_0 \in \mathcal X_0$ and $X_1 \in \mathcal X_1$ denote the vector of covariates in periods 0 and 1, respectively. The baseline information $I_0$ could be a subset of $X_0$ but does not have to be.

In this paper, we are interested in identifying the average treatment effect on the treated (ATT) defined as $$ATT\equiv \mathbb E[Y_1(1)-Y_1(0)\vert D=1].$$ We first focus on the case without covariates.

We start by defining the standard ordinary least squares (OLS) estimand, which is the same as the difference in means estimand, as $$\theta_{OLS} \equiv \mathbb E[Y_1\vert D=1] - \mathbb E[Y_1 \vert D=0].$$ We can rewrite this OLS estimand as the ATT plus a bias term. Indeed, we have

eqnarray[eqnarray omitted — 243 chars of source]

where $SB_t\equiv \mathbb E[Y_t(0) \vert D=1]-\mathbb E[Y_t(0) \vert D=0]$.

Equation ((ref)) shows that the standard OLS estimand in period 1 can be decomposed as equal to the ATT of interest plus a bias term that we call selection bias. Therefore, in order to identify the ATT with the help of the OLS estimand, we need to identify this selection bias. The main question we are asking at this point is how to obtain the selection bias $SB_1$. The literature provides at least two solutions to this problem. One can randomize the treatment and then get rid of the selection bias when there is full compliance, or one can rely on the PT assumption. While a successful randomized experiment yields zero selection bias ($SB_1=0$), it is often difficult and costly to implement (e.g., because of some ethical concerns, feasibility). On the other hand, the PT assumption could be too restrictive in some cases. For this reason, our approach aims at relaxing the PT assumption and provides credible bounds on the ATT instead of point-identifying this parameter.

Following Heckman_al1998, we reinterpret the PT assumption as a bias equality assumption: the selection bias in period 1 is equal to the selection bias in period 0, i.e., $SB_1=SB_0$, which is identified as the difference in the baseline outcome means between the treatment and control groups, under the no-anticipatory effects assumption. Indeed, we have

eqnarray*[eqnarray* omitted — 243 chars of source]

Why a selection-based relaxation?

The terminology parallel trends is more appropriate when the untreated potential outcome mean has linear trends in the treatment and control groups. However, when the trends are nonlinear, they may not be parallel even if the mathematical definition of the parallel trends assumption holds. Indeed, equality of the selection biases in periods 0 and 1 is sufficient for the mathematical definition of parallel trends to hold, regardless of what the untreated potential outcome mean trends are for the treatment and control groups. To illustrate this, consider a simple version of model ((ref)) where

eqnarray*[eqnarray* omitted — 204 chars of source]

and the information set $\mathcal I_0$ is the set of two pre-treatment periods $\mathcal T_0$, i.e., $\mathcal I_0=\mathcal T_0=\{-1,0\}$. In this model, the selection bias $SB_t=(1-2 \vert t \vert + 2 t^2)(\alpha_1-\alpha_0)$ where $\alpha_1=\frac{\phi(1)}{1-\Phi(1)}\approx 1.53$ and $\alpha_0=-\frac{\phi(1)}{\Phi(1)} \approx -0.29$. Therefore, we have $SB_0= \alpha_1-\alpha_0 =SB_{-1}=SB_1$. So, the standard parallel trends assumption holds. Yet, the trends of $\mathbb E[Y_t(0)\vert D=1]$ and $\mathbb E[Y_t(0)\vert D=0]$ are not parallel. We refer to these trends as spurious parallel trends. Figure (ref) displays those trends for $\theta=2$. Because of the existence of spurious trends, when a researcher is doubtful about the validity of the PT assumption, a selection-based relaxation approach could be more appropriate than a trends-based approach in some circumstances. In this sense, we view our approach as a complement to the existing trends-based relaxation approaches. For instance, suppose that the time unit on the $x$-axis is the year, and a semestral dataset is available. If a researcher is interested in identifying the treatment effect at period $t=1/2$ (first semester), the standard DID estimand will fail to identify the causal effect, as $SB_{\frac{1}{2}} \neq SB_0$. Our proposed approach will be robust to this kind of spurious parallel trends as long as our information set includes the period $t=-1/2$.

figure[figure omitted — 162 chars of source]

Before we present the formal results, we heuristically show the intuition behind our main identification strategy.

Overview of the main results

Suppose the information $I_0$ contains two pre-treatment periods such that $\mathcal I_0=\{-1,0\}$. In general, when $SB_{-1} \neq SB_0$, it is difficult to believe that $SB_0=SB_1$. Note that none of the conditions implies the other. Yet, researchers often rely on this pre-test to check the plausibility of PT. Our approach is to assume that $SB_1$ lies within the convex hull of $\{SB_{-1},SB_0\}$, that is, $SB_1 \in [\min\{SB_{-1}, SB_0\},\max\{SB_{-1}, SB_0\}]$. Under our assumption, we obtain the following bounds on the ATT: $$ATT \in \left[\theta_{OLS}-\max\{SB_{-1}, SB_0\}, \theta_{OLS}-\min\{SB_{-1}, SB_0\}\right].$$ Hence, our bounding approach is robust to violations of parallel trends that can be captured in the pre-treatment periods. However, this does not ensure that our identifying assumption is valid. One would have to justify why the selection bias in period 1 would lie within the convex hull of the selection biases in pre-treatment periods. As can be seen, the standard DID estimand ($\theta_{OLS}-SB_0$) lies within our bounds.

Suppose now that the baseline period 0 is the only pre-treatment period for which a data is available, and we observe a baseline covariate $X_0$. For simplicity, assume $I_0=X_0$. Unlike the standard approach which requires $X_0$ to be equal to $X_1$ i.e., $X_0=X_1=X$ abadie2005, we allow $X_0$ to be different from $X_1$ in our framework. Yet, to illustrate our contribution over the existing approaches, we consider the simple case where $X_0=X_1=X \in \{x_0,x_1\}$. Define $SB_t(x)\equiv \mathbb E[Y_t(0)\vert D=1, X=x]-\mathbb E[Y_t(0)\vert D=0, X=x]$. Existing methods assume $SB_0(x)=SB_1(x)$, while ours assumes $SB_1(x)\in [\min\{SB_0(x_0), SB_0(x_1)\},\max\{SB_0(x_0), SB_0(x_1)\}]$. As we can see, we allow for $SB_0(x) = SB_1(x)$ for some $x$, $SB_0(x)\neq SB_1(x)$ for all $x$, $SB_0(x) = SB_1(x')$ for some $(x,x')$ or $SB_0(x)\neq SB_1(x')$ for all $(x,x')$.

Given our above assumption, we partially identify $ATT(x)\equiv \mathbb E[Y_1(1)-Y_1(0)\vert D=1, X=x]$ as follows: $$ATT(x) \in \left[\theta_{OLS}(x)-\max\{SB_0(x_0), SB_0(x_1)\}, \theta_{OLS}(x)-\min\{SB_0(x_0), SB_0(x_1)\}\right],$$ where $\theta_{OLS}(x) \equiv \mathbb E[Y_1\vert D=1, X=x] - \mathbb E[Y_1 \vert D=0, X=x]$. Hence, we can integrate the bounds on $ATT(x)$ over the conditional distribution of $X$ given $D=1$ to obtain bounds on $ATT$.

The following example shows a data generating process where our identifying assumption holds in a situation where we only have two periods and a time-invariant covariate is available.

exampleConsider a the following model where \begin{eqnarray*} \left\{ \begin{array}{lcl} Y_t&=&(1+ 0.5^tX)U+\theta X D*t\mathbbm{1}\{t\geq 0\} \\ \\ D&=&\mathbbm{1}\{U\geq 1\}\\ \\ U &\sim& N(0,1), X \sim \mathcal U_{\left[0,1\right]}, and X\ \rotatebox[origin=c]{90}{$\models$}\ U \end{array} \right. \end{eqnarray*} where $\mathcal I_0=\mathcal X=[0,1]$. We have $SB_0(x)=(1+x)(\alpha_1-\alpha_0)$, and $SB_1(x)=(1+0.5x)(\alpha_1-\alpha_0)$ where $\alpha_1=\frac{\phi(1)}{1-\Phi(1)}\approx 1.53$ and $\alpha_0=-\frac{\phi(1)}{\Phi(1)} \approx -0.29$. We have $SB_0(x) \in [\alpha_1-\alpha_0, 2(\alpha_1-\alpha_0)]$ and $ SB_1(x) \in \left[\alpha_1-\alpha_0, 1.5(\alpha_1-\alpha_0)\right] \subseteq [\alpha_1-\alpha_0, 2(\alpha_1-\alpha_0)]\equiv \Delta_{SB_{0X}}$. So, the standard parallel trends assumption does not hold as $SB_0(x)\neq SB_1(x)$. However, the selection bias $SB_1(x) $ in period 1 belongs to the convex hull of all selection biases in period 0, i.e., $SB_1(x) \in \Delta_{SB_{0X}}$. Hence, our identifying assumption holds. We have $\theta_{OLS}(x)=(1+ 0.5 x)(\alpha_1-\alpha_0)+\theta x$, and our new bounds $ \Theta_I $ are obtained as $ATT(x) \in [\theta x - (1-0.5 x)(\alpha_1-\alpha_0),\theta x + 0.5 x(\alpha_1-\alpha_0)]$. The actual conditional ATT function is $ATT(x)=\theta x$, but the standard conditional DID estimand is $\theta_{DID}(x)=\theta x - 0.5 x (\alpha_1 - \alpha_0)$. Figure (ref) shows the bounds $ \Theta_I $, the true conditional ATT, and the conditional standard DID for different values of $x$ when $\theta=2$. The standard conditional DID is biased except for $x = 0$, whereas our bounds contain the true conditional ATT. \begin{figure}[h] \caption{Illustration of $\Theta_I$ for $\theta = 2$ and $x \in [0, 1]$} \end{figure}

Interpretation of our assumption. Let $X$ be the variable gender. The standard assumption $SB_0(x)=SB_1(x)$ states that the selection bias for females in period 0 is equal to selection bias for females in period 1, and similarly for males. Our assumption states that the selection bias for females (resp. males) in period 1 lies between those for males and females in period 0. Our assumption allows the selection bias for females in period 1 to be equal to that of males in period 0, and vice versa. Furthermore, we allow for the possibility that the selection bias for females (resp. males) in period 1 be different from those for males and females in period 0.

As we explain above, the standard DID estimand is defined as the difference between the OLS estimand in period 1 and the selection bias in period 0: $\theta_{DID} \equiv \theta_{OLS}-SB_0$. We introduce a generalized version of this estimand.

definitionGiven the baseline information set $\mathcal I_0$ and the selection bias $SB_0$, we define the generalized difference-in-differences (GDID) estimand as \begin{eqnarray} \theta_{GDID}\equiv \theta_{OLS}-SB_1(SB_0,\mathcal I_0), \end{eqnarray} where $SB_1(SB_0,\mathcal I_0)$ is a function/correspondence of the selection bias $SB_0$ in period 0 and the information set $\mathcal I_0$.

In the above definition, if $SB_1(SB_0,\mathcal I_0)=SB_0$, the generalized DID estimand is the same as the standard DID estimand. Note however that $SB_1(SB_0,\mathcal I_0)$ is allowed to be a set of values. In such a case, the generalized DID estimand will be a set instead of a single value.

In the next section, we formally discuss our assumptions and the main results.

Assumptions and main identification results

In this section, we state our identifying assumptions and present our main results.

Identification without covariates

Let us first consider the simple case with no covariates in the model. We now state our main assumption.

assumption[Bias set stability] \begin{eqnarray*} SB_1 \in \left[ \inf_{\iota_0\in \mathcal I_0} SB_0(\iota_0), \sup_{\iota_0 \in \mathcal I_0} SB_0(\iota_0)\right] \equiv \Delta_{SB_0}, \end{eqnarray*} where $SB_0(\iota_0)\equiv \mathbb E[Y_0\vert D=1, I_0=\iota_0] - \mathbb E[Y_0 \vert D=0, I_0=\iota_0]$ is the selection bias in the baseline period conditional on the information $\{I_0=\iota_0\}$.

Assumption (ref) is weaker than the standard “parallel/common trends” assumption. Indeed, if $\mathcal{I}_0$ is the singleton of a single baseline information $I_0=\{\iota_0\}$, then Assumption (ref) is equivalent to $SB_1= SB_0(\iota_0),$ which is equivalent to the parallel trends assumption, as discussed earlier. For example, suppose that the information set contains two pre-treatment periods such that $\mathcal I_0=\{-1,0\}$. The PT assumption $SB_1=SB_0$ implies $$SB_1 \in \left[\min\{SB_{-1},SB_0\}, \max\{SB_{-1},SB_0\}\right],$$ which is equivalent to Assumption (ref) in this example.

Instead of assuming parallel trends or bias equality, we assume that the convex hull of the set of selection biases in the pre-treatment periods is stable over time. This assumption could be violated in many situations. For example, when the information set is ordered (e.g., time, ordered covariates) and the baseline selection biases change monotonically with the elements in the information set, then the selection bias in the follow-up period will likely be outside the set of pre-treatment periods selection biases. In such a case, Assumption (ref) may not hold. We propose a solution for this context in Section (ref). Furthermore, when the potential outcome in the treatment period is linear (but not a random walk) in the baseline period (e.g., $Y_1(0)=\alpha Y_0(0) + \varepsilon$, where $\alpha \neq 1$, and $\varepsilon$ is exogenous), then neither parallel trends nor bias set stability holds.

Note that the set $\mathcal{I}_0$ could contain all pre-treatment periods, observed baseline characteristics, or information from other data sources. For example, suppose that $\mathcal I_0$ contains gender. The standard parallel trends assumption conditional on gender states that the selection bias for males in period 0 would be the same for males in period 1, and similarly for females. As discussed above, our assumption (ref) allows the selection bias for females in period 1 to be equal to that for males in period 0, and vice versa. We believe that this latter assumption is more flexible. Suppose that we collect information from multiple data sources that may not be representative of the population of interest. In this situation, $I_0$ could denote a categorical random variable for each data set. For example, for a study on the US population, a researcher can combine data from multiple states. Each of these data can be considered as a piece of information $\iota_0$.

Assumption (ref) defines a particular correspondence $SB_1(SB_0,\mathcal I_0)$ for the selection bias $SB_1$. Plugging this in the generalized DID estimand yields the following bounds on the ATT.

propositionSuppose that model ((ref)) along with Assumption (ref) holds. Then, the following bounds hold for the ATT: \begin{eqnarray*} ATT \in \left[ \theta_{OLS}- \sup_{\iota_0 \in \mathcal I_0} SB_0(\iota_0), \theta_{OLS}- \inf_{\iota_0 \in \mathcal I_0} SB_0(\iota_0) \right] \equiv \Theta_I. \end{eqnarray*} These bounds are sharp, and $\Theta_I$ is the identified set for the ATT.

The bounds in Proposition (ref) are never empty, as they always contain the standard DID estimand under the parallel trends assumption. However, they may not contain the OLS estimand in period 1, $\theta_{OLS}$, as 0 may not lie within the set $\Delta_{SB_0}$. If all pre-treatment periods selection biases are equal, i.e., $SB_0(\iota_0)=SB_0$ for all $\iota_0$, then our bounds collapse to a point, the standard DID estimand. In case the information set $\mathcal I_0$ is the set of pre-treatment periods, the above bounds are robust to violations of PT that can be captured in the pre-treatment periods. In this sense, our method can be seen as a way of salvaging (using the language of Masten_al2021) the standard DID model from violations of PT in the pre-treatment periods. However, our identification strategy does not rely on finding the falsification frontier.

An economic setting where our Assumption (ref) may hold would be the evaluation of a job training program in which the so-called Ashenfelter1978 dip occurs. As pointed out by Ashenfelter1978, individuals who participate in a job training program are usually those who have experienced a decline in employment and earnings prior to their enrollment in the program. If the decline were transitory, such individuals would normally experience a rebound in employment and earnings, even if they did not participate in the program. This phenomenon would likely make the PT assumption violated. We provide Example (ref) below that shows this pattern as depicted in Figure (ref). This Ashenfelter1978 dip has also been documented in the evaluation of incarceration on subsequent earnings and employment (e.g., see Lalonde_Cho2008, Jung2011).

The example below presents a data generating process (DGP) where the PT assumption fails, while Assumption (ref) holds. The DGP is inspired from a random-growth model with a factor structure discussed in HeckmanHotz1989. It shows how informative the bounds can be.

exampleConsider a simple version of model ((ref)) where \begin{eqnarray*} \left\{ \begin{array}{lcl} Y_t&=&(1+\vert t \vert + t^2)U+\theta D*t\mathbbm{1}\{t\geq 0\} \\ \\ D&=&\mathbbm{1}\{U\geq 1\}\\ \\ U &\sim& N(0,1) \end{array} \right. \end{eqnarray*} and $\mathcal I_0=\mathcal T_0=\{-2,-1,0\}$. In this model, $SB_t=(1+\vert t \vert + t^2)(\alpha_1-\alpha_0)$ where $\alpha_1=\frac{\phi(1)}{1-\Phi(1)}\approx 1.53$ and $\alpha_0=-\frac{\phi(1)}{\Phi(1)} \approx -0.29$. We have $SB_0= \alpha_1-\alpha_0 \neq 3(\alpha_1-\alpha_0) =SB_{-1} \neq SB_{-2}=7(\alpha_1-\alpha_0)$ and $SB_0=\alpha_1-\alpha_0 \neq 3(\alpha_1-\alpha_0)=SB_1$. So, the standard parallel trends assumption does not hold as illustrated in Figure (ref). However, the selection bias $SB_1$ in period 1 belongs to the convex hull of all selection biases in period 0, i.e., $SB_1 \in [\min\{SB_0,SB_{-1},SB_{-2}\}, \max\{SB_0,SB_{-1},SB_{-2}\}]=[\alpha_1-\alpha_0,7 (\alpha_1-\alpha_0)]$. Hence, our identifying assumption holds. \begin{figure}[h] \caption{Violations of parallel trends: Ashenfelter1978's dip ($\theta=9$)} \end{figure} We have $\theta_{OLS}=3(\alpha_1-\alpha_0)+\theta$, and $\Theta_I=[\theta-4(\alpha_1-\alpha_0),\theta+2(\alpha_1-\alpha_0)]$. The true $ATT=\theta$, and the DID estimand is $\theta_{DID}=\theta_{OLS}-SB_0=\theta+2(\alpha_1-\alpha_0)$. Thus, the DID estimand is upward biased and the bias is equal to $2(\alpha_1-\alpha_0)$. Figure (ref) shows that the bounds are generally informative about the magnitude of the ATT and can identify the sign of the ATT in some circumstances. For example, when the true ATT is equal to $-5$ or $9$, our bounds as well as the standard DID estimand identify the correct sign. On the other hand, when the true ATT is equal to $-1$, our bounds do not identify any sign, as they contain zero. But, the standard DID estimand identifies a wrong sign, it shows that the ATT is positive while it is actually negative. \begin{figure}[h] \caption{Illustration of $\Theta_I$ for $\theta \in [-10, 10]$} \end{figure}

A sufficient condition for Assumption (ref)

One question that comes to people's mind when they think about Assumption (ref) is: what are the conditions under which this assumption will hold? To this question, we provide a sufficient condition on the data generating process for the counterfactual untreated potential outcome under which this assumption holds.

assumption$\,$ \begin{enumerate}[(i)] • The untreated potential outcome satisfies: $Y_t(0)=g_t(\varepsilon) \lambda(U)+\gamma(V)+\eta_t$ where $(\varepsilon,U,V,\eta_t)$ is a random vector satisfying $(\varepsilon, \eta_t)\ \rotatebox[origin=c]{90}{$\models$}\ (U,V)$, and $g_t(.)$, $\lambda(.)$ and $\gamma(.)$ are three unknown (nontrivial) functions. • The function $g_t(.)$ is even in $t$ or there exists $t_0<0$ s.t. $\mathbb E[g_1(\varepsilon)]=\mathbb E[g_{t_0}(\varepsilon)]$; • The treatment receipt is defined as $D=h(U,V)$, where $h$ is a nontrivial function. \end{enumerate}

Assumption (ref).((ref)) postulates a factor (or interactive fixed effects) structure for the untreated potential outcome, which is commonly used in applied research. Assumption (ref).((ref)) imposes some symmetry condition on the factor function $g_t$. If the function $g_t(.)$ is even in $t$, then the selection bias is symmetric in $t$, and the set of biases before the baseline period will be identical to the set of biases after the baseline period. In such a case, our Assumption (ref) will hold. This symmetry condition is similar to the intuition behind the symmetric DID discussed in AshenfelterCard1985. This assumption can be relaxed. See Example (ref) in the appendix where $g_t=t$ is odd, but our Assumption (ref) holds. Assumption (ref).((ref)) postulates that the selection into treatment is function of the time-invariant unobservables in the model.

propositionSuppose $\mathcal I_0=\{-T_0,-T_0+1, \ldots, 0\}$, $\mathbb E[g_1(\varepsilon)]\neq \mathbb E[g_0(\varepsilon)]$, and Assumption (ref) holds. Then Assumption (ref) holds while PT fails to hold.

Our main assumption (Assumption (ref)) will generally hold in settings where there exist some common life-cycle factors that affect the untreated potential outcome. These life-cycle factors could translate into the symmetry condition or some periodicity in the potential outcome. Although we provide a sufficient condition for our main assumption, we believe that a deeper understanding of it through its connection to structural economic choice models, as discussed in Ghanem_al2022 and Marx_al2022 in the context of the PT assumption, would be an interesting direction for future research. A more general sufficient condition is provided below. The factor structure considered in Assumption (ref) is a special case of Assumption (ref) below.

assumption$\,$ \begin{enumerate}[(i)] • The untreated potential outcome satisfies: $$Y_t(0)=\varphi(t,U,\varepsilon)+\eta_t,$$ where $(U,\varepsilon, \eta_t)$ is a random vector of unobserved heterogeneity (U can be a vector); • The function $\varphi(t,u,e)$ is even in $t$ or there exists $t_0<0$: $\varphi(1,u,e)=\varphi(t_0,u,e)$ for all $(u,e)$; • The treatment receipt is defined as $D=h(U,V)$, where $h$ is a nontrivial function; • $(\varepsilon, \eta_t)\ \rotatebox[origin=c]{90}{$\models$}\ (U,V)$. \end{enumerate}

For example, Assumption (ref) holds in the following DGPs: $Y_t=\sqrt{t^2+U}+\varepsilon+\theta D*t\mathbbm{1}\{t\geq 0\}$, $D=\mathbbm{1}\{U \geq 1\}$, $U \sim N(0,\sigma^2)$, or $Y_t=\sqrt{(t+2)(t-1)+U}+\varepsilon+\theta D*t\mathbbm{1}\{t\geq 0\}$.

Identification with covariates

In this subsection, we include covariates in the analysis. We allow the baseline characteristics $X_0$ to be different from those in the follow-up period $X_1$. We denote the baseline information by $I(X_0)$ to explicitly show that it depends on the baseline covariates $X_0$. For the sake of clarity of the exposition, assume $I(X_0)=X_0$.

Define

eqnarray*[eqnarray* omitted — 229 chars of source]
assumption[Conditional bias set stability] \begin{eqnarray*} SB_1(x_1) \in \left[ \inf_{x_0\in \mathcal X_0} SB_0(x_0), \sup_{x_0 \in \mathcal X_0} SB_0(x_0)\right] \equiv \Delta_{SB_{0X}}, \end{eqnarray*} where $SB_0(x_0)\equiv \mathbb E[Y_0\vert D=1, X_0=x_0] - \mathbb E[Y_0 \vert D=0, X_0=x_0]$ is the selection bias in the baseline period conditional on the baseline information $\{X_0=x_0\}$.

The main idea behind Assumption (ref) is to use the baseline characteristics to help identify the set of possible values for the selection bias in the treatment period. The intuition is that observing different realizations of the baseline selection bias $SB_0$ can inform us about the range of possible values that the treatment period selection bias $SB_1$ can take. Assumption (ref) implies that the convex hull of all selection biases in the baseline period 0 is the same as that of all possible selection biases in period 1. Note that Assumption (ref) is weaker than the standard conditional PT in abadie2005, heckman1997, SantAnna_etal2020, etc. It allows for time-varying covariates as in Caetano_etal2022 but is different from their conditional PT assumption.

The next assumption is a common support assumption that requires that conditional on each period covariates, there exits at least a nonnegligible set of individuals in both treatment and control groups that share these characteristics. This assumption is standard when the covariates are time-invariant.

assumption[Overlap] $0< \mathbb P(D=1\vert X_t) < 1$ a.s. for $t=0,1.$

The identification results are summarized in Proposition (ref) below.

propositionSuppose that model ((ref)) along with Assumption (ref) and (ref) hold. Then, the following bounds hold for the $ATT(x_1)$: \begin{eqnarray*} ATT(x_1) \in \left[ \theta_{OLS}(x_1)- \sup_{x_0 \in \mathcal X_0} SB_0(x_0), \theta_{OLS}(x_1)- \inf_{x_0 \in \mathcal X_0} SB_0(x_0) \right] \equiv \Theta_I(x_1). \end{eqnarray*} These bounds are uniformly sharp across $x_1$, and $\Theta_I(x_1)$ is the identified set for the $ATT(x_1)$.

Using the results in Proposition (ref), we can then obtain sharp bounds on $ATT$ by integrating the bounds over the conditional distribution of $X_1$ in the treatment group ($D=1$):

eqnarray*[eqnarray* omitted — 208 chars of source]

The following Proposition (ref) provides a doubly robust estimand for $\int\theta_{OLS}(x_1)d F_{X_1\vert D=1}(x_1)$. This result is probably achieved in the literature, but since we could not find a closed-form expression of a doubly robust estimand for this quantity, we provide an estimand along with its proof for completeness.

propositionConsider the following estimand \begin{eqnarray*} \tau^{DR} &\equiv& \frac{1}{\mathbb E[D]}\mathbb E \bigg[ \frac{D-P(X_1)}{1-P(X_1)} \big(Y_1 - \mu_0(X_1) \big)\bigg], \end{eqnarray*} where $P(X_1)$ and $\mu_0(X_1)$ are postulated models for the true propensity score $\mathbb E[D \vert X_1]$ and the conditional outcome mean $\mathbb E[Y_1 \vert D=0, X_1]$, respectively. Then, $\tau^{DR} = \int\theta_{OLS}(x_1)d F_{X_1\vert D=1}(x_1)$ if either (but not necessarily both) $P(X_1)=\mathbb E[D \vert X_1]$ almost surely (a.s.) or $\mu_0(X_1) = \mathbb E[Y_1 \vert D=0, X_1]$ a.s.

Note that the proposed estimand $\tau^{DR}$ is equal to the desired quantity $\int\theta_{OLS}(x_1)d F_{X_1\vert D=1}(x_1)$ even if either the propensity score function or the conditional outcome mean function is misspecified. However, if both functions are misspecified, $\tau^{DR}$ is generally different from $\int\theta_{OLS}(x_1)d F_{X_1\vert D=1}(x_1)$.

exampleConsider another version model ((ref)) where \begin{eqnarray*} \left\{ \begin{array}{lcl} Y_t&=&(1+ X_t)U+\theta X_t D*t\mathbbm{1}\{t\geq 0\} \\ \\ D&=&\mathbbm{1}\{U\geq 1\}\\ \\ U &\sim& N(0,1), X_t \sim \mathcal U_{\left[0,\frac{1}{1+t^2}\right]}, and X_t\ \rotatebox[origin=c]{90}{$\models$}\ U \end{array} \right. \end{eqnarray*} and $\mathcal I_0=\mathcal X_0=[0,1]$. In this model, $SB_t(x_t)=(1+ x_t)(\alpha_1-\alpha_0)$ where $\alpha_1=\frac{\phi(1)}{1-\Phi(1)}\approx 1.53$ and $\alpha_0=-\frac{\phi(1)}{\Phi(1)} \approx -0.29$. We have $SB_0(x_0) \in [\alpha_1-\alpha_0, 2(\alpha_1-\alpha_0)]$ and $ SB_1(x_1) \in \left[\alpha_1-\alpha_0, 1.5(\alpha_1-\alpha_0)\right] \subseteq [\alpha_1-\alpha_0, 2(\alpha_1-\alpha_0)]\equiv \Delta_{SB_{0X}}$. So, the standard parallel trends assumption does not hold as $X_0\neq X_1$. However, the selection bias $SB_1(x_1) $ in period 1 belongs to the convex hull of all selection biases in period 0, i.e., $SB_1(x_1) \in \Delta_{SB_{0X}}$. Hence, our identifying assumption holds. We have $Y_1(1)=(1+X_1)U+\theta X_1$. Then, $\theta_{OLS}(x_1)=(1+ x_1)(\alpha_1-\alpha_0)+\theta x_1$, which implies the bounds $ATT(x_1) \in [(x_1-1)(\alpha_1-\alpha_0)+\theta x_1,x_1(\alpha_1-\alpha_0)+\theta x_1]$. The actual conditional ATT function is $ATT(x_1)=\theta x_1$. Figure (ref) shows the bounds for different values of $x_1$ when $\theta=2$. It appears that the bounds are informative and identify the sign of the ATT for values of $x_1$ bigger than $0.5$. \begin{figure}[h] \caption{Illustration of $\Theta_I$ for $\theta = 2$ and $x_1 \in [0, 1]$} \end{figure}

A sufficient condition for Assumption (ref)

As in the previous section, we are going to provide a sufficient condition under which Assumption (ref) holds. We slightly modify Assumption (ref) to the following.

assumption$\,$ \begin{enumerate}[(i)] • The untreated potential outcome satisfies: $Y_t(0)=g(X_t) \lambda(U)+\gamma(V)+\varepsilon_t,$ for $t\in \{0,1\},$ where $(X_t,U,V,\varepsilon_t)$ is a random vector satisfying $X_t\ \rotatebox[origin=c]{90}{$\models$}\ (U,V,\varepsilon_t)$, $\varepsilon_t\ \rotatebox[origin=c]{90}{$\models$}\ (U,V)$, and $g(.)$, $\lambda(.)$ and $\gamma(.)$ are three unknown (nontrivial) functions. • The function $g(.)$ is nondecreasing in $x$, and $Supp(X_1) \subseteq Supp(X_0)$. • The treatment receipt is defined as $D=h(U,V)$, where $h$ is a nontrivial function. \end{enumerate}

Assumption (ref) is a modified version of Assumption (ref) to allow the factor to depend on some time-varying covariate $X_t$. Assumption (ref).((ref)) postulates that in the interactive fixed effects structure for the untreated potential outcome, the time-varying factor is determined by some potentially time-varying covariate $X_t$. Assumption (ref).((ref)) imposes some monotonicity condition on the factor function $g$. It also imposes a support condition on the time-varying covariate $X_t$ over time, which holds if the covariate is not changing over time. Assumption (ref).((ref)) is the same as Assumption (ref).((ref)) and postulates that the selection into treatment is function of the time-invariant unobservables in the model.

propositionSuppose $\mathcal I_0=\mathcal X_0$, and Assumption (ref) holds. Then Assumption (ref) holds while conditional PT fails to hold.

We can broaden conditions (ref).((ref)) and (ref).((ref)) in Assumption (ref) by replacing them by $Y_t(0)=g_t(X_t) \lambda(U)+\gamma(V)+\varepsilon_t$ along with the other restrictions, and $Supp(g_1(X_1)) \subseteq Supp(g_0(X_0))$, respectively. The function $g$ in the potential outcome model now has a subscript $t$, which allows $X_t$ to be the same random variable across time periods ($X_0=X_1$).

Before we move on, let us elaborate on our contribution to the literature. As can be seen from Propositions (ref) and (ref), our approach does not require the support of the outcome variable to be bounded as it is customary in the literature on partial identification. Furthermore, the approach does not rely on a sensitivity parameter as in ManskiPepper2018, and RambachanRoth2020. Our bounds can still be informative in situations where there are only two periods, 0 (baseline) and 1 (follow-up), as long as there exists other information available from observed baseline characteristics, or multiple data sources that may not be representative of the target population. Below, we provide a deeper comparison of our method to that of RambachanRoth2020 when our information set contains only pre-treatment periods.

Comparison with RambachanRoth2020's (RambachanRoth2020) approach

First, for the sake of simplicity suppose the information set $I_0$ contains two pre-treatment periods $-1$ and $0$, such that $\mathcal I_0=\{-1,0\}$. Define $\delta\equiv(\delta_{-1},\delta_1)'$, where

eqnarray*[eqnarray* omitted — 192 chars of source]

Observe that $\delta_1=SB_1-SB_0$, and $\delta_{-1}=SB_{-1}-SB_0$.\footnote{In the terminology of RambachanRoth2020, $\delta_{-1}=\delta_{pre}$ and $\delta_{1}=\delta_{post}$.} Note that $\delta_{-1}$ is identified. Lemma 2.1 in RambachanRoth2020 provides a general characterization of the ATT if a researcher is willing to make a restriction that $\delta_1$ belongs to a closed and convex set. In our setting, we assume that $SB_1 \in [\min\{SB_{-1}, SB_0\},\max\{SB_{-1}, SB_0\}]$, which implies $\delta_{1} \in \left[\min\{SB_{-1}-SB_0,0\}, \max\{SB_{-1}-SB_0,0\}\right]$. Therefore, we can recast our framework in theirs where $\delta \in \{SB_{-1}-SB_0\}\times \left[\min\{SB_{-1}-SB_0,0\}, \max\{SB_{-1}-SB_0,0\}\right],$ which is a closed and convex set. Hence, our approach can be viewed as a special case of their method. However, they do not consider the type of restrictions we consider in this paper. We now compare our assumptions to the restrictions considered in RambachanRoth2020.

Smoothness restrictions

The differential trends evolve smoothly over time with slope changing by no more than $M$ between consecutive periods:

eqnarray*[eqnarray* omitted — 125 chars of source]

where $\delta_0$ is normalized to be equal to zero. We then have $\Delta^{SD}(M)\equiv \left\{\delta: \vert \delta_1 + \delta_{-1} \vert \leq M \right\}$. The parameter $M\geq 0$ is like a sensitivity parameter and governs the amount by which the slope of the differential trends can change between consecutive periods.

Under the smoothness restriction, we obtain the following bounds on the selection bias $SB_1$:

eqnarray*[eqnarray* omitted — 69 chars of source]

Our bounding approach yields the following bounds on $SB_1$:

eqnarray*[eqnarray* omitted — 76 chars of source]

In Appendix (ref), we show that if $SB_{-1} \neq SB_0$, there exists no value of $M$ such that the above two sets of bounds on $SB_1$ coincide. Furthermore, we show that there exist no values of $M$ for which RambachanRoth2020's (RambachanRoth2020) bounds are tighter than ours, while there exist values of $M$ for which our bounds are tighter than theirs $(M > 2\vert SB_0-SB_{-1}\vert)$.

Bounding relative magnitudes

This approach bounds the worst-case post-treatment violation of parallel trends in terms of the worst-case violation in the pre-treatment period:

eqnarray*[eqnarray* omitted — 144 chars of source]

where $\bar{M}\geq 0$ behaves as a sensitivity parameter. This implies the following bounds on $SB_1$:

eqnarray*[eqnarray* omitted — 112 chars of source]

In Appendix (ref), we show that if $SB_{-1} \neq SB_0$, there exists no value of $\bar{M}$ such that the above bounds on $SB_1$ coincide with ours. When $\bar{M }>1$, our bounds are tighter than RambachanRoth2020's (RambachanRoth2020), and there exist no positive values of $\bar{M}$ for which their bounds are tighter than ours.

Second, our approach covers the case where no pre-treatment trend exists, but there are multiple elements (baseline characteristics/data sets) available in the information set in period 0. Their methodology is silent about such a case.

Third, our approach does not require the knowledge of a sensitivity parameter, while the two restrictions they consider do. How to choose the values of the sensitivity parameters $M$ and $\bar{M}$ remains unclear in their approach.

Possibility of discordancy between RambachanRoth2020's (RambachanRoth2020) bounds and ours

In this subsection, we study the existence of possible discordancy between the restrictions we consider in the paper and those considered in RambachanRoth2020. We find that under the smoothness restrictions, when $SB_{-1} \neq SB_0$, the two bounds are discordant if $M < \vert SB_0-SB_{-1}\vert$, i.e., their intersection is empty if $M < \vert SB_0-SB_{-1}\vert$. Kedagni_etal2020 pointed out that when a full model is rejected, researchers should be cautious about the way they relax the model to avoid this kind of situations. We then recommend researchers not to use the smoothness restrictions with values of $M$ less than $\vert SB_0-SB_{-1}\vert$. This scenario never happens under the bounding relative magnitudes restriction, i.e., there exists no possible discordancy between our bounds and those obtained under this latter restriction. The two bounds always overlap.

Policy-oriented generalized DID estimand

The main question we are trying to answer is how to obtain the selection bias $SB_1$. Given the baseline information $I_0$, we are going to assume that the decision maker will choose the selection bias $SB_1$ in such a way that a loss function is minimized. By plugging such an optimal selection bias $SB_1$ into the definition of the generalized DID, we obtain what we call a policy-oriented generalized difference-in-differences (PO-GDID) estimand. This estimand may not have a causal interpretation, but it may help the policy-maker in her decision making process.

Best predictor of $SB_1$ based on a loss function

assumptionLet $\mathcal L(SB_1,SB_0, I_0)$ be the decision maker's loss function when she assumes that the selection bias is $SB_1$ in the presence of the baseline information $I_0$ and the selection bias $SB_0$. The decision maker chooses $SB_1$ to minimize the loss $\mathcal L(SB_1,SB_0, I_0)$: $SB_1(SB_0, I_0)=\arg\min \mathcal L(SB_1,SB_0, I_0)$.

In this paper, we consider the class of $p$-norm losses defined as: $$\mathcal L_p (SB_1,SB_0, I_0)=\left(\mathbb E_{I_0}\left[\vert SB_1 - SB_0(I_0)\vert ^p \right]\right)^{1/p},$$ where $1 \leq p \leq \infty$. We are going to derive the optimal selection bias $SB_1$ for $p\in \left\{1,2,\infty\right\}$. We consider those special loss functions because the solutions to the optimization problem have closed-form expressions. Other loss functions can also be considered.

$L1$ loss: Mean absolute error (MAE)

$\mathcal L_1 (SB_1,SB_0, I_0)=\mathbb E_{I_0}\left[\vert SB_1 - SB_0(I_0)\vert \right]$.

Given this $L1$ loss function, under Assumption (ref), the decision maker solves the following optimization problem:

eqnarray*[eqnarray* omitted — 86 chars of source]

The optimal decision is to set the selection $SB_1$ to be equal to the median selection bias in the baseline period, i.e., $SB_1=Med_{I_0}(SB_0(I_0))$. In such a case, the policy-oriented generalized DID estimand is given by $$\theta_{PO-GDID}=\theta_{OLS}-Med_{I_0}(SB_0(I_0)).$$

$L2$ loss: Root mean square error (RMSE)

$\mathcal L_2 (SB_1,SB_0, I_0)=\left(\mathbb E_{I_0}\left[\vert SB_1 - SB_0(I_0)\vert ^2 \right]\right)^{1/2}$.

Minimizing the RMSE is equivalent to minimizing the mean square error (MSE). Therefore, under Assumption (ref), the decision maker solves the following optimization problem:

eqnarray*[eqnarray* omitted — 90 chars of source]

This yields an optimal decision for the selection $SB_1$ to be set equal to the average selection bias in the baseline period, i.e., $SB_1=\mathbb E_{I_0}[SB_0(I_0)]$. Hence, we have $$\theta_{PO-GDID}=\theta_{OLS}-\mathbb E_{I_0}[SB_0(I_0)].$$

$L\infty$ loss: Maximal regret

$\mathcal L_{\infty} (SB_1,SB_0, I_0)=\text{ess}\sup_{\mathcal I_0}\vert SB_1 - SB_0(I_0)\vert$, where $\text{ess}\sup$ denotes essential supremum and is defined as follows: $$\text{ess}\sup_{\mathcal I_0} f=\inf\left\{M: \mathbb P(\iota_0 \in \mathcal I_0: f(\iota_0) \leq M)=1\right\}.$$

For simplicity, assume $\text{ess}\sup_{\mathcal I_0}\vert SB_1 - SB_0(I_0)\vert =\sup_{\iota_0 \in \mathcal I_0}\vert SB_1 - SB_0(\iota_0)\vert$. Then

eqnarray*[eqnarray* omitted — 348 chars of source]

Therefore, the minimum of $\mathcal L_{\infty} (SB_1, SB_0, I_0)$ is obtained when the two arguments of the max function are equal, i.e., $SB_1 - \inf_{\iota_0 \in \mathcal I_0}SB_0(\iota_0) = \sup_{\iota_0 \in \mathcal I_0}SB_0(\iota_0)-SB_1$. This implies $SB_1=\frac{1}{2}(\inf_{\iota_0 \in \mathcal I_0}SB_0(\iota_0)+\sup_{\iota_0 \in \mathcal I_0}SB_0(\iota_0))$, and $\mathcal L_{\infty} (SB_1)=\frac{1}{2}(\inf_{\iota_0 \in \mathcal I_0}SB_0(\iota_0)+\sup_{\iota_0 \in \mathcal I_0}SB_0(\iota_0))$.

This optimization problem with the $L\infty$ loss is equivalent to a minimax criterion, and yields the mid-point of the bounds on $SB_1$ stated in Assumption (ref). Hence, the PO-GDID estimand is given by $$\theta_{PO-GDID}=\theta_{OLS}-\frac{1}{2}(\inf_{\iota_0 \in \mathcal I_0}SB_0(\iota_0)+\sup_{\iota_0 \in \mathcal I_0}SB_0(\iota_0)).$$

Note that in all cases, if the information in the baseline period is a singleton, then the optimal $SB_1$ is the selection bias in the baseline period $SB_0$, which is equivalent to the parallel trends assumption. Unlike the PO-GDID estimand obtained from $L1$ and $L2$ loss functions, that obtained from the $L\infty$ loss function does not require the knowledge of the distribution of the information $I_0$ but only its support and is easy to compute. However, when the distribution of $SB_0(I_0)$ is uniform over $\left[ \inf_{\iota_0\in \mathcal I_0} SB_0(\iota_0), \sup_{\iota_0 \in \mathcal I_0} SB_0(\iota_0)\right]$, then the optimal selection bias $SB_1$ is the same in all three cases.

Let $\Lambda$ denote the set of possible distributions for $SB_0(I_0)$, and $SB_1(SB_0,\lambda)$ denote the optimal selection bias in period 1 given the distribution $\lambda \in \Lambda$ for $SB_0(I_0)$. Define $ATT_{\lambda} \equiv \theta_{OLS}-SB_1(SB_0,\lambda)$.

definitionWe define the robust GDID bounds as follows: \begin{eqnarray*} ATT \in \left[\inf_{\lambda \in \Lambda}ATT_{\lambda}, \sup_{\lambda \in \Lambda}ATT_{\lambda}\right]. \end{eqnarray*}

The following lemma holds.

lemmaThe robust GDID bounds coincide with the bounds in Proposition (ref) for the $L1$, $L2$, and $L\infty$ loss functions.

A sufficient condition for the policy-oriented generalized DID estimand to be equal to the ATT is that the potential outcome $Y_1(0)$ satisfies: $$Y_1(0)\sim \mathbb E[Y_1\vert D=0]+SB_1D +\varepsilon,$$ where $\mathbb E[\varepsilon \vert D]=0.$

Forecasting $SB_1$ when the baseline information is ordered

Suppose that the baseline information set $\mathcal I_0$ is ordered (e.g., a set of multiple pre-treatment periods $\mathcal I_0=\{-T_0, -T_0+1, \ldots, -1, 0\}$ or a continuous baseline covariate $X_0$ like age). We can regress $SB(I_0)$ on $\{I_0, I_0^2, \ldots\}$ and use this regression to predict $SB_1$.

For instance, if $\mathcal I_0=\{-T_0, -T_0+1, \ldots, -1, 0\}$ and the selection bias $SB(I_0)$ is increasing over time, Assumption (ref) may not hold. The researcher could instead use this increasing trend information about the selection bias to forcast the next period selection bias $\widehat{SB}_1$.

Estimation and inference

We briefly describe our estimation and inferential method. We assume that the information set $\mathcal I_0$ is finite. We can write the robust DID bounds $\Theta_I$ as the convex hull of the doubly-robust DID estimands as follows:

eqnarray*[eqnarray* omitted — 304 chars of source]

We can then take the convex hull of the confidence intervals of all DID estimands $\tau^{DR} - SB_0(\iota_0)$ to obtain valid confidence bounds for $\Theta$. More precisely, the confidence bounds can written as

eqnarray[eqnarray omitted — 244 chars of source]

where $CI^{1-\alpha}_{LB}(\tau^{DR}-SB_0(\iota_0))$ (resp. $CI^{1-\alpha}_{UB}(\tau^{DR}-SB_0(\iota_0))$) denotes the lower (resp. upper) bound of the $(1-\alpha)$-confidence interval of the parameter $\tau^{DR}-SB_0(\iota_0)$. But, these confidence bounds could be too conservative.

The proof of validity of this procedure is provided in Appendix (ref). The argument is similar to that of Berger1996 for union bounds. Indeed, Berger1996 showed that the union of the confidence regions has at least the same coverage rate as each confidence region. The confidence bounds in (ref) are similar to those in KolesarRothe2018 derived in a regression discontinuity design setting.

To implement these confidence bounds, we first estimate the propensity score function $P(X_1)$ (e.g., logit specification) and the outcome regression function $\mu_0(X_1)$ (e.g., linear or quadratic specification). Second, in order to obtain correct standard errors for each estimator $\hat{\tau}^{DR}-\widehat{SB}_0(\iota_0),$ we use a bootstrap method.\footnote{Note that we do not bootstrap $\min_{\iota_0 \in \mathcal I_0}\{\hat{\tau}^{DR}-\widehat{SB}_0(\iota_0)\}$ nor $\max_{\iota_0 \in \mathcal I_0}\{\hat{\tau}^{DR}-\widehat{SB}_0(\iota_0)\}$. As pointed out by Fang_al2019, the standard bootstrap is inconsistent in this case since the limiting distributions of these estimators are not Gaussian.}

Extensions

Extension to multiple treatment periods

In this subsection, we generalize our analysis to a setting where the treatment receipt occurs at multiple periods. We consider the following multiple treatment periods model:

eqnarray[eqnarray omitted — 329 chars of source]

where $Y_t$ denotes the observed outcome in period $t$, $D_t$ is the observed treatment status in period $t$ with $D_0=0$ by definition, while $Y_t(0,d_1, \ldots, d_T)$ is the potential outcome when the treatment path $(D_0,D_1, \ldots, D_T)$ is externally set to $(0,d_1,\ldots, d_T)$.\footnote{See Robins1986, Robins1987, and Han2021 for a similar definition of the potential outcome model.} Under the no-anticipation assumption, we have $Y_0(0,d_1,\ldots, d_T) =Y_0(0)$ for all $(d_1,\ldots,d_T) \in \{0,1\}^T.$ We assume that individuals do not anticipate any effects of the treatment before it occurs for the first time. However, we allow the individuals to anticipate the effects of the treatment for the rest of the period. This assumption is less restrictive than the commonly used no-anticipatory effects assumption.

Identification without covariates

In the above framework, the parameter of interest is the average treatment effect on the treated group following the path $(0,d_1',\ldots, d_T')$ to $(0,d_1,\ldots, d_T)$ in period $t$, which is defined as:

eqnarray*[eqnarray* omitted — 225 chars of source]

This parameter may help reveal some dynamic effect of the treatment. For example, in a non-staggered design framework, the parameter $ATT_t[(0,\ldots,0,d_s=0,0,\ldots, 0)\rightarrow (0,\ldots,0,d_s=1,0,\ldots, 0)]$ measures the dynamic effect of the treatment in period $t$ on people who were only treated in period $s$ compared to the status where they would have never been treated. Note that this setting requires a panel structure in the data. In a staggered design setting, the average treatment effect in period $t$ on units who are treated for the first time in period $g$ could be an interesting parameter, as considered in CallawaySantAnna2021: $$ATT_t[(0,\ldots,0,d_g=0,0,\ldots, 0)\rightarrow (0,\ldots,0,d_g=1,1,\ldots, 1)].$$

Similarly to what we have in the one post-treatment setting, we can write the difference-in-means estimand $(\theta_{DIM}^t)$ between the two groups $(0,d_1',\ldots, d_T')$ and $(0,d_1,\ldots, d_T)$ in period $t$ as:

eqnarray*[eqnarray* omitted — 316 chars of source]

where $SB_t[(0,d_1',\ldots, d_T')\rightarrow (0,d_1,\ldots, d_T)]\equiv \mathbb E\left[Y_t(0,d_1',\ldots, d_T') \vert (D_0,D_1, \ldots, D_T)=(0,d_1,\ldots, d_T)\right]-\mathbb E\left[Y_t(0,d_1',\ldots, d_T') \vert (D_0,D_1, \ldots, D_T)=(0,d_1',\ldots, d_T')\right]\equiv SB_t.$ We extend Assumption (ref) to the current setting.

assumption[Extended bias set stability] For each $t$, \begin{eqnarray*} SB_t \in \left[ \inf_{\iota_0\in \mathcal I_0} SB_{\iota_0}, \sup_{\iota_0 \in \mathcal I_0} SB_{\iota_0}\right] \equiv \Delta_{SB}, \end{eqnarray*} where $SB_{\iota_0}[(0,d_1',\ldots, d_T')\rightarrow (0,d_1,\ldots, d_T)]\equiv \mathbb E[Y_{\iota_0}(0)\vert (D_0,D_1, \ldots, D_T)=(0,d_1,\ldots, d_T)] - \mathbb E[Y_{\iota_0}(0)\vert (D_0,D_1, \ldots, D_T)=(0,d_1',\ldots, d_T')]\equiv SB_{\iota_0}$ is the baseline selection bias with respect to the treatment status in period $t$ when the information $I_0$ is equal to $\iota_0$.

Assumption (ref) is a generalization of Assumption (ref). In the appendix, we provide a sufficient condition for it to hold (Assumption (ref)).

propositionSuppose that model ((ref)) along with Assumption (ref) holds. Then, the following bounds are valid for $ATT_t$: \begin{eqnarray*} ATT_t[(0,d_1',\ldots, d_T')\rightarrow (0,d_1,\ldots, d_T)] \in \left[ \theta_{DIM}^t - \sup_{\iota_0 \in \mathcal I_0} SB_{\iota_0}, \theta_{DIM}^t- \inf_{\iota_0 \in \mathcal I_0} SB_{\iota_0} \right].\end{eqnarray*} These bounds are sharp, and $\Theta_I^t$ is the identified set for $ATT_t[(0,d_1',\ldots, d_T')\rightarrow (0,d_1,\ldots, d_T)]$.

One could be interested in a weighted average of all time periods treatment effects, i.e., $ATT[(0,d_1',\ldots, d_T')\rightarrow (0,d_1,\ldots, d_T)]=\sum_{t=1}^T \omega_t ATT_t[(0,d_1',\ldots, d_T')\rightarrow (0,d_1,\ldots, d_T)]$, with pre-specified weights $\omega_t \in [0,1]$. Bounds on this weighted average $ATT$ can then be obtained as follows:

eqnarray*[eqnarray* omitted — 268 chars of source]

For example, one can set $\omega_t=\frac{n_t}{\sum_{t=1}^T n_t}$, where $n_t$ denotes the cardinality of the treatment group in period $t$.

Another parameter that could be of interest is average treatment effect on people who are ever treated over the treatment period:

eqnarray*[eqnarray* omitted — 327 chars of source]

When the outcome variable is only a function of the current period treatment status, we denote by $Y_t(d_t)$ the potential outcome in period $t$ as $Y_t(0,d_1, \ldots, d_t, \ldots, d_T)=Y_t(0,d_1', \ldots, d_t, \ldots, d_T')$ for all $(0,d_1, \ldots, d_t, \ldots d_T)$ and $(0,d_1', \ldots, d_t, \ldots d_T')$. In such a case, without ambiguity, we denote $ATT_t=\mathbb E[Y_t(1)-Y_t(0) \vert D_t=1]$, $\theta^t_{OLS}=\mathbb E[Y_t \vert D_t=1]-\mathbb E[Y_t \vert D_t=0]$, $SB_t=\mathbb E[Y_t(0) \vert D_t=1]-\mathbb E[Y_t(0) \vert D_t=0]$, and $SB^t_{\iota_0}=\mathbb E[Y_{\iota_0}(0) \vert D_t=1]-\mathbb E[Y_{\iota_0}(0) \vert D_t=0]$.

Below, we propose a DGP in which parallel trends holds for each period.

example[PT holds] We consider a DGP in which there is selection on a time-invariant unobservable and there are no instrumental variables available. \begin{eqnarray*} \left\{ \begin{array}{lcl} Y_{t} &=& U+\varepsilon_t +\theta_t D_t\\ \\ D_t &=& \mathbbm{1}\{U \geq 2-\frac{t}{T}\} \end{array} \right. \end{eqnarray*} where $U\ \rotatebox[origin=c]{90}{$\models$}\ (\varepsilon_t, \theta_t)$, $\theta_t\sim \mathcal{U}_{[0,1+t^2]}$, $\mathcal I_0=\{0\}$, and $\varepsilon_t \sim \mathcal N(t^2,1)$. In this DGP, $\mathbb E[Y_t(0)-Y_0(0) \vert D_t=1]=\mathbb E[Y_t(0)-Y_0(0) \vert D_t=0]$. Therefore, PT holds. Hence, $ATT_t$ is point-identified as $\theta_{OLS}^t-SB^t_0=\frac{1+t^2}{2}$.

In the next example, we propose a DGP in which PT does not holds, but Assumption (ref) does.

example[PT is violated] We consider a DGP in which there is selection on a time-varying unobservable and there are no instrumental variables available. \begin{eqnarray*} \left\{ \begin{array}{lcl} Y_{t} &=& (\vert t \vert - 1) U_t+\theta_t D_t\\ \\ D_t &=& \mathbbm{1}\{U_t \geq 2-\frac{t}{T}\}\ for \ t=1,2 \end{array} \right. \end{eqnarray*} where $U_t\ \rotatebox[origin=c]{90}{$\models$}\ \theta_t$, $\theta_t\sim \mathcal{U}_{[0,4+t^2]}$, $(U_{-3}, U_{-2}, \cdots, U_2)' \sim \mathcal N ( \bm{\mu}, \bm{\Sigma})$ with \begin{eqnarray*} \bm{\mu} &=& (2, \cdots, 2)' \\ \bm{\Sigma} &=& \begin{pmatrix} 1 & \rho &\cdots & \rho^5 \\ \rho & 1 & \cdots& \rho^4 \\ \vdots & \vdots & \ddots & \vdots \\ \rho^5 & \rho^4 & \cdots & 1 \end{pmatrix}, \end{eqnarray*} $\rho= 0.9$, the baseline information set $\mathcal I_0=\{-3,-2,-1,0\}$ is the set of available pre-treatment periods, and $T=2$. In this DGP, $\mathbb E[Y_t(0)-Y_0(0) \vert D_t=1]\neq \mathbb E[Y_t(0)-Y_0(0) \vert D_t=0]$. In particular, we obtain $\mathbb E[Y_t(0)-Y_0(0) \vert D_t=1] = (\vert t \vert - 1 + \rho^{-t})\big[ \frac{\phi(-t/T)}{1-\Phi(-t/T)} \big]$ whereas $\mathbb E[Y_t(0)-Y_0(0) \vert D_t=0] = (\vert t \vert - 1 + \rho^{-t})\big[ - \frac{\phi(-t/T)}{\Phi(-t/T)} \big]$. Hence, PT fails to hold. However Assumption (ref) holds because we have $SB_t = (\vert t \vert - 1 )\big[ \frac{\phi(-t/T)}{(1-\Phi(-t/T))\Phi(-t/T)} \big]$ and $SB^t_{\iota_0} = \rho^{(t-\iota_0)}(\vert \iota_0 \vert - 1 )\big[ \frac{\phi(-t/T)}{(1-\Phi(-t/T))\Phi(-t/T)} \big]$. The following Figure (ref) shows the identified set $\Theta_{I, t}$ and $ATT_t$ for $t= 1, 2$, where the sign of $ATT_t$ is correctly identified for both periods. In each period, $\Theta_{I, t}$ is represented as a line interval, and a circle shows the true $ATT_t$. Note that $ATT_1 = 2.5 \in \Theta_{I, 1} \approx [0.33, 3.99]$ and $ATT_2 = 4 \in \Theta_{I, 2} \approx [3.67, 7.28]$. \begin{figure}[h] \caption{Illustration of $\Theta_I$ and $ATT_t$ for $t= 1, 2$} \end{figure}

It is important to note that the above framework applies to both staggered and non-staggered designs. In a non-staggered design DID framework, we have $2^T$ possible treatment paths, while in the staggered design case there are $T+1$ possible treatment paths.

A two-way fixed effects regression approach

Without covariates, our identified set for $ATT_t[(0,d_1',\ldots, d_T')\rightarrow (0,d_1,\ldots, d_T)]$ can be computed using a two-way fixed effects (TWFE) regression approach. Suppose that we observe $T$ treatment periods. Define $D^g\equiv \mathbbm{1}\{(D_0,D_1,\ldots,D_g,\ldots,D_T)=(0,0,\ldots,d_g=1,\ldots,1)\}$, and $D^0\equiv \mathbbm{1}\{(D_0,D_1,\ldots,D_T)=(0,0,\ldots,0)\}$. Consider the following regression for $t\in \{0,1,\ldots,T\}$

eqnarray[eqnarray omitted — 203 chars of source]

where the subscript $i$ refers to individual $i$, and $i=1\ldots,N$.

We have

eqnarray*[eqnarray* omitted — 430 chars of source]

Then,

eqnarray*[eqnarray* omitted — 373 chars of source]

Therefore under PT, $\mathbb E[\varepsilon_{is} \vert D^g_i=1]-\mathbb E[\varepsilon_{is} \vert D^0_i=1]=\mathbb E[\varepsilon_{i0} \vert D^g_i=1]-\mathbb E[\varepsilon_{i0} \vert D^0_i=1],$ and we have

eqnarray*[eqnarray* omitted — 201 chars of source]

That is, $\theta_{DIM}^s(D^g=1)-SB_0^s(D^0=1)=\theta^g_s$. For illustration, see Example (ref) in the appendix. This result is similar to the idea developed in Wooldridge2021 when PT holds.

Now, let us consider the case where PT may not hold. Suppose that the information set $\mathcal I_0$ is the set of pre-treatment periods. For each $\iota_0 \in \mathcal I_0$, we can run the TWFE regression for $t\in \{\iota_0,1,2, \ldots, T\}$. We then obtain a 95% confidence interval for $CI^{\iota_0}(\hat{\theta}^g_s)$ for $\theta^{g,\iota_0}_s$ from the TWFE regression. Therefore, we can obtain a 95% CI for $ATT_s(D^g=1 \rightarrow D^0=1)$ as $$\left[\min_{\iota_0 \in \mathcal I_0}CI^{\iota_0}_{LB}(\hat{\theta}^g_s), \max_{\iota_0 \in \mathcal I_0}CI^{\iota_0}_{UB}(\hat{\theta}^g_s)\right],$$ where $CI^{\iota_0}_{LB}(\hat{\theta}^g_s)$ and $CI^{\iota_0}_{UB}(\hat{\theta}^g_s)$ denote the lower and upper bounds on the confidence interval for $\theta^{g,\iota_0}_s$, respectively.

Identification with covariates

Let

eqnarray*[eqnarray* omitted — 104 chars of source]

where $X = (X_1, \cdots, X_T)$. Then, the result in Proposition (ref) holds, except that we replace $\theta_{OLS}(x_1)$ by $\theta_{DIM}^t(g, x),$ where $x=(x_1, \cdots, x_T)$.

Doubly-robust estimand for the staggered adoption case with covariates

The following proposition holds.

propositionConsider the following estimand \begin{eqnarray*} \tau_t^{g, DR} &\equiv& \frac{1}{\mathbb E[D^g]}\mathbb E \bigg[ \bigg(D^g - \frac{P^g(X)}{P^0(X)}D^0 \bigg) \big(Y_t - \mu_0^t(X) \big)\bigg], \end{eqnarray*} where $P^s(X)$ and $\mu_0^t(X)$ are postulated models for the true propensity scores $\mathbb{E}[D^s \vert X]$ for all $s = 0, \cdots, T$ and the conditional outcome mean $\mathbb E[Y_t \vert D^0=1, X]$, respectively, and $X = (X_1, \cdots, X_T)$. Then, $\tau_t^{g, DR} = \int \theta_{DIM}^t(g, x) d F_{X \vert D^g = 1} (x)$ if either (but not necessarily both) $P^s(X)=\mathbb{E}[D^s \vert X]$ a.s. for all $s\in\{0,1,\ldots,T\}$ or $\mu_0^t(X) = \mathbb E[Y_t \vert D^0=1, X]$ a.s.

Extension to synthetic control

Suppose we observe $J+1$ units, and without loss of generality only the first unit is exposed to the intervention. Let $\mathcal J=\{2, \ldots, J+1\}$ denote the donor pool. For simplicity, suppose first that we only have two periods, such that $\mathcal I_0=\{0\}$. We write the model as:\footnote{One can alternatively assume that $Y_1(0) \equiv \sum_{j\in \mathcal J} \lambda_j Y_1^j(0) + \varepsilon_1(0)$. As long as $\varepsilon_1(0)$ is exogenous, i.e., $\varepsilon_1(0)\ \rotatebox[origin=c]{90}{$\models$}\ D$, our approach would work, since we only need $\mathbb E[Y_1(0)] = \sum_{j\in \mathcal J} \lambda_j \mathbb E[Y_1^j(0)].$}

eqnarray[eqnarray omitted — 211 chars of source]

Define $ATT^j=\mathbb E[Y_1(1)-Y_1^j(0) \vert D=1]$. We can check that $ATT=\sum_{j\in \mathcal J} \lambda_j ATT^j$.

Before explaining how our approach can be extended to this synthetic control (SC) framework, we briefly discuss the SC method. Abadie_etal2003 and Abadie_etal2010 propose to choose $\lambda_2$, \ldots, $\lambda_{J+1}$ so that the resulting synthetic control best resembles the pre-intervention values for the treated unit of predictors of the outcome variable, subject to the restriction that the weights are nonnegative and sum to one. There are at least two issues with their approach. First, the weights obtained using their approach may not be the same weights that we are looking for in the intervention period. The approach implicitly relies on the assumption that the weights are stable across covariates and also between baseline and treatment periods. Second, a solution to their problem may not exist Shial2023. Our approach does not suffer from these above issues.

Each donor $j \in \mathcal J$ is a potential control group for the treatment group: $\lambda_j \geq 0$ for all $j\in \mathcal J$, and $\sum_{j \in \mathcal J} \lambda_j=1$. However, we do not know the weights $\lambda_j \geq 0$ for any donor $j$. Assuming that the selection bias when considering each donor $j$ as a control in period 1 lies within the convex hull of all selection biases in period 0, we obtain the worst-case bounds for the ATT as: $\left[\min_j \underline{\theta}_{ATT}^j, \max_j \overline{\theta}_{ATT}^j\right]$, where $\underline{\theta}_{ATT}^j=\theta^j_{OLS}-\max_j SB^j_0$, $\overline{\theta}_{ATT}^j=\theta^j_{OLS}-\min_j SB^j_0$, and $SB^j_t=\mathbb E[Y_t(0)\vert D=1]-\mathbb E[Y^j_t(0)]$. Indeed, we have:

eqnarray*[eqnarray* omitted — 489 chars of source]

where the third equality holds from the definition of $Y_1(0)$, the fourth holds from the definition of the potential outcome model, the fifth holds because $Y_1^j \equiv Y_1 \vert J=j$, and $\sum_{j \in \mathcal J} \lambda_j=1$. We are abusing the notation by considering $J$ as a random variable.

Similarly, we can show that $\theta_{OLS}=ATT + \sum_{j \in \mathcal J} \lambda_j SB^j_1$. Therefore, $$ATT=\sum_{j \in \mathcal J} \lambda_j (\theta_{OLS}^j - SB^j_1).$$ Hence, under the assumption $SB^j_1 \in [\min_j SB^j_0, \max_j SB^j_0]$, the above bounds for the ATT are valid.

When the information set $\mathcal I_0$ has more than one element, the bounds on the selection bias $SB^j_1$ become: $SB^j_1 \in \left[\min_{\iota_0 \in \mathcal I_0}\min_j SB^j_0(\iota_0), \max_{\iota_0 \in \mathcal I_0} \max_j SB^j_0(\iota_0)\right]$.

Empirical illustrations

In this section, we illustrate our framework using some empirical examples. First, using the dataset from kresch2020, we show our robust GDID bounds and our PO-GDID estimands with the various loss functions we discussed as well as the point-estimand obtained by forcasting the treatment period selection bias using a linear projection method. There are five pre-treatment periods that represent the information set we consider. We use our doubly-robust estimand throughout this illustration. Secondly, we consider the application of cawley2021SSB where we do not have covariates to control for. We construct the information set using the two pre-treatment periods available in the data, and some of the qualitative results can be shown to be robust to the relaxation of the standard PT assumption. Lastly, cai2016's cai2016 analysis is chosen to demonstrate the multiple treatment period extension as well as the robust bounds or the PO-GDID estimands.

kresch2020

kresch2020 analyzed the effect of the legal reform in Brazil that clarified the relationship between municipal (local) and state governments in the water and sanitation sector. The reform was designed to eliminate the takeover threat by state companies toward municipal companies. Accordingly, the author tried to examine if the risk before the reform caused sub-optimal investment by the municipal providers by investigating if the reform led to increased investment in self-run municipal systems. The original estimation equation is the following standard two-way fixed effects (TWFE) specification with covariates $X$ entering the model linearly:

eqnarray[eqnarray omitted — 188 chars of source]

where $Y_{mt}$ is the investment level of municipality $m$ in year $t$, and the data includes 12 years from 2001 to 2012. The reform $Reform_{mt}$ is equal to 1 for self-run municipalities after the legislation was proposed ($t > 2005$),\footnote{DentehKedagni2022 pointed out that using the proposed bill date instead of its passage date could introduce misclassification in the DID framework. Investigating the consequences of misclassification in our framework is beyond the scope of the current paper.} and there are 5 pre-treatment periods. kresch2020 considered 7 outcome variables of $Y_{mt}$, and we follow him by considering the same outcomes as described in Table (ref).

table[table omitted — 575 chars of source]

The covariates $X_{mt}$ includes municipality $m$’s population, gross domestic product (GDP) and taxes (including national and state shares), water-intensive industry variables (e.g., agriculture, livestock production), and annual temperature and rainfall measures. Since each of the covariates is continuous, for the sake of tractability, we define our information set $\mathcal I_0$ using only the 5 pre-treatment periods. In particular, we assume a slightly different version of Assumption (ref) as follows:\footnote{Note that this is different from the simple bias set stability assumption (Assumption (ref)), and the results from this assumptions are also presented below.}

eqnarray*[eqnarray* omitted — 159 chars of source]

for all $x_1 \in \mathcal X_1$, where $\mathcal I_0 \equiv \{2001, \cdots, 2005 \}$.

For the doubly robust estimand $\tau^{DR}$, we primarily consider a logit model for the true propensity score $P(X_1)\equiv \mathbb E[D \vert X_1]$ and a linear model for the conditional outcome mean $\mu_0(X_1)\equiv \mathbb E[Y_1 \vert D=0, X_1]$. We also present results from a quadratic specification for the conditional outcome mean.

figure[figure omitted — 199 chars of source]

Figure (ref) shows the scatter plot for the unconditional selection bias in the pre-treatment periods for the 7 outcome variables. Recall that for each outcome variable, we consider the convex hull of the pre-treatment periods selection biases to be stable before and after the reform proposal in order to (partially) identify the ATT. For the last outcome variable (other investments) in particular, we demonstrate the projection-based identification method in the end of this subsection as there seems to be a trend in the selection biases in the pre-treatment periods.

table[table omitted — 1,284 chars of source]

Table (ref) summarizes the results for the robust GDID bounds where each row represents the results of each outcome variable. The first two columns show point estimates of the GDID bounds, and the third and fourth columns are the confidence bounds obtained using a bootstrap method. The last four columns contain the $\tau^{DR}$ estimates, the constructed selection bias sets from which the GDID bounds are estimated, and the results from kresch2020.

table[table omitted — 1,241 chars of source]

On the other hand , Table (ref) shows a slightly different results (though qualitatively the same as those in Table (ref)) from the quadratic specification of the conditional outcome mean $\mathbb E[Y_1 \vert D=0, X_1]$. Note that the slight difference in the results is driven by the different estimates for $\tau^{DR}$ in the fifth column, since the constructed selection bias sets remain the same (the last two columns).

Discussion

The findings from Tables (ref)) and (ref) suggest that the increase in investment after the reform bill was introduced is less significant than what the results in kresch2020 suggest. Our point estimate bounds show that the magnitude of the increase in investment is much smaller for all seven outcomes. The confidence bounds for all outcomes contain 0, suggesting that there is no significant change in investment after the reform. One could argue that our confidence bounds are too conservative and this may be driving the results. This does not seem to be the main reason since the point estimate bounds which are usually tighter than any confidence bounds lead to a similar conclusion. Note that in contrast to what the theory suggests kresch2020's kresch2020 point estimates lie outside our point estimate bounds. As pointed out by drdid2020 in their Remark 1, the TWFE specification ((ref)) considered in kresch2020 imposes some additional restrictions on the data generating process when assuming conditional PT. More precisely, the treatment effect is homogeneous in $X_{mt}$, and it rules out X-specific trends in both treatment and control groups $(\mathbb E[Y_1-Y_0 \vert X, D=d]=\mathbb E[Y_1-Y_0 \vert D=d])$. When these restrictions do not hold (which is likely the case here), the parameter $\delta$ will not identify the ATT and may not have any causal interpretation. This might be the reason why the point estimates in kresch2020 do not lie within our point estimate bounds.

table[table omitted — 1,138 chars of source]

We also compute our bounds without the use of covariates. Table (ref) displays the GDID bounds estimates under Assumption (ref) to demonstrate how the results change when ignoring the covariates. In this estimation, $\theta_{OLS}$ estimate is used in lieu of $\tau^{DR}$. We notice that total investment as well as self-financed investment and investment from loans and debt have significantly increased, while the other types of investment remained statistically stable after the reform.

We now present the PO-GDID estimands considering the three loss functions discussed in Section (ref): $L1$, $L2$, and $L\infty$. We weight each element in the information set by the ratio of the number of observations in it to the total number of observations in the information set. Since each element in the information set has the same number of observations (balanced panel), the probability weight for each of the five elements in the information set is equal to $1/5$. However, the distribution of the selection bias $SB_0(I_0)$ is not uniform over the interval $\left[ \inf_{\iota_0\in \mathcal I_0} SB_0(\iota_0), \sup_{\iota_0 \in \mathcal I_0} SB_0(\iota_0)\right]$, and the PO-GDID estimates are different across the three loss functions.

table[table omitted — 1,391 chars of source]

Table (ref) shows the estimation results for the three PO-GDID estimands. Each set of three columns contains the point estimates and 95% confidence intervals for each of the loss functions. The results are consistent with the findings above. There is not a statistically significant increase in any of the types of investment considered.

Finally, we demonstrate another type of GDID estimand results where the selection bias $SB_1$ is forcast from a simple regression of pre-treatment periods selection biases on time. More precisely, we first regress the five available $SB_0(\iota_0), \iota \in \mathcal{I}_0$ on the time variable $t$ as follows

eqnarray*[eqnarray* omitted — 70 chars of source]

Visually, the regression line obtained from the last outcome variable “Other Investments” is illustrated in Figure (ref), and we predict $SB_1(t)$ for the middle point of the post-treatment periods $t=2009$. Hence, using the estimate for $\tau^{DR}$ and $\widehat{SB_1}$, we obtain the GDID estimate. For instance, using “Other Investments,” we have $\widehat{\theta_{GDID}} = \widehat{\tau^{DR}} - \widehat{SB_1} = 495 - ( - 138) = 633$. The results for all other types of investment are summarized in Table (ref).

figure[figure omitted — 213 chars of source]
table[table omitted — 1,286 chars of source]

We can still find that the reform effect was not statistically significant for any type of investment at 5%, but for “Other Investments,” it was statistically significant at 10% as the lower bound of the 90% confidence interval is greater than 0.

cawley2021SSB

The authors examine the pass-through of a tax of two cents per ounce on sugar-sweetened beverages (SSB tax) enacted in Boulder, Colorado, using the standard DID framework. They considered both store and restaurant prices and collected two different datasets for each of them: hand-collected data and Nielsen retail scanner data for the store prices, and hand-collected data and web-scrapped (OrderUp.com) data for the restaurant prices. Hence, this exercise could have been the best example for us to explore the information set consisting of the multiple datasets, but we focus on utilizing multiple pre-treatment periods of the hand-collected datasets in this subsection due to the data limitation.\footnote{Nielsen retail scanner data are proprietary, and the unit price information is not available in the web-scrapped (OrderUp.com) data.}

Each dataset is bimonthly-collected and has four periods April, June, August, and October, where the tax was imposed on July 1st of the same year. Thus, our information set has two elements April and June. Moreover, we can implement the event-study type DID analysis (static heterogeneous treatment effects in multiple treatment periods model as in Equation ((ref))) to capture the non-parametric evolution of the treatment effects over the post-treatment periods. For the first dataset of the store prices, we consider three different prices ($post\_tax$, $reg\_tax$, and $untax$) as our outcome variables, and $fount$ is selected from the hand-collected restaurant dataset as another outcome variable of interest. In particular, $post\_tax$ uses post prices on the shelves, $reg\_tax$ uses prices at the register,\footnote{cawley2021SSB found that not all retailers included the tax in the posted (or shelf) prices; i.e., some retailers added the tax at the register making it less salient.} and $untax$ uses prices of products irrelevant to SSB tax (e.g., diet soda, products in which milk is the primary ingredient, alcoholic mixers, or coffee drinks) for the blind test. On the other hand, $fount$ represents restaurant fountain drink prices.

The control community is Fort Collins, Colorado, which is geographically close to Boulder and similar in demographic characteristics as well. Hence, the standard PT assumption states that the average equilibrium beverage price differences between Boulder and Fort Collins in the post-treatment periods would have been the same as the average equilibrium price differences in the pre-treatment periods if there had not been the SSB tax in Boulder. On the other hand, our GDID model assumes that the average equilibrium price differences without the tax in the post-treatment periods would have lain between the average equilibrium price differences in April and June between the two cities. Our approach can be seen as a robustness check of the findings in cawley2021SSB.

figure[figure omitted — 205 chars of source]

Before we proceed to the estimation results, we show the scatter plot of the selection biases in the pre-treatment periods. Figure (ref) shows the selection biases in April and June for each of the outcome variables $post\_tax$, $reg\_tax$, $untax$, and $fount$. Accordingly, we construct a set that contains both of the selection biases in April and June for each panel or the outcome variable and assume that the set will be stable in the following post-treatment periods and contain the unobserved post-period selection bias.

table[table omitted — 776 chars of source]

Tables (ref) shows the standard DID results and our GDID bounds results. The first column presents the standard DID estimate for the ATT, and the second and third columns are corresponding 95% confidence intervals. The fourth and fifth columns show lower and upper bounds of our identified set for the ATT as presented in Proposition (ref), and the corresponding 95% confidence intervals are given in the sixth and seventh columns. Note that we are still able to reject the null hypothesis that the effect on the post prices is not different from zero under a significance level of 5% from our GDID model for $post\_tax$, $reg\_tax$, and $fount$, implying that the same qualitative conclusion can be drawn from the GDID model where we do not have to maintain the standard PT assumption. However, our results suggest that a pass-through rate higher than 100% cannot be rejected from the GDID model for $post\_tax$ whereas the standard DID estimates rule out that case; the market could be imperfectly competitive.

figure[figure omitted — 195 chars of source]

Figure (ref) shows the multiple treatment periods GDID bounds estimates over the post-treatment periods. The red and dark blue dashed lines are the upper and lower bound of the ATT over the periods 1 and 2, and their 95% confidence regions are depicted as gray areas with dotted lines. Although we have only two post-treatment periods, we observe the following patterns. First, the pass-through rates of SSB tax on store prices seem relatively stable over time compared to the restaurant fountain drink prices. Second, given the increasing pass-through rates on restaurant drinks, especially with the 100% pass-through rate within the bound estimates in Oct (the second post-treatment period), it would be interesting to examine further whether or not there is any excessive market power exercised through the restaurant drink prices in later periods. Finally, the figure for untaxed product prices shows that the impact of SSB tax seems to be transmitted to the other drinks in Boulder city over time, but it is not statistically significant.

table[table omitted — 924 chars of source]

Table (ref) summarizes the estimation results for the three PO-GDID estimands where each set of three columns contains point estimates and 95% confidence intervals for each of the loss functions. Here, we obtain more or less the same results as the standard DID estimates, but it is important to point out that those PO-GDID estimands are derived under stronger assumptions than the robust bounds.

table[table omitted — 762 chars of source]

Lastly, Table (ref) shows the results from applying the linear projection method that we discussed in the previous example kresch2020. We can see that every positive effect that we observed has disappeared because of the increasing trends shown in Figure (ref). Hence, even the GDID bounds results (Table (ref)) are to be taken with cautiousness if we cannot rule out the existence of trends in the selection biases.

cai2016

cai2016 investigates the impact of insurance provision on tobacco production using a household-level panel dataset provided by the Rural Credit Cooperative (RCC), the main rural bank in China. The regression equation used in cai2016 is as follows:

eqnarray[eqnarray omitted — 165 chars of source]

where $i, r, t$ are household, region, and year indices, respectively, and $Y$ is the outcome variable ($area\_tob$: area of tobacco production measured in mu,\footnote{1 mu corresponds to 1/15 ha.} $tobshare$: share of tobacco production in total area of agricultural production). The covariates $X$ linearly enters the equation to be controlled for and consist of the household size, education level, and age of the household head. Note that as is common in the applied research literature, the author interprets $\alpha_3$ in Equation ((ref)) as the ATT under the standard parallel trend assumption. As we previously discussed, this model specification can be too restrictive, especially when the treatment effect is heterogeneous in the covariates $X$.

figure[figure omitted — 201 chars of source]
figure[figure omitted — 201 chars of source]

Figure (ref) and (ref) show the selection biases in the pre-treatment periods for each of the outcome variables $area\_tob$, and $tobshare$. We do not see any clear pattern for the pre-treatment periods selection biases.

table[table omitted — 859 chars of source]
table[table omitted — 867 chars of source]

Tables (ref) and (ref) show the GDID bounds estimation results where each table uses either the linear or quadratic specifications for the outcome regression model $\mu_0(X_1)$. The first four columns represent the estimated GDID bounds and their 95% confidence intervals, and the following columns show the doubly-robust estimate, the standard OLS estimate, the constructed selection bias set $\Delta_{SB_0}$, and the original point estimates from cai2016. Note that the results are not significantly different across the specifications, and we still conclude that the effect of the insurance is positive on both the area and share of tobacco at 5% significance level. Different from kresch2020, on the other hand, $\tau^{DR}$ and $\theta_{OLS}$ are close to each other, and most of the estimated GDID bounds contain the original estimates from cai2016 except for $tobshare$ under the quadratic specification. These findings suggest that the specification in Equation ((ref)) is supported by the data.

figure[figure omitted — 205 chars of source]

Figure (ref) shows the event-study type DID analysis (static treatment effects in multiple treatment periods model as in Equation (ref)) to capture the non-parametric evolution of the treatment effects over the post-treatment periods. The red and dark blue dashed lines are respectively the upper and lower bounds of the time-specific treatment effects, and their 95% confidence regions are depicted as gray areas with dotted lines. From this analysis, we observe that the initial impact of the insurance provision on the tobacco production area is relatively small that the null hypothesis cannot be rejected at 10% significance levels. But, the effect becomes substantially significant over time.

table[table omitted — 796 chars of source]

Table (ref) shows the GDID estimates obtained from the three types of the loss functions. As can be seen, the results are not significantly different across the three loss functions and appear to be qualitatively the same.

table[table omitted — 533 chars of source]

Finally, Table (ref) summarizes the GDID estimation results from the linear projection. Note that the observed trend of selection bias in pre-treatment periods (Figure (ref)) has caused higher predicted selection bias in the post-treatment period ($\widehat{SB_1} = -0.062$), and the effect on $tobshare$ is now no longer statistically not significant at 5% level.

CallawaySantAnna2021

Following CallawaySantAnna2021, we applied our framework to investigate the impact of minimum wage increases on teen employment. Specifically, we used the Quarterly Workforce Indicators (QWI) data used in dube2016 to collect the first quarter teen employment as our outcome variable.

CallawaySantAnna2021 considered 7 years of periods between 2001 and 2007 where the federal minimum wage did not change over time, and 3 different control groups of $g=2004, 2006,$ and $2007$ with states that raised their minimum wage in or right before the beginning of years 2004, 2006, and 2007, respectively. The specific timing of the raise can be found in CallawaySantAnna2021, and it should be noted that there is some heterogeneity in the size of the minimum wage increase within each group. The control group consists of states that did not raise their minimum wage during this period, and the complete classification can be found in Table (ref).

table[table omitted — 635 chars of source]

For illustration purposes, we implement our bounding approach under Assumption (ref), where we defined the information set using the pre-treatment periods before the first treatment in 2004 (i.e., $\mathcal I_0 = { 2001, 2002, 2003 }$). Hence, by estimating (ref) three times and taking the convex hull of each estimate / confidence interval, we were able to estimate the bounds in Proposition (ref) with the corresponding confidence intervals. The results are summarized in Figure (ref) for each treatment group.

figure[figure omitted — 356 chars of source]

The vertical line in each panel of Figure (ref) represents the treatment timing, and the black dots shows upper/lower bound estimates of $ATT_t[(0, \ldots,0, d_g = 0, 0, \ldots, 0)\rightarrow (0, \ldots,0, d_g = 1, 1, \ldots, 1)]$ for each $g=2004, 2006, 2007$ and $t=2004, \cdots, 2007$ as well as the corresponding 95% confidence intervals. We confirm the similar and statistically significant treatment effect trends for $g=2004$ as those found in CallawaySantAnna2021. However, our estimates are not able to reject the null hypotheses of zero treatment effects after the treatment for $g=2006$ or $g=2007$ due to more dispersed and unstable selection biases in the pre-treatment periods, resulting in larger $\Delta_{SB}$. Note however that we do not include any covariates in this illustration.

Conclusion

In this paper, we propose a new DID method that is robust to violations of parallel trends that can be captured in the pre-treatment periods. Under a weaker assumption than the standard (conditional) parallel trends assumption, we derive novel bounds for the ATT, which we call the robust generalized DID bounds. These bounds always cover the standard DID estimand. If the PT assumption holds in the pre-treatment periods, our robust generalized DID bounds collapse to a point, the standard DID estimand. To construct the bounds, we define an information set in the baseline period where no individual was treated yet. This information set helps define the set of all pre-treatment periods selection biases. We therefore assume that the post-treatment period selection bias lies within the convex hull of all pre-treatment periods selection biases. We provide a sufficient condition for this assumption. We also show how baseline covariates can help in the identification strategy.

As the information set grows, our bounds become wider and may become less relevant for the policymaker. We therefore discuss different ways to select the post-treatment period selection bias optimally by minimizing a loss function chosen by the policymaker. Doing so will yield a point estimand that may not necessarily have a clear causal interpretation but could be relevant for the policymaker's decision making process. We call this parameter a policy-oriented generalized DID.

We show how our method can be extended to the multiple treatment periods DID designs and the synthetic control method. We illustrate our proposed method through some numerical and empirical examples. In the multiple treatment periods DID framework, our approach partially identifies various causal parameters that can help reveal some dynamic effects of the treatment. In this setup, we propose a two-way fixed effects regression inference method. Currently, the information set is static in our proposed approach as it does not change over time. Making the information set dynamic in order to allow past outcomes to influence current and future outcomes could be an interesting area for future research. For example, investment which is the outcome variable in one of our applications follows a dynamic process.