EconBase
← Back to paper

Selection and parallel trends

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

96,270 characters · 27 sections · 87 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Selection and parallel trends

\thispagestyle{empty}

abstractWe study the role of selection into treatment in difference-in-differences (DiD) designs. We derive necessary and sufficient conditions for parallel trends assumptions under general classes of selection mechanisms. These conditions characterize the empirical content of parallel trends and clarify the trade-offs between assumptions about selection into treatment and restrictions on the time series properties of the potential outcomes required for DiD methods. We use the necessary and sufficient conditions to provide a selection-based decomposition of the bias of DiD and provide easy-to-implement strategies for benchmarking its components. We also provide templates for justifying DiD in applications with and without covariates. Reanalyses of the causal effect of NSW training programs and the effect of the Medicaid expansion demonstrate the usefulness of our selection-based approach to benchmarking the bias of DiD. Keywords: causal inference, conditional parallel trends, covariates, difference-in-differences, selection mechanism, time-invariant and time-varying unobservables, treatment effects JEL Codes: C21, C23

\setcounter{page}{1}

quote\dots while the new papers [in the DiD literature] clarify very well the statistical assumptions needed for estimation, effective use of these methods also requires being able to understand what the threats to these assumptions are in different contexts, and to make a plausible rhetorical argument as to why we should think the assumptions hold.\\ --- David McKenzie, World Bank Development Impact Blog mckenzie2022a

Introduction

This paper provides a new perspective on difference-in-differences (DiD) identification through the lens of how units select into treatment. Parallel trends, the identifying assumption underlying DiD, requires the change in the expected (untreated) potential outcome over time to be the same in the treatment and control group. It is thus inherently a joint restriction on selection into treatment and the time series properties of the untreated potential outcome. Whether and to what extent there is a trade-off between these two types of restrictions is not well-understood, however. In particular, there is no general framework that characterizes the implications of parallel trends under selection-based restrictions. Such a framework is crucial for researchers to exploit contextual and economic information about how units select into treatment, which is available in many applications. It would enable researchers --- in the words of David McKenzie mckenzie2022a --- “to understand what the threats to these [parallel trends] assumptions are in different contexts, and to make a plausible rhetorical argument as to why we should think the assumptions hold.” By providing researchers with a selection-based framework for exploiting contextual and economic information to assess the plausibility of parallel trends, we complement the existing statistical and visual plausibility checks in DiD analyses.

Our goal is to provide a nonparametric selection-based characterization of the parallel trends assumption. To achieve this goal, we start by providing a general necessary and sufficient condition for parallel trends to hold for all selection mechanisms in a class defined by the unobservables that determine selection, which we denote by $\omega_i$.\footnote{From a theoretical perspective, this necessary and sufficient condition is the condition for nonparametric identification using DiD for a given class of selection mechanisms.} This general result allows us to provide necessary and sufficient conditions for many empirically relevant classes of selection mechanisms.

The focus on classes of selection mechanisms defined by $\omega_i$ is motivated by the reduced-form nature of DiD methods and the myriad of empirical contexts analyzed using these methods. While contextual and economic information about selection may allow researchers to posit what determines selection into treatment, it is generally difficult to specify a particular functional form for the selection mechanism. For instance, if we consider settings where selection into treatment is at the aggregate level, such as at the county or state level as in the Medicaid expansion miller2021medicaid,baker2026difference, specifying a selection mechanism might be intractable or prohibitively complex.\footnote{This echoes the point made by abadie2021using about the difficulty of specifying selection mechanisms in synthetic control analyses of comparative case studies.} The advantage of our necessary and sufficient conditions is that researchers can use them to assess and justify parallel trends without needing to defend a particular choice of selection mechanism.\footnote{If researchers have sufficient contextual and economic information to fully specify the selection mechanism, then we recommend they rely on identification strategies that directly exploit such information rather than DiD.}

Before considering various restricted classes of selection mechanisms, we take a step back and ask: What are the necessary and sufficient conditions for parallel trends if we do not impose any restrictions on selection? The first corollary of our general necessary and sufficient condition answers this question with a “no-free-lunch” result. When researchers are not willing to restrict the selection mechanism at all, $\omega_i$ can include time-invariant and time-varying determinants of the untreated potential outcomes in the pre- and post-treatment period, $Y_{i1}(0)$ and $Y_{i2}(0)$. We show that parallel trends holds for all selection mechanisms in this class if and only if the untreated potential outcome is constant across time up to deterministic mean shifts, $Y_{i2}(0)-Y_{i1}(0)=E[Y_{i2}(0)-Y_{i1}(0)]$. This result shows that if one is not willing to restrict selection into treatment, then one needs to essentially rule out time-varying unobservables.

This “no-free-lunch” result motivates considering restricted classes of selection mechanisms, such as selection on treatment effects (“Roy-style selection”), selection on pre-treatment unobservables (“imperfect foresight”), selection on lagged outcomes, and selection on time-invariant unobservables (“selection on fixed effects”). The necessary and sufficient conditions for these classes characterize the trade-offs between restrictions on selection and restrictions on the time series properties of the untreated potential outcome and its unobservable determinants. The more restrictions researchers are willing to impose on how units select into treatment, the weaker the time series restrictions necessary and sufficient for parallel trends to hold. Conversely, in settings where the available economic and contextual knowledge is not sufficient to justify strong restrictions on selection, researchers need to impose restrictive assumptions on the time series properties of the potential outcomes to justify parallel trends.

Consider, for example, a setting with Roy-style selection where the units select into treatment if their treatment effect exceed the costs. If the units know their treatment effect and costs, then the necessary and sufficient condition for parallel trends requires the difference between the potential outcomes over time, $Y_{i2}(0)-Y_{i1}(0)$, to be mean independent of those determinants of selection. Alternatively, suppose that the units select into treatment based on pre-treatment unobservables. In this case, the necessary and sufficient condition for parallel trends is a martingale-type condition on $Y_{it}(0)-E[Y_{it}(0)]$. Finally, in the context of selection on time-invariant unobservables, parallel trends is equivalent to a time homogeneity condition on the expectation of $Y_{it}(0)-E[Y_{it}(0)]$ conditional on the time-invariant unobservables. The advantage of our general necessary and sufficient condition is that researchers can use it to derive equivalent conditions for parallel trends for other empirically relevant settings.

The necessary and sufficient conditions we provide can serve as theory-based templates allowing researchers to assess and justify parallel trends in empirical applications based on contextual and economic information about how units select into treatment. To illustrate, we specialize our general results to the standard two-way fixed effects model for $Y_{it}(0)$. The resulting conditions explicitly allow for selection on time-invariant and time-varying unobservables, thus formalizing what “quasi-random” assignment means in the context of DiD analyses.

An appealing feature of our selection-based approach is that our necessary and sufficient conditions for parallel trends can help researchers better understand the bias of DiD. We provide a selection-based approach for benchmarking the bias of DiD when the validity of these necessary and sufficient conditions is questionable. To illustrate, we consider the leading case where selection on pre-treatment unobservables is a concern. We decompose the bias of DiD into two terms. The first component captures the bias resulting from selection on post-treatment unobservables. The second component captures the bias due to deviations from the martingale condition, which is necessary and sufficient for parallel trends under imperfect foresight. Under a linear relaxation of the martingale condition, the second bias component is equal to the product of the martingale deviation and the pre-treatment difference between the treatment and control group. From a theoretical perspective, our analysis allows us to better understand the bias of DiD and the extent to which pre-treatment information is useful for assessing the magnitude of this bias. For practitioners, we provide simple approaches to benchmark and sign both bias components, allowing them to assess the robustness of DiD in empirical applications.

We illustrate the usefulness of the selection-based approach for benchmarking the bias components of DiD based on two empirical applications. First, we consider the estimation of the causal effect of the NSW training programs. This application is well-suited for our purposes because there is an experimental benchmark allowing us to estimate the bias of DiD. Without covariates, the bias of DiD relative to the experimental benchmark is large and significant. The proposed benchmarking strategy yields the same sign of the bias as the experimental estimate. The decomposition further demonstrates that the bias of DiD is very sensitive to violations of the martingale property because there is a large pre-treatment difference between the treatment and control group. Incorporating covariates into the analysis reduces the estimated bias relative to the experimental benchmark and also renders DiD more robust by reducing the pre-treatment difference between the treatment and control group. Second, we revisit the DiD analysis of the causal effect of the Medicaid expansion. The proposed benchmarking strategy suggests that the sign of the bias depends on the degree of selection on post-treatment unobservables. As in the NSW application, DiD without covariates is sensitive to violations of the martingale condition. Controlling for covariates almost fully removes pre-treatment differences between the treated and control states, rendering DiD robust to violations of the martingale condition.

Related literature

This paper contributes to several branches of the literature on causal inference using panel data. First, we contribute to the classical literature on canonical DiD setups by providing a novel selection-based perspective on the parallel trends assumptions underlying this literature.\footnote{See, e.g., Ashenfelter1978, ashenfelter1985using, heckman1985alternative, Card1990, Card1994, Meyer1995, and angrist1999empirical for early developments in the DiD literature, and Lechner2010a for a historical perspective.}

Our second contribution is to the more recent literature on DiD methods. See, e.g., de_chaisemartin_two-way-survey_2021 and roth_et_al_DiD_survey for surveys. Within this strand of the literature, our paper is most closely related to roth2021when, Arkhangelsky2021_DRTWFE, and Arkhangelsky2022_DRId, though our focus greatly differs from theirs. roth2021when discuss necessary and sufficient conditions under which parallel trends holds for all (monotonic) transformations of the untreated potential outcome. We, on the other hand, take the outcome model (and thus the specific transformation) as given and study the connection between parallel trends and selection into treatment. Arkhangelsky2021_DRTWFE and Arkhangelsky2022_DRId propose doubly robust estimation methods that leverage restrictions on outcome models and/or selection models with unconfoundedness-type restrictions; see also Athey2021. Our results complement theirs as we maintain the parallel trends assumption and discuss the types of restrictions on selection compatible with it. Moreover, our analysis shows that parallel trends is compatible with various types of selection on unobservables, unlike standard unconfoundedness assumptions imbens2004nonparametric,imbens2009recent.

Our third contribution is to the literature on sensitivity analysis, partial identification, and robust inference under violations of parallel trends. Existing work has proposed different ways to use pre-treatment information to bound the ATT manski2018how,rambachan2023a,ban2023generalized. We complement this work by providing a selection-based decomposition of the bias of DiD and simple empirical strategies for benchmarking its components. Our approach uses pre-treatment periods to learn about the deviations from the selection-based necessary and sufficient conditions for parallel trends, whereas existing approaches rely on more “reduced-form” quantities, such as pre-treatment parallel trends violations. Our selection-based approach also differs from the analysis by marx2024parallel. They derive partial identification results under monotone treatment selection assumptions on the untreated potential outcome, which they motivate using an economic model of learning with binary outcomes. By contrast, we characterize the bias of DiD in terms of deviations from the necessary and sufficient conditions for parallel trends.

Our fourth contribution is to the literature imposing explicit selection and/or outcome models to develop and compare different methods for estimating treatment effects, including DiD.\footnote{See, e.g., ashenfelter1985using, heckman1985alternative, card2005estimating, chabeferret2015analysis, blundell2009alternative, dechaisemartin2018fuzzy, verdier2020average, dechaisemartin2022not, marx2024parallel, chabetferret2025should. Most papers in this literature examine sharp DiD designs, as we do, dechaisemartin2018fuzzy and marx2024parallel also consider fuzzy DiD designs.} We contribute to this literature by providing general necessary and sufficient conditions for parallel trends, which are derived for general selection and outcome models that nest models considered in this literature. Our conditions thus clarify trade-offs between assumptions on selection and time-varying unobservables that are relevant for those models. Within this strand of the literature, our paper is most closely related to contemporaneous work by marx2024parallel, though our focus markedly differs from theirs. marx2024parallel analyze parallel trends through the lens of various examples of dynamic choice models. In doing so, they focus on explicit models of selection, and many of their examples are for binary outcomes. By contrast, we provide general necessary and sufficient conditions for parallel trends without imposing restrictions on the nature of the outcome variables or explicit models of dynamic choice. The general classes of selection mechanisms we consider nest, and can be motivated by, dynamic choice models. Unlike marx2024parallel, we also incorporate covariates into our analysis and study settings with multiple groups and periods.

Finally, a byproduct of our analysis is an explicit connection between DiD and the literature on nonseparable panel models. In Appendix (ref), we show that our sufficient conditions for parallel trends imply combinations of identifying assumptions in this literature altonji2005cross,bester2009identification, hoderlein2012nonparametric,chernozhukov2013average.

Notation

For a random vector $W_{it}$, where $i=1,\dots,n$ and $t=1,2$, we let $\dot{W}_{it}\equiv W_{it}-E[W_{it}]$ and denote its time series by $W_i\equiv (W_{i1},W_{i2})$.\footnote{We define all vectors in this paper as row vectors.} We use $F_W$ to denote the distribution of the random vector $W$ and $\mathcal{W}$ to denote its support. Let $f(z,w)$ be a function defined on $\mathcal{Z}\times\mathcal{W}$. We say that $f(z,w)$ is a trivial function of $w$ if $f(z,w)=f(z,w')=h(z)$ for all $z\in\mathcal{Z}$, $w\neq w'$, and $(w,w')\in\mathcal{W}^2$. We say that $f(z,w)$ is a symmetric function in $z$ and $w$ if $f(z,w)=f(w,z)$ for all $(z,w)\in\mathcal{Z}\times\mathcal{W}$. We use the notation $\overset{d}{=}$ to denote equality of distribution. For random variables, $X_i$, $Z_i$, and $W_i$, $Z_i|W_i,X_i\overset{d}{=}Z_i|X_i,W_i$ denotes that $F_{Z_i|W_i,X_i}(z|w,x)=F_{Z_i|X_i,W_i}(z|w,x)$ for $(z,w,x)\in\mathcal{Z}\times\mathcal{W}\times\mathcal{X}$.

Setup, selection mechanism, and examples

In the main text, we consider the classical DiD setup with two groups and two periods, where the selection decision is made at the same level as the unit of observation. We extend our results to settings with disaggregate data (e.g., on individuals) where the selection decision is made at a more aggregate level (e.g., at the state level) in Appendix (ref) as well as to DiD designs with multiple groups and multiple periods in Appendix (ref).

Let $D_{it}$ and $Y_{it}$ denote the treatment status and outcome for unit $i\in \{1,\dots,n\}$ in period $t\in \{1,2\}$. Here the index $i$ refers to the unit making the decision to select into treatment. This could be an individual or a more aggregate administrative unit, such as a county or state. The treatment group ($G_i=1$) selects the treatment path $D_i=(0,1)$; the control group ($G_i=0$) selects $D_i=(0,0)$. The potential outcomes with and without the treatment are $Y_{it}(1)$ and $Y_{it}(0)$, respectively.\footnote{To focus attention on the role of the parallel trends assumption, we assume that there are no anticipatory effects. This is a standard assumption in the DiD literature. See, for example, roth_et_al_DiD_survey for a discussion.} We abstract from covariates for now to focus on the issues arising from selection on time-invariant and time-varying unobservables. We discuss the additional implications of including covariates in Appendix (ref).

We consider the standard parallel trends assumption. Throughout the paper, we assume that all relevant moments exist and $\{Y_{i1}(0),Y_{i2}(0),G_i\}$ is i.i.d. across $i$.

customass{PT} The (unconditional) parallel trends assumption holds: $$E[Y_{i2}(0)-Y_{i1}(0)|G_i=1]=E[Y_{i2}(0)-Y_{i1}(0)|G_i=0].$$

Under Assumption (ref), the average treatment effect on the treated group in period $t=2$, $\operatorname{ATT}\equiv E[Y_{i2}(1)-Y_{i2}(0)|G_i=1]$, is identified from the “difference-in-differences” as follows: $$ \operatorname{ATT}=E[Y_{i2}-Y_{i1}|G_i=1]-E[Y_{i2}-Y_{i1}|G_i=0]\equiv \operatorname{DiD}. $$

We work with a general nonseparable model for $Y_{it}(0)$,

equation[equation omitted — 123 chars of source]

where $\alpha_i$, $\varepsilon_{i1}$, and $\varepsilon_{i2}$ are finite-dimensional vector-valued random variables, and $\xi_t(\cdot)$ is an unrestricted time-varying function. The outcome model (ref), while not imposing any restrictions on $Y_{it}(0)$, allows us to distinguish between time-invariant and time-varying unobservables. This is necessary to define selection mechanisms that can directly depend on these unobservables. If, instead, we were to work directly with potential outcomes, this would rule out important examples of selection mechanisms, such as selection on fixed effects ashenfelter1985using.

The determinants of selection into treatment vary widely across DiD applications. In some applications, the unit $i$ making the selection decision is a state or county, whereas in other settings it is an individual economic agent, such as an individual, a household, or a firm. We therefore introduce a general selection mechanism that accommodates many different types of selection. To motivate this general selection mechanism, it is helpful to consider examples of selection mechanisms relevant for our setting.

example[Selection on untreated outcomes] Selection on untreated potential outcomes goes back to at least the seminal work of ashenfelter1985using in the context of individuals selecting into job training programs. It remains relevant in DiD applications where selection decisions are made by more aggregate units, which are prevalent in contemporary empirical research. For example, in a survey of governors on the reasons for expanding Medicaid, sommers2013us find that the health outcomes in their states were one of the factors affecting their support for Medicaid expansion. ashenfelter1985using studied the case where individuals select into the training programs if their pre-treatment earnings $Y_{i1}(0)$ fall below a fixed threshold, so that $G_i=1\left\{Y_{i1}(0)\le c\right\}$.\footnote{ashenfelter1985using consider a more general setting with multiple pre-treatment periods, where selection depends on earnings in the $k$th pre-treatment period.} In the context of Medicaid expansion, policymakers in a state might expand Medicaid if the health outcomes of their constituents (before the expansion) fall below a certain threshold. More generally, let $\omega_i$ denote the information set available to the units when deciding whether to select into the treatment and consider the following mechanism, \begin{align} G_i=1\left\{E[Y_{i1}(0)+\beta Y_{i2}(0)|\omega_i]\le E[\kappa_{i2}|\omega_i]\right\}, \end{align} where $\beta\in [0,1]$ is a discount factor and $\kappa_{i2}$ is a unit-specific random threshold. This example demonstrates the importance of allowing for additional unobservables in addition to the determinants of $Y_{it}(0)$, $(\alpha_i,\varepsilon_{i1},\varepsilon_{i2})$. \qed
example[Selection on treatment effects (Roy-style selection)] Let $\tau_{i2}$ denote the treatment effect in period $t=2$, $\tau_{i2}=Y_{i2}(1)-Y_{i2}(0)$. Suppose that units select into the treatment if the expected gains from treatment given the information set $\omega_i$, $E[\tau_{i2}|\omega_i]$, exceed the expected cost of treatment, $E[\kappa_{i2}|\omega_i]$, $G_i=1\{ E[\tau_{i2}|\omega_i]\ge E[ \kappa_{i2}|\omega_i]\}.$ In the context of the job training setting, units might decide to participate in the program if the expected gain in their earnings outweighs the expected costs, whereas in the Medicaid setting, policymakers might support a Medicaid expansion if the expected improvements in the health outcomes of their constituents outweigh the expected costs. Indeed, these expected improvements were among the factors mentioned in the survey in sommers2013us. \qed
example[Selection on fixed effects] DiD methods have traditionally been motivated using two-way fixed effects models. Fixed effects assumptions allow for unrestricted dependence between time-invariant unobservables and the regressors, thereby implicitly allowing for selection on time-invariant unobservables.\footnote{See, e.g., chamberlain1984panel,arellano2003panel,evdokimov2010identification,wooldridge2010econometric,hoderlein2012nonparametric,chernozhukov2013average.} A simple example is $G_i=1\{\alpha_i\le c\}$, where $\alpha_i$ is a scalar, which corresponds to the selection mechanism on p.650 in ashenfelter1985using. In the training program example, this mechanism captures settings where individuals select into the program if the permanent earnings component $\alpha_i$ falls below a threshold $c$, which is “based on potential trainees' discount rates, time horizons, and tastes for training” ashenfelter1985using. In the context of aggregate treatments, such as Medicaid expansion, policymakers might decide to expand Medicaid based on their political affiliations or state-specific characteristics, which are plausibly time-invariant. \qed

In the previous examples, all agents make their selection decision in the same way and based on the same information set $\omega_i$. This is likely an oversimplification in many contexts where DiD is used, especially in aggregate selection settings. In the following example, we consider a setting with heterogeneous units whose selection decisions depend on different unobservables.

example[Selection with heterogeneous units] For simplicity, we consider a setting with two types of units, noting that the example can be easily generalized to settings with more than two types. Let $\mu_i\in \{0,1\}$ be an indicator for the unit's type. If $\mu_i=1$, unit $i$ adopts the treatment if the expected benefits outweigh the expected costs given $\omega_i^1$; if $\mu_i=0$, unit $i$ selects into treatment if the expected discounted sum of untreated outcomes given $\omega_i^0$ falls below a certain threshold, so that \begin{equation*} G_i= 1\{E[Y_{i2}(1)-Y_{i2}(0)|\omega_i^{1}]\geq E[\kappa_{i2}|\omega_i^1]\}^{\mu_i}1\{E[Y_{i1}(0)+\beta Y_{i2}(0)|\omega_i^{0}]\leq E[\kappa_{i2}|\omega_i^0]\}^{1-\mu_i}. \end{equation*} Here, $G_i$ depends on the unobserved type $\mu_i$ as well as the information sets used by both types, $\omega_i^0$ and $\omega_i^1$. \qed

Motivated by these examples, we consider the following general selection mechanism,

eqnarray[eqnarray omitted — 88 chars of source]

We allow $\omega_i$ to be a function or a subvector of $(\alpha_i,\varepsilon_{i1},\varepsilon_{i2},\mu_i,\eta_{i1},\eta_{i2})$ and can thereby accommodate the above examples as special cases. Moreover, we can accommodate many other economic models of selection heckman1985alternative,chabeferret2015analysis, marx2024parallel. The additional unobservables $(\mu_i,\eta_{i1},\eta_{i2})$ capture any other determinants of selection that may be correlated with $Y_{it}(0)$, such as individual-specific thresholds in Example (ref) or treatment effects and costs in Example (ref).

The scalar unobservable, $\nu_i$, omitted for simplicity in the above examples, captures determinants of selection that are independent of $(\alpha_i,\varepsilon_{i1},\varepsilon_{i2},\mu_i,\eta_{i1},\eta_{i2})$ and allows us to accommodate the important special case of random assignment. We impose the following assumption.

customass{SEL} $P(\nu_i>c)\in(0,1)$ for some $c\in\mathbb{R}$ and $\nu_i\perp \!\!\! \perp (\alpha_i,\varepsilon_{i1},\varepsilon_{i2},\mu_i,\eta_{i1},\eta_{i2})$.

Note that since $G_i=D_{i2}$, $g(\cdot)$ can be equivalently viewed as the selection mechanism for $D_{i2}$. Let $\mathcal{G}_{\omega}$ denote the class of all selection mechanisms $g(\cdot)$ mapping from the support of $(\omega_i,\nu_i)$ to $\{0,1\}$. We index the class of selection mechanisms with $\omega_i$ since it could be correlated with the potential outcomes. Since $\omega_i$ can be interpreted as the units' information sets, the categorization of selection mechanisms based on $\omega_i$ is motivated by the usefulness of information sets for analyzing causal inference methods heckman2007econometric.

remark[Parallel trends and functional form] Throughout this paper, we take the functional form of the outcome as given. We thereby abstract from the issues arising from the sensitivity of DiD to functional form specification roth2021when. \qed

Necessary and sufficient conditions for parallel trends

In this section, we provide necessary and sufficient conditions for parallel trends. Section (ref) provides a general necessary and sufficient condition for parallel trends. We then apply this general condition in Sections (ref)--(ref) to derive necessary and sufficient conditions for various practically relevant scenarios by specifying what units select on.

A general necessary and sufficient condition for parallel trends

Here, we provide a general result that allows for specifying necessary and sufficient conditions for parallel trends for any class of selection mechanisms $\mathcal{G}_\omega$. Consistent with the nonparametric treatment of selection mechanisms in DiD analyses, we provide necessary and sufficient conditions for Assumption (ref) to hold for all $g\in \mathcal{G}_\omega$.\footnote{These conditions imply that Assumption (ref) holds for all $g\in \mathcal{G}_\omega$, which is in the spirit of standard notions of nonparametric identification hansen2022econometrics.}

Considering necessary and sufficient conditions for Assumption (ref) for all $g\in\mathcal{G}_\omega$ ensures that the invocation of Assumption (ref) is robust to the choice of selection mechanism within this class.\footnote{In Section (ref), we further demonstrate that, even in the case where a researcher knows the underlying selection mechanism, the necessary and sufficient condition for all $g\in\mathcal{G}_\omega$ rules out cases where Assumption (ref) holds due to peculiar combinations of parameters of the data-generating process (see Figure (ref)).} While robustness to the exact specification of the selection mechanism is important in all DiD applications, it is especially relevant in settings with aggregate selection (e.g., at the county or state level). When multiple entities or actors are involved, or when the decision is based on aggregating heterogeneous individual preferences, the selection mechanisms are often complex and difficult to model. Consider again the Medicaid expansion example. The survey of governors in sommers2013us shows that there are many different factors affecting the governors' views on expanding Medicaid, ranging from fiscal considerations to the potential health outcomes of their constituents and the potential benefits of the expansion in terms of health insurance coverage and health outcomes.

The following theorem provides a necessary and sufficient condition for any class of selection mechanisms $\mathcal{G}_\omega$. We focus on non-degenerate DiD designs, that is, designs with $P(G_i=1)\in (0,1)$. In the following, we will use the notation $\dot{Y}_{it}(0)= Y_{it}(0)-E[Y_{it}(0)]$.

theorem[Necessary and sufficient condition for (ref) to hold for all $g\in \mathcal{G}_\omega$] Suppose that Assumption (ref) holds. Suppose further that either $P(E[\dot{Y}_{i2}(0)-\dot{Y}_{i1}(0)|\omega_i]>0)<1$ or $P(E[\dot{Y}_{i2}(0)-\dot{Y}_{i1}(0)|\omega_i]<0)<1$. Then, Assumption (ref) holds for all $g\in \mathcal{G}_{\omega}$ satisfying $P(G_i=1)\in (0,1)$ if and only if $E[\dot{Y}_{i2}(0)-\dot{Y}_{i1}(0)|\omega_i]=0\text{ a.s.}$

The “if” direction of Theorem (ref) follows by the law of iterated expectation. The proof of the “only if” direction is constructive: we provide an explicit “least favorable” selection mechanism that yields the necessary condition. For the case where $P(E[\dot{Y}_{i2}(0)-\dot{Y}_{i1}(0)|\omega_i]>0)<1$, this selection mechanism takes the following form,

equation[equation omitted — 115 chars of source]

Assumption (ref) (together with $P(E[\dot{Y}_{i2}(0)-\dot{Y}_{i1}(0)|\omega_i]>0)<1$) ensures that this selection mechanism is non-degenerate when the necessary and sufficient condition holds.\footnote{Under the necessary and sufficient condition, $1\{E[\dot{Y}_{i2}(0)-\dot{Y}_{i1}(0)|\omega_i]\le 0\}=1$. In this case, the first component of the selection mechanism $1\{\nu_i>c\}$ plays the role of a tie-breaker, ensuring that mechanism (ref) is non-degenerate.}

Stating the necessary and sufficient condition in Theorem (ref) in terms of what the units select on and their information sets, $\omega_i$, rather than explicit selection mechanisms, has a key advantage: While it may be possible to determine (based on contextual and economic knowledge) what information units select on, there is likely uncertainty about the exact form of the selection mechanism, especially in settings with aggregate selection. The necessary and sufficient condition in Theorem (ref) therefore relieves the researcher from the need to impose additional structure on the selection mechanism that is not justified by their context.

Next, we apply Theorem (ref) to various practically relevant classes of selection mechanisms.

What if we impose no restrictions on selection?

Unlike other causal inference methods, DiD does not explicitly restrict selection into treatment. This begs the question: What if researchers are indeed not willing to impose any assumptions on selection so that parallel trends needs to hold for all selection mechanisms? To answer this question, we apply Theorem (ref) with $\omega_i$ including all unobservables that could enter the selection mechanism in (ref).

corollary[No restrictions on selection] Under the assumptions of Theorem (ref) with $\omega_i=(\alpha_i,\varepsilon_{i1},\varepsilon_{i2},\mu_i,\eta_{i1},\eta_{i2})$, Assumption (ref) holds for all $g\in\mathcal{G}_\omega$ satisfying $P(G_i=1)\in(0,1)$ if and only if $\dot{Y}_{i1}(0)=\dot{Y}_{i2}(0)$ a.s. The result continues to hold if $\omega_i=(\alpha_i,\varepsilon_{i1},\varepsilon_{i2})$.

To interpret the necessary and sufficient condition in Corollary (ref), it is helpful to rewrite it as $$ Y_{i2}(0)-Y_{i1}(0)=E[Y_{i2}(0)-Y_{i1}(0)]. $$ This shows that absent any restrictions on selection, parallel trends implies that the potential outcomes are constant over time, except for common mean shifts. This essentially rules out time-varying unobservables. To see this, consider the following standard two-way model

equation[equation omitted — 115 chars of source]

where $\lambda_t$ is a nonstochastic time trend. The necessary and sufficient condition specialized to this separable outcome model is $\varepsilon_{i1}=\varepsilon_{i2}$, implying that $\varepsilon_{it}$ is time-invariant.

Given that the necessary and sufficient condition for the unrestricted class of selection mechanisms is implausible in most applications, we next consider restricted classes of selection mechanisms.

Necessary and sufficient conditions for restricted classes of mechanisms

DiD applications differ substantially in terms of what determines selection into treatment. Corollary (ref) shows that restrictions on selection are unavoidable in realistic settings. In the following, we consider various restrictions on selection mechanisms that are practically relevant and well-established in the literature. The list of restrictions we consider is not exhaustive. The advantage of Theorem (ref) is that researchers can specialize the necessary and sufficient condition for the class of selection mechanisms relevant for their application.

Imperfect foresight

Selection on pre-treatment unobservables is likely in many applications. An example is when units make their selection decision based on expected future potential outcomes and costs, while only having access to pre-treatment information (e.g., $\omega_i=(\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1})$ in Examples (ref) and (ref)). Another example is when units select into treatment in response to negative (or positive) pre-treatment shocks or their pre-treatment outcome falling below (or above) a specific threshold (e.g., Example (ref) with $\beta=0$ and $\omega_i=(\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1})$) and more broadly when there is feedback bonhomme2025back,chamberlain2022feedback. Finally, imperfect foresight is relevant when individuals are myopic.

We first consider the case where selection depends on time-invariant and all pre-treatment unobservables, so that $\omega_i=(\alpha_i,\varepsilon_{i1},\nu_i,\eta_{i1})$. For this case, Theorem (ref) implies the following corollary.\footnote{We are grateful to Eric Mbakop for encouraging us to pursue necessary and sufficient conditions instead of necessary conditions only under imperfect foresight (and selection on fixed effects, discussed below).}

corollary[Imperfect foresight: Case 1] Under the assumptions of Theorem (ref) with $\omega_i=(\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1})$, Assumption (ref) holds for all $g\in \mathcal{G}_{\omega}$ satisfying $P(G_i=1)\in (0,1)$ if and only if $E[\dot{Y}_{i2}(0)|\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1}]=\dot{Y}_{i1}(0)$ a.s.

The necessary and sufficient condition in Corollary (ref) is a martingale-type condition on the untreated potential outcomes with respect to the unobservables that determine selection in this case. To build further intuition and to compare this result to Corollary (ref), note that the necessary and sufficient condition can be equivalently written as $$ Y_{i2}(0)-Y_{i1}(0)=E[Y_{i2}(0)-Y_{i1}(0)]+\zeta_{i2}, \quad E[\zeta_{i2}|\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1}]=0. $$ That is, the condition in Corollary (ref) allows the untreated potential outcomes to vary over time beyond deterministic mean shifts but requires the stochastic component of the change over time, $\zeta_{i2}$, to be mean-independent of the pre-treatment unobservables $(\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1})$.

In the two-way model (ref), the condition in Corollary (ref) becomes $E[\varepsilon_{i2}|\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1}]=\varepsilon_{i1}$, a martingale-type property that implies that $\varepsilon_{i2}-\varepsilon_{i1}+\zeta_{i2}$, where $\zeta_{i2}$ is an innovation satisfying $E[\zeta_{i2}|\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1}]=0$. This necessary and sufficient condition relates to the consistency of the first-difference estimator under violations of strict exogeneity when the idiosyncratic shocks follow a unit root.\footnote{We thank St\'ephane Bonhomme for pointing out this connection.}

The martingale-type condition in Corollary (ref) arises because the units select on the determinants of $Y_{i1}(0)$, $(\alpha_i,\varepsilon_{i1})$. As a result, the condition in Theorem (ref) with $\omega_i=(\alpha_i,\varepsilon_{i1},\nu_i,\eta_{i1})$, $ E[\dot{Y}_{i2}(0)-\dot{Y}_{i1}(0)|\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1}]=0$, simplifies to $E[\dot{Y}_{i2}(0)|\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1}]=\dot{Y}_{i1}(0)$, as stated in Corollary (ref). If selection is based on pre-treatment unobservables that do not include the determinants of $Y_{i1}(0)$, the martingale condition does not arise, as the next corollary shows.

corollary[Imperfect foresight: Case 2] Under the assumptions of Theorem (ref) with $\omega_i=(\mu_i,\eta_{i1})$, Assumption (ref) holds for all $g\in \mathcal{G}_{\omega}$ satisfying $P(G_i=1)\in (0,1)$ if and only if $E[\dot{Y}_{i2}(0)-\dot{Y}_{i1}(0)|\mu_i,\eta_{i1}]=0$ a.s.

The necessary and sufficient condition in Corollary (ref) resembles but is different from the standard definition of a martingale-difference condition on the difference $\Delta \dot{Y}_{i2}(0)=\dot{Y}_{i2}(0)-\dot{Y}_{i1}(0)$ hamilton1994time since the conditioning set does not include lagged values of $\Delta \dot{Y}_{i2}(0)$. It can therefore be consistent with a wider class of time series processes than the condition in Corollary (ref), which implies that $\dot{Y}_{it}(0)$ is a martingale.

Roy-style selection

Here, we consider settings with Roy-style selection, as in Example (ref). We first consider the case where the units know their treatment effect $\tau_{i2}=Y_{i2}(1)-Y_{i2}(0)$ and costs $\kappa_{i2}$ in period $t=2$. For this case, Theorem (ref) implies the following corollary.

corollary[Roy-style selection] Under the conditions of Theorem (ref) with $\omega_i=(\tau_{i2},\kappa_{i2})$, Assumption (ref) holds for all $g\in \mathcal{G}_{\omega}$ satisfying $P(G_i=1)\in (0,1)$ if and only if $E[\dot{Y}_{i2}(0)-\dot{Y}_{i1}(0)|\tau_{i2},\kappa_{i2}]=0$ a.s.

Rewriting the necessary and sufficient condition as $E[\dot{Y}_{i2}(0)|\tau_{i2},\kappa_{i2}]=E[\dot{Y}_{i1}(0)|\tau_{i2},\kappa_{i2}]$ demonstrates that it requires the conditional expectation of the demeaned untreated potential outcome given the treatment effects and costs to be equal across time. The condition would hold immediately if $(\tau_{i2},\kappa_{i2})$ were independent of the untreated potential outcomes. However, this is an arguably unrealistic restriction in many applications.

In the two-way model (ref), the necessary and sufficient condition in Corollary (ref) simplifies to $E[\varepsilon_{i2}|\tau_{i2},\kappa_{i2}]=E[\varepsilon_{i1}|\tau_{i2},\kappa_{i2}].$ The condition would be clearly violated if $\tau_{i2}$ is a monotonic transformation of $\varepsilon_{i2}$. The condition is more plausible if instead $\tau_{i2}$ and $\kappa_{i2}$ were determined by time-invariant factors. This discussion demonstrates that under Roy-style selection, parallel trends implies restrictions on treatment effect heterogeneity.

In some applications, assuming that the units know their treatment effects and costs in $t=2$ might not be plausible. Suppose instead that selection is based on expected treatment effects and costs conditional on all the available pre-treatment information, $E[\tau_{i2}|\omega_i]$ and $E[\kappa_{i2}|\omega_i]$, respectively, where $\omega_i=(\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1})$. For this case, the result in Corollary (ref) applies, and the necessary and sufficient condition for parallel trends is the martingale-type condition, $E[\dot{Y}_{i2}(0)|\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1}]=\dot{Y}_{i1}(0)$. If, instead, the expectations are conditional on the time-invariant unobservables $(\alpha_i,\mu_i)$, then the result in Corollary (ref) below applies. More generally, Theorem (ref) allows for considering many other variants of Roy-style selection.

Selection on fixed effects

Here, we consider the classical case of selection on fixed effects. Selection on fixed effects is plausible, for example, if the units' information sets only contain the time-invariant unobservables (in addition to $\nu_i$), so that $\omega_i=(\alpha_i,\mu_i)$, or if selection is directly based on fixed effects, as in Example (ref).

The following corollary provides the necessary and sufficient condition under selection on fixed effects.

corollary[Selection on fixed effects] Under the conditions of Theorem (ref) with $\omega_i=(\alpha_i,\mu_i)$, Assumption (ref) holds for all $g\in \mathcal{G}_{\omega}$ satisfying $P(G_i=1)\in (0,1)$ if and only if $E[\dot{Y}_{i2}(0)-\dot{Y}_{i1}(0)|\alpha_i,\mu_i]=0$ a.s.

Corollary (ref) shows that the necessary and sufficient condition for parallel trends under selection on fixed effects is a time-homogeneity restriction on the conditional mean of the untreated potential outcome. To interpret this necessary condition, note that it implies that $$ Y_{i2}(0)-Y_{i1}(0)=E[Y_{i2}(0)-Y_{i1}(0)]+\zeta_{i2}, \quad E[\zeta_{i2}|\alpha_i,\mu_i]=0. $$ This shows that Corollary (ref) implies a weaker mean-independence condition on the stochastic trend component $\zeta_{i2}$ than, for example, Corollary (ref), thus highlighting a trade-off between restrictions on selection and the evolution of the untreated potential outcomes over time.

In the context of the two-way model (ref), the necessary condition in Corollary (ref) simplifies to $E[\varepsilon_{i1}|\alpha_i,\mu_i]=E[\varepsilon_{i2}|\alpha_i,\mu_i]$, a time-homogeneity assumption on the conditional mean of the idiosyncratic shocks.\footnote{The time-homogeneity condition in Corollary (ref) relates to the strict exogeneity condition in fixed effects models. Suppose that $G_i=g(\alpha_i)$, then the strict exogeneity assumption $E[\varepsilon_{it}|G_i,\alpha_i]=0$ implies that $E[\varepsilon_{it}|\alpha_i]=0$, which in turn implies the time homogeneity condition in Corollary (ref) with $\omega_i=\alpha_i$.} While it is not surprising that the condition $E[\varepsilon_{i1}|\alpha_i,\mu_i]=E[\varepsilon_{i2}|\alpha_i,\mu_i]$ is sufficient for parallel trends if $G_i=g(\alpha_i,\mu_i,\nu_i)$, Corollary (ref) demonstrates that this condition is in fact necessary for parallel trends under model (ref).

Selection on lagged outcomes (unconfoundedness)

A popular alternative to the parallel trends assumption is to assume that selection is based on lagged dependent variables angrist2009mostly,ding2019bracketing, $Y_{i2}(0)\perp \!\!\! \perp G_i\mid Y_{i1}(0)$, which we will refer to as unconfoundedness. Here we apply Theorem (ref) to characterize the parallel trends assumption under unconfoundedness and shed light on the connection between these popular assumptions. Setting $\omega_i=Y_{i1}(0)$ in Theorem (ref), so that $G_i$ satisfies unconfoundedness, provides such a characterization.

corollary[Unconfoundedness] Under the assumptions of Theorem (ref) with $\omega_i=Y_{i1}(0)$, Assumption (ref) holds for all $g\in \mathcal{G}_{\omega}$ satisfying $P(G_i=1)\in (0,1)$ if and only if $E[\dot{Y}_{i2}(0)|Y_{i1}(0)]=\dot{Y}_{i1}(0)$ a.s.

Researchers imposing unconfoundedness typically do not impose additional explicit assumptions on the exact form of the selection mechanism. This provides an additional motivation for focusing on nonparametric conditions and deriving necessary and sufficient conditions for Assumption (ref) holding for all $g\in \mathcal{G}_{\omega}$.

Corollary (ref) shows that parallel trends is equivalent to the demeaned potential outcomes $\dot{Y}_{it}(0)$ satisfying a martingale property under unconfoundedness. Written in terms of original outcomes $Y_{it}(0)$, the condition becomes $ E[Y_{i2}(0)|Y_{i1}(0)]=Y_{i1}(0)+E[Y_{i2}(0)-Y_{i1}(0)] $, which holds, for example, if $Y_{it}(0)$ is a random walk with drift.

It is interesting to relate the result in Corollary (ref) to results in the existing literature.

remark[Connection to ding2019bracketing] Here, we connect the analysis to ding2019bracketing angrist2009mostly. ding2019bracketing assume that \begin{equation} E[Y_{i2}| Y_{i1},G_i]=\theta_1+\theta_2 Y_{i1}+\tau G_i. \end{equation} Proposition 1 in ding2019bracketing, written in terms of population coefficients, implies that the estimand under (ref) is \begin{equation} E[Y_{i2}|G_i=1]-E[Y_{i2}|G_i=1]-\theta_2(E[Y_{i1}|G_i=1]-E[Y_{i1}|G_i=0]) \end{equation} This estimand is equivalent to $\operatorname{DiD}$ if and only if $\theta_2=1$.\footnote{Note that main bracketing result, Theorem 1 in ding2019bracketing, does not apply in this case since Condition 1 (stationarity) is violated.} To relate the result in ding2019bracketing to Corollary (ref), note that under unconfoundedness, (ref) implies that $ E[Y_{i2}(0)| Y_{i1}(0)]=\theta_1+\theta_2 Y_{i1}(0). $ Corollary (ref) implies that $\theta_2=1$ is necessary and sufficient for Assumption (ref) to hold. Since $\operatorname{DiD}=\operatorname{ATT}$ under Assumption (ref), the result in ding2019bracketing is consistent with the necessary and sufficient condition in Corollary (ref) under the linearity assumption (ref). \qed

Other selection mechanisms and trade-offs

The previous subsections illustrate the implications of Theorem (ref) for various empirically relevant classes of selection mechanisms. Importantly, the result in Theorem (ref) is very general and allows us to characterize the empirical content of parallel trends for many other relevant classes of selection mechanisms, including many of the existing selection models discussed in the literature and reviewed in Section (ref). Applying Theorem (ref) only requires specifying $\omega_i$, that is, what information the units select on. In practice, the specification of $\omega_i$ should be guided by contextual and economic knowledge.

Varying $\omega_i$ allows us to characterize trade-offs between restrictions on selection, encoded in $\omega_i$, and restrictions on the time series properties of the untreated potential outcomes. The richer the information that the units select on, the more restrictive the time series restriction required for parallel trends to hold. The time series restrictions are particularly strong if selection is based on the unobservable determinants of $Y_{it}(0)$, that is, if $\omega_i$ includes or depends on $(\alpha_i,\varepsilon_{i1},\varepsilon_{i2})$.

The trade-offs between assumptions on selection and the time series properties of the untreated potential outcomes are particularly easy to see under explicit models for the untreated potential outcomes. To illustrate, suppose that $Y_{it}(0)$ is given by model (ref). Then, our necessary and sufficient conditions imply that parallel trends holds, for example, in the following scenarios:\footnote{Theorem (ref) provides a general framework for deriving sufficient conditions, depending on what unit select on. However, in some applications, researchers might be interested in imposing other types of assumptions on selection. In Appendix (ref), we provide a sufficient condition based on symmetry of the selection mechanism.}

enumerate[(a)]\setlength\itemsep{0pt} • Imperfect foresight (case 1): (i) $\omega_i=(\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1})$ and (ii) $E[\varepsilon_{i2}|\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1}]=\varepsilon_{i1}$ • Imperfect foresight (case 2): (i) $\omega_i=(\mu_i,\eta_{i1})$ and (ii) $E[\varepsilon_{i2}-\varepsilon_{i1}|\mu_i,\eta_{i1}]=0$ • Roy-style selection: (i) $\omega_i=(\tau_{i2},\kappa_{i2})$ and (ii) $E[\varepsilon_{i2}|\tau_{i2},\kappa_{i2}]=E[\varepsilon_{i1}|\tau_{i2},\kappa_{i2}]$ • Selection on fixed effects: (i) $\omega_i=(\alpha_i,\nu_i)$ and (ii) $E[\varepsilon_{i2}|\alpha_i,\nu_i]=E[\varepsilon_{i1}|\alpha_i,\nu_i]$

The conditions (a)--(d) provide practitioners with explicit theory-based templates for assessing and justifying parallel trends assumptions and can be used in conjunction with the selection mechanisms in Examples (ref), (ref), (ref), and (ref), or other selection mechanisms in the literature. These conditions allow researchers to provide, in the words of mckenzie2022a, “plausible rhetorical arguments as to why we should think the [parallel trends] assumptions hold.”

The relevance of parallel trends for all $g\in\mathcal{G}_{\omega}$ for empirical practice

Theorem (ref) and Corollaries (ref)--(ref) present necessary and sufficient conditions for parallel trends to hold for all $g\in\mathcal{G}_{\omega}$ because we are interested in nonparametric conditions, consistent with the nonparametric (model-agnostic) treatment of selection mechanisms in the DiD analyses. Here we elaborate on the practical relevance of focusing on parallel trends for all $g\in\mathcal{G}_{\omega}$. This is a crucial question, as it might not be obvious why practitioners should consider parallel trends for all $g\in\mathcal{G}_{\omega}$, when it is clearly stronger than parallel trends for a specific $g\in\mathcal{G}_{\omega}$. To keep our discussion concrete, we focus on the case of imperfect foresight where the units select on pre-treatment unobservables, so that $\omega_i=(\alpha_i,\varepsilon_{i1},\mu_i,\eta_{i1})$, but the arguments we make are relevant for any class of selection mechanisms.

First, in applications where contextual or economic knowledge suggests that units select on pre-treatment unobservables, there is likely uncertainty about the exact form of the selection mechanism. For instance, practitioners might not want to take a stance on (i) how expectations are formed (e.g., subjective expectations may be different from conditional expectations) or (ii) whether selection is based on the discounted sum of expected untreated outcomes (Example (ref)), on expected gains (Example (ref)), or other quantities motivated by economic models of selection. As discussed above, the uncertainty about the exact form of the selection mechanism can be particularly pronounced in settings with aggregate selection.

Second, even if one is certain about the exact parametric form of the units' selection mechanism, the exact distribution of unobservables, and how expectations are formed, one would typically not want parallel trends to depend on specific parameter choices. To illustrate, consider Example (ref) with $\omega_i=(\alpha_i,\varepsilon_{i1})$. Figure (ref) plots the parallel trends violation, $E[Y_{i2}(0)-Y_{i1}(0)|G_i=1]-E[Y_{i2}(0)-Y_{i1}(0)|G_i=0]$, for different discount factors $\beta$ against $Cov(\varepsilon_{i1},\varepsilon_{i2})$ for two different outcome models: (a) a separable model with autocorrelated shocks, (b) an autoregressive model with a drift. For both models, regardless of $\beta$, the parallel trends violation is exactly zero when $Cov(\varepsilon_{i1},\varepsilon_{i2})=1$, which corresponds to the martingale condition (Panels (a) and (b) in Figure (ref)). For the separable model, parallel trends additionally holds for very specific combinations of $\beta$ and $Cov(\varepsilon_{i1},\varepsilon_{i2})$ (Panel (a) in Figure (ref)). These additional instances of parallel trends, however, require a researcher to not only be willing to choose a parametric distribution, but also to rely on very particular combinations of the discount factor $\beta$ and $Cov(\varepsilon_{i1},\varepsilon_{i2})$ (in addition to a specific selection and outcome model). By focusing on parallel trends for all $g\in \mathcal{G}_{\omega}$, we rule out these additional cases and focus on “robust” instances of parallel trends.

figure[figure omitted — 2,143 chars of source]

Extensions

Here, we summarize three main extensions. We refer to the corresponding appendices for details.

\noindentDisaggregate data and aggregate decisions. In many DiD applications, the selection decisions are made at the aggregate level (e.g., at the county or state level), while outcome data are available at the disaggregate level (e.g., at the individual or firm level). In Appendix (ref), we consider a sharp DiD design with $n_s$ individuals, indexed by $i=1,\dots,n_s$, belonging to aggregate unit $s$. Let $Y_{st}(0)$ denote the untreated potential outcome for aggregate unit $s$ in period $t$ (e.g., the average of the disaggregate outcomes $Y_{ist}(0)$, $Y_{st}(0)=n_s^{-1}\sum_{i=1}^{n_s}Y_{ist}(0)$) and $G_s$ the selection decision of unit $s$. The necessary and sufficient conditions in this section directly apply to this setting by replacing $i$ with $s$ and interpreting the unobservables and potential outcomes as aggregate quantities. That said, being explicit about the aggregation can help “microfound” restrictions on selection, as we discuss in Appendix (ref).

\noindentMultiple periods and groups. In Appendix (ref), we extend our results to DiD designs with multiple periods and multiple groups.\footnote{Our setup and notation build on callaway_SantAnna_2021, sun_abraham_2021, and roth_et_al_DiD_survey.} Specifically, we consider a staggered adoption setting with $T$ periods, where no units are treated at $t=1$ and some units remain untreated at $t=T$. Appendix (ref) demonstrates that the necessary and sufficient condition in Theorem (ref) extends naturally to this setting, and based on it, it is straightforward to extend our theoretical results to DiD settings with multiple periods and staggered adoption.

\noindentCovariates. In many applications, parallel trends may only be plausible conditional on covariates heckman1997matching,abadie2005DiD,santanna2020drdid, callaway_SantAnna_2021. Therefore, we study the role of covariates through the lens of selection into treatment in Appendix (ref). We explicitly allow for a vector of both time-invariant and time-varying covariates, $X_{it}$, assuming that $X_{it}$ is not affected by the treatment. In Appendix (ref), we show that the necessary and sufficient conditions for conditional parallel trends imply separability requirements on how the covariates can enter the outcome model. Appendix (ref) provides selection-based templates for justifying conditional parallel trends assumptions for separable models. In Appendix (ref), we propose a weaker conditional parallel trends assumption that accommodates a rich class of nonseparable models and provide sufficient conditions for this assumption.

Selection-based bias decomposition with an application to imperfect foresight

The necessary and sufficient conditions in Section (ref) demonstrate that if we allow for selection on time-varying shocks and in particular on the determinants of the untreated potential outcomes, parallel trends implies strong restrictions on the time series properties of these outcomes. Here we analyze the bias of DiD when these necessary and sufficient conditions are violated.

The bias analysis accommodates, but does not require, data on additional pre-treatment periods. Suppose that there is one additional pre-treatment period, $t=0$, in which no units are treated, so that $Y_{i0}=Y_{i0}(0)$ for $i=1,\dots,n$. We allow selection to also depend on the shocks in period $t=0$, that is, we allow $\omega_i$ to be a function of $(\varepsilon_{i0},\eta_{i0})$.

Bias decomposition

The following lemma provides a decomposition of the bias of DiD.

lemma[Bias decomposition] Suppose that $P(G_i=1)\in (0,1)$. Then, for a given $\omega_i$, the bias of DiD can be decomposed as follows, $$ \operatorname{DiD}-\operatorname{ATT}=\Delta_{\operatorname{post}}^{\operatorname{sel}}+\Delta_{\operatorname{post}}^{\operatorname{dev}}, $$ where \begin{eqnarray*} \Delta_{\operatorname{post}}^{\operatorname{sel}}&=&\frac{E[(G_i-E[G_i|\omega_i])(\dot{Y}_{i2}(0)-\dot{Y}_{i1}(0))]}{P(G_i=1)P(G_i=0)},\\ \Delta_{\operatorname{post}}^{\operatorname{dev}}&=&\frac{E[E[G_i|\omega_i]E[\dot{Y}_{i2}(0)-\dot{Y}_{i1}(0)|\omega_i]]}{P(G_i=1)P(G_i=0)}. \end{eqnarray*}

The decomposition in Lemma (ref) shows that the bias of DiD is equal to the sum of two components. The component $\Delta_{\operatorname{post}}^{\operatorname{sel}}$ captures the parallel trends violation due to selection on unobservables not contained in $\omega_i$. The component $\Delta_{\operatorname{post}}^{\operatorname{dev}}$ captures the parallel trends violation due to deviations from the necessary and sufficient condition for Assumption (ref) when selection is based on $\omega_i$.

In the following, we apply the bias decomposition in Lemma (ref) to settings where selection based on pre-treatment unobservables is likely. Other classes of selection mechanisms could also be considered.

Application to imperfect foresight

Suppose that researchers deem selection on pre-treatment unobservables likely. In this case, there are two potential sources of bias: selection on post-treatment unobservables and violations of the martingale condition. In the following, we characterize these two bias terms and provide stratgies for benchmarking them using pre-treatment data.

To simplify the notation, define $\varepsilon_i^{t}\equiv (\varepsilon_{i0},\dots,\varepsilon_{it})$ and $\eta_i^{t}\equiv(\eta_{i0},\dots,\eta_{it})$ for $t>0$. Suppose that $\omega_i=(\alpha_i,\varepsilon_i^{1},\mu_i,\eta_i^{1})$. In this case, the necessary and sufficient condition in Corollary (ref) generalizes to

equation[equation omitted — 125 chars of source]

To aid interpretation and assess the magnitude of the bias components in Lemma (ref), we characterize them under the following linear relaxation of the martingale condition (ref).\footnote{We focus on linear relaxations for convenience. Extensions to nonparametric relaxations of the form $ E[\dot{Y}_{it}(0)|\alpha_i,\varepsilon_{i0},\dots, \varepsilon_{i(t-1)}]=\sigma_{t}\rho(\dot{Y}_{i(t-1)}(0)),$ where $\rho(\cdot)$ is an arbitrary nonparametric function and $\sigma_1$ is normalized to one, are straightforward.}

customass{REL} The following relaxation of the martingale condition holds:\footnote{Assumption (ref) yields a linear autoregressive model. This class of models has been studied extensively in the time series literature under restrictions on the heterogeneity of the coefficient nicholls1982random,regis2022random.} $$ E[\dot{Y}_{it}(0)|\alpha_i, \varepsilon_i^{t-1},\mu_i,\eta_i^{t-1}]=\rho_t\dot{Y}_{i(t-1)}(0), \quad i=1,\dots,n,\quad t=1,2 $$

Assumption (ref) imposes an AR(1) model with time-varying coefficients on $\dot{Y}_{it}(0)$,

equation[equation omitted — 154 chars of source]

Note that Assumption (ref) is imposed on the demeaned potential outcomes and thus allows for mean shifts in $Y_{it}(0)$. If $\rho_2=1$, Assumption (ref) reduces to the martingale assumption, $E[\dot{Y}_{i2}(0)|\alpha_i,\varepsilon_i^{1},\mu_i,\eta_i^{1}]=\dot{Y}_{i1}(0)$. As a result, deviations from this martingale property under Assumption (ref) are fully characterized by the deviation of $\rho_2$ from $1$, $(\rho_2-1)$.

The following proposition characterizes the bias components under Assumption (ref).

proposition[Bias characterization under linear martingale relaxation] Suppose that $P(G_i=1)\in(0,1)$, $\omega_i=(\alpha_i,\varepsilon_i^{1},\mu_i,\eta_i^{1})$, and Assumption (ref) holds. Then, \begin{eqnarray*} \Delta_{\operatorname{post}}^{\operatorname{sel}} &=& E[\zeta_{i2}|G_i=1]-E[\zeta_{i2}|G_i=0],\\ \Delta_{\operatorname{post}}^{\operatorname{dev}} &=& (\rho_2-1)(E[Y_{i1}|G_i=1]-E[Y_{i1}|G_i=0]). \end{eqnarray*}

Proposition (ref) shows that $\Delta_{\operatorname{post}}^{\operatorname{sel}}$ is equal to the mean difference in $\zeta_{i2}$ between both groups. Recall that $\zeta_{i2}$ is the difference between $\dot{Y}_{i2}(0)$ and its conditional expectation given pre-treatment unobservables $E[\dot{Y}_{i2}(0)|\alpha_i,\varepsilon_i^{1},\mu_i,\eta_i^{1}]$. If selection depends on post-treatment unobservables including $\varepsilon_{i2}$, then $\zeta_{i2}$ is correlated with selection $G_i$, so that $E[\zeta_{i2}|G_i=1]$ is not equal to $E[\zeta_{i2}|G_i=0]$.

Proposition (ref) further shows that $\Delta_{\operatorname{post}}^{\operatorname{dev}}$ is equal to the product of the martingale deviation, $(\rho_2-1)$, and the observed pre-treatment difference, $E[Y_{i1}|G_i=1]-E[Y_{i1}|G_i=0]$. This shows that the sensitivity of DiD with respect to violations of the martingale assumption depends on the pre-treatment group difference. It underscores that only if the martingale property holds exactly can we ignore pre-treatment differences (and the selection on unobservables they are indicative of). This discussion motivates using covariate adjustment for reducing the pre-treatment difference and thus the potential bias of DiD due to violations of the martingale assumption. We illustrate this point in Section (ref) and defer the formal analysis of the DiD bias decomposition with covariates to Appendix (ref). If the treatment is randomly assigned, then $E[Y_{i1}|G_i=1]-E[Y_{i1}|G_i=0]=0$ and $\Delta_{\operatorname{post}}^{\operatorname{dev}}=0$.

An important takeaway from Proposition (ref) is that the bias of DiD is an affine function of the martingale deviation $(\rho_2-1)$, where the slope is the pre-treatment difference, $E[Y_{i1}|G_i=1]-E[Y_{i1}|G_i=0]$, and the intercept is the post-treatment difference, $E[\zeta_{i2}|G_i=1]-E[\zeta_{i2}|G_i=0]$. The only component that is directly observable from the data is the pre-treatment difference. We next demonstrate how we can benchmark $\rho_2$ and $E[\zeta_{i2}|G_i=1]-E[\zeta_{i2}|G_i=0]$ using pre-treatment data.

Benchmarking DiD bias components

Here, we provide benchmarks for $\Delta_{\operatorname{post}}^{\operatorname{sel}}$ and $\Delta_{\operatorname{post}}^{\operatorname{dev}}$ based on pre-treatment data that allow practitioners to assess and sign the bias of DiD in applications.

The unobservable component in $\Delta_{\operatorname{post}}^{\operatorname{dev}}$ is $\rho_2$ for which there are two natural benchmarks. First, under Assumption (ref), $\rho_1$ can be identified from the pre-treatment data by noting that $E[\dot{Y}_{i1}(0)|\dot{Y}_{i0}(0)]=E[E[\dot{Y}_{i1}(0)|\alpha_i,\varepsilon_{i0},\mu_i,\eta_{i0}]|\dot{Y}_{i0}(0)]=\rho_1\dot{Y}_{i0}(0)$, such that $\rho_1$ is identified as the coefficient of a population regression of $\dot{Y}_{i1}$ on $\dot{Y}_{i0}$. This is a useful benchmark because $\rho_2=\rho_1$ under time-homogeneity of $\rho_t$ in Assumption (ref). Second, we can consider the persistence in the control group, $E[\tilde{Y}_{i2}(0)|\tilde{Y}_{i1}(0),G_i=0]=\rho^{0}_2\tilde{Y}_{i1}(0)$, where $\tilde{Y}_{it}(0)\equiv Y_{it}(0)-E[Y_{it}(0)|G_i=0]$, and use $\rho^{0}_2$ to inform $\rho_2$. This is a useful benchmark because $\rho_2=\rho_2^0$ under unconfoundedness.\footnote{Specifically, the unconfoundedness assumption $Y_{i2}(0)\perp \!\!\! \perp G_i|Y_{i1}(0)$ implies that $\rho_2=\rho^{0}_2$. This follows because $Y_{i2}(0)\perp \!\!\! \perp G_i|Y_{i1}(0)$ implies that $E[\dot{Y}_{i2}(0)|\dot{Y}_{i1}(0),G_i=0]=E[\dot{Y}_{i2}(0)|\dot{Y}_{i1}(0)]$. The proposed benchmarking strategy allows researchers to consider a range of values for $\rho_2$, including $\rho_2^0$, the value of $\rho_2$ identified under unconfoundedness.}

As for $\Delta_{\operatorname{post}}^{\operatorname{sel}}$, it is helpful to consider its observable pre-treatment analogue, $$\biasselpre\equiv E[\zeta_{i1}|G_i=1]-E[\zeta_{i1}|G_i=0], \quad \text{where}~~\zeta_{i1}=\dot{Y}_{i1}-\rho_1\dot{Y}_{i0}.$$ To relate $\Delta_{\operatorname{post}}^{\operatorname{sel}}$ and $\biasselpre$, note that (assuming $\biasselpre\neq 0$) $$\Delta_{\operatorname{post}}^{\operatorname{sel}}= \frac{\rho_{G,\zeta_2}\sigma_{\zeta_2}}{\rho_{G,\zeta_1}\sigma_{\zeta_1}}\biasselpre,$$ where $\rho_{G,\zeta_t}\equiv Corr(G_i,\zeta_{it})$ and $\sigma_{\zeta_t}^2\equiv Var(\zeta_{it})$.

In the case where $\sigma_{\zeta_1}=\sigma_{\zeta_2}$, the relative magnitude of $\Delta_{\operatorname{post}}^{\operatorname{sel}}$ to $\biasselpre$ is simply the ratio of the (scale-free) correlation coefficients, $\rho_{G,\zeta_2}/\rho_{G,\zeta_1}$. This ratio measures the relative degree of selection on pre- vs. post-treatment unobservables. For example, if contextual knowledge suggests that there is “more selection” on pre-treatment than on post-treatment unobservables, then $|\Delta_{\operatorname{post}}^{\operatorname{sel}}|\le |\biasselpre|$, assuming $\operatorname{sgn}\left(\rho_{G,\zeta_1}\right)=\operatorname{sgn}\left(\rho_{G,\zeta_2}\right)$. The edge case where $\Delta_{\operatorname{post}}^{\operatorname{sel}}=\biasselpre$ captures settings where the extent of selection on post- and pre-treatment unobservables is the same.

In practice, we recommend that researchers determine the robustness of DiD results based on the benchmarks we discuss above. We illustrate this approach in two empirical applications in Section (ref).

Empirical illustration of bias decomposition

Here we illustrate the bias decomposition in two applications. In the first application, we revisit the NSW training program, where we have access to an experimental estimate of the ATT and thus an estimate of the bias of DiD. This bias estimate allows us to directly evaluate the bias decomposition and benchmarking strategy using a lalonde1986evaluating-style exercise. In the second application, we consider the Medicaid expansion to demonstrate the usefulness of the bias decomposition to DiD applications with aggregate selection.

Both applications demonstrate that pre-treatment differences in mean outcomes between the treatment and control groups matter. While we can ignore such differences under parallel trends, our analysis underscores the central role they play once we entertain the possibility of violations of parallel trends. Both applications further highlight that covariates are crucial for reducing the pre-treatment difference in mean outcomes and thereby rendering DiD less sensitive to martingale violations.

NSW training program

Setup and DiD analysis

The evaluation of job training programs is one of the classical applications of DiD in economics. Here, we revisit the analysis of the causal effect of the NSW training programs on post-treatment earnings lalonde1986evaluating. We use the same dataset as santanna2020drdid and consider the “Dehejia1999,Dehejia2002 sample.”\footnote{The data are from the DRDID R-package santanna_zhao_DRDID.} This sample combines the experimental treatment group (185 individuals) with an observational control group (15,992 individuals).

The outcome of interest is earnings. We observe individual-level data on earnings for two pre-treatment periods, 1974 and 1975, and one post-treatment period, 1978. We also have access to a set of baseline covariates: age, years of education, and indicators for high school dropouts, married individuals, Black and Hispanic individuals.

The unconditional DiD estimate using 1975 as the pre-treatment period ($t=1$) and 1978 as the post-treatment period ($t=2$) is equal to $\widehat{\operatorname{DiD}}=$ 3,621 (s.e.\ 610). A comparison to the experimental benchmark, which is 1,794 (s.e.\ 671), shows that the unconditional DiD substantially overestimates the returns to the training program. The estimated bias relative to the experimental benchmark is statistically and economically significant at 1,827, comparable in magnitude to the experimental benchmark.

With covariates, the regression-adjusted DiD estimate under conditional parallel trends is equal to $E_n[\widehat{\operatorname{DiD}}(X_i)|G_i=1]=$ 2,436 (s.e.\ 653), where $E_n$ denotes the sample average and $\widehat{\operatorname{DiD}}(X_i)$ is the conditional DiD estimate obtained using the regression-adjusted DiD estimator. This shows that adjusting for differences in baseline covariates reduces the bias of DiD to 642, about a third of the bias of the unconditional DiD relative to the experimental benchmark.

It is standard to report the results from pre-trends tests when there are additional pre-treatment periods. Based on the pre-treatment data from 1974 and 1975, the unconditional and regression-adjusted DiD estimates are 198 (s.e.\ 280) and 335 (s.e.\ 309), respectively.

Despite the non-rejections of the pre-trends tests, the sensitivity of the DiD estimates to parallel trends violations remains a major concern for three reasons. First, pre-tests are, by construction, not direct tests of parallel trends assumptions. Second, these tests can be substantially underpowered roth_pre-test_2022. Finally, building on the necessary and sufficient conditions in Section (ref), ghanem2026when show that pre-trends can be uninformative under imperfect foresight. These issues are particularly evident in this application, where the pre-test does not reject, despite unconditional DiD being significantly biased relative to the experimental benchmark. Next, we demonstrate how our selection-based bias decomposition can help us better understand the difference between the DiD and the experimental estimate in this application.

Decomposing the bias of DiD

We start by illustrating the bias decomposition without covariates. Replacing the population expectations by sample averages, we obtain

eqnarray*[eqnarray* omitted — 344 chars of source]

There is a substantial pre-treatment difference: average earnings in 1975 are much lower in the treatment than in the control group.

Figure (ref) displays $\widehat{\Delta}_{\operatorname{post}}$ as a function of $\rho_2$ together with the bias estimate based on the experimental benchmark. Suppose first that $\Delta_{\operatorname{post}}^{\operatorname{sel}}=0$. In this case, the bias of DiD equals $\widehat{\Delta}_{\operatorname{post}}=(\rho_2-1)(-\text{12,119})$, depicted by the blue line. It is solely driven by violations of the martingale property (i.e., differences between $\rho_2$ and $1$). Alternatively, consider the edge case where $\Delta_{\operatorname{post}}^{\operatorname{sel}}=\biasselpre$. This corresponds to the case with equal sign and strength of selection on $\zeta_{i1}$ and $\zeta_{i2}$, as discussed in Section (ref). The sample analogue of $\biasselpre$ equals $-2,049$, resulting in the following bias estimate (red line in Figure (ref)), $$ \widehat{\Delta}_{\operatorname{post}}=-2,049+(\rho_2-1)(-\text{12,119}). $$

figure[figure omitted — 1,440 chars of source]

As we discussed in Section (ref), there are two natural benchmarks for $\rho_2$: $\rho_1$, the pre-treatment counterpart of $\rho_2$ and $\rho_2^0$, its control group counterpart. The corresponding estimates are $\widehat\rho_1=0.603$ and $\widehat\rho^0_2=0.695$ and are depicted in Figure (ref).\footnote{Recall that the post-treatment earnings are measured in 1978, so that $\rho_2$ measures the persistence over three years. To account for the difference in periodicity when estimating $\rho_1$, we proceed in two steps. First, we regress $\dot{Y}_{i1975}$ on $\dot{Y}_{i1974}$ to obtain an estimate of the yearly persistence in the pre-treatment period, $\tilde{\rho}_1=0.845$. Second, we adjust for the difference in periodicity by computing $\widehat\rho_1$ as $\widehat\rho_1=(\widehat{\tilde{\rho}}_1)^3=0.603$. This is justified under a linear AR(1) model for the demeaned outcomes in the pre-treatment period.} Both benchmark values would suggest that the unconditional DiD is upwardly biased, consistent with the experimental bias estimate.

The analysis without covariates demonstrates that the bias of DiD is very sensitive to deviations from the martingale property. The lack of robustness is driven by the treatment and control groups being very different before the treatment. This discussion suggests that we may reduce the pre-treatment difference and improve the robustness of DiD by adjusting for differences in baseline covariates.

We therefore incorporate covariates into our analysis in Figure (ref). In Appendix (ref), we show that under a linear relaxation of the conditional martingale property, the unconditional bias of DiD with covariates can be decomposed as

eqnarray*[eqnarray* omitted — 155 chars of source]

where $\Delta_{\operatorname{post}}^{\operatorname{sel}}\equiv E[\Delta_{\operatorname{post}}^{\operatorname{sel}}(X_i)|G_i=1]$.

Analogous to the unconditional bias decomposition, consider first the case where $\Delta_{\operatorname{post}}^{\operatorname{sel}}=0$. Using the regression-adjusted estimator for the pre-treatment difference described in Appendix (ref), we obtain the following bias estimate (blue line in Figure (ref)), $ \widehat{\Delta}_{\operatorname{post}}=(\rho_2-1)(-6,113). $ Adjusting for differences in baseline covariates reduces the magnitude of the pre-treatment difference by approximately 50%. As a result, incorporating covariates makes the bias of DiD less sensitive to violations of the martingale property.

Alternatively, consider the case where $\biasselpre=\Delta_{\operatorname{post}}^{\operatorname{sel}}$, which is implied by $\biasselpre(X_i)=\Delta_{\operatorname{post}}^{\operatorname{sel}}(X_i)$. Using the regression-adjusted estimator of $\biasselpre$ described in Appendix (ref), we obtain (red line in Figure (ref)) $$ \widehat{\Delta}_{\operatorname{post}}=-1,333+(\rho_2-1)(-6,113). $$ The estimates of $\rho_1$ and $\rho_2^0$ with covariates are $\widehat\rho_1=0.566$, which is somewhat smaller than without covariates, and $\widehat\rho_2^0=0.715$, which is somewhat larger than without covariates.\footnote{Under the linear relaxation of the martingale assumption, the yearly persistence in the pre-treatment period, $\tilde\rho_1$, can be estimated by regressing $\ddot{Y}_{i1975}$ on $\ddot{Y}_{i1974}$. The resulting estimate is $\widehat{\tilde{\rho}}_1=0.827$. Adjusting for the difference in periodicity yields $\widehat\rho_1=(\hat{\tilde{\rho}}_1)^3=0.566$.} Both benchmark values suggest the same sign and a similar magnitude of the bias as the experimental benchmark.

This analysis demonstrates how the proposed bias decomposition can help empirical practitioners assess the bias of DiD and its sensitivity. This is especially important in applications such as this one, where the (unconditional) pre-trends tests do not reject, even though DiD is biased relative to the experimental benchmark.

Medicaid Expansion

Setup and DiD analysis

We revisit the DiD evaluation of Medicaid expansion to illustrate the relevance of our selection-based bias decomposition to DiD settings with aggregate selection. We use the sample from the 2$\times$2 DiD implementation in baker2026difference, but consider one additional pre-treatment period.\footnote{The data are currently available in the following GitHub repository \url{https://github.com/pedrohcgs/JEL-DiD} and will soon be posted on OPENICPSR as part of the official replication package.} The treatment group consists of states that have expanded Medicaid in 2014, whereas the control group consists of states that have not expanded by 2019. The pre-treatment periods are 2012 and 2013 ($t=0,1$), and the post-treatment period is 2014 ($t=2$).

In the context of Medicaid expansion, the outcome of interest, observed at the county level, is the crude mortality rate for people aged 20-64 (measured per 100,000). In our conditional DiD analysis, we also include the percentages of a county’s population that are female, white, or Hispanic; the unemployment rate; the poverty rate; and county-level median income (in thousands of dollars)---all measured in 2012---in our regression adjustment.\footnote{We use covariate values from 2012, so we can treat them as time-invariant in our analysis, simplifying the exposition and avoiding the strong, possibly unrealistic assumption that our covariates are strictly exogenous in this application. In line with callaway_SantAnna_2021's implementation in the did R package, baker2026difference fixed covariate values at the 2013 values for post-treatment analysis, and at the 2012 values for pre-treatment periods, 2012-2013. We refer the reader to caetano2022difference, as well as to ghanem2026when, for additional discussion.} All estimates are weighted by county population in 2013.

We first examine the unconditional DiD estimate using 2013 and 2014, which equals $-2.6$ (s.e. 1.5), indicating a reduction in mortality due to Medicaid expansion that is statistically significant at the $10\%$ level. Once we account for covariates, however, the results are no longer significant with a regression-adjusted DiD estimate of $-2.1$ (s.e. 2.2).

Before we proceed to the bias decomposition, we conduct the pre-trends tests. We find that unconditional and conditional pre-trends tests are not rejected at the 5% level, with differences in pre-trends of $-2.8$ ($1.5$) and $-2.6$ (s.e. $2.5$), respectively. However, note that the pre-trends for the unconditional DiD are significant at the $10\%$ level.

Decomposing the bias of DiD

We next present the sample analogues of the bias decomposition, as described in Section (ref). For the unconditional DiD case, the sample analogue of the bias can be decomposed as follows

eqnarray*[eqnarray* omitted — 342 chars of source]

where $-53.7$ denotes the pre-treatment difference in means between the treatment and control group, statistically significant at the 1% level.

Figure (ref) demonstrates that this substantial pre-treatment difference translates to the bias of DiD being very sensitive to martingale violations. When considering the benchmark values for $\rho_2$, however, we note that both $\hat{\rho}_1$ and $\hat{\rho}_2^0$ are fairly close to 1, and therefore, the bias of DiD due to martingale deviations is relatively small for those values. If one is willing to assume that $\Delta_{\operatorname{post}}^{\operatorname{sel}}=0$, then our analysis suggests a positive bias of DiD for these benchmark values of $\rho_2$ (Figure (ref)). If we instead assume that $\Delta_{\operatorname{post}}^{\operatorname{sel}}=\biasselpre$, then our analysis indicates a negative bias of DiD for these benchmark values.

Once we adjust for covariates, the pre-treatment difference is no longer significant. Indeed, it is a negligible difference yielding the following sample analogue of the bias of the regression-adjusted DiD

eqnarray*[eqnarray* omitted — 230 chars of source]

As a result, the bias is insensitive to violations of the martingale condition and mostly driven by the magnitude of $\Delta_{\operatorname{post}}^{\operatorname{sel}}$, as illustrated in Figure (ref). If $\Delta_{\operatorname{post}}^{\operatorname{sel}}=0$, then our analysis suggests that the bias of DiD is negligible, whereas if $\Delta_{\operatorname{post}}^{\operatorname{sel}}=\biasselpre$, then the bias is negative.

The bias component $\Delta_{\operatorname{post}}^{\operatorname{sel}}$ captures the bias due to selection on post-treatment unobservables. Selection on post-treatment unobservables is unlikely if the time-varying determinants of mortality are difficult to predict. This is the case, for example, if mortality is determined by factors such as adverse weather shocks and disease outbreaks which are arguably difficult to perfectly foresee, even one year ahead. In this case, a researcher can argue that $\Delta_{\operatorname{post}}^{\operatorname{sel}}$ is close to zero, which implies that the bias of DiD is negligible.

figure[figure omitted — 1,193 chars of source]

Implications for empirical practice

In this paper, we study parallel trends assumptions through the lens of selection into treatment. We derive necessary and sufficient conditions that clarify the empirical content of parallel trends, shed light on the trade-offs between assumptions on selection and time series restrictions, motivate DiD bias decompositions and benchmarking strategies, and provide theory-based templates for assessing and justifying parallel trends in applications with and without covariates. Below, we summarize the main implications of our results for practitioners.

\noindentRestrictions on selection are unavoidable in DiD designs. The necessary and sufficient condition in Corollary (ref) underscores that if researchers are not willing to impose any restrictions on selection, then parallel trends is equivalent to the untreated potential outcomes being constant over time up to deterministic mean shifts. Therefore, in realistic settings, relying on parallel trends assumptions implicitly imposes restrictions on the time-varying unobservables and how selection depends on them.

Contextual and economic knowledge about selection can be used to assess and justify parallel trends. Our analysis provides a general approach to derive necessary and sufficient conditions for parallel trends with and without covariates. Importantly, these conditions do not require the researchers to specify explicit selection mechanisms, which may be difficult in practice. Instead, the researchers only need to specify what the units select on. When doing so, it is crucial for researchers to consider the periodicity of the data, the timing of the selection decision, the information set available to the units, who make the selection decision (e.g., the individuals themselves or caseworkers in the training program example), as well as at which level the selection decision is made (e.g., at the level of an individual economic agent vs. at the aggregate level via the political process).\footnote{The importance of the information available to units is underscored by the results in marx2024parallel, who study specific economic models of selection including learning and optimal stopping.} Another practical byproduct of our analysis is a menu of selection-based templates for assessing and justifying parallel trends with and without covariates, see Section (ref) and Appendix (ref).

Selection-based bias decompositions are useful to sign and benchmark the bias of DiD. In Section (ref), we provide a general selection-based decomposition of the bias of DiD. We then apply this decomposition to settings where selection on pre-treatment unobservables is likely. Exploiting a martingale relaxation, we show that the bias of DiD can be decomposed into two components: (i) the bias due to selection on post-treatment unobservables, (ii) the bias due to deviations from the martingale property necessary and sufficient for parallel trends under imperfect foresight. This characterization can be used in practice to sign and benchmark these two bias components, as we demonstrate in Section (ref). A practical implication of this characterization is that the pre-treatment difference between the treatment and control group is a key determinant of the bias of DiD when parallel trends is violated.

{0pt}