EconBase
← Back to paper

Matching $\leq$ Hybrid $\leq$ Difference in Differences

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

112,803 characters · 29 sections · 128 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Matching $$ Hybrid $$ Difference in Differences

\allowdisplaybreaks

abstract{6mm} Since lalonde1986evaluating's (lalonde1986evaluating) seminal paper, there has been ongoing interest in estimating treatment effects using pre- and post-intervention data. Scholars have traditionally used experimental benchmarks to evaluate the accuracy of alternative econometric methods, including Matching, Difference-in-Differences (DID), and their hybrid forms heckman1998_2matching,dehejia2002propensity,smith2005does. We revisit these methodologies in the evaluation of job training and educational programs using four datasets lalonde1986evaluating,heckman1998characterizing,smith2005does,chetty2014measuring1,athey2020combining, and show that the inequality relationship, Matching $\leq$ Hybrid $\leq$ DID, appears as a consistent norm, rather than a mere coincidence. We provide a formal theoretical justification for this puzzling phenomenon under plausible conditions such as negative selection, by generalizing the classical bracketing angrist2009mostly. Consequently, when treatments are expected to be non-negative, DID tends to provide optimistic estimates, while Matching offers more conservative ones. {\bf Keywords:} bias, difference in differences, educational program, job training program, matching.
comment\section{general outline} \begin{itemize} • - Main Part 1: Short term bracketing theory: - Introduce the DID and the matching and the hybrid version - Show the bracketing relationship (ofc, have to say the ding and li(2019) paper here, but should be fine) - Show the situations where the DIDM bias is greater than 0 or less than 0 • - Main Part2: Short term bracketing applicaton - revisit the Heckman et al type paper's application (JTPA, Lalonde's NSW) • - Main Part3: Long term bracketing theory ( • - Main Part4: Long term bracketing application - in addition to the project star and the star, we should also have the california gain application as well? ( \end{itemize}

Introduction

Since the seminal work by lalonde1986evaluating, there has been substantial interest in accurately estimating the short-term average treatment effects on the treated (ATT) and other causal parameters using panel data with two (pre-treatment and post-treatment) periods heckman1998_2matching,dehejia2002propensity,smith2005does. This body of literature has focused on a debate concerning which method -- matching (M), difference-in-differences (DID), or their hybrid (DIDM) -- is more effective at replicating experimental estimates when using observational data.

In the current era, where the DID has once again captured the attention of empirical practitioners, we contribute to the ongoing debate by examining the relationships among these three estimands from both empirical and theoretical perspectives. Our analysis clarifies the conditions under which one method provides the most conservative estimates while another offers the most optimistic estimates. By doing so, we provide a more nuanced understanding of the relative strengths and limitations of each approach, enabling researchers to make more informed methodological choices based on the specific characteristics of their data.

We begin by highlighting a puzzling pattern consistently observed in the empirical data used in both the aforementioned studies and other seminal works in the economics of education and labor economics. Specifically, the inequality relationship, M $\leq$ DIDM $\leq$ DID, frequently appears as a consistent norm, rather than a mere coincidence, when assessing the effectiveness of educational and job training programs.

For example, Figure (ref) illustrates the biases of the M, DIDM, and DID estimates for job training programs relative to experimental estimates perceived as benchmark truths. These estimates are excerpted from two seminal papers: heckman1998characterizing and smith2005does, which revisits the analysis of lalonde1986evaluating,dehejia1999causal,dehejia2002propensity. Despite differences in the job training programs studied, data sets, and estimation methods, the inequality relationship, M $\leq$ DIDM $\leq$ DID, holds robustly across both papers. This relationship remains robust even when considering alternative estimation methods, different programs, or other data sets, as will be demonstrated shortly in this paper.

figure[figure omitted — 735 chars of source]

This pattern is reminiscent of the so-called `bracketing' relationship between the lagged dependent variable (LDV) estimator and the fixed-effect (FE) estimator in panel regressions, as presented in angrist2009mostly. Specifically, angrist2009mostly outline plausible conditions under which the inequality relationship, LDV $\leq$ FE, holds on theoretical grounds. Their assumptions necessitate negative selection into treatment and the absence of explosive outcomes, which make plausible sense in many labor economic settings. This result has been elegantly extended to a nonparametric setup by ding2019bracketing.

We find that similar plausible conditions, in the spirit of angrist2009mostly and ding2019bracketing, also give rise to the aforementioned inequality relationship, M $\leq$ DIDM $\leq$ DID. In fact, a special case of our general double bracketing result, M $\leq$ DIDM $\leq$ DID, reduces to the conventional bracketing result, LDV $\leq$ FE, established by angrist2009mostly and ding2019bracketing.

Recall that M, DID, and DIDM identify the true causal parameter under the assumptions of observational unconfoundedness, parallel trends, and conditional parallel trends, respectively. In practice, an empirical researcher may not know which of these three alternative conditions is satisfied for an application of interest. Our double bracketing result, M $\leq$ DIDM $\leq$ DID, implies that the true causal parameter is bracketed below by M and above by DID, with DIDM between them, when one of the three alternative assumptions holds true. In other words, our theoretical prediction implies that the DID approach tends to yield the most optimistic estimates while the M approach tends to yield the most conservative estimates.

While our discussions primarily focus on the classic two-period framework heckman1998characterizing,smith2005does, we also extend our double bracketing result to a broader class that includes cases with multi-dimensional covariates, dynamic and multi-period settings, and impulse response functions in event studies, as explored in the recent literature callaway2018difference,acemoglu2019democracy,deChaisemartin2020two,dube2023local,imai2023matching.

Using four data sets that have been used for the evaluation of job training and educational programs in the literature lalonde1986evaluating,heckman1998characterizing,smith2005does,athey2020combining, we empirically examine our double bracketing hypothesis, M $\leq$ DIDM $\leq$ DID. In light of the robustness of this inequality relationship in all the empirical scenarios, we provide formal theoretical explanations for this intriguing phenomenon. Our assumptions required for the double bracketing relationship, motivated by angrist2009mostly and ding2019bracketing as mentioned earlier, is not only in line with the literature but are also empirically testable. Hence, we examine our assumptions, as well as the double bracketing relations per se, using these empirical data sets. It turns out that our assumptions, as well as the double bracketing relationship, M $\leq$ DIDM $\leq$ DID, indeed hold robustly in all these empirical cases.

comment- Extending these findings, we find that our bracketing results hold for multi-dimensional covariates as well as dynamic analogues of the classical two-time periods, significantly expanding the scope of the classical bracketing results, and widening our scope to analyses like acemoglu2019democracy,roth2023review. In addition to providing a way to intepret an important historical debate regarding the credibility revoluation, going forward, how can practitioners benefit from our genralized bracketing result? To see this concretely, consider an empiricist who has access to an observational data. She is considering the appropriate identification strategy for her work, and is contemplating whether a parallel trends assumption is appropriate in her context. She has to some extent a persuasive argument, but there is also some concerns for the violation of it due to some selection on pretreatment outcomes. While she conducted a pre-trends test and the test was not rejeted, since as roth2023parallel she is still not sure which one is Our empirical applications are also crucial. We consider three domains of application: starting from the classical application of Job training program heckman1998_2matching,heckman1998characterizing,smith2005does,dehejia2002propensity, we also consider the case of Educational program evaluation Krueger1999QJE,chetty2014measuring1,chetty2011does, and finally Political Economy acemoglu2019democracy. First, for the job training program, we consider two seminal job training program of NSW participation and also the JTPA. Both datasets come from Second, for the Educational Program, we use the small class size treatment on educational outcome data.

Relation to the Literature

The question we investigate relates to a long literature of econometrics including a number of seminal papers.

First, this paper closely relates to the classical debate in economics on what type of non-experimental estimates replicate the experimental estimates of lalonde1986evaluating. The pioneering papers in that literature are heckman1998characterizing, heckman1998_2matching, dehejia2002propensity, and smith2005does. There, the main purpose was to find the single best estimator, often measured by the absolute bias. We are interested in signs, as well as the magnitudes, of bias to identify which estimator is the most conservative/optimistic.

More recently, one notable work revisited a related problem. Specifically, Figure (ref) replicates chabe2017should and presents the absolute biases of the three estimands, M, DIDM, and DID, based on the estimates from heckman1998characterizing and smith2005does.

figure[figure omitted — 727 chars of source]

With the symbol `$\Delta$' denoting the bias, this figure essentially shows $|\Delta $M$|\geq|\Delta $DIDM$|\geq|\Delta $DID$|$, implying that DID achieves the smallest bias for these empirical cases. Notably, this Figure (ref) depicts the absolute biases corresponding to the signed biases shown in Figure (ref) in our introductory section.

Thus, while the existing literature primarily focuses on identifying the best estimator based on absolute bias, our paper distinguishes itself by examining the relative signs of biases. This signed characterization provides more nuanced insights; for instance, Figure (ref) illustrates that M underestimates the true causal effects, whereas DID overestimates them. Beyond this primary distinction between absolute and signed biases, our paper differs from chabe2017should in two other key aspects. First, we provide a formal theoretical result, whereas the existing work derives its conclusions from numerical studies. Second, the existing work relies on functional form and structural assumptions -- such as separability and distinctions between transitory and permanent income -- tailored to job training programs. In contrast, we adopt a more model-free, nonparametric approach.

Our theoretical results build on the existing literature on bracketing. For linear panel models, angrist2009mostly develop a bracketing relationship, showing that LDV $\leq$ FE. This result has been extended to a nonparametric context by ding2019bracketing. We contribute to this literature in two key ways. First, we consider a more general framework that encompasses the symmetric DID (DIDM) studied by heckman1998characterizing and smith2005does in particular. Second, within this general framework, we show that DIDM, which heckman1998characterizing recommend for minimizing bias, is further bracketed between M and DID, resulting in our double bracketing outcome. A special case of our general framework reduces M and DIDM to LDV, and thus, our double bracketing relationship, M $\leq$ DIDM $\leq$ DID, simplifies to the conventional bracketing relationship, LDV $\leq$ FE.

Finally, this paper is of course related to the burgeoning literature on event studies and DID today -- see recent surveys by deChaisemartin2023two and roth2023review among others. This paper is also related to the huge literature on matching, including covariate matching, propensity score matching, (augmented) inverse probability weighting, nearest neighbor methods, and regression adjustments, among others. In particular, the matching literature has emphasized the importance of conditioning on lagged outcomes as matching factors to improve the precision of causal inference lalonde1986evaluating,dehejia1999causal,dehejia2002propensity. Economically, matching on lagged outcomes helps to address the so-called “ashenfelter1978estimating dip.” The empirical work by acemoglu2019democracy confirms that conditioning on lagged outcomes indeed yields more credible estimates in event studies. Effectively, they advocate the M over the DID in our language. Later, dube2023local generalize acemoglu2019democracy to a hybrid framework which we refer to as the DIDM in this paper. See the DID${}_\text{M}$ estimator of deChaisemartin2020two and the (panel) matching estimator of imai2023matching, as well as dube2023local -- they all propose and study the properties of what we refer to as the DIDM. We contribute to the above literature by shedding light on the systematic relationship among M, DIDM, and DID, and empirically and theoretically examining this relationship.

The Setup and Definitions

Let $W$ denote the group of treatment assignment such that units with $W=1$ receive treatment between periods $t=0$ and $t=1$, while those with $W=0$ remain untreated. Thus, everyone is untreated for $t \leq 0$, and only those with $W=1$ are treated for $t \geq 1$. Let $Y_t(d)$ denote the potential outcome under treatment status $d$ at time $t$. With these notations, suppose that a researcher is interested in identifying the average treatment effect on the treated (ATT) at time $t=1$ defined by

align*[align* omitted — 64 chars of source]

The M, DID, and DIDM Estimands

Letting $Y_t$ denote the observed outcome at time $t$, the seminal paper by heckman1998characterizing proposes three alternative estimands to this goal:

align*[align* omitted — 338 chars of source]

which we call the matching (M), the difference-in-differences (DID), and the difference-in-differences matching (DIDM), respectively. Here, we focus on the past outcome $Y_{-s}=Y_{-s}(0)$ as the key matching criterion following heckman1998characterizing and smith2005does, but we present an extension to general vector-valued matching criteria $X$ in Section (ref) to include other auxiliary covariates as well as more lagged outcomes as in the local projection approaches to event studies acemoglu2019democracy.

The M estimand identifies the true ATT (i.e, ${\theta_{\text{ATT}}^{\text{M}}} = {\theta_{\text{ATT}}}$ holds) if the matching condition

align*[align* omitted — 87 chars of source]

is satisfied. The DID estimand identifies the true ATT (i.e, ${\theta_{\text{ATT}}^{\text{DID}}} = {\theta_{\text{ATT}}}$ holds) if the parallel trend condition

align*[align* omitted — 103 chars of source]

is satisfied. Finally, the DIDM estimand identifies the true ATT (i.e, ${\theta_{\text{ATT}}^{\text{DIDM}}} = {\theta_{\text{ATT}}}$ holds) if the conditional parallel trend condition

align*[align* omitted — 118 chars of source]

is satisfied.

As heckman1998characterizing and smith2005does stressed, in the absence of knowledge about the true data-generating process, there is a risk of bias associated with the three estimands, ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DID}}}$, and ${\theta_{\text{ATT}}^{\text{DIDM}}}$. For instance, when the true DGP satisfies Condition M but does not satisfy Condition DID or Condition DIDM, then only ${\theta_{\text{ATT}}^{\text{M}}}$ is guaranteed the identification while estimators for ${\theta_{\text{ATT}}^{\text{DID}}}$ and ${\theta_{\text{ATT}}^{\text{DIDM}}}$ are doomed to be biased in general. Hence, understanding the systematic relationship among ${\theta_{\text{ATT}}^{\text{M}}}$ ${\theta_{\text{ATT}}^{\text{M}}}$, and ${\theta_{\text{ATT}}^{\text{DID}}}$ can help researchers form insights into the possible range within which the true causal effect, ${\theta_{\text{ATT}}}$, may lie when one of the three alternative conditions holds. We will empirically confirm the systematic double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, in Section (ref) and provide theoretical explanations for it in Sections (ref)--(ref).

Mutual Non-Nestedness of The M, DID, and DIDM Conditions

{\color{black} The three identifying assumptions underlying the estimands \({\theta_{\text{ATT}}^{\text{M}}}\), \({\theta_{\text{ATT}}^{\text{DID}}}\), and \({\theta_{\text{ATT}}^{\text{DIDM}}}\) each restrict the data‐generating process (DGP) in a distinct way. No one assumption nests another; that is, none of the conditions implies or is implied by any of the others. In the following, we provide intuitive counterexamples—some leveraging differences in sample composition—to illustrate these distinctions.

Condition M Holds but Condition DIDM Fails

Consider the following data-generating process (DGP) for the potential outcomes:

align*[align* omitted — 113 chars of source]

where $(\eta_0(0),\eta_0(1),\eta_1(0),\eta_1(1)) \perp \!\!\! \perp (W, Y_{-s})$ and $\mu(\cdot)$ is an arbitrary function such that $\mu(0) \neq \mu(1)$. Under this DGP, $(Y_1(1),Y_1(0)) \perp \!\!\! \perp W | Y_{-s}$ holds, but ${\text{E}}[Y_1(0)-Y_0(0)|Y_{-s},W=0]={\text{E}}[\eta_1(0)-\eta_0(0)]-\mu(0) \neq {\text{E}}[\eta_1(0)-\eta_0(0)]-\mu(1) = {\text{E}}[Y_1(0)-Y_0(0)|Y_{-s},W=1]$. In other words, Condition M holds, but Condition DIDM fails.

commentCondition M, $ (Y_1(1), Y_1(0)) \perp \!\!\! \perp W \mid Y_{-s}, $ in particular implies \[ E[Y_1(0) \mid Y_{-s}, W=1] = E[Y_1(0) \mid Y_{-s}, W=0]. \] Suppose that the pre-treatment outcome \(Y_0(0)\) is generated by an additional component that differs by treatment status and is not captured by \(Y_{-s}\). For example, assume the following data‐generating process: \[ \begin{aligned} Y_1(0) &= Y_0(0) + \delta + \eta \quad\text{with } \eta \text{ satisfying } E[\eta\mid Y_{-s},W] = 0, \quad\text{and}\\ Y_0(0) &= \mu(Y_{-s}) + \xi(W), \end{aligned} \] where \(\mu(\cdot)\) and $\xi(\cdot)$ are arbitrary functions such that \(\xi(1) \neq \xi(0)\). Under this specification, Condition M holds because, conditional on \(Y_{-s}\), the level of \(Y_1(0)\) is determined solely by \(Y_0(0)\) and an independent shock \(\eta\). However, when we consider the untreated outcome change, \[ \begin{aligned} E[Y_1(0)-Y_0(0)\mid Y_{-s},W] &= \delta + E[\eta\mid Y_{-s},W] \\ &= \delta, \end{aligned} \] one might initially think that the conditional trends are identical. Yet note that the unobserved heterogeneity in \(Y_0(0)\) (via \(\xi(W)\)) means that the levels differ by treatment status. In many applications, the evolution from \(Y_0\) to \(Y_1\) might be modeled or estimated relative to the pre-treatment level. If, for instance, practitioners subtract \(Y_0(0)\) from \(Y_1(0)\) only after having adjusted for \(Y_{-s}\), any mis-specification in the baseline captured by \(\xi(W)\) can induce differences in the estimated changes even though \(E[\eta\mid Y_{-s},W]=0\). In other words, while the matching condition guarantees comparability in \(Y_1(0)\) given \(Y_{-s}\), the conditional difference \(Y_1(0)-Y_0(0)\) may still differ by treatment status if \(Y_0(0)\) embeds treatment-specific components. Thus, Condition DIDM, \[ E[Y_1(0)-Y_0(0)\mid Y_{-s},W=1] = E[Y_1(0)-Y_0(0)\mid Y_{-s},W=0], \] can fail even though Condition M is satisfied.

Condition DIDM Holds but Condition M Fails

Conversely, consider the following DGP for the potential outcomes:

align*[align* omitted — 133 chars of source]

where $(\eta_0(0),\eta_0(1),\eta_1(0),\eta_1(1)) \perp \!\!\! \perp (W, Y_{-s})$ and $\mu(\cdot)$ is an arbitrary function such that $\mu(0) \neq \mu(1)$. Under this DGP, ${\text{E}}[Y_1(w)|Y_{-s},W=0]=\mu(0)+{\text{E}}[\eta_1(0)] \neq \mu(1)+{\text{E}}[\eta_1(0)] = {\text{E}}[Y_1(w)|Y_{-s},W=1] $, and hence, $(Y_1(1),Y_1(0)) \not\perp \!\!\! \perp W | Y_{-s}$, but ${\text{E}}[Y_1(0)-Y_0(0)|Y_{-s},W=0]={\text{E}}[\eta_1(0)-\eta_0(0)] = {\text{E}}[Y_1(0)-Y_0(0)|Y_{-s},W=1]$ holds. In other words, Condition M fails, but Condition DIDM holds.

commentConversely, suppose that the untreated outcome changes are homogeneous across treatment groups when conditioning on \(Y_{-s}\), so that \[ E[Y_1(0)-Y_0(0)\mid Y_{-s},W=1] = E[Y_1(0)-Y_0(0)\mid Y_{-s},W=0] = \delta(Y_{-s}), \] for some function \(\delta(\cdot)\). However, let the level of the untreated outcome at time 1 differ across groups even after conditioning on \(Y_{-s}\). For instance, consider: \[ \begin{aligned} Y_1(0) &= \mu(Y_{-s}) + \zeta(W) + \epsilon_1, \\ Y_0(0) &= \mu(Y_{-s}) + \epsilon_0, \end{aligned} \] with \(E[\epsilon_1-\epsilon_0\mid Y_{-s},W]=\delta(Y_{-s})\) and \(\zeta(1)\neq \zeta(0)\). In this case, \[ E[Y_1(0)\mid Y_{-s},W=1] = \mu(Y_{-s}) + \zeta(1) \quad \text{and} \quad E[Y_1(0)\mid Y_{-s},W=0] = \mu(Y_{-s}) + \zeta(0), \] so that the matching condition \[ (Y_1(1),Y_1(0)) \perp \!\!\! \perp W\mid Y_{-s} \] fails due to the presence of the treatment-dependent term \(\zeta(W)\). Nonetheless, because \(\zeta(W)\) is constant over time, the difference \[ E[Y_1(0)-Y_0(0)\mid Y_{-s},W] = \delta(Y_{-s}) \] remains the same across groups. Hence, Condition DIDM holds even though Condition M is violated.

Condition DIDM Holds but Condition DID Fails

Consider the following DGP for untreated potential outcomes:

align*[align* omitted — 85 chars of source]

where $E[\eta_1(0)-\eta_0(0)|Y_{-1},W]=Y_{-1}$ and $E[Y_{-1}|W=0] \neq E[Y_{-1}|W=1]$. Under this DGP, we have $E[Y_1(0)-Y_0(0)|Y_{-1},W=0]=Y_{-1}=E[Y_1(0)-Y_0(0)|Y_{-1},W=1]$, but $ E[Y_1(0)-Y_0(0)|W=0] = E[E[\eta_1(0)-\eta_0(0)|Y_{-1},W=0]|W=0] = E[Y_{-1}|W=0] \neq E[Y_{-1}|W=1] = E[E[\eta_1(0)-\eta_0(0)|Y_{-1},W=1]|W=1] = E[Y_1(0)-Y_0(0)|W=1]. $ In other words, Condition DIDM holds, but Condition DID fails. Intuitively, even if the parallel trends hold at every level of \(Y_{-1}\), differences in the distribution of \(Y_{-1}\) between treatment groups can lead to unequal unconditional trends.

commentConsider the following DGP for untreated potential outcomes: \begin{align*} Y_0(0) &= \mu(Y_{-1}) + \eta_0,\\[1mm] Y_1(0) &= \mu(Y_{-1}) + \eta_1, \end{align*} where \(Y_{-1} \in \{0,1\}\) denotes lagged employment status (0: unemployed, 1: employed) and \(\eta_0,\eta_1\) are disturbances with \[ E[\eta_1-\eta_0 \mid Y_{-1},W] = \delta(Y_{-1}). \] Thus, conditioning on \(Y_{-1}\) yields \[ E\bigl[Y_1(0)-Y_0(0) \mid Y_{-1},W=1\bigr] = E\bigl[Y_1(0)-Y_0(0) \mid Y_{-1},W=0\bigr] = \delta(Y_{-1}), \] so that the conditional parallel trends (DIDM) hold. However, if the distribution of \(Y_{-1}\) differs between treatment groups, the unconditional trend is a weighted average: \[ E\bigl[Y_1(0)-Y_0(0) \mid W\bigr] = \sum_{y\in\{0,1\}}P(Y_{-1}=y\mid W)\,\delta(y). \] Hence, when \[ P(Y_{-1}=1\mid W=1) \neq P(Y_{-1}=1\mid W=0), \] we obtain \[ E\bigl[Y_1(0)-Y_0(0) \mid W=1\bigr] \neq E\bigl[Y_1(0)-Y_0(0) \mid W=0\bigr], \] and the unconditional DID fails despite valid conditional comparisons.
comment\subsubsection{Unconditional DID Holds but Conditional DIDM Fails} Alternatively, consider a DGP where \begin{align*} Y_0(0) &= \mu(Y_{-1},W) + \eta_0,\\[1mm] Y_1(0) &= \mu(Y_{-1},W) + \eta_1, \end{align*} with the disturbance term satisfying \[ E[\eta_1-\eta_0 \mid Y_{-1}] = \delta, \] a constant independent of \(W\). In this case, the unconditional trends are \[ E\bigl[Y_1(0)-Y_0(0) \mid W=1\bigr] = \delta = E\bigl[Y_1(0)-Y_0(0) \mid W=0\bigr], \] so that the unconditional DID correctly recovers the average trend difference. However, because \(\mu(Y_{-1},W)\) depends on \(W\), it follows that \[ E\bigl[Y_1(0)-Y_0(0) \mid Y_{-1},W=1\bigr] \neq E\bigl[Y_1(0)-Y_0(0) \mid Y_{-1},W=0\bigr], \] implying that the conditional DIDM fails.
comment\subsubsection{Unconditional DID Fails but Conditional DIDM Holds} Suppose that untreated outcomes exhibit gender-specific trends. Let \[ E[Y_1(0)-Y_0(0)\mid \text{male},W=1] = E[Y_1(0)-Y_0(0)\mid \text{male},W=0] \quad \text{and} \quad E[Y_1(0)-Y_0(0)\mid \text{female},W=1] = E[Y_1(0)-Y_0(0)\mid \text{female},W=0]. \] Thus, within each gender the parallel trend condition holds, implying that the conditional parallel trends assumption (DIDM) is satisfied when one conditions on gender. Now, imagine that the treated group (\(W=1\)) is composed of 5 males and 15 females, while the control group (\(W=0\)) consists of 15 males and 5 females. Because the gender proportions differ, the unconditional averages—weighted combinations of the gender-specific trends—will generally differ: \[ {\text{E}}[Y_1(0)-Y_0(0)\mid W=1] \neq {\text{E}}[Y_1(0)-Y_0(0)\mid W=0]. \] Hence, the unconditional parallel trends assumption (DID) fails despite the fact that conditioning on gender restores parallel trends.

Condition DID Holds but Condition DIDM Fails

Conversely, consider the following DGP for untreated potential outcomes:

align*[align* omitted — 85 chars of source]

where $E[\eta_1(0)-\eta_0(0)|Y_{-1},W] = (2W-1)Y_{-1}$ and $E[Y_{-1}|W=0]+E[Y_{-1}|W=1]=0$. Under this DGP, $ E[Y_1(0)-Y_0(0)|W=0] = E[E[Y_1(0)-Y_0(0)|Y_{-1},W=0]|W=0] = E[-Y_{-1}|W=0] = E[Y_{-1}|W=1] = E[E[Y_1(0)-Y_0(0)|Y_{-1},W=1]|W=1] = E[Y_1(0)-Y_0(0)|W=1], $ but $E[Y_1(0)-Y_0(0)|Y_{-1},W=0] = -Y_{-1} \neq Y_{-1} = E[Y_1(0)-Y_0(0)|Y_{-1},W=1].$ In other words, Condition DID holds but Condition DIDM fails. Intuitively, even if the overall (unconditional) trends are equal (due to cancellation when aggregating over \(Y_{-1}\)), the trends might differ for each subpopulation defined by \(Y_{-1}\).

commentConversely, consider a setting where untreated outcome trends differ by gender in opposite directions between treated and control groups. For example, suppose that for males \[ E[Y_1(0)-Y_0(0)\mid \text{male},W=1] = \delta_{\text{male}}^{(1)} \quad \text{and} \quad E[Y_1(0)-Y_0(0)\mid \text{male},W=0] = \delta_{\text{male}}^{(0)}, \] with \(\delta_{\text{male}}^{(1)} < \delta_{\text{male}}^{(0)}\), and for females \[ E[Y_1(0)-Y_0(0)\mid \text{female},W=1] = \delta_{\text{female}}^{(1)} \quad \text{and} \quad E[Y_1(0)-Y_0(0)\mid \text{female},W=0] = \delta_{\text{female}}^{(0)}, \] with \(\delta_{\text{female}}^{(1)} > \delta_{\text{female}}^{(0)}\). Now, if the proportions of males and females in the treated and control groups are such that the weighted averages of these gender-specific trends coincide, then one obtains \[ {\text{E}}[Y_1(0)-Y_0(0)\mid W=1] = {\text{E}}[Y_1(0)-Y_0(0)\mid W=0], \] so that the unconditional parallel trends (DID) hold by cancellation. However, within each gender the trends differ between treated and control groups, meaning the conditional parallel trends assumption is violated: \[ E[Y_1(0)-Y_0(0)\mid \text{gender},W=1] \neq E[Y_1(0)-Y_0(0)\mid \text{gender},W=0]. \] Thus, while the DID estimand may fortuitously recover the ATT at the aggregate level, the DIDM estimand—which requires valid conditional comparisons—would be biased.

\noindentImplication for Practice.\\ These examples underscore that each assumption addresses a different facet of the DGP. Practitioners must ex ante commit to one of these assumptions based on the context and the nature of available data. Whether one relies on matching (Condition M), unconditional parallel trends (DID), or conditional parallel trends (DIDM) is not a matter of nested robustness but of fundamentally different identifying restrictions, each carrying its own trade-offs and implications for causal inference.

}

The Lagged Outcome $Y_{-s}$ as A Key Matching Criterion

Following heckman1998characterizing and smith2005does, we place particular emphasis on lagged outcomes, like $Y_{-s}$, as the conditioning variable for M and DIDM. This focus is motivated by several key studies in the literature.

Firstly, ashenfelter1978estimating argues that participants in job training programs typically have permanently lower earnings than non-participants, but also experience a temporary decrease in earnings just before entering the program, a phenomenon now famously known as the “Ashenfelter dip.” Controlling for lagged outcomes is crucial to addressing this confounding factor, as also emphasized by heckman1999economics among others.

In lalonde1986evaluating and dehejia1999causal,dehejia2002propensity, indeed, including lagged outcomes in matching allowed for close replication of the baseline experimental estimates. The subsequent discourse between smith2005does and dehejia2005practical further underscored the significance of conditioning on lagged outcomes.

angrist2009mostly discuss the importance of conditioning on lagged outcomes in the context of lagged-dependent-variable (LDV) models. More recently, roth2023parallel, in their review of difference-in-differences methods, raise an open question about the validity of conditioning on lagged outcomes and whether it reduces or exacerbates bias.

Finally, influential empirical work has emphasized the special role that lagged outcomes play in reducing bias in policy evaluation. For example, chetty2014measuring1, in their work on value-added models in education, highlight the utility of lagged outcomes, specifically prior test scores, as essential covariates for obtaining unbiased value-added estimates.

These studies suggest the importance of focusing on lagged outcomes as a key conditioning variable and analyzing the sources of bias that may arise under different conditions. We thus follow this literature to focus on the lagged outcome $Y_{-s}$ for matching in the current section. With this said, Section (ref) will generalize this setup to include a general class of matching criteria in addition to lagged outcomes.

{\color{black} Finally, we briefly discuss the lag order, denoted as $-s$. In the special case where $s=0$, the DIDM estimand ${\theta_{\text{ATT}}^{\text{DIDM}}}$ simplifies to the M estimand ${\theta_{\text{ATT}}^{\text{M}}}$. As a result, the double bracketing relation ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$ that we present later reduces to ${\theta_{\text{ATT}}^{\text{M}}} = {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, which corresponds to the existing bracketing result, LDV $\leq$ FE, as demonstrated by angrist2009mostly and ding2019bracketing. Our framework, which accommodates $s \geq 0$, is more general and particularly includes the symmetric DID approach introduced by heckman1998characterizing and smith2005does. This approach, which aligns with our DIDM with $-s=-1$, was shown by heckman1998characterizing and subsequent studies to be more accurate than other specifications. Therefore, our arbitrary lag order $-s$ not only encompasses the classical bracketing scenario as a special case but also includes many important cases explored in foundational papers within the literature. }

Empirics of the Double Bracketing

This section revisits three empirical studies using four datasets from the literature, parts of which are derived from the seminal papers discussed in Section (ref). We are going to document the robustness of the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, across each of these scenarios, different data sets, estimation methods, and subpopulations defined by observed characteristics. The first two examples, presented in Section (ref), focus on job training programs, namely the NSW and JTPA. The final example, in Section (ref), examines educational programs.

Job Training Programs

The NSW Program

We begin our analysis with a reexamination of the National Supported Work (NSW) programs lalonde1986evaluating,dehejia1999causal,dehejia2002propensity,smith2005does through the lens of our double bracketing. We employ the same sets of the CPS and PSID data as those utilized by lalonde1986evaluating and smith2005does. Although these data sets have been extensively analyzed in the literature, we offer a brief description in Appendix (ref) for the sake of completeness and convenience of the readers.

The primary outcome variable, $Y_t$, represents participants' self-reported earnings, adjusted to 1982 dollars. The treatment variable, $W$, is a binary indicator denoting whether an individual was assigned to the NSW program. Additionally, demographic variables, including race and educational level, are included as auxiliary covariates in all M, DIDM, and DID estimations.

Let us revisit Figure (ref), which was initially introduced in Section (ref) to motivate our investigation. Recall that Figure (ref) illustrates the signed biases of the estimates for the three alternative estimands, ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DIDM}}}$, and ${\theta_{\text{ATT}}^{\text{DID}}}$, relative to the experimental estimates in percentage terms according to heckman1998characterizing and smith2005does.

In the current section, we focus on the black bars in Figure (ref), representing the estimates by smith2005does. Notably, the double bracketing inequality, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, is upheld when considering their point estimates. Since these represent the biases, note that ${\theta_{\text{ATT}}^{\text{M}}}$ and ${\theta_{\text{ATT}}^{\text{DIDM}}}$ are biased downward, whereas ${\theta_{\text{ATT}}^{\text{DID}}}$ is biased upward. These results suggest that the M estimand, ${\theta_{\text{ATT}}^{\text{M}}}$, tends to be conservative, while the DID estimand, ${\theta_{\text{ATT}}^{\text{DID}}}$, appears to be optimistic, assuming that the experimental estimates are the true values.

The observation above applies specifically to the estimates of ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DIDM}}}$, and ${\theta_{\text{ATT}}^{\text{DID}}}$ selected by chabe2017should from the various ST specifications. However, when we extend our analysis to include iterative estimations across all nine estimation methods employed by smith2005does using the CPS and PSID datasets, we find that the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, is robustly upheld across the vast majority of these specifications. This is illustrated in Figure (ref), where the top and bottom charts display the results for the CPS and PSID data, respectively.

figure[figure omitted — 823 chars of source]

Moreover, it is important to note that this ordering of estimates is never rejected for any of the nine specifications. This consistency ensures that the result is not an artifact of a specific estimation procedure. Consistently, we observe that the M estimand, ${\theta_{\text{ATT}}^{\text{M}}}$, tends to be conservative, while the DID estimand, ${\theta_{\text{ATT}}^{\text{DID}}}$, tends to be optimistic.

The JTPA Program

Next, we reexamine the Job Training Partnership Act (JTPA) program studied by heckman1998characterizing. Although the dataset they used has been extensively analyzed in the literature, a brief description is provided in Appendix (ref) for the sake of completeness and convenience of the readers.

The outcome variable, $Y_t$, represents participants' earnings adjusted for inflation. The treatment variable, $W$, indicates whether an individual was assigned to the JTPA program. Additionally, demographic covariates such as sex and age are included in all M, DIDM, and DID estimations.

Once again, we reiterate Figure (ref) that we provide in Section (ref). For the current section about the JTPA programs, focus on the gray bars in Figure (ref) based on the estimates by heckman1998characterizing. Notice that the inequality relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, is maintained when considering their point estimates. Specifically, the M estimand, ${\theta_{\text{ATT}}^{\text{M}}}$, is biased downward, while the DID estimand, ${\theta_{\text{ATT}}^{\text{DID}}}$, is biased upward.

Educational Programs

Now, we turn to the educational programs analyzed by athey2020combining, which were also studied in chetty2014measuring1, chetty2014measuring2. In their study, athey2020combining addressed the challenge of selection in observational by combining it with short-term experimental data. In contrast, our analysis focuses solely on observational data to examine the behavior of the observational estimands: ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DIDM}}}$, and ${\theta_{\text{ATT}}^{\text{DID}}}$. Details of the data are provided in Appendix (ref).

The outcome variable, $Y_t$, represents students' test scores, specifically standardized scores that average results from mathematics and English language arts. The treatment variable, $W$, indicates whether a student was assigned to a small class size. Additionally, covariates such as gender, race, and eligibility for free lunch are considered to define subpopulations in our analysis.

Given the large size of our dataset (see Appendix (ref) for details), we can precisely estimate the M, DIDM, and DID estimands even when the sample is divided into subpopulations based on observed attributes. Figure (ref) presents these estimates, along with their 95% confidence intervals, for each subpopulation.

figure[figure omitted — 343 chars of source]

Observe that the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, remains consistent across various subpopulations, despite notable differences in levels among these groups. Interestingly, the M estimand, ${\theta_{\text{ATT}}^{\text{M}}}$, generally indicates negative effects, except within the Black subpopulations. In contrast, both the DIDM estimand, ${\theta_{\text{ATT}}^{\text{DIDM}}}$, and the DID estimand, ${\theta_{\text{ATT}}^{\text{DID}}}$, typically suggest positive effects. This divergence in implications between estimation methods raises questions about relying solely on observational data, highlighting the value of combining experimental and observational data, as demonstrated by athey2020combining.

comment{\color{red} \section{Theory of the Double Bracketing: A Parametric Approach} We begin by showing how the double bracketing relationship, \[ {\theta_{\text{ATT}}^{\text{M}}} \;\leq\; {\theta_{\text{ATT}}^{\text{DIDM}}} \;\leq\; {\theta_{\text{ATT}}^{\text{DID}}}, \] arises under a familiar parametric setting. In particular, suppose that the true data-generating process takes the form: \begin{equation} Y_{i,t} \;=\; \alpha_i \;+\; \beta \,W_i \,\mathbbm{1}\{t \geq 1\} \;+\; \gamma \,W_i \;+\; \delta_t \;+\; \rho\,Y_{i,t-1} \;+\; \epsilon_{i,t}, \end{equation} where the error term satisfies \begin{equation} E\bigl[\epsilon_{i,1}\,\big|\,Y_{i,-1},W_i\bigr] \;=\; E\bigl[\epsilon_{i,0}\,\big|\,Y_{i,-1},W_i\bigr] \;=\; 0. \end{equation} Here, $\alpha_i$ and $\delta_t$ act as two-way fixed effects, while $\rho$ captures persistence in outcomes. In this framework, the parameter $\gamma$ can be interpreted as the difference \[ \gamma \;=\; E\bigl[Y_{i,0}\,\big|\,W_i=1,Y_{i,-1}\bigr] \;-\; E\bigl[Y_{i,0}\,\big|\,W_i=0,Y_{i,-1}\bigr]. \] Negative-selection arguments for many training or educational programs suggest that the treated population ($W_i=1$) often has weakly lower baseline outcomes. We formalize this with: \begin{equation} Negative Selection I: \quad \gamma \;=\; E\bigl[Y_{i,0}\,\big|\,W_i=1,Y_{i,-1}\bigr] \;-\; E\bigl[Y_{i,0}\,\big|\,W_i=0,Y_{i,-1}\bigr] \;\;\;\leq\;\; 0. \end{equation} Similarly, we impose a second negative-selection restriction based on the prior outcome $Y_{i,-1}$: \begin{equation} Negative Selection II: \quad E\bigl[Y_{i,-1}\,\big|\,W_i=1\bigr] \;\;\leq\;\; E\bigl[Y_{i,-1}\,\big|\,W_i=0\bigr]. \end{equation} Finally, we assume the underlying process is not explosive: \begin{equation} Non-Explosive Process: \quad 0 \;\;\leq\;\; \rho \;\;\leq\;\; 1. \end{equation} Under (ref)--(ref), one can derive explicit forms for the three estimands: \begin{align} {\theta_{ATT}^{M}} &=\; \beta + (1+\rho)\,\gamma, \\[6pt] {\theta_{\text{ATT}}^{\text{DIDM}}} &=\; \beta + \rho\,\gamma, \label{eq:parametric:didm} \\[6pt] {\theta_{\text{ATT}}^{\text{DID}}} &=\; \beta + \rho\,\gamma \;+\; \rho\,(1-\rho)\, \Bigl( E[Y_{i,-1}\,|\,W_i=0] \;-\; E[Y_{i,-1}\,|\,W_i=1] \Bigr). \label{eq:parametric:did} \end{align} A direct calculation gives \begin{equation} {\theta_{\text{ATT}}^{\text{DIDM}}} \;-\; {\theta_{\text{ATT}}^{\text{M}}} \;=\; -\,\gamma \;\;\geq\;\; 0, \quad\text{and}\quad {\theta_{\text{ATT}}^{\text{DID}}} \;-\; {\theta_{\text{ATT}}^{\text{DIDM}}} \;=\; \rho\,(1-\rho)\, \Bigl( E[Y_{i,-1}\,|\,W_i=0] \;-\; E[Y_{i,-1}\,|\,W_i=1] \Bigr) \;\;\geq\;\; 0, \end{equation} where the two inequalities follow from \eqref{eq:parametric:negative_selection_1}--\eqref{eq:parametric:non_explosive} and \eqref{eq:parametric:negative_selection_2}--\eqref{eq:parametric:non_explosive}, respectively. Hence, \[ {\theta_{\text{ATT}}^{\text{M}}} \;\;\;\leq\;\;\; {\theta_{\text{ATT}}^{\text{DIDM}}} \;\;\;\leq\;\;\; {\theta_{\text{ATT}}^{\text{DID}}}, \] which is precisely the double bracketing. This result holds whenever the plausible conditions (ref)--(ref) are satisfied, aligning with the discussion in angrist2009mostly. \paragraph{Special Cases.} We also see two immediate “collapses” of the triple into a double: \begin{itemize} • If $\gamma = 0$ (i.e., no negative selection on $Y_{i,0}$), then \[ {\theta_{\text{ATT}}^{\text{M}}} \;=\; {\theta_{\text{ATT}}^{\text{DIDM}}} \;\;\;\leq\;\;\; {\theta_{\text{ATT}}^{\text{DID}}}. \] • If $\rho=0$ or $\rho=1$, then \[ {\theta_{\text{ATT}}^{\text{M}}} \;\;\;\leq\;\;\; {\theta_{\text{ATT}}^{\text{DIDM}}} \;=\; {\theta_{\text{ATT}}^{\text{DID}}}. \] \end{itemize} Hence, in those extreme cases, DIDM equals one of the other estimands and thus becomes redundant. In all other situations, however, these three approaches provide distinct bounds. \section{Theory of the Double Bracketing: Nonparametric Framework} While the above derivation relies on a linear parametric model (ref), we now show that the core double bracketing relationship, \[ {\theta_{\text{ATT}}^{\text{M}}} \;\leq\; {\theta_{\text{ATT}}^{\text{DIDM}}} \;\leq\; {\theta_{\text{ATT}}^{\text{DID}}}, \] actually extends to much more general settings without imposing linearity or additivity. To highlight this, we adopt a nonparametric model-free framework, allowing for arbitrary functional forms and distributions subject only to minimal, empirically testable conditions. Let \begin{align*} \Delta({\theta_{ATT}^{\text{M}}}) &\;=\; {\theta_{\text{ATT}}^{\text{M}}}-{\theta_{\text{ATT}}}, \\ \Delta({\theta_{\text{ATT}}^{\text{DID}}}) &\;=\; {\theta_{\text{ATT}}^{\text{DID}}}-{\theta_{\text{ATT}}}, \\ \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) &\;=\; {\theta_{\text{ATT}}^{\text{DIDM}}}-{\theta_{\text{ATT}}}. \end{align*} Then the double bracketing relation, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}},$ is equivalent to \[ \Delta({\theta_{\text{ATT}}^{\text{M}}}) \;\;\leq\;\; \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \;\;\leq\;\; \Delta({\theta_{\text{ATT}}^{\text{DID}}}). \] We establish this in full generality under the following assumptions. \begin{assumption} \textbf{(Nonparametric Negative Selection & Stability)} \begin{enumerate}[(i)] • $E[Y_0 \mid W=0,Y_{-s}=y] \;\;\geq\;\; E[Y_0 \mid W=1,Y_{-s}=y]$ for all $y$. • $F_{Y_{-s}\mid W=0}$ first-order stochastically dominates $F_{Y_{-s}\mid W=1}$. • $y \mapsto \Phi(y) := E[Y_1 - Y_0 \mid W=0,Y_{-s}=y]$ is weakly decreasing in $y$. \end{enumerate} \end{assumption} Each of the three components, (ref)--(ref), of Assumption (ref) is empirically testable, as it does not involve unobserved latent variables such as $Y_t(d)$. Furthermore, they are plausible and align with the assumptions made in the literature angrist2009mostly,ding2019bracketing. Specifically, part (ref) requires the so-called “negative selection” that individuals who opt out from treatment (i.e., those with $W=0$) tend to have weakly higher pre-treatment (potential) outcomes $Y_0=Y_0(0)$ without treatment on average given $Y_{-s}=y$. Similarly, part (ref) requires the negative selection that individuals who opt out from treatment (i.e., those with $W=0$) tend to have no lower pre-treatment (potential) outcome $Y_{-s}=Y_{-s}(0)$ without treatment. Part (ref) requires the time series of $\{Y_t\}_t$ of outcomes for those with $W=0$ (i.e., $\{Y_t\}_t = \{Y_t(0)\}_t$) is stable over time. In real-world settings, especially among populations with limited socioeconomic status (SES), it may be reasonable to expect that outcomes do not escalate dramatically without intervention. For example, in educational programs targeting low-SES individuals, we wouldn't anticipate rapid, unsustainable improvements in outcomes without structured support. In the special case where $s=0$, this requirement essentially reflects the condition that the root of the AR(1) process is less than one, which is equivalent to stationarity, as deployed in, e.g., chetty2014measuring1. Each of these parts is analogous to the assumption made by angrist2009mostly in deriving their bracketing relationship between the DID and lagged dependent variable (LDV) in linear and additive models. In the special case of $s=0$, part (ref) will be trivially satisfied while parts (ref)--(ref) correspond to the assumptions invoked by ding2019bracketing. As argued at the end of Section (ref), our framework with $s \geq 0$ accommodates both this special case and other important scenarios considered by heckman1998characterizing and others. \begin{theorem} Suppose that Assumption (ref) holds. Then, we have the bracketing relationship \begin{align*} \Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}}). \end{align*} \end{theorem} \begin{proof}[Proof of Theorem (ref)] Note that the identification errors can be written as \begin{align*} \Delta\left({\theta_{\text{ATT}}^{\text{M}}}\right)= & E\left[E\left[Y_1(0) \mid W=1, Y_{-s}\right]-E\left[Y_1(0) \mid W=0, Y_{-s}\right] \mid W=1\right], \\ \Delta\left({\theta_{\text{ATT}}^{\text{DID}}}\right)= & E\left[Y_1(0)-Y_0(0) \mid W=1\right]-E\left[Y_1(0)-Y_0(0) \mid W=0\right], \qquad\text{and} \\ \Delta\left({\theta_{\text{ATT}}^{\text{DIDM}}}\right)= & E\left[E\left[Y_1(0)-Y_0(0) \mid W=1, Y_{-s}\right]-E\left[Y_1(0)-Y_0(0) \mid W=0, Y_{-s}\right] \mid W=1\right]. \end{align*} First, observe that \begin{align*} \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) - \Delta({\theta_{\text{ATT}}^{\text{M}}}) =& E[E[Y_0(0)|W=0,Y_{-s}] - E[Y_0(0)|W=1,Y_{-s}]|W=1] \\ =& E[E[Y_0|W=0,Y_{-s}] - E[Y_0|W=1,Y_{-s}]|W=1] \geq 0, \end{align*} where the second equality is due to $Y_t(0)=Y_t$ for all $t \leq 0$, and the last inequality follows from Assumption (ref) (ref). Second, observe that \begin{align*} &\Delta({\theta_{\text{ATT}}^{\text{DID}}}) - \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \\ =& E[E[Y_1(0)-Y_0(0)|W=0,Y_{-s}] | W=1] - E[Y_1(0)-Y_0(0)|W=0] \\ =& E[E[Y_1(0)-Y_0(0)|W=0,Y_{-s}] | W=1] - E[E[Y_1(0)-Y_0(0)|W=0,Y_{-s}] | W=0] \\ =& E[E[Y_1-Y_0|W=0,Y_{-s}] | W=1] - E[E[Y_1-Y_0|W=0,Y_{-s}] | W=0] \\ =& E[\Phi(Y_{-s}) | W=1] - E[\Phi(Y_{-s}) | W=0] \geq 0, \end{align*} where the second equality follows from the law of iterated expectations, the third equality is due to $Y_t(0)=Y_t$ given $W=0$, and the last inequality follows from Assumption (ref) (ref)--(ref). \end{proof} This theorem provides a theoretical explanation for the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, which was robustly observed in Section (ref). Furthermore, this theorem implies the following three consequences depending on the underlying DGP. First, when Condition M is true, then \begin{align*} 0 = \Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}}). \end{align*} In this case, ${\theta_{\text{ATT}}^{\text{M}}}$ identifies ${\theta_{\text{ATT}}}$ but ${\theta_{\text{ATT}}^{\text{DIDM}}}$ and ${\theta_{\text{ATT}}^{\text{DID}}}$ tend to be upwardly biased. Second, when Condition DID is true, then \begin{align*} \Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}}) = 0. \end{align*} In this case, ${\theta_{\text{ATT}}^{\text{DID}}}$ identifies ${\theta_{\text{ATT}}}$ but ${\theta_{\text{ATT}}^{\text{M}}}$ and ${\theta_{\text{ATT}}}$ tend to be downwardly biased. Finally, when Condition DIDM is true, then \begin{align*} \Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) = 0 \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}}). \end{align*} In this case, ${\theta_{\text{ATT}}^{\text{DIDM}}}$ identifies ${\theta_{\text{ATT}}}$ but ${\theta_{\text{ATT}}^{\text{M}}}$ tends to be downwardly biased while ${\theta_{\text{ATT}}^{\text{DID}}}$ tends to be upwardly biased. Hence, for applications where non-negative treatment effects are expected, the DID estimand will generally incur optimistic estimates while the M estimand will produce conservative estimates. The DIDM estimand can be optimistic or conservative, depending on the underlying DGP. In summary, the M estimand is conservatively the most robust. }

The Double Bracketing: A Parametric Approach

So far, we have observed that the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, appears to be more of a rule than a mere coincidence in empirical studies on educational and job training programs. In this section, and the two subsequent ones, we are going to provide theoretical explanations for this intriguing phenomenon.

While general theories will be presented in the subsequent sections, we begin by illustrating how the double bracketing relationship, $ {\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}} $ arises under a familiar parametric setting. In particular, suppose the true data-generating process takes the form:

equation[equation omitted — 151 chars of source]

where the error term satisfies

equation[equation omitted — 115 chars of source]

In this context, $\alpha_i$ and $\delta_t$ serve as two-way fixed effects and $\rho$ serves as a persistence parameter.

Note that $\gamma$ can be expressed as $ \gamma = E[Y_{i,0}|W_i=1,Y_{i,-1}] - E[Y_{i,0}|W_i=0,Y_{i,-1}] $ under (ref)--(ref). This value, $\gamma$, is non-positive if individuals with weakly lower $Y_{i,0}$ tend to select into treatment $W_i=1$ conditional on $Y_{i,-1}$ as is plausibly the case with job training programs. We postulate this negative selection assumption:

equation[equation omitted — 164 chars of source]

We also postulate a similar negative selection assumption for $Y_{i,-1}$:

equation[equation omitted — 135 chars of source]

Finally, we impose the assumption of a non-explosive earnings process:

equation[equation omitted — 107 chars of source]

Note that these three conditions (ref)--(ref) are analogous to the assumptions invoked by angrist2009mostly in developing their bracketing relationship, LDV $\leq$ FE.

Under the parametric data-generating process (ref), some calculations yield

align[align omitted — 350 chars of source]

See Appendix (ref) for detailed calculations that derive these expressions. From (ref)--(ref), we have

equation[equation omitted — 130 chars of source]

where the inequality follows from $\gamma \leq 0$ due to (ref). From (ref)--(ref), we have

equation[equation omitted — 179 chars of source]

where the inequality follows from $E[Y_{i,-1}|W_i=0] - E[Y_{i,-1}|W_i=1] \geq 0$ due to (ref) and $\rho \in [0,1]$ due to (ref).

Combining (ref)--(ref) yields the double bracketing relationship,

align*[align* omitted — 128 chars of source]

which holds true under the simple parametric model (ref) with the plausible restrictions (ref)--(ref) motivated by angrist2009mostly.

Finally, we highlight a couple of special cases in which the three estimands collapse into two. In the absence of the first negative selection (i.e., $\gamma = 0$) the double bracketing relationship reduces to

align*[align* omitted — 125 chars of source]

In other words, M and DIDM become equivalent. Similarly, in the absence of persistence (i.d., $\rho=0$) or under the unit root (i.e., $\rho=1$), the double bracketing relationship reduces to

align*[align* omitted — 125 chars of source]

In other words, DIDM and DID become equivalent. In these two special cases, DIDM indeed becomes redundant. However, it is generally distinct from the other two estimands otherwise.

The Double Bracketing in Nonparametric Frameworks

The previous section derived the double-bracketing relationship ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$ under a simple parametric model framework. To ensure generality, we adopt a non-parametric and model-free framework in the current section.

Let

align*[align* omitted — 345 chars of source]

be the identification errors of the estimands, ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DID}}}$, and ${\theta_{\text{ATT}}^{\text{DIDM}}}$, respectively. Observe that ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$ is equivalent to $\Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}})$ with these notations. Thus, we are going to establish the double bracketing relation, $\Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}})$, in terms of the identification error. To this end, we consider the following assumption.

assumptionThe following conditions hold. \begin{enumerate}[(i)] • $E[Y_0|W=0,Y_{-s}=y] \geq E[Y_0|W=1,Y_{-s}=y]$ for all $y$. • $F_{Y_{-s}|W=0}$ first-order stochastically dominates $F_{Y_{-s}|W=1}$. • $y \mapsto \Phi(y) := E[Y_1-Y_0|W=0,Y_{-s}=y]$ is weakly decreasing. \end{enumerate}

{\color{black} The three components, (ref)--(ref), of Assumption (ref) are analogous to (ref)--(ref), respectively. Each of the three components is empirically testable, as it does not involve unobserved latent variables such as $Y_t(d)$. Furthermore, they are plausible and align with the assumptions made in the literature angrist2009mostly,ding2019bracketing. Specifically, part (ref) requires the so-called “negative selection” that individuals who opt out from treatment (i.e., those with $W=0$) tend to have weakly higher pre-treatment (potential) outcomes $Y_0=Y_0(0)$ without treatment on average given $Y_{-s}=y$. Similarly, part (ref) requires the negative selection that individuals who opt out from treatment (i.e., those with $W=0$) tend to have no lower pre-treatment (potential) outcome $Y_{-s}=Y_{-s}(0)$ without treatment. Part (ref) requires the time series $\{Y_t\}_t$ of outcomes for those with $W=0$ (i.e., $\{Y_t\}_t = \{Y_t(0)\}_t$) is stable over time. In real-world settings, especially among populations with limited socioeconomic status (SES), it may be reasonable to expect that outcomes do not escalate dramatically without intervention. For example, in educational programs targeting low-SES individuals, we wouldn't anticipate rapid, unsustainable improvements in outcomes without structured support on average. In the special case where $s=0$, this requirement essentially reflects the condition that the root of the AR(1) process is less than one, which is equivalent to stationarity, as deployed in, e.g., chetty2014measuring1.

Each of these parts is analogous to the assumption made by angrist2009mostly in deriving their bracketing relationship between the DID and lagged dependent variable (LDV) in linear and additive models. In the special case of $s=0$, part (ref) will be trivially satisfied while parts (ref)--(ref) correspond to the assumptions invoked by ding2019bracketing. As argued at the end of Section (ref), our framework with $s \geq 0$ accommodates both this special case and other important scenarios considered by heckman1998characterizing and others. }

theoremSuppose that Assumption (ref) holds. Then, we have the bracketing relationship \begin{align*} \Delta({\theta_{ATT}^{M}}) \leq \Delta({\theta_{ATT}^{DIDM}}) \leq \Delta({\theta_{ATT}^{DID}}). \end{align*}
proof[Proof of Theorem (ref)] Note that the identification errors can be written as \begin{align*} \Delta\left({\theta_{ATT}^{M}}\right)= & E\left[E\left[Y_1(0) \mid W=1, Y_{-s}\right]-E\left[Y_1(0) \mid W=0, Y_{-s}\right] \mid W=1\right], \\ \Delta\left({\theta_{ATT}^{DID}}\right)= & E\left[Y_1(0)-Y_0(0) \mid W=1\right]-E\left[Y_1(0)-Y_0(0) \mid W=0\right], \qquadand \\ \Delta\left({\theta_{ATT}^{\text{DIDM}}}\right)= & E\left[E\left[Y_1(0)-Y_0(0) \mid W=1, Y_{-s}\right]-E\left[Y_1(0)-Y_0(0) \mid W=0, Y_{-s}\right] \mid W=1\right]. \end{align*} First, observe that \begin{align*} \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) - \Delta({\theta_{\text{ATT}}^{\text{M}}}) =& E[E[Y_0(0)|W=0,Y_{-s}] - E[Y_0(0)|W=1,Y_{-s}]|W=1] \\ =& E[E[Y_0|W=0,Y_{-s}] - E[Y_0|W=1,Y_{-s}]|W=1] \geq 0, \end{align*} where the second equality is due to $Y_t(0)=Y_t$ for all $t \leq 0$, and the last inequality follows from Assumption (ref) (ref). Second, observe that \begin{align*} &\Delta({\theta_{\text{ATT}}^{\text{DID}}}) - \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \\ =& E[E[Y_1(0)-Y_0(0)|W=0,Y_{-s}] | W=1] - E[Y_1(0)-Y_0(0)|W=0] \\ =& E[E[Y_1(0)-Y_0(0)|W=0,Y_{-s}] | W=1] - E[E[Y_1(0)-Y_0(0)|W=0,Y_{-s}] | W=0] \\ =& E[E[Y_1-Y_0|W=0,Y_{-s}] | W=1] - E[E[Y_1-Y_0|W=0,Y_{-s}] | W=0] \\ =& E[\Phi(Y_{-s}) | W=1] - E[\Phi(Y_{-s}) | W=0] \geq 0, \end{align*} where the second equality follows from the law of iterated expectations, the third equality is due to $Y_t(0)=Y_t$ given $W=0$, and the last inequality follows from Assumption (ref) (ref)--(ref).

This theorem provides a theoretical explanation for the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, which was robustly observed in Section (ref).

Furthermore, this theorem implies the following three consequences depending on the underlying DGP. First, when Condition M is true, then

align*[align* omitted — 156 chars of source]

In this case, ${\theta_{\text{ATT}}^{\text{M}}}$ identifies ${\theta_{\text{ATT}}}$ but ${\theta_{\text{ATT}}^{\text{DIDM}}}$ and ${\theta_{\text{ATT}}^{\text{DID}}}$ tend to be upwardly biased. Second, when Condition DID is true, then

align*[align* omitted — 156 chars of source]

In this case, ${\theta_{\text{ATT}}^{\text{DID}}}$ identifies ${\theta_{\text{ATT}}}$ but ${\theta_{\text{ATT}}^{\text{M}}}$ and ${\theta_{\text{ATT}}}$ tend to be downwardly biased. Finally, when Condition DIDM is true, then

align*[align* omitted — 156 chars of source]

In this case, ${\theta_{\text{ATT}}^{\text{DIDM}}}$ identifies ${\theta_{\text{ATT}}}$ but ${\theta_{\text{ATT}}^{\text{M}}}$ tends to be downwardly biased while ${\theta_{\text{ATT}}^{\text{DID}}}$ tends to be upwardly biased.

Hence, for applications where non-negative treatment effects are expected, the DID estimand will generally incur optimistic estimates while the M estimand will produce conservative estimates. The DIDM estimand can be optimistic or conservative, depending on the underlying DGP. In summary, the M estimand is conservatively the most robust.

Extension to General Cases

The main theoretical result presented in Section (ref) extends to a more general class of data-generating processes and associated estimands. In this section, we present an extension to general cases with a focus on the recent developments in the methods of event studies as principal examples.

Setup

The previous notations do not carry over to the current section. Suppose that a researcher is interested in identifying the average treatment effect on the treated (ATT) defined by

align[align omitted — 130 chars of source]

At this moment, we have not introduced the specific meanings of the notations. They will be discussed in the contexts of specific examples in Section (ref). With this said, we want to remark that they parallel with those notations introduced in Section (ref). Unlike the previous section, however, the subscripts no longer indicate the time in general.

Similarly to the previous section, we define the alternative estimands

align[align omitted — 656 chars of source]

called the matching (M), the difference-in-differences (DID), and the difference-in-differences matching (DIDM), respectively. The $p$-dimensional random vector $X$ is now used as a matching criterion.

The following conditions are imposed:

align[align omitted — 240 chars of source]

Condition (ref) requires that observed outcomes be the potential outcome without treatment for every unit prior to treatment. Condition (ref) requires that the observed outcome be the potential outcome without treatment for the control group.

Examples of the M, DID, and DIDM Estimands in Event Studies

In this section, we demonstrate that our general framework (ref)--(ref) encompasses alternative estimands studied in the literature of event studies as examples.

Example 1: M in Event Studies

acemoglu2019democracy consider the ATT

equation[equation omitted — 93 chars of source]

where $D_t$ denotes the indicator of democracy, $Y_t^s(1)$ denotes the potential GDP in period $t+s$ when a country is treated between periods $t-1$ and $t$ (i.e., $D_t=1$ and $D_{t-1}=0$), and $Y_t^s(0)$ denotes the potential GDP in period $t+s$ when such a treatment does not occur (i.e., $D_t=D_{t-1}=0$).\footnote{The original paper by acemoglu2019democracy considers $ E\left[ (Y_t^s(1)-Y_{t-1}) - (Y_t^s(0)-Y_{t-1}) | D_t=1, D_{t-1}=0\right] $ as the parameter of interest, where $Y_{t-1}$ denotes the realized GDP at period $t$, but this is equivalent to (ref).} acemoglu2019democracy identify this ATT by

equation[equation omitted — 224 chars of source]

where $Y_t^s$ denotes the observed GDP in period $t+s$, $Y_{t-1}$ denotes the observed GDP in period $t-1$, and $X := (Y_{t-1},\ldots,Y_{t-4})'$ in their baseline model with additional covariates in extended robustness analyses.

Since $X$ contains $Y_{t-1}$ in particular, the conditioning theorem\footnote{Specifically, the conditioning theorem yields $E[Y_{t-1}|D_t=0,D_{t-1}=0,X]=Y_{t-1}$ when $X$ contains $Y_{t-1}$.} cancels $Y_{t-1}$ between the two terms in (ref), so the identifying formula (ref) of acemoglu2019democracy boils down to

equation[equation omitted — 204 chars of source]

Our general framework encompasses this example. Specifically, our ATT (ref) reduces to (ref) and our M estimand (ref) reduces to (ref) by setting

align*[align* omitted — 204 chars of source]
align*[align* omitted — 178 chars of source]

The pre-treatment condition (ref) is satisfied by construction, as $\widetilde Y_0(0) = Y_{t-1} = \widetilde Y_0$. The comparison condition (ref) is also satisfied by construction via the definition of $Y_t^s(0)$ as the potential outcome under $D_t=D_{t-1}=0$. Namely, $\widetilde Y_1(0) = Y_t^s(0) = Y_t^s = \widetilde Y_1$ holds given $D_t=D_{t-1}=0$.

commentjorda2005estimation estimates the impulse response $\{{\text{E}}[Y_{t+h}(d)|Z_t]\}_h$ of a policy $D_t=d$ via the local projection \begin{equation*} E[Y_{t+h}|D_t,Z_t] = \beta_h D_t + \gamma_h' Z_t, \end{equation*} with a vector $Z_t$ of control variables. The effects of the policy $D_t=d$ with respect to a benchmark policy $D_t=d_0$ can be written by \begin{equation*} E[Y_{t+h}(d)|Z_t] - E[Y_{t+h}(d_0)|Z_t] = E[Y_{t+h}|D_t=d,Z_t] - E[Y_{t+h}|D_t=d_0,Z_t] = \beta_h (d-d_0), \end{equation*} where $\beta_h$ indicates the constant marginal policy effect. In a semi-parametric framework accommodating heterogeneous policy effects, angrist2018semiparametric consider the average treatment effect (ATE) \begin{equation*} E[ Y_{t+h}(d) - E[Y_{t+h}(d_0) ] = E[E[Y_{t+h}|D_t=d,Z_t] - E[Y_{t+h}|D_t=d_0,Z_t]] \end{equation*} by integrating the local projection difference with respect to the distribution of $Z_t$. More generally, if the integration is performed with respect to the conditional distribution of $Z_t$ given $D_t=d$, then one obtains \begin{equation} E[Y_{t+h}|D_t=d] - E[E[Y_{t+h}|D_t=d_0,Z_t]|D_t=d], \end{equation} which boils down to our M estimand (ref) for the ATT (ref), $E[ Y_{t+h}(d) - E[Y_{t+h}(d_0) |D_t=d]$, by setting \begin{align*} \widetilde Y_1(0) :=& Y_{t+h}(d_0), & \widetilde Y_1(1) :=& Y_{t+h}(d), & \widetilde Y_1 :=& Y_{t+h}, \end{align*} \begin{align*} and \qquad W := \begin{cases} 1 & if D_t=d \\ 0 & if D_t=d_0 \\ -1 & otherwise \end{cases} \end{align*} The pre-treatment condition (ref) is trivially satisfied since neither $\widetilde Y_0(0)$ nor $\widetilde Y_0$ is introduced for the M estimand. The comparison condition (ref) is also satisfied by construction, as $\widetilde Y_1(0) = Y_{t+h}(d_0) = Y_{t+h} = \widetilde Y_1$ given $D_t=d_0$.

Example 2: DID in Event Studies

callaway2018difference consider the ATT

equation[equation omitted — 94 chars of source]

where $G$ denotes the treatment period, $Y_t(g)$ denotes the potential outcome at period $t \geq g$ when an individual is treated at period $g$, and $Y_t(\infty)$ denotes the potential outcome at period $t$ when an individual does not receive a treatment. callaway2018difference identify this ATT by

equation[equation omitted — 140 chars of source]

for $g' \geq t+1$, where $Y_t$ denotes the observed outcome at period $t$. (The second term of (ref) may be aggregated over $G' \in \{t+1,t+2,\ldots\}$.)

Our general framework encompasses this example. Specifically, our ATT (ref) reduces to (ref) and our DID estimand (ref) reduces to (ref) by setting

align*[align* omitted — 214 chars of source]
align*[align* omitted — 129 chars of source]

The pre-treatment condition (ref) is satisfied by construction, as $\widetilde Y_0(0) = Y_{g-1}(\infty) = Y_{g-1} = \widetilde Y_0$ given $G=g$ or $G=g' \geq t+1 > g$. The comparison condition (ref) is also satisfied by construction, as $\widetilde Y_1(0) = Y_t(\infty) = Y_t = \widetilde Y_1$ given $G=g' \geq t+1$.

Example 3: DIDM in Event Studies

dube2023local consider the ATT

equation[equation omitted — 89 chars of source]

where $\Delta D_t$ denotes the indicator of policy change, $Y_{t+h}(1)$ denotes the potential outcome in period $t+h$ when a policy changes between periods $t-1$ and $t$ (i.e., $\Delta _t=1$), and $\Delta Y_{t+h}(0)$ denotes the potential outcome in period $t+h$ when such a change does not occur (i.e., $\Delta D_t=0$). dube2023local identify this ATT by

equation[equation omitted — 206 chars of source]

where $Y_{t+h}$ denotes the observed outcome in period $t+h$, $Y_{t-1}$ denotes the observed outcome in period $t-1$, and $X$ is a vector of general covariates.

Our general framework encompasses this example. Specifically, our ATT (ref) reduces to (ref) and our DIDM estimand (ref) reduces to (ref) by setting

align*[align* omitted — 210 chars of source]
align*[align* omitted — 146 chars of source]

The pre-treatment condition (ref) is satisfied by construction, as $\widetilde Y_0(0) = Y_{t-1} = \widetilde Y_0$. The comparison condition (ref) is also satisfied by construction via the definition of $Y_{t+h}(0)$ as the potential outcome under $\Delta D_t=0$. Namely, $\widetilde Y_1(0) = Y_{t+h}(0) = Y_{t+h} = \widetilde Y_1$ holds given $\Delta D_t=0$.

Also see the DID${}_\text{M}$ estimator of deChaisemartin2020two, and the (panel) matching estimator of imai2023matching, as well as the extended DID method of dube2023local -- they all propose and analyze the properties of what we refer to as the DIDM.

As pointed out by dube2023local, their framework encompasses acemoglu2019democracy as a special case. Indeed, when $X$ contains $\widetilde Y_0$, as is the case with acemoglu2019democracy presented in Section (ref), our DIDM framework reduces to our M framework. In general, however, the DIDM differs from the M.

comment\subsubsection{Example 3: DIDM in Event Studies} dube2023local consider the identification of \begin{equation} E\left[\left. Y_{t+h}(p) - Y_{t+h}(o) \right| P=p\right] \end{equation} by \begin{equation} E\left[\left. E\left[\left. Y_{t+h} - Y_{t-1} \right| P=p, X \right] - E\left[\left. Y_{t+h} - Y_{t-1} \right| D_{t+h}=o, X\right] \right| P=p \right], \end{equation} where $P$ denotes the treatment period, $Y_{t}(p)$ denotes the potential outcome at period $t$ when an individual starts receiving a treatment at period $p$, $Y_{t}(o)$ denotes the potential outcome at period $t$ when an individual remains untreated, $Y_t$ denotes the observed outcome at period $t$, and $D_t=o$ indicates that an individual has not been treated as of period $t+h$. Both (ref) and (ref) may be aggregated over $P < \infty$ as in dube2023local. Also, see imai2023matching for a closely related estimand. Our general framework encompasses this example. Specifically, our ATT (ref) reduces to (ref) and our DIDM estimand (ref) reduces to (ref) by setting \begin{align*} W := \begin{cases} 1 & if D_{t+h}=o \\ 0 & if P=p \\ -1 & otherwise \end{cases}, && \widetilde Y_0(0) := Y_{t-1}(o), && \widetilde Y_0(1) := Y_{t-1}(p), && \widetilde Y_0 := Y_{t-1}, \\ && \widetilde Y_1(0) := Y_{t+h}(o), && \widetilde Y_1(1) := Y_{t+h}(p), && \widetilde Y_1 := Y_{t+h}. \end{align*} Similar remarks to Section (ref) follow regarding the pre-treatment condition (ref) and the comparison condition (ref).

Summary and Discussions of the Three Examples

Albeit there are slight differences in their notations, the three examples presented above focus on similar setups. They fundamentally differ only in terms of the estimands: the three examples focus on the M, DID, and DIDM estimands in our language. In general, a researcher does not know which of them achieves the identification. The M estimand identifies the true ATT (i.e, ${\theta_{\text{ATT}}^{\text{M}}} = {\theta_{\text{ATT}}}$ holds) if the matching condition

align*[align* omitted — 127 chars of source]

is satisfied. The DID estimand identifies the true ATT (i.e, ${\theta_{\text{ATT}}^{\text{DID}}} = {\theta_{\text{ATT}}}$ holds) if the parallel trend condition

align*[align* omitted — 193 chars of source]

is satisfied. Finally, the DIDM estimand identifies the true ATT (i.e, ${\theta_{\text{ATT}}^{\text{DIDM}}} = {\theta_{\text{ATT}}}$ holds) if the conditional parallel trend condition

align*[align* omitted — 198 chars of source]

is satisfied. In the absence of knowledge of the underlying data-generating process, however, committing to a wrong assumption can lead to biased estimates by M, DID, or DIDM. It is therefore of interest to characterize the relation among the three estimands. The following subsection investigates this point.

The General Double Bracketing Result

Now, focus on the generic framework (ref)--(ref) again. Let

align*[align* omitted — 345 chars of source]

be the identification errors of the estimands, ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DID}}}$, and ${\theta_{\text{ATT}}^{\text{DIDM}}}$, respectively. We establish the double bracketing relation $\Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}})$ under the following assumption.

assumptionThe following conditions hold. \begin{enumerate}[(i)] • $E\left[\left.\widetilde Y_0 \right|W=0,X=x\right] \geq E\left[\left. \widetilde Y_0 \right|W=1,X=x\right]$ for all $x$. • $F_{X|W=0}$ multivariate first-order stochastically dominates $F_{X|W=1}$.\footnote{We say that $X$ multivariate first-order stochastically dominates $X^\ast$ if $P(X \leq x) \leq P(X^\ast \leq x)$ holds for all $x \in \mathbb{R}^p$.} • $x \mapsto \Phi(x) := E\left[\left.\widetilde Y_1-\widetilde Y_0\right|W=0,X=x\right]$ is weakly decreasing.\footnote{We say that $\Phi$ is weakly decreasing if $\Phi(x_1,\ldots,x_p) \geq \Phi(x^\ast_1,\ldots,x^\ast_p)$ holds whenever $x_1 \leq x^\ast_1$, $\ldots$ and, $x_p \leq x^\ast_p$.} \end{enumerate}

The three parts (ref)--(ref) of this assumption parallel those in Assumption (ref), albeit that $X$ is now possibly multi-dimensional. Hence, similar interpretations can be made especially when $X$ consists of lagged outcomes as in the first example a la acemoglu2019democracy presented in Section (ref). Such a convenient interpretation may not be feasible if $X$ contains other covariates, but we want to stress that each of the three conditions (ref)--(ref) of this assumption is still empirically testable.

The following theorem states the extended double bracketing result for the general cases.

theoremSuppose that Assumption (ref) holds for (ref)--(ref). Then, we have the bracketing relationship \begin{align*} \Delta({\theta_{ATT}^{M}}) \leq \Delta({\theta_{ATT}^{DIDM}}) \leq \Delta({\theta_{ATT}^{DID}}). \end{align*}
proof[Proof of Theorem (ref) ] Note that the identification errors can be written as \begin{align*} \Delta\left({\theta_{ATT}^{M}}\right)= & E\left[E\left[\widetilde Y_1(0) \mid W=1, X\right]-E\left[\widetilde Y_1(0) \mid W=0, X\right] \mid W=1\right], \\ \Delta\left({\theta_{ATT}^{DID}}\right)= & E\left[\widetilde Y_1(0)-\widetilde Y_0(0) \mid W=1\right]-E\left[\widetilde Y_1(0)-\widetilde Y_0(0) \mid W=0\right], \qquadand \\ \Delta\left({\theta_{ATT}^{\text{DIDM}}}\right)= & E\left[E\left[\widetilde Y_1(0)-\widetilde Y_0(0) \mid W=1, X\right]-E\left[\widetilde Y_1(0)-\widetilde Y_0(0) \mid W=0, X\right] \mid W=1\right]. \end{align*} First, observe that \begin{align*} \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) - \Delta({\theta_{\text{ATT}}^{\text{M}}}) =& E\left[E\left[\widetilde Y_0(0)|W=0,X\right] - E\left[\widetilde Y_0(0)|W=1,X\right]|W=1\right] \\ =& E\left[E\left[\widetilde Y_0|W=0,X\right] - E\left[\widetilde Y_0|W=1,\right]|W=1\right] \geq 0 \end{align*} where the second equality is due to (ref), and the last inequality follows from Assumption (ref) (ref). Second, observe that \begin{align*} &\Delta({\theta_{\text{ATT}}^{\text{DID}}}) - \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \\ =& E\left[E\left[\widetilde Y_1(0)-\widetilde Y_0(0)|W=0,X\right] | W=1\right] - E\left[\widetilde Y_1(0)-\widetilde Y_0(0)|W=0\right] \\ =& E\left[E\left[\widetilde Y_1(0)-\widetilde Y_0(0)|W=0,X\right] | W=1\right] - E\left[E\left[\widetilde Y_1(0)-\widetilde Y_0(0)|W=0,X\right] | W=0\right] \\ =& E\left[E\left[\widetilde Y_1-\widetilde Y_0|W=0,X\right] | W=1\right] - E\left[E\left[\widetilde Y_1-\widetilde Y_0|W=0,X\right] | W=0\right] \\ =& E\left[\Phi(X) | W=1\right] - E\left[\Phi(X) | W=0\right] \geq 0, \end{align*} where the second equality follows from the law of iterated expectations, the third equality is due to (ref)--(ref), and the last inequality follows from Assumption (ref) (ref)--(ref).

Empirical Evidence of the Assumptions

With Sections (ref)--(ref) providing theoretical justifications for the double bracketing relationship empirically observed in Section (ref), we will now examine whether the underlying assumption (Assumption (ref) or (ref)) of our theory is satisfied by the datasets used in Section (ref). Specifically, we revisit the CPS and PSID datasets from Section (ref) and the observational dataset from Section (ref).\footnote{Unfortunately, we are unable to present our analysis for the JTPA program discussed in Section (ref) due to discrepancies between the currently available microdata and the original microdata used by the authors, which we confirmed through repeated communications with them.}

Job Training Programs

Recall that Section (ref) demonstrates the robustness of the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \le {\theta_{\text{ATT}}^{\text{DID}}}$, for the NSW program using both the CPS data set and the PSID data set. In light of these consistent observations and our theoretical prediction provided in Theorem (ref), the current section examines each of the three parts, (i)--(iii), of Assumption (ref) in detail, using both data sets.

The first condition is Assumption (ref) (ref) requiring that $E[Y_0|W=0,Y_{-s}=y] \geq E[Y_0|W=1,Y_{-s}=y]$ holds for all $y$. To check this condition, we estimate the conditional expectation functions, $y \mapsto E[Y_0|W=0,Y_{-s}=y]$ and $y \mapsto E[Y_0|W=1,Y_{-s}=y]$, non-parametrically by the partitioning-based least squares regression.\footnote{We use the R package lspartition developed by cattaneo2019lspartition. We took all the default parameters.} Their estimates are plotted in Figure (ref).

figure[figure omitted — 719 chars of source]

The solid and dashed lines represent estimates of the conditional expectation functions $y \mapsto E[Y_0|W=0,Y_{-s}=y]$ and $y \mapsto E[Y_0|W=1,Y_{-s}=y]$, respectively. Shaded areas denote their 95% confidence bands.\footnote{We remark that the confidence bands are not centered around the estimates in lspartition, because of bias correction.} The top image in this figure is based on the CPS data set while the bottom one is based on the PSID data set. For both of the two data sets, the inequality $E[Y_0|W=0,Y_{-s}=y] \geq E[Y_0|W=1,Y_{-s}=y]$ is clearly satisfied for all $y$, providing evidence in support of our Assumption (ref) (ref).

The second condition, Assumption (ref) (ref), requires that $F_{Y_{-s}|W=0}$ first-order stochastically dominates $F_{Y_{-s}|W=1}$. We estimate the conditional cumulative distribution functions, $F_{Y_{-s}|W=0}$ and $F_{Y_{-s}|W=1}$, by their empirical counterparts, i.e., empirical CDFs conditionally on $W=0$ and $W=1$, respectively. Their estimates are plotted in Figure (ref).

figure[figure omitted — 683 chars of source]

The solid line represents the estimates of $F_{Y_{-s}|W=0}$ while the dotted line represents those of $F_{Y_{-s}|W=1}$. Their 95% confidence intervals are indicated by the shaded regions. The left figure is based on the CPS data set while the right one is based on the PSID data set. For both of the two data sets, the figure clearly shows that $F_{Y_{-s}|W=0}$ first-order stochastically dominates $F_{Y_{-s}|W=1}$, providing evidence in support of our Assumption (ref) (ref).

The third condition, Assumption (ref) (ref), requres the function $y \mapsto \Phi(y) := E[Y_1-Y_0|W=0,Y_{-s}=y]$ to be weakly decreasing. We estimate the conditional expectation function $\Phi$ non-parametrically by using the partitioning-based least squares regression -- see Footnote (ref) for details. The estimates, along with their 95% confidence bands, are plotted in Figure (ref).

figure[figure omitted — 724 chars of source]

The top figure is based on the CPS data set while the bottom one is based on the PSID data set. The confidence bands imply that the weak decreasingness of this function $\Phi$ cannot be refuted for either of the two data sets. Thus, they provide evidence in support of our Assumption (ref) (ref).

In the current section, we do not include auxiliary covariates in the current analysis. In Appendix (ref), we provide further evidence in support of our assumption even after accounting for auxiliary covariates.

In summary, all three components, (i)--(iii), of Assumption (ref) hold fairly robustly for both the CPS and PSID data sets regardless of whether auxiliary covariates are included or not. Hence, the double bracketing relationship $\Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}})$, as predicted by our Theorem (ref), is expected to hold, which we have empirically confirmed in Section (ref).

comment\subsubsection{The JTPA Programs} Given the empirical evidence for the conclusion in Theorem (ref), our next question is whether each of the three components of Assumption (ref) is satisfied for the JTPA program. To address this, we will examine each of the three parts, (i)--(iii), of Assumption (ref) in detail. Recall that the first condition is Assumption (ref) (ref), requiring that the inequality $E[Y_0|W=0,Y_{-s}=y] \geq E[Y_0|W=1,Y_{-s}=y]$ holds for all $y$. To check this condition, we estimate the conditional expectation functions, $y \mapsto E[Y_0|W=0,Y_{-s}=y]$ and $y \mapsto E[Y_1|W=0,Y_{-s}=y]$, non-parametrically using the Nadaraya-Watson estimator. The difference, $E[Y_0|W=0,Y_{-s}=y] - E[Y_0|W=1,Y_{-s}=y]$, of their estimates along with their 95% confidence intervals are plotted in Figure (ref). \begin{figure}[t] \caption{Evidence of Assumption (ref) (ref) for the JTPA program. The conditional expectation functions, $y \mapsto E[Y_0|W=0,Y_{-s}=y]$ and $y \mapsto E[Y_1|W=0,Y_{-s}=y]$, are estimated non-parametrically by the Nadaraya-Watson estimator. The difference $E[Y_0|W=0,Y_{-s}=y] - E[Y_1|W=0,Y_{-s}=y]$ and their 95% confidence intervals plotted. Both the vertical and horizontal ases are measured in thousands of U.S. dollars.}${}$ \end{figure} This figure showcases that the inequality $E[Y_0|W=0,Y_{-s}=y] \geq E[Y_0|W=1,Y_{-s}=y]$ indeed holds for all $y$ in terms of the estimates, providing evidence in support of our Assumption (ref) (ref). Next, recall that the second condition is Assumption (ref) (ref), which requires that $F_{Y_{-s}|W=0}$ first-order stochastically dominates $F_{Y_{-s}|W=1}$. We estimate the conditional cumulative distribution functions, $F_{Y_{-s}|W=0}$ and $F_{Y_{-s}|W=1}$, by their empirical counterparts, i.e., empirical CDFs conditionally on $W=0$ and $W=1$, respctively. Their estimates are plotted in Figure (ref). The blue line indicates the estimates of $F_{Y_{-s}|W=0}$ while red line indicates the estimates of $F_{Y_{-s}|W=1}$. Their 95% confidence intervals are indicated by the shaded regions in colors. This figure demonstrates that $F_{Y_{-s}|W=0}$ indeed first-order stochastically dominates $F_{Y_{-s}|W=1}$ in terms of the estimates, providing evidence in support of Assumption (ref) (ref). \begin{figure}[t] \caption{Evidence of Assumption (ref) (ref) for the JTPA program. The conditional cumulative distribution functions, $F_{Y_{-s}|W=0}$ and $F_{Y_{-s}|W=1}$, are estimated by their empirical counterparts, i.e., empirical CDFs. The estimates, along with their 95% confidence intervals, are plotted. The horizontal axes are measured in thousands of U.S. dollars.}${}$ \end{figure} The third condition is Assumption (ref) (ref), which requires $y \mapsto \Phi(y) := E[Y_1-Y_0|W=0,Y_{-s}=y]$ to be weakly decreasing. We estimate the conditional expectation function $\Phi$ non-parametrically by the Nadaraya-Watson estimator. The estimates along with their 95% confidence intervals are plotted in Figure (ref). The plot shows that this function $\Phi$, in terms of the estimates, is non-increasing for a large portion of the support except in a neighborhood of $Y_{-s}=5$. This provides a partial evidence in support of of our Assumption (ref) (ref). In Appendix (ref), we provide further evidence in support of our assumption even after accounting for auxiliary covariates. In summary, all the three components, (i)--(iii), of Assumption (ref) hold robustly for the JTPA program. Hence, the double bracketing relationship $\Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}})$, as predicted by our Theorem (ref), is expected to hold, which we have empirically confirmed in Section (ref). \begin{figure}[H] \caption{Evidence of Assumption (ref) (ref) for the JTPA program. The conditional expectation function $\Phi$ is non-parametrically estimated by the Nadaraya-Watson estimator. The estimates, along with their 95% confidence intervals, are plotted. Both the vertical and horizontal axes are measured in thousands of U.S. dollars.}${}$ \end{figure}

Educational Programs

Recall that Section (ref) demonstrates that the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, is robust for the educational program investigated by athey2020combining. In light of this observation and our theoretical prediction provided in Theorem (ref), our next question is whether Assumption (ref) for the double bracketing theory is satisfied for this educational program. To address this, we will examine each of the three parts, (i)--(iii), of Assumption (ref) in detail using the data set utilized in Section (ref).

Recall that the first condition is Assumption (ref) (ref), requiring that $E[Y_0|W=0,Y_{-s}=y] \geq E[Y_0|W=1,Y_{-s}=y]$ holds for all $y$. We estimate the conditional expectation functions, $y \mapsto E[Y_0|W=0,Y_{-s}=y]$ and $y \mapsto E[Y_0|W=1,Y_{-s}=y]$, non-parametrically using the partitioning-based least squares regression -- see Footnotes (ref)--(ref) for details. The estimates are plotted in Figure (ref).

figure[figure omitted — 535 chars of source]

The solid and dashed lines represent estimates of the conditional expectation functions $y \mapsto E[Y_0|W=0,Y_{-s}=y]$ and $y \mapsto E[Y_0|W=1,Y_{-s}=y]$, respectively. Shaded areas denote their 95% confidence bands, although they are nearly invisible due to the large sample size.\footnote{We remark that the estimates appear outside of the confidence bands because the bands are centered around bias-corrected estimates in the lspartition package -- see Footnote (ref).} This figure showcases that the inequality $E[Y_0|W=0,Y_{-s}=y] \geq E[Y_0|W=1,Y_{-s}=y]$ is satisfied for all $y$, providing evidence in support of our Assumption (ref) (ref).

Next, recall that the second condition is Assumption (ref) (ref), which requires that $F_{Y_{-s}|W=0}$ first-order stochastically dominates $F_{Y_{-s}|W=1}$. We estimate the conditional cumulative distribution functions, $F_{Y_{-s}|W=0}$ and $F_{Y_{-s}|W=1}$, by their empirical counterparts, i.e., empirical CDFs conditionally on $W=0$ and $W=1$, respectively. Their estimates are plotted in Figure (ref). The solid line indicates the estimates of $F_{Y_{-s}|W=0}$ while the dotted line indicates the estimates of $F_{Y_{-s}|W=1}$. Their 95% confidence intervals are indicated by the shaded regions in colors, although they are nearly invisible due to the large sample size again. This figure demonstrates that $F_{Y_{-s}|W=0}$ indeed first-order stochastically dominates $F_{Y_{-s}|W=1}$, providing evidence in support of Assumption (ref) (ref).

figure[figure omitted — 405 chars of source]

The third condition is Assumption (ref) (ref), which requires $y \mapsto \Phi(y) := E[Y_1-Y_0|W=0,Y_{-s}=y]$ to be weakly decreasing. We estimate the conditional expectation function $\Phi$ non-parametrically by using the partitioning-based least squares regression -- see Footnote (ref) for details. The estimates, along with their 95% confidence bands, are plotted in Figure (ref). The plot shows that this function $\Phi$, is indeed non-increasing, providing evidence in support of of our Assumption (ref) (ref).

Our analysis in the current section omits auxiliary covariates. In Appendix (ref), we provide further evidence in support of our assumption even after accounting for auxiliary covariates. In particular, accounting for the covariates will strengthen the empirical support of our Assumption (ref) (ref) as mentioned above.

In summary, all the three components, (i)--(iii), of Assumption (ref) hold robustly for the educational program regardless of whether auxiliary covariates are included or not. Hence, the double bracketing relationship $\Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}})$, as predicted by our Theorem (ref), is expected to hold, as we empirically confirmed in Section (ref).

figure[figure omitted — 404 chars of source]

Summary and Discussions

The paper evaluates the relative performance of three estimands -- Matching (M), Difference-in-Differences (DID), and a hybrid method (DIDM) -- for estimating causal effects in observational studies, particularly in the context of job training and educational programs. Our analysis reveals a consistent inequality: \( {\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}} \leq {\theta_{\text{ATT}}^{\text{DID}}} \). This indicates that Matching tends to produce the most conservative estimates, while DID often yields the most optimistic ones. When selecting a single method, it may be prudent to favor the more conservative Matching estimator. If the Matching estimator suggests a positive effect, this provides a compelling argument that the causal effect is indeed likely positive, leading to a more robust conclusion.

Moreover, these estimands can be utilized in a complementary manner. If practitioners believe that any one of the three identification assumptions -- unconfoundedness (for M), parallel trends (for DID), or conditional parallel trends (for DIDM) -- holds, then the true causal effect is bracketed by \( {\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}} \leq {\theta_{\text{ATT}}^{\text{DID}}} \). This approach offers a “triply robust bracketing” of the causal effect, providing valuable information regardless of which assumption is valid. By applying this framework, researchers can gain a clearer understanding of the potential range of treatment effects, with Matching yielding the most conservative estimate and DID providing the most optimistic one.

Furthermore, the results provide an intuitive strategy for managing uncertainty in identification assumptions. If the Matching (M) estimator indicates a positive effect, the sign of the treatment effect can be interpreted with confidence. The Difference-in-Differences (DID) estimator then provides an upper bound for the magnitude of this effect. This approach simplifies the analysis compared to traditional sensitivity methods (e.g., Manski and Pepper 2018), offering more interpretable bounds for researchers.