Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
112,803 characters · 29 sections · 128 citation commands
Matching $$ Hybrid $$ Difference in Differences
\allowdisplaybreaks
Since the seminal work by lalonde1986evaluating, there has been substantial interest in accurately estimating the short-term average treatment effects on the treated (ATT) and other causal parameters using panel data with two (pre-treatment and post-treatment) periods heckman1998_2matching,dehejia2002propensity,smith2005does. This body of literature has focused on a debate concerning which method -- matching (M), difference-in-differences (DID), or their hybrid (DIDM) -- is more effective at replicating experimental estimates when using observational data.
In the current era, where the DID has once again captured the attention of empirical practitioners, we contribute to the ongoing debate by examining the relationships among these three estimands from both empirical and theoretical perspectives. Our analysis clarifies the conditions under which one method provides the most conservative estimates while another offers the most optimistic estimates. By doing so, we provide a more nuanced understanding of the relative strengths and limitations of each approach, enabling researchers to make more informed methodological choices based on the specific characteristics of their data.
We begin by highlighting a puzzling pattern consistently observed in the empirical data used in both the aforementioned studies and other seminal works in the economics of education and labor economics. Specifically, the inequality relationship, M $\leq$ DIDM $\leq$ DID, frequently appears as a consistent norm, rather than a mere coincidence, when assessing the effectiveness of educational and job training programs.
For example, Figure (ref) illustrates the biases of the M, DIDM, and DID estimates for job training programs relative to experimental estimates perceived as benchmark truths. These estimates are excerpted from two seminal papers: heckman1998characterizing and smith2005does, which revisits the analysis of lalonde1986evaluating,dehejia1999causal,dehejia2002propensity. Despite differences in the job training programs studied, data sets, and estimation methods, the inequality relationship, M $\leq$ DIDM $\leq$ DID, holds robustly across both papers. This relationship remains robust even when considering alternative estimation methods, different programs, or other data sets, as will be demonstrated shortly in this paper.
This pattern is reminiscent of the so-called `bracketing' relationship between the lagged dependent variable (LDV) estimator and the fixed-effect (FE) estimator in panel regressions, as presented in angrist2009mostly. Specifically, angrist2009mostly outline plausible conditions under which the inequality relationship, LDV $\leq$ FE, holds on theoretical grounds. Their assumptions necessitate negative selection into treatment and the absence of explosive outcomes, which make plausible sense in many labor economic settings. This result has been elegantly extended to a nonparametric setup by ding2019bracketing.
We find that similar plausible conditions, in the spirit of angrist2009mostly and ding2019bracketing, also give rise to the aforementioned inequality relationship, M $\leq$ DIDM $\leq$ DID. In fact, a special case of our general double bracketing result, M $\leq$ DIDM $\leq$ DID, reduces to the conventional bracketing result, LDV $\leq$ FE, established by angrist2009mostly and ding2019bracketing.
Recall that M, DID, and DIDM identify the true causal parameter under the assumptions of observational unconfoundedness, parallel trends, and conditional parallel trends, respectively. In practice, an empirical researcher may not know which of these three alternative conditions is satisfied for an application of interest. Our double bracketing result, M $\leq$ DIDM $\leq$ DID, implies that the true causal parameter is bracketed below by M and above by DID, with DIDM between them, when one of the three alternative assumptions holds true. In other words, our theoretical prediction implies that the DID approach tends to yield the most optimistic estimates while the M approach tends to yield the most conservative estimates.
While our discussions primarily focus on the classic two-period framework heckman1998characterizing,smith2005does, we also extend our double bracketing result to a broader class that includes cases with multi-dimensional covariates, dynamic and multi-period settings, and impulse response functions in event studies, as explored in the recent literature callaway2018difference,acemoglu2019democracy,deChaisemartin2020two,dube2023local,imai2023matching.
Using four data sets that have been used for the evaluation of job training and educational programs in the literature lalonde1986evaluating,heckman1998characterizing,smith2005does,athey2020combining, we empirically examine our double bracketing hypothesis, M $\leq$ DIDM $\leq$ DID. In light of the robustness of this inequality relationship in all the empirical scenarios, we provide formal theoretical explanations for this intriguing phenomenon. Our assumptions required for the double bracketing relationship, motivated by angrist2009mostly and ding2019bracketing as mentioned earlier, is not only in line with the literature but are also empirically testable. Hence, we examine our assumptions, as well as the double bracketing relations per se, using these empirical data sets. It turns out that our assumptions, as well as the double bracketing relationship, M $\leq$ DIDM $\leq$ DID, indeed hold robustly in all these empirical cases.
The question we investigate relates to a long literature of econometrics including a number of seminal papers.
First, this paper closely relates to the classical debate in economics on what type of non-experimental estimates replicate the experimental estimates of lalonde1986evaluating. The pioneering papers in that literature are heckman1998characterizing, heckman1998_2matching, dehejia2002propensity, and smith2005does. There, the main purpose was to find the single best estimator, often measured by the absolute bias. We are interested in signs, as well as the magnitudes, of bias to identify which estimator is the most conservative/optimistic.
More recently, one notable work revisited a related problem. Specifically, Figure (ref) replicates chabe2017should and presents the absolute biases of the three estimands, M, DIDM, and DID, based on the estimates from heckman1998characterizing and smith2005does.
With the symbol `$\Delta$' denoting the bias, this figure essentially shows $|\Delta $M$|\geq|\Delta $DIDM$|\geq|\Delta $DID$|$, implying that DID achieves the smallest bias for these empirical cases. Notably, this Figure (ref) depicts the absolute biases corresponding to the signed biases shown in Figure (ref) in our introductory section.
Thus, while the existing literature primarily focuses on identifying the best estimator based on absolute bias, our paper distinguishes itself by examining the relative signs of biases. This signed characterization provides more nuanced insights; for instance, Figure (ref) illustrates that M underestimates the true causal effects, whereas DID overestimates them. Beyond this primary distinction between absolute and signed biases, our paper differs from chabe2017should in two other key aspects. First, we provide a formal theoretical result, whereas the existing work derives its conclusions from numerical studies. Second, the existing work relies on functional form and structural assumptions -- such as separability and distinctions between transitory and permanent income -- tailored to job training programs. In contrast, we adopt a more model-free, nonparametric approach.
Our theoretical results build on the existing literature on bracketing. For linear panel models, angrist2009mostly develop a bracketing relationship, showing that LDV $\leq$ FE. This result has been extended to a nonparametric context by ding2019bracketing. We contribute to this literature in two key ways. First, we consider a more general framework that encompasses the symmetric DID (DIDM) studied by heckman1998characterizing and smith2005does in particular. Second, within this general framework, we show that DIDM, which heckman1998characterizing recommend for minimizing bias, is further bracketed between M and DID, resulting in our double bracketing outcome. A special case of our general framework reduces M and DIDM to LDV, and thus, our double bracketing relationship, M $\leq$ DIDM $\leq$ DID, simplifies to the conventional bracketing relationship, LDV $\leq$ FE.
Finally, this paper is of course related to the burgeoning literature on event studies and DID today -- see recent surveys by deChaisemartin2023two and roth2023review among others. This paper is also related to the huge literature on matching, including covariate matching, propensity score matching, (augmented) inverse probability weighting, nearest neighbor methods, and regression adjustments, among others. In particular, the matching literature has emphasized the importance of conditioning on lagged outcomes as matching factors to improve the precision of causal inference lalonde1986evaluating,dehejia1999causal,dehejia2002propensity. Economically, matching on lagged outcomes helps to address the so-called “ashenfelter1978estimating dip.” The empirical work by acemoglu2019democracy confirms that conditioning on lagged outcomes indeed yields more credible estimates in event studies. Effectively, they advocate the M over the DID in our language. Later, dube2023local generalize acemoglu2019democracy to a hybrid framework which we refer to as the DIDM in this paper. See the DID${}_\text{M}$ estimator of deChaisemartin2020two and the (panel) matching estimator of imai2023matching, as well as dube2023local -- they all propose and study the properties of what we refer to as the DIDM. We contribute to the above literature by shedding light on the systematic relationship among M, DIDM, and DID, and empirically and theoretically examining this relationship.
Let $W$ denote the group of treatment assignment such that units with $W=1$ receive treatment between periods $t=0$ and $t=1$, while those with $W=0$ remain untreated. Thus, everyone is untreated for $t \leq 0$, and only those with $W=1$ are treated for $t \geq 1$. Let $Y_t(d)$ denote the potential outcome under treatment status $d$ at time $t$. With these notations, suppose that a researcher is interested in identifying the average treatment effect on the treated (ATT) at time $t=1$ defined by
Letting $Y_t$ denote the observed outcome at time $t$, the seminal paper by heckman1998characterizing proposes three alternative estimands to this goal:
which we call the matching (M), the difference-in-differences (DID), and the difference-in-differences matching (DIDM), respectively. Here, we focus on the past outcome $Y_{-s}=Y_{-s}(0)$ as the key matching criterion following heckman1998characterizing and smith2005does, but we present an extension to general vector-valued matching criteria $X$ in Section (ref) to include other auxiliary covariates as well as more lagged outcomes as in the local projection approaches to event studies acemoglu2019democracy.
The M estimand identifies the true ATT (i.e, ${\theta_{\text{ATT}}^{\text{M}}} = {\theta_{\text{ATT}}}$ holds) if the matching condition
is satisfied. The DID estimand identifies the true ATT (i.e, ${\theta_{\text{ATT}}^{\text{DID}}} = {\theta_{\text{ATT}}}$ holds) if the parallel trend condition
is satisfied. Finally, the DIDM estimand identifies the true ATT (i.e, ${\theta_{\text{ATT}}^{\text{DIDM}}} = {\theta_{\text{ATT}}}$ holds) if the conditional parallel trend condition
is satisfied.
As heckman1998characterizing and smith2005does stressed, in the absence of knowledge about the true data-generating process, there is a risk of bias associated with the three estimands, ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DID}}}$, and ${\theta_{\text{ATT}}^{\text{DIDM}}}$. For instance, when the true DGP satisfies Condition M but does not satisfy Condition DID or Condition DIDM, then only ${\theta_{\text{ATT}}^{\text{M}}}$ is guaranteed the identification while estimators for ${\theta_{\text{ATT}}^{\text{DID}}}$ and ${\theta_{\text{ATT}}^{\text{DIDM}}}$ are doomed to be biased in general. Hence, understanding the systematic relationship among ${\theta_{\text{ATT}}^{\text{M}}}$ ${\theta_{\text{ATT}}^{\text{M}}}$, and ${\theta_{\text{ATT}}^{\text{DID}}}$ can help researchers form insights into the possible range within which the true causal effect, ${\theta_{\text{ATT}}}$, may lie when one of the three alternative conditions holds. We will empirically confirm the systematic double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, in Section (ref) and provide theoretical explanations for it in Sections (ref)--(ref).
{\color{black} The three identifying assumptions underlying the estimands \({\theta_{\text{ATT}}^{\text{M}}}\), \({\theta_{\text{ATT}}^{\text{DID}}}\), and \({\theta_{\text{ATT}}^{\text{DIDM}}}\) each restrict the data‐generating process (DGP) in a distinct way. No one assumption nests another; that is, none of the conditions implies or is implied by any of the others. In the following, we provide intuitive counterexamples—some leveraging differences in sample composition—to illustrate these distinctions.
Consider the following data-generating process (DGP) for the potential outcomes:
where $(\eta_0(0),\eta_0(1),\eta_1(0),\eta_1(1)) \perp \!\!\! \perp (W, Y_{-s})$ and $\mu(\cdot)$ is an arbitrary function such that $\mu(0) \neq \mu(1)$. Under this DGP, $(Y_1(1),Y_1(0)) \perp \!\!\! \perp W | Y_{-s}$ holds, but ${\text{E}}[Y_1(0)-Y_0(0)|Y_{-s},W=0]={\text{E}}[\eta_1(0)-\eta_0(0)]-\mu(0) \neq {\text{E}}[\eta_1(0)-\eta_0(0)]-\mu(1) = {\text{E}}[Y_1(0)-Y_0(0)|Y_{-s},W=1]$. In other words, Condition M holds, but Condition DIDM fails.
Conversely, consider the following DGP for the potential outcomes:
where $(\eta_0(0),\eta_0(1),\eta_1(0),\eta_1(1)) \perp \!\!\! \perp (W, Y_{-s})$ and $\mu(\cdot)$ is an arbitrary function such that $\mu(0) \neq \mu(1)$. Under this DGP, ${\text{E}}[Y_1(w)|Y_{-s},W=0]=\mu(0)+{\text{E}}[\eta_1(0)] \neq \mu(1)+{\text{E}}[\eta_1(0)] = {\text{E}}[Y_1(w)|Y_{-s},W=1] $, and hence, $(Y_1(1),Y_1(0)) \not\perp \!\!\! \perp W | Y_{-s}$, but ${\text{E}}[Y_1(0)-Y_0(0)|Y_{-s},W=0]={\text{E}}[\eta_1(0)-\eta_0(0)] = {\text{E}}[Y_1(0)-Y_0(0)|Y_{-s},W=1]$ holds. In other words, Condition M fails, but Condition DIDM holds.
Consider the following DGP for untreated potential outcomes:
where $E[\eta_1(0)-\eta_0(0)|Y_{-1},W]=Y_{-1}$ and $E[Y_{-1}|W=0] \neq E[Y_{-1}|W=1]$. Under this DGP, we have $E[Y_1(0)-Y_0(0)|Y_{-1},W=0]=Y_{-1}=E[Y_1(0)-Y_0(0)|Y_{-1},W=1]$, but $ E[Y_1(0)-Y_0(0)|W=0] = E[E[\eta_1(0)-\eta_0(0)|Y_{-1},W=0]|W=0] = E[Y_{-1}|W=0] \neq E[Y_{-1}|W=1] = E[E[\eta_1(0)-\eta_0(0)|Y_{-1},W=1]|W=1] = E[Y_1(0)-Y_0(0)|W=1]. $ In other words, Condition DIDM holds, but Condition DID fails. Intuitively, even if the parallel trends hold at every level of \(Y_{-1}\), differences in the distribution of \(Y_{-1}\) between treatment groups can lead to unequal unconditional trends.
Conversely, consider the following DGP for untreated potential outcomes:
where $E[\eta_1(0)-\eta_0(0)|Y_{-1},W] = (2W-1)Y_{-1}$ and $E[Y_{-1}|W=0]+E[Y_{-1}|W=1]=0$. Under this DGP, $ E[Y_1(0)-Y_0(0)|W=0] = E[E[Y_1(0)-Y_0(0)|Y_{-1},W=0]|W=0] = E[-Y_{-1}|W=0] = E[Y_{-1}|W=1] = E[E[Y_1(0)-Y_0(0)|Y_{-1},W=1]|W=1] = E[Y_1(0)-Y_0(0)|W=1], $ but $E[Y_1(0)-Y_0(0)|Y_{-1},W=0] = -Y_{-1} \neq Y_{-1} = E[Y_1(0)-Y_0(0)|Y_{-1},W=1].$ In other words, Condition DID holds but Condition DIDM fails. Intuitively, even if the overall (unconditional) trends are equal (due to cancellation when aggregating over \(Y_{-1}\)), the trends might differ for each subpopulation defined by \(Y_{-1}\).
\noindentImplication for Practice.\\ These examples underscore that each assumption addresses a different facet of the DGP. Practitioners must ex ante commit to one of these assumptions based on the context and the nature of available data. Whether one relies on matching (Condition M), unconditional parallel trends (DID), or conditional parallel trends (DIDM) is not a matter of nested robustness but of fundamentally different identifying restrictions, each carrying its own trade-offs and implications for causal inference.
}
Following heckman1998characterizing and smith2005does, we place particular emphasis on lagged outcomes, like $Y_{-s}$, as the conditioning variable for M and DIDM. This focus is motivated by several key studies in the literature.
Firstly, ashenfelter1978estimating argues that participants in job training programs typically have permanently lower earnings than non-participants, but also experience a temporary decrease in earnings just before entering the program, a phenomenon now famously known as the “Ashenfelter dip.” Controlling for lagged outcomes is crucial to addressing this confounding factor, as also emphasized by heckman1999economics among others.
In lalonde1986evaluating and dehejia1999causal,dehejia2002propensity, indeed, including lagged outcomes in matching allowed for close replication of the baseline experimental estimates. The subsequent discourse between smith2005does and dehejia2005practical further underscored the significance of conditioning on lagged outcomes.
angrist2009mostly discuss the importance of conditioning on lagged outcomes in the context of lagged-dependent-variable (LDV) models. More recently, roth2023parallel, in their review of difference-in-differences methods, raise an open question about the validity of conditioning on lagged outcomes and whether it reduces or exacerbates bias.
Finally, influential empirical work has emphasized the special role that lagged outcomes play in reducing bias in policy evaluation. For example, chetty2014measuring1, in their work on value-added models in education, highlight the utility of lagged outcomes, specifically prior test scores, as essential covariates for obtaining unbiased value-added estimates.
These studies suggest the importance of focusing on lagged outcomes as a key conditioning variable and analyzing the sources of bias that may arise under different conditions. We thus follow this literature to focus on the lagged outcome $Y_{-s}$ for matching in the current section. With this said, Section (ref) will generalize this setup to include a general class of matching criteria in addition to lagged outcomes.
{\color{black} Finally, we briefly discuss the lag order, denoted as $-s$. In the special case where $s=0$, the DIDM estimand ${\theta_{\text{ATT}}^{\text{DIDM}}}$ simplifies to the M estimand ${\theta_{\text{ATT}}^{\text{M}}}$. As a result, the double bracketing relation ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$ that we present later reduces to ${\theta_{\text{ATT}}^{\text{M}}} = {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, which corresponds to the existing bracketing result, LDV $\leq$ FE, as demonstrated by angrist2009mostly and ding2019bracketing. Our framework, which accommodates $s \geq 0$, is more general and particularly includes the symmetric DID approach introduced by heckman1998characterizing and smith2005does. This approach, which aligns with our DIDM with $-s=-1$, was shown by heckman1998characterizing and subsequent studies to be more accurate than other specifications. Therefore, our arbitrary lag order $-s$ not only encompasses the classical bracketing scenario as a special case but also includes many important cases explored in foundational papers within the literature. }
This section revisits three empirical studies using four datasets from the literature, parts of which are derived from the seminal papers discussed in Section (ref). We are going to document the robustness of the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, across each of these scenarios, different data sets, estimation methods, and subpopulations defined by observed characteristics. The first two examples, presented in Section (ref), focus on job training programs, namely the NSW and JTPA. The final example, in Section (ref), examines educational programs.
We begin our analysis with a reexamination of the National Supported Work (NSW) programs lalonde1986evaluating,dehejia1999causal,dehejia2002propensity,smith2005does through the lens of our double bracketing. We employ the same sets of the CPS and PSID data as those utilized by lalonde1986evaluating and smith2005does. Although these data sets have been extensively analyzed in the literature, we offer a brief description in Appendix (ref) for the sake of completeness and convenience of the readers.
The primary outcome variable, $Y_t$, represents participants' self-reported earnings, adjusted to 1982 dollars. The treatment variable, $W$, is a binary indicator denoting whether an individual was assigned to the NSW program. Additionally, demographic variables, including race and educational level, are included as auxiliary covariates in all M, DIDM, and DID estimations.
Let us revisit Figure (ref), which was initially introduced in Section (ref) to motivate our investigation. Recall that Figure (ref) illustrates the signed biases of the estimates for the three alternative estimands, ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DIDM}}}$, and ${\theta_{\text{ATT}}^{\text{DID}}}$, relative to the experimental estimates in percentage terms according to heckman1998characterizing and smith2005does.
In the current section, we focus on the black bars in Figure (ref), representing the estimates by smith2005does. Notably, the double bracketing inequality, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, is upheld when considering their point estimates. Since these represent the biases, note that ${\theta_{\text{ATT}}^{\text{M}}}$ and ${\theta_{\text{ATT}}^{\text{DIDM}}}$ are biased downward, whereas ${\theta_{\text{ATT}}^{\text{DID}}}$ is biased upward. These results suggest that the M estimand, ${\theta_{\text{ATT}}^{\text{M}}}$, tends to be conservative, while the DID estimand, ${\theta_{\text{ATT}}^{\text{DID}}}$, appears to be optimistic, assuming that the experimental estimates are the true values.
The observation above applies specifically to the estimates of ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DIDM}}}$, and ${\theta_{\text{ATT}}^{\text{DID}}}$ selected by chabe2017should from the various ST specifications. However, when we extend our analysis to include iterative estimations across all nine estimation methods employed by smith2005does using the CPS and PSID datasets, we find that the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, is robustly upheld across the vast majority of these specifications. This is illustrated in Figure (ref), where the top and bottom charts display the results for the CPS and PSID data, respectively.
Moreover, it is important to note that this ordering of estimates is never rejected for any of the nine specifications. This consistency ensures that the result is not an artifact of a specific estimation procedure. Consistently, we observe that the M estimand, ${\theta_{\text{ATT}}^{\text{M}}}$, tends to be conservative, while the DID estimand, ${\theta_{\text{ATT}}^{\text{DID}}}$, tends to be optimistic.
Next, we reexamine the Job Training Partnership Act (JTPA) program studied by heckman1998characterizing. Although the dataset they used has been extensively analyzed in the literature, a brief description is provided in Appendix (ref) for the sake of completeness and convenience of the readers.
The outcome variable, $Y_t$, represents participants' earnings adjusted for inflation. The treatment variable, $W$, indicates whether an individual was assigned to the JTPA program. Additionally, demographic covariates such as sex and age are included in all M, DIDM, and DID estimations.
Once again, we reiterate Figure (ref) that we provide in Section (ref). For the current section about the JTPA programs, focus on the gray bars in Figure (ref) based on the estimates by heckman1998characterizing. Notice that the inequality relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, is maintained when considering their point estimates. Specifically, the M estimand, ${\theta_{\text{ATT}}^{\text{M}}}$, is biased downward, while the DID estimand, ${\theta_{\text{ATT}}^{\text{DID}}}$, is biased upward.
Now, we turn to the educational programs analyzed by athey2020combining, which were also studied in chetty2014measuring1, chetty2014measuring2. In their study, athey2020combining addressed the challenge of selection in observational by combining it with short-term experimental data. In contrast, our analysis focuses solely on observational data to examine the behavior of the observational estimands: ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DIDM}}}$, and ${\theta_{\text{ATT}}^{\text{DID}}}$. Details of the data are provided in Appendix (ref).
The outcome variable, $Y_t$, represents students' test scores, specifically standardized scores that average results from mathematics and English language arts. The treatment variable, $W$, indicates whether a student was assigned to a small class size. Additionally, covariates such as gender, race, and eligibility for free lunch are considered to define subpopulations in our analysis.
Given the large size of our dataset (see Appendix (ref) for details), we can precisely estimate the M, DIDM, and DID estimands even when the sample is divided into subpopulations based on observed attributes. Figure (ref) presents these estimates, along with their 95% confidence intervals, for each subpopulation.
Observe that the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, remains consistent across various subpopulations, despite notable differences in levels among these groups. Interestingly, the M estimand, ${\theta_{\text{ATT}}^{\text{M}}}$, generally indicates negative effects, except within the Black subpopulations. In contrast, both the DIDM estimand, ${\theta_{\text{ATT}}^{\text{DIDM}}}$, and the DID estimand, ${\theta_{\text{ATT}}^{\text{DID}}}$, typically suggest positive effects. This divergence in implications between estimation methods raises questions about relying solely on observational data, highlighting the value of combining experimental and observational data, as demonstrated by athey2020combining.
So far, we have observed that the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, appears to be more of a rule than a mere coincidence in empirical studies on educational and job training programs. In this section, and the two subsequent ones, we are going to provide theoretical explanations for this intriguing phenomenon.
While general theories will be presented in the subsequent sections, we begin by illustrating how the double bracketing relationship, $ {\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}} $ arises under a familiar parametric setting. In particular, suppose the true data-generating process takes the form:
where the error term satisfies
In this context, $\alpha_i$ and $\delta_t$ serve as two-way fixed effects and $\rho$ serves as a persistence parameter.
Note that $\gamma$ can be expressed as $ \gamma = E[Y_{i,0}|W_i=1,Y_{i,-1}] - E[Y_{i,0}|W_i=0,Y_{i,-1}] $ under (ref)--(ref). This value, $\gamma$, is non-positive if individuals with weakly lower $Y_{i,0}$ tend to select into treatment $W_i=1$ conditional on $Y_{i,-1}$ as is plausibly the case with job training programs. We postulate this negative selection assumption:
We also postulate a similar negative selection assumption for $Y_{i,-1}$:
Finally, we impose the assumption of a non-explosive earnings process:
Note that these three conditions (ref)--(ref) are analogous to the assumptions invoked by angrist2009mostly in developing their bracketing relationship, LDV $\leq$ FE.
Under the parametric data-generating process (ref), some calculations yield
See Appendix (ref) for detailed calculations that derive these expressions. From (ref)--(ref), we have
where the inequality follows from $\gamma \leq 0$ due to (ref). From (ref)--(ref), we have
where the inequality follows from $E[Y_{i,-1}|W_i=0] - E[Y_{i,-1}|W_i=1] \geq 0$ due to (ref) and $\rho \in [0,1]$ due to (ref).
Combining (ref)--(ref) yields the double bracketing relationship,
which holds true under the simple parametric model (ref) with the plausible restrictions (ref)--(ref) motivated by angrist2009mostly.
Finally, we highlight a couple of special cases in which the three estimands collapse into two. In the absence of the first negative selection (i.e., $\gamma = 0$) the double bracketing relationship reduces to
In other words, M and DIDM become equivalent. Similarly, in the absence of persistence (i.d., $\rho=0$) or under the unit root (i.e., $\rho=1$), the double bracketing relationship reduces to
In other words, DIDM and DID become equivalent. In these two special cases, DIDM indeed becomes redundant. However, it is generally distinct from the other two estimands otherwise.
The previous section derived the double-bracketing relationship ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$ under a simple parametric model framework. To ensure generality, we adopt a non-parametric and model-free framework in the current section.
Let
be the identification errors of the estimands, ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DID}}}$, and ${\theta_{\text{ATT}}^{\text{DIDM}}}$, respectively. Observe that ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$ is equivalent to $\Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}})$ with these notations. Thus, we are going to establish the double bracketing relation, $\Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}})$, in terms of the identification error. To this end, we consider the following assumption.
{\color{black} The three components, (ref)--(ref), of Assumption (ref) are analogous to (ref)--(ref), respectively. Each of the three components is empirically testable, as it does not involve unobserved latent variables such as $Y_t(d)$. Furthermore, they are plausible and align with the assumptions made in the literature angrist2009mostly,ding2019bracketing. Specifically, part (ref) requires the so-called “negative selection” that individuals who opt out from treatment (i.e., those with $W=0$) tend to have weakly higher pre-treatment (potential) outcomes $Y_0=Y_0(0)$ without treatment on average given $Y_{-s}=y$. Similarly, part (ref) requires the negative selection that individuals who opt out from treatment (i.e., those with $W=0$) tend to have no lower pre-treatment (potential) outcome $Y_{-s}=Y_{-s}(0)$ without treatment. Part (ref) requires the time series $\{Y_t\}_t$ of outcomes for those with $W=0$ (i.e., $\{Y_t\}_t = \{Y_t(0)\}_t$) is stable over time. In real-world settings, especially among populations with limited socioeconomic status (SES), it may be reasonable to expect that outcomes do not escalate dramatically without intervention. For example, in educational programs targeting low-SES individuals, we wouldn't anticipate rapid, unsustainable improvements in outcomes without structured support on average. In the special case where $s=0$, this requirement essentially reflects the condition that the root of the AR(1) process is less than one, which is equivalent to stationarity, as deployed in, e.g., chetty2014measuring1.
Each of these parts is analogous to the assumption made by angrist2009mostly in deriving their bracketing relationship between the DID and lagged dependent variable (LDV) in linear and additive models. In the special case of $s=0$, part (ref) will be trivially satisfied while parts (ref)--(ref) correspond to the assumptions invoked by ding2019bracketing. As argued at the end of Section (ref), our framework with $s \geq 0$ accommodates both this special case and other important scenarios considered by heckman1998characterizing and others. }
This theorem provides a theoretical explanation for the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, which was robustly observed in Section (ref).
Furthermore, this theorem implies the following three consequences depending on the underlying DGP. First, when Condition M is true, then
In this case, ${\theta_{\text{ATT}}^{\text{M}}}$ identifies ${\theta_{\text{ATT}}}$ but ${\theta_{\text{ATT}}^{\text{DIDM}}}$ and ${\theta_{\text{ATT}}^{\text{DID}}}$ tend to be upwardly biased. Second, when Condition DID is true, then
In this case, ${\theta_{\text{ATT}}^{\text{DID}}}$ identifies ${\theta_{\text{ATT}}}$ but ${\theta_{\text{ATT}}^{\text{M}}}$ and ${\theta_{\text{ATT}}}$ tend to be downwardly biased. Finally, when Condition DIDM is true, then
In this case, ${\theta_{\text{ATT}}^{\text{DIDM}}}$ identifies ${\theta_{\text{ATT}}}$ but ${\theta_{\text{ATT}}^{\text{M}}}$ tends to be downwardly biased while ${\theta_{\text{ATT}}^{\text{DID}}}$ tends to be upwardly biased.
Hence, for applications where non-negative treatment effects are expected, the DID estimand will generally incur optimistic estimates while the M estimand will produce conservative estimates. The DIDM estimand can be optimistic or conservative, depending on the underlying DGP. In summary, the M estimand is conservatively the most robust.
The main theoretical result presented in Section (ref) extends to a more general class of data-generating processes and associated estimands. In this section, we present an extension to general cases with a focus on the recent developments in the methods of event studies as principal examples.
The previous notations do not carry over to the current section. Suppose that a researcher is interested in identifying the average treatment effect on the treated (ATT) defined by
At this moment, we have not introduced the specific meanings of the notations. They will be discussed in the contexts of specific examples in Section (ref). With this said, we want to remark that they parallel with those notations introduced in Section (ref). Unlike the previous section, however, the subscripts no longer indicate the time in general.
Similarly to the previous section, we define the alternative estimands
called the matching (M), the difference-in-differences (DID), and the difference-in-differences matching (DIDM), respectively. The $p$-dimensional random vector $X$ is now used as a matching criterion.
The following conditions are imposed:
Condition (ref) requires that observed outcomes be the potential outcome without treatment for every unit prior to treatment. Condition (ref) requires that the observed outcome be the potential outcome without treatment for the control group.
In this section, we demonstrate that our general framework (ref)--(ref) encompasses alternative estimands studied in the literature of event studies as examples.
acemoglu2019democracy consider the ATT
where $D_t$ denotes the indicator of democracy, $Y_t^s(1)$ denotes the potential GDP in period $t+s$ when a country is treated between periods $t-1$ and $t$ (i.e., $D_t=1$ and $D_{t-1}=0$), and $Y_t^s(0)$ denotes the potential GDP in period $t+s$ when such a treatment does not occur (i.e., $D_t=D_{t-1}=0$).\footnote{The original paper by acemoglu2019democracy considers $ E\left[ (Y_t^s(1)-Y_{t-1}) - (Y_t^s(0)-Y_{t-1}) | D_t=1, D_{t-1}=0\right] $ as the parameter of interest, where $Y_{t-1}$ denotes the realized GDP at period $t$, but this is equivalent to (ref).} acemoglu2019democracy identify this ATT by
where $Y_t^s$ denotes the observed GDP in period $t+s$, $Y_{t-1}$ denotes the observed GDP in period $t-1$, and $X := (Y_{t-1},\ldots,Y_{t-4})'$ in their baseline model with additional covariates in extended robustness analyses.
Since $X$ contains $Y_{t-1}$ in particular, the conditioning theorem\footnote{Specifically, the conditioning theorem yields $E[Y_{t-1}|D_t=0,D_{t-1}=0,X]=Y_{t-1}$ when $X$ contains $Y_{t-1}$.} cancels $Y_{t-1}$ between the two terms in (ref), so the identifying formula (ref) of acemoglu2019democracy boils down to
Our general framework encompasses this example. Specifically, our ATT (ref) reduces to (ref) and our M estimand (ref) reduces to (ref) by setting
The pre-treatment condition (ref) is satisfied by construction, as $\widetilde Y_0(0) = Y_{t-1} = \widetilde Y_0$. The comparison condition (ref) is also satisfied by construction via the definition of $Y_t^s(0)$ as the potential outcome under $D_t=D_{t-1}=0$. Namely, $\widetilde Y_1(0) = Y_t^s(0) = Y_t^s = \widetilde Y_1$ holds given $D_t=D_{t-1}=0$.
callaway2018difference consider the ATT
where $G$ denotes the treatment period, $Y_t(g)$ denotes the potential outcome at period $t \geq g$ when an individual is treated at period $g$, and $Y_t(\infty)$ denotes the potential outcome at period $t$ when an individual does not receive a treatment. callaway2018difference identify this ATT by
for $g' \geq t+1$, where $Y_t$ denotes the observed outcome at period $t$. (The second term of (ref) may be aggregated over $G' \in \{t+1,t+2,\ldots\}$.)
Our general framework encompasses this example. Specifically, our ATT (ref) reduces to (ref) and our DID estimand (ref) reduces to (ref) by setting
The pre-treatment condition (ref) is satisfied by construction, as $\widetilde Y_0(0) = Y_{g-1}(\infty) = Y_{g-1} = \widetilde Y_0$ given $G=g$ or $G=g' \geq t+1 > g$. The comparison condition (ref) is also satisfied by construction, as $\widetilde Y_1(0) = Y_t(\infty) = Y_t = \widetilde Y_1$ given $G=g' \geq t+1$.
dube2023local consider the ATT
where $\Delta D_t$ denotes the indicator of policy change, $Y_{t+h}(1)$ denotes the potential outcome in period $t+h$ when a policy changes between periods $t-1$ and $t$ (i.e., $\Delta _t=1$), and $\Delta Y_{t+h}(0)$ denotes the potential outcome in period $t+h$ when such a change does not occur (i.e., $\Delta D_t=0$). dube2023local identify this ATT by
where $Y_{t+h}$ denotes the observed outcome in period $t+h$, $Y_{t-1}$ denotes the observed outcome in period $t-1$, and $X$ is a vector of general covariates.
Our general framework encompasses this example. Specifically, our ATT (ref) reduces to (ref) and our DIDM estimand (ref) reduces to (ref) by setting
The pre-treatment condition (ref) is satisfied by construction, as $\widetilde Y_0(0) = Y_{t-1} = \widetilde Y_0$. The comparison condition (ref) is also satisfied by construction via the definition of $Y_{t+h}(0)$ as the potential outcome under $\Delta D_t=0$. Namely, $\widetilde Y_1(0) = Y_{t+h}(0) = Y_{t+h} = \widetilde Y_1$ holds given $\Delta D_t=0$.
Also see the DID${}_\text{M}$ estimator of deChaisemartin2020two, and the (panel) matching estimator of imai2023matching, as well as the extended DID method of dube2023local -- they all propose and analyze the properties of what we refer to as the DIDM.
As pointed out by dube2023local, their framework encompasses acemoglu2019democracy as a special case. Indeed, when $X$ contains $\widetilde Y_0$, as is the case with acemoglu2019democracy presented in Section (ref), our DIDM framework reduces to our M framework. In general, however, the DIDM differs from the M.
Albeit there are slight differences in their notations, the three examples presented above focus on similar setups. They fundamentally differ only in terms of the estimands: the three examples focus on the M, DID, and DIDM estimands in our language. In general, a researcher does not know which of them achieves the identification. The M estimand identifies the true ATT (i.e, ${\theta_{\text{ATT}}^{\text{M}}} = {\theta_{\text{ATT}}}$ holds) if the matching condition
is satisfied. The DID estimand identifies the true ATT (i.e, ${\theta_{\text{ATT}}^{\text{DID}}} = {\theta_{\text{ATT}}}$ holds) if the parallel trend condition
is satisfied. Finally, the DIDM estimand identifies the true ATT (i.e, ${\theta_{\text{ATT}}^{\text{DIDM}}} = {\theta_{\text{ATT}}}$ holds) if the conditional parallel trend condition
is satisfied. In the absence of knowledge of the underlying data-generating process, however, committing to a wrong assumption can lead to biased estimates by M, DID, or DIDM. It is therefore of interest to characterize the relation among the three estimands. The following subsection investigates this point.
Now, focus on the generic framework (ref)--(ref) again. Let
be the identification errors of the estimands, ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DID}}}$, and ${\theta_{\text{ATT}}^{\text{DIDM}}}$, respectively. We establish the double bracketing relation $\Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}})$ under the following assumption.
The three parts (ref)--(ref) of this assumption parallel those in Assumption (ref), albeit that $X$ is now possibly multi-dimensional. Hence, similar interpretations can be made especially when $X$ consists of lagged outcomes as in the first example a la acemoglu2019democracy presented in Section (ref). Such a convenient interpretation may not be feasible if $X$ contains other covariates, but we want to stress that each of the three conditions (ref)--(ref) of this assumption is still empirically testable.
The following theorem states the extended double bracketing result for the general cases.
With Sections (ref)--(ref) providing theoretical justifications for the double bracketing relationship empirically observed in Section (ref), we will now examine whether the underlying assumption (Assumption (ref) or (ref)) of our theory is satisfied by the datasets used in Section (ref). Specifically, we revisit the CPS and PSID datasets from Section (ref) and the observational dataset from Section (ref).\footnote{Unfortunately, we are unable to present our analysis for the JTPA program discussed in Section (ref) due to discrepancies between the currently available microdata and the original microdata used by the authors, which we confirmed through repeated communications with them.}
Recall that Section (ref) demonstrates the robustness of the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \le {\theta_{\text{ATT}}^{\text{DID}}}$, for the NSW program using both the CPS data set and the PSID data set. In light of these consistent observations and our theoretical prediction provided in Theorem (ref), the current section examines each of the three parts, (i)--(iii), of Assumption (ref) in detail, using both data sets.
The first condition is Assumption (ref) (ref) requiring that $E[Y_0|W=0,Y_{-s}=y] \geq E[Y_0|W=1,Y_{-s}=y]$ holds for all $y$. To check this condition, we estimate the conditional expectation functions, $y \mapsto E[Y_0|W=0,Y_{-s}=y]$ and $y \mapsto E[Y_0|W=1,Y_{-s}=y]$, non-parametrically by the partitioning-based least squares regression.\footnote{We use the R package lspartition developed by cattaneo2019lspartition. We took all the default parameters.} Their estimates are plotted in Figure (ref).
The solid and dashed lines represent estimates of the conditional expectation functions $y \mapsto E[Y_0|W=0,Y_{-s}=y]$ and $y \mapsto E[Y_0|W=1,Y_{-s}=y]$, respectively. Shaded areas denote their 95% confidence bands.\footnote{We remark that the confidence bands are not centered around the estimates in lspartition, because of bias correction.} The top image in this figure is based on the CPS data set while the bottom one is based on the PSID data set. For both of the two data sets, the inequality $E[Y_0|W=0,Y_{-s}=y] \geq E[Y_0|W=1,Y_{-s}=y]$ is clearly satisfied for all $y$, providing evidence in support of our Assumption (ref) (ref).
The second condition, Assumption (ref) (ref), requires that $F_{Y_{-s}|W=0}$ first-order stochastically dominates $F_{Y_{-s}|W=1}$. We estimate the conditional cumulative distribution functions, $F_{Y_{-s}|W=0}$ and $F_{Y_{-s}|W=1}$, by their empirical counterparts, i.e., empirical CDFs conditionally on $W=0$ and $W=1$, respectively. Their estimates are plotted in Figure (ref).
The solid line represents the estimates of $F_{Y_{-s}|W=0}$ while the dotted line represents those of $F_{Y_{-s}|W=1}$. Their 95% confidence intervals are indicated by the shaded regions. The left figure is based on the CPS data set while the right one is based on the PSID data set. For both of the two data sets, the figure clearly shows that $F_{Y_{-s}|W=0}$ first-order stochastically dominates $F_{Y_{-s}|W=1}$, providing evidence in support of our Assumption (ref) (ref).
The third condition, Assumption (ref) (ref), requres the function $y \mapsto \Phi(y) := E[Y_1-Y_0|W=0,Y_{-s}=y]$ to be weakly decreasing. We estimate the conditional expectation function $\Phi$ non-parametrically by using the partitioning-based least squares regression -- see Footnote (ref) for details. The estimates, along with their 95% confidence bands, are plotted in Figure (ref).
The top figure is based on the CPS data set while the bottom one is based on the PSID data set. The confidence bands imply that the weak decreasingness of this function $\Phi$ cannot be refuted for either of the two data sets. Thus, they provide evidence in support of our Assumption (ref) (ref).
In the current section, we do not include auxiliary covariates in the current analysis. In Appendix (ref), we provide further evidence in support of our assumption even after accounting for auxiliary covariates.
In summary, all three components, (i)--(iii), of Assumption (ref) hold fairly robustly for both the CPS and PSID data sets regardless of whether auxiliary covariates are included or not. Hence, the double bracketing relationship $\Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}})$, as predicted by our Theorem (ref), is expected to hold, which we have empirically confirmed in Section (ref).
Recall that Section (ref) demonstrates that the double bracketing relationship, ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, is robust for the educational program investigated by athey2020combining. In light of this observation and our theoretical prediction provided in Theorem (ref), our next question is whether Assumption (ref) for the double bracketing theory is satisfied for this educational program. To address this, we will examine each of the three parts, (i)--(iii), of Assumption (ref) in detail using the data set utilized in Section (ref).
Recall that the first condition is Assumption (ref) (ref), requiring that $E[Y_0|W=0,Y_{-s}=y] \geq E[Y_0|W=1,Y_{-s}=y]$ holds for all $y$. We estimate the conditional expectation functions, $y \mapsto E[Y_0|W=0,Y_{-s}=y]$ and $y \mapsto E[Y_0|W=1,Y_{-s}=y]$, non-parametrically using the partitioning-based least squares regression -- see Footnotes (ref)--(ref) for details. The estimates are plotted in Figure (ref).
The solid and dashed lines represent estimates of the conditional expectation functions $y \mapsto E[Y_0|W=0,Y_{-s}=y]$ and $y \mapsto E[Y_0|W=1,Y_{-s}=y]$, respectively. Shaded areas denote their 95% confidence bands, although they are nearly invisible due to the large sample size.\footnote{We remark that the estimates appear outside of the confidence bands because the bands are centered around bias-corrected estimates in the lspartition package -- see Footnote (ref).} This figure showcases that the inequality $E[Y_0|W=0,Y_{-s}=y] \geq E[Y_0|W=1,Y_{-s}=y]$ is satisfied for all $y$, providing evidence in support of our Assumption (ref) (ref).
Next, recall that the second condition is Assumption (ref) (ref), which requires that $F_{Y_{-s}|W=0}$ first-order stochastically dominates $F_{Y_{-s}|W=1}$. We estimate the conditional cumulative distribution functions, $F_{Y_{-s}|W=0}$ and $F_{Y_{-s}|W=1}$, by their empirical counterparts, i.e., empirical CDFs conditionally on $W=0$ and $W=1$, respectively. Their estimates are plotted in Figure (ref). The solid line indicates the estimates of $F_{Y_{-s}|W=0}$ while the dotted line indicates the estimates of $F_{Y_{-s}|W=1}$. Their 95% confidence intervals are indicated by the shaded regions in colors, although they are nearly invisible due to the large sample size again. This figure demonstrates that $F_{Y_{-s}|W=0}$ indeed first-order stochastically dominates $F_{Y_{-s}|W=1}$, providing evidence in support of Assumption (ref) (ref).
The third condition is Assumption (ref) (ref), which requires $y \mapsto \Phi(y) := E[Y_1-Y_0|W=0,Y_{-s}=y]$ to be weakly decreasing. We estimate the conditional expectation function $\Phi$ non-parametrically by using the partitioning-based least squares regression -- see Footnote (ref) for details. The estimates, along with their 95% confidence bands, are plotted in Figure (ref). The plot shows that this function $\Phi$, is indeed non-increasing, providing evidence in support of of our Assumption (ref) (ref).
Our analysis in the current section omits auxiliary covariates. In Appendix (ref), we provide further evidence in support of our assumption even after accounting for auxiliary covariates. In particular, accounting for the covariates will strengthen the empirical support of our Assumption (ref) (ref) as mentioned above.
In summary, all the three components, (i)--(iii), of Assumption (ref) hold robustly for the educational program regardless of whether auxiliary covariates are included or not. Hence, the double bracketing relationship $\Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}})$, as predicted by our Theorem (ref), is expected to hold, as we empirically confirmed in Section (ref).
The paper evaluates the relative performance of three estimands -- Matching (M), Difference-in-Differences (DID), and a hybrid method (DIDM) -- for estimating causal effects in observational studies, particularly in the context of job training and educational programs. Our analysis reveals a consistent inequality: \( {\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}} \leq {\theta_{\text{ATT}}^{\text{DID}}} \). This indicates that Matching tends to produce the most conservative estimates, while DID often yields the most optimistic ones. When selecting a single method, it may be prudent to favor the more conservative Matching estimator. If the Matching estimator suggests a positive effect, this provides a compelling argument that the causal effect is indeed likely positive, leading to a more robust conclusion.
Moreover, these estimands can be utilized in a complementary manner. If practitioners believe that any one of the three identification assumptions -- unconfoundedness (for M), parallel trends (for DID), or conditional parallel trends (for DIDM) -- holds, then the true causal effect is bracketed by \( {\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}} \leq {\theta_{\text{ATT}}^{\text{DID}}} \). This approach offers a “triply robust bracketing” of the causal effect, providing valuable information regardless of which assumption is valid. By applying this framework, researchers can gain a clearer understanding of the potential range of treatment effects, with Matching yielding the most conservative estimate and DID providing the most optimistic one.
Furthermore, the results provide an intuitive strategy for managing uncertainty in identification assumptions. If the Matching (M) estimator indicates a positive effect, the sign of the treatment effect can be interpreted with confidence. The Difference-in-Differences (DID) estimator then provides an upper bound for the magnitude of this effect. This approach simplifies the analysis compared to traditional sensitivity methods (e.g., Manski and Pepper 2018), offering more interpretable bounds for researchers.