Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
62,019 characters · 14 sections · 76 citation commands
Choosing A Headline Estimand from Matching, DID, and Hybrid Designs: A Minimax-Regret Approach
\allowdisplaybreaks
Many policy evaluations in economics rely on non-experimental panel data. Large policy changes, such as mass layoffs, minimum wage reforms, the rollout of job-training programs, and major education policies, are typically not assigned at random. Instead, researchers observe repeated outcomes for the same units and exploit quasi-experimental variation in policy timing or exposure. The growing availability of large-scale and administrative data, together with a broader shift toward identification-focused empirical research, has made non-experimental panel and repeated-cross-section settings increasingly common in applied economics currie2020technology,goldsmithpinkham2024tracking.
When panel data are available, a natural and widely used way to address confounding is to exploit the lagged outcome itself. In job-training applications, recent earnings histories strongly predict both participation and future untreated earnings HeckmanSmith1999PreProgDip; in education settings, prior test scores play the same role chetty2014measuring1. The lagged outcome can be used in at least three distinct ways. A first approach is difference-in-differences (DID), which compares changes rather than levels lalonde1986evaluating,angrist2009mostly,bertrand2004did. A second approach conditions on lagged outcomes, for example via matching or flexible regression adjustment, so that identification rests on selection on lagged outcomes. We refer to this class as matching-type (M) estimands dehejia1999causal,dehejia2002propensity,heckman1997matching. A third, hybrid strategy combines the two ideas: it conditions on lagged outcomes and then differences over time, yielding a difference-in-differences-matching (DIDM) estimand heckman1998characterizing,abadie2005semiparametric,smith2005does,chetty2026creating.\footnote{We use the term difference-in-differences matching (DIDM) following heckman1997matching, smith2005does; recent applied work uses cognate labels, e.g., chetty2026creating term their design a “matched difference-in-differences.”} The three designs thus differ only in how they use the lagged outcome: DID differences it out, M conditions on its level, and DIDM does both. The identifying assumptions associated with these three designs are mutually non-nested restrictions on the joint distribution of potential outcomes and treatment.\footnote{In particular, the assumptions under which DID is valid do not imply those under which M is valid (and vice versa), and the assumptions for DIDM do not reduce to either DID or M. We illustrate this fact in Appendix (ref).}
These three designs account for a substantial share of published work using panel or repeated cross-section data to estimate causal effects of time-varying treatments. To gauge how common they are, we conducted a census of American Economic Review articles from 2020--2024 that use panel or repeated cross-section data.\footnote{We provide a detailed description of our methodology in Appendix (ref).} We find 77 studies employing panel-data identification strategies, and more than 80% use at least one of DID, M, or DIDM. About 70% implement some form of DID, roughly a third use matching or flexible conditioning on lagged outcomes, and just over 10% employ a hybrid DIDM design. More than a quarter of the in-scope papers employ more than one of these approaches within the same study.\footnote{The remaining papers primarily rely on other non-experimental strategies, such as regression adjustment under selection-on-covariates, dynamic panel instrumental variables, or synthetic control methods.}
There is, however, little formal guidance on how to choose among these alternative designs when experimental benchmarks are unavailable. In principle, applied researchers should select the estimand whose identifying assumptions are most credible in the application at hand. In practice, economic theory rarely delivers a single preferred specification. Dynamic models emphasize persistence and selection, making it natural to incorporate information on lagged outcomes, but they are typically too coarse to determine precisely how such information should enter the empirical specification.\footnote{A large literature on training programs, job search, and human capital emphasizes that both transitory shocks and persistent differences in ability shape participation decisions and earnings dynamics; see classic discussions of the “Ashenfelter dip” in training evaluations ashenfelter1978estimating and more recent work on dynamic selection and labor market histories HeckmanSmith1999PreProgDip,McCallSmithWunsch2016GovVocEd.} Reflecting this ambiguity, closely related empirical settings often make different choices among M, DID, and DIDM. Among job-displacement, minimum-wage, and value-added papers using similar administrative data and institutional environments, some studies rely on fixed-effects DID or event-study designs, while others explicitly condition on rich pre-treatment outcome histories or embed propensity-score matching and reweighting within DID-style estimators jacobson1993earnings,cengiz2019effect,rothstein2010teacher.\footnote{For displaced workers, jacobson1993earnings estimate long-run earnings losses using worker fixed-effects event-study DID in unemployment-insurance records, while couch2010earnings,hyslop2017impacts,illing2024gender,lachowska2020sources augment DID-type designs with propensity-score matching, reweighting, or controls for averages of pre-displacement earnings and hours. For minimum wages, health, and value-added, cengiz2019effect implement standard fixed-effects event-study DID, whereas kaminska2015effects,hafner2022minimum,lenhart2017uk,lenhart2017oecd,rothstein2010teacher,chetty2014measuring1,chetty2014measuring2,angrist2017leveraging,angrist2022methods use lagged-dependent-variable or matching-type estimators, propensity-score DID designs (whose scores often include lagged outcomes), or hybrids that combine lagged outcomes with gains-style differencing.} Even within a single paper, researchers often report multiple specifications, such as fixed-effects DID, lagged-outcome or matching estimators, and hybrid DIDM designs, and then compare them informally couch2010earnings,illing2024gender. We read this pattern as evidence that multiple strategies are considered credible for the same setting, and that a clear criterion for choosing among them is lacking.\footnote{For instance, couch2010earnings directly compare fixed-effects, random-growth, ATT, and “differenced ATT” estimators for displaced workers; hafner2022minimum present both fixed-effects DID and propensity-score DIDM estimators that explicitly match on lagged self-rated health before the German minimum wage reform; and kaminska2015effects,hyslop2017impacts,illing2024gender juxtapose regression-adjusted DID with matching- or reweighting-based DID in related administrative settings.}
In this paper, we provide formal criteria for choosing among M, DID, and DIDM in such environments. We consider a researcher who must commit to a single “headline” estimate but is uncertain about which of the three identifying assumptions (unconditional parallel trends for DID, selection on lagged outcomes for M, or conditional parallel trends for DIDM) is closest to the truth. Our main result shows that, under two economically interpretable conditions, the hybrid DIDM estimand is minimax-regret optimal among M, DID, and DIDM: it incurs the smallest worst-case loss, across the three possible identifying assumptions, relative to the estimand a researcher would have chosen had she known which assumption was correct. The two conditions are negative selection into treatment and stable (non-explosive) untreated outcome dynamics, which are common in labor and public economics applications. For a researcher seeking to limit the largest possible misspecification error without insisting that any one assumption holds exactly, committing to DIDM minimizes the worst-case deviation from the true average treatment effect on the treated.
These choices can matter substantively. In our empirical analysis in Section (ref), using canonical job-training settings such as NSW and JTPA lalonde1986evaluating,dehejia2002propensity,smith2005does,heckman1998characterizing, as well as the education application in athey2025combining, we show that switching among M, DID, and DIDM can materially change the estimated effects and, in some cases, even reverse the sign of the point estimates. A similar issue arises in recent work on banking deregulation. In a critique of boissel2022dividend, bach2023dividend identify the choice between DID and DIDM as one of two central points of contention: they report that, holding the data and most specification details fixed, replacing the hybrid DIDM design with a standard DID design renders the estimated positive treatment effect statistically insignificant.
Our minimax result is obtained through an intermediate analytical step that we believe is itself informative for applied work. We show that, under (i) negative selection into treatment (so that treated units would, on average, have had lower untreated outcomes than controls) and (ii) stable, non-explosive untreated outcome dynamics, the population estimands satisfy the same ordering: M $\le$ DIDM $\le$ DID. This result generalizes the insight in angrist2009mostly that the DID and lagged-dependent-variable estimands lie on opposite sides of the true effect, so that the truth is bracketed between them, by showing that the hybrid DIDM estimand lies systematically between the two endpoints. Across multiple program-evaluation settings spanning job training and educational interventions, and using four benchmark datasets lalonde1986evaluating,heckman1998characterizing,smith2005does,chetty2014measuring1,athey2025combining, we observe an empirical pattern consistent with our theory: matching-based estimates tend to be lower, DID estimates higher, and hybrid DIDM estimates lie in between.
To complement the estimand-level minimax-regret result, we consider a calibrated Monte Carlo design based on the NSW data from lalonde1986evaluating. The decision-theoretic question requires comparing the three candidate procedures across data-generating processes under which different identifying assumptions hold. We therefore construct three such environments, one each favoring M, DIDM, and DID, designed to remain observationally similar in the sense of matching key moments and cross-moment relationships in the data. This makes it plausibly difficult for a researcher to know which design is preferred from the observed data alone. The resulting \(3\times 3\) regret matrix then provides a compact summary of how each estimator performs across the three environments, and makes the minimax logic tangible: while no single procedure is pointwise best in every world, DIDM minimizes worst-case regret, that is, the largest amount by which it underperforms the best procedure for the realized world.
{\color{black}The framework has two main implications for applied work. First, it yields a principled default choice among common panel-data designs. When researchers must report a single headline estimate (for policy communication, executive summaries, or meta-analysis), the hybrid DIDM design provides a natural default because it minimizes worst-case regret across the three leading approaches. Second, it offers a structured way to interpret differences when multiple estimates are reported. }
This paper relates to three strands of literature.
First, it relates to the classical literature on nonexperimental evaluation following lalonde1986evaluating. Foundational contributions such as heckman1998characterizing,heckman1998_2matching, dehejia1999causal,dehejia2002propensity, and smith2005does study which nonexperimental methods best replicate experimental benchmarks in job-training settings. That literature compares specifications that, in our language, map naturally into M, DID, and DIDM-type estimands. Its main organizing question is typically which estimator has the smallest bias, often measured in absolute value, relative to the experimental ATT. Our paper asks a different question: how a researcher should choose among these competing observational estimands when the underlying identifying assumptions are mutually non-nested and no benchmark is available. The emphasis therefore shifts from ex post estimator comparison to ex ante design choice under model uncertainty.
This paper is also related to work such as chabe2017should,daw2018matching, which studies matching- and DID-based estimators in specific simulated or parametric environments. In particular, chabe2017should analyzes DID combined with conditioning on pre-treatment outcomes in a model with permanent and transitory confounders and in simulations calibrated to job-training settings, while daw2018matching uses Monte Carlo simulations to study regression-to-the-mean bias from matching on pre-period variables in DID designs. Relative to that literature, our contribution is twofold. First, our main comparative result is analytical and nonparametric. Second, whereas that literature studies performance within particular simulated environments, our contribution is decision-theoretic: we ask which headline estimand is safest when the researcher is uncertain which identifying restriction is closest to the truth. Under this model uncertainty, DIDM is minimax-regret optimal among the three leading panel-data estimands under a broad class of loss functions.
Second, our theory builds on the literature on bracketing between lagged-outcome and fixed-effects or DID estimands. In linear panel models, angrist2009mostly show that, in linear panel models, the lagged-dependent-variable and fixed-effects estimands bound the true effect from opposite sides, so that the truth lies between them. ding2019bracketing extend that insight to a nonparametric framework. Our paper contributes to this literature by introducing a third object, the hybrid DIDM estimand, and showing that under negative selection and stable untreated dynamics, \( \theta^M_{ATT} \;\le\; \theta^{DIDM}_{ATT} \;\le\; \theta^{DID}_{ATT}. \) DIDM reduces to M in the special case $s=0$, where the matching variable coincides with the differencing baseline ($Y_{-s}=Y_0$) and the $Y_0$ terms cancel. Our framework thus nests the familiar LDV-versus-FE bracketing as the case in which the middle object coincides with one endpoint, while extending the logic to the matched-DID designs (DIDM) that are common in practice.
Finally, the paper relates to the modern literature on event studies, DID, and panel matching. Recent surveys such as roth2023review and deChaisemartin2023two emphasize both the centrality of parallel-trends assumptions and the unresolved role of lagged outcomes and matching-type adjustments in DID practice. A related empirical literature uses lagged outcomes and other pre-treatment histories either to construct M-type estimands, as in acemoglu2019democracy, or to construct hybrid DIDM-type designs, as in deChaisemartin2020two, dube2023local, and imai2023matching. Our contribution to this literature is to place M, DIDM, and DID in a common framework, characterize when they are systematically ordered, and show how that ordering should guide design choice and interpretation in panel-data applications.
Let $W$ denote treatment-group status. Units with $W=1$ receive treatment between periods $t=0$ and $t=1$, whereas units with $W=0$ remain untreated throughout.\footnote{For expositional clarity, this section focuses on a two-group, single-treatment-timing setup. Appendix (ref) extends the framework to more general matching variables and to event-study or staggered-adoption settings by redefining $(\widetilde Y_0,\widetilde Y_1,W,X)$ appropriately; see especially Section (ref) and Sections (ref)--(ref). In such settings, cohort- and horizon-specific effects can be written as analogous DID- or DIDM-style contrasts and, when desired, aggregated by averaging the relevant comparisons across groups and horizons.} Hence, all units are untreated for $t \leq 0$, and only treated units are exposed to treatment for $t \geq 1$. Let $Y_t(w)$ denote the potential outcome at time $t$ under treatment status $w \in \{0,1\}$. The parameter of interest is the average treatment effect on the treated (ATT) at $t=1$, defined as \[ {\theta_{\text{ATT}}} \equiv E\!\left[ Y_1(1) - Y_1(0) \mid W = 1 \right]. \]
We write $Y_t$ for the observed outcome at time $t$. Throughout, we focus on the lagged untreated outcome $Y_{-s} = Y_{-s}(0)$, where $s \geq 0$, as the key matching variable. This choice reflects the emphasis in the evaluation literature on lagged outcomes as particularly informative predictors of both selection into treatment and the dynamics of untreated outcomes.\footnote{ Lagged outcomes play a central role in the empirical literatures motivating this paper. In job-training applications, recent earnings histories are highly predictive of both participation in treatment and future untreated earnings. In education settings, prior test scores play an analogous role. Our framework accommodates an arbitrary lag order $s \geq 0$. When $s=0$, DIDM reduces to M, thereby nesting the classical LDV-versus-FE bracketing logic of angrist2009mostly as a special case. When $s>0$, it encompasses the matched difference-in-differences strategies of heckman1998characterizing that condition on a lagged outcome measured $s>0$ periods before treatment, sometimes termed symmetric difference-in-differences.} Section (ref) extends the analysis to a general vector-valued matching variable $X$.
Following heckman1998characterizing, consider the following three observational estimands for ${\theta_{\text{ATT}}}$:
We refer to ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DID}}}$, and ${\theta_{\text{ATT}}^{\text{DIDM}}}$ as the matching (M), difference-in-differences (DID), and difference-in-differences matching (DIDM) estimands, respectively.
The three estimands differ only in how untreated outcomes are used to construct the missing counterfactual for treated units. The M estimand adjusts solely for selection on lagged outcomes, the DID estimand accounts only for average untreated outcome growth, and the DIDM estimand combines both approaches by conditioning DID-style comparisons on lagged outcomes.
Each estimand identifies ${\theta_{\text{ATT}}}$ under a distinct restriction on untreated potential outcomes.
First, the M estimand identifies ${\theta_{\text{ATT}}}$ if treatment assignment is conditionally independent of period-$1$ potential outcomes given $Y_{-s}$: \[ \text{Condition M:}\qquad (Y_1(1),Y_1(0)) \perp \!\!\! \perp W \mid Y_{-s}. \] Second, the DID estimand identifies ${\theta_{\text{ATT}}}$ under unconditional parallel trends: \[ \text{Condition DID:}\qquad E[Y_1(0)-Y_0(0)\mid W=1] = E[Y_1(0)-Y_0(0)\mid W=0]. \] Third, the DIDM estimand identifies ${\theta_{\text{ATT}}}$ under conditional parallel trends: \[ \text{Condition DIDM:}\qquad E[Y_1(0)-Y_0(0)\mid Y_{-s},W=1] = E[Y_1(0)-Y_0(0)\mid Y_{-s},W=0]. \]
\paragraph{Mutual Non-Nestedness of the M, DID, and DIDM Conditions:} The above three restrictions corresponding to M, DID, and DIDM are distinct and mutually non-nested. Formal arguments are provided in Appendix (ref). Here, we provide the intuition.
Condition M imposes a restriction on untreated levels conditional on lagged outcomes, whereas DID and DIDM impose restrictions on untreated growth rates. Accordingly, M may hold even when DIDM fails if treated and control units with the same $Y_{-s}$ share the same untreated outcome level at $t=1$ but exhibit different untreated growth between $t=0$ and $t=1$. Conversely, DIDM may hold while M fails if untreated growth is the same conditional on $Y_{-s}$, but untreated levels differ systematically across treatment status.
Likewise, DIDM and DID differ because DIDM imposes a conditional parallel-trends restriction, whereas DID imposes an unconditional one. DIDM may hold even when DID fails due to compositional differences across $Y_{-s}$ strata, while DID may hold even when DIDM fails if mutually offsetting conditional trend differences cancel out in the aggregate.
Thus, these assumptions are not ordered along a single robustness dimension, as they rule out different features of the data-generating process. In practice, a researcher ex ante does not know which restriction is most credible. The three estimands need not coincide, and each may be biased when its identifying condition fails. This gives rise to the design problem studied in this paper: when a researcher must select a single headline observational estimand from $\{{\theta_{\text{ATT}}^{\text{M}}}, {\theta_{\text{ATT}}^{\text{DIDM}}}, {\theta_{\text{ATT}}^{\text{DID}}}\}$, is there a principled default choice?
To address this question, the remainder of the paper proceeds in the following steps. Section (ref) documents that M, DID, and DIDM are recurring headline designs in applied work. Section (ref) then develops a nonparametric proposition giving conditions under which ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$, and documents (Section (ref)) that benchmark applications exhibit this ordering in practice. Section (ref) demonstrates that the ordering implies a minimax-regret rationale for DIDM as the default headline estimand.
Across several leading applied literatures, researchers addressing closely related causal questions employ empirical designs that map naturally into $ \theta_{ATT}^M,\, \theta_{ATT}^{DIDM},\, \text{ and } \theta_{ATT}^{DID}. $ Table (ref) summarizes canonical examples from job training, displaced workers, minimum wages, and teacher/school value-added. Our goal here is descriptive and selective: we use the table to show that the choice among M, DID, and DIDM recurs in applied work.
We classify a design as M when identification relies primarily on conditioning, matching, or reweighting based on lagged outcomes or other rich pre-treatment levels. We classify a design as DID when identification relies on a parallel-trends-type restriction implemented through first differences, fixed effects, or event-study specifications, without explicit conditioning on lagged outcomes. Finally, we classify a design as DIDM when it combines both elements, for example, by applying DID within a matched or reweighted sample, or by estimating a trends-based specification after explicitly conditioning on lagged outcomes. Here and throughout, we use M in a broad sense to include not only literal matching or reweighting estimators, but also lagged-outcome regression adjustments. While these specifications are not matching estimators in the narrow algorithmic sense, they share the same identifying logic: conditioning on pre-treatment outcomes so that treated units are compared to control units with similar outcome histories; see, for example, heckman1997matching and the balancing/weighting synthesis in doudchenko2016balancing.
The choice among M, DID, and DIDM has been central to program evaluation since its earliest non-experimental benchmarking exercises. Beginning with lalonde1986evaluating and continuing through heckman1998characterizing,heckman1998_2matching,dehejia1999causal,dehejia2002propensity,smith2005does, the job-training literature already compares all three designs against experimental benchmarks, making it a natural starting point for our framework.
The same design choice reappears in later applied work, which suggests that the question studied in this paper is relevant beyond the classical job-training context. In displaced-worker studies, canonical event-study specifications such as jacobson1993earnings employ DID, while later work incorporates matching and matched-DID hybrids; see, for example, couch2010earnings, hysloptownsend2019longer, lachowska2020sources, and schmieder2023costs. In minimum-wage applications, stacked event-study designs such as cengiz2019effect provide a canonical DID benchmark, whereas studies such as kaminska2015effects, hafner2022minimum, and arranz2025assessing combine matching or reweighting with DID-type comparisons. Finally, in teacher and school value-added research, lagged-score models map naturally into M, gains models into DID, and hybrid gains-plus-lagged-score specifications into DIDM.\footnote{In the value-added literature, empirical-Bayes or other shrinkage adjustments are conceptually distinct from the design taxonomy used here. Our classification concerns the identifying structure of the underlying causal signal--whether it is constructed from lagged-score conditioning (M), gains-style differencing (DID), or both (DIDM)--prior to any post-estimation shrinkage. Accordingly, the comparison in this paper is conducted in estimand space and speaks to identification bias across designs, rather than to a full finite-sample risk or MSE ranking that incorporates variance.} See, among others, kane2008estimating, rivkin2005teachers, rothstein2010teacher, chetty2014measuring1,chetty2014measuring2, and angrist2023methods.
The practical implication is that applied researchers repeatedly face the same design choice, often while targeting closely related causal parameters. This motivates a principled comparison among the three estimands. The next section establishes the ordering ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$ first theoretically and then empirically: a nonparametric proposition (Section (ref)) gives conditions under which it holds, and four benchmark applications (Section (ref)) show that it arises in practice. Section (ref) then shows that this ordering is what makes DIDM the minimax-regret choice among the three.
The goal of this section is to provide a formal condition under which the ordering $ {\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}} $ holds. This ordering serves as the key input for the minimax-regret decision problem studied in Section (ref). We begin with a nonparametric proposition, illustrate its intuition using a simple linear dynamic model, and then document the same pattern in benchmark applications.
Let
denote the identification errors of the three estimands relative to the target ${\theta_{\text{ATT}}}$. Then, the ordering \[ {\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}} \] is equivalent to the ordering \[ \Delta({\theta_{\text{ATT}}^{\text{M}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \leq \Delta({\theta_{\text{ATT}}^{\text{DID}}}). \]
We introduce the following assumption as a condition under which this ordering holds.
Assumption (ref) provides a nonparametric formulation of negative selection together with stable untreated dynamics. Part (ref) states that, conditional on lagged untreated outcomes, treated units have weakly lower period-$0$ untreated outcomes. Part (ref) requires that untreated units are positively shifted in terms of lagged untreated outcomes. Part (ref) stipulates that untreated growth is weakly smaller for units with higher lagged untreated outcomes. Importantly, the three conditions in Assumption (ref) have testable implications. In Appendix (ref), we show that these implications are not rejected in our job-training and education applications.
Proposition (ref) provides the main analytical input for the remainder of the paper. It delivers the ordering used in the minimax-regret argument developed below.
The proposition also yields an immediate interpretation under each candidate identifying restriction. If Condition M holds, then \[ 0 = \Delta({\theta_{\text{ATT}}^{\text{M}}}) \;\leq\; \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \;\leq\; \Delta({\theta_{\text{ATT}}^{\text{DID}}}), \] so ${\theta_{\text{ATT}}^{\text{M}}} = {\theta_{\text{ATT}}}$, while ${\theta_{\text{ATT}}^{\text{DIDM}}}$ and ${\theta_{\text{ATT}}^{\text{DID}}}$ weakly exceed it. If Condition DID holds, then \[ \Delta({\theta_{\text{ATT}}^{\text{M}}}) \;\leq\; \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) \;\leq\; \Delta({\theta_{\text{ATT}}^{\text{DID}}}) = 0, \] so ${\theta_{\text{ATT}}^{\text{DID}}} = {\theta_{\text{ATT}}}$, while ${\theta_{\text{ATT}}^{\text{M}}}$ and ${\theta_{\text{ATT}}^{\text{DIDM}}}$ weakly fall below it. If Condition DIDM holds, then \[ \Delta({\theta_{\text{ATT}}^{\text{M}}}) \;\leq\; \Delta({\theta_{\text{ATT}}^{\text{DIDM}}}) = 0 \;\leq\; \Delta({\theta_{\text{ATT}}^{\text{DID}}}), \] so ${\theta_{\text{ATT}}^{\text{DIDM}}} = {\theta_{\text{ATT}}}$ and lies between the other two candidate estimands. In all three cases, DIDM occupies the interior position implied by the double-bracketing structure.
To illustrate how Proposition (ref) can arise in a familiar setting, consider the following linear dynamic model:
with
Here, $\rho$ captures persistence in outcomes, $\gamma$ captures selection into treatment through baseline outcomes, and $\beta$ represents the causal effect of interest.
Assume:
Under (ref)--(ref), one obtains (with formal derivations found in Appendix (ref)) the gaps
Hence, ${\theta_{\text{ATT}}^{\text{M}}} \le {\theta_{\text{ATT}}^{\text{DIDM}}} \le {\theta_{\text{ATT}}^{\text{DID}}}$ holds.
The illustration also clarifies when DIDM coincides with an endpoint. If $\gamma = 0$, then ${\theta_{\text{ATT}}^{\text{M}}} = {\theta_{\text{ATT}}^{\text{DIDM}}} \leq {\theta_{\text{ATT}}^{\text{DID}}}$; if $\rho \in \{0,1\}$, then ${\theta_{\text{ATT}}^{\text{M}}} \leq {\theta_{\text{ATT}}^{\text{DIDM}}} = {\theta_{\text{ATT}}^{\text{DID}}}$. Outside such knife-edge cases, DIDM is generically distinct from both M and DID.
Section (ref) showed that empirical designs corresponding to \( \theta_{ATT}^M,\, \theta_{ATT}^{DIDM},\, \text{ and } \theta_{ATT}^{DID} \) appear across a wide range of applied fields. We particularly focus on four benchmark datasets: the NSW program with the CPS comparison sample, the NSW program with the PSID comparison sample, the JTPA program, and an education application based on athey2025combining and the related value-added literature chetty2014measuring1,chetty2014measuring2.
The first three datasets come from the job-training literature, the classical laboratory for comparing non-experimental estimators to experimental benchmarks; see lalonde1986evaluating, heckman1998characterizing,heckman1998_2matching, dehejia1999causal,dehejia2002propensity, and smith2005does. In these applications, the outcome variable is real earnings, and treatment corresponds to participation in a job-training program. The availability of experimental benchmarks makes it possible to assess the sign of the bias directly.
The education application serves a complementary role. In this setting, the outcome is student achievement, measured by standardized test scores, and treatment corresponds to assignment to smaller classes. Unlike the job-training benchmarks, its value does not primarily lie in comparison to an experimental benchmark. Rather, it provides both an external-domain validation of the same ordering and a high-precision environment in which the separation among ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DIDM}}}$, and ${\theta_{\text{ATT}}^{\text{DID}}}$ can be clearly observed, including across subgroups defined by observed characteristics.
Across these four datasets, we document a common empirical pattern: the estimates corresponding to M, DIDM, and DID satisfy \( {\theta_{\text{ATT}}^{\text{M}}} \;\leq\; {\theta_{\text{ATT}}^{\text{DIDM}}} \;\leq\; {\theta_{\text{ATT}}^{\text{DID}}} \) up to sampling uncertainty. In the job-training applications, this manifests as an ordering of signed biases relative to the experimental benchmark: matching-type estimands tend to be comparatively conservative, DID-type estimands comparatively optimistic, and DIDM lies in between. In the education application, the same ordering appears directly in the estimated effects across multiple subpopulations. Detailed descriptions of the data, institutional settings, variable construction, and estimation procedures are provided in Appendix (ref).
Figure (ref) consolidates the benchmark job-training evidence. The top panel presents the signed-bias version of the figure in chabe2017should using the NSW and JTPA experiments. The lower two panels report a broader collection of estimates from smith2005does for the NSW--CPS and NSW--PSID comparisons. Although the point estimates vary across specifications, the ordering of M, DIDM, and DID remains stable. For our purposes, the key point is that the same bracketing relationship repeatedly appears across benchmark job-training datasets and estimation choices.
Figure (ref) shows the corresponding pattern in the education application of athey2025combining. Across multiple subpopulations, the estimated effects continue to satisfy the same ordering, despite substantial variation in levels across groups.
Taken together, these four benchmark datasets point to a common empirical regularity. They show that the ordering in Proposition (ref) is a relationship that appears repeatedly in practice.
As noted earlier, applied researchers often report estimates of ${\theta_{\text{ATT}}^{\text{M}}}$, ${\theta_{\text{ATT}}^{\text{DIDM}}}$, and ${\theta_{\text{ATT}}^{\text{DID}}}$ side by side. If a researcher must select a single observational target as the headline estimand, a natural question is which choice is best when the most credible identifying restriction is unclear. We answer this by interpreting “best” as safest in the minimax sense.
Let ${\Theta_{\text{ATT}}} = \{{\theta_{\text{ATT}}^{\text{M}}}, {\theta_{\text{ATT}}^{\text{DIDM}}}, {\theta_{\text{ATT}}^{\text{DID}}}\}$, and let $L$ denote a loss function defined on ${\Theta_{\text{ATT}}} \times {\Theta_{\text{ATT}}}$. We interpret each element of ${\Theta_{\text{ATT}}}$ both as a possible action (the estimand a researcher reports as the headline effect) and as a possible state (the value of the true ATT), reflecting uncertainty about which identifying assumption is closest to the truth. For actions $a \in {\Theta_{\text{ATT}}}$ and states $\theta \in {\Theta_{\text{ATT}}}$, define the regret as \[ R(a,\theta) = L(a,\theta) - \min_{a' \in {\Theta_{\text{ATT}}}} L(a',\theta). \]
We impose the following condition on $L$.
This assumption accommodates a wide range of commonly used choices for $\ell$, including absolute loss, squared loss, power loss, Huber loss, $\varepsilon$-insensitive loss, exponential loss, logistic loss, Tukey's biweight loss, Cauchy loss, Welsch loss, and fair loss, among others. It requires only that the loss depends on the absolute deviation $|a - \theta|$ through a function $\ell$ that is weakly increasing in this deviation.
This theorem shows that ${\theta_{\text{ATT}}^{\text{DIDM}}}$ is the minimax-regret choice among the three alternatives in ${\Theta_{\text{ATT}}}$ for a researcher whose loss $L$ satisfies Assumption (ref). Hence, a researcher who seeks to minimize the worst-case regret over the three candidate identifying assumptions should adopt DIDM as the headline estimand. This recommendation is similar to, though distinct from, the use of midpoint estimators for interval-identified parameters song2014point in the partial-identification literature.
\paragraph{Interpretation:} M is safest only if one is primarily concerned about overstatement, while DID is safest only if one is primarily concerned about understatement. DIDM represents the middle option. Once regret is evaluated symmetrically in terms of the distance between the reported estimate and the true target, the middle option provides the best hedge against uncertainty regarding which identifying restriction is closest to the truth. \
Theorem (ref) is a statement about estimands. To make the regret ranking concrete, this section reports a simple Monte Carlo illustration calibrated to the NSW comparison samples. We construct three data-generating processes (“worlds”), one in which each of M, DID, and DIDM is the correctly specified design, and ask which estimand a researcher should report when she cannot tell the worlds apart.
\paragraph{The three worlds.} Let $Y_{-s}$ denote the lagged outcome, measured in thousands of dollars, and define a binary lagged-outcome index \[ L=2\cdot\mathbf{1}\{Y_{-s}>\operatorname{med}(Y_{-s})\}-1\in\{-1,1\}, \] so that $L=1$ for units with lagged outcome above the sample median and $L=-1$ otherwise. Let $U\in\{-1,1\}$ be an unobserved binary confounder, independent of $Y_{-s}$. The common assignment rule is \[ \Pr(W=1\mid Y_{-s},U)=\Lambda(aL+cLU), \qquad \Lambda(z)=\frac{1}{1+e^{-z}}. \] Untreated potential outcomes are \[ Y_0(0)=g_0(Y_{-s})+(m-q)U+\varepsilon_0, \qquad Y_1(0)=g_0(Y_{-s})+pL+mU+\varepsilon_1, \] where $g_0(y)=\alpha_0+\alpha_1 y$ is the linear regression of the period-0 untreated outcome on $Y_{-s}$, with $(\alpha_0,\alpha_1)$ estimated from the comparison sample under study (CPS or PSID). The implied untreated trend is \[ \Delta(0)=Y_1(0)-Y_0(0)=pL+qU+(\varepsilon_1-\varepsilon_0). \] Observed post-period outcomes are $Y_1=Y_1(0)+W\tau(g)$, where $g$ indexes the twelve subgroups defined by race, marital status, and high-school-degree status, and $\tau(g)$ is the treatment effect for subgroup $g$, set equal to the corresponding subgroup estimate in the NSW experimental sample. The DGP is calibrated once to the realized comparison sample; the Monte Carlo then varies only the simulation draws, so the exercise illustrates the population regret ranking of Theorem (ref) rather than quantifying estimation uncertainty.
The three candidate worlds use the same five coefficients $\theta=(a,c,p,q,m)$. One coefficient is then set to zero in each world: \[ \theta_M=(a,c,p,q,0), \qquad \theta_{DID}=(0,c,p,q,m), \qquad \theta_{DIDM}=(a,c,p,0,m). \] These restrictions encode the identifying logic of the construction. Setting $m=0$ removes the hidden post-period level channel, so matching is valid after conditioning on the lagged-outcome index $L$. Setting $a=0$ removes marginal selection on $L$, so the unconditional trend difference cancels by symmetry. Setting $q=0$ removes the hidden trend channel, so DIDM is valid after conditioning on $L$.
\paragraph{Calibration.} The free coefficients $(a,c,p,q,m)$ are calibrated to match a set of reduced-form moments of the comparison sample, chosen so that the simulated data resemble the real data on features an applied researcher could inspect. Because the structural coefficients are not separately observed, we anchor their magnitudes to sample quantities on the same scale. For each world $k$, we first reconstruct an estimated untreated post-period outcome by removing the subgroup treatment effect from the observed post-period outcome, \[ \widehat Y_{1,k}(0)=Y_1-W\,\widehat\tau_k(g), \qquad \widehat\Delta_k(0)=\widehat Y_{1,k}(0)-Y_0, \] where $\widehat\Delta_k(0)$ is the implied untreated trend. The calibration then targets: the standard deviations of $\widehat Y_{1,k}(0)$ and $\widehat\Delta_k(0)$, which anchor the scales of $m$, $p$, and $q$; the treatment shares within each value of $L$, which anchor the selection parameters $a$ and $c$; and the treated--control gaps in $\widehat Y_{1,k}(0)$ and $\widehat\Delta_k(0)$ within each $L$ cell, which anchor the hidden level and trend channels. Appendix (ref) states the exact moment vector and objective.
Table (ref) reports the calibrated coefficients. The unrestricted row gives the common coefficient vector before the world-specific zero restriction is imposed. The remaining rows show the actual coefficients used in each simulated world.
\paragraph{Results.} For each world $k$ we draw $B$ Monte Carlo samples. In draw $b$ we compute the sample analog of each candidate estimand and its mean squared error (MSE) relative to the draw-specific true ATT among the treated. Writing $\widehat\theta_{e,k,b}$ for the estimate of $e\in\{{\theta_{\text{ATT}}^{\text{M}}},{\theta_{\text{ATT}}^{\text{DIDM}}},{\theta_{\text{ATT}}^{\text{DID}}}\}$ in draw $b$ of world $k$ and ${\theta_{\text{ATT}}}_{k,b}$ for the corresponding true ATT, \[ \widehat r_{e,k}=\frac{1}{B}\sum_{b=1}^B \bigl(\widehat\theta_{e,k,b}-{\theta_{\text{ATT}}}_{k,b}\bigr)^2 \] is the MSE of estimand $e$ in world $k$. Regret subtracts the smallest MSE in the same world, and worst-case regret takes the maximum over worlds: \[ \widehat{\mathcal{R}}_{e,k} =\widehat r_{e,k}-\min_{e'\in\{{\theta_{\text{ATT}}^{\text{M}}},{\theta_{\text{ATT}}^{\text{DIDM}}},{\theta_{\text{ATT}}^{\text{DID}}}\}}\widehat r_{e',k}, \qquad \widehat{\overline{\mathcal{R}}}_{e}=\max_k\widehat{\mathcal{R}}_{e,k}. \] Table (ref) gives the resulting decision problem.
Table (ref) reports the underlying $3\times 3$ regret matrices. Each column subtracts the smallest MSE in that world, so every world has at least one zero. The “Worst” column gives the row maximum, and the minimax value is the smallest entry in that column (in bold).
In both comparison samples, ${\theta_{\text{ATT}}^{\text{DIDM}}}$ is the minimax-regret estimand. The diagonal zeros in Table (ref) show the intended pointwise pattern: ${\theta_{\text{ATT}}^{\text{M}}}$ is best in the $M$-world, ${\theta_{\text{ATT}}^{\text{DID}}}$ is best in the DID-world, and ${\theta_{\text{ATT}}^{\text{DIDM}}}$ is best in the DIDM-world. The decision criterion is minimax regret: \[ {\theta_{\text{ATT}}^{\text{DIDM}}} = \operatorname*{arg\,min}_{e\in\{{\theta_{\text{ATT}}^{\text{M}}},{\theta_{\text{ATT}}^{\text{DIDM}}},{\theta_{\text{ATT}}^{\text{DID}}}\}} \max_{k\in\{M,DID,DIDM\}} \widehat{\mathcal{R}}_{e,k}. \] The random-forest three-way accuracies, $0.367$ for NSW+CPS and $0.356$ for NSW+PSID, are close to the chance benchmark of $1/3$. The calibrated worlds therefore satisfy their intended identifying restrictions while remaining difficult to distinguish before the regret criterion is applied.
Researchers routinely choose among matching, DID, and hybrid DIDM designs in panel settings, yet applied work offers little formal guidance on which observational target should anchor the main result when experimental benchmarks are unavailable. This paper provides such guidance.
Our analysis delivers two related results. First, under two economically interpretable conditions, negative selection into treatment and stable untreated outcome dynamics, the three observational estimands satisfy the double-bracketing relation.
Second, once this ordering holds, DIDM is minimax-regret optimal among the three candidate headline estimands under a broad class of symmetric, distance-based loss functions. DIDM therefore emerges as a natural default when a researcher must report a single observational estimate while remaining uncertain about which identifying restriction is closest to the truth.
The main implication for applied work is that, when the double-bracketing logic is credible in a given setting, DIDM should be reported as the headline estimate, with matching and DID serving as lower and upper benchmarks.
The paper develops a decision-theoretic framework for choosing among common panel-data designs under uncertainty about identifying assumptions. The framework does not replace substantive judgment, but it shows that, in a large class of empirically relevant environments, researchers can make this choice in a disciplined way. When double bracketing is plausible, DIDM provides a robust default.
{6.8mm}