Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
150,178 characters · 14 sections · 164 citation commands
Efficient Difference-in-Differences and Event Study Estimators
{4pt} {3pt} {4pt} {3pt}
\thispagestyle{empty}
\setcounter{page}{1}
Difference-in-Differences (DiD) and Event Study (ES) designs are among the most widely used empirical strategies in economics and related fields. For instance, recent data indicates that over 30% of 2024 NBER applied microeconomics working papers mention DiD or ES---more than any other causal inference method,\footnote{The NBER data used in Goldsmith-Pinkham2024 is up to May 2024.} and their popularity is also rapidly expanding in empirical macroeconomics and finance Goldsmith-Pinkham2024. Although recent methodological advances have improved the robustness of DiD estimators to treatment effect heterogeneity,\footnote{See Roth2023_JoE_Review, deChaisemartin2022b_survey and Baker_etal_DiD_practioner_guide_2025 for recent reviews of DiD advances and a DiD practitioner's guide. See also Baker2022 for an overview with application in finance.} several empirically relevant econometric questions remain underexplored.
First, modern DiD and ES estimators are designed to accommodate rich treatment effect heterogeneity using short panel data. Unfortunately, most of them can have wide confidence intervals in applications with limited sample size.\footnote{See, e.g., Chiu2023_replicationDiD, Weiss2024_DiD_power, and Lal2025_TWFE for a discussion.} Whether the flexibility of modern DiD necessarily entails a meaningful loss in power compared to more restrictive estimators, such as two-way fixed effects (TWFE), remains an open question. Second, most heterogeneity-robust DiD implementations in short panels implicitly weight all pre-treatment periods equally or discard them entirely. These assumptions are generally unsupported by theory, data or subject-specific knowledge, and often yield imprecise estimators. Third, although it has become a common empirical practice to report several DiD and ES estimators,\footnote{See, e.g., Braghieri2022_facebook, Hansen_Wingender_2023_AERI, Arold_2024_QJE, and Mast2024_Restat.} there is limited formal guidance on how to compare them, especially so when applied researchers are unwilling to impose strong, arbitrary restrictions on temporal dependence or treatment effect heterogeneity. Sometimes, even basic questions such as whether different DiD estimators target the same causal parameters or if they rely on similar identification assumptions can be difficult to answer.
In this paper, we address these challenges by developing a unified framework for semiparametrically efficient estimation of DiD and ES causal parameters under point-identification assumptions---namely, parallel trends and no anticipation. We allow for several commonly-used DiD designs, including designs with a single treatment date and staggered treatment adoption, when covariates may or may not be important for identification. In particular, we (a) characterize the DiD potential outcome model in terms of equivalent restrictions on the joint distribution of observables, (b) derive the semiparametric efficient influence functions (EIF) for DiD and ES parameters and the efficient (i.e., smallest) variance bounds, (c) provide closed-form root-n asymptotically normal and efficient estimators as well as consistent estimators of the variances, and (d) show that semiparametric efficiency requires non-uniform weighting of pre-treatment periods and untreated cohorts. In addition, we highlight that the gains in efficiency/power can be empirically relevant in finite samples, even in designs without variation in treatment timing and covariates.
We start by providing an observationally equivalent characterization of the DiD potential outcome identification assumptions in terms of restrictions on the joint distribution of observables. This characterization links modern DiD designs to econometric models of sequential conditional moment restrictions with unknown functions of observables Ai_Chen_2012. It clarifies the informational content embedded in modern DiD identification assumptions, and enables our subsequent derivation of the semiparametric efficiency results for DiD models. Importantly, we show that DiD models are typically nonparametrically overidentified in the sense of Chen_Santos_2018_ECMA. To the best of our knowledge, this is the first result to establish such nonparametric overidentification in a causal inference setting without imposing parametric or semiparametric functional-form restrictions on nuisance components.
Next, we derive the semiparametric efficient influence functions for DiD and ES parameters in closed-forms, under various parallel trends assumptions commonly employed in the DiD literature, including settings with and without variation in treatment timing and cases where parallel trends hold only after conditioning on observed covariates. By definition, an EIF for a causal parameter has mean zero and its second moment is the semiparametric efficient variance bound, which is the smallest asymptotic variance across all possible root-$n$ consistent and asymptotically normal estimators under the DiD design identification conditions. Moreover, we provide closed-form estimators based on the EIFs to achieve the efficient variance bounds. To our knowledge, these are the first semiparametric efficiency results in DiD setups with multiple periods and without parametric functional-form assumptions. Importantly, these efficiency results only explore the DiD identification conditions and do not involve additional hard-to-justify restrictions on treatment effect heterogeneity (e.g., homoskedasticity) or the serial correlation of the outcomes.
Our efficient estimators for the DiD and ES are computed using sample EIFs, which are automatically Neyman orthogonal moments.\footnote{For nonparametrically overidentified models Chen_Santos_2018_ECMA, not all Neyman orthogonal moment-based doubly robust causal estimators are semiparametric efficient. Only the ones corresponding to the EIFs are semiparametric efficient.} The efficient estimators aggregate over multiple comparison groups and pre-treatment periods using analytically derived optimal weights that are proportional to the (conditional) covariances of outcome changes.
According to our semiparametric efficient variance bounds, DiD and ES estimators that treat all pre-treatment periods and comparison groups as equally informative or only use a particular comparison group and the last pre-treatment period as baseline, are generally inefficient, even in setups with a single treatment date. To achieve semiparametric efficiency, it is important to take into account that different untreated groups and pre-treatment periods are not equally informative, and that the best way to aggregate these information to gain precision involves constructing weights that depend on (conditional) covariance terms related to the outcome changes from different pre-treatment periods $t_{\text{pre}}$ to a given post-treatment period $t_{\text{post}}$ and different comparison groups. We introduce visualization tools that clarify how different baseline periods and comparison groups contribute to the estimator's efficiency, illustrating the structure implied by the semiparametric theory. Importantly, we stress that the geometry of these weights arises from our semiparametric efficiency results.\footnote{For large $n$ and large $T$ panel data with single treatment date design, Arkhangelsky2021_SDiD synthetic DiD exploits information from pre-treatment periods via a researcher-specified weighting criterion. Their weight is not for semiparametric efficiency consideration, however. See Remark (ref) for details.}
Our new semiparametric efficient variance bounds offer a rigorous benchmark for evaluating a broad class of DiD and ES estimators, including TWFE and those developed by deChaisemartin2020_AER, Callaway_Santanna_2021, Sun2021, Borusyak2023, Wooldridge2021a, Gardner2021, among others. Our proposed EIF-based estimators attain the efficient bounds, thereby dominating the existing estimators in terms of asymptotic efficiency. Simulation studies confirm the theoretical rankings.
We illustrate the practical relevance of our framework through empirically calibrated simulations and a real-data application. Our simulations build on DiD designs from Arkhangelsky2021_SDiD (for single-treatment) and Baker2022 (for staggered-adoption), calibrated to CPS and Compustat data, respectively. In both settings, our semiparametric efficient estimators deliver markedly lower root mean square error and narrower confidence intervals than alternative popular DiD estimators---often exceeding 40% gains in precision, with no loss in bias performance. We further revisit Dobkin_etal_2018_AER’s real-data analysis of hospitalization and out-of-pocket medical expenses, using the data compiled by Sun2021. In this application, alternative estimators would often require sample sizes at least 30% larger to achieve precision comparable to ours. These findings reinforce the central insight of the paper: aligning estimation strategies with the informational structure of the DiD model identification assumptions can yield substantial improvements in precision across a range of DiD settings.
The rest of the paper is as follows. Section (ref) formally introduces the general DiD designs and the parameters of interest. Section (ref) derives the semiparametric EIFs and the efficient variance bounds. Section (ref) proposes our semiparametric efficient estimation. Sections (ref) and (ref) present the simulation studies and real-data application, respectively. Section (ref) concludes with discussions of future extensions. Appendix (ref) introduces Hausman-type as well as incremental Sargan tests for the overidentification restrictions. Appendix (ref) extends the semiparametric efficiency results to an instrumented DiD setting. Appendix (ref) presents the proofs of the theoretical results.
Related Literature: This paper contributes to the popular DiD literature that accommodates treatment effect heterogeneity (see the references mentioned above). We provide the first semiparametric efficiency bounds for DiD and ES estimators in settings with multiple periods and varying forms of the parallel trends assumption, including covariate-conditional and staggered adoption designs. There are some recent work on efficient estimation under some extra conditions in DiD literature. Relative to Santanna2020, who studies a just-identified two-period, two-group DiD model, we generalize to more empirically relevant overidentified settings with multiple periods and staggered adoption. In contrast to Borusyak2023, Wooldridge2021a, and Harmon_2023_Efficiency, our efficiency results do not rely on auxiliary assumptions such as functional form, homoskedasticity, absence of serial correlation, or other time-series dependence restrictions. These assumptions are convenient but typically hard to justify on theoretical or empirical grounds, particularly in settings with heterogeneous treatment effects and dynamic responses.
Our semiparametric efficiency result builds on Ai_Chen_2012, as we characterize the DiD potential outcome model in terms of sequential conditional moment restrictions with unknown functions of observables.\footnote{Chamberlain_1992_ECMA_efficiency and Ai_Chen_2003 contain unknown functions of observables, but without sequential moments.} Without conditioning variables, the DiD potential outcome model can be equivalently transformed into Hansen's overidentified unconditional generalized moment restrictions Hansen1982_GMM, and the semiparametric efficient variance bound of Chamberlain_1987_JOE would also be applicable. Even in the case without covariates, our closed-form EIF expressions lead to simpler estimation procedure and provide new insights into how an efficient DiD estimator should weight different pre-treatment and comparison groups---to the best of our knowledge, these insights are new to the literature.
We start by describing our general setup. We consider a setting with $T$ periods indexed by $t \in \mathcal{T}= \{1,2,\dots,T\}$. At any given time $t>1$, units can start receiving a binary treatment, and different units can start receiving treatment at different points. We focus on setups where treatment is an absorbing state, so once a unit is treated, it remains treated until $t=T$. Let $D_{i,t}$ be an indicator for whether unit $i$ is treated by period $t$ and let $G_i = \min \{ t: D_{i,t} =1\}$ be a “group” variable (or cohort) that indicates the first period at which unit $i$ has received treatment. If unit $i$ does not receive treatment by $t=T$, we set $G_i = \infty$ and refer to these units as “never treated”. Since treatment is an absorbing state, $D_{i,t} = 1$ for all $t \geq G_i$. Let $\mathcal{G}$ denote the support of $G$ and let $\mathcal{G}_{\text{trt}} = \mathcal{G}\setminus \{\infty\}$ be the support of $G$ among “eventually-treated” units. Without loss of generality, we assume that a “never-treated” group always exists. If all units are eventually treated, we drop all the data from when the last cohort is treated, so the last-treated cohort becomes the “never-treated” cohort, and $T$ here denotes the number of available periods in the subset of the data that we will use in our analysis. Finally, we also assume that a vector of pre-treatment covariates $X_{i}$, whose support is denoted by $\mathcal{X}\subseteq \mathbb{R}^d$ is available. Throughout the paper, we adopt a “large-$n$, fixed-$T$” short panel data regime, as is typical in most DiD setups.
In this paper, we adopt the Neyman-Rubin potential outcome framework, indexing each potential outcome by the entire treatment path. Let $\textbf{0}_s$ and $\textbf{1}_s$ denote $s$-dimensional vectors of zeros and ones, respectively. We denote the potential outcome of unit $i$ in period $t$ if they were first treated at time $g$ by $Y_{i,t}(\textbf{0}_{g-1}, \textbf{1}_{T-g+1})$, and denote by $Y_{i,t}(\textbf{0}_{T})$ their “never-treated” potential outcome. To simplify notation, we can explore that treatment is an absorbing state and index potential outcomes by treatment starting time: $Y_{i,t}(g) = Y_{i,t}(\textbf{0}_{g-1}, \textbf{1}_{T-g+1})$ and $Y_{i,t}(\infty) = Y_{i,t}(\textbf{0}_{T})$. In practice, though, we observe $Y_{i,t} = \sum_{g \in \mathcal{G}} 1\{G_i = g\} Y_{i,t}(g)$, where $1\{A\}$ is the indicator function that takes value one if $A$ is true, and zero otherwise. Henceforth, we write $G_g = 1\{G = g\}$.
We assume that we observe a random sample of $(Y_{t=1},\dots,Y_{t=T}, X', G)'$.
We also impose overlap assumptions to avoid “extrapolation” and ensure that our identification arguments are nonparametrically valid.
We maintain the following no-anticipation condition throughout the paper.
Assumption (ref) states that, on average, units do not change their behavior before treatment starts, which essentially requires that the eventually treated units do not anticipate the treatment taking place. If treatment is announced in advance and units are expected to change their behavior due to it, this assumption suggests that researchers define the treatment date as the time of treatment announcements and not the time of treatment take-up. See Malani2015 for a discussion.
Many counterfactuals may be of interest in our context with potential variation in treatment timing. One particular family of causal parameters that have been popular in empirical work is the $ATT(g,t)$ defined by
Note that $ATT\left( g,t\right)$ measures the Average Treatment Effect at time $t$ of starting treatment at time $g$ versus not starting it, among the units that indeed started treatment at time $g$. Let $CATT(g,t,X) := \mathop{}\!\mathbb{E}\left[ Y_t(g) - Y_t(\infty) | G=g, X\right]$ denote the conditional $ATT(g,t)$ given covariates $X$.
In setups with a unique treatment date, there is a single treated group, $G=g$ with $g \neq \infty$, so tracking $ATT\left( g,t\right)$ over time for this treated group provides a measure of how treatment effects evolve with elapsed treatment time. This is usually referred to as event studies. With multiple groups defined by their treatment timing, $ATT\left( g,t\right)$ still provides information about treatment effect dynamics for each treated cohort $g\in \mathcal{G}_{\text{trt}}$. However, in such cases, researchers often want to summarize the average causal effects using weighted averages of $ATT\left( g,t\right)$ among units that started treatment $e$ periods ago, with $e \ge 0$. A natural aggregation uses the size of each treated cohort as weights, leading to the event study parameter that we denote by $ES(e)$,
One may also want to further aggregate the event study coefficients to recover a scalar summary measure. Let $\mathcal{E}$ denote the support of post-treatment event time $E=t-G$, $t\geq G$, and let $N_E$ denote its cardinality. Then,
provides a simple average of all post-treatment event study coefficients.
Under Assumptions (ref), (ref) and (ref) only, one cannot identify $ATT(g,t)$ for the post-treatment periods $t\ge g$, as $\mathop{}\!\mathbb{E}[Y_t(\infty)|G=g]$ is not observed or identified. DiD procedure tackles this problem by restricting the evolution of average untreated potential outcomes across the treated and comparison groups, a parallel trends (PT) assumption. Depending on the treatment timing and covariates, different versions of PT assumptions have been imposed in the literature. We discuss these below.
Two natural types of (conditional) PT have been adopted in the literature: one that effectively restricts trends of $Y_t(\infty)$ only in post-treatment periods and uses never-treated units as the relevant comparison group, and one that restricts trends of $Y_t(\infty)$ in both pre and post-treatment periods, and allow researchers to use any untreated group as a comparison group. We formally state these two PT assumptions below by allowing them to hold only after conditioning on covariates $X$, i.e., by allowing for covariate-specific trends. If covariates are unavailable or do not play an important role, one can take $X=1~ a.s.$ as a special case to resort to an unconditional DiD analysis.
Assumption (ref) only imposes PT in the post-treatment periods, {$t \geq g$}, and imposes that, conditional on covariates, the average untreated potential outcome evolution of any treated group $g\in \mathcal{G}_{\text{trt}}$ and the never-treated group is the same. This assumption has been effectively used by Callaway_Santanna_2021 and Sun2021 when constructing DiD estimators using never-treated units as the comparison group. In setups with a single treatment date without covariates, this is the assumption implicitly used in TWFE event study regressions Jacobson_etal_1993_AER_ES
where $ Y_{i,t}$ is the outcome for unit $i$ period $t$, $\alpha_t$ and $\eta_i$ are time and unit fixed effects often described as capturing time-invariant or common time-varying shocks, $E_{i,t} = t - G_i $ is the time relative to treatment (event-time), $\varepsilon_{i,t}$ is an idiosyncratic error term. The summation runs over all possible values of $E_{i,t}$ among eventually-treated units except for event-time $-1$.
Under Assumption (ref) (and Assumptions (ref), (ref), (ref)), the $ATT(g,t)$'s are identified as
In this case, $t_{\text{pre}} = g-1$ is the only reliable baseline period because the PT condition does not necessarily hold for the pre-treatment periods. In addition, Assumption (ref) prevents one from using untreated groups other than the never-treated as the comparison group. In such cases, even if more data on pre-treatment outcomes are available, or the relative size of the never-treated units is small compared to other untreated groups, this information cannot be used to improve the precision of DiD and ES estimators. But this raises the questions: When is PT plausible in post-treatment periods but not in pre-treatment periods? When does economic theory suggest we can only use a specific untreated group as a valid comparison group? The answers to these questions are probably application-dependent. But if we want to leverage more data to potentially get more informative inferences, we must strengthen Assumption (ref), leading us to Assumption (ref).
Assumption (ref) states that, conditional on covariates, the average untreated potential outcome follows the same path in all treatment groups and periods. It is perhaps the most popular identification assumption used in the literature, as invoked by deChaisemartin2020_AER, Sun2021, Gardner2021, Marcus2021, Borusyak2023, and Harmon_2023_Efficiency in the unconditional form, and Callaway_Santanna_2021, Wooldridge2021a, and Lee_Wooldridge_2023 in the conditional form. In setups with a single treatment date and no covariates, this assumption is arguably the one that justifies using a “static” TWFE regression specification
where $\beta = ES_{\text{avg}}$ is the parameter of interest. Unlike Assumption (ref), Assumption (ref) imposes a parallel pre-trends condition across all time periods.
As Assumption (ref) allows one to use different pre-treatments as baseline periods and leverage different sets of untreated units as a comparison group, the setup becomes much richer. The following lemma highlights how to combine comparison groups and pre-treatment periods to identify the $ATT(g,t)$'s, hinting that the DiD model becomes overidentified. Let $\mathcal{H}^{g,t} = \{(g',t_{\text{pre}},t_{\text{pre}}') \in \mathcal{G} \times \mathcal{T} \times \mathcal{T}: g>t_{\text{pre}}', g'>\max\{t_{\text{pre}}, t_{\text{pre}}'\} \}$.
Lemma (ref) leverages several restrictions on the indexes worth explaining. First, the restriction $t\ge g$ means we identify $ATT(g,t)$'s in group $g$ post-treatment periods. The restriction $g>t_{\text{pre}}'$ suggests using any of the group's $g$ pre-treatment period as a baseline, while the restriction $g'>\max\{t_{\text{pre}},t_{\text{pre}}'\}$ means that we are using the eventually treated cohort $g'$ that is treated after periods $t_{\text{pre}}, t_{\text{pre}}'$ as part of our effective comparison group. Interestingly, these restrictions allow $t_{\text{pre}}$ to be a post-treatment period for group $g$, as long as it is a pre-treatment period for cohort $g'$. Figure (ref) illustrates why this is possible in our setup using a stylized example with four groups and ten time periods---essentially, when parallel trends hold in all pre-treatment periods, one can “subtract back” some “excessive” corrections using the never-treated units.
Overall, Lemma (ref) indicates that one can use any group $g$'s pre-treatment period as a baseline, combine never-treated units with any not-yet-treated cohort $g'>g$ to form comparison groups, and use pre-treatment data for eventually-treated cohort $g'$ when identifying an $ATT(g,t)$. In practice, there are infinitely many ways one can combine these to identify the $ATT(g,t)$'s. In practice, however, one very relevant question is how to combine these different estimands to asymptotically obtain the most precise DiD estimator for the $ATT(g,t)$'s, or to obtain the most precise event study parameters $ES(e)$ as defined in (ref)? We discuss these points in the next section.
In this Section, we derive the semiparametric efficiency bound for the $ATT(g,t)$'s under Assumption (ref) and under Assumption (ref). We also provide analogous results for the event study parameters $ES(e), e\ge0$. These semiparametric efficiency bounds characterize the lowest possible asymptotic variance attainable by any regular, asymptotically linear estimator given the identification assumptions. Based on the efficient influence function, we derive closed-form expressions for estimands that can be used to construct efficient estimators under minimal smoothness assumptions. These results provide a benchmark for evaluating the asymptotic efficiency of existing DiD and ES estimators in the literature.
To build intuition, we first focus on the canonical case of a single treatment date. This simpler setting provides a transparent view of the key ideas and allows us to isolate the informational content of each pre-treatment period. We then generalize the results to the more complex case of staggered treatment timing. When covariates are uninformative for identification, the results further simplify to the unconditional DiD setup by setting $X=1$.
Consider the setting with a single treatment date, resulting in two groups: treated ($G=g$) and never-treated ($G=\infty$). In this case, the event study parameters $ES(e)$ are equivalent to $ATT(g,t)$'s with $t=g+e$. We proceed under either Assumption (ref) or Assumption (ref), noting that the former is a special case of the latter. With two groups, we have $G_\infty = 1-G_g$ and $p_\infty(X) = 1-p_g(X)$, with $p_g(X) = \mathop{}\!\mathbb{E}[G_g|X]$ being the propensity score, i.e., the probability of belonging to the treated group given covariates $X$.
Under Assumptions (ref), (ref), (ref), and (ref), Lemma (ref) implies that for any post-treatment period $t\ge g$, any pre-treatment period $t_{\text{pre}} \in \{1,\dots, g-1\}$, and any weight $w^{\text{att(g,t)}}_{t_{\text{pre}}}(\cdot)$ summing to one,
This expression shows how information from multiple pre-treatment periods can be used to identify our target parameters, but it does not yet indicate how to weigh each period to gain precision. To this end, the next step is to characterize all the information content in our identification assumptions. In the next lemma, we establish an equivalent representation of our identification assumptions that impose restrictions on potential outcomes using sequential conditional moment restrictions on observable variables.
Lemma (ref) lays the groundwork for deriving the efficient influence function and the semiparametric efficiency bound. Beyond that, it succinctly presents all the identifying power of our identification assumptions and highlights that one can construct many different DiD estimators based on it: it is just a matter of choosing an estimation method suitable for moment restrictions. Examples include (two-step) generalized methods of moments (GMM; see, e.g., Hansen1982_GMM, Ackerberg2014_two-step-gmm-Restud), generalized empirical likelihood Newey_Smith_2004_GEL, and minimum distance estimators Ai_Chen_2003, Ai_Chen_2007, Ai_Chen_2012.
A natural question is whether the choice of the estimation procedure matters in terms of asymptotic efficiency. As discussed in the previous section, the DiD model characterized in Lemma (ref) is overidentified in the sense of Chen_Santos_2018_ECMA. As such, not every nonparametric estimation procedure is asymptotically semiparametrically efficient, implying that it is relevant to characterize the minimum asymptotic variance for any regular, asymptotic linear (RAL) DiD estimator, and to consider estimators that achieve this bound.
Although different efficiency tools are available in the literature, many of them are not suitable for our DiD model in Lemma (ref). For instance, the DiD model in Lemma (ref) is based on sequential conditional moment restrictions with unknown nonparametric functions, making the efficiency results in Chamberlain_1992_ECMA_efficiency and Ai_Chen_2003 not directly applicable. In addition, note that the conditional average treatment effects among the treated units in period $t$ in our DiD model, $CATT(g,t,X)$, is given by
This suggests that our first-step nuisance function is overidentified, making the efficiency results of Ackerberg2014_two-step-gmm-Restud based on two-step GMM unsuitable for our context. On the other hand, we can build on Ai_Chen_2012, as their models are well-suited to our setup. Nonetheless, applying their results still requires substantial additional work: we need to first solve a calculus variations problem and then unravel several nested matrix inversions before arriving at the concise expression provided in Theorems (ref) and (ref).
In what follows, we present our semiparametric efficiency result under Assumption (ref), and then present as a special case the semiparametric efficiency under Assumption (ref). Let $\pi_g =\mathbb{E}[G_g]$ denote the population proportion of treated units, and denote the group-specific conditional expectations of the evolution of outcomes from period $t_{\text{pre}}$ to period $t$ by
Let $\mathbf{1}$ denote a column vector of ones with the appropriate length varying according to the context. Let $W = (Y_{t=1}, \dots, Y_{t=T}, X', G_g)$ denote the available random variables, and
For $1 \leq t_{\text{pre}} \leq g-1$, let
and denote the $g-1$ column vector that stacks of these $g-1$ influence functions by $\mathbb{IF}^{att(g,t)} =(\mathbb{IF}_{1}^{att(g,t)}, \mathbb{IF}_{2}^{att(g,t)},\cdots,\mathbb{IF}_{g-1}^{att(g,t)})'$. Finally, denote the $(g-1)\times (g-1)$ conditional covariance matrix of $\mathbb{IF}^{att(g,t)}$ as $V_{gt}(X) =\mathop{}\!\textnormal{Cov}(\mathbb{IF}^{att(g,t)}|X)$, and let $V_{gt}^*(X)$ be a matrix of the same dimension as $V_{gt}(X)$ with the $(j,k)$-th element being
In the proof of Theorem (ref), we show that including one post-treatment at a time and focusing on one $ATT(g,t)$ at a time leads to the same efficiency bound as including all post-treatments and estimating all $ATT(g,t)$'s. The main advantage of working with one $ATT(g,t)$ at a time is that it avoids dealing with a much larger variance-covariance matrix, which would preclude us from getting more intuition about efficiently aggregating each pre-treatment period. When it comes to estimation, this simpler characterization will allow us to avoid dealing with a much larger number of conditional moment restrictions that would lead to more computationally challenging procedures. Moreover, as a side result, this characterization allows us to present the semiparametric efficiency bound for the just-identified DiD setup when one imposes Assumption (ref) instead of Assumption (ref) in Corollary (ref) as a corollary of Theorem (ref).
Theorem (ref) and Corollary (ref) have several interesting implications. First, suppose we only use the pre-treatment period $t_{\text{pre}}=g-1$ as a baseline, so PT holds only for post-treatment periods (Assumption (ref)). Corollary (ref) shows that the efficient influence function is $\mathbb{IF}_{g,g-1}^{g,t}$ under Assumption (ref). This is effectively the same as dropping data from all periods but time $t$ (post-treatment) and time $g-1$ (pre-treatment), transforming the multi-period DiD model into a two-period DiD setup. Thus, under Assumption (ref), Corollary (ref) can be understood as generalizing the semiparametric efficiency bound result of Santanna2020 from the much simpler 2-group-2-period DiD setup to the multi-period DiD setup with a single treatment date. Importantly, Corollary (ref) highlights that pre-treatment data beyond the last pre-treatment period cannot be used to improve asymptotic efficiency, which can be hard to motivate empirically.
When one can leverage multiple pre-treatment periods (Assumption (ref)), Theorem (ref) highlights that the efficient influence function, in this case, is an optimally weighted average of all influence functions based on different baseline periods, where the optimal set of weights is computed by minimizing the conditional variance given $X$. Notably, the optimal way to aggregate pre-treatment periods for estimating $ATT(g,t)$ will generally differ from when estimating $ATT(g,t_{\text{pre}})$, with $t_{\text{pre}}\ne t$, with weights generally being non-uniform across pre-treatment periods and covariate strata. Of course, there are special cases where weighing all pre-treatment information equally is optimal. This is the case when outcome changes are (conditionally) uncorrelated across periods for both the treated and comparison group and have constant (conditional) variances across pre-treatment periods for both treatment groups (Wooldridge2021a, Borusyak2023). Outside this scenario, it is hard to see a setup where equal weights for all pre-treatment periods are asymptotically optimal. Our semiparametric efficiency results do not impose any of these strong, hard-to-empirically motivate assumptions. As highlighted in Lemma (ref), we do not impose additional restrictions beyond the identification assumptions, therefore being able to capture much richer notions of heterogeneity.
The efficient weights in Theorem (ref) also vary with covariates. This is intuitive, as the optimal way to combine pre-treatment periods may vary across covariate strata. For instance, it can be that for a particular partition of the covariate space, more recent pre-treatment periods are “more informative”, while for another partition, it can be that all pre-treatment periods are equally informative. The flexibility of the weights varying with $X$ is meant to capture exactly this. It is also worth mentioning that even when the covariance terms in (ref) are trivial functions of $X$, a type of homoscedasticity assumption within each treatment group, $V_{gt}^*(X)$ and, therefore, the optimal weights would still generally depend on $X$ via the propensity score $p_g(X)$.\footnote{It may be the case that the ratio of conditional covariance and the propensity score in (ref) is constant, even when all these terms are not trivial functions of the covariates. Nonetheless, this seems like a knife-edge case requiring a very specific covariance structure, which is unrealistic for most applications.} Thus, Theorem (ref) stresses the importance of choosing covariate-specific weights to achieve semiparametric efficiency.
The efficient influence function in Theorem (ref) also informs how one can construct DiD estimators that achieve the semiparametric efficiency bound by effectively weighting different pre-treatment periods. More specifically, we can leverage the fact that the efficient influence function has mean zero to get the following DiD estimand:
where $\widetilde{Y}_{g,t} = (\widetilde{Y}_{g,t,1}, \dots, \widetilde{Y}_{g,t,g-1})'$ is a $(g-1) \times 1$ column vector of transformed outcomes. Thus, the DiD estimand $ATT(g,t)$ uses a combination of pre-treatment periods as “effective” baselines, where the weights for each of these pre-treatment periods are given by $\mathbf{1}'V_{gt}^*(X)^{-1}\big/\mathbf{1}V_{gt}^*(X)^{-1} \mathbf{1}$. This expression highlights that, indeed, “more informative” baseline periods get weighted more than “less informative” periods, where the relevant notion of “informativeness” is given by the inverse of $V_{gt}^*(X)$ with elements defined in (ref).
We now extend our semiparametric efficiency bound results for ATT and ES estimators from the previous section to the staggered setups and provide guidelines on the most precise RAL DiD and ES estimators. We start studying the overidentified staggered DiD model for $ATT(g,t)$, and later discuss the just-identified one as a special case. As we discuss below, the efficiency of $ES(e)$, $e\ge0$, follows from the efficiency of $ATT(g,t)$'s and the Delta method.
As staggered treatment times imply the existence of multiple treatment groups, we now have that $G_\infty = 1 - \sum_{g \in \mathcal{G}_{\text{trt}}} G_g $ and $p_\infty(X) = 1-\sum_{g \in \mathcal{G}_{\text{trt}}} p_g(X)$, with $p_g(X) = \mathop{}\!\mathbb{E}[G_g|X]$ being a generalized propensity score, i.e., the probability of belonging to the group that is treated for the first time in period $g$ given covariates $X$.
The following lemma characterizes the information content of our staggered DiD identification assumptions based on potential outcomes using sequential conditional moment restrictions on observable variables.
The first two (unconditional) sets of moment conditions in Lemma (ref) define each cohort's relative size, $\pi_g$, and the $ATT(g,t)$ as a functional of the conditional $ATT(g,t)$. The third and fourth sets of (conditional) moment restrictions define the conditional $ATT(g,t)$ given covariates, $CATT(g,t,X)$, and the set of overidentification restrictions implied by our Assumptions (ref), (ref), (ref), and (ref). By combining the third and fourth conditional moment restriction, it is clear that any pre-treatment period $t_{\text{pre}}<g$ and multiple comparison groups can be used to identify the $CATT(g,t,X)$, and therefore, the $ATT(g,t)$. This is what essentially rationalizes Lemma (ref). The last set of conditional moment restrictions defines the generalized propensity scores, $p_g(X)$. Finally, also note that, from Lemma (ref), we can construct the event study parameters $ES(e), e\ge0$, as
where, for $e\ge 0$, $\mathcal{G}_{\text{trt},e} =\{g \in \mathcal{G}_{\text{trt}}: g + e \leq T\}$ denotes the set of eventually-treated groups that have data for $e$-periods after treatment started.
Similarly to Lemma (ref), Lemma (ref) presents all the identifying power of our identification assumptions in the staggered DiD setup and highlights that one can construct many different DiD estimators based on it. The relevant question is how to fully explore the empirical content of the moment conditions in Lemma (ref) to form DiD and ES estimators with appealing semiparametric efficiency guarantees. Here, in contrast to the setup in Lemma (ref), we can use different pre-treatment periods and different comparison groups when constructing DiD and ES estimators. Hence, the setup is more interesting, though it also requires additional notation to characterize how the “most precise” DiD and ES estimators should look.
Let the “generated outcome” for a given $ATT(g,t)$ be defined as
and, for $g' \in \mathcal{G}_{\text{trt}}$ and $1 \leq t_{\text{pre}} \leq g'-1$, let
Note that $\widetilde{Y}^{\text{att(g,t)}}_{g',t_{\text{pre}}}$ uses data from the earliest pre-treatment and from $t_{\text{pre}}$. It also leverages observations from the “never-treated” group $G_\infty$ and from $g'$ as the comparison group. When $g'=g$, though, it only uses data from $G_\infty$ as a comparison group, and, therefore, (ref) and (ref) reduce to (ref) and (ref), respectively.
Next, we collect all noncollinear $\mathbb{IF}_{g',t_{\text{pre}}}^{att(g,t)}$ to characterize the “relevant” set of influence functions to be used in the construction of the efficient influence function. For $g=g'$, let $\mathbb{IF}_{g}^{att(g,t)} =(\mathbb{IF}_{g,1}^{att(g,t)}, \mathbb{IF}_{g,2}^{att(g,t)},\cdots,\mathbb{IF}_{g,g-1}^{att(g,t)})'$, while for each $g' \ne g$, we do not include the very first time period and let $\mathbb{IF}_{g'}^{att(g,t)} =(\mathbb{IF}_{g',2}^{att(g,t)}, \mathbb{IF}_{g',3}^{att(g,t)},\cdots,\mathbb{IF}_{g',g'-1}^{att(g,t)})'$. Denote the vector stacking all these vectors together by $\mathbb{IF}^{att(g,t)}_{\text{stg}}$
and let its conditional covariance matrix be defined as $\Omega_{gt}(X) =\mathop{}\!\textnormal{Cov} (\mathbb{IF}^{att(g,t)}_{\text{stg}}|X)$. Let $\Omega_{gt}^{*}(X)$ be a matrix of the same dimension as $\Omega_{gt}$ with $(j,k)$-th element given by
where $(g'_s, t'_s)$ is the value that $g'$ and $t_{\text{pre}}$ takes in the $s$-th entry of $\mathbb{IF}^{att(g,t)}_{\text{stg}}$. Finally, let $q_{g,e} = \mathop{}\!\mathbb{P}(G=g|G +e \in [2,T])$.
The next Theorem establishes the semiparametric efficiency bound for $ATT(g,t)$ and $ES(e)$ parameters under (ref) and staggered treatment adoption. We present the semiparametric efficiency bound for the just-identified DiD setup when one imposes Assumption (ref) instead of Assumption (ref) in Corollary (ref).
Theorem (ref) and Corollary (ref) have some practical implications. First, when one is only willing to assume parallel trends for post-treatment periods, as in Assumption (ref), Corollary (ref) shows that this is essentially the same problem as in the single treatment setup and that $ATT(g,t)$ is just-identified: one can only use $g-1$ period as baseline, and the never-treated group as comparison group.
When one can leverage multiple pre-treatment periods as in Assumption (ref), Theorem (ref) highlights that the efficient influence function is an optimally weighted average of all influence functions based on different baseline periods and different comparison groups. Theorem (ref) also shows that the optimal way to pool information across periods and comparison groups is likely to be non-uniform, highlighting that not every comparison group and pre-treatment period has the same “relevance” when it comes to inference. Importantly, these weights are allowed to vary with the data-generating process and are derived as a consequence of our semiparametric efficiency considerations. As no currently available staggered DiD and ES estimator shares these weighting schemes, they do not generally achieve the semiparametric efficiency bound, suggesting that we can potentially conduct more informative inferences.
Theorem (ref) also provides a “blueprint” on how one can construct DiD and ES estimators that enjoy attractive semiparametric efficiency asymptotic guarantees. Similarly to the single-treatment date, we can explore the efficient influence function to get the following DiD $ATT(g,t)$ estimands:
where $\widetilde{Y}^{\text{att(g,t)}}$ is a column vector of all (non-collinear) generated outcomes $\widetilde{Y}^{\text{att(g,t)}}_{g',t_{\text{pre}}}$ that leveraged different $(g',t_{\text{pre}})$ pairs. The efficiency weights are given by $\mathbf{1}'\Omega_{gt}^*(X)^{-1}\big/\mathbf{1}\Omega_{gt}^*(X)^{-1} \mathbf{1}$. We recommend plotting the expected value of these weights so that one can have a better understanding of how each pre-treatment period and comparison group is leveraged for efficiency considerations. We do this in our simulations and empirical application.
Based on (ref), one can straightforwardly form estimands for event study parameters $ES(e), e \ge 0$, by plugging (ref) into (ref),
An analogous procedure follows for $ES_{\text{avg}}$.
In this section, we discuss how one can estimate and make inferences about the $ATT(g,t)$'s and $ES(e)$'s by leveraging the EIF-based estimands in (ref) and (ref). The estimator and inference procedures for DiD setups without covariates are straightforward, as they follow from standard plug-in arguments and can be seen as a special case of our proposal. The results from DiD setups with a single treatment time follow as special cases, too.
We first discuss estimation for $ATT(g,t)$'s. The EIF-based estimand $ATT_{\text{stg}}(g,t)$ in (ref) naturally suggests a two-step estimation procedure for $ATT(g,t)$'s, where one first estimate the nuisance parameters, $m(X) \coloneqq((m_{\infty,t,t_{\text{pre}}'}(X), t_{\text{pre}}' < t),(m_{g',t_{\text{pre}}',1}, g' \in \mathcal{G}_{\text{trt}}, t_{\text{pre}}' < g'))$ and $p_{\text{ratio}}(X) \coloneqq(p_g(X)\big/p_{g'}(X), (g,g') \in \mathcal{G}\times \mathcal{G})$, the conditional covariance matrix $\Omega_{gt}^*(X)$ (or $\Omega_{gt}(X)$), and then use their fitted values to form a plug-in estimators fo $ATT(g,t)$ based on (ref).
When one uses parametric working models for the nuisance parameters, it is easy to show that the resulting estimator is doubly robust in the sense that it remains consistent for the $ATT(g,t)$ as long as the working models for either (but not necessarily both) $p_{\text{ratio}}(X)$ or $m(X)$ are correctly specified, regardless of the weighting scheme adopted; see, e.g., Santanna2020 and Callaway_Santanna_2021.\footnote{In our case, as our estimator averages across different working models for the propensity score and regression adjustment for untreated units, it is possible to refine the notion of double robustness in the sense that our $ATT(g,t)$ estimator remains consistent if weighted averages of functionals of $p_{\text{ratio}}(X)$ or of $m(X)$ are consistently estimated. As we advocate for an efficiency-oriented weighting scheme, we would not require it to hold for all possible weighted averages (which would essentially require that $m(X)$ or $p_{\text{ratio}}(X)$ to be correctly specified).} Inference, in this case, follows from delta-method arguments.
In practice, however, parametric models may be too restrictive, and their misspecification can still lead to asymptotic biases. Thus, in what follows, we describe a flexible nonparametric estimation procedure that can leverage modern and traditional estimators for the nuisance functions.
First, recall that $m(X)$ involves conditional expectations of changes in outcomes over time among some untreated units, given covariates, i.e., it involves terms like $ m_{g',t_{\text{pre}}',t_{\text{pre}}}(X) = \mathbb{E}[Y_{t_{\text{pre}}'} - Y_{t_{\text{pre}}} |G=g',X]$ with $ g'> \max\{t_{\text{pre}},t_{\text{pre}}'\}$. This is nothing more than a nonparametric regression problem, and we can use a variety of estimators for it, including sieve-based Chen2007_Handbook, kernel-based Cattaneo_etal_npreg_JASA_2018, or recent machine learning methods such as random forests, lasso, ridge, deep neural nets, boosted trees, and the ensemble of these methods Chernozhukov2018_DML. Denote these estimators as $\widehat{m}(X)$.
Analogously, $p_{\text{ratio}}(X)$ involves terms like ratios of $\mathbb{E}[G_g|X]$, and each of these conditional expectations can also be estimated using these same estimation procedures.\footnote{It is often desirable to impose that $p_g(X)$ is bounded between zero and one and choose estimators that respect this constraint, such as multinomial series logit/probit and local multinomial logit/probit estimators. See, e.g., Chen2007_Handbook for a discussion of likelihood-based sieve estimators, and Staniswalis_kernel_logit_1989 for kernel-based ones.} Note that modeling each propensity score separately and then constructing their ratios can potentially lead to instabilities associated with estimated propensity scores close to zero. As the propensity scores enter into $ATT_{\text{stg}}(g,t)$ (ref) in a ratio format, it suffices to directly estimate $p_g(X)/p_{g'}(X)$, which is not bounded between zero and one, can be more flexibly modeled, and can lead to more stable procedures.
Towards this end, note that $p_g(X)/p_{g'}(X)$ can be found as the unique solution to the minimization problem:
This equivalence allows the use of \(\mathbb{E}\left[ r_{g,g'}(X)^2 G_{g'} - 2 r_{g,g'}(X) G_g \right]\) as a (convex) loss function to estimate $r_{g,g'}(X):=p_g(X)/p_{g'}(X)$, and different estimation procedures can be used. An easy-to-implement estimator for $r_{g,g'}(X)$ is the sieve-based estimator $\widehat{r}_{g,g'}(X) = {\left(\psi^K(X)'\widehat{\beta}_K\right)}$, where $\psi^K(x)$ is a $K$-dimensional vector of flexible transformations of the $X$ such as (tensor products of) cubic B-splines,
and we avoid indexing $\widehat{\beta}_K$ and ${\beta}_K$ by $(g,g')$ to simplify the notation burden. The ($(g,g')$-specific) sieve index \( K \) can be selected using some information criteria as follows: \[ \widehat{K} = \operatorname*{arg\,min}_{K} 2\mathbb{E}_n\left[G_{g'} \left(\psi^K(X)'\widehat{\beta}_K\right)^2 - 2G_g \left(\psi^K(X)'\widehat{\beta}_K\right) \right] + \frac{C_n K}{n}, \] where the information criterion corresponds to the Akaike information criterion (AIC) when \( C_n = 2 \), or the Bayesian information criterion (BIC) when \( C_n = \log(n) \). In the appendix, we show the consistency of the estimator based on $\widehat{K}$ following Chen_Liao_2014_JoE. Similarly to (ref), we can estimate $s_{g'} :=1/p_{g'}(X)$ by solving the empirical analog of the minimization problem
For instance, one can use sieve-estimators analogous to (ref) for this task.
To estimate the conditional covariance terms of $\Omega_{gt}^*$ as defined in (ref), we propose using a Nadaraya-Watson-type estimator based on kernel smoothing. Let $Ker$ be a kernel function on the covariates space and $h>0$ a bandwidth. Denote $K_h(\cdot) =Ker(\cdot / h)/h$. For $x = X_i$, we use $ \widehat{\mathop{}\!\textnormal{Cov}} (Y_t - Y_{t_{\text{pre}}},Y_t - Y_{t_{\text{pre}}'}|G=g',X=x)$ as an estimator for each the covariance terms $\mathop{}\!\textnormal{Cov} (Y_t - Y_{t_{\text{pre}}},Y_t - Y_{t_{\text{pre}}'}|G=g',X=x)$, where $ \widehat{\mathop{}\!\textnormal{Cov}} (Y_t - Y_{t_{\text{pre}}},Y_t - Y_{t_{\text{pre}}'}|G=g',X=x)$ is define as
Based on these estimators for the propensity score ratios and the conditional covariance terms, we estimate $\Omega^*_{gt}$ by $\widehat{\Omega}^*_{gt}$ with each $(j,k)$-th element given by the plug-in estimators of (ref).
Based on these estimators for the nuisance functions and weights, we can estimate $ATT(g,t)$ by\footnote{It is straightforward to consider variants of our estimator using sample-splitting arguments, too. }
where $\widehat{\widetilde{Y}}^{\text{att(g,t)}}_{\text{stg}}$ the estimated analog of ${\widetilde{Y}}^{\text{att(g,t)}}_{\text{stg}}$, with each $ \widehat{\widetilde{Y}}^{\text{att(g,t)}}_{g',t_{\text{pre}}}$ given by
and $\widehat{\pi}_g = \mathbb{E}_n[G_{g}]$.
The estimator for the event study parameter is
The following theorem establishes the large sample properties of our proposed DiD and ES estimators, and highlights that they achieve the semiparametric efficiency bound.
Theorem (ref) provides the pointwise asymptotic normality results for each $(g,t)$ pair, with $t\ge g$. In practice, researchers are often interested in making inferences about multiple groups and periods to better understand treatment effect dynamics and heterogeneity. In such cases, to avoid multiple-testing problems, it is recommended to construct simultaneous confidence bands. We omit the details of this procedure as it is standard.\footnote{See, e.g., Theorems 2 and 3, Algorithm 1, and Corollaries 1 and 2 in Callaway_Santanna_2021.} To compute standard errors, one can leverage a multiplier bootstrap procedure or take the square root of the average of the estimated EIF squared divided by the sample size.
When Assumptions (ref) and (ref) hold unconditionally, the efficiency results can be simplified by dropping $X$ from the conditioning set. This leads to an even simpler construction of the efficient estimator. In particular, ((ref)) reduces to
where each entry $\widetilde{Y}^{\text{att(g,t)}}_{g',t_{\text{pre}}}$ of $\widetilde{Y}^{\text{att(g,t)}}$ is given by
The covariance matrix $\Omega_{gt}^{*}$ has $(j,k)$-th element given by
where the indices $(g'_s, t'_s)$ are as defined in ((ref)).
In estimation, one can replace the expectations and covariances for each group with the corresponding within-group sample means and sample covariances. This procedure does not involve choosing tuning parameters or estimating conditional expectations.
When covariates are absent, the efficient GMM estimator also attains the semiparametric efficiency bound. Our DiD and ES estimators match that efficiency yet are substantially simpler to compute: they are available in closed form, whereas the optimal GMM approach entails solving a high-dimensional optimization problem that is often not practical. In addition, the construction of our efficient estimator provides new insights about how to weigh different pre-treatment and comparison groups to achieve efficiency---it is only a matter of analyzing the efficient weights ${\mathbf{1}'(\Omega_{gt}^*)^{-1}}\big/{\mathbf{1}'(\Omega_{gt}^*)^{-1} \mathbf{1} }$.
So far, we have motivated and established that our efficient DiD and ES estimators have attractive efficiency properties that allow researchers to get more precise inference procedures. In this section, we aim to see how these properties play out in realistic empirical settings. We build on Arkhangelsky2021_SDiD and Baker2022 and consider two sets of simulation studies calibrated to datasets people have used for DiD. The first set of simulations builds on Arkhangelsky2021_SDiD. It considers a single treatment date and leverages Current Population Survey (CPS) data for constructing a variety of outcomes and treatment assignments for a panel of US states. The second set of simulations builds on Baker2022. It explores Compustat data to calibrate outcomes and treatment assignments in staggered treatment setups with treatment timing varying across states. To simplify exposure and match the simulation designs in Arkhangelsky2021_SDiD and Baker2022, covariates play no important role in these Monte Carlo setups.
In our first set of simulations, we consider $ES_\text{avg}$ as in (ref) as the parameter of interest and compare our efficient DiD plug-in estimator based on (ref) (EDiD), traditional OLS-based TWFE estimator for $\beta$ based on (ref) (TWFE), average of post-treatment event-study TWFE estimators $\beta_e$ based on (ref) (DTWFE), and the Arkhangelsky2021_SDiD Synthetic DiD estimator (SDiD). We note that the DTWFE estimator would be the efficient estimator if parallel trends were to hold only in post-treatment periods.\footnote{In this non-staggered setup, the OLS estimate of $\beta$ based on (ref) coincides with the post-treatment average of the event-studies coefficients of the estimators proposed by Wooldridge2021a, Gardner2021, and Borusyak2023. Similarly, the post-treatment average of the OLS estimates of $\beta_e$ in (ref) coincides with Callaway_Santanna_2021, Sun2021, and deChaisemartin2020_AER estimators.}
In our second set of simulations, we also consider $ES_\text{avg}$ as in (ref) as the parameter of interest and compare our efficient (plug-in) estimator $\widehat{ES}_{\text{avg}}$ based on (ref) (EDiD), the average of post-treatment event-study estimates based on (i) Callaway_Santanna_2021 and Sun2021 DiD estimators using the “never-treated” as a comparison group (CS-SA), (ii) Callaway_Santanna_2021 and deChaisemartin2020_AER DiD estimators using the “not-yet-treated” as a comparison group (CS-dCDH), and (iii) Wooldridge2021a, Gardner2021, and Borusyak2023 “imputation” estimators (BJS-G-W). We do not include staggered synthetic DiD estimators in the comparison as we are not aware of a paper describing and establishing the statistical properties of synthetic DiD estimators for $ES_\text{avg}$. In this exercise, we note that the CS-SA estimator would be the efficient one if parallel trends were to hold only for post-treatment periods and the never-treated group was the only valid comparison group.
We compare all these estimators regarding bias, root-mean-squared-error (RMSE) relative to our efficient DiD estimator, $95\%$ empirical coverage, and $95\%$ confidence interval length relative to our efficient DiD. We consider analytical and bootstrapped standard errors for coverage and length of confidence intervals but use the Gaussian critical values. For each data generating process (DGP), we consider 1,000 Monte Carlo experiments, and for each experiment, we use 300 bootstrap draws.\footnote{For Synthetic DiD, Arkhangelsky2021_SDiD have not proposed a plug-in analytical standard error. As such, we do not report plug-in analytical standard errors from it.}
Our first set of simulations builds on Arkhangelsky2021_SDiD and explores CPS data to construct an empirically motivated data-generating process. We differ from Arkhangelsky2021_SDiD in some aspects: (i) we consider heterogeneous treatment effects across units; (ii) consistent with our DiD regime with “large $n$” and fixed $T$, we consider short panels with $T=7$---four pre-treatment and three post-treatment periods; (iii) we do not limit the maximum number of treated units in a given simulation; and (iv) all our outcomes are measured in log to avoid violating support restrictions in the simulation. Apart from these differences, the construction of our data-generating process follows from Arkhangelsky2021_SDiD, which we describe below for completeness.
As in Arkhangelsky2021_SDiD, we start designing our simulations using data on wages for women with positive wages in the March outgoing rotation groups in the CPS from 1979 to 2018. We then take logs and average the observations by state-year cells, leaving us with aggregate data from 50 states in 40 years. This will serve as the baseline dataset for our simulation designs.
Next, we consider a baseline specification for our untreated potential outcomes, $Y_{i,t}(\infty)$. We consider the same interactive fixed-effects specification as Arkhangelsky2021_SDiD,
with $\gamma_i$ and $\upsilon_t$ being 4-dimensional vectors of latent unit and time factors, and $\varepsilon_{i,t}$ being a mean-zero Gaussian error term that follows an $AR(2)$ process. This specification does not impose a two-way fixed effects structure on $Y_{i,t} (\infty)$ and allows for correlation over times within each state; accounting for serial correlation is very important when conducting valid inference using DiD and other panel data methods, see, e.g., Wooldridge2003 and Bertrand2004.
To construct a realistic set of simulations, we use the CPS aggregated data to estimate the $\gamma_i$'s, $\upsilon_t$'s, and the variance-covariance matrix of the error terms. In matrix notation, the interactive fixed-effects model (ref) is given by $ Y(\infty) = L + E$, where $L = \Gamma \Upsilon',$ and we estimate $L$ as
where $Y_{i,t}^*$ is the log wage in state $i$ in year $t$ in the CPS data. To help with interpretation, we decompose $L$ as unit-and-time fixed effects, $F$, and an interactive term $M$, with
We estimate the variance-covariance matrix of the residuals $\varepsilon_{i,t}$ by fitting an AR(2) model to the residuals of $Y_{i,t}^* - L_{i,t}$, and then compute the implied variance-covariance matrix of the error term for state $i$, $\Sigma$. Like Arkhangelsky2021_SDiD, we assume that the error terms are independent across units and $\Sigma$ is constant across states.
Next, we describe the treatment assignment process. In this set of simulations, units can start treatment in $t=2009$, so we only have two groups: $G_i=2009$ (treated units) and $G_i = \infty$ (untreated units). To simulate whether a state is treated, we follow Arkhangelsky2021_SDiD and consider the treatment status $D_{i,t} = 1\{t\ge 2009\} 1\{G_i = 2009\}$ with
We choose $\phi_\eta$ and $\phi_M$ as the coefficient estimates from a logistic regression of an observed binary characteristic of the state $i$ on $\eta_i$ and $M_i$.\footnote{These “unit factors” are the first four left singular vectors from single value decomposition of the outcome $Y$, as described in Arkhangelsky2021_SDiD.} We consider three different characteristics relating treatment groups to minimum wage laws (part of our baseline specification), abortion rights, and gun control laws. We also consider a completely random treatment assignment. Unlike Arkhangelsky2021_SDiD, though, we do not restrict the maximum number of treated units; only the minimum number of treated units to be two.\footnote{If we have less than two treated states in a simulation, we randomly select ten out of the 50 states to be treated completely at random. This is similar to Arkhangelsky2021_SDiD, though their minimum number of treated states was one.}
Based on these parameters, we simulate untreated and treated potential outcomes $Y_{i,t}(\infty)$ and $Y_{i,t}(2009)$ as
where $\varepsilon_{i} = (\varepsilon_{i,t=1},\dots, \varepsilon_{i,t=T})'$ have a multivariate Gaussian distribution with mean zero and variance $\Sigma$, and $\tau_i$ have a Gaussian distribution with mean zero and variance one. Note that we have a heterogeneous treatment effects model as the treatment effect varies across states. This differs from Arkhangelsky2021_SDiD, who sets $\tau_i = 0$ for all states. However, we stress that $ATT(2009,t) = 0$ for all periods. The observed outcomes for unit $i$ in time $t$ are $Y_{i,t} = 1\{G_i=2009\} Y_{i,t}(2009) + 1\{G_i=\infty\} Y_{i,t}(\infty)$. To respect our DiD setup with “large $n$ and fixed $T$”, we restrict $T=7$ and keep data from $t=2005$ until $t=2011$.
Note that we do not impose that the parallel trends assumption holds exactly across periods. If this assumption is violated, we should see a bias in estimates that rely on this assumption, namely our efficient DiD estimator and the DTWFE in (ref). If parallel trends in post-treatment periods is violated, we should also see a bias in the TWFE estimates based on (ref).
Like Arkhangelsky2021_SDiD, we consider different choices of treatment assignment and outcome variables. In addition, we also consider settings where we drop different components of the data generating process for the outcome, such as assuming that the error terms are serially independent (“No Corr”), that there is no interactive component $M$ (“No M”), no additive two-way fixed effects (“No F”), or there is $L$ (“Only noise”).\footnote{We omit results for DGPs in which there is no error term (“No noise”). We do it because the RMSE and the length of 95% confidence interval for our efficient estimators are close to zero, making it hard to report relative performance measures.} For the alternative outcome variables, we consider $log$ of hours worked and $log$ of unemployment rate; Arkhangelsky2021_SDiD consider these variables in levels, but that leads to some negative values during the simulation for these variables, violating their natural support restriction.
Finally, we consider two setups. We simulate data from the 50 states and seven periods for the first one. In this case, we expect that inference procedures based on asymptotic approximations may be challenging given the limited (effective) sample size of 50. In the second setup, we increase the number of cross-sectional units by drawing with replacement 200 states (each with its own error term, treatment assignment, and treatment effect). Here, we expect inference procedures based on asymptotic approximations to be better than in the setup with $n=50$, though we acknowledge that $n=200$ is still reasonably small.
Tables (ref) present the bias and relative RMSE of the four different estimators for $ES_\text{avg}$ when $n=50$ and $n=200$. At a high level, we have found that all estimators are (nearly) unbiased when $n=50$ and that the bias becomes even smaller when $n=200$. This suggests that, in these DGPs, the assumption that parallel trends hold in all periods is a good approximation. In terms of RMSE, however, there is a significant variation across estimators, with our efficient DiD estimator performing the best in nearly all DGPs, with Synthetic DiD coming as second, and the DTWFE estimator based on (ref) being the least precise. The result for DTWFE is expected, though, as this estimator does not leverage data from any pre-treatment period other than $t=2008$; all the other estimators leverage three additional pre-treatment periods. It is also important to stress the large RMSE gains that our efficient DiD estimator enjoys compared to the synthetic DiD estimator. For instance, when $n=50$, the RMSE of the synthetic DiD estimator is nearly 60% larger than that of our efficient DiD estimator in the baseline model. When treatment assignment is based on gun laws or abortion rights, their RMSE is approximately 5 times larger than ours when $n=50$, and this difference grows further when $n=200$. On the other hand, we notice that when the outcome of interest is the log unemployment rate, the synthetic DiD and the TWFE estimators have smaller RMSE than our efficient estimator. However, this difference reduces when $n=200$.
Table (ref) present the empirical coverage of $95\%$ confidence intervals for $ES_\text{avg}$ when $n=50$ and when $n=200$. We consider both nonparametric clustered-bootstrapped standard errors and analytical standard errors. Our simulation results show that using bootstrapped standard errors to construct confidence intervals yields very good coverage properties for all estimators across all considered DGPs, even in the challenging setup with $n=50$. On the other hand, our results also reveal that inference based on analytical standard errors with our efficient DiD estimator is anticonservative when $n=50$, though the results improve when $n=200$. Thus, in setups with very limited sample sizes, we recommend using cluster bootstrap standard errors.
Finally, Table (ref) presents the length of $95\%$ confidence intervals for all the estimators, using bootstrap and analytical standard errors. Since it is only reasonable to compare the length of confidence intervals across methods that control size, we focus our attention on the bootstrap results. Here, we note that our efficient DiD estimator tends to have substantially shorter confidence intervals than other available DiD estimators. For instance, in the baseline DGP, the bootstrapped confidence interval for the synthetic DiD estimator is more than $40\%$ larger than our efficient DiD estimator when $n=50$, and approximately two times larger than our efficient DiD estimator when $n=200$; these gains are even larger when compared to TWFE and DTWFE estimators. When treatment assignment is based on gun laws across states, the length of bootstrapped confidence intervals of all other considered estimators is more than four times larger than those based on our efficient estimator. We reach a qualitatively similar conclusion when using abortion rights as treatment. On the other hand, when the outcome of interest is the log unemployment rate, we note that bootstrap inference using synthetic DiD or traditional TWFE estimators is more precise than our efficient DiD estimator. However, this difference shrinks when $n=200$.
In practice, we anticipate that empirical researchers may also be interested in better understanding the contribution of each DiD component that uses a different pre-treatment period as a baseline period. In other words, one may be interested in understanding the “contribution” of each pre-treatment period to form efficient estimators. Figure (ref) plots this in a heatmap-style for four different DGPs, using a single simulation draw. As evident from this, the optimal way to aggregate pre-treatment periods to gain efficiency varies with the DGP, and the further it deviates from uniform weights, the higher the gains are compared to TWFE. This highlights that, in general, we can substantially improve upon TWFE.
Another point worth stressing is related to negative efficiency weights. In our context, negative efficiency weights do not lead to concerns related to treatment effect sign preservation as discussed in Goodmanbacon2021, Borusyak2023, and deChaisemartin2020_AER, for example. The reason for this is that our overidentification conditions imply that the $ATT(g,t)$'s are homogeneous across baseline periods, allowing us to potentially non-convex sums across different baseline periods. As illustrated in Figure (ref) and backed up by the results in Tables (ref) and (ref), this should not be interpreted as a concern. Negative efficiency weights arise simply as a consequence of the dependence structure of the changes in outcome in treatment and comparison groups, and de facto leveraging them can lead to substantial gains in precision, as illustrated in DGP 6.
Altogether, these simulation results highlight the excellent finite sample properties of our efficient DiD estimator in realistic DGPs: they tend to have much smaller RMSE and shorter confidence intervals than other available estimators. Importantly, the gains in precision can be very substantial.
Our second set of simulations builds on Baker2022 and considers DGPs with staggered treatment adoption calibrated to Compustat panel data. We differ from Baker2022 in some aspects: (i) consistent with our DiD framework with “large $n$” and fixed $T$, we consider a setup with $n=400$ firms and follow them for $T=11$ years; (ii) we consider that error terms are serially correlated with varying autocorrelation parameter $\rho$ to allow for richer outcome dynamics and assess its impact of the performance of different estimators; (iii) we follows a one-step treatment assignment in which firms are assigned to different cohorts. Apart from these differences, the construction of our data-generating process follows that of Baker2022, which allows for three treatment dates, with all units eventually being treated, and treatment effects evolving dynamically with heterogeneous trends. We describe the DGPs below for completeness.
As in Baker2022, we begin with a sample of all firms in Compustat over the 36-year period from 1980 to 2015 that are U.S. incorporated, non-financial, and contain at least five observations. This leaves us with a sample of 12,020 different firms. Using this unbalanced panel, we compute returns on assets (ROA) and then decompose it into year fixed effects, firm/unit fixed effects, and residuals:
The empirical (marginal) distribution of these terms serves as the benchmark distribution in our Monte Carlo simulations. We consider three possible treatment dates, $G=5$, $G=8$, and $G=11$, with each firm randomly assigned to each treatment group with probability $1/3$.
Based on these features, we simulate potential outcomes $Y_{i,t}(5)$, $Y_{i,t}(8)$, $Y_{i,t}(11)$, and $Y_{i,t}(\infty)$ as
where $\widetilde{\eta}_i$ and $\widetilde{\alpha}_t$ are drawn from the empirical distribution of $\{\widehat{\eta}_i\}_{i=1}^{12,020}$ and $\{\widehat{\alpha}_t\}_{t=1}^{36}$, respectively, $\widetilde{\varepsilon}_{i,t} = \rho \widetilde{\varepsilon}_{i,t-1} + \widetilde{u}_{i,t}$, with $\widetilde{u}_{i,t}$ being drawn from the empirical distribution of $\{\widehat{\varepsilon}_{i,t}\}_{it=1}^{176,622}$, and $\sigma_{ROA}$ refers to the sample standard deviation of ROA (0.309). To assess the impact of serial correlation on our results, we consider $\rho \in \{-1.1, -1, -0.5, 0, 0.5, 1, 1.1\}$.
The observed outcomes are given by $Y_{i,t} = \sum_{g \in \{5,8,11\}}Y_{i,t}(g)\times1\{G_i=g\}$. As we consider a setup where all units are eventually treated, we are unable to identify average treatment effects in period $ t=11$. Therefore, we effectively drop the data from the last period ($t=11$), and cohort $G=11$ acts as "never-treated" units. In our setup, we have that $ATT(5,t) = 0.154 (t-4)$ for $t=5\dots, 10$, and $ATT(8,t) = 0.093 (t-7)$ for $t=8,9,10$. As a consequence, we have that for $e=0,1,\dots, 5$, $ES(e)$ is equal to $0.123, 0.247, 0.370, 0.617, 0.772$, and $0.926$, respectively, and $ES_{\text{avg}} = 0.509$.
Table (ref) presents the bias and relative RMSE of the four different estimators for $ES_\text{avg}$. As expected, all estimators are unbiased. However, in terms of RMSE, there is a good amount of variation across estimators depending on the residual serial correlation. In all scenarios, our proposed efficient DiD estimator yields the lowest RMSE, which supports our theoretical results. When the errors are spherical ($\rho = 0)$, our efficient estimator performs very similarly to BJS-G-W. This is aligned with the efficiency results derived under these additional assumptions on the error term in Borusyak2023 and Wooldridge2021a. Outside this specific setup, though, the BJS-G-W estimator has no efficiency guarantees. Outside of our efficient estimators, our simulations also suggest that it is generally not possible to rank other estimators in terms of efficiency, as their performance depends on the specific DGP.
Table (ref) presents the empirical coverage of $95\%$ confidence intervals for $ES_\text{avg}$. All estimators display correct coverage across all considered DGPs. In terms of inferential precision measured by the length of the $95\%$ confidence intervals, the results once again highlight the practical gains of using our efficient DiD estimator. The relative size of these gains varies with the degree of serial correlation in the data, as expected and consistent with our theoretical results.
As we discussed before, an appealing feature of our proposal is that we can visualize the contribution of the treatment group, the effective comparison group, and the pre-treatment period used as baseline for the construction of our efficient DiD estimator for $ES_{\text{avg}}$. Figure (ref) plots this in a heatmap-style for three different DGPs, using a single simulation draw. Overall, we can see that the treatment group $G=5$ “contributes more” to the $\widehat{ES}_{\text{avg}}$, which is again expected, as this group has six available post-treatment periods; group $G=8$ has only three. Across DGPs, we can also see that the efficient weights vary across comparison groups and pre-treatment periods. When there is no serial correlation, most of the weights come from when $G=g'=5$, indicating that the efficient estimator is leveraging the “never-treated” units as a leading comparison group. In this case, we can also see that all pre-treatment periods have similar weighting schemes, again highlighting the appeal of exploring various pre-treatment periods as a baseline. When the error term follows a unit root process, we observe that the efficient estimator allocates almost all the weights to the last available pre-treatment periods, and earlier pre-treatment periods are less important for efficiency considerations. When the errors are negatively serially correlated with $\rho = -1$, we see that the weights are differently spread than in the previous case. Overall, these heatmap plots illustrate that the optimal way to aggregate information across treatment groups, comparison groups, and pre-treatment periods varies from DGP to DGP, and being able to (asymptotically) achieve the semiparametric efficiency bound without needing to know these additional features of the data can be very attractive. Similarly to the setup with a single treatment date, negative efficiency weights provide no reason to be concerned about bias in our context. These efficiency weights appear as a consequence of the semiparametric efficiency results, and leveraging their structure can lead to practical gains in precision.
We illustrate how our efficient DiD and ES estimators can be used by estimating the effect of hospitalization on out-of-pocket medical spending using publicly available survey data from the Health and Retirement Study (HRS) from the replication package of Dobkin_etal_2018_AER. Like Dobkin_etal_2018_AER, we explore the variation in the timing of hospitalization observed in the HRS to estimate the different treatment effects of interest. Dobkin_etal_2018_AER analyzes many other outcomes of interest and also leverage (not publicly available) hospitalization data linked to credit reports.
To conduct our illustrative analysis, we follow the same sample selection steps used by Sun2021, who also revisited Dobkin_etal_2018_AER. The sample selection closely follows Dobkin_etal_2018_AER, and restricts attention to non-pregnancy-related hospital admissions and to adults who are hospitalized at ages 50-59. Like Sun2021, we restrict our analysis to a subsample of individuals who appear throughout waves 7-11 (approximately spanning 2004-2012), so we maintain a balanced panel with 652 individuals in 5 waves, $t=7, \dots, 11$. Given the data construction, all individuals in the sample are hospitalized by $t=11$, but were not hospitalized in wave $t=7$. Thus, as we do not have a valid comparison group for wave $t=11$, we also drop data from that period.
Overall, we have 4 treatment groups: individuals (first) hospitalized in $t=8$ ($G_i = 8$, 252 individuals), $t=9$, ($G_i = 9$, 176 individuals), $t=10$ ($G_i = 10$, 163 individuals), and those hospitalized in $t=11$ ($G_i=\infty$, 65 individuals). Note that as we are only using data up to wave 10, we can re-label the units treated in wave $t=11$ as the “never-treated” group, so we match our notation.
As discussed in detail in Sun2021, (unconditional) parallel trends (in all periods and groups) and no-anticipation assumptions are plausible in this setting, and one should a priori expect that the effect of hospitalization is heterogeneous across individuals and periods. Figure (ref) plots the evolution of the average out-of-pocket medical spending across all treatment groups, and also provides suggestive evidence in favor of the plausibility of parallel trends assumptions across all periods and groups. Thus, the “heterogeneous robust” DiD and ES estimators for staggered treatment adoption proposed by deChaisemartin2020_AER, Callaway_Santanna_2021, Sun2021, Gardner2021, Wooldridge2021a, and Borusyak2023 are well-motivated. When parallel trends is plausible across all periods and groups, we should expect all modern DiD estimators to provide similar point estimates.
In this context, we are interested in different treatment effect parameters. As we discussed before, a natural starting point is to analyze the six different post-treatment $ATT(g,t)$'s that measure the evolution of ATTs since hospitalization for each group. For instance, $ATT(8,8)$ would give the average treatment effect of hospitalization on out-of-pocket medical spending in time $t=8$, among individuals in group $g=8$, i.e., those who were hospitalized in time $8$; $ATT(8,9)$ and $ATT(8,10)$ would provide the average treatment effect for these same individuals but in periods $t=9$ (one year after hospitalization) and $t=10$ (two years after hospitalization). The interpretation of the other $ATT(g,t)$'s is analogous. We may also want to summarize all these different $ATT(g,t)$'s into more aggregated summary measures to ease interpretations and potentially gain precision. A natural way to summarize the effects is to look at the event-study parameters $ES(e)$ as in (ref)---which provides an average treatment effect by length of treatment exposure $e$---and their averages, $ES_{\text{avg}}$, as in (ref).
Table (ref) reports the estimates of all these target parameters and displays their standard errors clustered at individual level in parenthesis. We report estimates using different methods: our efficient DiD procedure (EDiD), the DiD estimator of Callaway_Santanna_2021 and Sun2021 that uses the never-treated units as the comparison group (CS-SA), the one of Callaway_Santanna_2021 and deChaisemartin2020_AER that uses the not-yet-treated units as the comparison group (CS-dCDH), and the imputation DiD estimator of Borusyak2023, Gardner2021, and Wooldridge2021a (BJS-G-W). Several overall patterns arise. First, note that the point estimates for each target parameter are similar across all the different DiD estimators. This is expected, as parallel trends across all periods and groups and no-anticipation are plausible in this application. We also note that instantaneous treatment effects tend to be larger than later effects, across all cohorts; for instance, estimates of $ATT(8,8)$ tend to be larger than $ATT(8,9)$ and $ATT(8,10)$, and $ATT(9,9)$ larger than $ATT(9,10)$. This is well captured in the event-study estimates, $ES(0)$, $ES(1)$, and $ES(2)$, indicating that the effect of hospitalization on out-of-pocket medical spending is mostly transitory. This is in line with the findings in Dobkin_etal_2018_AER.
Another overall pattern that emerges from the results in Table (ref) relates to the inference precision of these estimators, as measured by their standard deviations. In line with our theoretical results, our efficient DiD estimators have the smallest standard errors across all target parameters. Outside our efficient DiD procedure, though, no other estimator uniformly dominates the remaining ones in terms of length of confidence intervals. In fact, CS-SA, CS-dCDH, and BJS-G-W take turns into leading to the second most precise estimator across target parameters, with CS-SA being “second best” for $ATT(9,9)$ and $ATT(9,10)$, CS-dCDH coming second for $ATT(8,8)$, $ATT(9,10)$, $ES(0)$, $ES(1)$, and $ES_{\text{avg}}$, and BJS-G-W for $ATT(8,9)$, $ATT(8,10)$, $ATT(10,10)$, and $ES(2)$. This highlights that relying on additional assumptions, such as homoskedasticity and restrictions on serial correlation, to justify efficiency gains may be practically hard. Our proposed efficient DiD estimator does not rely on these conditions and, in fact, delivers more precise inference procedures.
But one may wonder: do these efficiency gains matter in practice? To answer this practically relevant question, we display in Table (ref) estimates of the asymptotic relative efficiency (ARE) of our proposed efficient DiD estimator with respect to the other available DiD estimators.\footnote{For any parameter $\eta$ of a distribution $F$, and for estimators $\widehat{\eta}_{1}$ and $\widehat{\eta }_{2}$ approximately $N\left( \eta,V_{1}/n\right) $ and $N\left( \eta, V_{2}/n\right) $, respectively, the asymptotic relative efficiency of $\widehat{\eta}_{2}$ with respect to $\widehat{\eta}_{1}$ is given by $V_{1}/V_{2}$; see, e.g., Section 8.2 in VanderVaart1998.} Heuristically speaking, the ARE provides a relative measure of sample size needed for other DiD estimators to achieve the same precision as our efficient DiD estimator. The results in Table (ref) highlight that the gains in ARE of our efficient DiD estimator tend to be large. For example, for $ATT(8,8)$, CS-SA, CS-dCDH, and BJS-G-W DiD estimators would respectively need a sample size with $65\%$, $28\%$, and $29\%$ more hospitalized individuals than our EDiD estimator to achieve the same precision as EDiD. For $ATT(8,9)$, the gains are even larger, with CS-SA, CS-dCDH, and BJS-G-W DiD requiring $104\%$, $83\%$, and $45\%$ more hospitalized individuals than our EDiD estimator to equalize precision. These ARE gains remain large for the event-study coefficients, and their average.
To gain further insights into how our efficient DiD and ES estimators achieve efficiency gains, in Figure (ref) we report the weights attached to the different (effective) comparison groups and pre-treatment periods for each $ATT(g,t)$ of interest. Overall, there are four different (non-collinear) $ \widehat{\widetilde{Y}}^{\text{att(g,t)}}_{g',t_{\text{pre}}}$ as defined in (ref) that one can use to estimate each $ATT(g,t)$, and the weights attached to them sum up to one, i.e., the weights in Figure (ref) sum up to one in each row. If we were to weigh each of these uniformly, each of them would have a 0.25 weight. In practice, we can see that our efficient DiD estimators never use these uniform weights, and, in fact, these weights vary according to the parameter of interest. This, by itself, highlights the appeal of our procedure to “adapt” to these different scenarios with the researcher needing to take a stand on additional hard-to-justify assumptions.
It is also interesting to analyze these weights separately. We start with $ATT(8,8)$. For this parameter, the efficient estimator puts weights of 0.29 on the DiD estimator that uses only the never-treated (in our case, cohort treated in wave 11) as a comparison group ($G=g'=5)$, 0.54 on the DiD estimator that only uses cohort $G=9$ as a comparison group ($g'=9,t'=8)$, and 0.03 on the DiD estimator that uses cohort $G=10$ as the comparison group. Both of these use period $t=7$ as the baseline. Interestingly, the efficient DiD estimator for $ATT(8,8)$ puts 0.13 weight on the DiD estimator that uses “never-treated” and $G=10$ cohort as comparison group and leverages the first period as pre-treatment periods as well as $t_{\text{pre}}=9$, illustrating the “bridging” phenomenon discussed in the stylized example in Figure (ref). Here, note that $t_{\text{pre}}=9$ is a post-period for $G=5$ but a pre-treatment period for group $G=10$. The flexibility of our efficient DiD estimator to leverage this additional information from the data leads to the efficiency gains highlighted in Table (ref).
For cohort $G=8$ in period $t=9$, we note that the efficient DiD estimator weights the DiD estimator based on never-treated units (i.e., the CS-SA estimator) by 0.30, the DiD estimator that leverages never-treated and $G=9$ as comparison groups by 0.49 (0.48), and the DiD estimator that leverages only $G=10$ as a comparison group by 0.25. An interpretation for the weights for $ATT(8,10)$ is analogous. For $ATT(9,9)$ and $ATT(9,10)$, note that the efficient DiD estimator puts nearly 0.90 weight on the CS-SA DiD estimator, explaining why its performance is similar to theirs (but better than the other alternative DiD estimators). For $ATT(10,10)$, our efficient estimator put 0.58 weight on the CS-SA estimator, 0.04, and 0.19 weight on the DiD estimator that uses never-treated as comparison and period eight and seven as baseline period, respectively, and the remaining 0.20 weight on the DiD estimator that combines never-treated and the cohort treated in wave $9$ as comparison groups and use first two periods as baseline periods. Together, these flexible weighting schemes demonstrate that our efficient DiD estimators are flexible and capable of extracting information from the data to provide the gains in precision that our theoretical results indicate.
We conclude the session by discussing how one can use our results to assess the stability of our results. In Figure (ref), we plot our efficient ES estimators with their 95% confidence intervals in blue, and in gray, we plot all possible event-study estimators that attach unit weight to the different $ATT(g,t)$ estimators in Figure (ref) that uses a particular $(g', t_{\text{pre}})$ pair. For $ES(0)$, that are 64 possible estimators ($4\times4\times4$) that leverage different combinations of $ATT(8,8)$, $ATT(9,9)$, and $ATT(10,10)$. For $ES(1)$, we have 16 possible estimators, and for $ES(2)$, we have four possible estimators. As you can see in Figure (ref), all these different combinations lead to very stable event-study estimators, indicating the plausibility of Assumption (ref) in our context. Figure (ref) can be interpreted as an adaptation of the specification curve analysis Simonsohn_Simmons_Nelson_2020_spec_curve to our DiD context.
Empirical researchers routinely face a variety of alternative DiD and ES estimators to choose from. Existing approaches often overlook the fact that different pre-treatment periods and comparison groups may have varying identification power, leading them to treat all pre-treatment periods as equally informative or to use only the last pre-treatment period as a baseline. The semiparametric efficiency framework developed in this paper provides a principled approach to leveraging information from the pre-treatment period to form DiD and ES estimators with minimum variance under the maintained identification assumptions, while remaining agnostic to parametric functional form or parametric error structure. As our simulations and empirical illustration demonstrate, the gains in precision can be substantial, addressing the concern that heterogeneous robust DiD estimators may not be informative compared to more traditional estimators that assume treatment effect heterogeneity away.
In practice, practitioners may want to assess if there is any statistical evidence against using any pre-treatment period as the baseline in their DiD/ES procedure, or against using any specific available comparison group (when treatment is staggered). We present in Appendix (ref) how to construct Hausman-type tests for overidentification specification test of DiD models, or how to construct alternative estimators that aim to trade off bias and variance Armstrong_Kline_Sun_2024_adaptive. We also discuss visualization schemes to informally quantify the sensitivity of the results concerning the usage of different pre-treatment baseline periods, and different comparison groups, or to “eyeball” the magnitude of potential model violations.
Our framework opens avenues for several extensions. It would be useful to study semiparametric efficient estimation for nonlinear DiD-type setups such as the Changes-in-Changes model Athey2006, setups where treatment can turn on and off over multiple periods deChaisemartin2020_AER, deChaisemartin2024_intertemporal, and setups with unbalanced panel or repeated cross-sections Callaway_Santanna_2021---in this latter case, we expect new sources of overidentification related to event-study weights. Finally, it is important to stress that all results derived in this paper leveraged the large-$n$, fixed-$T$ panel data regime, as parallel trends are more likely to be satisfied when the number of periods is small. In setups where parallel trends are not plausible, we recommend that researchers use alternative estimators that do not rely on parallel trends assumptions, e.g., Arkhangelsky2021_SDiD, Viviano2021, and Imbens_Viviano_2023. Developing semiparametric efficient estimators in such frameworks is left for future research.