EconBase
← Back to paper

Cohort-Anchored Robust Inference for Event-Study with Staggered Adoption

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

136,139 characters · 19 sections · 67 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Cohort-Anchored Robust Inference for Event-Study with Staggered Adoption

\onehalfspacing

abstractThis paper proposes a cohort-anchored framework for robust inference in event studies with staggered adoption, building on rambachan2023more. Robust inference based on event-study coefficients aggregated across cohorts can be misleading due to the dynamic composition of treated cohorts, especially when pre-trends differ across cohorts. My approach avoids this problem by operating at the cohort-period level. To address the additional challenge posed by time-varying control groups in modern DiD estimators, I introduce the concept of block bias: the parallel-trends violation for a cohort relative to its fixed initial control group. I show that the biases of these estimators can be decomposed invertibly into block biases. Because block biases maintain a consistent comparison across pre- and post-treatment periods, researchers can impose transparent restrictions on them to conduct robust inference. In simulations and a reanalysis of minimum-wage effects on teen employment, my framework yields better-centered (and sometimes narrower) confidence sets than the aggregated approach when pre-trends vary across cohorts. The framework is most useful in settings with multiple cohorts, sufficient within-cohort precision, and substantial cross-cohort heterogeneity. \noindentKeywords: panel data, two-way fixed effects, parallel-trends, pre-trend, event-study plot, difference-in-differences, robust inference, confidence set

\thispagestyle{empty} \doublespace

\setcounter{page}{1} \abovedisplayskip=5pt \belowdisplayskip=5pt

Introduction

Difference-in-differences (DiD) methods are among the most widely used tools for causal inference with panel data in the social sciences. The credibility of these designs hinges on the parallel-trends (PT) assumption, which posits that treated and control groups would have evolved in parallel in the absence of treatment. Because the assumption is untestable, it has become standard practice for researchers to assess its plausibility by examining pre-treatment trends, either by eyeballing or testing whether pre-treatment event-study coefficients are jointly zero.

However, a growing literature has highlighted limitations of the pre-testing approach: such tests often have low power against meaningful violations of PT, and conditioning on passing a pre-test can introduce statistical distortions roth2022pretest. This motivates an inference procedure that does not require PT to hold exactly. The robust inference framework of rambachan2023more (hereafter \textcolor{NavyBlue}{RR (rambachan2023more)}) provides such a solution: rather than treating PT as a binary condition, it uses the observable pre-treatment series to impose restrictions on the magnitude of potential post-treatment PT violations (intuitively, post-treatment violations of PT cannot differ too much from those in pre-treatment periods). Under these restrictions, the average treatment effect is set-identified, and one can construct a confidence set for it.

While the \textcolor{NavyBlue}{RR (rambachan2023more)} framework is a crucial tool for robust inference, its reliance on event-study coefficients aggregated across cohorts creates complications under staggered adoption designs because of the dynamic composition of treated and control cohorts. I identify three distinct challenges for applying the robust inference framework in this setting.

The first challenge arises when applying robust inference to coefficients from a two-way fixed effects (TWFE) regression that includes relative-period dummies, which is not robust to heterogeneous treatment effects (HTE). The coefficient on a relative-period dummy can be contaminated by a weighted average of treatment effects across multiple cohorts and periods, with weights that may be negative sun2021-event. Because these coefficients are confounded by HTE, the pre-treatment coefficients do not validly measure violations of PT, undermining the rationale for robust inference. This issue can be addressed by a class of “HTE-robust” estimators de2020two,sun2021-event,callaway2021-did,gardner2022two,liu2024practical,borusyak2024revisiting,de2024difference, which estimate cohort-period level average treatment effects and then aggregate them by relative period across cohorts to obtain event-study coefficients.

Second, while these new estimators address the first challenge, applying the robust inference framework to these event-study coefficients aggregated across cohorts introduces a mechanical problem of dynamic treated composition. Under staggered adoption, the set of treated cohorts contributing to the identification of event-study coefficients changes across relative periods. As a result, aggregated pre-treatment and post-treatment coefficients are not directly comparable, since they are based on different treated cohort compositions. For example, coefficients in distant pre-treatment periods are identified exclusively from late-treated cohorts, whereas coefficients in distant post-treatment periods are identified exclusively from early-treated cohorts. Differences in event-study coefficients between adjacent periods may therefore reflect shifts in cohort composition rather than changes in PT violations or dynamic treatment effects. The problem is particularly severe when cohorts exhibit heterogeneous pre-trends, as aggregation can mask important cohort-level variation and produce a distorted pre-treatment benchmark that does not fit the cohort compositions in post-treatment periods.

Given the issue of dynamic treated composition when using event-study coefficients aggregated across cohorts, one might ask if robust inference can be conducted at the cohort level. However, a third challenge arises for some popular “HTE-robust” estimators: the problem of dynamic control group. These estimators use not-yet-treated cohorts as controls. Examples include the imputation estimator borusyak2024revisiting,liu2024practical---which, as I show in Section 2, employs not-yet-treated cohorts implicitly---and the callaway2021-did estimator with not-yet-treated controls (hereafter the CS-NYT estimator). For these estimators, the not-yet-treated comparison groups for earlier-treated cohorts shrink over time. Consequently, it is difficult to impose credible restrictions linking pre- and post-treatment violations of PT, as the very definition of the PT assumption changes as the control group evolves. This challenge does not arise for estimators that use a fixed, universal control group, such as sun2021-event and the callaway2021-did estimator that uses never-treated units as controls (hereafter the CS-NT estimator).

Figure (ref) uses a toy example to illustrate the latter two challenges. The left panel shows the treatment adoption status and relative periods for each observation, where $s$ denotes the relative period and $s=1$ is the first post-treatment period. The data includes two treated cohorts, $\mathcal{G}_{5}$ (treated at $t=5$) and $\mathcal{G}_{7}$ (treated at $t=7$), and a never-treated cohort, $\mathcal{G}_{\infty}$. I use the CS-NYT estimator to calculate cohort-period coefficients. The aggregated event-study coefficients are obtained by aggregating these cohort-period estimates by relative period, weighted by cohort sizes.

The center panel plots the aggregated event-study coefficients and illustrates the problem of dynamic treated composition. The composition of cohorts that identifies the aggregated coefficients changes over relative time. In distant pre-treatment periods ($s \le -4$), the coefficients are identified exclusively from the late-adopting cohort, $\mathcal{G}_7$. In periods $-3 \le s \le 2$, both cohorts contribute to the coefficients. In distant post-treatment periods ($s \ge 3$), the coefficients are identified exclusively from the early-adopting cohort, $\mathcal{G}_5$. This shifting composition makes it inconsistent to use pre-trends of one cohort composition to benchmark post-treatment PT violations for another composition. For example, the pre-trend in $s \le -4$ (identified from $\mathcal{G}_7$ alone) is not a valid benchmark for the potential violation of PT in post-treatment periods $s \ge 3$ (identified from $\mathcal{G}_5$ alone). Furthermore, this compositional dynamism means that the change in the aggregated coefficient between consecutive periods, such as from $s=-4$ to $s=-3$, is confounded by the entry of cohort $\mathcal{G}_5$ into the composition of treated cohorts, which makes such a change in coefficients an invalid benchmark for the period-to-period evolution of PT violations.

The right panel illustrates the dynamic control group issue for cohort $\mathcal{G}_5$ under the CS-NYT estimator. In its post-treatment periods, the control group for this cohort initially consists of $\mathcal{G}_7 \cup \mathcal{G}_\infty$ (for $s=1,2$), but shrinks to only $\mathcal{G}_\infty$ for later periods ($s=3,4$) after $\mathcal{G}_7$ itself becomes treated. Since its pre-treatment coefficients are based on a comparison between $\mathcal{G}_5$ and the larger, initial control group ($\mathcal{G}_7 \cup \mathcal{G}_\infty$), this pre-trend is not a suitable benchmark for potential PT violations in the later periods where the control group has changed.

figure[figure omitted — 745 chars of source]

To address these two challenges in robust inference under staggered adoption, I propose a cohort-anchored robust inference framework, building on the work of \textcolor{NavyBlue}{RR (rambachan2023more)}. The primary difference from \textcolor{NavyBlue}{RR (rambachan2023more)} is that my framework bases inference on cohort-period level coefficients rather than on aggregated event-study coefficients. To address the dynamic control group issue, I introduce the concept of the block bias. Each block bias is defined at the cohort-period level and compares a treated cohort in a given period to its initial control group---the units that are untreated when this cohort first adopts treatment---relative to a reference period (or set of reference periods). Conceptually, this comparison operates within a block-adoption structure formed by the cohort and its initial control group---hence the term “block bias.” Because a treated cohort's initial control group does not vary over time, its block biases are “anchored” to this control group and thus have a consistent interpretation in both pre- and post-treatment periods, capturing the PT violation of the treated cohort with respect to this fixed control group.

The block bias concept enables robust inference through a bias decomposition that links block biases to estimators' biases at the cohort-period level. Formally, I define the overall bias in any post-treatment cohort-period cell as the difference between the expected value of the estimated average treatment effect and the true average treatment effect. For both the imputation and CS-NYT estimators, this overall bias can be written as a linear combination of that cohort’s own block bias and the block biases of later-treated cohorts that have begun treatment by that period. Denoting the stacked cohort-period overall biases\footnote{In pre-treatment periods, overall biases are, for simplicity, set equal to the block biases in the same cohort-period cells. This does not affect the robust inference procedure.} by $\vec{\delta}$ and the stacked block biases by $\vec{\Delta}$, this decomposition implies an invertible linear mapping between them: $\vec{\delta} = \mathbf{W}\vec{\Delta}$. Appendix (ref) provides the explicit bias decomposition for the toy example discussed previously.

In each post-treatment cohort-period cell, what can be estimated is the sum of the true treatment effect and the overall bias; neither component is separately identified. The rationale for robust inference is to impose a restriction set on the possible values of the post-treatment overall bias and, under this restriction, set-identify the true effect. This restriction set bounds the post-treatment overall bias using a benchmark learned from pre-treatment periods. Ideally, this benchmark should have the same interpretation as the overall bias in post-treatment periods. However, for the imputation estimator and the CS-NYT estimator, such benchmarks for the overall bias do not exist because the control group underlying the definition of overall bias can vary over time as discussed above.

Unlike the overall bias, block biases are defined consistently in pre- and post-treatment periods; consequently, for each cohort, observable pre-treatment block biases provide valid benchmarks for their unobservable post-treatment counterparts. The cohort-anchored robust inference framework therefore imposes restrictions on cohort-period level block biases, requiring that for each cohort, the potential block bias in post-treatment periods should not differ dramatically from the observed block bias in its pre-treatment periods. Consistent with \textcolor{NavyBlue}{RR (rambachan2023more)}, these restriction sets on block biases are specified to take the form of a single polyhedron or a union of polyhedra. Using the invertible linear mapping $\vec{\delta}=\mathbf{W}\vec{\Delta}$, the framework then translates the restriction set on the block biases into a corresponding restriction set on the overall biases. Finally, it feeds these translated restriction sets into the robust inference algorithm in \textcolor{NavyBlue}{RR (rambachan2023more)}, such as the hybrid method developed by andrews2023inference, to construct a confidence set for the average treatment effects.

The restriction sets from \textcolor{NavyBlue}{RR (rambachan2023more)} carry over seamlessly to the cohort-anchored framework, where they become more transparent and interpretable when applied to block biases rather than to aggregated estimates. This paper focuses on two such restrictions: Relative Magnitudes (RM) and Second Differences (SD). The RM restriction bounds post-treatment changes in a cohort's block bias between consecutive periods by a benchmark learned from pre-treatment trends. I consider both a computationally simple global benchmark, based on the maximum pre-treatment change across all cohorts, and a more theoretically sound (but computationally intensive) cohort-specific benchmark based on each cohort's own pre-trends. The SD restriction, which is well-suited for linear pre-trends, bounds the change in the slope of the block bias path and is computationally simple as it forms a single polyhedron.

I use two simulated examples with heterogeneous pre-trends to compare my cohort-anchored framework to the robust inference framework that relies on aggregated event-study coefficients (hereafter the aggregated framework). First, in a simulation illustrating the RM restriction where one cohort has an oscillating pre-trend, the cohort-anchored framework yields markedly narrower confidence sets. This occurs because the aggregated framework adopts an overly conservative benchmark dictated by the cohort with the substantial pre-trend. Second, in a simulation illustrating the SD restriction where one cohort has a linear pre-trend, the cohort-anchored framework produces better-centered confidence sets. The aggregated framework, by contrast, uses an averaged pre-treatment slope ill-suited for any individual cohort, leading to sets centered away from the true effect. In both scenarios, the cohort-anchored approach provides more credible inference by properly accounting for cohort-level heterogeneity.

I revisit the empirical application from callaway2021-did on the effect of minimum-wage increases on teen employment, a setting with heterogeneous cohort-specific linear pre-trends that can mislead the robust inference based on aggregated coefficients. The advantages of the cohort-anchored framework are starkest under the SD restriction. While the aggregated framework produces a confidence set centered at zero, the cohort-anchored framework's set remains centered well below zero. This demonstrates that the negative employment effect is robust to cohort-specific linear trend violations, a conclusion obscured by the aggregated framework.

Compared with robust inference based on aggregated event–study coefficients \`a la \textcolor{NavyBlue}{RR (rambachan2023more)}, the proposed cohort-anchored framework resolves the problems of dynamic treated composition and dynamic control group in staggered adoption settings, imposes more transparent and interpretable restriction sets, and yields better-centered confidence sets. The approach is particularly useful when there are multiple cohorts, each large enough to deliver reasonably precise cohort–period estimates, and when violations of PT differ meaningfully across cohorts.

These gains come with trade-offs. Computation can be heavier when the restriction set is a union of polyhedra because inference must be carried out across multiple benchmark configurations. Second, cohort–period estimates typically exhibit higher statistical uncertainty than aggregated coefficients, which can widen confidence sets. Finally, aggregation tends to smooth volatility in the pre-treatment series; under a fixed sensitivity parameter, the pre-treatment benchmark used by the cohort-anchored framework can sometimes be larger in magnitude than the benchmark obtained from aggregated estimates, increasing identification uncertainty and further widening confidence sets. I do not view these trade-offs as “cons” of the method; they are the price of conducting robust inference in a more transparent and interpretable way. However, in practice, although the proposed framework has many advantages, researchers may face settings with many cohorts of small size—so that cohort–period coefficients have high statistical uncertainty—in which case robust inference based on aggregated coefficients may still be preferred.

The contribution of this paper is twofold. In addition to the proposed cohort-anchored robust inference framework, the bias decomposition itself is of independent interest. While a variant of this property is first documented in aguilar2025comparative, I derive it for the imputation estimator from a different perspective. My derivation leverages the equivalence between the imputation estimator and a sequential imputation procedure, as shown in arkhangelsky2024sequential. Furthermore, the pre-treatment block bias provides a way to visualize pre-treatment coefficients for the imputation estimator that are directly comparable to their post-treatment counterparts, addressing recent discussions about the asymmetric construction and interpretation of pre- versus post-treatment coefficients in event studies Roth2024interpret,li2025benchmarking.

The remainder of the paper proceeds as follows. Section (ref) derives the bias decompositions for the imputation estimator and the CS-NYT estimator. Section (ref) then formally lays out my robust inference framework and details how restrictions are placed on block biases. In Section (ref), I use two simulated examples to illustrate the practical implementation of my framework. Section (ref) applies the framework to the empirical application of callaway2021-did. Section (ref) concludes. All proofs appear in the Appendix. Table (ref) lists all notation.

Bias Decomposition

This section develops the bias decomposition for the imputation and CS-NYT estimators, two widely used and representative HTE-robust methods. I begin by reviewing these estimators, then formally define the concepts of block bias and overall bias, and show that the overall bias in any post-treatment cohort–period cell can be expressed as a linear combination of block biases. Next, I provide intuition for this decomposition, and finally, I discuss the estimation of block bias in pre-treatment periods and its implications for interpreting pre- and post-treatment coefficients in event studies.

Setup

I consider a DiD design with staggered treatment adoption. I observe a panel of $N$ units over $T$ time periods. Units are classified into $G$ treated cohorts, $\mathcal{G}_g$ for $g \in \{1, \ldots, G\}$, and a never-treated cohort $\mathcal{G}_\infty$. The treated cohort $\mathcal{G}_g$ starts to be treated in period $t_g$, and the number of units in each cohort is denoted by $N_g$ (and $N_\infty$ for the never-treated group). Without loss of generality, assume $1<t_1<\cdots<t_G\leq T$. The treatment status for a unit $i \in \mathcal{G}_g$ is thus $D_{it} = \mathbf{1}\{t \ge t_g\}$, while for never-treated units, $D_{it} = 0$ for all $t$. The observed outcome is $Y_{it} = D_{it}Y_{it}(1) + (1-D_{it})Y_{it}(0)$, where $Y_{it}(1)$ and $Y_{it}(0)$ are the potential outcomes under treatment and control, respectively. I use $t \in \{1, \ldots, T\}$ to denote the calendar period and $s$ to denote the period relative to the treatment date. For a given cohort $\mathcal{G}_{g}$, $s=1$ corresponds to the first post-treatment period ($t=t_g$), while $s=0$ corresponds to the last pre-treatment period ($t=t_g-1$).

Throughout this paper, I assume a balanced panel in a staggered adoption setting without covariates. My analysis therefore focuses on the unconditional, rather than the conditional, PT assumption. I also assume the existence of a never-treated cohort.

assumption[Model Setup] \begin{enumerate}[label=\arabic*., itemsep=-1pt, topsep=0pt, partopsep=0pt] • • Staggered Adoption: $1 < t_1 < \cdots < t_G \leq T$. • Never-Treated Cohort: The set of never-treated units is non-empty, i.e., $|\mathcal{G}_\infty| = N_\infty > 0$. • Balanced Panel: The data matrix $\{Y_{it}\}$ is complete for all units $i \in \bigcup_{g=1}^G \mathcal{G}_g \cup \mathcal{G}_\infty$ and all time periods $t \in \{1, \dots, T\}$. \end{enumerate}

The potential outcome notation, $Y_{it}(0)$, is a simplification that implicitly bundles some identifying assumptions as one unit's potential outcome could depend on the entire vector of treatment assignments across all units and time. I therefore formally separate the assumptions required to simplify this general case. They can be understood as two facets of the no interference principle: one that restricts interference across units (SUTVA) and another that restricts interference across time (no anticipation).

assumption[No Interference] \begin{enumerate}[label=\arabic*., itemsep=-1pt, topsep=0pt, partopsep=0pt] • • Across Units (SUTVA): The potential outcomes for any unit $i$ do not depend on the treatment status of any other unit $j \neq i$. This rules out spillover effects across units. • Across Time (No Anticipation): The potential outcomes for any unit $i$ in period $t$ do not depend on its treatment status in future periods. \end{enumerate}

\paragraph*{Imputation Estimator}

The imputation estimator, proposed by borusyak2024revisiting and liu2024practical,\footnote{Its numerically equivalent forms are documented in wooldridge2021two, gardner2022two and gardnerone.} is a popular approach for DiD designs. The procedure begins by estimating unit and time fixed effects, $(\hat{\alpha}_i, \hat{\xi}_t)$, from a two-way fixed effects model fitted using only the untreated observations in the panel. These estimates are then used to impute the counterfactual outcome for each treated observation: $\hat{Y}_{i,t_g+s-1}(0) = \hat{\alpha}_i + \hat{\xi}_{t_g+s-1}$. The estimated average treatment effect for cohort $\mathcal{G}_g$ at relative period $s$ is the average difference between the observed and imputed outcomes: $$ \hat{\tau}_{\mathcal{G}_g,s}^{\text{Imp}} = \frac{1}{N_g}\sum_{i\in\mathcal{G}_g} \left[ Y_{i,t_g+s-1} - \hat{Y}_{i,t_g+s-1}(0) \right]. $$ This estimator has several attractive properties that make it a popular alternative to traditional TWFE regression. Most importantly, because the fixed effects are estimated using only untreated observations, the resulting treatment effect estimates are robust to HTE and avoid the negative weighting problems that can bias standard TWFE estimators in staggered designs. Furthermore, borusyak2024revisiting show that the imputation estimator is the most efficient linear unbiased estimator under the assumption of homoskedastic and serially uncorrelated errors (spherical errors).

\paragraph*{The CS-NYT Estimator}

callaway2021-did provides alternative estimators for $\tau_{\mathcal{G}_g,s}$. One of them compares the change in the average outcome for a treated cohort $\mathcal{G}_g$ to the corresponding change for a control group consisting of not-yet-treated units, denoted $\mathcal{C}_{g,s}=\left(\cup_{k: t_k>t} \mathcal{G}_k\right) \cup \mathcal{G}_{\infty}$, where $t=t_g+s-1$. A key feature of this method is that the control group $\mathcal{C}_{g,s}$ is defined contemporaneously and therefore shrinks as later cohorts receive treatment. The initial control group for cohort $\mathcal{G}_g$ is the set of not-yet-treated units at its first period of treatment ($s=1$), which I denote as $\mathcal{C}_{g,1}$.

The estimator compares the change between the last pre-treatment period, $t_g-1$, and a given post-treatment period, $t_g+s-1$. The estimated average treatment effect for cohort $\mathcal{G}_g$ at relative period $s$ is $$\hat{\tau}_{\mathcal{G}_g,s}^{\text{CS-NYT}} = \frac{1}{N_g}\sum_{i\in\mathcal{G}_g}\bigl(Y_{i,t_g+s-1}-Y_{i,t_g-1}\bigr) - \frac{1}{N_{\mathcal{C}_{g,s}}}\sum_{i\in\mathcal{C}_{g,s}}\bigl(Y_{i,t_g+s-1}-Y_{i,t_g-1}\bigr),$$ where the time-varying control group $\mathcal{C}_{g,s}$ is defined on each relative post-treatment period $s\geq 1$. callaway2021-did also discuss extensions that incorporate covariates using outcome regression, inverse probability weighting, and doubly robust techniques.

\paragraph*{Other HTE-Robust Estimators} The landscape of HTE-robust DiD estimators also includes several other important methods. For instance, the CS-NT estimator in callaway2021-did that uses the never treated as controls and the regression-based estimator of sun2021-event both identify treatment effects using a fixed, never-treated control group. Another key approach, the $\text{DID}_l$ estimator of de2024difference, is closely related to the CS-NYT estimator I focus on.

In the setting of a balanced panel with staggered adoption and no covariates, the estimator of sun2021-event is equivalent to the CS-NT estimator. My framework’s core theoretical results can therefore be readily applied to this wider class of estimators. For methods that use a uniform control group, the bias decomposition becomes trivial, as adjustment for the varying control group is no longer needed.

Overall bias and Block Bias

I now formally define the overall bias ($\delta_{\mathcal{G}_g,s}$) and the block bias ($\Delta_{\mathcal{G}_g,s}$). Both concepts are defined on each cohort-period cell $(g,s)$ and form the foundation of the bias decomposition and robust inference framework that follows. \paragraph*{Overall bias} For the imputation estimator, I define the overall bias for cohort $\mathcal{G}_g$ in a post-treatment period $s \geq 1$, denoted $\delta^{\text{Imp}}_{\mathcal{G}_g,s}$, as the difference between the expectation of the estimated treatment effect and the true treatment effect in that cohort-period cell. As the derivation below shows, this simplifies to the difference between the true and the imputed counterfactuals:

equation[equation omitted — 631 chars of source]

The overall bias for the CS-NYT estimator can be defined analogously in post-treatment periods:

equation[equation omitted — 915 chars of source]

Note that $Y_{i,t_g+s-1}(0)=Y_{i,t_g+s-1}$ for $i \in \mathcal{C}_{g,s}$ and $Y_{i,t_g-1}(0)=Y_{i,t_g-1}$ for $i \in \mathcal{C}_{g,s}\bigcup\mathcal{G}_g$. The bias term $\delta^{\text{CS-NYT}}_{\mathcal{G}_g,s}$ therefore represents the difference between two trends from the reference period $t_g-1$ to the post-treatment period $t_g+s-1$: (1) the counterfactual trend for cohort $\mathcal{G}_g$ under no treatment, and (2) the observed trend for the not-yet-treated control group, $\mathcal{C}_{g,s}$. The condition that this bias term is zero, $\delta^{\text{CS-NYT}}_{\mathcal{G}_g,s}=0$, is precisely the PT assumption required for the unbiased estimation of $\tau_{\mathcal{G}_{g},s}$.

The overall bias terms, $\delta^{\text{Imp}}_{\mathcal{G}_g,s}$ and $\delta^{\text{CS-NYT}}_{\mathcal{G}_g,s}$, represent the bias in the estimated treatment effect for a given post-treatment cohort-period cell $(g,s)$. The goal of robust inference is to use observable pre-trends to form a restriction set on these unobservable post-treatment biases. For the restriction to be interpretable and credible, the benchmark derived from pre-trends should have the same interpretation as the post-treatment biases. However, the overall bias is ill-suited for this task. Specifically, one cannot find a measure of the pre-trend that best mirrors the overall bias in post-treatment periods. For the imputation estimator, the overall bias lacks a symmetrically defined pre-treatment analog. For the CS-NYT estimator, the control group's composition changes in post-treatment periods, meaning no single pre-treatment comparison can serve as a consistent benchmark for all post-treatment periods. These challenges motivate my introduction of the block bias.

\paragraph*{Block Bias}

The block bias, $\Delta_{\mathcal{G}_g,s}$, captures the violation of PT for a cohort $\mathcal{G}_g$ relative to its fixed, initial control group, $\mathcal{C}_{g,1}= \left(\bigcup_{k: t_k > t_g} \mathcal{G}_k\right) \cup \mathcal{G}_\infty$. This group consists of all units not-yet-treated at the moment cohort $\mathcal{G}_g$ first receives treatment ($s=1$). Because this comparison group is, by definition, fixed for all relative periods $s$, the block bias provides a consistent measure of the trend difference between $\mathcal{G}_g$ and $\mathcal{C}_{g,1}$ in both pre- and post-treatment periods. I use the term “block” because this comparison operates within a classic block-adoption structure formed by $\mathcal{G}_g$ and its initial control group $\mathcal{C}_{g,1}$.

The exact definition of the block bias depends on the estimator. For the imputation estimator, the block bias is defined as:

equation[equation omitted — 415 chars of source]

where $\text{pre}_g = \{1,\dots,t_g-1\}$ denotes the set of pre-treatment periods for cohort $\mathcal{G}_g$, and $\bar{Y}_{i,\text{pre}_g}(0)=\frac{1}{t_g-1}\sum_{t \in \text{pre}_g}Y_{it}(0)$ is the average potential outcome for unit $i$ over those periods. By definition, no units in cohort $\mathcal{G}_g$ or its initial control group $\mathcal{C}_{g,1}$ are treated during the $\text{pre}_g$ window. Therefore, their potential outcomes are equal to their observed outcomes during periods $\text{pre}_g$ (i.e., $Y_{i,t}(0) = Y_{i,t}$ for $i \in \mathcal{G}_g \bigcup \mathcal{C}_{g,1}$ and $t \in \text{pre}_g$).

The term $\Delta_{\mathcal{G}_g,s}^{\text{Imp}}$ takes a classic DiD-like structure. The first component, $\mathbb{E}[Y_{i,t_g+s-1}(0) \mid i\in\mathcal{G}_g] - \mathbb{E}[\bar{Y}_{i,\text{pre}_g}(0) \mid i\in\mathcal{G}_g]$ represents the trend in cohort $\mathcal{G}_g$'s potential outcome, measuring its value in a given period relative to the average over all its pre-treatment periods. The block bias, $\Delta_{\mathcal{G}_g,s}^{\text{Imp}}$, is the difference between this trend for cohort $\mathcal{G}_g$ and the corresponding trend for its initial control group, $\mathcal{C}_{g,1}$. A key feature of this definition is that the block bias is directly observable in all pre-treatment periods ($s \le 0$), while in post-treatment periods ($s > 0$) it becomes unobservable. This unobservability arises because not only the potential outcomes for the treated cohort $\mathcal{G}_g$ are unknown by definition, but also the potential outcomes for some units in the initial control group $\mathcal{C}_{g,1}$ may become unobservable once they themselves are treated at later dates.

Crucially, the definition and interpretation of the block bias remain identical across all pre- and post-treatment periods. It always compares the trend of the treated cohort with the trend of its anchored initial control group. This consistency allows observable pre-treatment block biases to serve as valid benchmarks for unobservable post-treatment block biases.

The block bias for the CS-NYT estimator takes a similar form, but uses the last pre-treatment period of the target cohort, $t_g-1$, as the reference period. Formally, it is defined as:

equation[equation omitted — 403 chars of source]

This definition closely resembles that of the overall bias in Equation (ref). The critical difference is that the block bias compares cohort $\mathcal{G}_g$ to its fixed initial control group, $\mathcal{C}_{g,1}$, whereas the overall bias uses the time-varying not-yet-treated control group, $\mathcal{C}_{g,s}$. The same as the imputation estimator, this definition allows the observable pre-treatment estimates of $\Delta_{\mathcal{G}_g,s}^{\text{CS-NYT}}$ to serve as a valid benchmark for its unobservable post-treatment counterparts.

For any treated cohort, the overall bias and block bias are equivalent in the first post-treatment period ($s=1$), $\delta_{\mathcal{G}_g,1} = \Delta_{\mathcal{G}_g,1}$. For the CS-NYT estimator, this result is straightforward. At the first post-treatment period, the not-yet-treated control group is, by definition, identical to the initial control group. For the imputation estimator, this equality follows directly from the lemma below.

lemmaFor any treated unit $i \in \mathcal{G}_g$ in its first post-treatment period $t_g$, the counterfactual imputed by the imputation estimator, $\hat{Y}^{\text{Imp}}_{i,t_g}(0)$, is algebraically equivalent to a simple block-DiD expression: \[ \hat{Y}_{i,t_g}^{\text{Imp}}(0) = \overline{Y}_{i,\text{pre}_g} + \bigl(\overline{Y}_{\mathcal{C}_{g,1},t_g} - \overline{Y}_{\mathcal{C}_{g,1},\text{pre}_g}\bigr), \] where $\mathcal{C}_{g,1}$ is the initial control group, $\overline{Y}_{\mathcal{C}_{g,1},t_g}$ is the average outcome for this group at time $t_g$, and $\overline{Y}_{\mathcal{C}_{g,1},\text{pre}_g}$ is this group's average outcome over cohort $\mathcal{G}_g$'s pre-treatment periods.

\noindentProof of Lemma (ref) is provided in the Appendix (ref).

This lemma establishes an exact algebraic identity for the imputed counterfactual in the first post-treatment period. Taking the expectation of the expression in Lemma (ref) implies that the overall bias, $\delta^{\text{Imp}}_{\mathcal{G}_g,1}$, is equivalent to the block bias, $\Delta^{\text{Imp}}_{\mathcal{G}_g,1}$, for every treated cohort. This equivalence provides a crucial link between the full, staggered panel and the simpler block structure consisting of cohort $\mathcal{G}_g$ and its initial control group, furnishing the first step for the bias decomposition that follows.

The Bias Decomposition

For relative periods beyond the first post-treatment period ($s>1$), the overall bias is no longer identical to the block bias. For both the imputation and the CS-NYT estimators, the control group for an early-adopting cohort changes over time as some cohorts in the initial control group receive treatment later. This dynamic produces a recursive structure where the overall bias of one cohort depends on the block biases of those treated after it.

The following proposition formalizes this relationship for the imputation estimator, decomposing the overall bias for any cohort-period cell into its own block bias plus a linear combination of the block biases of later-adopting cohorts. This result is a variant of the Corollary 1 in aguilar2025comparative, though my formulation and proof—which leverages the sequential nature of the imputation estimator—are distinct.

proposition[Bias Decomposition for the Imputation Estimator] Let $t = t_g + s - 1$ be the calendar time corresponding to post-treatment period $s\geq 1$ for cohort $\mathcal{G}_g$. Then the overall bias $\delta^{\text{Imp}}_{\mathcal{G}_g,s}$ satisfies\footnote{The cohort sizes $N_k$ are treated as fixed features of the design, and the expectations are conditional on this group structure. An alternative approach would be to define the weights using population shares, for which the sample proportions $N_k/N$ are consistent estimators.} \[ \delta^{\text{Imp}}_{\mathcal{G}_g,s} = \Delta^{\text{Imp}}_{\mathcal{G}_g,s} + \sum_{k \in \mathcal{K}_g(t)} \left( \frac{N_k}{\sum_{j=k}^{G} N_j + N_\infty} \right) \Delta^{\text{Imp}}_{\mathcal{G}_k,s_k(t)} \] where the set of later-treated cohorts that contributes to $\delta^{\text{Imp}}_{\mathcal{G}_g,s}$ is $\mathcal{K}_g(t) = \bigl\{k \,\big|\, t_g < t_k \le t\bigr\}$ and the relative period for such a cohort $\mathcal{G}_k$ is $s_k(t) = t - (t_k - 1)$.

\noindentProof. See Appendix (ref).

Proposition (ref) shows that the overall bias from the imputation estimator for cohort $\mathcal{G}_g$ at a post-treatment period $s$, $\delta^{\text{Imp}}_{\mathcal{G}_g,s}$, equals its own block bias, $\Delta^{\text{Imp}}_{\mathcal{G}_g,s}$, plus a weighted sum of block biases from a specific set of later-adopting cohorts. These cohorts, which I term adjustment cohorts, are those that begin treatment after $\mathcal{G}_g$ and no later than calendar time $t=t_g+s-1$, denoted by the set $\mathcal{K}_g(t) = \bigl\{k \,\big|\, t_g < t_k \le t\bigr\}$. The weight assigned to each adjustment cohort $\mathcal{G}_k$ depends on its size, $N_k$, relative to the total size of $\mathcal{G}_k$ and $\mathcal{G}_k$'s initial control group, $\mathcal{C}_{k,1}$, which is $\frac{N_k}{N_k+N_{\mathcal{C}_{k,1}}}=\frac{N_k}{\sum_{j=k}^{G} N_j + N_\infty}$. Finally, the relative period for each adjustment cohort, $s_k(t) = t - t_k+1$, ensures its block bias is evaluated at the same calendar time $t$ as the overall bias it adjusts.

figure[figure omitted — 744 chars of source]

To illustrate the bias-decomposition formula from Proposition (ref), Figure (ref) presents an illustrative panel with three treated cohorts and one never-treated group. Cohort $\mathcal{G}_4$ (unit 1) is treated at $t=4$, Cohort $\mathcal{G}_6$ (units 2–4) at $t=6$, and Cohort $\mathcal{G}_8$ (unit 5) at $t=8$. Each dark-blue cell denotes a treated observation, and the expression inside displays how the overall bias, $\delta^{\text{Imp}}$, for that observation's cohort, is decomposed into its constituent block biases, $\Delta^{\text{Imp}}$.

In each cohort’s first post-treatment period, as Lemma (ref) suggests, the overall bias equals the cohort’s own block bias. In later periods, the overall bias for early adopters becomes the sum of their own block bias and the block biases from their adjustment cohorts. For instance, at $t=6$, Cohort $\mathcal{G}_6$ begins to receive treatment, and the overall bias for Cohort $\mathcal{G}_4$ becomes $\delta^{\text{Imp}}_{\mathcal{G}_4,s=3} = \Delta^{\text{Imp}}_{\mathcal{G}_4,s=3} + 0.375\,\Delta^{\text{Imp}}_{\mathcal{G}_6,s=1}$. The weight of 0.375 is determined by Cohort $\mathcal{G}_6$’s size ($N_{6}=3$) relative to the total size of Cohort $\mathcal{G}_6$ and its initial control group, calculated as $0.375 = \frac{N_{6}}{N_{6} + N_{8} + N_\infty} = \frac{3}{3 + 1 + 4}$. Similarly, when $\mathcal{G}_8$ is treated at $t=8$, its block bias, $\Delta^{\text{Imp}}_{\mathcal{G}_8,s=1}$, contributes to the overall bias of $\mathcal{G}_4$ ($\delta^{\text{Imp}}_{\mathcal{G}_4,s=5}$) and $\mathcal{G}_6$ ($\delta^{\text{Imp}}_{\mathcal{G}_6,s=3}$) with a weight of $\frac{N_8}{N_8 + N_\infty} = \frac{1}{1 + 4} = 0.2$.

The bias decomposition for the CS-NYT estimator takes a slightly different form, as stated in Proposition (ref). This result is also a variant to the Corollary 3 in aguilar2025comparative, albeit with a different formulation and proof.

proposition[Bias Decomposition for the CS-NYT Estimator] Let $t = t_g + s - 1$ be the calendar time corresponding to post-treatment period $s\geq 1$ for cohort $\mathcal{G}_g$. Then the overall bias $\delta^{\text{CS-NYT}}_{\mathcal{G}_g,s}$ satisfies \[ \delta^{\text{CS-NYT}}_{\mathcal{G}_g,s} = \Delta^{\text{CS-NYT}}_{\mathcal{G}_g,s} + \sum_{k \in \mathcal{K}_g(t)} \left( \frac{N_k}{\sum_{j=k}^{G} N_j + N_\infty} \right) (\Delta^{\text{CS-NYT}}_{\mathcal{G}_k,s_k(t)}-\Delta^{\text{CS-NYT}}_{\mathcal{G}_k,s_k(t_g-1)}) \] where the set of adjustment cohorts is $\mathcal{K}_g(t) = \{k \mid t_g < t_k \le t\}$, and the relative periods for cohort $\mathcal{G}_k$ are $s_k(t) = t - t_k + 1$ and $s_k(t_g-1) = (t_g -1) - (t_k - 1)=t_g-t_k$.

\noindentProof. See Appendix (ref).

While the definitions of the bias terms ($\delta^{\text{CS-NYT}}$ and $\Delta^{\text{CS-NYT}}$) differ from the imputation estimator, the weights on the adjustment terms, $\frac{N_k}{\sum_{j=k}^{G} N_j + N_\infty}$, are identical. The distinction lies in the structure of the adjustment term. For the imputation estimator, the adjustment is a single block bias, $\Delta^{\text{Imp}}_{\mathcal{G}_k,s_k}$. For the CS-NYT estimator, the adjustment term from $\mathcal{G}_k$ is the difference between two of its block biases: $\Delta^{\text{CS-NYT}}_{\mathcal{G}_k,s_k(t)} - \Delta^{\text{CS-NYT}}_{\mathcal{G}_k,s_k(t_g-1)}$. This structure arises because the block biases of $\mathcal{G}_g$ and $\mathcal{G}_k$ are measured with respect to different reference periods. The block bias for cohort $\mathcal{G}_g$ is defined relative to its reference period ($t_g-1$), while the block bias for a later cohort $\mathcal{G}_k$ is defined relative to $\mathcal{G}_k$'s reference period ($t_k-1$). To make them comparable, the adjustment term for cohort $\mathcal{G}_k$ must be re-calibrated to cohort $\mathcal{G}_g$'s baseline. The term $\Delta^{\text{CS-NYT}}_{\mathcal{G}_k,s_k(t_g-1)}$ serves as this baseline correction. Since $s_k(t_g-1)=t_g-t_k<0$, this correction term is simply a pre-treatment block bias of $\mathcal{G}_k$. Subtracting this term isolates the component of the block bias from $\mathcal{G}_k$ that is properly aligned with $\mathcal{G}_g$'s block bias.

A key component of the bias decompositions in Propositions (ref) and (ref) is the weights placed on the adjustment cohort, $\mathcal{K}_g(t)$. The form, $w_k=\frac{N_k}{\sum_{j=k}^G N_j+N_{\infty}}$, implies that in $\mathcal{K}_g(t)$, (i) cohorts with larger population sizes $N_k$ receive larger weights; and (ii) with fixed cohort sizes, later-treated cohorts (larger $k$) in $\mathcal{K}_g(t)$ receive larger weights, since $\sum_{j=k}^G N_j$ decreases in $k$. The identical weights in both propositions highlight a similarity between the two estimators. Although they take different forms and have different implementations, both the imputation estimator and the CS-NYT estimator compare the trend of a treated cohort to that of its not-yet-treated control group. The primary distinction lies in how they use pre-treatment information: the imputation estimator uses the average of outcomes across all pre-treatment periods as the reference, while the CS-NYT estimator uses only the last pre-treatment period as the reference period. chen2025efficient provides a detailed discussion on the differences between these two estimators from this perspective.

Both propositions show that the overall bias of any post-treatment cohort-period cell can be expressed as a linear combination of block biases. For notational simplicity, I extend this relationship across all periods by defining the overall bias in pre-treatment periods to be equal to the block bias in the same periods. With this convention in place, the entire system of biases can be written compactly in matrix form. The stacked vectors of overall biases ($\vec{\delta}$) and block biases ($\vec{\Delta}$) are then related by the invertible linear mapping: $\vec{\delta} = \mathbf{W}\vec{\Delta}$.

Order the cohort–period cells $(g,s)$ by increasing calendar time $t=t_g+s-1$ and, within each $t$, by increasing adoption time $t_g$. Under this stacking, the mapping matrix $\mathbf W$ is block diagonal for the imputation estimator and block triangular for the CS-NYT estimator across calendar times: \[ \mathbf W =

cases\mathrm{diag}(A_1,\dots,A_T), & (Imputation)\\[6pt] \begin{bmatrix} A_1 & 0 & \cdots & 0\\ * & A_2 & \ddots & \vdots\\ \vdots & \ddots & \ddots & 0\\ * & \cdots & * & A_T \end{bmatrix}, & (CS-NYT),

\] where for each $t$ the diagonal block $A_t$ is unit upper–triangular. Hence $\mathrm{det}(A_t)=1$ for all $t$, and by the block–triangular determinant identity, $\mathrm{det}(\mathbf W)=\prod_{t=1}^T \mathrm{det}(A_t)=1\neq 0$. Therefore $\mathbf W$ is square and invertible (with $\mathbf W^{-1}$ block–diagonal in the imputation case and block lower–triangular in the CS-NYT case). Appendix (ref) provides the explicit form of the $\mathbf W$ matrix for the CS-NYT estimator as applied to the toy example in the introduction.

Intuition of the Bias Decomposition

The intuition behind the bias decomposition for the imputation estimator becomes clear by conceptualizing it as a sequential procedure that constructs the counterfactual panel one relative period at a time. The equivalence between the imputation estimator and sequential procedures is studied by arkhangelsky2024sequential that considers a more general form with unit- or period-specific weights. My derivation of the bias decomposition follows this sequential approach, providing an alternative perspective to the direct matrix algebra used in aguilar2025comparative.

figure[figure omitted — 994 chars of source]

Figure (ref) illustrates this sequential procedure with the example in Figure (ref). Imagine starting with a data matrix containing only the observed outcomes of untreated cells (light blue). The procedure begins by imputing all cells corresponding to the first post-treatment period, $s=1$ (highlighted in red in the left panel). For each such cohort-period cell, the counterfactual is imputed using a block structure that consists of the cohort and its initial controls, utilizing data from the current and all prior periods (marked by the green rectangle). For example, the counterfactual for unit 1 (cohort $\mathcal{G}_4$) at $t=4$ is constructed as $\hat{Y}_{1,4}(0) = \bar{Y}_{1,\text{pre}_4}(0) + [\bar{Y}_{\mathcal{C}_{4,1},4}(0) - \bar{Y}_{\mathcal{C}_{4,1},\text{pre}_4}(0)]$, where $\text{pre}_4 = \{1,2,3\}$ and the initial control group $\mathcal{C}_{4,1}$ consists of units 2 through 9. After this step, the imputed values for all $s=1$ cells fill their corresponding entries in the matrix.

The procedure continues by imputing cells for $s=2$ (middle panel), then $s=3$ (right panel), and so on. The key feature of these later rounds is that the construction of counterfactuals may rely on previously imputed values. For instance, to impute the counterfactual for unit 1 at $t=6$ (in the $s=3$ round), the procedure uses the formula: $\hat{Y}_{1,6}(0)=\bar{Y}_{1,\text{pre}_6}(0)+[\bar{Y}_{\mathcal{C}_{4,1},6}(0) - \bar{Y}_{\mathcal{C}_{4,1},\text{pre}_6}(0)]$ where $\text{pre}_6 = \{1,2,3,4,5\}$. However, units 2, 3, and 4 are treated at this time, the calculation of $\bar{Y}_{\mathcal{C}_{4,1},6}(0)$ thus uses their imputed counterfactuals---$\hat{Y}_{2,6}(0), \hat{Y}_{3,6}(0),$ and $\hat{Y}_{4,6}(0)$---which were already imputed during the $s=1$ round. Similarly, the calculation of $\bar{Y}_{1,\text{pre}_6}(0)$ also uses values for unit 1 that were imputed during the earlier $s=1$ and $s=2$ rounds.

This iterative process continues over all cohort-period cells from $s=1$ to the largest post-treatment relative periods. I prove in Appendix (ref) that this sequential procedure yields results that are numerically equivalent to the standard imputation estimator that solves for and uses all fixed effects simultaneously.

The equivalence between the imputation estimator and the sequential procedure implies that every cohort-period level imputed counterfactual can be expressed as a DiD-style comparison within the block structure formed by the cohort and its initial controls. Specifically, for any post-treatment cell $(g,s)$ at calendar time $t=t_g+s-1$, the average imputed counterfactual can be constructed with: $$\overline{\hat{Y}}_{\mathcal{G}_g, t}(0) = \bar{Y}^*_{\mathcal{G}_g, \text{pre}_t}(0) + [\bar{Y}^*_{\mathcal{C}_{g,1}, t}(0) - \bar{Y}^*_{\mathcal{C}_{g,1}, \text{pre}_t}(0)]$$ where the terms on the right-hand side are: (i) $\bar{Y}^*_{\mathcal{G}_g, \text{pre}_t}(0)$ is the average untreated outcome for cohort $\mathcal{G}_g$ over all prior periods; (ii) $\bar{Y}^*_{\mathcal{C}_{g,1}, t}(0)$ is the average untreated outcome for its initial control group at the current time $t$; and (iii) $\bar{Y}^*_{\mathcal{C}_{g,1}, \text{pre}_t}(0)$ is the average untreated outcome for the initial control group over all prior periods. The star ($^*$) signifies that these averages are taken over observed untreated outcomes or, if unavailable, over counterfactuals imputed in previous imputation rounds.

The overall bias for a post-treatment cell $(g,s)$ with calendar time $t$ is the expected difference between the true and imputed counterfactuals for cohort $\mathcal{G}_g$. Its finite-sample expression is: $$\bar{Y}_{\mathcal{G}_g,t}(0)-\overline{\hat{Y}}_{\mathcal{G}_g, t}(0) = [\bar{Y}_{\mathcal{G}_g,t}(0)-\bar{Y}^*_{\mathcal{G}_g, \text{pre}_t}(0)] - [\bar{Y}^*_{\mathcal{C}_{g,1}, t}(0) - \bar{Y}^*_{\mathcal{C}_{g,1}, \text{pre}_t}(0)]$$In contrast, the finite-sample expression for the block bias is:$$[\bar{Y}_{\mathcal{G}_g,t}(0)-\bar{Y}_{\mathcal{G}_g, \text{pre}_g}(0)] - [\bar{Y}_{\mathcal{C}_{g,1}, t}(0) - \bar{Y}_{\mathcal{C}_{g,1}, \text{pre}_g}(0)]$$ Comparing these two expressions reveals two key distinctions: the overall bias calculation uses different reference periods ($\text{pre}_t$ vs. $\text{pre}_g$) and relies on potentially imputed counterfactuals of $\mathcal{G}_g$ and its initial control group (indicated by the notation $^*$). As shown in the proof, the difference in reference periods cancels out algebraically. The remaining difference—the use of previously imputed values—is what introduces the biases from the adjustment cohorts $\mathcal{K}_g(t)$. Consequently, the overall bias for $\mathcal{G}_g$ is the sum of its own block bias and the weighted sum of biases inherited from these adjustment cohorts.

Crucially, the biases “inherited" from the adjustment cohorts are the overall biases of those cohorts at calendar time $t$. The weight placed on each adjustment cohort, given by $\frac{N_k}{N_{\mathcal{C}_{g,1}}}$, is its share within $\mathcal{G}_g$'s initial control group, reflecting the importance of $\mathcal{G}_k$ in constructing $\mathcal{G}_g$'s potential outcome. This relationship provides a decomposition expressing the overall bias of cohort $\mathcal{G}_g$ as its own block bias plus the weighted overall biases from adjustment cohorts. This kind of decomposition is then applied sequentially from the earliest-treated to the latest-treated cohort. Since the overall bias of each adjustment cohort can, in turn, be decomposed into its own block bias and the overall biases of its own adjustment cohorts, this process creates a nested, recursive structure that ultimately expresses any overall bias as a linear combination of only block biases.

Solving this nested recursion yields the final decomposition: any cohort's overall bias is the sum of its own block bias plus a linear combination of the block biases from its adjustment cohorts. The weights in this combination are given by $w_k = \frac{N_k}{\sum_{j=k}^G N_j + N_\infty}$. This weight, $w_k$, represents the total contribution of the block bias from a later cohort, $\mathcal{G}_k$, to the overall bias of an earlier cohort, $\mathcal{G}_g$. It captures both direct and indirect paths of influence from $\mathcal{G}_k$ that are mediated through the overall biases of intermediate cohorts during the sequential imputation.

The example in Figure (ref) makes this logic concrete. At $t=8$, we have the relative periods for each cohort as $s_8(8)=1$, $s_6(8)=3$, and $s_4(8)=5$, so \[ \delta^{\text{Imp}}_{\mathcal G_6,3} = \Delta^{\text{Imp}}_{\mathcal G_6,3} + w_8\,\Delta^{\text{Imp}}_{\mathcal G_8,1}, \qquad \delta^{\text{Imp}}_{\mathcal G_4,5} = \Delta^{\text{Imp}}_{\mathcal G_4,5} + w_6\,\Delta^{\text{Imp}}_{\mathcal G_6,3} + w_8\,\Delta^{\text{Imp}}_{\mathcal G_8,1}. \] With cohort sizes $(N_6,N_8,N_\infty)=(3,1,4)$, \[ w_8=\frac{N_8}{N_8+N_\infty}=\frac{1}{1+4}=0.2, \qquad w_6=\frac{N_6}{N_6+N_8+N_\infty}=\frac{3}{3+1+4}=0.375. \] For $\delta^{\text{Imp}}_{\mathcal G_4,5}$, the contribution of bias from $\mathcal G_8$ can be represented as the sum of a direct path and an indirect path via $\mathcal G_6$: \[ \underbrace{\frac{N_8}{N_6+N_8+N_\infty}}_{\text{direct at }t} \;+\; \underbrace{\frac{N_6}{N_6+N_8+N_\infty}\cdot \frac{N_8}{N_8+N_\infty}}_{\text{via }\mathcal G_6} \;=\; \frac{N_8}{N_8+N_\infty} \,=\, w_8, \]

Here, the indirect path captures the bias transmitted from $\mathcal{G}_8$ to $\mathcal{G}_4$ through the intermediate imputation of $\mathcal{G}_6$'s counterfactual. Its weight is the product of two components. The first fraction, $\frac{N_6}{N_6+N_8+N_\infty}$, is the share of cohort $\mathcal{G}_6$ in $\mathcal{G}_4$'s initial control group, representing the weight placed on $\mathcal{G}_6$ when imputing for $\mathcal{G}_4$. The second fraction, $\frac{N_8}{N_8+N_\infty}$, is the weight of $\mathcal{G}_8$ when imputing for $\mathcal{G}_6$. As the equation shows, the sum of this indirect path and the direct path simplifies to the total weight, $w_8$.

For the CS-NYT estimator, the mechanism is analogous but more immediate. The overall bias for an early-adopting cohort $\mathcal{G}_g$ compares its trend to that of its not-yet-treated control group, $\mathcal{C}_{g,s}$, while its block bias compares $\mathcal{G}_g$ to its full initial control group, $\mathcal{G}_{g,1}$. Both biases are defined with respect to the reference period $t_g-1$. The difference between these two biases is therefore driven entirely by the adjustment cohorts $\mathcal{K}_g(t)$—that is, the cohorts that have “dropped out" of the initial control group by becoming treated.

The difference arising from the “drop-outs" can be expressed as a weighted sum over these adjustment cohorts. Each term in the sum is the trend difference between an adjustment cohort ($\mathcal{G}_k$) and the currently not-yet-treated control group. Crucially, this not-yet-treated control group is the same for both $\mathcal{G}_g$ and $\mathcal{G}_k$ at given time $t$. The weight for each term is the adjustment cohort's share within the initial control group of $\mathcal{G}_g$.

This step reveals a recursive structure analogous to that of the imputation estimator: the overall bias of an early-adopting cohort can be written as its own block bias plus weighted sum of trends difference between its adjustment cohorts and the currently not-yet-treated group. Different from the imputation estimator, this trend difference of an adjustment cohort $\mathcal{G}_k$ is measured relative to cohort $\mathcal{G}_g$'s reference period rather than $\mathcal{G}_k$'s reference period, so it cannot be interpreted as the overall bias of cohort $\mathcal{G}_k$. This mismatch necessitates a correction to realign the reference periods, which is why the adjustment term in proposition (ref) takes the form of a difference between two block biases. Once this realignment correction is introduced, one can solve this recursive structure and express the overall bias as the linear combination of block biases.

Estimation of Pre-Treatment Block Biases

Having established that the overall bias is a linear combination of block biases, the goal of robust inference is to use the observable pre-treatment block biases to restrict their unobservable post-treatment counterparts. These restrictions on block biases can then be translated into restrictions on the overall biases. The first step in this process is to estimate the block biases in the pre-treatment periods.

The pre-treatment block biases for both the CS-NYT and imputation estimators are estimated by directly applying their respective definitions to the observed data. The calculation of block biases for each estimator is therefore comparing cohort $\mathcal{G}_g$'s outcome in each pre-treatment period to that of its initial control group $\mathcal{C}_{g,1}$, using either the last pre-treatment period as the reference period (CS-NYT) or the average outcome across all pre-treatment periods as reference (imputation estimator).

align*[align* omitted — 471 chars of source]

The did package didpackage incorporates the estimated pre-treatment block biases for the CS-NYT estimator as part of its output and provides ways to visualize them. In contrast, to my knowledge, no statistical program directly reports the corresponding pre-treatment block biases for the imputation estimator, let alone uses them to visualize the pre-trend for each cohort.

For the imputation estimator, the estimated pre-treatment block biases are not constructed strictly symmetrically with the post-treatment effects (in the sense of Roth2024interpret). To illustrate, consider the latest treated cohort, $\mathcal{G}_G$. For this cohort, the initial control group is the never-treated cohort, $\mathcal{G}_{\infty}$, and its overall bias is equal to its block bias in all post-treatment periods. Therefore, both its pre-treatment block biases and its post-treatment estimated effects for period $s$ can be expressed by the following formula: $\frac{1}{N_{G}}\sum_{i \in \mathcal{G}_{G}}[Y_{i,t_{G}+s-1}-\bar{Y}_{i,\text{pre}_{G}}]-\frac{1}{N_{\infty}}\sum_{i \in \mathcal{G}_{\infty}}[Y_{i,t_{G}+s-1}-\bar{Y}_{i,\text{pre}_{G}}]$. The asymmetry arises from the construction of the pre-treatment average, $\bar{Y}_{i,\text{pre}_{G}}$. For pre-treatment periods ($s \le 0$), the outcome $Y_{i,t_{G}+s-1}$ is included in the calculation of $\bar{Y}_{i,\text{pre}_{G}}$. For post-treatment periods ($s > 0$), it is not.

Despite this asymmetry in construction, the block bias can still serve as the building block for robust inference. First, the interpretation of the pre- and post-treatment block biases is still the same, as both compare the treated cohort to its initial controls relative to the same pre-treatment baseline. Second, for the RM and SD restriction sets considered in this paper, the benchmark is constructed using differences in block biases between consecutive or multiple periods. The asymmetry in construction will not distort these differences between periods. Appendix (ref) discusses alternative procedures for estimating pre-trends for the imputation estimator and their relationship with the pre-treatment block bias.

To facilitate a direct comparison between the cohort-anchored and aggregated frameworks in later sections, I construct aggregated pre-treatment coefficients from the estimated block biases for both estimators. The procedure involves aggregating the pre-treatment block biases by relative time period, using weights proportional to each cohort's size. The did package provides similar functionality for the CS-NYT estimator.

For inference, I compute standard errors and the full variance-covariance (VCOV) matrix for all cohort-period coefficients (including the estimated pre-treatment block biases and post-treatment treatment effects for each cohort-period cell) using a stratified cluster bootstrap. The procedure involves resampling units with replacement from within their original cohorts, thereby preserving the fixed cohort structure of the data. When a unit is selected, its entire time-series of observations is included in the bootstrap sample, which accounts for potential serial correlation in the errors. On each bootstrap sample, the full set of pre-treatment block biases and post-treatment treatment effects is re-estimated.

The Robust Inference Framework

The previous section has shown that the overall bias vector of the imputation estimator can be written as a linear transformation of the block bias vector, $\vec{\delta} = \mathbf{W}\vec{\Delta}$, and that the pre-treatment block biases can be estimated. Armed with these results, I now formally introduce the cohort-anchored robust inference framework. I will focus the main analysis on the imputation estimator, and the framework also works for the CS-NYT estimator once the definitions of overall bias and block bias are changed. Accordingly, for notational simplicity, I will omit the $\text{Imp}$ superscript from the overall and block bias terms unless I am explicitly discussing the differences between the two estimators.

First, I present the general setup and then describe how to impose credible restrictions on the block bias vector, $\vec{\Delta}$. I focus on two common restrictions introduced by \textcolor{NavyBlue}{RR (rambachan2023more)} —Relative Magnitudes (RM) and Second Differences (SD). Next, I adapt the machinery of \textcolor{NavyBlue}{RR (rambachan2023more)} to my setting and discuss the algorithm and computational considerations. Finally, I examine the respective roles of identification and statistical uncertainty.

Setup

The cohort–anchored robust inference framework operates on cohort–period estimates of the imputation estimator. For any treated cohort $\mathcal{G}_{g}$, the pre–treatment coefficients ($s \leq 0$) are the estimated block biases, $\hat{\Delta}_{\mathcal{G}_{g},s}$, as defined in the previous section: \[ \hat{\beta}_{\mathcal{G}_{g}}^{s}=\hat{\Delta}_{\mathcal{G}_{g},s}=\frac{1}{N_{g}}\sum_{i \in \mathcal{G}_{g}}\!\bigl[Y_{i,t_{g}+s-1}-\bar{Y}_{i,\text{pre}_{g}}\bigr]-\frac{1}{N_{\mathcal{C}_{g,1}}}\sum_{i \in \mathcal{C}_{g,1}}\!\bigl[Y_{i,t_{g}+s-1}-\bar{Y}_{i,\text{pre}_{g}}\bigr], \] where $\mathcal{C}_{g,1}$ is the initial control group for $\mathcal{G}_{g}$ and $\text{pre}_{g}=\{1,\ldots,t_g-1\}$ denotes its pre–treatment periods. For post–treatment periods ($s \geq 1$), the coefficients are the estimated cohort–period ATTs obtained from the imputation estimator: \[ \hat{\beta}_{\mathcal{G}_{g}}^{s}=\frac{1}{N_{g}}\sum_{i \in \mathcal{G}_{g}}\!\bigl[Y_{i,t_{g}+s-1}-\hat{Y}_{i,t_{g}+s-1}(0)\bigr], \] where $\hat{Y}_{i,t_{g}+s-1}(0)$ is the imputed counterfactual for unit $i$ in period $t_{g}+s-1$.

These coefficients are then stacked into vectors. Let $\vec{\hat{\beta}}_{\mathcal{G}_{g},\text{pre}}$ and $\vec{\hat{\beta}}_{\mathcal{G}_{g},\text{post}}$ be the vectors of pre- and post-treatment coefficients for cohort $\mathcal{G}_{g}$, respectively. These are then stacked across all $G$ cohorts into full pre- and post-treatment vectors: \[ \vec{\hat{\beta}}_{\text{pre}}= \bigl[\vec{\hat{\beta}}_{\mathcal{G}_{1},\text{pre}}',\ldots,\vec{\hat{\beta}}_{\mathcal{G}_{G},\text{pre}}'\bigr]' \quad\text{and}\quad \vec{\hat{\beta}}_{\text{post}}= \bigl[\vec{\hat{\beta}}_{\mathcal{G}_{1},\text{post}}',\ldots,\vec{\hat{\beta}}_{\mathcal{G}_{G},\text{post}}'\bigr]'. \] By construction, the estimated pre-treatment coefficients are precisely the pre-treatment block biases, so $\vec{\hat{\beta}}_{\text{pre}}=\vec{\hat{\Delta}}_{\text{pre}}$. The full vector of all estimated coefficients is therefore $\vec{\hat{\beta}}=[\vec{\hat{\beta}}_{\text{pre}}',\vec{\hat{\beta}}_{\text{post}}']'$. Let $\vec{\beta}$, $\vec{\beta}_{\text{pre}}$, $\vec{\beta}_{\text{post}}$, and $\vec{\Delta}_{\text{pre}}$ denote the corresponding population analogues. The framework's inferential approach relies on the asymptotic normality of these stacked coefficients, as formalized below.

assumption[Asymptotic Normality] Let $N$ be the total number of units. Assuming cohort shares remain fixed as $N {\rightarrow} \infty$, the stacked vector of all estimated cohort–period coefficients, $\vec{\hat{\beta}}$, is asymptotically normally distributed: \[ \sqrt{N}\,(\vec{\hat{\beta}}-\vec{\beta}) \stackrel{d}{{\rightarrow}} \mathcal{N}\!\left(0,\Sigma^*\right), \] where $\Sigma^*$ is the asymptotic variance–covariance matrix of the estimator. A consistent estimator $\hat{\Sigma}_N$ of the finite–sample variance approximation $\Sigma_N=\Sigma^*/N$ is available.

Let $\tau_{\mathcal{G}_g,s}$ denote the true treatment effect for cohort $\mathcal{G}_g$ in relative period $s$, and let $\delta_{\mathcal{G}_g,s}$ be the corresponding overall bias. Stacking these cohort-period level quantities into vectors $\vec{\tau}$ and $\vec{\delta}$, using the same ordering as for $\vec{\beta}$, the full vector of population coefficients, $\vec{\beta}$, decomposes as: \[ \vec{\beta} =\binom{\vec{\beta}_{\text{pre}}}{\vec{\beta}_{\text{post}}}= \underbrace{\binom{\vec{\tau}_{\text{pre}}}{\vec{\tau}_{\text{post}}}}_{=:\,\vec{\tau}} + \underbrace{\binom{\vec{\delta}_{\text{pre}}}{\vec{\delta}_{\text{post}}}}_{=:\,\vec{\delta}}, \qquad \text{with } \vec{\tau}_{\text{pre}}=0 \text{ and } \vec{\delta}_{\text{pre}}=\vec{\Delta}_{\text{pre}}. \] The restriction $\vec{\tau}_{\text{pre}}=0$ follows from the no-anticipation assumption. This decomposition implies that in pre-treatment periods, the population coefficients equal the block biases ($\vec{\beta}_{\text{pre}} = \vec{\Delta}_{\text{pre}}$), while in post-treatment periods, they are the sum of the true treatment effect and the overall bias ($\vec{\beta}_{\text{post}} = \vec{\tau}_{\text{post}} + \vec{\delta}_{\text{post}}$).

The parameter of interest is typically a linear combination of the true cohort-period level treatment effects $\theta \;=\; \ell'\vec{\tau}_{\text{post}}$, such as the average treatment effect across all cohorts and post–treatment periods or the average treatment effect across all cohorts at a certain relative period $s$.

The robust inference framework uses the observable pre-treatment block biases ($\vec{\beta}_{\text{pre}}=\vec{\Delta}_{\text{pre}}$) to impose restrictions on their unobserved post-treatment counterparts ($\vec{\Delta}_{\text{post}}$). It then connects the full vector of block biases to the overall biases via the bias decomposition $\vec{\delta}=\mathbf{W}\vec{\Delta}$. Formally, this restriction is expressed as the assumption that the full block bias vector lies within a set $\Lambda_{\Delta}$ (i.e., $\vec{\Delta} \in \Lambda_{\Delta}$). Following \textcolor{NavyBlue}{RR (rambachan2023more)}, this paper focuses on restriction sets that take the form of a single polyhedron or a union of polyhedra. A restriction set $\Lambda_\Delta$ is polyhedral if it can be written as $\Lambda_\Delta=\{\vec{\Delta} : A \vec{\Delta} \leq d\}$ for some known matrix $A$ and vector $d$.

Similar to \textcolor{NavyBlue}{RR (rambachan2023more)}, I formalize the robust inference by first defining the identified set, $\mathcal{S}(\vec{\beta}, \Lambda_{\Delta})$, as the set of all possible values for the target parameter $\theta$ that are consistent with the population coefficients ($\vec{\beta}$) and the assumed restrictions on the block biases ($\Lambda_{\Delta}$): \[\mathcal{S}(\vec{\beta}, \Lambda_{\Delta}) = \left\{ \theta : \exists \vec{\delta} \text{ s.t. } \theta = \ell'(\vec{\beta}_{\text{post}} - \vec{\delta}_{\text{post}}),\; \vec{\delta}_{\text{pre}} = \vec{\Delta}_{\text{pre}} = \vec{\beta}_{\text{pre}},\; \vec{\delta}=\mathbf{W}\vec{\Delta}, \text{ and } \vec{\Delta} \in \Lambda_{\Delta} \right\}\] The identified set for $\theta$ consists of all values of the parameter of interest $\ell'\vec{\tau}_{\text{post}}$ consistent with a plausible overall bias vector $\vec{\delta}$ that satisfies four conditions: (1) the true treatment effect is equal to the post-treatment population coefficients minus the overall bias, $\vec{\tau}_{\text{post}} = \vec{\beta}_{\text{post}} - \vec{\delta}_{\text{post}}$; (2) the pre-treatment overall bias is equal to the pre-treatment block bias and the pre-treatment coefficients, $\vec{\delta}_{\text{pre}}=\vec{\Delta}_{\text{pre}} = \vec{\beta}_{\text{pre}}$; (3) the overall bias $\vec{\delta}$ is a linear transformation of the block biases, $\vec{\delta}=\mathbf{W}\vec{\Delta}$; and (4) the underlying block biases $\vec{\Delta}$ fall within the assumed restriction set, $\vec{\Delta} \in \Lambda_{\Delta}$.

For classes of restrictions that take a polyhedral form or a union of polyhedra, my framework provides a straightforward way to map the restrictions on block biases ($\vec{\Delta}$) to equivalent restrictions on overall biases ($\vec{\delta}$). This is achieved using the invertible linear transformation matrix, $\mathbf{W}$, from the bias decomposition. Specifically, I can define the restriction set for overall biases, $\Lambda_{\delta}$, as the image of the block bias set $\Lambda_{\Delta}$ under the linear map $\mathbf{W}$: \[ \Lambda_{\delta} := \mathbf{W}\Lambda_{\Delta} = \bigl\{\vec{\delta}:\ \exists\,\vec{\Delta}\in\Lambda_{\Delta}\ \text{s.t.}\ \vec{\delta}=\mathbf{W}\vec{\Delta}\bigr\}. \] This transformation preserves the structure of the restriction set. If $\Lambda_{\Delta}$ is a polyhedron defined by $\{\vec{\Delta}:A\vec{\Delta}\le d\}$, then $\Lambda_{\delta}$ is also a polyhedron with the explicit representation $\{\vec{\delta}: A\mathbf{W}^{-1}\vec{\delta}\le d\}$. Similarly, if $\Lambda_{\Delta}$ is a union of polyhedra, $\Lambda_{\Delta}=\bigcup_{k=1}^K\left\{\vec{\Delta}: A_k \vec{\Delta} \leq d_k\right\}$, then $\Lambda_{\delta}$ is also a union of polyhedra given by $\Lambda_\delta = \bigcup_{k=1}^K\left\{\vec{\delta}: A_k \mathbf{W}^{-1} \vec{\delta} \leq d_k\right\}$.

Using this mapped set, the identified set for $\theta$ can be written with an explicit restriction on $\vec{\delta}$: \[ \mathcal{S}(\vec{\beta},\Lambda_{\delta}) =\Bigl\{\theta:\ \exists\,\vec{\delta}\in\Lambda_{\delta}\ \text{s.t.}\ \theta=\ell'(\vec{\beta}_{\text{post}}-\vec{\delta}_{\text{post}}),\ \vec{\delta}_{\text{pre}}=\vec{\beta}_{\text{pre}}\Bigr\}. \] This formulation is identical to $\mathcal{S}(\vec{\beta},\Lambda_{\Delta})$ because the condition $\vec{\delta}\in\Lambda_\delta$ is equivalent to the simultaneous conditions that $\vec{\delta}=\mathbf{W}\vec{\Delta}$ for $\vec{\Delta}\in\Lambda_{\Delta}$.

It is noteworthy that the definition of the identified set, $\mathcal{S}(\vec{\beta}, \Lambda_{\Delta})$, uses the population coefficients, $\vec{\beta}$, rather than their sample estimates, $\vec{\hat{\beta}}$. The identified set is therefore a conceptual object that captures only identification uncertainty arising from the partial identification of the model; it does not account for the statistical uncertainty inherent in the estimated coefficients.

Placing Restrictions on Block Biases

I now turn to the choice of the restriction set, $\Lambda_{\Delta}$, imposed on the block bias. My discussion focuses on two prominent restrictions introduced by \textcolor{NavyBlue}{RR (rambachan2023more)}—the RM restriction and the SD restriction—while noting that my framework is flexible and can readily accommodate other restrictions in \textcolor{NavyBlue}{RR (rambachan2023more)}, such as sign and monotonicity constraints, or combinations of these polyhedral restrictions.

RM Restriction

The RM restriction formalizes the intuition that while one cannot observe the post-treatment bias, it is unlikely to change more erratically or dramatically than the bias before treatment. As proposed in \textcolor{NavyBlue}{RR (rambachan2023more)}, the RM restriction operationalizes this by asserting that the change in violation of PT between any two consecutive post-treatment periods is bounded by a sensitivity parameter ($\overline{M}$) multiplied by the largest change in the violation of PT between any two consecutive pre-treatment periods. By varying $\overline{M}$, researchers can assess how their conclusions depend on the assumed severity of the post-treatment PT violation relative to what is observed in the pre-treatment data.

In the cohort-anchored robust inference framework, I impose the RM restriction directly to the block bias vector, $\vec{\Delta}$. The by-cohort nature of my approach allows for a natural extension that incorporates potential heterogeneity across cohorts when setting the benchmark. Specifically, I consider two versions of the RM restriction that differ in how they define this benchmark.

The first approach uses a single global benchmark, derived from the pre-treatment data of all cohorts combined. The restriction set is: \[ \Lambda_{\Delta}^{\text{RM, Global}}(\overline{M}) = \left\{ \vec{\Delta} : |\Delta_{\mathcal{G}_g,s} - \Delta_{\mathcal{G}_g,s-1}| \le \overline{M} \cdot \max_{k \in \{1..G\}, s' \leq 0} |\Delta_{\mathcal{G}_k,s'} - \Delta_{\mathcal{G}_k,s'-1}| \quad \forall s\geq1, \forall g \in \{1..G\} \right\} \] This restriction asserts that the change in block bias for any cohort between consecutive post-treatment periods cannot exceed $\overline{M}$ times the single largest such change between any two consecutive pre-treatment periods across all cohorts.

Alternatively, a more flexible approach is to use cohort-specific benchmarks, where each cohort’s post-treatment bias is benchmarked only against its own pre-treatment history: \[ \Lambda_{\Delta}^{\text{RM, Cohort}}(\overline{M}) = \left\{ \vec{\Delta} : |\Delta_{\mathcal{G}_g,s} - \Delta_{\mathcal{G}_g,s-1}| \le \overline{M} \cdot \max_{s' \le 0} |\Delta_{\mathcal{G}_g,s'} - \Delta_{\mathcal{G}_g,s'-1}| \quad \forall s\geq1, \forall g \in \{1..G\} \right\} \] This restriction requires that the change in a cohort's block bias between consecutive post-treatment periods cannot exceed $\overline{M}$ times the largest such change between any two consecutive pre-treatment periods of that cohort.

The cohort-specific benchmark is more theoretically sound, as it uses one cohort's own pre-treatment history to restrict its post-treatment block bias. It also leads to a narrower identified set for the same value of $\overline{M}>0$ (i.e., $\mathcal{S}(\vec{\beta},\Lambda_{\Delta}^{\text{RM, Cohort}}(\overline{M})) \subseteq \mathcal{S}(\vec{\beta},\Lambda_{\Delta}^{\text{RM, Global}}(\overline{M}))$), reflecting less identification uncertainty. However, this theoretical advantage comes with a practical drawback: the cohort-specific benchmark is more computationally demanding than the global benchmark, particularly in settings with a large number of treatment cohorts, as I will discuss in detail later.

Figure (ref) provides a geometric illustration of the two restriction sets for a simple case with two cohorts, $\mathcal{G}_{i}$ and $\mathcal{G}_{j}$. The x-axis represents the change in block bias for cohort $\mathcal{G}_{i}$ between its first post-treatment and last pre-treatment period, $\Delta_{\mathcal{G}_{i},1}-\Delta_{\mathcal{G}_{i},0}$, and the y-axis represents the same quantity for cohort $\mathcal{G}_{j}$.

The RM restriction set is the union over all possible pre-treatment benchmarks. For the global benchmark, this can be expressed as: $$\Lambda_{\Delta}^{\text{RM, Global}}(\overline{M}) = \bigcup_{b \in \mathcal{B}} \left\{ \vec{\Delta} : |\Delta_{\mathcal{G}_g,s} - \Delta_{g,s-1}| \le b \quad \forall s\geq1, \forall g \in \{i,j\} \right\}$$ where $\mathcal{B}$ is the set of all possible benchmark values, $\{\overline{M} \cdot |\Delta_{k,s'} - \Delta_{k,s'-1}|\}_{k \in \{i,j\}, s' \le 0}$. For any single benchmark value $b \in \mathcal{B}$, the restriction is a square. The RM restriction set (outlined in red) is thus the union of all these squares, which simplifies to the single largest square.

For the cohort-specific benchmark, the restriction set is a union over the Cartesian product of each cohort's independent set of benchmarks: $$\Lambda_{\Delta}^{\text{RM, Cohort}}(\overline{M}) = \bigcup_{(b_i, b_j) \in \mathcal{B}_i \times \mathcal{B}_j} \left\{ \vec{\Delta} : |\Delta_{\mathcal{G}_g,s} - \Delta_{\mathcal{G}_g,s-1}| \le b_g \quad \forall s\geq1, g \in \{i,j\} \right\}$$ where $\mathcal{B}_i = \{\overline{M} \cdot |\Delta_{\mathcal{G}_i,s'} - \Delta_{\mathcal{G}_i,s'-1}|\}_{s' \le 0}$, and likewise for $\mathcal{B}_j$. Each pair of benchmarks $(b_i, b_j)$ defines a rectangle. The RM restriction set is thus the union of all possible rectangles, which results in the largest rectangle. This rectangle has the same width as the largest square in the left, but with a smaller height, since the largest benchmark for Cohort $\mathcal{G}_j$ is smaller than the largest overall benchmark.

figure[figure omitted — 623 chars of source]

In the special case where $\overline{M}=0$, both the global and cohort-specific restrictions simplify to the same condition: $|\Delta_{\mathcal{G}_g,s} - \Delta_{\mathcal{G}_g,s-1}| = 0$ for all $s \ge 1$. This implies that for each cohort, the block bias must remain constant across all post-treatment periods, equal to its value in the last pre-treatment period ($s=0$).

A key feature of the imputation estimator is that the block bias in the last pre-treatment period is not mechanically normalized to zero (i.e., $\Delta^{\text{Imp}}_{\mathcal{G}_g,0} \neq 0$). This distinguishes it from the original framework in \textcolor{NavyBlue}{RR (rambachan2023more)}, which uses estimators where the coefficient for the last pre-treatment period is, by construction, normalized to zero. This feature implies that a non-zero “baseline” block bias, $\Delta^{\text{Imp}}_{\mathcal{G}_g,0}$, persists throughout the post-treatment periods for cohort $\mathcal{G}_g$ in my framework. For example, when $\overline{M}=0$, the restriction implies that $\Delta^{\text{Imp}}_{\mathcal{G}_g,s}=\Delta^{\text{Imp}}_{\mathcal{G}_g,0}$ for all $s \geq 1$. In this scenario, my robust inference procedure can be interpreted as a debiasing tool, subtracting the estimated block bias in the last pre-treatment period from the estimates in all post-treatment periods.

One can also incorporate a normalization condition inherent to the imputation estimator into the restriction set: for any given cohort, the sum of its pre-treatment block biases must be zero ($\sum_{s \leq 0}\Delta^{\text{Imp}}_{\mathcal{G}_{g},s} = 0$). This condition is a mechanical property of the estimator. Because this normalization only operates on pre-treatment block biases, it does not alter the identified set for the treatment effects. In my implementation, I also find that including this condition does not affect the constructed confidence sets.

The CS-NYT estimator normalizes the block bias in the last pre-treatment period to zero, making it more analogous to the original implementation in \textcolor{NavyBlue}{RR (rambachan2023more)}. Consequently, when $\overline{M}=0$, all post-treatment block biases for each cohort are restricted to zero. However, this does not imply that the overall bias in post-treatment periods is necessarily zero. The reason, as the bias decomposition for the CS-NYT estimator shows, is that the overall bias for an early cohort ($\mathcal{G}_g$) includes a correction term from each of its adjustment cohorts ($\mathcal{G}_k$). This correction term takes the form $\Delta^{\text{CS-NYT}}_{\mathcal{G}_k,s_k(t)}-\Delta^{\text{CS-NYT}}_{\mathcal{G}_k,s_k(t_g-1)}$, and it includes not only a post-treatment block bias (the first term) but also a pre-treatment block bias (the second term). While the RM restriction with $\overline{M}=0$ forces the post-treatment component to zero, it does not affect the non-zero pre-treatment block biases. Therefore, the overall bias for $\mathcal{G}_g$ can be non-zero even when all post-treatment block biases are restricted to zero. In effect, this non-zero overall bias quantifies the “hidden bias" inherited from the adjustment cohorts' pre-treatment trends, and the framework de-biases the original estimate by accounting for this contamination.

SD Restriction

The SD restriction offers an alternative approach that is particularly well-suited for settings where pre-trends are approximately linear. The intuition is that while the level of the bias path may change after treatment, its slope is unlikely to change dramatically. My framework operationalizes this by bounding the change in the slope of each cohort's block bias path—the second difference—by a sensitivity parameter, $M$: \[ \Lambda_{\Delta}^{\text{SD}}(M) = \left\{ \vec{\Delta} : |(\Delta_{\mathcal{G}_g,s} - \Delta_{\mathcal{G}_g,s-1}) - (\Delta_{\mathcal{G}_g,s-1} - \Delta_{\mathcal{G}_g,s-2})| \le M \quad \forall s\geq1, \forall g \in \{1..G\} \right\}. \] The special case of $M=0$ imposes a strict linear extrapolation of the trend observed in the last two pre-treatment periods. Positive values of $M$ relax this by allowing for some degree of slope change. As this restriction is a single polyhedron, it is computationally simple.

The block biases in the last two pre-treatment periods play the most critical role in the SD restriction, as they establish the linear trend that serves as the benchmark for post-treatment periods. As I will demonstrate with simulated and empirical examples, large discrepancies between the cohort-anchored and aggregated frameworks often arise when using the SD restriction. This is because the benchmark learned from aggregated coefficients can be highly ill-suited for any individual cohort, leading to large differences in the final robust inference, especially when cohorts have heterogeneous pre-trends.

The robust inference under SD restriction with $M=0$ can also be viewed as a debiasing procedure that differs for each estimator. For the CS-NYT estimator, since $\Delta^{\text{CS-NYT}}_{\mathcal{G}_g,0}=0$ by construction, the procedure corrects estimates by subtracting a linear trend determined solely by the value of the block bias in the second-to-last pre-treatment period, $\Delta^{\text{CS-NYT}}_{\mathcal{G}_g,-1}$. For the imputation estimator, where $\Delta^{\text{Imp}}_{\mathcal{G}_g,0} \neq 0$, the procedure corrects estimates by subtracting a linear trend defined by both an intercept ($\Delta^{\text{Imp}}_{\mathcal{G}_g,0}$) and a slope ($\Delta^{\text{Imp}}_{\mathcal{G}_g,0}-\Delta^{\text{Imp}}_{\mathcal{G}_g,-1}$).

While the framework allows for cohort-specific sensitivity parameters ($\{M_g\}_{g=1}^G$), this paper focuses on the simpler case of a common parameter, $M$, for all cohorts.

Inferential Goal and Computational Considerations

Having defined the restrictions placed upon the block biases, I now formally define the confidence sets. As discussed before, I assume the stacked vector of cohort-period level coefficients, $\vec{\hat{\beta}}$, is asymptotically normal, $\sqrt{N}(\vec{\hat{\beta}}-\vec{\beta}) \rightarrow_d \mathcal{N}\left(0, \Sigma^*\right)$. It suggests a finite-sample normal approximation $\overrightarrow{\hat{\beta}} \approx_d \mathcal{N}\left(\vec{\beta}, \Sigma_N\right)$, where $\Sigma_N=\frac{\Sigma^*}{N}$. Following \textcolor{NavyBlue}{RR (rambachan2023more)}, I construct confidence sets that are uniformly valid for all parameter values $\theta$ in the identified set when this normal approximation holds with the known $\Sigma_N$. That is, I construct confidence sets $\mathcal{C}_N(\vec{\hat{\beta}}, \Sigma_N)$ satisfying \[ \inf_{\vec{\delta} \in \Lambda_\delta,\vec{\tau}} \inf_{\theta \in \mathcal{S}(\vec{\tau}+\vec{\delta}, \Lambda_\delta)} \mathbb{P}_{\vec{\hat{\beta}} \sim \mathcal{N}\left(\vec{\tau}+\vec{\delta}, \Sigma_N\right)}\left(\theta \in \mathcal{C}_N\left(\vec{\hat{\beta}}, \Sigma_N\right)\right) \geq 1-\alpha. \] where $\Lambda_{\delta}$ is the image of the block bias set $\Lambda_{\Delta}$ under the linear map $\mathbf{W}$.

As shown in \textcolor{NavyBlue}{RR (rambachan2023more)}, this finite-sample size control under the normal approximation translates to a uniform asymptotic size control over a large class of data-generating processes when $\Sigma_N$ is replaced by a consistent estimate, $\hat{\Sigma}_N$. The constructed confidence sets therefore satisfy \[ \liminf_{N {\rightarrow} \infty} \inf_{P \in \mathcal{P}} \inf_{\theta \in \mathcal{S}\left(\vec{\delta}_P+\vec{\tau}_P, \Lambda_\delta\right)} \mathbb{P}_P\left(\theta \in \mathcal{C}_N\left(\vec{\hat{\beta}}, \hat{\Sigma}_N\right)\right) \geq 1-\alpha \] for a large class of distributions $\mathcal{P}$ such that the true bias vector $\vec{\delta}_P$ is in the assumed restriction set $\Lambda_\delta$ for all $P \in \mathcal{P}$. I refer readers to the Section 3.3 of \textcolor{NavyBlue}{RR (rambachan2023more)} for the formal proofs and a more detailed technical discussion of these properties.

The algorithm in \textcolor{NavyBlue}{RR (rambachan2023more)} focuses on constructing confidence sets when the restriction set on the overall biases, $\Lambda_\delta$, is either a single polyhedron (such as for the SD restriction) or a finite union of polyhedra (such as for the RM restriction). When $\Lambda_\delta$ is a finite union of polyhedra, $\Lambda_\delta = \bigcup_{k=1}^K \Lambda_{\delta,k}$, a valid confidence set can be constructed by taking the union of the confidence sets for each of its polyhedral components. As established by Lemma 2.2 in \textcolor{NavyBlue}{RR (rambachan2023more)}, if a confidence set $\mathcal{C}_{N,k}$ provides valid coverage for a restriction set $\Lambda_{\delta,k}$, then the union of these sets, $\mathcal{C}_N = \bigcup_{k=1}^K \mathcal{C}_{N,k}$, provides valid coverage for the union set $\Lambda_\delta$.

The procedure for constructing confidence sets under a single polyhedron constraint adapts the moment inequality approach detailed in \textcolor{NavyBlue}{RR (rambachan2023more)}, which builds on the conditional and hybrid methods developed by andrews2023inference. The core of this approach is to invert a hypothesis test for the parameter of interest, $\theta$. For each candidate value $\theta_0$ on a pre-specified grid, the algorithm tests the joint null hypothesis that the true parameter is $\theta_0$ and that the true bias vector $\vec{\delta}$ satisfies the restrictions. I use the recommended hybrid test in the implementation of my framework. This approach is attractive because it remains computationally tractable even with many nuisance parameters and has desirable asymptotic power properties. For a complete technical discussion of its properties, I refer readers to \textcolor{NavyBlue}{RR (rambachan2023more)} and andrews2023inference.

A key computational consideration arises for the RM restriction, which is a union of polyhedra. As shown in Figure (ref), each polyhedron in the union corresponds to a different pre-treatment benchmark. Because the population “max benchmark" is unknown— we only observe the estimated pre-treatment block biases and sampling noise can change which pre-treatment period attains the maximum—the algorithm must compute a confidence set for each polyhedron individually and then take their union. This necessitates a loop over all possible benchmark configurations, and the structure of this loop differs between the two benchmark specifications.

The global benchmark approach requires a single loop through every pre-treatment difference between consecutive periods across all cohorts. The number of benchmark configurations thus grows linearly with the total number of pre-treatment periods, $O(t_{1}+\cdots+t_{G})$. In contrast, the cohort-specific benchmark requires a nested loop structure to test every possible combination of cohort-specific benchmarks. The number of benchmark configurations therefore grows multiplicatively with the number of cohorts, $O(t_{1}\times\cdots\times t_{G})$. It can become infeasible in settings with many cohorts. This computational burden is the practical price of the more theoretically appealing cohort-specific restriction.

For the specific case of the SD restriction, \textcolor{NavyBlue}{RR (rambachan2023more)} recommends an alternative approach based on Fixed-Length Confidence Intervals (FLCIs), which can offer improved power in finite samples when the identified set is small, while the hybrid method remains asymptotically valid. For simplicity and consistency in my analysis, I use the hybrid method for both the RM and SD restrictions.

Identification Uncertainty and Statistical Uncertainty

Finally, I discuss the two sources of uncertainty inherent in the robust inference framework. The restriction set, $\Lambda_{\Delta}$, is imposed on the true, unobservable vector of block biases, $\vec{\Delta}$. We do not observe this vector; we only have a noisy estimate of its pre-treatment components, $\hat{\vec{\Delta}}_{\text{pre}}$. Therefore, the inference procedure must account for two distinct sources of uncertainty.

First, there is identification uncertainty: even if we knew the population pre-treatment block biases, the restriction imposed on the post-treatment block biases would only allow partial identification of the parameter of interest, yielding the identified set. Second, there is statistical uncertainty, which arises because we only have noisy estimates of the pre-treatment block biases and post-treatment estimated treatment effects. The inference procedure described previously constructs confidence sets that are uniformly valid by simultaneously accounting for both sources of uncertainty. The benefit of a smaller identified set (e.g., $\mathcal{S}(\vec{\beta}, \Lambda_{\Delta}^{\text{Cohort}}) \subseteq \mathcal{S}(\vec{\beta}, \Lambda_{\Delta}^{\text{Global}})$) is therefore most pronounced when statistical uncertainty is low. When the variance of the estimated coefficients $\vec{\hat{\beta}}$ is large, the gains from a smaller identified set can be outweighed by the large degree of statistical uncertainty.

To illustrate these two sources of uncertainty, I conduct a simulation study comparing the confidence sets under the RM restriction with the global and cohort-specific benchmarks under different levels of statistical noise. I construct a simple two-cohort staggered adoption design with one cohort treated at $t=3$ ($\mathcal{G}_3$) and another at $t=5$ ($\mathcal{G}_5$), observed over $T=6$ periods. I directly generate a vector of estimated coefficients, $\vec{\hat{\beta}}$, where the “good” cohort, $\mathcal{G}_3$, has no pre-treatment block bias (all its pre-treatment $\hat{\Delta}_{\mathcal{G}_3,s}$ are set to zero), while the “bad” cohort, $\mathcal{G}_5$, has a non-zero pre-treatment block bias, which I introduce by setting $\hat{\Delta}_{\mathcal{G}_5,s=-3}=-0.25$ and $\hat{\Delta}_{\mathcal{G}_5,s=-2}=0.25$.

I then construct confidence sets for a parameter $\theta(w) = w \cdot \tau_{\mathcal{G}_3, s=1} + (1-w) \cdot \tau_{\mathcal{G}_5, s=1}$, which is a weighted average of the first-period treatment effects for the two cohorts. I vary the weight $w$ from 0 to 1 to trace a path of parameters, from the effect solely on the “bad” cohort ($w=0$) to the effect solely on the “good” cohort ($w=1$). This entire exercise is repeated for two levels of statistical uncertainty: a high-noise scenario with a baseline variance-covariance matrix, $\Sigma=\text{Diag}(0.1)$, and a low-noise scenario with $\Sigma/100$. The sensitivity parameter $\overline{M}$ is fixed at 1 for both restrictions throughout the simulation.

figure[figure omitted — 881 chars of source]

Figure (ref) plots the confidence sets for $\theta(w)$ across the range $w \in [0,1]$ in light gray and, for comparison, the corresponding plug-in identified set in dark gray. The plug-in identified set, $\mathcal{S}(\vec{\hat{\beta}},\Lambda_{\Delta})$, is obtained by substituting the estimated coefficients, $\vec{\hat{\beta}}$, directly into the definition of the identified set. It represents a hypothetical scenario with no statistical uncertainty, where the sample estimates are treated as the true population values. The upper and lower panels of the figure show the results for the high-noise and low-noise scenarios, respectively.

As expected, the plug-in identified set under the global benchmark remains constant for all values of $w$, as it applies the same single benchmark to both the “good" and “bad" cohorts. Under the cohort-specific benchmark, however, the identified set changes with $w$. The benchmark for the “bad" cohort, $\mathcal{G}_5$, is determined by its own pre-treatment violation, $|\hat{\Delta}_{\mathcal{G}_5,s=-2}-\hat{\Delta}_{\mathcal{G}_5,s=-3}|=0.5$. This implies that the identified set for its first-period treatment effect, $\tau_{\mathcal{G}_5, s=1}$, is $[-0.5, 0.5]$. For the “good" cohort, $\mathcal{G}_3$, the corresponding benchmark is zero, so the identified set for its effect is the single point $\{0\}$. Consequently, as the weight $w$ on the good cohort in the parameter of interest, $\theta(w)$, increases from 0 to 1, the plug-in identified set for $\theta(w)$ smoothly shrinks from $[-0.5, 0.5]$ to the point $\{0\}$.

The confidence set is determined by both identification uncertainty and statistical uncertainty. In the high-noise scenario (upper panel), statistical uncertainty dominates; the confidence sets are substantially wider than the identified sets, and the benefit of the cohort-specific benchmark is obscured. In the low-noise scenario (lower panel), the confidence sets are only slightly wider than the identified sets, revealing identification uncertainty as the primary driver. Here, the advantage of the cohort-specific benchmark becomes clear, as its confidence set visibly shrinks with the underlying identified set, while the global benchmark's confidence set remains wide.

The degree of statistical uncertainty not only informs the choice between the global and cohort-specific benchmarks but also adds to the more fundamental question of whether to conduct robust inference on aggregated or cohort-period level coefficients. As previously discussed, the aggregated framework suffers from theoretical issues such as dynamic control group and dynamic treated composition, which can obscure interpretation. However, it sometimes benefits from lower statistical uncertainty of the aggregated coefficients. When the statistical variance of the cohort-period level coefficients is particularly large, the theoretical benefits of the cohort-anchored framework may be outweighed by the loss of precision. Given the importance of these two sources of uncertainty, I recommend that empirical researchers plot both the plug-in identified sets and the final confidence sets. This practice separately visualizes the contributions of identification and statistical uncertainty, as I will demonstrate in the following section.

Two Illustrative Simulated Examples

This section presents two simulated examples to demonstrate the practical implications of the cohort-anchored framework. I contrast its performance with the robust inference that relies on aggregated event-study coefficients (the aggregated framework) to highlight the relative merits of my approach.

The first example applies the RM restriction. I compare the confidence sets generated by three approaches: the cohort-anchored framework using a global benchmark, the cohort-anchored framework using cohort-specific benchmarks, and aggregated framework. The second example employs the SD restriction and also provide a comparison between the cohort-anchored framework and the aggregated framework.

Example I: The RM Restriction

I simulate a panel dataset spanning 11 time periods with two treated cohorts, $\mathcal{G}_{8}$ and $\mathcal{G}_{10}$, and a never-treated group, $\mathcal{G}_{\infty}$. The first cohort, $\mathcal{G}_{8}$, with size $N_{8} = 40$, receives treatment at $t=8$, while the second cohort, $\mathcal{G}_{10}$, with size $N_{10} = 40$, is treated at $t=10$. An additional $N_{\infty} = 60$ units constitute the never-treated group. The baseline potential outcome, $Y_{it}(0)$, is generated from a TWFE model, $Y_{it}(0) = \alpha_i + \xi_t + \epsilon_{it}$, where $\alpha_i$ and $\xi_t$ are unit and time fixed effects and $\epsilon_{it} \sim N(0,2)$ is an idiosyncratic error term. I introduce an oscillating trend, given by $(-1)^{t+1}$, only to the late-adopting cohort, $\mathcal{G}_{10}$, as a violation of PT. This adds 1 to its potential outcome in all odd-numbered periods and subtracts 1 in all even-numbered periods. Finally, the observed outcome is generated by adding a constant treatment effect of $+3$ to all units in cohorts $\mathcal{G}_{8}$ and $\mathcal{G}_{10}$ during all of their respective post-treatment periods.

From this data, I apply the imputation estimator to obtain the pre-treatment block biases and post-treatment ATTs for each cohort. These coefficients are then stacked into pre-treatment and post-treatment vectors, respectively:

align*[align* omitted — 446 chars of source]

The full vector of coefficients is formed by stacking these two components, $\vec{\hat{\beta}}=[\vec{\hat{\beta}}_{\text{pre}}^{\prime},\vec{\hat{\beta}}_{\text{post}}^{\prime}]'$. I compute the full variance-covariance matrix for this vector using a stratified cluster bootstrap. For comparison with the aggregated framework, I also construct the corresponding aggregated coefficients by taking a cohort-size-weighted aggregation of these estimates at each relative time period.

Panels (a) and (b) of Figure (ref) present the cohort-period coefficients for cohort $\mathcal{G}_{8}$ and cohort $\mathcal{G}_{10}$, respectively. In each plot, the pre-treatment coefficients (block biases) are highlighted in dark green. Panel (b) shows that, consistent with my data generating process, the pre-treatment coefficients for $\mathcal{G}_{10}$ exhibit a pronounced oscillating pattern. Notably, because the initial control group for $\mathcal{G}_{8}$ includes the cohort $\mathcal{G}_{10}$, the pre-treatment coefficients for $\mathcal{G}_{8}$ in panel (a) also display a faint oscillation.

Panel (c) of Figure (ref) plots the aggregated event-study coefficients. In early pre-treatment periods ($s\in\{-8,-7\}$), only the late-treated cohort, $\mathcal{G}_{10}$, contributes to these coefficients. The aggregated plot clearly exhibits the oscillating trend introduced into $\mathcal{G}_{10}$ and signals an obvious violation of the PT assumption. In periods $-6\leq s \leq 2$, the aggregated coefficients are determined by both cohorts, while in late post-treatment periods ($s\in\{3,4\}$), only the early-treated cohort, $\mathcal{G}_{8}$, contributes.

This observation highlights the fundamental dilemma for robust inference using aggregated coefficients. On the one hand, it is inappropriate to use the pre-treatment violations from periods $s\in\{-8,-7\}$---which are specific to $\mathcal{G}_{10}$---to benchmark the effects for $\mathcal{G}_{8}$ or the cohort average. This is because the overall bias for $\mathcal{G}_{8}$ in its initial post-treatment periods is theoretically distinct from the block biases of $\mathcal{G}_{10}$. On the other hand, discarding these early coefficients of $s\in\{-8,-7\}$ would mean ignoring the most direct evidence of a PT violation for cohort $\mathcal{G}_{10}$ itself. Furthermore, the bias decomposition shows that these same block biases from $\mathcal{G}_{10}$ are mathematically necessary for calculating the overall bias of $\mathcal{G}_{8}$ in its later post-treatment periods ($s \in \{3,4\}$).

Panel (d) of Figure (ref) displays the plug-in identified sets and 95% confidence sets for the ATT across all post-treatment periods and treated cohorts. These sets are constructed under the RM restriction and are plotted against the sensitivity parameter $\overline{M}$. The plug-in identified set reflects identification uncertainty, whereas the confidence set captures both identification uncertainty and statistical uncertainty.

Within the cohort-anchored framework, the cohort-specific benchmark (red) restricts the change in a cohort's post-treatment block bias between consecutive periods to be no larger than $\overline{M}$ times the maximum corresponding change within its own pre-treatment history. In contrast, the global benchmark (green) requires this change to remain within $\overline{M}$ times the largest pre-treatment change across all cohorts. In this simulation, cohort $\mathcal{G}_{10}$ has a significantly larger pre-treatment bias than $\mathcal{G}_{8}$. Consequently, the global benchmark is dictated by the large violations in $\mathcal{G}_{10}$ and is therefore overly conservative for $\mathcal{G}_{8}$. As a result, the plug-in identified sets and confidence sets generated by the cohort-specific method is considerably narrower than that from the global method for any $\overline{M} > 0$, while the two are equivalent at $\overline{M} = 0$.

figure[figure omitted — 1,820 chars of source]

Under the aggregated framework,\footnote{I make minor modifications to the original HonestDiD package to allow for a non-zero estimate at the last pre-treatment period, which is suitable for the imputation estimator; details are in chiu2025causal.} the plug-in identified sets and 95% confidence sets for the ATT are shown in blue in Panel (d). Here the RM restriction bounds post-treatment changes in the aggregated bias based on the maximum corresponding change in the pre-treatment aggregated series. Because the pre-treatment coefficients for $s \in \{-8, -7\}$ exhibit a large violation of PT, the benchmark in the RM restriction for the aggregated series is also large. This leads to a confidence set that is wider than the one produced by the cohort-anchored framework with the cohort-specific benchmark.

It is noteworthy that a direct comparison of the confidence sets obtained from the cohort-anchored and aggregated frameworks requires careful interpretation. These methods are not perfectly analogous, as they apply the RM restriction to different objects: the former to cohort-period block biases and the latter to the aggregated biases. Consequently, the benchmark values used to construct the RM restriction sets are not directly comparable, even under the same sensitivity parameter $\overline{M}$. For instance, in a scenario where individual cohort pre-trends are volatile but their average is smooth, the cohort-specific approach could yield a wider plug-in identified set and confidence set than the aggregated approach due to its more conservative (i.e., larger) benchmark. The relative width of the plug-in identified sets and confidence sets from the two frameworks is therefore data-dependent.

The confidence sets produced by the cohort-anchored and aggregated frameworks are also centered at slightly different values. This occurs because the cohort-anchored framework adjusts each cohort's estimates using its own block bias from the last pre-treatment period, while the aggregated framework uses a single adjustment based on the average of the last-period block biases across both cohorts. This averaged adjustment is ill-suited for cohort $\mathcal{G}_8$ in its later post-treatment periods ($s \in \{3,4\}$), which are periods in the aggregated series that are identified solely by that cohort. This difference in centering becomes even more pronounced under the SD restriction.

In a parallel analysis presented in the appendix, I apply the CS-NYT estimator to the same simulated data. Figure (ref) displays the resulting event-study plots and confidence sets, which are qualitatively similar to those from the imputation estimator discussed in the main text.

Example II: The SD Restriction

In this section, I present a second simulated example to illustrate the application of the SD restriction within the cohort-anchored. To better motivate this restriction, I modify the data generating process. Instead of an oscillating trend, I now introduce a linear trend violation for cohort $\mathcal{G}_{10}$.

Specifically, the simulation design mirrors the first example, featuring two treated cohorts, $\mathcal{G}_{8}$ ($N_{8}=40$, treated at $t=8$) and $\mathcal{G}_{10}$ ($N_{10}=40$, treated at $t=10$), and a never-treated group, $\mathcal{G}_{\infty}$ ($N_{\infty}=60$), over 11 time periods. The baseline potential outcome again follows a TWFE model, $Y_{it}(0) = \alpha_i + \xi_t + \epsilon_{it}$. The key difference lies in the PT violation: I add a linear time trend, $0.75 \cdot t$, to the potential outcomes of the late-adopting cohort, $\mathcal{G}_{10}$. The treatment effect remains a constant $+3$ for all treated units in their respective post-treatment periods.

Following the same procedure as in the first example, I estimate the cohort-period level and aggregated coefficients and their corresponding variance-covariance matrices. Panels (a), (b), and (c) of Figure (ref) display the resulting event-study plots for cohort $\mathcal{G}_{8}$, cohort $\mathcal{G}_{10}$, and the aggregated coefficients, respectively.

In panel (b), the pre-treatment coefficients for $\mathcal{G}_{10}$ exhibit a pronounced positive linear trend, corresponding directly to the violation I introduced. Because the initial control group for $\mathcal{G}_{8}$ includes $\mathcal{G}_{10}$, the violation of PT in $\mathcal{G}_{10}$ in turn induces a mild downward trend in the estimated pre-treatment block biases for $\mathcal{G}_{8}$, as seen in panel (a).

The SD restriction uses the trend in the last two pre-treatment periods ($s \in \{-1, 0\}$) as its benchmark. The case of $M=0$ constrains this violation of PT to be perfectly linear in the post-treatment period, with a slope identical to that between $s=-1$ and $s=0$. Allowing $M>0$ relaxes this assumption by further permitting the slope of the bias to change between consecutive post-treatment periods.

When applying the cohort-anchored framework, this restriction is tailored appropriately: the post-treatment bias for $\mathcal{G}_{8}$ is benchmarked against its own pre-trend, and likewise for $\mathcal{G}_{10}$. The aggregated framework, however, uses a single benchmark derived from the slope of the averaged coefficients. This aggregated upward benchmark—an average of $\mathcal{G}_{8}$'s mild downward trend and $\mathcal{G}_{10}$'s strong upward trend—is ill-suited for both individual cohorts. It is too large in magnitude (and of the wrong sign) to be a plausible benchmark for $\mathcal{G}_{8}$, yet too small in magnitude to properly capture the steep trend for $\mathcal{G}_{10}$.

Panel (d) of Figure (ref) contrasts the plug-in identified sets and 95% confidence sets for the ATT, using the cohort-anchored framework and the aggregated framework. This comparison starkly illustrates the consequences of using an aggregated benchmark. The confidence set from the aggregated framework, while slightly narrower—likely due to the reduced statistical uncertainty of averaged coefficients—is systematically biased downwards. This bias is a direct result of applying the ill-suited benchmark derived from the aggregated pre-trends. Consequently, the plug-in identified set constructed using the aggregated framework is centered far below the true ATT, and its confidence set barely covers the true value. In contrast, the cohort-anchored framework produces a plug-in identified set and confidence set that are better centered on the true ATT.

figure[figure omitted — 1,711 chars of source]

To clarify the differences between the two frameworks, Figure (ref) disaggregates the results by plotting confidence sets for the ATT at each relative post-treatment period. Panel (a) shows the aggregated framework; panel (b) presents the cohort-anchored framework. The main divergence between the two methods appears at $s=3$ and $s=4$.

The ATTs in these later periods are identified solely from the early-treated cohort, $\mathcal{G}_{8}$. As established above, $\mathcal{G}_{8}$ exhibits a mild downward pre-trend. The aggregated framework, however, applies a benchmark derived from the upward trend in the averaged pre-treatment series. This mismatch leads to erroneous inference, producing the downward-biased confidence sets shown in panel (a) of Figure (ref). In contrast, the cohort-anchored framework avoids this problem. As shown in panel (b), it correctly applies the benchmark from $\mathcal{G}_{8}$’s own pre-treatment history to its post-treatment coefficients, yielding more accurately centered confidence sets.

figure[figure omitted — 1,114 chars of source]

In the appendix, I demonstrate that my main findings are robust to the choice of estimator by conducting a parallel analysis with the CS-NYT estimator (Figures (ref) and (ref)). The qualitative results are consistent. However, the discrepancy between the aggregated and cohort-anchored frameworks is less pronounced than with the imputation estimator.

Empirical Example: Minimum Wage and Teen Employment

To demonstrate the practical utility and implications of my framework, I revisit a prominent empirical example from callaway2021-did: the effect of minimum wage increases on teen employment. I apply the cohort-anchored robust inference framework to their data and compare my findings to both the original study's conclusions and the results from the robust inference based on the aggregated coefficients.

The empirical analysis uses county-level data on teen employment from the Quarterly Workforce Indicators for the period 2001--2007, during which the federal minimum wage was held constant at \$5.15 per hour. The “treatment” is defined as a state-level increase in the minimum wage above this federal baseline. Because states implemented these increases at different times (in year 2004, 2006, 2007), the design is naturally staggered, creating three distinct treatment cohorts based on the year of adoption. The analysis focuses on counties in states that initially adhered to the federal minimum wage; those in states that later raised their minimum wage form treatment groups, while those in states that did not change their policy serve as the never-treated control group.

In the original analysis, callaway2021-did finds evidence that increasing the minimum wage led to a reduction in teen employment. I aim to assess the sensitivity of these conclusions to potential violations of PT that are not explicitly modeled by the original estimator. With the imputation estimator, I consider the unconditional PT assumption that is not conditional on pre-treatment covariates.

Following the same procedure, I use the imputation estimator to obtain the cohort-period level and aggregated coefficients. Panels (a)--(d) of Figure (ref) display the event-study plots for cohort $\mathcal{G}_{4}$ ($N_4=100$), cohort $\mathcal{G}_{6}$ ($N_6=223$), cohort $\mathcal{G}_{7}$ ($N_7=584$), and the aggregated series, respectively.\footnote{The data covers the period 2001--2007. For notational convenience, I denote the cohort first treated in 2004 as $\mathcal{G}_{4}$, in 2006 as $\mathcal{G}_{6}$, and in 2007 as $\mathcal{G}_{7}$.} The number of units in the never-treated cohort is $N_{\infty}=1377$. The plots reveal notable heterogeneity in the pre-trends across cohorts. Both cohort $\mathcal{G}_{4}$ and cohort $\mathcal{G}_{6}$ exhibit upward trends in their pre-treatment periods, while cohort $\mathcal{G}_{7}$ shows a downward trend in its last two pre-treatment periods. This observed heterogeneity motivates the application of the cohort-anchored framework.

Panel (e) of Figure (ref) plots the identified and confidence sets derived from the cohort-anchored framework (both cohort-specific and global benchmarks) alongside those from the aggregated framework, all under the RM restriction. In this empirical example, the confidence set from the aggregated framework is narrower than those from the cohort-anchored frameworks when $\overline{M}>0$. This outcome stems from two factors.

First and more importantly, the aggregated pre-treatment series is less volatile than the individual cohort series, which results in a less conservative (i.e., smaller) benchmark in the RM restriction. This can be seen by comparing the plug-in identified sets: the set from the aggregated approach is much narrower. Secondly, the cohort-period level coefficients have larger statistical uncertainty than the aggregated coefficients. The combination of a more conservative restriction and higher statistical variance results in the wider confidence sets for the cohort-anchored framework.

It is possible that these wider confidence sets appear to make the cohort-anchored framework less attractive. However, as I discussed before, a direct comparison of the set widths between the cohort-anchored and aggregated frameworks is inappropriate. The RM restrictions are imposed on different objects and imply different benchmarks for the same $\overline{M}$. I therefore view the resulting wider confidence set not as a drawback, but as the necessary price for conducting a more transparent and interpretable robust inference.

Panel (f) of Figure (ref) contrasts the identified and confidence sets under the SD restriction, comparing the cohort-anchored framework to the aggregated framework. While the former yields a slightly wider confidence set, a key difference emerges in their location. The confidence set from the aggregated framework is centered around zero even when $M=0$, suggesting that the estimated negative treatment effect is not robust to a linear violation of the PT assumption. In contrast, the set from the cohort-anchored framework remains centered below zero and only barely includes zero when $M$ is small. This indicates that the negative treatment effect is largely robust to cohort-specific linear PT violations.

Panels (g) and (h) of Figure (ref) display the by-period confidence sets, which clarify the divergence observed in Panel (f). The aggregated approach in Panel (g) produces inaccurate results because the pre-trend of the aggregated series is dominated by the largest cohort, $\mathcal{G}_{7}$, which exhibits a mild downward trend. This creates an ill-suited benchmark for cohorts $\mathcal{G}_{4}$ and $\mathcal{G}_{6}$, both of which have upward pre-trends. Applying this mismatched benchmark results in significant bias, particularly for the ATT estimates in later post-treatment periods ($s \geq 2$), which are identified exclusively from cohorts $\mathcal{G}_{4}$ and $\mathcal{G}_{6}$. In contrast, the cohort-anchored framework in Panel (h) applies the appropriate, cohort-specific benchmark to each series, resulting in more credible, better-centered confidence sets. This suggests that the original study's conclusion of a negative treatment effect holds, and may even be larger in magnitude for cohorts $\mathcal{G}_{4}$ and $\mathcal{G}_{6}$ after accounting for their specific pre-trends.

In the appendix, I repeat the analysis using the CS-NYT estimator and find qualitatively similar conclusions (Figure (ref)). One notable difference arises under the RM restriction: the confidence sets from the cohort-anchored framework with the cohort-specific benchmarks are of a similar width to those from the aggregated approach. This is because the underlying benchmarks for the RM restriction are more comparable across the two frameworks, a fact reflected in the similar widths of their respective plug-in identified sets. This contrasts with the imputation estimator, where the cohort-anchored framework produced a more conservative benchmark. This highlights that the choice of estimator—and therefore the definition of the block biases—can impact the benchmarks used in robust inference and the width of the resulting confidence sets.

Taken together, the results from the RM and SD restrictions illustrate a pitfall of the aggregated framework in this application: it can lead to qualitatively different—and likely erroneous—conclusions about the robustness of the estimated treatment effects. The cohort-anchored framework, while sometimes sacrificing statistical power in favor of wider, more conservative confidence sets, provides a more transparent, interpretable, and more credible tool for robust inference.

figure[figure omitted — 2,561 chars of source]

Conclusion

This paper develops a cohort-anchored framework for conducting robust inference in staggered-adoption settings, addressing the problems of dynamic treated composition and dynamic control group that challenge robust inference based on aggregated event-study coefficients. I operates at the more granular cohort–period level and introduce the concept of the block bias—a PT violation for a given cohort relative to its anchored initial control group. I derive a bias decomposition that expresses the overall bias as an invertible linear transformation of these block biases. This decomposition lets us impose transparent, theoretically grounded restrictions on interpretable block biases and then map them to overall biases for inference using the machinery of \textcolor{NavyBlue}{RR (rambachan2023more)}. Simulated and empirical results show that the cohort-anchored framework can yield more accurately centered—and sometimes narrower—confidence sets than the aggregated approach, especially when pre-trends are heterogeneous across cohorts.

The framework is most suitable when there are several cohorts (but not too many if applying RM with cohort-specific benchmarks) and each cohort is large enough that cohort–period coefficients can be estimated with reasonable precision. It offers the greatest advantages over robust inference based on aggregated coefficients when by-cohort heterogeneity is substantial. At present, the framework does not accommodate missing data or estimators that weight units using pre-treatment covariates (e.g., inverse probability weighting or the doubly robust estimators in callaway2021-did), because missingness and unit-level reweighting can alter the bias-decomposition identity linking overall and block biases. Extending the framework to conditional-PT settings with covariates—and to designs with missing data—is an important direction for future research.