EconBase
← Back to paper

Design-Robust Event-Study Estimation under Staggered Adoption Diagnostics, Sensitivity, and Orthogonalisation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

168,175 characters · 47 sections · 94 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Design-Robust Event-Study Estimation under Staggered Adoption: Diagnostics, Sensitivity, and Orthogonalisation

abstractThis paper develops a design-first econometric framework for event-study and difference-in-differences estimands under staggered adoption with heterogeneous effects, emphasising (i) exact probability limits for conventional two-way fixed effects event-study regressions, (ii) computable design diagnostics that quantify contamination and negative-weight risk, and (iii) sensitivity-robust inference that remains uniformly valid under restricted violations of parallel trends. The approach is accompanied by orthogonal score constructions that reduce bias from high-dimensional nuisance estimation when conditioning on covariates. Theoretical results and Monte Carlo experiments jointly deliver a self-contained methodology paper suitable for finance and econometrics applications where timing variation is intrinsic to policy, regulation, and market-structure changes.

\noindentKeywords: event study; staggered adoption; heterogeneous treatment effects; sensitivity analysis; partial identification; orthogonal scores.\\ JEL codes: C21; C23; C51; G10.

Introduction

Event-study practice in empirical finance and economics frequently relies on two-way fixed effects regressions with relative-time indicators. Under staggered adoption and heterogeneous dynamic effects, such regressions can converge to weighted averages that mix effects from multiple relative times and cohorts, generating apparent pre-trends and distorted dynamics even when identifying assumptions hold. This paper develops a design-first treatment of the event-study object, deriving exact decompositions of conventional estimators, proposing computable diagnostics that quantify design risk, and providing sensitivity-robust inference under restricted violations of parallel trends in the spirit of modern partially identified inference.

The methodological spine is grounded in recent advances in heterogeneous-treatment DiD and event studies SunAbraham2021,CallawaySantAnna2021,DeChaisemartinDHaultfoeuille2023 and in robust inference under deviations from parallel trends RambachanRoth2023. Orthogonal score and Riesz representer machinery is incorporated to accommodate high-dimensional covariates without compromising inference ChernozhukovNeweySingh2022. A complementary falsification-adaptive viewpoint on assumption relaxation motivates diagnostic reporting and transparent sensitivity regions MastenPoirier2021. Comprehensive review of staggered DiD advances Tominaga2024.

Recent syntheses organise the modern difference-in-differences and event-study literature around three core fault lines that are directly relevant here: the failure of conventional two-way fixed effects designs under heterogeneous dynamic effects, the distinction between design-based estimands and regression-implied mixtures, and the emergence of diagnostics and sensitivity analysis as first-class reporting objects rather than appendices to point estimates. The taxonomy in RothSantAnnaBilinskiPoe2023 is useful as an organising map for the contribution structure in this paper because it situates contamination, negative-weight risk, and robustness to restricted violations of parallel trends as a unified programme: separating what is identified by design, what is produced by projection geometry, and what remains credibly learnable after disciplined relaxations of the maintained counterfactual trend restrictions.

Framework and estimands

Panel structure, timing, and the design measure space

We formalize the staggering of treatment not as a regression specification, but as a restriction on the permissible paths of treatment history. Let $(\Omega,\mathcal{F},\mathbb{P})$ be a probability space supporting a triangular array \[ \left\{\big(Y_{it},X_{it},G_i\big): i=1,\dots,n,\ t=1,\dots,T\right\}, \] where $T$ is fixed (or slowly growing) relative to $n$ and the cross-sectional index $i$ is the asymptotic dimension. The realised panel is an element of the product space \[ \mathcal{W}_{n,T} \equiv \left(\mathbb{R}\times\mathbb{R}^{d_x}\times\{1,\dots,T,\infty\}\right)^{n\times T}, \] equipped with the product $\sigma$-field. The adoption time $G_i\in\mathcal{G}\equiv\{1,\dots,T,\infty\}$ is a latent characteristic of unit $i$ that generates treatment status through a known policy map.

Define the treatment indicator process $D_{it}\in\{0,1\}$ as the image of $(G_i,t)$ under a design rule $\pi:\mathcal{G}\times\{1,\dots,T\}\to\{0,1\}$. The canonical absorbing adoption design is

equation[equation omitted — 109 chars of source]

so that units are untreated up to and including $t<G_i$ and treated for all $t\ge G_i$ when $G_i<\infty$, while never-treated units satisfy $G_i=\infty$ and $D_{it}=0$ for all $t$. This formulation makes explicit that the empirical content of staggered adoption is not merely “timing variation” but a restriction on admissible treatment paths: the treatment trajectory lies in the monotone class \[ \mathcal{D}^{\uparrow}\equiv \left\{(d_1,\dots,d_T)\in\{0,1\}^T:\ d_{t}\le d_{t+1}\ \forall t\right\}, \] with $G_i$ acting as a sufficient index for the path in (ref).

The distribution of adoption times induces cohort shares \[ p_g \equiv \mathbb{P}(G_i=g),\qquad g\in\mathcal{G}, \] which define a design measure on $\mathcal{G}$. The empirical design is summarised by the random vector of cohort counts \[ N_g \equiv \sum_{i=1}^n \mathbbm{1}\{G_i=g\},\qquad \sum_{g\in\mathcal{G}} N_g=n, \] and by the induced treatment prevalence at each time $t$, \[ \bar D_t \equiv \frac{1}{n}\sum_{i=1}^n D_{it}=\sum_{g\le t}\frac{N_g}{n}. \] Both $p_g$ and the path $\{\bar D_t\}_{t=1}^T$ are primitive design objects: all weighting pathologies of two-way fixed effects event studies under heterogeneity can be expressed as functionals of these quantities and the chosen event window, independent of outcome realisations SunAbraham2021,DeChaisemartinDHaultfoeuille2023.

To accommodate empirically relevant departures from absorbing adoption, let $\pi$ be left unspecified beyond measurability and define a general treatment regime

equation[equation omitted — 60 chars of source]

where $\xi_i$ is a unit-specific regime type (e.g., temporary adoption, intensity, or reversion) that may be observed or latent. Non-absorbing paths correspond to allowing $\pi(\cdot)$ to map into the full $\{0,1\}^T$ rather than $\mathcal{D}^{\uparrow}$. The econometric implication is that the “comparison set” at time $t$ depends on both cohort composition and path geometry: units can be untreated at $t$ either because they are never-treated, not-yet-treated, or reverted. The proposed diagnostics and estimators are stated for (ref) to isolate the essential identification and weighting mechanics, then extended to (ref) by conditioning on regime types and redefining risk indices on the induced design matrix CallawaySantAnna2021,DeChaisemartinDHaultfoeuille2023.

Potential outcomes, dynamic effects, and economically interpretable aggregation

Let $\{Y_{it}(\bm{d}_{1:T}): \bm{d}_{1:T}\in\{0,1\}^T\}$ denote the collection of potential outcomes indexed by a full treatment path, and maintain the stable-unit, no-interference convention so that $Y_{it}(\bm{d}_{1:T})$ depends only on the own-unit path.\footnote{The path-indexed notation is the primitive object; the two-potential-outcome shorthand is recovered under absorbing adoption and a no-anticipation restriction.} Under absorbing adoption in (ref), write $\bm{d}^{(g)}_{1:T}$ for the path induced by adoption time $g$ and $\bm{0}_{1:T}$ for the never-treated path. Define \[ Y_{it}(g)\equiv Y_{it}\!\big(\bm{d}^{(g)}_{1:T}\big),\qquad Y_{it}(\infty)\equiv Y_{it}\!\big(\bm{0}_{1:T}\big), \] so that the observed outcome admits the selection equation $Y_{it}=Y_{it}(G_i)$.

A sharp dynamic causal object requires an explicit restriction excluding anticipation. Let $\mathcal{F}_{i,t}$ be the information set at time $t$ and impose:

assumption[No anticipation] For any $g<\infty$ and any $t<g$, $Y_{it}(g)=Y_{it}(\infty)$ almost surely.

Assumption (ref) elevates the event-time object from a notational convenience to a meaningful behavioural restriction, separating genuine pre-trends in $Y_{it}(\infty)$ from anticipation effects. Under Assumption (ref), a cohort- and event-time-specific causal effect is

equation[equation omitted — 163 chars of source]

This object is the cohort-specific event-study estimand that heterogeneity-robust designs target directly SunAbraham2021,CallawaySantAnna2021. It is useful to index the admissible event times by the finite observation window. Define \[ \mathcal{K}_g \equiv \{-g+1,\dots,T-g\},\qquad \mathcal{K}\equiv \bigcup_{g=1}^{T}\mathcal{K}_g, \] so that $k\in\mathcal{K}_g$ is observed for cohort $g$. The set of cohorts observed at event time $k$ is \[ \mathcal{G}(k)\equiv \{g\in\{1,\dots,T\}: 1\le g+k\le T\}. \]

Aggregation is not innocuous: it encodes the welfare or pricing object that the empirical design is intended to recover. A general aggregand is a weighting scheme $\omega_g(k)$ that is measurable with respect to the design $\sigma$-algebra generated by $\{G_i\}_{i=1}^n$---and thus strictly independent of outcomes---satisfying $\omega_g(k)\ge 0$ and $\sum_{g\in\mathcal{G}(k)}\omega_g(k)=1$. This yields the event-time average effect

equation[equation omitted — 95 chars of source]

The choice of $\omega_g(k)$ should be treated as part of the estimand, not as an afterthought. Three canonical choices are:

equation[equation omitted — 270 chars of source]

where $\omega_g^{\text{pop}}$ targets a population average over cohorts observed at $k$, $\omega_g^{\text{sample}}$ is its finite-sample analogue, and $\omega_g^{\text{risk}}$ can encode economically meaningful exposure weights such as market capitalisation or balance-sheet size in finance applications. The methodological point is that a transparent estimator reports $\omega_g(k)$ explicitly and confines the estimand to convex aggregation, in contrast with two-way fixed effects event studies whose implicit weights may be negative and may load on the wrong event times when effects are heterogeneous SunAbraham2021,DeChaisemartinDHaultfoeuille2023.

Finally, dynamic effects admit economically interpretable transforms. Define cumulative effects and long-run effects \[ \Delta_g(k)\equiv \sum_{j=0}^{k}\tau_g(j),\qquad \tau_g(\infty)\equiv \limsup_{k\to\infty}\tau_g(k), \] whenever these objects are well-defined. In finance settings, $\Delta_g(k)$ corresponds to the level effect on an integrated outcome (e.g., cumulative excess returns or accumulated spreads), while $\tau_g(k)$ corresponds to an impulse-response-type object at horizon $k$, providing a direct bridge between event studies and local-projection logic while retaining the staggered-adoption design discipline CallawaySantAnna2021,SunAbraham2021.

A design-first perspective is useful in staggered-adoption panels because identification and estimation hinge on the assignment structure implied by adoption timing, the availability of not-yet-treated comparisons at each date, and the set of feasible contrasts encoded by the sampling and adoption process rather than by modelling choices layered on top. Framing the analysis in terms of the underlying design clarifies which comparisons are licensed by the data-generating assignment, how the risk set of controls evolves over time, and why certain weighting or aggregation choices can inadvertently re-introduce prohibited comparisons when timing variation is substantial. This perspective also helps separate questions of design validity from questions of functional-form convenience, aligning the specification of the panel structure and treatment timing with the estimand definition and the admissible set of comparisons AtheyImbens2022.

Identifying assumptions, operators, and sensitivity-compatible primitives

Identification is stated at the level of counterfactual evolution of the untreated potential outcome. The objective is to recover $\tau_g(k)$ in (ref) and its convex aggregation $\tau(k)$ in (ref). Let $\mathcal{F}_{i,t}$ be the $\sigma$-field generated by observables up to time $t$ and maintain Assumption (ref). Write $\Delta Y_{it}(\infty)\equiv Y_{it}(\infty)-Y_{i,t-1}(\infty)$ for the untreated increment. Let $X_{it}$ denote covariates, possibly time-varying, that are predetermined with respect to contemporaneous shocks in the untreated outcome.

assumption[Predetermined covariates] For all $t$, $X_{it}$ is $\mathcal{F}_{i,t-1}$-measurable and satisfies $\mathbb{E}[\left\lvert \Delta Y_{it}(\infty) \right\rvert\mid X_{it}]<\infty$.

The baseline identifying restriction is a conditional parallel-trends statement in first differences, which is the natural multi-period analogue of classical difference-in-differences when treatment timing varies across cohorts CallawaySantAnna2021.

assumption[Conditional parallel trends] For every period $t\in\{2,\dots,T\}$ and every cohort $g\in\{1,\dots,T\}$, \begin{equation} \mathbb{E}\!\left[\Delta Y_{it}(\infty)\mid G_i=g,X_{it}\right] \;=\; \mathbb{E}\!\left[\Delta Y_{it}(\infty)\mid G_i=\infty,X_{it}\right] \quad a.s. \end{equation}

Assumption (ref) states that, after conditioning on $X_{it}$, the untreated trend for cohort $g$ matches that for the never-treated cohort. The condition is formulated in increments to avoid imposing equality of levels, and it is congenial to the construction of group-time estimators whose identifying variation arises from comparing changes in outcomes across groups over time CallawaySantAnna2021. It is also the appropriate starting point for sensitivity analysis because violations are naturally parameterised as deviations in the evolution of untreated increments rather than as arbitrary level shifts.

Two auxiliary primitives clarify what is, and is not, being assumed.

assumption[Overlap in covariates] For each $t$ and each $g<\infty$, the conditional distribution of $X_{it}$ given $G_i=g$ is absolutely continuous with respect to the conditional distribution of $X_{it}$ given $G_i=\infty$, with Radon--Nikodym derivative bounded above.
assumption[Regularity] The process $\{\Delta Y_{it}(\infty)\}$ has finite second moments conditional on $(G_i,X_{it})$, and cross-sectional dependence is weak enough that a law of large numbers and central limit theorem apply to the relevant influence functions.

Assumptions (ref)--(ref) are not decorative: overlap prevents identification from relying on off-support extrapolation, while regularity pins down the inferential regime used later for both point-identified and sensitivity-robust procedures. Under Assumptions (ref), (ref), and (ref), the cohort-time effect $\tau_g(k)$ is identified by contrasting observed changes for cohort $g$ with changes for the never-treated group (or, in alternative constructions, with not-yet-treated cohorts) while conditioning on $X_{it}$; this yields the multi-period cohort-time estimands emphasised in the heterogeneous-treatment event-study literature CallawaySantAnna2021,SunAbraham2021.

The methodological programme here does not end at point identification. Departures from Assumption (ref) are structured in a manner that preserves interpretability and enables uniform inference. Define the deviation (or “drift wedge”) for cohort $g$ at time $t$ as

equation[equation omitted — 197 chars of source]

Assumption (ref) is the sharp restriction $\delta_{g,t}(x)=0$ for all $(g,t,x)$. Sensitivity analysis replaces this knife-edge equality by membership of $\delta$ in a restriction class $\Delta(\mathcal{R})$ that bounds magnitude, smooths time variation, or links post-treatment deviations to empirically observable pre-treatment deviations, thereby generating identified sets for $\tau_g(k)$ and $\tau(k)$ rather than point estimates RambachanRoth2023. This approach parallels the broader econometric logic of salvaging empirical content under falsified baseline restrictions by reporting the smallest relaxations compatible with the observed data and the induced identified regions for the target parameter MastenPoirier2021.

In short, Assumption (ref) provides the baseline target for point identification in the group-time framework CallawaySantAnna2021, while (ref) provides the language needed to quantify and report controlled deviations in a manner that supports credible inference when strict parallel trends is too fragile to be treated as an article of faith RambachanRoth2023,MastenPoirier2021.

Conventional TWFE event-study regressions

Specification

Fix an event window $\mathcal{K}\subset\mathbb{Z}$ and a normalisation (baseline) period $k_0\in\mathcal{K}$, typically a pre-adoption event time. Define the relative-time indicators \[ D^{(k)}_{it}\;\equiv\;\mathbbm{1}\{G_i<\infty\}\mathbbm{1}\{t-G_i=k\}, \qquad k\in\mathcal{K}, \] so that $D^{(k)}_{it}$ selects the cohort-time cells in which unit $i$ is exactly $k$ periods away from its adoption time.

Let $W_{it}$ denote any additional controls included linearly (possibly empty). The canonical two-way fixed effects event-study regression is

equation[equation omitted — 146 chars of source]

where $\alpha_i$ and $\gamma_t$ are unit and time fixed effects and $k_0$ is omitted to impose a location normalisation on the event-time profile.

For subsequent derivations it is convenient to stack observations. Let $y\in\mathbb{R}^{nT}$ be the stacked outcome vector, let $F\in\mathbb{R}^{nT\times n}$ be the unit fixed effect design matrix, let $Tmat\in\mathbb{R}^{nT\times T}$ be the time fixed effect design matrix, and let $D\in\mathbb{R}^{nT\times (|\mathcal{K}|-1)}$ be the matrix whose columns are the stacked $D^{(k)}$ for $k\in\mathcal{K}\setminus\{k_0\}$. Writing $X=[F\;\;Tmat\;\;W]$, the OLS coefficient on event-time indicators can be expressed as the Frisch--Waugh--Lovell projection

equation[equation omitted — 111 chars of source]

where $M_X$ residualises with respect to unit and time fixed effects (and any controls). This representation makes explicit that the estimand is determined by the residualised design geometry, which is the object exploited in the weighting and diagnostic results in the next subsection SunAbraham2021,DeChaisemartinDHaultfoeuille2023.

The projection-driven interpretation of two-way fixed effects estimators under staggered treatment timing has a direct antecedent in the canonical decomposition of difference-in-differences with variation in adoption dates, which isolates how implicit comparison weights can arise from the timing structure rather than from an economically declared estimand. The decomposition in GoodmanBacon2021 provides the historical foundation for the weighting-pathology perspective developed in this section: once treatment effects are heterogeneous and timing varies, TWFE coefficients are best understood as design-dependent weighted averages of underlying group-time comparisons, so interpretability becomes a question of which comparisons receive weight and with what sign, and not merely a question of specification or standard errors.

Probability limit and contamination decomposition

This subsection fixes the estimand of the regression coefficient vector $\hat\beta$ in (ref) when dynamic treatment effects are heterogeneous across cohorts and across event time. The key point is that the event-study coefficient indexed by $k$ is generally not an average of $\{\tau_g(k)\}_g$ but a design-dependent linear functional of the entire collection $\{\tau_g(k')\}_{g,k'}$, with weights that can be negative and that can load on $k'\neq k$ SunAbraham2021,DeChaisemartinDHaultfoeuille2023.

Let $D_k$ denote the stacked regressor corresponding to $D^{(k)}_{it}$ (for $k\neq k_0$), and let $\tilde D_k \equiv M_X D_k$ be the residualised regressor after partialling out fixed effects (and any controls) via $M_X$ in (ref). Then

equation[equation omitted — 204 chars of source]

so $\hat\beta_k$ is mechanically coupled to the residualised correlation structure among the event-time indicators. Under staggered adoption, these correlations are generically non-zero because the event-time dummies are functions of a common adoption-time partition of the panel, and residualisation by two-way fixed effects induces sign changes in $\tilde D_k$ across cohort-time cells.

To express the probability limit in causal objects, define the cohort-time cell set \[ \mathcal{C}\equiv\{(g,t): g\in\{1,\dots,T\},\ t\in\{1,\dots,T\},\ 1\le t\le T\}, \] and the corresponding event time $k(t,g)\equiv t-g$. For each $(g,t)\in\mathcal{C}$ with $g<\infty$ and $t\ge g$, define the cell-level average causal effect \[ \tau(g,t)\equiv \mathbb{E}\!\left[Y_{it}(g)-Y_{it}(\infty)\mid G_i=g\right], \qquad \text{so that }\tau(g,t)=\tau_g(t-g). \] Stacking across cells and using linearity of projection, the regression coefficient has an asymptotic linear representation of the form

equation[equation omitted — 180 chars of source]

where the weights $\omega_{g,t}(k)$ depend only on the residualised design matrix $\tilde D = M_X D$ (and therefore only on $(G_i,t)$ and the chosen window $\mathcal{K}$, up to sampling error) SunAbraham2021,DeChaisemartinDHaultfoeuille2023. Grouping cells by event time $k' = t-g$ yields the implicit weighting scheme

equation[equation omitted — 146 chars of source]

with $w_{g,k'}(k)\equiv \sum_{t:\,t-g=k'} \omega_{g,t}(k)$. In general, $w_{g,k'}(k)$ can take negative values and can place mass on $k'\neq k$, so $\operatorname*{plim}\,\hat\beta_k$ can mix post-treatment effects into nominally pre-treatment coefficients, and can mix effects from other horizons into the coefficient labelled $k$ SunAbraham2021. This produces two distinct pathologies:

definition[Negative-weight mass and cross-horizon contamination] Fix $k\in\mathcal{K}\setminus\{k_0\}$. Define \[ \mathcal{N}(k)\equiv \sum_{g}\sum_{k'} \left\lvert w_{g,k'}(k) \right\rvert\,\mathbbm{1}\{w_{g,k'}(k)<0\}, \qquad \mathcal{C}(k)\equiv \sum_{g}\sum_{k'} \left\lvert w_{g,k'}(k) \right\rvert\,\mathbbm{1}\{k'\neq k\}. \]

$\mathcal{N}(k)$ quantifies the extent to which $\hat\beta_k$ relies on sign-reversing comparisons, while $\mathcal{C}(k)$ quantifies how much of $\hat\beta_k$ is, in probability limit, not an average of the target horizon $\{\tau_g(k)\}_g$ but an average of other horizons $\{\tau_g(k')\}_{k'\neq k}$ SunAbraham2021,DeChaisemartinDHaultfoeuille2023. These objects are computable from the residualised design (hence from the adoption schedule and the event window) and serve as design diagnostics reported alongside any TWFE event-study profile.

proposition[Design-driven contamination] Maintain absorbing adoption in (ref), Assumptions (ref) and (ref), and the TWFE event-study specification (ref) with event window $\mathcal{K}$ and baseline $k_0$. Suppose there exist at least two cohorts $g_1\neq g_2$ with $\mathbb{P}(G_i=g_j)>0$ and at least two calendar times $t_1\neq t_2$ such that the set of treated observations is non-degenerate over time, i.e. $\mathbb{P}(D_{it}=1)\in(0,1)$ for some $t$. If dynamic treatment effects are heterogeneous in the sense that either \[ \tau_{g_1}(k^\star)\neq \tau_{g_2}(k^\star)\ \text{for some }k^\star\in\mathcal{K}, \qquad\text{or}\qquad \tau_{g}(k_1)\neq \tau_{g}(k_2)\ \text{for some }g,\ k_1\neq k_2, \] then there exists at least one $k\in\mathcal{K}\setminus\{k_0\}$ such that the TWFE probability limit admits the representation (ref) with weights satisfying \[ \sum_{g}\sum_{k'\neq k}\left\lvert w_{g,k'}(k) \right\rvert \;>\; 0, \] so $\operatorname*{plim}\,\hat\beta_k$ loads on horizons $k'\neq k$. Moreover, there exists at least one (possibly different) $k\in\mathcal{K}\setminus\{k_0\}$ such that \[ \sum_{g}\sum_{k'} \left\lvert w_{g,k'}(k) \right\rvert\,\mathbbm{1}\{w_{g,k'}(k)<0\} \;>\; 0, \] so the implicit TWFE weighting scheme includes negative weights even though Assumption (ref) holds.
proof[Proof sketch] The Frisch--Waugh--Lovell representation (ref) implies $\hat\beta=(D'M_XD)^{-1}D'M_XY$, where $M_X$ residualises on fixed effects (and controls). Under staggered adoption, the columns of $D$ are deterministic functions of $(G_i,t)$ and are therefore linearly related through the cohort-time partition; after residualisation by two-way fixed effects, the residualised indicators $\tilde D_k=M_XD_k$ generally change sign across cohort-time cells. When treatment effects are heterogeneous across cohorts or horizons, the regression residualisation induces a mismatch between the label $k$ and the set of comparisons used to identify $\beta_k$, yielding a projection of $Y$ onto $\tilde D_k$ that necessarily depends on multiple cohort-by-horizon effects. This produces non-zero weights on $k'\neq k$ unless treatment effects are homogeneous in both cohort and horizon, and sign changes in $\tilde D_k$ generate negative weights in the implied linear functional of the cell-level effects, as established by the weighting decompositions in SunAbraham2021 and the TWFE heterogeneity analyses surveyed in DeChaisemartinDHaultfoeuille2023.

A computable representation

Let $D\in\mathbb{R}^{nT\times (|\mathcal{K}|-1)}$ be the stacked matrix of event-time indicators in (ref) (excluding the baseline $k_0$), and let $X$ collect unit and time fixed effects (and any additional controls). Define the residual-maker \[ M_X \equiv I_{nT}-X(X'X)^{-1}X', \] and the residualised event-time design \[ Z \equiv M_XD. \] Then the event-time coefficient vector satisfies the Frisch--Waugh--Lovell identity

equation[equation omitted — 56 chars of source]

where $Y\in\mathbb{R}^{nT}$ is the stacked outcome. Let \[ P_Z \equiv Z(Z'Z)^{-1}Z' \] denote the orthogonal projector onto the column space of $Z$.

The representation (ref) implies that every coefficient is a linear functional of $Y$:

equation[equation omitted — 149 chars of source]

where $e_k$ is the coordinate vector selecting column $k$ of $Z$ (i.e. the event-time regressor indexed by $k$ in the chosen ordering). The weights $\pi_{it}(k)$ are functions only of $Z$, hence only of the design $(G_i,t)$, the event window $\mathcal{K}$, and the fixed-effect structure (and any controls). In particular, $\pi_{it}(k)$ is computable without using $Y$.

To connect (ref) to the causal decomposition in (ref), write the observed outcome as \[ Y_{it} = Y_{it}(\infty) + \mathbbm{1}\{G_i<\infty\}\mathbbm{1}\{t\ge G_i\}\,\tau_{G_i}(t-G_i) + \varepsilon_{it}, \] where $\varepsilon_{it}$ is mean-zero noise orthogonal to the design. Substituting into (ref) and taking probability limits yields \[ \operatorname*{plim}\,\hat\beta_k = \sum_{i,t}\pi_{it}(k)\,\mathbbm{1}\{G_i<\infty\}\mathbbm{1}\{t\ge G_i\}\,\tau_{G_i}(t-G_i), \] so the implicit weight placed on a cohort-horizon pair $(g,k')$ is obtained by summing $\pi_{it}(k)$ over the corresponding cohort-time cells:

equation[equation omitted — 124 chars of source]

Equivalently, defining cohort-time cells $(g,t)$, one can work with $\omega_{g,t}(k)\equiv \mathbb{E}[\pi_{it}(k)\mid G_i=g]$ and then aggregate by event time $k'=t-g$ to obtain $w_{g,k'}(k)$.

This yields a fully computable diagnostic workflow: given the adoption schedule $\{G_i\}_{i=1}^n$, the event window $\mathcal{K}$, and the fixed-effect structure, form $Z$, compute $(Z'Z)^{-1}$, and recover $\pi(k)$ via (ref). The negative-weight mass $\mathcal{N}(k)$ and cross-horizon contamination $\mathcal{C}(k)$ in Definition (ref) follow directly from (ref) by summing over the implied cohort-horizon weights. In particular, these diagnostics depend only on $(D_{it},G_i,T)$ and the chosen event window, not on the realised outcomes, which is why they can be reported as design properties rather than as sample-specific artefacts SunAbraham2021,DeChaisemartinDHaultfoeuille2023.

Heterogeneity-robust estimation and transparent aggregation

Cohort-time and group-time estimators

The robust event-study object is the cohort-time average treatment effect, indexed by adoption cohort $g$ and calendar time $t$. Under absorbing adoption and no anticipation, define the group-time estimand

equation[equation omitted — 140 chars of source]

so that $\mathrm{GATT}(g,t)=\tau_g(t-g)$ in the event-time notation of (ref). Identification proceeds by constructing, for each $(g,t)$, a comparison between the observed change in outcomes for cohort $g$ and the contemporaneous change for an admissible control group that is untreated at $t$ and that shares the same untreated trend conditional on $X_{it}$ CallawaySantAnna2021,SunAbraham2021.

Let $\mathcal{C}(g,t)$ be a control set for cohort $g$ at time $t$. Two canonical choices are: (i) the never-treated group, $\mathcal{C}(g,t)=\{i:G_i=\infty\}$; and (ii) the not-yet-treated group, $\mathcal{C}(g,t)=\{i:G_i>t\}$, which enlarges the control pool at the cost of additional restrictions when treatment effects are dynamic CallawaySantAnna2021. Write $C_{it}(g,t)\equiv \mathbbm{1}\{i\in\mathcal{C}(g,t)\}$.

Define propensity-style reweighting functions to balance covariates between the treated cohort and the control set at time $t$. Let \[ p_{g,t}(x)\equiv \mathbb{P}(G_i=g\mid X_{it}=x,\, C_{it}(g,t)+\mathbbm{1}\{G_i=g\}=1), \] and define weights

equation[equation omitted — 211 chars of source]

Then a generic group-time estimating equation is

equation[equation omitted — 172 chars of source]

which is a weighted difference in outcomes between $t$ and a pre-period reference (often $g-1$) that differences out cohort-specific time-invariant components and isolates the post-adoption deviation in the treated cohort relative to controls CallawaySantAnna2021. Estimators replace expectations by sample averages and estimate $p_{g,t}(\cdot)$ by logit/probit or by flexible methods, provided the weights satisfy overlap and moment conditions.

The event-time estimand is obtained by mapping $(g,t)$ to $(g,k)$ with $k=t-g$ and aggregating across cohorts with transparent non-negative weights. Define the cohort-event-time estimator \[ \widehat{\tau}_g(k)\equiv \widehat{\mathrm{GATT}}(g,g+k), \qquad k\in\mathcal{K}_g, \] and for any chosen aggregation scheme $\{\omega_g(k)\}_{g\in\mathcal{G}(k)}$ satisfying $\omega_g(k)\ge 0$ and $\sum_{g\in\mathcal{G}(k)}\omega_g(k)=1$, define the robust event-time estimator

equation[equation omitted — 126 chars of source]

This construction makes explicit that (i) estimation targets $\mathrm{GATT}(g,t)$ cell-by-cell, and (ii) any averaging across cohorts is a declared choice of estimand, rather than an implicit by-product of regression geometry as in (ref) SunAbraham2021,CallawaySantAnna2021.

A closely related regression-based alternative is the two-stage DiD procedure, which replaces the direct TWFE event-time projection with a first stage that partials out unit and time effects and a second stage that regresses the residualised outcome on the treatment-path indicators of interest. This reframing clarifies which comparisons are being used and can mitigate some of the most opaque weighting behaviour of single-step TWFE in staggered settings. The cohort--time approach here differs in what is made explicit: estimation targets $\mathrm{GATT}(g,t)$ cell-by-cell and then aggregates with declared convex weights $\omega_g(k)$, which pins down the estimand as an interpretable functional of cohort--horizon effects rather than a regression-implied mixture; this transparency is exactly what enables design diagnostics and sensitivity regions to be reported as first-class objects alongside the event-time profile Gardner2021,CallawaySantAnna2021,SunAbraham2021.

Event-time aggregation

Event-time aggregation is treated as part of the estimand. For each event time $k\in\mathcal{K}$, the set of cohorts observed at $k$ is \[ \mathcal{G}(k)=\{g\in\{1,\dots,T\}:1\le g+k\le T\}. \] Let $\widehat{\tau}_g(k)\equiv \widehat{\mathrm{GATT}}(g,g+k)$ be the cohort-specific event-time estimate constructed in Section 4.1. An event-time estimator is defined by

equation[equation omitted — 189 chars of source]

The non-negativity and simplex constraints ensure that $\hat\tau(k)$ is a convex average of well-defined cohort-horizon objects; this rules out the negative and cross-horizon weighting pathologies of (ref) by construction SunAbraham2021,DeChaisemartinDHaultfoeuille2023.

Two aggregation schemes are canonical. The sample-share weights target the average effect among units observed at event time $k$:

equation[equation omitted — 164 chars of source]

Alternatively, population-share weights target the corresponding population average and replace $N_g/n$ by an estimate of $\mathbb{P}(G_i=g)$; in large $n$ panels the two coincide. When an economically motivated target is required, exposure weights can be used, for example

equation[equation omitted — 181 chars of source]

where $w_i\ge 0$ encodes a pre-specified exposure such as size or baseline risk. In all cases, the weighting scheme is explicit, interpretable, and verifiable.

The aggregation in (ref) eliminates cross-$k$ contamination in a precise sense: for each fixed $k$, $\hat\tau(k)$ depends only on estimators of $\tau_g(k)$ for cohorts observed at $k$, and does not mechanically load on $\tau_g(k')$ for $k'\neq k$. This property is the essential methodological contrast with TWFE event-study regressions, in which the coefficient labelled by $k$ is generally a design-dependent mixture over horizons and cohorts SunAbraham2021,DeChaisemartinDHaultfoeuille2023.

Comparison with imputation estimators

Imputation estimators constitute the main competing solution to the staggered-adoption event-study problem once heterogeneous dynamic effects are admitted and the objective is an interpretable event-time profile rather than the projection-defined object delivered by TWFE regressions. The imputation approach constructs the counterfactual untreated path for treated units by estimating an untreated-outcome model on observations that are untreated at each calendar time, then imputes $Y_{it}(\infty)$ for treated observations and averages the treated-minus-imputed residuals by event time. This replaces identification through the residualised correlation structure of overlapping relative-time indicators with identification through a counterfactual mapping for untreated outcomes, and it thereby removes the mechanical cross-horizon mixing that drives negative weights and contamination in TWFE event studies under heterogeneity.

A complementary approach is synthetic difference-in-differences, which combines outcome-modeling and weighting ideas by constructing synthetic controls within a DiD structure to improve robustness when treated and control units differ in levels and evolve differently over time. This method is closely aligned with the design-first emphasis on making the comparison structure explicit: it targets a transparent reweighting of control units so that the counterfactual path is a documented function of observed pre-treatment fit, rather than an implicit by-product of two-way fixed effects projection geometry. In the present framework, the connection is conceptual rather than notational: synthetic DiD can be read as another way to force the analyst to show the weights and the induced counterfactual trajectory, paralleling the role played here by explicit cohort--time estimation and declared convex aggregation in (ref)--(ref) ArkhangelskyAtheyHirshbergImbensWager2021.

To state the object precisely, fix an event time $k$ and define the treated set at that horizon as \[ \mathcal{T}(k)\equiv \{(i,t): G_i<\infty,\ t-G_i=k\}, \] and let $\widehat{Y}_{it}(\infty)$ denote an imputed untreated outcome generated by a first-stage model estimated using untreated observations. The imputation event-study estimand is operationally the event-time average of imputed residuals,

equation[equation omitted — 168 chars of source]

with the understanding that the first-stage estimation uses a training sample restricted to untreated observations at each $t$, typically \[ \{(i,t): D_{it}=0\}, \] possibly augmented by additional restrictions that exclude not-yet-treated observations close to adoption to enforce no-anticipation in the training set. The attraction of (ref) is interpretability: conditional on a valid imputation of the untreated counterfactual, the average residual at event time $k$ is directly an average effect among treated observations at that horizon, with no regression-geometry channel by which effects at other horizons mechanically enter.

The group--time approach adopted in this paper differs at the level of primitives and therefore at the level of failure modes. Group--time estimation begins from the cohort--time estimand $\mathrm{GATT}(g,t)$ and constructs $\tau(k)$ as an explicit convex aggregation over cohorts observed at horizon $k$, as in (ref). Identification is expressed through conditional parallel trends for untreated increments (Section (ref)) and overlap (Assumption (ref)) relative to a declared control set. Estimation is then performed cell-by-cell, so that the only averaging step is the declared convex aggregation with weights $\omega_g(k)$. This separation of (i) identification at the cohort--time level and (ii) aggregation at the event-time level is the mechanism that prevents cross-horizon contamination by construction. In contrast, the imputation approach collapses the problem into a single counterfactual modelling step: once $\widehat{Y}_{it}(\infty)$ is produced, aggregation by event time is mechanically straightforward.

The first-stage model that generates $\widehat{Y}_{it}(\infty)$ can be written generically as a regression for untreated outcomes. A common specification class is an interactive or additive structure with unit and time components and, optionally, covariates:

equation[equation omitted — 132 chars of source]

where $r(\cdot)$ can be high-dimensional or nonparametric and may be regularised, and where variants replace $(a_i,b_t)$ with more flexible low-rank or factor structures. The imputed counterfactual is then \[ \widehat{Y}_{it}(\infty)=\hat a_i+\hat b_t+\hat r(X_{it}), \] and the event-study effect is produced by (ref). This makes explicit the locus of identifying content: for treated units in post-treatment periods, the imputation method relies on extrapolating the untreated-outcome model from the untreated sample into treated states. When adoption is confounded by covariates or by latent types that also affect trends, this extrapolation can be fragile without strong overlap and without strong restrictions on the stability of the untreated-outcome model across the adoption boundary.

The comparison can be sharpened by mapping both approaches into the same design-first language. Under absorbing adoption and no anticipation, the cohort--time effect satisfies \[ \mathrm{GATT}(g,t)=\mathbb{E}[Y_{it}\mid G_i=g]-\mathbb{E}[Y_{it}(\infty)\mid G_i=g]. \] Group--time estimators identify $\mathbb{E}[Y_{it}(\infty)\mid G_i=g]$ through observed controls and a parallel-trends restriction stated for untreated increments, with explicit control sets and weighting. Imputation estimators identify the same object by supplying $\widehat{Y}_{it}(\infty)$ as a model-based proxy for $Y_{it}(\infty)$ and then averaging residuals. The key difference is therefore not the target but the instrument by which the counterfactual is obtained: explicit comparison sets and weighting versus an estimated counterfactual outcome process.

These distinctions imply complementary fragilities. Imputation removes TWFE regression-geometry contamination but introduces a modelling channel: the resulting estimator can be biased if the first-stage nuisance model in (ref) is misspecified or regularisation-biased in a way that does not vanish at the $n^{-1/2}$ rate relevant for inference on (ref). This concern is most acute when the adoption process depends on rich observables or latent states that also affect outcome trends, because the untreated sample used to estimate (ref) may differ systematically from the treated cohort in post-treatment periods, requiring stable extrapolation. The group--time approach instead places the burden on overlap and on the plausibility of conditional parallel trends within declared comparison sets, which is why Section (ref) explicitly parameterises deviations through $\delta_{g,t}(x)$ and why Section (ref) reports sensitivity regions under calibrated restriction classes. In this manuscript, the inferential workflow is organised around that explicit restriction language: identification fragility is surfaced and quantified directly, rather than absorbed implicitly into a counterfactual model.

The two approaches are not mutually exclusive, and the orthogonalisation programme in Section 7 can be interpreted as a bridge. When covariates are high-dimensional, imputation methods and group--time weighting methods both require estimation of nuisance functions whose errors can enter the target estimand at first order. Section 7 constructs Neyman-orthogonal moments for $\mathrm{GATT}(g,t)$ using the density-ratio Riesz representer, yielding debiased estimation under flexible nuisance learning. An analogous orthogonalisation logic can be applied to imputation estimators by treating the imputed counterfactual as a nuisance component and constructing scores whose Gateaux derivative with respect to that nuisance vanishes at the truth. The practical distinction retained in this paper is therefore one of reporting discipline: group--time estimation delivers a transparent convex-aggregation estimand, Section 5 delivers design-geometry diagnostics for the regression-based competitor, and Section 6 delivers falsification-adaptive sensitivity regions that remain meaningful precisely because the estimand is stated as an explicit functional of cohort--time causal objects rather than as a projection-defined regression coefficient.

A further regression-based competitor is the two-way Mundlak framework, which targets robustness to correlated heterogeneity by augmenting the regression with unit-specific covariate means (the Mundlak device) alongside time effects and a model appropriate to the outcome support, including Poisson pseudo-maximum-likelihood for nonnegative outcomes. The identifying content is shifted from fixed-effect residualisation of event-time indicators to an explicit conditional-mean model that allows the unobserved unit component to be correlated with covariates through its projection on covariate averages, thereby absorbing an important class of time-invariant selection and mitigating bias from heterogeneous levels that contaminate simpler two-way specifications. In staggered-adoption settings, this approach can be combined with event-time or group-time indicators, but its interpretation remains model-based: dynamic causal effects are recovered through functional-form restrictions on the conditional mean rather than through convex aggregation of cohort--time comparisons, and the quality of inference depends on whether the Mundlak-augmented specification adequately captures the relevant heterogeneity and trend structure in the untreated potential outcome process Wooldridge2021.

Design diagnostics

The case for reporting design diagnostics is especially acute in empirical finance, where staggered designs are common and where interpreted dynamic profiles frequently drive economic narratives. In that context, BakerLarckerWang2022 documents that staggered difference-in-differences regressions can generate materially misleading estimates in empirically standard settings, precisely because the regression object can behave as a non-convex mixture of heterogeneous effects with sign reversals and implicit comparisons that do not correspond to the intended timing interpretation. The diagnostics in this section operationalise that warning by turning the adoption schedule and event window into computable indices of negative-weight mass and cross-horizon loading, so that the credibility of a conventional event-study regression can be assessed as a property of design geometry before any economic interpretation is attached to the estimated profile WingEtAl2024.

Negative-weight risk index

Fix $k\in\mathcal{K}\setminus\{k_0\}$ and consider the TWFE probability-limit decomposition in (ref). The indices below are defined as functionals of the implied weight array $\{w_{g,k'}(k)\}_{g,k'}$ and are therefore design objects: they depend only on the adoption pattern, the event window, and the fixed-effect structure, not on the realised outcomes AbadieAngristFrandsen2025.

Define the signed and absolute weight masses \[ S(k)\equiv \sum_{g}\sum_{k'} w_{g,k'}(k)=1, \qquad A(k)\equiv \sum_{g}\sum_{k'} \left\lvert w_{g,k'}(k) \right\rvert\ \ge\ 1, \] where $A(k)=1$ holds if and only if all weights are non-negative. The excess absolute mass $A(k)-1$ is a direct measure of how far the TWFE coefficient departs from a convex average.

definition[Negative-weight mass] \[ \mathcal{N}(k)\equiv \sum_{g}\sum_{k'} \left\lvert w_{g,k'}(k) \right\rvert\,\mathbbm{1}\{w_{g,k'}(k)<0\}. \]

$\mathcal{N}(k)$ is zero if and only if the implicit weighting scheme is everywhere non-negative. Moreover, because $S(k)=1$, \[ A(k)=1+2\sum_{g,k'}\big(-w_{g,k'}(k)\big)\mathbbm{1}\{w_{g,k'}(k)<0\}, \] so $\mathcal{N}(k)$ is tightly linked to $A(k)$ and provides a scale-free measure of sign-reversing comparisons.

The second pathology is horizon contamination: the coefficient labelled by $k$ can load on causal effects at horizons $k'\neq k$.

definition[Cross-horizon contamination] \[ \mathcal{C}(k)\equiv \sum_{g}\sum_{k'} \left\lvert w_{g,k'}(k) \right\rvert\,\mathbbm{1}\{k'\neq k\}. \]

$\mathcal{C}(k)=0$ if and only if $\operatorname*{plim}\,\hat\beta_k$ is a weighted average of $\{\tau_g(k)\}_{g\in\mathcal{G}(k)}$ alone, with no mechanical loading on other horizons.

Both indices are computable without outcome data. Let $Z=M_XD$ be the residualised event-time design matrix in (ref). For each $k$, define the coefficient-weight vector \[ \pi(k)\equiv Z(Z'Z)^{-1}e_k, \] so that $\hat\beta_k=\pi(k)'Y$ by (ref). The implied cohort-horizon weights satisfy (ref) and are therefore functions only of $Z$ (hence only of $(G_i,t)$, the chosen $\mathcal{K}$, and the fixed-effect structure). Consequently, $\mathcal{N}(k)$ and $\mathcal{C}(k)$ are design diagnostics that can be reported ex ante as properties of the empirical design, and they quantify precisely the two mechanisms emphasised in the heterogeneous-effects critique of TWFE event studies SunAbraham2021,DeChaisemartinDHaultfoeuille2023.

Design geometry under staggered adoption

The weighting pathologies in (ref) are design-driven: they arise from the cohort--time partition induced by the adoption schedule and from the projection of event-time indicators onto the orthogonal complement of unit and time fixed effects. Figure (ref) provides a compact representation of this geometry. Each horizontal line corresponds to an adoption cohort $g$, the filled point marks the adoption date $G=g$, and the dashed ray to the right indicates the post-adoption region $(t\ge g)$ in which treatment is “on” under absorbing adoption. Every observed pair $(g,t)$ determines an event time $k=t-g$; the event-study regressors $D^{(k)}_{it}$ select diagonals in this cohort--time diagram. The residualisation implicit in (ref) transforms these diagonal selectors into residualised regressors $\tilde D_k=M_XD_k$ that generally change sign across cells, because unit and time fixed effects remove cohort means and time means. Those sign changes are the mechanical source of negative weights in the TWFE decomposition, while the non-orthogonality of the residualised diagonals is the mechanical source of cross-horizon contamination, both of which are summarised by $\mathcal{N}(k)$ and $\mathcal{C}(k)$.

figure[figure omitted — 1,465 chars of source]

Robust inference under restricted violations of parallel trends

A restriction class

Sensitivity analysis replaces the knife-edge restriction $\delta_{g,t}(x)\equiv 0$ in (ref) by a structured class of admissible deviations that is disciplined by pre-treatment behaviour and yields computable identified sets RambachanRoth2023. For notational economy, suppress covariates in this subsection and write the cohort-specific untreated increment as $\Delta Y_{it}(\infty)$ and its cohort-time mean as \[ m_{g,t}\equiv \mathbb{E}\!\left[\Delta Y_{it}(\infty)\mid G_i=g\right], \qquad m_{\infty,t}\equiv \mathbb{E}\!\left[\Delta Y_{it}(\infty)\mid G_i=\infty\right]. \] Define the deviation (trend wedge)

equation[equation omitted — 72 chars of source]

Assumption (ref) is the restriction $\delta_{g,t}=0$ for all $(g,t)$, which point-identifies the cohort-time effects under no anticipation.

A restriction class $\Delta(\mathcal{R})$ is a set of sequences $\{\delta_{g,t}\}_{g,t}$ consistent with a maintained regularity structure $\mathcal{R}$ and with the information contained in pre-treatment periods. The robust parallel-trends approach of RambachanRoth2023 can be implemented by imposing that post-treatment deviations are bounded by a function of pre-treatment deviations plus a curvature (smoothness) penalty. Concretely, fix a cohort $g$ and let $\mathcal{T}_g^{-}\equiv\{t: t<g\}$ be the pre-treatment times. Let $\Delta \delta_{g,t}\equiv \delta_{g,t}-\delta_{g,t-1}$ be the first difference and $\Delta^2 \delta_{g,t}\equiv \Delta\delta_{g,t}-\Delta\delta_{g,t-1}$ be the second difference. For constants $B\ge 0$ and $\Gamma\ge 0$, define the curvature-bounded class

equation[equation omitted — 259 chars of source]

The first restriction bounds the change in the deviation slope (a discrete curvature constraint) and limits how abruptly post-treatment trends can diverge, while the second uses pre-treatment periods to cap the magnitude of deviation levels. Alternative specifications replace $\max_{t<g}\left\lvert \delta_{g,t} \right\rvert\le B$ by data-driven calibration based on observed pre-trend estimates and their sampling error, which is the operational choice in robust event-study inference RambachanRoth2023.

Partial identification provides a clean way to formalise sensitivity analysis when point identification is not credible under strict assumptions, because the object of interest becomes a set of probability distributions and a corresponding identified set for the estimand, indexed by explicit, interpretable restrictions on departures from the maintained design assumptions. This framing clarifies that robustness exercises are not an afterthought but a disciplined shift from point statements to set statements, with the width of the identified set measuring the informational content delivered by the data under the chosen restriction class, and it situates restricted-violation frameworks within a broader econometric programme that treats assumptions as tunable constraints rather than as binary truths Manski2003, Roth2024.

More generally, write \[ \delta \in \Delta(\mathcal{R}), \] where $\mathcal{R}$ encodes one or more of: (i) level bounds anchored by pre-period estimates, (ii) slope bounds linking post-period drifts to pre-period drifts, and (iii) curvature or smoothness bounds limiting changes in drift. Each choice of $\mathcal{R}$ induces an identified set for $\tau(k)$ obtained by optimising the target estimand over $\delta\in\Delta(\mathcal{R})$, delivering sensitivity regions that are explicit functions of the maintained restriction rather than informal robustness claims RambachanRoth2023,MastenPoirier2021.

A further sensitivity channel concerns functional-form dependence: even when a parallel-trends restriction is stated in a conditional-mean or first-difference form, empirical assessments and pre-tests can be materially affected by how the counterfactual trend is parameterised, because distinct functional-form choices can imply distinct extrapolations from pre-treatment fit into post-treatment counterfactual evolution. This matters directly for restriction classes such as (ref), since calibration of $(B,\Gamma)$ from pre-period behaviour is only meaningful relative to a maintained mapping from observed pre-period deviations to admissible post-period paths, and functional-form choices can tighten or loosen that mapping in ways that are not innocuous. The practical implication is that restricted-violations inference should be interpreted as conditional on the maintained trend parameterisation used to construct the deviation process and its calibration, and robustness checks should treat functional form as part of the sensitivity analysis rather than as a separate specification decision RothSantAnna2023.

Identified sets and confidence regions

Fix an event time $k\in\mathcal{K}$ and let the target be the convexly aggregated effect $\theta\equiv\tau(k)$ in (ref). Under a restriction class $\delta\in\Delta(\mathcal{R})$ (for example (ref)), the target is generally no longer point identified. The restriction class induces an identified set

equation[equation omitted — 133 chars of source]

where $\theta(\delta)$ is the value of the estimand implied by a given admissible deviation process $\delta$. Operationally, $\Theta(\mathcal{R})$ is obtained by solving two optimisation problems that deliver sharp lower and upper bounds:

equation[equation omitted — 309 chars of source]

The defining feature of the robust parallel-trends approach is that $\Delta(\mathcal{R})$ is calibrated from pre-treatment behaviour and regularity restrictions, so the identified set expands smoothly as $\mathcal{R}$ relaxes, replacing binary acceptance or rejection of parallel trends by a quantitative description of what the data imply under controlled deviations RambachanRoth2023.

As illustrated in Figure (ref), the restriction class $\Delta(\mathcal{R})$ defines a geometry of admissible deviation paths in event time. The parameter $B$ imposes a rectangular bound on pre-treatment deviations, quantifying tolerance for pre-trend differences. In the post-treatment period, $\Gamma$ governs the rate at which the set of admissible biases expands. A higher $\Gamma$ allows the counterfactual trend to change curvature more abruptly, resulting in a wider `cone' of uncertainty and a larger identified set $\Theta(\mathcal{R})$.

Inference targets confidence regions that cover $\Theta(\mathcal{R})$ uniformly over the maintained restriction class. Let $\widehat{\Theta}(\mathcal{R})=[\underline{\hat\theta}(\mathcal{R}),\overline{\hat\theta}(\mathcal{R})]$ be the sample analogue of (ref), computed by substituting estimated reduced-form components and solving the corresponding sample optimisation problems. A $(1-\alpha)$ confidence region $\mathcal{C}_n(\mathcal{R})$ is said to be uniformly valid if it contains the true identified set with probability at least $1-\alpha$ uniformly over all data-generating processes satisfying $\mathcal{R}$. As illustrated in Appendix (ref) and Figure (ref), the restriction class $\Delta(\mathcal{R})$ defines a geometry of admissible deviation paths in event time, with pre-treatment tolerance governed by $B$ and post-treatment smoothness governed by $\Gamma$.

theorem[Uniformly valid confidence regions] Maintain Assumptions (ref) and (ref), and suppose the deviation process satisfies $\delta\in\Delta(\mathcal{R})$ where $\Delta(\mathcal{R})$ obeys the regularity and calibration structure in RambachanRoth2023. Then there exist confidence regions $\mathcal{C}_n(\mathcal{R})$ constructed from $\widehat{\Theta}(\mathcal{R})$ and an appropriate critical value sequence such that, for any $\alpha\in(0,1)$, \[ \liminf_{n\to\infty}\ \inf_{P\in\mathcal{P}(\mathcal{R})} \mathbb{P}_P\!\left(\Theta_P(\mathcal{R})\subseteq \mathcal{C}_n(\mathcal{R})\right)\ \ge\ 1-\alpha, \] where $\mathcal{P}(\mathcal{R})$ denotes the class of data-generating processes satisfying the maintained restrictions, and $\Theta_P(\mathcal{R})$ is the identified set for $\theta$ under $P$.

Calibration

\paragraph{Goal.} Calibrate sensitivity inputs $(\Delta(\mathcal{R}),B,\Gamma)$ using observables so that each restriction class $\mathcal{R}$ admits an interpretable and reproducible mapping from data diagnostics to parameter values.

\paragraph{Objects.} Let $\widehat{\beta}_{\ell}$ denote pre-period event-time coefficients for $\ell\in\mathcal{L}_{\mathrm{pre}}$ from the baseline event-study regression, with standard errors $\widehat{\sigma}_{\ell}$. Define the pre-trend magnitude statistic

equation[equation omitted — 294 chars of source]

Let $\widehat{\tau}$ denote the baseline target estimand and $\widehat{\mathcal{D}}$ denote a drift diagnostic (defined below).

Calibration rules for \texorpdfstring{$\Delta(\mathcal{R})$}{Delta(R)}

Let $\Delta(\mathcal{R})$ parameterise the radius of admissible violations within restriction class $\mathcal{R}$. For each $\mathcal{R}$, define a measurable diagnostic map

equation[equation omitted — 253 chars of source]

with $\phi_{\mathcal{R}}$ chosen so that $\widehat{\Delta}(\mathcal{R})$ is nondecreasing in $M_{\mathrm{pre}}$ and in placebo rejections.

\paragraph{Canonical rule.} For any $\mathcal{R}$ admitting a scalar violation radius, set

equation[equation omitted — 165 chars of source]

where $c_{\mathcal{R}}$ is selected by a holdout or placebo criterion (defined in Box (ref)).

Data-driven calibration of \texorpdfstring{$B$}{B}

Let $B$ denote the scale of admissible confounding or violation magnitude in the moment condition. Define $B$ in outcome units, anchored to observed pre-trend magnitudes:

equation[equation omitted — 126 chars of source]

Alternatively, define $B$ by a formal holdout rule using a pre-period split $\mathcal{L}_{\mathrm{pre}}=\mathcal{L}_{1}\cup\mathcal{L}_{2}$:

equation[equation omitted — 248 chars of source]

where $\widehat{\beta}_{\ell}(b)$ denotes the pre-coefficient under the sensitivity perturbation indexed by $b$.

Monotone family calibration of \texorpdfstring{$\Gamma$}{Gamma}

Let $\Gamma\ge 0$ index a monotone family of violations with a time-unit interpretation as drift per period. Define a drift diagnostic from pre-period differences:

equation[equation omitted — 214 chars of source]

Map $\Gamma$ to percent drift per period relative to the baseline effect scale:

equation[equation omitted — 166 chars of source]

The admissible violation set is then indexed by $\Gamma$ through a monotone constraint, for example

equation[equation omitted — 244 chars of source]

where $\{\Delta_{\ell}\}$ denotes the pathwise deviation sequence admitted under $\mathcal{R}$.

figure[figure omitted — 1,030 chars of source]
table[table omitted — 1,028 chars of source]
center[center omitted — 1,532 chars of source]

Orthogonalisation with covariates via Riesz representers

High-dimensional conditioning variables are standard in modern empirical work and arise naturally once treatment timing is allowed to depend on rich unit characteristics and time-varying states. In such environments, plug-in estimators that substitute flexible estimates of nuisance regressions into low-dimensional causal functionals typically suffer first-order bias because the estimation error in the nuisance component enters the target functional at the same order as sampling uncertainty. The orthogonalisation programme addresses this by replacing direct plug-in with estimating equations whose local (Gateaux) derivative with respect to the nuisance direction vanishes at the truth, rendering the target locally insensitive to small nuisance estimation errors and permitting valid inference after regularised or machine-learning estimation. Riesz representers provide a constructive route to Neyman-orthogonal scores and influence-function representations for broad classes of linear functionals, directly linking the implementation in this section to foundational semiparametric variance and efficiency calculations Newey1994 and to modern regularised Riesz and cross-fitting constructions ChernozhukovNeweySingh2022, SantAnnaZhao2025.

The event-study Riesz representer

This subsection specialises the generic Riesz/orthogonal-score construction to the group--time estimands $\mathrm{GATT}(g,t)$ in (ref). The purpose is not to reintroduce machine learning as an empirical flourish, but to obtain (i) an explicit representer whose population form is a covariate density ratio, and (ii) an orthogonal moment whose first-order behaviour is insensitive to regularised estimation of nuisance components, as in ChernozhukovNeweySingh2022.

\paragraph{Target functional.} Fix a cohort--time cell $(g,t)$ with $t\ge g$. Let the comparison sample be the union of cohort $g$ and the control set $\mathcal{C}(g,t)$ defined in Section 4.1. For concreteness, take the never-treated control set in this subsection, \[ \mathcal{C}(g,t)=\{i:G_i=\infty\}, \qquad S_{i}^{(g)} \equiv \mathbbm{1}\{G_i=g\}+\mathbbm{1}\{G_i=\infty\}. \] Let $Z\equiv X_{it}$ denote the covariates used to impose conditional parallel trends in Assumption (ref), and define the conditional regression functions \[ m_{d}(x)\equiv \mathbb{E}[Y_{it}\mid X_{it}=x,\,G_i\in\{g,\infty\},\,\mathbbm{1}\{G_i=g\}=d], \qquad d\in\{0,1\}, \] where $d=1$ corresponds to cohort $g$ and $d=0$ corresponds to the never-treated control group, within the restricted comparison sample $G_i\in\{g,\infty\}$. The cohort--time effect can be written as a linear functional of the control regression $m_0$ evaluated under the treated covariate distribution, together with observable treated moments. A convenient decomposition is

equation[equation omitted — 228 chars of source]

The second term in (ref) is the problematic component: it is an expectation of an unknown function $m_0(\cdot)$ under the treated covariate law. This is exactly the setting for a Riesz representer.

\paragraph{The Riesz representer as a density ratio.} Work on the Hilbert space $\mathcal{H}\subseteq L_2(\mathcal{Z},\mathbb{P}_{\infty})$ with inner product \[ \left\langle f, h \right\rangle_{\infty}\equiv \mathbb{E}\!\left[f(X_{it})\,h(X_{it})\mid G_i=\infty\right], \] so the reference measure is the control covariate distribution. Define the linear functional \[ \mathcal{L}(h)\equiv \mathbb{E}\!\left[h(X_{it})\mid G_i=g\right], \qquad h\in\mathcal{H}. \] Under overlap (Assumption (ref) specialised to the never-treated control), $\mathcal{L}$ is continuous on $\mathcal{H}$ and admits a unique Riesz representer $\alpha_{g,t}(\cdot)\in\overline{\mathcal{H}}$ satisfying

equation[equation omitted — 199 chars of source]

The population solution to (ref) is the Radon--Nikodym derivative of the treated covariate law with respect to the control covariate law:

equation[equation omitted — 109 chars of source]

where $f_{X\mid G=g,t}$ denotes the conditional density of $X_{it}$ given $G_i=g$ (and time $t$ if $X$ is time-varying) and likewise for $G_i=\infty$. Equivalently, writing \[ \pi_{g,t}(x)\equiv \mathbb{P}(G_i=g\mid X_{it}=x,\,G_i\in\{g,\infty\}), \] one obtains the odds-ratio form

equation[equation omitted — 192 chars of source]

which makes explicit that $\alpha_{g,t}$ is proportional to the inverse odds of being in the control group relative to cohort $g$.

\paragraph{Balancing property and interpretation.} Equation (ref) is a covariate balancing identity: weighting control observations by $\alpha_{g,t}(X_{it})$ transports expectations of any $h(X)$ from the control covariate law to the treated covariate law. When $\mathcal{H}$ is rich (for example, the closure of square-integrable functions), this implies that the reweighted control group matches the treated distribution in all $L_2$ moments. In finite-dimensional approximations, (ref) states that $\alpha_{g,t}$ learns precisely the weights that make the control group “look like” cohort $g$ with respect to the chosen dictionary of covariate functions.

\paragraph{Orthogonal score for $\mathrm{GATT}(g,t)$.} Substituting (ref) into (ref) yields an orthogonal moment condition. Define the nuisance functions \[ m_0(x)\equiv \mathbb{E}[Y_{it}\mid X_{it}=x,\,G_i=\infty], \qquad \alpha_{g,t}(x)\ \text{as in \eqref{eq:alpha-density}}, \] and define the score

equation[equation omitted — 241 chars of source]

with $W_{it}=(Y_{it},X_{it},G_i)$. At the truth, $\theta=\mathrm{GATT}(g,t)$, $\alpha=\alpha_{g,t}$, and $m_0$ as defined above, one has \[ \mathbb{E}\!\left[\psi_{g,t}\big(W_{it};\mathrm{GATT}(g,t),m_0,\alpha_{g,t}\big)\right]=0, \] and the corresponding moment is Neyman-orthogonal with respect to small perturbations in $m_0$ (and, under the regularity conditions in ChernozhukovNeweySingh2022, also with respect to perturbations in $\alpha$ when cross-fitting is used). The algebraic structure of (ref) makes the orthogonality transparent: the term $\mathbbm{1}\{G_i=\infty\}\alpha(X_{it})\{Y_{it}-m_0(X_{it})\}$ is a residual reweighted by the Riesz representer, while the remaining terms ensure the score has mean zero at the target and cancels first-order estimation error in the nuisance regression.

\paragraph{Operational implication.} Regularised Riesz regression estimates $\alpha_{g,t}$ by solving an empirical analogue of (ref) over a high-dimensional dictionary of functions of $X$, typically with a penalty that controls complexity; cross-fitting separates the estimation of $(m_0,\alpha_{g,t})$ from evaluation of (ref). The result is a plug-in estimator of $\mathrm{GATT}(g,t)$ (and therefore of $\tau_g(k)$ and $\tau(k)$ after aggregation) that retains valid first-order behaviour under flexible nuisance learning, matching the methodological template in ChernozhukovNeweySingh2022.

Debiased score construction

This subsection constructs a debiased estimating equation for $\mathrm{GATT}(g,t)$ using the cohort--time Riesz representer from Section (ref). The construction follows the orthogonal-score logic in ChernozhukovNeweySingh2022, specialised to the cohort--time setting.

Fix $(g,t)$ with $t\ge g$ and work in the restricted comparison sample $G_i\in\{g,\infty\}$. Let \[ p_g \equiv \mathbb{P}(G_i=g),\qquad p_\infty \equiv \mathbb{P}(G_i=\infty). \] Define the control outcome regression \[ m_0(x)\equiv \mathbb{E}\!\left[Y_{it}\mid X_{it}=x,\,G_i=\infty\right]. \] Let $\alpha_{g,t}(\cdot)$ denote the (scaled) Riesz representer that transports moments from the never-treated covariate law to the cohort-$g$ covariate law in the unconditional moment form

equation[equation omitted — 220 chars of source]

Under overlap, (ref) is satisfied by

equation[equation omitted — 134 chars of source]

which is equivalent to an odds-ratio representation in terms of $\pi_{g,t}(x)=\mathbb{P}(G_i=g\mid X_{it}=x,\,G_i\in\{g,\infty\})$.

\paragraph{Orthogonal score with an explicit double-robustness property.} Let $\theta\equiv \mathrm{GATT}(g,t)$. Maintain the restricted comparison sample $G_i\in\{g,\infty\}$ and define the control regression \[ m_0(x)\equiv \mathbb{E}\!\left[Y_{it}\mid X_{it}=x,\,G_i=\infty\right]. \] Let $\alpha_{g,t}(\cdot)$ denote the (scaled) Riesz representer that solves the unconditional balancing equation (ref). For generic candidates $(m,\alpha)$, define the estimating score

equation[equation omitted — 318 chars of source]

The labelling in (ref) is purely mnemonic. The substantive claim is a moment robustness statement: the moment condition $\mathbb{E}[\psi_{g,t}(W_{it};\theta,m,\alpha)]=0$ holds at the target $\theta=\mathrm{GATT}(g,t)$ under either of two nuisance correctness conditions, which is the precise meaning of the term double robustness in this setting.

To state the property explicitly, write $\theta_0\equiv \mathrm{GATT}(g,t)=\mathbb{E}[Y_{it}\mid G_i=g]-\mathbb{E}[m_0(X_{it})\mid G_i=g]$.

First, if the outcome regression is correct, $m=m_0$, then for any measurable $\alpha(\cdot)$, \[ \mathbb{E}\!\left[\mathbbm{1}\{G_i=\infty\}\,\alpha(X_{it})\big(Y_{it}-m_0(X_{it})\big)\right]=0 \] by iterated expectations conditional on $(X_{it},G_i=\infty)$, so $\mathbb{E}[\psi_{g,t}(W_{it};\theta_0,m_0,\alpha)]=0$.

Second, if the representer is correct, $\alpha=\alpha_{g,t}$ satisfying (ref), then for any measurable $m(\cdot)$, \[ \mathbb{E}\!\left[\mathbbm{1}\{G_i=\infty\}\,\alpha_{g,t}(X_{it})\,m(X_{it})\right] = \mathbb{E}\!\left[\mathbbm{1}\{G_i=g\}\,m(X_{it})\right], \] and also \[ \mathbb{E}\!\left[\mathbbm{1}\{G_i=\infty\}\,\alpha_{g,t}(X_{it})\,Y_{it}\right] = \mathbb{E}\!\left[\mathbbm{1}\{G_i=\infty\}\,\alpha_{g,t}(X_{it})\,m_0(X_{it})\right] = \mathbb{E}\!\left[\mathbbm{1}\{G_i=g\}\,m_0(X_{it})\right], \] so the $m(\cdot)$ terms cancel and $\mathbb{E}[\psi_{g,t}(W_{it};\theta_0,m,\alpha_{g,t})]=0$.

Hence the label double robustness refers to the exact statement \[ \mathbb{E}[\psi_{g,t}(W_{it};\theta_0,m,\alpha)]=0 \quad\text{if}\quad m=m_0 \ \text{or}\ \alpha=\alpha_{g,t}, \] with $\theta_0$ fixed at the cohort--time effect.

\paragraph{Orthogonality as a separate claim.} Neyman orthogonality is an additional property, not a synonym for double robustness. At $(m,\alpha)=(m_0,\alpha_{g,t})$, the Gateaux derivative of the moment map $\eta\mapsto \mathbb{E}[\psi_{g,t}(W_{it};\theta_0,m_0+\eta h,\alpha_{g,t})]$ vanishes for any direction $h\in\mathcal{H}$ because (ref) implies exact cancellation of the perturbation terms. This separates two concepts that can otherwise be conflated in exposition: double robustness is a correctness-under-one-nuisance property of the population moment in (ref), while orthogonality is the first-order insensitivity that delivers debiased behaviour when both nuisance components are estimated with regularisation and cross-fitting ChernozhukovNeweySingh2022.

\paragraph{Estimating equation with cross-fitting.} Let $\{\mathcal{I}_m\}_{m=1}^M$ be a fold partition. Construct nuisance estimates $\hat m^{(-m)}$ and $\hat\alpha^{(-m)}$ using only observations in $\mathcal{I}_m^{c}$. The debiased estimator $\widehat{\mathrm{GATT}}(g,t)$ solves

equation[equation omitted — 142 chars of source]

where $m(i)$ is the fold containing unit $i$. Under standard complexity control for the nuisance estimators and weak dependence conditions, (ref) yields asymptotically normal inference for $\mathrm{GATT}(g,t)$ while permitting high-dimensional or nonparametric learning of $m_0$ and $\alpha_{g,t}$ ChernozhukovNeweySingh2022.

Sensitivity regions and falsification-adaptive reporting

Section (ref) replaces the sharp restriction $\delta_{g,t}\equiv 0$ by a maintained restriction class $\delta\in\Delta(\mathcal{R})$ and thereby replaces point identification of $\theta=\tau(k)$ by the identified set $\Theta(\mathcal{R})$. The family $\{\Delta(\mathcal{R})\}_{\mathcal{R}}$ is naturally ordered by inclusion: tighter restrictions correspond to smaller admissible deviation sets and therefore smaller identified sets, while weaker restrictions enlarge $\Theta(\mathcal{R})$. Reporting $\Theta(\mathcal{R})$ across calibrated values of $\mathcal{R}$ therefore yields falsification-adaptive reporting in the precise sense of MastenPoirier2021: when the baseline restriction is inconsistent with the observed reduced-form implications, the analysis reports the minimal relaxation of $\mathcal{R}$ that restores non-falsification and the corresponding set of admissible values for the target parameter.

This set-based logic is conceptually distinct from, but complementary to, the design diagnostics in Section 5. Sensitivity regions quantify fragility with respect to counterfactual trend restrictions through admissible deviation paths $\delta$. By contrast, the indices $\mathcal{N}(k)$ and $\mathcal{C}(k)$ quantify fragility with respect to regression geometry. Even when Assumption (ref) holds exactly, TWFE event-study coefficients can be non-convex mixtures of heterogeneous cohort--horizon effects because residualisation by unit and time fixed effects induces sign changes and non-orthogonality among residualised event-time regressors, producing negative weights and cross-horizon loading in (ref) SunAbraham2021,DeChaisemartinDHaultfoeuille2023. Large values of $\mathcal{N}(k)$ and $\mathcal{C}(k)$ therefore constitute design-level evidence that the conventional estimator is not transparently interpretable as an event-time causal object, independent of any violations of parallel trends.

Taken together, the reporting protocol has two disciplined components. The first is a design report, which records $\mathcal{N}(k)$ and $\mathcal{C}(k)$ as ex ante properties of the adoption pattern and the event window. The second is an identification report, which records $\Theta(\mathcal{R})$ and its confidence region under calibrated restriction classes. The combination produces an empirically implementable, assumption-explicit workflow that is falsification-aware both at the level of the maintained identifying restrictions and at the level of estimator interpretability.

Monte Carlo designs

\noindentMonte Carlo replication law. Each replication is defined by a joint law for \\ \(\big(X_i,G_i,\{D_{it}\}_{t=1}^T,\{Y_{it}\}_{t=1}^T\big)\): covariates \(X_i\) determine the conditional adoption-time distribution \(\mathbb{P}(G_i=g\mid X_i)\), the adoption time \(G_i\) determines the treatment path \(D_{it}\), and the outcome process is generated by specifying an untreated potential outcome \(Y_{it}(0)\) together with a dynamic effect profile \(\tau_g(k)\) that maps \((G_i,t)\) into observed outcomes \(Y_{it}\); Subsection (ref) states the complete data-generating process.

The Monte Carlo section is defined by the replication law for \(\big(X_i,G_i,\{D_{it}\}_{t=1}^T,\{Y_{it}\}_{t=1}^T\big)\); covariates and confounded adoption are specified here, and the outcome data-generating process is specified explicitly in Subsection (ref) through equations (ref)--(ref).

\paragraph{Covariates and confounded adoption.} For each unit $i\in\{1,\dots,n\}$ generate a time-invariant covariate vector $X_i\in\mathbb{R}^{d_x}$, for example \[ X_i\sim \mathcal{N}(0,\Sigma_X), \qquad \Sigma_X\succ 0. \] To induce selection on observables, specify the conditional adoption-time distribution on \( \mathcal{G}\equiv\{1,\dots,T,\infty\} \) via a multinomial logit:

equation[equation omitted — 178 chars of source]

where $\{\gamma_g\}_{g\in\mathcal{G}}$ are nuisance coefficients chosen to create economically meaningful imbalance across cohorts (including non-trivial mass at $\infty$ through the $g=\infty$ index). Draw $G_i$ from (ref) and define absorbing treatment as \[ D_{it}=\mathbbm{1}\{G_i\le t\}\mathbbm{1}\{G_i<\infty\}. \] This construction makes the propensity score cohort- and covariate-dependent, so the density-ratio (odds-ratio) Riesz representer $\alpha_{g,t}(X)$ in (ref) is non-degenerate and the orthogonal-score machinery in Section (ref) is genuinely required for debiased estimation. Conditional on the realised adoption schedule $\{G_i\}$, the residualised design matrix $Z=M_XD$ (and hence the diagnostic indices $\mathcal{N}(k)$ and $\mathcal{C}(k)$) is fixed prior to simulating outcomes.

Monte Carlo Design

\paragraph{Panel, timing, and cohorts (fixed across all DGPs).} Units $i\in\{1,\dots,N\}$ with $N=5{,}000$; periods $t\in\{1,\dots,T\}$ with $T=12$. Adoption time $A_i\in\{4,6,8,10,\infty\}$ with cohort set $\mathcal{G}=\{1,2,3,4\}$ and never-treated $g=0$. Cohort shares: $\pi_0=0.20$ and $\pi_g=0.20$ for each $g\in\{1,2,3,4\}$. Define $D_{it}=\mathbf{1}\{t\ge A_i,\,A_i<\infty\}$ and event time $\ell_{it}=t-A_i$.

\paragraph{Outcome equation (common envelope).}

equation[equation omitted — 161 chars of source]

Fixed effects: $\alpha_i\sim\mathcal{N}(0,1)$ and $\lambda_t\sim\mathcal{N}(0,0.25)$, independent. Covariate: $X_{it}=\rho_X X_{i,t-1}+\xi_{it}$ with $\rho_X=0.50$, $X_{i1}\sim\mathcal{N}(0,1)$, $\xi_{it}\sim\mathcal{N}(0,1)$ i.i.d., and $\beta=1$.

\paragraph{Treatment-effect path (fixed functional form; cohort scaling).} Let $h_1=0.80$, $h_2=1.00$, $h_3=1.20$, $h_4=1.40$ and define the baseline event-time profile \[ m(\ell)=

cases0 & \ell<0,\\ 0.50 & \ell=0,\\ 0.75 & \ell=1,\\ 1.00 & \ell\ge 2.

\] Then $\tau_g(\ell)=h_g\,m(\ell)$ for $g\in\{1,2,3,4\}$.

\paragraph{Violation indices and grid (fixed; applied to each DGP).} \[ \Delta(\mathcal{R})\in\{0,\ 0.25,\ 0.50,\ 1.00\},\qquad B\in\{0,\ 0.05,\ 0.10,\ 0.20\},\qquad \Gamma\in\{0,\ 0.005,\ 0.010,\ 0.020\}. \] $B$ is in outcome units (calibrated to pre-trend magnitudes per Section (ref)); $\Gamma$ is drift-per-period (fractional) in a monotone family (Section (ref)). Total cells per DGP: $K_\Delta K_B K_\Gamma = 4\cdot 4\cdot 4 = 64$.

\paragraph{Construction of the violation term (monotone in each index).} Let $s_g\in\{+1,-1\}$ be fixed by cohort with $s_1=+1$, $s_2=-1$, $s_3=+1$, $s_4=-1$ and $s_0=0$ (never-treated). Define the pre-period normaliser $\bar{t}_{\mathrm{pre}}(g)=\frac{1}{A_g-1}\sum_{t=1}^{A_g-1}\frac{t-1}{T-1}$ and the post-period normaliser $\bar{t}_{\mathrm{post}}(g)=\frac{1}{T-A_g+1}\sum_{t=A_g}^{T}\frac{t-A_g}{T-1}$. Set

equation[equation omitted — 365 chars of source]

This yields: (i) $v_{it}=0$ for never-treated; (ii) sign-controlled, cohort-specific violations; (iii) weak set expansion as any of $\Delta(\mathcal{R})$, $B$, or $\Gamma$ increases (Appendix (ref) provides monotonicity statements and sufficient conditions).

\paragraph{DGPs (exact, fully specified).} All DGPs use (ref) and (ref); only $(u_{it},\lambda_t)$ dependence and selection differ.

DGP1 (i.i.d.\ Gaussian shocks).

equation[equation omitted — 77 chars of source]

DGP2 (serial correlation).

equation[equation omitted — 203 chars of source]

DGP3 (attrition/selection). Observed outcomes are $Y_{it}\mathbf{1}\{S_{it}=1\}$ with

equation[equation omitted — 147 chars of source]

where $(\eta_0,\eta_1,\eta_2)=(1.25,\ -0.35,\ -0.10)$ and $Y_{i0}=0$. Shock component is heavy-tailed:

equation[equation omitted — 88 chars of source]

\paragraph{Replications and randomisation.} For each DGP and each $(\Delta(\mathcal{R}),B,\Gamma)$ cell, run $R=2{,}000$ independent replications with fresh draws of $(\alpha_i,\lambda_t,X_{it},u_{it})$ and (for DGP3) $\{S_{it}\}$.

table[table omitted — 832 chars of source]
table[table omitted — 600 chars of source]
figure[figure omitted — 1,950 chars of source]
figure[figure omitted — 1,291 chars of source]
figure[figure omitted — 1,220 chars of source]

Data-generating processes

\paragraph{Common structure (all DGPs).}

equation[equation omitted — 175 chars of source]

where $A_i$ is adoption time, $\ell_{it}=t-A_i$ is event time, and $v_{it}(\Delta(\mathcal{R}),B,\Gamma)$ is the violation component satisfying monotone expansion in each index.

\paragraph{Treatment timing and cohort composition (all DGPs).} Cohorts $g\in\{1,\dots,G\}$ adopt at $a_g\in\{2,\dots,T\}$ with $a_1<\cdots<a_G$; never-treated have $A_i=\infty$; shares satisfy $\sum_{g=0}^G\pi_g=1$.

\paragraph{Violation component family (all DGPs).} $v_{it}(\Delta(\mathcal{R}),B,\Gamma)$ is constructed to satisfy: (i) $|v_{it}|$ weakly increases in $\Delta(\mathcal{R})$, $B$, and $\Gamma$; (ii) $B$ is tied to pre-trend magnitudes; (iii) $\Gamma$ is tied to a monotone drift-per-period family; (iv) $\Delta(\mathcal{R})$ indexes restriction-class relaxations. Calibration mappings are in Section (ref) and Appendix (ref).

\paragraph{DGP1: Baseline staggered adoption with i.i.d.\ shocks.}

equation[equation omitted — 89 chars of source]
equation[equation omitted — 182 chars of source]
equation[equation omitted — 218 chars of source]

with fixed sign pattern $(b^{(1)}_{g},\gamma^{(1)}_{g})$ scaled in calibrated units.

\paragraph{DGP2: Serial correlation and cohort-specific trends.}

equation[equation omitted — 160 chars of source]

Add a cohort trend term $\kappa_{g(i)}t$ to (ref). Treatment effects:

equation[equation omitted — 140 chars of source]

Violation:

equation[equation omitted — 202 chars of source]

where $r^{(2)}_{it}$ is admissible under the relaxed restriction class indexed by $\Delta(\mathcal{R})$.

\paragraph{DGP3: Attrition/selection with stronger heterogeneity.} Observed outcome is $Y_{it}$ only when $S_{it}=1$, with

equation[equation omitted — 156 chars of source]

Baseline shocks:

equation[equation omitted — 140 chars of source]

Violation:

equation[equation omitted — 143 chars of source]

with $(b^{(3)}_{it},\gamma^{(3)}_{it})$ constructed in calibrated units and monotone in $(B,\Gamma)$.

table[table omitted — 617 chars of source]
table[table omitted — 540 chars of source]
figure[figure omitted — 2,443 chars of source]

Performance Metrics and Reporting

\paragraph{Target estimands (fixed).} For each cohort $g\in\{1,2,3,4\}$ and event time $\ell\in\{0,1,2,3\}$ define the cohort-event ATT

equation[equation omitted — 131 chars of source]

Define the pooled post effect (reporting target) as an average over cohorts and post event times

equation[equation omitted — 181 chars of source]

Truth is known under the design because $\tau_g(\ell)$ is fixed and the violation term (ref) is fully specified; for each DGP and each $(\Delta(\mathcal{R}),B,\Gamma)$ cell,

equation[equation omitted — 192 chars of source]

so the Monte Carlo summaries benchmark against (ref).

\paragraph{Point-estimate accuracy metrics (per cell).} For replication $r\in\{1,\dots,R\}$ let $\widehat{\theta}^{(r)}$ denote the pooled estimate of $\theta^\star$ from a given estimator, and let $\widehat{\theta}^{(r)}_{g,\ell}$ denote the cohort-event estimate. Define cell-wise bias, RMSE, and median absolute error (MAE):

align[align omitted — 419 chars of source]

Parallel definitions apply to each $(g,\ell)$ component $\widehat{\theta}^{(r)}_{g,\ell}$ using $\theta^{\text{true}}_{g,\ell}=\tau_g(\ell)$.

\paragraph{Interval metrics (per cell).} Let $I^{(r)}=[L^{(r)},U^{(r)}]$ be the $1-\alpha$ confidence interval for $\theta^\star$ delivered by a method. Define empirical coverage, average length, and upper-tail excess length (to detect one-sided over-conservatism):

align[align omitted — 414 chars of source]

For component-wise reporting, define $\mathrm{Cov}_{g,\ell}$ and $\mathrm{Len}_{g,\ell}$ analogously.

\paragraph{Placebo pre-trend rejection (per cell).} Define a pre-period window for each cohort $g$ as $\ell\in\{-3,-2,-1\}$ and let $\widehat{\theta}^{(r)}_{g,\ell}$ denote placebo event-study coefficients. Let $T^{(r)}_{\mathrm{pre}}$ be the Wald test statistic for the joint null $H_0:\ \theta_{g,-3}=\theta_{g,-2}=\theta_{g,-1}=0$ (pooled across cohorts using the same weighting as (ref)). Define placebo rejection rate:

equation[equation omitted — 141 chars of source]

reported at $\alpha\in\{0.10,0.05,0.01\}$.

\paragraph{Robustness frontier (per DGP).} For each DGP and each estimator, define the robustness frontier as the set of $(\Delta(\mathcal{R}),B,\Gamma)$ cells achieving (i) target coverage and (ii) bounded interval length inflation. Let $\underline{c}=0.90$ and $\overline{\kappa}=2.5$. Let $\mathrm{Len}_0$ denote average length at $(\Delta(\mathcal{R}),B,\Gamma)=(0,0,0)$. Define admissibility indicator

equation[equation omitted — 224 chars of source]

Frontier plots (Figure (ref)) visualise $\mathbb{A}$ as a function of $(B,\Gamma)$ for each $\Delta(\mathcal{R})$.

table[table omitted — 621 chars of source]
figure[figure omitted — 1,798 chars of source]
figure[figure omitted — 1,948 chars of source]

Robustness Frontier and Placebo Pre-trends

equation[equation omitted — 117 chars of source]
equation[equation omitted — 158 chars of source]
equation[equation omitted — 175 chars of source]
equation[equation omitted — 334 chars of source]
table[table omitted — 543 chars of source]
figure[figure omitted — 1,034 chars of source]
table[table omitted — 480 chars of source]
figure[figure omitted — 926 chars of source]

Data-generating process

\noindentFactorisation of the replication law. For each replication, the joint distribution factorises as \[ f\!\left(X_i,G_i,\{D_{it}\}_{t=1}^T,\{Y_{it}\}_{t=1}^T\right) = f(X_i)\,f(G_i\mid X_i)\,f(\{D_{it}\}_{t=1}^T\mid G_i)\,f(\{Y_{it}\}_{t=1}^T\mid X_i,G_i,\{D_{it}\}_{t=1}^T), \] and the last term is pinned down by specifying \(Y_{it}(0)\) and \(\tau_g(k)\) together with the observation equation \(Y_{it}=Y_{it}(0)+D_{it}\tau_{G_i}(t-G_i)\).

\paragraph{Objects and joint law.} Each replication generates the quadruple \[ \left\{\big(X_i,G_i,\{D_{it}\}_{t=1}^T,\{Y_{it}\}_{t=1}^T\big): i=1,\dots,n\right\}, \] under a design in which adoption is confounded by covariates but conditional parallel trends holds by construction. The purpose is to force a separation between (i) design-driven TWFE contamination (Section 3 and Section 5) and (ii) bias from high-dimensional nuisance estimation when conditioning on covariates (Section 7).

\paragraph{Covariates.} For each unit $i\in\{1,\dots,n\}$ generate a time-invariant covariate vector $X_i\in\mathbb{R}^{d_x}$ as \[ X_i\sim \mathcal{N}(0,\Sigma_X), \qquad \Sigma_X\succ 0. \] Optionally include an intercept by augmenting $X_i$ with a leading 1 in implementation, but the theoretical description treats $X_i$ as mean-zero without loss of generality.

\paragraph{Confounded adoption (cohort assignment).} Let $\mathcal{G}\equiv\{1,\dots,T,\infty\}$ index adoption times and never-treated status. Specify the conditional cohort distribution using a multinomial logit:

equation[equation omitted — 189 chars of source]

where $\{\gamma_g\}_{g\in\mathcal{G}}$ are calibrated to generate non-trivial imbalance across cohorts and a non-trivial mass of never-treated units ($G_i=\infty$). Draw $G_i$ from (ref). Define absorbing treatment as

equation[equation omitted — 107 chars of source]

Under (ref)--(ref), adoption is confounded because the law of $G_i$ depends on $X_i$, so estimators that ignore $X_i$ face selection bias even when counterfactual trend restrictions hold conditionally.

The outcome data-generating process is defined by (i) a model for the untreated potential outcome \(Y_{it}(0)\) and (ii) a dynamic treatment-effect profile \(\tau_g(k)\) that maps cohort \(g\) and event time \(k=t-g\) into treated outcomes, with realised outcomes generated by the observation equation (ref).

\paragraph{Outcome process (untreated potential outcome).} Generate unit heterogeneity, time shocks, and idiosyncratic noise as \[ \mu_i\sim \mathcal{N}(0,\sigma_\mu^2), \qquad \lambda_t\sim \mathcal{N}(0,\sigma_\lambda^2)\ \text{independent over }t, \qquad \eta_{it}\sim \mathcal{N}(0,\sigma_\eta^2)\ \text{iid over }(i,t), \] mutually independent and independent of $\{X_i,G_i\}$. Define the untreated potential outcome by

equation[equation omitted — 128 chars of source]

This specification embeds both level heterogeneity ($X_i'\beta$) and trend heterogeneity ($t\,X_i'\kappa$). Because $G_i$ depends on $X_i$ via (ref), unconditional comparisons inherit differential trends correlated with adoption timing. At the same time, conditional on $X_i$, the untreated increment law is invariant across $G_i$ by construction, so conditional parallel trends holds: \[ \mathbb{E}\!\left[\Delta Y_{it}(0)\mid X_i,G_i=g\right]=\mathbb{E}\!\left[\Delta Y_{it}(0)\mid X_i,G_i=\infty\right] \quad \text{for all } g\in\{1,\dots,T\},\ t=2,\dots,T. \] Hence any remaining bias after conditioning on $X_i$ is attributable to nuisance estimation error rather than to a failure of the maintained conditional-trends structure.

\paragraph{Dynamic heterogeneous treatment effects.} Let the causal effect depend on both cohort and event time. For each cohort $g\in\{1,\dots,T\}$ define a dynamic profile $\{\tau_g(k)\}_{k\ge 0}$ on the observed horizons $k\in\{0,\dots,T-g\}$. A concrete calibration that generates heterogeneous dynamics is

equation[equation omitted — 159 chars of source]

with constants $(a_0,a_1,a_2,a_3,\ell,K)$ chosen so that effects vary meaningfully across cohorts and over horizons, including cases with slow build-up, plateauing, and mild non-monotonicity. Any alternative heterogeneous profile is admissible provided $\tau_g(k)$ varies in $g$ and $k$ so the TWFE weighting pathologies are operative.

\paragraph{Observed outcome and the joint law for $(X_i,G_i,D_{it},Y_{it})$.} Define the treated potential outcome for $t\ge g$ as

equation[equation omitted — 70 chars of source]

and generate observed outcomes as

equation[equation omitted — 97 chars of source]

Equations (ref), (ref), (ref), (ref), and (ref) jointly define the replication law. The adoption schedule determines the residualised TWFE design matrix $Z=M_XD$ and therefore the design diagnostics $\mathcal{N}(k)$ and $\mathcal{C}(k)$ prior to outcome simulation. The outcome process then ensures that (i) TWFE can be distorted even when conditional parallel trends holds, because heterogeneity activates design-driven cross-horizon mixing, and (ii) conditioning on $X_i$ is essential, because $X_i$ drives both adoption timing and untreated trends, which makes the covariate-tilting Riesz representer and orthogonal score construction in Section 7 non-optional when $d_x$ is moderate or large.

Evaluation targets

Let $\theta(k)\equiv\tau(k)$ denote the event-time target defined in (ref) for $k\in\mathcal{K}$. Each Monte Carlo replication $r=1,\dots,R$ generates $(Y^{(r)},D^{(r)},G^{(r)})$ from Section (ref) and produces (i) the TWFE event-study coefficients $\hat\beta_k^{(r)}$ from (ref), (ii) the heterogeneity-robust estimator $\hat\tau^{(r)}(k)$ from (ref) using the standard nuisance plug-in components, (iii) the orthogonal-score estimator $\hat\tau_{\mathrm{orth}}^{(r)}(k)$ obtained by estimating $\widehat{\mathrm{GATT}}(g,t)$ via (ref) and then aggregating to event time using (ref), (iv) the design diagnostics $\mathcal{N}^{(r)}(k)$ and $\mathcal{C}^{(r)}(k)$ computed from the residualised design matrix $Z^{(r)}$, and (v) a sensitivity-robust confidence region $\mathcal{C}_n^{(r)}(\mathcal{R};k)$ for the identified set of $\theta(k)$ under the restriction class in Section (ref).

\paragraph{Finite-sample accuracy.} For each estimator $\hat\theta^{(r)}(k)\in\{\hat\beta_k^{(r)},\hat\tau^{(r)}(k),\hat\tau_{\mathrm{orth}}^{(r)}(k)\}$, define Monte Carlo bias and root mean squared error

equation[equation omitted — 242 chars of source]

When the DGP includes controlled violations $\delta\neq 0$, $\theta(k)$ is interpreted as the pseudo-true target induced by the maintained estimand definition, and bias is interpreted relative to that target.

\paragraph{Design diagnostics and distortion.} Let the TWFE distortion at horizon $k$ be

equation[equation omitted — 82 chars of source]

and let $\mathcal{N}(k)$ and $\mathcal{C}(k)$ be the indices in Definition (ref) and Definition (ref). The diagnostic content is assessed by the strength of the cross-replication association between design risk and distortion, for example via Pearson correlations

equation[equation omitted — 164 chars of source]

and by regressions of $\mathrm{Dist}(k)$ on $(\mathcal{N}(k),\mathcal{C}(k))$ to quantify marginal predictive content holding the other index fixed. These objects evaluate whether large negative-weight mass and cross-horizon contamination are informative ex ante signals of estimator unreliability in finite samples.

\paragraph{Sensitivity-robust coverage and power.} Let $\Theta(\mathcal{R};k)$ denote the true identified set for $\theta(k)$ under the restriction class $\Delta(\mathcal{R})$, and let $\mathcal{C}_n^{(r)}(\mathcal{R};k)$ denote the computed confidence region in replication $r$. Uniform-coverage performance is summarised by

equation[equation omitted — 187 chars of source]

which targets $1-\alpha$ when the DGP satisfies the maintained restrictions RambachanRoth2023. Power is evaluated by considering alternatives in which $\delta$ violates the maintained restriction class (for example, curvature exceeding $\Gamma$), and reporting the frequency with which $\mathcal{C}_n^{(r)}(\mathcal{R};k)$ excludes economically relevant null hypotheses. A convenient null is $\theta(k)=0$; define the rejection indicator \[ \mathbbm{1}\!\left\{0\notin \mathcal{C}_n^{(r)}(\mathcal{R};k)\right\}, \] and define power as its Monte Carlo mean under the alternative.

These targets jointly evaluate (i) estimator accuracy, (ii) whether the diagnostics operationalise design fragility in a quantitatively meaningful manner, and (iii) whether sensitivity-robust confidence regions behave as advertised when the maintained restriction class is correct and when it is misspecified RambachanRoth2023.

Empirical Application: Bank Deregulation Shock (Replicable Panel)

Setting and policy timing

This section applies the proposed diagnostics and robust DID/event-study estimators to a canonical staggered-adoption policy setting in finance and macroeconomics: state-level banking deregulation. The empirical goal is not novelty of the application, but scale and interpretability of the deviations between standard TWFE event-studies and heterogeneity-robust counterparts under the sensitivity framework developed in Sections (ref)--(ref), with interpretation guided by recent practitioner and methodological syntheses WingEtAl2024,AbadieAngristFrandsen2025,Roth2024.

Let $s\in\{1,\dots,S\}$ index states and $t\in\{1,\dots,T\}$ index years. Each state adopts deregulation at an adoption year $A_s\in\mathbb{Z}$, with $A_s=\infty$ for never-treated states. Define the treatment indicator

equation[equation omitted — 54 chars of source]

and event-time $k=t-A_s$. For event-study specifications, define event-time indicators

equation[equation omitted — 110 chars of source]

where $\mathcal{K}=\{-K,\dots,-2,0,1,\dots,K\}$ and $k=-1$ is omitted as the reference period.

Data and panel construction (replicable)

The application is designed to be replicable from public sources with transparent construction rules and audit trails. The unit of analysis is a state-year panel $\{(Y_{st},D_{st},X_{st})\}$ with a single adoption date $A_s$ per state, a clearly defined set of never-treated states, and harmonised outcomes and covariates in constant units.

\paragraph{Outcomes.} Let $Y_{st}$ denote an economic outcome at the state-year level. The baseline outcome is real personal income per capita (log), with secondary outcomes used for robustness and interpretability: employment (log), unemployment rate (level), and real gross state product (log). Outcomes are deflated to constant dollars when applicable, then transformed as

equation[equation omitted — 119 chars of source]

with all transformations fixed ex ante and applied uniformly across states and years. Event-study interpretation and aggregation follow recent guidance on reading and comparing dynamic effects under staggered adoption Roth2024.

\paragraph{Treatment timing.} Adoption years $\{A_s\}_{s=1}^S$ are stored in a single state-level file with one row per state and a unique adoption year variable. The panel uses a single policy margin (one adoption clock). If multiple deregulation margins exist in the underlying legal history, the empirical design fixes one margin as primary and treats other margins as (i) alternative treatments for sensitivity checks, or (ii) exclusion criteria (dropping states with overlapping reforms within a pre-specified window). The timing file is validated by three mechanical checks:

equation[equation omitted — 221 chars of source]

\paragraph{Covariates and diagnostics inputs.} Let $X_{st}$ include time-varying observables used for balance diagnostics, placebo checks, and optional residualisation steps for orthogonal-score implementations. The baseline set is parsimonious: population (log), sector shares (levels or logs), and state fiscal capacity proxies (levels). Diagnostics emphasise staggered-adoption threats and cohort composition shifts WingEtAl2024,AbadieAngristFrandsen2025. When orthogonal scores are used, residualisation is performed by a pre-registered learner class $\mathcal{G}$ with cross-fitting, aligning the empirical implementation with the efficiency and robustness motivations in modern DID/event-study developments SantAnnaZhao2025.

\paragraph{Panel window and trimming.} Fix a sample window $t\in\{t_0,\dots,t_1\}$ with $t_0$ chosen to ensure adequate pre-treatment support for early adopters and $t_1$ chosen to avoid severe right-censoring for late adopters. Define the estimand event-time support by trimming event times outside $\mathcal{K}$ and trimming cohorts with insufficient pre-periods:

equation[equation omitted — 206 chars of source]

The analysis uses the restricted sample $\{(s,t): s\in\mathcal{S}_{\mathrm{keep}},\ t_0\le t\le t_1\}$, with $K_{\mathrm{pre}},K_{\mathrm{post}}$ fixed and reported.

\paragraph{Replicability checklist (machine-verifiable).} All construction steps output intermediate artefacts that can be re-run and diffed:

equation[equation omitted — 284 chars of source]

The empirical section then treats these artefacts as fixed inputs for (i) TWFE event-studies, (ii) heterogeneity-robust estimators, and (iii) the sensitivity-calibration objects $(\Delta(\mathcal{R}),B,\Gamma)$, with interpretation anchored in recent design guidance and modern event-study reading rules WingEtAl2024,Roth2024,AbadieAngristFrandsen2025.

Data and panel construction (replicable)

The application is designed to be fully replicable using publicly available data and a well-known staggered adoption design in the literature, with construction rules chosen to align with modern practical guidance on staggered DID and event-study designs WingEtAl2024,AbadieAngristFrandsen2025. The panel is constructed as follows.

\paragraph{Units, outcomes, and covariates.} Let $Y_{st}$ denote the outcome. The main outcomes are restricted to a small set to limit specification mining: (i) real GSP per capita growth, (ii) employment-to-population, and (iii) real wage growth, with one placebo outcome selected to be plausibly unaffected in the short run. State and year fixed effects are always included; optional covariates $X_{st}$ include lagged state macro controls (for example, demographic composition or sector shares) used only in robustness checks.

The outcome and covariate set also supports two extensions that are reported only as ancillary checks: a spillover-robust sensitivity screen (for example, allowing exposure in nearby states or economically linked states) and an instrument-based counterfactual-policy check that reinterprets adoption timing through an instrument for the policy margin when such an instrument is available Lee2025,RothKolesarMontielOlea2025. The empirical section does not treat these as new identification strategies; they are used to stress-test whether the diagnostics and sensitivity objects flag instability when interference or instrumented policy counterfactuals are plausible.

\paragraph{Sample window.} Two samples are used throughout:

equation[equation omitted — 112 chars of source]

where the windowed sample is used for dynamic plots and calibration of pre-trends, and the full sample is used for baseline estimates and computational scaling. Event-time availability is enforced mechanically by dropping cohort--event cells with insufficient support, and all trimming rules are fixed before estimation to avoid post hoc window choice Roth2024.

\paragraph{Treatment cohorts.} Define cohorts by adoption year $g\in\mathcal{G}$ where $\mathcal{G}=\{A_s: A_s<\infty\}$. Let $G_s\equiv A_s$ denote cohort membership. Controls at each $t$ are either not-yet-treated states ($t<A_s$) and (if present) never-treated states ($A_s=\infty$). All code must enforce a single control definition consistently across TWFE and robust estimators, and all reported estimates must declare the control definition explicitly WingEtAl2024,AbadieAngristFrandsen2025.

Baseline estimators and comparators

This subsection defines the estimators compared in the empirical application. The estimand is the dynamic average treatment effect at event time $k$:

equation[equation omitted — 123 chars of source]

with expectations taken over treated observations in event time $k$ and the relevant counterfactual defined by the design.

\paragraph{TWFE event-study (benchmark).} The conventional TWFE event-study is estimated as

equation[equation omitted — 149 chars of source]

with standard errors clustered at the state level. Equation (ref) is reported as a benchmark only, because the weights implicit in $\beta^{\text{TWFE}}_k$ can be non-convex under heterogeneity and staggered timing AbadieAngristFrandsen2025,Roth2024.

\paragraph{Heterogeneity-robust event-study (main).} The main estimator targets $\tau(k)$ using a heterogeneity-robust approach (group-time or interaction-weighted). Denote the robust estimates by $\widehat{\beta}^{\text{ROB}}_k$. Estimation proceeds by aggregating cohort-specific comparisons with an explicit control set (not-yet-treated and never-treated, as specified above), ensuring that each $k$ corresponds to a well-defined comparison and avoiding negative-weight pathologies WingEtAl2024,Tominaga2024.

\paragraph{Optional cross-check estimator.} A second robust estimator is included as a cross-check using the same estimand $\tau(k)$ but different weighting/aggregation. This is used to verify that results are not an artefact of a single implementation and to compare efficiency claims against the orthogonal-score framing when applicable SantAnnaZhao2025.

\paragraph{Ancillary stress tests (spillovers and instrumented counterfactuals).} Two ancillary stress tests are reported as diagnostics rather than headline estimates. First, a spillover screen evaluates whether estimated dynamics are sensitive to excluding potentially exposed controls or redefining exposure using a parsimonious proximity/economic-link rule; this addresses the possibility that staggered adoption induces interference that contaminates not-yet-treated controls Lee2025. Second, when the policy margin plausibly admits an instrument, an instrumented counterfactual-policy formulation is used to check whether the qualitative event-time pattern is stable when adoption timing is treated as endogenous and shifted by an instrument; this is reported only to show whether the proposed diagnostics flag fragility under plausible endogeneity channels RothKolesarMontielOlea2025.

Diagnostics: weights, pre-trends, and placebo

The empirical section demonstrates two objects: where and how TWFE differs from heterogeneity-robust estimators, and whether the differences matter at empirically relevant magnitudes WingEtAl2024,AbadieAngristFrandsen2025.

\paragraph{Weight diagnostics.} Compute and report the implied comparison weights for TWFE (or the equivalent decomposition for the chosen robust estimator), and summarise the mass on inadmissible comparisons (already-treated as controls) and the mass on negative weights, because these are the channels through which TWFE can generate sign reversals and non-interpretable dynamic patterns under staggered adoption AbadieAngristFrandsen2025,Tominaga2024. Concretely, report (i) the share of total absolute weight that is negative, (ii) the minimum and maximum cell weights, and (iii) the share of weight placed on comparisons between treated cohorts at different event times.

\paragraph{Pre-trend diagnostics.} Estimate the pre-treatment coefficients for $k\le -2$ under both TWFE and robust specifications and test

equation[equation omitted — 51 chars of source]

Record the joint $p$-value, the maximum absolute pre-coefficient magnitude, and the maximum absolute $t$-statistic over $k\le -2$, because these provide a transparent bridge from observed pre-period instability to the calibration objects used later WingEtAl2024,Roth2024.

\paragraph{Placebo adoption timing.} Construct placebo adoption years $\widetilde{A}_s=A_s+\Delta$ for a fixed $\Delta>0$ (for example, $\Delta=3$), re-define $\widetilde{D}_{st}\equiv\mathbbm{1}\{t\ge \widetilde{A}_s\}$, and re-estimate the full pipeline on the placebo-treated sample. The key output is the rejection frequency of placebo post coefficients, the distribution of placebo joint pre-trend $p$-values, and the distribution of placebo maximum pre-coefficient magnitudes, because these objects are used to interpret whether the empirical sensitivity calibrations are conservative relative to a falsification benchmark WingEtAl2024,Roth2024.

Calibration of \texorpdfstring{$(B,\Gamma,\Delta(\mathcal{R}))$}{(B,Gamma,Delta(R))} for the empirical panel

This subsection implements the calibration mappings introduced in Section (ref) and Appendix (ref). The calibration converts observables (pre-trends and imbalance) into the sensitivity parameters $(B,\Gamma,\Delta(\mathcal{R}))$ in interpretable units that match the empirical panel scale WingEtAl2024,RothKolesarMontielOlea2025.

\paragraph{Calibration of $B$ from empirical pre-trends.} Let $\widehat{\beta}_{k}$ denote the estimated pre-treatment event-study coefficients for $k\le -2$ under the robust estimator. Define the pre-trend magnitude statistic

equation[equation omitted — 128 chars of source]

Set $B$ by a deterministic rule

equation[equation omitted — 80 chars of source]

where $c_B\in\{1,2\}$ is reported transparently, and $c_B=2$ corresponds to a conservative doubling of the largest observed pre-trend magnitude. A holdout variant chooses $c_B$ to control false rejections under placebo timing by selecting the smallest $c_B$ such that placebo post-period rejections remain below a target level (for example, $5\%$) across outcomes WingEtAl2024,Roth2024.

\paragraph{Calibration of $\Gamma$ as monotone drift per period.} Interpret $\Gamma$ as a maximum per-period drift (in outcome units) in violations of the maintained restrictions, and map $\Gamma$ to a drift rate via

equation[equation omitted — 103 chars of source]

where $\widehat{\sigma}_{\Delta Y}$ is the empirical standard deviation of year-to-year changes in $Y_{st}$ within the pre-treatment sample, and $\gamma$ is a dimensionless grid value (for example, $\gamma\in\{0,0.25,0.5,1,2\}$). This mapping ensures that a value such as $\gamma=1$ corresponds to a drift bound equal to one pre-period innovation scale, and monotonicity of identified set expansion in $\Gamma$ follows from Appendix (ref) RothKolesarMontielOlea2025,SantAnnaZhao2025.

\paragraph{Calibration of restriction slack $\Delta(\mathcal{R})$ from imbalance.} Let $\widehat{\delta}$ denote a pre-period imbalance statistic, computed either as a standardised mean difference in pre-levels or as a standardised difference in pre-slopes between treated and control units (both computed using the same control definition as the main estimator). Define

equation[equation omitted — 96 chars of source]

with $d\in\{0,1,2\}$ a slack multiplier that is reported and varied in sensitivity plots. This anchoring ensures $\Delta(\mathcal{R})$ scales with an observed pre-treatment discrepancy rather than an abstract norm choice, and it matches the diagnostic logic that small measured imbalance should correspond to a small admissible relaxation of the maintained restriction class WingEtAl2024,Roth2024.

Sensitivity results and robustness frontier

The empirical sensitivity analysis reports how inference on $\tau(k)$ changes over a grid in $(B,\Gamma,\Delta(\mathcal{R}))$ and is designed to make the magnitudes required for sign changes explicit in outcome units WingEtAl2024,RothKolesarMontielOlea2025.

\paragraph{Robust intervals over the sensitivity grid.} For each $k\in\{0,1,2\}$ (and optionally all $k\in\mathcal{K}$), compute robust intervals $\mathcal{I}_k(B,\Gamma,\Delta(\mathcal{R}))$ over a Cartesian grid

equation[equation omitted — 326 chars of source]

and record, for each grid point, (i) sign stability relative to the baseline robust point estimate $\widehat{\tau}(k)$, (ii) interval length, and (iii) whether $0\in\mathcal{I}_k(B,\Gamma,\Delta(\mathcal{R}))$, because these are the three outputs that translate the sensitivity framework into a decision-relevant object Roth2024,SantAnnaZhao2025.

\paragraph{Breakdown frontier.} Define the breakdown point $\Gamma_k^\star$ as the smallest $\Gamma$ (given $B$ and $\Delta(\mathcal{R})$) at which the interval contains zero:

equation[equation omitted — 175 chars of source]

Compute $\Gamma_k^\star$ on the discrete grid $\mathcal{G}$ by the smallest $\Gamma\in\mathcal{G}$ satisfying $0\in\mathcal{I}_k(\cdot)$, and report $\Gamma_k^\star$ both in raw outcome units and as the dimensionless drift multiple $\gamma_k^\star=\Gamma_k^\star/\widehat{\sigma}_{\Delta Y}$ to preserve interpretability across outcomes WingEtAl2024,RothKolesarMontielOlea2025.

\paragraph{Frontier interpretation under heterogeneity and spillovers.} Report the frontier jointly with the diagnostic objects from Section (ref), because large negative-weight mass or placebo failures indicate that small $\Gamma$ values can be empirically plausible. If a spillover-robust variant is implemented, re-compute the frontier under spillover allowances to show how robustness changes when treatment affects nominal controls, as staggered finance policies can plausibly induce cross-state spillovers Lee2025,AbadieAngristFrandsen2025.

\paragraph{Minimum artefacts (non-negotiable).} Report the sensitivity outputs in Table (ref) and Figures (ref)--(ref) using TikZ only, with all numbers populated by the replication pipeline.

table[table omitted — 1,011 chars of source]
figure[figure omitted — 421 chars of source]
figure[figure omitted — 409 chars of source]

Implementation and replication pipeline

The empirical section is reproducible by a minimal pipeline that (i) builds the state-year panel, (ii) estimates TWFE and heterogeneity-robust event-studies, (iii) computes diagnostics, (iv) runs the sensitivity grid, and (v) exports the tables and TikZ figures listed in this section. Randomness is restricted to placebo and any resampling steps, with a fixed seed WingEtAl2024,Tominaga2024.

\paragraph{Pipeline contract and file outputs.} Define a single configuration file that pins $(K,\underline{t},\overline{t})$, the control definition (not-yet-treated only vs.\ not-yet-treated plus never-treated), the outcome transforms, and the sensitivity grids $\mathcal{B},\mathcal{G},\mathcal{D}$. The pipeline writes (a) a clean analysis panel, (b) estimator outputs for TWFE and robust methods, (c) diagnostic summaries, and (d) sensitivity outputs, then renders Table (ref) and Figures (ref)--(ref) by exporting numeric arrays to small .tex fragments that are \string\input into the main document.

table[table omitted — 702 chars of source]
table[table omitted — 801 chars of source]
table[table omitted — 823 chars of source]
figure[figure omitted — 891 chars of source]
figure[figure omitted — 729 chars of source]
figure[figure omitted — 888 chars of source]
figure[figure omitted — 475 chars of source]
figure[figure omitted — 667 chars of source]
figure[figure omitted — 766 chars of source]
figure[figure omitted — 626 chars of source]
figure[figure omitted — 484 chars of source]

Extensions

This section records extensions that preserve the central logic of the paper: (i) define an estimand as an explicit convex aggregation of cohort--time causal objects, (ii) exhibit how conventional regressions generate implicit, design-driven mixtures of heterogeneous effects, and (iii) construct estimators and diagnostics whose interpretation is enforced by construction rather than asserted ex post DeChaisemartinDHaultfoeuille2023.

Multiple treatments and overlapping policies

Let there be $J\ge 2$ policies indexed by $j\in\{1,\dots,J\}$. For each policy $j$, define an adoption time $G_i^{(j)}\in\{1,\dots,T,\infty\}$ and the absorbing treatment indicator \[ D_{it}^{(j)}=\mathbbm{1}\{t\ge G_i^{(j)}\}\mathbbm{1}\{G_i^{(j)}<\infty\}. \] Let $d_t\equiv(d_t^{(1)},\dots,d_t^{(J)})\in\{0,1\}^J$ denote the treatment vector at time $t$. Potential outcomes are indexed by treatment histories; for notational economy write $Y_{it}(\mathbf{d})$ for the potential outcome at time $t$ under a (possibly dynamic) treatment path $\mathbf{d}=\{d_s\}_{s\le t}$. A primitive estimand for policy $j$ at cohort--time $(g,t)$ is the incremental effect of switching on policy $j$ holding the remaining policies at their realised values:

equation[equation omitted — 206 chars of source]

where $d_t^{(-j)}$ denotes the vector of other policies at $t$. This is an economically interpretable partial effect, but its identification requires a strengthened no-anticipation restriction and a multi-policy parallel-trends condition formulated for the relevant counterfactual increments:

equation[equation omitted — 234 chars of source]

with the conditioning set enlarged as needed to render untreated increment evolution comparable across exposure patterns. The central warning is geometric rather than philosophical: a TWFE regression that includes event-time indicators for multiple policies simultaneously generally estimates an implicit linear combination of all policy effects across horizons, because the residualised indicators for distinct treatments are correlated through overlapping adoption patterns. Hence contamination can arise even when each policy separately satisfies its own parallel-trends restriction, because the projection operator loads policy-$j$ regressors onto residual variation that is itself a mixture of other-policy effects DeChaisemartinDHaultfoeuille2023.

A robust construction proceeds policy-by-policy with explicit conditioning on the realised exposure to other policies. For a fixed $j$, define a control set at time $t$ as units that have not yet adopted policy $j$ by $t$ and that match exposure to the other policies (exactly or through coarsened strata). Estimation then targets $\mathrm{GATT}^{(j)}(g,t)$ cell-by-cell and aggregates to event time using explicit convex weights as in (ref). The diagnostic indices generalise by computing, for each policy $j$, the implicit TWFE weights $w_{g,k'}^{(j)}(k)$ and measuring negative-weight mass and cross-horizon mass as in Section 5, with an additional cross-policy contamination index defined by the absolute weight placed on $\tau^{(\ell)}(\cdot)$ for $\ell\neq j$ induced by the joint regression geometry.

Non-absorbing adoption and switching

Absorbing adoption is not intrinsic to the event-study logic; it is a restriction on treatment paths. Let $D_{it}\in\{0,1\}$ follow an arbitrary path, allowing switching. Replace the adoption time $G_i$ by a treatment history $H_{it}=(D_{i1},\dots,D_{it})$. Potential outcomes are indexed by histories, and dynamic causal objects must be defined relative to a reference path. A natural estimand is the effect of initiating treatment at event time $k=0$ and maintaining it for $k\ge 0$ relative to remaining untreated, \[ \tau_g(k)\equiv \mathbb{E}\!\left[Y_{i,g+k}(1^{k+1},0^{g-1})-Y_{i,g+k}(0^{g+k})\mid G_i=g\right], \] where $1^{k+1}$ denotes $k+1$ consecutive treated periods after adoption and $0^{g-1}$ denotes untreated pre-periods. When switching occurs, this estimand is not equal to the effect of the realised treatment path unless additional restrictions connect realised switching behaviour to the maintained reference path.

Identification can still proceed with group--time objects if the control group at time $t$ is defined as units not treated at $t$ and if parallel-trends restrictions are stated for the appropriate counterfactual increments under the reference path. Estimation remains cell-by-cell; the crucial discipline is that aggregation must respect the chosen reference path. The diagnostic logic extends unchanged: any regression that encodes switching by contemporaneous $D_{it}$ and fixed effects can load on a mixture of past and future treatment episodes, so cross-horizon contamination generalises to cross-history contamination, which can be detected by computing the implicit weights for indicators of treatment histories rather than simple event-time indicators DeChaisemartinDHaultfoeuille2023.

Treatment intensity and continuous exposure

Let $A_{it}\in\mathbb{R}_+$ be a scalar treatment intensity (dose), such as a continuous policy exposure. Define potential outcomes $Y_{it}(a)$ for $a\in\mathcal{A}\subset\mathbb{R}_+$. A baseline estimand is the average causal response at exposure level $a$ relative to zero for cohort $g$, \[ \tau_g(k;a)\equiv \mathbb{E}\!\left[Y_{i,g+k}(a)-Y_{i,g+k}(0)\mid G_i=g\right], \] or, when differentiability is plausible, the average marginal response $\partial_a\mathbb{E}[Y_{i,g+k}(a)\mid G_i=g]$ evaluated at a reference exposure. Identification requires a parallel-trends restriction formulated for the counterfactual increment process under zero exposure and an exclusion restriction that maps observed exposure paths into potential outcomes in a stable way. Conventional TWFE specifications that treat $A_{it}$ as a linear regressor inherit the same projection-driven mixing problem: the slope coefficient is an implicit average of heterogeneous marginal responses across cohorts, times, and exposure levels, with sign-reversing weights possible once fixed effects are partialled out DeChaisemartinDHaultfoeuille2023.

A robust alternative is to discretise exposure into bins that define mutually exclusive exposure states and then apply the group--time logic to each bin relative to a baseline, with explicit convex aggregation across cohorts and horizons. When continuous exposure is retained, orthogonal-score methods of Section 7 can be applied to estimate average causal responses as linear functionals, with the Riesz representer corresponding to the appropriate exposure-tilting functional ChernozhukovNeweySingh2022. In both cases, the diagnostics generalise by computing the projection-induced weights for exposure indicators or basis expansions of $A_{it}$, and then measuring negative-weight mass and cross-horizon mass as in Section 5, now interpreted as properties of the exposure design rather than of binary adoption.

These extensions show that the central distinction is not between “TWFE” and “alternatives” as software choices, but between estimands that are explicit convex functionals of well-defined causal objects and estimands that are implicit, design-driven mixtures whose interpretation fails under heterogeneity, overlap of treatments, switching, or intensity variation DeChaisemartinDHaultfoeuille2023.