EconBase
← Back to paper

Stacked Triple Differences

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

117,594 characters · 29 sections · 65 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Stacked Triple Differences

abstractTriple differences (DDD) is a workhorse quasi-experimental design in applied economics. But, under staggered adoption, its conventional three-way fixed-effects (3WFE) implementation inherits the forbidden-comparison and interpretation issues now well understood in the difference-in-differences literature. To resolve these issues, I introduce stacked DDD. I extend the stacked difference-in-differences approach to the DDD setting by creating self-contained stacks, each consisting of four cells over an event window: treated and clean comparison cohorts, each with treatment-eligible and treatment-ineligible units. Appending these stacks yields a unified dataset for estimating treatment effects without making forbidden comparisons. I prove that, at each post-treatment event-time, a linear regression with fully saturated fixed-effects applied to the stacked dataset identifies a strictly positive, cell-size-weighted average of stack-level conditional average treatment effects, with stack weights proportional to stack-level cell sizes. Building on this characterization, I outline alternative weighting schemes that recover distinct, transparent causal estimands with clear interpretations. Stacked DDD complements recent GMM and imputation-based frameworks by trading efficiency for regression-based transparency, pairwise (rather than global) parallel trends, and direct control over aggregation weights. I provide two empirical illustrations where stacked DDD yields substantially different quantitative conclusions compared to existing procedures.

\@startsection{section}{2} \z@{.5\linespacing\@plus.7\linespacing}{-.5em} {\normalfont}{Motivation and Introduction}

Triple differences (DDD) is among the most widely used quasi-experimental research designs in applied economics. In settings where treatment requires satisfying two criteria---belonging to a group whose members become exposed to a policy and being eligible within that group---DDD permits both group-specific and eligibility-specific departures from parallel trends, and is therefore a more defensible identification strategy than standard difference-in-differences (DiD) whenever either margin alone is implausibly parallel. This flexibility explains its prominence in public, labor, health, and environmental economics.\footnote{See olden_triple_2022 for a systematic survey of applications.} The corresponding formal econometric theory, however, has not received as much attention compared to DiD, which has been the object of sustained methodological reexamination goodman-bacon_difference--differences_2021, callaway_difference--differences_2021-1, sun_estimating_2021, de_chaisemartin_two-way_2020, borusyak_revisiting_2023. Only recently have strezhnev_decomposing_2023, leventer_triple_2025, caron_triple_2025, and ortiz-santanna_triple_2025 examined the conventional DDD practice. They document that the three-way fixed effects (3WFE) specifications commonly used to implement DDD under staggered adoption target estimands contaminated by forbidden comparisons, mirroring the observation of goodman-bacon_difference--differences_2021 for DiD, and additional bias arises in the presence of covariates. Furthermore, they develop identification arguments under staggered adoption and heterogeneous treatment effects.

This paper develops the stacked triple-differences (stacked DDD) framework, which is complementary to existing approaches. This embodies the stacking approach now well known in the DiD literature, used by cengiz_effect_2019, deshpande_who_2019, butters_how_2022, matsuzawa_minimum_2025 and others, and discussed formally by wing_stacked_2024. The stacking approach works in the following manner. First, for each treated cohort, design a clean comparison group that consists of units that are not treated within the event window of interest. Second, create a dataset for each treated cohort and its clean comparison group; this is called a stack. Finally, concatenate all stacks into a unified dataset, which I call the stacked dataset. By construction, in the DDD setting, there are both treatment-eligible and treatment-ineligible units within both treatment and control groups in each stack.

Estimating treatment effects on the stacked dataset then exploits three sources of identifying variation simultaneously. First, within each treated group, eligible units are subject to the policy while ineligible units are not, so differencing them removes any confounding trend common to the group regardless of eligibility status. Second, across the treatment and clean comparison groups, differencing outcomes removes eligibility-specific shocks that would be shared by eligible units in both cohorts. Third, difference post- and pre-treatment outcomes then isolates the treatment effects within each group-by-eligibility cell. Restricting all three comparisons to uncontaminated units, the stacked construction delivers a $2 \times 2 \times 2$ sub-experiment for every cohort.\footnote{I will use “stack” and “cohort” interchangeably throughout, which are the nomenclature of the stacked DiD and staggered DiD literature, respectively.}

A distinctive feature of DDD is that units apparently lacking direct identifying power nonetheless play an essential role for identification. The ineligible units in the comparison cohort---neither treated nor eligible---anchor the baseline of the clean comparison group, and it is this anchor that permits the third difference to separate the treatment effect from an eligibility-specific trend. Without them, a researcher cannot distinguish a genuine treatment effect from a group-specific shock to eligible units in the treated cohort. The identifying content in the stacked DDD framework is, therefore, present by design.

A central question in the stacked DDD framework is identification of causal estimands. In practice, it is of interest to characterize the estimand that an event-study OLS regression targets when applied on the stacked dataset. As a starting point, I first characterize the estimand targeted by the standard DDD event-study OLS regression with unit, group-by-time, and eligibility-by-time fixed effects, and where the treatment timings are staggered. I show that its event-study coefficient at event-time $e$ is a linear combination of cohort- and event-time-specific conditional average treatment effects on the treated: the own-event-time cohort weights sum to one but are not guaranteed to be non-negative for individual treatment cohorts, and the weights on {other} event-times $e' \neq e$ are generically nonzero. Only under homogeneous treatment effects across cohorts does this estimand collapse to the conventional ATT; outside this case, the coefficient admits negative-weight contamination of the kind documented for DiD by goodman-bacon_difference--differences_2021.

In the stacked DDD framework, the answer parallels sun_estimating_2021's sun_estimating_2021 analysis of DiD: at each post-treatment event-time $e$, the fully saturated OLS regression on the stacked panel targets a strictly positive, cell-size-weighted convex combination of cohort-specific conditional average treatment effects on the treated. The implied weights from fully saturated OLS regression are a function of cell sizes within each stack, therefore this estimator assigns large weights to stacks whose stack-level treatment effect estimates are likely more precise.

However, it is not clear that the weights implied by fully saturated OLS regression estimated on the stacked dataset lead to a causally interesting aggregation of cohort-level treatment effects. Building on this characterization, I describe alternative aggregation schemes that can be applied to stacks to target estimands with more explicit population causal interpretations: cohort-size weights, which deliver a per-unit average treatment effect on the treated; and equal weights, which deliver a simple average of cohort-level average treatment effects on the treated. The stacked DDD framework, by design, addresses a concern raised by de_chaisemartin_two-way_2020 and de_chaisemartin_difference--differences_2024 that pooled estimators deliver estimands whose implicit weights are opaque and not necessarily non-negative. In this respect, stacked DDD sits alongside the doubly-robust DiD of sant2020doubly, the local-projection DiD of dube_local_2023, and the design-based framework of borusyak_revisiting_2023 as complementary estimators that avoid forbidden comparisons.

The stacked design also weakens the identifying assumption: parallel trends need hold only pairwise within each stack, between the treated group and its clean comparison group, and not globally across all pairs of treated and control groups. This is a weaker restriction than the global DDD parallel trends assumed by pooled estimators. In the stacked DDD framework, estimation of the causal parameters of interest follow from (weighted) fully saturated OLS regressions, and the asymptotic theory is large-$n$ regime with both the numbers of time and stacks held fixed. I further show that because the stacked estimator is exactly recovered by a fully saturated OLS regression, applying the cluster-robust standard errors clustered at the level of treatment delivers valid inference; specifically, I show that the cluster-robust standard errors automatically account for the cross-stack dependence induced by comparison units appearing in multiple sub-experiments. Therefore, I sidestep the need for deriving analytical standard error corrections abadie_when_2022, mackinnon_cluster-robust_2023 when OLS regression with fully saturated fixed effects is used. In summary, stacked DDD is complementary to the GMM procedure of ortiz-santanna_triple_2025 and the imputation-based approach of borusyak_revisiting_2023: it trades efficiency for regression-based transparency, pairwise rather than global parallel trends, and explicit control over the aggregation weights.

Two empirical applications illustrate the practical consequences from applying stacked DDD. First, applied to hansen_national_2023's study of genetically modified crop adoption, the stacked DDD estimator confirms the qualitative finding of positive yield effects but produces point estimates substantially smaller than the pooled 3WFE specification. Second, applied to shastry_vaccine_2025's evaluation of Gavi's vaccine program, the stacked DDD estimator reproduces the conclusion that Gavi-funded vaccines reduced child mortality from related causes, but attenuates the magnitude of the original study's estimates by roughly a third in the full sample and aligns with the authors' preferred vital-registry subset at a reduction of approximately a quarter of a death per 1,000 live births. In both applications, the stacked decomposition discloses which stack-level conditional average treatment effect on the treated drive the aggregated effect.

The remainder of the paper proceeds as follows. Section (ref) develops the notation and identifying assumptions. Section (ref) examines the current practice of 3WFE OLS regressions. Section (ref) defines the stacked sub-experiment, the target estimands, and the within-stack and pooled regression specifications that recover them. Section (ref) establishes identification under the DDD PCT restriction. Section (ref) develops the asymptotic theory and inference. Section (ref) reports the two illustrations of stacked DDD in practice. All proofs and extended discussions are collected in the appendices.

\@startsection{section}{2} \z@{.5\linespacing\@plus.7\linespacing}{-.5em} {\normalfont}{Setup and Assumptions}

Notation

I observe panel data on $n$ units indexed by $i$ over a common set of $T$ time indexed by $t$. Each unit is observed in a subset $\mathcal{T}_i \subseteq \{1, \ldots, T\}$; the panel need not be balanced. Throughout, $T$ is fixed and all asymptotics are with respect to $n \to \infty$.

Each unit belongs to a treatment-enabling group $S_i \in \mathcal{S}$, where

equation[equation omitted — 87 chars of source]

The variable $S_i$ records the period at which unit $i$'s group first becomes exposed to the policy; a unit with $S_i = g$ for $g \in \{2, \ldots, T\}$ belongs to a group whose treatment becomes enabled in period $g$, and a unit with $S_i = \infty$ belongs to a group whose treatment is never enabled within the sample. Let $\mathcal{G}_{\mathrm{trg}} = \mathcal{S} \setminus \{\infty\}$ denote the set of treated groups. Within each group, an eligibility indicator $Q_i \in \{0,1\}$ records whether the unit is itself eligible for treatment; I take $Q_i$ to be time-invariant throughout the panel, though the framework extends straightforwardly to time-varying eligibility. Treatment is received only by units in an activated enabling group who are themselves eligible, so the observed treatment status is

equation[equation omitted — 78 chars of source]

implying that units with $Q_i = 0$ are never treated irrespective of their group, and units with $S_i = \infty$ are never treated irrespective of their eligibility. Let $X_i \in \mathcal{X} \subseteq \mathbb{R}^d$ denote a vector of pre-treatment, time-invariant covariates available for adjustment across groups and eligibility categories. Finally, I write $Y_{i,t}$ to denote the observed outcome of interest, and I will later use the notation $\Delta Y_{i,t} \equiv Y_{i,t} - Y_{i,g-1}$ to denote the long differenced outcome relative to the baseline period (to be defined clearly later).

Following robins_new_1986, I index potential outcomes by treatment timing. Write $Y_{i,t}(g)$ for the potential outcome at time $t$ if unit $i$ is first treated at $g \in \mathcal{G}_{\mathrm{trg}}$, and $Y_{i,t}(\infty)$ for the potential outcome under no treatment. By convention, I set $Y_{i,t}(g) \equiv Y_{i,t}(\infty)$ for all $t < g$, so that treatment-timing indexing is active only at and after first exposure. The observed outcome satisfies

equation[equation omitted — 232 chars of source]

so that units that actually got treated ($S_i = g$, $Q_i = 1$) realize $Y_{i,t}(g)$ while all remaining units---including ineligible units in treated groups and every unit in the never-treated group---realize $Y_{i,t}(\infty)$. This formulation embeds two maintained restrictions: treatment effects operate through the timing of onset, and ineligible units within treated groups are unaffected by the policy. The observed data for unit $i$ are therefore

equation[equation omitted — 112 chars of source]

and $\{W_i\}_{i=1}^n$ is taken to be an i.i.d.\ sample from the population distribution of $W$.

DDD is ubiquitous in practice; I provide a few examples below.

ex[ACA Medicaid expansion] The Affordable Care Act (ACA) expanded Medicaid eligibility to adults with incomes below 138% of the Federal Poverty Level, with states adopting the expansion at different times beginning in 2014 courtemanche_early_2017, kaestner_effects_2017. In this application, $S_i$ records the year in which state $i$'s Medicaid expansion took effect, with $S_i = \infty$ for states that had not expanded by the end of the sample period. Eligibility is defined at the individual level: $Q_i = 1$ for adults with incomes below 138% FPL (the newly eligible population) and $Q_i = 0$ for higher-income adults in the same state who are ineligible for Medicaid. Outcomes include health insurance coverage rates, emergency department visits, and health expenditures. The three differences are: (i) pre- versus post-expansion, (ii) expansion versus non-expansion states, and (iii) income-eligible versus income-ineligible populations.
ex[WTO accession and trade] Countries have acceded to the WTO/GATT at different times since 1948, creating a staggered adoption setting for trade liberalization strezhnev_decomposing_2023. In this application the unit $i$ is an ordered exporter--importer directed pair $i = (e, m)$, with $e$ the exporting country and $m$ the importing country, so that the pair identity is fixed over the sample. $S_i$ records the year of country $i$'s WTO accession, with $S_i = \infty$ for non-member countries. Eligibility $Q_i \in \{0,1\}$ is an attribute of the directed pair: $Q_i = 1$ if the importer $m$ is also a WTO member throughout the sample window (so the exchange qualifies for Most Favored Nation treatment), and $Q_i = 0$ otherwise. Because both the exporter's WTO accession year $S_i$ and the importer's membership status are time-invariant characteristics of the directed pair, $Q_i$ is time-invariant as required by (ref). The outcome $Y_{i,t}$ is the log of bilateral trade flows. The three differences are: (i) pre- versus post-accession, (ii) new-member versus non-member countries, and (iii) eligible (both-member) versus ineligible (one-member) trading pairs.
ex[EPA nonattainment designations] Under the Clean Air Act, the U.S.\ Environmental Protection Agency (EPA) designates counties as “nonattainment” when they fail to meet national ambient air quality standards, with designations occurring at different times for different pollutants and counties greenstone_impacts_2002, walker_transitioning_2013. In this application, $S_i$ records the year of county $i$'s nonattainment designation, with $S_i = \infty$ for counties that remained in attainment throughout the sample period. Eligibility is defined at the industry level: $Q_i = 1$ for polluting manufacturing industries subject to heightened regulatory scrutiny under nonattainment, and $Q_i = 0$ for non-polluting industries in the same county. Outcomes include manufacturing employment, plant openings and closings, and total factor productivity. The three differences are: (i) pre- versus post-designation, (ii) nonattainment versus attainment counties, and (iii) regulated versus unregulated industries.

Identification Assumptions

I impose four assumptions in the ensuing identification analyses. The first three assumptions are standard in the difference-in-differences literature. The fourth is specific to the DDD setting.

assumption[Random Sampling] The observed data $\{W_i\}_{i=1}^n$ are independent and identically distributed: \begin{equation} \{W_i\}_{i=1}^n \stackrel{\overset{\mathrm{i.i.d.}}{\sim}}{\sim} F_W, \end{equation} where $W_i = (\{Y_{i,t}\}_{t=1}^T, S_i, Q_i, X_i) \in \mathbb{R}^T \times \mathcal{S} \times \{0,1\} \times \mathcal{X}$ and $F_W$ is the population distribution with $\mathbb{E}\|W_i\|^2 < \infty$.

(ref) requires that the observed data consist of an i.i.d.\ draw from the joint distribution of outcomes across all time periods, the treatment-enabling group indicator $S_i$, the eligibility indicator $Q_i$, and pre-treatment covariates $X_i$. This assumption is standard in the panel data literature and is automatically satisfied when units are sampled randomly from the population of interest. It rules out spatial or network dependence across units but can be relaxed to allow for cluster-level dependence when inference is conducted at the cluster level.

assumption[Overlap] Every cell in the group-by-eligibility partition has positive probability: \begin{equation} \min_{\substack{s \in \mathcal{S},\\ q \in \{0,1\}}} \mathbb{P}(S_i = s, Q_i = q) > 0 . \end{equation}

(ref) ensures that every cell in the $S \times Q$ partition of units is populated. This is the DDD analogue of the standard overlap (or positivity) condition in the treatment effects literature, stated unconditionally. It requires that no enabling group or eligibility status has zero probability in the population. When covariates are incorporated (Appendix (ref)), the overlap condition is strengthened to a conditional version requiring that the generalized propensity scores $\mathbb{P}(S_i = s, Q_i = q \mid X_i = x)$ are bounded away from zero uniformly over the covariate support.

assumption[No Anticipation] For all $g \in \mathcal{G}_{\mathrm{trg}}$ and $t < g$: \begin{equation} Y_{i,t}(g) = Y_{i,t}(\infty) \quad almost surely. \end{equation}

(ref) requires that in periods before treatment is enabled for cohort $g$, the potential outcomes under treatment and under no treatment coincide. That is, units do not alter their behavior in anticipation of future treatment. This assumption is standard in the difference-in-differences literature callaway_difference--differences_2021-1 and is plausible in settings where the exact timing of treatment enablement is not known in advance or where institutional constraints prevent anticipatory behavior.

assumption[DDD Parallel Changes-in-Trends (DDD-PCT)] For all $g \in \mathcal{G}_{\mathrm{trg}}$, all valid comparison groups $g_c > g$, and all $t \in \{2, \ldots, T\}$ with $t \leq g_c$: \begin{align} &\mathbb{E}[Y_{i,t}(\infty) - Y_{i,t-1}(\infty) \mid S_i = g, Q_i = 1] - \mathbb{E}[Y_{i,t}(\infty) - Y_{i,t-1}(\infty) \mid S_i = g, Q_i = 0] \notag \\ &= \mathbb{E}[Y_{i,t}(\infty) - Y_{i,t-1}(\infty) \mid S_i = g_c, Q_i = 1] - \mathbb{E}[Y_{i,t}(\infty) - Y_{i,t-1}(\infty) \mid S_i = g_c, Q_i = 0] . \end{align}
assumption[No Spillover] For every $g \in \mathcal{G}_{\mathrm{trg}}$ and every unit $i$ with $S_i = g$ and $Q_i = 0$, and for all $t \in \{1,\ldots,T\}$, \begin{equation} Y_{i,t} = Y_{i,t}(\infty) \quad almost surely, \end{equation} regardless of the treatment status of eligible units in the same enabling group.
assumption[Admissibility of Comparison Cohorts] For every treated cohort $g \in \mathcal{G}_{\mathrm{trg}}$, the set of admissible comparison cohorts \begin{equation} \mathcal{C}(g) \equiv \bigl\{ g_c \in \mathcal{S} : g_c > g + K and \mathbb{P}(S_i = s, Q_i = q, G_i = g_c)>0 for all (s,q)\in\{0,1\}^2 \bigr\} \end{equation} is non-empty.

(ref) is the core identifying assumption of the stacked DDD framework. It requires that the {difference} in untreated outcome trends between eligible ($Q_i = 1$) and ineligible ($Q_i = 0$) units is the same across the treated cohort $g$ and its comparison cohort $g_c$. This is a weaker condition than the parallel trends assumptions required in standard DiD settings. In particular, (ref) permits {group-specific trends} $\mathbb{E}[Y_{i,t}(\infty) - Y_{i,t-1}(\infty) \mid S_i = g]$ to differ freely across enabling groups $g$, which is precisely the type of heterogeneity that invalidates parallel trends in DiD settings. It also permits {eligibility-specific trends} to differ between eligible and ineligible units within the same treatment-enabling group. What is required is that this within-group differential trend be the {same} across groups.

(ref) as stated is already pairwise: it is indexed by a single treated cohort $g$ and a single comparison cohort $g_c$, and the identifying restriction is imposed one pair at a time. A pooled DDD regression compares eligible units against a single fixed comparison pool common to every treated cohort, which amounts to requiring (ref) to hold simultaneously across every $(g, g_c)$ pair the pooled estimator mixes; that is the joint restriction the stacked design relaxes. The stacked estimator, by contrast, invokes (ref) only for the $(g, g_c(g))$ pairs actually used to build each stack, and different stacks may use different $g_c$.

(ref) is the SUTVA-type restriction that makes the within-group comparison of eligible and ineligible units a clean third difference. It fails under within-group spillovers (e.g. general equilibrium effects) that would alter the outcomes of ineligibles once eligibles are treated. When this assumption is suspect, the stacked DDD design must either be reframed around a comparison group that is unambiguously outside the spillover's reach, or supplemented with a spillover-robust correction such as partial-population designs.

Finally, (ref) states that for every treated cohort $g$, there is at least one cohort $g_c$ that (i) remains untreated throughout the event window $[g-L, g+K]$, and (ii) has positive mass in every $(S,Q)$-cell, and the combination of the two conditions guarantees that DDD within stack $g$ is well-defined. When there are never-treated groups, condition (i) holds automatically.

remark[Conditional version] When pre-treatment covariates $X_i$ are available, (ref) can be strengthened to hold conditional on $X_i = x$ for almost all $x \in \mathcal{X}$. This extension is discussed in (ref).

\@startsection{section}{2} \z@{.5\linespacing\@plus.7\linespacing}{-.5em} {\normalfont}{The Practice of Using Triple-Difference Event-Study Specification}

hansen_national_2023 study the impact of genetically modified (GM) crop adoption on agricultural yields using a triple-differences design with staggered adoption across countries and crops. Their event-study specification is a leading example of a pooled DDD regression.

Mapping to the DDD Framework

In HW, the observational unit is a country--crop pair $(i,c)$, observed over years $t = 1, \ldots, T$. $Y_{ict}$ denotes an outcome of interest---for example, log agricultural yield. The event-study specification they use is

equation[equation omitted — 193 chars of source]

where $\delta_{it}$ are country-by-year fixed effects, $\gamma_{ci}$ are crop-by-country fixed effects, $\lambda_{ct}$ are crop-by-year fixed effects, $E_{ic}$ is the year when the first GM varieties of crop $c$ are harvested in country $i$, and the $\alpha_j$ are the event-study parameters of interest. The normalization is $\alpha_{-1} = 0$, anchoring the event study at the last pre-treatment period.

This specification maps into the framework of Section (ref) as follows. Relabel the country--crop pair as the “unit,” writing $i$ for a generic unit. Define the treatment-enabling group as the GM approval year, $S_i = E_{ic}$, and the eligibility indicator as $Q_i = \mathbb{1}\{\text{crop } c \text{ is a GM-eligible variety}\}$. The treatment indicator is $D_{i,t} = \mathbb{1}\{t \geq S_i\} Q_i$, which turns on when the country has adopted GM for that crop {and} the crop is GM-eligible.

Under this relabeling, two features are worth noting. First, the country-by-year fixed effects $\delta_{it}$ are finer than country-level fixed effects but are non-nested with the group-by-time fixed effects $\delta_{S_i,t}$ along the crop dimension, because the group index $S_i = E_{ic}$ varies within country (i.e. different crops in the same country may have different GM approval years). Country-by-year effects therefore absorb only the country-common piece of $\delta_{S_i,t}$; the residual group-by-time variation that differs across crops within a country is left unabsorbed. Second, the crop-by-year fixed effects $\lambda_{ct}$ are strictly finer than (and therefore nest the variation exploited by) the binary eligibility-by-time effects $\eta_{Q_i,t}$, since crop identity is more granular than the binary eligible/ineligible classification. Therefore, the HW specification in (ref) therefore includes {all three} sets of two-way interactions that the correct DDD specification requires. However, despite this correct structure, the specification pools the event-time indicators $\mathbb{1}\{t - E_{ic} = j\}$ across all treatment cohorts, thereby imposing a common event-study path $\alpha_j$ across cohorts potentially with heterogeneous treatment effects.

In the notation of this paper, the HW specification (ref) is the event-study version of the correct DDD regression (ref)

equation[equation omitted — 181 chars of source]

where $R_j(i,t) = \mathbb{1}\{t - S_i = j\} Q_i$ is the DDD event-time indicator and I have absorbed the finer country-by-year and crop-by-year effects into their coarser counterparts for notational economy. The critical difference from the 3WFE event-study specification (ref) is the inclusion of $\eta_{Q_i,t}$. I now characterize the estimand $\alpha_j$.

Estimand of the Event-Study Specification (ref)

For any variable $Z_{i,t}$, the residual after projecting out unit effects $\alpha_i$, group-by-time effects $\delta_{S_i,t}$, and eligibility-by-time effects $\eta_{Q_i,t}$ is

equation[equation omitted — 241 chars of source]

where the barred quantities are group means as defined in Section (ref). This is the inclusion-exclusion formula for three-way demeaning, to be stated later in (ref). For each cohort $g \in \mathcal{G}_{\mathrm{trg}}$ and event-time $\ell$, define the cohort-specific event-time indicator $R_{g,\ell}(i,t) = \mathbb{1}\{S_i = g, Q_i = 1, t = g + \ell\}$ and the {modified auxiliary regression}

equation[equation omitted — 201 chars of source]

The coefficient $\omega_{g,\ell}^{e,\star}$ measures how much of the variation in the cohort-specific indicator $R_{g,\ell}$ is captured by the aggregate event-time-$e$ indicator $R_e$ after partialling out all three sets of fixed effects. These weights are estimable from the data without any assumptions on the outcome model. The following results are described in the following object. Following sun_estimating_2021, I define the cohort-average treatment on the treated at event-time $e$ for cohort $g$ as

equation[equation omitted — 138 chars of source]

The CATT parameterization indexes treatment effects by (cohort, exposure duration) rather than (cohort, time), which is natural for studying how effects evolve with time since treatment onset, and the HW event-study specification (ref) is parameterized in event-time.

prop[Weights from auxiliary regression] The coefficients $\omega_{g,\ell}^{e,\star}$ from the auxiliary regression (ref) have the explicit form \begin{equation} \omega_{g,\ell}^{j,\star} = \mathbf{e}_j^{\top} \left(\sum_{t=1}^{T} \mathbb{E} \left[\ddot{\mathbf{R}}_{i,t} \ddot{\mathbf{R}}_{i,t}^{\top}\right]\right)^{-1} \mathbb{E}\left[\ddot{\mathbf{R}}_{i,g+\ell} R_{g,\ell}(i, g+\ell)\right] , \end{equation} where $\ddot{\mathbf{R}}_{i,t} = (\ddot{R}_e(i,t))_{e \neq -1}$ is the vector of three-way-demeaned event-time indicators, $\mathbf{e}_j$ is the unit vector selecting event-time $j$, and $\ddot{\cdot}$ denotes the three-way demeaning operator (ref). These weights satisfy the following properties: \begin{enumerate} • own-period weights sum to one, $\sum_{g \in \mathcal{G}_{\mathrm{trg}}} \omega_{g,j}^{j,\star} = 1$ ; • other included periods sum to zero, $\sum_{g \in \mathcal{G}_{\mathrm{trg}}} \omega_{g,\ell}^{j,\star} = 0$ for each $\ell \neq j$, $\ell \neq -1$ ; • excluded period sums to negative one, $\sum_{g \in \mathcal{G}_{\mathrm{trg}}} \omega_{g,-1}^{j,\star} = -1$ ; • never-treated units receive zero weight, $\omega_{\infty,\ell}^{j,\star} = 0$ for all $j, \ell$ . \end{enumerate}

(ref) follows from a direct application of the Frisch--Waugh--Lovell (FWL) theorem to (ref): $\omega_{g,\ell}^{j,\star}$ is the coefficient on the aggregate event-time-$j$ indicator $R_j$ when $R_{g,\ell}$ is regressed on the full vector of three-way-demeaned event-time indicators $\ddot{\mathbf{R}}_{i,t}$, so each weight is a linear functional of second moments of the $\ddot{R}_e$. No outcome model is invoked, and property {(iv)} is immediate: a never-treated unit has $R_\ell(i,t) \equiv 0$ for all $\ell$, so $\ddot{R}_\ell(i,t)$ aggregates to zero contribution in the numerator of (ref) at that cell.

remark[{Variance-ratio representation of the weights}] The matrix-form weight (ref) can be expressed as a transparent ratio of variances using a second application of the FWL theorem. By FWL applied {within} the auxiliary regression (ref), the coefficient $\omega_{g,\ell}^{j,\star}$ on $R_j$ equals the bivariate regression coefficient obtained by first partialling out all other regressors---the three sets of fixed effects and all event-time indicators $R_{e'}$ for $e' \neq j$, $e' \neq -1$---from both the dependent variable $R_{g,\ell}$ and the regressor of interest $R_j$. Let $\widetilde{R}_j(i,t)$ denote the partial residual of $R_j(i,t)$ after projecting out $(\alpha_i, \delta_{S_i,t}, \eta_{Q_i,t})$ and all other event-time indicators $\{R_{e'}\}_{e' \neq j, e' \neq -1}$. Then, \begin{equation} \omega_{g,\ell}^{j,\star} = \frac{\sum_{t=1}^{T} \mathbb{E}\left[\widetilde{R}_j(i,t) R_{g,\ell}(i,t)\right]}{\sum_{t=1}^{T} \mathbb{E}\left[\widetilde{R}_j(i,t)^2\right]} . \end{equation} Since $R_{g,\ell}(i,t) = \mathbb{1}\{S_i = g, Q_i = 1, t = g + \ell\}$ is nonzero only at time $t = g + \ell$, the numerator collapses to a single time period \begin{equation} \sigma_{j; g,\ell}^{\star} \equiv \sum_{t=1}^{T} \mathbb{E}\left[\widetilde{R}_j(i,t) R_{g,\ell}(i,t)\right] = \mathbb{E}\left[\widetilde{R}_j(i, g+\ell) \mathbb{1}\{S_i = g, Q_i = 1\}\right] = p_{g,1} \widetilde{r}_{j; g,1}^{(g+\ell)} , \end{equation} where $p_{g,1} = \mathbb{P}(S_i = g, Q_i = 1)$ is the population share of treated-eligible units in cohort $g$ and $\widetilde{r}_{j; g,1}^{(g+\ell)}$ is the value of the partial residual $\widetilde{R}_j$ evaluated at cell $(S_i = g, Q_i = 1)$ at time $t = g + \ell$. (The partial residual is constant within cells at a given time, since both $R_j$ and all projecting-out variables are cell-level indicators.) The denominator is the {partial residual variance} \begin{equation} \sigma_j^{2,\star} \equiv \sum_{t=1}^{T} \mathbb{E}\left[\widetilde{R}_j(i,t)^2\right] = \sum_{t=1}^{T} \sum_{s \in \mathcal{S}} \sum_{q \in \{0,1\}} p_{s,q} \left(\widetilde{r}_{j; s,q}^{(t)}\right)^2 , \end{equation} the total variation in the event-time-$j$ indicator that is {not} explained by the three sets of fixed effects or by any other event-time indicator. Therefore, \begin{equation} {\omega_{g,\ell}^{j,\star} = \frac{p_{g,1}\widetilde{r}_{j; g,1}^{(g+\ell)}}{\sigma_j^{2,\star}} } \end{equation} is the explicit form of the weight.
remark[{Why the partial residual can be negative}] The partial residual $\widetilde{R}_j(i,t)$ is obtained from the aggregate indicator $R_j(i,t) = \mathbb{1}\{t - S_i = j\} Q_i$ by two successive projections. First, the three-way demeaning (ref) removes unit effects, group-by-time effects, and eligibility-by-time effects, yielding the three-way-demeaned indicator $\ddot{R}_j(i,t)$. Second, the projection onto the space spanned by the other demeaned event-time indicators $\{\ddot{R}_{e'}\}_{e' \neq j, -1}$ is subtracted, yielding $\widetilde{R}_j(i,t) = \ddot{R}_j(i,t) - \sum_{e' \neq j, -1} \widehat{\gamma}_{j,e'} \ddot{R}_{e'}(i,t)$, where $\widehat{\gamma}_{j,e'}$ is the population regression coefficient from the projection. Consider a treated-eligible unit in cohort $g$ at time $t = g + \ell$ with $\ell \neq j$. At this time, the unit is at event-time $\ell$ relative to its own treatment, not at event-time $j$. The aggregate indicator $R_j(i, g+\ell)$ is zero for this unit (since $g + \ell - g = \ell \neq j$). After three-way demeaning, $\ddot{R}_j(i, g+\ell)$ can be nonzero because the demeaning subtracts group-time and eligibility-time means that are contaminated by other cohorts. After further partialling out other event-time indicators, $\widetilde{R}_j(i, g+\ell)$ can take either sign, depending on the correlation structure between $\ddot{R}_j$ and $\ddot{R}_{e'}$ at that time. The sign of $\widetilde{r}_{j; g,1}^{(g+\ell)}$ is driven by the {calendar-time composition} of the design. Under staggered adoption, at any given time $t$, different cohorts occupy different positions on the event-time axis. The indicator $R_j(i,t)$ lights up for the cohort $g_t$ satisfying $g_t + j = t$ (if such a cohort exists). After demeaning and partialling out, the residual at this time reflects the deviation of cohort $g_t$'s eligible share from the average eligible share across all cohorts that are “active” at time $t$. When this deviation is offset by the partialling-out step, the residual at a {different} cohort's cell can flip sign.
ex[{Explicit weights in a toy example}] {Consider a balanced panel with $T = 4$ and two treatment cohorts $g_1 = 2$, $g_2 = 3$ (no never-treated units), with equal cohort shares ($\pi_1 = \pi_2 = 1/2$) and equal eligibility shares ($p_1 = p_2 = p$). Both cohorts have a proper pre-period at $t = 1$, so the $e = -1$ normalization is well-defined for each.} The event-study specification includes only $e = 0$ and excludes $e = -1$ ($L = 1$, $K = 0$). I show below that \begin{equation} \alpha_0 = \frac{1}{2}\mathrm{CATT}(g_1, 0) + \frac{1}{2}\mathrm{CATT}(g_2, 0) - \left( \frac{1}{2}\mathrm{CATT}(g_1, 1) + \frac{1}{2}\mathrm{CATT}(g_2, -1) \right) . \end{equation} Under no anticipation ($\mathrm{CATT}(g_2, -1) = 0$), (ref) collapses to \[ \alpha_0 = \frac{1}{2}\mathrm{CATT}(g_1, 0) + \frac{1}{2}\mathrm{CATT}(g_2, 0) - \frac{1}{2}\mathrm{CATT}(g_1, 1)~. \] To see this, notice that since only $R_0$ is included, $\widetilde{R}_0 = \ddot{R}_0$, and (ref) reduces to \begin{equation} \omega_{g,\ell}^{0,\star} = \frac{p_{g,1} \ddot{r}_{0; g,1}^{(g+\ell)}}{\sigma_0^{2,\star}} , \end{equation} with $p_{g,1} = \mathbb{P}(S_i = g, Q_i = 1) = p/2$ under the parameters stated earlier. The indicator $R_0(i,t)$ equals $Q_i$ for cohort $g_1$ at $t = 2$ and for cohort $g_2$ at $t = 3$, and is zero otherwise. The relevant marginal shares of $R_0$ needed to form the three-way-demeaned residual are: the unit-level mean is $1/4$ for treated-eligible cells in either cohort and $0$ elsewhere; the group-by-time mean $\bar R_{0,g,t}$ equals $p$ at $(g_1, 2)$ and $(g_2, 3)$ and zero elsewhere; the eligibility-by-time mean $\bar R_{0,Q=1,t}$ equals $1/2$ at $t \in \{2, 3\}$ and zero elsewhere; the group all-time mean is $p/4$ for each group; the eligibility all-time mean is $1/4$ for $Q = 1$; the time-only mean is $p/2$ at $t \in \{2, 3\}$; and the total mean is $p/4$. Computing the three-way-demeaned residual for a treated-eligible unit in cohort $g_1$ at its own event-time $t = g_1 = 2$, \begin{align} \ddot{R}_0(g_1, 1, 2) &= \underbracket{1}_{R_0} - \underbracket{1/4}_{\overline{R}_{0,(g_1,1),\cdot}} - \underbracket{p}_{\overline{R}_{0,g_1,2}} - \underbracket{1/2}_{\overline{R}_{0,Q=1,2}} + \underbracket{p/4}_{\overline{R}_{0,g_1,\cdot}} + \underbracket{1/4}_{\overline{R}_{0,Q=1,\cdot}} + \underbracket{p/2}_{\overline{R}_{0,\cdot,2}} - \underbracket{p/4}_{\overline{R}_{0,\cdot\cdot}} = \frac{1-p}{2} , \end{align} and by symmetry $\ddot{r}_{0; g_2,1}^{(3)} = (1-p)/2$ at cohort $g_2$'s own event-time. The own-period weights are therefore equal across cohorts, and property {(i)} of Proposition (ref) forces each to equal $1/2$. Now turn to the cross-period cell: at $t = 3$, cohort $g_1$ is at event-time $\ell = 1$ (one period after its treatment), and the analogous calculation gives \begin{equation} \ddot{r}_{0; g_1,1}^{(3)} = {0 - \frac{1}{4} - 0 - \frac{1}{2} + \frac{p}{4} + \frac{1}{4} + \frac{p}{2} - \frac{p}{4}} = -\frac{1-p}{2} , \end{equation} exactly the negative of the own-period residual. {An identical calculation at $t = 2$ for cohort $g_2$ (its pre-period $\ell = -1$) yields $\ddot{r}_{0; g_2,1}^{(2)} = -(1-p)/2$. The partial-residual variance is $\sigma_0^{2,\star} = p(1-p)/2$ (summing over the four nonzero cells in each of the eligible and ineligible strata), so (ref) gives $\omega_{g_1,1}^{0,\star} = \omega_{g_2,-1}^{0,\star} = -1/2$, which yields (ref).} The cross-period weights have the same magnitude as the own-period weights but opposite sign. Under dynamic effects with $\mathrm{CATT}(g_1, 1) > 0$ and no anticipation, $\alpha_0$ is biased {downward} relative to the average of the two own-period CATTs, with the bias growing in $\mathrm{CATT}(g_1,1)$; only under event-time homogeneity $\mathrm{CATT}(g,\ell) = \mathrm{ATT}_\ell$ does the contamination vanish, as formalized in Proposition (ref) below.

Estimand Under Additional Assumptions

I now add identifying assumptions progressively and analyze the estimands.

prop[Under DDD-PCT only] Under {(ref)}, the population regression coefficient $\alpha_j$ is a linear combination of $\mathrm{CATT}(g,\ell)$ with the weights from Proposition {(ref)} \begin{equation} \alpha_j = \sum_{g \in \mathcal{G}_{\mathrm{trg}}}\sum_{\ell \neq -1} \omega_{g,\ell}^{j,\star} \mathrm{CATT}(g,\ell) . \end{equation} Cross-period contamination {(}nonzero weights $\omega_{g,\ell}^{j,\star}$ for $\ell \neq j${)} and cross-cohort contamination {(}possibly negative own-period weights $\omega_{g,j}^{j,\star}$ for individual $g${)} are both present.

(ref) says that when we estimate the HW event-study specification (ref) on staggered-adoption data, the coefficient $\alpha_j$ in a linear combination of cohort-level CATTs running over {all} included event-times, not only over time $j$. This is property (i) from (ref). This complicates interpretation of the estimand absent additional assumptions.

prop[Under DDD-PCT and no anticipation] Under {(ref)} and {(ref)}, the population coefficient $\alpha_j$ satisfies \begin{equation} \alpha_j = \sum_{g \in \mathcal{G}_{\mathrm{trg}}}\sum_{\ell \geq 0} \omega_{g,\ell}^{j,\star} \mathrm{CATT}(g,\ell) . \end{equation} $\mathrm{CATT}(g,\ell) = 0$ for $\ell < 0$, but for $j < 0$, $\alpha_j$ is generically nonzero because it depends on post-treatment $\mathrm{CATT}(g,\ell)$ for $\ell \geq 0$ through the weights $\omega_{g,\ell}^{j,\star}$.

(ref) has a striking implication for applied work. A researcher using the HW specification who finds $\widehat{\alpha}_j \neq 0$ for $j < 0$ may incorrectly conclude that DDD PCT is violated, when in fact the pre-period coefficient reflects heterogeneous post-treatment effects contaminating the pre-treatment window through the implicit weights. Conversely, $\widehat{\alpha}_j \approx 0$ for $j < 0$ does not validate DDD-PCT.

prop[Under DDD-PCT and treatment effect homogeneity] Under (ref), {(ref)} and the restriction that $\mathrm{CATT}(g,\ell) = \mathrm{ATT}_\ell$ for all $g$ (treatment effects depend on exposure duration but not on the cohort), the population coefficient $\alpha_j$ simplifies to \begin{equation} \alpha_j = \mathrm{ATT}_j . \end{equation}

(ref) identifies the knife-edge conditions under which the HW event-study specification recovers an interpretable causal parameter: $\alpha_j$ from (ref) requires DDD-PCT, no anticipation, and treatment effect homogeneity across treatment cohorts to admit a valid causal interpretation. Relating this to (ref), note that property {(i)} is reassuring {only} under treatment effect homogeneity in the event-time dimension, i.e., $\mathrm{CATT}(g,j) = \mathrm{ATT}_j$ for all $g$; in that case property {(ii)} guarantees that cross-period CATTs cancel and $\alpha_j = \mathrm{ATT}_j$ exactly. Under heterogeneous treatment effects, property {(ii)} of (ref) says that individual weights $\omega_{g,\ell}^{j,\star}$ for $\ell \neq j$ can be negative (and, as the variance-ratio representation in Remark (ref) makes explicit, own-period weights $\omega_{g,j}^{j,\star}$ can also be negative for particular cohorts), and $\alpha_j$ becomes a {linear}---not convex---combination of CATTs running over both the cohort and event-time axes. The population regression coefficient therefore does not admit the interpretation of any weighted {average} treatment effect at event-time $j$.

The Aggregated ATT

Applied researchers often summarize the event study by computing an aggregated ATT, averaging the post-treatment coefficients with researcher-chosen weights. Define

equation[equation omitted — 117 chars of source]

where $w_j \geq 0$ are non-negative weights with $\sum_{j=0}^{K} w_j = 1$. Common choices are equal weights ($w_j = 1/(K+1)$) or exposure-duration-specific weights. The following proposition characterizes the probability limit of this estimand.

prop[Aggregated ATT estimand] Under {(ref)}--{(ref)}, the aggregated ATT (ref) satisfies \begin{equation} \widehat{\mathrm{ATT}}_{\mathrm{agg}} \overset{\mathbb P}{\to} \sum_{g \in \mathcal{G}_{\mathrm{trg}}} \sum_{\ell \geq 0} \Omega_{g,\ell} \mathrm{CATT}(g,\ell) , \end{equation} where the aggregated weights are \begin{equation} \Omega_{g,\ell} = \sum_{j=0}^{K} w_j \omega_{g,\ell}^{j,\star} . \end{equation} The aggregated weights $\Omega_{g,\ell}$ need not be non-negative, but they satisfy the normalization \begin{equation} \sum_{g \in \mathcal{G}_{\mathrm{trg}}} \sum_{\ell \geq 0} \Omega_{g,\ell} = 1 . \end{equation}

\@startsection{section}{2} \z@{.5\linespacing\@plus.7\linespacing}{-.5em} {\normalfont}{Estimands, and What Does Stacked OLS Identify?}

Before developing the formal identification and estimation theory, I analyze the regression specifications that applied researchers commonly use for triple differences. This section characterizes exactly what each specification targets, under what conditions it recovers an interpretable causal parameter, and where the standard approach breaks down. I begin by deriving the two regression specifications that correctly target the causal estimands of interest, providing rigorous proofs for each claim. I then analyze commonly used specifications that fail to recover these estimands, diagnosing the precise source of each failure. Throughout this section I draw explicit parallels with the interaction-weighted framework of sun_estimating_2021 for difference-in-differences and the decomposition results of de_chaisemartin_two-way_2020. The analysis proceeds without covariates; the covariate-adjusted framework is developed in Appendix (ref).

Stacked sub-experiments

The central construction of the stacked DDD approach is that of the {stacked sub-experiment}. The following definition makes this precise.

defn[Stacked Sub-Experiment] For each treatment cohort $g \in \mathcal{G}_{\mathrm{trg}}$ and comparison group $g_c \in \mathcal{S}$ with $g_c > g$, the $g$-specific stack $\mathbb{S}_g = \mathbb{S}(g, g_c, L, K)$ consists of: \begin{enumerate} • Event window $\{g - L, \ldots, g - 1, g, \ldots, g + K\}$, where $L \geq 1$ pre-treatment and $K \geq 0$ post-treatment periods. • Treated units Units with $S_i = g$ and $Q_i = 1$ (the cohort receiving treatment); • Within-group controls Units with $S_i = g$ and $Q_i = 0$ (same enabling group, ineligible); • Clean comparison group, eligible Units with $S_i = g_c$ and $Q_i = 1$; • Clean comparison group, ineligible Units with $S_i = g_c$ and $Q_i = 0$; \end{enumerate} The clean comparison group satisfies $g_c > g + K$, so that no unit in the comparison group is treated during the time window of the stack.

Each stack $\mathbb{S}_g$ is a self-contained $2 \times 2 \times 2$-like experiment. The four cells are defined by the interaction of two binary dimensions:

equation[equation omitted — 368 chars of source]

Estimands of Interest

I now define the target parameters for the stacked DDD framework, building from cohort-specific treatment effects to aggregated event-study parameters.

\paragraph{Group-time average treatment effect.} The fundamental building block is the group-time average treatment effect on the treated. For each treatment cohort $g \in \mathcal{G}_{\mathrm{trg}}$ and post-treatment period $t \geq g$, define

equation[equation omitted — 128 chars of source]

This is the average causal effect of treatment for units in cohort $g$ who are eligible ($Q_i = 1$), measured at time $t$. The parameter $\mathrm{ATT}(g,t)$ is indexed by both the treatment cohort and time, and thus accommodates heterogeneity in treatment effects across cohorts and over time. This quantity is the DDD analogue of the group-time ATT defined by callaway_difference--differences_2021-1 in the standard DiD setting. Finally, note that $\mathrm{CATT}(g, e) = \mathrm{ATT}(g, g+e)$ under the identity $t = g + e$, so the two parameters refer to the same causal contrast.

remark[Heterogeneous treatment effects] The parameter $\mathrm{ATT}(g,t)$ averages over the joint distribution of individual treatment effects for treated-eligible units in cohort $g$. No assumption of constant treatment effects within a cohort-time cell is imposed. Individual effects $Y_{i,t}(g) - Y_{i,t}(\infty)$ may vary arbitrarily across units sharing the same $(g, t, X_i)$: the framework permits unrestricted within-cohort heterogeneity. This stands in contrast to the three-way fixed effects regression (ref) in (ref), which imposes a single coefficient $\theta$ and thereby requires treatment effect homogeneity across all $(g,t)$ pairs.

Within a given stack $\mathbb{S}_g = \mathbb{S}(g, g_c, L, K)$, I define the stack-specific treatment effect

equation[equation omitted — 77 chars of source]

which is the DDD estimand formed from the four cells in (ref) within the stack's time window. This is the difference-in-difference-in-differences contrast:

multline[multline omitted — 348 chars of source]

where $\Delta Y_{i,t} = Y_{i,t} - Y_{i,g-1}$ is the outcome change relative to the last pre-treatment period.

Applied researchers typically summarize treatment effects along the event-time axis. For event-time $e = t - g$ (periods relative to treatment onset), define the population event-study parameter

equation[equation omitted — 178 chars of source]

where $\omega_g(e) \geq 0$ are aggregation weights satisfying $\sum_g \omega_g(e) = 1$ for each $e$. Different weight choices yield different summary parameters.

The stacked event-study estimand aggregates stack-level treatment effect estimates across treated cohorts, which is defined as

equation[equation omitted — 213 chars of source]

where $g_c(g)$ denotes the comparison group used for cohort $g$ and the summation ranges over cohorts for which event-time $e$ falls within the stack's window. The constraint $-L \leq e \leq K$ restricts attention to event-times within the symmetric window shared by all stacks.

remark[Regression interpretation] The DDD estimand (ref) has an exact regression counterpart. Within each stack $\mathbb{S}_g$, the fully saturated OLS regression of $\Delta Y_{i,t}$ on the four cell indicators---intercept, group indicator $\mathbb{1}\{S_i = g\}$, eligibility indicator $\mathbb{1}\{Q_i = 1\}$, and their interaction $\mathbb{1}\{S_i = g, Q_i = 1\}$---recovers the sample triple difference as the coefficient on the interaction term. This equivalence means that applied researchers can implement the stacked DDD estimator by running a separate OLS regression within each stack and aggregating the resulting coefficients. The key requirement is that each stack-level regression must be fully saturated in the cell indicators; imposing common slopes across cells or pooling across stacks with a single treatment coefficient alters the target estimand, as I explore in Appendix (ref).

The Saturated Within-Stack Regression

Consider a single stack $\mathbb{S}_g$ with four cells defined by $(S_i, Q_i) \in \{(g,1), (g,0), (g_c,1), (g_c,0)\}$. Write $n_{s,q} = \sum_{i \in \mathbb{S}_g} \mathbb{1}\{S_i = s, Q_i = q\}$ for the number of units in cell $(s,q)$, and $n_g = \sum_{s,q} n_{s,q}$ for the total number of units in the stack. For a fixed post-treatment period $t \geq g$, the fully saturated regression on long differences $\Delta Y_{i,t}$ is

equation[equation omitted — 217 chars of source]

The coefficients $(\mu_{g,t}, \lambda_{g,t}, \eta_{g,t}, \tau_{g,t}^{\mathrm{sat}})$ are subscripted by $(g,t)$ to emphasize that they are population parameters specific to the stack built around cohort $g$ evaluated at time $t$. The regression has four parameters for four cell means and is exactly identified. The following proposition establishes that the OLS interaction coefficient recovers the sample triple difference.

prop[Saturated regression recovers the triple difference] In regression (ref), the OLS coefficients are \begin{align} \widehat{\mu}_{g,t} &= \overline{\Delta Y}_{g_c,0,t} , \\ \widehat{\lambda}_{g,t} &= \overline{\Delta Y}_{g,0,t} - \overline{\Delta Y}_{g_c,0,t} , \\ \widehat{\eta}_{g,t} &= \overline{\Delta Y}_{g_c,1,t} - \overline{\Delta Y}_{g_c,0,t} , \\ \widehat{\tau}_{g,t}^{\mathrm{sat}} &= \underbracket{\left(\overline{\Delta Y}_{g,1,t} - \overline{\Delta Y}_{g,0,t}\right)}_{within-group DD} - \underbracket{\left(\overline{\Delta Y}_{g_c,1,t} - \overline{\Delta Y}_{g_c,0,t}\right)}_{across-group DD} , \end{align} where $\overline{\Delta Y}_{s,q,t} = n_{s,q}^{-1}\sum_{i\colon S_i = s, Q_i = q} \Delta Y_{i,t}$ is the cell-specific sample mean. Under (ref)--(ref), $\widehat{\tau}_{g,t}^{\mathrm{sat}} \overset{\mathbb P}{\to} \mathrm{ATT}(g,t)$.

The saturated regression (ref) requires no functional form assumptions whatsoever. It is purely design-based; the four-cell structure of the stack, combined with the long-differencing that removes time-invariant heterogeneity, delivers the triple difference mechanically. Each coefficient has a transparent interpretation. The intercept $\mu_{g,t}$ is the mean outcome change for comparison-ineligible units, $\lambda_{g,t}$ is the group differential for ineligible units, $\eta_{g,t}$ is the eligibility premium in the comparison group, and $\tau_{g,t}^{\mathrm{sat}}$ is the excess eligibility premium in the treated group---the triple difference. This parallels the two-cell structure of the saturated DiD regression, where $\tau$ captures the excess treatment-group change relative to the control; here, the additional eligibility dimension requires four cells instead of two, but the logic is identical.

Pooled Stacked Regressions and Their Targets

The within-stack saturated regression (ref) can be pooled across stacks. I first analyze the fully unrestricted pooled regression, which correctly targets the cohort-time treatment effects, and then examine restricted specifications that impose various degrees of effect homogeneity.

{The unrestricted pooled regression.} Pool all stacks into a single dataset, where each observation $(i, t, g)$ belongs to stack $\mathbb{S}_g$ with $S_i \in \{g, g_c(g)\}$ and $t \in \{g - L, \ldots, g + K\}$. The fully unrestricted regression with stack fixed effects and stack-by-cell fixed effects is

equation[equation omitted — 229 chars of source]

where the subscript $g$ on the left indexes the stack from which the observation is drawn, $\alpha_g$ are stack fixed effects, $\mu_{s,q,g}$ are stack-by-cell fixed effects (with $(s,q) \in \{(g, 1), (g, 0), (g_c, 1), (g_c, 0)\}$ within each stack), and $\tau_{g',e}$ is a cohort-by-event-time coefficient. The summation now extends over all event-times $e \in \{-L, \ldots, K\}$, including the pre-treatment periods $e < 0$. This inclusion is deliberate---the pre-treatment coefficients $\tau_{g',e}$ for $e < 0$ serve as built-in pre-trend diagnostics, as I formalize below.

prop[Unrestricted pooled regression equivalence] Under (ref)--(ref), The fully unrestricted regression (ref) is numerically identical to running the saturated regression (ref) separately within each stack for each time period. Specifically, $\widehat{\tau}_{g',e} = \widehat{\tau}_{g',g'+e}^{\mathrm{sat}}$ for all $g' \in \mathcal{G}_{\mathrm{trg}}$ and $e \in \{-L, \ldots, K\}$.

(ref) establishes that the unrestricted pooled regression (ref) is a “master regression” that simultaneously delivers all within-stack triple differences and all within-stack pre-trend diagnostics. The pre-treatment coefficients provide a regression-based implementation of the pre-trend test---systematic departures of $\widehat{\tau}_{g',e}$ from zero for $e < 0$ signal violations of DDD-PCT.

The researcher forms the event-study parameter via post-estimation aggregation

equation[equation omitted — 165 chars of source]

with explicit, researcher-chosen weights $\widehat{\omega}_g(e) \geq 0$ summing to one. The practical question becomes how to transparently aggregate the cohort-level estimates. The upshot here is that estimating a {fully-saturated} event-study linear regression on the stacked dataset targets a weighted average of cohort-specific effects with strictly positive weights. A stacked event-study regression is fully saturated when it includes a full set of stack-by-group-by-time and stack-by-eligibility-by-time fixed effects, absorbing all cell-specific time shocks. In long differences, this specification is

equation[equation omitted — 206 chars of source]

where $\lambda_{s,t,g}$ is a fixed effect for group $s \in \{g, g_c\}$ at time $t$ in stack $g$, and $\eta_{q,t,g}$ is a fixed effect for eligibility type $q \in \{0, 1\}$ at time $t$ in stack $g$. Because $\lambda_{s,t,g}$ and $\eta_{q,t,g}$ consume exactly three degrees of freedom for the four cells in stack $g$ at time $t$, the addition of the treatment indicator perfectly saturates the four cell means at every period.

The following proposition establishes that this fully-saturated stacked event-study regression is sensible: at each post-treatment event-time $e$, its coefficient is a cell-size-weighted average of the cohort-specific treatment effects with strictly positive weights summing to one. This parallels the result of sun_estimating_2021 for the fully-saturated stacked DiD event-study regression, extended here to the triple-differences setting with the additional eligibility dimension.

prop[Estimands of the fully-saturated event-study regression] In the fully-saturated stacked event-study regression (ref), the OLS coefficient $\widehat{\tau}_e$ for each event-time $e \neq -1$ is numerically identical to a weighted average of the within-stack saturated coefficients $\widehat{\tau}_{g,g+e}^{\mathrm{sat}}$ \begin{equation} \widehat{\tau}_e = \sum_{g \in \mathcal{G}_{\mathrm{trg}}(e)} w_g^{\mathrm{FWL}}(e) \widehat{\tau}_{g,g+e}^{\mathrm{sat}} , \end{equation} where the weights $w_g^{\mathrm{FWL}}(e)$ are strictly positive {under (ref)} and sum to one across $g \in \mathcal{G}_{\mathrm{trg}}(e)$. The weight for stack $g$ is proportional to the FWL residual variance of the treatment indicator within that stack at time $t = g+e$. Under (ref)--(ref) and (ref), \begin{enumerate} • {No cross-cohort contamination.} The weights $w_g^{\mathrm{FWL}}(e) > 0$. • {No cross-period contamination.} For any {fixed} event-time $e$, $\widehat{\tau}_e \overset{\mathbb P}{\to} \sum_{g} w_g^{\mathrm{FWL}}(e) \mathrm{CATT}(g,e)$. Treatment effects from other event-times $e' \neq e$ receive exactly zero weight. \end{enumerate}

(ref) highlights a major advantage of the stacked DDD framework. While researchers are encouraged to compute the unrestricted cohort-specific estimates and explicitly aggregate them using transparent weights (such as cohort-size weights, as in Section (ref)), running the fully-saturated stacked event-study regression provides a safe, simple alternative. It guarantees a sensible weighted average of cohort-level treatment effects, in the sense that it avoids the negative weights phenomenon entirely.

remark[Implementing the stacked DDD as a linear regression] For practitioners, the stacked DDD event-study estimator can be implemented as a single linear regression on the stacked dataset. After constructing the stacked dataset (one observation per unit $\times$ time period $\times$ stack), the regression specification (ref) includes the following fixed effects and treatment indicators. \begin{enumerate} • Stack $\times$ treatment status $\times$ time $\lambda_{s,t,g}$; • Stack $\times$ eligibility $\times$ time fixed effects $\eta_{q,t,g}$; • Event-time treatment indicators. \end{enumerate} Within each stack $g$ at each time $t$, the four cells $(g,1), (g,0), (g_c,1), (g_c,0)$ have three degrees of freedom absorbed by FE1 and FE2 (the group indicator $\mathbb{1}\{S_i = g\}$, the eligibility indicator $\mathbb{1}\{Q_i = 1\}$, and a constant are jointly determined by $\lambda_{s,t,g}$ and $\eta_{q,t,g}$). The treatment indicator adds one more degree of freedom, thereby perfectly saturating the four cell means. This saturation is what ensures the regression recovers the triple difference, and guarantees a valid causal interpretation on the event-study coefficients (on the event-time treatment indicators).

Population estimands under different aggregation schemes

The stack DDD event-study estimator \[\widehat{\mathrm{ES}}_{\mathrm{stack}}(e) = \sum_{g \in \mathcal{G}_{\mathrm{trg}}(e)} \widehat{\omega}_g(e) \widehat{\mathrm{ATT}}_{\mathbb{S}_g}(g, g+e)\] depends on the aggregation weights $\widehat{\omega}_g(e) \geq 0$, $\sum_g \widehat{\omega}_g(e) = 1$. Different weight choices target different population parameters, and the stacked DDD framework makes these targets explicit. I now formally characterize what each weighting scheme identifies when the within-stack estimators are consistent for $\mathrm{ATT}(g, g+e) = \mathrm{CATT}(g,e)$ ((ref)).

Cohort-size weights.

Set $\widehat \omega_g^{\mathrm{cohort}}(e) = n_{g,1} / \sum_{g' \in \mathcal{G}_{\mathrm{trg}}(e)} n_{g',1}$, where $n_{g,1}$ is the number of treated-eligible units in stack $\mathbb{S}_g$ and $\mathcal{G}_{\mathrm{trg}}(e) = \{g \in \mathcal{G}_{\mathrm{trg}} \mid g + e \leq T\}$ is the set of cohorts observed at event-time $e$. The population analog is $\omega_g^{\mathrm{cohort}}(e) = \mathbb{P}(S_i = g \mid S_i \in \mathcal{G}_{\mathrm{trg}}(e), Q_i = 1)$, the share of cohort $g$ among treated-eligible units observed at event-time $e$. The resulting estimand is the per-capita average treatment effect

equation[equation omitted — 259 chars of source]

which averages the cohort-specific effects weighted by the number of treated-eligible individuals in each cohort. This parameter has a direct welfare interpretation as the average effect experienced by a randomly drawn treated-eligible unit at event-time $e$, and coincides with the average effect of switching defined and discussed in (ref). It is the natural estimand when the policy question concerns aggregate impact---how much did the treated population benefit, on average? As de_chaisemartin_difference--differences_2024 emphasize, this per-capita parameter is the most policy-relevant estimand in many applications, since it reflects the actual distribution of treatment effects across the affected population.

Equal weights.

Set $\omega_g^{\mathrm{eq}}(e) = 1/|\mathcal{G}_{\mathrm{trg}}(e)|$, assigning each cohort equal influence regardless of its sample size. The target is the simple average of cohort-specific effects

equation[equation omitted — 182 chars of source]

which treats each policy adoption event---each cohort's experience of treatment---as an equally informative observation about the causal mechanism. This estimand answers the question “what is the average treatment effect across the distinct adoption episodes observed at event-time $e$?”, weighting each cohort's evidence equally. It is appropriate when the research interest lies in the policy mechanism rather than the aggregate impact, or when one cohort is much larger than the others and the researcher does not want a single large cohort to dominate the estimate. Equal weights do not target a population-weighted parameter and should therefore be reported alongside cohort-size weights rather than as a primary specification.

Precision weights.

Set $\omega_g^{\mathrm{prec}}(e) = ({n_g/\sigma_g^2(e)})/({\sum_{g'\in\mathcal{G}_{\mathrm{trg}}(e)} n_{g'}/\sigma_{g'}^2(e)})$, where $\widehat{\sigma}_g^2(e)$ is a consistent estimator of the asymptotic variance of $\widehat{\mathrm{ATT}}_{\mathbb{S}_g}(g, g+e)$. This inverse-variance weighting minimizes the asymptotic variance of $\widehat{\mathrm{ES}}_{\mathrm{stack}}(e)$ when the within-stack estimators are asymptotically independent (an approximation when stacks share comparison units; see Section (ref) for the exact variance formula that accounts for shared controls). The population target of precision weights is

equation[equation omitted — 233 chars of source]

where $\sigma_g^2(e)$ is the asymptotic variance of the within-stack estimator. Unlike cohort-size and equal weights, the precision-weighted estimand depends on the data generating process through the variance terms $\sigma_g^2(e)$ and therefore does not have a fixed causal interpretation that is invariant to the sampling design. In particular, two different samples from the same population will in general target different weighted averages of $\mathrm{CATT}(g,e)$. Precision weights are the statistically optimal choice for minimizing estimation error, but the resulting estimand is harder to interpret substantively. Researchers using precision weights should report the realized weight shares $\widehat{\omega}_g^{\mathrm{prec}}(e)$ alongside point estimates.

General welfare weights.

The three schemes above are special cases of a general welfare-weighted aggregation. Suppose the researcher assigns welfare weights $v_g > 0$ reflecting the relative importance of cohort $g$'s treatment effect. The welfare-weighted event-study parameter is

equation[equation omitted — 187 chars of source]

The stacked DDD framework accommodates any choice of $v_g$ through the aggregation step, a flexibility that enhances interpretability in the case of stacked DDD. Making the welfare weights explicit is a key advantage of the stacked approach, as it separates the statistical problem (consistent estimation of $\mathrm{CATT}(g,e)$) from the aggregation problem (which cohorts matter more for the policy question at hand). This transparency echoes the recommendation of de_chaisemartin_two-way_2020 and de_chaisemartin_difference--differences_2024 that researchers should always report the weights underlying their treatment effect aggregation.

Connection to Existing Estimators

The interaction-weighted DDD estimator.

The stacked DDD is the natural triple-differences analogue of sun_estimating_2021's (sun_estimating_2021) interaction-weighted (IW) estimator for DiD. This connection reveals both the estimand targeted by the stacked estimator and the source of contamination in conventional 3WFE event-study specifications. The event-study parameter is

equation[equation omitted — 145 chars of source]

where $\mathcal{G}_{\mathrm{trg}}(e) = \{g \in \mathcal{G}_{\mathrm{trg}} \mid g + e \leq T\}$ is the set of cohorts observed for at least $e$ post-treatment periods.

I define the IW-DDD estimator as

equation[equation omitted — 173 chars of source]

where $\widehat{\mathrm{CATT}}(g,e)$ is any consistent estimator of $\mathrm{CATT}(g,e)$ and $\widehat{\omega}_g(e)$ are non-negative weights summing to one. When $\widehat{\mathrm{CATT}}(g,e)$ is the within-stack estimator and the weights coincide with those used in the stacked aggregation, the IW-DDD estimator is numerically identical to $\widehat{\mathrm{ES}}_{\mathrm{stack}}(e)$. More formally, by (ref), the within-stack saturated coefficient $\widehat\tau^{\mathrm{sat}}_{g,g+e}$ serves as the plug-in $\widehat\mathrm{CATT}(g,e)$ in (ref), and substituting into the IW-DDD definition gives

equation[equation omitted — 230 chars of source]

The within-stack triple difference as a DDD estimator for $\mathrm{CATT}$.

Following the DiD estimator for $\mathrm{CATT}_{e,\ell}$ in sun_estimating_2021, I define the analogous DDD estimator for $\mathrm{CATT}(g,\ell)$.

defn[DDD estimator for $\mathrm{CATT}$] For cohort $g \in \mathcal{G}_{\mathrm{trg}}$, event-time $\ell$, comparison group $g_c$ with $g_c > g + K$, and baseline period $s = g - 1$, the DDD estimator for $\mathrm{CATT}(g,\ell)$ is \begin{equation} \widehat{\delta}_{g,\ell}^{\mathrm{DDD}} = \left(\overline{\Delta Y}_{g,1,g+\ell} - \overline{\Delta Y}_{g,0,g+\ell}\right) - \left(\overline{\Delta Y}_{g_c,1,g+\ell} - \overline{\Delta Y}_{g_c,0,g+\ell}\right) , \end{equation} where $\overline{\Delta Y}_{s,q,t} = n_{s,q}^{-1}\sum_{i\colon S_i = s, Q_i = q}(Y_{i,t} - Y_{i,g-1})$ is the cell-specific sample mean of long differences.

This definition makes explicit that the within-stack triple difference is a cohort-specific estimator, targeting the treatment effect for a single cohort $g$ at a single event-time $\ell$, using a single comparison group $g_c$. The IW aggregation across cohorts is a separate, post-estimation step.

prop[DDD estimator consistency for $\mathrm{CATT}$] Under {(ref)}--{(ref)}, for any valid clean comparison group $g_c$, the DDD estimator $\widehat{\delta}_{g,\ell}^{\mathrm{DDD}}$ satisfies \begin{align} \widehat{\delta}_{g,\ell}^{\mathrm{DDD}} &\overset{\mathbb P}{\to} \mathrm{CATT}(g,\ell) \quad for all \ell \geq 0 . \end{align}

(ref) is the triple-differences analog of sun_estimating_2021's (sun_estimating_2021) Proposition 5 (DiD estimator for $\mathrm{CATT}_{e,\ell}$). The within-stack saturated regression coefficient $\widehat{\tau}_{g,g+\ell}^{\mathrm{sat}}$ from (ref) is identical to $\widehat{\delta}_{g,\ell}^{\mathrm{DDD}}$ when the comparison group used in the saturated regression is $g_c$ and the baseline period is $g - 1$. The stacked DDD framework thus has a natural connection to the IW estimator.

The average effect of switching

Following de_chaisemartin_difference--differences_2024, define the average effect of switching at event-time $e$ as

equation[equation omitted — 224 chars of source]

the population-share-weighted average of cohort-specific treatment effects at exposure $e$. This coincides with the cohort-size-weighted event-study parameter $\mathrm{ES}^{\mathrm{cohort}}(e)$ from (ref). Under the stacked DDD framework, $\mathrm{AS}(e)$ is estimated by the cohort-size-weighted stacked estimator.

\@startsection{section}{2} \z@{.5\linespacing\@plus.7\linespacing}{-.5em} {\normalfont}{Identification}

This section presents the identification strategy for stacked triple differences. I first establish nonparametric identification of the cohort-time average treatment effects within each stack, then show that the stacked construction mechanically eliminates the forbidden comparisons identified by strezhnev_decomposing_2023, and finally develop testable implications of the identifying assumptions.

Identification within Stacks

I now establish identification of $\mathrm{ATT}(g,t)$ within each stack. The main result shows that the triple difference of cell-mean outcome changes recovers the causal effect.

theorem[Identification in Stacks] Under (ref)--(ref), for each stack $\mathbb{S}_g$ with comparison group $g_c$, and for all post-treatment periods $t \geq g$ , \begin{align} \mathrm{ATT}(g,t) &= \mathbb{E}[\Delta Y_{i,t} \mid S_i = g, Q_i = 1] - \mathbb{E}[\Delta Y_{i,t} \mid S_i = g, Q_i = 0] \notag \\ &\quad - \left(\mathbb{E}[\Delta Y_{i,t} \mid S_i = g_c, Q_i = 1] - \mathbb{E}[\Delta Y_{i,t} \mid S_i = g_c, Q_i = 0]\right) . \end{align}
remark[Pairwise versus global parallel trends] A key advantage of the stacked DDD framework is that (ref) need hold only pairwise between each treatment cohort $g$ and its specific comparison group $g_c$, rather than globally across all cohort pairs simultaneously. This is a weaker requirement than the typical “global” parallel trends assumption. In the stacked DDD framework, we explicitly chooses which comparison group $g_c$ to pair with each treated cohort $g$, and the identifying assumption is tailored to each pair individually. If parallel trends fails for one particular comparison group---say, because a concurrent policy shock affects the eligibility-specific trend in that group---only that stack's estimates are affected; the remaining stacks retain their validity.
remark[Role of each cell in the stack] Each of the four cells in the stack $\mathbb{S}_g$ plays a distinct identifying role. The {treated-eligible} cell $(S_i = g, Q_i = 1)$ provides observed treated outcomes---these units are the target population for $\mathrm{ATT}(g,t)$. The {within-group ineligible} cell $(S_i = g, Q_i = 0)$ measures the group-specific trend for cohort $g$ among ineligible units, enabling the researcher to separate the treatment effect from group-specific confounds. The {comparison-eligible} cell $(S_i = g_c, Q_i = 1)$ measures the eligibility-specific trend absent treatment. Finally, the {comparison-ineligible} cell $(S_i = g_c, Q_i = 0)$ anchors the comparison group's baseline, enabling the third difference that isolates the causal effect from eligibility-specific trends.
remark[Identification without homogeneity] (ref) identifies $\mathrm{ATT}(g,t)$ under arbitrary within-cohort treatment effect heterogeneity. The individual-level effects $Y_{i,t}(g) - Y_{i,t}(\infty)$ may vary freely across units sharing the same $(g, t)$, and the identification formula (ref) recovers the average of these heterogeneous effects over the treated-eligible population without imposing any restriction on their joint distribution. In particular, no constant-effects assumption is needed.

Several features of (ref) merit emphasis. First, identification uses {only} data within the stack $\mathbb{S}_g$---that is, units with $S_i \in \{g, g_c\}$. No information from other treatment cohorts is needed, and no cross-stack restrictions are imposed. Second, the identification formula in (ref) is a {regression adjustment} representation. It recovers the counterfactual by modeling conditional outcome changes in each of the three comparison cells---$(g, 0)$, $(g_c, 1)$, and $(g_c, 0)$---and combining them to form the triple difference. Third, the use of the long difference $\Delta Y_{i,t}$ avoids the compositional issues that arise when the comparison group's treatment status changes over time. Finally, the identification formula (ref) is the unconditional triple difference of cell means. When pre-treatment covariates are available, one can leverage covariate-adjusted identification results developed in Appendix (ref).

\@startsection{section}{2} \z@{.5\linespacing\@plus.7\linespacing}{-.5em} {\normalfont}{Asymptotic Theory: Estimation and Inference}

This section establishes the asymptotic properties of the stacked DDD estimator. The key challenge is that comparison units may be shared across stacks (e.g., never-treated units appear in every stack), inducing cross-stack dependence that must be accounted for in variance estimation and inference.

Estimation

I begin with within-stack estimation of $\mathrm{ATT}(g,t)$, then describe aggregation across stacks to form the event-study parameter $\widehat{\mathrm{ES}}_{\mathrm{stack}}(e)$, discuss the treatment of multiple comparison groups, and outline the practical implementation algorithm.

Throughout, fix a treatment cohort $g \in \mathcal{G}_{\mathrm{trg}}$ with associated comparison group $g_c$ and stack $\mathbb{S}_g$. Let $n_g$ denote the total number of units in the stack. Write $n_{s,q} = \sum_{i \in \mathbb{S}_g} \mathbb{1}\{S_i = s, Q_i = q\}$ for the number of units in cell $(s,q)$, and $n_{g,1}$ for the number of treated-eligible units.

Within-Stack Estimation

The identification formula ((ref)) expresses $\mathrm{ATT}(g,t)$ as a triple difference of cell-mean outcome changes. The natural sample analog is the sample triple difference

equation[equation omitted — 231 chars of source]

where $\overline{\Delta Y}_{s,q,t} = n_{s,q}^{-1}\sum_{i\colon S_i = s, Q_i = q} \Delta Y_{i,t}$ is the cell-specific sample mean of long differences $\Delta Y_{i,t} = Y_{i,t} - Y_{i,g-1}$.

As shown in Section (ref), this estimator is numerically identical to the OLS coefficient on the interaction term in the fully saturated regression (ref) within the stack. It requires no functional form assumptions and no nuisance function estimation---it is a purely design-based estimator that exploits only the four-cell structure within each stack. For the following results, I maintain the following assumption.

assumption[Non-vanishing Stack Shares] As $n \to \infty$, \begin{enumerate} • Cohort shares. For each $g \in \mathcal{G} = \mathcal{G}_{\mathrm{trg}}\cup\{\infty\}$, $n_g/n \overset{\mathbb P}{\to} \lambda_g \equiv \mathbb{P}(G_i = g)$, with $\lambda_g > 0$. • Within-cohort cell shares. For each $g \in \mathcal{G}$ and each cell $(s,q) \in \{0,1\}^2$ with $\mathbb{P}(S_i = s, Q_i = q \mid G_i = g)>0$, the empirical share satisfies $n^{(g)}_{s,q}/n_g \overset{\mathbb P}{\to} \pi^{(g)}_{s,q} \equiv \mathbb{P}(S_i = s, Q_i = q \mid G_i = g) > 0$. • Stack-level shares. For each treated cohort $g \in \mathcal{G}_{\mathrm{trg}}$ with admissible comparison $g_c \in \mathcal{C}(g)$ and each cell $(s,q)$ in stack $g$, the within-stack share $n_{s,q}/n_g \overset{\mathbb P}{\to} \pi_{s,q}^{(g,g_c)} > 0$, where $n_g = \sum_{s,q} n_{s,q}$ aggregates over the four cells of stack $g$. \end{enumerate}

(ref) requires that each cohort, each within-cohort cell, and each stack-level cell retain non-vanishing population shares. The cohort-superscripted notation $\pi^{(g)}_{s,q}$ makes explicit that the cell share is a within-cohort conditional probability, not a marginal—different cohorts may have different $(S,Q)$-cell distributions, and the stack-level share $\pi^{(g,g_c)}_{s,q}$ derived in (iii) inherits this cohort-indexed structure. The primary sampling frame is i.i.d.\ at the unit level ((ref)), hence cell shares are limits of sample proportions. Under the i.i.d. sampling frame, (ref)(i)--(iii) follow from the weak law of large numbers.

prop[Within-Stack Consistency] Under (ref)--(ref) and (ref), the sample triple difference is consistent for the cohort-time average treatment effect on the treated: \begin{equation} \widehat{\mathrm{ATT}}_{\mathbb{S}_g, g_c}(g,t) \overset{\mathbb P}{\to} \mathrm{ATT}(g,t) \quad as n \to \infty . \end{equation}

(ref) extends to estimated weights satisfying $\widehat{\omega}_g(e) \overset{\mathbb P}{\to} \omega_g(e)$, which holds in particular for cohort-size weights $\widehat{\omega}_g(e) = n_{g,1}/\sum_{g'\in\mathcal{G}_{\mathrm{trg}}(e)} n_{g',1}$ under (ref).

Aggregation Across Stacks

The within-stack estimators $\widehat{\mathrm{ATT}}_{\mathbb{S}_g, g_c(g)}(g,t)$ recover cohort-time treatment effects separately for each $g$. To construct event-study parameters, I aggregate across cohorts at each relative event-time $e$. Define the stacked event-study estimator

equation[equation omitted — 253 chars of source]

where $L$ and $K$ are the pre- and post-treatment window lengths and $\widehat{\omega}_g(e) \geq 0$ with $\sum_g \widehat{\omega}_g(e) = 1$ for each $e$. The summation is over all cohorts $g$ for which event-time $e$ falls within the stack window. Following my earlier discussion, I consider two choices of weights.

\noindent1.\ Cohort-size weights. Set

equation[equation omitted — 181 chars of source]

so that each cohort's contribution is proportional to the number of treated-eligible units it contains. This weighting targets the population event-study parameter $\mathrm{ES}(e) = \mathbb{E}[\mathrm{ATT}(G_i, G_i + e) \mid G_i + e \in \{G_i - L, \ldots, G_i + K\}]$, which weights cohorts by their population shares. It is the natural choice when the goal is to estimate an average effect across the treated population.

\noindent2.\ Equal weights. Set

equation[equation omitted — 157 chars of source]

which assigns each cohort equal influence regardless of its sample size or estimation precision. Equal weighting is the most transparent aggregation scheme and places equal value on detecting treatment effects in small and large cohorts alike. However, it has no clear population interpretation and may be inefficient when cohort sizes vary substantially.

remarkThe choice among these two weights involves a familiar trade-off. Cohort-size weights have the most transparent causal interpretation, targeting the “per unit” (of treatment) ATT. Equal weights are the most transparent and are robust to concerns about one or two large cohorts dominating the aggregation, but they may sacrifice both efficiency and interpretability when cohort sizes are heterogeneous. In practice, I recommend reporting results under cohort-size weights as the primary specification, with equal weights as a robustness check. I note that precision weights are statistically efficient but may assign opaque weights that change across event-times, making causal interpretation difficult.

Inference

The stacked event-study regression (ref) delivers point estimates $\widehat{\tau}_e$ that are convex combinations of the within-stack triple differences ((ref)). However, standard OLS standard errors from this regression---the heteroskedasticity-robust standard errors---do not correctly account for the cross-stack dependence induced by shared comparison units. Luckily, the cluster-robust variance estimators (CRVE) clustered at the original unit level {do} correctly recover the asymptotic variance. In the results that follow, all limits are taken as $n \to \infty$ with the number of periods $T$ and the number of cohorts $|\mathcal{G}_{\mathrm{trg}}|$ held fixed.

assumption[Finite Second Moments] For each stack $\mathbb{S}_g$ with comparison group $g_c$, each cell $(s,q)$, and each $t$ in the stack window, $\mathbb{E}[Y_{i,t}^2 \mid S_i = s, Q_i = q] < \infty$.

(ref) requires that outcomes have finite variance in each cell.

theorem[Within-Stack Asymptotic Normality] Under (ref)--(ref), (ref), and (ref), for each stack $\mathbb{S}_g$ with comparison group $g_c$ and for $t \geq g$, {writing $\Delta Y_{i, t} \equiv Y_{i, t} - Y_{i, g - 1}$ for the long difference from the stack's baseline period $g - 1$,} \begin{equation} \sqrt{n_g}\Big(\widehat{\mathrm{ATT}}_{\mathbb{S}_g, g_c}(g,t) - \mathrm{ATT}(g,t)\Big) \overset{d}{\longrightarrow} \mathcal{N}\big(0, \Sigma_{g,t,g_c}\big) , \end{equation} where the asymptotic variance is \begin{equation} \Sigma_{g,t,g_c} = \sum_{(s,q)} \frac{1}{\pi_{s,q}} \mathrm{Var}\big(\Delta Y_{i,t} \mid S_i = s, Q_i = q\big) , \end{equation} {and $\pi_{s, q} \equiv \lim_{n \to \infty} n_{s, q} / n_g$ is the within-stack conditional share of cell $(s, q)$.} The influence function for unit $i \in \mathbb{S}_g$ is \begin{equation} \psi_{\mathbb{S}_g}(W_i; g, t, g_c) = \sum_{(s,q)} \frac{\mathbb{1}\{S_i = s, Q_i = q\}}{\pi_{s,q}} c_{s,q} \big(\Delta Y_{i,t} - \mathbb{E}[\Delta Y_{i,t} \mid S_i = s, Q_i = q]\big) , \end{equation} where $c_{g,1} = +1$, $c_{g,0} = -1$, $c_{g_c,1} = -1$, $c_{g_c,0} = +1$ are the signs of the triple difference.

The influence function (ref) has a transparent structure. Each unit contributes through its cell-specific deviation from the cell mean, weighted by the inverse of its cell's population share and signed according to the triple-difference pattern. The variance (ref) is the sum of cell-specific variance contributions, each scaled by the inverse cell share---larger cells contribute less variance per observation, as expected.

I now turn to the asymptotic distribution of the aggregated event-study estimator $\widehat{\mathrm{ES}}_{\mathrm{stack}}(e)$, which combines within-stack treatment effect estimates across cohorts using weights $\omega_g(e)$.

theorem[Stacked Event-Study CLT] Under the conditions of (ref), for deterministic weights $\omega_g(e) \geq 0$ satisfying $\sum_{g \in \mathcal{G}_{\mathrm{trg}}} \omega_g(e) = 1$ and for event-time $e \in \{0,\ldots,K\}$: \begin{equation} \sqrt{n}\Big(\widehat{\mathrm{ES}}_{\mathrm{stack}}(e) - \mathrm{ES}(e)\Big) \overset{d}{\longrightarrow} \mathcal{N}\big(0, V_{\mathrm{stack}}(e)\big), \end{equation} where $n = \sum_{g \in \mathcal{G}_{\mathrm{trg}}} n_g$ and the form of $V_{\mathrm{stack}}(e)$ depends on whether stacks share control units.

The key subtlety in the aggregated result is that when stacks share comparison units---the most common case in practice, where every stack uses the never-treated group $S_i = \infty$ as its comparison---the within-stack estimators are not independent. The shared comparison units induce cross-stack correlation that must be accounted for in the variance formula. To see why, consider a unit $i$ with $S_i = \infty$ that belongs to stacks $\mathbb{S}_g$ and $\mathbb{S}_{g'}$ for two distinct cohorts $g \neq g'$. The influence function contribution of this unit to the two within-stack estimators is generally non-zero for both, creating dependence between $\widehat{\mathrm{ATT}}_{\mathbb{S}_g, g_c}(g, g+e)$ and $\widehat{\mathrm{ATT}}_{\mathbb{S}_{g'}, g_c}(g', g'+e)$. The following result characterizes the resulting variance structure.

prop[Variance with Shared Controls] When stacks share comparison units (e.g., $g_c = \infty$ for all stacks), the asymptotic variance of the aggregated estimator is: \begin{equation} V_{\mathrm{stack}}(e) = \sum_{g \in \mathcal{G}_{\mathrm{trg}}} \sum_{g' \in \mathcal{G}_{\mathrm{trg}}} \omega_g(e) \omega_{g'}(e) \mathrm{Cov}\big(\psi_{\mathbb{S}_g}(W_i; g, g+e, g_c), \psi_{\mathbb{S}_{g'}}(W_i; g', g'+e, g_c)\big), \end{equation} where the covariance is non-zero when stacks $\mathbb{S}_g$ and $\mathbb{S}_{g'}$ share comparison units. When stacks use distinct comparison groups ($g_c(g) \neq g_c(g')$ with no shared units), the cross-terms vanish and: \begin{equation} V_{\mathrm{stack}}(e) = \sum_{g \in \mathcal{G}_{\mathrm{trg}}} \omega_g(e)^2 \frac{\Sigma_{g, g+e, g_c(g)}}{n_g} . \end{equation}
remark[Magnitude of cross-stack correlation] The cross-stack covariance in (ref) arises exclusively from shared comparison units. In applications where the never-treated pool is large relative to the treated cohorts, these units receive small inverse-probability weights, and the cross-stack correlation is modest. Conversely, when the never-treated pool is small and each comparison unit receives substantial weight, the cross-stack dependence is non-negligible and ignoring it leads to confidence intervals that are too narrow.

A natural estimator of $V_{\mathrm{stack}}(e)$ in (ref) is the sample analogue based on estimated influence functions. For each unit $i$ and each stack $\mathbb{S}_g$, define $\widehat{\psi}_{\mathbb{S}_g}(W_i; g, g+e, g_c)$ as the influence function (ref) evaluated at sample cell proportions and sample cell means, with $\widehat{\psi}_{\mathbb{S}_g}(W_i) = 0$ for units $i \notin \mathbb{S}_g$.\footnote{To verify that this is the correct influence function, note that $\widehat{\tau}_{g,g+e}^{\mathrm{sat}} = n_g^{-1}\sum_{i \in \mathbb{S}_g} \widehat{\psi}_{\mathbb{S}_g}(W_i) + \widehat{\tau}_{g,g+e}^{\mathrm{sat}}$, since $\sum_{i \in \mathbb{S}_g} \widehat{\psi}_{\mathbb{S}_g}(W_i) = 0$ by the orthogonality of cell-mean residuals. The influence function captures the first-order effect of perturbing unit $i$'s outcome on the triple difference---adding a unit to cell $(g,1)$ with outcome $\Delta Y_{i,g+e}$ shifts $\overline{\Delta Y}_{g,1,g+e}$ by approximately $(\Delta Y_{i,g+e} - \overline{\Delta Y}_{g,1,g+e})/n_{g,1}$, which changes $\widehat{\tau}_{g,g+e}^{\mathrm{sat}}$ by approximately $(\Delta Y_{i,g+e} - \overline{\Delta Y}_{g,1,g+e})/n_{g,1} = \widehat{\psi}_{\mathbb{S}_g}(W_i)/n_g$.} The plug-in variance estimator is

equation[equation omitted — 226 chars of source]

This estimator has two important features. First, by summing the weighted influence function contributions at the unit level before squaring, it automatically captures the cross-stack covariance for shared control units. When unit $i$ belongs to multiple stacks, its contributions from different stacks are summed inside the square, producing the correct cross-terms. Second, to establish consistency of this variance estimator, I require a strengthening of the moment condition in (ref).

assumption[Finite Fourth Moments] For each stack $\mathbb{S}_g$ with comparison group $g_c$, each cell $(s,q)$, and each $t$ in the stack window, $\mathbb{E}[Y_{i,t}^4 \mid S_i = s, Q_i = q] < \infty$.

(ref) strengthens the second-moment condition in (ref) to fourth moments. This is needed because the variance estimator (ref) is a sample average of squared influence functions: the law of large numbers requires that these squared terms have finite variance, which imply finite fourth moments of the outcome differences by the $c_r$ inequality.

prop[Variance Estimator Consistency] Under (ref)--(ref), (ref), and (ref), for deterministic weights $\omega_g(e)$ satisfying $\sum_g \omega_g(e) = 1$, and for each event-time $e \in \{0,\ldots,K\}$ \begin{equation} \widehat{V}_{\mathrm{stack}}(e) \overset{\mathbb P}{\to} V_{\mathrm{stack}}(e) . \end{equation}

(ref) extends to estimated weights satisfying $\widehat{\omega}_g(e) \overset{\mathbb P}{\to} \omega_g(e)$, which holds in particular for cohort-size weights $\widehat{\omega}_g(e) = n_{g,1}/\sum_{g'\in\mathcal{G}_{\mathrm{trg}}(e)} n_{g',1}$ under (ref). At a high-level, the proof involves showing that the plug-in variance estimator is a sample average of squared estimated influence functions that converges by the law of large numbers, with the estimation error in the influence functions controlled by all the regularity conditions above. As an immediate consequence, pointwise confidence intervals of the form

equation[equation omitted — 138 chars of source]

have asymptotically correct coverage. More precisely, (ref) combined with (ref) and Slutsky's theorem gives

equation[equation omitted — 201 chars of source]

so the interval (ref) has asymptotic coverage $1-\alpha$ for each $e$.

In practice, valid inference under the fully saturated OLS regression specification applied to stacked dataset can be done via the cluster-robust standard error, where the clustering is done at the unit level. This validity is ensured under the same assumptions maintained in (ref); I prove this in Appendix (ref). Hence, in the case of estimating fully saturated OLS regressions on the stacked dataset, the usual cluster-robust standard errors---clustered at the level of treatment---can be used without implementing the estimator proposed in (ref).

remark[Multiplier bootstrap] For uniform confidence bands across all event-times, we can implement the multiplier bootstrap proceeds as follows. For each bootstrap replication $b = 1, \ldots, B$: \begin{enumerate} • Draw i.i.d.\ multiplier weights $\{\xi_i^{(b)}\}_{i=1}^{n}$ from a distribution satisfying $\mathbb{E}[\xi_i] = 0$, $\mathbb{E}[\xi_i^2] = 1$, and $\mathbb{E}[\xi_i^4] < \infty$. Standard choices are $\xi_i \sim N(0,1)$ or $\xi_i \in \{-1, +1\}$ with equal probability. Importantly, the same weight $\xi_i^{(b)}$ is used for unit $i$ across all stacks containing that unit. • Define \begin{equation} \widehat\phi_i(e) \equiv \sum_{g \in \mathcal{G}_{\mathrm{trg}}(e)} \widehat\omega_g(e) \widehat\psi_{\mathbb{S}_g}\big(W_i; g, g + e, g_c(g)\big), \end{equation} the aggregated influence function of unit $i$ at event-time $e$, summed across all stacks the unit belongs to. Then, calculate \begin{equation} \widehat{\tau}_e^{*,(b)} = \frac{1}{n}\sum_{i=1}^{n} \xi_i^{(b)} \widehat{\phi}_i(e) . \end{equation} The bootstrap variance estimator is $\widehat{V}_{\mathrm{boot}}(e) = B^{-1}\sum_{b=1}^{B} [\widehat{\tau}_e^{*,(b)}]^2$, which is asymptotically equivalent to $\widehat{V}(e)/n$. \end{enumerate} To construct confidence bands, let $\mathcal{E} = \{-L, \ldots, K\} \setminus \{-1\}$ be the set of event-times. Compute the bootstrap critical value $c_\alpha$ as the $(1-\alpha)$-quantile of $\max_{e \in \mathcal{E}} |\widehat{\tau}_e^{*,(b)}| / \sqrt{\widehat{V}_{\mathrm{boot}}(e)}$ across $b = 1, \ldots, B$. The simultaneous $(1-\alpha)$th confidence band is $\{\widehat{\tau}_e \pm c_\alpha \sqrt{\widehat{V}_{\mathrm{boot}}(e)}\}_{e \in \mathcal{E}}$. The essential design feature of this bootstrap is the use of a {common} multiplier weight $\xi_i$ for each unit $i$ across all stacks. When unit $i$ belongs to stacks $\mathbb{S}_g$ and $\mathbb{S}_{g'}$, the bootstrap covariance between the two within-stack contributions is $\mathbb{E}[\xi_i^2] \widehat{\psi}_{\mathbb{S}_g}(W_i) \widehat{\psi}_{\mathbb{S}_{g'}}(W_i) = \widehat{\psi}_{\mathbb{S}_g}(W_i) \widehat{\psi}_{\mathbb{S}_{g'}}(W_i)$, which correctly mimics the population covariance arising from unit $i$'s shared membership. Drawing independent multipliers for each stack--unit pair would break this dependence and yield invalid inference.

\@startsection{section}{2} \z@{.5\linespacing\@plus.7\linespacing}{-.5em} {\normalfont}{Empirical Illustrations}

I illustrate the stacked DDD framework with two applications drawn from recent empirical work that exploits staggered rollouts over a country $\times$ eligibility-group panel: hansen_national_2023's study of the national impacts of genetically modified crops, and shastry_vaccine_2025's evaluation of Gavi's vaccine program. For both illustrations, I show the cohort-size weights in my stack DDD specification and use the never-treated groups as my clean comparison group in each stack.

National and Global Impacts of Genetically Modified Crops

hansen_national_2023 study the national and global impacts of genetically modified (GM) crops. In particular, they estimate the impact of genetically modified (GM) crops on countrywide yields, harvested area, and trade “using a triple-differences rollout design that exploits variation in the availability of GM seeds across crops, countries, and time”. In other words, there are staggered national approvals for GM cultivation across countries, a clear eligibility dimension defined at the crop level, and a multi-country panel dataset encompassing both treated and never-treated crops and countries. They find “positive impacts on yields, especially in poor countries,” and emphasize that “without GM crops, the world would have needed 3.4 percent additional cropland to keep global agricultural output at its 2019 level.”

Mapping their setting to my notations, the treatment-enabling group $S_i$ records the year a country $i$ first approved GM cultivation, with never-adopting countries assigned $S_i = \infty$. Eligibility $Q_c$ indicates whether crop $c$ has a commercially viable GM variety available globally (specifically: cotton, maize, rapeseed, and soybean, so $Q_c = 1$), whereas other field crops (e.g., rice, wheat) serve as ineligible within-country controls ($Q_c = 0$). The treatment indicator $D_{ict} = \mathbb{1}\{t \geq S_i\} Q_c$ takes the value one if country $i$ has approved GM cultivation by year $t$ and crop $c$ is a GM-eligible crop. Finally, I examine four primary outcomes at the country-crop-year level: the GM adoption rate, log yield, log harvested area, and net-export share.

I construct a stack for each GM approval cohort $g \in \mathcal{G}_{\mathrm{trg}}$. Within each stack, the never-treated countries ($S_i = \infty$) serve as the clean comparison group ($g_c = \infty$). For a given cohort $g$, the stack contains four cells: treated-eligible (GM crops in cohort $g$ countries), within-group ineligible (non-GM crops in cohort $g$ countries), across-group eligible (GM crops in never-adopting countries), and across-group ineligible (non-GM crops in never-adopting countries).

I now compare the results of the conventional 3WFE estimator (“H&W (pooled)”, (ref)) against a stacked OLS estimator (“Sat. OLS (stacked)”) and the triple-differences estimator of ortiz-santanna_triple_2025 (“triplediff (DR)”). For aggregating cohort-level ATTs into event-study, for my proposed method, I used cohort-size as weights, so the event-study estimates should be interpreted as “per unit” ATT. The event-study plots for all four outcomes are presented in Figure (ref) and Figure (ref).

Event-Study Estimates

Across all outcomes of interest, the stacked DDD estimator display more modest, but still statistically significant, treatment effect estimates. In particular, my results confirm most of the findings in hansen_national_2023. Unsurprisingly, my estimates and those of “triplediff (DR)” are qualitatively similar. The pooled specification of HW likely inflates the treatment effect estimates by using early-adopting countries as controls for later-adopting countries in the data.

Aggregated ATT

All estimated effects under the stacked design are uniformly less in magnitude compared to the HW results. In the case of log harvested area, the stacked OLS estimator reveals a statistically insignificant effect of GM adoption on percentage of harvest. In other cases, all the estimated aggregated ATTs are half of the estimated effects of HW. To illustrate treatment effect heterogeneity, I show cohort-level ATTs in Figure (ref). The vertical dotted lines inside the panels plot the cohort-size weighted aggregated ATTs across these stacks, which map exactly to the corresponding pooled point estimates generated by the “Sat. OLS (stacked)” estimator presented in Figure (ref).

Overall, the alternative estimation strategy I employ here mostly confirms the positive impact of the technology on outcomes of interest, with the exception of harvested area. The main advantage of the stacked design is it deals with identification transparently, allowing practitioners to design credible sub-experiments for each cohort and better argue that identification assumption holds.

Effective Health Aid through Gavi's Vaccine Program

I next revisit shastry_vaccine_2025's evaluation of Gavi, the Vaccine Alliance, which has distributed over US\$16 billion in vaccine aid to low- and middle-income countries since its founding in 1999. shastry_vaccine_2025 identify Gavi's causal effect from the differential timing of Gavi funding across countries, vaccines, and birth cohorts, comparing Gavi-funded vaccines to non-funded vaccines within the same country and birth cohort, and comparing the same vaccines across Gavi-adopting and non-adopting countries. They estimate a triple-differences specification

equation[equation omitted — 137 chars of source]

with country $\times$ vaccine, country $\times$ cohort, and vaccine $\times$ cohort fixed effects, clustering standard errors by country. They report that “Gavi's support for a vaccine increased coverage rates by $2-5$ percentage points across all vaccines, on average, and by $10-20$ percentage points for newer vaccines,” and that “Gavi's support for a vaccine reduced child mortality from related causes by 1 child per 1,000 live births,” implying roughly $1.5$ million lives saved at approximately US\$9,000 per life saved. shastry_vaccine_2025 acknowledge the sensitivity of staggered 3WFE estimators to treatment-effect heterogeneity and verify robustness using borusyak_revisiting_2023's imputation estimator.

Here is the mapping from the context of shastry_vaccine_2025 to my notations. The treatment-enabling group $S_i$ is the year country $i$ first received Gavi funding for the vaccine targeting cause $c$, with never-funded countries assigned $S_i = \infty$. Eligibility $Q_c$ is an indicator for the cause of death being the primary target of a Gavi-funded vaccine, so respiratory, diarrheal, measles, and meningitis/encephalitis causes have $Q_c = 1$ and other under-five causes have $Q_c = 0$. The treatment indicator is $D_{ict} = \mathbb{1}\{t \geq S_i\} Q_c$. The cross-sectional unit is the country-cause pair, and the outcome $Y_{ict}$ is the post-neonatal mortality rate (deaths per 1,000 live births) from cause $c$ in country $i$ in year $t$.

For each Gavi-funding cohort $g \in \mathcal{G}_{\mathrm{trg}}$ and each Gavi-linked cause $c$, I construct a stack with never-funded countries ($S_i = \infty$) as the clean comparison group. I set $L = K = 5$ to match the 11-year window used in shastry_vaccine_2025's mortality event study. I analyze two samples that parallel shastry_vaccine_2025's Figure 5: the full GHO sample (“All Countries”) and the vital-registry subset restricted to years from 2005 onwards (“Post-2005 Vital Registry”), which avoids imputed mortality data at the cost of a substantially smaller set of countries.

Event-study estimates

Figures (ref) and (ref) report event-study coefficients from three estimators applied to the mortality panel: the pooled OLS estimator in (ref) (“OLS (pooled)”), the saturated OLS stacked estimator (“Sat.\ OLS (stacked)”), and the triple-differences estimator of ortiz-santanna_triple_2025 (“triplediff (Reg)”).

In both samples, all three estimators display essentially flat pre-trends between $e = -5$ and $e = -2$, consistent with shastry_vaccine_2025's identifying assumption that coverage of Gavi-funded vaccines would have trended similarly to coverage of non-funded vaccines absent Gavi. Post-introduction, all three estimators show a downward trend in mortality rate.

However, the three estimators disagree in magnitude. In the full sample (Figure (ref)), the pooled OLS post-treatment ATT is $-0.945$ deaths per 1,000 live births, the saturated stacked OLS estimator gives $-0.633$, and triplediff produces $-0.248$. All aggregated ATTs estimated are statistically significant. In the cleaner vital-registry subset (Figure (ref)), the three estimates are to $-0.362$, $-0.248$, and $-0.257$ per 1,000 live births for the original specification, saturated stacked OLS estimator, and triplediff, respectively. The large gap between pooled OLS and the two stacked estimators in the full sample liekly reflects the forbidden-comparison problem in (ref): later-funded country-cause pairs are implicitly compared to already-funded ones whose mortality rates continue to decline in the post-treatment window, and the resulting negative weights inflate the pooled OLS point estimate. In the smaller vital-registry subset, where the treatment timing is more concentrated and the effective control pool is narrower, the three estimators roughly agree at approximately a quarter of a death per 1,000 live births. Figures (ref) and (ref) decompose the saturated stacked OLS estimate into its underlying cause-by-cohort stacks.

Applying the stacked DDD methodology to Gavi's vaccine rollout reinforces shastry_vaccine_2025's central conclusion that Gavi-funded vaccines meaningfully reduced child mortality from related causes. In the cleaner vital-registry sample---where shastry_vaccine_2025 themselves prefer the results---the stacked estimators agree with the pooled design on a reduction of roughly a quarter of a death per 1,000 live births, so the qualitative conclusion is robust. The stacked decomposition additionally shows transparently which country-vaccine-cohort stack drive the aggregate treatment effect estimate.

landscape\begin{figure}[ht] \caption{Event-Study Estimates by Outcome in hansen_national_2023} \end{figure} \begin{figure}[ht] \caption{Cohort-Level ATT by Outcome in hansen_national_2023} \end{figure} \begin{figure}[htbp] \caption{Event-Study Estimates: Impact of Gavi on Child Mortality, All Countries shastry_vaccine_2025} \end{figure} \begin{figure}[htbp] \caption{Event-Study Estimates: Impact of Gavi on Child Mortality, Post-2005 Vital Registry Subset shastry_vaccine_2025} \end{figure} \begin{figure}[htbp] \caption{Stack-Level ATT Decomposition: All Countries shastry_vaccine_2025} \end{figure} \begin{figure}[htbp] \caption{Stack-Level ATT Decomposition: Post-2005 Vital Registry Subset shastry_vaccine_2025} \end{figure}

\thispagestyle{empty}

center[center omitted — 133 chars of source]
spacing{1.5} \vskip1cm \startcontents[sections] \printcontents[sections]{l}{1}{Table of Contents\vskip3pt\hrule\vskip5pt\setcounter{tocdepth}{2}} \vskip5pt\hrule\vskip5pt

\thispagestyle{empty} \setcounter{page}{0} \setcounter{figure}{0}