The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
115,339 characters
Identification and Estimation of Staggered Difference-in-Differences with Network Spillovers
\maketitle
\begin{abstract}
This paper develops a difference-in-differences framework for staggered policy adoption when units can be affected by other units' adoption.
For each treated cohort and event time, the framework separates the effect of own adoption, the spillover effect generated by other adopters, and the total effect under the realized rollout.
Identification uses a prespecified summary of spillover exposure and parallel trends comparisons among units with the same exposure at the baseline and target dates.
Spillover effects are learned from never-treated units and evaluated for treated cohorts under the exposure distribution they face.
We construct estimators for these effects and an inference procedure that allows for spatial dependence.
Monte Carlo simulations illustrate that standard DID estimators that ignore spillovers can miss the total effect, whereas the proposed estimators have small bias for these effects and the associated confidence intervals have coverage close to the nominal level.
In an empirical study of the Community Health Centers rollout, estimated spillovers account for a substantial share of the effect on older-adult mortality.
\end{abstract}
Keywords: Difference-in-differences; staggered adoption; interference; spillovers; exposure mapping; spatial dependence.
\newpage
\section{Introduction}
Difference-in-differences designs use untreated units to approximate the counterfactual trend that treated units would have followed without adoption.
In the standard design, the stable unit treatment value assumption rules out spillovers: a unit's potential outcome depends on its own treatment path, not on other units' adoption.
Recent work on staggered adoption has made the treated-versus-control comparison explicit.
\citet{callaway2021difference} define group-time treatment effects and aggregation schemes.
\citet{goodman2021difference} shows how TWFE coefficients combine two-group comparisons, and \citet{sun2021estimating} show how event-study coefficients can mix effects across relative periods.
\citet{borusyak2024revisiting} and \citet{gardner2025two} estimate untreated potential outcomes before constructing dynamic effects.
\citet{ROTH20232218} provide a review of this recent DID literature.
Those results address the comparison and weighting problems created by staggered timing under this no-spillover assumption.
This paper studies staggered DID when adoption occurs on a network.
A unit can remain untreated by its own policy status and still be exposed to other adopters.
Treated units may also receive spillovers after their own adoption.
For later-treated cohorts, the baseline period may already include exposure generated by earlier adopters.
The treated-versus-control distinction therefore no longer coincides with the exposed-versus-unexposed distinction.
Once these states differ, a conventional DID contrast that compares treated cohorts with never-treated units combines several terms.
It subtracts a comparison-group trend that may already contain untreated-state spillovers.
It may also combine the effect of own adoption with spillovers received by treated units and with baseline exposure generated before the cohort's adoption date.
Under the stable unit treatment value assumption (SUTVA), these exposure terms are absent and the relevant causal objects reduce to the usual group-time treatment effect.
When that assumption fails because of spillovers, the effect of changing own adoption, the spillover effect that would arise without own adoption, and the total effect of the realized rollout are separate objects.
This separation follows the distinction between switching, spillover, and total effects in \citet{savje2021average} and \citet{butts2024difference}.
Following \citet{manski2013identification} and \citet{10.1214/16-AOAS1005}, we summarize spillover exposure with a prespecified exposure mapping.
Under that mapping, we define effects for each adoption cohort and event time under the exposure distribution generated by the realized staggered rollout.
The dynamic switching effect (DSE) changes own adoption for a treated cohort at a given event time, averages over that cohort, and holds exposure fixed at its realized state.
For a treated cohort, the control-state spillover effect (CSE) is the spillover contrast when own adoption is set to the untreated state, averaged over the exposure distribution faced by that cohort.
The dynamic total effect (DTE) compares a state with no adoption to the realized adoption path and equals the sum of DSE and CSE on the same admissible support.
This total effect is tied to the rollout-induced exposure distribution; it is not the effect of an arbitrary counterfactual rollout.
It is also distinct from the effect of own adoption when every treated unit has zero exposure, which requires potential outcomes for treated units that adopt but are not exposed to other adopters.
Identification follows the missing potential-outcome state required by each effect.
For the DSE, identification requires the trend that the treated cohort would have followed without own adoption while keeping its baseline and target exposure states fixed.
We recover this trend from never-treated units with the same covariates and the same exposure states at the baseline and target dates, under no anticipation, parallel trends after conditioning on those exposure states, and overlap.
For the CSE, the missing object is the untreated-state spillover response evaluated for the treated cohort.
Comparisons among never-treated units with different exposure states identify this response under a parallel trends assumption within never-treated units and exposure states.
A separate cross-cohort transportability assumption evaluates the response over the joint distribution of covariates and exposure in the treated cohort.
The DTE is obtained by adding the identified DSE and CSE on the same admissible support.
The same never-treated source contrast also identifies spillovers on never-treated units themselves.
With additional restrictions linking treated potential outcomes across exposure states, the framework can also identify the effect of own adoption when exposure is set to zero.
The estimators implement the identified effects.
The DSE estimator is a saturated long-difference comparison within retained cells defined by covariates and the two exposure states.
The CSE estimator fits the untreated-state spillover response in the never-treated source sample and evaluates the fitted contrast over the distribution of covariates and exposure in the treated cohort.
The DTE estimator is the post-estimation sum on the same support.
For inference, the estimating equations are stacked to obtain first-order approximations for DSE, CSE, and DTE.
Spatial heteroskedasticity- and autocorrelation-consistent (spatial HAC) covariance estimates are then applied to the resulting event-time rows.
This covariance step follows recent work on inference under spatial or network dependence, including \citet{leung2022causal} and \citet{xu2025difference}.
The paper contributes to methods for difference-in-differences with interference in three ways.
First, it defines policy effects for each cohort and event time that separate the effect of changing own adoption, spillovers when own adoption is held at the untreated state, and the total effect of the realized rollout.
Second, it gives identification results for the main effects.
Third, it provides estimators for DSE, CSE, and DTE and an inference method.
The procedure conditions on the realized design, uses first-order approximations to the DSE and CSE estimators, and estimates covariance with spatial HAC methods.
The simulations illustrate the finite-sample behavior of the proposed DSE, CSE, and DTE estimators.
They also show how standard DID benchmarks that assume no spillovers can miss the total effect of the realized rollout.
The empirical study revisits the rollout of Community Health Centers studied by \citet{bailey2015war}, using county-year mortality data to illustrate that the estimated total effect can differ materially once spillovers among counties without their own Community Health Center are estimated rather than absorbed into the comparison-group trend.
\subsection*{Related literature}\label{sec:related-literature}
The no-interference staggered-DID literature supplies the benchmark case in which exposure states do not enter potential outcomes.
\citet{callaway2021difference} identify group-time average treatment effects and aggregate them according to the empirical question of interest.
\citet{goodman2021difference} shows that the conventional TWFE DID estimator with variation in treatment timing is a weighted average of two-group, two-period comparisons, which complicates interpretation under heterogeneous and dynamic treatment effects.
\citet{sun2021estimating} show that standard event-study coefficients in staggered designs can be contaminated by treatment effects from other relative periods and propose interaction-weighted estimators.
\citet{borusyak2024revisiting} and \citet{gardner2025two} estimate untreated potential outcomes using untreated observations and then construct treatment effects from residualized treated outcomes.
Under the stable unit treatment value assumption, the dynamic switching effect and dynamic total effect reduce to a group-time treatment effect.
With interference, own-adoption switching effects, exposure effects in the untreated state, and total effects of the realized rollout are distinct objects.
A separate literature studies DID when treatment can affect nearby or connected units.
\citet{delgado2015difference} develop DID techniques for spatial data with local autocorrelation and spatial interaction.
\citet{clarke2017estimating} studies DID estimation when spillovers propagate from treated to nearby comparison areas.
\citet{berg2021spillover} discuss spillover-induced biases in empirical corporate finance designs.
These papers motivate treating exposure as part of the design rather than as an omitted nuisance.
Here, the exposure mapping defines cohort-event-time DSE, CSE, and DTE objects, and spillovers in the comparison group are separated from own-adoption switching.
\citet{10.1214/16-AOAS1005} formalize exposure mappings as a way to reduce the full assignment vector to interpretable exposure states and derive design-based estimators under randomized assignments.
\citet{savje2024causal} emphasizes that exposure mappings can serve two different roles: defining interpretable exposure-based effects and imposing assumptions about the true interference structure.
The exposure mapping in this paper plays both roles, which is why sensitivity to distance cutoffs, network weights, and exposure thresholds is part of the empirical interpretation.
Relatedly, \citet{leung2022causal} studies approximate neighborhood interference and network HAC inference, where the effect of distant assignments can decay rather than vanish exactly.
\citet{huber2021framework} distinguish individual-level treatment effects from spillover, interaction, and general-equilibrium effects.
Their setting allows spillover effects within aggregate units and permits individual treatment effects to interact with aggregate treatment intensity.
In the staggered network design studied here, timing varies by cohort and event time, and exposure enters through a unit-level network state.
The staggered-DID design here replaces randomized assignment with parallel trends assumptions, support conditions in the never-treated source sample, and transportability across cohorts.
\citet{fiorini2025simple} consider staggered DID with spillovers and target the ATT without interference.
Their identification strategy uses restrictions that remove spillovers from treated units after adoption and require a set of spillover-free controls.
The approach here uses the exposure mapping differently.
Rather than using it to recover the no-interference ATT, the mapping defines exposure states for each cohort and event time and the corresponding switching, spillover, and total effects.
Treated units may remain exposed after adoption, and the total effect of the realized rollout is decomposed into DSE and CSE on a common admissible support.
\citet{butts2024difference} develops DID methods for spatial spillovers in a two-period design and distinguishes switching effects from total effects, with an event-study extension to staggered adoption.
Relative to Butts, the object here is a decomposition of the total effect of the realized rollout for each cohort and event time.
The spillover effect is learned from never-treated units and evaluated over the exposure distributions faced by treated cohorts.
\citet{xu2025difference} analyzes a two-period DID design with interference under exposure mappings and develops doubly robust direct and spillover estimators, finite-population GMM, and spatial HAC inference.
This paper uses related finite-population and stacked-moment arguments.
The target differs because the stacked system is used for a staggered-adoption decomposition in which the total effect is built from switching and spillover effects on the same support.
\citet{sun2025differenceindifferencesnetworkinterference} focus on DID under network interference and direct and spillover ATT-type estimands while adjusting for network confounders.
The comparison is again mainly about the target effect.
Here, the exposure mapping defines switching, spillover, and total effects for the realized rollout, and the CSE uses a never-treated source response transported to treated cohorts.
The remainder of the paper is organized as follows.
Section~\ref{sec:setup} introduces the staggered-adoption environment, the exposure mapping, and the DSE, CSE, and DTE estimands.
Section~\ref{sec:identification} gives the identification results, including the decomposition of conventional DID contrasts and the non-identification result for the pure direct effect at zero exposure.
Sections~\ref{sec:estimation} and \ref{sec:inference} present the component estimators and the joint spatial HAC inference procedure.
Section~\ref{sec:monte-carlo} reports the Monte Carlo evidence, and Section~\ref{sec:application} applies the method to data on the rollout of Community Health Centers.
Section~\ref{sec:conclusion} concludes.
\section{Setup}\label{sec:setup}
\subsection{Framework}\label{sec:framework}
Consider a panel of $N$ units indexed by $i=1,\dots,N$ over periods $t=1,\dots,T$.
Let $D_{it}\in\{0,1\}$ denote own treatment status, and impose absorbing treatment, $D_{it}\le D_{i,t+1}$ for all $i$ and $t<T$.
Define the adoption time $G_i \equiv \min\{t:D_{it}=1\}\in\{2,\dots,T\}\cup\{\infty\}$, with $G_i=\infty$ for never-treated units.
Throughout, the never-treated state $G_i=\infty$ is the reference adoption status.
Comparisons involving untreated potential outcomes are therefore always taken with respect to $G_i=\infty$, rather than units that are untreated at time $t$ but adopt later.
Let $\mathbf G=(G_1,\dots,G_N)$ denote the adoption-time vector, and let $Y_{it}(\mathbf G)$ be the potential outcome for unit $i$ at time $t$ under adoption profile $\mathbf G$.
This notation permits unrestricted interference: unit $i$'s potential outcome may depend on the entire adoption-time vector.
The primitive objects in the identification analysis are $G_i$, baseline covariates $X_i$, the exposure path $\{H_{it}(\mathbf G_{-i})\}_{t=1}^T$, and the reduced-form potential-outcome collection $\{Y_{it}(g,h): g\in\mathcal G,\ h\in\mathcal H\}_{t=1}^T$, where $\mathcal G\equiv\{2,\dots,T\}\cup\{\infty\}$ and $\mathcal H$ is the support of the reduced-form exposure state.
Throughout the theoretical analysis, $\mathcal H$ is taken to be a finite set containing $0$.
This discretization ensures that the support conditions in the identification results are well defined.
The dependence of $Y_{it}(\mathbf G)$ on the full adoption-time vector rules out nonparametric identification without further structure.
Following \citet{10.1214/16-AOAS1005}, we use an exposure mapping that summarizes the relevant features of neighbors' adoption behavior into a scalar exposure state.
This is a maintained structural restriction: any channel of interference not encoded in the weights and temporal kernel is excluded by assumption, and the adequacy of the mapping must be defended on substantive grounds.
Let $\widetilde H_{it}(\mathbf G_{-i})$ denote the raw exposure index.
Let $w_{ii}=0$ for all $i$, and define $\widetilde H_{it}(\mathbf G_{-i})\equiv\sum_{j\ne i} w_{ij}\phi(t,G_j)$, where $\phi(t,G_j)=\psi(t-G_j)\mathbbm{1}\{t\ge G_j\}$, so that spillovers depend on event time since neighbors' adoption.
The theoretical exposure state is a coarsened version of this raw index.
Let $b:\mathbb R_+\to\mathcal H$ be a known coarsening map, where $\mathcal H$ is a finite set containing $0$, and define $H_{it}(\mathbf G_{-i})\equiv b(\widetilde H_{it}(\mathbf G_{-i}))$.
Throughout the theoretical analysis, $H_{it}$ refers to this finite-valued coarsened exposure state.
The raw index $\widetilde H_{it}$ is used only to construct $H_{it}$.
\begin{assumption}\label{assumption:exposure_mapping}
For each $(i,t)$, there exist reduced-form potential outcomes $\{Y_{it}(g,h):g\in\mathcal G,\ h\in\mathcal H\}$ such that $Y_{it}(\mathbf G)=Y_{it}\bigl(G_i,H_{it}(\mathbf G_{-i})\bigr)$ for every adoption profile $\mathbf G$.
\end{assumption}
\begin{assumption}\label{assumption:consistency}
For the realized adoption profile $\mathbf G$, the observed outcome satisfies $ Y_{it}=Y_{it}(\mathbf G)$.
\end{assumption}
Under Assumption~\ref{assumption:exposure_mapping}, this is equivalent to
$Y_{it}=Y_{it}(G_i,H_{it})$, $H_{it}=H_{it}(\mathbf G_{-i})$.
The state $(\infty,0)$ is the pure control state: the unit is untreated and unexposed.
The state $(\infty,h)$ with $h>0$ is the untreated-but-exposed state.
The state $(g,h)$ with $g<\infty$ and $h>0$ is the treated-and-exposed state.
In the treated-and-exposed state, observed outcomes combine own-treatment and spillover components, so separating direct and spillover effects requires restrictions beyond the exposure mapping itself.
Fix an anticipation window $\delta\ge 0$.
For cohort $g$, define the baseline period $t_0(g)\equiv g-\delta-1$.
For a given cohort-event-time pair $(g,l)$, write $t=g+l$.
When the baseline period is clear from context, write $t_0$ in place of $t_0(g)$.
\subsection{Policy effects}\label{sec:policy-effects}
For any calendar-time-indexed random variable $Z_{it}$, define $\mathbb E_g[Z_{i,g+l}]\equiv\mathbb E[Z_{i,g+l}\mid G_i=g]$ and $\mathbb E_{\infty}[Z_{it}]\equiv\mathbb E[Z_{it}\mid G_i=\infty]$.
The primary objects average over the realized distribution of exposure $H_{i,g+l}$ within cohort $g$.
For post-adoption event times $l\ge0$, define
\begin{align*}
\tau^{DSE}(g,l)
& =
\mathbb E_g\!\left[
Y_{i,g+l}(g,H_{i,g+l})-Y_{i,g+l}(\infty,H_{i,g+l})
\right], \\
\tau^{CSE}(g,l)
& =
\mathbb E_g\!\left[
Y_{i,g+l}(\infty,H_{i,g+l})-Y_{i,g+l}(\infty,0)
\right], \\
\tau^{DTE}(g,l)
& =
\mathbb E_g\!\left[
Y_{i,g+l}(g,H_{i,g+l})-Y_{i,g+l}(\infty,0)
\right].
\end{align*}
The dynamic switching effect (DSE) switches own adoption while holding realized exposure fixed.
The control-state spillover effect for cohort $g$ (CSE) is the untreated-state spillover contrast evaluated over the exposure distribution faced by that cohort.
The dynamic total effect (DTE) compares the pure-control regime to the realized adoption regime.
We allow $\tau^{CSE}(g,l)$ to be defined for $l\ge -1$ so that untreated-state spillovers can also be studied in pre-adoption periods; when $\delta=0$, the cell $l=-1$ coincides with the baseline period.
These definitions imply the maintained decomposition
\begin{align*}
\tau^{DTE}(g,l)
=
\tau^{DSE}(g,l)+\tau^{CSE}(g,l).
\end{align*}
The secondary objects state what is not identified under the maintained assumptions.
The pure direct effect is $\tau^{PDE}(g,l)=\mathbb E_g[Y_{i,g+l}(g,0)-Y_{i,g+l}(\infty,0)]$.
The treated-side spillover component is $\tau^{AST}(g,l)=\mathbb E_g[Y_{i,g+l}(g,H_{i,g+l})-Y_{i,g+l}(g,0)]$.
These objects satisfy $\tau^{DTE}(g,l)=\tau^{PDE}(g,l)+\tau^{AST}(g,l)$.
The control-group spillover effect among never-treated units is $\tau_{\infty}^{CSE}(t)\equiv\mathbb E_{\infty}[Y_{it}(\infty,H_{it})-Y_{it}(\infty,0)]$.
The full unit-level taxonomy generating these aggregate decompositions is given in Appendix~\ref{app:identification-proofs-subsection}.
The cohort qualifier in the CSE name refers to the target population over which the effect is averaged.
The phrase ``control-state'' refers to the no-own-treatment potential-outcome contrast $Y_{it}(\infty,H_{it})-Y_{it}(\infty,0)$.
The never-treated group is used as a source population for learning this response; it is not the target population of $\tau^{CSE}(g,l)$.
At $l=0$, these estimands coincide with cohort averages of the corresponding unit-level effects and summarize the immediate consequences of adoption for cohort $g$.
The primary identified targets in what follows are $\tau^{DSE}(g,l)$, $\tau^{CSE}(g,l)$, and therefore $\tau^{DTE}(g,l)$.
In empirical applications, fixed analysis weights can be used to define the target population over which these same potential-outcome contrasts are averaged.
The weighted empirical targets replace cohort expectations by fixed $\varpi_i$-weighted cohort expectations and reduce to these definitions when $\varpi_i=1$.
By contrast, $\tau^{AST}(g,l)$ and the global $\tau^{PDE}(g,l)$ are well-defined population objects but are not point-identified under the maintained assumptions without additional structure.
The local pure direct effect is identified only on isolated support.
The control-group spillover effect among never-treated units, $\tau_{\infty}^{CSE}(t)$, is a secondary estimand; its identification is discussed in Section~\ref{sec:cse}.
\begin{remark}\label{remark:no_spillover}\label{lem:no_spill}
If exposure does not affect potential outcomes, $Y_{it}(g,h)=Y_{it}(g,0)$ for all $(g,h)$.
Then, for every cohort $g$ and every post-treatment event time $l$,
\begin{align*}
\tau^{DTE}(g,l)=\tau^{DSE}(g,l)=\tau^{PDE}(g,l)
=
\mathbb E\!\left[
Y_{i,g+l}(g,0)-Y_{i,g+l}(\infty,0)
\mid
G_i=g
\right].
\end{align*}
This is an equality of estimands, not a DID identification result.
\end{remark}
Under the no-spillover specialization in Remark~\ref{remark:no_spillover}, the primary estimands collapse to the standard group-time treatment effect for cohort $g$ at calendar time $t=g+l$, up to the baseline and anticipation convention.
Conversely, any discrepancy between the spillover-robust objects here and a cohort-specific group-vs-never-treated DID contrast in the presence of interference reflects some combination of post-treatment spillovers and baseline contamination from earlier adopters.
\section{Identification}\label{sec:identification}
Table~\ref{tab:obs_missing_states} records which potential-outcome states are observed under the realized rollout and which states are counterfactual in the observed data.
For the DSE, treated cohorts reveal treated potential outcomes at realized exposure, while the comparison requires the untreated counterfactual trend for units with the same baseline and target exposure states.
For the CSE, exposure contrasts in the never-treated source sample recover the untreated spillover response, and transportability evaluates that response over the exposure distribution faced by the treated cohort.
Once DSE and CSE are available on the same admissible support, $\tau^{DTE}(g,l)=\tau^{DSE}(g,l)+\tau^{CSE}(g,l)$ follows.
The global PDE requires treated potential outcomes at zero exposure and is not point-identified for treated units observed at positive exposure without additional restrictions.
\begin{table}[!htbp]
\centering
\small
\caption{Observed and missing states in the realized rollout}
\label{tab:obs_missing_states}
\begingroup
\setlength{\tabcolsep}{6pt}
\begin{tabularx}{0.98\textwidth}{
@{}
>{\raggedright\arraybackslash}p{0.34\textwidth}
>{\raggedright\arraybackslash}X
@{}
}
\toprule
Object & Status in the observed data \\
\midrule
\(Y_{it}(G_i,H_{it})\)
& Observed for every unit-period by consistency. \\
\(Y_{it}(g,H_{it})\) for \(G_i=g\)
& Observed for treated cohort \(g\) at realized exposure. \\
\(Y_{it}(\infty,H_{it})\)
& Observed for never-treated units; counterfactual for treated cohorts after adoption. \\
\(Y_{it}(\infty,0)\)
& Observed only in never-treated zero-exposure cells; otherwise counterfactual. \\
\(Y_{it}(G_i,0)\)
& Observed only when realized exposure is zero; missing for treated-and-exposed observations. \\
\(Y_{it}(\infty,h)-Y_{it}(\infty,0)\)
& Not observed as a unit-level contrast; source-cell means are visible among never-treated units when supported. \\
\(Y_{it}(g,H_{it})-Y_{it}(\infty,H_{it})\)
& Not directly observed because the untreated state is missing for treated cohorts. \\
\(Y_{it}(g,H_{it})-Y_{it}(\infty,0)\)
& Not directly observed because the pure-control state is missing for treated cohorts. \\
\(Y_{it}(g,0)-Y_{it}(\infty,0)\)
& Component states are present only on isolated zero-exposure support; missing globally for treated-and-exposed units. \\
\bottomrule
\end{tabularx}
\endgroup
\end{table}
The assumptions below use its observed and missing states selectively.
\subsection{Identification assumptions}
\begin{assumption}\label{assumption:no_anticipation}\label{ass:no_anticipation}
For every cohort $g$, every $t\le t_0(g)$, and every $h\in\mathcal H$,
\begin{align*}
Y_{it}(g,h)=Y_{it}(\infty,h)
\quad\text{a.s.}
\end{align*}
conditional on $G_i=g$.
\end{assumption}
This no-anticipation condition makes the cohort-$g$ baseline outcome an untreated potential outcome at the realized exposure state.
For fixed $(g,l)$, let $t=g+l$ and $t_0=t_0(g)$.
Define the two-date exposure state
\begin{align*}
S_i^{g,l}\equiv S_{it}^{g,l}\equiv(H_{it},H_{i,t_0}).
\end{align*}
\begin{assumption}\label{assumption:dse_parallel_trends}\label{ass:pt_dse}
For each retained $(g,l)$, with $t=g+l$, $t_0=t_0(g)$, and $S_i^{g,l}=(H_{it},H_{i,t_0})$, the following equality holds:
\begin{align*}
& \mathbb E\!\left[
Y_{it}(\infty,H_{it})-Y_{i,t_0}(\infty,H_{i,t_0})
\mid G_i=g,\ X_i^d,\ S_i^{g,l}
\right] \\
& =
\mathbb E\!\left[
Y_{it}(\infty,H_{it})-Y_{i,t_0}(\infty,H_{i,t_0})
\mid G_i=\infty,\ X_i^d,\ S_i^{g,l}
\right].
\end{align*}
\end{assumption}
This is the DSE parallel trends assumption for units with the same baseline and target exposure states.
It compares treated and never-treated units after conditioning on $X_i^d$ and the baseline and target exposure states.
\begin{assumption}\label{assumption:cse_source_trends}\label{ass:pt_exposure}
Suppose no unit is treated at $t=1$, so that $H_{i1}=0$ for all $i$.
Then, for every $t\in\{2,\dots,T\}$, every $h\in\mathcal H$, and every finite-stratum value $x$,
\begin{align*}
& \mathbb E\!\left[
Y_{it}(\infty,0)-Y_{i1}(\infty,0)
\mid
G_i=\infty,\ X_i^d=x,\ H_{it}=h
\right] \\
& =
\mathbb E\!\left[
Y_{it}(\infty,0)-Y_{i1}(\infty,0)
\mid
G_i=\infty,\ X_i^d=x,\ H_{it}=0
\right].
\end{align*}
\end{assumption}
This is a parallel trends assumption within never-treated units and exposure states.
It lets zero-exposure never-treated cells remove pure-control trend differences when learning the control-state spillover response.
\begin{assumption}\label{assumption:cse_transport}\label{ass:transport}
For each retained $(g,l)$, with $t=g+l$,
\begin{align*}
r_{g,t}(X_i^d,H_{it})
=
r_{\infty,t}(X_i^d,H_{it})
\quad\text{a.s.}
\end{align*}
conditional on $G_i=g$, where
\begin{align*}
r_{a,t}(x,h)
\equiv
\mathbb E\!\left[
Y_{it}(\infty,h)-Y_{it}(\infty,0)
\mid
G_i=a,\ X_i^d=x,\ H_{it}=h
\right].
\end{align*}
\end{assumption}
This is a transportability condition, not a parallel trends assumption.
It carries the never-treated source response $r_{\infty,t}(x,h)$ to the cohort-$g$ target distribution.
The assumptions in this section bridge the missing states in Table~\ref{tab:obs_missing_states} in a selective way.
Assumptions~\ref{assumption:no_anticipation} and \ref{assumption:dse_parallel_trends} identify the DSE by supplying the untreated long difference for cohort $g$ among units with the same baseline and target exposure states.
Assumptions~\ref{assumption:cse_source_trends} and \ref{assumption:cse_transport} identify the control-state spillover response in the never-treated source sample and carry it to the cohort-$g$ target distribution.
What remains outside the maintained identification region is the global PDE, $\tau^{PDE}(g,l)$, and, more generally, the unrestricted treated-side zero-exposure state $Y_{it}(G_i,0)$ away from isolated support.
A local PDE can be recovered on isolated zero-exposure support under an additional zero-exposure parallel trends assumption and a positive-mass condition; Appendix~\ref{app:identification-assumptions} states the additional assumptions and Appendix~\ref{app:identification-proofs-subsection} gives the formal result.
Those isolated-support conditions are not used to identify the main DSE, CSE, or DTE objects.
\subsection{Conventional DID under spillovers}
To accommodate the general anticipation-window setup, consider the cohort-specific group-vs-never-treated DID comparison between the adoption period $t=g$ and the no-anticipation baseline $t_0=t_0(g)=g-\delta-1$.
Define
\begin{align*}
\Delta Y_{ig}\equiv Y_{ig}-Y_{i,t_0},
\\
\tau_g^{DID}\equiv \mathbb E_g[\Delta Y_{ig}]-\mathbb E_{\infty}[\Delta Y_{ig}].
\end{align*}
When $\delta=0$, this reduces to the adjacent-period DID comparison with $t_0(g)=g-1$.
Let
\begin{align*}
B_g^{base}
& \equiv
\mathbb E\!\left[
Y_{i,t_0}(\infty,H_{i,t_0})-Y_{i,t_0}(\infty,0)
\mid
G_i=g
\right], \\
B_g^{PT}
& \equiv
\mathbb E\!\left[
Y_{ig}(\infty,0)-Y_{i,t_0}(\infty,0)
\mid
G_i=g
\right]
-
\mathbb E\!\left[
Y_{ig}(\infty,0)-Y_{i,t_0}(\infty,0)
\mid
G_i=\infty
\right].
\end{align*}
Here $B_g^{base}$ is baseline exposure contamination for cohort $g$, and $B_g^{PT}$ is the pure-control trend difference between cohort $g$ and never-treated units.
\begin{prop}\label{prop:did_decomposition}\label{prop:DID_bias}
Suppose Assumptions~\ref{assumption:exposure_mapping}, \ref{assumption:consistency}, and \ref{assumption:no_anticipation} hold.
For $t=g$ and $t_0=t_0(g)$,
\begin{align*}
\tau_g^{DID}=\tau^{PDE}(g,0)+\tau^{AST}(g,0)-B_g^{base}-\tau_{\infty}^{CSE}(g)+\tau_{\infty}^{CSE}(t_0)+B_g^{PT}.
\end{align*}
\end{prop}
\begin{proof}
See Appendix~\ref{app:identification-proofs-subsection} for the separate treated and never-treated long-difference expansions.
\end{proof}
Thus, a DID comparison between treated cohorts and never-treated units combines own adoption, treated-side spillovers, comparison-group spillovers, baseline exposure contamination, and pure-control trend bias.
\begin{remark}\label{remark:no_spill_did}\label{prop:no_spill_did}
Suppose Assumptions~\ref{assumption:exposure_mapping}, \ref{assumption:consistency}, and \ref{assumption:no_anticipation} hold, and suppose the no-spillover specialization in Remark~\ref{remark:no_spillover} holds.
For cohort $g$, let $t_0=t_0(g)$.
Then
\begin{align*}
\tau_g^{DID}
=
\tau^{PDE}(g,0)+B_g^{PT},
\end{align*}
where $B_g^{PT}$ is the pure-control trend difference defined above.
Consequently, if $B_g^{PT}=0$, then
\begin{align*}
\tau_g^{DID}
=
\tau^{PDE}(g,0)
=
\tau^{DSE}(g,0)
=
\tau^{DTE}(g,0).
\end{align*}
The remaining requirement for a DID interpretation is the pure-control trend restriction $B_g^{PT}=0$.
\end{remark}
The decomposition above reveals two distinct distortions under interference.
In the post-treatment period, the terms $\tau^{AST}(g,0)$ and $\tau_{\infty}^{CSE}(g)$ reflect contamination from spillover exposure after adoption.
At the anticipation-safe baseline $t_0(g)$, the term $-B_g^{base}$ captures baseline contamination: by the time cohort $g$ reaches its reference period, earlier cohorts may already have shifted the baseline outcome away from the pure control state.
When $\delta=0$, $t_0(g)=g-1$ and $B_g^{base}$ coincides with $\tau^{CSE}(g,-1)$.
This distinction is specific to staggered adoption.
For the first treated cohort, or in a simultaneous-adoption design, the baseline period precedes any treatment in the population and the pre-treatment spillover terms vanish.
Later cohorts face a different environment because their baseline may already be contaminated by earlier adopters.
This cohort-specific DID contrast therefore mixes direct effects, post-treatment spillovers, and baseline contamination whenever interference is present.
\subsection{Identification results}\label{sec:cse}
For fixed $(g,l)$, let $t=g+l$, $t_0=t_0(g)$, and
\begin{align*}
\Delta Y_i(t,t_0)
& \equiv
Y_{it}-Y_{i,t_0}, \\
S_i^{g,l}
& \equiv
(H_{it},H_{i,t_0}).
\end{align*}
Define
\begin{align*}
m_{g,l}(x,s)
\equiv
\mathbb E\!\left[
\Delta Y_i(t,t_0)
\mid
G_i=\infty,\ X_i^d=x,\ S_i^{g,l}=s
\right],
\end{align*}
and let $F^D_{g,l}\ll F^D_{\infty,g,l}$ denote DSE overlap for $(X_i^d,S_i^{g,l})$.
For CSE, define
\begin{align*}
\Delta^{(1)}Y_{it}
& \equiv
Y_{it}-Y_{i1}, \\
\mu_t(x,h)
& \equiv
\mathbb E\!\left[
\Delta^{(1)}Y_{it}
\mid
G_i=\infty,\ X_i^d=x,\ H_{it}=h
\right],
\end{align*}
and
\begin{align*}
\mathcal O^C_t
\equiv
\left\{
(x,h):
p_{\infty,t}(h\mid x)>0
\text{ and }
p_{\infty,t}(0\mid x)>0
\right\}.
\end{align*}
Let $F_{g,l}\equiv\mathcal L(X_i^d,H_{it}\mid G_i=g)$.
The following propositions give identification formulas for DSE, CSE, and DTE.
The first proposition identifies the DSE by replacing the missing untreated long difference for cohort $g$ with the same-state long difference among never-treated units.
\begin{prop}\label{prop:dse}
Fix an admissible $(g,l)$ and let $t=g+l$ and $t_0=t_0(g)$.
Suppose Assumptions~\ref{assumption:exposure_mapping}, \ref{assumption:consistency}, \ref{assumption:no_anticipation}, and \ref{assumption:dse_parallel_trends} hold, and suppose $F^D_{g,l}\ll F^D_{\infty,g,l}$.
Then
\begin{align*}
\tau^{DSE}(g,l)
& =
\mathbb E\!\left[
\Delta Y_i(t,t_0)
\mid
G_i=g
\right]
-
\mathbb E\!\left[
m_{g,l}(X_i^d,S_i^{g,l})
\mid
G_i=g
\right].
\end{align*}
\end{prop}
\begin{proof}
See Appendix~\ref{app:identification-proofs-subsection}.
\end{proof}
The next proposition identifies the CSE by recovering the control-state spillover response in the never-treated source sample and integrating it over the cohort-$g$ exposure distribution.
\begin{prop}\label{prop:cse}\label{lem:spill_nt}
Fix an admissible $(g,l)$ and let $t=g+l$.
Suppose Assumptions~\ref{assumption:exposure_mapping}, \ref{assumption:consistency}, \ref{assumption:cse_source_trends}, and \ref{assumption:cse_transport} hold, and suppose $F_{g,l}(\mathcal O^C_t)=1$.
Then
\begin{align*}
\tau^{CSE}(g,l)
& =
\int
\Bigl(
\mu_t(x,h)-\mu_t(x,0)
\Bigr)\,
dF_{g,l}(x,h).
\end{align*}
\end{prop}
\begin{proof}
See Appendix~\ref{app:identification-proofs-subsection}.
\end{proof}
The following proposition identifies the DTE by adding the DSE and CSE formulas on the same admissible support.
\begin{prop}\label{prop:dte}\label{thm:main_identification}
Fix an admissible $(g,l)$ and let $t=g+l$ and $t_0=t_0(g)$.
Suppose Assumptions~\ref{assumption:exposure_mapping}, \ref{assumption:consistency}, \ref{assumption:no_anticipation}, \ref{assumption:dse_parallel_trends}, \ref{assumption:cse_source_trends}, and \ref{assumption:cse_transport} hold, with $F^D_{g,l}\ll F^D_{\infty,g,l}$ and $F_{g,l}(\mathcal O^C_t)=1$ on the same admissible support.
Then
\begin{align*}
\tau^{DTE}(g,l)
& =
\tau^{DSE}(g,l)+\tau^{CSE}(g,l) \\
& =
\mathbb E\!\left[
\Delta Y_i(t,t_0)
\mid
G_i=g
\right]
-
\mathbb E\!\left[
m_{g,l}(X_i^d,S_i^{g,l})
\mid
G_i=g
\right] \\
& +
\int
\Bigl(
\mu_t(x,h)-\mu_t(x,0)
\Bigr)\,
dF_{g,l}(x,h).
\end{align*}
\end{prop}
\begin{proof}
See Appendix~\ref{app:identification-proofs-subsection}.
\end{proof}
\begin{remark}\label{rem:asnc}
The same within-source comparison identifies the never-treated control-group spillover diagnostic when $F_{\infty,t}(\mathcal O^C_t)=1$:
\begin{align*}
\tau_{\infty}^{CSE}(t)
=
\int
\{\mu_t(x,h)-\mu_t(x,0)\}
\,dF_{\infty,t}(x,h),
\end{align*}
where $F_{\infty,t}\equiv\mathcal L(X_i^d,H_{it}\mid G_i=\infty)$.
This object is identified within the source population and is not part of the cohort-$g$ DTE decomposition.
\end{remark}
\begin{remark}\label{rem:no_transport}
Without transportability, DSE remains identified by the same-state comparison, but CSE and DTE are not point-identified for treated cohorts.
The never-treated-calibrated surrogate
\begin{align*}
\widetilde\tau^{CSE}(g,l)
\equiv
\int
\Bigl(
\mu_t(x,h)-\mu_t(x,0)
\Bigr)\,dF_{g,l}(x,h)
\end{align*}
is identified and equals the true $\tau^{CSE}(g,l)$ only under transportability.
\end{remark}
The identified dynamic total effect is indexed by the realized exposure distribution generated by the observed staggered-adoption process.
Thus,
\begin{align*}
\tau^{DTE}(g,l)
=
\mathbb E\!\left[
Y_{i,g+l}(g,H_{i,g+l})-Y_{i,g+l}(\infty,0)
\mid
G_i=g
\right]
\end{align*}
should be interpreted as the effect for cohort $g$ at event time $l$ of moving from the pure-control regime to the observed adoption regime, holding fixed the exposure environment generated by that regime.
It is not the effect of an arbitrary counterfactual rollout rule.
If all calendar-valid cells are support-valid, an event-time total effect can be written as
\begin{align*}
\tau^{DTE}(l)
=
\sum_{g:g+l\le T}
\omega_{g,l}\tau^{DTE}(g,l),
\\
\omega_{g,l}
=
\Pr(G_i=g\mid G_i<\infty,\ g+l\le T).
\end{align*}
With fixed analysis weights, the event-time aggregate replaces $\omega_{g,l}$ by the corresponding weighted cohort mass share.
If support failure removes some cohort-event-time cells, the aggregate is over the admissible cells and should be interpreted as an admissible-cell aggregate, denoted $\tau^{DTE}_{\mathcal A}(l)$ when that restricted population must be explicit.
\subsection{Pure direct effects}
The maintained identification results above do not, by themselves, separate the pure direct effect from the spillover effect on treated units.
The decomposition
\begin{align*}
\tau^{PDE}(g,l)=\tau^{DTE}(g,l)-\tau^{AST}(g,l),
\end{align*}
and the preceding propositions do not identify $\tau^{AST}(g,l)$.
We therefore treat $\tau^{PDE}(g,l)$ as a secondary object that requires additional restrictions.
\begin{prop}\label{prop:pde_nonidentification}\label{prop:global_pde_nonid}
Under the maintained assumptions, $\tau^{PDE}(g,l)$ is not point-identified for cohorts with positive realized exposure unless additional restrictions link $Y_{it}(g,H_{it})$ and $Y_{it}(g,0)$.
\end{prop}
\begin{proof}
See Appendix~\ref{app:identification-proofs-subsection}.
\end{proof}
One possible additional structure is
\begin{align*}
Y_{it}(a,h)
& =
Y_{it}(a,0)+s_t(h),
\\
a
& \in
\mathcal G.
\end{align*}
This restriction is a substantive no-interaction restriction between own treatment status and spillover exposure.
It is not implied by the exposure mapping or by the DID restrictions above.
Under this no-interaction restriction,
\begin{align*}
\tau^{AST}(g,l)
& =
\mathbb E_g\!\left[
Y_{i,g+l}(g,H_{i,g+l})-Y_{i,g+l}(g,0)
\right] \\
& =
\mathbb E_g\!\left[
s_{g+l}(H_{i,g+l})
\right],
\end{align*}
and
\begin{align*}
\tau^{CSE}(g,l)
& =
\mathbb E_g\!\left[
Y_{i,g+l}(\infty,H_{i,g+l})-Y_{i,g+l}(\infty,0)
\right] \\
& =
\mathbb E_g\!\left[
s_{g+l}(H_{i,g+l})
\right].
\end{align*}
Therefore $\tau^{AST}(g,l)=\tau^{CSE}(g,l)$, and
\begin{align*}
\tau^{PDE}(g,l)
=
\tau^{DTE}(g,l)-\tau^{CSE}(g,l)
=
\tau^{DSE}(g,l).
\end{align*}
Without that restriction, only a local PDE on isolated zero-exposure support is available under additional conditions.
Appendix~\ref{app:identification-assumptions} states the additional assumptions, and Appendix~\ref{app:identification-proofs-subsection} gives the formal result.
\section{Estimation}\label{sec:estimation}
DSE is estimated by saturated long differences within retained two-date exposure-state cells.
CSE is estimated by fitting a control-state spillover response in the never-treated source sample and evaluating the fitted contrast over the treated cohort target distribution.
DTE is formed as a post-estimation map using the same admissible cells and weights as the two primitive estimators.
The estimators in this section are implemented directly by the closed-form formulas below.
Section~\ref{sec:inference} stacks the corresponding estimating equations to derive joint influence rows for DSE, CSE, and DTE.
The stacked estimating-equation representation is used for joint linearization and covariance estimation; it does not change the closed-form point estimators.
For each cohort-event-time pair $(g,l)$, let $t=g+l$ and $t_0=t_0(g)$.
Define $\Delta_i^{g,l}\equiv Y_{it}-Y_{i,t_0}$, $D_i^g\equiv\mathbbm{1}\{G_i=g\}$, and $C_i\equiv\mathbbm{1}\{G_i=\infty\}$.
Let $\varpi_i>0$ denote a fixed analysis weight, such as a baseline older-adult population measure.
The unweighted case is obtained by setting $\varpi_i=1$.
For any random variable $A_i$, define the cohort-$g$ weighted average operator
\begin{align*}
\mathbb E_g^\varpi[A_i]
\equiv
\frac{
\mathbb E[\varpi_i A_i\mid G_i=g]
}{
\mathbb E[\varpi_i\mid G_i=g]
}.
\end{align*}
The weighted primitive causal targets are
\begin{align*}
\tau_{\varpi}^{DSE}(g,l)
& =
\mathbb E_g^\varpi[
Y_{i,g+l}(g,H_{i,g+l})-Y_{i,g+l}(\infty,H_{i,g+l})
], \\
\tau_{\varpi}^{CSE}(g,l)
& =
\mathbb E_g^\varpi[
Y_{i,g+l}(\infty,H_{i,g+l})-Y_{i,g+l}(\infty,0)
].
\end{align*}
For the parametric CSE estimator, let $V_i$ denote the first-stage covariate basis and let $c_t(v,h;\eta)$ denote a control-state spillover response contrast, normalized by $c_t(v,0;\eta)=0$.
The parametric transported CSE target is
\begin{align*}
\tau_{\varpi}^{CSE,param}(g,l)
=
\mathbb E_g^\varpi[
c_t(V_i,H_{it};\eta^0)
],
\end{align*}
for $t=g+l$ and the corresponding total target is
\begin{align*}
\tau_{\varpi}^{DTE,param}(g,l)
=
\tau_{\varpi}^{DSE}(g,l)+\tau_{\varpi}^{CSE,param}(g,l).
\end{align*}
When $\varpi_i$ is not constant, the DSE component is interpreted under the corresponding weighted parallel trends assumption within the retained $Z_i^{g,l}$ cells.
All identification results extend to the fixed-weight targets by replacing cohort and source expectations with the corresponding $\varpi_i$-weighted expectations; the application imposes the same identifying restrictions for these weighted target populations.
The unweighted case is the special case $\varpi_i=1$.
Let $S_i^{g,l}\equiv(H_{it},H_{i,t_0})$ denote the two-date exposure state.
If baseline covariates are used in the identifying parallel trends assumption, let $X_i^d$ denote a finite-dimensional discretization or stratification of $X_i$, fixed before estimation, and define
$Z_i^{g,l}\equiv(X_i^d,H_{it},H_{i,t_0})$.
If no covariate stratification is used, $Z_i^{g,l}=S_i^{g,l}$.
All saturated same-state comparisons below are saturated in $Z_i^{g,l}$, not merely in $S_i^{g,l}$, whenever the maintained identification assumption conditions on $X_i^d$.
Let $\mathcal Z_{g,l}\equiv\operatorname{supp}(Z_i^{g,l}\mid G_i=g)$ denote the target support for the DSE comparison.
\subsection{Main estimators}\label{subsec:main-estimators}
The DSE support condition requires every target state to also be observed among never-treated units:
\begin{align*}
\operatorname{supp}(Z_i^{g,l}\mid G_i=g)
\subseteq
\operatorname{supp}(Z_i^{g,l}\mid G_i=\infty).
\end{align*}
For $a\in\{g,\infty\}$ and $z\in\mathcal Z_{g,l}$, define
\begin{align*}
n_{a,z}^{g,l}
& \equiv
\sum_{i=1}^N
\mathbbm{1}\{G_i=a,Z_i^{g,l}=z\}, \\
n_g
& \equiv
\sum_{i=1}^N
\mathbbm{1}\{G_i=g\},
\end{align*}
and the corresponding weighted masses
\begin{align*}
W_{a,z}^{g,l}
& \equiv
\sum_{i=1}^N
\varpi_i\mathbbm{1}\{G_i=a,Z_i^{g,l}=z\}, \\
W_g
& \equiv
\sum_{i=1}^N
\varpi_i\mathbbm{1}\{G_i=g\}.
\end{align*}
The count variables determine empirical support and minimum-cell admissibility.
The point estimates use the analysis weights:
\begin{align*}
\bar\Delta_{a,z,\varpi}^{g,l}
\equiv
\frac{
\sum_{i=1}^N
\varpi_i\mathbbm{1}\{G_i=a,Z_i^{g,l}=z\}\Delta_i^{g,l}
}{
W_{a,z}^{g,l}
}.
\end{align*}
When $\varpi_i=1$, this reduces to the unweighted cell mean
\begin{align*}
\bar\Delta_{a,z}^{g,l}
\equiv
\frac{
\sum_{i=1}^N
\mathbbm{1}\{G_i=a,Z_i^{g,l}=z\}\Delta_i^{g,l}
}{
n_{a,z}^{g,l}
}.
\end{align*}
Let
\begin{align*}
\widehat{\mathcal Z}_{g,l}
\equiv
\{z:n_{g,z}^{g,l}>0,\ n_{\infty,z}^{g,l}>0\}
\end{align*}
be the empirical common support.
The saturated DSE estimator is
\begin{align*}
\widehat\tau_{\varpi}^{DSE}(g,l)
=
\sum_{z\in\widehat{\mathcal Z}_{g,l}}
\widehat\pi_{\varpi,g,l}(z)
\left(
\bar\Delta_{g,z,\varpi}^{g,l}
-
\bar\Delta_{\infty,z,\varpi}^{g,l}
\right),
\end{align*}
where
\begin{align*}
\widehat\pi_{\varpi,g,l}(z)
\equiv
\frac{W_{g,z}^{g,l}}{W_g}.
\end{align*}
In the empirical implementation, estimates are reported only when the stricter minimum-count rule in Section~\ref{subsec:estimation-aggregation} retains all target mass.
On reported cells, $\widehat{\mathcal Z}_{g,l}$ can therefore be replaced by the retained support $\widehat{\mathcal Z}_{g,l}^{D}$.
Equivalently, $\widehat\tau_{\varpi}^{DSE}(g,l)$ is obtained by weighted least squares on the retained comparison sample $\{i:G_i\in\{g,\infty\},\ Z_i^{g,l}\in\widehat{\mathcal Z}_{g,l}\}$,
\begin{align*}
\Delta_i^{g,l}
=
\sum_{z\in\widehat{\mathcal Z}_{g,l}}
\alpha_{z,g,l}\mathbbm{1}\{Z_i^{g,l}=z\}
+
\sum_{z\in\widehat{\mathcal Z}_{g,l}}
\beta_{z,g,l}D_i^g\mathbbm{1}\{Z_i^{g,l}=z\}
+
u_i^{g,l},
\end{align*}
and setting
\begin{align*}
\widehat\tau_{\varpi}^{DSE}(g,l)
=
\sum_{z\in\widehat{\mathcal Z}_{g,l}}
\widehat\pi_{\varpi,g,l}(z)\widehat\beta_{z,g,l}.
\end{align*}
Because the regression is saturated in the retained state $Z_i^{g,l}$, this regression estimator is numerically identical to the cell-mean estimator above.
The CSE first stage uses period $1$ as the zero-exposure baseline, because no unit is treated at $t=1$ and $H_{i1}=0$ under the maintained design.
For $t\ge 2$, define the never-treated conditional mean
\begin{align*}
\mu_t(x,h)
\equiv
\mathbb E[Y_{it}-Y_{i1}\mid G_i=\infty,X_i^d=x,H_{it}=h],
\end{align*}
and the conditional untreated-state spillover response
\begin{align*}
r_{\infty,t}(x,h)
\equiv
\mu_t(x,h)-\mu_t(x,0).
\end{align*}
The causal CSE estimand identified in Proposition~\ref{prop:cse} is
\begin{align*}
\tau_{\varpi}^{CSE}(g,l)
=
\frac{
\mathbb E[
\varpi_iD_i^g
r_{\infty,t}(X_i^d,H_{it})
]
}{
\mathbb E[\varpi_iD_i^g]
},
\end{align*}
for $t=g+l$ under the parallel trends assumption within the never-treated source sample and exposure states, together with the cross-cohort transportability restriction.
The identification result is nonparametric in the finite exposure-state strata: on the source support, $r_{\infty,t}(x,h)$ is identified by $\mu_t(x,h)-\mu_t(x,0)$.
The empirical estimator implements this identified contrast through a structured parametric model $c_t(V_i,H_{it};\eta)$.
Hence the plug-in CSE equals the causal CSE when the fitted contrast is correctly specified on the transported support; otherwise it is the corresponding transported parametric target.
The main CSE estimator uses the parametric transported plug-in obtained from the maintained first-stage specification in the never-treated source sample.
For the empirical implementation, the maintained first stage is a structured binary-positive exposure-response model estimated in the never-treated source sample.
The basis $V_i$ is distinct from the finite support strata $X_i^d$ used in saturated support checks.
For $t\ge2$, define
\begin{align*}
R_{it}
& \equiv
Y_{it}-Y_{i1}, \\
P_{it}
& \equiv
\mathbbm{1}\{H_{it}>0\}.
\end{align*}
The baseline period has $R_{i1}=0$ and $P_{i1}=0$; in the application code it anchors the baseline trend fit and contributes no exposure-response contrast.
The first stage writes the never-treated conditional mean as
\begin{align*}
\mathbb E[R_{it}\mid G_i=\infty,V_i,H_{it}]
& =
a_t(V_i;\alpha)+c_t(V_i,H_{it};\beta), \\
c_t(V_i,0;\beta)
& =
0 .
\end{align*}
In the binary-positive specification,
\begin{align*}
c_t(V_i,H_{it};\beta)
=
P_{it} B_t(V_i)^\top\beta .
\end{align*}
The term $a_t(V_i;\alpha)$ is the baseline trend component used for residualization, and $B_t(V_i)$ is the exposure-response basis.
Pooling over $t$ means that the finite-dimensional parameter is estimated using the stacked never-treated source observations; it imposes a time-invariant spillover response only if $B_t(V_i)$ is specified without calendar-time variation.
The maintained application includes period indicators in $B_t(V_i)$, so the binary-positive exposure contrast varies by calendar time after residualization.
The transported plug-in estimator evaluates the fitted contrast $c_t(V_i,H_{it};\widehat\beta)$ over the cohort-$g$ target distribution.
For the main CSE estimator, let $t=g+l$.
The estimator is
\begin{align*}
\widehat\tau_{\varpi}^{CSE}(g,l)
=
\frac{
\sum_{i=1}^N\varpi_iD_i^g
c_t(V_i,H_{it};\widehat\eta)
}{
W_g
}.
\end{align*}
Its model-implied population target is
\begin{align*}
\tau_{\varpi}^{CSE,param}(g,l)
\equiv
\frac{
\mathbb E[
\varpi_iD_i^g
c_t(V_i,H_{it};\eta^0)
]
}{
\mathbb E[\varpi_iD_i^g]
}.
\end{align*}
If the transported contrast is correctly specified on the target support,
\begin{align*}
c_t(v,h;\eta^0)
=
r_{\infty,t}(x,h)
\end{align*}
for all $(x,v,h)$ in $\operatorname{supp}(X_i^d,V_i,H_{it}\mid G_i=g)$,
then
\begin{align*}
\tau_{\varpi}^{CSE,param}(g,l)=\tau_{\varpi}^{CSE}(g,l).
\end{align*}
The same first-stage contrast also defines a diagnostic estimand for the never-treated comparison group.
This object is not a primary policy target in the DTE decomposition.
It measures spillover exposure in the comparison group and appears directly in the decomposition of conventional comparisons between treated cohorts and never-treated units.
For calendar time $t\ge2$, define the weighted never-treated control-group spillover target
\begin{align}
\tau_{\varpi,\infty}^{CSE}(t)
\equiv
\frac{
\mathbb E[
\varpi_i C_i
\{Y_{it}(\infty,H_{it})-Y_{it}(\infty,0)\}
]
}{
\mathbb E[\varpi_i C_i]
}.
\label{eq:cse-infty-target}
\end{align}
The unweighted case is obtained by setting $\varpi_i=1$.
By the source comparison underlying Proposition~\ref{prop:cse}, this target is identified on the never-treated source support as
\begin{align}
\tau_{\varpi,\infty}^{CSE}(t)
=
\frac{
\mathbb E[
\varpi_i C_i r_{\infty,t}(X_i^d,H_{it})
]
}{
\mathbb E[\varpi_i C_i]
},
\label{eq:cse-infty-identified}
\end{align}
provided
\begin{align*}
\frac{
\mathbb E[
\varpi_i C_i
\mathbbm{1}\{(X_i^d,H_{it})\in\mathcal O_t^C\}
]
}{
\mathbb E[\varpi_i C_i]
}
=
1.
\end{align*}
Unlike $\tau_{\varpi}^{CSE}(g,l)$, this object does not require the cross-cohort transportability restriction in Assumption~\ref{ass:transport}.
It is a within-source-population spillover estimand.
Under the parametric transported plug-in specification, define the corresponding model-implied target
\begin{align}
\tau_{\varpi,\infty}^{CSE,param}(t)
\equiv
\frac{
\mathbb E[
\varpi_i C_i c_t(V_i,H_{it};\eta^0)
]
}{
\mathbb E[\varpi_i C_i]
}.
\label{eq:cse-infty-param-target}
\end{align}
The plug-in estimator is
\begin{align}
\widehat\tau_{\varpi,\infty}^{CSE}(t)
=
\frac{
\sum_{i=1}^N
\varpi_i C_i c_t(V_i,H_{it};\widehat\eta)
}{
W_\infty
}.
\label{eq:cse-infty-estimator}
\end{align}
The denominator is
\begin{align*}
W_\infty
\equiv
\sum_{i=1}^N
\varpi_i C_i .
\end{align*}
If $c_t(v,h;\eta^0)=r_{\infty,t}(x,h)$ on the never-treated source support for $(X_i^d,V_i,H_{it})$, then
\begin{align*}
\tau_{\varpi,\infty}^{CSE,param}(t)
=
\tau_{\varpi,\infty}^{CSE}(t).
\end{align*}
For a cohort-specific DID comparison between $t_0(g)$ and $g$, when both calendar-time spillover effects are defined, define the never-treated spillover change
\begin{align}
\Delta_{\infty}^{CSE}(g)
\equiv
\tau_{\varpi,\infty}^{CSE}(g)
-
\tau_{\varpi,\infty}^{CSE}(t_0(g)).
\label{eq:cse-infty-change}
\end{align}
Its estimator is
\begin{align}
\widehat\Delta_{\infty}^{CSE}(g)
=
\widehat\tau_{\varpi,\infty}^{CSE}(g)
-
\widehat\tau_{\varpi,\infty}^{CSE}(t_0(g)).
\label{eq:cse-infty-change-estimator}
\end{align}
This quantity measures the spillover component embedded in the never-treated comparison group's long difference.
In the group-versus-never DID decomposition, it enters with the opposite sign because the never-treated long difference is subtracted.
For each admissible $(g,l)$, define
\begin{align*}
\widehat\tau_{\varpi}^{DTE}(g,l)
\equiv
\widehat\tau_{\varpi}^{DSE}(g,l)
+
\widehat\tau_{\varpi}^{CSE}(g,l).
\end{align*}
\subsection{Aggregation and admissibility}\label{subsec:estimation-aggregation}
For event-time aggregation, let $\widehat{\mathcal G}_l^\star$ be the set of cohorts for which both DSE and CSE are admissible at event time $l$.
For $g\in\widehat{\mathcal G}_l^\star$, define
\begin{align*}
\widehat\omega_{\varpi,g,l}
\equiv
\frac{W_g}
{\sum_{g'\in\widehat{\mathcal G}_l^\star}W_{g'}}.
\end{align*}
For $Q\in\{DSE,CSE,DTE\}$, define
\begin{align*}
\widehat\tau_{\varpi}^Q(l)
\equiv
\sum_{g\in\widehat{\mathcal G}_l^\star}
\widehat\omega_{\varpi,g,l}
\widehat\tau_{\varpi}^Q(g,l).
\end{align*}
Because the same admissible set and the same weights are used for all three components,
\begin{align*}
\widehat\tau_{\varpi}^{DTE}(l)
=
\widehat\tau_{\varpi}^{DSE}(l)
+
\widehat\tau_{\varpi}^{CSE}(l)
\end{align*}
holds exactly.
Let $m_N$ be a pre-specified minimum cell-count threshold.
Either $m_N=m$ for a fixed integer $m$, or $m_N\to\infty$ with $m_N/N\to0$.
For DSE, define
\begin{align*}
\widehat{\mathcal Z}_{g,l}^{D}
\equiv
\{z:n_{g,z}^{g,l}\ge m_N,\ n_{\infty,z}^{g,l}\ge m_N\}.
\end{align*}
Let $\widehat{\Pr}_{\varpi}(\cdot\mid G_i=g)$ denote the empirical distribution weighted by $\varpi_i$ within cohort $g$.
The DSE component for $(g,l)$ is reported only if
\begin{align*}
\sum_{z\in\widehat{\mathcal Z}_{g,l}^{D}}
\widehat{\Pr}_{\varpi}(Z_i^{g,l}=z\mid G_i=g)
=
1.
\end{align*}
For CSE, define
\begin{align*}
n_{\infty,t}(x,h)
\equiv
\sum_{i=1}^N
\mathbbm{1}\{G_i=\infty,X_i^d=x,H_{it}=h\}.
\end{align*}
The CSE component for $(g,l)$, with $t=g+l$, is reported only if, for every $(x,h)$ in the empirical support of $(X_i^d,H_{it})$ among units with $G_i=g$,
\begin{align*}
n_{\infty,t}(x,h)
& \ge
m_N, \\
n_{\infty,t}(x,0)
& \ge
m_N.
\end{align*}
For the never-treated control-group spillover effect, define the retained source support
\begin{align*}
\widehat{\mathcal XH}_{\infty,t}^{C}
\equiv
\left\{
(x,h):
n_{\infty,t}(x,h)\ge m_N
\text{ and }
n_{\infty,t}(x,0)\ge m_N
\right\}.
\end{align*}
The estimate $\widehat\tau_{\varpi,\infty}^{CSE}(t)$ is reported only if
\begin{align*}
\sum_{(x,h)\in\widehat{\mathcal XH}_{\infty,t}^{C}}
\widehat{\Pr}_{\varpi}
(X_i^d=x,H_{it}=h\mid G_i=\infty)
=
1.
\end{align*}
This rule ensures that the reported never-treated spillover target is evaluated on the same source support used to learn the control-state spillover response.
The admissible cohort set is
\begin{align*}
\widehat{\mathcal G}_l^\star
\equiv
\{g:(g,l)\text{ satisfies both the DSE and CSE admissibility rules}\}.
\end{align*}
All event-time aggregates are computed on $\widehat{\mathcal G}_l^\star$.
\begin{remark}
The component estimators also have regression representations.
The DSE estimator is numerically equivalent to a saturated same-state long-difference regression, and the CSE component can be implemented by an exposure-response regression in the never-treated source sample followed by transport to the cohort-$g$ exposure distribution.
These regression forms are implementation devices; the maintained estimators are the cohort-event-time component estimators defined above.
A lower-dimensional augmented TWFE specification can serve as a diagnostic benchmark, but it imposes linearity in exposure, homogeneity within included interactions, limited flexibility in covariate-specific trends, and the usual staggered-design restrictions on implicit weights.
Appendix~\ref{app:estimation-details} records the corresponding regression formulas.
\end{remark}
\section{Inference}\label{sec:inference}
For inference, the component estimators from Section~\ref{sec:estimation} are embedded in a finite-dimensional stacked estimating-equation system.
The just-identified DSE and CSE target blocks reproduce the closed-form component formulas; DTE is their post-estimation sum on retained cells, and event-time estimates aggregate retained cohort-event cells.
The asymptotic sequence is finite-population and conditional on the design variables.
The design sigma-field $\mathcal C_N$ contains fixed attributes, analysis weights, network weights, and locations; when the realized rollout is treated as fixed, it also contains the realized cohort and exposure paths.
Let $\mathcal A=\{(g,l):g\in\mathcal G_l^\star,\ l\in\mathcal L\}$ denote the retained cohort-event set.
The primitive parameter vector collects the retained DSE source-cell means, the CSE first-stage parameter, the DSE and CSE cell-level targets, the never-treated CSE diagnostic when reported, and the weighted cohort masses.
DTE is not a primitive component; it is the smooth map
\begin{align*}
\tau_{\varpi}^{DTE,param}(g,l)
=
\tau_{\varpi}^{DSE}(g,l)+\tau_{\varpi}^{CSE,param}(g,l).
\end{align*}
Define the stacked moment row
\begin{align*}
q_{i,N}(\theta)
=
\left(
q_i^{DSE,src}(\theta)^\top,
q_i^{CSE,1st}(\eta)^\top,
q_i^{DSE,tgt}(\theta)^\top,
q_i^{CSE,tgt}(\theta)^\top,
q_i^{CSE,\infty,tgt}(\theta)^\top,
q_i^{share}(\theta)^\top
\right)^\top,
\end{align*}
where the never-treated CSE diagnostic block is included only when that diagnostic is reported.
Each block collects the corresponding equations over retained cells, event times, and cohorts; Appendix~\ref{app:estimating-equation-array} gives the full block definitions.
Let $\bar q_N(\theta)=N^{-1}\sum_{i=1}^N q_{i,N}(\theta)$ and write $m_N(\theta)=\mathbb E[\bar q_N(\theta)\mid\mathcal C_N]$.
The population parameter $\theta_N^0$ satisfies the stacked moment condition $m_N(\theta_N^0)=0$.
The component estimators are represented by the stacked GMM first-order condition associated with
\begin{align*}
\bar q_N(\theta)^\top
\widehat\Psi_N
\bar q_N(\theta),
\end{align*}
where $\widehat\Psi_N$ is the weighting matrix for the stacked moment criterion.
The stack contains equations with the following roles.
One set estimates the never-treated trends used in the DSE comparison.
A second set estimates the CSE first stage in the never-treated source sample.
Two target equations map these quantities into the DSE and CSE components.
When the never-treated diagnostic is reported, an additional target equation evaluates the same first-stage response over the never-treated distribution.
The final equations estimate the cohort shares used in event-time aggregation.
In the just-identified target and cohort-share blocks, these equations reproduce the closed-form estimators.
If the CSE first stage is overidentified, the associated first-order score enters the stack under the maintained first-stage specification.
This construction parallels the stacked GMM representation in \citet{xu2025difference}.
Xu stacks nuisance and target moments for a two-period DID problem with interference.
Here, the stacked system is used for retained cohort-event cells in a staggered design and includes equations for the switching component, the spillover component, the first-stage response, the diagnostic spillover object, and cohort shares.
Event-time paths are smooth maps of the primitive components.
Because DTE uses DSE and CSE on the same retained cells, its influence row is formed by adding the DSE and CSE influence rows before spatial HAC covariance estimation.
This keeps the covariance between the two component rows in the DTE standard error.
Spatial dependence uses the $\psi$-dependence route of \citet{KOJEVNIKOV2021882}, and the increasing-domain spatial interpretation follows \citet{JENISH200986} and \citet{JENISH2012178}.
Spatial HAC covariance estimation is applied to the event-time influence rows.
\subsection{Assumptions}
Let $\mathcal C_N=\sigma\{\{X_i\}_{i=1}^N,\{\varpi_i\}_{i=1}^N,\{w_{ij}\}_{i,j=1}^N,\{s_i\}_{i=1}^N\}$ denote the sigma-field generated by fixed attributes, analysis weights, network weights, and locations.
All probability limits and expectations below are conditional on $\mathcal C_N$, unless otherwise stated.
Let $\mathcal L$ be the finite set of reported event times and let $\mathcal G_l^\star$ be the population admissible cohort set at event time $l$.
For a retained $(g,l)$, define
\begin{align*}
Z_i^{g,l}
=
(X_i^d,H_{i,g+l},H_{i,t_0(g)}).
\end{align*}
For a finite-valued random variable $A_i$ and event $B_i$, define the weighted support
\begin{align*}
\operatorname{supp}_{\varpi}(A_i\mid B_i,\mathcal C_N)
=
\left\{
a:
\mathbb E[
\varpi_i\mathbbm{1}\{B_i,A_i=a\}
\mid
\mathcal C_N
]>0
\right\}.
\end{align*}
For $a\in\{g,\infty\}$, define
\begin{align*}
p_{a,z}^{g,l}
=
\Pr(G_i=a,Z_i^{g,l}=z\mid\mathcal C_N).
\end{align*}
For the CSE source support, define
\begin{align*}
p_{\infty,t}(x,h)
=
\Pr(G_i=\infty,X_i^d=x,H_{it}=h\mid\mathcal C_N).
\end{align*}
Define weighted cohort masses
\begin{align*}
s_{\varpi,g}
& =
\mathbb E[\varpi_i\mathbbm{1}\{G_i=g\}\mid\mathcal C_N], \\
s_{\varpi,l}^{\star}
& =
\sum_{g\in\mathcal G_l^\star}
s_{\varpi,g}.
\end{align*}
Event-time weights are
\begin{align*}
\omega_{\varpi,g,l}
=
\frac{s_{\varpi,g}}{s_{\varpi,l}^{\star}}.
\end{align*}
When the realized rollout is treated as fixed, the conditional probabilities and expectations above are interpreted as finite-population shares on the realized design.
For $Q\in\{DSE,CSE,DTE\}$, the event-time target is $\tau_{\varpi}^Q(l)=\sum_{g\in\mathcal G_l^\star}\omega_{\varpi,g,l}\tau_{\varpi}^Q(g,l)$.
When CSE is estimated by a parametric transported plug-in, $Q=CSE$ and $Q=DTE$ refer to the corresponding parametric transported targets unless the fitted contrast is correctly specified on the transported support.
Let $\mathcal T_\infty$ denote the finite set of calendar times at which the never-treated control-group spillover effect is reported.
The never-treated control-group CSE diagnostic, when reported, is treated as an additional smooth map of the same CSE first stage and the never-treated target distribution.
The stacked vector $\widehat\theta_N$ collects the component estimates defined in Section~\ref{sec:estimation}: the DSE source-cell means, the CSE first-stage parameter, the DSE and CSE target estimates, the never-treated CSE diagnostic components when reported, and the cohort-mass estimates used in event-time aggregation.
Let $\theta_N^0$ be the corresponding population value.
Appendix~\ref{app:estimating-equation-array} defines the stacked estimating-equation array $q_{i,N}(\theta)$ for these components.
Write $\bar q_N(\theta)=N^{-1}\sum_{i=1}^N q_{i,N}(\theta)$.
\begin{assumption}\label{assumption:inf_support}\label{assumption:identification}
For $c>0$, define $\mathcal O_t^C(c)=\{(x,h):p_{\infty,t}(x,h)\ge c \text{ and } p_{\infty,t}(x,0)\ge c\}$.
There exists a constant $c>0$ such that the following conditions hold with probability approaching one.
\begin{enumerate}
\item[(i)] For every retained $(g,l)$, the weighted DSE target support $\mathcal Z_{\varpi,g,l}=\operatorname{supp}_{\varpi}(Z_i^{g,l}\mid G_i=g,\mathcal C_N)$ is finite, and $\inf_{g,l}\inf_{z\in\mathcal Z_{\varpi,g,l}}\min\{p_{g,z}^{g,l},p_{\infty,z}^{g,l}\}\ge c$.
\item[(ii)] For $t=g+l$, the weighted CSE target support satisfies $\operatorname{supp}_{\varpi}((X_i^d,H_{it})\mid G_i=g,\mathcal C_N)\subseteq\mathcal O_t^C(c)$.
For every $t\in\mathcal T_\infty$ at which the never-treated control-group CSE diagnostic is reported, $\operatorname{supp}_{\varpi}((X_i^d,H_{it})\mid G_i=\infty,\mathcal C_N)\subseteq\mathcal O_t^C(c)$.
\item[(iii)] The weighted cohort masses satisfy $\inf_{l\in\mathcal L}\inf_{g\in\mathcal G_l^\star}s_{\varpi,g}\ge c$ and $\inf_{l\in\mathcal L}s_{\varpi,l}^{\star}\ge c$.
The analysis weights are nonnegative and satisfy $\max_{1\le i\le N}\varpi_i/\sum_{j=1}^N\varpi_j\to0$.
\item[(iv)] The empirical reporting rule uses a minimum-count threshold satisfying either $m_N=m$ for a fixed positive integer $m$, or $m_N\to\infty$ and $m_N/N\to0$.
\item[(v)] The empirical reporting rule is stable:
\begin{align*}
\Pr
\left(
\widehat{\mathcal G}_l^\star=\mathcal G_l^\star
\text{ for all }l\in\mathcal L
\,\middle|\,
\mathcal C_N
\right)
\to
1.
\end{align*}
\end{enumerate}
\end{assumption}
\begin{assumption}\label{assumption:inf_regular}
The following conditions hold.
\begin{enumerate}
\item[(i)] The stacked parameter and moment dimensions satisfy $\sup_N\max\{\dim(\theta_N),\dim\{q_{i,N}(\theta_N)\}\}<\infty$.
\item[(ii)] There exists a population value $\theta_N^0$ such that $m_N(\theta_N^0)=0$.
For some $r_N\downarrow0$ with $r_N\sqrt N\to\infty$, $\Pr(\|\widehat\theta_N-\theta_N^0\|\le r_N\mid\mathcal C_N)\to1$.
\item[(iii)] On the neighborhood $\mathcal N_N=\{\theta:\|\theta-\theta_N^0\|\le r_N\}$, $\bar q_N(\theta)$ is continuously differentiable.
With $R_N=\nabla_\theta m_N(\theta)\big|_{\theta=\theta_N^0}$, the sample Jacobian is locally stable:
\begin{align*}
\sup_{\theta\in\mathcal N_N}
\left\|
\nabla_\theta\bar q_N(\theta)-R_N
\right\|
=
o_p(1).
\end{align*}
\item[(iv)] There exist constants $0<c_R<C_R<\infty$ such that
\begin{align*}
c_R
\le
\lambda_{\min}(R_N^\top\Psi_N R_N)
\le
\lambda_{\max}(R_N^\top\Psi_N R_N)
\le
C_R
\end{align*}
with probability approaching one, and $\|\widehat\Psi_N-\Psi_N\|=o_p(1)$.
\item[(v)] Let $\widehat R_N=\nabla_\theta\bar q_N(\widehat\theta_N)$.
The stacked GMM estimator satisfies the first-order condition
\begin{align*}
\sqrt N
\left\|
\widehat R_N^\top
\widehat\Psi_N
\bar q_N(\widehat\theta_N)
\right\|
=
o_p(1).
\end{align*}
\end{enumerate}
\end{assumption}
Assumption~\ref{assumption:inf_support} is the population-margin version of the empirical admissibility rules in Section~\ref{subsec:estimation-aggregation}, including stability of the reported cohort-event cells.
Assumption~\ref{assumption:inf_regular} is a local regularity condition for the stacked GMM representation.
It requires that the estimator defined by the stacked criterion admit the first-order expansion generated by the sample moments.
In the just-identified target and cohort-mass blocks, this representation coincides with the closed-form component estimators from Section~\ref{sec:estimation}.
\begin{assumption}\label{assumption:inf_spatial}
Let
\begin{align*}
U_{i,N}
=
\left(
q_{i,N}(\theta_N^0)^\top,
\operatorname{vec}\{\nabla_\theta q_{i,N}(\theta_N^0)\}^\top
\right)^\top .
\end{align*}
Conditional on $\mathcal C_N$, the triangular array $\{U_{i,N}:i=1,\ldots,N\}$ satisfies the spatial $\psi$-dependence conditions of \citet{KOJEVNIKOV2021882}.
For some $p>4$, $\sup_N\max_{i\le N}\mathbb E[\|U_{i,N}\|^p\mid\mathcal C_N]<\infty$ with probability approaching one.
The covariance matrix
\begin{align*}
\Omega_N
=
\operatorname{Var}
\left(
\frac{1}{\sqrt N}
\sum_{i=1}^N q_{i,N}(\theta_N^0)
\,\middle|\,
\mathcal C_N
\right)
\end{align*}
satisfies
\begin{align*}
0<c_\Omega
\le
\lambda_{\min}(\Omega_N)
\le
\lambda_{\max}(\Omega_N)
\le
C_\Omega<\infty
\end{align*}
with probability approaching one.
\end{assumption}
\begin{assumption}\label{assumption:inf_shac}\label{assumption:shac}
Let $d_\rho$ be the shell-growth dimension of the metric space underlying $\rho(i,j)$, and let $p>4$ be as in Assumption~\ref{assumption:inf_spatial}.
The kernel $K:\mathbb R_+\to\mathbb R$ is bounded, continuous at zero, satisfies $K(0)=1$, and satisfies $K(u)=0$ for $u>1$.
The bandwidth $b_N$ satisfies
\begin{align*}
b_N
& \to
\infty, \\
b_N
& =
o\!\left(N^{1/(2d_\rho)}\right).
\end{align*}
Let $\widetilde\kappa_{N,s}$ denote the shell-level dependence coefficients associated with the stacked event-time influence-row array.
They satisfy
\begin{align*}
\sum_{s=1}^{\infty}
s^{d_\rho-1}
\widetilde\kappa_{N,s}^{\,1-2/p}
=
O(1),
\end{align*}
\begin{align*}
\sum_{s=1}^{\infty}
|K(s/b_N)-1|
s^{d_\rho-1}
\widetilde\kappa_{N,s}^{\,1-2/p}
\to
0,
\end{align*}
and
\begin{align*}
\frac{b_N^{2d_\rho}}{N}
\sum_{s=0}^{\infty}
s^{d_\rho-1}
\widetilde\kappa_{N,s}^{\,1-4/p}
\to
0.
\end{align*}
For every reported component $Q\in\{DSE,CSE,DTE\}$ and every fixed pair of event times $l,l'$, the feasible influence rows are asymptotically equivalent to their population counterparts in the spatial HAC quadratic form:
\begin{align*}
&
\frac{1}{N}
\sum_{i=1}^{N}
\sum_{j=1}^{N}
K\!\left(\rho(i,j)/b_N\right)
\left\{
\widehat\varphi_{i,N}^Q(l)\widehat\varphi_{j,N}^Q(l')
-
\varphi_{i,N}^Q(l)\varphi_{j,N}^Q(l')
\right\} \\
& =
o_p(1).
\end{align*}
The same condition holds for the never-treated control-group spillover rows $\widehat\varphi_{i,N}^{CSE,\infty}(t)$.
For the conservative-variance statement, the kernel-weight matrix $\{K(\rho(i,j)/b_N)\}_{i,j=1}^{N}$ is positive semidefinite with probability approaching one.
\end{assumption}
Assumptions~\ref{assumption:inf_spatial} and \ref{assumption:inf_shac} are imposed directly on the stacked score and event-time influence arrays.
The conditional CLT used below is obtained by applying the $\psi$-dependence limit theory used in \citet{xu2025difference}; it is not imposed as a separate assumption.
Assumption~\ref{assumption:inf_shac} specifies the kernel, bandwidth, and dependence conditions used by the spatial HAC covariance estimator.
The bandwidth grows so that increasingly distant units enter the covariance estimate, but slowly enough relative to the expanding domain.
Together with Assumption~\ref{assumption:inf_spatial}, these conditions allow the spatial HAC covariance result of \citet{xu2025difference} to be applied to the finite-dimensional event-time influence-row array.
\subsection{Asymptotic properties}
Let $a_l^Q(\theta)$ denote the event-time aggregation map for $Q\in\{DSE,CSE,DTE\}$.
The influence row for the event-time path is
\begin{align*}
\varphi_{i,N}^Q(l)
=
-
\nabla_\theta a_l^Q(\theta_N^0)^\top
(R_N^\top\Psi_N R_N)^{-1}
R_N^\top\Psi_N
q_{i,N}(\theta_N^0).
\end{align*}
\begin{theorem}\label{thm:event_time_expansion}\label{prop:gmm_asymptotic}\label{prop:event_time_gmm}
Suppose Assumptions~\ref{assumption:inf_support}--\ref{assumption:inf_spatial} hold.
For each fixed $l\in\mathcal L$ and $Q\in\{DSE,CSE,DTE\}$,
\begin{align*}
\sqrt N
\{\widehat\tau_\varpi^Q(l)-\tau_\varpi^Q(l)\}
=
\frac{1}{\sqrt N}
\sum_{i=1}^N
\varphi_{i,N}^Q(l)
+
o_p(1).
\end{align*}
Moreover,
\begin{align*}
\varphi_{i,N}^{DTE}(l)
=
\varphi_{i,N}^{DSE}(l)+\varphi_{i,N}^{CSE}(l).
\end{align*}
\end{theorem}
\begin{proof}
See Appendix~\ref{app:inference-proofs}.
\end{proof}
\begin{remark}\label{prop:cse_infty_inference}
The same stacked representation covers the never-treated control-group spillover effect.
For $a_{\infty,t}^{CSE}(\theta)$ equal to the map that evaluates the fitted CSE contrast over the weighted never-treated distribution at calendar time $t$,
\begin{align}
\sqrt N
\left\{
\widehat\tau_{\varpi,\infty}^{CSE}(t)
-
\tau_{\varpi,\infty}^{CSE,param}(t)
\right\}
=
\frac{1}{\sqrt N}
\sum_{i=1}^N
\varphi_{i,N}^{CSE,\infty}(t)
+
o_p(1),
\label{eq:cse-infty-if}
\end{align}
where
\begin{align}
\varphi_{i,N}^{CSE,\infty}(t)
=
-
\nabla_\theta a_{\infty,t}^{CSE}(\theta_N^0)^\top
(R_N^\top\Psi_N R_N)^{-1}
R_N^\top\Psi_N
q_{i,N}(\theta_N^0).
\label{eq:cse-infty-if-row}
\end{align}
If $c_t(v,h;\eta^0)=r_{\infty,t}(x,h)$ on the never-treated source support for $(X_i^d,V_i,H_{it})$, the target is the causal never-treated control-group spillover effect.
\end{remark}
Define the feasible influence row
\begin{align*}
\widehat\varphi_{i,N}^Q(l)
=
-
\nabla_\theta a_l^Q(\widehat\theta_N)^\top
(\widehat R_N^\top\widehat\Psi_N\widehat R_N)^{-1}
\widehat R_N^\top\widehat\Psi_N
q_{i,N}(\widehat\theta_N),
\end{align*}
where
\begin{align*}
\widehat R_N
=
\frac{1}{N}
\sum_{i=1}^N
\nabla_\theta q_{i,N}(\widehat\theta_N).
\end{align*}
The feasible DTE row is formed as the sum of the DSE and CSE rows before covariance estimation:
\begin{align*}
\widehat\varphi_{i,N}^{DTE}(l)
=
\widehat\varphi_{i,N}^{DSE}(l)+\widehat\varphi_{i,N}^{CSE}(l).
\end{align*}
For the never-treated control-group spillover effect, define
\begin{align*}
\widehat\varphi_{i,N}^{CSE,\infty}(t)
=
-
\nabla_\theta a_{\infty,t}^{CSE}(\widehat\theta_N)^\top
(\widehat R_N^\top\widehat\Psi_N\widehat R_N)^{-1}
\widehat R_N^\top\widehat\Psi_N
q_{i,N}(\widehat\theta_N).
\end{align*}
For the spillover change in the never-treated comparison group, when both $g$ and $t_0(g)$ are included in $\mathcal T_\infty$, the feasible influence row is
\begin{align*}
\widehat\varphi_{i,N}^{\Delta CSE,\infty}(g)
=
\widehat\varphi_{i,N}^{CSE,\infty}(g)
-
\widehat\varphi_{i,N}^{CSE,\infty}(t_0(g)).
\end{align*}
The spatial HAC estimator is written in the uncentered form used in the finite-population covariance decomposition.
For $Q\in\{DSE,CSE,DTE\}$, define the spatial HAC covariance estimator for the event-time path by
\begin{align}
\widehat\Gamma_Q(l,l')
=
\frac{1}{N}
\sum_{i=1}^N
\sum_{j=1}^N
K(\rho(i,j)/b_N)
\widehat\varphi_{i,N}^Q(l)
\widehat\varphi_{j,N}^Q(l').
\label{eq:shac_event_cov}
\end{align}
Here $K(\cdot)$ is the spatial kernel, $b_N$ is the bandwidth, and $\rho(i,j)$ is the geographic or network distance from Assumptions~\ref{assumption:inf_spatial} and \ref{assumption:inf_shac}.
For the never-treated control-group spillover effect, define
\begin{align}
\widehat\Gamma_{\infty}^{CSE}(t,t')
=
\frac{1}{N}
\sum_{i=1}^N
\sum_{j=1}^N
K(\rho(i,j)/b_N)
\widehat\varphi_{i,N}^{CSE,\infty}(t)
\widehat\varphi_{j,N}^{CSE,\infty}(t').
\label{eq:shac-cse-infty}
\end{align}
The pointwise standard error is
\begin{align*}
\widehat{se}_{\infty}^{CSE}(t)
=
\sqrt{
\frac{
\widehat\Gamma_{\infty}^{CSE}(t,t)
}{
N
}
}.
\end{align*}
The corresponding pointwise interval is
\begin{align}
\widehat C_{\infty}^{CSE,pt}(t)
=
\left[
\widehat\tau_{\varpi,\infty}^{CSE}(t)
\pm
z_{1-\alpha/2}
\sqrt{
\frac{
\widehat\Gamma_{\infty}^{CSE}(t,t)
}{
N
}
}
\right].
\label{eq:cse-infty-ci}
\end{align}
The spatial HAC variance for $\widehat\Delta_{\infty}^{CSE}(g)$ is computed by replacing $\widehat\varphi_{i,N}^{CSE,\infty}(t)$ in \eqref{eq:shac-cse-infty} with $\widehat\varphi_{i,N}^{\Delta CSE,\infty}(g)$.
Let
\begin{align*}
\Gamma_Q(l,l')
=
\operatorname{Cov}
\left(
\frac{1}{\sqrt N}\sum_{i=1}^N\varphi_{i,N}^Q(l),
\frac{1}{\sqrt N}\sum_{i=1}^N\varphi_{i,N}^Q(l')
\,\middle|\,
\mathcal C_N
\right),
\end{align*}
and for the never-treated control-group spillover effect let
\begin{align*}
\Gamma_{\infty}^{CSE}(t,t')
=
\operatorname{Cov}
\left(
\frac{1}{\sqrt N}\sum_{i=1}^N\varphi_{i,N}^{CSE,\infty}(t),
\frac{1}{\sqrt N}\sum_{i=1}^N\varphi_{i,N}^{CSE,\infty}(t')
\,\middle|\,
\mathcal C_N
\right),
\end{align*}
and define
\begin{align*}
\mu_{i,N}^Q(l)
=
\mathbb E[
\varphi_{i,N}^Q(l)
\mid
\mathcal C_N
].
\end{align*}
Define the finite-population centering component
\begin{align*}
\Gamma_Q^E(l,l')
=
\frac{1}{N}
\sum_{i=1}^N
\sum_{j=1}^N
K(\rho(i,j)/b_N)
\mu_{i,N}^Q(l)
\mu_{j,N}^Q(l').
\end{align*}
Let
\begin{align*}
\Gamma_Q^+(l,l')
=
\Gamma_Q(l,l')+\Gamma_Q^E(l,l').
\end{align*}
For the never-treated control-group spillover effect, define $\mu_{i,N}^{CSE,\infty}(t)$, $\Gamma_{\infty}^{CSE,E}(t,t')$, and $\Gamma_{\infty}^{CSE,+}(t,t')$ analogously by replacing $\varphi_{i,N}^Q(l)$ with $\varphi_{i,N}^{CSE,\infty}(t)$ and $\Gamma_Q(l,l')$ with $\Gamma_{\infty}^{CSE}(t,t')$.
\begin{prop}\label{thm:shac_inference}\label{prop:shac_covariance}
Suppose Assumptions~\ref{assumption:inf_support}--\ref{assumption:inf_shac} hold.
For $Q\in\{DSE,CSE,DTE\}$, the spatial HAC estimator $\widehat\Gamma_Q$ converges to the spatial HAC covariance target in \citet{xu2025difference} for the event-time influence process.
For each fixed $l,l'\in\mathcal L$ and $Q\in\{DSE,CSE,DTE\}$,
\begin{align*}
\widehat\Gamma_Q(l,l')-\Gamma_Q^+(l,l')
=
o_p(1).
\end{align*}
If the finite-population centering component vanishes, or is consistently removed or estimated away, then
\begin{align*}
\widehat\Gamma_Q(l,l')-\Gamma_Q(l,l')
=
o_p(1).
\end{align*}
The positive-semidefinite kernel condition in Assumption~\ref{assumption:inf_shac} implies
\begin{align*}
\Gamma_Q^+(l,l)\ge \Gamma_Q(l,l)
\end{align*}
for each fixed $l$.
The usual uncentered spatial HAC variance estimate is therefore asymptotically conservative for the finite-population conditional variance.
The same statements hold for $\widehat\Gamma_{\infty}^{CSE}(t,t')$ over fixed $t,t'\in\mathcal T_\infty$, with $\Gamma_{\infty}^{CSE,+}(t,t')$ replacing $\Gamma_Q^+(l,l')$.
\end{prop}
\begin{proof}
See Appendix~\ref{app:inference-proofs}.
\end{proof}
For $Q\in\{DSE,CSE,DTE\}$, the pointwise interval is
\begin{align*}
\widehat C_Q^{pt}(l)
=
\left[
\widehat\tau_{\varpi}^Q(l)
\pm
z_{1-\alpha/2}
\sqrt{\frac{\widehat\Gamma_Q(l,l)}{N}}
\right].
\end{align*}
For $Q=CSE$ and $Q=DTE$, the target is the parametric transported target unless the CSE contrast is correctly specified on the transported target support.
\begin{cor}\label{prop:pointwise_ci}
Under the conditions of Theorem~\ref{thm:event_time_expansion} and Proposition~\ref{prop:shac_covariance}, for each fixed $l\in\mathcal L$ and $Q\in\{DSE,CSE,DTE\}$, the pointwise normal interval has asymptotic coverage at least $1-\alpha$.
If the centering correction vanishes or is consistently removed, the coverage converges to $1-\alpha$.
The same exact-versus-conservative coverage statements apply to $\widehat C_{\infty}^{CSE,pt}(t)$, with $\Gamma_{\infty}^{CSE}(t,t)$ and $\Gamma_{\infty}^{CSE,+}(t,t)$ replacing $\Gamma_Q(l,l)$ and $\Gamma_Q^+(l,l)$.
\end{cor}
\begin{proof}
See Appendix~\ref{app:inference-proofs}.
\end{proof}
\begin{remark}
The main tables report pointwise confidence intervals based on $\widehat\Gamma_Q(l,l)$.
If simultaneous inference over the finite set $\mathcal L$ is desired, one can simulate
\begin{align*}
Z_Q\sim \mathcal N(0,\widehat\Gamma_Q)
\end{align*}
and use the quantile of
\begin{align*}
\max_{l\in\mathcal L}
\left|
\frac{Z_Q(l)}{\sqrt{\widehat\Gamma_Q(l,l)}}
\right|
\end{align*}
to form Gaussian simultaneous bands.
This step uses the estimated Gaussian limit.
\end{remark}
\section{Monte Carlo simulation}\label{sec:monte-carlo}
\subsection{Design}
The simulations use finite populations in which DSE, CSE, and DTE are known functions of the simulated potential outcomes.
They ask whether the proposed estimators recover the targets on retained support and whether standard group-versus-never DID and \citet{callaway2021difference} ATT estimators implemented without exposure can be close to the DSE target while missing DTE.
DTE is formed on the same admissible support as the sum of DSE and CSE,
\begin{align*}
\tau^{DTE}(l)
=
\tau^{DSE}(l)+\tau^{CSE}(l).
\end{align*}
The simulations use the unweighted special case $\varpi_i=1$.
All main designs use a balanced six-period panel.
Cohorts are $G_i\in\{3,4,5,\infty\}$, treatment is absorbing with $D_{it}=\mathbbm{1}\{t\geq G_i\}$, there is no anticipation, and the baseline period is $t_0(g)=g-1$.
We report event times $l\in\{0,1,2\}$.
For a given event time $l$, only cohorts satisfying $g+l\leq T$ enter the corresponding event-time aggregate.
Units are located on an open line network with row-normalized weights $w_{ij}$.
The maintained exposure state is a coarsened line-network exposure share, $H_{it}=b(\widetilde H_{it})\in\{0,\text{low},\text{high}\}$.
The map keeps zero exposure at $0$, maps values in $(0,0.5]$ to low exposure, and maps values in $(0.5,1]$ to high exposure.
The dose score is $q(0)=0$, $q(\text{low})=1$, and $q(\text{high})=2$.
For all main designs, $X_i$, $\alpha_i$, and $\varepsilon_{it}$ are drawn independently from $N(0,1)$.
The time effect is $\lambda_t=0.2(t-1)$.
The control-state spillover slope is $\rho_t=0$ for $t\in\{1,2\}$ and $\rho_t=0.3$ for $t\geq3$.
The own-treatment profile is $\tau_0=1$, $\tau_1=1.5$, and $\tau_l=2$ for $l\geq2$.
The DGPs depend on $q(H_{it})$, not directly on the raw exposure index.
Untreated potential outcomes include the control-state spillover slope $\rho_t$, while treated potential outcomes add the own-treatment profile and, when $\kappa_d\neq0$, a treated-side exposure interaction.
Appendix~\ref{app:monte-carlo-details} gives the exact potential-outcome equations and finite-population truth formulas.
DGP1 sets $\kappa_d=0$, so own treatment and spillover exposure are additively separable.
In this design, $\tau^{PDE}(g,l)=\tau^{DSE}(g,l)$.
DGP2 sets $\kappa_d=0.4$.
This is the main interaction design: the switching effect, DTE, and the global PDE generally differ.
DGP3 uses the same potential-outcome structure as DGP2 but modifies the cohort assignment and support environment so that some exposure states faced by treated cohorts are less frequently represented among never-treated units.
DGP3 instead uses a weakly balanced support-aware line assignment with block size 12 and equal cohort shares across $G_i\in\{3,4,5,\infty\}$.
We interpret DGP3 as a support/admissibility stress design.
For each replication, the truth objects used for evaluation are finite-population averages of the simulated potential outcomes on the retained admissible support.
DTE is formed as the same-support sum of DSE and CSE, equivalently the direct contrast from the realized adoption regime to the pure-control regime.
\subsection{Estimators and retained support}
The reported estimators are the same-state DSE estimator, the parametric transported plug-in estimator for CSE, and the DTE estimator obtained from their sum.
For DSE, define $S_i^{g,l}=(H_{it},H_{it_0})$ and $Z_i^{g,l}=(X_i^d,S_i^{g,l})$, with $Z_i^{g,l}=S_i^{g,l}$ when no covariate stratification is used.
The DSE estimator compares treated and never-treated units within the retained support of $Z_i^{g,l}$.
For CSE, the never-treated sample is used to estimate an effect contrast $c_t(x,h;\eta)\approx\mu_t(x,h)-\mu_t(x,0)$, normalized by $c_t(x,0;\eta)=0$.
Let $t=g+l$.
The estimator then evaluates the fitted contrast over the cohort-$g$ exposure distribution:
\begin{align*}
\widehat\tau^{CSE}(g,l)
=
\frac{1}{n_g}
\sum_{i:G_i=g}
c_t(X_i^d,H_{it};\widehat\eta).
\end{align*}
The DTE estimator is the post-estimation sum:
\begin{align*}
\widehat\tau^{DTE}(g,l)
=
\widehat\tau^{DSE}(g,l)
+
\widehat\tau^{CSE}(g,l).
\end{align*}
At the event-time level, DSE, CSE, and DTE use the same admissible cohort set and the same weights, so $\widehat\tau^{DTE}(l)=\widehat\tau^{DSE}(l)+\widehat\tau^{CSE}(l)$ holds exactly in each replication.
The benchmark estimators are a cohort-specific DID contrast and the \citet{callaway2021difference} group-time ATT estimator implemented without exposure.
For each event time $l$, let $\widehat{\mathcal G}_l^\star$ be the set of cohorts for which both DSE and CSE support requirements are satisfied.
The support rule uses a pre-specified minimum cell count $m_N=5$.
A cohort-event cell is admissible only when all required treated and never-treated source cells for DSE and CSE meet this threshold.
For $g\in\widehat{\mathcal G}_l^\star$, event-time aggregates use
\begin{align*}
\widehat\omega_{g,l}
=
\frac{n_g}
{\sum_{g'\in\widehat{\mathcal G}_l^\star}n_{g'}}.
\end{align*}
For $Q\in\{DSE,CSE,DTE\}$,
\begin{align*}
\widehat\tau^Q(l)
=
\sum_{g\in\widehat{\mathcal G}_l^\star}
\widehat\omega_{g,l}\widehat\tau^Q(g,l).
\end{align*}
For event-time comparisons, benchmark cohort-time estimates are aggregated over the same $\widehat{\mathcal G}_l^\star$ and with the same weights $\widehat\omega_{g,l}$ whenever the corresponding cohort-time benchmark is available.
These benchmarks are reported as deviations from the relevant target on retained support: $\tau^{DSE}_{rep}(l)$ in the DSE table and $\tau^{DTE}_{rep}(l)$ in the DTE table.
Bias, RMSE, and coverage are computed relative to
\begin{align*}
\tau^Q_{rep}(l)
=
\sum_{g\in\widehat{\mathcal G}_l^\star}
\widehat\omega_{g,l}\tau^Q(g,l),
\end{align*}
rather than the full-cohort target that includes inadmissible cells.
The tables report bias or benchmark deviation, RMSE, retained treated mass or availability when relevant, and pointwise coverage.
\subsection{Inference and performance measures}
For DGP1--DGP3, pointwise intervals for DSE, CSE, and DTE use the sandwich covariance from the stacked estimating equations.
The spatial HAC kernel is applied over line-network distance with bandwidth $b_N=\lceil N^{1/3}\rceil$, which gives $b_N=6$ for $N=200$ and $b_N=8$ for $N=500$.
The covariance is computed after forming the event-time influence rows for DSE, CSE, and DTE; for DTE, this keeps the covariance between the DSE and CSE rows.
The main tables report bias or benchmark deviation, RMSE, and pointwise coverage.
For benchmark rows, the Bias column is the deviation from the target on retained support.
Coverage is computed conditional on interval availability.
The final main-text runs use $R=1000$ Monte Carlo replications for each design-size combination.
Tables~\ref{tab:sim_benchmark_dte_deviation} and \ref{tab:sim_proposed_componentwise} separate the benchmark comparison that ignores exposure from the performance of the proposed DSE, CSE, and DTE estimators on retained support.
First, DID benchmarks that ignore exposure generally differ from DTE when exposure affects outcomes.
Table~\ref{tab:sim_benchmark_dte_deviation} reports standard group-versus-never DID and \citet{callaway2021difference} benchmarks implemented while ignoring exposure.
These estimators answer no-interference DID questions; interpreting them as DTE estimators treats exposure as irrelevant both for the estimand and for comparison-group trends.
In the pooled DTE comparison, the benchmark deviations range from about $-0.31$ to $-0.36$ across the reported designs, and coverage for $N=500$ falls to $0.066$--$0.165$.
\begin{table}[!t]
\centering
\caption{Spillover-ignorant benchmark deviations from the dynamic total-effect target}
\label{tab:sim_benchmark_dte_deviation}
\begin{threeparttable}
\small
\begin{tabular}{llcccccc}
\toprule
DGP & Method & \multicolumn{3}{c}{$N=200$} & \multicolumn{3}{c}{$N=500$} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& & Bias & RMSE & Coverage & Bias & RMSE & Coverage \\
\midrule
DGP1 & Standard DID & -0.311 & 0.354 & 0.498 & -0.308 & 0.329 & 0.149 \\
& Callaway and Sant'Anna & -0.311 & 0.355 & 0.511 & -0.308 & 0.328 & 0.152 \\
DGP2 & Standard DID & -0.305 & 0.347 & 0.545 & -0.311 & 0.331 & 0.165 \\
& Callaway and Sant'Anna & -0.306 & 0.348 & 0.545 & -0.311 & 0.331 & 0.156 \\
DGP3 & Standard DID & -0.359 & 0.398 & 0.387 & -0.362 & 0.379 & 0.067 \\
& Callaway and Sant'Anna & -0.358 & 0.398 & 0.408 & -0.362 & 0.379 & 0.066 \\
\bottomrule
\end{tabular}
\begin{tablenotes}[flushleft]
\footnotesize
\item Notes: Bias and coverage are computed relative to the retained-support DTE target. Standard DID intervals use spatial HAC covariance estimates; the Callaway-Sant'Anna benchmark uses weighted influence-function intervals. Each design-size combination uses R=1,000 replications.
\end{tablenotes}
\end{threeparttable}
\end{table}
This comparison is not a claim that the proposed estimators dominate the benchmarks for the switching target; DTE includes CSE, which the benchmark estimators do not estimate.
Second, Table~\ref{tab:sim_proposed_componentwise} reports the proposed DSE, CSE, and DTE estimators.
Same-state comparisons estimate DSE, the never-treated first-stage contrast estimates the control-state spillover response used for CSE, and DTE is formed by adding the two components on the same admissible support.
Table~\ref{tab:sim_proposed_componentwise} shows small biases and pointwise coverage close to the nominal 95 percent level across the reported designs.
\begin{table}[!t]
\centering
\caption{Finite-sample performance of the proposed DSE, CSE, and DTE estimators}
\label{tab:sim_proposed_componentwise}
\begin{threeparttable}
\small
\begin{tabular}{llcccccc}
\toprule
DGP & Method & \multicolumn{3}{c}{$N=200$} & \multicolumn{3}{c}{$N=500$} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& & Bias & RMSE & Coverage & Bias & RMSE & Coverage \\
\midrule
DGP1 & Proposed DSE & -0.005 & 0.263 & 0.943 & -0.004 & 0.152 & 0.945 \\
& Proposed CSE & -0.006 & 0.296 & 0.921 & -0.003 & 0.174 & 0.944 \\
& Proposed DTE & -0.011 & 0.401 & 0.937 & -0.007 & 0.234 & 0.938 \\
DGP2 & Proposed DSE & 0.006 & 0.269 & 0.942 & -0.002 & 0.150 & 0.954 \\
& Proposed CSE & 0.006 & 0.299 & 0.926 & 0.009 & 0.182 & 0.934 \\
& Proposed DTE & 0.012 & 0.406 & 0.932 & 0.007 & 0.239 & 0.941 \\
DGP3 & Proposed DSE & -0.001 & 0.298 & 0.938 & 0.006 & 0.166 & 0.950 \\
& Proposed CSE & -0.021 & 0.350 & 0.931 & 0.002 & 0.231 & 0.941 \\
& Proposed DTE & -0.021 & 0.430 & 0.924 & 0.008 & 0.279 & 0.940 \\
\bottomrule
\end{tabular}
\begin{tablenotes}[flushleft]
\footnotesize
\item Notes: Bias, RMSE, and coverage are computed relative to each estimator's retained-support target. Confidence intervals use spatial HAC covariance estimates. Each design-size combination uses R=1,000 replications.
\end{tablenotes}
\end{threeparttable}
\end{table}
DGP3 is included to show the role of support restrictions.
When overlap is weak, the estimator reports common-support targets and availability becomes part of the interpretation.
\section{Application}\label{sec:application}
\subsection{Background and data}
The estimates in this section use the 50-mile exposure mapping and the identifying restrictions in Sections~\ref{sec:identification}--\ref{sec:inference}.
For CSE, the first stage is fit in the never-treated source sample and then transported to treated cohorts.
Under these maintained conditions, the CSE and DTE estimates have a causal interpretation; otherwise, they are transported parametric targets.
The empirical study revisits a staggered county-level policy setting in which nearby counties can be exposed before or after their own adoption date.
A conventional DID comparison between treated and never-treated counties in this setting subtracts trends from never-treated counties that may themselves be exposed to nearby Community Health Centers.
It also assigns any spillover received by treated counties after adoption to the same coefficient as own adoption.
The empirical exercise therefore reports the components in
\begin{align*}
\tau^{DTE}(g,l)=\tau^{DSE}(g,l)+\tau^{CSE}(g,l),
\end{align*}
rather than treating a single DID coefficient as an own-treatment effect.
We use the county-year mortality panel from the Community Health Centers replication files for \citet{bailey2015war}.
The maintained analysis sample covers 1959--1978, target cohorts adopt between 1965 and 1974, and exposure is a positive-exposure indicator for CHC adoption within 50 miles.
Older-adult population weights define the application target distribution.
The primary outcome is the mortality rate for individuals aged 50 and older, measured in deaths per 100,000 older-adult residents.
The panel also retains all-age mortality, age-specific mortality rates, cause-specific mortality rates, population weights, baseline demographic controls, physician supply, and the event-time variables used in the original replication files.
County identifiers are standardized to five-digit FIPS codes, and the analysis excludes Los Angeles County, Cook County, and New York County.
Treatment timing is defined by the first year in which county $i$ receives a Community Health Center, denoted $G_i$.
The target treated cohorts are $g\in\{1965,\dots,1974\}$, and never-treated counties are assigned $G_i=\infty$.
Counties first treated after 1974 are retained when constructing exposure to nearby adopters but are not used as target treated cohorts or never-treated comparison counties in the component estimators.
Treatment is absorbing, so the realized treatment indicator is
\begin{align*}
D_{it}=\mathbbm{1}\{t\ge G_i\}.
\end{align*}
Event time is indexed by $l=t-g$, so calendar time and event time satisfy $t=g+l$.
In the application, we set the anticipation window to zero, so the baseline period is
\begin{align*}
t_0(g)=g-1.
\end{align*}
The resulting empirical long difference is therefore
\begin{align*}
\Delta Y_i^{g,l}=Y_{i,g+l}-Y_{i,g-1}.
\end{align*}
Geographic exposure is constructed from the 2015 Census TIGER/Line county shapefile.
County representative points are obtained from polygon geometry, and distances are measured between representative points in miles.
The 50-mile exposure mapping is plausible in this setting because Community Health Centers may serve residents outside the adopting county.
Nearby counties may also be affected through travel to a center, provider networks, referral patterns, or reduced congestion in alternative sources of care.
The binary 50-mile indicator is a maintained exposure mapping for the DID design, not a claim that the true structural spillover function has a sharp cutoff at 50 miles.
Let $d_{ij}$ denote the distance between counties $i$ and $j$, and define the 50-mile neighbor set
\begin{align*}
\mathcal N_{50}(i)
=
\{j\neq i:d_{ij}\le 50\}.
\end{align*}
The maintained exposure count is
\begin{align*}
H_{it}^{raw}
=
\sum_{j\in\mathcal N_{50}(i)}
\mathbbm{1}\{t\ge G_j\}.
\end{align*}
The empirical exposure state is binary:
\begin{align*}
H_{it}
=
\begin{cases}
0, & \text{if } H_{it}^{raw}=0, \\
\text{positive}, & \text{if } H_{it}^{raw}>0.
\end{cases}
\end{align*}
For each cohort-event-time pair $(g,l)$, the empirical two-date exposure state is
\begin{align*}
S_i^{g,l}=(H_{i,g+l},H_{i,g-1}).
\end{align*}
Let $X_i^d$ denote the finite baseline strata used in the DSE comparison, fixed before estimation, and define
\begin{align*}
Z_i^{g,l}
=
(X_i^d,S_i^{g,l}).
\end{align*}
In the reported CHC estimates, no additional finite baseline covariate strata are imposed beyond the two-date exposure state, so $X_i^d$ is degenerate and $Z_i^{g,l}=S_i^{g,l}$.
The first-stage basis $V_i$ below is separate from this support object and is used to parameterize the CSE source response.
This construction allows outcomes to depend on a county's own adoption timing and on its realized exposure state before and after the relevant event time.
The current exposure state enters the potential outcome, while the two-date state $S_i^{g,l}$ conditions the same-state long-difference comparison.
\begin{figure}[!t]
\centering
\begin{minipage}{0.48\textwidth}
\centering
\includegraphics[width=\linewidth]{application/figures/fig_application_spatial_spillover_exposure_1964.png}
{\footnotesize (a) 1964}
\end{minipage}
\hfill
\begin{minipage}{0.48\textwidth}
\centering
\includegraphics[width=\linewidth]{application/figures/fig_application_spatial_spillover_exposure_1970.png}
{\footnotesize (b) 1970}
\end{minipage}
\par\medskip
\begin{minipage}{0.48\textwidth}
\centering
\includegraphics[width=\linewidth]{application/figures/fig_application_spatial_spillover_exposure_1975.png}
{\footnotesize (c) 1975}
\end{minipage}
\hfill
\begin{minipage}{0.48\textwidth}
\centering
\includegraphics[width=\linewidth]{application/figures/fig_application_spatial_spillover_exposure_1978.png}
{\footnotesize (d) 1978}
\end{minipage}
\caption{Spatial diffusion of spillover exposure in Community Health Centers data}
\label{fig:application_spatial_spillover_exposure}
\begin{minipage}{0.92\linewidth}
\footnotesize Notes: Gray counties are unexposed and untreated, orange counties are exposed but untreated, and blue counties are treated. Exposure is the positive-exposure indicator for CHC adoption within 50 miles. The panels classify in-sample counties by realized adoption and exposure status in the indicated calendar year.
\end{minipage}
\end{figure}
\subsection{Empirical estimation}
The empirical estimators use the cohort-event-time objects defined in Sections~\ref{sec:framework}--\ref{sec:inference}.
For each admissible pair $(g,l)$, the comparison sample contains counties in cohort $g$ and never-treated counties.
The DSE component is estimated with the saturated long-difference specification from Section~\ref{subsec:main-estimators}, using the retained state $Z_i^{g,l}$ and the older-adult population weight.
Thus the empirical DSE regression for each cell is
\begin{align*}
\Delta Y_i^{g,l}
=
\sum_{z}
\alpha_{z,g,l}\mathbbm{1}\{Z_i^{g,l}=z\}
+
\sum_{z}
\beta_{z,g,l}\mathbbm{1}\{G_i=g\}\mathbbm{1}\{Z_i^{g,l}=z\}
+
u_i^{g,l}.
\end{align*}
The estimated DSE is the weighted average of the fitted state-specific switching contrasts over the weighted empirical distribution of $Z_i^{g,l}$ in cohort $g$.
The reported estimands are older-adult-population-weighted, corresponding to $\varpi_i$ equal to the baseline older-adult population weight.
The reported DSE is therefore not a rich-controls re-estimation of the \citet{bailey2015war} event-study specification.
It is a comparison defined by exposure-state cells and older-adult population weights.
The CSE component is estimated with the structured residualized binary-positive first stage.
In the application, for $t\ge2$, the source outcome and exposure indicator are
\begin{align*}
R_{it} & = Y_{it}-Y_{i1}, \ P_{it} = \mathbbm{1}\{H_{it}>0\}.
\end{align*}
The structured first stage writes
\begin{align*}
\mathbb E[R_{it}\mid G_i=\infty,V_i,H_{it}]
=
a_t(V_i;\alpha)+P_{it}B_t(V_i)^\top\beta,
\end{align*}
so that
\begin{align*}
c_t(V_i,0;\beta)
& =
0, \\
c_t(V_i,H_{it};\beta)
& =
P_{it}B_t(V_i)^\top\beta.
\end{align*}
In the first-stage fit, $V_i$ consists of the mean pre-treatment outcome over 1960--1964, log 1960 older-adult population, physicians per capita, the baseline urban bucket, and Census region.
The baseline trend component $a_t(V_i;\alpha)$ includes these variables and a four-degree-of-freedom spline in calendar time.
The exposure-response basis $B_t(V_i)$ contains period indicators and positive-exposure interactions with the same baseline variables, so the binary-positive exposure contrast varies by calendar time after residualization.
The older-adult population weight is $\varpi_i$ in the weighted source criterion and target averages.
Together with transportability, the reported CSE and DTE estimates are causal only if this transported first-stage contrast is correctly specified on the treated-cohort support; otherwise they are transported parametric targets.
The empirical DTE is the post-estimation sum of DSE and CSE using the same admissible cells and weights.
Inference uses the sandwich covariance from the stacked estimating equations, with spatial HAC weighting.
Confidence intervals are pointwise normal intervals based on the estimated spatial HAC covariance matrix.
The same stacked covariance is used for DSE, CSE, and DTE, so the covariance between the DSE and CSE components is retained in inference for DTE.
The reported event-time summaries aggregate over the common admissible cohort set.
The main reporting windows are $l=0$ and $l=1,\dots,4$ for DSE, CSE, and DTE, together with a pre-adoption CSE window at $l=-1$.
\subsection{Results}
All estimates in this section use the 50-mile exposure mapping.
The CSE learns a spillover response from never-treated counties and evaluates that response over the exposure distributions faced by treated cohorts.
The CSE and DTE estimates therefore have a causal interpretation under the transportability condition and the maintained first-stage specification; otherwise, they should be read as transported parametric targets.
Figure~\ref{fig:application_spatial_spillover_exposure} shows that exposure to nearby CHCs is not rare in the analysis sample.
This pattern is consistent with the institutional setting: CHCs expanded access to low-cost primary care, reduced travel and access barriers, and could affect residents outside the county that formally received a center.
The estimates should therefore be interpreted as effects under the exposure distribution generated by the realized CHC rollout, not as effects of an arbitrary counterfactual rollout rule.
The mortality results are consistent with the conclusion in \citet{bailey2015war} that CHCs reduced older-adult mortality.
The spillover-aware decomposition adds that the DTE is larger in magnitude than the benchmark that ignores exposure.
In the adoption year, Table~\ref{tab:application_benchmark_comparison} reports a DTE of $-39.983$ deaths per 100,000, compared with $-24.023$ from the standard DID benchmark.
At event time $l=4$, the DTE is $-113.410$, compared with $-74.868$ for the benchmark.
Thus, the benchmark that ignores exposure is smaller in magnitude than the estimated DTE at each reported event time.
\begin{table}[!htbp]
\centering
\caption{Benchmark comparisons for older-adult mortality using Community Health Centers data}
\label{tab:application_benchmark_comparison}
\small
\setlength{\tabcolsep}{4pt}
\begin{tabular}{lcccc}
\toprule
Event time & DSE & DTE & standard DID & \shortstack{Bailey and\\Goodman-Bacon (2015)} \\
\midrule
$l=0$ & \shortstack[c]{-24.053\\{\footnotesize [-48.944, 0.838]}} & \shortstack[c]{-39.983\\{\footnotesize [-68.604, -11.362]}} & \shortstack[c]{-24.023\\{\footnotesize [-47.225, -0.820]}} & \shortstack[c]{-22.930\\{\footnotesize [-42.839, -3.021]}} \\
$l=1$ & \shortstack[c]{-33.902\\{\footnotesize [-57.921, -9.884]}} & \shortstack[c]{-55.823\\{\footnotesize [-84.730, -26.915]}} & \shortstack[c]{-36.850\\{\footnotesize [-59.430, -14.271]}} & \shortstack[c]{-27.788\\{\footnotesize [-49.226, -6.351]}} \\
$l=2$ & \shortstack[c]{-59.417\\{\footnotesize [-83.806, -35.027]}} & \shortstack[c]{-99.918\\{\footnotesize [-144.855, -54.980]}} & \shortstack[c]{-62.599\\{\footnotesize [-87.523, -37.674]}} & \shortstack[c]{-42.175\\{\footnotesize [-64.296, -20.054]}} \\
$l=3$ & \shortstack[c]{-60.958\\{\footnotesize [-86.854, -35.062]}} & \shortstack[c]{-107.305\\{\footnotesize [-157.481, -57.130]}} & \shortstack[c]{-63.986\\{\footnotesize [-88.614, -39.358]}} & \shortstack[c]{-43.054\\{\footnotesize [-70.108, -16.000]}} \\
$l=4$ & \shortstack[c]{-70.058\\{\footnotesize [-92.365, -47.750]}} & \shortstack[c]{-113.410\\{\footnotesize [-162.596, -64.224]}} & \shortstack[c]{-74.868\\{\footnotesize [-99.069, -50.668]}} & \shortstack[c]{-53.656\\{\footnotesize [-78.969, -28.343]}} \\
\bottomrule
\end{tabular}
\vspace{0.35em}
\parbox{0.95\linewidth}{\footnotesize Notes: Each cell reports a point estimate with a pointwise 95 percent confidence interval. Intervals for DSE and DTE use spatial HAC covariance estimates. The standard DID column aggregates cohort-event DID cells on the main admissible support. Bailey and Goodman-Bacon (2015) indicates current-sample event-study coefficients with county-clustered intervals.}
\end{table}
Table~\ref{tab:application_main_decomposition} shows that the total effect is not driven only by own-county adoption.
The DSE captures the effect of own-county adoption at the realized exposure state, whereas the CSE captures spillovers in the no-own-adoption state evaluated over the exposure distribution faced by treated cohorts.
In the adoption year, the DSE is negative but less precisely estimated, while the CSE is already negative.
For each of $l=1,2,3,4$, both DSE and CSE are negative.
The CSE accounts for about 40 percent of the magnitude of the DTE across the reported event times.\footnote{\citet{butts2024difference} reports near-zero spillover estimates for nearby untreated controls in a CHC event-study exercise using a 25-mile spillover indicator. The application here uses a 50-mile exposure mapping and reports CSE and DTE over the exposure distribution faced by treated cohorts.}
The same table reports the fitted source response over the never-treated comparison distribution.
Figure~\ref{fig:application_event_study_decomposition} reports the corresponding event-time paths for DSE, CSE, and DTE.
The decomposition attributes a substantial part of the mortality effect under the realized CHC rollout to spillovers in the no-own-adoption state rather than to own-adoption switching alone.
\begin{table}[!htbp]
\centering
\caption{DSE, CSE, DTE, and source-response estimates for older-adult mortality using Community Health Centers data}
\label{tab:application_main_decomposition}
\small
\setlength{\tabcolsep}{4pt}
\begin{tabular}{lcccc}
\toprule
Event time & DSE & CSE & \shortstack{Never-treated\\source response} & DTE \\
\midrule
$l=0$ & \shortstack[c]{-24.053\\{\footnotesize [-48.944, 0.838]}} & \shortstack[c]{-15.930\\{\footnotesize [-25.510, -6.350]}} & \shortstack[c]{-2.319\\{\footnotesize [-5.279, 0.640]}} & \shortstack[c]{-39.983\\{\footnotesize [-68.604, -11.362]}} \\
$l=1$ & \shortstack[c]{-33.902\\{\footnotesize [-57.921, -9.884]}} & \shortstack[c]{-21.920\\{\footnotesize [-33.482, -10.359]}} & \shortstack[c]{-0.832\\{\footnotesize [-3.722, 2.057]}} & \shortstack[c]{-55.823\\{\footnotesize [-84.730, -26.915]}} \\
$l=2$ & \shortstack[c]{-59.417\\{\footnotesize [-83.806, -35.027]}} & \shortstack[c]{-40.501\\{\footnotesize [-71.946, -9.057]}} & \shortstack[c]{-1.970\\{\footnotesize [-5.712, 1.771]}} & \shortstack[c]{-99.918\\{\footnotesize [-144.855, -54.980]}} \\
$l=3$ & \shortstack[c]{-60.958\\{\footnotesize [-86.854, -35.062]}} & \shortstack[c]{-46.347\\{\footnotesize [-80.280, -12.415]}} & \shortstack[c]{-4.365\\{\footnotesize [-8.748, 0.018]}} & \shortstack[c]{-107.305\\{\footnotesize [-157.481, -57.130]}} \\
$l=4$ & \shortstack[c]{-70.058\\{\footnotesize [-92.365, -47.750]}} & \shortstack[c]{-43.352\\{\footnotesize [-77.006, -9.698]}} & \shortstack[c]{-6.985\\{\footnotesize [-12.242, -1.728]}} & \shortstack[c]{-113.410\\{\footnotesize [-162.596, -64.224]}} \\
\bottomrule
\end{tabular}
\vspace{0.35em}
\parbox{0.95\linewidth}{\footnotesize Notes: Each cell reports a point estimate with a pointwise 95 percent confidence interval based on spatial HAC covariance estimates. DTE is computed as DSE plus CSE on the same admissible support. The never-treated source response is averaged over the never-treated source distribution.}
\end{table}
\begin{figure}[!htbp]
\centering
\includegraphics[width=0.9\textwidth]{application/figures/fig_application_event_study_decomposition.pdf}
\caption{Event-time decomposition of Community Health Center mortality effects.}
\label{fig:application_event_study_decomposition}
\begin{minipage}{0.9\textwidth}
\footnotesize Notes: Intervals are 95 percent pointwise intervals based on spatial HAC covariance estimates, and DTE is computed as DSE plus CSE on the same admissible support.
\end{minipage}
\end{figure}
The control-group CSE diagnostic further suggests that the never-treated comparison group is not a pure no-spillover group.
The source-response column in Table~\ref{tab:application_main_decomposition} evaluates the untreated-state spillover response over the never-treated source distribution.
This column is not part of the cohort-$g$ DTE decomposition; instead, it measures comparison-group exposure under the realized rollout.
The later-window estimate is negative, consistent with the possibility that residents in never-treated counties benefited from nearby CHC access, for example through cross-county use of primary care.
Figure~\ref{fig:application_control_spillover_source} reports the event-time path for the same never-treated source response.
This mechanism is suggestive rather than directly observed, but it is consistent with the design concern that never-treated units may themselves be exposed.
\begin{figure}[!htbp]
\centering
\includegraphics[width=0.9\textwidth]{application/figures/fig_application_control_spillover_source.png}
\caption{Control-group spillover effect.}
\label{fig:application_control_spillover_source}
\begin{minipage}{0.9\textwidth}
\footnotesize Notes: Gray points are cohort-event fitted contrasts in the never-treated source sample when available.
Intervals are 95 percent pointwise intervals based on spatial HAC covariance estimates.
\end{minipage}
\end{figure}
The pre-adoption CSE estimate is interpreted in the same spirit.
It is not a conventional placebo test.
In a staggered rollout with spillovers, later-adopting counties may already be exposed to earlier adopters before their own adoption.
A nonzero pre-adoption CSE therefore indicates possible baseline exposure contamination, which is one of the channels highlighted by the DID decomposition.
\section{Conclusion}\label{sec:conclusion}
This paper studies staggered-adoption difference-in-differences designs when exposure to other units' adoption can affect outcomes.
With a prespecified exposure mapping, untreated status need not coincide with zero exposure.
DSE, CSE, and DTE separate the cohort-event-time objects that a conventional comparison between treated cohorts and never-treated units can mix.
DSE switches own treatment while holding realized exposure fixed, CSE evaluates untreated-state spillovers over the exposure distribution faced by treated cohorts, and DTE compares the pure-control regime with the realized adoption regime.
On the retained admissible support, DTE is the sum of DSE and CSE.
The same decomposition shows that the pure direct effect at zero exposure is not point-identified without additional restrictions linking treated potential outcomes across exposure states.
Identification and estimation follow these three effects.
DSE is recovered from same-state comparisons with never-treated units, while CSE is estimated from the never-treated exposure response and evaluated over the treated cohort target distribution.
DTE is then formed on the same retained cells.
For inference, the estimating equations behind the DSE, CSE, first-stage, diagnostic, and cohort-share blocks are stacked only to form joint influence rows and spatial HAC covariance estimates.
The Monte Carlo designs illustrate the finite-sample behavior of the DSE, CSE, and DTE estimators on retained support, while standard DID benchmarks that assume no spillovers can miss DTE because they do not estimate CSE.
In the empirical study of Community Health Centers, the estimated total effect of the realized rollout is larger in magnitude than the benchmark that ignores exposure.
Under the 50-mile mapping and the maintained first-stage specification, the CSE accounts for about 40 percent of DTE across the reported event times.
Several extensions remain open.
The exposure mapping is maintained rather than estimated, so applied work requires sensitivity analysis over plausible distance cutoffs, network weights, exposure thresholds, and exposure histories.
Pre-adoption diagnostics also require care in staggered designs with spillovers.
A nonzero pre-adoption CSE need not be a conventional placebo failure; it can indicate baseline exposure contamination from earlier adopters.
This points to diagnostics for same-state trends, never-treated exposure responses, transportability, and baseline exposure before adoption.
Finally, the inference conditions treat the realized rollout as fixed and allow spatial or network dependence in the estimating-equation array.
An extension would treat adoption timing itself as network-dependent, as in policy diffusion or strategic adoption, where treatment timing and exposure are jointly shaped by the network.
\bibliographystyle{aer}
\bibliography{ref}