Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
115,339 characters · 25 sections · 45 citation commands
Identification and Estimation of Staggered Difference-in-Differences with Network Spillovers
Keywords: Difference-in-differences; staggered adoption; interference; spillovers; exposure mapping; spatial dependence.
Difference-in-differences designs use untreated units to approximate the counterfactual trend that treated units would have followed without adoption. In the standard design, the stable unit treatment value assumption rules out spillovers: a unit's potential outcome depends on its own treatment path, not on other units' adoption. Recent work on staggered adoption has made the treated-versus-control comparison explicit. callaway2021difference define group-time treatment effects and aggregation schemes. goodman2021difference shows how TWFE coefficients combine two-group comparisons, and sun2021estimating show how event-study coefficients can mix effects across relative periods. borusyak2024revisiting and gardner2025two estimate untreated potential outcomes before constructing dynamic effects. ROTH20232218 provide a review of this recent DID literature. Those results address the comparison and weighting problems created by staggered timing under this no-spillover assumption.
This paper studies staggered DID when adoption occurs on a network. A unit can remain untreated by its own policy status and still be exposed to other adopters. Treated units may also receive spillovers after their own adoption. For later-treated cohorts, the baseline period may already include exposure generated by earlier adopters. The treated-versus-control distinction therefore no longer coincides with the exposed-versus-unexposed distinction.
Once these states differ, a conventional DID contrast that compares treated cohorts with never-treated units combines several terms. It subtracts a comparison-group trend that may already contain untreated-state spillovers. It may also combine the effect of own adoption with spillovers received by treated units and with baseline exposure generated before the cohort's adoption date. Under the stable unit treatment value assumption (SUTVA), these exposure terms are absent and the relevant causal objects reduce to the usual group-time treatment effect. When that assumption fails because of spillovers, the effect of changing own adoption, the spillover effect that would arise without own adoption, and the total effect of the realized rollout are separate objects. This separation follows the distinction between switching, spillover, and total effects in savje2021average and butts2024difference.
Following manski2013identification and 10.1214/16-AOAS1005, we summarize spillover exposure with a prespecified exposure mapping. Under that mapping, we define effects for each adoption cohort and event time under the exposure distribution generated by the realized staggered rollout. The dynamic switching effect (DSE) changes own adoption for a treated cohort at a given event time, averages over that cohort, and holds exposure fixed at its realized state. For a treated cohort, the control-state spillover effect (CSE) is the spillover contrast when own adoption is set to the untreated state, averaged over the exposure distribution faced by that cohort. The dynamic total effect (DTE) compares a state with no adoption to the realized adoption path and equals the sum of DSE and CSE on the same admissible support. This total effect is tied to the rollout-induced exposure distribution; it is not the effect of an arbitrary counterfactual rollout. It is also distinct from the effect of own adoption when every treated unit has zero exposure, which requires potential outcomes for treated units that adopt but are not exposed to other adopters.
Identification follows the missing potential-outcome state required by each effect. For the DSE, identification requires the trend that the treated cohort would have followed without own adoption while keeping its baseline and target exposure states fixed. We recover this trend from never-treated units with the same covariates and the same exposure states at the baseline and target dates, under no anticipation, parallel trends after conditioning on those exposure states, and overlap.
For the CSE, the missing object is the untreated-state spillover response evaluated for the treated cohort. Comparisons among never-treated units with different exposure states identify this response under a parallel trends assumption within never-treated units and exposure states. A separate cross-cohort transportability assumption evaluates the response over the joint distribution of covariates and exposure in the treated cohort. The DTE is obtained by adding the identified DSE and CSE on the same admissible support.
The same never-treated source contrast also identifies spillovers on never-treated units themselves. With additional restrictions linking treated potential outcomes across exposure states, the framework can also identify the effect of own adoption when exposure is set to zero.
The estimators implement the identified effects. The DSE estimator is a saturated long-difference comparison within retained cells defined by covariates and the two exposure states. The CSE estimator fits the untreated-state spillover response in the never-treated source sample and evaluates the fitted contrast over the distribution of covariates and exposure in the treated cohort. The DTE estimator is the post-estimation sum on the same support. For inference, the estimating equations are stacked to obtain first-order approximations for DSE, CSE, and DTE. Spatial heteroskedasticity- and autocorrelation-consistent (spatial HAC) covariance estimates are then applied to the resulting event-time rows. This covariance step follows recent work on inference under spatial or network dependence, including leung2022causal and xu2025difference.
The paper contributes to methods for difference-in-differences with interference in three ways. First, it defines policy effects for each cohort and event time that separate the effect of changing own adoption, spillovers when own adoption is held at the untreated state, and the total effect of the realized rollout. Second, it gives identification results for the main effects. Third, it provides estimators for DSE, CSE, and DTE and an inference method. The procedure conditions on the realized design, uses first-order approximations to the DSE and CSE estimators, and estimates covariance with spatial HAC methods.
The simulations illustrate the finite-sample behavior of the proposed DSE, CSE, and DTE estimators. They also show how standard DID benchmarks that assume no spillovers can miss the total effect of the realized rollout. The empirical study revisits the rollout of Community Health Centers studied by bailey2015war, using county-year mortality data to illustrate that the estimated total effect can differ materially once spillovers among counties without their own Community Health Center are estimated rather than absorbed into the comparison-group trend.
The no-interference staggered-DID literature supplies the benchmark case in which exposure states do not enter potential outcomes. callaway2021difference identify group-time average treatment effects and aggregate them according to the empirical question of interest. goodman2021difference shows that the conventional TWFE DID estimator with variation in treatment timing is a weighted average of two-group, two-period comparisons, which complicates interpretation under heterogeneous and dynamic treatment effects. sun2021estimating show that standard event-study coefficients in staggered designs can be contaminated by treatment effects from other relative periods and propose interaction-weighted estimators. borusyak2024revisiting and gardner2025two estimate untreated potential outcomes using untreated observations and then construct treatment effects from residualized treated outcomes. Under the stable unit treatment value assumption, the dynamic switching effect and dynamic total effect reduce to a group-time treatment effect. With interference, own-adoption switching effects, exposure effects in the untreated state, and total effects of the realized rollout are distinct objects.
A separate literature studies DID when treatment can affect nearby or connected units. delgado2015difference develop DID techniques for spatial data with local autocorrelation and spatial interaction. clarke2017estimating studies DID estimation when spillovers propagate from treated to nearby comparison areas. berg2021spillover discuss spillover-induced biases in empirical corporate finance designs. These papers motivate treating exposure as part of the design rather than as an omitted nuisance. Here, the exposure mapping defines cohort-event-time DSE, CSE, and DTE objects, and spillovers in the comparison group are separated from own-adoption switching.
10.1214/16-AOAS1005 formalize exposure mappings as a way to reduce the full assignment vector to interpretable exposure states and derive design-based estimators under randomized assignments. savje2024causal emphasizes that exposure mappings can serve two different roles: defining interpretable exposure-based effects and imposing assumptions about the true interference structure. The exposure mapping in this paper plays both roles, which is why sensitivity to distance cutoffs, network weights, and exposure thresholds is part of the empirical interpretation. Relatedly, leung2022causal studies approximate neighborhood interference and network HAC inference, where the effect of distant assignments can decay rather than vanish exactly. huber2021framework distinguish individual-level treatment effects from spillover, interaction, and general-equilibrium effects. Their setting allows spillover effects within aggregate units and permits individual treatment effects to interact with aggregate treatment intensity. In the staggered network design studied here, timing varies by cohort and event time, and exposure enters through a unit-level network state. The staggered-DID design here replaces randomized assignment with parallel trends assumptions, support conditions in the never-treated source sample, and transportability across cohorts.
fiorini2025simple consider staggered DID with spillovers and target the ATT without interference. Their identification strategy uses restrictions that remove spillovers from treated units after adoption and require a set of spillover-free controls. The approach here uses the exposure mapping differently. Rather than using it to recover the no-interference ATT, the mapping defines exposure states for each cohort and event time and the corresponding switching, spillover, and total effects. Treated units may remain exposed after adoption, and the total effect of the realized rollout is decomposed into DSE and CSE on a common admissible support.
butts2024difference develops DID methods for spatial spillovers in a two-period design and distinguishes switching effects from total effects, with an event-study extension to staggered adoption. Relative to Butts, the object here is a decomposition of the total effect of the realized rollout for each cohort and event time. The spillover effect is learned from never-treated units and evaluated over the exposure distributions faced by treated cohorts.
xu2025difference analyzes a two-period DID design with interference under exposure mappings and develops doubly robust direct and spillover estimators, finite-population GMM, and spatial HAC inference. This paper uses related finite-population and stacked-moment arguments. The target differs because the stacked system is used for a staggered-adoption decomposition in which the total effect is built from switching and spillover effects on the same support.
sun2025differenceindifferencesnetworkinterference focus on DID under network interference and direct and spillover ATT-type estimands while adjusting for network confounders. The comparison is again mainly about the target effect. Here, the exposure mapping defines switching, spillover, and total effects for the realized rollout, and the CSE uses a never-treated source response transported to treated cohorts.
The remainder of the paper is organized as follows. Section (ref) introduces the staggered-adoption environment, the exposure mapping, and the DSE, CSE, and DTE estimands. Section (ref) gives the identification results, including the decomposition of conventional DID contrasts and the non-identification result for the pure direct effect at zero exposure. Sections (ref) and (ref) present the component estimators and the joint spatial HAC inference procedure. Section (ref) reports the Monte Carlo evidence, and Section (ref) applies the method to data on the rollout of Community Health Centers. Section (ref) concludes.
Consider a panel of $N$ units indexed by $i=1,\dots,N$ over periods $t=1,\dots,T$. Let $D_{it}\in\{0,1\}$ denote own treatment status, and impose absorbing treatment, $D_{it}\le D_{i,t+1}$ for all $i$ and $t<T$. Define the adoption time $G_i \equiv \min\{t:D_{it}=1\}\in\{2,\dots,T\}\cup\{\infty\}$, with $G_i=\infty$ for never-treated units. Throughout, the never-treated state $G_i=\infty$ is the reference adoption status. Comparisons involving untreated potential outcomes are therefore always taken with respect to $G_i=\infty$, rather than units that are untreated at time $t$ but adopt later. Let $\mathbf G=(G_1,\dots,G_N)$ denote the adoption-time vector, and let $Y_{it}(\mathbf G)$ be the potential outcome for unit $i$ at time $t$ under adoption profile $\mathbf G$. This notation permits unrestricted interference: unit $i$'s potential outcome may depend on the entire adoption-time vector.
The primitive objects in the identification analysis are $G_i$, baseline covariates $X_i$, the exposure path $\{H_{it}(\mathbf G_{-i})\}_{t=1}^T$, and the reduced-form potential-outcome collection $\{Y_{it}(g,h): g\in\mathcal G,\ h\in\mathcal H\}_{t=1}^T$, where $\mathcal G\equiv\{2,\dots,T\}\cup\{\infty\}$ and $\mathcal H$ is the support of the reduced-form exposure state. Throughout the theoretical analysis, $\mathcal H$ is taken to be a finite set containing $0$. This discretization ensures that the support conditions in the identification results are well defined.
The dependence of $Y_{it}(\mathbf G)$ on the full adoption-time vector rules out nonparametric identification without further structure. Following 10.1214/16-AOAS1005, we use an exposure mapping that summarizes the relevant features of neighbors' adoption behavior into a scalar exposure state. This is a maintained structural restriction: any channel of interference not encoded in the weights and temporal kernel is excluded by assumption, and the adequacy of the mapping must be defended on substantive grounds.
Let $\widetilde H_{it}(\mathbf G_{-i})$ denote the raw exposure index. Let $w_{ii}=0$ for all $i$, and define $\widetilde H_{it}(\mathbf G_{-i})\equiv\sum_{j\ne i} w_{ij}\phi(t,G_j)$, where $\phi(t,G_j)=\psi(t-G_j)\mathbbm{1}\{t\ge G_j\}$, so that spillovers depend on event time since neighbors' adoption. The theoretical exposure state is a coarsened version of this raw index. Let $b:\mathbb R_+\to\mathcal H$ be a known coarsening map, where $\mathcal H$ is a finite set containing $0$, and define $H_{it}(\mathbf G_{-i})\equiv b(\widetilde H_{it}(\mathbf G_{-i}))$. Throughout the theoretical analysis, $H_{it}$ refers to this finite-valued coarsened exposure state. The raw index $\widetilde H_{it}$ is used only to construct $H_{it}$.
Under Assumption (ref), this is equivalent to $Y_{it}=Y_{it}(G_i,H_{it})$, $H_{it}=H_{it}(\mathbf G_{-i})$. The state $(\infty,0)$ is the pure control state: the unit is untreated and unexposed. The state $(\infty,h)$ with $h>0$ is the untreated-but-exposed state. The state $(g,h)$ with $g<\infty$ and $h>0$ is the treated-and-exposed state. In the treated-and-exposed state, observed outcomes combine own-treatment and spillover components, so separating direct and spillover effects requires restrictions beyond the exposure mapping itself.
Fix an anticipation window $\delta\ge 0$. For cohort $g$, define the baseline period $t_0(g)\equiv g-\delta-1$. For a given cohort-event-time pair $(g,l)$, write $t=g+l$. When the baseline period is clear from context, write $t_0$ in place of $t_0(g)$.
For any calendar-time-indexed random variable $Z_{it}$, define $\mathbb E_g[Z_{i,g+l}]\equiv\mathbb E[Z_{i,g+l}\mid G_i=g]$ and $\mathbb E_{\infty}[Z_{it}]\equiv\mathbb E[Z_{it}\mid G_i=\infty]$.
The primary objects average over the realized distribution of exposure $H_{i,g+l}$ within cohort $g$. For post-adoption event times $l\ge0$, define
The dynamic switching effect (DSE) switches own adoption while holding realized exposure fixed. The control-state spillover effect for cohort $g$ (CSE) is the untreated-state spillover contrast evaluated over the exposure distribution faced by that cohort. The dynamic total effect (DTE) compares the pure-control regime to the realized adoption regime. We allow $\tau^{CSE}(g,l)$ to be defined for $l\ge -1$ so that untreated-state spillovers can also be studied in pre-adoption periods; when $\delta=0$, the cell $l=-1$ coincides with the baseline period. These definitions imply the maintained decomposition
The secondary objects state what is not identified under the maintained assumptions. The pure direct effect is $\tau^{PDE}(g,l)=\mathbb E_g[Y_{i,g+l}(g,0)-Y_{i,g+l}(\infty,0)]$. The treated-side spillover component is $\tau^{AST}(g,l)=\mathbb E_g[Y_{i,g+l}(g,H_{i,g+l})-Y_{i,g+l}(g,0)]$. These objects satisfy $\tau^{DTE}(g,l)=\tau^{PDE}(g,l)+\tau^{AST}(g,l)$. The control-group spillover effect among never-treated units is $\tau_{\infty}^{CSE}(t)\equiv\mathbb E_{\infty}[Y_{it}(\infty,H_{it})-Y_{it}(\infty,0)]$. The full unit-level taxonomy generating these aggregate decompositions is given in Appendix (ref).
The cohort qualifier in the CSE name refers to the target population over which the effect is averaged. The phrase “control-state” refers to the no-own-treatment potential-outcome contrast $Y_{it}(\infty,H_{it})-Y_{it}(\infty,0)$. The never-treated group is used as a source population for learning this response; it is not the target population of $\tau^{CSE}(g,l)$.
At $l=0$, these estimands coincide with cohort averages of the corresponding unit-level effects and summarize the immediate consequences of adoption for cohort $g$. The primary identified targets in what follows are $\tau^{DSE}(g,l)$, $\tau^{CSE}(g,l)$, and therefore $\tau^{DTE}(g,l)$. In empirical applications, fixed analysis weights can be used to define the target population over which these same potential-outcome contrasts are averaged. The weighted empirical targets replace cohort expectations by fixed $\varpi_i$-weighted cohort expectations and reduce to these definitions when $\varpi_i=1$. By contrast, $\tau^{AST}(g,l)$ and the global $\tau^{PDE}(g,l)$ are well-defined population objects but are not point-identified under the maintained assumptions without additional structure. The local pure direct effect is identified only on isolated support. The control-group spillover effect among never-treated units, $\tau_{\infty}^{CSE}(t)$, is a secondary estimand; its identification is discussed in Section (ref).
Under the no-spillover specialization in Remark (ref), the primary estimands collapse to the standard group-time treatment effect for cohort $g$ at calendar time $t=g+l$, up to the baseline and anticipation convention. Conversely, any discrepancy between the spillover-robust objects here and a cohort-specific group-vs-never-treated DID contrast in the presence of interference reflects some combination of post-treatment spillovers and baseline contamination from earlier adopters.
Table (ref) records which potential-outcome states are observed under the realized rollout and which states are counterfactual in the observed data. For the DSE, treated cohorts reveal treated potential outcomes at realized exposure, while the comparison requires the untreated counterfactual trend for units with the same baseline and target exposure states. For the CSE, exposure contrasts in the never-treated source sample recover the untreated spillover response, and transportability evaluates that response over the exposure distribution faced by the treated cohort. Once DSE and CSE are available on the same admissible support, $\tau^{DTE}(g,l)=\tau^{DSE}(g,l)+\tau^{CSE}(g,l)$ follows. The global PDE requires treated potential outcomes at zero exposure and is not point-identified for treated units observed at positive exposure without additional restrictions.
The assumptions below use its observed and missing states selectively.
This no-anticipation condition makes the cohort-$g$ baseline outcome an untreated potential outcome at the realized exposure state.
For fixed $(g,l)$, let $t=g+l$ and $t_0=t_0(g)$. Define the two-date exposure state
This is the DSE parallel trends assumption for units with the same baseline and target exposure states. It compares treated and never-treated units after conditioning on $X_i^d$ and the baseline and target exposure states.
This is a parallel trends assumption within never-treated units and exposure states. It lets zero-exposure never-treated cells remove pure-control trend differences when learning the control-state spillover response.
This is a transportability condition, not a parallel trends assumption. It carries the never-treated source response $r_{\infty,t}(x,h)$ to the cohort-$g$ target distribution.
The assumptions in this section bridge the missing states in Table (ref) in a selective way. Assumptions (ref) and (ref) identify the DSE by supplying the untreated long difference for cohort $g$ among units with the same baseline and target exposure states. Assumptions (ref) and (ref) identify the control-state spillover response in the never-treated source sample and carry it to the cohort-$g$ target distribution. What remains outside the maintained identification region is the global PDE, $\tau^{PDE}(g,l)$, and, more generally, the unrestricted treated-side zero-exposure state $Y_{it}(G_i,0)$ away from isolated support. A local PDE can be recovered on isolated zero-exposure support under an additional zero-exposure parallel trends assumption and a positive-mass condition; Appendix (ref) states the additional assumptions and Appendix (ref) gives the formal result. Those isolated-support conditions are not used to identify the main DSE, CSE, or DTE objects.
To accommodate the general anticipation-window setup, consider the cohort-specific group-vs-never-treated DID comparison between the adoption period $t=g$ and the no-anticipation baseline $t_0=t_0(g)=g-\delta-1$. Define
When $\delta=0$, this reduces to the adjacent-period DID comparison with $t_0(g)=g-1$. Let
Here $B_g^{base}$ is baseline exposure contamination for cohort $g$, and $B_g^{PT}$ is the pure-control trend difference between cohort $g$ and never-treated units.
Thus, a DID comparison between treated cohorts and never-treated units combines own adoption, treated-side spillovers, comparison-group spillovers, baseline exposure contamination, and pure-control trend bias.
The decomposition above reveals two distinct distortions under interference. In the post-treatment period, the terms $\tau^{AST}(g,0)$ and $\tau_{\infty}^{CSE}(g)$ reflect contamination from spillover exposure after adoption. At the anticipation-safe baseline $t_0(g)$, the term $-B_g^{base}$ captures baseline contamination: by the time cohort $g$ reaches its reference period, earlier cohorts may already have shifted the baseline outcome away from the pure control state. When $\delta=0$, $t_0(g)=g-1$ and $B_g^{base}$ coincides with $\tau^{CSE}(g,-1)$.
This distinction is specific to staggered adoption. For the first treated cohort, or in a simultaneous-adoption design, the baseline period precedes any treatment in the population and the pre-treatment spillover terms vanish. Later cohorts face a different environment because their baseline may already be contaminated by earlier adopters. This cohort-specific DID contrast therefore mixes direct effects, post-treatment spillovers, and baseline contamination whenever interference is present.
For fixed $(g,l)$, let $t=g+l$, $t_0=t_0(g)$, and
Define
and let $F^D_{g,l}\ll F^D_{\infty,g,l}$ denote DSE overlap for $(X_i^d,S_i^{g,l})$. For CSE, define
and
Let $F_{g,l}\equiv\mathcal L(X_i^d,H_{it}\mid G_i=g)$. The following propositions give identification formulas for DSE, CSE, and DTE.
The first proposition identifies the DSE by replacing the missing untreated long difference for cohort $g$ with the same-state long difference among never-treated units.
The next proposition identifies the CSE by recovering the control-state spillover response in the never-treated source sample and integrating it over the cohort-$g$ exposure distribution.
The following proposition identifies the DTE by adding the DSE and CSE formulas on the same admissible support.
The identified dynamic total effect is indexed by the realized exposure distribution generated by the observed staggered-adoption process. Thus,
should be interpreted as the effect for cohort $g$ at event time $l$ of moving from the pure-control regime to the observed adoption regime, holding fixed the exposure environment generated by that regime. It is not the effect of an arbitrary counterfactual rollout rule. If all calendar-valid cells are support-valid, an event-time total effect can be written as
With fixed analysis weights, the event-time aggregate replaces $\omega_{g,l}$ by the corresponding weighted cohort mass share. If support failure removes some cohort-event-time cells, the aggregate is over the admissible cells and should be interpreted as an admissible-cell aggregate, denoted $\tau^{DTE}_{\mathcal A}(l)$ when that restricted population must be explicit.
The maintained identification results above do not, by themselves, separate the pure direct effect from the spillover effect on treated units. The decomposition
and the preceding propositions do not identify $\tau^{AST}(g,l)$. We therefore treat $\tau^{PDE}(g,l)$ as a secondary object that requires additional restrictions.
One possible additional structure is
This restriction is a substantive no-interaction restriction between own treatment status and spillover exposure. It is not implied by the exposure mapping or by the DID restrictions above. Under this no-interaction restriction,
and
Therefore $\tau^{AST}(g,l)=\tau^{CSE}(g,l)$, and
Without that restriction, only a local PDE on isolated zero-exposure support is available under additional conditions. Appendix (ref) states the additional assumptions, and Appendix (ref) gives the formal result.
DSE is estimated by saturated long differences within retained two-date exposure-state cells. CSE is estimated by fitting a control-state spillover response in the never-treated source sample and evaluating the fitted contrast over the treated cohort target distribution. DTE is formed as a post-estimation map using the same admissible cells and weights as the two primitive estimators. The estimators in this section are implemented directly by the closed-form formulas below. Section (ref) stacks the corresponding estimating equations to derive joint influence rows for DSE, CSE, and DTE. The stacked estimating-equation representation is used for joint linearization and covariance estimation; it does not change the closed-form point estimators.
For each cohort-event-time pair $(g,l)$, let $t=g+l$ and $t_0=t_0(g)$. Define $\Delta_i^{g,l}\equiv Y_{it}-Y_{i,t_0}$, $D_i^g\equiv\mathbbm{1}\{G_i=g\}$, and $C_i\equiv\mathbbm{1}\{G_i=\infty\}$. Let $\varpi_i>0$ denote a fixed analysis weight, such as a baseline older-adult population measure. The unweighted case is obtained by setting $\varpi_i=1$. For any random variable $A_i$, define the cohort-$g$ weighted average operator
The weighted primitive causal targets are
For the parametric CSE estimator, let $V_i$ denote the first-stage covariate basis and let $c_t(v,h;\eta)$ denote a control-state spillover response contrast, normalized by $c_t(v,0;\eta)=0$. The parametric transported CSE target is
for $t=g+l$ and the corresponding total target is
When $\varpi_i$ is not constant, the DSE component is interpreted under the corresponding weighted parallel trends assumption within the retained $Z_i^{g,l}$ cells. All identification results extend to the fixed-weight targets by replacing cohort and source expectations with the corresponding $\varpi_i$-weighted expectations; the application imposes the same identifying restrictions for these weighted target populations. The unweighted case is the special case $\varpi_i=1$. Let $S_i^{g,l}\equiv(H_{it},H_{i,t_0})$ denote the two-date exposure state. If baseline covariates are used in the identifying parallel trends assumption, let $X_i^d$ denote a finite-dimensional discretization or stratification of $X_i$, fixed before estimation, and define $Z_i^{g,l}\equiv(X_i^d,H_{it},H_{i,t_0})$. If no covariate stratification is used, $Z_i^{g,l}=S_i^{g,l}$. All saturated same-state comparisons below are saturated in $Z_i^{g,l}$, not merely in $S_i^{g,l}$, whenever the maintained identification assumption conditions on $X_i^d$. Let $\mathcal Z_{g,l}\equiv\operatorname{supp}(Z_i^{g,l}\mid G_i=g)$ denote the target support for the DSE comparison.
The DSE support condition requires every target state to also be observed among never-treated units:
For $a\in\{g,\infty\}$ and $z\in\mathcal Z_{g,l}$, define
and the corresponding weighted masses
The count variables determine empirical support and minimum-cell admissibility. The point estimates use the analysis weights:
When $\varpi_i=1$, this reduces to the unweighted cell mean
Let
be the empirical common support. The saturated DSE estimator is
where
In the empirical implementation, estimates are reported only when the stricter minimum-count rule in Section (ref) retains all target mass. On reported cells, $\widehat{\mathcal Z}_{g,l}$ can therefore be replaced by the retained support $\widehat{\mathcal Z}_{g,l}^{D}$. Equivalently, $\widehat\tau_{\varpi}^{DSE}(g,l)$ is obtained by weighted least squares on the retained comparison sample $\{i:G_i\in\{g,\infty\},\ Z_i^{g,l}\in\widehat{\mathcal Z}_{g,l}\}$,
and setting
Because the regression is saturated in the retained state $Z_i^{g,l}$, this regression estimator is numerically identical to the cell-mean estimator above.
The CSE first stage uses period $1$ as the zero-exposure baseline, because no unit is treated at $t=1$ and $H_{i1}=0$ under the maintained design. For $t\ge 2$, define the never-treated conditional mean
and the conditional untreated-state spillover response
The causal CSE estimand identified in Proposition (ref) is
for $t=g+l$ under the parallel trends assumption within the never-treated source sample and exposure states, together with the cross-cohort transportability restriction. The identification result is nonparametric in the finite exposure-state strata: on the source support, $r_{\infty,t}(x,h)$ is identified by $\mu_t(x,h)-\mu_t(x,0)$. The empirical estimator implements this identified contrast through a structured parametric model $c_t(V_i,H_{it};\eta)$. Hence the plug-in CSE equals the causal CSE when the fitted contrast is correctly specified on the transported support; otherwise it is the corresponding transported parametric target. The main CSE estimator uses the parametric transported plug-in obtained from the maintained first-stage specification in the never-treated source sample. For the empirical implementation, the maintained first stage is a structured binary-positive exposure-response model estimated in the never-treated source sample. The basis $V_i$ is distinct from the finite support strata $X_i^d$ used in saturated support checks. For $t\ge2$, define
The baseline period has $R_{i1}=0$ and $P_{i1}=0$; in the application code it anchors the baseline trend fit and contributes no exposure-response contrast. The first stage writes the never-treated conditional mean as
In the binary-positive specification,
The term $a_t(V_i;\alpha)$ is the baseline trend component used for residualization, and $B_t(V_i)$ is the exposure-response basis. Pooling over $t$ means that the finite-dimensional parameter is estimated using the stacked never-treated source observations; it imposes a time-invariant spillover response only if $B_t(V_i)$ is specified without calendar-time variation. The maintained application includes period indicators in $B_t(V_i)$, so the binary-positive exposure contrast varies by calendar time after residualization. The transported plug-in estimator evaluates the fitted contrast $c_t(V_i,H_{it};\widehat\beta)$ over the cohort-$g$ target distribution. For the main CSE estimator, let $t=g+l$. The estimator is
Its model-implied population target is
If the transported contrast is correctly specified on the target support,
for all $(x,v,h)$ in $\operatorname{supp}(X_i^d,V_i,H_{it}\mid G_i=g)$, then
The same first-stage contrast also defines a diagnostic estimand for the never-treated comparison group. This object is not a primary policy target in the DTE decomposition. It measures spillover exposure in the comparison group and appears directly in the decomposition of conventional comparisons between treated cohorts and never-treated units. For calendar time $t\ge2$, define the weighted never-treated control-group spillover target
The unweighted case is obtained by setting $\varpi_i=1$. By the source comparison underlying Proposition (ref), this target is identified on the never-treated source support as
provided
Unlike $\tau_{\varpi}^{CSE}(g,l)$, this object does not require the cross-cohort transportability restriction in Assumption (ref). It is a within-source-population spillover estimand. Under the parametric transported plug-in specification, define the corresponding model-implied target
The plug-in estimator is
The denominator is
If $c_t(v,h;\eta^0)=r_{\infty,t}(x,h)$ on the never-treated source support for $(X_i^d,V_i,H_{it})$, then
For a cohort-specific DID comparison between $t_0(g)$ and $g$, when both calendar-time spillover effects are defined, define the never-treated spillover change
Its estimator is
This quantity measures the spillover component embedded in the never-treated comparison group's long difference. In the group-versus-never DID decomposition, it enters with the opposite sign because the never-treated long difference is subtracted.
For each admissible $(g,l)$, define
For event-time aggregation, let $\widehat{\mathcal G}_l^\star$ be the set of cohorts for which both DSE and CSE are admissible at event time $l$. For $g\in\widehat{\mathcal G}_l^\star$, define
For $Q\in\{DSE,CSE,DTE\}$, define
Because the same admissible set and the same weights are used for all three components,
holds exactly.
Let $m_N$ be a pre-specified minimum cell-count threshold. Either $m_N=m$ for a fixed integer $m$, or $m_N\to\infty$ with $m_N/N\to0$. For DSE, define
Let $\widehat{\Pr}_{\varpi}(\cdot\mid G_i=g)$ denote the empirical distribution weighted by $\varpi_i$ within cohort $g$. The DSE component for $(g,l)$ is reported only if
For CSE, define
The CSE component for $(g,l)$, with $t=g+l$, is reported only if, for every $(x,h)$ in the empirical support of $(X_i^d,H_{it})$ among units with $G_i=g$,
For the never-treated control-group spillover effect, define the retained source support
The estimate $\widehat\tau_{\varpi,\infty}^{CSE}(t)$ is reported only if
This rule ensures that the reported never-treated spillover target is evaluated on the same source support used to learn the control-state spillover response. The admissible cohort set is
All event-time aggregates are computed on $\widehat{\mathcal G}_l^\star$.
For inference, the component estimators from Section (ref) are embedded in a finite-dimensional stacked estimating-equation system. The just-identified DSE and CSE target blocks reproduce the closed-form component formulas; DTE is their post-estimation sum on retained cells, and event-time estimates aggregate retained cohort-event cells.
The asymptotic sequence is finite-population and conditional on the design variables. The design sigma-field $\mathcal C_N$ contains fixed attributes, analysis weights, network weights, and locations; when the realized rollout is treated as fixed, it also contains the realized cohort and exposure paths. Let $\mathcal A=\{(g,l):g\in\mathcal G_l^\star,\ l\in\mathcal L\}$ denote the retained cohort-event set. The primitive parameter vector collects the retained DSE source-cell means, the CSE first-stage parameter, the DSE and CSE cell-level targets, the never-treated CSE diagnostic when reported, and the weighted cohort masses. DTE is not a primitive component; it is the smooth map
Define the stacked moment row
where the never-treated CSE diagnostic block is included only when that diagnostic is reported. Each block collects the corresponding equations over retained cells, event times, and cohorts; Appendix (ref) gives the full block definitions. Let $\bar q_N(\theta)=N^{-1}\sum_{i=1}^N q_{i,N}(\theta)$ and write $m_N(\theta)=\mathbb E[\bar q_N(\theta)\mid\mathcal C_N]$. The population parameter $\theta_N^0$ satisfies the stacked moment condition $m_N(\theta_N^0)=0$. The component estimators are represented by the stacked GMM first-order condition associated with
where $\widehat\Psi_N$ is the weighting matrix for the stacked moment criterion. The stack contains equations with the following roles. One set estimates the never-treated trends used in the DSE comparison. A second set estimates the CSE first stage in the never-treated source sample. Two target equations map these quantities into the DSE and CSE components. When the never-treated diagnostic is reported, an additional target equation evaluates the same first-stage response over the never-treated distribution. The final equations estimate the cohort shares used in event-time aggregation. In the just-identified target and cohort-share blocks, these equations reproduce the closed-form estimators. If the CSE first stage is overidentified, the associated first-order score enters the stack under the maintained first-stage specification. This construction parallels the stacked GMM representation in xu2025difference. Xu stacks nuisance and target moments for a two-period DID problem with interference. Here, the stacked system is used for retained cohort-event cells in a staggered design and includes equations for the switching component, the spillover component, the first-stage response, the diagnostic spillover object, and cohort shares. Event-time paths are smooth maps of the primitive components. Because DTE uses DSE and CSE on the same retained cells, its influence row is formed by adding the DSE and CSE influence rows before spatial HAC covariance estimation. This keeps the covariance between the two component rows in the DTE standard error.
Spatial dependence uses the $\psi$-dependence route of KOJEVNIKOV2021882, and the increasing-domain spatial interpretation follows JENISH200986 and JENISH2012178. Spatial HAC covariance estimation is applied to the event-time influence rows.
Let $\mathcal C_N=\sigma\{\{X_i\}_{i=1}^N,\{\varpi_i\}_{i=1}^N,\{w_{ij}\}_{i,j=1}^N,\{s_i\}_{i=1}^N\}$ denote the sigma-field generated by fixed attributes, analysis weights, network weights, and locations. All probability limits and expectations below are conditional on $\mathcal C_N$, unless otherwise stated. Let $\mathcal L$ be the finite set of reported event times and let $\mathcal G_l^\star$ be the population admissible cohort set at event time $l$. For a retained $(g,l)$, define
For a finite-valued random variable $A_i$ and event $B_i$, define the weighted support
For $a\in\{g,\infty\}$, define
For the CSE source support, define
Define weighted cohort masses
Event-time weights are
When the realized rollout is treated as fixed, the conditional probabilities and expectations above are interpreted as finite-population shares on the realized design.
For $Q\in\{DSE,CSE,DTE\}$, the event-time target is $\tau_{\varpi}^Q(l)=\sum_{g\in\mathcal G_l^\star}\omega_{\varpi,g,l}\tau_{\varpi}^Q(g,l)$. When CSE is estimated by a parametric transported plug-in, $Q=CSE$ and $Q=DTE$ refer to the corresponding parametric transported targets unless the fitted contrast is correctly specified on the transported support. Let $\mathcal T_\infty$ denote the finite set of calendar times at which the never-treated control-group spillover effect is reported. The never-treated control-group CSE diagnostic, when reported, is treated as an additional smooth map of the same CSE first stage and the never-treated target distribution.
The stacked vector $\widehat\theta_N$ collects the component estimates defined in Section (ref): the DSE source-cell means, the CSE first-stage parameter, the DSE and CSE target estimates, the never-treated CSE diagnostic components when reported, and the cohort-mass estimates used in event-time aggregation. Let $\theta_N^0$ be the corresponding population value. Appendix (ref) defines the stacked estimating-equation array $q_{i,N}(\theta)$ for these components. Write $\bar q_N(\theta)=N^{-1}\sum_{i=1}^N q_{i,N}(\theta)$.
Assumption (ref) is the population-margin version of the empirical admissibility rules in Section (ref), including stability of the reported cohort-event cells. Assumption (ref) is a local regularity condition for the stacked GMM representation. It requires that the estimator defined by the stacked criterion admit the first-order expansion generated by the sample moments. In the just-identified target and cohort-mass blocks, this representation coincides with the closed-form component estimators from Section (ref).
Assumptions (ref) and (ref) are imposed directly on the stacked score and event-time influence arrays. The conditional CLT used below is obtained by applying the $\psi$-dependence limit theory used in xu2025difference; it is not imposed as a separate assumption. Assumption (ref) specifies the kernel, bandwidth, and dependence conditions used by the spatial HAC covariance estimator. The bandwidth grows so that increasingly distant units enter the covariance estimate, but slowly enough relative to the expanding domain. Together with Assumption (ref), these conditions allow the spatial HAC covariance result of xu2025difference to be applied to the finite-dimensional event-time influence-row array.
Let $a_l^Q(\theta)$ denote the event-time aggregation map for $Q\in\{DSE,CSE,DTE\}$. The influence row for the event-time path is
Define the feasible influence row
where
The feasible DTE row is formed as the sum of the DSE and CSE rows before covariance estimation:
For the never-treated control-group spillover effect, define
For the spillover change in the never-treated comparison group, when both $g$ and $t_0(g)$ are included in $\mathcal T_\infty$, the feasible influence row is
The spatial HAC estimator is written in the uncentered form used in the finite-population covariance decomposition. For $Q\in\{DSE,CSE,DTE\}$, define the spatial HAC covariance estimator for the event-time path by
Here $K(\cdot)$ is the spatial kernel, $b_N$ is the bandwidth, and $\rho(i,j)$ is the geographic or network distance from Assumptions (ref) and (ref). For the never-treated control-group spillover effect, define
The pointwise standard error is
The corresponding pointwise interval is
The spatial HAC variance for $\widehat\Delta_{\infty}^{CSE}(g)$ is computed by replacing $\widehat\varphi_{i,N}^{CSE,\infty}(t)$ in (ref) with $\widehat\varphi_{i,N}^{\Delta CSE,\infty}(g)$. Let
and for the never-treated control-group spillover effect let
and define
Define the finite-population centering component
Let
For the never-treated control-group spillover effect, define $\mu_{i,N}^{CSE,\infty}(t)$, $\Gamma_{\infty}^{CSE,E}(t,t')$, and $\Gamma_{\infty}^{CSE,+}(t,t')$ analogously by replacing $\varphi_{i,N}^Q(l)$ with $\varphi_{i,N}^{CSE,\infty}(t)$ and $\Gamma_Q(l,l')$ with $\Gamma_{\infty}^{CSE}(t,t')$.
For $Q\in\{DSE,CSE,DTE\}$, the pointwise interval is
For $Q=CSE$ and $Q=DTE$, the target is the parametric transported target unless the CSE contrast is correctly specified on the transported target support.
The simulations use finite populations in which DSE, CSE, and DTE are known functions of the simulated potential outcomes. They ask whether the proposed estimators recover the targets on retained support and whether standard group-versus-never DID and callaway2021difference ATT estimators implemented without exposure can be close to the DSE target while missing DTE. DTE is formed on the same admissible support as the sum of DSE and CSE,
The simulations use the unweighted special case $\varpi_i=1$.
All main designs use a balanced six-period panel. Cohorts are $G_i\in\{3,4,5,\infty\}$, treatment is absorbing with $D_{it}=\mathbbm{1}\{t\geq G_i\}$, there is no anticipation, and the baseline period is $t_0(g)=g-1$. We report event times $l\in\{0,1,2\}$. For a given event time $l$, only cohorts satisfying $g+l\leq T$ enter the corresponding event-time aggregate. Units are located on an open line network with row-normalized weights $w_{ij}$. The maintained exposure state is a coarsened line-network exposure share, $H_{it}=b(\widetilde H_{it})\in\{0,\text{low},\text{high}\}$. The map keeps zero exposure at $0$, maps values in $(0,0.5]$ to low exposure, and maps values in $(0.5,1]$ to high exposure. The dose score is $q(0)=0$, $q(\text{low})=1$, and $q(\text{high})=2$.
For all main designs, $X_i$, $\alpha_i$, and $\varepsilon_{it}$ are drawn independently from $N(0,1)$. The time effect is $\lambda_t=0.2(t-1)$. The control-state spillover slope is $\rho_t=0$ for $t\in\{1,2\}$ and $\rho_t=0.3$ for $t\geq3$. The own-treatment profile is $\tau_0=1$, $\tau_1=1.5$, and $\tau_l=2$ for $l\geq2$. The DGPs depend on $q(H_{it})$, not directly on the raw exposure index. Untreated potential outcomes include the control-state spillover slope $\rho_t$, while treated potential outcomes add the own-treatment profile and, when $\kappa_d\neq0$, a treated-side exposure interaction. Appendix (ref) gives the exact potential-outcome equations and finite-population truth formulas. DGP1 sets $\kappa_d=0$, so own treatment and spillover exposure are additively separable. In this design, $\tau^{PDE}(g,l)=\tau^{DSE}(g,l)$. DGP2 sets $\kappa_d=0.4$. This is the main interaction design: the switching effect, DTE, and the global PDE generally differ. DGP3 uses the same potential-outcome structure as DGP2 but modifies the cohort assignment and support environment so that some exposure states faced by treated cohorts are less frequently represented among never-treated units. DGP3 instead uses a weakly balanced support-aware line assignment with block size 12 and equal cohort shares across $G_i\in\{3,4,5,\infty\}$. We interpret DGP3 as a support/admissibility stress design.
For each replication, the truth objects used for evaluation are finite-population averages of the simulated potential outcomes on the retained admissible support. DTE is formed as the same-support sum of DSE and CSE, equivalently the direct contrast from the realized adoption regime to the pure-control regime.
The reported estimators are the same-state DSE estimator, the parametric transported plug-in estimator for CSE, and the DTE estimator obtained from their sum. For DSE, define $S_i^{g,l}=(H_{it},H_{it_0})$ and $Z_i^{g,l}=(X_i^d,S_i^{g,l})$, with $Z_i^{g,l}=S_i^{g,l}$ when no covariate stratification is used. The DSE estimator compares treated and never-treated units within the retained support of $Z_i^{g,l}$. For CSE, the never-treated sample is used to estimate an effect contrast $c_t(x,h;\eta)\approx\mu_t(x,h)-\mu_t(x,0)$, normalized by $c_t(x,0;\eta)=0$. Let $t=g+l$. The estimator then evaluates the fitted contrast over the cohort-$g$ exposure distribution:
The DTE estimator is the post-estimation sum:
At the event-time level, DSE, CSE, and DTE use the same admissible cohort set and the same weights, so $\widehat\tau^{DTE}(l)=\widehat\tau^{DSE}(l)+\widehat\tau^{CSE}(l)$ holds exactly in each replication.
The benchmark estimators are a cohort-specific DID contrast and the callaway2021difference group-time ATT estimator implemented without exposure. For each event time $l$, let $\widehat{\mathcal G}_l^\star$ be the set of cohorts for which both DSE and CSE support requirements are satisfied. The support rule uses a pre-specified minimum cell count $m_N=5$. A cohort-event cell is admissible only when all required treated and never-treated source cells for DSE and CSE meet this threshold. For $g\in\widehat{\mathcal G}_l^\star$, event-time aggregates use
For $Q\in\{DSE,CSE,DTE\}$,
For event-time comparisons, benchmark cohort-time estimates are aggregated over the same $\widehat{\mathcal G}_l^\star$ and with the same weights $\widehat\omega_{g,l}$ whenever the corresponding cohort-time benchmark is available. These benchmarks are reported as deviations from the relevant target on retained support: $\tau^{DSE}_{rep}(l)$ in the DSE table and $\tau^{DTE}_{rep}(l)$ in the DTE table. Bias, RMSE, and coverage are computed relative to
rather than the full-cohort target that includes inadmissible cells. The tables report bias or benchmark deviation, RMSE, retained treated mass or availability when relevant, and pointwise coverage.
For DGP1--DGP3, pointwise intervals for DSE, CSE, and DTE use the sandwich covariance from the stacked estimating equations. The spatial HAC kernel is applied over line-network distance with bandwidth $b_N=\lceil N^{1/3}\rceil$, which gives $b_N=6$ for $N=200$ and $b_N=8$ for $N=500$. The covariance is computed after forming the event-time influence rows for DSE, CSE, and DTE; for DTE, this keeps the covariance between the DSE and CSE rows.
The main tables report bias or benchmark deviation, RMSE, and pointwise coverage. For benchmark rows, the Bias column is the deviation from the target on retained support. Coverage is computed conditional on interval availability. The final main-text runs use $R=1000$ Monte Carlo replications for each design-size combination.
Tables (ref) and (ref) separate the benchmark comparison that ignores exposure from the performance of the proposed DSE, CSE, and DTE estimators on retained support. First, DID benchmarks that ignore exposure generally differ from DTE when exposure affects outcomes. Table (ref) reports standard group-versus-never DID and callaway2021difference benchmarks implemented while ignoring exposure. These estimators answer no-interference DID questions; interpreting them as DTE estimators treats exposure as irrelevant both for the estimand and for comparison-group trends. In the pooled DTE comparison, the benchmark deviations range from about $-0.31$ to $-0.36$ across the reported designs, and coverage for $N=500$ falls to $0.066$--$0.165$.
This comparison is not a claim that the proposed estimators dominate the benchmarks for the switching target; DTE includes CSE, which the benchmark estimators do not estimate. Second, Table (ref) reports the proposed DSE, CSE, and DTE estimators. Same-state comparisons estimate DSE, the never-treated first-stage contrast estimates the control-state spillover response used for CSE, and DTE is formed by adding the two components on the same admissible support. Table (ref) shows small biases and pointwise coverage close to the nominal 95 percent level across the reported designs.
DGP3 is included to show the role of support restrictions. When overlap is weak, the estimator reports common-support targets and availability becomes part of the interpretation.
The estimates in this section use the 50-mile exposure mapping and the identifying restrictions in Sections (ref)--(ref). For CSE, the first stage is fit in the never-treated source sample and then transported to treated cohorts. Under these maintained conditions, the CSE and DTE estimates have a causal interpretation; otherwise, they are transported parametric targets.
The empirical study revisits a staggered county-level policy setting in which nearby counties can be exposed before or after their own adoption date. A conventional DID comparison between treated and never-treated counties in this setting subtracts trends from never-treated counties that may themselves be exposed to nearby Community Health Centers. It also assigns any spillover received by treated counties after adoption to the same coefficient as own adoption. The empirical exercise therefore reports the components in
rather than treating a single DID coefficient as an own-treatment effect.
We use the county-year mortality panel from the Community Health Centers replication files for bailey2015war. The maintained analysis sample covers 1959--1978, target cohorts adopt between 1965 and 1974, and exposure is a positive-exposure indicator for CHC adoption within 50 miles. Older-adult population weights define the application target distribution. The primary outcome is the mortality rate for individuals aged 50 and older, measured in deaths per 100,000 older-adult residents. The panel also retains all-age mortality, age-specific mortality rates, cause-specific mortality rates, population weights, baseline demographic controls, physician supply, and the event-time variables used in the original replication files. County identifiers are standardized to five-digit FIPS codes, and the analysis excludes Los Angeles County, Cook County, and New York County.
Treatment timing is defined by the first year in which county $i$ receives a Community Health Center, denoted $G_i$. The target treated cohorts are $g\in\{1965,\dots,1974\}$, and never-treated counties are assigned $G_i=\infty$. Counties first treated after 1974 are retained when constructing exposure to nearby adopters but are not used as target treated cohorts or never-treated comparison counties in the component estimators. Treatment is absorbing, so the realized treatment indicator is
Event time is indexed by $l=t-g$, so calendar time and event time satisfy $t=g+l$. In the application, we set the anticipation window to zero, so the baseline period is
The resulting empirical long difference is therefore
Geographic exposure is constructed from the 2015 Census TIGER/Line county shapefile. County representative points are obtained from polygon geometry, and distances are measured between representative points in miles. The 50-mile exposure mapping is plausible in this setting because Community Health Centers may serve residents outside the adopting county. Nearby counties may also be affected through travel to a center, provider networks, referral patterns, or reduced congestion in alternative sources of care. The binary 50-mile indicator is a maintained exposure mapping for the DID design, not a claim that the true structural spillover function has a sharp cutoff at 50 miles. Let $d_{ij}$ denote the distance between counties $i$ and $j$, and define the 50-mile neighbor set
The maintained exposure count is
The empirical exposure state is binary:
For each cohort-event-time pair $(g,l)$, the empirical two-date exposure state is
Let $X_i^d$ denote the finite baseline strata used in the DSE comparison, fixed before estimation, and define
In the reported CHC estimates, no additional finite baseline covariate strata are imposed beyond the two-date exposure state, so $X_i^d$ is degenerate and $Z_i^{g,l}=S_i^{g,l}$. The first-stage basis $V_i$ below is separate from this support object and is used to parameterize the CSE source response. This construction allows outcomes to depend on a county's own adoption timing and on its realized exposure state before and after the relevant event time. The current exposure state enters the potential outcome, while the two-date state $S_i^{g,l}$ conditions the same-state long-difference comparison.
The empirical estimators use the cohort-event-time objects defined in Sections (ref)--(ref). For each admissible pair $(g,l)$, the comparison sample contains counties in cohort $g$ and never-treated counties. The DSE component is estimated with the saturated long-difference specification from Section (ref), using the retained state $Z_i^{g,l}$ and the older-adult population weight. Thus the empirical DSE regression for each cell is
The estimated DSE is the weighted average of the fitted state-specific switching contrasts over the weighted empirical distribution of $Z_i^{g,l}$ in cohort $g$. The reported estimands are older-adult-population-weighted, corresponding to $\varpi_i$ equal to the baseline older-adult population weight. The reported DSE is therefore not a rich-controls re-estimation of the bailey2015war event-study specification. It is a comparison defined by exposure-state cells and older-adult population weights.
The CSE component is estimated with the structured residualized binary-positive first stage. In the application, for $t\ge2$, the source outcome and exposure indicator are
The structured first stage writes
so that
In the first-stage fit, $V_i$ consists of the mean pre-treatment outcome over 1960--1964, log 1960 older-adult population, physicians per capita, the baseline urban bucket, and Census region. The baseline trend component $a_t(V_i;\alpha)$ includes these variables and a four-degree-of-freedom spline in calendar time. The exposure-response basis $B_t(V_i)$ contains period indicators and positive-exposure interactions with the same baseline variables, so the binary-positive exposure contrast varies by calendar time after residualization. The older-adult population weight is $\varpi_i$ in the weighted source criterion and target averages. Together with transportability, the reported CSE and DTE estimates are causal only if this transported first-stage contrast is correctly specified on the treated-cohort support; otherwise they are transported parametric targets. The empirical DTE is the post-estimation sum of DSE and CSE using the same admissible cells and weights.
Inference uses the sandwich covariance from the stacked estimating equations, with spatial HAC weighting. Confidence intervals are pointwise normal intervals based on the estimated spatial HAC covariance matrix. The same stacked covariance is used for DSE, CSE, and DTE, so the covariance between the DSE and CSE components is retained in inference for DTE. The reported event-time summaries aggregate over the common admissible cohort set. The main reporting windows are $l=0$ and $l=1,\dots,4$ for DSE, CSE, and DTE, together with a pre-adoption CSE window at $l=-1$.
All estimates in this section use the 50-mile exposure mapping. The CSE learns a spillover response from never-treated counties and evaluates that response over the exposure distributions faced by treated cohorts. The CSE and DTE estimates therefore have a causal interpretation under the transportability condition and the maintained first-stage specification; otherwise, they should be read as transported parametric targets. Figure (ref) shows that exposure to nearby CHCs is not rare in the analysis sample. This pattern is consistent with the institutional setting: CHCs expanded access to low-cost primary care, reduced travel and access barriers, and could affect residents outside the county that formally received a center. The estimates should therefore be interpreted as effects under the exposure distribution generated by the realized CHC rollout, not as effects of an arbitrary counterfactual rollout rule.
The mortality results are consistent with the conclusion in bailey2015war that CHCs reduced older-adult mortality. The spillover-aware decomposition adds that the DTE is larger in magnitude than the benchmark that ignores exposure. In the adoption year, Table (ref) reports a DTE of $-39.983$ deaths per 100,000, compared with $-24.023$ from the standard DID benchmark. At event time $l=4$, the DTE is $-113.410$, compared with $-74.868$ for the benchmark. Thus, the benchmark that ignores exposure is smaller in magnitude than the estimated DTE at each reported event time.
Table (ref) shows that the total effect is not driven only by own-county adoption. The DSE captures the effect of own-county adoption at the realized exposure state, whereas the CSE captures spillovers in the no-own-adoption state evaluated over the exposure distribution faced by treated cohorts. In the adoption year, the DSE is negative but less precisely estimated, while the CSE is already negative. For each of $l=1,2,3,4$, both DSE and CSE are negative. The CSE accounts for about 40 percent of the magnitude of the DTE across the reported event times.\footnote{butts2024difference reports near-zero spillover estimates for nearby untreated controls in a CHC event-study exercise using a 25-mile spillover indicator. The application here uses a 50-mile exposure mapping and reports CSE and DTE over the exposure distribution faced by treated cohorts.} The same table reports the fitted source response over the never-treated comparison distribution. Figure (ref) reports the corresponding event-time paths for DSE, CSE, and DTE. The decomposition attributes a substantial part of the mortality effect under the realized CHC rollout to spillovers in the no-own-adoption state rather than to own-adoption switching alone.
The control-group CSE diagnostic further suggests that the never-treated comparison group is not a pure no-spillover group. The source-response column in Table (ref) evaluates the untreated-state spillover response over the never-treated source distribution. This column is not part of the cohort-$g$ DTE decomposition; instead, it measures comparison-group exposure under the realized rollout. The later-window estimate is negative, consistent with the possibility that residents in never-treated counties benefited from nearby CHC access, for example through cross-county use of primary care. Figure (ref) reports the event-time path for the same never-treated source response. This mechanism is suggestive rather than directly observed, but it is consistent with the design concern that never-treated units may themselves be exposed.
The pre-adoption CSE estimate is interpreted in the same spirit. It is not a conventional placebo test. In a staggered rollout with spillovers, later-adopting counties may already be exposed to earlier adopters before their own adoption. A nonzero pre-adoption CSE therefore indicates possible baseline exposure contamination, which is one of the channels highlighted by the DID decomposition.
This paper studies staggered-adoption difference-in-differences designs when exposure to other units' adoption can affect outcomes. With a prespecified exposure mapping, untreated status need not coincide with zero exposure. DSE, CSE, and DTE separate the cohort-event-time objects that a conventional comparison between treated cohorts and never-treated units can mix. DSE switches own treatment while holding realized exposure fixed, CSE evaluates untreated-state spillovers over the exposure distribution faced by treated cohorts, and DTE compares the pure-control regime with the realized adoption regime. On the retained admissible support, DTE is the sum of DSE and CSE. The same decomposition shows that the pure direct effect at zero exposure is not point-identified without additional restrictions linking treated potential outcomes across exposure states.
Identification and estimation follow these three effects. DSE is recovered from same-state comparisons with never-treated units, while CSE is estimated from the never-treated exposure response and evaluated over the treated cohort target distribution. DTE is then formed on the same retained cells. For inference, the estimating equations behind the DSE, CSE, first-stage, diagnostic, and cohort-share blocks are stacked only to form joint influence rows and spatial HAC covariance estimates. The Monte Carlo designs illustrate the finite-sample behavior of the DSE, CSE, and DTE estimators on retained support, while standard DID benchmarks that assume no spillovers can miss DTE because they do not estimate CSE. In the empirical study of Community Health Centers, the estimated total effect of the realized rollout is larger in magnitude than the benchmark that ignores exposure. Under the 50-mile mapping and the maintained first-stage specification, the CSE accounts for about 40 percent of DTE across the reported event times.
Several extensions remain open. The exposure mapping is maintained rather than estimated, so applied work requires sensitivity analysis over plausible distance cutoffs, network weights, exposure thresholds, and exposure histories. Pre-adoption diagnostics also require care in staggered designs with spillovers. A nonzero pre-adoption CSE need not be a conventional placebo failure; it can indicate baseline exposure contamination from earlier adopters. This points to diagnostics for same-state trends, never-treated exposure responses, transportability, and baseline exposure before adoption. Finally, the inference conditions treat the realized rollout as fixed and allow spatial or network dependence in the estimating-equation array. An extension would treat adoption timing itself as network-dependent, as in policy diffusion or strategic adoption, where treatment timing and exposure are jointly shaped by the network.