Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
139,858 characters · 29 sections · 106 citation commands
Causal Graphs for Conditional Parallel Trends
Difference-in-Differences (DiD) is one of the most widely used research designs for causal inference in economics and related fields goldsmith-pinkham_tracking_2024. It is also studied in a rapidly growing methodological literature that examines its identifying assumptions, extensions, and limitations roth_whats_2023,dechaisemartin_credible_2023,arkhangelsky_causal_2024. The key identifying assumption in most DiD applications is that outcomes for treated and untreated units would have evolved in parallel in a counterfactual world with no treatment. In empirical practice, this assumption is often imposed conditional on observed covariates, corresponding to conditional parallel trends (CPT). For example, 19 of the papers published in the American Economic Review in 2024 and 2025 explicitly assume some form of parallel trends, 14 of which impose CPT, and 13 of which use time-varying controls. However, none of these papers provide a formal justification for the choice of conditioning variables beyond presenting evidence on parallel pre-trends.\footnote{Appendix (ref) provides details about the literature review.} This indicates that guidance on how to justify appropriate conditioning variables remains limited, especially when conditioning on time-varying covariates.
In settings with unconfoundedness or instrumental variables, causal graphs provide well-established frameworks for reasoning about valid conditioning variables and causal identification (see, e.g., textbooks by pearl_causal_2016, frolich_impact_2019, cunningham_causal_2021, huber_causal_2023, chernozhukov_applied_2024, hernan_causal_2024, wager_causal_2024, or reviews by heckman_econometric_2022, hunermund_causal_2025, cinelli_crash_2024, imbens_causal_2024, abbring_philip_2025). However, these tools do not directly apply to DiD settings, where identification is typically justified by functional form assumptions on time-invariant unobserved confounders that enter outcomes in an additively separable way. In particular, the recent literature typically states these assumptions only for the untreated potential outcome, while leaving treated potential outcomes, and thus the individual treatment effect, unrestricted (see Appendix (ref) for a review). We build on Single World Intervention Graphs (SWIGs) developed by richardson_single_2013 to represent this perspective graphically. In settings with unconfoundedness, SWIGs encode conditional independencies involving potential outcomes and treatment that justify standard identification of average treatment effects, which can be read off via $d$-separation, a standard tool in causal graphs richardson_single_2013a.
This paper introduces transformed SWIGs, the $\Delta$-SWIGs, and shows that they provide a graphical characterization of CPT via $d$-separation. In particular, $\Delta$-SWIGs encode conditional independencies involving differences in potential outcomes that correspond to conditional parallel trends and justify identification in DiD settings. This provides a general framework for reasoning about conditioning strategies in DiD settings. We use this framework to analyze how features of the causal structure—such as outcome dynamics, time-varying covariates, and treatment–covariate feedback—affect the existence and form of valid conditioning strategies in the standard 2x2 setting and the multi-period setting of callaway_differenceindifferences_2021. In particular, we characterize how CPT can be justified based on additive separability and the causal structure alone, show that outcome dynamics generally preclude it without further restrictions, and establish that when time-varying covariates affect the outcome, post-treatment variables must be controlled for. We also show that pre-treatment parallel trends are informative only about a subset of the assumptions required for unbiased post-treatment effects. Finally, we translate these insights into practical guidance for implementing conditional DiD with time-varying covariates.
To illustrate the main issues and findings, Figure (ref) presents results from a simple simulation with unobserved time-invariant confounders and time-varying observed covariates. The latter may be unaffected by the treatment (Figure (ref)) or affected by the treatment (Figure (ref)) (see Appendix (ref) for details). The figure reports conditional DiD estimates from a large sample for the group first treated in period four under three conditioning strategies: using only pre-treatment covariates (left), using covariates measured prior to each outcome in the respective DiD comparison (“pre-outcome” controls, middle), and using the full sequence of time-varying covariates (right). We highlight several patterns that may appear surprising but can all be rationalized within our framework: (i) using only pre-treatment covariates yields parallel pre-trends and unbiased short-term effects but biased dynamic effects, (ii) using pre-outcome controls provides the same results as using the full sequence of covariates, (iii) in the absence of treatment–covariate feedback, strategies involving post-treatment variables are unbiased, (iv) in the presence of treatment–covariate feedback, all dynamic effect estimates are biased, reflecting omitted variable bias or “wrong world control bias”, and (v) pre-trends do not diagnose post-treatment violations of CPT, while short-term effects remain unbiased even with treatment–covariate feedback. ghanem_when_2026 raise similar issues in a three-period setting.
Overall, this suggests that identification should be derived from explicit structural considerations complemented by statistical evidence on parallel pre-trends, rather than treating CPT as a primitive assumption justified solely by such evidence. Our framework provides a transparent and principled way to conduct this reasoning and is particularly useful for analyses with time-varying covariates and multiple periods, especially when researchers are reluctant to rely on parametric or distributional assumptions beyond standard additive separability of unobserved confounders.
The remainder of the paper is organized as follows. Section (ref) discusses the related literature and situates our contribution. Section (ref) introduces the necessary building blocks, including structural causal models, SWIGs, and $d$-separation. Section (ref) develops the $\Delta$-SWIG framework and its connection to conditional parallel trends. Section (ref) and Section (ref) apply this framework to characterize valid conditioning strategies in 2x2 and multi-period settings. Section (ref) discusses practical implications, and Section (ref) concludes.
This paper is related and contributes to several strands of literature. First, it contributes to the recent literature on justifying (conditional) parallel trends for DiD or DiD directly based on substantive knowledge about the data-generating process. Most prominently, ghanem_selection_2024 provide selection-based necessary and sufficient conditions for parallel trends largely focusing on the two-period case without covariates. They discuss extensions to more time periods without covariates, or to two periods with time-varying covariates that are not affected by the treatment. We focus on sufficient conditions but also cover multiple time periods with time-varying covariates that may be affected by the treatment. ghanem_when_2026 study the informativeness of pre-trends in work conducted independently and concurrently with ours. In line with our findings, they show that with time-varying covariates, pre-trends based on pre-treatment controls can be misleading for causal effects in a three-period setting. Again, our results also hold for more time periods and provide more details about what can and cannot be deduced from pre-trends trends testing. Therefore, our approach and results are complementary to those in ghanem_selection_2024 and ghanem_when_2026. marx_parallel_2024 focus on specific models of dynamic choice and the parallel trends assumption. chabe-ferret_should_2025 considers a linear model with selection on pre-treatment outcomes. Both do not address the role of covariates.
The second strand of literature is on causal graphs for DiD. renson_using_2026 is most closely related using SWIGs to provide necessary conditions for unconditional parallel trends in two time periods. In contrast, we introduce $\Delta$-SWIGs to obtain sufficient conditions for conditional parallel trends with time-varying covariates and multiple time periods. weber_assumption_2015 compare identifying assumptions, including conditional parallel trends, relying on DAGs rather than SWIGs. huber_joint_2024 uses DAGs to motivate a joint test of unconfoundedness and parallel trends. Both studies are restricted to two time periods and cannot read off standard CPT involving differences in untreated potential outcomes because they are not represented in the DAGs. kim_gain_2021 consider “gain scores” - where the outcome of interest is the difference between pre- and post-treatment outcomes - in linear models and DAGs using path tracing rules. zhang_exploiting_2021 show how to incorporate equality constraints in linear models, with DiD as an important special case. Unlike these last two contributions, we do not restrict attention to linear models in two time periods and incorporate covariates.
Third, our work relates to the literature on DiD with time-varying covariates. caetano_differenceindifferences_2024 establish identification under CPT conditional on the full sequence of time-varying covariates; our approach can be used to justify this assumption. We also show that controlling for fewer variables suffices under the same structural assumptions, thereby mitigating issues related to high dimensionality. caetano_difference_2024 consider CPT with post-treatment covariates in two periods, and provide conditions under which identification of treatment effects remains possible. Again, our approach can provide sufficient conditions for these assumptions grounded in the causal structure, and further extends the analysis to multiple time periods and the informativeness of pre-trend testing.
Taken together, our work advances the literature along several dimensions. We introduce and show the validity of $\Delta$-SWIGs as a graphical tool to derive sufficient conditions for conditional parallel trends, accommodating time-varying covariates, treatment-covariate feedback, covariate dynamics, and multiple time periods — settings not jointly covered by existing approaches. Beyond identification, we provide additional testable implications and clarify when pre-trend tests are and are not informative about the validity of the identifying assumptions.
Throughout the paper uppercase letters $V$ denote random variables, lowercase letters $v$ the corresponding realizations. Bold letters denote sets of random variables $\mathbf{V}$ with realizations $\mathbf{v}$.
The canonical 2x2 setting considers binary treatment $D$ and two periods $t=0,1$. It assumes potential outcomes $Y_t(d)$ in period $t$ under treatment $d$ and targets the average treatment effect on the treated $ATT := \mathbb{E}[Y_1(1) - Y_1(0) \mid D=1]$. It exploits that no unit is treated in period $t=0$ and some units are treated in $t=1$. Then, assuming parallel trends conditional on time-invariant control variables $X$
identifies the $ATT$ following standard arguments heckman_matching_1997, abadie_semiparametric_2005,lechner_estimation_2011. \footnote{Identification also requires overlap. However, we abstract from overlap considerations in this paper and implicitly assume in our discussions that overlap holds. We also omit the no anticipation assumption at this stage. See Remark (ref) for a detailed discussion how it follows naturally once the graphical perspective is introduced.} We stick to this well-established setting until we have introduced all relevant components to consider more complex scenarios.
A structural causal model (SCM) $\mathcal{M}$ consists of a set of endogenous variables $\mathbf V$, exogenous variables $\mathbf U$, and functions $\mathcal{F} := \{f_{V_j}\}_{V_j \in \mathbf V}$. For each $V_j \in \mathbf V$, the model specifies a structural equation $V_j \vcentcolon = f_{V_j}(Pa(V_j), U_{V_j})$, where $Pa(V_j) \subseteq \mathbf{V} \setminus \{V_j\} $ denotes the parents (direct causes) of $V_j$, and $U_{V_j} \in \mathbf U$ are exogenous random variables. The corresponding graph $\mathcal{G}$ contains a node for each $V_j \in \mathbf V$ and a directed edge from each member of $Pa(V_j)$ to $V_j$, with an arrowhead pointing to $V_j$. If there exists a directed path $V_a \rightarrow ... \rightarrow V_d$, i.e. with all arrows pointing towards $V_d$, then $V_d$ is call a descendant of $V_a$. We denote all descendants of a node $Desc(V_j)$. If there are no cycles, i.e. no directed sequences of edges from a node to itself, the graph is called a directed acyclic graph (DAG).\footnote{For a more in-depth introduction to SCMs and DAGs see, e.g., the textbooks by textbooks by pearl_causal_2016, frolich_impact_2019, cunningham_causal_2021, huber_causal_2023, chernozhukov_applied_2024, hernan_causal_2024, wager_causal_2024, or reviews by heckman_econometric_2022, hunermund_causal_2025, cinelli_crash_2024, imbens_causal_2024, abbring_philip_2025).}
Figure (ref) provides an example illustrating the conventions for causal graphs used throughout the paper: (i) exogenous variables like $U_{Y_0}$ are omitted unless their presence is useful; (ii) a missing circle around a variable indicates that it is unobservable. Figure (ref) depicts a setting with time-invariant observable confounders $X$ and unobservable time-invariant confounders $U$ where outcomes are measured in the pre- and post-treatment period. This is a setting where standard 2x2 DiD would often be considered to overcome unobserved confounding/heterogeneity indicated by the open backdoor path $D \leftarrow U \rightarrow Y_1$, which prevents identification of causal effects by controlling for $X$ pearl_causal_2016.
We note that common DAGs are useful to illustrate the problem in this setting, but not expressive enough to incorporate the DiD solution. In particular, they do not allow to read-off sufficient conditions for conditional parallel trends. This is in contrast to unconfoundedness settings where DAGs paired with the backdoor criterion are powerful to reason about good, bad and neutral controls cinelli_crash_2024. We build on Single World Intervention Graphs paired with $d$-separation to provide graph-based justifications for control variables in DiD. Both components are introduced in the following.
Single World Intervention Graphs (SWIGs) developed by richardson_single_2013 unify DAGs and the potential outcomes framework. In the underlying SCM, potential outcomes are obtained by fixing possibly multiple treatment variables $\mathbf D$ to a particular value $\mathbf d$ in the structural equations.\footnote{This fix operator is discussed, e.g. in heckman_causal_2015. See Table 3.1 in heckman_econometric_2022 for an illustration of the difference between the fix and the do operator, which is commonly applied in the context of DAGs.} This results in a new set of endogenous variables $\mathbf V (\mathbf d)$ consisting of observable and potential variables. For example, fixing $D = 0$ in Equation (ref) of Figure (ref) yields $Y_1(0) := f_Y(U,X,0,U_{Y_1})$ representing the potential outcome in a counterfactual world where every unit remains untreated. All other variables remain unchanged as their structural equations do not depend on $D$ or its descendants.
The SWIG $\mathcal{G}(\mathbf{d})$ graphically mirrors the hypothetical intervention of fixing $\mathbf D = \mathbf d$ by transforming the original DAG in two steps. First, each of the treatment nodes $D_j \in \mathbf D$ are split into a random part $D_j$ and a fixed part $d_j$. The random part of the split node inherits the incoming edges and the fixed part the outgoing edges. Second, all descendants of the variables in $\mathbf D$ are relabeled to become potential outcomes.\footnote{hernan_causal_2024 and chernozhukov_applied_2024 provide textbook introductions to SWIGs. andersen_guide_2023 use SWIGs to discuss sample selection.} We use a subscript to indicate if a specific expression refers to a SWIG, e.g. $Pa_{\mathcal{G}(\mathbf{d})}(V_j (\mathbf d))$ denotes the parents of node $V_j (\mathbf d)$ in SWIG $\mathcal{G}(\mathbf{d})$.
Figure (ref) shows the SWIG for the two period case with unobserved confounding where the treatment node is split and its descendant node is relabeled as potential outcome $Y_1(0)$. The relabeling of the pre-treatment outcome in gray is optional but useful as explained in the following remark:
The benefit of SWIGs is that they allow to read off (conditional) independencies involving observable and potential variables via $d$-separation, which is a standard tool in the causal graph literature verma_causal_1988. Two nodes $X,Y \in \mathbf{V}(\mathbf{d})$ are said to be $d$-separated by a set of nodes $\mathbf Z \subseteq \mathbf{V}(\mathbf{d}) \setminus \{X,Y\}$ in SWIG $\mathcal{G}(\mathbf{d})$ if every path between $X$ and $Y$ is blocked by the set $\mathbf Z$. Practically, a path can be blocked i) by including the middle node of chains ($V_i \rightarrow V_j \rightarrow V_k$) or forks ($V_i \leftarrow V_j \rightarrow V_k$) in $\mathbf Z$, and/or ii) by neither including the middle node of a collider structure ($V_i \rightarrow V_j \leftarrow V_k$) nor its descendants in $\mathbf{Z}$. We write $X \perp\!\!\!\perp_{\mathcal{G}(\mathbf{d})} Y \mid \mathbf Z$ to express that two nodes $X$ and $Y$ are $d$-separated by the set of nodes $\mathbf Z$ in SWIG $\mathcal{G}(\mathbf{d})$. Appendix (ref) provides a formal definition of $d$-separation and illustrates path blocking rules in an example that highlights rules frequently applied in the complex structures below.
Proposition 11 of richardson_single_2013 shows that the distribution of $\mathbf{V}(\mathbf{d})$ satisfies the Markov property with respect to SWIG $\mathcal{G}(\mathbf d)$ if the data is generated by the underlying SCM. Most importantly for our purposes this implies that the graphical $d$-separation criterion can be applied to find (conditional) independencies among the variables displayed in a SWIG richardson_single_2013. Concretely, let $X$, $Y$, and $\mathbf Z$ be nodes contained in SWIG $\mathcal{G}(\mathbf{d})$, then $X \perp\!\!\!\perp_{\mathcal{G}(\mathbf{d})} Y \mid \mathbf Z, \mathbf{d} \Rightarrow X \perp\!\!\!\perp Y \mid \mathbf Z$. In words, if the union of nodes $\mathbf Z$ and the fixed nodes $\mathbf{d}$ $d$-separate $X$ and $Y$, the corresponding random variables are conditionally independent given variables $\mathbf Z$ alone.
For example applying $d$-separation to Figure (ref) uncovers the conditional independence between the untreated potential outcome and the treatment $Y_1(0) \perp\!\!\!\perp D \mid X, U$. This implies mean independence $\mathbb{E}[Y_1(0) \mid X,U, D=1] = \mathbb{E}[Y_1(0) \mid X,U, D = 0]$ that could be used in the standard unconfoundedness identification $ATT = \mathbb{E}[Y_1 \mid D=1] - \mathbb{E}[\mathbb{E}[Y_1 \mid X,U, D = 0] \mid D=1]$ if $U$ would be observable rosenbaum_central_1983.\footnote{See richardson_single_2013a for a more detailed introduction to SWIGs in the unconfoundedness setting.} However, $U$ is unobservable and DiD assumes conditional parallel trends (ref) instead. A conditional independence like $\Delta Y_1(0) \perp\!\!\!\perp D \mid X$ would imply (ref). However, this conditional independence does not follow from SWIG (ref) without further restrictions and the introduction of $\Delta$-SWIGs.
DiD acknowledges the presence of time-invariant unobservable confounders but assumes that they enter additively separable in the untreated potential outcome $Y_t(0)$, as reviewed in Appendix (ref). Translated to the SCM in Figure (ref) this means that we restrict the functional form of the structural outcome equations:\footnote{For example, also blundell_alternative_2009 and bonhomme_back_2025 take this structural equation perspective instead of directly restricting untreated potential outcomes, which is most common in the recent DiD literature.}
The SCM perspective highlights a peculiarity of the DiD setting. While $U$ enters $Y_t(0)$ only via an additively separable and time-invariant function of time-invariant variables $\alpha(U,X)$, the individual treatment effect that is added in case of treatment $\tau(U,X,U_{Y_1})$ depends on $U$ in an arbitrary manner.\footnote{Note that in the unconditional case everything collapses to the standard two-way fixed-effects structure with $g_{Y_t}(U_{Y_t}) = \underbrace{\mathbb{E}[g_{Y_t}(U_{Y_t})]}_{\lambda_t} + \underbrace{g_{Y_t}(U_{Y_t}) - \mathbb{E}[g_{Y_t}(U_{Y_t})]}_{\varepsilon_t}$ such that $Y_t(0) = \alpha(U) + \lambda_t + \varepsilon_t$ (see also Appendix (ref)).}
We graphically represent the assumption that $U$ enters $Y_t(0)$ only via the additively separable function by annotating the respective edges with $+\alpha$ in the DAG and SWIG of Figure (ref) that is otherwise identical to Figure (ref).\footnote{Note that we do not label the edges emitted from $X$ in the same manner, as $X$ enters $Y_t(0)$ via the time-varying functions $g_{Y_t}$ in addition to time-invariant function $\alpha$.} This highlights that the DAG contains a single labeled edge $U \rightarrow Y_0$, indicating that $U$ enters $Y_0$ only via $\alpha(U,X)$, while the relation $U \rightarrow Y_1$ is nonparametric. It illustrates that $U$ is not additively separable in all observable outcomes. However, $U$ enters all untreated potential outcomes only via $\alpha(U,X)$ as depicted in Figure (ref). The observation that this structure is implied by standard DiD modeling assumptions motivates the development of $\Delta$-SWIGs in the next section to directly read off (ref) from a causal graph.
\FloatBarrier
In this section, we outline how to construct a $\Delta$-SWIG that allows to read off $\Delta Y_1(0) \perp\!\!\!\perp D \mid X$ and consequently (ref). We build on the concrete 2x2 SWIG in Figure (ref) to convey the main idea in a compact manner before Section (ref) provides the general procedure and proves its validity.
Note that Assumption \nameref{ass:swas-2x2} implies that the first difference of untreated potential outcomes does not depend on $U$.\footnote{This follows because $\Delta Y_1(0) = f_{Y_1}(U,X,0,U_{Y_1}) - f_{Y_0}(U,X,U_{Y_0}) = \cancel{\alpha(U,X)} + g_{Y_1}(X,U_{Y_1}) - \cancel{\alpha(U,X)} - g_{Y_0}(X,U_{Y_0}) =: f_{\Delta Y_1(0)} (\cancel{U},X,U_{Y_1},U_{Y_0})$.} This motivates adding a "difference node" to the SWIG that inherits all incoming edges from its two levels except the edge from $U$. Figure (ref) illustrates the "$\Delta$-SWIG" that results from augmenting SWIG (ref) accordingly.
Most importantly, Theorem (ref) below proves that $d$-separation also implies conditional independencies in such $\Delta$-SWIGs. Further, note that all paths between $D$ and $\Delta Y_1(0)$ in Figure (ref) are either blocked by colliders $Y_t(0)$ (e.g. $D\leftarrow U \rightarrow Y_t(0) \leftarrow U_{Y_t} \rightarrow \Delta Y_1(0)$) and/or by conditioning on $X$ (e.g. $D \leftarrow U \rightarrow X \rightarrow \Delta Y_1(0)$). Therefore, the $\Delta$-SWIG implies $\Delta Y_1(0) \perp\!\!\!\perp D \mid X$, which implies (ref) and therefore the standard identification result $ATT = \mathbb{E}[\Delta Y_1 \mid D=1] - \mathbb{E}[\mathbb{E}[\Delta Y_1 \mid X, D = 0] \mid D=1]$ holds.
As $Y_0(0)$ and $Y_1(0)$ have no descendants in Figure (ref), we can "prune" the full $\Delta$-SWIG by dropping $Y_0(0)$, $Y_1(0)$ and their incoming edges. Additionally, $U_{Y_0}$ and $U_{Y_1}$ do not need to be explicitly depicted as they only affect one variable in the pruned graph. The resulting pruned $\Delta$-SWIG in Figure (ref) establishes $\Delta Y_1(0) \perp\!\!\!\perp D \mid X$ in a more compact manner. The benefit of pruning is that paths that are already blocked by colliders $Y_t(0)$ disappear and we can focus on blocking the two confounding paths $D \leftarrow X \rightarrow \Delta Y_1(0)$ and $D \leftarrow U \rightarrow X \rightarrow \Delta Y_1(0)$. Theorem (ref) below shows that the described pruning keeps the validity of $d$-separation intact. Thus, unless the full $\Delta$-SWIG is instructive, we display appropriately pruned $\Delta$-SWIGs by default.
Procedure (ref) generalizes the transformation steps of the previous section to construct a full or a pruned $\Delta$-SWIG from any SCM:
The results provided in the following sections exploit that $\mathcal{G}_\Delta^{full}(\mathbf{d})$ or $\mathcal{G}^{prune}_\Delta(\mathbf{d})$ obtained via Procedure (ref) allow to read off conditional independencies involving difference nodes via $d$-separation, as Theorem (ref) shows under standard assumptions for causal graphs:
The assumption that $P(V(\mathbf{d}))$ satisfies the Markov property with respect to $(\mathcal{G}(\mathbf{d}))_{\mathbf{V}(\mathbf{d})}$ is the standard adaptation of the Markov property to SWIGs and is satisfied if the data is generated by the underlying SCM. Remark (ref) in Appendix (ref) discusses how the standard Markov property is connected to factorization according to a SWIG in richardson_single_2013, ensuring that $d$-separation implies conditional independence.
The proof of Theorem (ref) is provided in Appendix (ref) and exploits that a $\Delta$-SWIG is constructed by adding nodes $\Delta_{j,k}(\mathbf{d})$ that have no descendants. Thus, the variables represented by $\Delta_{j,k}(\mathbf{d})$ are deterministic functions of the variables represented by their parents in $\mathcal{G}_\Delta^{full}(\mathbf{d})$, which ensures that the probability distribution of the variables $\mathbf{V}_\Delta(\mathbf{d})$ satisfies the Markov property with respect to $\mathcal{G}_\Delta^{full}(\mathbf{d})$. The pruning step uses that removing nodes with no descendants leaves $d$-separation among the other nodes in a graph unchanged, as it can only remove paths containing unconditioned colliders or unconditioned descendants of nodes on other paths.
While we focus on difference nodes, the same arguments can be applied for other combinations of nodes, e.g. ratios of nodes to remove multiplicative separable unobservables. We leave such extensions for future research.
The example in Section (ref) illustrates how to deduce conditional parallel trends from a $\Delta$-SWIG in the 2x2 setting with time-invariant observable control variables, denoted $X$. Now, we introduce time-varying control variables $X_t$ in the 2x2 framework. This allows to demonstrate the application of the $\Delta$-SWIG technology in a relatively compact setting, while discussing implications of outcome dynamics and endogenous control variables for CPT that also occur in multi-period settings.
The DAG in Figure (ref) depicts a setting with time-varying but pre-treatment covariates, $X_0$ and $X_1$.\footnote{Time-invariant covariates can be subsumed in $X_0$ because an explicit node $X$ would emit the same arrows as node $X_0$.} The SCM spelled out below the figure incorporates an adapted single world additive separability assumption. Therefore, the $\Delta$-SWIG in Figure (ref) resulting from Procedure (ref) contains no edge $U \rightarrow \Delta Y_1(0)$.
In this $\Delta$-SWIG, all paths between $D$ and $\Delta Y_1(0)$ are blocked given $X_0,X_1$.\footnote{For $t = 0,1$, the paths $D \leftarrow X_t \rightarrow \Delta Y_1 (0)$ and $D \leftarrow U \rightarrow X_t \rightarrow \Delta Y_1(0)$ are all blocked by conditioning on $X_t$. Additionally, the paths $D \leftarrow U (\rightarrow X_0) \rightarrow Y_0 \leftarrow U_{Y_1} \rightarrow \Delta Y_1(0)$ are blocked by collider $Y_0$.} Therefore, they are $d$-separated and we can deduce $\Delta Y_1(0) \perp\!\!\!\perp D \mid X_0,X_1$ following Theorem (ref), which implies conditional parallel trends $\mathbb{E}[\Delta Y_1(0) \mid X_0, X_1, D = 0] = \mathbb{E}[\Delta Y_1(0) \mid X_0, X_1, D = 1]$ that enable standard identification of the $ATT$.
$\Delta$-SWIG (ref) explicitly includes node $Y_0$, although it does not have descendants and therefore could be pruned from the graph. However, it is instructive to revisit graphically why pre-treatment outcomes are considered bad controls although they are measured before the treatment daw_matching_2018,chabe-ferret_should_2025. Conditioning on collider $Y_0$ would open the path $D \leftarrow U \rightarrow Y_0 \leftarrow U_{Y_0} \rightarrow \Delta Y_1(0)$ and precludes reading off parallel trends conditional on $Y_0$. cinelli_crash_2024 refer to this type of bad control in the unconfoundedness setting as inducing M-bias.
Next, we graphically revisit why outcome dynamics are usually not compatible with CPT. Figure (ref) shows the DAG with the three possible instances of outcome dynamics depicted as dotted edges: state dependence ($Y_0 \rightarrow Y_1$), outcome-treatment feedback ($Y_0 \rightarrow D$), and outcome-covariate feedback ($Y_0 \rightarrow X_1$). We discuss them in separate $\Delta$-SWIGs for clarity, but the same conclusions follow from one $\Delta$-SWIG with all dynamics.\footnote{Note that $Y_0$ could not be pruned from the following $\Delta$-SWIGs because it now has a descendant. In contrast, level node $Y_1(0)$ can be pruned and is pruned.}
State-dependence: Following Procedure (ref), the difference node in $\Delta$-SWIG (ref) now inherits an arrow $Y_0 \rightarrow \Delta Y_1(0)$. This creates open path $D \leftarrow U \rightarrow Y_0 \rightarrow \Delta Y_1(0)$ that could be blocked by conditioning on $Y_0$. However, doing so opens path $D \leftarrow U \rightarrow Y_0 \leftarrow U_{Y_0} \rightarrow \Delta Y_1(0)$ and creates M-Bias, as discussed in the previous section. Therefore, $Y_0$ is a “dilemma node” that can block only one of two paths that can not be blocked otherwise. Consequently, $D$ and $\Delta Y_1(0)$ can not be $d$-separated and CPT can not be established in this structure renson_using_2026.
Outcome-treatment feedback: $\Delta$-SWIG (ref) displays selection on pre-treatment outcomes. The $\Delta$-SWIG now contains the open path $D \leftarrow Y_0 \leftarrow U_{Y_0} \rightarrow \Delta Y_1(0)$. $Y_0$ is again a dilemma node. Conditioning on it would block this additional path while opening path $D \leftarrow U \rightarrow Y_0 \leftarrow U_{Y_0} \rightarrow \Delta Y_1(0)$. Again, $d$-separation can not be established in this setting, which is in line with previous results showing that selection on pre-treatment outcomes is in conflict with parallel trends ashenfelter_using_1985,bonhomme_back_2025,marx_parallel_2024, renson_using_2026 unless a particular martingale condition holds, as discussed in ghanem_selection_2024.
Outcome-covariate feedback: $X_1$ is the dilemma node in $\Delta$-SWIG (ref). Conditioning on $X_1$ is required to block the path $D \leftarrow X_1 \rightarrow \Delta Y_1(0)$. However, this opens path $D \leftarrow U \rightarrow Y_0 \leftarrow U_{Y_0} \rightarrow \Delta Y_1(0)$ because $X_1$ is a descendant of collider $Y_0$. Again, there is no set of observable variables that $d$-separates $D$ and $\Delta Y_1(0)$ weber_assumption_2015.
Summary: The presence of path $D \leftarrow U \rightarrow Y_0 \leftarrow U_{Y_0} \rightarrow \Delta Y_1(0)$ prevents $d$-separation if collider node $Y_0$ emits an arrow because it (i) either creates an additional chain where $Y_0$ is the only observable variable that could block it, making itself a dilemma node, or (ii) it becomes a parent of an otherwise crucial control variable, making its descendant $X_1$ a dilemma node. This prevents the deduction of CPT in 2x2 settings with outcome dynamics but also in multi-period settings without imposing additional restrictions. Therefore, we abstract from outcome dynamics in the following.
Next, we discuss the role of post-treatment covariates in the 2x2 setting as considered in caetano_difference_2024. A corresponding DAG and $\Delta$-SWIG is depicted in Figure (ref). In contrast to Figure (ref), outcome dynamics are absent and $X_1$ is now affected by the treatment ($D \rightarrow X_1$).\footnote{Time-invariant covariates denoted $Z$ in caetano_difference_2024 are subsumed into $X_0$. Drawing the $Z$ node explicitly would not change the discussion.} Therefore, the $\Delta$-SWIG contains untreated potential covariate $X_1(0)$. $\Delta$-SWIG (ref) implies independence
and CPT $\mathbb{E}[\Delta Y_1(0) \mid X_0,X_1(0), D = 0] = \mathbb{E}[\Delta Y_1(0) \mid X_0,X_1(0), D = 1]$, which corresponds to Assumption 2 in caetano_difference_2024. However, $X_1(0)$ is unobservable for the treated units and identification based on this CPT is not feasible. Thus, caetano_difference_2024 discuss additional covariate unconfoundedness assumptions of the form
in their Corollary 1. Graphically, both (ref) and (ref) are implied by $\Delta$-SWIG (ref) after removing the edge $U \rightarrow X_1(0)$, i.e. assuming that $X_1$ is only indirectly affected by $U$ via $X_0$.\footnote{This represents the minimal modification to achieve (ref) and (ref). Additionally, also $U \rightarrow X_0$ could be removed without changing the result. Alternatively, $U \rightarrow D$ could be removed, while keeping the $U \rightarrow X_t$ edges. This would create a standard unconfoundedness scenario where $Y_1(0) \perp\!\!\!\perp D \mid X_0$ and a CPT is not required for identification.}\footnote{To see why (ref) holds, note that path $D \leftarrow U \rightarrow Y_0 \leftarrow U_{Y_1} \rightarrow \Delta Y_1(0) \leftarrow X_1(0)$ is blocked by collider $\Delta Y_1(0)$ even conditional on the other collider $Y_0$.} Then, conditional independencies (ref) and (ref) collapse to $\Delta Y_1(0) \perp\!\!\!\perp D \mid X_0$,\footnote{To see why, note that conditional independencies (ref) and (ref) imply $(X_1(0),\Delta Y_1(0)) \perp\!\!\!\perp D \mid X_0$ by contraction and $\Delta Y_1(0) \perp\!\!\!\perp D \mid X_0$ by decomposition.} which justifies standard DiD identification using CPT with pre-treatment covariates only, i.e. $\mathbb{E}[\Delta Y_1(0) \mid X_0, D = 0] = \mathbb{E}[\Delta Y_1(0) \mid X_0, D = 1]$. This observation provides an alternative way to prove Corollary 1 (1) of caetano_difference_2024.
Additionally, caetano_difference_2024 show in Corollary 1 (2) that combining (ref) and (ref) permits the non-standard identification result involving nested conditional expectations $ATT = \mathbb{E}[\Delta Y_1 | D=1] - \mathbb{E}[\mathbb{E}[\mathbb{E}[\Delta Y_1 \mid X_0, X_1,D = 0] \mid X_0, Y_0, D = 0] \mid D=1]$. This is conceptually interesting because it demonstrates that identification is possible through a particular conditioning strategy involving $Y_0$, even though CPT conditional on $Y_0$ is not attainable (see Section (ref)). However, the practical relevance of this identification result appears limited. Appendix (ref) shows that there exists no $\Delta$-SWIG with unobserved confounding in which (ref) and (ref) hold but (ref) and (ref) do not hold. Thus, there is at least no graphical motivation to prefer the alternative identification result over standard conditional DiD identification.
caetano_difference_2024 also consider a covariate exogeneity assumption $(X_1(0) \mid X_0, D=1) \sim (X_1(0) \mid X_0, D=1)$ that is combined with CPT (ref) to obtain additional identification results. A sufficient condition for covariate exogeneity is to remove edge $D \rightarrow X_1$ in DAG (ref). However, this would be stronger than required. Covariate exogeneity can therefore not be represented in a $\Delta$-SWIG without modifications that are beyond the scope of the paper. This shows that $\Delta$-SWIGs are not compatible with all identifying assumptions that appear in the DiD literature. In particular, distributional assumptions like covariate exogeneity or parametric functional form assumptions are not naturally captured by $\Delta$-SWIGs. However, in many cases we expect that $\Delta$-SWIGs are applicable to justify at least parts of the identification arguments. For example, independence (ref) can be justified by $\Delta$-SWIG (ref) but covariate exogeneity requires arguments going beyond the causal structure.
The previous section largely revisits known results in the well-studied 2x2 setting through the $\Delta$-SWIG lens. This section uses the graphical perspective to establish new results and to generalize or refine existing ones in the multi-period DiD setting of callaway_differenceindifferences_2021. The following Section (ref) discusses then practical implications.
Consider $T = \mathcal{T} + 1$ time periods indexed by $t = 0,...,\mathcal{T}$. Let $\overline{V} = \{V_0,...,V_\mathcal{T}\}$ denote a sequence of time-varying random variables, and write $\overline{V}_t = \{V_0,...,V_t\}$, $\underline{V}_t = \{V_t,...,V_\mathcal{T}\}$ for $t > 0$, $\overline V_{t,t'} = \{V_t,\ldots,V_{t'}\}$ for $0 < t \leq t'$ and $\varnothing$ otherwise. The same notation applies to realizations, e.g. $\overline{v}$. Denote general differences as $\Delta V_{s,t} = V_{t} - V_{s}$ such that $\Delta V_{t} = \Delta V_{t-1,t}$.
callaway_differenceindifferences_2021 consider the setting with staggered treatment, where no unit is treated in period 0 and treatment is irreversible:
As $D_0 = 0$ in “all worlds”, it plays no role in the following arguments and is absent from all graphs. Therefore, we define treatment sequences as excluding $D_0$, i.e. $\overline{D} := \{D_1,...,D_\mathcal{T}\}$, $\overline{D}_t := \{D_1,...,D_t\}$, and $\overline{D}_0 := \varnothing$. This has two advantages: (i) the 2x2 setting of the previous section is nested as the special case with $T=2$ and $D = D_1$, (ii) potential variables obtained by fixing a treatment sequence have an intuitive structure, e.g. the potential outcome for being first treated in period 3 is denoted $Y_t(\overline{0}_{2},\underline{1}_{3}) = Y_t(0,0,1,1,\dots)$.
The target parameter in callaway_differenceindifferences_2021 is the average treatment effect of the group first treated in period $g$ at time $t$:
The definition using the sequence notation is less common in the DiD literature. However, it follows the mechanics of SWIGs and is used throughout. The curly brackets in (ref) illustrate how it maps to the common group notation used, e.g., in roth_whats_2023 where $G := min\{t : D_t = 1\}$.
Identification of $ATT(g,t)$ is based on conditional parallel trends with the never-treated (ref) or the not-yet-treated (ref):
$\mathbf Z$ serves here as a placeholder. Like in the previous section, we use $\Delta$-SWIGs to obtain concrete sets of valid control variables or to explain the mechanisms that prevent reading off CPTs.
Generalizing the previous settings, we focus on the class of SCMs $\mathfrak{M} := \{\mathcal{M} : \mathbf{V} = \{ \overline{X},\overline{D}, \overline{Y}, U \} \}$, where $\overline{X}$, $\overline{D}$, and $\overline{Y}$ denote the covariate, binary treatment, and outcome sequences, respectively, and, $U$ denotes unobservable time-invariant confounders.\footnote{We abstract from time-invariant covariates $X$ to avoid notational clutter. All expressions that require conditioning on time-varying covariates can additionally condition on $X$ without changing the result.} We also generalize single world additive separability to avoid stating a tailored version for each scenario:
Like in Assumption \nameref{ass:swas-2x2}, unobservable confounders enter never-treated potential outcomes only via an additively separable function and individual treatment effects are unrestricted. In particular, note that the latter may depend on the previous treatment sequence $\overline{D}_{t-1}$, which implies under Assumption \nameref{ass:staggered-trmnt} that the effects may depend on the timing of the first treatment.\footnote{A general single world additive separability without Assumption \nameref{ass:staggered-trmnt} is provided in Appendix (ref).} Assumption \nameref{ass:SWAS-staggered} not only nests Assumption \nameref{ass:swas-2x2} but also the untreated potential outcome models used in the literature. We collect them and show how they are special cases of Assumption \nameref{ass:SWAS-staggered} in Appendix (ref).
We also streamline the discussion by abstracting from outcome dynamics because they prevent reading off CPTs in model class $\mathfrak{M}$ (see discussion in Section (ref)):
For expositional simplicity, we begin by considering the case $T=3$. The mechanics extend directly to $3 < T < \infty$, as discussed in Section (ref). DAG (ref) shows the $T=3$ setting under Assumptions \nameref{ass:SWAS-staggered} and \nameref{ass:no-y-dyn} with dotted edge $D_{1}\rightarrow X_2$ because we first consider the common setting without treatment-covariate feedback $D_{1}\not\rightarrow X_2$ and then allow for edge $D_{1}\rightarrow X_2$.
The absence of treatment-covariate feedback resembles the setting discussed, e.g., in caetano_differenceindifferences_2024. The respective $\Delta$-SWIG is depicted in Figure (ref).
Period 1: Consider the short-term effect for the group that is first treated in period 1:
where the second line follows by adding $Y_0\textcolor{gray}{(1,1)} - Y_0\textcolor{gray}{(0,0)} = 0$ and rearranging.
Note that $\Delta$-SWIG (ref) contains multiple open paths between $\Delta Y_1(0\textcolor{gray}{,0})$ and $D_1, D_2(0)$ that preclude reading off unconditional joint independence $\Delta Y_1(0\textcolor{gray}{,0}) \perp\!\!\!\perp D_1, D_2$, which would be sufficient to identify the counterfactual trend in (ref). However, the graph implies two useful independencies and the second one is further processed to prepare for the next step:\footnote{(ref) follows because $X_t$ blocks $D_1 \leftarrow (U \rightarrow) X_t \rightarrow \Delta Y_1(0\textcolor{gray}{,0})$, collider $\Delta Y_2(0,0)$ blocks paths like $D_1 \leftarrow U \rightarrow X_2 \rightarrow \Delta Y_2(0,0) \leftarrow U_{Y_1} \rightarrow \Delta Y_1(0\textcolor{gray}{,0})$, and fixed node 0 and collider $D_2(0)$ block paths like $D_1 \leftarrow U \rightarrow D_2(0) \leftarrow 0 \rightarrow \Delta Y_1(0\textcolor{gray}{,0})$. $X_2$ does not open any path conditional on $X_0,X_1$ and is therefore optional. (ref) follows from similar arguments. Notably, $D_1$ is a collider on paths like $D_2(0) \leftarrow X_t \rightarrow D_1 \leftarrow X_1 \rightarrow Y_1(0\textcolor{gray}{,0})$ but conditional on $X_t$ the path is always blocked and $D_1$ is optional.}
While (ref) contains potential treatment $D_2(0)$, it becomes observable $D_2$ by conditioning on $D_1=0$ and consistency. The final line uses that $\Delta Y_1(0\textcolor{gray}{,0}) \perp\!\!\!\perp D_2 \mid X_0,X_1\textcolor{gray}{,X_2},D_1 = d$ for all $d \in \{0,1\}$ holds by Assumption \nameref{ass:staggered-trmnt}.\footnote{Then, $D_2 = 1$ - and is thus degenerate - conditional on $D_1 = 1$ such that $\Delta Y_1(0\textcolor{gray}{,0}) \perp\!\!\!\perp D_2 \mid X_0,X_1\textcolor{gray}{,X_2},D_1 = 1$ holds. Combined with (ref), this observation yields (ref).} Combining (ref) and (ref) by contraction yields then joint conditional independence
and a variety of CPTs for different control groups:
Therefore, the counterfactual trend in (ref) can be identified, e.g., using the never-treated as control group $\mathbb{E}[\mathbb{E}[\Delta Y_1 \mid X_0,X_1\textcolor{gray}{,X_2}, D_1 = 0, D_2 = 0] \mid D_1 = 1, D_2 = 1]$. This means that $ATT(1,1)$ is identified and can be estimated by controlling for $X_0,X_1$ (and $X_2$) via standard regression or (augmented) inverse probability estimators.
Additionally, conditional pre-trends for $ATT(2,2)$ are parallel. In particular, never-treated trend (ref) and later-treated trend (ref) coincide. Consequently, pre-trend analyses are not expected to reject parallel pre-trends if they control for $X_0,X_1$ (and $X_2$).
Period 2: $\Delta$-SWIG (ref) further implies $\Delta Y_2(0,0) \perp\!\!\!\perp D_1 \mid X_0,X_1,X_2$ and $\Delta Y_2(0,0) \perp\!\!\!\perp D_2(0) \mid X_0,X_1,X_2\textcolor{gray}{,D_1}$. Following similar arguments as before they contract to joint conditional independence
and justify identification of dynamic effect $ATT(1,2)$ and short-term effect $ATT(2,2)$ using the never-treated as control group:
Figure (ref) shows the $\Delta$-SWIG with treatment-covariate feedback. The main difference to Figure (ref) is that it includes now potential covariate $X_2(0)$. This has minor consequences for CPTs in period 1 but major consequences in period 2.
Period 1: First, note that $\Delta$-SWIG (ref) implies
which is nearly identical to (ref) but $X_2$ is not optional anymore. At first glance this suggests that CPT does not hold anymore conditional on the full covariate sequence $X_0,X_1,X_2$. However, the additional independence
implies that all results regarding identifiability of $ATT(1,1)$ and parallel pre-trends conditional on $X_0,X_1$ (and $X_2$) from the previous section apply identically with treatment-covariate feedback.\footnote{To see this formally, note that $\mathbb{E}[\Delta Y_1(0\textcolor{gray}{,0}) \mid X_0,X_1, D_1 = 1, D_2 = 1] = \mathbb{E}[\Delta Y_1(\textcolor{blue}{0}\textcolor{gray}{,0}) \mid X_0,X_1, D_1 = \textcolor{blue}{0}, D_2 = d] = \mathbb{E}[\Delta Y_1(\textcolor{blue}{0}\textcolor{gray}{,0}) \mid X_0,X_1,X_2, D_1 = \textcolor{blue}{0}, D_2 = d]$ where the first equality follows by (ref) and the second by (ref). Therefore, the counterfactual trend for $ATT(1,1)$ is $\mathbb{E}[\Delta Y_1(0\textcolor{gray}{,0}) \mid X_0,X_1, D_1 = 1, D_2 = 1] = \eqref{eq:nt-trend} = \eqref{eq:lt-trend} = \eqref{eq:nyt-trend}$.}
Period 2: $\Delta$-SWIG (ref) implies two crucial independencies that could not be simplified further:
Independence (ref) conditions on observable variables but only shows marginal independence regarding $D_2$. This leads to the following (non-)CPTs:
The short-term effect $ATT(2,2)$ can still be identified using the never-treated trend. However, the dynamic effect $ATT(1,2)$ is not identified in the presence of treatment-covariate feedback.
Identification of $ATT(1,2)$ would require a joint independence conditional on observable variables. However, the minimal conditioning set to achieve joint independence is shown in (ref) and contains unobservable $X_2(0)$. This poses a dilemma where researchers can basically choose between two biases: only controlling for $X_0,X_1$ leads to confounding bias, but controlling additionally for $X_2 = D_1 X_2(1) + (1-D_1) X_2(0)$ instead of $X_2(0)$ leads to “wrong world control bias”.
Summary: Short-term effects $ATT(g,g)$ are still identified in the presence of treatment-covariate feedback. However, dynamic effects are not identified without further assumptions. Strikingly, pre-trends hold by construction in this setting and would not flag problems with the dynamic effect estimates. This rationalizes the patterns shown in the introduction.
This section extends the analysis beyond $T = 3$ and first differences. The latter is motivated by the conditional DiD estimand targeting $ATT(g,t)$ that involves general differences $\Delta Y_{g-1,t}$:\footnote{We focus here on the estimand using never-treated but note that everything discussed in this section applies identically for the not-yet-treated, as we show in Appendix (ref).} {
}
for some adjustment set $\mathbf{Z}$. This section discusses which adjustment sets ensure, under which model restrictions, that $\Delta Y_{g-1,t} (\overline{0}) \perp\!\!\!\perp \underline{D}_{g} \mid \mathbf{Z}, \overline{D}_{g-1} = \overline{0}_{g-1}$ and therefore $ATT(g,t) = DiD_{g,t}\left(\mathbf{Z}\right)$.\footnote{The conditional independence implies CPT $\mathbb{E}[\Delta Y_{g-1,t} (\overline{0}) \mid \mathbf{Z}, \overline{D} = \overline{0}] = \mathbb{E}[\Delta Y_{g-1,t} (\overline{0}) \mid \mathbf{Z}, \overline{D}_{g-1} = \overline{0}_{g-1}, \underline{D}_g = \underline{1}_g]$, which is an essential part of the standard identification proof that is provided in Appendix (ref) for completeness.} To this end, we first show how to graphically obtain the minimal sufficient adjustment set, which can contain unobservable potential covariates and might be infeasible. Nevertheless, this set provides a constructive starting point for discussing the feasible adjustment sets in the following.
The previous sections establish how independencies leading to sufficient adjustment sets can be deduced from $\Delta$-SWIGs for $T \leq 3$, e.g. (ref), (ref), (ref), (ref). The same strategies apply to larger $T$ but require to either draw many $\Delta$-nodes within one $\Delta$-SWIG or many $\Delta$-SWIGs. To circumvent this complexity, we show that the single SWIG from step 3 in Procedure (ref) suffices to compactly reason about minimal sufficient adjustment sets under Assumptions \nameref{ass:SWAS-staggered} and \nameref{ass:no-y-dyn}, building on the following result:
The proposition shows that the minimal sufficient adjustment set consists of graphical parents $Pa_{\mathcal{G}(\overline{0})}$, excluding exogenous variables and $U$, of the two level nodes in the intermediate SWIG that map to $\Delta Y_{g-1,t} (\overline{0})$ in a $\Delta$-SWIG. $\mathbf{S}_{g,t}(\overline{0})$ might contain unobservable potential covariates and we define their observable analogue as $\mathbf{S}_{g,t} := \{X_t \in \overline{X} : X_t(\overline{0}) \in \mathbf{S}_{g,t}(\overline{0})\}$, which plays an important role in characterizing feasible sets below.
Consider Figure (ref) for illustration. It depicts the intermediate SWIGs between DAG (ref) and $\Delta$-SWIGs (ref) and (ref). For example, Proposition (ref) applied to the SWIG without treatment-covariate feedback in Figure (ref) leads to
or applied to SWIG (ref) with treatment-covariate feedback yields
The last example is an instance where the minimal sufficient adjustment set is not enough for identification because potential covariate $X_{2}(0)$ is not observable. However, Proposition (ref) still provides a useful starting point to discuss feasible adjustment sets under different model restrictions.
To organize the discussion, we adapt the concept of valid adjustment sets (VAS) from the causal graph literature on unconfoundedness for the DiD setting shpitser_validity_2012. Define VAS as all sets of time-varying covariates ensuring that the conditional DiD estimand is unbiased for $ATT(g,t)$ in all SCMs that obey restrictions $R$, i.e. $\mathcal{Z}_{g,t}(R) := \{\mathbf{Z} \subseteq \overline{X} : ATT(g,t) = DiD_{g,t}(\mathbf{Z}) ~\forall~ \mathcal{M} \in \mathfrak{M} \text{ satisfying }R \}$. For example, the discussion in Section (ref) implies already that $\mathcal{Z}_{g,t}(\text{\nameref{ass:SWAS-staggered}}) = \varnothing ~\forall~g,t$, i.e. no VAS exists under additive separability Assumption \nameref{ass:SWAS-staggered} alone. We gradually impose now additional model restrictions to study existence and form of the respective minimal VAS $\mathbf{Z}^{min}_{g,t}(R)$, which represents the minimal element of $\mathcal{Z}_{g,t}(R)$. The focus on the minimal VAS has two reasons: First, it is arguably the most interesting VAS from a practical perspective. Second, the maximal VAS and sets in between follow directly once the minimal VAS is established, as we show at the end of the section.
Minimal valid adjustment sets: Proposition (ref) provides the minimal sufficient adjustment set under Assumptions \nameref{ass:SWAS-staggered} and \nameref{ass:no-y-dyn}. It also implies the existence of minimal VAS for $g>t$ without further restrictions. This follows directly from (ref) because in pre-treatment periods any potential variables in $\mathbf{S}_{g,t}(\overline{0})$ are observable conditional on $\overline{D}_{g-1} = \overline{0}_{g-1}$ such that the independence still holds if $\mathbf{S}_{g,t}(\overline{0})$ is replaced by observable $\mathbf{S}_{g,t}$. Thus, $\mathbf{Z}^{min}_{g,t}(\text{\nameref{ass:SWAS-staggered}},\text{\nameref{ass:no-y-dyn}}) = \mathbf{S}_{g,t} = \overline{X}_{g-1} ~\forall~g>t$.\footnote{For example, in both SWIGs of Figure (ref) $\mathbf{Z}^{min}_{2,0}(\text{\nameref{ass:SWAS-staggered}},\text{\nameref{ass:no-y-dyn}}) = \mathbf{S}_{2,0} = \{X_0, X_1\}$.} In contrast, for all $g \leq t$, no VAS exists without further restrictions. This shows that parallel pre-trends require fewer restrictions than identifying the treatment effects of interest.\footnote{Note that $ATT(g,t) = 0, ~\forall~g>t$ by construction (see end of the identification proof in Appendix (ref)). Thus, $ATT(g,t) = DiD_{g,t}(\mathbf{Z}) = 0$ by definition if $\mathbf Z$ is a VAS. As $DiD_{g,t}(\mathbf{Z}) = 0, g > t$ is interpreted as parallel pre-trends, a VAS for all $g > t$ implies parallel pre-trends.} We illustrate and discuss practical implications of this finding below.
To identify the short-term effect $ATT(g,g)$ via a VAS, it suffices to assume that $D_t$ does not affect $X_t$:
This assumption holds naturally if $X_t$ precedes $D_t$ within time period $t$. This is the case in most settings depicted above. The exception is the setting of caetano_difference_2024 discussed in Section (ref) where Assumption \nameref{ass:x-d-ordering} holds after removing edge $D_t \rightarrow X_t$. Proposition (ref) yields $\mathbf{Z}^{min}_{g,t}(\text{\nameref{ass:SWAS-staggered}},\text{\nameref{ass:no-y-dyn}},\text{\nameref{ass:x-d-ordering}}) = \mathbf{S}_{g,t} = \overline{X}_{g}~\forall~g \geq t$ such that $ATT(g,g)$ is identifiable under Assumptions \nameref{ass:SWAS-staggered}, \nameref{ass:no-y-dyn}, and \nameref{ass:x-d-ordering}. Example 3 in Section (ref) illustrates this result via SWIG (ref) for $ATT(2,2)$ where $\mathbf{Z}^{min}_{2,2}(\text{\nameref{ass:SWAS-staggered}},\text{\nameref{ass:no-y-dyn}},\text{\nameref{ass:x-d-ordering}}) = \mathbf{S}_{2,2} = \{X_{0}, X_{1}, X_{2}\}$.
Identification also of the dynamic effects $ATT(g,t)$ for all $g,t$ requires further restrictions. The most common restriction is to rule out treatment-covariate feedback from $D_t$ to future covariates caetano_differenceindifferences_2024, ghanem_selection_2024:
Then, the minimal sufficient adjustment set equals the minimal VAS, i.e $\mathbf{Z}^{min}_{g,t}(\text{\nameref{ass:SWAS-staggered}},\text{\nameref{ass:no-y-dyn}},\text{\nameref{ass:x-d-ordering}},\text{\nameref{ass:no-d-x}}) = \mathbf{S}_{g,t} = \mathbf{S}_{g,t}(\overline{0}) = \overline{X}_{t} ~\forall~g,t$ because SWIG $\mathcal{G}(\overline{0})$ contains no potential covariates. Example 2 in Section (ref) illustrates this result, while Example 4 shows why dynamic effects are not identified without a further restriction.
So far, the added restrictions lead to the existence of additional minimal VAS. However, restrictions can also further shrink already existing minimal VAS. One prominent example is to rule out that $X_t$ directly affects future outcomes borusyak_revisiting_2024, ghanem_selection_2024, caetano_differenceindifferences_2024:
The SWIGs in Figure (ref) fulfill this assumption if all dotted edges are absent. Then, each potential outcome node is affected by only one contemporary covariate leading to adjustment sets with two elements, one element for each level of the respective difference.\footnote{Consequently, Examples 1 and 3 in Section (ref) would not include $X_0$, and Examples 2 and 4 would not include $X_1$. Appendix (ref) discusses also a $T=4$ example for illustration.}
Finally, also the remaining contemporary covariate-outcome effect can be ruled out:
Combined with \nameref{ass:no-x-dyn} this would rule out any direct covariate-outcome effects and would be encoded by removing all edges from $X_t$ to outcome nodes in the graphs above. Then, the empty set is the minimal VAS regardless of the presence of treatment-covarate feedback, i.e. $\mathbf{Z}^{min}_{g,t}(\text{\nameref{ass:SWAS-staggered}},\text{\nameref{ass:no-y-dyn}},\text{\nameref{ass:no-x-dyn}},\text{\nameref{ass:no-x-y}}) = \varnothing ~\forall~g,t.$
The remaining possible restrictions within model class $\mathfrak{M}$ are (i) covariate-covariate, (ii) covariate-treatment, (iii) treatment-treatment, and (iv) treatment-outcome. (i) and (ii) would not refine the presented results as they neither affect the parents of $Y_t(\overline{0})$ nor the presence of potential covariates. (iii) would rule out staggered treatment. (iv) would rule out the effects of interest. We therefore ignore these possibilities. Table (ref) in the next section summarizes the minimal VAS under the relevant restrictions and discusses practical implications.
All valid adjustment sets $\mathcal{Z}_{g,t}(R)$: In the $T=3$ settings of Section (ref), whenever identification holds conditional on a subsequence of $\overline{X}$, it also holds conditional on the full sequence. This result generalizes and allows to compactly characterize all VAS:
The proof in Appendix (ref) uses that $\Delta Y_{g-1,t}(\overline{0}) \perp\!\!\!\perp \overline{X} \setminus \mathbf{Z}^{min}_{g,t}(R) \mid \mathbf{Z}^{min}_{g,t}(R), \overline{D} = \overline{0}$ holds whenever $\mathbf{Z}^{min}_{g,t}(R)$ exists and generalizes the arguments for $T=3$ around (ref).
The key insight from Proposition (ref) is that the full sequence $\overline{X}$ is always the maximal VAS if a minimal VAS exists. Also, any subset of $\overline{X}$ that contains the minimal VAS remains a VAS.
Table (ref) summarizes the minimal VAS for different questions under the different restrictions. It also implicitly characterizes all VAS according to Proposition (ref). Each cell containing a minimal VAS can be augmented with any combination of other time-varying covariates up to the full sequence and still remains a VAS.\footnote{Table (ref) in Appendix (ref) provides such an augmented table for completeness.} Table (ref) highlights how additional restrictions expand the range of identifiable parameters or shrink the size of the minimal VAS. In the following we discuss how the table can be used to justify adjustment strategies, how it assist in interpreting pre-trend diagnostics, and conclude with practical recommendations.
Table (ref) allows to read off valid adjustment sets after researchers commit to a particular set of assumptions. For illustration, consider the least restrictive potential outcome model in caetano_differenceindifferences_2024 that reads in our notation $Y_t(\overline{0}) = \alpha(U) + g_{Y_t}(X_t) + U_{Y_t}$ such that Assumptions \nameref{ass:SWAS-staggered}, \nameref{ass:x-d-ordering}, \nameref{ass:no-d-x} and \nameref{ass:no-x-dyn} hold by construction, and \nameref{ass:no-y-dyn} is usually also directly or indirectly imposed.\footnote{ We interpret the fact that observable $X_t$ appears in an equation of a potential outcome as ruling out treatment-covariate feedback because otherwise $X_t(0)$ would enter. Either way, caetano_differenceindifferences_2024 rule out treatment-covariate feedback explicitly in the text such that Assumptions \nameref{ass:x-d-ordering}/\text{\nameref{ass:no-d-x}} hold. Also, single $X_t$ and not sequence $\overline{X}_t$ enters the equation such that Assumption \text{\nameref{ass:no-x-dyn}} holds. State dependence holds because previous outcomes do not enter the equation. Outcome-covariate feedback is also ruled because $X_t$ would otherwise become a potential covariate due to a $D_{t-1} \rightarrow Y_{t-1} \rightarrow X_t$ path. Thus, only outcome-treatment feedback is not ruled out by the functional form. However, it is usually ruled out as well bonhomme_back_2025. Assumption \text{\nameref{ass:no-y-dyn}} is therefore not covered directly implied by the functional from but often implicitly assumed.} The second last row of Table (ref) shows that $\{X_{g-1},X_t\}$ is the minimal VAS to identify $ATT(g,t)$ for all $g \leq t$ under these restrictions and thus all valid adjustment sets are characterized by $\mathcal{Z}_{g,t}(\text{\nameref{ass:SWAS-staggered}},\text{\nameref{ass:no-y-dyn}},\text{\nameref{ass:x-d-ordering}},\text{\nameref{ass:no-d-x}},\text{\nameref{ass:no-x-dyn}}) = \{\mathbf{Z} : \{X_{g-1},X_t\} \subseteq \mathbf{Z} \subseteq \overline{X} \}$ by Proposition (ref). This “forward engineering” of VAS has the advantage that the assumptions to be defended are the starting point of the argument and therefore very transparent. It also cleanly separates identification and estimation issues.
However, Table (ref) can also be consulted to “backwards engineer” assumptions under which a particular adjustment strategy could be justified. To this end, we determine those minimal VAS that are a subset of the adjustment strategy and read off sufficient assumptions on the left. We apply this for five adjustment strategies that are discussed in the literature and briefly comment on implications of the results:
Full covariate sequence: caetano_differenceindifferences_2024 assume parallel trends conditional on the full covariate sequence $\overline{X}$ and discuss estimation based on the identification result $ATT(g,t) = DiD_{g,t}(\overline{X})$. As all minimal VAS in Table (ref) are subsets of the full sequence, the restrictions in rows 4, 7 and 8 justify controlling for $\overline{X}$ to identify $ATT(g,t)$ for all $g,t$. Row 4 imposes the weakest restrictions, ruling out treatment-covariate feedback while allowing for unrestricted covariate-outcome effects. Even then the maximal VAS $\overline{X}$ only coincides with the minimal VAS $\overline{X}_t$ for $ATT(g,\mathcal{T})$. For $t < \mathcal{T}$, the same assumptions would justify controlling for $\overline{X}_t$ instead of $\overline{X}$. Using the minimal VAS therefore partly addresses the covariate dimensionality issues discussed in caetano_differenceindifferences_2024 Section 6 without imposing further assumptions.\footnote{The fraction of variables that can be removed without losing identification of $ATT(g,t)$ is $1 -(t+1)/T$, which is largest in early periods and increases with $T$.}
Early level and first difference: caetano_differenceindifferences_2024 further discuss to condition on $\{X_{g-1},\Delta X_{g-1,t}\}$ instead of $\overline{X}$ to reduce the covariate dimension for estimation. This is equivalent to conditioning on the two levels $\{X_{g-1},X_{t}\}$. Thus, Table (ref) reveals that this strategy can be justified by the restrictions in the last two rows that both limit covariate-outcome dynamics/effects. The reduced dimensionality requires therefore stronger assumptions compared to using $\overline{X}$. This highlights that the decision to move from $\overline{X}$ to $X_{g-1},X_{t}$ is not only an estimation issue but also has implications for identification.\footnote{The third option discussed in caetano_differenceindifferences_2024 is to control for the covariate average $T^{-1}\sum_t X_t$. We cannot provide a graphical way of justifying this strategy without violating the “the future cannot affect the past” principle and therefore do not discuss this strategy further.}
Period 0 controls: One common strategy is to match observations on $X_0$ and to run unconditional DiD on the matched sample dupas_women_2024,fenizia_organized_2024,humlum_changing_2025. This is equivalent to controlling for $X_0$. Only the empty set in the last row of Table (ref) is a subset of this one variable. Therefore, $X_0$ is a non-minimal VAS under this strictest set of assumptions. However, Assumption \nameref{ass:no-x-y} and \nameref{ass:no-x-dyn} could be relaxed to only hold for $t > 0$ making $X_0$ the minimal VAS.
Earlier period control: The description of function att_gt() in the popular did R package callaway_did_man states that "[...] in each 2x2 comparison, the covariates are taken to be the value of the covariates in the earlier time period [...]".\footnote{The quote is taken from version 2.3.0.} This corresponds to single covariate $X_{min(g-1,t)}$. Again, this is a non-minimal VAS under the strictest set of assumptions in Table (ref).\footnote{Similar to period 0 controls, Assumption \nameref{ass:no-x-y} could technically be relaxed for one period $t'$ such that $X_{min(g-1,t)}$ is the minimal VAS for all $ATT(g,t)$ with $g-1 = t'$ or $t = t'$.}
Pre-treatment controls: callaway_differenceindifferences_2021 discuss controlling for “pre-treatment” variables $X$. The absence of a time index suggests that $X$ is limited to time-invariant covariates. However, it might be interpreted differently as in ghanem_when_2026 Section III that discusses controlling for $X_{g}$. Again, this can be justified by ruling out that time-varying covariates directly affect the outcome at all. ghanem_when_2026 discuss alternative sufficient conditions without imposing $R_t^{XY}$ but adding restrictions beyond the causal structure.
Summary: The results in Table (ref) enable a principled analysis of different adjustment strategies. The most compact discussion is possible if assumptions are stated first to directly derive all VAS that ensure identification. Selecting the concrete VAS for estimation requires then statistical considerations. The minimal VAS is a plausible default choice but selecting the optimal VAS in terms of statistical efficiency in a principled manner is a promising direction for future research.
The “covariates first, assumptions second” direction is less transparent but reveals noteworthy insights. The benefit of controlling for one pre-treatment covariate as in the final three strategies appears limited. Their inclusion can be justified under a causal structure without direct covariate-outcome effects. Then, $X_0$, $X_{min(g-1,t)}$ or $X_{g}$ are all VAS, but not the minimal VAS. If such covariate-outcome restrictions are considered implausible, the alternative is to rule out treatment-covariate feedback. Then, even the minimal VAS contains post-treatment but pre-outcome covariates. In particular, controlling for pre-outcome covariates $\overline{X}_t$ requires the least additional restrictions to identify $ATT(g,t)$ via DiD.
It is important to note that the graphical perspective is only one possibility to justify an adjustment strategy. For example, ghanem_selection_2024 and ghanem_when_2026 provide alternative templates based on the treatment selection mechanism that go beyond the causal structure and are thus not naturally covered in the graphical approach. One advantage of our approach is that Table (ref) also implies testable implications to complement graphical reasoning with statistical evidence, as we discuss in the following.
As discussed in the introduction, insignificant parallel pre-trends (PPT) often serve as the only justification for CPT in practice. However, bonhomme_back_2025 notes that while rejections of PPT are informative, non-rejections should be interpreted with caution. Table (ref) provides a basis for clarifying what exactly can be learned from such rejections and where substantive reasoning is still required. For a compact discussion, define five conditional DiD estimands with superscripts related to the number of involved time-varying covariates: $\delta_{g,t}^{0} := DiD_{g,t}(\varnothing)$, $\delta_{g,t}^{2} := DiD_{g,t}(X_t,X_{g-1})$, $\delta_{g,t}^g := DiD_{g,t}(\overline{X}_{g-1})$, $\delta_{g,t}^t := DiD_{g,t}(\overline{X}_{t})$, and $\delta_{g,t}^{T} := DiD_{g,t}(\overline{X})$.
Rows 2, 5 and 8 of Table (ref) inform about which assumptions can be rejected if PPT are rejected. The restrictions in the left part of the table are sufficient for PPT with the adjustment set on the right. Therefore, rejecting the respective conditional PPT makes a violation of at least one of the assumptions on the left necessary. Concretely,
This reveals a nested structure where specifications with more covariates are more informative about which assumptions are violated.\footnote{Note that hypotheses adding even more covariates, e.g. $H_0^{T} : \delta_{g,t}^T = 0 ~\forall~g > t = 0 ~\forall~g > t$, hold under all three sets of sufficient assumptions and they all contain \{{\nameref{ass:SWAS-staggered}, \nameref{ass:no-y-dyn}}\}. Consequently, these two assumptions alone are sufficient for $H_0^T$, implying that it does not assess additional restrictions beyond those that justify $H_0^g$.} Rejecting hypotheses corresponding to more assumptions can be informative about the violated assumptions through comparisons with nested, less restrictive specifications:
The practically most relevant question is to what extent different outcomes of the hypothesis tests can replace or complement economic reasoning. Rejections are highly informative and provide strong evidence against specifications that rely on assumption sets containing the rejected assumptions, potentially overruling economic arguments in favor of these specifications. The implications of non-rejections are more ambiguous, and using them to justify the credibility of post-treatment results requires additional arguments:
$H_0^{0}$ not rejected: Arguing in favor of specifications without time-varying covariates based on a non-rejection of $H_0^{0}$ is relatively credible, as identification of post-treatment effects and pre-treatment PPT relies on the same set of assumptions (see final row of Table (ref)).\footnote{This observations is closely related to the discussion in Section II.A in ghanem_when_2026.} A remaining, well-documented obstacle is statistical in nature: failure to reject may reflect insufficient power rather than validity of the null roth_pretest_2022. Thus, arguing for credibility based on not rejected PPT still requires arguments that the tests are sufficiently powered.
$H_0^{0}$ rejected, $H_0^{2}$/$H_0^{g}$ not rejected: Credibility of $\delta_{g,t}^t$ and $\delta_{g,t}^2$ is at least not rejected by the data. However, only a subset of the assumptions that are sufficient for their unbiasedness is tested. In particular, both estimands require to assume \{\nameref{ass:x-d-ordering},\nameref{ass:no-d-x}\} to be unbiased. These assumptions are by construction untestable using pre-treatment trends, as treatment-covariate feedback unfolds only post-treatment. Therefore, statistical arguments that violations of the tested assumptions would be detected should be complemented by economic arguments supporting the absence of treatment-covariate feedback to rule out wrong world control bias when applying $\delta_{g,t}^t$ or $\delta_{g,t}^2$.
$H_0^{0}/H_0^{2}$ rejected, $H_0^{g}$ not rejected: $\delta_{g,t}^t$ is the only option for which at least parts of the sufficient assumptions are not rejected by the data. However, assessing its credibility still requires reasoning about the plausibility of the untested assumptions \{{\nameref{ass:x-d-ordering}, \nameref{ass:no-d-x}}\}.
$H_0^{g}$ rejected: This is the most informative but also the most destructive testing outcome. It cleanly rejects at least one of the baseline assumptions \{\nameref{ass:SWAS-staggered},\nameref{ass:no-y-dyn}\} that are required to justify all VAS implied by Table (ref) and Proposition (ref). Thus, no specification with or without time-varying covariates is expected to deliver credible results. Importantly, even if $H_0^{2}$ and/or $H_0^{0}$ are not rejected, this is likely due to bias cancelling and/or power differences and does not change this conclusion. Justification of (conditional) DiD requires then arguments beyond additive separability and causal structure.
We conclude with some practical recommendations for researchers who consider using time-varying covariates in their conditional DiD that are informed by the previous findings and discussions:
Graphically represent domain knowledge: Drawing the DAG, SWIG and/or $\Delta$-SWIG based on domain knowledge facilitates communication and reasoning about the causal structure. Once a structure is established, Table (ref) in combination with Proposition (ref) provides the sets of good control variables that follow directly from the structure.
Test pre-trends with all pre-treatment covariates: Rejecting parallel pre-trends conditional on $\overline{X}_{g-1}$ rejects standard identification arguments involving additive separability and absence of outcome dynamics. Conditional DiD should then be justified by alternative arguments or discarded.
Consider three specifications: Running conditional DiD (i) without time-varying covariates, (ii) with $\overline{X}_t$ and (iii) with $\{X_{g-1}, X_t\}$ provides a principled way to gather statistical evidence and may help to narrow down plausible specifications. In particular, rejected parallel pre-trends for some or all specifications are informative about crucial assumptions, as discussed in the previous two sections.
Use full covariate sequence only for computational convenience: Controlling for full covariate sequence $\overline{X}$ instead of pre-outcome sequence $\overline{X}_t$ does not require weaker assumptions. This means that the covariate dimension would be much larger than required for early $t$. However, applying the pre-outcome sequence requires to run a different specification for each $t$, e.g. leading to $\mathcal{T}$ propensity scores to be estimated. Running one specification with $\overline{X}$ can therefore be attractive from a computational perspective.
This paper introduces $\Delta$-SWIGs as a graphical tool to reason about conditional parallel trends in a transparent and principled manner. They provide a natural starting point to scrutinize and justify control variables in DiD analyses based on the causal structure but without distributional or parametric assumptions. The only functional form assumption that must be defended outside of the graph is the single world additive separability that generalizes a standard functional form assumption in the literature. A transparent tool to reason about this substantive assumption is still missing.
We apply $\Delta$-SWIGs to study time-varying control variables in a model class with unrestricted unobservables and binary staggered treatment. However, the $\Delta$-SWIG technology can similarly be applied to model classes with more variable categories, restrictions on unobservables, beyond staggered treatment, or beyond binary treatments. For the latter two, the single world additive separability most likely must be strengthened to hold for all treatment levels that serve as control group at some point, thereby restricting treatment effect heterogeneity. A detailed discussion along these lines is left for future research.
We also abstract from estimation and testing issues in this paper. One interesting follow-up question is regarding the interplay between different valid adjustment sets and overlap, which most likely interacts with the question of the most efficient valid adjustment set. Another interesting direction is an exhaustive conceptual and statistical study of the testable implications of the model structure, and the development of powerful tests building on them.
Finally, the paper reinforces recent reservations about relying solely on parallel pre-trends to justify a DiD analysis. In particular, when time-varying covariates are used in the analysis, they are not informative about all required assumptions for unbiased post-treatment estimates, even in large data sets.
\printbibliography