Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
100,693 characters · 22 sections · 110 citation commands
When are time series predictions causal? The potential system and dynamic causal effects
Keywords: Causality, design-based inference, impulse response function, potential outcomes, sequential assignment, time series.
\baselineskip=20pt
Let $Y_{t+h}$ be an outcome at time $t+h$, where $h \ge 0$ is a horizon and $t$ is the current time, $X_t$ is a feature (which may be, e.g., an observed confounder), $A_t$ is an assignment and $D_{1:t-1}$ are past outcomes, features and assignments. When do time series data-based predictions, such as $$ \textnormal{E}[Y_{t+h}\mid X_t,D_{1:t-1},A_t=a_t] - \textnormal{E}[Y_{t+h}\mid X_t,D_{1:t-1},A_t=a_t'], $$ measure how changes in the assignment at time $t$ cause the outcomes at time $t+h$ to move? This paper provides sufficient nonparametric conditions to answer this type of question.
Our approach is based on defining a foundational “potential system,” denoted {\tt PS}. It directly connects familiar time series objects like impulse response functions to average treatment effects and, more generally, the time series causality literature to the nonparametric causal inference literature based either on population or design-based inference strategies.
This paper is closely related to five time series papers as well as a stream of panel data papers associated with James M. Robins. BojinovShephard(19), RambachanShephard(21) and LinDing(25) worked with a potential outcome based time series model, while AngristKuersteiner(11) and AngristJordaKuersteiner(18) define and work with what we call “branch potential outcomes.” Both potential outcomes and branch potential outcomes appear as a part of our {\tt PS} (and hence our system could be thought of as providing the primitives to these four papers).
BojinovShephard(19) provide many references to the literature on dynamic causal effects using potential outcome type objects. The vast majority of the work in this type of literature on dynamic causal effects considers panel data, not pure time series, which is the subject of this paper. The panel data literature is reviewed in, for example, HernanRobins(18), ArkhangelskyImbens(24), and ChernozhukovNeweySinghSyrgkanis(2023).
A linear special case of the {\tt PS} is a structural vector autoregression (SVAR), the workhorse of modern applied linear time series methods. Reviews of some of this work focusing on macroeconomics include KilianLutkepohl(17), FernandezVillaverdeRubioRamirez(10), StockWatson(18), Ramey(16) and JordaTaylor(24).
“Granger causality” has played an important role in time series over the last 50 years. Though inspiring, Granger causality is not really about causality, but about prediction. (“The definition of causality used above is based entirely on the predictability of some series” Granger(69).) Some of the literature on Granger causality is discussed in, for example, Kuersteiner(10), WhiteLu(10) and ShojaieFox(22).
HarveyDurbin(86) tried to assess the causal impact of a one-time assignment using a time series model, applying it to assess the causal impact of the introduction of compulsion of seat belt wearing on driver deaths in the UK. Synthetic control (AbadieGardeazabal(03) and AbadieDiamondHainmueller(10)) is a similar idea, but enriched. There a multivariate set of outcomes move together but only one is impacted by the intervention. The multivariate data can help pin down the intervention under some assumed model. Abadie(21) provides a review.
A separate ocean of work on time series causality focuses on “control,” where an engineer builds a system which collects data to control an output in some optimal way to their benefit. The most famous version of this is linear/quadratic controller, e.g., Whittle(82),Whittle(83),Whittle(90a),Whittle(96), HansenSargent(14) and HerbstSchorfheide(15). Under the {\tt PS} the assignments can be selected to minimize expected loss given the past data --- as you would see in the control literature. Hence the {\tt PS} bridges observational reduced form models, experiments for time series, and control models of dynamic decision making.
Much of the more modern material on control is phrased in terms of Markov decision processes (e.g., Puterman(05)). Sometimes the researcher uses the data to learn the Markov decision process itself. That area is usually called reinforcement learning (e.g., SuttonBarto(18)). The stationary Markov version of the {\tt PS}, augmented with a loss function, again forms a bridge to this literature.
Learning optimal dynamic treatment rules (or “policies” or “regimes”) is often phrased using potential outcomes (e.g., Murphy(03), NieBrunskillWager(2021), HeckmanNavarro(2007), ChernozhukovNeweySinghSyrgkanis(2023), VivianoBradic(24), BradicJiZhang(2024)). This literature connects to our work in the case of sequences of interventions, but it focuses on panel data. A notable recent exception is KitagawaWangXu(23), which learns optimal policies based on a single time series.
BojinovSimchiLeviZhao(22) and BasseDingToulis(20) look at “switchback designs” to optimally learn sequences of treatment effects from time series. There is a large other literature on sequential experiments which is not phrased in terms of potential outcomes, e.g., Efron(71) and GlynnJohariRasouli(20), as well as the substantial literature on so-called $N$-of-$1$ trials which appear prominently in, for example, the study of personalized medicine (LilliePatayDiamantIssellTopolSchork(11)). The design-based content of the potential system provides a nonparametric foundation for these settings as well. Related recent work includes LiangRecht(25), SchnaffeAiroldi(26) and LinDing(25). The latter relates potential outcomes to regression in the design-based context.
Although phrased using potential outcomes, our system can also be written using directed acyclic graph (DAG) theory, expressed using the tools developed in the pioneering efforts of Pearl(09) and coauthors. Important related causal graph theory topics include the “Single World Intervention Graph” (SWIG) associated with RichardsonRobins(2013). We use SWIG graphs to illustrate the {\tt PS} and various constraints on the equential assignment mechanism.
The rest of this paper has six sections. The potential system and different measures of the dynamic causal effects are defined in Section (ref). Section (ref) explores several important examples of the {\tt PS} and its relationship to various other common models of causality in the time series literature. Section (ref) focuses on constraints on the sequential assignment mechanism and how they allow us to identify different measures of the dynamic causal effects. Section (ref) considers various extensions of the potential system, applying the framework to settings featuring instrumental variables, consecutive sequences of assignments, design-based causal inference, and stochastic dynamic programming (control). Section (ref) concludes. There is also an Appendix containing proofs.
Throughout, for any (random or deterministic) sequence $\{x_1, x_2, ..., x_T\}$ we denote for $T\ge s>t \ge 1$ the $x_{t:s} := \{x_t,...,x_{s}\}$, while $(A \perp \!\!\!\perp B)\mid C$ denotes variables $A$ and $B$ are conditionally independent given $C$.
The entire paper is based on the potential system, which we now define.
First, the data generating process is set up using two assumptions.
Second, the counterfactual process is set up using two assumptions.
The left-hand side of Figure (ref) visualizes the counterfactual paths of potential outcomes and confounders defined in {\tt CP.1} for binary assignments. The right-hand side shows $Z_{1:3}$ corresponding to $A_{1:3}=(1,1,0)$, highlighting an assigned path in bold.
The left-hand side of Figure (ref) visualizes the counterfactual paths of potential branches, again for binary assignments. The right-hand side shows $D_{t:t+2}$ corresponding to $A_t=1$, highlighting the assigned path.
Finally, the data generating and counterfactual processes are linked by one assumption, completing the definition of the potential system.
Having set up the {\tt PS}, it is now possible to define what a dynamic causal effect is and, in turn, how it can be summarized in the case it is stochastic.
We may summarize the dynamic causal effect in a number of ways.
The following are important special cases of the {\tt PS}. Going forward, for notational convenience, we define $D_t(a_{1:t}) := \{X_t(a_{1:t-1}),a_t,Y_t(a_{1:t})\}$ and $D_{1:t}(a_{1:t}) := \{D_1(a_1),...,D_t(a_{1:t})\}$.
We begin by discussing the nonparametric structural equation model (SEM) potential system, a highly general example of a {\tt PS} that adds useful additional structure to the potential outcomes and features.
This model can be viewed as requiring that Assumption {\tt DGP.2} defines a nonparametric structural equation model (NPSEM) or structural causal model (SCM), formalized by Pearl(95),Pearl(09). This framework for causal inference has many direct ties to potential outcomes frameworks for inference on counterfactuals (see, e.g., Imbens(2020) or Geffner(2022) for linkages).
We may further specialize Example (ref) by incorporating linearity, delivering the homogeneous linear Markov {\tt PS}.
Under the model of Example (ref), Assumption {\tt LP} implies that the DGP is \[
\] We may thus compactly write the DGP as a VAR(1): $$ D_{t}=\phi D_{t-1}+B\varepsilon _{t} $$ where \[ \phi := \left(
\right),\quad B:=\left(
\right). \] Writing the system this way makes clear it is Markovian.
For the counterfactual process, we start by writing $$
. $$ As $A_{t,0}(a_{t})=a_{t}$, \[ D_{t,0}(a_{t})= \left(
\right) D_{t-1}+\left(
\right) a_{t} + \left(
\right) \varepsilon _{t}, \] and, for $h=1,2,...,H$, we define \[ D_{t,h}(a_{t}):=\phi D_{t,h-1}(a_{t})+B\varepsilon _{t+h}. \]
The dynamic causal effect on all variables at horizion $h\geq 0$ is given by
which is non-stochastic. Thus (abusing notation slightly) $$ D_{t,h}(a_{t})-D_{t,h}(a_{t}^{\prime }) = {\tt ATE}_{t,h}(a_t,a_t') = {\tt CATE}_{t,h}(a_t,a_t') = {\tt FTE}_{t,h}(a_t,a_t') = {\tt CFTE}_{t,h}(a_t,a_t'). $$
Another important special case of the {\tt PS} occurs when there are no features and the assignments are independent through time. We may further specialize Example (ref) to explore this setting: the Homogeneous Markov news impact {\tt PS}.
The DGP under Example (ref) is then \[
\] Note further that the $h=0$ potential branch and $h > 0$ potential branch are, respectively, $$ Y_{t,0}(a_t) = \gamma(Y_{t-1},a_{t},W_t), \quad Y_{t,h}(a_t) = \gamma(Y_{t,h-1}(a_t),V_{t+h},W_{t+h}), $$ observing that $A_{t+h}(a_t) = V_{t+h}$ for all $h > 0$. The causal effect at $h=0$ is then $$ Y_{t,0}(a_t) - Y_{t,0}(a_t') = \gamma(Y_{t-1},a_{t},W_t) - \gamma(Y_{t-1},a_{t}',W_t) $$ and the causal effect at $h > 0$ is $$ Y_{t,h}(a_t) - Y_{t,h}(a_t') = \gamma(Y_{t,h-1}(a_{t}),V_{t+h},W_{t+h})-\gamma(Y_{t,h-1}(a_{t}'),V_{t+h},W_{t+h}). $$
If it exists, the marginal dynamic causal effect is, for $h \ge 0$,
which is typically stochastic.
News impact causal studies appear in financial econometrics, but are typically not expressed in causal language, and instead discuss “parameterized mechanisms.” In that literature, a major topic is understanding how time-varying volatility of speculative assets (e.g., BollerslevEngleNelson(94) and Shephard(05)) change in response to news (e.g., CampbellHentschel(92) and EngleNg(93)).
A further special case of this structure is the homogeneous Markov partially linear news impact {\tt PS}, which sets $Y_{t}(a_{1:t}) = \gamma Y_{t-1}(a_{1:t-1}) + \zeta(a_{t})$. Then $$ Y_t = \gamma Y_{t-1} + \zeta(V_{t}),\quad Y_{t,0}(a_t) = \gamma Y_{t-1} + \zeta(a_{t}), \quad Y_{t,h}(a_t) = \gamma Y_{t,h-1}(a_t) + \zeta(V_{t+h}) $$ for $h>0$. Thus, for $h\ge 0$, the dynamic causal effect is non-stochastic with $$ Y_{t,h}(a_t) - Y_{t,h}(a_t') = \gamma \{Y_{t,h-1}(a_t)- Y_{t,h-1}(a_t')\} = \gamma^h \{\zeta(a_{t}) - \zeta(a_{t}')\}. $$
We may also consider examples of the {\tt PS} that further restrict the temporal impact of assignments on potential outcomes and features.
For $m$-th order causal {\tt PS} we write the time $t$ potential outcome and feature using the shorthand $$Z_{t}(a_{t-m:t}),$$ burying the irrelevance of $a_{1:t-m-1}$. Again, this is a type of non-interference assumption Cox(58book).
In the important case of Definition (ref) where $\{\varepsilon_t\} = \{(U_t^\mathtt{T},V_t^\mathtt{T},W_t^\mathtt{T})^\mathtt{T}\}$ is a sequence of independent random vectors, then an $m$-order PS is $m$-order Markovian. Hence statistical methods designed for $m$-order Markov processes can be used in this setting, but now they have causal content.
The idea of exogeneity has a long history in econometrics and is defined in many different ways. Some of time series literature on this topic is discussed in EngleHendryRichard(83). Here we give a definition in the context of a {\tt PS}, viewing exogeneity as a form of invariance with respect to a possible intervention. That line of thought goes back at least to Simon(53).
{\tt PS}-exogeneity is an important condition for extending the potential system to applications in, e.g., design-based causal inference, discussed further in Section (ref).
With definitions and measures of dynamic causality now in place, we investigate assumptions and results that allow data-based predictions to have nonparametric causal interpretations.
A major way of progressing from data-based predictions to causality is to make assumptions that constrain the behavior of the {\tt SAM} from Assumption {\tt DGP.2}. To do this compactly we use the notation $$ Y_{t,0:H}(\mathcal{A}_t) = \{Y_{t,h}(a_{t}): a_t \in \mathcal{A}_t, h =0,1,...,H\}, $$ collecting the $a_t$-potential branch at different time horizons.
Under special cases of the potential system, the conditions {\tt SAM.BSU}, {\tt SAM.BSR}, {\tt SAM.BU} and {\tt SAM.BR} follow under more primitive conditions.
Assumption {\tt SAM.BSU} is stated as the conditional independence of $A_{t}$ and all the elements of $\{Y_{t,h}(a_{t}): a_t \in \mathcal{A}_t, h =0,1,...,H\}$. A formally weaker alternative condition is to require many pairs of conditional independence rather than a single very large joint conditional independence. This kind of pairwise assumption appears often in cross-sectional and panel population-based inference, for example HernanRobins(18). We define such conditions below.
Under Assumption {\tt SAM.BSU-}, Figure (ref) depicts the potential branch of a {\tt PS} as part of a Single World Intervention Template (SWIT), which is a concise representation of a set of Single World Intervention Graphs (SWIGs). The SWIG framework takes DAGs (Directed Acyclic Graphs) as inputs and “splits” them into SWIGs at nodes being intervened on, in this case at the node that represents $A_t$. Nodes downstream of the intervention become potential branches (the $Y_{t,0}(a_t)$ and $D_{t,h}(a_t)$ for $h=0,1,...,H$). (For other properties of SWIGs, see RichardsonRobins(2013).) At a glance, Figure (ref) shows that conditioning on $(X_t, D_{1:t-1})$ (which are acting as observed confounders) makes $A_t$ independent of branch potential outcomes (per standard analysis of probabilistic graphical models, blocking outgoing arrows from $(X_t, D_{1:t-1})$ separates $A_t$ from all other nodes), which is exactly what Assumption {\tt SAM.BSU-} states.
Using the conditions introduced in Definition (ref), the following Theorem (ref) shows that the causal summaries introduced in Definition (ref) can be expressed in terms of population-based predictive quantities, delivering a version of the promise at the start of this paper: providing conditions where the difference of two data-based predictions are causal at horizon $h\ge0$.
Such predictive quantities are not, in general, easy to estimate or approximate in practice. However, we have made progress, moving from counterfactuals to observables which can be modeled and predicted: we have “identified” the causal objects defined in the earlier sections.
Assumption {\tt SAM.BSR-} yields the SWIT in Figure (ref), which has pruned off the confounders compared to Figure (ref). The corresponding result relating causal quantities to data quantities under {\tt SAM.BSR-} is given in the third part of Theorem (ref). Assumption {\tt SAM.BU-} yields the SWIT given in the left-hand side of Figure (ref). This is the same as the SWIT for {\tt SAM.BSU-} except the dependence on the history is removed. The corresponding result relating causal quantities to data quantities under {\tt SAM.BU-} is given in the second part of Theorem (ref). Assumption {\tt AM.BR-} yields the SWIT for branch-randomization, shown on the right-hand side of Figure (ref). Confounders no longer appear in the graph. The corresponding result relating causal quantities to data quantities under {\tt SAM.BR-} is given in the first part of Theorem (ref).
We saw in Theorem (ref) various conditions on the {\tt PS} that allow the ${\tt ATE}_{t,h}(a_t,a_t')$, ${\tt CATE}_{t,h}(a_t,a_t')$, ${\tt FTE}_{t,h}(a_t,a_t')$ and ${\tt CFTE}_{t,h}(a_t,a_t')$ to be written as data-based (conditional) expectations. In some applications it is helpful to replace the data-based conditional expectations by projections. Economists often use linear projections (see, e.g., AngristPischke(09)), which we now consider in detail.
Before proceeding, to establish notation, denote the usual linear projection of a generic random variable $A$ on 1 and a generic random variable $B$ as $\operatorname{LP}[A\mid 1,B] = \kappa + \beta B$, where $\beta=\textnormal{Cov}(A,B)\textnormal{Var}(B)^{-1}$ and $\kappa = \textnormal{E}[A] - \beta \textnormal{E}[B]$, so that
For this setup to make sense, $A,B$ must both be in $L^2$ and $\textnormal{Var}(B)>0$. Here $\textnormal{E}_{A,B}$ is the expectation with respect to both $A$ and $B$. $\textnormal{E}_{B}$ is the expectation solely with respect to $B$.
Assume the {\tt PS} is in $L^2$ and work under the {\tt SAM.BR-} condition, so $$\textnormal{E}[Y_{t+h}(a_t)] = \textnormal{E}[Y_{t+h} \mid A_t=a_t].$$ Then the linear projection of $Y_{t+h}$ on 1, $A_t$ is $$ \operatorname{LP}[Y_{t+h} \mid 1,A_t=a_t] = \kappa_{t,h} + \beta_{t,h} a_t,\quad \text{where} \quad \beta_{t,h} = \textnormal{Cov}(Y_{t+h},A_t) \textnormal{Var}(A_t)^{-1} $$ assuming $\textnormal{Var}(A_t)>0$. So the linear projection of the conditional average treatment effect ${\tt ATE}_{t,h}(a_t,a_t') = \textnormal{E}[Y_{t+h}(a_t)] - \textnormal{E}[Y_{t+h}(a_t')]$ (which is non-stochastic) is $$ \operatorname{LP}[Y_{t+h} \mid 1,A_t=a_t] - \operatorname{LP}[Y_{t+h} \mid 1,A_t=a_t'] = \beta_{t,h}(a_t - a_t'), $$ (which is non-stochastic).
We now turn to a subtler case. Assume the {\tt PS} is in $L^2$ and work under the {\tt SAM.BU-} condition, so $$\textnormal{E}[Y_{t+h}(a_t) \mid X_t] = \textnormal{E}[Y_{t+h} \mid A_t=a_t,X_t].$$ Then the linear projection of $Y_{t+h}$ on 1, $A_t$ and $X_t$ is $$ \operatorname{LP}[Y_{t+h} \mid 1,A_t=a_t,X_t] = \kappa_{t,h} + \beta_{t,h} a_t + \delta_{t,h}X_t, $$ where
assuming $\textnormal{Var}((A_t^\mathtt{T},X_t^\mathtt{T})^\mathtt{T})>0$. So the linear projection of the conditional average treatment effect ${\tt CATE}_{t,h}(a_t,a_t') = \textnormal{E}[Y_{t+h}(a_t)\mid X_t] - \textnormal{E}[Y_{t+h}(a_t')\mid X_t]$ (which is stochastic) is $$ \operatorname{LP}[Y_{t+h} \mid 1,A_t=a_t,X_t] - \operatorname{LP}[Y_{t+h} \mid 1,A_t=a_t',X_t] = \beta_{t,h}(a_t - a_t') $$ (which is non-stochastic, but the first two moments of $X_t$ influence $\beta_{t,h}$).
Under covariance stationarity of $\{D_{t}\}$, one may write $\beta_{t,h}:=\beta_h$ for all $t$. A weaker assumption is to assume $\{D_{t}\}$ is only locally covariance stationary, where the dependence through time changes slowly (e.g., Dahlhaus(12)).
Covariance stationarity of $\{D_{t}\}$ is a sufficient condition to produce a time-invariant linear projected causal effect, but it is not necessary. By the Frisch-Waugh-Lovell theorem Yule(1907), we may also write \[ \beta_{t,h} = \textnormal{Cov}(Y_{t+h},A_t^\perp) \textnormal{Var}(A_t^\perp)^{-1} \] for $A_t^\perp := A_t - \operatorname{LP}[A_t \mid 1, X_t]$. As such, we can see it is also sufficient to only require that $\{(Y_{t}, A_t^\perp)\}$ is covariance stationary to yield $\beta_{t,h}:=\beta_h$ for all $t$. Moreover, for any random $B_{t+h}$ such that $\textnormal{Cov}(B_{t+h}, A_t^\perp)=0$, we have that \[ \beta_{t,h} = \textnormal{Cov}(Y_{t+h}-B_{t+h},A_t^\perp)\textnormal{Var}(A_t^\perp)^{-1}, \] i.e., regressing $Y_{t+h}-B_{t+h}$ on $A_t^\perp$ yields the same $\beta_{t,h}$ as regressing $Y_{t+h}$ on $A_t^\perp$. Certain choices of $B_{t+h}$ may help improve precision in downstream estimation, or it may be more plausible that $\{(Y_{t+h}-B_{t+h},A_t^\perp)\}$ is covariance stationary. An example of this is where $\{Y_{t}\}$ is an integrated variable but $\{B_{t}\}$ is a detrender or synthetic control.
Assume throughout this subsection strict stationarity of $\{D_t\}$ (so the dimensions of features, assignments and outcomes are time invariant) and the {\tt PS} is in $L^1$. Further assume $\mathcal{A}_t = \mathcal{A}$ for all $t$, i.e., the assignment space is not changing over time.
Under the condition {\tt SAM.BR-} the $$ \textnormal{E}[Y_{t+h}(a)] = \textnormal{E}[Y_{t+h}|A_t=a] = \mu_h(a), $$ where $\mu_h = \{\mu_h(a):a\in \mathcal{A}\}$ is a deterministic function. Thus the ${\tt ATE}_h(a,a') =\mu_h(a)-\mu_h(a')$. Typically, for strictly stationary processes, $\mu_h$ would be estimated as a nonparametric regression of $Y_{t+h}$ on $A_t$, e.g., through Nadaraya-Watson kernel regression, local linear regressions FanYao(05), splines, or neural networks.
Under the weaker condition {\tt SAM.BU-} the $$ \textnormal{E}[Y_{t+h}(a)|X_t] = \textnormal{E}[Y_{t+h}|A_t=a,X_t] = \mu_h(a,X_t), $$ where $\mu_h = \{\mu_h(a,x):a \in \mathcal{A},x \in \mathcal{X}\}$ is a deterministic function. Thus $$ {\tt CATE}_h(a,a') =\mu_h(a,X_t)-\mu_h(a',X_t)$$ which implies that, by iterated expectations $${\tt ATE}_h(a,a') =E_{X_1}[\mu_h(a,X_1)]-E_{X_1}[\mu_h(a',X_1)]. $$ The $\mu_h$ can be estimated by a non-parametric regression of $Y_{t+h}$ on $A_t$ and $X_t$.
Under the condition {\tt SAM.BSU-} and imposing the {\tt PS} is $m$-order Markovian (see Section (ref)), then $$ \textnormal{E}[Y_{t+h}(a)|X_t,D_{t-m:t-1}] = \textnormal{E}[Y_{t+h}|A_t=a,X_t,D_{t-m:t-1}]=\mu_h(a,X_t,D_{t-m:t-1}), $$ where $\mu_h = \{\mu_h(a,x,d):a\in \mathcal{A},x \in \mathcal{X}, d \in \mathcal{D}^m\}$ is a deterministic function. Thus $$ {\tt CFTE}_h(a,a') =\mu_h(a,X_t,D_{t-m:t-1})-\mu_h(a',X_t,D_{t-m:t-1})$$ which implies that, by iterated expectations $${\tt ATE}_h(a,a') =E_{X_{m+1},D_{1:m}}[\mu_h(a,X_{m+1},D_{1:m})]-E_{X_{m+1},D_{1:m}}[\mu_h(a',X_{m+1},D_{1:m})]. $$ The $\mu_h$ can again be estimated by a non-parametric regression of $Y_{t+h}$ on $A_t$, $X_t$ and $D_{t-m:t-1}$.
Under the condition {\tt SAM.BSR-} plus imposing the {\tt PS} is $m$-order Markovian, then $$ \textnormal{E}[Y_{t+h}(a)|D_{t-m:t-1}] = \textnormal{E}[Y_{t+h}|A_t=a,D_{t-m:t-1}]=\mu_h(a,D_{t-m:t-1}), $$ where $\mu_h = \{\mu_h(a,d):a\in \mathcal{A},d \in \mathcal{D}^m\}$ is a deterministic function. Thus $$ {\tt FTE}_h(a,a') =\mu_h(a,D_{t-m:t-1})-\mu_h(a',D_{t-m:t-1})$$ which implies that, by iterated expectations $${\tt ATE}_h(a,a') =E_{D_{1:m}}[\mu_h(a,D_{1:m})]-E_{D_{1:m}}[\mu_h(a',D_{1:m})]. $$ Here $\mu_h$ can be estimated by a non-parametric regression of $Y_{t+h}$ on $A_t$ and $D_{t-m:t-1}$.
Assume that the {\tt PS} is in $L^1$, that $\mathcal{A}_t$ is made up of a finite number of atoms, and that {\tt SAM.BSU-} holds. Define $\lambda_{a_t}(X_t,D_{1:t-1}):=P(A_{t}=a_t \mid X_t,D_{1:t-1})$, the propensity score, and assume it is bounded away from zero and one for all $a_t \in \mathcal{A}_t$. Recall that \[ {\tt CFTE}_{t,h}(a_t,a_t') = \textnormal{E}[Y_{t+h}\mid X_t,D_{1:t-1},A_{t}=a_{t}] - \textnormal{E}[Y_{t+h}\mid X_t,D_{1:t-1},A_{t}=a_{t}']. \] As such we have that \[ {\tt ATE}_{t,h}(a_t,a_t') =\textnormal{E}[{\tt CFTE}_{t,h}(a_t,a_t')] = \textnormal{E}[\textnormal{E}[Y_{t+h}\mid X_t,D_{1:t-1},A_{t}=a_{t}]] - \textnormal{E}[\textnormal{E}[Y_{t+h}\mid X_t,D_{1:t-1},A_{t}=a_{t}']]. \] From the literature on semiparametric inference, we know that this is a classic missing data functional Kennedy(2024), for which the influence curve (or “efficient influence function”) in a fully nonparametric observed-data model for $\left(Y_{t+h}, A_t, X_t, D_{1: t-1}\right)$ is
Influence curves are random objects that can be used to construct semiparametric efficient estimators, e.g., the doubly robust estimators that are familiar from RobinsRotnitzkyZhao(94) and HernanRobins(18), or the double/debiased machine learning estimators familiar from Chernozhukov(2018). Semiparametric efficient inference on nonparametrically defined impulse response functions is explored in part in, e.g., BallinariWehrli(24), building from RambachanShephard(21), and can be grounded in the {\tt PS}. The desirable properties of influence function-based estimators are discussed in many works (e.g., RobinsRotnitzkyZhao(94) or Chernozhukov(2018), or see Kennedy(2024) for an overview).
The {\tt PS} also accommodates identification of local summaries of causal effects using instrumental variables, in the spirit of ImbensAngrist(94).
The definition of the IV {\tt PS} imposes further (exclusion) restrictions on the causal relationships of the variables in the {\tt PS}. Under the IV {\tt PS}, we may write
where $\tilde D_{t}:=\left(A_{t}^{\mathtt{T}},Y_{t}^{\mathtt{T}}\right)^{\mathtt{T}}$ and $\tilde D_{1:t-1}(a_{1:t-1}) := \{\tilde D_1(a_1),...,\tilde D_{t-1}(a_{1:t-1})\}$ where $\tilde D_t(a_{1:t}) := \{a_t,Y_t(a_{1:t})\}$. Figure (ref) depicts the instrumental variables {\tt PS} in a SWIT.
Under this system definition, we have that feature $X_t$ is a valid instrumental variable for the time $t$ assignment conditional on $\tilde D_{1:t-1}$, so long as (sufficiently) the $U_t \perp \!\!\!\perp (V_t, W_t) \mid \tilde D_{1:t-1}$ and an instrument relevance condition holds. By further making a monotonicity assumption familiar from ImbensAngrist(94), letting $\mathcal{A}_t = \mathcal{X}_t = \{0,1\}$ for all $t$, and, for all $h$, defining $A_t(x_t):=\alpha_t(\tilde D_{1:t-1}, x_t,V_t)$, we can identify a local summary of the causal effect, \[ E[Y_{t,h}(1) - Y_{t,h}(0) \mid \mathbf{1}\{A_t(1) > A_t(0) \}=1, \tilde D_{1:t-1}]. \] The event $\{A_t(1) > A_t(0)\}$ can be thought of as the single unit in the time series “complying” with the instrument.
The first condition in Theorem (ref), in conjunction with the exclusion restrictions imposed by the IV {\tt PS}, grants that \[ X_t \perp \!\!\!\perp \big[A_t(1),A_t(0), Y_{t,h}(1), Y_{t,h}(0)\big] \mid \tilde D_{1:t-1}. \] The second and third conditions of Theorem (ref) mirror the relevance and monotonicity assumptions, respectively, introduced in ImbensAngrist(94).
The empirical setting represented by the IV {\tt PS} may be relevant to, e.g., a health system that wants to know the causal effect of ingesting a drug on a patient's health over time, but can only randomly encourage the patient to do so with a text reminder; or a ride-share application company that wants to understand the causal effect of augmenting some aspect of city-wide driver behavior on app engagement, but can only provide that city's drivers with randomized incentives to encourage desired behavior at scale.
The same primitives in the {\tt PS} can be used to define potential branches based on many consecutive periods of intervention. We may call these objects $s$-potential branches, and define them in the following alternative to assumption {\tt CP.2}. Naturally, the $0$-potential branch is a potential branch, and so potential branches are a special case of $s$-potential branches.
Under {\tt CP.2$^\prime$}, analyzing dynamic causal effects in the setting of Example (ref), we see that $D_{t,0}(a_{t:t+s}) = D_{t,0}(a_{t})$ and then recursively, for $h \leq s$, the \[ D_{t,h}(a_{t:t+s})=\left(
\right) =\left(
\right) D_{t,h-1}(a_{t:t+s})+\left(
\right) a_{t+h}+\left(
\right) \varepsilon _{t+h}, \] while for $h>s$, \[ D_{t,h}(a_{t:t+s}) = \phi D_{t,h-1}(a_{t:t+s}) + B \varepsilon_{t+h}. \] Then the dynamic causal effect is again non-stochastic. For $h \le s$, it is determined by the recursion $$ D_{t,h}(a_{t:t+s}) - D_{t,h}(a_{t:t+s}') = \left(
\right) \{D_{t,h-1}(a_{t:t+s}) - D_{t,h-1}(a_{t:t+s}')\} + \left(
\right) (a_{t+h} - a_{t+h}') $$ and then, for $h>s$, $$D_{t,h}(a_{t:t+s}) - D_{t,h}(a_{t:t+s}') = \phi \{D_{t,h-1}(a_{t:t+s}) - D_{t,h-1}(a_{t:t+s}')\}. $$
Design-based inference is extremely influential in randomized control trials and observational studies (e.g., Fisher(25),Fisher(35) and ImbensRubin(15)). In these settings, researchers choose to condition on the potential outcomes. The importance of the design-based approach in panel data is highlighted in, for example, ArkhangelskyImbens(24).
In time series this design-based approach was introduced by BojinovShephard(19) in their simpler setting with no features and randomized assignments. They condition on all the potential outcomes $$ Y_{1:T}(\mathcal{A}_{1:T}):= \{Y_{1:T}(a_{1:T}):a_{1:T}\in \mathcal{A}_{1:T}\}, $$ where, generically, we write $Z_{s:t}(A_{1:s-1},\mathcal{A}_{s:t}) = \{Z_{s:t}(A_{1:s-1},a_{s:t}):a_{s:t}\in \mathcal{A}_{s:t}\}$. One of the attractions of the design-based approach is that it allows some forms of causal inference without specifying a detailed model for the outcomes. This is very compelling, as time series has no direct form of replication. LinDing(25) further develop time series design-based studies, relating it to regression.
A key object of interest in design-based causal inference is the distribution of assignments conditional on the potential outcomes. This law is naturally called the “assignment mechanism,” taken from the cross-sectional literature.
Our goal is to sample from the {\tt AM} under a particular null hypothesis in order to perform design-based (causal) inference. Towards this goal, we state the following theorem.
Note that the joint law of $A_{1:T},Z_{1:T}(\mathcal{A}_{1:T})$ is determined by the marginal law of $Z_{1:T}(\mathcal{A}_{1:T})$ and the {\tt AM}. From Theorem (ref) and the time series prediction decomposition, the law of $Z_{1:T}(\mathcal{A}_{1:T})$ is entirely determined by the sequence $Z_t(\mathcal{A}_{1:t}) \mid Z_{1:t-1}(\mathcal{A}_{1:t-1})$, and thus the {\tt AM} is solely determined by the sequence $A_t \mid [A_{1:t-1},Z_{1:t}(A_{1:t-1},\mathcal{A}_t)]$. Notice this is close to the {\tt SAM}, but differs as {\tt AM} additionally conditions on $Y_t(A_{1:t-1},\mathcal{A}_t)$. This observation then motivates the following infeasible theorem. It uses the assumption that the features are {\tt PS}-exogenous, for all $t$ and $h$, which implies that $ X_{1:T}(a_{1:T}) = X_{1:T}$ for all $a_{1:T} \in \mathcal{A}_{1:T}.$
Now the equation ((ref)) has the same structure as the sequential assignment mechanism, but for the simulated path of the assignment. However, the law of ((ref)) is still infeasible to sample from as the $Y_{1:t-1}(A^*_{1:t-1})$ are counterfactuals we do not see: we only see the outcomes $Y_{1:t-1} = Y_{1:t-1}(A_{1:t-1})$.
To make sampling feasible, we define the null composite hypothesis $$ H_0: \ Y_{t}(a_{1:t}) = Y_{t}(a_{1:t}') + g_t(a_{1:t},a_{1:t}';\theta),\ \forall \ a_{1:t},a_{1:t}' \in \mathcal{A}_{1:t}, \ t=1,...,T, $$ where $g_t$ is a deterministic, known, parameterized function where $\theta \in \Theta$. Under the null, the $$ Y_{t}(A^*_{1:t}) = Y_{t} + g_t(A^*_{1:t},A_{1:t};\theta),\quad t=1,...,T. $$ Hence in the setting of Example (ref), with the features being {\tt PS}-exogenous, knowledge of the {\tt SAM} allows simulating $A^*_{1:T}$ from the {\tt AM} under the null.
The advantage of this design-based approach is that, for any value of $\theta$, we can simulate under the null $B$ copies of the triple $$ D_{1:T}^* = \{X_{1:T}^{\mathtt{T}},A^{*\mathtt{T}}_{1:T},Y_{1:T}(A_{1:T}^*)^{\mathtt{T}}\}^{\mathtt{T}}, $$ which can then be compared to the actual data $D_{1:T}.$ The comparison is made through a low dimensional statistic $$ T(D_{1:T}^*) $$ designed by the researcher, comparing it to $T(D_{1:T})$: rejecting the null if $T(D_{1:T})$ is large compared to the $B$ simulated versions of $T(D_{1:T}^*)$. Consequently we can find an exact confidence region $\mathcal{C}$ for $\theta$ (and so in turn causal effects) by inverting the distribution of the test, so that $ P(\theta \in \mathcal{C}) = 1-\alpha, $ where $\alpha\in (0,1)$ is selected by the researcher.
A significant virtue of this approach is that we are parametrically modeling the causal effects, but are entirely agnostic about the underlying dynamics of the outcomes. A downside is that it requires the ability to simulate from {\tt SAM} through ((ref)), the law of $$ A_t^* \mid [A_{1:t-1}^*,X_{1:t},\{Y_{1:t-1}=y_{1:t-1} + g_{1:t-1}(A_{1:t-1}^*,A_{1:t-1})\}]. $$ For experimental data this should be known. For observational data, the {\tt SAM} can be estimated from the data, using Assumption {\tt CP.2} to extrapolate from the assignments based on lagged observables (that is Assumption {\tt DGP.2}) to assignments based on lagged counterfactuals. In the important special case of the news impact {\tt PS} this difficulty entirely disappears as $A_t$ is an i.i.d. sequence.
There is a large literature on control, e.g., AndersonMoore(79), Whittle(81),Whittle(82),Whittle(96), Bertsekas(87) and HansenSargent(14). Here, the discussion will connect the control literature to the potential system, using the above notation, but with no features. \ Often the control literature is collected under the label of stochastic dynamic programming.
Define the sequence which would minimize expected future loss $J_t$ from taking actions $a_{t:T}\in \mathcal{A}_{t:T}$ as
give information at time $t-1$. Then take the time-t assignment as
ignoring the other $\widehat{a}_{t+1:T|t-1}$. Recall that the potential system's “consistency” means that
Here, the $A_{t}$ is previsible --- it is stochastic but entirely dependent on past outcome data. \ Thus the {\tt SAM} is probabilistically degenerate.
The time $t$ \textquotedblleft value\textquotedblright\ is defined as
the best possible expected future loss, given past data. In stochastic dynamic programming, the sequence $\left\{ \widehat{a}_{t|t-1}\right\} $ is typically called the control sequence, as it is assumed to be under the control of the researcher. Under the {\tt PS}, it determines the {\tt SAM}.
The causal effect of moving the assignment to $a_{t}$, away from the optimal value $a_{t}^{\prime }=A_{t}=\widehat{a}_{t|t-1}$, is particularly interesting. The immediate dynamic causal effect on the outcome is
which spills over to the effect at horizon $h=1$,
where $A_{t,1}(a_{t})=\widehat{a}_{t+1|t}(a_{t})$ with
being the optimal assignment path from time $t+1$ to time $T$ thinking we had seen the data $D_{1:t-1},D_{t,0}(a_{t})$.
In the control context, the researcher typically assumes losses are time separable, that is
where the $\ell_t= \{\ell_t(y,a):y \in \mathcal{Y}, a\in \mathcal{A}\}$ are known deterministic functions for all $t=1,...,T$, and in the final period for all $a,a' \in\mathcal{A}$, the $\ell_{T+1}(y,a)=\ell_{T+1}(y,a'):=\ell_{T+1}(y)$. Hence $L_{t}(a_{1:t})$ is random only because of the potential outcome, while $L_{t+h}(A_{1:t-1},a_{t:t+h})$ is random because of the assignment sequence and the potential outcome.
Looking forward from time $t$ to time $T$, the overall loss for a sequence of future assignments $a_{t:T}$ sums up future individual period losses. This can be written as a backward recursion
This recursive structure sets the stage for defining and solving Bellman equations.
This paper defines and works on the potential system ({\tt PS}), a foundational nonparametric model for studying how an assignment at time $t$ causally affects an outcome at time $t+h$, possibly in the presence of confounders. It yields familiar measures of causality --- explored through examples connected to various other literatures in time series causality --- that can be mapped to data-based predictions under familiar assumptions from cross-sectional causal inference. Because this foundational model is defined in terms of low-level nonparametric primitives, it can be readily extended to numerous other time series causality settings, such as design-based based inference, the study of more exotic causal effects, control, and beyond.
This paper does not discuss the intricate details of estimation nor inference --- our focus is on identification. PlagborgMullerKolesar(24) and BallinariWehrli(24) are population-based inference papers which can sit on top of our {\tt PS}, giving them causal meaning, covering inference for the relationship between what we call the branch potential outcomes and the assignments. This builds off inference results in RambachanShephard(21). Other recent work on non-linear impulse response functions include GoncalvesHerreraKillianPesavento(24),GoncalvesHerreraKillianPesavento(21) and GourierouxLee(23).
\baselineskip=12pt
\baselineskip=20pt