Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
85,303 characters · 17 sections · 51 citation commands
Estimating the Causal Effect of an Intervention in a Time Series Setting: the C-ARIMA Approach
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
{\it Keywords:} Business research, causal inference, econometrics, intervention analysis, potential outcomes, time series
\spacingset{1.45}
The potential outcomes approach is a framework that allows to define the causal effect of a treatment (or “intervention”) as a contrast of potential outcomes, to discuss assumptions enabling to identify such causal effects from available data, as well as to develop methods for estimating causal effects under these assumptions Rubin:1974,Rubin:1975,Rubin:1978,Imbens:Rubin:2015. Following Holland:1986, we refer to this framework as the Rubin Causal Model (RCM). Under the RCM, the causal effect of a treatment is defined as a contrast of potential outcomes, only one of which will be observed while the others will be missing and become counterfactuals once the treatment is assigned. For example, in a study investigating the impact of a new legislation to reduce carbon emissions, the incidence of lung cancer after the enforcement of the new law is the observed outcome, whereas the counterfactual outcome is the incidence that would have been observed if the legislation had not been enforced.
Having its roots in the context of randomized experiments, several methods have been developed to define and estimate causal effects under the RCM in the most diverse settings, including networks VanderWeele:2010, Forastiere:Airoldi:Mealli:2020, Noirjean:Mariani:Mattei:Mealli:2020, time series Robins:1986, Robins:Greenland:Hu:1999, Bojinov:Shephard:2019 and panel data Rambachan:Shephard:2019, Bojinov:Rambachan:Shephard:2020.
Focusing on a time series setting, a different approach extensively used in the econometric literature is intervention analysis, introduced by Box:Tiao:1975,Box:Tiao:1976 to assess the impact of shocks occurring on a time series. Since then, it has been successfully applied to estimate the effect of interventions in many fields, including economics Larcker:Gordon:1980, Balke:Fomby:1994, social science Bhattacharyya:Layton:1979,Murry:Stam:Lastovicka:1993 and terrorism Cauley:Iksoon:1988,Enders:Sandler:1993. The effect is generally estimated by fitting an ARIMA-type model with the addition of an intervention component whose structure should capture the effect generated on the series (e.g., level shift, slope change and similar). However, this approach fails to define the causal estimands and to discuss the assumptions enabling the attribution of the uncovered effect to the intervention.
Closing the gap between causal inference under the RCM and intervention analysis, in this paper we propose a novel approach, Causal-ARIMA (C-ARIMA), to estimate the causal effect of an intervention in observational time series settings where units receive a single persistent treatment over time. In particular, we introduce the assumptions needed to define and estimate the causal effect of an intervention under the RCM; we then define the causal estimands of interest and derive a methodology to perform inference. Laying its foundation in the potential outcomes framework, the proposed approach can be successfully used to estimate properly defined causal effects, whilst making use of ARIMA-type models that are widely employed in the intervention analysis literature.
We also test C-ARIMA with an extensive simulation study to explore its performance in uncovering causal effects as compared to a standard intervention analysis approach. The results indicate that C-ARIMA performs well when the true effect takes the form of a level shift; furthermore, it outperforms the standard approach in the estimation of irregular, time-varying effects.
Finally, we illustrate how the proposed approach can be conveniently applied to solve real inferential issues by estimating the causal effect of a permanent price reduction on supermarket sales. More specifically, on October 4, 2018 the Florence branch of an Italian supermarket chain introduced a new price policy that permanently lowered the price of 707 store brands. The main goal is to assess whether the new policy has influenced the sales of those products. In addition, we want to estimate the indirect effect of the permanent price reduction on the perfect substitutes, i.e., products sharing the same characteristics of the discounted goods but selling under a competitor brand. Our results suggest that store brands' sales increased due to the permanent price discount; interestingly, we find little evidence of a detrimental effect on competitor brands, suggesting that unobserved factors may drive competitor-brand sales more than price. We therefore believe that our approach can support decision making at firms, as it can be used to perform business research whose implications can be of great interest to marketing professionals.
The remainder of the paper is organized as follows: Section (ref) surveys the literature; Section (ref) presents the causal framework; Section (ref) illustrates the proposed C-ARIMA approach; Sections (ref) and (ref) describe, respectively, the simulations and the empirical study; Section (ref) concludes.
In time series settings, the identification and the estimation of causal effects using potential outcomes have been formalized in the context of randomized experiments Bojinov:Shephard:2019, Rambachan:Shephard:2019, Bojinov:Rambachan:Shephard:2020. However, unlike randomized experiments, in an observational study the assignment mechanism, i.e., the process that determines which units receive treatment and which receive control, is unknown. Thus, observational studies pose additional challenges to the identification and estimation of causal effects, especially in time series settings, due to the presence of a single series receiving the intervention: estimands usually employed in panel settings like the ATT (average treatment effect on the treated units) are not applicable and sometimes it might be difficult to find suitable control series.
A method that has been extensively used to evaluate the effect of interventions in the absence of experimental data is Difference-in-Difference (DiD) Card:Krueger:1993, Meyer:Viscusi:Durbin:1995, Garvey:Hanka:1999, Angrist:Pischke:2008, Anger:Kvasnicka:Siedler:2011. In its simplest formulation, this method requires to observe a treated and a control group at a single point in time before and after the intervention; the effect is then estimated by contrasting the change in the average outcome for the treated group with that of the control group under the assumption that, in the absence of treatment, the outcomes of the treated and control units would have followed parallel paths. However, many applications differ from this canonical setup: treatments may occur at different times, and the parallel trend assumption is often unrealistic Abadie:2005, Ryan:Burgess:Dimick:2015, ONeill:Kreif:Grieve:Sutton:Sekhon:2016.\footnote{For example, in case of time-varying unobserved confounders the parallel trend assumption is invalid. In such cases, researchers may switch to methods relying on the assumption of ignorability conditional on past outcomes and covariates, like the lagged dependent variable estimator (LDV). However, LDV and DiD estimators are related by a bracketing relationship Angrist:Pischke:2008, Ding:Li:2019, namely, if either parallel trends or ignorability holds, the true effect is bounded by the LDV and the DiD estimators, so when we rely on the wrong assumption we will either overestimate or underestimate the effect. In practice, when multiple pre-treatment periods are available, researchers test for parallel trends before employing DiD. Roth:2018 also suggests to adjust for the result of pretesting and Rambachan:Roth:2019 propose new methods that weakens the reliance on the parallel trend assumption by imposing restrictions on the differences in trends between treated and control units.}
Emerging literature on heterogeneous treatment effects in DiD with staggered adoption and variation in timing have partially overcome these limitations. For example, Callaway:Santanna:2020 introduce a weaker version of the parallel trend assumption that holds after conditioning on covariates; in addition, by focusing on a design-based perspective in which the assignment date is assumed to be randomized, Athey:Imbens:2021 do not need any functional form assumption for the potential outcomes. Furthermore, in a staggered adoption setting where the treatment time varies across units and they remain exposed to this treatment at all times afterwards, it is possible to estimate the effect of the intervention even when all units are eventually treated by using as a control group the set of last treated units Sun:Abraham:2020 or the set of not-yet treated units Callaway:Santanna:2020.
Another popular method to infer the causal effect of an intervention from panel data under the RCM is constructing a synthetic control from a set of time series that are not directly impacted by the treatment and have weighted pre-treatment variables matching those of the treated unit Abadie:Gardeazabal:2003, Abadie:Diamond:Hainmueller:2010, Abadie:Diamond:Hainmueller:2015. In contrast to DiD, synthetic control methods compensate for the lack of parallel trends by re-weighting control units so that the weighted pre-intervention outcomes and covariates are as close as possible to the average pre-intervention outcomes and covariates of the treated units. For example, in a study investigating the impact of a new legislation to reduce pollution levels, a suitable set of control series could be the evolution of carbon emissions in neighboring states that did not activate the new regulation: the synthetic control would then be constructed such that, in the pre-intervention period, the weighted average of the emissions and characteristics (e.g., population density, number of industries) of the neighboring states are similar to the emissions and characteristics of the treated state. Since their introduction, synthetic control methods have been successfully applied in a wide range of research areas, including healthcare Kreif:Grieve:Hangartner:Turner:Nikolova:Sutton:2016,Choirat:Dominici:Mealli:Papadogeorgou:Wasfy:Zigler:2018,Viviano:Bradic:2019, economics Billmeier:Nannicini:2013,Abadie:Diamond:Hainmueller:2015,Dube:Zipperer:2015,Gobillon:Magnac:2016,Benmichael:Feller:Rothstein:2018, marketing and online advertising Brodersen:Gallusser:Koehler:Remy:Scott:2015,Li:2019. Recently, Arkhangelsky:Athey:Hirshberg:Imbens:Wager:2019 introduced the synthetic difference in difference estimator (SDID) combining the attractive features of both DiD and synthetic controls: like synthetic controls, this method re-weights control units such that their weighted pre-intervention trend and the (average) trend of the treated unit(s) are parallel (but not necessarily identical), thereby weakening the reliance on the parallel trend assumption; then, it uses DiD on the re-weigted panel by also focusing on the time periods that are more similar to the post-intervention periods. Therefore, SDID also addresses the concerns on pre-trend adjustments expressed in Roth:2018.
Nevertheless, DiD estimators, synthetic control methods and their combinations usually require to observe at least one suitable control unit, which is often impractical. For example, in our application, appropriate control series could be the sales of products that are not impacted by the new policy. However, since the supermarket chain implemented an extensive price policy change addressing at least one product in each category, all products were impacted directly or indirectly by the intervention, preventing us to find suitable controls. In addition, all products received the intervention simultaneously, thereby precluding the adoption of the DiD estimators developed under variation in timing. Finally, by focusing on a few time points, DiD estimators and synthetic controls have a limited ability to exploit the information provided by pre-treatment temporal dynamics; thus, they are not the best option when there are few units observed over a large period of time (small N, large T panels).
A recent approach overcoming these limitations is proposed by Brodersen:Gallusser:Koehler:Remy:Scott:2015. Their methodology share several features with DiD and synthetic control methods but, instead of using control units or external characteristics, it only requires to learn the dynamics of the treated unit prior to the intervention. In other words, it builds a synthetic control by forecasting the counterfactual series in the absence of intervention based on a model estimated on the pre-intervention data. In particular, the authors employ Bayesian Structural Time Series models West:Harrison:2006, Harvey:1989 since they allow to add the components (e.g., trend, seasonality, cycle) that better describe the characteristics of the time series, whilst incorporating prior knowledge in the estimation process. Borrowing the name from the associated R package, from now on we refer to their method as “CausalImpact”. An extension of this approach to a multivariate time series setting is proposed in Menchetti:Bojinov:2020, where the authors employ Multivariate Bayesian Structural Time Series model to assess the impact of an intervention on statistical units showing interactions with one another.
Our work is closely related to DiD with staggered adoption, to synthetic control methods and to the methodology proposed by Brodersen:Gallusser:Koehler:Remy:Scott:2015. In the same vein as CausalImpact, we propose C-ARIMA as a novel approach to build a synthetic control for a time series subject to an intervention by learning its time dynamics in the pre-intervention period and then forecasting the series in the absence of intervention. Both CausalImpact and C-ARIMA brings several improvements over DiD estimators and synthetic control methods: they are tailored to estimate the effect of an intervention when no control series is available; and they are well suited to the case of a single time series or small N large T panels, since they allow to fully exploit useful information provided by the pre-intervention dynamics.\footnote{On the contrary, when there are multiple units observed over a short period of time, DiD estimators, synthetic control methods and SDID can be a better choice, since they allow to exploit the panel dimension; moreover, when there is variation in treatment timing it would be possible to construct estimators for the average treatment effect even in the absence of untreated units.} Furthermore, compared to CausalImpact, our methodology is based on ARIMA models and thus can be used as an alternative by researchers and practitioners in a wide range of fields that are not familiar to (or are not willing to adopt) Bayesian inference. In addition, ARIMA models are able to describe a wide variety of time series generated by complex, non-stationary processes and are already implemented in a large number of statistical software programs, which makes C-ARIMA very easy to use in practice.
Finally, the C-ARIMA also shares many features with the approach described in Box:Tiao:1976, where the authors suggest to compare the observed data after an intervention with the forecasts from a model fitted to the pre-intervention period. In particular, the $k$-step ahead forecast error in Box:Tiao:1976 is equivalent to our point causal effect $\tau_{t+k}$ and, as a result, our test statistic as defined by equation ((ref)) in Section (ref) relates to their Q test statistic. However, in addition to Box:Tiao:1976 we also provide test statistics for two additional effects: the cumulative and the temporal average effect. Indeed, oftentimes researchers are interested in a cumulative sum of point effects. For example, in a study evaluating the effect of the Hospital Readmission Reduction Program on hospital readmissions and mortality, Choirat:Dominici:Mealli:Papadogeorgou:Wasfy:Zigler:2018 focus on estimating the total number of additional readmissions due to the new law over the entire post-intervention period. Most importantly, we frame our estimators in the potential outcome framework, defining the effects and discussing the assumptions enabling their attribution to the intervention: both Box:Tiao:1976 and canonical intervention analysis (see Box:Tiao:1975 and Bhattacharyya:Layton:1979 among others) fail to do that and thus it not clear whether their effect might also have a causal interpretation or not. The distinction between C-ARIMA and canonical intervention analysis is even more apparent: in those setups, researchers need to make an assumption on the structural form of the effect (e.g., level shift); then, the resulting ARIMA model is estimated to the entire time series and the assumed structure is finally checked by analysing model adequacy. Essentially, this is a trial and error process: if model adequacy is poor, it needs to be re-estimated under a different assumed effect structure (e.g., slope change). Conversely, with C-ARIMA we do not need to impose any structure on the effect of the intervention.
On October 4, 2018 the Florence branch of an Italian supermarket chain introduced a new price policy that permanently lowered the price of $707$ store brands in several product categories; the empirical analysis focuses on the goods belonging to the “cookies” category. The supermarket chain also sells competitor brand cookies with the same characteristics (e.g., ingredients, flavour, shape) as their store brand equivalent; starting from the intervention date, these products became more expensive compared to the store brands. Therefore, we can consider all of them to be treated: the treatment on the store brand is the permanent price reduction, whereas the treatment on the competitor brands is the resulting relative price increase. The main goal is to assess the overall impact of the price policy, which is done by estimating the causal effect of the two treatments on the sales of store and competitor brand cookies.
In this section we present the notation and discuss some assumptions allowing the estimation of the causal effect and its attribution to the intervention. We also deduce some examples from our empirical context, so as to clarify both the theoretical concepts and the application. Finally, we define the causal estimands we are interested in.
Let $\operatorname{W}_{j,t} \in \{0,1\}$ be a random variable describing the treatment assignment of unit $j \in \{1,\dots,J\}$ at time $t \in \{1,\dots, T\}$, where $1$ denotes that a “treatment” (or “intervention”) has taken place and $0$ denotes control. The first assumption is about the treatment process.
In a randomized experiment, the treatment can be administered at any point in time Bojinov:Shephard:2019, whereas in observational studies it is not uncommon to observe a single persistent treatment, as in the case of a policy change Callaway:Santanna:2020, a new government law Choirat:Dominici:Mealli:Papadogeorgou:Wasfy:Zigler:2018 or a price promotion Brodersen:Gallusser:Koehler:Remy:Scott:2015. This is also the case of our empirical application and, as a result, we restrict our attention to a setting where all units are subject to a simultaneous and persistent intervention and we denote with $t^{*}$ the intervention date. Assumption (ref) is equivalent to the irreversibility of treatment assumption in Callaway:Santanna:2020; such a persistent treatment is also analogous to the absorbing treatment in Sun:Abraham:2020.
Whether a unit is assigned to treatment or control may impact its outcome. For example, a brand of cookies will likely sell more under a $50\%$ price discount compared to a scenario where it is not discounted. Under the RCM, the sales in these two alternative scenarios are known as potential outcomes.
Denote with $\operatorname{W}_{1:J,1:T} = (\operatorname{W}_{1,1:T}, \dots, \operatorname{W}_{J,1:T})$ the assignment paths of all units up to time $T$ and let $\operatorname{w}_{1:J,1:T}$ be a realization of $\operatorname{W}_{1:J,1:T}$. In general, the potential outcomes of a unit $j$ at time $t$ are function of the entire assignment panel, i.e. $\operatorname{Y}_{j,t}(\operatorname{w}_{1:J, 1:T})$ Bojinov:Rambachan:Shephard:2020. However, under Assumption (ref) we are able to restrict this dependence structure by focusing on non-anticipating potential outcomes.
In words, pre-intervention outcomes are not impacted by the future intervention; an implication is that there is no treatment effect in the pre-treatment period. Assumption (ref) is analogous to the non-anticipation assumptions usually made in the literature Bojinov:Shephard:2019, Callaway:Santanna:2020, Sun:Abraham:2020 and it is plausible when the statistical units have no knowledge of the future intervention. This is also the case of our empirical application, since the supermarket chain did not advertise the price reduction in advance. Furthermore, in our empirical setting, we can also rule out any form of interference between the units (i.e., cookies) in each group of store and competitor brands.
It means that whether the other units receive the treatment at time $t^*$ or not, this doesn't affect unit $j$'s potential outcomes. This assumption, which is also known as Temporal Stable Unit Treatment Value Assumption or TSUTVA Bojinov:Shephard:2019,Bojinov:Rambachan:Shephard:2020, is the time series equivalent of the cross-sectional SUTVA Rubin:1974 and contributes to reduce the number of potential outcomes. In our empirical setting, the cookies selected for the permanent price discount differ on many characteristics, such as the shape, flavor and the ingredients, meaning that they appeal to different customers. Therefore, the temporal no-interference assumption is plausible within each group of store and competitor brands. \\
Although playing a partially different role, above assumptions are crucial for a proper definition of the causal effect: Assumptions (ref) and (ref) pose a limit to the set of potential outcomes; Assumption (ref) excludes any treatment effect anticipation. Moreover, they allow us to ease notation. Indeed, a single persistent intervention implies $\operatorname{w}_{j,t} = \operatorname{w}_{j,t'} = \operatorname{w}_j$ for all $j$ and all $t,t' \in \{t^*+1, \dots, T \}$, whereas for all $t \in \{1, \dots, t^* \}$ we simply have $\operatorname{w}_j = 0$. Under temporal no-interference we can also drop the $j$ subscript, so from now on we use $\operatorname{Y}_{t}(\operatorname{w})$ to indicate the potential outcomes time series of a generic unit at time $t$.
Recall that there are two possible treatments for each unit in the post intervention period: $\operatorname{w} = 1$ if the unit is treated; $\operatorname{w} = 0$ if it is assigned to control. However, we can only observe one of them and in our empirical application, the observed path is $\operatorname{w} = 1$ for all units. Therefore, the observed potential outcome time series for all $t\in \{t^*+1, \dots, T \}$ is $\operatorname{Y}_{t}(1)$, whereas $\operatorname{Y}_{t}(0)$ is the missing or counterfactual potential outcome time series and needs to be estimated in order to compute the causal effect of the intervention. Including covariates that are linked to the outcome can improve the estimation of the missing potential outcome; however, if the covariates are influenced by the treatment, the estimates will be biased. We therefore make the following assumption.
As a result, we can use the known covariates values at time $t \in \{t^*+1, \dots, T \}$ to improve the prediction of the outcome in the absence of intervention $\operatorname{Y}_{t}(0)$. As detailed in Section (ref), our set of covariates includes a holiday dummy, some day-of-the-week dummies and the price per unit. While it is quite obvious that all the dummies are unaffected by the intervention, things get trickier for price. For the analysis on competitor brands we used their actual price, since it is not directly affected by the intervention; conversely, for the analysis on store brands, to avoid dependencies between treatment and price we considered a “modified” price that, starting from the intervention date, assumes a constant value equal to the price of day before the intervention, which is the most likely price that the item would have had if the discount hadn't been introduced.\footnote{The supermarket chain sometimes run temporary promotions reducing the price of selected goods for a limited period of time. The time interval after the permanent price discount spans from October 4, 2018 to April 30, 2019 and in the corresponding period before the intervention (October 4, 2017 - April 30, 2018) there were no temporary promotions on the store brands that are part of this analysis. Thus, the assumption of a constant price level in the period following the intervention is plausible.}
Our last assumption entails the assignment mechanism and it is essential to ensure that the uncovered causal effect can be attributed to the intervention.
In words, the decision of lowering the price of a product is informed only by its past sales performances and past covariates or, at most, by general beliefs on the sales evolution under active treatment. Instead, the assumption would be violated, for example, if prices were lowered to discourage the opening of a competing supermarket chain store in the neighborhood: in this case we would be uncertain if a positive effect on sales was due to the price reduction or to the deferred market entrance of the competitor. This assumption can not be tested, but based on the statements made by the supermarket chain, it is plausible in our empirical setting. Notice that Assumption (ref) is a re-statement of the non-anticipating treatment assumption of Bojinov:Shephard:2019 in a setting with a single intervention occurring at time $t^*$. Furthermore, a non-anticipating treatment in a time series setting is the analogous of the unconfounded assignment mechanism in the cross-sectional setting Imbens:Rubin:2015. Whilst a classical randomized experiment is unconfounded by design, we focus on observational studies where we have no control on the assignment mechanism. Thus, the non-anticipating treatment assumption is essential to ensure that any difference in the potential outcomes is due to the treatment.
We now introduce three related causal estimands: the point effect (an instantaneous effect at each point in time after the intervention), the cumulative effect (a partial sum of the point effects) and the temporal average effect (the average of the point effects in a given time period).
In words, the point effect measures the causal effect at a specific point in time and can be defined at every $t \in \{t^*+1, \dots, T\}$, thereby originating a vector of causal effects. The cumulative effect is then obtained by summing the point effects up to a predefined time point. For example, in our application, the cumulative effect would be the total number of cookies sold due to the permanent price reduction from the day when the new policy became effective until the end of the analysis period. Finally, we can also compute an average causal effect for the period. In our case, the temporal average effect indicates the number of cookies sold daily, on average, due to the permanent price reduction.
Notice that the point effect is analogous to the general causal effect defined in Bojinov:Shephard:2019, with the difference that our estimand is referred to a special setting were the units are subject to a single persistent treatment.
In the next section, we introduce the C-ARIMA model and we describe how it can be used to estimate the causal quantities of interest.
We propose a causal version of the widely used ARIMA model, which we indicate as C-ARIMA. We first introduce a simplified version of this model for stationary data generating processes; then, we relax the stationarity assumption and we extend the model to encompass seasonality and external regressors. Finally, we provide a theoretical comparison of the proposed approach with a standard ARIMA model with no causal connotation, REG-ARIMA henceforth.
We start with a simplified model for stationary data generating processes: this allows us to illustrate the building blocks of our approach with a clear and easy-to-follow notation. So, let us assume the potential outcome series $\{ \operatorname{Y}_t(\operatorname{w})\}$ evolving as
where: $c$ is a constant term; $\phi_p(L)$ and $\theta_q(L)$ are lag polynomials with $\phi_p(L)$ having all roots outside the unit circle; $\varepsilon_t$ is white noise with mean $0$ and variance $\sigma^2_{\varepsilon}$; $\tau_t = 0$ $\forall t \leq t^*$ and $\mathds{1}_{\{\operatorname{w} = 1\}}$ is an indicator function which is one if $\operatorname{w} = 1$. As a result, $\tau_t$ can be interpreted as the point causal effect at time $t \in \{t^*+1, \dots, T \}$, since it is defined as a contrast of potential outcomes,
$$\tau_t(\operatorname{w} = 1; \operatorname{w} = 0) \equiv \operatorname{Y}_t(\operatorname{w} = 1) - \operatorname{Y}_t(\operatorname{w} = 0) = \tau_t.$$
Notice that under Assumptions (ref)-(ref), $\tau_t$ is a properly defined causal effect in the RCM and, as such, it should not be confused with additive outliers or any other kind of intervention component typically used in the econometric literature (e.g., Box:Tiao:1975, Chen:Liu:1993). Indeed, we can show that Equation ((ref)) encompasses all types of interventions. For example, consider the following model specification (innovation-type effect),
$$\operatorname{Y}_t(\operatorname{w}) = c + \frac{\theta_q(L) }{\phi_p(L)}(\varepsilon_t + \tau_t\mathds{1}_{\{\operatorname{w} = 1\}})$$
and define $\tau_t^{\dagger} = \frac{\theta_q(L) }{\phi_p(L)} \tau_t$. Then, we have
$$\operatorname{Y}_t(\operatorname{w}) = c + \frac{\theta_q(L) }{\phi_p(L)}\varepsilon_t + \tau_t^{\dagger}\mathds{1}_{\{\operatorname{w} = 1\}}$$
where $\tau_t^{\dagger}(1;0) = \operatorname{Y}_t(1) - \operatorname{Y}_t(0) = \tau_t^{\dagger}$ is the point causal effect at time $t$. As it will be clear in Section (ref), our model is estimated on the pre-intervention data, thus in the C-ARIMA approach we do not need to find the structure that better represents the effect of the intervention (e.g., additive outlier, transient change, innovation outlier). Conversely, such effect emerges as a contrast of potential outcomes in the post-intervention period and the proposed approach allows us to estimate $\tau_t$ whatever structure it has.
To improve readability of the model equations, from now on we use $\operatorname{Y}_t$ to indicate $\operatorname{Y}_t(\operatorname{w})$; the usual notation is resumed in Section (ref). Thus, Equation ((ref)) can be written as,
Setting
Equation ((ref)) becomes,
Assuming perfect knowledge of the parameters ruling $\{z_{t}\}$, indicating with $\mathcal{I}_{t^*}$ the information up to time $t^*$ and denoting with $H_{0}$ the situation where the intervention has no effect (namely, $\tau_t = 0$ for all $t > t^*$) we have that for a positive integer $k$, the $k$-step ahead forecast of $\operatorname{Y}_t$ under $H_{0}$ conditionally on $\mathcal{I}_{t^*}$ is
where $\hat{z}_{t^* + k|t^*} = E[z_{t^*+k} | \mathcal{I}_{t^*}, H_{0}]$. Thus, $\hat{z}_{t^* + k|t^*}$ represents an estimate of the missing potential outcomes in the absence of intervention, i.e., $\hat{\operatorname{Y}}_{t^*+k}(0) = \hat{z}_{t^* + k|t^*}$.
We now generalize the above framework to a setting where $\{ \operatorname{Y}_t \}$ is non-stationary and possibly includes seasonality as well as external regressors.
Let $\{ \operatorname{Y}_t \}$ follow a regression model with ARIMA errors and the addition of the point effect $\tau_t$,
where $\Theta_Q (L^s)$, $\Phi_P(L^s)$ are the lag polynomials of the seasonal part of the model with $\Phi_P(L^s)$ having roots all outside the unit circle; $\operatorname{X}_t$ is a set of external regressors satisfying Assumption (ref); $(1-L^s)^D$ and $(1-L)^d$ are contributions of the differencing operators to ensure stationarity, and $s$ is the seasonal period. Notice that the intercept defined in model ((ref)) is now included in the set of regressors. To ease notation, defining
and indicating with $T(\cdot)$ the transformation of $\operatorname{Y}_t$ needed to achieve stationarity, i.e. $T(\operatorname{Y}_t) =(1-L^s)^D (1-L)^d \operatorname{Y}_t $, model ((ref)) becomes
where $T(\operatorname{X}_t)' = (1-L^s)^D (1-L)^d \operatorname{X}_t'$ indicates that the same transformation is applied to the vector of regressors. Thus, the $k$-step ahead forecast of $S_t$ under $H_0$, given the information up to time $t^*$ is
$$\hat{S}_{t^*+k} = \operatorname{E}[S_{t^*+k} | \mathcal{I}_{t^*}, H_0] = \operatorname{E}[T(\operatorname{Y}_{t^*+k}) - T(\operatorname{X}_{t^*+k})' \beta | \mathcal{I}_{t^*}, H_0] = \hat{z}_{t^*+k|t^*} .$$
We now derive estimators for the causal effects defined in Section (ref) based on the C-ARIMA model and we discuss their properties.
Inference on the point, cumulative and temporal average effects can be conducted using hypothesis tests based on the above estimators. The following theorem illustrates their distributional properties.
Theorem (ref) implies that the testing procedure can be lead in two different ways. If one safely relies on the normality of the error term, the test can be based on Equations ((ref))--((ref)) and the corresponding Gaussian quantiles. Otherwise, one can compute empirical critical values from Equations ((ref))--((ref)) by bootstrapping the errors from the model residuals. In both cases, the needed parameters are replaced by their estimated counterpart. The implicit assumption is that the model can be misspecified in the real practice, but the possible misspecification is not so severe to jeopardize the consistency of the estimators of the model parameters and, in turn, of the $\psi_i$ estimators.
So far, we derived estimators for three causal effects defined for the transformed variable $S_t = T(\operatorname{Y}_t)- T(\operatorname{X}_t)'\beta$, where $T(\cdot) = (1-L^s)^D (1-L)^d$ is a transformation to achieve stationarity. By doing a further step we can also estimate the effect for the original (untransformed) variable $\operatorname{Y}_t$. Indeed, model ((ref)) can also be written as,
where
is the point causal effect on the original variable, defined by a contrast of two potential outcomes,
$$\tau^Y_t(\operatorname{w} = 1; \operatorname{w} = 0) \equiv \operatorname{Y}_t(\operatorname{w} = 1) - \operatorname{Y}_t(\operatorname{w} = 0).$$
Again, we can define the estimators for the point, cumulative and temporal average effects on the original variable under the null hypothesis that the intervention has no effect, $H_0: \tau^Y_t (1;0) = 0$ for all $t \in \{t^*+1, \dots, T \}$.
To perform inference on ((ref)) we introduce Theorem (ref), whose derivation follows directly from Theorem (ref).
Summarizing, in order to estimate the causal effects ((ref)), ((ref)) and ((ref)) we need to follow a three-step process: i) estimate the ARIMA model only in the pre-intervention period, so as to learn the dynamics of the dependent variable and the links with the covariates without being influenced by the treatment; ii) based on the process learned in the pre-intervention period, perform a prediction step and obtain an estimate of the counterfactual outcome during the post-intervention period in the absence of intervention; iii) by comparing the observations with the corresponding forecasts at any time point after the intervention, evaluate the resulting differences, which represent the estimated point causal effects. To perform inference, we can then use the results presented in Theorem (ref) and Theorem (ref).
Conversely, REG-ARIMA model is fitted to the full time series (pre- and post-intervention) and the estimated coefficient for the dummy variable $D_t$ gives a measure of the association between the intervention (in the form of a level shift) and the outcome.
To measure the effect of an intervention on an outcome repeatedly observed over time, a widely used approach is fitting a linear regression with ARIMA errors (REG-ARIMA). This method uses the entire time series and a dummy variable activating after the intervention; then, SARIMA-type errors are added to the model to account for autocorrelation and possible seasonality. In its simplest formulation, such a model can be written as,
where $z_t$ is a stationary ARMA$(p,q)$; $D_t$ is a dummy variable taking value $1$ after the intervention and $0$ otherwise and $\beta_0$ is the regression coefficient. Generalizing to a possibly seasonal and non-stationary ARIMA, above model can be re-written as,
where $\operatorname{X}_t$ is a set of regressors, including the intercept and the dummy variable and $\beta$ is a vector of regression coefficients.\footnote{Notice that under REG-ARIMA, $\operatorname{Y}_t$ is not a potential outcome, therefore, it does not depend on the treatment path.}
Essentially, REG-ARIMA is a standard intervention analysis approach that is used when the intervention is supposed to have produced a level shift on the outcome, i.e. a fixed change in the level of the outcome during the post-intervention period. Thus, there are two main differences between C-ARIMA and REG-ARIMA. First, without a critical discussion of the fundamental assumptions, the effect grasped by $\beta_0$ can not be attributed with certainty to the intervention. For example, it might be driven by an undetected confounder, biased by the inclusion of a regressor linked to the treatment, or even be the anticipated result of a future intervention. Second, the size of the effect is given by the estimated coefficient of a dummy variable activating after the intervention, so that REG-ARIMA can only capture effects in the form of level shifts. Conversely, C-ARIMA assumes no structure on $\tau_t$ and, as such, it can capture any form of effects (level shift, slope change and even irregular time-varying effects). Furthermore, the estimation of the effect is done in a very natural way under C-ARIMA. Indeed, intervention analysis requires the estimation of two models: the first learns the structure of the effect and the second measures its size; by letting the intervention component free to vary, C-ARIMA can instead estimate any form of effect in only one step.
In Section (ref) we report a simulation study where we compare the empirical performance of both approaches (C-ARIMA and REG-ARIMA) in inferring causal effects.
We perform a simulation study to check the ability of the C-ARIMA approach to uncover causal effects. Furthermore, in order to show its merits over a more standard approach, we also assess the performance of REG-ARIMA. We remark, however, that the comparison is purely methodological, since the theoretical limitations of REG-ARIMA do not allow the attribution of such effects to the intervention. Sections (ref) and (ref) illustrate, respectively, the simulations design and the results.
We generate $1000$ replications from the following ARIMA$(1,0,1)(1,0,1)_7$ model,\footnote{Notice that $D=d=0$ implies $\tau_t \equiv \tau^Y_t$.}
The two covariates of the regression equation are generated as $\operatorname{X}_{1,t} = \alpha_1t + u_{1,t}$ and $\operatorname{X}_{2,t} = \sin (\alpha_2t) + u_{2,t}$, with $\alpha_1 = \alpha_2 = 0.01$, $u_{1,t} \sim N(0,0.02)$, $u_{2,t} \sim N(0,0.5)$ and coefficients $\beta_1 = 0.7$ and $\beta_2 = 2$, respectively; regarding the ARIMA parameters, they are set to $\phi_1 = 0.7$, $\Phi_1 = 0.6$, $\theta_1 = 0.6$ and $\Theta_1 = 0.5$. Finally, $\varepsilon_t \sim N(0, \sigma)$ with $\sigma = 5$. Figure (ref) shows the evolution of the generated covariates and their linear combination according to the above model.
We assume that each generated time series starts on January 1, 2017 and ends on December 31, 2019 and that a fictional intervention takes place on June 30, 2019. In particular, we tested two types of intervention: i) a level shift with $5$ different magnitudes, i.e., $+1 \%$, $+10 \%$, $+25 \%$, $+50 \%$, $+100 \%$ ; ii) an intervention producing an immediate shock of $+10\%$ followed by a steady increase up to $+40\%$, a regular decline afterwards and a second increase near the end of the analysis period. As an example, Figure (ref) provides a graphical representation of the two interventions for one of the simulated time series.
The estimation of the causal effect is performed under two different models: the proposed C-ARIMA approach and REG-ARIMA, i.e, a linear regression with ARIMA errors and the addition of a dummy variable, as in Equation ((ref)). Recall from previous Section (ref) that the C-ARIMA approach requires that the model is estimated on the pre-intervention data and the effect is given by direct comparison of the observed series and the corresponding forecasts post-intervention. Conversely, REG-ARIMA is fitted on the full time series and the estimated coefficient of the dummy variable provides a measure of the impact of the intervention. In addition, we estimate two versions of each model: a correctly specified model, denoted respectively with C-ARIMA$^{TRUE}$ and REG-ARIMA$^{TRUE}$, and the best fitting model selected by BIC minimization, denoted with C-ARIMA$^{BIC}$ and REG-ARIMA$^{BIC}$. Finally, in order to evaluate the performance of both approaches in uncovering causal effects at longer time horizons, we perform predictions at $1$ month, $3$ months and $6$ months from the intervention. As a result, the total number of estimated models in the pre-intervention period is $4000$ and the total number of estimated causal effects is $72,000$ (one for each time series, model, tested intervention and time horizon).
We measure the performance of the four models in terms of the following indicators:
Table (ref) shows the simulation results in terms of the length of the $95 \%$ confidence intervals around $\tau_t(1;0)$ and $\beta_0$, respectively. As expected, for the C-ARIMA models the interval length is independent of the impact; we can also notice that it reduces as the time horizon increases, whereas the interval length estimated under REG-ARIMA is stable over time. Finally, we can observe that REG-ARIMA yields shorter confidence intervals than C-ARIMA.
Table (ref) reports the absolute percentage errors resulting from the simulations. When the intervention takes the form of a level shift, the error decreases with the size of the effect and, unsurprisingly, REG-ARIMA yields slightly better results than C-ARIMA. Indeed, the former model is especially suited for interventions in the form of level shifts. However, when we consider a level shift $> +50\%$ or an irregular intervention, the estimation errors of REG-ARIMA are $2$ to $4$ times higher than those coming from C-ARIMA.
The interval coverage is reported in Table (ref). Again, the coverage of the C-ARIMA approach does not change with the impact size and it is very close to the nominal $95 \%$ level. Instead, the coverage of REG-ARIMA decreases with the impact size and, with the only exception of the first two impacts, the results are quite far from the nominal $95\%$ level. This can be explained by the short confidence intervals achieved by REG-ARIMA, suggesting that even though the estimation error is small, the confidence intervals are not wide enough to contain the true effect. More importantly, when the effect is irregular, the estimated confidence intervals never contain the true effect.
Concluding, REG-ARIMA approach fails to detect irregular interventions and, most of the times, it does not achieve the desired interval coverage. As expected, REG-ARIMA model is suited only when there is reason to believe that the intervention produced a fixed change in the outcome level. Otherwise, should the researcher fail to identify the structure of the effect, using REG-ARIMA on irregular patterns produces biased estimates. Conversely, the C-ARIMA approach does a reasonably good job in detecting both type of interventions. Moreover, C-ARIMA does not require an investigation of the effect type prior to the estimation step; in addition, when the intervention is in the form of a level shift, the reliability of the C-ARIMA estimates increases with the impact size. Finally, we can observe that the results of the models selected through BIC minimization are very similar to those of the correct model specifications (the BIC correctly identifies the true model $74\%$ of times).
In this section we describe the results of our empirical application; the goal is to estimate the impact of the permanent price reduction performed by an Italian supermarket chain.
Data consists of daily sales counts of $11$ store brands and their corresponding competitor brand cookies in the period September 1, 2017, April 30, 2019.\footnote{We excluded the last competitor brand because $62 \%$ of observations were missing. Thus, we analyzed $11$ store and $10$ competitor brands.} The permanent price reduction on the store brand cookies was introduced by the supermarket chain on October 4, 2018. As an example, Figure (ref) shows the time series of units sold, the evolution of price per unit and the autocorrelation function of one store brand and its direct competitor. The plots for the remaining store-brand and competitor-brand cookies are provided in Appendix (ref). The occasional price drops before the intervention date indicate temporary promotions run regularly by the supermarket chain. The products exhibit a clear weekly seasonal pattern, evidenced by the spikes in the autocorrelation functions. In the panel referred to the direct competitor brand, we can also observe the evolution of the relative price per unit (the ratio between the prices of the competitor brand and the corresponding store brand). Unsurprisingly, despite the occasional drops due to temporary promotions, the price of the competitor brand relative to the corresponding store brand has increased after the intervention.
To estimate the causal effect of the permanent price discount on the sales of store-brands cookies, we follow the approach outlined in Section (ref). In particular, under Assumption (ref), we analyze each cookie separately, thereby fitting $11$ independent models. In order to improve model diagnostics, the dependent variable is the natural log of the daily sales count. This also means that we are postulating the existence of a multiplicative effect of the new price policy on the sales of cookies. Since in terms of the original variable the cumulative sum of daily effects is equivalent to their product, we focused our attention on estimating the temporal average causal effect, which can still be interpreted as an average multiplicative effect. Furthermore, we included covariates to improve prediction of the missing potential outcomes in the absence of intervention. In particular, to take care of the seasonality we included six dummy variables corresponding to the day of the week and one dummy denoting December Sundays.\footnote{In principle, we may also have a monthly seasonal pattern on top of the weekly cycle but the reduced length of the pre-intervention time series ($398$ observations) does not allow us to assess whether a double seasonality is present.} Indeed, the policy of the supermarket chain implies that all shops are closed on Sunday afternoon except during Christmas holidays. Thus, we may have two opposite “Sunday effects”: a positive effect in December, when the shops are open all the day; a negative effect during the rest of the year, since all shops are closed in the afternoon. We also included a holiday dummy taking value $1$ before and after a national holiday and $0$ otherwise. This is to account for consumers' tendency to increase purchases before and after a closure day.\footnote{To be precise, on the day of a national holiday we have a missing value (so there is no holiday effect), whereas the dummy variable should capture the effect of additional purchases before and after the closure day(s).} Finally, we included a modified version of the unit price, that after the intervention day and during all the post-period is taken equal to the last price before the permanent discount. As explained in discussing Assumption (ref), this is the most likely price that the unit would have had in the absence of intervention. In addition, to estimate the average causal effect of the intervention on store brands, we are also interested in evaluating how this effect evolves with time. Thus, we repeated the analysis by making predictions at three different time horizons: $1$ month, $3$ months and $6$ months after the intervention.
The same methodology is applied to the competitor brands, with a slight modification on the set of covariates. Indeed, this time the unit price is not directly influenced by the intervention, which instead affects the relative price (as shown in Figure (ref)); so, to forecast competitor sales in the absence of intervention we directly used the actual price.
Again, to illustrate the merits of the proposed approach, the results obtained from C-ARIMA are then compared to those of REG-ARIMA, as described by Equation ((ref)). More specifically, we fitted independent linear regressions with ARIMA errors for each of the $11$ store brands and their competitors.
Table (ref) shows the results of the C-ARIMA and the REG-ARIMA approaches applied to the store brands.\footnote{For the C-ARIMA approach, Table (ref) shows that the results with the empirical critical values computed from the bootstrapped errors are in line with those reported in Table (ref).} Figure (ref) illustrates the causal effect, the observed time series and the forecasted series in the absence of intervention for one selected item.\footnote{The same plots for the remaining store-brand and competitor-brand cookies are provided, respectively, in Appendix (ref) and Appendix (ref).} At the $1$-month time horizon, the causal effect is significantly positive for $8$ out of $11$ items; three months after the intervention, the causal effect is significantly positive for $10$ items; after six months, the effect is significant and positive for all items. Conversely, REG-ARIMA fails to detect some of the effects: compared to the C-ARIMA results, the effect on items $4$ and $11$ at the first time horizon, on item $11$ at the second horizon and on items $5$ and $11$ at the third horizon are not significant.
Table (ref) reports the results for the competitor brands and Figure (ref) plots the causal effect, the observed series and the forecasted series for one selected item.\footnote{For the C-ARIMA approach, Table (ref) shows that the results with the empirical critical values computed from the bootstrapped errors are in line with those reported in Table (ref).} Again, the causal effect seems to strengthen as we proceed far away from the intervention. At $1$-month horizon no significant effect is observed; three months after the intervention, we find a significant and negative effect on item $10$; at $6$-month horizon we find significant negative effects on items $8$ and $10$ and a significant positive effect on item $5$. A negative effect suggests that following the permanent price discount, consumers have changed their behavior by privileging the cheaper store brand. Instead, a positive effect might indicate that the price policy has determined an increase in the customer base, i.e. new clients have entered the shop and eventually bought the items at full price. Again, REG-ARIMA model leads to partially different results: at $6$-month horizon, a positive effect is found on item $6$ and no effect is detected on item $8$.
Summarizing, the intervention seems to have produced a significant and positive effect on the sales of store brand cookies. Conversely, we do not find considerable evidence of a detrimental effect on competitor cookies (the only exceptions being items $8$ and $10$). This indicates that, even though each store-competitor pair is formed by perfect substitutes, price might not be the only factor driving sales. For example, unobserved factors such as individual preferences or brand faithfulness may have a role as well.
We propose a novel approach, C-ARIMA, to estimate the effect of interventions in a time series setting under the Rubin Causal Model. After a detailed illustration of the assumptions underneath our causal framework, we defined three causal estimands of interest, i.e., the point, cumulative and average causal effects. Then, we introduced a methodology to perform inference.
To measure the performance of C-ARIMA in uncovering causal effects, we presented a simulation study showing that this approach performs well in comparison with a standard intervention analysis approach (REG-ARIMA) where the true effect is in the form of a level shift; it also outperforms the latter in case of irregular, time-varying effects.
Finally, we applied the proposed methodology to estimate the causal effect of the new price policy introduced by a big supermarket chain in Italy, which addressed a selected subset of store-brands products by permanently lowering their price. The empirical analysis was carried out on the goods belonging to the “cookies” category: the results show that the permanent price reduction was effective in increasing store-brand cookies' sales. Little evidence of a detrimental effect on the corresponding competitor-brand cookies is found.