EconBase
← Back to paper

Identification and Inference for Synthetic Control Methods with Spillover Effects: Estimating the Economic Cost of the Sudan Split

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

67,532 characters · 26 sections · 70 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Bayesian Inference for Synthetic Control Methods with Spillover Effects

abstractThe synthetic control method (SCM) is widely used for causal inference with panel data, particularly when the number of treated units is small. It relies on the stable unit treatment value assumption (SUTVA), ruling out spillover effects. However, interventions often affect not only treated but also untreated units. This study proposes a novel panel data method that extends standard SCM to account for spillovers and estimate both treatment and spillover effects. The approach extends the SCM framework by incorporating a spatial autoregressive (SAR) panel data model that captures spillover patterns across units. We also develop a Bayesian inference procedure using horseshoe priors for regularization. We apply the proposed method to two empirical studies: (i) evaluating the effect of the California tobacco tax on cigarette consumption, and (ii) assessing the economic impact of the 2011 Sudan division on GDP per capita. \\ Keywords: Bayesian inference; national breakup; spatial autoregressive model; spatial panel data; SUTVA

\setstretch{1.4}

Introduction

The synthetic control method (SCM) abadie2003economic,abadie2010synthetic is a causal inference approach used to estimate treatment effects when the number of units undergoing intervention is very small, such as a single unit. This approach identifies treatment effects from panel data by substituting the counterfactual outcomes of the treated unit in the absence of treatment with weighted averages of outcomes of untreated units. The method has been widely used in various fields, including political economy and marketing. athey2017state describe SCM as “arguably the most important innovation in the policy evaluation literature in the last 15 years.”

The SCM is typically applied to country- or district-level data. For instance, Abadie_et_al_2015 apply SCM to country-level data to estimate the economic impact of the 1990 German reunification, while bifulco2017using apply it to school district data to evaluate the effect of an educational program. In such contexts, interventions may affect untreated units through geographic or socioeconomic connections among countries or districts, which are known as spillovers. However, SCM relies on the stable unit treatment value assumption (SUTVA) rubin1978bayesian, which posits that one unit's outcome is unaffected by the treatment status of other units, excluding spillovers. If spillovers occur to untreated units, SUTVA is violated, which can bias treatment effect estimation by the SCM.

This paper tackles this problem by extending the standard SCM to allow for spillover effects. We leverage the spatial autoregressive (SAR) model, a workhorse model in spatial data analysis, to address this issue with its capability to capture dependence of outcomes among units. We propose a novel approach to identify and estimate treatment effects by incorporating an SAR panel data model into SCM, where we characterize the outcomes of untreated units with the SAR model. This approach allows for spillover effects arising from the dependence of outcomes among treated and untreated units, thereby relaxing SUTVA in SCM. Furthermore, the extended SCM can identify and estimate spillover effects on untreated units, which are also parameters of interest in many empirical studies adopting SCM.

Building on the identification results, we propose a Bayesian inference method for the proposed SCM. Data targeted by SCM often involves panel data, but the number of units or length of pretreatment periods is not always sufficiently large. In such cases, frequentist approaches may not provide accurate inference. Bayesian inference offers several advantages. First, it can yield more accurate estimates than frequentist methods when the sample size is small. Second, it facilitates statistical inference via the Markov Chain Monte Carlo (MCMC) procedure. Third, Bayesian modeling is flexible, allowing models to be less dependent on the assumptions about prior distributions.

We apply Bayesian regularization in the construction of synthetic controls. Following kim2020bayesian, we use Bayesian horseshoe priors carvalho2010horseshoe to model the synthetic weights. The Bayesian horseshoe prior has strong regularization effects, making it effective in avoiding overfitting bias. This is particularly important in SCM, as it is often applied to data with a relatively large number of control units and short pretreatment periods, which can easily cause overfitting bias in the estimation of synthetic control outcomes. The simulation study in the paper examines the finite sample performance of the proposed Bayesian inference.

This paper presents two substantive empirical studies applying the proposed SCM. The first revisits abadie2010synthetic to evaluate the impact of California’s tobacco tax on cigarette consumption, and the second estimates the effect of Sudan’s north–south split in 2011 on GDP per capita. Both studies involve spillover effects among US states and African countries, respectively. Regarding California's tobacco tax, our results show a negative impact on tobacco consumption in California, supporting the findings of abadie2010synthetic. We also find evidence of spillover effects, showing that the tax reduced consumption in other states.

To the best of our knowledge, the second empirical study is the first to examine the economic impact of Sudan's north-south split on the Sudans (the region of the former united Sudan). In a related study, mawejje2021economic estimate the economic impact of the 2012 oil production halt in South Sudan on South Sudan itself using the standard SCM. They intentionally exclude countries neighboring South Sudan from the control group to address spillover concerns. This is a common practice in dealing with spillover in SCM abadie2021using; however, excluding untreated units, particularly neighboring countries, can lead to poor fitting of synthetic controls and cause additional bias. The SCM proposed in this paper does not need to exclude any untreated units. Our empirical results show that Sudan's north-south split led to a 9.5% decline in GDP per capita in the Sudans, with a cumulative reduction of 34% between 2011 and 2015. We also find evidence of negative spillover effects on other African countries with strong economic ties to Sudan, such as Egypt and Kenya.

Related Literature

Since abadie2003economic and abadie2010synthetic pioneered SCM, along with its broad application in the social sciences, SCM has been theoretically and methodologically advanced by many works. abadie2021penalized present a penalized synthetic control method to reduce interpolation biases. arkhangelsky2021synthetic combine SCM with difference-in-differences (DID) and propose a new estimator with robustness properties. ben2021augmented propose an augmented synthetic control method to correct bias resulting from imperfect pre-treatment fit. li2020statistical and chernozhukov2021exact propose inference methods for SCM. ferman_pinto_2021 point out a potential bias in SCM when the perfect fit assumption is not satisfied.

Several works have proposed Bayesian estimation in SCM to facilitate statistical inference through MCMC sampling and/or to apply flexible modeling. kim2020bayesian introduce a Bayesian synthetic control method that uses the Bayesian horseshoe prior proposed by carvalho2010horseshoe for the synthetic weights. Their simulations demonstrate more accurate predictions of counterfactual outcomes than the standard SCM, particularly when the number of units is relatively large compared to the length of the pretreatment periods. brodersen2015inferring propose a method based on Bayesian structural time-series models. pang2022bayesian and klinenberg2023synthetic propose approaches to correct bias in SCM induced by the dynamic characteristics of untreated units. pang2022bayesian propose a Bayesian posterior predictive approach with a latent factor term that incorporates unit-specific time trends. klinenberg2023synthetic incorporates time-varying parameters by using a state space framework.

While the related literature mentioned so far assumes SUTVA in SCM, few studies have considered its relaxation. cao2019estimation allow for spillover effects in SCM by assuming a specified structure of spillover effects. Our approach differs from theirs in two aspects: (i) our approach supposes a spatial correlation structure for the outcomes of units rather than a direct structure for spillover effects, and (ii) our approach supposes the existence of a synthetic control only for the treated unit, whereas their approach supposes synthetic controls for both the treated and untreated units. menchetti2022estimating assume that a unit's potential outcome depends on its own and neighboring treatment assignments. They propose a Bayesian structural time series method to identify and estimate treatment and spillover effects by estimating the predictive distribution of counterfactual outcomes. grossi2020synthetic focus on the average spillover effects within specified clusters and provide a method for constructing confidence intervals for treatment and average spillover effects. In the context of DID, miguel2004worms propose an approach for estimating spillover effects using RCT data, which utilizes non-compliance in treatment groups to identify spillover effects.

This work also relates and contributes to the economic analysis of national breakups alesina2000economic,bolton1997breakup through its original empirical study. The economic consequences of major political reconfigurations, such as national breakups or unifications, have long been the subject of rigorous empirical and theoretical research. For example, redding2008costs exploit the division and reunification of Germany as a natural experiment to quantify the impact of border changes on the economic development of border regions. Seminal works by alesina2000economic and bolton1997breakup explore the interplay between economic integration, political economy conflicts, and the formation or dissolution of nations, providing foundational frameworks for understanding such large-scale transformations. This paper contributes to this literature by empirically demonstrating the macroeconomic impacts of national division, using unique data on Sudan's split and a novel SCM approach that accounts for spillover effects arising from the breakup.

Structure of the Paper

The remainder of the paper proceeds as follows. Section (ref) describes the SCM setup with spillover effects. Section (ref) introduces the SAR panel data model and presents the identification results. Section (ref) proposes a Bayesian SCM procedure. Section (ref) shows simulation results to demonstrate the finite-sample performance of the proposed SCM. Section (ref) presents the results of two empirical applications. Section (ref) concludes the paper. The supplementary appendix provides additional details of the Bayesian SCM procedure.

Setup

Consider units $i=0,1,\cdots,N$ and time points $t=1,2,\cdots,T$. We suppose that $i=0$ is the unit receiving the treatment, and $i=1,\cdots,N$ are the units in the control group. Let $T_{0}$ be the number of pretreatment periods with $1 \leq T_0 < T$. The treatment status of a unit $i$ at time $t$ is denoted as $d_{it} \in \{0,1\}$, where $d_{it}=1$ means that unit $i$ receives treatment at time $t$, and $d_{it}=0$ means otherwise. In our context, $d_{it}=1$ when $i=0$ and $t > T_0$ and $d_{it}=0$ otherwise. We define $\bm{d}_{t} \equiv (d_{0t}, d_{1t},\cdots,d_{Nt})\in\{0,1\}^{N+1}$, the treatment status vector for time $t$.

We consider the potential outcome framework. While many studies assume that one's potential outcome is not influenced by the treatment status of others (SUTVA) rubin1978bayesian, this study allows for the potential outcome to be affected by the treatment status of others. Thus, for treatment status $\bm{d}_{t}$ at time $t$, the potential outcome of unit $i$ is denoted as $Y_{it}(\bm{d}_{t})$.\footnote{SUTVA means that the potential outcome of each unit $i$ is represented by $Y_{it}(d_{it})$, which does not depend on the treatment status of others $\bm{d}_{t} \backslash d_{it}$,} The observed outcomes are

align*[align* omitted — 184 chars of source]

for all $i=0,1,\cdots,N$ and $t=1,2,\ldots,T$.

We define the treatment effect $\xi_{0t}$ for the treated unit $0$ at time $t$ $(> T_{0})$ as follows:

align[align omitted — 168 chars of source]

where $Y_{0t}(1,0,\cdots,0)$ is the outcome when unit $0$ is treated, and $Y_{0t}(0,0,\cdots,0)$ is the outcome when the unit is not treated, a counterfactual outcome not observed in the data.

The framework of this study can capture the spillover effects on units in the control group. Although no unit in the control group receives treatment, the influence of unit $0$ receiving treatment may affect the outcomes of the control group units, which is referred to as spillover effects. We define the spillover effect $\xi_{it}$ for any untreated unit $i\ (\geq 1)$ and time $t\ (>T_{0})$ as follows:

align[align omitted — 168 chars of source]

The spillover effect is the counterfactual difference in the outcomes of an untreated unit when unit $i=0$ receives treatment versus when unit $i=0$ does not.

For the outcomes of the treated unit, following abadie2010synthetic, we suppose the following assumption.

assumption[Perfect Fit] There exists a vector of weights $\bm{\alpha} = (\alpha_{1},\alpha_{2},\cdots,\alpha_{N})^{\top}\in\mathbb{R}^{N}$ that satisfies the following: For each $t=1,2,\cdots,T$, \begin{align} Y_{0t}(0,0,\cdots,0) = \sum_{i=1}^{N}\alpha_{i}Y_{it}(0,0,\cdots,0) \ \ a.s. \end{align}

This assumption implies that, in the absence of treatment, the control outcome of the treated unit can be replicated by the weighted average of the control outcomes of the untreated units. This is a fundamental assumption for SCM and is either explicitly or implicitly imposed in most SCM studies. Since our framework allows outcomes to depend on the treatment status of other units, Assumption (ref) pertains to a state where all units are untreated.

Under Assumption (ref), we can estimate the synthetic weights $\bm{\alpha}$ by solving the following least-squares problem:

align[align omitted — 194 chars of source]

Using $\widehat{\bm{\alpha}}$, the standard SCM abadie2010synthetic estimates the treatment effect $\xi_{0t}$ ($t > T_0$) by $Y_{0t} - \sum_{i=1}^{N}\hat{\alpha}_{i}Y_{it}$. Under SUTVA and Assumption (ref) (perfect fit), abadie2010synthetic show the identification and unbiased estimation of treatment effects using SCM. However, when SUTVA is violated, the standard SCM can be biased in its estimation of treatment effects. The subsequent section clarifies the sources of this bias and introduces a new approach to identify both treatment and spillover effects.

Identification

This section begins by clarifying the bias in the standard SCM that arises from spillover effects. We then introduce the SAR panel data model and discuss its capability to account for spillover effects. Finally, we present our main identification result for both treatment and spillover effects.

Bias for the Standard SCM

We first show that when SUTVA is violated, the standard SCM can be biased due to spillover. Assumption (ref) implies that the treatment effects $\xi_{0t}$ can be expressed as $Y_{0t}(1,0,\cdots,0)-\sum_{i=1}^{N}\alpha_{i}Y_{it}(0,0,\cdots,0)$. However, when the outcomes depend on the treatment statuses of other units, we cannot observe counterfactual control outcomes $Y_{it}(0,0,\cdots,0)$ after the pre-treatment periods ($t > T_0$). This causes bias in the standard SCM as follows:

align*[align* omitted — 251 chars of source]

where $Y_{0t} - \sum_{i=1}^{N}\alpha_{i}Y_{it}$ is the standard SCM estimator of $\xi_{0t}$ given knowledge of $\bm{\alpha}$. This result shows the existence of bias in the standard SCM when SUTVA is not satisfied.

Spatial Autoregressive Model

To address the issue of bias caused by spillover, we introduce a model that captures spillover effects between treated and untreated units. Specifically, for the outcomes of units in the control group, we assume the following SAR panel data model for each $t\geq 1$:

align[align omitted — 218 chars of source]

where $\bm{Y}_{t}^{c}(\bm{d}_{t}) \equiv (Y_{1t}(\bm{d}_{t}), Y_{2t}(\bm{d}_{t}),\cdots, Y_{Nt}(\bm{d}_{t}))^{\top}\in\mathbb{R}^{N}$. The vector $\bm{w}\in\mathbb{R}^{N}$ and matrix $\bm{W}\in\mathbb{R}^{N\times N}$ comprise spatial weights, which are to be specified a priori. For $i=1,2,\ldots,N$ and $j=0,1,\ldots,N$, we denote by $w_{ij}$ the spatial weight between units $i$ and $j$; that is, the $(i,j+1)$-th element of the matrix $(\bm{w},\bm{W}) \in \mathbb{R}^{N\times (1+N)}$. Typical examples of spatial weights include adjacent weights, where $w_{ij}$ is 1 if units $i$ and $j$ are adjacent and 0 otherwise. Geographic distance (e.g., distance between the capitals of two countries) and economic distance (e.g., trade amount between two countries) are also often employed as spatial weights in the literature on spatial data analysis lesage2009introduction. The matrix $\bm{X}_{t}\equiv (\bm{X}_{1t},\ldots,\bm{X}_{Nt})^{\top}\in\mathbb{R}^{N\times k}$ denotes the covariate matrix for control units, where $\bm{X}_{it}$ is the covariate vector for unit $i$ at time $t$. We suppose that $\bm{X}_{t}$ does not depend on the treatment status $\bm{d}_{t}$. The vector $\bm{u}_{t}^{\bm{d}_{t}} \equiv (u_{1t}^{\bm{d}_{t}},\cdots,u_{Nt}^{\bm{d}_{t}})^{\top}\in\mathbb{R}^{N}$ contains the error terms.

The SAR panel data model (ref) captures spillover effects arising from spatial dependence in outcomes between treated and untreated units. The magnitude of these spillovers is determined by the spatial autoregressive coefficient $\rho$ as well as the spatial weights $\bm{w}$ and $\bm{W}$, which encode the strength and structure of spatial correlations. Several methods have been proposed to estimate $\rho$ in SAR panel data models fingleton2008generalized,lee2010estimation,su2012semiparametric,glass2016spatial,liang2022semiparametric.

Identification

We now turn to the identification of treatment and spillover effects under Assumption (ref) and the SAR panel data model (ref). For notational simplicity, since the treatment state $d_{it}$ for all $i \geq 1$ and $t$ is always $0$, we write

align[align omitted — 155 chars of source]

We impose the following assumption.

assumptionFor $t=1,2,\cdots,T$ and $i=1,2,\cdots,N$, the error term $u_{it}^{\bm{d}_{t}}$ depends only on its own treatment state $d_{it}$; that is, $u_{it}^{\bm{d}_{t}} = u_{it}^{\bm{d}_{t}^{\prime}}$ a.s. for any $\bm{d}_{t}=(d_{0t},d_{1t},\ldots,d_{1T})^{\top}$ and $\bm{d}_{t}^{\prime}=(d_{0t}^{\prime},d_{1t}^{\prime},\ldots,d_{1T}^{\prime})^{\top}$ such that $d_{it} = d_{it}^{\prime}$.

This assumption rules out spillover effects arising from unobserved factors, focusing instead on those generated by outcome dependence. Addressing the former would require additional structural assumptions that may not align with the standard SCM framework. Given this assumption and the fact that $d_{it}=0$ for any control unit $i$ at any time $t$, the error terms for the control units always satisfy $\bm{u}_{t}^{\bm{d}_{t}}=\bm{u}_{t}^{\bm{0}_{N+1}}$, where $\bm{0}_{N+1}$ denotes an $(N+1)$-dimensional zero vector. For notational simplicity, we denote $\bm{u}_{t}^{\bm{d}_{t}}$($=\bm{u}_{t}^{\bm{0}_{N+1}}$) by $\bm{u}_{t}$.

We further assume that the model ((ref)) and the synthetic weights $\bm{\alpha}$ in Assumption (ref) satisfy the following condition.

assumption$\bm{I}_{N}-\rho \bm{w}\bm{\alpha}^{\top} -\rho \bm{W}$ is full rank.

This assumption ensures that the matrix $\bm{I}_{N}-\rho \bm{w}\bm{\alpha}^{\top} -\rho \bm{W}$ is invertible. Assumption (ref) is testable given estimators of $\bm{\alpha}$ and $\rho$.

Given the parameters $\bm{\alpha}$ and $\rho$, the treatment and spillover effects can be identified as follows. For each $t>T_{0}$, under Assumption (ref) and the SAR panel data model (ref), we have

align[align omitted — 315 chars of source]

Since the matrix $(\bm{I}_{N}-\rho \bm{w}\bm{\alpha}^{\top} -\rho \bm{W})$ is invertible under Assumption (ref), rearranging the above equation yields

align[align omitted — 150 chars of source]

The treatment effect can then be written as

align[align omitted — 349 chars of source]

Furthermore, under the model (ref) and Assumption, (ref), we have for each $t > T_0$:

align[align omitted — 134 chars of source]

Therefore, for $t > T_0$, we obtain

align[align omitted — 427 chars of source]

If the parameters $\bm{\alpha}$ and $\rho$ are estimated using data from the pretreatment periods ($t\leq T_{0}$), then $\xi_{0t}$ ($t> T_{0}$) can be estimated as equation ((ref)). Note that $\bm{\alpha}$ can be consistently estimated by the least-squares method ((ref)) under Assumption (ref), and several methods are available to estimate $\rho$ in the literature of the SAR panel data model fingleton2008generalized,lee2010estimation,su2012semiparametric,glass2016spatial,liang2022semiparametric.

The following theorem summarizes the identification result for $\xi_{0t}$.

theorem[Identification of the Treatment Effect] Suppose that Assumptions (ref), (ref), and (ref) hold. Then, given $\rho$ and $\bm{\alpha}$, the treatment effect $\xi_{0t}$ for $t>T_{0}$ can be identified as follows: \begin{align} \xi_{0t} = Y_{0t} - \bm{\alpha}^{\top} \big(\bm{I}_{N}-\rho \bm{w}\bm{\alpha}^{\top} -\rho \bm{W}\big)^{-1}\big(\big(\bm{I}_{N} - \rho \bm{W}\big)\bm{Y}_{t}^{c} - \rho \bm{w}Y_{0t}\big). \end{align}
proofThe result follows from the discussion above.

The discussion above shows that the counterfactual outcome $\bm{Y}_{t}^{c}(0)$ for each untreated unit for $t>T_{0}$ can be identified as $\big(\bm{I}_{N}-\rho \bm{w}\bm{\alpha}^{\top} -\rho \bm{W}\big)^{-1}\big(\big(\bm{I}_{N} - \rho \bm{W}\big)\bm{Y}_{t}^{c} - \rho \bm{w}Y_{0t}\big)$. Therefore, given the parameters $\bm{\alpha}$ and $\rho$, the spillover effects on the untreated units are also identifiable.

theorem[Identification of the Spillover Effects] Suppose that Assumptions (ref), (ref), and (ref) hold. Then, given $\rho$ and $\bm{\alpha}$, the spillover effects $\bm{\xi}_{t}^{c}=(\xi_{1t},\xi_{2t},\cdots,\xi_{Nt})^{\top}\in\mathbb{R}^{N}$ on the $N$ control units for $t>T_{0}$ can be identified as follows: \begin{align} \bm{\xi}_{t}^{c} & =\bm{Y}_{t}^{c} - \big(\bm{I}_{N}-\rho \bm{w}\bm{\alpha}^{\top} -\rho \bm{W}\big)^{-1}\big(\big(\bm{I}_{N} - \rho \bm{W}\big)\bm{Y}_{t}^{c} - \rho \bm{w}Y_{0t}\big). \end{align}
proofNote that $\bm{Y}_{t}^{c} = \bm{Y}_{t}^{c}(1)$ for $t>T_0$ and that $\big(\bm{I}_{N}-\rho \bm{w}\bm{\alpha}^{\top} -\rho \bm{W}\big)^{-1}\big(\big(\bm{I}_{N} - \rho \bm{W}\big)\bm{Y}_{t}^{c} - \rho \bm{w}Y_{0t}\big) = \bm{Y}_{t}^{c}(0)$. This leads to the result ((ref)).

Theorems (ref) and (ref) show that by incorporating the SAR model into the SCM framework, treatment and spillover effects can be identified without relying on SUTVA. The results in Theorems (ref) and (ref) also suggest that although the SAR model ((ref)) includes the covariates $\bm{X}_t$ and error terms $\bm{u}_{t}$, it is not necessary to estimate $\bm{\beta}$ or the distribution of $\bm{u}_{t}$ to estimate treatment and spillover effects. For their estimation, only the spatial autoregressive parameter $\rho$ and the spatial weights $(\bm{w},\bm{W})$ in the SAR model ((ref)) are relevant.

Remarkably, when there is no spillover (i.e., $\rho=0$), our identification result ((ref)) for $\xi_{0t}$ corresponds to the identification result of the standard SCM abadie2010synthetic. That is, when $\rho=0$, the identification result ((ref)) simplifies to

align*[align* omitted — 404 chars of source]

where the final line is identical to the identification equation for the standard SCM abadie2010synthetic. This result indicates that the identification result obtained from the standard SCM is the special case of our identification result when $\rho=0$ (no spillover).

Inference

In this section, we propose a Bayesian inference method for treatment effects $\xi_{0t}$ and spillover effects $\xi_{it}$ ($i =1,2,\ldots,N$), building on the identification results presented in Section (ref). SCM is typically applied to data with small or no large pretreatment periods, where frequentist methods may not provide accurate estimations. On the other hand, Bayesian methods can provide precise statistical inference (e.g., credible intervals) through MCMC procedures even with short pretreatment periods. Moreover, by examining the posterior samples of $\rho$, one can test for the presence of spatial correlation.

This section presents a Bayesian inference method for the model described in Sections (ref) and (ref). Let $\bm{\theta}= (\bm{\alpha},\rho,\bm{\beta},\bm{\eta},\{\bm{\gamma}_{t}\}_{t=1}^{T_{0}},\sigma^{2})$ denote the parameter vector. Under the model in Sections (ref)--(ref), the pre-treatment outcomes satisfy

align*[align* omitted — 209 chars of source]

Letting $\tilde{\bm{u}}_t$ denote the transformed disturbance term, the conditional density of $\bm{Y}_t^c$ given $(Y_{0t},\rho,\bm{\beta},\bm{\eta},\bm{\gamma}_t,\sigma^2)$ is Gaussian with Jacobian factor $\lvert \bm{I}_N - \rho \bm{W} \rvert$. Thus, the joint likelihood for the pre-treatment sample is

align[align omitted — 378 chars of source]

Expression ((ref)) makes explicit that (i) information from the treated unit $Y_{0t}$ enters entirely through the synthetic-control restriction, while (ii) the spatial autoregressive structure governs the distribution of $\bm{Y}_t^{c}$. However, the likelihood contains the multiplicative interaction term $\rho\,\bm{w}\bm{\alpha}^\top$, which produces weak identification of $(\bm{\alpha},\rho)$ and leads to extremely slow mixing when implementing full joint MCMC. Furthermore, jointly sampling all latent factor components $(\bm{\eta},\bm{\gamma}_t)$ over $T_0$ periods imposes substantial computational burden without improving post-treatment inference. In practice, we found that full joint estimation results in unstable chains, poor effective sample sizes, and unreliable inference for the treatment effects. To overcome these identification and computational difficulties, we decompose the joint likelihood in ((ref)) into two components,

align[align omitted — 328 chars of source]

and estimate the two components sequentially. We therefore adopt a two-step Bayesian inference strategy in which (i) the synthetic control weights $\bm{\alpha}$ are sampled from (ref), and the spatial parameter $\rho$ is sampled from (ref) conditional on $\hat{\bm{\alpha}}$. The detailed MCMC procedures for each step are described in the following subsections. The validity of the MCMC implementation including the joint distribution test of geweke2004getting, prior predictive checks, sensitivity analyses, and convergence diagnostics is documented in detail in the Supplementary Appendix.

Synthetic Weights

Regarding the estimation of the synthetic weights $\bm{\alpha}$, we adopt kim2020bayesian's (kim2020bayesian) method, which utilizes the Bayesian horseshoe prior as the prior distribution for parameters. The Bayesian horseshoe prior places a hierarchical prior distribution on parameters, characterized by a strong shrinkage effect around zero. park2008bayesian show that LASSO regression in linear models corresponds to estimation with a Laplace prior distribution on coefficients, while the Bayesian horseshoe prior proposed by carvalho2010horseshoe suggests an even more selective coefficient choice. Thus, by placing this prior distribution on the synthetic weights, it is anticipated that the estimated weights will select the units in the control group that are more related to the treated unit.

The following hierarchical prior distributions are placed on $\alpha_{1},\alpha_{2},\cdots,\alpha_{N}$:

align*[align* omitted — 373 chars of source]

When these prior distributions are set, as shown by makalic2015simple, the full conditionals of each parameter can be derived analytically, enabling parameter sampling using the Gibbs sampler. Full conditional distributions for each parameter are presented in the supplementary appendix.

Spatial Autoregressive Panel Data Model

Regarding the model ((ref)), we set the Bayesian horseshoe prior for $\bm{\beta}$ as in estimating the synthetic weights, that is,

align*[align* omitted — 357 chars of source]

We assume that $\bm{u}_{t}$ follows a multilevel latent factor model $\bm{u}_{t}=\bm{\eta}\bm{\gamma}_{t}+\bm{e}_{t}$, where $\bm{\gamma}_{t}\in\mathbb{R}^{p}$ denote factors at time $t$, $\bm{\eta}\in\mathbb{R}^{N \times p}$ are factor loadings, and $\bm{e}_{t} \in\mathbb{R}^{N}$ is an error vector. We assume $\bm{\gamma}_{t}$ follow the AR($1$) model $\bm{\gamma}_{t}=\phi_{\gamma}\bm{\gamma}_{t-1} + \bm{v}_{t}$, where $\bm{v}_{t}\sim N(\bm{0}_{p},\sigma^{2}_{\gamma}\bm{I}_{p})$. For each control unit $i$ ($=1,\cdots,N$), $\eta_{i}$ follows $p$-dimensional normal distribution $N_{p}(\bm{0}_{p},\sigma^{2}_{\eta}\bm{\Sigma}_{\eta})$, where $\bm{\Sigma}_{\eta}=\mathrm{diag}(\omega_{1}^{2},\cdots,\omega_{p}^{2})$. The priors of $\sigma_{\eta}$ and $\omega_{1}, \cdots,\omega_{p}$ are $\text{half-Cauchy}(0,10)$. These settings are the same as those used by pang2022bayesian.

We adopt the method proposed by lesage1997bayesian and lesage2009introduction for modeling $\rho$. Given all parameters except $\rho$, the conditional distribution of $\rho$ is

align*[align* omitted — 164 chars of source]

where $\tilde{\bm{u}}_{t} = \bm{A}\bm{Y}_{t}^{c} - \rho\bm{w}Y_{0t}-\bm{X}_{t}\bm{\beta} - \bm{\gamma}_{t}^{\top}\bm{\eta}$ and $\bm{A}=\bm{I}_{N}-\rho\bm{W}$. However, since this is not a standard form, it is impossible to analytically derive the full conditional distribution. Therefore, sampling is conducted using the Metropolis algorithm. If the value of $\rho$ in the $m$-th sampling is $\rho^{(m)}$, then the proposal value $\rho^{*}$ for the ($m+1$)-th sampling is determined by $\rho^{*}=\rho^{(m)}+ k \cdot N(0,1)$, where $k$ is a tuning parameter adjusted to achieve an acceptance rate of approximately $40\%$ to $60\%$. The MCMC procedure for all parameters is detailed in the supplementary appendix.

Using the sampled parameters from the posterior distributions, we can estimate the treatment and spillover effects through the identification results ((ref)) and ((ref)) in Theorems (ref) and (ref). Specifically, if $M$ samplings are performed, the estimated values of the treatment and spillover effects for the $m$-th sample are as follows:

align*[align* omitted — 429 chars of source]

Calculating this for each $m\ (=1,2,\cdots,M)$ allows us to construct the distribution of the treatment and spillover effects. Then the posterior means of $\xi_{0t}$ and $\bm{\xi}_{t}$ can be obtained by $(1/M)\sum_{m=1}^{M}\xi_{0t}^{(m)}$ and $(1/M)\sum_{m=1}^{M}\bm{\xi}_{t}^{(m)}$, respectively. Given the estimators of $\bm{\alpha}$ and $\rho$, the estimation $\xi_{0t}$ and $\bm{\xi}_{t}$ depends only on the observed outcomes $Y_{0t}$ and $\bm{Y}_{t}^{c}$ over time. This suggests that time-varying components in the SAR panel data model, such as factors $\bm{\gamma}_{t}$, are not required to be estimated in the post-treatment periods $t\ (> T_{0})$.

Simulation Study

Simulation Design

We conduct a simulation study to examine the finite sample performance of the proposed Bayesian inference method. Three scenarios are considered for the number of untreated units $N$: $N=16$, $36$, and $64$. For the length of time periods, we consider two scenarios: $(T,T_{0})=(30,20)$ and $(60,50)$, where each scenario has a post-treatment length of ten.

In each scenario of $(N,T,T_0)$, we set up the DGPs as follows. We consider a network of $N$ untreated units represented by a rook matrix.\footnote{We use a rook matrix based on an $r$ board (so that $N=r^2$) to represent the network of $N$ untreated units. The rook matrix represents a square tessellation with connectivity of four for the inner fields on the chessboard, and two and three for the corner and border fields, respectively. } Subsequently, we set the spatial weight $\omega_{ij}$ of any two untreated units to be $1$ if units $i$ and $j$ are connected in the network and $0$ otherwise. Each element of the adjacency vector $\bm{w}$ takes the value of $1$ if $i \in \{1,2,3,4\}$ and $0$ otherwise. The synthetic weights $\bm{\alpha}$ are set to provide large weights only to control units adjacent to the treatment unit, as follows:

align*[align* omitted — 319 chars of source]

For each $t$ $(=1,2,\cdots,T)$, the control outcomes of untreated units $\bm{Y}_{t}^{c}(0)$ are distributed from a SAR panel data model as follows: $ \bm{Y}_{t}^{c}(0) = \rho\bm{w}Y_{0t}(0) + \rho\bm{W}\bm{Y}_{t}^{c}(0) + \bm{X}_{t}\bm{\beta} + \bm{u}_{t}$, where $\bm{X}_{t} = (X_{1t},\ldots,X_{NT})^{\top}$ with each element being i.i.d as $N(0,1)$, $\bm{u}_{t} = (u_{1t},\ldots,u_{NT})^{\top}$ with each element being i.i.d as $N(0,1)$, and $\beta = 1.0$. The control outcome of the treated unit $Y_{0t}(0)$ is generated as $Y_{0t}(0) =\sum_{i=1}^{N}\alpha_{i}Y_{i,t}(0)$. We consider seven scenarios of $\rho$: $\rho=-0.8,\ -0.3,\ -0.1,\ 0.0,\ 0.1,\ 0.3,\ 0.8$. A larger absolute value of $\rho$ implies a stronger spatial correlation among the units.

We generate the treatment outcomes of the treated unit $Y_{0t}(1)$ as $Y_{0t}(1) = Y_{0t}(0) + N(1, 1)$. We also set $ \bm{Y}_{t}^{c}(1) = \rho\bm{w}Y_{0t}(1) + \rho\bm{W}\bm{Y}_{t}^{c}(1) + \bm{X}_{t}\bm{\beta} + \bm{u}_{t}.$ The observed outcome for each $i$ and $t$ is $Y_{it} = Y_{it}(0)\cdot 1\{t\leq T_0\} + Y_{it}(1)\cdot 1\{t > T_0\}$.

Following the DGPs, we conduct 1000 Monte Carlo simulations. For each simulation $r(=1,\cdots,1000)$, we draw $M(=5000)$ samples by MCMC and compute bias and root mean squared error (RMSE) of the posterior mean of treatment effect as follows:

align*[align* omitted — 283 chars of source]

where $\xi_{0t}^{(r)}$ denotes a true treatment effect in the $r$-th simulation, $\widehat{\xi}_{0t}^{(r)}=\sum_{m=1}^{M}\xi_{0t}^{(r,m)}/M$ is the estimated treatment effect with $\xi_{0t}^{(r,m)}$ being the estimate of the treatment effect in the $m$-th iteration for the $r$-th simulation.

We compare the performance of the proposed method (labeled “Proposed”) with those of the standard SCM (labeled “SCM”) abadie2010synthetic and the Bayesian SCM of kim2020bayesian (labeled “BSCM”).\footnote{SCM does not involve MCMC sampling.} We also calculate the 95% coverage rate of treatment effect for the proposed method.

Results

Table (ref) presents the simulation results for bias and RMSE. The key finding is that the error in estimating the treatment effect using SCM and BSCM becomes large as the absolute value of $\rho$ increases. Particularly in the case of a strong positive spatial correlation ($\rho=0.8$), SCM and BSCM exhibit substantial biases and large RMSEs. These results indicate that, when spillovers exist, the treatment effects estimated by SCM and BSCM are biased. Conversely, the proposed method exhibits a small bias and RMSE in each simulation scenario. The bias and RMSE of the proposed method do not increase with the magnitude of the spatial correlation. These results indicate that the proposed method is robust to the spillover effects arising from the spatial correlation of the outcomes.

Table (ref) presents the coverage rate of $95$% credible interval of the treatment effect for the proposed method. In each scenario, the coverage rate is close to 95%, indicating that the proposed inference method performs adequately. The inference performs well even when the length of pretreatment periods is not very large ($T_0 = 20$) and/or the spatial correlation is strong (i.e., $|\rho|$ is large).

Empirical Application

We conduct two empirical studies applying the proposed method. The first estimates the impact of California’s tobacco tax on cigarette consumption abadie2010synthetic, and the second examines the economic impact of Sudan’s 2011 division on GDP per capita. We present the results of both applications below.

Application I: California Tobacco Tax

In this section, we apply the proposed method to estimate the effect of California tobacco tax (Proposition 99) on cigarette consumption abadie2010synthetic. Proposition 99 is an anti-tobacco law issued in California in 1988 to promote awareness of the health risks associated with tobacco. It raised the excise tax on cigarettes by 25 cents per pack. abadie2010synthetic estimate the treatment effect of Proposition 99 on cigarette sales using the standard SCM, without accounting for spillover effects. Our proposed method allows for spillovers and enables the estimation of both treatment and spillover effects of Proposition 99.

Data

We use annual state-level cigarette sales data in the U.S. from 1970 to 2000, as used in abadie2010synthetic.\footnote{ This dataset is available from cunningham2021causal on GitHub. It has been preprocessed according to the procedures described in abadie2010synthetic. } The outcome of interest is annual per capita cigarette consumption at the state level. We compare our proposed method with the standard SCM abadie2010synthetic (labeled as “SCM”), which is not robust to the existence of spillover effects. For SCM, we use the synthetic weights estimated by abadie2010synthetic.\footnote{The estimated values are listed in Table 2 of abadie2010synthetic.}

In implementing the proposed method, we include the average retail cigarette prices for each year $t$ as covariates $\bm{X}_{t}$. For the spatial weights $w_{ij}$ in the SAR model ((ref)), we construct a contiguity-based adjacency matrix using U.S. state boundary shapefiles from the Census TIGER/Line data. Specifically, we set $w_{ij}=1$ if states $i$ and $j$ share a common land border and $w_{ij}=0$ otherwise, and we set all diagonal elements $w_{ii}$ to zero. We then row-normalize $\bm{W}$ so that each row sums to one. We also construct the vector $\bm{w}$ by the same way.

Results

Figure (ref) presents the estimation results for the synthetic control outcomes for California and the treatment effects of the California tobacco tax on cigarette consumption, with the posterior means reported for the proposed method. Panel (a) shows that both the SCM and the proposed method fit well with per-capita cigarette sales in California during the pretreatment periods. Panel (b) indicates that each estimation method suggests Proposition 99 reduced cigarette consumption in California, with the proposed method exhibiting larger treatment effect estimates than SCM. The 90% credible interval for the proposed method is negative for all post-treatment years, beginning in 1988. This result suggests that Proposition 99 negatively impacted cigarette sales for a decade, supporting the findings of abadie2010synthetic. Figure (ref) illustrates the posterior mean synthetic weights from our method along with the weights reported by abadie2010synthetic.

Figure (ref) reports the estimated spillover effects for all states in the control group. The results align closely with economic and geographic intuition: the estimated spillover effects are largest in states geographically adjacent to California. The most pronounced negative effect appears in Nevada, a direct neighbor of California, while smaller yet persistent negative effects are also evident for Idaho and Utah.

Application II: The Economic Cost of the 2011 Sudan Split

In this section, we assess the impact of Sudan's north-south split in 2011 on GDP per capita in the Sudans (the region of the former united Sudan) and other African countries.

Background

Sudan has long been divided along ethnic and religious lines, with Arabs (primarily Muslims) predominantly in the north and Africans (primarily Christians) in the south, leading to many conflicts. Particularly, the Darfur conflict, driven by the Arab versus non-Arab ethnic divide, has persisted for many years in western Sudan. This conflict has been marked by large-scale atrocities, including mass killings carried out by Arab militias known as “Janjaweed.”

Amid these unending conflicts, the Sudanese government and the Sudan People's Liberation Army (SPLA), the main rebel force in Southern Sudan, signed the Comprehensive Peace Agreement (CPA) in 2005. This agreement, aimed at ending Sudan's civil wars, allowed South Sudan to establish its own government, achieve autonomy, and pursue independence through a referendum. Following the CPA, South Sudan voted for independence in January 2011 and was officially recognized as a nation on July 9, 2011.

Since 2011, South Sudan's independence has caused several economic disturbances in both Sudan and South Sudan. In particular, South Sudan's oil production shutdown in 2012 and its relapse into conflict in 2013 provoked a severe macroeconomic crisis in the region mawejje2020macroeconomic. This study estimates the economic impact of Sudan's south-north split, with a focus on GDP per capita in the Sudans and other African countries.\footnote{Using the standard SCM, mawejje2021economic estimate the economic losses in South Sudan, owing to the oil production halt in 2012, finding a nearly 70% loss in per capita real GDP from 2012 to 2018. They excluded neighboring countries of South Sudan from their SCM analysis to avoid bias caused by spillovers; however, this practice can result in a poorly fitting synthetic control, making the perfect-fit assumption (Assumption (ref)) less plausible and potentially causing additional bias.}

Data

We use data from African countries obtained from the World Bank DataBank. Our outcome of interest is “GDP per capita (constant 2015 US\$)” in the post-division period. The covariates we use include “exports of goods and services (% of GDP)”, “merchandise trade (% of GDP)”, “access to electricity (% of population)”, “inflation measured by the consumer price index (annual %)”, “net migration”, and “trade (% of GDP)”. Countries with missing values for the outcome or any covariates were excluded.

We focus on GDP per capita in the Sudans (i.e., the region of the former united Sudan) as the treated unit's outcome $Y_{0t}$.\footnote{The GDP per capita in the Sudans after the Sudan split is calculated by dividing the sum of GDP in North and South Sudan by the sum of their populations. Before the split, it corresponds to the GDP per capita in (the united) Sudan.} The control group consists of $N = 29$ African countries with complete data from 2000 to 2015.\footnote{The control group includes: Algeria, Angola, Benin, Botswana, Burundi, Cameroon, Central African Republic, Chad, Egypt, Gabon, Ghana, Ivory Coast, Kenya, Madagascar, Mali, Mauritania, Mauritius, Morocco, Niger, Nigeria, Republic of the Congo, Rwanda, Senegal, South Africa, Tanzania, Togo, Tunisia, Uganda, and Zambia.} Since South Sudan's independence occurred in July 2011, we define the pre-treatment periods as 2000–2010 and the post-treatment periods as 2011–2015. Although GDP per capita in the Sudans cannot be computed for 2011 due to incomplete data following the split, this does not affect the estimation of synthetic control outcomes in the pre-treatment periods or the treatment effects from 2012 onward.

Regarding the spatial weights in model ((ref)), we specify $w_{ij}$ as the average bilateral trade volume between countries $i$ and $j$, normalized by the total trade of country $i$:

align*[align* omitted — 185 chars of source]

where the trade data are obtained from the IMF and the averages are calculated over the pre-intervention periods. Each weight $w_{ij}$ captures the strength of the economic ties between the two countries. Figure (ref) presents the resulting weights $w_{ij}$ between the former united Sudan and each control country during the pre-treatment period. Countries with stronger economic connections to the former unified Sudan are expected to experience more pronounced spillover effects from the Sudan split.

Results

Figure (ref) presents the estimation results for the synthetic outcomes of the Sudans and the treatment effects of the Sudan split. Across all methods, the estimates indicate a negative impact of South Sudan’s independence on GDP per capita in the Sudans. In particular, the proposed method estimates a decline of approximately 100USD in 2012, corresponding to about 7.8% of the actual GDP per capita for that year. In 2015, the estimated GDP per capita in the synthetic Sudan is about $9.5\%$ higher than the actual GDP per capita. Furthermore, the cumulative losses in the Sudans are estimated to have reached $34\%$ from 2011 to 2015.\footnote{The losses in the Sudans are defined by $100\times(\xi_{0t}/Y_{0t}(0))$ (%) for each $t>T_{0}$ and those in control countries are defined by $100\times\xi_{it}/Y_{it}(0)$ (%) for each $i=1,2,\cdots,N$ and $t>T_{0}$. The cumulative losses are computed by $100\times \sum_{t=2011}^{2015}\xi_{it}/Y_{it}(0)$ for each $i=0,1,\cdots,29$. We estimate these by the posterior means of these losses.} Figure (ref) displays the posterior mean weights from our Bayesian SCM alongside the weights estimated by the standard SCM.

Figure (ref) presents the estimated spillover effects of the Sudan split on other African countries. Countries with substantial trade volumes with the former united Sudan, such as Egypt and Kenya, experienced pronounced negative spillover effects from the split.

The political instability and north–south split in Sudan significantly altered the country’s industrial structure, which in turn had a profound impact on trade. For example, although South Sudan is rich in oil and other natural resources, the conflict between the North and South led to a halt in oil production in 2012, resulting in the loss of critical export commodities. This disruption also affected countries with close economic ties to Sudan through trade channels. Thus, the north–south split not only had adverse economic consequences for the Sudans themselves but also generated negative spillover effects on other countries. This empirical study illustrates that political and economic changes in one country can have broader consequences for other nations with strong economic linkages.

Conclusion

This study extends the SCM to account for spillover effects. Although SCM is frequently applied to spatial data where spillovers may be present, the conventional approach relies on SUTVA, which can yield biased estimates when spillovers are present. To address this limitation, we propose a novel SCM that incorporates the SAR panel data model to capture spillover effects. We also develop a Bayesian inference procedure that estimates both treatment and spillover effects, using horseshoe priors for regularization. We apply the method to two empirical studies: (i) evaluating the impact of the California tobacco tax on cigarette consumption abadie2010synthetic, and (ii) assessing the economic impact of the 2011 Sudan division on GDP per capita. The first application reveals a negative impact of the tax on cigarette consumption in California and other US states. The second shows that the Sudan split substantially reduced GDP per capita in the Sudans and caused negative spillover effects on other African countries with strong economic ties to the former united Sudan.

Acknowledgments

We thank the editor and two anonymous referees for helpful comments that greatly improved the quality of the paper. We also thank Kaoru Irie, Ryo Okui, Yasuyuki Sawada, and participants in various seminars and workshops for valuable comments. The authors gratefully acknowledge the financial support from JSPS KAKENHI Grant (number 24K16342).

\setstretch{1.5}

Tables

table[table omitted — 3,703 chars of source]
table[table omitted — 1,031 chars of source]

\newgeometry{left=1cm, right=1cm, top=2cm, bottom=2cm}

Figures

figure[figure omitted — 1,267 chars of source]
figure[figure omitted — 586 chars of source]
figure[figure omitted — 511 chars of source]
figure[figure omitted — 556 chars of source]
figure[figure omitted — 555 chars of source]
figure[figure omitted — 1,332 chars of source]
figure[figure omitted — 609 chars of source]

\restoregeometry