EconBase
← Back to paper

Estimation and Inference for Synthetic Control Methods with Spillover Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

153,290 characters · 42 sections · 89 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimation and Inference for Synthetic Control Methods with Spillover Effects

\abstract{ Estimation and inference procedures for synthetic control methods often do not allow for the existence of spillover effects, which are plausible in many applications. In this paper, we consider estimation and inference for synthetic control methods, allowing for spillover effects. We propose estimators for both direct treatment effects and spillover effects and show that they are asymptotically unbiased. In addition, we propose an inferential procedure and show that it is asymptotically unbiased. Our estimation and inference procedure applies to cases with multiple treated units and/or multiple post-treatment periods, and to ones where the underlying factor model is either stationary or cointegrated. We discuss the bias from misspecified spillover structures and propose a test for correct specification. We apply our method to a classic empirical example that investigates the effect of California's tobacco control program as in Abadie2010 and find evidence of spillovers. We contrast our method with the pure-donor approach through a sensitivity analysis. }

Introduction

The synthetic control method (SCM), introduced by Abadie2003, has gained popularity for estimating treatment effects in settings with few treated units and post-treatment periods, such as studying state-level policies. By leveraging pre-treatment data, SCM often provides better counterfactual estimates than methods like difference-in-differences in comparative case studies Abadie2018,Ferman2016. However, like many other methods, SCM relies on the Stable Unit Treatment Value Assumption (SUTVA), which assumes untreated units are unaffected by the treatment. This assumption is often violated, and the structure of SCM makes it particularly vulnerable to bias when SUTVA does not hold. For instance, when California increases cigarette taxes, SUTVA implies that consumers won't shift purchases to neighboring states like Nevada, an assumption we show is likely to be violated in Section (ref). Similar SUTVA violations are common in many applications using geographically aggregated data.

The presence of spillover effects can severely bias SCM in estimating the direct treatment effect (formally defined later by $\alpha_1=y_{1,T+1}(1,0,\dots,0)-y_{1,T+1}(0,\dots,0)$). Post-treatment control contamination leads to a biased counterfactual estimate and, consequently, a biased treatment effect estimate. Compared to difference-in-differences, this contamination-induced bias can be substantially larger, if spillovers are concentrated in heavily-weighted control units. Moreover, if spillovers propagate along the same channels as the underlying factor model, SCM may actively select for bias-inducing units. We illustrate these bias effects in our simulation section (Section (ref)), showing a scenario where SCM disproportionately assigns weights to control units with higher spatial correlation, amplifying the bias when spillover effects also increase with spatial correlation.

The problem of spillover effects can be partially solved by eliminating contaminated units in estimation, known as the pure-donor method, where a pure donor refers to a control unit not experiencing spillover effects. Similar methods include those developed for multiple treated units, which are often equivalent to the pure-donor method Cavallo2013, Firpo2018, Kreif2016, Robbins2017, Xu2017. However, these methods may be concerning for several reasons. First, there are cases where most, if not all, control units are affected by the spillover, making it impractical to exclude all such units. We aim to develop a practical solution for these scenarios that allows effective analysis even when spillovers are widespread. Second, contaminated units are often crucial in forming the synthetic control. Their exclusion can lead to a potential loss of efficiency, as these units often contribute valuable information to the estimation DiStefano2020. In addition, by using fewer control units, the pure-donor method exhibits a large worst-case scenario bias when the spillover structure is misspecified. Specifically, this method presents a larger identified set of potential biases. We argue that this bias can be significantly mitigated by using a full-sample method. We illustrate this through a sensitivity analysis conducted in our empirical application.

The goal of this paper is to relax the SUTVA condition and to perform estimation and testing. Particularly, we look at the cases where there are spillover effects, which are defined by a Rubin model as the difference between the actual outcomes and the counterfactual ones. To facilitate estimation, we assume some knowledge about the spillover effects is known. Specifically, the treatment effect and the spillover effects are linear in some unknown parameters. We give examples where this assumption is plausible. Thanks to the known spillover structure, we can propose an asymptotically unbiased estimator for the treatment and spillover effects. We also characterize the asymptotic distribution of the estimator. Compared with existing methods, our method uses information from all known units in estimation. Our method can also often deal with cases where all units are contaminated by spillovers, at the cost of assuming more structure on the spillovers.

We follow the setup in Ferman2016, where we focus on cases with an imperfect pre-treatment fit. We will proceed under the assumption that treatment assignment is not correlated with time-varying unobservables, under which Ferman2016 show that the demeaned SCM ensures asymptotically unbiased estimation of treatment effects. We also rely on the imperfect pre-treatment fit to identify the null distribution of the proposed statistic. In terms of the asymptotic framework, we consider cases with many pre-treatment periods and a fixed number of control units. As suggested by Ferman2016, even in cases where large-$T$ asymptotics is not justified, our results can be interpreted as the SCM weights not converging to weights that reconstruct the factor loadings of the treated unit when the number of pre-treatment periods is large. Besides, Monte-Carlo simulation shows that our methods produce reasonable estimation and testing results as long as there is a moderate number of pre-treatment periods.

Additionally, we propose an inferential procedure based on Andrews2003's end-of-sample instability test, or $P$-test. We generalize the $P$-test to the synthetic control method with and without spillover effects. Similar to the $P$-test, our testing procedures use the idea of approximating the null distribution of the statistic using pre-treatment data. We show the validity of the proposed method and compare it with the standard placebo test through a simulation study. Tangent to the main idea of this paper, our method alleviates the problem of selection into treatment, which is a major threat to the placebo test.

We give high-level conditions under which our methods are valid. In addition, our conditions adapt to factor models with either stationary or cointegrated common factors, which are often used to justify the usage of synthetic control methods. Furthermore, we consider extensions where treatment applies to multiple units or periods, and where there are extra covariates.

Throughout, we assume the spillover structure is known. We illustrate through our empirical application in Section (ref) how contextual knowledge can be used to inform the structure of the spilovers. Moreover, we propose a test for correct specification ($\kappa_A$-statistic) to alleviate concern regarding structure misspecification. In Section (ref), we extend this test to settings with multiple post-treatment periods, and give conditions under which consistent structure estimation is possible.

Another assumption we adhere to in the paper is that there is no selection on time-varying unobservables, a condition similarly employed by Ferman2016. For potential extension to cases without this assumption, readers may explore settings with a perfect pre-treatment fit Abadie2010,DiStefano2020, or settings where the number of controls approaches infinity arkhangelsky_synthetic_2021, ferman_properties_2021.

We revisit the empirical example from Abadie2010 on California's 1989 cigarette tax. Abadie2010 used data from 1970 onward, excluding 12 states potentially affected by spillovers or interventions, leaving 38 control states. Despite this, we find evidence of SUTVA violations in most post-tax years, even within the 38 control states, and in the 12 excluded states. Our estimates of the tax's impact are consistently smaller than Abadie2010's for the first four post-treatment years, potentially due to adjustments for significant early spillovers in Nevada, which was heavily weighted in standard SCM. We also conduct a sensitivity analysis of worst-case bias from misspecified spillover structures, examining how the identified bias set expands with increasing spillover magnitude. The pure-donor method exhibits a larger identified bias set than our method.

This paper contributes to three developing literatures. First, it complements the literature on statistical inference in SCM by providing formal results without assuming SUTVA. Many works have developed inference methods for settings with few treated units and short pre- and post-treatment periods, such as difference-in-differences Conley2011, which can be viewed as a special case of SCM with equal weights, and placebo tests that permute across units Hahn2017. Most related work to ours is Chernozhukov2021, who also use variation across periods rather than units to perform testing. Li2019 proposes a testing procedure based on projection onto convex sets and results from Fang2018. cao_synthetic_2025 consider synthetic control methods in the context of staggered adoption. However, few works consider spillovers. To our knowledge, the only inferential methods allowing for spillovers are DiStefano2020 and Grossi2020. DiStefano2020 require perfect pre-treatment fit, while our method allows for imperfect fit. Grossi2020 use a penalized SCM similar to Abadie2021 and assume exchangeable groups, which we do not require. Furthermore, our method applies to cointegrated factor models, of interest even without spillovers.

We also contribute to the growing literature on estimating treatment and spillover effects. This fast-growing literature looks into both estimation of treatment effects in the presence of spillover effects, as well as estimation of spillover effects themselves. For example, Vazquez-Bare2017 considers a framework with clusters, allowing spillovers within but not across clusters, and estimates heterogeneous treatment and spillover effects. Basse2019 and Rosenbaum2007 use randomization tests for inference with spillovers. See Basse2019 and Vazquez-Bare2017 for literature reviews. However, this literature rarely considers panel data settings with few treated units and short post-treatment periods, partly due to insufficient information about spillovers. We address this by assuming pre-specified spillover structures linear in underlying parameters, enabling estimation and testing of spillover effects. We also introduce a test for correct specification of spillover structures and provide conditions for consistent spillover structure estimation with multiple post-treatment periods, contributing to existing work on spillover structure estimation in panel data de_paula_identifying_2023,Manresa2013-cr.

Third, our results extend the literature on Andrews2003's end-of-sample instability tests. Andrews2003 uses data across time periods to approximate the null distribution of the test statistic and applies this idea to OLS, IV, and GMM. Chernozhukov2021 propose a permutation test that is more general, but similar in cases where serial correlation matters. We extend this idea to the SCM case, and further to more complicated cases with spillover effects. As Andrews2006 extend Andrews2003's results to the cointegrated cases, we also show that our method is still valid for a cointegrated factor model.

The rest of this paper is organized as follows. Section (ref) introduces a potential outcome framework with spillover effects and defines a known spillover structure that will be useful later. Section (ref) proposes an estimator and derives its asymptotic distribution. It also discusses an example of a factor model and an invertibility assumption of the proposed estimator. Section (ref) considers the $P$-test introduced by Andrews2003 and Andrews2006 and explains how it can be applied in our setting. In Section (ref), we discuss spillover misspecification and comparison with the pure-donor method. An empirical application of our method is presented in Section (ref). Section (ref) further discusses (i) an example where all control units are affected by the treatment and how to interpret a relevant parameter of interest, and (ii) a comparison between the proposed inference procedure and existing ones. Section (ref) discusses some extensions of our methods, including a similar estimator with a small asymptotic variance, cases with multiple treated units and/or multiple post-treatment periods, and cases with additional covariates. Section (ref) presents all the Monte Carlo simulation results, and Section (ref) contains all the proofs.

Model Specification

A Rubin model with spillover effects

We start our discussion with a Rubin potential outcome framework. We consider a standard synthetic control setting where only one unit is treated and only one period is available after the treatment is implemented. We consider cases with multiple treated units and multiple post-treatment periods in Section (ref).

In Rubin's model with a violation of SUTVA, the potential outcomes are functions of treatment assignments on all units. Assume the outcome of unit $i$ at time $t$ is $y_{i,t}=y_{i,t}(d_t),$ where $d_t=(d_{1,t},\dots,d_{N,t})'$ and $d_{i,t}=1$ if unit $i$ has been treated at time $t$. Assume unit 1 is treated between time $T$ and $T+1$, and there are another $N-1$ units that are not directly treated by the policy. Thus, we observe an $N\times (T+1)$ panel as shown in Figure (ref).

figure[figure omitted — 710 chars of source]

Note that we only observe outcomes with $d_{t}=(0,\dots,0)'$ for $t=1,\dots,T$ and $d_{T+1}=(1,0,\dots,0)'$. This is the fundamental limitation of the dataset we are currently studying. Unless other homogeneity conditions are assumed, we cannot say anything about $y_{i,T+1}(d_{T+1})$ for $d_{T+1}\not\in \{(0,\dots,0)',(1,0,\dots,0)' \}$, because only one unit is treated and only one post-treatment period is available. For notation simplicity, let

equation*[equation* omitted — 116 chars of source]

for each $(i,t)$. Let $\alpha_i=y_{i,T+1}(1)-y_{i,T+1}(0)$ be the potential deviation from unit $i$'s counterfactual outcome $y_{i,T+1}(0)$ where no unit is treated at time $T+1$. That is, $\alpha_1$ is the direct treatment effect on unit 1, while $\alpha_i$ with $i\ne 1$ is the indirect effect or spillover effect. The whole effect vector $\alpha=(\alpha_1,\dots,\alpha_N)'$ can be of interest in our setting. Throughout, we consider the case where $N$ is fixed and $T$ goes to infinity.

Our analysis will start with the following linear model, which will be formally defined later:

equation[equation omitted — 78 chars of source]

where $Y_{t}(0)=(y_{1,t}(0),\dots, y_{N,t}(0))'$ is the stacked untreated outcomes, $a\in \mathbb R^{N\times 1}$, $B=\mathbb R^{N\times N}$ with diagonal entries being zero, and $u_{t}\in \mathbb R^{N\times 1}$ is a mean-zero stationary process. This equation characterizes the relationship between any unit, including both treated and controls, and the rest. The plan is to use the pre-treatment data to estimate $a$ and $B$, and hope this relationship carries over to the post-treatment period, from which we can learn an asymptotically unbiased estimator of both treatment and spillover effects. Importantly, although $a$ and $B$ will be learned using SCM in this paper, all high-level conditions discussed later do not rely on a specific restriction, allowing for alternative estimation strategies.

Known spillover structures

Throughout the paper, we assume that some knowledge about the spillover effects is known.

Namely, assume that the full effect vector $\alpha$ is a linear transformation of some unknown parameter $\gamma\in \mathbb{R}^k$, i.e., $\alpha = A\gamma$, where $\alpha\in \mathbb R^{N\times 1}$, $A\in \mathbb R^{N\times k}$, and $\gamma\in \mathbb R^{k\times 1}$. Here, $A$ contains the information about the spillover structure, and $\gamma$ contains sufficient information about the magnitude of spillover effects. Typically, $\gamma$ has fewer dimensions than $\alpha$ does ($k<N$). Note that the linearity is not particularly restrictive of actual spillover structures, but rather allows the researcher to incorporate available information about the spillover patterns to facilitate the estimation. Consider the extreme example where we impose no restrictions on the possible spillovers. This is a special case of $\alpha=A\gamma$ with $A$ being the identity matrix and $\gamma=\alpha$.

Our leading example is given as follows, where we specify the range of spillover units without restricting their magnitude:

eg(Limited range) Assume that the spillover effect is likely to take place at some known locations, but not at other locations, while the sizes of spillover effects are allowed to vary across those units. For example, assume there are potential spillovers at locations whose distance to unit $1$ is less than $\bar{d}$. Then, the treatment and spillover effect vector can be represented by $A\gamma$. WLOG order the units by increasing distance from unit $1$, and let $p$ be the number of units experiencing spillovers. Then, $\alpha=A\gamma = (\alpha_1,\alpha_{k_1},\dots,\alpha_{k_p},0_{1\times (N-p-1)})'$, where $$A=\begin{bmatrix} I_{1+p}\\ 0_{(N-p-1)\times (1+p)} \end{bmatrix} \text{ and } \gamma=\begin{bmatrix} \alpha_1, \alpha_{k_1},\dots,\alpha_{k_p} \end{bmatrix}'. $$ Thus, the units indexed $2,...,(p+1)$ each experience their own size spillover effect.

The assumptions in Example (ref) are often plausible. Later we will illustrate how contextual knowledge can be used to inform the structure of the spilovers through our empirical application in Section (ref). If misspecification of the spillover structure is a concern, one can choose an $A$ matrix that incorporates more potential spillovers, i.e., a larger $p$. The consequences of misspecification are discussed in Section (ref).

Besides Example (ref), we will explore other possibilities, including the case where some units experience equal spillover effects, as discussed in Example (ref) in Section (ref), and the case where spillover effects decrease exponentially as the distance from the main treated unit increases, as in Example (ref) in Section (ref).

To alleviate concern regarding structure misspecification, we propose a test for correct specification ($\kappa_A$-statistic) in Section (ref). In Section (ref), we extend this test to settings with multiple post-treatment periods, and give conditions under which consistent structure estimation is possible.

An Asymptotically Unbiased Estimator

SCM without spillover effects

We consider a version of SCM proposed in Ferman2016, which starts with obtaining the synthetic control weights by solving the optimization problem

equation[equation omitted — 161 chars of source]

where $Y_t=(y_{1,t},\dots,y_{N,t})'$ and $W^{(1)}=\{(w_1,\dots,w_N)'\in\mathbb{R}^{N}_+: w_1=0, \sum_{j=2}^Nw_j=1 \}$. This restricts the estimation weights such that they sum to 1, are all non-negative, and the same-unit weight is 0, as well as including an intercept term. An estimator of the treatment effect $\alpha_{1}$ is given by $ \widehat{\alpha}_{1}=y_{1,T+1}-(\widehat{a}+Y_{T+1}'\widehat{b}),$ i.e., the counterfactual value $y_{1,T+1}(0)$ is approximated by $\widehat{a}+Y_{T+1}'\widehat{b}$. Here, we do not restrict the intercept but require the other coefficients to be positive and sum up to one. See Doudchenko2016 for a discussion of other choices of restriction sets. The intercept term $a$ is important in our setting because it takes out the bias by recentering the estimator. Ferman2016 show that when the pre-treatment is imperfect and there is no selection on time-varying unobservables, this estimator is asymptotically unbiased, while the original SCM as in Abadie2010 has bias.

The proposed estimator under spillovers

In order to back out the spillover effects, we first define the individual synthetic control weights and their limits. Namely, for each $i$, let the individual-specific synthetic control weights (and the intercept) be

equation[equation omitted — 189 chars of source]

where $W^{(i)}=\{(w_1,\dots,w_N)'\in \mathbb{R}^{N}_+: w_i=0, \sum_{j=1}^Nw_j=1 \}$. These are the same restrictions as above (sum-to-one, non-negativity, same-weight is 0). Then, let the probability limit of the intercept and weights be

equation*[equation* omitted — 127 chars of source]

and we only consider cases where they are well-defined. We show later by Lemma (ref) in Section (ref) that $a_i$ and $b_i$ exist for each $i$ in factor models with stationary or cointegrated common factors. In general, $a_i$ and $b_i$ do not coincide with the weights that reconstruct the factor loadings Ferman2016.

For each $(i,t)$, define the specification error by

equation[equation omitted — 86 chars of source]

Define $a=(a_1,\dots,a_N)'$ and $B=(b_1,\dots,b_N)'$. Stacking Equation (ref) for all $i$'s gives $u_t=Y_t(0)-(a+BY_t(0)),$ where $u_t=(u_{1,t},\dots,u_{N,t})'$. This gives us Equation (ref). Since $\alpha = Y_{T+1}(1)-Y_{T+1}(0)$, we have at period $T+1$

equation[equation omitted — 73 chars of source]

where $Y_{T+1}=(y_{1,T+1},\dots,y_{N,T+1})'$. We will use this equation to estimate the whole effect vector $\alpha$.

We form estimators for $(a,B)$ using synthetic control methods as in (ref). We do that for each $i=1,\dots, N$, as if each $i$ is the treated unit and other units are controls. Then, the estimators for $a$ and $B$ are $\widehat{a}=(\widehat{a}_1,\dots,\widehat{a}_N)'$ and $\widehat{B}=(\widehat{b}_1,\dots,\widehat{b}_N)'$, respectively. Define $M=(I-B)'(I-B)$ and let $\widehat{M}=(I-\widehat{B})'(I-\widehat{B})$ be an estimator for $M$.

Recall that the effect vector is $\alpha=A\gamma$. Let an estimator of $\gamma$ be such that

equation[equation omitted — 230 chars of source]

Note that the first-order condition implies $ A'(I-\widehat{B})'\widehat u_{T+1}=0,$ where $\widehat u_{T+1} = (I-\widehat B) (Y_{T+1}-\widehat\alpha)-\widehat a$, i.e., it requires that some weighted sum of the residuals be zero. Under that condition, the treatment and spillover effect vector $\alpha$ can be estimated by $\widehat{\alpha}=A\widehat{\gamma}$. One way to interpret this estimator is that the proposed $\widehat \alpha$ is the most consistent with the (constrained) linear model learned using pre-treatment data.

assum(a) $\{ u_t\}_{t\ge 1}$ is stationary, and has mean zero; (b) $\Vert \widehat{a}-a\Vert=o_p(1)$, $\Vert \widehat{B}-B\Vert =o_p(1)$; (c) $\Vert (\widehat{B}-B)Y_{T+1}(0)\Vert =o_p(1)$; (d) $A'MA$ is non-singular.

Part (a) generally requires that there is no regime shift or structural break. It also corresponds to the assumption of no selection on time-varying unobservables as in Ferman2016. Part (b) and (c) requires that there are at least a moderate number of pre-treatment periods so that the synthetic control weights are well-estimated. We show later that Assumption (ref)(a)-(c) are satisfied in factor models with either stationary or cointegrated common factors. We will discuss Part (d) later in Section (ref).

thmSuppose Assumption (ref) holds. Then, $\widehat{\alpha}-(\alpha+Gu_{T+1})\rightarrow_p0$ as $T\rightarrow \infty$, where $G=A(A'MA)^{-1}A'(I-B)'$. Moreover, $E[Gu_{T+1}]=0$.

The structure of the limiting distribution is similar to the case in Ferman2016, as it is inconsistent but asymptotically unbiased. Note that consistent estimators are impossible because only one post-treatment period with one treated unit is available, so the error term for that one unit and period does not shrink in any limit we consider.

The factor model as an example

Factor models are often used to justify the usage of synthetic control methods. See cao_principal_2020 for a review. Here we show that our assumptions are satisfied by factor models with stationary and cointegrated common factors. We follow Ferman2016 and consider a factor model such that for $i=1,\dots,N$ and $t=1,\dots,T+1$,

equation[equation omitted — 96 chars of source]

where $\lambda_t$ is $F$-dimensional common factors with a fixed $F$, and $\varepsilon_{i,t}$ is the noise that is uncorrelated with $\lambda_t$. For notation simplicity, we write $Y_t(0) = (y_{1,t}(0),\dots,y_{N,t}(0))'$, $Y_t = (y_{1,t},\dots,y_{N,t})'$, and $\varepsilon_t=(\varepsilon_{1,t},\dots,\varepsilon_{n,t})'$.

We focus on two sets of conditions in our discussion.

conditionst[model with stationary common factors] Assume $\{(\lambda_t,\varepsilon_t)\}_{t\ge 1}$ is stationary, ergodic for the first and second moments, and has finite $(2+\delta)$-moment for some $\delta>0$. The time fixed effect $\eta_t$ satisfies $\Vert (\widehat B-B)\eta_{T+1}\Vert=o_p(1)$. Assume $\Omega_y=Cov[(\lambda_t'\mu_1+\varepsilon_{1,t},\dots,\lambda_t'\mu_N+\varepsilon_{N,t})']$ is positive definite.

\noindentRemarks: 1. Stationarity implies that there is no selection on time-varying unobservables. The time fixed effect $\eta_t$ is allowed to be non-stationary or to explode, but not faster than the rate of $\widehat B$. We can do this because the time fixed effect always gets canceled out, given the simplex restriction of the synthetic control methods.

2. We show in the proof of Lemma 1 that in this case $a_i=E[y_{i,1}(0)-Y_{1}(0)'b_i]$ and $b_i=\underset{w\in W^{(i)}}{\arg\min}\ (w-e_i)'\Omega_y (w-e_i)$, where $e_i$ is a unit vector with one at the $i$-th entry and zeros everywhere else, and $W^{(i)}=\{(w_1,\dots,w_N)\in \mathbb{R}_+^N: w_i=0,\sum_{j\ne i }w_j=1 \}$. Note that in general $b_i$ does not recover the factor structure, because $\mu_i\ne (\mu_1,\dots,\mu_N)b_i$ in general.

3. We do not impose any restriction on the factor loadings $\{\mu_i \}_{i=1}^N$ except for $\Omega_y$ being positive definite. In the stationary case, the key for the treatment estimator to be asymptotically unbiased and the test proposed below to be valid is to include an intercept in the optimization problem (ref).

conditionns[model with cointegrated $\mathcal{I}(1)$ common factors] Rewrite Equation (ref) as $ y_{i,t}(0)= (\lambda_t^1)'\mu_i^1+ (\lambda_t^0)'\mu_i^0+\varepsilon_{i,t}, $ and $\eta_t$ can be either in $\lambda_t^1$ or $\lambda_t^0$. Assume $\{(\lambda_t^0,\varepsilon_t) \}_{t\ge 1}$ is stationary, ergodic for the first and second moments, and has a finite $4$-th moment. Without loss of generality, $E[\varepsilon_{i,t}]=0$. Assume $\{\lambda_t^1\}_{t\ge 1}$ is $\mathcal{I}(1)$. Further assume for each $i$, $y_{i,t}(0)$ is such that weak convergence holds for $T^{-1/2}y_{i,[rT]}(0)\Rightarrow \nu_i(r)$, where $\Rightarrow$ is weak convergence and process $\nu_i(r)$ is defined on $[0,1]$ and has bounded continuous sample path almost surely. For each $i$, let $W^{(i)}=\{(w_1,\dots,w_N)\in \mathbb{R}_+^N:w_i=0,\sum_{j\ne i}w_j=1 \}$. Assume for each $i$, there exists $w^{(i)}\in W^{(i)}$ such that $\mu_i^1=\sum_{j=1}^Nw^{(i)}_j\mu_j^1$. That is, $(w^{(i)}-e_i)$ is a cointegrating vector for $Y_t(0)$, where $e_i$ is a unit vector with $i$-th entry being one and zeros everywhere else.

Note that Condition CO puts restrictions on the factor loadings. The restrictions are similar to those in Ferman2016.

The relevance of the factor model is given by the following lemma:

lemmaSuppose $A'MA$ is non-singular. Then, either Condition ST or Condition CO implies Assumption (ref).

Thus, results derived in Theorem 1 apply to factors models with Condition ST or Condition CO.

Invertibility of $A'MA$

In Assumption (ref)(d), we require that the matrix \(A'MA\) must be invertible. This section explores the implications of this key assumption, providing both examples and counterexamples. In addition, we discuss the connection to a network framework by giving conditions under which \(I-B\) has rank \(N-1\), where a single pure donor unit is sufficient to ensure the invertibility of \(A'MA\). Extending beyond this section, we discuss how the invertibility of $A'MA$ parallels the full rank assumption of the design matrix in a linear model in Section (ref).

Intuition and a toy example

First, note the invertibility of $A'MA$ is testable in principle. Recall that $M=(I-B)'(I-B)$, so that $A'MA = A'(I-B)'(I-B)A$. $A$ is defined by the econometrician ahead of time. We can consistently estimate $B$ so the data informs us of the validity of this assumption.

To understand this assumption better, we replace $\alpha$ by $A\gamma$ in Equation (ref) and have

equation[equation omitted — 74 chars of source]

Equation (ref) is the key to learning $\alpha$. Under mild regularity conditions, $a$ and $B$ are identified from the model and learned by the synthetic control method. We do not observe $u_{T+1}$, but the distribution of $u_{T+1}$ can be learned using pre-treatment data under stationarity of $\{u_t \}_{t\ge 1}$. Therefore, if $A'MA$ is non-singular, or equivalently, $(I-B)A$ has full rank, we can form an estimator of $\gamma$ whose limiting distribution is identified by multiplying both sides of Equation (ref) by $(A'MA)^{-1}A'(I-B)'$. Note that we do not point-identify $\gamma$ or $\alpha$, because we have only one observation of the outcome in the post-treatment period.

The following example is useful in illustrating the invertibility assumption:

eg(Homogeneous spillovers) Assume that a subset of control units, but not all of them, are equally affected by the spillover effects, i.e. $\alpha=A\gamma = (\alpha_1,b,\dots, b,0_{1\times l})'$, where $$ A=\begin{bmatrix} 1&0\\0_{k\times 1}&1_{k\times 1}\\ 0_{l\times 1} & 0_{l\times 1} \end{bmatrix}, \ \ \gamma=\begin{bmatrix} \alpha_1\\b \end{bmatrix}, $$ $\alpha_1$ is the treatment effect, and $b$ is the homogeneous spillover effect.

For illustration, consider a three-unit case, where unit 1 is treated. WLOG, let the synthetic control weight matrix $B$ be $$B=

bmatrix[bmatrix omitted — 73 chars of source]

.$$ Suppose the researcher first assumes unit 2 and 3 are equally exposed to the spillover effects. That is, we have $\alpha=(\gamma_1,\gamma_2,\gamma_2)'$ with $$A_1=

bmatrix[bmatrix omitted — 44 chars of source]

and \gamma=

bmatrix[bmatrix omitted — 40 chars of source]

, and thus (I-B)A_1=

bmatrix[bmatrix omitted — 50 chars of source]

, $$ leading to a non-invertible $A'MA$. Intuitively, the problem here is there are two control observations we want to take a difference from, to determine the treatment effect.

If they instead assume only one of the controls is exposed to the spillover effects, $A'MA$ is non-singular in general. In this case, we have $\alpha=(\gamma_1,\gamma_2,0)'$ with $$A_2=

bmatrix[bmatrix omitted — 44 chars of source]

, \gamma=

bmatrix[bmatrix omitted — 40 chars of source]

, and thus (I-B)A_2=

bmatrix[bmatrix omitted — 56 chars of source]

.$$ It can be shown that $(I-B)A_2$ always has full rank for $(w_1,w_2,w_3)\in [0,1]^3$.

This applies to more general settings. That is, if all controls are equally hit by the spillover effects, then $(I-B)A$ does not have full rank and $A'MA$ is non-invertible. Allowing a few units to be exempt from the spillover effects makes $(I-B)A$ have full rank in general.

Results on rank of $I-B$

We extend the main idea of the previous section to the more interesting case of Example (ref), where the range of spillover effects is bounded but their magnitudes can vary. In this case, the matrix $(I-B)A$ is obtained by eliminating columns corresponding to units that are neither treated nor exposed to spillovers. The invertibility assumption is more plausible when a moderate number of columns are eliminated from $(I-B)$, suggesting that only a limited range of units are exposed to spillovers.

We explore the properties of $I-B$ here. Note that $I-B$ has a maximum rank of $N-1$, implying a necessary condition for invertibility in Example (ref) is having at least one pure donor. For the sufficient condition, we show its connections to network frameworks, which is formalized below.

Let $\mathcal{X} = \{1, \dots, N\}$ be the set of nodes corresponding to units. The matrix $B$ can then define an adjacency matrix, with the directed edge $e: \mathcal{X}^2 \rightarrow \{0,1\}$ defined by $e(i,j) = \mathbbm{1}\{b_{i,j} > 0\}$. As a result, $(\mathcal{X}, e)$ defines a directed network. Let a path from $i_1 \in \mathcal{X}$ to $i_k \in \mathcal{X}$ be a sequence of nodes $p(i_1, i_k) = (i_1, \dots, i_k)$ where $e(i_1, i_2) = e(i_2, i_3) = \dots = e(i_{k-1}, i_k) = 1$.

propIf for each ordered pair $(i, j) \in \mathcal{X}^2$ there exists a path $p(i,j)$, then $I-B$ has rank $N-1$.

This result helps us understand when the invertibility of $A'MA$ is satified. The connectedness assumption excludes approximately two types of networks. First, it rules out disconnected networks, which imply that $I-B$ is a block-diagonal matrix. This results in at least two groups of nodes, each of which is capable of predicting outcomes within the group without outside data, e.g., \[ I-B=

bmatrix[bmatrix omitted — 84 chars of source]

. \] Here, the blocks without the parameter of interest are redundant. Normally, the rank of $I-B$ is $N$ minus the number of blocks (so $N-1$ in a connected network with only one block). Second, it rules out networks with sinks, where a few units can predict all other units, e.g., \[ I-B=

bmatrix[bmatrix omitted — 84 chars of source]

. \] This is a case where some units can be predicted, but not predicting relevant units such as the treated one. Again, those units are redundant in learning the parameter of interest.

Statistical Inference

In this section, we discuss formal results on inference. At a high level, our test uses pre-treatment data to form the null distribution of a pre-specified post-treatment quantity. We only consider cases with imperfect pre-treatment fit to facilitate the identification of the null distribution. In Section (ref), we consider the case without spillover effects, and state the assumptions under which Andrews' $P$ test Andrews2003 is valid. This result is of independent interest. In Section (ref), we generalize $P$ test to cases where spillover effects cannot be ignored, and allow for a more general set of hypotheses.

Cases without spillover effects

Suppose for now that there are no spillover effects ($\alpha_2=\dots=\alpha_N=0$). We want to test for the existence of treatment effect on unit 1. The null and alternative hypotheses of interest are $H_0:\alpha_1 = 0$ and $H_1:\alpha_1\ne 0$, respectively. The test procedure we consider here is the end-of-sample instability test ($P$-test) in Andrews2003. The usage of Andrews' test in the context of synthetic control methods is mentioned in Ferman2018, where they focus on the difference-in-differences estimator. We formalize this idea and derive conditions under which Andrews' test delivers valid inference results.

We assume that $\alpha_1$ is not a function of $T$ under $H_1$. That is, we consider fixed, not local, alternatives, as in Andrews2003 and Andrews2006. Specifically, $\alpha_1$ does not change as $T$ grows, which facilitates our analysis of the test statistic under $H_1$.

Now we translate our hypothesis into the linear formulation considered in Andrews2003. Namely, we have $ y_{1,t}=a_1+(a_1^*-a_1)\mathbbm 1\{ t=T+1\}+Y_t'b_1+u_{1,t}.$ A non-zero treatment effect is equivalent to a shift in the intercept $a_1$ (or equivalently, change of the distribution of $u_{1,t}$, at $t=T+1$). The null and alternative hypotheses become $H_0:a_1^*=a_1$ and $ H_1:a_1^*\ne a_1$, respectively. Let the synthetic control regression residuals be $\widehat{u}_{1,t}=y_{1,t}-\widehat{a}_1-Y_t'\widehat{b}_1$. If there is no treatment effect, the distribution of $\widehat{u}_{1,T+1}$ should be asymptotically equivalent to that of $\widehat{u}_{1,t}$ for $t\le T$. Using this idea, define the test statistic by $ P=\widehat{u}_{1,T+1}^2.$ For notational simplicity, let $\widehat{\beta}_1=(\widehat{a}_1,\widehat{b}_1')'$ and $x_t=(1,Y_t')'$. For any $\beta\in\mathbb{R}^{N+1}$, define $P_t(\beta)=(y_{1,t}-x_t'\beta)^2.$ Then, $P=(y_{1,T+1}-x_{T+1}'\widehat{\beta}_1)^2=P_{T+1}(\widehat{\beta}_1)$. The pre-treatment counterparts are defined by $P_t=P_t(\widehat{\beta}_1^{(t)})$, where $\widehat{\beta}_1^{(t)}=\widehat{\beta}_1$ for each $t$. For a significance level of $\tau$, we reject $H_0$ if $P$ is larger than the $(1-\tau)$-quantile of $\{P_t \}_{t=1}^T$. See Andrews2003 and Andrews2006 for other ways of constructing $P_t$ such as the leave-one-out method.

To establish the validity of the proposed test, let $P_\infty$ be a random variable with the same distribution as $P_{T+1}(\beta_1)$ with $\beta_1=(a_1,b_1')'$. Define the empirical CDF of $\{P_t \}_{t=1}^T$ by $ \widehat{F}_{P,T}(x)=T^{-1}\sum_{t=1}^{T}\mathbbm{1}\{P_t\le x\},$ and let $F_P(x)$ be the distribution function of $P_1(\beta_1)$. We reject $H_0$ if $P > \widehat{q}_{P,1-\tau}$, where $\widehat{q}_{P,1-\tau}=\inf \{x\in \mathbb{R}:\widehat{F}_{P,T}(x)\ge 1-\tau \}$. Finally, let $q_{P,1-\tau}$ be the $(1-\tau)$-quantile of $P_1(\beta_1)$.

assum(a) $\{ u_t\}_{t\ge 1}$ are stationary, ergodic, and have mean zero; \\ (b) $E[|u_t|]<\infty$; \\ (c) $\exists$ positive definite $\{C_T \}_{T\ge 1}$ such that $\max_{t\le T+1}\Vert C_T^{-1} x_t\Vert=O_p(1) $;\\ (d) $\Vert C_T(\widehat{\beta}_1-\beta_1)\Vert=o_p(1)$, and $\max_{t=1,\dots,T}\Vert C_T(\widehat{\beta}_1^{(t)}-\beta_1)\Vert=o_p(1)$; \\ (e) The distribution function of $P_1(\beta_1)$ is continuous and increasing at its $(1-\tau)$-quantile.

Assumption (ref) is similar to those in Andrews2003. Part (a) does not allow for a structural break. Part (b) and (c) are moment conditions. Part (d) requires at least a moderate number of pre-treatment periods so that the synthetic control weights are well-estimated. Part (e) generally requires that $u_t$ follows a continuous distribution.

thmSuppose Assumption (ref) holds. Then, as $T\rightarrow \infty$, \\ (a) $P\rightarrow_d P_\infty$ under $H_0$ and $H_1$;\\ (b) $\widehat{F}_{P,T}(x)\rightarrow_p F_P(x)$ for all $x$ in a neighborhood of $q_{P,1-\tau}$ under $H_0$ and $H_1$; \\ (c) $\widehat{q}_{P,1-\tau}\rightarrow_p q_{P,1-\tau} $ under $H_0$ and $H_1$; \\ (d) $\Pr(P>\widehat{q}_{P,1-\tau})\rightarrow \tau$ under $H_0$.

Theorem (ref) states that the distribution of our test statistic $P$ can be approximated by the empirical distribution of $\{ P_t\}_{t=1}^T$. Specifically, Part (d) shows that the proposed test is asymptotically valid in the sense that the rejection probability under the null is convergent to the nominal level.

Combining this result with the Dominated Convergence Theorem, we have that under the alternative hypothesis \(H_1\), the rejection probability \(\Pr(P > \widehat{q}_{P, 1-\tau})\) converges to \(\Pr((u_{1,1} + \alpha_1)^2 > q_{P, 1-\tau})\), where \(q_{P, 1-\tau}\) is the \((1-\tau)\)-quantile of \(u_{1,1}^2\). Since the distribution of \(u_{1,1}\) can be estimated from the pre-treatment data, this result provides a practical method to approximate the power curves of the test.

We also show the relevance of the factor model in this context by the following lemma:

lemmaSuppose the distribution function of $P_1(\beta_1)$ is continuous and increasing at its $(1-\tau)$-quantile. Then, either Condition ST or CO implies Assumption (ref).

Cases with spillover effects

In this section, we generalize Section (ref) to cases allowing for non-zero spillover effects. We propose a testing procedure that is based on Andrews' $P$-test and accounts for the spillover effect. The null and alternative hypotheses we consider are $H_0:C\alpha=d$ and $H_1:C\alpha\ne d$, with known $C$ and $d$. For example, we want to test for the hypothesis that there is no treatment effect at the treated unit (unit 1), then we let $C=(1,0,0,\dots,0)\in \mathbb{R}^{1\times N}$ and $d=0$. This effectively makes Section (ref) a special case of our test, although Theorem 2 has slightly stronger results than Theorem 3 does. If we want to test that there is a spillover, then we can let $C=[0_{(N-1)\times 1} \ I_{N-1}]\in \mathbb{R}^{(N-1)\times N}$ and $d=(0,\dots,0)'\in \mathbb{R}^{(N-1)\times 1}$.

The test statistic we consider here is $P=(C\widehat{\alpha}-d)'W_T(C\widehat{\alpha}-d)$ for some weighting matrix $W_T\rightarrow_p W$. Recall $G=A(A'MA)^{-1}A'(I-B)$ and can be consistently estimated by $\widehat{G}=A(A'\widehat{M}A)^{-1}A'(I-\widehat{B})$ if $\widehat{B}\rightarrow_p B$. By Theorem (ref), $P$ is asymptotically equivalent to $u_{T+1}'G'C'W CGu_{T+1}$. To construct critical values, define

equation*[equation* omitted — 102 chars of source]

and its population version $P_t(\theta)=(Y_t-\theta x_t)'{G}'C'WC{G}(Y_t-\theta x_t)$, for some $\theta\in \mathbb{R}^{N\times (N+1)}$, $x_t=(1,Y_t')'$, and $\widehat{G}=A(A'\widehat{M}A)^{-1}A'(I-\widehat{B})'$. Let $\widehat{P}_t=\widehat{P}_t(\widehat{\theta}^{(t)})$, where $\widehat{\theta}^{(t)}=\widehat{\theta}$ for each $t$. For a significance level of $\tau$, we reject $H_0$ if $P$ is larger than the $(1-\tau)$-quantile of $\{\widehat{P}_t \}_{t=1}^T$.

To establish the validity of the proposed test, let $P_\infty=P_1(\theta_0)$ for $\theta_0=[a\ B]$. Define $\widehat{F}_{P,T}(x)=T^{-1}\sum_{t=1}^{T}\mathbbm{1}\{\widehat{P}_t\le x\},$ and let $F_P(x)$ be the distribution function of $P_\infty$. Finally, let $\widehat{q}_{P,1-\tau}=\inf \{x\in \mathbb{R}:\widehat{F}_{P,T}(x)\ge 1-\tau \}$, and $q_{P,1-\tau}$ be the $(1-\tau)$-quantile of $P_\infty$. The assumptions and validity of the proposed testing procedure are given as follows.

assum(a) Assumption 1 holds; (b) $\{ u_t\}_{t\ge 1}$ is ergodic and $E[\Vert u_t\Vert]<\infty$; (c) There exists a non-random sequence of positive definite matrices $\{D_T \}_{T\ge1}$ such that $\max_{t\le T+1}\Vert D_T^{-1}x_t\Vert =O_p(1)$; (d) $\Vert (\widehat{\theta}-\theta_0)D_T\Vert_F=o_p(1)$, and $\max_{t=1,\dots,T}\Vert (\widehat{\theta}^{(t)}-\theta_0)D_T\Vert_F=o_p(1)$, where $\Vert\cdot\Vert_F$ is the Frobenius norm; (e) The distribution function of $P_1(\theta_0)$ is continuous and increasing at its $(1-\tau)$-quantile; (f) $W_T\rightarrow_p W$ as $T\rightarrow \infty$.

Assumption (ref)(b)-(f) are similar to Assumption (ref) as well as those in Andrews2003.

thmSuppose Assumption (ref) holds. Then, under $H_0$, as $T\rightarrow \infty$, \\ (a) $P\rightarrow_d P_\infty$,\\ (b) $\widehat{F}_{P,T}(x)\rightarrow_p F_P(x)$ for all $x$ in a neighborhood of $q_{P,1-\tau}$, \\ (c) $\widehat{q}_{P,1-\tau}\rightarrow_p q_{P,1-\tau} $ , \\ (d) $\Pr(P>\widehat{q}_{P,1-\tau})\rightarrow \tau$.

Just like Theorem (ref), Theorem (ref) shows that we can approximate the null distribution of $P$ using its pre-treatment counterparts. Part (d) shows the asymptotic validity of the test proposed in this section.

Similar to Theorem (ref), we can extend the results to derive power curves. Under the alternative hypothesis \(H_1\), $\Pr(P > \widehat{q}_{P,1-\tau}) \rightarrow \Pr\left((u_1 + \alpha)' G' C' W C G (u_1 + \alpha) > q_{P,1-\tau}\right),$ where \(q_{P,1-\tau}\) is the \((1-\tau)\)-quantile of \(u_1' G' C' W C G u_1\). The right hand side is estimable from the pre-treatment data, allowing us to approximate the power curves for the test.

We show the relevance of the factor model in this context by the following lemma:

lemmaSuppose that $A'MA$ is non-singular and the distribution function of $P_1(\theta_0)$ is continuous and increasing at its $(1-\tau)$-quantile. Then, Assumption (ref) is satisfied if either of these holds: (i) Condition ST with $W_T=I$ or $W_T=(C\widehat{G}(T^{-1}\sum_{t=1}^T\widehat{u}_t\widehat{u}_t')\widehat{G}'C')^{-1} $; (ii) Condition CO with $W_T=I$.

Discussion

Structure misspecification

In this section, we first characterize the bias resulting from spillover structure misspecification and demonstrate that the misspecification bias in our proposed method is linear in the overlooked spillover effects. In contrast, the bias in the usual synthetic control method depends on all spillover effects. Subsequently, we introduce a statistic, $\kappa_A$, to assess spillover specification and propose an associated test.

Misspecification bias

We follow Example (ref) and assume that the spillover structure specifies the range of exposed units. Assume the effect vector is $\alpha=(\alpha_1,\dots,\alpha_{k},\dots,\alpha_{k+p},0,\dots,0)'$. The (asymptotic) bias for the usual synthetic control method (SCM) is

equation[equation omitted — 103 chars of source]

When applying the proposed method in the paper, suppose we include only units 2 through $k$ in the spillover structure. The bias of our estimator for the entire vector $\alpha$ then becomes

equation[equation omitted — 132 chars of source]

Consequently, the bias for the treatment effect estimator is a linear combination of the omitted spillovers $$\delta_{1,SP}=\sum_{i=k+1}^{k+p} c_i \alpha_i,$$ where each $c_i$ is determined by the first entry of Equation (ref).

For concreteness, consider the case where units 2 and 3 are affected by spillover effects, but unit 3 is omitted in our method. The biases are then given by $$\delta_{1,SCM}=-b_{1,2}\alpha_2 - b_{1,3}\alpha_3, \quad \text{and} \quad \delta_{1,SP}=-\frac{\mathrm{det}\left( [\tilde{b}_1,\tilde{b}_2]'[\tilde{b}_2,\tilde{b}_3] \right) }{\mathrm{det}(A'MA)}\alpha_3, $$ where $\tilde{b}_i=e_i-(b_{1,i},b_{2,i},\dots,b_{N,i})'$ and $e_i$ is the unit vector with one in the $i$-th entry and zeros elsewhere. There is generally no guarantee that either $\delta_{1,SCM}$ or $\delta_{1,SP}$ will be smaller. We suggest that the researcher be conservative about choosing the structure. That is, if a certain unit is suspected to be affected by spillover effects, it should be included in the spillover structure.

$\kappa_A$-statistic

In addition to results from the previous section, we note that the goodness of fit is informative about the accuracy of the spillover structure. To see this, define $$\kappa_A = \Vert (I-\widehat{B})(Y_{T+1} - \widehat{\alpha}) - \widehat{a} \Vert$$ as a function of $A$, where $\widehat{\alpha}$ is the estimated effect vector under the assumption that $A$ correctly specifies the spillover effects. This statistic is useful in evaluating the correctness of spillover specifications, which is motivated by the observation that $\kappa_A$ is small when $A$ is correctly specified. We formalize this idea by the proposition below.

Given $A$, define $\Gamma_A = (I-B)A(A'(I-B)'(I-B)A)^{-1}A'(I-B)'$, the projection onto the span of columns of $(I-B)A$. The sample analog is defined as $\widehat{\Gamma}_A = (I-\widehat{B})A(A'(I-\widehat{B})'(I-\widehat{B})A)^{-1}A'(I-\widehat{B})'$. Andrews' test can be applied to test for correct spillover structure specification. Similar to Section (ref), we approximate the null distribution of $\kappa_A$ using $\{\Vert (I-\widehat \Gamma_A)\widehat u_t\Vert\}_{t=1}^T$. Define $\widehat{q}_{\kappa, 1-\tau}^A = \mathbbm{1}\{x \in \mathbb{R} : \widehat F_{\kappa, T}^A(x) \geq 1-\tau\},$ where $\widehat F_{\kappa, T}^A(x) = T^{-1} \sum_{t=1}^T \mathbbm{1}\{\Vert (I-\widehat \Gamma_A)\widehat u_t\Vert \leq x\}$. We reject the null hypothesis that $A$ correctly specifies the spillover effects if $\kappa_A > \widehat q_{\kappa, 1-\tau}^A$. Let $q_{\kappa, 1-\tau}^A$ be the $(1-\tau)$-quantile of $\Vert (I-\Gamma_A)u_1\Vert$.

propSuppose Assumption (ref) holds. Then, $\Pr(\kappa_A > \widehat q_{\kappa, 1-\tau}^A) \rightarrow \Pr(\Vert(I-\Gamma_A)u_{T+1} + (I-\Gamma_A)(I-B)\alpha\Vert \geq q_{\kappa, 1-\tau}^A)$. Specifically, when $A$ is a correct specification, $\Pr(\kappa_A > \widehat q_{\kappa, 1-\tau}^A) \rightarrow \tau$.

When $A$ is correctly specified, $(I-\Gamma_A)(I-B)\alpha = 0$, making $q_{\kappa, 1-\tau}^A$ precisely the $(1-\tau)$-quantile of $\Vert(I-\Gamma_A)u_{T+1}\Vert$. Conversely, $(I-\Gamma_A)(I-B)\alpha \neq 0$ generally when $A$ is not a correct specification, which enable the test to have power.

A heuristic usage of this statistic is to select $A$ by solving $\widehat{A} = \arg\min_{A \in \mathcal{A}} \kappa_A,$ where $\mathcal{A}$ is a set of potential spillover specifications. For instance, if $\mathcal{A}$ contains only two elements, we choose the $A$ corresponding to the smaller $\kappa_A$ between the two competing structures. It is important to note this data-dependent procedure may lead to model selection error, since only one post-treatment period is available. In Section (ref), we extend this method to multi-period settings, where consistent model selection becomes possible.

Comparison with the pure-donor estimator

In Example (ref), we consider a known spillover structure where the range of spillover is assumed to be known. In this setting, an alternative estimation approach other than ours is to use only the “pure donors,” which means excluding all units potentially exposed to spillovers and treating the remaining units as the sole control group. This section characterizes the (asymptotic) variance of our method and the pure donor approach. Additionally, we propose a sensitivity analysis to assess the robustness of both methods.

Variance comparison

If the spillover structure is correctly specified, both the pure-donor and our method are asymptotically unbiased. While intuitively, discarding information might seem unfavorable DiStefano2020, there is no guarantee that one method is superior to the other in terms of asymptotic variance. To illustrate this, we first note that a corollary of Theorem (ref) indicates that as $T$ approaches infinity, the variance of the estimator $\widehat{\alpha}$ is $ \lim_{T\rightarrow \infty} {Var}[\widehat{\alpha}] = G(I-B)\Omega_Y(I-B)'G', $ where $\Omega_Y=\lim_{T\rightarrow\infty} Var[Y_{T+1}]$ and $G=A(A'MA)^{-1}A'(I-B)'$. Let $g'$ be the first row of $G(I-B)$. Then, the variance of the first treatment effect estimator, $\widehat \alpha_1$, is

equation*[equation* omitted — 86 chars of source]

The minimization problem $ \min_{\tilde{a}\in \mathbb{R}, \tilde{b}\in \Delta_{PD}} \sum_{t=1}^T(y_{i,t}-\tilde{a}-Y_t\tilde{b}')^2$ defines the pure-donor estimator $(\widehat{a}_{1}^{PD}, \widehat{b}_{1}^{PD} )$, where $\Delta_{PD} = \{ w\in \mathbb{R}^N:w_1=0,w_i\ge 0\text{ for each }i,\sum w_i=1, w_i=0\text{ if }i\in S \}$ and $S$ is the set of indices with potential spillovers. Assuming the estimator converges to $(a_{1}^{PD},b_{1}^{PD})$, the variance is

equation*[equation* omitted — 113 chars of source]

This suggests that neither method is definitively superior. However, both $g$ and $b_{1,PD}$ are estimable. If we are willing to estimate a covariance model for $Y_{t}$ (for example, by assuming stable spatial correlation over time), we can estimate $\Omega_Y$. This information can be used to determine which estimator has a smaller variance, contingent on the validity of the covariance estimation.

A sensitivity analysis of misspecification bias

In this section, we provide another perspective on comparing the two methods, which involves spillover structure misspecification. Structure misspecification could be common in comparative case studies, especially with datasets of moderate sizes. As a result, it is crucial to evaluate the robustness of these methods under structure misspecification. To this end, we introduce a sensitivity analysis to assess the impact when the spillover structure is incorrectly specified.

Suppose the effect vector is $\alpha = (\alpha_1, \dots, \alpha_N)'$. We proceed with estimation assuming units $i = 2, \dots, k$ are exposed to spillovers, though in reality additional units in $i = k+1, \dots, N$ may also be contaminated. Let $p$ be the number of units mistakenly assumed not to be exposed to spillovers, and let $\bar{\alpha}$ denote the maximum possible spillover effects in absolute value. We want to assess the asymptotic bias for both methods, $\delta_{1,SP} = \sum_{i=k+1}^N c_i \alpha_i$ and $\delta_{1,PD} = -\sum_{i=k+1}^N b_{1,i}^{PD} \alpha_i$ as characterized in Section (ref). Since the identities of the neglected spillover units are unknown, the bias for each method is characterized by the identified set $\Delta_{1,SP}(\bar\alpha) = \{ \delta_{1,SP} : \alpha \in \mathcal{A}_p(\bar{\alpha}) \}$ and $\Delta_{1,PD}(\bar\alpha) = \{ \delta_{1,PD} : \alpha \in \mathcal{A}_p(\bar{\alpha}) \}$, where $\mathcal{A}_p(\bar\alpha) = \{ \alpha : \exists S \subset \{k+1, \dots, N\} \text{ s.t. } |S| = p, |\alpha_i| \leq \bar\alpha \text{ for } i \in S, \alpha_i = 0 \text{ for } i \notin S \}.$ The sensitivity analysis is conducted by plotting these identified sets against $\bar{\alpha}$ to compare the potential biases under varying levels of misspecification.

In the application section, we demonstrate this method and observe that the pure-donor method exhibits a significantly larger identified set of bias compared to our method (Figure (ref)). This discrepancy arises because the pure-donor method uses fewer control units, which results in more concentrated weights and consequently, a greater worst-case scenario bias.

Estimating the Effects of California's Proposition 99

figure[figure omitted — 1,713 chars of source]

To demonstrate our method, we apply it to the classic SCM example from Abadie2010 (Abadie2010, thereafter ADH), which looks at the effect of Proposition 99 on California cigarette consumption. In this section, we will walk through the results from our method, with interruptions to point out key features and issues. Furthermore, we compare the proposed method and the pure-donor one in a sensitivity analysis.

Set-up

Proposition 99 intended to disincentivize smoking, which was primarily achieved by introducing a \$$0.25$ tax on each pack of cigarettes. By measuring sales in California, ADH and others have attempted to determine the effect of the policy on cigarette consumption. However, traditional SCM is not guaranteed to produce an unbiased treatment effect estimator in the presence of spillover effects. In this tobacco control program example, we are concerned about two kinds of spillover effects. The first spillover is based on concerns about “leakage”. A common problem with cigarette taxes (and other vice taxes like gambling and alcohol) is that measured local consumption might fall as people move their purchasing behavior across legal boundaries, particularly in early years. In order to accommodate this, we allow for a spillover affecting states neighboring California, i.e., Arizona, Nevada, and Oregon. One might also think that there could be policy contamination whereby culturally close states also enact policies with similar targets. Our method can allow for this kind of spillover in our estimation. ADH took that type of problem into account, and 12 states which experienced legislative changes in the ensuing years were removed in that paper.

The data used is per capita cigarette consumption in the 50 states plus the District of Columbia running from 1970 to 2000. In 1989 California enacted Proposition 99, so all periods from 1989 onwards are considered post-treatment periods. We replicate this program evaluation using the method introduced in previous sections, allowing for possible spillover effects. We use the spillover structure as in Example (ref). That is, we allow for spillover effects in states that are geographically close or have experienced policy contamination, but not the others. The states that are considered exposed to spillovers include AK, AZ, DC, FL, HI, MA, MD, MI, NJ, NV, NY, OR, and WA. Those spillover effects are allowed to be different for different states and different time periods. We also perform hypothesis testing on both treatment effects and spillover effects. Since we have multiple post-treatment periods, we treat each post-treatment period as if it is the year right after the policy implementation, as outlined in Section (ref).

Main results

The results are shown in Figure (ref). The standard synthetic control method that is similar to Abadie2010 is indexed by SCM, and our method is SP. Figure (ref) shows the “synthetic California” and (ref) elaborates on this by specifically looking at the estimated treatment effects. Figure (ref) plots the estimated spillover effects for the three neighboring states of California. The error bars denote 95% confidence intervals that are built by inverting the test proposed in Section (ref). The scale of the error bars is visually larger than the amount of variation in the pre-period, which a consequence of adjusting for spillovers.

As Figure (ref) shows, our estimated consumption in the “synthetic California” does not differ substantially from what a standard SCM would predict, especially for later periods. However, our estimates of the first two post-treatment periods (1989 and 1990) are not significantly different from zero at a 95% level, in contrast with SCM. This difference may result from the over-estimation of the scale of the treatment effects by SCM in the presence of spillover effects. From the tests of spillover effects (shaded area of Figure (ref)), we see that likely there were substantial spillover effects. One potential cause of spillovers may be that consumers in California shifted their purchasing to the nearby states, Arizona, Nevada, and Oregon. Since similar laws with a tax increase on cigarettes were passed in Arizona and Oregon in 1994 and 1996, respectively, it is difficult to distinguish the spillover effects of Proposition 99 from anticipation effects as well as direct effects of their own laws. Nevada, however, did not pass any such laws in this period, so the spillover effect estimates for Nevada are more reliable. From Figure (ref), we observe that Nevada experienced significant spillover effects in the first two periods after the passage of Proposition 99 and mostly insignificant effects afterward, with the exception of 1997. This is consistent with our conjecture and may provide more evidence on how the effects of the policy have propagated.

The non-significant effects of policy right after the implementation could be explained by the addictive behavior of cigarette consumption. The persistence of cigarette consumption has been extensively studied and well-understood by both the rational addiction and the medical literature Baltagi2001,Baumeister2017,Becker1994,Benowitz1992,Christelis2011,Labeaga1999,Miura2019,Vleeming2002. Compared with SCM, our results are more consistent with an addiction story, that tobacco consumption is addictive and unlikely to drop immediately after the policy, but rather slowly transition to a lower equilibrium.

figure[figure omitted — 1,422 chars of source]

Sensitivity analysis

We use this application to illustrate the comparison between our proposed method and the pure-donor method using a sensitivity analysis discussed in Section (ref). Figure (ref) presents the comparison. The figure contains two panels that delineate the upper and lower bounds of the bias under spillover structure misspecification scenarios with $p=1$ and $p=2$. Here, $p$ is the number of incorrectly omitted units that are actually subject to spillover effects. The solid lines in both panels represent the bounds for the bias of our proposed method ($\delta_{1,SP}$), while the dashed lines represent that of the pure-donor method ($\delta_{1,PD}$). Key reference points, $\alpha_{max}^{CA}$ and $\alpha_{min}^{CA}$, represent the largest and smallest estimated treatment effects observed post-treatment, which are 3.71 (2 years post-treatment) and -18.96 (11 years post-treatment), respectively. The largest estimated spillover effect in absolute value, $\alpha_{max}^{NV}=26.86$, is recorded two years post-treatment in Nevada, which serves as a benchmark measure of the magnitude of spillover effects. The intersection of $\alpha_{max}^{CA}$ with the upper bound of $\delta_{1,SP}$ is also shown (17.07 and 9.66 for the $p=1$ and the $p=2$ case, respectively).

We have two key findings. First, we find that the pure-donor method has a much larger identified bias. The reason is that the pure-donor method uses fewer units and therefore has more concentrated weights. As a result, in the worst-case scenario where the spillover occurs at the missed unit with the largest weight, the treatment effect estimator has a larger bias because of the more concentrated weights. Second, the intersection of $\alpha_{max}$ and the upper bound of $\delta_{1,SP}$ can be interpreted as the smallest possible spillover effect needed to invalidate the non-zero treatment effect estimate. For example, when $p=1$, they intersect at 17.07. This means that in the extreme case where a spillover of 17.07 occurs at the missed control with the largest weight, our estimate should have been zero, but is not so due to the misspecification of the spillover structure. Of course, in the application we are more interested in the negative effect. We see that $\alpha_{min}^{CA}$ intersects the lower bound of $\delta_{1,SP}$ at a point much larger than $\alpha_{max}^{NV}$, suggesting the robustness of our results even allowing for some level of spillover structure misspecification.

Further Discussion

An example without pure donors

In this section, we explore a case where all control units are affected by the treatment, leaving no pure donors available. We provide an example of spillover structure $\alpha=A\gamma$ that satisfies this condition. Interestingly, this discussion shows a link to the continuous treatment scenario.\footnote{We thank an anonymous reviewer for providing this insight.}

eg(Exponential decay) Assume that the spillover effect shrinks as the geometric distance goes up. For $i=2,\dots, N$, $\alpha_i=b\exp(-d_i)$ where $d_i$ is the distance between unit $1$ and unit $i$, and $b$ is an unknown parameter of interest. Then, we have $\alpha = A\gamma =(\alpha_1,b\exp(-d_2),\dots, b\exp(-d_N))'$, where $$A=\begin{bmatrix} 1&0\\0&\exp(-d_2)\\\vdots&\vdots\\0&\exp(-d_N) \end{bmatrix} \text{ and } \gamma=\begin{bmatrix} \alpha_1\\b \end{bmatrix}. $$

Example (ref) introduces an interesting case where all control units are affected by spillover effects. This is a challenging scenario, and a sensible estimation strategy requires strong assumptions. As in Example (ref), one such assumption is that the spillover effects are linear in distance from the treated unit. All previously introduced estimation and inference strategies would work, given the assumptions are valid.

Another perspective on this example is its connection to a continuous treatment model. Namely, $\exp(-d_i)$ is a variable that summarizes the intensity of spillover effects. This links directly to the continuous or moderated treatment literature callaway_event-studies_2024,erten_trade_2023. In certain contexts, it may be more useful to focus on intensity as the primary variable of interest. In this scenario, no units are exempt from effects, meaning that the pure counterfactual outcome $y_{i,T+1}(0)$ does not appear for any $i$ and may not be particularly relevant. This situation leads to an alternative parameter of interest, which we discuss below.

In our formulation, \(\gamma = (\alpha_1, b)\), where \(\alpha_1\) is the treatment effect, defined by \(y_{1,T+1} - y_{1,T+1}(0)\), and \(b\) is a candidate parameter of interest. Specifically, \(b\) can be interpreted as the marginal contribution to the outcome associated with a one unit increase in the implied intensity metric, \(\exp(-d_i)\).

To illustrate this, consider the proposed estimator for \(\gamma\), \[ \widehat{\gamma} = (A'\widehat{M}A)^{-1}A'(I-\widehat{B})'((I-\widehat{B})Y_{T+1} - \widehat{a}), \] which is the OLS estimator for the linear model (in matrix formulation)

equation[equation omitted — 136 chars of source]

Equation (ref) corresponds to the linear model

equation[equation omitted — 106 chars of source]

where \(A\) is the data matrix of some independent variables and the \(i\)-th observation vector of $A$ is \((\mathbbm{1}\{i=1\}, \mathbbm{1}\{i \ne 1\}\exp(-d_i))'\). Thus, \(b\) is the coefficient associated with \(\mathbbm{1}\{i \ne 1\}\exp(-d_i)\) and has the standard interpretation of a coefficient in a linear model. Note that Equation (ref) can be interpreted as a residualized version of Equation (ref), although the residualization is through the linear model learned using pre-treatment data.

The above analysis implies that \(b\) has the usual good properties of a linear model coefficient, as well as the limitations associated with those, such as linearity, measurement error, etc. This formulation provides one perspective on what we are estimating. Of course, we face greater challenges than those typically encountered in a standard linear regression scenario, due to data limitations. Moreover, the invertibility of $A'MA$ is now an analog to the common linear model assumption that the design matrix has full rank in the residualized model (ref).

Statistical inference can be conducted similarly as before. Since \(\widehat{\gamma} - \gamma = (A'MA)^{-1}A(I-B)u_{T+1} + o_p(1)\), we can use the pre-treatment residuals to approximate the null distribution of \(\widehat{\gamma}\). For the hypothesis \(H_0: b = b_0\), let the test statistic be \(P = (\widehat{b} - b_0)^2\), where \(\widehat{b}\) is the second entry of \(\widehat{\gamma}\). Its null distribution can be approximated by \[ P_t = \left[(0,1)(A'\widehat{M}A)^{-1}A'(I-\widehat{B})\widehat{u}_{t}\right]^2, \] for \(t = 1, \dots, T\). Confidence intervals can be constructed by inverting the test.

table[table omitted — 707 chars of source]

For illustrative purposes, we apply this method to California's tobacco control policy as in Section (ref). We emphasize its illustrative nature due to the strong and debatable assumption that treatment intensity decreases exponentially. The summary statistics of the intensity metric $\exp(-d_i)$ are provided in Table (ref). The results are shown in Figure (ref), showing that the estimate $\widehat{b}$ is reasonably stable over time, with most estimates significantly larger than zero. This suggests that closer states suffer more from the spillover effects, leading to higher cigarette consumption increase.

figure[figure omitted — 224 chars of source]

Comparison to other testing procedures

\pgfmathdeclarefunction{gauss}{2}{ \pgfmathparse{1/(#2*sqrt(2*pi))*exp(-((x-#1)^2)/(2*#2^2))} }

figure[figure omitted — 2,695 chars of source]

In this section, we compare the proposed inference method with the existing one. When we allow for the existence of non-zero spillover effects, the existing testing procedures will have poor performance. Here we intuitively explain what happens to the placebo test as in Abadie2010 and Andrews' test as in Andrews2003 in the presence of spillover effects.

Suppose we want to test for the treatment effect being zero and are not aware of the spillover effects. Placebo test and Andrews' test are similar in the sense that they use data to form the null distribution of $u_{1,T+1}$ in order to perform hypothesis testing. The difference is that the placebo test exploits variations of $\{\widehat{u}_{i,T+1}\}_{i=1}^N$, while Andrews' test uses variations of $\{\widehat{u}_{1,t} \}_{t=1}^{T+1}$.

We look at the placebo test first. When there is no spillover effect, the distribution of $\widehat{u}_{1,T+1}$ and that of any element in $\{\widehat{u}_{i,T+1} \}_{i=2}^N$ coincide asymptotically. As shown in Figure (ref)(b), when there are positive spillover effects, we will underestimate the treatment effect, and the density function of $\widehat{u}_{1,T+1}$ moves to the left. At the same time, some of the control units shift to the right because of the positive spillovers, so density of $\{\widehat{u}_{i,T+1} \}_{i=2}^N$ moves to the right and gets wider. In terms of test performance, the shift of $\widehat{u}_{1,T+1}$ is offset by the wider density of $\{\widehat{u}_{i,T+1} \}_{i=2}^N$ (harder to reject $H_0$), which explains why the empirical sizes of placebo test often do not deviate too much from the nominal size, even in the presence of spillovers (see Table (ref)). In essence, the placebo test becomes much more conservative and has low power.

Now we consider Andrews' test. When there is no spillover effect, the distribution of $\widehat{u}_{1,T+1}$ and that of any element in $\{\widehat{u}_{1,t} \}_{t=1}^T$ coincide asymptotically. As shown in Figure (ref)(b), when there is a positive spillover effect, we underestimate the treatment effect, and the density function of $\widehat{u}_{1,T+1}$ shifts to the left, while the density of $\{\widehat{u}_{1,t} \}_{t=1}^T$ doesn't, since they are pre-treatment and the spillover only happens after the treatment. This results in an invalid test.

Although not the main focus of this paper, selection into treatment can be a threat to the placebo test in practice. For example, if a unit is more likely to be treated when its own outcome $y$ is higher than those of other units, then the placebo test tends to over-reject the zero treatment effect hypothesis, even without spillover effects. This form of selection is not a problem for Andrews' test and the test proposed in this paper, since they use the variation across time periods instead of different units.

figure[figure omitted — 2,410 chars of source]

Acknowledgments

We thank Max Farrell and Christian Hansen for their invaluable guidance. We are grateful to Tetsuya Kaji, Giovanni Mellace, Alexander Torgovitsky, Ruey Tsay, Yinchu Zhu, and other seminar participants at the the 2018 Midwest Econometrics Group Conference, the 2019 China Meeting of the Econometric Society, and the 2019 North America Summer Meeting of the Econometric Society. We thank Ziyao Wang for his excellent research assistantship.

\setcounter{section}{0} \setcounter{equation}{0} \setcounter{figure}{0} \setcounter{table}{0} \setcounter{prop}{0} \setcounter{assum}{0} \setcounter{eg}{0}

Extensions

An estimator with a smaller variance

In this section, we show that it is possible to form an estimator of $\alpha$ with variance lower than $\widehat{\alpha}$ proposed in Section (ref). The idea is to minimize $\Vert W^{1/2}\widehat{u}_{T+1}\Vert$ instead of $\Vert \widehat{u}_{T+1}\Vert$, where $W\in \mathbb{R}^N$ is some positive definite matrix and $\widehat{u}_{T+1}=(Y_{T+1}-\widehat{\alpha})-\widehat{a}-\widehat{B}(Y_{T+1}-\widehat{\alpha})$. The resulting estimator, as a function of $W$, is

align*[align* omitted — 228 chars of source]

where $\widehat{M}_W=(I-\widehat{B})'W(I-\widehat{B})$. The corresponding estimator for $\alpha$ is $\widehat{\alpha}_W=A\widehat{\gamma}_W$. In the spirit of GMM with an efficient weighting matrix, let $\Omega=Cov[u_1]$ and $W_T^e$ be a consistent estimator of $\Omega^{-1}$. Then an estimator of $\alpha $ with lower variance can be achieved by $\widehat{\alpha}^e=\widehat{\alpha}_{W_T^e}. $

Let $M_W=(I-B)'W(I-B)$, $G_W=A(A'M_WA)^{-1}A'(I-B)'W$ for some weighting matrix $W$, $W^e=\Omega^{-1}$, $M^e=M_{W^e}$, and $G^e=G_{W^e}$. Then, we have the following results.

propSuppose Assumption (ref) holds, $W_T$ is a consistent estimator for $W$, and $W_T^e$ is a consistent estimator for $W^e$. Then, $\widehat{\alpha}_{W_T}-(\alpha+G_Wu_{T+1})\rightarrow_p0$, and specifically, $\widehat{\alpha}^e-(\alpha+G^eu_{T+1})\rightarrow_p0$, as $T\rightarrow \infty$. Moreover, $(Cov[G_Wu_{T+1}]-Cov[G^eu_{T+1}])$ is positive semi-definite.

Proposition (ref) states that $\widehat{\alpha}^e$ always has a smaller asymptotic variance than $\widehat{\alpha}$. In practice, we need to estimate $\Omega$, and for that we would need a relatively large sample size (large $T$) and/or a reasonable covariance model to have a good approximation.

Multiple treated units

Our method readily extends to cases where multiple units are treated. In our setting, the treatment and spillover effects can be estimated in the same way, since the spillover effects can be interpreted as the indirect treatment effect. With a correctly specified structure matrix $A$, we can perform estimation and inference just as in previous sections. For example, suppose $N=4$, unit 1 and unit 2 are treated, unit 3 is affected by spillover effects, and unit 4 is neither treated nor exposed to spillover effects. Then we can specify $A=[I_3, 0_{3\times 1}]'$, and the resulting estimator $\widehat{\gamma}=(\widehat{\gamma}_1,\widehat{\gamma}_2,\widehat{\gamma}_3)'$ by (ref) is such that $\widehat{\gamma}_1$ and $\widehat{\gamma}_2$ are the treatment effect estimator for unit 1 and unit 2, respectively, and $\widehat{\gamma}_3$ is the spillover effect estimator for unit 3. Tests can be performed accordingly. If the researcher wants to test for the hypothesis that there are no spillover effects, the null is then $H_0:C\alpha=d$, where $C=(0,0,1,0)$ and $d=0$.

Multiple post-treatment time periods

The availability of multiple post-treatment periods facilitates a more flexible set of estimands, such as dynamic treatment effects, as illustrated in the empirical application section. More importantly, assuming a stable spillover structure over time, the spillover structure is in principle estimable Manresa2013-cr,de_paula_identifying_2023. In this section, we introduce estimation and inference strategies when multiple post-treatment periods are available and propose a procedure to choose the spillover structure, assuming a stable structure over time.

Suppose we have observations $\{y_{i,t}\}$ for $i=1,\dots,N$ and $t=1,\dots,T+T_1$. Treatment is received at $t=T+1$. The model becomes \[ Y_{t}=

casesY_{t}(0), & if t \le T,\\ Y_{t}(0) + \alpha_t, & otherwise.

\] Note that we do not allow for spillovers in time, meaning the treatment effect or spillover effects cannot affect future selves. For each $t=T+1,\dots,T+T_1$, we need to specify the spillover structure matrix $A_t$.\footnote{In the empirical application, we assume a stable structure such that $A_t = A$ for all $t$.} Then, an estimator of $\alpha_t$ is \[ \widehat{\alpha}_t = A_t(A_t'\widehat{M}A_t)^{-1}A_t'(I-\widehat{B})'((I-\widehat{B})Y_{t}-\widehat{a}). \] For each $t=T+1,\dots, T+T_1$, we can perform separate tests as introduced in previous sections.

To answer questions such as whether there is a spillover effect at all, we can extend Andrews' instability test discussed above. Consider the null hypothesis $H_0: C_t\alpha_t=d_t$ for $t=T+1,\dots,T+T_1$. Let $\widehat{P}_t$ be constructed as in Section (ref) for $t=1,\dots, T$. For $t=T+1,\dots,T+T_1$, let $\widehat{P}_{t}=(C_t\widehat{\alpha}_t-d_t)'W_T(C_t\widehat{\alpha}_t-d_t)$. For $t=1,\dots, T+1$, let $P^{(t)}=\sum_{s=0}^{T_1-1}\widehat{P}_{t+s}$, each of which contains information from period $t$ through $t+T_1-1$. The test statistic is then $P^{(T+1)}$, and we use its pre-treatment counterparts $\{P^{(t)} \}_{t=1}^T$ to form the null distribution.

If a stable spillover structure is assumed such that $A_t = A$ for all post-treatment periods, it becomes feasible to estimate the structure. Following Section (ref), given a predefined structure $A$, the test statistic is defined as \[ \kappa_A = \left\| \frac{1}{\sqrt{T_1}} \sum_{s=1}^{T_1} [(I - \widehat{B}) (Y_{T+s} - \widehat{\alpha}_{T+s}) - \widehat{a}] \right\|, \] where $\widehat{\alpha}_{T+s}$ depends on $A$. The selection of $A$ is based on minimizing $\kappa_A$: \[ \widehat{A} = {\arg\min}_{A \in \mathcal{A}} \kappa_A, \] where $\mathcal{A}$ represents a set of potential spillover structures. For instance, if the range of spillover effects is restricted as in Example (ref), $\mathcal{A}$ includes all matrices of the form $A = [I_{1+p}, 0_{(1+p) \times (N-p-1)}]'$ for some $p$. This approach is justified by the observation that $\kappa_A$ increases significantly when $A$ does not accurately represent the spillover structure. The proposition below formalizes this observation.

Given $A$, define $\Gamma_A = (I-B)A(A'(I-B)'(I-B)A)^{-1}A'(I-B)'$, the projection onto the span of columns of $(I-B)A$. Define the sample version accordingly by $\widehat \Gamma_A = (I-\widehat{B})A(A'(I-\widehat{B})'(I-\widehat{B})A)^{-1}A'(I-\widehat{B})'$. Define $\bar{\alpha}_{T_1} = T_1^{-1}\sum_{s=1}^{T_1}\alpha_{T+s}$ as the average effects over time.

assum(a) $\{ u_t\}_{t\ge 1}$ is stationary, has mean zero, and satisfies $\sum_{s=0}^\infty |E[u_t u_{t+s}]|<\infty$. (b) $\widehat \Gamma_A-\Gamma_A=o_p(1)$, $\sqrt{T_1}[(I-\widehat \Gamma_A)(I-\widehat{B})-(I-\Gamma_A)(I-B)]\bar{\alpha}_{T_1}=O_p(1)$; (c) $(I-\widehat \Gamma_A)(\widehat{B}-B) T_1^{-1/2}\sum_{s=1}^{T_1}Y_{T+s}(0)=O_p(1)$, $\sqrt{T_1}(I-\widehat \Gamma_A)(a-\widehat{a})=O_p(1)$; (d) $A'MA$ is non-singular.

This is a multiperiod counterpart of Assumption (ref). Parts (b) and (c) restrict the rate of $\widehat{B}$ and $\widehat{a}$. Under weak conditions, they imply $T$ and $T_1$ are of similar magnitude.

propUnder Assumption (ref), $\kappa_A=O_p(1)$ if and only if $\sqrt{T_1}(I-\Gamma_A)(I-B)\bar{\alpha}_{T_1}=O(1)$.

Consider the specification of model spillovers. If $A$ accurately captures the spillover effects, then $\sqrt{T_1}(I-\Gamma_A)(I-B)\bar{\alpha}_{T_1}=0$. Conversely, when $A$ misspecifies the spillover structure and $\alpha_t$ exhibits a persistent component in the spillover unit, the term $(I-\Gamma_A)(I-B)\bar{\alpha}_{T_1}$ generally does not vanish, leading to a divergence in $\kappa_A$. This result facilitates consistent structure estimation, when (i) $\sqrt{T_1}(I-\Gamma_A)(I-B)\bar{\alpha}_{T_1}$ diverges for misspecifed structures, and (ii) the set of candidate spillover structures is finite, as in Example (ref).

Including covariates

Many empirical researchers are interested in including extra covariates when using synthetic control methods. Our framework can be combined with existing methods such as Abadie2010 and Li2019, and be readily adapted to settings with covariates. For example, suppose we have a vector of observable variables $z_{i,t}$ and want to estimate the treatment effects, while being worried about spillover effects. Following Li2019, we estimate the least square coefficients for the model $$y_{i,t}(0)=a_i+\sum_{j\ne i}b_{i,j}y_{j,t}(0)+z_{i,t}'\pi+u_{i,t},$$ with the simplex constraints on $b_{i,j}$ and obtain coefficient estimates $(\widehat{a}_i,\widehat{b}_i,\widehat{\pi}_i)$. This is done for each $i$. Let $\widehat{g}_t=(z_{1,t}'\widehat{\pi}_i,\dots,z_{N,t}'\widehat{\pi}_N)'$. Under appropriate regularity conditions, the results of the paper apply when the intercept estimator $\widehat{a}$ is replaced by $\widehat{a}+\widehat{g}_{t}$ at time $t$. The treatment effects estimator now becomes $\widehat{\gamma}=(A'\widehat{M}A)^{-1}A'(I-\widehat{B})'((I-\widehat{B})Y_{T+1}-\widehat{a}-\widehat{g}_{T+1}).$

Monte Carlo Simulations

We present the Monte Carlo simulation results in this section. For each case considered, we use 1000 simulation repetitions.

Estimation with spillover effects

In this section, we examine the finite sample performance of our estimation procedure proposed in Section 2.2. The model considered here is similar to Li2019, where $y_{i,t}(0)$ follows a factor model structure. We show both stationary and $\mathcal{I}(1)$ case.

table[table omitted — 2,690 chars of source]
figure[figure omitted — 434 chars of source]
table[table omitted — 2,702 chars of source]

Stationary case

The underlying factor model is

equation*[equation* omitted — 70 chars of source]

where $\lambda_t=(\lambda_{1,t},\lambda_{2,t},\lambda_{3,t})'$,

align*[align* omitted — 206 chars of source]

and $\varepsilon_{i,t}$ and $\nu_{j,s}$ is i.i.d. $N(0,1)$ for each $(i,j,s,t)$. Each entry of $\mu_i$ is drawn from an independent uniform distribution on $[0,1]$ and fixed for all repetitions. At $t=T+1$, the observed outcome is $y_{i,T+1}=y_{i,T+1}(0)+\alpha_i$, where $\alpha_i$ is either treatment effect or spillover effect and is specified below. The treatment effect is set to 5 and the spillover effect is 3.

We consider three spillover patterns. No spillover effects is the case where unit 1 receives a treatment effect of 5 at $t=T+1$ and other units are not affected. Concentrated spillover effects is the case where 1/3 of the control units receive a spillover effect of 3. Spreadout spillover effects is the case where 2/3 of the control units receive a spillover effect of 3. SCM is the original synthetic control method, and SP is the corrected synthetic control method proposed in Section (ref). Throughout the simulations, assume that we know the coverage of spillover effects but no other information, so $A$ is constructed as in Example (ref). For No spillover effects, we are being conservative in our use of the SP estimator and run it as if 1/3 of the control units are exposed to spillover effects. To better compare results, we also fit the simulation results using kernel density for the $(N,T)=(10,50)$ case with concentrated spillover effects and plot it in Figure (ref).

The empirical bias and variance (in parentheses) of the treatment effect estimator using two methods are shown in Table (ref). Throughout, SP produces virtually unbiased estimates, while the usual SCM has a bias that increases as spillovers propagate. For all cases in Concentrated spillover effects and five out of nine cases in Spreadout spillover effects, SP has a smaller empirical variance than SCM does.

$\mathcal{I}(1)$ case

For the $\mathcal{I}(1)$ case, the underlying factor model follows

equation*[equation* omitted — 63 chars of source]

where $\lambda_t=(\lambda_{1,t},\lambda_{2,t},\lambda_{3,t})'$,

align*[align* omitted — 155 chars of source]

and $\varepsilon_{i,t}$ and $\nu_{j,s}$ follows i.i.d. $N(0,1)$ for each $(i,j,s,t)$. The factor loadings are constructed such that Condition CO is satisfied. Namely, we let $\mu_1=(1,0,0)'$, $\mu_2=(0,1,0)'$, $\mu_3=(1,0,0)'$, $\mu_4=(0,1,0)'$, and for $\mu_j$ with $j=5,\dots,N$, we draw independent uniform distribution on $[0,1]$ for each entry and then normalize each loading vector such that three entries of each $\mu_j$ sum up to one. The constructed factor loadings are fixed for each repetition while other settings are the same as the stationary case.

The results are shown in Table (ref). Similarly as in the stationary case, SP produces virtually unbiased results, while SCM is biased. One thing different here is that SP often has a larger variance than SCM does, except for four out of nine cases in Concentrated spillover effects. SP has an especially large variance when $N=10$ and $T=200$.

Test for treatment effects

table[table omitted — 2,095 chars of source]
table[table omitted — 2,093 chars of source]

In this section we compare test procedures against the null hypothesis $H_0:\alpha_1=0$, i.e., the treatment effect is zero. The results are shown in Table (ref) and Table (ref). The DGP is exactly the same as in Section (ref) (the stationary case), except that $\alpha_1=0$ (the null) for Table (ref) and $\alpha_1=5$ (the alternative) for Table (ref). Placebo test is as in Abadie2010 and Hahn2017. Andrews' test is as in Andrews2003. SP is the spillover-adjust test proposed in Section (ref).

Among the three testing procedures, SP test has mostly correct sizes and outperforms the other two methods in power. The placebo test has correct sizes in some cases but has lower power, and Andrews' test over-rejects under the null. The reasons are discussed in Section (ref).

It is worth mentioning that Andrews' test and SP test may experience over-rejection in cases with small $T$. For example, in the case with $(N,T)=(50,15)$ in Table (ref), Andrews' test rejects the null 14.1% of the time in No spillover effects, and SP test rejects the null 10.9% of the time in Concentrated spillover effects. This is because Andrews-type tests rely on variation across time periods to deliver valid inference, and may experience over-rejection when it observes insufficient variation.

Test for existence of spillover effects

figure[figure omitted — 572 chars of source]

In this section, we examine the power of the proposed test against the null hypothesis that there are no spillover effects. We also look into its behavior when the range of the spillover effect is not correctly specified. In this set of experiments, the level of spillover effects varies from 0 to 2, corresponding to the strength of alternative hypotheses. We set $(N,T)=(20,50)$ and $\alpha_1=5$. There are 9 units that are affected by spillover effects. Other settings follow exactly as in Section (ref) (the stationary case). The model for the range of spillover is as in Example (ref).

The empirical rejection rates against various levels of spillover effects using our method proposed in Section (ref) are plotted in Figure (ref). Here Include too few misses half of the units that are actually affected by the treatment (assuming that unit 1 as well as four other units are affected), Correct specification assumes we know exactly which units are affected, and Include too many assumes 15 units are affected in estimation, five of which are actually not affected by spillover effects.

Among the three cases, Include too many is still a correct specification but is supposed to be more conservative, so it has less power than Correct specification. Note that the range of spillover effects for Include too few is correctly specified only when the level of spillover effects is zero. Its power curve is similar to that of Include too many.

Spatial correlation and SCM weights

Spatial settings are common in economics cao_inference_2025,cao_neighborhood_2025. In this section, we conduct simulation studies to illustrate a spatially correlated scenario in which spillover effects increase as the distance between treated and control units decreases. At the same time, the synthetic control method is found to assign greater weights to units that are closer to the treated unit. Although this example is not exhaustive, it highlights the potential risks associated with ignoring spillovers in SCM settings. In particular, SCM may assign greater weights to units that are significantly affected by spillovers due to their relative importance and proximity to the treated unit.

The data-generating process used is similar to the stationary case described in Section (ref), with $N=10$ and $T=15$. The difference is the generation of factor loadings, which are modeled as $\mu_i=\rho h_i+(1-\rho)\mu_{i,0}$, where $\rho$ is the degree of spatial correlation, $h_i$ includes location data, and $\mu_{i,0}$ is free of spatial information. Specifically, $h_i=(h_{i,1}, h_{i,2}, h_{i,3})'$ is generated such that $h_{i,1}$ and $h_{i,2}$ are independently and uniformly distributed over $[0,3]$, and $h_{i,3}$ follows an independent standard normal distribution. Here, $(h_{i,1}, h_{i,2})$ are the geographic coordinates of unit $i$. Control units are sorted by their distance to unit 1, so that unit 2 is the closest one. The non-spatial component $\mu_{i,0}$ is drawn from an independent uniform distribution on $[0,1]$. All factor loadings are predetermined and remain fixed across replications.

The configuration of the spillover effects also reflects a spatial pattern. Following the framework in Example (ref), the spillover effect vector is $\alpha =(\alpha_1, b\exp(-d_2), \ldots, b\exp(-d_N))'$, where $d_i=\sqrt{(h_{1,1}-h_{i,1})^2+(h_{1,2}-h_{i,2})^2}$ represents the distance to the treated unit. The treatment effect $\alpha_1$ is set to 5, and the largest spillover effect $b\exp(-d_2)$ is set to 3.

figure[figure omitted — 742 chars of source]

The results are shown in Figure (ref) and Table (ref). SCM is the synthetic control method neglecting the possibilities of spillover effects. SP is our estimator, where we assume the spillover structure is known as described in Example (ref). We find that the bias of the usual synthetic control estimator increases in absolute value as spatial correlation increases. In contrast, our method remains relatively unbiased and stable across different levels of spatial correlation.

table[table omitted — 2,259 chars of source]

A detailed examination of the Table (ref) provides insight into the underlying mechanism. At a spatial strength ($\rho$) of zero, the correlations between the treated unit and the controls are relatively equally-distributed, resulting in an (approximately) uniformly-distributed set of SCM weights. However, as $\rho$ increases, closer units tend to have a stronger correlation with unit 1 than more distant units, leading to more disproportionate weights to closer units. In our simulation, the unit that receives the highest average SCM weight (unit 3) also has the strongest correlation with the treated unit, except when $\rho=0$. This finding is consistent with the premise that the SCM may put more weights on units that are more affected by spillovers.

Proofs

\renewenvironment{proof}{{ Proof of Proposition (ref).}}{\qed}

proofSuppose a non-zero vector $w\in \mathbb R^{N}$ satisfies $(I-B)w=0$. Let $i\in\arg\max_j |w_j|$. Suppose for contradiction that there exists $j$ where $e(i,j)=1$ and $|w_i|\ne |w_j|$. Then, by the $i$-th row of $(I-B)w=0$, $$|w_i|=\left|\sum_{j\ne i} b_{i,j} w_j\right|\le \sum_{e(i,j)=1}b_{i,j}|w_j|< \sum_{e(i,j)=1}b_{i,j}|w_i|=|w_i|,$$ leading to contradiction, so $w_j$ is either $w_i$ or $-w_i$ for $e(i,j)=1$. However, if $w_j=-w_i$ for any $j$ with $e(i,j)=1$, $w_i=\sum_{e(i,j)=1}b_{i,j}w_j$ will not hold. Therefore, $w_j=w_i$ for any $j$ with $e(i,j)=1$. Since we have a connected network, induction implies $w_j=w_i$ for any $j$. This means the null space of $I-B$ has only one dimension, concluding our proof.

\renewenvironment{proof}{{ Proof of Theorem (ref).}}{\qed}

proofUsing formula of $\widehat{\gamma}$ in Equation (ref), we have \begin{align*} \widehat{\gamma} &= (A'\widehat{M}A)^{-1}A'(I-\widehat{B})'((I-\widehat{B})Y_{T+1}(0)+(I-\widehat{B})\alpha-\widehat{a})\notag\\ &= (A'\widehat{M}A)^{-1}A'(I-\widehat{B})'(u_{T+1}+(B-\widehat{B})Y_{T+1}(0)+(a-\widehat{a})+(I-\widehat{B})A\gamma)\notag\\ &= (A'\widehat{M}A)^{-1}A'(I-\widehat{B})'u_{T+1}+o_p(1)+o_p(1)+\gamma. \end{align*} The first equality is by $Y_{T+1}=Y_{T+1}(0)+\alpha$. The second equation is because $Y_{T+1}(0)=a+BY_{T+1}(0)+u_{T+1}$. The third equation is by (b) and (c) in Assumption (ref). Therefore, \begin{align*} \widehat{\alpha}-(\alpha+Gu_{T+1})&=A(A'\widehat{M}A)^{-1}A'(I-\widehat{B})'u_{T+1}+A\gamma +o_p(1)-\alpha-Gu_{T+1}\notag\\ &= (A(A'\widehat{M}A)^{-1}A'(I-\widehat{B})-G)'u_{T+1}+o_p(1)\notag\\ &= o_p(1)O_p(1)+o_p(1)\notag\\ &= o_p(1). \end{align*} The third equality is by (b) in Assumption (ref) and stationarity of $\{u_t \}_{t\ge 1}$.

\renewenvironment{proof}{{ Proof of Lemma (ref).}}{\qed}

proof(i) First assume that $A'MA$ is non-singular and Condition ST holds. The proof follows Ferman2016, except that we do not assume that there is a set of weights that reconstruct the factor loadings and belong to the simplex. Also, we proceed assuming $\eta_t=0$ for proof of part (a) and (b). This is without loss of generality because none of $\widehat{a}$, $\widehat{B}$, and $u_t$ depends on $\{\eta_t\}$. This also implies $\Omega_y=Cov[Y_t(0)]$. We first show part (b). It suffices to show $|\widehat{a}_i-a_i|=o_p(1)$ and $\Vert \widehat{b}_i-b_i\Vert=o_p(1)$ for each $i$, i.e., $a_i$ and $b_i$ are well-defined. We show it for the $i=1$ case, and other cases follow the same strategy. Let $\bar{y}_j=T^{-1}\sum_{t=1}^Ty_{j,t}$. Write down an (equivalent) optimization problem \begin{equation*} \widehat{v}=\underset{v\in V}{\arg\min}\left( (y_{1,t}-\bar{y}_1)-\sum_{j=2}^N(y_{j,t}-\bar{y}_{j})v_j\right) ^2, \end{equation*} where $V=\{v=(v_2,\dots,v_N)\in \mathbb{R}_+^{N-1}:\sum_{j=2}^Nv_j=1 \}$. The objective is strictly convex (with probability approaching one), so the solution is unique. Note that it implies $\widehat{b}_1$ is numerically equivalent to $(0,\widehat{v}')'$, otherwise the minimization problem in forming $\widehat{a}_1$ and $\widehat{b}_1$ may have a lower objective evaluated at $(\bar{y}_1-\sum_{j=2}^N\bar{y}_j\widehat{v}_j,0,\widehat{v}')'$. Now we let $\widehat{Q}(v)$ denote the objective function such that \begin{equation*} \widehat{Q}(v)=\frac{1}{T}\sum_{t=1}^T\left( (y_{1,t}-\bar{y}_1)-\sum_{j=2}^N(y_{j,t}-\bar{y}_{j})v_j\right) ^2, \end{equation*} and its population analog be $Q(v)=(-1,v')\Omega_y(-1,v')'$. Let $v_0$ be a minimizer of $Q(v)$ in $V$. We verify the conditions for consistency Newey1994 : (i) Since $\Omega_y$ is positive definite, $Q(v)$ is strictly convex. Also, $V$ is convex. Therefore, $Q(v)$ is uniquely minimized at $v_0$. (ii) $V$ is compact since it is an $(N-1)$-dimensional simplex. (iii) $Q(v)$ is continuous since it has a quadratic form. (iv) To see uniform convergence, note \begin{align*} \sup_{v\in V}|\widehat{Q}(v)-Q(v)| &=\sup_{v\in V}\left| \begin{bmatrix} -1\\v \end{bmatrix}'\left(\frac{1}{T}\sum_{t=1}^T(Y_t-\bar{Y})(Y_t-\bar{Y})'-\Omega_y \right) \begin{bmatrix} -1\\v \end{bmatrix}\right| \notag\\ &\le \sup_{v\in V}\left\Vert \begin{bmatrix} -1\\v \end{bmatrix}\right\Vert^2\left\Vert\frac{1}{T}\sum_{t=1}^T(Y_t-\bar{Y})(Y_t-\bar{Y})'-\Omega_y \right\Vert_F \notag\\ &\le N\cdot o_p(1)\notag\\ &= o_p(1), \end{align*} where $\Vert\cdot\Vert_F$ is the Frobenius norm. The second inequality is by ergodicity for the second moments. Therefore, $\widehat{v}\rightarrow_p v_0$. This implies $\Vert \widehat{b}_1-b_1\Vert=o_p(1)$. By ergodicity, \begin{equation*} \widehat{a}_1=\bar{y}_{1}-[\bar{y}_2\ \bar{y}_3\ \dots\ \bar{y}_N]\widehat{v}\rightarrow_p E[y_{1,t}(0)-Y_t(0)'b_1]=a_1. \end{equation*} This shows part (b) and $E[u_{1,t}]=0$ by definition of $u_{i,t}$. We also have that $\{u_t \}_{t\ge 1}$ is stationary since it is a linear combination of stationary and ergodic processes. This shows part (a) in Assumption (ref). Part (c) follows from part (b), the stationarity of $\{(\lambda_t,\varepsilon_t)\}_{t\ge 1}$, and $\Vert (\widehat B-B)\eta_{T+1}\Vert=o_p(1)$. Part (d) is assumed. Thus, Assumption (ref) holds under the invertibility of $A'MA$ and Condition ST. \ (ii) Now, we assume that $A'MA$ is non-singular and Condition CO holds. We first show part (c). We will show $\Vert Y_{T+1}(0)'(\widehat{b}_1-b_1)\Vert =o_p(1)$ and other $i$'s follows the same strategy. Since the synthetic control estimator can be written as a projection of the OLS estimator onto a closed convex set, we will first derive the asymptotic properties of the OLS estimator, and then use the properties of projections to obtain the desired results. For examples of this strategy, see Li2019 and Yu2019. For some positive definite matrix $D\in \mathbb{R}^{N}$, let $\mathbb{R}^{N}$ be a Hilbert space with the inner product $\langle\cdot,\cdot\rangle_D$ such that for $\theta_1,\theta_2\in \mathbb{R}^{N}$, $ \langle \theta_1,\theta_2\rangle_D=\theta_1'D\theta_2.$ The norm $\Vert \cdot\Vert_D$ is defined accordingly, i.e. $\Vert\theta\Vert_D=\sqrt{\theta'D\theta}$, for $\theta\in \mathbb{R}^{N}$. For a closed convex set $\Lambda\subset\mathbb{R}^{N}$, define a projection $\Pi_D$ such that for each $\theta\in \mathbb{R}^{N}$, $\Pi_D\theta = \arg\min_{\theta'\in \Lambda}\Vert \theta-\theta'\Vert_D$. Zarantonello1971 shows that for each $\theta,\theta'\in \mathbb{R}^{N}$, \begin{equation} \Vert \Pi_D\theta-\Pi_D\theta'\Vert_D\le \Vert \theta-\theta'\Vert_D. \end{equation} With some abuse of notation, let $x_t=Y_t-T^{-1}\sum_{s=1}^TY_s$. Then, $\widehat{b}_1$ is the synthetic control weight estimators of regressing $(y_{1,t}-T^{-1}\sum_{s=1}^Ty_{1,s})$ on $x_t$, subject to $\{0 \}\times \Delta_{N-1}$ with $\Delta_{N-1}$ being an $(N-1)$-dimensional simplex. Let $\tilde{b}_1$ be the OLS estimator of regressing $(y_{1,t}-T^{-1}\sum_{s=1}^Ty_{1,s})$ on $x_t$. Let $\Sigma_T=T^{-1}\sum_{t=1}^Tx_tx_t'$. Appendix A.2 in Li2019 establishes that $\widehat{b}_1=\Pi_{\Sigma_T}\tilde{b}_1$. Thus, we have \begin{eqnarray} \Vert\widehat{b}_1-b_1\Vert&=& \Vert\Sigma_T^{-1/2}\Sigma_T^{1/2}(\widehat{b}_1-b_1)\Vert\notag\\ &\le&\Vert \Sigma_T^{-1/2}\Vert_F\cdot\Vert\Sigma_T^{1/2}(\widehat{b}_1-b_1)\Vert\notag\\ &=&\Vert \Sigma_T^{-1/2}\Vert_F\cdot\Vert\widehat{b}_1-b_1\Vert_{\Sigma_T}\notag\\ &=&\Vert \Sigma_T^{-1/2}\Vert_F\cdot\Vert \Pi_{\Sigma_T}\tilde{b}_1-\Pi_{\Sigma_T}b_1\Vert_{\Sigma_T}\notag\\ &\le & \Vert \Sigma_T^{-1/2}\Vert_F\cdot \Vert \tilde{b}_1-b_1\Vert_{\Sigma_T}\notag\\ &=&\Vert \Sigma_T^{-1/2}\Vert_F\cdot\Vert \Sigma_T^{1/2}\Vert_F\cdot \Vert \tilde{b}_1-b_1\Vert\notag\\ &=& O_p(1)o_p(T^{-1/2})\notag\\ &=&o_p(T^{-1/2}), \end{eqnarray} where $\Vert \cdot\Vert_F$ is the Frobenius norm of a matrix. The third equality is because $b_1\in \{0 \}\times \Delta_{N-1}$. The second inequality is by (ref). To see the fifth equality, note \begin{equation*} \Sigma_T=T\left( \frac{1}{T^2}\sum_{t=1}^TY_tY_t'-\left( \frac{1}{T^{3/2}}\sum_{t=1}Y_t\right)\left( \frac{1}{T^{3/2}}\sum_{t=1}Y_t\right)' \right), \end{equation*} so \begin{equation*} \Vert \Sigma_T^{-1/2}\Vert_F\cdot\Vert \Sigma_T^{1/2}\Vert_F=tr(\Sigma_T^{-1})tr(\Sigma_T)=O_p(1)\cdot \frac{1}{T}\cdot T\cdot O_p(1)=O_p(1), \end{equation*} where the second equality is standard results for $\mathcal{I}_1$ process Hamilton1994. Also, $\Vert \tilde{b}_1-b_1\Vert =o_p(T^{-1/2})$ is by Proposition 19.2 in Hamilton1994. This shows (ref). Apply part (a) of Proposition 18.1 in Hamilton1994, we have \begin{equation*} \Vert Y_{T+1}(0)'(\widehat{b}_1-b)\Vert = \Vert (T^{-1/2}Y_{T+1}(0))'(T^{-1/2}(\widehat{b}_1-b))\Vert = O_p(1)o_p(1)=o_p(1). \end{equation*} Now we show part (b). Again, it suffices to show $|\widehat{a}_i-a_i|=o_p(1)$ and $\Vert \widehat{b}_i-b_i\Vert=o_p(1)$. We consider the $i=1$ case and other cases follow the same strategy. We have showed $\Vert \widehat{b}_i-b_i\Vert =o_p(1)$ in part (c) of the proof. Section A.6.1 in Ferman2016 establishes that \begin{equation} [\mu_1^1\ \mu_2^1\ \dots \ \mu_N^1](b_1-e_1)=0, \end{equation} where $e_i$ is the unit vector with one at the $i$-th entry. Thus, \begin{eqnarray} \widehat{a}_1 &=&[\bar{y}_1\ \bar{y}_2\ \dots\ \bar{y}_N](e_1-\widehat{b}_1)\notag\\ &=&[\bar{y}_1\ \bar{y}_2\ \dots\ \bar{y}_N](e_1-b_1)+[\bar{y}_1\ \bar{y}_2\ \dots\ \bar{y}_N](b_1-\widehat{b}_1)\notag\\ &=&\left\lbrace \frac{1}{T}\sum_{t=1}^T\left( (\lambda_t^0)'[\mu_1^0\ \dots\ \mu_N^0]+[\varepsilon_{1,t}\ \dots\ \varepsilon_{N,t}]\right) \right\rbrace (e_1-b_1)+\notag\\ &&\left( \frac{1}{\sqrt{T}}[\bar{y}_1\ \bar{y}_2\ \dots\ \bar{y}_N]\right) \sqrt{T}(b_1-\widehat{b}_1)\notag\\ &=& E[\lambda_t^0]'[\mu_1^0\ \dots\ \mu_N^0](e_1-b_1)+o_p(1)+O_p(1)o_p(1)\notag\\ &\rightarrow_p& E[\lambda_t^0]'[\mu_1^0\ \dots\ \mu_N^0](e_1-b_1).\notag\\ &=& a_1. \end{eqnarray} The third equality is by (ref). The fourth equality is by stationarity of $\{(\lambda_t^0,\varepsilon_{t})\}_{t\ge 1}$ and results in part (d) of the proof. This shows part (b) of the Assumption (ref) . Combining (ref) and (ref), we have part (a) in Assumption (ref). Part (d) is assumed.

\renewenvironment{proof}{{ Proof of Theorem (ref).}}{\qed}

proofWe follow the proof of Theorem 2 in Andrews2006. Let \begin{equation*} \begin{aligned} &L_{1,T}(\epsilon)=\left\lbrace \Vert C_T(\widehat{\beta}_1-\beta_1)\Vert\le \epsilon, \max_{t=1,\dots,T}\Vert C_T(\widehat{\beta}_1^{(t)}-\beta_1) \Vert \le \epsilon \right\rbrace, \\ &L_{2,T}(c)=\left\lbrace \max_{t\le T+1}\Vert C^{-1}_Tx_t\Vert\le c \right\rbrace . \end{aligned} \end{equation*} By Assumption 2(d), there exists a positive sequence $\{\epsilon_T \}_{T\ge 1}$ such that $\epsilon_T\rightarrow 0$ and $\Pr(L_{1,T}(\epsilon_T))\rightarrow 1$. Let $c_T=1/\sqrt{\epsilon_T}$. So we have $c_T\rightarrow \infty$ and $c_T\epsilon_T\rightarrow 0$. By Assumption 2(c), we must have $\Pr(L_{2,T}(c_T))\rightarrow 1$. Let $L_T=L_{1,T}(\epsilon_T)\cap L_{2,T}(c_T)$, then we have $\Pr(L_T)\rightarrow 1$ and $\Pr(L_T^c)\rightarrow 0$. Suppose $L_T$ holds. Then, for $\beta=\widehat{\beta}_1$ or $\beta=\widehat{\beta}_{1}^{(t)}$ for some $t=1,\dots,T$, we have \begin{eqnarray*} |P_t(\beta)-P_t(\beta_1)|&=&\left| (\beta-\beta_1)'x_tx_t'(\beta-\beta_1)-2x_t'(\beta-\beta_1)u_{1,t}\right| \notag\\ &=&\left| (\beta-\beta_1)'C_T'(C_T')^{-1}x_tx_t'C_T^{-1}C_T(\beta-\beta_1)-2x_t'C_T^{-1}C_T(\beta-\beta_1)u_{1,t}\right| \notag\\ &\le & \Vert C_T(\beta-\beta_1)\Vert^2\Vert C_T^{-1}x_t\Vert^2+2\Vert C_T^{-1}x_t\Vert \Vert C_T(\beta-\beta_1)\Vert |u_{1,t}|\notag\\ &\le & \epsilon_T^2c_T^2+2\epsilon_Tc_T|u_{1,t}|. \end{eqnarray*} Define $g_t(\epsilon_T,c_T)=\epsilon_T^2c_T^2+2\epsilon_Tc_T|u_{1,t}|$. Note that $g_t(\epsilon_T,c_T)$ is identically distributed across $t$ for a fixed $T$, by Assumption 2(a). We first prove part (a). Let $x$ be some continuous point of distribution function of $P_{T+1}(\beta_1)$. Then, \begin{eqnarray*} \Pr(P_{T+1}(\widehat{\beta}_1)\le x)&=&\Pr(\{P_{T+1}(\widehat{\beta}_1)\le x \}\cap L_T)+\Pr(\{P_{T+1}(\widehat{\beta}_1)\le x \}\cap L_T^c)\notag\\ &\le & \Pr(P_{T+1}(\widehat{\beta}_1)\le x +g_t(\epsilon_T,c_T))+\Pr(L_T^c)\notag\\ &\le &\Pr(P_{T+1}(\beta_1)\le x)+o(1). \end{eqnarray*} To see the last equality, pick $\epsilon>0$. By continuity, $\exists \delta>0$ such that for each $y\in (x-\delta,x+\delta)$, $|\Pr(P_{T+1}(\beta_1)\le y)-\Pr(P_{T+1}(\beta_1)\le x)|<\epsilon$. Therefore, \begin{eqnarray*} &&\Pr(P_{T+1}(\widehat{\beta}_1)\le x +g_t(\epsilon_T,c_T))\\ &=&\Pr(\{P_{T+1}(\widehat{\beta}_1)\le x +g_t(\epsilon_T,c_T)\}\cap \{ |g_t(\epsilon_T,c_T)|\ge \delta\})\notag\\ &&+\Pr(\{P_{T+1}(\widehat{\beta}_1)\le x +g_t(\epsilon_T,c_T)\}\cap \{|g_t(\epsilon_T,c_T)|<\delta\})\notag\\ &\le &\Pr(|g_t(\epsilon_T,c_T)|\ge \delta)+\Pr(P_{T+1}(\widehat{\beta}_1)\le y)\notag\\ &<&\Pr(P_{T+1}(\beta_1)\le x)+o(1). \end{eqnarray*} Similarly, $ \Pr(P_{T+1}(\widehat{\beta}_1)\le x)\ge \Pr(P_{T+1}(\beta_1)\le x)+o(1).$ This shows part (a). To see part (b), let $k:\mathbb{R}\rightarrow\mathbb{R}$ be a monotonically decreasing and everywhere differentiable function that has bounded derivative and satisfies $k(x)=1$ for $x\le 0$, $k(x)\in [0,1]$ for $x\in (0,1)$, and $k(x)=0$ for $x\ge 1$. For example, let $k(x)=\cos(\pi x)/2+1/2$ for $x\in (0,1)$. Given some $\{\beta^{(t)} \}_{t=1}^T$, a smoothed df is defined by \begin{equation*} \widehat{F}_T(x,\{\beta^t \},h_T)=\frac{1}{T}\sum_{t=1}^{T}k\left( \frac{P_t(\beta^{(t)})-x}{h_T} \right), \end{equation*} for some sequence of positive constants $\{h_T \}$ such that $h_T\rightarrow 0$ and $c_T\epsilon_T/h_T\rightarrow 0$. For example, we let $h_T=\epsilon_T^{1/4}$ when $c_T=1/\sqrt{\epsilon_T}$. Also, define, \begin{equation*} \widehat{F}_T(x,\{\beta_1\})=\frac{1}{T}\sum_{t=1}^{T}\mathbbm{1} \{ P_t(\beta_1)\le x \}, \end{equation*} i.e., $\widehat{F}_T(x,\{\beta_1\})$ is the empirical cdf of $P_t$ as if the true parameter $\beta_1$ is known. We write \begin{equation*} |\widehat{F}_{P,T}(x)-F_P(x)|\le \sum_{i=1}^{4}D_{i,T}, \end{equation*} for \begin{align*} D_{1,T}&= |\widehat{F}_{P,T}(x)-\widehat{F}_T(x,\{\widehat{\beta}_j \},h_T)|,\notag\\ D_{2,T}&= |\widehat{F}_T(x,\{\widehat{\beta}_j \},h_T)-\widehat{F}_T(x,\{\beta_1 \},h_T)|,\notag\\ D_{3,T}&= |\widehat{F}_T(x,\{\beta_1 \},h_T)-\widehat{F}_T(x,\{\beta_1\})|, and\notag\\ D_{4,T}&= |\widehat{F}_T(x,\{\beta_1\})-{F}_P(x)|. \end{align*} We want to show that all four terms vanish. First note that \begin{equation*} D_{1,T}\le \frac{1}{T}\sum_{t=1}^{T}\mathbbm{1}\left\lbrace \frac{P_t(\widehat{\beta}_1^{(t)})-x}{h_T}\in (0,1) \right\rbrace . \end{equation*} Thus, for any $\delta>0$, \begin{align} \Pr(D_{1,T}>\delta)&\le \Pr(\{D_{1,T}>\delta \}\cap L_T)+\Pr(L_T^c)\notag\\ &\le \Pr \left( \frac{1}{T}\sum_{t=1}^{T}\mathbbm{1}\left\lbrace P_t(\widehat{\beta}_1^{(t)})-x\in (-g_t(\epsilon_T,c_T),h_T+g_t(\epsilon_T,c_T) \right\rbrace>\delta \right) +o(1)\notag\\ &\le \frac{E\mathbbm{1}\left\lbrace P_t(\widehat{\beta}_1^{(t)})-x\in (-g_t(\epsilon_T,c_T),h_T+g_t(\epsilon_T,c_T) \right\rbrace}{\delta}+o(1), \end{align} where the last inequality is by Markov's inequality. Recall $\Pr(P_1(\beta_1)\ne x)=1$ and $g_t(\epsilon_T,c_T)\rightarrow 0$ a.s., so $\mathbbm{1}\{P_t(\beta_1)-x\in \{-g_t(\epsilon_T,c_T),h_T+g_t(\epsilon_T,c_T) \}\rightarrow 0$ a.s.. By the dominated convergence theorem, (ref) implies $\Pr(D_{1,T}>\delta)\le o(1)$ and thus $D_{1,T}=o_p(1)$. For $D_{2,T}$, we have \begin{align*} D_{2,T}= \left| \frac{1}{T} \sum_{t=1}^{T} k'\left( \frac{\tilde{P}_t-x}{h_T} \right) \frac{P_t(\widehat{\beta}_1^{(t)})-P_t(\beta_1)}{h_T} \right| \le \frac{\bar{k}}{T}\sum_{t=1}^{T}\frac{g_t(\epsilon_T,c_T)}{h_T}. \end{align*} The equality is by the mean value theorem and we have $\tilde{P}_t$ lies between $P_t(\widehat{\beta}_1^{(t)})$ and $P_t(\beta_1)$. In the inequality, $\bar{k}$ is a bound for the derivative of $k$. Also, note \begin{equation*} E\left[ \frac{g_t(\epsilon_T,c_T)}{h_T}\right] =\frac{\epsilon_T^2c_T^2}{h_T}+2\frac{\epsilon_Tc_T}{h_T}E|u_{1,t}|=o(1). \end{equation*} Therefore, \begin{align*} \Pr(D_{2,T}>\delta)&\le \Pr (\{D_{2,T}>\delta \}\cap L_T)+\Pr (L_T^c)\notag\\ &\le \Pr \left( \frac{\bar{k}}{T}\sum_{t=1}^{T}\frac{g_t(\epsilon_T,c_T)}{h_T}>\delta \right) +o(1)\notag\\ &\le \bar{k} \frac{Eg_t(\epsilon_T,c_T)}{\delta h_T}\notag\\ &\rightarrow 0. \end{align*} The third inequality is by Markov's inequality. This shows $D_{2,T}=o_p(1)$. $D_{3,T}$ is similar to the $D_{1,T}$ case. Finally, by stationary and ergodicity of $u_{1,t}$, we have $D_{4,T}=o_p(1)$. This shows part (b). Now we show part (c). Pick any small $\epsilon$ such that $\widehat{F}_{P,T}(x)\rightarrow_pF_P(x)$ for $x\in (q_{P,1-\tau}-\epsilon,q_{P,1-\tau}+\epsilon)$. Note \begin{eqnarray*} && \Pr (\widehat{q}_{P,1-\tau}>q_{P,1-\tau}+\epsilon)\\ &\le &\Pr (\widehat{F}_{P,T}({q}_{P,1-\tau}+\epsilon)<1-\tau)\notag\\ &= &\Pr (\widehat{F}_{P,T}({q}_{P,1-\tau}+\epsilon)-{F}_P({q}_{P,1-\tau}+\epsilon)<(1-\tau)-{F}_P({q}_{P,1-\tau}+\epsilon))\notag\\ &\rightarrow &0. \end{eqnarray*} The inequality is by definition of $\widehat{q}_{P,1-\tau}$. The convergence is because of part (e) of Assumption 2 and part (b) of Theorem 2. Similarly, \begin{eqnarray*} && \Pr (\widehat{q}_{P,1-\tau}<q_{P,1-\tau}-\epsilon)\\ &\le &\Pr (\widehat{F}_{P,T}({q}_{P,1-\tau}-\epsilon)\ge 1-\tau)\notag\\ &=& \Pr (\widehat{F}_{P,T}({q}_{P,1-\tau}-\epsilon)-{F}_P({q}_{P,1-\tau}-\epsilon)\ge (1-\tau)-{F}_P({q}_{P,1-\tau}-\epsilon))\notag\\ &\rightarrow& 0. \end{eqnarray*} Again, the inequality is by definition of $\widehat{q}_{P,1-\tau}$, and the convergence is because of part (e) of Assumption 2 and part (b) of Theorem 2. Finally, we show part (d). Under null, $P_\infty$ and $P_1(\beta_1)$ have the same distribution, so $q_{P,1-\tau}$ is $(1-\tau)$-quantile of $P_\infty$. Therefore, $$ \Pr(P>\widehat{q}_{P,1-\tau})=1-\Pr (P\le \widehat{q}_{P,1-\tau})=1-\Pr (P+( {q}_{P,1-\tau}- \widehat{q}_{P,1-\tau})\le {q}_{P,1-\tau})\rightarrow \tau,$$ where the convergence is by combining part (a) and (c). This concludes our proof.

\renewenvironment{proof}{{ Proof of Lemma (ref).}}{\qed}

proofSince Assumption 3 implies Assumption 2, we only need to show Lemma 3.

\renewenvironment{proof}{{ Proof of Theorem (ref).}}{\qed}

proofWe use similar strategy as we do in the proof of Theorem (ref). Let \begin{equation*} \begin{aligned} &L_{1,T}(\epsilon)=\left\lbrace \Vert (\widehat{\theta}-\theta_0)D_T\Vert_F\le \epsilon, \max_{t=1,\dots,T}\Vert (\widehat{\theta}^{(t)}-\theta_0)D_T \Vert_F \le \epsilon \right\rbrace, \\ &L_{2,T}(c)=\left\lbrace \max_{t\le T+1}\Vert D^{-1}_Tx_t\Vert\le c \right\rbrace,\\ &L_{3,T}(\eta)=\left\lbrace\Vert \widehat{G}'C'W_TC\widehat{G}-G'C'WCG\Vert_F<\eta \right\rbrace. \end{aligned} \end{equation*} By Assumption (ref)(d), there exists a positive sequence $\{\epsilon_T \}_{T\ge 1}$ such that $\epsilon_T\rightarrow 0$ and $\Pr(L_{1,T}(\epsilon_T))\rightarrow 1$. Let $c_T=1/\sqrt{\epsilon_T}$. So we have $c_T\rightarrow \infty$ and $c_T\epsilon_T\rightarrow 0$. By Assumption 2(c), we must have $\Pr(L_{2,T}(c_T))\rightarrow 1$. By Assumption 1(c) and Assumption 2(f), there exists a positive sequence $\{\eta_T \}_{T\ge 1}$ such that $\eta_T\rightarrow 0$ and $\Pr(L_{3,T}(\eta_T))\rightarrow 1$. Let $L_T=L_{1,T}(\epsilon_T)\cap L_{2,T}(c_T)\cap L_{3,T}(\eta_T)$, then we have $\Pr(L_T)\rightarrow 1$ and $\Pr(L_T^c)\rightarrow 0$. Suppose $L_T$ holds. Then, for some $\theta=\widehat{\theta}$ or $\theta=\widehat{\theta}^{(t)}$ and for some $t=1,\dots, T$, we have \begin{equation} |\widehat{P}_t(\theta)-P_t(\theta_0)|\le |\widehat{P}_t(\theta)-P_t(\theta)|+|{P}_t(\theta)-P_t(\theta_0)|. \end{equation} Note that \begin{align} |\widehat{P}_t(\theta)-P_t(\theta)| &= \left| (Y_t-\theta x_t)'(\widehat{G}'C'W_TC\widehat{G})-G'C'WCG)(Y_t-\theta x_t) \right| \notag\\ &\le \Vert Y_t-\theta x_t\Vert^2 \Vert (\widehat{G}'C'W_TC\widehat{G}-G'C'WCG\Vert_F\notag\\ &\le \Vert u_t+(\theta_0-\theta)x_t\Vert^2 \cdot \eta_T\notag\\ &\le (\Vert u_t\Vert+\Vert (\theta_0-\theta)D_TD_T^{-1}x_t\Vert)^2\eta_T\notag\\ &\le (\Vert u_t\Vert +\Vert (\theta_0-\theta)D_T\Vert_F \Vert D_T^{-1}x_t\Vert)^2\eta_T\notag\\ &\le (\Vert u_t\Vert +\epsilon_Tc_T)^2\eta_T \end{align} and \begin{align} |{P}_t(\theta)-P_t(\theta_0)| =&\ |(Y_t-\theta x_t)'G'C'WCG(Y_t-\theta x_t)-(Y_t-\theta_0 x_t)'G'C'WCG(Y_t-\theta_0 x_t)|\notag\\ \le&\ |(Y_t-\theta x_t)'G'C'WCG(Y_t-\theta x_t)-(Y_t-\theta x_t)'G'C'WCG(Y_t-\theta_0 x_t)|\notag\\ &\ + |(Y_t-\theta x_t)'G'C'WCG(Y_t-\theta_0 x_t)-(Y_t-\theta_0 x_t)'G'C'WCG(Y_t-\theta_0 x_t)|\notag\\ =&\ |(u_t+(\theta_0-\theta) x_t)'G'C'WCG(\theta_0-\theta) x_t|+|((\theta_0-\theta)x_t)'G'C'WCGu_t|\notag\\ \le &\ \Vert u_t+(\theta_0-\theta)D_TD_T^{-1} x_t\Vert \Vert G'C'WCG\Vert_F \Vert (\theta_0-\theta)D_TD_T^{-1} x_t\Vert \notag\\ &\ +\Vert (\theta_0-\theta)D_TD_T^{-1}x_t\Vert \Vert G'C'WCG\Vert_F \Vert u_t\Vert \notag\\ \le &\ (\Vert u_t\Vert +\epsilon_Tc_T) \Vert G'C'WCG\Vert_F \epsilon_Tc_t+\epsilon_Tc_T \Vert G'C'WCG\Vert_F\Vert u_t\Vert\notag\\ = &\ (2\Vert u_t\Vert +\epsilon_Tc_T) \Vert G'C'WCG\Vert_F \epsilon_Tc_t. \end{align} Combining (ref), (ref), and (ref), we have $ |\widehat{P}_t(\theta)-P_t(\theta_0)|\le g(\epsilon_T,c_T,\eta_T),$ where $ g_t(\epsilon_T,c_T,\eta_T)= (\Vert u_t\Vert +\epsilon_Tc_T)^2\eta_T+(2\Vert u_t\Vert +\epsilon_Tc_T) \Vert G'C'WCG\Vert_F \epsilon_Tc_t.$ By Assumption 1(a), $g_t(\epsilon_T,c_T,\eta_T)$ is identically distributed across $t$ for a fixed $T$. To show part (a), note that under null, \begin{align*} P&=(C\widehat{\alpha}-d)'W_T(C\widehat{\alpha}-d)\notag\\ &=(C(\alpha+Gu_{T+1}+o_p(1))-d)'(W+o_p(1))(C(\alpha+Gu_{T+1}+o_p(1))-d)\notag\\ &= (CGu_{T+1}+o_p(1))'(W+o_p(1))(CGu_{T+1}+o_p(1))\notag\\ &= u_{T+1}'G'C'WCGu_{T+1}+o_p(1). \end{align*} The second equality is by Theorem 1. Since $P_\infty=u_1'G'C'WCGu_1$, we have $P\rightarrow_d P_\infty$ by stationary of $\{u_t \}_{t\ge 1}$. Part (b)-(d) can be shown using the same strategy as in the proof of Theorem 2, with $g_t(\epsilon_T,c_T,\eta_T)$ in place of $g_t(\epsilon_T,c_T)$, and $\theta$ in place of $\beta$, so is omitted here.

\renewenvironment{proof}{{ Proof of Lemma (ref).}}{\qed}

proof(i) Assume Condition ST holds. By Lemma 1, part (a) of Assumption (ref) holds. Part (b) is because $u_t$ is a linear combination of $\eta_t,\lambda_t,\varepsilon_t$. For part (c), pick some $\tau$ such that $1/(2+\delta)<\tau<1/2$, where $\delta$ is defined in Condition ST. Let \begin{equation} D_T=\begin{bmatrix} 1 & 0 \\ 0 & T^\tau I_N \end{bmatrix}. \end{equation} Then, we have \begin{equation} \max_{t\le T+1}\Vert D_T^{-1}x_t\Vert = \max_{t\le T+1}\left\Vert\begin{bmatrix} 1\\T^{-\tau}Y_t \end{bmatrix} \right\Vert =\sqrt{1+\left( \max_{t\le T+1}\Vert T^{-\tau}Y_t\Vert\right)^2 }. \end{equation} Also, for any $\epsilon>0$, note \begin{align} \Pr\left( \max_{t\le T+1}\Vert T^{-\tau}Y_t\Vert >\epsilon \right) &=\Pr \left( \bigcup_{t\le T+1}\Vert Y_t\Vert >T^\tau \epsilon \right) \notag\\ &\le\left( \sum_{t=1}^T\Pr (\Vert Y_t\Vert >T^\tau \epsilon)\right) +\Pr (\Vert Y_{T+1}(0)+\alpha\Vert >T^\tau \epsilon )\notag\\ &=\frac{TE[\Vert Y_t\Vert ^{2+\delta}]}{T^{\tau (2+\delta)}\epsilon^{2+\delta}}+o(1)\notag\\ &=o(1). \end{align} The second equality is due to Markov inequality and stationarity of $\{Y_{T+1}(0)\}_{t+1}$. The last equality is because $\tau>1/(2+\delta)$. Combining (ref) and (ref), we obtain part (c). For part (d), we use $D_T$ defined in (ref). Following the same reasoning as in (ref), for each $i=1,\dots,N$, we have \begin{align} \Vert \widehat{b}_i-b_i\Vert& \le \Vert \Sigma_T^{-1/2}\Vert_F\cdot \Vert \Sigma_T^{1/2}\Vert_F\cdot\Vert \tilde{b}_i-b_i\Vert \notag\\ &=O_p(1) O_p(T^{-1/2})\notag\\ &=O_p(T^{-1/2}). \end{align} The first equality is because $\{Y_t(0) \}_{t\ge 1}$ is ergodic for the second moment, and $\tilde{b}_i$ is the OLS estimator for $b_i$. Thus, \begin{align*} \Vert D_T(\widehat{\beta}_i-\beta_i)\Vert &=\left\Vert \begin{bmatrix} 1 & 0 \\ 0 & T^{\tau-1/2}I_N \end{bmatrix} \begin{bmatrix} 1 & 0 \\ 0 & T^{1/2}I_N \end{bmatrix} (\widehat{\beta}_i-\beta_i) \right\Vert \notag\\ & \le \left\Vert \begin{bmatrix} 1 & 0 \\ 0 & T^{\tau-1/2}I_N \end{bmatrix} \right\Vert_F\left\Vert \begin{bmatrix}\widehat{a}_i-a_i\\\sqrt{T}(\widehat{b}_i-b_i) \end{bmatrix} \right\Vert \notag\\ &=\sqrt{1+NT^{2\tau-1}}\Vert O_p(1)\Vert\notag\\ &=o_p(1). \end{align*} The second equality is due to (ref). The last equality is because $\tau<1/2$. Therefore, $ \Vert (\widehat{\theta}-\theta_0)D_T\Vert_F=\sqrt{\sum_{i=1}^N\Vert D_T(\widehat{\beta}_i-\beta_i) \Vert^2 }=o_p(1). $ Also, since $\widehat{\theta}^{(t)}=\widehat{\theta}$ for each $t$, $ \max_{t=1,\dots,T}\Vert (\widehat{\theta}^{(t)}-\theta_0)D_T\Vert_F= \Vert (\widehat{\theta}-\theta_0)D_T\Vert_F=o_p(1). $ This shows part (d). Part (e) is assumed. Part (f) is trivial if $W_T=I$. Assume now $W_T=(C\widehat{G}(T^{-1}\sum_{t=1}^T\widehat{u}_t\widehat{u}_t')\widehat{G}'C')^{-1}$. Then, \begin{eqnarray*} &&\frac{1}{T}\sum_{t=1}^T\widehat{u}_t\widehat{u}_t'\\ &=&(I-\widehat{B})\left( \frac{1}{T}\sum_{t=1}^TY_tY_t'\right) (I-\widehat{B})'-(I-\widehat{B})\left( \frac{1}{T}\sum_{t=1}^TY_t\right)\widehat{a}'\\ &&-\widehat{a}\left( \frac{1}{T}\sum_{t=1}^TY_t'\right)(I-\widehat{B})'+\widehat{a}\widehat{a}'\notag\\ &\rightarrow &E[u_tu_t'], \end{eqnarray*} by ergodicity and Assumption 1(b). Therefore, $\widehat{W}_T\rightarrow_p W=(CGE[u_tu_t']G'C')^{-1}$. This concludes part (i) of Lemma 3. \ (ii) Assume Condition CO holds. By Lemma 1, Assumption 1 holds. This shows Part (a). By (ref), $u_t$ is a linear combination of $\lambda_t^o$ and $\varepsilon_t$, so $\{u_t \}_{t\ge 1}$ is ergodic and has finite first moment. This shows Part (b). Now we show Part (c). Let \begin{equation*} D_T=\begin{bmatrix} 1 & 0 \\ 0 & \sqrt{T}\cdot I_N \end{bmatrix}. \end{equation*} Then, we have \begin{align*} \max_{t\le T+1} \Vert D_T^{-1}x_t\Vert &=\sqrt{1+\left( \max_{t\le T+1}\Vert T^{-1/2}Y_t\Vert \right)^2 }\notag\\ &\le \sqrt{1+\sum_{i=1}^N\left( \max_{t\le T+1}|T^{-1/2}y_{i,t}| \right)^2 }\notag\\ &\le \sqrt{1+\sum_{i=1}^N\left( T^{-1/2}|\alpha_i|+\max_{t\le T+1}|T^{-1/2}y_{i,t}(0)| \right)^2 }\notag\\ &= \sqrt{1+\sum_{i=1}^N\left( o(1)+O_p(1) \right)^2 }\notag\\ &=O_p(1) \end{align*} The second equality is because \begin{equation*} \max_{t\le T+1}|T^{-1/2}y_{i,t}(0)|= \max_{r\in [0,1]}|(T+1)^{-1/2}y_{i,[r(T+1)]}(0)| \Rightarrow \max_{r\in[0,1]}\nu_i(r) \end{equation*} by the continuous mapping theorem. To show Part (d), we combine (ref) and (ref), and have \begin{equation*} \Vert D_T(\widehat{\beta}_i-\beta_i)\Vert =\left\Vert \begin{bmatrix}\widehat{a}_i-a_i\\\sqrt{T}(\widehat{b}_i-b_i) \end{bmatrix} \right\Vert=o_p(1). \end{equation*} Therefore, $ \Vert (\widehat{\theta}-\theta_0)D_T\Vert_F=\sqrt{\sum_{i=1}^N\Vert D_T(\widehat{\beta}_i-\beta_i) \Vert^2 }=o_p(1). $ The second half of Part (d) is also satisfied since $\widehat{\theta}^{(t)}=\widehat{\theta}$ for each $t$. Part (e) is assumed and Part (f) is trivial for $W_T=I$.

\renewenvironment{proof}{{ Proof of Proposition (ref).}}{\qed}

proofThe main statement follows from the same argument as in the proof of Theorem (ref), so is omitted. For the correct specification case, note that \begin{align*} \kappa_A = & \Vert (I-\widehat \Gamma_A)u_{T+1}+(I-\Gamma_A)(I-B)\alpha+[(I-\widehat \Gamma_A)(I-\widehat{B})-(I-\Gamma_A)(I-B)]\alpha \\ & +(I-\widehat \Gamma_A)(\widehat{B}-B) Y_{T+1}(0)+(I-\widehat \Gamma_A)(a-\widehat{a})\Vert \\ =& \Vert (I-\widehat \Gamma_A)u_{T+1}+(I-\Gamma_A)(I-B)\alpha+o_p(1)\Vert. \end{align*} When $A$ is correctly specified, $(I-\Gamma_A)(I-B)\alpha=0$, which concludes the proof.

\renewenvironment{proof}{{ Proof of Proposition (ref).}}{\qed}

proofThe proof for the first half of the proposition is similar to the proof for Theorem (ref), and thus is omitted. To see the second half, note \begin{equation*} Cov[G_Wu_{T+1}]=A(Q'WQ)^{-1}Q'W\Omega WQ(Q'WQ)^{-1}A' \end{equation*} and \begin{equation*} Cov[G^eu_{T+1}]=A(Q'\Omega Q)^{-1}A', \end{equation*} where $Q=(I-B)A$. It suffices to show $((Q'WQ)^{-1}Q'W\Omega WQ(Q'WQ)^{-1}-(Q'\Omega Q)^{-1})$ is positive semi-definite. Note that the first term is asymptotic variance of using $W$ as the weighting matrix in GMM exercise and the second term is the one using the efficient weighting matrix Hayashi2000. Thus, $(Cov[G_Wu_{T+1}]-Cov[G^eu_{T+1}])$ is positive semi-definite.

\renewenvironment{proof}{{ Proof of Proposition (ref).}}{\qed}

proofLet $\widehat{G} = A(A'\widehat{M}A)^{-1}A'(I-\widehat{B})'$. Note that \[ \alpha_{T+s}-\widehat{\alpha}_{T+s}=(I-\widehat{G}(I-\widehat{B}))\alpha_{T+s}-\widehat{G}(I-\widehat{B})Y_{T+s}(0)+\widehat{G}\widehat{a}, \] so we have \begin{align*} &(I-\widehat{B})(Y_{T+s}-\widehat{\alpha}_{T+s})-\widehat{a}\\ = & (I-\widehat{B})Y_{T+s}(0)+(I-\widehat{B})(\alpha_{T+s}-\widehat{\alpha}_{T+s})-\widehat{a} \\ = & (I-\widehat \Gamma_A)u_{T+s}+(I-P_A)(I-B)\alpha_{T+s}+[(I-\widehat \Gamma_A)(I-\widehat{B})-(I-P_A)(I-B)]\alpha_{T+s} \\ & +(I-\widehat \Gamma_A)(\widehat{B}-B) Y_{T+s}(0)+(I-\widehat \Gamma_A)(a-\widehat{a}). \end{align*} Thus, we can write \begin{align*} \kappa_A &= (I-\widehat \Gamma_A)\frac{1}{\sqrt{T_1}}\sum_{s=1}^{T_1} u_{T+t}+\sqrt{T_1}(I-P_A)(I-B)\bar{\alpha}_{T_1}\\ &\quad +\sqrt{T_1}[(I-\widehat \Gamma_A)(I-\widehat{B})-(I-P_A)(I-B)]\bar{\alpha}_{T_1} \\ &\quad +(I-\widehat \Gamma_A)(\widehat{B}-B) \frac{1}{\sqrt{T_1}}\sum_{s=1}^{T_1}Y_{T+s}(0)+\sqrt{T_1}(I-\widehat \Gamma_A)(a-\widehat{a})\\ &= (I-\widehat \Gamma_A)\frac{1}{\sqrt{T_1}}\sum_{s=1}^{T_1} u_{T+s}+\sqrt{T_1}(I-P_A)(I-B)\bar{\alpha}_{T_1}+O_p(1). \end{align*} Since $I-\widehat{\Gamma}\rightarrow_p I-\Gamma$ and $u_{T+t}$ is stationary, $(I-\widehat{\Gamma})T_1^{-1/2}\sum_t u_{T+t} =O_p(1)$ by Proposition 7.5 of Hamilton1994. This concludes the proof.