Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
97,656 characters · 7 sections · 87 citation commands
Synthetic Control As Online Linear Regression
Synthetic control abadie2003economic,abadie2015comparative is an increasingly popular method for causal inference among policymakers, private institutions, and social scientists alike. In parallel, there is a rapidly growing methodological literature providing statistical guarantees for synthetic control methods. \oldFootnote{See the review by abadie2021using as well as the special section on synthetic control methods in the Journal of the American Statistical Association abadie2021introduction.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi \Copy{introsent}{Existing results for synthetic control---and for modifications thereof---are typically derived under a low-rank linear factor model or a vector autoregressive model of the outcomes abadie2010synthetic,ben2019synthetic,ben2021augmented,ferman2021synthetic,viviano2019synthetic.} \oldFootnote{Notably, like this paper, bottmer2021design consider a design-based framework which conditions on the outcomes and considers randomness arising solely from assignment of the treated unit or the treatment time period. }\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi While these statistical guarantees formally hold under these outcome models, a number of authors have expressed optimism that the synthetic control method is robust to these modeling assumptions. \oldFootnote{For instance, ben2019synthetic write, “Outcome modeling can also be sensitive to model mis-specification, such as selecting an incorrect number of factors in a factor model. Finally, [... synthetic control] can be appropriate under multiple data generating processes (e.g., both the autoregressive model and the linear factor model) so that it is not necessary for the applied researcher to take a strong stand on which is correct.” jaume write, “Synthetic controls are intuitive, transparent, and produce reliable estimates for a variety of data generating processes.”}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi
On the other hand, in empirical settings where synthetic control is commonly applied---where the treated unit is an aggregate entity like a country or a U.S. state---plausible outcome modeling may be challenging. manski2018right, in studying the effect of gun laws in the United States using state-level crime rates, provocatively ask, “what random process should be assumed to have generated the existing United States, with its realized state-year crime rates?” \Copy{tension}{Granted, the low-rank linear factor model is a general class of data-generating processes and may even arise under finer-grained models on the individual outcomes contained in the aggregate data shi2022assumptions. But to pessimists and skeptics, perhaps even such a model is implausible for the settings considered by many synthetic control studies. Indeed, if practitioners were willing to fully commit to an outcome model, perhaps they should estimate the outcome model directly---e.g., use factor model-based methods bai2002determining,bai2003inferential,xu2017generalized,athey2021matrix---instead of using synthetic control?}
As a result, existing methodological results seem to leave practitioners in a somewhat awkward position. On the one hand, synthetic control is intuitively appealing, and it is conjectured to have good properties under a variety of outcome models. On the other hand, perhaps existing outcome models that have so far proved sufficiently analytically tractable are not always compelling in common empirical settings. To address this tension, this paper provides a few theoretical results and offers a novel interpretation of synthetic control methods. In particular, we seek guarantees for synthetic control that do not rely on any outcome model. Consequently, our results complement existing, model-based ones.
It is unlikely that nontrivial guarantees on the {performance} of synthetic control exist without any structure on the outcomes. However, we {can} derive guarantees of synthetic control's performance relative to a class of alternatives, such as weighted matching or weighted difference-in-differences (DID) estimators, which practitioners may otherwise choose. Our first main result shows that, on average over hypothetical treatment timings, synthetic control predictions are never much worse than the predictions made by any weighted matching estimator. Our second main result shows that the same is true for synthetic control on differenced data versus any weighted DID estimator. These results imply that if there is a weighted matching or DID estimator that performs well, synthetic control likewise performs well. To be clear, these regret guarantees average over hypothetical treatment timings, which can be interpreted as expected loss under random treatment timing, a design-based assumption.
\Copy{intropractice}{Taken together, our results provide reassurances for practitioners, as they offer justifications for synthetic control that do not rely on particular statistical models of the outcomes. At least on average over hypothetical treatment timings, regardless of outcomes, variations of synthetic control are competitive against common estimators, such as weighted matching and weighted DID estimators. Additionally, our second result introduces a novel version of synthetic control that is competitive against DID. Since DID is extremely popular in practice currie2020technology and is thus a natural benchmark, this version of synthetic control may be particularly attractive.}
\Copy{typo1}{We derive our results by casting prediction with panel data as an instance of online convex optimization, and by recognizing synthetic control as an online regression algorithm known as Follow-The-Leader kalai2005efficient. \oldFootnote{For an introduction to online convex optimization, see hazan2019introduction, orabona2019modern, cesa2006prediction, and shalev2011online.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi} Regret guarantees on FTL in the online convex optimization literature translate directly to guarantees for synthetic control against a class of alternative estimators. Since most results in online convex optimization have been derived under an adversarial model---where an imagined adversary generates the data---these results translate to guarantees on synthetic control without any structure on the outcome process.
This paper is perhaps closest to viviano2019synthetic. They propose an ensemble scheme to aggregate predictions from multiple predictive models, which can include synthetic control, interactive fixed effects models, and random forests. Using results from the online learning literature, viviano2019synthetic's ensemble scheme has the no-regret property, making the ensemble predictions competitive against the predictions of any fixed predictive model in the ensemble. Under sampling processes that yield good performance for some predictive model in the ensemble, viviano2019synthetic then derive performance guarantees for the ensemble learner. In contrast, we study synthetic control directly in the worst-case setting, and connect corresponding worst-case results to guarantees on statistical risk in a design-based framework. We show that synthetic control algorithms themselves are no-regret online algorithms and are in fact competitive against a wide class of matching or DID estimators.
(ref) sets up the notation and the decision protocol and presents our main results for synthetic control. (ref) presents several extensions that show alternative guarantees on modifications of synthetic control; in particular, we show that synthetic control on differenced data is competitive against a class of difference-in-differences estimators. (ref) concludes the paper.
Consider a simple setup for synthetic control, following doudchenko2016balancing. There are $T$ time periods and $N + 1$ units. \Copy{tgen}{To simplify convergence rate expressions, we assume $T > N$ unless noted otherwise, but this assumption is not strictly necessary for our results.} Let unit $0$ be the only treated unit, first treated at some time $S \in \{1,\ldots, T\} \equiv[T]$. The other $N$ units are referred to as control units. Since we observe the treated potential outcomes for the treated unit after $S$, estimating causal effects for unit 0 amounts to predicting the unobserved, post-$S$ untreated potential outcomes of this unit. Thus, we focus on untreated potential outcomes.
Let the full panel of untreated potential outcomes be $\mathbf{Y}$ with representative entry $y_{it}$, where (i) $\mathbf{Y}_ {1:s} = (y_{0t},\ldots,y_{Nt})_{t=1}^s$ collects all untreated potential outcomes until and including time $s$, and (ii) $\mathbf{y}_t = (y_{1t}, \ldots, y_{Nt})'$ is the vector of control unit outcomes at time $t$. Additionally, we let $\mathbf{y}(1) = (y_{1}(1),\ldots, y_{T}(1))'$ denote the treated potential outcomes of unit $0$, which are only observable for times $t \ge S$. Similarly, we let $\mathbf{y}(0) = (y_{01},\ldots, y_ {0T})'$ denote the untreated potential outcomes of unit $0$, which are observable for $t < S$. The analyst is tasked with predicting $y_ {0S}$ from observed data, which typically consist of pre-treatment outcomes of unit $0$ and outcomes of untreated units. \Copy{firstpara}{Like the main analysis in doudchenko2016balancing, we do not consider covariates extensively, though (ref) considers matching on covariates as a form of regularization. \oldFootnote{To extend our analysis to cases with covariates, at a minimum, we can interpret $\mathbf{Y}$ as the residuals of the untreated potential outcomes against some fixed regression function of the covariates, i.e. $y_ {it} = y_{it}^* - h_t(x_i)$, for fixed $h_t$ (perhaps estimated from auxiliary data), outcomes $y_{it}^*$, and covariate vectors $x_i$. The residualization is similar to Section 5.5 in doudchenko2016balancing and expression (16) in abadie2021using, but is stronger due to $h_t$ being fixed for different adversarial choices of $\mathbf{Y}$. Our results apply so long as these residuals obey the boundedness assumption $\norm{\mathbf{Y}}_\infty \le 1$ that we impose later.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi }
Synthetic control abadie2003economic,abadie2010synthetic, in its basic form, chooses some convex weights $\smash{\hat\theta_S}$ that minimize past prediction errors
where $\Theta \equiv \{(\theta_1,\ldots,\theta_N) \in \R^N \colon \theta_i \ge 0, 1'\theta = 1\}$ is the simplex. For a one-step-ahead forecast for $y_{0S}$, synthetic control outputs the weighted average $\hat y_{S} \equiv \hat\theta_S'\mathbf{y}_S$, and forms the treatment effect estimate $\hat{\tau}_{S} \equiv y_S(1) - \hat y_S$.
Theoretical guarantees for treatment effect estimates $\hat\tau_S$ often rely on statistical models of the outcomes $\mathbf{Y}$. \Copy{repeatsamp}{While synthetic control has good performance under a range of outcome models, one may still doubt whether these models are plausible---and whether the underlying repeated sampling thought experiments are appropriate---in the spirit of comments by manski2018right.} In contrast to the usual outcome modeling approach, we instead consider a worst-case setting where the outcomes are generated by an adversary. \oldFootnote{The adversarial framework, popular in online learning, dates to the works of hannan20164 and blackwell1956analog.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi Doing so has the appeal of giving decision-theoretic justification for methods while being entirely agnostic towards the data-generating process. Since a dizzying range of reasonable data-generating models and identifying assumptions are possible in panel data settings---yet perhaps none are unquestionably realistic---this worst-case view is valuable, and worst-case guarantees can be comforting.
\Copy{online}{ In particular, we assume an adversary picks the outcomes $\mathbf{Y}$---or, equivalently, we derive results that hold uniformly over $\br{\mathbf{Y}: \norm{\mathbf{Y}}_\infty \le 1}$. Specifically, we consider the following protocol between an analyst and an adversary:
Under such a protocol, the analyst's average squared loss, averaging over hypothetical values of $S$, is
Most results in this paper are guarantees in terms of the decision criterion (ref) for synthetic control, where synthetic control (ref) is viewed as a particular strategy $\sigma$ under (ref).
As the second equality in (ref) indicates, under an additional assumption that treatment timing is uniformly random, $S \sim \Unif [T]$, the average loss over hypothetical treatment timings is equal to the expected squared loss over $S$. This additional assumption is a design-based perspective doudchenko2016balancing,bottmer2021design on the panel causal inference problem. This perspective enables us to interpret average prediction loss over hypothetical treatment timings as expected prediction loss under the random treatment time $S$. The latter can in turn be thought of as design-based risk. Uniformly random assignment of $S$ is restrictive, but we shall relax this requirement in (ref). \oldFootnote{The protocol (ref) easily generalizes when we replace $f(\mathbf{y}_t, \theta_t)$ with any known scalar function and $\ell (\cdot,\cdot)$ with any loss function, so long as $\theta \mapsto \ell(f (\mathbf{y}_t, \theta), y_{0t})$ is convex and bounded. Our results in (ref) allow for general loss functions.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi
We now make clear the connection with online convex optimization hazan2019introduction. Online convex optimization works with the following general protocol. Time $t$ increments sequentially for $T$ periods, and at time $t$:
At the end of the game, the online player suffers total loss $\sum_{t=1}^T \ell_t (\theta_t)$.
Our setup of the panel prediction protocol, (ref), is then an instance of online convex optimization, (ref). To see this, the most important step is to recognize that the analyst's loss (ref) is analogous to the online player's loss, and therefore to think of the analyst as making {sequential decisions} where $\mathbf{Y}$ is sequentially revealed to them. This change in perspective relies on (i) our choice of decision criterion (ref) and (ii) the fact that the analyst's decisions $\theta_t (\cdot)$ only require outcomes prior to $t$. Indeed, by fixing $\mathbf{Y}$ and considering the hypothetical values of $S=1,\ldots, T$ sequentially, we can treat the analyst as if they were solving an online problem and learning from data in the past---even though, for any particular value of $S$, they are only confronted with a static, offline problem. To be clear, we are not considering some online version of synthetic control; the connection to online convex optimization comes from considering hypothetical, unrealized values of $S$.
After viewing the analyst's problem as an online problem, we may straightforwardly establish the remaining correspondences. First, note that the simplex $\Theta$ is convex and bounded. Second, note that we may imagine the adversary in the panel prediction game as picking loss functions $\ell_t (\cdot)$ of the form $\theta \mapsto (y_{0t} - \theta'\mathbf{y}_t)^2$, parametrized by the potential outcomes $(y_{0t}, \mathbf{y}_t)$. These loss functions are indeed convex in $\theta$ and bounded, since both $\theta$ and $\mathbf{Y}$ are bounded. Finally, note that the average loss (ref) is equal to $\frac1T\sum_{t=1}^T \ell_t(\theta_t)$, which is simply the total loss in the online protocol scaled by $\frac1T$. \oldFootnote{\Copy{transpose}{It may be tempting to ask whether the same argument applies to “horizontal regression” athey2021matrix, where one regresses $y_{iS}$ on $y_{i1},\ldots, y_ {iS-1}$, perhaps constraining the coefficients to some bounded, convex set. Since synthetic control can be viewed as a “vertical regression,” where one regresses $y_ {0t}$ on $y_{1t},\ldots, y_{Nt}$, it seems we may apply our argument to the transposed $\mathbf{Y}$ matrix. Indeed, we may formulate analogous claims by replacing $t$ with $i$, $s$ with $j$, $S$ with some randomly chosen unit $M \in [N]$, and $T$ with $N$. However, a difficulty with this interpretation is that synthetic control (ref) naturally only uses information in the past ($t<S$), but the analogous restriction in horizontal regression, $i < M$, for a randomly chosen treated unit $M \in [N]$, is much less natural.}}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi
Having recognized our setup as an instance of online convex optimization, the main observation of this paper recognizes that synthetic control is an online learning algorithm known as Follow-the-Leader (FTL). FTL, under (ref), is the algorithm that, when prompted for a decision in (ref), simply chooses $\theta_t$ to minimize past losses: \oldFootnote{FTL is also known as fictitious play in game theory brown1951iterative. The name “follow-the-leader,” coined by kalai2005efficient, is popular in the recent computer science literature. For an introduction to FTL and similar algorithms, see Chapter 5 in hazan2019introduction and Chapters 1 and 7 in orabona2019modern.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi \oldFootnote{\Copy{ftlunique}{When there are multiple minima, the choice of $\theta_t$ does not affect our theoretical guarantees. Nevertheless, it seems sensible in practice to take the minimum that is smallest in some norm, e.g. $\norm{\cdot}_2$.}}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi \[ \theta_t \in \argmin_{\theta \in \Theta} \sum_{s < t} \ell_s(\theta). \] }
Standard online convex optimization results on regret then apply to synthetic control as well. Before introducing these results, let us define regret as the gap between the total loss of a strategy $\sigma$ and the best fixed weights $\theta$ in hindsight:
(ref) observes that, in our setting, regret is the difference between total squared prediction error of a strategy $\sigma$ and that of the best fixed weights $\theta$ chosen in hindsight, summing over hypothetical treatment times $S$. (ref) interprets the sum of losses as $T$ times the expected loss under random treatment timing. Finally, (ref) observes that regret is an upper bound of the expected error gap between the strategy $\sigma$ and any fixed weights $\theta$. We refer to $\argmin_{\theta\in\Theta} \sum_ {S=1}^T (y_{0S} - \theta'\mathbf{y}_S)^2$ as the oracle weighted match---the best set of weights for a given realization of the data $\mathbf{Y}$.
Focusing on regret rather than loss shifts the goalposts from performance to competition, which is a more fruitful perspective in our adversarial setting. After all, we cannot hope to obtain meaningful loss control as the all-powerful adversary can make the analyst miserable. However, the crucial insight of regret analysis is that, for certain strategies $\sigma$, the adversary cannot simultaneously make the analyst suffer high loss while letting some fixed strategy $\theta$ perform well---in other words, if any fixed $\theta$ performs well, then $\sigma$ performs almost as well over time. Indeed, if regret is sublinear, i.e., $\mathrm{Regret}_T \le o(T)$, \oldFootnote{We mean $\mathrm{Regret}_T \le o(T)$ in the sense that $\limsup_{T\to\infty}\frac{1}{T} \mathrm{Regret}_T \le 0$, since it is possible for $\mathrm{Regret}_T$ to be negative. Following the online convex optimization literature, we sometimes refer to $\sigma$ as no-regret if it has sublinear regret. }\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi then the strategy $\sigma$ never performs much worse than any fixed weights $\theta$, on average over hypothetical treatment timing $S$. In this case, we can interpret $\sigma$ as a strategy that is competitive against the class of weighted matching estimators.
It may seem surprising that these no-regret strategies $\sigma$ exist in the first place. We emphasize that $\sigma$ can output different weights $\theta_t$, chosen adaptively over time, while $\sigma$ is compared to an oracle that uses the best fixed weights. As a result, $\sigma$ can compensate for its lack of oracle access by changing its choices judiciously over time.
The main result of this paper shows that the regret of synthetic control under quadratic loss is logarithmic in $T$. The result follows from a direct application of hazan2007logarithmic's regret bound for FTL (Theorem 5 in their paper, reproduced as (ref) in the appendix).
(ref) shows that the synthetic control strategy (ref) achieves logarithmic regret---and as a result, the average difference between the losses of synthetic control and losses of the oracle weighted match vanishes quickly as a function of $T$. \oldFootnote{Restricting $\theta$ to the simplex $\Theta$---a debated choice in the synthetic control literature---is somewhat important for the dependence on $N$, in so far as the simplex is bounded in $\norm{\cdot}_1$. This is a consequence of the assumption that the outcomes $\mathbf{Y}$ are bounded in the dual norm $\norm {\cdot}_\infty$, which implies a bound on $\theta'\mathbf{y}_t$ that is free of $N,T$. In contrast, if we let $\Theta = \{\theta : \norm{\theta}_2 \le D/2\}$ be an $\ell_2$-ball, then the regret bound worsens to $O(D^2N^2\log (T))$. }\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi In particular, if there exists a weighted average of the untreated units' outcomes that tracks $\mathbf{y}(0)$ well, then the average one-step-ahead loss of synthetic control estimates is only worse by $O\pr{\frac{N\log T}{T}}$.
On its own, (ref) is purely an optimization result; we now offer a few comments on its statistical implications. As a preview, under random treatment timing, (ref) implies that the risk of estimating the causal effect at time $S$ for synthetic control is not too much higher than that for any weighted matching estimator. Indeed, if any weighted matching estimator performs well, then synthetic control achieves low risk as well. Our discussion below translates (ref) into guarantees on the expected loss at treatment time---expressing regret as (ref)---which relies on the design assumption that $S$ is randomly assigned. Nevertheless, we stress that we could view (ref) purely as guarantees of average loss over hypothetical timings $S$---expressing regret only as (ref)---which does not require a treatment timing assumption.
We can interpret regret as a gap in the design-based {risk} of estimating treatment effects. Specifically, we can interpret the expected loss of predicting the untreated outcome as the risk of estimating the treatment effect:
Hence, (ref) and (ref), combined with (ref), imply that the risk of using synthetic control is no more than $N\log T/T$ worse than the risk of the oracle weighted match, \oldFootnote{We slightly abuse notation and use $\theta$ to denote the strategy that outputs $\theta$ every period.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi regardless of the potential outcomes $\mathbf{Y}, \mathbf{y}(1)$:
This observation connects regret on prediction of the untreated potential outcome with differences in the risk of estimating treatment effects. Roughly speaking, (ref) shows that synthetic control estimates of one-step-ahead causal effects are competitive against that of any fixed weighted match, for any realization of $\mathbf{Y}, \mathbf{y}(1)$, on average over $S$.
Of course, since the guarantee (ref) holds for every $\mathbf{Y}$, it continues to hold when we average over $\mathbf{Y}$ and $\mathbf{y}(1)$, over a joint distribution $P$ that respects the boundedness condition $\norm{\mathbf{Y}}_\infty \le 1$. In this sense, analyzing regret in the adversarial framework not only does not preclude statistical interpretations, but rather {facilitates} analysis in a wide range of outcome models. \oldFootnote{The technique of “online-to-batch conversion” in the online learning literature exploits this intuition to prove results in batch (i.i.d.) settings via results in online adversarial settings. }\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi Formally, let $\mathcal P$ be a family of distributions for $\mathbf{Y}, \mathbf{y}(1)$ such that $P (\norm{\mathbf{Y}}_\infty \le 1) = 1$ for all $P \in \mathcal P$. Under an outcome model $P$, we may understand $\mathrm{Risk}(\sigma, \mathbf{Y}, \mathbf{y}(1))$ as conditional risk and $\E_P \mathrm{Risk}(\sigma, \mathbf{Y}, \mathbf{y}(1))$ as unconditional risk. Then, (ref) implies that \oldFootnote{abernethy2009stochastic show that a minimax theorem applies, and \[ \sup_P \inf_\sigma \E_P\bk{\mathrm{Risk}(\sigma, \mathbf{Y}, \mathbf{y}(1))] - \min_{\theta\in \Theta} \mathrm{Risk}(\theta, \mathbf{Y}, \mathbf{y}(1))} = \frac{1}{T} \inf_\sigma \sup_\mathbf{Y} \mathrm{Regret}_T(\sigma, \mathbf{Y}). \] Note that the $\le$ direction is immediate via the min-max inequality. This result shows that the worst-case optimal risk differences in a stochastic setting (i.e. the analyst knows $P$ and responds to it optimally) is equal to minimax regret. In this sense, worst-case regret analysis is not by itself conservative for a stochastic setting---minimax regret is a tight upper bound for performance in stochastic settings. }\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi
Therefore, the unconditional risk of synthetic control is never much worse than the risk of the oracle weighted match \[R_\Theta^* \equiv \E_P\bk{\min_ {\theta\in \Theta} \mathrm{Risk}(\theta, \mathbf{Y}, \mathbf{y} (1))}.\] Hence, if the data-generating process $P$ guarantees that $R_\Theta^*$ is small, then synthetic control achieves low expected risk as well. Concretely speaking, this latter requirement is that, for most realizations of the data, had we observed all the potential outcomes, we could find a weighted match that tracks the potential outcomes $y_{01},\ldots, y_{0T}$ well, so that \oldFootnote{Also, observe that $ \E_P[\min_{\theta \in \Theta} \frac{1}{T}\sum_{t=1}^T (y_{0t} - \theta' \mathbf{y}_t)^2] \le \min_{\theta \in \Theta} \E_P[\frac{1}{T}\sum_{t=1}^T (y_ {0t} - \theta' \mathbf{y}_t)^2], $ and thus the guarantee (ref) is stronger in the sense that it allows the oracle $\theta$ to depend on the realization of the data.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi \[ \E_P\bk{\min_{\theta \in \Theta} \frac{1}{T}\sum_{t=1}^T (y_{0t} - \theta' \mathbf{y}_t)^2} \approx 0. \]
In many empirical settings, it seems plausible that the oracle weighted match performs well. \oldFootnote{We recognize that under many data-generating models, there is unforecastable, idiosyncratic randomness in $y_{0t}$. As a result, there may not exist a synthetic match that perfectly tracks the realized series $y_{0t}$ (even though such a match may exist that tracks various conditional expectations of $y_{0t}$ quite well). In many such cases, since squared error can be orthogonally decomposed, risk differences for estimating $y_ {0t}$ are also risk differences for estimating conditional means $\mu_{t}$ of $y_ {0t}$. We discuss these results in (ref). }\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi abadie2021using states the following intuition in many comparative case studies: “[T]he effect of an intervention can be inferred by comparing the evolution of the outcome variables of interest between the unit exposed to treatment and a group of units that are similar to the exposed unit but were not affected by the treatment.” More formally speaking, a well-fitting oracle weighted match also resembles---and implies---abadie2010synthetic's assumption that there exists a perfect pre-treatment fit of the outcomes. When the oracle weighted match performs well, our regret guarantees imply a guarantee on the loss of the feasible synthetic control estimator, making it an attractive option for causal inference in comparative case studies.
Even if no weighted average of the untreated units tracks $y_{0t}$ closely, synthetic control continues to enjoy the assurance that it performs almost as well as the best weighted match. Moreover, in the general online learning setup (ref), this no-regret property cannot be attained without choosing $\theta_t$ in some data-dependent manner. \oldFootnote{See (ref) for a simple argument in a general setup with unspecified $\ell(\cdot)$. Since simple DID does not choose weights adaptively, it fails to control regret against the class of weighted DID estimators that we discuss in (ref).}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi This observation rules out alternatives such as simple difference-in-differences, which does not aggregate the control units in a data-dependent manner. In contrast, in (ref), we additionally show that synthetic control on differenced data performs almost as well as the best weighted difference-in-differences estimator, a popular class of estimators in practice.
The previous interpretations---in (ref) and (ref)---rely on interpreting average loss over hypothetical values of $S$ as expected loss over $S$, which requires uniform treatment timing $S \sim \Unif[T]$. Despite being plausible in certain settings and appearing elsewhere in the literature doudchenko2016balancing,bottmer2021design, this assumption is perhaps crude. \oldFootnote{doudchenko2016balancing discuss inference in synthetic control via randomization of the treatment timing in their Section 6.2. bottmer2021design consider randomization of the treated period in their Assumption 2, though, in their setting, the treatment lasts only one period. We also note that the randomness per se of $S$ conditional on $\mathbf{Y}$ can be realistic, but that its distribution is uniform and known is restrictive.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi To some extent this is inevitable: Since we are agnostic on the outcome generation process, it is unavoidable to make treatment timing assumptions in order to obtain nontrivial statistical results on estimation of causal quantities. Nevertheless, note that such an assumption is only necessary for interpreting average losses as expected losses. The a priori proposition that it is reasonable to expect a causal estimator to predict well relative to some oracle, at least on average over hypothetical treatment timings, strikes us as defensible. Accepting this dictum relieves us of any need to model treatment timing.
Even if we wish to maintain the interpretation of average loss as expected loss, we can relax the uniform treatment timing assumption. In this subsection, we show that if the treatment timing distribution is known, then a weighted version of synthetic control achieves logarithmic weighted regret. Moreover, even if the treatment timing distribution is non-uniform, unknown, and possibly chosen by the adversary, we continue to show that synthetic control performs well if some weighted average of untreated units predicts $y_{0S}$ accurately. Both results have constants that worsen if the treatment timing distribution deviates far from $\Unif[T]$.
Suppose the conditional distribution $ (S \mid \mathbf{Y})$ is denoted by $\pi =(\pi_1,\ldots,\pi_T)'$, which may depend on $\mathbf{Y}$. Note that, for a known $\pi$, we may apply the same argument in (ref) to the following weighted synthetic control estimator:
by redefining the loss functions $\ell_t(\cdot)$. This argument shows that (ref) achieves $\log T$ weighted regret, stated in the following corollary. Note that (ref) implements FTL with loss functions $\ell_t(\theta) \equiv \pi_t (y_{0t} - \theta'\mathbf{y}_t)^2$, and hence the argument of hazan2007logarithmic applies.
(ref) shows that the weighted regret---a difference in $\pi$-expected loss---is logarithmic in $T$, thereby controlling the worst-case gap between weighted synthetic control and the oracle weighted match for the expected loss. Assuming a known $\pi$ could be reasonable. With a known dynamic treatment regime, $\pi$ can depend on $\mathbf{Y}_{1:S-1}$, but is known whenever the analyst is prompted for a prediction at time $S$. \oldFootnote{Since the bound is for a fixed $\mathbf{Y}$, we can allow $\pi$ to depend on $\mathbf{Y}$, so long as $\pi_t(\mathbf{Y})$ is known at time $t+1$ so that the analyst can compute (ref). This allows for (ref) to be applied in the following example, which is a more realistic design-based setting. There is a known dynamic treatment regime chakraborty2014dynamic parametrizing the treatment hazard: That is, \[ \P(S = t \mid S \ge t, \mathbf{Y}) = r_t(\mathbf{Y}_{1:t-1}) \] for some known $r_t(\cdot)$. Then $ \pi_t(\mathbf{Y}) = \P(S = t \mid \mathbf{Y}) = (1-r_1)\cdots(1-r_{t-1}) r_t $ is a function of $\mathbf{Y}_{1:t-1}$. We thank Davide Viviano for suggesting this extension. }\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi We can also interpret (ref) as providing guarantees on differences in Bayes risk under the analyst's prior $S \sim \pi$, independent of $\mathbf{Y}$.
Even when $\pi$ is unknown and chosen by the adversary, we can bound the loss of unweighted synthetic control, so long as $\pi$ is not too far from uniform.
The result (ref) shows that, uniformly over all bounded $\mathbf{Y}$ and bounded treatment distributions $\pi$, the expected squared error is bounded by the average loss of the oracle weighted match plus the regret, all scaled with a constant $C$ that indexes how far $\pi$ deviates from the uniform distribution. Under the same assumption that the oracle weighted match performs well on average, (ref) continues to show that the treatment estimation risk of synthetic control is small. Since such a result is valid for all $\mathbf{Y}$ and $\pi$, we may understand (ref) as a bound that holds even in a setting where the adversary picks both $\mathbf{Y}$ and $\pi$, with the restriction that $\pi_t \le C/T$, but otherwise unrestricted in the dependence between $\mathbf{Y}$ and $\pi$.
As before, since (ref) is a guarantee uniformly over $\mathbf{Y}$, it is also a guarantee when we average over $\mathbf{Y}$ under an outcome model, yielding (ref). Again, (ref) shows that for any joint distribution of the bounded outcomes and the treatment timing, the unconditional risk of synthetic control is small when the expected oracle conditional risk, $\E_Q[\min_ {\theta \in \Theta} \frac{1}{T} \sum_{t=1}^T (y_{0t} - \theta'\mathbf{y}_t)^2]$, is small---so long as $S$ has sufficient randomness conditional on $\mathbf{Y}$ so that $C$ is not too large.
So far, we have considered weighted averages of untreated units as the class of competing estimators. These competing estimators are matching estimators. However, a more common class of competing estimators in applications are difference-in-differences (DID) estimators. It turns out that synthetic control on preprocessed data has regret guarantees against a class of DID estimators, which we turn to in the next subsection.
(ref) shows that the original synthetic control estimator is competitive against a class of matching estimators that use weighted averages of untreated units as matches for the treated unit. However, in many applications in economics, matching estimators are much less popular than DID estimators, since the latter accounts for unobserved confounders that are additive and constant over time. In this subsection, we show that synthetic control on differenced data is competitive against a large class of DID estimators. Additionally, (ref) offers regret guarantees against other flavors of DID estimators.
In practice, a common DID specification is the following two-way fixed effects regression: \[ \min_{\mu_i, \alpha_t, \lambda} \sum_{i=0}^N \sum_{t=1}^S \pr{y_{it}^{\text{obs}} - \mu_i - \alpha_t - \lambda \one\bk{(i,t) = (0, S)}}^2, \] where the observed outcome $y_{it}^{\text{obs}} = y_{it}$ for all $(i,t) \neq (0,S)$, and $y_{0S}^{\text{obs}} = y_S(1)$. This specification regresses the observed outcomes on unit and time fixed effects, and uses the estimated coefficient $\lambda$ as an estimate of the treatment effect $y_ {S}(1) - y_{0S}$. Implicitly, this regression uses the estimated fixed effects $\mu_0 +\alpha_S$ as a forecast for the unobserved $y_{0S}$. \Copy{sdid1}{ We consider a weighted generalization of this regression, a special case of the synthetic DID estimators in arkhangelsky2021synthetic: \oldFootnote{ The weight $w_0$ does not affect $\mu_0 +\alpha_S$ achieving the optimum in the least-squares problem, per the calculation in (ref). As a result, we normalize $w_0 = 1$. Moreover, specifically, (ref) is a special case of synthetic DID, (1) in arkhangelsky2021synthetic, with only unit-level weights and no time-level weights. }\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi \oldFootnote{(ref) is underdetermined if $S = 1$. The ensuing discussion assumes $\sum_{i=1}^N w_i y_{i1}$ is the weighted two-way fixed effects prediction for $y_ {01}$.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi
} For convex weights $w = (w_1,\ldots, w_N)'$, denote by $\sigma_{ \mathrm{TWFE}}(w)$ the strategy that estimates (ref) on the data $(\mathbf{Y}_{1:t-1}, \mathbf{y}_t)$ at time $t$, \oldFootnote{The value of $y_{0t}$ does not enter $\alpha_S + \mu_0$ since it is absorbed by the coefficient $\lambda$.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi and outputs the estimated coefficients $\mu_0 + \alpha_t$ as a prediction for $y_{0t}$. By varying over $w \in \Theta$, we obtain a class of competing DID strategies, where conventional DID corresponds to picking uniform weights $w = (1/N,\ldots, 1/N)'$. We calculate in (ref) that the prediction that $\sigma_{\mathrm{TWFE}}(w)$ makes is \[ \hat y_{t}(\sigma_{\mathrm{TWFE}}(w)) = \frac{1}{t-1} \sum_{s=1}^{t-1} y_{0s} + w' \pr{\mathbf{y}_t - \frac{1}{t-1} \sum_{s=1}^{t-1} \mathbf{y}_s}\qquad t \ge 2, \] which simply uses the outcome difference against historical averages of untreated units to forecast that of unit $0$. Note that this strategy amounts to using a weighted match with weight $w$ on the differenced data \[\tilde y_{i1} = y_{i1}\qquad \tilde y_{it} \equiv y_{it} - \frac{1}{t-1} \sum_{s=1}^ {t-1} y_ {is} \qquad |\tilde y_{it}| \le 2\] to forecast the same differences of unit $0$, $\tilde y_{0t}$. Therefore, we may apply (ref) and show the following regret bound.
(ref) shows that synthetic control on differenced data controls regret against the class of DID estimators (ref). \oldFootnote{The benchmark class of DID estimators in (ref) output predictions in a sequential manner, in so far as the coefficients in the regression (ref) depend on $S$. In contrast, (ref) compares synthetic control against a class of static DID estimators that do not exhibit this feature.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi In particular, the class of DID benchmarks corresponds to weighted two-way fixed effects regressions, and synthetic control is competitive against any fixed weighting. In this sense, (ref) builds on the intuition that synthetic control is a generalization of DID doudchenko2016balancing to show that a version of synthetic control performs as well as any weighted DID estimator. Again, if any weighted DID estimator performs well, then (ref) becomes a performance guarantee on synthetic control. Moreover, since (ref) is a popular alternative for many practitioners---setting aside whether there is a weighted DID that performs well---(ref) shows that it is without much loss to use synthetic control in such settings instead. \Copy{recommendation}{Since DID is more popular in practice than weighted matching, competitive performance against DID is a more relevant consideration, which suggests prioritizing synthetic control on differenced data $\tilde y_{it}$ over classic synthetic control (ref). \oldFootnote{This comment is with the caveat that the constant in (ref) is worse than that in (ref). It seems possible to further improve the guarantee in (ref), since in our proof, we solely use the implication $|\tilde y_{it}| \le 2$ and do not restrict the adversary from choosing $\tilde y_{it}$ where the implied $|y_ {it}|>1$. We leave such a refinement to future work.
Of course, this observation also implies that (ref) holds without bounded outcomes $ \norm{\mathbf{Y}}_\infty \le 1$ and solely with bounded differences $\max_{i, t} |\tilde y_{it}| \le 2$. }\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi}
\Copy{sdid}{ To the best of our knowledge, the difference scheme $\tilde y_{it}$ has yet to be considered in the literature. We do note that since the resulting predictions are equivalent to a weighted two-way fixed effects regression, this proposed synthetic control scheme can be thought of as synthetic DID arkhangelsky2021synthetic with weights chosen by constrained least-squares on $\tilde y_{it}$. } We also note that $\tilde y_ {it}$ is slightly different from ferman2021synthetic's demeaned synthetic control, which takes the difference $\dot y_{it} \equiv y_{it} - \frac{1}{t} \sum_{s=1}^{t} y_{is}$. In (ref), we show that ferman2021synthetic's demeaned synthetic control achieves logarithmic regret against a different class of DID estimators that we call static DID estimators. \oldFootnote{Under certain conditions, ferman2021synthetic (Proposition 3) show that the demeaned synthetic control in (ref) dominates DID with uniform weighting $\theta_i = 1/N$. The results (ref) are in a similar flavor, and show that synthetic control is competitive against DID with any fixed weighting, on average over random assignment of treatment time. Of course, (ref) are not generalizations of ferman2021synthetic's result---for one, we consider average loss under random treatment timing, and ferman2021synthetic consider a fixed treatment time under an outcome model, with the number of pre-treatment periods tending to infinity.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi Another popular alternative is first-differencing abadie2021using, which by similar arguments may be shown to control regret against a class of two-period weighted DID strategies that output $ \hat y_{t} (\sigma_{\text{2P-DID}}(\theta)) \equiv y_{0 t-1} + \theta'\pr{\mathbf{y}_t - \mathbf{y}_ {t-1}}$ as successive predictions.
(ref) shows that synthetic control, as FTL, gives logarithmic regret when we consider quadratic loss. However, to some extent this bound is an artifact of using squared losses, whose curvature ensures that the FTL predictions do not move around excessively over time. If we replace the loss function with the absolute loss $|\hat y - y|$, then the regret may be linear in $T$---no better than that of the trivial prediction $\hat y_t \equiv 0$ orabona2019modern.
Motivated by the lack of general sublinear regret guarantees in FTL, the online learning literature proposes a large class of algorithms called Follow-The-Regularized-Leader (FTRL), where regularization helps stabilize the FTL predictions. With linear prediction functions $f(\mathbf{y}; \theta) = \theta'\mathbf{y}$, such strategies take the form
for some convex penalty $\Phi$ and regularization strength $1/\eta > 0$. Here, we let $\ell (\cdot, \cdot)$ denote a generic convex and bounded loss function, generalizing our previous framework. Many regularized variants of synthetic control have been proposed chernozhukov2021exact,doudchenko2016balancing,hirshberg2021least. These regularized estimators have the form (ref), though most such estimators are based on quadratic loss.
\Copy{cov}{ Moreover, we can think of synthetic control with covariates as regularized synthetic control as well. With time-invariant covariates $\mathbf{x}_j = (x_{1j},\ldots,x_{Nj})'$ for $j=1,\ldots, J$, synthetic control may choose weights $\theta$ to additionally match the covariates abadie2021using:
for some given $\eta_j$ that indexes the importance of matching covariate $j$. Observe that, for fixed $x_{0j}, \mathbf{x}_{j}$, (ref) is a special case of (ref); in particular, (ref) uses a quadratic penalty of the form \[\Phi (\theta) = \frac{1}{2}(\mathbf{x}-\mathbf{X}\theta)'H(\mathbf{x}-\mathbf{X}\theta)\] for some positive definite $H$, vector $\mathbf{x}$, and conformable matrix $\mathbf{X}$. Thus, under the assumption that the covariates $x_{0j}, \mathbf{x}_j$ are fixed and not chosen by the adversary, we may analyze synthetic control with time-invariant covariates as a special case of FTRL.}
Motivated by the importance of loss function curvature, we slightly generalize and consider regularized synthetic control estimators using generic loss functions. A standard result in online convex optimization (e.g. Corollary 7.9 in orabona2019modern, Theorem 5.2 in hazan2019introduction) shows that choices of $\eta$ exist to obtain $\sqrt{T}$ regret. \oldFootnote{This rate matches the lower bound for linear losses. See Chapter 5 of orabona2019modern.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi The conditions for this result are highly general, explaining the popularity of FTRL in online convex optimization. We specialize to a few choices of the penalty function $\Phi$ in the synthetic control setting; see (ref) for a general statement.
\Copy{reg}{ Naturally, these choices correspond to regularized variants of synthetic control. As we discuss above, quadratic penalties generalize ridge penalization hirshberg2021least and matching on covariates. \oldFootnote{Ridge penalties are a special case of elastic net penalties proposed by doudchenko2016balancing. (ref) applies to elastic net penalties with nonzero $\ell_2$ component as well.
Note that when $\mathbf{X} \in \R^{J \times N}$ represents pre-treatment covariates of the control units, $\mathbf{X}'H\mathbf{X}$ being positive definite requires that the dimension of the covariates is at least the number of control units. }\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi The entropy penalty, which is very natural when the parameters lie on the simplex, is a special case of the proposal in robbins2017framework; the resulting regret bound has better dependence on $N$ and obtains the no-regret property as long as $\frac{\log N}{T} \to 0$. \oldFootnote{ Interestingly, $\ell_1$-penalty chernozhukov2021exact alone is not strongly convex boyd2004convex, and (ref) does not apply. However, (ref) only contains sufficient conditions, and so this alone is not a criticism of $\ell_1$-penalty.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi For these guarantees, the choice of $\eta$ does require knowledge on the total number of periods $T$. This may be relaxed via the “doubling trick” (see shalev2011online, Section 2.3.1), if we allow for different regularization strengths $\eta_S$ for different realizations of $S$.}
We conclude this section by pointing out a few other extensions. First, another weakening of the uniform treatment timing requirement can be achieved by considering the maximal regret over subperiods of $[T]$, also known as adaptive regret. We show in (ref) that a modification to the synthetic control algorithm---which still outputs a weighted average of untreated units---achieves worst subperiod regret of order $\log T$. Such a result implies that if we additionally let the adversary pick a subperiod of length $T'$, and treatment is uniformly randomly assigned on this subperiod, then modified synthetic control is at most $\frac{\log T}{T'}$-worse on expected loss than the oracle weighted match. Of course, this regret guarantee is meaningful only when the subperiod is sufficiently long, i.e., $T'\gg \log T$. Second, under a design-based framework on treatment timing, we can test sharp hypotheses of the form $H_0 : \mathbf{y}(1) - \mathbf{y}(0) = \mathbf{z}$ by leveraging symmetries induced by random treatment timing. We briefly discuss inference in (ref).
This paper notes a simple connection between synthetic control methods and online convex optimization. Synthetic control is an instance of Follow-The-Leader, which are well-studied strategies in the online learning literature. We present standard regret bounds for FTL that apply to synthetic control, which have interpretations as bounds for expected regret under random treatment timing. These regret bounds translate to bounds on expected risk gap under outcome models and imply that synthetic control is competitive against a wide class of matching estimators. In cases where some weighted match of untreated units predict the unobserved potential outcomes, these results show that synthetic control achieves low expected loss. Moreover, the regret bounds can be adapted to be regret bounds against difference-in-differences strategies. Lastly, we draw an analogous connection between regularized synthetic control and Follow-the-Regularized-Leader, a popular class of strategies in online learning.
\Copy{limit}{ We now point out a few limitations of this paper and directions for future work. First and foremost, the approach we have taken in this paper is deliberately pessimistic. Living in fear of an adversary constrained solely by bounded outcomes is perhaps too paranoid for sound decision-making. For instance, this worst-case perspective is not particularly amenable to incorporating covariates, since matching on covariates is inherently based on the hope that the covariates are predictive of potential outcomes. Further constraining the adversary rakhlin2011online may be an interesting direction for future research. For instance, it may be fruitful to consider an adversary with a fixed budget for how much $y_{0t}, \mathbf{y}_t$ deviate from $y_{0,t-1}, \mathbf{y}_{t-1}$. Constraining the adversary may also render covariates useful, even in a worst-case framework. }
It may also be interesting to consider alternative online protocols. So far, we have considered a thought experiment where, before each step $t$, the analyst only has access to data $\mathbf{Y}_{1:t-1}$ to output a prediction function. In practice, the analyst typically does have access to $\mathbf{y}_1,\ldots, \mathbf{y}_T$. Alternative protocols have been considered in the online learning literature. One example is the Vovk--Azoury--Warmuth forecaster orabona2019modern, where we assume the analyst additionally has access to $\mathbf{y}_t$ before they are prompted for a prediction at time $t$. In this case, regularized strategies can also achieve $\log T$ regret. Additionally, bartlett2015minimax consider the fixed design setting in which $\mathbf{y}_1,\ldots, \mathbf{y}_T$ is fully accessible to the analyst before they are prompted for a prediction. bartlett2015minimax give a simple and explicit minimax regret strategy for online linear regression, which we may adapt into a synthetic control estimator.
\Copy{multi}{We have only considered regret on one-step-ahead prediction for $y_ {0S}$, but synthetic control estimates are often extrapolated multiple time periods ahead in practice. In attempting to extend our results to $k$-step-ahead prediction, it is natural to consider $\check y_{it} = (y_{it},\ldots, y_{i,t + k})$, and to attempt a similar argument on $\check \mathbf{Y}$. The chief difficulty in doing so is one of delayed feedback, where the analyst cannot update their time-$S$ decision based on loss from times $1,\ldots, S-1$. That is, for $k$-step-ahead prediction, the analyst, viewed as an online player who is prompted for a forecast of $\check y_{0,S} = (y_{0S}, y_{0,S+1}, \ldots, y_ {0,S+k-1})$, does not have access to their prediction loss for $\check y_{0,S-1} = (y_{0,S-1}, y_{0S}, \ldots, y_{0,S+k-2})$, since $y_{0, S+k-2}$ is not yet observed. As a result, unlike (ref) in the standard online convex optimization protocol, the analyst does not have access to $\ell_1 (\cdot),\ldots, \ell_{S-1} (\cdot)$ when making decisions $\theta_S$---rendering our results here insufficient. That said, delayed feedback---where the online player only has knowledge of the loss function after $k$ periods---is studied in online learning weinberger2002delayed,korotin2018aggregating,flaspohler2021online, and we leave an exploration to future work.}