EconBase
← Back to paper

Synthetic Controls for Experimental Design

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

118,621 characters · 0 sections · 63 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\setcounter{page}{0} \thispagestyle{empty}

center[center omitted — 1,101 chars of source]
abstractThis article studies experimental design in settings where the experimental units are large aggregate entities (e.g., markets), and only one or a small number of units can be exposed to the treatment. In such settings, randomization of the treatment may result in treated and control groups with substantially different baseline characteristics, inducing biases. We propose a variety of experimental non-randomized synthetic control designs abadie2003synthetic, abadie2010synthetic that select the units to be treated, as well as the untreated units to be used as a control group. Average potential outcomes with treatment are estimated as weighted averages of observed outcomes for treated units, and average potential outcomes without treatment as weighted averages of observed outcomes for control units. We analyze the properties of estimators based on synthetic control designs and propose new inferential techniques. We show that in experimental settings with aggregate units, synthetic control designs can substantially reduce estimation biases in comparison to randomization of the treatment.

\setcounter{page}{1} {0.5\baselineskip}

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\center\normalfont\bf}{Introduction}

Consider the problem of a ride-sharing company choosing between two compensation plans for drivers (doudchenkonddesigning, n.d.; jones2019uber, jones2019uber). The company can either keep the current compensation plan or adopt a new one with higher incentives. In order to estimate the effect of a change in compensation plans on profits, the company's data science unit designs an experimental evaluation where the new plan is deployed at a small scale, say, in one of the local markets (cities) in the country. In this setting, a randomized control trial\,---\,or A/B test, where drivers in a local market are randomized into the new plan (active treatment arm) or the status quo (control treatment arm)\,---\,is problematic. On the one hand, such an experiment raises equity concerns, as drivers in the same local market but in different treatment arms obtain different compensations for the same jobs. On the other hand, if drivers in the active treatment arm respond to higher incentives by working longer hours, they will effectively steal business from drivers in the control arm of the experiment, resulting in biased experimental estimates.\footnote{A randomized evaluation across many markets is a potential solution to the problem of experimental interference between drivers. In practice, however, large-scale market-level randomized evaluations are often unfeasible. In the context of the ridesharing company example, large-scale market-level randomized evaluations {\em (i)} could be prohibitively expensive, {\em (ii)} could still raise substantial equity concerns, {\em (iii)} could negatively affect morale for the large number of drivers in the treated cities if the program is rolled back after experimentation, and {\em (iv)} in some cases, the number of cities where the company operates could be too small for effective randomization.}

One possible approach to this problem is to assign an entire local market to treatment, and use the rest of the local markets, which remain under the current compensation plan during the experimental periods, as potential comparison units. In this setting, using randomization to assign the active treatment allows ex-ante (i.e., pre-randomization) unbiased estimation of the effect of the active treatment. However, ex-post (i.e., post-randomization) biases can be large if, at baseline, the treated unit differs from the untreated units in the values of the features that affect the outcomes of interest. We document the magnitude and practical relevance of these biases in Sections (ref) and (ref).

As in the ride-sharing example where there is only one treated local market, large biases may arise more generally in randomized studies when either the treatment arm or the control arm contains a small number of units, so randomized treatment assignment may not produce treated and control groups that are similar in their features. In those cases, the fact that estimation biases would have averaged out over alternative treatment assignments is of little comfort to a researcher who, in practice, is limited to one assignment only.

To address these challenges, we propose using the synthetic control method abadie2003synthetic, abadie2010synthetic as an experimental design to select treated units in non-randomized experiments, as well as the untreated units to serve as a comparison group. We adopt the name {\em synthetic control designs} to refer to the resulting experimental designs.\footnote{While we leave the “experimental” qualifier implicit in “synthetic control design”, it should be noted that the synthetic control designs proposed in this article differ from observational synthetic control designs abadie2003synthetic, abadie2010synthetic, doudchenko2016balancing, for which the identity of the treated unit(s) is taken as given.}$^,$\footnote{See, e.g., abadie2021using, amjad2018robust, arkhangelsky2019synthetic, doudchenko2016balancing for background material on synthetic controls and related methods.}

In our framework, the choice of the treated unit (or treated units, if multiple treated units are desired) aims to accomplish two goals. First, the treated units should be representative of an aggregate of interest, such as a national market, so that the estimated effect reflects the aggregate impact of the treatment. Second, the treated units should not be idiosyncratic in the sense that the untreated units cannot closely approximate their features. Otherwise, the reliability of the estimate of the effect on the treated unit may be questionable. We show how to achieve these two objectives, whenever they are possible to achieve, using synthetic control methods.

While we are aware of the extensive use of synthetic control methods for experimental design in data science units, especially in the technology industry,\footnote{See, in particular, jones2019uber, which also provides the basis for the ride-sharing example above.} the academic literature on this subject is at a nascent stage. There are, however, a few publicly available studies that are connected to this article. Aside from the present article, to our knowledge, doudchenkonddesigning (n.d.) and doudchenko2021synthetic are the only other publicly available studies on the topic of experimental design with synthetic controls. The focus of doudchenkonddesigning (n.d.) is on statistical power, which they calculate by simulating the estimated effects of placebo interventions using historical (pre-experimental) data. That is, the selection of treated units is based on a measure of statistical power implied by the distribution of the placebo estimates for each unit. As a result, estimates based on the procedure in doudchenkonddesigning (n.d.) target the effect of the treatment for the unit or units that are most closely tracked in the placebo distribution. In the same spirit, the target parameter in doudchenko2021synthetic is the treatment effect for a weighted average of treated units that can be closely matched in their pre-treatment outcomes by a weighted average of untreated units. In the present article, we aim to take a different perspective on the problem of unit selection in experiments with synthetic controls; one that takes into account the extent to which different sets of treated and control units approximate an aggregate causal effect of interest chosen by the analyst, such as the average treatment effect for the relevant population.\footnote{Consistent with the majority of literature on synthetic controls, our focus is primarily on average treatment effects. For an analysis of distributional effects using synthetic controls, see gunsilius2023distributional.} The inferential methods in the present article also differ from those in the related literature. In particular, doudchenko2021synthetic proposes a permutation procedure for inference that requires that potential outcomes without the treatment are independent and identically distributed (i.i.d.) in time. In contrast, the inferential procedure proposed in the present article allows for time series dependence and non-stationarity in outcomes, which are pervasive features of time-series data. Another important difference between the present article and doudchenkonddesigning (n.d.) and doudchenko2021synthetic is that doudchenkonddesigning (n.d.) and doudchenko2021synthetic make use of pre-treatment outcomes only to select treated and control units, while our method allows the use of other observed features of the units.

A related literature applies synthetic control methods and nearest-neighbor matching methods to select experimental sites in multi-site designs egamidesigning,olea2024externally. In contrast, we examine settings where treatment occurs at the aggregate (i.e., site) level: each site receives only treatment or only control, precluding the estimation of site-level treatment effects.

agarwal2021synthetic proposes synthetic interventions, a framework related to synthetic controls, and applies it to estimate treatment effect heterogeneity in an experimental setting with multiple treatments. Their work primarily focuses on the analysis of experimental data, but not on the design of experiments. bottmer2021designbased is also related to the present article in the sense that it studies synthetic control estimation in an experimental setting. Their article, however, considers only the case when the treatment is randomized, and is not concerned with issues of experimental design.

The research designs in this article are also related to ex-ante synthetic control designs for observational studies (see abadie2021using, abadie2021using, kasy2023employing, kasy2023employing, and chen2023synthetic, chen2023synthetic, the latter for an online version of the problem) in that the synthetic weights are computed and can be pre-registered before post-intervention outcomes are realized. However, our methods differ significantly in one key aspect: they confront the challenge that experimenters face when selecting specific units for exposure to the intervention of interest. In a wider context, our methods are rooted in the broader framework of experimental non-randomized designs kasy2016why,armstrong2018finite,thorlund2020synthetic. Yet, they diverge by addressing a distinct challenge: estimating synthetic control-based aggregate counterfactuals in experimental settings where only a limited number of aggregate units can be treated.

An alternative approach to control post-randomization bias involves stratifying units based on covariate values prior to randomization of treatment within each stratum. Stratification can significantly reduce post-randomization biases if units have similar covariate values within strata. However, traditional stratification methods do not adapt to the setting considered in this article, which features a limited number of large aggregate entities as units of analysis and a single unit or a handful chosen for treatment. Because every stratum in stratified designs must have at least one unit randomized into treatment, the number of strata cannot exceed the desired number of treated units in the experiment. In the case of only one treated unit, we would be limited to a single stratum. This may lead to significant variation in units' characteristics within strata, reducing the appeal of stratification procedures. Additionally, a stratified design with a small number of strata or treated units may result in the selection of a set of treated units that is not truly representative of the target population.

The rest of the article is organized as follows: Section (ref) presents and discusses the synthetic control designs proposed in this article. Section (ref) details the formal properties of estimators based on synthetic control designs and proposes inferential methods. In Section (ref), we report the findings from an empirical validation of synthetic control designs using sales data from a sample of Walmart stores. Section (ref) discusses the results of simulation studies. Finally, Section (ref) provides concluding remarks. The Appendix contains proofs and supplemental materials.

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\center\normalfont\bf}{Synthetic Control Designs} We consider a setting with $T$ time periods and $J$ units, which may represent $J$ local markets as in the ride-sharing example in the previous section. Let $T_0$ be the number of pre-experimental periods, with $1\leq T_0<T$. At the end of period $T_0$, a researcher designs an experiment to conduct during periods $T_0+1, T_0+2, \ldots, T$. Using the information available at $T_0$, the experimenter aims to select the set of units that will receive the treatment (intervention) during the experimental periods.

To define causal parameters, we formally adopt a potential outcomes framework. For any $j \in \{1,\ldots, J\}$ and any $t \in \{T_0+1, \ldots, T\}$, $Y^I_{jt}$ is the potential outcome for unit $j$ at time $t$ when unit $j$ is exposed to treatment starting at $T_0+1$. Similarly, for any $j \in \{1,\ldots, J\}$ and any $t \in \{1, \ldots, T\}$, $Y^N_{jt}$ is the potential outcome for unit $j$ at time $t$ under no treatment. In the ridesharing example, $Y^I_{jt}$ and $Y^N_{jt}$ could measure net revenue divided by market size under the active and the control treatment, respectively. Unit-level treatment effects are defined as \[ Y^I_{jt} - Y^N_{jt}, \] for $j = 1,\ldots, J$ and $t = T_0+1, \ldots, T$. They represent the effect of switching unit $j$ to the active treatment at time $T_0+1$ on the outcome of unit $j$ at time $t> T_0$. We aim to estimate the average treatment effect

align[align omitted — 86 chars of source]

for $t=T_0+1, \ldots, T$. In this expression, $f_1, \ldots, f_J$ are known positive weights that define the average of interest. In the ride-sharing example from the previous section, $f_j$ may represent the size of local market $j$ as a share of the national market. Without loss of generality, and because it is often the case in applications, we can assume that the weights $f_j$ sum to one, \[ \sum_{j=1}^J f_j = 1. \] When units are equally weighted, we set $f_j=1/J$ for $j=1,\ldots, J$. We use the notation $\bm f$ for a vector that collects the values of $f_j$ for all the units, i.e., $\bm f=(f_1,\ldots, f_J)$.

At time $T_0$, in order to estimate the treatment effect $\tau_t$ for $t=T_0+1, \ldots, T$, the experimenter chooses $\bm w=(w_1, \ldots, w_J)$ and $\bm v = (v_1, \ldots, v_J)$, such that

gather[gather omitted — 185 chars of source]

Units with $w_j>0$ are units that will be assigned to the intervention of interest from $T_0+1$ to $T$, and will be used to estimate average outcomes under the intervention. Units with $w_j=0$ constitute an untreated reservoir of potential control units (a “donor pool”). Among units with $w_j=0$, those with $v_j>0$ will be used to estimate average outcomes under no intervention.

The first goal of the experimenter is to choose $w_1,\ldots, w_J$ such that

align[align omitted — 96 chars of source]

for $t=T_0+1, \ldots, T$. If equation (ref) holds, a weighted average of outcomes for the units selected for treatment reproduces the average outcome with treatment for the entire population of $J$ units. In practice, however, the choice of $w_1,\ldots, w_J$ cannot directly rely on matching the population average of $Y_{jt}^I$, as in equation (ref). The quantities $Y_{jt}^I$ are unobserved before time $T_0+1$, and will remain unobserved in the experimental periods for the units that are not exposed to the treatment. Instead, we aim to approximate equation ((ref)) using predictors observed at $T_0$ of the values of $Y_{jT_0+1}^I, \ldots, Y_{jT}^I$. Note also that it is not possible to use the weights $w_1=f_1, \ldots, w_J=f_J$, because it would leave no units in the donor pool, making the set of units with $v_j>0$ empty and violating equation (ref).

The second goal of the experimenter is to choose $v_1,\ldots, v_J$ such that

align[align omitted — 97 chars of source]

or, alternatively,

align[align omitted — 98 chars of source]

If equations (ref) or (ref) hold, a weighted average of outcomes for the units in the donor pool reproduces the average outcome without treatment for the entire population of $J$ units (equation (ref)), or for the units selected for treatment (equation (ref)). Like in the previous case with treated outcomes, it is not feasible to directly choose $v_1, \ldots, v_J$ so that equation (ref) or (ref) is satisfied. Instead, we propose a variety of methods to approximate either ((ref)) or ((ref)) based on predictors of $Y_{jT_0+1}^N, \ldots, Y_{jT}^N$.

For the treated units, we define $Y_{jt}=Y^N_{jt}$ if $t=1,\ldots, T_0$, and $Y_{jt}=Y^I_{jt}$ if $t=T_0+1,\ldots, T$. For the untreated units, we define $Y_{jt}=Y^N_{jt}$, for all $t=1, \ldots, T$. That is, $Y_{jt}$ is the outcome observed for unit $j=1,\ldots, J$ at time $t=1,\ldots, T$. We say that \[ \sum_{j=1}^J w_jY_{jt}\quad\mbox{ and }\quad \sum_{j=1}^J v_jY_{jt} \] are the synthetic treated and synthetic control outcomes, respectively. The difference between these two quantities is \[ \tau_t(\bm w,\bm v) = \sum_{j=1}^J w_jY_{jt}- \sum_{j=1}^J v_jY_{jt}, \] for $t=T_0+1, \ldots, T$. Suppose that equations ((ref)) and ((ref)) hold. Then, $\tau_t(\bm w,\bm v)$ is equal to the average treatment effect, $\tau_t$. If equation ((ref)) holds instead, then $\tau_t(\bm w,\bm v)$ is equal to the average effect of the treatment on the treated ($\bm w$-weighted), \[ \tau^T_t = \sum_{j=1}^J w_j \cdot (Y^I_{jt}-Y^N_{jt}) \] doudchenko2021synthetic.

We choose $\bm w=(w_1, \ldots, w_J)$ and $\bm v=(v_1, \ldots, v_J)$ to match the pre-intervention values of predictors of the potential outcomes $Y_{jt}^N$ and $Y_{jt}^I$ for $t > T_0$.

Let $\bm X_j$ be a column vector of pre-intervention features of unit $j$. We view the features in $\bm X_j$ as predictors of the values of $Y^N_{jt}$ and $Y^I_{jt}$ in the experimental periods, in a sense that will be made precise in Section (ref). We use the notation \[

\hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm X} \kern-0.0em } } } =\sum_{j=1}^J f_j \bm X_j. \] That is, $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm X} \kern-0.0em } } } $ is the vector of population values for the predictors in $\bm X_j$. For any real vector $\bm x$, $\|\bm x\|$ is the Euclidean norm of $\bm x$, and $\|\bm x\|_0$ is the number of non-zero coordinates of $\bm x$. Let $\underline m$ and $\widebar m$ be positive integers such that $1\leq \underline m\leq \widebar m\leq J-1$. A simple selector of $\bm w=(w_1, \ldots, w_J)$ and $\bm v=(v_1, \ldots, v_J)$ is

align[align omitted — 982 chars of source]

The first term of the objective function in ((ref)) measures the discrepancies between the population average of the features in $\bm X_j$ ($\bm f$-weighted) and the averages of the features for units assigned to the treatment group ($\bm w$-weighted). The second term is analogous but with the second average taken over the units assigned to no intervention ($\bm v$-weighted). The first four constraints require that the weights in $\bm w$, as well as the weights in $\bm v$, are non-negative and sum to one. They also require that any unit selected for treatment cannot be utilized as a control unit --- so, if $w_j>0$, then $v_j=0$. The last constraint allows a minimum and maximum number of units assigned to treatment. This restriction is of practical importance in a variety of contexts, especially when experimentation is costly and the experimenter is restricted in the number of units that may receive the treatment. We say that the design is Unconstrained if $\underline m=1$ and $\widebar m=J-1$; otherwise, we say the design is Constrained. The last constraint in ((ref)) is not the only conceivable restriction to the size or cost of the experiment. An explicit upper bound on the cost of an experiment would be given by $\bm c'\bm d\leq B$, where the $j$-th coordinate of $\bm c$ is equal to the cost of assigning unit $j$ to treatment, $\bm d$ is a $J$-dimensional vector with ones at coordinates where $w_j>0$, and zeros otherwise, and $B$ is the experimenter's budget.

Let $\bm w^*=(w^*_1, \ldots, w^*_J)$ and $\bm v^*=(v^*_1, \ldots, v^*_J)$ be a solution to the optimization problem in ((ref)). In practice, we do not require optimality of $(\boldsymbol w^*, \boldsymbol v^*)$, as long as $(\boldsymbol w^*, \boldsymbol v^*)$ is feasible and satisfies $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm X} \kern-0.0em } } } -\sum_{j=1}^J w^*_j\bm{X}_j\approx \boldsymbol 0$ and $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm X} \kern-0.0em } } } -\sum_{j=1}^J v^*_j\bm{X}_j\approx \boldsymbol 0$, where $\boldsymbol 0$ is a vector of zeros of the same dimension as $\boldsymbol X_j$. Suppose that units with $w^*_j>0$ are assigned to treatment in the experiment, and units with $w^*_j=0$ are kept untreated. A synthetic control estimator of $\tau_t$ is $\widehat\tau_t=\tau_t(\bm w^*,\bm v^*)$, i.e.,

equation[equation omitted — 110 chars of source]

This estimator is based on approximations to equations ((ref)) and ((ref)) that rely on $\bm X_j$, the observed predictors of the potential outcomes, $Y_{jt}^N$, and $Y_{jt}^I$. Note that for every solution to ((ref)) with $\underline m \leq \|\bm v\|_0 \leq \widebar m$, there exists another solution that swaps the roles of the treated and the untreated in the experiment without altering the value of the objective function.

In what follows, we take the weight selector in ((ref)) as a starting point for synthetic control designs and modify it in several ways. A second formulation of the synthetic control design is based on equations (ref) and (ref),

align[align omitted — 833 chars of source]

The parameter $\beta > 0$ reflects the trade-off between selecting treated units to fit the aggregate value of the predictors $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm{X}} \kern-0.0em } } } $, and selecting control units to fit the aggregate value of the treated units. A small value of $\beta$ favors designs with treated units that closely match $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm{X}} \kern-0.0em } } } $. A large value of $\beta$, on the other hand, favors designs with aggregate treated and aggregate control units that closely match each other. While it is possible to use data-driven selectors of $\beta$, the rule of thumb $\beta = 1$ provides a natural choice that equally weights the two terms in the objective function in (ref). For this formulation of the synthetic control design, the treatment assignment and estimation procedures follow the same steps as those used in the previous formulation. Large values for $\beta$ produce estimators that target the $\bm w$-weighted average effect of the treatment on the treated, $\tau_t^T$ of doudchenko2021synthetic. Small values of $\beta$ prioritize estimation of the average treatment effect, $\tau_t$.

In our third formulation of the synthetic control design, the experimenter selects a synthetic treated unit to match the average values of the characteristics in the population. However, unlike the design in (ref), the experimenter chooses multiple synthetic controls, one for each unit that contributes to the synthetic treated unit. For any $J$-dimensional vector of non-negative coordinates, $\bm w=(w_1, \ldots, w_J)$, let $\mathcal J_{\bm w}$ be the set of the indices with non-zero coordinates, $\mathcal J_{\bm w} = \{j: w_j>0\}$. Our next version of the synthetic control design is:

align[align omitted — 1,085 chars of source]

The parameter $\xi>0$ arbitrates potential trade-offs between selecting treated units to fit the aggregate value of the predictors $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm X} \kern-0.0em } } } $ and selecting control units to fit the values of the predictors for the treated units. A small value of $\xi$ favors experimental designs with treated units that closely match $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm X} \kern-0.0em } } } $. A large value of $\xi$, on the other hand, favors designs where the values of the predictors for the treated units are closely matched by those of their synthetic controls.

Let $\{w^*_j, v^*_{ij}\}_{i,j=1,\ldots,J}$ be a solution of the optimization problem in ((ref)). As before, we do not strictly require optimality of $\{w^*_j, v^*_{ij}\}_{i,j=1,\ldots,J}$, provided $\{w^*_j, v^*_{ij}\}_{i,j=1,\ldots,J}$ is feasible and $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm X} \kern-0.0em } } } -\sum_{j=1}^J w^*_j\bm{X}_j\approx \boldsymbol 0$ and $\bm X_j-\sum_{j=1}^J v^*_{ij}\bm{X}_j\approx \boldsymbol 0$ for all $j$ such that $w^*_j>0$. Assign units with $w^*_j>0$ to treatment in the experiment, and keep units with $w^*_j=0$ untreated. Let

align[align omitted — 89 chars of source]

Then,

align[align omitted — 196 chars of source]
figure[figure omitted — 3,169 chars of source]

Our next adjustment to the synthetic control design is motivated by settings where experimental units may be naturally divided into clusters with similar values in the predictors, $\bm X_1, \ldots, \bm X_J$. For example, weather patterns, which may be highly dependent across cities in the same region (e.g., Northeast, Midwest, etc., in the US), may influence the seasonality of the demand for ride-sharing services. In those cases, it is natural to treat each cluster (each region, in our example) as a distinct experimental design to ameliorate interpolation biases. Figure (ref) illustrates this point. Panels (a) and (b) depict identical samples in the space of the predictors. In this simple example, we have two predictors only, and their values for each unit are represented by the coordinates of the dots in the figure. Red dots represent units assigned to treatment. All other units are plotted as black dots. Panel (a) visualizes the result of treating the entire sample as one cluster. Three units are assigned to treatment. They closely reproduce the value of $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm X} \kern-0.0em } } } $, but they all fall in the same central cluster, far away from observations in other clusters. In panel (b), assignment to treatment takes into account the clustered nature of the data, and one unit is treated per cluster. This provides a better approximation of the distribution of the predictor values for the entire sample, ameliorating concerns of interpolation biases.

Suppose we divide the set of $J$ available units into $K$ clusters. Let $\mathcal I_k$ be the set of indices for the units in cluster $k$. The cluster mean is \[

\hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm X} \kern-0.0em } } } _k=\sum_{j\in \mathcal I_k} f_j \bm X_j \Big/\sum_{j\in \mathcal I_k} f_j, \] for each cluster $k=1,\ldots, K$. For each index $i = 1,\ldots,J$, let $k(i)$ be the cluster to which unit $i$ belongs, i.e., $i \in \mathcal{I}_{k(i)}$. A clustered version of the synthetic control design in (ref) is given by:

align[align omitted — 1,306 chars of source]

We conclude this section by discussing other possible extensions to the synthetic control design. First, it is well known that synthetic control estimators may not be unique. Lack of uniqueness is typical in settings where the values of the predictors that a synthetic control is targeting (i.e., $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm X} \kern-0.0em } } } $ in equation (ref), or $\bm X_j$ for a treated unit in equation (ref)) fall inside the convex hull of the values of $\bm X_j$ for the units in the donor pool. To address the potential lack of uniqueness, we adapt the penalized estimator of abadie2021penalized to the synthetic control designs proposed in this article. The penalized synthetic control estimator of abadie2021penalized is unique provided that predictor values for the units in the donor pool are in general quadratic position abadie2021penalized. Moreover, penalized synthetic controls favor solutions where the synthetic units are composed of units that have predictor values, $\bm X_j$, similar to the target values. Applying the penalized synthetic control of abadie2021penalized to the objective function of (ref), we obtain

align[align omitted — 1,450 chars of source]

Here, $\lambda_1$ and $\lambda_2$ are positive constants that penalize discrepancies between the target values of the predictor $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm{X}} \kern-0.0em } } } $ and the values of the predictors for the units that contribute to their synthetic counterparts.\footnote{See abadie2021penalized for details on penalized synthetic control estimators. The synthetic control design in (ref) is a penalized version of (ref). Section (ref) in the Online Appendix discusses how to apply the abadie2021penalized penalty to the other synthetic designs proposed in this article.}

Other types of penalization are possible. In particular, doudchenko2016balancing, doudchenko2021synthetic, and others have proposed synthetic control estimators that use ridge or elastic net regularization on the synthetic control weights (e.g., on $w_j$ and $v_j$ in design (ref)). The synthetic control designs proposed in this article can be modified to incorporate regularization on the weights.

Finally, abadie2021penalized, arkhangelsky2019synthetic, and augmented2021feller have proposed bias-correction techniques for synthetic control methods. Section (ref) in the Online Appendix provides details on how to apply bias correction techniques in a synthetic control design.

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\center\normalfont\bf}{Formal Results}

We introduce an extension of the linear factor model commonly employed in the synthetic control literature and use it to analyze the properties of estimators based on synthetic control designs.

assumptionPotential outcomes follow a linear factor model, \begin{subequations} \begin{align} Y^N_{jt} & = \delta_t + \bm{\theta}_t' \bm{Z}_j + \bm{\lambda}_t' \bm{\mu}_j + \epsilon_{jt}, \\ Y^I_{jt} & = \upsilon_t + \bm{\gamma}_t' \bm{Z}_j + \bm{\eta}_t' \bm{\mu}_j + \xi_{jt}, \end{align} \end{subequations} where $\bm{Z}_j$ is a $(R \times 1)$ vector of observed covariates, $\bm{\theta}_t$ and $\bm{\gamma}_t$ are $(R \times 1)$ vectors of unknown parameters, $\bm{\mu}_j$ is a $(F \times 1)$ vector of unobserved covariates, $\bm{\lambda}_t$ and $\bm{\eta}_t$ are $(F \times 1)$ vectors of unknown parameters, and $\epsilon_{jt}$ and $\xi_{jt}$ are unobserved random shocks.

Equation ((ref)) is the linear factor model for potential outcomes under no treatment, a benchmark commonly used in the literature to analyze the properties of synthetic control estimators abadie2010synthetic, ferman2021properties. Equation ((ref)) extends the linear factor structure to potential outcomes under treatment. The reason for this extension is that, in contrast to synthetic control estimation with observational data, synthetic control designs require the choice of a treatment group in addition to the choice of a comparison group.

We employ the covariates in $\bm Z_j$ as well as pre-experimental values of the outcome variable $Y_{jt}$ to construct the vectors of predictors, $\bm X_j$. In particular, let $\mathcal E\subseteq \{1,\ldots, T_0\}$, let $T_\mathcal{E}=|\mathcal E|$, and let $\bm Y^{\mathcal E}_j$ be the $(T_\mathcal{E}\times 1)$ vector of $T_\mathcal{E}$ pre-experimental outcomes for unit $j$ and time indices in $\mathcal E$. We define \[ \bm X_j = \left(

array[array omitted — 43 chars of source]

\right), \] for $j=1, \ldots, J$. That is, the vector of predictors $\bm X_j$ collects the covariates in $\bm Z_j$ and the pre-experimental outcome values $Y_{jt}$ for the {\em fitting periods} in $\mathcal E$. In practice, the values in $\bm X_j$ are often scaled to make them independent of units of measurement or to reflect the relative importance of each of the predictors abadie2021using.

The next assumption gathers regularity conditions on model primitives.

assumption\leavevmode \begin{enumerate}[label=(\roman*)] • $F \leq T_\mathcal{E}$. Moreover, let $\bm{\lambda}_{\mathcal E}$ be the $(T_\mathcal{E}\times F)$ matrix with rows equal to the $\bm\lambda_t$'s indexed by $\mathcal E$. Let $\zeta_{\mathcal E}$ be the smallest eigenvalue of $\bm{\lambda}_{\mathcal E}' \bm{\lambda}_{\mathcal E}$. Then, $\underline{\zeta}=\zeta_{\mathcal E}/T_\mathcal{E}>0$. • For each $j=1, \ldots, J$, $\epsilon_{j1}, \ldots, \epsilon_{jT}$ is a sequence of i.i.d. sub-Gaussian random variables with mean zero and variance proxy $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\sigma} \kern-0.0em } } } ^2$. For any $j=1, \ldots, J$, $\xi_{jT_0+1}, \ldots, \xi_{jT}$ is a sequence of i.i.d. sub-Gaussian random variables with mean zero, variance proxy $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\sigma} \kern-0.0em } } } ^2$, and independent of $\epsilon_{j1}, \ldots, \epsilon_{jT}$. \end{enumerate}

Assumption (ref){\it (i)} is similar to conditions in abadie2010synthetic. Assumption (ref){\it (ii)} is similar to conditions in abadie2010synthetic, doudchenko2016balancing, chernozhukov2021exact, and arkhangelsky2019synthetic. Sub-Gaussianity is not strictly necessary, but it simplifies the form of our results. It can be relaxed by assuming bounded finite-order moments (instead of bounding the entire moment generating function). At the same time, sub-Gaussianity is a relatively mild assumption. It holds for any Gaussian distribution, as well as any distribution with a bounded support. Distributions with heavy tails, such as the Cauchy distribution, are not sub-Gaussian. Notably, Assumption (ref){\it (ii)} allows for dependence of $\epsilon_{jt}$ and $\xi_{jt}$ across units.

Unless otherwise noted, all probability statements are over the joint distribution of $\epsilon_{jt}$ and $\xi_{jt}$ and conditional on the values of the other components on the right-hand sides of equations (ref) and (ref). The next assumption pertains to the quality of the synthetic control fit. For concreteness, we focus on the base design in (ref), and choose where $\bm w^*=(w^*_1, \ldots, w^*_J)$ and $\bm v^*=(v^*_1, \ldots, v^*_J)$ so that the synthetic treated and synthetic control units reproduce the average values of $\bm X_j$.

assumptionWith probability one, {\it (i)} \begin{subequations} \begin{align} \sum_{j = 1}^J w^*_j \bm{Z}_{j} = \sum_{j = 1}^J v^*_j \bm{Z}_{j}=\sum_{j = 1}^J f_j \bm{Z}_{j}, \end{align} and {\it (ii)} \begin{align} \sum_{j = 1}^{J} w^*_j \bm{Y}^{\mathcal E}_j=\sum_{j = 1}^{J} v^*_j \bm{Y}^{\mathcal E}_j = \sum_{j = 1}^{J} f_j \bm{Y}^{\mathcal E}_j. \end{align} \end{subequations}

Assumption (ref) implies that the synthetic treated and control units defined by $\bm w^*$ and $\bm v^*$ provide a perfect fit for $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm X} \kern-0.0em } } } $. Assumption (ref) is a strong restriction, which may only hold approximately in practice. The next assumption relaxes the perfect fit condition in Assumption (ref).

assumptionThere exists a positive constant $d > 0$, such that with probability one, {\it (i)} \begin{subequations} \begin{align} \Big\| \sum_{j = 1}^J w^*_j \bm{Z}_{j} - \sum_{j = 1}^J f_j \bm{Z}_{j} \Big\|_2^2 \leq R d^2, \qquad \Big\| \sum_{j = 1}^J v^*_j \bm{Z}_{j} - \sum_{j = 1}^J f_j \bm{Z}_{j} \Big\|_2^2 \leq R d^2, \end{align} and {\it (ii)} \begin{align} \Big\|\sum_{j = 1}^{J} w^*_j \bm{Y}^{\mathcal E}_j - \sum_{j = 1}^{J} f_j \bm{Y}^{\mathcal E}_j \Big\|_2^2 \leq T_\mathcal{E} d^2, \qquad \Big\|\sum_{j = 1}^{J} v^*_j \bm{Y}^{\mathcal E}_j - \sum_{j = 1}^{J} f_j \bm{Y}^{\mathcal E}_j \Big\|_2^2 \leq T_\mathcal{E} d^2. \end{align} \end{subequations}

Let $\lambda_{t,f}$ be the $f$-th coordinate of $\bm\lambda_t$, and \[ \widebar\lambda = \max_{\substack{t=1, \ldots , T\\f=1, \ldots, F}} |\lambda_{tf}|. \] We define $\eta_{t f}$, $\theta_{t r}$, $\gamma_{t r}$, $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\eta} \kern-0.0em } } } $, $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\theta} \kern-0.0em } } } $ and $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\gamma} \kern-0.0em } } } $ analogously, so $|\eta_{t f} | \leq \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\eta} \kern-0.0em } } } $ for $t=T_0+1,\ldots, T$, $f=1, \ldots ,F$, $| \theta_{t r} | \leq \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\theta} \kern-0.0em } } } $ for $t=1,\ldots, T$, $r=1, \ldots ,R$, and $| \gamma_{t r} | \leq \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\gamma} \kern-0.0em } } } $, for $t=T_0+1,\ldots, T$, $r=1, \ldots ,R$. Next theorem extends results on the bias of synthetic control estimators abadie2010synthetic, vives2022predictor to the experimental set-up of Section (ref).

theoremIf Assumptions (ref) -- (ref) hold, then for any $t \geq T_0+1$, \begin{align} |E \left[ \widehat\tau_t - \tau_t \right]| \leq \frac{ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\lambda} \kern-0.0em } } } ( \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\eta} \kern-0.0em } } } + \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\lambda} \kern-0.0em } } } ) F}{\zeta} \sqrt{2\log{(2J)}}\frac{ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\sigma} \kern-0.0em } } } }{\sqrt{T_\mathcal{E}}}. \end{align} If Assumptions (ref), (ref), and (ref) hold, then for any $t \geq T_0+1$, \begin{align} |E \left[ \widehat\tau_t - \tau_t \right]| \leq \Big(( \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\gamma} \kern-0.0em } } } + \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\theta} \kern-0.0em } } } )R + \frac{ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\lambda} \kern-0.0em } } } ( \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\eta} \kern-0.0em } } } + \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\lambda} \kern-0.0em } } } ) F}{\zeta} (1+ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\theta} \kern-0.0em } } } R)\Big) d + \frac{ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\lambda} \kern-0.0em } } } ( \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\eta} \kern-0.0em } } } + \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\lambda} \kern-0.0em } } } ) F}{\zeta} \sqrt{2\log{(2J)}}\frac{ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\sigma} \kern-0.0em } } } }{\sqrt{T_\mathcal{E}}}. \end{align}

Note that, while the factor model in equations (ref) and (ref) leave the sign and scale of $\bm{\lambda}_t$ and $\bm{\eta}_t$ free (e.g., multiplying $\bm{\lambda}_t$ and dividing $\bm{\mu}_t$ by the same non-zero constant does not change the value of $\bm{\lambda}_t'\bm{\mu}_j$), the value of the bound in Theorem (ref) is invariant to changes in the sign or the scale of $\bm{\lambda}_t$ and $\bm{\eta}_t$. Moreover, the bound in (ref) does not depend on the scale of ${\bm Z}_j$, because changing the scale of $\bm{Z}_j$ leaves the product $\widebar\theta d$ unchanged. The scale of $Y_{jt}$ does affect the bound in (ref) because the treatment effect $\tau_t$ is measured in the same units as $Y_{jt}$. The results in Theorem (ref) do not depend on the specific formulation of the synthetic control design (e.g., {\it Constrained} vs. {\it Unconstrained}).

The bias bounds (ref) and (ref) depend on the ratio between the scale of $\epsilon_{jt}$, represented by $\widebar\sigma$, and the number of fitting periods $T_\mathcal{E}$. Intuitively, the bias of the synthetic control estimator is small when a good fit in pre-experimental outcomes (Assumption (ref)) is obtained by implicitly fitting the values of the latent variables, $\mu_j$. Overfitting happens when pre-experimental outcomes are instead fitted out of the variability in the individual transitory shocks, $\epsilon_{jt}$. A small number of fitting periods $T_\mathcal{E}$ combined with large variability in $\epsilon_{jt}$ increases the risk of overfitting and, as a result, increases the bias bound. Similarly, for any fixed value of $T_\mathcal{E}$, the bias bound increases with $J$, reflecting the increased risk of over-fitting caused by increased variability in $\epsilon_{jt}$ over larger donor pools. Finally, the number of unobserved factors $F$ enters the bound (ref) linearly, which highlights the importance of including the observed predictors $\bm Z_j$ ---\,other than pre-experimental outcomes\,--- in the vector of fitting variables $\bm X_j$. Under the factor model in equations (ref) and (ref), observed predictors not included in ${\bm Z}_j$ are shifted to ${\bm\mu}_j$, increasing $F$ and the magnitude of the bound.\footnote{Shifting predictors from ${\bm Z}_j$ to ${\bm\mu}_j$ changes the bias bound (ref) in a more complex manner than what might be inferred from a cursory look at the bias formula. First, moving predictors from ${\bm Z}_j$ to ${\bm\mu}_j$ also means shifting components of ${\bm\theta}_t$ to ${\bm\lambda}_t$, which can change the value of $\underline\zeta$. Poincar\'{e}'s separation theorem implies that $\underline\zeta$ cannot increase as a result of this shift. Moreover, moving predictors from ${\bm Z}_j$ to ${\bm\mu}_j$ cannot decrease the values of $\widebar\lambda$ and $\widebar\eta$. Overall, the value of the bias bound in (ref) cannot decrease and will typically increase by moving predictors from ${\bm Z}_j$ to ${\bm\mu}_j$. This is not necessarily true for the bound in (ref), because a shift of components from ${\bm Z}_j$ to ${\bm\mu}_j$ decreases the value of $R$.}

We next turn our attention to inference. We utilize a set of {\em blank periods}, $\mathcal B\subseteq \{1,\ldots, T_0\}\setminus {\mathcal E}$, which comprise pre-experimental periods whose outcomes $Y_{jt}$ have not been used to calculate $\bm w^*$ or $\bm v^*$. Because pre-experimental periods that are not in $\mathcal E$ or $\mathcal B$ are not used in our procedure, without loss of generality, we consider $\mathcal B= \{1,\ldots, T_0\}\setminus {\mathcal E}$. We, therefore, assume that the number of elements of $\mathcal B$ is $T_{\mathcal B}=|\mathcal B|=T_0-T_\mathcal{E}$. We aim to test the null hypothesis:

tcolorbox[breakable, boxrule=.5pt, colback=white, right=0pt, left=-.4cm, enlarge left by=.8cm, enlarge right by=-.1cm, width=\linewidth-.7cm, arc=0pt] \begin{quote} For $t=T_0+1, \ldots, T$, and $j=1, \ldots, J$,\end{quote} \begin{equation} Y^I_{jt} = \delta_t+\bm\theta_t'\bm Z_j + \bm\lambda_t'\bm\mu_j + \xi_{jt}, \end{equation} \begin{quote}where $\xi_{jt}$ has the same distribution as $\epsilon_{jt}$.\tikz[overlay,remember picture] \node (br) ; \end{quote}

Under the null hypothesis in (ref), the distribution of $Y^{I}_{jt}$ is the same as the distribution of $Y^{N}_{jt}$, for $t=T_0+1, \ldots, T$, and $j=1, \ldots, J$. But the realized values of $Y^{I}_{jt}$ and $Y^{N}_{jt}$ may differ.

Recall from (ref) that, for $t \in \{T_0+1, \ldots, T\}$, a synthetic control estimator is defined as

align*[align* omitted — 85 chars of source]

Let $\widehat{u}_t = \widehat{\tau}_t, \forall t \in \{T_0+1, \ldots, T\}$ be the synthetic control estimator in the experimental periods. Similarly, for each $t\in \mathcal B$ in the blank periods, let

align*[align* omitted — 81 chars of source]

Such $\widehat{u}_t$ for $t \in \mathcal{B}$ are placebo treatment effects estimated for the blank periods. We study the properties of a test based on combinations from the set $\{\widehat{u}_t : t\in\mathcal B\cup \{T_0+1, \ldots, T\}\}$.

We define $\Pi$ as the set of all $(T-T_0)$-combinations of $\mathcal{B} \cup \{T_0+1, \ldots, T\}$. That is, for each $\pi\in\Pi$, $\pi$ is a subset of indices from the blank periods and the experimental periods $\mathcal{B} \cup \{T_0+1, \ldots, T\}$, such that $|\pi| = T-T_0$. The cardinality of $\Pi$ is $|\Pi| = (T-T_\mathcal{E})!/((T-T_0)!(T_0-T_\mathcal{E})!)$. For each $\pi \in \Pi$, let $\pi(i)$ be the $i^{\text{th}}$ smallest value in $\pi$, and

align*[align* omitted — 115 chars of source]

In addition, let $\widehat{\bm e} = (\widehat{u}_{T_0+1},\ldots, \widehat{u}_{T}) = (\widehat \tau_{T_0+1}, \ldots, \widehat \tau_{T})$. This is a vector of treatment effect estimates from the experimental periods. For any $(T-T_0)$-dimensional vector $\bm e =(e_1,\ldots, e_{T-T_0})$, we adopt the test statistic,

align[align omitted — 105 chars of source]

Other choices of test statistics are possible, such as those based on an $L_p$-norm of $\bm e$ chernozhukov2021exact and one-sided versions of the resulting test statistics (i.e., with the positive or the negative parts of $e_t$ replacing $|e_t|$ in equation ((ref))).

The $p$-value of a permutation test on (ref) is

align[align omitted — 178 chars of source]

Theorem (ref) below shows that if $\bm{\lambda}_t$ are exchangeable random variables for $t\in \mathcal B\,\cup\,\{T_0+1, \ldots, T\}$, then a test of the null hypothesis in (ref) based on the $p$-value in (ref) is exact.

theoremSuppose that Assumptions (ref), (ref){\it (ii)}, and (ref){\it (i)} hold. Assume that $\{\bm{\lambda}_t\}_{t \in \mathcal{B} \cup \{T_0+1,...,T\}}$ is a sequence of exchangeable random variables independent of $\{\epsilon_{jt}\}_{t \in \mathcal{B} \cup \{T_0+1,...,T\}}$ and $\{\xi_{jt}\}_{t \in \{T_0+1,...,T\}}$. Then under the null hypothesis (ref), we have \begin{align} \alpha - \frac{1}{|\Pi|} \leq \Pr(\widehat{p} \leq \alpha) \leq \alpha, \end{align} for any $\alpha\in [0,1]$, where $\Pr(\widehat{p} \leq \alpha)$ is taken over the distribution of $\{\xi_{jt},\epsilon_{jt},\bm{\lambda}_t\}$.

Note that, under the assumptions of Theorem (ref), the potential outcome series $Y^N_{jt}$ is allowed to be non-stationary through the term $\delta_t+\bm{\theta}_t' \bm{Z}_j$ in equation (ref). This is in contrast to a related result in doudchenko2021synthetic, which requires that the potential outcomes $Y^N_{jt}$ are i.i.d. over time.

The assumptions in Theorem (ref) build upon those in Theorem (ref). Although these assumptions are simple and sufficient for the result of the theorem, they can be substantially relaxed. Under exchangeability of $\bm{\lambda}_t$, if Assumption (ref){\it (ii)} is violated, the result for Theorem (ref) holds if for each $j=1, \ldots, J$, $\{\epsilon_{jt}\}_{t \in \mathcal{B} \cup \{T_0+1,...,T\}}$ and $\{\xi_{jt}\}_{t \in \{T_0+1,...,T\}}$ are sequences of exchangeable random variables. Second, if Assumption (ref){\it (i)} is violated, the result for Theorem (ref) holds if $\{({\bm \theta}_t, {\bm\lambda}_t)\}_{t \in \mathcal{B} \cup \{T_0+1,...,T\}}$ is a sequence of exchangeable random variables independent of $\{\epsilon_{jt}\}_{t \in \mathcal{B} \cup \{T_0+1,...,T\}}$ and $\{\xi_{jt}\}_{t \in \{T_0+1,...,T\}}$. In the above two cases under exchangeability of $\bm{\lambda}_t$, we still have exact $p$-value. Finally, exchangeability of $\bm{\lambda}_t$ is a strong restriction. Theorem (ref) in the Online Appendix relaxes this restriction by showing that for fixed $\bm{\lambda}_t$ (i.e., without resorting to exchangeability of $\bm{\lambda}_t$), the $p$-value in (ref) is still approximately valid for large $T_\mathcal{E}$.

In some settings, the number of possible combinations, $|\Pi|$, could be very large, making exact calculation of $\widehat p$ computationally expensive. In those instances, random samples from $\Pi$ can be used to approximate the $p$-value in equation (ref).

The inferential technique proposed in this article is related to, but distinct from, the permutation methods in abadie2010synthetic, chernozhukov2021exact, chernozhukov2019distributional, lei2021conformal, firpo2018synthetic, and others. Inferential methods that reassign treatment across units abadie2010synthetic are not appropriate for the designs of Section (ref), which explicitly select treated and control units to satisfy an optimality criterion.

Similar to chernozhukov2021exact, our method is based on rearrangements of estimated treatment effects across time periods. But unlike chernozhukov2021exact, which proposes permutations over all periods, including the pre-intervention periods, our inferential method permutes only over the blank periods and post-intervention periods, which are not used to estimate the weights in the synthetic control design. Relative to chernozhukov2021exact, the generative models of equations (ref) and (ref), which allow for unobserved factors, and the finite sample nature of the results require a novel testing procedure that, similar to split conformal prediction methods vovk2005algorithmic, lei2018distribution, takes advantage of the availability of blank periods.

Confidence intervals for $\tau_t$ can be constructed using split conformal inference methods. For any $\alpha \in (0,1)$, let

align[align omitted — 301 chars of source]

be the empirical $(1-\alpha)$-quantile on the absolute values of placebo treatment effects in the blank periods, and

align[align omitted — 257 chars of source]

We next show that the confidence interval defined in (ref) approximately achieves correct point-wise coverage in large samples if treatment does not change the distribution of the idiosyncratic noises.

theoremAssume that Assumptions (ref)-- (ref) hold. Assume there exists a constant $\kappa<\infty$, such that for all $j=1, \ldots, J$, $t=1, \ldots, T$, $\epsilon_{jt}$ are continuously distributed with the probability density function upper bounded by $\kappa$. Assume that for $t=T_0+1, \ldots, T$, and $j=1, \ldots, J$, $\xi_{jt}$ has the same distribution as $\epsilon_{jt}$. Then the confidence interval defined in (ref) approximately achieves point-wise coverage, i.e., for any $\alpha \in (0,1)$ and any $t \in \{T_0+1,...,T\}$, as $(T_0-T_\mathcal{E}), T_\mathcal{E} \to +\infty$, \begin{multline*} \bigg\vert \Pr\Big( \tau_t \in \widehat{C}_{1-\alpha}(Y_{1t}, Y_{2t},...,Y_{Jt}) \Big) - (1-\alpha)\bigg\vert \\ = O\Big(\big(\log{(T_0-T_\mathcal{E})}/(T_0-T_\mathcal{E})\big)^{1/2} + \big(\log{T_\mathcal{E}}/T_\mathcal{E}\big)^{1/2}\Big) \longrightarrow 0. \end{multline*}

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\center\normalfont\bf}{Empirical Illustration Using Walmart Data}

In this section, we illustrate the applicability of the methods in this article using store-level data from Walmart prakash2023. The dataset is a balanced panel of weekly sales for $J=45$ Walmart stores and $T=143$ weeks, spanning the period from the week of February 5, 2010, to the week of October 26, 2012. We estimate the effect of a placebo intervention and show that, in the presence of a good pre-intervention fit, the methods of Section (ref) produce point estimates that are close to zero and a test result that does not reject the null hypothesis in (ref) for the placebo intervention.

We consider the design of a fictitious experiment across stores taking place on July 20, 2012 (week $129$ in the data). Out of the $T_0 = 128$ pre-experimental weeks, we take the first $T_\mathcal{E} = 100$ weeks as the fitting period, and the last $(T_0-T_\mathcal{E}) = 28$ weeks as the blank period. The number of weeks in the experimental period is $T-T_0=15$. The outcomes $\{Y_{jt}\}_{j=1,...,J, t=1,...,T}$ are weekly sales (units of revenue are undisclosed in the data). We use uniform weights $f_j = 1/J$ for $j=1,...,J$, to average sales across all stores. For the purpose of estimating the synthetic treated and synthetic control weights, we normalize each of the 100 pre-experimental outcomes to have a unit variance.

figure[figure omitted — 602 chars of source]
figure[figure omitted — 712 chars of source]

We compute synthetic treated and control units that apply the {\it Constrained} formulation in (ref) with $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{m} \kern-0.0em } } } = 2$. We adopt $\widebar m =2$ because using only one store for the synthetic treated fails to produce a good fit between the resulting synthetic treated and synthetic control units during the fitting period. Increasing to $\widebar m =3$ brings only marginal improvements in fit. Figures (ref) and (ref) report results for $\widebar m=2$. Results for $\widebar m=1$ and $\widebar m=3$ appear in the Online Appendix.

Figure (ref) reports the time series of weekly sales for the synthetic treated unit (black solid line), the synthetic control unit (black dashed line), and for each individual store in the dataset (blue dashed lines). Weekly sales for the synthetic treated and the synthetic control units closely follow each other during the fitting period. The gap between the two synthetic units remains small after the fitting period, indicating good out-of-sample predictive power in the absence of intervention.

Figure (ref) reports the difference in weekly sales between the synthetic treated and the synthetic control units. The $p$-value of equation (ref), calculated over the residuals of Figure (ref), is equal to $0.933$, which results in a failure to reject the null hypothesis (ref). Confidence intervals based on equation (ref) cover zero for all $t$ in the experimental period.

Table (ref) compares the performance of the synthetic control design to those of straight randomization followed by difference-in-means, randomization after stratification on pre-intervention outcomes followed by difference-in-means, and 1- and 5-nearest neighbor adjustment after randomization. In particular, Table (ref) reports out-of-sample root mean square error (RMSE) over the post-intervention period, normalized by the post-intervention outcome mean (see Section (ref) for a precise definition of the estimators and RMSE performance metric). For each of the three randomization-based estimators, the reported RMSE is the average over 1000 randomized treatment assignments. The synthetic control design dominates all other alternatives, even when it uses only the outcomes in the fitting periods to construct the synthetic treated and synthetic control units, whereas stratification and nearest-neighbor adjustment utilize all pre-intervention outcomes.

table[table omitted — 2,473 chars of source]

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\center\normalfont\bf}{Simulation Study}

This section presents simulation results that showcase the behavior of estimators based on synthetic control designs. We consider a setting with $J = 15$ units, $R=7$ observable covariates, and $F=11$ unobservable covariates. We simulate data for a total of $T=30$ periods, comprising $T_0=25$ pre-experimental periods and $T-T_0=5$ experimental or post-intervention periods. We compute weights during the first $T_\mathcal{E}=20$ periods and leave periods $t=21, \ldots, 25$ as blank periods. We set the weights $f_j$ in expression (ref) to be $f_j = 1/J$, for all $j=1,...,J$.

For our baseline simulation design, we use the factor model in Assumption (ref) to generate potential outcomes. For $t = 1, \ldots, T$, we generate the series $\delta_t$ and $\upsilon_t$ as small-to-large re-arrangements of $T$ i.i.d. Uniform $(0,20)$ random variables. For $j = 1, \ldots, J$, we set both $\bm{Z}_j$ and $\bm{\mu}_j$ to be random vectors of i.i.d. Uniform $(0,1)$ random variables. For $t = 1, \ldots, T$, we set $\bm{\theta}_t$, $\bm{\gamma}_t$, $\bm{\lambda}_t$, and $\bm{\eta}_t$ to be random vectors of i.i.d. Uniform $(0,10)$ random variables. Finally, for $j = 1, \ldots, J$, and any $t = 1, \ldots, T$, we set $\epsilon_{jt}$ and $\xi_{jt}$ to be i.i.d. Normal $(0,\sigma^2)$ random variables, with $\sigma^2=1$. We present additional simulation results of alternative values of the noise parameter $\sigma^2$ in Section (ref) in the Online Appendix.

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Results for a Single Simulation}

Using the data generating process described above, we draw a single sample and conduct the synthetic control design in (ref), with parameters $\underline{m}=1$ and $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{m} \kern-0.0em } } } =14$, i.e., no constraint on the number of treated units. We report the results in Figures (ref) and (ref). In Figure (ref), each blue dashed line represents an outcome trajectory $Y_{jt}$, for $t=1,\ldots,T$ and $j=1, \ldots J$. The solid black line represents the trajectory of the synthetic treated unit $\sum_{j=1}^J w^*_j Y_{jt}$, for $t=1,\ldots,T$. The black dashed line represents the trajectory of the synthetic control unit $\sum_{j=1}^J v^*_j Y_{jt}$, for $t=1,\ldots,T$. The synthetic treated and synthetic control units closely track each other in the pre-experimental periods. They diverge during the experimental periods, when a treatment effect emerges as a result of the differences in the parameters of the data-generating processes for $Y^N_{jt}$ and $Y^I_{jt}$. Figure (ref) reports the difference between the synthetic treated and the synthetic control outcomes. The inferential procedure of Section (ref) produces $p$-value equal to $0.004$ for the null hypothesis of no treatment effect in (ref).

figure[figure omitted — 408 chars of source]
figure[figure omitted — 421 chars of source]

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Performance Across Many Simulations}

This section compares the performance of the different varieties of the synthetic control designs over 1000 simulations that independently generate the model primitives (i.e., the factor loadings, covariates, and error terms) of Assumption (ref). The data generating process is the same as in Section (ref).

We consider five varieties of the synthetic control design:

enumerate[noitemsep,topsep=0pt] • {\it Unconstrained design}: This is the design in (ref) without a cardinality constraint, so $\underline{m}=1$ and $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{m} \kern-0.0em } } } =J-1=14$. • {\it Constrained design}: Same as the design in (ref), but with $\underline{m}=1$ and $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{m} \kern-0.0em } } } =1,\ldots, 7$. • {\it $\textit{Weakly-targeted}$ design}: This is the design in (ref). We vary $\beta$ from $0.01$ to $100$. • {\it Unit-level design}: This is the design in (ref), which fits a different synthetic control to each unit assigned to treatment. We vary $\xi$ from $0.01$ to $100$. • {\it Penalized design}: This is the design in (ref), with $\lambda=\lambda_1=\lambda_2$. We vary $\lambda$ from $0.01$ to $100$.

The Constrained design imposes sparsity in the synthetic treatment weights through a hard cardinality constraint specified by the integer $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{m} \kern-0.0em } } } $. The Weakly-targeted design targets the average treatment effect for small values of $\beta$ and a weighted average effect for the treated for large values of $\beta$. For the Unit-level design, large values of $\xi$ generate sparsity in the synthetic treated weights. A sufficiently large value of $\xi$ produces a Unit-level design where the only single treated unit can be closely fitted by a convex combination of the other units. For large values of $\lambda$, the Penalized design behaves like a one-to-one matching design, assigning all the weight to one treated and one control unit.

For the {\it Unit-level} design, synthetic control weights are aggregated as in (ref). For the {\it Unconstrained} and {\it Penalized} designs, the synthetic treated and synthetic control weights can always be swapped without changing the objective values for their respective designs. For the {\it Constrained} design, the weights can be swapped when $\|\bm{v}^*\|_0\leq \widebar m$. When it is possible to swap synthetic treated and synthetic control weights, we choose the treated units so that the number of units with positive weights in $\bm w^*$ is smaller than the number of units with positive weights in $\bm v^*$. When $\|\bm w^*\|_0=\|\bm v^*\|_0$, we determine whether to swap using a specific rule described in Section (ref) of the Online Appendix.

\@startsection{subsubsection}{3}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont}{Average Treatment Effects}

table[table omitted — 6,444 chars of source]

The first panel of Table (ref) reports average treatment effects, $\tau_t$, over $1000$ simulations. The second panel reports estimates of the average treatment effects, mean absolute error, root mean square error, and $p$-value, all averaged over 1000 simulations, as well as rejection rates. Mean absolute error (MAE) and root mean square error are defined as

align[align omitted — 190 chars of source]

and the $p$-value is defined as in (ref). Because the treatment effect is not equal to zero in the simulation of Table (ref), smaller $p$-values and larger rejection rates reflect better performance of the testing procedure for a particular design.

In Table (ref), the Unconstrained design has a strong relative performance. The performance of the Constrained design improves for larger $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{m} \kern-0.0em } } } $, and is virtually identical to the performance of the Unconstrained design when $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{m} \kern-0.0em } } } = 7$. The performance of $\textit{Weakly-targeted}$ and Unit-level designs is best when $\beta$ and $\xi$ take intermediate values. The Penalized design yields results similar to those of the Unconstrained design for small values of the penalization parameter $\lambda$.

\@startsection{subsubsection}{3}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont}{Performance with Nonlinearities}

We now examine the behavior of estimators based on synthetic control designs under deviations from the linear model in (ref) and (ref). We consider a nonlinear data generating process,

subequations\begin{align} Y^N_{jt} & = \delta_t + \exp{(\bm{\theta}_t' \bm{Z}_j)} + \exp{(\bm{\lambda}_t' \bm{\mu}_j)} + \epsilon_{jt}, \\ Y^I_{jt} & = \upsilon_t + \exp{(\bm{\gamma}_t' \bm{Z}_j)} + \exp{(\bm{\eta}_t' \bm{\mu}_j)} + \xi_{jt}. \end{align}

The motivation to study a nonlinear model is that nonlinearities may induce interpolation biases, affecting the relative performance of the different designs. All parameter values are the same as in the simulation setup of section (ref), except for the values of $\bm{\theta}_t$, $\bm{\gamma}_t$, $\bm{\lambda}_t$, and $\bm{\eta}_t$, which are chosen to be random vectors of i.i.d. Uniform $(0,3)$ random variables, instead of Uniform $(0,10)$, to control the magnitude of the exponential components in the nonlinear design.

table[table omitted — 6,449 chars of source]

Table (ref) reports the results for $\tau_t$. In comparison to the results in Table (ref), we now see that the Unit-level and Penalized designs can easily match and in some cases improve the performance of the Unconstrained design. By fitting each treated unit with a unit-specific synthetic control, the Unit-level design can ameliorate interpolation biases induced by the aggregation of ${\bm X}_j$. The Penalized design selects synthetic treated and control units close to $ \hbox{ \vbox{ \hrule height 0.5pt \kern0.35ex \hbox{ \kern-0.0em \ensuremath{\bm X} \kern-0.0em } } } $ in the space of the predictors, which can reduce interpolation biases at the potential cost of lower precision for large values of $\lambda$ (in which case, the Penalized design employs a small number of units in the synthetic treated and synthetic control).

\@startsection{subsubsection}{3}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont}{Test size}

table[table omitted — 6,503 chars of source]

In this section, we generate the model primitives under the null hypothesis (ref). That is, we employ a data generating process such that the values of the common factors and the distributions of the idiosyncratic error variables are unaffected by the intervention.

We report the simulation results in Table (ref), which organizes information in the same way as in Table (ref). Because the data are generated from the same distribution under treatment and under no treatment, the average treatment effects in Table (ref) are close to zero. The same is true for the averages of $\widehat\tau_t$ for all designs. Under the null hypothesis (ref), the $p$-value should approximately follow a uniform distribution between zero and one. The results in Table (ref) show good behavior of our testing procedure under the null hypothesis: average $p$-values and rejection rates are close to $0.5$ and $0.05$, respectively.

\@startsection{subsubsection}{3}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont}{Comparison to Randomized Treatment Assignment}

Randomized treatment assignment produces ex-ante (pre-randomization) unbiased estimation of the average treatment effect. As we show below, however, ex-post (post-randomization) biases can be large, especially when only a small number of units are treated.

table[table omitted — 3,113 chars of source]

In this section, we adopt the same set-up as for Table (ref). We consider randomized treatment assignment with $\widebar m$ treated units. $D_j$ is a treatment indicator that equals one if unit $j$ is randomized into the treated group and zero otherwise. We study the performance of the following estimation strategies:

enumerate• SC: Constrained formulation of the synthetic control design. The results reproduce those of Table (ref). • RND: Randomized assignment of $\widebar m$ units to treatment followed by the difference in means estimator, \begin{align*} \frac{1}{\widebar m} \sum_{j=1}^J D_j Y_{jt} - \frac{1}{J-\widebar m} \sum_{j=1}^J (1-D_j)Y_{jt}. \end{align*} • STR: Divide the sample in $\widebar m$ strata, such that each stratum has at least two units. In each stratum, one unit is assigned to treatment at random. The composition of the strata is chosen to minimize the maximal within-strata discrepancy in the covariates, $\bm{Z}_j$, and pre-experimental outcomes (all normalized to have unit variance). Let $B_{jk}$ be a binary variable that equals one if and only if unit $j$ belongs to cluster $k$. Let $J_k$ be the number of units in stratum $k$. \begin{align*} \sum_{k=1}^{\widebar m} \frac{J_k}{J}\bigg(\sum_{j=1}^J B_{jk}D_jY_{jt} - \frac{1}{J_k-1}\sum_{j=1}^k B_{jk}(1-D_j)Y_{jt}\bigg), \end{align*} where $J_k$ represents the number of units within the $k$-th block. • REG: Randomized assignment of $\widebar m$ units to treatment followed by regression adjustment on the covariates, $\bm{Z}_j$. Ordinary least-squares adjustment on all pre-treatment outcomes is unfeasible as the number of pre-treatment outcomes exceeds the number of units in the sample. • 1-NN and \textit{5-NN}: Randomized assignment of $\widebar m$ units to treatment followed by $1$-nearest neighbor and $5$-nearest neighbor matching, respectively, on all pre-experimental outcomes and covariates. In both cases, predictors are rescaled to have unit variance.

Results are reported in Table 5. Across all values of $\widebar m$, the synthetic control design outperforms randomized assignment, including variants that incorporate pre-stratification, post-stratification, or regression adjustment. Taken together with the findings in Table (ref), these results underscore the potential of synthetic controls as a more effective design strategy in experiments involving aggregate units and a limited number of treated units.

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\center\normalfont\bf}{Conclusions}

Experimental design methods have largely been concerned with settings where a large number of experimental units are randomly assigned to a treatment arm, and a similarly large number of experimental units are assigned to a control arm. This focus on large samples and randomization has proven to be enormously useful in various classes of problems but becomes inadequate when treating more than a few units is unfeasible, as is often the case in experimental studies with large aggregate units (e.g., markets). In that case, randomized designs may produce estimators that are substantially biased (post-randomization) relative to the average treatment effect or to the average treatment effect on the treated. Large biases can be expected when the unit or units assigned to treatment fail to approximate average outcomes under treatment for the entire population or when the units in the control arm fail to approximate the outcomes that treated units would experience without treatment.

In this article, we have proposed synthetic control techniques, widely used in observational studies, to design experiments when the treatment can only be applied to a small number of experimental units. The synthetic control design optimizes jointly over the identities of the units assigned to the treatment and the control arms and over the weights that determine the relative contribution of those units to reproduce the counterfactuals of interest. We propose various designs to estimate average treatment effects, analyze the properties of such designs and the resulting estimators, and devise inferential methods to test a null hypothesis of no treatment effects and construct confidence intervals. In addition, we report results from an application to retail sales data and simulation results that demonstrate the applicability and computational feasibility of the methods proposed in this article. We show that synthetic control design can substantially outperform randomized designs in experimental settings with a small number of treated units.

Corporate researchers, policymakers, and academic investigators are often confronted with settings where interventions at the micro-unit level (e.g., customers, workers, or families) are unfeasible, impractical, or ineffective duflo2007using, jones2019uber. Consequently, there is broad scope for experimental design methods targeting large aggregate entities (such as regional markets, school districts, or states), a setting where synthetic control designs offer a powerful tool for data-driven evaluation of treatment effects.