EconBase
← Back to paper

Identification and Inference for Synthetic Controls with Confounding

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

69,395 characters · 16 sections · 60 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Inference for Synthetic Controls with Confounding

\if00 \fi

\if10 { \ \\

\ \\

} \fi

abstractThis paper studies inference on treatment effects in panel data settings with unobserved confounding. We model outcome variables through a factor model with random factors and loadings. Such factors and loadings may act as unobserved confounders: when the treatment is implemented depends on time-varying factors, and who receives the treatment depends on unit-level confounders. We study the identification of treatment effects and illustrate the presence of a trade-off between time and unit-level confounding. We provide asymptotic results for inference for several Synthetic Control estimators and show that different sources of randomness should be considered for inference, depending on the nature of confounding. We conclude with a comparison of Synthetic Control estimators with alternatives for factor models.

{\it Keywords:} Synthetic Control, Panel Data, Difference-in-Differences, Causal Inference. \\ {\it JEL Codes:} C33

\onehalfspacing

Introduction

This paper studies identification and inference in panel or longitudinal data settings where a (possibly small) number of units is exposed to treatment from one period $T_0$ onwards. We will refer broadly to these scenarios as Synthetic Control settings. The existing literature has proposed a variety of methods for estimating counterfactual outcomes: controlling for the lagged outcomes (known as horizontal regression), controlling for the control units outcomes abadie2003economic, abadie2010synthetic, or using information both across time and units as for the Synthetic DiD arkhangelsky2021synthetic and matrix completion methods and factor models athey2021matrix, gobillon2016regional, xu2017generalized. Most of this literature has focused on settings where the treatment assignment process is not random, and unobservable characteristics of treated and control units can be matched almost surely.\footnote{See for example, the discussion in ferman2021properties. } This approach allows the researcher to interpret the task of counterfactual estimation as a prediction problem, while inference must only consider variation in idiosyncratic shocks. This paper studies the properties of synthetic control-type estimators in the presence of unobservable (random) confounders. These confounders cannot be matched exactly. We derive identification conditions that formalize (i) which regression strategy is best suited in the presence of different confounding mechanisms and (ii) which sources of randomness should be considered for inference.

We study settings where potential outcomes follow an interactive factor model, with possibly high dimensional factors and loadings and additive (nonrandom) treatment effects. Unlike previous literature on synthetic control methods, here, both the factors and loadings are random variables and can act as confounders. The loadings act as unobserved confounders for whom receives the treatment, and the factors act as confounders for when the treatment is implemented. For instance, consider the problem of studying the effect of state-level regulation abadie2010synthetic. When the regulation is implemented may depend on aggregate factors of the economy, and which state implements the regulation may depend on state-specific (unobserved) characteristics. Because of confounding, conditional on the assignment mechanism, the distribution of the factors and loadings can be arbitrary.

We provide identification restrictions corresponding to the vertical, horizontal, and Synthetic DiD regressions within this confounding model. Vertical regression (i.e., controlling for a weighted combination of control units’ outcomes) allows for arbitrary time-level confounders, and it imposes restrictions on unit confounders. It assumes that conditional on who receives the treatment, unit-level confounders over the treated and control units match in expectations after appropriately reweighting (for weights summing to one). The leading case is the following: we can match all unit confounders' conditional expectations that are endogenous. The remaining loadings are exogenous with respect to the assignment mechanism. We refer to the latter identification restriction as “no high dimensional unit confounders" since these scenarios are attained only when a few loadings are endogenous. Vice-versa, horizontal regression (i.e., controlling for a weighted combination of lagged outcomes) allows for arbitrary endogenous unit-level confounders and assumes “no high dimensional time confounders", inverting the role of the factors and loadings. This finding illustrates the trade-off between the confounding restriction and the regression strategy.

We then turn to identification strategies robust to either high dimensional units or time confounders. We show that the Synthetic Difference-in-Differences in arkhangelsky2021synthetic is robust to either specification: its bias depends on the product of two differences; the difference of the time confounders' weighted expectations between treatment and control periods, and the difference of the unit level confounders. Therefore, its bias is zero if there are either low dimensional unit confounders or time confounders. This characterization of the bias formalizes double-robustness in Synthetic Control settings with confounding, and, while building on the intuition in arkhangelsky2021synthetic, is novel to the literature.

Taking these results as stock, we draw their implications for inference. Inference with time confounders for Synthetic Controls must consider the randomness generated across units by the (exogenous) high dimensional loadings but can condition on time variation induced by the (endogenous) factors. The reason is that the estimator is unbiased only unconditional on the loadings. For the horizontal regression, instead, we should account for the randomness generated by the factors but not the loadings. Finally, inference with Synthetic DiD must account for randomness over both unit and time dimensions because robust to confounding occurring either across units or time, motivating the construction of standard errors that capture both sources of randomness.

We derive asymptotic properties of the estimators and standard errors corresponding to confounding over time, units, or either of the two. We assume that the post-treatment period and the number of treated units grow with the sample but are small relative to the number of control periods and units. (We return to the case of one or few treated units in Appendix (ref).) We characterize the convergence rates of each estimator as a function of the treated units $N_1$ and treatment periods $T_1$. The convergence rate of Synthetic Control is between $1/\sqrt{N_1}$, and $1/\sqrt{N_1 T_1}$, with $1/\sqrt{N_1}$ if the factors are arbitrarily endogenous, and $1/\sqrt{N_1 T_1}$ if the factors are exogenous. Surprisingly, the rate of convergence of Synthetic DiD can be faster than the one of Synthetic Control with confounding. The key insight is that the Synthetic control's convergence rate is of order $1/\sqrt{N_1 T_1}$ only if both unit and time confounders concentrate around zero. However, for Synthetic DiD, a convergence rate of order $1/\sqrt{N_1 T_1}$ only requires that the time confounders before and after the treatment concentrate around the same expectation after reweighting.

We conclude with a discussion on Synthetic Control and factor models. Under distributional restrictions on the factors and loadings, for a Synthetic Control regression, convergence rates depend on the rank of a matrix of (low) dimensional unit confounders, as opposed to the rank of the (possibly high dimensional) matrix of factors and loadings. However, sufficiently fast convergence rates of the estimated weights also require some restrictions on the distribution of factors before the treatment time occurs since, otherwise we would not be able to estimate consistently such weights.

Our paper relates to an extensive literature on panel data models and Synthetic Controls. We build on the literature on synthetic controls abadie2003economic, abadie2010synthetic, doudchenko2016balancing, abadie2022synthetic, matrix completion methods and synthetic difference-in-differences abadie2010synthetic, athey2021matrix, arkhangelsky2021synthetic, liu2022practical, arkhangelsky2023large. We contribute to this literature by studying confounding through the unit and time-level confounders. Different from existing analysis with factor models athey2021matrix, arkhangelsky2021synthetic, here, both factors and loadings are random instead of deterministic and low rank, motivating different identification properties and inference. ben2021augmented provide useful finite sample upper bound with fixed loadings and factors using a balancing procedure. Here, we clarify conditions for different estimation strategies and variance estimators in the presence of random confounders.

We relate to studies on confounding with panel data, including ferman2016revisiting, hahn2017synthetic, kellogg2021combining, and agarwal2022synthetic for recent works on identification with dynamic treatments. shi2021theory, imbens2021controlling study instead proximal methods with panel data. Here, we allow confounding over either or both time and units, whereas these references implicitly condition on either (or both) the factors and loadings. Our focus on inference and its connection to confounding is a further distinction from these references. shen2022tale show that horizontal and vertical regression provide the same point estimates in the absence of an intercept and constraint on the weights. Here, we show that these estimators would be biased without an intercept and weights summing to one. Also, the authors consider random assignments while here we motivate the choice of regression strategies and confidence intervals based on the confounding mechanisms.

We complement the literature on inference with synthetic control and two-way fixed effects, including bottmer2021design,chernozhukov2019practical, cattaneo2021prediction, chernozhukov2021exact, viviano2023synthetic, imai2021use. Different from the references above, here we explicitly model the confounding mechanism in the construction of confidence intervals. shaikh2021randomization allow for random treatment timing but does not consider unobserved confounders. Finally, we more broadly relate to a larger strand of literature on factor models bai2009panel, moon2017dynamic, bai2019rank among others, which, however, does not directly tackle the problem of confounding for inference on causal effects. We defer an extensive discussion of theoretical comparisons of synthetic control weights and estimators for factor models to Section (ref).

Setup

We consider a panel with units and periods $$ i \in \{1, \cdots, N_0, \cdots, N_0 + N_1\}, \quad t \in \{1, \cdots, T_0, \cdots, T_0 + T_1\}, $$ respectively. Define $W_{i,t} \in \{0,1\}$ the treatment assignment for unit $i$ at time $t$, with $$ W_{i,t} = 1\{i \ge N_0\} 1\{t \ge T_0\}. $$ We let both $T_0$ and $N_0$ be random variables. Denote $\Big(Y_{i,t}(0), Y_{i,t}(1)\Big)$ the potential outcomes under treatment and control of unit $i$ at time $t$. Throughout our discussion, we assume constant and homogeneous treatment effects of the form

equation[equation omitted — 63 chars of source]

where $\tau$ denotes the average treatment effect. We return to heterogeneous treatment effects in Remark (ref). Researchers observe a matrix of potential outcomes under control depicted below, and potential outcomes $\mathbf{Y}_{i,t}(1)$ only for units $i \ge N_0, t \ge T_0$.

$$

aligned\mathbf{Y}(0) = \quad \quad \begin{tikzpicture}[baseline,decoration=brace] \matrix (m) [table] { Y_{1,1} & Y_{1,2} & \cdots & Y_{1, T_0 - 1} & & Y_{1, T_0} & \cdots & Y_{1, T_0 +T} \\ \vdots & \vdots & \ddots & \cdots & & \vdots & \vdots & \vdots \\ \vdots & \vdots & \ddots & \cdots & & Y_{N_0 - 1, T_0} & \vdots & Y_{N_0 - 1, T_0} \\ Y_{N_0,1} & Y_{N_0,2} & \cdots & Y_{N_0, T_0 - 1} & & ? & \cdots & ? \\ \vdots & \vdots & \ddots & \cdots & & \vdots & \vdots & \vdots \\ Y_{N_0 + N_1,1} & Y_{N_0 + N_1,2} & \cdots & Y_{N_0 + N_1, T_0 - 1} & & ? & \cdots & ? \\ }; \draw[decorate,transform canvas={xshift=-1.4em},thick] (m-3-1.south west) -- node[left=2pt] {$N_0$} (m-1-1.north west); \draw[decorate,transform canvas={yshift=0.5em},thick] (m-1-1.north west) -- node[above=2pt] {$T_0 - 1$} (m-1-4.north east); \end{tikzpicture}.

$$ Our goal is to estimate the average treatment effect $\tau$, which entails estimating the counterfactual outcomes. The literature has proposed three approaches for this problem: \textit{horizontal regression}, \textit{vertical regression} (i.e., Synthetic Control methods, \cite{abadie2010synthetic}), and \textit{matrix completion methods} \citep{athey2021matrix}. The horizontal regression uses \textit{past outcomes} to predict future outcomes of treated units. The vertical regression uses the control units' outcomes to predict the treated units' outcomes in periods $t > T_0$. The matrix completion method leverages the assumption that $\mathbf{Y}(0)$ can be approximated by a low-rank factor model bai2003inferential. In this paper, we study how different sources of confounding justify different regression strategies and standard errors.

Data generating process and confounding

We consider the following factor model.

ass[High-dimensional factor model] Assume that for all $(i,t)$, $$ \small \begin{aligned} Y_{i,t}(0) & = \Big(\lambda_t + \tilde{\Lambda}_t\Big)^\top \Big(\gamma_i + \tilde{\Gamma}_i\Big) + \iota_{0,i} + \iota_{1, t} + \varepsilon_{i,t}, \\ ||\mathbb{E}[\tilde{\Gamma}_i]||_{\infty} & = ||\mathbb{E}[\tilde{\Lambda}_t ]||_{\infty} = \mathbb{E}[\varepsilon_{i,t}] = 0, \end{aligned} $$ for random variables $\tilde{\Gamma}_i, \tilde{\Lambda}_t \in \mathbb{R}^{r}, \varepsilon_{i,t} \in \mathbb{R}$, and constants $\lambda_t, \gamma_i \in \mathbb{R}^r, \iota_{0,i}, \iota_{0,t} \in \mathbb{R}$, and $r$ possibly unknown to researchers. Let $\varepsilon_{i,t} | N_0, T_0 \sim_{i.i.d.} \mathcal{P}$ with $\mathbb{E}[\varepsilon_{i,t}^3] < \infty, \mathbb{E}[\varepsilon_{i, t}^2] = \sigma_\varepsilon^2 > 0$.

Assumption (ref) postulates that outcomes follow a factor model. The factor model depends on factors $\tilde{\Lambda}_t, \lambda_t$, and loadings $\tilde{\Gamma}_i, \gamma_i$. Here, $\varepsilon_{i,t}$ is an exogenous shock. The main difference between the model in Assumption (ref) with methods discussed in previous literature athey2021matrix, abadie2010synthetic, arkhangelsky2021synthetic, shen2022tale is that that we treat the factors and loadings as random variables, possibly acting as confounders.

Confounding over time and across units follows two independent processes.

ass[Independent processes] $\Big[\Big(\tilde{\Gamma}_i\Big)_{i\ge 1}, N_0\Big] \perp \Big[\Big(\tilde{\Lambda}_t\Big)_{t\ge 1}, T_0\Big]$, and $\Big[(\tilde{\Gamma}_i)_{i \ge 1}, (\tilde{\Lambda}_t)_{t \ge 1}\Big] \perp (\varepsilon_{i,t})_{i \ge 1, t \ge 1}$.

Assumption (ref) states that who receives the treatment depends on individual confounders and is independent of when the treatment is implemented. Assumption (ref) simplifies our theoretical analysis and allows us to study independently two sources of confounding over time and across units.

exmp[Effects of minimum wage on wages] For state-level regulations, such as studying the effect of minimum wage, Assumption (ref) states that which states may implement the policy depends on state-level characteristics, whereas when the policy is implemented depends on aggregate (time-varying) factors of the economy. \qed
assFor all $i, t$, $||\tilde{\Gamma}_i||_2, ||\gamma_i||_2, ||\tilde{\Lambda}_t||_2 , ||\lambda_t||_2$ are uniformly bounded almost surely. Let $\tilde{\Gamma}_i \sim_{i.i.d.} \mathcal{P}_\Gamma, \tilde{\Lambda}_t \sim_{i.i.d} \mathcal{P}_{\Lambda}$, with $\mathrm{Var}(\tilde{\Lambda}_t) = \Sigma_{\tilde{\Lambda}}, \mathrm{Var}(\tilde{\Gamma}_i) = \Sigma_{\tilde{\Gamma}}$ being positive definite.

Assumption (ref) states that (i) factors and loadings have bounded $l_2$-norm; (ii) factors and loadings are $i.i.d.$ unconditional on the treatment assignment mechanism. Assumption (ref) does not impose restrictions on how factors and loadings relate to the assignment mechanism. Therefore, it does not rule out dependence between factors, loadings, and the assignment mechanism. The bounded $l_2$-norm restriction allows for a growing number of possibly weak factors and allows us to control the variance of potential outcomes.\footnote{The assumption that factors are uniformly bounded can be relaxed by subgaussianity at the expense of additional notation.} Finally, note that time trends or seasonality components can be directly incorporated in a time fixed effect, assuming the separability of non-stationary components.

Research questions

Given Equation (ref), it follows that

equation[equation omitted — 623 chars of source]

Namely, researchers may estimate counterfactual outcomes conditional on different information sets to recover the same estimand $\tau$. This raises the following questions:

itemize• Which conditions on factors and loadings guarantee the identification of $\tau$, and how do these conditions relate to existing regression strategies? Since the factors and loadings are random, identification requires conditions on how they relate to treatment assignments. In Section (ref), we argue that different regression strategies are motivated by different identification strategies. • Which source of randomness should we consider for inference? Equation (ref) shows that we can take differences of expectations conditional on different information sets and recover the same estimand $\tau$. This creates confusion for inference on $\tau$: estimators of different conditional expectations will present different variances, not only because such estimators are different but also because we can condition on different information sets. In Section (ref), we relate confidence intervals (and sources of randomness) to the underlying identification assumptions. • Synthetic Control or factor models? Equation (ref) $(i)$ shows that it suffices to estimate the factors and loadings to recover $\tau$. However, standard estimators for factor models require a low-rank representation. In Section (ref), we provide a discussion and comparison with Synthetic Controls in settings with high dimensional (weak) factors.

Identification

This section studies identification conditional on $(N_0, T_0) = (N, T)$ for given $(N,T)$.

Vertical and horizontal identification

We first discuss identification for the horizontal regression.

ass[Limited confoundedness over time] Suppose that \begin{itemize} • For all $i, t$, for fixed $\gamma_i, \lambda_t$, $$ Y_{i,t}(0) \perp \Big[N_0, T_0\Big] \Big| \tilde{\Gamma}_i. $$ • There exist weights $w_{h, 1}, \cdots, w_{h, T}, v_{h,T+1}, \cdots, v_{h, T + T_1}$ such that for $T_0 = T$ $$ \Big|\Big|\sum_{s< T} w_{h, s} \lambda_s - \sum_{t \ge T} v_{h, t} \lambda_t\Big|\Big|_2 = 0, \quad \sum_{s < T} w_{h, s} = 1, \quad \sum_{t \ge T} v_{h,t} = 1. $$ \end{itemize}

Assumption (ref) states that we can match the endogenous factors ($\lambda_t$) exactly for some pre and post-treatment weights $w_h, v_h$. The remaining (random) factors $\tilde{\Lambda}_t$ are exogenous. Because we require exact matching for each entry of $\lambda_t$, we interpret the second restriction as imposing a sparsity restriction on $\lambda_t$.\footnote{We can relax the second condition in Assumption (ref) up to a small error of order $o(N_1^{-1/2} T_1^{-1/2})$.} As noted in Table (ref) and discussed more extensively in Section (ref), Assumption (ref) differs from usual low-rank restrictions in factor models, that instead typically assume that we can consistently estimate all interactive fixed effects. Instead, Assumption (ref) states that we can find a set of weights such that we can match the (low-dimensional) endogenous factors $\lambda_t$, whereas the remaining ones are exogenous.

prop[Horizontal identification] Let Assumptions (ref), (ref), (ref) hold. Then for all $i \in \{1, \cdots, N, \cdots, N + N_1\}$ \begin{equation} \begin{aligned} & \sum_{t \ge T} v_{h, t} \mathbb{E}\Big[Y_{i,t}(0) | N_0 = N, T_0 = T, \tilde{\Gamma}_i\Big] = \sum_{s < T} w_{h, s} \mathbb{E}\Big[Y_{i,s} \Big| N_0 = N, T_0 = T, \tilde{\Gamma}_i\Big] + \beta_{0, h}(w_h, v_h) \end{aligned} \end{equation} for some constant $\beta_{0, h}(w_h, v_h)$, which depends on $T$, $w_h, v_h$ and any weights $w_{h}, v_h$ satisfying Assumption (ref).
proof[Proof of Proposition (ref)] See Appendix (ref)

For known weights $w_h, v_h$, Proposition (ref) suggests regressing the outcomes of the control units onto a weighted average of the outcomes in the previous period as discussed in Appendix (ref). A similar result holds also for a Synthetic Control (vertical) regression, after inverting the role of the factor and loadings.

ass[Limited confoundedness over units] Suppose that \begin{itemize} • For all $i, t$, for fixed $\gamma_i, \lambda_t$, $$ Y_{i,t}(0) \perp \Big[N_0, T_0\Big] \Big| \tilde{\Lambda}_t. $$ • There exist weights $w_{v, 1}, \cdots, w_{v, N-1}, v_{v, N}, \cdots, v_{v, N+1}$ such that for $N_0 = N$, $$ \Big|\Big|\sum_{j < N} w_{v, j} \gamma_j - \sum_{n \ge N} v_{v, n} \gamma_n \Big|\Big|_2 = 0, \quad \sum_{j < N} w_{v, j} = 1, \quad \sum_{n \ge N} v_{v, n} = 1. $$ \end{itemize}
prop[Vertical identification] Let Assumptions (ref), (ref), (ref) hold. Then for all $t \in \{1, \cdots, T, \cdots, T + T_1\}$, $$ \small \begin{aligned} \sum_{n \ge N} v_{v, n} \mathbb{E}\Big[Y_{n, t}(0) | N_0 = N, T_0 = T, \tilde{\Lambda}_t\Big] = \sum_{j < N} w_{v, j} \mathbb{E}\Big[Y_{j,t} | N_0 = N, T_0 = T, \tilde{\Lambda}_t\Big] + \beta_{0,v}(w_v, v_v). \end{aligned} $$ for some constant $\beta_{0,v}(w_v, v_v)$ which only depends on $N$, $w_v, v_v$, and any weights $w_v, v_v$ satisfying Assumption (ref).

The proof mimics the one of Proposition (ref) once we invert the loadings and factors.

Proposition (ref) and (ref) provide identification restrictions for horizontal and vertical regression. The former assumes no high-dimensional time confounders, and the latter assumes no high-dimensional unit confounders.\footnote{Note that Proposition (ref), (ref) hold under (weaker) moment restrictions which only require mean exogeneity of $\tilde{\Lambda}_t, \tilde{\Gamma}_i$ instead of complete exogeneity of such random variables. In that case we can interpret, $\lambda_t = \mathbb{E}[\tilde{\Lambda}_t | T_0 = T], \gamma_i = \mathbb{E}[\tilde{\Gamma}_i | N_0 = N]$, as the conditional expectations of the factors and loadings, conditional on the treatment assignment.}

In the context of Example (ref), if we interpret $\tilde{\Gamma}_i$ as sectors in the economy, Assumption (ref) states that the decision to introduce the minimum wage may only depend on a few (aggregate) sectors but not on each of the sectors separately.

rem[Intercept and weights summing to one] The presence of an intercept in both vertical and horizontal identification strategies and the weights summing to one guarantee unbiasedness in the presence of time and unit fixed effects. Without such conditions, we would be unable to guarantee unbiasedness in the presence of fixed effects. This differs from shen2022tale, who show that vertical and horizontal regressions lead to the same point estimates assuming no intercepts and unconstrained weights. \qed

Double-robust identification

Next, we study settings where either unconfoundedness over time or units may occur. Consider the population equivalent of the Synthetic Differences-in-Difference (sDiD) in arkhangelsky2021synthetic, here augmented with post-treatment weights $v_{v}, v_h$: $$

aligned\bar{\tau}^{dr}(w_h, w_v, v_h,v_v) & = \sum_{j \ge N, t \ge T} v_{v, j} v_{h, t} \mathbb{E}\Big[\bar{\tau}_{j, t}^{dr}(w_h, w_v) \Big|N_0 = N, T_0 = T\Big]

$$ where $$

aligned\bar{\tau}_{N,T}^{dr}(w_h, w_v, v_h, v_v) & = Y_{N,T} - \Big\{\sum_{s < T} w_{h, s} \mathbb{E}\Big[Y_{N,s} \Big| N_0 = N, T_0 = T\Big] + \sum_{j < N} w_{v, j} \mathbb{E}\Big[Y_{j,T} \Big| N_0 = N, T_0 = T\Big] \\ &\quad \quad \quad \quad \quad \quad - \sum_{s<T} \sum_{j < N} w_{h,s} w_{v,j} \mathbb{E}\Big[Y_{j,s} \Big| N_0 = N, T_0 = T\Big]\Big\}.

$$

In the following proposition, we formalize double-robustness of $\bar{\tau}^{dr}(\cdot)$.

propSuppose that Assumptions (ref), (ref) hold. Then \begin{equation} \begin{aligned} \tau - & \bar{\tau}^{dr}(w_h, w_v, v_h, v_v) = \Big(\bar{\Gamma}_{pre}(w_v) - \bar{\Gamma}_{post}(v_v) \Big)^\top\Big(\bar{\Lambda}_{pre}(w_h) - \bar{\Lambda}_{post}(v_h)\Big), \end{aligned} \end{equation} where $$ \small \begin{aligned} \bar{\Gamma}_{pre}(w_v) = \sum_{i < N} w_{v, i} \Big(\mathbb{E}[\tilde{\Gamma}_i | N_0 = N] + \gamma_i\Big), \quad \bar{\Gamma}_{post}(v_v) = \sum_{i \ge N} v_{v, i} \Big(\mathbb{E}[\tilde{\Gamma}_i | N_0 = N] + \gamma_i\Big) \\ \bar{\Lambda}_{pre}(w_h) = \sum_{t < T} w_{h,t} \Big( \mathbb{E}[\tilde{\Lambda}_t | T_0 = T] + \lambda_t\Big), \quad \bar{\Lambda}_{post}(v_h) = \sum_{t \ge T} v_{h,t} \Big( \mathbb{E}[\tilde{\Lambda}_t | T_0 = T] + \lambda_t\Big). \end{aligned} $$
proofSee Appendix (ref).

Proposition (ref) shows the robustness properties of the Synthetic DiD method: the method is robust to imperfect match over time and units. Here, $\bar{\Gamma}_{pre}(w_v), \bar{\Gamma}_{post}(v_v)$ denote the pre and post treatment expectation of the loadings, and $\bar{\Lambda}_{pre}(w_h)$, $\bar{\Lambda}_{post}(v_h)$ of the factors (conditional on the event $N_0 = N, T_0 = T$). Under Assumption (ref), $\Big(\bar{\Lambda}_{pre}(w_h) - \bar{\Lambda}_{post}(v_h)\Big) = 0$, and under Assumption (ref), $\Big(\bar{\Gamma}_{pre}(w_v) - \bar{\Gamma}_{post}(v_v) \Big) = 0$.

Equation (ref) connects to previous results in the double robust literature farrell2015robust. However, here we intend double-robustness at the population level for $\bar{\tau}^{dr}$, as a function of the time and unit level mismatch, instead of a function of the estimators' convergence rates. It also differs from the analysis in arkhangelsky2021synthetic because we characterize double robustness as a function of the distributional properties of the factors and loadings.

Implications for estimators

We conclude this discussion with a short summary of our findings in this section. Assuming that we know the weights and intercepts $w_h, v_h, \beta_h, w_v, v_v, \beta_v$ for the moment, we can construct three estimators

equation[equation omitted — 529 chars of source]

corresponding to the horizontal regression for unit $n$, vertical for unit $t$, and double robust estimator. For the horizontal and vertical regression, we take

equation[equation omitted — 347 chars of source]

where $q_h, q_v$ are arbitrary weights that sum to one. A simple example is $q_{h, n} = 1/N_1, q_{v,t} = 1/T_1$. The additional weights $q_h, q_v$ and the post-treatment weights $v_h, v_v$ are motivated treatment effect homogeneity, see Remark (ref). In the following proposition, we illustrate when each of these estimators is unbiased.

prop[Conditioning sets] Under Assumptions (ref), (ref), for $n \ge N, t \ge T$, $$ \begin{aligned} (A) \quad & \tau - \mathbb{E}\left[\hat{\tau}_n^h(w_h, v_h, \beta_h) \Big| N_0 = N, T_0 = T, \Big(\tilde{\Gamma}_i\Big)_{i\ge 1}\right] = 0, \text{ if Assumption } \ref{ass:unc1a} \text{ holds}. \\ (B) \quad & \tau - \mathbb{E}\left[\hat{\tau}_t^v(w_v, v_v, \beta_v)\Big| N_0 = N, T_0 = T, \Big(\tilde{\Lambda}_t\Big)_{t\ge 1} \right] = 0 , \text{ if Assumption } \ref{ass:unc2a} \text{ holds}. \\ (C) \quad & \tau - \mathbb{E}\left[\hat{\tau}^{dr}(w_h, w_v, v_h, v_v)\Big| N_0 = N, T_0 = T \right] = 0, \text{ if either Assumption } \ref{ass:unc1a} \text{ or } \ref{ass:unc2a} \text{ holds}. \end{aligned} $$
proof[Proof of Proposition (ref)] The proof follows directly from Propositions (ref), (ref) and (ref).

Proposition (ref) has important implications for inference. It shows that we can condition on the loadings for a horizontal regression and factors for a vertical regression, but not vice-versa. Inference with the Synthetic DiD should take into account the randomness generated by both the factors and loadings. We formalize these intuitions in the following section.

rem[Heterogeneous treatment effects] Our analysis assumes that treatment effects are constant, i.e., $Y_{i,t}(1) = Y_{i,t}(0) + \tau$. Consider instead settings with heterogeneous (additive) treatment effects $\tau_{i,t} = Y_{i,t}(1) - Y_{i,t}(0)$. Effects heterogeneity may, for example, be relevant when treatment intensity varies by state due to different regulatory systems. Estimated treatment effects denote weighted averages, which depend on such heterogeneity. In the context of Synthetic Controls (vertical regression), researchers may be able to identify a weighted combination of treatment effects $\sum_{t \ge T} q_{v, t} \sum_{n \ge N} v_{v, n} \tau_{n, t}$ where the weights $v_{v, n}$ must satisfy the balancing restriction in Assumption (ref). Intuitively, if we can guarantee that Assumption (ref) holds for $v_{v, n} = 1/N_1$, we can recover the average effect on the treated unit when choosing $q_{v,t} = 1/T$. In contrast, if Assumption (ref) holds for some other weighting mechanism, we can only recover a weighted combination of treatment effects. This intuition formalizes the trade-offs of matching restrictions on the weights $(v_h, v_v)$ and the identified class of estimands. On the other hand, because $q_h, q_v$ can be chosen arbitrarily by the researcher, we recommend as choices for $q_h, q_v$, $q_h = 1/N_1 \mathbf{1}, q_v = 1/T_1 \mathbf{1}$, i.e., imposing equal weights on the estimators.\footnote{This choice is recommended under the failure of homogeneous treatment effects when the goal is to study the average effect on the treated unit. A second possible choice is to minimize the variance of the estimator, which, however, would require estimating the factors and loadings (and would affect the estimand of interest in the absence of homogeneous treatment effects).} \qed
table[table omitted — 1,817 chars of source]

Inference over the post-treatment period

This section studies inference for vertical, horizontal regression and Synthetic DiD. We consider the estimators in Equation (ref) for the horizontal and vertical regression and $\hat{\tau}^{dr}(\cdot)$ in Equation (ref) for the Synthetic DiD. We will condition on the event $ N_0 = N, T_0 = T$. We study asymptotics as $N, T, N_1, T_1 \rightarrow \infty, $ with $N_1, T_1$ growing at an appropriate slower rate than $N, T$. We return to settings with a finite number of treated units in Appendix (ref).

We study inference given some known (not data-dependent) weights $(w_h^\star, v_h^\star,\beta_h^\star), (w_v^\star, v_v^\star,\beta_v^\star)$. We will assume that these weights satisfy the restrictions in Assumptions (ref) and (ref), respectively, and regularity conditions presented below. Corollary (ref) shows that estimation in the absence of estimation error is asymptotically equivalent to inference with estimated weights -- assuming that the post-treatment period is sufficiently shorter than the pre-treatment period. We study the estimation of the weights in Section (ref).

As shown in Proposition (ref), inference must account for different sources of randomness depending on the nature of confounding: estimators are unbiased only unconditional (but not necessarily conditional) on either time or unit-level confounders, depending on the confounding assumption. We formalize this intuition below.

Inference with horizontal and vertical regression

We present assumptions and results for the horizontal regression first.

ass[Horizontal regression conditions] Assume the following: \begin{itemize} • $T_1 N_1/T_0^{1/3}= o(1)$; • $||w_h^\star||_{\infty} = \mathcal{O}(T_0^{-2/3}), ||v_h^\star||_{\infty} = \mathcal{O}(T_1^{-2/3})$, and $w^\star, v^\star$ satisfy condition (B) in Assumption (ref). \end{itemize}

Condition (A) states that the post-treatment period is much shorter than the pre-treatment period. Condition (B) assumes that the oracle weights guarantee pre and post treatment balance in the factors. Condition (B) assumes that no single (or few) units receive all the weights similar to arkhangelsky2018role, ruling out sparse settings.

thmSuppose that Assumptions (ref), (ref), (ref), (ref), (ref) hold. Then as $T_1 \rightarrow \infty$, for any $q_h: \sum_{n \ge N} q_{h,n} = 1$, $$ \small \begin{aligned} \frac{ \Big(\hat{\tau}^h(w_h^\star, v_h^\star, \beta_h^\star, q_h) - \tau\Big)}{\bar{\mathbb{V}}_h^{1/2}(N, T, q_h)} \rightarrow_d \mathcal{N}(0,1) \end{aligned} $$ where $\bar{\mathbb{V}}_h(N, T, q_h) = \mathbb{V}\left( \sum_{n \ge N} q_{h, n} \sum_{t \ge T} v_{h,t}^\star Y_{n,t} \Big| N_0 = N, T_0 = T, \Big(\tilde{\Gamma}_i\Big)_{i\ge 1}\right)$.
proofSee Appendix (ref).

Theorem (ref) shows that the horizontal estimator, for any weighted combination of treated units converges in distribution to a Gaussian. The convergence rates also depend on unit-level confounders. The variance can be estimated directly using the empirical moments of the post-treatment outcomes.

cor[Rate of convergence] Let the conditions in Theorem (ref) hold. Then $$ \small \begin{aligned} \frac{\bar{\mathbb{V}}_h(N, T, q_h)}{||v_h^\star||_2^2} = \Big(\sum_{n \ge N} (\tilde{\Gamma}_n + \gamma_n) q_{h, n}\Big)^\top \Sigma_{\tilde{\Lambda}} \Big(\sum_{n \ge N} (\tilde{\Gamma}_n + \gamma_n) q_{h, n}\Big) + \sigma_{\varepsilon}^2 ||q_h||_2^2, \end{aligned} $$ where $\Sigma_{\tilde{\Lambda}} = \mathbb{V}(\tilde{\Lambda}_t), \sigma_\varepsilon^2 = \mathbb{V}(\varepsilon_{i,t})$.

Corollary (ref) shows that the variance (and rate of convergence) depends on the norm of the weights and on the loadings. The rate of convergence is weakly slower than $\frac{1}{\sqrt{T_1 N_1}}$, and equal to $\frac{1}{\sqrt{T_1 N_1}}$ if the loadings are exogenous. Specifically, we can write $$ \bar{\mathbb{V}}_h(N,T, q_H) = \mathcal{O}\Big(||q_h||^2 || v_h^\star||_2^2 \sigma_\varepsilon^2 + \rho_{N_1}(q_h) ||v_h^\star||_2^2 \Big), \quad \rho_{N_1}(q_h) = \Big| \Big|\sum_{n \ge N} q_{h,n} (\tilde{\Gamma}_n + \gamma_n) \Big| \Big|_2^2. $$ Here, $||v_h^\star||_2^1 = 1/T$, and $\rho_{N_1}$ characterizes the strength of the unit-level confounding, with $\rho_{N_1}$ either bounded away from zero, or $\rho_{N_1} \rightarrow 0$. For $\rho_{N_1} \rightarrow 0$ to converge to zero, we would need no unit-level confounding, that is $\frac{1}{N} \sum_{n \ge N} \tilde{\Gamma}_n$ concentrates around its unconditional expectation (unconditional on $N_0$) and $\gamma_n$ is local to zero. This result illustrates properties of convergence rates that depend on the strength of confounding.

The Synthetic Control case follows similarly to the horizontal regression case, where we invert the role of the loadings and factors, with $$

aligned\frac{\Big(\hat{\tau}^v(w_v^\star, v_v^\star, \beta_v^\star, q_v) - \tau\Big)}{\bar{\mathbb{V}}_v^{1/2}(N, T, q_v)} \rightarrow_d \mathcal{N}(0,1)

$$ where $\bar{\mathbb{V}}_v(N, T, q_v) = \mathbb{V}\left( \sum_{t \ge T} q_{v, t} \sum_{n \ge N} v_{v, n}^\star Y_{n,t} \Big| N_0 = N, T_0 = T, \Big(\tilde{\Lambda}_t\Big)_{t\ge 1}\right)$. Here, $$

aligned\frac{ \bar{\mathbb{V}}_v(N, T, q_v)}{||v_v^\star||_2^2} = \Big(\sum_{t \ge T} (\tilde{\Lambda}_t + \lambda_t) q_{v, t}\Big)^\top \Sigma_\Gamma \Big(\sum_{t \ge T} (\tilde{\Lambda}_t + \lambda_t) q_{v, t}\Big) + \sigma_\varepsilon^2 ||q_v||_2^2,

$$ where $\Sigma_{\tilde{\Gamma}} = \mathbb{V}(\tilde{\Gamma}_i)$. With diagonal $\Sigma_{\tilde{\Gamma}}$, the rate of convergence is of order $||v_v^\star||_2^2 ||q_v||_2^2 \sigma_\varepsilon^2 + ||v_v^\star||_2^2 \Big| \Big| \sum_{t \ge T} q_{v,t} (\tilde{\Lambda}_t + \lambda_t)\Big| \Big|_2^2$, which depends on time-level confounders. Therefore, whereas each regression provides consistent estimates as $N, T \rightarrow \infty$ under their corresponding unconfoundedness restriction (either no high dimensional time or unit level confounders), a faster rate of convergence can be obtained if both no high dimensional time and unit level confounders hold.

We conclude this discussion by showing how our results in Theorem (ref) apply in the presence of estimated weights.

corSuppose that the conditions in Theorem (ref) hold, and consider any data dependent weights $(\hat{w}_h, \hat{v}_h, \hat{\beta}_h)$ such that \begin{equation} \hat{\tau}_{n,t}^h(\hat{w}_h, \hat{v}_h, \hat{\beta}_h) - \hat{\tau}_{n,t}^h(w_h^\star, v_h^\star, \beta_h^\star) = o_p\left(N_1^{-1/2} T_1^{-1/2}\right). \end{equation} Then Theorem (ref) holds with $\hat{\tau}^h(\hat{w}_h, \hat{v}_h, \hat{\beta}_h)$ in lieu of $\hat{\tau}^h(w_h^\star, v_h^\star, \beta_h^\star)$.
proof[Proof of Corollary (ref)] See Appendix (ref).

An important insight of Corollary (ref) is that even if the variance is conditional on the loadings for the horizontal regression, the estimated weights only need to converge to their population counterpart unconditionally on the factors and loadings. The reason is that the variance $\bar{\mathbb{V}}_h$ converges to zero at a rate at most $1/\sqrt{N_1 T_1}$, while the estimated weights can converge at a faster rate when the pre-treatment period and number of control units is larger than the post-treatment period and number of treated units. Appendix (ref) presents the details. Conditions on the convergence rates of the weights are discussed in Section (ref), where we review different weights estimators.

rem[Inference with a single treated unit] In Appendix (ref), we study inference in the presence of a single treated unit. We show how placebo tests in abadie2010synthetic can be used in our setting under additional unconfoundedness restrictions on the endogenous loadings, assuming that we can find a donor pool of “placebo treated units" such that we can exactly match their loadings with the loadings of the remaining control units. \qed

Robust inference

In the following lines, we study inference that is robust to confoundedness of either the loadings or factors. We impose the following conditions.

ass[Robust regression conditions] Assume the following: \begin{itemize} • $T_1 N_1/T_0^{1/3} = o(1), N_1 T_1/N_0^{1/3} = o(1)$; • $||w_h^\star||_{\infty} = \mathcal{O}(T_0^{-2/3}), ||w_v^\star||_{\infty} = \mathcal{O}(N_0^{-2/3})$, $||v_h^\star||_{\infty} = \mathcal{O}(T_1^{-2/3}), ||v_v^\star||_{\infty} = \mathcal{O}(N_1^{-2/3})$, and either $(w_h^\star, v_h^\star)$ satisfy (B) in Assumption (ref) or $(w_v^\star, v_v^\star)$ satisfy (B) in Assumption (ref). \end{itemize}

Assumption (ref) formalizes double robustness properties, for which either no high dimensional unit confounders or no high dimensional time confounders exist.

thmLet Assumptions (ref), (ref), (ref), (ref) hold, and either Assumption (ref) or Assumption (ref) (or both) hold. Then as $N_1, T_1 \rightarrow \infty$, $$ \small \begin{aligned} P\left(\frac{\Big(\hat{\tau}^{dr}(w_h^\star, w_v^\star, v_h^\star, v_v^\star) - \tau\Big)}{\max\{\mathbb{V}_h(v_h^\star, v_v^\star), \mathbb{V}_v(v_v^\star, v_h^\star)\}^{1/2}} \le z_{1 -\alpha} \Big| N_0 = N, T_0 = T\right) \ge 1 - \alpha, \end{aligned} $$ where $\Phi(z_{1-\alpha}) = 1- \alpha$, $\Phi$ is the Gaussian CDF, and $$ \small \begin{aligned} \mathbb{V}_h(v_h^\star, v_v^\star) & = \mathbb{V}\left(\sum_{t \ge T} \sum_{j \ge N} v_{h,j}^\star v_{v, t}^\star \Big\{ Y_{j,t} - \sum_{i < N} w_{v, i}^\star Y_{i, t}\Big\}\Big| N_0 = N, T_0 = T, (\tilde{\Gamma}_i)_{i \ge 1}\right), \\ \mathbb{V}_v(v_v^\star, v_h^\star) & = \mathbb{V}\left( \sum_{j \ge N} \sum_{t \ge T} v_{v, t}^\star v_{h,t}^\star \Big\{ Y_{j,t} - \sum_{s < T} w_{v,s}^\star Y_{j, s}\Big\} \Big| N_0 = N, T_0 = T, (\tilde{\Lambda}_t)_{t \ge 1}\right). \end{aligned} $$
proof[Proof of Theorem (ref)] See Appendix (ref).

Theorem (ref) shows that confidence intervals for Synthetic DiD depend on the worst-case variance conditional on either the factors or loadings. The variance calculation differs from settings with non-random factors and loadings arkhangelsky2021synthetic, and can be computed using the empirical moments of weighted combinations of the outcomes.\footnote{Because the $\max\{\}$ operator is a continuous function, we can use the estimated variances $\mathbb{V}_h, \mathbb{V}_u$ and invoke Slutsky theorem to provide asymptotically valid confidence intervals. } This difference is because of the population double-robustness property we derived.

A direct corollary of Theorem (ref) is that the estimator $\hat{\tau}^{dr}(\cdot)$'s convergence rate depends on both $N_1, T_1$, and on the unconfoundedness restriction. The rate of converge is $\max\{\mathbb{V}_h(v_h^\star, v_v^\star), \mathbb{V}_v(v_v^\star, v_h^\star)\}$, namely (let $\wedge$ denote the maximum operator, $\Gamma_i = \gamma_i + \tilde{\Gamma}_i, \Lambda_t = \lambda_t + \tilde{\Lambda}_t$)

equation[equation omitted — 612 chars of source]

If both $\Lambda_t \perp T_0, \Gamma_i \perp T_0$, the right-hand side of Equation (ref) converges almost surely to zero, with rate $1/\sqrt{N_1 T_1}$. On the other hand, under lack of either unconfoundedness restriction (over time or units), the right-hand-side expression in Equation (ref) does not converge to zero and the rate of convergence is $\max\{T_1^{-1/2}, N_1^{-1/2}\}$. Convergence rates faster than $1/\sqrt{T_1}$ or $1/\sqrt{N_1}$ do not require that factors or loadings are exogenous. Convergence rates of order $1/\sqrt{N_1 T_1}$ allow the conditional expectations of the loadings and factors to be different from zero but require that the expectations match before and after the treatment after reweighting.

table[table omitted — 1,418 chars of source]

Estimation of the weights and factors: a discussion

In this section, we review existing methods for estimating the weights and discuss properties and assumptions that these methods require in the context of our model of confounding.

Weights estimation with penalized $l_2$-norm minimization

First, we review estimation of the weights as in arkhangelsky2021synthetic, which are similar to those in abadie2010synthetic with the additional norm constraint. For the horizontal regression, we can construct weights' estimators as follows:

equation[equation omitted — 456 chars of source]

where we minimize the $l_2$-distance between the post-treatment average control outcomes and a weighted combination of the pre-treatment outcomes. Similarly, for the vertical regression

equation[equation omitted — 462 chars of source]

Here, $p_v, p_h$ denote choosen penalty parameters discussed in arkhangelsky2021synthetic.

The constraints are as discussed in Sections (ref), (ref). Algorithm (ref) presents a summary.

algorithm[algorithm omitted — 781 chars of source]

The minimization problem in Section (ref) corresponds to a error-in-variables optimizaton problem. By letting $\Gamma_i = \gamma_i + \tilde{\Gamma}_i$

equation[equation omitted — 535 chars of source]

where

equation[equation omitted — 127 chars of source]

Following arkhangelsky2021synthetic, we define “oracle" counterpart of such a problem for the horizontal regression solves the following optimization problem.

equation[equation omitted — 593 chars of source]

where $p_h^\star$ is a penalty that depends on the residuals' variance as in arkhangelsky2021synthetic. The penalization of the oracle solution is different from the penalization used for estimation as motivated in hirshberg2021least. We can define oracle weights for Synthetic Control $(w_v^\star, v_v^\star, \beta_v^\star)$ in a similar manner.

Consider estimating the horizontal regression weights, and let $||\cdot||_{op}$ be the operator norm for a matrix,\footnote{For a matrix $A$ the operator norm is defined as $\sup_{u, ||u|| = 1} ||A u||$.} and $$

aligned\Sigma_\eta & = \frac{1}{T_0 + T_1} \mathbb{E}\Big[\tilde{\eta} \tilde{\eta}^\top \Big| (\tilde{\Gamma}_i)_{i \ge 1}, N_0 = N, T_0 = T\Big], \quad \mu^2 = ||\Sigma_\eta||_{op}.

$$ Under Assumptions (ref), (ref), (ref), (ref)

equation[equation omitted — 418 chars of source]

It follows that conditional the treated units $N_0 = N$, $\mu^2$ measures the amount of endogeneity of $\tilde{\Gamma}_i$. Specifically, let $$

alignedA & \in \mathbb{R}^{ (N_0 - 1) \times (T_0 + T_1) }, \quad A_{i,t} = \begin{cases} \lambda_t^\top \Gamma_i + \iota_{0, t} & if t < T_0 \\ - \lambda_t^\top \Gamma_i - \iota_{0,t} & otherwise. \end{cases}, \\ \tilde{\eta} & \in \mathbb{R}^{(N_0 - 1) \times (T_0 + T_1)}, \quad \tilde{\eta}_{i,t} = \begin{cases} \eta_{i,t}, & if t < T_0 \\ - \eta_{i,t} , & otherwise. \end{cases}

$$ Here the matrix $A$ depends on the endogenous factors; $\tilde{\eta}$ is the noise matrix for the \textit{control} units. Define $$ \theta_h = (w_h, v_h) and \theta_h^\star = (w_h^\star, v_h^\star). $$

In Appendix (ref), using rate-properties of error in variable models in hirshberg2021least applied to our model of confounding, we show that the convergence rate depends on three main components: $$ \frac{\mathrm{rank}(A) \mu^2 ||\theta^\star||_2^2}{N_0}, \quad \mu^2 ||\theta^\star||_2 \log(T_0 + T_1) N_0^{-1/2}, \quad \frac{\mu^2 \log(T_0 + T_1)}{N_0}. $$ Importantly, here, the rank of the matrix $A$ only depends on the non-zero number of endogenous factors $\lambda_t$, i.e., $||\lambda_t||_0$. In the presence of few endogenous factors, the rank of $A$ is small, even in settings where $\tilde{\Lambda}_t$ is high-dimensional. The component $\mu^2$ depends on the degree of endogeneity of the loadings (unit-level confounders). The rate of convergence of $\mu^2$ is of order $\sqrt{N_0}$ from standard properties of matrix concentration inequalities van2017spectral for exogenous $\tilde{\Gamma}_i$ and slower otherwise. The component $||\theta_h^\star||$ instead converges to zero as $T_1 \rightarrow \infty$. The error does not necessarily converge to zero for arbitrarly endogenous loadings $\tilde{\Gamma}_i$. It converges to zero under restrictions on $\mu^2$ (e.g., only few loadings $\tilde{\Gamma}_i$ are endogenous and the remaining ones are exogenous). The same results holds for vertical weights once we exchange the role of the factors and loadings.

This result illustrates the benefits of the Synthetic control methods in the presence of high-rank factor models.

rem[Balancing] We can gain further intuition if we interpret the weights' estimator in Section (ref) as the dual of a balancing problem. For given constraint $\nu = o(1)$ (as a function of $N, T$) we can formulate the dual of Equation (ref) as $$ \small \begin{aligned} \min_{w_h, v_h, \beta} ||w_h||^2 + ||v_h||_2^2, \quad \text{ s.t. } & \underbrace{\frac{1}{N} \sum_{i < N} \Big(\sum_{t \ge T_0} v_{h,t} Y_{i,t} - \sum_{s > T} w_{h,s} Y_{i,s} + \beta\Big)^2}_{:=\frac{1}{N} ||(A + \tilde{\eta}) \theta_h + \beta||_2^2} \le \nu, \quad (w_h, v_h) \ge 0, \\ & ||w_h||_{\infty} \le T_0^{-2/3}, ||v_h||_{\infty} \le T_1^{-2/3}, 1^\top w_h = 1^\top v_h = 1. \end{aligned} $$ The dual formulation clarifies the role of the constraint on the weights $w_h$: The constraint on the weights norm guarantees the variance of the error component $\tilde{\eta} \theta_h$ converges to zero (its variance is of order $||\theta_h||^2$). Intuitively, the Synthetic Control method averages over the component $\tilde{\eta}$ to obtain consistent weight's estimators without estimating the factors directly.\footnote{It is possible to consider alternative balancing estimators. For example, we might balance the means of treated and control units by imposing $$ \Big|\frac{1}{N} \sum_{i < N} \sum_{t \ge T} v_{h,t} Y_{i,t} - \frac{1}{N} \sum_{i < N} \sum_{s > T} w_{h,s} Y_{i,s} \Big | \le \nu', $$ for some constraint $\nu' = o(1)$. Imposing such balancing restriction assumes that we can find a set of weights $v_h^\star, v_h^\star$ such that $\sum_{t < T} v_{h,t}^\star \lambda_t = \sum_{t \ge T} w_{h,t}^\star \lambda_t$ and $\sum_{t < T} v_{h,t}^\star \iota_{1,t} = \sum_{t \ge T} w_{h,t}^\star \iota_{1, t}$, i.e., we can also match fixed effects before and after the treatment.} \qed

Proximal methods for weights estimation

It is possible to use proximal methods to estimate the weights in the spirit of shi2021theory, imbens2021controlling. Consider estimating synthetic control weights first. Suppose we can find variables $Z_t$ independent of $\varepsilon_{i,t}$ such that for all $w_{v}, v_v$

equation[equation omitted — 251 chars of source]

Assuming that the number of such variables $Z_t$ is sufficiently larger than the number of parameters (weights) to be estimated, we can use such moment restrictions to estimate the weights. The main advantage of the proximal method is that it does not impose restrictions on the distribution of the endogenous factors $\tilde{\Lambda}_t$. However, it requires finding a set of proximal variables that satisfy Equation (ref). A similar approach follows for horizontal weights, where we should find a (recenter) instrument $X_i$ such that $$ \mathbb{E}\Big[X_i \Big(\sum_{t < T} w_{h,t} \tilde{\Lambda}_t^\top - \sum_{s \ge T} v_{h, s} \tilde{\Lambda}_s\Big)^\top \tilde{\Gamma}_i \Big| N_0 = N, T_0 = T, \tilde{\Gamma}_i\Big] = 0, \quad \forall i \le N_0. $$

Factor models estimators

We conclude this section with a discussion on factor models. We write our model as (for $\Gamma_i = \tilde{\Gamma}_i + \gamma_i$)

equation[equation omitted — 163 chars of source]

where $\lambda_t^\top \Gamma_i$ is low rank when $||\lambda_t||_0$ is finite, and $\tilde{\Lambda}_t^\top \Gamma_i$ can be an high rank component.

Our discussion above suggests that horizontal regression (and similarly vertical regression after exchanging the factors with the loadings) leverage the assumption that endogenous factors $\lambda_t$ are low dimensional (e.g., $||\lambda_t||_0$ is uniformly bounded).

It is interesting to contrast the convergence rate of the weights of Synthetic controls (described in Section (ref) and formally presented in Proposition (ref)) with those that we would obtain using standard methods for estimating factor models. Two main differences from properties of least squares estimators as in bai2009panel is that Synthetic Control does not require a (i) low-rank assumption on the factors, but only a low-rank assumption on the endogenous factors; (ii) it does not require to specify or estimate the number of (endogenous) factors. For example, if we consider a factor model where the factors are aggregate sectoral shocks and the loadings are sectoral weights, a sparsity assumption on loadings implies that the policy implementation may only depend on a subset of sectoral weights.

More closely related to the results in Proposition (ref), moon2017dynamic show that conditions for least squares factor model can be expressed as a function of the operator norm ($||\eta||_{op}$). The main distinction from Proposition (ref), however, is that the number of factors (the degree of sparsity of $\lambda_t$ in our framework) must be known for moon2017dynamic's results to hold. With an unknown number of factors (as in the case of Proposition (ref)), stronger conditions on the error term $\eta$, such as rank restrictions, are imposed moon2015linear.

The model in Equation (ref) also connects to the approximate factor model discussed in chamberlain1982arbitrage. The main distinction, however, is that low dimensional factor structure $\lambda_t^\top \gamma_i$ cannot be necessarily separated by the (exogenous and high dimensional) structure $\tilde{\Gamma}_i^\top \tilde{\Lambda}_t$, different from chamberlain1982arbitrage. This is an important distinction also from work in the literature of causal inference with panel data as in athey2021matrix (see, e.g., Section 8.2).

Interestingly, Equation (ref) may justify alternative estimators if additional conditions are imposed on the loadings and factors. For example, under low rank $\lambda_t^\top \Gamma_i$ and sparse $\tilde{\Lambda}_t^\top \Gamma_i$ we could use estimators in candes2011robust, and for bounded eigenvalues of $\tilde{\Lambda}_t^\top \Gamma_i$ as in bai2019rank.

Finally, Table (ref) presents numerical studies that support the benefit of Synthetic control methods over least square regression.

table[table omitted — 1,605 chars of source]

Conclusion

This paper studies inference on treatment effects in panel data settings in the presence of confounding. We model confounding through unobserved factors -- which might affect when the treatment occurs -- and unobserved loadings -- which might affect which units receive the treatment. We illustrate the existence of a trade-off between assuming no (high dimensional) confounding across units or time and introducing notions of double robustness in this setting. We relate notions of confounding to the source of randomness for confidence intervals.

This paper opens new questions on trade-offs between the choice of the estimator and robustness to confounding. Different sources of confounding justify different estimators, including Synthetic DiD, Synthetic Control, factor models, or proxy variable methods. A comprehensive comparison of such estimators remains an open question. Future research should also study trade-offs between weak factors and the choice of the estimator. Synthetic control methods implicitly leverage the low-rank representation of the confounders, whereas the least-squares method estimates all such confounders. Finally, a further avenue for future avenue of research is to study augmented inverse probability weight estimators in the spirit of robins1994estimation in contexts with synthetic controls and high dimensional factor models.