The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
82,593 characters
A Relaxation Approach to Synthetic Control
\title{A Relaxation Approach to Synthetic Control\thanks{Liao: \protect\url{[email removed]}. Shi (corresponding author):
\protect\url{[email removed]}; 9F Esther Lee Building, The
Chinese University of Hong Kong, Shatin, New Territories, Hong Kong
SAR, China. Zheng: \protect\url{[email removed]}. Shi
acknowledges the partial financial support from the Research Grants
Council of Hong Kong (Project No.14617423) and the National Natural
Science Foundation of China (Project No.72425007, 72133002).}}
\author{Chengwang Liao\quad{}Zhentao Shi\quad{}Yapeng Zheng \\
\emph{The Chinese University of Hong Kong} \bigskip{}
}
\maketitle
\begin{abstract}
The synthetic control method (SCM) is widely used for constructing
the counterfactual of a treated unit based on data from control units
in a donor pool. Allowing the donor pool contains more control units
than time periods, we propose a novel machine learning algorithm,
named SCM-relaxation, for counterfactual prediction. Our relaxation
approach minimizes an information-theoretic measure of the weights
subject to a set of relaxed linear inequality constraints in addition
to the simplex constraint. When the donor pool exhibits a group structure,
SCM-relaxation approximates the equal weights within each group to
diversify the prediction risk. Asymptotically, the proposed estimator
achieves oracle performance in terms of out-of-sample prediction accuracy.
We demonstrate our method by Monte Carlo simulations and by an empirical
application that assesses the economic impact of Brexit on the United
Kingdom’s real GDP.
{\large\bigskip{}
}{\large\par}
\noindent\textbf{Key words}: counterfactual prediction, high dimension,
latent groups, machine learning, optimization
\noindent\textbf{JEL classification}\textsc{: C22, C31, C53, C55}
\end{abstract}
\newpage{}
\section{Introduction}
A fundamental challenge in policy evaluation is that the potential
outcomes in the absence of treatment are unobservable in the treated
group. When resources are available to conduct randomized controlled
trials, carefully designed experiments can identify the average treatment
effect (ATE) by comparing outcomes between the treatment group and
the control group. However, when only observational data are available,
one prominent solution is the synthetic control method (SCM) \citep{abadie2003economic,abadieSyntheticControlMethods2010}.
A typical application involves a panel data with a single treated
unit of interest and a set of control units, where they are connected
by sharing latent common factors. SCM has been successfully applied
to many economic studies, for example, \citet{kleven2013taxation},
\citet{bohn2014did}, \citet{acemoglu2016value}, and \citet{cunningham2018decriminalizing}.
\citet{athey2017state} hailed SCM as ``arguably the most important
innovation in the policy evaluation literature in the last 15 years.''
A pivotal step in SCM is assigning weights to the units in the donor
pool to construct a “synthetic” version of the treated unit. These
weights are determined by minimizing the discrepancy between the treated
unit and the synthetic control in the pre-treatment period. The resulting
synthetic unit is then used to extrapolate outcomes in the post-treatment
period, with the estimated policy effect given by the difference between
the observed outcome of the treated unit and the counterfactual prediction.
In a standard SCM implementation, the weights are constrained into
the simplex---that is, they must be non-negative and sum to one.
This constraint often leads to sparse solutions, where only a few
weights are non-zero. While such sparsity enhances interpretability,
it may not always be desirable. In many practical economic settings,
there is no compelling reason to assume that only a small subset of
variables are relevant to the target variable \citep{giannone2021economic}.
When researchers invest substantial effort in collecting data from
a large pool of potential controls, they may prefer dense weighting
schemes that make a full use of the available information. A dense
solution can also help diversify prediction risk, potentially improving
the robustness of the counterfactual estimates \citep{liao2024does}.
To encourage dense solutions, in this paper we propose a relaxation
approach within the synthetic control framework, termed \emph{SCM-relaxation},
for estimating the weights and predicting the counterfactual outcomes.
Our approach is built on the $L_{2}$-relaxation method introduced
by \citet{shi$ell_2$relaxationApplicationsForecast2022}. It formulates
the weight estimation as a constrained optimization problem, where
the objective function is an information-theoretic divergence measure
of the weight vector. Staying true to the spirit of SCM, we retain
the simplex constraint. In addition, to address challenges arising
from many potential control units, we introduce a relaxed form of
the first-order condition (FOC) derived from the original SCM formulation.
In high-dimensional settings, developing asymptotic theory requires
structural assumptions that enable effective dimension reduction.
The search for a convex combination in SCM parallels the classical
econometric problem of forecast combination, making it natural to
leverage the aggregated information from multiple control units. Recent
advancements in the literature on latent groups in panel data---\citet{bonhomme2015grouped},
\citet{su2016identifying}, and \citet{vogt2017classification}---provide
a useful framework to bridge the two extremes of full homogeneity
(identical units) and full heterogeneity (distinct units). Under the
latent group structures in the loadings of a factor model, this paper's
\ref{thm:w_convergence} shows that our estimated weight vector converges
sufficiently fast to a desirable target, even when the number of control
units ($J$) is much larger than the number of pre-treatment time
periods ($T_{0}$). Once the weights are estimated, the ATE can be
readily obtained through the predicted counterfactual. \ref{thm:oracle_inequalities}
establishes that the prediction risk of the counterfactuals is asymptotically
equal to that under the oracle weights.
Our use of divergence measures is motivated by the literature on \emph{generalized
empirical likelihood} (GEL) \citep{kitamuraEmpiricalLikelihoodMethods2007,neweyHigherOrderProperties2004}.
We develop a unified asymptotic framework that accommodates any convex
divergence function with non-negative support and restricted strong
convexity. This generalization, shown in \ref{thm:wg_convergence},
allows our method to encompass several popular divergence-based estimators,
including empirical likelihood \citep{owen1988empirical} and entropy
balancing \citep{hainmuellerEntropyBalancingCausal2012}.
Our approach, as a convex optimization programming, is easy to implement
numerically. We conduct Monte Carlo simulations to demonstrate the
finite sample performance of our proposal in comparison with other
off-the-shelf machine learning methods. Finally, as an empirical example
we evaluate the impact of Brexit on the real GDP growth of the United
Kingdom. We find substantial economic loss after Brexit, and the estimated
weights of the control units from our method are more interpretable
than those from SCM.
In the development of the asymptotics, we make three original theoretical
contributions. Firstly, as mentioned above, we set up a general framework
that accommodates the information-theoretic approach in formulating
SCM by maintaining the simplex constraint. Secondly, we overcome a
key technical challenge in handling many non-negativity constraints
rising from the simplex constraint. The same issue appeared in risk
management for large portfolios \citep{jagannathan2003risk}, but
to the best of our knowledge we are unaware of a theoretical investigation
of the consequence. We provide a rigorous proof by taking the non-negativity
constraints into consideration. In particular, \ref{rem:key-tech-remark}
explains our innovative idea in pushing the estimated vector to its
desirable target by establishing the compatibility condition \citep{buhlmann2011statistics}
that connects the $L_{2}$-norm of the high-dimensional weight vector
with its low-dimensional counterpart. Last but not least, while the
group structures and the low-dimensional factor model are dimension-reduction
assumptions, we allow both the number of groups ($K$) and the number
of factors ($r$) to diverge to infinity, regardless of their relative
order. It offers a flexible accommodation for approximate group features
\citep{bonhomme2022discretizing} and factor propagation \citep{feng2020taming},
and it breaks the severe restriction $K\leq r$ in \citet{shi$ell_2$relaxationApplicationsForecast2022}.
\medskip
This paper stands on several strands of vast literature. We follow
\citet{abadieUsingSyntheticControls2021}'s proposal by maintaining
the simplex constraint, while in the meantime take advantage of the
latent group structures. \citet{doudchenkoBalancingRegressionDifferenceDifferences2016}
develop a constrained regression framework for balancing, regression,
difference-in-difference, and SCM. Ours serves as an complementary
method for estimating individual weights in their procedure when the
donor pool exhibits latent groups. When the pre-treatment outcomes
of the treated unit lie in the convex hull of the pre-treatment outcomes
of the control units, the estimated weights induced by SCM may not
be unique. This occurs particularly when $J$ is larger than $T_{0}$.
\citet{abadiePenalizedSyntheticControl2021} utilize a penalization
method to guarantee a unique solution, where the penalty is imposed
on the pairwise distance of covariates between units. \citet{arkhangelsky2021synthetic}
and \citet{chernozhukov2021exact} propose different versions of penalized
SCM, including $L_{1}$- (Lasso), $L_{2}$- (Ridge), and the elastic
net. Departing from the conventional penalty-based regularization,
this paper takes the relaxation approach to guarantee the uniqueness
of the solution.
Our method is closely related to forecast combination \citep{batesCombinationForecasts1969}.
Many generalizations have been proposed to improve the robustness
of the optimization-based combination methods, in particular machine
learning techniques such as regularization \citep{dieboldMachineLearningRegularized2019}.
Another popular method for constructing counterfactuals is the panel
data approach (PDA) proposed by \citet{hsiao2012panel}. PDA also
uses linear combination of outcomes from control units to predict
the counterfactual. It does not impose the simplex constraint as in
SCM. The focal issue of PDA concerns the selection of control units.
\citet{li2017estimation} propose a Lasso method, while \citet{shi2023forward}
suggest forward selection.
A variety of asymptotic frameworks have appeared to justify SCM. \citet{abadieSyntheticControlMethods2010}
derive bounds on the SCM estimator when perfect pre-treatment fit
is met under fixed $J$ and diverging $T_{0}$. \citet{ferman2021synthetic}
and \citet{fermanPropertiesSyntheticControl2021} analyze the asymptotic
property of estimated weights under the imperfect pre-treatment fit
assumption when $T_{0}$ goes to infinity while $J$ is held fixed
and diverges, respectively. A key message is that when $J$ is fixed,
the SCM estimator will be biased even $T_{0}$ goes to infinity. If
$J$ diverges, the estimator will be unbiased as weights spread out
across the units. In a panel of two-way fixed effects, \citet{arkhangelsky2023large}
show consistency and asymptotic normality of SCM where the treatment
is selected on permanent and time-varying components. Our theory adds
to this literature by demonstrating the consistency of SCM-relaxation
when $J$ diverges into the high-dimensional regime.
Finally, \citet{wang2020minimal}, \citet{hainmuellerEntropyBalancingCausal2012}
and \citet{zheng2024dynamic} estimate the weights by minimizing an
information-theoretic divergence measure. While the first paper is
for general causal inference, the latter two focus on SCM. \citet{hainmuellerEntropyBalancingCausal2012}
works with the entropy, and \citet{zheng2024dynamic} use the empirical
likelihood (EL). Our estimator is different from theirs in two aspects.
First, we do not attempt to exactly match the synthetic control's
outcomes with those of the treated unit but instead take the route
of relaxation. Second, our estimator could be adapted to a class of
general convex functions including both entropy and EL. We provide
in-depth theoretical analysis about these specific choices.
\textbf{Organization. }The rest of the paper is organized as the following.
Section~\ref{sec:Model-and-estimation}\textbf{ }sets up the data
generating process (DGP) and motivates the SCM-relaxation method.
Section \ref{sec:Asymptotic-theory} then establishes the asymptotic
theory of the SCM-relaxation estimators under the group structures
of factor loadings. We conduct Monte Carlo experiments in Section~\ref{sec:Monte-Carlo-experiments},
and in Section~\ref{sec:Empirical-applications} we apply our method
to evaluate the Brexit shock. Proofs of all theoretical results are
relegated to the online appendix.
\textbf{Notations.} For a positive integer $N$, we use $[N]$ as
shorthand for $\{1,2,\dots,N\}$. Let $a\wedge b:=\min(a,b)$. \emph{Absolute
constants }are positive finite numbers independent of sample size.
Vectors and matrices are in bold case. For a generic $n\times1$ vector
$\bm{a}$, let $\left\Vert \bm{a}\right\Vert _{p}=(\sum_{i\in[n]}|a_{i}|^{p})^{1/p}$
and $\left\Vert \bm{a}\right\Vert _{\infty}=\max_{i\in[n]}|a_{i}|$
be its $L_{p}$- and $L_{\infty}$-norm, respectively. We denote $\bm{1}_{n}$
and $\bm{0}_{n}$ as the $n\times1$ vector of ones and zeros, respectively.
The $n$-dimensional simplex is $\Delta_{n}=\{\bm{a}\in\mathbb{R}^{n}\colon\min_{i\in[n]}a_{i}\geq0,\bm{a}'\bm{1}_{n}=1\}$.
For an $m\times n$ matrix $\bm{A}$, let $\varsigma_{\max}(\bm{A})$
be its largest singular value and $\varsigma_{\min}(\bm{A})$ its
minimum \emph{nonzero} singular value. Let $\|\bm{A}\|_{2}=\varsigma_{\max}(\bm{A})$
be the spectral norm. For a symmetric matrix $\bm{B}$, let $\phi_{\min}(\bm{B})$
and $\phi_{\max}(\bm{B})$ be its minimum and maximum eigenvalues,
respectively. The symbol ``$\to_{p}$'' signifies \emph{convergence
in probability}.
\section{Setup }\label{sec:Model-and-estimation}
\subsection{Model}
A researcher possesses data of $J+1$ cross-sectional units and $T_{0}+T_{1}$
time periods, where $T_{0}$ is the number of pre-treatment periods
and $T_{1}$ the number of post-treatment periods. Let $\mathcal{T}_{0}:=[T_{0}]$
denote the index set of the pre-treatment periods, $\mathcal{T}_{1}:=\{T_{0}+1,\dots,T_{0}+T_{1}\}$
denote that of the post-treatment periods, and $\mathcal{T}:=\mathcal{T}_{0}\cup\mathcal{T}_{1}$.
Unit $j=0$ is the sole individual that received a treatment during
$t\in\mathcal{T}_{1}$. The remaining $J$ units, indexed by $j\in[J]$,
make the donor pool of potential controls. In the potential outcome
framework, let $y_{jt}^{I}$ be the outcome of unit $j$ being treated
at time $t$, and $y_{jt}^{I}$ if untreated.\footnote{Following the literature, the superscripts ``$N$'' and ``$I$''
stand for ``non-intervened'' (interchangeable for ``untreated'')
and ``intervened'' (interchangeable for ``treated''), respectively.} In reality, we observe $y_{jt}=d_{jt}y_{jt}^{I}+(1-d_{jt})y_{jt}^{N}$,
where $d_{jt}=1$ for treatment and $d_{jt}=0$ otherwise. We consider
the simple case that only one individual is treated after an event,
and therefore the treatment indicator is $d_{jt}=1\left\{ j=0,t\in\mathcal{T}_{1}\right\} $,
where $1\left\{ \cdot\right\} $ is the indicator function. The treatment
effect at time $t\in\mathcal{T}_{1}$ is
\[
\tau_{0t}=y_{0t}^{I}-y_{0t}^{N}.
\]
Since $y_{0t}^{I}$ is observable while $y_{0t}^{N}$ is unobserved
in the post-treatment periods, the counterfactual $y_{0t}^{N}$ must
be estimated from the data.
This paper presents the simple forms of SCM with the outcome variable
only, as in \citet{ferman2021synthetic}, \citet{fermanPropertiesSyntheticControl2021},
and \citet{chenSyntheticControlOnline2023}. SCM takes a linear combination
$\sum_{j\in[J]}w_{j}y_{jt}$ of the control units' outcomes, where
the weights $w_{j}$ fall into the $J$-dimensional simplex $\Delta_{J}$---they
are non-negative and sum to one. The synthetic control estimator for
$\bm{w}=(w_{1},\dots,w_{J})'$ is obtained by minimizing the $L_{2}$-distance
between the treated unit and synthetic control in the pre-treatment
periods, i.e.,
\begin{equation}
\hat{\bm{w}}^{\mathrm{SC}}:=\arg\min_{\bm{w}\in\Delta_{J}}\left\Vert \bm{y}_{0}-\bm{Y}\bm{w}\right\Vert _{2}^{2},\label{eq:SCM}
\end{equation}
where $\bm{y}_{j}=(y_{j1},\dots,y_{jT_{0}})'$, $j\in\{0\}\cup[J]$
and $\bm{Y}=(\bm{y}_{1},\dots,\bm{y}_{J})$.\footnote{When additional covariates are available, the synthetic control can
incorporate them into the quadratic form. \citet{abadie2003economic}
minimizes $(\bm{x}_{0}-\bm{X}\bm{w})'\bm{V}(\bm{x}_{0}-\bm{X}\bm{w})$,
where $\bm{X}=(\bm{x}_{1},\cdots,\bm{x}_{J})$ consists of individual
characteristics $\bm{x}_{j}$ which can include $\boldsymbol{y}_{j}$,
and $\bm{V}$ is a $T_{0}\times T_{0}$ symmetric and positive semi-definite
matrix.} \citet{abadieSyntheticControlMethods2010} motivate (\ref{eq:SCM})
by a factor model. Suppose that the untreated potential outcomes $y_{jt}^{N}$
of each individual are generated from the linear factor model
\begin{align}
y_{jt}^{N} & =\bm{\lambda}_{j}'\bm{f}_{t}+u_{jt}\quad\text{for \ensuremath{j\in\{0\}\cup[J]} and \ensuremath{t\in\mathcal{T}}},\label{eq:model}
\end{align}
where $u_{jt}$ is the idiosyncratic error, $\bm{f}_{t}$ is an $r\times1$
vector of latent common factors at time $t$, and it is multiplied
by an $r\times1$ vector $\bm{\lambda}_{j}$ of factor loadings for
individual $j$.\footnote{The factor model (\ref{eq:model}) can incorporate individual- and
time-specific fixed effects by setting $\bm{\lambda}_{j}=(1,\alpha_{j},\tilde{\bm{\lambda}}_{j}')'$
and $\bm{f}_{t}=(\gamma_{t},1,\tilde{\bm{f}}_{t}')'$.} Throughout this paper, we view the factors $f_{t}$ as random whereas
factor loadings $\bm{\lambda}_{j}$ as deterministic. The common factors
shared by the treated unit and the untreated ones generate correlations
to be utilized for counterfactual prediction. \citet{abadieUsingSyntheticControls2021}
stresses the importance of the simplex constraint. It ensures that
all members in the donor pool make non-negative contributions as a
weighted average, and in practice the solution $\hat{\bm{w}}^{\mathrm{SC}}$
is often sparse with many zeros.
In practical observational studies with panel data, researchers may
have many potential control units in the donor pool, while the time
dimension, mostly collected in low frequency, is not too long. In
\citet{abadieSyntheticControlMethods2010}'s example of California
tobacco control program $(J,T_{0})=(16,30)$, and \citet{abadie2015comparative}'s
example of German reunification has $(J,T_{0})=(38,19)$. The number
of control units $J$ is often comparable to, or even exceeds, the
number of pre-treatment periods $T_{0}$. To prevent in-sample over-fitting
that may occur when $J>T_{0}$, we propose the SCM-relaxation scheme.
\subsection{SCM-relaxation}
To motivate SCM-relaxation, we consider the Lagrangian of (\ref{eq:SCM}).
When no confusion arises, we suppress the superscript $N$ for the
pre-treatment potential outcome $y_{jt}^{N}$ for $t\in\mathcal{T}_{0}$
as they are observable. We temporarily ignore the non-negativity constraint
and write the Lagrangian as
\[
\frac{1}{2}\left\Vert \bm{y}_{0}-\bm{Y}\bm{w}\right\Vert _{2}^{2}+\gamma(\bm{w}'\bm{1}_{J}-1),
\]
where $\gamma$ is the Lagrangian multiplier associated with the equality
constraint $\bm{w}'\bm{1}_{J}=1$. This is \citet{doudchenkoBalancingRegressionDifferenceDifferences2016}
and \citet{li2020statistical}'s modified SCM. The corresponding Karush--Kuhn--Tucker
(KKT) conditions are
\begin{equation}
\hat{\bm{\Sigma}}\bm{w}-\hat{\bm{\Upsilon}}+\gamma\bm{1}_{J}=\bm{0}_{J}\quad\text{and}\quad\bm{w}'\bm{1}_{J}=1.\label{eq:KKT_SCM}
\end{equation}
where $\hat{\bm{\Sigma}}:=T_{0}^{-1}\bm{Y}'\bm{Y}$ is an $J\times J$
cross moment matrix of the control units, and $\hat{\bm{\Upsilon}}:=T_{0}^{-1}\bm{Y}'\bm{y}_{0}$.
A unique closed-form solution to (\ref{eq:KKT_SCM}) is
\[
\hat{\bm{\Sigma}}^{-1}\biggl(\hat{\bm{\Upsilon}}-\frac{\bm{1}_{J}'\hat{\bm{\Sigma}}^{-1}\hat{\bm{\Upsilon}}-1}{\bm{1}_{J}'\hat{\bm{\Sigma}}^{-1}\bm{1}_{J}}\bm{1}_{J}\biggr)
\]
when $\hat{\bm{\Sigma}}$ is invertible, but if $J$ is close to $T_{0}$,
some eigenvalues of $\hat{\bm{\Sigma}}$ are close to zero, leading
to numerical instability. Moreover, the invertibility of $\hat{\bm{\Sigma}}$
must be violated if $J>T_{0}$.
To stabilize the solution, we estimate the weights by solving the
following convex optimization problem
\begin{equation}
\min_{\bm{w}\in\Delta_{J},\gamma\in\mathbb{R}}\left\Vert \bm{w}\right\Vert _{2}^{2}\quad\text{s.t.}\ \Vert\hat{\bm{\Sigma}}\bm{w}-\hat{\bm{\Upsilon}}+\gamma\bm{1}_{J}\Vert_{\infty}\leq\eta,\label{eq:SCM-relaxation}
\end{equation}
where $\eta=\eta_{T_{0},J}$ is a user-specified tuning parameter
that depends on $T_{0}$ and $J$ but we suppress the subscript for
conciseness. Here we follow \citet{abadieUsingSyntheticControls2021}
to keep the simplex constraint $\bm{w}\in\Delta_{J}$. We call (\ref{eq:SCM-relaxation})
the\emph{ $L_{2}$-SCM-relaxation} problem, for $\Vert\hat{\bm{\Sigma}}\bm{w}-\hat{\bm{\Upsilon}}+\gamma\bm{1}_{J}\Vert_{\infty}\leq\eta$
is a ``relaxation'' of the FOC (\ref{eq:KKT_SCM}).
The problem (\ref{eq:SCM-relaxation}) is strictly convex. In operational
research, the set
\begin{equation}
S_{\eta}:=\left\{ \left(\bm{w},\gamma\right)\in\Delta_{J}\times\mathbb{R}:\Vert\hat{\bm{\Sigma}}\bm{w}-\hat{\bm{\Upsilon}}+\gamma\bm{1}_{J}\Vert_{\infty}\leq\eta\right\} \label{eq:feasible-set}
\end{equation}
is called the \emph{feasible set}. If $S_{\eta}\neq\emptyset$, then
the solution to (\ref{eq:SCM-relaxation}) is unique. The $L_{2}$-norm
encourages diversified weights across similar control units to lower
the prediction variance. It is inspired by the Dantzig selector \citep{candes2007dantzig}
which minimizes the $L_{1}$-norm of parameters for sparsity, and
$L_{2}$-relaxation \citep{shi$ell_2$relaxationApplicationsForecast2022}
in the forecast combination problem.
Let $(\hat{\bm{w}},\hat{\gamma})$ be the the solution to (\ref{eq:SCM-relaxation})
if $S_{\eta}\neq\emptyset$, which holds if $\eta$ is sufficiently
large. On the one hand, if $\eta\geq$$\Vert\hat{\bm{\Sigma}}\boldsymbol{1}_{J}/J-\hat{\bm{\Upsilon}}\Vert_{\infty}$,
then $\left(\hat{\bm{w}}=J^{-1}\boldsymbol{1}_{J},\widehat{\gamma}=0\right)$
is feasible and the objective function reaches the lowest bound.\footnote{The seemingly naive equal weight solution is a surprisingly robust
estimator in empirical applications in economics and finance, leading
to the ``forecast combination puzzle'' \citep{Stock2004combination}
and equal-weight portfolio \citep{demiguel2009optimal}.} On the other hand, if $S_{\eta=0}$ is nonempty, then (\ref{eq:SCM-relaxation})
entails exact satisfaction of the FOC. With a proper choice of $\eta\in[0,\Vert\hat{\bm{\Sigma}}\boldsymbol{1}_{J}/J-\hat{\bm{\Upsilon}}\Vert_{\infty}]$,
the relaxation scheme balances the bias and variance and mitigates
the sensitivity of SCM weights to the sampling noise in $(\hat{\bm{\Sigma}},\hat{\bm{\Upsilon}})$,
effectively preventing in-sample over-fitting.
Next, we explore the asymptotic properties of the relaxation scheme.
\section{Asymptotic Theory }\label{sec:Asymptotic-theory}
Statistical analysis of high-dimensional problems typically postulates
certain structures on the DGP for dimension reduction. For example,
variable selection methods such as Lasso \citep{tibshirani1996regression}
and SCAD \citep{fan2001variable} are suitable for linear regressions
with sparse coefficients, meaning most of the parameters are either
exactly zero or approximately zero. Similarly, in large covariance
matrix estimation, various structures have been proposed: \citet{bickel2008covariance}
impose sparse covariance entries, and \citet{tong2025cluster} assume
a block correlation matrix.
The group pattern in panel data has been widely adopted as a bridge
between homogeneous coefficients and fully heterogeneous coefficients
\citep{bonhomme2015grouped,su2016identifying,vogt2017classification}.
Were the factors $\boldsymbol{f}_{t}$ observable, the factor loadings
play the role as the slope coefficients. Based on this resemblance,
we assume latent groups across the factor loadings of the control
units.\footnote{In analysis of large factor models, \citet{he2024penalized} also
borrows this setting for dimension reduction. } The donor pool is partitioned into $K$ disjoint groups $\mathcal{G}_{1},\dots,\mathcal{G}_{K}$,
where $\mathcal{G}_{k}\cap\mathcal{G}_{\ell}=\emptyset$ for any $k\neq\ell$
and $\cup_{k=1}^{K}\mathcal{G}_{k}=[J]$. The factor loadings are
homogeneous within each group, that is, for each $k\in[K]$, for any
$i,j\in\mathcal{G}_{k}$, $\bm{\lambda}_{i}=\bm{\lambda}_{j}=\bm{\lambda}_{\mathcal{G}_{k}}$.
To encode the group identities, we introduce a $J\times K$ membership
matrix $\bm{Z}$, where $\bm{Z}_{jk}=1\{j\in\mathcal{G}_{k}\}$. Then
the group structure can be written as
\begin{equation}
\underset{(J\times r)}{\bm{\Lambda}}=\underset{(J\times K)}{\bm{Z}}\underset{(K\times r)}{\bm{\Lambda}^{\mathrm{co}}},\label{eq:grouped_structure}
\end{equation}
where $\bm{\Lambda}=(\bm{\lambda}_{1},\dots,\bm{\lambda}_{J})'$ is
the matrix of factor loadings and $\bm{\Lambda}^{\mathrm{co}}=(\bm{\lambda}_{1}^{\mathrm{co}},\dots,\bm{\lambda}_{K}^{\mathrm{co}})^{\prime}$
collects all group-specific loadings; here the superscript ``co''
stands for ``core.'' Let $J_{k}:=|\mathcal{G}_{k}|$ be the number
of controls in the $k$-th group; obviously $J=\sum_{k\in[K]}J_{k}$.
\begin{rem}
The group structure is common in real data. For instance, the clustered
variance is ubiquitous in microeconometrics---within each cluster
the covariances are equicorrelated whereas the cross-cluster covariances
are zero \citep{abadie2023should}. In time series forecast combination,
\citet{chan2018some} model the presence of the ``best'' unbiased
forecast and thus all forecasters incurs the same unpredictable idiosyncratic
error rising from the target variable to form one equicorrelated group.
In finance, \citet{engle2012dynamic} employ the Standard Industrial
Classification to assign the asset groups in the volatility matrix.
While each of these papers explicitly specifies a prior criterion
for group allocation, no knowledge about the membership is required
to implement $L_{2}$-SCM-relaxation; block equicorrelation is taken
as a latent structure in this paper, and we do not intend to recover
the group membership.
\end{rem}
\subsection{Conditions}
While the literature mostly assumes finite groups and finite factors,
here we allow $K$ and $r$ to be either fixed or diverging, which
accommodates ``many groups'' and ``many factors.'' In asymptotic
statements, we will explicitly send $T_{0}\to\infty$, whereas $K$,
$r$, $J$, and $T_{1}$ are viewed as deterministic functions of
$T_{0}$ that go to infinity. We impose the following regularity conditions.
\begin{assumption}[Factors and loadings]
\label{assu:DGP} When $T_{0}$ is sufficiently large, there exist
two absolute constants $\underline{c}$ and\/ $\bar{c}$ such that:
\begin{enumerate}[label=(\alph*),parsep=0pt,topsep=0pt]
\item \label{enu:factors}Factors: $\hat{\bm{\Omega}}_{\bm{F}}:=T_{0}^{-1}\bm{F}^{\prime}\bm{F}$
is invertible, and there exists an $r\times r$ deterministic positive
definite matrix $\bm{\Omega}_{\bm{F}}$ with $\underline{c}\leq\phi_{\min}(\bm{\Omega}_{\bm{F}})\leq\phi_{\max}(\bm{\Omega}_{\bm{F}})\leq\bar{c}$
such that $\|\hat{\bm{\Omega}}_{\bm{F}}-\bm{\Omega}_{\bm{F}}\|_{2}\to_{p}0$
as $T_{0}\to\infty$.
\item \label{enu:Loadings}Loadings: $\|\bm{\lambda}_{0}\|_{2}+\sup_{k\in[K]}\|\bm{\lambda}_{k}^{\mathrm{co}}\|_{2}\leq\bar{c}$,
$\mathrm{rank}(\bm{\Lambda}^{\mathrm{co}})=K\wedge r$, and $\underline{c}K/(K\wedge r)\leq\varsigma_{\min}^{2}(\bm{\Lambda}^{\mathrm{co}})\leq\varsigma_{\max}^{2}(\bm{\Lambda}^{\mathrm{co}})\leq\bar{c}K/(K\wedge r)$.
\end{enumerate}
\end{assumption}
Due to their multiplicative form, the factors and loadings cannot
be uniquely identified. \ref{assu:DGP} requires the existence of desirable
factors and the loadings. Part~\ref{enu:factors} is standard in
factor models \citep{bai2002determining,bai2003inferential}. The
invertibility condition and the boundedness of eigenvalues ensure
that the factors are non-degenerate. We require $\hat{\bm{\Omega}}_{\bm{F}}$
to be invertible itself, not just have an invertible limit, because
the expression of oracle weight (introduced later) implicitly relies
on a nonsingular $\hat{\bm{\Omega}}_{\bm{F}}$, though in an asymptotic
sense, an invertible limit would suffice. Part~\ref{enu:Loadings}
requires bounded loadings $\boldsymbol{\lambda}_{0}$ for the treated
unit. Regarding the control units, a sufficient condition for a bounded
$\|\bm{\lambda}_{k}^{\mathrm{co}}\|_{2}$ is assuming the maximum
loading of order $O(1/\sqrt{r})$, as what \citet{liDeterminingNumberFactors2017}
do, to accommodate a diverging number of factors. The rank condition
implies that $\bm{\Lambda}^{\mathrm{co}}$ has $K\wedge r$ nonzero
singular values, and thus $\sum_{k\in[K\wedge r]}\varsigma_{k}^{2}(\bm{\Lambda}^{\mathrm{co}})=\mathrm{trace}(\bm{\Lambda}^{\mathrm{co}\prime}\bm{\Lambda}^{\mathrm{co}})=\sum_{k\in[K]}\|\bm{\lambda}_{k}^{\mathrm{co}}\|_{2}^{2}=O(K)$.
The assumption implies all singular values are of the same order.
Intuitively, these conditions make sure that the core loadings are
immune from collinearity and no single group makes dominant contribution
to the variances of the outcome variables.
Under the latent group structure (\ref{eq:grouped_structure}), the
sample covariance matrix for $\bm{Y}$ can be decomposed as $\hat{\bm{\Sigma}}=\hat{\bm{\Sigma}}^{*}+\hat{\bm{\Sigma}}^{\mathrm{e}}$,
where
\[
\hat{\bm{\Sigma}}^{*}:=\frac{1}{T_{0}}\bm{Z}\bm{\Lambda}^{\mathrm{co}}\bm{F}'\bm{F}\bm{\Lambda}^{\mathrm{co}}{}'\bm{Z}',
\]
and thus $\hat{\bm{\Sigma}}^{\mathrm{e}}:=(\bm{Z}\bm{\Lambda}^{\mathrm{co}}\bm{F}'\bm{U}+\bm{U}'\bm{F}\bm{\Lambda}^{\mathrm{co}}{}'\bm{Z}'+\bm{U}^{\prime}\bm{U})/T_{0}$,
with $\bm{U}=(\bm{u}_{1},\dots,\bm{u}_{J})$ and $\bm{u}_{j}=(u_{j1},\dots,u_{jT_{0}})'$
for $j\in[J]$. Here the superscript ``$*$'' denotes the leading
component---the sample covariance of the infeasible signal $\boldsymbol{\Lambda}\bm{f}_{t}$,
whereas ``e'' stands for the remaining idiosyncratic error. Similarly,
the covariance between $\bm{Y}$ and $\bm{y}_{0}$ can be written
as $\hat{\bm{\Upsilon}}=\hat{\bm{\Upsilon}}^{*}+\hat{\bm{\Upsilon}}^{\mathrm{e}}$,
where
\[
\hat{\bm{\Upsilon}}^{*}:=\frac{1}{T_{0}}\bm{Z}\bm{\Lambda}^{\mathrm{co}}\bm{F}'\bm{F}\bm{\lambda}_{0},
\]
and $\hat{\bm{\Upsilon}}^{\mathrm{e}}:=(\bm{Z}\bm{\Lambda}^{\mathrm{co}}\bm{F}'\bm{u}_{0}+\bm{U}^{\prime}\bm{F}\bm{\lambda}_{0}+\bm{U}^{\prime}\bm{u}_{0})/T_{0}$
with $\bm{u}_{0}=(u_{01},\dots,u_{0T_{0}})^{\prime}$. Next, we specify
an asymptotic target for the $L_{2}$-SCM-relaxation estimator $\hat{\bm{w}}$.
Consider the oracle problem of $L_{2}$-SCM-relaxation with infeasible
data $y_{jt}^{*}\coloneqq\bm{\lambda}_{j}'\bm{f}_{t}$ that is free
of the idiosyncratic errors:
\begin{align}
\min_{\bm{w}\in\Delta_{J},\gamma\in\mathbb{R}}{}\left\Vert \bm{w}\right\Vert _{2}^{2}\quad & \text{s.t. }\hat{\bm{\Sigma}}^{*}\bm{w}-\hat{\bm{\Upsilon}}^{*}+\gamma\bm{1}_{J}=\bm{0}_{J}.\label{eq:oracle_relaxation_problem}
\end{align}
Denote the solution to the above problem as $\bm{w}^{*}$.
\begin{rem}
\label{rem:recovery-loading}Let $S^{*}:=\{(\bm{w},\gamma)\in\Delta_{J}\times\mathbb{R}:\hat{\bm{\Sigma}}^{*}\bm{w}-\hat{\bm{\Upsilon}}^{*}+\gamma\bm{1}_{J}=\bm{0}_{J}\}$
be the feasible set of the oracle problem (\ref{eq:oracle_relaxation_problem}).
A sufficient condition to ensure $S^{*}\neq\emptyset$ is the existence
of some $\bm{\pi}\in\Delta_{K}$ such that $\bm{\Lambda}^{\mathrm{co}\prime}\bm{\pi}=\bm{\lambda}_{0}$,
under which $(\bm{w}=\bm{Z}(\bm{Z}'\bm{Z})^{-1}\bm{\pi},\gamma=0)$
is feasible. This condition means that the true factor loadings of
the treated can be exactly recovered from the loadings of the controls.
\end{rem}
The following lemma characterizes the oracle solution $\bm{w}^{*}$.
\begin{lem}[Oracle weight]
\label{lem:oracle_target_relaxation}Suppose \ref{assu:DGP} holds.
If the solution to (\ref{eq:oracle_relaxation_problem}) exists, it
must be within-group equal, i.e., $w_{i}^{*}=w_{j}^{*}$ if $i$ and
$j$ belong to the same group. Furthermore, if the solution is in
the interior of the simplex, i.e., $w_{j}^{*}>0$ for all $j\in[J]$,
then it has the following closed form
\[
\bm{w}^{*}=\bm{Z}(\bm{Z}'\bm{Z})^{-1}\bm{w}_{\mathcal{G}}^{*},
\]
where the expression of $\bm{w}_{\mathcal{G}}^{*}=(w_{\mathcal{G}_{1}}^{*},\dots,w_{\mathcal{G}_{K}}^{*})'$
is given by (\ref{eq:w_G_oracle}) for $K\leq r$, and (\ref{eq:wg_oracle_r_less_K_in_col})
or (\ref{eq:wg_oracle_r_less_K_not_in_col}) for $K>r$ in Appendix~\ref{subsec:Proof-of-Lemma}.
\end{lem}
\ref{lem:oracle_target_relaxation} demonstrates that for $j\in\mathcal{G}_{k}$,
the oracle weights $w_{j}^{*}=J_{k}^{-1}w_{\mathcal{G}_{k}}^{*}$
are evenly distributed within group $\mathcal{G}_{k}$. We allow the
group weight $w_{\mathcal{G}_{k}}^{*}=0$ for some $k$. Such group
does not contribute to the prediction of $\bm{y}_{0}$.
We will show that $\hat{\bm{w}}$ and $\bm{w}^{*}$ are sufficiently
close in both $L_{1}$ and $L_{2}$ distance. Define $\phi_{J,T_{0}}:=\sqrt{(\log J)/T_{0}}$,
which will serve as an upper bound of sampling errors and average
cross-sectional dependence in the following assumption. Note that
$\phi_{J,T_{0}}$ can shrink to 0 even if $J$ is much larger than
$T_{0}$; for example, if $J=T_{0}^{2}$, then $\phi_{J,T_{0}}=\sqrt{2(\log T_{0})/T_{0}}\to0$
as $T_{0}\to\infty$.
\begin{assumption}
\label{assu:errors}The factors and idiosyncratic errors satisfy
\begin{enumerate}[label=(\alph*),parsep=0pt,topsep=0pt]
\item \label{enu:uncorr}$(u_{jt})_{j\in\{0\}\cup[J]}$ and $\bm{f}_{t}$
are uncorrelated for all $t\in\mathcal{T}_{0}$;
\item \label{enu:Fu}$\sup_{j\in\{0\}\cup[J]}\|T_{0}^{-1}\bm{F}'\bm{u}_{j}\|_{\infty}=O_{p}(\phi_{J,T_{0}})$;
\item \label{enu:uu}$\sup_{j\in[J]}J^{-1}\sum_{i=1}^{J}|\mathbb{E}(T_{0}^{-1}\bm{u}_{i}^{\prime}\bm{u}_{j})|+\sup_{j\in[J]}|\mathbb{E}(T_{0}^{-1}\bm{u}_{0}^{\prime}\bm{u}_{j})|=O(\phi_{J,T_{0}})$;
\item \label{enu:uu_dev}$\sup_{i,j\in\{0\}\cup[J]}\bigl|T_{0}^{-1}[\bm{u}_{i}^{\prime}\bm{u}_{j}-\mathbb{E}(\bm{u}_{i}^{\prime}\bm{u}_{j})]\bigr|=O_{p}(\phi_{J,T_{0}})$.
\end{enumerate}
\end{assumption}
\ref{assu:errors} imposes high-level conditions on idiosyncratic errors.
Specifically, Condition~\ref{enu:uncorr} means that $\mathbb{E}(\bm{f}_{t}u_{jt})=\bm{0}_{r\times1}$
and Condition~\ref{enu:Fu} ensures that the corresponding sampling
errors is controlled by $\phi_{J,T_{0}}$ uniformly for all $j$ and
$t$. Condition~\ref{enu:uu} allows for mild cross-sectional dependence
in idiosyncratic errors of the control units; note that this is a
weak assumption since it only requires that on average $\mathbb{E}(T_{0}^{-1}\bm{u}_{i}^{\prime}\bm{u}_{j})$
should be controlled by $\phi_{J,T_{0}}$. On the other hand, the
weak correlation of the idiosyncratic errors between the treated unit
and the controls is for simplicity, so that the predictability is
dominantly due to the factors. Sampling errors are restricted in Condition~\ref{enu:uu_dev}.
The order $\phi_{J,T_{0}}=\sqrt{(\log J)/T_{0}}$ for the sup-norm
of the sampling errors is commonplace in high-dimensional statistics,
and can be established from low-level assumptions; see, for instance,
\citet[Lemma~4]{fan2013large} and \citet[Section 6.5]{wainwright2019high}.
\begin{assumption}
\label{assu:eta}As $T_{0}\to\infty$,
\begin{enumerate}[label=(\alph*),parsep=0pt,topsep=0pt]
\item \label{enu:group-size}Group size: $\min_{k\in[K]}KJ_{k}/J\geq\underline{c}$
for some absolute constant $\underline{c}$;
\item \label{enu:tuning}Tuning parameter: $(K\wedge r)^{2}\eta+K\sqrt{r}\phi_{J,T_{0}}/\eta\to0$.
\end{enumerate}
\end{assumption}
Part~\ref{enu:group-size} of \ref{assu:eta} ensures that none of
the groups is negligible in terms of its size; otherwise a tiny group
contributes little in prediction but its associate weight can be disproportionately
large. Part~\ref{enu:tuning} requires that the tuning parameter
$\eta$ should shrink to 0 in a rate faster than $1/(K\wedge r)^{2}$
to guarantee the convergence of $\hat{\bm{w}}$, but slower than $K\sqrt{r}\phi_{J,T_{0}}$
so that the oracle weight $\bm{w}^{*}$ satisfies the relaxation $\Vert\hat{\bm{\Sigma}}\bm{w}^{*}-\hat{\bm{\Upsilon}}+\gamma^{*}\bm{1}_{J}\Vert_{\infty}\leq\eta$
with high probability.
\begin{rem}
\citet{shi$ell_2$relaxationApplicationsForecast2022}'s theory hinges
on the assumption $K\leq r$. Here we break this restriction. A large
$K$ makes it possible to use the group structures to approximate
continuously distributed factor loadings \citep{bonhomme2022discretizing},
and a large $r$ can deal with propagation of factors, as is extensively
studied in the financial asset pricing literature \citep{feng2020taming}.
\end{rem}
The tuning parameter $\eta$ plays an important role in the finite
sample behavior of the estimator. In asymptotic analysis we specify
the rate for a proper $\eta$ as the sample size increases. In practice,
the tuning parameter is chosen by cross validation; this is what we
do in the numerical work throughout this paper.
\subsection{$L_{2}$-norm Objective Function}
The assumptions in the previous section are prepared for the consistency
of $L_{2}$-SCM-relaxation.
\begin{thm}[Convergence of $\hat{\bm{w}}$]
\label{thm:w_convergence} If \assuref[s]{DGP}--\ref{assu:eta}
hold, then
\begin{align*}
\|\hat{\bm{w}}-\bm{w}^{*}\|_{1} & =O_{p}\bigl([(K\wedge r)^{2}\eta]^{1/3}\bigr)=o_{p}(1),\text{ and }\\
\|\hat{\bm{w}}-\bm{w}^{*}\|_{2} & =O_{p}\left(\frac{[(K\wedge r)^{2}\eta]^{1/3}}{\sqrt{J}}\right)=o_{p}\left(\frac{1}{\sqrt{J}}\right).
\end{align*}
\end{thm}
\ref{thm:w_convergence} shows that the estimator $\hat{\bm{w}}$ converges
to $\bm{w}^{*}$ under both $L_{1}$- and $L_{2}$-norm with desirable
rates of convergence. Its proof deviates substantially from that in
\citet{shi$ell_2$relaxationApplicationsForecast2022}, which resorts
to the dual problem (unconstrained optimization) of their primal $L_{2}$-relaxation
(constrained optimization). Had we mimic their proof, as many as $J$
non-negativity constraints from the simplex would make the dual formulation
intractable. Therefore, here we must come up with a new strategy that
directly works with the primal problem.
\begin{rem}
\label{rem:key-tech-remark} We briefly discuss the road map of proofs.
Our idea is to first show in \ref{lem:w_star_feasibility} that the
oracle weights are feasible with high probability. The oracle target
$\bm{w}^{*}=\bm{Z}(\bm{Z}'\bm{Z})^{-1}\bm{w}_{\mathcal{G}}^{*}$ lies
in the low-dimensional space spanned by the membership matrix $\bm{Z}$,
which implies an inequality that links $\left\Vert \hat{\bm{w}}-\bm{w}^{*}\right\Vert _{2}$
with $\|\hat{\bm{w}}_{\mathcal{G}}-\bm{w}_{\mathcal{G}}^{*}\|_{2}$
in \ref{lem:link_two_norms}, following Proposition 9.13 in \citet[Section 9.2]{wainwright2019high}.
Then we establishes a version of the \emph{restricted eigenvalue}
(RE),\footnote{While the Lasso literature usually \emph{assumes} the RE \citep{bickel2009simultaneous,buhlmann2011statistics},
we instead \emph{derive} the RE from the group structures. } and it further yields a \emph{compatibility inequality} \citep[Chapter 6.13]{buhlmann2011statistics}
that links $\|\hat{\bm{w}}-\bm{w}^{*}\|_{2}$ and $\|\hat{\bm{w}}-\bm{w}^{*}\|_{1}$
via $\|\hat{\bm{w}}_{\mathcal{G}}-\bm{w}_{\mathcal{G}}^{*}\|_{2}$
as a bridge. These preparatory results allows us to borrow existing
inequalities from the theory of Lasso in high dimension and tailor
them for our purpose in the proof of \ref{thm:w_convergence} in Appendix~\ref{subsec:proof-thm1}.
\end{rem}
An unexpected windfall of this strategy of proof is that the asymptotic
analysis can be carried over into general convex objective functions,
to be elaborated in Section~\ref{subsec:Other-objective-functions}.
Moreover, it overcomes a longstanding technical challenge of establishing
asymptotic results with a high-dimensional simplex constraint, which
was also encountered as the no-short-sell constraint in large portfolios
\citep{jagannathan2003risk}.
\begin{rem}
We compare our \ref{thm:w_convergence} with \citet{fermanPropertiesSyntheticControl2021},
which analyzes the convergence of $\hat{\bm{\lambda}}^{\mathrm{SC}}:=\bm{\Lambda}'\hat{\bm{w}}^{\mathrm{SC}}$
and shows $\|\bm{f}_{t}'(\hat{\bm{\lambda}}^{\mathrm{SC}}-\bm{\lambda}_{0})\|_{2}\to_{p}0$
under a key condition that there exists a sequence of ``oracle''
weight $\bm{w}^{*}$ such that $\bm{\Lambda}'\bm{w}^{*}\to\bm{\lambda}_{0}$,
the asymptotic unbiasedness condition. We instead\emph{ prove} that
the SCM-relaxation estimator $\hat{\bm{w}}$ converges to $\bm{w}^{*}$
at a sufficiently fast rate. Another interesting result in \citet{fermanPropertiesSyntheticControl2021}
is that the SCM weight $\hat{\bm{w}}^{\mathrm{SC}}$ will spread out
over many control units as $J\to\infty$, giving rise to $\Vert\hat{\bm{w}}^{\mathrm{SC}}\Vert_{2}\to_{p}0$,
which motivates \citet{ferman2021synthetic} to suggest reporting
$\Vert\hat{\bm{w}}^{\mathrm{SC}}\Vert_{2}$ in empirical applications
to check whether many control units help reduce the bias. Our \ref{thm:w_convergence},
moreover, delivers an even faster rate $\left\Vert \hat{\bm{w}}\right\Vert _{2}=o_{p}(J^{-1/2})$.
\end{rem}
\ref{thm:w_convergence} has established that SCM-relaxation diversifies
prediction risk over many control units. Next, we continue with the
empirical risk. For a generic weight $\bm{w}\in\Delta_{J}$, define
the in-sample empirical risk as $R_{\mathcal{T}_{0}}(\bm{w}):=T_{0}^{-1}\sum_{t\in\mathcal{T}_{0}}(\bm{w}'\bm{y}_{t}-y_{0t})^{2}$
and similarly the out-of-sample empirical risk as $R_{\mathcal{T}_{1}}(\bm{w})=T_{1}^{-1}\sum_{t\in\mathcal{T}_{1}}(\bm{w}'\bm{y}_{t}-y_{0t})^{2}$.
\begin{thm}[Oracle empirical risks]
\label{thm:oracle_inequalities} Under \assuref[s]{DGP}--\ref{assu:eta},
we have
\begin{enumerate}[label=(\roman*),parsep=0pt,topsep=0pt]
\item \label{enu:in-sample}$R_{\mathcal{T}_{0}}(\hat{\bm{w}})=R_{\mathcal{T}_{0}}(\bm{w}^{*})+o_{p}(1)$
as $T_{0}\to\infty$.
\item \label{enu:out-sample}If $(y_{jt}^{N})_{j\in\{0\}\cup[J],t\in\mathcal{T}_{1}}$
follow the same DGP as that in the pre-treatment periods, and $r\log(J)/T_{1}=O(1)$,
then $R_{\mathcal{T}_{1}}(\hat{\bm{w}})=R_{\mathcal{T}_{1}}(\bm{w}^{*})+o_{p}(1)$
as $T_{0},T_{1}\to\infty$.
\end{enumerate}
\end{thm}
\ref{thm:oracle_inequalities} \ref{enu:in-sample} is an in-sample
oracle equality, and \ref{enu:out-sample} is an out-of-sample oracle
equality, where the extra condition $r\log(J)/T_{1}=O(1)$ is mild
in that it allows $T_{1}$ to diverge at a much slower speed than
$T_{0}$. This theorem shows that empirical risk under the weights
estimated by $L_{2}$-SCM-relaxation from the training data is asymptotically
as low as the oracle $\bm{w}^{*}$, up to an asymptotically negligible
$o_{p}(1)$ gap. It highlights the effectiveness of the relaxation
scheme: even though the relaxation method does not seek to identify
the group membership, the risk of our estimator would be as good as
if we were informed of the infeasible oracle group identities of the
control units.
\subsection{Information-theoretic Objective Functions }\label{subsec:Other-objective-functions}
The $L_{2}$-norm objective in (\ref{eq:SCM-relaxation}) is one of
the information-theoretic divergence measures. Shall we prefer the
$L_{2}$-norm, or it is equivalent if another member of the family
is summoned? We will provide an in-depth analysis in this section.
Denote a generic divergence measure as $g(\cdot)$, which is strictly
convex and sufficiently smooth, and then the associated $g$-SCM-relaxation
problem solves
\begin{equation}
\min_{(\bm{w},\gamma)\in S_{\eta}}{}\sum_{j\in[J]}g(w_{j}),\label{eq:general_SCM_relaxation}
\end{equation}
where $S_{\eta}$ is the feasible set defined by (\ref{eq:feasible-set}),
and the $\boldsymbol{w}$-part of the solution is denoted as $\hat{\bm{w}}_{(g)}$.
Conformably, the oracle weights, denoted $\bm{w}_{(g)}^{*}$, are
the solution to
\begin{equation}
\min_{(\bm{w},\gamma)\in S^{*}}{}\sum_{j\in[J]}g(w_{j}).\label{eq:g-oracle}
\end{equation}
Similar to \ref{lem:oracle_target_relaxation}, it is easy to show
that $\bm{w}_{(g)}^{*}$ is within-group equal, though a closed-form
expression is unavailable for a generic $g$ function.
We study the asymptotic properties of $g$-SCM-relaxation, represented
by two popular choices of $g(\cdot)$. The function $g(x)=-\log x$
implicitly keeps weights positive. This choice is closely related
to the empirical likelihood \citep{owen1988empirical}. We call the
associate problem \emph{EL-SCM-relaxation}, and the solution is denoted
as $\hat{\bm{w}}^{\mathrm{EL}}$. Alternatively, the entropy function
$g(x)=x\log x$ also ensures $x>0$, and it leads to \emph{entropy-SCM-relaxation}
with the solution $\hat{\bm{w}}^{\mathrm{entr}}$. In SCM with fixed
$J$, \citet{zheng2024dynamic} use EL and \citet{hainmuellerEntropyBalancingCausal2012}
employs entropy in association with moment \emph{equalities}. Their
theory cannot be generalized to our high-dimensional contexts, where
the moment conditions must be relaxed by inequalities.
\begin{rem}
The Cressie-Read (CR) discrepancies \citep{Cressie1984multi}, $g(x)=(x^{\gamma+1}-1)/[\gamma(\gamma+1)]$
where $\gamma$ is the parameter indexing the family, include the
$L_{2}$-norm, EL, and entropy functions as special cases: (i) $g(x)=-\log x$
if $\gamma=-1$; (ii) $g(x)=x\log x$ if $\gamma=0$; (iii) $g(x)=x^{2}$
if $\gamma=1$. Indeed, for all $\gamma\in[-1,1]$ the CR discrepancy
functions are strictly convex, and can serve as the objective functions
for the SCM-relaxation formulation.
\end{rem}
\begin{rem}
The choice of $g$ is related to \emph{Bregman divergence}. For any
differentiable convex function $\psi\colon\mathcal{C}\to\mathbb{R}$,
the Bregman divergence between points $\bm{x}$ and $\bm{y}$ in $\mathcal{C}$
is defined as
\[
D_{\psi}(\bm{x},\bm{y})=\psi(\bm{x})-\psi(\bm{y})-\langle\nabla\psi(\bm{y}),\bm{x}-\bm{y}\rangle.
\]
The objective functions we study can all be recast as a Bregman divergence
between $\bm{w}\in\Delta_{J}$ and $\bm{1}_{J}/J$ induced by respective
functions. Specifically, the $L_{2}$-norm is associated with $D_{\psi_{1}}(\bm{w},J^{-1}\bm{1}_{J})=\frac{1}{2}\|\bm{w}-J^{-1}\bm{1}_{J}\|_{2}^{2}$,
the EL function with $D_{\psi_{2}}(\bm{w},J^{-1}\bm{1}_{J})=-\sum_{j\in[J]}(\log w_{j}-\log J^{-1})$,
and the entropy function with $D_{\psi_{3}}(\bm{w},J^{-1}\bm{1}_{J})=\sum_{j\in[J]}w_{j}(\log w_{j}-\log J^{-1})$.
Notice that the $L_{2}$-norm is symmetric, whereas EL and entropy
impose asymmetric penalty on positive and negative deviations from
the simple average $\bm{1}_{J}/J$ since the divergence is asymmetric.
\end{rem}
To carry out analysis in a unified framework, we introduce a few notions
that characterize the shape of function $g$. A differentiable function
$g\colon\mathcal{D}\to\mathbb{R}$ is called $\alpha_{g}$\emph{-strongly
convex} on $\mathcal{D}\subseteq\mathbb{R}$ if for any $x,y\in\mathcal{D}$,
we have $g(x)-g(y)\geq\frac{\mathrm{d}g(y)}{\mathrm{d}y}(x-y)+\frac{\alpha_{g}}{2}(x-y)^{2}$.
Since a function $f$ is convex if and only if $f(x)-f(y)\geq\frac{\mathrm{d}f(y)}{\mathrm{d}y}(x-y)$
for any $x$ and $y$, an equivalent condition for $\alpha_{g}$-strong
convexity is that $f(x)=g(x)-\frac{\alpha_{g}}{2}x^{2}$ is convex.
Clearly, an $\alpha_{g}$-strongly convex function has its Bregman
divergence bounded from below by $\frac{\alpha_{g}}{2}(x-y)^{2}$.
A function $g\colon\mathcal{D}\to\mathbb{R}$ is called $\beta_{g}$-\emph{Lipschitz}
on $\mathcal{D}\subseteq\mathbb{R}$ if for any $x,y\in\mathcal{D}$,
we have $|g(x)-g(y)|\leq\beta_{g}|x-y|$. A differentiable $\beta_{g}$-Lipschitz
function $g$ must posses bounded derivative $|\mathrm{d}g(x)/\mathrm{d}x|\leq\beta_{g}$
for all $x\in\mathcal{D}$.
\begin{assumption}[Tuning parameter, as a surrogate for \ref{assu:eta}\ref{enu:tuning}]
\label{assu:add_assu}Let $\mathcal{W}_{K,J}\subseteq[0,1]$ be a
deterministic interval that can depend on $K$ and $J$ such that
$\hat{w}_{(g),j},w_{(g),k}^{*}\in\mathcal{W}_{K,J}$ with probability
approaching one uniformly all $j\in[J]$ and $k\in[K]$. Suppose
the function $g$ is $\alpha_{g}$-strongly convex and $\beta_{g}$-Lipschitz
on $\mathcal{W}_{K,J}$. The tuning parameter $\eta$ satisfies $(K\wedge r)J^{2}(\beta_{g}/\alpha_{g})^{2}\eta+K\sqrt{r}\phi_{J,T_{0}}/\eta\to0$.
\end{assumption}
Typically, the weights $\hat{w}_{(g),j}$ and $w_{(g),k}^{*}$ lie
in some interval $[\underline{c}/J,\bar{c}K/J]$. \ref{assu:add_assu}
basically requires the ratio $\beta_{g}/\alpha_{g}$ to be small enough
in that interval. We discuss this for $g(x)=-\log x$ and $g(x)=x\log x$.
\begin{example}[EL]
If $g(x)=-\log x$, then for fixed $0<\varepsilon_{1}<\varepsilon_{2}\leq1$,
we can show that $g(x)$ is $\varepsilon_{2}^{-2}$-strongly convex
on $[\varepsilon_{1},\varepsilon_{2}]$ and $\sup_{x\in[\varepsilon_{1},\varepsilon_{2}]}|\mathrm{d}g(x)/\mathrm{d}x|\leq\varepsilon_{1}^{-1}$.
If we take $\varepsilon_{1}=O(1/J)$ and $\varepsilon_{2}=O(K/J)$,
then $\beta_{g}/\alpha_{g}=\varepsilon_{1}^{-1}/\varepsilon_{2}^{-2}=O(K^{2}/J)$.
Therefore \ref{assu:add_assu} holds if $K^{5}\eta+K\sqrt{r}\phi_{J,T_{0}}/\eta\to0$.
\end{example}
\begin{example}[Entropy]
If $g(x)=x\log x$, for fixed $0<\varepsilon_{1}<\varepsilon_{2}\leq1$,
the function $g(x)$ is $\varepsilon_{2}^{-1}$-strongly convex on
$[\varepsilon_{1},\varepsilon_{2}]$ and $\sup_{x\in[\varepsilon_{1},\varepsilon_{2}]}|\mathrm{d}g(x)/\mathrm{d}x|\leq1+|\log\varepsilon_{1}|$.
Taking $\varepsilon_{1}=O(1/J)$ and $\varepsilon_{2}=O(K/J)$, we
get $\beta_{g}/\alpha_{g}=(1+|\log\varepsilon_{1}|)\varepsilon_{2}=O(K[\log J]/J)$.
\ref{assu:add_assu} is satisfied if $K^{3}(\log J)^{2}\eta+K\sqrt{r}\phi_{J,T_{0}}/\eta\to0$.
\end{example}
Given a proper choice of $\eta$, the consistency of $\hat{\bm{w}}_{(g)}$
follows.
\begin{thm}
\label{thm:wg_convergence} If \assuref[s]{DGP}, \ref{assu:errors},
\ref{assu:eta}\ref{enu:group-size}, and \ref{assu:add_assu} hold,
we have
\begin{align*}
\|\hat{\bm{w}}_{(g)}-\bm{w}_{(g)}^{*}\|_{1} & =O_{p}\left(\left[\frac{(K\wedge r)J^{2}\beta_{g}^{2}\eta}{\alpha_{g}^{2}}\right]^{1/3}\right)=o_{p}(1),\text{ and}\\
\|\hat{\bm{w}}_{(g)}-\bm{w}_{(g)}^{*}\|_{2} & =O_{p}\left(\left[\frac{(K\wedge r)J^{2}\beta_{g}^{2}\eta}{\alpha_{g}^{2}}\right]^{1/3}\frac{1}{\sqrt{J}}\right)=o_{p}(J^{-1/2}).
\end{align*}
\end{thm}
\ref{thm:wg_convergence} guarantees that $\hat{\bm{w}}_{(g)}$ behaves
as well as the oracle weight $\bm{w}^{*}$ in large sample, as in
the following corollary.
\begin{cor}[Oracle equalities]
\label{cor:mse_wg}Under the conditions in Theorem \ref{thm:wg_convergence},
we have $R_{\mathcal{T}_{0}}(\hat{\bm{w}}_{(g)})=R_{\mathcal{T}_{0}}(\bm{w}^{*})+o_{p}(1)$.
In addition, if the conditions in \ref{thm:oracle_inequalities}\ref{enu:out-sample}
also hold, then $R_{\mathcal{T}_{1}}(\hat{\bm{w}}_{(g)})=R_{\mathcal{T}_{1}}(\bm{w}^{*})+o_{p}(1)$.
\end{cor}
Although \ref{cor:mse_wg} shows SCM-relaxation with the EL or entropy
objective function can also achieve the out-of-sample oracle performance
parallel to the $L_{2}$-norm objective as in \ref{thm:oracle_inequalities},
we recommend $L_{2}$-SCM-relaxation for practical use. The EL and
entropy cousins have the following drawbacks. First, to achieve the
same convergence rate, both EL and entropy rely on extra technical
conditions to ensure $\left\Vert \hat{\bm{w}}_{(g)}\right\Vert _{\infty}\leq\bar{c}K/J$.\footnote{This upper bound, determined by $K$ and $J$, is proven to hold for
the $L_{2}$-norm objective.} Second, EL and entropy implicitly require that the weights be strictly
positive, which incurs bias when some true group weights $\boldsymbol{w}^{*}$
are zero. This is particularly undesirable, as the true weight can
lie on the boundary of the simplex, which is stressed by \citet{abadieUsingSyntheticControls2021}.
\begin{rem}
In the large literature of the GEL \citep{neweyHigherOrderProperties2004},
researchers often view the choices of the divergence measures exchangeable---after
all, the empirical likelihood, exponential tilting \citep{schennachPointEstimationExponentially2007},
and continuous-updated GMM \citep{Hansen1996finite} are asymptotically
equivalent. This perspective is valid when the sample are independent
and identically distributed, in that each of the implied probability
weight tend to the equal weight asymptotically. In sharp contrast,
here we allocate weights on heterogeneous cross-sectional units. If
there are at least two groups, then the group oracle weight is in
general unequal, and it is possible that one of the groups has zero
weight. This highlights the key difference between the knowledge we
learn from the GEL literature and the current context of SCM where
we are faced with heterogeneous individual units.
\end{rem}
Before we conclude this section, we would like to mention that feature
engineering matters for all machine learning methods, including ours.
If the scales of the empirical data is heterogeneous across the individuals,
then a scale-standardization step can be desirable prior to feeding
the data into the the SCM-relaxation algorithm. All theoretical results
discussed above can be extended to the standardized version in a straightforward
fashion; see Appendix~\ref{sec:Standardized-SCM-relaxation} for
details.
\section{Monte Carlo Experiments }\label{sec:Monte-Carlo-experiments}
In this section, we conduct Monte Carlo simulations to study the finite
sample performance. We consider two sets of simulations: (i) the factor
loadings follow an exact group structure as (\ref{eq:grouped_structure})
specifies; and (ii) the factor loadings fluctuate around the group
means so that (\ref{eq:grouped_structure}) is violated. We estimate
the weights by the relaxation methods proposed in this paper, including
$L_{2}$-, EL-, and entropy-SCM-relaxation, and compare them with
SCM. We also compute the weights by the off-the-shelf Lasso and Ridge
methods:
\begin{align*}
\hat{\bm{w}}^{\mathrm{Lasso}} & =\arg\min_{\bm{w}\in\Delta_{J}}\left\{ T_{0}^{-1}\left\Vert \bm{y}_{0}-\bm{Y}\bm{w}\right\Vert _{2}^{2}+\lambda^{\mathrm{Lasso}}\|\bm{w}-J^{-1}\bm{1}_{J}\|_{1}\right\} ,\\
\hat{\bm{w}}^{\mathrm{Ridge}} & =\arg\min_{\bm{w}\in\Delta_{J}}\left\{ T_{0}^{-1}\left\Vert \bm{y}_{0}-\bm{Y}\bm{w}\right\Vert _{2}^{2}+\lambda^{\mathrm{Ridge}}\|\bm{w}-J^{-1}\bm{1}_{J}\|_{2}^{2}\right\} ,
\end{align*}
where $\lambda^{\mathrm{Lasso}}$ and $\lambda^{\mathrm{Ridge}}$
are tuning parameters, and the weights are shrunken toward the simple
average for a fair comparison. We further report the performance of
\citet{shi2023forward}'s forward-selected PDA (fsPDA). For every
method, the tuning parameters are selected via a two-fold cross-validation
(CV) if $T_{0}<50$, and four-fold CV otherwise. For relaxation methods,
the grid for $\eta$ spans $[0,\bar{\eta}]$, where the upper bound
$\bar{\eta}=\min_{\gamma\in\mathbb{R}}\Vert\hat{\bm{\Sigma}}\bm{1}_{J}/J-\hat{\bm{\Upsilon}}+\gamma\bm{1}_{J}\Vert_{\infty}$
so that $J^{-1}\boldsymbol{1}_{J}$ is a feasible solution.
In the two sets of experiments, $T_{1}$ is fixed at $50$, $J\in\left\{ 50,100,200\right\} $
and $T_{0}\in\{J/2,J,2J\}$. To account for diverging number of factors
and groups, we set $r=\lfloor\log T_{0}\rfloor$ and $K\in\{\lfloor0.8r\rfloor,r,\lfloor1.2r\rfloor+1\}$
which corresponds to $K<r$, $K=r$, and $K>r$, respectively. The
experiment is repeated 1000 times in each DGP. We calculate the oracle
weights $\bm{w}^{*}$ following (\ref{eq:oracle_relaxation_problem})
(or (\ref{eq:g-oracle}) for a generic objective function). Define
the oracle synthetic control as $y_{0t}^{N,*}:=\sum_{j=1}^{J}w_{j}^{*}y_{jt}^{N}$.
For a synthetic control estimator $\hat{\bm{w}}$ that yields the
outcome estimate $\widehat{y}_{0t}^{N}$, the post-treatment prediction
error is $\sum_{t\in\mathcal{T}_{1}}(\widehat{y}_{0t}^{N}-y_{0t}^{N,*})^{2}$.
For ease of comparison, we treat SCM's $\hat{\bm{w}}^{\mathrm{SC}}$
as the benchmark; with the corresponding estimated counterfactual
$\widehat{y}_{0t}^{N,\mathrm{SC}}$, we assess the out-of-sample performance
of each estimator by the ratio
\[
\sum_{t\in\mathcal{T}_{1}}\left(\widehat{y}_{0t}^{N}-y_{0t}^{N,*}\right){}^{2}\bigg/\sum_{t\in\mathcal{T}_{1}}\left(\widehat{y}_{0t}^{N,\mathrm{SC}}-y_{0t}^{N,*}\right){}^{2}.
\]
Moreover, we present the the $L_{1}$- and $L_{2}$-distance ratios,
$\left\Vert \hat{\bm{w}}-\bm{w}^{*}\right\Vert _{1}/\Vert\hat{\bm{w}}^{\mathrm{SC}}-\bm{w}^{*}\Vert_{1}$
and $\left\Vert \hat{\bm{w}}-\bm{w}^{*}\right\Vert _{2}/\Vert\hat{\bm{w}}^{\mathrm{SC}}-\bm{w}^{*}\Vert_{2}$,
to check the quality of the weight estimation.
\subsection{Exact Group Structures}
The simulated data consist of $J+1$ units and $T_{0}+T_{1}$ time
periods. The treatment to the unit $j=0$ occurs immediately after
time $T_{0}$ and sustains from $T_{0}+1$ on. The potential outcome,
if untreated, follows a factor model, as (\ref{eq:model}):
\begin{align*}
y_{jt}^{N} & =\bm{\lambda}_{j}'\bm{f}_{t}+u_{jt}\quad\text{for \ensuremath{j\in\{0\}\cup[J]} and \ensuremath{t\in[T_{0}+T_{1}]}}.
\end{align*}
The idiosyncratic errors $u_{jt}\stackrel{\text{i.i.d.}}{\sim}\mathcal{N}(0,1)$
and are independent of the common factors. The $r$ factors $f_{\ell t}$
($\ell\in[r]$) are mutually independent, and each follows an AR(1)
process:
\[
f_{\ell t}=0.5f_{\ell t-1}+u_{\ell t}^{f},\qquad\ell\in[r],\ t\in[T_{0}+T_{1}],
\]
where $u_{\ell t}^{f}\stackrel{\text{i.i.d.}}{\sim}\mathcal{N}(0,1)$.
Each entry in the core factor loadings $\bm{\Lambda}^{\mathrm{co}}$
is independently drawn from $\mathcal{N}(0,3/r)$. The loadings for
the control units are given by $\bm{\Lambda}=\bm{Z}\bm{\Lambda}^{\mathrm{co}}$.
The loading for the treated united is generated as $\bm{\lambda}_{0}=\bm{\Lambda}^{\mathrm{co}\prime}\bm{w}_{\mathcal{G}}^{*}+\bm{\varepsilon}$,
where $\bm{\varepsilon}$ have entries drawn independently from $\text{Uniform}(-0.1/\sqrt{r},0.1/\sqrt{r})$
and $\bm{w}_{\mathcal{G}}^{*}=(w_{\mathcal{G}_{1}}^{*},\dots,w_{\mathcal{G}_{K}}^{*})^{\prime}$
has its first entry set to be zero and the other $K-1$ entries jointly
generated from a Dirichlet distribution so that $\bm{w}_{\mathcal{G}}^{*}$
is assured to live in the simplex. The noise $\bm{\varepsilon}$ implies
that $\bm{\lambda}_{0}$ may not be perfectly recovered as a convex
combination of the core loadings. The loadings and weights, once drawn,
are fixed over the replications.
\begin{table}[t]
\centering
\caption{Average post-treatment prediction error ratio under exact group structures
}\label{tab:Mean-MSE-ratios-DGP1}
\medskip{}
{\footnotesize\begin{tabular*}{\linewidth}{@{\enspace\extracolsep{\fill}}cccccccc@{\enspace}}
\toprule
& & & & & \multicolumn{3}{c}{Relaxation} \\
\cmidrule{6-8}
{$J$} & {$T_0$} & {Lasso} & {Ridge} & {fsPDA} & {$L_2$} & {EL} & {Entropy} \\
\midrule
\multicolumn{8}{c}{Panel A: $K<r$} \\
50 & 25 & 0.8136 & 0.5035 & 3.7740 & \textbf{0.3019} & 0.8441 & 0.3349 \\
50 & 50 & 0.7324 & 0.4503 & 3.0109 & \textbf{0.1657} & 0.6993 & 0.1857 \\
50 & 100 & 0.6152 & 0.4225 & 3.0252 & \textbf{0.2451} & 0.4784 & 0.2453 \\
100 & 50 & 0.7116 & 0.4096 & 3.4506 & \textbf{0.1509} & 0.7274 & 0.1686 \\
100 & 100 & 0.6263 & 0.4218 & 3.2039 & \textbf{0.2232} & 0.4908 & 0.2420 \\
100 & 200 & 0.6947 & 0.4167 & 2.8071 & \textbf{0.2503} & 0.5395 & 0.2812 \\
200 & 100 & 0.6133 & 0.3592 & 3.5371 & \textbf{0.2145} & 0.5266 & 0.2397 \\
200 & 200 & 0.6930 & 0.4134 & 3.0715 & \textbf{0.2535} & 0.5469 & 0.2902 \\
200 & 400 & 0.5963 & 0.3376 & 3.3766 & \textbf{0.1847} & 0.4590 & 0.2161 \\
\midrule
\multicolumn{8}{c}{Panel B: $K=r$} \\
50 & 25 & 0.8044 & 0.6530 & 2.8289 & \textbf{0.5290} & 0.8095 & 0.5895 \\
50 & 50 & 0.7025 & 0.5374 & 1.9234 & \textbf{0.3206} & 0.5601 & 0.3401 \\
50 & 100 & 0.6845 & 0.5524 & 2.7681 & \textbf{0.4717} & 0.7560 & 0.5148 \\
100 & 50 & 0.6915 & 0.4891 & 2.3996 & \textbf{0.2973} & 0.5707 & 0.3307 \\
100 & 100 & 0.6928 & 0.5437 & 2.9114 & \textbf{0.4626} & 0.7735 & 0.5072 \\
100 & 200 & 0.7014 & 0.5372 & 2.8514 & \textbf{0.3423} & 0.4892 & 0.3518 \\
200 & 100 & 0.6771 & 0.5039 & 3.3672 & \textbf{0.4202} & 0.7268 & 0.4599 \\
200 & 200 & 0.6953 & 0.5086 & 3.0984 & \textbf{0.3296} & 0.4972 & 0.3350 \\
200 & 400 & 0.6595 & 0.4477 & 3.3109 & \textbf{0.2623} & 0.4102 & 0.2676 \\
\midrule
\multicolumn{8}{c}{Panel C: $K>r$} \\
50 & 25 & 0.8214 & 0.6131 & 2.7519 & 0.5075 & 0.5937 & \textbf{0.4890} \\
50 & 50 & 0.7275 & 0.5341 & 1.8694 & 0.3418 & 0.4638 & \textbf{0.3245} \\
50 & 100 & 0.7461 & 0.5661 & 2.4932 & \textbf{0.4158} & 0.5045 & 0.4267 \\
100 & 50 & 0.7409 & 0.5038 & 2.1725 & 0.3695 & 0.4846 & \textbf{0.3327} \\
100 & 100 & 0.7196 & 0.5501 & 2.7311 & \textbf{0.3639} & 0.4577 & 0.3728 \\
100 & 200 & 0.7028 & 0.4758 & 2.0562 & 0.4254 & \textbf{0.3827} & 0.3965 \\
200 & 100 & 0.7111 & 0.5036 & 3.0201 & \textbf{0.3705} & 0.4585 & 0.3753 \\
200 & 200 & 0.7034 & 0.4518 & 2.2439 & 0.4167 & \textbf{0.3687} & 0.3915 \\
200 & 400 & 0.6119 & 0.3936 & 2.4515 & \textbf{0.3177} & 0.3853 & 0.3321 \\
\bottomrule
\end{tabular*}
}{\footnotesize\par}
\end{table}
Table~\ref{tab:Mean-MSE-ratios-DGP1} reports the average (over the
replications) post-treatment prediction error ratio. It reveals that
all methods except fsPDA outperform the canonical SCM for all combinations
of $K$ and $r$. The unfavorable performance of fsPDA is anticipated,
as the oracle weights are dense and constrained on the simplex while
the estimated weights of fsPDA do not lie on the simplex, and are
typically sparse. Lasso and Ridge surpass SCM substantially by penalizing
solutions toward the simple average. Moreover, Ridge also encourages
dense solutions, which gains an edge over Lasso. Leveraging latent
group structures and dense weights, the SCM-relaxation methods further
improve on Ridge in all settings. Among the relaxation approaches,
the $L_{2}$norm objective excels when $K\leq r$, due to its advantage
of allowing for exact zero weights. Furthermore, entropy is uniformly
better than EL in estimating weights close to zero, in view of the
fact that $\lim_{x\to0^{+}}x\log x=0$ while $\lim_{x\to0^{+}}\log x=-\infty$.
When $K>r$, the oracle weights vary with the choice of objective
function. As a result, weight estimates by $L_{2}$, EL, and entropy
can have different oracle targets. Overall, $L_{2}$-SCM-relaxation
stands out as the best performer and is recommended in practice.
\begin{sidewaystable}
\centering
\caption{Average distance ratios }\label{tab:Mean-distance-ratios-DGP1}
\medskip{}
\resizebox{0.95\linewidth}{!}{{\footnotesize\begin{tabular*}{\linewidth}{@{\enspace\extracolsep{\fill}}cccccccccccc@{\enspace}}
\toprule
& & \multicolumn{5}{c}{$L_1$-distance} & \multicolumn{5}{c}{$L_2$-distance} \\
\cmidrule{3-7}\cmidrule{8-12}
{$J$} & {$T_0$} & {Lasso} & {Ridge} & {$L_2$} & {EL} & {Entropy} & {Lasso} & {Ridge} & {$L_2$} & {EL} & {Entropy} \\
\midrule
\multicolumn{12}{c}{Panel A: $K<r$} \\
50 & 25 & 0.689 & 0.458 & \textbf{0.144} & 0.474 & 0.175 & 0.742 & 0.413 & \textbf{0.115} & 0.535 & 0.148 \\
50 & 50 & 0.695 & 0.432 & \textbf{0.120} & 0.467 & 0.148 & 0.748 & 0.408 & \textbf{0.108} & 0.549 & 0.138 \\
50 & 100 & 0.535 & 0.378 & \textbf{0.101} & 0.306 & 0.113 & 0.640 & 0.356 & \textbf{0.090} & 0.376 & 0.108 \\
100 & 50 & 0.663 & 0.442 & \textbf{0.141} & 0.431 & 0.155 & 0.741 & 0.384 & \textbf{0.110} & 0.559 & 0.132 \\
100 & 100 & 0.534 & 0.396 & \textbf{0.098} & 0.280 & 0.111 & 0.650 & 0.348 & \textbf{0.077} & 0.367 & 0.099 \\
100 & 200 & 0.628 & 0.441 & \textbf{0.108} & 0.263 & 0.128 & 0.700 & 0.375 & \textbf{0.087} & 0.318 & 0.114 \\
200 & 100 & 0.514 & 0.422 & \textbf{0.104} & 0.263 & 0.109 & 0.643 & 0.318 & \textbf{0.067} & 0.371 & 0.084 \\
200 & 200 & 0.597 & 0.422 & \textbf{0.113} & 0.234 & 0.126 & 0.684 & 0.324 & \textbf{0.078} & 0.292 & 0.104 \\
200 & 400 & 0.588 & 0.395 & \textbf{0.088} & 0.211 & 0.102 & 0.677 & 0.311 & \textbf{0.067} & 0.266 & 0.090 \\
\midrule
\multicolumn{12}{c}{Panel B: $K=r$} \\
50 & 25 & 0.537 & 0.453 & \textbf{0.177} & 0.352 & 0.199 & 0.625 & 0.396 & \textbf{0.140} & 0.369 & 0.167 \\
50 & 50 & 0.525 & 0.433 & \textbf{0.140} & 0.339 & 0.159 & 0.633 & 0.399 & \textbf{0.120} & 0.393 & 0.147 \\
50 & 100 & 0.605 & 0.508 & \textbf{0.360} & 0.516 & 0.385 & 0.637 & 0.458 & \textbf{0.300} & 0.467 & 0.331 \\
100 & 50 & 0.519 & 0.461 & \textbf{0.157} & 0.337 & 0.174 & 0.646 & 0.391 & \textbf{0.115} & 0.407 & 0.146 \\
100 & 100 & 0.604 & 0.525 & \textbf{0.334} & 0.474 & 0.350 & 0.639 & 0.432 & \textbf{0.242} & 0.404 & 0.273 \\
100 & 200 & 0.619 & 0.490 & \textbf{0.238} & 0.333 & 0.245 & 0.644 & 0.424 & \textbf{0.186} & 0.352 & 0.206 \\
200 & 100 & 0.577 & 0.536 & \textbf{0.307} & 0.427 & 0.311 & 0.622 & 0.397 & \textbf{0.185} & 0.353 & 0.219 \\
200 & 200 & 0.599 & 0.502 & 0.221 & 0.307 & \textbf{0.219} & 0.647 & 0.392 & \textbf{0.150} & 0.339 & 0.170 \\
200 & 400 & 0.567 & 0.467 & \textbf{0.178} & 0.267 & 0.183 & 0.615 & 0.375 & \textbf{0.132} & 0.306 & 0.151 \\
\midrule
\multicolumn{12}{c}{Panel C: $K>r$} \\
50 & 25 & 0.624 & 0.457 & 0.191 & 0.381 & \textbf{0.185} & 0.681 & 0.395 & \textbf{0.140} & 0.365 & 0.150 \\
50 & 50 & 0.630 & 0.436 & 0.169 & 0.376 & \textbf{0.152} & 0.696 & 0.393 & 0.132 & 0.362 & \textbf{0.131} \\
50 & 100 & 0.565 & 0.457 & \textbf{0.165} & 0.261 & 0.171 & 0.615 & 0.417 & \textbf{0.147} & 0.278 & 0.164 \\
100 & 50 & 0.627 & 0.459 & 0.200 & 0.370 & \textbf{0.165} & 0.712 & 0.385 & 0.133 & 0.361 & \textbf{0.128} \\
100 & 100 & 0.542 & 0.478 & 0.168 & 0.233 & \textbf{0.162} & 0.608 & 0.404 & \textbf{0.130} & 0.248 & 0.145 \\
100 & 200 & 0.638 & 0.435 & 0.178 & 0.235 & \textbf{0.174} & 0.692 & 0.362 & 0.130 & 0.181 & \textbf{0.128} \\
200 & 100 & 0.514 & 0.497 & 0.194 & 0.233 & \textbf{0.174} & 0.591 & 0.369 & \textbf{0.121} & 0.248 & 0.138 \\
200 & 200 & 0.613 & 0.415 & 0.152 & 0.203 & \textbf{0.149} & 0.679 & 0.312 & \textbf{0.094} & 0.142 & 0.094 \\
200 & 400 & 0.603 & 0.398 & \textbf{0.119} & 0.198 & 0.131 & 0.672 & 0.308 & \textbf{0.080} & 0.140 & 0.087 \\
\bottomrule
\end{tabular*}
}}\vspace{.5\baselineskip}
\begin{minipage}[t]{0.95\linewidth}
{\footnotesize Note: We omit the weight estimates by fsPDA as they
are not restricted to the simplex and thus far away from the oracle
as we define.}
\end{minipage}
\end{sidewaystable}
\begin{figure}
\centering
\includegraphics[width=1\textwidth]{simu_DGP1_weight_comparisons}
\caption{Estimated weight vectors in a typical simulation }\label{fig:weight-comparison-simu}
\end{figure}
The performance of the estimated weights is displayed in \ref{Tab:Mean-distance-ratios-DGP1}.
Lasso and Ridge are better in recovering the oracle weights than SCM,
and Ridge outperforms Lasso in terms of $L_{2}$-distance as it tends
to produce weights with a smaller $L_{2}$-norm. Among SCM-relaxation
methods, $L_{2}$-SCM-relaxation consistently excels in estimating
the weights for $K\leq r$. Again, because of potentially different
oracle targets when $K>r$, the $L_{2}$-objective is not guaranteed
to yield the best performance compared with EL and entropy. For evaluation
in terms of $L_{1}$-distance, the three relaxation methods perform
comparably well. These findings align with our theory.
To see how well these methods recover the group structure of oracle
weights, \ref{Fig:weight-comparison-simu} illustrates the estimated
weights in a typical simulation with 3 groups, 3 factors, 200 control
units and 100 pre-treatment periods. The oracle weights (the black
dash line) are clustered by groups for visualization. As is clear
in the figure, SCM produces sparse weights, concentrating on groups
with nonzero oracle weights but deviating far from the within-group
equal oracle weight. Lasso tends to collapse its weights to the simple
average, the target of shrinkage. Ridge, with the denser solution,
better approximates the group structure. As predicted by our theory,
$L_{2}$-SCM-relaxation nearly perfectly traces the group pattern
with the smallest estimation errors.
While EL and entropy objectives partially recover the group patterns,
they exhibit large fluctuations in estimating weights, particularly
for groups with higher oracle weights. This variability stems from
the nonuniform curvature of their objective functions’ penalty terms,
as measured by second derivatives \citep[Section 9.3]{wainwright2019high}.
For example, the second derivative of the entropy function $x\log x$
is $1/x$, which decreases from $+\infty$ to 100 as the oracle weight
increases from 0 to 0.01. Consequently, the entropy function imposes
disproportionately heavy penalties on groups with near-zero oracle
weights compared to those with higher weights. This is why, in the
last subgraph of \ref{Fig:weight-comparison-simu}, the volatility
of weights keeps growing from the first group to the third group.
The same reasoning applies to the ``EL'' subgraph. In contrast,
the constant curvature of the $L_{2}$ penalty ensures uniform penalization
across all groups.
\subsection{Approximate Group Structures}
\begin{table}[t]
\centering
\caption{Average post-treatment prediction error ratio under approximate group
structures }\label{tab:Mean-MSE-ratios-DGP3}
\medskip{}
{\footnotesize\begin{tabular*}{\linewidth}{@{\enspace\extracolsep{\fill}}cccccccc@{\enspace}}
\toprule
& & & & & \multicolumn{3}{c}{Relaxation} \\
\cmidrule{6-8}
{$J$} & {$T_0$} & {Lasso} & {Ridge} & {fsPDA} & {$L_2$} & {EL} & {Entropy} \\
\midrule
\multicolumn{8}{c}{Panel A: $K<r$} \\
50 & 25 & 0.8080 & 0.5189 & 3.8169 & \textbf{0.3191} & 0.8353 & 0.3394 \\
50 & 50 & 0.7233 & 0.4510 & 2.9675 & \textbf{0.1705} & 0.6970 & 0.1897 \\
50 & 100 & 0.6225 & 0.4148 & 2.9786 & \textbf{0.2332} & 0.4965 & 0.2438 \\
100 & 50 & 0.7353 & 0.4404 & 3.3038 & \textbf{0.1714} & 0.7353 & 0.1882 \\
100 & 100 & 0.6225 & 0.4018 & 3.1202 & \textbf{0.2293} & 0.4905 & 0.2387 \\
100 & 200 & 0.6970 & 0.4219 & 2.6966 & \textbf{0.2516} & 0.5709 & 0.2901 \\
200 & 100 & 0.6302 & 0.3836 & 3.4238 & \textbf{0.2012} & 0.5052 & 0.2242 \\
200 & 200 & 0.6863 & 0.3951 & 2.9600 & \textbf{0.2349} & 0.5645 & 0.2744 \\
200 & 400 & 0.6179 & 0.3476 & 3.3050 & \textbf{0.1913} & 0.5156 & 0.2221 \\
\midrule
\multicolumn{8}{c}{Panel B: $K=r$} \\
50 & 25 & 0.8051 & 0.6582 & 2.9176 & \textbf{0.5732} & 0.8293 & 0.6153 \\
50 & 50 & 0.6970 & 0.5387 & 1.8485 & \textbf{0.3217} & 0.5600 & 0.3548 \\
50 & 100 & 0.6865 & 0.5628 & 2.7738 & \textbf{0.4809} & 0.7663 & 0.5196 \\
100 & 50 & 0.7074 & 0.4966 & 2.1994 & \textbf{0.3078} & 0.5953 & 0.3491 \\
100 & 100 & 0.6942 & 0.5533 & 2.9537 & \textbf{0.4470} & 0.7318 & 0.4883 \\
100 & 200 & 0.7118 & 0.5358 & 2.8092 & \textbf{0.3547} & 0.5148 & 0.3640 \\
200 & 100 & 0.6785 & 0.5070 & 3.3877 & \textbf{0.4170} & 0.6874 & 0.4523 \\
200 & 200 & 0.6955 & 0.4961 & 3.0983 & \textbf{0.3293} & 0.5019 & 0.3348 \\
200 & 400 & 0.6462 & 0.4470 & 3.1869 & \textbf{0.2805} & 0.4690 & 0.2859 \\
\midrule
\multicolumn{8}{c}{Panel C: $K>r$} \\
50 & 25 & 0.8401 & 0.6034 & 2.7073 & 0.5243 & 0.6166 & \textbf{0.5087} \\
50 & 50 & 0.7460 & 0.5477 & 1.9607 & 0.3706 & 0.5034 & \textbf{0.3526} \\
50 & 100 & 0.7559 & 0.5655 & 2.5131 & \textbf{0.4087} & 0.5033 & 0.4149 \\
100 & 50 & 0.7279 & 0.5054 & 2.0952 & 0.3666 & 0.4928 & \textbf{0.3334} \\
100 & 100 & 0.7104 & 0.5433 & 2.6731 & \textbf{0.3832} & 0.4650 & 0.3861 \\
100 & 200 & 0.6998 & 0.4891 & 2.0388 & 0.4428 & \textbf{0.3802} & 0.4010 \\
200 & 100 & 0.7187 & 0.5274 & 2.8991 & \textbf{0.3680} & 0.4624 & 0.3792 \\
200 & 200 & 0.6970 & 0.4502 & 2.2148 & 0.4254 & \textbf{0.3795} & 0.3997 \\
200 & 400 & 0.6273 & 0.3941 & 2.3705 & \textbf{0.3246} & 0.3903 & 0.3412 \\
\bottomrule
\end{tabular*}
}{\footnotesize\par}
\end{table}
In this section, we consider an approximate group design to check
the robustness of the estimation methods when the exact group structures
do not hold. All the settings are the same as those in the previous
subsection except that the loadings are generated as $\bm{\Lambda}=\bm{Z}\bm{\Lambda}^{\mathrm{co}}+\bm{\Xi}$
where $\bm{\Xi}$ consists of entries drawn independently from $\text{Uniform}(-0.2/\sqrt{r},0.2/\sqrt{r})$
and is independent of other random variables. \ref{Tab:Mean-MSE-ratios-DGP3}
presents the average post-treatment prediction error ratios. The patterns
are similar to those in \ref{Tab:Mean-MSE-ratios-DGP1} in exact group
settings: the relaxation methods outperform substantially SCM as well
as Lasso and Ridge; moreover, $L_{2}$SCM-relaxation is the best when
$K\leq r$.
The Monte Carlo simulation evidence shows that the relaxation methods
are robust against mild deviation from group structure, manifesting
their practical usefulness.
\section{Empirical Application }\label{sec:Empirical-applications}
In this section, we examine the impact of the 2016 Brexit referendum
on the real GDP of the United Kingdom (UK). On June 23, 2016, a referendum
on the UK's membership in the European Union (EU) resulted in 51.89\%
of voters favoring departure, triggering a withdrawal process that
would conclude on January 31, 2020. The referendum's outcome was largely
unexpected by markets. We designate the third quarter of 2016 as the
starting point of the treatment.
\begin{figure}
\centering
\includegraphics[width=1\textwidth]{GDP_Growth_time_series_UK}
\caption{Real GDP and growth rates of the UK}\label{fig:UK_GDP_GR}
\end{figure}
We collect quarterly real GDP data for all available economies from
the CEIC database, a subsidiary of Caixin providing business information.
The completeness of the quarterly GDP time series varies across the
countries and regions, with a fraction of them having records only
in recent years. We shape the available data into a balanced panel,
with a comprehensive donor pool of 57 countries from 2002Q4 to 2024Q2.
To remove time trends and seasonality, we construct the year-over-year
(YoY) GDP growth rate $\tilde{y}_{jt}=y_{jt}/y_{j(t-4)}-1$ for all
economies. Figure \ref{fig:UK_GDP_GR} plots UK's quarterly real GDP
(blue line) and its YoY growth rate (red line).
\begin{figure}
\centering
\begin{raggedright}
\includegraphics[width=1\textwidth]{growth_fit_diff_UK_time_series_2016}
\par\end{raggedright}
\begin{raggedright}
\footnotesize{Note: ER is the in-sample empirical risk.}
\par\end{raggedright}
\caption{Gap between realized growth rate and the fitted value (Treatment time
2016Q3) }\label{fig:Fitted-GDP-growth}
\end{figure}
\begin{figure}[ph]
\centering
\includegraphics[width=0.9\textwidth]{weight_comparison_L2_SCM_histogram_2016}
\caption{$L_{2}$-SCM-relaxation weights versus SCM weights (Treatment time
2016Q3) }\label{fig:weight_L2_SCM_2016}
\end{figure}
We fit and predict the GDP growth rate by SCM and the three relaxation
methods. Shown in Figure~\ref{fig:Fitted-GDP-growth}, all methods
exhibit similar patterns with pre-treatment fit. SCM has the smallest
in-sample empirical risk, which is exactly the objective function
it minimizes; it leads to an aggressive post-treatment extrapolation
that visibly deviates from the realization. The relaxation approaches
are more conservative in both the in-sample fit and the out-of-sample
prediction.
Given that the counterfactuals are constructed by weighting the control
units, we report in Figure \ref{fig:weight_L2_SCM_2016} the weights
from SCM and those of $L_{2}$-SCM-relaxation.\footnote{The weights from the EL and entropy objectives are similar to those
of $L_{2}$ and thus omitted.} The country/region names are sorted by the size of the $L_{2}$ weights,
according to which we \emph{manually} classified into Group 1 ($\geq0.04$,
dark blue), Group 2 ($[0.01,0.04)$, light blue), Group 3 ($(0,0.01)$,
red), and Group 4 (exactly zero, brown). Group 1 includes major EU
countries Italy, Germany, Denmark, and France, as well as big English-speaking
countries Canada and the United States. They are important economic
partners of the UK. Group 2 further covers several middle-sized EU
economies. Those in Group 3 and 4 are mostly geographically distant
from the UK. In contrast, the sparse weights produced by SCM are less
interpretable. While USA alone takes 30\% of the weight, the estimate
excludes Germany, France, and Italy, the top 3 EU countries by GDP,
but instead places non-trivial weights on Lithuania and Luxembourg,
countries that play a minor economic role.
\begin{figure}
\centering
\begin{centering}
\includegraphics[width=1\textwidth]{GDP_extrapolation_UK_2016}
\par\end{centering}
\caption{Real GDP versus counterfactual GDP (Treatment time 2016Q3) }\label{fig:(Log)-Difference}
\end{figure}
To assess the GDP loss in levels, we use the estimated GDP growth
rate counterfactuals to reconstruct the counterfactual real GDP. Let
$\widehat{\tilde{y}}_{0t}^{N}=\sum_{j\in[J]}\hat{w}_{j}\tilde{y}_{jt}$
in the post-treatment period, and we compute the level GDP as
\[
\hat{y}_{0t}^{N}=(1+\widehat{\tilde{y}}_{0t}^{N})z_{0(t-4)},\quad t\in\mathcal{T}_{1},
\]
where the base $z_{0t}=y_{0t}$ if $t\leq T_{0}$ and $z_{0t}=\hat{y}_{0t}^{N}$
for $t>T_{0}$.
We calculate $\sum_{t\in\mathcal{T}_{1}}(y_{0t}^{I}-\hat{y}_{0t}^{N})/\sum_{t\in\mathcal{T}_{1}}y_{0t}^{I}$,
the treatment effects relative to the UK's total GDP after the treatment.
SCM predicts a substantial loss of $4.93\%$ of the total realized
GDP. $L_{2}$-, EL- and entropy-SCM-relaxation estimate $3.85\%$,
$3.64\%$ and $3.90\%$, respectively. All numbers suggest that Brexit
has shrunken the UK's economy. In Figure~\ref{fig:(Log)-Difference},
the gap between the counterfactual and the realization is visible
from 2016Q3 to 2020Q1, and the negative impact quickly worsens after
2020Q1. This underscores the delayed but substantial economic consequences
of Brexit, which were not immediately evident in the aftermath of
the referendum but emerged clearly following the formal exit.\footnote{The first quarter of the official departure coincided with the breakout
of Covid-19, which triggered a worldwide economic recession. The pandemic
was a global event that affected all countries. Given that the UK's
public health system was relatively well developed compared to other
countries, the GDP gap could be even larger in absence of the pandemic.} The substantial effect after official departure is huge; if we compute
the corresponding ratio after 2020Q1, SCM yields a shocking loss of
$7.74\%$, whereas the relaxation method outcomes range from $5.58\%$
to $5.98$\%.
\begin{figure}
\centering
\begin{raggedright}
\includegraphics[width=1\textwidth]{growth_fit_diff_UK_time_series_2020}
\par\end{raggedright}
\begin{raggedright}
\footnotesize{Note: ER is the in-sample empirical risk.}
\par\end{raggedright}
\caption{Gap between realized growth rate and the fitted value (Treatment time
2020Q1)}\label{fig:Fitted-GDP-growth-1}
\end{figure}
\begin{figure}[ph]
\centering
\includegraphics[width=0.9\textwidth]{weight_comparison_L2_SCM_histogram_2020}
\caption{$L_{2}$-SCM-relaxation weights versus SCM weights (Treatment time
2020Q1) }\label{fig:weight_L2_SCM_2020}
\end{figure}
\begin{figure}
\centering
\includegraphics[width=1\textwidth]{GDP_extrapolation_UK_2020}
\caption{Real GDP versus counterfactual GDP (Treatment time 2020Q1) }\label{fig:Difference-between-real}
\end{figure}
While the outcome of referendum was largely unexpected, the economic
decoupling has been working in progress after 2016, with 2020 in anticipation.
As a robustness check, we carry out re-estimation by setting the treatment
starting point at 2020Q1. In Figure \ref{fig:Fitted-GDP-growth-1},
the in-sample empirical risks are similar to those of Figure \ref{fig:Fitted-GDP-growth},
whereas the variation of SCM weights in Figure \ref{fig:weight_L2_SCM_2020}
is substantial relative to those in Figure \ref{fig:weight_L2_SCM_2016}.
For the loss of total GDP, SCM estimates $5.22\%$ of the UK's total
GDP after 2020Q1, while the $L_{2}$, EL and entropy estimate $4.42\%$,
$3.61\%$ and $4.02\%$, respectively. These enlarged effects are
due to the drop of the ``mild loss'' period between 2016 and 2020
as shown in Figure \ref{fig:Difference-between-real}. They are smaller
than the numbers for the same time period reported at the end of the
last paragraph. This is due to the fact that this exercise uses the
2016--2020 data for in-sample fit, where the downward anticipation
is leaked into the pre-treatment period and contaminates the weight
estimation. Therefore we view that the results are cleaner with 2016Q1
as the treatment point. Overall, regardless of the choices of treatment
point, the departure of the UK has inflicted a profound economic loss.
\section{Conclusion}
We propose a machine learning algorithm, termed SCM-relaxation, to
estimate the weights for counterfactual prediction of a treated unit
at the presence of many potential control units. Given a factor model
with the latent group structures, $L_{2}$-SCM-relaxation produces
weights that converge asymptotically to the oracle weights. The convergence
rate is sufficiently fast to ensure oracle prediction accuracy as
if the group membership were known. If one chooses the EL or the entropy
objective function to conduct the relaxation, under extra regularization
conditions they possess similar properties as the $L_{2}$ counterpart
if the true weight is non-zero. Our theoretical study suggests that
$L_{2}$-SCM-relaxation enjoys desirable properties under fewer conditions
and is thus preferred. We apply them to evaluate the impact of Brexit
on the UK's real GDP, and find that the estimated counterfactual without
Brexit would be substantially higher than the realized values.
\bigskip \bigskip
\bibliographystyle{ecta}
\bibliography{emplikeli,ref_books,scm}
\newpage\setcounter{equation}{0}
\setcounter{thm}{0}
\setcounter{lem}{0}
\setcounter{page}{1}
\begin{center}
{\LARGE\textbf{Online Appendices}}{\LARGE\par}
\par\end{center}