EconBase
← Back to paper

Same Root Different Leaves: Time Series and Cross-Sectional Methods in Panel Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

165,190 characters · 35 sections · 69 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Same Root Different Leaves: Time Series and Cross-Sectional Methods in Panel Data

frontmatter\runtitle{ Time Series and Cross-Sectional Methods in Panel Data } \begin{aug} \address[id=add1]{ \orgdiv{Simons Institute for the Theory of Computing}, \orgname{University of California, Berkeley}} \address[id=add2]{ \orgdiv{Department of Statistics}, \orgname{University of California, Berkeley}} \address[id=add3]{ \orgdiv{Departments of Statistics & Data Science and Political Science}, \orgname{Yale University}} \address[id=add4]{ \orgdiv{Departments of Statistics and EECS}, \orgname{University of California, Berkeley}} \end{aug} \support{ We sincerely thank Alberto Abadie, Avi Feller, Guido Imbens, and Devavrat Shah for their thoughtful comments and insightful feedback. We gratefully acknowledge support from NSF grants 1945136, 1953191, 2022448, 2023505 on Collaborative Research: Foundations of Data Science Institute (FODSI), and ONR grant N00014-17-1-2176. The data and code to reproduce the results in this article are available at \href{https://github.com/deshen24/panel-data-regressions}{https://github.com/deshen24/panel-data-regressions}.} \begin{abstract} A central goal in social science is to evaluate the causal effect of a policy. One dominant approach is through panel data analysis in which the behaviors of multiple units are observed over time. The information across time and space motivates two general approaches: (i) horizontal regression (i.e., unconfoundedness), which exploits time series patterns, and (ii) vertical regression (e.g., synthetic controls), which exploits cross-sectional patterns. Conventional wisdom states that the two approaches are fundamentally different. We establish this position to be partly false for estimation but generally true for inference. In particular, we prove that both approaches yield identical point estimates under several standard settings. For the same point estimate, however, each approach quantifies uncertainty with respect to a distinct estimand. The confidence interval developed for one estimand may have incorrect coverage for another. This emphasizes that the source of randomness that researchers assume has direct implications for the accuracy of inference. \end{abstract} \begin{keyword} \kwd{horizontal regression} \kwd{vertical regression} \kwd{unconfoundedness} \kwd{synthetic controls} \kwd{causal inference} \kwd{minimum norm estimators} \end{keyword}

Introduction

In a seminal paper, abadie1 set out to investigate the economic impact of terrorism in Basque Country. Prior to the outset of terrorist activity in the early 1970's, Basque Country was considered to be one of the wealthiest regions in Spain. After thirty years of turmoil, however, its economic activity dropped substantially relative to its neighboring regions. Although intuition affirms that Basque Country's economic downturn can be attributed, at least partially, to its political and civil unrest, it is difficult to quantitatively isolate the economic costs of conflict. In response to this challenge, abadie1 introduced the synthetic controls framework. At its core, synthetic controls constructs a synthetic Basque Country from a weighted composition of control regions that are largely unaffected by the instability to estimate Basque Country's economic evolution in the absence of terrorism. This novel concept has inspired an entire subliterature within econometrics that is “arguably the most important innovation in the policy evaluation literature in the last 15 years” athey_imbens.

Researchers have historically tackled problems of this flavor using repeated observations of units across time, i.e., panel data, where a subset of units are exposed to a treatment during some time periods while the other units are unaffected. In the study above, the per capita gross domestic product (GDP) of $17$ Spanish regions are measured from 1955--1998. Basque Country is the sole treated unit and the remaining regions are the control units; the pre- and post-treatment periods are defined as the time horizons before and after the first wave of terrorist activity, respectively.

Synthetic controls has become a cornerstone for panel studies in recent years and across numerous fields. Beforehand, the unconfoundedness approach rubin_rosenbaum, imbens_wooldridge served as a common workhorse. Whereas synthetic controls posits a relation between treated and control units that is stable across time, unconfoundedness posits a relation between treated and pretreatment periods that is stable across units. Accordingly, synthetic controls exploits cross-sectional correlation patterns while unconfoundedness exploits time series correlation patterns. Considering the panel data format, unconfoundedness and synthetic controls based methods are commonly referred to as horizontal (HZ) and vertical (VT) regressions, respectively. Given their conceptual and computational distinctions, the two approaches are considered to be fundamentally different mc_panel.

Yet, contrary to conventional wisdom, it turns out that HZ and VT regressions can yield identical point estimates. As Figure (ref) shows, when the regression models are learned via ordinary least squares (OLS) or principal component regression (PCR), then the two approaches produce the same economic evolution for Basque Country in the absence of terrorism. Figure (ref), by contrast, shows that when the regression models are learned via lasso or lie within the simplex---as proposed by abadie1 for VT regression---then the two approaches output contrasting economic trajectories.

figure[figure omitted — 836 chars of source]

Curiously, Figure (ref) indicates that even when the two regressions arrive at the same point estimate, the confidence intervals can be markedly different under different sources of randomness.

figure[figure omitted — 1,093 chars of source]

The juxtaposition of these figures beg two questions:

tcolorbox[colback=gray!10!white,colframe=black!75!black] \begin{center} { {\bf Q1:} “When are HZ and VT point estimates identical?” \\ {\bf Q2:} “When the point estimates are identical, how does the source of randomness impact inference?” } \end{center}

{\bf Contribution.} This article tackles Q1--Q2 from first principles. In this endeavor, we begin by classifying several widely studied regression formulations into (i) a symmetric class that yields identical point estimates and (ii) an asymmetric class that yields contrasting point estimates. Within the symmetric class, we study properties of the estimator with randomness stemming from (i) time series patterns, (ii) cross-sectional patterns, and (iii) both patterns. We conduct our analysis from a (i) model-based perspective, which attributes randomness to the potential outcomes, and a (ii) design-based perspective, which attributes randomness to the treatment assignment mechanism. In both frameworks, we find that the source of randomness has large implications for the estimand and inference. Under the model-based framework, we construct confidence intervals for each source of randomness. Through data-inspired simulations and empirical applications, we demonstrate that the confidence interval developed for one estimand often has incorrect coverage for another estimand. Taken together, our results emphasize that the source of randomness that researchers assume has direct implications for the accuracy of the inference that can be conducted.

{\bf Organization.} Section (ref) overviews the panel data framework. Sections (ref)--(ref) provide one set of answers for Q1--Q2. Section (ref) illustrates concepts developed in this article. Section (ref) summarizes our findings. Details of simulations and empirical applications, select discussions, and mathematical proofs are relegated to the appendix.

{\bf Notation.} Let $\boldsymbol{I}$ be the identity matrix. Let $\boldsymbol{1}$ and $\boldsymbol{0}$ be the vectors of ones and zeros, respectively. The curled inequality denotes $\succeq$ the generalized inequality, i.e., componentwise inequality between vectors and matrix inequality between symmetric matrices. Let $\circ$ denote the componentwise product. For vectors $\boldsymbol{a}$ and $\boldsymbol{b}$, let $\langle a, b \rangle = a' b$ denote the inner product. Let $\operatorname{tr}(\bA)$ denote the trace of $\bA$. We define $0/0 = 0$ when applicable.

The Panel Data Framework

We anchor on the Basque study to introduce the panel data framework. Panel data contains observations of $N$ units over $T$ time periods. The Basque study, for instance, consists of per capita GDP across $N=17$ Spanish regions over $T=43$ years. In each time period $t$, each unit $i$ is characterized by two potential outcomes, $Y_{it}(0)$ and $Y_{it}(1)$, which correspond to its outcome in the absence and presence of a binary treatment, respectively. The potential outcomes framework posits that each region possesses two possible levels of economic activity each year, one that is immune to terrorism and another that is affected by terrorism. In reality, however, we can only observe one economic state, $Y_{it}(0)$ or $Y_{it}(1)$---this is the fundamental challenge of causal inference.

Let $Y_{it}$ be the observed outcome. Often, we observe all $N$ units without treatment (control) for $T_0$ time periods, i.e., $Y_{it} = Y_{it}(0)$ for all $i \le N$ and $t \le T_0$. For the remaining $T_1 = T-T_0$ time periods, $N_1$ units receive treatment while the remaining $N_0 = N-N_1$ units remain under control, i.e., if we arbitrarily label the first $N_0$ units as the control group, then $Y_{it} = Y_{it}(1)$ for all $i > N_0$ and $t > T_0$, and $Y_{it} = Y_{it}(0)$ for all $i \le N_0$ and $t > T_0$. In our study, Basque Country is the single treated unit, thus $N_1 = 1$ and $N_0=16$. The first wave of terrorist activity partitions the time horizon into pre- and post-treatment periods of lengths $T_0=15$ and $T_1=28$ years, respectively.

For ease of exposition, this article considers a single treated unit and single treated period indexed by the $N$th unit and $T$th time period, respectively. However, our results hold for any $(i,t)$ pair where $i > N_0$ is a treated unit and $t > T_0$ is a treated period. We organize our observed control data into an $N \times T$ matrix, $\boldsymbol{Y} = [Y_{it}]$, as shown in Figure (ref). In our example, $\boldsymbol{y}_N = [Y_{Nt}: t \le T_0] \in \Rb^{T_0}$ represents Basque Country's economic evolution prior to the outset of terrorism; $\boldsymbol{Y}_0 = [Y_{it}: i \le N_0, t \le T_0] \in \Rb^{N_0 \times T_0}$ represents the control regions' economic evolution prior to the outset of terrorism; and $\boldsymbol{y}_T = [Y_{iT}: i \le N_0] \in \Rb^{N_0}$ represents the control regions' economic evolution after the outset of terrorism. Our object of interest is Basque Country's counterfactual GDP in the absence of terrorism, $Y_{NT}(0)$.

figure[figure omitted — 208 chars of source]

Time Series Versus Cross-Sectional Based Regressions

The information across time and space motivates two natural ways to impute the missing $(N,T)$th entry. These perspectives are explored in two large and mostly separate bodies of work mc_panel.

Horizontal Regression and Unconfoundedness

The unconfoundedness literature operates on the concept that “history is a guide to the future”. As such, unconfoundedness methods express outcomes in the treated period as a weighted composition of outcomes in the pretreatment periods. This is carried out by regressing the control units' treated period outcomes $\boldsymbol{y}_T$ on its lagged outcomes $\boldsymbol{Y}_0$ and applying the learned regression coefficients to the treated unit's lagged outcomes $\boldsymbol{y}_N$ to predict the missing $(N,T)$th outcome. Following mc_panel, we refer to such methods as horizontal (HZ) regression.

Vertical Regression and Synthetic Controls

The synthetic controls literature is built on the concept that “similar units behave similarly”. Therefore, synthetic controls methods express the treated unit's outcomes as a weighted composition of control units' outcomes. This is carried out by regressing the treated unit's lagged outcomes $\boldsymbol{y}_N$ on the control units' lagged outcomes $\boldsymbol{Y}'_0$ and applying the learned regression coefficients to the control units' treated period outcomes $\boldsymbol{y}_T$ to predict the missing $(N,T)$th outcome. Following mc_panel, we refer to such methods as vertical (VT) regression.

Conventional Wisdom

The asymmetry between HZ and VT regressions has created the conception that they are fundamentally different approaches mc_panel. In fact, the unregularized forms of HZ and VT regressions are cautioned against when $T > N$ and $N > T$, respectively abadie4, sc_enet, li_bell, mc_panel. With regularization, however, mc_panel argues the two approaches can be applied to the same setting. In turn, this allows the two approaches to be systematically compared through methods such as cross-validation.

In parallel, the growth rates of the two literatures have also exhibited asymmetry. While the development of the unconfoundedness literature has seemingly plateaued, the synthetic controls literature continues to rapidly expand. Across many domains, synthetic controls based methods are arguably the de facto approach for panel studies.

Point Estimation

tcolorbox[colback=gray!10!white,colframe=black!75!black] \begin{center} { {\bf Q1:} “When are HZ and VT point estimates identical?” } \end{center}

We tackle Q1 by studying the finite-sample estimation properties of HZ and VT regressions. We denote the singular value decomposition of $\boldsymbol{Y}_0$ as $\boldsymbol{Y}_0 = \sum_{\ell =1}^{R} s_\ell \boldsymbol{u}_\ell \boldsymbol{v}'_\ell = \bU \bS \bV'$, where $\boldsymbol{u}_\ell \in \Rb^{N_0}$ and $\boldsymbol{v}_\ell \in \Rb^{T_0}$ are the left and right singular vectors, respectively, $s_\ell \in \Rb$ are the ordered singular values, and $R = \text{rank}(\boldsymbol{Y}_0) \le \min\{N_0, T_0\}$. $\bU \in \Rb^{N_0 \times R}$ and $\bV \in \Rb^{T_0 \times R}$ denote the matrices formed by the left and right singular vectors, respectively, and $\bS \in \Rb^{R \times R}$ is the diagonal matrix of singular values. The Moore-Penrose pseudoinverse of $\boldsymbol{Y}_0$ is $\boldsymbol{Y}_0^\dagger = \sum_{\ell=1}^R (1/s_\ell) \boldsymbol{v}_\ell \boldsymbol{u}'_\ell = \bV \bS^{-1} \bU'$. Critically, we do not place any assumptions on the relative magnitudes of $N$ and $T$.

Classifying Notable Regression Formulations

We present several of the most widely studied regression formulations in the HZ and VT literatures. This list is far from exhaustive given the vastness of these literatures.

Description of Estimation Strategies

{\bf Penalized regression.} A large class of penalized regressions are expressed as follows:

itemize• HZ regression: for $\lambda_1, \lambda_2 \ge 0$, \begin{align} &\widehat{\boldsymbol{\alpha}} = \operatorname*{\arg\!\min}_{\boldsymbol{\alpha}} \| \boldsymbol{y}_T - \boldsymbol{Y}_0 \boldsymbol{\alpha} \|_2^2 + \lambda_1 \| \boldsymbol{\alpha} \|_1 + \lambda_2 \|\boldsymbol{\alpha} \|_2^2 \\ &\widehat{Y}_{NT}^hz(0) = \langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}} \rangle. \end{align} • VT regression: for $\lambda_1, \lambda_2 \ge 0$, \begin{align} &\widehat{\boldsymbol{\beta}} = \operatorname*{\arg\!\min}_{\boldsymbol{\beta}} \| \boldsymbol{y}_N - \boldsymbol{Y}'_0 \boldsymbol{\beta} \|_2^2 + \lambda_1 \| \boldsymbol{\beta} \|_1 + \lambda_2 \| \boldsymbol{\beta} \|_2^2 \\ &\widehat{Y}_{NT}^vt(0) = \langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle. \end{align}

We overview common choices for $(\lambda_1, \lambda_2)$ and describe the corresponding strategy.

{\em I: Ordinary least squares (OLS).} Arguably, the mother of all regressions is OLS, where $\lambda_1 = \lambda_2 = 0$. OLS is an unconstrained problem with possibly infinitely many solutions. OLS has been analyzed in numerous works on panel studies, including hcw, li_bell, and li2020.

{\em II: Principal component regression (PCR).} To formalize PCR, let

align[align omitted — 120 chars of source]

denote the rank $k < R$ approximation of $\boldsymbol{Y}_0$ that retains the top $k$ principal components. HZ and VT PCR corresponds to replacing $\boldsymbol{Y}_0$ with $\boldsymbol{Y}_0^{(k)}$ within (ref) and (ref), respectively, with $\lambda_1 = \lambda_2 = 0$. In words, PCR first finds a $k$ dimensional representation of the covariate matrix via principal component analysis; then, PCR performs OLS with the compressed $k$ dimensional covariates. Within the synthetic controls literature, rsc, mrsc and si utilize PCR.

{\em III: Ridge regression.} Consider ridge regression, where $\lambda_1=0$ and $\lambda_2 > 0$. When $\boldsymbol{Y}_0$ is rank deficient, the gram matrix, i.e., $\boldsymbol{Y}'_0 \boldsymbol{Y}_0$ for HZ regression and $\boldsymbol{Y}_0 \boldsymbol{Y}'_0$ for VT regression, is ill-conditioned. This often discourages the usage of OLS. In these settings, ridge provides a remedy by adding a {\em ridge} on the diagonal of the gram matrix, which increases all eigenvalues by $\lambda_2$, thus removing the singularity problem. sc_aug explores the properties of a doubly robust estimator that utilizes HZ ridge regression.

{\em IV: Lasso regression.} Consider lasso regression, where $\lambda_1 > 0$ and $\lambda_2 = 0$. Lasso has become a popular tool for estimating sparse linear coefficients in high-dimensional regimes. Because the lasso criterion not strictly convex, there are possibly infinitely many solutions. Thus, for our analysis of lasso only, we make the mild assumption that the entries of $\boldsymbol{Y}_0$ are drawn from a continuous distribution. As established in lasso_ryan, this guarantees the lasso solution to be unique. Several notable works in the synthetic controls literature, e.g., li_bell, arco, and sc_lasso1, analyze the lasso.

{\em V: Elastic net regression.} Mixing both $\ell_1$ and $\ell_2$-penalties, i.e., $\lambda_1, \lambda_2 > 0$, is known as elastic net. At a high level, elastic net selects variables similar to the lasso, but deals with correlated variables more gracefully as with ridge. When $\lambda_2 > 0$, the criterion is strictly convex so the solution is unique. sc_enet propose an elastic net synthetic controls variant.

{\bf Constrained regression.}

{\em VI: Simplex regression.} The next formulation constrains the regression weights to lie within the simplex, i.e., the weights are nonnegative and sum to one:

itemize• HZ regression: for $\lambda \ge 0$, \begin{align} &\widehat{\boldsymbol{\alpha}} = \operatorname*{\arg\!\min}_{\boldsymbol{\alpha}} \| \boldsymbol{y}_T - \boldsymbol{Y}_0 \boldsymbol{\alpha} \|_2^2 + \lambda \| \boldsymbol{\alpha} \|_2^2 \, \, subject to \, \boldsymbol{\alpha}' \boldsymbol{1} = 1, \boldsymbol{\alpha} \succeq \boldsymbol{0} \\ &\widehat{Y}_{NT}^hz(0) = \langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}} \rangle. \end{align} • VT regression: for $\lambda \ge 0$, \begin{align} &\widehat{\boldsymbol{\beta}} = \operatorname*{\arg\!\min}_{\boldsymbol{\beta}} \| \boldsymbol{y}_N - \boldsymbol{Y}_0' \boldsymbol{\beta} \|_2^2 + \lambda \| \boldsymbol{\beta} \|_2^2 \, \, subject to \, \boldsymbol{\beta}' \boldsymbol{1} = 1, \boldsymbol{\beta} \succeq \boldsymbol{0} \\ &\widehat{Y}_{NT}^vt(0) = \langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle. \end{align}

We consider a vanishing $\ell_2$ penalty since $\lambda = 0$ (standard formulation) can induce multiple minima abadie_hour. When $\lambda > 0$, the criterion becomes strictly convex and the solution is unique. Simplex regression is the original formulation set forth in the pioneering works of abadie1, abadie2, abadie4, and its properties continue to be actively studied today. Attractive aspects of simplex regression include interpretability, sparsity, and transparency abadie_survey.

Classification Results

To answer Q1, we classify the regression formulations into (i) a symmetric class, where HZ and VT point estimates agree, and (ii) an asymmetric class, where HZ and VT point estimates disagree. We use the shorthand $\textsf{HZ} = \textsf{VT}$ if the two approaches produce identical point estimates and $\textsf{HZ} \neq \textsf{VT}$ otherwise.

{\bf I: Symmetric class.} We first state the symmetric formulations.

theorem$\emph{\textsf{HZ}} = \emph{\textsf{VT}}$ for (i) OLS with $(\widehat{\boldsymbol{\alpha}}, \widehat{\boldsymbol{\beta}})$ as the minimum $\ell_2$-norm solutions: \begin{align} \widehat{Y}_{NT}^{hz}(0) &= \widehat{Y}_{NT}^{vt}(0) = \langle \boldsymbol{y}_N, \boldsymbol{Y}_0^\dagger \boldsymbol{y}_T \rangle = \sum_{\ell =1}^R (1/s_\ell) \langle \boldsymbol{y}_N, \boldsymbol{v}_\ell \rangle \langle \boldsymbol{u}_\ell, \boldsymbol{y}_T \rangle; \end{align} (ii) PCR with the same choice of $k < R$: \begin{align} \widehat{Y}_{NT}^{hz}(0) &= \widehat{Y}_{NT}^{vt}(0) = \langle \boldsymbol{y}_N, (\boldsymbol{Y}_0^{(k)})^\dagger \boldsymbol{y}_T \rangle = \sum_{\ell =1}^k (1/s_\ell) \langle \boldsymbol{y}_N, \boldsymbol{v}_\ell \rangle \langle \boldsymbol{u}_\ell, \boldsymbol{y}_T \rangle; \end{align} (iii) ridge regression with the same choice of $\lambda_2>0$: \begin{align} \widehat{Y}_{NT}^{hz}(0) &= \widehat{Y}_{NT}^{vt}(0) = \langle \boldsymbol{y}_N, (\boldsymbol{Y}_0' \boldsymbol{Y}_0 + \lambda_2 \boldsymbol{I})^{-1} \boldsymbol{Y}_0' \boldsymbol{y}_T \rangle = \sum_{\ell =1}^R \frac{s_\ell}{s_\ell^2 + \lambda_2} \langle \boldsymbol{y}_N, \boldsymbol{v}_\ell \rangle \langle \boldsymbol{u}_\ell, \boldsymbol{y}_T \rangle. \end{align}

Theorem (ref) might seem familiar at first glance. As observed in abadie4 and sc_aug, the point estimates associated with HZ OLS and HZ ridge can be written as linear combinations of the elements in $\boldsymbol{y}_T$, which take the same {\em linear forms} as the corresponding VT point estimates. However, their results stop short of establishing numerical equivalence as in Theorem (ref). From this perspective, Theorem (ref) is perhaps surprising as it takes the next step forward in contradicting the notion that the two regressions are fundamentally different. In particular, Theorem (ref) proves the same point estimate is derived from both approaches whenever the regression model belongs to the symmetric class, namely (i) OLS with minimum $\ell_2$-norm, (ii) PCR for the same choice of $k$, and (iii) ridge for the same choice of $\lambda_2$. In this view, HZ and VT regressions are not perspectives at duel---they are dual perspectives.

We emphasize that Theorem (ref) holds for any data configuration. As such, it clarifies that HZ and VT OLS are {\em not} invalid when $T > N$ and $N > T$, respectively, as previously believed. In fact, the OLS estimate can even be written as $\langle \widehat{\boldsymbol{\alpha}}, \boldsymbol{Y}'_0 \widehat{\boldsymbol{\beta}} \rangle$, which incorporates both regression models. The origin of the prior misconception may have come from the fact that infinitely many solutions exist when $\boldsymbol{Y}_0$ is rank deficient. Among these solutions, however, is the unique minimum $\ell_2$-norm model, which is arguably sufficient for inference shao_deng. This is also the solution when the problem is optimized via gradient descent, a ubiquitous optimizer in practice. Phenomena of this form are known as “implicit regularization”, where the optimization algorithm is biased towards a particular solution even though the bias is not explicit in the objective function ireg2, ireg4.

Through its connection to the $\ell_2$-penalty, the minimum $\ell_2$-norm also offers a high-level intuition for the root of symmetry. More specifically, observe that the ridge model converges to the OLS model with minimum $\ell_2$-norm as $\lambda_2 \rightarrow 0$. Since the PCR model is precisely the OLS minimum $\ell_2$-norm model that is restricted to the space spanned by the top $k$ principal components, we conjecture that the geometry of the $\ell_2$-ball is a likely source for HZ and VT estimation symmetry.

{\bf II: Asymmetric class.} Next, we state the class of formulations that fracture symmetry.

theorem$\emph{\textsf{HZ}} \neq \emph{\textsf{VT}}$ for (i) lasso, (ii) elastic net, and (iii) simplex regression.

To examine the implications of Theorem (ref), we observe that the common thread between the objective functions in the asymmetric class is a penalty or constraint that promotes sparse models. Such regularizers are noticeably absent in the symmetric formulations. This highlights an interesting trade-off---while sparsity is widely considered to be a salient feature, the geometries of the $\ell_1$-ball and simplex that encourage sparsity are also the likely sources of HZ and VT estimation asymmetry.

Doubly Robust Estimators

In recent years, there has been a surge of interest in doubly robust estimators. Within panel data, we discuss two prominent works that are rising in popularity.

Synthetic Difference-in-Differences

An important approach that continues to dominate empirical work in panel data is the difference-in-differences (DID) estimator did1. At a high level, DID posits an additive outcome model with unit- and time-specific fixed effects, known more colloquially as the “parallel trends” assumption. The recent work of sdid anchors on the DID principle and brings in concepts from the unconfoundedness and synthetic controls literatures to derive a doubly robust estimator called synthetic difference-in-differences (SDID). In our setting, the SDID prediction for the missing $(N,T)$th potential outcome can be written as

align[align omitted — 465 chars of source]

where $\widehat{\boldsymbol{\alpha}}$ and $\widehat{\boldsymbol{\beta}}$ represent general HZ and VT models, respectively. The authors note that $\widehat{\boldsymbol{\alpha}} = (1/T_0) \boldsymbol{1}$ and $\widehat{\boldsymbol{\beta}} = (1/N_0) \boldsymbol{1}$ recovers DID. Moving beyond simple DID to performing a weighted two-way bias removal, the authors propose to learn $\widehat{\boldsymbol{\alpha}}$ via simplex regression and $\widehat{\boldsymbol{\beta}}$ via simplex regression with an $\ell_2$-penalty.

Augmented Synthetic Controls

Another notable work is that of sc_aug. The authors introduce the augmented synthetic control (ASC) estimator, which uses an outcome model to correct the bias induced by the classical synthetic controls estimator.\footnote{See abadie_hour for a bias correction of synthetic controls through matching.} Concretely, the ASC estimator predicts the missing $(N,T)$th potential outcome as

align[align omitted — 153 chars of source]

where $\widehat{M}_{iT}(0)$ is the estimator for the $(i,T)$th entry. The authors instantiate

align[align omitted — 103 chars of source]

Plugging the HZ outcome model in (ref) into (ref) then gives

align[align omitted — 291 chars of source]

We consider this particular variant of ASC since it takes the same form as SDID, as seen in (ref). In contrast to sdid, sc_aug learns $\widehat{\boldsymbol{\alpha}}$ via ridge regression and $\widehat{\boldsymbol{\beta}}$ via simplex regression.

When Doubly Robust Estimators are No Longer Doubly Robust

We leverage Theorem (ref) to study the estimation properties of SDID and ASC when $(\widehat{\boldsymbol{\alpha}}, \widehat{\boldsymbol{\beta}})$ are learned via OLS and PCR.

cor$\emph{\textsf{SDID}} = \emph{\textsf{ASC}} = \emph{\textsf{HZ}} = \emph{\textsf{VT}}$ for (i) $(\widehat{\boldsymbol{\alpha}}, \widehat{\boldsymbol{\beta}})$ as the OLS minimum $\ell_2$-norm solutions and (ii) $(\widehat{\boldsymbol{\alpha}}, \widehat{\boldsymbol{\beta}})$ as the PCR solutions with the same choice of $k < R$.

Corollary (ref) states that a researcher who uses the implicitly regularized (but explicitly unconstrained) variants of SDID and ASC inevitably arrive at the same point estimate as their colleague who simply uses HZ or VT OLS. The same phenomena occurs for PCR. From this, we observe that SDID and ASC can lose their weighted double-differencing effects under certain formulations. However, Corollary (ref) is not to discredit either approach. The fact remains that both methods allow researchers to naturally and simultaneously encode their knowledge of time and unit specific structures into the estimator. \textsf{SDID} and \textsf{ASC} warrant further studies and careful consideration.

Intercepts

Intercepts can be included in the HZ regression model by modifying the $\ell_2$-errors in (ref) and (ref) as $\| \boldsymbol{y}_T - \boldsymbol{Y}_0 \boldsymbol{\alpha} - \alpha_0 \boldsymbol{1} \|_2^2$; similarly, they can be included in the VT regression model by modifying $\ell_2$-errors in (ref) and (ref) as $\| \boldsymbol{y}_N - \boldsymbol{Y}'_0 \boldsymbol{\beta} - \beta_0 \boldsymbol{1} \|_2^2$. To date, there is still a lack of general consensus on the inclusion of intercepts within the synthetic controls literature. We attempt to shed light on the role of intercepts.

cor$\emph{\textsf{HZ}} \neq \emph{\textsf{VT}}$ for (i) OLS, (ii) PCR, and (iii) ridge with intercepts.

We develop an intuition for Proposition (ref) by interpreting intercepts in panel studies. A nonzero time intercept, $\alpha_0$, imposes a permanent constant difference between the treated and pretreatment periods; a nonzero unit intercept, $\beta_0$, imposes a permanent constant difference between the treated and control units. These systematic structures then create an asymmetry between the two regressions. Below, we propose a methodology based on centering the data that allows for intercepts yet retains symmetry.

Including Intercepts and Retaining Symmetry through Data Centering

Let $\boldsymbol{Y}_0$ be twice centered, i.e., the rows and columns of $\boldsymbol{Y}_0$ are mean zero. This can be satisfied by applying $\boldsymbol{I} - (1/N_0) \boldsymbol{1} \boldsymbol{1}'$ and $\boldsymbol{I} - (1/T_0) \boldsymbol{1} \boldsymbol{1}'$ to the left and right, respectively, of $\boldsymbol{Y}_0$. Consider the following modified formulations.

itemize• HZ regression: for $\lambda \ge 0$, \begin{align} &(\widehat{\alpha}_0, \widehat{\alpha}_1, \widehat{\boldsymbol{\alpha}}) = \operatorname*{\arg\!\min}_{(\alpha_0, \alpha_1, \boldsymbol{\alpha})} \| \boldsymbol{y}_T - \boldsymbol{Y}_0 \boldsymbol{\alpha} - \alpha_0 \boldsymbol{1} \|_2^2 + \| \boldsymbol{y}_N - \alpha_1 \boldsymbol{1} \|_2^2 + \lambda \|\boldsymbol{\alpha} \|_2^2 \\ &\widehat{Y}_{NT}^hz(0) = \langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}} \rangle + \widehat{\alpha}_0 + \widehat{\alpha}_1. \end{align} • VT regression: for $\lambda \ge 0$, \begin{align} &(\widehat{\beta}_0, \widehat{\beta}_1, \widehat{\boldsymbol{\beta}}) = \operatorname*{\arg\!\min}_{(\beta_0, \beta_1, \boldsymbol{\beta})} \| \boldsymbol{y}_N - \boldsymbol{Y}'_0 \boldsymbol{\beta} - \beta_0 \boldsymbol{1} \|_2^2 + \| \boldsymbol{y}_T - \beta_1 \boldsymbol{1} \|_2^2 + \lambda \|\boldsymbol{\beta} \|_2^2 \\ &\widehat{Y}_{NT}^vt(0) = \langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle + \widehat{\beta}_0 + \widehat{\beta}_1. \end{align}

Similar to before, OLS corresponds to $\lambda = 0$, PCR corresponds to OLS with $\boldsymbol{Y}_0^{(k)}$ for $k < R$ in place of $\boldsymbol{Y}_0$, and ridge regression corresponds to any $\lambda > 0$.

cor$\emph{\textsf{HZ}} = \emph{\textsf{VT}}$ for the symmetric estimators in Theorem (ref) under the formulations set in (ref) and (ref) with $\boldsymbol{Y}_0$ being twice centered.

We inspect (ref) and (ref) to understand the implications of Corollary (ref). First, we recall Theorem (ref), which establishes that the HZ and VT estimates share the same “base” estimate, i.e., $\langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}} \rangle = \langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle$. Next, we note that $\widehat{\alpha}_0 = \widehat{\beta}_1 = (1/N_0) \boldsymbol{1}' \boldsymbol{y}_T$ and $\widehat{\beta}_0 = \widehat{\alpha}_1 = (1/T_0) \boldsymbol{1}' \boldsymbol{y}_N$, which correspond to the time and unit fixed effects, respectively. Intuitively, the modified point estimates in (ref) and (ref) include both fixed effect models to compensate for $\boldsymbol{Y}_0$ being twice centered. Putting everything together, the modified HZ and VT point estimates are identical.

Inference

tcolorbox[colback=gray!10!white,colframe=black!75!black] \begin{center} { {\bf Q2:} “When the HZ and VT point estimates are identical, how does the source of randomness impact inference?” } \end{center}

To answer Q2, we study the inferential properties of the counterfactual prediction. Formal discussions for classical inference require an explicit postulation of where the randomness arises from. This article takes both a (i) model-based approach, which makes assumptions about the distribution of the potential outcomes, and a (ii) design-based approach, which makes assumptions about the assignment mechanism of treatment. We start by taking a model-based route that is similar in parts to the recent works of li2020, cattaneo, and sc_lasso1, sc_t_test from the synthetic controls literature. We then transition to a design-based route that is paved by the ideas introduced in guido_inf.

For ease of exposition, we focus on OLS and its minimum $\ell_2$-norm solutions, i.e., $\widehat{\boldsymbol{\alpha}} = \boldsymbol{Y}_0^\dagger \boldsymbol{y}_T$ and $\widehat{\boldsymbol{\beta}} = (\boldsymbol{Y}'_0)^\dagger \boldsymbol{y}_N$. As such, we define $\widehat{Y}_{NT}(0) = \widehat{Y}^\text{hz}_{NT}(0) = \widehat{Y}^\text{vt}_{NT}(0)$, which is justified under Theorem (ref). To improve readability, we present several results informally and provide their precise versions in Appendix (ref).

Model-Based Inference

Within the model-based framework, we consider a classical regression model. This postulation is not always plausible but it is useful to dissect its implications for the role of randomness in conducting inference.

Generative Models

We now study properties of $\widehat{Y}_{NT}(0)$ from three different sources of randomness.

{\bf I: Horizontal model.} The HZ model focuses on time series correlation patterns.

assumptionConditional on $(\boldsymbol{y}_N, \boldsymbol{Y}_0)$, we have \begin{align} &Y_{iT} = \sum_{t \le T_0} \alpha^*_t Y_{it} + \varepsilon_{iT}, \quad i=1, \dots, N_0, \end{align} where $\boldsymbol{\alpha}^*$ is a vector of unknown coefficients and $\varepsilon_{iT}$ is an idiosyncratic error term.

Assumption (ref) motivates the HZ approach, which models time series patterns as the source of randomness. Under the HZ model, the statistical uncertainty of $\widehat{Y}_{NT}(0)$ is governed by the construction of $\widehat{\boldsymbol{\alpha}}$ from $(\boldsymbol{y}_T, \boldsymbol{Y}_0)$, i.e., the in-sample uncertainty.

assumption$\{\varepsilon_{iT}\}_{i=1}^{N_0}$ has zero mean and is independent over $i = 1, \dots, N_0$, conditional on $(\boldsymbol{y}_N, \boldsymbol{Y}_0)$.

Assumption (ref) states that the errors have zero mean and thus the regressors, i.e., lagged outcomes, are uncorrelated with the errors; this is known in the literature as strict exogeneity. The errors are also conditionally independent across space.

{\bf II: Vertical model.} The VT model focuses on cross-sectional correlation patterns.

assumptionConditional on $(\boldsymbol{y}_T, \boldsymbol{Y}_0)$, we have \begin{align} &Y_{Nt} = \sum_{i \le N_0} \beta^*_i Y_{it} + \varepsilon_{Nt}, \quad t=1, \dots, T_0, \end{align} where $\boldsymbol{\beta}^*$ is a vector of unknown coefficients and $\varepsilon_{Nt}$ is an idiosyncratic error term.

Assumption (ref) is analogous to Assumption (ref) with the source of randomness here emanating from cross-sectional patterns. Hence, the statistical uncertainty of $\widehat{Y}_{NT}(0)$ under the VT model is governed by the construction of $\widehat{\boldsymbol{\beta}}$ from $(\boldsymbol{y}_N, \boldsymbol{Y}_0)$.

assumption$\{\varepsilon_{Nt}\}_{t=1}^{T_0}$ has zero mean and is independent over $t = 1, \dots, T_0$, conditional on $(\boldsymbol{y}_T, \boldsymbol{Y}_0)$.

Assumption (ref) is analogous to Assumption (ref) with the errors here being conditionally independent across time.

{\bf III: Mixed model.} We introduce a new model that mixes aspects of the HZ and VT models. At a high level, the mixed model accounts for randomness along both dimensions of the data. This comes at the price of placing additional constraints on the stochastic properties of the errors.

assumptionConditional on $\boldsymbol{Y}_0$, we have \begin{align} &Y_{iT} = \sum_{t \le T_0} \alpha^*_t Y_{it} + \varepsilon_{iT}, \quad i = 1, \dots, N_0, \\ &Y_{Nt} = \sum_{i \le N_0} \beta^*_i Y_{it} + \varepsilon_{Nt}, \quad t = 1, \dots, T_0, \end{align} where $(\boldsymbol{\alpha}^*, \varepsilon_{iT})$ and $(\boldsymbol{\beta}^*, \varepsilon_{Nt})$ are defined as in Assumptions (ref) and (ref), respectively.

(ref) and (ref) correspond to Assumptions (ref) and (ref), respectively. Collectively, they model time series and cross-sectional patterns as two distinct sources of randomness. Thus, the statistical uncertainty of $\widehat{Y}_{NT}(0)$ under the mixed model is governed by the constructions of both $\widehat{\boldsymbol{\alpha}}$ and $\widehat{\boldsymbol{\beta}}$.

assumption$\{\varepsilon_{iT}\}_{i=1}^{N_0}$ and $\{\varepsilon_{Nt}\}_{t=1}^{T_0}$ have zero mean and are independent over $i =1, \dots, N_0$ and $t = 1, \dots, T_0$, conditional on $\boldsymbol{Y}_0$.

Assumption (ref) combines Assumptions (ref) and (ref). In words, it states that $\boldsymbol{Y}_0$ contains all measured confounders and the errors are independent across both time and space. We leave a proper analysis under dependent errors as future work.

Inferential Properties

Equipped with our three models, we are ready to provide one set of answers to Q2. In what follows, we define the error covariance matrices as (i) $\bSigma^\text{hz}_T = \operatorname{Cov}(\boldsymbol{\varepsilon}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0)$, (ii) $\bSigma^\text{vt}_N = \operatorname{Cov}(\boldsymbol{\varepsilon}_N | \boldsymbol{y}_T, \boldsymbol{Y}_0)$, (iii) $\bSigma^\text{mix}_T = \operatorname{Cov}(\boldsymbol{\varepsilon}_T | \boldsymbol{Y}_0)$, and (iv) $\bSigma^\text{mix}_N = \operatorname{Cov}(\boldsymbol{\varepsilon}_N | \boldsymbol{Y}_0)$, where $\boldsymbol{\varepsilon}_T = [\varepsilon_{iT}: i \le N_0]$ and $\boldsymbol{\varepsilon}_N = [\varepsilon_{Nt}: t \le T_0]$.

theorem(i) [HZ model] Under Assumptions (ref)--(ref) and suitable moment conditions, we have \begin{align} & \frac{\widehat{Y}_{NT}(0) - \mu_0^hz}{ \sqrt{v_0^hz}} \xrightarrow{d} \mathcal{N}(0,1), \end{align} where $\mu_0^\emph{hz} = \langle \boldsymbol{y}_N, \bH^v \boldsymbol{\alpha}^* \rangle$ and $v_0^\emph{hz} = \widehat{\boldsymbol{\beta}}' \bSigma^\emph{hz}_T \widehat{\boldsymbol{\beta}}$. (ii) [VT model] Under Assumptions (ref)--(ref) and suitable moment conditions, we have \begin{align} \frac{\widehat{Y}_{NT}(0) - \mu_0^\emph{vt}}{ \sqrt{v^\emph{vt}_0}} \xrightarrow{d} \mathcal{N}(0,1), \end{align} where $\mu_0^\emph{vt} = \langle \boldsymbol{y}_T, \bH^u \boldsymbol{\beta}^* \rangle$ and $v^\emph{vt}_0 = \widehat{\boldsymbol{\alpha}}' \bSigma^\emph{vt}_N \widehat{\boldsymbol{\alpha}}$. (iii) \emph{[Mixed model]} Under Assumptions (ref)--(ref) and suitable moment conditions, we have \begin{align} \frac{ \widehat{Y}_{NT}(0) - \mu^\emph{mix}_0}{\sqrt{v^\emph{mix}_0}} \xrightarrow{d} \mathcal{N}(0,1), \end{align} where $\mu^\emph{mix}_0 = \langle \boldsymbol{\alpha}^*, \boldsymbol{Y}'_0 \boldsymbol{\beta}^* \rangle$ and $$v^\emph{mix}_0 = (\bH^u \boldsymbol{\beta}^*)' \bSigma^\emph{mix}_T (\bH^u \boldsymbol{\beta}^*) + (\bH^v \boldsymbol{\alpha}^*)' \bSigma^\emph{mix}_N (\bH^v \boldsymbol{\alpha}^*) + \operatorname{tr}( \boldsymbol{Y}^\dagger_0 \bSigma^\emph{mix}_T (\boldsymbol{Y}'_0)^\dagger \bSigma^\emph{mix}_N).$$

Theorem (ref) highlights that each model measures uncertainty with respect to a different estimand. In particular, Theorem (ref) states that the asymptotic variance is controlled by time series patterns under the HZ model, cross-sectional patterns under the VT model, and both correlation patterns under the mixed model. This clarifies that the source of randomness has substantive implications for the estimand and inference.

Confidence Intervals

Theorem (ref) motivates separate HZ, VT, and mixed confidence intervals: for $\theta \in (0,1)$,

align[align omitted — 364 chars of source]

where $z_{\frac \theta 2}$ is the upper $\theta/2$ quantile of $\mathcal{N}(0,1)$, and $(\widehat{v}^\text{hz}_0, \widehat{v}^\text{vt}_0, \widehat{v}^\text{mix}_0)$ are the estimators of $(v^\text{hz}_0, v^\text{vt}_0, v^\text{mix}_0)$. We construct

align[align omitted — 465 chars of source]

where $\widehat{\bSigma}_T$ and $\widehat{\bSigma}_N$ are the estimators of $(\bSigma^\text{hz}_T, \bSigma^\text{mix}_T)$ and $(\bSigma^\text{vt}_N, \bSigma^\text{mix}_N)$, respectively. We precisely define them under homoskedastic and heteroskedastic errors below. To reduce ambiguity, we index $(\widehat{v}^\text{hz}_0, \widehat{v}^\text{vt}_0, \widehat{v}^\text{mix}_0)$ by the covariance estimator. We also denote $\bH^u = \bU \bU'$ and $\bH^u_\perp = \boldsymbol{I} - \bH^u$. With this notation, the HZ and VT in-sample errors can be written as $\bH^u_\perp \boldsymbol{y}_T = \boldsymbol{y}_T - \boldsymbol{Y}_0 \widehat{\boldsymbol{\alpha}}$ and $\bH^v_\perp \boldsymbol{y}_N = \boldsymbol{y}_N - \boldsymbol{Y}'_0 \widehat{\boldsymbol{\beta}}$, respectively.

It is clear from (ref) that $(\widehat{v}^\text{hz}_0, \widehat{v}^\text{vt}_0)$ are plug-in estimators for $(v^\text{hz}_0, v^\text{vt}_0)$. As such, we discuss $\widehat{v}^\text{mix}_0$ with respect to $v^\text{mix}_0$. Recall $\widehat{\boldsymbol{\alpha}} = \bH^v \widehat{\boldsymbol{\alpha}}$ and $\widehat{\boldsymbol{\beta}} = \bH^u \widehat{\boldsymbol{\beta}}$ by construction. To justify the negative trace in $\widehat{v}^\text{mix}_0$, note that $\widehat{v}_0^\text{hz}$ is a quadratic involving $(\boldsymbol{y}_N, \boldsymbol{y}_T)$. Since both quantities are random, the expectation of $\widehat{v}_0^\text{hz}$ induces an additional term that precisely corresponds to the trace term in $v_0^\text{mix}$. The same property holds for $\widehat{v}_0^\text{vt}$. Thus, $\widehat{v}^\text{mix}_0$ corrects for this bias via the negative trace.

{\bf Homoskedastic errors.} Consider $\bSigma^\text{hz}_T$ with identical diagonal elements, i.e., $\bSigma^\text{hz}_T = (\sigma^\text{hz}_T)^2 \boldsymbol{I}$, where $(\sigma^\text{hz}_T)^2 = \operatorname{Var}(\varepsilon_{iT} | \boldsymbol{y}_N, \boldsymbol{Y}_0)$ for $i = 1, \dots, N_0$. Let $(\bSigma^\text{vt}_N, \bSigma^\text{mix}_T, \bSigma^\text{mix}_N)$ be defined analogously. We use the standard variance estimators

align[align omitted — 275 chars of source]

where $R = \text{rank}(\boldsymbol{Y}_0)$, which can be computed as $R = \operatorname{tr}(\bH^u) = \operatorname{tr}(\bH^v)$.

lemmaConsider homoskedastic errors. (i) [HZ model] Under Assumptions (ref)--(ref), we have \begin{align} \mathbb{E}[\widehat{\bSigma}_T^homo | \boldsymbol{y}_N, \boldsymbol{Y}_0] = \bSigma^hz_T \quad and \quad \mathbb{E}[ \widehat{v}^{hz, \emph{homo}}_0 | \boldsymbol{y}_N, \boldsymbol{Y}_0] = v^\emph{hz}_0. \end{align} (ii) \emph{[\textsf{VT} model]} Under Assumptions (ref)--(ref), we have \begin{align} \mathbb{E}[ \widehat{\bSigma}_N^\emph{homo} | \boldsymbol{y}_T, \boldsymbol{Y}_0] = \bSigma^\emph{vt}_N \quad \text{and} \quad \mathbb{E}[ \widehat{v}^{\emph{vt}, \emph{homo}}_0 | \boldsymbol{y}_T, \boldsymbol{Y}_0] = v^\emph{vt}_0. \end{align} (iii) \emph{[Mixed model]} Under Assumptions (ref)--(ref), we have \begin{align} \mathbb{E}[\widehat{\bSigma}_T^\emph{homo} | \boldsymbol{Y}_0] = \bSigma^\emph{mix}_T, \quad \mathbb{E}[ \widehat{\bSigma}_N^\emph{homo} | \boldsymbol{Y}_0] = \bSigma^\emph{mix}_N, \quad \text{and} \quad \mathbb{E}[ \widehat{v}^{\emph{mix}, \emph{homo}}_0 | \boldsymbol{Y}_0] = v^\emph{mix}_0. \end{align}

Lemma (ref) is a well known result within the OLS literature, albeit it is typically formalized under the stricter full column rank assumption.

{\bf Heteroskedastic errors.} We adopt two strategies for the heteroskedastic setting.

{\em I: Jackknife.} The first estimator is based on the jackknife. Traditionally, the jackknife estimates the covariance of the regression coefficients $(\widehat{\boldsymbol{\alpha}}, \widehat{\boldsymbol{\beta}})$. By analyzing said estimates, we derive the following:

align[align omitted — 494 chars of source]
lemmaConsider heteroskedastic errors. (i) [HZ model] Let Assumptions (ref)--(ref) hold. If $(\bH^u_\perp \circ \bH^u_\perp \circ \boldsymbol{I})$ is nonsingular, then \begin{align} \mathbb{E}[\widehat{\bSigma}^jack_T | \boldsymbol{y}_N, \boldsymbol{Y}_0] \succeq \bSigma^hz_T \quad and \quad \mathbb{E}[ \widehat{v}^{hz, \emph{jack}}_0 | \boldsymbol{y}_N, \boldsymbol{Y}_0] \ge v^\emph{hz}_0. \end{align} (ii) \emph{[\textsf{VT} model]} Let Assumptions (ref)--(ref) hold. If $(\bH^v_\perp \circ \bH^v_\perp \circ \boldsymbol{I})$ is nonsingular, then \begin{align} \mathbb{E}[ \widehat{\bSigma}_N^\emph{jack} | \boldsymbol{y}_T, \boldsymbol{Y}_0] \succeq \bSigma^\emph{vt}_N \quad \text{and} \quad \mathbb{E}[ \widehat{v}^{\emph{vt}, \emph{jack}}_0 | \boldsymbol{y}_T, \boldsymbol{Y}_0] \ge v^\emph{vt}_0. \end{align} (iii) \emph{[Mixed model]} Let Assumptions (ref)--(ref) hold. If $(\bH^u_\perp \circ \bH^u_\perp \circ \boldsymbol{I})$ and $(\bH^v_\perp \circ \bH^v_\perp \circ \boldsymbol{I})$ are nonsingular, then \begin{align} \mathbb{E}[\widehat{\bSigma}_T^\emph{jack} | \boldsymbol{Y}_0] \succeq \bSigma^\emph{mix}_T, \quad \mathbb{E}[ \widehat{\bSigma}_N^\emph{jack} | \boldsymbol{Y}_0] \succeq \bSigma^\emph{mix}_N, \quad \text{and} \quad \mathbb{E}[ \widehat{v}^{\emph{mix}, \emph{jack}}_0 | \boldsymbol{Y}_0] \ge v^\emph{mix}_0. \end{align}

Lemma (ref) establishes that the jackknife is conservative, provided $(\bH^u_\perp \circ \bH^u_\perp \circ \boldsymbol{I})$ and $(\bH^v_\perp \circ \bH^v_\perp \circ \boldsymbol{I})$ are nonsingular. Strictly speaking, the jackknife is well defined if these quantities are singular, as seen through the pseudoinverse in (ref) and (ref). Lemma (ref) considers the nonsingular case for simplicity. We remark that $\max_\ell H^u_{\ell \ell} < 1$ and $\max_\ell H^v_{\ell \ell} < 1$ are sufficient conditions for invertibility.

{\em II: HRK-estimator.} Next, we consider the covariance estimator proposed by variance. We index this estimator by the authors, Hartley-Rao-Kiefer:

align[align omitted — 442 chars of source]
lemmaConsider heteroskedastic errors. (i) [HZ model] Let Assumptions (ref)--(ref) hold. If $(\bH^u_\perp \circ \bH^u_\perp)$ is nonsingular, then \begin{align} \mathbb{E}[\widehat{\bSigma}^HRK_T | \boldsymbol{y}_N, \boldsymbol{Y}_0] = \bSigma^hz_T \quad and \quad \mathbb{E}[ \widehat{v}^{hz, \emph{HRK}}_0 | \boldsymbol{y}_N, \boldsymbol{Y}_0] = v^\emph{hz}_0. \end{align} (ii) \emph{[\textsf{VT} model]} Let Assumptions (ref)--(ref) hold. If $(\bH^v_\perp \circ \bH^v_\perp)$ is nonsingular, then \begin{align} \mathbb{E}[ \widehat{\bSigma}_N^\emph{HRK} | \boldsymbol{y}_T, \boldsymbol{Y}_0] = \bSigma^\emph{vt}_N \quad \text{and} \quad \mathbb{E}[ \widehat{v}^{\emph{vt}, \emph{HRK}}_0 | \boldsymbol{y}_T, \boldsymbol{Y}_0] = v^\emph{vt}_0. \end{align} (iii) \emph{[Mixed model]} Let Assumptions (ref)--(ref) hold. If $(\bH^u_\perp \circ \bH^u_\perp)$ and $(\bH^v_\perp \circ \bH^v_\perp)$ are nonsingular, then \begin{align} \mathbb{E}[\widehat{\bSigma}_T^\emph{HRK} | \boldsymbol{Y}_0] = \bSigma^\emph{mix}_T, \quad \mathbb{E}[ \widehat{\bSigma}_N^\emph{HRK} | \boldsymbol{Y}_0] = \bSigma^\emph{mix}_N, \quad \text{and} \quad \mathbb{E}[ \widehat{v}^{\emph{mix}, \emph{HRK}}_0 | \boldsymbol{Y}_0] = v^\emph{mix}_0. \end{align}

Lemma (ref) establishes that the HRK estimator is unbiased, provided $(\bH^u_\perp \circ \bH^u_\perp)$ and $(\bH^v_\perp \circ \bH^v_\perp)$ are invertible. For the former quantity, $\max_\ell H^u_{\ell \ell} < 1/2$ is a sufficient condition for invertibility. Since $\operatorname{tr}(\bH^u) = R$, this restricts $R < N_0/2$. A similar conclusion is drawn for VT regression.

Design-Based Inference

This section studies the counterfactual prediction from a design-based perspective, whereby the potential outcomes are considered fixed and the treatment assignments are considered stochastic.

Assumptions

In congruence with the article thus far, we focus on a single treated unit and treated period. We will find it useful to separate the assignment mechanism into the selection of each component. Accordingly, let $\bA \in \{0,1\}^T$ with $\boldsymbol{1}' \bA = 1$ and $\bB \in \{0,1\}^N$ with $\boldsymbol{1}' \bB = 1$ be the indicator vectors for the treated time period and treated unit, respectively, e.g., $B_N = 1$ and $A_T = 1$ if unit $N$ is treated at time $T$. With this notation, we denote the realized outcome as $Y_{it} = B_i A_t Y_{it}(1) + (1 - B_i A_t) Y_{it}(0)$. Following guido_inf, we consider the following assignment mechanisms:

assumption[Random assignment of time period] \begin{align} \Pb(\bA = \boldsymbol{a}) = \begin{cases} & 1/T, \quad if a_t \in \{0,1\} \forall t, \boldsymbol{1}' \boldsymbol{a} = 1 \\ & 0, \quad otherwise. \end{cases} \end{align}
assumption[Random assignment of unit] \begin{align} \Pb(\bB = \boldsymbol{b}) = \begin{cases} & 1/N, \quad if b_i \in \{0,1\} \forall i, \boldsymbol{1}' \boldsymbol{b} = 1 \\ & 0, \quad otherwise. \end{cases} \end{align}

Assumption (ref) considers the treated period to be randomly selected while Assumption (ref) considers the treated unit to be randomly selected. As guido_inf notes, these assumptions are not always plausible, but they underlie the placebo tests that are commonly used in synthetic controls applications.

Estimator

To conduct design-based analysis, we consider all possible treatment assignments, not only the realized assignment. Let $Y^*_{it}(0)$ be the OLS fit of $Y_{it}(0)$ based on $[Y_{j\tau} : j \neq i, \tau < t]$, $[Y_{i \tau} : \tau < t]$, and $[Y_{j t}: j \neq i]$. We define the design-based estimator as

align[align omitted — 105 chars of source]

In words, (ref) predicts the mean counterfactual outcome under control for unit $i$ at time $t$ if unit $i$ is treated at time period $t$, i.e., $B_i = 1$ and $A_t = 1$. We reemphasize that the stochasticity of $\widehat{Y}(0)$ stems from the treatment assignment mechanism since $Y^*_{it}(0)$ is a fixed quantity. As such, while the model-and design-based estimators share the same point estimate for the realized assignment, they differ in their formulations and attributions of randomness.

Inferential Properties

In Table (ref), we summarize the estimands associated with the model-based and design-based estimators under three sources of randomness: (i) time, (ii) unit, and (iii) time and unit. Within the model-based framework, mechanisms (i)--(iii) correspond to the HZ model (Assumptions (ref)--(ref)), VT model (Assumptions (ref)--(ref)), and mixed model (Assumptions (ref)--(ref)), respectively. Within the design-based framework, mechanisms (i)--(iii) correspond to Assumption (ref), Assumption (ref), and both Assumptions (ref) and (ref), respectively. Accordingly, the model-based and design-based expectations are taken over different probability measures.

table[table omitted — 1,187 chars of source]

Let us compare the estimands in Table (ref). In words, $\mathbb{E}[\widehat{Y}_{NT}(0) | \boldsymbol{y}_N, \boldsymbol{Y}_0]$ is a weighted combination of outcomes under control for the treated unit across all $T_0$ pretreatment periods; $\mathbb{E}[\widehat{Y}(0) | \bB]$ is a simple average of fitted outcomes under control for the treated unit across all $T$ periods, which guido_inf calls the “HZ” effect. Similarly, $\mathbb{E}[\widehat{Y}_{NT}(0) | \boldsymbol{y}_T, \boldsymbol{Y}_0]$ is a weighted combination of outcomes under control for the treated period across all $N_0$ control units; $\mathbb{E}[\widehat{Y}(0) | \bA]$ is a simple average of fitted outcomes under control for the treated period across all $N$ units, also called the “VT” effect. Finally, $\mathbb{E}[\widehat{Y}_{NT}(0) | \boldsymbol{Y}_0]$ is a weighted combination of outcomes under control across all $N_0$ control units and $T_0$ pretreatment periods; $\mathbb{E}[\widehat{Y}(0)]$ is a simple average of fitted outcomes under control across all unit and time period pairs, which we coin the “mixed” effect.

Though our design-based analysis is brief, there two important takeaways: (i) the model-based and design-based estimators recover similar estimands for each source of randomness; and (ii) different sources of randomness lead to different estimands, which is consistent with our model-based insights. The connection between assumptions of randomness and resulting estimand has been previously noticed in related contexts, e.g., guido_inf0, guido_inf and jas_inf.

Discussion

Several remarks on this section's results and extensions are in order.

remark[Correct Specification] For expositional convenience, we assume correct specification of the outcome model, as in li2020, to explain the panel data intuition and main theoretical results in a simple and transparent fashion. Alternatively, we can interpret Assumptions (ref), (ref), and (ref) as linear prediction models that are detached from structural meanings a la cattaneo and sc_lasso1, sc_t_test. For instance, within the mixed framework, we can redefine $\boldsymbol{\alpha}^* = \operatorname*{\arg\!\min} \mathbb{E}[ \| \boldsymbol{y}_T - \boldsymbol{Y}_0 \boldsymbol{\alpha} \|_2^2 | \boldsymbol{Y}_0]$ as the best linear approximation of $\boldsymbol{y}_T$ based on $\boldsymbol{Y}_0$, conditional on $\boldsymbol{Y}_0$. Such models can be justified via factor models and vector autoregressive models sc_t_test.
remark[Mixed Model] Any estimator, whether it falls within the symmetric or asymmetric class, can be studied under the mixed model (Assumptions (ref)--(ref)). There are also numerous ways to encode the mixed perspective beyond our postulation, e.g., random assignment of periods and units (Assumptions (ref)--(ref)) of guido_inf.
remark[Mixed Variance Estimators] We highlight that Lemmas (ref)--(ref) only hold in expectation. For any particular realization, $\widehat{v}^\text{mix}_0$ may exhibit unexpected properties. For instance, if $\operatorname{tr}(\boldsymbol{Y}_0^\dagger \widehat{\bSigma}_T (\boldsymbol{Y}'_0)^\dagger \widehat{\bSigma}_N) > \max\{\widehat{v}_0^\text{hz}, \widehat{v}_0^\text{vt} \}$, then $\widehat{v}^\text{mix}_0 < \min\{\widehat{v}_0^\text{hz}, \widehat{v}_0^\text{vt} \}$; thus, the mixed coverage will be smaller than both HZ and VT coverages. In fact, $\widehat{v}^\text{mix}_0$ can be negative if $\operatorname{tr}(\boldsymbol{Y}_0^\dagger \widehat{\bSigma}_T (\boldsymbol{Y}'_0)^\dagger \widehat{\bSigma}_N) > \widehat{v}_0^\text{hz} + \widehat{v}_0^\text{vt}$, which may occur if both HZ and VT in-sample errors are “too large”. For these scenarios, one na\"ive solution is to modify $\widehat{v}^\text{mix}_0$ as $\widehat{v}_0^{\text{mix}} \leftarrow \widehat{v}_0^\text{hz} + \widehat{v}_0^\text{vt}$, which is conservative by Lemmas (ref)--(ref). However, this case is arguably better resolved with a different point estimator altogether.
remark[Bridging Model-Based and Design-Based Inferences] The takeaways from the model-based and design-based analyses are consistent with one another. However, a formal and complete connection between the two perspectives on inference remains to be established. Towards this, we highlight that lin_ols demonstrates that one set of model-based confidence intervals (i.e., the Huber-White sandwich estimator) are justifiable under design-based arguments. In this view, a fascinating line of future inquiry is to analyze whether similar arguments hold for the heteroskedastic confidence intervals presented in Section (ref) or other related model-based constructions.
remark[On the Role of Randomness] This section underscores that the source of randomness plays a pivotal role for conducting inference. Translated to practice, our results stress that researchers should be scrupulous in reasoning through where the randomness in their data comes from. For instance, in the Basque study, some researchers may find it more plausible that the randomness is over time rather than over space, e.g., it is more conceivable that the onset of terrorism could have occurred in a different year but less conceivable that it could have occurred in a different region of Spain but not in the Basque Country. We do not take a substantive view on the matter, but we do stress that these decisions are meaningful for proper analysis as they have immediate implications for the resulting estimand and inferential procedure.
remark[Model Checking] It is possible in an application that substantive knowledge does not make clear on the source of randomness and hence which model to use, i.e., HZ, VT, or mixed. In such scenarios, the “in-space” and “in-time” placebo tests (a la cross-validation) proposed in abadie2, abadie4 are attractive tools to analyze the prediction properties of the various estimators under consideration. More generally, practices established within pcs provide an organized framework based on the principles of predictability, computability, and stability (PCS) to conduct rigorous comparison analyses.
remark[Extension to PCR] The previous results immediately extend to PCR by replacing $\boldsymbol{Y}_0$ with $\boldsymbol{Y}_0^{(k)}$, as defined in (ref), for any $k < R$. Intuitively, PCR-based models operate under the belief that the data is inherently low-dimensional. We comment on several benefits of PCR over OLS. To begin, the HZ and VT OLS variance estimators constructed in Section (ref) can suffer from degeneracy when $N$ and $T$ are of different sizes. That is, if $N < T$, then the HZ in-sample error is likely zero (otherwise known as overfitting), which causes the HZ coverage to collapse on the point estimate; analogous statements hold for the VT coverage when $N > T$. The PCR-based variance estimators, on the other hand, can avoid degeneracy through the number of chosen principal components $k$ (regularization). On a related note, the nonsingularity conditions required for the jackknife and HRK variance estimators can also be by controlled by $k$.

Illustrations

This section illustrates key concepts developed in this article. Our report is based on three canonical synthetic controls studies: (i) terrorism in Basque Country, (ii) California's Proposition 99 abadie2, and (iii) the reunification of West Germany abadie4. In particular, we will conduct a model-based analysis using the confidence intervals developed in Section (ref). We provide an overview of the results and relegate details (e.g., implementation) to Appendix (ref).

Background on Case Studies

{\bf Basque study.} See Sections (ref)--(ref) for details.

{\bf California study.} This study examines the effect of California's Proposition 99, an anti-tobacco legislation, on its tobacco consumption. The panel data contains per capita cigarette sales of $N = 39$ U.S. states over $T=31$ years. There are $T_0=18$ pretreatment observations and $N_0 = 38$ control units. Our interest is to estimate California's cigarette sales in the absence of Proposition 99.

{\bf West Germany study.} This study examines the economic impact of the 1990 reunification in West Germany. The panel data contains per capita GDP of $N=17$ countries over $T=44$ years. There are $T_0 = 30$ pretreatment observations and $N_0=16$ control units. Our interest is to estimate West Germany's GDP in the absence of reunification.

Data-Inspired Simulation Studies

We look to better understand the trade-offs in conducting inference under different sources of randomness. In an attempt to document our analysis in a realistic environment, we calibrate our simulations to our three studies.

Data Generating Process

We consider the single treated unit and time period setting. Specifically, we consider the actual treated unit, e.g., Basque Country, and focus on the first post-treatment period, e.g., one year after the outset of terrorism; hence, $T \leftarrow T_0 + 1$. Using the actual data, we generate the underlying regression models as

align[align omitted — 325 chars of source]

where $\boldsymbol{y}^*_N = [Y_{Nt}: t \le T_0]$, $\boldsymbol{y}^*_T = [Y_{iT}: i \le N_0]$, and $\boldsymbol{Y}^*_0 = [Y_{it}: i \le N_0, t \le T_0]$.

Observationally, we have access to the following quantities. Let $\boldsymbol{Y}_0$ be the rank $r$ approximation of $\boldsymbol{Y}^*_0$, where $r$ is chosen as the minimum number of singular values needed to capture at least $99.9\%$ of $\boldsymbol{Y}^*_0$'s spectral energy. Next, we sample $\boldsymbol{y}_T \sim \mathcal{N}(\boldsymbol{Y}_0 \boldsymbol{\alpha}^*, (N_0 - r)^{-1} \| \boldsymbol{y}^*_T - \boldsymbol{Y}^*_0 \boldsymbol{\alpha}^* \|_2^2 \boldsymbol{I})$ and $\boldsymbol{y}_N \sim \mathcal{N}(\boldsymbol{Y}'_0 \boldsymbol{\beta}^*, (T_0 - r)^{-1} \| \boldsymbol{y}^*_N - (\boldsymbol{Y}^*_0)' \boldsymbol{\beta}^* \|_2^2 \boldsymbol{I})$. We then define three estimands: (i) $\mu_0^\text{hz} = \langle \boldsymbol{y}_N, \bH^v \boldsymbol{\alpha}^* \rangle$, (ii) $\mu_0^\text{vt} = \langle \boldsymbol{y}_T, \bH^u \boldsymbol{\beta}^* \rangle$, and (iii) $\mu_0^\text{mix} = \langle \boldsymbol{\alpha}^*, \boldsymbol{Y}'_0 \boldsymbol{\beta}^* \rangle$, where $(\bH^u, \bH^v)$ are computed from $\boldsymbol{Y}_0$.

Simulation Results

For the purposes of stability, we conduct 500 replications of the above DGP for each study. In the $\ell$th simulation repeat, we learn the regression coefficients as

align[align omitted — 367 chars of source]

The corresponding point estimate is defined as $\widehat{Y}^{(\ell)}_{NT}(0) = \langle \boldsymbol{y}^{(\ell)}_N, \widehat{\boldsymbol{\alpha}}^{(\ell)} \rangle = \langle \boldsymbol{y}_T^{(\ell)}, \widehat{\boldsymbol{\beta}}^{(\ell)} \rangle$. We then construct separate HZ, VT, and mixed homoskedastic confidence intervals based on (ref) and (ref) around the point estimate.

In Table (ref), we report the coverage probabilities (CP) and average lengths (AL) for each confidence interval with respect to each estimand at the $95\%$ nominal mark. With respect to $\mu_0^\text{hz}$, the coverage of the HZ confidence interval is closer to the nominal coverage than that of the VT and mixed intervals as the latter two can substantially under- or over-cover. This storyline is consistent for the VT interval with respect to $\mu_0^\text{vt}$ and the mixed interval with respect to $\mu_0^\text{mix}$.

Collectively, our formal results and simulations demonstrate that (i) the choice of estimand directly affects the accuracy of the inference; and (ii) the variance formulas developed for one estimand may not have the correct coverage for another estimand. Accordingly, researchers should carefully consider the source of randomness in their data as it can have a significant influence over their ability to conduct valid inference. We comment that these conclusions are in line with those drawn in jas_inf, which analyzes the classical difference-in-means estimator with respect to standard estimands for randomized control trials.

table[table omitted — 1,480 chars of source]

Empirical Applications

Next, we analyze our three case studies of interest. All regression models are built on pretreatment data only, and the point and variance estimation formulas are separately applied for the treated unit at each post-treatment period $t > T_0$.

Point Estimation

Figure (ref) visualizes the counterfactual trajectories generated by the estimators in Section (ref). Our findings reinforce Theorems (ref) and (ref). On a separate note, we observe that, within the Basque study, the OLS estimates are wildly different from the other estimates and HZ simplex regression reduces to the last observation carried forward (LOCF) estimator. In the California and West Germany studies, the estimates are all qualitatively similar with the exception of the HZ simplex regression, which again reduces to LOCF. In fact, the OLS and ridge estimates appear to overlap, as well as the lasso and elastic net estimates.

figure[figure omitted — 1,470 chars of source]

Inference

To reduce visual redudancy, we only present the figures associated with the jackknife-based intervals for OLS and PCR in Figures (ref) and (ref), respectively. Consider the Basque study. The top row of plots in both figures demonstrates that $\mu_0^\text{vt}$ is more accurately estimated than both $\mu_0^\text{hz}$ and $\mu_0^\text{mix}$. Put differently, there is less uncertainty about conducting inference on $\mu_0^\text{vt}$ relative to the other estimands. At the same time, these plots indicate that if $\mu_0^\text{hz}$ or $\mu_0^\text{mix}$ are the estimands of interest, then the VT confidence interval will undercover in both settings. Analogous statements can be made for the remaining subfigures. As with our simulations, the large potential differences in coverage reinforce the importance of properly reasoning through the source of randomness in the data.

figure[figure omitted — 2,041 chars of source]
figure[figure omitted — 1,943 chars of source]

Conclusion

This article rewrites the conventional wisdom on HZ and VT regressions for panel data analysis. Contrary to standard notions, we show that the two regressions yield identical point estimates under several standard settings. At the same time, we articulate that the source of randomness directly affects the accuracy of the inference that can be conducted. From a practical standpoint, this stresses that researchers should carefully consider where the randomness in their data stems from as this decision will then guide their choice of estimand and inferential procedure.

appendix\section{Inference} \subsection{Model-Based Inference: Asymptotic Properties} First, we state the precise form of Theorem (ref). Towards this, let $(\sigma^\text{hz}_{iT})^2 = \operatorname{Var}(\varepsilon_{iT} | \boldsymbol{y}_N, \boldsymbol{Y}_0)$ for $i=1, \dots, N_0$. We define $((\sigma^\text{vt}_{Nt})^2, (\sigma^\text{mix}_{iT})^2, (\sigma^\text{mix}_{Nt})^2)$ for $i=1, \dots, N_0$ and $t = 1, \dots, T_0$ with respect to $(\bSigma^\text{vt}_N, \bSigma^\text{mix}_T, \bSigma^\text{mix}_N)$ analogously. \begin{theorem} (i) [HZ model] Let Assumptions (ref)--(ref) hold. If \begin{align} \left(\sum_{i \le N_0} \mathbb{E}\left[ |\widehat{\beta}_i \varepsilon_{iT} |^3 | \boldsymbol{y}_N, \boldsymbol{Y}_0 \right] \right)^2 &= o \left( \Big( \sum_{i \le N_0} \widehat{\beta}^2_i (\sigma^hz_{iT})^2 \Big)^3 \right), \end{align} then \begin{align} & \frac{\widehat{Y}_{NT}(0) - \mu_0^hz}{ \sqrt{v^hz_0}} \xrightarrow{d} \mathcal{N}(0,1), \end{align} where $\mu_0^\emph{hz} = \langle \boldsymbol{y}_N, \bH^v \boldsymbol{\alpha}^* \rangle$ and $v_0^\emph{hz} = \widehat{\boldsymbol{\beta}}' \bSigma^\emph{hz}_T \widehat{\boldsymbol{\beta}}$. (ii) [\textsf{VT} model] Let Assumptions (ref)--(ref) hold. If \begin{align} \left(\sum_{t \le T_0} \mathbb{E}\left[ | \widehat{\alpha}_t \varepsilon_{Nt} |^3 | \boldsymbol{y}_T, \boldsymbol{Y}_0 \right] \right)^2 &= o \left( \Big( \sum_{t \le T_0} \widehat{\alpha}^2_t (\sigma_{Nt}^\emph{vt})^2 \Big)^3 \right), \end{align} then \begin{align} \frac{\widehat{Y}_{NT}(0) - \mu_0^\emph{vt}}{ \sqrt{v_0^\emph{vt}}} \xrightarrow{d} \mathcal{N}(0,1), \end{align} where $\mu^\emph{vt}_0 = \langle \boldsymbol{y}_T, \bH^u \boldsymbol{\beta}^* \rangle$ and $v^\emph{vt}_0 = \widehat{\boldsymbol{\alpha}}' \bSigma^\emph{vt}_N \widehat{\boldsymbol{\alpha}}$. (iii) \emph{[Mixed model]} Let Assumptions (ref)--(ref) hold. If \begin{align} &\left( \sum_{i \le N_0} \sum_{t \le T_0} \mathbb{E}\Big[ | (\boldsymbol{Y}_0^\dagger)_{it} \left\{ \mathbb{E}[Y_{iT} | \boldsymbol{Y}_0] \varepsilon_{Nt} + \mathbb{E}[Y_{Nt} | \boldsymbol{Y}_0] \varepsilon_{iT} + \varepsilon_{iT} \varepsilon_{Nt} \right\} |^3 | \boldsymbol{Y}_0 \Big] \right)^2 \\ & = o \left( \left( \sum_{i \le N_0} \sum_{t \le T_0} (\boldsymbol{Y}_0^\dagger)^2_{it} \left\{ \mathbb{E}[Y_{iT} | \boldsymbol{Y}_0]^2 (\sigma^\emph{mix}_{Nt})^2 + \mathbb{E}[Y_{Nt} | \boldsymbol{Y}_0]^2 (\sigma^\emph{mix}_{iT})^2 + (\sigma^\emph{mix}_{iT})^2 (\sigma^\emph{mix}_{Nt})^2 \right\} \right)^3 \right), \end{align} then \begin{align} \frac{ \widehat{Y}_{NT}(0) - \mu^\emph{mix}_0}{\sqrt{v^\emph{mix}_0}} \xrightarrow{d} \mathcal{N}(0,1), \end{align} where $\mu^\emph{mix}_0 = \langle \boldsymbol{\alpha}^*, \boldsymbol{Y}'_0 \boldsymbol{\beta}^* \rangle$ and \begin{align} v^\emph{mix}_0 = (\bH^u \boldsymbol{\beta}^*)' \bSigma^\emph{mix}_T (\bH^u \boldsymbol{\beta}^*) + (\bH^v \boldsymbol{\alpha}^*)' \bSigma^\emph{mix}_N (\bH^v \boldsymbol{\alpha}^*) + \operatorname{tr}( \boldsymbol{Y}^\dagger_0 \bSigma^\emph{mix}_T (\boldsymbol{Y}'_0)^\dagger \bSigma^\emph{mix}_N). \end{align} \end{theorem} If $(\varepsilon_{iT}, (\sigma^\text{hz}_{iT})^2)$ are bounded, then (ref) translates to $\sum_{i \le N_0} | \widehat{\beta}_i |^3 = o( \| \widehat{\boldsymbol{\beta}} \|_2^3)$, which rules out outlier coefficients; a similar interpretation can be derived for (ref). Similarly, if $(\varepsilon_{iT}, \varepsilon_{Nt})$ and $(\sigma^2_{iT}, \sigma^2_{Nt})$ are bounded for all $(i,t)$, then (ref) loosely translates to $$ \sum_{i \le N_0} | \widehat{\beta}_i|^3 + \sum_{t \le T_0} | \widehat{\alpha}_t|^3 + \sum_{i \le N_0} \sum_{t \le T_0} | (\boldsymbol{Y}^\dagger_0)_{it} |^3 = o\left( \| \widehat{\boldsymbol{\beta}} \|_2^3 + \| \widehat{\boldsymbol{\alpha}} \|_2^3 + \| \boldsymbol{Y}_0^\dagger \|_F^3 \right), $$ which effectively bounds the magnitudes of the \textsf{HZ} and \textsf{VT} OLS coefficients and pseudoinverse matrix entries. We note that (ref)--(ref) are known as Lyapunov's condition, and we refer the interested reader to lehmann for details. Next, we provide a bound on the trace term in $v^\text{mix}_0$. Beginning with the upper bound, notice that \begin{align} \operatorname{tr}( \boldsymbol{Y}^\dagger_0 \bSigma^\text{mix}_T (\boldsymbol{Y}'_0)^\dagger \bSigma^\text{mix}_N) \le \max_{i \le N_0} (\sigma^\text{mix}_{iT})^2 \max_{t \le T_0} (\sigma^\text{mix}_{Nt})^2 \operatorname{tr}( \boldsymbol{Y}^\dagger_0 (\boldsymbol{Y}'_0)^\dagger). \end{align} By the cyclic property of the trace operator, \begin{align} \operatorname{tr}( \boldsymbol{Y}^\dagger_0 (\boldsymbol{Y}'_0)^\dagger) = \operatorname{tr}( \bV \bS^{-2} \bV') = \operatorname{tr}(\bS^{-2} \bV' \bV) = \operatorname{tr}(\bS^{-2}). \end{align} Putting everything together, we obtain \begin{align} \operatorname{tr}( \boldsymbol{Y}^\dagger_0 \bSigma^\text{mix}_T (\boldsymbol{Y}'_0)^\dagger \bSigma^\text{mix}_N) \le \max_{i \le N_0} (\sigma^\text{mix}_{iT})^2 \max_{t \le T_0} (\sigma^\text{mix}_{Nt})^2 \operatorname{tr}(\bS^{-2}). \end{align} The same arguments can be applied to derive the lower bound, which yields \begin{align} &\operatorname{tr}( \boldsymbol{Y}^\dagger_0 \bSigma^\text{mix}_T (\boldsymbol{Y}'_0)^\dagger \bSigma^\text{mix}_N) \\ &\qquad \in \operatorname{tr}(\bS^{-2}) \left[ \min_{i \le N_0} (\sigma^\text{mix}_{iT})^2 \min_{t \le T_0} (\sigma^\text{mix}_{Nt})^2, \max_{i \le N_0} (\sigma^\text{mix}_{iT})^2 \max_{t \le T_0} (\sigma^\text{mix}_{Nt})^2 \right]. \end{align} From (ref), we see that the lower and upper bounds match under homoskedasticity. \subsection{Model-Based Inference: Confidence Intervals} Next, we discuss technical aspects of the variance estimators in Section (ref). {\bf Homoskedastic errors.} We take note of the recent work of si in the synthetic controls literature. si propose a \textsf{VT} PCR estimator under the homoskedastic setting and provide a similar confidence interval to that of (ref) via large-sample approximations. Under a closely related \textsf{VT} model, they propose $\widehat{\boldsymbol{\beta}}' \widehat{\bSigma}_N^\text{homo} \widehat{\boldsymbol{\beta}}$ in place of $\widehat{\boldsymbol{\alpha}}' \widehat{\bSigma}_N^\text{homo} \widehat{\boldsymbol{\alpha}}$. While the point estimate of si also takes the form $\langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle$, their variance estimator only depends on $(\boldsymbol{y}_N, \boldsymbol{Y}_0)$; in comparison, ours depends on $(\boldsymbol{y}_N, \boldsymbol{y}_T, \boldsymbol{Y}_0)$. Therefore, the confidence interval as per si is numerically identical for every post-treatment point estimate while ours can vary across the post-treatment periods, which may be favorable. {\bf Heteroskedastic errors.} Consider the heteroskedastic setting. {\em I: Jackknife.} In the following lemma, we quantify the bias in Lemma (ref). \begin{lemma} [Detailed restatement of Lemma (ref)] (i) \emph{[\textsf{HZ} model]} Let Assumptions (ref)--(ref) hold. If $(\bH^u_\perp \circ \bH^u_\perp \circ \boldsymbol{I})$ is nonsingular, then \begin{align} \mathbb{E}[\widehat{\bSigma}^\emph{jack}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0] = \bSigma^\emph{hz}_T + \boldsymbol{\Delta}^\emph{hz} \quad \text{and} \quad \mathbb{E}[ \widehat{v}^{\emph{hz}, \emph{jack}}_0 | \boldsymbol{y}_N, \boldsymbol{Y}_0] = v^\emph{hz}_0 + \widehat{\boldsymbol{\alpha}}' \boldsymbol{\Delta}^\emph{hz} \widehat{\boldsymbol{\alpha}}, \end{align} where $\Delta^\emph{hz}_{\ell \ell} = \sum_{j \neq \ell} (\sigma^\emph{hz}_{jT})^2 (H^u_{\ell j})^2 (1 - H^u_{\ell \ell})^{-2}$ for $\ell=1, \dots, N_0$. (ii) \emph{[\textsf{VT} model]} Let Assumptions (ref)--(ref) hold. If $(\bH^v_\perp \circ \bH^v_\perp \circ \boldsymbol{I})$ is nonsingular, then \begin{align} \mathbb{E}[ \widehat{\bSigma}_N^\emph{jack} | \boldsymbol{y}_T, \boldsymbol{Y}_0] = \bSigma^\emph{vt}_N + \boldsymbol{\Gamma}^\emph{vt} \quad \text{and} \quad \mathbb{E}[ \widehat{v}^{\emph{vt}, \emph{jack}}_0 | \boldsymbol{y}_T, \boldsymbol{Y}_0] = v^\emph{vt}_0 + \widehat{\boldsymbol{\beta}}' \boldsymbol{\Gamma}^\emph{vt} \widehat{\boldsymbol{\beta}}, \end{align} where $\Gamma^\emph{vt}_{\ell \ell} = \sum_{j \neq \ell} (\sigma^\emph{vt}_{Nj})^2 (H^v_{\ell j})^2(1-H^v_{\ell \ell})^{-2}$ for $\ell=1, \dots, T_0$. (iii) \emph{[Mixed model]} Let Assumptions (ref)--(ref) hold. If $(\bH^u_\perp \circ \bH^u_\perp \circ \boldsymbol{I})$ and $(\bH^v_\perp \circ \bH^v_\perp \circ \boldsymbol{I})$ are nonsingular, then \begin{align} &\mathbb{E}[\widehat{\bSigma}_T^\emph{jack} | \boldsymbol{Y}_0] = \bSigma^\emph{mix}_T + \boldsymbol{\Delta}^\emph{mix}, \quad \mathbb{E}[ \widehat{\bSigma}_N^\emph{jack} | \boldsymbol{Y}_0] = \bSigma^\emph{mix}_N + \boldsymbol{\Gamma}^\emph{mix}, \\ &\mathbb{E}[\widehat{v}^{\emph{mix}, \emph{jack}}(\boldsymbol{Y}_0) | \boldsymbol{Y}_0] \&\qquad = v^\emph{mix}_0 + (\bH^u \boldsymbol{\beta}^*)' \boldsymbol{\Delta}^\emph{mix} (\bH^u \boldsymbol{\beta}^*) + (\bH^v \boldsymbol{\alpha}^*)' \boldsymbol{\Gamma}^\emph{mix} (\bH^v \boldsymbol{\alpha}^*) + \operatorname{tr}(\boldsymbol{Y}_0^\dagger \boldsymbol{\Delta}^\emph{mix} (\boldsymbol{Y}'_0)^\dagger \boldsymbol{\Gamma}^\emph{mix}), \end{align} where $\Delta^\emph{mix}_{\ell \ell}$ and $\Gamma^\emph{mix}_{\ell \ell}$ are defined analogously to $\Delta^\emph{hz}_{\ell \ell}$ and $\Gamma^\emph{vt}_{\ell \ell}$, respectively, with $(\sigma^\emph{mix}_{jT})^2$ and $(\sigma^\emph{mix}_{Nj})^2$ in place of $(\sigma^\emph{hz}_{jT})^2$ and $(\sigma^\emph{vt}_{Nj})^2$, respectively. \end{lemma} Without loss of generality, we consider the magnitudes of $(\Delta^\text{mix}_{\ell \ell}, \Gamma^\text{mix}_{\ell \ell})$. Towards bounding the former quantity, notice that $\bH^u$ is an orthogonal projector and is thus idempotent, i.e., $(\bH^u)^2 = \bH^u$, and symmetric. Therefore, \begin{align} H^u_{\ell \ell} = (H^u_{\ell \ell})^2 + \sum_{j \neq \ell} (H^u_{\ell j})^2 \implies \sum_{j \neq \ell} (H^u_{\ell j})^2 = H^u_{\ell \ell} ( 1 - H^u_{\ell \ell}). \end{align} This yields $\Delta^\text{mix}_{\ell \ell} \in H^u_{\ell \ell} (1 - H^u_{\ell \ell})^{-1} [ \min_{j \neq \ell} (\sigma^\text{mix}_{jT})^2, \max_{j \neq \ell} (\sigma^\text{mix}_{jT})^2]$. Since $\sum_{j \neq \ell} (H_{\ell j}^u)^2 \ge 0$, (ref) implies $H^u_{\ell \ell} \in [0,1]$. Thus, $\Delta^\text{mix}_{\ell \ell} = 0$ if $H^u_{\ell \ell} = 0$ and diverges if $H^u_{\ell \ell} = 1$. If $H^u_{\ell \ell}$ takes the average value $N_0^{-1} \sum_{\ell} H^u_{\ell \ell} = R N_0^{-1}$, then it follows that $\Delta^\text{mix}_{\ell \ell} \in R (N_0-R)^{-1} [ \min_{j \neq \ell} (\sigma^\text{mix}_{jT})^2, \max_{j \neq \ell} (\sigma^\text{mix}_{jT})^2]$. A similar result is derived for $\Gamma^\text{mix}_{\ell \ell}$. {\em II: HRK-estimator.} Recall Lemma (ref). Consider the \textsf{HZ} estimator and the invertibility of $(\bH^u \circ \bH^u)$. A sufficient condition is strict diagonal dominance varga1969matrix: $(1 - H^u_{\ell \ell})^2 > \sum_{j \neq \ell} (H^u_{\ell j})^2$. Using (ref), we simplify this condition as $(1 - H^u_{\ell \ell})^2 > H^u_{\ell \ell} - (H^u_{\ell \ell})^2$. Thus, $\max_\ell H^u_{\ell \ell} < 1/2$ is a sufficient condition for invertibility. The same arguments apply for \textsf{VT} regression. {\bf Mixed variance estimator.} Let us bound $\widehat{v}^\text{mix}_0$. From (ref), it follows that $\widehat{v}^\text{mix}_0 \in [\widehat{v}^\text{mix}_{0, \min}, \widehat{v}^\text{mix}_{0, \max}]$, where \begin{align} & \widehat{v}^\text{mix}_{0, \min} = \min_{i \le N_0} \widehat{\sigma}^2_{iT} \| \widehat{\boldsymbol{\beta}} \|_2^2 + \min_{t \le T_0} \widehat{\sigma}^2_{Nt} \| \widehat{\boldsymbol{\alpha}} \|_2^2 - \max_{i \le N_0} \widehat{\sigma}^2_{iT} \max_{t \le T_0} \widehat{\sigma}^2_{Nt} \operatorname{tr}(\bS^{-2}), \\ & \widehat{v}^\text{mix}_{0, \max} = \max_{i \le N_0} \widehat{\sigma}^2_{iT} \| \widehat{\boldsymbol{\beta}} \|_2^2 + \max_{t \le T_0} \widehat{\sigma}^2_{Nt} \| \widehat{\boldsymbol{\alpha}} \|_2^2 - \min_{i \le N_0} \widehat{\sigma}^2_{iT} \min_{t \le T_0} \widehat{\sigma}^2_{Nt} \operatorname{tr}(\bS^{-2}). \end{align} Here, $\widehat{\sigma}^2_{iT}$ and $\widehat{\sigma}^2_{Nt}$ are the $i$th and $t$th diagonal elements of $\widehat{\bSigma}_T$ and $\widehat{\bSigma}_N$, respectively. Observe that $\widehat{v}^\text{mix}_{0, \min} = \widehat{v}^\text{mix}_{0, \max}$ for homoskedastic errors. \section{Illustrations} This section provides details on the simulations that were absent in the main article. \subsection{Implementation Details} For ridge, lasso, and elastic net regressions, we use the default \texttt{scikit-learn} hyperparameters ($\lambda_1, \lambda_2$). For PCR, we choose the number of principal components $k$ via the approach described in Section (ref). This yields $k=$ for the Basque study, $k=3$ for the California study, and $k = 4$ for the West Germany study. We implement simplex regression using the code made available at \href{https://matheusfacure.github.io/python-causality-handbook/15-Synthetic-Control.html}{https://matheusfacure.github.io/python-causality-handbook/15-Synthetic-Control.html}. \subsection{Data-Inspired Simulation Studies} Formally, we define average length (AL) as \begin{align} \text{AL} = \frac{1}{m} \sum_{\ell=1}^m \frac{(2 \cdot 1.96) \sqrt{\widehat{v}^{(\ell)}_0}}{|\widehat{Y}^{(\ell)}_{NT}(0) |}, \end{align} where $\widehat{v}^{(\ell)}_0$ and $\widehat{Y}^{(\ell)}_{NT}(0)$ are the variance and point estimates for the $\ell$th repeat. \section{Proofs for Point Estimation} {\bf Helper lemmas.} To establish Theorems (ref) and (ref), we first state the following useful lemmas for the collection of regression formulations presented in Section (ref). We provide their proofs in Appendix (ref). \begin{lemma} [OLS] $\emph{\textsf{HZ}} = \emph{\textsf{VT}}$ for OLS with $(\widehat{\boldsymbol{\alpha}}, \widehat{\boldsymbol{\beta}})$ as the minimum $\ell_2$-norm solutions: \begin{align} \widehat{Y}_{NT}^{\emph{hz}}(0) &= \widehat{Y}_{NT}^{\emph{vt}}(0) = \langle \boldsymbol{y}_N, \boldsymbol{Y}_0^\dagger \boldsymbol{y}_T \rangle = \sum_{\ell=1}^R (1/s_\ell) \langle \boldsymbol{y}_N, \boldsymbol{v}_\ell \rangle \langle \boldsymbol{u}_\ell, \boldsymbol{y}_T \rangle. \end{align} \end{lemma} \begin{lemma} [PCR] $\emph{\textsf{HZ}} = \emph{\textsf{VT}}$ for PCR with the same choice of $k < R$: \begin{align} \widehat{Y}_{NT}^{\emph{hz}}(0) &= \widehat{Y}_{NT}^{\emph{vt}}(0) = \langle \boldsymbol{y}_N, (\boldsymbol{Y}_0^{(k)})^\dagger \boldsymbol{y}_T \rangle = \sum_{\ell=1}^k (1/s_\ell) \langle \boldsymbol{y}_N, \boldsymbol{v}_\ell \rangle \langle \boldsymbol{u}_\ell, \boldsymbol{y}_T \rangle. \end{align} \end{lemma} \begin{lemma} [Ridge] $\emph{\textsf{HZ}} = \emph{\textsf{VT}}$ for ridge regression with the same choice of $\lambda_2 > 0$: \begin{align} \widehat{Y}_{NT}^{\emph{hz}}(0) &= \widehat{Y}_{NT}^{\emph{vt}}(0) = \langle \boldsymbol{y}_N, (\boldsymbol{Y}_0' \boldsymbol{Y}_0 + \lambda_2 \boldsymbol{I})^{-1} \boldsymbol{Y}_0' \boldsymbol{y}_T \rangle = \sum_{\ell=1}^R \frac{s_\ell}{s_\ell^2 + \lambda_2} \langle \boldsymbol{y}_N, \boldsymbol{v}_\ell \rangle \langle \boldsymbol{u}_\ell, \boldsymbol{y}_T \rangle. \end{align} \end{lemma} \begin{lemma} [Lasso] $\emph{\textsf{HZ}} \neq \emph{\textsf{VT}}$ for lasso regression. \end{lemma} \begin{lemma} [Elastic net] $\emph{\textsf{HZ}} \neq \emph{\textsf{VT}}$ for elastic net regression. \end{lemma} \begin{lemma} [Simplex regression] $\emph{\textsf{HZ}} \neq \emph{\textsf{VT}}$ for simplex regression. \end{lemma} \subsection{Proof of Theorem (ref)} \begin{proof} The proof is immediate from Lemmas (ref)--(ref). \end{proof} \subsection{Proof of Theorem (ref)} \begin{proof} The proof is immediate from Lemmas (ref)--(ref). \end{proof} \subsection{Proof of Corollary (ref)} \begin{proof} We begin with the OLS. Recall that $\widehat{Y}^\text{hz}_{NT}(0) = \langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}} \rangle$ with $\widehat{\boldsymbol{\alpha}} = \boldsymbol{Y}_0^\dagger \boldsymbol{y}_T$ and $\widehat{Y}^\text{vt}_{NT}(0) = \langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle$ with $\widehat{\boldsymbol{\beta}} = (\boldsymbol{Y}_0')^\dagger \boldsymbol{y}_N$. By Theorem (ref), we have \begin{align} \widehat{Y}^\text{hz}_{NT}(0) = \widehat{Y}^\text{vt}_{NT}(0) = \langle \boldsymbol{y}_N, \boldsymbol{Y}_0^\dagger \boldsymbol{y}_T \rangle. \end{align} Returning to (ref), we obtain \begin{align} \widehat{Y}^\text{sdid}_{NT}(0) &= \widehat{Y}^\text{asc}_{NT}(0) = \langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle + \langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}} \rangle - \langle \widehat{\boldsymbol{\alpha}}, \boldsymbol{Y}_0' \widehat{\boldsymbol{\beta}} \rangle = 2 \langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle - \langle \widehat{\boldsymbol{\alpha}}, \boldsymbol{Y}_0' \widehat{\boldsymbol{\beta}} \rangle. \end{align} Recall $(\boldsymbol{Y}_0')^\dagger = (\boldsymbol{Y}_0^\dagger)'$. Therefore, \begin{align} \langle \widehat{\boldsymbol{\alpha}}, \boldsymbol{Y}_0' \widehat{\boldsymbol{\beta}} \rangle &= \boldsymbol{y}'_T(\boldsymbol{Y}_0')^\dagger \boldsymbol{Y}_0' (\boldsymbol{Y}_0')^\dagger \boldsymbol{y}_N = \boldsymbol{y}'_T(\boldsymbol{Y}_0')^\dagger \boldsymbol{y}_N = \boldsymbol{y}'_T \widehat{\boldsymbol{\beta}}. \end{align} Plugging (ref) into (ref), we conclude \begin{align} \widehat{Y}^\text{sdid}_{NT}(0) &= \widehat{Y}^\text{asc}_{NT}(0) = 2 \langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle - \langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle = \widehat{Y}^\text{vt}_{NT}(0) = \widehat{Y}^\text{hz}_{NT}(0). \end{align} Now, observe that the same arguments above hold when $\boldsymbol{Y}_0^{(k)}$ takes the place of $\boldsymbol{Y}_0$ for any $k < R$. Therefore, the same reduction can be derived for PCR. \end{proof} \subsection{Proof of Corollary (ref)} \begin{proof} Let $\boldsymbol{Y}^\text{hz}_0 = [\boldsymbol{1}, \boldsymbol{Y}_0]$ and $\boldsymbol{Y}^\text{vt}_0 = [\boldsymbol{1}, \boldsymbol{Y}'_0]$. The proof is immediate from Theorem (ref) by noting that $(\boldsymbol{Y}^\text{hz}_0)' \neq \boldsymbol{Y}^\text{vt}_0$. \end{proof} \subsection{Proof of Corollary (ref)} \begin{proof} Consider \textsf{HZ} ridge regression. To begin, the optimality conditions give \begin{align} \nabla_{(\alpha_0, \alpha_1, \boldsymbol{\alpha})} \left\{ \| \boldsymbol{y}_T - \boldsymbol{Y}_0 \boldsymbol{\alpha} - \alpha_0 \boldsymbol{1}\|_2^2 + \| \boldsymbol{y}_N - \boldsymbol{\alpha}_1 \boldsymbol{1} \|_2^2 + \lambda \| \boldsymbol{\alpha} \|_2^2 \right\} &= 0. \end{align} Solving for $\alpha_0$, we have $\widehat{\alpha}_0 = (1/N_0) \langle \boldsymbol{y}_T, \boldsymbol{1} \rangle$, where we have used the fact that $\boldsymbol{Y}'_0 \boldsymbol{1} = \boldsymbol{0}$. Solving for $\alpha_1$, we have $\widehat{\alpha}_1 = (1/T_0) \langle \boldsymbol{y}_N, \boldsymbol{1} \rangle$. Finally, solving for $\boldsymbol{\alpha}$, we have $\widehat{\boldsymbol{\alpha}} = (\boldsymbol{Y}'_0 \boldsymbol{Y}_0 + \lambda \boldsymbol{I})^{-1} \boldsymbol{Y}'_0 \boldsymbol{y}_T$. Switching gears to \textsf{VT} ridge regression, similar arguments yield $\widehat{\beta}_0 = (1/T_0) \langle \boldsymbol{y}_N, \boldsymbol{1} \rangle$, $\widehat{\beta}_1 = (1/N_0) \langle \boldsymbol{y}_T, \boldsymbol{1} \rangle$, and $\widehat{\boldsymbol{\beta}} = (\boldsymbol{Y}_0 \boldsymbol{Y}'_0 + \lambda \boldsymbol{I}) \boldsymbol{Y}_0 \boldsymbol{y}_N$. We establish our desired result by invoking Theorem (ref) to obtain $\langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}} \rangle = \langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle$. Next, we observe that $\widehat{\alpha}_0 = \widehat{\beta}_1$ and $\widehat{\beta}_0 = \widehat{\alpha}_1$. This proves that \textsf{HZ} and \textsf{VT} ridge yield numerically identical point estimates. Setting $\lambda = 0$ and using the pseudoinverse, we have our result for OLS. The result for PCR is then established from the OLS result by substituting $\boldsymbol{Y}_0^{(k)}$ for $\boldsymbol{Y}_0$. This completes the proof. \end{proof} \subsection{Proofs of Point Estimation Lemmas} \subsubsection{Proof of Lemma (ref): OLS} \begin{proof} Consider \textsf{HZ} regression. By the optimality conditions, \begin{align} \nabla_{\boldsymbol{\alpha}} \| \boldsymbol{y}_T - \boldsymbol{Y}_0 \boldsymbol{\alpha} \|_2^2 &= 0. \end{align} Solving for $\boldsymbol{\alpha}$, we derive the well-known “normal equations” \begin{align} \boldsymbol{Y}_0' \boldsymbol{Y}_0 \boldsymbol{\alpha} = \boldsymbol{Y}_0' \boldsymbol{y}_T. \end{align} Using the pseudoinverse, we obtain \begin{align} \widehat{\boldsymbol{\alpha}} &= (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \boldsymbol{Y}'_0 \boldsymbol{y}_T = \boldsymbol{Y}_0^\dagger \boldsymbol{y}_T \end{align} Observe that this corresponds to the unique minimum $\ell_2$-norm solution that lies within the rowspace of $\boldsymbol{Y}_0$. Therefore, the \textsf{HZ} prediction is given by \begin{align} \widehat{Y}^\text{hz}_{NT}(0) &= \langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}} \rangle = \langle \boldsymbol{y}_N, \boldsymbol{Y}^\dagger_0 \boldsymbol{y}_T \rangle. \end{align} Following the arguments above for \textsf{VT} regression, it follows that \begin{align} \widehat{\boldsymbol{\beta}} &= (\boldsymbol{Y}_0 \boldsymbol{Y}_0')^\dagger \boldsymbol{Y}_0 \boldsymbol{y}_N = (\boldsymbol{Y}_0')^\dagger \boldsymbol{y}_N, \end{align} which corresponds to the unique minimum $\ell_2$-norm solution that lies within the columnspace of $\boldsymbol{Y}_0$. Therefore, \begin{align} \widehat{Y}^\text{vt}_{NT}(0) &= \langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle = \langle \boldsymbol{y}_T, (\boldsymbol{Y}'_0)^\dagger \boldsymbol{y}_N \rangle. \end{align} Given that $(\boldsymbol{Y}_0')^\dagger = (\boldsymbol{Y}_0^\dagger)'$, we conclude \begin{align} \widehat{Y}^\text{hz}_{NT}(0) &= \langle \boldsymbol{y}_N, \boldsymbol{Y}_0^\dagger \boldsymbol{y}_T \rangle = \langle \boldsymbol{y}_T, (\boldsymbol{Y}_0')^\dagger \boldsymbol{y}_N \rangle = \widehat{Y}^\text{vt}_{NT}(0). \end{align} \end{proof} \subsubsection{Proof of Lemma (ref): PCR} \begin{proof} Consider \textsf{HZ} regression with any $k < R$. Let $\bU_k \in \Rb^{N_0 \times k}$ and $\bV_k \in \Rb^{T_0 \times k}$ denote the matrices formed by the top $k$ left and right singular vectors, respectively, and $\bS_k \in \Rb^{k \times k}$ denote the matrix of top $k$ singular values. Observe that \begin{align} \left((\boldsymbol{Y}_0^{(k)})' \boldsymbol{Y}_0^{(k)} \right)^\dagger (\boldsymbol{Y}_0^{(k)})' &= (\bV_k \bS_k^{-2} \bV'_k) \bV_k \bS_k \bU'_k = \bV_k \bS_k^{-1} \bU'_k = (\boldsymbol{Y}_0^{(k)})^\dagger. \end{align} Therefore, $\widehat{\boldsymbol{\alpha}} = (\boldsymbol{Y}_0^{(k)})^\dagger \boldsymbol{y}_T$, which corresponds to the unique minimum $\ell_2$-norm solution that lies within the rowspace of $\boldsymbol{Y}^{(k)}_0$. Following the proof of Lemma (ref), we conclude that \begin{align} \widehat{Y}^\text{hz}_{NT}(0) &= \langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}} \rangle = \langle \boldsymbol{y}_N, (\boldsymbol{Y}_0^{(k)})^\dagger \boldsymbol{y}_T \rangle. \end{align} Similarly, for \textsf{VT} regression, we note that \begin{align} \left(\boldsymbol{Y}_0^{(k)} (\boldsymbol{Y}_0^{(k)})' \right)^\dagger \boldsymbol{Y}_0^{(k)} &= (\bU_k \bS_k^{-2} \bU'_k) \bU_k \bS_k \bV'_k = \bU_k \bS_k^{-1} \bV'_k = ((\boldsymbol{Y}_0^{(k)})')^\dagger. \end{align} In turn, we have $\widehat{\boldsymbol{\beta}} = ((\boldsymbol{Y}_0^{(k)})')^\dagger \boldsymbol{y}_N$, which corresponds to the unique minimum $\ell_2$-norm solution that lies within the columnspace of $\boldsymbol{Y}^{(k)}_0$. Moreover, \begin{align} \widehat{Y}^\text{vt}_{NT}(0) &= \langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle = \langle \boldsymbol{y}_T, ((\boldsymbol{Y}_0^{(k)})')^\dagger \boldsymbol{y}_N \rangle. \end{align} We finish by establishing \begin{align} \widehat{Y}^\text{hz}_{NT}(0) = \langle \boldsymbol{y}_N, (\boldsymbol{Y}_0^{(k)})^\dagger \boldsymbol{y}_T \rangle = \langle \boldsymbol{y}_T, ((\boldsymbol{Y}_0^{(k)})')^\dagger \boldsymbol{y}_N \rangle = \widehat{Y}^\text{vt}_{NT}(0). \end{align} \end{proof} {\bf Helper lemma.} To establish Lemmas (ref)--(ref), we first establish a general result in Lemma (ref) for $\ell_p$-penalties, where $p = 2/K$ and $K$ is an integer $\ge 1$, based on the contributions of hoff. More formally, consider \begin{itemize} • \textsf{HZ} regression: for $K \ge 1$ and $\lambda > 0$, \begin{align} &\widehat{\boldsymbol{\alpha}} = \operatorname*{\arg\!\min}_{\boldsymbol{\alpha}} \| \boldsymbol{y}_T- \boldsymbol{Y}_0 \boldsymbol{\alpha} \|_2^2 + \lambda \| \boldsymbol{\alpha} \|_p^p \\ &\widehat{Y}_{NT}^\text{hz}(0) =\langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}} \rangle. \end{align} • \textsf{VT} regression: for $K \ge 1$ and $\lambda > 0$, \begin{align} &\widehat{\boldsymbol{\beta}} = \operatorname*{\arg\!\min}_{\boldsymbol{\beta}} \| \boldsymbol{y}_N - \boldsymbol{Y}_0' \boldsymbol{\beta} \|_2^2 + \lambda \| \boldsymbol{\beta} \|_p^p \\ &\widehat{Y}_{NT}^\text{vt}(0) =\langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}} \rangle. \end{align} \end{itemize} We remark that $K=1$ and $K=2$ yield ridge and lasso regression, respectively, while $K > 2$ yields non-convex penalties. We relegate the proof of Lemma (ref) to Appendix (ref). \begin{lemma} For any $K \ge 1$ and $\lambda > 0$, a \emph{\textsf{HZ}} and \emph{\textsf{VT}} regression solution is \begin{align} \widehat{Y}^\emph{hz}_{NT}(0) &=\langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}}_1 \circ \dots \circ \widehat{\boldsymbol{\alpha}}_K \rangle \\ \widehat{Y}^\emph{vt}_{NT}(0) &=\langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}}_1 \circ \dots \circ \widehat{\boldsymbol{\beta}}_K \rangle, \end{align} where for every $k \le K$, \begin{align} \widehat{\boldsymbol{\alpha}}_k &= \left( \bD(\widehat{\boldsymbol{\alpha}}_{\sim k}) \boldsymbol{Y}_0' \boldsymbol{Y}_0 \bD(\widehat{\boldsymbol{\alpha}}_{\sim k}) + \frac{\lambda}{K} \boldsymbol{I} \right)^{-1} \bD(\widehat{\boldsymbol{\alpha}}_{\sim k}) \boldsymbol{Y}_0' \boldsymbol{y}_T, \\ \widehat{\boldsymbol{\beta}}_k &= \left( \bD(\widehat{\boldsymbol{\beta}}_{\sim k}) \boldsymbol{Y}_0 \boldsymbol{Y}_0' \bD(\widehat{\boldsymbol{\beta}}_{\sim k})+ \frac{\lambda}{K} \boldsymbol{I} \right)^{-1} \bD(\widehat{\boldsymbol{\beta}}_{\sim k}) \boldsymbol{Y}_0 \boldsymbol{y}_N, \end{align} $\widehat{\boldsymbol{\alpha}}_{\sim k} = \widehat{\boldsymbol{\alpha}}_1 \circ \dots \circ \widehat{\boldsymbol{\alpha}}_{k-1} \circ \widehat{\boldsymbol{\alpha}}_{k+1} \circ \dots \circ \widehat{\boldsymbol{\alpha}}_K$, $\widehat{\boldsymbol{\beta}}_{\sim k} = \widehat{\boldsymbol{\beta}}_1 \circ \dots \circ \widehat{\boldsymbol{\beta}}_{k-1} \circ \widehat{\boldsymbol{\beta}}_{k+1} \circ \dots \circ \widehat{\boldsymbol{\beta}}_K$, and $\bD(\widehat{\boldsymbol{\alpha}}_{\sim k})$ and $\bD(\widehat{\boldsymbol{\beta}}{\sim k})$ are diagonal matrices formed from $\widehat{\boldsymbol{\alpha}}_{\sim k}$ and $\widehat{\boldsymbol{\beta}}_{\sim k}$, respectively. \end{lemma} \subsubsection{Proof of Lemma (ref): Ridge Regression} \begin{proof} By Lemma (ref) for $K=1$ and $\lambda = \lambda_2 > 0$, the \textsf{HZ} regression solution is \begin{align} \widehat{Y}^\text{hz}_{NT}(0) &=\langle \boldsymbol{y}_N, (\boldsymbol{Y}_0' \boldsymbol{Y}_0 + \lambda_2 \boldsymbol{I})^{-1} \boldsymbol{Y}_0' \boldsymbol{y}_T\rangle. \end{align} Similarly, the \textsf{VT} regression solution is given by \begin{align} \widehat{Y}^\text{vt}_{NT}(0) &=\langle \boldsymbol{y}_T, (\boldsymbol{Y}_0 \boldsymbol{Y}_0' + \lambda_2 \boldsymbol{I})^{-1} \boldsymbol{Y}_0 \boldsymbol{y}_N \rangle. \end{align} Since $(\boldsymbol{Y}_0' \boldsymbol{Y}_0 + \lambda \boldsymbol{I})^{-1} \boldsymbol{Y}_0' = \boldsymbol{Y}_0' (\boldsymbol{Y}_0 \boldsymbol{Y}_0' + \lambda \boldsymbol{I})^{-1}$, it follows that $\widehat{Y}^\text{hz}_{NT}(0) = \widehat{Y}^\text{vt}_{NT}(0)$. \end{proof} \subsubsection{Proof of Lemma (ref): Lasso Regression} \begin{proof} By Lemma (ref) for $K=2$ and $\lambda = \lambda_1 > 0$, a \textsf{HZ} regression solution is \begin{align} \widehat{Y}^\text{hz}_{NT}(0) &=\langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}}_1 \circ \widehat{\boldsymbol{\alpha}}_2 \rangle, \end{align} where \begin{align} \widehat{\boldsymbol{\alpha}}_{1+k} &= \left( \bD(\widehat{\boldsymbol{\alpha}}_{2-k}) \boldsymbol{Y}_0' \boldsymbol{Y}_0 \bD(\widehat{\boldsymbol{\alpha}}_{2-k}) + \frac{\lambda_1}{2} \boldsymbol{I} \right)^{-1} \bD(\widehat{\boldsymbol{\alpha}}_{2-k}) \boldsymbol{Y}_0' \boldsymbol{y}_T \end{align} for $k \in \{0,1\}$. Similarly, a \textsf{VT} regression solution is given by \begin{align} \widehat{Y}^\text{vt}_{NT}(0) &=\langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}}_1 \circ \widehat{\boldsymbol{\beta}}_2 \rangle, \end{align} where \begin{align} \widehat{\boldsymbol{\beta}}_{1+k} &= \left( \bD(\widehat{\boldsymbol{\beta}}_{2-k}) \boldsymbol{Y}_0 \boldsymbol{Y}_0' \bD(\widehat{\boldsymbol{\beta}}_{2-k}) + \frac{\lambda_1}{2} \boldsymbol{I} \right)^{-1} \bD(\widehat{\boldsymbol{\beta}}_{2-k}) \boldsymbol{Y}_0 \boldsymbol{y}_N \end{align} for $k \in \{0,1\}$. Leveraging (ref) and (ref), we find that the \textsf{HZ} regression solution can be linear in $y$ and at least quadratic in $q$. On the other hand, the \textsf{VT} regression solution can be linear in $q$ and at least quadratic in $y$. Since the lasso solution is unique under the assumption the entries of $\boldsymbol{Y}_0$ are drawn from a continuous distribution, this implies that \textsf{HZ} and \textsf{VT} regressions do not yield matching solutions in general. \end{proof} \subsubsection{Proof of Lemma (ref): Elastic Net Regression} \begin{proof} Consider \textsf{HZ} regression. We rewrite (ref) in a lasso formulation: \begin{align} \widehat{\boldsymbol{\alpha}}^* &= \operatorname*{\arg\!\min}_{\boldsymbol{\alpha}^*} \| \boldsymbol{y}_T^* - \boldsymbol{Y}_0^* \boldsymbol{\alpha}^* \|_2^2 + \lambda^* \| \boldsymbol{\alpha}^* \|_1, \end{align} where \begin{align} &q^* = \begin{pmatrix} \boldsymbol{y}_T \\ \boldsymbol{0} \end{pmatrix}, \quad \boldsymbol{Y}_0^* = \frac{1}{\sqrt{1 + \lambda_2}} \begin{pmatrix} \boldsymbol{Y}_0 \\ \sqrt{\lambda_2} \boldsymbol{I} \end{pmatrix}, \quad \lambda^* = \frac{\lambda_1}{\sqrt{1 + \lambda_2}}, \quad \boldsymbol{\alpha}^* = (\sqrt{1 + \lambda_2}) \boldsymbol{\alpha}. \end{align} We apply Lemma (ref) to (ref) with $K=2$ and $\lambda = \lambda^* > 0$ to obtain \begin{align} \widehat{Y}^\text{hz}_{NT}(0) &= \frac{\langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}}^*_1 \circ \widehat{\boldsymbol{\alpha}}^*_2 \rangle}{ \sqrt{1 + \lambda_2}}, \end{align} where \begin{align} \widehat{\boldsymbol{\alpha}}^*_{1+k} &= \left( \frac{1}{\sqrt{1 + \lambda_2}} \bD(\widehat{\boldsymbol{\alpha}}^*_{2-k}) (\boldsymbol{Y}_0' \boldsymbol{Y}_0 + \lambda_2 \boldsymbol{I}) \bD(\widehat{\boldsymbol{\alpha}}^*_{2-k}) + \frac{\lambda_1}{2} \boldsymbol{I} \right)^{-1} \bD(\widehat{\boldsymbol{\alpha}}^*_{2-k}) \boldsymbol{Y}_0' \boldsymbol{y}_T \end{align} for $k \in \{0,1\}$. Similarly, for \textsf{VT} regression, we proceed as above to obtain \begin{align} \widehat{Y}^\text{vt}_{NT}(0) &= \frac{\langle \boldsymbol{y}_T, \widehat{\boldsymbol{\beta}}^*_1 \circ \widehat{\boldsymbol{\beta}}^*_2 \rangle}{\sqrt{1 + \lambda_2}}, \end{align} where \begin{align} \widehat{\boldsymbol{\beta}}^*_{1+k} &= \left( \frac{1}{\sqrt{1 + \lambda_2}} \bD(\widehat{\boldsymbol{\beta}}^*_{2-k}) (\boldsymbol{Y}_0 \boldsymbol{Y}_0' + \lambda_2 \boldsymbol{I}) \bD(\widehat{\boldsymbol{\beta}}^*_{2-k}) + \frac{\lambda_1}{2} \boldsymbol{I} \right)^{-1} \bD(\widehat{\boldsymbol{\beta}}^*_{2-k}) \boldsymbol{Y}_0 \boldsymbol{y}_N \end{align} for $k \in \{0,1\}$. Leveraging (ref) and (ref), we find that the \textsf{HZ} regression solution can be linear in $y$ and at least quadratic in $q$. On the other hand, the \textsf{VT} regression solution can be linear in $q$ and at least quadratic in $y$. Since the elastic net regression solution is unique, provided $\lambda_2 > 0$, this implies that \textsf{HZ} and \textsf{VT} regressions do not yield matching solutions in general. \end{proof} \subsubsection{Proof of Lemma (ref): Simplex Regression} \begin{proof} Consider \textsf{HZ} regression. We write the Lagrangian of (ref) as \begin{align} \widehat{\boldsymbol{\alpha}} &= \operatorname*{\arg\!\min}_{\boldsymbol{\alpha}} \| \boldsymbol{y}_T - \boldsymbol{Y}_0 \boldsymbol{\alpha} \|_2^2 + \lambda \| \boldsymbol{\alpha} \|_2^2 - (\boldsymbol{\theta}^\text{hz})' \boldsymbol{\alpha} + \nu^\text{hz} (\boldsymbol{1}' \boldsymbol{\alpha} - 1), \end{align} where $\boldsymbol{\theta}^\text{hz} \in \Rb^{T_0}$ and $\nu^\text{hz} \in \Rb$. By the Karush-Kuhn-Tucker (KKT) conditions, optimality is achieved if the following are satisfied: \begin{align} &\widehat{\boldsymbol{\alpha}} \succeq \boldsymbol{0}, \quad \boldsymbol{1}' \widehat{\boldsymbol{\alpha}} = 1, \\ & \widehat{\boldsymbol{\theta}}^\text{hz} \succeq \boldsymbol{0}, \\ & \widehat{\theta}^\text{hz}_i \widehat{\alpha}_i = 0 \quad \text{for } i=1, \dots, T_0, \\ & \widehat{\boldsymbol{\alpha}} = (\boldsymbol{Y}_0' \boldsymbol{Y}_0 + \lambda \boldsymbol{I})^{-1} \left( \boldsymbol{Y}_0' \boldsymbol{y}_T + \frac{1}{2} \widehat{\boldsymbol{\theta}}^\text{hz} - \frac{\widehat{\nu}^\text{hz}}{2} \boldsymbol{1} \right). \end{align} Therefore, given primal and dual feasible variables $(\widehat{\boldsymbol{\alpha}}, \widehat{\boldsymbol{\theta}}^\text{hz}, \widehat{\nu}^\text{hz})$, we can write the final \textsf{HZ} prediction as \begin{align} \widehat{Y}^\text{hz}_{NT}(0) &= \widehat{Y}^{\text{hz}, \text{ols}}_{NT}(0) + (1/2) \boldsymbol{y}'_N (\boldsymbol{Y}_0' \boldsymbol{Y}_0 + \lambda \boldsymbol{I})^{-1} (\widehat{\boldsymbol{\theta}}^\text{hz} - \widehat{\nu}^\text{hz} \boldsymbol{1} ), \end{align} where $\widehat{Y}^{\text{hz}, \text{ols}}_{NT}(0) = \boldsymbol{y}'_N (\boldsymbol{Y}_0' \boldsymbol{Y}_0 + \lambda \boldsymbol{I})^{-1} \boldsymbol{Y}_0' \boldsymbol{y}_T$ converges to the prediction corresponding to the OLS solution with minimum $\ell_2$-norm as $\lambda \rightarrow 0^+$. Similarly, for \textsf{VT} regression, the KKT conditions are \begin{align} &\widehat{\boldsymbol{\beta}} \succeq \boldsymbol{0}, \quad \boldsymbol{1}' \widehat{\boldsymbol{\beta}} = 1, \\ & \widehat{\boldsymbol{\theta}}^\text{vt} \succeq \boldsymbol{0}, \\ & \widehat{\theta}^\text{vt}_i \widehat{\beta}_i = 0 \quad \text{for } i=1, \dots, N_0, \\ & \widehat{\boldsymbol{\beta}} = (\boldsymbol{Y}_0 \boldsymbol{Y}_0' + \lambda \boldsymbol{I})^{-1} \left(\boldsymbol{Y}_0 \boldsymbol{y}_N + \frac{1}{2} \widehat{\boldsymbol{\theta}}^\text{vt} - \frac{\widehat{\nu}^\text{vt}}{2} \boldsymbol{1}\right). \end{align} For primal and dual feasible variables $(\widehat{\boldsymbol{\beta}}, \widehat{\boldsymbol{\theta}}^\text{vt}, \widehat{\nu}^\text{vt})$, this yields \begin{align} \widehat{Y}^\text{vt}_{NT}(0) &= \widehat{Y}^{\text{vt}, \text{ols}}_{NT}(0) + (1/2) \boldsymbol{y}'_T (\boldsymbol{Y}_0 \boldsymbol{Y}_0' + \lambda \boldsymbol{I})^{-1} (\widehat{\boldsymbol{\theta}}^\text{vt} - \widehat{\nu}^\text{vt} \boldsymbol{1} ), \end{align} where $\widehat{Y}^{\text{vt}, \text{ols}}_{NT}(0) = \boldsymbol{y}'_T (\boldsymbol{Y}_0 \boldsymbol{Y}_0' + \lambda \boldsymbol{I})^{-1} \boldsymbol{Y}_0 \boldsymbol{y}_N$ converges to the prediction corresponding to the OLS solution with minimum $\ell_2$-norm as $\lambda \rightarrow 0^+$. Notably, as per Theorem (ref), $\widehat{Y}^{\text{hz}, \text{ols}}_{NT}(0) = \widehat{Y}^{\text{vt}, \text{ols}}_{NT}(0) = \widehat{Y}^{\text{ols}}_{NT}(0)$ for any $\lambda \ge 0$. As a result, \begin{align} \widehat{Y}^\text{hz}_{NT}(0) &= \widehat{Y}^{\text{ols}}_{NT}(0) + (1/2) \boldsymbol{y}'_N (\boldsymbol{Y}_0' \boldsymbol{Y}_0 + \lambda \boldsymbol{I})^{-1} (\widehat{\boldsymbol{\theta}}^\text{hz} - \widehat{\nu}^\text{hz} \boldsymbol{1} ) \\ \widehat{Y}^\text{vt}_{NT}(0) &= \widehat{Y}^{\text{ols}}_{NT}(0) + (1/2) \boldsymbol{y}'_T (\boldsymbol{Y}_0 \boldsymbol{Y}_0' + \lambda \boldsymbol{I})^{-1} (\widehat{\boldsymbol{\theta}}^\text{vt} - \widehat{\nu}^\text{vt} \boldsymbol{1} ). \end{align} As seen from (ref) and (ref), the leading terms in the \textsf{HZ} and \textsf{VT} simplex regression predictions are identical. The remaining terms, however, can differ from one another. As an example, consider $N = T$ with \begin{align} \boldsymbol{Y}_0 = \boldsymbol{I}, \quad \boldsymbol{y}_N = \boldsymbol{0}, \quad \boldsymbol{y}_T = (1+ \lambda) (\widehat{\boldsymbol{\theta}}^\text{vt} - \widehat{\nu}^\text{vt} \boldsymbol{1}). \end{align} By construction, observe that \begin{align} \widehat{\boldsymbol{\beta}} = \frac{1}{2(1+ \lambda)} (\widehat{\boldsymbol{\theta}}^\text{vt} - \widehat{\nu}^\text{vt} \boldsymbol{1}). \end{align} Recall from the KKT conditions for \textsf{VT} regression that $\widehat{\boldsymbol{\beta}} \succeq \boldsymbol{0}$ and $\boldsymbol{1}' \widehat{\boldsymbol{\beta}} = 1$. Therefore, at least one entry of $(\widehat{\boldsymbol{\theta}}^\text{vt} - \widehat{\nu}^\text{vt} \boldsymbol{1})$ must be strictly positive. This yields \begin{align} (1+ \lambda)^{-1} \boldsymbol{y}'_T (\widehat{\boldsymbol{\theta}}^\text{vt} - \widehat{\nu}^\text{vt} \boldsymbol{1}) &= (\widehat{\boldsymbol{\theta}}^\text{vt} - \widehat{\nu}^\text{vt} \boldsymbol{1})' (\widehat{\boldsymbol{\theta}}^\text{vt} - \widehat{\nu}^\text{vt} \boldsymbol{1}) > 0. \end{align} Plugging (ref) and (ref) into (ref) and (ref), we obtain \begin{align} \widehat{Y}^\text{hz}_{NT}(0) &= 0 \quad \text{and} \quad \widehat{Y}^\text{vt}_{NT}(0) > 0, \end{align} which concludes our proof. \end{proof} \subsubsection{Proof of Lemma (ref): $\ell_p$-penalties} \begin{proof} We recall the Hadamard product parametrization (HPP): for any vector $\boldsymbol{z}$ and integer $K \ge 1$, \begin{align} \| \boldsymbol{z} \|_p^p &= \min_{\boldsymbol{z}_1\circ \dots \circ \boldsymbol{z}_K = \boldsymbol{z}} \frac{1}{K} \sum_{k=1}^K \| \boldsymbol{z}_k \|_2^2, \end{align} where $\circ$ denotes the Hadamard (componentwise) product. We rewrite our subclass of $\ell_p$-penalties, i.e., (ref) and (ref), as sums of $\ell_2$-penalties via the HPP technique: \begin{align} (\widehat{\boldsymbol{\alpha}}_1, \dots, \widehat{\boldsymbol{\alpha}}_K) &= \operatorname*{\arg\!\min}_{\boldsymbol{\alpha}_1, \dots, \boldsymbol{\alpha}_K} \| \boldsymbol{y}_T- \boldsymbol{Y}_0 (\boldsymbol{\alpha}_1 \circ \dots \circ \boldsymbol{\alpha}_K) \|_2^2 + \frac{\lambda}{K} \sum_{k=1}^K \| \boldsymbol{\alpha}_k \|_2^2 \\ (\widehat{\boldsymbol{\beta}}_1, \dots, \widehat{\boldsymbol{\beta}}_K) &= \operatorname*{\arg\!\min}_{\boldsymbol{\beta}_1, \dots, \boldsymbol{\beta}_K} \| \boldsymbol{y}_N - \boldsymbol{Y}_0' (\boldsymbol{\beta}_1 \circ \dots \circ \boldsymbol{\beta}_K) \|_2^2 + \frac{\lambda}{K} \sum_{k=1}^K \| \boldsymbol{\beta}_k \|_2^2, \end{align} where $\widehat{\boldsymbol{\alpha}} = \widehat{\boldsymbol{\alpha}}_1 \circ \dots \circ \widehat{\boldsymbol{\alpha}}_K$ and $\widehat{\boldsymbol{\beta}} = \widehat{\boldsymbol{\beta}}_1 \circ \dots \circ \widehat{\boldsymbol{\beta}}_K$. Below, we leverage the results of hoff, which provides an alternating ridge regression algorithm to solve for (ref)--(ref). Consider \textsf{HZ} regression. Let us solve for $\boldsymbol{\alpha}_k$ for $k \in [K]$ by fixing $\boldsymbol{\alpha}_{k'}$ for $k' \neq k$. By the optimality conditions, \begin{align} \nabla_{\boldsymbol{\alpha}_k} \left\{ (\boldsymbol{\alpha}_1 \circ \dots \circ \boldsymbol{\alpha}_K)' \boldsymbol{Y}_0' \boldsymbol{Y}_0 (\boldsymbol{\alpha}_1 \circ \dots \circ \boldsymbol{\alpha}_K) - 2 (\boldsymbol{\alpha}_1 \circ \dots \circ \boldsymbol{\alpha}_K)' \boldsymbol{Y}_0' \boldsymbol{y}_T+ \frac{\lambda}{K} \boldsymbol{\alpha}_k' \boldsymbol{\alpha}_k \right\}= 0. \end{align} In order to solve for (ref), observe that \begin{align} & (\boldsymbol{\alpha}_1 \circ \dots \circ \boldsymbol{\alpha}_K)' \boldsymbol{Y}_0' \boldsymbol{Y}_0 (\boldsymbol{\alpha}_1 \circ \dots \circ \boldsymbol{\alpha}_K) = \boldsymbol{\alpha}_k' (\boldsymbol{Y}_0' \boldsymbol{Y}_0 \circ \boldsymbol{\alpha}_{\sim k} \boldsymbol{\alpha}_{\sim k}') \boldsymbol{\alpha}_k \\ & (\boldsymbol{\alpha}_1 \circ \dots \circ \boldsymbol{\alpha}_K)' \boldsymbol{Y}_0' \boldsymbol{y}_T= \boldsymbol{\alpha}_k' (\boldsymbol{\alpha}_{\sim k} \circ \boldsymbol{Y}_0' \boldsymbol{y}_T), \end{align} where $\boldsymbol{\alpha}_{\sim k} = \boldsymbol{\alpha}_1 \circ \dots \circ \boldsymbol{\alpha}_{k-1} \circ \boldsymbol{\alpha}_{k+1} \circ \dots \circ \boldsymbol{\alpha}_K$. This allows us to rewrite (ref) as \begin{align} \nabla_{\boldsymbol{\alpha}_k} \left\{ \boldsymbol{\alpha}_k' \left(\boldsymbol{Y}_0' \boldsymbol{Y}_0 \circ \boldsymbol{\alpha}_{\sim k} \boldsymbol{\alpha}_{\sim k}' + \frac{\lambda}{K} \boldsymbol{I} \right) \boldsymbol{\alpha}_k - 2 \boldsymbol{\alpha}_k' (\boldsymbol{\alpha}_{\sim k} \circ \boldsymbol{Y}_0' \boldsymbol{y}_T) \right\} = 0. \end{align} This is quadratic in $\boldsymbol{\alpha}_k$ for fixed $\boldsymbol{\alpha}_{\sim k}$. Thus, the unique minimizer at convergence is \begin{align} \widehat{\boldsymbol{\alpha}}_k &= \left(\boldsymbol{Y}_0' \boldsymbol{Y}_0 \circ \widehat{\boldsymbol{\alpha}}_{\sim k} \widehat{\boldsymbol{\alpha}}_{\sim k}' + \frac{\lambda}{K} \boldsymbol{I} \right)^{-1} (\widehat{\boldsymbol{\alpha}}_{\sim k} \circ \boldsymbol{Y}_0' \boldsymbol{y}_T), \end{align} where $\widehat{\boldsymbol{\alpha}}_{\sim k} = \widehat{\boldsymbol{\alpha}}_1 \circ \dots \circ \widehat{\boldsymbol{\alpha}}_{k-1} \circ \widehat{\boldsymbol{\alpha}}_{k+1} \circ \dots \circ \widehat{\boldsymbol{\alpha}}_K$. Leveraging properties of the Hadamard product noted in styan, we rewrite \begin{align} \boldsymbol{Y}_0' \boldsymbol{Y}_0 \circ \widehat{\boldsymbol{\alpha}}_{\sim k} \widehat{\boldsymbol{\alpha}}_{\sim k}' &= \bD(\widehat{\boldsymbol{\alpha}}_{\sim k}) \boldsymbol{Y}_0' \boldsymbol{Y}_0 \bD(\widehat{\boldsymbol{\alpha}}_{\sim k}) \\ \boldsymbol{Y}_0' \boldsymbol{y}_T\circ \widehat{\boldsymbol{\alpha}}_{\sim k} &= \bD(\widehat{\boldsymbol{\alpha}}_{\sim k}) \boldsymbol{Y}_0' \boldsymbol{y}_T, \end{align} where $\bD(\widehat{\boldsymbol{\alpha}}_{\sim k})$ is the diagonal matrix formed from $\widehat{\boldsymbol{\alpha}}_{\sim k}$. Leveraging these equalities, we simplify (ref) as \begin{align} \widehat{\boldsymbol{\alpha}}_k &= \Big( \bD(\widehat{\boldsymbol{\alpha}}_{\sim k}) \boldsymbol{Y}_0' \boldsymbol{Y}_0 \bD(\widehat{\boldsymbol{\alpha}}_{\sim k}) + \frac{\lambda}{K} \boldsymbol{I} \Big)^{-1} \bD(\widehat{\boldsymbol{\alpha}}_{\sim k}) \boldsymbol{Y}_0' \boldsymbol{y}_T. \end{align} We now turn to \textsf{VT} regression. Following the arguments above, for every $k \in [K]$, \begin{align} \widehat{\boldsymbol{\beta}}_k &= \Big( \bD(\widehat{\boldsymbol{\beta}}_{\sim k}) \boldsymbol{Y}_0 \boldsymbol{Y}_0' \bD(\widehat{\boldsymbol{\beta}}_{\sim k}) + \frac{\lambda}{K} \boldsymbol{I} \Big)^{-1} \bD(\widehat{\boldsymbol{\beta}}_{\sim k}) \boldsymbol{Y}_0 \boldsymbol{y}_N, \end{align} where $\widehat{\boldsymbol{\beta}}_{\sim k} = \widehat{\boldsymbol{\beta}}_1 \circ \dots \circ \widehat{\boldsymbol{\beta}}_{k-1} \circ \widehat{\boldsymbol{\beta}}_{k+1} \circ \dots \circ \widehat{\boldsymbol{\beta}}_K$ and $\bD(\widehat{\boldsymbol{\beta}}_{\sim k})$ is the diagonal matrix formed from $\widehat{\boldsymbol{\beta}}_{\sim k}$. This completes the proof. \end{proof} \section{Proofs for Inference} \subsection{Proof of Theorem (ref)} To establish Theorem (ref), we first state a few useful results. \begin{lemma} [Theorem 2.7.1 of lehmann] Let $X_i$ for $i=1, \dots, n$ be independently distributed with means $\mathbb{E}[X_i] = \zeta_i$ and variances $\sigma^2_i$, and with finite third moments. Let $\bar{X} = (1/n) \sum_{i=1}^n X_i$. Then \begin{align} \frac{\bar{X} - \mathbb{E}[\bar{X}]}{\operatorname{Var}(\bar{X})^{1/2}} \xrightarrow{d} \mathcal{N}(0,1), \end{align} provided \begin{align} \left( \sum_{i=1}^n \mathbb{E} \left[ | X_i - \zeta_i |^3 \right] \right)^2 = o\left( \Big( \sum_{i=1}^n \sigma_i^2 \Big)^3 \right). \end{align} \end{lemma} \begin{lemma} Consider a random vector $\boldsymbol{x}$ and random matrix $\bA$. Let $\mathbb{E}[\boldsymbol{x} | \bA] = \boldsymbol{0}$ and $\operatorname{Cov}(\boldsymbol{x} | \bA) = \bSigma$. Then $\mathbb{E}[ \boldsymbol{x}' \bA \boldsymbol{x} | \bA] = \operatorname{tr}(\bA \bSigma)$. \end{lemma} \begin{proof} (i) [\textsf{HZ} model] Let Assumptions (ref)--(ref) hold. By (ref), Lemma (ref) yields \begin{align} \frac{ \widehat{Y}_{NT}(0) - \mathbb{E}[\widehat{Y}_{NT}(0) |\boldsymbol{y}_N, \boldsymbol{Y}_0] }{\operatorname{Var}(\widehat{Y}_{NT}(0) | \boldsymbol{y}_N, \boldsymbol{Y}_0)^{1/2}} \xrightarrow{d} \mathcal{N}(0,1). \end{align} To evaluate $\mathbb{E}[\widehat{Y}_{NT}(0) | \boldsymbol{y}_N, \boldsymbol{Y}_0]$, we first observe that \begin{align} \mathbb{E}[\widehat{Y}_{NT}(0) | \boldsymbol{y}_N, \boldsymbol{Y}_0] &= \mathbb{E}[ \langle \boldsymbol{y}_N, \widehat{\boldsymbol{\alpha}} \rangle | \boldsymbol{y}_N, \boldsymbol{Y}_0] \&= \mathbb{E}[ \langle \boldsymbol{y}_N, \boldsymbol{Y}_0^\dagger \boldsymbol{y}_T \rangle | \boldsymbol{y}_N, \boldsymbol{Y}_0] \&= \boldsymbol{y}'_N \boldsymbol{Y}_0^\dagger \mathbb{E}[ \boldsymbol{y}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0] \&= \boldsymbol{y}'_N \boldsymbol{Y}_0^\dagger \boldsymbol{Y}_0 \boldsymbol{\alpha}^* \&= \boldsymbol{y}'_N \bH^v \boldsymbol{\alpha}^*. \end{align} Moving to the variance term, we note that \begin{align} \operatorname{Var}(\widehat{Y}_{NT}(0) |\boldsymbol{y}_N, \boldsymbol{Y}_0) &= \boldsymbol{y}'_N \operatorname{Cov}(\widehat{\boldsymbol{\alpha}} | \boldsymbol{y}_N, \boldsymbol{Y}_0) \boldsymbol{y}_N. \end{align} Towards evaluating the above, we note that \begin{align} \operatorname{Cov}(\widehat{\boldsymbol{\alpha}} | \boldsymbol{y}_N, \boldsymbol{Y}_0) &= \operatorname{Cov}( \boldsymbol{Y}_0^\dagger \boldsymbol{y}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0) \&= \boldsymbol{Y}_0^\dagger \operatorname{Cov}(\boldsymbol{y}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0) (\boldsymbol{Y}'_0)^\dagger \&= \boldsymbol{Y}_0^\dagger \operatorname{Cov}(\boldsymbol{\varepsilon}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0) (\boldsymbol{Y}'_0)^\dagger \&= \boldsymbol{Y}_0^\dagger \bSigma^\text{hz}_T (\boldsymbol{Y}'_0)^\dagger. \end{align} Plugging (ref) into (ref), we obtain \begin{align} \operatorname{Var}(\widehat{Y}_{NT}(0) |\boldsymbol{y}_N, \boldsymbol{Y}_0) = \boldsymbol{y}_N' \boldsymbol{Y}_0^\dagger \bSigma^\text{hz}_T (\boldsymbol{Y}'_0)^\dagger \boldsymbol{y}_N = \widehat{\boldsymbol{\beta}}' \bSigma^\text{hz}_T \widehat{\boldsymbol{\beta}}, \end{align} where we recall that $\widehat{\boldsymbol{\beta}} = (\boldsymbol{Y}_0')^\dagger \boldsymbol{y}_N$. Putting it all together, we conclude \begin{align} \frac{ \widehat{Y}_{NT}(0) - \langle \boldsymbol{y}_N, \bH^v \boldsymbol{\alpha}^* \rangle }{(\widehat{\boldsymbol{\beta}}' \bSigma^\text{hz}_T \widehat{\boldsymbol{\beta}})^{1/2}} \xrightarrow{d} \mathcal{N}(0,1). \end{align} (ii) [\textsf{VT} model] Let Assumptions (ref)--(ref) hold. Following the arguments above, we have \begin{align} \frac{ \widehat{Y}_{NT}(0) - \langle \boldsymbol{y}_T, \bH^u \boldsymbol{\beta}^* \rangle }{(\widehat{\boldsymbol{\alpha}}' \bSigma^\text{vt}_N \widehat{\boldsymbol{\alpha}})^{1/2}} \xrightarrow{d} \mathcal{N}(0,1). \end{align} (iii) [Mixed model] Let Assumptions (ref)--(ref) hold. We will find it useful to write \begin{align} \widehat{Y}_{NT}(0) &= \boldsymbol{y}'_N \boldsymbol{Y}_0^\dagger \boldsymbol{y}_T = \sum_{i \le N_0} \sum_{t \le T_0} (\boldsymbol{Y}_0^\dagger)_{it} Y_{iT} Y_{Nt}. \end{align} By Assumption (ref), (ref) is a sum of independent random variables with \begin{align} \mathbb{E}[Y_{iT} Y_{Nt} | \boldsymbol{Y}_0] &= \mathbb{E}[Y_{iT} | \boldsymbol{Y}_0] \mathbb{E}[Y_{Nt} | \boldsymbol{Y}_0] \\ \operatorname{Var}(Y_{iT} Y_{Nt} | \boldsymbol{Y}_0) &= \mathbb{E}[Y_{iT} | \boldsymbol{Y}_0]^2 \sigma^2_{Nt} + \mathbb{E}[Y_{Nt} | \boldsymbol{Y}_0]^2 \sigma^2_{iT} + \sigma^2_{iT} \sigma^2_{Nt}. \end{align} Lemma (ref) then establishes that \begin{align} \frac{ \widehat{Y}_{NT}(0) - \mathbb{E}[ \widehat{Y}_{NT}(0) | \boldsymbol{Y}_0]}{ \operatorname{Var}(\widehat{Y}_{NT}(0) | \boldsymbol{Y}_0)^{1/2}} \xrightarrow{d} \mathcal{N}(0,1). \end{align} Our aim is to evaluate $\mathbb{E}[\widehat{Y}_{NT}(0) | \boldsymbol{Y}_0]$ and $\operatorname{Var}(\widehat{Y}_{NT}(0) | \boldsymbol{Y}_0)$. Towards the former, we use Assumptions (ref)--(ref) with the law of total expectation to obtain \begin{align} \mathbb{E}[\widehat{Y}_{NT}(0) | \boldsymbol{Y}_0] &= \mathbb{E} \left[ \mathbb{E}[ \langle \boldsymbol{y}_N, \boldsymbol{Y}_0^\dagger \boldsymbol{y}_T \rangle | \boldsymbol{\varepsilon}_N, \boldsymbol{Y}_0] | \boldsymbol{Y}_0 \right] \\ &= \mathbb{E} \left[ \mathbb{E}[ \boldsymbol{y}_N' \boldsymbol{Y}_0^\dagger (\boldsymbol{Y}_0 \boldsymbol{\alpha}^* + \boldsymbol{\varepsilon}_T) | \boldsymbol{\varepsilon}_N, \boldsymbol{Y}_0] | \boldsymbol{Y}_0 \right] \\ &= \mathbb{E} \left[ (\boldsymbol{Y}'_0 \boldsymbol{\beta}^* + \boldsymbol{\varepsilon}_N)' \boldsymbol{Y}_0^\dagger \boldsymbol{Y}_0 \boldsymbol{\alpha}^* | \boldsymbol{Y}_0 \right] \\ &= \langle \boldsymbol{\beta}^*, \boldsymbol{Y}_0 \boldsymbol{\alpha}^* \rangle. \end{align} Note that we have used the fact that $\boldsymbol{y}_N$ is deterministic given $(\boldsymbol{\varepsilon}_N, \boldsymbol{Y}_0)$. Similarly, by the law of total variance, \begin{align} \operatorname{Var}(\widehat{Y}_{NT}(0) | \boldsymbol{Y}_0) &= \mathbb{E}[ \operatorname{Var}(\widehat{Y}_{NT}(0) | \boldsymbol{\varepsilon}_N, \boldsymbol{Y}_0) | \boldsymbol{Y}_0 ] + \operatorname{Var}( \mathbb{E}[\widehat{Y}_{NT}(0) | \boldsymbol{\varepsilon}_N, \boldsymbol{Y}_0] | \boldsymbol{Y}_0 ). \end{align} Following the derivation of (ref), we have \begin{align} &\mathbb{E}[ \operatorname{Var}(\widehat{Y}_{NT}(0) | \boldsymbol{\varepsilon}_N, \boldsymbol{Y}_0) | \boldsymbol{Y}_0] \\ &\qquad \qquad = \mathbb{E}[ \boldsymbol{y}'_N \boldsymbol{Y}_0^\dagger \bSigma^\text{mix}_T (\boldsymbol{Y}'_0)^\dagger \boldsymbol{y}_N | \boldsymbol{Y}_0] \\ &\qquad \qquad = (\boldsymbol{Y}'_0 \boldsymbol{\beta}^*)' \bA (\boldsymbol{Y}'_0 \boldsymbol{\beta}^*) + \mathbb{E}[ \boldsymbol{\varepsilon}'_N \bA \boldsymbol{\varepsilon}_N | \boldsymbol{Y}_0] + 2 \mathbb{E}[ \boldsymbol{\varepsilon}'_N \boldsymbol{Y}'_0 \boldsymbol{\beta}^* | \boldsymbol{Y}_0], \end{align} where $\bA = \boldsymbol{Y}^\dagger_0 \bSigma^\text{mix}_T (\boldsymbol{Y}'_0)^\dagger$. Notice that Assumption (ref) gives $\mathbb{E}[\boldsymbol{\varepsilon}'_N \boldsymbol{Y}'_0 \boldsymbol{\beta}^* | \boldsymbol{Y}_0] = 0$. Since $\bA$ is deterministic given $\boldsymbol{Y}_0$, Lemma (ref) yields \begin{align} \mathbb{E}[ \boldsymbol{\varepsilon}'_N \bA \boldsymbol{\varepsilon}_N | \boldsymbol{Y}_0] &= \operatorname{tr}( \bA \bSigma^\text{mix}_N). \end{align} Following the arguments that led to the derivation of (ref), we have \begin{align} \operatorname{Var}( \mathbb{E}[\widehat{Y}_{NT}(0) | \boldsymbol{\varepsilon}_N, \boldsymbol{Y}_0] | \boldsymbol{Y}_0 ) &= \operatorname{Var}( \boldsymbol{y}'_N \bH^v \boldsymbol{\alpha}^* | \boldsymbol{Y}_0) = (\bH^v \boldsymbol{\alpha}^*)' \bSigma^\text{mix}_N (\bH^v \boldsymbol{\alpha}^*). \end{align} Plugging (ref), (ref), and (ref) into (ref), we arrive at \begin{align} &\operatorname{Var}(\widehat{Y}_{NT}(0) | \boldsymbol{Y}_0) \\ &\qquad = (\bH^v \boldsymbol{\alpha}^*)' \bSigma^\text{mix}_N (\bH^v \boldsymbol{\alpha}^*) + (\bH^u \boldsymbol{\beta}^*)' \bSigma^\text{mix}_T (\bH^u \boldsymbol{\beta}^*) + \operatorname{tr}( \boldsymbol{Y}^\dagger_0 \bSigma^\text{mix}_T (\boldsymbol{Y}'_0)^\dagger \bSigma^\text{mix}_N). \end{align} This completes the proof. \end{proof} \subsection{Proofs for Model-Based Confidence Intervals} We first state a useful lemma to prove Lemmas (ref)--(ref). \begin{lemma} \emph{[Mixed model]} Let Assumptions (ref)--(ref) hold. Then, \begin{align} \mathbb{E}[ \widehat{v}^\emph{mix}_0 | \boldsymbol{Y}_0] &= (\bH^u \boldsymbol{\beta}^*)' \mathbb{E}[ \widehat{\bSigma}_T | \boldsymbol{Y}_0] (\bH^u \boldsymbol{\beta}^*) + \operatorname{tr}( \boldsymbol{Y}_0^\dagger \mathbb{E}[ \widehat{\bSigma}_T | \boldsymbol{Y}_0] (\boldsymbol{Y}'_0)^\dagger \bSigma^\emph{mix}_N) \\ &\quad + (\bH^v \boldsymbol{\alpha}^*)' \mathbb{E}[ \widehat{\bSigma}_N | \boldsymbol{Y}_0] (\bH^v \boldsymbol{\alpha}^* ) + \operatorname{tr}( \boldsymbol{Y}_0^\dagger \bSigma^\emph{mix}_T (\boldsymbol{Y}'_0)^\dagger \mathbb{E}[\widehat{\bSigma}_N | \boldsymbol{Y}_0]) \\ &\quad - \operatorname{tr}( \boldsymbol{Y}_0^\dagger \mathbb{E}[ \widehat{\bSigma}_T | \boldsymbol{Y}_0] (\boldsymbol{Y}'_0)^\dagger \mathbb{E}[\widehat{\bSigma}_N | \boldsymbol{Y}_0]). \end{align} \end{lemma} \subsubsection{Proof of Lemma (ref)} \begin{proof} (i) [\textsf{HZ} model] Let Assumptions (ref)--(ref) hold. Taking note that $\bH^u_\perp \boldsymbol{Y}_0 = \boldsymbol{0}$, \begin{align} \| \bH^u_\perp \boldsymbol{y}_T \|_2^2 &= \boldsymbol{y}'_T \bH^u_\perp \boldsymbol{y}_T \\ &= (\boldsymbol{Y}_0 \boldsymbol{\alpha} + \boldsymbol{\varepsilon}_T)' \bH^u_\perp (\boldsymbol{Y}_0 \boldsymbol{\alpha} + \boldsymbol{\varepsilon}_T) \\ &= \boldsymbol{\varepsilon}'_T \bH^u_\perp \boldsymbol{\varepsilon}_T. \end{align} Applying Lemma (ref) then gives \begin{align} \mathbb{E}[\boldsymbol{\varepsilon}'_T \bH^u_\perp \boldsymbol{\varepsilon}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0] = \operatorname{tr}(\bH^u_\perp) (\sigma^\text{hz}_T)^2 = (N_0 - R) (\sigma^\text{hz}_T)^2, \end{align} where the final equality follows because the trace of a projection matrix equals its rank. Taken altogether, we have $\mathbb{E}[\widehat{\bSigma}_T^\text{homo} | \boldsymbol{y}_N, \boldsymbol{Y}_0] = \bSigma^\text{hz}_T$. Therefore, \begin{align} \mathbb{E}[ \widehat{v}_0^{\text{hz}, \text{homo}} | \boldsymbol{y}_N, \boldsymbol{Y}_0] &= \widehat{\boldsymbol{\beta}}' \mathbb{E}[ \widehat{\bSigma}^\text{homo}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0] \widehat{\boldsymbol{\beta}} = v_0^\text{hz}. \end{align} (ii) [\textsf{VT} model] Let Assumptions (ref)--(ref) hold. Following the arguments above, we conclude that $\mathbb{E}[\widehat{\bSigma}_N^\text{homo} | \boldsymbol{y}_T, \boldsymbol{Y}_0] = \bSigma^\text{vt}_N$ and $\mathbb{E}[\widehat{v}_0^{\text{vt}, \text{homo}} | \boldsymbol{y}_T, \boldsymbol{Y}_0] = v_0^\text{vt}$. (iii) [Mixed model] Let Assumptions (ref)--(ref) hold. Following the arguments that led to (ref), we obtain $\mathbb{E}[\widehat{\bSigma}^\text{homo}_T | \boldsymbol{Y}_0] = \bSigma^\text{mix}_T$ and $\mathbb{E}[\widehat{\bSigma}^\text{homo}_N | \boldsymbol{Y}_0] = \bSigma^\text{mix}_N$. Applying Lemma (ref) then gives $\mathbb{E}[ \widehat{v}^{\text{mix}, \text{homo}}_0 | \boldsymbol{Y}_0] = v^\text{mix}_0$. The proof is complete. \end{proof} \subsection{Proof of Lemma (ref)} \begin{proof} Before we establish the biases of $(\widehat{\bSigma}^\text{jack}_T, \widehat{\bSigma}^\text{jack}_N)$, we first justify their forms. As noted in Section (ref), jackknife is a popular approach to estimate the covariances of $(\widehat{\boldsymbol{\alpha}}, \widehat{\boldsymbol{\beta}})$. Below, we follow the standard techniques to derive the jackknife estimate of these objects, which will then be used to derive $(\widehat{\bSigma}^\text{jack}_T, \widehat{\bSigma}^\text{jack}_N)$. Without loss of generality, we begin with $\widehat{\boldsymbol{\alpha}}$. Notably, while standard derivations consider $\boldsymbol{Y}_0$ with full column rank, we consider a general matrix $\boldsymbol{Y}_0$ that may be rank deficient. This difference is subtle so the following proof is by no means novel. We provide it simply for completeness. To describe the jackknife, we define $\widehat{\boldsymbol{\alpha}}_{\sim i}$ as the minimum $\ell_2$-norm solution to (ref), where $\lambda_1 = \lambda_2 = 0$, without the $i$th observation, i.e., \begin{align} \widehat{\boldsymbol{\alpha}}_{\sim i} &= (\boldsymbol{Y}'_{0, \sim i} \boldsymbol{Y}_{0, \sim i})^\dagger \boldsymbol{Y}'_{0, \sim i} \boldsymbol{y}_{T, \sim i}, \end{align} where $\boldsymbol{Y}_{0, \sim i}$ and $\boldsymbol{y}_{T, \sim i}$ correspond to $\boldsymbol{Y}_0$ and $\boldsymbol{y}_T$ without the $i$th observation. We define the pseudo-estimator as $\tilde{\boldsymbol{\alpha}}_i = T_0 \widehat{\boldsymbol{\alpha}} - (T_0-1) \widehat{\boldsymbol{\alpha}}_{\sim i}$. With these quantities defined, we write the jackknife variance estimator as \begin{align} \bhV^\text{jack} = \frac{1}{(T_0-1)^2} \sum_{i \le N_0} (\tilde{\boldsymbol{\alpha}}_i - \widehat{\boldsymbol{\alpha}}) (\tilde{\boldsymbol{\alpha}}_i - \widehat{\boldsymbol{\alpha}})'. \end{align} To evaluate this quantity, we will rewrite $\widehat{\boldsymbol{\alpha}}_{\sim i}$ in a more convenient form. In particular, \begin{align} &\boldsymbol{Y}'_{0, \sim i} \boldsymbol{Y}_{0, \sim i} = \boldsymbol{Y}'_0 \boldsymbol{Y}'_0 - \boldsymbol{y}_i \boldsymbol{y}'_i \\ &\boldsymbol{Y}'_{0, \sim i} \boldsymbol{y}_{T, \sim i} = \boldsymbol{Y}'_0 \boldsymbol{y}_T - \boldsymbol{y}_i Y_{iT}, \end{align} where $\boldsymbol{y}_i = [Y_{it}: t \le T_0]$ is the $i$th row of $\boldsymbol{Y}_0$. We do not assume that $\boldsymbol{Y}'_0 \boldsymbol{Y}_0$ is nonsingular. As such, we use a generalized form of the Sherman-Morrison formula cline, meyer to obtain \begin{align} (\boldsymbol{Y}'_{0, \sim i} \boldsymbol{Y}_{0, \sim i})^\dagger = (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger + (1 - H^u_{ii})^{-1} (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \boldsymbol{y}_i \boldsymbol{y}'_i (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger. \end{align} Recall $\widehat{\boldsymbol{\alpha}} = (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \boldsymbol{Y}'_0 \boldsymbol{y}_T$ and note $Y_{iT} - \boldsymbol{y}'_i \widehat{\boldsymbol{\alpha}}$ is the $i$th element of $\widehat{\boldsymbol{\varepsilon}}_T = \bH^u_\perp \boldsymbol{y}_T$. Using these facts, we plug (ref) into (ref) to yield \begin{align} \widehat{\boldsymbol{\alpha}}_{\sim i} &= \left[ (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger + (1 - H^u_{ii})^{-1} (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \boldsymbol{y}_i \boldsymbol{y}'_i (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \right] ( \boldsymbol{Y}'_0 \boldsymbol{y}_T - \boldsymbol{y}_i Y_{iT}) \\ &= \widehat{\boldsymbol{\alpha}} - (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \boldsymbol{y}_i Y_{iT} + (1 - H^u_{ii})^{-1} (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \boldsymbol{y}_i \boldsymbol{y}'_i \widehat{\boldsymbol{\alpha}} - H^u_{ii} (1 - H^u_{ii})^{-1} (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \boldsymbol{y}_i Y_{iT} \\ &= \widehat{\boldsymbol{\alpha}} - (1-H^u_{ii})^{-1} (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \boldsymbol{y}_i \widehat{\varepsilon}_{iT}. \end{align} Inserting (ref) into our pseudo-estimate, we have \begin{align} \tilde{\boldsymbol{\alpha}}_i &= T_0 \widehat{\boldsymbol{\alpha}} - (T_0-1) \left( \widehat{\boldsymbol{\alpha}} - (1-H^u_{ii})^{-1} (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \boldsymbol{y}_i \widehat{\varepsilon}_{iT} \right) \\ &= \widehat{\boldsymbol{\alpha}} + (T_0-1)(1-H^u_{ii})^{-1} (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \boldsymbol{y}_i \widehat{\varepsilon}_{iT}. \end{align} Inserting (ref) into (ref), we have \begin{align} \bhV^\text{jack} &= (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \left( \sum_{i \le N_0} \frac{\widehat{\varepsilon}_{iT}^2}{(1-H^u_{ii})^2} \boldsymbol{y}_i \boldsymbol{y}'_i \right) (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \\ &= (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger \boldsymbol{Y}'_0 \boldsymbol{\Omega} \boldsymbol{Y}_0 (\boldsymbol{Y}'_0 \boldsymbol{Y}_0)^\dagger, \end{align} where $\boldsymbol{\Omega}$ is a diagonal matrix with $\Omega_{ii} = \widehat{\varepsilon}_{iT}^2 (1-H^u_{ii})^{-2}$. Equivalently, $\boldsymbol{\Omega} = \text{diag}( [ \bH^u_\perp \circ \bH^u_\perp \circ \boldsymbol{I} ]^\dagger [\widehat{\boldsymbol{\varepsilon}}_T \circ \widehat{\boldsymbol{\varepsilon}}_T])$. It then follows that \begin{align} \boldsymbol{y}'_N \bhV^\text{jack} \boldsymbol{y}_N &= \widehat{\boldsymbol{\beta}}' \boldsymbol{\Omega} \widehat{\boldsymbol{\beta}}. \end{align} To arrive at (ref), we define $\widehat{\bSigma}^\text{jack}_T = \boldsymbol{\Omega}$. This corresponds to the EHW estimator with the jackknife correction. We derive (ref) for $\widehat{\boldsymbol{\beta}}$ by applying the same arguments above. Now, we will evaluate the biases of $(\widehat{\bSigma}^\text{jack}_T, \widehat{\bSigma}^\text{jack}_N)$. (i) [\textsf{HZ} model] Let Assumptions (ref)--(ref) hold. We define $(\sigma^\text{hz}_{iT})^2 = \operatorname{Var}(\varepsilon_{iT} | \boldsymbol{y}_N, \boldsymbol{Y}_0)$ for $i = 1, \dots, N_0$. Observe that \begin{align} \mathbb{E}[ ( \bH^u_\perp \circ \bH^u_\perp \circ \boldsymbol{I} )^\dagger (\widehat{\boldsymbol{\varepsilon}}_T \circ \widehat{\boldsymbol{\varepsilon}}_T) | \boldsymbol{y}_N, \boldsymbol{Y}_0] &= ( \bH^u_\perp \circ \bH^u_\perp \circ \boldsymbol{I} )^\dagger \mathbb{E}[ \widehat{\boldsymbol{\varepsilon}}_T \circ \widehat{\boldsymbol{\varepsilon}}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0]. \end{align} To evaluate (ref), we follow the derivations of (ref) and (ref) to obtain \begin{align} \mathbb{E}[\widehat{\boldsymbol{\varepsilon}}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0] &= \bH^u_\perp \boldsymbol{Y}_0 \boldsymbol{\alpha}^* = \boldsymbol{0} \\ \operatorname{Cov}(\widehat{\boldsymbol{\varepsilon}}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0) &= \bH^u_\perp \bSigma^\text{hz}_T \bH^u_\perp. \end{align} Recall that $\mathbb{E}[X^2] = \operatorname{Var}(X) + \mathbb{E}[X]^2$ for any random variable $X$. Thus, combining (ref) with (ref) gives \begin{align} \mathbb{E}[ \widehat{\boldsymbol{\varepsilon}}_T \circ \widehat{\boldsymbol{\varepsilon}}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0] &= ( \bH^u_\perp \bSigma^\text{hz}_T \bH^u_\perp \circ \boldsymbol{I} ) \boldsymbol{1}. \end{align} Let $\widehat{\boldsymbol{\gamma}} = \mathbb{E}[ \widehat{\boldsymbol{\varepsilon}}_T \circ \widehat{\boldsymbol{\varepsilon}}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0]$. By (ref), the $\ell$th entry of $\widehat{\boldsymbol{\gamma}}$ can be written as \begin{align} \widehat{\gamma}_\ell = \sum_{j \neq \ell} (H^u_{j\ell})^2 (\sigma^\text{hz}_{jT})^2 + (1 - H^u_{\ell\ell})^2 (\sigma^\text{hz}_{\ell T})^2, \end{align} where $H^u_{j\ell}$ is the $(j,\ell)$th entry of $\bH^u$. In turn, this allows us to rewrite (ref) as \begin{align} \widehat{\boldsymbol{\gamma}} &= ( \bH^u_\perp \circ \bH^u_\perp ) \bSigma^\text{hz}_T \boldsymbol{1}. \end{align} Next, let $\widehat{\boldsymbol{\zeta}} = (\bH^u_\perp \circ \bH^u_\perp \circ \boldsymbol{I} )^{-1} \widehat{\boldsymbol{\gamma}}$. Notice that the $\ell$th entry of $\widehat{\boldsymbol{\zeta}}$ is given by \begin{align} \widehat{\zeta}_\ell &= (\sigma^\text{hz}_{\ell T})^2 + \sum_{j \neq \ell} \frac{(H^u_{\ell j})^2}{(1-H^u_{\ell \ell})^2} (\sigma^\text{hz}_{jT})^2. \end{align} Therefore, $\text{diag}( \widehat{\boldsymbol{\zeta}} ) = \bSigma^\text{hz}_T + \boldsymbol{\Delta}^\text{hz}$, where $\Delta^\text{hz}_{\ell \ell} = \sum_{j \neq \ell} (\sigma^\text{hz}_{jT})^2 (H^u_{\ell j})^2 (1 - H^u_{\ell \ell})^{-2}$ for $\ell=1, \dots, N_0$. Notice if $\max_\ell H^u_{\ell \ell} < 1$, then $(\bH^u_\perp \circ \bH^u_\perp \circ \boldsymbol{I})$ is nonsingular, i.e., the pseudo-inverse is precisely the inverse. In this situation, plugging the above into (ref) gives \begin{align} \mathbb{E}[\widehat{\bSigma}^\text{jack}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0] &= \text{diag}\left((\bH^u_\perp \circ \bH^u_\perp \circ \boldsymbol{I} )^{-1} \mathbb{E}[ \widehat{\boldsymbol{\varepsilon}}_T \circ \widehat{\boldsymbol{\varepsilon}}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0] \right) \\ &= \text{diag}\left((\bH^u_\perp \circ \bH^u_\perp \circ \boldsymbol{I} )^{-1} \widehat{\boldsymbol{\gamma}} \right) \\ &= \text{diag}(\widehat{\boldsymbol{\zeta}}) \\ &= \bSigma^\text{hz}_T + \boldsymbol{\Delta}^\text{hz}. \end{align} From this, we conclude that \begin{align} \mathbb{E}[\widehat{v}_0^{\text{hz}, \text{jack}} | \boldsymbol{y}_N, \boldsymbol{Y}_0] &= \widehat{\boldsymbol{\beta}}' \mathbb{E}[\widehat{\bSigma}_T^\text{jack} | \boldsymbol{y}_N, \boldsymbol{Y}_0] \widehat{\boldsymbol{\beta}} \\ &= \widehat{\boldsymbol{\beta}}' (\bSigma^\text{hz}_T + \boldsymbol{\Delta}^\text{hz}) \widehat{\boldsymbol{\beta}} \\ &= v_0^\text{hz} + \widehat{\boldsymbol{\beta}}' \boldsymbol{\Delta}^\text{hz} \widehat{\boldsymbol{\beta}}, \end{align} where we note that $\widehat{\boldsymbol{\beta}}' \boldsymbol{\Delta}^\text{hz} \widehat{\boldsymbol{\beta}} \ge 0$. (ii) [\textsf{VT} model] Let Assumptions (ref)--(ref) hold. Following the arguments above, we conclude $\mathbb{E}[\widehat{\bSigma}^\text{jack}_N | \boldsymbol{y}_T, \boldsymbol{Y}_0] = \bSigma^\text{vt}_N + \boldsymbol{\Gamma}^\text{vt}$, where $\Gamma^\text{vt}_{\ell \ell} = \sum_{j \neq \ell} (\sigma^\text{vt}_{Nj})^2 (H^v_{\ell j})^2(1-H^v_{\ell \ell})^{-2}$ for $\ell=1, \dots, T_0$. Thus, $\mathbb{E}[\widehat{v}_0^{\text{vt}, \text{jack}} | \boldsymbol{y}_T, \boldsymbol{Y}_0] = v^\text{vt}_0 + \widehat{\boldsymbol{\alpha}}' \boldsymbol{\Gamma}^\text{vt} \widehat{\boldsymbol{\alpha}}$, where we note that $\widehat{\boldsymbol{\alpha}}' \boldsymbol{\Gamma}^\text{vt} \widehat{\boldsymbol{\alpha}} \ge 0$. (ii) [Mixed model] Let Assumptions (ref)--(ref) hold. We define $(\sigma^\text{mix}_{iT})^2 = \operatorname{Var}(\varepsilon_{iT} | \boldsymbol{Y}_0)$ for $i = 1, \dots, N_0$ and $(\sigma^\text{mix}_{Nt})^2 = \operatorname{Var}(\varepsilon_{Nt} | \boldsymbol{Y}_0)$ for $t = 1, \dots, T_0$. Following the arguments that led to (ref), we obtain $\mathbb{E}[\widehat{\bSigma}^\text{jack}_T | \boldsymbol{Y}_0] = \bSigma^\text{mix}_T + \boldsymbol{\Delta}^\text{mix}$, where $\Delta^\text{mix}_{\ell \ell} = \sum_{j \neq \ell} (\sigma^\text{mix}_{jT})^2 (H^u_{\ell j})^2 (1 - H^u_{\ell \ell})^{-2}$ for $\ell=1, \dots, N_0$. Similarly, we obtain $\mathbb{E}[\widehat{\bSigma}^\text{jack}_N | \boldsymbol{Y}_0] = \bSigma^\text{mix}_N + \boldsymbol{\Gamma}^\text{mix}$, where $\Gamma^\text{mix}_{\ell \ell} = \sum_{j \neq \ell} (\sigma^\text{mix}_{Nj})^2 (H^v_{\ell j})^2(1-H^v_{\ell \ell})^{-2}$ for $\ell=1, \dots, T_0$. Applying Lemma (ref) then gives \begin{align} &\mathbb{E}[\widehat{v}^{\text{mix}, \text{jack}}_0 | \boldsymbol{Y}_0] \\ &= v^\text{mix}_0 + (\bH^u \boldsymbol{\beta}^*)' \boldsymbol{\Delta}^\text{mix} (\bH^u \boldsymbol{\beta}^*) + (\bH^v \boldsymbol{\alpha}^*)' \boldsymbol{\Gamma}^\text{mix} (\bH^v \boldsymbol{\alpha}^*) + \operatorname{tr}( \boldsymbol{Y}_0^\dagger \boldsymbol{\Delta}^\text{mix} (\boldsymbol{Y}'_0)^\dagger \boldsymbol{\Gamma}^\text{mix} ). \end{align} The proof is complete. \end{proof} \subsection{Proof of Lemma (ref)} \begin{proof} We adopt the strategy of variance to prove our desired result. (ii) [\textsf{HZ} model] Let Assumptions (ref)--(ref) hold. As in the proof of Lemma (ref), we define $\widehat{\boldsymbol{\varepsilon}}_T = \bH^u_\perp \boldsymbol{y}_T$. Observe \begin{align} \mathbb{E}[ ( \bH^u_\perp \circ \bH^u_\perp )^{-1} (\widehat{\boldsymbol{\varepsilon}}_T \circ \widehat{\boldsymbol{\varepsilon}}_T) | \boldsymbol{y}_N, \boldsymbol{Y}_0] &= ( \bH^u_\perp \circ \bH^u_\perp )^{-1} \mathbb{E}[ \widehat{\boldsymbol{\varepsilon}}_T \circ \widehat{\boldsymbol{\varepsilon}}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0]. \end{align} To evaluate (ref), we plug in (ref) to obtain \begin{align} \mathbb{E}[ ( \bH^u_\perp \circ \bH^u_\perp )^{-1} (\widehat{\boldsymbol{\varepsilon}}_T \circ \widehat{\boldsymbol{\varepsilon}}_T) |\boldsymbol{y}_N, \boldsymbol{Y}_0] &= ( \bH^u_\perp \circ \bH^u_\perp )^{-1} ( \bH^u_\perp \circ \bH^u_\perp ) \bSigma^\text{hz}_T \boldsymbol{1} = \bSigma^\text{hz}_T \boldsymbol{1}. \end{align} Plugging (ref) into (ref) yields \begin{align} \mathbb{E}[\widehat{\bSigma}^\text{HRK}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0] &= \text{diag} \left( (\bH^u_\perp \circ \bH^u_\perp )^{-1} \mathbb{E}[ \widehat{\boldsymbol{\varepsilon}}_T \circ \widehat{\boldsymbol{\varepsilon}}_T | \boldsymbol{y}_N, \boldsymbol{Y}_0] \right) = \bSigma^\text{hz}_T. \end{align} It then follows that $\mathbb{E}[ \widehat{v}_0^{\text{hz}, \text{HRK}} | \boldsymbol{y}_N, \boldsymbol{Y}_0] = v_0^\text{hz}$. (ii) [\textsf{VT} model] Let Assumptions (ref)--(ref) hold. Following the same arguments as above, we conclude $\mathbb{E}[\widehat{\bSigma}^\text{HRK}_N | \boldsymbol{y}_T, \boldsymbol{Y}_0] = \bSigma^\text{vt}_N$ and $\mathbb{E}[ \widehat{v}_0^{\text{vt}, \text{HRK}} | \boldsymbol{y}_T, \boldsymbol{Y}_0] = v^\text{vt}_0$. (ii) [Mixed model] Let Assumptions (ref)--(ref) hold. Following the arguments that led to (ref), we obtain $\mathbb{E}[\widehat{\bSigma}^\text{HRK}_T | \boldsymbol{Y}_0] = \bSigma^\text{mix}_T$ and $\mathbb{E}[\widehat{\bSigma}^\text{HRK}_N | \boldsymbol{Y}_0] = \bSigma^\text{mix}_N$. Applying Lemma (ref) then gives $\mathbb{E}[ \widehat{v}^{\text{mix}, \text{HRK}}_0 | \boldsymbol{Y}_0] = v^\text{mix}_0$. The proof is complete. \end{proof} \subsection{Proof of Lemma (ref)} \begin{proof} By linearity of expectations, \begin{align} \mathbb{E}[\widehat{v}^\text{mix}_0 | \boldsymbol{Y}_0] &= \mathbb{E}[ \widehat{v}_0^\text{hz} | \boldsymbol{Y}_0] + \mathbb{E}[ \widehat{v}_0^\text{vt} | \boldsymbol{Y}_0] - \mathbb{E}[ \operatorname{tr}(\boldsymbol{Y}_0^\dagger \widehat{\bSigma}_T (\boldsymbol{Y}'_0)^\dagger \widehat{\bSigma}_N) | \boldsymbol{Y}_0]. \end{align} We evaluate each term in (ref). Beginning with the first term, note that the randomness in $\widehat{\bSigma}_T$ stems from $\boldsymbol{\varepsilon}_T$ and $\widehat{\boldsymbol{\beta}}$ is deterministic given $(\boldsymbol{\varepsilon}_N, \boldsymbol{Y}_0)$. As such, Assumptions (ref)--(ref) with Lemma (ref) gives \begin{align} \mathbb{E}[ \widehat{v}_0^\text{hz} | \boldsymbol{Y}_0] &= \mathbb{E}[ \widehat{\boldsymbol{\beta}}' \widehat{\bSigma}_T \widehat{\boldsymbol{\beta}} | \boldsymbol{Y}_0] \\ &= \mathbb{E}\left[ \mathbb{E}[ \widehat{\boldsymbol{\beta}}' \widehat{\bSigma}_T \widehat{\boldsymbol{\beta}} | \boldsymbol{\varepsilon}_N, \boldsymbol{Y}_0] | \boldsymbol{Y}_0 \right] \\ &= \mathbb{E}\left[ \boldsymbol{y}'_N \boldsymbol{Y}^\dagger_0 \mathbb{E}[ \widehat{\bSigma}_T | \boldsymbol{Y}_0] (\boldsymbol{Y}'_0)^\dagger \boldsymbol{y}_N | \boldsymbol{Y}_0 \right] \\ &= \mathbb{E}\left[ (\boldsymbol{Y}'_0 \boldsymbol{\beta}^* + \boldsymbol{\varepsilon}_N) \boldsymbol{Y}^\dagger_0 \mathbb{E}[ \widehat{\bSigma}_T | \boldsymbol{Y}_0] (\boldsymbol{Y}'_0)^\dagger (\boldsymbol{Y}'_0 \boldsymbol{\beta}^* + \boldsymbol{\varepsilon}_N) | \boldsymbol{Y}_0 \right] \\ &= (\bH^u \boldsymbol{\beta}^*)' \mathbb{E}[ \widehat{\bSigma}_T | \boldsymbol{Y}_0] (\bH^u \boldsymbol{\beta}^*) + \operatorname{tr}( \boldsymbol{Y}_0^\dagger \mathbb{E}[ \widehat{\bSigma}_T | \boldsymbol{Y}_0] (\boldsymbol{Y}'_0)^\dagger \bSigma^\text{mix}_N). \end{align} By an analogous argument, we derive \begin{align} \mathbb{E}[ \widehat{v}_0^\text{vt} | \boldsymbol{Y}_0] = (\bH^v \boldsymbol{\alpha}^*)' \mathbb{E}[ \widehat{\bSigma}_N | \boldsymbol{Y}_0] (\bH^v \boldsymbol{\alpha}^* ) + \operatorname{tr}( \boldsymbol{Y}_0^\dagger \bSigma^\text{mix}_T (\boldsymbol{Y}'_0)^\dagger \mathbb{E}[\widehat{\bSigma}_N | \boldsymbol{Y}_0]). \end{align} Finally, we use the linearity of the trace operator with Assumption (ref) to obtain \begin{align} \mathbb{E}[ \operatorname{tr}(\boldsymbol{Y}_0^\dagger \widehat{\bSigma}_T (\boldsymbol{Y}'_0)^\dagger \widehat{\bSigma}_N) | \boldsymbol{Y}_0] &= \mathbb{E}\left[ \mathbb{E}[ \operatorname{tr}(\boldsymbol{Y}_0^\dagger \widehat{\bSigma}_T (\boldsymbol{Y}'_0)^\dagger \widehat{\bSigma}_N) | \boldsymbol{\varepsilon}_N, \boldsymbol{Y}_0] | \boldsymbol{Y}_0 \right] \\ &= \mathbb{E}\left[ \operatorname{tr}(\boldsymbol{Y}_0^\dagger \mathbb{E}[\widehat{\bSigma}_T | \boldsymbol{Y}_0] (\boldsymbol{Y}'_0)^\dagger \widehat{\bSigma}_N) | \boldsymbol{Y}_0 \right] \\ &= \operatorname{tr}(\boldsymbol{Y}_0^\dagger \mathbb{E}[\widehat{\bSigma}_T | \boldsymbol{Y}_0] (\boldsymbol{Y}'_0)^\dagger \mathbb{E}[\widehat{\bSigma}_N | \boldsymbol{Y}_0 ]). \end{align} Putting everything together completes the proof. \end{proof}