EconBase
← Back to paper

Synthetic Interventions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

146,151 characters · 28 sections · 70 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Synthetic Interventions

The synthetic controls (SC) methodology is a prominent tool for policy evaluation in panel data applications. Researchers commonly justify the SC framework with a low-rank matrix factor model that assumes the potential outcomes are described by low-dimensional unit and time specific latent factors. In the recent work of abadie_survey, one of the pioneering authors of the SC method posed the question of how the SC framework can be extended to multiple treatments. This article offers one resolution to this open question that we call synthetic interventions (SI). Fundamental to the SI framework is a low-rank tensor factor model, which extends the matrix factor model by including a latent factorization over treatments. Under this model, we propose a generalization of the standard \textsf{SC}-based estimators. We prove the consistency for one instantiation of our approach and provide conditions under which it is asymptotically normal. Moreover, we conduct a representative simulation to study its prediction performance and revisit the canonical \textsf{SC} case study of abadie2 on the impact of anti-tobacco legislations by exploring related questions not previously investigated.

{ Keywords: synthetic controls, tensor completion, tensor factor model, principal component regression, covariate shift }

\footnotetext[1]{We sincerely thank Alberto Abadie, Abdullah Alomar, Peter Bickel, Victor Chernozhukov, Romain Cosson, Peng Ding, Esther Duflo, Avi Feller, Guido Imbens, Anna Mikusheva, Jasjeet Sekhon, Rahul Singh, and Bin Yu for invaluable feedback, guidance, and discussion. We also thank members within the MIT Economics department and the Laboratory for Information and Decisions Systems (LIDS) for useful discussions.}

Introduction

In 1988, California passed the first large-scale anti-tobacco program in the United States known as Proposition 99. Early reports of the success of Proposition 99 sparked a wave of anti-tobacco measures across the country: from 1989--2000, four states (Arizona, Massachusetts, Oregon, and Florida) introduced similar anti-tobacco programs and seven states (Alaska, Hawaii, Maryland, Michigan, New Jersey, New York, Washington) raised their state cigarette taxes by 50 cents or more. The remaining 38 states (e.g., Kansas or Virginia) kept their statewide status quos. To isolate the impact of these anti-tobacco measures, researchers sought an answer to the following question:

center[center omitted — 174 chars of source]

The seminal papers of abadie1, abadie2 introduced the synthetic controls (SC) framework that provides an elegant approach to tackle this challenge. In the case of California, the SC method constructs a synthetic California from a weighted composition of “control” states that did not impose any anti-tobacco measures to estimate California's counterfactual cigarette sales (outcome variable of interest) in the absence of Proposition 99. The weights are chosen such that the trajectory of cigarette sales of the synthetic California mirrors that of the observed California prior to the passage of Proposition 99 (“pre-treatment” period). The difference between the observed California with Proposition 99 and the synthetic California without Proposition 99 during the “post-treatment” period of 1989--2000 sheds insight into the causal effect of the initiative. Following this procedure, abadie2 found that tobacco consumption fell markedly in California under Proposition 99 compared to its synthetic counterpart. Given its immense popularity, the SC framework is often regarded as one of the most important innovations in the policy evaluation literature athey.

However, even though the SC framework offers a resolution to the question above, several related questions remain: {\em what would have happened if a...}

center[center omitted — 368 chars of source]

The SC framework, in its current form, is not equipped to explore these set of questions. To see this, consider how SC would address Q2. Following Q1, the SC method constructs a synthetic California from a weighted composition of {\em treated} states that raised taxes during the post-treatment period, states such as Kansas and Virginia. These weights, as in Q1, would be learned during the pre-treatment period whence all outcomes are observed under {\em control}. In contrast to Q1, however, these weights would then be applied on the {\em treated} outcomes associated with those states that raised taxes during the post-treatment period. Similar approaches can be taken to answer Q3-Q5, e.g., the solution to Q4 would mimic that of Q2 with the control states in place of the anti-tobacco program states.

The juxtaposition of Q1 and Q2-Q5 reveals their fundamental difference: whereas the SC solution to Q1 applies the weights learnt from the pre-treatment outcomes observed under {\em control} to post-treatment outcomes also observed under {\em control}, the SC solutions to Q2-Q5 apply the weights learnt from the pre-treatment outcomes observed under {\em control} to post-treatment outcomes observed under {\em treatment}. Restated differently, Q1 inherits “data symmetry” between the pre- and post-treatment periods while Q2-Q5 faces “data asymmetry” between the pre- and post-treatment periods. The stark contrast in data usage between Q1 versus Q2-Q5 begs the question: can the SC framework be extended to the setting of multiple treatments such that the same data used to answer questions a la Q1 can also be used to answer questions a la Q2-Q5? More formally, can a set of weights learned from pre-treatment outcomes observed under one setting (e.g., control) be applied toward post-treatment outcomes observed under another setting (e.g., treatment)? This is posed as an open question by one of the pioneering authors of the SC framework abadie_survey.

Contributions

In this article, we propose the synthetic interventions (SI) framework, which provides one resolution to the question asked by abadie_survey. Building on the canonical matrix factor model that is widely used to justify the SC framework, we consider a {\em tensor} factor model, which serves as the bedrock for the SI framework.

Intuitively, the matrix factor model postulates that the potential cigarette sales of California under their statewide status quo (i.e., absence of any anti-tobacco measure) in the year 1993, for instance, can be characterized by latent factors that describe California and the year 1993. Notably, only unit (e.g., California) and time (e.g., year 1993) specific latent factors are required since the SC framework is solely concerned with outcomes observed under a single setting, typically control. The tensor factor model, on the other hand, postulates that the same potential outcome is characterized by not only the latent factors of California and the year 1993, but also the latent factor associated with the status quo (control). More generally, the tensor factor model posits that each potential outcome is described by three latent factors tailored to the unit, time period, and treatment of interest, where control can be viewed as the “default” treatment. With the added assumption of separate latent factors for treatments, the SI framework can justifiably exploit the same data pattern considered in the SC framework yet provide answers to far more questions.

Under the tensor factor model, we establish the identification of several causal estimands related to questions such as Q2-Q5. Moreover, we propose a generalization of the standard SC-based estimators to answer said questions of interest. We analyze one instantiation of this methodology and prove its consistency. In Appendix (ref), we provide conditions under which our estimation strategy is asymptotically normal and show that a minor modification to the instantiation analyzed in the main body satisfies those conditions and thus enables inference. Notably, our results remain valid in the presence of hidden confounders, provided these confounders can be explained by the latent factors associated with our tensor formulation. We corroborate our theoretical results with simulations and revisit the popular California Proposition 99 study of abadie2 by examining the other treated states; in the latter, we find that (i) anti-tobacco programs and raised taxes had similar effects and (ii) both measures led to a consequential decrease in tobacco consumption relative to the status quo, which reaffirms the conclusion drawn in abadie2 for California.

Organization of Paper

Section (ref) presents the SI framework and describes the causal estimand. Section (ref) proposes an extension of the SC-based estimators for our setting. Section (ref) analyzes one instantiation of our estimator and proves its consistency. Section (ref) reinforces our theoretical results with a simulation study. Section (ref) revisits the classical California Proposition 99 study of abadie2. Section (ref) summarizes this article. Additional discussions and proofs are relegated to the appendix.

Notation

For any positive integer $a$, let $[a] = \{1, \dots, a\}$ and $[a]_0 = \{0, \dots, a-1\}$. For a vector $v \in \mathbb{R}^a$, let $\|v\|_p$ denote its $\ell_p$-norm. We define the inner product between vectors $v, x \in \mathbb{R}^a$ as $\langle v, x \rangle = v^\top x = \sum_{\ell=1}^a v_\ell x_\ell$. For a matrix $\boldsymbol{A} \in \mathbb{R}^{a \times b}$, we denote its Frobenius norm as $\|\boldsymbol{A}\|_F$. Let $\| \cdot \|_{\psi_2}$ denote the Orlicz norm. Let $O_p$ and $o_p$ denote the probabilistic versions of the deterministic big-$O$ and little-$o$ notations. Denote $\boldsymbol{I}$ as the identity matrix and $(\cdot)^{\dagger}$ as the Moore-Penrose pseudoinverse.

Synthetic Interventions Framework

We consider the canonical data format within the SC literature known as {\em panel data}, a collection of observations on multiple units over multiple time periods. Formally, we observe $N$ units over $T$ time periods, indexed by $i \in [N]$ and $t \in [T]$, respectively. As an example, the tobacco case study tracks the cigarette sales of $N = 50$ states over $T = 31$ years. Following the potential outcomes framework of neyman, rubin, we philosophize that in each time period $t$, each unit $i$ is characterized by $D$ potential outcomes denoted as $Y_{ti}^{(d)} \in \mathbb{R}$ for $d \in [D]_0$; consistent with standard notation, we reserve $d = 0$ for control. In the context of the tobacco study, the potential outcomes framework asserts that three versions of each state simultaneously exist each year: one that keeps the statewide status quo ($d=0$), one that imposes an anti-tobacco program ($d=1$), and one that raises taxes ($d=2$). Of course, each state can only adopt one policy in any year and hence, researchers can never observe all possible potential outcomes---this is the fundamental challenge of causal inference.

Counterfactual Estimation as Tensor Completion

The synthetic interventions (SI) framework views the problem of counterfactual estimation through the lens of an order-three tensor. Specifically, we encode the potential outcomes into a $T \times N \times D$ tensor, $\boldsymbol{Y}^* = [Y_{ti}^{(d)}]$, whose dimensions correspond to time, units, and treatments. Accordingly, the $(t, i, d)$-th entry of $\boldsymbol{Y}^*$ is the potential outcome of unit $i$ at time $t$ under treatment $d$. If $\boldsymbol{Y}^*$ is known to the researcher, then any causal quantity of interest can be calculated by accessing the relevant entries of $\boldsymbol{Y}^*$. The challenge, however, is that the researcher can only observe a sparse subset of $\boldsymbol{Y}^*$.

In the classical SC setup, all $N$ units are observed under control for the first $T_0$ time periods. We refer to this duration as the pre-intervention or pre-treatment period. For the remaining $T_1 = T - T_0$ time periods, known as the post-intervention or post-treatment period, each unit is assigned to one of the $D$ treatments. Let $\mathcal{I}^{(d)} \subset [N]$ denote the set of units that are assigned to treatment $d$ during the post-treatment period with $N_d = |\mathcal{I}^{(d)}|$. In the tobacco study, $N_0 = 38$ states kept their statewide status quo, $N_1=5$ states imposed a tobacco control program, and $N_2=7$ states raised taxes during the post-treatment period.

We denote $Y_{tid} \in \{\mathbb{R} \cup \star\}$, where $\star$ represents a missing entry, as the recorded outcome for unit $i$ at time $t$ under treatment $d$, and store these outcomes into a $T \times N \times D$ tensor $\boldsymbol{Y} = [Y_{tid}]$. The following assumption formalizes our observation pattern in $\boldsymbol{Y}$ and establishes its connection with $\boldsymbol{Y}^*$.

assumption[Stable Unit Treatment Value Assumption] We observe \begin{itemize} • Pre-treatment period: $Y_{ti0} = Y_{ti}^{(0)}$ for $t \le T_0$ and $i \in [N]$. • Post-treatment period: $Y_{tid} = Y_{ti}^{(d)}$ for $t > T_0$, $i \in \mathcal{I}^{(d)}$, and $d \in [D]_0$. \end{itemize} For all other entries of $\boldsymbol{Y}$, we have $Y_{tnd} = \star$.

In view of Assumption (ref), estimating counterfactual potential outcomes is equivalent to imputing missing entries in $\boldsymbol{Y}$; in the machine learning community, the latter problem is known as {\em tensor completion}. For a visualization of $\boldsymbol{Y}^*$ and $\boldsymbol{Y}$, see Figure (ref). We remark that Assumption (ref) implicitly rules out spillover (network) effects between units imbens_rubin_2015.

figure[figure omitted — 944 chars of source]

Tensor Factor Model

To enable faithful recovery of the missing entries in $\boldsymbol{Y}$, we impose structure on $\boldsymbol{Y}^*$. We motivate our model by first revisiting the dominant idea behind the SC framework---the low-rank matrix factor model, also known as the interactive fixed effects model bai03. Recall that the SC framework is solely focused on a single treatment setting, most commonly control. Restated through our tensor formulation, the SC framework aims to recover the single tensor slice (matrix) corresponding to control. The potential outcomes under control are posited to obey the following: for all $t \in [T]$ and $i \in [N]$, we have

align[align omitted — 153 chars of source]

where (i) $u_t \in \mathbb{R}^r$ and $v_t \in \mathbb{R}^r$ are latent factors specific to time $t$ and unit $i$, respectively, with $r$ sufficiently small relative to $N_0$ and $T_0$, and (ii) $\varepsilon_{ti} \in \mathbb{R}$ represents the (mean-zero) idiosyncratic shock. The key structure in (ref) is that the latent unit factors are {\em invariant} across time. As such, the relationship between units in the pre-treatment period continues to hold in the post-treatment period. In turn, researchers can learn a set of weights from pre-treatment outcomes observed under one setting (e.g., status quo) and apply them toward post-treatment outcomes observed under the {\em same} setting (e.g., status quo) in order to answer questions a la Q1 of Section (ref). The SC framework is thus justified under (ref) abadie_survey.

The tensor factor model builds upon (ref) by introducing a latent factorization across treatments: for all $t \in [T]$, $i \in [N]$, and $d \in [D]_0$, we have

align[align omitted — 129 chars of source]

where (i) $\lambda_d \in \mathbb{R}^r$ is a latent factor specific to treatment $d$ with $r$ sufficiently small relative to $N_d$ and $T_0$, and (ii) $\varepsilon_{ti}^{(d)}$ is the (mean-zero) idiosyncratic shock. Under the tensor factor model, the latent unit factors are invariant across both time {\em and treatments}. As a result, the relationship between units in the pre-treatment period under one treatment continues to hold in the post-treatment period under {\em another treatment}. Translated to practice, researchers can learn a set of weights from pre-treatment outcomes observed under one setting (e.g., status quo) and apply them toward post-treatment outcomes observed under {\em another setting} (e.g., raised taxes) to answer questions a la Q2-Q5 of Section (ref). The SI framework, therefore, is justified under (ref). We highlight that the low-rank tensor factor model stands as the de-facto assumption in the tensor completion literature tensor_missing, gandy, anandkumar2014tensor, barak2016noisy.

With this concept in mind, we present our primary structural assumption on $\boldsymbol{Y}^*$.

assumption[Tensor factor model] Each entry in $\boldsymbol{Y}^*$ satisfies the following: for all $(t,i,d)$, we have \begin{align} Y_{ti}^{(d)} = \langle u^{(d)}_t, v_i \rangle + \varepsilon_{ti}^{(d)} = \sum_{\ell=1}^r u^{(d)}_{t\ell} v_{i\ell} + \varepsilon_{ti}^{(d)}, \end{align} where $u^{(d)}_t \in \mathbb{R}^r$ is a latent factor specific to $(t, d)$, $v_i \in \mathbb{R}^r$ is a latent factor specific to $i$, and $\varepsilon_{ti}^{(d)} \in \mathbb{R}$ is the (mean-zero) idiosyncratic shock.

Note that (ref) is implied by (ref) since we can express $u_{t\ell}^{(d)} = u_{t\ell} \lambda_{d\ell}$, where $u_{t\ell}$ and $\lambda_{d\ell}$ are defined as in (ref). Since the invariance of the unit factors is the critical characteristic needed to validate the SI framework, we operate with the weaker condition stated in (ref) as opposed to (ref). We acknowledge that (ref) adds to the traditional matrix factor model assumption within the SC framework by imposing that each slice of $\boldsymbol{Y}^*$ satisfies its own matrix factor model and that the slices share an invariant set of latent unit factors.

Causal Estimand

Anchoring on the tensor factor model stated in Assumption (ref), we are ready to define our causal estimand. Formally, for a given unit $i$ and treatment $d$, we are interested in estimating

align[align omitted — 151 chars of source]

In words, (ref) represents the average expected potential outcome for unit $i$ under treatment $d$ during the post-treatment period. Since (ref) conditions on the latent factors, its expectation is taken over the measurement shocks. Returning to Q2 in Section (ref) with California as the unit of interest, (ref) translates as California's average expected cigarette sales from 1989--2000 had the state raised taxes instead of imposing an anti-tobacco program.

We take this opportunity to underscore once more the contrast between the SC and SI objectives: whereas the SC framework aims to recover $\theta_i^{(0)}$, the SI framework aims to recover $\theta_i^{(d)}$ for any treatment $d$. Put another way, the SC objective is to impute the first tensor slice in Figure (ref) while the SI framework looks to impute the remaining slices during the post-treatment period. Next, we discuss how our causal estimand can be identified.

Additional Assumptions for Identification

Let $\mathcal{D} = \{ (t, i, d) : ~Y_{tid} \neq \star \}$ denote the treatment assignments. Ideally, $\mathcal{D}$ is independent of $\boldsymbol{Y}^*$ as in the case of randomized experiments, the gold standard mechanism for learning causal effects. With observational data, however, $\mathcal{D}$ is often dependent on $\boldsymbol{Y}^*$---this is known as {\em confounding}. For example, a state policy-maker may favor an anti-tobacco program as they expect their statewide tobacco consumption to fall more rapidly under the program than under the continued status quo or raised taxes.

We consider a setting where the confounders can be latent but their impact on $\mathcal{D}$ and $\boldsymbol{Y}^*$ is mediated through the latent factors, $\mathcal{LF} = \{ u^{(d)}_t, v_i: t \in [T], i \in [N], d \in [D]_0 \}$.

assumption[Selection on latent factors] For all $(t,i,d)$, we have $\mathbb{E}[\varepsilon^{(d)}_{ti}| \mathcal{LF}] = 0$ and $\varepsilon^{(d)}_{ti} \protect\mathpalette{\protect\independenT}{\perp} \mathcal{D} ~|~ \mathcal{LF}$.

Strictly speaking, we only require $\mathbb{E}[\varepsilon^{(d)}_{tn}| \mathcal{LF}, \mathcal{D}] = 0$, which is known as conditional mean independence. We state it as $\varepsilon^{(d)}_{tn} \protect\mathpalette{\protect\independenT}{\perp} \mathcal{D} ~|~ \mathcal{LF}$ to increase the interpretability of our conditional exogeneity assumption. Collectively, Assumptions (ref) and (ref) imply that $\boldsymbol{Y}^* \protect\mathpalette{\protect\independenT}{\perp} \mathcal{D} ~|~ \mathcal{LF}$. Hence, Assumption (ref) can be interpreted as {\em selection on latent factors}, which is analogous to the classical {\em selection on observables} assumption. That is, Assumption (ref) states that the latent factors can act as our latent confounders. Similar conditional independence assumptions have been utilized in athey1, kallus2018causal, and asc.

To state our next assumption, let $\mathcal{E} = \{\mathcal{LF}, \mathcal{D}\}$.

assumption[Latent unit factors] Given $(i,d)$, we have $v_i \in \emph{span}(\{v_j: j \in \mathcal{I}^{(d)}\})$, conditional on $\mathcal{E}$. Specifically, there exists weights $w^{(i,d)} \in \mathbb{R}^{N_d}$ such that $v_i = \sum_{j \in \mathcal{I}^{(d)}} w^{(i,d)}_j v_j$.

Assumption (ref) states that, within the space of latent unit factors, unit $i$ can be reconstructed as a weighted average of units that adopt treatment $d$ during the post-treatment period. The critical feature to observe is that $w^{(i,d)}$ is time- and treatment-invariant. Returning to our tobacco study, if Assumption (ref) holds for California under raised taxes ($d=2$), then California's latent factor can be expressed as a weighted sum of latent factors for the seven states that raised taxes from 1989--2000. If Assumption (ref) were to also hold between California under the status quo ($d=0$), then California's latent factor can equivalently be written as another weighted sum of latent factors associated with the 38 states that kept their statewide status quo from 1989--2000. Due to its invariance property, the weights in both settings remain static regardless of the time period or treatment.

When the tensor factor model stated in Assumption (ref) admits a low-dimensional structure, i.e., $r$ as defined in (ref) is sufficiently small (formalized in Theorem (ref)), Assumption (ref) effectively follows as an immediate consequence. To see this, let $\boldsymbol{V}_{\mathcal{I}^{(d)}} = [v_j: j \in \mathcal{I}^{(d)}] \in \mathbb{R}^{N_d \times r}$. Note that if $\text{rank}(\boldsymbol{V}_{\mathcal{I}^{(d)}}) = r \ll N_d$, then Assumption (ref) becomes a direct byproduct. In other words, a low-dimensional tensor factor model implies that the latent unit factors are linearly dependent, which is the structure exploited in Assumption (ref). From a practical standpoint, Assumption (ref) requires that sufficiently many units are assigned to treatment $d$ and that their latent factors are collectively “diverse” enough so that their span includes the latent factor of unit $i$. In this view, Assumption (ref) is analogous to the classical {\em common support} assumption. We finish by remarking that low-dimensional factor models have been shown to be a ubiquitous phenomenon in practice and are justified under many natural data generating processes Chatterjee15, Udell2017NiceLV, udell2018big, xu2017rates.

Identification Result

With our assumptions stated, we establish our identification result for (ref).

theoremGiven $(i,d)$, let Assumptions (ref) to (ref) hold. Then we have \begin{align} \mathbb{E} [ Y^{(d)}_{ti} \big| u^{(d)}_t, v_i ] &= \sum_{j \in \mathcal{I}^{(d)}} w_j^{(i,d)} \cdot \mathbb{E}\left[ Y_{tjd} | \mathcal{E} \right] \quad for all t \in [T], \\ \theta^{(d)}_i &= \frac{1}{T_1} \sum_{t > T_0} \sum_{j \in \mathcal{I}^{(d)}} w_j^{(i,d)} \cdot \mathbb{E}\left[ Y_{tjd} | \mathcal{E} \right]. \end{align}

Theorem (ref) involves two claims: (ref) states that our causal estimand in (ref), which is defined as an average over all $t > T_0$, can be written as a function of estimable quantities; (ref) is a stronger version of (ref) that holds for each time $t \in [T]$. As a consequence of Theorem (ref), for a given $(i, d)$ pair, we look to estimate $w^{(i,d)}$, which we reemphasize is invariant across time and treatments. Section (ref) proposes an estimator to achieve this objective.

Remarks on Assumptions

remark[On Assumption (ref)] {\em In Appendix (ref), we show that sufficiently continuous non-linear latent variable models are arbitrarily well-approximated by linear factor models of the form (ref) in Assumption (ref) as $N$ and $T$ grow. }
remark[On Assumption (ref)] {\em As with any assumption on confounding, it is difficult---if not impossible---to verify Assumption (ref) for observational data. Rationalizing the validity of Assumption (ref), therefore, requires subject matter knowledge on the application of interest. }
remark[On Assumption (ref)] {\em Assumption (ref) cannot be fully verified since the unit factors are latent. With that said, we can examine its validity through the pre-treatment fit, as prescribed in abadie1, abadie2, abadie_survey. See asc, SDID for a bias-correction variant of the standard SC-based estimators when the pre-treatment fit is poor. }

Synthetic Interventions Estimator

Motivated by Theorem (ref), we propose the synthetic interventions (SI) estimator. Without loss of generality, we focus on estimating $\theta^{(d)}_i$ for a given $(i,d)$ pair.

Description of Estimator

To facilitate the description of the SI estimator, we introduce some useful notation. Let $Y_{\text{pre}, i} = [Y_{ti0} : t \le T_0] \in \mathbb{R}^{T_0}$ collect the pre-treatment outcomes for unit $i$. Let $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} = [ Y_{tj0} : t \le T_0, \ j \in \mathcal{I}^{(d)} ] \in \mathbb{R}^{T_0 \times N_d}$ and $\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}} = [ Y_{tjd} : t > T_0, \ j \in \mathcal{I}^{(d)} ] \in \mathbb{R}^{T_1 \times N_d}$ collect the pre- and post-treatment outcomes, respectively, for the units in $\mathcal{I}^{(d)}$. Note that $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$ is observed under the “default” treatment (control) while $\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}}$ is observed under the treatment $d$ of interest. The SI estimator is implemented via the following two steps:

enumerate• {\em Model identification}: given a constraint set $\mathcal{W}$, define \begin{align} \widehat{w}^{(i,d)} &\in \underset{w \in \mathcal{W}}{argmin} \| Y_{pre, i} - \boldsymbol{Y}_{pre, \mathcal{I}^{(d)}} w \|_2^2. \end{align} • {\em Point estimation}: let $\widehat{\mathbb{E}}[Y^{(d)}_{ti}] = \sum_{j \in \mathcal{I}^{(d)}} \widehat{w}^{(i,d)}_j \cdot Y_{tjd}$ for each $t > T_0$. Then, we have \begin{align} \widehat{\theta}^{(d)}_i &= \frac{1}{T_1} \sum_{t > T_0} \widehat{\mathbb{E}}[Y^{(d)}_{ti}]. \end{align}

For an illustration of the SI estimator, consider Q2 of Section (ref) with California and raised taxes as unit $i$ and treatment $d$, respectively. In (ref), the SI estimator outputs a set of weights that minimizes the pre-treatment fit between California and a weighted average of the seven states that raised taxes (e.g., New Jersey) subject to the constraint set $\mathcal{W}$. The weighted combination of these taxed states acts as our synthetic California. Accordingly, in (ref), the SI estimator proceeds to re-weight the post-treatment outcomes of those seven states as per the learnt weights to predict California's counterfactual trajectory of cigarette sales from 1989--2000 had its policy-makers opted to raise taxes.

We reemphasize that the weights are constructed from pre-treatment outcomes observed under control and applied toward post-treatment outcomes observed under treatment $d$, which can be the same or different from control. In this view, the SI estimator is a generalization of the standard SC-based estimators that operate only on control outcomes.

Different Formulations of the SI Estimator

There are numerous formulations of the SC estimator, each characterized by a particular choice of $\mathcal{W}$ in (ref). The pioneering works of abadie1, abadie2 restrict the weights to lie within the simplex, i.e., $\mathcal{W} = \{w \in \mathbb{R}^{N_d}: w_j \ge 0, \sum_{j \in \mathcal{I}^{(d)}} w_j = 1\}$. Attractive aspects of the simplex formulation include interpretability, sparsity, and transparency abadie_survey. Other common approaches include the lasso LiBell17, arco, chernozhukov2020practical, sc_lasso1, ridge regression asc, elastic-net imbens16, principal component regression rsc, mrsc, and unconstrained linear regression hcw, li2020. Given its connection with the SC estimator, the SI estimator also allows for the same flexibility in defining $\mathcal{W}$.

A Variation Based on Principal Component Regression

In the remainder of this article, we consider the principal component regression (PCR) variant of the SI estimator, which we refer as SI-PCR. To describe this variation, we denote the singular value decomposition of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$ as $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} = \sum_{\ell \ge 1} \widehat{s}_{\ell} \widehat{u}_{\ell} \widehat{v}^\top_{\ell}$, where $\widehat{s}_\ell \in \mathbb{R}$ are the singular values arranged in decreasing order, and $\widehat{u}_\ell \in \mathbb{R}^{T_0}, \widehat{v}_\ell \in \mathbb{R}^{N_d}$ are the corresponding left and right singular vectors, respectively. Letting $\widehat{\boldsymbol{V}}_\text{pre} = [\widehat{v}_1, \dots, \widehat{v}_k] \in \mathbb{R}^{N_d \times k}$ collect the top $k$ right singular vectors of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$, we define $\mathcal{W} = \{w \in \mathbb{R}^{N_d}: (\boldsymbol{I} - \widehat{\boldsymbol{V}}_\text{pre} \widehat{\boldsymbol{V}}_\text{pre}^\top) w = 0\}$. This yields a closed-form expression for the linear model:

align[align omitted — 161 chars of source]

In words, SI-PCR first pre-processes $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$ by obtaining its rank-$k$ approximation and then regresses $Y_{\text{pre}, i}$ on the resulting matrix.

\paragraph{Benefits.} PCR yields several benefits. To begin, PCR tackles overfitting by regularizing the weights through a spectral sparsity constraint on $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$ that sets all of its singular values smaller than $\widehat{s}_k$ to zero. Further, Assumptions (ref) and (ref) imply that we would ideally regress $Y_{\text{pre}, i}$ on $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}]$. In reality, however, we only observe its noisy instantiation, $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$; this is known as {\em error-in-variables} regression. When $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}]$ has low-dimensional structure, agarwal2020robustness and pcr_aos show that the first step of PCR is an effective approach to de-noise $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$. For this reason, PCR is particularly suited under our model.

\paragraph{Choosing $k$.} Popular data-driven approaches to choose $k$ include cross-validation and “universal” thresholding schemes that preserve singular values above a precomputed value Gavish_2014, usvt. A human-in-the-loop approach is to choose $k$ as the “elbow” point (see Figure (ref)) that partitions the singular values of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$ into those of large and small magnitudes. Note that if $\boldsymbol{E}_{\text{pre}, \mathcal{I}^{(d)}} = \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}]$ has independent sub-Gaussian rows, then the singular values of $\boldsymbol{E}_{\text{pre}, \mathcal{I}^{(d)}}$ scale as $O_p(\sqrt{T_0} + \sqrt{N_d})$ vershynin2018high. If the entries of $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}]$ are $\Theta(1)$ and its nonzero singular values are of the same magnitude, then they will scale as $\Theta(\sqrt{T_0 N_d / r_\text{pre}} )$, where $r_\text{pre} = \rank(\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}])$. Such a model leads to the aforementioned elbow point. Thus, one feasibility check for the suitability of PCR is to empirically examine if the smallest retained singular value $\widehat{s}_k$ scales quicker than $\sqrt{T_0} + \sqrt{N_d}$.

figure[figure omitted — 998 chars of source]

Formal Results

In this section, we analyze the SI-PCR estimator described in Section (ref).

Additional Assumptions for Consistency

We state additional assumptions required to establish the consistency of the SI-PCR estimator. Please refer to Sections (ref) and (ref) for a refresher of our earlier assumptions and notation.

assumption[Sub-Gaussian noise] For all $(t,i,d)$, $\varepsilon^{(d)}_{ti}$ are independent sub-Gaussian random variables with $\mathrm{Var}(\varepsilon^{(d)}_{ti} | \mathcal{LF} ) = \sigma^2$ and $\|\varepsilon^{(d)}_{ti} | \mathcal{LF} \|_{\psi_2} \le C\sigma$ for some constant $C > 0$.
assumption[Boundededness] For all $(t,i,d)$, we have $\mathbb{E}[Y_{ti}^{(d)} | \mathcal{LF} ] \in [-1,1]$.
assumption[Well-balanced spectra] Given $d$, we have $\kappa^{-1} \ge c$ and $\| \mathbb{E}[\boldsymbol{Y}_{\emph{pre}, \mathcal{I}^{(d)}} | \mathcal{E}] \|_F^2 \ge c' N_d T_0$, where $\kappa$ is the condition number of $\mathbb{E}[\boldsymbol{Y}_{\emph{pre}, \mathcal{I}^{(d)}} | \mathcal{E}]$ and and constants $c, c' > 0$.
assumption[Latent time-treatment factors] Given $d$, we have $u_t^{(d)} \in \emph{span}(\{u^{(0)}_\tau: \tau \le T_0\})$ for every $t > T_0$, conditional on $\mathcal{E}$.

We defer our discussions of Assumptions (ref)--(ref) to Section (ref).

Consistency Result

We now formally state our result characterizing the prediction accuracy of SI-PCR. For ease of notation, we absorb dependencies on $\sigma$ into the constant within $O_p(\cdot)$.

theoremGiven $(i,d)$, let Assumptions (ref) to (ref) hold. Consider the SI-PCR estimator given by (ref) and suppose $k = r_{\emph{pre}} = \emph{rank}(\mathbb{E}[ \boldsymbol{Y}_{\emph{pre}, \mathcal{I}^{(d)}} | \mathcal{E}])$. Then conditioned on $\mathcal{E}$, we have \begin{align} \widehat{\theta}_i^{(d)} - \theta_i^{(d)} = O_p \left(\sqrt{\log(T_0 N_d)} \left[ \frac{r^{3/4}_pre}{T_0^{1/4}} + r^{2}_pre \max\left\{\frac{\sqrt{N_d}}{T^{3/2}_0}, \frac{1}{\sqrt{T_0}}, \frac{1}{\sqrt{N_d}}\right\} \right] \right). \end{align} Here, we define $O_p(\cdot)$ with respect to the sequence $\min\{N_d, T_0\}$.

Theorem (ref) establishes that the SI-PCR estimator yields a consistent estimate of the causal estimand. More precisely, for a fixed $r_\text{pre}$, the estimation error decays as $N_d$ and $T_0$ grow, provided $T_0 = \omega(N_d^{1/3})$. Notably, the duration of the post-treatment period, $T_1$, can be of fixed length. In fact, if $T_1 = 1$, then Theorem (ref) establishes {\em point-wise consistency}. We retain the flexibility of allowing $T_1 > 1$ since a common goal in practice is to estimate the average counterfactual outcome over the post-treatment period.

remark[Asymptotic normality] {\em Appendix (ref) provides conditions that lead to asymptotic normality, which then allows for the construction of valid confidence intervals. There, we show that a minor modification to the SI-PCR estimator satisfies these conditions. }

Remarks on Additional Assumptions

remark[On Assumption (ref)] {\em While the latent time-treatment factors, $u_t^{(d)}$, can be arbitrarily correlated, Assumption (ref) enforces the idiosyncratic shocks to be independent. This can be restrictive. However, just as abadie2 used this independence structure to provide the first analysis of the SC estimator, we also adopt this independence structure to provide the first analysis of the SI estimator. A formal analysis of the SI estimator under more general noise models is left as important future work. }
remark[On Assumption (ref)] {\em The precise bound $[-1,1]$ is without loss of generality. That is, it can be extended to $[a, b]$ for $a, b \in \mathbb{R}$ with $a \le b$. }
remark[On Assumption (ref)] {\em Assumption (ref) requires that the nonzero singular values of $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}]$ are well-balanced, which we can empirically inspect through the spectral profile of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$. Within the econometrics factor model and matrix completion literatures, Assumption (ref) is analogous to incoherence-style conditions bai2020matrix and notions of pervasiveness fan2018eigenvector. Assumption (ref) has also been shown to hold w.h.p. for the canonical probabilistic generating process used to analyze probabilistic principal component analysis in bayesianpca and probpca; here, the observations are a high-dimensional embedding of a low-rank matrix with independent sub-Gaussian entries agarwal2020robustness. We highlight that the assumption of a gap between the top few singular values of observed matrix of interest and the remaining singular values has been widely adopted in the econometrics literature of large dimensional factor analysis dating back to chamberlainfactor. }
remark[On Assumption (ref)] {\em Assumption (ref) is analogous to Assumption (ref). As such, similar interpretations and implications carry over. Another interpretation of Assumption (ref) views $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$ as the training set and $\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}}$ as the test set. Intuitively, generalization relies on a notion of “similarity” between the train and test sets. Standard arguments from the statistical learning literature formalizes this notion through the assumption that the two sets are constructed from i.i.d. draws. Since potential outcomes from different treatments are likely to arise from different distributions, such a postulation is ill-suited for our setting. Instead, Assumption (ref) implies that generalization is achievable if the rowspace of $\mathbb{E}[\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}} | \mathcal{E}]$ is contained within that of $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}| \mathcal{E}]$ (see Lemma (ref) of Appendix (ref)), i.e., each test point is a linear combination of the training set---a purely linear algebraic condition. Though $\mathbb{E}[\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}}| \mathcal{E}]$ and $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}| \mathcal{E}]$ are not observable, we can still investigate the validity of Assumption (ref) using $\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}}$ and $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$ as proxies. To this end, we first recall that Assumption (ref) for unit $i$ is empirically tested through its pre-treatment fit, defined as \begin{align} \rho_i = \frac{ \| (\boldsymbol{I} - \widehat{\boldsymbol{U}}_pre \widehat{\boldsymbol{U}}_pre^\top) Y_{pre, i} \|_2} { \| Y_{pre, i} \|_2}, \end{align} where $\widehat{\boldsymbol{U}}_\text{pre} \in \mathbb{R}^{T_0 \times k}$ collects the top $k$ left singular vectors of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$; note that the numerator of (ref) measures the $\ell_2$-size of the in-sample residuals. Given the symmetry between Assumptions (ref) and (ref), it is natural to define for every $t > T_0$, \begin{align} \phi_t = \frac{ \| (\boldsymbol{I} - \widehat{\boldsymbol{V}}_pre \widehat{\boldsymbol{V}}_pre^\top) Y_{t, \mathcal{I}^{(d)}} \|_2} { \| Y_{t, \mathcal{I}^{(d)}} \|_2}, \end{align} where $\widehat{\boldsymbol{V}}_\text{pre} \in \mathbb{R}^{N_d \times k}$ collects the top $k$ right singular vectors of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$ and $Y_{t, \mathcal{I}^{(d)}} = [Y_{tjd}: j \in \mathcal{I}^{(d)}]$. Observe that $\rho_i$ and $\phi_t$ largely depend on the “richness” of the row and column spaces of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$, respectively. The row space, generated by $\mathcal{I}^{(d)}$, determines the existence of a synthetic unit $i$ formed as a weighted average of $\mathcal{I}^{(d)}$; the column space, generated by the pre-treatment outcomes (training set), controls the \textsf{SI}-\texttt{PCR} estimator's ability to generalize to the post-treatment period (test set). By construction, $\rho_i$ and $\phi_t$ take values in $[0,1]$ with lower values being more desirable but higher values being more informative: a value close to zero does not guarantee the tested assumption holds but values close to one suggest that the tested assumption is likely violated. Thus, $\rho_i$ and $\phi_t$ are useful one-sided tests for the suitability of the \textsf{SI} framework toward a study. The concept of generalization under a change in distribution between the training and test sets is known in the machine learning community as {\em covariate shift}. We hope that those exploring this area of research will find resonance with Assumption (ref) and its implications, which Section (ref) explores in greater depth. }

Simulation Study

This section presents a simulation on the SI-PCR estimator that complements the statistical claims of Theorem (ref). In particular, we evaluate the {\em point-wise} prediction accuracy of the SI-PCR estimator and study the role of Assumption (ref).

Data Generating Process

We set $T_1 = 1$ and vary $N_d = T_0 \in \{25, 50, 75, \dots, 200\}$. We set $r =15$ and $r_\text{pre} = 10$.

Latent factors

We generate the latent factors associated the units in $\mathcal{I}^{(d)}$, denoted as $\boldsymbol{V}_{\mathcal{I}^{(d)}} \in \mathbb{R}^{N_d \times r}$, by independently sampling its entries from a standard Normal distribution. Then, we sample $w^{(i,d)} \in \mathbb{R}^N_d$ by drawing its entries from a uniform distribution over $[0,1]$ and normalizing it to have unit norm. We define the target unit latent factor as $v_i = \boldsymbol{V}_{\mathcal{I}^{(d)}}^\top w^{(i,d)}$.

Define the latent factors for the pre-treatment outcomes under control as $\boldsymbol{U}^{(0)}_\text{pre} = \boldsymbol{A} \boldsymbol{B}^\top$, where the entries of $\boldsymbol{A} \in \mathbb{R}^{T_0 \times r_\text{pre}}$ and $\boldsymbol{B} \in \mathbb{R}^{r \times r_\text{pre}}$ are i.i.d. samples from a standard Normal. Note $\rank(\boldsymbol{U}^{(0)}_\text{pre}) = r_\text{pre}$. Next, we sample $\phi \in \mathbb{R}^r$ whose entries are i.i.d. draws from a uniform distribution over $[0, 1]$. Consider two post-treatment latent factors under treatment $d$: $u_\text{post}^{(d, \text{\ding{51}})} = (\boldsymbol{U}_\text{pre}^{(0)})^\dagger \boldsymbol{U}^{(0)}_\text{pre} \phi$ such that Assumption (ref) holds ($\text{\ding{51}}$) with respect to $\boldsymbol{U}^{(0)}_\text{pre}$ and $u_\text{post}^{(d, \text{\ding{55}})} = [ \boldsymbol{I} - (\boldsymbol{U}_\text{pre}^{(0)})^\dagger \boldsymbol{U}^{(0)}_\text{pre} ] \phi$ such that Assumption (ref) fails ($\text{\ding{55}}$) with respect to $\boldsymbol{U}^{(0)}_\text{pre}$. Thus, we define two causal estimands, $\theta_i^{(d, \text{\ding{51}})} = \langle u_\text{post}^{(d, \text{\ding{51}})}, v_i \rangle$ and $\theta_i^{(d, \text{\ding{55}})} = \langle u_\text{post}^{(d, \text{\ding{55}})}, v_i \rangle$.

Observations

Let (i) $Y_{\text{pre}, i} = \boldsymbol{U}^{(0)}_\text{pre} v_i + \varepsilon_{\text{pre}, i}$; (ii) $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} =\boldsymbol{U}^{(0)}_\text{pre} \boldsymbol{V}^\top_{\mathcal{I}^{(d)}} + \boldsymbol{E}_{\text{pre}, \mathcal{I}^{(d)}}$; (iii) $Y_{\text{post}, \mathcal{I}^{(d)}}^{(\text{\ding{51}})} = \boldsymbol{V}_{\mathcal{I}^{(d)}} u_\text{post}^{(d, \text{\ding{51}})} + \varepsilon_{\text{post}, \mathcal{I}^{(d)}}$; and (iv) $Y_{\text{post}, \mathcal{I}^{(d)}}^{(\text{\ding{55}})} = \boldsymbol{V}_{\mathcal{I}^{(d)}} u_\text{post}^{(d, \text{\ding{55}})} + \varepsilon_{\text{post}, \mathcal{I}^{(d)}}$, where the entries of $\varepsilon_{\text{pre}, i}$, $\boldsymbol{E}_{\text{pre}, \mathcal{I}^{(d)}}$, and $\varepsilon_{\text{post}, \mathcal{I}^{(d)}}$ are i.i.d. samples from a standard Normal.

Simulation Results

From $Y_{\text{pre}, n}$ and $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$, we learn a {\em single} linear model $\widehat{w}^{(i,d)}$ via the SI-PCR estimator with the hyper-parameter $k$ chosen via the approach prescribed in Gavish_2014. We then produce two estimates, $\widehat{\theta}_i^{(d, \text{\ding{51}})} = \langle Y_{\text{post}, \mathcal{I}^{(d)}}^{(\text{\ding{51}})}, \widehat{w}^{(i,d)} \rangle$ and $\widehat{\theta}_i^{(d, \text{\ding{55}})} = \langle Y_{\text{post}, \mathcal{I}^{(d)}}^{(\text{\ding{55}})}, \widehat{w}^{(i,d)} \rangle$. Figure (ref) visualizes the point-wise prediction errors of $| \widehat{\theta}_i^{(d, \text{\ding{51}})} - \theta_i^{(d, \text{\ding{51}})}|$ and $| \widehat{\theta}_i^{(d, \text{\ding{55}})} - \theta_i^{(d, \text{\ding{55}})}|$ across different dimensions and averaged over $50$ trials, where each trial consists of an independent draw of the latent factors and the observations of size $100$; this amounts to $5000$ total simulation repeats per dimension. We observe that the $\text{\ding{51}}$ point-wise errors decay with increasing dimensions, which corroborates with Theorem (ref). The $\text{\ding{55}}$ point-wise errors fail to converge, which suggests that Assumption (ref) can be critical for generalization.

figure[figure omitted — 772 chars of source]

Empirical Application

We revisit the classical Proposition 99 study of abadie2 introduced in Section (ref). Recall that the focus of abadie2 was to answer Q1. By contrast, our attention rests with Q2-Q5 given our interest in the SI-PCR estimator when its linear model is learned on outcomes observed under the status quo and applied on outcomes observed under an anti-tobacco program or raised taxes. More formally, whereas abadie2 aimed to estimate $\theta_i^{(0)}$, we set out to estimate $\theta_i^{(d)}$ for $d \in \{1,2\}$.

Setup

While we would ideally like to respond to Q2-Q5 directly, the answers to these questions are counterfactual, which precludes any evaluation. Accordingly, we emulate these questions via a leave-one-out (LOO) analysis. Specifically, for every treatment $d$, we iteratively designate a state $i \in \mathcal{I}^{(d)}$ as the target unit. We withhold its post-treatment outcomes, treating them as the ground truth, and define our causal estimand as $\theta_i^{(d)} = (1/T_1) \sum_{t > T_0} Y_{tid}$, acknowledging that our estimand is imperfect since it inherently incorporates the idiosyncratic shocks. We retain access to the target unit's pre-treatment outcomes, $Y_{ti0}$ for $t \le T_0$, as well as the outcomes for the remaining units in $\mathcal{I}^{(d)}$ for the entire time horizon, i.e., for all $j \in \mathcal{I}^{(d)} \setminus \{i\}$, we observe $Y_{tj0}$ for $t \le T_0$ and $Y_{tjd}$ for $t > T_0$. Our objective is to estimate the causal estimand from these observations with the SI-PCR estimator.

For each treatment $d$, we summarize the performance of the SI-PCR estimator as $\texttt{errors}(d) = [| (\widehat{\theta}_i^{(d)} - \theta_i^{(d)}) / \theta_i^{(d)}|: i \in \mathcal{I}^{(d)}]$, which collects the absolute normalized LOO prediction errors across $\mathcal{I}^{(d)}$. Practically speaking, $\texttt{errors}(d)$ measures the SI-PCR estimator's ability to recreate the post-treatment outcomes under treatment $d$ using a linear model learned from pre-treatment outcomes under control (status quo). With this interpretation and given the widespread acceptance of the SC framework---particularly, in the context of this dataset---we view $\texttt{errors}(0)$ as the baseline; hence, the difference between $\texttt{errors}(0)$ and $\texttt{errors}(d)$ for $d \in \{1,2\}$ serves as a measure for the validity of the SI framework. In what follows, the hyper-parameter $k$ of the \textsf{SI}-\texttt{PCR} estimator is chosen as the minimum number of singular values needed to capture $99\%$ of the spectral energy in $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$.

Empirical Results

Table (ref) reports on the average and standard deviation of $\texttt{errors}(d)$ across all treatments $d \in \{0,1,2\}$. Compared to the baseline, the average error and standard deviation associated with $d=2$ (raised taxes) is lower. At the same time, the average error for $d=1$ (anti-tobacco program) matches that of the baseline but its standard deviation is noticeably larger. This may be an artifact of there being significantly less states that imposed an anti-tobacco program ($N_1=5$) compared to keeping the status quo ($N_0 = 38$). However, if we redefine the treatments to take on two values as opposed to three, where we maintain the status quo as control ($d'=0$) but combine anti-tobacco programs and raised taxes into a single anti-tobacco measure ($d'=1$), then we find the resulting error summary (not reported in Table (ref)) to be $0.077 \pm 0.079$. Figure (ref) provides a visualization of the estimated trajectory of potential cigarette sales from 1989--2000 under the statewide status quo ($d=0$), an anti-tobacco program ($d=1$), and raised taxes ($d=2$) for three example states.

table[table omitted — 579 chars of source]

Collectively, our results provide empirical evidence that the SI methodology, like the SC methodology, is suitable for this case study. Moreover, our findings suggest that (i) anti-tobacco programs and raised taxes yielded similar effects with (ii) both leading to a significant decrease in tobacco consumption compared to the status quo, which is consistent with the conclusions drawn in abadie2 for California.

figure[figure omitted — 1,245 chars of source]

Conclusion

In this article, we reinterpret the problem of counterfactual estimation as one of tensor completion. To enable faithful recovery of the underlying tensor, we anchor on the low-rank tensor factor model, which is arguably the de-facto assumption in the tensor completion literature. The tensor factor model is a natural generalization of the canonical matrix factor model used to justify the SC framework with the added assumption of a latent factorization over treatments and shared latent unit factors across not only time (as in the matrix factor model) but treatments as well. Accordingly, we propose the SI framework, which rests upon this structural assumption. As a consequence, our proposed framework extends the SC framework to the setting of multiple treatments and, therefore, provides one resolution to the open question posed in abadie_survey.

We couple our framework with an estimation strategy that generalizes the standard SC-based estimators, and show that one variation of our approach, SI-PCR, is a consistent estimator of our causal estimand of interest. Our statistical claims are supplemented by a simulation study that highlights the substantive implications of a linear algebraic condition (Assumption (ref)) for the prediction accuracy of the SI-PCR variant. Finally, our empirical study provides evidence that the SI methodology, like the \textsf{SC} methodology, may be suitable for the classical tobacco study of abadie2. Complementing the conclusions of abadie2, our findings suggest that anti-tobacco programs and raised taxes had similar effects with both leading to a significant reduction in tobacco consumption compared to the status quo.

We envision several interesting avenues of future research. Keeping with our tensor factor model, an analysis of the SI-PCR estimator under less restrictive assumptions on the noise model (Assumption (ref)) would offer deeper insights into its predictive behavior for more general settings. More broadly, formal analyses on different formulations of the SI estimator may provide guidance to researchers on the choice of formulation for their study, should the SI framework be suitable. In light of the recent work of bottmer that provided the first design-based analysis of the SC method, it would also be fascinating to understand the design-based properties of the the SI method to study its utility for randomized experiments.

appendix\setcounter{equation}{0} \begin{center} {\bf SUPPLEMENTARY MATERIAL} \end{center} The supplementary material is structured as follows. Appendices (ref), (ref), and (ref) prove Theorems (ref) and (ref). Appendix (ref) expounds on Remark (ref) regarding non-linear latent variable models and low-rank factor models. Appendix (ref) extends the SI framework to incorporate auxiliary covariates. Appendix (ref) continues our discussion on a connection between causal inference and tensor completion. Appendix (ref) elaborates on Remark (ref) regarding inference and a modified version of the SI-PCR estimator. \section{Proof of Theorem (ref)} \begin{proof} We have that \begin{align} \mathbb{E} [ Y^{(d)}_{ti} \big| u^{(d)}_t, v_i ] & = \mathbb{E} [\langle u^{(d)}_t, v_i \rangle + \varepsilon^{(d)}_{ti} \big| u^{(d)}_t, v_i ] &&\because Assumption (ref) \& = \langle u^{(d)}_t, v_i \rangle \big| \{u^{(d)}_t, v_i\} &&\because Assumption (ref) \&= \langle u^{(d)}_t, v_i \rangle \big| \mathcal{E} \&= \langle u^{(d)}_t, \sum_{j \in \mathcal{I}^{(d)}} w_j^{(i,d)} v_j \rangle \big| \mathcal{E} &&\because Assumption (ref) \& = \sum_{j \in \mathcal{I}^{(d)}} w_j^{(i,d)} \cdot \mathbb{E} \left[ \langle u^{(d)}_t, v_j \rangle + \varepsilon^{(d)}_{tj} \big| \mathcal{E} \right] &&\because \text{Assumption (ref)} \\ & = \sum_{j \in \mathcal{I}^{(d)}} w_j^{(i,d)} \cdot \mathbb{E} [ Y^{(d)}_{tj} \big| \mathcal{E} ] &&\because \text{Assumption (ref)} \& = \sum_{j \in \mathcal{I}^{(d)}} w_j^{(i,d)} \cdot \mathbb{E} \left[ Y_{tjd} | \mathcal{E} \right]. &&\because \text{Assumption (ref)} \end{align} The third equality follows since $\langle u^{(d)}_t, v_i \rangle$ is deterministic conditioned on $\{u^{(d)}_t, v_i\}$. Recalling (ref), one can verify (ref) using the same the argument used to prove (ref) but conditioning on the set of latent factors $\{u^{(d)}_t, v_i: t > T_0 \}$ at the start rather than just $\{u^{(d)}_t, v_i\}$. \end{proof} \section{Helper Lemmas to Prove Theorem (ref)} In this section, we state and prove several helper lemmas to establish Theorem (ref). All $O_p(\cdot)$ statements are with respect to the sequence $\min\{T_0, N_d\}$. \subsection{Perturbation of Singular Values} In Lemma (ref), we argue that the singular values of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$ and $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}]$ are close. Recall that $s_\ell$ and $\widehat{s}_\ell$ are the singular values of $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}]$ and $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$, respectively. \begin{lemma} Let Assumptions (ref), (ref), (ref), (ref), (ref) and (ref) hold. Conditioned on $\mathcal{E}$, we have for any $d$, $\zeta > 0$, and $\ell \le \min\{T_0, N_d\}$, \begin{align} \abs{s_\ell - \widehat{s}_\ell} \le C\sigma \left( \sqrt{T_0} + \sqrt{N_d} + \zeta \right) \end{align} w.p. at least $1-2\exp(-\zeta^2)$, where $C>0$ is an absolute constant. \end{lemma} \begin{proof} {Proof of Lemma (ref).} For ease of notation, we suppress the conditioning on $\mathcal{E}$ for the remainder of the proof. To bound the gap between $s_\ell$ and $\widehat{s}_\ell$, we recall the following well-known results. \begin{lemma} [Weyl's inequality] Given $\boldsymbol{A}, \boldsymbol{B} \in \mathbb{R}^{m \times n}$, let $\sigma_\ell$ and $\widehat{\sigma}_\ell$ be the $\ell$-th singular values of $\boldsymbol{A}$ and $\boldsymbol{B}$, respectively, in decreasing order and repeated by multiplicities. Then for all $i \le \min\{m, n\}$, $\abs{ \sigma_\ell - \widehat{\sigma}_\ell} \le \|\boldsymbol{A} - \boldsymbol{B} \|_\emph{op}.$ \end{lemma} \begin{lemma} [Sub-Gaussian Matrices: Theorem 4.4.5 of vershynin2018high] Let $\boldsymbol{A} = [A_{ij}]$ be an $m \times n$ random matrix where the entries $A_{ij}$ are independent, mean zero, sub-Gaussian random variables. Then for any $t > 0$, we have $\| \boldsymbol{A} \|_{\emph{op}} \le CK \left(\sqrt{m} + \sqrt{n} + t \right)$ w.p. at least $1-2\exp(-t^2)$. Here, $K = \max_{i,j} \|A_{ij}\|_{\psi_2}$ and $C>0$ is an absolute constant. \end{lemma} By Lemma (ref), we have for any $\ell \le \min\{T_0, N_d\}$, $\abs{s_\ell - \widehat{s}_\ell} \leq \| \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}]\|_\text{op}$. Recalling Assumption (ref) and applying Lemma (ref), we conclude for any $t>0$ and some absolute constant $C>0$, $\abs{s_\ell - \widehat{s}_\ell} \le C\sigma\left(\sqrt{T_0} + \sqrt{N_d} + t \right)$ w.p. at least $1-2\exp(-t^2)$. This completes the proof. \end{proof} \subsection{On the Rowspaces of the Pre- and Post-Intervention Expected Outcomes} Similar to Assumption (ref), observe that Assumption (ref) implies that there exists a $\beta^{(d)} \in \mathbb{R}^{T_0}$ such that \begin{align} u_t^{(d)} = \sum_{\tau \le T_0} \beta_\tau^{(d)} \cdot u^{(0)}_\tau. \end{align} With this interpretation, we present a similar result to Theorem (ref) that relates the rowspaces between the pre- and post-intervention expected outcomes. \begin{lemma} Given $d$, let Assumptions (ref), (ref), (ref), and (ref) hold. Then given $\beta^{(d)}$ as defined in (ref), we have for all $t > T_0$ and $j \in [N]$, \begin{align} \mathbb{E} [ Y_{tjd} | \mathcal{E}] &= \sum_{\tau \le T_0} \beta_\tau^{(d)} \cdot \mathbb{E}[ Y_{\tau j 0} | \mathcal{E}]. \end{align} \end{lemma} \begin{proof} We have that \begin{align} \mathbb{E}[Y_{tjd} | \mathcal{E}] &= \mathbb{E}[Y^{(d)}_{tj} | \mathcal{E}] && \because \text{Assumption (ref)} \\ &= \mathbb{E}[ \langle u_t^{(d)}, v_j \rangle + \varepsilon_{tj}^{(d)} | \mathcal{E}] && \because \text{Assumption (ref)} \\ &= \langle u_t^{(d)}, v_j \rangle | \mathcal{E} && \because \text{Assumption (ref)} \\ &= \langle \sum_{\tau \le T_0} \beta^{(d)}_\tau \cdot u_\tau^{(0)}, v_j \rangle | \mathcal{E} && \because \text{Assumption (ref)} \\ &= \sum_{\tau \le T_0} \beta^{(d)}_\tau \cdot \langle u_\tau^{(0)}, v_j \rangle | \mathcal{E} \\ &= \sum_{\tau \le T_0} \beta^{(d)}_\tau \cdot \mathbb{E}[ \langle u_\tau^{(0)}, v_j \rangle + \varepsilon_{\tau j}^{(0)} | \mathcal{E} ] && \because \text{Assumption (ref)} \\ &= \sum_{\tau \le T_0} \beta^{(d)}_\tau \cdot \mathbb{E}[Y^{(0)}_{\tau j} | \mathcal{E}] && \because \text{Assumption (ref)} \\ &= \sum_{\tau \le T_0} \beta^{(d)}_\tau \cdot \mathbb{E}[Y_{\tau j 0} | \mathcal{E}]. && \because \text{Assumption (ref)} \end{align} This completes the proof. \end{proof} \subsection{Rewriting Theorem (ref) with the Minimum $\ell_2$-Norm Solution} Let $\boldsymbol{V}_\text{pre} \in \mathbb{R}^{N_d \times r_{\text{pre}}}$ and $\boldsymbol{V}_\text{post} \in \mathbb{R}^{N_d \times r_{\text{post}}}$ denote the right singular vectors of $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E} ]$ and $\mathbb{E}[\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}} | \mathcal{E} ]$, respectively. The next result rewrites Theorem (ref) via the following: \begin{align} \widetilde{w}^{(i,d)} = \boldsymbol{V}_\text{pre} \boldsymbol{V}_\text{pre}^\top w^{(i,d)}. \end{align} In words, $\widetilde{w}^{(i,d)}$ is the orthogonal projection of $w^{(i,d)}$ onto the rowspace of $\mathbb{E}[ \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}]$ and is the minimum $\ell_2$-norm solution to (ref) and (ref) of Theorem (ref). \begin{lemma} Given $(i, d)$, let the setup of Theorem (ref) and Assumption (ref) hold. Then, we have \begin{align} \theta^{(d)}_i = \frac{1}{T_1} \sum_{t > T_0} \sum_{j \in \mathcal{I}^{(d)}} \widetilde{w}_j^{(i,d)} \cdot \mathbb{E}[Y_{tjd} | \mathcal{E}]. \end{align} \end{lemma} \begin{proof} From Lemma (ref), it follows that the rowspan of $\mathbb{E}[\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}} | \mathcal{E}]$ lies within that of $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}]$. Hence, \begin{align} \boldsymbol{V}_\text{post} &= \boldsymbol{V}_\text{pre} \boldsymbol{V}_\text{pre}^\top \boldsymbol{V}_\text{post}. \end{align} Therefore, we have \begin{align} \mathbb{E}[\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}} | \mathcal{E}] \cdot \widetilde{w}^{(i,d)} &= \mathbb{E}[\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}} | \mathcal{E}] \boldsymbol{V}_\text{pre} \boldsymbol{V}_\text{pre}^\top \cdot w^{(i,d)} \nonumber \\ & = \mathbb{E}[\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}} | \mathcal{E}] \cdot w^{(i,d)}, \end{align} where we use (ref) in the last equality. Hence, we conclude \begin{align} \theta^{(d)}_i &= \frac{1}{T_1} \sum_{t > T_0} \sum_{j \in \mathcal{I}^{(d)}} w^{(i,d)}_j \cdot \mathbb{E}[Y_{tjd} | \mathcal{E} ] = \frac{1}{T_1} \sum_{t > T_0} \sum_{j \in \mathcal{I}^{(d)}} \widetilde{w}^{(i,d)}_j \cdot \mathbb{E}[Y_{tjd} | \mathcal{E} ]. \end{align} The first equality follows from (ref) in Theorem (ref) and the second equality follows from (ref). \end{proof} \subsection{Model Identification} \begin{lemma} [Corollary 4.1 of pcr_aos] Given $(i,d)$, let Assumptions (ref) to (ref) hold. Further, suppose $k = r_\emph{pre} = \rank(\mathbb{E}[\boldsymbol{Y}_{\emph{pre}, \mathcal{I}^{(d)}} \, | \, \mathcal{E}])$, where $k$ is defined as in (ref). Then conditioned on $\mathcal{E}$, we have \begin{align} \widehat{w}^{(i,d)} - \widetilde{w}^{(i,d)} &= O_p\left(\sqrt{\log(T_0 N_d)} \left[ \frac{r^{3/4}_\emph{pre} }{T^{1/4}_0 N_d^{1/2}} + \frac{r^{3/2}_\emph{pre} }{\min\{T_0,N_d\}} \right] \right). \end{align} \end{lemma} \begin{proof} We first restate pcr_aos. \begin{lemma}[Corollary 4.1 of pcr_aos] Let the setup of Lemma (ref) hold. Then w.p. at least $1 - O(1 / (T_0 N_d)^{10})$, \begin{align} \|\widehat{w}^{(i,d)} - \widetilde{w}^{(i,d)}\|^2_2 \le C(\sigma) \cdot \log(T_0 N_d) \left( \frac{r^{3/2}_\emph{pre} }{T^{1/2}_0 N_d} + \frac{r^{3}_\emph{pre} }{\min\{T_0,N_d\}^2}\right), \end{align} where $C(\sigma)$ is a constant that only depends on $\sigma$. \end{lemma} The proof follows by adapting the notation in pcr_aos to that used in this article. In particular, $y = Y_{\text{pre}, i}, \boldsymbol{X} = \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}], \widetilde{\boldsymbol{Z}} = \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}, \widehat{\beta} = \widehat{w}^{(i,d)}, \beta^* = \widetilde{w}^{(i,d)}, T_0 = n, N_d = p, \rho = 1, d = 1$, where $y, \boldsymbol{X}, \widetilde{\boldsymbol{Z}}, \widehat{\beta}, \beta^*, n, p, \rho, d$ are the notations used in pcr_aos; since we do not consider missing values in this work, $\widetilde{\boldsymbol{Z}} = \boldsymbol{Z}$, where $\boldsymbol{Z}$ is the notation used in pcr_aos. We also use the fact that $\boldsymbol{X} \beta^*$ in the notation of pcr_aos, i.e., $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E} ] \cdot \widetilde{w}^{(i,d)}$ in our notation, equals $\mathbb{E}[Y_{\text{pre}, i} | \mathcal{E}]$. This follows from \begin{align} \mathbb{E}[Y_{\text{pre}, i} | \mathcal{E}] = \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}] \cdot \widetilde{w}^{(i,d)}, \end{align} which follows from (ref) in Theorem (ref) and (ref). We conclude by noting (ref) gives our desired result. \end{proof} \begin{lemma}[Lemma 19 of synthetic_combinations] Given $(i,d)$, let Assumptions (ref) to (ref) hold. Then conditioned on $\mathcal{E}$, we have that $$\| \widetilde{w}^{(i,d)} \|_2 \le C \sqrt{\frac{r_\emph{pre}}{N_d}}$$ for some constant $C > 0$. This immediately implies that $\| \widetilde{w}^{(i,d)} \|_1 \le C \sqrt{r_\emph{pre}}.$ \end{lemma} \begin{proof} As in the proof of Lemma (ref), the result immediately holds by matching notation to that of synthetic_combinations. \end{proof} \section{Proof of Theorem (ref)} Armed with our helper lemmas in Section (ref), we are ready to formally establish Theorem (ref). We adopt the following notation. For any $t > T_0$, let $Y_{t, \mathcal{I}^{(d)}} = [Y_{tjd}: j \in \mathcal{I}^{(d)}] \in \mathbb{R}^{N_d}$ and $\varepsilon_{t, \mathcal{I}^{(d)}} = [\varepsilon^{(d)}_{tj}: j \in \mathcal{I}^{(d)}] \in \mathbb{R}^{N_d}$. Note that the rows of $\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}}$ are formed by $\{Y_{t, \mathcal{I}^{(d)}}: t > T_0\}$. Additionally, let $\Delta^{(i,d)} = \widehat{w}^{(i,d)} - \widetilde{w}^{(i,d)}$. Finally, for any matrix $\boldsymbol{A}$ with orthonormal columns, let $\mathcal{P}_A = \boldsymbol{A} \boldsymbol{A}^\top$ denote the projection matrix onto the subspace spanned by the columns of $\boldsymbol{A}$. We recall that all $O_p(\cdot)$ statements are with respect to the sequence $\min(T_0, N_d)$. \begin{proof} For ease of notation, we suppress the conditioning on $\mathcal{E}$ throughout the proof. By (ref) and Lemma (ref) (implied by Theorem (ref)), we have \begin{align} \widehat{\theta}^{(d)}_i - \theta^{(d)}_i &= \frac{1}{T_1} \sum_{t > T_0} \left( \langle Y_{t, \mathcal{I}^{(d)}}, \widehat{w}^{(i,d)} \rangle - \langle \mathbb{E}[ Y_{t, \mathcal{I}^{(d)}}], \widetilde{w}^{(i,d)} \rangle \right) \\ &= \frac{1}{T_1} \sum_{t > T_0} \left( \langle \mathbb{E}[Y_{t, \mathcal{I}^{(d)}}], \Delta^{(i,d)} \rangle + \langle \varepsilon_{t, \mathcal{I}^{(d)}}, \widetilde{w}^{(i,d)} \rangle + \langle \varepsilon_{t, \mathcal{I}^{(d)}},\Delta^{(i,d)} \rangle \right), \end{align} where we have used $Y_{t, \mathcal{I}^{(d)}} = \mathbb{E}[ Y_{t, \mathcal{I}^{(d)}}] + \varepsilon_{t, \mathcal{I}^{(d)}}$. From Lemma (ref), it follows that $\boldsymbol{V}_\text{post}= \mathcal{P}_{V_\text{pre}} \boldsymbol{V}_\text{post}$, where $\boldsymbol{V}_\text{pre}$ and $\boldsymbol{V}_\text{post}$ are the right singular vectors of $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}]$ and $\mathbb{E}[\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}}]$, respectively. Hence, $\mathbb{E}[\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}}] = \mathbb{E}[\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}}] \mathcal{P}_{V_\text{pre}}$. As such, for any $t > T_0$, \begin{align} \langle \mathbb{E}[Y_{t, \mathcal{I}^{(d)}}], \Delta^{(i,d)} \rangle &= \langle \mathbb{E}[Y_{t, \mathcal{I}^{(d)}}], \mathcal{P}_{V_\text{pre}} \Delta^{(i,d)} \rangle. \end{align} Plugging (ref) into (ref) yields \begin{align} \widehat{\theta}^{(d)}_i - \theta^{(d)}_i &= \frac{1}{T_1}\sum_{t > T_0} \left(\langle \mathbb{E}[Y_{t, \mathcal{I}^{(d)}}], \mathcal{P}_{V_\text{pre}} \Delta^{(i,d)} \rangle + \langle \varepsilon_{t, \mathcal{I}^{(d)}}, \widetilde{w}^{(i,d)} \rangle + \langle \varepsilon_{t, \mathcal{I}^{(d)}}, \Delta^{(i,d)} \rangle \right). \end{align} Below, we bound the three terms on the right-hand side (RHS) of (ref) separately. {\em Bounding term 1.} By Cauchy-Schwartz inequality, observe that $$ \langle \mathbb{E}[Y_{t, \mathcal{I}^{(d)}}], \mathcal{P}_{V_\text{pre}} \Delta^{(i,d)} \rangle \le \| \mathbb{E}[Y_{t, \mathcal{I}^{(d)}}] \|_2 \cdot \| \mathcal{P}_{V_\text{pre}} \Delta^{(i,d)} \|_2. $$ Under Assumption (ref), we have $\| \mathbb{E}[Y_{t, \mathcal{I}^{(d)}}] \|_2 \le \sqrt{N_d}$. As such, \begin{equation} \frac{1}{T_1} \sum_{t > T_0} \langle \mathbb{E}[Y_{t, \mathcal{I}^{(d)}}], \mathcal{P}_{V_\text{pre}} \Delta^{(i,d)} \rangle \le \sqrt{N_d} \cdot \| \mathcal{P}_{V_\text{pre}} \Delta^{(i,d)} \|_2. \end{equation} Hence, it remains to bound $\| \mathcal{P}_{V_\text{pre}} \Delta^{(i,d)} \|_2$. Towards this, we state the following lemma. Its proof can be found in Appendix (ref). \begin{lemma} Consider the setup of Theorem (ref). Then, $$ \mathcal{P}_{V_\emph{pre}} \Delta^{(i,d)} = O_p \left( \frac{\sqrt{r_\emph{pre}}}{\sqrt{N_d} T_0^{\frac 1 4}} + \frac{r^{3/2}_\emph{pre} \sqrt{\log(T_0 N_d)}}{\sqrt{N_d} \cdot \min\{\sqrt{T_0}, \sqrt{N_d} \}} + \frac{r^{2}_\emph{pre} \sqrt{\log(T_0 N_d)}}{\min\{T_0^{3/2},N_d^{3/2}\}} \right). $$ \end{lemma} Using Lemma (ref), we obtain \begin{align} & \frac{1}{T_1} \sum_{t > T_0} \langle \mathbb{E}[Y_{t, \mathcal{I}^{(d)}}], \mathcal{P}_{V_\text{pre}} \Delta^{(i,d)} \rangle \\ &= O_p \Bigg( \frac{r^{1/2}_\text{pre}}{T_0^{\frac 1 4}} + \frac{r^{3/2}_\text{pre} \sqrt{\log(T_0 N_d)}}{\min\{\sqrt{T_0}, \sqrt{N_d} \}} + \frac{r^{2}_\text{pre} \sqrt{N_d \log(T_0 N_d)}}{\min\{T_0^{3/2},N_d^{3/2}\}} \Bigg). \qquad \end{align} This concludes the analysis for the first term. {\em Bounding term 2. } We begin with a simple lemma, which immediately holds by Hoeffding's Lemma (Lemma (ref)). \begin{lemma} Let $\gamma_t$ be a sequence of independent mean zero sub-Gaussian random variables with variance less than $\sigma^2$. Then, for any vector $\boldsymbol{a} = (a_1, \dots, a_T) \in \mathbb{R}^T$, we have $ \sum^T_{t = 1} a_t \gamma_t = O_p\left( \sigma \| \boldsymbol{a} \|_2 \right). $ \end{lemma} By Assumptions (ref) and (ref), as well as Lemma (ref), we have \begin{align} \sum_{t > T_0} \langle \varepsilon_{t, \mathcal{I}^{(d)}}, \widetilde{w}^{(i,d)} \rangle &= O_p \left( \sqrt{T_1} \| \widetilde{w}^{(i,d)} \|_2 \right). \end{align} Here, we have used the fact that the components of $\langle \varepsilon_{t, \mathcal{I}^{(d)}}, \widetilde{w}^{(i,d)} \rangle$ are independent mean-zero sub-Gaussians and that $\varepsilon_{t, \mathcal{I}^{(d)}}$ are independent across $t$. Then using Lemma (ref), we have that \begin{align} \frac{1}{T_1}\sum_{t > T_0} \langle \varepsilon_{t, \mathcal{I}^{(d)}}, \widetilde{w}^{(i,d)} \rangle &= O_p \left( \sqrt{\frac{r_\text{pre}}{N_d T_1}} \right). \end{align} {\em Bounding term 3. } First, we define the event $\mathcal{E}_1$ as $$ \mathcal{E}_1 = \left\{ \|\Delta^{(i,d)} \|_2 = O\left(\sqrt{\log(T_0 N_d)} \left[ \frac{r^{3/4}_\text{pre} }{T^{1/4}_0 N_d^{1/2}} + \frac{r^{3/2}_\text{pre} }{\min\{T_0,N_d\}} \right] \right) \right\}. $$ By Lemma (ref), $\mathcal{E}_1$ occurs w.h.p. Next, we define $\mathcal{E}_2$ as $$ \mathcal{E}_2 = \Bigg\{ \sum_{t > T_0} \langle \varepsilon_{t, \mathcal{I}^{(d)}}, \Delta^{(i,d)} \rangle = O\left( \sqrt{T_1 \log(T_0 N_d)} \left[ \frac{r^{3/4}_\text{pre} }{T^{1/4}_0 N_d^{1/2}} + \frac{r^{3/2}_\text{pre} }{\min\{T_0,N_d\}}\right] \right) \Bigg\}. $$ Now, condition on $\mathcal{E}_1$. Since $\widehat{w}^{(i,d)}$ only depends on $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$ and $Y_{\text{pre}, i}$, it is independent of $\varepsilon_{t, \mathcal{I}^{(d)}}$ for all $t > T_0$. Hence, Lemmas (ref) and (ref) imply $\mathcal{E}_2 | \mathcal{E}_1$ occurs w.h.p. Further, we note \begin{align} \mathbb{P}(\mathcal{E}_2) = \mathbb{P}(\mathcal{E}_2 | \mathcal{E}_1) \mathbb{P}(\mathcal{E}_1) + \mathbb{P}(\mathcal{E}_2 | \mathcal{E}_1^c) \mathbb{P}(\mathcal{E}_1^c) \ge \mathbb{P}(\mathcal{E}_2 | \mathcal{E}_1) \mathbb{P}(\mathcal{E}_1). \end{align} Since $\mathcal{E}_1$ and $\mathcal{E}_2 | \mathcal{E}_1$ occur w.h.p, it follows from (ref) that $\mathcal{E}_2$ occurs w.h.p. As a result, \begin{align} \frac{1}{T_1} \sum_{t > T_0} \langle \varepsilon_{t, \mathcal{I}^{(d)}}, \Delta^{(i,d)} \rangle &= O_p \left( \frac{\sqrt{\log(T_0 N_d)}}{T^{1/2}_1} \left[ \frac{r^{3/4}_\text{pre} }{T^{1/4}_0 N_d^{1/2}} + \frac{r^{3/2}_\text{pre} }{\min\{T_0,N_d\}} \right] \right). \end{align} {\em Collecting terms. } Plugging (ref), (ref), (ref) into (ref), and simplifying gives our desired result. \end{proof} \subsection{Proof of Lemma (ref)} We introduce some helpful notation. Let $\boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} = \sum_{\ell=1}^{r_{\text{pre}}} \widehat{s}_\ell \widehat{u}_\ell \widehat{v}^\top_\ell$ be the rank-$r_\text{pre}$-approximation of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$. More compactly, $\boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} = \widehat{\boldsymbol{U}}_\text{pre} \widehat{\boldsymbol{S}}_\text{pre} \widehat{\boldsymbol{V}}^\top_\text{pre}$. \begin{proof} To establish Lemma (ref), consider the following decomposition: $$ \mathcal{P}_{V_\text{pre}} \Delta^{(i,d)} = (\mathcal{P}_{V_\text{pre}} - \mathcal{P}_{\widehat{V}_\text{pre}}) \Delta^{(i,d)} + \mathcal{P}_{\widehat{V}_\text{pre}} \Delta^{(i,d)}. $$ We proceed to bound each term separately. {\em Bounding term 1. } Recall $\| \boldsymbol{A} v \|_2 \le \| \boldsymbol{A} \|_\text{op} \cdot \| v \|_2$ for any $\boldsymbol{A} \in \mathbb{R}^{a \times b}$ and $v \in \mathbb{R}^b$. Thus, \begin{align} \| (\mathcal{P}_{V_\text{pre}} - \mathcal{P}_{\widehat{V}_\text{pre}}) \Delta^{(i,d)} \|_2 \le \| \mathcal{P}_{V_\text{pre}} - \mathcal{P}_{\widehat{V}_\text{pre}}\|_\text{op} \cdot \| \Delta^{(i,d)} \|_2. \end{align} To control (ref), we state a helper lemma that bounds the distance between the subspaces spanned by the columns of $\boldsymbol{V}_\text{pre}$ and $\widehat{\boldsymbol{V}}_\text{pre}$. Its proof is given in Appendix (ref). \begin{lemma} Consider the setup of Theorem (ref). Then for any $\zeta > 0$, we have w.p. at least $1 - 2\exp(-\zeta^2)$ \begin{align} \| \mathcal{P}_{\widehat{V}_\emph{pre}} - \mathcal{P}_{V_\emph{pre}} \|_\emph{op} \le \frac{C\sigma \left(\sqrt{T_0} + \sqrt{N_d} + \zeta \right)}{s_{r_\emph{pre}}}, \end{align} where $s_{r_\emph{pre}}$ is the $r_\emph{pre}$-th singular value of $\mathbb{E}[\boldsymbol{Y}_{\emph{pre}, \mathcal{I}^{(d)}}]$ with $r_\emph{pre} = \rank(\mathbb{E}[\boldsymbol{Y}_{\emph{pre}, \mathcal{I}^{(d)}}])$, and $C>0$ is an absolute constant. \end{lemma} Applying Lemma (ref) with Assumption (ref), we have \begin{align} \| \mathcal{P}_{\widehat{V}_\text{pre}} - \mathcal{P}_{V_\text{pre}} \|_\text{op} = O_p \left( \frac{\sqrt{r_\text{pre}}}{\min\{\sqrt{T_0}, \sqrt{N_d}\}}\right). \end{align} Substituting (ref) and the bound in Lemma (ref) into (ref), we obtain \begin{align} (\mathcal{P}_{V_\text{pre}} - \mathcal{P}_{\widehat{V}_\text{pre}}) \Delta^{(i,d)} &= O_p\left(\frac{\sqrt{r_\text{pre} \log(T_0 N_d)}}{\min\{\sqrt{T_0}, \sqrt{N_d}\}} \left[ \frac{r^{3/4}_\text{pre} }{T^{1/4}_0 N_d^{1/2}} + \frac{r^{3/2}_\text{pre} }{\min\{T_0,N_d\}} \right] \right). \end{align} {\em Bounding term 2. } Since $\widehat{\boldsymbol{V}}_{\text{pre}}$ is an isometry, it follows that \begin{align} \| \mathcal{P}_{\widehat{V}_{\text{pre}}} \Delta^{(i,d)} \|_2^2 & = \|\widehat{\boldsymbol{V}}_{\text{pre}}^\top \Delta^{(i,d)} \|_2^2. \end{align} We upper bound $\|\widehat{\boldsymbol{V}}_{\text{pre}}^\top \Delta^{(i,d)} \|_2^2$ as follows: consider \begin{align} &\| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \Delta^{(i,d)} \|_2^2 = (\widehat{\boldsymbol{V}}_{\text{pre}}^\top \Delta^{(i,d)})^\top \widehat{\boldsymbol{S}}_{r_\text{pre}}^2 (\widehat{\boldsymbol{V}}_{\text{pre}}^\top \Delta^{(i,d)}) \geq \widehat{s}_{r_\text{pre}}^2 \cdot \| \widehat{\boldsymbol{V}}_{\text{pre}}^\top \Delta^{(i,d)} \|_2^2 . \end{align} Using (ref) and (ref) together implies \begin{align} \| \mathcal{P}_{\widehat{V}_{\text{pre}}} \Delta^{(i,d)} \|_2^2 \le \frac{\| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \Delta^{(i,d)} \|_2^2}{\widehat{s}_{r_\text{pre}}^2}. \end{align} To bound the numerator in (ref), note \begin{align} \| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \Delta^{(i,d)} \|_2^2 & \leq 2\| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widehat{w}^{(i,d)} -\mathbb{E}[Y_{\text{pre}, i}] \|_2^2 + 2 \| \mathbb{E}[Y_{\text{pre}, i}] -\boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widetilde{w}^{(i,d)}\|_2^2 \\ & = 2 \| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widehat{w}^{(i,d)} -\mathbb{E}[Y_{\text{pre}, i}] \|_2^2 + 2 \| (\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}] - \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}}) \widetilde{w}^{(i,d)}\|_2^2, \end{align} where we have used (ref). To further upper bound the the second term on the RHS above, we use the following inequality: for any $\boldsymbol{A} \in \mathbb{R}^{a \times b}$, $v \in \mathbb{R}^{b}$, \begin{align} \| \boldsymbol{A} v \|_2 & = \| \sum_{j=1}^b \boldsymbol{A}_{\cdot j} v_j \|_2 \leq \left(\max_{j \le b} \| \boldsymbol{A}_{\cdot j}\|_2 \right) \cdot \left(\sum_{j=1}^b |v_j|\right) = \| \boldsymbol{A}\|_{2, \infty} \cdot \|v\|_1, \end{align} where $\|\boldsymbol{A}\|_{2,\infty} = \max_{j} \| \boldsymbol{A}_{\cdot j}\|_2$ and $\boldsymbol{A}_{\cdot j}$ represents the $j$-th column of $\boldsymbol{A}$. Thus, \begin{align} \| (\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}] - \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}}) \widetilde{w}^{(i,d)}\|_2^2 \le \| \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}] - \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \|_{2,\infty}^2 \cdot \|\widetilde{w}^{(i,d)}\|_1^2. \end{align} Substituting (ref) into (ref) and subsequently using (ref) implies \begin{align} \| \mathcal{P}_{\widehat{V}_{\text{pre}}} \Delta^{(i,d)} \|_2^2 &\leq \frac{2}{\widehat{s}_{r_\text{pre}}^2} \Big( \| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widehat{w}^{(i,d)} -\mathbb{E}[Y_{\text{pre}, i}] \|_2^2 + \| \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}] - \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \|_{2,\infty}^2 \cdot \|\widetilde{w}^{(i,d)}\|_1^2\Big). \end{align} Next, we bound $\| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widehat{w}^{(i,d)} -\mathbb{E}[Y_{\text{pre}, i}] \|_2^2$. To this end, observe that \begin{align} \| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widehat{w}^{(i,d)} - Y_{\text{pre}, i} \|_2^2 & = \| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widehat{w}^{(i,d)} - \mathbb{E}[Y_{\text{pre}, i}] - \varepsilon_{\text{pre}, i}\|_2^2 \& = \| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widehat{w}^{(i,d)} - \mathbb{E}[Y_{\text{pre}, i}] \|_2^2 + \|\varepsilon_{\text{pre}, i}\|_2^2 \\ &\quad - 2 \langle \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widehat{w}^{(i,d)} - \mathbb{E}[Y_{\text{pre}, i}], \varepsilon_{\text{pre}, i} \rangle. \end{align} By the optimality conditions of $\widehat{w}^{(i,d)}$ and (ref), we have \begin{align} \| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widehat{w}^{(i,d)} - Y_{\text{pre}, i} \|_2^2 &\leq \| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widetilde{w}^{(i,d)} - Y_{\text{pre}, i} \|_2^2 \\ &= \|\boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widetilde{w}^{(i,d)}- \mathbb{E}[Y_{\text{pre}, i}] - \varepsilon_{\text{pre}, i} \|_2^2 \\ &= \|\boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widetilde{w}^{(i,d)}- \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}] \widetilde{w}^{(i,d)} - \varepsilon_{\text{pre}, i} \|_2^2 \\ &= \|( \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}])\widetilde{w}^{(i,d)} \|_2^2 + \| \varepsilon_{\text{pre}, i} \|_2^2 \\ &\quad - 2 \langle \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widetilde{w}^{(i,d)}- \mathbb{E}[Y_{\text{pre}, i}], \varepsilon_{\text{pre}, i}\rangle. \end{align} From (ref) and (ref), we have \begin{align} \| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widehat{w}^{(i,d)} - \mathbb{E}[Y_{\text{pre}, i}] \|_2^2 &\leq \| ( \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}])\widetilde{w}^{(i,d)} \|_2^2 + 2 \langle \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \Delta^{(i,d)}, \varepsilon_{\text{pre}, i} \rangle \\ &\leq \| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}] \|_{2,\infty}^2 \cdot \|\widetilde{w}^{(i,d)}\|_1^2 \\ &\quad + 2 \langle \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \Delta^{(i,d)}, \varepsilon_{\text{pre}, i} \rangle, \end{align} where we used (ref) in the second inequality above. Using (ref) and (ref), we obtain \begin{align} \| \mathcal{P}_{\widehat{V}_{\text{pre}}} \Delta^{(i,d)} \|_2^2 & \leq \frac{4}{\widehat{s}_{r_\text{pre}}^2} \left( \| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}] \|_{2,\infty}^2 \cdot \|\widetilde{w}^{(i,d)}\|_1^2 +\langle \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \Delta^{(i,d)}, \varepsilon_{\text{pre}, i} \rangle\right). \end{align} We now state two helper lemmas that will help us conclude the proof. The proof of Lemmas (ref) and (ref) are given in Appendices (ref) and (ref), respectively. \begin{lemma} Let Assumptions (ref), (ref), (ref), (ref), (ref), (ref) hold. Suppose $k = r_\emph{pre}$, where $k$ is defined as in (ref). Then conditioned on $\mathcal{E}$, we have $$ \| \boldsymbol{Y}^{r_{\emph{pre}}}_{\emph{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[\boldsymbol{Y}_{\emph{pre}, \mathcal{I}^{(d)}}] \|_{2,\infty} = O_p \left( \frac{\sqrt{r_\emph{pre} T_0 \log(T_0 N_d)}}{\min\{\sqrt{T_0}, \sqrt{N_d}\}} \right). $$ \end{lemma} \begin{lemma} Let Assumptions (ref) to (ref) hold. Then, given $\boldsymbol{Y}^{r_{\emph{pre}}}_{\emph{pre}, \mathcal{I}^{(d)}}$ and conditioned on $\mathcal{E}$, the following holds with respect to the randomness in $\varepsilon_{\emph{pre}, i}$: $$ \langle \boldsymbol{Y}^{r_{\emph{pre}}}_{\emph{pre}, \mathcal{I}^{(d)}} \Delta^{(i,d)}, \varepsilon_{\emph{pre}, i} \rangle = O_p \left( r_\emph{pre} + \sqrt{T_0} + \| \boldsymbol{Y}^{r_{\emph{pre}}}_{\emph{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[\boldsymbol{Y}_{\emph{pre}, \mathcal{I}^{(d)}}] \|_{2,\infty} \cdot \| \widetilde{w}^{(i,d)}\|_1 \right). $$ \end{lemma} Incorporating Lemmas (ref), (ref), (ref), (ref), and Assumption (ref) into (ref), we conclude \begin{align} &\mathcal{P}_{\widehat{V}_{\text{pre}}} \Delta^{(i,d)} \\ &= O_p \left(\frac{r_\text{pre} \| \widetilde{w}^{(i,d)}\|_1 \sqrt{\log(T_0 N_d)}}{\sqrt{N_d} \cdot \min\{\sqrt{T_0}, \sqrt{N_d} \}} + \frac{r_\text{pre}}{\sqrt{N_d T_0}} + \frac{\sqrt{r_\text{pre}}}{\sqrt{N_d} T_0^{\frac 1 4}} + \frac{r^{3/4}_\text{pre} \| \widetilde{w}^{(i,d)}\|^{1/2}_1 \log^{1/4}(T_0 N_d)}{\sqrt{N_d} T^{1/4}_0 \min\{T^{1/4}_0, N^{1/4}_d \}}\right) \\ &= O_p \left( \frac{\sqrt{r_\text{pre}}}{\sqrt{N_d} T_0^{\frac 1 4}} + \frac{r^{3/2}_\text{pre} \sqrt{\log(T_0 N_d)}}{\sqrt{N_d} \cdot \min\{\sqrt{T_0}, \sqrt{N_d} \}} \right). \end{align} {\em Collecting terms. } Combining (ref) and (ref), we conclude \begin{align} &\mathcal{P}_{V_\text{pre}} \Delta^{(i,d)} \&= O_p \left( \frac{\sqrt{r_\text{pre}}}{\sqrt{N_d} T_0^{\frac 1 4}} + \frac{r^{3/2}_\text{pre} \sqrt{\log(T_0 N_d)}}{\sqrt{N_d} \cdot \min\{\sqrt{T_0}, \sqrt{N_d} \}} + \frac{\sqrt{r_\text{pre} \log(T_0 N_d)}}{\min\{\sqrt{T_0}, \sqrt{N_d}\}} \left[ \frac{r^{3/4}_\text{pre} }{T^{1/4}_0 N_d^{1/2}} + \frac{r^{3/2}_\text{pre} }{\min\{T_0,N_d\}} \right] \right) \&= O_p \left( \frac{\sqrt{r_\text{pre}}}{\sqrt{N_d} T_0^{\frac 1 4}} + \frac{r^{3/2}_\text{pre} \sqrt{\log(T_0 N_d)}}{\sqrt{N_d} \cdot \min\{\sqrt{T_0}, \sqrt{N_d} \}} + \frac{r^{2}_\text{pre} \sqrt{\log(T_0 N_d)}}{\min\{T_0^{3/2},N_d^{3/2}\}} \right). \end{align} \end{proof} \subsection{Proof of Lemma (ref)} We recall the well-known singular subspace perturbation result by Wedin1972PerturbationBI. \begin{lemma} [Wedin's Theorem Wedin1972PerturbationBI] Given $\boldsymbol{A}, \boldsymbol{B} \in \mathbb{R}^{m \times n}$, let $\boldsymbol{V}, \widehat{\boldsymbol{V}} \in \mathbb{R}^{n \times n}$ denote their respective right singular vectors. Further, let $\boldsymbol{V}_{k} \in \mathbb{R}^{n \times k}$ (respectively, $\widehat{\boldsymbol{V}}_{k} \in \mathbb{R}^{n \times k}$) correspond to the truncation of $\boldsymbol{V}$ (respectively, $\widehat{\boldsymbol{V}}$), respectively, that retains the columns corresponding to the top $k$ singular values of $\boldsymbol{A}$ (respectively, $\boldsymbol{B}$). Let $s_{i}$ represent the $i$-th singular value of $\boldsymbol{A}$. Then, $$ \| \mathcal{P}_{V_{k}} - \mathcal{P}_{\widehat{V}_{k}} \|_\emph{op} \leq \frac{2\|\boldsymbol{A} - \boldsymbol{B}\|_\emph{op}} { s_{k} - s_{k+1}}. $$ \end{lemma} \begin{proof} Recall that $\widehat{\boldsymbol{V}}_\text{pre}$ is formed by the top $r_\text{pre}$ right singular vectors of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$ and $\boldsymbol{V}_\text{pre}$ is formed by the top $r_\text{pre}$ singular vectors of $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}]$. Therefore, Lemma (ref) gives $$ \| \mathcal{P}_{\widehat{V}_\text{pre}} - \mathcal{P}_{V_\text{pre}} \|_\text{op} \le \frac{2 \| \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}] \|_\text{op}} {s_{r_\text{pre}}}, $$ where we used $\rank(\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}]) = r_\text{pre}$ and hence $s_{r_{\text{pre}}+1}=0$. By Assumption (ref) and Lemma (ref), we can further bound the inequality above. In particular, for any $\zeta >0$, we have w.p. at least $1-2\exp(-\zeta^2)$, $$ \| \mathcal{P}_{\widehat{V}_\text{pre}} - \mathcal{P}_{V_\text{pre}} \|_\text{op} \le \frac{C \sigma (\sqrt{T_0} + \sqrt{N_d} + \zeta)}{s_{r_\text{pre}}}, $$ where $C>0$ is an absolute constant. \end{proof} \subsection{Proof of Lemma (ref)} We first re-state Lemma 7.2 of pcr_aos. \begin{lemma} [Lemma 7.2 of pcr_aos] Let the setup of Lemma (ref) hold. Then w.p. at least $1 - O(1 / (T_0 N_d)^{10})$, \begin{align} &\|\mathbb{E}[\boldsymbol{Y}_{\emph{pre}, \mathcal{I}^{(d)}}] - \boldsymbol{Y}^{r_{\emph{pre}}}_{\emph{pre}, \mathcal{I}^{(d)}} \|^2_{2, \infty} \& \quad \le C(\sigma) \left\{\frac{(T_0 + N_d) \cdot (T_0 + \sqrt{T_0}\log(T_0 N_d))}{s_{r_\emph{pre}}^2} + r_\emph{pre} + \sqrt{r_\emph{pre} }\log(T_0 N_d) \right\} + \frac{\log(T_0 N_d)}{N_d}, \qquad \end{align} where $C(\sigma)$ is a constant that only depends on $\sigma$. \end{lemma} \begin{proof} We adapt the notation in pcr_aos to that used in this paper. In particular, $\boldsymbol{X} = \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}]$, $\widetilde{\boldsymbol{Z}}^{r} = \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} $, where $\boldsymbol{X}, \widetilde{\boldsymbol{Z}}^{r}$ are the notations used in pcr_aos with $r = r_\text{pre}$. Next, we simplify (ref) using Assumption (ref). W.p. at least $1 - O(1 / (T_0 N_d)^{10})$, we have \begin{align} &\|\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}] - \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \|^2_{2, \infty} \le C(\sigma) \left(\frac{r_\text{pre} T_0\log(T_0 N_d)}{\min\{T_0, N_d\}}\right). \end{align} This concludes the proof of Lemma (ref). \end{proof} \subsection{Proof of Lemma (ref)} Throughout this proof, $C, c >0$ will denote absolute constants, which can change from line to line or even within a line. \begin{proof} Recall $\widehat{w}^{(i,d)} = \widehat{\boldsymbol{V}}_\text{pre} \widehat{\boldsymbol{S}}^{-1}_\text{pre} \widehat{\boldsymbol{U}}^\top_\text{pre} Y_{\text{pre}, i}$ and $Y_{\text{pre}, i} = \mathbb{E}[Y_{\text{pre}, i}] + \varepsilon_{\text{pre}, i}$. Thus, \begin{align} \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widehat{w}^{(i,d)} = \widehat{\boldsymbol{U}}_\text{pre}\widehat{\boldsymbol{S}}_\text{pre} \widehat{\boldsymbol{V}}_\text{pre}^\top \widehat{\boldsymbol{V}}_\text{pre} \widehat{\boldsymbol{S}}_\text{pre}^{-1} \widehat{\boldsymbol{U}}_\text{pre}^\top Y_{\text{pre}, i} = \mathcal{P}_{\widehat{U}_\text{pre}} \mathbb{E}[Y_{\text{pre}, i}] + \mathcal{P}_{\widehat{U}_\text{pre}} \varepsilon_{\text{pre}, i}. \end{align} We then obtain \begin{align} \langle \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} (\widehat{w}^{(i,d)} - \widetilde{w}^{(i,d)}), \varepsilon_{\text{pre}, i} \rangle &= \langle \mathcal{P}_{\widehat{U}_\text{pre}} \mathbb{E}[Y_{\text{pre}, i}], \varepsilon_{\text{pre}, i} \rangle + \langle \mathcal{P}_{\widehat{U}_\text{pre}} \varepsilon_{\text{pre}, i}, \varepsilon_{\text{pre}, i} \rangle \\ &\quad - \langle \widehat{\boldsymbol{U}}_\text{pre}\widehat{\boldsymbol{S}}_\text{pre} \widehat{\boldsymbol{V}}_\text{pre}^\top \widetilde{w}^{(i,d)}, \varepsilon_{\text{pre}, i} \rangle. \end{align} Note that $\varepsilon_{\text{pre}, i}$ is independent of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$, and thus also independent of $\widehat{\boldsymbol{U}}_\text{pre}, \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}}$. Therefore, \begin{align} \mathbb{E}\Big[\langle \mathcal{P}_{\widehat{U}_\text{pre}} \mathbb{E}[Y_{\text{pre}, i}], \varepsilon_{\text{pre}, i} \Big] &= 0, \\ \mathbb{E}\Big[\langle \widehat{\boldsymbol{U}}_\text{pre}\widehat{\boldsymbol{S}}_\text{pre} \widehat{\boldsymbol{V}}_\text{pre}^\top \widetilde{w}^{(i,d)}, \varepsilon_{\text{pre}, i}\rangle \Big] &= 0 \end{align} Moreover, using the cyclic property of the trace operator, we obtain \begin{align} \mathbb{E}\big[\langle \mathcal{P}_{\widehat{U}_\text{pre}} \varepsilon_{\text{pre}, i}, \varepsilon_{\text{pre}, i} \rangle\big] &= \mathbb{E}\big[ \varepsilon^\top_{\text{pre}, i} \mathcal{P}_{\widehat{U}_\text{pre}} \varepsilon_{\text{pre}, i} \big] = \mathbb{E} \big[\tr(\varepsilon^\top_{\text{pre}, i} \mathcal{P}_{\widehat{U}_\text{pre}} \varepsilon_{\text{pre}, i})\big] \\ &= \tr(\mathbb{E}\big[ \varepsilon_{\text{pre}, i} \varepsilon^\top_{\text{pre}, i}\big] \mathcal{P}_{\widehat{U}_\text{pre}}) = \tr(\sigma^2 \mathcal{P}_{\widehat{U}_\text{pre}}) \\ &= \sigma^2 \cdot r_\text{pre}. \end{align} Accordingly, we have \begin{align} \mathbb{E} \left[\langle \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \Delta^{(i,d)}, \varepsilon_{\text{pre}, i} \rangle \right] & = \sigma^2 \cdot r_\text{pre}. \end{align} Next, to obtain high probability bounds for the inner product term, we use the following lemmas, the proofs of which can be found in pcr_aos. \begin{lemma} [Modified Hoeffding's Lemma] Let $X \in \mathbb{R}^n$ be r.v. with independent mean-zero sub-Gaussian random coordinates with $\| X_i \|_{\psi_2} \le K$. Let $a \in \mathbb{R}^n$ be another random vector that satisfies $\|a\|_2 \le b$ for some constant $b \ge 0$. Then for any $\zeta \ge 0$, $ \mathbb{P} \Big( \Big| \sum_{i=1}^n a_i X_i\Big| \ge \zeta \Big) \le 2 \exp\Big(-\frac{c\zeta^2}{K^2 b^2} \Big), $ where $c > 0$ is a universal constant. \end{lemma} \begin{lemma} [Modified Hanson-Wright Inequality] Let $X \in \mathbb{R}^n$ be a r.v. with independent mean-zero sub-Gaussian coordinates with $\|X_i\|_{\psi_2} \le K$. Let $\boldsymbol{A} \in \mathbb{R}^{n \times n}$ be a random matrix satisfying $\|\boldsymbol{A}\|_{\emph{op}} \le a$ and $\|\boldsymbol{A}\|_F^2 \, \le b$ for some $a, b \ge 0$. Then for any $\zeta \ge 0$, \begin{align*} \mathbb{P} \left( \abs{ X^\top \boldsymbol{A} X - \mathbb{E}[X^\top \boldsymbol{A} X] } \ge \zeta \right) &\le 2 \exp \left\{ -c \min\left(\frac{\zeta^2}{K^4 b}, \frac{\zeta}{K^2 a} \right) \right\}. \end{align*} \end{lemma} Using Lemma (ref) and (ref), and Assumptions (ref) and (ref), it follows that for any $\zeta > 0$ \begin{align} \mathbb{P}\left( \langle \mathcal{P}_{\widehat{U}_\text{pre}} \mathbb{E}[Y_{\text{pre}, i}], \varepsilon_{\text{pre}, i} \rangle \geq \zeta \right) & \leq \exp\left( - \frac{c \zeta^2}{T_0 \sigma^2 } \right). \end{align} Note that the above also uses $ \| \mathcal{P}_{\widehat{U}_\text{pre}} \mathbb{E}[Y_{\text{pre}, i}] \|_2 \le \| \mathbb{E}[Y_{\text{pre}, i}] \|_2 \le \sqrt{T_0}, $ which follows from the fact that $\|\mathcal{P}_{\widehat{U}_\text{pre}}\|_\text{op} \le 1$ and Assumption (ref). Further, (ref) implies \begin{align} \langle \mathcal{P}_{\widehat{U}_\text{pre}} \mathbb{E}[Y_{\text{pre}, i}], \varepsilon_{\text{pre}, i} \rangle = O_p \left( \sqrt{T_0} \right). \end{align} Similarly, using (ref), we have for any $\zeta > 0$ \begin{align} &\mathbb{P}\left( \langle \widehat{\boldsymbol{U}}_\text{pre}\widehat{\boldsymbol{S}}_\text{pre} \widehat{\boldsymbol{V}}_\text{pre}^\top\widetilde{w}^{(i,d)}, \varepsilon_{\text{pre}, i} \rangle \geq \zeta \right) \\ &\quad \leq \exp\left\{ - \frac{c \zeta^2}{\sigma^2 (T_0 +\| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[ \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}]\|^2_{2,\infty} \cdot \|\widetilde{w}^{(i,d)}\|_1^2) } \right\}, \end{align} where we use the fact that \begin{align} \| \widehat{\boldsymbol{U}}_\text{pre}\widehat{\boldsymbol{S}}_\text{pre} \widehat{\boldsymbol{V}}_\text{pre}^\top\widetilde{w}^{(i,d)}\|_2 &= \| \boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} \widetilde{w}^{(i,d)} \pm \mathbb{E}[Y_{\text{pre}, i}] \|_2 \\ & = \| (\boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}])\widetilde{w}^{(i,d)} + \mathbb{E}[Y_{\text{pre}, i}]\|_2 \\ &\leq \| (\boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}]) \widetilde{w}^{(i,d)}\|_2 + \|\mathbb{E}[Y_{\text{pre}, i}] \|_2 \nonumber \\ & \leq \|\boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}]\|_{2,\infty} \cdot \| \widetilde{w}^{(i,d)}\|_1 + \sqrt{T_0}. \end{align} In the inequalities above, we use $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}] \cdot \widetilde{w}^{(i,d)} = \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}] \cdot w^{(i,d)} = \mathbb{E}[Y_{\text{pre}, i}]$, which follows from (ref) and (ref). Then, (ref) implies \begin{align} \langle \widehat{\boldsymbol{U}}_\text{pre}\widehat{\boldsymbol{S}}_\text{pre} \widehat{\boldsymbol{V}}_\text{pre}^\top\widetilde{w}^{(i,d)}, \varepsilon_{\text{pre}, i} \rangle = O_p\left( \|\boldsymbol{Y}^{r_{\text{pre}}}_{\text{pre}, \mathcal{I}^{(d)}} - \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}]\|_{2,\infty} \cdot \| \widetilde{w}^{(i,d)}\|_1 + \sqrt{T_0} \right) \end{align} Finally, using Lemma (ref), (ref), Assumptions (ref) and (ref), we have for any $\zeta > 0$ \begin{align} \mathbb{P}\left( \langle \mathcal{P}_{\widehat{U}_\text{pre}} \varepsilon_{\text{pre}, i}, \varepsilon_{\text{pre}, i} \rangle \geq \sigma^2 r_\text{pre} + \zeta \right) & \leq \exp\left\{ - c \min\left(\frac{\zeta^2}{\sigma^4 r_\text{pre}}, \frac{\zeta}{\sigma^2}\right)\right\}, \end{align} where we have used $\|\mathcal{P}_{\widehat{U}_\text{pre}}\|_\text{op} \le 1$ and $\|\mathcal{P}_{\widehat{U}_\text{pre}}\|^2_F = r_\text{pre}$. Then, (ref) implies \begin{align} \langle \mathcal{P}_{\widehat{U}_\text{pre}} \varepsilon_{\text{pre}, i}, \varepsilon_{\text{pre}, i}\rangle = O_p\left( r_\text{pre} \right). \end{align} From (ref), (ref), (ref), and (ref), we arrive at our desired result. \end{proof} \section{Smooth Non-linear Latent Variable Models} We expound on Remark (ref). Specifically, we show that sufficiently continuous non-linear latent variable models (LVMs) are arbitrarily well-approximated by linear factor models of appropriate dimension as $N, T$ grow. In this sense, the factor model in Assumption (ref) serves as a good approximation for non-linear LVMs when we have many units and time periods. \subsection{Definitions} We first precisely define the notion of sufficiently continuous. \begin{definition}[H\"older continuity] {\em The H\"older class $\mathcal{H}(q,S,C_H)$ on $[0, 1)^q$ is the set of functions $h: [0, 1)^q \to \mathbb{R}$ whose partial derivatives satisfy $$ \sum_{s: |s| = \lfloor S \rfloor} \frac{1}{s!} |\nabla_s h(\mu) - \nabla_s h(\mu') | \le C_H \norm{\mu - \mu'}_{\max}^{S - \lfloor S \rfloor},\quad \forall \mu, \mu' \in [0, 1)^q, $$ where $\lfloor S \rfloor$ denotes the largest integer strictly smaller than $S$. Here, $s = (s_1, \dots, s_q)$ is a multi-index with $s_\ell \in \mathbb{N}$, $|s| = \sum^q_{\ell = 1} s_\ell$, and $\nabla_s h(\cdot) = \frac{\partial^{s_1} \dots \partial^{s_q}}{\partial \mu_{1} \dots \partial \mu_{q}} h(\cdot)$. } \end{definition} Next, we define “smooth” non-linear LVMs as follows. \begin{definition}[“Smooth” non-linear LVMs] {\em The potential outcomes are said to satisfy a smooth non-linear factor model if (i) $\mathbb{E}[Y_{ti}^{(d)}] = f(\rho^{(d)}_{t}, \phi_i)$, where $f: [0, 1)^{2q} \to \mathbb{R}$, and $\phi_i, \rho^{(d)}_{t} \in [0, 1)^q$; (ii) $f(\rho^{(d)}_{t}, \cdot) \in \mathcal{H}(q,S,C_H)$ for all $t, d \in [T] \times [D]_0$. } \end{definition} Intuitively, Definition (ref) states that for a pair of units $(i_1, i_2)$, we have $\mathbb{E}[Y_{ti_1}^{(d)}] \approx \mathbb{E}[Y_{ti_2}^{(d)}]$ for all $(t,d)$ if $\phi_{i1} \approx \phi_{i2}$. The dimension $q$ of the latent variables $(\phi_i, \rho^{(d)}_{t})$ encodes the complexity of representing the various characteristics of units, time, and treatments that affect the potential outcomes. Note that a linear factor model is a special case of a non-linear LVM where $f(\rho^{(d)}_{t}, \phi_i) = \langle \rho^{(d)}_{t}, \phi_i \rangle$. Such a model satisfies Definition (ref) for all $S \in \mathbb{N}$ and $C_H = C$, for some absolute finite, positive constant $C$. \subsection{Formal Result} Proposition (ref) establishes that linear factor models of appropriate dimension also provide a good approximation for non-linear LVMs. \begin{proposition}[pcr_aos, xu2017rates] Suppose the potential outcomes satisfy Definition (ref) for some fixed $\mathcal{H}(q,S,C_H)$. Then for any $\delta>0$, there exists latent variables $u^{(d)}_t, v_i \in \mathbb{R}^r$ such that for all $(t,i,d)$, we have $| \mathbb{E}[Y_{ti}^{(d)}] - \sum^r_{\ell = 1} u^{(d)}_{t\ell} v_{i\ell} | \leq \Delta_E$, where $r \le C \cdot \delta^{-q}$ and $\Delta_E \le C_H \cdot \delta^S$; here, $C$ is allowed to depend on $(q,S)$. \end{proposition} By choosing $\delta = \min\{N, T\}^{-c / q}$ for $c \in (0, 1)$, Proposition (ref) implies that $\Delta_E \to 0$ as $N, T$ grow and $r \ll \min\{N, T\}$. From this, we identify that linear latent factor models of appropriate dimension serve as a good approximation for non-linear LVMs as $N, T$ grow. Further, we note that H\"older continuous functions are closed under composition. For example, if $\mathbb{E}[Y^{(d)}_{ti}]$ satisfies Definition (ref), then $\log(\mathbb{E}[Y^{(d)}_{ti}])$ also satisfies it as $\log(\cdot)$ is an analytic function. In contrast, a two-way fixed effects model, which is the canonical model for methods such as Difference-in-Differences, is of the form \begin{align} \mathbb{E}[Y^{(d)}_{ti}] = u^{(d)}_t + v_i. \end{align} This is not robust to the choice of parameterization, e.g., if (ref) holds for $\mathbb{E}[Y^{(d)}_{ti}]$ then it is not guaranteed to hold for $\log(\mathbb{E}[Y^{(d)}_{ti}])$. \section{Incorporating Covariates} In this section, we discuss how meaningful unit-level covariates can help improve estimation. Let $\boldsymbol{X} = [X_{ki}] \in \mathbb{R}^{K \times N}$ denote the matrix of covariates across units, i.e., $X_{ki}$ denotes the $k$-th descriptor (or feature) of unit $i$. \subsection{Assumptions} One approach towards incorporating covariates into the \textsf{SI} estimator described in Section (ref), is to impose the following structure on $\boldsymbol{X}$. \begin{assumption} [Covariate structure] For any $k \in [K]$ and $i \in [N]$, let $X_{ki} = \langle \varphi_k, v_i \rangle + \zeta_{ki}$. Here, $\varphi_k \in \mathbb{R}^r$ represents a latent factor specific to descriptor $k$, $v_i \in \mathbb{R}^r$ is the unit latent factor as defined in (ref), and $\zeta_{ki} \in \mathbb{R}$ is a mean zero measurement noise specific to descriptor $k$ and unit $i$. \end{assumption} The key structure of Assumption (ref) is that the covariates $X_{ki}$ share the same latent unit factors, $v_i$, as the potential outcomes $Y^{(d)}_{tn}$. Thus, given a target unit $i$ and subgroup $\mathcal{I}^{(d)}$, this allows us to express unit $i$'s covariates as a linear combination of the covariates associated with units within $\mathcal{I}^{(d)}$ via the {\em same} linear model that describes the relationship between their respective potential outcomes (formalized in Proposition (ref) below). One notable flexibility of our covariate model is that the observations of covariates can be {\em noisy} due to the presence of measurement noise $\zeta_{ki}$. \begin{proposition} Given $(i,d)$, let Assumptions (ref) and (ref) hold. Then, conditioned on $\mathcal{E}_\varphi = \mathcal{E} \cup \{\varphi_k: k \in [K]\}$, we have for all $k$, \begin{align} \mathbb{E}[X_{ki} \,| \, \mathcal{E}_\varphi] = \sum_{j \in \mathcal{I}^{(d)}} w_j^{(i,d)} \cdot \mathbb{E}[X_{kj} \,|\, \mathcal{E}_\varphi], \end{align} where $w^{(i,d)}$ defined as in Assumption (ref). \end{proposition} \begin{proof} Plugging Assumption (ref) into Assumption (ref) completes the proof. \end{proof} \subsection{Description of Estimator with Covariates} Proposition (ref) suggests a natural modification to step (i) of the \textsf{SI} estimator. Similar to the estimation procedure of abadie2, we propose to append the covariates to the pre-treatment outcomes. Formally, let $X_i = [X_{ki}] \in \mathbb{R}^{K}$ denote the vector of covariates associated with the unit $i$; analogously, let $\boldsymbol{X}_{\mathcal{I}^{(d)}} = [X_{kj}: j \in \mathcal{I}^{(d)}] \in \mathbb{R}^{K \times N_d}$ denote the matrix of covariates for the units in $\mathcal{I}^{(d)}$. Define $Z_i = [Y_{\text{pre}, i}^\top, X_i^\top]^\top \in \mathbb{R}^{T_0 + K}$ and $\boldsymbol{Z}_{\mathcal{I}^{(d)}} = [\boldsymbol{Y}^\top_{\text{pre}, \mathcal{I}^{(d)}}, \boldsymbol{X}^\top_{\mathcal{I}^{(d)}}]^\top \in \mathbb{R}^{(T_0 + K) \times N_d}$ as the concatenation of pre-treatment outcomes and covariates for unit $i$ and subgroup $\mathcal{I}^{(d)}$, respectively. The \textsf{SI} estimator then proceeds by learning $\widehat{w}^{(i,d)}$ from $(Z_i, \boldsymbol{Z}_{\mathcal{I}^{(d)}})$ and continues with step (ii) as described in Section (ref). \subsection{Theoretical Implications} Consider the \textsf{SI}-\texttt{PCR} estimator with covariates. Then, it can be verified that Theorem (ref) continues to hold with $T_0$ replaced with $T_0 + K$. Therefore, adding covariates into the model-learning stage augments the data size, which can improve the estimation rates. \section{Causal Inference and Tensor Completion} We elaborate on our connection between causal inference and tensor completion. Recall Assumption (ref), which connects our observed tensor, $\boldsymbol{Y}$, with the underlying potential outcomes tensor, $\boldsymbol{Y}^*$. As elucidated in Section (ref), the fundamental task in causal inference of estimating unobserved potential outcomes can be achieved through tensor completion. Thus, as suggested by Figure (ref), different observational and experimental studies can be equivalently posed as tensor completion under different sparsity patterns. \subsection{Tensor Factor Models and Algorithm Design} Recall that the factorization given by (ref) in Assumption (ref) is implied by the factorization assumed by a low-rank tensor given by (ref). In particular, Assumption (ref) does not require the additional factorization of the time-treatment factor $u^{(d)}_t$ as $ u_t \circ \lambda_d$, where $\circ$ denotes the entry-wise Hadamard product. Hence, we ask whether it is feasible to design estimators that exploit an implicit factorization of $u^{(d)}_t = u_t \circ \lambda_d $. Indeed, a recent follow-up work of squires2021causal finds that a variant of \textsf{SI} that regresses across treatments as opposed to units leads to better empirical results in the setting of single cell therapies. At the same time, christina_tensor directly exploits the tensor factor structure in (ref) to provide max-norm error bounds for tensor completion under a missing completely at random (MCAR) setting. We leave as an open question whether one can exploit the latent factor structure in (ref) over that in Assumption (ref) to operate under less stringent causal assumptions or data requirements and/or identify and estimate different causal estimands. \subsection{A New Tensor Completion Toolkit for Causal Inference} The tensor completion problem provides an expressive formal framework for a large number of emerging applications. In particular, this literature quantifies the number of samples required and the computational complexity of different estimators to achieve theoretical guarantees for a given error metric (e.g., Frobenius-norm error over the entire tensor). Indeed, this trade-off is of central importance in the emerging sub-discipline at the interface of computer science, statistics, and machine learning. Given our preceding discussions connecting causal inference with tensor completion, a natural question to ask is whether we can apply the current toolkit of tensor completion methods to understand statistical and computational trade-offs in causal inference. We believe a direct transfer of tensor completion techniques and analyses to causal inference problems is not immediately possible in general for a couple of reasons. First, most tensor completion results assume MCAR entries. As seen in Figure (ref), causal inference settings often induce missing not at random (MNAR) sparsity patterns. Second, typical tensor completion results hold for the Frobenius-norm error across the entire tensor. By contrast, a variety of meaningful causal estimands require more refined error metrics over subsets of the tensor. Hence, we pose two important and related open questions: (i) what block sparsity patterns and structural assumptions on the potential outcomes tensor allow for faithful recovery with respect to a meaningful error metric for causal inference, and (ii) if recovery is possible, what are the fundamental statistical and computational trade-offs that are achievable? An answer to these questions can aid in bridging causal inference with tensor completion, as well as computer science and machine learning more broadly. \section{Inference} \subsection{Asymptotic Normality} We show how asymptotic normality can be established under our operating assumptions. \begin{theorem} Given $(i,d)$, let the setup of Theorem (ref) hold. Let $ \widetilde{w}^{(i,d)}$ be defined as in (ref) and suppose we condition on $\mathcal{E}$. Let the following conditions hold: (i) $T_0, N_d, T_1 \to \infty$, (ii) \begin{align} \frac{ r^{3/2}_\emph{pre} \sqrt{\log(T_0 N_d)}}{\| \widetilde{w}^{(i,d)}\|_2 \cdot \min\{T_0,N_d, T^{1/4}_0 N_d^{1/2}\} } = o(1), \end{align} and (iii) \begin{align} \frac{1}{\sqrt{T_1} \|\widetilde{w}^{(i,d)}\|_2} \sum_{t > T_0} \left \langle \mathbb{E}[Y_{t, \mathcal{I}^{(d)}}], \boldsymbol{V}_\emph{pre} \boldsymbol{V}_\emph{pre}^\top \left(\widehat{w}^{(i,d)} - \widetilde{w}^{(i,d)} \right) \right \rangle &= o_p(1), \end{align} where $\widehat{w}^{(i,d)}$ is given by (ref). Then, we have \begin{align} \frac{\sqrt{T_1}}{\sigma \|\widetilde{w}^{(i,d)}\|_2} \Big(\widehat{\theta}_i^{(d)} - \theta_i^{(d)}\Big) & \xrightarrow{d} \mathcal{N} \left(0, 1\right). \end{align} Here, $O_p(\cdot)$ is defined with respect to the sequence $\min\{N_d, T_0\}$. \end{theorem} A few remarks are in order. First, as opposed to Theorem (ref) that establishes consistency under a fixed $T_1$, Theorem (ref) proves asymptotic normality with a growing $T_1$. Second, we note that (ref) can be equivalently stated as $$ \| \widetilde{w}^{(i,d)}\|_2 \gg \left((r^{3/2}_\text{pre} \sqrt{\log(T_0 N_d)}) / \min\{T_0,N_d, T^{1/4}_0 N_d^{1/2}\} \right),$$ i.e., we require the $\ell_2$-norm of $\widetilde{w}^{(i,d)}$ to be sufficiently large. Next, for (ref) to hold, we require that the estimation error between $\widehat{w}^{(i,d)}$ and $\widetilde{w}^{(i,d)}$, projected onto the rowspace of $\mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}]$, to be vanishing sufficiently quickly. Finally, by (ref), we see that any formulation of the \textsf{SI} estimator in which (ref) and (ref) simultaneously hold leads to an estimate that is asymptotically normal around the causal estimand. In light of this, we describe one such formulation. \subsection{Modified \textsf{SI}-\texttt{PCR} Estimator} We present a slight modification to the \textsf{SI}-\texttt{PCR} estimator and show that this strategy obeys our desired condition in (ref). \subsubsection{Description of Estimator} Recall the notation set in Section (ref). At a high-level, the modified estimator proceeds in the exact same manner as described in Section (ref) with one change to (ref) of step (i). Let $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^k = \sum_{\ell=1}^k \widehat{s}_\ell \widehat{u}_\ell \widehat{v}_\ell^\top$ denote the rank-$k$-approximation of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$. We then construct a sub-sampled matrix that retains the columns of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^k$ within an index set $\Omega \subset \mathcal{I}^{(d)}$. We denote the resulting matrix as $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{(k,\Omega)} \in \mathbb{R}^{T_0 \times |\Omega|}$. Next, we replace (ref) with \begin{align} \widehat{w}^{(i,d,\Omega)} = \left( \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{(k, \Omega)} \right)^\dagger Y_{\text{pre}, i} \in \mathbb{R}^{|\Omega|}. \end{align} The estimation strategy continues as per step (ii) of Section (ref) with $\widehat{w}^{(i,d,\Omega)}$ and $\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}}^\Omega \in \mathbb{R}^{T_1 \times |\Omega|}$, which is defined as the sub-matrix of $\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}}$ that retains the columns within $\Omega$. \subsubsection{Theoretical Implications} Let $\boldsymbol{V}_{\mathcal{I}^{(d)}} = [v_j: j \in \mathcal{I}^{(d)}] \in \mathbb{R}^{N_d \times r}$. We define $\boldsymbol{V}_{\mathcal{I}^{(d)}}^\Omega \in \mathbb{R}^{|\Omega| \times r}$ as the sub-sampled matrix that retains the rows of $\boldsymbol{V}_{\mathcal{I}^{(d)}}$ within $\Omega$. To adapt the results in Section (ref) to our new setting, note that if $\Omega$ is chosen such that $\text{span}(\{v_j: j \in \Omega\})$ is an $r_\text{pre}$-dimensional subspace of $\mathbb{R}^{N_d}$ (i.e., $\rank(\boldsymbol{V}_{\mathcal{I}^{(d)}}^\Omega) = r_\text{pre}$), then there exists a $w^{(i,d,\Omega)} \in \mathbb{R}^{|\Omega|}$ such that \begin{align} v_i = \boldsymbol{V}_{\mathcal{I}^{(d)}}^\top \cdot w_j^{(i,d)} = \left(\boldsymbol{V}_{\mathcal{I}^{(d)}}^\Omega \right)^\top \cdot w_j^{(i,d,\Omega)}, \end{align} where $w^{(i,d)} \in \mathbb{R}^{N_d}$ is defined as in Assumption (ref) (and Theorem (ref)). Intuitively, our modified strategy anchors on the concept that only $r_\text{pre}$ donors are needed to recreate our target unit $i$ since $\boldsymbol{V}_{\mathcal{I}^{(d)}}$ has rank $r_\text{pre}$; hence, it suffices to use a subset of donors defined over $\Omega$ rather than utilizing all $N_d$ donor units, provided $\Omega$ is chosen such that (ref) holds. With this intuition, we are now ready to present the main result of this section. Below, let $\boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}} = \mathbb{E}[\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} | \mathcal{E}]$ and define $\boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}}^{\Omega} \in \mathbb{R}^{T_0 \times |\Omega|}$ as the sub-sampled matrix that retains the columns of $\boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}}$ within $\Omega$. \begin{lemma} Given $(i,d)$, let Assumptions (ref) to (ref) hold. Consider the modified \textsf{SI}-\texttt{PCR} estimator with $\widehat{w}^{(i,d,\Omega)}$ and $w^{(i,d,\Omega)}$ as defined as in (ref) and (ref), respectively. Suppose $k$ and $\Omega$ are chosen such that for appropriate constant $C$, we have \begin{itemize} • $k = r_\emph{pre} = \rank(\boldsymbol{M}_{\emph{pre}, \mathcal{I}^{(d)}})$, • $|\Omega| = r_\emph{pre}$ with $\rank(\boldsymbol{Y}_{\emph{pre}, \mathcal{I}^{(d)}}^{(r_\emph{pre}, \Omega)}) = r_\emph{pre}$ and $\rank(\boldsymbol{M}_{\emph{pre}, \mathcal{I}^{(d)}}^\Omega) = r_\emph{pre}$, • Assumption (ref) holds with respect to $\boldsymbol{M}_{\emph{pre}, \mathcal{I}^{(d)}}^\Omega$ and $s^\Omega_{r_\emph{pre}} \ge C \sigma \sqrt{T_0}$, • $N_d < T_0$, • $\| w^{(i,d,\Omega)}) \|_2 \ge C r^{-1/2}_\emph{pre}$, • $r^2_\emph{pre} = o\left( \frac{\min\{T_0,N_d, T^{1/4}_0 N_d^{1/2}\}}{\sqrt{\log(T_0 N_d)}} \right)$, • $T_1 = o\left( \min\left\{ \frac{N_d}{\log(T_0 N_d) r^4_\emph{pre}}, \frac{T^{1/2}_0}{r^2_\emph{pre}} \right\} \right)$. \end{itemize} Then conditioned on $\mathcal{E}$ and $\Omega$, we have \begin{align} \frac{ r^{3/2}_\emph{pre} \sqrt{\log(T_0 N_d)}}{\| w^{(i,d,\Omega)}\|_2 \cdot \min\{T_0,N_d, T^{1/4}_0 N_d^{1/2}\} } = o(1), \end{align} and \begin{align} \frac{1}{\sqrt{T_1} \|w^{(i,d,\Omega)}\|_2} \sum_{t > T_0} \left \langle \mathbb{E}[Y^\Omega_{t, \mathcal{I}^{(d)}}], \left( \boldsymbol{M}^\Omega_{\emph{pre}, \mathcal{I}^{(d)}} \right)^\dagger \boldsymbol{M}^\Omega_{\emph{pre}, \mathcal{I}^{(d)}}\left(\widehat{w}^{(i,d,\Omega)} - w^{(i,d,\Omega)} \right) \right \rangle &= o_p(1), \end{align} where $\mathbb{E}[Y^\Omega_{t, \mathcal{I}^{(d)}}] = [\mathbb{E}[Y_{tjd}]: j \in \Omega] \in \mathbb{R}^{r_\emph{pre}}$ and $O_p(\cdot)$ is defined with respect to $N_d$. \end{lemma} Lemma (ref) states that the modified \textsf{SI}-\texttt{PCR} estimator satisfies conditions (ref) and (ref) of Theorem (ref) to achieve asymptotic normality. This enables us to conduct valid inference via the following confidence interval: for $\gamma \in (0,1)$, \begin{align} \mathcal{CI}(\gamma) = \left[ \widehat{\theta}_i^{(d, \Omega)} \pm \frac{z_{\gamma/2} \cdot \widehat{\sigma}^\Omega \cdot \|\widehat{w}^{(n, d, \Omega)}\|_2}{\sqrt{T_1}} \right], \end{align} where $z_{\gamma/2}$ is the upper $\gamma/2$ quantile of the standard Normal and \begin{align} \widehat{\sigma}^\Omega = \frac{1}{T_0} \| Y_{\text{pre}, i} - \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{(r, \Omega)} \cdot \widehat{w}^{(i,d,\Omega)} \|_2^2. \end{align} We study properties of the proposed confidence interval below. \subsubsection{Simulation Study} This section presents a simulation to study the coverage properties of the confidence interval in (ref) based on the modified \textsf{SI}-\texttt{PCR} estimator. \paragraph{Data generating process.} We vary $T_0 \in \{200, 400, 600, 800, 1000\}$ and choose $r = 5$. Motivated by our conditions in Lemma (ref), we choose $N_d = T_0/2$ and $T_1 = \sqrt{T_0}$ for each $T_0$. {\em Latent factors.} As in Section (ref), we generate the latent factors associated with the donor units, $\boldsymbol{V}_{\mathcal{I}^{(d)}} \in \mathbb{R}^{N_d \times r}$, by independently sampling its entries from a standard Normal. Then, we sample $w^{(i,d)} \in \mathbb{R}^N_d$ by drawing its entries from a uniform distribution over $[0,1]$ and normalizing it to have unit norm. We define the target unit latent factor as $v_i = \boldsymbol{V}_{\mathcal{I}^{(d)}}^\top \cdot w^{(i,d)}$. We define the latent factors for the pre-treatment outcomes under control as $\boldsymbol{U}^{(0)}_\text{pre}$, where its entries are i.i.d. samples from a standard Normal. Next, we sample $\boldsymbol{\Phi} \in \mathbb{R}^{T_1 \times r}$ whose entries are i.i.d. draws from a uniform distribution over $[0, 1]$. We construct the post-treatment latent factors of treatment $d$ as $\boldsymbol{U}_\text{post}^{(d)} = \boldsymbol{\Phi} \cdot (\boldsymbol{U}_\text{pre}^{(0)})^\dagger \boldsymbol{U}^{(0)}_\text{pre}$. We define our causal estimand as $\theta_i^{(d)} = \textsf{average}(\boldsymbol{U}_\text{post}^{(d)} \cdot v_i )$. {\em Observations.} Let (i) $Y_{\text{pre}, i} = \boldsymbol{U}^{(0)}_\text{pre} \cdot v_i + \varepsilon_{\text{pre}, i}$; (ii) $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} =\boldsymbol{U}^{(0)}_\text{pre} \cdot \boldsymbol{V}^\top_{\mathcal{I}^{(d)}} + \boldsymbol{E}_{\text{pre}, \mathcal{I}^{(d)}}$; and (iii) $\boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}} = \boldsymbol{U}^{(d)}_\text{post} \cdot \boldsymbol{V}^\top_{\mathcal{I}^{(d)}} + \boldsymbol{E}_{\text{post}, \mathcal{I}^{(d)}}$, where the entries of $(\varepsilon_{\text{pre}, i}, \boldsymbol{E}_{\text{pre}, \mathcal{I}^{(d)}}, \boldsymbol{E}_{\text{post}, \mathcal{I}^{(d)}} )$ are i.i.d. samples from a standard normal. \paragraph{Simulation results.} From $(Y_{\text{pre}, i}, \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}})$, we learn $\widehat{w}^{(i,d,\Omega)}$ via the modified \textsf{SI}-\texttt{PCR} estimator, where $\Omega$ is chosen to index the first $r$ donors of $\mathcal{I}^{(d)}$; hence, $| \Omega | = r$. We define our estimate as $\widehat{\theta}_i^{(d, \Omega)} = \textsf{average}( \boldsymbol{Y}_{\text{post}, \mathcal{I}^{(d)}}^{\Omega} \cdot \widehat{w}^{(i,d,\Omega)} )$ and construct its corresponding confidence interval as per (ref). Table (ref) reports on the empirical coverage probabilities and average interval lengths for $\theta_i^{(d)}$ at each $T_0$ over $50$ trials, where each trial consists of an independent draw of the latent factors and the observations of size $100$; this amounts to $5000$ total simulation repeats per dimension. As suggested by Lemma (ref), we observe that our confidence interval consistently achieves close to the nominal targets of $90\%$ and $95\%$ across all values of $T_0$, albeit the interval lengths can be large due to the poorer prediction quality compared to the standard \textsf{SI}-\texttt{PCR} estimator of (ref). \begin{table} [!t] \caption{Coverage results for the modified \textsf{SI}-\texttt{PCR} estimator across 5000 replications.} \begin{tabular}{@lcccc@} \toprule \midrule \multirow{2}{*}{Size of $\mathcal{T}_\text{pre}$} & \multicolumn{2}{c}{$90\%$ nominal coverage} & \multicolumn{2}{c}{$95\%$ nominal coverage} \\ \cmidrule(l){2-3} \cmidrule(l){4-5} & Coverage probability & Interval length & Coverage probability & Interval length \\ \midrule $T_0 = 200$ & 0.88 & 15.66 & 0.94 & 18.66 \\ \midrule $T_0 = 400$ & 0.88 & 18.90 & 0.94 & 22.52 \\ \midrule $T_0 = 600$ & 0.88 & 8.90 & 0.94 & 10.61 \\ \midrule $T_0 = 800$ & 0.89 & 7.25 & 0.94 & 8.64 \\ \midrule $T_0 = 1000$ & 0.90 & 15.58 & 0.95 & 18.57 \\ \bottomrule \end{tabular} \end{table} \subsection{Proof of Theorem (ref)} \begin{proof} For ease of notation, we suppress the conditioning on $\mathcal{E}$ for the remainder of the proof. We scale the left-hand side (LHS) of (ref) by $\sqrt{T_1}/ (\sigma \|\widetilde{w}^{(i,d)}\|_2)$ and analyze each of the three terms on the right-hand side (RHS) of (ref) separately. To address the first term on the RHS of (ref), observe that (ref) immediately gives \begin{align} \frac{1}{\sqrt{T_1}\sigma \|\widetilde{w}^{(i,d)}\|_2} \sum_{t > T_0} \langle \mathbb{E}[Y_{t, \mathcal{I}^{(d)}}], \mathcal{P}_{V_\text{pre}} \Delta^{(i,d)} \rangle = o_p(1). \end{align} For the second term on the RHS of (ref), note that $\langle \varepsilon_{t, \mathcal{I}^{(d)}}, \widetilde{w}^{(i,d)} \rangle$ are independent across $t$. Note that $ \langle \varepsilon_{t, \mathcal{I}^{(d)}}, \widetilde{w}^{(i,d)} / (\sigma \|\widetilde{w}^{(i,d)}\|_2) \rangle$ is a sub-Gaussian random variable with variance equal to $1$ and sub-Gaussian norm less than some constant $C > 0$. It is easy to verify that the Lyapunov condition Billingsley for $\delta = 1$ holds, and hence the Lyapunov Central Limit Theorem yields \begin{align} \frac{1}{\sqrt{T_1}} \sum_{t > T_0} \Big\langle \varepsilon_{t, \mathcal{I}^{(d)}}, \frac{\widetilde{w}^{(i,d)}}{\sigma \|\widetilde{w}^{(i,d)}\|_2} \Big\rangle \xrightarrow{d} \mathcal{N}\big( 0, 1\big). \end{align} as $T_1 \to \infty$. For the third term on the RHS of (ref), we scale the RHS of (ref) by $\sqrt{T_1}/ (\sigma \|\widetilde{w}^{(i,d)}\|_2)$ and recall (ref). This yields \begin{align} \frac{1}{\sqrt{T_1} \sigma \|\widetilde{w}^{(i,d)}\|_2 } \sum_{t > T_0} \langle \varepsilon_{t, \mathcal{I}^{(d)}}, \Delta^{(i,d)} \rangle &= o_p(1). \end{align} Collecting (ref), (ref), (ref), we establish (ref). \end{proof} \subsection{Proof of Lemma (ref)} We state the proof of Lemma (ref), which is an adaptation of the proof of Lemma (ref) in Appendix (ref). \begin{proof} Below we establish (ref) and (ref) hold. \begin{itemize} • {\em Proof of (ref).} Follows from conditions (v) and (vi) in the lemma statement. • {\em Proof of (ref).} By condition (ii), the rowspaces of both $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{(r_\text{pre}, \Omega)}$ and $\boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}}^\Omega$ span all of $\mathbb{R}^{r_\text{pre}}$. Hence, the orthogonal projection matrices onto their rowspaces obey \begin{align} \left( \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{(r_\text{pre}, \Omega)}\right)^\dagger \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{(r_\text{pre}, \Omega)} = \left(\boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}}^\Omega \right)^\dagger \boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}}^\Omega = \boldsymbol{I}_{r_\text{pre}}, \end{align} where $\boldsymbol{I}_{r_\text{pre}}$ denotes the $r_\text{pre} \times r_\text{pre}$-dimensional identity matrix. The remainder of this proof is dedicated to bounding $\Delta^{(i,d,\Omega)} = \| \widehat{w}^{(i,d,\Omega)} - w^{(i,d,\Omega)} \|_2$ using the proof of Lemma (ref) in Appendix (ref) as a guide. Using (ref) and the arguments that led to (ref), we obtain \begin{align} \| \Delta^{(i,d,\Omega)} \|_2^2 & \leq \frac{4}{(\widehat{s}^\Omega_{r_\text{pre}})^2} \left( \| \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{(r_\text{pre}, \Omega)} - \boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}}^\Omega \|_{2,\infty}^2 \cdot \|\widetilde{w}^{(i,d,\Omega)}\|_1^2 +\langle \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{(r_\text{pre}, \Omega)} \cdot \Delta^{(i,d,\Omega)}, \varepsilon_{\text{pre}, n} \rangle\right). \end{align} where $\widehat{s}^\Omega_{k}$ is defined as the $k$-th singular value of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{(r_\text{pre}, \Omega)}$ while $\varepsilon_{\text{pre},n}$ maintains the same definition as set in Appendix (ref). By construction, it follows that \begin{align} \| \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{(r_\text{pre}, \Omega)} - \boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}}^\Omega \|_{2,\infty}^2 \le \| \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{r_\text{pre}} - \boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}} \|_{2,\infty}^2 \end{align} since $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{(r_\text{pre}, \Omega)}$ and $\boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}}^\Omega$ are sub-matrices of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{r_\text{pre}}$ and $\boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}}$, respectively. Let $s^\Omega_{k}$ denote the $k$-th singular value of $\boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}}^\Omega$. Further, without loss of generality, suppose $\Omega$ corresponds to the first $r_\text{pre}$ entries of $\mathcal{I}^{(d)}$. This allows us to write \begin{align} \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{(r_\text{pre}, \Omega)} = \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{r_\text{pre}} \boldsymbol{A}, \qquad \boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}}^\Omega = \boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}} \boldsymbol{A}, \end{align} where $\boldsymbol{A} = [\boldsymbol{I}_{r_\text{pre}}, 0_{r_\text{pre} \times (N_d-r_\text{pre})}]^\top$. Combining (ref) with Lemma (ref), we have \begin{align} | \widehat{s}^\Omega_{r_\text{pre}} - s^\Omega_{r_\text{pre}} | &\le \| \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{(r_\text{pre}, \Omega)} - \boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}}^\Omega \|_\text{op} \\ &= \| (\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{r_\text{pre}} - \boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}} ) \cdot \boldsymbol{A} \|_\text{op} \\ &\le \| \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{r_\text{pre}} - \boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}} \|_\text{op} &&\because \| \boldsymbol{A} \|_\text{op} = 1 \\ &\le \| \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}^{r_\text{pre}} - \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} \|_\text{op} + \| \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} - \boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}} \|_\text{op} \\ &= \widehat{s}_{r_\text{pre} + 1} + \| \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} - \boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}} \|_\text{op}, \end{align} where we recall $\widehat{s}_k$ is the $k$-th singular value of $\boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}}$. Let us define $s_k$ analogously to $\widehat{s}_k$ with respect to $\boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}}$. By condition (i), $s_{r_\text{pre}+1} = 0$ and thus \begin{align} \widehat{s}_{r_\text{pre}+1} &= \widehat{s}_{r_\text{pre}+1} - s_{r_\text{pre}+1} \\ &\le \| \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} - \boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}} \|_\text{op}. &&\because \text{Lemma (ref)} \end{align} Applying Assumption (ref) and Lemma (ref) to (ref) and (ref) yields w.p. at least $1-2\exp(-\zeta^2)$, \begin{align} | \widehat{s}^\Omega_{r_\text{pre}} - s^\Omega_{r_\text{pre}} | &\le 2 \cdot \| \boldsymbol{Y}_{\text{pre}, \mathcal{I}^{(d)}} - \boldsymbol{M}_{\text{pre}, \mathcal{I}^{(d)}} \|_\text{op} \\ &\le C \sigma \left( \sqrt{T_0} + \sqrt{N_d} + \zeta \right) \end{align} for any $\zeta > 0$. Next, we follow the arguments that led to (ref). In particular, we plug (ref) and (ref) into (ref), and invoke Lemmas (ref), (ref), (ref), and conditions (iii)--(iv) to conclude \begin{align} \Delta^{(i,d,\Omega)} &= O_p \left( \frac{r_\text{pre} \sqrt{ \log(T_0 N_d)} }{ \sqrt{N_d}} + \frac{1}{T_0^{1/4}} \right). \end{align} Then, \begin{align} &\frac{1}{\sqrt{T_1} \|w^{(i,d,\Omega)}\|_2} \sum_{t > T_0} \left \langle \mathbb{E}[Y^\Omega_{t, \mathcal{I}^{(d)}}], \left( \boldsymbol{M}^\Omega_{\text{pre}, \mathcal{I}^{(d)}} \right)^\dagger \boldsymbol{M}^\Omega_{\text{pre}, \mathcal{I}^{(d)}}\left(\widehat{w}^{(i,d,\Omega)} - w^{(i,d,\Omega)} \right) \right \rangle, \\ &\quad \le \frac{1}{\sqrt{T_1} \|w^{(i,d,\Omega)}\|_2} \sum_{t > T_0} \| \mathbb{E}[Y^\Omega_{t, \mathcal{I}^{(d)}}] \|_2 \cdot \| \Delta^{(i,d,\Omega)} \|_2, \\ &\quad\le \frac{1}{\sqrt{T_1} \|w^{(i,d,\Omega)}\|_2} \sum_{t > T_0} \sqrt{r_\text{pre}} \cdot \| \Delta^{(i,d,\Omega)} \|_2, \\ &\quad= O_p \left( \frac{ \sqrt{T_1} r^{2}_\text{pre} \sqrt{ \log(T_0 N_d)} }{ \sqrt{N_d}} + \frac{ \sqrt{T_1} r_\text{pre}}{T_0^{1/4}} \right), \\ &\quad= o_p(1), \end{align} where the second to last line follows from (ref) and condition (v), and the last line follows from condition (vii). \end{itemize} \end{proof}