EconBase
← Back to paper

A Causal Inference Framework for Data Rich Environments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

93,803 characters · 0 sections · 19 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Causal Inference Framework for Data Rich Environments

abstractWe propose a formal model for counterfactual estimation with unobserved confounding in “data-rich” settings, i.e., where there are a large number of units and a large number of measurements per unit. Our model provides a bridge between the structural causal model view of causal inference common in the graphical models literature with that of the latent factor model view common in the potential outcomes literature. We show how classic models for potential outcomes and treatment assignments fit within our framework. We provide an identification argument for the average treatment effect, the average treatment effect on the treated, and the average treatment effect on the untreated. For any estimator that has a fast enough estimation error rate for a certain nuisance parameter, we establish it is consistent for these various causal parameters. We then show principal component regression is one such estimator that leads to consistent estimation, and we analyze the minimal smoothness required of the potential outcomes function for consistency.

\setcounter{page}{1} {0.5\baselineskip}

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\center\normalfont\bf}{Introduction} One of the central goals of empirical economic research is to ascertain the effects of treatments (policies, treatments) on the outcomes of interest. A fundamental challenge for the estimation of treatment effects is the pervasive presence of unobserved confounders. For example, in a study of the effects of health insurance on healthcare utilization, unobserved or latent health determinants may differ between insured and uninsured individuals, biasing treatment effect estimates. Several approaches have been put forward to estimate treatment effects in the presence of confounders, including explicit randomization of the treatment, controlling for measured confounders, and instrumental variable methods. Traditionally, these methods are not designed to operate in data-rich environments where the curse of dimensionality creates challenges for estimation and inference, and do not take advantage of the information contained in high-dimensional data to identify treatment effects.

In recent times, the availability of high-dimensional data on economic behavior has become commonplace. Modern data harvesting technologies, based on digitization and pervasive sensors, enable the collection of detailed high-frequency attribute and outcome information on individuals (or other observational units; e.g., geo-locations) concurrently undergoing different treatments. For example, electronic health records contain rich information about patients' medical history over time. Similarly, internet retailers and marketing firms use scanner data to collect high-dimensional information on customers' purchases. The goal of this article is to provide a framework for causal inference that takes advantage of modern data-rich environments to counter the effect of unobserved or latent confounders.

Given this goal, we consider a setting where we have access to data for $N$ units (e.g., individuals, sub-populations, firms, geographic locations) and $T$ measurements of outcomes per unit. Different measurements may represent the same outcome metric at different time periods, different outcome metrics (e.g., customers' expenditures in different product categories) for the same time period, or a combination of both. We argue that (high-dimensional) data-rich environments---i.e., large $N$ and large $T$---make it possible to estimate treatment effects in the presence of unobserved or latent confounding, without needing to make parametric assumptions in the manner in which unobserved confounders affect selection for treatment and the outcome metrics.

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Contributions and Related Work} {\bf Contributions.} We propose a formal model for counterfactual estimation with unobserved confounding in “data-rich” settings, i.e., where there are a large number of units and a large number of measurements per unit. We posit a general data-generating process (DGP) for how potential outcomes are defined and how treatments are assigned, allowing for unobserved confounding. We provide a structural causal model view of the conditional independence conditions required in our DGP that imply that the treatment assignments are exogenous of the potential outcomes conditional on these unobserved confounders. We establish that if the unobserved confounders are low-dimensional, relative to the number of units and measurements, and the potential outcomes are a smooth non-linear function of them, it {\em implies} that an approximate linear latent factor model of appropriate dimension holds, where the approximation error decays as the number of units and measurements increase. In doing so, we believe this model provides a formal bridge between the structural causal model view of causal inference common in the graphical models literature with that of the latent factor model view common in the potential outcomes literature.

We formalize how classic models for potential outcomes and treatment assignments fit within our framework. For the potential outcomes, we show how two-way fixed effects, interactive fixed effects, binary choice, and dictionary basis expansions fit within our framework. For treatment assignments, we show how randomized control trials (RCTs), selection on (un)observables, regression discontinuity, random utility models, and staggered adoption settings fit within our framework. Theoretically, we provide an identification argument for the average treatment effect (ATE), the average treatment effect on the treated (ATT), and the average treatment effect on the untreated (ATU). For any estimator that has a fast enough estimation error rate for a certain nuisance parameter, we establish it is consistent for these various causal parameters. We then show principal component regression (PCR) is one such estimator that leads to consistent estimation, and we analyze the minimal smoothness required of the potential outcomes function for consistency.

{\bf Related work.} This model builds upon the latent factor model literature studied in the growing literature on causal panel data models (chamberlain_factor, SC, bai2009interactive, athey2021matrix, bai2021matrix, SDID, CMC, SI, dwivedi2022doubly, synth_combo). The model we propose can be viewed as a generalization of the exact linear factor model studied in these works to a non-linear factor model. Our model allows for both panel data and cross-sectional data, and combinations thereof. Importantly, we argue that beginning from a general structural causal model, if the outcomes are a smooth function of the unobserved confounders, then it {\em implies} that a factor model of appropriate dimension holds if there are large number of units and measurements, thereby hopefully providing a bridge between these two frameworks. In terms of the estimator we propose, to the best of our knowledge, this is also the first theoretical analysis of consistency for the ATE with a (smooth) non-linear factor model. It is also the first analysis of PCR (PCR1, PCR2, CICD, PCR3) for such target causal estimands with unobserved confounding. This requires dealing with the novel technical challenge of error-in-variables, only an approximate low-rank model holding on the noiseless covariates, and linear misspecification error.

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Notation} For a matrix $\bA \in \Rb^{a \times b}$, we denote its transpose as $\bA^T \in \Rb^{b \times a}$. We denote the operator (spectral) and Frobenius norms of $\bA$ as $\|\bA\|_{\text{op}}$ and $\|\bA\|_F$, respectively. The columnspace (or range) of $\bA$ is the span of its columns, which we denote as $\Rc(\bA) = \{v \in \Rb^a: v = \bA x, x \in \Rb^b\}$. The rowspace of $\bA$, given by $\Rc(\bA^T)$, is the span of its rows. Recall that the nullspace of $\bA$ is the set of vectors that are mapped to zero under $\bA$. For any vector $v \in \Rb^a$, let $\|v\|_p$ denote its $\ell_p$-norm, and let $\|v\|_\infty$ denote its max-norm. The inner product between vectors $v, x \in \Rb^a$ is $\langle v, x \rangle = \sum_{\ell=1}^a v_\ell x_\ell$. If $v$ is a random variable, we denote its sub-Gaussian (Orlicz) norm as $\|v\|_{\psi_2}$. For any positive integer $a$, we use the notation $[a] = \{1, \dots, a\}$.

Let $f$ and $g$ be two real-valued functions defined on $\mathcal X$, an unbounded subset of $[0,\infty)$. We say that $f(x)$ = $O(g(x))$ if and only if there exists a positive real number $M$ and $x_0\in \mathcal X$ such that, for all $x \ge x_0$, we have $|f (x)| \le M|g(x)|$. Analogously, we say that $f (x) = \Theta(g(x))$ if and only if there exist positive real numbers $m, M$ and $x_0\in \mathcal X$ such that for all $x \ge x_0$, we have $m|g(x)| \le |f(x)| \le M|g(x)|$; $f (x) = o(g(x))$ if for any $m > 0$, there exists $x_0\in \mathcal X$ such that for all $x \ge x_0$, we have $|f(x)| \le m|g(x)|$.

We adopt standard notation and definitions for stochastic convergence. We employ $\xrightarrow{d}$ and $\xrightarrow{p}$ to indicate convergence in distribution and probability, respectively. For any sequence of random vectors, $X_n$, and any sequence of positive real numbers, $a_n$, we say $X_n = O_p(a_n)$ if for every $\varepsilon>0$, there exists constants $C_\varepsilon$ and $n_\varepsilon$ such that $\mathbb{P}( \| X_n \|_2 > C_\varepsilon a_n) < \varepsilon$ for every $n \ge n_\varepsilon$; equivalently, we say $(1/a_n) X_n$ is uniformly tight or bounded in probability. $X_n = o_p(a_n)$ means $X_n/a_n\xrightarrow{p} 0$. We say a sequence of events $\Ec_n$, indexed by $n$, holds “with probability approaching one” (w.p.a.1) if $\mathbb{P}(\Ec_n) \rightarrow 1$ as $n \rightarrow \infty$, i.e., for any $\varepsilon > 0$, there exists a $n_\varepsilon$ such that for all $n > n_\varepsilon$, $\mathbb{P}(\Ec_n) > 1 - \varepsilon$. More generally, a multi-indexed sequence of events $\Ec_{n_1,\dots, n_d}$, with indices $n_1,\dots, n_d$ with $d \geq 1$, is said to hold w.p.a.1 if $\mathbb{P}(\Ec_{n_1,\dots, n_d}) \rightarrow 1$ as $\min\{n_1,\dots, n_d\} \rightarrow \infty$. We also use $\Nc(\mu, \sigma^2)$ to denote a normal or Gaussian distribution with mean $\mu$ and variance $\sigma^2$---we call it {\em standard} normal or Gaussian if $\mu = 0$ and $\sigma^2 = 1$. We use $C$ to denote a positive constant, with a value that can change across instances.

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\center\normalfont\bf}{Model} We are interested in evaluating the effect of treatments on outcomes of interest. Specifically, we observe $T$ outcomes or measurements for $N$ units. For each measurement $t \in [T]$ and unit $n \in [N]$, we observe $Y_{n, t} \in \Rb$ under treatment $A_{n, t} \in \Ac$, where $|\Ac| = A$.

Let $\bA = [A_{n, t}]_{n \in [N], t \in [T]} \in \Ac^{N \times T}$ and $\boldsymbol{Y} = [Y_{n, t}]_{n \in [N], t \in [T]} \in \Rb^{N \times T}$ collect the matrix of treatment assignments and outcomes, respectively. We now define how the treatment assignments and outcomes are generated. We define the random variables $\bU \in \mathcal{U}, \bE \in \mathcal{E}$ and functions $h: \mathcal{U} \to \Ac^{N \times T}$, $f: \Ac^{N \times T} \times \mathcal{U} \times \mathcal{E} \to \Rb^{N \times T}$ such that

align[align omitted — 112 chars of source]

We allow the functions $h, f$ and the variables $\bU, \bE$ to be unobserved; that is, we only observe the treatment assignments $\bA$ and the outcomes $\boldsymbol{Y}$. $\bU$ contains all potential confounders that can affect both the treatment assignment $\bA$ and the outcomes $\boldsymbol{Y}$. $\bE$ is the random variation in $\boldsymbol{Y}$ not explained by $\bU$. The question we study in this paper is:

center[center omitted — 218 chars of source]

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Data Generating Process} Towards answering the question above, we assume the following data-generating process (DGP). As stated earlier, we hope this DGP serves a bridge between the SCM and latent factor model view of causal inference.

assumption[Data generating process]{\color{white}.} \begin{enumerate} • We assume the following factorization of $\bU, \bE$: \begin{align} \bU = [U_n]_{n \in [N]}, \quad \bE = [\varepsilon^{(a)}_{n, t}]_{n \in [N], t \in [T], a \in \Ac} \end{align} where $U_n \in \Rb^q$, and $\varepsilon^{(a)}_{n, t} \in \Rb$ for some $q \geq 1$. • We do not make any distributional assumptions about $\bU$ and it can be thought to be conditioned on for the remainder of the paper. Conditional on $\bU$, for all $n\in [N]$ and $t\in [T]$, we assume the vector $(\varepsilon^{(a)}_{n, t})_{a \in \Ac}$ is sampled independently. Hence, the only source of uncertainty in our model is due to $\bE$. • We assume $f$ has the following factorization: for $n \in [N], t \in [T]$, potential and observed outcomes are generated as \begin{align} Y^{(a)}_{n, t} &= f_{t, a}\Big(U_n\Big) + \varepsilon^{(a)}_{n, t}, for a \in \Ac, \\ Y_{n, t} &= Y^{(A_{n,t})}_{n, t}, \end{align} where $f_{t, a}: \Rb^q \to \Rb$, $\varepsilon^{(a)}_{n, t} \in \Rb$, and we assume $\mathbb{E}[\varepsilon^{(a)}_{n, t} \mid \bU] = 0$. $\varepsilon^{(a)}_{n, t}$ can be interpreted as capturing the random variation in the potential outcomes $Y^{(a)}_{n, t}$ that is not captured by $f_{t, a}$. As discussed in Section (ref), various models for potential outcomes considered in the literature can be captured via Eq (ref) as long as $f_{t, a}$ is sufficiently smooth. \end{enumerate}
remarkThe DAG in Figure (ref) is consistent with the independence assumptions we make above in the DGP. Further, Assumption (ref) implies the following conditional exogeneity condition \begin{align} Y^{(a)}_{n, t} \mathbin{ \mathpalette{\@indep} } \bA \mid U_n. \end{align}
figure[figure omitted — 501 chars of source]
remarkA potential outcome function of the form $Y^{(a)}_{n, t} = \bar{f}_{t, a}(U_n, \bar{\varepsilon}^{(a)}_{n, t})$ can be nested into (ref) under additional conditions on the distribution of $\bar{\varepsilon}^{(a)}_{n, t}$. In particular, assume that the distribution of $\bar{\varepsilon}^{(a)}_{n, t}$ is independent of $n$, conditional on $\bU$. Then $$\mathbb{E}[Y^{(a)}_{n, t} \mid \bU] = \mathbb{E}[\bar{f}_{t, a}(U_n, \bar{\varepsilon}^{(a)}_{n, t}) \mid \bU] = f_{t,a}(U_n),$$ where the expectation is taken with respect to $\bar{\varepsilon}^{(a)}_{n, t}$. In particular, because the distribution of $\bar{\varepsilon}^{(a)}_{n, t}$ is not dependent on $n$, the conditional expectation $\mathbb{E}[Y^{(a)}_{n, t} \mid U_n]$ can be written as only a function of $U_n$, and $t, a$. Then by defining $\varepsilon^{(a)}_{n, t} = Y^{(a)}_{n, t} - \mathbb{E}[Y^{(a)}_{n, t} \mid U_n],$ (ref) holds.

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\center\normalfont\bf}{Outcome and Treatment Assignment Functions Within our Framework} Thus far, the setup has been quite general. To make progress, we impose relatively generic smoothness conditions on $f_{t, a}$, and argue this encompasses familiar models for potential outcomes considered in the econometric literature. Further, we show how various models for the treatment assignment functions studied in the literature can be encompassed within our framework.

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Outcome Functions}

We first formally define what we mean by smoothness.

definition[H\"older continuity, e.g., jiaming_spectral_convergence, jiaming_spectral_convergence] For $k \geq 1$, let $s = (s_1, \dots, s_k)$ be a $k$-tuple of non-negative integers with $|s| = \sum^k_{\ell = 1} s_\ell$. For $S \in \Nb$ and $C_H>0$, the H\"older class $\Hc(k,S,C_H)$ on $[0, 1)^k$ is the set of functions $g: [0, 1)^k \to \mathbb{R}$ with partial derivatives that satisfy \[ \sum_{s: |s| = S - 1} \frac{1}{s!} |\nabla_{\!s} g(\mu) - \nabla_{\!s} g(\mu') | \le C_H \norm{\mu - \mu'}_{\infty},\quad \forall \mu, \mu' \in [0, 1)^k. \footnote{Note that for any compact set $\mathcal{X} \in \Rb^k$, we have $\mathcal{X} \subset [-c, c)^k$, where $c \le \infty$. Then, $[-c, c)^k$ can be replaced by $[0, 1)^k$, without loss of generality by re-scaling.} \]

In essence, Definition (ref) requires that the $(S-1)$-th derivatives of $g$ are Lipchitz continuous. For example, it is easy to verify that an analytic function with compact domain is H\"older continuous for all $S \in \Nb$.

assumptionFor all $n \in [N]$, recall $U_n \in [0,1)^q$. For all $t \in [T], a \in \Ac$, we assume $f_{t, a}$ is H\"older continuous, i.e., $f_{t, a} \in \Hc(q,S,C_H)$, where $C_H < C < \infty$.

Informally, Assumption (ref) is a continuity condition that posits that if latent variables $U_{n_1}$ and $U_{n_2}$ for any two units $n_1$ and $n_2$ are close ($U_{n_1}\approx U_{n_2}$), then their average potential outcomes are close as well, ($\mathbb{E}[Y^{(a)}_{n_1, t}] \approx \mathbb{E}[Y^{(a)}_{n_2, t}]$, for all $t \in [T], a \in \Ac$), where the expectation is taken with respect to $\varepsilon^{(a)}_{n_1, t}$ and $\varepsilon^{(a)}_{n_2, t}$, respectively.

A linear factor model, $f_{t, a}(U_n) = \left\langle U_n, \tilde{U}_{t, a} \right\rangle$, is a special case of Assumption (ref) and one can verify it satisfies Definition (ref) for all $S \in \mathbb{N}$ . Proposition (ref) establishes that linear factor models of sufficiently large dimension also provide a “universal” representation for smooth non-linear factor models.

proposition[H\"older low rank matrix approximation, jiaming_spectral_convergence, jiaming_spectral_convergence] Suppose Assumption (ref) holds. Then, for all $n \in [N], t \in [T], a \in \Ac$, and any $\delta>0$, there exist latent variables $\lambda_n, \rho_{t, a} \in \Rb^r$ such that: \begin{align} \left| f_{t, a}(U_n) - \left\langle \lambda_n, \rho_{t, a} \right\rangle \right| \leq \Delta_E, \end{align} where for $\bar{C}$ that is allowed to depend on $(q,S)$, \begin{align} \Delta_E \le C_H \cdot \delta^S\quadwith\quad r \le \bar{C} \cdot \delta^{-q}. \end{align}

Proposition (ref) establishes that if $f_{t, a}$ has a H\"older smooth latent variable representation, then it is uniformly well-approximated by a linear factor model of finite dimension, $r$. For $\delta < 1$ we have that as the latent dimension $q$ of the confounder $U_n$ increases, the bound on the rank $r$ increases, and as smoothness $S$ of $f_{t, a}$ increases, the bounds on the approximation error $\Delta_E$ decreases. If we take $\delta = (\mbox{min}\{N, T\})^{-c}$ for some constant $c$, such that $0<c<1/q$, we obtain $r \ll \mbox{min}\{N, T\}$ and $\Delta_E = o(1)$ as $N, T \to \infty$.

\@startsection{subsubsection}{3}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont}{Classic Econometric Models that Fit our Framework}

Our framework nests some classical econometric models, which we describe below.

example[Two-way fixed effects model] Suppose \[ Y^{(a)}_{n, t} = \left\langle a, \beta_n \right\rangle + \mu_n + w_t + \varepsilon^{(a)}_{n, t} \] Here, $\mu_n$, $v_t$ are scalars, and $a$ and $\beta_n$ have dimension $p$. This corresponds to a two-way fixed effect model with heterogeneous coefficients on the treatment. Now let \[ U_n = \left(\begin{array}{c} \beta_n\\ \mu_n\\operatorname{\mathbbm 1} \end{array}\right), \quad \tilde{U}_{t, a} = \left(\begin{array}{c}a \\ 1 \\ w_t \end{array}\right),\vspace{.2cm} \] and \[ f_{t, a}(U_n) = \left\langle U_n, \tilde{U}_{t, a} \right\rangle. \] Then \[ Y^{(a)}_{n, t} = f_{t, a}(U_n) + \varepsilon^{(a)}_{n, t}, \] where $f_{t, a}$ is linear in $U_n$ (i.e., is H\"older continuous), and $r = p+2$. Time-varying treatment coefficients can be easily accommodated in the same setting. Now suppose \[ Y^{(a)}_{n, t} = \left\langle a, \beta_{n, t} \right\rangle+\mu_n + w_t + \varepsilon^{(a)}_{n, t} \] where $\beta_{n,t}=F_t\beta_n$, and $F_t$ is a $(p\times p)$ matrix of time-varying coefficients and the dimensions of the other components are unchanged. Now let \[ U_n = \left(\begin{array}{c} \beta_n\\ \mu_n\\operatorname{\mathbbm 1} \end{array}\right), \quad \tilde{U}_{t, a} = \left(\begin{array}{c} \left\langle F_t, a \right\rangle \\ 1 \\ w_t \end{array}\right),\vspace{.2cm} \] and \[ f_{t, a}(U_n) = \left\langle U_n, \tilde{U}_{t, a} \right\rangle. \] Then again \[ Y^{(a)}_{n, t} = f_{t, a}(U_n) + \varepsilon^{(a)}_{n, t}, \] where $f_{t, a}$ is linear in $U_n$, and $r = p+2$.
example[Interactive fixed effects model, bai2009interactive, bai2009interactive] Suppose \[ Y^{(a)}_{n, t} = \left\langle a, \beta \right\rangle + \left\langle \mu_n, w_t \right\rangle + \varepsilon^{(a)}_{n, t}, \] where $\mu_n$ and $v_t$ are factors of dimension $k$. Let \[ U_n = \left(\begin{array}{c} 1\\ \mu_n\end{array}\right), \quad \tilde{U}_{t, a} = \left(\begin{array}{c} \left\langle a, \beta \right\rangle \\ w_t \end{array}\right), \] and \[ f_{t, a}(U_n) = \left\langle U_n, \tilde{U}_{t, a} \right\rangle, \] Then \[ Y^{(a)}_{n, t} = f_{t, a}(U_n) + \varepsilon^{(a)}_{n, t}, \] where $f_{t, a}$ is linear in $U_n$, and $r = k+1$.
example[Tensor factor model, SI] Suppose \[ Y^{(a)}_{n, t} = \left\langle \mu_n, w^{(a)}_t \right\rangle + \varepsilon^{(a)}_{n, t}, \] where $\mu_n, w^{(a)}_t$ are factors of dimension $k$. One can verify the models in Examples (ref) and (ref) are special cases of the model above. Here there are unit-specific heterogeneous coefficients on the treatment, and in addition the treatment can be time-varying. Let \[ U_n = \mu_n, \quad \tilde{U}_{t, a} = w^{(a)}_t, \] and \[ f_{t, a}(U_n) = \left\langle U_n, \tilde{U}_{t, a} \right\rangle, \] Then \[ Y^{(a)}_{n, t} = f_{t, a}(U_n) + \varepsilon^{(a)}_{n, t}, \] where $f_{t, a}$ is linear in $U_n$, and $r = k$.
example[Dictionary basis expansion] Consider \[ Y_{n, t}^{(a)}= \gamma_n(a, X_t) + \varepsilon^{(a)}_{n, t}, \] where $X_t, a \in \Rb^p$, and $\gamma_n: \Rb^{2p} \to \Rb$ has the following dictionary representation \[ \gamma_n(a, X_t) = \sum^L_{\ell = 1} \alpha_{n, \ell} b_{\ell}(a, X_t), \] where $b_{\ell}: \Rb^{2p} \to \Rb$ are dictionary basis functions, and $\alpha_{n, \ell} \in \Rb$, the corresponding linear coefficients. For example, $b_{\ell}$ could be a polynomial of $a$ and $X_t$. Then, we can let \[ U_n = \left(\begin{array}{c} \alpha_{n, 0}\\ \vdots \\ \alpha_{n, L} \end{array}\right), \quad \tilde{U}_{t, a} = \left(\begin{array}{c}b_{0}(a, X_t)\\ \vdots \\ b_{L}(a, X_t) \end{array}\right), \vspace{.2cm} \] and \[ f_{t, a}(U_n) = \left\langle U_n, \tilde{U}_{t, a} \right\rangle. \] Then \[ Y^{(a)}_{n, t} = f_{t, a}(U_n) + \varepsilon^{(a)}_{n, t}, \] where $f_{t, a}$ is linear in $U_n$, and $r = L$. Note that in our model the dictionary basis functions $(b_{\ell})_{\ell \in [L]}$ and the covariates $X_t$ can be unobserved. Further, our consistency results allow for $L$ to be increasing in $N, T$, as long as $L = o(\min(N, T))$.
example[Binary choice] Let $I_{\mathcal S}$ be the indicator function for set $\mathcal S$. Suppose \[ Y_{n, t}^{(a)}=I_{[0,\infty)}\left(F\left(\left\langle a, \beta \right\rangle + \mu_n+w_t\right) - e_{n, t}^{(a)}\right), \] where $F: \Rb \to [0, 1]$ is a H\"{o}lder continuous function, and for every $(n, t, a)$, $e_{n, t}^{(a)}$ is an independent realization of a continuous random variable. Without loss of generality, we can assume that $e_{n, t}^{(a)}$ is uniformly distributed on $[0,1]$---if $e_{n, t}^{(a)}$ is not uniform on $[0,1]$, we can apply the probability integral transform, $f$, to both $F\left(\left\langle a, \beta \right\rangle + \mu_n + w_t\right)$ and $e_{n, t}^{(a)}$, where $f$ is the cumulative distribution function of $e_{n, t}^{(a)}$. Then, by Remark (ref), we can let \[ U_n = \left(\begin{array}{c} 1\\ \mu_n\\operatorname{\mathbbm 1} \end{array}\right), \quad \tilde{U}_{t, a} = \left(\begin{array}{c}\left\langle a, \beta \right\rangle \\ 1 \\ w_t \end{array}\right).\vspace{.2cm} \] and \[ f_{t, a}(U_n) = F\left(\left\langle U_n, \tilde{U}_{t,a} \right\rangle\right). \] Then \[ Y^{(a)}_{n, t} = f_{t, a}(U_n) + \varepsilon^{(a)}_{n, t}, \] where $f_{t, a}$ is H\"older continuous in $U_n$. Similar to the examples above, we can easily generalize this model to where we have $F\Big(\left\langle \mu_n, w^{(a)}_{t} \right\rangle\Big)$, i.e., we have unit-specific heterogeneous coefficients on treatments, and in addition, the treatments can be time-varying.

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Treatment Assignment Functions} Without loss of generality, we will let $\Ac = \{0, 1\}$. Below we show how various models for treatment assignment functions studied in the literature fit within our framework. In particular, our proposed DGP requires unconfoundedness conditional on $U_n$. We formally establish how and when this condition holds for commonly studied treatment assignment functions.

We denote $h_{n, t}(\bU)$ as the $(n, t)$-th output of $h(\bU)$, i.e., the treatment $A_{n, t}$.

example[Randomized trial] Consider a setup of a randomized trial where the $N$ units are assigned one of the two treatments at random, where the probability can differ across measurements. Specifically, for all $n \in [N], t \in [T]$ \begin{align} A_{n, t} = \begin{cases} 1 & with probability p_t, \\ 0 & otherwise, \end{cases} \end{align} independent of everything else. Here, we can take \[ h_{n, t}(\nu_{n, t}) = \mathbbm{1}\{v_{n,t}\leq p_t\}, \] where $\nu_{n,t}$ is a random variable uniformly distributed in $[0, 1]$. In this case, there is no confounding as the treatment assignment is not correlated with the outcomes, and so $h_{n, t}$ is not a function of $U_n$.
example[Selection on (un)observables] Suppose there are unobserved (or partially observed) covariates $U_n \in \Rb^q$ such that \[ Y^{(a)}_{n, t} = f_{t, a}(U_n) + \varepsilon^{(a)}_{n, t}. \] Further, the treatment assignment for all $n \in [N], t \in [T]$ is given by \begin{align} A_{n, t} = \begin{cases} 1 & with probability \sigma_{t}(U_n), \\ 0 & otherwise, \end{cases} \end{align} where $\sigma_t$ is a function mapping to $[0, 1]$ (e.g., the logistic function). Here $U_n$ is an unobserved confounder as it affects both the potential outcome $Y^{(a)}_{n, t}$, and is the input to the treatment assignment function $\sigma_{t}$. Here, we can take \[ h_{n, t}(U_n) = \mathbbm{1}\{v_{n,t}\leq\sigma_t(U_n)\}, \] where $\nu_{n,t}$ is a random variable uniformly distributed in $[0, 1]$. Our framework allows for treatment assignments where positivity does not hold, i.e., $\sigma_{t}(U_n)$ equals $0$ or $1$, by exploiting the smoothness of $f_{t, a}$ as given by Assumption (ref).
example[Regression discontinuity] Suppose \[ Y^{(a)}_{n, t} = f_{t, a}(U_n) + \varepsilon^{(a)}_{n, t}, \] as in Eq (ref). Further, suppose that treatment assignment for unit $n$ and measurement $t$ is a function of a score $X_{n, t}$. In particular, treatment is given if the score $X_{n, t} \in \Rb^p$ is lower than some threshold $\theta_{n, t} \in \Rb^p$, i.e., \begin{align} A_{n, t} = \begin{cases} 1, & if X_{n, t} > \theta_{n, t} \\ 0 & otherwise \end{cases} \end{align} where \[ X_{n, t} = \ell_{n, t}(U_n), \] with $\ell_{n, t}: \Rb^q \to \Rb^p$. Here $U_n$ is an unobserved confounder as it affects both the potential outcome $Y^{(a)}_{n, t}$, and the score $X_{n, t}$, which in turn deterministically affects treatment assignment. Here, we can take \begin{align} h_{n, t}(U_n) = \mathbbm{1}(\ell_{n, t}(U_n) > \theta_{n, t}). \end{align} As we show later, despite a regression discontinuity treatment assignment function, our framework allows for the estimation of treatment effects for units away from the threshold $\theta_{n, t}$ by exploiting the smoothness of $f_{t, a}$ as given by Assumption (ref).
example[Random utility model] Suppose \[ Y^{(a)}_{n, t} = f_{t, a}(U_n) + \varepsilon^{(a)}_{n, t}, \] as in Eq (ref). Further, suppose that treatment assignment for unit $n$ and measurement $t$ is given as follows: \begin{align} A_{n, t} = \begin{cases} 1, & if \ell_{t, 1}(U_n) - \ell_{t, 0}(U_n) + \nu_{n, t} > \delta_{n, t} \\ 0 & otherwise \end{cases} \end{align} where $\ell_{t, 0}, \ell_{t, 1}: \Rb^q \to \Rb$, $\nu_{n, t}, \delta_{n, t} \in \Rb$. If $\nu_{n, t}$ has a logistic distribution, then this recovers the Luce model (luce_model). Here, we can simply take \begin{align} h_{n, t}(U_n) = \mathbbm{1}(\ell_{t, 1}(U_n) - \ell_{t, 0}(U_n) + \nu_{n, t} > \delta_{n, t}). \end{align}
example[Staggered adoption] Suppose we have a panel data model where $t$ corresponds to a time point and potential outcomes are given by \[ Y^{(a)}_{n, t} = f_{t, a}(U_n) + \varepsilon^{(a)}_{n, t}, \] as in Eq (ref). Further, suppose that treatment assignment for unit $n$ and time $t$ is given as follows: \begin{align} A_{n, t} = \begin{cases} 1 & if there exists\ \, t' \le t, \ such that\ \, U_n > \theta_{n, t'} \\ 0 & otherwise, \end{cases} \end{align} where $\theta_{n, t'} \in \Rb^q$. That is, unit $n$ receives treatment $A_{n,t} = 1$ for time period $t$ if there existed a time point $t' \le t$ such that $U_n$ is less than the threshold $\theta_{n, t'}$, which is both unit and time specific. Here assignment of intervention $1$ is an absorbing state. It is easy to see that such an assignment scheme leads to a staggered adoption observation pattern. Here, we can take \begin{align} h_{n, t}(U_n) = \mathbbm{1}(U_n > \bar{\theta}_{n, t}), \end{align} where $\bar{\theta}_{n, t} = \min\{\theta_{n, 1}, \ldots, \theta_{n, t}\}$.

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\center\normalfont\bf}{Identification and Estimation of Treatment Effects}

{\bf Causal parameters of interest.} We restrict our attention to the binary treatment setting where for all $t \in [T]$, we let $\Ac = \{0, 1\}$. Our analysis easily extends for any finite $\Ac$. We focus on the estimation of the average treatment effect for a given measurement $t^*$, and for a subset of units $\Mc \subset [N]$, with $|\Mc| = M$:

align[align omitted — 159 chars of source]

where expectations are taken over the distribution $\varepsilon^{(a)}_{n, t^*}$. We note our target causal parameter is defined conditional on the unobserved confounders $\bU$.

Let

align[align omitted — 81 chars of source]

That is, $\Ic^{(a)}$ is the set of units that received intervention $a$ for measurement $t^*$, and $N_a$ is the number of units in that set. For different subsets $\Mc$, $\mathsf{ATE}_{\Mc}$ nests a variety of causal parameters of interest:

itemize• If $\Mc = \Ic^{(1)}$ (all the units that were treated during measurement $t^*$), then $\mathsf{ATE}_{\Mc}$ corresponds to the {\em average treatment effect on the treated}, which we denote as $\mathsf{ATT}$. • If $\Mc = \Ic^{(0)}$ (all the units that were untreated during measurement $t^*$), then $\mathsf{ATE}_{\Mc}$ corresponds to the {\em average treatment effect on the untreated}, which we denote as $\mathsf{ATU}$. • If $\Mc = [N]$ , then $\mathsf{ATE}_{\Mc}$ corresponds to the {\em average treatment effect}, which we denote as $\mathsf{ATE}$.

For concreteness, we focus our estimation results on $\mathsf{ATT}, \mathsf{ATU}$, and $\mathsf{ATE}$. However, our results easily extend to any set of units $\Mc \subset [N]$.

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Identification} We now show how the model for treatment assignment and potential outcomes, summarized in Section (ref) leads to a novel identification argument for $\mathsf{ATE}_{\Mc}$. Motivated by Proposition (ref), we define the linear factor model approximation error to $f_{t, a}(U_n)$ as follows.

definition[Linear factor model approximation] For $r \in \Nb$, let $\{\lambda_n \}_{n \in [N]} \cup \{ \rho_{t, a} \}_{t \in [T], a \in \Ac}$, with $\lambda_n, \rho_{t, a} \in \Rb^r,$ be (one of) the linear factor model approximations of $f_{t, a}(U_n)$ that minimizes $\Delta_E$ where, \begin{align} \Delta_E = \max_{n \in [N], t \in [T], a \in \Ac}| \eta^{(a)}_{n, t} |, \quad and \eta^{(a)}_{n, t} = f_{t, a}(U_n) - \left\langle \lambda_{n}, \rho_{t, a} \right\rangle. \end{align}

Recall that if $f_{t, a}$ is H\"older continuous, then Proposition (ref) implies that both $r$ and $\Delta_E$ can be simultaneously controlled.

We define two subsets of $\Mc$: for $a \in \{0, 1\}$,

align[align omitted — 83 chars of source]

Note that $\Mc^{(a)} \subset \Ic^{(a)}$. We are now equipped to define the key assumption we require for identification of the causal parameter of interest.

assumption[Linear span inclusion] For $a \in \{0, 1\}$, let $\lambda_{\Mc^{(1 - a)}} = \sum_{n \in \Mc^{(1 - a)}} \lambda_n$. We assume there exists linear weights $\beta^{(a)} \in \Rb^{N_{a}}$ such that, \begin{align} \lambda_{\Mc^{(1 - a)}} = \sum_{n \in \Ic^{(a)}} \beta^{(a)}_n \lambda_n. \end{align} That is, $\lambda_{\Mc^{(1 - a)}}$ lies in the linear span of $\{\lambda_n\}_{n \in \Ic^{(a)}}$. In settings where there are multiple weights that satisfy condition (ref), we define $\beta^{(a)}$ to be the unique one with minimum $\ell_2$-norm.

This assumption implicitly adds a restriction on the treatment assignment. For example, Assumption (ref) does not allow for a treatment assignment mechanism such that the latent factors associated with the units in $\Ic^{(a)}$ and $\Mc^{(1 - a)}$ live in orthogonal spaces. Hence the assignment mechanism needs to be diverse enough, so that the latent factors associated with the units in different treatments are linearly expressible in terms of each other. We only require the weaker condition that this linear span inclusion holds for the {\em sum} of the unit latent factors associated with $\Mc^{(1 - a)}$, as opposed to it holding for {\em each} latent factor $\lambda_n$ for $n \in \Mc^{(1 - a)}$.

theorem[Identification] Let Assumptions (ref), (ref) and (ref) hold. Then, given $\beta^{(a)}$ for $a \in \{0, 1\}$, \begin{align} \sum_{n \in \Mc} \mathbb{E}[Y^{(a)}_{n, t^*} \mid \bU] &= \sum_{n \in \Mc^{(a)}} \mathbb{E}[Y_{n, t^*} \mid \bA, \bU] + \sum_{n \in\Ic^{(a)}} \beta^{(a)}_n \mathbb{E}\left[Y_{n, t^*} \mid \bA, \bU\right] - \sum_{n \in \Ic^{(a)}} \beta^{(a)}_n \eta^{(a)}_{n, t^*} + \sum_{n \in \Mc^{(1-a)}} \eta^{(a)}_{n, t^*},\quad \end{align} where expectations are taken over the distribution of $\varepsilon^{(a)}_{n, t^*}$.

We note that an explicit representation of $\beta^{(a)}$ in terms of the observed data is given in (ref) below, and we establish $\widehat{\beta}^{(a)}$ as given in (ref) is a consistent estimator for it in Proposition (ref). Next, we establish how Theorem (ref) helps establish identification of our causal parameter of interest.

corollary[Identification] Let \begin{align} \mathsf{Observed}_{a} = \frac{1}{M}\left(\sum_{n \in \Mc^{(a)}} \mathbb{E}[Y_{n, t^*} \mid \bA, \bU] + \sum_{n \in \Ic^{(a)}} \beta^{(a)}_n \mathbb{E}\left[ Y_{{n, t^*}} \mid \bA, \bU \right]\right), \end{align} and \begin{align} \mathsf{Observed} = \mathsf{Observed}_{1} - \mathsf{Observed}_{0}. \end{align} Then, under the conditions of Theorem (ref), \begin{align} \Big| \mathbb{E}[\mathsf{ATE}_{\mathcal{M}} \mid \bU] - \mathsf{Observed} \Big| \le \Delta_E \left(1 + \frac{\| \beta^{(0)}\|_1 + \| \beta^{(1)}\|_1}{M} \right). \end{align}

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Estimator} The identification result in Corollary (ref) suggests an estimator of the form

align[align omitted — 321 chars of source]

where $\widehat{\beta}^{(1)}_j$ and $\widehat{\beta}^{(0)}_j$ are estimates of $\beta^{(1)}_j$ and $\beta^{(0)}_j$, respectively. That is, $\widehat{\mathsf{ATE}}_{\Mc}$ imputes the sums of the unobserved potential outcomes with and without treatment by $\sum_{n \in \Ic^{(1)}} \widehat{\beta}^{(1)}_n Y_{n, t^*}$ and $\sum_{n \in \Ic^{(0)}} \widehat{\beta}^{(0)}_n Y_{n, t^*}$, respectively.

Below, we provide sufficient conditions on any estimator $\widehat{\beta}^{(a)}$ of $\beta^{(a)}$, which establish the finite-sample consistency of $\widehat{\mathsf{ATE}}_{\Mc}$. Hence, we denote

align[align omitted — 72 chars of source]

In Section (ref), we provide explicit conditions for consistency and normality when the estimator for $\widehat{\beta}^{(a)}$ is PCR.

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Finite-Sample Consistency} To establish consistency, we make the (mild) assumption that $\varepsilon^{(a)}_{n, t}$ has a sub-Gaussian distribution.

assumption[Sub-Gaussian potential outcomes] For all $(n, t, a)$, $\varepsilon^{(a)}_{n, t} \mid \bU$ is a sub-Gaussian random variable with standard deviation $\sigma^{(a)}_{n, t}$. Let $\sigma_\max = \max_{n \in [N], t \in [T], a \in \{0, 1\}} \sigma^{(a)}_{n, t}$, and assume $\sigma_\max < C$.
proposition[Conditions for consistency for any linear estimator] Let Assumptions (ref), (ref), (ref), and (ref) hold. Let $Y_{\Ic^{(a)}} = [Y_{n, t^*}]_{n \in \Ic^{(a)}}, \varepsilon_{\Ic^{(a)}} \coloneqq [\varepsilon^{(a)}_{n, t^*}]_{n \in \Ic^{(a)}}$. Then, \begin{align} \widehat{\mathsf{ATE}}_{\Mc} &- \mathbb{E}[\mathsf{ATE}_{\Mc} \mid \bU] \le Bias + Variance \end{align} where \begin{align} Bias &= \Delta_E \left(1 + \frac{\| \beta^{(0)}\|_1 + \| \beta^{(1)}\|_1}{M} \right) + O_p\left( \sum_{a \in \{0, 1\}} \frac{\sigma_\max \|\Delta_{\beta^{(a)}}\|_2 + \left\langle \Delta_{\beta^{(a)}} , \mathbb{E}\left[ Y_{\Ic^{(a)}} \right] \right\rangle}{M} \right) \\ Variance &= O_p\left( \sum_{a \in \{0, 1\}} \frac{\sigma_\max \left(\sqrt{M_a} + \|\beta^{(a)} \|_2 \right)}{M} \right) \end{align}

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\center\normalfont\bf}{Estimation results using Principal Component Regression (PCR)}

In Section (ref) below, we provide explicit bounds on $\Delta_{\beta^{(a)}}$ for the case when PCR is used to estimate the coefficients $\widehat{\beta}^{(a)}$. Theorem (ref) collects sufficient conditions for consistency of $\widehat{\mathsf{ATE}}_{\Mc}$.

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Estimating Linear Weights via PCR} We now show how to estimate $\widehat{\beta}^{(a)}$ via PCR, and subsequently control $\Delta_{\beta^{(a)}}$ for the case when there are many measurements of all units under a common set of interventions.

{\bf Necessary notation to define PCR.} We introduce additional notation that we will use to discuss PCR estimation of $\widehat{\beta}^{(a)}$. Let $\bar{\mathcal{T}} \subset [T]$ be defined as follows:

align[align omitted — 133 chars of source]

That is, $\bar{\mathcal{T}}$ is the of measurements for which all units are seen under the same intervention. Let $a_t$ be the common treatment value for $t \in \bar{\mathcal{T}}$. For $a \in \{0, 1\}$, define

align[align omitted — 502 chars of source]

where to reduce notational burden we suppress dependence on $a$ in the notation for $\boldsymbol{Y}$, $\bZ$, and $\bX$. $\boldsymbol{Y}$ is a vector of summed outcomes of the units in $\Mc^{(1 - a)}$ for the measurements in $\bar{\mathcal{T}}$, $\bZ$ is a matrix of outcomes for the units in $\Ic^{(a)}$, and measurements in $\bar{\mathcal{T}}$, and $\bX$ is defined analogously to $\bZ$, but with respect to the expected observed outcomes. $\bX^{\text{lr}}$ is the low-rank approximation of $\bX$; note $\bX - \bX^{\text{lr}} = [\eta^{(a_t)}_{n, t}]_{t \in \bar{\mathcal{T}}, n \in \Ic^{(a)}}$.

{\bf PCR estimator for $\widehat{\beta}^{(a)}$.} Define the singular value decomposition (SVD) of $\bZ$ as

align[align omitted — 101 chars of source]

where $\hat{s}_\ell, \hat{u}_\ell, \hat{v}_\ell$ refer to the $\ell$-th singular value, left singular vector, and right singular vector, respectively. For any SVD, we order the singular values by decreasing magnitude. Given hyper-parameter $k \in [\min(\bar{T}, N_a)]$, we define $\widehat{\bX}^{\text{lr}}$ as follows:

align[align omitted — 126 chars of source]

That is, $\widehat{\bX}^{\text{lr}}$ is a low-rank approximation of $\bZ$. $\widehat{\beta}^{(a)}$ is then estimated by simply doing ordinary least squares (OLS) on $\boldsymbol{Y}$ and $\widehat{\bX}^{\text{lr}}$ as follows:

align[align omitted — 108 chars of source]

Here $\Big(\widehat{\bX}^{\text{lr}})^{+}$ denotes the Moore-Penrose pseudoinverse of $\widehat{\bX}^{\text{lr}}$ defined as

align[align omitted — 137 chars of source]

That is, PCR can be seen as doing ordinary least squares (OLS) on the best $k$-rank approximation of $\bZ$, given by $\Big(\widehat{\bX}^{\text{lr}})^{+}.$ If the ordinary least squares problem has multiple solutions, it is well-known that $\widehat{\beta}^{(a)}$ in equation (ref) is the minimum $\ell_2$-norm solution. We can then use $\widehat{\beta}^{(a)}$, to estimate $\widehat{\mathsf{ATE}}_{\Mc}$ as shown in (ref).

{\bf Interpreting PCR.} Using Assumption (ref) and Definition (ref), we have that for all $n \in [N], t \in \bar{\mathcal{T}}$,

align[align omitted — 153 chars of source]

Hence, using (ref) and the definitions of $\boldsymbol{Y}, \bZ, \bX, \bX^{\text{lr}}$, we have

align[align omitted — 459 chars of source]

By Assumption (ref), we have $\lambda_{\Mc^{(1 - a)}} = \sum_{n \in \Ic^{(a)}} \beta^{(a)}_{n} \lambda_n$. Hence, we can write

align[align omitted — 232 chars of source]

where $\phi^{\text{lr}} = \left[\sum_{n \in \Mc^{(1 - a)}} \eta^{(a_t)}_{n, t} \right]_{t \in \bar{\mathcal{T}}} \ $, $\bar{\varepsilon} = \left[\sum_{n \in \Mc^{(1 - a)}} \varepsilon^{(a_t)}_{n, t} \right]_{t \in \bar{\mathcal{T}}} \ $, $\bE^{\text{lr}} = [\eta^{(a_t)}_{n, t}]_{t \in \bar{\mathcal{T}}, n \in \Ic^{(a)}} \ $, $\bH = [\varepsilon^{(a_t)}_{n, t}]_{t \in \bar{\mathcal{T}}, n \in \Ic^{(a)}}$. Thus we have reduced our problem of estimating $\beta^{(a)}$ to that of linear regression where: (i) the covariates are noisily observed (i.e., error-in-variables regression), i.e. (ref) holds; (ii) the noiseless covariate matrix has an approximate low-rank representation, i.e. (ref) holds; (iii) an approximate linear model holds between the approximate low-rank representation of the noiseless covariates and the response variable, i.e. (ref) holds.

Hence, we can interpret PCR as follows: (1) the first step of doing PCA given in (ref) creates an estimate of the approximate low-rank approximation $\bX^{\text{lr}}$; (2) the second step of doing OLS given in (ref) creates an estimate of $\beta^{(a)}$ by regressing $\boldsymbol{Y}$ on $\widehat{\bX}^{\text{lr}}$, which is motivated by (ref).

The novel technical challenge in analyzing this setting is that there are four sources of error: the noise on the covariates given by $\bH$, the low-rank approximation error given by $\bE^{\text{lr}}$, the linear model approximation error given by $\phi^{\text{lr}}$, and the error on the response given by $\bar{\varepsilon}$.

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Additional Assumptions for Estimation Results with PCR} We make the following additional assumptions to state our consistency results. Note by Assumption (ref) and Proposition (ref), for all $\delta > 0$, \[ \text{rank}(\bX^{\text{lr}}) := \bar{r} \le r \le \bar{C} \cdot \delta^{-q}, \quad \|\bX - \bX^{\text{lr}} \|_{\infty} = \Delta_E \le C_H \cdot \delta^S. \]

remarkPreviewing our consistency results, we will pick $\delta = \left( \frac{1}{(\min(N_0, N_1, \bar{T})}\right)^{\frac{1}{2S}}$. Then, $r \le \bar{C} \min(N_0, N_1, \bar{T})^{\frac{q}{2S}}$ and $\Delta_E \le C_H \left(\frac{1}{(\min(N_0, N_1, \bar{T})}\right)^{\frac{1}{2}}.$ Hence, if $q < 2S$, then as $\min(N_0, N_1, \bar{T}) \to \infty$, $r \ll \min(N_0, N_1, \bar{T})$ and $\Delta_E = o(1)$.
assumption[Well-balanced spectra.] Given the SVD of $\bX^{\text{lr}} = \sum^{\bar{r}}_{\ell = 1} s_\ell u_{\ell} v^T_{\ell},$ we assume \begin{align} s_{\bar{r}} \ge C \sqrt{\frac{ \bar{T} N_a }{\bar{r}}}. \end{align}

An interpretation of Assumption (ref) is as follows. Suppose that each entry of $\bX^{\text{lr}} \ge c > 0$, i.e. is bounded below by an absolute constant $c$. Then since $\bX^{\text{lr}} \in \Rb^{\bar{T} \times N_a}$ we have that $\sum^{\bar{r}}_{\ell = 1} s^2_\ell = \| \bX^{\text{lr}} \|^2_F \ge C \bar{T} N_a$. If all the singular values of $\bX^{\text{lr}}$ are of the same order of magnitude, i.e., $\frac{s_{\bar{r}}}{s_1} \ge C$, this immediately implies that $s_r \ge C \sqrt{\frac{ \bar{T} N_a }{\bar{r}}}$.

assumption[Subspace inclusion.] For intervention $a \in \{0, 1\}$, $\Big[\left\langle \lambda_n, \rho_{t^*, a} \right\rangle\Big]_{n \in \Ic^{(a)}}$ lies in the rowspace of $\bX^{\text{lr}} = \Big[\left\langle \lambda_{n}, \rho_{t, a_t} \right\rangle \Big]_{t \in \bar{\mathcal{T}}, n \in \Ic^{(a)}}$.

Note a sufficient condition for Assumption (ref) is for $a \in \{0, 1\}$

align[align omitted — 87 chars of source]

Hence, intuitively, we require that the target measurement $t^*$ for which we wish to compute $\mathsf{ATE}_{\Mc}$, the latent factors $\rho_{t^*, 0}, \rho_{t^*, 1}$, are linearly expressible in terms of the latent factors $\{\rho_{t, a_t} \}_{t \in \bar{\mathcal{T}}}$ corresponding to the measurements under which all units are under a common intervention. This is the key condition that lets us {\em generalize} from the measurements $\bar{T}$ we learn on to the measurement $t^*$ we make counterfactual predictions on.

Below, we provide exact conditions if the linear estimator is PCR, the appropriateness of which was motivated in Section (ref).

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Finite-sample Consistency using PCR}

theorem[ATT, ATU, ATE consistency using PCR] Let Assumptions (ref), (ref) , (ref), (ref), (ref), and (ref) hold. Let $\widehat{\beta}^{(a)}$ be estimated via PCR as in (ref) and (ref). Assume the following additional conditions hold. \begin{enumerate} • Correct rank estimation for PCR: $k$ in (ref) is such that $k = \bar{r}$. • Smooth outcome model: Let $\alpha = S / q > 0$, where recall $S$ is smoothness parameter of $f_{t, a}(U_n)$ and $q$ is the latent dimension of $U_n$. Assume $\alpha > 1.$ • Disperse weights: For $a \in \{0, 1\}$, assume $\|\beta^{(a)}\|_2 = O\left( \frac{M_{1 - a}}{N^{w}_a} \right)$, where $\frac{1}{2\alpha} < w \le \frac{1}{2}$. • Growing common measurements, units: $\min(N_0, N_1, \bar{T}) \to \infty.$ \end{enumerate} Then we have the following consistency results: \begin{itemize} • {\bf $\mathsf{ATT}$ consistency:} If $\frac{N_0^{1 - w}}{\bar{T}^{1 - \frac{1}{2 \alpha}}} = o(1)$, we have, \begin{align} \widehat{\mathsf{ATT}} - \mathbb{E}[\mathsf{ATT} \mid \bU] &= o_p(1). \end{align} Further, we can take \begin{align} r \le \bar{C} \min(N_0, \bar{T})^{\frac{1}{2\alpha}}, \quad \Delta_E \le C \min(N_0, \bar{T})^{- \frac{1}{2}}. \end{align} • {\bf $\mathsf{ATU}$ consistency:} If $\frac{N_1^{1 - w}}{\bar{T}^{1 - \frac{1}{2 \alpha}}} = o(1)$, we have, \begin{align} \widehat{\mathsf{ATU}} - \mathbb{E}[\mathsf{ATU} \mid \bU] &= o_p(1). \end{align} Further, we can take \begin{align} r \le \bar{C} \min(N_1, \bar{T})^{\frac{1}{2\alpha}}, \quad \Delta_E \le C \min(N_1, \bar{T})^{- \frac{1}{2}}. \end{align} • {\bf $\mathsf{ATE}$ consistency:} If $N_0, N_1 = \Theta(N)$, and $\frac{N^{1 - w}}{\bar{T}^{1 - \frac{1}{2 \alpha}}} = o(1)$, we have, \begin{align} \widehat{\mathsf{ATE}} - \mathbb{E}[\mathsf{ATE} \mid \bU] &= o_p(1). \end{align} Further, we can take \begin{align} r \le \bar{C} \min(N, \bar{T})^{\frac{1}{2\alpha}}, \quad \Delta_E \le C \min(N, \bar{T})^{- \frac{1}{2}}. \end{align} \end{itemize}

Theorem (ref) establishes exact conditions such that PCR is a consistent estimate for $\widehat{\mathsf{ATT}}, \widehat{\mathsf{ATU}}$, and $\widehat{\mathsf{ATE}}$, which are quantified by: the smoothness of $f_{t, a}$; the dimension of $U_n$; the number of common measurements $\bar{T}$; the number of units undergoing interventions $a \in \{0, 1\}$ given by $N_0, N_1$; and the magnitude of the linear weights $\beta^{(a)}$. We recall from Section (ref) that if $f_{t, a}$ is an analytic function, then we can take $S$ to be an arbitrarily large integer, i.e., it is H\"older continuous for all $S \in \Nb$.

Below we provide a natural sufficient condition for which the disperse weights condition holds.

propositionAssume for every set $\Ic \subset [N]$ where $|\Ic| = N^{\theta}$, with $0 < \theta < 1 - \frac{3}{2\alpha}$, there exists a subset $\tilde{\Ic}$, where $|\tilde{\Ic}| = r$ and $\text{span}\{ \lambda_n\}_{j \in \tilde{\Ic}} = \Rb^r$. Assume $r \le \bar{C}\min(N_a, \bar{T})^{\frac{1}{2\alpha}}$. Then the minimum $\ell_2$-norm $\beta^{(a)}$ is such that $\|\beta^{(a)}\|_2 =o\left(\frac{M_{1 - a}}{N_{a}^{{\frac{1}{2\alpha}}}}\right)$. That is, the property, $\frac{1}{2 \alpha} < w$, in Condition 3 of Theorem (ref) holds.

Proposition (ref) establishes that if for any given set of units of sufficient size, their associated latent factors space the entire space $\Rb^r$, then the disperse weights condition must hold.

appendix\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\center\normalfont\bf}{ATE: Identification Proofs} \@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Proof of Theorem (ref)} From Assumption (ref), we have that for $a \in \{0, 1\},$ \begin{align} \sum_{n \in \Mc} \mathbb{E}[Y^{(a)}_{n, t^*} \mid U_n] &= \sum_{n \in \Mc^{(a)}} \mathbb{E}[Y^{(a)}_{n, t^*} \mid U_n] + \sum_{n \in \Mc^{(1 - a)}} \mathbb{E}[Y^{(a)}_{n, t^*} \mid U_n]\\ &= \sum_{n \in \Mc^{(a)}} \mathbb{E}[Y^{(a)}_{n, t^*} \mid \bA, \bU] + \sum_{n \in \Mc^{(1 - a)}} \mathbb{E}[Y^{(a)}_{n, t^*} \mid \bU] \&= \sum_{n \in \Mc^{(a)}} \mathbb{E}[Y_{n, t^*} \mid \bA, \bU] + \sum_{n \in \Mc^{(1 - a)}} \mathbb{E}[Y^{(a)}_{n, t^*} \mid \bU]. \end{align} What remains to be tackled is the second term in (ref). Assumptions (ref) and (ref), and Definition (ref) implies $Y^{(a)}_{n, t^*} = \left\langle \lambda_n, \rho_{t^*, a} \right\rangle + \eta^{(a)}_{n, t^*} + \varepsilon^{(a)}_{n, t^*}$, where $\mathbb{E}[\varepsilon^{(a)}_{n, t^*} \mid \bU] = 0$. Hence, from Assumption (ref), \begin{align} \sum_{n \in \Mc^{(1 - a)}} \mathbb{E}[Y^{(a)}_{n, t^*} \mid \bU] &= \sum_{n \in \Mc^{(1 - a)}} \left(\left\langle \lambda_n, \rho_{t^*, a} \right\rangle + \eta^{(a)}_{n, t^*}\right) \nonumber\\ &= \left\langle \lambda_{\Mc^{(1 - a)}}, \rho_{t^*, a} \right\rangle + \sum_{n \in \Mc^{(1-a)}} \eta^{(a)}_{n, t^*}\nonumber\\ &= \sum_{n \in \Ic^{(a)}} \beta^{(a)}_n \left\langle \lambda_n, \rho_{t^*, a} \right\rangle + \sum_{n \in \Mc^{(1-a)}} \eta^{(a)}_{n, t^*} \end{align} where in the last line we have used the definition of $\lambda_{\Mc^{(1 - a)}}$. In addition, \begin{align} \sum_{n \in \Ic^{(a)}} \beta^{(a)}_n \left\langle \lambda_n, \rho_{t^*, a} \right\rangle &= \sum_{n \in\Ic^{(a)}} \beta^{(a)}_n \mathbb{E}\left[Y^{(a)}_{n, t^*} \mid \bU\right] - \sum_{n \in \Ic^{(a)}} \beta^{(a)}_n \eta^{(a)}_{n, t^*}\nonumber\\ &=\sum_{n \in\Ic^{(a)}} \beta^{(a)}_n \mathbb{E}\left[Y_{n, t^*} \mid \bA, \bU\right] - \sum_{n \in \Ic^{(a)}} \beta^{(a)}_n \eta^{(a)}_{n, t^*}. \end{align} Combining (ref), (ref), and (ref), we conclude the proof. \@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Proof of Corollary (ref).} Using Theorem (ref), we have \begin{align} \Big| \mathbb{E}[\mathsf{ATE}_{\mathcal{M}} \mid \bU] - \mathsf{Observed} \Big| &\le \frac{1}{M} \left(\left| \sum_{j \in \Mc^{(0)}} \eta^{(1)}_{n, t^*}\right| + \left| \sum_{j \in \Ic^{(1)}} \beta^{(1)}_j \eta^{(1)}_{n, t^*}\right| + \left| \sum_{j \in \Mc^{(1)}} \eta^{(0)}_{n, t^*}\right| + \left| \sum_{j \in \Ic^{(0)}} \beta^{(0)}_j \eta^{(0)}_{n, t^*}\right|\right) \&\le \Delta_E \left(1 + \frac{\| \beta^{(0)}\|_1 + \| \beta^{(1)}\|_1}{M} \right). \end{align} \@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\center\normalfont\bf}{ATE: Estimation Proofs} \@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Proof of Proposition (ref).} We recall notation required for the proofs of this section. Let $\varepsilon_{\Ic^{(a)}} \coloneqq [\varepsilon^{(a)}_{n, t^*}]_{n \in \Ic^{(a)}}$, $Y_{\Ic^{(a)}} \coloneqq [Y_{n, t^*}]_{n \in \Ic^{(a)}}$. From Corollary (ref) and the definition of a linear estimator in (ref), we have that \begin{align} |&\widehat{\mathsf{ATE}}_{\Mc}-\mathbb{E}\left[\mathsf{ATE}_{\Mc} \mid \bU \right] | \\ &\le \Delta_E \left(1 + \frac{\| \beta^{(0)}\|_1 + \| \beta^{(1)}\|_1}{M} \right)\\ &+ \sum_{a \in \{0, 1\}} \left| \frac{1}{M}\left(\sum_{n \in \Mc^{(a)}} \mathbb{E}[Y_{n, t^*}\mid \bA, \bU] + \sum_{n \in \Ic^{(a)}} \beta^{(a)}_n \mathbb{E}\left[ Y_{n, t^*} \mid \bA, \bU \right]\right) - \frac{1}{M} \left(\sum_{n \in \Ic^{(a)}} Y_{n, t^*} + \sum_{n \in \Ic^{(a)}} \widehat{\beta}^{(a)}_j Y_{n, t^*} \right) \right|. \end{align} From (ref), it suffices to bound the following terms for $a \in \{0, 1\}$, \begin{align} \left| \frac{1}{M} \left(\sum_{n \in \Mc^{(a)}} Y_{n, t^*} \right) - \frac{1}{M}\left(\sum_{n \in \Mc^{(a)}} \mathbb{E}[Y_{n, t^*} \mid \bA, \bU ] \right) \right|, \\ \left| \frac{1}{M} \left(\sum_{n \in \Ic^{(a)}} \widehat{\beta}^{(a)}_n Y_{n, t^*} \right) - \frac{1}{M}\left(\sum_{n \in \Ic^{(a)}} \beta^{(a)}_n \mathbb{E}\left[ Y_{n, t^*} \mid \bA, \bU \right]\right) \right|. \end{align} For (ref), using Assumptions (ref), \begin{align} \frac{1}{M} \left(\sum_{n \in \Mc^{(a)}} Y_{n, t^*} \right) - \frac{1}{M}\left(\sum_{n \in \Mc^{(a)}} \mathbb{E}[Y_{n, t^*} \mid \bA, \bU] \right) &= \frac{1}{M} \left(\sum_{n \in \Mc^{(a)}} \varepsilon^{(a)}_{n, t^*} \right) \end{align} For (ref), Using the definition of $\Delta_{\beta^{(a)}}$ we have \begin{align} &\frac{1}{M} \left(\sum_{n \in \Ic^{(a)}} \widehat{\beta}^{(a)}_n Y_{n, t^*} \right) - \frac{1}{M}\left(\sum_{n \in \Ic^{(a)}} \beta^{(a)}_n \mathbb{E}\left[ Y_{n, t^*} \mid \bA, \bU \right]\right) \&= \frac{1}{M}\left( \left\langle \beta^{(a)}, \ \varepsilon_{\Ic^{(a)}} \right\rangle + \left\langle \Delta_{\beta^{(a)}}, \ \varepsilon_{\Ic^{(a)}} \right\rangle + \left\langle \Delta_{\beta^{(a)}}, \ \mathbb{E}\left[ Y_{\Ic^{(a)}} \mid \bA, \bU \right] \right\rangle \right) \end{align} From (ref), (ref), (ref) we can write \begin{align} |&\widehat{\mathsf{ATE}}_{\Mc}-\mathbb{E}\left[\mathsf{ATE}_{\Mc} \mid \bU \right] | \le Bias + Variance \end{align} where \begin{align} Bias &= \Delta_E \left(1 + \frac{\| \beta^{(0)}\|_1 + \| \beta^{(1)}\|_1}{M} \right) + \sum_{a \in \{0, 1\}} \frac{1}{M} \left( \left\langle \Delta_{\beta^{(a)}}, \ \varepsilon_{\Ic^{(a)}} \right\rangle + \left\langle \Delta_{\beta^{(a)}}, \ \mathbb{E}\left[ Y_{\Ic^{(a)}} \mid \bA, \bU \right] \right\rangle \right) \\ Variance &= \sum_{a \in \{0, 1\}} \frac{1}{M} \left(\sum_{n \in \Mc^{(a)}} \varepsilon^{(a)}_{n, t^*} \right) +\frac{1}{M}\left\langle \beta^{(a)}, \ \varepsilon_{\Ic^{(a)}} \right\rangle \end{align} We now further bound the $\textsf{Bias}$ and $\textsf{Variance}$ terms. {\em Bounding $\textsf{Variance}$.} We consider the two terms in $\textsf{Variance}$ separately. To bound these two terms, we apply Hoeffding's inequality, which we restate next. \begin{lemma}[Hoeffding's inequality, e.g., vershynin_2018, vershynin_2018] Let $X_1, \dots, X_N$ be independent mean zero sub-Gaussian random variables, and $a = (a_1, \dots, a_N) \in \Rb^N$. Then, for every $t \ge 0,$ we have \begin{align} \mathbb{P}\left( \left| \sum^N_{n = 1} a_n X_n \right| \ge t \right) \le 2 \exp\left( - \frac{Ct^2}{K^2 \|a \|_2^2}\right) \end{align} where $K = \max \{\| X_n \|_{\psi_2}\}_{n=1}^N$. \end{lemma} Using Assumptions (ref), and (ref), and applying Hoeffding's inequality from Lemma (ref) (with $X_n = \varepsilon^{(a)}_{n, t^*}$, $a_n = 1$, $K = \sigma_{\max}, t = \sigma_{\max}\sqrt{M_a}$), we have that (ref) is bounded by \begin{align} \frac{1}{M} \left(\sum_{n \in \Mc^{(a)}} \varepsilon^{(a)}_{n, t} \right) = O_p\left( \frac{ \sigma_\max \sqrt{M_a} }{M}\right) \end{align} Similarly, we have \begin{align} \frac{1}{M}\left\langle \beta^{(a)}, \ \varepsilon_{\Ic^{(a)}} \right\rangle = O_p\left(\frac{\sigma_\max\|\beta^{(a)} \|_2}{M}\right) \end{align} {\em Bounding $\textsf{Bias}$.} Again, using Assumptions (ref) and (ref), and using Lemma (ref), we have \begin{align} &\frac{1}{M}\left\langle \Delta_{\beta^{(a)}}, \ \varepsilon_{\Ic^{(a)}} \right\rangle = O_p\left( \frac{\sigma_\max\|\Delta_{\beta^{(a)}}\|_2 }{M}\right) \end{align} Collecting terms completes the proof. \@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Proof of Theorem (ref).} \@startsection{subsubsection}{3}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont}{Bounding linear parameter estimation error of PCR.} We first state and prove two key propositions required to establish Theorem (ref) that bound $\Delta_{\beta^{(a)}}$. \begin{proposition} Let the conditions of Theorem (ref), and Assumptions (ref), (ref) hold. Suppose we estimate $\widehat{\beta}^{(a)}$ via PCR (i.e., (ref) and (ref)) and $k = \bar{r}$. Then with probability $1 - O((N_a \bar{T})^{-10})$ \begin{align} \|\Delta_{\beta^{(a)}}\|_2 \le C \cdot \sigma_\max^3 \cdot \ln^3(\bar{T} N_a) \cdot \left[\left\| \beta^{(a)} \right\|_2 \cdot \left( \frac{r}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + r \Delta_E \right) + \frac{M_{1-a} \sqrt{r} \Delta_E}{\sqrt{N_a}} \right]. \end{align} \end{proposition} {\bf Proof of Proposition (ref).} \begin{table}[h] \begin{tabular}{||c c||} \hline Notation of agarwal2021causal & Our Notation \\ [1ex] \hline $\boldsymbol{Y}$ & $\frac{\boldsymbol{Y}}{M_{1-a}}$ \\ [1ex] \hline $\bX$ & $\bX$ \\ [1ex] \hline $\bZ$ & $\bZ$ \\ [1ex] \hline $\bX^{(lr)}$ & $\bX^{lr}$ \\ [1ex] \hline $n$ & $\bar{T}$ \\ \hline $p$ & $N_a$ \\ \hline $\boldsymbol{\beta}^*$ & $\frac{\beta^{(a)}}{M_{1-a}}$ \\ \hline $\hat{\boldsymbol{\beta}}$ & $\frac{\widehat{\beta}^{(a)}}{M_{1-a}}$ \\ \hline $\Delta_E$ & $\Delta_E$ \\ \hline $r$ & $\bar{r} \ (\le r)$ \\ \hline $\phi^{(lr)}$ & $\frac{\phi^{lr}}{M_{1-a}}$ \\ [1ex] \hline $\varepsilon$ & $\frac{\bar{\varepsilon}}{M_{1-a}}$ \\ [1ex] \hline $\Bar{A}$ & $C$ \\ \hline $\Bar{K}$ & 0 \\ [1ex] \hline $K_{a}, \kappa, \Bar{\sigma}$ & $C\sigma_{\max}$ \\ \hline $\rho_{min}$ & 1 \\ \hline \end{tabular} \caption{A summary of the main notational differences between our setting and that of agarwal2021causal.} \end{table} Using (ref), (ref), (ref), and (ref), we have reduced our problem of estimating $\beta^{\Ic_t^{(a)}}$ to that of linear regression where: (i) the covariates are noisily observed (i.e., error-in-variables regression), i.e. (ref) holds; (ii) the noiseless covariate matrix has an approximate low-rank representation, i.e. (ref) holds; (ii) an approximate linear model holds between the approximate low-rank representation of the noiseless covariates and the response variable, i.e. (ref) holds. We observe that bounding $\|\Delta_{\beta^{(a)}}\|_2 $ in such a setting is exactly the setup considered in Proposition E.3 of agarwal2021causal, where they also analyze PCR. We match notation with that of agarwal2021causal as seen in Table (ref). We then get \begin{align} \left\|\frac{\Delta_{\beta^{(a)}}}{M_{1 - a}}\right\|_2 &\le C \cdot (\sigma_\max)(2\sigma_\max) \cdot \sigma_\max \cdot \ln^3(\bar{T} N_a) \cdot \sqrt{r} \cdot \left( \frac{\|\phi^{lr}\|_2}{\sqrt{N_a \bar{T}}} + \sqrt{r} \cdot \left\| \frac{\beta^{(a)}}{M_{1 - a}} \right\|_2 \cdot \left( \frac{1}{\sqrt{\bar{T}}} + \frac{1}{\sqrt{N_a}} + \Delta_E \right)\right) \end{align} Using $\|\phi^{\text{lr}}\|_2 \le \Delta_E \sqrt{\bar{T}}$ and simplifying (ref) completes the proof \begin{proposition} Let the conditions of Proposition (ref) hold. Let $\mathsf{Proj}$ be the projection operator onto the rowspace of $\bX^{\text{lr}}$, i.e., $\mathsf{Proj} = \bV_r \bV_r^T$, where $\bV_r$ are the right singular vectors of $\bX^{\text{lr}}$. Then with probability $1 - O((N_a \bar{T})^{-10})$ \begin{align} &\|\mathsf{Proj}(\Delta_{\beta^{(a)}})\|_2 \ \le \& \quad C \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{ \left\| \beta^{(a)} \right\|_2 \cdot \left( \frac{r^{3/2}}{\min(\bar{T}, N_a)} + \frac{r^{3/2} \Delta_E}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + r^{3/2} \Delta^2_E \right) \right\}, \& + C \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{ \left\| \beta^{(a)} \right\|_1 \cdot \left( \frac{\sqrt{r}}{\left\| \frac{\beta^{(a)}}{M_{1 - a}} \right\|^{1/2}_1 \bar{T}^{\frac{1}{4}} \sqrt{N_a}} + \frac{r}{\min(\bar{T}, N_a)} + \frac{r \Delta_E}{\sqrt{N_a}} \right) \right\}, \& + C \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot M_{1-a} \cdot \left\{\frac{\sqrt{r} \Delta_E}{\sqrt{N_a}} + \frac{r \Delta^2_E}{\sqrt{N_a}} \right\}. \end{align} \end{proposition} {\bf Proof of Proposition (ref).} As in the proof of Proposition (ref), we use (ref), (ref), (ref) and observe that bounding $\|\textsf{Proj}(\Delta_{\beta^{(a)}})\|_2$ in such a setting is exactly the setup considered in Corollary E.1 of agarwal2021causal, where they also analyze PCR. Matching notation with that of agarwal2021causal, \footnote{ The additional notation compared to Proposition (ref) that needs to be matched here is $\bV_r \bV^T_r = \textsf{Proj}(\cdot)$, where $\bV_r \bV^T_r$ is the notation used in agarwal2021causal. } we get \begin{align} \left\|\frac{Proj(\Delta_{\beta^{(a)}})}{M_{1-a}} \right\|_2 &\le C \cdot (\sigma_\max)(2\sigma_\max)^2 \cdot \sigma_\max \cdot \ln^{9/2}(\bar{T} N_a) \cdot \sqrt{r} \cdot \Big[ (A) + (B) + (C)\Big] \end{align} where \begin{align} (A) &\coloneqq \frac{1}{\sqrt{\bar{T}}} \| \phi^{\text{lr}} \|_2 \left( \frac{1}{\sqrt{N_a}} + \frac{\sqrt{r}}{N_a} + \frac{\sqrt{r}}{\sqrt{\bar{T} N_a}} + \frac{\sqrt{r}}{\sqrt{N_a}} \Delta_E \right) \\ (B) &\coloneqq \left\| \frac{\beta^{(a)}}{M_{1 - a}} \right\|_1 \left( \frac{\bar{T}^{1/4}}{\left\| \frac{\beta^{(a)}}{M_{1 - a}} \right\|^{1/2}_1 \sqrt{\bar{T} N_a}} + \frac{\sqrt{r}}{\sqrt{\bar{T} N_a}} + \frac{\sqrt{r}}{N_a}+ \frac{\sqrt{r}}{\sqrt{N_a}} \Delta_E \right) \\ (C) &\coloneqq \left\| \frac{\beta^{(a)}}{M_{1 - a}} \right\|_2 \cdot r \cdot \left( \frac{1}{\bar{T}} + \frac{1}{N_a} + \frac{1}{\sqrt{\bar{T} N_a}} + \left( \frac{1}{\sqrt{\bar{T}}} + \frac{1}{\sqrt{N_a}} \right) \Delta_E + \Delta^2_E \right) \end{align} Using $\|\phi^{\text{lr}}\|_2 \le \sqrt{\bar{T}}\Delta_E$ and $r \le \min(N_a, \bar{T})$, we have \begin{align} (A) &\le \frac{\Delta_E}{\sqrt{N_a}} + \frac{\sqrt{r} \Delta^2_E}{\sqrt{N_a}}. \end{align} Simplifying (B) and (C), we have \begin{align} (B) &\le \left\| \frac{\beta^{(a)}}{M_{1 - a}} \right\|_1 \left( \frac{1}{\left\| \frac{\beta^{(a)}}{M_{1 - a}} \right\|^{1/2}_1 \bar{T}^{\frac{1}{4}} \sqrt{N_a}} + \frac{\sqrt{r}}{\min(N_a, \bar{T})} + \frac{\sqrt{r} \Delta_E}{\sqrt{N_a}} \right) \\ (C) &\le \left\| \frac{\beta^{(a)}}{M_{1 - a}} \right\|_2 \cdot r \cdot \left( \frac{1}{\min(\bar{T}, N_a)} + \frac{\Delta_E}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + \Delta^2_E \right) \end{align} Collecting the various bounds completes this section. \@startsection{subsubsection}{3}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont}{General conditions for ATE consistency.} For simplicity, we suppress the conditioning on $\bA, \bU$ in the remainder of the proof. \begin{proposition} Let the conditions of Proposition (ref) and Assumption (ref) hold. For $a \in \{0, 1\}$, assume $$\|\beta^{(a)}\|_2 = O\left( \frac{M_{1-a}}{N^{w}_a} \right),$$ where $0 \le w \le \frac{1}{2}$. Then, \begin{align} &\widehat{\mathsf{ATE}_{\Mc}} - \mathbb{E}[\mathsf{ATE}_{\Mc} \mid \bU] \& = \sum_{a \in \{0, 1\}} C \cdot \Delta_E \left( \frac{M_{1 - a} \cdot N_a^{0.5 - w}}{M} \right), \& \quad + C \cdot \frac{1}{\sqrt{M}}, \& \quad + \sum_{a \in \{0, 1\}} C \cdot \frac{M_{1-a} \cdot N_a^{- w}}{M}, \& \quad + \sum_{a \in \{0, 1\}}C \cdot \frac{M_{1-a}}{M N^w_a} \cdot \sigma_\max^3 \cdot \ln^3(\bar{T} N_a) \cdot \left[r \left( \frac{1}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + \Delta_E \right) \right], \& \quad + \sum_{a \in \{0, 1\}} C \cdot \frac{M_{1-a}}{M} \cdot N_a^{1/2 - w} \cdot \sigma_\max^3 \cdot \ln^3(\bar{T} N_a) \cdot \left[r \left( \frac{1}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + \Delta_E \right) \right]\Delta_E, \& \quad + \sum_{a \in \{0, 1\}} C \cdot \frac{1}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{ \frac{M_{1-a}}{N^{w - 0.5}_a} \cdot \left( \frac{r^{3/2}}{\min(\bar{T}, N_a)} + \frac{r^{3/2} \Delta_E}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + r^{3/2} \Delta^2_E \right) \right\}, \& \quad + \sum_{a \in \{0, 1\}} C \cdot \frac{\sqrt{N_a} M_{1-a}}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{ \frac{\sqrt{r}}{ \bar{T}^{\frac{1}{4}} N_a^{0.5w + 0.25}} + \frac{ r}{N_a^{w - 0.5} \cdot \min(\bar{T}, N_a)} + \frac{ r \cdot \Delta_E}{N_a^{w}} \right\}, \& \quad +\sum_{a \in \{0, 1\}} C \cdot \frac{ M_{1-a}}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{\sqrt{r}\Delta_E + r \Delta^2_E \right\}. \end{align} \end{proposition} {\bf Proof of Proposition (ref).} From Proposition (ref), we have that \begin{align} &\widehat{\mathsf{ATE}_{\Mc}} - \mathbb{E}[\mathsf{ATE}_{\Mc} \mid \bU] \&\le C \Delta_E \left(1 + \frac{\| \beta^{(0)}\|_1 + \| \beta^{(1)}\|_1}{M} \right) + O_p\left( \sum_{a \in \{0, 1\}} \frac{\sigma_\max \left(\sqrt{M_a} + \|\beta^{(a)} \|_2 + \|\Delta_{\beta^{(a)}}\|_2\right) + \left\langle \Delta_{\beta^{(a)}} , \mathbb{E}\left[ Y_{\Ic^{(a)}} \right] \right\rangle}{M} \right) \end{align} We consider the various terms on the right-hand side above separately. {\em 1. Bounding the $\Delta_E \left(1 + \frac{\| \beta^{(0)}\|_1 + \| \beta^{(1)}\|_1}{M} \right)$ term.} Given the assumption that $ \| \beta^{(a)} \|_2 = O\left( \frac{M_{1-a}}{N^w_a} \right)$, we have $\| \beta^{(a)} \|_1 = O\left( \frac{M_{1-a} N_a^{0.5}}{N_a^w} \right).$ Therefore \begin{align} \Delta_E \left(\frac{\| \beta^{(0)}\|_1 + \| \beta^{(1)}\|_1}{M} \right) &= C \cdot \Delta_E \left( \frac{M_0 \cdot N_1^{0.5 - w} + M_1 \cdot N_0^{0.5 - w}}{M} \right) \&= \sum_{a \in \{0, 1\}} C \cdot \Delta_E \left( \frac{M_{1 - a} \cdot N_a^{0.5 - w}}{M} \right) \end{align} {\em 2. Bounding the $\sum_{a \in \{0, 1\}} \frac{\sigma_\max (\sqrt{M_a} + \|\beta^{(a)} \|_2)}{M}$ term.} Note that since $M_a < M$, \begin{align} \sum_{a \in \{0, 1\}} \frac{\sigma_\max \sqrt{M_a}}{M} = O\left(\frac{1}{\sqrt{M}}\right). \end{align} Further, \begin{align} \sum_{a \in \{0, 1\}} \frac{\|\beta^{(a)} \|_2}{M} = C \cdot \frac{M_0 \cdot N_1^{- w} + M_1 \cdot N_0^{- w}}{M} = \sum_{a \in \{0, 1\}} C \cdot \frac{M_{1-a} \cdot N_a^{- w}}{M} \end{align} {\em 3. Bounding the $\sum_{a \in \{0, 1\}} \frac{\sigma_\max \left\|\Delta_{\beta^{(a)}}\right\|_2}{M}$ term.} Using Proposition (ref) and $ \| \beta^{(a)} \|_2 = O\left( \frac{M_{1-a}}{N^w_a} \right)$, we have \begin{align} &\sum_{a \in \{0, 1\}}\frac{\|\Delta_{\beta^{(a)}}\|_2}{M} \\ &\le \sum_{a \in \{0, 1\}}\frac{1}{M} \cdot C \cdot \sigma_\max^3 \cdot \ln^3(\bar{T} N_a) \cdot \left[\left\| \beta^{(a)} \right\|_2 \cdot \left( \frac{r}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + r \Delta_E \right) + \frac{M_{1-a} \sqrt{r} \Delta_E}{\sqrt{N_a}} \right], \\ &\le \sum_{a \in \{0, 1\}}\frac{M_{1-a}}{M} \cdot C \cdot \sigma_\max^3 \cdot \ln^3(\bar{T} N_a) \cdot \left[\frac{1}{N^w_a} \cdot \left( \frac{r}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + r \Delta_E \right) + \frac{\sqrt{r} \Delta_E}{\sqrt{N_a}} \right], \\ &\le \sum_{a \in \{0, 1\}} C \cdot \frac{M_{1-a}}{M N^w_a} \cdot \sigma_\max^3 \cdot \ln^3(\bar{T} N_a) \cdot \left[r \left( \frac{1}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + \Delta_E \right) \right], \end{align} where in the third inequality we have used that $w \le \frac{1}{2}$. {\em 4. Bounding the $\sum_{a \in \{0, 1\}} \frac{\left\langle \Delta_{\beta^{(a)}} , \mathbb{E}\left[ Y_{\Ic^{(a)}} \right] \right\rangle}{M}$ term.} From Definition (ref), we have that $\mathbb{E}[Y^{(a)}_{j, t^*}] = \left\langle \lambda_j, \rho_{t^*, a} \right\rangle + \eta^{(a)}_{j, t^*}$. Hence \begin{align} \sum_{a \in \{0, 1\}} \left| \left\langle \Delta_{\beta^{(a)}} , \mathbb{E}\left[ Y_{\Ic^{(a)}} \right] \right\rangle \right| &= \sum_{a \in \{0, 1\}} \left| \left\langle \Delta_{\beta^{(a)}} , [\left\langle \lambda_j, \rho_{t^*, a} \right\rangle + \eta^{(a)}_{j, t^*}]_{n \in \Ic^{(a)}} \right\rangle]_{n \in \Ic^{(a)}}} \right| \\ &= \sum_{a \in \{0, 1\}} \left| \left\langle \Delta_{\beta^{(a)}} , [\left\langle \lambda_j, \rho_{t^*, a} \right\rangle]_{n \in \Ic^{(a)}} \right\rangle]_{n \in \Ic^{(a)}}} + \left\langle \Delta_{\beta^{(a)}} , [\eta^{(a)}_{j, t^*}]_{n \in \Ic^{(a)}} \right\rangle \right| \\ &\le \sum_{a \in \{0, 1\}} \left| \left\langle \Delta_{\beta^{(a)}} , [\left\langle \lambda_j, \rho_{t^*, a} \right\rangle]_{n \in \Ic^{(a)}} \right\rangle]_{n \in \Ic^{(a)}}} \right| + \|\Delta_{\beta^{(a)}}\|_2 \| [\eta^{(a)}_{j, t^*}]_{n \in \Ic^{(a)}} \|_2 \\ &\le \sum_{a \in \{0, 1\}} \left| \left\langle \Delta_{\beta^{(a)}} , [\left\langle \lambda_j, \rho_{t^*, a} \right\rangle]_{n \in \Ic^{(a)}} \right\rangle]_{n \in \Ic^{(a)}}} \right| + \|\Delta_{\beta^{(a)}}\|_2 \sqrt{N_a} \Delta_E \end{align} Using Assumption (ref), we have \begin{align} \left| \left\langle \Delta_{\beta^{(a)}} , [\left\langle \lambda_j, \rho_{t^*, a} \right\rangle]_{n \in \Ic^{(a)}} \right\rangle]_{n \in \Ic^{(a)}}} \right| &= \left| \left\langle \Delta_{\beta^{(a)}} , \mathsf{Proj}([\left\langle \lambda_j, \rho_{t^*, a} \right\rangle]_{n \in \Ic^{(a)}} \right\rangle]_{n \in \Ic^{(a)}}}) \right| \&= \left| \left\langle \mathsf{Proj}(\Delta_{\beta^{(a)}}), [\left\langle \lambda_j, \rho_{t^*, a} \right\rangle]_{n \in \Ic^{(a)}} \right\rangle]_{n \in \Ic^{(a)}}} \right| \& \le \| \mathsf{Proj}(\Delta_{\beta^{(a)}}) \|_2 \| [\left\langle \lambda_j, \rho_{t^*, a} \right\rangle]_{n \in \Ic^{(a)}} \|_2 \& \le C \| \mathsf{Proj}(\Delta_{\beta^{(a)}}) \|_2 \sqrt{N_a}, \end{align} where recall $\mathsf{Proj} = \bV_r \bV_r^T$, and $\bV_r$ are the right singular vectors of $\bX^{\text{lr}}$. Hence, we have \begin{align} \frac{1}{M}\sum_{a \in \{0, 1\}} \left| \left\langle \Delta_{\beta^{(a)}} , \mathbb{E}\left[ Y_{\Ic^{(a)}} \right] \right\rangle \right| &\le \sum_{a \in \{0, 1\}} \frac{1}{M}\|\Delta_{\beta^{(a)}}\|_2 \sqrt{N_a} \Delta_E + \frac{C}{M} \| \mathsf{Proj}(\Delta_{\beta^{(a)}}) \|_2 \sqrt{N_a} \end{align} We bound each term above separately. {\em 4a. Bounding the $ \sum_{a \in \{0, 1\}} \frac{1}{M}\|\Delta_{\beta^{(a)}}\|_2 \sqrt{N_a} \Delta_E$ term.} For the first term, by applying a similar logic used to derive (ref), we have that \begin{align} &\sum_{a \in \{0, 1\}} \frac{1}{M}\|\Delta_{\beta^{(a)}}\|_2 \sqrt{N_a} \Delta_E \&\le \sum_{a \in \{0, 1\}}\frac{M_{1-a}}{M N^w_a} \cdot C \cdot \sigma_\max^3 \cdot \ln^3(\bar{T} N_a) \cdot \left[r \left( \frac{1}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + \Delta_E \right) \right]\sqrt{N_a}\Delta_E \&\le \sum_{a \in \{0, 1\}}\frac{M_{1-a}}{M} \cdot N_a^{1/2 - w} \cdot C \cdot \sigma_\max^3 \cdot \ln^3(\bar{T} N_a) \cdot \left[r \left( \frac{1}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + \Delta_E \right) \right]\Delta_E \end{align} {\em 4b. Bounding the $\sum_{a \in \{0, 1\}} \frac{C}{M} \| \mathsf{Proj}(\Delta_{\beta^{(a)}}) \|_2 \sqrt{N_a}$ term.} Using Proposition (ref) and $ \| \beta^{(a)} \|_2 = O\left( \frac{M_{1-a}}{N^w_a} \right)$, we have \begin{align} &\sum_{a \in \{0, 1\}} \frac{C}{M} \| \mathsf{Proj}(\Delta_{\beta^{(a)}}) \|_2 \sqrt{N_a} \& \le \sum_{a \in \{0, 1\}} \quad \frac{C\sqrt{N_a}}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{ \left\| \beta^{(a)} \right\|_2 \cdot \left( \frac{r^{3/2}}{\min(\bar{T}, N_a)} + \frac{r^{3/2} \Delta_E}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + r^{3/2} \Delta^2_E \right) \right\}, \& \quad \quad + \frac{C\sqrt{N_a}}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{ \left\| \beta^{(a)} \right\|_1 \cdot \left( \frac{\sqrt{r}}{\left\| \frac{\beta^{(a)}}{M_{1 - a}} \right\|^{1/2}_1 \bar{T}^{\frac{1}{4}} \sqrt{N_a}} + \frac{r}{\min(\bar{T}, N_a)} + \frac{r \Delta_E}{\sqrt{N_a}} \right) \right\}, \& \quad \quad + \frac{C\sqrt{N_a}}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot M_{1-a} \cdot \left\{\frac{\sqrt{r}\Delta_E}{\sqrt{N_a}} + \frac{r \Delta^2_E}{\sqrt{N_a}} \right\}. \end{align} We bound the three terms on the r.h.s above separately. \begin{enumerate} • First term of (ref). \begin{align} &\sum_{a \in \{0, 1\}} \frac{C\sqrt{N_a}}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{ \left\| \beta^{(a)} \right\|_2 \cdot \left( \frac{r^{3/2}}{\min(\bar{T}, N_a)} + \frac{r^{3/2} \Delta_E}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + r^{3/2} \Delta^2_E \right) \right\} \& \le \sum_{a \in \{0, 1\}} \quad \frac{C}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{ \frac{M_{1-a}}{N^{w - 0.5}_a} \cdot \left( \frac{r^{3/2}}{\min(\bar{T}, N_a)} + \frac{r^{3/2} \Delta_E}{\min(\sqrt{\bar{T}}, \sqrt{N_a})} + r^{3/2} \Delta^2_E \right) \right\}. \end{align} • Second term of (ref). \begin{align} &\sum_{a \in \{0, 1\}} \frac{C\sqrt{N_a}}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{ \left\| \beta^{(a)} \right\|_1 \cdot \left( \frac{\sqrt{r}}{\left\| \frac{\beta^{(a)}}{M_{1 - a}} \right\|^{1/2}_1 \bar{T}^{\frac{1}{4}} \sqrt{N_a}} + \frac{r}{\min(\bar{T}, N_a)} + \frac{r \Delta_E}{\sqrt{N_a}} \right) \right\} \\ &\le \sum_{a \in \{0, 1\}} \frac{C\sqrt{N_a}}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{ \frac{\sqrt{r} \left\| \beta^{(a)} \right\|^{0.5}_1 M_{1 - a}^{0.5}}{ \bar{T}^{\frac{1}{4}} \sqrt{N_a}} + \frac{ \left\| \beta^{(a)} \right\|_1 r}{\min(\bar{T}, N_a)} + \frac{ \left\| \beta^{(a)} \right\|_1 r \Delta_E}{\sqrt{N_a}} \right\}, \\ &\le \sum_{a \in \{0, 1\}} \frac{C\sqrt{N_a}}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{ \frac{\sqrt{r} \cdot M_{1-a} }{ \bar{T}^{\frac{1}{4}} N_a^{0.5w + 0.25}} + \frac{ M_{1-a} \cdot r}{N_a^{w - 0.5} \cdot \min(\bar{T}, N_a)} + \frac{ M_{1-a} \cdot r \cdot \Delta_E}{N_a^{w}} \right\}, \\ &\le \sum_{a \in \{0, 1\}} \frac{C\sqrt{N_a} M_{1-a}}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{ \frac{\sqrt{r} }{ \bar{T}^{\frac{1}{4}} N_a^{0.5w + 0.25}} + \frac{ r}{N_a^{w - 0.5} \cdot \min(\bar{T}, N_a)} + \frac{ r \cdot \Delta_E}{N_a^{w}} \right\}. \end{align} • Third term of (ref). \begin{align} &\sum_{a \in \{0, 1\}} \frac{C\sqrt{N_a}}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot M_{1-a} \cdot \left\{\frac{\sqrt{r}\Delta_E}{\sqrt{N_a}} + \frac{r \Delta^2_E}{\sqrt{N_a}} \right\} \\ &= \sum_{a \in \{0, 1\}} C \cdot \frac{ M_{1-a}}{M} \cdot \sigma_\max^4 \cdot \ln^{9/2}(\bar{T} N_a) \cdot \left\{\sqrt{r}\Delta_E + r \Delta^2_E \right\} \end{align} \end{enumerate} {\em Summarizing all terms.} Using (ref), (ref), (ref), (ref), (ref), (ref), (ref), (ref), we complete the proof of the proposition. \@startsection{subsubsection}{3}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont}{Finishing proof of Theorem (ref).} Recall from Proposition (ref), we have that \begin{align} r \le C \cdot \delta^{-q}, \quad \Delta_E \le C \cdot \delta^S. \end{align} {\bf $\textsf{ATT}$ consistency.} For estimation of $\textsf{ATT}$, we have that $M = M_1 = N_1$ and $M_0 = 0$. Hence, by simplifying the result in Proposition (ref), we get that \begin{align} &\widehat{\mathsf{ATE}_{\Mc}} - \mathbb{E}[\mathsf{ATE}_{\Mc} \mid \bU] \&\le C \cdot \Delta_E \left( N_0^{0.5 - w}\right), \& \quad + C \cdot \frac{1}{\sqrt{M}}, \& \quad + C \cdot N_0^{- w}, \& \quad + C \cdot \frac{1}{N^w_0} \cdot \ln^3(\bar{T} N_0) \cdot \left[r \left( \frac{1}{\min(\sqrt{\bar{T}}, \sqrt{N_0})} + \Delta_E \right) \right], \& \quad + C \cdot N_0^{1/2 - w} \cdot \ln^3(\bar{T} N_0) \cdot \left[r \left( \frac{1}{\min(\sqrt{\bar{T}}, \sqrt{N_0})} + \Delta_E \right) \right]\Delta_E, \& \quad + C \cdot \ln^{9/2}(\bar{T} N_0) \cdot \left\{ \frac{1}{N^{w - 0.5}_0} \cdot \left( \frac{r^{3/2}}{\min(\bar{T}, N_0)} + \frac{r^{3/2} \Delta_E}{\min(\sqrt{\bar{T}}, \sqrt{N_0})} + r^{3/2} \Delta^2_E \right) \right\}, \& \quad + C \cdot \sqrt{N_0} \cdot \ln^{9/2}(\bar{T} N_0) \cdot \left\{ \frac{\sqrt{r}}{ \bar{T}^{\frac{1}{4}} N_0^{0.5w + 0.25}} + \frac{ r}{N_0^{w - 0.5} \cdot \min(\bar{T}, N_0)} + \frac{ r \cdot \Delta_E}{N_0^{w}} \right\}, \& \quad + C \cdot \ln^{9/2}(\bar{T} N_0) \cdot \left\{\sqrt{r}\Delta_E + r \Delta^2_E \right\}. \end{align} We deal with the seven terms above separately. Let $G = \min(N_0, \bar{T})$. For $\gamma > 0$, take $\delta = \left(\frac{1}{G}\right)^{\gamma / q}$. Then (ref) implies \begin{align} r \le C G^\gamma, \quad \Delta_E &\le C \left(\frac{1}{G}\right)^{ \gamma \alpha}. \end{align} We set $\gamma = \frac{1}{2\alpha}$ and so we have $r \le C G^{\frac{1}{2\alpha}}$, and that $\Delta_E \le C G^{-0.5}$. {\em Term (ref)}. \begin{align} &C \cdot \Delta_E \left( N_0^{0.5 - w}\right) \le G^{-\gamma \alpha} N_0^{0.5 - w} = G^{-0.5} N_0^{0.5 - w} = o_p(1) \end{align} where in the last line we have used $\alpha > 1$ and the assumption $\frac{N_0^{1 - w}}{\bar{T}^{1 - \frac{1}{2 \alpha}}} = o(1) \implies \frac{N_0^{0.5 - w}}{\bar{T}^{0.5}} = o(1)$; we also use the assumption that $w > 0$. {\em Term (ref) and (ref)}. Given the assumption that $M (= N_1), N_0 \to \infty$, and that $w > 0$, \begin{align} C \cdot \frac{1}{\sqrt{M}} &= o_p(1), \\ C \cdot N_0^{- w} &= o_p(1). \end{align} {\em Term (ref)}. \begin{align} &C \cdot \frac{1}{N^w_0} \cdot \ln^3(\bar{T} N_0) \cdot \left[r \left( \frac{1}{\min(\sqrt{\bar{T}}, \sqrt{N_0})} + \Delta_E \right) \right] \&\le C \cdot \frac{1}{N^w_0} \cdot \ln^3(\bar{T} N_0) \cdot \left[G^{\gamma} \left( \frac{1}{\min(\sqrt{\bar{T}}, \sqrt{N_0})} + G^{-\gamma \alpha} \right) \right] \&\le C \cdot \ln^3(\bar{T} N_0) \cdot \left[G^{\gamma - w - 0.5} + G^{\gamma (1 - \alpha) - w} \right] \&= C \cdot \ln^3(\bar{T} N_0) \cdot \left[G^{0.5(\frac{1}{\alpha} - 1) - w} \right] \&= o_p(1) \end{align} where in the last line we use the inequality that $w > \frac{1}{2 \alpha} - \frac{1}{2}$. {\em Term (ref)}. \begin{align} &C \cdot N_0^{1/2 - w} \cdot \ln^3(\bar{T} N_0) \cdot \left[r \left( \frac{1}{\min(\sqrt{\bar{T}}, \sqrt{N_0})} + \Delta_E \right) \right]\Delta_E \&\le C \cdot \ln^3(\bar{T} N_0) \cdot \left[G^{\frac{1}{2 \alpha} - 0.5} \right] \cdot G^{-0.5} \cdot N_0^{0.5 - w} \&= o_p(1) \end{align} where in the last line we have used $\alpha > 1$ and the assumption $\frac{N_0^{1 - w}}{\bar{T}^{1 - \frac{1}{2 \alpha}}} = o(1) \implies \frac{N_0^{0.5 - w}}{\bar{T}^{0.5}} = o(1)$; we also use the assumption that $w > 0$. {\em Term (ref)}. \begin{align} &C \cdot \ln^{9/2}(\bar{T} N_0) \cdot \left\{ \frac{1}{N^{w - 0.5}_0} \cdot \left( \frac{r^{3/2}}{\min(\bar{T}, N_0)} + \frac{r^{3/2} \Delta_E}{\min(\sqrt{\bar{T}}, \sqrt{N_0})} + r^{3/2} \Delta^2_E \right) \right\} \&\le C \cdot \ln^{9/2}(\bar{T} N_0) \cdot \left\{ \frac{1}{N^{w - 0.5}_0} \cdot \left( G^{1.5\gamma - 1} + G^{\gamma(1.5 - \alpha) - 0.5} + G^{\gamma(1.5 - 2\alpha)}\right) \right\} \&\le C \cdot \ln^{9/2}(\bar{T} N_0) \cdot \left\{ \frac{ N_0 ^{0.5 - w} }{G^{1 - \frac{0.75}{\alpha}}} \right\} \&= o_p(1) \end{align} where in the last line we have used the fact that $w > \frac{1}{2 \alpha}$, $\alpha > \frac{1}{2}$ and the assumption that $\frac{N_0^{1 - w}}{\bar{T}^{1 - \frac{1}{2 \alpha}}} = o(1)$. {\em Term (ref)}. \begin{align} &C \cdot \sqrt{N_0} \cdot \ln^{9/2}(\bar{T} N_0) \cdot \left\{ \frac{\sqrt{r}}{ \bar{T}^{\frac{1}{4}} N_0^{0.5w + 0.25}} + \frac{ r}{N_0^{w - 0.5} \cdot \min(\bar{T}, N_0)} + \frac{ r \cdot \Delta_E}{N_0^{w}} \right\} \&\le C \cdot \ln^{9/2}(\bar{T} N_0) \cdot \left\{ \frac{\sqrt{r}}{ \bar{T}^{\frac{1}{4}} N_0^{0.5w - 0.25}} + \frac{r}{N_0^{w -1} \cdot \min(\bar{T}, N_0)} + \frac{ r \cdot \Delta_E}{N_0^{w - 0.5}} \right\} \&\le C \cdot \ln^{9/2}(\bar{T} N_0) \cdot \left\{ \frac{G^{\frac{1}{4 \alpha}} N_0^{0.25 - 0.5w}}{ \bar{T}^{\frac{1}{4}} } + \frac{N_0^{1 - w}}{G^{1 - \frac{1}{2 \alpha}}} + \frac{N_0^{0.5 - w}}{G^{0.5 - \frac{1}{2\alpha}}} \right\} \&= o_p(1) \end{align} where in the last line we have used the fact that $w > \frac{1}{2 \alpha}$ and the assumption that $\frac{N_0^{1 - w}}{\bar{T}^{1 - \frac{1}{2 \alpha}}} = o(1)$. {\em Term (ref)}. Using the assumption that $\alpha > \frac{1}{2}$, we have that \begin{align} &\sqrt{r}\Delta_E \le C G^{0.5 \gamma} \cdot G^{- \gamma \alpha} = G^{\gamma(0.5 - \alpha)} = o_p(1), \end{align} which also implies that $r\Delta^2_E = o_p(1)$. {\em Completing the proof of $\mathsf{ATT}$ consistency.} The above inequalities establish that $\widehat{\mathsf{ATT}} - \mathbb{E}[\mathsf{ATT} \mid \bU] = o_p(1)$. {\bf $\textsf{ATU}$ consistency.} The proof follows in an analogous manner to that of $\mathsf{ATT}$, where we switch the roles of the treated and untreated. That is, $M = M_0 = N_0$ and $M_1 = 0$. {\bf $\textsf{ATE}$ consistency.} The proof follows in an analogous manner to that of $\mathsf{ATT}$, where now $M_0, M_1 = \Theta(M)$, and we have $M_0 = N_0$, $M_1 = N_1$, $M = N$. \@startsection{subsection}{2}{0mm}{-\baselineskip}{0.05\baselineskip}{\normalfont\bf}{Proof of Proposition (ref).} For simplicity and without loss of generality, we let the $N_a$ units in $\Ic^{(a)}$ be the indexed as the first $N_a$ units. Now, given the assumption in the statement of Proposition (ref), we have that for $k \in [N_{a}^{1 - \theta}]$ there exists $\beta^{n, k} \in \Rb^{N_{a}^{\theta}}$ such that for all $n \in \Mc^{(1-a)}$ \begin{align} \lambda_n = \sum^{k N_{a}^\theta}_{i = 1 + (k - 1)N_{a}^\theta} \beta^{n, k}_i \lambda_i \end{align} and \begin{align} \| \beta^{n, k} \|_2 = O(\sqrt{r}). \end{align} Define $\tilde{\beta}^n \in \Rb^{N_a}$ as follows \begin{align} \tilde{\beta}^n = \frac{1}{N_{a}^{1 - \theta}}[\beta^{n, 1}, \dots, \beta^{n, N_{a}^{1 - \theta}}]. \end{align} Note \begin{align} \lambda_n = \sum^{N_a}_{i = 1} \tilde{\beta}^n_i \lambda_i. \end{align} Then using $1 - \frac{3}{2\alpha} > \theta \Longleftrightarrow \frac{1 - \theta}{2} - \frac{1}{4\alpha} > \frac{1}{2\alpha}$, we have \begin{align} \| \tilde{\beta}^n \|_2 = O\left( \frac{\sqrt{r}}{N_{a}^{\frac{1 - \theta}{2}}} \right) = O\left( \frac{N_{a}^{\frac{1}{4\alpha}}}{N_{a}^{\frac{1 - \theta}{2}}} \right) = o\left(\frac{1}{N_{a}^{{\frac{1}{2\alpha}}}}\right). \end{align} Then define \begin{align} \tilde{\beta}^{\Ic^{(a)}} &= \sum_{n \in \Mc^{(1-a)}} \tilde{\beta}^n, \\\implies \sum^{N_a}_{i = 1} \tilde{\beta}^{\Ic^{(a)}}_i \lambda_i &= \sum_{n \in \Mc^{(1-a)}} \sum^{N_a}_{i = 1} \tilde{\beta}^n_i \lambda_i = \sum_{n \in \Mc^{(1-a)}} \lambda_n = \lambda_{\Mc^{(1-a)}}. \end{align} Hence, \begin{align} \| \tilde{\beta}^{\Ic^{(a)}} \|_2 = o\left(\frac{M_{1-a}}{N_{a}^{{\frac{1}{2\alpha}}}}\right). \end{align} Since we define $\beta^{(a)}$ to be linear weight with minimum $\ell_2$-norm in Assumption (ref), it follows that $\| \beta^{(a)} \|_2 = o\left(\frac{M_{1-a}}{N_{a}^{{\frac{1}{2\alpha}}}}\right)$. This completes the proof.