EconBase
← Back to paper

Semiparametric Estimation of Long-Term Treatment Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

259,982 characters · 29 sections · 137 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametric Estimation of Long-Term Treatment Effects$^*$

titlepage\begin{spacing}{1} \begin{abstract} \smalltonormalsize{Long-term outcomes of experimental evaluations are necessarily observed after long delays. We develop semiparametric methods for combining the short-term outcomes of experiments with observational measurements of short-term and long-term outcomes, in order to estimate long-term treatment effects. We characterize semiparametric efficiency bounds for various instances of this problem. These calculations facilitate the construction of several estimators. We analyze the finite-sample performance of these estimators with a simulation calibrated to data from an evaluation of the long-term effects of a poverty alleviation program.}\\ \\ Keywords: Long-Term Treatment Effects, Semiparametric Efficiency, Experimentation \\ JEL Codes: C01, C13, C14, O10 \end{abstract} \end{spacing}

\thispagestyle{empty} \setcounter{page}{1}

spacing{1.4} \section{Introduction} Empirical researchers often aim to estimate the long-term effects of policies or interventions. Randomized experimentation provides a simple and interpretable approach to this problem \citep*{fisher1925statistical, duflo_glennerester_2007, athey_imbens_2017}. However, long-term outcomes of experimental evaluations are necessarily observed after long and potentially costly delays. Consequently, there is relatively limited experimental evidence on the long-term effects of economic and social policies.\footnote{bouguen_etal_2019 provide a systematic review of randomized control trials in development economics and report that only a small proportion evaluate long-term effects. Similarly, in a review on experimental evaluations of the effects of early childhood educational interventions, tanner_etal_2015 identify only one randomized evaluation that reports long-term employment and labor market effects gertler_etal_2014.} In this paper, we develop methods for estimating long-term treatment effects by combining short-term experimental and long-term observational data sets. We consider two closely related models, proposed by athey2020combining and athey2020estimating. In both cases, a researcher is interested in the average effect of a binary treatment on a scalar long-term outcome. The researcher observes two samples of data, an experimental sample and an observational sample. The experimental sample measures short-term outcomes of a randomized evaluation of the treatment. the treatment The observational sample measures both short-term and long-term outcomes, but may be subject to unmeasured confounding. The two settings that we consider are distinguished by whether treatment is observed in the observational sample. Similar identifying assumptions are required in each case. We review these assumptions in (ref). Methods for combining experimental and observational data to estimate long-term treatment effects are widely applicable in applied microeconomics and online platform experimentation. The estimators proposed in athey2020estimating have been used to estimate the long-term effects of free tuition on college completion \citep*{dynarski2021closing}, of an agricultural monitoring technology on farmer revenue in Paraguay \citep*{dal2021information}, and of changes to Twitter's platform on user engagement twitter_engineering_2021. As a result, there is considerable interest in advancing a statistical and methodological foundation for this problem gupta2019top. To that end, this paper offers two contributions. First, we develop semiparametric theory for estimation of long-term average treatment effects in (ref). In particular, we derive the semiparametric efficiency bound, and the corresponding efficient influence function, for estimating long-term average treatment effects in each of the models that we consider. We then demonstrate that the efficient influence function is the unique influence function in each model, indicating that all regular and asymptotically linear estimators have the same asymptotic variance and achieve the semiparametric efficiency bound chen2018overidentification, newey1994asymptotic. In both cases, we find that the efficient influence functions possess a “double-robust” structure commonly found in causal estimation problems scharfstein1999adjusting, kang2007demystifying. These results are novel. In particular, our calculations correct statements concerning efficient influence functions and semiparametric efficiency bounds given in a working paper draft of athey2020estimating.\footnote{The results given in athey2020estimating were obtained through a non-rigorous calculation related to standard heuristics involving a discretization of the sample space (see e.g., kennedy2022semiparametric and ichimura2022influence for discussion). Our arguments are rigorous, following the method developed in Section 3.4 of bickel1993efficient.} Analogous results for the model considered in athey2020combining have not appeared before in the literature. Second, in (ref), we establish the consistency and asymptotic normality of a suite of estimators. These estimators differ according to whether they are based on moment conditions associated with the efficient influence functions derived in (ref). Moment conditions defined by efficient influence functions are often referred to as Neyman orthogonal moment conditions, due to their insensitivity to local perturbations of nuisance parameters chernozhukov2018double,foster2019orthogonal. The estimators that we formulate can be viewed as instances of one-step estimators le1956asymptotic, bickel1982adaptive, newey1994asymptotic. Roughly, the conditional expectations that compose an efficient influence function are estimated by augmenting standard machine learning algorithms with cross-fitting. Estimates of long-term average treatment effects are then formed by treating estimated efficient influence functions as identifying moment functions. Estimators obtained from Neyman orthogonal moment conditions admit a very general analysis, under high-level sufficient conditions, due to chernozhukov2018double. We adapt these arguments to our setting. We compare these estimators to a variety of alternative estimators based on non-orthogonal moment conditions. In particular, we consider a set of estimators that are analogous to standard inverse propensity score weighting and outcome regression estimators for average treatment effects under unconfoundedness (see e.g., imbens2004nonparametric for a review). These estimators include, but are not limited to, many of the estimators proposed in athey2020combining and athey2020estimating, and applied by e.g., dynarski2021closing. Here, to obtain theoretical guarantees, we restrict attention to estimators that plug-in nuisance parameter estimates derived from the method of sieves chen2015sieve,chen2014sieve. Relative to those for estimators based on orthogonal moments, sufficient conditions for estimators based on non-orthogonal moments appear more stringent. (ref) assesses the finite sample performance of these semiparametric estimators with a simulation calibrated to data from a randomized evaluation of the long-term effects of a poverty alleviation program originally analyzed in banerjee2015multifaceted. We find that estimators based on orthogonal moments are substantially more accurate than estimators based on non-orthogonal moments. (ref) concludes. Proofs for all results stated in the main text are provided in (ref). (ref) give additional results or details and will be introduced at appropriate points throughout the paper. Code implementing the estimators developed in this paper is available on GitHub at the link \href{https://github.com/DavidRitzwoller/longterm}{https://github.com/DavidRitzwoller/longterm}. \subsection{Related literature} Settings closely related to our own include rosenman2018propensity, rosenman2020combining, and kallus2020role. In rosenman2018propensity, treatment assignment is unconfounded in both samples. In rosenman2020combining, the outcome of interest is observed in both samples. In kallus2020role, treatment is unconfounded in both samples considered jointly, but not necessarily in either sample considered individually. yang2020targeting, gui2020combining, and gechter2021combining give complementary analyses of related models. hou2021efficient consider a model in which measurements of long-term outcomes and treatments are missing completely at random. imbens2022longterm propose several alternative identification strategies for models similar to the model that we consider. In contemporaneous work, singh2021finite and singh2022generalized propose estimators and confidence intervals similar to those developed in (ref) for the case in which treatment is not observed in the observational data set. This paper contributes to the literature on missing data models \citep*{little2019statistical, ridder2007econometrics, hotz2005predicting}, and more specifically to the literature on semiparametric efficiency in missing data models \citep*{chen2008semiparametric, graham2011efficiency, muris2020efficient, bia2020double}. The models that we consider are related to the literatures on statistical surrogacy \citep*{prentice1989surrogate, begg2000use} and mediation analysis \citep*{van2004estimation, imai2010general}. \subsection{Notation} Let $\lambda$ be a $\sigma$-finite measure on the measurable space $\left(\Omega,\mathcal{F}\right)$, and let $\mathcal{M}_\lambda$ be the set of all probability measures on $\left(\Omega,\mathcal{F}\right)$ that are absolutely continuous with respect to $\lambda$. For an arbitrary random variable $D$ defined on $\left(\Omega,\mathcal{F}, Q\right)$ with $Q\in \mathcal{M}_\lambda$, we let $\E_{Q}[D]$ denote its expected value. The quantity $p(d \mid E)$ will denote the density of the random variable $D$ at $d$ conditional on the event $E\in\Omega$ and $l(d \mid E)$ will denote the corresponding log-likelihood. This notation leaves the law of the random variable $D$ implicit, but will cause no ambiguity. We let $\| \cdot \|_{Q,q}$ denote the $L^q(Q)$ norm and $[n]$ denote the set $\{1,\ldots,n\}$. \section{Problem Formulation} Consider a researcher who conducts a randomized experiment aimed at assessing the effects of a policy or intervention. For each individual in the experiment, they measure a $q$-vector $X_i$ of pre-treatment covariates, a binary variable $W_i$ denoting assignment to treatment, and a $d$-vector $S_i$ of short-term post-treatment outcomes. The researcher is interested in the effect of the treatment on a scalar, long-term, post-treatment outcome $Y_i$ that is not measured in their experimental data. They are able to obtain an auxiliary, observational data set containing measurements, for a separate population of individuals, of the long-term outcome of interest, in addition to records of the same pre-treatment covariates and short-term outcomes that were measured in the experimental data set. This observational data set may or may not record whether each individual was exposed to the treatment of interest. In this paper, we develop methods for estimating the effect of a treatment on long-term outcomes that combine experimental and observational data sets with this structure. To fix ideas, consider dynarski2021closing, who randomize grants of free college tuition to a population of high achieving, low income high school students. They estimate the average effects of these free tuition grants on college application and enrollment rates. The long-term effects of free tuition grants on college completion rates may be of more direct interest to policy-makers considering the expansion of college aid programs. However, it will take several years before college completion is observed for the cohort of students in the experimental sample. Given an observational data set that records college enrollment and completion for a comparable population of high school students, the methods developed in this paper may facilitate a more timely quantification of the effect of free tuition grants on college completion. We begin this section by defining the data structure and estimands that we consider. We then review, and restate in a common notation, two sets of closely related sets of identifying assumptions proposed by athey2020combining and athey2020estimating. We refer to these settings as the Latent Unconfounded Treatment and Statistical Surrogacy Models, respectively. \subsection{Data} Consider the collection of random variables \[ \{{A}_i\}_{i=1}^n = \{(Y_i(0), Y_i(1), {S}_i(0), S_i(1), W_i, G_i, {X}_i)\}_{i=1}^n \] consisting of the potential outcomes and characteristics of a sample of individuals drawn independently and identically from a distribution $P_\star$. Here, $G_i$ is a binary indicator denoting whether the observation was acquired in the observational sample ($G_i=1$) or the experimental sample ($G_i=0$). The variables $Y_i(\cdot)$ are long-term potential outcomes, $S_i(\cdot)$ are short-term {potential} outcomes, $W_i$ is a binary treatment indicator, and $X_i$ are covariates. The data observable to the researcher are denoted by $\{{B}_i\}_{i=1}^n$ and are i.i.d.\ according to a distribution denoted by $P$. The observed outcomes $S_i$ and $Y_i$ are given by $ S_i=W_iS_i(1)+(1-W_i)S_i(0) $ and $ Y_i=W_iY_i(1)+(1-W_i)Y_i(0), $ respectively. The short-term outcomes $S_i$ are observed in both the observational and experimental data sets. The long-term outcome $Y_i$ is observed only in the observational data set. Treatment $W_i$ may or may not be observed in the observational sample. Thus, the observable data $B_i$ are given by $(G_iY_i, S_i, W_i, G_i, X_i)$ if treatment is measured in the observational data set and by $(G_iY_i, S_i, (1-G_i)W_i, G_i, X_i)$ if treatment is not measured in the observational data set. \subsection{Estimands} In the main text, we consider estimation of the long-term average treatment effect in the observational population, given by \begin{equation} \tau_1 = \E_{P_\star}\left[Y_i(1) - Y_i(0) \mid G_i = 1\right] . \end{equation}athey2020combining note that there is often reason to believe that features of the observational population ($G_i=1$) are more “externally valid,” in that they of greater interest to policymakers. In (ref), we give results analogous to those presented in the main text for the long-term average treatment effect in the experimental population \begin{equation} \tau_0 = \E_{P_\star}\left[Y_i(1) - Y_i(0) \mid G_i = 0\right] , \end{equation} which may be of interest in some contexts.\footnote{An alternative estimand is the unconditional long-term average treatment effect $\tau = \E_{P_\star}\left[Y_i(1) - Y_i(0)\right]$. The practical interpretation of this parameter is somewhat nebulous, as it is unclear why it would be desirable to weight the two samples in the definition of the parameter according to their sizes.} athey2020combining and athey2020estimating consider estimation of $\tau_1$ and $\tau_0$, respectively. \subsection{Identifying Assumptions} We consider two sets of assumptions, proposed in athey2020combining and athey2020estimating. In both models, the long-term average treatment effect $\tau_1$ is identified. The models differ according to whether they are applicable to contexts in which treatment is or is not measured in the observational data set. Both models assume that treatment assignment is unconfounded in the experimental data set and that the probabilities of being assigned treatment or of being measured in the observational data set satisfy a strict overlap condition. \begin{assumption}[Experimental Unconfounded Treatment] In the experimental data set, treatment is independent of short-term and long-term potential outcomes conditional on pre-treatment covariates, in the sense that \[ W_i \indep (Y_i(0), S_i(0), Y_i(1), S_i(1)) \mid {X}_i, G_i = 0~. \] \end{assumption} \begin{assumption}[Strict Overlap] The probability of being assigned to treatment or of being measured in the observational data set is strictly bounded away from zero and one, in the sense that, for each $w$ and $g$ in $\{0,1\}$, the conditional probabilities \[ P(W = w \mid S, X , G=g) \quad\text{and}\quad P(G= g \mid S, X, W=w) \] are bounded between $\varepsilon$ and $1-\varepsilon$, $\lambda$-almost surely, for some fixed constant $0<\varepsilon<1/2$. \end{assumption} Unconfounded treatment and strict overlap are often satisfied in the experimental sample by design. Assessing and accounting for violations of overlap in observational data are important empirical and methodological issues crump2009dealing. In particular, strong overlap conditions can place stringent restrictions on the data generating process when there are many covariates or short-term outcomes d2021overlap. We view systematic consideration of these issues in our context as an important area for further research. In addition, as our aim is to use the experimental sample to estimate a feature of the observational population, we require an assumption limiting the differences between the two populations. \begin{assumption}[Experimental Conditional External Validity] The distribution of the potential outcomes is invariant to whether the data belong to the experimental or observational data sets, in the sense that \[ G_i \indep \left(Y_i(1), Y_i(0), S_i(1), S_i(0)\right) \mid {X}_i~. \] \end{assumption} (ref) implies that adjustments to the distribution of covariates in the experimental data set are sufficient to obtain approximations to features of the observational population, thereby ruling out unobserved systematic differences between the two populations conditional on covariates. \subsubsection{Latent Unconfounded Treatment} If treatment is measured in the observational data set, then the key identifying assumption is that treatment assignment is unconfounded with respect to the long-term outcome if conditioned on the short-term potential outcomes. We term this restriction “Latent Unconfounded Treatment” following athey2020combining. \begin{assumption}[Observational Latent Unconfounded Treatment] In the observational data set, treatment is independent of the long-term potential outcomes conditional on the short-term potential outcomes and pre-treatment covariates, in the sense that, for $w\in\{0,1\}$, \[ W_i \indep Y_i(w) \mid S_i(w), {X}_i, G_i = 1~. \] \end{assumption} Informally, (ref) states that all unobserved confounding in the observational sample is mediated through the short-term outcomes. (ref) are sufficient for identification of the long-term treatment effect $\tau_1$. We summarize these assumptions with the following shorthand. \begin{defn} The collection of (ref), in a setting where treatment is measured in the observational data set, is referred to as the Latent Unconfounded Treatment Model. \end{defn} Panel A of (ref) displays a causal Directed Acyclic Graph (DAG) that is consistent with the restrictions that the Latent Unconfounded Treatment Model place on the data generating process for the observational data set.\footnote{See pearl1995causal for further discussion of the applications of manipulation of graphical models to causal inference. We do not make use of do-calculus methodology to establish identification.} The following proposition is stated as Theorem 1 in athey2020combining. We state and prove the result for completeness. \begin{prop}[athey2020combining] Under the Latent Unconfounded Treatment Model, the long-term average treatment effect $\tau_1$ is point identified. \end{prop} \begin{figure} \begin{centering} \caption{Causal DAGs Consistent with Assumptions on Observational Data} \begin{tabular}{cc} Panel A: Latent Unconfounded Treatment & \textit{Panel B: Statistical Surrogacy} \tabularnewline \begin{tikzpicture}[yscale=2, xscale = 2] \node (w) at (-1.3,0) [label=left:{W},point,scale = 2]; \node (s) at (0,0) [label=above:{S},point,scale = 2]; \node (y) at (1.3,0) [label=right:{Y},point,scale = 2]; \path (w) edge[line width=1mm, color = black!30] (s); \path (s) edge[line width=1mm, color = black!30] (y); \path (w) edge[bend left=60, line width=1mm, color = black!30] (y); \path[bidirected] (w) edge[bend left=-60, line width=1mm, color = black!30] (s); \path[bidirected] (y) edge[bend left=60, color=red, line width=1mm] (s); \path[bidirected] (y) edge[bend left=60, color=red, line width=1mm] (w); \node at (.60,-.40) [red] {$\text{\Large{X}}$}; \node at (0,-0.73) [red] {$\text{\Large{X}}$}; \end{tikzpicture} & \begin{tikzpicture}[yscale=2, xscale = 2] \node (w) at (-1.3,0) [label=left:{W},point,scale = 2]; \node (s) at (0,0) [label=above:{S},point,scale = 2]; \node (y) at (1.3,0) [label=right:{Y},point,scale = 2]; \path (w) edge[line width=1mm, color = black!30] (s); \path (s) edge[line width=1mm, color = black!30] (y); \path (w) edge[bend left=60, color=red, line width=1mm] (y); \path[bidirected] (w) edge[bend left=-60, line width=1mm, color = black!30] (s); \path[bidirected] (y) edge[bend left=60, color=red, line width=1mm] (s); \path[bidirected] (y) edge[bend left=60, color=red, line width=1mm] (w); \node at (.60,-.40) [red] {$\text{\Large{X}}$}; \node at (0,.73) [red] {$\text{\Large{X}}$}; \node at (0,-0.73) [red] {$\text{\Large{X}}$}; \end{tikzpicture} \tabularnewline \end{tabular} \end{centering} \justifying {Notes: Panels A and B of Figure (ref) display causal DAGs that describe the restrictions on the data generating process for the observational data set ($G=1$) implied by the Latent Unconfounded Treatment and Statistical Surrogacy Models, respectively. Light grey arrows denote the existence of an effect of the tail variable on the head variable. Dark red arrows with x's denote that an effect of the the tail variable on the head variable is ruled out. Dashed bidirectional arrows represent existence of some unobserved common causal variable $U$, where we have a fork $\leftarrow U \rightarrow$. } \end{figure} \subsubsection{Statistical Surrogacy} If treatment is not measured in the observational data set, an alternative “Statistical Surrogacy” assumption, in the spirit of prentice1989surrogate, is required in the place of (ref). Under this restriction, the short-term outcomes can be interpreted as “proxies” or “surrogates” for the long-term outcome. \begin{assumption}[Experimental Statistical Surrogacy] In the experimental data set, treatment is independent of the long-term observed outcomes conditional on the short-term observed outcomes and pre-treatment covariates, in the sense that \[ W_i \indep Y_i \mid S_i, {X}_i, G_i = 0~. \] \end{assumption} Informally, (ref) additionally rules out a causal link from treatment to the long-term outcome that is not mediated by the short-term outcomes. In addition to (ref), a supplementary restriction is required to ensure that the experimental and observational data sets are suitably comparable, conditional on the observed outcomes. \begin{assumption}[Long-Term Outcome Comparability] The distribution of the long-term outcome is invariant to whether the data belong to the experimental or observational data sets conditional on the short-term outcome and covariates, in the sense that \[ G_i \indep Y_i \mid {X}_i, S_i~.\] \end{assumption} (ref) is not strictly stronger than (ref), as belonging to the experimental or observational data sets is not necessarily independent of treatment assignment. Statistical Surrogacy, in addition to (ref), is sufficient for identification of the long-term treatment effect $\tau_1$ in settings where treatment is not measured in the observational data set. Again, we summarize these conditions with the following shorthand. \begin{defn} The collection of (ref), in a setting where treatment is not measured in the observational data set, is referred to as the Statistical Surrogacy Model. \end{defn} Panel B of (ref) describes the restrictions that the Statistical Surrogacy Model place on the data generating process for the observational data set. The following proposition restates Theorem 1 of athey2020estimating. Again, we state and prove the result for completeness.\footnote{We note that (ref) is unnecessary for the identification of $\tau_0$ in the Statistical Surrogacy Model.} \begin{prop}[athey2020estimating] Under the Statistical Surrogacy Model, $\tau_1$ is point identified. \end{prop} \section{Semiparametric Efficiency} In this section, we derive efficient influence functions and corresponding semiparametric efficiency bounds for estimation of long-term average treatment effects $\tau_1$ in observational populations. In (ref), we state comparable results for long-term average treatment effects $\tau_0$ in experimental populations. \subsection{Nuisance Functions} The efficient influence functions that we derive are expressed in terms of a set of unknown, but identified, conditional expectations. We classify each of these objects as being either a “long-term outcome mean” or a “propensity score.” Each long-term outcome mean is an expectation of the long-term outcome conditioned on other features of the data. Under the Latent Unconfounded Treatment Model, where treatment is measured in the observational data set, the long-term outcome means that appear in the efficient influence function are given by \begin{align} \mu_{w}(s,x) &= \E_P[Y \mid W=w, S=s, X=x, G=1]\quad\text{and} \\ \bar{\mu}_w(x) & = \E_{P}[\mu_w(S,X) \mid W=w, X=x, G=0] . \end{align} The function $\mu_w(s,x)$ is the mean of the long term outcomes in the observational sample, conditioned on treatment $w$, the short-term outcomes $s$, and the covariates $x$. The function $\bar \mu_w(x)$ is the projection of $\mu_w(s,x)$ onto the experimental population conditional on $w$ and $x$, integrated over $s$. In turn, under the Statistical Surrogacy model, where treatment is not measured in the observational data set, the analogous nuisance functions appearing in the efficient influence function are given by \begin{align} \nu(s,x) &= \E_P[Y \mid S=s, X=x, G=1]\quad\text{and} \\ \bar{\nu}_w(x) & = \E_{P}[\nu(S,X) \mid W=w, X=x, G=0] . \end{align} The function $\nu(s,x)$ is similar to $\mu_w(s,x)$, but does not condition on treatment, as treatment is not observed in the observational sample under the Statistical Surrogacy Model. The function $\bar \nu_w(x)$ projects $\nu(s,x)$ on the observational sample, conditional on $x$ and $w$, and integrates over $s$. It useful to note that the proofs of (ref) operate by establishing, in their respective models, that the equalities \[ \mathbb{E}_{P_*}\left[Y(1)\mid G=1\right] = \mathbb{E}_{P_*}\left[\bar{\mu}_1(X) \mid G=1\right] \quad\text{and}\quad \mathbb{E}_{P_*} \left[Y(1)\mid G=1\right] = \mathbb{E}_{P_*}\left[\bar{\nu}_1(X) \mid G=1\right] \] hold and that the objects (ref) and (ref) are identified from the observable data. (ref) presents a set of alternative moment conditions that analogously identify $\tau_1$. Each member of the second class of nuisance functions---propensity scores---expresses either the probability of treatment or the probability of inclusion in the observational sample conditioned on other features of the data. In particular, let \begin{align} \rho_w(s,x) & = P_\star(W=w \mid S(w) = s, X = x, G=1) , \\ \varrho(s,x) & = P(W=1 \mid S = s, X = x, G=0) ,\quad\text{and} \\ \varrho(x) & = P(W=1\mid X=x, G=0) \end{align} denote the probabilities of treatment conditional on various features of the data. Similarly, let \begin{align} \gamma(s,x) & = P(G=1\mid S = s, X=x) , \gamma(x) = P(G=1\mid X = x) , \text{and} \pi = P(G=1) \end{align} denote the probabilities of inclusion in the observational sample conditional on various features of the data.\footnote{ athey2020estimating term $\nu(s,x)$ the \emph{surrogate index}, $\varrho(s,x)$ the \emph{surrogate score}, and $1-\gamma(s,x)$ the \emph{sampling score}.} Each of these objects is bounded away from zero and one $\lambda$-almost surely by strict overlap ((ref)). Although the propensity score $\rho_w(s,x)$ includes the unobserved random variable $S(w)$ in its conditioning set, we may write $\rho_w(s,x)$ in terms of observables as \begin{align} \rho_w(s,x) &= \frac{P(G=1 \mid S=s, W=w, X=x) }{P(G=0 \mid S=s, W=w, X=x)} \nonumber \\ &\quad \times \frac{\P( G=0 \mid W=w, X=x)}{\P(G=1 \mid W=w, X=x)} P(W=w\mid X=x, G=1) \end{align} by successive applications of Bayes' rule, (ref), and (ref). \subsection{Influence Functions and Efficiency Bounds} We characterize efficient influence functions for $\tau_1$ with the approach developed in Section 3.4 of bickel1993efficient. In particular, we characterize the tangent spaces for the classes of distributions restricted by the maintained assumptions. We then verify that the conjectured efficient influence functions are pathwise derivatives of $\tau_1$ and are elements of their respective tangent spaces. Recall that the Latent Unconfounded Treatment Model and Statistical Surrogacy Model differ according to the assumptions they impose on the data generating distribution $P_\star$ in addition to whether treatment is assumed to have been measured in the observational sample. \begin{theorem} Let $b = (y,s,w,g,x)$ denote a possible value for the observed data. \begin{enumerate} • Under the Latent Unconfounded Treatment Model, given in (ref), the efficient influence function for the parameter $\tau_1$ is given by \begin{align} \psi_1(b,\tau_1,\eta) &= \frac{g}{\pi} \left( \frac{w(y-\mu_1(s,x))}{\rho_1(s,x)} - \frac{(1-w)(y-\mu_0(s,x))}{\rho_0 (s,x)} + (\bar{\mu}_1(x) - \bar{\mu}_0(x)) - \tau_1 \right) \\ & + \frac{1-g}{\pi} \left(\frac{\gamma (x)}{1-\gamma(x)}\left( \frac{w(\mu_1(s,x)-\bar{\mu}_1(x))}{\varrho(x)} - \frac{(1-w)(\mu_0(s,x)-\bar{\mu}_0(x))}{1-\varrho(x)}\right)\right),\nonumber \end{align} where the parameter $\eta $ collects the nuisance functions appearing in (ref). • Under the Statistical Surrogacy Model, given in (ref), the efficient influence function for the parameter $\tau_1$ is given by \begin{align} & \xi_1(b,\tau_1,\varphi) = \frac{g}{\pi} \left( \frac{\gamma(x)}{\gamma(s,x)} \frac{1- \gamma(s,x)}{1-\gamma(x)}\frac{(\varrho(s,x)-\varrho(x))(y-\nu(s,x))}{\varrho(x)(1-\varrho(x))} + (\bar{\nu}_1(x) - \bar{\nu}_0(x)) - \tau_1\right) \nonumber \\ & \quad \quad \quad + \frac{1-g}{\pi} \left(\frac{\gamma(x)}{1-\gamma(x)} \left(\frac{w(\nu(s,x)-\bar{\nu}_1(x))}{\varrho(x)} -\frac{(1-w)(\nu(s,x)-\bar{\nu}_0(x))}{1-\varrho(x)}\right)\right) , \end{align} where the parameter $\varphi$ collects nuisance functions appearing in (ref).\footnote{We thank Rahul Singh for noting a typo in the statement of $\xi_1(\cdot)$ in a previous draft of this paper, which was missing the factor $\gamma(x)/(1-\gamma(x))$ in the second term.} \end{enumerate} \end{theorem} In our discussion, we will frequently reference the nuisance parameters $\eta$ and $\varphi$ introduced in (ref). We introduce the following notation to expedite our exposition. \begin{defn} Partition the nuisance function $\eta$ and $\varphi$ defined in (ref) into the long-term outcome means, propensity scores, and $\pi$ by \[ \eta = (\omega_\psi, \kappa_\psi, \pi) \quad\text{and}\quad \varphi = (\omega_\xi, \kappa_\xi, \pi). \] In particular, the parameters $\omega_\psi$ and $\omega_\xi$ collect long-term outcome means \[ \omega_\psi = (\bar{\mu}_1(\cdot), \bar{\mu}_0(\cdot), \mu_1(\cdot,\cdot), \mu_0(\cdot,\cdot)) \quad\text{and}\quad \omega_\xi = (\bar{\nu}_1(\cdot), \bar{\nu}_0(\cdot), \nu(\cdot,\cdot))~, \] respectively, and the parameters $\kappa_\psi$ and $\kappa_\xi$ collect propensity scores \[ \kappa_\psi = (\rho_1(\cdot,\cdot), \rho_0(\cdot,\cdot), \varrho(\cdot), \gamma(\cdot)) \quad\text{and}\quad \kappa_\xi = (\varrho(\cdot, \cdot), \varrho(\cdot), \gamma(\cdot, \cdot), \gamma(\cdot))~, \] respectively. \end{defn} \begin{remark} The efficient influence functions $\psi_1(b,\tau_1,\eta)$ and $\xi_1(b,\tau_1,\varphi)$ derived in (ref) are additively separable into two terms associated with the observational and experimental data sets, respectively. The structure of the terms associated with the observational data set resembles the structure of the efficient influence function for the average treatment effect under unconfoundedness hahn1998role; we discuss the relationship between these objects in (ref). \qed \end{remark} \begin{remark} The efficient influence functions $\psi_1(b,\tau_1,\eta)$ and $\xi_1(b,\tau_1,\varphi)$ possess a “double-robust” structure that is prevalent in causal inference and missing data problems kang2007demystifying, bang2005doubly. In particular, the mean-zero property of $\psi_{1} (b,\tau_1,\eta)$ is maintained even if some of the nuisance functions are misspecified. Suppose that we let arbitrary measurable functions $\tilde{\omega}$ replace the long-term outcome means $\omega_\psi$, then $ \E_P\bk{\psi_1(B, \tau_1, (\tilde{\omega},\kappa_\psi, \pi))} = 0. $ Similarly, if the arbitrary measurable functions $\tilde{\kappa}$ replace the propensity scores $\kappa_\psi$, then we have that $ \E_P\bk{\psi_1(B, \tau_1, (\omega_\psi,\tilde{\kappa},\pi))} = 0. $ The efficient influence function $\xi_{1}(b,\tau_1,\varphi)$ is similarly robust to misspecification of either the long-term outcome means or the propensity scores. This result echoes an analogous double-robustness property of the efficient influence function of the average treatment effect under ignorability \citep*{scharfstein1999adjusting}, in which the efficient influence function is mean zero under misspecification of either the conditional means of the outcome variable or the propensity score. The form of the double-robustness entailed here is slightly more general, as $\omega$ and $\kappa$ each collect several nuisance function. Double robustness, in this form, has a useful implication for estimation. We demonstrate in (ref) that, under appropriate regularity conditions, estimators based on the efficient influence functions $\psi_1(b,\tau_1,\eta)$ or $\xi_1(b,\tau_1,\varphi)$ are consistent for $\tau_1$ if \emph{either} the outcome means ($\omega_\psi$ or $\omega_\xi$) or the propensity scores ($\kappa_\psi$ or $\kappa_\xi$) are estimated consistently. We analyze estimators of this form in further detail in (ref). \qed \end{remark} The population variance of the efficient influence function is the semiparametric efficiency bound. The respective bounds are presented in (ref). \begin{cor} Define the conditional variances \begin{align*} \sigma_w^2(s,x) &= \E_{P_\star}[(Y(w) - \mu_w(S,X))^2\mid S=s,X=x]\quad\text{and}\\ \sigma^2(s,x) &= \E_{P_\star}[(Y - \nu(S,X))^2\mid S=s,X=x] \end{align*} as well as the expressions \begin{align*} \Gamma_{w,1}(s,x) & = \frac{\gamma (x)}{1-\gamma(x)} \frac{\left(\mu_w(s,x) - \bar{\mu}_w (x)\right)^2}{\varrho(x)^w(1-\varrho(x))^{1-w}} \text{and} \Lambda_{w,1}(s,x) =\frac{\gamma(x)}{1-\gamma(x)} \frac{\left(\nu(s,x) - \bar{\nu}_w(x)\right)^2}{\varrho(x)^w(1-\varrho(x))^{1-w}} . \end{align*} \begin{enumerate} • Under the Latent Unconfounded Treatment Model, given in (ref), the semiparametric efficiency bound for $\tau_1$ is given by \begin{align} V_1^{\star} & = \E_P\Bigg[\frac{\gamma(X)}{\pi^2} \Bigg( \frac{\sigma_1^2(S,X)}{\rho_1(S,X)} + \frac{\sigma_0^2(S,X)}{\rho_0(S,X)} \nonumber \\ & \quad \quad\quad \quad \quad \quad +(\bar{\mu}_1(X) - \bar{\mu}_0(X) - \tau_1)^2 + \Gamma_{0,1}(S,X) + \Gamma_{1,1}(S,X)\Bigg) \Bigg] . \end{align} • Under the Statistical Surrogacy Model, given in (ref), the semiparametric efficiency bound for $\tau_1$ is given by \begin{align} V_1^{\star\star} & = \E_P\Bigg[\frac{\gamma(X)}{\pi^2} \Bigg( \left(\frac{\gamma(X)}{\gamma(S,X)} \frac{1-\gamma(S,X)}{1-\gamma(X)} \frac{\varrho(S,X)-\varrho(X)}{\varrho(X)(1-\varrho(X))} \right)^2\sigma^2(S,X) \nonumber \\ & \quad \quad\quad \quad \quad\quad + (\bar{\nu}_1(X) - \bar{\nu}_0(X) - \tau_1)^2 + \Lambda_{0,1}(S,X) + \Lambda_{1,1}(S,X)\Bigg) \Bigg] . \end{align} \end{enumerate} \end{cor} \begin{remark} In (ref), we analyze how the semiparametric efficiency bounds derived in (ref) change if different components of the nuisance parameters $\eta$ or $\varphi$ are known a priori. In both models, if the classical propensity score $\varrho (X)$, i.e., the probability of being assigned to treatment in the experimental sample as a function of covariates, is known, then the semiparametric efficiency bounds are unchanged.\footnote{Invariance of the semiparametric efficiency bounds to knowledge of the propensity score $\varrho(X)$ would no longer hold if the estimands of interest were average long-term effects for the treated population.} This echoes an analogous ancillarity result for estimation of average treatment effects under unconfoundedness given in hahn1998role. By contrast, both semiparametric efficiency bounds change if the propensity score $\gamma(x)$, i.e., the probability of being assigned to the observational sample as a function of covariates, is known. This result is relevant for settings where the experimental sample is known to be drawn from the same population as the observational sample and indicates that development of estimators tailored to this setting may be fruitful.\footnote{On the other hand, in that context, consideration of the unconditional long-term treatment effect $\tau=\mathbb{E}[Y(1) - Y(0)]$ is tenable and natural. We expect the efficiency bound for this functional to be invariant to knowledge of the propensity score $\gamma(x)$.} In the Statistical Surrogacy Model, somewhat curiously, the efficient influence function $\xi_{1}(b,\tau_1,\varphi)$ is additionally invariant to knowledge of the distribution, in the observational sample, of the short-term outcomes conditional on covariates, i.e. the law $S \mid X, G=1$. This invariance is a consequence of the choice (ref) in construction of the efficient influence function $\xi_{1}(b,\tau_1,\varphi)$ in the proof of (ref). \qed \end{remark} Next, we demonstrate that the efficient influence functions $\psi_1(b,\tau_1,\eta)$ and $\xi_1(b,\tau_1,\varphi)$ expressed in (ref) are, in fact, the unique influence functions for $\tau_1$ in their respective models. We recall that an influence for $\tau_1$ is any mean-zero and square integrable function function $\tilde{\psi}(b)$ that satisfies the condition \begin{align*} \tau_1^\prime = \mathbb{E}_P[\tilde{\psi}(B)l^{\prime}(B)] , \end{align*} where $\tau_1^\prime$ is the pathwise derivative of $\tau_1$ along an arbitrary parametric submodel evaluated at zero and $l^{\prime}(B)$ is the score function of this submodel; see e.g., Chapter 25 of van2000asymptotic for further discussion. \begin{theorem} There are unique influence functions in each model: \begin{enumerate} • Under the Latent Unconfounded Treatment Model, given in (ref), $\psi_1(b,\tau_1,\eta)$ is the unique influence function for $\tau_1$. • Under the Statistical Surrogacy Model, given in (ref), $\xi_1(b,\tau_1,\varphi)$ is the unique influence function for $\tau_1$. \end{enumerate} \end{theorem} \begin{remark} Let $\mathcal{P}\subset\mathcal{M}_\lambda$ denote the set of probability distributions that satisfy either the Latent Unconfounded Treatment Model or the Statistical Surrogacy Model. In the terminology of chen2018overidentification, (ref), Part (1), demonstrates that $P$ is locally just-identified by $\mathcal{P}$. As a result, by Theorem 3.1 of chen2018overidentification, all regular and asymptotically linear (RAL) estimators of $\tau_1$ are first-order equivalent under the maintained assumptions. In particular, there are no RAL estimators of $\tau_1$ that have smaller asymptotic variances than others. Consequently, semiparametric efficiency is equivalent to regularity and asymptotic linearity under the maintained assumptions. We note that if the propensity score $\varrho(x)$ admits known restrictions, then the resultant model would be semiparametrically over-identified. In this case, not all RAL estimators are first-order equivalent. However, since the efficiency bound does not change, the estimators we propose in the following section remain efficient with known propensity score.\footnote{The semiparametric literature on average treatment effect and local average treatment effect estimation hirano2003efficient,frolich2007nonparametric discusses efficiency as well, despite the fact that the models in question are similarly just-identified.} Moreover, again by Theorem 3.1 of chen2018overidentification, the model $\mathcal {P}$ does not have any locally testable restrictions in the sense that there are no specification tests of the maintained identifying assumptions with nontrivial local asymptotic power. Analogous statements follow from (ref), Part (2). In that sense, the maintained identifying assumptions are minimal.\qed \end{remark} \section{Estimation} We now consider estimation of the long-term average treatment effect $\tau_1$.\footnote{In (ref), we provide an analogous treatment of estimators of the long-term average treatment effect $\tau_0$ in the experimental population.} The estimators that we consider can each be viewed as semiparametric $Z$-estimators associated with an identifying moment function. That is, each estimator is premised on determining the value $\tau_1$ that solves a sample analogue of a moment condition \[ \mathbb{E}_P\left[g(B_i, \tau_1, \zeta)\right]= 0~, \] for some identifying moment function $g(\cdot)$, where $\zeta$ is an unknown nuisance parameter, replaced with its estimated counterpart in practice. Our treatment differs by whether the identifying moment function $g(\cdot)$ is given by the efficient influence functions $\psi_1(\cdot)$ or $\xi_1(\cdot)$, derived in (ref) or given by some other moment condition. Moment conditions defined by influence functions are often referred to as Neyman orthogonal moment conditions. We adapt very general arguments from chernozhukov2018double to establish that these estimators are consistent and asymptotically normal. Our consideration of estimators based on non-orthogonal moments is selective and is more specialized. To obtain theoretical guarantees, we restrict attention to estimators that plug-in nuisance parameter estimates derived from the method of sieves chen2015sieve,chen2014sieve. Sufficient conditions for estimators with this structure are generally available, but are more delicate and difficult to verify. We state and verify these conditions for one of the estimators that we consider.\footnote{The general high-level conditions in chen2015sieve and chen2014sieve apply to each of estimators that we consider. Lower-level conditions, analogous to those discussed in (ref) will exist for these estimators as well. However, as the derivation of these conditions is lengthy and cumbersome, we provide an illustration of this argument for only one estimator.} Throughout, it is useful to keep in mind that the efficient influence functions $\psi_1(\cdot)$ and $\xi_1(\cdot)$ are the only influence functions in their respective models (i.e., (ref)). Consequently, all regular and asymptotically linear estimators are first-order equivalent, that is, their asymptotic variances are all equal to the semiparametric efficiency bound. Thus, the asymptotic variances of different estimators are the same, but the conditions under which they are asymptotically normal may be different. \subsection{Orthogonal Moments} We begin by considering estimators that are built directly on the influence functions $\psi_1(\cdot)$ or $\xi_1(\cdot)$, derived in (ref) with the “Double/Debiased Machine Learning” (DML) construction developed in chernozhukov2018double. \subsubsection{Construction} The DML construction proceeds in two steps. First, estimates of the nuisance functions $\eta$ or $\varphi$, defined in (ref), are computed with cross-fitting. Second, estimates of $\tau_1$ are obtained by plugging the estimated values of $\eta$ and $\varphi$ into their respective efficient influence functions and solving for the values of $\tau_1$ that equate the sample means of these estimates of the efficient influence functions with zero. \begin{defn}[DML Estimators] Let $\hat{\eta}(I)$ and $\hat{\varphi}(I)$ denote generic estimates of $\eta$ and $\varphi$ based on the data $\{B_i\}_{i\in I}$ for some subset $I\subseteq [n]$. Let $ \{I_l\}_{l=1}^k$ denote a random $k$-fold partition of $[n]$ such that the size of each fold is $m=n/k$. The estimator $\hat{\tau}_{1,\mathsf{DML}}$ is defined as the solution to \[ \frac{1}{k}\sum_{l=1}^k \frac{1}{m} \sum_{i\in I_l} \psi_1(B_i,\hat{\tau}_{1,\mathsf{DML}},\hat{\eta}(I_l^c)) = 0 \quad\text{or}\quad \frac{1}{k}\sum_{l=1}^k \frac{1}{m} \sum_{i\in I_l} \xi_1(B_i,\hat{\tau}_{1,\mathsf{DML}},\hat{\varphi}(I_l^c)) = 0 \] for the Latent Unconfounded Treatment and Statistical Surrogacy Models, respectively. \end{defn} \begin{remark} The fundamental structures underlying standard estimators of average treatment effects under unconfoundedness can be classified as being based on either “inverse propensity score weighting” or “outcome regression” imbens2004nonparametric; elements of each structure appear in the estimators formulated in (ref).\footnote{imbens2004nonparametric also discusses estimators based on matching and Bayesian calculations. We do not develop estimators with these structures in this paper, and view their consideration as an interesting extension.} Inverse propensity score weighted (IPW) estimators, also referred to as horvitz1952generalization estimators, are constructed by weighting the observed values of outcomes by their inverse propensity scores; see e.g., rosenbaum1983central and hirano2003efficient. By contrast, outcome regression estimators are constructed by imputing unobserved potential outcomes with estimates of their expectation conditioned on covariates. The estimators formulated in (ref) combine IPW and outcome regression components with an error-correcting structure comparable to the augmented inverse propensity weighted (AIPW) estimator of robins1995analysis. To illustrate, observe that $\psi_1(b,\tau_1,\eta)$ can be interpreted as first approximating $\tau_1$ with the outcome regression $\bar{\mu}_1(x) - \bar{\mu}_0(x)$ in the observational sample. Heuristically, the biases in this approximation, e.g., induced by regularization, are then corrected by applying IPW to the residuals of the approximation of $\bar{\mu}_w(x)$ to $\mu_w(s,x)$ with \[ \frac{w(\mu_1(s,x)-\bar{\mu}_1(x))}{\varrho(x)} - \frac{(1-w)(\mu_0(s,x)-\bar{\mu}_0(x))}{1-\varrho(x)} \] computed in the experimental sample and reweighed by $\gamma(x)/(1-\gamma(x))$ to represent an expectation over the observational sample. However, the correction above may introduce additional biases through the estimation of $\mu_w(s,x)$; these are in turn corrected by applying IPW to the residuals of the approximation of $\mu_w(s,x)$ to $Y(w)$ with \[ \frac{w(y-\mu_1(s,x))}{\rho_1(s,x)} - \frac{(1-w)(y-\mu_0(s,x))}{\rho_0(s,x)} \] computed in the experimental sample. An analogous interpretation can be formulated for the structure of the efficient influence function $\xi_1(b,\tau_1,\varphi)$.\footnote{Note that in the term corresponding to the observational sample in $\xi_1(\cdot)$, the unobserved treatment indictor $w$ is replaced by the probability of treatment conditional on short-term outcomes and covariates.} Further discussion, at varying levels of rigor, of this “bias-correction” interpretation of the structure of estimators based on efficient influence functions is given in Section 4 of kennedy2023semiparametric and Chapter 7 of bickel1993efficient.\qed \end{remark} \begin{remark} At a high-level, the cross-fitting construction used in (ref) is implemented so that the estimation errors, e.g., $\mu_w(S_i, X_i) - \hat{\mu}_w(S_i, X_i)$, and model errors, e.g., $Y(W_i) - \mu_w(S_i, X_i)$, are unrelated for a given observation. Association between these two forms of error may have particularly pernicious effects in finite-samples if estimates of nuisance functions suffer from over-fitting. More technically, cross-fitting allows us to avoid imposing Donsker-type regularity conditions in our asymptotic analysis, which would exclude estimators with non-negligible asymptotic regularization. Standard implementations of popular machine learning algorithms may feature such regularization; see chernozhukov2016locally for detailed discussion and illustration of this point. Further discussion of cross-fitting methods in semiparametric estimation is given in klassen1987consistent and newey2018crossfitting.\qed \end{remark} \begin{remark} We require specialized approaches to estimate the nuisance functions $\rho_w(s,x)$, $\bar{\mu}_w(x)$, and $\bar{\nu}_w(x)$. We estimate $\rho_w(s,x)$ by combining separate estimates of each of the objects displayed in (ref). We estimate $\bar{\mu}_w(x)$ and $\bar{\nu}_w(x)$ by first computing estimates of $\mu_w(s,x)$ and $\nu(s,x)$ in the observational sample, denoted by $\hat{\mu}_w(s,x)$ and $\hat{\nu}(s,x)$, and then computing estimates of $\mathbb{E}_P[\hat{\mu}_w(s,x)\vert W = w, X = x, G = 0] $ and $\mathbb{E}_P[ \hat{\nu}(s,x)\vert W = w,X = x, G = 0]$ in the experimental sample. In (ref), we derive the rate of convergence for particular implementations of estimators with this structure based on linear sieves. \qed \end{remark} \subsubsection{Large-Sample Theory} We now study the asymptotic behavior of the estimators formulated in (ref). First, in (ref), we demonstrate that, under weak regularity conditions and under both models, the estimator $\hat{\tau}_{1,\mathsf{DML}}$ is consistent for $\tau_1$ if either the long-term outcome means or the propensity scores are estimated consistently. Second, we establish asymptotic normality by providing conditions sufficient for the application of Theorem 3.1 of chernozhukov2018double. We impose a set of standard bounds on moments of the data, and a set of conditions on the uniform rates of convergence of nuisance parameter estimators. Throughout, for a collection of scalar-valued nuisance parameters $\theta = (\theta_1, \ldots, \theta_l)$, we let $\|\theta\|_{P,q}=\max_{i\in[l]} \pr{\E_P |\theta_i(B)|^q}^{1/q}$. \begin{assumption}[Moment Bounds] Let $C, c > 0$ be constants. Under the Latent Unconfounded Treatment Model, the moment bounds \begin{align*} \|Y(w)\|_{P,q}\leq C, & \quad \mathbb{E}_P \left[ \sigma^2_w (S,X)\mid X\right]\leq C,\\ \mathbb{E}_P \left[ (Y(w) - \mu_w(S,X))^2\right] \geq c, & \quad \text{ and } \quad c \le \mathbb{E}_P \left[ (\mu_w(S,X) - \bar{\mu}_w(x))^2\mid X\right] \leq C \end{align*} hold for each $w\in\{0,1\}$ and $\lambda$-almost every $X$. Analogous bounds hold for the Statistical Surrogacy Model, where $\sigma(S,X)$, $\nu(S,X)$, and $\bar{\nu}_w(x)$ replace $\sigma_w(S,X)$, $\mu_w(S,X)$, and $\bar{\mu}_w(x)$, respectively. \end{assumption} \begin{assumption}[Convergence Rates] Let $\mathcal{P}\subset\mathcal{M}_\lambda$ be the set of all probability distributions $P$ that satisfy the Latent Unconfounded Treatment Model stated in (ref). Consider a sequence of estimators $\hat\eta_{n}(I_n^c) = (\hat \omega_{\psi,n}, \hat\kappa_{\psi,n}, \hat\pi_n)$ indexed by $n$, where $I_n \subset [n]$ is a random subset of size $m = n/k$ and $\hat{\pi}_n = \frac{1}{n-m} \sum_{i\in I_n^c} G_i$. For some sequences $\Delta_n \to 0$ and $\delta_n \to 0$ and constants $\varepsilon, C > 0$ and $q>2$, that do not depend on $P$, with $P$-probability at least $1-\Delta_n$, \begin{enumerate} • $n^{-1/2} \le \norm{\hat \eta_n - \eta}_ {P,2} \le \delta_n$, • $\norm{\hat \eta_n - \eta}_ {P,q} < C$, • $\varepsilon \le \hat \kappa_n \le 1-\varepsilon$, where the inequalities apply entry-wise, and • $\| \hat{\omega}_{\psi,n} - \omega_\psi \|_ {P,2} \cdot \| \hat{\kappa}_{\psi,n} - \kappa_\psi \|_{P,2} \leq \delta_n n^{-1/2}$. \end{enumerate} Analogous conditions hold for the Statistical Surrogacy Model, where $\varphi$ replaces $\eta$. \end{assumption} \begin{remark} (ref) imposes the restriction that the product of the estimation errors for the long-term outcome means and propensity scores converges at the rate $o(n^{-1/2})$.\footnote{If the true propensity score $\varrho(x)$ is known, then using this information when constructing $\hat{\tau}_{1,\mathsf{DML}}$ may produce an estimator that performs well in finite-samples. However, plugging the true propensity score into an estimator based off of a non-orthogonal moment would probably be inefficient. For example, in hahn1998role, an IPW type estimator for the average treatment effect based off of the true propensity score is shown to be inefficient.} These rates can be achieved, even if the dimensionality of the covariates or the short-term outcomes is increasing with $n$, by many standard machine learning algorithms including the Lasso and Dantzig selector bickel2009simultaneous,belloni2014pivotal, boosting algorithms ye2017boosting, regression trees and random forests wager2015adaptive, and neural networks chen1999improved,farrell2021deep under appropriate conditions on the structure or sparsity of the underlying model. In (ref), we verify that estimates of $\bar{\mu}_w(x)$ and $\bar{\nu}_w(x)$ based on linear sieves can achieve these rates under sufficiently stringent restrictions on the smoothness of the long-term outcome means. It is reasonable to expect that analogous results should be available for more complicated estimators, e.g., featuring penalization or more complicated bases. Results of this form are an interesting direction for further research. \qed \end{remark} (ref) establishes the asymptotic properties of the estimators formulated in (ref). For the sake of brevity, we state the result only for the Latent Unconfounded Treatment Model. An analogous result holds for the Statistical Surrogacy Model. We provide proofs for both results in (ref). \begin{theorem} Let $\mathcal{P}\subset\mathcal{M}_\lambda$ be the set of all probability distributions $P$ for $\{B_i\}_{i=1}^n$ that satisfy the Latent Unconfounded Treatment Model stated in (ref) in addition to (ref). If (ref) holds for $\mathcal P$, then \begin{align} \sqrt{n}(\hat{\tau}_{1,\mathsf{DML}} -\tau_1) \overset{d}{\to} \mathcal{N}(0,V_1^\star) \end{align} uniformly over $P\in\mathcal{P}$, where $\hat{\tau}_{1,\mathsf{DML}}$ is defined in (ref), $V_1^\star$ is defined in (ref), and $\overset{d}{\to}$ denotes convergence in distribution. Moreover, we have that \begin{align} \hat{V}_1^\star = \frac{1}{k}\sum_{l=1}^k \frac{1}{m} \sum_{i\in I_l} \left(\psi_1(B_i,\hat{\tau}_{1,\mathsf{DML}},\hat{\eta}(I_l^c))\right)^2 \overset{p}{\to} V_1^\star \end{align} uniformly over $P\in\mathcal{P}$, where $\overset{p}{\to}$ denotes convergence in probability. As a result, we obtain the uniform asymptotic validity of the confidence intervals \begin{align} \lim_{n\to\infty} \sup_{P\in\mathcal{P}} \Big\vert P\left(\tau_1 \in \left[\hat{\tau}_{1,\mathsf{DML}} \pm z_{1-\alpha/2}\sqrt{\hat{V}_1^\star/n} \right]\right) - (1-\alpha) \Big\vert = 0 , \end{align} where $z_{1-\alpha/2}$ is the $1-\alpha/2$ quantile of the standard normal distribution. \end{theorem} \subsection{Non-orthogonal Moments} The moment functions $\psi_1(\cdot)$ and $\xi_1(\cdot)$ considered in (ref) are quite complicated. Constructing the estimator $\hat{\tau}_{1,\mathsf{DML}}$ requires estimating several propensity scores and long-term outcome means. It is natural to ask whether it suffices to consider simpler estimators based on moment conditions with fewer nuisance functions. In this section, we consider a suite of estimators that do not use the efficient influence functions $\psi_1(\cdot)$ or $\xi_1(\cdot)$ as identifying moment functions. We emphasize estimators that bear a similarity to standard IPW or outcome regression estimators for average treatment effects, several of which were initially proposed in athey2020combining and athey2020estimating. We verify the asymptotic normality of one of these estimators, when nuisance parameters are estimated with the method of sieves, following ideas developed in chen2015sieve,chen2014sieve. \subsubsection{Construction} Each of the estimators that we consider can be viewed as a $Z$-estimator based on a moment function $g(\cdot,\tau_1,\zeta)$, taking as an argument an unknown nuisance parameter $\zeta$. \begin{defn}[Non-orthogonal Moment Estimators] The estimator $\hat{\tau}_{1}(g)$ is defined as the solution to the sample moment condition \[ \frac{1}{n}\sum_{i=1}^n g(B_i,\hat{\tau}_{1}(g),\hat{\zeta}) = 0~, \] where $\hat{\zeta}$ is a generic estimate of $\zeta$ based on the data $\{B_i\}_{i=1}^n$. \end{defn} We consider two classes of moment functions, differing in whether or not the estimators that they entail resemble IPW or outcome regression estimators for average treatment effects. In the Latent Unconfounded Treatment Model, the random variable \begin{align} g_{\mathsf{w}}(B_i, \tau_1, \zeta_{\mathsf{w}}) &= \frac{G_i}{\pi} \left( \frac{W_i Y_i}{\rho_1(S_i,X_i)} - \frac{(1-W_i) Y_i}{\rho_0(S_i,X_i)} \right) - \tau_1 , \end{align} where $\zeta_{\mathsf{w}}=(\pi, \rho_1,\rho_0)$, has mean zero under $P$. The estimator $\hat{\tau}_{1}(g_{\mathsf{w}})$, proposed originally by athey2020combining, only requires estimation of the nuisance parameters in $\zeta_{\mathsf{w}}$ and can be viewed as an analogue to the IPW estimator for average treatment effects. In turn, the functions \begin{align} g_{\mathsf{or,1}}(B_i, \tau_1,\zeta_{\mathsf{or,1}}) &= \frac{G_i}{\pi} \left(\bar{\mu}_1(X_i) - \bar{\mu}_0(X_i)\right) \quad \text{and} \\ g_{\mathsf{or,0}}(B_i, \tau_1,\zeta_{\mathsf{or,0}}) &= \frac{1-G_i}{\pi} \frac{\gamma(X_i)}{1-\gamma(X_i)} \left(\mu_1(S_i,X_i) - \mu_0(S_i,X_i)\right) , \end{align} where $\zeta_{\mathsf{or,1}}$ and $\zeta_{\mathsf{or,0}}$ collect nuisance parameters, yield the outcome regression type estimators $\hat{\tau}_{1}(g_{\mathsf{or,1}})$ and $\hat{\tau}_{1}(g_{\mathsf{or,0}})$. The estimator $\hat{\tau}_{1}(g_{\mathsf{or,0}})$ was originally proposed by athey2020combining. Analogously, in the Statistical Surrogacy Model, the moment function \begin{align} h_{\mathsf{w}}(B_i, \tau_1, \varsigma_{\mathsf{w}}) &= \frac{G_i Y_i}{\pi} \frac{\gamma(X_i)}{1-\gamma(X_i)} \frac{1-\gamma(S_i, X_i)}{\gamma(S_i, X_i)}\pr{ \frac{\varrho(S_i, X_i)}{\varrho(X_i)} - \frac{1-\varrho(S_i, X_i)}{1-\varrho(X_i)} } - \tau_1 , \end{align} where $\varsigma_\mathsf{w}$ collects nuisance parameters, yields an IPW-type estimator $\hat{\tau}_{1}(h_{\mathsf{w}})$ that is similar to an estimator proposed by athey2020estimating. The moment functions \begin{align} h_{\mathsf{or,1}}(B_i, \tau_1,\varsigma_{\mathsf{or,1}}) &= \frac{G_i}{\pi} \left(\bar{\nu}_1(X_i) - \bar{\nu}_0(X_i)\right) \quad \text{and}\\ h_{\mathsf{or,0}}(B_i, \tau_1,\varsigma_{\mathsf{or,0}}) &= \frac{1-G_i}{\pi} \frac{\gamma(X_i)}{1-\gamma(X_i)} \left(\frac{W_i}{\varrho(X_i)} - \frac{1-W_i}{1-\varrho(X_i)}\right)\nu(S_i,X_i) , \end{align} result in the outcome regression type estimators $\hat{\tau}_{1}(h_{\mathsf{or,1}})$ and $\hat{\tau}_{1}(h_{\mathsf{or,0}})$, respectively. The estimator $\hat{\tau}_{1}(h_{\mathsf{or,0}})$ was originally proposed by athey2020estimating. dynarski2021closing use an estimator closely related to $\hat{\tau}_{1}(h_{\mathsf{or,0}})$ with an estimate of $\nu(s,x)$ based on linear regression, in their analysis of the effects of college tuition grants on college complteion rates. \subsubsection{Large-Sample Theory} Theoretical analysis of the large-sample performance of estimators based on non-orthogonal moment conditions requires a more specialized treatment. Sufficient conditions for their asymptotic normality in the literature are often more delicate and stronger than those for the DML estimators considered in (ref). We provide details of this analysis for the estimator $\hat{\tau}_{1}(g_{\mathsf{w}})$ only, as stating and verifying sufficient conditions for asymptotic linearity is cumbersome. Nevertheless, the basic structure of the conditions that we pose, and the method of their verification, is applicable to each of the estimators formulated above. Constructing the estimator $\hat{\tau}_{1}(g_{\mathsf{w}})$ requires an estimate of the nuisance parameter $\zeta_{\mathsf{w}}$. We restrict attention to procedures that estimate the remaining components of $\zeta_{\mathsf{w}}$, i.e., $\rho_1$ and $\rho_0$, with linear sieves. We detail this procedure in (ref). (ref) establishes the asymptotic normality of the estimator $\hat{\tau}_{1}(g_{\mathsf{w}})$ when the nuisance parameter $\zeta_{\mathsf{w}}$ is estimated with the method of sieves. Stating the precise sufficient conditions for this Theorem requires additional notation and definitions, which we defer to (ref). \begin{theorem} Under the Latent Unconfounded Treatment Model and (ref) stated in (ref). we have that \[ \sqrt{n} (\hat{\tau}_{1}(g_{\mathsf{w}}) - \tau_1) \overset{d}{\to} \Norm(0, V_1^\star)~, \] where $V_1^\star$ is the semiparametric efficiency bound for $\tau_1$ defined in (ref). \end{theorem} We again emphasize that, by (ref), \emph{any} regular and asymptotically linear estimator $\hat\tau_1$ for $\tau_1$ achieves the semiparametric efficiency bound $V_1^\star$, if the assumptions that define the Latent Unconfounded Treatment Model are the \emph{only} set of restrictions that are imposed. Thus, the semiparametric efficiency of $\hat{\tau}_{1}(g_{\mathsf{w}})$, per se, is to be expected. The substance of (ref) is the asymptotic linearity. \begin{remark} Informally, Assumption A.3 requires that: \begin{enumerate} • $\hat \zeta_{\mathsf{w}}$ is $o(n^{-1/4})$-consistent; • The parameter space containing $\zeta_\mathsf{w}$ is Donsker; and • The sieve space chosen to approximate $\zeta_\mathsf{w}$ has limited complexity. \end{enumerate} Condition (1) is analogous to the product rate condition in (ref). It is weaker in the sense that no consistent estimators for the nuisance parameters in $\eta$ that are not in $\zeta_{\mathsf{w}}$ are needed. On the other hand, Condition (1) places more stringent conditions on the rate that $\hat \zeta_{\mathsf{w}}$ estimates $\zeta_{\mathsf{w}}$. In particular, (ref) will still hold in situations where some elements of $\zeta_{\mathsf{w}}$ are estimated at a rate slower than $o(n^{-1/4})$, so long as the product of the errors in estimation of the long-term outcome means and propensity scores is smaller than $o(n^{-1/2})$.\footnote{imbens2004nonparametric outlines similar heuristics for average treatment effect estimation under unconfoundedness.} Condition (2) is imposed in order to ensure stochastic equicontinuity for the moment condition $\zeta_\mathsf{w} \mapsto g_\mathsf{w}(\cdot, \tau_1, \zeta_\mathsf{w})$ treated as a process indexed by $\zeta_\mathsf{w}$. This condition ensures that estimating $\tau_1$ and $\zeta_\mathsf{w}$ using the same data does not induce errors that are excessively large. Using sample-splitting would eliminate the need for this condition. Condition (3) is specific to the sieve approach for estimating nuisance parameters. It is not directly imposed in (ref), but may be needed to justify rate conditions when one estimates nuisance parameters with sieves.\qed \end{remark} \section{Simulation} We now compare the estimators formulated in (ref) with a simulation calibrated to data from banerjee2015multifaceted. We find that the DML estimators considered in (ref) are more accurate than the estimators based on non-orthogonal moments considered in (ref), particularly if a nonparametric approach is taken to nuisance parameter estimation. \subsection{Data, Calibration, and Design} banerjee2015multifaceted study randomized evaluations of several similar poverty-alleviation programs implemented by BRAC, a large non-governmental organization. These programs allocated productive assets (typically livestock) to participating households and measured both short-term and long-term economic outcomes. We restrict our attention to data from the evaluation of the program implemented in Pakistan. For each of the $854$ households in our cleaned sample, survey measurements of the consumption levels, food security, assets, savings, and outstanding loans were taken prior to, as well as two and three years after, treatment. We use the pre-treatment measurements as covariates (i.e., $X_i$), the two-year post-treatment measurements as short-term outcomes (i.e., $S_i$), and the three-year post-treatment measurements as long-term outcomes (i.e., $Y_i$). There are $20$ pre-treatment covariates and $21$ short-term outcomes. In the main text, the long-term outcome of interest is total household assets; we give analogous results for total household consumption in (ref). (ref) gives further information on the construction and content of these data. We calibrate a generative model to these data with a Generative Adversarial Network goodfellow2014generative, following a method for simulation design developed in athey2021using. The details of this calibration are given in (ref).\footnote{In (ref), we demonstrate that the joint distribution of data drawn from this model matches the leading moments of joint distribution of the data from banerjee2015multifaceted remarkably closely.} Crucially, a sample drawn from this model consists of covariates and both treated and untreated short-term and long-term potential outcomes for a hypothetical household (the vector $(X_i, S_i(1), S_i(0), Y_i(1), Y_i(0))$ in our notation). That is, we observe the true short-term and long-term treatment effects for each household sampled from this model, and can measure true long-term average treatment effects by averaging over many simulation draws. With this calibrated model, we generate a collection of hypothetical data sets that satisfy either the Latent Unconfounded Treatment Model or the Statistical Surrogacy Model, as desired. The quality of various estimators is then determined by measuring their average accuracy in recovering long-term average treatment effects in a variety of metrics. To generate a hypothetical data set, we draw a collection of samples from the generative mode of size $h\cdot n$, where $h$ is some multiplier that we vary and $n$ is the sample size of the banerjee2015multifaceted data. Each observation is assigned to being either “experimental” or “observational” with probability $1/2$, and so the experimental and observational samples have identical distribution of covariates. Treatment for experimental samples is always assigned uniformly at random. Treatment for observational samples is assigned with possible confounding. Specifically, we determine treatment probabilities with an increasing function, indexed by a parameter $\phi$, of each hypothetical household's true short-term treatment effects. Larger values of $\phi$ indicate more confounding. Details of this simulation design and parameterization of confounding are given in (ref). \subsection{Comparison of Methods} We begin by comparing the estimators formulated in (ref) with a simple difference between the mean long-term treated and untreated outcomes in the observational sample \begin{equation} \hat{\tau}_{1,DM} = \frac{\sum_i G_iW_i Y_i }{\sum_i G_iW_i} - \frac{\sum_i G_i(1-W_i) Y_i }{\sum_i G_i(1-W_i)} . \end{equation} This estimator is naive, making no adjustment for confounding in the observational sample, and is infeasible if treatment is not observed in the observational sample. (ref) compares the absolute bias and root mean squared error of the estimators formulated in (ref) with the naive estimator (ref).\footnote{Measurements of the variance of each estimator are displayed in (ref). We note that in the Latent Unconfounded Treatment Model, one of the outcome regression estimators has very small variance when nuisance parameters are estimated with linear regression. However, the bias of this estimator is very large.} We compare the use of the Lasso tibshirani1996regression, Generalized Random Forests athey2019generalized, and XGBoost chen2016xgboost for nuisance parameter estimation. Details about the implementation of these estimators are given in (ref). Panels A and B display results for the Latent Unconfounded Treatment and Statistical Surrogacy Models, respectively. Columns within each panel vary the sample size multiplier $h$. Each column displays results for the confounding parameter $\phi$ set to zero, indicating no confounding in the observational sample and labeled as “Baseline," in addition to two parameterizations indicating non-zero confounding. \begin{figure} \begin{centering} \caption{Comparison of Estimators} \begin{tabular}{c} \textit{Panel A: Latent Unconfounded Treatment}\tabularnewline \tabularnewline \textit{Panel B: Statistical Surrogacy}\tabularnewline \tabularnewline \end{tabular} \end{centering} \justifying {Notes: (ref) compares measurements of the absolute bias and root mean squared error for the estimators formulated in (ref), in addition to the difference in means estimator defined in (ref). The y-axes are displayed in logs, base 10. The long-term outcome is total household assets. Panels A and B display results for the Latent Unconfounded Treatment and Statistical Surrogacy Models defined in (ref) and (ref), respectively. The columns of each panel vary the sample size multiplier $h$. Each sub-panel displays results for the baseline, unconfounded, case, as well as for the cases that the confounding parameter $\phi$ has been set to $1/2$ and $2/3$. Results for each estimator are displayed with dots of different colors. Results for different nuisance parameter estimators are displayed with dots of different shapes.} \end{figure} In both the Latent Unconfounded Treatment Model and the Statistical Surrogacy Model, the biases of the DML estimators considered in (ref) tend to be substantially smaller than the biases of the alternative outcome regression or weighting estimators considered in (ref), so long as the nuisance parameters are not estimated by linear regression. The weighting estimator has performance more comparable to the DML estimator in the Latent Unconfounded Treatment Model. The performances of the weighting estimator and the outcome regression estimators in the Statistical Surrogacy Model are more similar. (ref) displays measurements of estimator quality for just the DML estimators in a format analogous to (ref), with the addition of a third row measuring one minus the coverage probability of confidence intervals constructed around each estimator. We report these estimates of coverage probabilities in Supplementary (ref). Confidence intervals constructed with the intervals given in (ref) around the estimators formulated in (ref) have coverage probabilities that are reasonably close to the nominal level. \begin{figure} \begin{centering} \caption{Finite-Sample Performance with Different Nuisance Parameter Estimators} \begin{tabular}{c} \textit{Panel A: Latent Unconfounded Treatment}\tabularnewline \tabularnewline \textit{Panel B: Statistical Surrogacy}\tabularnewline \tabularnewline \end{tabular} \end{centering} \justifying{Notes: (ref) displays measurements of the quality of the estimators formulated in (ref) implemented with several alternative choices of nuisance parameter estimators. The long-term outcome is total household assets. Panels A and B display results for the estimators defined in (ref) for the Latent Unconfounded Treatment and Statistical Surrogacy Models defined in (ref) and (ref), respectively. The columns of each panel vary the sample size multiplier $h$. The rows of each panel display the absolute value of the bias, one minus the coverage probability, and the root mean squared error of each estimator, from top to bottom, respectively. A dotted line denoting one minus the nominal coverage probability, 0.05, is displayed in each sub-panel in the third row. Each sub-panel displays a bar graph comparing measurements of the performance of the estimator defined in (ref), constructed with three types of nuisance parameter estimators, with the difference in means estimator (ref). Each sub-panel displays results for the baseline, unconfounded, case, as well as for the cases that the confounding parameter $\phi$ has been set to $1/2$ and $2/3$.} \end{figure} \section{Conclusion} We study the estimation of long-term treatment effects through the combination of short-term experimental and long-term observational data sets. We derive efficient influence functions and calculate corresponding semiparametric efficiency bounds for this problem. These calculations facilitate the development of estimators that accommodate the applications of standard machine learning algorithms for estimating nuisance parameters. We demonstrate with simulation that these estimators are able to recover long-term treatment effects in realistic settings. Important unresolved practical issues remain. Methods for choosing valid and informative short-term outcomes and assessing of the sensitivity of estimates to violations of identifying assumptions would be valuable. Additional useful extensions include the incorporation of instruments and continuous treatments and the accommodation of settings with limited covariate overlap. In the case of estimation of average treatment effects under unconfoundedness, recent promising work athey2018approximate, bradic2019sparsity, tan2020model has developed estimators that are specifically optimized to handle high-dimensional covariates and are able to attain various notions of optimality under weak assumptions on, e.g., the sparsity of the outcome regression or propensity score models. It is not immediately clear how to apply these ideas to long-term average treatment effects. Further consideration of this problem would be a potentially valuable extension, as the resultant estimators may be particularly well-suited to the case where there are many short-term outcomes. Some progress on a related problem (where, effectively, there is a single short-term outcome) has been made by viviano2021dynamic.
spacing{1.172}
appendix\begin{center} {\it Supplemental Appendix to:} \vskip0.2cm \begin{spacing}{1} {Semiparametric Estimation of \\Long-Term Treatment Effects\daggerfootnote{Date: \today}} \begin{tabular}[t]{c@{\extracolsep{2em}}c} {Jiafeng Chen} & {David M. Ritzwoller}\\ {\it Harvard Business School} & {\it Stanford Graduate School of Business} \\ {[email removed]} & {[email removed]} \end{tabular} \end{spacing} \end{center} \begin{spacing}{1.13} \DoToC \end{spacing} \thispagestyle{empty} \setcounter{page}{0} \begin{spacing}{1.4}

Proofs for Results Presented in the Main Text

Proof of Proposition (ref)

Unless otherwise noted, we drop the dependence on the individual $i$. It will suffice to consider the parameter \[ \theta_{1,1} = \mathbb{E}_{P_\star}[Y(1) \mid G = 1]~, \] as the point-identification of $\theta_{1,0} = \mathbb{E}[Y(0) \mid G = 1]$ will follow by an analogous argument.

We consider Latent Unconfounded Treatment Model, specified in (ref), and repeat the argument provided in support of Theorem 1 of athey2020combining. Define the functions

align*[align* omitted — 141 chars of source]

and

align*[align* omitted — 149 chars of source]

and observe that, by (ref), $\mu_1(s,x)$ is identified from the observational sample and $\bar{\mu}_1(x)$ is identified from the experimental sample given the identification of $\mu_1(s,x)$. Observe that

align*[align* omitted — 804 chars of source]

Hence, $\tau_1$ is identified as $\bar{\mu}_1(x)$ is identified.

Proof of Proposition (ref)

We consider the Statistical Surrogacy Model, specified in (ref). Define the function \[ \nu(s, x) = \mathbb{E}_{P}[ Y \mid S=s, X=x, G = 1] \] and observe that $\nu(s, x)$ is identified from the observational sample by (ref). Observe that

align*[align* omitted — 347 chars of source]

and that

align*[align* omitted — 445 chars of source]

It is clear that $\mathbb{E}_{P}[ \nu(S, X) \mid X, W=1, G=0 ]$ is identified from the experimental sample by (ref), and that therefore $\tau_1$ is identified. \qed

Proof of Theorem (ref). Part 1

Throughout, for a general functional $\varphi$ on $\mathcal{M}_\lambda$ and an arbitrary random variable $D$ with distribution $Q$ in $\mathcal{M}_\lambda$, we write $ \varphi(Q)=\varphi(D) $ with a minor abuse of notation.We let \[ \varphi^\prime(D) = \varphi^\prime(D; 0) = \frac{d}{d\varepsilon} \varphi (Q_\varepsilon) {\small\evalbar}_{\varepsilon=0} \] denote the pathwise derivative of $\varphi(\cdot)$ on some regular parametric submodel $\mathcal{Q} = \{Q_\varepsilon : \varepsilon \in [-1,1]\}$ of $\mathcal{M}_\lambda$ evaluated at $\varepsilon=0$.

It suffices to derive the efficient influence function for the parameter $\theta_{1,1}$, defined more generally by $\theta_{w,g}=\mathbb{E}[Y_i (w)\mid G_i = g]$, as an analogous argument will hold for $\theta_{0,1}$. Unless otherwise noted, we drop the dependence on the individual $i$. We omit the subscript $w$ on the nuisance functions $\mu_w(s,x)$, $\bar{\mu}_w(x)$, and $\varrho_w(s,x)$, as only the case $W = 1$ will be relevant.

The presentation of the proof will be somewhat constructive, as we hope to convey the rationale behind the form of the efficient influence function, for which we have had to take a few educated guesses in forming a conjecture. Nonetheless, we state that the efficient influence function $\psi_{1,1}(b)$ takes the form \[ \psi_{1,1}(b) = \frac{g w (y-\mu(s,x))}{\pi \rho(s,x)} +\frac{(1-g)w}{(1-\gamma(x)) \varrho(x)} \frac{\gamma(x)}{\pi} (\mu(s,x) - \bar\mu(x)) + \frac{g}{\pi} (\bar \mu(x) - \theta_{1,1}) \] The reader can of course verify (ref) directly by using this stated form of the influence function and checking a series of conditions that we lay out.

The argument proceeds as follows. First, we factorize the density of the observed data and relate each factor to conditional densities of the complete data distribution. Second, for an arbitrary smooth parametric submodel, we characterize the tangent space of the distribution of the observed data as a linear space of mean-zero square-integrable functions. Third, we conjecture a functional form that explicitly resides in the tangent space, where we note that the efficient influence function is the unique influence function for the target statistical functional $\theta_{1,1}$ in the tangent space of $\theta_{1,1}$ van2000asymptotic. Fourth, we state the conditions an influence function of the pathwise derivative of $\theta_{1,1}$ necessarily satisfies, and we verify that the conjectured efficient influence function satisfies these conditions.

Density Factorization. Consider the random variable \[ A_1 = (Y(1), S(1), W, G, X) \equiv (Y^*, S^*, W, G, X) \] where we observe $B_1 = (WGY^*, WS^*, W, G, X)$. Let $Y = WGY^*$ and $S = WS^*$. Recall that the distribution of the complete data is denoted by $P_\star$, with density $p_\star$ with respect to $\lambda$, and that the distribution of the observed data is denoted by $P$, with density $p$ with respect to $\lambda$. Observe that the densities under the complete data distribution and the observed data distribution agree when data isn't missing. That is, for all $(y,s)\in\mathbb{R}\times\mathbb{R}^d$ and events $E\in\mathcal A$, we have that \[ p_\star(y \mid W = 1, G=1, E) = p(y \mid W=1, G=1, E) \] and \[ p_\star(s \mid W = 1, E) = p(s \mid W=1, E) \] and that for all events $E$ measurable with respect to the $\sigma$-algebra generated by $(W,G,X)$, we have that $P(E) = P_\star(E)$.

By (ref), the density of the observed data $p$ admits the factorization

align*[align* omitted — 213 chars of source]

Each of the terms of this expression can be rewritten in terms of the density of the complete data $p_\star$. In particular, we have that the first term can be written

align[align omitted — 274 chars of source]

the second and third terms can be written

align*[align* omitted — 205 chars of source]

and

align*[align* omitted — 140 chars of source]

the fourth term can be written

align*[align* omitted — 94 chars of source]

and finally the fifth term can be written

align*[align* omitted — 182 chars of source]

As a result, we obtain a factorization of the density of the observed data in terms of conditional distributions of the complete data

align[align omitted — 261 chars of source]

Characterization of the Tangent Space. Let $\mathcal{P}$ be any regular parametric submodel of $\mathcal{M}_\lambda$ indexed by $\varepsilon \in\R$, with densities $ p_\varepsilon := dP_\varepsilon/d\lambda $ with $p_0=dP/d\lambda$ and such that (ref) hold for each $P_\varepsilon \in \mathcal{P}$. Let the subscript $_\star$ distinguish log densities of the complete data from log densities of the observed data. By the factorization (ref), we obtain the score

align[align omitted — 557 chars of source]

Thus, the tangent space $\mathcal{T}$ is given by mean-square closure of the linear span of the functions

align*[align* omitted — 468 chars of source]

where the functions $s_1$ through $s_5$ range over the space of mean-zero and square integrable functions that satisfy the restrictions

align*[align* omitted — 241 chars of source]

Pathwise Differentiability Conditions. In the following calculations, it is often easier to work with the complete data distribution. To that end, we note that for an arbitrary measurable function $h$, we have that certain conditional means of the observed distribution equal certain conditional means of the complete data distribution:

align*[align* omitted — 287 chars of source]

and, as a result, for $h$,

align*[align* omitted — 190 chars of source]

Note that $\theta_{1,1}$ is such a functional with $h(y,s,x) = y$.

Observe that the parameter of interest can be written

align*[align* omitted — 272 chars of source]

The pathwise derivative of this parameter at 0 on $\mathcal{P}$ is given by

align[align omitted — 215 chars of source]

Now, observe that

align*[align* omitted — 80 chars of source]

and that

align*[align* omitted — 66 chars of source]

Thus, each of the terms in (ref) can be written in terms of conditional scores by

align*[align* omitted — 497 chars of source]

respectively.

An influence function for $\theta_{1,1}$ is a mean-zero and square-integrable function $\tilde{\psi}_{1,1}(B_1)$ that satisfies the condition

equation[equation omitted — 119 chars of source]

Hence, by ((ref)) and ((ref)), in order to establish that a mean-zero and square-integrable function $\tilde{\psi}_{1,1} (B_1)$ is an influence function for $\theta_{1,1}$, it suffices to show that the terms in the right-hand side of (ref) match the terms in the right-hand side of (ref):

align[align omitted — 1,097 chars of source]

Conjectured Efficient Influence Function. We begin by considering an element of the tangent space. For the unspecified functions $a(s,x)$ and $b(x)$, consider the choices

align[align omitted — 281 chars of source]

These choices correspond to the following conjectured efficient influence function

align[align omitted — 322 chars of source]

In particular, $s_4$ is chosen so that certain terms cancel to construct the term $(1-g)w s_2(s,x)$, which does not, prima facie, belong in the tangent space (ref).

In order to ensure that $\psi_{1,1}(b_1)$ belongs to the tangent space, we require

align[align omitted — 166 chars of source]

In the sequel, we verify that the choices

align*[align* omitted — 136 chars of source]

satisfy these conditions, and, moreover, give rise to the efficient influence function by satisfying the conditions (ref)--(ref).

Verifying Pathwise Differentiability Conditions. We begin by stating and proving a useful lemma for handling terms of the form $GWh(Y,S,X)$.

lemmaUnder the assumptions of (ref), if $h(y,s,x)$ is any measurable function, then \[ \E_P\bk{\frac{GW}{\rho(S,X)} h(Y,S,X)} = \E_{P_\star}\bk{G h(Y^*,S^*,X)} = \E_{P_\star}\bk{G \E[h(Y^*,S^*,X) \mid S^*, X]}~. \]
proofObserve that \begin{align*} \E_P\bk{\frac{GW}{\rho(S,X)} h(Y,S,X)} & = \E_P\bk{ \pi \frac{W}{\rho(S,X)} h(Y,S,X)\mid G=1 } \\ &= \pi \E_{P_\star}\bk{ \frac{W}{\rho(S^*,X)} h(Y^*,S^*,X)\mid G=1 } \\ &= \pi \E_{P_\star}\bigg[ \rho^{-1}(S^*, X) \E_{P_\star}[W \mid S^*, X, G=1] \\ &\quad\quad\quad \E_{P_\star}[h(Y^*, S^*, X) \mid S^*, X, G=1] \mid G=1 \bigg] \tag{(ref)} \\ &= \E_{P_\star}[\pi h(Y^*, S^*, X) \mid G=1]\\ &=\E_{P_\star}[G h(Y^*, S^*, X)] , \end{align*} giving the first equality. The second equality then follows from \begin{align*} \E_{P_\star}[G h(Y^*, S^*, X)] &= \pi \E_{P_\star}[\E_{P_\star}[h(Y^*, S^*, X) \mid S^*, X, G=1] \mid G=1] \\ &= \E_{P_\star}[G \E_{P_\star}[h(Y^*, S^*, X) \mid S^*, X, G=1]] \\ &= \E_{P_\star}[G \E_{P_\star}[h(Y^*, S^*, X) \mid S^*, X]] , \end{align*} completing the proof.

It now suffices to verify the conditions (ref)--(ref) for choices of $a(S,X)$ and $b(X)$ that satisfy (ref) and (ref). The remainder of the proof verifies these conditions sequentially.

Condition (ref): Observe that

align*[align* omitted — 225 chars of source]

as required.

Condition (ref): We derive the restrictions on $a(s,x)$ implied by (ref). Observe that the inner product from (ref) can be written

align*[align* omitted — 473 chars of source]

Note that the first term can be eliminated by (ref):

align*[align* omitted — 82 chars of source]

Observe that \[\E_{P_\star}[G(1-W) \mid X] = P_\star(W \mid G=1, X) \gamma(X) = (1-\E_{P_\star}[\rho(S,X)\mid X]) \gamma(X),\] and thus the second term is given by

align*[align* omitted — 294 chars of source]

Next, observe that \[\E_{P_\star}[(1-G)W \mid S^*, X] = (1-\gamma(X)) P_\star(W=1 \mid G=0, S^*, X) = (1-\gamma(X))\varrho(X),\] and therefore the third term can be written \[ \E_P\bk{ \frac{(1-G)W}{(1-\gamma(X)) \varrho(x)} a(S,X)l'_*(S \mid X)} = \E_{P_\star}\bk{a(S^*,X) l'_*(S^* \mid X)}~. \] The fourth term is given by \[ \E_P[GWb(X) l'_*(S \mid X)] = \E_{P_\star}[\gamma(X) q(S^*, X) b(X) l'_*(S \mid X)]~, \] which cancels with the second term.

Thus, the condition (ref) is equivalent to

equation[equation omitted — 137 chars of source]

We proceed by verifying this condition for a choice of $a(S,X)$ that satisfies (ref), as required. Consider the function \[ a(s,x) = \frac{\gamma(x)}{\pi}(\mu(s,x) - \bar{\mu}(x))~. \] Then $\E_{P_\star}[a(S^*,X) \mid X] = 0$, which satisfies (ref). Moreover,

align*[align* omitted — 313 chars of source]

where the second to last equality follows by (ref), which verifies (ref) and thereby (ref).

Condition (ref): We derive the restrictions on $b(x)$ implied by (ref). Note that the conjectured influence function (ref) can be written as \[ \psi(b_1) = gw \cdot \alpha(y,s,x) + (1-g)w\cdot\beta(s,x) + gb(x), \] where

align*[align* omitted — 103 chars of source]

We first show that the terms in (ref) involving $\alpha$ and $\beta$ are zero. Indeed, for terms involving $\alpha$, we have that \[ \E_P[GW\cdot\alpha(Y,S,X)l'_*(G, X)] = \pi \E_P[\alpha(Y,S,X)l'_*(G, X)\mid G=1, W=1] P(W=1\mid G=1)=0 \] since $\E_P[\alpha(Y,S,X)l'_*(G, X; 0)\mid G=1, W=1] = 0$. Analogously, we have that \[ \E_P[(1-G) W \beta(S,X) l'_*(G, X; 0)] = 0. \] Therefore, (ref) is equivalent to

equation[equation omitted — 168 chars of source]

We proceed by verifying this condition for a choice of $b(X)$.

Consider the function \[b(x) = \frac{1}{\pi}(\bar{\mu}(x) - \theta_{1,1})~,\] which satisfies (ref). We can compute

align*[align* omitted — 592 chars of source]

and \[ -\E_{P_\star}[G\pi^{-1}\theta_{1,1} l'_*(G,X;0)] = -\theta_{1,1} \E_{P_\star}[l'_*(1,X;0) \mid G=1] \] which verifies (ref) and thereby (ref).

Condition (ref): Since the first and second terms $\alpha(Y,S,X)$ and $\beta(S,X)$ in the conjectured efficient influence function (ref) are mean zero conditional on $S^*$ and $X$ and $X$, respectively, the right-hand side of (ref) is equal to

align*[align* omitted — 516 chars of source]

by iterated expectation conditional on $X$.

Condition (ref): The only term on the right-hand side of (ref) is \[ \E_P[(1-G)W\cdot\beta(S,X)\cdot (W-\rho(X)) \delta(X)] \] for some function $\delta(\cdot)$, where $(1-G) W \cdot\beta(S,X)$ is the second term in the conjectured efficient influence function (ref). We have:

align*[align* omitted — 248 chars of source]

by the condition $\E_{P_\star}[\beta(S^*,X) \mid X] = 0$, completing the proof. \qed

Proof of (ref). Part 2

We maintain the notation of the proof of (ref), Part 1. In particular, we omit the $w$ subscript in the nuisance function $\bar{\nu}_w(x)$ and let \[ A_1 = (Y(1), S(1), W, G, X) = (Y^*, S^*, W, G, X) \] and $B_1 = (Y,S,W,G,X)$ denote the complete and observed data, respectively. The efficient influence function we construct is

align*[align* omitted — 314 chars of source]

Density Factorization. By (ref), the density $p$ of the observed data distribution admits the factorization

align[align omitted — 242 chars of source]

Characterization of the Tangent Space. Let $\mathcal{P}$ be any regular parametric submodel of $\mathcal{M}_\lambda$ indexed by $\varepsilon \in\R$, with densities $ p_\varepsilon := dP_\varepsilon /d\lambda $ with $p_0=dP/d\lambda$ and such that (ref) hold for each $P_\varepsilon \in \mathcal{P}$. By the factorization (ref), we obtain the score

align[align omitted — 267 chars of source]

Thus, the tangent space $\mathcal{T}$ is given by the mean-square closure of the linear span of the functions

align[align omitted — 285 chars of source]

where the functions $s_1$ through $s_6$ range over the space of mean-zero and square integrable functions that satisfy the restrictions

align[align omitted — 371 chars of source]

and $s_5,$ and $s_6$ are unconstrained, where (ref) and (ref) follow from (ref) and (ref), respectively.

Pathwise Differentiability Representation. Recall from the proof of (ref) that, by (ref), the target parameter $\theta_{1,1}$ can be written as

align*[align* omitted — 254 chars of source]

The pathwise derivative of this parameter is given by

align*[align* omitted — 99 chars of source]

with

align*[align* omitted — 240 chars of source]

where final three terms follow from

align*[align* omitted — 58 chars of source]

and Bayes' rule in the form of \[ l(x \mid G=1) = \log \gamma(x) + l(x) - \log \int \gamma(x)p(x)\,d\lambda(x)~. \] Observe that, using the definitions of $\nu(s,x)$ and $\bar{\nu}(x)$, we may further simplify the pathwise derivative to give

align[align omitted — 371 chars of source]

An influence function for $\theta_{1,1}$ is a mean-zero and square-integrable function $\tilde{\xi}_{1,1}(B_1)$ that satisfies the condition

equation[equation omitted — 89 chars of source]

We make a conjecture for such a function and verify that it satisfies this condition and is an element of the tangent space $\mathcal{T}$, establishing that it is the efficient influence function.

Conjectured Efficient Influence Function. We conjecture that the efficient influence function takes the form

equation[equation omitted — 153 chars of source]

for functions $f_1(s,x)$, $f_2(x)$, and $f_3(x)$ to be specified, where $f_3$ will be chosen to satisfy

equation[equation omitted — 85 chars of source]

Observe that for this choice, $\xi_{1,1}(b)$ will be an element of the tangent space $\mathcal{T}$. This follows as \[s_1(y\mid s,x, G=1) = (y-\nu(s,x)) f_1 (s,x)\] satisfies (ref), we make the choice \[s_2(s \mid X, G=1)=0 \numberthis \label{eq:surprising_ancillarity}\] satisfying (ref), \[s_3(s\mid W=1, x, G=0) = (\nu(s,x) - \bar{\nu}(x))f_2(x)\] satisfies (ref), and that setting $s_4(x) = \gamma(x) f_3(x)$ and $s_5(x) = \gamma(x) (1-\gamma(x)) f_3(x)$ yields \[ g f_3(x)= s_4(x) + \frac{g-\gamma(x)}{\gamma(x) (1-\gamma(x))} s_5(x)~, \] satisfying (ref), the mean zero conditions for $s_4$ and $s_5$ by (ref), and implicitly satisfying (ref) and the mean zero condition for $s_6$.

Verification of Pathwise Differentiability Representation. In the sequel, we make choices of $f_1(s,x)$, $f_2(x)$, and $f_3(x)$ such that the pathwise derivative of $\theta_{1,1}$ satisfies (ref) and (ref), completing the proof.

Our conjecture for $\xi_{1,1}(b_1)$ allows for some simplification of the right-hand side of (ref). Note that, for any measurable functions $h(S,X)$ and $k(X)$, we have

align*[align* omitted — 107 chars of source]

by iterated expectations. Thus, by additionally applying the mean-zero property of the scores $l'(Y \mid S, X, G=1)$, $l'(S \mid X, G=1)$, we can simplify the pathwise differentiability condition to give

align[align omitted — 273 chars of source]

Therefore, we obtain a set of sufficient conditions for (ref) by matching each term in (ref) with (ref), giving

align[align omitted — 619 chars of source]

We complete the proof by considering each condition sequentially, making appropriate choices of $f_1(s,x)$, $f_2(x)$, and $f_3(x)$, where we note that each function appears in only one term.

Condition (ref): Evaluating the left-hand side of (ref) we find that

align*[align* omitted — 213 chars of source]

Thus, if we choose \[ f_1(s,x) = \frac{p(s \mid x, W=1, G=0)}{\gamma(x) p(s\mid x, G=1)} \frac{p (x \mid G=1)}{p(x)} \] to change the measure of the two outer expectations, (ref) is satisfied. Bayes' rule calculations show that \[ f_1(s,x) = \frac{1}{\pi} \frac{\varrho(s,x)}{\varrho(x)} \frac{1-\gamma(s,x)}{\gamma(s,x)}\frac{\gamma(x)}{1-\gamma(x)}~. \]

Condition (ref): Observe that, for any measurable function $h(s,x)$, we have that \[ \E_P[W(1-G) h(S,X)] = \E_P[\varrho(X) (1-\gamma(X)) \E_P[h(S,X) \mid X, W=1, G=0]]. \] Also, recall that since $l'(S \mid X, W=1, G=0)$ is a conditional score, it is conditionally orthogonal to any function of $X$; that is, for any measurable function $h(x)$, \[ \E_P[h(X) l'(S\mid X, W=1, G=0) \mid X, W=1, G=0] = 0~. \]

With the above two observations in mind, evaluating the left-hand side of (ref), we find that

align*[align* omitted — 297 chars of source]

Thus, again, if we choose \[ f_2(x) = \frac{p(x\mid G=1)}{p(x) \varrho(x) (1-\gamma(x)) } = \frac{1}{\pi} \frac{1}{\varrho(x)} \frac{\gamma(x)}{1-\gamma(x)} ~, \] again to change the measure of the outer expectation, then (ref) is satisfied.

Condition (ref): Lastly, consider the choice \[f_3(x) = \frac{\bar{\nu}(X) - \theta_{1,1}}{\pi}~,\] which satisfies (ref). In this case, we find that the left-hand side of (ref) is given by

align*[align* omitted — 180 chars of source]

which is exactly the right-hand side of (ref).

Collecting the above choices for $f_1(s,x)$, $f_2(x)$, and $f_3(x)$ gives the efficient influence function

align*[align* omitted — 302 chars of source]

as required. \qed

Proof of Theorem (ref). Part 1

We maintain the notation from the proof of (ref). Let $\tilde{\psi}_{1,1}(b_1)$ be an influence function for $\theta_{1,1}$. It is necessarily that case that \[ \tilde{\psi}_{1,1} (b_1) = \psi_{1,1}(b_1) + f(b_1)~, \] where $\psi_{1,1}(b_1)$ is the efficient influence function in derived in the proof of (ref) and $f(b_1)$ is mean-zero and orthogonal to the tangent space $\mathcal{T}$, also defined in the proof of (ref). Moreover, we may write

align[align omitted — 131 chars of source]

where $A$ is a constant making $f$ mean-zero.

The orthogonality between $f(b_1)$ and the tangent space $\mathcal{T}$ implies the following set of conditions that are analogous to the conditions (ref) through (ref) defined in the proof of (ref):

align[align omitted — 671 chars of source]

In the sequel, we characterize $f$ by deriving restrictions from each of these conditions for particular choices of the scores $l_*'(\cdot)$, $\varrho'(x)$, and $\rho'(s,x)$.

Combining (ref) and (ref) yields \[ \E_P\bk{GW f_1(Y,S,X) l'_*(Y \mid S,X)} = 0~. \] As $l'_*(Y \mid S,X)$ ranges over all conditionally mean-zero, square integrable functions, we may pick \[ l'_*(Y \mid S,X) = f_1(Y,S,X) - \E_P[f_1(Y,S,X) \mid G=1, W=1, S, X], \] from which we find that

align*[align* omitted — 142 chars of source]

and so $f_1(Y,S,X) = f_1(S,X)$ does not depend on $Y$.

Next, observe that plugging (ref) into (ref), yields \[ 0 = \E_{P_\star}\bk{\gamma(X) \pr{ f_1(S^*, X) \rho'(S^*,X) - f_3(X) \E_{P_\star}[\rho'(S^*, X) \mid X]}} \] after simplification via law of iterated expectations. Choosing \[ \rho'(S^*, X) = \gamma(X)^{-1}\pr{f_1(S^*,X) - \E_{P_\star}[f_1(S^*,X)\mid X]} \] establishes that \[ 0 = \E_{P_\star}[f_1^2(S^*, X)-f_1(S^*, X)\E[f_1(S^*,X)\mid X]]~, \] and so $f_1(S,X) = f_1(X)$ does not depend on $S$. This then implies that $f_1(x) = f_3(x)$ by law of iterated expectations.

Now, plugging (ref) into (ref) yields \[ 0 = \E_{P_\star}\bk{(1-\gamma(X)) \varrho(X) f_2(S^*,X) l'_*(S^* \mid X)}~, \] again after simplification via law of iterated expectations. As we may choose \[l'_*(S^* \mid X) = f_2(S^*, X) - \E_{P_\star}[f_2(S^*, X) \mid X]~,\] we find similarly that $f_2$ is a function only of $X$.

Plugging (ref) into (ref), gives the condition \[ 0 = \E_P\bk{(1-\gamma(X)) (f_2(X) - f_4(X)) \varrho'(X)} \] where picking $\varrho'(X) = f_2(X) - f_4(X)$ implies that $f_4 (x) = f_2(x)$.

Lastly, plugging (ref) into (ref), we find \[ 0 = \E_P\bk{ l'_*(G, X) (G f_3(X) + (1-G) f_2(X)) }. \] We then find that that \[ \E_P[l'_*(G, X)] = 0 = \E_P[\gamma(X) l'_*(1,X) + (1-\gamma(X)) l'_*(0,X)]~, \] and so $f_1 (x) = f_2(x) = f_3(x) = f_4(x) = C$ almost everywhere, for some constant $C$. Hence, $f(b) = 0$ almost everywhere and the space of influence functions is a singleton.\qed

Proof of (ref). Part 2

We demonstrate that $\xi_{1,1}$ is the unique influence function for $\theta_{1,1}$. An analogous argument verifies the statement for $\theta_{0,1}$, and therefore for $\tau_1$. We maintain the notation from the proof of (ref), Part 2.

Let $\tilde{\xi}(b_1)$ be an influence function for $\theta_{1,1}$. It is necessarily the case that \[ \tilde{\xi}_{1,1}(b_1) = \xi_{1,1}(b_1) + f(b_1)~, \] where $\xi_{1,1}(b_1)$ is the efficient influence function derived in (ref), Part 2, and $f(b_1)$ is mean-zero and orthogonal to the tangent space $\mathcal{T}$, defined in the proof of (ref), Part 2. We may without loss of generality write \[ f(b) = g \cdot f_1(y,s,x) + (1-g) w \cdot f_2(s,x) + (1-g)(1-w) \cdot f_3(x) + A~, \] where $A$ is some constant making $f$ mean-zero.

The orthogonality between $f$ and the tangent space $\mathcal{T}$ implies that

equation[equation omitted — 57 chars of source]

for any $l'(\cdot)$ corresponding to a score of a parametric submodel.

We first argue that $f_1(y,s,x) = f_1(x)$ does not depend on $y$ or $s$. Consider all free components in the score (ref), and set everything except for $l'(y\mid s, x, G=1)$ to zero. Then the orthogonality condition (ref) becomes \[ \E_P[G f_1(Y,S,X) l'(Y \mid S, X, G=1)] = 0~. \] Picking $l'(Y\mid S, X, G=1) = f_1(Y,S,X) - \E_P[f_1(Y) \mid S, X, G=1]$ implies that $\E_P[G f_1(Y,S,X) l'(Y \mid S, X, G=1)] > 0$ unless $f_1(Y,S,X) = \E[f_1(Y) \mid S, X, G=1]$ almost surely. Therefore $f_1(Y,S,X)$ does not depend on $Y$. Similarly, if we set every free component in (ref) except for $l'(S \mid X, G=1)$ to zero, then choosing $l'(S\mid X, G=1) = f_1 (S,X) - \E[f(S,X) \mid X, G=1]$ again implies that $f_1$ only depends on $X$. A similar argument with the score component $l'(S \mid W=1, X, G=0)$ shows that $f_2(s,x)=f_2(x)$ depends only on $x$. Therefore, $f(b)$ must take the form \[ f(b) = g\cdot f_1(x) + (1-g)w \cdot f_2(x) + (1-g)(1-w) \cdot f_3(x). \]

The orthogonality condition (ref) then takes the form

align*[align* omitted — 224 chars of source]

for any valid choices of $l'(x), \gamma'(x), \rho'(x)$, as the other terms are zero by the mean-zero conditions on the conditional scores.

Observe that the third term can be simplified as

align[align omitted — 222 chars of source]

Picking $\varrho'(x) = f_2(x) - f_3(x), l'(x) = \gamma'(x) = 0$ implies that $f_2 (x) = f_3(x)$. Thus $f(b) = gf_1(x) + (1-g) f_2(x)$. Working with the second term in (ref) (i.e. setting $l'(x) = 0 = \rho'(x)$ and considering choices of $\gamma'(x)$) yields that \[ \E_P\bk{ f(B)\frac{G-\gamma(X)}{\gamma(X) (1-\gamma(X))}\gamma'(X) } = \E_P[(f_1(X) - f_2(X)) \gamma'(X)] \] Picking $\gamma'(x) = f_1(x) - f_2(x)$ shows that $f_1(x) = f_2(x)$. Therefore, from the first term in (ref), we find that for all $l'$ s.t. $\E_P[l'(X)] = 0$, \[ \E_P[f(B_1) l'(X)] = \E_P[f_1(X) l'(X)]= 0 \implies f_1(X) = 0. \] implies that $f_1(X) = 0$ almost everywhere. Hence, $f(b_1) = 0$ almost everywhere, completing the proof. \qed

Proof of (ref)

In this section, we provide proofs for both (ref) and for the analogous result for the statistical surrogacy model, stated as follows.

theoremLet $\mathcal{P}\subset\mathcal{M}_\lambda$ be the set of all probability distributions $P$ for $\{B_i\}_{i=1}^n$ that satisfy the Statistical Surrogacy Model stated in (ref) in addition to (ref). If (ref) holds for $\mathcal P$, then \begin{align} \sqrt{n}(\hat{\tau}_{1,\mathsf{DML}} -\tau_1) \overset{d}{\to} \mathcal{N}(0,V_1^{\star\star}) \end{align} uniformly over $P\in\mathcal{P}$, where $\hat{\tau}_{1,\mathsf{DML}}$ is defined in (ref), $V_1^{\star\star}$ is defined in (ref), and $\overset{d}{\to}$ denotes convergence in distribution. Moreover, we have that \begin{align} \hat{V}_1^{\star\star} = \frac{1}{k}\sum_{l=1}^k \frac{1}{m} \sum_{i\in I_l} \left(\xi_1(B_i,\hat{\tau}_{1,\mathsf{DML}},\hat{\varphi}(I_l^c))\right)^2 \overset{p}{\to} V_1^{\star\star} \end{align} uniformly over $P\in\mathcal{P}$, where $\overset{p}{\to}$ denotes convergence in probability. As a result, we obtain the uniform asymptotic validity of the confidence intervals \begin{align} \lim_{n\to\infty} \sup_{P\in\mathcal{P}} \Big\vert P\left(\tau_1 \in \left[\hat{\tau}_{1,\mathsf{DML}} \pm z_{1-\alpha/2}\sqrt{\hat{V}_1^{\star\star}/n} \right]\right) - (1-\alpha) \Big\vert = 0 , \end{align} where $z_{1-\alpha/2}$ is the $1-\alpha/2$ quantile of the standard normal distribution.

To ease notation for the purposed of this proof, we change the notation introduced in (ref). In particular, partition the collection of nuisance functions $\eta$ appearing as an argument in $\psi_1(B,\tau_1,\eta)$ into long-term outcome means and propensity scores by $\eta=(\omega,\kappa)$, where we include the parameter $\pi$ in the propensity scores $\kappa$ to ease notation. Similarly partition the collection of nuisance function $\varphi$ appearing as an argument in $\xi_1(B,\tau_1,\varphi)$ into $\varphi=(\vartheta,\zeta)$.

Observe that the efficient influence functions $\psi_1(b,\tau_1,\eta)$ and $\xi_1(b,\tau_1,\varphi)$ can be decomposed linearly into \[ \psi_1(b,\tau_1,\eta) = \psi_{1}^\alpha(b,\eta)\cdot\tau_1 + \psi_{1}^\beta(b,\eta) \quad\text{and}\quad \xi_1(b,\tau_1,\varphi = \xi_{1}^\alpha(b,\varphi)\cdot\tau_1 + \xi_{1}^\beta(b,\varphi) \] where

align*[align* omitted — 889 chars of source]

Consequently, it will suffice to verify the conditions of Theorem 3.1 of chernozhukov2018double. In particular, it will suffice to verify Assumption 3.1 and Assumption 3.2 therein.

Verification of Assumption 3.1 of chernozhukov2018double. Both $\psi_1(b,\tau_1,\eta)$ and $\xi_1(b,\tau_1,\varphi)$ are mean-zero, satisfying condition (a). By the above decomposition, both $\psi_1(b,\tau_1,\eta)$ and $\xi_1(b,\tau_1,\varphi)$ are linear in $\tau_1$ satisfying condition (b). It is clear that the mappings $\eta \mapsto \mathbb{E}_P\left[\psi_1(b,\tau_1,\eta)\right]$ and $\varphi \mapsto \mathbb{E}_P\left[\xi_1(b,\tau_1,\varphi)\right]$ are twice continuously Gateaux differentiable, as they are both smooth functions by the strict-overlap (ref), satisfying condition (c). Both $\psi_1(b,\tau_1,\eta)$ and $\xi_1(b,\tau_1,\varphi)$ satisfy the “Neyman orthogonality” condition of chernozhukov2018double by the orthogonality conditions (\refeq{eq: pathwise deriv}) and (\refeq{acik:pathwise}), satisfying condition (d). Finally, as \[ \mathbb{E}_P\left[ -\frac{g}{\pi} \right] = -1~, \] the singular values of $\mathbb{E}_P\left[ \psi_1(b,\tau_1,\eta) \right]$ and $\mathbb{E}_P\left[\xi_1(b,\tau_1,\varphi)) \right]$ are bounded away from zero, satisfying condition (e).

Verification of Assumption 3.2 of chernozhukov2018double. We note that (ref) can be restated as: For all $P \in \mathcal P$, with probability at least $1-\Delta_n$, the nuisance estimator $\hat \eta_n$ (resp. $\hat \varphi_n$) belongs to a realization set $\mathcal{R}_{P,n}$. The realization set $ \mathcal{R}_{P,n}$ contains nuisance parameter values $\tilde \eta = (\tilde \omega, \tilde \kappa)$ such that

enumerate$\| \tilde{\eta} - \eta \|_{P,q} \leq C$, • $\| \tilde{\eta} - \eta \|_{P,2} \leq \delta_n$, $\| \tilde{\kappa} - 1/2 \|_{P,\infty} \leq 1/2 - \epsilon$, and • $\| \tilde{\omega} - \omega \|_{P,2} \cdot \| \tilde{\kappa} - \kappa \|_{P,2} \leq \delta_n / n^{1/2}$

Define the realization set $\mathcal{R}_{P,n}$ analogously for the Statistical Surrogacy Model, where the nuisance function $\varphi$ replaces $\eta$.

Condition (a). The nuisance parameter estimators are elements of their respective realization sets with probability $1-\Delta_N$ by (ref), verifying condition (a).

Condition (b). To verify condition (b), in the Latent Unconfounded Treatment Model, we must demonstrate that

flalign*\sup_{\eta\in\mathcal{R}_{P,n}}\left(\mathbb{E}_{P}\left[\|\psi_{1}^{\alpha}\left(B;\eta\right)\|^{q}\right]\right)^{1/q} & \leq C^{\prime}, and\\ \sup_{\eta\in\mathcal{R}_{P,n}}\left(\mathbb{E}_{P}\left[\|\psi_{1}\left(B;\tau_{1},\eta\right)\|^{q}\right]\right)^{1/q} & \leq C^{\prime}

for some $C^{\prime}$. Analogous conditions suffice for the Statistical Surrogacy Model. Verification of the first inequality is immediate, as \[ \sup_{\eta\in\mathcal{R}_{P,n}}\left(\mathbb{E}_{P}\left[\|\psi_{1}^{\alpha}\left(B;\eta\right)\|^{q}\right]\right)^{1/q}=1 \quad\text{and}\quad \sup_{\varphi\in\mathcal{R}_{P,n}}\left(\mathbb{E}_{P}\left[\|\xi_{1}^{a}\left(B;\varphi\right)\|^{q}\right]\right)^{1/q}=1~. \] For the Latent Unconfounded Treatment Model, by the equalities \[ \mu_{w}\left(s,x\right)=\mathbb{E}_{P_{\star}}\left[Y\left(w\right)\mid S\left(w\right)=s,X=x\right]\quad\text{and}\quad\bar{\mu}_{w}\left(X\right)=\mathbb{E}_{P_{\star}}\left[Y\left(w\right)\mid X=x\right]~, \] we have that

alignat*{1} \|\mu_{w}\left(S,X\right)\|_{P,q}\vee\|\bar{\mu}_{w}\left(X\right)\|_{P,q} & \leq\|Y\left(w\right)\|_{P,q}\leq C

by Jensen's inequality and (ref), where $\vee$ denotes the join binary operation. Thus, for any $\tilde{\eta}\in\mathcal{R}_{P,n}$, we have that

flalign*\|\psi_{1}\left(B;\tau_{1},\tilde{\eta}\right)\|_{P,q} & \leq C_{1}( \|\tilde{\mu}_{1}\left(S,X\right)\|_{P,q}+\|\tilde{\mu}_{0}\left(S,X\right)\|_{P,q}\\ &\quad+\|\tilde{\bar{\mu}}_{1}\left(X\right)\|_{P,q}+\|\tilde{\bar{\mu}}_{0}\left(X\right)\|_{P,q} +\|Y\|_{P,q}+\vert\tau_{1}\vert)\\ & \leq C_{1}(\|\mu_{1}\left(S,X\right)-\tilde{\mu}_{1}\left(S,X\right)\|_{P,q} +\|\tilde{\mu}_{0}\left(S,X\right)-\tilde{\mu}_{0}\left(S,X\right)\|_{P,q}\\ &\quad+\|\mu_{1}\left(S,X\right)\|_{P,q}+\|\mu_{0}\left(S,X\right)\|_{P,q} + \|\bar{\mu}_{1}\left(X\right)-\tilde{\bar{\mu}}_{1}\left(X\right)\|_{P,q}\\ &\quad +\|\bar{\mu}_{1}\left(X\right)-\tilde{\bar{\mu}}_{0}\left(X\right)\|_{P,q} +\|\tilde{\bar{\mu}}_{1}\left(X\right)\|_{P,q}+\|\tilde{\bar{\mu}}_{0}\left(X\right)\|_{P,q}\\ &\quad +\|Y\|_{P,q}+\vert\tau_{1}\vert) \leq C_{2},

where the constants $C_{1}$ and $C_{2}$ depend only on $C$ and $\epsilon$, by the (ref), the triangle inequality, and (ref).

For the Statistical Surrogacy Model, by the equalities \[ \nu\left(s,x\right)=\mathbb{E}_{P}\left[Y\mid S=x,X=x,G=1\right]\quad\text{and}\quad\bar{\nu}_{w}\left(X\right)=\mathbb{E}_{P_{\star}}\left[Y\left(w\right)\mid X=x,W=w,G=0\right] \] and the bounds

align*[align* omitted — 348 chars of source]

again we have that

flalign*\|\nu\left(S,X\right)\|_{P,q} & \leq \epsilon^{-1/q} \|Y\|_{P,q}\leq2\epsilon^{-1/q}\left(1-\epsilon\right)C \quadand\quad \|\bar{\nu}_{w}\left(X\right)\|_{P,q}\le\epsilon^{-2/q}C

by Jensen's inequality, (ref), and (ref). Thus, for any $\tilde{\varphi}\in\mathcal{R}_{P,n}$, we have by an analogous argument that

flalign*\|\xi_{1}\left(B;\tau_{1},\tilde{\varphi}\right)\|_{P,q} & \leq C_{4},

where $C_{4}$ depends only on $C$, $\epsilon$, and $q$, by Assumption 3.1, the triangle inequality, and (ref).

Condition (c). To verify condition (c), in the Latent Unconfounded Treatment Model, we must demonstrate that

flalign\sup_{\tilde{\eta}\in\mathcal{R}_{P,n}}\|E_{P}\left[\psi_{1}^{\alpha}\left(B;\tilde{\eta}\right)-\psi_{1}^{\alpha}\left(B;\eta\right)\right]\|_{P,2} & \leq\delta^\prime_{n},\\ \sup_{\tilde{\eta}\in\mathcal{R}_{P,n}}\left(E_{P}\left[\psi_{1}\left(B;\tau_{1},\tilde{\eta}\right)-\psi_{1}\left(B;\tau_{1},\eta\right)\right]^{2}\|^{1/2}\right) & \leq\delta^\prime_{n}, and\\ \sup_{h_{0}\in\left(0,1\right),\tilde{\eta}\in\mathcal{R}_{P,n}} \norm[\bigg]{ \frac{\partial^ {2}}{\partial h^{2}}\mathbb{E}_{P}\left[\psi_{1}\left(B;\tau_{1},\eta+h\left(\tilde{\eta}-\eta\right)\right)\right]\vert_{h=h_{0}} } & \leq\delta^\prime_{n}/\sqrt{n} ,

for some sequence of positive constants $\{\delta^\prime_n\}_{n\geq1}$ converging to zero. Analogous conditions suffice for the Statistical Surrogacy Model. Verification of ((ref)) is immediate as \[ \|E_{P}\left[\psi_{1}^{\alpha}\left(B;\tilde{\eta}\right)-\psi_{1}^{\alpha}\left(B;\eta_{0}\right)\right]\|_{P,2}=\|\pi^{-1}-\tilde{\pi}^{-1}\|_{P,2}\leq\epsilon^{-2}\delta_{n} \] and \[ \|E_{P}\left[\psi_{1}^{a}\left(B;\tilde{\eta}\right)-\psi_{1}^{a}\left(B;\eta_{0}\right)\right]\|=\|\pi^{-1}-\tilde{\pi}^{-1}\|_{P,2}\leq\epsilon^{-2}\delta_{n} \] for any $\tilde{\eta}$ and $\tilde{\varphi}$ in their respective realization sets.

For the Latent Unconfounded Treatment Model, observe that for any $\tilde{\eta}\in\mathcal{R}_{P,n}$, the triangle inequality gives the bound \[ \|\psi_{1}\left(B;\tau_{1},\tilde{\eta}\right)-\psi_{1}\left(B;\tau_{1},\eta_{0}\right)\|_{P,2}\leq\mathcal{I}_{1,1}+\mathcal{I}_{1,0}+\mathcal{I}_{2}+\mathcal{I}_{3}+\mathcal{I}_{4,1}+\mathcal{I}_{4,0} \] where

flalign*\mathcal{I}_{1,1} & = \norm[\bigg]{\frac{G}{\tilde{\pi}}\left(\frac{W\left (Y-\tilde{\mu}_{1}\left(S,X\right)\right)}{\tilde{\rho}_{1}\left(S,X\right)}\right)-\frac{G}{\pi}\left(\frac{W\left(Y-\mu_{1}\left(S,X\right)\right)}{\rho_{1}\left(S,X\right)}\right)}_{P,2},\\ \mathcal{I}_{1,0} & = \norm[\bigg]{ \frac{G}{\tilde{\pi}}\left(\frac{\left (1-W\right)\left(Y-\tilde{\mu}_{0}\left(S,X\right)\right)}{\tilde{\rho}_ {0}\left(S,X\right)}\right)-\frac{g}{\pi}\left(\frac{\left(1-W\right)\left (Y-\mu_{0}\left(S,X\right)\right)}{\rho_{0}\left(S,X\right)}\right) } _{P,2},\\ \mathcal{I}_{2} & = \norm[\bigg]{\frac{G}{\tilde{\pi}}\tilde{\bar{\mu}}_ {1}\left(X\right)-\frac{G}{\pi}\bar{\mu}_{1}\left(x\right)\bigg\|_{P,2}+\bigg\| \frac{G}{\tilde{\pi}}\tilde{\bar{\mu}}_{0}\left(x\right)-\frac{G}{\pi}\bar{\mu}_{0}\left(x\right) } _{P,2},\\ \mathcal{I}_{3} & = \norm[\bigg]{\left(\pi^{-1}-\tilde{\pi}^ {-1}\right)G\tau_{1} } _{P,2},\\ \mathcal{I}_{4,1} & = \bigg\|\frac{G}{\tilde{\pi}}\left(\frac{ \tilde{\gamma}\left(X\right)}{1-\tilde{\gamma}\left(X\right)}\frac{W\left(\tilde{\mu}_{1}\left(S,X\right)-\tilde{\bar{\mu}}_{1}\left(X\right)\right)}{\tilde{\varrho}\left(X\right)}\right)\\ &-\frac{G}{\pi}\left(\frac{\gamma\left(X\right)}{1-\gamma\left(X\right)} \frac{W\left(\mu_{1}\left(S,X\right)-\bar{\mu}_{1}\left(X\right)\right)} {\varrho\left(X\right)}\right)\bigg \|_{P,2}, and\\ \mathcal{I}_{4,0} & = \bigg\|\frac{G}{\tilde{\pi}}\left(\frac{ \tilde{\gamma}\left(X\right)}{1-\tilde{\gamma}\left(X\right)}\frac{\left(1-W\right)\left(\tilde{\mu}_{0}\left(S,X\right)-\tilde{\bar{\mu}}_{0}\left(X\right)\right)}{\tilde{\varrho}\left(X\right)}\right)\\ & -\frac{G}{\pi}\left(\frac{\gamma\left(X\right)}{1-\gamma\left(X\right)} \frac{\left(1-W\right)\left(\mu_{0}\left(S,X\right)-\bar{\mu}_{0}\left (X\right)\right)}{1-\tilde{\varrho}\left(X\right)}\right) \bigg\|_{P,2} .

Observe that

flalign*\mathcal{I}_{1,1} & \leq\epsilon^{-2} \norm[\bigg]{\pi\left(\frac{W\left(Y- \tilde{\mu}_{1}\left(S,X\right)\right)}{\tilde{\rho}_{1}\left(S,X\right)}\right)-\tilde{\pi}\left(\frac{W\left(Y-\mu_{1}\left(S,X\right)\right)}{\rho_{1}\left(S,X\right)}\right)}_{P,2}\\ & \leq\epsilon^{-2} \norm[\bigg]{\frac{W\left(Y-\tilde{\mu}_{1}\left (S,X\right)\right)}{\tilde{\rho}_{1}\left(S,X\right)}-\frac{W\left(Y-\mu_ {1}\left(S,X\right)\right)}{\rho_{1}\left(S,X\right)}}_{P,2}+ \norm [\bigg]{\left(\pi-\tilde{\pi}\right)\left(\frac{W\left(Y-\mu_{1}\left (S,X\right)\right)}{\rho_{1}\left(S,X\right)}\right)}_{P,2}\\ & \leq\epsilon^{-4}\left(\|\left(\rho_{1}\left(S,X\right)-\tilde{\rho}_{1}\left(S,X\right)\right)\left(Y-\mu_{1}\left(S,X\right)\right)\|_{P,2}+\|\rho_{1}\left(S,X\right)\left(\mu_{1}\left(S,X\right)-\tilde{\mu}_{1}\left(S,X\right)\right)\|_{P,2}\right)\\ & +\epsilon^{-3} \norm[\bigg]{\left(\pi-\tilde{\pi}\right)\left(Y-\mu_ {1}\left(S,X\right)\right)}_{P,2}\\ & \leq\epsilon^{-4}\left(\sqrt{C}\|\rho_{1}\left(S,X\right)-\tilde{\rho}_{1}\left(S,X\right)\|_{P,2}+\|\mu_{1}\left(S,X\right)-\tilde{\mu}_{1}\left(S,X\right)\|_{P,2}\right)+\epsilon^{-3}\sqrt{C}\|\pi-\tilde{\pi}\|_{P,2}\\ & \leq\epsilon^{-3}\left(\epsilon^{-1}\left(\sqrt{C}+1\right)+\sqrt{C}\right)\delta_{n},

where the third inequality follows from $\mathbb{E}\left[\sigma^{2}\left(S,X\right)\mid X\right]\leq C$ and $\rho_{1}\left(S,X\right)\leq1$. An analogous argument verifies that $\mathcal{I}_{1,0}\leq\epsilon^{-3}\left(\epsilon^{-1}\left(\sqrt{C}+1\right)+\sqrt{C}\right)\delta_{n}$. Likewise, we can see that

flalign*\mathcal{I}_{2} & \leq\epsilon^{-2}\left(\|\tilde{\bar{\mu}}_{1}\left(X\right)-\bar{\mu}_{1}\left(x\right)\|_{P,2}+\|\tilde{\bar{\mu}}_{0}\left(X\right)-\bar{\mu}_{0}\left(X\right)\|_{P,2}\right)\leq2\epsilon^{-2}\delta_{n}

and that \[ \mathcal{I}_{3}\leq\epsilon^{-2}\|\left(\tilde{\pi}-\pi\right)\tau_{1}\|_{P,2}\leq\epsilon^{-2}2C\delta_{n}. \] Finally, if we define $\upsilon\left(X\right)$ by \[ \frac{1}{\upsilon\left(X\right)}=\frac{\gamma\left(X\right)}{\pi\left(1-\gamma\left(X\right)\right)\varrho\left(X\right)}, \] and observe that $\upsilon\left(X\right)\geq\epsilon^{3}/\left(1-\epsilon\right)=\epsilon_{\star}$, we have that

flalign*\mathcal{I}_{4,1} & \leq\epsilon_{\star}^{-2}\|\upsilon\left(X\right)\left(\tilde{\mu}_{1}\left(S,X\right)-\tilde{\bar{\mu}}_{1}\left(X\right)\right)-\tilde{\upsilon}\left(X\right)\left(\mu_{1}\left(S,X\right)-\bar{\mu}_{1}\left(X\right)\right)\|_{P,2}\\ & \leq\epsilon_{\star}^{-2}\|\upsilon\left(X\right)\left(\tilde{\mu}_{1}\left(S,X\right)-\mu_{1}\left(S,X\right)+\mu_{1}\left(S,X\right)-\bar{\mu}_{1}\left(X\right)+\bar{\mu}_{1}\left(X\right)-\tilde{\bar{\mu}}_{1}\left(X\right)\right)\\ & \quad\quad-\tilde{\upsilon}\left(X\right)\left(\mu_{1}\left(S,X\right)-\bar{\mu}_{1}\left(X\right)\right)\|_{P,2}\\ & \leq\epsilon_{\star}^{-4}\Big(\|\upsilon\left(X\right)\left(\tilde{\mu}_{1}\left(S,X\right)-\mu_{1}\left(S,X\right)\right)\|_{P,2}+\|\left(\upsilon\left(X\right)-\tilde{\upsilon}\left(X\right)\right)\left(\mu_{1}\left(S,X\right)-\bar{\mu}_{1}\left(X\right)\right)\|_{P,2}\\ & \quad\quad+\|\upsilon\left(X\right)\left(\bar{\mu}_{1}\left(X\right)-\tilde{\bar{\mu}}_{1}\left(X\right)\right)\|_{P,2}\Big)\leq C_{5}\delta_{n} ,

where $C_{5}$ depends only on $\epsilon$ and $C$, with an analogous argument giving $\mathcal{I}_{4,1}\leq C_{5}\delta_{N}$, verifying ((ref)) for the Latent Unconfounded Treatment Model.

In turn, for the Statistical Surrogacy Model, observe that for any $\tilde{\varphi}\in\mathcal{R}_{P,n}$, the triangle inequality gives the bound \[ \|\xi_{1}\left(B;\tau_{1},\tilde{\varphi}\right)-\xi_{1}\left(B;\tau_{1},\varphi_{0}\right)\|_{P,2}\leq\mathcal{J}_{1,1}+\mathcal{J}_{1,0}+\mathcal{J}_{2}+\mathcal{J}_{3}+\mathcal{J}_{4,1}+\mathcal{J}_{4,0} \] where

flalign*\mathcal{J}_{1,1} & =\bigg \|\frac{G}{\tilde{\pi}}\left(\frac{\tilde{\gamma}\left(X\right)}{\tilde{\gamma}\left(S,X\right)}\frac{1-\tilde{\gamma}\left(X\right)}{1-\tilde{\gamma}\left(S,X\right)}\frac{\tilde{\varrho}\left(S,X\right)\left(Y-\tilde{\nu}\left(S,X\right)\right)}{\tilde{\varrho}\left(X\right)}\right)\\ & \quad-\frac{G}{\pi}\left(\frac{\gamma\left(X\right)}{\gamma\left (S,X\right)}\frac{1-\gamma\left(X\right)}{1-\gamma\left(S,X\right)} \frac{\varrho\left(S,X\right)\left(Y-\nu\left(S,X\right)\right)} {\varrho\left(X\right)}\right)\bigg \|_{P,2},\\ \mathcal{J}_{1,0} & =\bigg \|\frac{G}{\tilde{\pi}}\left(\frac{\tilde{\gamma}\left(X\right)}{\tilde{\gamma}\left(S,X\right)}\frac{1-\tilde{\gamma}\left(X\right)}{1-\tilde{\gamma}\left(S,X\right)}\frac{\left(1-\tilde{\varrho}\left(S,X\right)\right)\left(Y-\tilde{\nu}\left(S,X\right)\right)}{1-\tilde{\varrho}\left(X\right)}\right),\\ & \quad-\frac{G}{\pi}\left(\frac{\gamma\left(X\right)}{\gamma\left (S,X\right)}\frac{1-\gamma\left(X\right)}{1-\gamma\left(S,X\right)} \frac{\left(1-\varrho\left(S,X\right)\right)\left(y-\nu\left(S,X\right)\right)}{1-\varrho\left(X\right)}\right)\bigg\|_{P,2}\\ \mathcal{J}_{2} & =\bigg \|\frac{G}{\tilde{\pi}}\tilde{\bar{\nu}}_{1}\left (X\right)-\frac{G}{\pi}\bar{\nu}_{1}\left(x\right) \bigg{\|}_{P,2}+\bigg{\|}\frac{G}{ \tilde{\pi}}\tilde{\bar{\nu}}_{0}\left(x\right)-\frac{G}{\pi}\bar{\nu}_{0}\left(x\right)\bigg\|_{P,2},\\ \mathcal{J}_{3} & =\bigg \|\left(\pi^{-1}-\tilde{\pi}^{-1}\right)G\tau_ {1}\bigg\|_{P,2},\\ \mathcal{J}_{4,1} & =\bigg \|\frac{G}{\tilde{\pi}}\left(\frac{W\left( \tilde{\nu}_{1}\left(S,X\right)-\tilde{\bar{\nu}}_{1}\left(X\right)\right)} {\tilde{\varrho}\left(X\right)}\right)-\frac{G}{\pi}\left(\frac{W\left(\nu_{1}\left(S,X\right)-\bar{\nu}_{1}\left(X\right)\right)}{\varrho\left(X\right)}\right)\bigg\|_{P,2}, and\\ \mathcal{J}_{4,0} & =\bigg \|\frac{G}{\tilde{\pi}}\left(\frac{\left (1-W\right)\left(\tilde{\nu}_{0}\left(S,X\right)-\tilde{\bar{\nu}}_{0}\left (X\right)\right)}{\tilde{\varrho}\left(X\right)}\right)-\frac{G}{\pi}\left( \frac{\left(1-W\right)\left(\nu_{0}\left(S,X\right)-\bar{\nu}_{0}\left(X\right)\right)}{1-\tilde{\varrho}\left(X\right)}\right) \bigg\|_{P,2}.

Bounds for each term follow similar arguments to their analogues for the case where treatment is observed in the observational sample. We omit their explicit derivation to avoid repetition, concluding the verification of ((ref)).

The expressions for the first and second Gateaux derivatives \[ \frac{\partial}{\partial h}\mathbb{E}_{p}\left[\psi_{1}\left(B;\tau_{1},\eta+h\left(\tilde{\eta}-\eta\right)\right)\right]\quad\text{and} \quad\frac{\partial^2}{\partial h^2}\mathbb{E}_{p}\left[\psi_{1}\left(B;\tau_{1},\eta+h\left(\tilde{\eta}-\eta\right)\right)\right] \] as well as \[ \frac{\partial}{\partial h}\mathbb{E}_{p}\left[\xi_{1}\left(B;\tau_{1},\varphi+h\left(\tilde{\varphi}-\varphi\right)\right)\right] \quad\text{and}\quad \frac{\partial^2}{\partial h^2}\mathbb{E}_{p}\left[\xi_{1}\left(B;\tau_{1},\varphi+h\left(\tilde{\varphi}-\varphi\right)\right)\right] \] are very tedious. In particular, the second Gateaux derivatives each take more than a page to display in the present formatting. They are available upon request, but are omitted for the sake of clarity. The second Gateaux derivatives can be written as a sum of several terms, each upper bounded, by Hölder's inequality, by the product between the $L_{2}\left(P\right)$ norms of the differences of components of $\tilde{\omega}-\omega$ and $\tilde{\kappa}-\kappa$, if treatment is observed in the observational sample, and $\tilde{\vartheta}-\vartheta$ and $\tilde{\zeta}-\zeta$, if treatment is not observed in the observational sample. Thus, by (ref), ((ref)) is satisfied in each case, completing the verification of condition (c). We refer the reader to the final passage in the proof of Theorem 5.1 of chernozhukov2018double for an example of an argument with an identical structure for the substantially simpler context of the average treatment effect under ignorability.

Condition (d). Finally, to verify condition (d), we must demonstrate that $V_{1}^{\star}$ and $V_{1}^{\star\star}$ are non degenerate. Observe that

flalign*V_{1}^{\star} & =\mathbb{E}_{P}\left[\frac{\gamma\left(X\right)}{\pi}\left(\frac{\sigma_{1}^{2}\left(S,X\right)}{\rho_{1}\left(S,X\right)}+\frac{\sigma_{0}^{2}\left(S,X\right)}{\rho_{0}\left(S,X\right)}+\left(\bar{\mu}_{1}\left(X\right)+\bar{\mu}_{0}\left(X\right)-\tau_{1}\right)^{2}+\Gamma_{0}\left(S,X\right)+\Gamma_{1}\left(S,X\right)\right)\right]\\ & \geq\frac{\epsilon}{\pi\left(1-\epsilon\right)}\mathbb{E}_{P}\left[\sigma_{1}^{2}\left(S,X\right)+\sigma_{0}^{2}\left(S,X\right)\right]\\ & +\frac{\epsilon^{2}}{\pi\left(1-\epsilon\right)^{2}}\mathbb{E}_{P}\left[\left(\mu_{1}\left(S,X\right)-\bar{\mu}_{1}\left(X\right)\right)^{2}+\left(\mu_{0}\left(S,X\right)-\bar{\mu}_{0}\left(X\right)\right)^{2}\right]\\ & \geq2\left(\frac{\epsilon}{\pi\left(1-\epsilon\right)}+\frac{\epsilon^{2}}{\pi\left(1-\epsilon\right)^{2}}\right)c,

as required. An analogous argument holds for $V_{1}^{\star\star}$, completing the proof.\qed

Proof of (ref)

We begin by giving an explicit construction for the estimator \[ \hat{\theta}_{1,1}(g_{\mathsf{w}}) = \frac{1}{n} \sum_{i=1}^n \frac{G_i}{\hat \pi_n} \frac{W_i Y_i}{\hat{\rho}_{1n}(S_i,X_i)} \] for the parameter $\theta_{1,1} = \E[Y(1) \mid G=1]$. We then state sufficient conditions for its asymptotic linearity and Gaussianity. The construction and sufficient conditions for the corresponding estimator of the parameter $\theta_{1,0} =\E[Y(0) \mid G=1]$ are analogous. To ease notation, we drop the subscript $\mathsf{w}$ from the parameter $\zeta_{\mathsf{w}}$. We will use the standard empirical process notation $\Pn[n] f = \frac{1}{n} \sum_{i=1}^n f(B_i)$, $\Pn[] f = \E_P[f(B_i)]$, and $\Gn = \sqrt{n}(\Pn[n] - \Pn[])$, where expectations are taken over $B_i$ and not over $f$ if $f$ is random. Our assumptions and argument closely follow chen2015sieve.

Observe that \[ \rho_1(s,x) = \frac{\zeta_{1}(s,x)}{1-\zeta_{1}(s,x)} \zeta_{2}(x) \frac{1-\zeta_{3}(x)}{\zeta_{3}(x)}~, \numberthis \label{eq:rho_1_def ap} \] where

align*[align* omitted — 158 chars of source]

The parameter $\zeta = (\zeta_{1}, \zeta_{2}, \zeta_{3})$ can be estimated with the sieve extremum estimator

align*[align* omitted — 212 chars of source]

where $\mathcal H_n$ is a sieve approximation of the parameter space for $\zeta \in \mathcal H$. The estimator $\hat \rho_{1n}$ is then derived from plugging $\hat \zeta_n$ into (ref). The parametric nuisance parameter $\pi$ can be estimated by its sample analogue $\hat \pi_n = \frac{1}{n} \sum_{i=1}^n G_i.$

We begin by imposing a set of regularity conditions on the data generating distribution $P$, the estimator $\hat\zeta_{n}$, and the parameter space $\mathcal H$.

assumptionAssume that \begin{enumerate} • The parameter $\theta_{1,1} \in \mathrm{int}(\Theta) \subset \R$, where $\Theta$ is compact. • The nuisance parameters $\pi$ and $\zeta$ are bounded between $\varepsilon$ and $1-\varepsilon$, $\lambda$-almost surely, for some fixed constant $0<\varepsilon<1/2$. • The long-term outcome $Y_i$ has a finite fourth moment, i.e., $\E_P[|Y_i|^4] < \infty$. • The estimator $\hat\zeta_n$ is $n^{-1/4}$-consistent in $\|\cdot\|_\infty$, where $\|\zeta\|_\infty = \max_j \|\zeta_j\|_\infty$, in the sense that there exists a sequence $0 \le \delta_n = o(n^{-1/2})$ such that \[ \delta_n^{-1} \norm{\hat \zeta_n - \zeta}^2_\infty \overset{p}{\to} 0 \] and entries of $\hat \zeta_{n}(\cdot)$ belong to $[\epsilon, 1-\epsilon]$, $\lambda$-almost surely. • The parameter space $\mathcal H = \mathcal H_1 \times \mathcal H_2 \times \mathcal H_3$ is a product of Donsker classes. \end{enumerate}

Next, we introduce more structure on the sieve space $\mathcal H_n$. We assume that $\mathcal H_n$ is dense in $\mathcal H$ in $\norm{\cdot}_\infty$ as $n \to \infty$. Define

align*[align* omitted — 454 chars of source]

as the pathwise derivative of $\theta_{1,1}$ in $\zeta$ in the direction of $v = (v_1, v_2, v_3)$. Similarly, define the inner product \[ \ip{v,u} = \E_P[W(v_1 u_1 + v_3 u_3) + G v_2 u_2] = \diff{^2}{t_1 \partial t_2} \E_P[\L (B, \zeta+t_1 v + t_2 u)]\evalbar_{t_1=t_2=0} \numberthis \label{eq:inner_product_def} \] as the cross-derivative of the first-step criterion function $L$ and let $\norm{\cdot}$ denote the norm induced by $\ip{\cdot, \cdot}$. Define \[\mathcal V = \mathrm{CL}\pr{\br{h - \zeta = (h_1 - \zeta_1, h_2 - \zeta_2, h_3 - \zeta_3) : h \in \mathcal H}}\] as the closed linear span of the re-centered parameter space under $\norm{\cdot}$.\footnote{$ \mathrm{CL}$ takes the closed linear span under $\norm{\cdot}$.} Note that $\bm{\Gamma} : \mathcal V \to \R$ is a bounded linear functional on $\mathcal V$. Define $v^* \in \mathcal V$ as the Riesz representer of $\bm{\Gamma}$ in $\ip{\cdot, \cdot}$: That is, for all $w \in \mathcal V$, we have that \[ \ip{v^*, w} = \Gamma[w]. \] Similarly, let $\zeta_n = \argmin_{h \in \mathcal H_n} \norm{h- \zeta}$ be the projection of $\zeta$ to the sieve space in $\norm{\cdot}$. Define \[ \mathcal V_n = \mathrm{CL}\pr{\br{ h - \zeta_n : h \in \mathcal H_n }} \] as the re-centered sieve space. Lastly, let $v^*_n$ be the Riesz representer in $\mathcal V_n$ of $\bm{\Gamma} [\cdot]$ in $\ip{\cdot,\cdot}$. We refer to $v^*$ as the Riesz representer and $v^*_n$ as the sieve Riesz representer. Note that these objects are defined under the inner product $ \ip{\cdot, \cdot}$, which is different from how Riesz representers are defined in our efficiency calculations.

With this notation in place, we impose the following conditions.

assumption$\text{ }$ \begin{enumerate} • The rate condition $\norm{\hat \zeta_n - \zeta} \norm{v_n^* - v^*} = o_p(n^{-1/2})$ is satisfied. • Recall $\delta_n$ from (ref)(4). Assume that the function classes \[\mathcal F_n = \br{h - \zeta : h \in \mathcal H_n, \norm{h-\zeta}_\infty < \delta_n^{1/2}}\] have finite uniform entropy integrals $J(1, \mathcal F_n)$. The entropy integral $J(1, \mathcal F_n)$ is defined as \[ J(1, \mathcal F_n) = \sup_Q \int_0^1 \sqrt{1 + \log N\pr{ \epsilon \delta_n^{1/2}, \mathcal F_n, L_2(Q)}}\, d \epsilon < C < \infty, \] where $Q$ ranges over all discrete distributions, $N(\epsilon \delta_n^{1/2}, \mathcal F, L_2(Q))$ is a covering number in $L_2(Q)$-norm. We choose the envelope function $F$ for $\mathcal F_n$ as the constant function $F = \delta_n^{1/2}$. \end{enumerate}

(ref) are sufficient for the asymptotic linearity and Gaussianity of $\hat{\theta}_{1,1}(g_{\mathsf{w}})$. The following assumption collects these conditions for reference in the main text.

assumption(ref), in addition to their analogues that pertain to the corresponding weighted estimator of $\theta_{1,0}=\E[Y(0) \mid G=1]$, hold.

Under (ref), we can readily verify Theorem (ref), stated in the main text using Corollary (ref) stated in (ref).

proof[Proof of Theorem (ref)] By an application of (ref) (and its analogue for $\E[Y(0) \mid G=1]$), we have that \[ \sqrt{n} (\hat{\tau}_{1}(g_{\mathsf{w}}) - \tau_1) = \frac{1}{\sqrt{n}} \sum_{i=1}^n \psi_1 (B_i, \tau_1, \eta) + o_p(1). \] where $\psi_1(\cdot)$ is the efficient influence function for $\tau_1$ under the Latent Unconfounded Treatment Model, derived in (ref). As we have assumed that $\E_P[Y_i^4] < \infty$ and by and strict overlap, i.e., (ref), we can readily verify that \[ \E_P[|\psi_1(B_i, \tau_1, \eta)|^{2+\epsilon}] < \infty \] for some $\epsilon > 0$. As a result, by the Lyapunov central limit theorem, \[ \frac{1}{\sqrt{n}} \sum_{i=1}^n \psi_1 (B_i, \tau_1, \eta) \dto \Norm(0, V_1^\star), \] by the definition of $V_1^\star$ given in (ref). The proof is complete with an application of the continuous mapping theorem and Slutsky's theorem.

Asymptotic linearity

Define \[ \Delta_i[v] = -\diff{\L(B_i, \zeta)}{\zeta}[v] = W_i(G_i-\zeta_1) v_1 + G_i(W_i-\zeta_2) v_2 + W_i (G_i-\zeta_3) v_3. \numberthis \label{eq:delta_def} \] as the negative pathwise derivative of the first-stage loss function with respect to $\zeta$ in the direction of $v$, evaluated at the true nuisance $\zeta$. For a test value $\tilde\zeta$ of the nuisance parameter $\zeta$, let an infeasible moment condition be \[ \tilde{g}_1(B_i, \theta_{1,1}, \chi) = \frac{G_iW_iY_i}{\pi \tilde\rho_1(S, X)} - \theta_{1,1}, \numberthis \label{eq:moment_defn_sieve} \] where $\tilde \rho_1$ is derived from plugging $\chi$ into (ref). $\tilde{g}_1$ is a moment condition that assumes knowledge of the true $\pi$.

Let $\tilde\theta_{1,1}$ be the $Z$ estimator based on the moment condition

align[align omitted — 95 chars of source]

The following theorem applies Theorem 2.1 from chen2015sieve to the moment condition (ref). The result is key in that it shows that the influence of the sieve estimation of $\zeta$ on $\theta_{1n}$ is asymptotically linear and is given by $\Delta_i[v^*]$.

theoremUnder (ref), we have that \[ \sqrt{n} (\tilde\theta_{1,1} - \theta_{1,1}) = \frac{1}{\sqrt{n}} \sum_{i=1}^n \tilde{g}_1(B_i, \theta, \zeta) + \Delta_i[v^*] + o_p(1). \]
corUnder (ref), we have that \begin{align} \sqrt{n} (\hat\theta_{1,1} - \theta_{1,1}) &= \frac{1}{\sqrt{n}} \sum_{i=1}^n \bk{\tilde{g}_1(B_i, \theta, \zeta) + \Delta_i[v^*] + \theta_{1,1} - \frac{G_i \theta_{1,1} }{\pi}} + o_p(1) \\ &= \frac{1}{\sqrt{n}} \sum_{i=1}^n \psi_{1,1}(B_i) + o_p(1) , \end{align} where $\psi_{1,1}(\cdot)$ is the efficient influence function for $\theta_{1,1}$ under the Latent Unconfounded Treatment Model, derived in the the proof of (ref).
proof[Proof of (ref)] Note that \begin{align*} \sqrt{n}(\hat\theta_{1,1} - \theta_{1,1} ) &= \sqrt{n} \pr{\frac{\pi}{\hat \pi_n} - 1} \theta_{1,1} + \sqrt{n} (\tilde\theta_{1,1} - \theta_{1,1}) + \sqrt{n} (\tilde\theta_{1,1} - \theta_{1,1}) \pr{\frac{\pi}{\hat \pi_n} - 1} \\ &= \frac{1}{\sqrt{n}} \sum_{i=1}^n \bk{\tilde{g}_1(B_i, \theta_{1,1}, \zeta) + \Delta_i[v^*] + \theta_{1,1} - \frac{G_i\theta_{1,1}}{\pi}} + o_p(1) , \end{align*} where the second equality follows from (ref), stated in (ref), and (ref). This verifies (ref). The equality (ref) follows from (ref), stated in (ref).

Towards verifying (ref), we state the assumptions and main result of chen2015sieve. For that purpose, let $g(\cdot, \theta, \zeta)$ be a generic moment condition, with true parameter values $(\theta, \zeta)$ and generic parameter values $ (\vartheta, h)$. Let $L(\cdot, \cdot)$ a generic first-stage loss function. Define the generic counterparts of other objects analogously, relative to $g$ and to $L$. Let $\bm{\Gamma}_1(\vartheta)$ be the derivative of $\E[g(\cdot, \theta, \zeta)]$ in $\theta$, evaluated at $\vartheta$. Let $\bm{\Gamma}_2(\vartheta)[v]$ be the pathwise derivative of $\E[g (\cdot, \theta, \zeta)]$ in $\zeta$, evaluated at $\vartheta$ and $\zeta$, in the direction $v$. chen2015sieve maintain the following high-level assumptions.

assumption[Assumption A.1 in chen2015sieve] Assume that $\lim_{n \to \infty} \norm{v^*_n} < \infty$. Additionally assume that the estimator $\hat\zeta_n$ satisfies: \begin{enumerate} • $\abs{\frac{1}{n} \sum_{i=1}^n \Delta_i[v^*_n] - \ip{v^*_n, \hat \zeta_n - \zeta}} = o_p(n^{-1/2})$$\norm{\hat \zeta_n - \zeta} \cdot \norm{v^*_n - v^*} = o_p(n^{-1/2})$ \end{enumerate}
assumption[Assumption A.2 in chen2015sieve] Assume $\theta$ is in the interior of $\Theta$ and $\tilde\theta \overset{p}{\to} \theta$. Assume additionally that \begin{enumerate} • The derivative \[\bm{\Gamma}_1(\vartheta) = \diff{\E_P[g (B_i,\vartheta, \zeta)]}{\vartheta}\] exists in a neighborhood of $\theta$ and is continuous at $\theta$. • The pathwise derivative \[ \bm{\Gamma}_2(\vartheta)[v] = \diff{\E_P[g(B_i, \vartheta, \zeta)]}{\zeta}[v] \] exists in all directions $v$ and satisfies \[ \abs{\bm{\Gamma}_2(\vartheta)[v] - \bm{\Gamma}_2(\theta)[v]} \le |\vartheta - \theta| \cdot o(1). \] • We have \[ \abs{ \Pn[n] g(\cdot, \vartheta, \hat\zeta_n) - \Pn[] g(\cdot, \vartheta, \zeta) - \bm{\Gamma}_2(\vartheta)[\hat\zeta_n - \zeta]} = o_p(n^{-1/2}). \] for all $\vartheta = \theta + o(1)$ • For all sequences $0 < \kappa_n = o(1)$, \[ \sup_{\abs{\vartheta - \theta} < \kappa_n, \norm{h-\zeta}_\infty < \kappa_n} \abs{\Gn[n] [g(\cdot, \vartheta, h) - g(\cdot, \theta, \zeta)]} = o_p(1). \] \end{enumerate}
theorem[Theorem 2.1 in chen2015sieve] Let $\tilde\theta_n$ be the estimator based on $g$: That is, \[ \Pn[n] g(\cdot, \tilde\theta_n, \hat\zeta_n) = 0. \] Under (ref), \[ \sqrt{n}(\tilde\theta_n - \theta) = \frac{1}{\sqrt{n}} \sum_{i=1}^n g(B_i, \theta, \zeta) + \Delta_i[v^*]. \]
proof[Proof intuition for (ref)] The influence of the first-step estimation is of the form \[ \sqrt{n} \bm{\Gamma}_2(\theta)[\hat\zeta_n - \zeta] = \sqrt{n} \ip{\hat\zeta_n - \zeta, v^*_k}, \] for which we use (ref). The linearity of $ \ip{\hat\zeta_n - \zeta, v^*_k}$ is assumed in (ref)(1). As $n\to \infty$, $v_k^* \to v^*$, and (ref)(2) controls the corresponding difference between $\ip{v_n^*, \hat \zeta_n - \zeta}$ and $\ip{v^*, \hat \zeta_n - \zeta}$.
proof[Proof of (ref)] It suffices to show that, in our setting, with $g$ chosen as $\tilde g$, (ref) follow from (ref). (ref)(1) follows from (ref), stated in (ref). (ref)(2) is assumed as (ref)(1). Lastly, note that $\mathcal V_n$ consist of functions that are uniformly bounded, since entries of $\zeta$ are functions taking values in $[\varepsilon, 1-\varepsilon]$. Thus, $\lim_{n\to\infty} \norm{v_n^*} \le \limsup_{n\to\infty} \norm{v_n^*}_\infty < \infty$. This verifies (ref). Next, the interiority and consistency in (ref) are assumed by (ref)(1) and shown in (ref). Our moment function is (ref), whose derivatives are particularly simple. The derivative $\bm{\Gamma}_1(\vartheta) = -1$ is a constant, and thus (ref)(1) is satisfied. The derivative $\bm{\Gamma}_2 (\vartheta)[v]$ in our setting does not depend on $\vartheta$, and so (ref)(2) is satisfied. Since $\bm{\Gamma}_2$ does not depend on $\vartheta$ and neither does $\Pn[n] g (\cdot, \vartheta, \hat\zeta_n) - \Pn[] g(\cdot, \vartheta, \zeta)$, we can simplify (ref)(3) into \[ \abs{ \Pn[] \tilde{g}_1(\cdot, \theta_{1,1}, \hat\zeta_n) - \Pn[] \tilde{g}_1(\cdot, \theta_{1,1}, \zeta) - \bm{\Gamma}[\hat\zeta_n - \zeta]} = o_p(n^{-1/2}) \] Consider the functional \[ t \mapsto \Pn[] \tilde{g}_1(\cdot, \theta, \zeta + t (\hat\zeta_n - \zeta)). \] By Taylor's theorem, there exists a $\tilde t \in (0,1)$ such that \[ \Pn[] \tilde{g}_1(\cdot, \theta_{1,1}, \hat\zeta_n) - \Pn[] \tilde{g}_1(\cdot, \theta_{1,1}, \zeta) - \bm{\Gamma}[\hat\zeta_n - \zeta] = \diff{^2\Pn[] \tilde{g}_1(\cdot, \theta, \zeta + t (\hat\zeta_n - \zeta))}{t^2}\evalbar_ {t=\tilde t}. \] It is tedious, but not difficult, to see that the second derivative term is quadratic in the deviation $\hat\zeta_n - \zeta$, and therefore satisfies \[ \abs[\bigg]{\diff{^2\Pn[] \tilde{g}_1(\cdot, \theta_{1,1}, \zeta + t (\hat\zeta_n - \zeta))}{t^2}\evalbar_ {t=\tilde t}} \le C(\epsilon) \E|Y| \cdot \norm{\hat\zeta_n -\zeta}_\infty^2 = o_p(n^ {-1/2}). \] for some $C(\epsilon)$ that depends on $\epsilon$. Thus, it suffices to assume (ref)(2,4). Lastly, the stochastic equicontinuity condition (ref)(4) is implied by the condition that\[ \mathcal G = \br{\tilde{g}_1(\cdot, \vartheta, h) - g(\cdot, \theta_{1,1}, \zeta) : \vartheta \in \Theta, h\in \mathcal H } \] is Donsker. Note that by (ref)(2) and the fact that $\Theta$ is compact, $g(\cdot, \vartheta, h) - g(\cdot, \theta_{1,1}, \zeta)$ is a composition of Lipschitz functions (i.e. addition, multiplication, and division) of $\vartheta, h_1, h_2 ,h_3$. Therefore, since Lipschitz functions of Donsker classes are Donsker (Example 19.20 in van2000asymptotic), it suffices to assume that $\mathcal H$ is a product of Donsker classes ((ref)(5)).

The influence function for $\hat\theta_{1,1}$ is the efficient influence function

In this section, we show that the influence function expression given in (ref) matches the efficient influence function for $\theta_{1,1}$, derived in the the proof of (ref), yielding (ref).

lemmaUnder (ref), we have that \[\psi_{1,1}(B_i) = \tilde{g}_1(B_i, \theta_{1,1}, \zeta) + \Delta_i[v^*] + \theta_{1,1} - \frac{G_i\theta_{1,1}}{\pi}.\]
proofBy inspection of $\psi_{1,1}(B_i)$, it will suffice to show that \[ -\Delta_i[v^*] = \frac{G_i W_i \mu_1(S_i,X_i)}{\pi\rho_1 (S_i,X_i)} - (1-G_i)W_i \frac{1}{\pi \varrho(X_i)} \frac{\gamma(X_i)}{1-\gamma(X_i)} (\mu_1(S_i, X_i) - \bar\mu_1(X_i)) - \frac{G_i}{\pi}\bar\mu_1(X). \numberthis \label{eq:missing_terms} \] It is not difficult to verify that the choices \begin{align*} -v_1^*(S, X) &= \frac{1}{\zeta_1(S, X) \pi} \frac{\gamma(X)}{1-\gamma(X)}\frac{1}{\varrho(S, X)} \mu_1(S,X) = \frac{\mu (S,X)}{\pi \rho_1(S,X)} \frac{1}{1-\zeta_1(S,X)} , \\ -v_2^*(X) &= \frac{\bar \mu_1(X) }{\zeta_2(X) \pi} ,\quadand \\ -v_3^*(X) &= -\frac{1}{\varrho(X)} \frac{\gamma(X)}{1-\gamma(X)} \frac{1}{\zeta_3(X) \pi} \bar\mu_1(X) , \end{align*} ensure that (ref) is satisfied by (ref). These terms can be derived by inspecting terms in (ref) that are of the form $W f(X), W f(S, X), G f(X)$, and matching them with corresponding expressions in (ref). The following identity (ref) is useful for this verification. Note that, by Bayes rule, \[ \zeta_3(x) = \frac{\gamma(x) \zeta_2(x)}{\gamma(x) \zeta_2(x) + (1-\gamma(x))\varrho(x)} \implies \frac{\zeta_3(x)}{1-\zeta_3(x)} = \frac{\gamma(x) \zeta_2(x)}{ (1-\gamma(x)) \varrho(x)} \numberthis \label{eq:ratio} \] and so it is convenient to note that \[ \rho_1(s,x) \frac{1-\zeta_1(s,x)}{\zeta_1(s,x)} = \zeta_2(x) \frac{1-\zeta_3(x)}{\zeta_3(x)} = \varrho(x) \frac{1-\gamma(x)}{\gamma(x)}. \] Lastly, we need to verify that this $v^*$ is actually the Riesz representer for $\bm{\Gamma} [\cdot]$. To do so, we need to check that for any $u = (u_1(s,x), u_2(x), u_3(x))$, \begin{align*} &\E_P[W(v_1^* u_1 + v_3^* u_3) + G v_2^* u_2] \\ &= -\E_P\bk{ \frac{GWY}{\pi \rho_1(S, X)} \pr{ \frac{u_1(S,X)}{\zeta_1(S,X)(1-\zeta_1(S,X))} + \frac{u_2(X)}{\zeta_2(X)} + \frac{v_3(X)}{\zeta_3(X)(1-\zeta_3(X))} } } . \end{align*} The remainder of the proof verifies that \[ \E[W v_1^* u_1] = -\E_P\bk{ \frac{GWY}{\pi \rho_1(S, X)} \frac{u_1(S,X)}{\zeta_1(S,X)(1-\zeta_1(S,X))}. } \] The other two terms can be verified analogously. We first analyze the left-hand side. By law of total probability, we can break the left-hand side into \begin{align*} -\E_P\bk{ W v_1^* u_1 } &= -\pi\E_P[W v_1^* u_1 \mid G=1] -(1-\pi) \E_P[W v_1^* u_1 \mid G=0] \\ where -\pi\E_P[W v_1^* u_1 \mid G=1] &= \pi\E_{P^*}[\rho_1(S(1), X) v_1^*(S(1), X) u_1(S(1), X) \mid G=1] \\ &=\E_{P^*}\bk{\frac{\mu(S(1), X) u_1(S(1), X)}{(1-\zeta_1)} \mid G=1}\end{align*} and where \begin{align*} -(1-\pi) \E_P[W v_1^* u_1 \mid G=0] &= \frac{1-\pi}{\pi} \E_{P^*}\bk{ \varrho(X) \frac{\mu(S(1), X) u_1(S(1), X)}{(1-\zeta_1(S(1), X)) \rho_1(S(1), X)} \mid G=0} \\ &= \E_{P^*}\bk{ \frac{1-\pi}{\pi} \frac{\gamma(X)}{1-\gamma(X)} \frac{1}{\zeta_1(S(1), X)} \mu_1(S(1), X) u_1(S(1), X) \mid G=0 } \\ &= \E_{P^*}\bk{ \frac{1}{\zeta_1(S(1), X)} \mu_1(S(1), X) u_1(S(1), X) \mid G=1 } \end{align*} where the last step follows from the observation that \[ p(s(1), x \mid G=1) = \frac{\gamma(x)}{1-\gamma(x)} \frac{1-\pi}{\pi} p(x \mid G=0) p(s(1) \mid x). \] We can also compute from the right-hand side that \[ \E\bk{\frac{GWY}{\pi \rho_1} \frac{1}{\zeta_1(1-\zeta_1)} u_1} = \E\bk{ \frac{\mu(S(1), X) u_1(S(1), X)}{\zeta_1(1-\zeta_1)} \mid G=1}~. \] Therefore, the terms involving $u_1$ in $\bm{\Gamma}[u]$ and $\ip{v^*, u}$ do equal.

Auxiliary Lemmas

lemmaUnder (ref)(1--3), \[\sqrt{n}(\hat \pi_n / \pi-1) = \frac{1}{\sqrt{n}}\sum_ {i=1}^n \pr{1-\frac{G_i}{\pi}} + o_p(1) \] and $ \hat\theta_{1,1} -\theta = o_p(1). $
proofThe normality of $\hat \pi_n/\pi$ follows from the delta method. The consistency of $\hat\theta_{1n}$ follows from the sup-norm consistency of $\hat \zeta_n$. Both claims rely on (ref)(2) to enforce continuity.
lemmaUnder (ref), the condition (ref)(1) is satisfied.
proofDefine $\epsilon_n = \pm \delta_n$. For some $h - \zeta \in \mathcal F_n$, defined in (ref)(2), let $\tilde h = h + \epsilon_n v_n^*$. We first consider the quantity \[ \sup_{h-\zeta \in \mathcal F_n} \Gn[n] \bk{ L(\cdot, \tilde h) - L(\cdot, h) + \Delta_i[\epsilon_n v_n^*] } \numberthis \label{eq:emp_proc_term_sieve}. \] Note that we can compute \begin{align*} &L(\cdot, \tilde h) - L(\cdot, h) + \Delta_i[\epsilon_n v_n^*] \&= \epsilon_n \pr{W(h_1 - \zeta_1) v_{n1}^* + G(h_2 - \zeta_2) v_{n2}^* + W(h_3 - \zeta_3) v_{n3}^*} + \frac{1}{2}\epsilon_n^2 ({v_{n1}^*}^2 + {v_{n2}^*}^2 + {v_{n3}^*}^2) \\ & \equiv \epsilon_n R_2(\cdot, h, v_n^*) + \epsilon_n^2 R_3(\cdot, v_n^*) \end{align*} Note that $R_3$ does not depend on $h$, and hence \[ \sup_{h - \zeta \in \mathcal F_n} \Gn[n] R_3(\cdot, v_n^*) \le \abs{\Gn[n] R_3(\cdot, v_n^*)} = \sqrt{n}O_p\pr{\E \norm{v_n^*}_2^2} = O_p(\sqrt{n}) \] where we note that $ \E \norm{v_n^*}_2^2 < \infty$ since \[ \infty > \norm{v_n^*}^2 = \E\bk{ \P(W=1 \mid S, X) ({v_{n1}^*}^2 + {v_{n3}^*}^2) + \P(G=1 \mid X) {v_{n2}^*}^2 } \ge \varepsilon \E[{v_{n1}^*}^2 + {v_{n2}^*}^2 + {v_{n3}^*}^2 ]. \] To bound $R_2$, observe that \begin{align*} \norm{R_2(\cdot, h_1, v_n^*) - R_2(\cdot, h_2, v_n^*)}_{L_2(Q)}^2 &\lesssim \norm{v_n^*}_\infty^2 \norm{h_1 - h_2}_{L_2(Q)}^2 \\ &\lesssim \norm{h_1 - h_2}_{L_2(Q)}^2 \end{align*} where the implicit constant does not depend on $Q$. Observe too that for $h - \zeta \in \mathcal F_n$ \[ |R_2(\cdot, h, v_n^*)| \lesssim \delta_n^{1/2} \] uniformly, and thus $C \delta_n^{1/2}$ serves as an envelope function for $\mathcal R_n$, defined below. Hence, letting \[\mathcal R_n = \br{ R_2(\cdot, h, v_n^*) : h - \zeta \in \mathcal F_n },\] we have that \[ \sqrt{1 + \log N\pr{ C_1 \epsilon \delta_n^{1/2}, \mathcal R_n, L_2(Q)}} \le C_2 \sqrt{1 + \log N\pr{ \epsilon \delta_n^{1/2}, \mathcal F_n, L_2(Q)}} < C < \infty. \] Hence, by Theorem 2.14.1 in van1996weak, \[ \E \sup_{h-\zeta \in \mathcal F_n} \abs{\Gn[n] R_2(\cdot, h, v_n^*)} \lesssim \delta_n^ {1/2} \] This implies that \[ \eqref{eq:emp_proc_term_sieve} = O_p(\delta_n^{3/2} + \delta_n^2 \sqrt{n}). \] Now, consider some estimator $\hat \zeta$ and the event that $\norm{\hat\zeta_n - \zeta}_\infty < \delta_n$. On this event, which occurs with probability tending to 1, $\norm{\hat\zeta_n - \zeta} \le C \delta_n$. Let $\tilde \zeta_n = \hat\zeta_n + \epsilon_n v_n^*$. Let $A_n = \one\pr{\norm{\hat\zeta_n - \zeta}_\infty < \delta_n}$. Since $\hat \zeta_n$ minimizes the empirical criterion \begin{align*} 0 &\le A_n \sqrt{n} \Pn[n][\L(\cdot, \tilde\zeta_n) - \L(\cdot, \hat\zeta_n)] \\ &= \sqrt{n} A_n \Pn[]\bk{\L(\cdot, \tilde\zeta_n) - \L(\cdot, \hat\zeta_n)} + A_n \Gn[n] \bk{\L(\cdot, \tilde\zeta_n) - \L(\cdot, \hat\zeta_n)} \\ &= \sqrt{n} \frac{A_n}{2}\pr{\norm{\tilde\zeta_n - \zeta}^2 - \norm{\hat\zeta_n - \zeta}^2} + A_n \Gn[n]\bk{\L(\cdot, \tilde\zeta_n) - \L(\cdot, \hat\zeta_n)} \tag{Note that $\Pn[][\L(\cdot, h) - \L(\cdot, \zeta)] = \frac{1}{2} \norm{h - \zeta}^2$ by definition of $\ip{\cdot,\cdot}$}\\ &\le \sqrt{n} A_n \ip{\hat\zeta_n - \zeta, \epsilon_n v_n^*} + \sqrt{n} \frac{1}{2} \epsilon_n^2 + A_n \Gn[n]\bk{\L(\cdot, \tilde\zeta_n) - \L(\cdot, \hat\zeta_n)} \\ &\le \sqrt{n} A_n \ip{\epsilon_n v_n^*, \hat\zeta_n - \zeta} + \Gn[n] [-\Delta_i[\epsilon_n v_n^*]] + O_p (\delta_n^{3/2} + \delta_n^2 \sqrt{n}), \end{align*} where the last step follows from our bound on (ref). Finally, this implies that \[ \Gn[n] [\Delta_i[\epsilon_n v_n^*]] \le \sqrt{n} A_n \ip{\hat\zeta_n - \zeta, \epsilon_n v_n^*} + O_p (\delta_n^{3/2} + \delta_n^2 \sqrt{n}) \numberthis \label{eq:pos_neg_condition}. \] When $\epsilon_n = \delta_n$, (ref) implies \[ \Gn[n] \Delta_i[v_n^*] \le \sqrt{n} A_n \ip{\hat\zeta_n - \zeta, v_n^*} + O_p (\delta_n^{1/2} + \delta_n \sqrt{n}) = \sqrt{n} A_n\ip{\hat\zeta_n - \zeta, v_n^*} + o_p(1) \] When $\epsilon_n = -\delta_n$, (ref) implies \[ \sqrt{n} A_n \ip{\hat\zeta_n - \zeta, v_n^*} \le \Gn[n] \Delta_i[v_n^*] + o_p(1). \] Thus, taken together \[ |\sqrt{n} A_n \ip{\hat\zeta_n - \zeta, v_n^*} - \Gn[n] \Delta_i[v_n^*]| = o_p(1) \] Note that (i) $\Pn[] \Delta_i[v] = 0$ by the first-order condition of the first-step problem and (ii) $\sqrt{n} A_n \ip{\hat\zeta_n - \zeta, v_n^*} = \sqrt{n} \ip{\hat\zeta_n - \zeta, v_n^*} + o_p(1)$. As a result, we can rewrite the above display as \[ \abs{\ip{\hat\zeta_n - \zeta, v_n^*} - \Pn[n] \Delta_i[v_n^*] } = o_p(n^{-1/2})~, \] completing the proof.

Long-Term Treatment Effect for the Experimental Population

In this section, we develop results analogous to those presented in the main text for long-term average treatment effect for the experimental population, given by

equation[equation omitted — 102 chars of source]

This estimand was considered in athey2020estimating for the Statistical Surrogacy Model. The efficient influence functions and efficiency bounds we state in this section for that context correct those given in Theorem 1 and Theorem 3 of the February, 2020 draft of that paper.

Identification

The identifying assumptions for $\tau_0$ are slightly less restrictive than for $\tau_1$. In particular, in the case that treatment is not measured in the observational data set, it is unnecessary to impose the restriction (ref). Thus, the set of assumptions that compose the Statistical Surrogacy Model will not include (ref) in this section.

prop\begin{enumerate} • athey2020combining Under the Latent Unconfounded Treatment Model, $\tau_0$ is point identified. • athey2020estimating Under the Statistical Surrogacy Model, $\tau_0$ is point identified. \end{enumerate}
proofThe result follows almost immediately from inspection of the proof of (ref) by considering the parameter \[ \theta_{0,1} = \mathbb{E}_{P_\star}[Y(1)\mid G=0]~. \] In the Latent Unconfounded Treatment Model, the only difference will occur in the final step where it is apparent that \[ \mathbb{E}_{P_\star}[\mu_1(X) \mid G=0] \] is identified. In the Statistical Surrogacy Model, the only difference is that (ref) need not be invoked in order to condition on $G=0$, where again the expectation of $\mathbb{E}_{P}[ \mu(S, X) \mid X, W=1, G=0 ]$ conditional on $G=0$ remains identified.

Semiparametric Efficiency

Next, we state a theorem analogous to (ref) for the case where the estimand of interest is the average long-term treatment effect $\tau_0$ in the experimental population.

theorem\begin{enumerate} • Under the Latent Unconfounded Treatment Model, where treatment is observed in the observational data set, the efficient influence function for the parameter $\tau_1$ is given by \begin{align} \psi_0(b,\tau_0,\eta) &= \frac{g}{1-\pi} \left(\frac{1-\gamma(x)}{\gamma(x)}\left( \frac{w(y-\mu_1(s,x))}{\rho_1(s,x)} - \frac{(1-w)(y-\mu_0(s,x))}{\rho_0 (s,x)}\right)\right)\\ & + \frac{1-g}{1-\pi} \left(\frac{w(\mu_1(s,x)-\bar{\mu}_1(x))}{\varrho(x)} - \frac{(1-w)(\mu_0(s,x)-\bar{\mu}_0(x))}{1-\varrho(x)} + (\bar{\mu}_1(x) - \bar{\mu}_0(x)) -\tau_0\right),\nonumber \end{align} where $\eta = (\omega,\kappa)$ collects nuisance functions with \[ \omega = \left\{\mu_w, \bar{\mu}_w\right\}_{w\in\{0,1\}} \quad\text{and}\quad \kappa = \{\{\rho_w\}_{w\in\{0,1\}}, \varrho(\cdot), \gamma(\cdot), \pi\}~. \] collecting long-term outcome means and propensity scores, respectively. • Under the Statistical Surrogacy Model, where treatment is not observed in the observational data set, the efficient influence function for the parameter $\tau_1$ is given by \begin{align} \xi_0(b,\tau_0,\varphi) & = \frac{g}{1-\pi} \left(\frac{1- \gamma(s,x)}{\gamma(s,x)}\frac{(\varrho(s,x)-\varrho(x))(y-\nu(s,x))}{\varrho(x)(1-\varrho(x)}\right) \\ & + \frac{1-g}{1-\pi}\left(\frac{w(\nu(s,x)-\bar{\nu}_1(x))}{\rho(x)} -\frac{(1-w)(\nu(s,x)-\bar{\nu}_0(x))}{1-\rho(x)}+ (\bar{\nu}_1(x) - \bar{\nu}_0(x)) - \tau_0\right) ,\nonumber \end{align} where $\varphi = (\vartheta,\zeta)$ collects nuisance functions with \[ \vartheta = \left\{\nu, \{\bar{\nu}_w\}_{w\in\{0,1\}}\right\} \quad\text{and}\quad \zeta = \{\varrho(\cdot,\cdot), \varrho(\cdot), \gamma(\cdot,\cdot), \gamma(\cdot), \pi\}~. \] collecting long-term outcome means and propensity scores, respectively. \end{enumerate}
corDefine the functionals \begin{align*} \Gamma_{w,0}(s,x) & = \frac{\left(\mu_w(s,x) - \bar{\mu}_w(x)\right)^2}{\varrho(x)^w(1-\varrho(x))^{1-w}} and \Lambda_{w,0}(s,x) = \frac{\left(\nu(s,x) - \bar{\nu}_w(x)\right)^2}{\varrho(x)^w(1-\varrho(x))^{1-w}} . \end{align*} \begin{enumerate} • Under the Latent Unconfounded Treatment Model, the semiparametric efficiency bound for $\tau_0$ is given by \begin{align} V_0^{\star} & = \E_P\Bigg[\frac{1-\gamma(X)}{(1-\pi)^2} \Bigg( \frac{1-\gamma(X)}{\gamma(X)}\left( \frac{\sigma_1^2(S,X)}{\rho_1(S,X)} + \frac{\sigma_0^2(S,X)}{\rho_0(S,X)}\right) \nonumber \\ & \quad \quad\quad \quad \quad \quad +(\bar{\mu}_1(X) - \bar{\mu}_0(X) - \tau_0)^2 + \Gamma_{0,0}(S,X) + \Gamma_{1,0}(S,X)\Bigg) \Bigg] . \end{align} • Under the Statistical Surrogacy Model, the semiparametric efficiency bound for $\tau_0$ is given by \begin{align} V_0^{\star\star} & = \E_P\Bigg[\frac{\gamma(X)}{(1-\pi)^2} \Bigg( \left(\frac{1-\gamma(S,X)}{\gamma(S,X)} \frac{\varrho(S,X)-\varrho(X)}{\varrho(X)(1-\varrho(X))} \right)^2\sigma^2(S,X)\Bigg) \nonumber \\ & \quad \quad + \frac{1-\gamma(X)}{(1-\pi)^2} \Bigg((\bar{\nu}_1(X) - \bar{\nu}_0(X) - \tau_0)^2 + \Lambda_{0,0}(S,X) + \Lambda_{1,0}(S,X)\Bigg) \Bigg] . \end{align} \end{enumerate}
proofWe maintain the same notation and follow the same structure as the proof of (ref), emphasizing differences without repeating shared steps in the argument. We begin by proving Part 1. The analogue of (ref) for $\theta_{1,0}$ is \begin{align} \theta_{1,0}' & = \E_{P_\star}[Y^* l'(Y^* \mid S^*, X) \mid G=0] \\ & + \E_{P_\star}[Y^* l'(S^* \mid X) \mid G=0] + \E_{P_\star}[Y^*l'(X \mid G=0) \mid G=0] , \end{align} with each term now conditioned on $G=0$ and $l'(X \mid G=0)$ replacing $l'(X \mid G=1)$. Note also that \[ l'_*(x \mid G=0) = l'_*(x, G=0) - \E_{P_\star}[l'_*(X, G=0) \mid G=0]~. \] For a candidate influence function $\tilde{\psi}_{1,0}(b_1)$, the analogues of the pathwise differentiability conditions (ref) through (ref) are \begin{align} \frac{1}{1-\pi}\E_{P_\star}[(1-G)Y^* l'_*(Y^*\mid S^*, X)] &= \E_P\bk{\tilde \psi_{1,0}(B_1) GW l'_*(Y\mid S, X)} \\ \frac{1}{1-\pi}\E_{P_\star}[(1-G)Y^* l'_*(S^* \mid X)] &= \E_P\Bigg[\tilde{\psi}_{1,0}(B_1) \Bigg( W l'_*(S \mid X) \nonumber \\ & \quad - G(1-W) \frac{\int \rho(s,X) p'_*(s \mid X)\,d\lambda(s)}{1-\int \rho(s,X)p_\star(s \mid X)\,d\lambda(s)}\Bigg)\Bigg] \\ \E_P[\tilde{\psi}_{1,0}(B_1) l'_*(G, X)] &= \frac{1}{1-\pi}\E_{P_\star}[(1-G)Y^* l'_*(G=0, X) ] \nonumber\\ & \quad - \theta_{1,0} \E_{P_\star}\left[l'_*(G=0,X) \mid G=0\right] \\ 0 &= \E_P\Bigg[\tilde{\psi}_{1,0}(B_1) G\Bigg(W \frac{\rho^\prime(S,X)}{\rho(S,X)} \nonumber \\ &\quad \quad - (1-W) \frac{\int \rho^\prime(s,X) p_\star(s\mid X) d\lambda(s)}{1-\int \rho (s,X)p_\star(s\mid X)\,d\lambda(s)}\Bigg) \Bigg] \\ 0 &= \E_P\bk{\tilde{\psi}_{1,0}(B_1)(1-G)\frac{W - \varrho(X)}{\varrho(X)(1-\varrho(X))} \varrho'(X)} \end{align} where the only differences relative to (ref) through (ref) are the replacement of $G$ with $1-G$ and $\pi$ with $1-\pi$ in (ref) through (ref) on the sides of the equalities that correspond to $\theta'_{1,0}$. Since the density of the data doesn't change with the estimand, the sides of the equalities that correspond to orthogonality condition \[\E_P\left[\tilde{\psi}_{1,0}(B_1) l'(B_1)\right]\] remains unchanged relative to (ref) through (ref). We make the same choices $s_4(s,x) = -\rho(s,x) s_2(s,x)$ and $s_5(x) = 0$ as in (ref), resulting in an conjectured efficient influence function of the form \[ \psi_{1,0}(b_1) = gw \cdot s_1 (y, s,x) + (1-g)w \cdot s_2(s,x) + s_3(g,x) \] We make the conjecture that \begin{align} s_1 (y,s,x) & = f_1(x) \cdot \frac{y-\mu(s,x)}{\rho(s,x)} , \nonumber & s_2 (s,x) = \frac{\mu(s,x) - \bar{\mu}(x)}{(1-\pi) \varrho(x)} , \quadand\quad \nonumber\\ s_3(g,x) &= \frac{1-g}{1-\pi}(\bar{\mu}(x) - \theta_{1,0}) , \end{align} where \[f_1(x) = \frac{1}{1-\pi}\frac{1-\gamma(x)}{\gamma(x)}~.\] These choices satisfy the conditional mean-zero conditions for scores. We verify the conditions (ref) through (ref) sequentially, completing the proof. To verify condition (ref), we note that by the mean-zero property of the conditional scores, the right-hand side of (ref) simplifies to \begin{align*} \E_{P_\star}\bk{f_1(X) \frac{GWY^*}{\rho(s,x)} l'_*(Y^*\mid S^*,X) } &= \E_{P_\star}[\pi f_1(X)Y^* l'_*(Y^*\mid S^*,X) \mid G=1] \\ & = \E_{P_\star}\bk{\frac{1-\gamma}{\gamma} \frac{\pi}{1-\pi} Y^*l'(Y^* \mid S^*,X) \mid G=1} \\ & = \E\bk{Y^* l'(Y^* \mid S^*, X) \mid G=0}\\ &= \frac{1}{1-\pi}\E_{P_\star}[(1-G)Y^* l'_*(Y^*\mid S^*, X)] \\ &= \E_P\bk{\tilde \psi_{1,0}(B_1) GW l'_*(Y\mid S, X)} . \end{align*} where the third equality follows from the importance sampling argument \[ p_\star(y,s,x \mid G=1) = p_\star(y,s\mid x)p_\star(x\mid G=1) = p_\star(y,s\mid x) p_\star(x \mid G=0) \frac{1-\gamma}{\gamma}\frac{\pi}{1-\pi}~. \] To verify condition (ref), we note that by the mean-zero property of the conditional scores, the right-hand side of (ref) simplifies to \begin{align*} \E_{P_\star}\bk{ \frac{1-G}{1-\pi} W \frac{\mu(S,X)}{\varrho(x)}l'(S^* \mid X) } & = \E_{P_\star}[\mu(S,X) l'(S^* \mid X) \mid G=0]\\ & = \E_{P_\star}[Y^* l'(S^* \mid X) \mid G=0] , \end{align*} which is equal to the left-hand side of (ref). To verify condition (ref), we note again that by the mean-zero properties of the conditional scores, the left-hand side of (ref) simplifies to \begin{align*} \E_P\bk{\frac{1-G}{1-\pi} (\bar{\mu}(X) - \theta_{1,0}) l'(G=0, X)} &= \E_{P_\star}\bk{(\bar{\mu}(X) - \theta_{1,0}) l'_*(G=0, X) \mid G=0} \\ &= \E_{P_\star}[\bar{\mu}(X) l'_*(G=0, X) \mid G=0] \\ &- \theta_{1,0} \E_{P_\star}[l'_*(G=0, X) \mid G=0] \end{align*} which is equal to the right-hand side of (ref). Finally, condition (ref) holds by mean-zero properties of $s_1(y,s,x)$ and condition (ref) holds by mean-zero properties of $s_2(y,s,x)$ and the fact that $\E_P[W-\varrho(x) \mid X, G=0] =0 $, completing the argument. Next, we prove Part 2. The analogue of the pathwise derivative (ref) for $\theta_{1,0}$ is now \begin{align} \theta_{1,0}' =\,\,& \E_P[\E_P[\E_P[Y l'(Y \mid S, X, G=1) \mid S, X, G=1] \mid X, W=1 ,G=0] \mid G=0] \nonumber \\ & + \E_P[\E_P[\nu(S,X)l'(S\mid X, G=0, W=1) \mid X, W=1, G=0] \mid G=0] \nonumber \\ & + \E_P\bk{(1-\pi)^{-1} \bar{\nu}(X)\pr{-\gamma'(X) + l'(X)(1-\gamma(X))} } \nonumber \\ & + (1-\pi)^{-1} \E_P[\gamma'(X) + \gamma(X) l'(X)] \theta_{1,0}. \end{align} The conjectured efficient influence function is given by \[ \xi_{1,0}(b_1) = g(y - \nu(s,x))\cdot f_1(s,x) + (1-g) w(\nu(s,x) - \bar{\nu}(x))\cdot f_2(x) +(1-g)\cdot f_3(x) \] with $f_1(s,x)$, $f_2(x)$, and $f_3(x)$ to be specified, where $f_3(x)$ is chosen such that \[ \E[f_3(X) \mid G=0] = 0~. \] An argument very similar to that given in (ref) demonstrates that $\xi_{1,0}(b_1)$ is in the tangent space. Again, following a similar argument, we may verify that the choices \begin{align*} f_1(s,x) &= \frac{\varrho(s,x)}{(1-\pi) \varrho(x)} \frac{1-\gamma(s,x)}{\gamma(s,x)} , \\ f_2(x) &= \frac{1}{(1-\pi) \varrho(x)} , \\ f_3(x) &= \frac{\nu(x) - \theta_{1,0}}{1-\pi} ,and \end{align*} satisfy the pathwise differentiability conditions. The conditions for $f_1(s,x)$ and $f_2(x)$ are such that multiplication of the density ratio \[\frac{p (x\mid G=0)}{p(x \mid G=1)} = \frac{\pi} {1-\pi} \frac{1-\gamma(x)}{\gamma(x)}\] with the choices for $f_1(s,x)$ and $f_2(x)$ given in (ref) yields the corresponding choices here.

Estimation

In this section, we define a semiparametric estimator of the long-term treatment effect $\tau_0$ in the experimental population. This estimator is analogous to the estimator formulated in (ref). Again, we use the “Double/Debiased Machine Learning” (DML) construction developed in chernozhukov2018double.

defn[DML Estimators] Let $\hat{\eta}(I)$ and $\hat{\varphi}(I)$ denote generic estimates of $\eta$ and $\varphi$ based on the data $\{B_i\}_{i\in I}$ for some subset $I\in[n]$. Let $\{I_l\}_{l=1}^k$ denote a random $k$-fold partition of $[n]$ such that the size of each fold is $m=n/k$. The estimator $\hat{\tau}_0$ is defined as the solution to \[ \frac{1}{k}\sum_{l=1}^k \frac{1}{m} \sum_{i\in I_l} \psi_0(B_i,\hat{\tau}_0,\hat{\eta}(I_l^c)) = 0 \quad\text{or}\quad \frac{1}{k}\sum_{l=1}^k \frac{1}{m} \sum_{i\in I_l} \xi_0(B_i,\hat{\tau}_0,\hat{\varphi}(I_l^c)) = 0 \] for the Latent Unconfounded Treatment and Statistical Surrogacy Models, respectively.

(ref) demonstrates that $\hat{\tau}_0$ is semiparametrically efficient for $\tau_0$ under the Latent Unconfounded Treatment Model.

theorem[DML Estimation and Inference] Let $\mathcal{P}\subset\mathcal{M}_\lambda$ be the set of all probability distributions $P$ for $\{B_i\}_{i=1}^n$ that satisfy the Latent Unconfounded Treatment Model stated in (ref) in addition to (ref). If (ref) holds for every $\mathcal P$, then \begin{align} \sqrt{n}(\hat{\tau}_0 -\tau_0) \overset{d}{\to} \mathcal{N}(0,V_0^\star) \end{align} uniformly over $P\in\mathcal{P}$, where $\hat{\tau}_0$ is defined in (ref), $V_0^*$ is defined in (ref), and $\overset{d}{\to}$ denotes convergence in distribution. Moreover, we have that \begin{align} \hat{V}_0^* = \frac{1}{k}\sum_{l=1}^k \frac{1}{m} \sum_{i\in I_l} \left(\psi_0(B_i,\hat{\tau}_0,\hat{\eta}(I_l^c))\right)^2 \overset{p}{\to} V_0^\star \end{align} uniformly over $P\in\mathcal{P}$, where $\overset{p}{\to}$ denotes convergence in probability, and as a result we obtain the uniform asymptotic validity of the confidence intervals \begin{align} \lim_{n\to\infty} \sup_{P\in\mathcal{P}} \Big\vert P\left(\tau_0 \in \left[\hat{\tau}_1 \pm z_{1-\alpha/2}\sqrt{\hat{V}_0^\star/n} \right]\right) - (1-\alpha) \Big\vert = 0 , \end{align} where $z_{1-\alpha/2}$ is the $1-\alpha/2$ quantile of the standard normal distribution.
theorem[DML Estimation and Inference] Let $\mathcal{P}\subset\mathcal{M}_\lambda$ be the set of all probability distributions $P$ for $\{B_i\}_{i=1}^n$ that satisfy the Statistical Surrogacy Model stated in (ref) in addition to (ref). If (ref) holds for every $\mathcal P$, then \begin{align} \sqrt{n}(\hat{\tau}_0 -\tau_0) \overset{d}{\to} \mathcal{N}(0,V_0^{\star\star}) \end{align} uniformly over $P\in\mathcal{P}$, where $\hat{\tau}_0$ is defined in (ref), $V_0^{\star\star}$ is defined in (ref), and $\overset{d}{\to}$ denotes convergence in distribution. Moreover, we have that \begin{align} \hat{V}_0^{\star\star} = \frac{1}{k}\sum_{l=1}^k \frac{1}{m} \sum_{i\in I_l} \left(\xi_0(B_i,\hat{\tau}_0,\hat{\varphi}(I_l^c))\right)^2 \overset{p}{\to} V_0^{\star\star} \end{align} uniformly over $P\in\mathcal{P}$, where $\overset{p}{\to}$ denotes convergence in probability, and as a result we obtain the uniform asymptotic validity of the confidence intervals \begin{align} \lim_{n\to\infty} \sup_{P\in\mathcal{P}} \Big\vert P\left(\tau_0 \in \left[\hat{\tau}_1 \pm z_{1-\alpha/2}\sqrt{\hat{V}_0^{\star\star}/n} \right]\right) - (1-\alpha) \Big\vert = 0 , \end{align} where $z_{1-\alpha/2}$ is the $1-\alpha/2$ quantile of the standard normal distribution.
proof[Proof of (ref) and (ref)] The proof of (ref) is very similar to the proof of (ref), being obtained by verifying the conditions of Theorem 3.1 of chernozhukov2018double. The arguments in support of the verification of Assumption 3.1 are identical, with $1-g$ replacing $g$ in the statement supporting condition (e). Similarly, the arguments supporting the verification of Assumption 3.2 are very similar, where the constants premultiplying the various bounds derived in support of conditions (b), (c), and (d) will be slightly different, but will be obtained with the same arguments, which we omit to avoid repetition. The second Gateaux derivatives of the efficient influence functions have the same structure, and so the third inequality in condition (c) will again follow from Hölder's inequality and (ref).

Additional Results

Consistency Double Robustness

Recall that, in (ref), we divide the nuisance parameters $\eta$ or $\varphi$ into $(\omega, \kappa, \pi)$, where $\omega$ consists of a set of outcome regression nuisance parameters, and $\kappa$ consists of a set of propensity score-type nuisance parameters. We note that the mean-zero properties of the efficient influence function is preserved as long as one of $\omega, \kappa$ is set at its respective true value. In this section, we show a consistency analogue of this double robustness result.

To economize on notation, we note that the solutions $\hat\tau_{1n}$ to both $\psi_1 = 0$ and $\xi_1 = 0$ take the form \[ \hat\tau_{1n} = \tau_{1n}(\hat \omega_n, \hat\kappa_n) = \frac{1}{n} \sum_{i=1}^n g(B_i; \hat \omega_n(B_i), \hat \kappa_n(B_i), \hat{\pi}_n) \numberthis \label{eq:consistency_tau_def} \] when $\hat\pi_n = \frac{1}{n} \sum_i G_i$ is used. Here, $\hat\omega_n (B_i), \hat\kappa_n(B_i)$ evaluates the nuisance parameters at $B_i$. Define $\norm{\cdot}_\infty$ entrywise as $\norm{\theta}_\infty = \max_j \norm{\theta_j}_\infty$ for a vector of nuisance parameters $\theta$.

theoremLet $\hat\tau_{1n}$ be as in (ref), derived either from either $g$ equal to $\psi_1$ or $\xi_1$ under either the Latent Unconfoundedness or the Statistical Surrogacy Models. Assume that $\hat\pi_n = \frac{1}{n} \sum_i G_i$. Assume additionally that (i) entries in $\kappa$ are bounded uniformly between $[1-\varepsilon, \varepsilon]$ for some $\varepsilon > 0$, (ii) entries in $\omega$ are bounded uniformly by some finite $C > 0$, (iii) entries in $\hat\kappa_n$ are bounded uniformly between $[1-\varepsilon, \varepsilon]$ almost surely, (iv) entries in $\hat\omega_n$ are bounded uniformly by $C$ almost surely, (v) all nuisance parameters, except for $\hat\pi_n$, are estimated on a hold-out sample, and (vi) $\E_P |Y|^2 < \infty$. Then: \begin{enumerate} • If $\norm{\hat\omega_n - \omega}_\infty \pto 0$, then $\hat\tau_{1n} \pto \tau_1$. • If $\norm{\hat\kappa_n - \kappa}_\infty \pto 0$, then $\hat\tau_{1n} \pto \tau_1$. \end{enumerate}
proofLet $A_n = \one(\hat\pi_n \in [\varepsilon/2, 1-\varepsilon/2])$ be an event that occurs with probability tending to one. For either (1) or (2), let $\eta_C$ collect the consistently estimated nuisance parameters, and let $\eta_I$ collect the rest. Let $\hat\eta_{Cn}$, $\hat\eta_{In}$ be their estimated analogues. Note that $\pi$ is always in $\eta_C$. Note that under either model, by Taylor's theorem, on the event $A_n$, \[ g(B_i; \hat\omega_n(B_i), \hat\kappa_n(B_i)) = g(B_i; \eta_C(B_i), \hat\eta_{In}(B_i)) + \diff{}{\eta_{Ci}} g(B_i; \eta_{Ci}, \hat\eta_{In}(B_i))\evalbar_{\eta_{Ci} = \tilde \eta_{Ci}} (\hat \eta_{Cn} (B_i) - \eta_C(B_i)) \] where $\tilde \eta_{Ci}$ lies somewhere on the line segment connecting $\hat \eta_{Cn} (B_i)$ and $ \eta_C(B_i)$. Hence, by inspecting the derivative of $g$, we find that, for some constant $M$ that depends on $(C, \varepsilon)$, \begin{align*} &\abs[\bigg]{\hat\tau_{1n} - \frac{1}{n} \sum_{i=1}^n g(B_i; \eta_C(B_i), \hat\eta_{In}(B_i)) }\\ &= A_n \abs[\bigg]{\hat\tau_{1n} - \frac{1}{n} \sum_{i=1}^n g(B_i; \eta_C(B_i), \hat\eta_{In}(B_i)) } + o_p(1) \&\le \norm{\hat \eta_{Cn} - \eta_C}_\infty \frac{A_n}{n} \sum_{i=1}^n \abs[\bigg]{\diff{\eta_{Ci}} g(B_i; \eta_{Ci}, \hat\eta_{In}(B_i))\evalbar_{\eta_{Ci} = \tilde \eta_{Ci}}} + o_p(1) \\ &\le M(C, \varepsilon) \norm{\hat \eta_{Cn} - \eta_C}_\infty \frac{1}{n} \sum_{i=1}^n |Y_i| + o_p(1) \tag{(i)--(iv)}\\ &=o_p(1) \tag{vi}. \end{align*} It suffices to then show that \[ \abs[\bigg]{\tau_1 - \frac{1}{n} \sum_{i=1}^n g(B_i; \eta_C(B_i), \hat\eta_{In}(B_i))} = o_p(1). \] Note that, conditional on $\hat\eta_{In}$, $\frac{1}{n} \sum_{i=1}^n g(B_i; \eta_C(B_i), \hat\eta_{In}(B_i))$ is an average of i.i.d. random variables with mean $\tau_1$, by the double robust property and (v). By the law of large numbers applied to triangular arrays, which uses (i)--(iv) and (vi), for almost every sequence $\hat\eta_{In}$, for every $\epsilon > 0$, \[ \P\pr{\abs[\bigg]{\tau_1 - \frac{1}{n} \sum_{i=1}^n g(B_i; \eta_C(B_i), \hat\eta_{In}(B_i))} > \epsilon \mid \hat\eta_{In}} \to 0. \] Thus, by the dominated convergence theorem applied to left-hand side of the previous display, treating it as a random variable indexed by $n$, \begin{align*} &\P\pr{\abs[\bigg]{\tau_1 - \frac{1}{n} \sum_{i=1}^n g(B_i; \eta_C(B_i), \hat\eta_{In}(B_i))} > \epsilon } \\ &= \E\bk{ \P\pr{\abs[\bigg]{\tau_1 - \frac{1}{n} \sum_{i=1}^n g(B_i; \eta_C(B_i), \hat\eta_{In}(B_i))} > \epsilon \mid \hat\eta_{In}} } \to 0 , \end{align*} as required.

Known Nuisance Functions

In this section, we assess how the efficient influence functions derived in (ref) change if different components of the nuisance functions $\eta$ and $\varphi$ are known. We consider only estimation of the long-term treatment effect in the observational sample $\tau_1$. The case of the treatment effect in the experimental sample $\tau_1$ is analogous.

First, we consider the Latent Unconfounded Treatment Model. We show that the efficient influence function, and thereby, the semiparametric efficiency bound, is unchanged if the propensity score in the experimental sample $\varrho(x)$ is known. This result echoes an analogous result for estimation of the average treatment effects under ignorability given in hahn1998role. On the other hand, if the probability of being included in observational sample $\gamma(x)$ is known, then there is a change in the efficient influence function. In particular, the term $(g/\pi)((\bar{\mu}_1(x) - \bar{\mu}_0(x)) - \tau_1)$ in $\psi_1(b,\tau_1,\eta)$ is replaced by $(\gamma(x)/\pi)((\bar{\mu}_1(x) - \bar{\mu}_0(x)) - \tau_1)$. Proofs for these results are given in (ref), respectively.

theoremConsider the Latent Unconfounded Treatment Model, given in (ref). (i) If the propensity score in the experimental sample $\varrho(x)$ is known, then the efficient influence function for the parameter $\tau_1$ is unchanged. (ii) If the probability of being included in observational sample $\gamma(x)$ is known, then the efficient influence function for the parameter $\tau_1$ is given by \begin{align} \check{\psi}_1(b,\tau_1,\eta) &= \frac{1}{\pi} \Bigg(g\left( \frac{w(y-\mu_1(s,x))}{\rho_1(s,x)} - \frac{(1-w)(y-\mu_0(s,x))}{\rho_0 (s,x)}\right) + \gamma(x)\left( (\bar{\mu}_1(x) - \bar{\mu}_0(x)) - \tau_1\right) \nonumber\\ & + \frac{(1-g)\gamma (x)}{1-\gamma(x)}\left( \frac{w(\mu_1(s,x)-\bar{\mu}_1(x))}{\varrho(x)} - \frac{(1-w)(\mu_0(s,x)-\bar{\mu}_0(x))}{1-\varrho(x)}\right)\Bigg), \end{align} where, again, the parameter $\eta $ collects the nuisance functions appearing in (ref).

Next, we provide analogous results for the Statistical Surrogacy Model. Here, however, the efficient influence function is additionally invariant to knowledge of the distribution in the observational sample of the short-term outcomes conditional on the covariates, i.e. the law $S\mid X, G=1$. Proofs are given in (ref).

theoremConsider the Statistical Surrogacy Model, given in (ref). (i) If the propensity score in the experimental sample $\varrho(x)$ is known, then the efficient influence function for the parameter $\tau_1$ is unchanged. (ii) If the distribution of the short-term outcomes, conditional on the pretreatment covariates, is known in observational dataset, then the efficient influence function for the parameter $\tau_1$ is unchanged. (iii) If the probability of being included in observational sample $\gamma(x)$ is known, then the efficient influence function for the parameter $\tau_1$ is given by \begin{align} \check{\xi}_1(b,\tau_1,\eta) &= \frac{g}{\pi} \left( \frac{\gamma(x)}{\gamma(s,x)} \frac{1- \gamma(s,x)}{1-\gamma(x)}\frac{(\varrho(s,x)-\varrho(x))(y-\nu(s,x))}{\varrho(x)(1-\varrho(x)}\right) + \frac{\gamma(x)}{\pi}\left((\bar{\nu}_1(x) - \bar{\nu}_0(x)) - \tau_1\right) \nonumber \\ & + \frac{1-g}{\pi} \left(\frac{\gamma(x)}{1-\gamma(x)} \left(\frac{w(\nu(s,x)-\bar{\nu}_1(x))}{\varrho(x)} -\frac{(1-w)(\nu(s,x)-\bar{\nu}_0(x))}{1-\varrho(x)}\right)\right) , \end{align} where, again, the parameter $\varphi$ collects the nuisance functions appearing in (ref).

Proof of (ref), Part (i)

Let $\mathcal{P}$ be a regular parametric submodel of $\mathcal{M}_{\lambda}$, indexed by $\varepsilon\in\mathbb{R}$ and such that (ref) hold for each $P_{\varepsilon}\in\mathcal{P}$. Additionally, suppose that \[ p_{\varepsilon}\left(W=1\mid x,G=0\right)=\varrho\left(x\right) \] for all $\varepsilon\in\mathbb{R}$ and some fixed function $\varrho\left(x\right)$.

In this case, we have the score

flalign*\ell^{\prime}\left(b_{1}\right) & =wg\cdot\ell_{*}^{\prime}\left(y\mid x,s\right)+w\cdot\ell_{*}^{\prime}\left(s\mid x\right)+\ell_{*}^{\prime}\left(g,x\right)\\ & -g\left(1-w\right)\left(\frac{\mathbb{E}_{P_{\star}}\left[\rho\left(S^{*},X\right)\ell_{*}^{\prime}\left(S^{*}\mid X\right)\mid G=0,W=1,X=x\right]}{1-\mathbb{E}_{P_{\star}}\left[\rho\left(S^{*},X\right)\mid G=0,W=1,X=x\right]}\right)\\ & +g\left(w\cdot\frac{\rho^{\prime}\left(s,x\right)}{\rho\left(s,x\right)}-\left(1-w\right)\frac{\mathbb{E}_{P_{\star}}\left[\rho^{\prime}\left(S^{*},X\right)\mid G=0,W=1,X=x\right]}{1-\mathbb{E}_{P_{\star}}\left[\rho\left(S^{*},X\right)\mid G=0,W=1,X=x\right]}\right)

and so the tangent space $\mathcal{T}$ is given by the mean-square closure of the linear space of the functions

align*[align* omitted — 610 chars of source]

where the functions $s_{1}$ through $s_{4}$ range over the space of mean-zero and square integrable functions that additionally satisfy the restrictions

flalign\mathbb{E}_{P}\left[s_{1}\left(Y\mid X,S\right)\mid W=1,G=1,S,X\right] & =0\quadand\\ \mathbb{E}_{P}\left[s_{2}\left(S\mid X\right)\mid W=1,G=0,X\right] & =0.

Recall the integral representation

equation[equation omitted — 245 chars of source]

Observe that $p\left(w=1\mid x,G=0\right)$ does not appear in ((ref)). Thus, the pathwise derivative of $\theta_{1,1}$ at $0$ on the submodel of $\mathcal{P}$ is again given by (ref). Thus, it will suffice to choose functions $s_{1}\left(\cdot\right)$ through $s_{4}\left(\cdot\right)$ satisfying the above conditions and whose resultant score function satisfies (ref) through (ref).

Consider the choices

align*[align* omitted — 448 chars of source]

which again yield the conjectured influence function \[ \psi_{1,1}\left(b_{1}\right)=\frac{gw\left(y-\bar{\mu}\left(s,x\right)\right)}{\pi\rho\left(s,x\right)}+\frac{\left(1-g\right)w}{\left(1-\gamma\left(x\right)\right)\varrho\left(x\right)}\frac{\gamma\left(x\right)}{\pi}\left(\mu\left(s,x\right)-\bar{\mu}\left(x\right)\right)+\frac{g}{\pi}\left(\bar{\mu}\left(x\right)-\theta_{1,1}\right). \] Each of these choices are mean-zero and square integrable, and satisfy the conditions ((ref)) and ((ref)), as before. The arguments verifying the conditions (ref) through (ref) given in proof of (ref), Part 1, completeing the proof. \qed

Proof of (ref), Part (ii)

Let $\mathcal{P}$ be a regular parametric submodel of $\mathcal{M}_{\lambda}$, indexed by $\varepsilon\in\mathbb{R}$ and such that (ref) hold for each $P_{\varepsilon}\in\mathcal{P}$. Additionally, suppose that \[ p_{\varepsilon}\left(G=1\mid X=x\right)=\gamma\left(x\right) \] for all $\varepsilon\in\mathbb{R}$ and some fixed function $\gamma\left(x\right)$.

In this case, we have the score

flalign*\ell^{\prime}\left(b_{1}\right) & =wg\cdot\ell_{*}^{\prime}\left(y\mid x,s\right)+w\cdot\ell_{*}^{\prime}\left(s\mid x\right)+\ell_{*}^{\prime}\left(x\right)\\ & -g\left(1-w\right)\left(\frac{\mathbb{E}_{P_{\star}}\left[\rho\left(S^{*},X\right)\ell_{*}^{\prime}\left(S^{*}\mid X\right)\mid G=0,W=1,X=x\right]}{1-\mathbb{E}_{P_{\star}}\left[\rho\left(S^{*},X\right)\mid G=0,W=1,X=x\right]}\right)\\ & +g\left(w\cdot\frac{\rho^{\prime}\left(s,x\right)}{\rho\left(s,x\right)}-\left(1-w\right)\frac{\mathbb{E}_{P_{\star}}\left[\rho^{\prime}\left(S^{*},X\right)\mid G=0,W=1,X=x\right]}{1-\mathbb{E}_{P_{\star}}\left[\rho\left(S^{*},X\right)\mid G=0,W=1,X=x\right]}\right)\\ & +\left(1-g\right)\varrho^{\prime}\left(x\right)\left(\frac{w-\varrho\left(x\right)}{\varrho\left(x\right)\left(1-\varrho\left(x\right)\right)}\right)

and so the tangent space $\mathcal{T}$ is given by the mean-square closure of the linear space of the functions

flalign*s\left(b_{1}\right) & =wg\cdot s_{1}\left(y\mid x,s\right)+w\cdot s_{2}\left(s\mid x\right)+s_{3}\left(x\right)\\ & -g\left(1-w\right)\left(\frac{\mathbb{E}_{P_{\star}}\left[\rho\left(S^{*},X\right)s_{2}\left(S^{*}\mid X\right)\mid G=0,W=1,X=x\right]}{1-\mathbb{E}_{P_{\star}}\left[\rho\left(S^{*},X\right)\mid G=0,W=1,X=x\right]}\right)\\ & +g\left(w\cdot\frac{s_{4}\left(s,x\right)}{\rho\left(s,x\right)}-\left(1-w\right)\frac{\mathbb{E}_{P_{\star}}\left[s_{4}\left(S^{*},X\right)\mid G=0,W=1,X=x\right]}{1-\mathbb{E}_{P_{\star}}\left[\rho\left(S^{*},X\right)\mid G=0,W=1,X=x\right]}\right)\\ & +\left(1-g\right)\left(s_{5}\left(x\right)\cdot\frac{w-\varrho\left(x\right)}{\varrho\left(x\right)\left(1-\varrho\left(x\right)\right)}\right)

where the functions $s_{1}$ through $s_{5}$ range over the space of mean-zero and square integrable functions that additionally satisfy the restrictions ((ref)) and ((ref)). By the integral representation ((ref)) and the fact that \[ p_{\varepsilon}\left(x,G=1\right)=\gamma\left(x\right)p_{\varepsilon}\left(x\right) \] for all $\varepsilon\in\mathbb{R}$, the pathwise derivative of $\theta_{1,1}$ at $0$ on $\mathcal{P}$ is given by

flalign\theta_{1,1}^{\prime} & =\mathbb{E}_{P_{\star}}\left[Y^{*}\ell^{\prime}\left(Y^{*}\mid S^{*},X\right)\mid G=1\right]\nonumber \\ & +\mathbb{E}_{P_{\star}}\left[Y^{*}\ell^{\prime}\left(S^{*}\mid X\right)\mid G=1\right]+\mathbb{E}_{P_{\star}}\left[Y^{*}\ell^{\prime}\left(X\right)\mid G=1\right].

Each of the terms in ((ref)) can be written in terms of conditional scores by

flalign*\mathbb{E}_{P_{\star}}\left[Y^{*}\ell^{\prime}\left(Y^{*}\mid S^{*},X\right)\mid G=1\right] & =\mathbb{E}_{P_{\star}}\left[\pi^{-1}GY^{*}\ell_{*}^{\prime}\left(Y^{*}\mid S^{*},X\right)\right],\\ \mathbb{E}_{P_{\star}}\left[Y^{*}\ell^{\prime}\left(S^{*}\mid X\right)\mid G=1\right] & =\mathbb{E}_{P_{\star}}\left[\pi^{-1}GY^{*}\ell_{*}^{\prime}\left(S^{*}\mid X\right)\right],\quadand\\ \mathbb{E}_{P_{\star}}\left[Y^{*}\gamma\left(X\right)\ell^{\prime}\left(X\right)\mid G=1\right]. & =\mathbb{E}_{P_{\star}}\left[\pi^{-1}GY^{*}\ell_{*}^{\prime}\left(X\right)\right],

respectively. Thus, in order to establish that a mean-zero and square-integrable function $\tilde{\psi}_{1,1}\left(B_{1}\right)$ is an influence function for $\theta_{1,1}$ it suffices to verify the conditions (ref) through (ref), with (ref) replaced by

equation[equation omitted — 234 chars of source]

Consider the choices

align*[align* omitted — 483 chars of source]

which yield the conjectured influence function

flalign*\check{\psi}_{1,1}\left(b_{1}\right) & =\frac{gw\left(y-\mu\left(s,x\right)\right)}{\pi\rho\left(s,x\right)}+\frac{\left(1-g\right)w}{\pi\varrho\left(x\right)}\frac{\gamma\left(x\right)}{1-\gamma\left(x\right)}\left(\mu\left(s,x\right)-\bar{\mu}\left(x\right)\right)+\frac{\gamma\left(x\right)}{\pi}\left(\bar{\mu}\left(x\right)-\theta_{1,1}\right).

These choices are again mean-zero and square integrable, and satisfy the conditions ((ref)) and ((ref)). Thus, it suffices to check that the conditions (ref), (ref), (ref), (ref), and ((ref)) continue to hold.

The conditions (ref), (ref), and (ref) continue to hold through iterated expectation conditional on $X$ by the fact the only term changed relative to $\psi_{1,1}\left(b_{1}\right)$ is $s_{3}\left(x\right)$. The condition (ref) continues to hold as

flalign*& \mathbb{E}_{P}\left[\frac{1-G}{\pi}\left(\gamma\left(X\right)\bar{\mu}\left(X\right)-\theta_{1,1}\right)\frac{W-\varrho\left(X\right)}{\varrho\left(X\right)\left(1-\varrho\left(X\right)\right)}\varrho^{\prime}\left(X\right)\right]\\ & =\mathbb{E}_{P}\left[\frac{\left(1-G\right)}{\pi}\left(\mathbb{E}_{P}\left[G\mid X\right]\gamma\left(X\right)\bar{\mu}\left(X\right)-\theta_{1,1}\right)\frac{W-\varrho\left(X\right)}{\varrho\left(X\right)\left(1-\varrho\left(X\right)\right)}\varrho^{\prime}\left(X\right)\right]\\ & =-\mathbb{E}_{P}\left[\frac{\theta_{1,1}\left(1-G\right)}{\pi}\frac{\theta_{1,1}\left(W-\varrho\left(X\right)\right)}{\varrho\left(X\right)\left(1-\varrho\left(X\right)\right)}\varrho^{\prime}\left(X\right)\right]=0.

Thus, it remains to verify the new condition ((ref)). This condition reduces to \[ \mathbb{E}_{P_{\star}}\left[\pi^{-1}GY^{*}\ell_{*}^{\prime}\left(X\right)\right]=\mathbb{E}_{P}\left[\frac{1}{\pi}\gamma\left(X\right)\bar{\mu}\left(X\right)\ell_{*}^{\prime}\left(X\right)\right], \] as before. Observe that

flalign*\mathbb{E}_{P}\left[\pi^{-1}\gamma\left(X\right)\bar{\mu}\left(X\right)\ell_{*}^{\prime}\left(X\right)\right] & =\mathbb{E}_{P}\left[\pi^{-1}\gamma\left(X\right)\mathbb{E}_{P}\left[\mu\left(S,X\right)\mid G=0,W=1,X\right]\ell_{*}^{\prime}\left(X\right)\right]\\ & =\mathbb{E}_{P}\left[\pi^{-1}\gamma\left(X\right)\mathbb{E}_{P^{*}}\left[\mu\left(S^{*},X\right)\mid X\right]\ell_{*}^{\prime}\left(X\right)\right]\\ & =\mathbb{E}_{P}\left[\pi^{-1}\gamma\left(X\right)\mathbb{E}_{P^{*}}\left[\mathbb{E}_{P^{*}}\left[Y^{*}\mid W=1,G=1,S,X\right]\mid X\right]\ell_{*}^{\prime}\left(X\right)\right]\\ & =\mathbb{E}_{P}\left[\pi^{-1}\mathbb{E}\left[G\mid X,Y^{*}\right]Y^{*}\ell_{*}^{\prime}\left(X\right)\right]\\ & =\mathbb{E}_{P_{\star}}\left[\pi^{-1}GY^{*}\ell_{*}^{\prime}\left(X\right)\right]

where the second to last equality follows from Assumption 2.3, completing the proof. \qed

Proof of (ref), Part (i)

Let $\mathcal{P}$ be a regular parametric submodel of $\mathcal{M}_{\lambda}$, indexed by $\varepsilon\in\mathbb{R}$ and such that (ref) hold for each $P_{\varepsilon}\in\mathcal{P}$. Additionally, suppose that \[ p_{\varepsilon}\left(W=1\mid x,G=0\right)=\varrho\left(x\right) \] for all $\varepsilon\in\mathbb{R}$ and some fixed function $\varrho\left(x\right)$.

In this case, we have the score

flalign*\ell^{\prime}\left(b_{1}\right) & =g\cdot\ell^{\prime}\left(y\mid x,s,G=1\right)+g\cdot\ell^{\prime}\left(s\mid x,G=1\right)+w\left(1-g\right)\cdot\ell^{\prime}\left(s\mid W=1,x,G=0\right)\\ & +\ell^{\prime}\left(x\right)+\frac{g-\gamma\left(x\right)}{\gamma\left(x\right)\left(1-\gamma\left(x\right)\right)}\gamma^{\prime}\left(x\right)

and so the tangent space $\mathcal{T}$ is given by the mean-square closure of the linear space of the functions

flalign*s\left(b_{1}\right) & =g\cdot s_{1}\left(y\mid x,s,G=1\right)+g\cdot s_{2}\left(s\mid x,G=1\right)+w\left(1-g\right)s_{3}\left(s\mid W=1,x,G=0\right)\\ & +s_{4}\left(x\right)+\frac{g-\gamma\left(x\right)}{\gamma\left(x\right)\left(1-\gamma\left(x\right)\right)}s_{5}\left(x\right)

where the functions $s_{1}$ through $s_{5}$ range over the space of mean-zero and square integrable functions that additionally satisfy the restrictions

flalign\mathbb{E}_{P}\left[s_{1}\left(Y\mid X,S,G=1\right)\mid S,X,G=1\right] & =0,\\ \mathbb{E}_{P}\left[s_{2}\left(S\mid X,G=1\right)\mid X,G=1\right] & =0,\quadand\\ \mathbb{E}_{P}\left[s_{3}\left(S\mid X,W=1G=0\right)\mid X,W=1G=0\right] & =0.

Consider the choices

flaligns_{1}\left(y\mid s,x,G=1\right) & =\frac{1}{\pi}\frac{\varrho\left(s,x\right)}{\varrho\left(x\right)}\frac{1-\gamma\left(s,x\right)}{\gamma\left(s,x\right)}\frac{\gamma\left(x\right)}{1-\gamma\left(x\right)}\left(y-\nu\left(s,x\right)\right),\\ s_{2}\left(s\mid X,G=1\right) & =0,\\ s_{3}\left(s\mid x,W=1,G=1\right) & =\frac{1}{\pi}\frac{1}{\varrho\left(x\right)}\frac{\gamma\left(x\right)}{1-\gamma\left(x\right)}\left(\nu\left(s,x\right)-\bar{\nu}\left(x\right)\right),\\ s_{4}\left(x\right) & =\gamma\left(x\right)\left(\frac{\bar{\nu}\left(x\right)-\theta_{1,1}}{\pi}\right),\quadand\\ s_{5}\left(x\right) & =\gamma\left(x\right)\left(1-\gamma\left(x\right)\right)\left(\frac{\bar{\nu}\left(x\right)-\theta_{1,1}}{\pi}\right).

These choices are again mean-zero and square integrable, and satisfy the conditions ((ref)), ((ref)), and ((ref)).

Recall the integral representation

equation[equation omitted — 250 chars of source]

Observe that $p\left(w=1\mid x,G=0\right)$ does not appear in ((ref)). Thus, the pathwise derivative of $\theta_{1,1}$ at $0$ on the submodel of $\mathcal{P}$ is again given by (ref). Thus, it will suffice to verify the conditions

flalign& \mathbb{E}_{P}\left[Gs_{1}\left(Y\mid S,X,G=1\right)\ell^{\prime}\left(Y\mid X,S,G=1\right)\right]\\ & \quad\quad=\mathbb{E}_{P}\left[\mathbb{E}_{P}\left[\mathbb{E}_{P}\left[Y\ell^{\prime}\left(Y\mid S,X,G=1\right)\mid S,X,G=1\right]\mid X,W=1,G=0\right]\mid G=1\right]\nonumber \\ & \mathbb{E}_{P}\left[W\left(1-G\right)s_{3}\left(S\mid X,W=1,G=1\right)\ell^{\prime}\left(y\mid x,s,G=1\right)\right]\\ & \quad\quad=\mathbb{E}_{P}\left[\mathbb{E}_{P}\left[\nu\left(S,X\right)\ell^{\prime}\left(S\mid X,W=1,G=0\right)\mid X,W=1,G=0\right]\mid G=1\right]\nonumber \\ & \mathbb{E}_{P}\left[\left(G\ell^{\prime}\left(X\right)+\gamma^{\prime}\left(X\right)\right)\left(\bar{\nu}\left(x\right)-\theta_{1,1}\right)\right]\\ & \quad\quad=\mathbb{E}_{P}\left[\bar{\nu}\left(X\right)\left(\gamma^{\prime}\left(X\right)+\ell^{\prime}\left(X\right)\gamma\left(X\right)\right)\right]-\mathbb{E}_{P}\left[\gamma^{\prime}\left(X\right)+\ell^{\prime}\left(X\right)\gamma\left(X\right)\right]\theta_{1,1}\nonumber

These are identical to the conditions (ref), (ref), and (ref) and so are verified in the Proof of (ref), Part 2. \qed

Proof of (ref), Part (ii)

Let $\mathcal{P}$ be a regular parametric submodel of $\mathcal{M}_{\lambda}$, indexed by $\varepsilon\in\mathbb{R}$ and such that (ref) hold for each $P_{\varepsilon}\in\mathcal{P}$. Additionally, suppose that \[ p_{\varepsilon}\left(s\mid x,G=1\right)=f\left(s\mid x\right) \] for all $\varepsilon\in\mathbb{R}$ and some fixed function $f\left(s\mid x\right)$.

In this case, we have the score

flalign*\ell^{\prime}\left(b_{1}\right) & =g\cdot\ell^{\prime}\left(y\mid x,s,G=1\right)+w\left(1-g\right)\cdot\ell^{\prime}\left(s\mid W=1,x,G=0\right)\\ & +\ell^{\prime}\left(x\right)+\frac{g-\gamma\left(x\right)}{\gamma\left(x\right)\left(1-\gamma\left(x\right)\right)}\gamma^{\prime}\left(x\right)+\left(1-g\right)\frac{w-\varrho\left(x\right)}{\varrho\left(x\right)\left(1-\varrho\left(x\right)\right)}\varrho^{\prime}\left(x\right)

and so the tangent space $\mathcal{T}$ is given by the mean-square closure of the linear space of the functions

flalign*s\left(b_{1}\right) & =g\cdot s_{1}\left(y\mid x,s,G=1\right)+w\left(1-g\right)s_{3}\left(s\mid W=1,x,G=0\right)\\ & +s_{4}\left(x\right)+\frac{g-\gamma\left(x\right)}{\gamma\left(x\right)\left(1-\gamma\left(x\right)\right)}s_{5}\left(x\right)+\left(1-g\right)\frac{w-\varrho\left(x\right)}{\varrho\left(x\right)\left(1-\varrho\left(x\right)\right)}s_{6}\left(x\right)

where the functions $s_{1}$ through $s_{6}$ range over the space of mean-zero and square integrable functions that additionally satisfy the restrictions ((ref)) and ((ref)). Consider the choices ((ref)) through ((ref)), noting that in this case $s_{2}\left(\cdot\right)$ does not appear, and additionally choose $s_{6}\left(x\right)=0$. These choices are again mean-zero and square integrable, and satisfy the conditions ((ref)) and ((ref)).

Observe that in the the integral representation ((ref)), the quantity $p\left(s\mid x,G=1\right)$ does not appear. Thus, the pathwise derivative of $\theta_{1,1}$ at $0$ on the submodel of $\mathcal{P}$ is again given by (ref) and it will again suffice to verify the conditions ((ref)) through ((ref)). These are identical to the conditions (ref), (ref), and (ref) and so are verified in the Proof of (ref), Part 2. \qed

Proof of (ref), Part (iii)

Let $\mathcal{P}$ be a regular parametric submodel of $\mathcal{M}_{\lambda}$, indexed by $\varepsilon\in\mathbb{R}$ and such that (ref) hold for each $P_{\varepsilon}\in\mathcal{P}$. Additionally, suppose that \[ p_{\varepsilon}\left(G=1\mid X=x\right)=\gamma\left(x\right) \] for all $\varepsilon\in\mathbb{R}$ and some fixed function $\gamma\left(x\right)$.

In this case, we have the score

flalign*\ell^{\prime}\left(b_{1}\right) & =g\cdot\ell^{\prime}\left(y\mid x,s,G=1\right)+g\cdot\ell^{\prime}\left(s\mid x,G=1\right)+w\left(1-g\right)\cdot\ell^{\prime}\left(s\mid W=1,x,G=0\right)\\ & +\ell^{\prime}\left(x\right)+\left(1-g\right)\frac{w-\varrho\left(x\right)}{\varrho\left(x\right)\left(1-\varrho\left(x\right)\right)}\varrho^{\prime}\left(x\right)

and so the tangent space $\mathcal{T}$ is given by the mean-square closure of the linear space of the functions

flalign*s\left(b_{1}\right) & =g\cdot s_{1}\left(y\mid x,s,G=1\right)+g\cdot s_{2}\left(s\mid x,G=1\right)+w\left(1-g\right)s_{3}\left(s\mid W=1,x,G=0\right)\\ & +s_{4}\left(x\right)+\left(1-g\right)\frac{w-\varrho\left(x\right)}{\varrho\left(x\right)\left(1-\varrho\left(x\right)\right)}s_{6}\left(x\right)

where the functions $s_{1}$ through $s_{6}$ range over the space of mean-zero and square integrable functions that additionally satisfy the restrictions ((ref)) through ((ref)). Consider the choices ((ref)) through ((ref)) as well as

flalign*s_{4}\left(x\right) & =\frac{\gamma\left(x\right)\bar{\nu}\left(x\right)}{\pi}-\theta_{1,1}\quadand\quad s_{6}\left(x\right)=0,

which yield the conjectured influence function

flalign*\check{\xi}_{1,1}\left(b_{1}\right) & =\frac{g}{\pi}\frac{\varrho\left(s,x\right)}{\varrho\left(x\right)}\frac{1-\gamma\left(s,x\right)}{\gamma\left(s,x\right)}\frac{\gamma\left(x\right)}{1-\gamma\left(x\right)}\left(y-\nu\left(s,x\right)\right)\\ & +\frac{1}{\pi}\frac{\gamma\left(x\right)}{1-\gamma\left(x\right)}\frac{w\left(\nu\left(s,x\right)-\bar{\nu}\left(x\right)\right)}{\varrho\left(x\right)}+\frac{\gamma\left(x\right)\bar{\nu}\left(x\right)}{\pi}-\theta_{1,1}.

These choices are again mean-zero and square integrable, and satisfy the conditions ((ref)) and ((ref)).

Observe that \[ p_{\varepsilon}\left(x\mid G=1\right)=\frac{\gamma\left(x\right)p_{\varepsilon}\left(x\right)}{\pi}. \] Thus, by the integral representation ((ref)) the pathwise derivative of $\theta_{1,1}$ at $0$ on the submodel of $\mathcal{P}$ is given by \[ \theta_{1,1}^{\prime}=\mathbb{E}_{P}\left[\mathbb{E}_{P}\left[\mathbb{E}_{P}\left[YU\left(Y,S,X\right)\mid S,X,G=1\right]\mid X,W=1,G=0\right]\mid G=1\right], \] where \[ U\left(y,s,x\right)=\ell^{\prime}\left(y\mid s,x,G=1\right)+\ell^{\prime}\left(s\mid x,W=1,G=0\right)+\ell^{\prime}\left(x\right). \] Thus, we have that

flalign*\theta_{1,1}^{\prime} & =\mathbb{E}_{P}\left[\mathbb{E}_{P}\left[\mathbb{E}_{P}\left[Y\ell^{\prime}\left(Y\mid S,X,G=1\right)\mid S,X,G=1\right]\mid X,W=1,G=0\right]\mid G=1\right]\\ & +\mathbb{E}_{P}\left[\mathbb{E}_{P}\left[\nu\left(S.X\right)\ell^{\prime}\left(S\mid X,W=1,G=0\right)\mid X,W=1,G=0\right]\mid G=1\right]\\ & +\pi^{-1}\mathbb{E}_{P}\left[\bar{\nu}\left(X\right)\gamma\left(X\right)\ell^{\prime}\left(X\right)\right].

Note that the first two lines in this expression are unchanged relative to (ref). By the choices of $s_{1}\left(\cdot\right)$ through $s_{6}\left(\cdot\right)$ given above, we have that

flalign*\mathbb{E}\left[\check{\xi}_{1,1}\left(B_{1}\right)\ell^{\prime}\left(B_{1}\right)\right] & =\mathbb{E}_{P}\left[Gs_{1}\left(Y\mid S,X,G=1\right)\ell^{\prime}\left(Y\mid X,S,G=1\right)\right]\\ & +\mathbb{E}\left[W\left(1-G\right)s_{3}\left(S\mid X,W=1,G=1\right)\ell^{\prime}\left(Y\mid X,S,G=1\right)\right]+\mathbb{E}\left[s_{4}\left(X\right)\ell^{\prime}\left(X\right)\right].

The equality \[ \mathbb{E}_{P}\left[\check{\xi}_{1,1}\left(B_{1}\right)\ell^{\prime}\left(B_{1}\right)\right]=\theta_{1,1}^{\prime}, \] follows from the fact that ((ref)) and ((ref)) hold, as $s_{1}\left(\cdot\right)$ and $s_{3}\left(\cdot\right)$ are unchanged relative to the proof of (ref), Part 2, and \[ \mathbb{E}_{P}\left[s_{4}\left(X\right)\ell^{\prime}\left(X\right)\right]=\pi^{-1}\mathbb{E}_{P}\left[\bar{\nu}\left(X\right)\gamma\left(X\right)\ell^{\prime}\left(X\right)\right] \] holds by definition, completing the proof. \qed

Nested Condition Expectation Estimation

In this section, we derive explicit rates of convergence for estimators of the nested condition expectation functions $\bar{\mu}_w(x)$ and $\bar{\nu}_w(x)$, defined in (ref), based on linear sieves. To simplify exposition, we focus attention on the conditional expectation functions

flalign*\mu\left(s,x\right) & =\mathbb{E}_{P}\left[Y\mid S=s,X=x,G=1\right]\quadand\quad\bar{\mu}\left(x\right)=\mathbb{E}_{P}\left[\mu\left(S,x\right)\mid X=x,G=0\right] ,

and let $Z_i=(S_i,X_i)$ collect the short-term outcomes and pretreatment covariates.

Estimation

We consider the following simple linear sieve estimators for $\mu(\cdot,\cdot)$ and $\bar{\mu}(\cdot)$. For functions $h$ and $\bar{h}$ in $L_{2}\left(P_{Z}\right)$ and $L_{2}\left(P_{X}\right)$, define the loss functions

equation[equation omitted — 205 chars of source]

as well as the population risk functions

equation[equation omitted — 212 chars of source]

It is clear that the functions $\mu\left(\cdot,\cdot\right)$ and $\bar{\mu}\left(\cdot\right)$ minimize the risks ((ref)) over $L^{2}\left(P_{Z}\right)$ and $L^{2}\left(P_{X}\right)$.

Throughout, we assume that the functions $\mu(\cdot,\cdot)$ and $\bar{\mu}(\cdot)$ are elements of the linear subspaces $\Theta\subseteq L_{2}\left(P_{Z}\right)$ and $\bar{\Theta}\subseteq L_{2}\left(P_{X}\right)$, respectively. Define the sequences of finite-dimensional parameter spaces $\Theta_{1}\subseteq\Theta_{2}\subseteq\cdots\subseteq\Theta$ and $\bar{\Theta}_{1}\subseteq\bar{\Theta}_{2}\subseteq\cdots\subseteq\bar{\Theta}$. Let $\Pi_{m}:\Theta\to\Theta_{m}$ denote the $L_{2}\left(P_{Z}\right)$ projection of $\Theta$ onto $\Theta_{m}$ and let $\bar{\Pi}_{m}:\bar{\Theta}\to\bar{\Theta}_{m}$ denote the $L_{2}\left(P_{X}\right)$ projection of $\bar{\Theta}$ onto $\bar{\Theta}_{m}$.

Define the empirical criterion and associated sieve extremum estimator

equation[equation omitted — 210 chars of source]

Proceeding analogously, define the infeasible empirical criterion and associated sieve extremum estimator

equation[equation omitted — 284 chars of source]

in addition to the feasible empirical criterion and associated sieve extremum estimator

equation[equation omitted — 316 chars of source]

Consistency

The following assumptions are sufficient to show that the estimators $\hat{\mu}_{m,n}$ and $\quad\hat{\bar{\mu}}_{m,n}$ are consistent in the norm $\|\cdot\|_{P,2}$ for their respective estimands.

assumption$\text{ }$\\ (i) The variables $Z_{i}$ have support $\mathcal{Z}=\mathcal{S\times\mathcal{X}},$ where $\mathcal{S}\subset\mathbb{R}^{d}$ and $\mathcal{X}\subset\mathbb{R}^{k}$ are compact. (ii) The sieve spaces $\Theta_{m}$ and $\bar{\Theta}_{m}$ satisfy \[ \|\Pi_{m}\left(\mu\right)-\mu\|_{P_{Z},2}\to0\quad\text{and}\quad\|\bar{\Pi}_{m}\left(\bar{\mu}\right)-\bar{\mu}\|_{P_{X},2}\to0 \] as $n\to\infty$.

Consistency follows from (ref), which are stated and proved in (ref).

theoremUnder (ref), the estimators $\hat{\mu}_{m,n}$ and $\hat{\bar{\mu}}_{m,n}$ satisfy \[ \|\hat{\mu}_{m,n}-\mu\|_{P,2}\overset{p}{\to}0\quad\text{and}\quad\|\hat{\bar{\mu}}_{m,n}-\bar{\mu}\|_{P,2}\overset{p}{\to}0, \] respectively, as $n,m\to\infty$.
proofBy strict convexity, the population risk function $R_{P}\left(h\right)$ has a unique minimizer over the convex subspace $\Theta_{m}$, denoted by $\mu_{m}^{*}$. We have that $\|\hat{\mu}_{m,n}-\mu_{m}^{*}\|_{P,2}\overset{p}{\to}0$ as $n\to\infty$ and $\|\mu_{m}^{*}-\mu\|_{P,2}\to0$ as $m\to\infty$ by Lemmas (ref) and (ref), respectively. The triangle inequality then gives \[ \|\hat{\mu}_{m,n}-\mu\|_{P,2}\leq\|\hat{\mu}_{m,n}-\mu_{m}^{*}\|_{P,2}+\|\mu_{m}^{*}-\mu\|_{P,2}\overset{p}{\to}0, \] as $m,n\to\infty$. An identical argument establishes the convergence of $\hat{\bar{\mu}}_{m,n}$.

Rate of Convergence

Next, derive explicit rates of convergence for the linear sieve estimators $\hat{\mu}_{m,n}$ and $\hat{\bar{\mu}}_{m,n}$. We discipline this exercise by assuming that the functions $\mu(\cdot)$ and $\bar{\mu}(\cdot)$ are elements of Hölder classes with known parameterizations.

To this end, we fix the following notation. If $\mathcal{X}$ is a compact set in $\mathbb{R}^{d}$ and $0<\gamma\leq1$ is a constant, then a real-valued function $h$ on $\mathcal{X}$ is said to satisfy a Hölder condition with exponent $\gamma$ if there is a positive number $c$ such that $\vert h\left(x\right)-h\left(y\right)\vert\leq c\|x-y\|_{2}^{\gamma}$ for all $x,y\in\mathcal{X}$. For a $d$-dimensional multi-index $\beta$, let $D^{\beta}$ denote the associated Differential operator and let $\vert\beta\vert=\beta_{1}+\cdots+\beta_{d}$. A real-valued function $h$ on $\mathcal{X}$ is said to be $p$-smooth if it is $m$ times continuously differentiable on $\mathcal{X}$, $D^{\beta}h$ satisfies a Hölder condition with exponent $\gamma$ for all $\vert\beta\vert=m$, and $m+\gamma=p$. The Hölder class $\Lambda^{p}\left(\mathcal{X}\right)$ denotes the space of all $p$-smooth functions on $\mathcal{X}$. Let $C^{m}\left(\mathcal{X}\right)$ denote the space of all $m$-times continuously differentiable functions on $\mathcal{X}$. The Hölder ball on $\mathcal{X}$ with width $c$ and smoothness $p$ is given by

flalign*\Lambda_{c}^{p}\left(\mathcal{X}\right) & =\left\{ h\in C^{m}\left(\mathcal{X}\right):\sup_{\vert\beta\vert\leq m}\sup_{x\in\mathcal{X}}\vert D^{\beta}h\left(x\right)\vert\leq c,\sup_{\vert\beta\vert=m}\sup_{x,y\in\mathcal{X},x\neq y}\frac{\vert D^{\beta}h\left(x\right)-D^{\beta}h\left(y\right)\vert}{\vert x-y\vert_{2}^{\gamma}}\vert\leq c\right\} .

We make the following assumptions on the regularity of the data, smoothness of the target functions, the quality of their approximation by the chosen linear sieve spaces.

assumption$\text{ }$ (i) The moment bounds \begin{flalign} \sup_{z\in\mathcal{Z}} \mathbb{E}_{P}\left[\left(Y-\mu\left(z\right)\right)^{2}\mid Z=z,G=1\right] & <C\quadand\\ \sup_{z\in\mathcal{Z}} \mathbb{E}_{P}\left[\left(\mu\left(S,x\right)-\bar{\mu}\left(x\right)\right)^{2}\mid X=x,G=0\right] & <\bar{C}\nonumber \end{flalign} hold for $\lambda$-almost every $z\in\mathcal{Z}$ and $x\in\mathcal{X}$, respectively. (ii) The densities of $X_{i}$ and $Z_{i}$ are uniformly bounded away from zero and infinity. (iii) The function classes $\Theta$ and $\bar{\Theta}$ are given by $\Lambda_{c}^{p}\left(\mathcal{Z}\right)$ and $\bar{\mu}\in\Lambda_{\bar{c}}^{\bar{p}}\left(\mathcal{X}\right)$, respectively. \textbf{(iv) }The sieve spaces $\Theta_{n}$ and $\bar{\Theta}_{n}$ are linear and satisfy \[ \mathsf{dim}\left(\Theta_{n}\right)=J_{n}^{d+k}\quad\text{and}\quad\mathsf{dim}\left(\bar{\Theta}_{n}\right)=J_{n}^{k} \] as well as \[ \inf_{h\in\Theta_{n}}\|h-\mu\|=O\left(J_{n}^{-p}\right)\quad\text{and}\quad\inf_{h\in\Theta_{n}}\|h-\bar{\mu}\|=O\left(\bar{J}_{n}^{-\bar{p}}\right) \] for some sequences $J_{n}$ and $\bar{J}_{n}$, respectively.

The following Theorem, stating rates of convergence for the estimators $\hat{\mu}_{m,n}$ and $\hat{\bar{\mu}}_{m,n}$, is a consequence of Theorem 3.2 of chen2007large. The proof uses Lemma 2 of chen1998sieve, in addition to (ref), which are stated and proved in (ref).

theoremSuppose that (ref), (i), and (ref) hold, and that the sieve spaces $\Theta_{n}$ and $\bar{\Theta}_{n}$ are chosen such that \begin{equation} J_{n}\asymp\left(\frac{n}{\log n}\right)^{\frac{1}{2p+d+k}}\quadand\quad \bar{J}_{n}\asymp\left(\frac{n}{\log n}\right)^{\frac{1}{2\bar{p}+k}}. \end{equation} (i) The estimator $\hat{\mu}_{n,n}$ satisfies \begin{flalign} \|\hat{\mu}_{n,n}-\mu\|_{P,2} & =O_{P}\left(\left(\frac{\log n}{n}\right)^{\frac{p}{2p+d+k}}\right) \end{flalign} as $n\to\infty$. (ii) If \[ \frac{\bar{p}}{2\bar{p}+k}\leq\frac{1}{2}\frac{p}{2p+d+k}, \] then the estimator $\hat{\bar{\mu}}_{n,n}$ satisfies \begin{equation} \|\hat{\bar{\mu}}_{n,n}-\bar{\mu}\|_{P,2}=O_{P}\left(\left(\frac{\log n}{n}\right)^{\frac{p}{2p+d+k}}\right) \end{equation} as $n\to\infty$.
remarkAssumption (ref), (iv), will hold if $\Theta_{n}$ is the $d+k$ tensor product of the space of polynomials of degree $J_{n}$ or less. Similar statements hold for different choices of sieve spaces, e.g., tensor products of univariate splines or orthogonal wavelets. See e.g., Section 2.3 of chen2007large for further discussion. $\blacksquare$
proofFirst, we verify the conditions of Theorem 3.2 of chen2007large, re-stated here as Lemma (ref), for the loss functions $L_{P}\left(\cdot,A\right)$ and $\bar{L}_{P}\left(\cdot,A\right)$, defined in ((ref)). These conditions will allow for the application of (ref) in both cases. For this task, we will make frequent use of the functions $K(y,z)$ and $\bar{K}(s,x)$, defined in (ref). In both cases, Condition (i) follows immediately from the smoothness and convexity of the population risk functions $R_{P}\left(\cdot\right)$ and $\bar{R}_{P}\left(\cdot\right)$ about their minimizers. To verify Condition (ii), fix $h\in\Theta_{n}$ such that $\|h-\mu\|_{P,2}\leq\eta$, and observe that \begin{flalign*} \Var_{P}\left(L\left(h,A\right)-L\left(\mu,A\right)\right) & \leq\mathbb{E}_{P}\left[\left(L\left(h,A\right)-L\left(\mu,A\right)\right)^{2}\right]\\ & \leq\mathbb{E}_{P}\left[\left(\left[K\left(Y,Z\right)\left(h\left(Z\right)-\mu\left(Z\right)\right)\right]\right)^{2}\right]\lesssim\|h-\mu\|_{p,2}^{2}\lesssim\eta^{2}, \end{flalign*} as required. Here, the second and third inequalities follow from Lemma (ref). An identical argument verifies Condition (ii) for the loss $\bar{L}_{P}\left(\cdot,A\right)$. To verify Condition (iii) of Lemma (ref), again fix $h\in\Theta_{n}$ with $\|h-\mu\|_{P,2}\leq\eta$, and observe that \begin{flalign*} \vert L\left(h,A\right)-L\left(\mu,A\right)\vert & \leq K\left(Y,Z\right)\|h-\mu\|_{\infty},\\ & \lesssim K\left(Y,Z\right)\|h-\mu\|_{\lambda,2}^{\frac{2p}{2p+d+k}}\lesssim K\left(Y,Z\right)\|h-\mu\|_{P,2}^{\frac{2p}{2p+d+k}} , \end{flalign*} where the first inequality follows from Lemma (ref), the second inequality follows from Lemma (ref), and the third inequality follows from (ref), Part (ii), which implies that $\|\cdot\|_{\lambda,2}$ and $\|\cdot\|_{P,2}$ are equivalent norms on $\mathcal{Z}$. Similarly, if $\bar{h}\in\bar{\Theta}_{n}$ with $\|\bar{h}-\bar{\mu}\|_{P,2}\leq\eta$, then \begin{flalign*} \vert\bar{L}\left(\bar{h},A\right)-\bar{L}\left(\bar{\mu},A\right)\vert & \leq\bar{K}\left(S,X\right)\|\bar{h}-\bar{\mu}\|_{\infty},\\ & \lesssim\bar{K}\left(S,X\right)\|\bar{h}-\bar{\mu}\|_{\lambda,2}^{\frac{2\bar{p}}{2\bar{p}+k}}\lesssim\bar{K}\left(S,X\right)\|\bar{h}-\bar{\mu}\|_{P,2}^{\frac{2\bar{p}}{2\bar{p}+k}} , \end{flalign*} again by Lemmas (ref) and (ref) and (ref), Part (ii). Thus, in the case of $L_{P}\left(\cdot,A\right)$, Condition (iii) is verified by setting $U\left(A\right)=C\cdot K\left(Y,Z\right)$ for some constant $C$ and $s=\frac{2p}{2p+d+k}$. Similarly, in the case of $\bar{L}_{P}\left(\cdot,A\right)$, Condition (iii) is verified by setting $U\left(A\right)=\bar{C}\cdot\bar{K}\left(S,X\right)$ for some constant $\bar{C}$ and $s=\frac{2\bar{p}}{2\bar{p}+k}$. In both cases, the required bound deterministic on $\mathbb{E}\left[U\left(A\right)^{2}\right]$ follows from Assumption (ref), Part (i), and the definitions of $K\left(Y,Z\right)$ and $\bar{K}\left(S,X\right)$. Now, we apply (ref) to verify the rate ((ref)) for the estimator $\hat{\mu}_{n,n}$. Observe that (ref), Part (iv) implies (ref), Part (ii). Consequently, we have that $\|\hat{\mu}_{n,n}-\mu\|=o_{P}\left(1\right)$ as $n\to\infty$ by Theorem (ref). Thus, as $\hat{\mu}_{n,m}$ is a sequence of estimators that satisfy \[ \frac{1}{n}R_{n}\left(\hat{\mu}_{n,n}\right)=\sup_{h\in\Theta_{n}}\frac{1}{n}R_{n}\left(h\right)~, \] the rate ((ref)) follows from (ref), Part (iv), and Lemma (ref). It remains to verify the rate ((ref)) for the estimator $\hat{\bar{\mu}}_{n,n}$. This result is achieved through an inductive application of (ref). First note that, (ref), Part (iv) again implies that $\|\hat{\bar{\mu}}_{n,n}-\bar{\mu}\|=o_{P}\left(1\right)$ as $n\to\infty$ by Theorem (ref). Observe that the rate ((ref)) implies that for any $\bar{h}\in\bar{\Theta}_{n}$, \begin{flalign} \frac{1}{n}\bar{R}_{n}\left(\bar{h}\right) & =\frac{1}{n}\sum_{i=1}^{n}\left(1-G_{i}\right)\left(\left(\mu\left(S_{i},X_{i}\right)-\hat{\mu}_{n,n}\left(S_{i},X_{i}\right)\right)+\left(\hat{\mu}_{n,n}\left(S_{i},X_{i}\right)-\bar{h}\left(X_{i}\right)\right)\right)^{2}\nonumber \\ & =\frac{1}{n}\bar{R}_{n,n}\left(\bar{h}\right)+\frac{1}{n}\sum_{i=1}^{n}\left(1-G_{i}\right)\left(\mu\left(S_{i},X_{i}\right)-\hat{\mu}_{n,n}\left(S_{i},X_{i}\right)\right)^{2}\nonumber \\ & +\frac{2}{n}\sum_{i=1}^{n}\left(1-G_{i}\right)\left(\mu\left(S_{i},X_{i}\right)-\hat{\mu}_{n,n}\left(S_{i},X_{i}\right)\right)\left(\hat{\mu}_{n,n}\left(S_{i},X_{i}\right)-\bar{h}\left(X_{i}\right)\right)\nonumber\\ & =\frac{1}{n}\bar{R}_{n,n}\left(\bar{h}\right)+O_{P}\left(\left(\frac{\log n}{n}\right)^{\frac{2p}{2p+d+k}}\right)\nonumber \\ & +\frac{2}{n}\sum_{i=1}^{n}\left(1-G_{i}\right)\left(\mu\left(S_{i},X_{i}\right)-\hat{\mu}_{n,n}\left(S_{i},X_{i}\right)\right)\left(\hat{\mu}_{n,n}\left(S_{i},X_{i}\right)-\bar{h}\left(X_{i}\right)\right). \end{flalign} The term ((ref)) can be written $C_n + D_n(\bar{h})$, where \begin{align} C_n &=\frac{2}{n}\sum_{i=1}^{n}\left(1-G_{i}\right)\left(\mu\left(S_{i},X_{i}\right)-\hat{\mu}_{n,n}\left(S_{i},X_{i}\right)\right)\cdot \left(\hat{\mu}_{n,n}\left(S_{i},X_{i}\right)-\bar{\mu}\left(X_{i}\right)\right)\quadand \nonumber\\ D_n(\bar{h}) &= \frac{2}{n}\sum_{i=1}^{n} \left(1-G_{i}\right)\left(\mu\left(S_{i},X_{i}\right)-\hat{\mu}_{n,n}\left(S_{i},X_{i}\right)\right) \cdot \left(\bar{\mu}\left(X_{i}\right)-\bar{h}\left(X_{i}\right)\right)\nonumber . \end{align} Observe that the term $C_n$ does not depend on $\bar{h}$. Consequently, if we define the auxillary loss function \begin{align} \tilde{R}_n(\bar{b}) = \bar{R}_{n}\left(\bar{h}\right) - nC_n ,\nonumber \end{align} then the minimization problem is unchanged, in the sense that \begin{equation} \hat{\bar{\mu}}_{m,n}^{*}=\underset{h\in\Theta_{m}}{\arg\min }\tilde{R}_{n}\left(\bar{h}\right) ,\nonumber \end{equation} as in (ref), and we obtain the expansion \begin{align} \frac{1}{n}\tilde{R}_{n}\left(\bar{h}\right) =\frac{1}{n}\bar{R}_{n,n}\left(\bar{h}\right) + D_n(\bar{h}) + O_{P}\left(\left(\frac{\log n}{n}\right)^{\frac{2p}{2p+d+k}}\right) , \end{align} from the expression (ref). Observe that the relations \begin{flalign} \frac{1}{n}\bar{R}_{n,n}\left(\hat{\bar{\mu}}_{n,n}\right) & =\frac{1}{n}\tilde{R}_{n}\left(\hat{\bar{\mu}}_{n,n}\right)-D_n(\hat{\bar{\mu}}_{n,n})+O_{P}\left(\left(\frac{\log n}{n}\right)^{\frac{2p}{2p+d+k}}\right),\\ \frac{1}{n}\bar{R}_{n,n}\left(\hat{\bar{\mu}}_{n,n}\right) & \geq\frac{1}{n}\bar{R}_{n,n}\left(\hat{\bar{\mu}}_{n,n}^{*}\right),\quadand\nonumber \\ \frac{1}{n}\bar{R}_{n,n}\left(\hat{\bar{\mu}}_{n,n}^{*}\right) & =\frac{1}{n}\tilde{R}_{n}\left(\hat{\bar{\mu}}_{n,n}^{*}\right)-D_n(\hat{\bar{\mu}}_{n,n}^{*})+O_{P}\left(\left(\frac{\log n}{n}\right)^{\frac{2p}{2p+d+k}}\right), \end{flalign} where $\hat{\bar{\mu}}_{n,n}$ is defined in ((ref)), imply the inequality \begin{equation} \frac{1}{n}\bar{R}_{n}\left(\hat{\bar{\mu}}_{n,n}\right)\geq \frac{1}{n}\bar{R}_{n}\left(\hat{\bar{\mu}}_{n,n}^{*}\right) +\left(D_n(\hat{\bar{\mu}}_{n,n}) - D_n(\hat{\bar{\mu}}_{n,n}^{*})\right) +O_{P}\left(\left(\frac{\log n}{n}\right)^{\frac{2p}{2p+d+k}}\right) . \end{equation} Moreover, the term $D_n(\hat{\bar{\mu}}_{n,n}^{*}))$ is bounded from above by \begin{equation} \frac{2}{n}\sum_{i=1}^{n}\|\mu\left(S_{i},X_{i}\right)-\hat{\mu}_{n,n}\left(S_{i},X_{i}\right)\|_{2}\|\bar{\mu}\left(X_{i}\right)-\bar{h}\left(X_{i}\right)\|_{2}=O_{P}\left(\left(\frac{\log n}{n}\right)^{\frac{p}{2p+d+k}+\frac{\bar{p}}{2\bar{p}+k}}\right) , \end{equation} by the Cauchy-Schwarz inequality, the rate ((ref)), and the choice (ref). Plugging this bound into (ref) implies \begin{equation} \frac{1}{n}\bar{R}_{n}\left(\hat{\bar{\mu}}_{n,n}\right)\geq \frac{1}{n}\bar{R}_{n}\left(\hat{\bar{\mu}}_{n,n}^{*}\right) +D_n(\hat{\bar{\mu}}_{n,n}) +O_{P}\left(\left(\frac{\log n}{n}\right)^{\frac{p}{2p+d+k} + \min\big\{ \frac{p}{2p+d+k},\frac{\bar{p}}{2\bar{p}+k}\big\}}\right) . \end{equation} Thus, in order to apply (ref), must must control the term $D_n(\hat{\bar{\mu}}_{n,n})$. To this end, suppose that it were known that \begin{align} \|\hat{\bar{\mu}}_{n,n}-\bar{\mu}\|_{2,P}=O_{P}\left(\left(\frac{\log n}{n}\right)^{s}\right) \end{align} for some constant $0\leq s <1$. In this case, the term $D_n(\hat{\bar{\mu}}_{n,n})$ is similarly bounded from above by \begin{equation} \frac{2}{n}\sum_{i=1}^{n}\|\mu\left(S_{i},X_{i}\right)-\hat{\mu}_{n,n}\left(S_{i},X_{i}\right)\|_{2}\|\bar{\mu}\left(X_{i}\right)-\bar{h}\left(X_{i}\right)\|_{2}=O_{P}\left(\left(\frac{\log n}{n}\right)^{\frac{p}{2p+d+k}+s}\right) , \end{equation} by the Cauchy-Schwarz inequality and the rate ((ref)), which again implies \begin{equation} \frac{1}{n}\bar{R}_{n}\left(\hat{\bar{\mu}}_{n,n}\right)\geq \frac{1}{n}\bar{R}_{n}\left(\hat{\bar{\mu}}_{n,n}^{*}\right) +O_{P}\left(\left(\frac{\log n}{n}\right)^ {\frac{p}{2p+d+k} + \min\big\{s,\frac{p}{2p+d+k},\frac{\bar{p}}{2\bar{p}+k}\big\}}\right) . \end{equation} Hence, by (ref), Part (iv), Lemma (ref) implies that \begin{align} \|\hat{\bar{\mu}}_{n,n}-\bar{\mu}\|_{2,P}=O_{P}\left(\left(\frac{\log n}{n}\right)^{r(s)}\right) , \end{align} where \begin{align} r(s) = \max\bigg\{ \frac{1}{2}\left(\frac{p}{2p+d+k} + \min\bigg\{s,\frac{p}{2p+d+k},\frac{\bar{p}}{2\bar{p}+k}\bigg\}\right), \frac{\bar{p}}{2\bar{p}+k} \bigg\} . \end{align} Observe that we have shown $\bar{\mu}_{n,n}$ converges to $\bar{\mu}$ in $\|\cdot\|_{2,P}$ at rate $s$, then it also converges at the rate $r(s)$. To conclude the proof, observe that (ref) must be satisfied by $s=0$, as $\hat{\bar{\mu}}_{n,n}$ and $\bar{\mu}$ are subsets of $\Lambda^{\bar{p}}_{\bar{c}(\mathcal{X})}$. Thus, if we define the sequence $(s_i)_{i=0}^\infty$ by $s_i = r(s_{i-1})$, with initial condition $s_0=1$, we find that \[ \lim_{i\to\infty} s_i = \frac{p}{2p+d+k}\quad\text{if}\quad\frac{\bar{p}}{2\bar{p}+k}\leq\frac{1}{2}\frac{p}{2p+d+k}~, \] and so the rate of covergence (ref) is obtained by induction.

Auxiliary Lemmas

lemma[{newey1994large}] If there is a function $Q_{0}\left(\theta\right)$ defined on a convex set $\Theta$ such that (i) $Q_{0}\left(\theta\right)$ is uniquely minimized at $\theta_{0}$, (ii) $\theta_{0}$ is an element of the interior of $\Theta$, (ii) the function $\hat{Q}_{n}\left(\theta\right)$ defined on $\Theta$ is concave, and (iv) $\hat{Q}_{n}\left(\theta\right)\to Q_{0}\left(\theta\right)$ for all $\theta\in\Theta$, then \[ \hat{\theta}_{n}=\arg\max_{\theta\in\Theta}\text{ }\hat{Q}_{n}\left(\theta\right) \] exists with probability one and $\hat{\theta}_{n}\overset{p}{\to}\theta_{0}$ as $n\to\infty$.
lemmaUnder (ref), the estimators $\hat{\mu}_{m,n}$ and $\hat{\bar{\mu}}_{m,n}$ satisfy \[ \|\hat{\mu}_{m,n}-\mu_{m}^{*}\|_{P,2}\overset{p}{\to}0\quad\text{and}\quad\|\hat{\bar{\mu}}_{m,n}-\bar{\mu}_{m}^{*}\|_{P,2}\overset{p}{\to}0 \] as $n\to\infty$ and $m$ is sufficiently large, respectively.
proofIt will suffice to verify the conditions of Lemma (ref). Note first that both $\Theta_{m}$ and $\bar{\Theta}_{m}$ are convex. The population criteria $R_{P}\left(\cdot\right)$ and $\bar{R}_{P}\left(\cdot\right)$ are both uniquely minimized at $\mu_{m}^{*}$ and $\bar{\mu}_{m}^{*}$ over $\Theta_{m}$ and $\bar{\Theta}_{m}$, respectively, by convexity. Both $\mu_{m}^{*}$ and $\bar{\mu}_{m}^{*}$ are elements of the interiors of $\Theta_{m}$ and $\bar{\Theta}_{m}$ for sufficiently large $m$ by Lemma (ref). It is clear that the functions $R_{n}\left(h\right)$ and $\bar{R}_{m,n}\left(\bar{h}\right)$ are strictly convex. The pointwise convergence \[ R_{n}\left(h\right)\overset{\text{p}}{\to}R_{p}\left(h\right) \] as $n\to\infty$ for $m$ sufficiently large follows from the weak law of large numbers, and so \begin{equation} \|\hat{\mu}_{m,n}-\mu_{m}^{*}\|_{P,2}\overset{p}{\to}0 \end{equation} as $n\to\infty$ for $m$ sufficiently large follows from Lemma (ref). Consequently, the pointwise convergence \[ \bar{R}_{n}\left(\bar{h}\right)\overset{\text{p}}{\to}\bar{R}_{p}\left(\bar{h}\right) \] follows from ((ref)) and the weak law of large numbers, again giving that $\|\hat{\bar{\mu}}_{m,n}-\bar{\mu}_{m}^{*}\|_{P,2}\overset{p}{\to}0$ as $n\to\infty$ for $m$ sufficiently large by Lemma (ref).
lemmaUnder (ref), the unique minimizers $\mu_{m}^{*}$ and $\bar{\mu}_{n}^{*}$ of $R_{P}\left(h\right)$ and $\bar{R}_{P}\left(h\right)$ over $\Theta_{m}$ and $\bar{\Theta}_{m}$ satisfy \[ \|\mu_{m}^{*}-\mu\|_{P,2}\to0\quad\text{and}\quad\|\bar{\mu}_{n}^{*}-\bar{\mu}\|_{P,2}\to0 \] as $m\to\infty$, respectively.
proofWe prove only $\|\mu_{m}^{*}-\mu\|_{P,2}\to0$ as $m\to\infty$, as the argument for $\|\bar{\mu}_{n}^{*}-\bar{\mu}\|_{P,2}\to0$ as $m\to\infty$ is identical. Assume, by the way of contradiction, that there exists some $\delta>0$ such that $\|\mu_{m}^{*}-\mu\|_{P,2}>\delta$ for all $m$ greater than some integer $m_{1}$. As $R_{P}\left(\cdot\right)$ is strictly convex on $\Theta$, we have that \begin{equation} R_{P}\left(\mu_{m}^{*}\right)>R_{P}\left(\mu\right)+\varepsilon \end{equation} for all $m\geq\max\left\{ m_{0},m_{1}\right\} $ for some $\varepsilon>0$. By the continuity of $R_{P}\left(\cdot\right)$, there exists some $\delta^{\prime}>0$ such that $\|\mu^{\prime}-\mu\|_{P,2}<\delta^{\prime}$ implies \begin{equation} R_{P}\left(\mu^{\prime}\right)<R_{P}\left(\mu\right)+\varepsilon. \end{equation} By (ref), (ii), there exists some integer $m_{1}$ such that $\|\Pi_{m}\left(\mu\right)-\mu\|_{P,2}<\delta^{\prime}$ for all $m\geq m_{1}$. Thus, by ((ref)) and ((ref)), we have that \[ R_{P}\left(\Pi_{m}\left(\mu\right)\right)<R_{P}\left(\mu\right)+\varepsilon<R_{P}\left(\mu_{m}^{*}\right), \] for all $m\geq\max\left\{ m_{0},m_{1}\right\} $, giving a contradiction, as $\mu_{m}^{*}$ minimizes $R_{P}\left(\cdot\right)$ over $\Theta_{m}$.
lemma[{chen2007large}] Let $\left\{ X_{i}\right\} _{i=1}^{n}$ be an independent sequence of random variables with an identical distribution $P$. Fix a sieve $\Theta_{1}\subseteq\Theta_{2}\subseteq\cdots\subseteq\Theta$. For some loss function $l\left(x,\theta\right)$, let $\theta^{*}\in\Theta$ be the population risk minimizer \[ \theta^{*}=\underset{\theta\in\Theta}{\arg\min}\text{ }\mathbb{E}_{P}\left[l\left(\theta,X_{i}\right)\right] \] and let $\|\cdot\|$ be some norm on $\Theta$. Define the function class \[ \mathcal{F}_{n}=\left\{ l\left(\theta,\cdot\right)-l\left(\theta^{*},\cdot\right):\|\theta-\theta^{*}\|\leq\delta,\theta\in\Theta_{n}\right\} . \] For some constant $b>0$, let \[ \delta_{n}=\underset{\delta\in\left(0,1\right)}{\inf}\frac{1}{\sqrt{n}\delta^{2}}\int_{b\delta^{2}}^{\delta}\sqrt{H_{[]}\left(w,\mathcal{F}_{n},\|\cdot\|\right)}\text{d}w\leq1, \] where $H_{[]}\left(w,\mathcal{F}_{n},\|\cdot\|\right)$ is the $L^{r}\left(P\right)$ metric entropy with bracketing of the class $\mathcal{F}_{n}$. Assume that the following conditions hold: (i) There exists constants $c_{1},c_{2}>0$ such that \[ c_{1}\mathbb{E}\left[l\left(\theta,X_{i}\right)-l\left(\theta^{*},X_{i}\right)\right]\leq\|\theta-\theta^{*}\|^{2}\leq c_{2}\mathbb{E}\left[l\left(\theta,X_{i}\right)-l\left(\theta^{*},X_{i}\right)\right] \] for any $\theta\in\Theta$. (ii) There is a constant $C_{1}>0$ such that for sufficiently small $\eta>0$, \[ \sup_{\theta\in\Theta_{n}:\|\theta-\theta^{*}\|\leq\eta}\Var\left(l\left(\theta,X_{i}\right)-l\left(\theta^{*},X_{i}\right)\right)\leq C_{1}\eta^{2}. \] (iii) For any $\eta>0$, there exists a constant $s\in\left(0,2\right)$ such that \[ \sup_{\theta\in\Theta_{n}:\|\theta-\theta^{*}\|\leq\eta}\vert l\left(\theta,X_{i}\right)-l\left(\theta^{*},X_{i}\right)\vert\leq\eta^{s}U\left(X_{i}\right) \] with $\mathbb{E}\left[U\left(Z_{i}\right)^{\gamma}\right]\leq C_{2}$ for some $\gamma\geq2$. If $\hat{\theta}_{n}$ is a sequence of estimators satisfying \begin{flalign} \|\hat{\theta}_{n}-\theta^{*}\| & =o_{P}\left(1\right),\\ \frac{1}{n}\sum_{i=1}^{n}l\left(\hat{\theta}_{n},X_{i}\right) & \geq\sup_{\theta\in\Theta_{n}}\frac{1}{n}\sum l\left(\theta,X_{i}\right)+O_{P}\left(\varepsilon_{n}^{2}\right),\quad and\\ \max\left\{ \delta_{n},\inf_{\theta\in\Theta_{n}}\|\theta^{*}-\theta\|\right\} & =O\left(\varepsilon_{n}\right), \end{flalign} as $n\to\infty$ for some sequence $\varepsilon_{n}$, then $\|\hat{\theta}_{n}-\theta^{*}\|=O_{P}\left(\varepsilon_{n}\right)$ as $n\to\infty$.
lemmaSuppose that (ref) holds. (i) There exists a function $K\left(Y,Z\right)$ and a constant $C$ such that such that \[ \vert L\left(h,a\right)-L\left(\mu,a\right)\vert\leq K\left(y,z\right)\vert h\left(z\right)-\mu\left(z\right)\vert, \] for each $a=\left(y,z,g\right)\in\mathbb{R}\times\mathcal{S}\times\left\{ 0,1\right\} $, where \[ \sup_{z\in\mathcal{Z}}\text{ }\mathbb{E}_{P}\left[K\left(Y,Z\right)^{2}\mid Z=z\right]\leq C. \] (ii) There exists a function $\bar{K}\left(S,Z\right)$ and a constant $\bar{C}$ such that such that \[ \vert\bar{L}\left(\bar{h},a\right)-\bar{L}\left(\bar{\mu},a\right)\vert\leq\bar{K}\left(s,z\right)\vert\bar{h}\left(x\right)-\bar{\mu}\left(x\right)\vert, \] for each $a=\left(y,z,g\right)\in\mathbb{R}\times\mathcal{S}\times\left\{ 0,1\right\} $, where \[ \sup_{x\in\mathcal{X}}\text{ }\mathbb{E}_{P}\left[\bar{K}\left(S,X\right)^{2}\mid X=x\right]\leq\bar{C}. \]
proofWe provide the details of the verification of the first claim, as verification of the second claim follows by an identical argument. For any $h\in\Lambda_{c}^{p}\left(\mathcal{Z}\right)$ and $z\in\mathcal{Z}$, we have that \begin{flalign*} \vert L\left(h,a\right)-L\left(\mu,a\right)\vert & =g\vert\left(y-h\left(z\right)\right)^{2}-\left(y-\mu\left(z\right)\right)^{2}\vert\\ & =g\cdot\vert2\left(y-\lambda\left(z\right)h\left(z\right)+\left(1-\lambda\left(z\right)\right)\mu\left(z\right)\right)\vert\cdot\vert h\left(z\right)-\mu\left(z\right)\vert \end{flalign*} for some function $\lambda\left(z\right)\in\left[0,1\right]$, by the Mean Value Theorem. Thus, we can set \[ K\left(y,z\right)=g\cdot\vert2\left(y-\lambda\left(z\right)h\left(z\right)+\left(1-\lambda\left(z\right)\right)\mu\left(z\right)\right)\vert. \] We can verify that, for some constant $C$, we have that \begin{flalign*} \mathbb{E}_{P}\left[K\left(Y,Z\right)^{2}\mid Z=z\right] & =4\mathbb{E}_{P}\left[G\left(Y-\lambda\left(Z\right)h\left(Z\right)+\left(1-\lambda\left(Z\right)\right)\mu\left(Z\right)\right)^{2}\mid Z=z\right]\\ & =\mathbb{E}_{P}\left[G\left(\left(Y-\mu\left(Z\right)\right)-\lambda\left(Z\right)\left(h\left(Z\right)-\mu\left(Z\right)\right)\right)^{2}\mid Z=z\right]\leq C \end{flalign*} for all $z\in\mathcal{Z}$, by the bound ((ref)) in Assumption (ref) and the fact that both $h\left(z\right)$ and $\mu\left(z\right)$ are in $\Lambda_{c}^{p}\left(\mathcal{Z}\right)$, and are therefore bounded on $\mathcal{Z}$.
lemma[{chen1998sieve}] If $\mathcal{X}$ is a compact subset of $\mathbb{R}^{d}$ and $h\in\Lambda_{c}^{p}\left(\mathcal{X}\right)$, then \[ \|h\|_{\infty}\leq2\|h\|_{2}^{\frac{2p}{2p+d}}c^{1-\frac{2p}{2p+d}}, \] where $\|\cdot\|_{\infty}$ denotes the $L_{\infty}$ norm and $\|\cdot\|_{2}$ denotes the $L_{2}$ norm under the Lebesgue measure.
lemmaContinue the notation of Lemma (ref). Assume that the conditions (i)-(iii) of Lemma (ref) are satisfied. Suppose that the function class $\Theta=\Lambda_{c}^{p}\left(\mathcal{X}\right)$, where $\mathcal{X}$ is some compact subset of $\mathbb{R}^{d}$, that the sieve spaces $\Theta_{n}$ are linear and satisfy \begin{equation} \mathsf{dim}\left(\Theta_{n}\right)\asymp\left(\frac{n}{\log n}\right)^{\frac{d}{2p+d}}\quadand\quad\inf_{h\in\Theta_{n}}\|h-\mu\|\asymp\left(\frac{\log n}{n}\right)^{\frac{p}{2p+d}}. \end{equation} If $\hat{\theta}_{n}$ is a sequence of estimators satisfying the consistency condition ((ref)) and optimization condition ((ref)) for the sequence $\varepsilon_{n}$, then \begin{equation} \|\hat{\theta}_{n}-\theta^{*}\|_{P,2}=O_{P}\left(\max\left\{ \varepsilon_{n},\left(\frac{\log n}{n}\right)^{\frac{p}{2p+d}}\right\} \right) \end{equation} as $n\to\infty$.
proofAs the conditions of Lemma (ref) are satisfied, we have that \begin{flalign} \|\hat{\theta}_{n}-\theta^{*}\|_{P,2} & =O_{P}\left(\max\left\{ \varepsilon_{n},\delta_{n},\inf_{h\in\Theta_{n}}\|h-\mu\|\right\} \right)\nonumber \\ & =O_{P}\left(\max\left\{ \varepsilon_{n},\delta_{n},\left(\frac{\log n}{n}\right)^{\frac{p}{2p+d}}\|\right\} \right), \end{flalign} where \[ \delta_{n}=\underset{\delta\in\left(0,1\right)}{\inf}\frac{1}{\sqrt{n}\delta^{2}}\int_{b\delta^{2}}^{\delta}\sqrt{H_{[]}\left(w,\mathcal{F}_{n},\|\cdot\|_{2}\right)}\text{d}w\leq1 \] and $H_{[]}\left(w,\mathcal{F}_{n},\|\cdot\|_{2}\right)$ is the $L^{2}\left(P\right)$ metric entropy with bracketing of the function class \[ \mathcal{F}_{n}=\left\{ L\left(h,\cdot\right)-L\left(\mu,\cdot\right):\|h-\mu\|_{P,2}\leq\delta,h\in\Theta_{n}\right\} . \] Now, Condition (iii) of (ref) implies that \[ H_{[]}\left(w,\mathcal{F}_{n},\|\cdot\|_{2}\right)\leq\log N\left(w^{1+\frac{d}{2p}},\Theta_{n},\|\cdot\|_{P,2}\right), \] where $N\left(w,\mathcal{F}_{n},\|\cdot\|_{2}\right)$ is the covering number of $\mathcal{F}_{n}$. Moreover, we have that \[ \log N\left(w^{1+\frac{d}{2p}},\mathcal{F}_{n},\|\cdot\|_{2}\right)\lesssim\mathsf{dim}\left(\Theta_{n}\right)\log\left(\frac{1}{w}\right) \] as $\Theta_{n}$ is an element of a finite dimensional linear sieve (see e.g., Section 3.2 of chen2007large). Thus, we have that \begin{flalign*} \frac{1}{\sqrt{n}\delta^{2}}\int_{b\delta^{2}}^{\delta}\sqrt{H_{[]}\left(w,\mathcal{F}_{n},\|\cdot\|_{2}\right)}dw & \leq\frac{1}{\sqrt{n}\delta^{2}}\int_{b\delta^{2}}^{\delta}\sqrt{\log N\left(w^{1+\frac{d}{2p}},\Theta_{n},\|\cdot\|_{P,2}\right)}dw\\ & \lesssim\frac{1}{\delta}\sqrt{\frac{\mathsf{dim}\left(\Theta_{n}\right)}{n}\log\frac{1}{\delta}} \end{flalign*} and that therefore \begin{equation} \delta_{n}=O\left(\sqrt{\frac{\mathsf{dim}\left(\Theta_{n}\right)\log n}{n}}\right)=O\left(\left(\frac{\log n}{n}\right)^{\frac{p}{2p+d}}\right) \end{equation} by Assumption (ref) and the choice ((ref)). The proof is then complete by plugging the rate ((ref)) into the bound ((ref)).