Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Multiply Robust Causal Mediation Analysis with Continuous Treatments
\def\spacingset#1{
{#1}} \spacingset{1}
}
abstractIn many applications, researchers are interested in the direct and indirect causal effects of a treatment or exposure on an outcome of interest. Mediation analysis offers a rigorous framework for identifying and estimating these causal effects. For binary treatments, efficient estimators for the direct and indirect effects are presented by tchetgen2012semiparametric based on the influence function of the parameter of interest. These estimators possess desirable properties such as multiple-robustness and asymptotic normality while allowing for slower than root-n rates of convergence for the nuisance parameters. However, in settings involving continuous treatments, these influence function-based estimators are not readily applicable without making strong parametric assumptions.
In this work, utilizing a kernel-smoothing approach, we propose an estimator suitable for settings with continuous treatments inspired by the influence function-based estimator of tchetgen2012semiparametric. Our proposed approach employs cross-fitting, relaxing the smoothness requirements on the nuisance functions and allowing them to be estimated at slower rates than the target parameter. Additionally, similar to influence function-based estimators, our proposed estimator is multiply robust and asymptotically normal, allowing for inference in settings where parametric assumptions may not be justified.
{\bf Keywords:} Causal inference; Debiased estimator; Kernel smoothing; Semiparametric estimators.
\spacingset{1.7}
Introduction
Estimating the causal effect of a treatment, policy, or intervention on an outcome of interest is a fundamental task in various fields such as epidemiology, economics, medicine, and sociology. A common parameter of interest is the average causal effect (ACE), which has been extensively studied hernan2020causal, imbens2015causal. However, in addition to estimating the ACE, one can also be interested in the pathways and mechanisms through which the treatment affects the outcome of interest. Causal mediation analysis offers a precise and rigorous mathematical framework to answer such questions robins1992identifiability, tchetgen2012semiparametric,
pearl2001direct,
vanderweele2009marginal, goetgeluk2008estimation, imai2010identification, van2008direct, lange2011direct, lange2012simple.
Much of the literature on mediation analysis assumes that the treatment of interest is binary. However, interventions involving the dosage of a drug, and the duration or frequency of an activity are better described as continuous variables. In such cases, mediation effects are naturally represented by a multi-dimensional surface rather than a scalar parameter.
This learning task is challenging if a priori shape constraints are not imposed on the surface.
Additionally, the presence of continuous treatments complicates the estimation of nuisance parameters, making the estimation of causal parameters more challenging.
The challenges related to estimating ACE in the continuous treatment setting have been addressed in multiple work kennedy2017nonparametric, AI2021, imbens-continuous, kreif2015evaluation, 10.2307/2673642, su2019non,kallus2018policy, colangelo2020double, hill2011bayesian. A common method
is based on outcome regression, which requires the correct specification of the relevant models, and hence machine learning methods such as Bayesian additive regression trees (BART) hill2011bayesian are often used. However, this inherits the rate of the outcome regression estimation, and complex machine learning methods tend to have a slower convergence rate than simple parametric methods wasserman2006all, tsybakov2009nonparametric. An alternative approach involves specifying a parametric form for the dose-response curve or projecting the true curve onto a parametric model, as presented in robins2000marginal, van1998locally, neugebauer2007nonparametric. However, these methods may suffer from bias when the dose-response curve is misspecified. In contrast to approaches involving parametric assumptions on the dose-response curve, kennedy2017nonparametric leverage semiparametric theory by utilizing a two-stage estimator that first constructs a doubly robust pseudo-outcome in the first stage, and then regresses the pseudo-outcome on the treatment in the second stage using non-parametric regression methods. colangelo2020double utilize double machine learning along with applying kernel smoothing to the augmented inverse propensity weighted (AIPW) score robins1995semiparametric. This allows a slower convergence rate of nuisance parameters, while still guaranteeing fast rates for the target parameter.
However, these approaches are not investigated for mediation analysis in the presence of continuous treatments.
In this paper, we propose a kernel smoothing approach inspired by influence function-based estimators tsiatis2007semiparametric,newey1994asymptotic, bickel1993efficient,ichimura2015influence, tchetgen2012semiparametric to deal with continuous treatments for causal mediation analysis. We propose an estimator that, under mild regularity conditions, is consistent and asymptotically normal. Our work aims to extend the results for the continuous treatment ACE to the case of mediation analysis involving continuous treatments in the presence of high-dimensional nuisance functions. huber2020direct tackle this problem by weighting the observations with a generalized propensity score that involves two nuisance functions, which are the conditional density of treatment given covariates and the conditional density of treatment given mediators and covariates.
In their proposed approach, the nuisance functions can be estimated either parametrically or non-parametrically.
However, their estimator for the causal parameter is not robust with respect to the misspecification of the two nuisance functions and also inherits the rate of the nuisance function estimators, which could be slow. In contrast, we propose an approach motivated by influence functions and hence obtain many of the desirable properties of influence functions, namely allowing for slower estimation of nuisance functions, as well as robustness properties. Our work draws heavily on the existing causal mediation literature that discusses the identification and estimation of such effects pearl2001direct, imai2010identification, tchetgen2012semiparametric. Additionally, we utilize the cross-fitting strategy to relax the smoothness assumptions on the nuisance functions chernozhukov2018double.
The remainder of this paper is organized as follows. Section (ref) introduces the formal mediation framework, describes its identifying assumptions, and discusses an influence function-based estimator of mediation effects for binary treatments. Section (ref) extends the influence function-based approach to continuous treatment settings and describes the sample-splitting and smoothing procedures. In Section (ref), we provide our main results along with the required regularity conditions. Sections 5 and 6 present simulation and application results. We provide conclusions and discussion of future directions in Section 7.
Mediation Analysis
figure[figure omitted — 2,171 chars of source]
Let $A\in\mathcal{A}$ be the continuous treatment variable, $Y\in\mathcal{Y}$ be the outcome variable, and $M\in\mathcal{M}$ be a mediator variable that relays part of the causal effect of $A$ on $Y$. In addition, let $X\in\mathcal{X}$ denote the observed pre-treatment covariates in the system. See Figure (ref)$(c)$ for a graphical representation of the causal relationships between the variables. To describe the causal effect of the treatment on the outcome, we use the potential outcome framework rubin1974estimating. Let $Y^{(A=a)}$ be the random variable representing the potential outcome when the treatment is set to value $a$. Suppose we are interested in comparing the treatment values of $a$ and $a'$. A popular way to measure the causal effect of this change in treatment is to use the average causal effect (ACE), which captures the difference in the expected value of the potential outcome variables, that is
\[
ACE=\mathbb{E}[Y^{(a)}-Y^{(a')}],
\]
where $\mathbb{E}[\cdot]$ denotes the population expectation operator.
The total ACE of the treatment $A$ on the outcome $Y$ can be partitioned into the part mediated by variable $M$ and the part directly affecting outcome $Y$ (see Figure (ref)).
To formally define this partitioning, let $Y^{(a,m)}$ denote the potential outcome variable corresponding to the outcome when the treatment is set to value $a$ and the mediator is set to value $m$, and $M^{(a)}$ denote the mediator variable when the treatment is set to value $a$.
robins1992identifiability and pearl2001direct proposed the following partitioning of the ACE into the natural direct and indirect effects.
equation[equation omitted — 364 chars of source]
The natural direct effect (NDE) and natural indirect effect (NIE) can be described as follows. NDE captures the change in the expectation of the outcome if the value of the treatment variable is switched between the two arms of the experiment, while the mediator behaves as if the treatment has not changed. NIE captures the change in the expectation of the outcome if the value of the treatment variable is fixed, while the mediator behaves as if the treatment has been switched between the two arms of the experiment.
In the following subsection, we discuss the estimation of NDE and NIE from observational data.
Estimating Natural Direct and Indirect Effects
To estimate the natural direct and indirect effects, from the partitioning in Display (ref), it suffices to focus on estimating the parameter
\[
\psi_0(a,a')=\mathbb{E}[Y^{(a,M^{(a')})}],
\]
for $a,a'\in\mathcal{A}$. Suppose i.i.d. data from a distribution $f$ on variables $O=\{A,X,M,Y\}$ are given.
In general, the estimand $\psi_0$ is not identified from observational data, and identification assumptions are required to relate the distribution of the observational data to that of counterfactual variables. We require the following assumptions for the identification of $\psi_0$ from the observed distribution on variables, $f(O)$.
assumption[Identification Assumptions]
Let $X_1 \perp X_2 \mid X_3$ indicate that the random variables $X_1$ and $X_2$ are conditionally independent given the random variable $X_3$.
\begin{itemize}
• {\bf Consistency.} For all $a\in\mathcal{A}$ and $m\in\mathcal{M}$,
\begin{align*}
Y^{(a,m)}=Y&\quadif A=a and M=m,\\
M^{(a)}=M&\quadif A=a.
\end{align*}
• {\bf Sequential Exchangeability.} For all $a,a'\in\mathcal{A}$, and $m\in\mathcal{M}$,
\begin{align*}
&(Y^{(a,m)},M^{(a')})\perp A | X,\\
&Y^{(a,m)}\perp M^{(a')} | A=a',X.
\end{align*}
• {\bf Positivity.} For all $a\in\mathcal{A}$, $m\in\mathcal{M}$ and $x\in\mathcal{X}$,
\begin{align*}
&f_{M|A,X}(m|A,X)>0,\\
&f_{A|X}(a|X)>0,
\end{align*}
where $f_{M|A,X}$ and $f_{A|X}$ are the conditional density of $M$ given $A$ and $X$, and the conditional density of $A$ given $X$, respectively.
\end{itemize}
Under Assumption (ref), the estimand $\psi_0(a,a')$ can be identified from the observed distribution $f(O)$ using the following expression called the mediation formula, originally proposed in imai2010identification and pearl2001direct.
equation[equation omitted — 134 chars of source]
where $f_X$ is the marginal distribution of $X$.
Using Equation (ref), one can estimate the parameter of interest $\psi_0(a,a')$ by first estimating the nuisance functions $\mathbb{E}[Y|A,M,X]$ and $f_{M|A,X}$, and then using a plug-in estimator to estimate $\psi_0$ as follows
\[
\frac{1}{n}\sum_{i=1}^n
\int_\mathcal{M}\hat{\mathbb{E}}[Y_i|A=a,M=m,X_i]\hat{f}_{M|A,X}(m|A=a',X_i)dm.
\]
Unfortunately, this estimator is sensitive to bias in the estimation of nuisance functions. That is, misspecifying either of the nuisance functions induces bias in the estimation of the parameter of interest.
As an alternative approach, in the case of binary treatment, that is, $\mathcal{A}=\{0,1\}$, tchetgen2012semiparametric developed a general semiparametric framework for inference.
They derived the efficient influence function for $\psi_0(a,a')$ as
equation[equation omitted — 256 chars of source]
where $\lambda(a,X) := 1/{f(a|X)}$, $\alpha(a,M,X) := f(M| a, X)$, and $\gamma(X,M,a) := \mathbb{E}[Y|A=a,M,X]$ are the nuisance functions, $a,a'\in\{0,1\}$, $I(\cdot)$ denotes the indicator function, and
comment\\for poster
\begin{align*}
\hat{\psi}_0(a,a')
= \frac{1}{n}\sum^n_{i=1}&\bigg\{\frac{I(A_i = a)\hat{f}_{M|A,X}(M_i | A_i = a', X_i)}{\hat{f}_{A|X}(a | X_i)\hat{f}_{M|A,X}(M_i | A_i = a, X_i)}\{Y_i - \hat{\mathbb{E}}[Y | A_i = a,M_i,X_i]\}\\
\quad+& \frac{I(A_i = a')}{\hat{f}_{A|X}(a' | X)}\{\hat{\mathbb{E}}[Y | A_i = a,M_i,X_i] - \hat{\eta}(a, a', X_i)\} + \hat{\eta}(a, a', X_i) \bigg\}\\
&\hat{\eta}(a, a', X) = \int_\mathcal{M} \hat{\mathbb{E}}[Y | A = a, M = m,X]\hat{f}_{M|A,X}(m | A = a', X) dm.
\end{align*}
\[
\eta(a, a', X) = \int_\mathcal{M} \gamma(X, m, a) \alpha(a', m,X) dm.
\]
Note that $IF_{\psi_0}$ is a function of three nuisance functions: $\lambda$, $
\alpha$, and $\gamma$. tchetgen2012semiparametric showed that the estimator based on this influence function has the triple robustness property, that is, it is consistent even if the model for one (but not more than one) nuisance function is misspecified. Formally, let
itemize• $\mathfrak{M}_{ym}$ be the sub-model in which the model for $\gamma$ and $\alpha$ are correctly specified.
• $\mathfrak{M}_{ya}$ be the sub-model in which the model for $\gamma$ and $\lambda$ are correctly specified.
• $\mathfrak{M}_{ma}$ be the sub-model in which the model for $\alpha$ and $\lambda$ are correctly specified.
The estimator for $\psi_0$ based on the influence function $IF_{\psi_0}$ defined as
align*[align* omitted — 318 chars of source]
is consistent when the truth lies in the submodel union $\mathfrak{M}_{ym}\cup\mathfrak{M}_{ya}\cup\mathfrak{M}_{ma}$ and the estimators of nuisance functions in the correctly specified submodel(s) are consistent, where $$\hat{\eta}(a, a', X) = \int_\mathcal{M} \hat{\gamma}(X,m,a) \hat{\alpha}(a',m, X)dm.$$
Inspired by this result, in the following section, we propose a kernel-based estimator for meditation effects in settings with continuous treatment variables, while preserving multiple robustness and allowing for the nuisance parameters to be estimated at a slower rate than the parameter of interest.
Continuous Mediation Analysis
In the case of continuous treatments, the parameter of interest, $\psi_0(a,a')$, is no longer regular bickel1993efficient,tsiatis2007semiparametric. Therefore, the method of tchetgen2012semiparametric cannot be applied directly.
However, their estimator can be modified to be suitable for inference in the case of continuous treatments, while still obtaining desirable properties such as asymptotic normality, robustness to misspecification of nuisance functions, and valid inference while allowing for
the nuisance parameters to be estimated at a slower rate than the parameter of interest.
Specifically, we modify $\hat{\psi}^{TTS}(a,a')$ by employing a kernel smoothing technique, wherein the indicators in the calculation of $\hat{\psi}^{TTS}(a, a')$
are replaced by kernel-based weights. The weights are functions of treatment values falling within a neighborhood (defined by the bandwidth parameter $h$) of $a$ and $a'$. This modification introduces several challenges in the inference, which we will present and address in Section (ref).
Let $d_A$ denote the dimension of the treatment variable, and let
\[
K_h(a):=\frac{1}{h^{d_A}}\prod_{j=1}^{d_A}k\Big(\frac{a_j}{h}\Big),
\]
where $k(\cdot)$ is a kernel function, and $h$ denotes the bandwidth parameter.
We propose to use the following modification of the efficient influence function in Equation (ref) for any $a$ and $a' \in \mathcal{A}$:
equation[equation omitted — 284 chars of source]
To derive desired results on consistency, asymptotic normality, and multiple robustness, we require the kernel $k(\cdot)$ to satisfy the following conditions.
assumption[Kernel & Bandwidth Assumptions]
The kernel function $k(\cdot)$ satisfies
\begin{enumerate}
• $\int k(u) du = 1$
• $\int uk(u) du = 0$
• $0 < \int u^6 k(u) du < \infty$
• $\int k^2(u) du < \infty$
• $0 < \int u^2 k^2(u) du < \infty$
\end{enumerate}
Additionally, the kernel bandwidth $h$ is assumed to be a function of the sample size $n$ and satisfies $h \rightarrow 0$, $nh^{d_A} \rightarrow \infty$ and $nh^{d_A + 4} \rightarrow C_h$, for a constant $C_h$, as $n\rightarrow \infty$.
These assumptions are satisfied by common kernels such as the Gaussian kernel and Epanechnikov kernel. Note that in the moment function in Equation (ref), the nuisances are not functions of the parameter of interest $\psi_0$. Therefore, having estimators for nuisance functions suffices for obtaining an estimator for $\psi_0$. Next, we describe the estimation procedure for utilizing Equation (ref) to estimate $\psi_0(a, a')$.
Estimation Procedure. We use the cross-fitting estimation approach of chernozhukov2018double for separating the estimation of the nuisance functions from the parameter of interest.
This approach is beneficial since weaker smoothness requirements are needed for the estimation of nuisance functions. In the cross-fitting approach, we partition the sample indices into $L$ folds $\{I_1,...,I_L\}$ of roughly equal size. Data from the $\ell$-th fold is denoted by $O_{I_\ell}$, and the data in the rest of the folds is denoted by $O^c_{I_\ell}$.
For $\ell\in\{1,...,L\}$, we estimate the nuisance functions $\alpha$, $
\lambda$, $\gamma$ by $\hat{\alpha}_\ell,\hat{\lambda}_\ell,\hat{\gamma}_\ell$ based on data $O^c_{I_{\ell}}$.
For all $\ell$, let $\hat{\psi}_\ell$ be the estimator for $\psi_0$ obtained by solving
\[
\frac{1}{|I_\ell|}\sum_{i\in I_\ell}m(O_i;\hat{\alpha}_\ell,\hat{\lambda}_\ell,\hat{\gamma}_\ell,\hat{\psi}_\ell(a,a'))=0.
\]
Our proposed estimator for $\psi_0$ is
equation[equation omitted — 111 chars of source]
where MR stands for multiply robust. In the next section, we present the asymptotic properties of our proposed estimator, along with the required regularity conditions.
Asymptotic Analysis
In this section, we provide asymptotic properties of our proposed estimator $\hat{\psi}^{MR}(a,a')$ in Equation (ref). We start by stating the required regularity conditions.
assumption[Regularity Conditions]
\\
\begin{enumerate}
\begin{sloppypar}
• For all $Y$, $M$, and $X$, the functions $f(a \mid Y, M, X)$, $f(a \mid M, X)$, $f(a \mid X)$, $\gamma(X,M,a)$ as a function of $a$ are three times continuously differentiable, and the functions and their first, second, and third derivatives with respect to $a$ are bounded.
\end{sloppypar}
• The nuisance functions $\alpha,\lambda, \gamma$ and the estimators $\hat{\alpha},\hat{\lambda},\hat{\gamma}$ are bounded. Additionally, $\alpha$, $\lambda$ and their estimators $\hat{\alpha}$, $\hat{\lambda}$ are bounded away from zero.
• $Y$'s conditional variance $var(Y|a,m,x)$ and its first and second derivative with respect to $a$ are bounded for any $a\in \mathcal{A}$, $m\in \mathcal{M}$, and $x\in \mathcal{X}$.
\end{enumerate}
In addition to the regularity conditions, we require the following conditions regarding the convergence of the estimators of the nuisance functions.
assumption[Convergence of Nuisance Estimators]
\\
For any value $a \in \mathcal{A}$, the estimators $\hat{\alpha}(a,M, X)$, $\hat{\lambda}(a,X)$, and $\hat{\gamma}(X,M,a)$ satisfy the following conditions
\begin{enumerate}
• $\int\left(\hat{\lambda}(a, x) - \lambda(a, x) \right)^2 f_{X}(x) dx \xrightarrow[]{P} 0$,
• $\int \left(\hat{\alpha}(a, m, x) - \alpha(a, m, x) \right)^2 f_{M, X}(m, x) dm dx \xrightarrow[]{P} 0$, and
• $\int\left(\hat{\gamma}(x,m,a) - \gamma(x,m,a) \right)^2 f_{M, X}(m, x) dm dx \xrightarrow[]{P} 0$,
\end{enumerate}
where $\xrightarrow[]{P}$ indicates convergence in probability.
Similar to influence function-based estimators, in Assumption (ref), we do not require individual nuisance estimators to satisfy convergence rate conditions. However, in our proposed method, we have requirements on the convergence rate of the product for the nuisance estimators as follows.
assumption[Nuisance Convergence Rates]
\\
For any value $a, a' \in \mathcal{A}$, the estimators $\hat{\alpha}(a,M, X)$, $\hat{\lambda}(a,X)$, and $\hat{\gamma}(X,M,a)$ satisfy the following conditions
{
\begin{enumerate}
• \begin{align*}
&\sqrt{nh^{d_A}}\left(\int \left(\hat{\alpha}(a', m, x) - \alpha(a', m, x) \right)^2f_{M, X}(m,x)dm dx\right)^{\frac{1}{2}}\\
&\times \left(\int \left(\hat{\gamma}(x,m,a) - \gamma(x,m,a) \right)^2f_{M, X}(m,x) dm dx\right)^{\frac{1}{2}} \xrightarrow[]{P}0,
\end{align*}
• \[
\sqrt{nh^{d_A}}\left(\int \left(\hat{\lambda}(a', x) - \lambda(a', x) \right)^2f_{X}( X)dx\right)^{\frac{1}{2}}
\left(\int \left(\hat{\gamma}(x,m,a) - \gamma(x,m,a) \right)^2f_{M, X}(m,x) dm dx\right)^{\frac{1}{2}} \xrightarrow[]{P} 0,
\]
• \[
\sqrt{nh^{d_A}}\left(\int \left(\hat{\lambda}(a',x) - \lambda(a', x) \right)^2f_{X}(X) dx\right)^{\frac{1}{2}}\left(\int \left(\hat{\alpha}(a, m, x) - \alpha(a, m, x) \right)^2f_{M, X}(m,x)dm dx\right)^{\frac{1}{2}}
\xrightarrow[]{P} 0.
\]
\end{enumerate}
}
comment{\tiny
\begin{enumerate}[label=\roman*]
• $\sqrt{nh^{d_A}}\left(\int_{\mathcal{X}}\int_{\mathcal{M}}\left(\left[\frac{\hat{\alpha}(a^\prime, \cdot, \cdot)}{\hat{\alpha}(a, \cdot, \cdot)} - \frac{\alpha(a^\prime, \cdot, \cdot)}{\alpha(a, \cdot, \cdot)} \right]\right)^2f_{M, A, X}(M, a, X)dm dx\right)^{\frac{1}{2}}
\left(\int_{\mathcal{X}}\int_{\mathcal{M}} \left(\hat{\gamma}(a, \cdot, \cdot) - \gamma(a, \cdot, \cdot) \right)^2f_{M, A, X}(M, a, X) dm dx\right)^{\frac{1}{2}} \xrightarrow[]{P}0$
• $\sqrt{nh^{d_A}}\left(\int_{\mathcal{X}}\int_{\mathcal{M}}\left(\left[\hat{\lambda}(a, \cdot, \cdot) - \lambda(a, \cdot, \cdot) \right]\right)^2f_{A, X}(a, X)dm dx\right)^{\frac{1}{2}}
\left(\int_{\mathcal{X}}\int_{\mathcal{M}} \left(\hat{\gamma}(a, \cdot, \cdot) - \gamma(a, \cdot, \cdot) \right)^2f_{M, A, X}(M, a_, X) dm dx\right)^{\frac{1}{2}} \xrightarrow[]{P} 0$
• $\sqrt{nh^{d_A}}\left(\int_{\mathcal{X}}\int_{\mathcal{M}}\left(\left[\hat{\lambda}(a, \cdot) - \lambda(a, \cdot) \right]\right)^2f_{A, X}(a, X)dm dx\right)^{\frac{1}{2}}\left(\int_{\mathcal{X}}\int_{\mathcal{M}}\left(\left[\hat{\eta}(a, a^\prime, \cdot) - \eta(a, a^\prime, \cdot) \right]\right)^2f_{A, X}(a, X)dm dx\right)^{\frac{1}{2}}
\xrightarrow[]{P} 0$
\end{enumerate}
}
As seen in Assumption (ref),
our requirements on the convergence rate of nuisance function estimators are on the product of the error rates, rather than on the individual nuisance function estimators. Therefore, if one of the estimators converges at a slow rate, the other estimator can compensate. This is a desirable property when working with non-parametric estimators since they typically have slow rates of convergence.
In the case of binary treatment variables, the combination of Assumptions (ref) and (ref) can lead to asymptotic normality, which is used to construct Wald-style confidence intervals. However, when the treatments are continuous, the Central Limit Theorem (CLT) cannot be directly applied to our proposed method because the bandwidth $h$ varies as a function of the sample size $n$, implying that the distribution of equation (ref) changes with $n$.
Instead, we impose additional assumptions stated below to satisfy the Lyapunov's condition for CLT and achieve asymptotic normality.
assumption[Assumptions for Lyapunov CLT]
\begin{enumerate}
• $\mathbb{E}\left[|Y-\gamma(X,M,a)|^3 \big| A = a', M = m, X = x\right]$ is bounded over any $(a,a',m,x)\in \mathcal{A}\times \mathcal{A}\times\mathcal{M} \times \mathcal{X}$.
• $\int^{\infty}_{-\infty} k(u)^{c_1}k(u+\tilde{c})^{c_2} du < \infty$ and $\int^{\infty}_{-\infty} u^2k(u)^{c_1}k(u+\tilde{c})^{c_2} du < \infty$ for $\tilde{c} \in \mathbb{R}$ and $c_1+c_2 \in \{2,3\}$, where $c_1,c_2\in\{0,1,2,3\}$.
\end{enumerate}
Having stated the assumptions in our setting, we now provide the following result regarding the asymptotic behavior of the proposed estimator in Equation (ref).
theoremSuppose Assumptions (ref)-(ref) hold. Then for any values of $a,a'\in\mathcal{A}$,
\[
\sqrt{nh^{d_A}}(\hat{\psi}^{MR}(a,a')-\psi_0(a,a'))=
\sqrt{\frac{h^{d_A}}{n}}
\sum_{i=1}^nm(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))+o_p(1).
\]
Additionally, if Assumption (ref) holds, then $\sqrt{nh^{d_A}}(\hat{\psi}^{MR}(a,a')-\psi_0(a,a')-h^2 B(a,a'))$ converges to the Gaussian distribution $\mathcal{N}(0,V(a,a'))$, where $B(a,a')$ and $V(a, a^\prime)$ are defined as
{
\begin{align*}
B(a,a') =& \left[\int u^2 k(u)du\right]\\
&\times \mathbb{E}\bigg[ \frac{\alpha(a',M,X)}{\alpha(a,M,X)\lambda(a,X)}\bigg(\sum^{d_A}_{j=1} \partial_{a_j} \gamma(X,M,a) \partial_{a_j} f(a|X,M) + \frac{1}{2} \left[\sum^{d_A}_{j=1} \partial^2_{a_j} \gamma(X,M,a)\right]f(a|X,M) \bigg)\\
&+ \{\gamma(X,M,a) - \eta(a, a', X)\}\frac{1}{2}\frac{\sum_{j = 1}^{h_{d_A}} \partial^2_{a_j} f(a'|X,M)}{\lambda(a',X)}\bigg]+O(h),
\end{align*}
}
and
{
\begin{align*}
V(a, a^\prime) = &\left[\int k(u)^2 du\right]^{d_A}\\
&\times\mathbb{E}\Bigg\{\frac{\alpha^2(a', M, X) f(a|X,M)}{\alpha^2(a, M, X)} \lambda^2(a|X)var(Y|X,M,a)+ \lambda(a'|X)var[E(Y|X,M,a)|X,a']\Bigg\}.
\end{align*}
}
All proofs are provided in the Supplementary Material.
Theorem (ref) provides results on the point-wise convergence of $\hat{\psi}^{MR}(a,a')$ and establishes the asymptotic normality of our estimator. Additionally, $\hat{\psi}^{MR}(a,a')$ has a multiple robustness property analogous to the $\hat{\psi}^{TTS}(a,a')$, formally stated in Lemma (ref).
lemmaUnder Assumptions (ref), (ref), and (ref), the proposed $\hat{\psi}^{MR}(a, a')$ will be a consistent estimator for $\psi_0(a, a')$ as long as any two out of the three conditions in Assumption (ref) hold.
While Theorem (ref) and Lemma (ref) establish properties of $\hat{\psi}^{MR}(a,a')$ that are desirable for point estimation, uncertainty quantification through the calculation of valid confidence intervals requires the estimation of $V(a, a')$ and $B(a, a')$. However, these are hard to estimate well because of their complicated analytical form. Nevertheless, by choosing an undersmoothing bandwidth $h$ that satisfies $\sqrt{nh^{d_A + 4} } \rightarrow 0$, valid confidence intervals can still be constructed without estimating $B(a, a')$. In this case, we only need to focus on estimating $V(a, a')$. Naturally, an estimator for $V(a, a')$ can be constructed as follows.
\[\widehat{V}(a, a') = h^{d_A}\frac{1}{L}\sum_{\ell = 1}^L \frac{1}{|I_\ell|}\sum_{i \in I_\ell} m^2(O_i;\hat{\alpha},\hat{\lambda},\hat{\gamma},\hat{\psi}_\ell(a,a')).
\]
We present the additional assumptions necessary for the consistency of $\widehat{V}(a, a')$ below.
assumption[Consistency of $\widehat{V}(a, a')$]
\begin{enumerate}
• $\mathbb{E}\{[Y-\gamma(X,M,a)]^4 | A = a', M = m, X = x\}$ is bounded over any $(a,a',m,x)\in \mathcal{A}\times \mathcal{A}\times\mathcal{M} \times \mathcal{X}$.
• $\int^{\infty}_{-\infty} k(u)^{c_1}k(u+\tilde{c})^{c_2} du < \infty$ and $\int^{\infty}_{-\infty} u^2k(u)^{c_1}k(u+\tilde{c})^{c_2} du < \infty$ for $\tilde{c} \in \mathbb{R}$ and $c_1+c_2 \in \{2,3,4\}$ for $c_1,c_2\in\{0,1,2,3,4\}$.
• \begin{comment}
\[
\int (\gamma(X,M,a) - \hat{\gamma}(X,M,a))^2\Big(\hat{\lambda}(a,X) - \lambda(a,X)\Big)^2 f(X, M) dX dM \xrightarrow[]{P} 0,
\]
\[
\int \Big(\frac{\hat{\alpha}(a',M,X)}{\hat{\alpha}(a, M, X)} - \frac{\alpha(a',M,X)}{\alpha(a,M,X)}\Big)^2(\gamma(X,M,a) - \hat{\gamma}(X,M,a))^2f(X, M) dM dX \xrightarrow[]{P} 0.
\]
\end{comment}
\[
\int (\gamma(x,m,a) - \hat{\gamma}(x,m,a))^2\Big(\hat{\lambda}(a,x) - \lambda(a,x)\Big)^2 f(x,m) dx dm \xrightarrow[]{P} 0,
\]
\[
\int \Big(\frac{\hat{\alpha}(a',m,x)}{\hat{\alpha}(a,m,x)} - \frac{\alpha(a',m,x)}{\alpha(a,m,x)}\Big)^2(\gamma(x,m,a) - \hat{\gamma}(x,m,a))^2f(x,m) dm dx \xrightarrow[]{P} 0.
\]
\end{enumerate}
Next, we prove that $\widehat{V}(a, a')$ is a consistent estimator of $V(a, a')$ in Lemma (ref).
lemmaSuppose Assumptions (ref)-(ref) hold. Then for any values of $a,a'\in\mathcal{A}$, $\widehat{V}(a, a')$ is a consistent estimator for $V(a, a')$.
Using the result of Lemma (ref), $\widehat{V}(a, a')$ can be used to construct asymptotically valid confidence intervals as follows. Choose an undersmoothing bandwidth $h$ that satisfies $\sqrt{nh^{d_A + 4}} \rightarrow 0$, so that $\sqrt{nh^{d_A}}h^2 B(a, a')$ is asymptotically negligible. Then, the $(1 - \alpha)$ confidence interval is given as
equation[equation omitted — 136 chars of source]
where $\Phi$ is the CDF of $\mathcal{N}(0,1)$. However, there is little guidance about how to undersmooth in practice and it is mostly a technical device kennedy2017nonparametric. In simulation and application, we express uncertainty about the estimator $\hat{\psi}(a,a')$ using a bandwidth chosen under the Silverman rule of thumb van2000asymptotic, silverman2018density using pointwise confidence intervals (ref). As stated by wasserman2006all, such an approach does not necessitate the choice of an undersmoothing bandwidth, but the constructed confidence intervals center around the parameter of interest $\psi_0(a,a')$ with an additional asymptotic bias.
Simulation Study
We now present a simulation study to demonstrate that the proposed estimator is consistent and multiply robust, and a sensitivity analysis to assess the uncertainty of the proposed estimator under different bandwidths and sample sizes. The data-generating process is as follows:
align*[align* omitted — 292 chars of source]
The parameter of interest is $\psi_0(a, a')$ at $a = 4.5$ and $a' = 6$. Under the described simulation setting, the true parameter value is $9.1$, calculated based on Monte Carlo approximation of $\psi_0(a, a') = \int_\mathcal{X}\eta(a,a',X=x)dx$. To demonstrate the multiple robustness property, we consider various types of model misspecification in Table (ref), where we also compare the proposed estimator $\hat{\psi}^{MR}(a, a')$ to the estimator in huber2020direct, $\hat{\psi}^H(a, a')$, and the estimator without bias correction, $\hat{\psi}^{\eta}(a,a')$. To ensure comparability, we calculate the three estimators under the cross-fitting approach and define $\hat{\psi}^H(a, a')$ and $\hat{\psi}^{\eta}(a, a'))$ as
$$\hat{\psi}^H(a, a') = \frac{1}{L}\sum^L_{\ell = 1}\hat{\psi}^H_\ell(a, a')\quad\text{and}\quad \hat{\psi}^\eta(a, a') = \frac{1}{L}\sum^L_{\ell = 1}\hat{\psi}^\eta_\ell(a, a'),$$
where
align*[align* omitted — 512 chars of source]
$\hat{\eta}(a,a',X) = \int_\mathcal{M} \hat{\gamma}(X, m, a) \hat{\alpha}(a',m,X)dm$, and we set $L=3$. We use 1000 simulation replicates for each of the sample sizes 2000, 5000, and 8000, and chose the kernel bandwidth using the Silverman rule of thumb silverman2018density under Gaussian kernels.
The types of model misspecification considered include the scenario where all three models, $\mathbb{E}[Y|A,M,X]$, $f(M|A,X)$, and $f(A | X)$, are correctly specified, scenarios where only two out of the three models are correctly specified, and the scenario where all three models are misspecified. As shown in Table (ref), our proposed estimator has minimal or close to minimal bias for all scenarios except when all models are misspecified, demonstrating its theoretically proven multiple robustness property. Additionally, bias and the root mean square error (RMSE) across simulation replicates reduce as the sample size gets larger, showing the consistency of our estimator. The bias becomes significant when all models are misspecified for all sample sizes and considered estimators.
table[table omitted — 1,429 chars of source]
table[table omitted — 1,586 chars of source]
Table (ref) summarizes the sensitivity of the estimated $\hat{\psi}^{MR}(a,a')$ under correct model specifications to different sample sizes and kernel smoothing bandwidths, by reporting absolute average bias, average of $\hat{V}(a,a')^{1/2}$, and Monte Carlo coverage. We define coverage as the proportion of simulation replicates that include the true value $\psi_0(a, a')$ in the estimated 95$\%$ confidence interval. Recall that $\hat{V}(a, a')$ is the empirical variance of the estimating functions; we construct pointwise 95% confidence intervals at significance level $\alpha = 0.05$ via a Wald approach as in equation (ref). The Silverman kernel bandwidths for n = 2000, 5000, 8000, and under $L=3$ are approximately 0.28, 0.23, and 0.20. For the estimated mediation function $\hat{\psi}^{MR}(a,a')$, the presence of bias becomes apparent with reduced length of confidence intervals and decreasing coverage when the kernel bandwidths are larger than the Silverman-suggested optimal bandwidths, i.e., when the bandwidths are greater or equal to $0.3$. This pattern persists across different sample sizes. Coverage is theoretically guaranteed when the sample size $n$ goes to infinity and $\sqrt{nh^{d_A + 4}} \rightarrow 0$. For example, when $d_A = 1$, an undersmoothing bandwidth would satisfy the requirement. We can see from Table (ref) that when bandwidth equals 0.1 (undersmoothed), coverage of the proposed pointwise 95% confidence interval is indeed at least 0.95 for all sample sizes. In practice, choosing an undersmoothed kernel bandwidth can guarantee a relatively smaller bias in finite sample settings (for the price of conservative coverage). However, as previously discussed, there is no guidance on the choice of an undersmoothing bandwidth in practice.
Application
table[table omitted — 4,942 chars of source]
figure[figure omitted — 318 chars of source]
figure[figure omitted — 811 chars of source]
The proposed approach is applied to the Job Corps study huber2020direct,schochet2008does,schochet2001national. Study participants were enrolled between 16 and 24 years old and from low-income households. The program provides eight months or approximately 1,200 hours of training on average. We aim to study the effect of the duration of Job Corp training $(A)$ on the binary outcome of the occurrence of any criminal arrests in the fourth year following program participation $(Y)$, with the proportion of weeks employed in the second year being the mediator $(M)$. Our study design follows huber2020direct, who considered a similar causal mechanism and focused on the actual number of arrests in the fourth year as the outcome.
We consider a rich set of time-invariant socioeconomic variables as pretreatment confounders $X$, similar to the study in huber2020direct. Table (ref) presents summary statistics of the following variables: outcome $Y$, mediator $M$, treatment $A$, and confounders $X$. Missing values in confounders are addressed by including the indicators of missingness as covariates. Moreover, following previous work on this dataset huber2020direct,flores2012estimating, we apply our evaluation to the
4,000 individuals in the dataset who received training, i.e., with a training duration in the program strictly greater than zero. Table (ref) shows that on average, 5.1$\%$ of the participants had a history of imprisonment, and 23.75$\%$ had been arrested at least once before joining the study. Additionally, 8.7$\%$ of the individuals included in the study were arrested for criminal activities during the fourth year after study participation.
We investigate whether longer Job Corps training reduces criminal behavior by evaluating the natural direct effect for treatment durations at $a\in\{100,200,\ldots,2000\}$ hours versus just $a'=40$ hours. Following huber2020direct, we assume treatment to follow a log-normal distribution and parametric linear models for the outcome, mediator, and log-treatment. Let $\hat{f}(a|X)$ be a model-based propensity score estimator at treatment $a$.
To improve the stability of our estimator, we use a H\'ajek-type stabilized weighted propensity score hernan2020causal, which is defined as $\hat{f}(a|X_i) \times \frac{1}{|I_{-\ell}|}\sum^{|I_{-\ell}|}_{j\in I_{-\ell}} \frac{K_h (A_j-a)}{\hat{f}(a|X_j)}$.
This prevents a large influence from observations that have extremely small propensities $\hat{f}(a|X)$.
In Supplementary Material Section 5, we show that the proposed Hajek-type stabilized weighted propensity score is a consistent estimator of the conditional treatment probability when the model-based propensity estimator is consistent.
Figure (ref) displays the mean and 95$\%$ confidence interval of natural direct effect
over the range of values for $a$, under H\'ajek-type stabilized weighted propensities clipped at 0.01 ionides2008truncated and Gaussian kernels with bandwidth chosen using the Silverman-type rule of thumb silverman2018density. The confidence interval is obtained via Equation (ref). Figure (ref) reports the sensitivity analysis and demonstrates how the estimated mean and empirical standard deviation
vary by different clipping thresholds and smoothing bandwidths.
Both schochet2008does and huber2020direct identified significant effects of the training program in reducing criminal arrests, especially when the training duration is over 1000 hours. Similar to huber2020direct, we observed a nonlinear effect of the training duration on the occurrence of arrests. However, our sensitivity analysis showed that a significant reduction in the likelihood of arrests (at 5$\%$ level) does not always occur after a certain number of hours in Job Corps. As shown in the sensitivity analysis, the significance and size of the natural direct effect also depend on two factors: the choice of bandwidth for smoothing the continuous treatment and the threshold for clipping the estimated propensity score. Regardless of the actual treatment effect level, we generally noticed a decline in the probability of experiencing at least one arrest in year 4 for those with training durations between approximately 500 and 1,300 hours. From summary descriptions in Table (ref), we know that the interquartile range (IQR) of the training durations in the observed data falls within the range of approximately 405 to 1,767 hours. Hence, it can be inferred that extended training has a tendency to diminish criminal behaviors among the majority of the study participants.
Conclusion
In this paper, we present an estimator for the direct and indirect components of the causal effect in the presence of continuous treatments, while allowing for complex data generating processes. Our estimator is inspired by influence function based estimation strategies from the literature on semiparametric statistics bickel1993efficient, tsiatis2007semiparametric. We provide results on the convergence rate and asymptotic normality of our proposed estimator, which allows for the construction of confidence intervals and hypothesis tests. Extending the point-wise results to uniform results and providing optimal rates is left for future work.
in the presence of high-dimensional nuisance parameters, the estimator may inherit the slow rates of the nuisance estimators, adversely affecting the inference of the target parameter.
center[center omitted — 48 chars of source]
description• Technical derivations and proofs that are used to support the results. (.pdf file)
center[center omitted — 129 chars of source]
Before we start with the proofs, we establish some lemmas that will help us with the proofs in the rest of the Supplementary Material.
lemmaLet $\{X_m\}$ and $\{Y_m\}$ be a sequence of random variables. Then under conditions outlined in Lemma 6.1 in chernozhukov2018double, $\mathbb{E}[|X_m| \mid Y_m] = o_p(1)$ implies $X_m = o_p(1)$.
proofBy the Conditional Markov Inequality, for any $\epsilon > 0$,
\begin{align*}
p(|X_m| \geq \epsilon \mid Y_m) \leq \frac{\mathbb{E}[|X_m| \mid Y_m]}{\epsilon}
\end{align*}
By $\mathbb{E}[|X_m| \mid Y_m] = o_p(1)$, there is $p(|X_m| \geq \epsilon \mid Y_m) = o_p(1)$. An application of Lemma 6.1 then yields $p(|X_m| > \epsilon) \rightarrow 0$, therefore $X_m = o_p(1)$.
lemmaUnder Assumption 2, for a twice continuously differentiable function $f$ with bounded first and second derivative, we have
\[
\int_A K_h(A-a)f(A)dA=f(a)+O(h^2).
\]
proof{
\begin{align*}
\int_A K_h(A-a)f(A)dA
&=\int \left[\prod_{j = 1}^{d_A}k(u_j)\right]f(uh+a)du_1\dots du_{d_A}\\
&=\int \left[\prod_{j = 1}^{d_A}k(u_j)\right]\Bigg\{f(a)+\sum_{j = 1}^{d_A} u_jh \partial_{a_j}f(a)+
\frac{1}{2}\sum_{j = 1}^{d_A} \sum_{j^\prime = 1}^{d_A} u_ju_{j^\prime}h^2\partial_{a_j}\partial_{a_{j^\prime}}\partial_{a_{j^{\prime}}}f(a)|_{\Bar{a}}\Bigg\}du_1 \dots du_{d_A}\\
&=f(a)+O(h^2),
\end{align*}
}
Where $\Bar{a}$ is in between $A$ and $a$.
Remark: We assume the second derivative is bounded over the support of the function $f(a)$, which is a stronger assumption than $O(1)$ since the bound holds everywhere as opposed to only for $a \geq c$ where $c$ is a constant. If $\nu(x)$ and $\omega(x)$ are two arbitrary functions, then $\int \nu(x) \omega(x) dx = O(1) \int |\omega(x)|dx$ is true when $\nu(x)$ is bounded, but not when $\nu(x) = O(1)$, e.g. when $\nu(x) = 1/x$ and $\omega(x) = \mathbb{I}\{0\le x \le 1\}$ .
\section{Proof for Theorem 1}
We follow a similar outline as colangelo2020double and chernozhukov2018double. The proof for this theorem is split into two parts. The first part establishes that the proposed estimator satisfies
\begin{align*}
\sqrt{\frac{h^{d_A}}{n}}\sum^L_{\ell=1} \sum_{i \epsilon I_\ell} \left\{m(O_i;\hat{\alpha}_\ell,\hat{\lambda}_\ell,\hat{\gamma}_\ell,\psi_0(a,a')) - m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a')) \right\} = o_p(1),
\end{align*} and the second part establishes that $\sqrt{nh^{d_A}}(\hat{\psi}^{MR}(a,a')-\psi_0(a,a')-B(a,a'))$ converges to the Gaussian distribution $\mathcal{N}(0,V(a,a'))$.
Starting with the first part of the proof, note that
\begin{align*}
&\sqrt{nh^{d_A}} \frac{1}{n}\sum^L_{\ell=1}\sum_{i\in I_\ell} \Big\{m(O_i; \hat{\alpha}_\ell,\hat{\lambda}_\ell, \hat{\gamma}_\ell, \hat{\psi}_\ell(a,a'))- m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\Big\}\\
=& \sqrt{\frac{h^{d_A}}{n}}\sum^L_{\ell=1}\sum_{i\in I_\ell} \Big\{m(O_i; \hat{\alpha}_\ell,\hat{\lambda}_\ell, \hat{\gamma}_\ell, \hat{\psi}_\ell(a,a'))-m(O_i;\hat{\alpha}_\ell,\hat{\lambda}_\ell,\hat{\gamma}_\ell,\psi_0(a,a'))\\
&+m(O_i;\hat{\alpha}_\ell,\hat{\lambda}_\ell,\hat{\gamma}_\ell,\psi_0(a,a'))- m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\Big\}\\
=& -\sqrt{nh^{d_A}} (\hat{\psi}^{MR}(a,a')- \psi_0(a,a') )\\ &+\sqrt{\frac{h^{d_A}}{n}}\sum^L_{\ell=1}\sum_{i\in I_\ell}\Big\{m(O_i;\hat{\alpha}_\ell,\hat{\lambda}_\ell,\hat{\gamma}_\ell,\psi_0(a,a'))- m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\Big\}.
\end{align*}
Since $\frac{1}{n}\sum^L_{\ell=1}\sum_{i\in I_\ell} m(O_i; \hat{\alpha}_\ell,\hat{\lambda}_\ell, \hat{\gamma}_\ell, \hat{\psi}_\ell(a,a'))=0$, we have
\begin{align*}
&\sqrt{nh^{d_A}} (\hat{\psi}^{MR}(a,a')- \psi_0(a,a') )\\
=& \sqrt{\frac{h^{d_A}}{n}}\sum^L_{\ell=1}\sum_{i\in I_\ell} \Big\{ m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\Big\}\\
&+\sqrt{\frac{h^{d_A}}{n}}\sum^L_{\ell=1}\sum_{i\in I_\ell}\Big\{m(O_i;\hat{\alpha}_\ell,\hat{\lambda}_\ell,\hat{\gamma}_\ell,\psi_0(a,a'))- m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\Big\}.
\end{align*}
In order to establish an asymptotically linear representation for our proposed estimator, it suffices to to show that for all $1\le \ell\le L$ we have
\begin{align*}
\sqrt{\frac{h^{d_A}}{n}} \sum_{i \epsilon I_\ell} \left\{m(O_i;\hat{\alpha}_\ell,\hat{\lambda}_\ell,\hat{\gamma}_\ell,\psi_0(a,a')) - m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a')) \right\} = o_p(1).
\end{align*}
Next, we expand $m(O_i;\hat{\alpha}_\ell,\hat{\lambda}_\ell,\hat{\gamma}_\ell,\psi_0(a,a')) - m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))$ into multiple terms and bound each term individually. Note that
\begin{align}
&m(O_i;\hat{\alpha}_\ell,\hat{\lambda}_\ell,\hat{\gamma}_\ell,\psi_0(a,a')) - m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\nonumber\\
&=K_h(A_i - a)\big\{\hat{\lambda}(a,X_i)\frac{\hat{\alpha}(a',M_i,X_i)}{\hat{\alpha}(a,M_i,X_i)}\{Y_i - \hat{\gamma}(a,M_i,X_i)\}-\lambda(a,X_i)\frac{\alpha(a',M_i,X_i)}{\alpha(a,M_i,X_i)}\{Y_i - \gamma(X_i,M_i,a)\}\big\}\\
&\quad+K_h(A_i - a')\big\{\hat{\lambda}(a',X_i)\{\hat{\gamma}(a,M_i,X_i) - \hat{\eta}(a, a', X_i)\}
-\lambda(a',X_i)\{\gamma(X_i,M_i,a) - \eta(a, a', X_i)\}\big\}\\
&\quad+\hat{\eta}(a, a', X_i)-\eta(a, a', X_i)\tag{R1}.
\end{align}
Defining $R(M_i,X_i) := \frac{\alpha(a^\prime, M_i, X_i)}{\alpha(a, M_i, X_i)}$, terms (ref) and (ref) can be expanded additionally. Expanding term (ref), we get
\begin{align*}
&K_h(A_i - a)\big\{\hat{\lambda}(a, X_i)\hat{R}(M_i,X_i)\{Y_i - \hat{\gamma}(X_i,M_i,a)\}-\lambda(a,X_i)R(M_i,X_i)\{Y_i - \gamma(X_i,M_i,a)\}\big\}\\
&=-K_h(A_i - a)\big(\hat{R}(M_i,X_i)-R(M_i,X_i)\big)\big(\hat{\lambda}(a, X_i)-\lambda(a,X_i)\big)\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\tag{CS1}\\
&\quad+K_h(A_i - a)\big(\hat{R}(M_i,X_i)-R(M_i,X_i)\big)\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\big(Y_i - \gamma(X_i,M_i,a)\big)\tag{CS2}\\
&\quad-K_h(A_i - a)\big(\hat{R}(M_i,X_i)-R(M_i,X_i)\big)\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a,X_i)\tag{CS3}\\
&\quad-K_h(A_i - a)\big(\hat{\lambda}(a, X_i)-\lambda(a,X_i)\big)\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)R(M_i,X_i)\tag{CS4}\\
&\quad+\Big\{K_h(A_i - a)\big(\hat{R}(M_i,X_i)-R(M_i,X_i)\big)\lambda(a,X_i)\big(Y_i - \gamma(X_i,M_i,a)\big)\\
&\quad\quad\quad-\mathbb{E}\big[
K_h(A_i - a)\big(\hat{R}(M_i,X_i)-R(M_i,X_i)\big)\lambda(a,X_i)\big(Y_i - \gamma(X_i,M_i,a)\big) \mid O^c_{I_\ell} \big]\Big\}\tag{E1}\\
&\quad+\mathbb{E}\big[
K_h(A_i - a)\big(\hat{R}(M_i,X_i)-R(M_i,X_i)\big)\lambda(a,X_i)\big(Y_i - \gamma(X_i,M_i,a)\big) \mid O^c_{I_\ell}\big]\tag{TR1}\\
&\quad+\Big\{K_h(A_i - a)\big(\hat{\lambda}(a, X_i)-\lambda(a,X_i)\big)R(M_i,X_i)\big(Y_i - \gamma(X_i,M_i,a)\big)\\
&\quad\quad\quad-\mathbb{E}\big[
K_h(A_i - a)\big(\hat{\lambda}(a, X_i)-\lambda(a,X_i)\big)R(M_i,X_i)\big(Y_i - \gamma(X_i,M_i,a)\big) \mid O^c_{I_\ell}\big]\Big\}\tag{E2}\\
&\quad+\mathbb{E}\big[
K_h(A_i - a)\big(\hat{\lambda}(a, X_i)-\lambda(a,X_i)\big)R(M_i,X_i)\big(Y_i - \gamma(X_i,M_i,a)\big) \mid O^c_{I_\ell} \big]\tag{TR2}\\
&\quad-K_h(A_i - a)\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a,X_i)R(M_i,X_i)\tag{R2}.
\end{align*}
For term (ref), note that
\begin{align*}
&K_h(A_i - a')\big\{\hat{\lambda}(a', X_i)\{\hat{\gamma}(X_i,M_i,a) - \hat{\eta}(a, a', X_i)\}
-\lambda(a', X_i)\{\gamma(X_i,M_i,a) - \eta(a, a', X_i)\}\big\}\\
&=K_h(A_i - a')\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\tag{CS5}\\
&\quad-K_h(A_i - a')\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)\big(\hat{\eta}(a, a', X_i)-\hat{\eta}(a, a', X_i)\big)\tag{CS6}\\
&\quad+K_h(A_i - a')\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)\gamma(X_i,M_i,a)\tag{R3}\\
&\quad+K_h(A_i - a')\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a', X_i)\tag{R4}\\
&\quad-K_h(A_i - a')\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)\hat{\eta}(a, a', X_i)\tag{R5}\\
&\quad-K_h(A_i - a')\big(\hat{\eta}(a, a', X_i)-\hat{\eta}(a, a', X_i)\big)\lambda(a', X_i)\tag{R6}.
\end{align*}
Next, we group terms (R1)-(R6) as follows. We pair (R1) with (R6), (R2) with (R4), and (R3) with (R5). Note that every expectation introduced here is only over $O_i$, conditional on $O^c_{I_\ell}$, i.e., $\mathbb{E}(\cdot| O^c_{I_\ell})$, and hence all the terms are random variables. For (R1)+(R6) we have
\begin{align*}
&(R1)+(R6)\\
&=\big(\hat{\eta}(a, a', X_i)-\eta(a, a', X_i)\big)
-K_h(A_i - a')\big(\hat{\eta}(a, a', X_i)-\eta(a, a', X_i)\big)\lambda(a', X_i)\\
&=\big(\hat{\eta}(a, a', X_i)-\eta(a, a', X_i)\big)
-\mathbb{E}\big[\hat{\eta}(a, a', X_i)-\eta(a, a', X_i) \big]\tag{E3}\\
&\quad-\Big\{
K_h(A_i - a')\big(\hat{\eta}(a, a', X_i)-\eta(a, a', X_i)\big)\lambda(a', X_i)\\
&-\mathbb{E}\big[K_h(A_i - a')\big(\hat{\eta}(a, a', X_i) -\eta(a, a', X_i)\big)\lambda(a', X_i) \mid O^c_{I_\ell} \big]\Big\}\tag{E4}\\
&\quad+\mathbb{E}\big[
\big(\hat{\eta}(a, a', X_i)-\eta(a, a', X_i)\big)\big(1-K_h(A_i - a')\lambda(a', X_i)\big) \mid O^c_{I_\ell}
\big]\tag{TR3}.
\end{align*}
For (R2)+(R4) we have
\begin{align*}
&(R2)+(R4)\\
&=
-K_h(A_i - a)\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a,X_i)R(M_i,X_i)
+\\
& K_h(A_i - a')\big(\hat{\gamma}_a(M_i,X_i)-\gamma(X_i,M_i,a)\big)\lambda(a', X_i)\\
&=-\Big\{
K_h(A_i - a)\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a,X_i)R(M_i,X_i)\\
&\quad\quad\quad-
\mathbb{E}\big[
K_h(A_i - a)\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a,X_i)R(M_i,X_i)
\mid O^c_{I_\ell} \big]
\Big\}\tag{E5}\\
&\quad+\Big\{
K_h(A_i - a')\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a', X_i)\\
&\quad\quad\quad-
\mathbb{E}\big[
K_h(A_i - a')\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a', X_i) \mid O^c_{I_\ell} \big]
\Big\}\tag{E6}\\
&\quad+\mathbb{E}\big[
\big(\hat{\gamma}_a(M_i,X_i)-\gamma(X_i,M_i,a)\big)
\big\{
K_h(A_i - a')\lambda_{a'}(X_i)
-K_h(A_i - a)\lambda(a,X_i)R(M_i,X_i)
\big\} \mid O^c_{I_\ell} \big]
\tag{TR4}.
\end{align*}
For (R3)+(R5) we have
\begin{align*}
&(R3)+(R5)\\
&=
K_h(A_i - a')\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)\gamma(X_i,M_i,a)
-K_h(A_i - a')\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)\eta(a, a', X_i)\\
&=\Big\{
K_h(A_i - a')\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)\gamma(X_i,M_i,a)\\
&\quad\quad\quad-
\mathbb{E}\big[
K_h(A_i - a')\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)\gamma(X_i,M_i,a) \mid O^c_{I_\ell} \big]
\Big\}\tag{E7}\\
&\quad-\Big\{
K_h(A_i - a')\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)\eta(a, a', X_i)\\
&\quad\quad\quad-
\mathbb{E}\big[
K_h(A_i - a')\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)\eta(a, a', X_i) \mid O^c_{I_\ell} \big]
\Big\}\tag{E8}\\
&\quad+\mathbb{E}\big[
K_h(A_i - a')\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)
\big\{
\gamma(X_i,M_i,a)-\eta(a, a', X_i)
\big\} \mid O^c_{I_\ell} \big]\tag{TR5}.
\end{align*}
And so, to prove $\sqrt{\frac{h^{d_A}}{n}} \sum_{i \epsilon I_\ell} \left\{m(O_i;\hat{\alpha}_\ell,\hat{\lambda}_\ell,\hat{\gamma}_\ell,\psi_0(a,a')) - m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a')) \right\} = o_p(1)$, we provide proofs for the convergence of the terms (CS1) - (CS6), (E1) - (E8) and (TR1) - (TR5) in the following sub-sections.
\subsubsection{Proof for Terms (CS1)-(CS6)}
All of these terms contain the product of two or more errors and can be treated similarly. We provide a detailed proof for (CS2), and a similar method can be followed for the rest of the terms.
For (CS2), write $\Delta_{i\ell}=K_h(A_i - a)\big[\hat{R}(M_i,X_i)-R(M_i,X_i)\big]\big[\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big]\big[Y_i - \gamma(X_i,M_i,a)\big]$. Following Lemma (ref), it suffices to bound $\mathbb{E}\left[|\sqrt{\frac{h^{d_A}}{n}} \sum_{i \in I_\ell} \Delta_{i\ell} | \Big| O^c_{I_\ell}\right]$ as $o_p(1)$ in order to show that $$\sqrt{\frac{h^{d_A}}{n}} \sum_{i \in I_\ell}\Delta_{i\ell} = o_p(1).$$
First, from the triangle inequality, $\mathbb{E}\left[\left|\sqrt{\frac{h^{d_A}}{n}} \sum_{i \in I_\ell} \Delta_{i\ell} \right| \mid O^c_{I_\ell}\right] \leq \frac{1}{L}\sqrt{nh^{d_A}} \mathbb{E}\left[\left|\Delta_{i\ell} \right| \mid O^c_{I_\ell}\right]$, and so it suffices to bound $\sqrt{nh^{d_A}}\mathbb{E} \bigg[\big|\Delta_{i\ell} \big|\bigg| O^c_{I_\ell}\bigg]$. In the interest of space, we introduce the following notation $\Tilde{k}(u) = \prod_{j = 1}^{d_A} k(u_j)$, where $u$ is a vector in $\mathbb{R}^{d_A}$.
{
\begin{align*}
&\sqrt{nh^{d_A}}\mathbb{E} \bigg[\big|\Delta_{i\ell} \big|\bigg| O^c_{I_\ell}\bigg] \\
=& \sqrt{nh^{d_A}}\int\bigg| K_h(A_i - a)\big[\hat{R}(M_i,X_i)-R(M_i,X_i)\big]\big[\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big]\big[Y_i - \gamma(X_i,M_i,a)\big]\bigg| \\
&f(Y_i, A_i, M_i, X_i)dO_i\\
=&\sqrt{nh^{d_A}} \int\bigg|\Tilde{k}(u)\big[\hat{R}(M_i,X_i)-R(M_i,X_i)\big]\big[\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big]\left[Y_i - \gamma(X_i,M_i,a)\right]\bigg|\\
&f(Y_i, uh+a, M_i, X_i) dudY_idM_idX_i\\
=& \sqrt{nh^{d_A}}\int\left\{\int \bigg|\Tilde{k}(u)f(uh+a|M_i,X_i)\bigg\{\int \bigg|\left[Y_i - \gamma(X_i,M_i,a)\right]\bigg| f(Y_i | uh+a, M_i, X_i) dY_i\bigg\}\bigg|du\right\}\\
&\bigg|\big[\hat{R}(M_i,X_i)-R(M_i,X_i)\big]\big[\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big]f(M_i,X_i)\bigg|dM_idX_i\\
\end{align*}
}
Next, Assumption 3.1 on the boundedness of $\gamma(X,M,a)$ and Assumption 3.3 on the boundedness of $var(Y_i|a,m,x)$, along with an application of Lemma (ref) on $f(a \mid M, X)$, we get
\begin{align*}
=& O(\sqrt{nh^{d_A}} )\int
\left\{f(a \mid M_i, X_i) + O(h^2) \right\}\bigg|\big[\hat{R}(M_i,X_i)-R(M_i,X_i)\big]\big[\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big]\bigg|\\
&\qquad \times f(M_i,X_i)dM_idX_i\\
=& O(\sqrt{nh^{d_A}} )\int
f(a \mid M_i, X_i) \bigg|\big[\hat{R}(M_i,X_i)-R(M_i,X_i)\big]\big[\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big]\bigg|f(M_i,X_i)dM_idX_i\\
&+ O(\sqrt{nh^{d_A + 4}})\int
\bigg|\big[\hat{R}(M_i,X_i)-R(M_i,X_i)\big]\big[\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big]\bigg|f(M_i,X_i)dM_idX_i\\
\overset{(a)}{\le} & O(\sqrt{nh^{d_A}} )\int
\bigg|\big[\hat{R}(M_i,X_i)-R(M_i,X_i)\big]\big[\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big]\bigg|f(M_i,X_i)dM_idX_i\\
&+ O(\sqrt{nh^{d_A + 4}})\int
\bigg|\big[\hat{R}(M_i,X_i)-R(M_i,X_i)\big]\big[\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big]\bigg|f(M_i,X_i)dM_idX_i\\
\overset{(b)}{\le} & O\bigg(\sqrt{nh^{d_A}} \bigg\{ \int\big[\hat{R}(M_i,X_i)-R(M_i,X_i)\big] ^2f(M_i,X_i)dM_idX_i \\
&\int\big[\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big] ^2f(M_i,X_i)dM_idX_i\bigg\}^{1/2}\bigg) + o_p(1)\\
&= o_p(1).
\end{align*}
Where $(a)$ follows from an application of Holder's inequality combined with Assumption 3.1 on the boundedness of $f(a \mid M, X)$, and (b) and the last equality follows from an application of Cauchy-Schwartz, combined with Assumption 5.1 and $nh^{d_A + 4} \rightarrow C_h$ by Assumption 2.
\subsubsection{Proof for Terms (E1)-(E8)}
Terms (E1)-(E8) are normalized terms of the form of a bias times a bounded quantity; they can all be treated similarly. We only provide the proof of the convergence in probability to zero for the term (E2). (E2) is given as
\begin{align*}
&K_h(A_i - a)\big(\hat{\lambda}(a, X_i)-\lambda(a,X_i)\big)R(M_i,X_i)\big(Y_i - \gamma(X_i,M_i,a)\big)\\
-&\mathbb{E}\big[K_h(A_i - a)\big(\hat{\lambda}(a, X_i)-\lambda(a,X_i)\big)R(M_i,X_i)\big(Y_i - \gamma(X_i,M_i,a)\big)\mid O^c_{I_\ell}\big]
\end{align*}
To prove $\sqrt{nh^{d_A}}$ times (E2) is $o_p(1)$, we set $\hat{\Delta}_{i\ell}$ as (E2).
By construction, $O^c_{I_\ell}$ and $O_i$ are independent, $i\in I_\ell$, and consequently $\mathbb{E}\left[\hat{\Delta}_{i\ell}|O^c_{I_\ell}\right]=0$ and $\mathbb{E}\left[\hat{\Delta}_{i\ell}\hat{\Delta}_{j\ell}|O^c_{I_\ell}\right]=0$ for $i,j \in I_\ell$ and all $a', a\in \mathcal{A}_0$. Next we note that
\begin{align*}
&h^{d_A}\mathbb{E}\left[\hat{\Delta}^2_{i\ell}|O^c_{I_\ell}\right]\\
= & h^{d_A}\int K^2_h(A_i-a) \left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]^2R^2_i(M_i,X_i)\left[Y_i - \gamma(X_i,M_i,a)\right]^2 f(Y_i, A_i, M_i, X_i) dO_i\\
=& \int \Tilde{k}(u)^2\left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]^2R^2_i(M_i,X_i)\left[Y_i - \gamma(X_i,M_i,a)\right]^2 f(Y_i, uh+a, M_i, X_i) dudY_idM_idX_i\\
=& \int \int \Tilde{k}(u)^2f(uh+a|M_i,X_i)\bigg\{\int \left[Y_i - \gamma(X_i,M_i,a)\right]^2 f(Y_i | uh+a, M_i, X_i) dY_i\bigg\}du\\
&\left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]^2R^2_i(M_i,X_i)f(M_i,X_i)dM_idX_i\\
\overset{(a)}{=} & O\bigg(\int\Tilde{k}(u)^2du \int\left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]^2R^2_i(M_i,X_i)f(M_i,X_i)dM_idX_i\bigg)\\
\overset{(b)}{=} & O(1)\int\left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]^2R^2_i(M_i,X_i)f(M_i,X_i)dM_idX_i\\
\overset{(c)}{=} & o_p(1)
\end{align*}
Where (a) follows from Assumption 3.1 on the boundedness of $f(a \mid M, X)$, along with Assumption 3.1 and Assumption 3.3 combined with the derivation provided below
\begin{align*}
&\int_{\mathcal{Y}}\left[Y_i - \gamma(X_i,M_i,a)\right]^2 f(Y_i | uh+a, M_i, X_i) dY_i\\
=& \int_{\mathcal{Y}}\left[Y^2_i + \gamma^2_a(M_i, X_i)-2\gamma(X_i,M_i,a)Y_i\right] f(Y_i | uh+a, M_i, X_i) dY_i\\
=& \mathbb{E}[Y^2_i | uh+a, M_i, X_i] + \gamma^2_a(M_i, X_i) -2\gamma(X_i,M_i,a) \int_{\mathcal{Y}} Y_i f(Y_i | uh+a, M_i, X_i) dY_i\\
=&\mathbb{E}[Y^2_i | uh+a, M_i, X_i] + \gamma^2_a(M_i, X_i) -2\gamma(X_i,M_i,a) - 2\gamma(X_i,M_i,a)\gamma_{uh+a}(M_i,X_i)\\
=& O(1).
\end{align*}
Next, (b) follows from Assumption 2.4, and finally, (c) follows Assumption 3.2 along with Assumption 4.1. Then
$\mathbb{E} \bigg[\left(\sqrt{h^{d_A}/n} \sum^L_{l=1}\sum_{i\in I_\ell} \hat{\Delta}_{i\ell}\right)^2\bigg| O^c_{I_\ell}\bigg] =h^{d_A} /n \sum^L_{\ell=1}\sum_{i\in I_\ell} \mathbb{E}\left[\hat{\Delta}^2_{i\ell}|O^c_{I_\ell}\right] = h^{d_A}\mathbb{E}\left[\hat{\Delta}^2_{i\ell}|O^c_{I_\ell}\right] = o_p(1).$
Applying Lemma 1 to the above gives $\sqrt{h^{d_A}/n} \sum^L_{l=1}\sum_{i\in I_\ell} \hat{\Delta}_{i\ell}\xrightarrow[]{P} 0$, i.e. $\sqrt{nh^{d_A}}$ times (E2) being $o_p(1)$.
\subsubsection{Proof for Terms (TR1)-(TR5)}
The proofs of the convergence in probability to zero for the terms (TR1)-(TR5) require extra considerations, and we prove them on a case by case basis below.
Proof for Terms TR1 and TR2
Terms (TR1) and (TR2) are similar; we only provide the proof of the convergence in probability to zero for the term (TR2).
To bound TR2, first set $\hat{\Delta}_{i\ell}=K_h(A_i - a) \left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]R(M_i, X_i)\left[Y_i - \gamma(X_i,M_i,a)\right]$. Bounding (TR2) amounts to showing $\sqrt{nh^{d_A}}\mathbb{E}[\hat{\Delta}_{i\ell}|O^c_{I_\ell}] = o_p(1)$.
\begin{align*}
&\sqrt{nh^{d_A}}\mathbb{E} \bigg[\hat{\Delta}_{i\ell} \bigg| O^c_{I_\ell}\bigg] \\
=& \sqrt{nh^{d_A}}\mathbb{E}\bigg\{K_h(A_i - a)
\left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]
R(M_i, X_i)
\left[Y_i - \gamma(X_i,M_i,a)\right]\bigg| O^c_{I_\ell}\bigg\}\\
=& \sqrt{nh^{d_A}}\int_{\mathcal{O}} K_h(A_i-a) \left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]R(M_i, X_i)\left[Y_i - \gamma(X_i,M_i,a)\right] f(Y_i, A_i, M_i, X_i) dO_i\\
=&\sqrt{nh^{d_A}}\int \left[ \int_{\mathcal{A}} K_h(A_i-a)f(A_i \mid Y_i, M_i, X_i) dA_i\right] \left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]\\
&R(M_i, X_i)\left[Y_i - \gamma(X_i,M_i,a)\right]f(Y_i,M_i, X_i) dY_idM_idX_i\\
\intertext{Applying Lemma (ref) under Assumption 3.1}
=&\sqrt{nh^{d_A}}\int \left[ f(a \mid Y_i, M_i, X_i) +O(h^2) \right] \\
&\left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]R(M_i, X_i)\left[Y_i - \gamma(X_i,M_i,a)\right]f(Y_i,M_i, X_i) dY_idM_idX_i\\
\overset{(a)}{=}&\sqrt{nh^{d_A}}\int O(h^2) \left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]\\
&R(M_i, X_i)\left[Y_i - \gamma(X_i,M_i,a)\right]f(Y_i,M_i, X_i) dY_idM_idX_i\\
\overset{(b)}{=} &O(\sqrt{nh^{d_A + 4}})\int \Big|\left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]R(M_i, X_i)\Big|\\
&\left[\int | Y_i-\gamma(X_i,M_i,a) | f(Y_i\mid M_i, X_i) dY_i \right] f(M_i,X_i)dM_idX_i\\
\overset{(c)}{=}& o_p(1)
\end{align*}
where (a) follows from
\begin{align*}
&\int [Y_i -\gamma(X_i,M_i,a)] f(Y_i\mid a, M_i, X_i) dY_i \\
=& \int Y_i f(Y_i\mid a, M_i, X_i) dY_i-\gamma(X_i,M_i,a) = 0,
\end{align*}
(b) is from the exchange of $O(\cdot)$ and integration,
(c) follows from Assumption 2 ($nh^{d_A + 4} \rightarrow C_h$, $h \rightarrow 0$), Assumption 3 and Assumption 4.1, Cauchy-Schwartz combined with the boundedness of $\int |Y_i -\gamma(X_i,M_i,a)| f(Y_i\mid M_i, X_i) dY_i $ derived from Assumption 3.1 shown below
\begin{align*}
&\int |Y_i -\gamma(X_i,M_i,a)| f(Y_i\mid M_i, X_i) dY_i \\
=& \int \bigg[\int |Y_i -\gamma(X_i,M_i,a)| f(Y_i\mid a, M_i, X_i) dY_i \bigg]f(a|M_i,X_i)da \\
\le & \int \bigg[Var(Y_i|a,M_i, X_i)\bigg]^{1/2} f(a|M_i,X_i)da < \infty,
\end{align*}
where the last line also comes from the Cauchy-Schwartz inequality.
\begin{comment}
Equality $(a)$ follows from that
\begin{align*}
&\sqrt{nh^{d_A}}\int f(a \mid Y_i, M_i, X_i)\left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]R(M_i, X_i)Y_i f(Y_i,M_i, X_i) dY_idM_idX_i\\
=& \sqrt{nh^{d_A}}\int \left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]R(M_i, X_i)
\bigg[ \intY_i f(Y_i \mid a, M_i, X_i)dY_i \bigg]f(a, M_i, X_i) dM_idX_i\\
=& \sqrt{nh^{d_A}}\int \left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]R(M_i, X_i)
\gamma(X_i,M_i,a)f(a, M_i, X_i) dM_idX_i\\
=& \sqrt{nh^{d_A}}\int \left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]R(M_i, X_i)
\gamma(X_i,M_i,a)f(a, Y_i, M_i, X_i) dY_idM_idX_i\\
=&\sqrt{nh^{d_A}}\int f(a \mid Y_i, M_i, X_i)\left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]R(M_i, X_i)\gamma(X_i,M_i,a) f(Y_i,M_i, X_i) dY_idM_idX_i\\
\end{align*}
\end{comment}
Proof for TR3
For Term (TR3), we have
\begin{align*}
&\sqrt{nh^{d_A}}\mathbb{E}\left[
\big(\hat{\eta}(a, a', X_i)-\eta(a, a', X_i)\big)\big(1-K_h(A_i - a')\lambda(a', X_i)\big)
\big| O^c_{I_\ell}\right] \\
&=\sqrt{nh^{d_A}}\int
\big(\hat{\eta}(a, a', X_i)-\eta(a, a', X_i)\big)\big(1-K_h(A_i - a')\lambda(a', X_i)\big)f(A_i,X_i)dA_idX_i \\
&=\sqrt{nh^{d_A}}\int
\big(\hat{\eta}(a, a', X_i)-\eta(a, a', X_i)\big)\big(1-\bigg\{\int K_h(A_i - a')f(A_i\mid X_i)dA_i\bigg\}\lambda(a', X_i)\big)f(X_i)dX_i \\
&\overset{(a)}{=}\sqrt{nh^{d_A}}\int
\big(\hat{\eta}(a, a', X_i)-\eta(a, a', X_i)\big)\big(1- f(a'\mid X_i)\lambda_{a'}(X_i)\big)f(X_i)dX_i \\
&\quad+\sqrt{nh^{d_A}}\int
\big(\hat{\eta}(a, a', X_i)-\eta(a, a', X_i)\big)O(h^2)\lambda_{a'}(X_i)f(X_i)dX_i\\
&\overset{(b)}{=} o_p(1).
\end{align*}
where $(a)$ follows from Lemma (ref), and (b) follows from the definition of $\lambda_{a^\prime}(X_i)$, $nh^{d_A+4}\rightarrow C_h$, Assumption 4 (convergence of $\hat{\eta}(X_i)$), Assumption 3 (boundedness of $\lambda$) combined with an application of Cauchy-Schwartz inequality.
Proof For TR4
Demonstrating the bound for (TR4), we have
\begin{align*}
&\sqrt{nh^{d_A}}\mathbb{E}\big[
\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)
\big\{
K_h(A_i - a')\lambda(a', X_i)
-K_h(A_i - a)\lambda(a,X_i)R(M_i,X_i)
\big\}
\big]\\
&= \sqrt{nh^{d_A}}\mathbb{E}\big[
\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)
\big\{
K_h(A_i - a')\lambda(a', X_i)
\big\}
\big]\tag{TR4-1}\\
- &\sqrt{nh^{d_A}}\mathbb{E}\big[
\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)
\big\{
K_h(A_i - a)\lambda(a,X_i)R(M_i,X_i)
\big\}
\big]\tag{TR-4-2}
\end{align*}
TR-4-1 can be written as
{
\begin{align*}
&\sqrt{nh^{d_A}}\mathbb{E}\big[\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\big\{K_h(A_i - a')\lambda(a', X_i)\big\}\big]\\
&= \sqrt{nh^{d_A}} \int \big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a', X_i)\left\{ \int K_h(A_i - a') f(A_i \mid M_i, X_i) dA_i \right\} f(M_i, X_i) dM_i dX_i
\end{align*}
}
An application of Lemma (ref) to TR-4-1 gives
\begin{align*}
&\sqrt{nh^{d_A}}\mathbb{E}\big[
\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)
\big\{
K_h(A_i - a')\lambda(a', X_i)
\big\}
\big]\\
&=\sqrt{nh^{d_A}} \int \big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a', X_i)f(a' \mid M_i, X_i)f(M_i, X_i) dM_i dX_i\\
& + \sqrt{nh^{d_A}}\int \big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a', X_i)O(h^2)f(M_i, X_i) dM_i dX_i
\end{align*}
A similar approach applied to TR-4-2 gives
\begin{align*}
&\sqrt{nh^{d_A}}\mathbb{E}\big[
\big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)
\big\{
K_h(A_i - a)\lambda(a,X_i)R(M_i,X_i)
\big\}
\big]\\
=&\sqrt{nh^{d_A}} \int \big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a,X_i)R(M_i,X_i)f(a \mid M, X) f(M_i, X_i) dM_i dX_i\\
&+ \sqrt{nh^{d_A}}\int \big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a,X_i)R(M_i,X_i)O(h^2)f(M_i, X_i) dM_i dX_i
\end{align*}
Now, the first terms of TR-4-1 and TR-4-2 cancel out with each other, shown below
{
\begin{align*}
&\int \big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\left\{\lambda(a', X_i)f(a' \mid M, X)
- \lambda(a,X_i)R(M_i,X_i)f(a \mid M, X)\right\} f(M_i, X_i) dM_i dX_i = 0
\end{align*}
}
This can be seen from
\begin{align*}
\lambda(a', X_i)f(a' \mid M_i, X_i) = \frac{f(X_i)}{f(a', X_i)}\frac{f(a', M_i, X_i)}{f(M_i, X_i)}
\end{align*}
Along with
\begin{align*}
\lambda(a,X_i)R(M_i,X_i)f(a \mid M_i, X_i) &= \frac{f(X_i)}{f(a, X_i)}\frac{f(M_i, a', X_i)}{f(a', X_i)}
\frac{f(a, X_i)}{f(M_i, a, X_i)}\frac{f(a, M_i, X_i)}{f(M_i, X_i)}\\
&= \frac{f(X_i)}{f(a', X_i)}\frac{f(M_i , a', X_i)}{f(M_i, X_i)}
\end{align*}
Consequently the first terms in TR4-1 and TR4-2 cancel each other out, and this leaves us to bound the remaining terms.
\[
\sqrt{nh^{d_A}} \int \big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a', X_i) O(h^2)f(M, X) dM_i dX_i = o_p(1)
\]
The second term in TR-4-1 and TR-4-2 can be bounded by an application of Cauchy-Schwartz, combined with Assumption 4 (consistency of $\hat{\gamma}$) and boundedness of $\lambda$ in Assumption 3.1.
Proof For TR5
Finally, for term (TR5), we note that
\begin{align*}
&\sqrt{nh^{d_A}}\mathbb{E}\left[
K_h(A_i - a')\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)
\big\{
\gamma(X_i,M_i,a)-\eta(X_i)
\big\}\big| O^c_{I_\ell}
\right]\\
&=\sqrt{nh^{d_A}}\int
K_h(A_i - a')\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)
\big\{
\gamma(X_i,M_i,a)-\eta(a, a', X_i)
\big\}f(A_i,M_i,X_i)dA_idM_idX_i\\
&=\sqrt{nh^{d_A}}\int
\Bigg\{\int K_h(A_i - a')f(A_i\mid M_i,X_i)dA_i\Bigg\}
\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)\\
&\times \big\{
\gamma(X_i,M_i,a)-\eta(a, a', X_i)
\big\}f(M_i,X_i)dM_idX_i\\
&\overset{(a)}{=}\sqrt{nh^{d_A}}\int \big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)
\big\{
\gamma(X_i,M_i,a)-\eta(a, a', X_i)
\big\}f(a',M_i,X_i)dM_idX_i\\
&\quad+\sqrt{nh^{d_A}}\int \big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)
\big\{
\gamma(X_i,M_i,a)-\eta(a, a', X_i)
\big\}O(h^2)f(M_i,X_i)dM_idX_i\\
&\overset{(b)}{=} 0 +O(\sqrt{nh^{d_A + 4}})\int \Bigg|\big(\hat{\lambda}(a', X_i)-\lambda(a', X_i)\big)
\big\{
\gamma(X_i,M_i,a)-\eta(a, a', X_i)
\big\}\Bigg|f(M_i,X_i)dM_idX_i\\
&\overset{(c)}{=} o_p(1)
\end{align*}
Where $(a)$ follows from an application of Lemma (ref), (b) follows from the definition of $\eta$, and (c) follows from an application of Cauchy-Schwartz combined with the consistency of $\hat{\lambda}$.
Proof For Asymptotic Normality
The proof for asymptotic normality follows from an application of the Lyapunov Central Limit theorem to the terms $\sqrt{nh^{d_A}}n^{-1} m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))$. We first prove the Lyapunov condition holds for $\delta = 1$, i.e.
\[
\lim_{n \rightarrow \infty} \frac{1}{s^{3}_n} \sum_{i = 1}^n \mathbb{E}\left[\Big|\sqrt{nh^{d_A}}n^{-1} m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a')) - \mu_{i} \Big|^3\right] = 0
\]
Where $\mu_i$ equals $\mathbb{E}\left[\sqrt{nh^{d_A}}n^{-1} m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\right]$ and $s^2_n = \sum_{i = 1}^n \sigma^2_i$ where $\sigma^2_i$ is the variance of of $\sqrt{nh^{d_A}}n^{-1} m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))$. To prove the Lyaponuv condition holds, we first derive $\mu_i$ and $\sigma^2_i$.
Calculation for $B(a, a')$ and $\mu_i$
Given
align*[align* omitted — 324 chars of source]
Since $\mathbb{E}[\eta(a, a', X_i) - \psi_0(a,a')] = 0$, we focus on
$\frac{K_h(A - a)f(M \mid A = a', X)}{f(M \mid A = a, X)f(a \mid X)}\{Y - \mathbb{E}[Y \mid X, M , A = a]\}+ \frac{K_h(A - a')}{f(a' \mid X)}\{\mathbb{E}[Y \mid X, M, A = a] - \eta(a, a', X)\}$.
We start by computing the expectation of each the individual terms one at a time.
Expectation Part 1
$$\frac{K_h(A - a)f(M \mid A = a', X)}{f(M \mid A = a, X)f(a \mid X)}\{Y - \mathbb{E}[Y \mid X, M , A = a]\}$$
From $\mathbb{E}\{\mathbb{E}[\gamma(X,M,a)|X,M]\} = \mathbb{E}\{\mathbb{E}[Y|X,M]\}$ and $\gamma(X,M,a) = \mathbb{E}(Y|X,M,A)$, expectation of the first term
align*[align* omitted — 477 chars of source]
The inner product further expands as follows,
align*[align* omitted — 988 chars of source]
for all $X,M$ in respective range. Inserting this back into the original expectation we get,
{
align*[align* omitted — 482 chars of source]
}
Expectation Part 2
$$\frac{K_h(A - a')}{f(a' \mid X)}\{\gamma(X,M,a) - \eta(a, a', X)\}$$
align*[align* omitted — 483 chars of source]
The inner expectation can be written as
align*[align* omitted — 724 chars of source]
Plugging this back into the above expectation
{
align*[align* omitted — 611 chars of source]
}
from having the first term in this expectation equal to zero, which we prove below
align*[align* omitted — 410 chars of source]
Hence, letting
align*[align* omitted — 447 chars of source]
we have $\mathbb{E}\left[m(O_i;\alpha,\lambda,\gamma, \psi_0(a,a'))\right] = h^2 B(a,a')$. Additionally from this derivation $\mathbb{E}\left[\sqrt{nh^{d_A}}n^{-1} m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\right] = O(\sqrt{\frac{h^{d_A + 4}}{n}})$. Next, we prove the properties of variance.
Calculation for $V(a, a')$ and $s^2_n$
From the definition of $s^2_n$, we have
align*[align* omitted — 215 chars of source]
Consequently, we calculate
align*[align* omitted — 313 chars of source]
Using the property that $var(X) = \mathbb{E}[X^2] - \mathbb{E}[X]^2$ and constant values do not contribute to the variance, the variance term above can be re-written as
align*[align* omitted — 570 chars of source]
Examining each of the terms above one by one, the first term can be expanded as
align*[align* omitted — 974 chars of source]
We analyze each of these terms part by part.
Variance Part 1
align*[align* omitted — 1,251 chars of source]
Because $0 < \int u^6 k(u) du < \infty$ from Assumption 2 (3), we also have boundedness of $\int u^6k^2(u)du$. The inner expectation can be written as
align*[align* omitted — 1,179 chars of source]
where $\Bar{a}_v,\Bar{a}_\gamma$, and $\Bar{a}_f$ are between $a$ and $a+uh$. Hence, part 1 of the variance
align*[align* omitted — 310 chars of source]
Variance Part 2
align*[align* omitted — 303 chars of source]
The inner expectation can be written as
align*[align* omitted — 939 chars of source]
the last equation is from
align*[align* omitted — 168 chars of source]
Hence, the part 2 of variance
align*[align* omitted — 271 chars of source]
Variance Part 3
align*[align* omitted — 78 chars of source]
This holds because we assume $\eta$ is bounded.
Variance Part 4
align*[align* omitted — 912 chars of source]
The inner expectation
align*[align* omitted — 875 chars of source]
Hence, the part 4 of variance
align*[align* omitted — 252 chars of source]
Variance Part 5
align*[align* omitted — 381 chars of source]
Applying the same expansion in Expectation Part 1, we can write the inner expectation as
align*[align* omitted — 212 chars of source]
Inserting this back into the full expectation, combined with the boundedness of $\eta$, $f(M \mid a, X)$ and $f(a \mid X)$, we get
align*[align* omitted — 201 chars of source]
Variance Part 6
align*[align* omitted — 470 chars of source]
Using the expansion from Part 2 of the expectation on $\mathbb{E}[K_h(A - a')| X,M]$, we get
align*[align* omitted — 144 chars of source]
Plugging this back into the full expectation, we get
align*[align* omitted — 178 chars of source]
And using the calculation for the bias,
align*[align* omitted — 314 chars of source]
Finally, putting the pieces of the variance together, we have
align*[align* omitted — 330 chars of source]
where the term converges to $V(a, a^\prime)$ as $h\rightarrow 0$ and
align*[align* omitted — 226 chars of source]
Having derived the bias and variance terms, we now prove the Lyapunov condition for $\delta = 1$.
Proof for Lyapunov Condition
We now prove the Lyapunov condition
\[
\lim_{n \rightarrow \infty} \frac{1}{s^{3}_n} \sum_{i = 1}^n \mathbb{E}\left[\Big|\sqrt{nh^{d_A}}n^{-1} m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a')) - \mu_{i} \Big|^3\right] = 0
\]
Note that
$$\Big|\sqrt{nh^{d_A}}n^{-1} m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a')) - \mu_{i} \Big| \leq \Big|\sqrt{nh^{d_A}}n^{-1} m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\Big| + \Big| \mu_{i} \Big|$$
Since both sides are positive,
align*[align* omitted — 405 chars of source]
From the monotonicity of the expected value, we have
align*[align* omitted — 601 chars of source]
Since $\sum_{i = 1}^n\Big| \mu_{i} \Big|^3 = O(h^{(d_A + 4)3/2}n^{-1/2}) = o(1)$, $\sum_{i = 1}^n 3h^{d_A}n^{-1}\mathbb{E}\left[\Big| m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\Big|^2\right]\Big|\mu_i\Big| = O(\sqrt{\frac{h^{d_A + 4}}{n}}) = o(1)$, and $\sum_{i = 1}^n 3(h^{d_A}n^{-1})^{1/2}\mathbb{E}\left[\Big| m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\Big|\right]\Big|\mu_i\Big|^2 = o(1)$, it suffices to prove the following condition
\[
\lim_{n \rightarrow \infty} \frac{1}{s^{3/2}_n} \sum_{i = 1}^n \mathbb{E}\left[\Big|\sqrt{nh^{d_A}}n^{-1} m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a')) \Big|^3\right] = 0
\]
Following a similar derivation as in the proof for consistency of $\hat{V}(a, a')$, from the assumption $\mathbb{E}\{[Y-\gamma(X,M,a)]^3 | A = a', M = m, X = x\}$ over any $(a,a',m,x)\in \mathcal{A}\times\mathcal{A}\times \mathcal{M} \times \mathcal{X}$, along with $\int^{\infty}_{-\infty} k(u)^{c_1}k(u+\tilde{c})^{c_2} du < \infty$ and $\int^{\infty}_{-\infty} u^2k(u)^{c_1}k(u+\tilde{c})^{c_2} du < \infty$ for $\tilde{c} \in \mathcal{R}$ and $c_1+c_2 \in \{2,3\}$ for $c_1,c_2\in\{0,1,2,3\}$, we can bound $\mathbb{E}\left[\Big|m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a')) \Big|^3\right] = O(\frac{1}{h^{2d_A}})$. Hence
\[
\sum_{i = 1}^n \mathbb{E}\left[\Big|\sqrt{nh^{d_A}}n^{-1} m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a')) \Big|^3\right] = O((nh^{d_A})^{-1/2}) = o(1)
\]
Combining this with $s^2_n = V(a, a') + o(1)$ proves the Lyapunov condition. Hence,
$$
\frac{1}{s_n} \sum_{i = 1}^n \left(\sqrt{\frac{h^{d_A}}{n}}m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a')) - \mu_i\right) \xrightarrow[]{d} \mathcal{N}(0, 1)
$$
An application of Slutsky's theorem provides the desired result that
$$\sqrt{nh^{d_A}} (\hat{\psi}^{MR}(a,a')- \psi_0(a,a') -B(a,a')) \xrightarrow[]{d} N(0, V(a,a'))$$
comment\section{Proof for Consistency of $\widehat{B}(a, a^\prime)$}
Claim: The estimator $\widehat{B}(a, a^\prime)$ defined as
\begin{align*}
\widehat{B}(a, a^\prime) &= \frac{\hat{\psi}^{MR}_b(a,a') - \hat{\psi}^{MR}_{\rho b}(a,a')}{b^2(1 - \rho^2)}
\end{align*}
Where $a \in (0, 1)$, and $\hat{\psi}^{MR}_{ab}(a,a')$ denotes our estimate for $\psi_0(a, a^\prime)$ using kernel bandwidth $ab$. Then $\widehat{B}(a, a^\prime)$ is a consistent estimator for $B(a, a^\prime)$.
Assumptions
\begin{enumerate}
• $b \rightarrow 0$ and $nb^{d_A + 4} \rightarrow \infty$
• $\int k(u) k(u^\prime)k(\frac{u}{a})k(\frac{u^\prime}{a})$ is bounded.
\end{enumerate}
Proof
Our final estimator of $\psi_0$ is obtained by
\begin{align*}
\hat{\psi}^{MR}(a,a'&)=\frac{1}{L}\sum_{\ell=1}^L\frac{1}{|I_\ell|}\sum_{i\in I_\ell} \frac{K_h(A_i - a)\hat{f}(M_i \mid A = a', X_i)}{\hat{f}(M_i \mid A = a, X_i)\hat{f}(a \mid X_i)}\{Y_i - \hat{\gamma}(X, M , a)\}\\
&+ \frac{K_h(A_i - a')}{\hat{f}(a' \mid X_i)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X_i)\} + \hat{\eta}(a, a', X_i)
\end{align*}
To prove consistency of $\widehat{B}(a, a^\prime)$, we must show that $\forall \epsilon > 0$
\begin{align*}
\lim_{n \rightarrow \infty} P(|\hat{B}(a, a^\prime) - \mathbb{E}[\hat{B}(a, a^\prime)]| > \epsilon) \rightarrow 0.
\end{align*}
Using Chebyshev's inequality, we have
\begin{align*}
P(|\hat{B}(a, a^\prime) - \mathbb{E}[\hat{B}(a, a^\prime)]| > \epsilon) \leq \frac{var(\hat{B}(a, a^\prime))}{\epsilon}
\end{align*}
Using the properties of variance,
\begin{align*}
var(\hat{B}(a, a^\prime)) &= \frac{1}{b^4(1 - a^2)^2}\Bigg[var(\hat{\psi}^{MR}_b(a, a^\prime)) + var(\hat{\psi}^{MR}_{ab}(a, a^\prime)) - 2cov(\hat{\psi}^{MR}_b(a, a^\prime), \hat{\psi}^{MR}_{\rho b}(ab, a^\prime))\Bigg]
\end{align*}
Following a similar calculation for the variance, $var(\hat{\psi}^TR_ab(a, a^\prime)) = O((nb^{d_A})^{-1})$, and we calculate the covariance as
\begin{align*}
cov(\hat{\psi}^{MR}_b(a, a^\prime), \hat{\psi}^{MR}_{\rho b}(a, a^\prime)) = \mathbb{E}[\hat{\psi}^{MR}_b(a, a^\prime), \hat{\psi}^{MR}_{\rho b}(ab, a^\prime)] - \mathbb{E}[\hat{\psi}^{MR}_b(a, a^\prime)]\mathbb{E}[\hat{\psi}^{MR}_{\rho b}(b, a^\prime)]
\end{align*}
\subsubsection{Calculation of $\mathbb{E}[\hat{\psi}^{MR}_{ab}(a,a')]$}
Following a similar calculation for $B(a, a^\prime)$,
\begin{align*}
\mathbb{E}[\hat{\psi}^{MR}_{ab}(a,a')] &= \frac{1}{L}\sum_{\ell=1}^L\frac{1}{|I_\ell|}\sum_{i\in I_\ell} \mathbb{E}\Bigg[ \frac{K_h(A_i - a)\hat{f}(M_i \mid A = a', X_i)}{\hat{f}(M_i \mid A = a, X_i)\hat{f}(a \mid X_i)}\{Y_i - \hat{\gamma}(X, M , a)\}\\
&+ \frac{K_h(A_i - a')}{\hat{f}(a' \mid X_i)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X_i)\} + \hat{\eta}(a, a', X_i)\Bigg]\\
&= \frac{1}{L}\sum_{\ell=1}^L \mathbb{E}\Bigg[ \frac{K_h(A_i - a)\hat{f}(M_i \mid A = a', X_i)}{\hat{f}(M_i \mid A = a, X_i)\hat{f}(a \mid X_i)}\{Y_i - \hat{\gamma}(X, M , a)\}\\
&+ \frac{K_h(A_i - a')}{\hat{f}(a' \mid X_i)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X_i)\} + \hat{\eta}(a, a', X_i)\Bigg]
\end{align*}
From linearity of expectation, we can write the above as
\begin{align*}
&\frac{1}{L}\sum_{\ell=1}^L \mathbb{E}\Bigg[ \frac{K_h(A_i - a)\hat{f}(M_i \mid A = a', X_i)}{\hat{f}(M_i \mid A = a, X_i)\hat{f}(a \mid X_i)}\{Y_i - \hat{\gamma}(X, M , a)\}\Bigg]\\
&+ \mathbb{E}\Bigg[\frac{K_h(A_i - a')}{\hat{f}(a' \mid X_i)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X_i)\}\Bigg] + \mathbb{E}\Bigg[\hat{\eta}(a, a', X_i)\Bigg]
\end{align*}
Each of the terms are examined and simplified individually below.
\begin{enumerate}
• Estimated Expectation Part 1
\end{enumerate}
\begin{align*}
&\frac{1}{L}\sum_{\ell=1}^L \mathbb{E}\Bigg[ \frac{K_h(A_i - a)\hat{f}( M_i|a', X_i)}{\hat{f}(M_i|a, X_i)\hat{f}(a| X_i)}\{Y_i - \hat{\gamma}(X, M , a)\}\Bigg]\\
=& \mathbb{E}\Bigg\{\mathbb{E}\bigg\{\frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)} \mathbb{E}\bigg[K_h(A - a) (\gamma(X,M,A) - \hat{\gamma}(X,M,a))\bigg| X,M, O^c_{I_\ell}\bigg]\bigg|O^c_{I_\ell}\bigg\}\Bigg\}.
\end{align*}
The innermost expectation further expands as follows,
\begin{align*}
&\mathbb{E}\bigg[K_h(A - a) (\gamma(X,M,A) - \hat{\gamma}(X,M,a))\bigg| X,M,O^c_{I_\ell}\bigg]\\
=& \int K_h(A - a) (\gamma(X,M,A) - \hat{\gamma}(X,M,a)) f(A|X,M) dA\\
=& \int \bigg[\prod^{d_A}_{j=1}\frac{1}{h}k\Big(\frac{A_j-a}{h}\Big)\bigg] (\gamma(X,M,A) - \hat{\gamma}(X,M,a)) f(A|X,M) dA\\
=& \int \bigg[\prod^{d_A}_{j=1} k(u_j)\bigg] (\gamma(X,M,a+uh) - \hat{\gamma}(X,M,a)) f(a+uh|X,M) du\\
=& \int \bigg[\prod^{d_A}_{j=1} k(u_j)\bigg] \bigg(\gamma(X,M,a) - \hat{\gamma}(X,M,a)+\sum^{d_A}_{j=1} u_j h\partial_{a_j} \gamma(X,M,a) +\sum^{d_A}_{j=1}\sum^{d_A}_{j^\prime=1}\frac{u_ju_{j^\prime}h^2}{2} \partial_{a_j}\partial_{a_{j^\prime}} \gamma(X,M,a) \bigg)\\
& \times \bigg( f(a|X,M)+\sum^{d_A}_{j=1} u_j h\partial_{a_j} f(a|X,M) +\sum^{d_A}_{j=1}\sum^{d_A}_{j^\prime=1}\frac{u_ju_{j^\prime}h^2}{2} \partial_{a_j}\partial_{a_{j^\prime}}f(a|X,M)\bigg)du_1\cdots du_{d_A} + O(h^3)\\
=& \big[\gamma(X,M,a) - \hat{\gamma}(X,M,a)\big]f(a|X,M)\\
&+ \int \bigg[\prod^{d_A}_{j=1} k(u_j)\bigg] \big[\gamma(X,M,a) - \hat{\gamma}(X,M,a)\big]\frac{u^2_j h^2}{2} \partial^2_{a_j} f(a|X,M)du_1\cdots du_{d_A}\\
&+h^2 \int u^2 k(u)du \bigg(\sum^{d_A}_{j=1} \partial_{a_j} \gamma(X,M,a) \partial_{a_j} f(a|X,M) + \frac{1}{2} \left[\sum^{d_A}_{j=1} \partial^2_{a_j} \gamma(X,M,a)\right]f(a|X,M) \bigg) + O(h^3)
\end{align*}
for all $X,M$ in respective range. Inserting this back into the original expectation we get,
\begin{align*}
&\mathbb{E}\Bigg\{\mathbb{E}\bigg\{\frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)} \mathbb{E}\bigg[K_h(A - a) (\gamma(X,M,A) - \hat{\gamma}(X,M,a))\bigg| X,M, O^c_{I_\ell}\bigg]\bigg|O^c_{I_\ell}\bigg\}\Bigg\} \\
=& \mathbb{E}_{O^c_{I_\ell}}\Bigg\{\int \frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\big[\gamma(X,M,a) - \hat{\gamma}(X,M,a)\big]f(a|X,M)f(X,M)dXdM\\
&+ \iint \frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)} \bigg[\prod^{d_A}_{j=1} k(u_j)\bigg] \big[\gamma(X,M,a) - \hat{\gamma}(X,M,a)\big]\frac{u^2_j h^2}{2} \partial^2_{a_j} f(a|X,M)f(X,M)du_1\cdots du_{d_A}dXdM \\
&+h^2 \int u^2 k(u)du \\
&\times \mathbb{E}\bigg\{ \frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)} \bigg(\sum^{d_A}_{j=1} \partial_{a_j} \gamma(X,M,a) \partial_{a_j} f(a|X,M) + \frac{1}{2} \left[\sum^{d_A}_{j=1} \partial^2_{a_j} \gamma(X,M,a)\right]f(a|X,M) \bigg)\bigg\} + O(h^3)\Bigg\}
\end{align*}
The first term gets canceled from the subtraction of the estimators.
\begin{align*}
&\iint \frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)} \bigg[\prod^{d_A}_{j=1} k(u_j)\bigg] \big[\gamma(X,M,a) - \hat{\gamma}(X,M,a)\big]\frac{u^2_j h^2}{2} \partial^2_{a_j} f(a|X,M)f(X,M)du_1\cdots du_{d_A}dXdM\\
=& O(h^2) \iint \frac{\hat{f}(M \mid A = a',X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)} \big[\gamma(X,M,a) - \hat{\gamma}(X,M,a)]\big\partial^2_{a_j} f(a|X,M)f(X,M) dXdM\\
\le& O(h^2) \bigg\{\iint \frac{\hat{f}^2(M \mid A = a',X)}{\hat{f}^2(M \mid A = a, X)\hat{f}^2(a \mid X)} \big[\partial^2_{a_j} f(a|X,M)\big]^2f(X,M) dXdM\bigg\}^{1/2}\\
& \times \bigg\{\iint \big[\gamma(X,M,a) - \hat{\gamma}(X,M,a)]^2 f(X,M) dXdM\bigg\}^{1/2}\\
=& o_p(h^2)
\end{align*}
where the last equality comes from the consistency of $\hat{\gamma}$ and the boundedness of $\hat{\alpha}$, $\hat{\lambda}$, and the derivative of $f(a|M,X)$. For the last term, the consistency assumptions imply convergence in probability as below,
$$\frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)} \rightarrow_p \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)}\text{ and }\frac{1}{\hat{f}(a \mid X)} \rightarrow_p \frac{1}{f(a \mid X)}.$$
As a result, we obtain the following from the continuous mapping theorem,
\begin{align*}
& h^2 \int u^2 k(u)du \\
&\times \mathbb{E}\bigg\{ \frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)} \bigg(\sum^{d_A}_{j=1} \partial_{a_j} \gamma(X,M,a) \partial_{a_j} f(a|X,M) + \frac{1}{2} \left[\sum^{d_A}_{j=1} \partial^2_{a_j} \gamma(X,M,a)\right]f(a|X,M) \bigg)\bigg\} \\
=& h^2 \int u^2 k(u)du \\
&\times \mathbb{E}\bigg\{ \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)f(a \mid X)} \bigg(\sum^{d_A}_{j=1} \partial_{a_j} \gamma(X,M,a) \partial_{a_j} f(a|X,M) + \frac{1}{2} \left[\sum^{d_A}_{j=1} \partial^2_{a_j} \gamma(X,M,a)\right]f(a|X,M) \bigg)\bigg\} + o_p(h^2)
\end{align*}
from the boundedness of $\int u^2 k(u)du$ and that $\mathbb{E}\{A(X,M)B(X,M)\}$ is a continuous function of $A(X,M)$ for any functionals $A(\cdot)$ and $B(\cdot)$.
As a result, this estimated expectation part 1 can be written as
\begin{align*}
&\frac{1}{L}\sum_{\ell=1}^L \mathbb{E}\Bigg[ \frac{K_h(A_i - a)\hat{f}( M_i|a', X_i)}{\hat{f}(M_i|a, X_i)\hat{f}(a| X_i)}\{Y_i - \hat{\gamma}(X, M , a)\}\Bigg]\\
=&\mathbb{E}_{O^c_{I_\ell}}\Bigg\{\int \frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\big[\gamma(X,M,a) - \hat{\gamma}(X,M,a)\big]f(a|X,M)f(X,M)dXdM\Bigg\} \\
& +h^2 \int u^2 k(u)du \\
&\times \mathbb{E}\bigg\{ \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)f(a \mid X)} \bigg(\sum^{d_A}_{j=1} \partial_{a_j} \gamma(X,M,a) \partial_{a_j} f(a|X,M) + \frac{1}{2} \left[\sum^{d_A}_{j=1} \partial^2_{a_j} \gamma(X,M,a)\right]f(a|X,M) \bigg)\bigg\} + o_p(h^2),
\end{align*}
where the second term in the summation is exactly the expectation part 1 under known truth.
Estimated Expectation Part 2
$\frac{K_h(A - a')}{\hat{f}(a' \mid X)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\}$
\begin{align*}
&\mathbb{E}\bigg[\frac{K_h(A - a')}{\hat{f}(a' \mid X)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\}\bigg]\\
=&\mathbb{E}\bigg\{\mathbb{E}\bigg[\frac{K_h(A - a')}{\hat{f}(a' \mid X)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\}\bigg| X,M\bigg]\bigg\} \\
=&\mathbb{E}\bigg\{\frac{1}{\hat{f}(a' \mid X)}\mathbb{E}\bigg[K_h(A - a')\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\}\bigg| X,M\bigg]\bigg\}\\
=&\mathbb{E}\bigg\{\frac{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)}{\hat{f}(a' \mid X)}\mathbb{E}[K_h(A - a')| X,M]\bigg\}
\end{align*}
The inner expectation can be written as
\begin{align*}
&\mathbb{E}\bigg[K_h(A - a')\bigg| X,M\bigg]\\
=& \int \bigg[\prod^{d_A}_{j=1}\frac{1}{h}k\Big(\frac{A_j-a'}{h}\Big)\bigg] f(A|X,M) dA \\
=& \int k(u_1)\cdots k(u_{d_A}) \bigg( f(a'|X,M)+ \sum^{d_A}_{j=1} u_jh\partial_{a_j} f(a'|X,M) +\sum^{d_A}_{j=1} \sum^{d_A}_{j^\prime=1}\frac{u_ju_{j^\prime}h^2}{2} \partial_{a_j}\partial_{a_{j^\prime}} f(a^\prime|X,M) \\
&+ \sum^{d_A}_{j=1} \sum^{d_A}_{j^\prime=1}\sum^{d_A}_{j^{\prime\prime}=1}\frac{u_ju_{j^\prime}u_{j^{\prime\prime}}h^3}{2} \partial_{a_j}\partial_{a_{j^\prime}}\partial_{a_{j^{\prime\prime}}} f(\Bar{a}|X,M)\bigg)du_1\cdots du_{d_A}\\
=&f(a'|X,M) + \frac{1}{2}h^2\int u^2k(u)du \sum^{d_A}_{j=1}\partial^2_{a_j} f(a'|X,M) + O(h^3)
\end{align*}
Plugging this back into the above expectation
\begin{align*}
&\mathbb{E}\Bigg\{ \frac{1}{\hat{f}(a'|X)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\}\times \bigg(f(a'|X,M) + \frac{1}{2}h^2\int u^2k(u)du \sum_{j = 1}^{h_{d_A}}\partial^2_{a_j} f(a'|X,M) \bigg)\Bigg\} + O(h^3)\\
=& \mathbb{E}\Bigg\{ \{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\} \bigg(\frac{f(a'|X,M)}{\hat{f}(a'|X)} + \frac{1}{2}h^2\int u^2k(u)du \frac{\sum_{j = 1}^{h_{d_A}} \partial^2_{a_j} f(a'|X,M)}{\hat{f}(a'|X)} \bigg)\Bigg\} + O(h^3)\\
=& \mathbb{E}\Bigg\{ \{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\}\frac{f(a'|X,M)}{\hat{f}(a'|X)}\Bigg\}\\
&+ \mathbb{E}\Bigg\{ \frac{1}{2}h^2\int u^2k(u)du \{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\} \frac{\sum_{j = 1}^{h_{d_A}} \partial^2_{a_j} f(a'|X,M)}{\hat{f}(a'|X)} \bigg)\Bigg\} + O(h^3)\\
\intertext{An application of the continuous mapping theorem allows us to write the second term as, which matches the originally calculated bias term. Since the second term is not a function of the bandwidth, we leave it intact in order to have it canceled out later}
=& h^2\left[\int u^2k(u)du\right]\mathbb{E}\Bigg\{ \{\gamma(X,M,a) - \eta(a, a', X)\}\frac{1}{2}\frac{\sum_{j = 1}^{h_{d_A}} \partial^2_{a_j} f(a'|X,M)}{f(a'|X)}\Bigg\}+ o_p(h^2)\\
\end{align*}
Estimated Expectation Part 3: This term in the expectation is given as
\[
\mathbb{E}[\hat{\eta}(a, a^\prime, X_i)]
\]
And application of the continuous mapping theorem shows this converges to $\psi_0(a, a')$. Putting these together, the $\mathbb{E}[\hat{\psi}^{MR}]$ is given as
\begin{align*}
\mathbb{E}[\hat{\psi}^{MR}] = &\mathbb{E}_{O^c_{I_\ell}}\Bigg\{\int \frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\big[\gamma(X,M,a) - \hat{\gamma}(X,M,a)\big]f(a|X,M)f(X,M)dXdM\Bigg\} \\
& +h^2 \int u^2 k(u)du \\
&\times \mathbb{E}\bigg\{ \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)f(a \mid X)} \bigg(\sum^{d_A}_{j=1} \partial_{a_j} \gamma(X,M,a) \partial_{a_j} f(a|X,M) + \frac{1}{2} \left[\sum^{d_A}_{j=1} \partial^2_{a_j} \gamma(X,M,a)\right]f(a|X,M) \bigg)\bigg\} \\
&+ h^2\left[\int u^2k(u)du\right]\mathbb{E}\Bigg\{ \{\gamma(X,M,a) - \eta(a, a', X)\}\frac{1}{2}\frac{\sum_{j = 1}^{h_{d_A}} \partial^2_{a_j} f(a'|X,M)}{f(a'|X)}\Bigg\}\\
&+ \mathbb{E}\Bigg\{ \{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\}\frac{f(a'|X,M)}{\hat{f}(a'|X)}\Bigg\}\\
&+ \psi_0(a, a') + o_p(h^2)
\end{align*}
And this can be written as
\begin{align*}
\mathbb{E}[\hat{\psi}^{MR}] &= B(a, a') + \mathbb{E}_{O^c_{I_\ell}}\Bigg\{\int \frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\big[\gamma(X,M,a) - \hat{\gamma}(X,M,a)\big]f(a|X,M)f(X,M)dXdM\Bigg\} \\
& + \mathbb{E}\Bigg\{ \{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\}\frac{f(a'|X,M)}{\hat{f}(a'|X)}\Bigg\} + \psi_0(a, a') + o_p(h^2)
\end{align*}
Consequently, $\widehat{B}(a, a')$ obeys the identity in Colenaglo, but the $h^2$ needs to be taken out in order for this to work.
Next, we compute $var(\hat{\psi}^{MR}(a, a'))$
\subsubsection{Calculation of Estimated Variance}
\textbf{Estimated Variance Part 1}
\begin{align*}
&h^{d_A}\mathbb{E}\Bigg\{\bigg[\frac{K_h(A - a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\{Y - \hat{\gamma}(X,M,a)\} \bigg]^2\Bigg\}\\
=&h^{d_A}\mathbb{E}\Bigg\{\mathbb{E}\bigg\{\bigg[\frac{K_h(A - a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\{Y - \hat{\gamma}(X,M,a)\} \bigg]^2\bigg|X,M\bigg\}\Bigg\}\\
=&h^{d_A}\mathbb{E}\Bigg\{\frac{\hat{f}(M \mid A = a', X)^2}{\hat{f}(M \mid A = a, X)^2\hat{f}(a \mid X)^2}\mathbb{E}\bigg\{K_h(A - a)^2(Y - \hat{\gamma}(X,M,a))^2\bigg|X,M\bigg\}\Bigg\}\\
=&h^{d_A}\mathbb{E}\Bigg\{\frac{\hat{f}(M \mid A = a', X)^2}{\hat{f}(M \mid A = a, X)^2\hat{f}(a \mid X)^2}\times \mathbb{E}\bigg\{K_h(A - a)^2\mathbb{E}\{(Y - \hat{\gamma}(X,M,a))^2|X,M,A\}\bigg|X,M\bigg\}\Bigg\}\\
=&h^{d_A}\mathbb{E}\Bigg\{\frac{\hat{f}(M \mid A = a', X)^2}{\hat{f}(M \mid A = a, X)^2\hat{f}(a \mid X)^2}\\
&\times \mathbb{E}\bigg\{K_h(A - a)^2 \bigg[var(Y|X,M,A) + \gamma(X,M,A)^2 -2\gamma(X,M,A)\hat{\gamma}(X,M,a)+\hat{\gamma}(X,M,a)^2\bigg]
\bigg|X,M\bigg\}\Bigg\}\\
=&h^{d_A}\mathbb{E}\Bigg\{\frac{\hat{f}(M \mid A = a', X)^2}{\hat{f}(M \mid A = a, X)^2\hat{f}(a \mid X)^2}\mathbb{E}\bigg\{K_h(A - a)^2 \bigg[var(Y|X,M,A) + (\gamma(X,M,A)-\hat{\gamma}(X,M,a))^2\bigg]
\bigg|X,M\bigg\}\Bigg\}
\end{align*}
Because we assume $0 < \int u^6 k(u) du < \infty$, we also have boundedness of $\int u^6k^2(u)du$. The inner expectation can be written as
\begin{align*}
&h^{d_A}\mathbb{E}\bigg\{K_h(A - a)^2 \bigg[var(Y|X,M,A) + (\gamma(X,M,A)-\hat{\gamma}(X,M,a))^2\bigg]
\bigg|X,M\bigg\}\\
= &h^{d_A} \int \bigg[\prod^{d_A}_{j=1}\frac{1}{h^2}k\Big(\frac{A_j-a_j}{h}\Big)^2\bigg] \bigg[var(Y|X,M,A) + (\gamma(X,M,A)-\hat{\gamma}(X,M,a))^2\bigg]
f(A|X,M) dA\\
= &\int \bigg[\prod^{d_A}_{j=1} k(u_j)^2 \bigg]\times \bigg[var(Y|X,M,a+uh) + (\gamma(X,M,a+uh)-\hat{\gamma}(X,M,a))^2\bigg]
f(a+uh|X,M) du_1 \dots du_{d_A}\\
= & \int k(u_1)^2\cdots k(u_{d_A})^2 \\
&\times \bigg[var(Y|X,M,a)+\sum^{d_A}_{j=1}u_jh\partial_{a_j} var(Y|X,M,a) + \sum^{d_A}_{j=1}\sum^{d_A}_{j^\prime=1}u_ju_{j^\prime}h^2\partial_{a_j}\partial_{a_j^\prime} var(Y|X,M,\Bar{a}) + \\
&\bigg( \gamma(X,M,a)-\hat{\gamma}(X,M,a)+\sum^{d_A}_{j=1} u_jh\partial_{a_j} \gamma(X,M,\Bar{a}) \bigg)^2\bigg]\\
&\times \bigg(f(a|X,M)+\sum^{d_A}_{j=1}u_jh\partial_{a_j} f(a|X,M) + \sum^{d_A}_{j=1}\sum^{d_A}_{j^\prime=1}u_ju_{j^\prime}h^2\partial_{a_j}\partial_{a_{j^\prime}} f(\Bar{a}|X,M) \bigg)du_1\cdots du_{d_A}
\end{align*}
Expanding the terms,
\begin{align*}
&h^{d_A}\mathbb{E}\bigg\{K_h(A - a)^2 \bigg[var(Y|X,M,A) + (\gamma(X,M,A)-\hat{\gamma}(X,M,a))^2\bigg]
\bigg|X,M\bigg\}\\
= &\int\bigg[\prod^{d_A}_{j=1} k(u)^2 \bigg] var(Y|X,M,a)f(a|X,M)du_1\cdots du_{d_A} \\
&+\int\bigg[\prod^{d_A}_{j=1} k(u)^2 \bigg] \bigg( \gamma(X,M,a)-\hat{\gamma}(X,M,a)\bigg)^2f(a|X,M)du_1\cdots du_{d_A} \\
&+\int\bigg[\prod^{d_A}_{j=1} k(u)^2 \bigg] var(Y|X,M,a)\bigg( \sum^{d_A}_{j=1}u_jh\partial_{a_j} f(a|X,M) \bigg)du_1\cdots du_{d_A} \\
&+\int\bigg[\prod^{d_A}_{j=1} k(u)^2 \bigg] \bigg( \gamma(X,M,a)-\hat{\gamma}(X,M,a)\bigg)^2\bigg( \sum^{d_A}_{j=1}u_jh\partial_{a_j} f(a|X,M) \bigg)du_1\cdots du_{d_A} \\
&+\int\bigg[\prod^{d_A}_{j=1} k(u)^2 \bigg]
\bigg[\sum^{d_A}_{j=1}u_jh\partial_{a_j} var(Y|X,M,a)\bigg]f(a|X,M)du_1\cdots du_{d_A}\\
&+\int\bigg[\prod^{d_A}_{j=1} k(u)^2 \bigg] 2\bigg( \gamma(X,M,a)-\hat{\gamma}(X,M,a)\bigg)\bigg(\sum^{d_A}_{j=1} u_jh\partial_{a_j} \gamma(X,M,\Bar{a})\bigg)f(a|X,M)du_1\cdots du_{d_A} \\
&+\int\bigg[\prod^{d_A}_{j=1} k(u)^2 \bigg] \bigg(\sum^{d_A}_{j=1} u_jh\partial_{a_j} \gamma(X,M,\Bar{a})\bigg)^2f(a|X,M)du_1\cdots du_{d_A} \\
&+\int\bigg[\prod^{d_A}_{j=1} k(u)^2 \bigg] \bigg[var(Y|X,M,a)+\sum^{d_A}_{j=1}u_jh\partial_{a_j} var(Y|X,M,a) + \sum^{d_A}_{j=1}\sum^{d_A}_{j^\prime=1}u_ju_{j^\prime}h^2\partial_{a_j}\partial_{a_j^\prime} var(Y|X,M,\Bar{a}) + \\
&\bigg( \gamma(X,M,a)-\hat{\gamma}(X,M,a)+\sum^{d_A}_{j=1} u_jh\partial_{a_j} \gamma(X,M,\Bar{a}) \bigg)^2\bigg]\bigg( \sum^{d_A}_{j=1}\sum^{d_A}_{j^\prime=1}u_ju_{j^\prime}h^2\partial_{a_j}\partial_{a_{j^\prime}} f(\Bar{a}|X,M) \bigg)du_1\cdots du_{d_A} \\
&+\int\bigg[\prod^{d_A}_{j=1} k(u)^2 \bigg] \bigg[\sum^{d_A}_{j=1}u_jh\partial_{a_j} var(Y|X,M,a) + \sum^{d_A}_{j=1}\sum^{d_A}_{j^\prime=1}u_ju_{j^\prime}h^2\partial_{a_j}\partial_{a_j^\prime} var(Y|X,M,\Bar{a}) + \\
&\bigg( \sum^{d_A}_{j=1} u_jh\partial_{a_j} \gamma(X,M,\Bar{a}) \bigg)^2
+2\bigg( \sum^{d_A}_{j=1} u_jh\partial_{a_j} \gamma(X,M,\Bar{a}) \bigg)
\bigg( \gamma(X,M,a)-\hat{\gamma}(X,M,a) \bigg)\bigg]\bigg( \sum^{d_A}_{j=1}u_jh\partial_{a_j} f(a|X,M) \bigg)du_1\cdots du_{d_A} \\
&+\int\bigg[\prod^{d_A}_{j=1} k(u)^2 \bigg] \bigg( \sum^{d_A}_{j=1}\sum^{d_A}_{j^\prime=1}u_ju_{j^\prime}h^2\partial_{a_j}\partial_{a_j^\prime} var(Y|X,M,\Bar{a}) \bigg)f(a|X,M)du_1\cdots du_{d_A}
\end{align*}
Note that by the properties of the kernel function, terms 3 to 6 are zero and also, terms 7 to 10 are $O(h^2)$. Thus,
\begin{align*}
&h^{d_A}\mathbb{E}\bigg\{K_h(A - a)^2 \bigg[var(Y|X,M,A) + (\gamma(X,M,A)-\hat{\gamma}(X,M,a))^2\bigg]
\bigg|X,M\bigg\}\\
&= \Big\{\int \bigg[\prod^{d_A}_{j=1} k(u)^2 \bigg] du_1\dots du_j\Big\} \times \bigg(var(Y|X,M,a)+\bigg( \gamma(X,M,a)-\hat{\gamma}(X,M,a)\bigg)^2\bigg)f(a|X,M) + O(h^2).
\end{align*}
\textbf{Estimated Variance Part 2}:
\begin{align*}
&h^{d_A}\mathbb{E}\Bigg\{\bigg[\frac{K_h(A - a')}{\hat{f}(a' \mid X)}\Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big) \bigg]^2\Bigg\}\\
=&h^{d_A}\mathbb{E}\Bigg\{\frac{1}{\hat{f}(a' \mid X)^2}\mathbb{E}\bigg[K_h(A - a')^2\Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big)^2\Big|X\bigg]\Bigg\}\\
\end{align*}
The inner expectation can be written as
\begin{align*}
&h^{d_A}\mathbb{E}\bigg[K_h(A - a')^2\Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big)^2\Big|X\bigg]\\
=& h^{d_A}\int\int \bigg[\prod^{d_A}_{j=1}\frac{1}{h^2}k\Big(\frac{A_j-a'}{h}\Big)^2\bigg] \Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big)^2 f(A| M,X)f(M\mid X) dAdM \\
=&\int\int \bigg[\prod^{d_A}_{j=1} k(u_j)^2 \bigg]\Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big)^2 f(a'+uh|M, X)f(M\mid X) dudM\\
=&\int\int \bigg[\prod^{d_A}_{j=1} k(u_j)^2 \bigg]\Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big)^2 \Big\{f(a'|M, X) + \sum_{j = 1}^{d_A}u_jh\partial_{a_j}f(a \mid X, M)\\
&+ \sum_{j = 1}^{d_A}\sum_{j^\prime = 1}^{d_A}u_ju_{j^\prime}h^2\partial_{a_j}\partial_{a_{j^\prime}}f(\Bar{a} \mid X, M)\Big\}f(M\mid X) dudM\\
=& \int\int k^2(u_1)\cdots k^2(u_{d_A})\Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big)^2f(a'|X, M)f(M\mid X)
du_1\cdots du_{d_A}dM + O(h^2)\\
\end{align*}
And an application of the continuous mapping theorem gives
\begin{align*}
&h^{d_A}\mathbb{E}\Bigg\{\frac{1}{f(a' \mid X)^2}\mathbb{E}\bigg[K_h(A - a')^2\Big(\gamma(X,M,a) - \eta(a, a', X)\Big)^2\Big|X\bigg]\Bigg\}\\
=& \int k(u)^2 du \times \mathbb{E}\bigg\{\frac{1}{f(a'|X)}var[E(Y|X,M,a)|X,a']\bigg\} + o_p(1)
\end{align*}
\textbf{Estimated Variance Part 3}
\begin{align*}
h^{d_A}\mathbb{E}\Bigg\{\hat{\eta}^2(a, a', X)\Bigg\} &= O(h^{d_A})
\end{align*}
This holds because we assume the nuisance estimators are bounded, hence $\hat{\eta}$ is bounded.
\textbf{Estimated Variance Part 4}
\begin{align*}
&h^{d_A}\mathbb{E}\Bigg\{\bigg[\frac{K_h(A - a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\{Y - \hat{\gamma}(X,M,a)\} \bigg]\\
&\qquad \times \bigg[\frac{K_h(A - a')}{\hat{f}(a' \mid X)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\} \bigg]\Bigg\}\\
=&h^{d_A}\mathbb{E}\Bigg\{\frac{K_h(A-a)K_h(A-a')}{\hat{f}(a|X)\hat{f}(a'|X)}\frac{\hat{f}(M|a',X)}{\hat{f}(M|a,X)}\Big[Y-\hat{\gamma}(X,M,a)\Big]\Big[\hat{\gamma}(X,M,a)-\hat{\eta}(a,a',X)\Big]\Bigg\}\\
=&h^{d_A}\mathbb{E}\Bigg\{\frac{1}{\hat{f}(a|X)\hat{f}(a'|X)}\frac{\hat{f}(M|a',X)}{\hat{f}(M|a,X)} \Big[\hat{\gamma}(X,M,a)-\hat{\eta}(a,a',X)\Big]\\
& \qquad\times\mathbb{E}\Big\{K_h(A-a)K_h(A-a')\Big[Y-\hat{\gamma}(X,M,a)\Big]\Big|X,M\Big\}\Bigg\}\\
=&h^{d_A}\mathbb{E}\Bigg\{\frac{1}{\hat{f}(a|X)\hat{f}(a'|X)}\frac{\hat{f}(M|a',X)}{\hat{f}(M|a,X)}
\Big[\hat{\gamma(X,M,a)} - \hat{\eta}(a,a',X)\Big]\\
& \qquad\times\mathbb{E}\Big\{K_h(A-a)K_h(A-a')\Big[\gamma(X,M,A)-\hat{\gamma}(X,M,a)\Big]\Big|X,M\Big\}\Bigg\}
\end{align*}
The inner expectation
\begin{align*}
&h^{d_A}\mathbb{E}\Big\{K_h(A-a)K_h(A-a')\Big[\gamma(X,M,A)-\hat{\gamma}(X,M,a)\Big]\Big|X,M\Big\}\\
=& h^{d_A}\int \bigg[\prod^{d_A}_{j=1}\frac{1}{h^2}k\Big(\frac{A_j-a}{h}\Big)k\Big(\frac{A_j-a'}{h}\Big)\bigg]
\Big[\gamma(X,M,A)-\gamma(X,M,a)\Big] f(A|X,M)dA\\
=& \int k(u_1)\cdots k(u_{d_A}) k(u_1+\frac{a-a'}{h})\cdots k(u_{d_A}+\frac{a-a'}{h})\Big[\gamma(X,M, uh + a)-\hat{\gamma}(X,M,a)\Big] f(uh+ a|X,M)dA\\
=& \int k(u_1)\cdots k(u_{d_A}) k(u_1+\frac{a-a'}{h})\cdots k(u_{d_A}+\frac{a-a'}{h})\\ &\times \Big[(\gamma(X,M,a) - \hat{\gamma}(X,M,a)) + \sum^{d_A}_{j=1} u_j h\partial_{a_j} \gamma(X,M,a)+\frac{u^2_j h^2}{2}\partial^2_{a_j}\gamma(X,M,a)+\frac{u^3_jh^3}{6}\partial^3_{a_j}\gamma(X,M,\Bar{a})\Big]\\
&\times \Big[f(a|X,M) +\sum^{d_A}_{j=1} u_jh\partial_{a_j} f(a|X,M)+\frac{u^2_jh^2}{2}\partial^2_{a_j} f(\Bar{a}|X,M)\Big]du_1\cdots du_{d_A}\\
&= o_P(1)
\end{align*}
\textbf{Estimated Variance Part 5}
\begin{align*}
&2h^{d_A}\mathbb{E}\Bigg\{\hat{\eta}(a, a', X)\bigg[\frac{K_h(A - a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\{Y - \hat{\gamma}(X,M,a)\} \bigg]\Bigg\}\\
=& 2h^{d_A}\mathbb{E}\Bigg\{\frac{\hat{\eta}(a, a', X)K_h(A-a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\mathbb{E}\bigg[Y - \hat{\gamma}(X,M,a) \mid X, M \bigg]\Bigg\}\\
=& 2h^{d_A}\mathbb{E}\Bigg\{\mathbb{E}\bigg\{\frac{\hat{\eta}(a, a', X)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)} \mathbb{E}\bigg[K_h(A - a) (\gamma(X,M,A) - \hat{\gamma}(X,M,a))\bigg| X,M, O^c_{I_\ell}\bigg]\bigg|O^c_{I_\ell}\bigg\}\Bigg\}.
\end{align*}
Similar to the derivation in Estimated Expectation Part 1, by $\hat{\eta}(a, a', X)\rightarrow_p\eta(a, a', X)$, there is
\begin{align*}
&\mathbb{E}\Bigg\{\mathbb{E}\bigg\{\frac{\hat{\eta}(a, a', X)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)} \mathbb{E}\bigg[K_h(A - a) (\gamma(X,M,A) - \hat{\gamma}(X,M,a))\bigg| X,M, O^c_{I_\ell}\bigg]\bigg|O^c_{I_\ell}\bigg\}\Bigg\} \\
=&\mathbb{E}_{O^c_{I_\ell}}\Bigg\{\int \frac{\hat{\eta}(a, a', X)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\big[\gamma(X,M,a) - \hat{\gamma}(X,M,a)\big]f(a|X,M)f(X,M)dXdM\Bigg\} \\
& +h^2 \int u^2 k(u)du \\
&\times \mathbb{E}\bigg\{ \frac{\eta(a, a', X)f(M \mid A = a', X)}{f(M \mid A = a, X)f(a \mid X)} \bigg(\sum^{d_A}_{j=1} \partial_{a_j} \gamma(X,M,a) \partial_{a_j} f(a|X,M) + \frac{1}{2} \left[\sum^{d_A}_{j=1} \partial^2_{a_j} \gamma(X,M,a)\right]f(a|X,M) \bigg)\bigg\}\\
&+ o_p(h^2),
\end{align*}
By the boundedness of $\frac{\hat{\eta}(a, a', X)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}f(a|X,M)$ and the consistency of $\hat{\gamma}(X,M,a)$, the first term satisfies
$\mathbb{E}_{O^c_{I_\ell}}\Bigg\{\int \frac{\hat{\eta}(a, a', X)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\big[\gamma(X,M,a) - \hat{\gamma}(X,M,a)\big]f(a|X,M)f(X,M)dXdM\Bigg\} = o_p(1)$.
Inserting this back into the full expectation, combined with the boundedness of $\eta$, $f(M \mid a, X)$ and $f(a \mid X)$, we get
\begin{align*}
2h^{d_A}\mathbb{E}\Bigg\{\hat{\eta}(a, a', X)\bigg[\frac{K_h(A - a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\{Y - \hat{\gamma}(X,M,a)\} \bigg]\Bigg\} = O_p(h^{d_A+2} )+o_p(1)
\end{align*}
\textbf{Estimated Variance Part 6}
\begin{align*}
&2h^{d_A}\mathbb{E}\Bigg\{\eta(a, a', X)
\bigg[\frac{K_h(A - a')}{\hat{f}(a' \mid X)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\} \bigg]\Bigg\}\\
=2h^{d_A}&\mathbb{E}\bigg\{\mathbb{E}\bigg[\frac{K_h(A - a')\hat{\eta}(a, a^\prime, X)}{\hat{f}(a' \mid X)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\}\bigg| X,M\bigg]\bigg\} \\
=2h^{d_A}&\mathbb{E}\bigg\{\frac{\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\}\hat{\eta}(a, a^\prime, X)}{\hat{f}(a' \mid X)}\mathbb{E}[K_h(A - a')| X,M]\bigg\}
\end{align*}
Using the expansion from Part 2 of the expectation on $\mathbb{E}[K_h(A - a')| X,M]$, we get
\begin{align*}
\mathbb{E}[K_h(A - a')| X,M] =&f(a'|X,M) + \frac{1}{2}h^2\int u^2k(u)du \sum^{d_A}_{j=1}\partial^2_{a_j} f(a'|X,M) + O(h^3)
\end{align*}
Plugging this back into the full expectation, we get
\begin{align*}
2h^{d_A}\mathbb{E}\Bigg\{\eta(a, a', X)
\bigg[\frac{K_h(A - a')}{f(a' \mid X)}\{\mathbb{E}[Y \mid X, M, A = a] - \eta(a, a', X)\} \bigg]\Bigg\} = O(h^{d_A})
\end{align*}
And using the calculation for the bias,
{
\begin{align*}
&h^{d_A}\mathbb{E}\Bigg\{\bigg[\frac{K_h(A - a)f(M \mid A = a', X)}{f(M \mid A = a, X)f(a \mid X)}\{Y - \mathbb{E}[Y \mid X, M , A = a]\} + \frac{K_h(A - a')}{f(a' \mid X)}\{\mathbb{E}[Y \mid X, M, A = a] - \eta(a, a', X)\} + \eta(a, a', X)\bigg]\Bigg\}^2\\
&= O(\frac{1}{nh^{d_A}})\\
\end{align*}
}
Next, under the product of kernels being bounded, the covariance will be of the same order as the variance. Then following a similar proof to Coleangelo, this will prove consistency.
Proof for Multiple Robustness
lemmaUnder Assumptions 1, 2, 3, and 6, the proposed $\hat{\psi}^{MR}(a, a')$ will be a consistent estimator for $\psi(a, a')$ as long as any two out of three conditions in Assumption 4 hold.
Following a similar breakdown as that in Theorem 1, $\hat{\psi}^{MR}(a, a') - \psi(a, a')$ can be expanded as
align*[align* omitted — 439 chars of source]
Following the result in Theorem 1 on asymptotic normality by application of the Lyapunov CLT, the term $\sqrt{\frac{h^{d_A}}{n}}\sum^L_{\ell=1}\sum_{i\in I_\ell} \Big\{ m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\Big\} = O_p(1)$. Since $\frac{1}{\sqrt{nh^{d_A}}} = o_p(1)$, the following holds
\[\frac{1}{\sqrt{nh^{d_A}}} \times \sqrt{\frac{h^{d_A}}{n}}\sum^L_{\ell=1}\sum_{i\in I_\ell} \Big\{ m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\Big\} = o_p(1).
\]
The remainder of the proof demonstrates the remaining the remaining terms are $o_p(1)$, i.e.
\[
\frac{1}{\sqrt{nh^{d_A}}} \times \sqrt{\frac{h^{d_A}}{n}}\sum^L_{\ell=1}\sum_{i\in I_\ell}\Big\{m(O_i;\hat{\alpha}_\ell,\hat{\lambda}_\ell,\hat{\gamma}_\ell,\psi_0(a,a'))- m(O_i;\alpha,\lambda,\gamma,\psi_0(a,a'))\Big\}.
\]
To see this, we first expand these terms identically as the proof of Theorem 1, and provide proofs for the convergence of the terms (CS1) - (CS6), (E1) - (E8) and (TR1) - (TR5) under the assumption that any two out of three nuisance models are correctly specified in the following sub-sections.
Proof for Terms (CS1)-(CS6)
All of these terms contain the product of two or more errors and can be treated similarly. We provide a detailed proof for (CS2), and a similar method can be followed for the rest of the terms.
For (CS2), write $\Delta_{i\ell}=K_h(A_i - a)\big[\hat{R}(M_i,X_i)-R(M_i,X_i)\big]\big[\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big]\big[Y_i - \gamma(X_i,M_i,a)\big]$. Following Lemma (ref), it suffices to bound $\mathbb{E}\left[|\frac{1}{\sqrt{nh^{d_A}}} \times \sqrt{\frac{h^{d_A}}{n}} \sum_{i \in I_\ell} \Delta_{i\ell} | \Big| O^c_{I_\ell}\right]$ as $o_p(1)$ in order to show that $$\frac{1}{\sqrt{nh^{d_A}}} \times \sqrt{\frac{h^{d_A}}{n}} \sum_{i \in I_\ell}\Delta_{i\ell} = o_p(1).$$
First, from the triangle inequality, $\mathbb{E}\left[\left|\frac{1}{\sqrt{nh^{d_A}}} \times \sqrt{\frac{h^{d_A}}{n}} \sum_{i \in I_\ell} \Delta_{i\ell} \right| \mid O^c_{I_\ell}\right] \leq \frac{1}{L}\mathbb{E}\left[\left|\Delta_{i\ell} \right| \mid O^c_{I_\ell}\right]$, and so it suffices to bound $\mathbb{E} \bigg[\big|\Delta_{i\ell} \big|\bigg| O^c_{I_\ell}\bigg]$.
{
align*[align* omitted — 825 chars of source]
}
Next, Assumption 3.1 on the boundedness of $\gamma(X,M,a)$ and Assumption 3.3 on the boundedness of $var(Y_i|a,m,x)$, along with an application of Lemma (ref) on $f(a \mid M, X)$, we get
{
align*[align* omitted — 760 chars of source]
}
As long as either Assumption 4.2 or Assumption 4.3 hold, then combined with Assumption 3.2,
(CS2) will be $o_p(1)$. A similar approach can be used to bound the remaining CS terms.
Proof for Terms (E1)-(E8)
Terms (E1)-(E8) are normalized terms of the form of a bias times a bounded quantity; they can all be treated similarly. We only provide the proof of the convergence in probability to zero for the term (E2). (E2) is given as
align*[align* omitted — 260 chars of source]
To prove this, we set $\hat{\Delta}_{i\ell}$ as (E2). By construction, $O^c_{I_\ell}$ and $O_i$ are independent, $i\in I_\ell$, and consequently $\mathbb{E}\left[\hat{\Delta}_{i\ell}|O^c_{I_\ell}\right]=0$ and $\mathbb{E}\left[\hat{\Delta}_{i\ell}\hat{\Delta}_{j\ell}|O^c_{I_\ell}\right]=0$ for $i,j \in I_\ell$ and all $a', a\in \mathcal{A}_0$. Next we note that
align*[align* omitted — 1,041 chars of source]
Where (a) follows from Assumption 3.1 on the boundedness of $f(a \mid M, X)$, along with Assumption 3.1 and Assumption 3.3 combined with the derivation provided below
align*[align* omitted — 507 chars of source]
Next, (b) follows from Assumption 2.4, and finally, (c) follows Assumption 3.2 along with Assumption 4.1. Then
$\mathbb{E} \bigg[\left(\frac{1}{\sqrt{nh^{d_A}}} \times \sqrt{h^{d_A}/n} \sum_{i\in I_\ell} \hat{\Delta}_{i\ell}\right)^2\bigg| O^c_{I_\ell}\bigg] =\frac{1}{n^2} \sum_{i\in I_\ell} \mathbb{E}\left[\hat{\Delta}^2_{i\ell}|O^c_{I_\ell}\right] = O(\frac{1}{n})\mathbb{E}\left[\hat{\Delta}^2_{i\ell}|O^c_{I_\ell}\right] = O_p(\frac{1}{nh^{d_A}}) = o_p(1)$.
Applying Lemma 1 to the above gives $\frac{1}{\sqrt{nh^{d_A}}} \times \sqrt{h^{d_A}/n} \sum^L_{l=1}\sum_{i\in I_\ell} \hat{\Delta}_{i\ell}\xrightarrow[]{P} 0$.
Proof for Terms (TR1)-(TR5)
The proofs of the convergence in probability to zero for the terms (TR1)-(TR5) follows a similar outline as Theorem 1, and we prove convergence on a case by case below.
Proof for Terms TR1 and TR2
Terms (TR1) and (TR2) are similar; we only provide the proof of the convergence in probability to zero for the term (TR2).
To bound TR2, first set $\hat{\Delta}_{i\ell}=K_h(A_i - a) \left[\hat{\lambda}(a, X_i) - \lambda(a,X_i) \right]R(M_i, X_i)\left[Y_i - \gamma(X_i,M_i,a)\right]$. Bounding (TR2) amounts to showing $\mathbb{E}[\hat{\Delta}_{i\ell}|O^c_{I_\ell}] = o_p(1)$.
align*[align* omitted — 977 chars of source]
where the equalities follow identically as in the proof of Theorem 1, and the final equality follows from $h \rightarrow 0$ along with the boundedness assumptions in Assumption 3.
Proof for TR3
For Term (TR3), we have
align*[align* omitted — 474 chars of source]
where (b) follows from the definition of $\lambda_{a^\prime}(X_i)$, $h\rightarrow 0$, Assumption 3 (boundedness of $\lambda$, $\hat{\eta}$ and $\eta$).
Proof For TR4
Demonstrating the bound for (TR4), we have
align*[align* omitted — 451 chars of source]
TR-4-1 can be written as
align*[align* omitted — 296 chars of source]
An application of Lemma (ref) to TR-4-1 gives
align*[align* omitted — 364 chars of source]
A similar approach applied to TR-4-2 gives
align*[align* omitted — 382 chars of source]
Now, the first terms of TR-4-1 and TR-4-2 cancel out with each other, with an identical proof to that used in the proof of Theorem 1.
Consequently only the remaining terms must be bounded.
\[
\int \big(\hat{\gamma}(X_i,M_i,a)-\gamma(X_i,M_i,a)\big)\lambda(a', X_i) O(h^2)f(M, X) dM_i dX_i = o_p(1)
\]
The second term in TR-4-1 and TR-4-2 can be bounded from $h \rightarrow 0$, combined with the boundedness assumptions in Assumption 3.2.
Proof For TR5
Finally, for term (TR5), we note that
align*[align* omitted — 1,074 chars of source]
Where $(a)$ follows from an application of Lemma (ref), (b) follows from the definition of $\eta$, and (c) follows from an application of Cauchy-Schwartz combined with the consistency of $\hat{\lambda}$.
Proof for Consistency of $\widehat{V}(a, a')$
commentAssumptions:
\begin{enumerate}
• $\mathbb{E}\{[Y-\gamma(X,M,a)]^4 | A = a', M = m, X = x\}$ is bounded uniformly over any $(a,a',m,x)\in \mathcal{A}^2\times \mathcal{M} \times \mathcal{X}$.
• $\int^{\infty}_{-\infty} k(u)^{c_1}k(u+\tilde{c})^{c_2} du < \infty$ and $\int^{\infty}_{-\infty} u^2k(u)^{c_1}k(u+\tilde{c})^{c_2} du < \infty$ for $\tilde{c} \in \mathcal{R}$ and $c_1+c_2 \in \{2,3,4\}$ for $c_1,c_2\in\{0,1,2,3,4\}$
• The following hold:
\[
\int (\gamma(X,M,a) - \hat{\gamma}(X,M,a))^2\Big(\frac{1}{\hat{f}(a \mid X)} - \frac{1}{f(a \mid X)}\Big)^2 f(X, M) dX dM \rightarrow_P 0
\]
\[
\int \Big(\frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)} - \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)}\Big)^2(\gamma(X,M,a) - \hat{\gamma}(X,M,a))^2f(X, M) dM dX \rightarrow_P 0
\]
\end{enumerate}
Recall that
equation*[equation* omitted — 237 chars of source]
To prove consistency of $\widehat{V}(a, a')$, we first prove propositions (I), (II) and (III), which together prove the desired result.
(I) $h^{d_A}n^{-1} \sum_{i\in I_\ell} m^2(O_i;\alpha,\lambda,\gamma,\psi(a,a')) - V(a,a') = o_p(1)$
To simplify notation, denote $m(O_i;\alpha,\lambda,\gamma,\psi(a,a'))$ as $m_i(a,a')$. From the proof of Theorem 1, we have $h^{d_A} \mathbb{E}[m^2_i(a,a')] = V(a,a') + o_p(1)$.
We write
align*[align* omitted — 207 chars of source]
Then,
equation*[equation* omitted — 512 chars of source]
We only need to investigate the terms $\mathbb{E}(U_1^{c_1} U_2^{c_2}U_3^{c_3})$ for any $c_1 \ge 0$, $c_2 \ge 0$ and $c_3 \ge 0$ with $c_1 + c_2 + c_3 = 4$. To be specific, dropping the terms with power index being zero, we will be studying $\mathbb{E}(U_1^{c_1})$, $\mathbb{E}(U_2^{c_2})$, $\mathbb{E}(U_1^{c_1} U_2^{c_2})$, $\mathbb{E}(U_1^{c_1} U_3^{c_3})$, $\mathbb{E}(U_2^{c_2} U_3^{c_3})$, and $\mathbb{E}(U_1^{c_1} U_2^{c_2}U_3^{c_3})$ for positive $c_1, c_2,$ and $c_3$.
enumerate• $\mathbb{E}(U_1^{c_1})$: By the assumed boundedness of $\lambda(a,X)$, $R(M,X)$, and $\mathbb{E}\{[Y-\gamma(X,M,a)]^4 | A = a', M = m, X = x\}$ over any $(a,a',m,x)\in \mathcal{A}\times\mathcal{A}\times \mathcal{M} \times \mathcal{X}$ from Assumption 7,
\begin{align*}
&\mathbb{E}(U_1^{c_1}) = \int \bigg\{K_h(A - a)\lambda(a,X)R(M,X)[Y - \gamma(X,M,a)]\bigg\}^{c_1} f(Y,A, M,X) dO\\
=& O(\frac{1}{h^{(c_1-1)d_A}})\int \tilde{k}(u)^{c_1} \bigg\{\int |Y - \gamma(X,M,a)|^{c_1} f(Y|A = uh+a, M, X)dY \bigg\}f(uh+a,M,X) dudMdX\\
=& O(\frac{1}{h^{(c_1-1)d_A}})\int \tilde{k}(u)^{c_1} \mathbb{E}\{ |Y - \gamma(X,M,a)|^{c_1} |A = uh+a, M, X\}f(uh+a,M,X) dudMdX\\
=& O(\frac{1}{h^{(c_1-1)d_A}})\int \tilde{k}(u)^{c_1} f_{MX}(M,X)\\ &\bigg\{f(a|M,X)+ \sum^{d_A}_{j=1}u_jh\frac{\partial}{\partial a} f(a|M,X) + \sum^{d_A}_{j=1}\sum^{d_A}_{j'=1}u_ju_{j'}h^2\frac{\partial^2}{\partial a_j\partial a_{j'}} f(\bar{a}|M,X) \bigg\}dudMdX\\
=& O(\frac{1}{h^{(c_1-1)d_A}})\int \tilde{k}(u)^{c_1} du + o(\frac{1}{h^{(c_1-1)d_A}}) = O(\frac{1}{h^{(c_1-1)d_A}}).
\end{align*}
where $\bar{a}$ is between $a$ and $a+uh$.
• $\mathbb{E}(U_2^{c_2})$: From the boundedness of $\lambda(a',X)$, $\gamma(X,M,a)$ and $\eta(a, a', X)$ over any $(a,a',a'',m,x)\in \mathcal{A}^3\times \mathcal{M} \times \mathcal{X}$,
\begin{align*}
&\mathbb{E}(U_2^{c_2}) = \int \bigg\{K_h(A - a')\lambda(a',X)[\gamma(X,M,a) - \eta(a, a', X)]\bigg\}^{c_2} f(A, M,X) dO\\
=& O(\frac{1}{h^{(c_2-1)d_A}})\int \tilde{k}(u)^{c_2}f_{MX}(M,X) f(uh+a'|M,X) dudMdX\\
=& O(\frac{1}{h^{(c_2-1)d_A}})\int \tilde{k}(u)^{c_2}f_{MX}(M,X) \\
&\bigg\{f(a'|M,X)+ \sum^{d_A}_{j=1}u_jh\frac{\partial}{\partial a'} f(a'|M,X) + \sum^{d_A}_{j=1}\sum^{d_A}_{j'=1}u_ju_{j'}h^2\frac{\partial^2}{\partial a'_j\partial a'_{j'}} f(\bar{a}|M,X) \bigg\}dudMdX\\
=& O(\frac{1}{h^{(c_2-1)d_A}})\int \tilde{k}(u)^{c_2} du + o(\frac{1}{h^{(c_2-1)d_A}}) = O(\frac{1}{h^{(c_2-1)d_A}}).
\end{align*}
where $\bar{a}$ is between $a'$ and $a'+uh$.
• $\mathbb{E}(U_1^{c_1} U_2^{c_2})$
\begin{align*}
&\mathbb{E}(U_1^{c_1} U_2^{c_2})\\
=& \int \bigg\{K_h(A - a)\lambda(a,X)\frac{\alpha(a',M,X)}{\alpha(a,M,X)}[Y - \gamma(X,M,a)]\bigg\}^{c_1}\\
&\bigg\{K_h(A - a')\lambda(a',X)[\gamma(X,M,a) - \eta(a, a', X)]\bigg\}^{c_2} f(Y,A, M,X) dO\\
=& O(\frac{1}{h^{(c_1+c_2-1)d_A}})\int \tilde{k}(u)^{c_1} \bigg\{\int |Y - \gamma(X,M,a)|^{c_1} f(Y|A = uh+a, M, X)dY \bigg\}\\
&\tilde{k}(u+\frac{a-a'}{h})^{c_2} f_{MX}(M, X) f(uh+a|M,X) dudMdX \\
=& O(\frac{1}{h^{(c_1+c_2-1)d_A}})\int \bigg[\prod^{d_A}_{j=1} k(u_j)^{c_1} k(u_j + \frac{a_j-a'_j}{h})^{c_2}\bigg] \mathbb{E}\{ |Y - \gamma(X,M,a)|^{c_1} |A = uh+a, M, X\}\\
& f_{MX}(M, X) f(uh+a|M,X) dudMdX \\
=& O(\frac{1}{h^{(c_1+c_2-1)d_A}})\int \bigg[\prod^{d_A}_{j=1} k(u_j)^{c_1} k(u_j + \frac{a_j-a'_j}{h})^{c_2}\bigg] f_{MX}(M, X) \\
&\bigg\{f(a|M,X)+ \sum^{d_A}_{j=1}u_jh\frac{\partial}{\partial a} f(a|M,X) + \sum^{d_A}_{j=1}\sum^{d_A}_{j'=1}u_ju_{j'}h^2\frac{\partial^2}{\partial a_j\partial a_{j'}} f(\bar{a}|M,X) \bigg\}dudMdX\\
=& O(\frac{1}{h^{(c_1+c_2-1)d_A}})\int \bigg[\prod^{d_A}_{j=1} k(u_j)^{c_1} k(u_j + \frac{a_j-a'_j}{h})^{c_2}\bigg] du + o(\frac{1}{h^{(c_1+c_2-1)d_A}})\\
=& O(\frac{1}{h^{(c_1+c_2-1)d_A}}).
\end{align*}
where $\bar{a}$ is between $a$ and $a+uh$.
• $\mathbb{E}(U_1^{c_1} U_3^{c_3})$
\begin{align*}
&\mathbb{E}(U_1^{c_1} U_3^{c_3})\\
=& \int \bigg\{K_h(A - a)\lambda(a,X)\frac{\alpha(a',M,X)}{\alpha(a,M,X)}[Y - \gamma(X,M,a)]\bigg\}^{c_1}\bigg\{\eta(a, a', X) - \psi(a,a') \bigg\}^{c_2} f(Y,A, M,X) dO\\
=&O(1) \int K_h(A - a)^{c_1}|Y - \gamma(X,M,a)|^{c_1}f(Y,A, M,X) dO\\
=&O(1) \int K_h(A - a)^{c_1}\mathbb{E}\bigg[|Y - \gamma(X,M,a)|^{c_1} \bigg| A, M, X \bigg]f(A, M,X) dO\\
=&O(1) \int K_h(A - a)^{c_1}f(A \mid M,X) dA f_{MX}(M, X) dMdX\\
=&O(\frac{1}{h^{(c_1 - 1)d_A}}) \int \tilde{k}(u)^{c_1}f_{MX}(M, X)\\
&\bigg\{f(a|M,X)+ \sum^{d_A}_{j=1}u_jh\frac{\partial}{\partial a} f(a|M,X) + \sum^{d_A}_{j=1}\sum^{d_A}_{j'=1}u_ju_{j'}h^2\frac{\partial^2}{\partial a_j\partial a_{j'}} f(\bar{a}|M,X) \bigg\}dudMdX\\
=& O(\frac{1}{h^{(c_1 - 1)d_A}})
\end{align*}
where $\bar{a}$ is between $a$ and $a+uh$, the second equality is from from the boundedness of $\lambda$, $\eta$, $\alpha$ and $\psi$, and the fourth equality comes from the assumed boundedness of $\mathbb{E}[|Y - \gamma|^4 \mid A, M, X]$.
• $\mathbb{E}(U_2^{c_2} U_3^{c_3})$
\begin{align*}
&\mathbb{E}(U_2^{c_2} U_3^{c_3})\\
=& \int \bigg\{K_h(A - a')\lambda(a',X)[\gamma(X,M,a) - \eta(a, a', X)]\bigg\}^{c_2}\bigg\{\eta(a, a', X) - \psi(a,a')\bigg\}^{c_3} f(Y,A, M,X) dO\\
=& O(\frac{1}{h^{(c_2-1)d_A}})\int \tilde{k}(u+\frac{a-a'}{h})^{c_2} f_{MX}(M, X) f_{A|X}(uh+a|M,X) dudMdX \\
=& O(\frac{1}{h^{(c_2-1)d_A}})\int \tilde{k}(u+\frac{a-a'}{h})^{c_2} f_{MX}(M, X) \\
&\bigg\{f(a'|M,X)+ \sum^{d_A}_{j=1}u_jh\frac{\partial}{\partial a'} f(a'|M,X) + \sum^{d_A}_{j=1}\sum^{d_A}_{j'=1}u_ju_{j'}h^2\frac{\partial^2}{\partial a'_j\partial a'_{j'}} f(\bar{a}|M,X) \bigg\}dudMdX\\
=& O(\frac{1}{h^{(c_2-1)d_A}})\int \tilde{k}(u+\frac{a-a'}{h})^{c_2} du + o(\frac{1}{h^{(c_2-1)d_A}})\\
=& O(\frac{1}{h^{(c_2-1)d_A}}).
\end{align*}
where $\bar{a}$ is between $a'$ and $a'+uh$.
• $\mathbb{E}(U_1^{c_1} U_2^{c_2} U_3^{c_3} )$
\begin{align*}
&\mathbb{E}(U_1^{c_1} U_2^{c_2}U_3^{c_3})\\
=& \int \bigg\{K_h(A - a)\lambda(a,X)\frac{\alpha(a',M,X)}{\alpha(a,M,X)}[Y - \gamma(X,M,a)]\bigg\}^{c_1}\\
&\bigg\{K_h(A - a')\lambda(a',X)[\gamma(X,M,a) - \eta(a, a', X)]\bigg\}^{c_2}\bigg\{\eta(a, a', X) - \psi(a,a')\bigg\}^{c_3} f(Y,A, M,X) dO\\
=& O(\frac{1}{h^{(c_1+c_2-1)d_A}})\int \tilde{k}(u)^{c_1} \bigg\{\int |Y - \gamma(X,M,a)|^{c_1} f(Y|A = uh+a, M, X)dY \bigg\}\\
&\tilde{k}(u+\frac{a-a'}{h})^{c_2} f_{MX}(M, X) f(uh+a|M,X) dudMdX\\
=& O(\frac{1}{h^{(c_1+c_2-1)d_A}}).
\end{align*}
where the last equality is obtained as in the calculation for ${E}(U_1^{c_1} U_2^{c_2})$.
Combining all the terms, we obtain $\mathbb{E} (m^4_i) = O(h^{-3d_A})$. Then by Markov inequality, for any $\epsilon > 0$,
align*[align* omitted — 692 chars of source]
where the equality in the last row comes from $var( m^2_i) = O(\mathbb{E} (m^4_i)) = O(h^{-3d_A})$.
(II): $h^{d_A}|I_\ell|^{-1}\sum_{i\in I_\ell}\mathbb{E}[ m^2(O_i;\hat{\alpha}_\ell,\hat{\lambda}_\ell,\hat{\gamma}_\ell,\hat{\psi}_\ell(a,a')) - m^2(O_i;\alpha,\lambda,\gamma,\psi(a,a')) \mid O^c_{I_\ell}] = o_p(1)$
For simplicity in notation, we ignore the subscripts $\ell$ below for nuisance parameters estimated from $O^c_{I_\ell}$.
First, we analyze $h^{d_A}\mathbb{E}[m^2(O_i;\hat{\alpha},\hat{\lambda},\hat{\gamma},\hat{\psi}(a,a'))\mid O^c_{I_\ell}]$ as follows.
We write
align*[align* omitted — 273 chars of source]
Denote $m(O_i;\hat{\alpha},\hat{\lambda},\hat{\gamma},\hat{\psi}(a,a'))$ as $\hat{m}_i$. Then,
equation*[equation* omitted — 426 chars of source]
enumerate• $h^{d_A}\mathbb{E}(\hat{U}_1^{2} \mid O^c_{I_\ell})$ \begin{align*}
&h^{d_A}\mathbb{E}\Bigg(\bigg\{\frac{K_h(A - a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}[Y - \hat{\gamma}(X,M,a)] \bigg\}^2 \Bigg| O^c_{I_\ell}\Bigg)\\
=&h^{d_A}\mathbb{E}\Bigg[\mathbb{E}\bigg(\bigg\{\frac{K_h(A - a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}[Y - \hat{\gamma}(X,M,a)] \bigg\}^2\bigg|X,M, O^c_{I_\ell}\bigg)\Bigg| O^c_{I_\ell}\Bigg]\\
=&h^{d_A}\mathbb{E}\Bigg\{\frac{\hat{f}(M \mid A = a', X)^2}{\hat{f}(M \mid A = a, X)^2\hat{f}(a \mid X)^2}\times\\
&\mathbb{E}\bigg[K_h(A - a)^2\mathbb{E}\{[Y - \hat{\gamma}(X,M,a)]^2|X,M,A, O^c_{I_\ell}\}\bigg|X,M, O^c_{I_\ell}\bigg] \Bigg| O^c_{I_\ell} \Bigg\}
\end{align*}
After adding and subtracting $\mathbb{E}[Y \mid X, M, A]$, the middle expectation can be written as
\begin{align*}
&h^{d_A}\mathbb{E}\bigg(K_h(A - a)^2 \bigg\{var(Y|X,M,A) + [\gamma(X,M,a)-\hat{\gamma}(X,M,a)]^2\bigg\}
\bigg|X,M, O^c_{I_\ell} \bigg)\\
= &\int \bigg[\prod^{d_A}_{j=1} k(u_j)^2 \bigg]\times \bigg\{var(Y|X,M,a+uh) + [\gamma(a+uh,M,X)-\hat{\gamma}(X,M,a)]^2\bigg\}
\\
&\times f(a+uh|X,M) du_1 \dots du_{d_A}\\
= & \int k(u_1)^2\cdots k(u_{d_A})^2
\times \bigg\{var(Y|X,M,a)+\sum^{d_A}_{j=1}u_jh\partial_{a_j} var(Y|X,M,a) +\\
&\sum^{d_A}_{j=1}\sum^{d_A}_{j^\prime=1}u_ju_{j^\prime}h^2\partial_{a_j}\partial_{a_j^\prime} var(Y|X,M,\Bar{a}_v) + \\
&\bigg[ \gamma(X,M,a)-\hat{\gamma}(X,M,a)+\sum^{d_A}_{j=1} u_jh\partial_{a_j} \gamma(X,M, a) + \sum^{d_A}_{j=1}\sum^{d_A}_{j'=1} u_ju_{j'}h^2\partial_{a_j}\partial_{a_{j'}} \gamma(X,M,\Bar{a}_{\gamma})\bigg]^2\bigg\}\\
&\times \bigg[f(a|X,M)+\sum^{d_A}_{j=1}u_jh\partial_{a_j} f(a|X,M) + \sum^{d_A}_{j=1}\sum^{d_A}_{j^\prime=1}u_ju_{j^\prime}h^2\partial_{a_j}\partial_{a_{j^\prime}} f(\Bar{a}_f|X,M) \bigg]du_1\cdots du_{d_A}\\
\overset{(a)}{=}& \Big[\int \tilde{k}(u)^2 du\Big] \times \Bigg\{ var(Y|X,M,a) + [\gamma(X,M,a) - \hat{\gamma}(X,M,a)]^2\Bigg\}f(a|X,M) + O(h^2)\\
=& \Big[\int \tilde{k}(u)^2 du\Big] \times \mathbb{E}\Big\{[Y - \hat{\gamma}(X,M,a)]^2\Big|X,M, a, O^c_{I_\ell}\Big\}f(a|X,M) + O(h^2)
\end{align*}
where $\bar{a}_v, \bar{a}_{\gamma}$, and $\bar{a}_f$ are between $a$ and $a+h$. Equality (a) comes from the boundedness of $\int u^6k^2(u)du$, which is true because we assume $0 < \int u^6 k(u) du < \infty$. Plugging this back into the original expectation gives
\begin{align*}
\mathbb{E}\Bigg\{\frac{\hat{f}(M \mid A = a', X)^2}{\hat{f}(M \mid A = a, X)^2\hat{f}(a \mid X)^2}
\Big[\int \tilde{k}(u)^2 du\Big] \times \mathbb{E}\{[Y - \hat{\gamma}(X,M,a)]^2|X,M, a, O^c_{I_\ell}\}\Bigg| O^c_{I_\ell} \Bigg\} + o_p(1)
\end{align*}
• $h^{d_A}\mathbb{E}(\hat{U}_2^{2} \mid O^c_{I_\ell})$
\begin{align*}
&h^{d_A}\mathbb{E}\Bigg\{\bigg[\frac{K_h(A - a')}{\hat{f}(a' \mid X)}\Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big) \bigg]^2 \Bigg| O^c_{I_{\ell}}\Bigg\}\\
=&h^{d_A}\mathbb{E}\Bigg\{\frac{1}{\hat{f}(a' \mid X)^2}\mathbb{E}\bigg[K_h(A - a')^2\Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big)^2\Big|X, O^c_{I_{\ell}}\bigg] \Bigg| O^c_{I_{\ell}}\Bigg\}
\end{align*}
The inner expectation can be written as
\begin{align*}
&h^{d_A}\mathbb{E}\bigg[K_h(A - a')^2\Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big)^2\Big|X, O^c_{I_{\ell}} \bigg]\\
\intertext{Following a similar kernel expansion as before}
=& \int k^2(u_1)\cdots k^2(u_{d_A})\Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big)^2f(a'|X, M)f(M\mid X)
du_1\cdots du_{d_A}dM + O(h^2)
\end{align*}
Plugging this back into the original expectation leads to
\begin{align*}
& \int \tilde{k}(u)^2 du \times \mathbb{E}\bigg\{\frac{1}{\hat{f}(a'|X)^2}\Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big)^2 \bigg| O^c_{I_\ell}\bigg\} + o_p(1)
\end{align*}
• $h^{d_A}\mathbb{E}(\hat{U}_3^{2} \mid O^c_{I_\ell})$
\begin{align*}
h^{d_A}\mathbb{E}\Bigg\{[\hat{\eta}(a, a', X)-\hat{\psi}(a,a')]^2 \bigg| O^c_{I_\ell}\Bigg\} & = o_p(1)
\end{align*}
This holds because we assume the nuisance estimators are bounded, and following a similar calculation as the variance it can be seen that $h^{d_A}\mathbb{E}[\hat{\psi}^2(a, a') \mid O^c_{I_\ell}] = o_p(1)$, which combined with Jensen's inequality can be used to obtain the desired result.
• $h^{d_A}\mathbb{E}(\hat{U}_1\hat{U}_2|O^c_{I_\ell})$
\begin{align*}
&h^{d_A}\mathbb{E}\Bigg\{\bigg[\frac{K_h(A - a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\{Y - \hat{\gamma}(X,M,a)\} \bigg]\\
&\qquad \times \bigg[\frac{K_h(A - a')}{\hat{f}(a' \mid X)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\} \bigg] \bigg| O^c_{I_\ell} \Bigg\}\\
\end{align*}
\begin{align*}
&h^{d_A}\mathbb{E}\Bigg\{\bigg[\frac{K_h(A - a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\{Y - \hat{\gamma}(X,M,a)\} \bigg]\\
&\qquad \times \bigg[\frac{K_h(A - a')}{\hat{f}(a' \mid X)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\} \bigg]\Bigg\}\\
=&h^{d_A}\mathbb{E}\Bigg\{\frac{K_h(A-a)K_h(A-a')}{\hat{f}(a|X)\hat{f}(a'|X)}\frac{\hat{f}(M|a',X)}{\hat{f}(M|a,X)}\Big[Y-\hat{\gamma}(X,M,a)\Big]\Big[\hat{\gamma}(X,M,a)-\hat{\eta}(a,a',X)\Big]\Bigg\}\\
=&h^{d_A}\mathbb{E}\Bigg\{\frac{1}{\hat{f}(a|X)\hat{f}(a'|X)}\frac{\hat{f}(M|a',X)}{\hat{f}(M|a,X)} \Big[\hat{\gamma}(X,M,a)-\hat{\eta}(a,a',X)\Big]\\
& \qquad\times\mathbb{E}\Big\{K_h(A-a)K_h(A-a')\Big[Y-\hat{\gamma}(X,M,a)\Big]\Big|X,M\Big\}\Bigg\}\\
=&h^{d_A}\mathbb{E}\Bigg\{\frac{1}{\hat{f}(a|X)\hat{f}(a'|X)}\frac{\hat{f}(M|a',X)}{\hat{f}(M|a,X)}
\Big[\hat{\gamma}(X,M,a) - \hat{\eta}(a,a',X)\Big]\\
& \qquad\times\mathbb{E}\Big\{K_h(A-a)K_h(A-a')\Big[\gamma(X,M,A)-\hat{\gamma}(X,M,a)\Big]\Big|X,M\Big\}\Bigg\}
\end{align*}
The inner expectation
\begin{align*}
&h^{d_A}\mathbb{E}\Big\{K_h(A-a)K_h(A-a')\Big[\gamma(X,M,A)-\hat{\gamma}(X,M,a)\Big]\Big|X,M\Big\}\\
=& h^{d_A}\int \bigg[\prod^{d_A}_{j=1}\frac{1}{h^2}k\Big(\frac{A_j-a}{h}\Big)k\Big(\frac{A_j-a'}{h}\Big)\bigg]
\Big[\gamma(X,M,A)-\hat{\gamma}(X,M,a)\Big] f(A|X,M)dA\\
=& \int \tilde{k}(u) \tilde{k}(u+\frac{a-a'}{h})\Big[\gamma(X,M, uh + a)-\hat{\gamma}(X,M,a)\Big] f(uh+ a|X,M)du\\
=& \int k(u_1)\cdots k(u_{d_A}) k(u_1+\frac{a-a'}{h})\cdots k(u_{d_A}+\frac{a-a'}{h})\\
&\times \Big[(\gamma(X,M,a) - \hat{\gamma}(X,M,a)) + \sum^{d_A}_{j=1} u_j h\partial_{a_j} \gamma(X,M,a)+\\
&\qquad \frac{u^2_j h^2}{2}\partial^2_{a_j}\gamma(X,M,a)+\frac{u^3_jh^3}{6}\partial^3_{a_j}\gamma(X,M,\Bar{a}_\gamma)\Big]\\
&\times \Big[f(a|X,M) +\sum^{d_A}_{j=1} u_jh\partial_{a_j} f(a|X,M)+\frac{u^2_jh^2}{2}\partial^2_{a_j} f(\Bar{a}_f|X,M)\Big]du_1\cdots du_{d_A}
\end{align*}
where $\bar{a}_{\gamma}$ and $\bar{a}_f$ are between $a$ and $a+h$. Inserting this back into the full expectation combined with Assumption 4 bounds this term as $o_p(1)$.
• $h^{d_A}\mathbb{E}(\hat{U}_1\hat{U}_3|O^c_{I_\ell})$
\begin{align*}
&2h^{d_A}\mathbb{E}\Bigg\{\bigg[\hat{\eta}(a, a', X) - \hat{\psi}(a,a')\bigg]\bigg[\frac{K_h(A - a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\{Y - \hat{\gamma}(X,M,a)\} \bigg] \bigg| O^c_{I_\ell} \Bigg\} = o_p(1)\\
\end{align*}
Expanding this into two terms,
\begin{align*}
&2h^{d_A}\mathbb{E}\Bigg\{\hat{\eta}(a, a', X) \bigg[\frac{K_h(A - a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\{Y - \hat{\gamma}(X,M,a)\} \bigg] \bigg| O^c_{I_\ell} \Bigg\} \\
&- 2h^{d_A}\mathbb{E}\Bigg\{\hat{\psi}(a,a') \bigg[\frac{K_h(A - a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\{Y - \hat{\gamma}(X,M,a)\} \bigg] \bigg| O^c_{I_\ell} \Bigg\}
\end{align*}
The first term can be bounded as $o_p(1)$ using a similar approach used above, and for the second term, from the i.i.d assumption on the data we can re-write it as
\begin{align*}
&2h^{d_A}\mathbb{E}\Bigg\{\hat{\psi}(a,a') \bigg[\frac{K_h(A - a)\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\{Y - \hat{\gamma}(X,M,a)\} \bigg] \bigg| O^c_{I_\ell} \Bigg\} \\
&= 2h^{d_A} |I_\ell|^{-1}\Bigg(\sum_{i \in I_\ell}\mathbb{E}\Bigg\{\bigg[\frac{K_h(A_i - a)\hat{f}(M_i \mid A = a', X_i)}{\hat{f}(M_i \mid A = a, X_i)\hat{f}(a \mid X_i)}\{Y_i - \hat{\gamma}(X_i,M_i,a)\} \bigg]^2 \bigg| O^c_{I_\ell} \Bigg\}\\
&+ \sum_{i \in I_\ell}\mathbb{E}\Bigg\{\bigg[\frac{K_h(A_i - a)\hat{f}(M_i \mid A = a', X_i)}{\hat{f}(M_i \mid A = a, X_i)\hat{f}(a \mid X_i)}\{Y_i - \hat{\gamma}(X_i,M_i,a)\} \bigg]\times \\
&\qquad \bigg[\frac{K_h(A - a')}{\hat{f}(a' \mid X)}\Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big) \bigg] \bigg| O^c_{I_\ell} \Bigg\}\\
&+ \sum_{i \in I_\ell}\mathbb{E}\Bigg\{\bigg[\frac{K_h(A_i - a)\hat{f}(M_i \mid A = a', X_i)}{\hat{f}(M_i \mid A = a, X_i)\hat{f}(a \mid X_i)}\{Y_i - \hat{\gamma}(X_i,M_i,a)\} \bigg]\hat{\eta}(a, a', X) \bigg| O^c_{I_\ell} \Bigg\}
\end{align*}
By the boundedness of$$\mathbb{E}\{[Y - \hat{\gamma}(X,M,a)]^2|X,M,A, O^c_{I_\ell}\} = var(Y|X,M,A) + [\gamma(X,M,a)-\hat{\gamma}(X,M,a)]^2$$ from Assumption 3, and following the results in the first part of (II), we know that $h^{d_A}\mathbb{E}\bigg[K_h(A - a)^2[Y - \hat{\gamma}(X,M,a)]^2\bigg|X,M, O^c_{I_\ell}\bigg]$ is bounded.
Thus, the first term is $O(|I_\ell|^{-1}) = o_p(1)$ from the law of total expectation. Because $\hat{f}(a'|X)$, $\hat{\gamma}$, and $\hat{\eta}$ are bounded by assumptions, the boundedness of $h^{d_A}\mathbb{E}\bigg[K_h(A_i - a)K_h(A_j-a')[Y - \hat{\gamma}(X,M,a)]\bigg|X,M, O^c_{I_\ell}\bigg]$ can be obtained similar to the third part of (I). Hence, the second term also has $O(|I_\ell|^{-1}) = o_p(1)$. From the boundedness of $h^{d_A/2}\mathbb{E}\bigg[K_h(A - a)[Y - \hat{\gamma}(X,M,a)]\bigg|X,M, O^c_{I_\ell}\bigg]$ based on Jensen's inequality and the boundedness of $\hat{\eta}$, the third term satisfies $O(h^{d_A/2}|I_\ell|^{-1}) = o_p(1)$. As a result, $h^{d_A}\mathbb{E}(\hat{U}_1\hat{U}_3|O^c_{I_\ell}) = o_p(1)$.
• $h^{d_A}\mathbb{E}(\hat{U}_2\hat{U}_3|O^c_{I_\ell})$
\begin{align*}
&h^{d_A}\mathbb{E}\Bigg\{\bigg[\hat{\eta}(a, a', X) - \hat{\psi}(a,a')\bigg]
\bigg[\frac{K_h(A - a')}{\hat{f}(a' \mid X)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\} \bigg] \bigg| O^c_{I_\ell} \Bigg\}\\
\end{align*}
From the boundedness of $\hat{\gamma}$, $\hat{\eta}$, and $\hat{f}(a'|X)$, there is $$h^{d_A}\mathbb{E}\Bigg\{\hat{\eta}(a, a', X)
\bigg[\frac{K_h(A - a')}{\hat{f}(a' \mid X)}\{\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\} \bigg] \bigg| O^c_{I_\ell} \Bigg\} = O(h^{d_A}) = o_p(1).$$
A similar proof as the fifth part of (II) above can show that the second term is also $o_p(1)$.
Combining all the six parts, we have $h^{d_A}\mathbb{E}[m^2(O_i;\hat{\alpha},\hat{\lambda},\hat{\gamma},\hat{\psi}(a,a'))]$ equal to
align*[align* omitted — 427 chars of source]
Next, $h^{d_A}\mathbb{E}[m^2(O_i;\alpha,\lambda,\gamma,\psi(a,a'))]$ can be written as
align*[align* omitted — 383 chars of source]
Define $\int \tilde{k}(u)^2 du = R^2_{d_A}$,
align*[align* omitted — 712 chars of source]
Then $h^{d_A}\mathbb{E}[m^2(O_i;\hat{\alpha},\hat{\lambda},\hat{\gamma},\hat{\psi}(a,a'))]- h^{d_A}\mathbb{E}[m^2(O_i;\alpha,\lambda,\gamma,\psi(a,a'))] = \omega_1 + \omega_2 + o_p(1).$ First, we focus on simplifying $\omega_2$, which equals
align*[align* omitted — 214 chars of source]
From expressing $\frac{1}{\hat{f}(a'|X)}\Big(\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big)$ as
align*[align* omitted — 325 chars of source]
there is
align*[align* omitted — 1,480 chars of source]
We show each of these terms are $o_p(1)$ as follows. Because $f^2(a'| X)$ is bounded away from $0$ based on Assumption 3 (ii) and the consistency of $\hat{\gamma}$ from Assumption 4(iii), there is $$\mathbb{E}\bigg\{ \frac{1}{f^2(a'|X)}\Big[\hat{\gamma}(X,M,a) - \gamma(X,M,a)\Big]^2\Big| O^c_{I_\ell} \Big\} = o_p(1).$$
Under a similar argument and from Assumption 4(iv), $$\mathbb{E}\bigg\{ \frac{1}{f^2(a'|X)}\Big[\eta(a, a', X) - \hat{\eta}(a, a', X)\Big]^2\Big| O^c_{I_\ell} \Big\} = o_p(1).$$
Based on the boundedness of nuisance estimators from Assumption 3(ii) and the consistency of $\hat{f}(a'|X)$ from Assumption 4(i), there is
$$\mathbb{E}\bigg\{ \Big[\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\Big]^2\Big[\frac{1}{\hat{f}(a'|X)} - \frac{1}{f(a'|X)}\Big]^2\Big| O^c_{I_\ell} \Big\} = o_p(1).$$
Each of the remaining cross terms is a product of a term that is $o_p(1)$ from the estimator's consistency and a term that is bounded. Hence, we have $\omega_2 = o_p(1)$.
Next, we employ a similar derivation to simplify $\omega_1$.
commentBy adding and subtracting $\frac{f(M \mid A = a', X)(Y - \gamma(X,M,a))}{f(M \mid A = a, X)f(a \mid X)}$ and $\frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\gamma(X,M,a)$,
\begin{align*}
&\frac{\hat{f}(M \mid A = a', X)(Y - \hat{\gamma}(X,M,a))}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\\
=&
\frac{f(M \mid A = a', X)(Y - \gamma(X,M,a))}{f(M \mid A = a, X)f(a \mid X)}
- \frac{f(M \mid A = a', X)(Y - \gamma(X,M,a))}{f(M \mid A = a, X)f(a \mid X)}\\
&+ \frac{\hat{f}(M \mid A = a', X)(Y - \gamma(X,M,a))}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\\
&+ \frac{\hat{f}(M \mid A = a', X)(\gamma(X,M,a) - \hat{\gamma}(X,M,a)) } {\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}
\end{align*}
Adding and subtracting $\frac{f(M \mid A = a', X)}{f(M \mid A = a, X)\hat{f}(a \mid X)}(\gamma(X,M,a) - \hat{\gamma}(X,M,a))$, the above can be re-written as
\begin{align*}
\frac{\hat{f}(M \mid A = a', X)(Y - \hat{\gamma}(X,M,a))}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)} &=
\frac{f(M \mid A = a', X)(Y - \gamma(X,M,a))}{f(M \mid A = a, X)f(a \mid X)}
- \frac{f(M \mid A = a', X)(Y - \gamma(X,M,a))}{f(M \mid A = a, X)f(a \mid X)}\\
&+ \frac{\hat{f}(M \mid A = a', X)(Y - \gamma(X,M,a))}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\\
&+ \frac{1}{\hat{f}(a \mid X)}\Big(\frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)} - \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)}\Big)(\gamma(X,M,a) - \hat{\gamma}(X,M,a))\\
&+ \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)\hat{f}(a \mid X)}(\gamma(X,M,a) - \hat{\gamma}(X,M,a))
\end{align*}
Adding and subtracting $\frac{f(M \mid A = a', X)}{f(M \mid A = a, X)f(a \mid X)}(\gamma(X,M,a) - \hat{\gamma}(X,M,a))$, we get
\begin{align*}
\frac{\hat{f}(M \mid A = a', X)(Y - \hat{\gamma}(X,M,a))}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)} &=
\frac{f(M \mid A = a', X)(Y - \gamma(X,M,a))}{f(M \mid A = a, X)f(a \mid X)}
- \frac{f(M \mid A = a', X)(Y - \gamma(X,M,a))}{f(M \mid A = a, X)f(a \mid X)}\\
&+ \frac{\hat{f}(M \mid A = a', X)(Y - \gamma(X,M,a))}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)}\\
&+ \frac{1}{\hat{f}(a \mid X)}\Big(\frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)} - \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)}\Big)(\gamma(X,M,a) - \hat{\gamma}(X,M,a))\\
&+ \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)}(\gamma(X,M,a) - \hat{\gamma}(X,M,a))\Big(\frac{1}{\hat{f}(a \mid X)} - \frac{1}{f(a \mid X)}\Big)\\
&+ \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)f(a \mid X)}(\gamma(X,M,a) - \hat{\gamma}(X,M,a))
\end{align*}
Finally, adding and subtracting $\frac{\hat{f}(M \mid A = a', X)(Y - \gamma(X,M,a))}{\hat{f}(M \mid A = a, X)f(a \mid X)}$, the above can be written as
\begin{align*}
\frac{\hat{f}(M \mid A = a', X)(Y - \hat{\gamma}(X,M,a))}{\hat{f}(M \mid A = a, X)\hat{f}(a \mid X)} &=
\frac{f(M \mid A = a', X)(Y - \gamma(X,M,a))}{f(M \mid A = a, X)f(a \mid X)}\\
&+\frac{(Y - \gamma(X,M,a))}{f(a \mid X)}\Big(\frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)} - \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)}\Big)\\
&+ \frac{\hat{f}(M \mid A = a', X)(Y - \gamma(X,M,a))}{\hat{f}(M \mid A = a, X)}\Big(\frac{1}{\hat{f}(a \mid X)} - \frac{1}{f(a \mid X)}\Big)\\
&+ \frac{1}{\hat{f}(a \mid X)}\Big(\frac{\hat{f}(M \mid A = a', X)}{\hat{f}(M \mid A = a, X)} - \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)}\Big)(\gamma(X,M,a) - \hat{\gamma}(X,M,a))\\
&+ \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)}(\gamma(X,M,a) - \hat{\gamma}(X,M,a))\Big(\frac{1}{\hat{f}(a \mid X)} - \frac{1}{f(a \mid X)}\Big)\\
&+ \frac{f(M \mid A = a', X)}{f(M \mid A = a, X)f(a \mid X)}(\gamma(X,M,a) - \hat{\gamma}(X,M,a))
\end{align*}
Note that
align*[align* omitted — 648 chars of source]
Hence,
align*[align* omitted — 1,622 chars of source]
After further expansions, we can show that the squared terms contain a component that is bounded based on Assumption 3 and another component that is $o_p(1)$ from Assumption 4. The $(Y - \gamma(X,M,a))^2$ in some squared terms is integrated out as a bounded component due to $var(Y|X,M,a)$ being bounded as assumed in Assumption 3(3). For interaction terms, those containing $(Y - \gamma(X,M,a))$ equals zero because $\int (Y - \gamma(X,M,a))f(Y|X,M,a)dY = 0$. All of the interaction terms contain a bounded component and a $o_p(1)$ component. Consequently, $\omega_1 = o_p(1)$, leading to $h^{d_A}\mathbb{E}[m^2(O_i;\hat{\alpha},\hat{\lambda},\hat{\gamma},\hat{\psi}(a,a'))]- h^{d_A}\mathbb{E}[m^2(O_i;\alpha,\lambda,\gamma,\psi(a,a'))] = o_p(1).$
(III) $h^{d_A}|I_\ell|^{-1} \sum_{i\in I_\ell} \Delta_{i\ell} = o_p(1)$, where $$
\Delta_{i\ell} = m^2(O_i;\hat{\alpha},\hat{\lambda},\hat{\gamma},\hat{\psi}(a,a')) - m^2(O_i;\alpha,\lambda,\gamma,\psi(a,a')) - \mathbb{E}\Big\{m^2(O_i;\hat{\alpha},\hat{\lambda},\hat{\gamma},\hat{\psi}(a,a')) - m^2(O_i;\alpha,\lambda,\gamma,\psi(a,a'))\Big| O^c_{I_\ell}\Big\}.
$$
By Lemma (ref), it suffices to bound $\mathbb{E}\left[\Big(h^{d_A}|I_\ell|^{-1} \sum_{i\in I_\ell} \Delta_{i\ell} \Big)^2 \Big| O^c_{I_\ell}\right] =h^{2d_A}|I_\ell|^{-1} \mathbb{E}\left[\Delta^2_{i\ell} \Big| O^c_{I_\ell}\right] $ as $o_p(1)$. Note that $\mathbb{E}[\Delta_{i\ell}] = 0$ and interaction terms are zero due to conditional independence. We start with analyzing $\mathbb{E}[\Delta^2_{i\ell} \mid O^c_{I_\ell}]$ as follows. For simplicity of notation, we adopt the notation definitions in parts (I) and (II), ignoring the subscripts $\ell$ for nuisance estimators. We have
align*[align* omitted — 350 chars of source]
From (II), we know that $\mathbb{E} \left(\hat{m}^2_{i} - m^2_i \mid O^c_{I_\ell}\right)^2 = o_p( h^{-2d_A})$. To bound $\mathbb{E}\bigg[\big(\hat{m}^2_{i} - m^2_i \big)^2\Big| O^c_{I_\ell}\bigg]$, by $\hat{m}_{i} = \hat{U}_1 + \hat{U}_2 + \hat{U}_3$ and $m_i = U_1 + U_2 + U_3$, we can rewrite the term as
align*[align* omitted — 871 chars of source]
where $\bar{c} = (c_1, \ldots, c_6)$ and $\mathcal{W}$ represents the possible combinations of $\bar{c}$ from the decomposition. We will prove that $\mathbb{E}\bigg\{\hat{U}^{c_1}_1\hat{U}^{c_2}_2\hat{U}^{c_3}_3 U^{c_4}_1U^{c_5}_2U^{c_6}_3 \Big|O^c_{I_\ell}\bigg\} = O(h^{-(c_1 + c_2 + c_4 + c_5 - 1)d_A})$. Note that
$$\mathbb{E}\bigg\{\hat{U}^{c_1}_1\hat{U}^{c_2}_2\hat{U}^{c_3}_3 U^{c_4}_1U^{c_5}_2U^{c_6}_3 \Big| O^c_{I_\ell}\bigg\} = \iint \hat{U}^{c_1}_1\hat{U}^{c_2}_2\hat{U}^{c_3}_3 U^{c_4}_1U^{c_5}_2U^{c_6}_3 f(Y, A, M, X \mid O^c_{I_\ell})dYdMdAdX.$$ By the boundedness of nuisance parameters and their estimates (Assumption 3(ii)), the above term equals
{ $$ O\Big(\iint K_h(A - a)^{c_1 + c_4}K_h(A - a')^{c_2 + c_5}\Big|[Y - \hat{\gamma}(X,M,a)]^{c_{1}}[Y - \gamma(X,M,a)]^{c_{4}}\Big| f(Y, A, M, X \mid O^c_{I_\ell})dYdMdAdX \Big).$$}
The possible combinations of $c_1, c_4$ in $\bar{c}$ are $\{(c_1, c_4): (1, 1), (2, 0), (0, 2), (2, 1), (1, 2), (3, 0), (0, 3)\}$. Similar to the derivation in part (I), we will prove that the rate is $O(h^{-(c_1 + c_2 + c_4 + c_5 - 1)d_A})$ case-by-case. For the terms with $c_1 = 0$, the boundedness of $\mathbb{E}[|Y - \gamma|^4 \mid X, M, A]$ from Assumption 7 provides the boundedness of lower moments by separately considering the regions on which $|Y-\gamma|^{c_4}$ is $\ge$ or $< 1$. Next, we prove for the remaining terms.
enumerate• $c_1 > 0$ and $c_4 = 0$. The integral can be written as
\begin{align*}
&\iint K_h^{c_1 + c_4}(A - a)K_h^{c_2 + c_5}(A - a')\big|Y - \hat{\gamma}(X,M,a)\big|^{c_1} f(Y, A, M, X \mid O^c_{I_\ell})dYdMdAdX \\
=& \iint K_h^{c_1 + c_4}(A - a)K_h^{c_2 + c_5}(A - a')\mathbb{E}\left[\big|Y - \hat{\gamma}(X,M,a)\big|^{c_1} \Big| A, M, X, O^c_{I_\ell}\right] f(A, M, X \mid O^c_{I_\ell})dMdAdX \end{align*}
The inner expectation $\mathbb{E}\left[\big|Y - \hat{\gamma}(X,M,a)\big|^{c_1} \big| A, M, X, O^c_{I_\ell}\right]$ can be bounded as follows,
\begin{align*}
&\mathbb{E}\left[\big|Y - \hat{\gamma}(X,M,a)\big|^{c_1} \Big| A, M, X, O^c_{I_\ell}\right]\\
=& \mathbb{E}\left[\big|Y - \gamma(X,M,a) + \gamma(X,M,a) - \hat{\gamma}(X,M,a)\big|^{c_1} \Big| A, M, X, O^c_{I_\ell}\right]\\
\leq& \mathbb{E}[|Y - \gamma(X,M,a)|^{c_1} \mid A, M, X] + \mathbb{E}[|\gamma(X,M,a) - \hat{\gamma}(X,M,a)|^{c_1} \mid A, M, X, O^c_{I_\ell}]\\
&\quad + \sum^{c_1-1}_{k=1} \binom{c_1}{k} \mathbb{E}[|Y - \gamma(X,M,a)|^{k}|\gamma(X,M,a) - \hat{\gamma}(X,M,a)|^{c_1-k} \big| A, M, X, O^c_{I_\ell}].
\end{align*}
Each of the terms in the expansion can be bounded from Assumption 3(ii) combined with the boundedness of $\mathbb{E}\left[(Y - \gamma(X, M , a))^4 \big| X, M, A\right]$ from Assumption 7. Hence, the original integral equals
\begin{align*}
&O\left(\iint K_h(A - a)^{c_1 + c_4}K_h(A - a')^{c_2 + c_5} f(A, M, X \mid O^c_{I_\ell}) dMdAdX\right)\\
&= O(h^{-(c_1 + c_2 + c_4 + c_5 - 1)d_A})
\end{align*}
where the last equality holds from the boundedness of the integrals of the kernels.
• $c_1 > 0$ and $c_4 > 0$. The integral is
\begin{align*}
&\iint K_h(A - a)^{c_1 + c_4}K_h(A - a')^{c_2 + c_5}\Big|[Y - \hat{\gamma}(X,M,a)]^{c_1}[Y - \gamma(X,M,a)]^{c_4}\Big| \\
&\qquad \times f(Y, A, M, X \mid O^c_{I_\ell})dYdAdMdX\\
=& \iint K_h(A - a)^{c_1 + c_4}K_h(A - a')^{c_2 + c_5}\\
&\mathbb{E}\left[\big|[Y - \hat{\gamma}(X,M,a)]^{c_1}[Y - \gamma(X,M,a)]^{c_4}\big| \Big| A, M, X, O^c_{I_\ell}\right] f(A, M, X \mid O^c_{I_\ell})dAdMdX
\end{align*}
The inner expectation $\mathbb{E}\left[\big|[Y - \hat{\gamma}(X,M,a)]^{c_1}[Y - \gamma(X,M,a)]^{c_4}\big| \Big| A, M, X, O^c_{I_\ell}\right]$ can be bounded with
\begin{align*}
&\mathbb{E}\left[\big|[Y - \hat{\gamma}(X,M,a)]^{c_1}[Y - \gamma(X,M,a)]^{c_4}\big| \Big| A, M, X, O^c_{I_\ell}\right]\\
=& \mathbb{E}\left\{\big|[Y - \gamma(X,M,a) + \gamma(X,M,a) - \hat{\gamma}(X,M,a)]^{c_1}[Y - \gamma(X,M,a)]^{c_4}\big| \Big| A, M, X, O^c_{I_\ell}\right\}\\
\leq& \mathbb{E}\{|Y - \gamma(X,M,a)|^{c_1+c_4} \mid A, M, X\} \\
& \quad + \mathbb{E}[|\gamma(X,M,a) - \hat{\gamma}(X,M,a)|^{c_1} |Y - \gamma(X,M,a)|^{c_4} \mid A, M, X, O^c_{I_\ell}]\\
&\quad +\sum^{c_1-1}_{k=1} \binom{c_1}{k} \mathbb{E}[|\gamma(X,M,a) - \hat{\gamma}(X,M,a)|^{c_1-k}|Y - \gamma(X,M,a)|^{k+c_4} \mid A, M, X, O^c_{I_\ell}]
\end{align*}
The first term is the conditional variance, which is bounded by Assumption 3(3). The second and third terms can be bounded from Assumption 3(ii) and Assumption 7.
The bound of the integral follows similarly as before.
The remaining terms to bound are $\mathbb{E}[\hat{U}^2_1-U^2_1)^2\mid O^c_{I_\ell}]$, $ \mathbb{E}[(\hat{U}^2_2-U^2_2)^2 \mid O^c_{I_\ell}]$, and $ \mathbb{E}[(\hat{U}^2_3-U^2_3)^2 \mid O^c_{I_\ell}]$. First, $\mathbb{E}[(\hat{U}^2_3-U^2_3)^2 \mid O^c_{I_\ell}]$ can be bounded from Assumption 3(ii). Next, we demonstrate the boundedness of $\mathbb{E}[(\hat{U}^2_2-U^2_2)^2 \mid O^c_{I_\ell}]$; a similar derivation applies to $\mathbb{E}[\hat{U}^2_1-U^2_1)^2\mid O^c_{I_\ell}]$. To start with, we re-express the term
align*[align* omitted — 569 chars of source]
From Assumption 7,
align*[align* omitted — 230 chars of source]
Hence,
align*[align* omitted — 286 chars of source]
From the expansion of $\omega_2$ in proving term 6 of the part (II), we can express $\frac{1}{\hat{f}(a'|X)^2}\big[\hat{\gamma}(X,M,a) - \hat{\eta}(a, a', X)\big]^2 - \frac{1}{f(a'|X)^2}\big[\gamma(X,M,a) - \eta(a, a', X)\big]^2$ as a summation of 9 components, i.e.
align*[align* omitted — 1,364 chars of source]
For the multiplication of any two of the nine components chosen with replacement, the corresponding conditional expectation $\mathbb{E}(\cdot|O^c_{I_\ell})$ is a construct of a subcomponent that is $o_p(1)$ from the consistency of nuisance parameters multiplied by other subcomponents that are bounded from Assumption 3. As a consequence, we obtained that $\mathbb{E}[(\hat{U}^2_2-U^2_2)^2 \mid O^c_{I_\ell}] = o_p(h^{-3d_A})$. A similar argument can be used to prove $\mathbb{E}[\hat{U}^2_1-U^2_1)^2\mid O^c_{I_\ell}] = o_p(h^{-3d_A})$ by utilizing the boundedness of $\mathbb{E}[(Y - \gamma)^4 \mid X, M, A]$ from Assumption 7 (i).
Because $ c_1+c_2+c_4+c_5 \le 4$, $O(1) \le O(h^{-(c_1 + c_2 + c_4 + c_5 - 1)d_A}) \le O(h^{-3d_A})$.
We conclude that $\mathbb{E}\left[\Delta^2_{i\ell} \Big| O^c_{I_\ell}\right] = O(h^{-3d_A})$ and
$$\mathbb{E}\left[\Big(h^{d_A}|I_\ell|^{-1} \sum_{i\in I_\ell} \Delta_{i\ell} \Big)^2 \Big| O^c_{I_\ell}\right] =h^{2d_A}|I_\ell|^{-1} \mathbb{E}\left[\Delta^2_{i\ell} \Big| O^c_{I_\ell}\right] =O\left([nh^{d_A}]^{-1}\right) = o_p(1).$$
Consistency of Hajek-type Propensity Estimator in Cross Validation
Given a consistent estimator of propensity score at treatment value $a$ for person $i$ in cross validation fold $I_\ell$, $\hat{f}(a|X_i)$, we define the corresponding Hajek-type stabilized weighted propensity score as follows, $$\hat{f}(a|X_i) \times \frac{1}{|I_{-\ell}|}\sum^{|I_{-\ell}|}_{j\in I_{-\ell}} \frac{K_h (A_j-a)}{\hat{f}(a|X_j)}. $$
The goal here is to prove that
$$\lim_{|I_{-\ell}| \rightarrow \infty}\hat{f}(a|X_i) \times \frac{1}{|I_{-\ell}|}\sum^{|I_{-\ell}|}_{j\in I_{-\ell}} \frac{K_h (A_j-a)}{\hat{f}(a|X_j)}\stackrel{p}{=} f(a|X_i). $$
enumerate• \begin{align*}
&\lim_{|I_{-\ell}| \rightarrow \infty}\frac{1}{|I_{-\ell}|}\sum^{|I_{-\ell}|}_{j\in I_{-\ell}} \frac{K_h (A_j-a)}{\hat{f}(a|X_j)} \stackrel{p}{=} \iint \frac{K_h(A-a)}{f(a|X)} f(A,X) dAdX \%{\mathcal{A}_{-\ell}\times\mathcal{X}_{-\ell}}
=&\int_{\mathcal{X}} \bigg\{ \int_{\mathcal{A}} K_h(A-a) \frac{f(A,X)}{f(a|X)}dA\bigg\}dX\\
& by Lemma 4\\
=& \iint \prod^{d_A}_{k=1} k(u_k)\bigg\{ \frac{f(a,X)}{f(a|X)} + \sum^{d_A}_{k=1} u_k h \frac{\partial_{a_k}f(a,X)}{f(a|X)} + \frac{1}{2} \sum^{d_A}_{k=1} \sum^{d_A}_{k'=1} u_k u_{k'}h^2\frac{\partial_{a_k}\partial_{a_{k'}}f(a,X)}{f(a|X)}\bigg|_{\bar{a}}\bigg\} \\
&\qquad du_1\ldots du_{d_A} dX\\
& assume \int_{\mathcal{X}}\partial_{a_k}\partial_{a_{k'}}f(a,X)|_{\bar{a}}dX < \infty for \bar{a} between a and a+uh, then\\
=& \int_{\mathcal{X}} f(X) dX + O(h^2) = 1 + O(h^2)
\end{align*}
• $\lim_{|I_{-\ell}| \rightarrow \infty}\hat{f}(a|X_i) = f(a|X_i)$ from the consistency of the propensity estimator $\hat{f}$
• Combining the first two bullets, we get
\begin{align*}
&\lim_{|I_{-\ell}| \rightarrow \infty}\hat{f}(a|X_i) \times \frac{1}{|I_{-\ell}|}\sum^{|I_{-\ell}|}_{j\in I_{-\ell}} \frac{K_h (A_j-a)}{\hat{f}(a|X_j)}\\
=& \lim_{|I_{-\ell}| \rightarrow \infty}\hat{f}(a|X_i) \times
\lim_{|I_{-\ell}| \rightarrow \infty}\frac{1}{|I_{-\ell}|}\sum^{|I_{-\ell}|}_{j\in I_{-\ell}} \frac{K_h (A_j-a)}{\hat{f}(a|X_j)}\\
\stackrel{p}{=}& f(a|X_i)
\end{align*}