Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
On the Asymptotic Properties of Debiased Machine Learning Estimators
\relax
\hypersetup{pageanchor=false}
\hypersetup{pageanchor=true}
\thispagestyle{empty}
spacing{1.2}
\begin{abstract}
This paper studies the properties of debiased machine learning (DML) estimators under a novel asymptotic framework, offering insights for improving the performance of these estimators in applications. DML is an estimation method suited to economic models where the parameter of interest depends on unknown nuisance functions that must be estimated. It requires weaker conditions than previous methods while still ensuring standard asymptotic properties. Existing theoretical results do not distinguish between two alternative versions of DML estimators, DML1 and DML2. Under a new asymptotic framework, this paper demonstrates that DML2 asymptotically dominates DML1 in terms of bias and mean squared error, formalizing a previous conjecture based on simulation results regarding their relative performance. Additionally, this paper provides guidance for improving the performance of DML2 in applications.
\end{abstract}
\thispagestyle{empty}
Introduction
Debiased machine learning (DML) has become a popular method for estimating parameters in economic models. DML is particularly suited to cases where the parameter of interest depends on unknown nuisance functions that require estimation chernozhukov2018double. DML offers standard asymptotic properties (e.g., asymptotic normality and parametric convergence rate) under milder conditions compared to previous methods (e.g., newey1994asymptotic, andrews1994asymptotics, newey1994large). In practice, two versions of DML ---introduced by chernozhukov2018double---can be used, DML1 and DML2. Both versions randomly divide the data into $K$ equal-sized folds (samples) to estimate the nuisance function, but they differ in how these estimates are used to construct an estimator for the parameter of interest. While DML2 is believed to perform better than DML1 based on simulation results about their relative performance, DML1 and DML2 yield estimators with the same asymptotic distribution when $K$ remains fixed as the sample size $n$ diverges to infinity.
This paper studies the properties of DML1 and DML2 under a novel asymptotic framework, where the number of folds $K$ diverges to infinity as $n$ diverges to infinity. Under this asymptotic framework, I show that DML2 offers theoretical advantages over DML1 in terms of bias and mean-square error (MSE). This result suggests that practitioners should adopt DML2 to achieve more accurate and reliable results.
Additionally, it provides practical recommendations for improving the performance of DML2 in applications, specifically conditions under which setting $K$ equals $n$ minimizes asymptotic bias and MSE for DML2.
DML is useful to estimate a parameter $\theta_0$ that satisfies a moment condition of the form:
equation[equation omitted — 77 chars of source]
where $m$ is a known moment function, $W$ is an observed random vector, and $\eta_0$ is an unknown nuisance function. Examples of a parameter $\theta_0$ that can be identified by the moment condition (ref) include several treatment effect parameters, such as the average treatment effect (ATE), the average treatment effect on the treated in difference-in-differences designs (ATT-DID), the local average treatment effect (LATE), the weighted average treatment effects (w-ATE), the average treatment effect on the treated (ATT), and the treatment effect coefficient in the partial linear model (PLM), which have been studied in the literature of semi-parametric models (e.g., robinson1988root, robins1994estimation, hahn1998role, hirano2003efficient, frolich2007nonparametric, farrell2015robust, chernozhukov2017double, sant2020doubly, chang2020double). In all these examples, the moment function $m$ is linear in a real-valued parameter $\theta_0$, and the nuisance function $\eta_0$ consists of conditional expectations, such as the propensity score. This paper considers a setup that includes all these examples.
DML relies on two ingredients to guarantee that the estimation of $\theta_0$ is as accurate as if the true $\eta_0$ had been used. The first ingredient is the Neyman orthogonality condition on the moment function $m$. This condition reduces the sensitivity of the estimation of $\theta_0$ to errors in the estimation of $\eta_0$; see Remarks (ref) and (ref) for additional details. The second ingredient is the cross-fitting procedure, a sample-splitting method used to construct estimators for $\eta_0$. This procedure and the Neyman orthogonality condition on $m$ remove the “own observation” bias, which arises when the same data is used to estimate both $\eta_0$ and $\theta_0$.
DML1 and DML2 estimate $\theta_0$ by first randomly dividing the data into $K$ equal-sized folds, denoted by $\mathcal{I}_k$ for $k=1,\ldots,K$. For each fold $\mathcal{I}_k$, an estimator $\hat{\eta}_k$ of $\eta_0$ is constructed using all the data except the data in fold $\mathcal{I}_k$.
Then, DML1 first calculates preliminary estimators $\tilde{\theta}_k$ by solving the moment condition (ref) within each fold $\mathcal{I}_k$ using the estimator $\hat{\eta}_k$. It then combines the information across the folds by averaging the $\tilde{\theta}_k$'s to obtain the proposed estimator for $\theta_0$.
In contrast, DML2 first combines the information across the folds by averaging moment conditions based on (ref), where each fold uses estimates $\hat{\eta}_k$, and then $\theta_0$ is estimated as the solution in $\theta$ of the average of moment conditions, $K^{-1} \sum_{k=1}^K \left( (n/K)^{-1} \sum_{i \in \mathcal{I}_k} m(W_i, \theta, \hat{\eta}_k) \right) = 0$.
chernozhukov2018double conjectured that DML2 performs better than DML1 in small samples based on simulation results about their relative performance. However, existing asymptotic theory is insufficient to validate this conjecture, as it predicts that both DML1 and DML2 lead to estimators with the same limiting distribution, assuming that the number of folds $K$ remains fixed as the sample size $n$ diverges to infinity.
This paper studies the properties of the estimators based on DML1 and DML2 under a new asymptotic framework, aiming to understand which version has theoretical advantages. I consider an asymptotic framework where the number of folds $K \to \infty$ as $n \to \infty$. This framework offers a better description of finite sample situations where the practitioner implementing DML desires to increase $K$ to improve the precision of the estimators $\hat{\eta}_k$'s, which use a fraction $(K-1)/K$ of the data. The use of alternative asymptotic approximations incorporating features of the finite sample problem is not new in the literature; an incomplete list of similar and recent approaches in different econometric problems include cattaneo2018kernel, bugni2021testing, and cai2022linear, among many other authors.
This paper makes three contributions. First, it shows that DML2 offers theoretical advantages over DML1 in terms of bias and MSE, formalizing a previous conjecture based on simulation evidence. More concretely, it is shown that the first-order asymptotic distribution of DML1 may exhibit an asymptotic bias, which is not the case for DML2. This asymptotic bias is proportional to a parameter $\Lambda$ that only depends on the moment function $m$, $\theta_0$, and $\eta_0$.
When $\Lambda$ equals zero, DML1 and DML2 exhibit similar first-order asymptotic
properties. However, as $\Lambda$ deviates from zero, DML1 becomes increasingly sensitive to
large values of $K$ regarding bias and MSE, while DML2 remains unaffected by the choice of
$K$. For several treatment effect parameters ---such as the ATE, ATT-DID, ATT, and PLM--- $\Lambda$ equals zero, but for others, like the LATE and w-ATE, it is typically nonzero. The distinction between DML1 and DML2 through $\Lambda$ emerges under the proposed asymptotic framework, providing insights not captured by existing asymptotic theory or simulation-based evidence.
Second, this paper provides conditions that guarantee the asymptotic validity of implementing DML2 with any number of folds. Specifically, under these conditions, the first-order asymptotic distribution of estimators based on DML2 when $K \to \infty$ as $n \to \infty$ is the same as in the existing first-order asymptotic theory (where $K$ is fixed). In particular, it is possible to set the number of folds $K$ equal to the sample size $n$ for DML2 and obtain a leave-one-out estimator with the same asymptotic distribution. This result is particularly useful for practitioners. It shows that dividing the data into many folds to implement DML2 is asymptotically valid, a common practice to improve the precision of the nuisance function estimates $\hat{\eta}_k$'s; see Remark (ref) for additional discussion on increasing the number of folds. Furthermore, the conditions that guarantee the asymptotic validity for DML2 allow for the study of higher-order properties for all the DML2 estimators, including the leave-one-out.
Third, this paper provides conditions under which the leave-one-out estimator is asymptotically optimal in terms of bias and MSE among the estimators based on DML2. More concretely, under these conditions, the absolute value of the leading term in the higher-order asymptotic bias of DML2 estimators decreases as $K$ increases, with the minimum achieved at $K = n$. Therefore, the leave-one-out estimator is optimal in terms of bias among the estimators based on DML2. Moreover, it holds that the leave-one-out estimator is optimal with respect to the second-order asymptotic MSE whenever certain data-dependent conditions hold, making it the most efficient choice for practitioners.
Finally, the previous results offer several lessons for practitioners.
First, DML2 is the recommended option for implementing DML, especially in small-sample situations when increasing the number of folds is desired to improve the precision of the estimators $\hat{\eta}_k$.
Second, choosing the number of folds $K$ equal to the sample size $n$ is optimal for implementing DML2 to reduce the asymptotic bias, where the asymptotic bias refers to the leading term of the higher-order asymptotic bias of the DML2 estimator. Third, choosing $K=n$ is also optimal for the asymptotic accuracy of DML2 to estimate $\theta_0$ when certain data-dependent condition holds, where asymptotic accuracy refers to the second-order asymptotic MSE.
The previous two lessons reveal that the common recommendations of choosing 5, 10 or 20 folds for the cross-fitting procedure in DML (e.g., ahrens2024ddml,ahrens2024model, bach2022doubleml, and JSSv108i03) are suboptimal in terms of asymptotic bias and asymptotic accuracy. The next lesson concerns the relative loss a practitioner can face by choosing a $K$ different than the optimal choice.
Fourth, if the optimal choice for minimizing bias and second-order asymptotic MSE is $K=n$, then choosing $K=10$ to implement DML2 guarantees that the maximum relative loss compared to the optimal choice in terms of asymptotic bias and asymptotic accuracy is around 10% and 5%, respectively.
The conditions provided in this paper include stronger assumptions about the estimators of $\eta_0$ compared to the existing first-order asymptotic theory of DML to address two challenging situations. The first concerns the asymptotic properties of the estimators of $\theta_0$, as the proof strategy for fixed $K$ cannot be directly adapted to the case where $K \to \infty$ as $n \to \infty$. The second challenge involves analyzing the higher-order properties of the DML2 estimator when $K \to \infty$ as $n \to \infty$, which requires additional structure on the estimators of $\eta_0$.
Related Literature
This paper contributes to the growing literature on DML, where different estimators have been proposed based on DML for addressing semi-parametric estimation problems without requiring strong conditions on the estimators of $\eta_0$ (e.g., without invoking a Donsker class assumption). An incomplete list in this literature includes chernozhukov2017double, chernozhukov2018double, chernozhukov2022locally,chernozhukov2022automatic,chernozhukov2022debiased, semenova2021debiased, semenova2023adaptive,semenova2023debiased, escanciano2023machine, rafi2023efficient, cheng2023weight, ji2023model, noack2021flexible, fava2024predicting, kennedy2024minimax, and jin2024structure. All these papers use DML2 with some exceptions, such as chernozhukov2017double, ji2023model, and cheng2023weight, which used DML1.\footnote{In many of these papers, such as rafi2023efficient and semenova2023adaptive, DML1 and DML2 are numerically equivalent; see Remark (ref) for an explanation.} Except for kennedy2024minimax and jin2024structure, all these papers derive the first-order asymptotic theory for their estimators, assuming that $K$ remains fixed as $n \to \infty$. kennedy2024minimax and jin2024structure use a structure-agnostic framework to show the optimality of estimators based on DML. In contrast, I study the properties of estimators based on DML1 and DML2 when $K \to \infty$ as $n \to \infty$, showing that DML2 offers theoretical advantages over DML1 and providing conditions under which the leave-one-out estimator, defined as DML2 using $K=n$, is optimal in terms of bias and MSE. To the best of my knowledge, this literature does not provide theoretical results on selecting $K$, which is also addressed as part of my results.
This paper also contributes to the literature on double-robust estimators, which includes robins1994estimation, robins1995semiparametric, scharfstein1999adjusting, farrell2015robust, sant2020doubly, chang2020double, callaway2021difference, rothe2019properties, singh2024double, among others. Except rothe2019properties, all these papers study first-order asymptotic theory for their estimators that remain consistent even if some components of $\eta_0$ are misspecified. rothe2019properties studies the higher-order properties of double-robust estimators in a missing-data setting, where $\eta_0$ is estimated by a leave-one-out approach. My results complement their findings. First, DML versions of the double-robust estimators allow flexible estimation of the components of $\eta_0$ (e.g., nonparametric methods). Second, this paper presents the higher-order properties of estimators based on DML2. Third, among these DML2 estimators, the leave-one-out estimator is optimal in terms of bias and MSE whenever certain conditions hold. Interestingly, the leave-one-out estimator in this paper is the estimator studied in rothe2019properties.
More broadly, this paper contributes to the literature on semi-parametric models, which has a long tradition in econometrics and statistics (e.g., bickel1982adaptive, robinson1988root, newey1990efficient, andrews1994asymptotics, newey1994large, newey1994asymptotic, linton1995second, bickel2003nonparametric). Many of the papers in this literature provide conditions to study the estimators based on a plug-in approach (i.e., the same data is used to estimate $\eta_0$ and $\theta_0$). In contrast, I provide conditions to study the (higher-order) properties of estimators based on DML2.
Structure of the rest of the paper
The remainder of the paper is organized as follows. Section (ref) describes the setup, notation, and estimators based on DML.
Section (ref) presents the formal results: Section (ref) states the first-order asymptotic properties of the DML1 and DML2 estimators when $K \to \infty$ as $n \to \infty$, and Section (ref) finds the higher-order properties of the DML2 estimators when $K \to \infty$ as $n \to \infty$. Section (ref) presents the lessons for practitioners based on the formal results obtained in Section (ref). Section (ref) revisits Monte-Carlo simulations for the ATT-DID (sant2020doubly) and the LATE (hong2010semiparametric), using estimators based on DML1 and DML2, to examine the relevance of my asymptotic analysis in finite samples.
Finally, Section (ref) presents concluding remarks. Appendix (ref) presents additional examples and results. Appendix (ref) collects the proof of the main results. For brevity, the proofs of the auxiliary results that appear in Appendix (ref) have been placed in the Online Appendix.\footnote{ \href{https://www.amilcarvelez.com/JMP/DML/online_appendix.pdf}{https://www.amilcarvelez.com/JMP/DML/online_appendix.pdf}}
Setup and Notation
This section presents the setup for the parameter of interest and the estimators based on DML. It contains examples previously studied in the literature that illustrate the setup. It also states the results of the existing asymptotic theory.
The parameter of interest is $\theta_0 \in \Theta \subseteq \mathbf{R}$ and satisfies the following moment condition:
equation[equation omitted — 79 chars of source]
where $m: \mathcal{W} \times \Theta \times \mathcal{T} \to \mathbf{R}$ is a known moment function, $W \in \mathcal{W} \subseteq \mathbf{R}^{d_w}$ is a random vector with distribution $F_0$, and $X \in \mathcal{X} \subseteq \mathbf{R}^{d_x}$ is a sub-vector of $W$. The nuisance parameter $\eta_0: \mathcal{X} \to \mathcal{T} \subseteq \mathbf{R}^p$ is an unknown function of the covariates $X$.
This paper considers moment functions $m$ that are linear in the parameter of interest:
equation[equation omitted — 102 chars of source]
where $\psi^b$ and $\psi^a$ are functions that satisfy conditions specified in Assumption (ref), which includes the identification condition $E[\psi^a(W,\eta_0(X))] \neq 0$ and guarantees a Neyman orthogonality condition,
$$ E\left[ \partial_\eta m(W,\theta_0,\eta_0(X)) \mid X \right] = 0~, \quad a.e.$$
where $\partial_\eta m$ denotes the partial derivative of the function $m$ with respect to the values of $\eta$ and $\partial_\eta m(W,\theta_0,\eta_0(X))$ is the $\partial_\eta m$ evaluated at $\eta = \eta_0(X)$.
A wide range of parameters of interest can be identified through moment conditions such as (ref) using a moment function like (ref).
Examples of $\theta_0$ include the average treatment effect (Example (ref)), the average treatment effect on the treated in difference-in-differences designs (Example (ref)), and the local average treatment effect (Example (ref)), among others. Further examples are presented in Section (ref) of Appendix (ref).
example[Average Treatment Effect]
Let $A \in \{0,1\}$ denote a binary treatment status, $Y(a)$ denote the potential outcome under treatment $a \in \{0,1\}$, $X$ denote a vector of covariates, and
$$Y = A Y(1) + (1-A) Y(0)$$
denote the observed outcome. The available data is modeled by the vector $W = (Y,A,X)$.
The parameter of interest is
$$\theta_0 = E[Y(1) - Y(0)]~,$$
which is the expectation of the treatment effect when the treatment is mandated across the entire population, also known as the ATE.
A standard assumption used to identify $\theta_0$ is the selection-on-observables assumption,
$$(Y(1),Y(0)) \perp A \mid X~.$$
Under the selection-on-observables assumption, the ATE can be identified by a moment condition such as (ref) using a moment function like (ref), which is defined by
\begin{align*}
\psi^b(W,\eta) &= \eta_1 - \eta_2 + A(Y-\eta_1) \eta_3 - (1-A)(Y-\eta_2) \eta_4 ,\\
\psi^a(W,\eta) &= 1 ,
\end{align*}
for $\eta \in \mathbf{R}^4$, and where the nuisance parameter $\eta_0(X)$ has four components:
\begin{align*}
\eta_{0,1}(X) &= E[Y \mid X, A=1] , \\
\eta_{0,2}(X) &= E[Y \mid X, A=0] , \\
\eta_{0,3}(X) &= (E[A \mid X])^{-1} , \\
\eta_{0,4}(X) &= (E[1 -A \mid X])^{-1} .
\end{align*}
This moment function corresponds to the augmented inverse propensity weighted (AIPW) estimator (robins1994estimation, scharfstein1999adjusting). It also appears as the efficient influence function for the ATE in hahn1998role and hirano2003efficient.
example[Difference-in-Differences]
This example considers the average treatment effect on the treated in difference-in-differences research designs with two periods and panel data, as studied in sant2020doubly. Let $A \in \{0,1\}$ denote a binary treatment status on the post-period treatment, $Y_1(a)$ denote the potential outcome on the post-period treatment under treatment status $ a \in \{0,1\}$, $Y_0$ denote the outcome of interest in a pre-treatment period, $X$ denote a vector of covariates, and
$$ Y_1 = A Y_1(1) + (1-A) Y_1(0)~$$
denote the observed outcome in the post-treatment period. The available data is modeled by the vector $W = (Y_0, Y_1, A, X)$. The parameter of interest is
$$\theta_0 = E[Y_1(1) - Y_1(0) \mid A=1]~,$$
which represents the treatment effect for the treated group in the post-treatment period, also known as ATT-DID. sant2020doubly used the following conditional parallel trend assumption,
$$ E[Y_1(0) - Y_0 \mid X, A=1 ] = E[Y_1(0) - Y_0 \mid X, A=0 ]~, $$
to identify the ATT-DID by a moment condition, such as (ref), using a moment function like (ref), which is defined by
\begin{align*}
\psi^b(W,\eta) &= A( Y_1 - Y_0 - \eta_1) + (1-A) (1-\eta_2) (Y_1-Y_0 - \eta_1) ,\\
\psi^a(W,\eta) &= A ,
\end{align*}
for $\eta \in \mathbf{R}^2$, and where the nuisance parameter $\eta_0(X)$ has two components:
\begin{align*}
\eta_{0,1}(X) &= E[ Y_1 - Y_0 \mid X, A=0] \\
\eta_{0,2}(X) &= (E[1 - A \mid X])^{-1} .
\end{align*}
This moment function is the efficient influence function for the ATT-DID under the conditions in
sant2020doubly.
example[Local Average Treatment Effect]
This example considers a framework where individuals can decide their treatment status as in imbens1994identification and frolich2007nonparametric.
Let $Z \in \{0,1\}$ denote a binary instrumental variable (e.g., treatment assignment), $D(z)$ denote potential treatment decisions under the intervention $z \in \{0,1\}$, and assume the observed treatment decision is given by
$$D = Z D(1) + (1-Z) D(0)~.$$
Let $X$ denote a vector of covariates, $Y(d)$ denote the potential outcome under treatment decision $d \in \{0,1\}$, and $Y = D Y(1) + (1-D) Y(0)$ denote the observed outcome. The available data is modeled by the vector $W = (Y,Z,D,X)$. The parameter of interest is
$$\theta_0 = E[ Y(1) - Y(0) \mid D(1) > D(0)]~,$$
which is the expected treatment effect for the sub-population that complies with the assigned treatment, also known as LATE. A sufficient assumption for identification is the following selection-on-observables assumption,
$$(Y(1),Y(0),D(1),D(0)) \perp Z \mid X~.$$
Using this assumption and similar assumptions as in frolich2007nonparametric, singh2024double identified the LATE by a moment condition, such as (ref), using a moment function like (ref), which is defined by
\begin{align*}
\psi^b(W,\eta) &= \eta_1 - \eta_2 + Z(Y-\eta_1)\eta_5 - (1-Z)(Y-\eta_2)\eta_6 \\
\psi^a(W,\eta) &= \eta_3 - \eta_4 + Z(D-\eta_3)\eta_5 - (1-Z) (D-\eta_4) \eta_6
\end{align*}
for $\eta \in \mathbf{R}^6$, and where the nuisance parameter $\eta_0(X)$ has six components:
\begin{align*}
\eta_{0,1}(X) &= E[Y \mid X, Z = 1] , \\
\eta_{0,2}(X) &= E[Y \mid X, Z = 0] , \\
\eta_{0,3}(X) &= E[D \mid X, Z = 1] , \\
\eta_{0,4}(X) &= E[D \mid X, Z = 0] , \\
\eta_{0,5}(X) &= (E[Z \mid X])^{-1} , \\
\eta_{0,6}(X) &= (E[1-Z \mid X])^{-1} .
\end{align*}
This moment function appears in frolich2007nonparametric as the efficient influence function for the LATE. This moment function corresponds to the estimators proposed in tan2006regression.
Estimators based on DML
Consider the goal of estimating $\theta_0$ using a random sample $\{W_i : 1 \le i \le n\}$ drawn from the distribution $F_0$. The parameter $\theta_0$ based on (ref) and (ref) can be identified as a ratio of expected values,
equation[equation omitted — 103 chars of source]
Accordingly, an ideal estimator for $\theta_0$ is defined by replacing the expected values in (ref) with sample analogs. That is,
equation[equation omitted — 151 chars of source]
where $\eta_i = \eta_0(X_i)$ is the value of the nuisance parameter $\eta_0$ for the observation $i$, and $X_i$ is a sub-vector of $W_i$. However, the values of the $\eta_i$'s are unknown. As a result, the oracle estimator $\hat{\theta}_n^*$ is infeasible. For this reason, it is common to calculate first estimates $\hat{\eta}_i$ of $\eta_i$ that can be used later to compute an estimator of $\theta_0$.
For instance, an estimator $\hat{\eta}$ of $\eta_0$ can be obtained by using all the data, and then an estimator of $\theta_0$ can be defined by replacing $\eta_i$ by the estimates $\hat{\eta}_i$ in (ref), where $\hat{\eta}_i = \hat{\eta}(X_i)$. An estimator of $\theta_0$ based on this approach is known as the plug-in estimator, and the conditions under which it has standard properties (e.g., asymptotic normality and parametric convergence rates) have been studied in the literature on semi-parametric models (e.g., andrews1994asymptotics, newey1994asymptotic, newey1994large). However, this approach is sensitive to the “own observation” bias, which arises when the same data is used to estimate both $\eta_0$ and $\theta_0$ (newey2018cross); see also Remark (ref). To attenuate the first-order effect of this bias on the plug-in estimators, stronger conditions are required on the estimator $\hat{\eta}$ (e.g., Donsker conditions). In contrast, DML ---the approach considered in this paper--- relies on general and simple conditions, such as a certain mean-square consistency condition, to obtain the standard properties (chernozhukov2018double, chernozhukov2022locally).
In what follows, I explain how DML estimates $\eta_0$ and $\theta_0$ by relying on cross fitting to avoids the “own observation” bias.
Estimates of the Nuisance Parameter
DML proposes to calculate the estimates $ \hat{\eta}_i$ of $\eta_i$ using a cross-fitting procedure, which is a form of sample-splitting.
This procedure has two steps and implicitly assumes that $n$ can be divided by $K$:
enumerate• Sample splitting: Randomly split the indices into $K$ equal-sized folds $\mathcal{I}_k$, i.e., $\cup_{k=1}^K \mathcal{I}_k = \{1,2,\ldots,n\}$. The number of observations in fold $\mathcal{I}_k$ is denoted $n_k = n/K$.\footnote{When $n$ is not divisible by $K$, the number of observations in some folds will be $\lfloor n/K \rfloor$ while in others $\lfloor n/K \rfloor + 1$, where $\lfloor n/K \rfloor$ is the greatest integer less than or equal to $n/K$.}
• Nuisance Parameter Estimates: For each fold $\mathcal{I}_k$, the estimates $ \hat{\eta}_i $ of $\eta_i$ are defined by
\begin{equation}
\hat{\eta}_i = \hat{\eta}_k(X_i) , \quad \forall i \in \mathcal{I}_k ,
\end{equation}
where $\hat{\eta}_k(\cdot)$ is an estimator of the nuisance parameter $\eta_0(\cdot)$ using $\{ W_i : i \notin \mathcal{I}_k\}$, which is all the data except the ones with indices on the fold $\mathcal{I}_k$. All the estimates $\hat{\eta}_i$ are calculated by repeating the process for all the $k=1,\ldots,K$.
Both DML estimators use these estimates $\hat{\eta}_i$, but they differ in how they combine information across the different folds defined above. I explain this next.
DML1
The estimator based on DML1 first calculates preliminary estimators $\tilde{\theta}_k$ by solving the moment condition (ref) within each fold $\mathcal{I}_k$ using the estimates $\hat{\eta}_i$,
$$ \tilde{\theta}_k \quad \text{solve} \quad n_k^{-1} \sum_{i \in \mathcal{I}_k} m(W_i, \theta, \hat{\eta}_i) = 0 ~,$$
it then combines the information across the folds by averaging the $\tilde{\theta}_k$'s to obtain the proposed estimator for $\theta_0$,
equation[equation omitted — 102 chars of source]
Explicit expressions for $\tilde{\theta}_{k}$ can be obtained since the moment function $m$ is as in (ref),
equation*[equation* omitted — 208 chars of source]
Note that $\tilde{\theta}_{k}$ is similar to (ref) but using only observations in the fold $\mathcal{I}_k$ and the estimates $\hat{\eta}_i$ instead of $\eta_i$.
DML2
In contrast, the estimator based on DML2 first combines the information across the folds $\mathcal{I}_k$ by averaging the sample analog of moment conditions like (ref) using the estimates $\hat{\eta}_i$, and then estimates $\theta_0$ by solving the average of moment conditions,
$$ \hat{\theta}_{n,2} \quad \text{solve} \quad K^{-1} \sum_{k=1}^K \left( n_k^{-1} \sum_{i \in \mathcal{I}_k} m(W_i, \theta, \hat{\eta}_i) \right) = 0 ~.$$
An explicit expression for $\hat{\theta}_{n,2}$ is obtained by using that the moment function $m$ is as in (ref),
equation[equation omitted — 159 chars of source]
Note that $\hat{\theta}_{n,2}$ is similar to (ref) but using the estimates $\hat{\eta}_i$ instead of $\eta_i$.
remarkThe estimators based on DML1 and DML2 can be equal under certain conditions.
If $\psi^a(W_i,\hat{\eta}_i)$ has zero variance (e.g., $\psi^a$ is a constant $\psi^a_0$ as in Example (ref)) and the $K$-fold partition $\{I_k: 1 \le k \le K\}$ divides the data into exactly $K$ subsets with equal size, then both DML1 and DML2 estimators defined in (ref) and (ref) are equal. In particular,
$$ \hat{\theta}_{n,1} = K^{-1} \sum_{k=1}^K \frac{ n_k^{-1}\sum_{i \in \mathcal{I}_k} \psi^b(W_i,\hat{\eta}_i)}{ \psi^a_0} = \frac{n^{-1}\sum_{i=1}^n \psi^b(W_i,\hat{\eta}_i)}{ \psi^a_0 } = \hat{\theta}_{n,2} ~.$$
Therefore, the DML1 and DML2 estimators for the ATE (Example (ref)) are numerically the same when the data are divided in exactly $K$ folds.
In contrast, if $\psi^a(W_i,\hat{\eta}_i)$ has positive variance, then $\hat{\theta}_{n,1} \neq \hat{\theta}_{n,2}$ in general. This occurs in all the other examples.
remark[Oracle version of DML]
The oracle version of the DML1 estimator depends on random splitting, while this is not the case for the oracle version of the DML2. By oracle version, I refer to the case where the DML estimators are calculated assuming perfect knowledge of the $\eta_i$. More concretely, the oracle version of the estimator based on DML1 is defined as
\begin{equation}
\hat{\theta}_{n,1}^* = K^{-1} \sum_{k=1}^K \frac{ n_k^{-1}\sum_{i \in \mathcal{I}_k} \psi^b(W_i,{\eta}_i)}{ n_k^{-1} \sum_{i \in \mathcal{I}_k} \psi^a(W_i,{\eta}_i)} ,
\end{equation}
which depends on sample splitting, i.e., the $K$-fold partition $\{\mathcal{I}_k : 1\le k \le K\}$. In contrast, the oracle version of the estimator based on DML2 is defined as
\begin{equation}
\hat{\theta}_{n,2}^* = \frac{n^{-1}\sum_{i=1}^n \psi^b(W_i,{\eta}_i)}{n^{-1}\sum_{i=1}^n \psi^a(W_i,{\eta}_i)} ,
\end{equation}
which does not depend on sample splitting.
Moreover, the oracle version of the DML2 estimator $\hat{\theta}_{n,2}^*$ is exactly the same as the one defined in (ref), but this is not typically the case for the oracle version of the DML1 estimator $\hat{\theta}_{n,1}^*$ with some exceptions. When $\psi^a(W_i,\eta_i)$ has zero variance, then both $\hat{\theta}_{n,1}^*$ and $\hat{\theta}_{n,2}^*$ are equal to the one defined in (ref).
remarkSimulation evidence has reported that increasing the number of folds $K$ improves the performance of the estimators based on DML2 in terms of bias and mean square error (ahrens2024ddml,ahrens2024model and chernozhukov2018double).
The cross-fitting procedure produces $K$ possible different estimators $\hat{\eta}_k(\cdot)$ for the nuisance parameter $\eta_0(\cdot)$ and each of them uses a fraction $(K-1)/K $ of the data. For instance, these estimators use $50\%$, $80\%$, and $90\%$ of the data when $K$ is 2, 5, and 10, respectively. Therefore, the accuracy of these estimators increases with the values of $K$. However, it is theoretically unknown if the improvement in the accuracy of the estimation of ${\eta}_0$ translated into more precise estimates for $\theta_0$.
Previous Results
Under some conditions (including $K$ fixed as $n \to \infty$), chernozhukov2018double showed that both DML estimators $\hat{\theta}_{n,1}$ and $\hat{\theta}_{n,2}$ have the same asymptotic distribution,
equation[equation omitted — 141 chars of source]
where the variance of the asymptotic distribution is given by
equation[equation omitted — 125 chars of source]
which only depends on the moment function $m$, the true nuisance parameter $\eta_0$, and the data distribution $F_0$. This result implies that the existing theoretical framework cannot distinguish between estimators based on DML1 and DML2, as discussed in the introduction. Moreover, this asymptotic theory provides no direct guidance for implementing DML.
The proof of (ref) relies on a first-order equivalent condition. More concretely, both DML estimators $\hat{\theta}_{n,1}$ and $\hat{\theta}_{n,2}$ are first-order equivalent to the oracle estimator $\hat{\theta}_n^*$, which means
equation[equation omitted — 160 chars of source]
This result is particularly useful since it implies that the estimation of $\theta_0$ using DML is as accurate as if the true $\eta_0$ had been used.
The first-order equivalence condition (ref) was obtained when $K$ is fixed as $n \to \infty$ by using (i) a Neyman orthogonality condition (which is necessary condition to obtain (ref); see Remark (ref)) and (ii) a conditional independence property due to construction of the estimates $\hat{\eta}_i$ using the cross-fitting procedure (e.g.,
conditional on $X_i$, the estimation error $\hat{\eta}_i - \eta_i$ and $W_i$ are independent). Importantly, the proof technique presented in chernozhukov2018double relies on $K$ being fixed as $n \to \infty$.
Although the existing asymptotic theory shows that DML1 and DML2 are asymptotically equivalent, DML2 is conjectured to perform better than DML1 based on simulation results about their relative performance. To investigate whether DML2 offers theoretical advantages, this paper considers an asymptotic framework where $K \to \infty$ as $n \to \infty$ in the next section. There, it will be shown that the accuracy (bias and MSE) of DML1 is sensitive to $K$, while this is not the case for DML2, implying that DML2 asymptotically dominates DML1 in terms of bias and MSE. Furthermore, it presents conditions under which setting $K = n$ minimizes bias and MSE for the DML2 estimators, suggesting that practitioners should implement DML2 with $K=n$ in settings well-approximated by these conditions.
Main Results
This section presents the asymptotic properties of estimators based on DML1 and DML2 when $K \to \infty$ as $n \to \infty$.
Assumptions
This section presents and discusses the conditions required for the moment function and the nuisance function estimators.
Assumption (ref) specifies the formal conditions on the moment function $m$ defined in (ref), including a (strong) Neyman orthogonality condition, while Assumption (ref) provides the details of the stochastic expansion satisfied by the nuisance parameter estimators. The technical conditions in parts (c) and (d) of Assumptions (ref) are stronger than in the existing first-order asymptotic theory of DML to address the technical difficulties that arise when $K \to \infty$ as $n \to \infty$. Assumption (ref) presents joint conditions on the moment function and the nuisance parameter estimators to conduct appropriate analysis of the leading terms of the higher-order bias and variance of the DML2 estimators.
The next assumption imposes conditions on the known functions $\psi^a$ and $\psi^b$, which define the moment function $m$. These conditions are presented below and depend on the following finite positive constants $M$, $C_0$, $C_1$, $C_2$, and $C_3$.
assumptionThe functions $\psi^a$ and $\psi^b$ are three-times continuously differentiable on $\eta \in \mathcal{T} \subseteq \mathbf{R}^p$ and satisfy for $z = a, b$,
\begin{enumerate}[(a)]
• $|E\left[ \psi^a(W_i,\eta_i)\right]|> C_0$.
• $ E\left[\partial_\eta \psi^z(W_i,\eta_i) \mid X_i \right] = 0~,~a.e.$
• $E\left[ \psi^z(W_i,\eta_i)^{4} \right] <M$ and $E\left[ ||\partial_\eta \psi^z(W_i,\eta_i)||^{4} \right] < M $.
• $|| E\left[ ( \partial_\eta \psi^z(W_i,\eta_i) ) ( \partial_\eta \psi^z(W_i,\eta_i) )^{\top} \mid X_i \right] ||_{\infty} \le C_1$.
• $\sup_{{\eta} \in \mathcal{T}} || \partial_\eta^2 \psi^z(W_i, \eta) ||_{\infty} \le C_2$ and $\sup_{{\eta} \in \mathcal{T}} || \partial_\eta^3 \psi^z(W_i, \eta) ||_{\infty} \le C_3$ for $z=a,b$.
\end{enumerate}
where $\partial_\eta \psi^z(W_i,\eta_i)$ is the partial derivative of $\psi^z$ with respect to $\eta$ evaluated at $\eta_i = \eta_0(X_i)$ and the $|| \cdot ||_{\infty}$-norm is the maximum of the absolute value of the matrix entries.
Part (a) of Assumption (ref) is an identification condition for the parameter $\theta_0$. It implies that $\theta_0$ can be written as a ratio of expected values as in (ref).
Part (b) of Assumption (ref) guarantees that a (strong) Neyman orthogonality condition holds for the moment function $m$ defined in (ref). That is,
equation[equation omitted — 118 chars of source]
The Neyman orthogonality condition is a necessary condition to guarantee first-order equivalence conditions, such as the one presented in (ref); see Remark (ref) for a further explanation. It has been used to remove the effects of the estimation of the nuisance parameter $\eta_0$ on the asymptotic distribution of estimators for $\theta_0$. Many estimators of the treatment parameters of interest have associated moment functions that satisfy this condition, including the ones in Examples (ref) (ATE), (ref) (ATT-DID), and (ref) (LATE).
Weaker forms of this condition have been used in the literature for a similar purpose; see, for instance, Assumption 5.1 in belloni2017program, Assumption 3.1 in chernozhukov2018double, and Equation (2.12) in andrews1994asymptotics. The Neyman orthogonality condition is helpful in studying the asymptotic properties of DML; however, it is not a restrictive requirement. Under certain conditions, it is possible to transform a moment function into a moment function that satisfies a Neyman orthogonality condition; see Remark (ref) for additional details.
Part (c) of Assumption (ref) is a regularity condition. It is used (i) to obtain a first-order equivalence property of the DML estimators and their oracle versions in Section (ref) and (ii) to guarantee that the high-order asymptotic approximation and quantities that appear in Section (ref) are well-defined.
Parts (d) and (e) of Assumption (ref) are mild technical conditions.
It is possible to set $C_3 = 0$ when the moment function is a quadratic polynomial in $\eta$ (the values of the nuisance parameter). This occurs when a doubly-robust moment condition defines the parameter of interest; see Theorem 4 in chernozhukov2022locally.
All the examples presented in Section (ref) and in Appendix (ref) satisfy Assumption (ref).
remarkWithout suitable assumptions about the moment functions, a first-order equivalent condition between a feasible estimator and its oracle version (as in (ref)) is not true in general. To see this, consider the following example. Suppose $\psi^a(W,\eta) = 1$ and $\psi^b(W,\eta)$ is a linear function in $\eta$. In addition, assume that the nuisance parameter $\eta_0$ is an unknown finite-dimensional parameter. Consider the estimator $\hat{\theta}_n = n^{-1} \sum_{i=1}^n \psi^b(W_i,\hat{\eta})$ and its oracle version $\hat{\theta}^*_n = n^{-1} \sum_{i=1}^n \psi^b(W_i,{\eta}_0)$, where $\hat{\eta}$ is an estimator of $\eta_0$ such that $n^{1/2}(\hat{\eta} - \eta_0) \overset{d}{\to} N(0,\Sigma)$ and $\Sigma$ is an invertible matrix. It can be shown that
$$ n^{1/2}(\hat{\theta}_n - \hat{\theta}_n^*) = n^{1/2} (\hat{\eta} - \eta_0)^\top E[ \partial_\eta m(W_i, \theta_0, \eta_0)] + o_p(1)~,$$
which implies $n^{1/2}(\hat{\theta}_n - \hat{\theta}_n^*) $ is $o_p(1)$ if and only if $E[ \partial_\eta m(W_i, \theta_0, \eta_0)] = 0$. In other words, the first-order equivalence condition in this example holds if and only if a Neyman orthogonality condition as in (ref) holds.
remarkMoment functions satisfying a Neyman orthogonality condition can be obtained by adding an adjustment term to the original moment functions. Specifically, under certain conditions on the nuisance parameter $\eta_0$, there exists $\alpha_0$ (function of covariates $X$) and $\phi(W,\theta,\eta,\alpha)$ (adjustment term) such that
\begin{itemize}
• $E[\phi(W,\theta,\eta_0,\alpha_0)] = 0 $
• the augmented moment function $\tilde{m}(W,\theta,\eta,\alpha) = m(W,\theta,\eta) + \phi(W,\theta,\eta,\alpha)$ satisfies a Neyman orthogonality condition
$$ E \left[ \partial_{\tilde{\eta}} \tilde{m}(W,\theta_0, \tilde{\eta}) |_{\tilde{\eta} = \tilde{\eta}_0(X)} \mid X \right] = 0~,\quad a.e. $$
where $\tilde{\eta}_0(X) = (\eta_0(X),\alpha_0(X))$.
\end{itemize}
The augmented term $\phi$ is an influence function of a particular parameter. It can be obtained using methods previously developed in the literature, such as ichimura2022influence and newey1994asymptotic. Recently, chernozhukov2022locally used the approach of ichimura2022influence and estimators based on DML2 to propose debiased GMM estimators.
Stochastic Expansion for the Nuisance Parameter Estimator
The next assumption imposes additional structure on the nuisance parameter estimators compared to the existing DML framework. These stronger conditions have a twofold purpose: addressing the technical challenges that arise when $K \to \infty$ as $n \to \infty$ in Section (ref) and allowing the analysis of higher-order properties in Section (ref).
To fix ideas, consider an estimator $\hat{\eta}$ of $\eta_0$ at a given point $x$ such that (i) $n^{-2\varphi_1}$ is the convergence rate of the variance of $\hat{\eta}$ and (ii) $n^{-\varphi_2}$ is the convergence rate of the bias of $\hat{\eta}$. To have more information about the variance and bias part, the next assumption imposes an asymptotic linear representation for both parts. More concretely, it takes as given the positive constants $\varphi_1$ and $\varphi_2$ and assumes the existence of two sequences of functions $\delta_n$ (for the variance part) and $b_n$ (for the bias part) that satisfy the asymptotic linear representation. Additional assumptions on these functions are imposed and depend on the following finite positive constants $M_1$, $C_\varphi$, and $C_b$, and a sequence of positive constants $\tau_{n}$ converging to zero (i.e., $\tau_{n} = o(1)$). Before continuing, let $n_0 = ((K-1)/K)n$ be the number of observations in the sample $\{W_i : i \notin \mathcal{I}_k\}$ used by $\hat{\eta}_k(\cdot)$ to estimate $\eta_0(\cdot)$.
assumptionThere exist two sequences of functions $\delta_{n_0}: \mathcal{W} \times \mathcal{X} \to \mathbf{R}^p$ and $b_{n_0}: \mathcal{X} \times \mathcal{X} \to \mathbf{R}^p$, such that
\begin{enumerate}[(a)]
• For any given $x \in \mathcal{X}$ and $k \in \{1,\ldots,K\}$,
$$ \hat{\eta}_k(x) - \eta_0(x) = n_0^{-1/2 } \sum_{\ell \notin \mathcal{I}_k} n_0^{-\varphi_1}\delta_{n_0}(W_{\ell},x) + n_0^{-1} \sum_{\ell \notin \mathcal{I}_k} n_0^{-\varphi_2} b_{n_0}(X_{\ell},x) + n_0^{-2 \min\{\varphi_1, \varphi_2\} } \hat{R}(x)~,$$
where $\hat{R}(x) = O_p(1)$, $E[\delta_{n_0}(W_{\ell},x) \mid X_{\ell}] = 0$ a.e., and $ E \left[ ||\delta_{n_0}(W_{\ell},x)||^{2} \right] > C_\delta$.
• For any $i \neq j$,
\begin{enumerate}[(b.1)]
• $E\left[ E[ ||\delta_{n_0}(W_j,X_i)||^2 \mid X_i ]^2 \right] \le M_1$ and $E \left[ E \left[|| n_0^{-\varphi_2} b_{n_0}(X_{j},X_i)||^2 \mid X_i \right ]^{2} \right] \le n_0^{2(1-2\varphi_1)} \tau_{n_0}$.
• $ E \left[ ||\delta_{n_0}(W_{j},X_i)||^{2s} \right] < n_0^{(s-1)(1-2\varphi_1)}M_1 $ for $s=1,2$ .
• $E \left[ ||E \left[b_{n_0}(X_{j},X_i) \mid X_i \right ]||^{4} \right] \in (C_b, M_1)$.
• $ E \left[ || n_0^{-\varphi_2}b_{n_0}(X_{j},X_i)||^{2s} \right] < n_0^{(2s-1)(1-2\varphi_1)} \tau_{n_0}$ for $s=1,2$.
\end{enumerate}
• $E[ ||\hat{R}(X_i)||^2 ] = O(1)$.
• $ n^{-1} \sum_{i=1}^n ||\hat{R}(X_i)||^4 = O_p(n^{4 \min\{ \varphi_1,\varphi_2 \} })$.
\end{enumerate}
Part (a) of Assumption (ref) present a stochastic expansion for the estimation error of the nuisance parameter estimator $\hat{\eta}_k(x)$. It assumes that this approximation has two terms that model the variance ($ n_0^{-1/2 } \sum_{\ell \notin \mathcal{I}_k} n_0^{-\varphi_1}\delta_{n_0}(W_{\ell},x)$) and bias ($n_0^{-1} \sum_{\ell \notin \mathcal{I}_k} n_0^{-\varphi_2} b_{n_0}(X_{\ell},x)$) components of the estimator $\hat{\eta}_k(x)$. Nonparametric kernel estimators satisfy this asymptotic expansion under mild regularity conditions on the nuisance parameter $\eta_0$. Appendix (ref) presents $\delta_{n_0}$ and $b_{n_0}$ for a class of nonparametric kernel estimators and the nuisance parameters $\eta_0$ (conditional expectations) that appear in the examples.
Part (b) of Assumption (ref) presents regularity conditions on $\delta_{n_0}(W_j,X_i)$ and $b_{n_0}(X_j,X_i)$ for $j \neq i$. Part (b.1) is helpful to establish that the leading terms in the stochastic approximation for the estimation error of $\hat{\eta}_k$ have a finite fourth moment; see Lemma (ref) in Appendix (ref).
Part (b.2) and the condition $ E \left[ ||\delta_{n_0}(W_{\ell},x)||^{2} \right] > C_\delta$ in part (a) guarantee that the variance of the nuisance parameter estimator has a convergence rate $O(n^{-2\varphi_1})$, while part (b.3) establishes that the bias term has a convergence rate $O(n^{-\varphi_2})$. Finally, part (b.4) considers additional regularity conditions. These regularity conditions can be verified for Nadaraya-Watson estimators and the nuisance parameters $\eta_0$ (conditional expectations) that appear in the examples.
Parts (c) and (d) of Assumption (ref) are helpful high-level conditions to establish results in Section (ref), where the asymptotic framework considers $K \to \infty$ as $n \to \infty$. This assumption can be verified for Nadaraya-Watson estimators under additional suitable conditions on the nuisance parameter $\eta_0$ (conditional expectations) that appear in the examples. Part (c) is sufficient to guarantee that $E[||\hat{\eta}_i-\eta_i||^2]$ is finite and has convergence rate $O(n^{-2\varphi_1}) + O(n^{-2\varphi_2})$. Part (d) is sufficient to guarantee that $n^{-1} \sum_{i=1}^n || \hat{\eta}_i - \eta_i||^4$ is $O_p(n^{-4\min\{\varphi_1,\varphi_2\}})$. Both intermediate results are formally established in Lemma (ref) in Appendix (ref).
Stochastic Expansion for the DML2 Estimator
Finally, the next assumption imposes joint conditions on the functions $\delta_{n_0}$ and $b_{n_0}$, defined in Assumption (ref), and the moment functions $m$, defined in (ref). These conditions are important to (i) derive a valid stochastic expansion for the estimators based on DML2 and (ii) conduct an appropriate analysis of the leading terms of the higher-order bias and variance of the DML2 estimator. Before continuing, define $\tilde{b}_{n_0}(X_i) = E\left[ b_{n_0}(X_j,X_i) \mid X_i \right]$ for $j \neq i$.
assumption\begin{enumerate}[(a)]
• The following limits exist and are finite,
\begin{align}
G_\delta &= \lim_{n_0 \to \infty} E\left[ E\left[ \delta_{n_0}(W_{j},X_{i})^\top \left(\partial_\eta^2 m(W_i,\theta_0,\eta_i) \right) \delta_{n_0}(W_{\ell},X_{i}) \mid W_{j}, W_{\ell} \right]^2 \right] /E[\psi^a(W_i,\eta_i)]^2, \\
F_\delta &= \lim_{n_0 \to \infty} \frac{1}{2} E\left[ \delta_{n_0}(W_j,X_i)^\top\left(\partial_\eta^2 m(W_i,\theta_0,\eta_i) \right) \delta_{n_0}(W_j,X_i) \right] /E[\psi^a(W_i,\eta_i)] \\
F_b &= \lim_{n_0 \to \infty} \frac{1}{2} E\left[ \tilde{b}_{n_0}(X_i)^\top \left(\partial_\eta^2 m(W_i,\theta_0,\eta_i) \right) \tilde{b}_{n_0}(X_i) \right]/E[\psi^a(W_i,\eta_i)] , \\
G_b &= \lim_{n_0 \to \infty} E\left[ m(W_j,\theta_0,\eta_j) \delta_{n_0}(W_j,X_i)^\top \left(\partial_\eta^2 m(W_i,\theta_0,\eta_i) \right) \tilde{b}_{n_0}(X_i) \right] /E[\psi^a(W_i,\eta_i)]^2 ,
\end{align}
• $G_\delta >0$, $F_\delta + F_b \neq 0$, and $G_b \neq 0$.
\end{enumerate}
Part (a) of Assumption (ref) is a regularity condition. Assumptions (ref) and (ref) ensure that the sequences on the right-hand side of (ref)--(ref) are bounded, implying that, if the limit exists, it is finite. Whether the limits in (ref)--(ref) exist depends on the estimator of the nuisance function, $\eta_0$. For instance, this can be verified for Nadaraya-Watson estimators ---using the same kernel as $n_0 \to \infty$--- under additional suitable conditions on $\eta_0$. The limit may not exist when $\eta_0$ is estimated by two different types of estimators based on $n_0$, which is the number of observations in the estimation of $\eta_0$. For instance, if the estimator of $\eta_0$ uses a given kernel when $n_0$ is odd and a different one when $n_0$ is even, the limit can exist for each subsequence, but they may be different.
Part (b) of Assumption (ref) is a sufficient condition to conduct an appropriate analysis of the leading terms of the higher-order bias and variance of the DML2 estimators in Section (ref). Determining whether part (b) of Assumption (ref) hold may require a case-by-case analysis, as it depends on the interaction between the second-order partial derivatives of the moment function $m$ with respect to $\eta$ and the estimator of the nuisance function. Importantly, a necessary condition is that the moment function $m$ is a nonlinear function on $\eta$, i.e., the matrix of second-order partial derivatives of $m$ with respect to $\eta$ is different than zero.
For Examples (ref) (ATE), (ref) (ATT-DID), (ref) (ATT), and (ref) (PLM), it can be verified that $G_\delta > 0 $, $|F_\delta + F_b|>0$, and $G_b \neq 0$ when $\eta_0$ is estimated using Nadaraya-Watson estimators.
A First-Order Asymptotic Theory when $K$ increases
This section presents the asymptotic distribution of the estimators based on DML1 and DML2 in an asymptotic framework where the number of folds $K$ increases with the sample size $n$. Specifically, it shows that DML1 may exhibit a first-order asymptotic bias (i.e., its asymptotic distribution may not be centered at zero), which is not the case for DML2.
These results show that DML2 offers theoretical advantages over DML1 in terms of bias and MSE. Furthermore, the conditions used in this section guarantee that DML2 remains asymptotically valid for any $K \le n$.
The asymptotic framework considered in this section for DML is, to the best of my knowledge, new and approximates finite sample situations often faced by practitioners. For instance, small-sample situations where the practitioner implementing DML desires to increase $K$ to improve the precision of the nuisance parameter estimators $\hat{\eta}_k$'s, which use a fraction $(K-1)/K$ of the data; see Remark (ref). This finite sample situation is not well approximated by the available asymptotic framework in the literature, which considers that $K$ is fixed as $n \to \infty$.
The next results present the asymptotic distributions of the estimators based on DML1 and DML2 when the number of folds $K \to \infty$ as $n \to \infty$.
theoremSuppose Assumptions (ref) and (ref) hold. In addition, assume $K$ is such that $K \le n$, $K \to \infty $ and $K/\sqrt{n} \to c \in [0,\infty)$ as $n \to \infty$. If $\varphi_1 \le 1/2$ and $1/4 < \min \{\varphi_1, \varphi_2 \}$, then
$$ n^{1/2}\left( \hat{\theta}_{n,1} - \theta_0 \right) \overset{d}{\to} N(c \Lambda,\sigma^2) ~,$$
where $\hat{\theta}_{n,1}$ and $\sigma^2$ are as in (ref), (ref), respectively, and
\begin{equation}
\Lambda = \frac{Cov\left[m(W_i,\theta_0,\eta_i) , -\psi^a(W_i,\eta_i)\right]}{E[\psi^a(W_i,\eta_i)]^2} .
\end{equation}
theoremSuppose Assumptions (ref) and (ref) hold. In addition, assume $K$ is such that $K \le n$, $K \to \infty $ as $n \to \infty$. If $\varphi_1 \le 1/2$ and $1/4 < \min \{\varphi_1, \varphi_2 \}$, then
$$ n^{1/2}\left( \hat{\theta}_{n,2} - \theta_0 \right) \overset{d}{\to} N(0,\sigma^2) ~, $$
where $\hat{\theta}_{n,2}$ and $\sigma^2$ are as in (ref) and (ref), respectively.
Theorems (ref) and (ref) explain why DML1 and DML2 behave similarly in simulations for many models previously studied in the literature, including the ones presented in Example (ref) (ATE), Example (ref) (ATT-DID), and Example (ref) (PLM). All these examples have $\Lambda = 0$, and, therefore, by these theorems, it follows that both estimators have the same asymptotic distribution whenever $K \to \infty$ slowly as $n \to \infty$ (i.e., $K = O( n^{1/2} ) $).
Theorem (ref) shows that (first-order) asymptotic properties of DML1 can be sensitive to $K$ when $\Lambda \neq 0$. Intuitively, this theorem shows the distribution of
$ n^{1/2}\left( \hat{\theta}_{n,1} - \theta_0 \right)$ can be approximated by
$ N( \Lambda K/\sqrt{n} ,\sigma^2)$, which is sensitive to the choice of $K$ when $n$ is small. In particular, when $K \sim \sqrt{n}$ and $\Lambda \neq 0$, the asymptotic distribution of estimators based on DML1 is not centered at zero, i.e., there is an asymptotic bias proportional to $\Lambda$ affecting the reliability of the inference procedure and the accuracy of the estimation. Some examples where $\Lambda$ is typically nonzero are Example (ref) (LATE) and Example (ref) (w-ATE).
The previous implications make clear that $\Lambda$ is a discrepancy measure between DML1 and DML2. More concretely, when $\Lambda$ equals zero, DML1 and DML2 exhibit similar first-order asymptotic properties, but as $\Lambda$ deviates from zero, DML1 becomes more sensitive to large values of $K$ in terms of bias and MSE, while this is not the case of DML2. See Remark (ref) for additional discussion on why DML1 has asymptotic bias proportional to $\Lambda$.
The conditions presented in Theorem (ref) guarantee that estimators based on DML2 are asymptotically valid for any $K \le n$. More concretely, under the conditions presented in this theorem, all the DML2 estimators using different $K$ share the same asymptotic distribution $N(0,\sigma^2)$. This class of DML2 estimators includes the leave-one-out estimator defined by setting $K=n$. These results have practical consequences for the practitioner that desires to increase $K$ to improve the precision of the nuisance parameter estimators. This theorem indicates that these DML2 estimators obtained by increasing $K$ are asymptotically valid.
Finally, the proofs of Theorems (ref) and (ref) rely on two intermediate results. The first intermediate result is presented in Theorem (ref). It states that a first-order equivalence condition holds for the estimators based on DML and their oracle versions defined in Remark (ref) when $K \to \infty$ as $n \to \infty$. The second intermediate result calculates the asymptotic distribution of the oracle version of the DML estimators defined in Remark (ref). This result is presented in Proposition (ref). It can be shown that $\Lambda$ appears as an asymptotic bias for DML1 because the oracle version of the DML1 estimators depends on sample splitting, while this is not the case for the oracle version of the DML2 estimators.
remarkIf $K/\sqrt{n} \to \infty$ and $\Lambda > 0$, then the estimators based on DML1 may have a degenerate asymptotic distribution.
More concretely, when $K \sim n^{1/2+\epsilon}$ for some sufficiently small $\epsilon>0$, it is possible to guarantee that (i) $n^{1/2}\left( \hat{\theta}_{n,1} - \hat{\theta}_{n,1}^* \right) = o_p(1)$ by extending Theorem (ref) and Lemma (ref), and (ii) $ n^{1/2}\left( \hat{\theta}_{n,1}^* - \theta_0 \right)$ converge to $\infty$ with probability approaching one, which follows by the proof of Proposition (ref). Therefore, $ n^{1/2}\left( \hat{\theta}_{n,1} - \theta_0 \right)$ can converge to $\infty$, which is a degenerate distribution.
remarkThe first-order equivalence condition between $\hat{\theta}_{n,j}$ and its oracle version $\hat{\theta}_{n,j}^*$ relies on stronger assumptions compared to the existing DML framework to accommodate that $K \to \infty$ and $n \to \infty$. These assumptions are presented in parts (c) and (d) of Assumption (ref), and they are used to prove two important intermediate results that appear in Lemma (ref),
\begin{equation}
n^{-1/2} \sum_{i=1}^n (\hat{\eta}_i - \eta_i)^\top \partial_\eta m (W_i,\theta_0,\eta_i) = o_p(1) ,
\end{equation}
and in Lemma (ref),
\begin{equation}
\max_{k=1,\ldots,K} \left|n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{\eta}_i - \eta_i)^\top \partial_\eta m(W_i, \theta_0, \eta_i) \right|= o_p(1) .
\end{equation}
When $K$ is fixed as $n \to \infty$, the previous intermediate results, (ref) and (ref), follow from a Bonferroni correction argument and
\begin{equation}
\left|n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{\eta}_i - \eta_i)^\top \partial_\eta m(W_i, \theta_0, \eta_i) \right|= o_p(1) .
\end{equation}
The proof of (ref) when $K$ is fixed follows by (i) a Neyman orthogonality condition and (ii) a conditional independence property due to the construction of the estimates $\hat{\eta}_i$ using cross-fitting (e.g.,
conditional on $X_i$, the estimation error $\hat{\eta}_i - \eta_i$ and $W_i$ are independent). However, this proof cannot be adapted to the case where $K \to \infty$ as $n \to \infty$.
remarkThe condition (ref) is important to obtain asymptotically valid estimators with the same asymptotic distribution as the oracle estimator in (ref). Remark (ref) pointed out this is the case for DML estimators. Equivalent formulations of (ref) as a high-level condition have been used in the literature to establish a first-order equivalence condition between a plug-in estimator---defined in Section (ref)--- and its oracle version (e.g., andrews1994asymptotics, farrell2015robust).
In general, the verification of (ref) for plug-in estimators is difficult since it is unclear whether $E[ (\hat{\eta}_i - \eta_i)^\top \partial_\eta m (W_i,\theta_0,\eta_i)]$ is zero (or asymptotically zero) due to the correlation between $\hat{\eta}_i - \eta_i$ and $\partial_\eta m (W_i,\theta_0,\eta_i)$, which is a manifestation of the “own observation” bias that arise because the same data is used to estimate $\eta_0$ and $\theta_0$.
remarkThe discrepancy measure $\Lambda$ is proportional to the first-order asymptotic bias of the DML1 estimator because its oracle version depends on sample splitting, which is not the case for the DML2 estimator. For the sake of explanation, suppose the nuisance parameter is known. In this case, the oracle estimator $\hat{\theta}_{n,1}^*$ defined in Remark (ref) is equal to the average of $K$ preliminary estimators $\tilde{\theta}_k^*$, that is
$$ \hat{\theta}_{n,1}^* = K^{-1} \sum_{k=1}^K \tilde{\theta}_k^* ~,$$
where each $\tilde{\theta}_k^*$ is as in (ref) but using only observations in the fold $\mathcal{I}_k$, which has $n/K$ observations. Therefore, each of these preliminary estimators has a (higher-order) asymptotic bias equal to $\Lambda (n/k)^{-1}$ since it uses $n/K$ observations. An explicit expression for $\Lambda$ can be obtained based on standard arguments (e.g., newey2004higher). Since the bias of the average of the estimators is the same as the average of the bias of the estimators, it follows that the (higher-order) asymptotic bias of $\hat{\theta}_{n,1}^*$ is $\Lambda K/n$. Intuitively, when $K \sim \sqrt{n}$, this asymptotic bias become proportional to $\Lambda / \sqrt{n}$ and shows up in the first-order asymptotic distribution of the oracle estimator $\hat{\theta}_{n,1}^*$. Finally, since the feasible DML1 estimator $\hat{\theta}_{n,1}$ and $\hat{\theta}_{n,1}^*$ are first-order equivalent, their asymptotic distributions are the same.
High-Order Asymptotic Theory for DML2 estimators
This section presents higher-order asymptotic properties (e.g., bias, variance, and MSE) of the estimators based on DML2 in an asymptotic framework where the number of folds $K$ increases with the sample size $n$. Specifically, it shows that the leading term of the higher-order bias is a decreasing function of $K$.
Moreover, it presents conditions under which setting $K$ equals $n$ minimizes the second-order MSE of DML2 estimators.
The goal of this section is to propose asymptotic approximations that offer a better description of the finite sample behavior of the DML2 estimators. The approximations based on the first-order asymptotic distribution are insufficient for this goal since all the DML2 estimators share the same asymptotic distribution for any $K \le n$ (Theorem (ref)). In this section, I obtain better approximations by considering stochastic expansions up to a smaller remainder error term than in the existing first-order asymptotic theory.
The main idea is to use these better approximations to study the higher-order asymptotic properties of the estimators based on DML2, with the hope that they are reliable enough to explain the finite sample behavior of the estimators. This approach has a long history in econometrics to compare estimators that are first-order equivalent (e.g., rothenberg1984approximating, linton1995second, newey2004higher, graham2012inverse).
Stochastic Expansions for DML2 estimators
The next result presents a stochastic expansion for the estimators based on DML2 that is asymptotically valid when $K$ increases with the sample size $n$. This stochastic expansion focuses on the nuisance parameter estimators satisfying Assumption (ref) with $\varphi_1 = \varphi_2$; see Remark (ref) for additional discussion of other cases. For the remainder of this section, I set $\varphi = \varphi_1 = \varphi_2$.
theoremSuppose Assumptions (ref), (ref), and (ref) hold. In addition, assume $K$ is such that $K \le n$ and $K \to \infty$ as $n \to \infty$. If $\varphi \in (1/4, 1/2)$, then
\begin{equation}
n^{1/2} (\hat{\theta}_{n,2} - \theta_0) = \mathcal{T}_{n}^* + \mathcal{T}_{n,K}^{nl} + R_{n,K} ,
\end{equation}
where $\hat{\theta}_{n,2}$ is as (ref), $\mathcal{T}_{n}^*$ is defined in (ref) and $ \mathcal{T}_{n}^* \overset{d}{\to} N(0,\sigma^2)$ as $n \to \infty$ with $\sigma^2$ defined as in (ref), $T_{n,K}^{nl}$ is defined in (ref) and satisfies (i) $ \lim_{n \to \infty} \inf_{K \le n} Var[ n^{2\varphi-1/2} T_{n,K}^{nl}] > 0$ and (ii) $ \lim_{n \to \infty} \sup_{K \le n} E\left[ \left(n^{2\varphi-1/2} T_{n,K}^{nl} \right)^2 \right] < \infty $, and
\begin{align*}
\lim_{n \to \infty} \sup_{K \le n} &P( n^{ 2\varphi - 1/2}|R_{n,K}| > \epsilon) = 0 ,
\end{align*}
for any given $\epsilon >0$.
Theorem (ref) presents a stochastic expansion more accurate than the available first-order asymptotic theory for any given sequence $K \to \infty$ as $n \to \infty$. More concretely, equation (ref) presents a remainder error term $R_{n,k}$ that is stochastically smaller than $ T_{n,K}^{nl}$, i.e., $n^{2\varphi-1/2} R_{n,k}$ converges to zero in probability for any sequence $K \to \infty$, while this is not the case for $n^{2\varphi-1/2} T_{n,k}^{nl}$ since its variance is positive for any $n$ sufficiently large. Therefore, under the conditions of this theorem, $R_{n,K}$ in (ref) is smaller than the remainder error term $R_{n,K}^*= T_{n,K}^{nl}+R_{n,K}$ obtained by the first-order approximation, denoted by $\mathcal{T}_{n}^*$,
$$ n^{1/2} (\hat{\theta}_{n,2} - \theta_0) = \mathcal{T}_{n}^* + R_{n,K}^*~, $$
where
equation[equation omitted — 147 chars of source]
The approximation in Theorem (ref) includes the additional term $\mathcal{T}_{n,K}^{nl}$ to accommodate for the errors of the nuisance parameter estimators. More concretely, under the conditions of Theorem (ref), $\mathcal{T}_{n,K}^{nl}$ is defined as the leading term in the scaled difference between the feasible estimator $\hat{\theta}_{n,2}$ and oracle estimators $\hat{\theta}_{n,2}^*$ defined in Remark (ref),
$$ n^{1/2} \left( \hat{\theta}_{n,2} - \hat{\theta}_{n,2}^*\right) = \mathcal{T}_{n,K}^{nl} + R_{n,K}^{nl}~,$$
where
$$\lim_{n \to \infty} \sup_{K \le n} P( n^{ 2\varphi - 1/2}|R_{n,K}^{nl}| > \epsilon) = 0$$
and
equation[equation omitted — 264 chars of source]
with $\Delta_i $ defined below for $i \in \mathcal{I}_k$,
equation[equation omitted — 222 chars of source]
where $\delta_{n_0}$ and $b_{n_0}$ are the functions in Assumption (ref). The approximation of the scaled difference between $\hat{\theta}_{n,2}$ and $\hat{\theta}_{n,2}^*$ is obtained by using Taylor expansions to approximate $\psi^z(W_i,\hat{\eta}_i)$ by $\psi^z(W_i,\eta_i)$ for $z=a,b$.
remarkWhen $K$ is fixed as $n \to \infty$, a similar stochastic expansion can be derived for DML1,
\begin{equation*}
n^{1/2} (\hat{\theta}_{n,1} - \theta_0) = \mathcal{T}_{n}^* + \mathcal{T}_{n,K}^{nl} + o_p(n^{1/2-2\varphi}) .
\end{equation*}
Furthermore, the stochastic expansion in (ref) for DML2 remains valid when $K$ is fixed as $n \to \infty$. These expressions show that when $K$ is fixed and $\varphi = \varphi_1 = \varphi_2$, the two leading terms in the stochastic approximation are the same.
remarkTheorem (ref) presents a stochastic expansion for the case $ \varphi = \varphi_1 = \varphi_2$, representing situations where the nuisance function estimator balances bias and variance. For example, this occurs when a bandwidth with an optimal convergence rate is used in Nadaraya-Watson estimators to estimate the nuisance function. Appendix (ref) discusses the case where $\varphi_1 < \varphi_2$, which represents situations where the bias of the nuisance function estimator converges faster than the variance components (e.g., undersmoothing). Theorem (ref) extends Theorem (ref) by providing the stochastic expansion for the alternative cases of $(\varphi_1,\varphi_2)$.
High-Order Asymptotic Properties
I calculate the higher-order asymptotic bias, variance, and mean square error (MSE) of the estimator $\hat{\theta}_{n,2}$ by using the asymptotic approximation $\mathcal{T}_{n,K}$ defined next,
equation[equation omitted — 117 chars of source]
More concretely, the higher-order bias, variance, and MSE of $\hat{\theta}_{n,2}$ are respectively defined as
align*[align* omitted — 246 chars of source]
Theorems (ref) and (ref), and Corollary (ref) present explicit expressions for the leading terms of $E[\mathcal{T}_{n,K}]$, $\text{Var}[\mathcal{T}_{n,K}]$, and $E[\mathcal{T}_{n,K}^2]$.
Similar definitions of higher-order asymptotic bias and variance have been used to compare alternative estimators with the same asymptotic distribution, including rothenberg1984approximating, linton1995second, and newey2004higher. As discussed in rothenberg1984approximating, these definitions are valid asymptotic approximations of the bias and variance of the estimators whenever additional regularity conditions hold.
An alternative interpretation discussed in linton1995second suggests that these definitions could be interpreted as a form of approximations of the bias and variance of $\hat{\theta}_{n,2}$ since they are based on the moments of the approximation $\mathcal{T}_{n,K}$, which has a distribution that asymptotically approximates the distribution of $ n^{1/2} (\hat{\theta}_{n,2} - \theta_0)$ up to an error $o(n^{1/2-2\varphi})$ under certain regularity conditions.
The next theorem presents the expected value and variance of $\mathcal{T}_{n,K}$, which can be used to calculate the higher-order bias and variance of the DML2 estimator.
theorem[Higher-Order Bias]
Suppose Assumptions (ref), (ref), and (ref) hold. In addition, assume $K$ is such that $K \le n$ and $K \to \infty$ as $n \to \infty$. If $\varphi \in (1/4, 1/2)$, then
\begin{equation*}
E[\mathcal{T}_{n,K}]
= (F_\delta + F_b) \left( 1 + \frac{1}{K-1} \right)^{2\varphi} n^{1/2-2\varphi} + \nu_{n,K} ,
\end{equation*}
where
$$ \sup_{K \le n } | \nu_{n,K} | = o(n^{1/2-2\varphi})~,$$
with $\mathcal{T}_{n,K}$, $F_{\delta}$ and $F_{b}$ defined as in (ref), (ref) and (ref), respectively.
theorem[Higher-Order Variance]
Suppose Assumptions (ref), (ref), and (ref) hold. In addition, assume $K$ is such that $K \le n$ and $K \to \infty$ as $n \to \infty$. If $\varphi \in (1/4, 1/2)$, then
\begin{equation*}
Var[\mathcal{T}_{n,K}] = \sigma^2 + G_{b} \left( 1 + \frac{1}{K-1} \right)^{2\varphi-1/2} n^{1/2-2\varphi} + r_{n,K} ,
\end{equation*}
where
$$ \sup_{K \le n} |r_{n,K}| = o(n^{1/2-2\varphi})~,$$
with $\mathcal{T}_{n,K}$, $\sigma^2$, and $G_b$ defined as in (ref), (ref), and (ref), respectively.
Theorems (ref) and (ref) can be used to find the higher-order bias and variance of $\hat{\theta}_{n,2}$,
equation*[equation* omitted — 155 chars of source]
and
equation*[equation* omitted — 166 chars of source]
The leading term of the higher-order bias depends on $F_\delta$ and $F_b$ defined in (ref) and (ref), respectively. While the second leading term of the higher-order variance depends on $G_b$ defined in (ref). These are quantities that depend on three elements: (i) the functions $\delta_{n_0}$ and $b_{n_0}$ that appear in the stochastic expansion for the nuisance parameter estimator presented in Assumption (ref), (ii) the second-order derivatives of the moment function $m$ with respect to $\eta$, and (iii) the data distribution.
The previous expression reveals that the absolute value of the leading term in the higher-order bias, $|F_\delta + F_b|(1+1/(K-1))^{2\varphi} n^{-2\varphi}$, decreases as $K$ increases. Therefore, the leave-one-out estimator, defined as the DML2 estimator with $K=n$, minimizes the absolute value of the leading term in the higher-order asymptotic bias. The two leading terms of the higher-order variance, $\sigma^2 n^{-1}$ and $G_b (K/(K-1))^{2\varphi-1/2}n^{-1/2 -2\varphi}$, define a decreasing function of $K$ when $G_b >0$. Similarly, when $G_b>0$, the leave-one-out minimize the two leading terms of the higher-order variance.
The next result is a corollary derived from Theorems (ref) and (ref). It presents the second moment of $\mathcal{T}_{n,K}$, which can be used to calculate the higher-order MSE of the DML2 estimator.
corollarySuppose the conditions of Theorems (ref) and (ref) holds. Then,
\begin{equation*}
E[ \mathcal{T}_{n,K}^2] = \sigma^2 + G_{b} \left( 1 + \frac{1}{K-1} \right)^{2\varphi-1/2} n^{1/2-2\varphi} + \tilde{r}_{n,K} ,
\end{equation*}
where
$$ \sup_{K \le n} |\tilde{r}_{n,K}| = o(n^{1/2-2\varphi})~,$$
with $\mathcal{T}_{n,K}$, $\sigma^2$ and $G_b$ defined as in (ref), (ref) and (ref), respectively.
Corollary (ref) can be used to find the higher-order MSE,
$$ \text{HO-MSE}[\hat{\theta}_{n,2}] = \sigma^2n^{-1} + G_{b} \left( 1 + \frac{1}{K-1} \right)^{2\varphi-1/2} n^{-1/2-2\varphi} + \tilde{r}_{n,K} n^{-1}~.$$
The two leading terms of $\text{HO-MSE}[ \hat{\theta}_{n,2} ]$ define the second-order asymptotic MSE,
equation*[equation* omitted — 150 chars of source]
which is a decreasing function on $K$ when $G_b$ is positive. Therefore, the leave-one-out estimator can minimize the second-order asymptotic MSE by setting $K = n$ when $G_b >0$.
remarkWhen $K$ is fixed as $n \to \infty$, it can be shown that the two leading terms of the higher-order MSE, $\sigma^2/n $ and $ G_{b} \left( 1 + 1/(K-1) \right)^{2\varphi-1/2} n^{-1/2-2\varphi} $, are the same for DML1 and DML2 up to an error of size $o(n^{-1/2-2\varphi})$. For this conclusion is important that $K$ is fixed as $n \to \infty$, since some higher-order terms in the asymptotic approximation for DML1 may become large terms for larger values of $K$. Remark (ref) illustrate this point by considering the higher-order MSE of the oracle version of the DML1 and DML2 estimators defined in Remark (ref).
remarkWhen $K$ is fixed as $n \to \infty$, an explicit expression for the higher-order MSE of the oracle estimators $\hat{\theta}_{n,1}^*$ and $\hat{\theta}_{n,2}^*$ defined in Remark (ref) can be derived based on standard arguments (e.g., newey2004higher):
\begin{align*}
HO-MSE[\hat{\theta}_{n,1}^*] &= \sigma^2/n + \left( K^2 \Lambda^2 + K \Lambda_1 \right)/n^2 + o(n^{-2}) \\
HO-MSE[\hat{\theta}_{n,2}^*] &= \sigma^2/n + \left( \Lambda + \Lambda_1 \right)/n^2 + o(n^{-2})
\end{align*}
where
\begin{equation*}
\Lambda_1 = 5\Lambda^2 + \sigma^2 \left\{ 3\frac{E\left[ \psi^a(W,\eta_0(X)) ^2 \right] }{E\left[ \psi^a(W,\eta_0(X)) \right]^2} - 1 \right \}- 2 \frac{ E\left[m(W,\theta_0,\eta_0(X))^2 \psi^a(W,\eta_0(X)) \right]}{E\left[ \psi^a(W,\eta_0(X)) \right]^3} ,
\end{equation*}
with $\sigma^2$ and $\Lambda$ defined as in (ref) and (ref), respectively.
Two main differences with respect to the results in Corollary (ref) deserve further discussion. First, the remainder errors for the oracle versions are $o(n^{-2})$ and the second leading terms have a convergence rate of order $n^{-2}$, implying these terms are smaller than the second leading term of the higher-order MSE of $\hat{\theta}_{n,2}$ that is of order $n^{-1/2-2\varphi}$. Therefore, the second leading term that appears for the oracle estimators is a higher-order term included in $o(n^{-1/2-2\varphi})$. Second, the second leading term, $\left( K^2 \Lambda + K \Lambda_1 \right)/n^2$, in the higher-order MSE of $\hat{\theta}_{n,1}^*$ depends on $K$; therefore, for large values of $K$ the accuracy of $\hat{\theta}_{n,1}^*$ is worse than the accuracy of $\hat{\theta}_{n,2}^*$.
Lessons for Practitioners
This section presents some lessons for practitioners to implement DML based on the formal results of Section (ref). These lessons have theoretical support and, to the best of my knowledge, are new in the literature of DML. Furthermore, these lessons are possible due to the asymptotic framework considered in this paper, providing insights not captured by the existing first-order asymptotic theory or simulation-based evidence.
Before presenting them, it is important to remember that DML provides estimators as good as if the true nuisance function $\eta_0(\cdot)$ has been used. However, DML provides a wide array of alternatives to practitioners that may seem roughly equivalent. Among these alternatives, there are two available estimators, DML1 and DML2, both presented in Section (ref). In addition, each estimator depends on the number of equal-sized folds $K$ in which the data is randomly split. In what follows, I present and discuss the recommendations for implementing DML, including how to select $K$.
First lesson: DML2 is the recommended option for implementing DML, especially in small-sample situations when increasing the number of folds is desired to improve the precision of the estimators $\hat{\eta}_k(\cdot)$, which use a fraction $(K-1)/K$ of the sample size $n$. This recommendation is not new ---it was presented in chernozhukov2018double---, but now it has a theoretical justification in terms of bias and MSE. Results in Section (ref) show that the asymptotic distribution of DML2 is insensitive in terms of bias and MSE to the values of $K$, which is not the case of DML1 ---which becomes increasingly sensitive in terms of bias and MSE to large values of $K$ whenever the discrepancy measure $\Lambda$ ---defined in (ref)--- deviates from zero. Moreover, the conditions presented in Section (ref) guarantee that estimators based on DML2 are asymptotically valid for any $K \le n$, including the leave-one-out estimator defined as DML2 with $K=n$.
Number of Folds for Cross-Fitting
The previous lesson recommends the use of DML2 estimators, but it is not clear how to choose the number of folds since the results of Section (ref) show that all DML2 estimators share the same asymptotic distribution. This question is addressed in what follows by considering two different criterion based on the (higher-order) asymptotic bias and the second-order MSE. I use the explicit formulas presented in Section (ref), where the higher-order properties of DML2 estimators were studied when $K$ is fixed or increases with the sample size $n$.
Second lesson: Choosing the number of folds equal to the sample size to implement DML2 is asymptotically optimal to reduce the (higher-order) asymptotic bias. The explicit formulas presented in Section (ref) show that the absolute value of the leading term of the higher-order asymptotic bias for DML2 estimators is decreasing on $K$. For convenience, this explicit formula is presented next,
equation[equation omitted — 122 chars of source]
where $\varphi \in (1/4,1/2)$.
Figure (ref) presents this explicit formula scaled by $\sqrt{n}$ as a function of the number of folds $K$, which appears as a blue line with circular markers, and the optimal level (minimal) obtained when $K=n$, which appears as a constant red line with square markers. Figure (ref) uses $n =1,000 $, $F=1$ and $\varphi = 2/5$ for illustrative purposes.
Therefore, choosing $K=n$ to implement DML2 is optimal to reduce the asymptotic bias, while the common recommendations of choosing $K=5, 10$ or $20$ is suboptimal. Finally, Figure (ref) reveals that the discrepancy between the asymptotic bias with the common recommendations for $K$ and the optimal choice can be small. However, this conclusion may depend on the values of $n$, $F$, and $\varphi$ that Figure (ref) is using. In the fourth lesson, I discuss the relative loss of the asymptotic bias with respect to the optimal choice for the common choice of $K$ for arbitrary values of $F$ and $\varphi$.
figure[figure omitted — 241 chars of source]
Third lesson: Choosing the number of folds equal to the sample size to implement DML2 can be asymptotically optimal to reduce the second-order asymptotic MSE. In other words, the leave-one-estimator defined as the DML2 estimator using $K = n$ can be the most asymptotically accurate estimator among the class of DML2 estimators when a certain data-dependent condition holds.
To illustrate this lesson, I present next the second-order asymptotic MSE when the variance and the bias of the nuisance parameter estimator have the same convergence rate (i.e., $\varphi = \varphi_1 = \varphi_2$), that is
equation[equation omitted — 172 chars of source]
where $\varphi \in (0,1/2)$ and $G_b$ is a complex object that depends on the data. Note that $\text{SO-MSE of } \hat{\theta}_{n,2}$ in (ref) is a decreasing function of $K$ whenever $G_b$ is positive, which is the data-dependent condition mentioned above. Figure (ref) presents the second-order MSE scaled by $n$ as a function of the number of folds $K$, which appears as a blue line with circular markers, and the optimal second-order MSE obtained by $K=n$ (under the assumption that $G_b >0$), which appears as a constant red line with square markers. Figure (ref) uses $n = 1,000$, $\sigma = 1$, $G_b = 1$, and $\varphi = 2/5$ for illustrative purposes.
figure[figure omitted — 257 chars of source]
Therefore, choosing $K=n$ to implement DML2 is optimal to reduce the second-order MSE, while the common recommendations of $K=5, 10,$ or $20$ are suboptimal under this criterion. Finally, the relative loss of the second-order MSE with respect to the optimal choice looks small for $K \ge 10$. However, a general conclusion may depend on the values of the parameters ($n$, $\sigma$, $G_b$, $\varphi$). In the fourth lesson, I discuss the relative loss of the second-order asymptptoc MSE with respect to the optimal choice for the common choice of $K$ for arbitrary values of the parameters used in Figure (ref).
remarkFigure (ref) ---and more concretely, the explicit expression for the second-order MSE presented in (ref)--- explains several of the findings obtained through simulations, such as the relatively large gains in accuracy by increasing $K$ from 2 to 5 folds in comprising with increasing $K$ from 5 to 10 or 10 to 20.
The previous two lessons reveal that the common recommendation of choosing 5, 10, or 20 folds for the cross-fitting procedure in DML (e.g., ahrens2024ddml,ahrens2024model, bach2022doubleml, and JSSv108i03) is suboptimal in terms of (higher-order) bias and second-order MSE. In the next lesson, I will discuss the relative loss a practitioner can face by choosing $K = 5, 10, 20$ instead of the optimal choice $K = n$.
Fourth lesson: If the optimal choice in terms of bias and accuracy is $K=n$, then choosing $K=10$ to implement DML2 guarantees that the maximum relative loss with respect to the optimal choice in terms of bias and second-order MSE is around 10% and 5%, respectively. In other words, the practitioner implementing DML2 with $K=10$ has an estimator with a (higher-order) bias that is at most $10\%$ larger than the bias obtained by implementing DML2 with the optimal $K=n$. Similarly, the second-order MSE of the DML2 estimator with $K=10$ is at most $5\%$ larger than the one obtained by implementing DML2 with the optimal $K=n$. In what follows, I explain in more detail these results and present simple expressions to calculate these maximum relative losses as a function of $K$, $n$, and $\varphi$.
The relative loss of the (higher-order) bias with respect to the optimal choice is presented next as a function of $K$,
equation[equation omitted — 126 chars of source]
which represents the percentage change of (ref) with respect to the optimal value of (ref) ($K=n$). Panel (a) in Figure (ref) presents the previous expression in percentages as a function of $K$. The blue line with circular markers represents the case of nuisance function estimators with slower convergence rates ($\varphi = 1/4$), while the red line with square markers corresponds to nuisance function estimators with faster convergence rates ($\varphi = 1/2$). It follows from (ref) and the figure that when $K=10$, the relative loss of the (higher-order) bias is between $5\%$ and $10\%$ for different values of $\varphi \in (1/4, 1/2)$; therefore, it is at most $10\%$. Figure (ref) considered $n=1,000$; nevertheless, the results do not change whenever $n \ge 1,000$.
figure[figure omitted — 759 chars of source]
The relative loss of the second-order MSE with respect to the optimal choice is presented next as a function of $K$,
equation*[equation* omitted — 196 chars of source]
which represents the percentage change of (ref) with respect to the optimal value of (ref) when $G_b$ is positive ($K=n$), and where $\upsilon = G_b/\sigma^2$. This previous expression depends on $\upsilon$, which may be difficult to know in empirical applications. When $\upsilon$ is positive, the expression in (ref) is an increasing function of $\upsilon$ for any $K \le n$. In particular, it can be bounded by the following expression,
equation[equation omitted — 164 chars of source]
which only depends on $K$, $n$, and $\zeta = 2\varphi - 1/2$. Panel (b) in Figure (ref) presents the previous expression in percentages as a function of $K$. It follows from (ref) and the figure that when $K=10$, the relative loss of the second-order asymptotic MSE is between $0.2\%$ and $5\%$; therefore, it is at most $5\%$. Figure (ref) considered $n=1,000$; nevertheless, the results do not change whenever $n \ge 1,000$.
remarkThe third lesson uses the second-order asymptotic MSE as an optimality criterion to evaluate the implementation of DML2. In particular, it was used to guide how to select the number of folds $K$ by minimizing the second-order MSE of $\hat{\theta}_{n,2}$. This criterion can be interpreted as an accuracy criterion since $\text{SO-MSE} [\hat{\theta}_{n,2}]$ can be interpreted as an approximation to the MSE of $\hat{\theta}_{n,2}$ based on the following informal derivations
$$ E\left[ \left( \hat{\theta}_{n,2} - \theta_0 \right)^2 \right] \overset{(1)}{\approx} n^{-1} E\left[ \mathcal{T}_{n,K}^2 \right] \overset{(2)}{\approx} \text{SO-MSE}[\hat{\theta}_{n,2}]~,$$
where (1) is motivated by the stochastic expansion in Theorem (ref) and (2) by the calculations based on the Corollary (ref).
A similar criterion has been used by linton1995second to select an optimal bandwidth and by donald2001choosing to select the optimal number of instruments, while newey2004higher used a similar idea to compare estimators sharing the same asymptotic distribution.
remarkSimulation evidence presented in Section (ref) is consistent with $G_b > 0$ in (ref). However, testing whether $G_b$ is positive is left for future research.
remarkThe second-order MSE described in Remark (ref) defines an optimality criterion that can be used to compare different decisions regarding the implementation of DML. In this paper, I used this criterion to select $K$. However, it can also be applied to select bandwidths or estimators for the nuisance function in applications. These alternative uses, however, are beyond the scope of this paper and are left for future research.
Monte-Carlo Simulations
This section examines the asymptotic results presented in Section (ref) in finite samples. Specifically, I present the bias and mean square error (MSE) of estimators based on DML1 and DML2 building on the Monte Carlo simulation designs for ATT-DID from sant2020doubly and for LATE from hong2010semiparametric. Additionally, I present the coverage probability of confidence intervals constructed as in chernozhukov2018double. For the sake of readability, these confidence intervals are defined next
equation[equation omitted — 220 chars of source]
where $z_{1-\alpha}$ is the $1-\alpha$ quantile of the standard normal distribution,
$$ \hat{\sigma}_{n,j}^2 = \frac{ n^{-1} \sum_{i=1}^n m(W_i,\hat{\theta}_{n,j},\hat{\eta}_i)^2 }{ \left(n^{-1} \sum_{i=1}^n \psi^a(W_i,\hat{\eta}_i) \right)^2} ~,$$
$\hat{\theta}_{n,j} $ is as in (ref) and (ref) for DML1 and DML2, respectively, and $\hat{\eta}_i$ is as in (ref).
Difference-in-Difference
This section is based on Example (ref). I built on the simulation design presented in sant2020doubly. The observed outcome in the pre-treatment period and the potential outcomes in the post-period treatment are defined by
align*[align* omitted — 164 chars of source]
where $f_{reg}(X) = 210 + 6.85 X_1 + 3.425(X_2 + X_3 + X_4)$ and $v(X_i,A_i) = A_i f_{reg}(X) + \varepsilon_{v,i}$,
and $(\varepsilon_{0,i}, \varepsilon_{1,i}(0), \varepsilon_{1,i}(1), \varepsilon_{v,i})$ is distributed as $N(0,I_4)$, $I_4$ is the $4 \times 4$ identity matrix.
The treatment assignment is defined by $A_i \sim \text{Bernoulli}( p(X_i) )$, where
align*[align* omitted — 149 chars of source]
Finally, the vector of covariates is $X_i = (X_{1,i}, X_{2,i}, X_{3,i}, X_{4,i}) \in [0,1]^4$ and all its coordinates are independent uniform random variables (e.g., $X_{1,i} \sim \text{Uniform}[0,1])$.
The estimators for the ATT-DID are defined as in (ref) and (ref) using $\psi^a$ and $\psi^b$ presented in Example (ref). To estimate the $ j$th component of the vector of nuisance functions $\eta_0$, I use the Nadaraya-Watson estimator with a 6th-order Gaussian kernel and common bandwidth $h_j = c n_0^{-1/16}$ for all coordinates, where $n_0 = (K-1)/K n$.\footnote{I also considered a 2nd order Gaussian Kernel in the simulations. The results are presented in Figure (ref) and (ref) in Appendix (ref), and they are similar to the ones presented using a 6th order Gaussian kernel. }
I consider a sample size $n = 3,000$, different values for the choice of the number of folds $K \in \{2, 5, 10,\ldots,30\}$, and perform $5,000$ simulations. I additionally consider different values for the constant $c \in \{0.37, 0.62, 0.86, 1.11\}$ in the bandwidth.
figure[figure omitted — 959 chars of source]
figure[figure omitted — 1,007 chars of source]
Figures (ref) and (ref) report the results of the simulations in terms of bias, MSE, and coverage probability.
Figure (ref) compares the performance between the DML1 and DML2 estimators across different values of $K$, with $c = 0.62$. Panel (a) presents the absolute value of the scaled bias. It shows that the biases of DML1 and DML2 are similar. This result is consistent with the findings of Section (ref) since the discrepancy measure $\Lambda = 0$ for Example (ref) (ATT-DID). Furthermore, the bias of DML2 decreases as $K$ increases, consistent with Theorem (ref). Panel (b) presents the scaled MSE. It shows that DML1 and DML2 exhibit similar values. Moreover, the scaled MSE of DML2 decreases as $K$ increases, which aligns with Corollary (ref). Finally, Panel (c) highlights the similarities between DML1 and DML2 in terms of the coverage probability of their confidence intervals. Overall, Figure (ref) aligns with the findings in Section (ref), suggesting that DML1 and DML2 behave similarly when $\Lambda=0$.
Figure (ref) compares the results of DML2 estimators for different values of $c$. It presents results qualitatively similar to the ones in Figure (ref), with some exceptions for $c=0.37$ that present a non-monotonic scaled bias and large values for the scaled MSE. It also reveals that the scaled bias and MSE are sensitive to the values of $c$. For instance, the scaled MSE for $c=1.11$ is larger than twice the scaled MSE for $c=0.62$ due to a larger bias.
Local Average Treatment Effects
This section is based on Example (ref). I built on the simulation design presented in hong2010semiparametric. The potential treatment decisions are defined as
align*[align* omitted — 103 chars of source]
where $X_i \sim \text{Uniform}[0,1]$ and $V_i \sim N(0,1)$ are independent random variables.
The potential outcomes are defined by
align*[align* omitted — 228 chars of source]
where $\xi_{1,i} \sim \text{Poisson}(\exp(X_i/2)) $, $\xi_{2,i} \sim \text{Poisson}(\exp(X_i/2)) $, $\xi_{3,i} \sim \text{Poisson}(2) $, and $\xi_{1,i} \sim \text{Poisson}(1) $, and all these random variables are independent conditional on $X_i$. The treatment assignment is defined by $Z_i \sim \text{Bernoulli}(\Phi(X_i-0.5)) $. As in Example (ref), the observed treatment decision and the observed outcome are defined by $D_i = Z_i D_i(1) + (1-Z_i) D_i(0)$ and by $Y_i = D_i Y_i(1) + (1-D_i) Y_i(0)$, respectively.
The estimators for the LATE are defined as in (ref) and (ref) using $\psi^a$ and $\psi^b$ presented in Example (ref). To estimate the $j$th component of the vector of nuisance functions $\eta_0$, I use the Nadaraya-Watson estimator with a 2th-order Gaussian kernel and common bandwidth $h_j = c n_0^{-1/5}$, where $n_0 = (K-1)/K n$.
I consider a sample size $n = 3,000$, different values for the choice of the number of folds $K \in \{2, 5, 10,\ldots,30\}$, and perform $5,000$ simulations. I additionally consider different values for the constant $c \in \{0.32, 0.53, 0.74, 0.95\}$ in the bandwidth.
figure[figure omitted — 941 chars of source]
figure[figure omitted — 934 chars of source]
Figures (ref) and (ref) present the results of the simulations in terms of bias, MSE, and coverage probability. Figure (ref) compares the performance between the DML1 and DML2 estimators across different values of $K$, with $c=0.53$. Panel (a) presents the absolute value of the scaled bias. It shows that the bias of DML1 is approximately an increasing linear function of $K$, which is consistent with the intuition presented in Remark (ref) since the discrepancy measure $\Lambda$ for the LATE in Example (ref) is different than zero. In contrast, the bias of DML2 decreases as $K$ increases, consistent with Theorem (ref). Panel (b) presents the scaled MSE. It reveals that the scaled MSE for DML1 is increasing and approximately quadratic on $K$. This finding aligns with the expressions presented in Remark (ref) for the oracle version of DML1. Additional simulation results presented in Figure (ref) in Appendix (ref) show that the DML1 estimator and its oracle version exhibit similar values. In contrast, the scaled MSE of DML2 is a decreasing function of $K$. Finally, Panel (c)shows dramatic discrepancies between DML1 and DML2 in terms of the coverage probability of the confidence intervals associated with them. In particular, it evidences that inference based on DML1 deteriorates as $K$ increases. In contrast, DML2 does not have this problem, and inference is reliable for all the values of $K$. Overall, Figure (ref) is consistent with the findings in Section (ref), suggesting that DML1 and DML2 behave differently when $\Lambda \neq 0$.
Figure (ref) compares the results of DML2 estimators for different values of $c$. It presents results qualitatively similar to the ones in Figure (ref), with some exceptions for $c=0.32$ that present a non-monotonic scaled bias. It also reveals that the scaled bias and MSE are less sensitive to the values of $c$. For instance, the scaled MSE exhibits values between 123 and 124.5 for all the values of $c$ and $K$ between 5 and 30.
Concluding Remarks
This paper studies the properties of debiased machine learning (DML) estimators under a novel asymptotic framework. DML is an estimation method suited to economic models where the parameter of interest depends on unknown nuisance functions that must be estimated. In practice, two versions of DML ---introduced by chernozhukov2018double---can be used, DML1 and DML2. Both versions randomly divide the data into $K$ equal-sized folds to estimate the nuisance function, but they differ in how these estimates are used to estimate the parameters of interest. In this paper, I consider an asymptotic framework in which $K$ diverges to infinity as $n$ diverges to infinity, accommodating small-sample situations where the practitioner may wish to increase $K$, situations not well approximated by the existing framework in which $K$ is fixed as $n$ diverges.
Under this framework, this paper makes several contributions. First, it shows that DML2 asymptotically dominates DML1 in terms of bias and mean square error. Furthermore, it characterizes the first-order asymptotic difference between DML1 and DML2 using a discrepancy measure, $\Lambda$, which can be computed for several treatment effect parameters. Second, it provides conditions under which all DML2 estimators, regardless of $K$, are asymptotically valid and share the same limiting distribution. To distinguish among them, this paper uses higher-order asymptotic approximations, leading to the final contribution: setting $K=n$ to implement DML2 can be asymptotically optimal in terms of higher-order asymptotic bias and second-order asymptotic MSE within the class of DML2 estimators.
appendix\section{Additional Examples and Results}
\subsection{More Examples}
\begin{example}[Weighted Average Treatment Effect]
This example is built on the setup of Example (ref) (ATE) and the parameter considered in Equation (2) of hirano2003efficient. The parameter of interest is defined by
$$ \theta_0 = E\left[ E[ (Y(1) - Y(0) ) \mid X] g(X) \right]/E[g(X)]~,$$
where $g(\cdot)$ is a known function of covariates $X$, such that $|g(X)|$ is bounded and $E[g(X)] >0$. When $g(X)$ equals the propensity score $ E[ A \mid X]$, the parameter $\theta_0$ equals the average treatment effect on the treated, which implicitly assumes perfect knowledge of the propensity score. Under the selection-on-observables assumptions, the parameter $\theta_0$ can be identified by a moment condition, such as (ref), using a moment function like (ref), where
\begin{align*}
\psi^b(W,\eta) &= g(X) \left( \eta_1 - \eta_2 + A(Y-\eta_1)\eta_3 - (1-A) (Y-\eta_2) \eta_4 \right) ,\\
\psi^a(W,\eta) &= g(X) ,
\end{align*}
for $\eta \in \mathbf{R}^4$, and where the nuisance parameter $\eta_0(X)$ is exactly the same as in Example (ref). This moment function appears as the efficient influence function in hirano2003efficient for the weighted average treatment effect. In this example,
$$ \Lambda = E[ g(X)^2 \left \{ \eta_{0,1}(X) - \eta_{0,2}(X) - \theta_0 \right \} ]/E[g(X)]^2~,$$
which is typically different than zero.
\end{example}
\begin{example}[Average Treatment Effect on the Treated]
This example is built on the setup of Example (ref) (ATE). It is assumed that there is no knowledge of the propensity score $E[A \mid X]$, and it has to be estimated. The parameter of interest is
$$ \theta_0 = E[ Y(1) - Y(0) \mid A = 1]~,$$
which is the treatment effect for the treated group, also known as ATT. Under selection-on-observable assumptions, the parameter $\theta_0$ can be identified by a moment condition, such as (ref), using a moment function like (ref), where
\begin{align*}
\psi^b(W,\eta) &= A( Y - \eta_1) + (1-A) (1-\eta_2) (Y - \eta_1) ,\\
\psi^a(W,\eta) &= A ,
\end{align*}
for $\eta \in \mathbf{R}^2$, and where the nuisance parameter $\eta_0(X)$ has two components:
\begin{align*}
\eta_{0,1}(X) &= E[ Y \mid X, A=0] ,\\
\eta_{0,2}(X) &= (E[1 - A \mid X])^{-1} .
\end{align*}
When there is no knowledge of the propensity score, this moment function appears as the efficient influence function for the ATT in hahn1998role and hirano2003efficient. In this example, $\Lambda = 0$.
\end{example}
\begin{example}[Partial Linear Model]
This example presents the model studied in robinson1988root and linton1995second. Consider the following model:
$$Y = D \theta_0 + g(X) + U~, $$
where $E[ U \mid D, X] = 0$. Here $W= (Y,D,X)$. The parameter of interest is $\theta_0$. In this example, $\theta_0$ can be identified by (ref) and (ref), where
\begin{align*}
\psi^a(W,\eta) &= (D-\eta_2)^2 ,\\
\psi^b(W,\eta) &= (Y-\eta_1)(D - \eta_2) ,
\end{align*}
for $\eta \in \mathbf{R}^2$, and where the nuisance parameter $\eta_0(X)$ has two components:
\begin{align*}
\eta_{0,1}(X) &= E[Y \mid X] ,\\
\eta_{0,2}(X) &= E[D \mid X] .
\end{align*}
In this example, $\Lambda = 0$.
\end{example}
\begin{example}[Partial Linear IV Model]
This example presents the extended PLM. Consider the following model:
\begin{align*}
Y &= D \theta_0 + g(X) + U , \\
Z &= m(X) + V ,
\end{align*}
where $E[ U \mid D, Z] = 0$ and $E[V \mid X] = 0$. Here $W = (Y,D,X,Z)$. The parameter of interest is $\theta_0$. In this example, $\theta_0$ can be identified by (ref) and (ref), where
\begin{align*}
\psi^a(W,\eta) &= (D-\eta_2)(Z-\eta_3) ,\\
\psi^b(W,\eta) &= (Y-\eta_1)(Z - \eta_3) ,
\end{align*}
for $\eta \in \mathbf{R}^3$, and where the nuisance parameter $\eta_0(X)$ has three components:
\begin{align*}
\eta_{0,1}(X) &= E[Y \mid X] ,\\
\eta_{0,2}(X) &= E[D \mid X]) ,\\
\eta_{0,3}(X) &= E[Z \mid X]) .
\end{align*}
In this example, $\Lambda$ is typically different than zero.
\end{example}
\subsection{Additional Results for DML2 estimators}
Let $\delta_{n_0}$ and $b_{n_0}$ be the functions defined in Assumption (ref). For $j \neq i$, define
\begin{equation*}
\tilde{b}_{n_0}(X_i) = E\left[ b_{n_0}(X_j,X_i) \mid X_i \right] .
\end{equation*}
Recall $\eta_i = \eta_0(X_i)$ and consider the following notation:
\begin{align}
J_0 &= E\left[ \psi^a(W_i,\eta_i) \right] ,\\
D_i &= J_0^{-1}\left(\partial_\eta m(W_i,\theta_0,\eta)|_{\eta = \eta_i}\right) .
\end{align}
Using the previous notation, consider the stochastic approximation terms:
\begin{itemize}
• $\mathcal{T}_{n,K}^l $ is the asymptotic second-order linear term and
\begin{equation}
\mathcal{T}_{n,K}^l = n^{-1/2} \sum_{i=1}^n \Delta_{i} ^{\top} D_i ,
\end{equation}
where $\Delta_i$ and $D_i$ are as in (ref) and (ref), respectively.
• $\mathcal{T}_{n}^{dml2}$ is the asymptotic high-order DML term of the DML2 estimator and
\begin{equation}
\mathcal{T}_{n}^{dml2} = -n^{-1/2} \left( n^{-1/2} \sum_{i=1}^n m_i/J_0 \right) \left( n^{-1/2} \sum_{i=1}^n ({\psi}^a_i-J_0)/J_0 \right) .
\end{equation}
• $\mathcal{T}_{n,K}^{dml1}$ is the asymptotic high-order DML term of the DML1 estimator
\begin{equation}
\mathcal{T}_{n,K}^{dml1} = -n^{-1/2} \sum_{k=1}^K \left( n_k^{-1/2} \sum_{i \in \mathcal{I}_k} m_i/J_0 \right) \left( n_k^{-1/2} \sum_{i \in \mathcal{I}_k} ({\psi}^a_i-J_0)/J_0 \right) ,
\end{equation}
where $m_i = m(W_i,\theta_0,\eta_i)$, $ \psi^a_i = \psi^a(W_i,\eta_i)$, and $n_k = n/K$.
\end{itemize}
The next assumption complements Assumption (ref) to derive valid stochastic expansions for DML2 estimators when $\varphi_1 < \varphi_2$.
\begin{assumption}
\begin{enumerate}[(a)]
• The following limits exist and are finite,
\begin{align}
G_{\delta}^l &= \lim_{n_0 \to \infty} E \left[\left( \delta_{n_0}(W_j,X_i)^{\top} D_i \right) \left( \delta_{n_0}(W_i,X_j)^{\top} D_j + \delta_{n_0}(W_j,X_i)^{\top} D_i \right) \right] , \\
G_{b}^l &= \lim_{n_0 \to \infty} E\left[ \left( m_{i}/J_0 \right) \tilde{b}_{n_0}(X_i)^{\top} D_i \right] .
\end{align}
• $G_\delta^l >0$ .
\end{enumerate}
\end{assumption}
The next theorem is an extension of Theorem (ref). It considers valid stochastic expansions for $(\varphi_1,\varphi_2) \in (1/4, 1/2) \times (1/4, 1)$. It uses the following notation:
\begin{itemize}
• $\mathcal{R}_1 = \{ (\varphi_1,\varphi_2) \in (1/4, 1/2) \times (1/4, 1): \varphi_1 < 1/3 \text{ or } \varphi_2 < 1/2 \}$.
• $\mathcal{R}_2 = \{ (\varphi_1, \varphi_2) \in (1/4, 1/2) \times (1/4,1) : \varphi_1 \ge 1/3, \varphi_2 \ge 1/2, (\varphi_1, \varphi_2) \notin \mathcal{R}_3 \}$.
• $\mathcal{R}_3 = \{ (\varphi_1, \varphi_2) \in (1/4, 1/2) \times (1/4,1) : \varphi_1 \ge 3/8, \varphi_2 \ge 1/2, \varphi_1 + \varphi_2 \ge 1 \}$.
\end{itemize}
\begin{theorem}
Suppose Assumptions (ref), (ref), and (ref) hold. In addition, assume that $K$ is such that $K \le n$ and $K \to \infty$ as $n \to \infty$. If $\varphi_1 \in (1/4, 1/2)$, and $\varphi_1 \le \varphi_2$, then
\begin{equation}
n^{1/2} (\hat{\theta}_{n,2} - \theta_0) = \mathcal{T}_{n,K} + R_{n,K} ,
\end{equation}
where $\zeta = \min \{ 4\varphi_1-1, \varphi_1 + \varphi_2 - 1/2 \}$, $\hat{\theta}_{n,2}$ is as in (ref), and
\begin{itemize}
• Case 1: $ \mathcal{T}_{n,K} = \mathcal{T}_{n}^* + \mathcal{T}_{n,K}^{nl} $ if $(\varphi_1, \varphi_2) \in \mathcal{R}_1$, where $\mathcal{T}_{n}^*$ is defined in (ref) and $\mathcal{T}_{n}^* \overset{d}{\to} N(0,\sigma^2)$ with $\sigma^2$ defined in (ref), $\mathcal{T}_{n,K}^{nl}$ is defined in (ref) and satisfies (i) $\lim_{n \to \infty} \inf_{K \le n } Var[n^{2\varphi_1 - 1} \mathcal{T}_{n,K}^{nl}] > 0$ and (ii) $\lim_{n \to \infty} \sup_{K \le n } E[(n^{2\varphi_1 - 1} \mathcal{T}_{n,K}^{nl})^2 ] < \infty$.
For the next two cases, suppose in addition that Assumption (ref) holds.
• Case 2: $ \mathcal{T}_{n,K} = \mathcal{T}_{n}^* + \mathcal{T}_{n,K}^{nl} + \mathcal{T}_{n,K}^l $ if $(\varphi_1, \varphi_2) \in \mathcal{R}_2$, where $\mathcal{T}_{n}^*$ and $ \mathcal{T}_{n,K}^{nl} $ are defined as in Case 1, $\mathcal{T}_{n,K}^l$ is defined in (ref) and satisfies (i) $\lim_{n \to \infty} \inf_{K \le n } Var[n^{\varphi_1} \mathcal{T}_{n,K}^{l}] > 0$ and (ii) $\lim_{n \to \infty} \sup_{K \le n } E[(n^{\varphi_1} \mathcal{T}_{n,K}^{l})^2 ] < \infty$.
• Case 3: $ \mathcal{T}_{n,K} = \mathcal{T}_{n}^* + \mathcal{T}_{n,K}^{nl} + \mathcal{T}_{n,K}^l + \mathcal{T}_{n}^{dml2} $ if $(\varphi_1, \varphi_2) \in \mathcal{R}_3$, where $\mathcal{T}_{n}^*$, $\mathcal{T}_{n,K}^{nl}$, $\mathcal{T}_{n,K}^{l} $ are defined as in Case 1 and 2, $ \mathcal{T}_{n}^{dml2}$ is defined in (ref) and $n^{1/2} \mathcal{T}_{n}^{dml2}$ has a non-degenerate limit distribution.
\end{itemize}
with
$$ \lim_{n \to \infty} \sup_{K \le n } P\left( n^{\zeta} |R_{n,K}| > \epsilon \right) = 0~,$$
for any given $\epsilon>0$.
\end{theorem}
\begin{remark}
When $K$ is fixed as $n \to \infty$ and $(\varphi_1,\varphi_2) \in \mathcal{R}_3$, the following stochastic expansion can be derived for DML1 estimators,
\begin{equation*}
n^{1/2} (\hat{\theta}_{n,1} - \theta_0) = \mathcal{T}_{n}^* + \mathcal{T}_{n,K}^{nl} + \mathcal{T}_{n,K}^{nl} + \mathcal{T}_{n}^{dml1} + o_p(n^{-\zeta}) ,
\end{equation*}
where $\zeta = \min \{ 4\varphi_1-1, \varphi_1 + \varphi_2 - 1/2 \}$ and $\mathcal{T}_{n}^{dml1}$ is defined in (ref). Furthermore, the stochastic expansion presented for DML2 in Theorem (ref) for case 3 ($(\varphi_1,\varphi_2) \in \mathcal{R}_3$) is also valid when $K$ is fixed as $n \to \infty$. For convenience, it is presented below
$$ n^{1/2} (\hat{\theta}_{n,2} - \theta_0) = \mathcal{T}_{n}^* + \mathcal{T}_{n,K}^{nl} + \mathcal{T}_{n,K}^{nl} + \mathcal{T}_{n}^{dml2} + o_p(n^{-\zeta})~.$$
These expressions show that, for the values of $(\varphi_1,\varphi_2) \in \mathcal{R}_3$, both stochastic expansions are only different in $\mathcal{T}_{n}^{dml1} $ and $\mathcal{T}_{n}^{dml2} $, which capture the effects of implementing DML.
\end{remark}
\begin{theorem}
Suppose Assumptions (ref), (ref), and (ref) hold. In addition, assume that $K$ is such that $K \le n$ and $K \to \infty$ as $n \to \infty$. If $\varphi_1 \in (1/4, 1/2)$, $\varphi_2 < 1$, and $\varphi_1 \le \varphi_2$, then
\begin{equation*}
E[\mathcal{T}_{n,K}] = F_K n^{1/2-2\varphi_1} + \nu_{n,K} ,
\end{equation*}
where
\begin{equation}
F_{K} =
\begin{cases}
(F_\delta + F_b) \left( 1 + \frac{1}{K-1} \right)^{2\varphi_1} & if \varphi_1 = \varphi_2 , \\
F_\delta \left( 1 + \frac{1}{K-1} \right)^{2\varphi_1} & if \varphi_1 < \varphi_2 ,
\end{cases}
\end{equation}
and
$$ \sup_{K \le n } |\nu_{n,K}| = o(n^{1/2-2\varphi_1})~,$$
with $\mathcal{T}_{n,K}$ defined as in Theorem (ref), and $F_{\delta}$ and $F_{b}$ defined as in (ref) and (ref), respectively.
\end{theorem}
\begin{theorem}
Suppose Assumptions (ref), (ref), and (ref) hold. In addition, assume that $K$ is such that $K \le n$ and $K \to \infty$ as $n \to \infty$. If $\varphi_1 \in (1/4, 1/2)$, $\varphi_2 < 1$, and $\varphi_1 \le \varphi_2$, then
\begin{equation*}
Var[\mathcal{T}_{n,K}] = \sigma^2 + \Omega_{K}/n^{\zeta} + r_{n,K} ,
\end{equation*}
where
\begin{equation}
\Omega_{K} =
\begin{cases}
G_{b} \left( \frac{K}{K-1} \right)^{\zeta} & if 3\varphi_1-1/2 > \varphi_2 , \\
\left( G_{\delta} \frac{ K^2-3K+3 }{(K-1)^2 } + G_{b} \right) \left( \frac{K}{K-1} \right)^{\zeta} & if 3\varphi_1-1/2 = \varphi_2 , \\
G_{\delta} \left( \frac{ ( K^2-K+3) }{(K-1)^{2 } } \right) \left( \frac{K}{K-1} \right)^{\zeta} & if 3\varphi_1-1/2 < \varphi_2 ,
\end{cases}
\end{equation}
and
\begin{equation*}
\sup_{K \le n } |r_{n,K}| = o(n^{-\zeta}) ,
\end{equation*}
with $\zeta = \min \{ 4\varphi_1-1, \varphi_1 + \varphi_2 - 1/2 \}$, $\mathcal{T}_{n,K}$ defined as in Theorem (ref), and $\sigma^2$, $G_\delta$ and $G_b$ defined as in (ref), (ref), and (ref), respectively.
\end{theorem}
\begin{corollary}
Suppose Assumptions (ref), (ref), and (ref) hold. In addition, assume that $K$ is such that $K \le n$ and $K \to \infty$ as $n \to \infty$. If $\varphi_1 \in (1/4, 1/2)$, $\varphi_2 < 1$, and $\varphi_1 \le \varphi_2$, then
\begin{equation*}
E[\mathcal{T}_{n,K}^2] = \sigma^2 + \Tilde{\Omega}_{K}/ n^{\zeta} + \tilde{r}_{n,K} ,
\end{equation*}
where
\begin{equation}
\Tilde{\Omega}_{K} =
\begin{cases}
G_{b} \left( \frac{K}{K-1} \right)^{\zeta} & \text{if } 3\varphi_1-1/2 > \varphi_2 , \\
\left( G_{\delta} \frac{ K^2-3K+3 }{(K-1)^2 } + G_{b} + F_{\delta}^2 \left( \frac{K}{K-1} \right)\right) \left( \frac{K}{K-1} \right)^{\zeta} & \text{if } 3\varphi_1-1/2 = \varphi_2 , \\
\left( G_{\delta} \frac{ ( K^2-3K+3) }{(K-1)^{2 } } + F_{\delta}^2 \left( \frac{K}{K-1} \right) \right) \left( \frac{K}{K-1} \right)^{\zeta} & \text{if } 3\varphi_1-1/2 < \varphi_2 ,
\end{cases}
\end{equation}
\begin{equation*}
\sup_{K \le n } |\tilde{r}_{n,K}| = O(n^{1/2-2\varphi_1}) ,
\end{equation*}
with $\zeta = \min \{ 4\varphi_1-1, \varphi_1 + \varphi_2 - 1/2 \}$, $\mathcal{T}_{n,K}$ defined as in Theorem (ref), and $\sigma^2$, $F_\delta$, $G_\delta$, and $G_b$ defined as in (ref), (ref), (ref), and (ref), respectively.
\end{corollary}
\subsubsection*{Tuning of Nuisance Parameters}
The second-order asymptotic MSE can be defined by
\begin{equation}
\text{SO-MSE} [\hat{\theta}_{n,2}] = \sigma^2/n + \Tilde{\Omega}_{K} /n^{\zeta + 1} ,
\end{equation}
where $\sigma^2$ is the variance of the asymptotic distribution of the estimator $\hat{\theta}_{n,2}$ defined in (ref), $\Tilde{\Omega}_{K}$ is the higher-order asymptotic MSE defined in (ref), $\zeta = \min\{ 4 \varphi_1 - 1, \varphi_1 + \varphi_2 - 1/2 \}$, $\varphi_1$ is such that $n^{-2\varphi_1}$ is the convergence rate of the variance of the nuisance parameter estimator, and $\varphi_2$ is such that $n^{-\varphi_2}$ is the convergence rate of the bias of the nuisance parameter estimator.
The explicit formulas for the second-order asymptotic MSE show that tuning the nuisance parameter estimators optimally may still be suboptimal for estimating $\theta_0$. Available recommendations for tuning the nuisance parameter estimators often rely on minimizing an out-of-sample prediction error, which can be interpreted as recommendations that minimize the asymptotic integrated mean square error and guarantee optimal convergence rates for the nuisance parameter estimators. However, it is unclear if these recommendations are optimal for estimating the parameter of interest. In particular, the optimal tuning of the estimators for the nuisance parameter $\eta_0$ alone may be suboptimal for the parameter of interest $\theta_0$. To illustrate this point, I consider Example (ref) (ATE) using the Nadaraya-Watson (N-W) estimators for the nuisance parameter $\eta_0$. More concretely, I first obtain the values of $\varphi_1$ and $\varphi_2$ associated with the N-W estimators, and I then use them in (ref) to derive the optimal convergence rate of the bandwidth based on the criterion.
For the sake of exposition I consider the case where the covariate $X$ is univariate (e.g., $d_x=1$), and the N-W uses a second-order kernel; see Appendix (ref) for additional details for the general case. Let $h_j = c_j n_0^{-\varphi_0}$ be the bandwidth used for estimating the component $\eta_{0,j}$ of $\eta_0$, where $c_j$ is a given positive constant and $n_0 = ((K-1)/K)n$ is the sample size used for the estimation. In this case, it can be shown that
\begin{equation}
\varphi_1 = (1 - \varphi_0) /2 \quad \text{and} \quad \varphi_2 = 2 \varphi_0 ,
\end{equation}
where $\varphi_0 \in [1/5, 1/2)$ to guarantee that $\varphi_1 \in (1/4,1/2)$ and $\varphi_1 \le \varphi_2$. Using this notation, it follows that
\begin{equation}
\zeta = \min \{ 1 - 2\varphi_0, 3\varphi_0/2 \} .
\end{equation}
which is lower or equal than $ 3/7$.
In this example, the optimal convergence rate of the second term in (ref) is $O(n^{-10/7})$ since $\zeta \le 3/7$. This convergence rate is achieved when the bandwidth has a convergence rate $n^{-2/7}$ (i.e., $\varphi_0 = 2/7$). Importantly, this convergence rate is different than the optimal convergence rate of the bandwidth for the N-W estimators, which is $n^{-1/5}$. Therefore, the optimal tuning of the estimators for the nuisance parameter $\eta_0$ alone is suboptimal for the parameter of interest in this case.
\begin{remark}
When $d_x = 1$ and the N-W estimator uses a second order kernel, then (ref) also holds for Examples 2 and 3, and the other examples in Appendix (ref). Therefore, the conclusions derived here also apply to these examples. Appendix (ref) presents additional details and further discussion on the derivation of a general version of (ref) when covariates $X$ have dimension $d_x$ and the N-W estimator uses a kernel of order $s > d_x/2$.
\end{remark}
\subsection{Results for Nadaraya-Watson Estimators}
This section presents explicit expressions for $\varphi_1$, $\varphi_2$ , $\delta_{n}$, and $b_n$ when the nuisance parameter estimator $\hat{\eta}$ is the Nadaraya-Watson (N-W) estimator.
Specifically, it considers the N-W estimators with a common bandwidth for all coordinates, where the kernel function is the product of the same univariate kernels:
$$ K_h(x) = h^{-d_x} \prod_{\ell=1}^{d_x} K(x_\ell/h)~, \quad x \in \mathbf{R}^{d_x}$$
where $h = C_h n^{-\varphi_0}$ is the bandwidth and $K(\cdot)$ is a bounded symmetric kernel of order $s$.
Let $\eta_{0,j}(x)$ be the $j$th component of the nuisance parameter $\eta_0$. Let $f(x)$ be the density function of the covariates $X$ at $x$.
In what follows, I consider three possible types for this component:
\textbf{Type 1:} Simple conditional expectation, $\eta_{0,j}(x) = E[ Y \mid X=x]$. In this case,
$$ \hat{\eta}_{0,j}(x) = \frac{\sum_{\ell = 1 }^n Y_\ell K_h(x-X_\ell) }{ \sum_{\ell = 1 }^n K_h(x-X_\ell)} ~.$$
By matching the convergence rate of the variance of this estimator ($(nh^d)^{-1})$ with $n^{-2\varphi_1}$, it follows that:
$$ \varphi_1 = (1-d_x \varphi_0)/2~,$$
and by matching the convergence rate of the bias of this estimator ($h^s$) with $n^{-\varphi_2}$, it follows that:
$$ \varphi_2 = s \varphi_0~.$$
Finally, the $j$th component of the functions $\delta_n$ and $b_n$ are given by
\begin{align*}
\delta_{n,j}(W,x) &= C_h^{-d_x/2} h^{d_x/2} \left(Y - \eta_{0,j}(X) \right) K_h\left(x-X\right)/f(x) , \\
b_{n,j}(W,x) &= h^{-s} \left(\eta_{0,j}(X) - \eta_{0,j}(x) \right) K_h\left(x-X\right)/f(x) ,
\end{align*}
where $X$ is a sub-vector of $W = (Y,X)$. Note that the dependence of these functions on $n$ is due to the definition of bandwidth $h = C_h n^{-\varphi_0}$.\\
\textbf{Type 2:} Conditional expectation of a sub-group, $\eta_{0,j}(x) = E[ Y \mid A = 1, X=x]$. In this case,
$$ \hat{\eta}_{0,j}(x) = \frac{\sum_{\ell = 1 }^n Y_\ell A_\ell K_h(x-X_\ell) }{ \sum_{\ell = 1 }^n A_\ell K_h(x-X_\ell)} ~.$$
The same arguments presented for type 1 imply
\begin{align*}
\varphi_1 &= (1-d_x \varphi_0)/2 ,\\
\varphi_2 &= s \varphi_0 .
\end{align*}
Finally, the $j$th component of the functions $\delta_n$ and $b_n$ are given by
\begin{align*}
\delta_{n,j}(W,x) &=C_h^{-d_x/2} h^{d_x/2} \left \{ \left(Y A - g_1(X) \right) - \eta_{0,j}(x) \left(A - g_2(X) \right) \right\} K_h\left(x-X\right)/f(x) ,\\
b_{n,j}(W,x) &= h^{-s} \left \{ \left(g_1(X) - g_1(x) \right) - \eta_{0,j}(x) \left(g_2(X) - g_2(x) \right) \right\}K_h\left(x-X\right)/f(x) ,
\end{align*}
where $g_1(x) = E[ Y A \mid X=x ]$ and $g_2(x) = E[ A \mid X = x ]$, $X$ is a sub-vector of $W = (Y,A,X)$.\\
\textbf{Type 3:} Inverse of propensity score, $\eta_{0,j}(x) = \left( E[ A \mid X=x] \right)^{-1}$. In this case,
$$ \hat{\eta}_{0,j}(x) = \frac{\sum_{\ell = 1 }^n K_h(x-X_\ell) }{ \sum_{\ell = 1 }^n A_\ell K_h(x-X_\ell)} ~.$$
Under standard assumptions, it can be shown that $\hat{\eta}_{0,j}$ has the same convergence rates for the variance and bias. Therefore, the arguments presented for type 1 also apply and imply,
\begin{align*}
\varphi_1 &= (1-d_x \varphi_0)/2 ,\\
\varphi_2 &= s \varphi_0 .
\end{align*}
Finally, the $j$th component of the functions $\delta_n$ and $b_n$ are given by
\begin{align*}
\delta_{n,j}(W,x) &= - C_h^{-d_x/2} h^{d_x/2} \left(\eta_{0,j}(x) \right)^2 \left( A - g_2(X) \right) K_h\left(x-X\right)/f(x) \\
b_{n,j}(W,x) &= - h^{-s} \left(\eta_{0,j}(x) \right)^2 \left(g_2(X) - g_2(x) \right) K_h\left(x-X\right)/f(x)
\end{align*}
where $g_2(x) = E[ A \mid X = x ]$ and $X$ is sub-vector of $W = (A,X)$.
\begin{remark}
Any component of the nuisance parameters considered in Examples (ref), (ref), (ref), (ref), (ref), (ref), and (ref) is of Type 1, 2, or 3, with minor modification (e.g., $\eta_{0,4} = E[ 1- A \mid X]$ in Example (ref) is of type 3).
\end{remark}
\section{Proofs of Main Results}
\subsection{Proof of Theorems (ref) and (ref) }
The proof of these theorems relies on the following decomposition,
\begin{equation}
n^{1/2} \left( \hat{\theta}_{n,j} - \theta_0 \right) = n^{1/2} \left( \hat{\theta}_{n,j} - \hat{\theta}_{n,j}^* \right) + n^{1/2} \left( \hat{\theta}_{n,j}^* - \theta_0\right) .
\end{equation}
and three intermediate results. The first two results are Theorems (ref) and (ref) in Appendix (ref) that imply $n^{1/2} \left( \hat{\theta}_{n,j} - \hat{\theta}_{n,j}^* \right) = o_p(1)$ for $j=1,2$, respectively. These intermediate results rely on part (c) and (d) of Assumption (ref) to accommodate the challenging situation that arises in the proof due to $K\to \infty$ as $n \to \infty$.
The third result is Proposition (ref) in Appendix (ref) that calculates the asymptotic distribution of $n^{1/2} \left( \hat{\theta}_{n,j}^* - \theta_0\right) $ for $j=1,2$.
These three results and (ref) complete the proof of the theorem.
\subsection{Proof of Theorem (ref) }
This follows from Theorem (ref) when $\varphi_1 = \varphi_2$.
\subsection{Proof of Theorem (ref)}
This follows from Theorem (ref) when $\varphi_1 = \varphi_2$.
\subsection{Proof of Theorem (ref)}
This follows from Theorem (ref) when $\varphi_1 = \varphi_2$.
\subsection{Proof of Theorem (ref)}
The proof of (ref) in this theorem relies on the following decomposition,
\begin{equation}
n^{1/2} \left( \hat{\theta}_{n,2} - \theta_0 \right) = n^{1/2} \left( \hat{\theta}_{n,2} - \hat{\theta}_{n,2}^* \right) + n^{1/2} \left( \hat{\theta}_{n,2}^* - \theta_0\right) .
\end{equation}
and two intermediate results. The first result is Theorem (ref) in Appendix (ref) that implies
$$ n^{1/2} \left( \hat{\theta}_{n,2} - \hat{\theta}_{n,2}^* \right) = \mathcal{T}_{n,K}^{nl} + \mathcal{T}_{n,K}^{l} + \hat{R}_{n,K}~,$$
where $n^{\zeta } \hat{R}_{n,K}$ converges to zero in probability uniformly on $K \to \infty$ as $n \to \infty$ (equivalently, $\lim_{n \to \infty } \sup_{K \le n} P( n^{\zeta } |\hat{R}_{n,K}| > \epsilon) = 0$ for any given $\epsilon>0$). For case 1, it follows that $ n^{\zeta } \mathcal{T}_{n,K}^{l}$ converges to zero in probability uniformly on $K \to \infty$ as $n \to \infty$. Proposition (ref) implies (i) $\lim_{n \to \infty} \inf_{K \le n } Var[n^{2\varphi_1 - 1} \mathcal{T}_{n,K}^{nl}] > 0$ and $\lim_{n \to \infty} \sup_{K \le n } E[(n^{2\varphi_1 - 1} \mathcal{T}_{n,K}^{nl})^2 ] < \infty$. For case 2 and 3, under Assumption (ref), Proposition (ref) implies that (i) $\lim_{n \to \infty} \inf_{K \le n } Var[n^{\varphi_1} \mathcal{T}_{n,K}^{l}] > 0$ and $\lim_{n \to \infty} \sup_{K \le n } E[(n^{\varphi_1} \mathcal{T}_{n,K}^{l})^2 ] < \infty$.
The second result is Proposition (ref) in Appendix (ref) that implies
\begin{equation}
n^{1/2} \left( \hat{\theta}_{n,2}^* - \theta_0\right) = \mathcal{T}_n^* + \mathcal{T}_{n}^{dml2} + O_p(n^{-1}) .
\end{equation}
This expansion is independent of the number of folds $K$ since the oracle version of DML2 defined in (ref) does not depend on sample splitting. Furthermore, it is valid for a larger class of parameters identified by (ref) and can be obtained by standard arguments (e.g., newey2004higher). Central Limit Theorem implies $\mathcal{T}_n^* \to N(0,\sigma^2)$, and Proposition (ref) implies $n^{1/2} \mathcal{T}_{n}^{dml2}$ has a non-degenerate limit distribution. Finally, note that in case 1 and 2, $n ^{\zeta }\mathcal{T}_{n}^{dml2}$ converges to zero in probability uniformly on $K \to \infty$ as $n \to \infty$ (since $\mathcal{T}_{n}^{dml2}$ does not depend on $K$). In case 3, the remainder error term in (ref) scaled by $n^{\zeta}$, $n^{\zeta} O_p(n^{-2})$ converges to zero in probability uniformly on $K \to \infty $ as $n \to \infty$, since equation (ref) does not depend on $K$.
The proof of (ref) is completed by adding the previous two results in their respective case.
\subsection{Proof of Theorem (ref)}
The proof of this theorem considers the definition of $\mathcal{T}_{n,K}$ with more terms (case 3),
$$ \mathcal{T}_{n,K} = \mathcal{T}_{n}^* + \mathcal{T}_{n,K}^l + \mathcal{T}_{n,K}^l + \mathcal{T}_{n}^{dml2}~,$$
and two intermediate results.
The first result is Proposition (ref) that shows $E[\mathcal{T}_{n,K}^l] = 0$ and Proposition (ref) that shows
$$ E[\mathcal{T}_{n,K}^{nl}] = F_\delta \left( \frac{K}{K-1} \right)^{2\varphi_1} n^{1/2-2\varphi_1} + F_b \left( \frac{K}{K-1} \right)^{2\varphi_2} n^{1/2-2\varphi_2} + \nu_{n,K}~.$$
The calculation of these terms relies on
the structure imposed by part (a) of Assumption (ref), and the Neyman orthogonality condition implied by part (b) of Assumption (ref).
The second result is Proposition (ref) in Appendix (ref) that shows $E[\mathcal{T}_{n}^*] = 0$ and $E[\mathcal{T}_{n}^{dml2}] = \Lambda n^{-1/2}$.
These two results and the definition of $\mathcal{T}_{n,K}$ complete the proof. For $\mathcal{T}_{n,K}$ in case 1 or 2, the proof is analogous and uses that $E[\mathcal{T}_{n}^{dml2}] = \Lambda n^{-1/2}$ do not depend on $K$ and is $o(n^{1/2-2\varphi_1})$.
\subsection{Proof of Theorem (ref)}
In the proof of the next theorem, $x_{n,K} = o(1)$ denotes a real valued sequence $x_{n,K}$ converging to zero uniformly on $K \to \infty$ as $n \to \infty$ (equivalently, $\lim_{n \to \infty} \sup_{K \le n} |x_{n,K}| = 0$).
The proof of this theorem considers the definition of $\mathcal{T}_{n,K}$ with more terms (case 3) since the other two cases follow similarly.
\begin{align}
\text{Var}[\mathcal{T}_{n,K}] &= \text{Var}[\mathcal{T}_{n}^* + \mathcal{T}_{n,K}^l + \mathcal{T}_{n,K}^l + \mathcal{T}_{n}^{dml2}] , \notag \\
&= \text{Var}[\mathcal{T}_{n}^* + \mathcal{T}_{n,K}^l ] + \text{Var}[ \mathcal{T}_{n,K}^l + \mathcal{T}_{n}^{dml2}] + 2 \text{Cov}(\mathcal{T}_{n}^* + \mathcal{T}_{n,K}^l , \mathcal{T}_{n,K}^l + \mathcal{T}_{n}^{dml2}) \notag \\
&\overset{(1)}{=} \text{Var}[\mathcal{T}_{n}^* + \mathcal{T}_{n,K}^l ] + \text{Var}[ \mathcal{T}_{n}^{dml2}] + 2 \text{Cov}(\mathcal{T}_{n}^*, \mathcal{T}_{n}^{dml2}) + n^{-\zeta} o(1) ,
\end{align}
where (1) holds by the auxiliary results presented in Appendix (ref). More concretely, Proposition (ref) implies $\sup_{K \le n} \text{Var}[\mathcal{T}_{n,K}^l] = o(n^{-\zeta})$ and $\sup_{K \le n} \text{Cov}(\mathcal{T}_{n,K}^l, \mathcal{T}_{n}^{dml2}) = o(n^{-\zeta})$, respectively, part 3 and 4 of Proposition (ref) imply $\sup_{K \le n} \text{Cov}(\mathcal{T}_{n,K}^{nl}, \mathcal{T}_{n,K}^l + \mathcal{T}_{n}^{dml2}) = o(n^{-\zeta})$, and part 1 of Proposition (ref) implies $\sup_{K \le n} \text{Cov}(\mathcal{T}_n^*, \mathcal{T}_{n,K}^{l}) = o(n^{-\zeta})$.
To complete the proof, I present below explicit expressions for each of the first three terms in (ref). Part 4 of Proposition (ref) is used to compute the second and third terms in (ref). It shows
$$ \text{Var}[ \mathcal{T}_{n}^{dml2}] + 2 \text{Cov}(\mathcal{T}_{n}^*, \mathcal{T}_{n}^{dml2}) = \Lambda_1 n^{-1} + O(n^{-2}) ~.$$
These calculations are independent of the assumptions on the nuisance parameter estimators and the number of folds since the oracle version of DML2 does not depend on sample splitting.
Part 3 of Proposition (ref) calculates $\text{Var}[\mathcal{T}_{n}^* ] $, part 2 of Proposition (ref) calculates $\text{Var}[\mathcal{T}_{n,K}^l ]$, and part 2 of Proposition (ref) calculates $\text{Cov}(\mathcal{T}_{n}^* , \mathcal{T}_{n,K}^l )$. All these expressions together compute the first term in (ref). That is
\begin{align}
\text{Var}[\mathcal{T}_{n}^* + \mathcal{T}_{n,K}^l ]
&= \sigma^2 + G_\delta \left( \frac{K^2-3K+3}{(K-1)^2} \right) n_0^{1-4\varphi_1} + G_b n_0^{1/2-\varphi_1-\varphi_2} + o(n^{-\zeta})
\end{align}
These calculations all depend on the structure imposed by part (a) of Assumption (ref).
Finally, there are three possible cases based on $\zeta = \min\{ 4\varphi_1 - 1, \varphi_1 + \varphi_2 - 1/2\}$.
First, if $3\varphi_1 - 1/2 > \varphi_2$ it follows that $\zeta = \varphi_1 + \varphi_2 - 1/2 < 4\varphi_1 - 1$ and the second term in (ref) is $o(n^{-\zeta})$. In this case, the third term in (ref) equals $\Omega_K n^{-\zeta}$. Second, if $3\varphi_1 - 1/2 = \varphi_2$, then $\zeta = 4\varphi_1 - 1$. In this case, the sum of the second and third terms in (ref) equals $\Omega_K n^{-\zeta}$. Third, if $3\varphi_1 - 1/2 < \varphi_2$, then $\zeta = 4\varphi_1 - 1 < \varphi_1 + \varphi_2 - 1/2$ and the third term in (ref) is $o(n^{-\zeta})$. In this case, the second term in (ref) equals $\Omega_K n^{-\zeta}$.
\section{Auxiliary Results}
The next result guarantees that a first-order equivalence property holds for the estimators based on DML and their oracle versions, even when $K$ grows with the sample size.
\begin{theorem}
Suppose Assumptions (ref) and (ref) hold. In addition, assume $K$ is such that $K \le n$, $K \to \infty $ and $K/\sqrt{n} \to c \in [0,\infty)$ as $n \to \infty$. If $\varphi_1 \le 1/2$ and $1/4 < \min \{\varphi_1, \varphi_2 \}$, then
\begin{equation*}
n^{1/2} \left( \hat{\theta}_{n,1} - \hat{\theta}_{n,1}^* \right) = o_p(1)
\end{equation*}
where $\hat{\theta}_{n,1}$ and $\hat{\theta}_{n,1}^*$ are as in (ref) and (ref), respectively.
\end{theorem}
\begin{proof}
See Section (ref) in Appendix (ref)
\end{proof}
\begin{theorem}
Suppose Assumptions (ref) and (ref) hold. In addition, assume that $K$ is such that $K \le n$ and $K \to \infty$ as $n \to \infty$. If $\varphi_1 \in (1/4, 1/2)$, and $\varphi_1 \le \varphi_2$, then
\begin{equation*}
n^{1/2} \left( \hat{\theta}_{n,2} - \hat{\theta}_{n,2}^* \right) = o_p(1) ,
\end{equation*}
where $\hat{\theta}_{n,2}$ and $\hat{\theta}_{n,2}^*$ are defined as (ref) and (ref), respectively. Furthermore, if Assumption (ref) holds, then
\begin{equation*}
n^{1/2} \left( \hat{\theta}_{n,2} - \hat{\theta}_{n,2}^* \right) = \mathcal{T}_{n,K}^l + \mathcal{T}_{n,K}^{l} + \hat{R}_{n,K}
\end{equation*}
where
$$ \lim_{n \to \infty} \sup_{K \le n} P(n^{\zeta} |\hat{R}_{n,K}| > \epsilon) = 0~,$$
for any fixed $\epsilon>0$, with $\zeta = \min\{ 4\varphi_1-1, \varphi_1 + \varphi_2 -1/2\}$, $\mathcal{T}_{n,K}^l$ defined as in (ref) and satisfies (i) $\lim_{n \to \infty} \inf_{K \le n } Var[n^{2\varphi_1 - 1} \mathcal{T}_{n,K}^{nl}] > 0$ and (ii) $\lim_{n \to \infty} \sup_{K \le n } E[(n^{2\varphi_1 - 1} \mathcal{T}_{n,K}^{nl})^2 ] < \infty$, and $\mathcal{T}_{n,K}^l$ defined as in (ref) and satisfies $\lim_{n \to \infty} \sup_{K \le n } E[(n^{\varphi_1} \mathcal{T}_{n,K}^{l})^2 ] < \infty$.
\end{theorem}
\begin{proof}
See Section (ref) in Appendix (ref).
\end{proof}
The next proposition calculates the asymptotic distribution of the oracle version of the DML estimators defined in Remark (ref).
\begin{proposition}
Suppose Assumption (ref). In addition, assume $K$ is such that $K \le n$, $K \to \infty $ and $K/n^{\gamma} \to c \in [0,\infty)$ as $n \to \infty$. Then,
\begin{enumerate}
• $n^{1/2}\left( \hat{\theta}_{n,1}^* - \theta_0 \right) \overset{d}{\to} N(c \Lambda ,\sigma^2)$ when $\gamma= 1/2$,
• $n^{1/2}\left( \hat{\theta}_{n,2}^* - \theta_0 \right) \overset{d}{\to} N(0,\sigma^2)$ when $\gamma = 1$,
\end{enumerate}
where $\hat{\theta}_{n,1}^*$, $\hat{\theta}_{n,2}^*$, $\sigma^2$, and $\Lambda$ are as in (ref), (ref), (ref), and (ref), respectively.
\end{proposition}
\begin{proof}
See Section (ref) in Appendix (ref)
\end{proof}
The next proposition presents a stochastic expansion for the oracle version of the DML2 estimator, which does not depend on sample splitting or the number of folds $K$.
\begin{proposition}
Suppose Assumption (ref) holds. Then,
\begin{enumerate}
• $n^{1/2} \left( \hat{\theta}_{n,2}^* - \theta_0\right) = \mathcal{T}_n^* + \mathcal{T}_{n}^{dml2} + O_p(n^{-1})$ and $E[\mathcal{T}_{n}^{dml2}] = \Lambda n^{-1/2} $
• $E[\mathcal{T}_n^*] = 0$ and $E[(\mathcal{T}_n^*)^2 ] = \sigma^2$
• $\text{Cov}(\mathcal{T}_n^*, \mathcal{T}_{n}^{dml2}) = - \Xi_1 n^{-1} $ and $Var[ \mathcal{T}_{n}^{dml2}] = ( \sigma^2 \sigma_a^2 + \Lambda^2) n^{-1} + O(n^{-2}) $
\end{enumerate}
where
\begin{align}
\Xi_1 & = E\left[ \left( m(W_i,\theta_0,\eta_i)/J_0 \right)^2 (\psi^a(W_i,\eta_i)-J_0)/J_0\right] \\
\sigma_a^2 &= E\left[ \left((\psi^a(W_i,\eta_i) - J_0)/J_0 \right)^2 \right]
\end{align}
with $J_0 = E[\psi^a(W_i, \eta_i]$, $\eta_i = \eta_0(X_i)$, and $\hat{\theta}_{n,2}^*$, $\sigma^2$, $\Lambda$, $\mathcal{T}_n^*$, and $\mathcal{T}_{n}^{dml2}$ defined as in (ref), (ref), (ref), (ref), and (ref), respectively.
\end{proposition}
\begin{proof}
See Section (ref) in Appendix (ref).
\end{proof}
\begin{proposition}
Suppose Assumptions (ref), (ref), and (ref) hold. In addition, assume that $K$ is such that $K \le n$ and $K \to \infty$ as $n \to \infty$. If $\varphi_1 \le \varphi_2$, then
\begin{enumerate}
• $E[ \mathcal{T}_{n,K}^{l} ] = 0$
• $\text{Var}[\mathcal{T}_{n,K}^{l}] = G_{\delta}^l \left( \frac{K}{K-1} \right)^{2\varphi_1} n^{-2\varphi_1} + r_{n,K}^l$, where $\sup_{ K \le n} | r_{n,K}^l | = o(n^{-2\varphi_1})$.
• $ \left |\text{Cov}(\mathcal{T}_{n}^{dml2}, \mathcal{T}_{n,K}^{l}) \right| = o(n^{-2\varphi_1})$.
\end{enumerate}
where $\mathcal{T}_{n,K}^l$, $\mathcal{T}_{n}^{dml2}$, and $G_\delta^l$ are as in (ref), (ref), and (ref), respectively.
\end{proposition}
\begin{proof}
See Section (ref) in Appendix (ref).
\end{proof}
\begin{proposition}
Suppose Assumptions (ref), (ref), and (ref) hold. In addition, assume that $K$ is such that $K \le n$ and $K \to \infty$ as $n \to \infty$. If $\varphi_1 \le \varphi_2$ and $\varphi_1 \in (1/4, 1/2)$, then
\begin{enumerate}
• $E[\mathcal{T}_{n,K}^{nl}] = F_\delta \left( \frac{K}{K-1} \right)^{2\varphi_1} n^{1/2-2\varphi_1} + F_b \left( \frac{K}{K-1} \right)^{2\varphi_2} n^{1/2-2\varphi_2} + \nu_{n,K}^{nl}$, where $\sup_{K \le n} |\nu_{n,K}| = o(n^{1/2-2\varphi_1})$.
• $Var[\mathcal{T}_{n,K}^{nl}] = G_\delta \left(\frac{(K^2-3K+3)}{(K-1)^2} \right) \left( \frac{K}{K-1} \right)^{4\varphi_1-1} n^{1-4\varphi_1} + r_{n,K}^{nl}$, where $ |r_{n,K}^{nl} | = o(n^{-\zeta})$ .
• $ \sup_{K \le n} \left | \text{Cov}(\mathcal{T}_{n}^{dml2}, \mathcal{T}_{n,K}^{nl}) \right| = o(n^{-\zeta})$
• $ \sup_{K \le n} \left | \text{Cov}(\mathcal{T}_{n,K}^{l}, \mathcal{T}_{n,K}^{nl}) \right| = o(n^{-\zeta})$
\end{enumerate}
where $\zeta = \min\{ 4\varphi_1-1, \varphi_1 + \varphi_2 -1/2\}$, and $\mathcal{T}_{n,K}^{l}$, $\mathcal{T}_{n,K}^{nl}$, $\mathcal{T}_{n}^{dml2}$, $F_\delta$, $F_b$, and $G_\delta$ are as in (ref), (ref), (ref), (ref), (ref), and (ref), respectively.
\end{proposition}
\begin{proof}
See Section (ref) in Appendix (ref).
\end{proof}
\begin{proposition}
Suppose Assumptions (ref), (ref), and (ref) hold. In addition, assume that $K$ is such that $K \le n$ and $K \to \infty$ as $n \to \infty$. If $\varphi_1 \in (1/4, 1/2)$, $\varphi_2 < 1$, and $\varphi_1 \le \varphi_2$. Then,
\begin{enumerate}
• $ \text{Cov}(\mathcal{T}_n^*, \mathcal{T}_{n,K}^{l}) = G_{b}^l \left( \frac{K}{K-1}\right)^{\varphi_2} n^{-\varphi_2} + r_{n,K}^{cov,l}$, where $\sup_{K \le n} |r_{n,K}^{cov,l}| = o(n^{-\varphi_2})$.
• $\text{Cov}(\mathcal{T}_n^*, \mathcal{T}_{n,K}^l) = \frac{G_{b}}{2} \left( \frac{K}{K-1}\right)^{1/2 - \varphi_1 - \varphi_2} n^{1/2-\varphi_1-\varphi_2} + r_{n,K}^{cov,nl}$, where $\sup_{K \le n} | r_{n,K}^{cov,nl}| = o(n^{-\zeta})$.
\end{enumerate}
where $\zeta = \min\{ 4\varphi_1-1, \varphi_1+\varphi_2-1/2\}$, and $\mathcal{T}_n^*$, $\mathcal{T}_{n,K}^l$, $\mathcal{T}_{n,K}^{nl}$, $G_{b}$, and $G_{b}^l$ are as in (ref), (ref), (ref), (ref), and (ref), respectively.
\end{proposition}
\begin{proof}
See Section (ref) in Appendix (ref).
\end{proof}
The next lemma is useful to prove intermediate results such as Lemmas (ref) and (ref).
\begin{lemma}
Let Assumption (ref) hold. Then, there exists a positive constant $C = C(p,M_1) $ such that for any $i \in \mathcal{I}_k$ and $k \in \{1,\ldots,K\}$
\begin{enumerate}
• $E\left[ || n_0^{-1/2} \sum_{\ell \notin \mathcal{I}_k} n_0^{-\varphi_1} \delta_{n_0}(W_\ell,X_i) ||^4 \right] \le C n_0^{-4\varphi_1} $
• $E\left[ || n_0^{-1} \sum_{\ell \notin \mathcal{I}_k} n_0^{-\varphi_2} b_{n_0}(W_\ell,X_i) ||^4 \right] \le C ( n_0^{-4\varphi_1} + n_0^{-4\varphi_2} ) $
\end{enumerate}
\end{lemma}
\begin{proof}
See Section (ref) in Appendix (ref).
\end{proof}
\begin{lemma}
Suppose that Assumptions (ref) and (ref) hold. In addition, assume that $K$ is such that $K \le n$ and $K \to \infty$ as $n \to \infty$. If $\varphi_1 \le 1/2$ and $1/4 < \min\{ \varphi_1, \varphi_2\}$, then
\begin{enumerate}
• For $z= a, b$,
$$ \lim_{\tilde{M} \to \infty} \lim_{ n \to \infty} \sup_{K \le n }P \left( n^{\min\{ \varphi_1, \varphi_2\} +1/2 }\left| n^{-1} \sum_{i=1}^n (\hat{\eta}_i - \eta_i)^\top \partial_\eta \psi^z (W_i,\eta_i)\right| > \tilde{M} \right) = 0~,$$
• For $z = a, b$,
$$ \lim_{\tilde{M} \to \infty} \lim_{ n \to \infty} \sup_{K \le n }P \left( n^{2 \min\{ \varphi_1, \varphi_2\}} \left| n^{-1} \sum_{i=1}^n \psi^z(W_i,\hat{\eta}_i) -\psi^z(W_i,\eta_i) \right| > \tilde{M} \right) = 0~,$$
• $$ \lim_{\tilde{M} \to \infty} \lim_{ n \to \infty} \sup_{K \le n }P \left( n^{\min\{ \varphi_1, \varphi_2\}-1/2 } \left | n^{-1} \sum_{i=1}^n (\hat{\eta}_i - \eta_i)^\top \partial_\eta m (W_i,\theta_0,\eta_i) \right| > \tilde{M} \right) = 0~,$$
• Let $ Z_{n,K} = n^{-1/2} \sum_{i=1}^n \left( m(W_i, \theta_0,\hat{\eta}_i) - m(W_i,\theta_0,\eta_i) \right)/J_0- \left( \mathcal{T}_{n,K}^l + \mathcal{T}_{n,K}^{nl} \right) $, then
$$ \lim_{\tilde{M} \to \infty} \lim_{ n \to \infty} \sup_{K \le n }P \left( n^{-1/2+3 \min\{ \varphi_1, \varphi_2\} } \left | Z_{n,K} \right| > \tilde{M} \right) = 0~,$$
\end{enumerate}
where $\eta_i = \eta_0(X_i)$, $\hat{\eta}_i$ is as in (ref), $J_0 = E[\psi^a(W_i,\eta_i)]$, and $\mathcal{T}_{n,K}^l$ and $\mathcal{T}_{n,K}^{nl}$ are as in (ref) and (ref), respectively. Moreover, it holds $\lim_{n \to \infty} \sup_{K \le n } E[(n^{2\varphi_1 - 1} \mathcal{T}_{n,K}^{nl})^2 ] < \infty$ and $\lim_{n \to \infty} \sup_{K \le n } E[(n^{\varphi_1} \mathcal{T}_{n,K}^{l})^2 ] < \infty$.
Furthermore, if Assumption (ref) holds, then $\lim_{n \to \infty} \inf_{K \le n } Var[n^{2\varphi_1 - 1} \mathcal{T}_{n,K}^{nl}] > 0$; and if Assumption (ref) holds, then
$\lim_{n \to \infty} \inf_{K \le n } Var[n^{\varphi_1} \mathcal{T}_{n,K}^{l}] > 0$.
\end{lemma}
\begin{proof}
See Section (ref) in Appendix (ref).
\end{proof}
\begin{lemma}
Suppose Assumptions (ref) and (ref) hold. In addition, assume $K$ is such that $K \le n$, $K \to \infty $ and $K/\sqrt{n} \to c \in [0,+\infty)$ as $n \to \infty$. If $1/4 < \min\{\varphi_1,\varphi_2\}$ and $\varphi_1 \le 1/2$, then,
\begin{equation}
\lim_{ n \to \infty} \sup_{K \le n }P \left( \max_{k=1,\ldots,K} \left | n_k^{-1/2} \sum_{i \in \mathcal{I}_k} (\hat{\eta}_i -\eta_i)^\top \partial_\eta \psi^z(W_i,\eta_i) \right| > \epsilon \right) = 0 ,
\end{equation}
and
\begin{equation}
\lim_{ n \to \infty} \sup_{K \le n }P \left( \max_{k=1,\ldots,K} n_k^{-1} \sum_{i \in \mathcal{I}_k } || \hat{\eta}_i - \eta_i ||^2 > \epsilon \right) = 0 ,
\end{equation}
for $z = a,b$, where $\hat{\eta}_i$ is as in (ref) and $\eta_i = \eta_0(X_i)$. In particular,
\begin{equation}
\lim_{ n \to \infty} \sup_{K \le n }P \left( \max_{k=1,\ldots,K} \left|n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{\eta}_i - \eta_i)^\top \partial_\eta m(W_i, \theta_0, \eta_i) \right| > \epsilon \right) = 0 ,
\end{equation}
and
\begin{equation}
\lim_{ n \to \infty} \sup_{K \le n }P \left( \max_{k=1,\ldots,K} \left| n_k^{-1} \sum_{ i \in \mathcal{I}_k} \psi^a(W_i, \hat{\eta}_i) - \psi^a(W_i,\eta_i) \right|> \epsilon \right) = 0
\end{equation}
for any given $\epsilon>0$.
\end{lemma}
\begin{proof}
See Section (ref) in Appendix (ref).
\end{proof}
\begin{lemma}
Suppose Assumption (ref) holds. In addition, assume $K$ is such that $K \le n$, $K \to \infty $ as $n \to \infty$. If $\varphi_1 \le 1/2$, then
\begin{enumerate}
• $ \lim_{\tilde{M} \to \infty} \lim_{ n \to \infty} \sup_{K \le n } P \left( n^{4\min\{\varphi_1, \varphi_2\}} (n^{-1} \sum_{i=1}^n || \hat{\eta}_i - \eta_i||^4 ) > \tilde{M} \right) = 0$ ,
• $ \lim_{n \to \infty} \sup_{K \le n} n^{2\min\{\varphi_1, \varphi_2\}} E\left[ || \hat{\eta}_i - \eta_i ||^2 \right] < \infty $ ,
• $ \lim_{\tilde{M} \to \infty} \lim_{ n \to \infty} \sup_{K \le n } P \left( n^{2\min\{\varphi_1, \varphi_2\}} n^{-1} \sum_{i=1}^n || \hat{\eta}_i - \eta_i||^2 > \tilde{M} \right) = 0$ ,
• $ \lim_{ n \to \infty} \sup_{K \le n }P \left( n^{-1/2} \sum_{i=1}^n || \hat{\eta}_i - \eta_i||^2 > \epsilon \right) = 0$ when $1/4 < \min \{\varphi_1, \varphi_2\}$ , for any given $\epsilon>0$.
\end{enumerate}
where $\eta_i = \eta_0(X_i)$ and $\hat{\eta}_i$ is as in (ref).
\end{lemma}
\begin{proof}
See Section (ref) in Appendix (ref).
\end{proof}
\section{Additional Simulation Results}
\begin{figure}[h!]
\begin{subfigure}{0.32\textwidth}
\caption{Bias}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{MSE}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{Coverage probability ($\%$)}
\end{subfigure}
\caption{Bias and MSE of estimators for the ATT-DID based on DML1 as in (ref) for different values of $c$ in $h = c n_0^{-1/5}$. Coverage probability of confidence intervals as in (ref) for the ATT-DID with a nominal level of $95\%$. Sample size $n = 3,000$ and 5,000 simulations. It uses a Second Order Gaussian Kernel}
\end{figure}
\begin{figure}[h!]
\begin{subfigure}{0.32\textwidth}
\caption{Bias}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{MSE}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{Coverage probability ($\%$)}
\end{subfigure}
\caption{Bias and MSE of estimators for the ATT-DID based on DML2 as in (ref) for different values of $c$ in $h = c n_0^{-1/5}$. Coverage probability of confidence intervals as in (ref) for the ATT-DID with a nominal level of $95\%$. Sample size $n = 3,000$ and 5,000 simulations. It uses a Second Order Gaussian Kernel}
\end{figure}
\begin{figure}[t!]
\begin{subfigure}{0.32\textwidth}
\caption{Bias}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{MSE}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{Coverage Prob.($\%$)}
\end{subfigure}
\caption{Bias and MSE of estimators for the LATE based on DML1 and DML2 as in (ref) and (ref), respectively. Coverage probability of confidence intervals as in (ref) but using true $\sigma^2$ for the LATE with a nominal level of $95\%$. Discrepancy measure $\Lambda \neq 0$, sample size $n = 3,000$ and $5,000$ simulations.}
\end{figure}
\begin{figure}[h!]
\begin{subfigure}{0.32\textwidth}
\caption{Bias}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{MSE}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{Coverage probability ($\%$)}
\end{subfigure}
\caption{Bias and MSE of estimators for the LATE based on DML1 as in (ref) for different values of $c$ in $h = c n_0^{-1/5}$. Coverage probability of confidence intervals as in (ref) for the LATE with a nominal level of $95\%$. Discrepancy measure $\Lambda \neq 0$, sample size $n = 3,000$ and $5,000$. }
\end{figure}
\section{Proofs of Auxiliary Results}
\textbf{Notation:} Recall $\hat{\eta}_i = \hat{\eta}_k(X_i)$ for $i \in \mathcal{I}_k$ and $n_k = n/K$ is the number of observations on the fold $\mathcal{I}_k$. Denote ${\psi}^z_i = \psi^z(W_i,{\eta}_i)$ and $\hat{\psi}^z_i = \psi^z(W_i,\hat{\eta}_i)$ for $z=a,b$; $m_i = m(W_i,\theta_0,\eta_i)$, and $\hat{m}_i = m(W_i,\theta_0,\hat{\eta}_i)$; $\partial_\eta m_i = \partial_\eta m(W_i, \theta_0, \eta_i)$ and $\partial_\eta \hat{m}_i = \partial_\eta m(W_i, \theta_0, \hat{\eta}_i)$; $\partial_\eta^2 m_i = \partial_\eta^2 m(W_i, \theta_0, \eta_i)$ and $\partial_\eta^2 \hat{m}_i = \partial_\eta^2 m(W_i, \theta_0, \hat{\eta}_i)$ . Here, $|| \cdot || $ is the euclidean norm ($\ell_2$ norm), $J_0 = E[\psi^a_i]$, CLT is for Central Limit Theorem, LLN is for Law of Large Numbers, LIE is for Law of Iterated Expectations, C-S is for Cauchy-Schwartz inequality, RHS is for right-hand side.
\subsection{Proof of Theorem (ref)}
\begin{proof}
Using the notation of this section, the definitions of the DML1 estimator in (ref) and the moment function $m$ in (ref), it follows
$$ n^{1/2} \left( \hat{\theta}_{n,1} - \theta_0 \right) = K^{-1/2} \sum_{k=1}^K \frac{n_k^{-1/2} \sum_{i \in \mathcal{I}_K} \hat{m}_i }{n_k^{-1} \sum_{i \in \mathcal{I}_K} \hat{\psi}^a_i}~,$$
and similarly for the oracle version defined in (ref),
$$ n^{1/2} \left( \hat{\theta}_{n,1}^* - \theta_0 \right) = K^{-1/2} \sum_{k=1}^K \frac{n_k^{-1/2} \sum_{i \in \mathcal{I}_K} {m}_i }{n_k^{-1} \sum_{i \in \mathcal{I}_K} {\psi}^a_i}~.$$
Using the previous two expressions,
\begin{align*}
n^{1/2} \left( \hat{\theta}_{n,1} - \hat{\theta}_{n,1}^* \right)
&= I_1 + I_2
\end{align*}
where
\begin{align*}
I_1 &= K^{-1/2} \sum_{k=1}^K \frac{ n_k^{-1/2} \sum_{i \in \mathcal{I}_K} (\hat{m}_i - {m}_i) }{n_k^{-1} \sum_{i \in \mathcal{I}_K} \hat{\psi}^a_i }
I_2
&= K^{-1/2} \sum_{k=1}^K \frac{ \left( n_k^{-1/2} \sum_{i \in \mathcal{I}_K} {m}_i \right) \left( n_k^{-1}\sum_{i \in \mathcal{I}_k} {\psi}^a_i - \hat{\psi}^a_i \right)}{ \left(n_k^{-1} \sum_{i \in \mathcal{I}_K} \hat{\psi}^a_i \right) \left( n_k^{-1}\sum_{i \in \mathcal{I}_k} {\psi}^a_i \right) }
\end{align*}
In what follows, I will show that both $I_1$ and $I_2$ are $o_p(1)$, which is sufficient to complete the proof of the theorem.
\textit{Claim 1:} $I_1 = o_p(1)$. I first rewrite $I_1$ using the identity $a (1 + b)^{-1} = a - ab(1 + b)^{-1}$ with $a = \hat{I}_{1,k} $ and $b = I_{1,k} $, where
\begin{align*}
\hat{I}_{1,k} &= n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{m}_i - {m}_i)/J_0 ,\\
I_{1,k} &= n_k^{-1} \sum_{i \in \mathcal{I}_K} (\hat{\psi}^a_i-J_0)/J_0 .
\end{align*}
This implies
\begin{equation}
I_1 = K^{-1/2} \sum_{k=1}^K \hat{I}_{1,k}
- \hat{I}_{1,k} I_{1,k} \left( 1+I_{1,k} \right)^{-1}
\end{equation}
To show the claim, consider the following derivations
\begin{align*}
|I_1| &\overset{(1)}{\le} \left| K^{-1/2} \sum_{k=1}^K \hat{I}_{1,k} \right| + \left(K^{-1/2} \sum_{k=1}^K | \hat{I}_{1,k}| | I_{1,k}| \right) \times \max_{k = 1,\ldots, K} \left|n_k^{-1} \sum_{i \in \mathcal{I}_K } \hat{\psi}^a_i/J_0\right|^{-1} \\
&\overset{(2)}{=} \left|n^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{m}_i - {m}_i)/J_0\right| + \left(K^{-1/2} \sum_{k=1}^K | \hat{I}_{1,k}| | I_{1,k}| \right) \times O_p(1) ,\\
&\overset{(3)}{=} o_p(1) + o_p(1) \times O_p(1) ,
\end{align*}
where (1) holds by triangular inequality used on (ref) and definition of $I_{1,k}$, (2) holds by definition of $\hat{I}_{1,k}$ and part 2 of Lemma (ref), and (3) hold by part 3 of Lemma (ref) and (ref) presented below,
\begin{equation}
K^{-1/2} \sum_{k=1}^K | \hat{I}_{1,k}| | I_{1,k}| = o_p(1) .
\end{equation}
I use Taylor expansion and the mean value theorem to write $\hat{I}_{1,k} = \hat{I}_{1,1,k} + \hat{I}_{1,2,k}$ and $I_{1,k} = n_{k}^{-1/2} (I_{1,1,k} + I_{1,2,k} + I_{1,3,k} )$, where
\begin{align*}
\hat{I}_{1,1,k} &= n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{\eta}_i - \eta_i)^\top \partial_\eta m_i/J_0 ,\\
\hat{I}_{1,2,k} &= n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{\eta}_i - \eta_i)^\top (\partial_\eta^2 \tilde{m}_i/(2J_0)) (\hat{\eta}_i - \eta_i)/J_0 ,\\
I_{1,1,k} &= n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\psi^a_i - J_0)/J_0 ,\\
I_{1,2,k} &= n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{\eta}_i - \eta_i)^\top \partial_\eta \psi^a_i/J_0 ,\\
I_{1,3,k} &= n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{\eta}_i - \eta_i)^\top \partial_\eta^2 \tilde{\psi}^a_i/(2J_0) (\hat{\eta}_i - \eta_i) ,
\end{align*}
with $\partial_\eta^2 \tilde{m}_i = \partial_\eta^2 m(W_i, \theta_0, \tilde{\eta}_i)$ for some $\tilde{\eta}_i$, due to mean value theorem, and similar for $\partial_\eta^2 \tilde{\psi}^a_i$.
In what follows I prove $ K^{-1/2} \sum_{k=1}^K n_k^{-1/2} | \hat{I}_{1,j_1,k}| | I_{1,j_2,k}| = o_p(1)$ for $j_1 = 1, 2$ and $j_2 = 1,2,3$, which is sufficient to prove (ref).
\textit{Claim 1.1:} $K^{-1/2} \sum_{k=1}^K n_k^{-1/2} |\hat{I}_{1,1,k}| |I_{1,1,k}| = o_p(1) $. Consider the following
\begin{align*}
K^{-1/2} \sum_{k=1}^K n_k^{-1/2} |\hat{I}_{1,1,k}| |I_{1,1,k}| &\overset{(1)}{\le} \max_{k=1,\ldots,K} \left|n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{\eta}_i - \eta_i)^\top \partial_\eta m_i/J_0 \right| \times n^{-1/2} \sum_{k=1}^K |I_{1,1,k}|\\
&\overset{(2)}{=} o_p(1) \times (n^{-1/2} K) \times O_p(1) \\
&\overset{(3)}{=} o_p(1) ,
\end{align*}
where (1) holds by definition of $\hat{I}_{1,1,k}$, (2) holds by Lemma (ref) and the derivation presented below, and (3) holds since $K = O(n^{1/2})$.
\begin{align*}
E\left[ n^{-1/2} \sum_{k=1}^K |I_{1,1,k}| \right] & \overset{(1)}{=} n^{-1/2} K E[ |I_{1,1,k}|] \\
&\overset{(2)}{\le} n^{-1/2} K E\left[ \left(n_k^{-1/2} \sum_{i \in \mathcal{I}_k} (\psi^a_i - J_0)/J_0 \right)^2 \right]^{1/2} \\
&\overset{(3)}{\le} n^{-1/2} K O(1)
\end{align*}
where (1) holds since $I_{1,1,k}$ are i.i.d. random variables, (2) holds by Jensen's inequality and definition of $I_{1,1,k}$, and (3) holds since $\{ \psi^a_i - J_0 : i \in \mathcal{I}_k \}$ are zero mean i.i.d. random variables and by parts (a) and (c) of Assumption (ref).
\textit{Claim 1.2:} $K^{-1/2} \sum_{k=1}^K n_k^{-1/2}|\hat{I}_{1,1,k}| \times |I_{1,2,k}| = o_p(1) $. It follows by
\begin{align*}
K^{-1/2} \sum_{k=1}^K n_k^{-1/2}|\hat{I}_{1,1,k}| |I_{1,2,k}| &\le \max_{k=1,\ldots,K} |\hat{I}_{1,1,k}| \times \max_{k=1,\ldots,K} |I_{1,2,k}| \times n^{-1/2} K \\
&\overset{(1)}{=} o_p(1) \times o_p(1) \times O(1)
\end{align*}
where (1) holds by Lemma (ref) and because $K = O(n^{1/2})$.
\textit{Claim 1.3:} $K^{-1/2} \sum_{k=1}^K n_k^{-1/2}|\hat{I}_{1,1,k}| |I_{1,3,k}| = o_p(1) $. It follows by
\begin{align*}
K^{-1/2} \sum_{k=1}^K n_k^{-1/2}|\hat{I}_{1,1,k}| |I_{1,3,k}| &\le \max_{k=1,\ldots,K} |\hat{I}_{1,1,k}| \times n^{-1/2} \sum_{k=1}^K \left| n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{\eta}_i - \eta_i)^\top \partial_\eta^2 \tilde{\psi}^a_i/(2J_0) (\hat{\eta}_i - \eta_i) \right | \\
&\overset{(1)}{\le} \max_{k=1,\ldots,K} |\hat{I}_{1,1,k}| \times n^{-1/2} \sum_{k=1}^K (C_2 p/(2|J_0|)) \times n_k^{-1/2} \sum_{i \in \mathcal{I}_k } || \hat{\eta}_i - \eta_i ||^2 \\
&= \max_{k=1,\ldots,K} |\hat{I}_{1,1,k}| \times (C_2 p /(2 |J_0|) (K n^{-1/2})^{1/2} n^{1/4} n^{-1} \sum_{i=1}^n || \hat{\eta}_i - \eta_i ||^2 \\
&\overset{(2)}{=} o_p(1) \times O(1) \times n^{1/4} \times O_p(n^{-2\min\{ \varphi_1,\varphi_2\} } ) \\
&\overset{(3)}{=} o_p(1)
\end{align*}
where (1) holds by part (e) of Assumption (ref) and Loeve’s inequality (Davidson1994), (2) holds by Lemmas (ref) and (ref) and because $K = O(n^{1/2})$, and (3) holds since $\min\{\varphi_1, \varphi_2 \} > 1/4$.
\textit{Claim 1.4:} $K^{-1/2} \sum_{k=1}^K n_k^{-1/2}|\hat{I}_{1,2,k}| |I_{1,1,k}| = o_p(1) $. Consider the following derivations,
\begin{align*}
& n^{-1/2} \sum_{k=1}^K|\hat{I}_{1,2,k}| |I_{1,1,k}| \\
&\overset{(1)}{\le} n^{-1/2} \left( \sum_{k=1}^K \left| n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{\eta}_i - \eta_i)^\top (\partial_\eta^2 \tilde{m}_i/(2J_0)) (\hat{\eta}_i - \eta_i)/J_0 \right|^2 \right)^{1/2} \left( \sum_{k=1}^K |I_{1,1,k}|^2 \right)^{1/2} \\
&\overset{(2)}{\le} (C_2 p/2) n_k^{-1/2} \left( \sum_{k=1}^K | n_k^{-1/2} \sum_{i \in \mathcal{I}_K } ||\hat{\eta}_i - \eta_i||^2 |^2 \right)^{1/2} K^{-1/2}\left( \sum_{k=1}^K |I_{1,1,k}|^2 \right)^{1/2} \\
&\overset{(3)}{\le} (C_2 p/2) K^{1/2} \left( \sum_{k=1}^K n^{-1} \sum_{i \in \mathcal{I}_K } ||\hat{\eta}_i - \eta_i||^4 \right)^{1/2} \left( K^{-1} \sum_{k=1}^K |I_{1,1,k}|^2 \right)^{1/2} \\
&\overset{(4)}{=} (K n^{-1/2})^{1/2} n^{1/4} \times O_p(n^{-2\min \{ \varphi_1,\varphi_2 \}} ) \times \left( K^{-1} \sum_{k=1}^K |I_{1,1,k}|^2 \right)^{1/2} \\
&\overset{(5)}{=} O(1) \times n^{1/4} \times O_p(n^{-2\min \{ \varphi_1,\varphi_2 \}} ) \times O_p(1) \\
&\overset{(6)}{=} o_p(1)
\end{align*}
where (1) holds by Cauchy-Schwartz and definition of $\hat{I}_{1,2,k}$, (2) holds by part (e) of Assumption (ref) and Loeve’s inequality (Davidson1994), (3) holds by Jensen's inequality, (4) holds by Lemma (ref), (5) holds because $K = O(n^{1/2})$, $E[K^{-1} \sum_{k=1}^K |I_{1,1,k}|^2] = O(1)$ by definition of $I_{1,1,k}$ and due to parts (a) and (c) of Assumption (ref), and (6) holds since $\min\{\varphi_1, \varphi_2 \} > 1/4$.
\textit{Claim 1.5:} $K^{-1/2} \sum_{k=1}^K n_k^{-1/2}|\hat{I}_{1,2,k}| |I_{1,2,k}| = o_p(1) $. The proof is similar to the proof of Claim 1.3; therefore, it is omitted.
\textit{Claim 1.6:} $K^{-1/2} \sum_{k=1}^K n_k^{-1/2}|\hat{I}_{1,2,k}| \times |I_{1,3,k}| = o_p(1) $. Consider the derivations,
\begin{align*}
K^{-1/2} \sum_{k=1}^K n_k^{-1/2}|\hat{I}_{1,2,k}| |I_{1,3,k}|
&\overset{(1)}{\le} n^{-1/2} \sum_{k=1}^K (C_2 p /(2|J_0|))^2 \times \left(n_k^{-1/2} \sum_{i \in \mathcal{I}_k } || \hat{\eta}_i - \eta_i ||^2 \right)^2 \\
&\overset{(2)}{\le} (C_2 p /(2J_0))^2 \times n^{1/2} n^{-1} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k } || \hat{\eta}_i - \eta_i ||^4 \\
&\overset{(3)}{=} (C_2 p /(2J_0))^2 \times n^{1/2} \times O_p(n^{-4\min \{ \varphi_1,\varphi_2 \} } ) \\
&\overset{(4)}{=} o_p(1) ,
\end{align*}
where (1) holds by using the definition of $\hat{I}_{1,2,k}$ and $I_{1,3,k}$, part (e) of Assumption (ref), and Loeve’s inequality (Davidson1994), (2) holds by Jensen's inequality, (3) holds by Lemma (ref), and (4) holds since $\min\{\varphi_1, \varphi_2 \} > 1/4$.
\textit{Claim 2:} $I_2 = o_p(1)$. Consider the following representation of $I_2$,
\begin{align*}
I_2 &= K^{-1/2} \sum_{k=1}^K \frac{ \left( n_k^{-1/2} \sum_{i \in \mathcal{I}_K} {m}_i \right) \left( n_k^{-1}\sum_{i \in \mathcal{I}_k} {\psi}^a_i - \hat{\psi}^a_i \right)}{ \left(n_k^{-1} \sum_{i \in \mathcal{I}_K} \hat{\psi}^a_i \right) \left( n_k^{-1}\sum_{i \in \mathcal{I}_k} {\psi}^a_i \right) } \\
&= K^{-1/2} \sum_{k=1}^K \frac{ n_k^{-1/2} I_{2,k} \hat{I}_{2,k}}{ \left(n_k^{-1} \sum_{i \in \mathcal{I}_K} \hat{\psi}^a_i/J_0 \right) \left( n_k^{-1}\sum_{i \in \mathcal{I}_k} {\psi}^a_i/J_0 \right) } ,
\end{align*}
where
\begin{align*}
I_{2,k} &= n_k^{-1/2} \sum_{i \in \mathcal{I}_K} {m}_i/J_0 \\
\hat{I}_{2,k} &= n_k^{-1/2}\sum_{i \in \mathcal{I}_k} ({\psi}^a_i - \hat{\psi}^a_i)/J_0 .
\end{align*}
To show the claim, consider the following derivation
\begin{align*}
|I_2| &\overset{(1)}{\le} \max_{k=1,\ldots,K} \left|n_k^{-1} \sum_{i \in \mathcal{I}_K} \hat{\psi}^a_i/J_0 \right|^{-1} \times \max_{k=1,\ldots,K} \left|n_k^{-1} \sum_{i \in \mathcal{I}_K} {\psi}^a_i/J_0 \right|^{-1} \times K^{-1} \sum_{k=1}^K n_k^{-1/2} |I_{2,k}| |\hat{I}_{2,k}| \\
&\overset{(2)}{=} O_p(1) \times O_p(1) \times o_p(1) ,
\end{align*}
where (1) holds by triangular inequality and definition of $I_2$, and (2) by Lemma (ref) and (ref) presented below,
\begin{equation}
K^{-1} \sum_{k=1}^K n_k^{-1/2} |\hat{I}_{2,k}| |I_{2,k}| = o_p(1) .
\end{equation}
As in the proof of claim 1, I use Taylor approximation and mean value theorem to write $\hat{I}_{2,k} = \hat{I}_{2,1,k} + \hat{I}_{2,2,k}$, where
\begin{align*}
\hat{I}_{2,1,k} &= n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{\eta}_i - \eta_i)^\top \partial_\eta \psi^a_i/J_0 ,\\
\hat{I}_{2,2,k} &= n_k^{-1/2} \sum_{i \in \mathcal{I}_K } (\hat{\eta}_i - \eta_i)^\top (\partial_\eta^2 \tilde{\psi}^a_i/(2J_0)) (\hat{\eta}_i - \eta_i)/J_0 .
\end{align*}
Finally, in what follows I prove $ K^{-1/2} \sum_{k=1} n_k^{-1/2} | \hat{I}_{2,j,k}| | I_{2,k}| = o_p(1)$ for $j = 1, 2$, which is sufficient to prove (ref).
\textit{Claim 2.1:} $K^{-1/2} \sum_{k=1}^K n_k^{-1/2} | \hat{I}_{2,1,k}| | I_{2,k}| = o_p(1) $. The proof is similar to the one in Claim 1.1; therefore, it is omitted.
\textit{Claim 2.2:} $K^{-1/2} \sum_{k=1}^K n_k^{-1/2} | \hat{I}_{2,2,k}| | I_{2,k}| = o_p(1) $. The proof is similar to the one in Claim 1.4; therefore, it is omitted.
\end{proof}
\subsection{Proof of Theorem (ref)}
\begin{proof}
Notation: In the proof of this theorem, $x_{n,K} = o_p(1)$ denotes a sequence of random variables $x_{n,K}$ converging to zero uniformly on $K \to \infty$ as $n \to \infty$ (equivalently, $\lim_{n \to \infty} \sup_{K \le n} P( |x_{n,K}| > \epsilon) = 0$ for any given $\epsilon>0$).
Using the definitions of the DML2 estimator in (ref) and the moment function $m$ in (ref), it follows
$$ n^{1/2} \left( \hat{\theta}_{n,2} - \theta_0 \right) = \frac{ n^{-1/2} \sum_{i =1}^n \hat{m}_i }{ n^{-1} \sum_{i =1}^n \hat{\psi}^a_i}~,$$
and similarly for the oracle version defined in (ref),
$$ n^{1/2} \left( \hat{\theta}_{n,2}^* - \theta_0 \right) = \frac{ n^{-1/2} \sum_{i =1}^n {m}_i }{ n^{-1} \sum_{i =1}^n {\psi}^a_i}~.$$
Using the previous two expressions, it follows
\begin{equation*}
n^{1/2} \left( \hat{\theta}_{n,2} - \hat{\theta}_{n,2}^* \right) = I_1 + I_2
\end{equation*}
where
\begin{align}
I_1
&= \frac{ n^{-1/2}\sum_{i=1}^n (\hat{m}_i - m_i )/J_0}{ n^{-1} \sum_{i=1}^n \hat{\psi}^a_i/J_0 } \\
I_2
&= \frac{ \left ( n^{-1/2}\sum_{i=1}^n {m}_i/J_0 \right) \left( n^{-1} \sum_{i=1}^n ({\psi}^a_i - \hat{\psi}^a_i)/J_0 \right) }{ \left( n^{-1} \sum_{i=1}^n \hat{\psi}^a_i/J_0 \right) \left( n^{-1} \sum_{i=1}^n {\psi}^a_i/J_0 \right)}
\end{align}
In what follows, I show that $I_1 = \mathcal{T}_{n,K}^{l} + \mathcal{T}_{n,K}^{nl} + o_p(n^{-\zeta})$ and $I_2 = o_p(n^{-\zeta})$, which is sufficient to complete the proof of the theorem since both $\mathcal{T}_{n,K}^{l}$ and $\mathcal{T}_{n,K}^{nl}$ are $O_p(n^{-\varphi_1})$ and $O_p(n^{1/2-2\varphi_1})$, respectively, under Assumptions (ref) and (ref) and by the proof of Propositions (ref) and (ref). Furthermore, if Assumption (ref) holds, part 2 of Proposition (ref) implies $\text{Var}[n^{2\varphi_1-1/2} \mathcal{T}_{n,K}^{nl}] = G_\delta (K^2-3K + 3)(K-1)^{-1-4\varphi_1} K^{4\varphi_1-1} + n^{4\varphi_1-1} r_{n,K}^{nl}$, which implies that $\lim_{ n \to \infty} \inf_{K \le n } \text{Var}[n^{2\varphi_1-1/2} \mathcal{T}_{n,K}^{nl}] > 0$. Part 1 of Proposition (ref) implies that $ \sup_{K \le n } |n^{2\varphi_1-1/2} E[ \mathcal{T}_{n,K}^{nl}]| < \infty $; then, $\lim_{ n \to \infty} \sup_{K \le n } E[ (n^{2\varphi_1-1/2} \mathcal{T}_{n,K}^{nl})^2] < \infty$. Similarly, the proof of Propositions (ref) guarantees $\lim_{ n \to \infty} \sup_{K \le n } E[ (n^{\varphi_1} \mathcal{T}_{n,K}^{l})^2] < \infty$.
\textit{Claim 1:} $I_1 = \mathcal{T}_{n,K}^{l} + \mathcal{T}_{n,K}^{nl} + o_p(n^{-\zeta})$. I first rewrite the RHS of (ref) using the identity $a(1+b)^{-1} = a - a b(1+b)^{-1}$, where $a = n^{-1/2} \sum_{ i= 1}^n (\hat{m}_i - m_i)/J_0$ and $b = n^{-1} \sum_{i =1}^n (\hat{\psi}^a_i - J_0)/J_0$. That is
$$ I_1 = a - ab(1+b)^{-1}$$
I then conclude the proof of the claim by using claims 1.1 and 1.2, stated below.
\textit{Claim 1.1:} $a = \mathcal{T}_{n,K}^{l} + \mathcal{T}_{n,K}^{nl} + o_p(n^{-\zeta})$. This result holds by part 4 of Lemma (ref) since $a = n^{-1/2}\sum_{i=1}^n (\hat{m}_i - m_i )/J_0$.
\textit{Claim 1.2:} $ab(1+b)^{-1} = o_p(n^{-\zeta})$. Note that part 4 of Lemma (ref) implies $a = O_p(n^{1/2-2\varphi_1})$. Note also that part 2 of Lemma (ref) and CLT imply $b = O_p(n^{-1/2})$, which guarantees that $(1+b)^{-1} = O_p(1)$; therefore, $ab(1+b)^{-1} = O_p(n^{-2\varphi_1})$, which is $o_p(n^{-\zeta})$ since $\varphi_1 < 1/2$.
\textit{Claim 2:} $I_2 = o_p(n^{-\zeta})$. I first rewrite $I_2$ defined in (ref) as follows,
$$ I_2 = a b (1+c-b)^{-1} (1 + c)^{-1}~,$$
where $a = n^{-1/2}\sum_{i=1}^n m_i/J_0 $, $b = n^{-1} \sum_{i=1}^n (\psi^a_i - \hat{\psi}^a_i)/J_0$, and $c = n^{-1} \sum_{i=1}^n ({\psi}^a_i-J_0)/J_0$. CLT implies that $a = O_p(1)$ and $c = O_p(n^{-1/2})$. Part 2 of Lemma (ref) implies $b = O_p(n^{-2\varphi_1})$. Therefore, $ab = O_p(n^{-2\varphi_1})$, and both $(1+c-b)^{-1}$ and $ (1 + c)^{-1}$ are $O_p(1)$. This implies $I_2 $ is $O_p(n^{-2\varphi_1})$, which is $o_p(n^{-\zeta})$ since $\varphi_1 < 1/2$.
\end{proof}
\subsection{Proof of Proposition (ref)}
\begin{proof}
\textbf{Part 1:} By the definition of the oracle version of the DML1 estimator in (ref) and the moment function $m$ in (ref), it follows
\begin{equation}
n^{1/2} \left( \hat{\theta}_{n,1}^* - \theta_0 \right) = K^{-1/2} \sum_{k=1}^K \frac{n_k^{-1/2} \sum_{i \in \mathcal{I}_K} {m}_i }{n_k^{-1} \sum_{i \in \mathcal{I}_K} {\psi}^a_i} .
\end{equation}
I first rewrite the RHS of (ref) using the identity $a_k(1+b_k)^{-1} = a_k - a_k b_k + a_k b_k^2(1+b_k)^{-1}$ with $a_k = n_k^{-1/2} \sum_{i \in \mathcal{I}_k} m_i/J_0$ and $b_k = n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0$. That is
\begin{align*}
n^{1/2} \left( \hat{\theta}_{n,1}^* - \theta_0 \right)
&= I_1 + I_2 + I_3 ,
\end{align*}
where
\begin{align*}
I_1 &= K^{-1/2} \sum_{k=1}^K a_k \\
I_2 &= - K^{-1/2} \sum_{k=1}^K a_k b_k \\
I_3 &= K^{-1/2} \sum_{k=1}^K a_k b_k^2(1+b_k)^{-1}
\end{align*}
By CLT, it follows $I_1 = n^{-1/2} \sum_{i=1}^n m_i/J_0 \overset{d}{\to} N(0,\sigma^2)$ as $n \to \infty$, where $\sigma^2$ is as in (ref). Therefore, if $I_2 - K/\sqrt{n} \Lambda $ and $I_3$ are $o_p(1)$, then
$$ n^{1/2} \left( \hat{\theta}_{n,1}^* - \theta_0 \right) = n^{-1/2} \sum_{i=1}^n m_i/J_0 + K/\sqrt{n} \Lambda + o_p(1)~, $$
which is sufficient to complete the proof of part 1 since $K/\sqrt{n} \to c$ as $n \to \infty$. In what follows, Claim 1 shows $I_2 - K/\sqrt{n} \Lambda = o_p(1)$ and Claim 2 shows $I_3$ are $o_p(1)$.
\textit{Claim 1:} $I_2 - K/\sqrt{n} \Lambda = o_p(1)$. First, note that $E[-a_k b_k] = K^{1/2}/\sqrt{n} \Lambda$ due to the following derivations,
\begin{align*}
E[a_k b_k] &\overset{(1)}{=} E\left[ \left( n_k^{-1/2} \sum_{i \in \mathcal{I}_k} m_i/J_0 \right) \left( n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right) \right] \\
&\overset{(2)}{=} n_k^{-1/2} E\left[ \left( m_i/J_0 \right) \left( (\psi^a_i-J_0)/J_0 \right) \right] \\
&\overset{(3)}{=} -n^{-1/2} K^{1/2} \Lambda
\end{align*}
where (1) holds by definition of $a_k$ and $b_k$, (2) holds since $\{ (m_i, \psi^a_i - J_0) : i \in \mathcal{I}_k\}$ are zero mean i.i.d. random vectors, and (3) holds by the definition of $\Lambda$ in (ref) and condition (ref).
Therefore, $ E[I_2] = -K^{-1/2} \sum_{k=1}^K E[a_k b_k] = K/\sqrt{n} \Lambda$, which implies that the claim is equivalent to show that $I_2 - E[I_2]$ is $o_p(1)$, which follows by the following derivations
\begin{align*}
E\left[ \left(I_2 - E[I_2] \right)^2 \right]
&\overset{(1)}{=} E\left[ \left( K^{-1/2} \sum_{k=1}^K (a_k b_k - E[a_k b_k]) \right)^2 \right] \\
&\overset{(2)}{=} K^{-1} \sum_{k=1}^K E\left[ \left(a_k b_k - E[a_k b_k] \right)^2 \right] \\
&\overset{(3)}{\le} E\left[ \left(a_k b_k \right)^2 \right] \\
&\overset{(4)}{=} n_k^{-1} E\left[ \left( n_k^{-1/2} \sum_{i \in \mathcal{I}_k} m_i/J_0 \right)^2 \left( n_k^{-1/2} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right)^2 \right] \\
&\overset{(5)}{\le} n_k^{-1} E\left[ \left( n_k^{-1/2} \sum_{i \in \mathcal{I}_k} m_i/J_0 \right)^4 \right]^{1/2} \times E\left[ \left( n_k^{-1/2} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right)^4 \right]^{1/2} \\
&\overset{(6)}{=} n_k^{-1} \times O(1) \times O(1)
\end{align*}
where (1) holds by definition of $I_2$; (2) and (3) hold since $\{a_kb_k - E[a_kb_K] : 1 \le k \le K \}$ are zero mean i.i.d random variables due to the definition of $a_k$ and $b_k$; (4) holds by definition of $a_k$ and $b_k$; (5) holds by Cauchy-Schwartz; and (6) holds since $\{(m_i,\psi^a_i-J_0) : i \in \mathcal{I}_k\}$ are zero mean i.i.d. random vectors, parts (a) and (c) of Assumption (ref), and $n_k \to \infty$. This completes the proof of Claim 1.
\textit{Claim 2:} $I_3 = o_p(1)$. Consider the following derivation
\begin{align*}
|I_3| &\overset{(1)}{\le} \max_{k=1,\ldots,K} \left| 1 + n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right|^{-1} \times K^{-1/2} \sum_{k=1}^K |a_k| b_k^2 \\
&= O_p(1) \times o_p(1) ,
\end{align*}
where (1) holds by definition of $I_3$ and triangular inequality, and (2) holds by Lemma (ref) and (ref) presented below,
\begin{equation}
K^{-1/2} \sum_{k=1}^K |a_k| b_k^2 = o_p(1) .
\end{equation}
To prove (ref), consider the following
\begin{align*}
E\left[ K^{-1/2} \sum_{k=1}^K |a_k| b_k^2 \right] &\overset{(1)}{\le} K^{-1/2} \sum_{k=1}^K E[|a_k|^2]^{1/2} E[|b_k|^4]^{1/2} \\
&\overset{(2)}{=} K^{1/2}n_k^{-1} E\left[ \left( n_k^{-1/2} \sum_{i \in \mathcal{I}_k} m_i/J_0 \right)^{2} \right]^{1/2} \times E\left[ \left( n_k^{-1/2} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right)^4 \right]^{1/2} \\
&\overset{(3)}{=} (K n^{-1/2} )^{3/2} n^{-1/4} O(1) \times O(1) \\
&\overset{(4)}{=} o(1) ,
\end{align*}
where (1) holds by Cauchy-Schwartz, (2) holds since $\{(a_k,b_k\}: 1 \le k\le K\}$ are i.i.d random vectors and the definition of $a_k$ and $b_k$, (3) holds since $\{(m_i,\psi^a_i-J_0) : i \in \mathcal{I}_k\}$ are zero mean i.i.d. random vectors, and parts (a) and (c) of Assumption (ref), and (4) holds since $K=O(n^{1/2})$. This completes the proof of Claim 2.
\textbf{Part 2:} By the definition of $\hat{\theta}_{n,2}^*$ in (ref), and the moment function $m$ in (ref), it follows
\begin{align*}
n^{1/2} \left( \hat{\theta}_{n,2}^* - \theta_0 \right) &= \frac{n^{-1/2} \sum_{i = 1}^n {m}_i/J_0 }{n^{-1} \sum_{i = 1}^n {\psi}^a_i/J_0}
\end{align*}
Since the denominator converges to 1 in probability by the LLN and the numerator converges to $N(0,\sigma^2)$ in distribution due to the CLT, it follows that $n^{1/2} \left( \hat{\theta}_{n,2}^* - \theta_0 \right)$ converges in distribution to $N(0,\sigma^2)$. This completes the proof of part 2
\end{proof}
\subsection{Proof of Proposition (ref)}
\begin{proof}
\textbf{Part 1:} By the definition of the oracle version of DML2 estimator in (ref), and the moment function $m$ in (ref), it follows
\begin{equation}
n^{1/2} \left( \hat{\theta}_{n,2}^* - \theta_0 \right) = \frac{n^{-1/2} \sum_{i =1}^n {m}_i }{n^{-1} \sum_{i=1}^n {\psi}^a_i} .
\end{equation}
I rewrite the RHS of (ref) using the identity $a(1+b)^{-1} = a - a b + a b^2(1+b)^{-1}$ with $a = n^{-1/2} \sum_{i=1}^n m_i/J_0$ and $b = n^{-1} \sum_{ i = 1}^n (\psi^a_i-J_0)/J_0$. That is
\begin{align*}
n^{1/2} \left( \hat{\theta}_{n,2}^* - \theta_0 \right) &= a - ab + ab^2 (1+b)^{-1} \\
&\overset{(1)}{=} \mathcal{T}_n^* + \mathcal{T}_{n}^{dml2} + ab^2 (1+b)^{-1}
\end{align*}
where (1) holds by the definition of $ \mathcal{T}_n^*$ and $\mathcal{T}_{n}^{dml2} $. It is sufficient to show $ab^2 (1+b)^{-1} $ is $ O_p(n^{-1})$ to complete the proof, which follows by CLT that implies $a = O_p(1)$ and $b = O_p(n^{-1/2})$, and $(1+b)^{-1} = O_p(1)$.
Finally, consider the following derivations
\begin{align*}
E[\mathcal{T}_{n}^{dml2} ] &\overset{(1)}{=} - n^{-1/2} E\left[ \left( n^{-1/2} \sum_{i=1}^n m_i/J_0 \right) \left( n^{-1/2} \sum_{ i=1}^n (\psi^a_i-J_0)/J_0 \right) \right] \\
&\overset{(2)}{=} - n^{-1/2} \Lambda
\end{align*}
where (1) holds by the definition of $\mathcal{T}_{n}^{dml2}$, and (2) holds since $\{(m_i, (\psi^a_i - J_0)/J_0 ): 1\le i \le n \}$ are zero mean i.i.d. random vectors and by the definition of $\Lambda$ in (ref).
\textbf{Part 2:} By definition of $\sigma^2$ and since $\{(m_i/J_0): 1 \le i \le n \}$ are zero mean i.i.d. random variables, it follows that $E[\mathcal{T}_n^*] = 0$ and $E[(\mathcal{T}_n^*)^2] = \sigma^2$.
\textbf{Part 3:} First note that $\text{Cov}(\mathcal{T}_n^*, \mathcal{T}_{n}^{dml2}) = E[(\mathcal{T}_n^*) (\mathcal{T}_{n}^{dml2})]$. Now, consider the following derivations,
\begin{align*}
E[(\mathcal{T}_n^*) (\mathcal{T}_{n}^{dml2})] &\overset{(1)}{=} - n^{-1/2} E\left[ \left( n^{-1/2} \sum_{i=1}^n m_i/J_0 \right)^2 \left( n^{-1/2} \sum_{ i = 1}^n (\psi^a_i-J_0)/J_0\right) \right] \\
&\overset{(2)}{=} - n^{-1} E\left[ \left( m_i/J_0 \right)^2 \left( (\psi^a_i-J_0)/J_0\right) \right] \\
&\overset{(3)}{=} - n^{-1} \Xi_1
\end{align*}
where (1) holds by definition of $\mathcal{T}_n^*$ and $\mathcal{T}_{n}^{dml2}$, (2) holds since $\{(m_i/J_0, (\psi^a_i-J_0)/J_0 ): 1 \le i \le n \}$ are zero mean i.i.d. random vectors, and (3) holds by definition of $\Xi_1$ in (ref).
Similarly, consider the following derivations
\begin{align*}
Var[\mathcal{T}_{n}^{dml2}] &\overset{(1)}{=} n^{-1} E\left[ \left( n^{-1/2} \sum_{i=1}^n m_i/J_0 \right)^2 \left( n^{-1/2} \sum_{i=1}^n (\psi^a_i-J_0)/J_0 \right)^2 \right ] - n^{-1} \Lambda^2 \\
&\overset{(2)}{=} n^{-1} \left( E[(m_i/J_0)^2] E[((\psi^a_i - J_0)/J_0)^2] + 2 \Lambda^2 + O(n^{-1}) \right) - n^{-1} \Lambda^2 \\
&\overset{(3)}{=} n^{-1} ( \sigma^2 \sigma_a^2 + \Lambda^2) + O(n^{-2})
\end{align*}
where (1) holds by definition of $\mathcal{T}_{n}^{dml2}$ and $\Lambda$ in (ref) and (ref), respectively, (2) holds since $\{(m_i/J_0, (\psi^a_i-J_0)/J_0) : 1 \le i \le n \}$ are zero mean i.i.d. random vectors and by definition of $\Lambda$, and (3) holds by definition of $\sigma^2$ and $\sigma_a^2$ in (ref) and (ref), respectively.
\end{proof}
\subsection{Proof of Proposition (ref)}
\begin{proof}
For $i \in \mathcal{I}_k$, denote $\Delta_{i} = \Delta_{i}^b + \Delta_{i}^l$, where
\begin{align*}
\Delta_{i}^l &= n_0^{-\varphi_1} n_0^{-1/2} \sum_{j \notin \mathcal{I}_k} \delta_{n_0,j,i} ,\\
\Delta_{i}^b &= n_0^{-\varphi_2} n_0^{-1} \sum_{j \notin \mathcal{I}_k} b_{n_0,j,i} ,
\end{align*}
Here, $\delta_{n_0,j,i} = \delta_{n_0}(W_{j},X_i)$ and $b_{n_0,j,i} = b_{n_0}(X_{j},X_i)$, and $\delta_{n_0}$ and $b_{n_0}$ are functions satisfying Assumption (ref).
\textbf{Part 1:} Using the previous notation, it holds that $E[(\Delta_{i})^{\top} \partial_\eta m_i/J_0] = 0$. To see this, consider the following derivations
\begin{align*}
E[(\Delta_{i})^{\top} \partial_\eta m_i/J_0] &\overset{(1)}{=} E\left[ (\Delta_{i})^{\top} E\left[\partial_\eta m_i/J_0 \mid (W_j : j \notin \mathcal{I}_k), X_i \right] \right] \\
&\overset{(2)}{=} E\left[ (\Delta_{i})^{\top} E\left[\partial_\eta m_i/J_0 \mid X_i \right] \right] \\
&\overset{(3)}{=} 0 ,
\end{align*}
where (1) holds by the law of interactive expectations and because $\Delta_{i}$ is non-stochastic conditional on $(W_j : j \notin \mathcal{I}_k)$ and $X_i$, (2) holds since $\{ W_j : 1 \le j \le n \}$ are i.i.d. random vectors and $i \in \mathcal{I}_k$, and (3) holds by the Neyman orthogonality condition (part (b) of Assumption (ref)).
Therefore, $E[\mathcal{T}_{n,K}^l] = 0$ holds due to the definition of $\mathcal{T}_{n,K}^l$ in (ref) and the previous result,
\begin{align*}
E[\mathcal{T}_{n,K}^l] &= n^{-1/2} \sum_{i=1}^n E[(\Delta_{i})^{\top} \partial_\eta m_i/J_0] \\
&= 0 .
\end{align*}
\textbf{Part 2:} By part 1, $\text{Var}[\mathcal{T}_{n,K}^{l}] = E[ (\mathcal{T}_{n,K}^{l})^2 ]$. Now, consider the following decomposition:
\begin{align*}
E[ (\mathcal{T}_{n,K}^{l})^2 ] &= E \left[ \left( n^{-1/2} \sum_{i=1}^n (\Delta_{i})^{\top} \partial_\eta m_i/J_0 \right)^2 \right] \\
&= n^{-1} \sum_{i_1=1}^n \sum_{i_2=1}^n E \left[\left( (\Delta_{i_1})^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( (\Delta_{i_2})^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
&= I_{1} + 2 I_{2} + I_{3} ,
\end{align*}
where I use $\Delta_i = \Delta_i^l +\Delta_i^b$ in the last equality, with $I_1$, $I_2$, and $I_3$ defined below,
\begin{align*}
I_{1} &= n^{-1} \sum_{i_1=1}^n \sum_{i_2=1}^n E \left[\left( (\Delta_{i_1}^l)^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( (\Delta_{i_2}^l)^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
I_{2} &= n^{-1} \sum_{i_1=1}^n \sum_{i_2=1}^n E \left[\left( (\Delta_{i_1}^l)^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( (\Delta_{i_2}^b)^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
I_{3} &= n^{-1} \sum_{i_1=1}^n \sum_{i_2=1}^n E \left[\left( (\Delta_{i_1}^b)^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( (\Delta_{i_2}^b)^{\top} \partial_\eta m_{i_2}/J_0 \right) \right]
\end{align*}
In what follows, I show $I_{1} = n_0^{-2\varphi_1} G_\delta^l + o(n^{-2\varphi_1})$ with $G_\delta^l$ defined as in (ref), $I_{2} = 0$, and $I_{3} = O(n^{-2\varphi_1})$, which is sufficient to complete the proof of Part 2.
\textit{Claim 1:} $I_{1} = n_0^{-2\varphi_1} G_\delta^l + o(n^{-2\varphi_1})$. Consider the following derivations,
\begin{align*}
& I_{1}\\
&= n^{-1} \sum_{i_1=1}^n \sum_{i_2=1}^n E \left[\left( (\Delta_{i_1}^l)^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( (\Delta_{i_2}^l)^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
&= n^{-1} \sum_{k_1,k_2=1}^K \sum_{i_1 \in \mathcal{I}_{k_1}} \sum_{i_2 \in \mathcal{I}_{k_2}} n_0^{-2\varphi_1} n_0^{-1} \sum_{j_1 \notin \mathcal{I}_{k_1}} \sum_{j_2 \notin \mathcal{I}_{k_2}} E \left[\left( \delta_{n_0,j_1,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( \delta_{n_0,j_2,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
&\overset{(1)}{=} n^{-1} \sum_{k_1,k_2=1}^K \sum_{i_1 \in \mathcal{I}_{k_1}} \sum_{i_2 \in \mathcal{I}_{k_2}} n_0^{-2\varphi_1} n_0^{-1} \sum_{j_1 \notin \mathcal{I}_{k_1}} \sum_{j_2 \notin \mathcal{I}_{k_2}} E \left[\left( \delta_{n_0,j_1,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( \delta_{n_0,j_2,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] I\{ k_1 \neq k_2 \} \\
& + n^{-1} \sum_{k=1}^K \sum_{i_1 \in \mathcal{I}_{k}} \sum_{i_2 \in \mathcal{I}_{k}} n_0^{-2\varphi_1} n_0^{-1} \sum_{j_1 \notin \mathcal{I}_{k}} \sum_{j_2 \notin \mathcal{I}_{k}} E \left[\left( \delta_{n_0,j_1,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( \delta_{n_0,j_2,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] I\{ i_1 \neq i_2 \} \\
& + n^{-1} \sum_{k=1}^K \sum_{i \in \mathcal{I}_{k}} n_0^{-2\varphi_1} n_0^{-1} \sum_{j_1 \notin \mathcal{I}_{k}} \sum_{j_2 \notin \mathcal{I}_{k}} E \left[\left( \delta_{n_0,j_1,i}^{\top} \partial_\eta m_{i}/J_0 \right) \left( \delta_{n_0,j_2,i}^{\top} \partial_\eta m_{i}/J_0 \right) \right] \\
&\overset{(2)}{=} n^{-1} \sum_{k_1,k_2=1}^K \sum_{i_1 \in \mathcal{I}_{k_1}} \sum_{i_2 \in \mathcal{I}_{k_2}} n_0^{-2\varphi_1} n_0^{-1} E \left[\left( \delta_{n_0,i_2,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( \delta_{n_0,i_1,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] I\{ k_1 \neq k_2 \} \\
& + n^{-1} \sum_{k=1}^K \sum_{i_1 \in \mathcal{I}_{k}} \sum_{i_2 \in \mathcal{I}_{k}} n_0^{-2\varphi_1} n_0^{-1} \sum_{j_1 \notin \mathcal{I}_{k}} \sum_{j_2 \notin \mathcal{I}_{k}} E \left[\left( \delta_{n_0,j_1,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( \delta_{n_0,j_2,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] I\{ i_1 \neq i_2 \} \\
& + n^{-1} \sum_{k=1}^K \sum_{i \in \mathcal{I}_{k}} n_0^{-2\varphi_1} n_0^{-1} \sum_{j \notin \mathcal{I}_{k}} E \left[\left( \delta_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) \left( \delta_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) \right] \\
&\overset{(3)}{=} n^{-1} \sum_{k_1,k_2=1}^K \sum_{i_1 \in \mathcal{I}_{k_1}} \sum_{i_2 \in \mathcal{I}_{k_2}} n_0^{-2\varphi_1} n_0^{-1} E \left[\left( \delta_{n_0,i_2,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( \delta_{n_0,i_1,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] I\{ k_1 \neq k_2 \} \\
& + n^{-1} \sum_{k=1}^K \sum_{i \in \mathcal{I}_{k}} n_0^{-2\varphi_1} n_0^{-1} \sum_{j \notin \mathcal{I}_{k}} E \left[\left( \delta_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) \left( \delta_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) \right] \\
&\overset{(4)}{=} n^{-1} \sum_{k_1 = 1}^K \sum_{i_1 \in \mathcal{I}_{k_1}} \sum_{i_2 \notin \mathcal{I}_{k_1}} n_0^{-2\varphi_1} n_0^{-1} E \left[\left( \delta_{n_0,i_2,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( \delta_{n_0,i_1,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
& + n_0^{-2\varphi_1} E \left[ \left( \delta_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) \left( \delta_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) \right] \\
&\overset{(5)}{=} n_0^{-2\varphi_1} \left( E \left[\left( \delta_{n_0,i_2,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( \delta_{n_0,i_1,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] + E \left[ \left( \delta_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) \left( \delta_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) \right] \right)\\
&\overset{(6)}{=} n_0^{-2\varphi_1} G_\delta^l + o(n^{-2\varphi_1}) ,
\end{align*}
where (1) holds because there are 3 possible situations for $i_1 \in \mathcal{I}_{k_1}$ and $i_2 \in \mathcal{I}_{k_2}$: i) $k_1 \neq k_2$, ii) $k_1 = k_2$ but $i_1 \neq i_2$, and iii) $i_1 = i_2$, (2) holds by the law of iterative expectations and since
$$ E \left[\left( \delta_{n_0,i_2,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( \delta_{n_0,i_1,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \mid X_{i_1}, W_{i_2}, W_{j_1}, W_{j_2} \right] = 0~,~\text{when } i_1 \neq j_2$$
and
$$ E \left[\left( \delta_{n_0,i_2,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( \delta_{n_0,i_1,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \mid X_{i_2}, W_{i_1}, W_{j_1}, W_{j_2} \right] = 0~,~\text{when } i_2 \neq j_1~,$$
(3) holds by the law of iterative expectations and since
$$ E \left[\left( \delta_{n_0,j_1,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( \delta_{n_0,j_2,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \mid X_{i_2}, W_{i_1}, W_{j_1}, W_{j_2} \right] = 0~,\text{when } i_1 \neq i_2~,$$
(4) holds since $\{ \left( \delta_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) : j \notin \mathcal{I}_k \} $ are i.i.d. random variables conditional on $W_i$ (here I use $i \in \mathcal{I}_k$), by noting that $ \sum_{k_2 = 1 }^K I_{k_2 \neq k_1} \sum_{i_2 \in \mathcal{I}_{k_2}} (\cdot) = \sum_{i_2 \notin \mathcal{I}_{k_1}} (\cdot)$, and recalling that $n_0$ is the number of observations outside the fold $\mathcal{I}_k$, (5) holds because the random variables $\{ \left( \delta_{n_0,i_2,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( \delta_{n_0,i_1,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) : i_1 \neq i_2 \}$ are identically distributed, and (6) holds by the definition of $G_\delta^l$ in (ref) and $n/2 \le n_0 \le n$. This completes the proof of claim 1.
\textit{Claim 2:} $I_{2} = 0$. First, consider the following derivations
\begin{align*}
& E \left[\left( (\Delta_{i_1}^l)^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( (\Delta_{i_2}^b)^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
&\overset{(1)}{=} n_0^{-\varphi_1-\varphi_2} n_0^{-3/2} \sum_{j_1 \notin \mathcal{I}_{k_1}} \sum_{j_2 \notin \mathcal{I}_{k_2}} E \left[\left( \delta_{n_0,j_1,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( b_{n_0,j_2,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
&\overset{(2)}{=} 0 ,
\end{align*}
where (1) holds by definition of $\Delta_{i}^l$ and $\Delta_{i}^b$, and (2) holds by considering 3 possible cases:
\begin{itemize}
• If $j_1 \neq i_2$ and $j_1 \neq j_2$ ($j_1$ different than all other sub-indices), then
$$ E \left[\left( \delta_{n_0,j_1,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( b_{n_0,j_2,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \mid W_{i_1}, W_{i_2}, W_{j_2}\right] = 0~,$$
since $E[ \delta_{n_0,j_1,i_1} \mid W_{i_1}, W_{i_2}, W_{j_2} ] = 0$ due to part (a) of Assumption (ref).
• If $j_1 = i_2$, then $i_2 \neq i_1$ (otherwise $j_1 \in \mathcal{I}_k$) and
$$ E \left[\left( \delta_{n_0,i_2,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( b_{n_0,j_2,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \mid W_{i_2}, X_{i_1}, W_{j_2} \right] = 0~,$$
since $E[ \delta_{n_0,i_2,i_1} \mid W_{i_1}, X_{i_2}, W_{j_2} ] = 0$ due to part (a) of Assumption (ref).
• If $j_1 = j_2 = j$, then
$$ E \left[\left( \delta_{n_0,j,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( b_{n_0,j,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \mid X_j, W_{i_1}, W_{i_2} \right] = 0~,$$
since $E[ \delta_{n_0,j,i_1} \mid X_j, W_{i_1}, W_{i_2} ] = 0$ due to part (a) of Assumption (ref).
\end{itemize}
Therefore,
\begin{align*}
I_{1,2} &= n^{-1} \sum_{i_1=1}^n \sum_{i_2=1}^n E \left[\left( (\Delta_{i_1}^l)^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( (\Delta_{i_2}^b)^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
&= 0 ,
\end{align*}
which completes the proof of claim 2.
\textit{Claim 3:} $I_{3} = O(n^{-2\varphi_1})$. Algebra shows
\begin{align*}
I_{3} &= n^{-1} \sum_{i_1=1}^n \sum_{i_2=1}^n E \left[\left( (\Delta_{i_1}^b)^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( (\Delta_{i_2}^b)^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
&= n^{-1} \sum_{k_1,k_2=1}^K \sum_{i_1 \in \mathcal{I}_{k_1}} \sum_{i_2 \in \mathcal{I}_{k_2}} n_0^{-2\varphi_2} n_0^{-2} \sum_{j_1 \notin \mathcal{I}_{k_1}} \sum_{j_2 \notin \mathcal{I}_{k_2}} E \left[\left( b_{n_0,j_1,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( b_{n_0,j_2,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
&\overset{(1)}{=} n^{-1} \sum_{k_1 = 1}^K \sum_{i_1 \in \mathcal{I}_{k_1}} \sum_{i_2 \notin \mathcal{I}_{k_1}} n_0^{-2\varphi_2} n_0^{-2} E \left[\left( b_{n_0,i_2,i_1}^{\top} \partial_\eta m_{i_1}/J_0 \right) \left( b_{n_0,i_1,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
& + n^{-1} \sum_{k=1}^K \sum_{i \in \mathcal{I}_{k}} n_0^{-2\varphi_2} n_0^{-2} \sum_{j \notin \mathcal{I}_{k}} E \left[\left( b_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) \left( b_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) \right] \\
&\overset{(2)}{=} n_0^{-1} E \left[\left( n_0^{-\varphi_2} b_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) \left( n_0^{-\varphi_2} b_{n_0,i,j}^{\top} \partial_\eta m_{j}/J_0 \right) \right] \\
& + n_0^{-1} E \left[\left( n_0^{-\varphi_2} b_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) \left( n_0^{-\varphi_2} b_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right) \right] , \\
&\overset{(3)}{\le} n_0^{-1} \left (E \left[\left( n_0^{-\varphi_2} b_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right)^2 \right] \right)^{1/2} \left (E \left[\left( n_0^{-\varphi_2} b_{n_0,i,j}^{\top} \partial_\eta m_{i}/J_0 \right)^2 \right] \right)^{1/2} \\
& + n_0^{-1} E \left[\left( n_0^{-\varphi_2} b_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right)^2 \right] \\
&\overset{(4)}{\le} 2 (p \tilde{C}_1/C_0^2) n_0^{-1} E \left[ || n_0^{-\varphi_2} b_{n_0,j,i}||^2 \right] \\
&\overset{(5)}{\le} 2 (p \tilde{C}_1/C_0^2) n_0^{-1} n_0^{1-2\varphi_1} \tau_{n_0} \\
&\overset{(6)}{=} o(n^{-2\varphi_1}) ,
\end{align*}
where (1) uses the same argument to calculate $I_{1}$, (2) holds since the random vectors $\{ \left( b_{n_0,i_2,i_1}^{\top} \partial_\eta m_{i_1}/J_0 , b_{n_0,i_1,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right) : i_1 \neq i_2 \}$ are identically distributed, (3) holds by Cauchy-Schwartz inequality, (4) holds by the inequalities (ref) and (ref) presented below where $\tilde{C}_1$ is a constant depending only on $(C_0,C_1,M)$, (5) holds by part (b.4) in Assumption (ref), and (6) holds since $n/2 \le n_0 \le n$ and $\tau_{n_0} = o(1)$.
\begin{align}
E \left[\left( n_0^{-\varphi_2} b_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right)^2 \right] &\le (p \tilde{C}_1/C_0^2) E \left[ || n_0^{-\varphi_2} b_{n_0,j,i}||^2 \right] \\
E \left[\left( n_0^{-\varphi_2} b_{n_0,i,j}^{\top} \partial_\eta m_{i}/J_0 \right)^2 \right] &\le (p \tilde{C}_1/C_0^2) E \left[ || n_0^{-\varphi_2} b_{n_0,j,i}||^2 \right]
\end{align}
To verify (ref) consider the following derivation,
\begin{align*}
E \left[\left( n_0^{-\varphi_2} b_{n_0,j,i}^{\top} \partial_\eta m_{i}/J_0 \right)^2 \right] &\overset{(1)}{=} E \left[ n_0^{-\varphi_2} b_{n_0,j,i}^{\top} E \left[ (\partial_\eta m_{i}/J_0)(\partial_\eta m_{i}/J_0)^\top \mid X_j, X_i \right] n_0^{-\varphi_2} b_{n_0,j,i} \right] \\
&\overset{(2)}{\le} (1/C_0^2) E \left[ n_0^{-\varphi_2} b_{n_0,j,i}^{\top} E \left[ (\partial_\eta m_{i})(\partial_\eta m_{i})^\top \mid X_i \right] n_0^{-\varphi_2} b_{n_0,j,i} \right] \\
&\overset{(3)}{\le} p (\tilde{C}_1/C_0^2) E \left[ ||n_0^{-\varphi_2} b_{n_0,j,i}||^2 \right]
\end{align*}
where (1) holds by LIE and since $b_{n_0,j,i}$ is non-random conditional on $X_j$ and $X_i$, (2) holds by part (a) of Assumption (ref) and independence between $X_i$ and $X_j$ since $i \neq j$, and (3) holds by definition of euclidean norm and since $|| E[(\partial_\eta m_{i})(\partial_\eta m_{i})^\top \mid X_i] ||_{\infty} \le \tilde{C}_1 = C_1(1+M^{1/4}/C_0)^2$ due to parts (d) in Assumption (ref) and $|\theta_0| \le M^{1/4}/C_0$ (which holds by definition of $\theta_0$ and parts (a) and (c) of Assumption (ref)).
The verification of (ref) follows the same previous derivations but reverting the role of $i$ and $j$. Lastly, it uses that $E \left[ || n_0^{-\varphi_2} b_{n_0,i,j}||^2 \right] = E \left[ || n_0^{-\varphi_2} b_{n_0,j,i}||^2 \right]$ since $b_{n_0,j,i}$ and $b_{n_0,i,j}$ have the same distribution for $i \neq j$.
\textbf{Part 3:} By Cauchy-Schwartz, parts 3 of Proposition (ref), and part 2 of this proposition,
\begin{equation*}
\text{Cov}(\mathcal{T}_n^{dml}, \mathcal{T}_{n,K}^{l}) \le \left( (G_\delta^l)^{1/2} (K/(K-1))^{\varphi_1} + o(1) \right)^{1/2} \left( \sigma^2 \sigma^2_a + \Lambda^2 + o(1) \right)^{1/2}n^{-\varphi_1 -1/2} ,
\end{equation*}
which implies the RHS is $O(n^{-\varphi_1 -1/2})$, and this is $o(n^{-2\varphi_1})$ since $\varphi_1 < 1/2$.
\end{proof}
\subsection{Proof of Proposition (ref)}
\begin{proof}
For $i \in \mathcal{I}_k$, denote $\Delta_{i} = \Delta_{i}^b + \Delta_{i}^l$, where
\begin{align*}
\Delta_{i}^l &= n_0^{-\varphi_1} n_0^{-1/2} \sum_{j \notin \mathcal{I}_k} \delta_{n_0,j,i} ,\\
\Delta_{i}^b &= n_0^{-\varphi_2} n_0^{-1} \sum_{j \notin \mathcal{I}_k} b_{n_0,j,i} ,
\end{align*}
Here, $\delta_{n_0,j,i} = \delta_{n_0}(W_{j},X_i)$ and $b_{n_0,j,i} = b_{n_0}(X_{j},X_i)$, and $\delta_{n_0}$ and $b_{n_0}$ are functions satisfying Assumption (ref). Denote $H_i = \partial_\eta^2 m_i/(2J_0)$.
\textbf{Part 1:} Consider the following decomposition using the definition of $\mathcal{T}_{n,K}^{nl}$ in (ref),
\begin{align*}
E[\mathcal{T}_{n,K}^{nl}] &= n^{-1/2 }\sum_{i = 1}^n E\left[ \Delta_{i}^\top H_i \Delta_{i} \right] \\
&= I_1 + 2I_2 + I_3
\end{align*}
where
\begin{align*}
I_1 &= n^{-1/2 }\sum_{i = 1}^n E\left[ (\Delta_{i}^l)^\top H_i \Delta_{i}^l \right] \\
I_2 &= n^{-1/2 }\sum_{i = 1}^n E\left[ (\Delta_{i}^b)^\top H_i \Delta_{i}^l \right] \\
I_3 &= n^{-1/2 }\sum_{i = 1}^n E\left[ (\Delta_{i}^b)^\top H_i \Delta_{i}^b \right]
\end{align*}
In what follows, I show $I_1 = n^{1/2} n_0^{-2\varphi_1}F_\delta + o(n^{1/2-2\varphi_1})$, $I_2 = 0$, $I_3 = n^{1/2} n_0^{-2\varphi_2}F_b + o(n^{1/2-2\varphi_1})$, which is sufficient to complete the proof of Part 1 since $n_0 = ((K-1)/K) n$.\\
\textit{Claim 1:} $I_1 = n^{1/2} n_0^{-2\varphi_1}F_\delta + o(n^{1/2-2\varphi_1}) $. Consider the following derivations,
\begin{align*}
E[ (\Delta_{i}^l)^\top H_i \Delta_{i}^l ]
&\overset{(1)}{=} n_0^{-2\varphi_1} n_0^{-1} \sum_{j_1 \notin \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_k} E[( \delta_{n_0,j_1,i})^\top H_i ( \delta_{n_0,j_2,i} ) ] \\
&\overset{(2)}{=} n_0^{-2\varphi_1} n_0^{-1} \sum_{j\notin \mathcal{I}_k} E[( \delta_{n_0,j,i})^\top H_i ( \delta_{n_0,j,i} ) ] \\
&\overset{(3)}{=} n_0^{-2\varphi_1} E[( \delta_{n_0,j,i})^\top H_i ( \delta_{n_0,j,i} ) ] ,
\end{align*}
where (1) holds by definition of $\Delta_{i}^l$, and (2) and (3) hold since $\{ \delta_{n_0,j,i}: j \notin \mathcal{I}_k \} $ are zero mean i.i.d. random vectors conditional on $W_i$ due to part (a) of Assumption (ref) (here I use that $i \in \mathcal{I}_k$). Therefore,
\begin{align*}
I_1 &= n^{-1/2 }\sum_{i = 1}^n E\left[ (\Delta_{i}^l)^\top H_i ( \Delta_{i}^l ) \right] \\
&= n^{1/2} n_0^{-2\varphi_1} E[( \delta_{n_0,j,i})^\top H_i ( \delta_{n_0,j,i} ) ] \\
&\overset{(1)}{=} n^{1/2} n_0^{-2\varphi_1} F_\delta + o( n^{1/2-2\varphi_1})
\end{align*}
where (1) holds by definition of $F_\delta $ in (ref), Assumption (ref), and because $n/2 \le n_0 \le n$. This completes the proof of claim 1.
\textit{Claim 2:} $I_2 = 0$. Consider the following derivations,
\begin{align*}
E[( \Delta_{i}^b)^\top H_i ( \Delta_{i}^l) ]
&\overset{(1)}{=} n_0^{-\varphi_2-\varphi_1} n_0^{-3/2} \sum_{j_1 \notin \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_k} E[( b_{n_0,j_1,i} )^\top H_i ( \delta_{n_0,j_2,i} ) ] \\
&\overset{(2)}{=} 0
\end{align*}
where (1) holds by definition of $\Delta_{i}^b$ and $\Delta_{i}^l$, and (2) holds since
$$E[( b_{n_0,j_1,i} )^\top (\partial_\eta^2 m_i/(2 J_0)) ( \delta_{n_0,j_2,i} ) \mid X_{j_2}, X_{j_1}, W_i] = 0 $$
due to part (a) of Assumption (ref) ( $E[\delta_{n_0,j_2,i} \mid X_{j_2}, W_i ] = 0$). Therefore,
\begin{align*}
I_2 &= n^{-1/2 }\sum_{i = 1}^n E\left[ (\Delta_{i}^b)^\top (\partial_\eta^2 m_i/(2 J_0)) ( \Delta_{i}^l ) \right] \\
&= 0 ,
\end{align*}
which completes the proof of claim 2.
\textit{Claim 3:} $I_3 = n^{1/2} n_0^{-2\varphi_2} F_b + o(n^{1/2-2\varphi_1})$. Denote $\tilde{b}_{n_0,i} = E[b_{n_0,j,i} \mid X_i]$ for $j \neq i$. Consider the following derivations,
\begin{align*}
E[( \Delta_{i}^b)^\top H_i ( \Delta_{i}^l) ]
&\overset{(1)}{=} n_0^{-2\varphi_2} n_0^{-2} \sum_{j_1 \notin \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_k} E[( b_{n_0,j_1,i} )^\top H_i ( b_{n_0,j_2,i} ) ] \\
&\overset{(2)}{=} n_0^{-2\varphi_2} n_0^{-2} \sum_{j_1 \notin \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_k} E[( b_{n_0,j_1,i} - \tilde{b}_{n_0,i} )^\top H_i ( b_{n_0,j_2,i} - \tilde{b}_{n_0,i} ) ] \\
& + n_0^{-2\varphi_2} n_0^{-2} \sum_{j_1 \notin \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_k} E[( b_{n_0,j_1,i} - \tilde{b}_{n_0,i} )^\top H_i ( \tilde{b}_{n_0,i} ) ] \\
& + n_0^{-2\varphi_2} n_0^{-2} \sum_{j_1 \notin \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_k} E[( \tilde{b}_{n_0,i} )^\top H_i ( b_{n_0,j_2,i} - \tilde{b}_{n_0,i} ) ] \\
& + n_0^{-2\varphi_2} n_0^{-2} \sum_{j_1 \notin \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_k} E[( \tilde{b}_{n_0,i} )^\top H_i ( \tilde{b}_{n_0,i} ) ] \\
&\overset{(3)}{=} n_0^{-2\varphi_2} n_0^{-1} E[( b_{n_0,j,i} - \tilde{b}_{n_0,i} )^\top H_i ( b_{n_0,j,i} - \tilde{b}_{n_0,i} ) ] \\
& + n_0^{-2\varphi_2} E[( \tilde{b}_{n_0,i} )^\top H_i ( \tilde{b}_{n_0,i} ) ]
\end{align*}
where (1) holds by definition of $\Delta_{i}^b$, (2) holds by adding and subtracting $\tilde{b}_{n_0,i}$, and (3) holds since $\{ b_{n_0,j,i} - \tilde{b}_{n_0,i} : j \notin \mathcal{I}_k \}$ are zero mean i.i.d. random vectors conditional on $W_i$, which implies $E[( \tilde{b}_{n_0,i} )^\top H_i ( b_{n_0,j,i} - \tilde{b}_{n_0,i} ) \mid W_i ] = 0$. Therefore,
\begin{align*}
I_3 &= n^{-1/2 }\sum_{i = 1}^n E\left[ (\Delta_{i}^b)^\top H_i ( \Delta_{i}^b ) \right] \\
&= n^{1/2 } n_0^{-1} n_0^{-2\varphi_2}E[( b_{n_0,j,i} - \tilde{b}_{n_0,i} )^\top H_i ( b_{n_0,j,i} - \tilde{b}_{n_0,i} ) ]
+ n^{1/2 } n_0^{-2\varphi_2} E[( \tilde{b}_{n_0,i} )^\top H_i ( \tilde{b}_{n_0,i} ) ]\\
&\overset{(1)}{=} o(n^{1/2-2\varphi_1}) + n^{1/2 } n_0^{-2\varphi_2} E[( \tilde{b}_{n_0,i} )^\top H_i ( \tilde{b}_{n_0,i} ) ] \\
&\overset{(2)}{=} o(n^{1/2-2\varphi_1}) + n^{1/2} n_0^{-2\varphi_2} F_b + o(n^{1/2-2\varphi_2}) \\
&\overset{(3)}{=} n^{1/2} n_0^{-2\varphi_2} F_b + o(n^{1/2-2\varphi_1})
\end{align*}
where (1) holds by (ref) presented below and since $n/2 \le n_0 \le n$, (2) holds by definition of $F_b$ in (ref) and Assumption (ref), and (3) since $\varphi_1 \le \varphi_2$ and $n/2 \le n_0 \le n$.
\begin{equation}
n_0^{-2\varphi_2} E[( b_{n_0,j,i} - \tilde{b}_{n_0,i} )^\top H_i ( b_{n_0,j,i} - \tilde{b}_{n_0,i} ) ] = o(n^{1 - 2\varphi_1}) .
\end{equation}
To verify (ref) consider the following derivations,
\begin{align*}
| n_0^{-2\varphi_2} E[( b_{n_0,j,i} - \tilde{b}_{n_0,i} )^\top H_i ( b_{n_0,j,i} - \tilde{b}_{n_0,i} ) ] |
&\overset{(1)}{\le} n_0^{-2\varphi_2} (2C_0)^{-1} E[ |( b_{n_0,j,i} - \tilde{b}_{n_0,i} )^\top \partial_\eta^2 m_i ( b_{n_0,j,i} - \tilde{b}_{n_0,i} ) |] \\
&\overset{(2)}{\le} (2C_0)^{-1} (p \tilde{C}_2) E[ || n_0^{-\varphi_2} b_{n_0,j,i} - n_0^{-\varphi_2} \tilde{b}_{n_0,i} ||^2 ] \\
&\overset{(3)}{\le} 2 (2C_0)^{-1} (p \tilde{C}_2) \left( E[ || n_0^{-\varphi_2} b_{n_0,j,i}||^2] + E[ ||n_0^{-\varphi_2} \tilde{b}_{n_0,i} ||^2 ] \right) \\
&\overset{(4)}{\le} (p \tilde{C}_2/C_0) \left( n_0^{1-2\varphi_1} \tau_{n_0} + n_0^{-2\varphi_2} M_1^{1/2} \right) \\
&\overset{(5)}{=} o(n^{1 - 2\varphi_1}) ,
\end{align*}
where (1) holds by triangular inequality and part (a) of Assumption (ref), (2) holds by definition of euclidean norm and since $||E[\partial_\eta^2 m_i \mid X_i]||_{\infty} \le \tilde{C}_2 = C_2(1+ M^{1/4}/C_0)$ due to part (e) of Assumption (ref) and $|\theta_0| \le M^{1/4}/C_0$ (which holds by definition of $\theta_0$ and parts (a) and (c) of Assumption (ref)), (3) holds by standard properties of euclidean norm, (4) holds by parts (b.3) and (b.4) of Assumption (ref) with $\tau_{n_0} = o(1)$, and (5) holds since $\varphi_1 \le \varphi_2$ and $n/2 \le n_0 \le n$. \\
\textbf{Part 2:} Consider the following decomposition,
\begin{equation*}
\mathcal{T}_{n,K}^{nl} - E[\mathcal{T}_{n,K}^{nl}]= I_{l,l} + 2 I_{l,b} + I_{b,b}
\end{equation*}
where
\begin{align*}
I_{l,l} &= n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_1-1} \sum_{j_1 \notin \mathcal{I}_k } \sum_{j_2 \notin \mathcal{I}_k } \left( \delta_{n_0,j_1,i}^\top H_i \delta_{n_0,j_2,i} - E[\delta_{n_0,j_1,i}^\top H_i \delta_{n_0,j_2,i}] \right) \\
I_{l,b} &= n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1-\varphi_2-3/2} \sum_{j_1 \notin \mathcal{I}_k } \sum_{j_2 \notin \mathcal{I}_k } \left( \delta_{n_0,j_1,i}^\top H_i b_{n_0,j_2,i} - E[\delta_{n_0,j_1,i}^\top H_i b_{n_0,j_2,i}] \right) \\
I_{b,b} &= n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_2-2}
\sum_{j_1 \notin \mathcal{I}_k } \sum_{j_2 \notin \mathcal{I}_k } \left( b_{n_0,j_1,i}^\top H_i b_{n_0,j_2,i} - E[b_{n_0,j_1,i}^\top H_i b_{n_0,j_2,i}] \right)
\end{align*}
which implies
\begin{align*}
Var[\mathcal{T}_{n,K}^{nl}] &= E \left[ \left( I_{l,l} + 2 I_{l,b} + I_{b,b} \right)^2 \right] \\
&= E[I_{l,l}^2] + E[I_{b,b}^2] + 4E[I_{l,b}^2] + 2E[I_{l,l} I_{b,b}] + 4E[(I_{l,l}+I_{b,b}) I_{l,b}]
\end{align*}
In what follows, I show $E[I_{l,l}^2] = G_\delta (K^2-3K+3)(K-1)^{-2} n_0^{1-4\varphi_1} + o(n^{-\zeta}) $, $E[I_{b,b}^2] = o(n^{-\zeta})$, and $E[I_{l,b}^2] = o(n^{-\zeta})$, which is sufficient to complete the proof of Part 2 since $n_0 = ((K-1)/K) n$ and by Cauchy-Schwartz it holds $E[I_{l,l} I_{b,b}] = o(n^{-\zeta})$, and $E[(I_{l,l}+I_{b,b}) I_{l,b}] = o(n^{-\zeta})$.\\
\textit{Claim 1:} $E[I_{l,l}^2] = G_\delta (K^2-3K+3)(K-1)^{-2} n_0^{1-4\varphi_1} + o(n^{-\zeta})$. Consider the following notation
$$ \Gamma_{j_1,j_2,i}^{l,l} = \left( \delta_{n_0,j_1,i}^\top H_i \delta_{n_0,j_2,i} - E[\delta_{n_0,j_1,i}^\top H_i \delta_{n_0,j_2,i}] \right)~.$$
Note that $E[\Gamma_{j_1,j_2,i}^{l,l}] = 0$ by construction, and $j_1 \neq j_2$ implies
$$ E[\delta_{n_0,j_1,i}^\top H_i \delta_{n_0,j_2,i}] = 0 \quad \text{and} \quad \Gamma_{j_1,j_2,i}^{l,l} = \delta_{n_0,j_1,i}^\top H_i \delta_{n_0,j_2,i} ~.$$
Therefore, $E[\Gamma_{j_1,j_2,i}^{l,l} \mid W_i,W_{j_1},X_{j_2}] = 0 $ and $E[\Gamma_{j_1,j_2,i}^{l,l} \mid W_i,W_{j_2},X_{j_1}] = 0$ when $j_1 \neq j_2$ due to part (a) of Assumption (ref). Furthermore,
\begin{equation}
\left| E\left[ \left( \Gamma_{j_1,j_2,i_1}^{l,l} \right) \left( \Gamma_{j_3,j_4,i_2}^{l,l} \right) \right] \right| \le (p\tilde{C}_2/C_0)^2 n_0^{1-2\varphi_1} M_1 ,
\end{equation}
which follows by Cauchy-Schwartz, part (e) of Assumption (ref), and part (b.2) of Assumption (ref), with $\tilde{C}_2 = C_2(1+ M^{1/4}/C_0)$.
Using the previous notation, $I_{l,l}$ can be written as follows
\begin{equation}
I_{l,l}= n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_1-1} \sum_{j_1 \notin \mathcal{I}_k } \sum_{j_2 \notin \mathcal{I}_k } \Gamma_{j_1,j_2,i}^{l,l}
\end{equation}
and $E[I_{l,l}^2]$ can be decompose in three terms
\begin{align*}
E[I_{l,l}^2] &= E\left[ \left( n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_1-1} \sum_{j_1 \notin \mathcal{I}_k } \sum_{j_2 \notin \mathcal{I}_k } \Gamma_{j_1,j_2,i}^{l,l} \right)^2 \right] \\
&= E\left[ \left( n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_1-1} \sum_{j \notin \mathcal{I}_k } \Gamma_{j,j,i}^{l,l}
+ n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_1-1} \sum_{j_1 \notin \mathcal{I}_k } \sum_{j_2 \notin \mathcal{I}_k } \Gamma_{j_1,j_2,i}^{l,l} I\{j_1 \neq j_2\} \right)^2 \right] \\
&= I_1 + I_2 + 2I_3 ,
\end{align*}
where
\begin{align*}
I_1 &= E\left[ \left( n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_1-1} \sum_{j \notin \mathcal{I}_k } \Gamma_{j,j,i}^{l,l} \right)^2 \right] \\
I_2 &= E\left[ \left( n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_1-1} \sum_{j_1 \notin \mathcal{I}_k } \sum_{j_2 \notin \mathcal{I}_k } \Gamma_{j_1,j_2,i}^{l,l} I\{j_1 \neq j_2\} \right)^2 \right] \\
I_3 &= E\left[ \left( n^{-1/2} \sum_{k_1=1}^K \sum_{i_1 \in \mathcal{I}_{k_1}} n_0^{-2\varphi_1-1} \sum_{j_1 \notin \mathcal{I}_{k_1} } \Gamma_{j_1,j_2,i_1}^{l,l} \right) \left( n^{-1/2} \sum_{k_2=1}^K \sum_{i_2 \in \mathcal{I}_{k_2}} n_0^{-2\varphi_1-1} \sum_{j_3 \notin \mathcal{I}_{k_2} } \sum_{j_4 \notin \mathcal{I}_{k_2} } \Gamma_{j_3,j_4,i_2}^{l,l} I\{j_3 \neq j_4\} \right) \right]
\end{align*}
In what follows, I show that $I_1 = o(n^{-\zeta})$, $I_2 = G_\delta (K^2-3K+3)(K-1)^{-2} n_0^{1-4\varphi_1} + o(n^{-\zeta})$, and $I_3 = o(n^{-\zeta})$, which is sufficient to complete the proof of Claim 1.\\
\textit{Claim 1.1} $I_1 = o(n^{-\zeta})$. Consider the following expansion
\begin{align*}
I_1 &= n^{-1} n_0^{-4\varphi_1-2} \sum_{k_1=1}^K \sum_{i_1 \in \mathcal{I}_{k_1}} \sum_{j_1 \notin \mathcal{I}_{k_1} } \sum_{k_2=1}^K \sum_{i_2 \in \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_{k_2} } E\left[ \left( \Gamma_{j_1,j_1,i_1}^{l,l} \right) \left( \Gamma_{j_2,j_2,i_2}^{l,l} \right) \right] \\
&= n^{-1} n_0^{-4\varphi_1-2} \sum_{(i_1,i_2,j_1,i_2) \in \mathcal{E}} E\left[ \left( \Gamma_{j_1,j_1,i_1}^{l,l} \right) \left( \Gamma_{j_2,j_2,i_2}^{l,l} \right) \right] ,
\end{align*}
where $\mathcal{E} = \{ (i_1,i_2,j_1,j_2) \in [n]^4 : i_1 \in \mathcal{I}_{k_1}, i_2 \in \mathcal{I}_{k_2}, j_1 \notin \mathcal{I}_{k_1}, j_2 \notin \mathcal{I}_{k_2}, k_1 \in [K], k_2 \in [K] \}$, with $[n]$ denoting $\{1,\ldots, n\}$. Let $\mathcal{E}_4 \subseteq [n]^4$ be the subset of indices with distinct entries. (e.g., $i_1 \notin \{i_2,j_1,j_2\}$, $i_2 \notin \{j_1,j_2\}$, $j_1 \neq j_2$). Let $\mathcal{E}_{\le 3} \subset [n]^4$ be the subset of indices with at most three distinct entries.
Now, take $(i_1,i_2,j_1,j_2) \in \mathcal{E} \cap \mathcal{E}_4$. It follows that
$$ E\left[ \left( \Gamma_{j_1,j_1,i_1}^{l,l} \right) \left( \Gamma_{j_2,j_2,i_2}^{l,l} \right) \right] = E\left[ \left( \Gamma_{j_1,j_1,i_1}^{l,l} \right) \right] E\left[ \left( \Gamma_{j_2,j_2,i_2}^{l,l} \right) \right] = 0~,$$
due to independence and by definition of $\Gamma_{j_1,j_1,i_1}^{l,l}$.
Therefore,
\begin{align*}
|I_1| &= \left| n^{-1} n_0^{-4\varphi_1-2} \sum_{(i_1,i_2,j_1,i_2) \in \mathcal{E}\setminus \mathcal{E}_4 } E\left[ \left( \Gamma_{j_1,j_1,i_1}^{l,l} \right) \left( \Gamma_{j_2,j_2,i_2}^{l,l} \right) \right] \right| \\
&\overset{(1)}{\le} n^{-1} n_0^{-4\varphi_1-2} \sum_{(i_1,i_2,j_1,i_2) \in \mathcal{E}_{\le 3} } \left| E\left[ \left( \Gamma_{j_1,j_1,i_1}^{l,l} \right) \left( \Gamma_{j_2,j_2,i_2}^{l,l} \right) \right] \right|\\
&\overset{(2)}{\le} n^{-1} n_0^{-4\varphi_1-2} \sum_{(i_1,i_2,j_1,i_2) \in \mathcal{E}_{\le 3} } (p\tilde{C}_2/C_0)^2 n_0^{1-2\varphi_1} M_1 \\
&\overset{(3)}{\le} n^{-1} n_0^{-4\varphi_1-2} \times 3^4 n^3 \times (p\tilde{C}_2/C_0)^2 n_0^{1-2\varphi_1} M_1 \\
&\overset{(4)}{=} O(n^{1-6\varphi_1}) ,
\end{align*}
where (1) holds by triangular inequality and since $\mathcal{E}\setminus \mathcal{E}_4 \subset \mathcal{E}_{\le 3}$, (2) holds by (ref), (3) holds since the number of elements of $\mathcal{E}_3 $ is at most $3^4n^3$ (for each 3-tuple $(a,b,c) \in [n]^3$ consider the functions from the positions $\{1,2,3,4\}$ into the possible values $\{a,b,c\}$, the number of all these functions is $3^4$ and there number of 3-tuple is $n^3$), and (4) hold since $n/2 \le n \le n$. Therefore, $I_1$ is $O(n^{1-6\varphi_1})$, which is $o(n^{-\zeta})$ since $6\varphi_1 -1 > 4\varphi_1 - 1 \ge \zeta$. This completes the proof of Claim 1.1.\\
\textit{Claim 1.2:} $I_2 = G_\delta (K^2-3K+3)(K-1)^{-2} n_0^{1-4\varphi_1} + o(n^{-\zeta})$. Consider the following expansion
\begin{align*}
I_2 &= n^{-1} n_0^{-4\varphi_1-2} \sum_{k_1=1}^K \sum_{i_1 \in \mathcal{I}_{k_1}}\sum_{j_1,j_2 \notin \mathcal{I}_{k_1} } \sum_{k_2=1}^K \sum_{i_2 \in \mathcal{I}_{k_2}} \sum_{j_3, j_4 \notin \mathcal{I}_{k_2} } E\left[ \left( \Gamma_{j_1,j_2,i_1}^{l,l} I\{j_1 \neq j_2\} \right)
\left( \Gamma_{j_3,j_4,i_2}^{l,l} I\{j_3 \neq j_4\} \right) \right] \\
&= n^{-1} n_0^{-4\varphi_1-2} \sum_{ (i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} } E\left[ \left( \Gamma_{j_1,j_2,i_1}^{l,l} \right)
\left( \Gamma_{j_3,j_4,i_2}^{l,l} \right) \right] ,
\end{align*}
where $\mathcal{E} = \{ (i_1,i_2,j_1,j_2,j_3,j_4) \in [n]^6 : i_1 \in \mathcal{I}_{k_1}, i_2 \in \mathcal{I}_{k_2}; j_1,j_2 \notin \mathcal{I}_{k_1}; j_3, j_4 \notin \mathcal{I}_{k_1}; j_1 \neq j_2; j_3 \neq j_4; k_1, k_2 \in [K] \}$. Let $\mathcal{E}_6 \subseteq [n]^6$ be the subset of indices with distinct entries. Let $\mathcal{E}_5 \subseteq [n]^6 $ be the subset of indices with exactly five distinct entries, meaning that two entries are identical while the remaining entries are distinct. Let $\mathcal{E}_4 \subseteq [n]^6 $ be the subset of indices with exactly four distinct entries. Let $\mathcal{E}_{\le 3} \subset [n]^6$ be the subset of indices with at most three distinct entries. Note that $[n]^6 = \mathcal{E}_{\le 3} \cup \mathcal{E}_{4} \cup \mathcal{E}_{5} \cup \mathcal{E}_{6}$.
Now, take $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_6 $. It follows that
\begin{equation}
E\left[ \left( \Gamma_{j_1,j_2,i_1}^{l,l} \right)
\left( \Gamma_{j_3,j_4,i_2}^{l,l} \right) \right] = 0 ,
\end{equation}
since $\Gamma_{j_1,j_2,i_1}^{l,l}$ and $\Gamma_{j_3,j_4,i_2}^{l,l}$ are independent zero mean random variables.
Now take $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_5$. Without loss of generality, assume that $j_1$ is different than all the other indices (otherwise, this statement holds with $j_2$ or $j_3$ or $j_4$). Then,
\begin{align}
E\left[ \left( \Gamma_{j_1,j_2,i_1}^{l,l} \right)
\left( \Gamma_{j_3,j_4,i_2}^{l,l} \right) \right] &\overset{(1)}{=} E\left[ \left( \delta_{n_0,j_1,i_1}^\top H_{i_1} \delta_{n_0,j_2,i_1} \right)
\left( \delta_{n_0,j_3,i_2}^\top H_{i_2} \delta_{n_0,j_4,i_2} \right) \right] \notag \\
&\overset{(2)}{=} E\left[ E\left[ \delta_{n_0,j_1,i_1}^\top \mid W_{i_1}, W_{i_2}, W_{j_2}, W_{j_3}, W_{j_4} \right]\left( H_{i_1} \delta_{n_0,j_2,i_1} \right)
\left( \delta_{n_0,j_3,i_2}^\top H_{i_2} \delta_{n_0,j_4,i_2} \right) \right] \notag \\
&\overset{(3)}{=} 0
\end{align}
where (1) holds by definition of $\Gamma_{j_1,j_2,i_1}^{l,l}$ and $ \Gamma_{j_3,j_4,i_2}^{l,l}$ since $j_1 \neq j_2$ and $j_3 \neq j_4$, (2) holds by LIE, and (3) holds by part (a) of Assumption (ref). Note that this argument can be used whenever one $j_s$ is different than all the other indices, for some $s \in \{1,2,3,4\}$.
Now take $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_4$. Suppose $\{a,b,c,d\}$ are four different indices, then there are two possible distributions for the 6-tuples: (i) two pairs, e.g., $(a,a,b,b,c,d)$, or (ii) one triple, e.g., $(a,a,a,b,c,d)$. Notice that for 6-tuples in (ii), there exists one $j_s$ different than all the other indices, for some $s \in \{1,2,3,4\}$. In this case, $E\left[ \left( \Gamma_{j_1,j_2,i_1}^{l,l} \right) \left( \Gamma_{j_3,j_4,i_2}^{l,l} \right) \right] $ equals zero due to the argument described above.
Therefore, in what follows, I consider only 6-tuples in (i), specifically, the cases where $j_s$ appears in a pair for all $s=1,2,3,4$.
\begin{itemize}
• Case 1: $j_1 = j_3 $, $j_2 = j_4$, and $i_1 \neq i_2$. Then,
\begin{equation}
E\left[ \left( \Gamma_{j_1,j_2,i_1}^{l,l} \right)
\left( \Gamma_{j_3,j_4,i_2}^{l,l} \right) \right] = E\left[ \left( \delta_{n_0,j_1,i_1}^\top H_{i_1} \delta_{n_0,j_2,i_1} \right)
\left( \delta_{n_0,j_1,i_2}^\top H_{i_2} \delta_{n_0,j_2,i_2} \right) \right]
\end{equation}
To compute the number of indices $(i_1,i_2,j_1,j_2,j_1,j_2) \in \mathcal{E}$ in this case, recall that $i_1 \in \mathcal{I}_{k_1}$ and $i_2 \in \mathcal{I}_{k_2}$, therefore, there are two situations (i) $k_1 = k_2$ or (ii) $k_1 \neq k_2$. For the first situation, $i_1$ can take $n$ values, $i_2$ can take $n_k-1$ values (since it is different than $i_1$ but is in the same fold $\mathcal{I}_k$), and $j_1$ and $j_2$ can take $n_0$ and $n_0-1$ values (since they are different but not in $\mathcal{I}_k$). That is $ n (n_k-1) n_0 (n_0-1)$ combinations.
For the second situation, $i_1$ can take $n$ values, $i_2$ can take $n_0$ values, then $j_1$ and $j_2$ take values in all the data except into the two folds that contain $i_1$ and $i_2$ (since $j_1,j_2 \notin \mathcal{I}_{k_1}$ and $j_1=j_3,j_2=j_4 \notin \mathcal{I}_{k_2}$), that is $(n_0-n_k)(n_0-n_k-1)$. That is $n(n-n_k) (n_0-n_k)(n_0-n_k-1)$ combinations. Therefore, the total number of indices is equal to
\begin{equation}
n n_0^3 \left( \frac{K^2-3K+3}{(K-1)^2} -2n_0^{-1} + n_0^{-2} \right)
\end{equation}
• Case 2: $j_1 = j_4$, $j_2 = j_3$, and $i_1 \neq i_2$. Then,
\begin{align*}
E\left[ \left( \Gamma_{j_1,j_2,i_1}^{l,l} \right)
\left( \Gamma_{j_3,j_4,i_2}^{l,l} \right) \right] = E\left[ \left( \delta_{n_0,j_1,i_1}^\top H_{i_1} \delta_{n_0,j_2,i_1} \right)
\left( \delta_{n_0,j_1,i_2}^\top H_{i_2} \delta_{n_0,j_2,i_2} \right) \right] .
\end{align*}
The number of indices $(i_1,i_2,j_1,j_2,j_1,j_2) \in \mathcal{E}$ in this case is exactly the same as in the previous case, which is presented in (ref).
\end{itemize}
Finally, note that
\begin{equation}
\left | \sum_{ (i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cup \mathcal{E}_{\le 3} } E\left[ \left( \Gamma_{j_1,j_2,i_1}^{l,l} \right)
\left( \Gamma_{j_3,j_4,i_2}^{l,l} \right) \right] \right| \le 3^6 n^3 (p\tilde{C}_2/C_0) n_0^{1-2\varphi_1} M_1 ,
\end{equation}
which follows by triangular inequality, (ref), and by using that the number of elements in $\mathcal{E}_{\le 3}$ is lower or equal to $3^6 n^3$ (for each 3-tuple $(a,b,c) \in [n]^3$, consider the functions from the positions $\{1,2,3,4,5,6\}$ into the possible values $\{a,b,c\}$, the number of all these functions is $3^4$, while the number of 3-tuple is $n^3$).
In what follows, I use the preliminary findings to calculate $I_2$ up to an error of size $o(n^{-\zeta})$,
\begin{align*}
I_2 &= n^{-1} n_0^{-4\varphi_1-2} \sum_{ (i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} } E\left[ \left( \Gamma_{j_1,j_2,i_1}^{l,l} \right)
\left( \Gamma_{j_3,j_4,i_2}^{l,l} \right) \right] \\
&\overset{(1)}{=} n^{-1} n_0^{-4\varphi_1-2} \sum_{ (i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_4 } E\left[ \left( \Gamma_{j_1,j_2,i_1}^{l,l} \right)
\left( \Gamma_{j_3,j_4,i_2}^{l,l} \right) \right] \\
& + n^{-1} n_0^{-4\varphi_1-2} \sum_{ (i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cup \mathcal{E}_{\le 3} } E\left[ \left( \Gamma_{j_1,j_2,i_1}^{l,l} \right)
\left( \Gamma_{j_3,j_4,i_2}^{l,l} \right) \right] \\
&\overset{(2)}{=} n^{-1} n_0^{-4\varphi_1-2} \sum_{k_1=1}^K \sum_{ (i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_4 } E\left[ \left( \Gamma_{j_1,j_2,i_1}^{l,l} \right)
\left( \Gamma_{j_3,j_4,i_2}^{l,l} \right) \right] + O(n^{1-6\varphi_1}) \\
&\overset{(3)}{=} n_0^{1-4\varphi_1} \left( \frac{K^2-3K+3}{(K-1)^2} -2n_0^{-1} + n_0^{-2} \right) 2 E\left[ \left( \delta_{n_0,j_1,i_1}^\top H_{i_1} \delta_{n_0,j_2,i_1} \right)
\left( \delta_{n_0,j_1,i_2}^\top H_{i_2} \delta_{n_0,j_2,i_2} \right) \right] + O(n^{1-6\varphi_1}) \\
&\overset{(4)}{=} n_0^{1-4\varphi_1} \left( \frac{K^2-3K+3}{(K-1)^2} \right) G_\delta + o(n^{-\zeta}) ,
\end{align*}
where (1) holds by the derivations in (ref) and (ref), (2) holds by (ref), (3) holds by (ref) that computes the expected value and (ref) that calculates the number of indices to consider, and (4) holds by definition of $G_\delta$ in (ref), Assumption (ref), $n/2 \le n_0 \le n$, and since $6\varphi_1-1 > \zeta$. This completes the proof of Claim 1.2. \\
\textit{Claim 1.3:} $I_3 = o(n^{-\zeta})$. First, claim 1.1 implies $I_1$ is $o(n^{-\zeta})$. Second, claim 1.2 implies $I_2$ is $O(n^{-\zeta})$ since $4\varphi_1 - 1 \ge \zeta$. Finally, then $I_3$ is $o(n^{-\zeta})$ due to Cauchy-Schwartz ( $|I_3| \le |I_1|^{1/2} |I_2|^{1/2} $). This completes the proof of Claim 1.3.\\
\textit{Claim 2:} $E[I_{b,b}^2] = o(n^{-\zeta})$. Consider the following notation,
$$ \Gamma_{j_1,j_2,i}^{b,b} = b_{n_0,j_1,i}^\top H_i b_{n_0,j_2,i} - E[ b_{n_0,j_1,i}^\top H_i b_{n_0,j_2,i} ] ~,$$
where by construction $E[\Gamma_{j_1,j_2,i}^{b,b}] = 0$. Denote $\tilde{b}_{n_0,i} = E[ b_{n_0,j,i} \mid X_i]$. Note that if $j_1 \neq j_2$, then
$$ \Gamma_{j_1,j_2,i}^{b,b} = b_{n_0,j_1,i}^\top H_i b_{n_0,j_2,i} - E[\tilde{b}_{n_0,i}^\top H_i \tilde{b}_{n_0,i}] ~,$$
Furthermore,
\begin{equation}
n_0^{-4\varphi_2} E\left[ |\Gamma_{j_1,j_1,i}^{b,b}|^2 \right] \le (p \Tilde{C}_2/C_0) n_0^{3(1-2\varphi_1)} \tau_{n_0} ,
\end{equation}
which follows by C-S, part (e) of Assumption (ref), and part (b.4) of Assumption (ref), with $\tilde{C}_2 = C_2(1+M^{1/4}/C_0)$. And, if $j_1 \neq j_2$,
\begin{equation}
n_0^{-4\varphi_2} E\left[ |\Gamma_{j_1,j_2,i}^{b,b}|^2 \right] \le (p \Tilde{C}_2/C_0) n_0^{2(1-2\varphi_1)} {\tau}_{n_0} ,
\end{equation}
which holds by C-S, part (e) of Assumption (ref), and part (b.1) of Assumption (ref).
The previous notation can be used to rewrite $I_{b,b}$ as follows
\begin{align}
E[I_{b,b}^2] &= E\left[ \left( n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_2-2} \sum_{j_1 \notin \mathcal{I}_k } \sum_{j_2 \notin \mathcal{I}_k } \Gamma_{j_1,j_2,i}^{b,b} \right)^2 \right] \notag \\
&= E\left[ \left( n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_2-2} \sum_{j \notin \mathcal{I}_k } \Gamma_{j,j,i}^{b,b} + n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_2-2} \sum_{j_1 \notin \mathcal{I}_k } \sum_{j_2 \notin \mathcal{I}_k } \Gamma_{j_1,j_2,i}^{b,b} I\{j_1 \neq j_2\} \right)^2 \right] \notag \\
&= I_1 + I_2 + 2I_3 \notag
\end{align}
where
\begin{align*}
I_1 &= E\left[ \left( n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_2-2} \sum_{j \notin \mathcal{I}_k } \Gamma_{j,j,i}^{b,b} \right)^2 \right] \\
I_2 &= E\left[ \left( n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_2-2} \sum_{j_1 \notin \mathcal{I}_k } \sum_{j_2 \notin \mathcal{I}_k } \Gamma_{j_1,j_2,i}^{b,b} I\{j_1 \neq j_2\} \right)^2 \right] \\
I_3 &= E\left[ \left( n^{-1/2} \sum_{k_1=1}^K \sum_{i \in \mathcal{I}_{k_1}} n_0^{-2\varphi_2-2} \sum_{j_1 \notin \mathcal{I}_{k_1} } \Gamma_{j_1,j_1,i_1}^{b,b} \right) \left( n^{-1/2} \sum_{k_2 =1}^K \sum_{i_2 \in \mathcal{I}_{k_2}} n_0^{-2\varphi_2-2} \sum_{j_3 \notin \mathcal{I}_{k_2} } \sum_{j_4 \notin \mathcal{I}_{k_2} } \Gamma_{j_3,j_4,i}^{b,b} I\{j_3 \neq j_4\} \right) \right]
\end{align*}
In what follows, I show that $I_1 = o(n^{-\zeta})$, $I_2 = o(n^{-\zeta})$, and $I_3 = o(n^{-\zeta})$, which is sufficient to complete the proof of Claim 2.\\
\textit{Claim 2.1:} $I_1 = o(n^{-\zeta})$. Consider the following expansion,
\begin{align*}
I_1 &= n^{-1} n_0^{-4\varphi_2-4} \sum_{k_1=1}^K \sum_{k_2=1}^K \sum_{i_1 \in \mathcal{I}_{k_1}} \sum_{i_2 \in \mathcal{I}_{k_2}}\sum_{j_1 \notin \mathcal{I}_{k_1} } \sum_{j_2 \notin \mathcal{I}_{k_2} } E\left[ \Gamma_{j_1,j_1,i_1}^{b,b} \Gamma_{j_2,j_2,i_2}^{b,b} \right] \\
&= n^{-1} n_0^{-4\varphi_2-4} \sum_{(i_1,i_2,j_1,j_2) \in \mathcal{E}} E\left[ \Gamma_{j_1,j_1,i_1}^{b,b} \Gamma_{j_2,j_2,i_2}^{b,b} \right] ,
\end{align*}
where $\mathcal{E} = \{ (i_1,i_2,j_1,j_2) \in [n]^4 : i_1 \in \mathcal{I}_{k_1}, i_2 \in \mathcal{I}_{k_2}, j_1 \notin \mathcal{I}_{k_1}, j_2 \notin \mathcal{I}_{k_2}, k_1 \in [K], k_2 \in [K] \}$, with $[n]$ denoting $\{1,\ldots, n\}$. Let $\mathcal{E}_4 \subseteq [n]^4$ be the subset of indices with distinct entries. Let $\mathcal{E}_{ \le 3} \subset [n]^4$ be the subset of indices with at most three distinct entries.
Now, take $(i_1,i_2,j_1,j_2) \in \mathcal{E} \cap \mathcal{E}_4$. It follows that
$$ E\left[ \Gamma_{j_1,j_1,i_1}^{b,b} \Gamma_{j_2,j_2,i_2}^{b,b} \right] = 0~,$$
since $\Gamma_{j_1,j_1,i_1}^{b,b} $ and $ \Gamma_{j_2,j_2,i_2}^{b,b} $ are zero mean independent random variables.
Now, take $(i_1,i_2,j_1,j_2) \in \mathcal{E} \cap \mathcal{E}_{\le 3}$. Consider the following derivation,
\begin{align*}
n_0 ^{-4\varphi_2} \left| E\left[ \Gamma_{j_1,j_1,i_1}^{b,b} \Gamma_{j_2,j_2,i_2}^{b,b} \right] \right|
\le (p \Tilde{C}_2/C_0) n_0^{3(1-2\varphi_1)} \tau_{n_0} ,
\end{align*}
which follows by C-S and (ref).
Therefore,
$$ \left | n^{-1} n_0^{-4\varphi_2-4} \sum_{(i_1,i_2,j_1,j_2) \in \mathcal{E} \cap \mathcal{E}_{\le 3}} E\left[ \Gamma_{j_1,j_1,i_1}^{b,b} \Gamma_{j_2,j_2,i_2}^{b,b} \right] \right| \le n^{-1} n_0^{-4} 3^4 n^3 (p \Tilde{C}_2/C_0) n_0^{3(1-2\varphi_1)} \tau_{n_0} ~,$$
which uses that $\mathcal{E}_{\le 3}$ has at most $3^4 n^3$ elements (as in the proof of claim 1.1).
Using these two preliminary results, it follows that
\begin{align*}
I_1 = o(n^{1-6\varphi_1}) ,
\end{align*}
since $\tilde{\tau}_{n_0} = o(1)$ and $n/2 \le n_0 \le n$. This completes the proof of Claim 2.2 since $6\varphi_1 - 1 > \zeta$. \\
\textit{Claim 2.2}: $I_2 = o(n^{-\zeta})$. Consider the following expansion,
\begin{align*}
I_2 &= n^{-1} n_0^{-4\varphi_2-4} \sum_{k_1=1}^K \sum_{i_1 \in \mathcal{I}_{k_1}} \sum_{k_2=1}^K \sum_{i_2 \in \mathcal{I}_{k_2}} \sum_{j_1,j_2 \notin \mathcal{I}_{k_1} } \sum_{j_3,j_4 \notin \mathcal{I}_{k_2} } E\left[ \Gamma_{j_1,j_2,i_1}^{b,b} \Gamma_{j_3,j_4,i_2}^{b,b} \right] I\{j_1 \neq j_2\} I\{j_3 \neq j_4\} \\
&= n^{-1} n_0^{-4\varphi_2-4} \sum_{(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E}} E\left[ \Gamma_{j_1,j_2,i_1}^{b,b} \Gamma_{j_3,j_4,i_2}^{b,b} \right] ,
\end{align*}
where $\mathcal{E} = \{ (i_1,i_2,j_1,j_2,j_3,j_4) \in [n]^6 : i_1 \in \mathcal{I}_{k_1}, i_2 \in \mathcal{I}_{k_2}; j_1,j_2 \notin \mathcal{I}_{k_1}; j_3, j_4 \notin \mathcal{I}_{k_1}; j_1 \neq j_2; j_3 \neq j_4; k_1, k_2 \in [K] \}$. Let $\mathcal{E}_6 \subseteq [n]^6$ be the subset of indices with distinct entries. Let $\mathcal{E}_5 \subseteq [n]^6 $ be the subset of indices with exactly five distinct entries, meaning that two entries are identical while the remaining entries are distinct. Let $\mathcal{E}_4 \subseteq [n]^6 $ be the subset of indices with exactly four distinct entries. Let $\mathcal{E}_{\le 3} \subset [n]^6$ be the subset of indices with at most three distinct entries. Note that $[n]^6 = \mathcal{E}_{\le 3} \cup \mathcal{E}_{4} \cup \mathcal{E}_{5} \cup \mathcal{E}_{6}$.
Note that for all $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E}_6$, it follows $E\left[ \Gamma_{j_1,j_2,i_1}^{b,b} \Gamma_{j_3,j_4,i_2}^{b,b} \right] =0$ since $\Gamma_{j_1,j_2,i_1}^{b,b} $ and $ \Gamma_{j_3,j_4,i_2}^{b,b}$ are independent zero mean random variables.
Now, take $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_5$. There are three possible cases:
\begin{itemize}
• Case 1: $j_s = j_r$ for some $s \in \{1,2\}$ and $r \in \{3, 4\} $. Since all the sub-cases are similar, without loss of generality, take $(s,r) = (1,3)$. It follows that
\begin{align*}
n_0^{-2\varphi_2} \left|E\left[ \Gamma_{j_1,j_2,i_1}^{b,b} \Gamma_{j_3,j_4,i_2}^{b,b} \right]\right| &\overset{(1)}{\le} \left| E\left[( n_0^{-\varphi_2}b_{n_0,j_1,i_1})^\top H_{i_1} \tilde{b}_{n_0,i_1} ( n_0^{-\varphi_2}b_{n_0,j_1,i_1})^\top H_{i_2} \tilde{b}_{n_0,i_2} \right] \right|\\
& + n_0^{-2\varphi_2} \left| E\left[ \tilde{b}_{n_0,i_1}^\top H_{i_1} \tilde{b}_{n_0,i_1} \right] E\left[ \tilde{b}_{n_0,i_2}^\top H_{i_2} \tilde{b}_{n_0,i_2} \right] \right| \\
&\overset{(2)}{\le} \tilde{C} E\left[ |( n_0^{-\varphi_2}b_{n_0,j_1,i_1})|^2 | \tilde{b}_{n_0,i_1}|^2 \right] + n_0^{-2\varphi_2} \tilde{C} M_1 \\
&\overset{(3)}{\le} \tilde{C} E\left[ E\left[|( n_0^{-\varphi_2}b_{n_0,j_1,i_1})|^2\mid X_{i_1} \right]^2 \right]^{1/2} E\left[ | \tilde{b}_{n_0,i_1}|^4 \right]^{1/2} + n_0^{-2\varphi_2} \tilde{C} M_1\\
&\overset{(4)}{\le} \tilde{C} n_0^{1-2\varphi_1} \tau_{n_0} M_1^{1/2} + n_0^{-2\varphi_2} \tilde{C} M_1
\end{align*}
where (1) holds by triangular inequality, LIE and definition of $\Gamma_{j_1,j_2,i_1}^{b,b} $ and $ \Gamma_{j_3,j_4,i_2}^{b,b}$ when $j_1 = j_3$ and $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_5$, (2) holds by part (e) of Assumption (ref) with $\tilde{C}$ as a function of $(C_0,C_2,M,p)$ and part (b.3) of Assumption (ref) with C-S, (3) holds by LIE and C-S, and (4) holds by parts (b.1) and (b.3) of Assumption (ref).
• Case 2: $j_s = i_2$ for some $s \in \{1,2\}$ or $j_r = i_1$ for some $r \in \{3, 4 \}$. Since all the sub-cases are similar, take $s=1$. It follows that
\begin{align*}
n_0^{-\varphi_2} \left|E\left[ \Gamma_{j_1,j_2,i_1}^{b,b} \Gamma_{j_3,j_4,j_1}^{b,b} \right]\right|
&\overset{(1)}{\le} \left|E\left[ (n_0^{-\varphi_2}b_{n_0,j_1,i_1})^\top H_{i_1} \tilde{b}_{n_0,i_1} \tilde{b}_{n_0,j_1}^\top H_{j_1} \tilde{b}_{n_0,j_1} \right] \right|\\
& + n_0^{-\varphi_2} \left | E\left[ \tilde{b}_{n_0,i_1}^\top H_{i_1} \tilde{b}_{n_0,i_1} \right] E\left[ \tilde{b}_{n_0,j_1}^\top H_{j_1} \tilde{b}_{n_0,j_1} \right] \right| \\
&\overset{(2)}{\le} \tilde{C} E\left[ |n_0^{-\varphi_2}b_{n_0,j_1,i_1}|^4 \right]^{1/4} E\left[ |\tilde{b}_{n_0,j_1}|^4 \right]^{3/4} +n_0^{-\varphi_2} \tilde{C} M_1 \\
&\overset{(3)}{\le} \tilde{C} n_0^{3(1-2\varphi_1)/4} \tau_{n_0}^{1/4} M_1^{3/4} + n_0^{-\varphi_2} \tilde{C} M_1
\end{align*}
where (1) holds by triangular inequality, LIE, and definition of $\Gamma_{j_1,j_2,i_1}^{b,b} $ and $ \Gamma_{j_3,j_4,i_2}^{b,b}$ when $j_1 = i_2$ and $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_5$; (2) holds by part (e) of Assumption (ref) with $\tilde{C}$ as a function of $(C_0,C_2,M,p)$, C-S, and part (b.3) of Assumption (ref) with C-S; (3) holds by parts (b.3) and (b.4) of Assumption (ref).
• Case 3: $i_1 = i_2$. It follows that
\begin{align*}
\left|E\left[ \Gamma_{j_1,j_2,i_1}^{b,b} \Gamma_{j_3,j_4,j_1}^{b,b} \right]\right|
&\overset{(1)}{\le} \left| E\left[\tilde{b}_{n_0,i_1}^\top H_{i_1} \tilde{b}_{n_0,i_1} \tilde{b}_{n_0,i_1}^\top H_{i_1} \tilde{b}_{n_0,i_1} \right] \right| \\
& + \left| E\left[ \tilde{b}_{n_0,i_1}^\top H_{i_1} \tilde{b}_{n_0,i_1} \right] E\left[ \tilde{b}_{n_0,i_1}^\top H_{i_1} \tilde{b}_{n_0,i_1} \right] \right| \\
&\overset{(2)}{\le} 2 \tilde{C} M_1
\end{align*}
where (1) holds by triangular inequality, LIE, and definition of $\Gamma_{j_1,j_2,i_1}^{b,b} $ and $ \Gamma_{j_3,j_4,i_2}^{b,b}$ when $i_1 = i_2$ and $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_5$; and (2) holds by part (e) of Assumption (ref) with $\tilde{C}$ as a function of $(C_0,C_2,M,p)$, C-S, and part (b.3) of Assumption (ref) with C-S.
\end{itemize}
Therefore, for the indices on $\mathcal{E} \cap \mathcal{E}_5$, it follows that
\begin{align}
n^{-1} n_0^{-4\varphi_2-4} \sum_{(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_5 } E\left[ \Gamma_{j_1,j_2,i_1}^{b,b} \Gamma_{j_3,j_4,i_2}^{b,b} \right] &\overset{(1)}{=} o(n^{1-2\varphi_1-2\varphi_2}) + o(n^{3/4-3\varphi_1/2-3\varphi_2}) + O(n^{-4\varphi_2}) \notag \\
&\overset{(2)}{=} o(n^{-\zeta})
\end{align}
where (1) holds since the number of elements of $\mathcal{E} \cap \mathcal{E}_5$ is lower than $5^6 n^5$ and the preliminary findings in cases 1, 2, and 3, and (2) holds since $2\varphi_1+2\varphi_2 - 1 > \varphi_1 + \varphi_2 - 1/2 \ge \zeta$, $3\varphi_1/2+3\varphi_2-3/4 > \varphi_1 + \varphi_2 - 1/2 \ge \zeta$, and $4\varphi_2 > 4\varphi_1 - 1 \ge \zeta$.
Now, take $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_4$. There are three cases.
\begin{itemize}
• Case 1: $j_1 = j_s$ and $j_2 = j_r$ for $\{r,s\} = \{3,4\}$. Without loss of generality, consider $(s,r) = (3,4)$. It follows
\begin{align*}
n_0^{-4\varphi_2} &\left|E\left[ \Gamma_{j_1,j_2,i_1}^{b,b} \Gamma_{j_1,j_2,i_2}^{b,b} \right]\right| \\
&\overset{(1)}{\le} \left| E\left[( n_0^{-\varphi_2}b_{n_0,j_1,i_1})^\top H_{i_1} ( n_0^{-\varphi_2}b_{n_0,j_2,i_1}) ( n_0^{-\varphi_2}b_{n_0,j_1,i_2})^\top H_{i_2} ( n_0^{-\varphi_2}b_{n_0,j_2,i_2}) \right] \right| \\
& + n_0^{-4\varphi_2} \left| E\left[ \tilde{b}_{n_0,i_1}^\top H_{i_1} \tilde{b}_{n_0,i_1} \right] E\left[ \tilde{b}_{n_0,i_2}^\top H_{i_2} \tilde{b}_{n_0,i_2} \right] \right| \\
&\overset{(2)}{\le} \tilde{C} E\left[ |n_0^{-\varphi_2}b_{n_0,j_1,i_1})| |n_0^{-\varphi_2}b_{n_0,j_2,i_1}| |n_0^{-\varphi_2}b_{n_0,j_1,i_2}| | n_0^{-\varphi_2}b_{n_0,j_2,i_2} | \right] + n_0^{-4\varphi_2} \tilde{C} M_1 \\
&\overset{(3)}{=} \tilde{C} E\left[ E\left[ |n_0^{-\varphi_2}b_{n_0,j_1,i_1})| |n_0^{-\varphi_2}b_{n_0,j_1,i_2}| \mid X_{i_1}, X_{i_2} \right] E\left[ |n_0^{-\varphi_2}b_{n_0,j_2,i_1}| | n_0^{-\varphi_2}b_{n_0,j_2,i_2} | \mid X_{i_1}, X_{i_2} \right] \right] \\
& + n_0^{-4\varphi_2} \tilde{C} M_1 \\
&\overset{(4)}{\le} \tilde{C} E\left[ E\left[ |n_0^{-\varphi_2}b_{n_0,j_1,i_1})|^2 \mid X_{i_1} \right] E\left[ | n_0^{-\varphi_2}b_{n_0,j_2,i_2} |^2 \mid X_{i_2} \right] \right] + n_0^{-4\varphi_2} \tilde{C} M_1 \\
&\overset{(5)}{\le} \tilde{C} E\left[ E\left[ |n_0^{-\varphi_2}b_{n_0,j_1,i_1})|^2 \mid X_{i_1} \right]^2 \right]^{1/2} E\left[ E\left[ | n_0^{-\varphi_2}b_{n_0,j_2,i_2} |^2 \mid X_{i_2} \right]^{2} \right]^{1/2} + n_0^{-4\varphi_2} \tilde{C} M_1 \\
&\overset{(6)}{\le} \tilde{C} n_0^{2(1-2\varphi_1)} \tau_{n_0} + n_0^{-4\varphi_2} \tilde{C} M_1
\end{align*}
where (1) holds by triangular inequality, LIE, and definition of $\Gamma_{j_1,j_2,i_1}^{b,b} $ and $ \Gamma_{j_3,j_4,i_2}^{b,b}$ when $j_1 = j_3$, $j_2 = j_4$ and $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_5$; (2) holds by part (e) of Assumption (ref) with $\tilde{C}$ as a function of $(C_0,C_2,M,p)$, C-S, and part (b.3) of Assumption (ref) with C-S; (3) holds by LIE; (4) and (5) holds by C-S; and (6) holds by part (b.1) of Assumption (ref).
• Case 2: $j_1 = j_s$ for $s \in \{3,4\}$ and $j_2 = i_2$. Without loss of generality, $s=3$. It follows
\begin{align*}
n_0^{-3\varphi_2} \left|E\left[ \Gamma_{j_1,j_2,i_1}^{b,b} \Gamma_{j_1,j_2,i_2}^{b,b} \right]\right|
&\overset{(1)}{\le} \left| E\left[( n_0^{-\varphi_2}b_{n_0,j_1,i_1})^\top H_{i_1} ( n_0^{-\varphi_2}b_{n_0,j_2,i_1}) ( n_0^{-\varphi_2}b_{n_0,j_1,j_2})^\top H_{j_2} \tilde{b}_{n_0,j_2}) \right] \right| \\
& + n_0^{-3\varphi_2} \left| E\left[ \tilde{b}_{n_0,i_1}^\top H_{i_1} \tilde{b}_{n_0,i_1} \right] E\left[ \tilde{b}_{n_0,j_2}^\top H_{j_2} \tilde{b}_{n_0,j_2} \right] \right| \\
&\overset{(3)}{\le} \tilde{C} n_0^{9(1-2\varphi_1)/4} \tau_{n_0}^{3/4} M_1^{1/4} + \tilde{C} n_0^{-3\varphi_2} M_1
\end{align*}
where (1) holds by triangular inequality, LIE, and definition of $\Gamma_{j_1,j_2,i_1}^{b,b} $ and $ \Gamma_{j_3,j_4,i_2}^{b,b}$ when $j_1 = j_3$, $j_2 = i_2$ and $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_5$; and (2) holds by C-S and parts (b.3) and (b.4) of Assumption (ref).
• Case 3: $j_1 = j_s$ for $s \in \{3, 4\}$ and $i_1 = i_2$. Without loss of generality, consider $s = 3$. It follows that
\begin{align*}
n_0^{-2\varphi_2} \left|E\left[ \Gamma_{j_1,j_2,i_1}^{b,b} \Gamma_{j_1,j_4,i_1}^{b,b} \right]\right|
&\overset{(1)}{\le} \left| E\left[( n_0^{-\varphi_2}b_{n_0,j_1,i_1})^\top H_{i_1} \tilde{b}_{n_0,i_1} ( n_0^{-\varphi_2}b_{n_0,j_1,i_1})^\top H_{i_1} \tilde{b}_{n_0,i_1} \right] \right|\\
& + n_0^{-2\varphi_2} \left| E\left[ \tilde{b}_{n_0,i_1}^\top H_{i_1} \tilde{b}_{n_0,i_1} \right] E\left[ \tilde{b}_{n_0,i_1}^\top H_{i_1} \tilde{b}_{n_0,i_1} \right] \right| \\
&\overset{(2)}{\le} \tilde{C} n_0^{1-2\varphi_1} \tau_{n_0} M_1^{1/2} + n_0^{-2\varphi_2} \tilde{C} M_1
\end{align*}
where (1) holds by triangular inequality, LIE and definition of $\Gamma_{j_1,j_2,i_1}^{b,b} $ and $ \Gamma_{j_3,j_4,i_2}^{b,b}$ when $j_1 = j_3$ and $i_1=i_2$ and $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_4$; and (2) holds by the same derivations presented in Case 1 when $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_5$; therefore, it is omitted.
\end{itemize}
Therefore, for the indices on $\mathcal{E} \cap \mathcal{E}_4$, it follows that
\begin{align}
n^{-1} n_0^{-4\varphi_2-4} \sum_{(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_4 } E\left[ \Gamma_{j_1,j_2,i_1}^{b,b} \Gamma_{j_3,j_4,i_2}^{b,b} \right] &\overset{(1)}{=} o(n^{1-4\varphi_1}) + o(n^{5/4-9\varphi_1/2-\varphi_2}) + O(n^{-2\varphi_1-2\varphi_2}) \notag \\
&\overset{(2)}{=} o(n^{-\zeta})
\end{align}
where (1) holds since the number of elements of $\mathcal{E} \cap \mathcal{E}_4$ is lower than $4^6 n^4$ and the preliminary findings in cases 1, 2, and 3; and (2) holds since $4\varphi_1 - 1 \ge \zeta$, $ 9\varphi_1/2 + \varphi_2 - 5/4 > \varphi_1 + \varphi_2 - 1/2 \ge \zeta$, and $2\varphi_1+2\varphi_2 >\varphi_1 + \varphi_2 - 1/2 \ge \zeta$.
Now, take $(i_1,i_2,j_1,j_2,j_3,j_4) \in \mathcal{E} \cap \mathcal{E}_{\le 3}$. Similar to the proof of Claim 2.1 but using (ref) instead of (ref), it follows that
\begin{equation}
\left | n^{-1} n_0^{-4\varphi_2-4} \sum_{(i_1,i_2,j_1,j_2) \in \mathcal{E} \cap \mathcal{E}_{\le 3}} E\left[ \Gamma_{j_1,j_1,i_1}^{b,b} \Gamma_{j_2,j_2,i_2}^{b,b} \right] \right| = o(n^{-\zeta})
\end{equation}
Finally, using (ref), (ref), and (ref), it follows that $I_2 = o(n^{-\zeta})$. This completes the proof of Claim 2.2.\\
\textit{Claim 2.3:} $I_3 = o(n^{-\zeta})$. This result is a consequence of C-S and Claims 2.1 and 2.2.\\
\textit{Claim 3:} $E[I_{l,b}^2] = o(n^{-\zeta})$. Consider the following notation,
$$ \Gamma_{j_1,j_2,i}^{l,b} = \delta_{n_0,j_1,i}^\top H_i b_{n_0,j_2,i}$$
where it holds $E[\Gamma_{j_1,j_2,i}^{l,b}] = 0$ due to part (a) of Assumption (ref).
Furthermore,
\begin{equation}
n_0^{-2\varphi_2} E\left [ |\Gamma_{j_1,j_2,i}^{l,b}|^2 \right] \le \tilde{C} n_0^{2(1-2\varphi_1) } \tau_{n_0}^{1/2} M_1^{1/2} ,
\end{equation}
which follows by C-S, part (e) of Assumption (ref), and parts (b.2) and (b.4) of Assumption (ref), with $\tilde{C}$ function of $(C_2, M, C_0,p)$.
The previous notation can be used to rewrite $E[I_{l,b}^2]$ as follows
\begin{align*}
&= E\left[ \left(n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1-\varphi_2-3/2} \sum_{j_1 \notin \mathcal{I}_k } \sum_{j_2 \notin \mathcal{I}_k } \Gamma_{j_1,j_2,i}^{l,b} \right)^2 \right] \\
&= E\left[ \left(n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1-\varphi_2-3/2} \sum_{j \notin \mathcal{I}_k } \Gamma_{j,i}^{l,b} + n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1-\varphi_2-3/2} \sum_{j_1,j_2 \notin \mathcal{I}_k } \Gamma_{j_1,j_2,i}^{l,b} I\{j_1 \neq j_2\} \right)^2 \right] \\
&= I_1 + I_2 + 2I_3
\end{align*}
where
\begin{align*}
I_1 &= E\left[ \left(n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1-\varphi_2-3/2} \sum_{j \notin \mathcal{I}_k } \Gamma_{j,j,i}^{l,b} \right)^2 \right] \\
I_2 &= E\left[ \left( n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1-\varphi_2-3/2} \sum_{j_1,j_2 \notin \mathcal{I}_k } \Gamma_{j_1,j_2,i}^{l,b} I\{j_1 \neq j_2\} \right)^2 \right] \\
I_3 &= E\left[ \left(n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1-\varphi_2-3/2} \sum_{j \notin \mathcal{I}_k } \Gamma_{j,j,i}^{l,b} \right)\left(n^{-1/2} \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1-\varphi_2-3/2} \sum_{j_1,j_2 \notin \mathcal{I}_k } \Gamma_{j_1,j_2,i}^{l,b} I\{j_1 \neq j_2\} \right) \right]
\end{align*}
In what follows, I show that $I_1 = o(n^{-\zeta})$, $I_2 = o(n^{-\zeta})$, and $I_3 = o(n^{-\zeta})$, which is sufficient to complete the proof of Claim 2.\\
\textit{Claim 3.1:} $I_1 = o(n^{-\zeta})$. Consider the following expansion,
\begin{align*}
I_1 &= n^{-1} n_0^{-2\varphi_1-2\varphi_2-3} \sum_{k_1,k_2=1}^K \sum_{i_1 \in \mathcal{I}_{k_1}} \sum_{i_1 \in \mathcal{I}_{k_1}}\sum_{j_1 \notin \mathcal{I}_{k_1} } \sum_{j_2 \notin \mathcal{I}_{k_2} } E[\Gamma_{j_1,j_1,i_1}^{l,b} \Gamma_{j_2,j_2,i_2}^{l,b}] .
\end{align*}
Now, $E[\Gamma_{j_1,j_1,i_1}^{l,b} \Gamma_{j_2,j_2,i_2}^{l,b}] $ is calculated under the two possible cases based on the indices $(i_1,j_1,i_2,j_2)$.
\begin{itemize}
• Case 1: $(i_1,j_1)$ and $(i_2,j_2)$ have no element in common. Then $E[\Gamma_{j_1,j_1,i_1}^{l,b} \Gamma_{j_2,j_2,i_2}^{l,b}] $ is zero since $\Gamma_{j_1,j_1,i_1}^{l,b} $ and $ \Gamma_{j_2,j_2,i_2}^{l,b}$ are independent zero mean random variables.
• Case 2: $(i_1,j_1)$ and $(i_2,j_2)$ have at least one element in common. In this case, there are at most $3^4 n^3$ possible indices. Moreover, due to (ref) and C-S, it follows
$$ |E[\Gamma_{j_1,j_1,i_1}^{l,b} \Gamma_{j_2,j_2,i_2}^{l,b}]| \le \tilde{C} n_0^{3(1-2\varphi_1)/2} \tau_{n_0}^{1/2} n_0^{(1-2\varphi_1)} M_1^{1/2}~. $$
\end{itemize}
Therefore,
\begin{align*}
| I_1| &\le n^{-1} n_0^{-2\varphi_1-3} 3^4 n^3 \tilde{C} n_0^{2(1-2\varphi_1) } \tau_{n_0}^{1/2} M_1^{1/2} \\
&= o(n^{1-6\varphi_1}) ,
\end{align*}
which is sufficient to conclude that $I_1$ is $o(n^{-\zeta})$ since $6\varphi_1 - 1 > 4\varphi_1-1 \ge \zeta$. This completes the proof of Claim 3.1.\\
\textit{Claim 3.2:} $I_2 = o(n^{-\zeta})$. Consider the following expansion,
\begin{align*}
I_2 = n^{-1} \sum_{k_1,k_2=1}^K \sum_{i_1 \in \mathcal{I}_{k_1}} \sum_{i_2 \in \mathcal{I}_{k_2}} n_0^{-2\varphi_1-2\varphi_2-3} \sum_{j_1,j_2 \notin \mathcal{I}_{k_1} } \sum_{j_3,j_4 \notin \mathcal{I}_{k_2} } E\left[ \Gamma_{j_1,j_2,i_1}^{l,b} \Gamma_{j_3,j_4,i_2}^{l,b} \right] I\{j_1 \neq j_2\} I\{j_3 \neq j_4\}
\end{align*}
Now, $ E\left[ \Gamma_{j_1,j_2,i_1}^{l,b} \Gamma_{j_3,j_4,i_2}^{l,b} \right]$ is calculated under four possible cases based on the indices.
\begin{itemize}
• Case 1: all indices are different. Then , $ E\left[ \Gamma_{j_1,j_2,i_1}^{l,b} \Gamma_{j_3,j_4,i_2}^{l,b} \right] $ equals zero since $\Gamma_{j_1,j_1,i_1}^{l,b} $ and $ \Gamma_{j_2,j_2,i_2}^{l,b}$ are independent zero mean random variables.
• Case 2: there are exactly five different indices. Then, consider the four different sub-cases:
\begin{itemize}
• $j_1 = j_3$, then
\begin{align*}
\left| E\left[ \Gamma_{j_1,j_2,i_1}^{l,b} \Gamma_{j_1,j_4,i_2}^{l,b} \right] \right|
&= \left| E\left[ \delta_{n_0,j_1,i_1}^\top H_{i_1} \tilde{b}_{n_0,i_1} \delta_{n_0,j_1,i_2}^\top H_{i_2} \tilde{b}_{n_0,i_2}\right]\right| \\
&\le \tilde{C} E\left[ |\delta_{n_0,j_1,i_1}| | \tilde{b}_{n_0,i_1}| | \delta_{n_0,j_1,i_2}| | \tilde{b}_{n_0,i_2}|\right] \\
&\le \tilde{C} E\left[ E\left[ |\delta_{n_0,j_1,i_1}|^2 \mid X_{i_1} \right] | \tilde{b}_{n_0,i_1}|^2 \right] \\
&=O(1)
\end{align*}
which holds by C-S, LIE, and part (b.1) and (b.3) of Assumption (ref). Since there are at most $5^6 n^5$ terms, it follows these terms contributed to $I_2$ with $ O(n^{1-2\varphi_1 -2\varphi_2})$ which is larger than $o(n^{-\zeta})$ since $2\varphi_1+2\varphi_1-1 > \varphi_1 + \varphi_2 - 1/2 \ge \zeta$.
• $j_1 = j_4$, then
\begin{align*}
E\left[ \Gamma_{j_1,j_2,i_1}^{l,b} \Gamma_{j_3,j_1,i_2}^{l,b} \right] &= E\left[ \delta_{n_0,j_1,i_1}^\top H_{i_1} b_{n_0,j_2,i_1} \delta_{n_0,j_3,i_2}^\top H_{i_2} b_{n_0,j_1,i_2}\right] \\
&= E[ E[ \delta_{n_0,j_1,i_1}^\top \mid X_{j_1}, W_{i_1}, W_{i_2}, W_{j_2}, W_{j_3}] H_{i_1} b_{n_0,j_2,i_1} \delta_{n_0,j_3,i_2}^\top H_{i_2} b_{n_0,j_1,i_2}] \\
& = 0 ,
\end{align*}
which holds due to part (a) of Assumption (ref).
• $j_1 = i_2$, then
\begin{align*}
E\left[ \Gamma_{j_1,j_2,i_1}^{l,b} \Gamma_{j_3,j_4,j_1}^{l,b} \right] &= E\left[ \delta_{n_0,j_1,i_1}^\top H_{i_1} b_{n_0,j_2,i_1} \delta_{n_0,j_3,j_1}^\top H_{j_1} b_{n_0,j_4,j_1} \right] \\
&= E\left[ \delta_{n_0,j_1,i_1}^\top H_{i_1} b_{n_0,j_2,i_1} E\left[\delta_{n_0,j_3,j_1}^\top \mid W_{i_1}, W_{j_1}, W_{j_2}, W_{j_4} \right] H_{j_1} b_{n_0,j_4,j_1}\right] \\
&= 0.
\end{align*}
• $i_1 = i_2$, then $j_3$ is different than all and the previous argument used for $j_1 = i_2$ applies and implies
$$ E\left[ \Gamma_{j_1,j_2,i_1}^{l,b} \Gamma_{j_3,j_4,i_2}^{l,b} \right] = 0$$
\end{itemize}
• Case 3: there are exactly four different indices. Then
\begin{itemize}
• if $j_2 $ or $j_4$ is different than all, then
\begin{align*}
\left| E\left[ \Gamma_{j_1,j_2,i_1}^{l,b} \Gamma_{j_3,j_4,i_2}^{l,b} \right] \right| &= \left| E\left[ \delta_{n_0,j_1,i_1}^\top H_{i_1} \tilde{b}_{n_0,i_1} \delta_{n_0,j_3,i_2}^\top H_{i_2} \tilde{b}_{n_0,i_2} \right] \right|\\
&= \tilde{C} E\left[ |\delta_{n_0,j_1,i_1}| |\tilde{b}_{n_0,i_1}| |\delta_{n_0,j_3,i_2}| |\tilde{b}_{n_0,i_2}| \right] \\
&= O(1)
\end{align*}
which holds by C-S, LIE, and parts (b.1) and (b.3) of Assumption (ref). Since there are at most $4^6 n^4$ terms, it follows these terms, in this case, contributed to $I_2$ with $O(n^{-2\varphi_1 - 2\varphi_2})$, which is $o(n^{-\zeta})$ since $2\varphi_1 + 2\varphi_2>4\varphi_1-1$.
• if $j_1 = j_4$ and $j_2 = j_3$, then
\begin{align*}
E\left[ \Gamma_{j_1,j_2,i_1}^{l,b} \Gamma_{j_2,j_1,i_2}^{l,b} \right] &= E\left[ \delta_{n_0,j_1,i_1}^\top H_{i_1} b_{n_0,j_2,i_1} \delta_{n_0,j_2,i_2}^\top H_{i_2} b_{n_0,j_1,i_2}\right] \\
&= E[ E[ \delta_{n_0,j_1,i_1}^\top \mid X_{j_1}, W_{i_1}, W_{j_2}, W_{i_2} ] H_{i_1} b_{n_0,j_2,i_1} \delta_{n_0,j_2,i_2}^\top H_{i_2} b_{n_0,j_1,i_2}]\\
&= 0 ,
\end{align*}
which follows by part (a) of Assumption (ref).
• if $j_2 = j_3$ and $i_1 = j_4$, then
\begin{align*}
E\left[ \Gamma_{j_1,j_2,i_1}^{l,b} \Gamma_{j_2,i_1,i_2}^{l,b} \right] &= E\left[ \delta_{n_0,j_1,i_1}^\top H_{i_1} b_{n_0,j_2,i_1} \delta_{n_0,j_2,i_2}^\top H_{i_2} b_{n_0,i_1,i_2}\right] \\
&= E[ E[ \delta_{n_0,j_1,i_1}^\top \mid X_{j_1}, W_{i_1}, W_{j_2}, W_{i_2}] H_{i_1} b_{n_0,j_2,i_1} \delta_{n_0,j_2,i_2}^\top H_{i_2} b_{n_0,i_1,i_2}] \\
&= 0 ,
\end{align*}
which follows by part (a) of Assumption (ref).
• if $j_1 = j_3$ and $j_2 = j_4$, then
\begin{align*}
n_0^{-2\varphi_2}\left| E\left[ \Gamma_{j_1,j_2,i_1}^{l,b} \Gamma_{j_1,j_2,i_2}^{l,b} \right] \right| &= \left| E\left[ \delta_{n_0,j_1,i_1}^\top H_{i_1} b_{n_0,j_2,i_1} \delta_{n_0,j_1,i_2}^\top H_{i_2} b_{n_0,j_2,i_2}\right]\right| \\
&\le \tilde{C} E\left[ E\left[ |\delta_{n_0,j_1,i_1}|^2 \mid X_{i_1} \right] E\left[ | n_0^{-\varphi_2}b_{n_0,j_2,i_1}|^2 \mid X_{i_1} \right] \right] \\
&\le \tilde{C} E\left[ E\left[ |\delta_{n_0,j_1,i_1}|^2 \mid X_{i_1} \right]^2 \right] ^{1/2} E\left[ E\left[ | n_0^{-\varphi_2}b_{n_0,j_2,i_1}|^2 \mid X_{i_1} \right]^2 \right]^{1/2} \\
&\le \tilde{C} M_1^{1/2} n_0^{(1-2\varphi_1)} \tau_{n_0}^{1/2}
\end{align*}
which holds due to C-S, LIE, part (b.1) of Assumption (ref). Since there are at most $4^6 n^4$ terms, it follows these terms contributed to $I_2$ with $o(n^{1-4\varphi_1})$, which is $o(n^{-\zeta})$ since $4\varphi_1-1 \ge \zeta$.
• $j_1 = i_2$ and $j_2 = j_4$, then
\begin{align*}
E\left[ \Gamma_{j_1,j_2,i_1}^{l,b} \Gamma_{j_3,j_2,j_1}^{l,b} \right]
&= E\left[ \delta_{n_0,j_1,i_1}^\top H_{i_1} b_{n_0,j_2,i_1} \delta_{n_0,j_3,j_1}^\top H_{j_1} b_{n_0,j_2,j_1}\right] \\
&= E[ \delta_{n_0,j_1,i_1}^\top H_{i_1} b_{n_0,j_2,i_1} E[ \delta_{n_0,j_3,j_1}^\top \mid X_{j_3}, W_{i_1}, W_{j_2}, W_{j_1}] H_{j_1} b_{n_0,j_2,j_1}] \\
&= 0 ,
\end{align*}
which holds due to part (a) of Assumption (ref).
\end{itemize}
• Case 4: there are exactly three different indices. All the terms in this case contributed to $I_2$ with $o(n^{-\zeta})$ by a similar argument as Case 2 in the proof of Claim 3.1.
\end{itemize}
All the previous cases imply that $I_2 = o(n^{-\zeta})$, which completes the proof of Claim 3.2.\\
\textit{Claim 3.3:} $I_3 = o(n^{-\zeta})$. This result is a consequence of C-S and Claims 3.1 and 3.2.\\
\textbf{Part 3:} It follows by Cauchy-Schwartz, using part 2 of this proposition and part 1 of Proposition (ref).
\textbf{Part 4:} It follows by Cauchy-Schwartz, using part 2 of this proposition and part 2 of Proposition (ref).
\end{proof}
\subsection{Proof of Proposition (ref)}
\begin{proof}
For $i \in \mathcal{I}_k$, denote $\Delta_{i} = \Delta_{i}^b + \Delta_{i}^l$, where
\begin{align*}
\Delta_{i}^l &= n_0^{-\varphi_1} n_0^{-1/2} \sum_{j \notin \mathcal{I}_k} \delta_{n_0,j,i} ,\\
\Delta_{i}^b &= n_0^{-\varphi_2} n_0^{-1} \sum_{j \notin \mathcal{I}_k} b_{n_0,j,i} ,
\end{align*}
Here, $\delta_{n_0,j,i} = \delta_{n_0}(W_{j},X_i)$ and $b_{n_0,j,i} = b_{n_0}(X_{j},X_i)$, and $\delta_{n_0}$ and $b_{n_0}$ are functions satisfying Assumption (ref).
\textbf{Part 1:}
Consider the following decomposition
\begin{align*}
E[ \mathcal{T}_n^{*} \mathcal{T}_{n,K}^{l} ] &= E\left[ \left( n^{-1/2} \sum_{i_1=1}^n m_{i_1}/J_0 \right) \left( n^{-1/2} \sum_{i_2=1}^n (\Delta_{i_2})^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
&\overset{(1)}{=} n^{-1} \sum_{i_1=1}^n \sum_{i_2=1}^n E\left[ \left( m_{i_1}/J_0 \right) \left( (\Delta_{i_2}^l + \Delta_{i_2}^b)^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
&= n^{-1} \sum_{i_1=1}^n \sum_{k=1}^K \sum_{i_2 \in \mathcal{I}_k} E\left[ \left( m_{i_1}/J_0 \right) \left( (\Delta_{i_2}^l )^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] + E\left[ \left( m_{i_1}/J_0 \right) \left( ( \Delta_{i_2}^b)^{\top} \partial_\eta m_{i_2}/J_0 \right) \right]\\
&= I_{1} + I_{2} ,
\end{align*}
where (1) holds since $\Delta_{i} = \Delta_{i}^l + \Delta_{i}^b$. Claim 1.1 below shows that $ I_{1} = 0$, while Claim 1.2 shows $ I_{2} = (G_b^l/2) n_0^{-\varphi_2} + o(n^{-\varphi_2})$.
\textit{Claim 1:} $ I_{1} = 0$. To see this, consider the following derivations
\begin{align*}
I_{1} &= n^{-1} \sum_{i_1=1}^n \sum_{k=1}^K \sum_{i_2 \in \mathcal{I}_k} E\left[ \left( m_{i_1}/J_0 \right) \left( (\Delta_{i_2}^l )^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
&\overset{(1)}{=} n^{-1} \sum_{i_1=1}^n \sum_{k=1}^K \sum_{i_2 \in \mathcal{I}_k} n_0^{-\varphi_1} n_0^{-1/2} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{i_1}/J_0 \right) \delta_{n,j,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right] \\
&\overset{(2)}{=} n^{-1} \sum_{i_1=1}^n \sum_{k=1}^K \sum_{i_2 \in \mathcal{I}_k} n_0^{-\varphi_1} n_0^{-1/2} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{i_1}/J_0 \right) E\left[ \delta_{n,j,i_2}^{\top} \mid W_{i_1}, W_{i_2} \right] \partial_\eta m_{i_2}/J_0\right] \\
&\overset{(3)}{=} n^{-1} \sum_{k=1}^K \sum_{i_2 \in \mathcal{I}_k} n_0^{-\varphi_1} n_0^{-1/2} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{j}/J_0 \right) \delta_{n,j,i_2}^{\top} E\left[\partial_\eta m_{i_2}/J_0 \mid X_{i_2}, W_j \right] \right] \\
&\overset{(4)}{=} 0
\end{align*}
where (1) holds by definition of $\Delta_{i}^l$, (2) holds by the law of iterative expectations, (3) holds since $ E\left[ \delta_{n,j,i_2}^{\top} \mid W_{i_1}, W_{i_2} \right] = 0$ when $i_1 \neq j$ due to part (a) of Assumption (ref) and by the law of iterative expectations, and (4) holds by the Neyman orthogonality condition implied by part (b) of Assumption (ref).
\textit{Claim 2:} $ I_{2} = (G_b^l/2) n_0^{-\varphi_2} + o(n^{-\varphi_2})$. To see this, consider the following derivations
\begin{align*}
I_{2} &= n^{-1} \sum_{i_1=1}^n \sum_{k=1}^K \sum_{i_2 \in \mathcal{I}_k} E\left[ \left( m_{i_1}/J_0 \right) \left( (\Delta_{i_2}^b )^{\top} \partial_\eta m_{i_2}/J_0 \right) \right] \\
&\overset{(1)}{=} n^{-1} \sum_{i_1=1}^n \sum_{k=1}^K \sum_{i_2 \in \mathcal{I}_k} n_0^{-\varphi_2} n_0^{-1} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{i_1}/J_0 \right) b_{n_0,j,i_2}^{\top} \partial_\eta m_{i_2}/J_0 \right] \\
&\overset{(2)}{=} n^{-1} \sum_{i_1=1}^n \sum_{k=1}^K \sum_{i_2 \in \mathcal{I}_k} n_0^{-\varphi_2} n_0^{-1} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{i_1}/J_0 \right) b_{n_0,j,i_2}^{\top} E\left[ \partial_\eta m_{i_2}/J_0 \mid X_{i_2}, W_{i_1}, X_j \right] \right] \\
&\overset{(3)}{=} n^{-1} \sum_{k=1}^K \sum_{i_2 \in \mathcal{I}_k} n_0^{-\varphi_2} n_0^{-1} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{i_2}/J_0 \right) E\left[ b_{n_0,j,i_2}^{\top} \mid W_{i_2} \right] \partial_\eta m_{i_2}/J_0 \right] \\
&\overset{(4)}{=} n_0^{-\varphi_2} E\left[ \left( m_{i_2}/J_0 \right) \tilde{b}_{n_0}(X_{i_2}) \partial_\eta m_{i_2}/J_0 \right] \\
&\overset{(4)}{=} n_0^{-\varphi_2} (G_b^l/2) + o(n_0^{-\varphi_2}) ,
\end{align*}
where (1) holds by definition of $\Delta_{i}^b$, (2) holds by the law of iterative expectations, (3) holds since $ E\left[ \partial_\eta m_{i_2}/J_0 \mid X_{i_2}, W_{i_1}, X_j \right] = 0$ when $i_1 \neq i_2$ due to the Neyman orthogonality condition implied by part (b) of Assumption (ref) and the law of iterative expectations, (4) holds by definitions of $\tilde{b}_{n_0,i} = E[b_{n_0,j,i} \mid X_i]$ which is equal to $ E\left[ b_{n_0,j,i} \mid W_{i} \right] $, and (5) holds by definition of $G_b^l$ in (ref) and Assumption (ref).\\
\textbf{Part 2:} Consider the following decomposition,
\begin{align*}
E[\mathcal{T}_n^* \mathcal{T}_{n,K}^{nl}] &= E\left[ \left( n^{-1/2 }\sum_{i_1 = 1}^n m_{i_1}/J_0 \right) \left( n^{-1/2 }\sum_{i_2 = 1}^n (\Delta_{i_2})^\top H_{i_2} ( \Delta_{i_2} ) \right) \right] \\
&= I_{1} + 2I_2 + I_3
\end{align*}
where
\begin{align*}
I_1 &= E\left[ \left( n^{-1/2 }\sum_{i_1 = 1}^n m_{i_1}/J_0 \right) \left( n^{-1/2 }\sum_{i_2 = 1}^n (\Delta_{i_2}^l )^\top H_{i_2} ( \Delta_{i_2}^l ) \right) \right] \\
I_2 &= E\left[ \left( n^{-1/2 }\sum_{i_1 = 1}^n m_{i_1}/J_0 \right) \left( n^{-1/2 }\sum_{i_2 = 1}^n (\Delta_{i_2}^l )^\top H_{i_2} ( \Delta_{i_2}^b ) \right) \right] \\
I_3 &= E\left[ \left( n^{-1/2 }\sum_{i_1 = 1}^n m_{i_1}/J_0 \right) \left( n^{-1/2 }\sum_{i_2 = 1}^n (\Delta_{i_2}^b )^\top H_{i_2} ( \Delta_{i_2}^b ) \right) \right]
\end{align*}
In what follows, I show that $I_1 = o(n^{-\zeta})$, $I_2 = (G_b/2) n_0^{1/2-\varphi_1-\varphi_2} + o(n^{-\zeta})$, and $I_3 = o(n^{-\zeta})$.\\
\textit{Claim 1:} $I_1 = o(n^{-\zeta})$. Consider the following derivations,
\begin{align*}
I_1
&\overset{(1)}{=} n^{-1 }\sum_{i_1 = 1}^n \sum_{k=1}^K \sum_{i_2 \in \mathcal{I}_k} n_0^{-2\varphi_1} n_0^{-1} \sum_{j_1 \notin \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_k} E\left[ \left( m_{i_1}/J_0 \right) \left((\delta_{n,j_1,i_2} )^\top H_{i_2} ( \delta_{n,j_2,i_2} ) \right) \right] \\
&\overset{(2)}{=} n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_1} n_0^{-1} \sum_{j_1 \notin \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_k} E\left[ \left( m_{i}/J_0 \right) (\delta_{n,j_1,i} )^\top H_{i} ( \delta_{n,j_2,i} ) \right] \\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_1} n_0^{-1} \sum_{j_1 \notin \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_k} E\left[ \left( m_{j_1}/J_0 \right) (\delta_{n,j_1,i} )^\top H_{i} ( \delta_{n,j_2,i} ) \right] \\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_1} n_0^{-1} \sum_{j_1 \notin \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_k} E\left[ \left( m_{j_2}/J_0 \right) (\delta_{n,j_1,i} )^\top H_{i} ( \delta_{n,j_2,i} ) \right] \\
&\overset{(3)}{=} n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_1} n_0^{-1} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{i}/J_0 \right) (\delta_{n,j,i} )^\top H_{i} ( \delta_{n,j,i} ) \right] \\
& + 2 n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_1} n_0^{-1} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{j}/J_0 \right) (\delta_{n,j,i} )^\top H_{i} ( \delta_{n,j,i} ) \right] \\
&= n_0^{-2\varphi_1} E\left[ \left( m_{i}/J_0 \right) (\delta_{n,j,i} )^\top H_{i} ( \delta_{n,j,i} ) \right] + 2 n_0^{-2\varphi_1} E\left[ \left( m_{j}/J_0 \right) (\delta_{n,j,i} )^\top H_{i} ( \delta_{n,j,i} ) \right] \\
&\overset{(4)}{=} O(n^{-2\varphi_1})
\end{align*}
where (1) holds by definition of $\Delta_{i}^l$, (2) holds since $i_1 \notin \{ i_2, j_1,j_2 \}$ implies
$$ E\left[ \left( m_{i_1}/J_0 \right) \left((\delta_{n,j_1,i_2} )^\top H_{i_2} ( \delta_{n,j_2,i_2} ) \right) \right] = 0~,$$
which follows since $m_{i_1}$ is a zero mean random variable independent of $W_{i_2}$, $W_{j_1}$ and $W_{j_2}$, (3) holds since $j_1 \neq j_2$ implies
\begin{align*}
E\left[ \left( m_{i}/J_0 \right) (\delta_{n,j_1,i} )^\top H_{i} ( \delta_{n,j_2,i} ) \right] &= 0\\
E\left[ \left( m_{j_1}/J_0 \right) (\delta_{n,j_1,i} )^\top H_{i} ( \delta_{n,j_2,i} ) \right] &= 0 \\
E\left[ \left( m_{j_2}/J_0 \right) (\delta_{n,j_1,i} )^\top H_{i} ( \delta_{n,j_2,i} ) \right] &=0 ,
\end{align*}
which follows by the law of iterative expectations and noting $E[ \delta_{n,j_2,i} \mid W_{i}, W_{j_1}] =0$ and $E[ \delta_{n,j_1,i} \mid W_{i}, W_{j_2}] =0$ (due to part (a) of Assumption (ref)), and (4) holds by Holder's inequality, part (e) of Assumption (ref), and part (a) of Assumption (ref).\\
\textit{Claim 2:} $I_2 = (G_b/2) n_0^{1/2-\varphi_1-\varphi_2} + o(n^{-\zeta})$. Consider the following derivations,
\begin{align*}
I_2
&\overset{(1)}{=} n^{-1 }\sum_{i_1 = 1}^n \sum_{k=1}^K \sum_{i_2 \in \mathcal{I}_k} n_0^{-\varphi_1 - \varphi_2} n_0^{-3/2} \sum_{j_1 \notin \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_k} E\left[ \left( m_{i_1}/J_0 \right) \left((\delta_{n,j_1,i_2} )^\top H_{i_2} ( b_{n_0,j_2,i_2} ) \right) \right] \\
&\overset{(2)}{=} n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1 - \varphi_2-3/2} \sum_{j_1, j_2 \notin \mathcal{I}_k} E\left[ \left( m_{i}/J_0 \right) (\delta_{n,j_1,i} )^\top H_{i} ( b_{n_0,j_2,i} ) \right] \\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1 - \varphi_2-3/2} \sum_{j_1, j_2 \notin \mathcal{I}_k} E\left[ \left( m_{j_1}/J_0 \right) (\delta_{n,j_1,i} )^\top H_{i} ( b_{n_0,j_2,i} ) \right] I\{j_1 \neq j_2\}\\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1 - \varphi_2-3/2} \sum_{j_1, j_2 \notin \mathcal{I}_k} E\left[ \left( m_{j_2}/J_0 \right) (\delta_{n,j_1,i} )^\top H_{i} ( b_{n_0,j_2,i} ) \right] I\{j_1 \neq j_2\}\\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1 - \varphi_2-3/2} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{j}/J_0 \right) (\delta_{n,j,i} )^\top H_{i} ( b_{n_0,j,i} ) \right] \\
&\overset{(3)}{=} n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1 - \varphi_2-3/2} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{i}/J_0 \right) (\delta_{n,j,i} )^\top H_{i} ( b_{n_0,j,i} ) \right] \\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1 - \varphi_2-3/2} \sum_{j_1, j_2 \notin \mathcal{I}_k} E\left[ \left( m_{j_1}/J_0 \right) (\delta_{n,j_1,i} )^\top H_{i} ( \tilde{b}_{n_0,i} ) \right] I\{j_1 \neq j_2\}\\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_1 - \varphi_2-3/2} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{j}/J_0 \right) (\delta_{n,j,i} )^\top H_{i} ( b_{n_0,j,i} ) \right] \\
&\overset{(4)}{=} n_0^{-\varphi_1 - \varphi_2 + 1/2} E\left[ \left( m_{j_1}/J_0 \right) (\delta_{n,j_1,i} )^\top H_{i} ( \tilde{b}_{n_0,i} ) \right] \\
& - n_0^{-\varphi_1 - \varphi_2 - 1/2} E\left[ \left( m_{j_1}/J_0 \right) (\delta_{n,j_1,i} )^\top H_{i} ( \tilde{b}_{n_0,i} ) \right] \\
& + n_0^{-\varphi_1 -1/2} E\left[ \left( m_{j}/J_0 \right) (\delta_{n,j,i} )^\top H_{i} ( n_0^{- \varphi_2} b_{n_0,j,i} ) \right] \\
&\overset{(5)}{=} (G_b/2) n_0^{1/2-\varphi_1 - \varphi_2 } + o(n^{-\zeta}) ,
\end{align*}
where (1) holds by definition of $\Delta_{1,i}^l$ and $\Delta_{1,i}^b$, (2) holds since $i_1 \notin \{i_2, j_1,j_2\}$ implies
$$ E\left[ \left( m_{i_1}/J_0 \right) \left((\delta_{n,j_1,i_2} )^\top H_{i_2} ( b_{n_0,j_2,i_2} ) \right) \right]~$$
by using that $m_{i_1}$ is zero mean and independent of $ W_{i_2}$, $W_{j_1}$, and $ W_{j_2}$, (3) holds by definition of $\tilde{b}_{n_0,i} = E[b_{n_0,j,i} \mid X_i]$, and since $ j_1 \neq j_2$ and the law of iterative expectations implies
$$ \left( m_{\ell}/J_0 \right) E\left[ (\delta_{n,j_1,i} )^\top \mid W_{i}, W_{j_2} \right] H_{i_2} ( b_{n_0,j_2,i} ) = 0$$
for $\ell = i, j_2$ by using part (a) of Assumption (ref), (4) holds by the law of the iterative expectations,
$$ \left( m_{i}/J_0 \right) E\left[(\delta_{n,j,i} )^\top \mid W_i, X_j \right]H_{i}( b_{n_0,j,i} )= 0~, $$
and by parts (a) of Assumption (ref), and (5) holds by definition of $G_b$ in (ref) and Assumption (ref) and because
\begin{align*}
& n_0^{-\varphi_1 -1/2} E\left[ \left( m_{j}/J_0 \right) (\delta_{n,j,i} )^\top H_{i} ( n_0^{- \varphi_2} b_{n_0,j,i} ) \right] \\
& \le C n_0^{-\varphi_1 -1/2} E[|m_j/J_0|^2]^{1/2} E[||\delta_{n,j,i}||^4]^{1/4} E[||n_0^{-\varphi_2} b_{n_0,j,i} ||^4] \\
&= O(n^{-\varphi_1-1/2}) \times O(1) \times O(n_0^{1/4-\varphi_1/2}) \times o(n_0^{3/4-3\varphi_1/2}) \\
&= o(n^{1/2 -3 \varphi_1})
\end{align*}
where the inequality uses part (d) of Assumption (ref) and Cauchy-Schwartz inequality, and the equalities follows by part (b) of Assumption (ref). The proof of the claim is completed since $\zeta < 1/2 - 3\varphi_1$ whenever $\varphi_1 < 1/2$.\\
\textit{Claim 3:} $I_3 = o(n^{-\zeta})$. Consider the following derivations,
\begin{align*}
I_4
&\overset{(1)}{=} n^{-1 }\sum_{i_1 = 1}^n \sum_{k=1}^K \sum_{i_2 \in \mathcal{I}_k} n_0^{-2\varphi_2} n_0^{-2} \sum_{j_1 \notin \mathcal{I}_k} \sum_{j_2 \notin \mathcal{I}_k} E\left[ \left( m_{i_1}/J_0 \right) \left((b_{n_0,j_1,i_2} )^\top H_{i_2} ( b_{n_0,j_2,i_2} ) \right) \right] \\
&\overset{(2)}{=} n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2 \varphi_2-2} \sum_{j_1, j_2 \notin \mathcal{I}_k} E\left[ \left( m_{i}/J_0 \right) (b_{n_0,j_1,i} )^\top H_{i} ( b_{n_0,j_2,i} ) \right] \\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2 \varphi_2-2} \sum_{j_1, j_2 \notin \mathcal{I}_k} E\left[ \left( m_{j_1}/J_0 \right) (b_{n_0,j_1,i} )^\top H_{i} ( b_{n_0,j_2,i} ) \right] I\{j_1 \neq j_2\} \\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_2-2} \sum_{j_1, j_2 \notin \mathcal{I}_k} E\left[ \left( m_{j_2}/J_0 \right) (b_{n_0,j_1,i} )^\top H_{i} ( b_{n_0,j_2,i} ) \right] I\{j_1 \neq j_2\} \\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2\varphi_2-2} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{j}/J_0 \right) (b_{n_0,j,i} )^\top H_{i} ( b_{n_0,j,i} ) \right] \\
&\overset{(3)}{=} n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2 \varphi_2-2} \sum_{j_1, j_2 \notin \mathcal{I}_k} E\left[ \left( m_{i}/J_0 \right) (\tilde{b}_{n_0,i} )^\top H_{i} ( \tilde{b}_{n_0,i} ) \right] I\{j_1 \neq j_2\} \\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{ -2} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{i}/J_0 \right) (n_0^{-\varphi_2}b_{n_0,j,i} )^\top H_{i} ( n_0^{-\varphi_2}b_{n_0,j,i} ) \right] \\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{- \varphi_2-2} \sum_{j_1, j_2 \notin \mathcal{I}_k} E\left[ \left( m_{j_1}/J_0 \right) (n_0^{-\varphi_2} b_{n_0,j_1,i} )^\top H_{i} ( \tilde{b}_{n_0,i} ) \right] I\{j_1 \neq j_2\} \\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-\varphi_2-2} \sum_{j_1, j_2 \notin \mathcal{I}_k} E\left[ \left( m_{j_2}/J_0 \right) (\tilde{b}_{n_0,i} )^\top H_{i} ( n_0^{-\varphi_2} b_{n_0,j_2,i} ) \right] I\{j_1 \neq j_2\} \\
& + n^{-1 } \sum_{k=1}^K \sum_{i \in \mathcal{I}_k} n_0^{-2} \sum_{j \notin \mathcal{I}_k} E\left[ \left( m_{j}/J_0 \right) (n_0^{-\varphi_2} b_{n_0,j,i} )^\top H_{i} ( n_0^{-\varphi_2} b_{n_0,j,i} ) \right] \\
&\overset{(4)}{=} O(n^{-2\varphi_2}) + o(n^{1/2-3\varphi_1}) + o(n^{1/2-\varphi_1 - \varphi_2}) + o(n^{1/2-3\varphi_1})\\
&\overset{(5)}{=} o(n^{-\zeta}) ,
\end{align*}
where (1) holds by definition of $\Delta_{1,i}^b$, (2) holds since $i_1 \notin \{j_1,j_2,i_2\}$ implies
$$E\left[ \left( m_{i_1}/J_0 \right) \left((b_{n,j_1,i_2} )^\top (\partial_\eta^2 m_{i_2}/(2 J_0)) ( b_{n,j_2,i_2} ) \right) \right] = 0$$
by using that $m_{i_1}$ is zero mean and independent of $ W_{i_2}$, $W_{j_1}$, and $ W_{j_2}$; (3) holds by the law of iterative expectations; (4) holds by parts (b) and (c) of Assumption (ref), parts (c) and (e) of Assumption (ref), part (b.1) of Assumption (ref), and Holder's inequality; and (5) holds since $\zeta < 1/2 - 3\varphi_1$ and $\zeta < 2\varphi_2$ (because $\varphi_1 < 1/2$ and $\varphi_1 \le \varphi_2$).
\end{proof}
\subsection{Proof of Lemma (ref)}
\begin{proof} It is sufficient to show the result for the case when $\delta_{n_0}$ and $b_{n_0}$ are real-valued functions since for any $x = (x_1,\ldots,x_p) \in \mathbf{R}^p$ it holds
$|| x ||^4 = \left(\sum_{\ell=1}^p x_\ell^2 \right)^2 \le p \sum_{\ell=1}^p |x_\ell|^4 $.
In the proof, I use that $E[ (\sum_{\ell \notin \mathcal{I}_k} Z_\ell)^4] \le n_0 E[Z_\ell^4] + 3 n_0^2 E[Z_\ell^2]^2$, which holds for zero mean i.i.d. random variables $Z_\ell$.
\textbf{Part 1:} Fix $i \in \mathcal{I}_k$ and denote $Z_\ell = \delta_{n_0}(W_\ell,X_i)$ for any $ \ell \notin \mathcal{I}_k$. Conditional on $X_i$, it holds that $ \{Z_\ell: \ell \notin \mathcal{I}_k\}$ is a zero mean i.i.d. sequence of random variables due to part (a) of Assumption (ref). Therefore,
\begin{align*}
E\left[ | n_0^{-1/2} \sum_{\ell \notin \mathcal{I}_k} n_0^{-\varphi_1} Z_\ell |^4 \mid X_i \right] &\le n_0^{-2-4\varphi_1} \left( n_0 E[Z_\ell^4 \mid X_i] + 3 n_0^2 E[Z_\ell^2 \mid X_i ]^2 \right)
\end{align*}
Using the previous inequality and the law of iterative expectations, it follows
\begin{align*}
E\left[ | n_0^{-1/2} \sum_{\ell \notin \mathcal{I}_k} n_0^{-\varphi_1} Z_\ell |^4 \right] &\le n_0^{ -4\varphi_1} \left( n_0^{-1} E[Z_\ell^4 ] + 3 E[ E[Z_\ell^2 \mid X_i ]^2 ] \right) \\
&\overset{(1)}{\le} n_0^{ -4\varphi_1} \left( n_0^{-2\varphi_1} M_1 + 3 M_1 \right) ,
\end{align*}
where (1) holds by parts (b.1) and (b.2) of Assumption (ref), and the definition of $Z_\ell$. Taking $C \ge 4M_1$ completes the proof of part 1.
\textbf{Part 2:} Fix $i \in \mathcal{I}_k$ and denote $Z_\ell = n_0^{-\varphi_2} (b_{n_0}(X_\ell,X_i) - \tilde{b}_{n_0}(X_i))$ for any $ \ell \notin \mathcal{I}_k$, where $\tilde{b}_{n_0}(X_i) = E[b_{n_0}(X_\ell,X_i) \mid X_i] $. As in part 1, $\{Z_\ell: \ell \notin \mathcal{I}\}$ conditional on $X_i$ are zero mean i.i.d. random variables. Therefore,
\begin{align*}
E\left[ | n_0^{-1} \sum_{\ell \notin \mathcal{I}_k } (Z_\ell + n_0^{-\varphi_2}\tilde{b}_{n_0}(X_i) ) |^4 \right] &\overset{(1)}{\le} 2^3 E\left[ | n_0^{-1} \sum_{\ell \notin \mathcal{I}_k } Z_\ell |^4 \right] + 2^3 E\left[ | n_0^{-1} \sum_{\ell \notin \mathcal{I}_k } n_0^{-\varphi_2} \tilde{b}_{n_0}(X_i) |^4 \right] \\
&\overset{(2)}{\le} 8n_0^{ -2} \left( n_0^{-1} E[Z_\ell^4 ] + 3 E[ E[Z_\ell^2 \mid X_i ]^2 ] \right) + 8 n_0^{-4\varphi_2} E[|\tilde{b}_{n_0}(X_i)|^4] \\
&\overset{(3)}{\le} 8 n_0^{-6 \varphi_1} \tau_{n_0} + 24 n_0^{-4\varphi_1} \tau_{n_0} + 8 n_0^{-4\varphi_2} M_1
\end{align*}
where (1) holds by Loeve's inequality (Davidson1994), (2) holds by the same arguments as in part 1, and (3) holds by part (b.1), (b.3) and (b.4) of Assumption (ref). Taking $C \ge 32 + 8 M_1 $ completes the proof of part 2.
\end{proof}
\subsection{Proof of Lemma (ref)}
\begin{proof}
For $i \in \mathcal{I}_k$, denote $\Delta_{i} = \Delta_{i}^b + \Delta_{i}^l$, where
\begin{align*}
\Delta_{i}^l &= n_0^{-\varphi_1} n_0^{-1/2} \sum_{j \notin \mathcal{I}_k} \delta_{n_0,j,i} ,\\
\Delta_{i}^b &= n_0^{-\varphi_2} n_0^{-1} \sum_{j \notin \mathcal{I}_k} b_{n_0,j,i} ,
\end{align*}
Here, $\delta_{n_0,j,i} = \delta_{n_0}(W_{j},X_i)$ and $b_{n_0,j,i} = b_{n_0}(X_{j},X_i)$, and $\delta_{n_0}$ and $b_{n_0}$ are functions satisfying Assumption (ref). In what follows, the results are proved for any given sequence $K$ that diverges to infinity as $n$ diverges to infinity, which is sufficient to guarantee the results of this lemma.
\textbf{Part 1:} Using Assumption (ref), it follows
\begin{equation*}
n^{-1} \sum_{i=1}^n \left( \hat{\eta}_i - \eta_i \right)^\top \partial_\eta \psi^z_i = I_1 + I_2 + I_3 ,
\end{equation*}
where
\begin{align*}
I_1 &= n^{-1} \sum_{i=1}^n (\Delta_i^l)^{\top} \partial_\eta \psi^z_i \\
I_2 &= n^{-1} \sum_{i=1}^n (\Delta_i^b)^{\top} \partial_\eta \psi^z_i \\
I_3 &= n^{-1} n_0^{-2 \min\{\varphi_1, \varphi_2\} } \sum_{i=1}^n \hat{R}_1(X_i)^{\top} \partial_\eta \psi^z_i
\end{align*}
and $n_0 = ((K-1)/K) n$.
\textit{Claim 1:} $I_1 = O_p(n^{-1/2-\min\{ \varphi_1,\varphi_2 \} })$. I first show that $E[I_1] = 0$ (claim 1.1). I then show that $E[I_1^2] = O(n^{-1-2\min\{ \varphi_1,\varphi_2 \} })$ (claim 1.2), which is sufficient to conclude the claim. \\
\textit{Claim 1.1:} $E[I_1] = 0$. Consider the following derivations,
\begin{align*}
E[I_1] &= n^{-1} \sum_{i=1}^n E[(\Delta_i^l)^{\top} \partial_\eta \psi^z_i] \\
&= n^{-1} \sum_{k=1}^K \sum_{ i \in \mathcal{I}_k} E\left[ E\left[(\Delta_i^l)^{\top} \partial_\eta \psi^z_i \mid X_i, (W_j : j \notin \mathcal{I}_k)\right] \right] \\
&\overset{(1)}{=} n^{-1} \sum_{k=1}^K \sum_{ i \in \mathcal{I}_k} E\left[ (\Delta_i^l)^{\top} E\left[ \partial_\eta \psi^z_i \mid X_i, (W_j : j \notin \mathcal{I}_k)\right] \right] \\
&\overset{(2)}{=} 0 ,
\end{align*}
where (1) holds since $\Delta_i^l=\Delta_1^l(X_i)$ is a function of $X_i$ and the data $(W_j : j \notin \mathcal{I}_k)$ used to estimate $\hat{\eta}_k(\cdot)$, and (2) holds by part (b) in Assumption (ref).
\textit{Claim 1.2:} $E[I_1^2] = O(n^{-1-2\min\{ \varphi_1,\varphi_2 \} })$. Recall that I use the following notation $\delta_{n_0,j,i} = \delta_{n_0}(W_j,X_i)$ and $\Delta_i^l = \Delta_1^l(X_i) = n_0^{-\varphi_1} n_0^{-1/2} \sum_{ j \notin \mathcal{I}_k} \delta_{n_0,j,i}$ for $i \in \mathcal{I}_k$. To show $E[I_1^2]= O(n^{-1-2\min\{ \varphi_1,\varphi_2 \} })$, consider the following derivations,
\begin{align*}
E[I_1^2] &\overset{(1)}{=} n^{-2} \sum_{i_1,i_2=1}^n E\left[ (\Delta_{i_1}^l)^{\top} \partial_\eta \psi^z_{i_1} (\Delta_{i_2}^l)^{\top} \partial_\eta \psi^z_{i_2} \right] \\
&\overset{(2)}{\le} n_0^{-2\varphi_1} n^{-2} \sum_{i=1}^n E\left[ ( (\Delta_{i}^l)^{\top} \partial_\eta \psi^z_{i})^2 \right]
+ n_0^{-2\varphi_1} n^{-2} \sum_{i_1 \neq i_2}^n \left| E\left[ (\Delta_{i_1}^l)^{\top} \partial_\eta \psi^z_{i_1} (\Delta_{i_2}^l)^{\top} \partial_\eta \psi^z_{i_2} \right] \right| \\
& \overset{(3)}{\le} n_0^{-2\varphi_1-1} E\left[ \left( \delta_{n_0,j,i}^{\top} \partial_\eta \psi^z_{i} \right)^2 \right]
+ \frac{(n-1)n^{-1}}{n_0^{1+2\varphi_1}} \left| E\left[ \delta_{n_0,i_2,i_1}^{\top} \partial_\eta \psi^z_{i_1} \delta_{n_0,i_1,i_2}^{\top} \partial_\eta \psi^z_{i_2} \right] \right| \\
& \overset{(4)}{\le} n_0^{-2\varphi_1-1} E\left[ \left( \delta_{n_0,j,i}^{\top} \partial_\eta \psi^z_{i} \right)^2 \right]
+ (n-1)n^{-1} n_0^{-(1+2\varphi_1)} E\left[ \left( \delta_{n_0,i_2,i_1}^{\top} \partial_\eta \psi^z_{i_1} \right)^2 \right] ^{1/2} E\left[ \left( \delta_{n_0,i_1,i_2}^{\top} \partial_\eta \psi^z_{i_2} \right)^2 \right] ^{1/2} \\
& = n_0^{-2\varphi_1-1} E\left[ \left( \delta_{n_0,j,i}^{\top} \partial_\eta \psi^z_{i} \right)^2 \right] + (n-1)n^{-1} n_0^{-(1+2\varphi_1)} E\left[ \left( \delta_{n_0,j,i}^{\top} \partial_\eta \psi^z_{i} \right)^2 \right] \\
& \overset{(5)}{\le} n_0^{-2\varphi_1-1} M_1^{1/2} C_1 \times p \left( 1 + (n-1) n^{-1 } \right) ,
\end{align*}
where (1) holds by definition of $I_1$, (2) holds by triangular inequality, (3) holds by (ref) and (ref) presented below, (4) holds by Cauchy-Schwartz inequality, and (5) holds by the derivations presented next,
\begin{align*}
E\left[ \left( \delta_{n_0,j,i}^{\top} \partial_\eta \psi^z_{i} \right)^2 \right] &= E\left[ \delta_{n_0,j,i}^{\top} E\left[ \left( \partial_\eta \psi^z_{i} (\partial_\eta \psi^z_{i})^\top \right) \mid X_i, W_j \right] \delta_{n_0,j,i}\right] \\
&\overset{(1)}{=} E\left[ \delta_{n_0,j,i}^{\top} E\left[ \left( \partial_\eta \psi^z_{i} (\partial_\eta \psi^z_{i})^\top \right) \mid X_i \right] \delta_{n_0,j,i} \right] \\
&\overset{(2)}{\le} E\left[ ||\delta_{n_0,j,i}||^2 \right] C_1 \times p \\
&\overset{(3)}{\le} M_1^{1/2} C_1 \times p
\end{align*}
where (1) holds since $i \neq j$, (2) holds by part (d) of Assumption (ref) and Loeve’s inequality (Davidson1994), and (3) holds by Jensen's inequality (e.g., $ E\left[ ||\delta(W_{j},X_{i})||^2 \right]^{1/2} \le E\left[ ||\delta(W_{j},X_{i})||^4 \right]^{1/4}$) and part (b) of Assumption (ref). Note these derivations complete the proof of claim 1.2.
The previous derivations used the following claims:
\begin{align}
E\left[ ((\Delta_{i}^l)^{\top} \partial_\eta \psi^z_{i})^2 \right] &= n_0^{-1} \sum_{j \notin \mathcal{I}_{k}} E\left[ \delta(W_{j},X_{i})^{\top} \partial_\eta \psi^z_{i} \delta(W_{j},X_{i})^{\top} \partial_\eta \psi^z_{i} \right]\\
E\left[ (\Delta_{i_1}^l)^{\top} \partial_\eta \psi^z_{i_1} (\Delta_{i_2}^l)^{\top} \partial_\eta \psi^z_{i_2} \right] &= n_0^{-1} E\left[ \delta_{n_0,i_2,i_1}^{\top} \partial_\eta \psi^z_{i_1} \delta_{n_0,i_1,i_2}^{\top} \partial_\eta \psi^z_{i_2} \right] I\{k_1 \neq k_2\}
\end{align}
To show (ref), consider the following derivations.
\begin{align*}
E\left[ (\Delta_{i}^l)^{\top} \partial_\eta \psi^z_{i} (\Delta_{i}^l)^{\top} \partial_\eta \psi^z_{i} \right] &= n_0^{-1} \sum_{j_1 \notin \mathcal{I}_{k}} \sum_{j_2 \notin \mathcal{I}_{k}} E\left[ \delta_{n_0,j_1,i}^{\top} \partial_\eta \psi^z_{i} \delta_{n_0,j_2,i}^{\top} \partial_\eta \psi^z_{i} \right]\\
&\overset{(1)}{=} n_0^{-1} \sum_{j \notin \mathcal{I}_{k}} E\left[ \delta(W_{j},X_{i})^{\top} \partial_\eta \psi^z_{i} \delta(W_{j},X_{i})^{\top} \partial_\eta \psi^z_{i} \right]
\end{align*}
where (1) holds due to the following: if $j_1 \neq j_2$, then
\begin{align*}
E\left[ \delta_{n_0,j_1,i}^{\top} \partial_\eta \psi^z_{i} \delta_{n_0,j_1,i}^{\top} \partial_\eta \psi^z_{i} \right] &= E\left[ E\left[ \delta_{n_0,j_1,i}^{\top} \mid W_i, W_{j_2} \right] \partial_\eta \psi^z_{i} \delta_{n_0,j_2,i}^{\top} \partial_\eta \psi^z_{i} \right] \\
&= E\left[ E\left[ \delta_{n_0,j_1,i}^{\top} \mid X_i \right] \partial_\eta \psi^z_{i} \delta_{n_0,j_2,i}^{\top} \partial_\eta \psi^z_{i} \right] \\
&\overset{(1)}{=} 0 ,
\end{align*}
where (1) holds by definition of $\delta_{n_0}$ in part (a) of Assumption (ref).
To show (ref), consider $i_1 \neq i_2$ where $i_1 \in \mathcal{I}_{k_1}$ and $i_2 \in \mathcal{I}_{k_2}$, therefore
\begin{align*}
E\left[ (\Delta_{i_1}^l)^{\top} \partial_\eta \psi^z_{i_1} (\Delta_{i_2}^l)^{\top} \partial_\eta \psi^z_{i_2} \right] &= n_0^{-1} \sum_{j_1 \notin \mathcal{I}_{k_1}} \sum_{j_2 \notin \mathcal{I}_{k_2}} E\left[ \delta_{n_0,j_1,i_1}^{\top} \partial_\eta \psi^z_{i_1} \delta_{n_0,j_2,i_2}^{\top} \partial_\eta \psi^z_{i_2} \right] \\
&\overset{(1)}{=} n_0^{-1} E\left[ \delta_{n_0,i_2,i_1}^{\top} \partial_\eta \psi^z_{i_1} \delta_{n_0,i_1,i_2}^{\top} \partial_\eta \psi^z_{i_2} \right] I\{k_1 \neq k_2\}
\end{align*}
where (1) holds since $k_1 = k_2$ implies $j_2 \neq i_1$ and $j_1 \neq i_2$, and because the conditions $j_2 \neq i_1$ or $j_1 \neq i_2$ imply that $E\left[ \delta_{n_0,j_1,i_1}^{\top} \partial_\eta \psi^z_{i_1} \delta_{n_0,j_2,i_2}^{\top} \partial_\eta \psi^z_{i_2} \right]$ is zero. To see this, suppose $j_2 \neq i_1$ and consider the following derivations,
\begin{align*}
E\left[ \delta_{n_0,j_1,i_1}^{\top} \partial_\eta \psi^z_{i_1} \delta_{n_0,j_2,i_2}^{\top} \partial_\eta \psi^z_{i_2} \right] &= E\left[ E\left[ \delta_{n_0,j_1,i_1}^{\top} \partial_\eta \psi^z_{i_1} \delta_{n_0,j_2,i_2}^{\top} \partial_\eta \psi^z_{i_2} \mid X_{i_1}, W_{i_2}, W_{j_1}, W_{j_2} \right] \right] \\
&= E\left[ \delta_{n_0,j_1,i_1}^{\top} E\left[ \partial_\eta \psi^z_{i_1} \mid X_{i_1}, W_{i_2}, W_{j_1}, W_{j_2} \right] \delta_{n_0,j_2,i_2}^{\top} \partial_\eta \psi^z_{i_2} \right] \\
&\overset{(1)}{=} E\left[ \delta_{n_0,j_1,i_1}^{\top} E\left[ \partial_\eta \psi^z_{i_1} \mid X_{i_1} \right] \delta_{n_0,j_2,i_2}^{\top} \partial_\eta \psi^z_{i_2} \right] \\
&\overset{(2)}{=} 0 ,
\end{align*}
where (1) holds since $i_1 \notin \{i_2, j_1, j_2\}$ (since $i_1 \neq j_2$) and (2) holds by part (b) in Assumption (ref). Similar derivations conclude the same for $j_1 \neq i_2$.\\
\textit{Claim 2:} $I_2 = O(n^{-1/2-\min\{ \varphi_1,\varphi_2 \} })$. Define $X^{(n)} = \{X_i : 1 \le i \le n \}$. I first show $E[I_2 \mid X^{(n)}] =0$. I then show $E[I_2^2 ] \le n^{-1} E[ ||\Delta_i^b||^2] C_1 p $, which is sufficient to conclude due to Lemma (ref) that implies that $ E[ ||\Delta_i^b||^2] $ is $O(n^{-2\min\{ \varphi_1,\varphi_2 \} })$ due to Cauchy-Schwartz.
The first part holds due to the following derivations,
\begin{align*}
E[I_2 \mid X^{(n)}] &= E\left[ n^{-1} \sum_{i=1}^n (\Delta_i^b)^{\top} \partial_\eta \psi^z_i \mid X^{(n)} \right] \\
&\overset{(1)}{=} n^{-1} \sum_{i=1}^n (\Delta_i^b)^{\top} E\left[ \partial_\eta \psi^z_i \mid X^{(n)} \right] \\
&\overset{(2)}{=} n^{-1} \sum_{i=1}^n (\Delta_i^b)^{\top} E\left[ \partial_\eta \psi^z_i \mid X_i \right] \\
&\overset{(3)}{=} 0 ,
\end{align*}
where (1) holds since $\Delta_i^b$ is function of $X^{(n)}$ and $\Delta_i^b = n_0^{-\varphi_2} n_0^{-1} \sum_{i_0 \notin \mathcal{I}_k} b(X_{i_0},X_i)$ for $i \in \mathcal{I}_k$ due to part (a) of Assumption (ref), (2) holds since the observations are i.i.d., and (3) follows due to part (b) of Assumption (ref).
To prove that $E[I_2^2 ] \le n^{-1} E[ ||\Delta_i^b||^2] C_1 p $, first note that
\begin{align*}
E[I_2^2 \mid X^{(n)}] &= E\left[ \left( n^{-1} \sum_{i=1}^n (\Delta_i^b)^{\top} \partial_\eta \psi^z_i \right)^2 \mid X^{(n)} \right] \\
&\overset{(1)}{=} E\left[ n^{-2} \sum_{i=1}^n \left( (\Delta_i^b)^{\top} \partial_\eta \psi^z_i \right)^2 \mid X^{(n)} \right] \\
&\overset{(2)}{=} n^{-2} \sum_{i=1}^n (\Delta_i^b)^{\top} E\left[ (\partial_\eta \psi^z_i) (\partial_\eta \psi^z_i)^{\top} \mid X_i \right] \Delta_i^b \\
&\overset{(1)}{\le} n^{-2} \sum_{i=1}^n ||\Delta_i^b||^2 C_1 \times p
\end{align*}
where (1) holds because $ E\left[ \left( (\Delta_i^b)^{\top} \partial_\eta \psi^z_i \right) \left( (\Delta_j^b)^{\top} \partial_\eta \psi^z_j \right) \mid X^{(n)} \right] = 0$ when $ i \neq j$ (since $\Delta_i^b$ and $\Delta_j^b$ are functions of $X^{(n)}$, and part (b) of Assumption (ref)), (2) holds since $\Delta_i^b$ and $\Delta_j^b$ are functions of $X^{(n)}$ and the observations are i.i.d., and (3) holds by part (d) of Assumption (ref). Then,
$$ E[ E[I_2^2 \mid X^{(n)}] ] \le E[ n^{-2} \sum_{i=1}^n ||\Delta_i^b||^2 C_1 \times p ] = n^{-1} E[ ||\Delta_i^b||^2] C_1 p $$
which completes the proof of this claim.\\
\textit{Claim 3:} $I_3 = O_p(n_0^{-2\min\{\varphi_1, \varphi_2\}})$. Algebra shows
\begin{align*}
|I_3| &= | n_0^{-2\min\{\varphi_1, \varphi_2\} } n^{-1} \sum_{i=1}^n \hat{R}_1(X_i)^{\top} \partial_\eta \psi^z_i | \\
&\le n_0^{-2\min\{\varphi_1, \varphi_2\} } \left( n^{-1} \sum_{i=1}^n ||\hat{R}_1(X_i)||^2 \right)^{1/2} \left( n^{-1} \sum_{i=1}^n || \partial_\eta \psi^z_i||^2 \right)^{1/2} \\
&\overset{(1)}{=} n_0^{-2\min\{\varphi_1, \varphi_2\} } \times O_p(1) \times \left( n^{-1} \sum_{i=1}^n || \partial_\eta \psi^z_i||^2 \right)^{1/2} ,\\
&\overset{(2)}{=} n_0^{-2\min\{\varphi_1, \varphi_2\}} \times O_p(1) \times O_p(1) ,\\
&\overset{(3)}{=} O_p(n_0^{-2\min\{\varphi_1, \varphi_2\}}) ,
\end{align*}
where (1) holds by part (c) of Assumption (ref), (2) holds by the law of large numbers, Jensen's inequality (e.g., ($E[|| \partial_\eta \psi^z_i||^2] \le E[|| \partial_\eta \psi^z_i||^4]^{1/2}$), and part (c) of Assumption (ref), and (3) holds since $n/2 \le n \le n$
\textbf{Part 2:} By Taylor approximation and mean-value theorem (since $\psi^z(w,\eta)$ is twice continuously differentiable on $\eta$ by Assumption (ref)), it follows
$$ \hat{\psi}^z_i - \psi^z_i = (\hat{\eta}_i - \eta_i)^\top \partial_\eta \psi^z_i + \frac{1}{2} (\hat{\eta}_i - \eta_i)^\top \partial_\eta^2 \tilde{\psi}^z_i (\hat{\eta}_i - \eta_i)$$
where $ \partial_\eta^2 \tilde{\psi}^z_i = \partial_\eta^2 \psi^z(W_i, \eta)|_{\eta = \tilde{\eta}_i}$ for some $\hat{\eta}_i$ (due to mean-value theorem). Using this
\begin{align*}
n^{-1}\sum_{i=1}^n (\hat{\psi}^z_i - \psi^z_i ) &= n^{-1}\sum_{i=1}^n (\hat{\eta}_i - \eta_i)^\top \partial_\eta \psi^z_i + \frac{1}{2} n^{-1}\sum_{i=1}^n (\hat{\eta}_i - \eta_i)^\top \partial_\eta^2 \tilde{\psi}^z_i (\hat{\eta}_i - \eta_i) \\
&\overset{(1)}{=} O_p(n^{-\min\{\varphi_1, \varphi_2\} -1/2}) + \frac{1}{2} n^{-1}\sum_{i=1}^n (\hat{\eta}_i - \eta_i)^\top \partial_\eta^2 \tilde{\psi}^z_i (\hat{\eta}_i - \eta_i) \\
&\overset{(2)}{=} O_p(n^{-2\min\{\varphi_1, \varphi_2\} }) ,
\end{align*}
where (1) holds due to Part 1, and (2) holds due to the derivations presented next,
\begin{align*}
| n^{-1}\sum_{i=1}^n (\hat{\eta}_i - \eta_i)^\top \partial_\eta^2 \tilde{m}_i (\hat{\eta}_i - \eta_i) | &\le n^{-1}\sum_{i=1}^n |(\hat{\eta}_i - \eta_i)^\top \partial_\eta^2 \tilde{m}_i (\hat{\eta}_i - \eta_i)| \\
&\overset{(1)}{\le} n^{-1}\sum_{i=1}^n ||\hat{\eta}_i - \eta_i||^2 C_2 \times p ,\\
&\overset{(2)}{=} O_p(n^{-2\min\{\varphi_1, \varphi_2\} }) ,
\end{align*}
where (1) holds due to part (e) of Assumption (ref) and Loeve’s inequality (Davidson1994), and (2) holds due to part 4 of Lemma (ref). \\
\textbf{Part 3:} It follows from part 1, by using that $\partial_\eta m_i = \partial_\eta \psi^b_i - \partial_\eta \psi^a_i \theta_0 $ and $|\theta_0| \le M_1^{1/4}/C_0$ (due to parts (a) and (c) of Assumptions (ref) and the representation of $\theta_0$ as a ratio of expected values in (ref)). \\
\textbf{Part 4:} By Taylor expansion and mean value theorem,
$$ \hat{m}_i - m_i = (\hat{\eta}_i - \eta_i)^\top \partial_\eta m_i + (\hat{\eta}_i - \eta_i)^\top (\partial_\eta^2 m_i/2) (\hat{\eta}_i - \eta_i) + \tilde{r}_i~, $$
where $\tilde{r}_i $ is the Lagrange's remainder error term (since $m$ is three-times continuous differentiable on $\eta$ by assumption on $\psi^z$). Therefore,
\begin{equation}
|\tilde{r}_i| \le (1/6) p^{3/2} C_3 ||\hat{\eta}_i - \eta_i||^3 ,
\end{equation}
where the bound follows by part (e) of Assumption (ref), Jensen's inequality, and the definition of Euclidean norm. It follows
$$ n^{-1/2} \sum_{i=1}^n (\hat{m}_i - m_i)/J_0 = I_1 + I_2 + I_3 ~,$$
where
\begin{align*}
I_1 &= n^{-1/2} \sum_{i=1}^n (\hat{\eta}_i - \eta_i)^\top \partial_\eta m_i / J_0 \\
I_2 &= n^{-1/2} \sum_{i=1}^n (\hat{\eta}_i - \eta_i)^\top (\partial_\eta^2 m_i/(2J_0)) (\hat{\eta}_i - \eta_i) \\
I_3 &= n^{-1/2} \sum_{i=1}^n \tilde{r}_i/J_0
\end{align*}
In the claims below I show that $I_1 = \mathcal{T}_{n,K}^{l} + o_p(n^{-\zeta})$, $I_2 = \mathcal{T}_{n,K}^{nl} + o_p(n^{-\zeta})$, and $I_3 = o_p(n^{-\zeta})$, which is sufficient to complete the proof of part 4. Furthermore, if Assumption (ref) holds, then Proposition (ref) implies $\lim_{n \to \infty} \inf_{K \le n } Var[n^{2\varphi_1 - 1} \mathcal{T}_{n,K}^{nl}] > 0$; and if Assumption (ref) holds, then Proposition (ref) implies $\lim_{n \to \infty} \inf_{K \le n } Var[n^{\varphi_1} \mathcal{T}_{n,K}^{l}] > 0$.\\
\textit{Claim 1:} $I_1 = \mathcal{T}_{n,K}^{l} + O_p(n^{-2\min\{\varphi_1,\varphi_2\} })$. By part (a) of Assumption (ref), it follows
$$ I_1 = I_{1,1} + I_{1,2}~,$$
where $\hat{R}_{i} = \hat{R}(X_i)$ for $i \in \mathcal{I}_k$ and
\begin{align*}
I_{1,1} &= n^{-1/2} \sum_{i=1}^n (\Delta_{i})^\top \partial_\eta m_i / J_0 \\
I_{1,2} &= n^{-1/2} \sum_{i=1}^n ( n_0^{-2\varphi_1} \hat{R}_{i} )^\top \partial_\eta m_i / J_0
\end{align*}
By definition of $\mathcal{T}_{n,K}^{l}$ in (ref), it follows that $I_{1,1} = \mathcal{T}_{n,K}^{l}$. Since $\partial_\eta m_i = \partial_\eta \psi^b_i - \theta_0 \partial_\eta \psi^a_i $ and $|\theta_0| \le M^{1/4}/C_0$, it follows that $I_{1,2}$ is $O_p(n^{-2\min\{\varphi_1,\varphi_2\} })$ due to proof of Claim 3 in Part 1 of this lemma. \\
\textit{Claim 2:} $I_2 = \mathcal{T}_{n,K}^{nl} + O_p(n^{1/2-3\min\{\varphi_1, \varphi_2\}})$. By part (a) of Assumption (ref), it follows
$$ I_2 = I_{2,1} + 2 I_{2,2} + I_{2,3} $$
where $\hat{R}_{i} = \hat{R}(X_i)$ for $i \in \mathcal{I}_k$ and
\begin{align*}
I_{2,1} &= n^{-1/2} \sum_{i=1}^n (\Delta_{i})^\top (\partial_\eta^2 m_i/(2J_0)) ( \Delta_{i} ) \\
I_{2,2} &= n^{-1/2} \sum_{i=1}^n (\Delta_{i})^\top (\partial_\eta^2 m_i/(2J_0)) ( n_0^{-2\varphi_1}\hat{R}_{i} ) \\
I_{2,3} &= n^{-1/2} \sum_{i=1}^n ( n_0^{-2\varphi_1}\hat{R}_{i} ) ^\top (\partial_\eta^2 m_i/(2J_0)) ( n_0^{-2\varphi_1}\hat{R}_{i} )
\end{align*}
By definition of $\mathcal{T}_{n,K}^{nl}$ in (ref), it follows that $I_{2,1} = \mathcal{T}_{n,K}^{nl}$. In what follows, I prove claims that imply $I_{2,j} = o_p(n^{-\zeta})$ for $j=2,3$ using that $\zeta < 3\varphi_1 - 1/2$ since $\varphi_1 < 1/2$, which is sufficient to complete the proof of claim 2. \\
\textit{Claim 2.1:} $I_{2,2} = O_p(n^{1/2-3\min\{\varphi_1, \varphi_2\}})$. To see this, consider the following derivations,
\begin{align*}
|I_{2,2}| &\overset{(1)}{\le} n^{1/2} \times p C_2 \left(n^{-1} \sum_{i=1}^n || \Delta_{i}|| \times || n_0^{-2\min\{\varphi_1, \varphi_2\}}\hat{R}_{i}|| \right) \\
&\overset{(2)}{\le} n^{1/2} n_0^{-2\min\{\varphi_1, \varphi_2\}} \times p C_2 \left(n^{-1} \sum_{i=1}^n || \Delta_{i}||^2 \right)^{1/2} \times \left(n^{-1} \sum_{i=1}^n ||\hat{R}_{i}||^2 \right)^{1/2}\\
&\overset{(3)}{=} n^{1/2} n_0^{-2\min\{\varphi_1, \varphi_2\}} \times O_p(n^{-\min\{\varphi_1, \varphi_2\}}) \times O_p(1) \\
&= O_p(n^{1/2 - 3\min\{\varphi_1, \varphi_2\}})
\end{align*}
where (1) holds by triangle inequality, part (e) of Assumption (ref), Jensen's inequality, and definition of Euclidean norm, (2) holds by Cauchy-Schwartz inequality, (3) holds by Lemma (ref) and by part (c) of Assumption (ref) and Markov's inequality.
\textit{Claim 2.2:} $I_{2,3} = O_p(n^{1/2-4\min\{\varphi_1, \varphi_2\}})$. The proof is similar to Claim 2.1; therefore, it is omitted. \\
\textit{Claim 3:} $I_3 = O_p(n^{1/2-3\min\{\varphi_1, \varphi_2\}})$. Using (ref), it follows
$$ |I_3| \le (1/6) p^{3/2} C_3/J_0 n^{-1/2} \sum_{i=1}^n || \hat{\eta}_i - \eta_i||^3~.$$
In what follows, I prove that $n^{-1/2} \sum_{i=1}^n || \hat{\eta}_i - \eta_i||^3$ is $O_p(n^{1/2-3\varphi_1})$.
By part (a) of Assumption (ref) and since $\varphi_1 \le \varphi_2$, it follows
$$\hat{\eta}_i - \eta_i = \Delta_{i} + n_0^{-2\min\{\varphi_1, \varphi_2\}} \hat{R}_{i} $$
where $\Delta_{i} = \Delta_{i}^l + \Delta_{i}^b$ and $\hat{R}_{i} = \hat{R}(X_i)$. Using triangle inequality and Loeve’s inequality (Davidson1994) in the previous expression, it follows
$$ || \hat{\eta}_i - \eta_i ||^3 \le 2^2 \left ( || \Delta_{i} ||^3 + n_0^{-6 \min\{\varphi_1, \varphi_2\}}|| \hat{R}_{i} ||^3 \right) $$
which implies
$$ n^{-1} \sum_{i=1}^n ||\hat{\eta}_i -\eta_i||^3 \le 2^2(I_{3,1} + I_{3,2} )$$
where
\begin{align*}
I_{3,1} &= n^{-1} \sum_{i=1}^n || \Delta_{i} ||^3\\
I_{3,2} &= n^{-1} \sum_{i=1}^n n_0^{-6 \varphi_1}|| \hat{R}_{i} ||^3
\end{align*}
To complete the proof of claim 3, it is sufficient to show $I_{3,1} = O_p(n^{-3\varphi_1})$ and $I_{3,2} = O_p(n^{1/2-6\varphi_1})$, since they imply $n^{-1/2} \sum_{i=1}^n || \hat{\eta}_i - \eta_i||^3$ is $O_p(n^{1/2 - 3\varphi_1})$.\\
\textit{Claim 3.1:} $I_{3,1} = O_p(n^{-3\min\{\varphi_1, \varphi_2\}})$. The proof is a direct result of Lemma (ref) and Markov's inequality; therefore, it is omitted.\\
\textit{Claim 3.2:} $I_{3,2} = O_p(n^{1/2-6\min\{\varphi_1, \varphi_2\}})$. Consider the following derivations,
\begin{align*}
I_{3,2} &\overset{(1)}{\le} n^{1/2} n^{-6\min\{\varphi_1, \varphi_2\}} \left( n^{-1} \sum_{i=1}^n ||\hat{R}_{i}||^2 \right)^{3/2} \\
&\overset{(2)}{=} n^{1/2} \times n^{-6\min\{\varphi_1, \varphi_2\}} \times O_p(1)
\end{align*}
where (1) holds by Loeve’s inequality (Davidson1994), and (2) holds by part (c) of Assumption (ref) with Markov's inequality. This completes the proof of claim 3.3.
Claim 1 and Claim 2 in the proof of Part 1 imply that $\mathcal{T}_{n,K}^{l}$ is $O_p(n^{-\min\{\varphi_1,\varphi_2\}})$. By the same argument used in the proof of Part 2 to bound the non-linear expression (but using Lemma (ref) instead of part 4 of Lemma (ref)), it follows that $\mathcal{T}_{n,K}^{nl}$ is $O_p(n^{1/2-2\min\{\varphi_1,\varphi_2\}})$.
\end{proof}
\subsection{Proof of Lemma (ref)}
\begin{proof}
In what follows, the results are proved for any given sequence $K$ that diverges to infinity as $n$ diverges to infinity, which is sufficient to guarantee the result of this lemma.
\textbf{Part 1:} The proof of (ref) has two steps. The first step shows
\begin{equation}
E\left[ \left ( n_k^{-1/2} \sum_{i \in \mathcal{I}_k} (\hat{\eta}_i -\eta_i)^\top \partial_\eta \psi^z_i \right)^2 \right] \le C n^{-2\min\{ \varphi_1,\varphi_2\}} ,
\end{equation}
for some positive constant $C = C(p,C_1,M_1,M_2)$.
The second step shows
\begin{equation}
E\left[ \max_{k=1,\ldots,K} \left ( n_k^{-1/2} \sum_{i \in \mathcal{I}_k} (\hat{\eta}_i -\eta_i)^\top \partial_\eta \psi^z_i \right)^2 \right] \le C K n^{-1/2} n^{-2\min\{ \varphi_1,\varphi_2\}+1/2}
\end{equation}
which is sufficient to prove (ref) by using Markov's inequality and $1/2 < 2 \min \{\varphi_1,\varphi_2\}$.
\textit{Step 1:} Consider the following derivation
\begin{align*}
E \left[ \left( n_k^{-1/2} \sum_{i \in \mathcal{I}_k } (\hat{\eta}_i - \eta_i)^\top \partial_\eta \psi^z_i \right)^2 \mid (W_j: j \notin \mathcal{I}_k )\right]
\end{align*}
which is equal to
\begin{align*}
&\overset{(1)}{=}E \left[ n_k^{-1} \sum_{i \in \mathcal{I}_k } \left( (\hat{\eta}_i - \eta_i)^\top \partial_\eta \psi^z_i \right)^2 \mid (W_j: j \notin \mathcal{I}_k) \right] \\
&\overset{(2)}{=} n_k^{-1} \sum_{i \in \mathcal{I}_k } (\hat{\eta}_i - \eta_i)^\top E\left[ (\partial_\eta \psi^z ) (\partial_\eta \psi^z )^\top \mid (W_j: j \notin \mathcal{I}_k) \right] (\hat{\eta}_i - \eta_i) \\
&\overset{(3)}{\le} C_1 \times p \times n_k^{-1} \sum_{i \in \mathcal{I}_k } || \hat{\eta}_k(X_i) - \eta_0(X_i) ||^2
\end{align*}
where (1) by the i.i.d zero mean of the random vectors $\{ (\hat{\eta}_i - \eta_i)^\top \partial_\eta \psi^z_i : i \in \mathcal{I}_k \} $ conditional on $ (W_j: j \notin \mathcal{I}_k)$ that holds by part (b) of Assumption (ref), (2) holds since $\hat{\eta}_i - \eta_i$ are not random conditional on $ (W_j: j \notin \mathcal{I}_k)$, and (3) by part (d) of Assumption (ref) and Loeve’s inequality (Davidson1994).
Using the previous derivations, it follows
\begin{align*}
E\left[ \left ( n_k^{-1/2} \sum_{i \in \mathcal{I}_k} (\hat{\eta}_i -\eta_i)^\top \partial_\eta \psi^z_i \right)^2 \right] &\le E\left[ C_1 \times p \times n_k^{-1} \sum_{i \in \mathcal{I}_k } || \hat{\eta}_k(X_i) - \eta_0(X_i) ||^2 \right] \\
&\overset{(1)}{=} C_1 \times p \times E\left[ || \hat{\eta}_k(X_i) - \eta_0(X_i) ||^2 \right] \\
&\overset{(2)}{\le} C n^{-2\min\{ \varphi_1,\varphi_2\}}
\end{align*}
for some positive constant $C = C(p,C_1,M_1,M_2)$, where (1) holds since $\hat{\eta}_k(X_i) - \eta_0(X_i)$ are i.i.d. for $i \in \mathcal{I}_k$, and (2) by part 2 of Lemma (ref) and (ref), which defines the constant $C$. This completes the proof of step 1.
\textit{Step 2:} Note that the maximum of $K$ positive number is bounded by their sum. Using this observation and (ref), it follows (ref).
\textbf{Part 2:} The proof of (ref) is similar to the proof part 1. It follows from the following inequality:
\begin{align*}
E\left[ \max_{k=1,\dots,K} n_k^{-1} \sum_{i \in \mathcal{I}_K} ||\hat{\eta}_i - \eta_i ||^2 \right] &\le E\left[ \sum_{k=1}^K \left(n_k^{-1} \sum_{i \in \mathcal{I}_K} ||\hat{\eta}_i - \eta_i ||^2 \right) \right] \\
&= K E\left[ ||\hat{\eta}_i - \eta_i ||^2 \right] \\
&\le K n^{-1/2} O(n^{-2\min\{\varphi_1,\varphi_2\} + 1/2}) ,
\end{align*}
which goes to zero since $ 1/2 < 2\min\{\varphi_1,\varphi_2\} $, and this is sufficient to prove (ref) by using Markov's inequality.
\textbf{Part 3:} The proof of (ref) follows from (ref), by using $\partial_\eta m_i = \partial_\eta \psi^b - \theta_0 \partial_\eta \psi^a$ and $|\theta_0| \le M_1^{1/4}/C_0$ (due to parts (a) and (c) of Assumptions (ref) and (ref)).
\textbf{Part 4:} The proof of (ref) follows from (ref) and (ref) and the following inequality
\begin{align*}
\left| n_k^{-1} \sum_{ i \in \mathcal{I}_k} \hat{\psi}^a_i - \psi^a_i \right| \le n_k^{-1/2} \left | n_k^{-1/2} \sum_{i \in \mathcal{I}_k} (\hat{\eta}_i -\eta_i)^\top \partial_\eta \psi^z_i \right| + C_2 p n_k^{-1} \sum_{i \in \mathcal{I}_k } || \hat{\eta}_i - \eta_i ||^2
\end{align*}
which holds due to Taylor expansion and mean valued theorem, part (e) of Assumption (ref), and Loeves' inequality.
\end{proof}
\subsection{Proof of Lemma (ref)}
\begin{proof}
For $i \in \mathcal{I}_k$, denote $\Delta_{i} = \Delta_{i}^b + \Delta_{i}^l$, where
\begin{align*}
\Delta_{i}^l &= n_0^{-\varphi_1} n_0^{-1/2} \sum_{j \notin \mathcal{I}_k} \delta_{n_0,j,i} ,\\
\Delta_{i}^b &= n_0^{-\varphi_2} n_0^{-1} \sum_{j \notin \mathcal{I}_k} b_{n_0,j,i} ,
\end{align*}
Here, $\delta_{n_0,j,i} = \delta_{n_0}(W_{j},X_i)$ and $b_{n_0,j,i} = b_{n_0}(X_{j},X_i)$, and $\delta_{n_0}$ and $b_{n_0}$ are functions satisfying Assumption (ref). In what follows, the results are proved for any given sequence $K$ that diverges to infinity as $n$ diverges to infinity, which is sufficient to guarantee the result of this lemma.
\textbf{Part 1:} By part (a) of Assumption (ref),
\begin{equation*}
\hat{\eta}_i - \eta_i = \Delta^l_i + \Delta^b_i + n_0^{-2\min\{ \varphi_1, \varphi_2 \} } \hat{R}_i
\end{equation*}
and by Loeve’s inequality (Davidson1994),
\begin{align*}
|| \hat{\eta}_i - \eta_i||^4 \le 3^3\left( || \Delta_i^l ||^4 + || \Delta_i^b ||^4 + || n_0^{-2\min\{ \varphi_1, \varphi_2 \}} \hat{R}_i||^4 \right) .
\end{align*}
Using the previous inequality, it follows
\begin{align}
n^{-1} \sum_{i=1}^n || \hat{\eta}_i - \eta_i||^4 &\le 3^3\left( n^{-1} \sum_{i=1}^n || \Delta_i^l||^4 + n^{-1} \sum_{i=1}^n || \Delta_i^b||^4 + n_0^{-8\min\{ \varphi_1, \varphi_2 \}} n^{-1} \sum_{i=1}^n || \hat{R}_i ||^4 \right) \notag \\
&\overset{(1)}{\le} 3^3\left( O_p(n^{-4\min\{\varphi_1, \varphi_2\}}) + n_0^{-8\min\{ \varphi_1, \varphi_2 \}} n^{-1} \sum_{i=1}^n || \hat{R}_i||^4 \right) \notag \\
&\overset{(2)}{=} O_p(n^{-4\min\{\varphi_1, \varphi_2\}}) ,
\end{align}
where (1) holds by Markov's inequality and Lemma (ref), and (2) holds by part (d) of Assumption (ref) and since $n_0 = ((K-1)/K) n$, which completes the proof of part 1.
\textbf{Part 2:} By part (a) of Assumption (ref),
\begin{equation*}
\hat{\eta}_i - \eta_i = \Delta_i^l + \Delta_i^b + n_0^{-2\min\{\varphi_1, \varphi_2\} } \hat{R}_i
\end{equation*}
and by Loeve’s inequality (Davidson1994),
\begin{align*}
|| \hat{\eta}_i - \eta_i||^2 \le 3 || \Delta_i^l||^2 + 3|| \Delta_i^b ||^2 + 3 || n_0^{-2\varphi_1} \hat{R}_i ||^2 .
\end{align*}
Using the previous inequality, it follows
\begin{align}
E[ || \hat{\eta}_i - \eta_i||^2 ] &\le 3 E[ || \Delta_i^l ||^2] + 3 E[ || \Delta_i^b ||^2]+ 3 E [ || n_0^{-2\varphi_1} \hat{R}_i ||^2] \notag \\
&\overset{(1)}{\le} 3 E[ || \Delta_i^l ||^4]^{1/2} + 3 E[ || \Delta_i^b ||^4]^{1/2}+ 3 n_0^{-2 \min\{ \varphi_1,\varphi_2\}} O(1)\notag \\
&\overset{(2)}{=} O(n^{-2\min\{\varphi_1, \varphi_2\}}) ,
\end{align}
where (1) holds by Jensen's inequality and part (c) of Assumption (ref), and (2) holds by Lemma (ref) and since $n_0 = ((K-1)/K) n$. This completes the proof of part 2.
\textbf{Part 3:} It follows from parts 2 and Markov's inequality.
\textbf{Part 4:} It follows from part 2 and by using that $ n^{1/2-\min\{ \varphi_1,\varphi_2\}} = o(1)$.
\end{proof}
\begin{lemma}
Suppose Assumptions (ref) and (ref) hold. In addition, assume $K$ is such that $K \le n$, $K \to \infty $ and $K/n^{\gamma} \to c \in [0,+\infty)$ as $n \to \infty$.
\begin{enumerate}
• If $\gamma = 1/2$, then
$$ \max_{k=1,\ldots,K} \left| 1 + n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a(W_i,\eta_i)-J_0)/J_0 \right|^{-1} = O_p(1)$$
• If $\gamma = 1/2$, $1/4 < \min\{\varphi_1,\varphi_2\}$ and $\varphi_1 \le 1/2$, then
$$ \max_{k=1,\ldots,K} \left| 1 + n_k^{-1} \sum_{ i \in \mathcal{I}_k} ( {\psi}^a(W_i,\hat{\eta}_i)-J_0)/J_0 \right|^{-1} = O_p(1)$$
• If $\gamma =1$, then
$$ \left| n^{-1} \sum_{i=1}^n {\psi}^a(W_i,\eta_i)/J_0 \right|^{-1}= O_p(1)$$
• If $\gamma = 1$ and $\varphi_1 > 1/4$, then
$$ \left| n^{-1} \sum_{i=1}^n {\psi}^a(W_i, \hat{\eta}_i) /J_0 \right|^{-1} = O_p(1) $$
\end{enumerate}
where $\eta_i = \eta_0(X_i)$, $\hat{\eta}_i$ is as in (ref), and $J_0 = E[\psi^a(W_i,\eta_i)]$.
\end{lemma}
\begin{proof}
\textbf{Part 1:} Consider $M>1$ and the following derivations
\begin{align*}
& P\left( \max_{k=1,\ldots,K} \left| 1 + n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right|^{-1} \le M \right) \\
&\overset{(1)}{=} P\left( \left| 1 + n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right|^{-1} \le M \right)^K \\
&= P\left( 1/M \le \left| 1 + n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right| \right)^K \\
&\ge P\left( 1/M \le 1 + n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right)^K \\
&= \left\{ 1 - P\left( n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 < -(M-1)/M \right) \right\}^K\\
&\overset{(2)}{\ge} \left\{ 1 - P\left( \left| n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right| > (M-1)/M \right) \right\}^K \\
&\overset{(3)}{\ge} \left\{ 1 - ((M-1)/M)^4 n_k^{-2} E\left[ \left| n_k^{-1/2} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right|^4\right] \right\}^K \\
&\overset{(4)}{\ge} \left\{ 1 - ((M-1)/M)^4 n_k^{-2} 10 E\left[ \left| (\psi^a_i-J_0)/J_0 \right|^4\right] \right\}^K \\
&\overset{(5)}{\ge} 1 - ((M-1)/M)^4 n^{-2} K^2 10 E\left[ \left| (\psi^a_i-J_0)/J_0 \right|^4\right] K \\
&\overset{(6)}{\ge} 1 - n^{-2 + 3/2 } (K n^{-1/2})^3 ((M-1)/M)^4 O(1) \\
&\overset{(7)}{\ge} 1 - o(1)
\end{align*}
(1) holds since $\{ \psi^a_i : 1 \le i \le n \}$ are i.i.d. random variables and because $\{ \mathcal{I}_k : 1 \le k \le K \}$ defines a partition of $\{1,\ldots,n\}$, (2) holds since $M>1$, (3) holds by Markov's inequality, (4) holds since $\{ \psi^a_i-J_0 : i \in \mathcal{I}_k\}$ are zero mean i.i.d. random variables, (5) holds by Bernoulli's inequality, (6) holds by parts (a) and (c) of Assumption (ref), and (7) holds since $K = O(n^{1/2})$.\\
\textbf{Part 2:} Define the event $E_{n,\epsilon} = \{ \max_{k=1,\ldots,K} \left| n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\hat{\psi}^a_i - \psi^a_i)/J_0 \right| < \epsilon \} $. Now, consider an small $\epsilon>0$ and $M>1$ such that $(1/M + \epsilon)^{-1} > 1$ and the following derivations
\begin{align*}
& P\left( \max_{k=1,\ldots,K} \left| 1 + n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 + (\hat{\psi}^a_i - \psi^a_i)/J_0 \right|^{-1} \le M \right)\\
& = P\left( \min_{k=1,\ldots,K} \left| 1 + n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 + (\hat{\psi}^a_i - \psi^a_i)/J_0 \right| \ge 1/M \right) \\
&\ge P\left( \min_{k=1,\ldots,K} \left| 1 + n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 + (\hat{\psi}^a_i - \psi^a_i)/J_0 \right| \ge 1/M, E_{n,\epsilon} \right) \\
&\overset{(1)}{\ge} P\left( \min_{k=1,\ldots,K} \left| 1 + n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right| \ge 1/M + \epsilon , E_{n,\epsilon} \right) \\
&\ge P\left( \min_{k=1,\ldots,K} \left| 1 + n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right| \ge 1/M + \epsilon \right) - P(E_{n,\epsilon}^c) \\
&\overset{(2)}{\ge} P\left( \max_{k=1,\ldots,K} \left| 1 + n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\psi^a_i-J_0)/J_0 \right|^{-1} \le (1/M + \epsilon)^{-1} \right) - o(1) \\
&\overset{(3)}{\ge} 1 -o(1) - o(1)
\end{align*}
where (1) holds because $\min_{k=1,\ldots,K} - \left| n_k^{-1} \sum_{ i \in \mathcal{I}_k} (\hat{\psi}^a_i - \psi^a_i)/J_0 \right| > - \epsilon$ conditional on the event $E_{n,\epsilon}$ and triangular inequality, (2) holds since $P(E_{n,\epsilon}^c) = o(1)$ due to Lemma (ref) (here I use $ 1/2 < 2 \min \{ \varphi_1, \varphi_2\}$ and $\varphi_1 \le 1/2$), and (3) holds by the same arguments presented in the proof of Part 1 by using $(1/M + \epsilon)^{-1}$ instead of $M$; therefore, it is omitted. \\
\textbf{Part 3:} Consider $M>1$ and $\tilde{M}> 1$ such that $\tilde{M} < ((M-1)/M) n^{1/2} $ and the following derivations,
\begin{align*}
P\left( \left| n^{-1} \sum_{i=1}^n {\psi}^a_i/J_0 \right|^{-1} >M \right) &= P\left( \left| 1 + n^{-1} \sum_{i=1}^n ({\psi}^a_i-J_0)/J_0 \right| < 1/M \right) \\
&\overset{(1)}{\le} P\left( n^{-1/2} \sum_{i=1}^n ({\psi}^a_i-J_0)/J_0 < -((M-1)/M) n^{1/2} \right) \\
&\overset{(2)}{\le} P\left( n^{-1/2} \sum_{i=1}^n ({\psi}^a_i-J_0)/J_0 < -\tilde{M} \right)\\
&\overset{(3)}{=} \Phi(-\tilde{M} / \sigma_a) + o(1)
\end{align*}
where (1) holds since $M>1$, (2) holds by definition of $\tilde{M}$, and (3) holds by CLT as $n \to \infty$ (here, $\sigma_a^2$ is as in (ref)). To complete the proof, note that $ \Phi(-\tilde{M} / \sigma_a) \to 0$ as $\tilde{M} \to \infty$. \\
\textbf{Part 4:} Define the event $E_{n,\epsilon} = \{ \left| n^{-1} \sum_{i = 1}^n (\hat{\psi}^a_i - \psi^a_i)/J_0 \right| < \epsilon \}$. Now consider an small $\epsilon>0$ and $M>1$ such that $(1/M + \epsilon)^{-1} > 1$. Note that $P(E_{n,\epsilon}^c) = o(1)$ due to Lemma (ref) (here I use $\min \{\varphi_1, \varphi_2\} > 1/4$). The proof is completed by similar arguments presented in part 2 and part 3; therefore, it is omitted.
\end{proof}