The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
229,196 characters
Supplemental Material for ``A Unified Framework for Efficient Estimation of General Treatment Models ''
\title{Supplemental Material for \\ ``\textit{A Unified Framework for Efficient Estimation of General
Treatment Models }''}
\author{ \ \ Chunrong Ai\thanks{
Department of Economics, University of Florida.
E-mail: \texttt{[email removed]}} -- University of Florida \\
Oliver Linton\thanks{
Faculty of Economics, University of Cambridge.
E-mail: \texttt{[email removed]}} -- University of Cambridge \\
Kaiji Motegi\thanks{
Graduate School of Economics, Kobe University. E-mail: \texttt{[email removed]}} -- Kobe University \\
Zheng Zhang\thanks{
Institute of Statistics and Big Data, Renmin University of China. E-mail: \texttt{[email removed]}} -- Renmin University of China}
\date{{\large This draft:} \today
}
\maketitle
\setstretch{1.3}
\newpage
\tableofcontents
\listoftables
\newpage
\section{Complete Simulation Results}
In this supplemental material, we present complete simulation results that are partly omitted in the main paper \cite{Ai_Motegi_Zhang_cts_treat}.
See Section \ref{sec:sim_design} for simulation designs and Section \ref{sec:sim_results} for results.
\subsection{Simulation Design \label{sec:sim_design}}
We consider four data generating processes (DGPs):
\begin{description}
\setlength{\leftskip}{1.0cm}
\item[DGP-1] $T_{i} = X_{1i} + X_{2i} + 0.2 X_{3i} + 0.2 X_{4i} + \xi_{i}$ and $Y_{i} = 1 + X_{1i} + 0.1 X_{2i} + 0.1 X_{3i} + 0.1 X_{4i} + T_{i} + \epsilon_{i}$.
\item[DGP-2] $T_{i} = (X_{1i} + 0.5)^{2} + 0.4 X_{2i} + 0.4 X_{3i} + 0.4 X_{4i} + \xi_{i}$ and $Y_{i} = 1 + X_{1i} + 0.1 X_{2i} + 0.1 X_{3i} + 0.1 X_{4i} + T_{i} + \epsilon_{i}$.
\item[DGP-3] $T_{i} = X_{1i} + X_{2i} + 0.2 X_{3i} + 0.2 X_{4i} + \xi_{i}$ and $Y_{i} = 0.75 X_{1i}^{2} + 0.2 (X_{2i} + 0.5)^{2} + T_{i} + \epsilon_{i}$.
\item[DGP-4] $T_{i} = (X_{1i} + 0.5)^{2} + 0.4 X_{2i} + 0.4 X_{3i} + 0.4 X_{4i} + \xi_{i}$ and $Y_{i} = 0.75 X_{1i}^{2} + 0.2 (X_{2i} + 0.5)^{2} + T_{i} + \epsilon_{i}$.
\end{description}
For each DGP, $\xi_{i} \stackrel{i.i.d.}{\sim} N(0, 9)$, $\epsilon_{i} \stackrel{i.i.d.}{\sim} N(0, 25)$,
and $\boldsymbol{X}_{i} = [X_{1i}, X_{2i}, X_{3i}, X_{4i}]^{T} \stackrel{i.i.d.}{\sim} N(0, \boldsymbol{\Sigma})$, where the diagonal elements of $\boldsymbol{\Sigma}$ are all 1 and the off-diagonal elements are all $\rho \in \{ 0.0, 0.2, 0.4 \}$.
We draw $J = 1000$ Monte Carlo samples with sample size $N \in \{ 100, 500, 1000 \}$.
Our DGPs have similar structures with \cite{Fong_Hazlett_Imai_2018}.
The characteristics of the four DGPs can be summarized as follows.
\begin{description}
\setlength{\leftskip}{1.0cm}
\item[DGP-1] $T_{i}$ is linear in $\boldsymbol{X}_{i}$; $Y_{i}$ is linear in $\boldsymbol{X}_{i}$ and $T_{i}$.
\item[DGP-2] $T_{i}$ is nonlinear in $\boldsymbol{X}_{i}$; $Y_{i}$ is linear in $\boldsymbol{X}_{i}$ and $T_{i}$.
\item[DGP-3] $T_{i}$ is linear in $\boldsymbol{X}_{i}$; $Y_{i}$ is nonlinear in $\boldsymbol{X}_{i}$ and linear in $T_{i}$.
\item[DGP-4] $T_{i}$ is nonlinear in $\boldsymbol{X}_{i}$; $Y_{i}$ is nonlinear in $\boldsymbol{X}_{i}$ and linear in $T_{i}$.
\end{description}
In the main paper \cite{Ai_Motegi_Zhang_cts_treat}, we discuss only DGP-1 and DGP-4 in order to save space.
(DGP-4 here is called ``DGP-2'' in the main paper since DGP-2 and DGP-3 here are skipped.)
As shown in the main paper, a parametric version of Fong, Hazlett, and Imai's \citeyearpar{Fong_Hazlett_Imai_2018} covariate balancing generalized propensity score (CBGPS) estimators
produces no bias under DGP-1 and considerable bias under DGP-4.\footnote{
A non-parametric version of the CBGPS estimators is also biased under a data generating process which is similar to DGP-4,
and the magnitude of the bias is as large as the parametric version \citep[cf. Figure 2,][]{Fong_Hazlett_Imai_2018}.
}
Here we also consider DGP-2 and DGP-3, which serve as intermediate cases where \textit{either} $T_{i}$ is nonlinear in $\boldsymbol{X}_{i}$ \textit{or} $Y_{i}$ is nonlinear in $\boldsymbol{X}_{i}$.
Since $Y_{i}$ is linear in $T_{i}$ under all four DGPs, the link function is always of the form $\mathbb{E} [Y^*(t)] = \beta_{1} + \beta_{2} t$.
It is straightforward to see that the true coefficients are $(\beta_{1}, \beta_{2}) = (1, 1)$ for each DGP.
$\beta_{2}$ is of greater interest than $\beta_{1}$ since $\beta_{2}$ measures the average treatment effect.
To estimate $(\beta_{1}, \beta_{2})$ via our proposed method, we should decide which variables to include in polynomials with respect to $T_{i}$ and $\boldsymbol{X}_{i}$.
For each DGP, we use $u_{K_{1}} (t_{i}) = [1, t_{i}, t_{i}^{2}]^{T}$ (i.e. $K_{1} = 3$) and $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}]^{T}$ (i.e. $K_{2} = 9$)
or $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}, X_{1i} X_{2i}, X_{1i} X_{3i}, X_{1i} X_{4i}, X_{2i} X_{3i}, X_{2i} X_{4i}, X_{3i} X_{4i}]^{T}$ (i.e. $K_{2} = 15$).
In the literature of non-parametric estimation of treatment effects, it is common to include the first and second moments of $\boldsymbol{X}_{i}$ with or without cross terms \citep[see e.g.][]{chan2016globally}.
Hence we consider two cases with and without the cross terms of $\boldsymbol{X}_{i}$.
We compare our estimators with the parametric version of Fong, Hazlett, and Imai's \citeyearpar{Fong_Hazlett_Imai_2018} CBGPS estimator.
The non-parametric version of CBGPS is omitted since the parametric and non-parametric versions exhibit similar performance according to the simulation results of \cite{Fong_Hazlett_Imai_2018}.
Computation of the parametric CBGPS estimator involves two steps.
The first step is the estimation of stabilized weights, and the second step is the estimation of average treatment effects.
Our covariates are chosen to be $\boldsymbol{X}_{i} = [X_{1i}, X_{2i}, X_{3i}, X_{4i}]^{T}$ for the first step
and $\boldsymbol{Z}_{i} = [1, T_{i}, \boldsymbol{X}_{i}^{T}]^{T}$ for the second step.\footnote{
In extra simulations not reported here, we added some irrelevant covariates to the model, as \cite{Fong_Hazlett_Imai_2018} did, in order to see whether the finite sample performance is sensitive to redundant covariates.
Since the results were not so sensitive to the inclusion of irrelevant covariates, we only report results without the irrelevant covariates here.
}
\subsection{Simulation Results \label{sec:sim_results}}
In this section we present our simulation results.
See Tables \ref{table:sim_results_beta1_DGP1}-\ref{table:sim_results_beta2_DGP1} for DGP-1;
Tables \ref{table:sim_results_beta1_DGP2}-\ref{table:sim_results_beta2_DGP2} for DGP-2;
Tables \ref{table:sim_results_beta1_DGP3}-\ref{table:sim_results_beta2_DGP3} for DGP-3;
Tables \ref{table:sim_results_beta1_DGP4}-\ref{table:sim_results_beta2_DGP4} for DGP-4.
For each table, we report the bias, standard deviation, root mean squared error, and coverage probability based on the 95\% confidence bands.
As discussed in the main paper \cite{Ai_Motegi_Zhang_cts_treat}, both of our stabilized-weight (SW) estimator and Fong, Hazlett, and Imai's \citeyearpar{Fong_Hazlett_Imai_2018} CBGPS estimator are unbiased under DGP-1.
The latter often has smaller bias and standard deviation across all sample sizes $N \in \{ 100, 500, 1000 \}$.
That is not a surprising result since DGP-1 satisfies the linearity assumption underpinning the CBGPS estimator.
The coverage probability, however, is sometimes comparable between the two estimators.
See, for example, Table \ref{table:sim_results_beta2_DGP1} ($\beta_{2}$) with $\rho = 0.2$ and $N = 500$,
where the coverage probability is 0.926 for the SW estimator with $K_{2} = 9$ and 0.890 for the CBGPS estimator.
The simulation results are not sensitive to the correlation coefficient $\rho \in \{ 0.0, 0.2, 0.4 \}$.
The SW estimators with and without the cross terms of $\boldsymbol{X}_{i}$ lead to roughly similar performance.
Similar implications hold for DGP-2.
The CBGPS estimator performs well arguably due to a negligibly small degree of nonlinearity.
There is a nonlinear term $(X_{1i} + 0.5)^{2}$ in the DGP of $T_{i}$, but the DGP of $Y_{i}$ is kept to be linear.
Hence the CBGPS estimator maintains the sharp performance as in DGP-1.
Under DGP-3, the SW estimator begins to be more comparable with the CBGPS estimator.
See, for example, Table \ref{table:sim_results_beta2_DGP3} ($\beta_{2}$) with $\rho = 0.4$ and $N = 1000$,
where the bias is 0.003 for the SW estimator with $K_{2} = 9$ and 0.004 for the CBGPS estimator;
the standard deviation is 0.174 for SW and 0.090 for CBGPS;
the coverage probability is 0.932 for SW and 0.892 for CBGPS.
Under DGP-3, there are two nonlinear terms $0.75 X_{1i}^{2}$ and $0.2 (X_{2i} + 0.5)^{2}$ in the DGP of $Y_{i}$.
It is likely that the CBGPS estimator is negatively affected by those nonlinear features while the SW estimator is more robust against nonlinearity.
As discussed in the main paper \cite{Ai_Motegi_Zhang_cts_treat}, our SW estimator dominates the CBGPS estimator under DGP-4.
See, for example, Table \ref{table:sim_results_beta1_DGP4} ($\beta_{1}$) with $\rho = 0.4$ and $N = 1000$,
where the bias is 0.016 for the SW estimator with $K_{2} = 9$ and -0.182 for the CBGPS estimator;
the standard deviation is 0.605 for SW and 0.194 for CBGPS;
the coverage probability is 0.923 for SW and 0.800 for CBGPS.
Also see Table \ref{table:sim_results_beta2_DGP4} ($\beta_{2}$) with $\rho = 0.4$ and $N = 1000$,
where the bias is 0.038 for SW and 0.156 for CBGPS;
the standard deviation is 0.170 for SW and 0.080 for CBGPS;
the coverage probability is 0.916 for SW and 0.304 for CBGPS.
Under DGP-4, the CBGPS estimator keeps producing negative bias in $\beta_{1}$ and positive bias in $\beta_{2}$.
The bias is considerably large, and it does not vanish as sample size $N$ increases.
Our estimator, in contrast, produces virtually no bias for any sample size $N \in \{ 100, 500, 1000 \}$.
It also achieves accurate coverage probability for larger sample sizes $N \in \{ 500, 1000 \}$.
The relative advantage of our estimator stems from the strong degree of nonlinearity contained in DGP-4, where both $T_{i}$ and $Y_{i}$ have nonlinear terms.
In summary, our estimator is by construction robust against the functional form of underlying DGPs,
while Fong, Hazlett, and Imai's \citeyearpar{Fong_Hazlett_Imai_2018} estimator is sensitive to model misspecification.
In fact, our approach produces unbiased estimates under all DGPs considered,
while Fong, Hazlett, and Imai's \citeyearpar{Fong_Hazlett_Imai_2018} approach produces biased estimates under DGP-4.
In reality, the functional form of an underlying DGP is unknown to the researcher.
It is therefore of practical use to employ our estimator in order to accomplish correct inference for dose-response curves.
\clearpage
\begin{table}[th]
\begin{center}
\caption{Simulation Results on $\beta_{1}$ (DGP-1) \label{table:sim_results_beta1_DGP1}}
{\fontsize{9.5pt}{16.5pt} \selectfont
\begin{tabular}{c|cccc|cccc|cccc} \hline \hline
\multicolumn{13}{c}{$\rho = 0.0$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.019 & 0.866 & 0.866 & 0.872 & 0.019 & 0.527 & 0.523 & 0.951 & -0.004 & 0.544 & 0.544 & 0.937 \\
SW ($K_{2} = 15$) & 0.093 & 0.979 & 0.983 & 0.862 & 0.017 & 0.666 & 0.666 & 0.961 & 0.008 & 0.630 & 0.630 & 0.980 \\
CBGPS & -0.012 & 0.590 & 0.590 & 0.940 & -0.009 & 0.276 & 0.276 & 0.943 & 0.006 & 0.198 & 0.198 & 0.944 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.2$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.011 & 0.934 & 0.934 & 0.874 & -0.010 & 0.620 & 0.620 & 0.947 & 0.019 & 0.561 & 0.561 & 0.946 \\
SW ($K_{2} = 15$) & -0.018 & 1.200 & 1.200 & 0.872 & -0.005 & 0.633 & 0.633 & 0.978 & -0.004 & 0.641 & 0.641 & 0.976 \\
CBGPS & 0.013 & 0.643 & 0.643 & 0.937 & 0.002 & 0.296 & 0.296 & 0.930 & 0.000 & 0.211 & 0.211 & 0.927 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.4$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.031 & 1.012 & 1.012 & 0.855 & 0.002 & 0.642 & 0.642 & 0.927 & -0.021 & 0.637 & 0.638 & 0.942 \\
SW ($K_{2} = 15$) & 0.013 & 0.975 & 0.975 & 0.876 & 0.029 & 0.693 & 0.694 & 0.963 & -0.022 & 0.666 & 0.667 & 0.980 \\
CBGPS & 0.001 & 0.634 & 0.634 & 0.944 & -0.001 & 0.318 & 0.318 & 0.924 & 0.009 & 0.227 & 0.227 & 0.916 \\ \hline \hline
\end{tabular}
}
\end{center}
{\fontsize{9.5pt}{12pt} \selectfont
The DGP is $T_{i} = X_{1i} + X_{2i} + 0.2 X_{3i} + 0.2 X_{4i} + \xi_{i}$ and $Y_{i} = 1 + X_{1i} + 0.1 X_{2i} + 0.1 X_{3i} + 0.1 X_{4i} + T_{i} + \epsilon_{i}$.
The covariates follow $\boldsymbol{X}_{i} \stackrel{i.i.d.}{\sim} N(0, \boldsymbol{\Sigma})$, where the diagonal elements of $\boldsymbol{\Sigma}$ are all 1 and the off-diagonal elements are all $\rho \in \{ 0.0, 0.2, 0.4 \}$.
We report the bias, standard deviation, root mean squared error, and coverage probability based on the 95\% confidence band across $J = 1000$ Monte Carlo samples.
'SW' signifies our own stabilized-weight estimator, while 'CBGPS' signifies Fong, Hazlett, and Imai's \citeyearpar{Fong_Hazlett_Imai_2018} parametric covariate balancing generalized propensity score (CBGPS) estimator.
For SW, the polynomials are $u_{K_{1}} (t_{i}) = [1, t_{i}, t_{i}^{2}]^{T}$ (i.e. $K_{1} = 3$) and $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}]^{T}$ (i.e. $K_{2} = 9$)
or $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}, X_{1i} X_{2i}, X_{1i} X_{3i}, X_{1i} X_{4i}, X_{2i} X_{3i}, X_{2i} X_{4i}, X_{3i} X_{4i}]^{T}$ (i.e. $K_{2} = 15$).
For CBGPS, covariates are chosen to be $\boldsymbol{X}_{i} = [X_{1i}, X_{2i}, X_{3i}, X_{4i}]^{T}$ for the first step of estimating stabilized weights
and $\boldsymbol{Z}_{i} = [1, T_{i}, \boldsymbol{X}_{i}^{T}]^{T}$ for the second step of estimating average treatment effects.
}
\end{table}
\clearpage
\begin{table}[th]
\begin{center}
\caption{Simulation Results on $\beta_{2}$ (DGP-1) \label{table:sim_results_beta2_DGP1}}
{\fontsize{9.5pt}{16.5pt} \selectfont
\begin{tabular}{c|cccc|cccc|cccc} \hline \hline
\multicolumn{13}{c}{$\rho = 0.0$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.023 & 0.291 & 0.292 & 0.853 & 0.016 & 0.189 & 0.189 & 0.928 & 0.017 & 0.167 & 0.168 & 0.922 \\
SW ($K_{2} = 15$) & 0.036 & 0.316 & 0.318 & 0.819 & 0.022 & 0.201 & 0.202 & 0.969 & 0.035 & 0.186 & 0.189 & 0.979 \\
CBGPS & 0.001 & 0.205 & 0.205 & 0.922 & -0.007 & 0.099 & 0.099 & 0.914 & -0.001 & 0.072 & 0.072 & 0.906 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.2$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.031 & 0.290 & 0.292 & 0.846 & 0.031 & 0.186 & 0.188 & 0.926 & 0.027 & 0.174 & 0.176 & 0.917 \\
SW ($K_{2} = 15$) & 0.050 & 0.313 & 0.317 & 0.814 & 0.040 & 0.202 & 0.206 & 0.964 & 0.046 & 0.190 & 0.195 & 0.982 \\
CBGPS & -0.013 & 0.219 & 0.219 & 0.911 & 0.002 & 0.106 & 0.106 & 0.890 & 0.000 & 0.083 & 0.083 & 0.909 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.4$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.044 & 0.287 & 0.290 & 0.841 & 0.038 & 0.190 & 0.194 & 0.916 & 0.038 & 0.167 & 0.172 & 0.923 \\
SW ($K_{2} = 15$) & 0.050 & 0.300 & 0.304 & 0.811 & 0.042 & 0.201 & 0.205 & 0.955 & 0.037 & 0.188 & 0.192 & 0.975 \\
CBGPS & 0.015 & 0.227 & 0.227 & 0.895 & 0.007 & 0.115 & 0.115 & 0.911 & -0.002 & 0.087 & 0.087 & 0.882 \\ \hline \hline
\end{tabular}
}
\end{center}
{\fontsize{9.5pt}{12pt} \selectfont
The DGP is $T_{i} = X_{1i} + X_{2i} + 0.2 X_{3i} + 0.2 X_{4i} + \xi_{i}$ and $Y_{i} = 1 + X_{1i} + 0.1 X_{2i} + 0.1 X_{3i} + 0.1 X_{4i} + T_{i} + \epsilon_{i}$.
The covariates follow $\boldsymbol{X}_{i} \stackrel{i.i.d.}{\sim} N(0, \boldsymbol{\Sigma})$, where the diagonal elements of $\boldsymbol{\Sigma}$ are all 1 and the off-diagonal elements are all $\rho \in \{ 0.0, 0.2, 0.4 \}$.
We report the bias, standard deviation, root mean squared error, and coverage probability based on the 95\% confidence band across $J = 1000$ Monte Carlo samples.
'SW' signifies our own stabilized-weight estimator, while 'CBGPS' signifies Fong, Hazlett, and Imai's \citeyearpar{Fong_Hazlett_Imai_2018} parametric covariate balancing generalized propensity score (CBGPS) estimator.
For SW, the polynomials are $u_{K_{1}} (t_{i}) = [1, t_{i}, t_{i}^{2}]^{T}$ (i.e. $K_{1} = 3$) and $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}]^{T}$ (i.e. $K_{2} = 9$)
or $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}, X_{1i} X_{2i}, X_{1i} X_{3i}, X_{1i} X_{4i}, X_{2i} X_{3i}, X_{2i} X_{4i}, X_{3i} X_{4i}]^{T}$ (i.e. $K_{2} = 15$).
For CBGPS, covariates are chosen to be $\boldsymbol{X}_{i} = [X_{1i}, X_{2i}, X_{3i}, X_{4i}]^{T}$ for the first step of estimating stabilized weights
and $\boldsymbol{Z}_{i} = [1, T_{i}, \boldsymbol{X}_{i}^{T}]^{T}$ for the second step of estimating average treatment effects.
}
\end{table}
\clearpage
\begin{table}[th]
\begin{center}
\caption{Simulation Results on $\beta_{1}$ (DGP-2) \label{table:sim_results_beta1_DGP2}}
{\fontsize{9.5pt}{16.5pt} \selectfont
\begin{tabular}{c|cccc|cccc|cccc} \hline \hline
\multicolumn{13}{c}{$\rho = 0.0$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & -0.050 & 0.983 & 0.985 & 0.839 & -0.055 & 0.584 & 0.586 & 0.922 & -0.042 & 0.563 & 0.565 & 0.928 \\
SW ($K_{2} = 15$) & -0.003 & 1.033 & 1.033 & 0.843 & -0.055 & 0.682 & 0.684 & 0.953 & -0.089 & 0.735 & 0.740 & 0.965 \\
CBGPS & -0.025 & 0.586 & 0.586 & 0.934 & 0.010 & 0.255 & 0.255 & 0.957 & -0.001 & 0.181 & 0.181 & 0.942 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.2$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & -0.051 & 0.937 & 0.938 & 0.868 & -0.061 & 0.645 & 0.648 & 0.904 & -0.061 & 0.592 & 0.595 & 0.928 \\
SW ($K_{2} = 15$) & -0.100 & 1.086 & 1.091 & 0.839 & -0.062 & 0.737 & 0.739 & 0.962 & -0.076 & 0.722 & 0.726 & 0.960 \\
CBGPS & -0.014 & 0.604 & 0.604 & 0.938 & -0.006 & 0.265 & 0.266 & 0.936 & -0.008 & 0.185 & 0.185 & 0.950 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.4$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & -0.114 & 0.989 & 0.996 & 0.864 & -0.052 & 0.664 & 0.666 & 0.921 & -0.075 & 0.591 & 0.596 & 0.919 \\
SW ($K_{2} = 15$) & -0.111 & 1.064 & 1.070 & 0.851 & -0.100 & 0.700 & 0.707 & 0.944 & -0.095 & 0.715 & 0.721 & 0.958 \\
CBGPS & -0.006 & 0.625 & 0.625 & 0.923 & 0.011 & 0.271 & 0.271 & 0.934 & 0.004 & 0.190 & 0.190 & 0.937 \\ \hline \hline
\end{tabular}
}
\end{center}
{\fontsize{9.5pt}{12pt} \selectfont
The DGP is $T_{i} = (X_{1i} + 0.5)^{2} + 0.4 X_{2i} + 0.4 X_{3i} + 0.4 X_{4i} + \xi_{i}$ and $Y_{i} = 1 + X_{1i} + 0.1 X_{2i} + 0.1 X_{3i} + 0.1 X_{4i} + T_{i} + \epsilon_{i}$.
The covariates follow $\boldsymbol{X}_{i} \stackrel{i.i.d.}{\sim} N(0, \boldsymbol{\Sigma})$, where the diagonal elements of $\boldsymbol{\Sigma}$ are all 1 and the off-diagonal elements are all $\rho \in \{ 0.0, 0.2, 0.4 \}$.
We report the bias, standard deviation, root mean squared error, and coverage probability based on the 95\% confidence band across $J = 1000$ Monte Carlo samples.
'SW' signifies our own stabilized-weight estimator, while 'CBGPS' signifies Fong, Hazlett, and Imai's \citeyearpar{Fong_Hazlett_Imai_2018} parametric covariate balancing generalized propensity score (CBGPS) estimator.
For SW, the polynomials are $u_{K_{1}} (t_{i}) = [1, t_{i}, t_{i}^{2}]^{T}$ (i.e. $K_{1} = 3$) and $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}]^{T}$ (i.e. $K_{2} = 9$)
or $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}, X_{1i} X_{2i}, X_{1i} X_{3i}, X_{1i} X_{4i}, X_{2i} X_{3i}, X_{2i} X_{4i}, X_{3i} X_{4i}]^{T}$ (i.e. $K_{2} = 15$).
For CBGPS, covariates are chosen to be $\boldsymbol{X}_{i} = [X_{1i}, X_{2i}, X_{3i}, X_{4i}]^{T}$ for the first step of estimating stabilized weights
and $\boldsymbol{Z}_{i} = [1, T_{i}, \boldsymbol{X}_{i}^{T}]^{T}$ for the second step of estimating average treatment effects.
}
\end{table}
\clearpage
\begin{table}[th]
\begin{center}
\caption{Simulation Results on $\beta_{2}$ (DGP-2) \label{table:sim_results_beta2_DGP2}}
{\fontsize{9.5pt}{16.5pt} \selectfont
\begin{tabular}{c|cccc|cccc|cccc} \hline \hline
\multicolumn{13}{c}{$\rho = 0.0$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.029 & 0.275 & 0.276 & 0.821 & 0.020 & 0.177 & 0.178 & 0.909 & 0.021 & 0.159 & 0.160 & 0.914 \\
SW ($K_{2} = 15$) & 0.038 & 0.301 & 0.303 & 0.806 & 0.020 & 0.183 & 0.184 & 0.966 & 0.024 & 0.181 & 0.182 & 0.983 \\
CBGPS & 0.012 & 0.187 & 0.187 & 0.888 & -0.001 & 0.086 & 0.086 & 0.918 & 0.000 & 0.060 & 0.060 & 0.905 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.2$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.028 & 0.299 & 0.300 & 0.805 & 0.020 & 0.174 & 0.175 & 0.915 & 0.031 & 0.164 & 0.167 & 0.916 \\
SW ($K_{2} = 15$) & 0.037 & 0.295 & 0.298 & 0.796 & 0.029 & 0.199 & 0.201 & 0.969 & 0.030 & 0.183 & 0.186 & 0.966 \\
CBGPS & -0.009 & 0.191 & 0.191 & 0.874 & 0.000 & 0.087 & 0.087 & 0.897 & 0.001 & 0.066 & 0.066 & 0.902 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.4$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.038 & 0.288 & 0.290 & 0.808 & 0.030 & 0.181 & 0.183 & 0.887 & 0.029 & 0.159 & 0.162 & 0.898 \\
SW ($K_{2} = 15$) & 0.051 & 0.300 & 0.305 & 0.791 & 0.044 & 0.196 & 0.200 & 0.939 & 0.036 & 0.173 & 0.177 & 0.964 \\
CBGPS & -0.009 & 0.192 & 0.192 & 0.884 & -0.003 & 0.090 & 0.090 & 0.888 & -0.001 & 0.069 & 0.069 & 0.870 \\ \hline \hline
\end{tabular}
}
\end{center}
{\fontsize{9.5pt}{12pt} \selectfont
The DGP is $T_{i} = (X_{1i} + 0.5)^{2} + 0.4 X_{2i} + 0.4 X_{3i} + 0.4 X_{4i} + \xi_{i}$ and $Y_{i} = 1 + X_{1i} + 0.1 X_{2i} + 0.1 X_{3i} + 0.1 X_{4i} + T_{i} + \epsilon_{i}$.
The covariates follow $\boldsymbol{X}_{i} \stackrel{i.i.d.}{\sim} N(0, \boldsymbol{\Sigma})$, where the diagonal elements of $\boldsymbol{\Sigma}$ are all 1 and the off-diagonal elements are all $\rho \in \{ 0.0, 0.2, 0.4 \}$.
We report the bias, standard deviation, root mean squared error, and coverage probability based on the 95\% confidence band across $J = 1000$ Monte Carlo samples.
'SW' signifies our own stabilized-weight estimator, while 'CBGPS' signifies Fong, Hazlett, and Imai's \citeyearpar{Fong_Hazlett_Imai_2018} parametric covariate balancing generalized propensity score (CBGPS) estimator.
For SW, the polynomials are $u_{K_{1}} (t_{i}) = [1, t_{i}, t_{i}^{2}]^{T}$ (i.e. $K_{1} = 3$) and $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}]^{T}$ (i.e. $K_{2} = 9$)
or $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}, X_{1i} X_{2i}, X_{1i} X_{3i}, X_{1i} X_{4i}, X_{2i} X_{3i}, X_{2i} X_{4i}, X_{3i} X_{4i}]^{T}$ (i.e. $K_{2} = 15$).
For CBGPS, covariates are chosen to be $\boldsymbol{X}_{i} = [X_{1i}, X_{2i}, X_{3i}, X_{4i}]^{T}$ for the first step of estimating stabilized weights
and $\boldsymbol{Z}_{i} = [1, T_{i}, \boldsymbol{X}_{i}^{T}]^{T}$ for the second step of estimating average treatment effects.
}
\end{table}
\clearpage
\begin{table}[th]
\begin{center}
\caption{Simulation Results on $\beta_{1}$ (DGP-3) \label{table:sim_results_beta1_DGP3}}
{\fontsize{9.5pt}{16.5pt} \selectfont
\begin{tabular}{c|cccc|cccc|cccc} \hline \hline
\multicolumn{13}{c}{$\rho = 0.0$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.005 & 0.832 & 0.832 & 0.888 & 0.065 & 0.572 & 0.576 & 0.937 & 0.096 & 0.570 & 0.579 & 0.950 \\
SW ($K_{2} = 15$) & -0.088 & 0.925 & 0.929 & 0.871 & 0.073 & 0.688 & 0.692 & 0.984 & 0.155 & 0.686 & 0.704 & 0.979 \\
CBGPS & -0.030 & 0.606 & 0.606 & 0.949 & 0.007 & 0.302 & 0.302 & 0.915 & 0.016 & 0.234 & 0.235 & 0.932 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.2$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & -0.007 & 0.883 & 0.883 & 0.886 & 0.066 & 0.634 & 0.638 & 0.936 & 0.130 & 0.595 & 0.609 & 0.952 \\
SW ($K_{2} = 15$) & -0.029 & 1.005 & 1.005 & 0.857 & 0.090 & 0.686 & 0.691 & 0.966 & 0.201 & 0.694 & 0.722 & 0.981 \\
CBGPS & -0.009 & 0.645 & 0.645 & 0.942 & 0.037 & 0.318 & 0.320 & 0.947 & 0.027 & 0.253 & 0.255 & 0.927 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.4$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & -0.028 & 1.009 & 1.009 & 0.834 & 0.057 & 0.681 & 0.684 & 0.928 & 0.085 & 0.619 & 0.625 & 0.937 \\
SW ($K_{2} = 15$) & -0.017 & 1.006 & 1.006 & 0.872 & 0.065 & 0.684 & 0.687 & 0.965 & 0.111 & 0.745 & 0.753 & 0.978 \\
CBGPS & 0.030 & 0.662 & 0.663 & 0.955 & 0.078 & 0.341 & 0.350 & 0.936 & 0.051 & 0.281 & 0.286 & 0.914 \\ \hline \hline
\end{tabular}
}
\end{center}
{\fontsize{9.5pt}{12pt} \selectfont
The DGP is $T_{i} = X_{1i} + X_{2i} + 0.2 X_{3i} + 0.2 X_{4i} + \xi_{i}$ and $Y_{i} = 0.75 X_{1i}^{2} + 0.2 (X_{2i} + 0.5)^{2} + T_{i} + \epsilon_{i}$.
The covariates follow $\boldsymbol{X}_{i} \stackrel{i.i.d.}{\sim} N(0, \boldsymbol{\Sigma})$, where the diagonal elements of $\boldsymbol{\Sigma}$ are all 1 and the off-diagonal elements are all $\rho \in \{ 0.0, 0.2, 0.4 \}$.
We report the bias, standard deviation, root mean squared error, and coverage probability based on the 95\% confidence band across $J = 1000$ Monte Carlo samples.
'SW' signifies our own stabilized-weight estimator, while 'CBGPS' signifies Fong, Hazlett, and Imai's \citeyearpar{Fong_Hazlett_Imai_2018} parametric covariate balancing generalized propensity score (CBGPS) estimator.
For SW, the polynomials are $u_{K_{1}} (t_{i}) = [1, t_{i}, t_{i}^{2}]^{T}$ (i.e. $K_{1} = 3$) and $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}]^{T}$ (i.e. $K_{2} = 9$)
or $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}, X_{1i} X_{2i}, X_{1i} X_{3i}, X_{1i} X_{4i}, X_{2i} X_{3i}, X_{2i} X_{4i}, X_{3i} X_{4i}]^{T}$ (i.e. $K_{2} = 15$).
For CBGPS, covariates are chosen to be $\boldsymbol{X}_{i} = [X_{1i}, X_{2i}, X_{3i}, X_{4i}]^{T}$ for the first step of estimating stabilized weights
and $\boldsymbol{Z}_{i} = [1, T_{i}, \boldsymbol{X}_{i}^{T}]^{T}$ for the second step of estimating average treatment effects.
}
\end{table}
\clearpage
\begin{table}[th]
\begin{center}
\caption{Simulation Results on $\beta_{2}$ (DGP-3) \label{table:sim_results_beta2_DGP3}}
{\fontsize{9.5pt}{16.5pt} \selectfont
\begin{tabular}{c|cccc|cccc|cccc} \hline \hline
\multicolumn{13}{c}{$\rho = 0.0$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.019 & 0.273 & 0.274 & 0.840 & 0.014 & 0.187 & 0.188 & 0.936 & 0.006 & 0.161 & 0.161 & 0.956 \\
SW ($K_{2} = 15$) & 0.011 & 0.296 & 0.297 & 0.842 & 0.017 & 0.212 & 0.213 & 0.971 & -0.003 & 0.184 & 0.184 & 0.979 \\
CBGPS & -0.000 & 0.224 & 0.224 & 0.905 & 0.005 & 0.100 & 0.100 & 0.918 & 0.001 & 0.073 & 0.073 & 0.915 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.2$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & -0.003 & 0.281 & 0.282 & 0.836 & 0.005 & 0.188 & 0.188 & 0.920 & 0.012 & 0.170 & 0.170 & 0.942 \\
SW ($K_{2} = 15$) & 0.008 & 0.304 & 0.304 & 0.809 & 0.002 & 0.200 & 0.200 & 0.969 & 0.010 & 0.182 & 0.182 & 0.990 \\
CBGPS & 0.011 & 0.222 & 0.222 & 0.909 & 0.002 & 0.107 & 0.107 & 0.907 & -0.005 & 0.081 & 0.082 & 0.912 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.4$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.002 & 0.288 & 0.288 & 0.830 & 0.008 & 0.184 & 0.185 & 0.917 & 0.003 & 0.174 & 0.174 & 0.932 \\
SW ($K_{2} = 15$) & 0.011 & 0.298 & 0.299 & 0.820 & 0.018 & 0.198 & 0.199 & 0.968 & 0.007 & 0.188 & 0.188 & 0.982 \\
CBGPS & 0.004 & 0.220 & 0.220 & 0.915 & 0.001 & 0.117 & 0.117 & 0.892 & 0.004 & 0.090 & 0.090 & 0.892 \\ \hline \hline
\end{tabular}
}
\end{center}
{\fontsize{9.5pt}{12pt} \selectfont
The DGP is $T_{i} = X_{1i} + X_{2i} + 0.2 X_{3i} + 0.2 X_{4i} + \xi_{i}$ and $Y_{i} = 0.75 X_{1i}^{2} + 0.2 (X_{2i} + 0.5)^{2} + T_{i} + \epsilon_{i}$.
The covariates follow $\boldsymbol{X}_{i} \stackrel{i.i.d.}{\sim} N(0, \boldsymbol{\Sigma})$, where the diagonal elements of $\boldsymbol{\Sigma}$ are all 1 and the off-diagonal elements are all $\rho \in \{ 0.0, 0.2, 0.4 \}$.
We report the bias, standard deviation, root mean squared error, and coverage probability based on the 95\% confidence band across $J = 1000$ Monte Carlo samples.
'SW' signifies our own stabilized-weight estimator, while 'CBGPS' signifies Fong, Hazlett, and Imai's \citeyearpar{Fong_Hazlett_Imai_2018} parametric covariate balancing generalized propensity score (CBGPS) estimator.
For SW, the polynomials are $u_{K_{1}} (t_{i}) = [1, t_{i}, t_{i}^{2}]^{T}$ (i.e. $K_{1} = 3$) and $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}]^{T}$ (i.e. $K_{2} = 9$)
or $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}, X_{1i} X_{2i}, X_{1i} X_{3i}, X_{1i} X_{4i}, X_{2i} X_{3i}, X_{2i} X_{4i}, X_{3i} X_{4i}]^{T}$ (i.e. $K_{2} = 15$).
For CBGPS, covariates are chosen to be $\boldsymbol{X}_{i} = [X_{1i}, X_{2i}, X_{3i}, X_{4i}]^{T}$ for the first step of estimating stabilized weights
and $\boldsymbol{Z}_{i} = [1, T_{i}, \boldsymbol{X}_{i}^{T}]^{T}$ for the second step of estimating average treatment effects.
}
\end{table}
\clearpage
\begin{table}[th]
\begin{center}
\caption{Simulation Results on $\beta_{1}$ (DGP-4) \label{table:sim_results_beta1_DGP4}}
{\fontsize{9.5pt}{16.5pt} \selectfont
\begin{tabular}{c|cccc|cccc|cccc} \hline \hline
\multicolumn{13}{c}{$\rho = 0.0$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & -0.100 & 0.954 & 0.960 & 0.847 & -0.041 & 0.639 & 0.640 & 0.919 & 0.002 & 0.572 & 0.572 & 0.931 \\
SW ($K_{2} = 15$) & -0.105 & 1.050 & 1.055 & 0.852 & 0.010 & 0.690 & 0.690 & 0.956 & 0.072 & 0.716 & 0.720 & 0.975 \\
CBGPS & -0.207 & 0.627 & 0.661 & 0.902 & -0.173 & 0.268 & 0.319 & 0.869 & -0.174 & 0.181 & 0.251 & 0.817 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.2$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & -0.029 & 1.004 & 1.004 & 0.847 & -0.008 & 0.651 & 0.651 & 0.908 & 0.018 & 0.593 & 0.593 & 0.910 \\
SW ($K_{2} = 15$) & -0.109 & 1.126 & 1.131 & 0.839 & 0.004 & 0.727 & 0.727 & 0.965 & 0.033 & 0.675 & 0.676 & 0.956 \\
CBGPS & -0.171 & 0.641 & 0.664 & 0.912 & -0.181 & 0.260 & 0.317 & 0.882 & -0.176 & 0.188 & 0.258 & 0.821 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.4$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & -0.057 & 0.969 & 0.970 & 0.839 & -0.029 & 0.641 & 0.641 & 0.914 & 0.016 & 0.605 & 0.605 & 0.923 \\
SW ($K_{2} = 15$) & -0.043 & 1.091 & 1.092 & 0.846 & 0.047 & 0.726 & 0.728 & 0.935 & 0.081 & 0.734 & 0.738 & 0.958 \\
CBGPS & -0.162 & 0.652 & 0.672 & 0.920 & -0.180 & 0.266 & 0.322 & 0.893 & -0.182 & 0.194 & 0.266 & 0.800 \\ \hline \hline
\end{tabular}
}
\end{center}
{\fontsize{9.5pt}{12pt} \selectfont
The DGP is $T_{i} = (X_{1i} + 0.5)^{2} + 0.4 X_{2i} + 0.4 X_{3i} + 0.4 X_{4i} + \xi_{i}$ and $Y_{i} = 0.75 X_{1i}^{2} + 0.2 (X_{2i} + 0.5)^{2} + T_{i} + \epsilon_{i}$.
The covariates follow $\boldsymbol{X}_{i} \stackrel{i.i.d.}{\sim} N(0, \boldsymbol{\Sigma})$, where the diagonal elements of $\boldsymbol{\Sigma}$ are all 1 and the off-diagonal elements are all $\rho \in \{ 0.0, 0.2, 0.4 \}$.
We report the bias, standard deviation, root mean squared error, and coverage probability based on the 95\% confidence band across $J = 1000$ Monte Carlo samples.
'SW' signifies our own stabilized-weight estimator, while 'CBGPS' signifies Fong, Hazlett, and Imai's \citeyearpar{Fong_Hazlett_Imai_2018} parametric covariate balancing generalized propensity score (CBGPS) estimator.
For SW, the polynomials are $u_{K_{1}} (t_{i}) = [1, t_{i}, t_{i}^{2}]^{T}$ (i.e. $K_{1} = 3$) and $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}]^{T}$ (i.e. $K_{2} = 9$)
or $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}, X_{1i} X_{2i}, X_{1i} X_{3i}, X_{1i} X_{4i}, X_{2i} X_{3i}, X_{2i} X_{4i}, X_{3i} X_{4i}]^{T}$ (i.e. $K_{2} = 15$).
For CBGPS, covariates are chosen to be $\boldsymbol{X}_{i} = [X_{1i}, X_{2i}, X_{3i}, X_{4i}]^{T}$ for the first step of estimating stabilized weights
and $\boldsymbol{Z}_{i} = [1, T_{i}, \boldsymbol{X}_{i}^{T}]^{T}$ for the second step of estimating average treatment effects.
}
\end{table}
\clearpage
\begin{table}[th]
\begin{center}
\caption{Simulation Results on $\beta_{2}$ (DGP-4) \label{table:sim_results_beta2_DGP4}}
{\fontsize{9.5pt}{16.5pt} \selectfont
\begin{tabular}{c|cccc|cccc|cccc} \hline \hline
\multicolumn{13}{c}{$\rho = 0.0$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.033 & 0.288 & 0.290 & 0.810 & 0.038 & 0.187 & 0.191 & 0.913 & 0.040 & 0.165 & 0.170 & 0.930 \\
SW ($K_{2} = 15$) & 0.045 & 0.293 & 0.300 & 0.822 & 0.028 & 0.197 & 0.199 & 0.959 & 0.044 & 0.180 & 0.185 & 0.976 \\
CBGPS & 0.121 & 0.189 & 0.224 & 0.814 & 0.146 & 0.090 & 0.172 & 0.470 & 0.151 & 0.069 & 0.166 & 0.269 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.2$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.030 & 0.280 & 0.282 & 0.801 & 0.025 & 0.173 & 0.175 & 0.905 & 0.025 & 0.162 & 0.164 & 0.918 \\
SW ($K_{2} = 15$) & 0.051 & 0.425 & 0.428 & 0.779 & 0.027 & 0.189 & 0.191 & 0.956 & 0.041 & 0.182 & 0.187 & 0.975 \\
CBGPS & 0.123 & 0.203 & 0.237 & 0.805 & 0.146 & 0.096 & 0.174 & 0.488 & 0.152 & 0.071 & 0.168 & 0.279 \\ \hline \hline
\multicolumn{13}{c}{$\rho = 0.4$} \\ \hline
& \multicolumn{4}{c|}{$N = 100$} & \multicolumn{4}{c|}{$N = 500$} & \multicolumn{4}{c}{$N = 1000$} \\ \hline
& Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP & Bias & Stdev & RMSE & CP \\ \hline
SW ($K_{2} = 9$) & 0.048 & 0.273 & 0.277 & 0.821 & 0.041 & 0.178 & 0.182 & 0.910 & 0.038 & 0.170 & 0.174 & 0.916 \\
SW ($K_{2} = 15$) & 0.046 & 0.301 & 0.304 & 0.789 & 0.025 & 0.192 & 0.194 & 0.946 & 0.039 & 0.182 & 0.186 & 0.965 \\
CBGPS & 0.121 & 0.209 & 0.242 & 0.811 & 0.145 & 0.098 & 0.176 & 0.514 & 0.156 & 0.080 & 0.176 & 0.304 \\ \hline \hline
\end{tabular}
}
\end{center}
{\fontsize{9.5pt}{12pt} \selectfont
The DGP is $T_{i} = (X_{1i} + 0.5)^{2} + 0.4 X_{2i} + 0.4 X_{3i} + 0.4 X_{4i} + \xi_{i}$ and $Y_{i} = 0.75 X_{1i}^{2} + 0.2 (X_{2i} + 0.5)^{2} + T_{i} + \epsilon_{i}$.
The covariates follow $\boldsymbol{X}_{i} \stackrel{i.i.d.}{\sim} N(0, \boldsymbol{\Sigma})$, where the diagonal elements of $\boldsymbol{\Sigma}$ are all 1 and the off-diagonal elements are all $\rho \in \{ 0.0, 0.2, 0.4 \}$.
We report the bias, standard deviation, root mean squared error, and coverage probability based on the 95\% confidence band across $J = 1000$ Monte Carlo samples.
'SW' signifies our own stabilized-weight estimator, while 'CBGPS' signifies Fong, Hazlett, and Imai's \citeyearpar{Fong_Hazlett_Imai_2018} parametric covariate balancing generalized propensity score (CBGPS) estimator.
For SW, the polynomials are $u_{K_{1}} (t_{i}) = [1, t_{i}, t_{i}^{2}]^{T}$ (i.e. $K_{1} = 3$) and $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}]^{T}$ (i.e. $K_{2} = 9$)
or $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}, X_{1i} X_{2i}, X_{1i} X_{3i}, X_{1i} X_{4i}, X_{2i} X_{3i}, X_{2i} X_{4i}, X_{3i} X_{4i}]^{T}$ (i.e. $K_{2} = 15$).
For CBGPS, covariates are chosen to be $\boldsymbol{X}_{i} = [X_{1i}, X_{2i}, X_{3i}, X_{4i}]^{T}$ for the first step of estimating stabilized weights
and $\boldsymbol{Z}_{i} = [1, T_{i}, \boldsymbol{X}_{i}^{T}]^{T}$ for the second step of estimating average treatment effects.
}
\end{table}
\clearpage
\section{Assumptions}
\begin{assumption}[\emph{Unconfounded Treatment Assignment}]
\label{as:TYindep} For all $t\in \mathcal{T}$, given $\boldsymbol{X}$ , $T$
is independent of $Y^{\ast }(t)$, i.e., $Y^{\ast }(t)\perp T|\boldsymbol{X,}$
for all $t\in \mathcal{T}$.
\end{assumption}
\begin{assumption}
\label{as:suppX} The support $\mathcal{X}$ of $\boldsymbol{X}$ is a compact
subset of $\mathbb{R}^{r}$. The support $\mathcal{T}$ of the treatment
variable $T$ is a compact subset of $\mathbb{R}$.
\end{assumption}
\begin{assumption}
\label{as:pi0} There exist two positive constants $\eta_1$ and $\eta_2$ such
that
\begin{equation*}
0 < \eta_1 \leq \pi_0(t,\boldsymbol{x}) \leq \eta_2 <\infty \ ,\ \forall (t,
\boldsymbol{x}) \in \mathcal{T}\times\mathcal{X}\ .
\end{equation*}
\end{assumption}
\begin{assumption}
\label{as:smooth_pi} There exist $\Lambda _{K_{1}\times K_{2}}\in \mathbb{R}
^{K_{1}\times K_{2}}$ and a positive constant $\alpha >0$ such that
\begin{equation*}
\sup_{(t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X}}\left\vert (\rho
^{\prime -1}\left( \pi _{0}(t,\boldsymbol{x})\right) -u_{K_{1}}(t)^{\top
}\Lambda _{K_{1}\times K_{2}}v_{K_{2}}(\boldsymbol{x})\right\vert
=O(K^{-\alpha }).
\end{equation*}
\end{assumption}
\begin{assumption}
\label{as:u&v} For every $K_1$ and $K_2$, the smallest eigenvalues of $
\mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^\top\right]$ and $\mathbb{E}\left[
v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^\top\right]$ are bounded away
from zero uniformly in $K_1$ and $K_2$.
\end{assumption}
\begin{assumption}
\label{as:K&N_consistency}There are two sequences of constants $\zeta
_{1}(K_{1})$ and $\zeta _{2}(K_{2})$ satisfying\newline
$\sup_{t\in \mathcal{T}}\Vert u_{K_{1}}(t)\Vert \leq \zeta _{1}(K_{1})$ and $
\sup_{\boldsymbol{x}\in \mathcal{X}}\Vert v_{K_{2}}(\boldsymbol{x})\Vert
\leq \zeta _{2}(K_{2})$, $K=K_{1}(N)K_{2}(N)$ and $\zeta (K)=\zeta
_{1}(K_{1})\zeta _{2}(K_{2})$, such that $\zeta (K)K^{-\alpha }\rightarrow 0$
and $\zeta (K)\sqrt{K/N}\rightarrow 0$ as $N\rightarrow \infty $.
\end{assumption}
\begin{assumption}
\label{as:Theta} The parameter space $\Theta\subset\mathbb{R} ^{p}$ is a
compact set and the true parameter $\boldsymbol{\beta}_0$ is in the interior
of $\Theta$ , where $p\in\mathbb{N}$.
\end{assumption}
\begin{assumption}
\label{as:solution} There exists a unique solution $\boldsymbol{\beta }_{0}$
for the optimization problem
\begin{equation*}
\min_{\boldsymbol{\beta }\in \Theta }\int_{\mathcal{T}}\mathbb{E}\left[
L(Y^{\ast }(t)-g(t;\boldsymbol{\beta }))\right] dF_{T}(t)\ .
\end{equation*}
\end{assumption}
\begin{assumption}
\label{as:EY2} $\mathbb{E}\left[\sup_{\boldsymbol{\beta} \in
\Theta }|L\left( Y-g(T;\boldsymbol{\beta} )\right)|^{2}\right]<\infty $.
\end{assumption}
\begin{assumption}
\label{as:m_smooth} The following conditions hold true:
\begin{enumerate}
\item $g(t;\boldsymbol{\beta })$ is twice continuously differentiable in $
\boldsymbol{\beta }\in \Theta $;
\item $L(Y-g(T;\boldsymbol{\beta }))$ is differentiable in $\boldsymbol{\beta }$ with probability one, i.e., for any directional vector $\boldsymbol{
\ \eta }\in \mathbb{R}^{p}$, there exists an integrable random variable $
L^{\prime }(Y-g(T;\boldsymbol{\beta }))$ such that{\footnotesize \
\begin{equation*}
\mathbb{P}\left( \lim_{\epsilon \rightarrow 0}\frac{L(Y-g(T;\boldsymbol{\
\beta }+\epsilon \boldsymbol{\eta }))-L(Y-g(T;\boldsymbol{\beta }))}{
\epsilon }=L^{\prime }(Y-g(T;\boldsymbol{\beta }))\cdot \left\langle m(T;
\boldsymbol{\beta }),\boldsymbol{\eta }\right\rangle _{\mathbb{R}
^{p}}\right) =1,
\end{equation*}
} where $\left\langle \cdot ,\cdot \right\rangle _{\mathbb{R}^{p}}$ is the
inner product in Euclidean space $\mathbb{R}^{p}$;
\item $\mathbb{E}\left[ L^{\prime }(Y-g(T;\boldsymbol{\beta }_{0}))^{2}
\right] <\infty $.
\end{enumerate}
\end{assumption}
\begin{assumption}
\label{as:first_order} Suppose that
\begin{equation*}
\frac{1}{N}\sum_{i=1}^{N}\hat{\pi}_{K}(T_{i},\boldsymbol{X}_{i})L^{\prime
}\left( Y_{i}-g(T_{i};\hat{\boldsymbol{\beta }})\right) m(T_{i};\hat{
\boldsymbol{\beta }})=0
\end{equation*}
holds with probability approaching one.
\end{assumption}
\begin{assumption}
\label{as:H} $\mathbb{E}\left[ \pi _{0}(T,\boldsymbol{X})L^{\prime }(Y-g(T;
\boldsymbol{\beta }))m(T;\boldsymbol{\beta })\right] $ is differentiable
with respect to $\boldsymbol{ \beta}$ and $H_{0}:=-\nabla _{\beta }\mathbb{E}\left[ \pi
_{0}(T,\boldsymbol{X})L^{\prime }(Y-g(T;\boldsymbol{\beta }))m(T;\boldsymbol{
\beta })\right] \Big|_{\boldsymbol{\beta }=\boldsymbol{\beta }_{0}}$ is
nonsingular.
\end{assumption}
\begin{assumption}
\label{as:smooth_varepsilon} $\varepsilon (t,\boldsymbol{x};{\boldsymbol{
\beta }}_{0}):=\mathbb{E}[L^{\prime }(Y-g(T;\boldsymbol{\beta }_{0}))|T=t,
\boldsymbol{X}=\boldsymbol{x}]$ is continuously differentiable in $(t,
\boldsymbol{x})$.
\end{assumption}
\begin{assumption}
\label{as:entropy} \
\begin{enumerate}
\item $\mathbb{E}\left[ \sup_{\boldsymbol{\beta} \in \Theta }|L^{\prime }(Y-g(T;\boldsymbol{\beta}
))^{2+\delta }\right] <\infty $ for some $\delta >0$;
\item The function class $\{L^{\prime }(y-g(t;\boldsymbol{\beta} )):\boldsymbol{\beta} \in \Theta \}$
satisfies:
\begin{equation*}
\mathbb{E}\left[ \sup_{\boldsymbol{\beta} _{1}:\Vert \boldsymbol{\beta}_{1}-\boldsymbol{\beta} \Vert <\delta
}\left\vert L^{\prime }(Y-g(T;\boldsymbol{\beta} _{1}))-L^{\prime }(Y-g(T;\boldsymbol{\beta}
))\right\vert ^{2}\right] ^{1/2}\leq a\cdot \delta ^{b}
\end{equation*}
for any $\forall \boldsymbol{\beta}\in \Theta $ and any small $\delta >0$ and for some
finite positive constants $a$ and $b$.
\end{enumerate}
\end{assumption}
\begin{assumption}
\label{as:K&N_c} $\zeta (K)\sqrt{K^{4}/N}\rightarrow 0$ and $\sqrt{N}
K^{-\alpha}\to 0$ as $N\to \infty$.
\end{assumption}
\section{Efficiency Bound}
\subsection{Proof of Theorem 3.1 \label{sec:proof_efficiency_theorem}}
Without loss of generality, we only consider the distribution of $(T,\boldsymbol{X},Y)$ to be absolutely continuous with respect to Lebesgue measure, i.e., there exists a density function $f_{T,X,Y}(t,\boldsymbol{x},y)$ such that $dF_{T,X,Y}(t,\boldsymbol{x},y)=f_{T,X,Y}(t,\boldsymbol{x},y)dtd\boldsymbol{x}dy$. For discrete cases, the proof can be established by using a similar argument.\\
We follow the approach of \citet[Section 3.3]{Bickel_Klaassen_Ritov_Wellner_1993} to derive the variance bound of $\boldsymbol{\beta}_0$, see also \cite{tchetgen2012semiparametric}. Let $\left\{f^{\alpha}_{Y,T,X}(y,t,\boldsymbol{x})\right\}_{\alpha \in\mathbb{R}}$ denote a one dimensional regular parametric submodel with $
f^{\alpha=0}_{Y,T,X}(y,t,\boldsymbol{x})=f_{Y,T,X}(y,t,\boldsymbol{x})$.
By definition, $\boldsymbol{\beta}_0$ solves following equation:
\begin{align}\label{def_beta^*_eff_0}
\int_{\mathcal{T}}\mathbb{E}\left[ m(t;\boldsymbol{\beta}_0)L'\left(Y^*(t)-g\left(t;\boldsymbol{\beta}_0\right)\right)\right]f_T(t)dt=0 \ .
\end{align}
By Assumption \ref{as:TYindep}, \eqref{def_beta^*_eff_0} is equivalent to
\begin{align} \notag
\int_{\mathcal{T}} \int_{\mathcal{X}} \mathbb{E}\left[m(T;\boldsymbol{\beta}_0)L'\left(Y-g\left(T;\boldsymbol{\beta}_0\right)\right)|T=t,\boldsymbol{X}=\boldsymbol{x}\right]f_{X}(\boldsymbol{x}) f_T(t)d\boldsymbol{x}dt=0 \ .
\end{align}
Therefore, the parameter $\boldsymbol{\beta}(\alpha)$ induced by the submodel $f^{\alpha}_{Y,T,X}(y,t,\boldsymbol{x})$ satisfies:
\begin{align}
&\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}(\alpha))\cdot \mathbb{E}^{\alpha}\left[L'\left(Y-g\left(t;\boldsymbol{\beta}(\alpha)\right)\right)|T=t,\boldsymbol{X}=\boldsymbol{x}\right]f^{{\alpha}}_{T}(t) f^{{\alpha}}_{X}(\boldsymbol{x})d\boldsymbol{x}dt=0\ , \label{perturbated_model}
\end{align}
where $\mathbb{E}^{\alpha}\left[\cdot|T=t,\boldsymbol{X}=\boldsymbol{x}\right]$ denotes taking expectation with respect to the submodel $f^{\alpha}_{Y|T,X}(\cdot|t,\boldsymbol{x})$. \\
Differentiating both sides of \eqref{perturbated_model} with respect to $\alpha$, evaluating at $\alpha = 0$ and using the condition $Y^*(t)\perp T|\boldsymbol{X}$, we can deduce that
\begin{align*}
0=&\int_{\mathcal{T}}\int_{\mathcal{X}} \frac{\partial}{\partial \alpha} \bigg|_{\alpha=0} \left\{m(t;\boldsymbol{\beta}(\alpha))\mathbb{E}^{\alpha}\left[L'(Y-g(t;\boldsymbol{\beta}(\alpha)))|T=t,\boldsymbol{X}=\boldsymbol{x}\right] f^{{\alpha}}_{T}(t)f^{{\alpha}}_{X}(\boldsymbol{x})\right\}d\boldsymbol{x}dt\\
=&\int_{\mathcal{T}}\int_{\mathcal{X}} \mathbb{E}\left[L'(Y-g(t;\boldsymbol{\beta}_0))|T=t,\boldsymbol{X}=\boldsymbol{x}\right] f_{T}(t)f_{X}(\boldsymbol{x})\nabla_{\boldsymbol{\beta}}m(t;\boldsymbol{\beta}_0) d\boldsymbol{x}dt\cdot \frac{\partial}{\partial \alpha} \bigg|_{\alpha=0}\boldsymbol{\beta}(\alpha) \\
&+\int_{\mathcal{X}\times \mathcal{T}} \mathbb{E}[L'(Y-g(t;\boldsymbol{\beta}_0))|T=t,\boldsymbol{X}=\boldsymbol{x}]m(t;\boldsymbol{\beta}_0)\cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{X}(\boldsymbol{x})\bigg|_{\alpha=0}f_{T}(t)d\boldsymbol{x}dt \notag\\
&+\int_{\mathcal{Y}\times \mathcal{X}\times \mathcal{T}} m(t;\boldsymbol{\beta}_0)L'(y-g(t;\boldsymbol{\beta}_0)) \cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{Y|T,X}(y|t,\boldsymbol{x})\bigg|_{\alpha=0}f_{X}(\boldsymbol{x})f_T(t)dyd\boldsymbol{x}dt \notag \\
& + \int_{\mathcal{X}\times \mathcal{T}}m(t;\boldsymbol{\beta}_0) \cdot \nabla_{\beta}\mathbb{E}[L'(Y^*(t)-g(t;\boldsymbol{\beta}))|T=t,\boldsymbol{X}=\boldsymbol{x}]\Bigg|_{\beta=\boldsymbol{\beta}_0} \cdot \frac{\partial}{\partial \alpha} \bigg|_{\alpha=0}\boldsymbol{\beta}(\alpha)\cdot f_{T}(t)f_{X}(\boldsymbol{x})d\boldsymbol{x}dt \\
&+ \int_{\mathcal{X}\times \mathcal{T}} \mathbb{E}[L'(Y-g(t;\boldsymbol{\beta}_0))|T=t,\boldsymbol{X}=\boldsymbol{x}]m(t;\boldsymbol{\beta}_0)\cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{T}(t)\bigg|_{\alpha=0}f_{X}(\boldsymbol{x})d\boldsymbol{x}dt \\
=&\int_{\mathcal{T}}\int_{\mathcal{X}} \mathbb{E}\left[L'(Y^*(t)-g(t;\boldsymbol{\beta}_0))|\boldsymbol{X}=\boldsymbol{x}\right] f_{T}(t)f_{X}(\boldsymbol{x})\nabla_{\boldsymbol{\beta}}m(t;\boldsymbol{\beta}_0) d\boldsymbol{x}dt\cdot \frac{\partial}{\partial \alpha} \bigg|_{\alpha=0}\boldsymbol{\beta}(\alpha) \\
&+\int_{\mathcal{X}\times \mathcal{T}} \mathbb{E}[L'(Y-g(t;\boldsymbol{\beta}_0))|T=t,\boldsymbol{X}=\boldsymbol{x}]m(t;\boldsymbol{\beta}_0)\cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{X}(\boldsymbol{x})\bigg|_{\alpha=0}f_{T}(t)d\boldsymbol{x}dt \notag\\
& +\int_{\mathcal{Y}\times \mathcal{X}\times \mathcal{T}} m(t;\boldsymbol{\beta}_0)L'(y-g(t;\boldsymbol{\beta}_0)) \cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{Y|T,X}(y|t,\boldsymbol{x})\bigg|_{\alpha=0}f_{X}(\boldsymbol{x})f_T(t)dyd\boldsymbol{x}dt \notag \\
& + \int_{\mathcal{X}\times \mathcal{T}}m(t;\boldsymbol{\beta}_0) \cdot \nabla_{\beta}\mathbb{E}[L'(Y^*(t)-g(t;\boldsymbol{\beta}))|\boldsymbol{X}=\boldsymbol{x}]\Bigg|_{\boldsymbol{\beta}=\boldsymbol{\beta}_0} \cdot f_{T}(t)f_{X}(\boldsymbol{x})d\boldsymbol{x}dt \cdot \frac{\partial}{\partial \alpha} \bigg|_{\alpha=0}\boldsymbol{\beta}(\alpha) \\
& + \int_{\mathcal{X}\times \mathcal{T}} \mathbb{E}[L'(Y-g(t;\boldsymbol{\beta}_0))|T=t,\boldsymbol{X}=\boldsymbol{x}]m(t;\boldsymbol{\beta}_0)\cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{T}(t)\bigg|_{\alpha=0}f_{X}(\boldsymbol{x})d\boldsymbol{x}dt \\
=&\int_{\mathcal{T}} \mathbb{E}[L'(Y^*(t)-g(t;\boldsymbol{\beta}_0))]\cdot f_{T}(t)\nabla_{\boldsymbol{\beta}}m(t;\boldsymbol{\beta}_0) dt\cdot \frac{\partial}{\partial \alpha} \bigg|_{\alpha=0}\boldsymbol{\beta}(\alpha) \\
&+\int_{\mathcal{X}\times \mathcal{T}} \mathbb{E}[L'(Y-g(t;\boldsymbol{\beta}_0))|T=t,\boldsymbol{X}=\boldsymbol{x}]m(t;\boldsymbol{\beta}_0)\cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{X}(\boldsymbol{x})\bigg|_{\alpha=0}f_{T}(t)d\boldsymbol{x}dt \notag\\
& +\int_{\mathcal{Y}\times \mathcal{X}\times \mathcal{T}} m(t;\boldsymbol{\beta}_0)\cdot L'(y-g(t;\boldsymbol{\beta}_0)) \cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{Y|T,X}(y|t,\boldsymbol{x})\bigg|_{\alpha=0}f_{X}(\boldsymbol{x})f_T(t)dyd\boldsymbol{x}dt \notag \\
& +\int_{\mathcal{T}} \nabla_{\beta}\mathbb{E}[L'(Y^*(t)-g(t;\boldsymbol{\beta}))]\bigg|_{\boldsymbol{\beta}=\boldsymbol{\beta}_0}m(t;\boldsymbol{\beta}_0) \cdot f_{T}(t)dt \cdot \frac{\partial}{\partial \alpha} \bigg|_{\alpha=0}\boldsymbol{\beta}(\alpha) \\
& + \int_{\mathcal{X}\times \mathcal{T}} \mathbb{E}[L'(Y-g(t;\boldsymbol{\beta}_0))|T=t,\boldsymbol{X}=\boldsymbol{x}]m(t;\boldsymbol{\beta}_0)\cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{T}(t)\bigg|_{\alpha=0}f_{X}(\boldsymbol{x})d\boldsymbol{x}dt\\
=&\nabla_{\beta}\left\{\int_{\mathcal{T}} \mathbb{E}[L'(Y^*(t)-g(t;\boldsymbol{\beta}))]\cdot m(t;\boldsymbol{\beta})f_{T}(t) dt\right\}\Bigg|_{\boldsymbol{\beta}=\boldsymbol{\beta}_0}\cdot \frac{\partial}{\partial \alpha} \bigg|_{\alpha=0}\boldsymbol{\beta}(\alpha) \\
&+\int_{\mathcal{X}\times \mathcal{T}} \mathbb{E}[L'(Y-g(t;\boldsymbol{\beta}_0))|T=t,\boldsymbol{X}=\boldsymbol{x}]m(t;\boldsymbol{\beta}_0)\cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{X}(\boldsymbol{x})\bigg|_{\alpha=0}f_{T}(t)d\boldsymbol{x}dt \notag\\
& +\int_{\mathcal{Y}\times \mathcal{X}\times \mathcal{T}} m(t;\boldsymbol{\beta}_0)\cdot L'(y-g(t;\boldsymbol{\beta}_0)) \cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{Y|T,X}(y|t,\boldsymbol{x})\bigg|_{\alpha=0}f_{X}(\boldsymbol{x})f_T(t)dyd\boldsymbol{x}dt \notag \\
& + \int_{\mathcal{X}\times \mathcal{T}} \mathbb{E}[L'(Y-g(t;\boldsymbol{\beta}_0))|T=t,\boldsymbol{X}=\boldsymbol{x}]m(t;\boldsymbol{\beta}_0)\cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{T}(t)\bigg|_{\alpha=0}f_{X}(\boldsymbol{x})d\boldsymbol{x}dt.
\end{align*}
Since $H_0=- \nabla_{\beta}\left\{\int_{\mathcal{T}} \mathbb{E}[L'(Y^*(t)-g(t;\boldsymbol{\beta}))]\cdot m(t;\boldsymbol{\beta})f_{T}(t) dt\right\}\Bigg|_{\boldsymbol{\beta}=\boldsymbol{\beta}_0}$ is invertible by Assumption \ref{as:EY2}, we get
\begin{align}
&\frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}\boldsymbol{\beta}(\alpha) =H_0^{-1}\cdot \bigg\{\int_{\mathcal{X}\times \mathcal{T}} \mathbb{E}[L'(Y-g(t;\boldsymbol{\beta}_0))|T=t,\boldsymbol{X}=\boldsymbol{x}]m(t;\boldsymbol{\beta}_0)\cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{X}(\boldsymbol{x})\bigg|_{\alpha=0}f_{T}(t)d\boldsymbol{x}dt \notag\\
&\qquad +\int_{\mathcal{Y}\times \mathcal{X}\times \mathcal{T}} m(t;\boldsymbol{\beta}_0) \cdot L'(y-g(t;\boldsymbol{\beta}_0)) \cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{Y|T,X}(y|t,\boldsymbol{x})\bigg|_{\alpha=0}f_{X}(\boldsymbol{x})f_T(t)dyd\boldsymbol{x}dt \notag \\
&\qquad + \int_{\mathcal{X}\times \mathcal{T}} \mathbb{E}[L'(Y-g(t;\boldsymbol{\beta}_0))|T=t,\boldsymbol{X}=\boldsymbol{x}]m(t;\boldsymbol{\beta}_0)\cdot \frac{\partial }{\partial \alpha}f^{\alpha}_{T}(t)\bigg|_{\alpha=0}f_{X}(\boldsymbol{x})d\boldsymbol{x}dt \bigg\} \notag \ .
\end{align}
The efficient influence function of $\boldsymbol{\beta}_0$, denoted by $S_{eff}(Y,T,\boldsymbol{X};\boldsymbol{\beta}_0)$, is a unique function satisfying the following equation:
\begin{align}\label{def_efficient_influence}
\frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}\boldsymbol{\beta}(\alpha)=\mathbb{E}\left[S_{eff}(Y,T,\boldsymbol{X};\boldsymbol{\beta}_0)\frac{\partial}{\partial \alpha}\bigg|_{\alpha=0}\log f^{\alpha}_{Y,X,T}(Y,\boldsymbol{X},T)\right] \ .
\end{align}
Therefore, to justify our theorem, it suffices to substitute
$S_{eff}(Y,T,\boldsymbol{X};\boldsymbol{\beta}_0) = H_0 ^{-1} \psi(Y,T,\boldsymbol{X};\boldsymbol{\beta}_0)$
into \eqref{def_efficient_influence} and check the validity.
Note that
\begin{align}
&\mathbb{E}\left[S_{eff}(Y,T,\boldsymbol{X};\boldsymbol{\beta}_0)\frac{\partial}{\partial \alpha}\bigg|_{\alpha=0}\log f^{\alpha}_{Y,X,T}(Y,\boldsymbol{X},T)\right] \notag\\
=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}\times \mathcal{Y}}\psi(y,t,\boldsymbol{x};\boldsymbol{\beta}_0)\frac{\partial }{\partial \alpha}\bigg|_{\alpha=0} f^{\alpha}_{Y|X,T}(y|\boldsymbol{x},t)f_{T,X}(t,\boldsymbol{x})dyd\boldsymbol{x}dt \label{eff_1}\\
+&H_0^{-1}\int_{\mathcal{X}\times\mathcal{T}\times \mathcal{Y}}\psi(y,t,\boldsymbol{x};\boldsymbol{\beta}_0) f_{Y|X,T}(y|\boldsymbol{x},t)\frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}f^{\alpha}_{T|X}(t|\boldsymbol{x})f_X(\boldsymbol{x})dyd\boldsymbol{x}dt \label{eff_2}\\
+&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}\times \mathcal{Y}}\psi(y,t,\boldsymbol{x};\boldsymbol{\beta}_0) f_{Y|X,T}(y|\boldsymbol{x},t)f_{T|X}(t|\boldsymbol{x})\frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}f^{\alpha}_X(\boldsymbol{x})dyd\boldsymbol{x}dt. \label{eff_3} \end{align}
For the term \eqref{eff_1}, we have
\begin{align*}
\eqref{eff_1}=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}\times \mathcal{Y}} \bigg\{ \frac{f_T(t)}{f_{T|X}(t|\boldsymbol{x})}m(t;\boldsymbol{\beta}_0) \cdot L'(y-g(t;\boldsymbol{\beta}_0))-\frac{f_T(t)}{f_{T|X}(t|\boldsymbol{x})}m(t;\boldsymbol{\beta}_0) \cdot\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\\
&\qquad +\mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|\boldsymbol{X}=\boldsymbol{x}\right]+\mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|T=t\right]\bigg\} \\
& \times \frac{\partial }{\partial \alpha}\bigg|_{\alpha=0} f^{\alpha}_{Y|X,T}(y|\boldsymbol{x},t)f_{T,X}(t,\boldsymbol{x})dyd\boldsymbol{x}dt\\
=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}\times \mathcal{Y}} \frac{f_T(t)}{f_{T|X}(t|\boldsymbol{x})}m(t;\boldsymbol{\beta}_0) \cdot L'(y-g(t;\boldsymbol{\beta}_0))\cdot \frac{\partial }{\partial \alpha}\bigg|_{\alpha=0} f^{\alpha}_{Y|X,T}(y|\boldsymbol{x},t)f_{T,X}(t,\boldsymbol{x})dyd\boldsymbol{x}dt\\
=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}\times \mathcal{Y}}m(t;\boldsymbol{\beta}_0)\cdot L'(y-g(t;\boldsymbol{\beta}_0))\cdot \frac{\partial }{\partial \alpha}\bigg|_{\alpha=0} f^{\alpha}_{Y|X,T}(y|\boldsymbol{x},t)f_{T}(t)f_{X}(\boldsymbol{x})dyd\boldsymbol{x}dt.
\end{align*}
For the term \eqref{eff_2}, we have
\begin{align*}
\eqref{eff_2}=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}\times \mathcal{Y}} \bigg\{ \frac{f_T(t)}{f_{T|X}(t|\boldsymbol{x})}m(t;\boldsymbol{\beta}_0) \cdot L'(y-g(t;\boldsymbol{\beta}_0))-\frac{f_T(t)}{f_{T|X}(t|\boldsymbol{x})}m(t;\boldsymbol{\beta}_0) \cdot\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\\
&\qquad +\mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|\boldsymbol{X}=\boldsymbol{x}\right]+\mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|T=t\right]\bigg\} \\
& \times f_{Y|X,T}(y|\boldsymbol{x},t)\frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}f^{\alpha}_{T|X}(t|\boldsymbol{x})f_{X}(\boldsymbol{x})dyd\boldsymbol{x}dt\\
=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}} \bigg\{\mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|\boldsymbol{X}=\boldsymbol{x}\right]+\mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|T=t\right]\bigg\} \\
& \qquad \cdot \frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}f^{\alpha}_{T|X}(t|\boldsymbol{x})f_{X}(\boldsymbol{x})d\boldsymbol{x}dt\\
=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}} \mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|T=t\right]\cdot \frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}f^{\alpha}_{T|X}(t|\boldsymbol{x})f_{X}(\boldsymbol{x})d\boldsymbol{x}dt\\
=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}} \mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|T=t\right]\cdot \frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}f^{\alpha}_{T}(t)dt\\
=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}} \varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\frac{f_{T}(t)}{f_{T|X}(t|\boldsymbol{x})}m(t;\boldsymbol{\beta}_0)\cdot \frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}f^{\alpha}_{T}(t)\cdot f_{X|T}(\boldsymbol{x}|t)d\boldsymbol{x}dt\\
=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}} \varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0) m(t;\boldsymbol{\beta}_0)\cdot \frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}f^{\alpha}_{T}(t)\cdot f_{X}(\boldsymbol{x})d\boldsymbol{x}dt,
\end{align*}
where the first equality holds in accordance with the definition of $\int_{\mathcal{Y}}L'(y-g(t;\boldsymbol{\beta}_0))f_{Y|X,T}(y|\boldsymbol{x},t)dy=:\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)$. \\
For the term \eqref{eff_3}, we have
\begin{align*}
\eqref{eff_3}=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}\times \mathcal{Y}} \bigg\{ \frac{f_T(t)}{f_{T|X}(t|\boldsymbol{x})}m(t;\boldsymbol{\beta}_0) \cdot L'(y-g(t;\boldsymbol{\beta}_0))-\frac{f_T(t)}{f_{T|X}(t|\boldsymbol{x})}m(t;\boldsymbol{\beta}_0) \cdot\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\\
&\qquad +\mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|\boldsymbol{X}=\boldsymbol{x}\right]+\mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|T=t\right]\bigg\} \\
&\times f_{Y|X,T}(y|\boldsymbol{x},t)f_{T|X}(t|\boldsymbol{x})\frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}f^{\alpha}_{X}(\boldsymbol{x})dyd\boldsymbol{x}dt\\
=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}} \bigg\{\mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|\boldsymbol{X}=\boldsymbol{x}\right]+\mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|T=t\right]\bigg\} \\
&\times f_{T|X}(t|\boldsymbol{x})\cdot \frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}f^{\alpha}_{X}(\boldsymbol{x})d\boldsymbol{x}dt\\
=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}} \mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|\boldsymbol{X}=\boldsymbol{x}\right]\cdot f_{T|X}(t|\boldsymbol{x})\cdot \frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}f^{\alpha}_{X}(\boldsymbol{x})d\boldsymbol{x}dt\\
=&H_0 ^{-1}\int_{\mathcal{X} } \mathbb{E}\left[\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|\boldsymbol{X}=\boldsymbol{x}\right]\cdot \frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}f^{\alpha}_{X}(\boldsymbol{x})d\boldsymbol{x}\\
=&H_0 ^{-1}\int_{\mathcal{X}\times\mathcal{T}} \varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)m(t;\boldsymbol{\beta}_0) \cdot f_{T}(t)\cdot \frac{\partial }{\partial \alpha}\bigg|_{\alpha=0}f^{\alpha}_{X}(\boldsymbol{x})d\boldsymbol{x}dt.
\end{align*}
We have proved \eqref{def_efficient_influence} holds, hence $S_{eff}$ is the efficient influence function of $\boldsymbol{\beta}_0$.
\subsection{Particular Case I: Binary Treatment Effects \label{eff_bound:binary}}
In this section, we show that when $T\in\{0,1\}$, $g(t;\boldsymbol{\beta})=\beta_0+\beta_1\cdot t$ and $L(v)=v^2$, our general efficiency bound derived in Theorem 3.1 reduces to the well-known efficiency bound for average treatment effects in \cite{robins1994estimation} and \cite{hahn1998role}. In accordance with our identification condition, $\beta_0$ and $\beta_1$ are identified by minimizing the following loss function
$$\sum_{t\in\{0,1\}}\mathbb{E}[(Y^*(t)-\beta_0-\beta_1\cdot t)^2]\cdot \mathbb{P}(T=t).$$
The solutions are given by
$$\beta_0=\mathbb{E}[Y^*(0)], \ \beta_1=\mathbb{E}[Y^*(1)-Y^*(0)].$$
Here $\beta_1$ is the average treatment effects.
\begin{cor}\label{cor_effbound_binary}
Suppose $T\in\{0,1\}$, $L(v)=v^2$, $g(t;\boldsymbol{\beta})=\beta_0+\beta_1\cdot t$ and the conditions in Theorem 3.1 hold, the efficient influence functions of $\beta_0$ and $\beta_1$ given by Theorem 3.1 reduce to
\begin{align*}
&S_{eff}(T,\boldsymbol{X},Y;\beta_0)=\phi_2(T,\bold{X},Y;\beta_0), \\
&S_{eff}(T,\boldsymbol{X},Y;\beta_1,\beta_0)=\phi_2(T,\bold{X},Y;\beta_0)-\phi_1(T,\bold{X},Y;\beta_1,\beta_0),
\end{align*}
where
\begin{align*}
&\phi_1(T,\boldsymbol{X},Y;\boldsymbol{\beta})=\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot Y^*(1)- \left\{\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}-1\right\}\cdot \mathbb{E}[Y^*(1)|\boldsymbol{X}] -\beta_0-\beta_1, \\
&\phi_2(T,\boldsymbol{X},Y;\boldsymbol{\beta})=\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot Y^*(0)- \left\{\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}-1\right\}\cdot \mathbb{E}[Y^*(0)|\boldsymbol{X}] -\beta_0,
\end{align*}
and they are the same as the efficient influence functions given in \cite{robins1994estimation} and \cite{hahn1998role}.
\end{cor}
\begin{proof}
Using our notation, we have
\begin{align*}
&\boldsymbol{\beta}_0=(\beta_0,\beta_1)^{\top}\ ,\ g(t;\boldsymbol{\beta}_0)=\beta_0+\beta_1 \cdot t , \quad m(t;\boldsymbol{\beta}_0)=\begin{bmatrix}
1 \\ t
\end{bmatrix}\ ,\ H_0=\mathbb{E}\left[m(T;\boldsymbol{\beta}_0)m(T;\boldsymbol{\beta}_0)^{\top}\right]\ ,\\ &\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)=T\cdot\left\{ \mathbb{E}[Y^*(1)-Y^*(0)|\boldsymbol{X}]-\beta_1\right\}+ \mathbb{E}[Y^*(0)|\boldsymbol{X}]-\beta_0, \\
&\pi_0(T,\boldsymbol{X})=\frac{T\cdot p+(1-T)\cdot q}{T\cdot \mathbb{P}(T=1|\boldsymbol{X})+T\cdot \mathbb{P}(T=0|\boldsymbol{X})}=\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p +\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q,
\end{align*}
where $p=\mathbb{P}(T=1)$ and $q=\mathbb{P}(T=0)$.
In accordance with our Theorem 3.1, the efficient influence function of $(\beta_0,\beta_1)$ is
\begin{align*}
H^{-1}_0\bigg\{\pi_0 (T,
\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\left\{ Y-\mathbb{E}[Y|
\boldsymbol{X},T]\right\} +\mathbb{E}\left[ \varepsilon (T,\boldsymbol{X};
\boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|
\boldsymbol{X}\right]\bigg\}\ .
\end{align*}
With some computation, we have
\begin{align} \label{m_inverse}
H^{-1}_0 =\begin{bmatrix}
1& p \\
p& p
\end{bmatrix}^{-1}=\frac{1}{pq}\cdot \begin{bmatrix}
p& -p \\
-p& 1
\end{bmatrix}=\begin{bmatrix}
\frac{1}{q}& -\frac{1}{q} \\
-\frac{1}{q}& \frac{1}{pq}
\end{bmatrix}\ .
\end{align}
and
\begin{align}
&\pi_0 (T,
\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\left\{ Y-\mathbb{E}[Y|
\boldsymbol{X},T]\right\} \notag\\
=&\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p \cdot \begin{bmatrix}
1 \\ T
\end{bmatrix}\cdot\bigg\{Y-T\cdot \mathbb{E}[Y^*(1)|\boldsymbol{X}]-(1-T)\cdot \mathbb{E}[Y^*(0)|\boldsymbol{X}]\bigg\} \notag\\
&+\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q\cdot \begin{bmatrix}
1 \\ T
\end{bmatrix}\cdot\bigg\{Y-T\cdot \mathbb{E}[Y^*(1)|\boldsymbol{X}]-(1-T)\cdot \mathbb{E}[Y^*(0)|\boldsymbol{X}]\bigg\} \notag\\
=&\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p \cdot \begin{bmatrix}
1 \\ 1
\end{bmatrix}\cdot\bigg\{Y^*(1)- \mathbb{E}[Y^*(1)|\boldsymbol{X}]\bigg\}
+\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q \cdot \begin{bmatrix}
1 \\ 0
\end{bmatrix}\cdot\bigg\{Y^*(0)- \mathbb{E}[Y^*(0)|\boldsymbol{X}]\bigg\} \notag\\
=&\begin{bmatrix}
\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot \left\{Y^*(1)- \mathbb{E}[Y^*(1)|\boldsymbol{X}]\right\}\cdot p+\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot \left\{Y^*(0)- \mathbb{E}[Y^*(0)|\boldsymbol{X}]\right\}\cdot q \\[4mm]
\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot \left\{Y^*(1)- \mathbb{E}[Y^*(1)|\boldsymbol{X}]\right\}\cdot p
\end{bmatrix} \label{eff:pi}
\end{align}
and
\begin{align}
&\mathbb{E}\left[ \varepsilon (T,\boldsymbol{X};
\boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|
\boldsymbol{X}\right] \notag\\
=&\mathbb{E}\Bigg[\bigg(T\cdot\left\{ \mathbb{E}[Y^*(1)-Y^*(0)|\boldsymbol{X}]-\beta_1\right\}+ \mathbb{E}[Y^*(0)|\boldsymbol{X}]-\beta_0\bigg)\cdot \frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p \cdot \begin{bmatrix}
1 \\ T
\end{bmatrix} \bigg|\boldsymbol{X}\Bigg] \notag\\
&+\mathbb{E}\Bigg[\bigg(T\cdot\left\{ \mathbb{E}[Y^*(1)-Y^*(0)|\boldsymbol{X}]-\beta_1\right\}+ \mathbb{E}[Y^*(0)|\boldsymbol{X}]-\beta_0\bigg)\cdot \frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q \cdot \begin{bmatrix}
1 \\ T
\end{bmatrix} \bigg|\boldsymbol{X} \Bigg] \notag \\
=&\mathbb{E}\Bigg[\bigg( \mathbb{E}[Y^*(1)|\boldsymbol{X}]-\beta_1-\beta_0\bigg)\cdot \frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p \cdot \begin{bmatrix}
1 \\ 1
\end{bmatrix} \bigg|\boldsymbol{X}\Bigg] \notag\\
&+\mathbb{E}\Bigg[\bigg( \mathbb{E}[Y^*(0)|\boldsymbol{X}]-\beta_0\bigg)\cdot \frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q \cdot \begin{bmatrix}
1 \\ 0
\end{bmatrix} \bigg|\boldsymbol{X} \Bigg] \notag\\
=&\begin{bmatrix}
\bigg(\mathbb{E}[Y^*(1)|\boldsymbol{X}]-\beta_1-\beta_0\bigg)\cdot p +\bigg(\mathbb{E}[Y^*(0)|\boldsymbol{X}]-\beta_1\bigg)\cdot q
\\[4mm]
\bigg(\mathbb{E}[Y^*(1)|\boldsymbol{X}]-\beta_1-\beta_0\bigg)\cdot p
\end{bmatrix}. \label{eff:epsilon}
\end{align}
Therefore, with \eqref{m_inverse}, \eqref{eff:pi}, and \eqref{eff:epsilon} we can obtain that
\begin{align*}
&\pi_0 (T,
\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\left\{ Y-\mathbb{E}[Y|
\boldsymbol{X},T]\right\}+\mathbb{E}\left[ \varepsilon (T,\boldsymbol{X};
\boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|
\boldsymbol{X}\right]\\
=&\begin{pmatrix}
p\cdot \phi_1(T,\boldsymbol{X},Y;\boldsymbol{\beta}_0)+q\cdot \phi_2(T,\boldsymbol{X},Y;\boldsymbol{\beta}_0) \\[2mm]
p\cdot \phi_1(T,\boldsymbol{X},Y;\boldsymbol{\beta}_0)
\end{pmatrix},
\end{align*}
and the efficient influence functions of $\beta_1$ and $\beta_2$ are given by
\begin{align*}
&\begin{bmatrix}
\frac{1}{q}& -\frac{1}{q} \\
-\frac{1}{q}& \frac{1}{pq}
\end{bmatrix}\cdot \begin{pmatrix}
p\cdot \phi_1(T,\boldsymbol{X},Y;\boldsymbol{\beta})+q\cdot \phi_2(T,\boldsymbol{X},Y;\boldsymbol{\beta}) \\[2mm]
p\cdot \phi_1(T,\boldsymbol{X},Y;\boldsymbol{\beta})
\end{pmatrix}=\begin{pmatrix}
\phi_2(T,\boldsymbol{X},Y;\boldsymbol{\beta}) \\[2mm]
\phi_1(T,\boldsymbol{X},Y;\boldsymbol{\beta})- \phi_2(T,\boldsymbol{X},Y;\boldsymbol{\beta})\ .
\end{pmatrix}.
\end{align*}
\end{proof}
\subsection{Particular Case II: Multiple Treatment Effects \label{eff_bound:multiple}}
In this section, we show that when $T\in\{0,1,...,J\}$, $J\in\mathbb{N}$, $g(t;\boldsymbol{\beta})=\sum_{j=0}^J \beta_j \cdot I(t=j)$ and $L(v)=v^2$, our general efficiency bound derived in Theorem 3.1 reduces to the efficiency bound of multi-level treatment effects given in \cite{cattaneo2010efficient}. In accordance with our proposed identification condition, $\{\beta_j\}_{j=0}^J$ are identified by minimizing the following loss function
$$\sum_{j=0}^J\mathbb{E}\left[(Y^*(j)-\beta_j)^2\right]\cdot \mathbb{P}(T=j).$$
The solutions are $\beta_j=\mathbb{E}[Y^*(j)]$ for $j\in\{0,...,J\}$.
\begin{cor}
\label{cor_effbound_multiple}
Suppose $T\in\{0,1,...,J\}$, $J\in\mathbb{N}$, $g(t;\boldsymbol{\beta})=\sum_{j=0}^J \beta_j \cdot I(t=j)$, $L(v)=v^2$, and the conditions in Theorem 3.1 hold, the efficient influence functions of $\{\beta_j\}_{j=0}^J$ given by Theorem 3.1 reduce to
\begin{align*}
S_{eff}(T,\boldsymbol{X},Y;\beta_j)=\frac{I(T=j)}{\mathbb{P}(T=j|\boldsymbol{X})}\cdot \left\{Y^*(j)-\mathbb{E}[Y^*(j)|\boldsymbol{X}]\right\} + \mathbb{E}[Y^*(j)|X]- \beta_j,\ j\in\{0,...,J\},
\end{align*}
and they are the same as the efficient influence functions given in \cite{cattaneo2010efficient}.
\end{cor}
\begin{proof}
Using our notation, we have
\begin{align*}
\boldsymbol{\beta}_0=(\beta_0,...,\beta_J)^{\top}, \ g(t;\boldsymbol{\beta}_0)=\sum_{j=0}^J \beta_j \cdot I(t=j), \ m(t;\boldsymbol{\beta}_0)=\begin{bmatrix}
I(t=0) \\ I(t=1)\\ \vdots \\ I(t=J)
\end{bmatrix}, \ H_0=\mathbb{E}\left[m(T;\boldsymbol{\beta}_0)m(T;\boldsymbol{\beta}_0)^\top\right].
\end{align*}
Then
\begin{align*}
\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)=&\mathbb{E}[Y|T,X]-g(T;\boldsymbol{\beta}_0)\\
=&\sum_{j=0}^J \mathbb{E}[Y^*(j)|X] \cdot I(t=j)-\sum_{j=0}^J \beta_j \cdot I(T=j)\\
=&\sum_{j=0}^J\left( \mathbb{E}[Y^*(j)|X]- \beta_j\right) \cdot I(T=j)
\end{align*}
and
\begin{align*}
\pi_0(T,X)=\sum_{j=0}^J\frac{I(T=j)}{\mathbb{P}(T=j|\boldsymbol{X})}\cdot p_j, \ \text{where} \ p_j=\mathbb{P}(T=j).
\end{align*}
Then we have
\begin{align*}
H_0^{-1}=\mathbb{E}\left[m(T;\boldsymbol{\beta}_0)m(T;\boldsymbol{\beta}_0)^\top\right]^{-1}=\begin{bmatrix}
& p_0^{-1} & & & \\
& & p_1^{-1} & & \\
& & \cdots & & \\
& & & & p_J^{-1}
\end{bmatrix},
\end{align*}
and
\begin{align}
&\pi_0 (T,
\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\left\{ Y-\mathbb{E}[Y|
\boldsymbol{X},T]\right\} \notag\\
=&\left\{\sum_{j=0}^J\frac{I(T=j)}{\mathbb{P}(T=j|\boldsymbol{X})}\cdot p_j\right\} \cdot \begin{bmatrix}
I(T=0) \\ I(T=1)\\ \vdots \\I(T=J)
\end{bmatrix}\cdot\bigg\{Y-\sum_{j=0}^J I(T=j)\cdot \mathbb{E}[Y^*(j)|\boldsymbol{X}]\bigg\} \notag\\
=&\begin{bmatrix}
I(T=0) \\ I(T=1)\\ \vdots \\I(T=J)
\end{bmatrix} \left\{\sum_{j=0}^J\frac{I(T=j)}{\mathbb{P}(T=j|\boldsymbol{X})}\cdot p_j\cdot Y^*(j)-\sum_{j=0}^J\frac{I(T=j)}{\mathbb{P}(T=j|\boldsymbol{X})}\cdot p_j\cdot \mathbb{E}[Y^*(j)|\boldsymbol{X}] \right\}\notag\\
=& \begin{bmatrix}
\frac{I(T=0)}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot p_0\cdot \left\{Y^*(0)-\mathbb{E}[Y^*(0)|\boldsymbol{X}]\right\} \\[2mm]
\frac{I(T=1)}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p_1\cdot \left\{Y^*(1)-\mathbb{E}[Y^*(1)|\boldsymbol{X}]\right\}\\ \vdots \\[2mm] \frac{I(T=J)}{\mathbb{P}(T=J|\boldsymbol{X})}\cdot p_J\cdot \left\{Y^*(j)-\mathbb{E}[Y^*(j)|\boldsymbol{X}]\right\}
\end{bmatrix}\label{eff:pi_mul}
\end{align}
and
\begin{align*}
& \varepsilon (T,\boldsymbol{X};
\boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\\
=&\left\{\sum_{j=0}^J\left( \mathbb{E}[Y^*(j)|X]- \beta_j\right) \cdot I(T=j)\right\}\left\{\sum_{j=0}^J\frac{I(T=j)}{\mathbb{P}(T=j|\boldsymbol{X})}\cdot p_j\right\}\begin{bmatrix}
I(T=0) \\ I(T=1)\\ \vdots \\ I(T=J)
\end{bmatrix} \\
=&\begin{bmatrix}
\frac{I(T=0)}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot p_0\cdot \left\{ \mathbb{E}[Y^*(0)|X]- \beta_0\right\} \\[2mm] \frac{I(T=1)}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p_1\cdot \left\{ \mathbb{E}[Y^*(1)|X]- \beta_1\right\}\\ \vdots \\ \frac{I(T=J)}{\mathbb{P}(T=J|\boldsymbol{X})}\cdot p_J\cdot \left\{ \mathbb{E}[Y^*(j)|X]- \beta_J\right\}
\end{bmatrix}
\end{align*}
and
\begin{align}
\mathbb{E}\left[ \varepsilon (T,\boldsymbol{X};
\boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|
\boldsymbol{X}\right]= \begin{bmatrix}
p_0\cdot \left\{ \mathbb{E}[Y^*(0)|X]- \beta_0\right\} \\[2mm] p_1\cdot \left\{ \mathbb{E}[Y^*(1)|X]- \beta_1\right\}\\ \vdots \\ p_J\cdot \left\{ \mathbb{E}[Y^*(j)|X]- \beta_J\right\}
\end{bmatrix}. \label{eff:epsilon_mul}
\end{align}
From Theorem 3.1, the efficient influence function of $\boldsymbol{\beta }_0=(\beta_0,...,\beta_J)$ is given by
\begin{align*}
&H_0^{-1}\left\{\pi_0 (T,
\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\left\{ Y-\mathbb{E}[Y|
\boldsymbol{X},T]\right\}+ \mathbb{E}\left[ \varepsilon (T,\boldsymbol{X};
\boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)|
\boldsymbol{X}\right]\right\}\\
=&\begin{bmatrix}
\frac{I(T=0)}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot \left\{Y^*(0)-\mathbb{E}[Y^*(0)|\boldsymbol{X}]\right\} + \mathbb{E}[Y^*(0)|X]- \beta_0\\[2mm]
\frac{I(T=1)}{\mathbb{P}(T=1|\boldsymbol{X})} \cdot \left\{Y^*(1)-\mathbb{E}[Y^*(1)|\boldsymbol{X}]\right\} +\mathbb{E}[Y^*(1)|X]- \beta_1\\ \vdots \\[2mm] \frac{I(T=J)}{\mathbb{P}(T=J|\boldsymbol{X})} \cdot \left\{Y^*(j)-\mathbb{E}[Y^*(j)|\boldsymbol{X}]\right\}+ \mathbb{E}[Y^*(j)|X]- \beta_J
\end{bmatrix},
\end{align*}
which is the same as the efficient influence function developed in Corollary 1 of \cite{cattaneo2010efficient}.
\end{proof}
\subsection{Particular Case III: Quantile Treatment Effects}
In this section, we show that when $T\in\{0,1\}$ is a binary treatment variable, $L(v)=v(\tau-I(v\leq 0))$ is the check function with $\tau\in(0,1)$, and $g(t;\boldsymbol{\beta}_0)=\beta_0\cdot (1-t)+\beta_1\cdot t$, where $\boldsymbol{\beta}_0=(\beta_0,\beta_1)$, our general efficiency bound derived in Theorem 3.1 reduces to the efficiency bound of quantile treatment effects given in \cite{Firpo2007Efficient}. In accordance with our identification condition, $\beta_0$ and $\beta_1$ are identified by minimizing the following loss function
\begin{align*}
\sum_{j\in\{0,1\}}\mathbb{P}(T=j)\cdot \mathbb{E}\left[(Y^*(j)-\beta_j)\left\{\tau-I(Y^*(j)\leq \beta_j) \right\}\right].
\end{align*}
The solutions are $\beta_0=\inf\{q: \mathbb{P}(Y^*(0)\leq q)\geq \tau\}$ and $\beta_1=\inf\{q: \mathbb{P}(Y^*(1)\leq q)\geq \tau\}$, which are the $\tau^{th}$ quantiles of potential outcomes.
\begin{cor}\label{cor_effbound_QTE}
Let $T\in\{0,1\}$, $f_{Y^*(1)}$ and $f_{Y^*(0)}$ be the probability densities of the potential outcomes $Y^*(1)$ and $Y^*(0)$ respectively, $g(t;\boldsymbol{\beta}_0)=\beta_0\cdot (1-t)+\beta_1\cdot t$, $L(v)=v(\tau-I(v\leq 0))$, and the conditions in Theorem 3.1 hold, then the efficient influence function of $\boldsymbol{\beta}_0$ given by Theorem 3.1 reduces to
\begin{align*}
S_{eff}(Y,T,\boldsymbol{X};\boldsymbol{\beta}_0)
=\begin{bmatrix}
\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot \left\{\frac{\tau-I(Y^*(0)\leq \beta_0) }{f_{Y^*(0)}(\beta_0)}\right\}- \left(\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}-1\right)\cdot \mathbb{E}\left[\frac{\tau-I(Y^*(0)\leq \beta_0)}{f_{Y^*(0)}(\beta_0)} \big|\boldsymbol{X}\right]\\[2mm]
\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})} \cdot \left\{\frac{\tau-I(Y^*(1)\leq \beta_1)}{f_{Y^*(1)}(\beta_1)} \right\}- \left(\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}-1\right)\cdot \mathbb{E}\left[\frac{\tau-I(Y^*(1)\leq \beta_1)}{f_{Y^*(1)}(\beta_1)} \big|\boldsymbol{X}\right]
\end{bmatrix},
\end{align*}
which is the same as the efficient influence function given in \cite{Firpo2007Efficient}.
\end{cor}
\begin{proof}
Using our notation, we have
\begin{align*}
&\boldsymbol{\beta}_0=(\beta_0,\beta_1)^{\top},\ g(t;\boldsymbol{\beta}_0)=\beta_0\cdot (1-t)+\beta_1 \cdot t , \ m(t;\boldsymbol{\beta}_0)=\begin{bmatrix}
1-t \\ t
\end{bmatrix},\\
&L(v)=v(\tau-I(v\leq 0)), \ L'(v)=\tau-I(v\leq 0)\ \text{a.s.}, \\
&\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)=T\cdot \mathbb{E}[\tau-I(Y^*(1)\leq \beta_1)|\boldsymbol{X}]+(1-T)\cdot \mathbb{E}[\tau-I(Y^*(0)\leq \beta_0)|\boldsymbol{X}], \\
&\pi_0(T,\boldsymbol{X})=\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p +\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q\ , \ p=\mathbb{P}(T=1), \ q=\mathbb{P}(T=0).
\end{align*}
Direct computation yields
\begin{align*}
&\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)L'(Y-g(T;\boldsymbol{\beta}_0))\\
=&\left\{\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p +\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q\right\}\cdot \begin{bmatrix}
1-T \\ T
\end{bmatrix}\cdot \bigg\{\tau-I(Y\leq \beta_0\cdot (1-T)+\beta_1\cdot T )\bigg\}\\
=& \begin{bmatrix}
\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q\cdot \left\{\tau-I(Y^*(0)\leq \beta_0) \right\} \\[2mm]
\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p\cdot \left\{\tau-I(Y^*(1)\leq \beta_1) \right\}
\end{bmatrix}
\end{align*}
and
\begin{align*}
\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)=\begin{bmatrix}
\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q\cdot \mathbb{E}\left[\tau-I(Y^*(0)\leq \beta_0)|\boldsymbol{X} \right] \\[2mm]
\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p\cdot \mathbb{E}\left[\tau-I(Y^*(1)\leq \beta_1) |\boldsymbol{X}\right]
\end{bmatrix}
\end{align*}
and
\begin{align*}
\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)|\boldsymbol{X}\right]=\begin{bmatrix}
q\cdot \mathbb{E}\left[\tau-I(Y^*(0)\leq \beta_0)|\boldsymbol{X} \right] \\[2mm]
p\cdot \mathbb{E}\left[\tau-I(Y^*(1)\leq \beta_1) |\boldsymbol{X}\right]
\end{bmatrix}
\end{align*}
and
\begin{align*}
H_0=&\nabla_{\boldsymbol{\beta}}\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta})L'(Y-g(T;\boldsymbol{\beta}))\right]
=\begin{bmatrix}
-q\cdot f_{Y^*(0)}(\beta_0) & 0 \\
0 &-p\cdot f_{Y^*(1)}(\beta_1)
\end{bmatrix}.
\end{align*}
Therefore, by Theorem 3.1, the efficient influence function of $\boldsymbol{\beta}_0$ is
\begin{align*}
&S_{eff}(Y,T,\boldsymbol{X};\boldsymbol{\beta}_0)\\
=&H_0^{-1}\cdot \bigg\{ \pi_0 (T,
\boldsymbol{X})m(T;\boldsymbol{\beta }_{0})L'(Y-g(T;\boldsymbol{\beta}_0))-\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta }_{0})\varepsilon (T,\boldsymbol{X};
\boldsymbol{\beta }_{0})\\
& \qquad \qquad+\mathbb{E}\left[ \varepsilon (T,\boldsymbol{X};
\boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta }_{0})|
\boldsymbol{X}\right] \bigg\}\\
=&\begin{bmatrix}
q^{-1}\cdot \frac{1}{f_{Y^*(0)}(\beta_0)} & 0 \\
0 &p^{-1}\cdot \frac{1}{f_{Y^*(1)}(\beta_1)}
\end{bmatrix}\\
&\times \begin{bmatrix}
\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q\cdot \left\{\tau-I(Y^*(0)\leq \beta_0) \right\}-q\cdot \left(\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}-1\right)\cdot \mathbb{E}\left[\tau-I(Y^*(0)\leq \beta_0) |\boldsymbol{X}\right]\\[2mm]
\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p\cdot \left\{\tau-I(Y^*(1)\leq \beta_1) \right\}-p\cdot \left(\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}-1\right)\cdot \mathbb{E}\left[\tau-I(Y^*(1)\leq \beta_1) |\boldsymbol{X}\right]
\end{bmatrix}\\
=&\begin{bmatrix}
\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot \left\{\frac{\tau-I(Y^*(0)\leq \beta_0) }{f_{Y^*(0)}(\beta_0)}\right\}- \left(\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}-1\right)\cdot \mathbb{E}\left[\frac{\tau-I(Y^*(0)\leq \beta_0)}{f_{Y^*(0)}(\beta_0)} \bigg|\boldsymbol{X}\right]\\[2mm]
\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})} \cdot \left\{\frac{\tau-I(Y^*(1)\leq \beta_1)}{f_{Y^*(1)}(\beta_1)} \right\}- \left(\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}-1\right)\cdot \mathbb{E}\left[\frac{\tau-I(Y^*(1)\leq \beta_1)}{f_{Y^*(1)}(\beta_1)} \bigg|\boldsymbol{X}\right]
\end{bmatrix},
\end{align*}
which coincides with efficiency bound derived in \cite{Firpo2007Efficient}.
\end{proof}
\section{Convergence Rate of Estimated Stabilized Weights \label{sec:key_lemmas}}
In this section, we establish the convergence rate of estimated stabilized weights $\hat{\pi}_K(T,\boldsymbol{X})$. Let $G_{K_1\times K_2}^*$, $\Lambda_{K_1\times K_2}^*$ and $\pi_K^*(t,\boldsymbol{x})$ be the theoretical
counterparts of $\hat{G}_{K_1\times K_2}$, $\hat{\Lambda}_{K_1\times K_2}$ and $\hat{\pi}_K(t,\boldsymbol{x})$ respectively:
\begin{align*}
G_{K_1\times K_2}^*(\Lambda):= & \mathbb{E}[\hat{G}_{K_1\times
K_2}(\Lambda)]=\mathbb{E}\left[\rho\left(u_{K_1}(T)^{\top}\Lambda v_{K_2}(
\boldsymbol{X})\right)\right]-\mathbb{E}[u_{K_1}(T)^{\top}]\cdot\Lambda\cdot
\mathbb{E}[v_{K_2}(\boldsymbol{X})], \\[2mm]
\Lambda_{K_1\times K_2}^* := & \arg\max G_{K_1\times K_2}^*(\Lambda), \\[2mm]
\pi_K^*(t,\boldsymbol{x}):= &\rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right).
\end{align*}
As discussed in Appendix A.3, we assume the sieve basises $u_{K_1}(T)$ and $v_{K_2}(\boldsymbol{X})$ are orthonormalized, i.e.,
\begin{align}\label{eq:orthonormal_basis}
\mathbb{E}\left[u_{K_1}(T)u_{K_1}^{\top}(T)\right]=I_{K_1\times K_1}, \quad \mathbb{E}\left[v_{K_2}(\boldsymbol{X})v_{K_2}^{\top}(\boldsymbol{X})\right]=I_{K_2\times K_2} .
\end{align}
Let
\begin{align*}
& \zeta_1(K_1) :=\sup_{t\in\mathcal{T}}\|u_{K_1}(t)\|\ , \
\zeta_2(K_2):=\sup_{\boldsymbol{x}\in\mathcal{X}}\|v_{K_2}(\boldsymbol{x})\| \ , \ K=K_1\cdot K_2 \ , \ \zeta(K)=\zeta_1(K_1)\zeta_2(K_2).
\end{align*}
We also recall the following property satisfied by $\pi_0(T,\boldsymbol{X})$: for any integrable functions $u(t)$ and $v(\boldsymbol{X})$,
\begin{equation}
\mathbb{E}\left[ \pi _{0}(T,\boldsymbol{X})u(T)v(\boldsymbol{X})\right] =
\mathbb{E}[u(T)]\cdot \mathbb{E}[v(\boldsymbol{X})]. \label{moment1}
\end{equation}
\subsection{Lemma $\ref{lemma_pi^*}$}
The first lemma states that $\pi_K^*(t,\boldsymbol{x})$ is arbitrarily close to the true stabilized weights $\pi_0(t,\boldsymbol{x})$.
\begin{lemma}\label{lemma_pi^*}
Under Assumption \ref{as:suppX}-\ref{as:K&N_consistency}, we have
$$ \sup_{(t,\boldsymbol{x}) \in \mathcal{T} \times \mathcal{X}}|\pi_0(t,\boldsymbol{x}) - \pi^*_K(t,\boldsymbol{x}) | = O\left(K^{-\alpha}\zeta(K)\right),$$
and
$$ \mathbb{E}\left[|\pi_0(T,\boldsymbol{X}) - \pi^*_K(T,\boldsymbol{X}) |^2\right] = O\left(K^{-2\alpha}\right),$$
and
$$\frac{1}{N}\sum_{i=1}^N|\pi_0(T_i,\boldsymbol{X}_i) - \pi^*_K(T_i,\boldsymbol{X}_i) |^2=O_p\left(K^{-2\alpha}\right).$$
\end{lemma}
\begin{proof} By Assumption \ref{as:pi0}, $\pi_0(t,\boldsymbol{x})\in [\eta_1, \eta_2], \ \forall (t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}$ and $(\rho')^{-1}$ is strictly decreasing.
Define $$ \overline{\gamma} := \sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times\mathcal{X}} (\rho')^{-1}\left(\pi_0(t,\boldsymbol{x})\right)
\leq (\rho')^{-1}(\eta_1) ~~ \text{and} ~~
\underline{\gamma} := \inf_{(t,\boldsymbol{x}) \in \mathcal{T}\times\mathcal{X}} (\rho')^{-1}\left(\pi_0(t,\boldsymbol{x})\right)
\geq(\rho')^{-1}(\eta_2),$$
which are two finite constants.
By Assumptions \ref{as:smooth_pi},
there exist a constant $C> 0$ and a $K_1\times K_2$ matrix $\Lambda_{K_1\times K_2} \in \mathbb{R}^{K_1\times K_2}$ such that
\begin{align*}
\sup_{(t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X}}\left|(\rho')^{-1}\left(\pi_0(t,\boldsymbol{x})\right) - u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right| < CK^{-\alpha},
\end{align*}
which implies
\begin{align}\label{eq:resubmit_lamdau_K}
u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) \in & \left((\rho')^{-1}\left(\pi_0(t,\boldsymbol{x})\right) - CK^{-\alpha}, (\rho')^{-1}\left(\pi_0(t,\boldsymbol{x})\right) + CK^{-\alpha} \right) \\
\subset & \left[\underline{\gamma} - CK^{-\alpha}, \overline{\gamma} + CK^{-\alpha}\right], \ \forall (t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}, \notag
\end{align}
and
\begin{align*}
&\rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) +CK^{-\alpha}\right) - \rho'(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{X}))\\
< &\pi_0(t,\boldsymbol{x}) - \rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right) \\
< & \rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})-CK^{-\alpha}\right) - \rho'(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})) \ , \forall (t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X}.
\end{align*}
Let $\Gamma_1:= [\underline{\gamma} -1, \overline{\gamma} + 1]$,
by Mean Value Theorem, for large enough $K$, there exist
\begin{align*}
\xi_1(t,\boldsymbol{x}) \in & \left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}), u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) + CK^{-\alpha}\right)\\
\subset &\left[\underline{\gamma} -CK^{-\alpha}, \overline{\gamma} + 2CK^{-\alpha}\right]\subset\Gamma_1 \ , \\
\xi_2(t,\boldsymbol{x}) \in& \left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) -CK^{-\alpha}, u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)\\
\subset & \left[\underline{\gamma} - 2CK^{-\alpha}, \overline{\gamma} + CK^{-\alpha}\right]\subset \Gamma_1,
\end{align*}
such that
\begin{align*}
\rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) + CK^{-\alpha}\right) - \rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) \right)= \rho''(\xi_1(t,x))CK^{-\alpha}\geq -a_1CK^{-\alpha}
\end{align*}
and
\begin{align*}
\rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) - CK^{-\alpha}\right) - \rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) \right)= -\rho''(\xi_2(t,\boldsymbol{x}))CK^{-\alpha} \leq a_2CK^{-\alpha},
\end{align*}
where $ -a_1 :=\inf_{\gamma \in \Gamma_1} \rho''(\gamma)$ and $a_2 := \sup_{\gamma \in \Gamma_1}\left( -\rho''(\gamma)\right)$.
Let $a := \max\{a_1, a_2\}$, we have
\begin{align} \label{eq:SW1}
\sup_{(t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X}}\left| {\pi_0(t,\boldsymbol{x})} - \rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) \right)\right| < aCK^{-\alpha}.
\end{align}
For some fixed $C_2 > 0$ (to be chosen later), define
$$ \Upsilon_{K_1\times K_2} := \left\{\Lambda \in \mathbb{R}^{K_1\times K_2}: \|\Lambda - \Lambda_{K_1\times K_2}\| \leq C_2K^{-\alpha}\right\}.$$
For sufficiently large $K_1$ and $K_2$,
we have that $\forall \Lambda \in \Upsilon_{K_1\times K_2}$, $\forall (t,\boldsymbol{x}) \in \mathcal{T}\times\mathcal{X}$,
\begin{align*}
&\left|u_{K_1}(t)^{\top}\Lambda v_{K_2}(\boldsymbol{x}) - u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right|\\
\leq &\|\Lambda - \Lambda_{K_1\times K_2}\| \cdot \sup_{\boldsymbol{x}\in\mathcal{X}}\|v_{K_2}(\boldsymbol{x}) \| \cdot \sup_{t\in\mathcal{T}}\|u_{K_1}(t)\|
\leq C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2) .
\end{align*}
Then in light of \eqref{eq:resubmit_lamdau_K} and Assumption \ref{as:K&N_c}, for large enough $K_1$ and $K_2$, $\forall\Lambda \in \Upsilon_{K_1\times K_2}$ and $\forall (t,\boldsymbol{x}) \in \mathcal{T}\times\mathcal{X}$, we can deduce that
\begin{align}\label{eq:gamma3}
& u_{K_1}(t)^{\top}\Lambda v_{K_2}(\boldsymbol{x}) \in \left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) - C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2), \right. \\
&\left. \quad \quad \quad \quad \quad \quad \quad \quad \quad u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})+ C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2)\right) \notag \\
& ~~~~~~~~~~~~~~~~~~~~~\subset \left[\underline{\gamma} - CK^{-\alpha} - C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2), \right. \notag\\
&~~~~~~~~~~~~~~~~~~~~~~~~~~~\left. \overline{\gamma} + CK^{-\alpha} + C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2)\right]\subset \Gamma_1.\notag
\end{align}
By definition
\begin{align*}
G_{K_1\times K_2}^*\left(\Lambda\right)= \mathbb{E}\left[\rho\left(u_{K_1}(T)^{\top}\Lambda v_{K_2}(\boldsymbol{X})\right)\right] -\mathbb{E}[u_{K_1}(T)]^{\top}\Lambda \mathbb{E}[v_{K_2}(\boldsymbol{X})],
\end{align*}
is a strictly concave function of $\Lambda$. By \eqref{moment1}, the formula $\tr(AB)=\tr(BA)$ for matrices $A$ and $B$, the facts $\mathbb{E}\left[v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^\top\right]=I_{K_2\times K_2}$ and $\mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^\top\right]=I_{K_1\times K_1}$, we can deduce that{\footnotesize
\begin{align}
&\|\nabla {G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\|^2\notag\\
=&\left\|\mathbb{E}\left[\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] -\mathbb{E}[u_{K_1}(T)]\mathbb{E}[v_{K_2}(\boldsymbol{X})]^{\top}\right\|^2 \notag\\
=&\left\|\mathbb{E}\left[\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] -\mathbb{E}[\pi_0(T,\boldsymbol{X})u_{K_1}(T)v_{K_2}(\boldsymbol{X})]^{\top}\right\|^2\quad \text{(by \eqref{moment1})} \notag\\
=&\left\|\mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] \right\|^2 \notag\\
=&\tr\Bigg\{\mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] \notag \\
&\quad \times \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}v_{K_2}(\boldsymbol{X})u_{K_1}(T)^\top\right]\Bigg\} \notag\\
=&\tr\Bigg\{\mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] \cdot \mathbb{E}\left[u_{K_2}(\boldsymbol{X})u_{K_2}(\boldsymbol{X})^\top\right]\notag \\
&\quad \times \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}v_{K_2}(\boldsymbol{X})u_{K_1}(T)^\top\right]\cdot \mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^\top\right]\Bigg\} \notag\\
=&\mathbb{E}\Bigg[\tr\Bigg\{u_{K_1}(T)^\top\cdot \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] \cdot \mathbb{E}\left[u_{K_2}(\boldsymbol{X})u_{K_2}(\boldsymbol{X})^\top\right]\notag \\
&\qquad \times \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}v_{K_2}(\boldsymbol{X})u_{K_1}(T)^\top\right]\cdot u_{K_1}(T)\Bigg\} \Bigg] \notag\\
=&\mathbb{E}\Bigg[\pi_0(T,\boldsymbol{X}) \cdot u_{K_1}(T)^\top\cdot \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] \cdot u_{K_2}(\boldsymbol{X})\notag \\
&\qquad\times \cdot u_{K_2}(\boldsymbol{X})^\top \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}v_{K_2}(\boldsymbol{X})u_{K_1}(T)^\top\right]\cdot u_{K_1}(T) \Bigg] \quad \text{(by \eqref{moment1})} \notag\\
=&\mathbb{E}\Bigg[\bigg|{\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}u_{K_1}(T) \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] {\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}v_{K_2}(\boldsymbol{X}) \bigg|^2\Bigg]\label{eq:nablaG^*_LambdaK}.
\end{align}}
Note that the term in the last expression {\footnotesize
$${\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}u_{K_1}(T)\cdot \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] {\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}v_{K_2}(\boldsymbol{X})$$}
is the $L^2(dF_{T,X})$-projection of $\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}$ on the space spanned by {\footnotesize$\{{\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}u_{K_1}(T)$, ${\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}v_{K_2}(\boldsymbol{X})\}$}, which implies that{\footnotesize
\begin{align}
&\mathbb{E}\Bigg[\bigg|{\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}u_{K_1}(T) \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] {\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}v_{K_2}(\boldsymbol{X}) \bigg|^2\Bigg]\notag\\
\leq &\mathbb{E}\left[\bigg|\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}\bigg|^2\right]. \label{eq:L^2}
\end{align}}
Now, with \eqref{eq:nablaG^*_LambdaK}, \eqref{eq:L^2}, we can obtain that
\begin{align}
&\|\nabla {G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| \notag\\
\leq& \mathbb{E}\left[\bigg|\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}\bigg|^2\right]^{\frac{1}{2}}\notag\\
\leq & \frac{1}{\sqrt{\eta_1}} \sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}\left|\rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{x})\right)-\pi_0(t,\boldsymbol{x})\right| \qquad (\text{by Assumption \ref{as:pi0}}) \notag \\
\leq & \frac{aC}{\sqrt{\eta_1}}\cdot K^{-\alpha} \qquad (\text{by \eqref{eq:SW1}}. \label{order:nabla_G_K*}
\end{align}
Note that for any $\Lambda \in \partial \Upsilon_{K_1\times K_2}$, i.e. $\|\Lambda - \Lambda_{K_1\times K_2}\|= C_2K^{-\alpha}$, by Mean Value Theorem and the fact $\rho''(y)=-\rho'(y)$, we can deduce that{\small
\begin{align*}
& \quad {G}^*_{K_1\times K_2}(\Lambda) - {G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\\
&= \sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} \frac{\partial}{\partial \lambda_i}{G}^*_{K_1\times K_2}(\lambda_1^K,\ldots,\lambda_{K_2}^K)\\
&\qquad + \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}\frac{1}{2}(\lambda_j - \lambda_j^K)^{\top} \frac{\partial^2}{\partial \lambda_i \partial \lambda_l}{G}^*_{K_1\times K_2}(\bar{\lambda}^K_1,\ldots,\bar{\lambda}^K_{K_2})(\lambda_l - \lambda_l^K) \notag\\
& \leq \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| \\
&
+\frac{1}{2} \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} \mathbb{E}\left[\rho''\left(u_{K_1}^\top(T)\bar{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right)u_{K_1}(T)u_{K_1}(T)^\top v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right](\lambda_l - \lambda_l^K) \notag\\
& = \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| \\
&
\quad -\frac{1}{2} \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} \mathbb{E}\left[\frac{\rho'\left(u_{K_1}^\top(T)\bar{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right)}{\pi_0(T,\boldsymbol{X})}\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^\top v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right](\lambda_l - \lambda_l^K) \notag\\
& \leq \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| \\
&\quad-\frac{a_3}{2\eta_2} \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} \mathbb{E}\left[\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^\top v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right](\lambda_l - \lambda_l^K) \notag \quad \text{(by $a_3=\inf_{y\in\Gamma_1}\{\rho'(y)\}$)}\\
& = \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| \\
&\quad-\frac{a_3}{2\eta_2} \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} \mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^\top\right]\mathbb{E}\left[ v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right](\lambda_l - \lambda_l^K) \quad \text{(by \eqref{moment1})}\notag\\
& = \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| -\frac{a_3}{2\eta_2} \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} \mathbb{E}\left[ v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right](\lambda_l - \lambda_l^K) \notag\\
& = \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| -\frac{a_3}{2\eta_2} \sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} (\lambda_j - \lambda_j^K) \notag\\
&= \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\|
- \frac {a_3}{2\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|^2 \notag\\
& = \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\left(\|\nabla {G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| -
\frac {a_3}{2\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\right)\\
&\leq \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\left(\frac{aC}{\sqrt{\eta_1}}K^{-\alpha}-\frac {a_3}{2\eta_2}\cdot C_2K^{-\alpha}\right), \quad \text{(by \eqref{order:nabla_G_K*})}
\end{align*}}
where $\bar{\Lambda}_{K_1\times K_2}=(\bar{\lambda}_1^K,...,\bar{\lambda}_{K_2}^K)$ lies on the line joining ${\Lambda}=(\lambda_1,...,\lambda_{K_2})$ and ${\Lambda}_{K_1\times K_2}=(\lambda_1^K,...,\lambda_{K_2}^K)$, which implies $u_{K_1}^\top(t)\bar{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\in \Gamma_1$ by \eqref{eq:gamma3}; $a_3=\inf_{y\in\Gamma_1}\{\rho'(y)\}>0$ is a finite positive constant;
the fourth and fifth equalities follow from $\mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^\top\right]=I_{K_1\times K_1}$ and $\mathbb{E}\left[v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^\top\right]=I_{K_2\times K_2}$ respectively. Therefore, by choosing
$$C_2>\frac{2\eta_2}{a_3}\cdot \frac{aC}{{\sqrt{\eta_1}}},$$
we can obtain the following conclusion:
\begin{align}\label{eq:concave}
G^*_{K_1\times K_2}(\Lambda_{K_1\times K_2}) > G^*_{K_1\times K_2}(\Lambda) \ ,~~ \forall \Lambda \in \partial\Upsilon_{K_1\times K_2}\ .
\end{align}
Since $G^*_{K_1\times K_2}$ is continuous, \eqref{eq:concave} implies that there exists a local maximum of $G^*_{K_1\times K_2}$ in the interior of $\Upsilon_{K_1\times K_2}$. Note that $G^*_{K_1\times K_2}$ is strictly concave with a unique global maximum point $\Lambda^*_{K_1\times K_2}$, therefore we can claim that
\begin{align}\label{Lambda_interior}
\Lambda^*_{K_1\times K_2} \in \Upsilon_{K_1\times K_2}^\circ, \ \text{i.e.}\ \|\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\|=O(K^{-\alpha}) \ .
\end{align}
By Mean Value Theorem, \eqref{eq:gamma3} and \eqref{Lambda_interior}, we can deduce that
\begin{align} \notag
&|\rho'\left(u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right) - \rho'\left(u_{K_1}(t)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)| \\
= &|\rho''(\xi^*(t,\boldsymbol{x}))|\left|u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})-u_{K_1}(t)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right| \notag\\
\leq & -\rho''(\xi^*(t,\boldsymbol{x})) \times \|\Lambda_{K_1\times K_2}-\Lambda_{K_1\times K_2}^* \| \times \sup_{t\in\mathcal{T}}\|u_{K_1}(t)\| \times \sup_{\boldsymbol{x}\in \mathcal{X}}\|v_{K_2}(\boldsymbol{x})\| \notag \\
\leq & a_2C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2) \ ,\notag
\end{align}
where $ a_2 = \sup_{\gamma \in \Gamma_1} \{ -\rho''(\gamma) \} <\infty$ is a finite positive constant, and $\xi^*(t,\boldsymbol{x})$
lies between the point $u_{K_1}(t)^{\top} \Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})$ and $u_{K_1}(t)^{\top} \Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})$
(note \eqref{eq:gamma3} implies $\xi^*(t,\boldsymbol{x}) \in \Gamma_1 $ for all $(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}$ and large enough $K$).
Therefore, using the triangle inequality, and Assumption \ref{as:K&N_c}, we can have
\begin{align}
&\sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}}\left|{\pi_0(t,\boldsymbol{x})} - \pi^*_K(t,\boldsymbol{x})\right| \notag \\
\leq &~ \sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}}\left|{\pi_0(t,\boldsymbol{x})} - \rho'\left(u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)\right|\notag\\
&+ \sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}}\left|\rho'\left(u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right) - \rho'\left(u_{K_1}(t)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)\right| \notag \\
\leq &~ aCK^{-\alpha} + a_2C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2)=O\left(K^{-\alpha}\zeta(K)\right), \notag
\end{align}
where $\zeta(K)=\zeta_1(K_1)\zeta_2(K_2)$. \\
We next prove $\mathbb{E}\left[\left|{\pi_0(T,\boldsymbol{X})} - \pi^*_K(T,\boldsymbol{X})\right|^2\right]=O\left(K^{-2\alpha}\right)$. By Assumption \ref{as:K&N_c}, we can deduce that
{\footnotesize
\begin{align*}
&\mathbb{E}\left[\left|{\pi_0(T,\boldsymbol{X})} - \pi^*_K(T,\boldsymbol{X})\right|^2\right]\\
\leq &2\cdot \mathbb{E}\left[\left|{\pi_0(T,\boldsymbol{X})} - \rho'\left(u_{K_1}(T)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right)\right|^2\right]+2\cdot \mathbb{E}\left[\left|\rho'\left(u_{K_1}(T)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right) - \rho'\left(u_{K_1}(T)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right)\right|^2\right]\\
\leq & 2\cdot \sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}}\left|{\pi_0(t,\boldsymbol{x})} - \rho'\left(u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)\right|^2+2\sup_{\gamma\in \Gamma_1}|\rho''(\gamma)|^2\cdot \mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]\\
\leq & O(K^{-2\alpha})+ O(1)\cdot \mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right].
\end{align*}}
We next compute the order of $\mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]$. Note that $\mathbb{E}[u_{K_1}(T)u_{K_1}(T)^\top]=I_{K_1\times K_1}$, $\mathbb{E}[v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^\top]=I_{K_2\times K_2}$, \eqref{moment1}, \eqref{Lambda_interior} and Assumption \ref{as:pi0}, we can deduce that {\footnotesize
\begin{align}
&\mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]\notag\\
=& \mathbb{E}\left[ u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^{\top}\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top u_{K_1}(T)\right] \notag\\
=& \mathbb{E}\left[\frac{1}{\pi_0(T,\boldsymbol{X})}\pi_0(T,\boldsymbol{X}) u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^{\top}\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top u_{K_1}(T)\right] \notag\\
\leq& \frac{1}{\eta_1}\cdot \mathbb{E}\left[\pi_0(T,\boldsymbol{X}) u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^{\top}\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top u_{K_1}(T)\right] \notag\\
=&\frac{1}{\eta_1}\cdot \int_{\mathcal{T}} u_{K_1}^\top(t)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}\mathbb{E}\left[v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^{\top}\right]\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top u_{K_1}(t)dF_T(t)\quad (\text{by \eqref{moment1}}) \notag\\
=& \frac{1}{\eta_1}\cdot \int_{\mathcal{T}} u_{K_1}^\top(t)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}\cdot\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top u_{K_1}(t)dF_T(t) \notag\\
=&\frac{1}{\eta_1}\cdot \int_{\mathcal{T}} \tr\Bigg(\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}\cdot\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top u_{K_1}(t)u_{K_1}^\top(t)\Bigg)dF_T(t) \notag\\
=&\frac{1}{\eta_1}\cdot \tr\Bigg(\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}\cdot\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top\Bigg) \notag\\
\leq & \frac{1}{\eta_1}\cdot\|\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\|^2= O(K^{-2\alpha}). \quad (\text{by \eqref{Lambda_interior}}) \label{eq:u_lam_v}
\end{align}}
Therefore, we can obtain
\begin{align*}
\mathbb{E}\left[\left|{\pi_0(T,\boldsymbol{X})} - \pi^*_K(T,\boldsymbol{X})\right|^2\right]=O\left(K^{-2\alpha}\right).
\end{align*}
We finally prove $N^{-1}\sum_{i=1}^N\left|{\pi_0(T_i,\boldsymbol{X}_i)} - \pi^*_K(T_i,\boldsymbol{X}_i)\right|^2=O_p\left(K^{-2\alpha}\right)$. Note that by \eqref{eq:u_lam_v}, we can have{\footnotesize
\begin{align*}
&\mathbb{E}\left[\left\{\frac{1}{N}\sum_{i=1}^N \left|u_{K_1}^\top(T_i)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}_i) \right|^2-\mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]\right\}^2\right]\\
\leq &\frac{1}{N}\cdot \mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^4\right]\\
\leq& \frac{1}{N}\cdot \mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]\cdot \sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}\left|u_{K_1}^\top(t)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x}) \right|^2\\
\leq & \frac{1}{N}\cdot O(K^{-2\alpha})\cdot \zeta_1(K_1)^2\zeta_2(K_2)^2\cdot \left\|\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\|^2\leq \frac{1}{N}\cdot \zeta(K)^2\cdot O( K^{-4\alpha}),
\end{align*}}
then in light of Chebyshev's inequality and Assumption \ref{as:K&N_consistency}, we have
\begin{align}\label{eq:uv-E[uv]}
&\frac{1}{N}\sum_{i=1}^N \left|u_{K_1}^\top(T_i)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}_i) \right|^2-\mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]\notag \\
=&O_p\left(\frac{\zeta(K)}{\sqrt{N}}K^{-2\alpha}\right)=o_p\left(K^{-2\alpha}\right).
\end{align}
With \eqref{Lambda_interior}, \eqref{eq:u_lam_v}, \eqref{eq:uv-E[uv]}, and Assumption \ref{as:pi0}, we can deduce that
\begin{align*}
&\frac{1}{N}\sum_{i=1}^N\left|{\pi_0(T_i,\boldsymbol{X}_i)} - \pi^*_K(T_i,\boldsymbol{X}_i)\right|^2\\
\leq &\frac{2}{N}\sum_{i=1}^N \left|{\pi_0(T_i,\boldsymbol{X}_i)} - \rho'\left(u_{K_1}(T_i)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)\right|^2\\ &+\frac{2}{N}\sum_{i=1}^N \left|\rho'\left(u_{K_1}(T_i)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right) - \rho'\left(u_{K_1}(T_i)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)\right|^2 \\
\leq &2\sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}}\left|{\pi_0(t,\boldsymbol{x})} - \rho'\left(u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)\right|^2\\
&+\sup_{\gamma\in \Gamma_1}|\rho''(\gamma)|^2\cdot \frac{2}{N}\sum_{i=1}^N \left|u_{K_1}^\top(T_i)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}_i) \right|^2 \\
\leq &2\sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}}\left|{\pi_0(t,\boldsymbol{x})} - \rho'\left(u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)\right|^2\\
&+2\cdot \sup_{\gamma\in \Gamma_1}|\rho''(\gamma)|^2\cdot \mathbb{E}\left[ \left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]+o_p\left(K^{-2\alpha}\right) \\
=& O(K^{-2\alpha})+ O(K^{-2\alpha})+o_p\left(K^{-2\alpha}\right) = O_p\left(K^{-2\alpha}\right). \quad (\text{by \eqref{Lambda_interior}})
\end{align*}
\end{proof}
\subsection{Lemma $\ref{lemma_pi^hat}$}
\begin{lemma}\label{lemma_pi^hat} Under Assumption \ref{as:suppX}-\ref{as:K&N_consistency}, we have
$$ \left\|\hat{\Lambda}_{K_1\times K_2}- \Lambda_{K_1\times K_2}^*\right\| = O_p\left(\sqrt{\frac {K} {N}} \right)\ .$$
\end{lemma}
\begin{proof}
Define $$ \hat{S}_N := \frac {1} {N} \sum_{i=1}^N \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \pi_0(T_i,\boldsymbol{X}_i)u_{K_1}(T_i)u_{K_1}(T_i)^{\top}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X}_i)v_{K_2,l}(\boldsymbol{X}_i),$$
where $\lambda_j$ and $\lambda^*_j$ are the $j$-th column of $\Lambda$ and $\Lambda_{K_1\times K_2}^*$ respectively.
Since $\hat{S}_N$ is symmetric, using \eqref{moment1} and the facts that $\mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^{\top }\right]=I_{K_1\times K_1}$ and $\mathbb{E}\left[v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^{\top}\right]=I_{K_2\times K_2}$, we can have
\begin{align*}
\mathbb{E}\left[\hat{S}_N\right]=&\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \mathbb{E}\left[\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right](\lambda_l-\lambda_l^*)\\
=&\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^{\top}\right]\mathbb{E}[v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})](\lambda_l-\lambda_l^*)\\
=&\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top}(\lambda_j-\lambda_j^*)=\left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|.
\end{align*}
Then we can further deduce that
\begin{align*}
&\mathbb{E}\left[\left|\hat{S}_N - \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\| \right|^2\right]\notag\\
= & \mathbb{E}[\hat{S}_N^2] - 2\mathbb{E}[\hat{S}_N]\left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\| + \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2\notag \\
= &\frac{N}{N^2}\cdot \mathbb{E}\left[\left(\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right)^2\right]\\
&+2\cdot \frac{C_N^2}{N^2}\cdot \mathbb{E}\left[\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right]^2 \notag \\
&-\left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2 \notag \\
= &\frac{1}{N}\mathbb{E}\left[\left(\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right)^2\right]\\
&+\frac{N(N-1)}{N^2}\cdot \mathbb{E}\left[\hat{S}_N\right]^2 -\left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2 \notag \\
=&\frac{1}{N}\mathbb{E}\left[\left(\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right)^2\right]\\
& \quad -\frac{1}{N}\left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2\\
< & \frac{1}{N}\mathbb{E}\left[\left(\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right)^2\right].
\end{align*}
In light of the fact that
$$ 0 \leq y^{\top} \left\{\pi_0(t,\boldsymbol{x})u_{K_1}(t)u_{K_1}(t)^{\top} \right\}y \leq \eta_2 \zeta_1(K_1)^2 y^{\top}y \ , \ \forall y\in\mathbb{R}^{K_1},\ \forall (t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X} \ ,$$
we can deduce that
\begin{align*}
&\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \left\{ \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}\right\}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\\
=&\left[\sum_{j=1}^{K_2}v_{K_2,j}(\boldsymbol{X})(\lambda_j-\lambda_j^*)^{\top}\right] \cdot \left\{ \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}\right\}\cdot \left[\sum_{l=1}^{K_2}(\lambda_l-\lambda_l^*)v_{K_2,l}(\boldsymbol{X})\right] \\
\leq & \eta_2 \cdot \|u_{K_1}(T)\|^2\cdot \left\|\sum_{i=1}^{K_2}(\lambda_i-\lambda_i^*)^{\top}v_{K_2,i}(\boldsymbol{X}) \right\|^2\\
\leq & \eta_2 \cdot \|u_{K_1}(T)\|^2\cdot \left(\sum_{i=1}^{K_2}\|\lambda_i-\lambda_i^*\|^2\right)\left(\sum_{i=1}^{K_2}v_{K_2,i}(\boldsymbol{X})^2\right)\\
=& \eta_2 \cdot \|u_{K_1}(T)\|^2\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^2\|v_{K_2}(\boldsymbol{X})\|^2 .
\end{align*}
Therefore, we can obtain that{\small
\begin{align}\label{eq:resubmit}
&\mathbb{E}\left[\left|\hat{S}_N - \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\| \right|^2\right] \notag\\
\leq
& \frac{1}{N} \eta^2_2 \cdot \mathbb{E}\left[ \|u_{K_1}(T)\|^4\cdot \|v_{K_2}(\boldsymbol{X})\|^4 \right]\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4 \notag\\
\leq& \frac{1}{N} \eta^2_2 \cdot \zeta_1(K_1)^2\cdot \zeta_2(K_2)^2\cdot \mathbb{E}\left[ \|u_{K_1}(T)\|^2\cdot \|v_{K_2}(\boldsymbol{X})\|^2 \right]\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4 \notag\\
=& \frac{1}{N} \eta^2_2 \cdot \zeta_1(K_1)^2\cdot \zeta_2(K_2)^2\cdot \mathbb{E}\left[\frac{1}{\pi_0(T,\boldsymbol{X})}\cdot\pi_0(T,\boldsymbol{X}) \|u_{K_1}(T)\|^2\cdot \|v_{K_2}(\boldsymbol{X})\|^2 \right]\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4 \notag\\
\leq& \frac{1}{N} \frac{\eta^2_2}{\eta_1} \cdot \zeta_1(K_1)^2\cdot \zeta_2(K_2)^2\cdot \mathbb{E}\left[\pi_0(T,\boldsymbol{X}) \|u_{K_1}(T)\|^2\cdot \|v_{K_2}(\boldsymbol{X})\|^2 \right]\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4\notag \quad (\text{by Assumption}\ \ref{as:pi0})\\
=&\frac{1}{N} \frac{\eta^2_2}{\eta_1} \cdot \zeta_1(K_1)^2\cdot \zeta_2(K_2)^2\cdot \mathbb{E}\left[ \|u_{K_1}(T)\|^2 \right]\cdot\mathbb{E}\left[ \|v_{K_2}(\boldsymbol{X})\|^2 \right]\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4\notag \quad (\text{by} \eqref{moment1})\\
=&\frac{1}{N} \frac{\eta^2_2}{\eta_1} \cdot \zeta_1(K_1)^2\cdot \zeta_2(K_2)^2\cdot K_1\cdot K_2\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4 \notag (\text{since}\ \mathbb{E}[ \|u_{K_1}(T)\|^2]=K_1\ \text{and}\ \mathbb{E}[ \|v_{K_2}(\boldsymbol{X})\|^2]=K_2)\\
=&\frac{1}{N} \frac{\eta^2_2}{\eta_1} \cdot \zeta(K)^2\cdot K\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4\ . \quad (\text{since}\ \zeta(K)=\zeta_1(K_1)\zeta_2(K_2)\ \text{and} \ K=K_1\cdot K_2)
\end{align}}
Considering the event set
$$E_{N}:= \left\{\hat{S}_N > \frac{1}{2}\left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2,\ \Lambda\neq \Lambda_{K_1\times K_2}^*\right\}\ ,$$
by Chebyshev's inequality, \eqref{eq:resubmit}, and Assumption \ref{as:K&N_c} we can get
\begin{align*}
&\mathbb{P}\left(\left|\hat{S}_N- \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2\right|\geq \frac{1}{2} \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2,\ \Lambda\neq \Lambda_{K_1\times K_2}^*\right)\\
&\leq \frac{4\mathbb{E}\left[\left|\hat{S}_N - \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\| \right|^2\right]}{\left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4} \\
& \leq\frac{4}{N} \frac{\eta^2_2}{\eta_1}\cdot \zeta(K)^2\cdot K \leq O\left( \frac{\zeta(K)^2K}{N}\right)=o(1),
\end{align*}
which implies that for any $\epsilon>0$, there exists $N_0(\epsilon)\in\mathbb{N}$ such that $N>N_0(\epsilon)$ large enough
\begin{align}\label{eq:resubmit_E}
\mathbb{P}\left((E_{N})^c\right)<\mathbb{P}\left(\left|\hat{S}_N- \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2\right|\geq \frac{1}{2} \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2,\ \Lambda\neq \Lambda_{K_1\times K_2}^*\right) <\frac{\epsilon}{2} \ .
\end{align}
Note that
\begin{align*}
&\frac{\partial}{\partial {\lambda_j}}\hat{G}_{K_1\times K_2}(\lambda_1,\ldots,\lambda_{K_2})\\
=&\frac{1}{N}\sum_{i=1}^N\rho'\left(u_{K_1}^{\top}(T_i)\Lambda v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)-\frac{1}{N^2}\sum_{i=1}^N\sum_{l=1}^Nu_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_l)\\
=& \frac{1}{N}\sum_{i=1}^N \left\{\rho'\left(u_{K_1}^{\top}(T_i)\Lambda v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})]u_{K_1}(T_i)\right\} \\
&-\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)\left\{\frac{1}{N}\sum_{l=1}^Nv_{K_2,j}(\boldsymbol{X}_l)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})] \right\}
\end{align*}
and
\begin{align*}
&\frac{\partial}{\partial {\lambda_j}}{G}^*_{K_1\times K_2}(\lambda_1,\ldots,\lambda_{K_2})=\mathbb{E}\left[\rho'\left(u_{K_1}^{\top}(T_i)\Lambda v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)\right]-\mathbb{E}[u_{K_1}(T)]\cdot \mathbb{E}[v_{K_2,j}(\boldsymbol{X})] .
\end{align*}
Since $\Lambda_{K_1\times K_2}^*$ is the unique maximizer of $G^*_{K_1\times K_2}(\cdot)$, then for each $j\in \{1,\ldots,K_2\}$,
\begin{align*}
&\frac{\partial}{\partial {\lambda_j}}{G}^*_{K_1\times K_2}(\lambda_1^*,\ldots,\lambda^*_{K_2})\\
=&\mathbb{E}\left[\rho'\left(u_{K_1}^{\top}(T)\Lambda_{K_1\times K_2}^* v_{K_2}(\boldsymbol{X})\right)u_{K_1}(T)v_{K_2,j}(\boldsymbol{X})\right]-\mathbb{E}[u_{K_1}(T)]\mathbb{E}[v_{K_2,j}(\boldsymbol{X})]=0.
\end{align*}
Therefore, for large enough $K$, we can deduce that{\small
\begin{align} \label{eq:EG2<}
&\mathbb{E}\left[\|\nabla\hat{G}_{K_1\times K_2}(\Lambda^*_{K_1\times K_2})\|^2\right] = \sum_{j=1}^{K_2}\mathbb{E}\left[\left\|\frac{\partial}{\partial {\lambda_j}}\hat{G}_{K_1\times K_2}(\lambda_1^*,\ldots,\lambda^*_{K_2})\right\|^2\right] \\
\leq & 2 \sum_{j=1}^{K_2}\mathbb{E}\left[\left\|\frac{1}{N}\sum_{i=1}^N \left\{\rho'\left(u_{K_1}^{\top}(T_i)\Lambda^*_{K_1\times K_2} v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})]u_{K_1}(T_i)\right\} \right\|^2\right]\notag \\
&+ 2\sum_{j=1}^{K_2}\mathbb{E}\left[\left\|\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)\left\{\frac{1}{N}\sum_{l=1}^Nv_{K_2,j}(\boldsymbol{X}_l)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})] \right\}\right\|^2\right] \notag \\
=& \frac{2}{N^2} \sum_{j=1}^{K_2}\sum_{i=1}^N \mathbb{E}\left[\left\|\rho'\left(u_{K_1}^{\top}(T_i)\Lambda^*_{K_1\times K_2} v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})]u_{K_1}(T_i)\right\|^2\right] \notag \\
&+ 2\sum_{j=1}^{K_2}\mathbb{E}\left[\left\|\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)\left\{\frac{1}{N}\sum_{l=1}^Nv_{K_2,j}(\boldsymbol{X}_l)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})] \right\}\right\|^2\right] \notag \\
\leq & \frac{2}{N^2} \sum_{j=1}^{K_2}\sum_{i=1}^N \mathbb{E}\left[\left\|\rho'\left(u_{K_1}^{\top}(T_i)\Lambda^*_{K_1\times K_2} v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})]u_{K_1}(T_i)\right\|^2\right] \notag \\
&+ 2\sum_{j=1}^{K_2}\mathbb{E}\left[\left|\frac{1}{N}\sum_{l=1}^Nv_{K_2,j}(\boldsymbol{X}_l)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})]\right|^2\right]\cdot \mathbb{E}\left[\left\|\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)\right\|^2\right] \notag \\
\leq & \frac{4}{N} \sum_{j=1}^{K_2} \left\{ \mathbb{E}\left[\left\|\rho'\left(u_{K_1}^{\top}(T)\Lambda^*_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)u_{K_1}(T)v_{K_2,j}(\boldsymbol{X})\right\|^2\right]\right. +\mathbb{E}[v_{K_2,j}(\boldsymbol{X})^2]\mathbb{E}\left[\|u_{K_1}(T)\|^2\right] \bigg\}\notag \\
&+ \frac{2}{N}\sum_{j=1}^{K_2}\mathbb{E}\left[v_{K_2,j}(\boldsymbol{X})^2\right]\cdot \mathbb{E}\left[\left\|u_{K_1}(T)\right\|^2\right] \notag \\
= & \frac{4}{N} \sum_{j=1}^{K_2} \left\{ \mathbb{E}\left[\frac{\left|\rho'\left(u_{K_1}^{\top}(T)\Lambda^*_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)\right|^2}{\pi_0(T,\bold{X})}\cdot \pi_0(T,\bold{X})\cdot \left\|u_{K_1}(T)v_{K_2,j}(\boldsymbol{X})\right\|^2\right]\right. +\mathbb{E}[v_{K_2,j}(\boldsymbol{X})^2]\mathbb{E}\left[\|u_{K_1}(T)\|^2\right] \bigg\}\notag \\
&+ \frac{2}{N}\sum_{j=1}^{K_2}\mathbb{E}\left[v_{K_2,j}(\boldsymbol{X})^2\right]\cdot \mathbb{E}\left[\left\|u_{K_1}(T)\right\|^2\right] \notag \\
\leq & \frac{4}{N} \sum_{j=1}^{K_2} \left\{ \frac{\left(\sup_{\gamma\in\Gamma_1}\rho'(\gamma)\right)^2}{\eta_1}\cdot \mathbb{E}\left[ \pi_0(T,\bold{X})\cdot \left\|u_{K_1}(T)v_{K_2,j}(\boldsymbol{X})\right\|^2\right]\right. +\mathbb{E}[v_{K_2,j}(\boldsymbol{X})^2]\mathbb{E}\left[\|u_{K_1}(T)\|^2\right] \bigg\}\notag \\
&+ \frac{2}{N}\sum_{j=1}^{K_2}\mathbb{E}\left[v_{K_2,j}(\boldsymbol{X})^2\right]\cdot \mathbb{E}\left[\left\|u_{K_1}(T)\right\|^2\right] \notag\\
= & \frac{4}{N} \sum_{j=1}^{K_2} \left\{ \frac{\left(\sup_{\gamma\in\Gamma_1}\rho'(\gamma)\right)^2}{\eta_1}\cdot \mathbb{E}\left[ v_{K_2,j}(\boldsymbol{X})^2\right] \mathbb{E}\left[ \left\|u_{K_1}(T)\right\|^2\right]\right. +\mathbb{E}[v_{K_2,j}(\boldsymbol{X})^2]\mathbb{E}\left[\|u_{K_1}(T)\|^2\right] \bigg\}\notag \\
&+ \frac{2}{N}\sum_{j=1}^{K_2}\mathbb{E}\left[v_{K_2,j}(\boldsymbol{X})^2\right]\cdot \mathbb{E}\left[\left\|u_{K_1}(T)\right\|^2\right] \notag\\
\leq & \frac{1}{N}\left\{ \frac{4}{\eta_1}\left(\sup_{\gamma\in\Gamma_1}\rho'(\gamma)\right)^2+4+2\right\}\cdot \mathbb{E}\left[\left\|u_{K_1}(T)\right\|^2\right]\sum_{j=1}^{K_2}\mathbb{E}\left[v_{K_2,j}(\boldsymbol{X})^2\right] \notag \\
=&\frac{1}{N}\left\{ \frac{4}{\eta_1}\left(\sup_{\gamma\in\Gamma_1}\rho'(\gamma)\right)^2+6\right\} K_1K_2= C_4^2 \frac{K}{N}, \notag
\end{align}}
where the last inequality follows by Assumption \ref{as:K&N_c} and $C_4:=\sqrt{ \frac{4}{\eta_1}\left(\sup_{\gamma\in\Gamma_1}\rho'(\gamma)\right)^2+6}$ is a finite universal constant.\\
Let $\epsilon>0$, fix $\dps C_5(\epsilon) > 0$ (to be chosen later) and define
\begin{align} \notag
\hat{\Upsilon}_{K_1\times K_2}(\epsilon):= \left\{\Lambda \in \mathbb{R}^{K_1\times K_2}: \|\Lambda - \Lambda_{K_1\times K_2}^*\|
\leq C_5(\epsilon) C_4 \sqrt{\frac {K} {N}} \right\}.
\end{align}
For $\forall \Lambda \in \hat{\Upsilon}_{K_1\times K_2}(\epsilon), \forall
(t,\boldsymbol{x}) \in\mathcal{T}\times \mathcal{X}$, we can have
\begin{align*}
& \left|u_{K_1}(t)^{\top}\Lambda v_{K_2}(\boldsymbol{x}) - u_{K_1}(t)^{\top}\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right|\\
\leq& \|\Lambda - \Lambda_{K_1\times K_2}^* \|\sup_{t\in\mathcal{T}}\|u_{K_1}(t)\| \sup_{\boldsymbol{x}\in\mathcal{X}}\|v_{K_2}(\boldsymbol{x})\| \leq C_5(\epsilon)C_4\sqrt{\frac {K} {N}} \zeta_1(K_1)\zeta_2(K_2) ,
\end{align*}
thus for large enough $N$, in accordance with Assumption \ref{as:K&N_c} and \eqref{eq:resubmit_lamdau_K}, we have
\begin{align}
&u_{K_1}(t)^{\top}\Lambda v_{K_2}(\boldsymbol{x})
\in \Bigg[ u_{K_1}(t)^{\top}\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})
- C_5(\epsilon)C_4 \zeta_1(K_1)\zeta_2(K_2) \sqrt{\frac {K} {N}}, \notag\\
& \quad \quad \quad \quad \quad \quad \quad \quad \quad u_{K_1}(t)^{\top}\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})
+ C_5(\epsilon)C_4 \zeta_1(K_1)\zeta_2(K_2) \sqrt{\frac {K} {N}} \Bigg] \notag \\
&\quad \quad \quad \quad \quad \quad \quad \subset \Bigg[ \underline{\gamma} - CK^{-\alpha} - C_5(\epsilon)C_4 \zeta_1(K_1)\zeta_2(K_2) \sqrt{\frac {K} {N}}, \notag \\
&~~~~~~~~~~~~~~~~~~~~~~~~~~~~\quad \quad \overline{\gamma} + CK^{-\alpha} + C_5(\epsilon)C_4 \zeta_1(K_1)\zeta_2(K_2) \sqrt{\frac {K} {N}}\Bigg] \subset \Gamma_2(\epsilon)\ , \label{Gamma_2_epsilon}
\end{align}
where $\Gamma_2(\epsilon):= \left[\underline{\gamma}-1-C_5(\epsilon), \overline{\gamma}+1+C_5(\epsilon)\right]$
is a compact set and independent of $(t,\boldsymbol{x})$. \\
For any $\Lambda \in \partial \hat{\Upsilon}_{K_1\times K_2}(\epsilon)$, there exists
$\bar{\Lambda}$ on the line joining $\Lambda$ and $\Lambda_{K_1\times K_2}^*$ such that
\begin{align*}
\hat{G}_{K_1\times K_2}(\Lambda) =& \hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*) + \sum_{j=1}^{K_2}(\lambda_j - \lambda_j^*)^{\top} \frac{\partial}{\partial \lambda_i}\hat{G}_{K_1\times K_2}(\lambda_1^*,\ldots,\lambda_{K_2}^*)\\
&+ \frac{1}{2}\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^*)^{\top} \frac{\partial^2}{\partial \lambda_i \partial \lambda_l}\hat{G}_{K_1\times K_2}(\bar{\lambda}_1,\ldots,\bar{\lambda}_{K_2})(\lambda_l - \lambda_l^*)\ ,
\end{align*}
where $\bar{\lambda}_j$ denotes the $j$-th column of $\bar{\Lambda}$.
For the second order term in above equality, note that $u_{K_1}^{\top}(t)\bar{\Lambda}v_{K_2}(\boldsymbol{x})\in\Gamma_2(\epsilon)$ for all $(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}$, we can further deduce that
\begin{align}\label{eq:resubmit_secondorder}
&\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^*)^{\top} \frac{\partial^2}{\partial \lambda_i \partial \lambda_l}\hat{G}_{K_1\times K_2}(\bar{\lambda}_1,\ldots,\bar{\lambda}_{K_2})(\lambda_l - \lambda_l^*) \\
= &\frac {1} {N} \sum_{i=1}^N\sum_{j=1}^{K_2}\sum_{l=1}^{K_2}(\lambda_j-\lambda^*_{j})^{\top}u_{K_1}(T_i)\rho''\left(u_{K_1}^{\top}(T_i)\bar{\Lambda}v_{K_2}(\boldsymbol{X}_i)\right)(\lambda_l - \lambda_l^*)^{\top}u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)v_{K_2,l}(\boldsymbol{X}_i) \notag\\
\leq & - \frac {\bar{b}(\epsilon)} {N} \sum_{i=1}^N\sum_{j=1}^{K_2}\sum_{l=1}^{K_2} (\lambda_j-\lambda^*_{j})^{\top}u_{K_1}(T_i) u_{K_1}(T_i)^{\top}(\lambda_l - \lambda_l^*)v_{K_2,j}(\boldsymbol{X}_i)v_{K_2,l}(\boldsymbol{X}_i)\notag\\
= & - \frac {\bar{b}(\epsilon)} {N} \sum_{i=1}^N\sum_{j=1}^{K_2}\sum_{l=1}^{K_2} \frac{1}{\pi_0(T_i,\boldsymbol{X}_i)} (\lambda_j-\lambda^*_{j})^{\top} \pi_0(T_i,\boldsymbol{X}_i)u_{K_1}(T_i) u_{K_1}(T_i)^{\top}(\lambda_l - \lambda_l^*)v_{K_2,j}(\boldsymbol{X}_i)v_{K_2,l}(\boldsymbol{X}_i)\notag\\
\leq & - \frac {\bar{b}(\epsilon)} {N \eta_2} \sum_{i=1}^N\sum_{j=1}^{K_2}\sum_{l=1}^{K_2} (\lambda_j-\lambda^*_{j})^{\top}\pi_0(T_i,\boldsymbol{X}_i) u_{K_1}(T_i) u_{K_1}(T_i)^{\top}(\lambda_l - \lambda_l^*)v_{K_2,j}(\boldsymbol{X}_i)v_{K_2,l}(\boldsymbol{X}_i)\notag\\
=& -\frac{\bar{b}(\epsilon)}{\eta_2}\hat{S}_N\ , \notag
\end{align}
where $-\bar{b}(\epsilon):=
\sup_{\gamma \in \Gamma_2(\epsilon)} \rho''(\gamma)<\infty$ for each fixed $\epsilon$.
Therefore, on the event $E_{N}$ and for large enough $N$, we can deduce that for any $\Lambda\in \partial\hat{\Upsilon}_{K_1\times K_2}(\epsilon)$,
\begin{align}\label{eq:resubmit_convex}
\text{on the event}\ E_N: &\quad \ \hat{G}_{K_1\times K_2}(\Lambda) - \hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*)\\
&= \sum_{j=1}^{K_2}(\lambda_j - \lambda_j^*)^{\top} \frac{\partial}{\partial \lambda_i}\hat{G}_{K_1\times K_2}(\lambda_1^*,\ldots,\lambda_{K_2}^*) \notag\\
&\quad + \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}\frac{1}{2}(\lambda_j - \lambda_j^*)^{\top} \frac{\partial^2}{\partial \lambda_i \partial \lambda_l}\hat{G}_{K_1\times K_2}(\bar{\lambda}_1,\ldots,\bar{\lambda}_{K_2})(\lambda_l - \lambda_l^*) \notag\\
& \leq \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\|\nabla\hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*)\|
-\frac{\bar{b}(\epsilon)}{2\eta_2} \hat{S}_N \ \text{(by \eqref{eq:resubmit_secondorder})} \notag\\
&\leq \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\|\nabla\hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*)\|
- \frac {\bar{b}(\epsilon)}{4\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|^2 \notag\\
& = \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\left(\|\nabla\hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*)\| -
\frac {\bar{b}(\epsilon)}{4\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\right) \ , \notag
\end{align}
where the second inequality follows from definition of the event $E_{N}$. \\
Note that for sufficiently large $N$, by Chebyshev's inequality and \eqref{eq:EG2<} we have
\begin{align} \label{eq:G^1<}
&\mathbb{P}\left\{\|\nabla\hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*)\| \geq
\frac {\bar{b}(\epsilon)}{4\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\right\}\\ \leq&\frac{16\eta_2^2}{\bar{b}(\epsilon)^2}\cdot \frac{\mathbb{E}\left[\left\|\nabla\hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*)\right\|^2\right]}{ \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|^2}
\leq\frac{16\eta_2^2}{\bar{b}(\epsilon)^2C_5^2(\epsilon)} \leq \frac{\epsilon}{2}\ ,\notag
\end{align}
where the last inequality holds by choosing $$C_5(\epsilon) \geq \sqrt{\frac{32\eta_2^2}{\bar{b}(\epsilon)^2\epsilon}}\ .$$
Therefore, for sufficiently large $N$, by \eqref{eq:resubmit_E} and \eqref{eq:G^1<} we can derive
\begin{align} \label{eq:resubmit_events}
&\mathbb{P}\left( (E_{N})^c \ \text{or} \ \|\nabla\hat{G}_{K_1\times K_2}(\Lambda^*_{K_1\times K_2})\| \geq
\frac {\bar{b}(\epsilon)}{2\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\right)
\leq \frac {\epsilon} {2} + \frac {\epsilon} {2} = \epsilon \notag \\
\Rightarrow &\mathbb{P}\left( E_{N} \ \text{and} \ \|\nabla\hat{G}_{K_1\times K_2}(\Lambda^*_{K_1\times K_2})\| <
\frac {\bar{b}(\epsilon)}{2\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\right)
> 1-\epsilon.
\end{align}
With \eqref{eq:resubmit_convex} and \eqref{eq:resubmit_events}, we can obtain that
$$\mathbb{P}\left\{\hat{G}_{K_1\times K_2}(\Lambda) - \hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*) < 0 ,~~ \forall \Lambda \in \partial\hat{\Upsilon}_{K_1\times K_2}(\epsilon) \right\} \geq 1 - \epsilon\ .$$
Note that the event $\left\{ \hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*) > \hat{G}_{K_1\times K_2}(\Lambda) ,~ \forall \Lambda \in \partial\hat{\Upsilon}_{K_1\times K_2}(\epsilon) \right\}$ implies that there exists a local maximizer in the interior of $\hat{\Upsilon}_{K_1\times K_2}(\epsilon)$. Since $\hat{G}_{K_1\times K_2}(\cdot)$ is strictly concave and $\hat{\Lambda}_{K_1\times K_2} $ is the unique global maximizer of $\hat{G}_{K_1\times K_2}$, then
\begin{align}\label{eq:Lambdahat_in_Upsilonhat}
\mathbb{P}\left(\hat{\Lambda}_{K_1\times K_2} \in \hat{\Upsilon}_{K_1\times K_2}(\epsilon)\right)>1-\epsilon ,
\end{align}
i.e.
$ \left\|\hat{\Lambda}_{K_1\times K_2}- \Lambda_{K_1\times K_2}^*\right\| = O_p\left(\sqrt{\frac {K} {N}} \right)$.
\end{proof}
\subsection{Corollary $\ref{cor:pi^-pi*}$}
The next corollary states that $\hat{\pi}_K(t,\boldsymbol{x})$ is arbitrarily close to ${\pi}^*_K(t,\boldsymbol{x})$.
\begin{cor}\label{cor:pi^-pi*}
Under Assumption \ref{as:suppX}-\ref{as:K&N_consistency}, we have
$$\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\hat{\pi}_K(t,\boldsymbol{x})-{\pi}^*_K(t,\boldsymbol{x})|=O_p\left(\zeta(K)\sqrt{\frac{K}{N}}\right),$$
and
$$\int_{\mathcal{T}\times \mathcal{X}}|\hat{\pi}_K(t,\boldsymbol{x})-{\pi}^*_K(t,\boldsymbol{x})|^2dF_{T,X}(t,\boldsymbol{x})=O_p\left(\frac{K}{N}\right),$$
and
$$\frac{1}{N}\sum_{i=1}^N|\hat{\pi}_K(T_i,\boldsymbol{X}_i)-{\pi}^*_K(T_i,\boldsymbol{X}_i)|^2=O_p\left(\frac{K}{N}\right).$$
\end{cor}
\begin{proof}
From the proof of Lemma \ref{lemma_pi^hat}, we know the facts $\mathbb{P}\left(\hat{\Lambda}_{K_1\times K_2}\in \hat{\Upsilon}_{K_1\times K_2}(\epsilon)\right)>1-\epsilon$ and \eqref{Gamma_2_epsilon}. Then for any element $\tilde{\Lambda}_{K_1\times K_2}$ lying on the line joining $\hat{\Lambda}_{K_1\times K_2}$ and $\Lambda_{K_1\times K_2}^*$, we can have that $\mathbb{P}(u_{K_1}(t)^{\top}\tilde{\Lambda}_{K_1\times K_2} v_{K_2}(\boldsymbol{x})\in \Gamma_{2}(\epsilon)$ for all $(t,\boldsymbol{x})\in \mathcal{T}\times\mathcal{X})\geq 1-\epsilon$, which implies
\begin{align}\label{eq:rho''_tilde_Op(1)}
\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\rho''(u_{K_1}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x}))|=O_p(1).
\end{align}
Using Mean Value Theorem, Lemma \ref{lemma_pi^*}, and \eqref{eq:rho''_tilde_Op(1)}, we can obtain that
\begin{align} \notag
&\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\hat{\pi}_K(t,\boldsymbol{x})-{\pi}^*_K(t,\boldsymbol{x})|\\
=&\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\rho'\left(u_{K_1}(t)\hat{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right) - \rho'\left(u_{K_1}(t)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)| \notag \\
\leq &\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\rho''(u_{K_1}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x}))|\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}\left|u_{K_1}(t)\hat{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})-u_{K_1}(t)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right| \notag\\
\leq & O_p(1) \cdot \|\hat{\Lambda}_{K_1\times K_2}-\Lambda_{K_1\times K_2}^* \| \cdot \sup_{t\in\mathcal{T}}\|u_{K_1}(t)\| \cdot\sup_{\boldsymbol{x}\in \mathcal{X}}\|v_{K_2}(\boldsymbol{x})\| \notag \\
\leq & O_p(1)\cdot O_p\left(\sqrt{\frac {K} {N}} \right) \zeta_1(K_1)\cdot \zeta_2(K_2)=O_p\left(\zeta(K)\sqrt{\frac{K}{N}}\right).\notag
\end{align}
Note that by Mean Value Theorem and \eqref{eq:rho''_tilde_Op(1)}, we can deduce that
\begin{align*}
&\int_{\mathcal{T}\times \mathcal{X}}|\hat{\pi}_K(t,\boldsymbol{x})-{\pi}^*_K(t,\boldsymbol{x})|^2dF_{T,X}(t,\boldsymbol{x})\\
\leq&\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\rho''(u_{K_1}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x}))|^2\int_{\mathcal{T}\times \mathcal{X}}\left|u_{K_1}(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right|^2dF_{T,X}(t,\boldsymbol{x})\\
\leq &O_p(1)\cdot \int_{\mathcal{T}\times \mathcal{X}}\left|u_{K_1}(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right|^2dF_{T,X}(t,\boldsymbol{x}) .
\end{align*}
We estimate $\int_{\mathcal{T}\times \mathcal{X}}\left|u_{K_1}(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right|^2dF_{T,X}(t,\boldsymbol{x})$. Note that $\mathbb{E}[u_{K_1}(T)[u_{K_1}(T)^\top]=I_{K_1\times K_1}$, $\mathbb{E}[v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^\top]=I_{K_2\times K_2}$, \eqref{moment1} and Assumption \ref{as:pi0}, we can deduce that{\small
\begin{align}
&\int_{\mathcal{T}\times \mathcal{X}}\left|u_{K_1}(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right|^2dF_{T,X}(t,\boldsymbol{x}) \notag\\
\leq & \int_{\mathcal{T}\times \mathcal{X}}u_{K_1}^\top(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})v_{K_2}(\boldsymbol{x})^\top\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top u_{K_1}(t) dF_{T,X}(t,\boldsymbol{x}) \notag\\
=& \int_{\mathcal{T}\times \mathcal{X}}\frac{1}{\pi_0(t,\boldsymbol{x})}\pi_0(t,\boldsymbol{x}) u_{K_1}^\top(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})v_{K_2}(\boldsymbol{x})^\top\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top u_{K_1}(t) dF_{T,X}(t,\boldsymbol{x}) \notag\\
\leq & \frac{1}{\eta_1} \int_{\mathcal{T}\times \mathcal{X}}\pi_0(t,\boldsymbol{x})\cdot u_{K_1}^\top(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})v_{K_2}(\boldsymbol{x})^\top\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top u_{K_1}(t) dF_{T,X}(t,\boldsymbol{x})\notag \\
= & \frac{1}{\eta_1} \int_{\mathcal{T}} u_{K_1}^\top(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}\left(\int_{\mathcal{X}}v_{K_2}(\boldsymbol{x})v_{K_2}(\boldsymbol{x})^\top dF_X(x)\right)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top u_{K_1}(t) dF_{T}(t) \notag\\
= & \frac{1}{\eta_1} \int_{\mathcal{T}} u_{K_1}^\top(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top u_{K_1}(t) dF_{T}(t) \notag\\
= & \frac{1}{\eta_1} \tr\Bigg( \left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top \int_{\mathcal{T}}u_{K_1}(t)u_{K_1}^\top(t) dF_{T}(t)\Bigg) \notag\\
= & \frac{1}{\eta_1} \tr\Bigg( \left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top \Bigg)\notag \\
=& \frac{1}{\eta_1}\cdot \left\|\hat{\Lambda}_{K_2\times K_2}-\Lambda^*_{K_1\times K_2}\right\|^2 =O_p\left(\frac{K}{N}\right). \label{eq:u_Lambda_v}
\end{align}}
Then we obtain
\begin{align*}
\int_{\mathcal{T}\times \mathcal{X}}|\hat{\pi}_K(t,\boldsymbol{x})-{\pi}^*_K(t,\boldsymbol{x})|^2dF_{T,X}(t,\boldsymbol{x})=O_p\left(\frac{K}{N}\right).
\end{align*}
Similar to \eqref{eq:uv-E[uv]}, we have
{\small
\begin{align}\label{eq:uv-E[uv]_hat}
&\frac{1}{N}\sum_{i=1}^N \left|u_{K_1}^\top(T_i)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}_i) \right|^2-\int_{\mathcal{T}\times \mathcal{X}}\left|u_{K_1}(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right|^2dF_{T,X}(t,\boldsymbol{x})\notag \\
=&O_p\left(\frac{\zeta(K)}{\sqrt{N}}\cdot \|\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\|^2\right)=O_p\left(\frac{\zeta(K)}{\sqrt{N}}\cdot \frac{K}{N}\right)=o_p\left(\frac{K}{N}\right).
\end{align}}
where the last equality holds in light of Assumption \ref{as:K&N_consistency}. Hence, with \eqref{eq:u_Lambda_v} and \eqref{eq:uv-E[uv]_hat}, we have
\begin{align*}
&\frac{1}{N}\sum_{i=1}^N|\hat{\pi}_K(T_i,\boldsymbol{X}_i)-{\pi}^*_K(T_i,\boldsymbol{X}_i)|^2 \\
\leq&\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\rho''(u_{K_1}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x}))|^2\cdot \frac{1}{N}\sum_{i=1}^N\left|u_{K_1}(T_i)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}_i)\right|^2\\
\leq &O_p(1)\cdot \int_{\mathcal{T}\times \mathcal{X}}\left|u_{K_1}(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right|^2dF_{T,X}(t,\boldsymbol{x})+o_p\left(\frac{K}{N}\right) \\
\leq &O_p\left(\frac{K}{N}\right)+o_p\left(\frac{K}{N}\right)=O_p\left(\frac{K}{N}\right).
\end{align*}
\end{proof}
\section{Efficient Estimation}
\subsection{Proof of Theorem 5.17 \label{sec:proof_main_theorem}}
Since we have proved $\|\hat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0\|\xrightarrow{p}0$ in Theorem 10,
we now begin to prove the asymptotic efficiency of $\hat{\boldsymbol{\beta}}$. By Assumption \ref{as:first_order}, $\hat{\boldsymbol{\beta}}$ is a unique solution of the following equation:
\begin{align}\label{eq:beta^hat}
\frac{1}{N}\sum_{i=1}^{N}\hat{\pi}_{K}(T_{i},\boldsymbol{X}_{i})m(T_{i};
\hat{\boldsymbol{\beta }})L'\left\{ Y_{i}-g\left( T_{i};\hat{\boldsymbol{\beta
}}\right) \right\} =0 ,
\end{align}
with probability approaching to one. Note that $L'(\cdot)$ may be a non-differentiable function, e.g. $L'(v)=\tau-I(v\leq 0)$ in quantile regression, we cannot simply apply Mean Value Theorem on \eqref{eq:beta^hat} to obtain the expression for $\sqrt{N}(\hat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0)$. To solve this problem, we define
\begin{align*}
f(\boldsymbol{\beta}):=\mathbb{E}\left[\pi_0 (T,\boldsymbol{X})L'(Y
-g(T;\boldsymbol{\beta} ))m(T;\boldsymbol{\beta})\right],
\end{align*}
which is a differentiable function in $\boldsymbol{\beta}$ and by definition $f(\boldsymbol{\beta}_0)=0$. Using Mean Value Theorem, we can obtain that
\begin{align*}
0=\sqrt{N}f(\boldsymbol{\beta}_0)=\sqrt{N}f(\hat{\boldsymbol{\beta}})-\nabla_{\beta}f(\tilde{\boldsymbol{\beta}})\cdot \sqrt{N}(\hat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0)\ ,
\end{align*}
where $\tilde{\boldsymbol{\beta}}$ lies on the line joining $\hat{\boldsymbol{\beta}}$ and $\boldsymbol{\beta}_0$. Because $\nabla_{\beta}f(\boldsymbol{\beta})$ is continuous in $\boldsymbol{\beta}$ at $\boldsymbol{\beta}_0$, and $\|\hat{\boldsymbol{\beta}}-{\boldsymbol{\beta}}_0\|\xrightarrow{p}0$, then we have
\begin{align*}
\sqrt{N}(\hat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0)=&\nabla_{\beta}f(\boldsymbol{\beta}_0)^{-1}\cdot \sqrt{N}f(\hat{\boldsymbol{\beta}}).
\end{align*}
Define the empirical process:
\begin{align*}
\mu_N(\boldsymbol{\beta})=\frac{1}{\sqrt{N}}\sum_{i=1}^N \bigg\{\hat{\pi}_K (T_{i},\boldsymbol{X}_{i})L'(Y
_{i} -g(T_{i};\boldsymbol{\beta} ))m(T_i;\boldsymbol{\beta})-\mathbb{E}\left[\pi_0 (T,\boldsymbol{X})L'(Y
-g(T;\boldsymbol{\beta} ))m(T;\boldsymbol{\beta})\right]\bigg\}\ ,
\end{align*}
and by Assumption \ref{as:first_order} we can have
\begin{align*}
&\sqrt{N}(\hat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0)\\
=&\nabla_{\beta}f(\boldsymbol{\beta}_0)^{-1}\cdot \Bigg\{ \sqrt{N}f(\hat{\boldsymbol{\beta}})-\frac{1}{\sqrt{N}}\sum_{i=1}^N\hat{\pi}_K (T_{i},\boldsymbol{X}_{i})L'(Y
_{i} -g(T_{i};\hat{\boldsymbol{\beta}} ))m(T_i;\hat{\boldsymbol{\beta}})\\ &\qquad \qquad \qquad +\frac{1}{\sqrt{N}}\sum_{i=1}^N\hat{\pi}_K (T_{i},\boldsymbol{X}_{i})L'(Y
_{i} -g(T_{i};\hat{\boldsymbol{\beta}} ))m(T_i;\hat{\boldsymbol{\beta}}) \Bigg\}\\
=&-\nabla_{\beta}f(\boldsymbol{\beta}_0)^{-1}\cdot \mu_N(\hat{\boldsymbol{\beta}})+o_p(1)\\
=&H_0^{-1}\cdot \bigg\{\left(\mu_N(\hat{\boldsymbol{\beta}}) -\mu_N(\boldsymbol{\beta}_0)\right) +\mu_N(\boldsymbol{\beta}_0)\bigg\}+o_p(1).
\end{align*}
By Assumption \ref{as:m_smooth} and \ref{as:entropy}, Theorems 4 and 5 of \cite{andrews1994empirical}, we can conclude that $\mu_N(\cdot)$ is stochastically equicontinuous, which implies $\mu_N(\hat{\boldsymbol{\beta}}) -\mu_N(\boldsymbol{\beta}_0)\xrightarrow{p}0$. We also note that $\mathbb{E}[\pi_0 (T,\boldsymbol{X})L'(Y
-g(T;\boldsymbol{\beta}_0 ))m(T;\boldsymbol{\beta}_0)]=0$, then
\begin{align}\label{eq:beta^hat-beta0}
\sqrt{N}(\hat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0)=H_0^{-1}\frac{1}{\sqrt{N}}\sum_{i=1}^N \hat{\pi}_K (T_{i},\boldsymbol{X}_{i})L'(Y
_{i} -g(T_{i},\beta_0 ))m(T_i;\boldsymbol{\beta}_0)+o_p(1).
\end{align}
We next claim the following important Lemma \ref{asy:equivalent}, and leave its proof to Section \ref{sec:proof_equivalent}.
\begin{lemma}
\label{asy:equivalent} Under Assumption \ref{as:TYindep}-\ref{as:K&N_c}, we have
\begin{align} \label{asy:equivalent_1}
\frac{1}{\sqrt{N}}\sum_{i=1}^N\hat{\pi}_K(T_i,\boldsymbol{X}_i)m(T_i;{
\boldsymbol{\beta}}_0)L'\left\{Y_i-g\left(T_i;{\boldsymbol{\beta}}
_0\right)\right\}=\frac{1}{\sqrt{N}}\sum_{i=1}^N\psi(Y_i,T_i,\boldsymbol{X}
_i;\boldsymbol{\beta}_0) +o_p(1),
\end{align}
where \begin{align*}
&\psi (Y,T,\boldsymbol{X};\boldsymbol{\beta }_{0}):= \pi_0 (T,
\boldsymbol{X})m(T;\boldsymbol{\beta }_{0})L'(Y-g(T;\boldsymbol{\beta}_0))-\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta }_{0})\varepsilon (T,\boldsymbol{X};
\boldsymbol{\beta }_{0}) \\
&\qquad \qquad\qquad+\mathbb{E}\left[ \varepsilon (T,\boldsymbol{X};
\boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta }_{0})|
\boldsymbol{X}\right] +\mathbb{E}\left[ \varepsilon (T,\boldsymbol{X};
\boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta }_{0})|
T\right] \ ,
\end{align*}
and $\varepsilon (T,\boldsymbol{X};\boldsymbol{\beta }_{0}):= \mathbb{E}
[L'(Y-g(T;\boldsymbol{\beta}_0))|T,\boldsymbol{X}]$.
\end{lemma}
Lemma \ref{asy:equivalent} is the most important
step for establishing the efficiency of our proposed estimator. A key technique in proving Lemma \ref
{asy:equivalent} is a use of a weighted least square projection of
$L'(Y-g(T;\boldsymbol{\beta}_0))$ onto the space linearly spanned by the approximation
basis $\{u_{K_{1}}(T),v_{K_{2}}(\boldsymbol{X})\}$.
Combining \eqref{eq:beta^hat-beta0} and Lemma \ref
{asy:equivalent}, we can obtain the asymptotic expression for $\sqrt{N}(\hat{
\boldsymbol{\beta }}-\boldsymbol{\beta }_{0})$:
\begin{align*}
\sqrt{N}(\hat{\boldsymbol{\beta }}-\boldsymbol{\beta }_{0})= H_0^{-1}\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\psi (T_{i},\boldsymbol{X}
_{i},Y_{i};{\boldsymbol{\beta }}_{0})+o_{p}(1) = \frac{1}{\sqrt{N}}\sum_{i=1}^{N}S_{eff}(T_{i},\boldsymbol{X}_{i},Y_{i};{
\boldsymbol{\beta }}_{0})+o_{p}(1),
\end{align*}
which leads to our Theorem 5.17.
\subsection{Proof of Lemma \ref{asy:equivalent} \label{sec:proof_equivalent}}
Before proving Lemma \ref{asy:equivalent}, we prepare some preliminary notation and results that will be used later. Since $\hat{\Lambda}_{K_1\times K_2}$ is a unique maximizer of the concave function $\hat{G}_{K_1\times K_2}$, then
\begin{align*}
\frac{1}{N}\sum_{i=1}^N\rho'\left(u_{K_1}(T_i)^{\top}\hat{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2}(\boldsymbol{X}_i)^{\top}-\frac{1}{N^2}\sum_{i=1}^N\sum_{l=1}^Nu_{K_1}(T_l)v_{K_2}(\boldsymbol{X}_i)^{\top}=0.
\end{align*}
Using Mean Value Theorem, we can have
\begin{align}\label{aa}
&\frac{1}{N}\sum_{i=1}^N\rho'\left(u_{K_1}(T_i)^{\top}{\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2}(\boldsymbol{X}_i)^{\top} \notag\\
+& \frac{1}{N}\sum_{i=1}^N\rho''\left(u_{K_1}(T_i)^{\top}\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)u_{K_1}(T_i)^{\top} \left\{\hat{\Lambda}_{K_1\times K_2}-{\Lambda}^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}_i)v_{K_2}(\boldsymbol{X}_i)^{\top} \notag\\
=&\frac{1}{N^2}\sum_{i=1}^N\sum_{l=1}^N u_{K_1}(T_l)v_{K_2}(\boldsymbol{X}_i)^{\top} \ ,
\end{align}
where $\tilde{\Lambda}_{K_1\times K_2}$ lies on the line joining from $\hat{\Lambda}_{K_1\times K_2}$ to ${\Lambda}^*_{K_1\times K_2}$. We define the following notation:
\begin{align}
&\hat{A}_{K_1\times K_2} := \hat{\Lambda}_{K_1\times K_2}-\Lambda_{K_1\times K_2}^* , \label{def:A} \\
&\tilde{A}_{K_1\times K_2} := \tilde{\Lambda}_{K_1\times K_2}-\Lambda_{K_1\times K_2}^* , \label{def:A_tilde}
\end{align}
and
\begin{align}\label{def:A*}
&A_{K_1\times K_2}^*:= \nabla\hat{G}_{K_1\times K_2}\left(\Lambda_{K_1\times K_2}^*\right) \notag \\
=&\frac{1}{N}\sum_{i=1}^N\rho'\left(u_{K_1}(T_i)^{\top}\Lambda_{K_1\times K_2}^* v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2}(\boldsymbol{X}_i)^{\top}-\left(\frac{1}{N}\sum_{l=1}^N u_{K_1}(T_l)\right)\left(\frac{1}{N}\sum_{i=1}^N v_{K_2}(\boldsymbol{X}_i)^{\top}\right).
\end{align}
In light of \eqref{eq:EG2<} we have
$$\left\|A^*_{K_1\times K_2}\right\|= O_p\left(\sqrt{\frac{K}{N}}\right). $$
From \eqref{aa}, $A^*_{K_1\times K_2}$ can also be written as
\begin{align} \label{eq:A^*_{K_1timesK_2}}
A^*_{K_1\times K_2}=-\frac{1}{N}\sum_{i=1}^N\rho''\left(u_{K_1}(T_i)^{\top}\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)u_{K_1}(T_i)^{\top} \left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda_{K_1\times K_2}^*\right\}v_{K_2}(\boldsymbol{X}_i)v_{K_2}(\boldsymbol{X}_i)^{\top}.
\end{align}
We now start to prove Lemma \ref{asy:equivalent}. We decompose $\frac{1}{\sqrt{N}}\sum_{i=1}^N\hat{\pi}_K(T_i,\boldsymbol{X}_i)\left\{Y_i-g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m(T_i;\boldsymbol{\beta}_0)$ as follows:
\begin{align}
&\frac{1}{\sqrt{N}}\sum_{i=1}^N\hat{\pi}_K(T_i,\boldsymbol{X}_i)L'\left\{Y_i-g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m(T_i;\boldsymbol{\beta}_0)\notag\\
=&\label{eq:WK} \frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg\{ \left(\hat{\pi}_K(T_i,\boldsymbol{X}_i) - \pi_K^*(T_i,\boldsymbol{X}_i)\right)L'\left\{Y_i-g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m(T_i;\boldsymbol{\beta}_0) \\
&\quad \quad \quad \quad \quad -\int_{\mathcal{T}}\int_{\mathcal{X}} \left(\hat{\pi}_K(t,\boldsymbol{x}) - \pi_K^*(t,\boldsymbol{x})\right)\varepsilon(\boldsymbol{x},t;\boldsymbol{\beta}_0)m(t;\boldsymbol{\beta}_0) dF_{X,T}(\boldsymbol{x},t)\Bigg\} \notag \\[2mm]
\label{eq:VK} &~ + \frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg\{\left(\pi_K^*(T_i,\boldsymbol{X}_i) - {\pi_0(T_i,\boldsymbol{X}_i)}\right)L'\left\{Y_i-g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m(T_i;\boldsymbol{\beta}_0) \\
&\quad \quad \quad \quad \quad - \int_{\mathcal{T}}\int_{\mathcal{X}} m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\left(\pi_K^*(t,\boldsymbol{x}) - {\pi_0(t,\boldsymbol{x})}\right) dF_{X,T}(\boldsymbol{x},t) \Bigg\} \notag \\
\label{eq:Lemma1}&~ + \sqrt{N}\int_{\mathcal{T}}\int_{\mathcal{X}} m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\left(\pi_K^*(t,\boldsymbol{x}) - {\pi_0(t,\boldsymbol{x})}\right) dF_{X,T}(\boldsymbol{x},t) \\[2mm]
&~ \label{eq:tauto}+ {\sqrt{N}}\int_{\mathcal{T}}\int_{\mathcal{X}} \left(\hat{\pi}_K(t,\boldsymbol{x}) - \pi_K^*(t,\boldsymbol{x})\right)\varepsilon(\boldsymbol{x},t;\boldsymbol{\beta}_0)m(t;\boldsymbol{\beta}_0) dF_{X,T}(\boldsymbol{x},t) \\
& \qquad - \sqrt{N}\int_{\mathcal{T}}\int_{\mathcal{X}} \varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) m(t;\boldsymbol{\beta}_0) dF_{X,T}(\boldsymbol{x},t) \notag \\[4mm]
& \label{eq:Q} +\sqrt{N} \int_{\mathcal{T}}\int_{\mathcal{X}} \varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) m(t;\boldsymbol{\beta}_0) dF_{X,T}(\boldsymbol{x},t)\\
&\qquad - \sqrt{N} \int_{\mathcal{T}}\int_{\mathcal{X}} \varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)A^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) m(t;\boldsymbol{\beta}_0) dF_{X,T}(\boldsymbol{x},t)\notag \\[2mm]
& \label{eq:proj}+{\sqrt{N}}\int_{\mathcal{X}} \varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)A^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) m(t;\boldsymbol{\beta}_0) dF_{X,T}(\boldsymbol{x},t)\\
&\qquad +\frac{1}{\sqrt{N}}\sum_{i=1}^N \bigg\{ \pi_0(T_i,\boldsymbol{X}_i)m(T_i;\boldsymbol{\beta}_0)\varepsilon(T_i,\boldsymbol{X}_i;\boldsymbol{\beta}_0)-\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{x};\boldsymbol{\beta}_0)|\boldsymbol{X}=\boldsymbol{X}_i\right]\notag\\
&\qquad\qquad \qquad \qquad -\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{x};\boldsymbol{\beta}_0)|T=T_i\right]\bigg\} \notag \\[2mm]
&~ + \frac {1} {\sqrt{N}} \sum_{i=1}^N \bigg\{ \pi_0(T_i,\boldsymbol{X}_i)L'\left\{Y_i-g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m(T_i;\boldsymbol{\beta}_0)-\pi_0(T_i,\boldsymbol{X}_i)m(T_i;\boldsymbol{\beta}_0)\varepsilon(T_i,\boldsymbol{X}_i;\boldsymbol{\beta}_0)\label{eq:Normal}\\
& \quad \quad \quad +\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)|\boldsymbol{X}=\boldsymbol{X}_i\right]+\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{x};\boldsymbol{\beta}_0)|T=T_i\right] \bigg\}, \notag
\end{align}
where $\hat{A}_{K_1\times K_2}$ and $A_{K_1\times K_2}^*$ are defined in \eqref{def:A} and \eqref{eq:A^*_{K_1timesK_2}}. We show that the terms \eqref{eq:WK}-\eqref{eq:proj} are all of $o_p(1)$, while the term \eqref{eq:Normal} is asymptotically normal.\\
\noindent \textbf{\emph{For term \eqref{eq:WK}}:}
\noindent Denoting \eqref{eq:WK} by $W_K$ and applying Mean Value Theorem twice, we can obtain
\begin{align*}
W_K = &\frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg[L'\left\{Y_i-g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m(T_i;\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(T_i)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i) \\
&-\int_{\mathcal{T}} \int_{\mathcal{X}} m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0) \rho''\left(u_{K_1}^{\top}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\Bigg]\\
=&\frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg[L'\left\{Y_i-g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m(T_i;\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(T_i){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right) u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i) \\
&- \int_{\mathcal{T}}\int_{\mathcal{X}} m(t;\boldsymbol{\beta}_0) \varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\Bigg]\\
&+\frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg[L'\left\{Y_i -g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m(T_i;\boldsymbol{\beta}_0)\rho'''\left(\xi_3(T_i,\boldsymbol{X}_i)\right)\left\{u_{K_1}(T_i)^{\top}\tilde{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right\}\\
& \quad \quad \quad \quad \quad \times u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\bigg] \\
&-\sqrt{N}\int_{\mathcal{T}} \int_{\mathcal{X}} m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho'''\left(\xi_3(t,\boldsymbol{x})\right)\left\{u_{K_1}(t)^{\top}\tilde{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right\} \\
&\qquad \qquad \qquad\times u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\\
= & W_{1K} + W_{2K} + W_{3K},
\end{align*}
where $\tilde{A}_{K_1\times K_2}$ is defined in \eqref{def:A_tilde}, and{\footnotesize
\begin{align*}
W_{1K} := &~\frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg[L'\left\{Y_i-g\left( T_i;\boldsymbol{\beta}_0\right)\right\}m(T_i;\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(T_i){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right) u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i) \\
&\qquad \qquad - \int_{\mathcal{T}}\int_{\mathcal{X}} m(t;\boldsymbol{\beta}_0) \varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\Bigg]\ ,\\
W_{2K} := &~\frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg[L'\left\{Y_i-g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m(T_i;\boldsymbol{\beta}_0)\rho'''\left(\xi_3(T_i,\boldsymbol{X}_i)\right)\left\{u_{K_1}(T_i)^{\top}\tilde{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right\} u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\bigg]\ , \\
W_{3K} := &~ -{\sqrt{N}} \int_{\mathcal{T}}\int_{\mathcal{X}} m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho'''\left(\xi_3(t,\boldsymbol{x})\right)\left\{u_{K_1}(t)^{\top}\tilde{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right\} u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\ ,
\end{align*}}
and $\xi_3(t,\boldsymbol{x})$ lies between $u_{K_1}(t)\tilde{\Lambda}_{K_1\times K_2}^{\top}v_{K_2}(\boldsymbol{x})$ and $u_{K_1}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})$. \\
For the term $W_{1K}$, we denote its $k^{th}$ component by $W_{1K,k}$, $k=1,\ldots,p$, and we also let $m_k(T_i;\boldsymbol{\beta}_0)$ be the $k^{th}$ component of $m(T_i;\boldsymbol{\beta}_0)$, i.e.,
\begin{align*}
W_{1K,k} &:= \frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg[L'\left\{Y_i-g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m_{k}(T_i;\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(T_i){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right) u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i) \\
&- \int_{\mathcal{T}}\int_{\mathcal{X}} m_k(t;\boldsymbol{\beta}_0) \varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\Bigg]\\
=&\tr \bigg\{ \frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg[L'\left\{Y_i-g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m_{k}(T_i;\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(T_i){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right) v_{K_2}(\boldsymbol{X}_i)u_{K_1}(T_i)^{\top} \\
&- \int_{\mathcal{T}}\int_{\mathcal{X}} m_k(t;\boldsymbol{\beta}_0) \varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)v_{K_2}(\boldsymbol{x})u_{K_1}^{\top}(t)dF_{X,T}(\boldsymbol{x},t)\Bigg]\hat{A}_{K_1\times K_2} \bigg\}\\
=& \tr \left\{U_{K_2\times K_1}(k) \hat{A}_{K_1\times K_2}\right\} ,
\end{align*}
where
\begin{align*}
U_{K_2\times K_1}(k):= & \frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg[L'\left\{Y_i-g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m_{k}(T_i;\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(T_i){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)v_{K_2}(\boldsymbol{X}_i)u_{K_1}(T_i)^{\top} \\
&- \int_{\mathcal{T}} \int_{\mathcal{X}} m_k(t;\boldsymbol{\beta}_0) \varepsilon (t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)v_{K_2}(\boldsymbol{x})u_{K_1}^{\top}(t)dF_{X,T}(\boldsymbol{x},t)\Bigg].
\end{align*}
We compute the second moment of $U_{K_2\times K_1}(k)$ to get that
\begin{align*}
&\mathbb{E}\left[\|U_{K_2\times K_1}(k)\|^2\right] = \mathbb{E}[\tr\{(U_{K_2\times K_1}(k))^{\top}U_{K_2\times K_1}(k)\}]\\
=& \mathbb{E}\left[L'\left\{Y -g\left(T;\boldsymbol{\beta}_0\right)\right\}^2m_{k}(T;\boldsymbol{\beta}_0)^2\rho''\left(u_{K_1}^{\top}(T){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right)^2\|v_{K_2}(\boldsymbol{X})\|^2\|
u_{K_1}(T)\|^2\right] \\
&~ -\tr \bigg\{ \mathbb{E}[m_k(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}^*)\rho''\left(u_{K_1}^{\top}(T){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right)
u_{K_1}(T)v^{\top}_{K_2}(\boldsymbol{X})]\\
& \quad \quad \quad \quad \times \mathbb{E}[m_k(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(T){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right)v_{K_2}(\boldsymbol{X})u_{K_1}(T)^{\top}] \bigg\} \\
\leq &\mathbb{E}\left[L'\left\{Y-g\left(T;\boldsymbol{\beta}_0\right)\right\}^2m_{k}(T;\boldsymbol{\beta}_0)^2\rho''\left(u_{K_1}^{\top}(T){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right)^2\|v_{K_2}(\boldsymbol{X})\|^2\|u_{K_1}(T)\|^2\right] \\
\leq & \mathbb{E}\left[L'\left\{Y -g\left(T;\boldsymbol{\beta}_0\right)\right\}^2\right]\left( \sup_{t\in\mathcal{T}}m_k(t;\boldsymbol{\beta}_0)^2\right) \cdot a_3\cdot \mathbb{E}\left[\|v_{K_2}(\boldsymbol{X})\|^2\|u_{K_1}(T)\|^2\right]\\
= &\mathbb{E}\left[L'\left\{Y -g\left(T;\boldsymbol{\beta}_0\right)\right\}^2\right]\left( \sup_{t\in\mathcal{T}}m_k(t;\boldsymbol{\beta}_0)^2\right) \cdot a_3\cdot \mathbb{E}\left[\frac{1}{\pi_0(T,\boldsymbol{X})}\cdot \pi_0(T,\boldsymbol{X})\cdot\|v_{K_2}(\boldsymbol{X})\|^2\|u_{K_1}(T)\|^2\right]\\
\leq &\mathbb{E}\left[L'\left\{Y -g\left(T;\boldsymbol{\beta}_0\right)\right\}^2\right]\left( \sup_{t\in\mathcal{T}}m_k(t;\boldsymbol{\beta}_0)^2\right) \cdot a_3\cdot \frac{1}{\eta_1}\cdot \mathbb{E}\left[ \pi_0(T,\boldsymbol{X})\cdot\|v_{K_2}(\boldsymbol{X})\|^2\|u_{K_1}(T)\|^2\right]\\
= &\mathbb{E}\left[L'\left\{Y -g\left(T;\boldsymbol{\beta}_0\right)\right\}^2\right]\left( \sup_{t\in\mathcal{T}}m_k(t;\boldsymbol{\beta}_0)^2\right) \cdot a_3\cdot \frac{1}{\eta_1}\cdot \mathbb{E}\left[ \|v_{K_2}(\boldsymbol{X})\|^2\right] \mathbb{E}\left[\|u_{K_1}(T)\|^2\right] \quad (\text{by using}\ \eqref{moment1})\\
\leq& O(1) \cdot O(K_2)\cdot O(K_1)=O(K),
\end{align*}
where $\dps a_3 := \sup_{\gamma \in \Gamma_1} |\rho''(\gamma)|^2 < +\infty$, the second inequality follows from this definition and the fact that
$u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) \in \Gamma_1,\ \forall (t,\boldsymbol{x}) \in \mathcal{T}\times\mathcal{X}$ when $K$ is large enough; the third inequality follows from Assumption \ref{as:pi0}; the forth inequality follows from Assumption \ref{as:EY2} and the facts
\begin{align}\label{eq:E[u^2]}
&\mathbb{E}[\|u_{K_1}(T)\|^2]=\mathbb{E}[\tr(u_{K_1}(T)u_{K_1}^\top(T))]=\tr(I_{K_1\times K_1})=K_1, \\
&\mathbb{E}[\|v_{K_2}(\boldsymbol{X})\|^2]=\mathbb{E}[\tr(v_{K_2}(\boldsymbol{X})v_{K_2}^\top(\boldsymbol{X}))]=\tr(I_{K_2\times K_2})=K_2. \label{eq:E[v^2]}
\end{align}
Then in light of Chebyshev's inequality, Lemma \ref{lemma_pi^hat} and Assumption \ref{as:K&N_c}, we have $$|W_{1K,k}|\leq \|U_{K_2\times K_1}\|\|\hat{A}_{K_1\times K_2}\| =O_p(\sqrt{K})O_p\left(\sqrt{\frac{K}{N}}\right)=O_p\left(\sqrt{\frac{K^2}{N}}\right) ,$$ which implies
\begin{align} \notag
\|W_{1K}\|^2= \sum_{k=1}^{p}|W_{1K,k}|^2= O_p\left({\frac{K^2}{N}}\right) .
\end{align}
For the term $W_{3K}$, since $\xi_3(t,\boldsymbol{x})$ lies between $u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})$ and $u_{K_1}(t)^{\top}\tilde{\Lambda}_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})$, which implies $\xi_3(t,\boldsymbol{x})$ lies between $u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})$ and $u_{K_1}(t)^{\top}\hat{\Lambda}_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})$. Then in light of \eqref{Gamma_2_epsilon} and \eqref{eq:Lambdahat_in_Upsilonhat}, we have $\mathbb{P}\left(\xi_3(t,\boldsymbol{x})\in \Gamma_2(\epsilon),\ \forall (t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X}\right)>1-\epsilon$, therefore,
\begin{align}\label{xi_3}
\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times \mathcal{X}}|\rho'''\left(\xi_3(t,\boldsymbol{x})\right)|=O_p(1) \ .
\end{align}
With \eqref{eq:u_Lambda_v}, \eqref{xi_3}, the fact $\|\tilde{A}_{K_1\times K_2}\|\leq \|\hat{A}_{K_1\times K_2}\|$, Lemma \ref{lemma_pi^hat}, and Assumption \ref{as:K&N_c}, we can derive that{\footnotesize
\begin{align}
&\left\|W_{3K} \right\| =\bigg\|{\sqrt{N}}\int_{\mathcal{T}}\int_{\mathcal{X}} m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho'''\left(\xi_3(t,\boldsymbol{x})\right)\left\{u_{K_1}(t)^{\top}\tilde{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right\} u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t) \bigg\| \notag\\
\leq&\sqrt{N}\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times \mathcal{X}}\left|\rho'''\left(\xi_3(t,\boldsymbol{x})\right)\right| \sup_{t\in\mathcal{T}}\|m(t;\boldsymbol{\beta}_0)\|\cdot\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)| \notag\\
&\qquad \qquad \cdot \int_{\mathcal{T}}\int_{\mathcal{X}} \left|u_{K_1}(t)^{\top}\tilde{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right|\cdot \left| u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right|dF_{X,T}(\boldsymbol{x},t) \notag\\
\leq & \sqrt{N}\cdot O_p(1)\cdot O(1)\cdot O(1)\cdot \left\{\int_{\mathcal{T}}\int_{\mathcal{X}} \left|u_{K_1}(t)^{\top}\tilde{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right|^2dF_{X,T}(\boldsymbol{x},t)\right\}^{\frac{1}{2}} \notag\\
&\qquad \qquad \cdot \left\{\int_{\mathcal{T}}\int_{\mathcal{X}}\left| u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right|^2dF_{X,T}(\boldsymbol{x},t)\right\}^{\frac{1}{2}}\notag\\
=& \sqrt{N}\cdot O_p(1)\cdot O(1)\cdot O(1)\cdot O_p\left(\sqrt{\frac{K}{N}}\right)\cdot O_p\left(\sqrt{\frac{K}{N}}\right)=O_p\left(\sqrt{\frac{K^2}{N}}\right) \qquad \text{(by (\eqref{eq:u_Lambda_v}))} . \label{order:W_3K}
\end{align}}
For the term $W_{2K}$, we can deduce that
{\footnotesize
\begin{align*}
&\bigg\|\frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg[L'\left\{Y_i -g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m(T_i;\boldsymbol{\beta}_0)\rho'''\left(\xi_3(T_i,\boldsymbol{X}_i)\right) \left\{u_{K_1}(T_i)^{\top}\tilde{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right\} u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\bigg]\bigg\|\\
\leq & \left\{ \frac {1} {\sqrt{N}} \sum_{i=1}^N |L'(Y_i-g\left(T_i;\boldsymbol{\beta}_0\right))|\cdot \|u_{K_1}(T_i)\|^2 \|v_{K_2}(\boldsymbol{X}_i)\|^2\right\}\cdot \sup_{t\in\mathcal{T}}\|m(t;\boldsymbol{\beta}_0)\|\cdot \sup_{(t,\boldsymbol{x})\in\mathcal{T}\times \mathcal{X}}|\rho'''\left(\xi_3(t,\boldsymbol{x})\right)|\cdot \|\hat{A}_{K_1\times K_2}\|^2\\
\leq& \sqrt{N} \left\{ \frac {1} {N} \sum_{i=1}^N |L'(Y_i-g\left(T_i;\boldsymbol{\beta}_0\right))|^2\right\}^{\frac{1}{2}} \left\{\frac{1}{N}\sum_{i=1}^N \|u_{K_1}(T_i)\|^4 \|v_{K_2}(\boldsymbol{X}_i)\|^4\right\}^{\frac{1}{2}} \sup_{t\in\mathcal{T}}\|m(t;\boldsymbol{\beta}_0)\| \sup_{(t,\boldsymbol{x})\in\mathcal{T}\times \mathcal{X}}|\rho'''\left(\xi_3(t,\boldsymbol{x})\right)|\cdot \|\hat{A}_{K_1\times K_2}\|^2\\
\leq& \sqrt{N}\cdot O_p(1)\cdot \{\zeta_1(K_1)\zeta_2(K_2)\}\cdot \left\{\frac{1}{N}\sum_{i=1}^N \|u_{K_1}(T_i)\|^2 \|v_{K_2}(\boldsymbol{X}_i)\|^2\right\}^{\frac{1}{2}}\cdot O(1)\cdot O_p(1)\cdot O_p\left(\frac{K}{N}\right)\\
\leq &\sqrt{N}\cdot O_p(1)\cdot\zeta(K)\cdot \left\{\mathbb{E}\left[ \|u_{K_1}(T)\|^2\|v_{K_2}(\boldsymbol{X})\|^2\right]+O_p\left(\zeta(K)\sqrt{\frac{K}{N}}\right) \right\}^{\frac{1}{2}}\cdot O(1)\cdot O_p(1)\cdot O_p\left(\frac{K}{N}\right)\\
\leq &\sqrt{N}\cdot O_p(1)\cdot\zeta(K)\cdot O_p(\sqrt{K})\cdot O(1)\cdot O_p(1)\cdot O_p\left(\frac{K}{N}\right)=O_p\left(\zeta(K)\sqrt{\frac{K^3}{N}}\right)
\end{align*}}
where the fourth inequality follows from the fact that{\footnotesize
\begin{align*}
&\mathbb{E}\left[\left(\frac {1} {N} \sum_{i=1}^N \|u_{K_1}(T_i)\|^2 \|v_{K_2}(\boldsymbol{X}_i)\|^2-\mathbb{E}\left[ \|u_{K_1}(T)\|^2 \|v_{K_2}(\boldsymbol{X})\|^2\right]\right)^2\right]\\
= & \frac{1}{N}\cdot \mathbb{E}\left[\left(\|u_{K_1}(T)\|^2 \|v_{K_2}(\boldsymbol{X})\|^2-\mathbb{E}\left[\|u_{K_1}(T)\|^2 \|v_{K_2}(\boldsymbol{X})\|^2\right]\right)^2\right]\\
\leq &\frac{1}{N}\cdot \mathbb{E}\left[ \|u_{K_1}(T)\|^4 \|v_{K_2}(\boldsymbol{X})\|^4\right]\leq\frac{1}{N}\cdot \zeta_1(K_1)^2\zeta_2(K_2)^2\cdot \mathbb{E}\left[ \|u_{K_1}(T)\|^2 \|v_{K_2}(\boldsymbol{X})\|^2\right]\\
= &\frac{1}{N}\cdot \zeta_1(K_1)^2\zeta_2(K_2)^2\cdot \mathbb{E}\left[\frac{1}{\pi_0(T,\boldsymbol{X})}\cdot \pi_0(T,\boldsymbol{X}) \|u_{K_1}(T)\|^2 \|v_{K_2}(\boldsymbol{X})\|^2\right]\\
\leq &\frac{1}{N}\cdot \frac{1}{\eta_1}\cdot \zeta_1(K_1)^2\zeta_2(K_2)^2\cdot \mathbb{E}\left[ \pi_0(T,\boldsymbol{X}) \|u_{K_1}(T)\|^2 \|v_{K_2}(\boldsymbol{X})\|^2\right]\\
=&\frac{1}{N}\cdot \frac{1}{\eta_1}\cdot \zeta_1(K_1)^2\zeta_2(K_2)^2\cdot \mathbb{E}\left[ \|u_{K_1}(T)\|^2 \right]\cdot \mathbb{E}\left[ \|v_{K_2}(\boldsymbol{X})\|^2\right] =O\left(\frac{K}{N}\zeta(K)^2\right) .
\end{align*}}
Therefore, we can obtain that
\begin{align*}
\eqref{eq:WK}=W_{1K}+W_{2K}+W_{3K}=O_p\left(\sqrt{\frac{K^2}{N}}\right)+O_p\left(\zeta(K)\sqrt{\frac{K^3}{N}}\right)+O_p\left(\sqrt{\frac{K^2}{N}}\right)=O_p\left(\zeta(K)\sqrt{\frac{K^3}{N}}\right)\ .
\end{align*}
Finally, it follows that the term \eqref{eq:WK} is of $o_p(1)$ in light of Assumption \ref{as:K&N_c}.\\
\noindent \textbf{\emph{For term \eqref{eq:VK}}:}
Note that
\begin{align*}
&\mathbb{E}\left[\Bigg\|\frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg\{\left(\pi_K^*(T_i,\boldsymbol{X}_i) - {\pi_0(T_i,\boldsymbol{X}_i)}\right)m(T_i;\boldsymbol{\beta}_0)L'\left\{Y_i-g\left(T_i;\boldsymbol{\beta}_0\right)\right\}\right.\\
& \left. \quad \quad \quad \quad \quad \quad \quad \quad - \mathbb{E}\left[m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\left(\pi_K^*(T,\boldsymbol{X}) - {\pi_0(T,\boldsymbol{X})}\right)\right] \Bigg\}\Bigg\|^2\right]\\
\leq & \mathbb{E}\left[\left(\pi_K^*(T,\boldsymbol{X}) - {\pi_0(T,\boldsymbol{X})}\right)^2\|m(T;\boldsymbol{\beta}_0)\|^2\cdot L'\left\{Y-g\left(T;\boldsymbol{\beta}_0\right)\right\}^2\right]\\
\leq &\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}\left|\pi_K^*(t,\boldsymbol{x}) - {\pi_0(t,\boldsymbol{x})}\right|^2\cdot \sup_{t\in\mathcal{T}}\|m(t;\boldsymbol{\beta}_0)\|^2\cdot \mathbb{E}\left[L'\left\{Y-g\left(T;\boldsymbol{\beta}_0\right)\right\}^2\right]\\
\leq & O(\zeta(K)^2K^{-2\alpha}) \ ,
\end{align*}
where the last equality follows from Lemma \ref{lemma_pi^*}.
Then by Chebyshev's inequality, we can claim that the term \eqref{eq:VK} is of $O_p(\zeta(K)K^{-\alpha})$.\\
\noindent\textbf{\emph{For term \eqref{eq:Lemma1}}:} By Lemma \ref{lemma_pi^*} and Assumption \ref{as:K&N_c}, we can deduce that
\begin{align*}
&\left\|{\sqrt{N}}\cdot \mathbb{E}\left[m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\left(\pi_K^*(T,\boldsymbol{X}) - {\pi_0(T,\boldsymbol{X})}\right)\right]\right\|\\
\leq& \sqrt{N} \sup_{t\in\mathcal{T}}\|m(t;\boldsymbol{\beta}_0)\|\cdot \mathbb{E}[|\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)|^2]^{\frac{1}{2}}\cdot \mathbb{E}\left[|\pi_K^*(T,\boldsymbol{X}) - {\pi_0(T,\boldsymbol{X})}|^2\right]^{\frac{1}{2}}=O\left(\sqrt{N}K^{-\alpha}\right) \ .
\end{align*}
\noindent \textbf{\emph{For term \eqref{eq:tauto}}:} By Mean Value Theorem and the definition of $\hat{A}_{K_1\times K_2}$ in \eqref{def:A}, the term \eqref{eq:tauto} is exactly equal to zero.
\\
\noindent \textbf{\emph{For term \eqref{eq:Q}}:} We can telescope \eqref{eq:Q} as follows:
\begin{align}
&{\sqrt{N}}\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right) u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t) \notag\\
&-{\sqrt{N}}\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right) u_{K_1}^{\top}(t)A^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t) \notag\\
=&{\sqrt{N}}\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\left\{\rho''\left(u_{K_1}^{\top}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)-\rho''\left(u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right) \right\} \notag\\
&\quad \quad \quad \quad \quad \quad \quad \quad \quad \quad \qquad \qquad \times u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\label{A_star1}\\
&+{\sqrt{N}}\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right) u_{K_1}^{\top}(t)\notag\\
&\quad \quad \quad \quad \quad \quad \quad \quad \quad \quad \qquad \qquad \times\left\{\hat{A}_{K_1\times K_2}-A^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x}) dF_{X,T}(\boldsymbol{x},t) .\label{A_star2}
\end{align}
For the term \eqref{A_star1}, by Mean Value Theorem, {\footnotesize
\begin{align*}
\eqref{A_star1}=&{\sqrt{N}}\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho'''\left(\xi_3(t,\boldsymbol{x})\right)\left\{u_{K_1}^{\top}(t)\tilde{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right\} \left\{u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right\}dF_{X,T}(\boldsymbol{x},t)\\
=&-W_{3K} \ ,
\end{align*}}
which is $O_p\left(\sqrt{\frac{K^2}{N}}\right)$ from \eqref{order:W_3K}.\\
For the term \eqref{A_star2}, we first compute the probability order of $\|A^*_{K_1\times K_2}-\hat{A}_{K_1\times K_2}\|$. Using \eqref{eq:A^*_{K_1timesK_2}}, the fact $\rho''(v)=-\rho'(v)$ and Mean Value Theorem, we have{\footnotesize
\begin{align}
&A_{K_1\times K_2}^*-\hat{A}_{K_1\times K_2} \notag\\
=&-\frac{1}{N}\sum_{i=1}^N\rho''\left(u_{K_1}^{\top}(T_i){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)v^{\top}_{K_2}(\boldsymbol{X}_i)\notag \\
&-\frac{1}{N}\sum_{i=1}^N\rho'''\left(\xi_3(T_i,\boldsymbol{X}_i)\right)\left\{u_{K_1}(T)^{\top}\tilde{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right\}u_{K_1}(T_i)u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)v^{\top}_{K_2}(\boldsymbol{X}_i)\notag\\
&-\hat{A}_{K_1\times K_2} \notag \\
=&\frac{1}{N}\sum_{i=1}^N\left\{\rho'\left(u_{K_1}^{\top}(T_i){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)u_{K_1}(T_i)^{\top} \hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)v^{\top}_{K_2}(\boldsymbol{X}_i) -\hat{A}_{K_1\times K_2} \right\} \label{A_difference_1} \\
-&\frac{1}{N}\sum_{i=1}^N\rho'''\left(\xi_3(T_i,\boldsymbol{X}_i)\right)\left\{u_{K_1}(T_i)^{\top}\tilde{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right\}u_{K_1}(T_i)u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)v^{\top}_{K_2}(\boldsymbol{X}_i) .\label{A_difference_2}
\end{align}}
For the term \eqref{A_difference_1}, by \eqref{moment1} we can write $\hat{A}_{K_1\times K_2}$ as
$$ \hat{A}_{K_1\times K_2}=\mathbb{E}_{T,\boldsymbol{X}}\left[\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X})v^{\top}_{K_2}(\boldsymbol{X})\right], $$
where $\mathbb{E}_{T,\boldsymbol{X}}[\cdot]$ denotes taking expectation with respect to $(T,\boldsymbol{X})$.
We telescope \eqref{A_difference_1} as follows:
\begin{align}
&\frac{1}{N}\sum_{i=1}^N\bigg\{\rho'\left(u_{K_1}^{\top}(T_i){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)v^{\top}_{K_2}(\boldsymbol{X}_i)- \hat{A}_{K_1\times K_2}\bigg\} \notag\\
=&\frac{1}{N}\sum_{i=1}^N\bigg\{\left\{\rho'\left(u_{K_1}^{\top}(T_i){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)-\pi_0(T_i,\boldsymbol{X}_i)\right\}u_{K_1}(T_i)u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)v^{\top}_{K_2}(\boldsymbol{X}_i)\bigg\} \label{A_diff_1_1}\\
&-\frac{1}{N}\sum_{i=1}^N\bigg\{\pi_0(T_i,\boldsymbol{X}_i)u_{K_1}(T_i)u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)v^{\top}_{K_2}(\boldsymbol{X}_i) \notag \\
&\quad \quad \quad \quad \quad \quad-\mathbb{E}_{T,\boldsymbol{X}}\left[\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X})v^{\top}_{K_2}(\boldsymbol{X})\right] \bigg\}. \label{A_diff_1_2}
\end{align}
For the term \eqref{A_diff_1_1}, by Lemmas \ref{lemma_pi^*} and \ref{lemma_pi^hat}, we have that{\small
\begin{align*}
&\bigg\|\frac{1}{N}\sum_{i=1}^N \left\{\rho'\left(u_{K_1}^{\top}(T_i){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)-\pi_0(T_i,\boldsymbol{X}_i)\right\}u_{K_1}(T_i)u_{K_1}(T_i)^{\top} \hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)v^{\top}_{K_2}(\boldsymbol{X}_i)\bigg\|\\
\leq & \frac{1}{N}\sum_{i=1}^N \left|\rho'\left(u_{K_1}^{\top}(T_i){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)-\pi_0(T_i,\boldsymbol{X}_i)\right|\cdot \|u_{K_1}(T_i)\|^2 \cdot \|v_{K_2}(\boldsymbol{X}_i)\|^2\cdot \|\hat{A}_{K_1\times K_2}\|\\
=&\bigg\{\mathbb{E}\left[\left|\rho'\left(u_{K_1}^{\top}(T){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right|\cdot \|u_{K_1}(T)\|^2 \cdot \|v_{K_2}(\boldsymbol{X})\|^2\right]+O_p\left(\zeta(K)\sqrt{{\frac{K}{N}}}\right)\bigg\}\cdot \|\hat{A}_{K_1\times K_2}\|\\
\leq & \left\{\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times \mathcal{X}}\left|\pi_K^*(t,\boldsymbol{x})-\pi_0(t,\boldsymbol{x})\right|\cdot \frac{1}{\eta_1}\cdot \mathbb{E}\left[\pi_0(T,\boldsymbol{X})\|u_{K_1}(T)\|^2\|v_{K_2}(\boldsymbol{X})\|^2\right] +O_p\left(\zeta(K)\sqrt{{\frac{K}{N}}}\right) \right\}\cdot \|\hat{A}_{K_1\times K_2}\| \\
= & \left\{\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times \mathcal{X}}\left|\pi_K^*(t,\boldsymbol{x})-\pi_0(t,\boldsymbol{x})\right|\cdot \frac{1}{\eta_1}\cdot \mathbb{E}\left[\|u_{K_1}(T)\|^2\right]\cdot \mathbb{E}\left[\|v_{K_2}(\boldsymbol{X})\|^2\right] +O_p\left(\zeta(K)\sqrt{{\frac{K}{N}}}\right) \right\}\cdot \|\hat{A}_{K_1\times K_2}\| \\
\leq & \left\{O\left(K^{-\alpha}\zeta(K)\right) \cdot O(K_1)\cdot O(K_2)+O_p\left(\zeta(K)\sqrt{{\frac{K}{N}}}\right)\right\}\cdot O_p\left(\sqrt{\frac{K}{N}}\right)\\
\leq & O_p\left(N^{-\frac{1}{2}}\zeta(K)\cdot K^{\frac{3}{2}-\alpha}\right) .
\end{align*}}
For the term \eqref{A_diff_1_2}, define the linear map $\mathcal{J}(\cdot):\mathbb{R}^{{K_1\times K_2}}\rightarrow \mathbb{R}$ by {\footnotesize
\begin{align*}
\mathcal{J}(M):= \frac{1}{N}\sum_{i=1}^N\bigg\{\pi_0(T_i,\boldsymbol{X}_i)u_{K_1}(T_i)u_{K_1}(T_i)^{\top}Mv_{K_2}(\boldsymbol{X}_i)v^{\top}_{K_2}(\boldsymbol{X}_i) -\mathbb{E}_{T,\boldsymbol{X}}\left[\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}Mv_{K_2}(\boldsymbol{X})v^{\top}_{K_2}(\boldsymbol{X})\right] \bigg\},
\end{align*}}
then $ \eqref{A_diff_1_2}=\mathcal{J}(\hat{A}_{K_1\times K_2})$.
For any fixed $M \in \mathbb{R}^{{K_1\times K_2}}$, by \eqref{moment1} and
$M=\mathbb{E}[\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}M$ $\cdot v_{K_2}(\boldsymbol{X})v^{\top}_{K_2}(\boldsymbol{X})]$,
then we have
\begin{align*}
&\mathbb{E}\left[\mathcal{J}(M)^2\right]\\=&\frac{1}{N}\cdot\mathbb{E}\bigg[\bigg\|\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}Mv_{K_2}(\boldsymbol{X})v^{\top}_{K_2}(\boldsymbol{X})-\mathbb{E}\left[\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}Mv_{K_2}(\boldsymbol{X})v^{\top}_{K_2}(\boldsymbol{X})\right]\bigg\|^2\bigg]\\
\leq &\frac{1}{N}\cdot\mathbb{E}\bigg[\bigg\|\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}Mv_{K_2}(\boldsymbol{X})v^{\top}_{K_2}(\boldsymbol{X})\bigg\|^2\bigg]\\
\leq &\frac{1}{N} \cdot \eta_2\cdot \mathbb{E}\bigg[ \pi_0(T,\boldsymbol{X})\cdot \|u_{K_1}(T)\|^4\|v_{K_2}(\boldsymbol{X})\|^4\bigg]\cdot \left\|M\right\|^2\\
= &\frac{1}{N} \cdot \eta_2\cdot \mathbb{E}[ \|u_{K_1}(T)\|^4]\cdot \mathbb{E}[\|v_{K_2}(\boldsymbol{X})\|^4]\cdot \left\|M\right\|^2 \\
\leq & \frac{1}{N} \cdot \eta_2\cdot \zeta_1(K)^2\cdot \zeta_2(K)^2\cdot \mathbb{E}[ \|u_{K_1}(T)\|^2]\cdot \mathbb{E}[\|v_{K_2}(\boldsymbol{X})\|^2]\cdot \left\|M\right\|^2\\
=& \left\|M\right\|^2\cdot O\left(\zeta(K)^2\frac{K}{N}\right).
\end{align*}
Using Chebyshev's inequality we have
$$ |\mathcal{J}(M)| =\|M\| O_p\left(\zeta(K)\sqrt{\frac{K}{N}}\right)\ ,$$
then in light of Lemma \ref{lemma_pi^hat},
$$\eqref{A_diff_1_2}=\mathcal{J}(\hat{A}_{K_1\times K_2})=\|\hat{A}_{K_1\times K_2}\| O_p\left(\zeta(K)\sqrt{\frac{K}{N}}\right)= O_p\left(\zeta(K)\frac{K}{N}\right)\ .$$
Therefore,
$$\eqref{A_difference_1}=\eqref{A_diff_1_1}+\eqref{A_diff_1_2}= O_p\left(N^{-\frac{1}{2}}\zeta(K)\cdot K^{\frac{3}{2}-\alpha}\right)+O_p\left(\zeta(K)\frac{K}{N}\right) \ .$$
For the term \eqref{A_difference_2}, we can deduce that
\begin{align*}
&\bigg\|\frac{1}{N}\sum_{i=1}^N\rho'''\left(\xi_3(T_i,\boldsymbol{X}_i)\right)\left\{u_{K_1}(T_i)^{\top}\tilde{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right\} \left\{u_{K_1}(T_i)u_{K_1}(T_i)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)v^{\top}_{K_2}(\boldsymbol{X}_i)\right\}\bigg\| \\
\leq & \sup_{(t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X}}\left|\rho'''\left(\xi_3(t,\boldsymbol{x})\right)\right|\cdot \|\hat{A}_{K_1\times K_2}\|^2 \cdot \frac{1}{N}\sum_{i=1}^N\|u_{K_1}(T_i)\|^3\cdot \|v_{K_2}(\boldsymbol{X}_i)\|^3 \\
\leq & \sup_{(t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X}}\left|\rho'''\left(\xi_3(t,\boldsymbol{x})\right)\right|\cdot \|\hat{A}_{K_1\times K_2}\|^2 \cdot \zeta_1(K_1)\cdot \zeta_2(K_2)\cdot \frac{1}{N}\sum_{i=1}^N\|u_{K_1}(T_i)\|^2\cdot \|v_{K_2}(\boldsymbol{X}_i)\|^2 \\
\leq & \sup_{(t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X}}\left|\rho'''\left(\xi_3(t,\boldsymbol{x})\right)\right|\cdot \|\hat{A}_{K_1\times K_2}\|^2\cdot \zeta(K) \left\{\mathbb{E}\left[\|u_{K_1}(T)\|^2\|v_{K_2}(\boldsymbol{X})\|^2\right]+O_p\left(\zeta(K)\sqrt{\frac{K}{N}}\right)\right\}\\
\leq & O_p(1)\cdot O_p\left(\frac{K}{N}\right)\cdot \zeta(K) \cdot O(K)=O_p\left(\zeta(K)\frac{K^2}{N}\right) \ ,
\end{align*}
where the fourth inequality follows from \eqref{xi_3} and Lemma \ref{lemma_pi^hat}. Now, we can obtain
\begin{align}\label{A_K-A_K^*}
\|\hat{A}_{K_1\times K_2}-A_{K_1\times K_2}^*\|&=\eqref{A_difference_1}+\eqref{A_difference_2}=O_p\left(N^{-\frac{1}{2}}\zeta(K)K^{\frac{3}{2}-\alpha}\right)+O_p\left(\zeta(K)\frac{K}{N}\right)+O_p\left(\zeta(K)\frac{K^2}{N}\right) \notag\\
&=O_p\left(N^{-\frac{1}{2}}\zeta(K)\cdot K^{\frac{3}{2}-\alpha}\right)+O_p\left(\zeta(K)\frac{K^{2}}{N}\right) \ .
\end{align}
Using \eqref{A_K-A_K^*}, Assumptions \ref{as:EY2} and \ref{as:K&N_c}, for large enough $N$, we have
{\footnotesize\begin{align*}
&\eqref{A_star2}=\bigg\|{\sqrt{N}}\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t) \left\{\hat{A}_{K_1\times K_2}-A^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x}) dF_{X,T}(\boldsymbol{x},t) \bigg\|\\
\leq & \sqrt{N} \sup_{t\in\mathcal{T}}\| m(t;\boldsymbol{\beta}_0)\|\sup_{\gamma\in\Gamma_1}|\rho''\left(\gamma\right)|\cdot \mathbb{E}\left[\left|\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\right|^2\right]^{\frac{1}{2}}\cdot \left[\int_{\mathcal{T}\times\mathcal{X}}\left(u_{K_1}(t)\left\{\hat{A}_{K_1\times K_2}-A^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right)^2dF_{T,X}(t,\boldsymbol{x})\right]^{\frac{1}{2}} \\
\leq& \sqrt{N}\cdot O(1)\cdot O(1)\cdot O(1)\cdot O(1)\cdot O( \|\hat{A}_{K_1\times K_2}-A_{K_1\times K_2}^*\|)\\
\leq &O_p\left(\zeta(K)\cdot K^{\frac{3}{2}-\alpha}\right)+O_p\left(\zeta(K)\frac{K^{2}}{\sqrt{N}}\right),
\end{align*}}
where the second inequality holds since by using the same argument of establishing \eqref{eq:u_Lambda_v}, we have
$$\int_{\mathcal{T}\times\mathcal{X}}\left(u_{K_1}(t)\left\{\hat{A}_{K_1\times K_2}-A^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right)^2dF_{T,X}(t,\boldsymbol{x})= O( \|\hat{A}_{K_1\times K_2}-A_{K_1\times K_2}^*\|).$$
Therefore, we can obtain that
\begin{align*}
\eqref{eq:Q}=\eqref{A_star1}+\eqref{A_star2}=&O_p\left(\sqrt{\frac{K^2}{N}}\right)+O_p\left(\zeta(K)\cdot K^{\frac{3}{2}-\alpha}\right)+O_p\left(\zeta(K)\frac{K^{2}}{\sqrt{N}}\right)\\
=&O_p\left(\zeta(K)\cdot K^{\frac{3}{2}-\alpha}\right)+O_p\left(\zeta(K)\frac{K^{2}}{\sqrt{N}}\right).
\end{align*}
\noindent \textbf{\emph{For term \eqref{eq:proj}}:} By the definition of $A^*_{K_1\times K_2}$ in \eqref{def:A*}, we have{\footnotesize
\begin{align}
\eqref{eq:proj} &=\frac{1}{\sqrt{N}}\sum_{i=1}^N\bigg\{\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\label{projection_1}\\
& \qquad \times \left\{u_{K_1}(T_i)\rho'\left(u_{K_1}(T_i)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{X}_i)\right)v^{\top}_{K_2}(\boldsymbol{X}_i)\right\}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)+ m(T_i;\boldsymbol{\beta}_0)\varepsilon(T_i,\boldsymbol{X}_i;\boldsymbol{\beta}_0)\pi_0(T_i,\boldsymbol{X}_i)\bigg\} \notag \\
& -\frac{1}{\sqrt{N}}\sum_{i=1}^N\left\{\int_{\mathcal{T}}\int_{\mathcal{X}} m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}^*)\rho''\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v^{\top}_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\right.\label{projection_2}\left(\frac{1}{N}\sum_{l=1}^Nu_{K_1}(T_l)\right)\\
& \quad \quad \times \left(\frac{1}{N}\sum_{j=1}^Nv^{\top}_{K_2}(\boldsymbol{X}_j)\right)v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t) + \mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)|\boldsymbol{X}=\boldsymbol{X}_i\right]\notag\\
&\qquad \qquad + \mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)|T=T_i\right]\bigg\}. \notag
\end{align}}
We shall show that both \eqref{projection_1} and \eqref{projection_2} are of $o_p(1)$. Noting $\rho''=-\rho'$, we can telescope \eqref{projection_1} as follows:
{\footnotesize
\begin{align}
\eqref{projection_1}=&\frac{1}{\sqrt{N}}\sum_{i=1}^N\bigg\{\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\label{Projection_1_1}\\
& \qquad \qquad \qquad \times \left\{u_{K_1}(T_i)\bigg[-\rho'\left(u_{K_1}(T_i)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{X}_i)\right)+\pi_0(T_i,\boldsymbol{X}_i)\bigg]v^{\top}_{K_2}(\boldsymbol{X}_i)\right\}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\bigg\}\notag\\
&-\frac{1}{\sqrt{N}}\sum_{i=1}^N\bigg\{\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\left\{\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)-\pi_0(t,\boldsymbol{x})\right\}u_{K_1}^{\top}(t) \label{Projection_1_2}\\
& \qquad \qquad \qquad \times \left\{u_{K_1}(T_i)\pi_0(T_i,\boldsymbol{X}_i)v^{\top}_{K_2}(\boldsymbol{X}_i)\right\}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\bigg\}\notag\\
&- \frac{1}{\sqrt{N}}\sum_{i=1}^N\bigg\{\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\pi_0(t,\boldsymbol{x})u_{K_1}^{\top}(t) \left\{u_{K_1}(T_i)\pi_0(T_i,\boldsymbol{X}_i)v^{\top}_{K_2}(\boldsymbol{X}_i)\right\}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\notag\\
&\qquad \qquad \qquad +m(T_i;\boldsymbol{\beta}_0)\varepsilon(T_i,\boldsymbol{X}_i;\boldsymbol{\beta}_0)\pi_0(T_i,\boldsymbol{X}_i)\bigg\}\label{Projection_1_3}\ .
\end{align}
}
We shall show that \eqref{Projection_1_1}, \eqref{Projection_1_2} and \eqref{Projection_1_3} are all of $o_p(1)$. Note that second moment of \eqref{Projection_1_1} is
{\footnotesize
\begin{align*}
&\mathbb{E}[|\eqref{Projection_1_1}|^2]=\mathbb{E}\Bigg[\Bigg|\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\\
& \qquad \qquad \times \left\{u_{K_1}(T_i)\bigg[-\rho'\left(u_{K_1}(T_i)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{X}_i)\right)+\pi_0(T_i,\boldsymbol{X}_i)\bigg]v^{\top}_{K_2}(\boldsymbol{X}_i)\right\}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\Bigg|^2\Bigg]\\
= &\mathbb{E}\Bigg[\Bigg|\int_{\mathcal{T}}\int_{\mathcal{X}}\pi_0(t,\boldsymbol{x})\cdot m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\bigg[\frac{\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)}{\pi_0(t,\boldsymbol{x})}\bigg]u_{K_1}^{\top}(t)\\
& \qquad \qquad \times \left\{u_{K_1}(T_i)\bigg[-\rho'\left(u_{K_1}(T_i)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{X}_i)\right)+\pi_0(T_i,\boldsymbol{X}_i)\bigg]v^{\top}_{K_2}(\boldsymbol{X}_i)\right\}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\Bigg|^2\Bigg]\\
\leq&\mathbb{E}\Bigg[\Bigg|\int_{\mathcal{T}}\int_{\mathcal{X}}\pi_0(t,\boldsymbol{x})\cdot m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\bigg[\frac{\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)}{\pi_0(t,\boldsymbol{x})}\bigg]u_{K_1}^{\top}(t) \left\{u_{K_1}(T_i)v^{\top}_{K_2}(\boldsymbol{X}_i)\right\}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\Bigg|^2\Bigg]\\
&\qquad \times\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}} \left\{-\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)+\pi_0(t,\boldsymbol{x})\right\}^2 \\
=& \left\{ \mathbb{E}\Bigg[\Bigg| m(T_i;\boldsymbol{\beta}_0)\varepsilon(T_i,\boldsymbol{X}_i;\boldsymbol{\beta}_0)\bigg[\frac{\rho'\left(u_{K_1}^{\top}(T_i)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{X}_i)\right)}{\pi_0(T_i,\boldsymbol{X}_i)}\bigg]\Bigg|^2\Bigg]+o(1)\right\}\times\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}} \left\{-\pi_K^*(t,\boldsymbol{x})+\pi_0(t,\boldsymbol{x})\right\}^2 \\
=&O(1)\cdot O(K^{-\frac{2s}{r}}\zeta(K)^2)=O(K^{-\frac{2s}{r}}\zeta(K)^2)\ ,
\end{align*}}
where the third equality holds because
$$\int_{\mathcal{T}}\int_{\mathcal{X}}\pi_0(t,\boldsymbol{x})\cdot m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\bigg[\frac{\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)}{\pi_0(t,\boldsymbol{x})}\bigg]u_{K_1}^{\top}(t) \left\{u_{K_1}(T)v^{\top}_{K_2}(\boldsymbol{X})\right\}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t) $$
is the weighted $L^2$-projection of $m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\bigg[\frac{\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)}{\pi_0(t,\boldsymbol{x})}\bigg]$ on the space linearly spanned by $\{u_{K_1}(t),v_{K_2}(\boldsymbol{x})\}$ with the weighted measure $\pi_0(t,\boldsymbol{x})dF_{T,X}(t,\boldsymbol{x})$. Similarly, we can also show \eqref{Projection_1_2} and \eqref{Projection_1_3} are of $o_p(1)$. Therefore, \eqref{projection_1} is of $o_p(1)$.\\
For the term \eqref{projection_2}, since $\rho''(v)=-\rho'(v)$ and the fact $\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\right]=0$, we telescope it as follows:{\footnotesize
\begin{align}
\eqref{projection_2} =& {\sqrt{N}}\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t) \left( \frac{1}{N}\sum_{l=1}^N u_{K_1}(T_l)-\mathbb{E}\left[u_{K_1}(T)\right]\right) \notag\\
& \quad \quad \quad \times \left(\frac{1}{N}\sum_{j=1}^Nv^{\top}_{K_2}(\boldsymbol{X}_j)-\mathbb{E}[v^{\top}_{K_2}(\boldsymbol{X})]\right)v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t) \label{projection_2_1} \\
+&\frac{1}{\sqrt{N}}\sum_{i=1}^N\left\{\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\mathbb{E}\left[ u_{K_1}(T)\right]v^{\top}_{K_2}(\boldsymbol{X}_i)v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\right. \notag\\
&\qquad \qquad \qquad -\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)|\boldsymbol{X}=\boldsymbol{X}_i\right]\bigg\} \label{projection_2_4} \\
+&\frac{1}{\sqrt{N}}\sum_{l=1}^N \bigg\{ \int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right) u_{K_1}^{\top}(t)u_{K_1}(T_l) \mathbb{E}[v^{\top}_{K_2}(\boldsymbol{X})]v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t) \notag\\
&\qquad \qquad\qquad -\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)|T=T_i\right] \bigg\} \label{projection_2_3} \\
&-\frac{1}{\sqrt{N}}\sum_{i=1}^N\bigg\{\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\mathbb{E}\left[u_{K_1}(T)\right] \mathbb{E}[v^{\top}_{K_2}(\boldsymbol{X})]v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t)\notag\\
&\qquad \qquad \qquad -\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\right] \bigg\} . \label{projection_2_5}
\end{align}}
For the term \eqref{projection_2_1}, since
\begin{align*}
&\left\|\frac{1}{N}\sum_{l=1}^N u_{K_1}(T_l)-\mathbb{E}\left[u_{K_1}(T)\right]\right\|=O_p\left(\sqrt{\frac{K_1}{N}}\right) \ ,
\\
&\left\|\frac{1}{N}\sum_{j=1}^N v_{K_2}(\boldsymbol{X}_j)-\mathbb{E}\left[v_{K_2}(\boldsymbol{X})\right]\right\|=O_p\left(\sqrt{\frac{K_2}{N}}\right) \ ,\\
& \sup_{(t,\boldsymbol{x})\in\mathcal{T}\times \mathcal{X}}\left|\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)\right|=O(1) \ ,
\end{align*}
and by Assumptions \ref{as:m_smooth}, \ref{as:EY2}, and \ref{as:K&N_c}, we can deduce that
\begin{align*}
\eqref{projection_2_1}=\sqrt{N}\cdot O(\zeta(K)) O_p\left(\sqrt{\frac{K_1}{N}}\right)O_p\left(\sqrt{\frac{K_2}{N}}\right)=O_p\left(\zeta(K)\sqrt{\frac{K}{N}}\right) =o_p(1) \ .
\end{align*}
For the term \eqref{projection_2_4}, noting the fact that $\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)|\boldsymbol{X}\right]=\int_{\mathcal{T}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{X};\boldsymbol{\beta}_0)$ $dF_T(t)$, we can rewrite \eqref{projection_2_4} as follows:
\begin{align}
\eqref{projection_2_4}=&\frac{1}{\sqrt{N}}\sum_{j=1}^N\left\{\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\frac{\pi_K^*(t,\boldsymbol{x})}{\pi_0(t,\boldsymbol{x})}u_{K_1}^{\top}(t)\mathbb{E}\left[ u_{K_1}(T)\right]\right. v^{\top}_{K_2}(\boldsymbol{X}_j)v_{K_2}(\boldsymbol{x})dF_{X}(\boldsymbol{x})dF_T(t)\notag \\
&\qquad \qquad\qquad - \int_{\mathcal{T}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{X}_j;\boldsymbol{\beta}_0) dF_T(t)\Bigg\}. \notag
\end{align}
By computing the second moment of \eqref{projection_2_4}, we can obtain that{\footnotesize
\begin{align*}
&\mathbb{E}\bigg[\bigg\|\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\frac{\pi_K^*(t,\boldsymbol{x})}{\pi_0(t,\boldsymbol{x})}u_{K_1}^{\top}(t)\mathbb{E}\left[ u_{K_1}(T)\right] v^{\top}_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{x})dF_{X}(\boldsymbol{x})dF_T(t) - \int_{\mathcal{T}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{X};\boldsymbol{\beta}_0) dF_T(t)\bigg\|^2 \bigg]\\
\leq &\mathbb{E}\bigg[\bigg\|\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\frac{\pi_K^*(t,\boldsymbol{x})}{\pi_0(t,\boldsymbol{x})}u_{K_1}^{\top}(t) u_{K_1}(T^*) v^{\top}_{K_2}(\boldsymbol{X}^*)v_{K_2}(\boldsymbol{x})dF_{X}(\boldsymbol{x})dF_T(t) - m(T^*;\boldsymbol{\beta}_0)\varepsilon(T^*,\boldsymbol{X}^*;\boldsymbol{\beta}_0) \bigg\|^2 \bigg]\\
\leq & 2\cdot \mathbb{E}\bigg[\bigg\| \int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)u_{K_1}^{\top}(t)u_{K_1}(T^*) v^{\top}_{K_2}(\boldsymbol{X}^*)v_{K_2}(\boldsymbol{x})dF_{X}(\boldsymbol{x})dF_T(t) -m(T^*;\boldsymbol{\beta}_0)\varepsilon(T^*,\boldsymbol{X}^*;\boldsymbol{\beta}_0) \bigg\|^2\bigg] \\
&+2\cdot \mathbb{E}\bigg[\bigg\| \int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\frac{\pi^*_K(t,\boldsymbol{x})-\pi_0(t,\boldsymbol{x})}{\pi_0(t,\boldsymbol{x})} u_{K_1}^{\top}(t)u_{K_1}(T^*) v^{\top}_{K_2}(\boldsymbol{X}^*)v_{K_2}(\boldsymbol{x})dF_{X}(\boldsymbol{x})dF_T(t)\bigg\|^2\bigg]\\
=& o(1),
\end{align*}}
where $T^*\sim F_T$, $\boldsymbol{X}^*\sim F_X$, and $T^*$ is independent of $\boldsymbol{X}^*$; the first inequality holds by Jensen's inequality; the last equality follows from Lemma \ref{lemma_pi^*} and the fact that
\begin{align*}
\int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)u_{K_1}^{\top}(t)u_{K_1}(T^*) v^{\top}_{K_2}(\boldsymbol{X}^*)v_{K_2}(\boldsymbol{x})dF_{X}(\boldsymbol{x})dF_T(t) \end{align*}
is the $L^2$-projection of $m(T^*;\boldsymbol{\beta}_0)\varepsilon(T^*,\boldsymbol{X}^*;\boldsymbol{\beta}_0)$ on the space spanned by $\{u_{K_1}(T^*),v_{K_2}(\boldsymbol{X}^*)\}$, which implies {\footnotesize
\begin{align*}
&\mathbb{E}\bigg[\bigg\| \int_{\mathcal{T}}\int_{\mathcal{X}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)u_{K_1}^{\top}(t)u_{K_1}(T^*) v^{\top}_{K_2}(\boldsymbol{X}^*)v_{K_2}(\boldsymbol{x})dF_{X}(\boldsymbol{x})dF_T(t) -m(T^*;\boldsymbol{\beta}_0)\varepsilon(T^*,\boldsymbol{X}^*;\boldsymbol{\beta}_0) \bigg\|^2\bigg]\rightarrow 0.
\end{align*}}
Thus \eqref{projection_2_4} is of $o_p(1)$ by Chebyshev's inequality. Similar argument can be applied to show that both \eqref{projection_2_3} and \eqref{projection_2_5} are of $o_p(1)$. Therefore, we can have that
\begin{align*}
|\eqref{projection_2}|\leq |\eqref{projection_2_1}|+|\eqref{projection_2_4}|+|\eqref{projection_2_3}|=o_p(1) \ .
\end{align*}
Then, we can obtain that
\begin{align*}
|\eqref{eq:proj}|\leq |\eqref{projection_1}|+|\eqref{projection_2}|=o_p(1)\ .
\end{align*}
Summing up all orders \eqref{eq:WK}-\eqref{eq:proj} and using Assumption \ref{as:K&N_c}, we have{\small
\begin{align*}
&\eqref{eq:WK}+\eqref{eq:VK}+\eqref{eq:Lemma1}+\eqref{eq:tauto}+\eqref{eq:Q}+\eqref{eq:proj}\\
=&O_p\left(\zeta(K)\sqrt{\frac{K^3}{N}}\right)+ O_p\left(\zeta(K)K^{-\alpha}\right)+ O(\sqrt{N}K^{-\alpha})+0+ \left\{O_p\left(\zeta(K)\cdot K^{\frac{3}{2}-\alpha}\right)+O_p\left(\zeta(K)\frac{K^{2}}{\sqrt{N}}\right)\right\}+o_p(1)\\
=&o_p(1).
\end{align*}}
\section{Variance Estimation}
\cite{masry1996multivariate} studies the strong consistency of kernel regression estimation. The following conditions are imposed so that Theorem 6 of \cite{masry1996multivariate} applies. Let
$\boldsymbol{u}=(u_1,...,u_{r+1})$ and $K(\boldsymbol{u})=\prod_{j=1}^{r+1}k(u_j)$.\\
\noindent\textbf{Condition 1}. The kernel $K(\boldsymbol{u})\in L_1$ satisfies $\|\boldsymbol{u}\|K(\boldsymbol{u})\in L_{1}$ and $\|\boldsymbol{u}\|^2K(\boldsymbol{u})\in L_{1}$.
\noindent\textbf{Condition 2}. The density function $f_{Y,X,T}(y,\boldsymbol{x},t)$ is uniformly bounded away from zero and above, and also it is uniformly continuous on $\mathbb{R}^{r+1}$.
\noindent\textbf{Condition 3}. (a) The kernel $K(\cdot)$ is bounded with compact support; (b) let $H_j(\boldsymbol{u}):=\boldsymbol{u}^jK(\boldsymbol{u})$, and $|H_j(\boldsymbol{u})-H_j(\boldsymbol{v})|\leq C\|\boldsymbol{u}-\boldsymbol{v}\|$ for all $j$ with $0\leq j\leq 3$.
\noindent\textbf{Condition 4}. The functions $f_{Y,X,T}(y,\boldsymbol{x},t)$, $\mathbb{E}\left[\pi_0(T,\boldsymbol{X})L'(Y-g(T;\boldsymbol{\beta}))|T=t,\boldsymbol{X}=\boldsymbol{x}\right]$, $\mathbb{E}[\pi_0(T,\boldsymbol{X})L'(Y-g(T;\boldsymbol{\beta}))|T=t]$ and $\mathbb{E}\left[\pi_0(T,\boldsymbol{X})L'(Y-g(T;\boldsymbol{\beta}))|\boldsymbol{X}=\boldsymbol{x}\right]$ are twice differentiable, and the derivatives are Lipschitz continuous and uniformly bounded.
\noindent\textbf{Condition 5}. $\mathbb{E}\left[\sup_{\boldsymbol{ \beta}\in \Theta}|L'(Y-g(T;\boldsymbol{\beta}))|^{\sigma}\right]<\infty$ for some $\sigma>2$.
\noindent\textbf{Condition 6}.The bandwidths $h_1\asymp \cdots \asymp h_r \asymp h_Y \asymp h_T\asymp h_N$ go to zero slowly enough such that
\begin{align*}
\frac{N^{1-2/\sigma}h_N^{r+1}}{(\log N)\{(\log N)(\log \log N)^{1+\sigma}\}^{2/\sigma}}\to \infty \ \text{as} \ N\to \infty.
\end{align*}
\section{Some Extensions}
\subsection{Proof of Theorem 7.1}
\textbf{(Proof of Consistency)}.
Let
\begin{align*}
\hat{\gamma}=\left[\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\sum_{i=1}^Nu_{K_1}(T_i)\hat{\pi}_K(T_i,\boldsymbol{X}_i)Y_i\right]
\end{align*}
then $\hat{\theta}_K(t)=\hat{\gamma}^{\top}u_{K_1}(t)$. By assumption, there exists $\gamma^*\in\mathbb{R}^{K_1}$ such that
\begin{align}\label{eq:approximate_error}
\sup_{t\in\mathcal{T}}\left|\mathbb{E}[\pi_0(T,X)Y|T=t]-(\gamma^*)^{\top}u_{K_1}(t)\right|=O(K_1^{-\tilde{\alpha}}).
\end{align}
We first claim that
\begin{align}\label{claim:gamma_hat-star}
\|\hat{\gamma}-\gamma^*\|=O_p\left(\zeta(K)\sqrt{\frac{K}{N}}+\zeta(K)K^{-\alpha}+K_1^{-\tilde{\alpha}}\right)\ ,
\end{align}
and the proof will be established later. With the claim \eqref{claim:gamma_hat-star}, we first show that $\int_{\mathcal{T}}|\hat{\theta}_K(t)-\theta(t)|^2dF_T(t)=O_p\left(\frac{\zeta(K)^2K}{N}+\zeta(K)^2K^{-2\alpha}+K_1^{-2\tilde{\alpha}}\right)$. Note that
\begin{align*}
&\int_{\mathcal{T}}[\hat{\theta}_K(t)-\theta(t)]^2dF_T(t)\\
=&\int_{\mathcal{T}}[\hat{\gamma}^{\top}u_{K_1}(t)-(\gamma^*)^{\top}u_{K_1}(t)+(\gamma^*)^{\top}u_{K_1}(t)-\theta(t)]^2dF_T(t)\\
\leq & 2(\hat{\gamma}-\gamma^*)^{\top}\left[\int_{\mathcal{T}}u_{K_1}(t)u_{K_1}(t)^{\top}dF_T(t)\right](\hat{\gamma}-\gamma^*)+2\int_{\mathcal{T}}[(\gamma^*)^{\top}u_{K_1}(t)-\theta(t)]^2dF_T(t)\\
\leq &2 \|\hat{\gamma}-\gamma^*\|^2\cdot \lambda_{\max}\left(\mathbb{E}[u_{K_1}(T)u_{K_1}(T)^{\top}]\right)+2\sup_{t\in\mathcal{T}}|(\gamma^*)^{\top}u_{K_1}(t)-\theta(t)|^2\\
=&O_p\left(\zeta(K)^2\frac{K}{N}+\zeta(K)^2K^{-2\alpha}+K_1^{-2\tilde{\alpha}}\right).
\end{align*}
With the claim \eqref{claim:gamma_hat-star}, we next show that $\sup_{t\in\mathcal{T}}|\hat{\theta}_K(t)-\theta(t)|=O_p[\zeta_1(K_1)(\zeta(K)\sqrt{K/N}+\zeta(K)K^{-\alpha}+K_1^{-\alpha})]$. Note that
\begin{align*}
&\sup_{t\in\mathcal{T}}|\hat{\theta}_K(t)-\theta(t)|\\
=&\sup_{t\in\mathcal{T}}\left|\hat{\gamma}^{\top}u_{K_1}(t)-(\gamma^*)^{\top}u_{K_1}(t)+(\gamma^*)^{\top}u_{K_1}(t)-\theta(t)\right|\\
\leq & \sup_{t\in\mathcal{T}}\|u_{K_1}(t)\|\cdot \left\|\hat{\gamma}-\gamma^*\right\|+\sup_{t\in\mathcal{T}}|(\gamma^*)^{\top}u_{K_1}(t)-\theta(t)|\\
\leq &\zeta_1(K_1)\cdot \left\{O_p\left(\sup_{(t,x)\in\mathcal{T}\times\mathcal{X}}|\hat{\pi}_K(t,x)-\pi_0(t,x)|\right)+O(K_1^{-\tilde{\alpha}})\right\}+O(K_1^{-\tilde{\alpha}}) \\
\leq &O_p\left[\zeta_1(K_1)\left(\zeta(K)\sqrt{\frac{K}{N}}+\zeta(K)K^{-\alpha}+K_1^{-\alpha}\right)\right]+O(K_1^{-\tilde{\alpha}})\\
=&O_p\left[\zeta_1(K_1)\left(\zeta(K)\sqrt{\frac{K}{N}}+\zeta(K)K^{-\alpha}+K_1^{-\tilde{\alpha}}\right)\right].
\end{align*}
Finally, we turn back to prove the claim \eqref{claim:gamma_hat-star}. Note that
\begin{align*}
\hat{\gamma}-\gamma^* =&\left[\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\sum_{i=1}^N\hat{\pi}_K(T_i,\boldsymbol{X}_i)u_{K_1}(T_i)Y_i\right]-\gamma^*\\
=&\left[\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\sum_{i=1}^Nu_{K_1}(T_i)\left\{\hat{\pi}_K(T_i,\boldsymbol{X}_i)-{\pi}_0(T_i,\boldsymbol{X}_i)\right\}Y_i\right]\\
&+\left[\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\sum_{i=1}^Nu_{K_1}(T_i)\{\pi_0(T_i,\boldsymbol{X}_i)Y_i-\mathbb{E}[\pi_0(T_i,\boldsymbol{X}_i)Y_i|T_i]\}\right]\\
&+\left[\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\sum_{i=1}^Nu_{K_1}(T_i)\left\{\mathbb{E}[{\pi}_0(T_i,\boldsymbol{X}_i)Y_i|T_i]-(\gamma^*)^{\top}u_{K_1}(T_i)\right\}\right]\\
=&A_{1N}+A_{2N}+A_{3N}
\end{align*}
where
\begin{align*}
&A_{1N}=\left[\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\sum_{i=1}^Nu_{K_1}(T_i)\left\{\hat{\pi}_K(T_i,\boldsymbol{X}_i)-{\pi}_0(T_i,\boldsymbol{X}_i)\right\}Y_i\right] , \\
&A_{2N}=\left[\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\sum_{i=1}^Nu_{K_1}(T_i)\{\pi_0(T_i,\boldsymbol{X}_i)Y_i-\mathbb{E}[\pi_0(T_i,\boldsymbol{X}_i)Y_i|T_i]\}\right] ,\\
&A_{3N}=\left[\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\sum_{i=1}^Nu_{K_1}(T_i)\left\{\mathbb{E}[{\pi}_0(T_i,\boldsymbol{X}_i)Y_i|T_i]-(\gamma^*)^{\top}u_{K_1}(T_i)\right\}\right] .
\end{align*}
We first compute the probability order of $A_{1N}$. We use the following notation:
\begin{align*}
&\hat{H}_N:=\left(\left\{\hat{\pi}_K(T_1,X_1)-{\pi}_0(T_1,X_1)\right\}Y_1,...,\left\{\hat{\pi}_K(T_N,X_N)-{\pi}_0(T_N,X_N)\right\}Y_N\right)^\top \ ,\\
&U_{N\times K_1}:=(u_{K_1}(T_1),...,u_{K_1}(T_N))^\top\ , \\
&\hat{\Phi}_{K_1\times K_1}:=\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T)u_{K_1}^{\top}(T).
\end{align*}
Then we can obtain that
\begin{align*}
\|A_{1N}\|^2=&\left\|\left[\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\sum_{i=1}^Nu_{K_1}(T_i)\left\{\hat{\pi}_K(T_i,\boldsymbol{X}_i)-{\pi}_0(T_i,\boldsymbol{X}_i)\right\}Y_i\right]\right\|^2\\
=& N^{-2}\text{tr}\left(\hat{\Phi}_{K_1\times K_1}^{-1}U_{N\times K_1}^{\top}\hat{H}_N\hat{H}_N^{\top}U_{N\times K_1}\hat{\Phi}_{K_1\times K_1}^{-1}\right)\\
=& N^{-2}\text{tr}\left(U_{N\times K_1}^{\top}\hat{H}_N\hat{H}_N^{\top}U_{N\times K_1}\hat{\Phi}_{K_1\times K_1}^{-1}\hat{\Phi}_{K_1\times K_1}^{-1}\right)\\
=& N^{-2}\text{tr}\left(\hat{\Phi}_{K_1\times K_1}^{-1/2}U_{N\times K_1}^{\top}\hat{H}_N\hat{H}_N^{\top}U_{N\times K_1}\hat{\Phi}_{K_1\times K_1}^{-1/2}\hat{\Phi}_{K_1\times K_1}^{-1}\right)\\
\leq &\lambda_{\max}(\hat{\Phi}_{K_1\times K_1}^{-1})N^{-2}\text{tr}\left(\hat{\Phi}_{K_1\times K_1}^{-1/2}U_{N\times K_1}^{\top}\hat{H}_N\hat{H}_N^{\top}U_{N\times K_1}\hat{\Phi}_{K_1\times K_1}^{-1/2}\right)\\
=&\lambda_{\max}(\hat{\Phi}_{K_1\times K_1}^{-1})N^{-1}\text{tr}\left(\hat{H}_N\hat{H}_N^{\top}U_{N\times K_1}(U^{\top}_{N\times K_1}U_{N\times K_1})^{-1}U_{N\times K_1}^{\top}\right)\\
\leq &[\lambda_{\min}(\hat{\Phi}_{K_1\times K_1})]^{-1}N^{-1}\|\hat{H}_N\|^2\\
=&[\lambda_{\min}(\hat{\Phi}_{K_1\times K_1})]^{-1}\cdot \frac{1}{N}\sum_{i=1}^N\left\{\hat{\pi}_K(T_i,\boldsymbol{X}_i)-{\pi}_0(T_i,\boldsymbol{X}_i)\right\}^2Y_i^2\\
\leq &[\lambda_{\min}(\hat{\Phi}_{K_1\times K_1})]^{-1}\sup_{(t,x)\in\mathcal{T}\times \mathcal{X}}|\hat{\pi}_K(t,x)-{\pi}_0(t,x)|^2\cdot \frac{1}{N}\sum_{i=1}^NY_i^2\\
\leq & O_p(1)\cdot O_p\left(\zeta(K)^2K^{-2\alpha}+\frac{\zeta(K)^2K}{N}\right) \cdot O_p(1)\\
=&O_p\left(\zeta(K)^2K^{-2\alpha}+\frac{\zeta(K)^2K}{N}\right),
\end{align*}
where the first inequality follows from the fact that $tr(AB)\leq \lambda_{\max}(B)tr(A)$ for any symmetric matrix $B$ and positive semidefinite matrix $A$,
the second inequality follows from the same fact and the fact that $U_{N\times K_1}(U^{\top}_{N\times K_1}U_{N\times K_1})^{-1}U_{N\times K_1}^{\top}$ is a projection matrix with maximum eigenvalue 1,
and the fourth inequality follows from the facts that $|\lambda_{\min}(\hat{\Phi}_{K_1\times K_1})|^{-1}=O_p(1)$,
$\sup_{(t,x)\in\mathcal{T}\times \mathcal{X}}|\hat{\pi}_K(t,x)-{\pi}_0(t,x)|=O_p\left(\zeta(K)K^{-\alpha}+\zeta(K)\sqrt{K/N}\right)$ and $N^{-1}\sum_{i=1}^NY_i^2=O_p(1)$. \\
Next, we compute the probability order of $A_{2N}$. We can deduce that
\begin{align*}
\|A_{2N}\|^2=&\left\|\left[\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\sum_{i=1}^Nu_{K_1}(T_i)\varepsilon_i\right]\right\|^2\\
=&N^{-2}\text{tr}\left(\hat{\Phi}_{K_1\times K_1}^{-1}U_{N\times K_1}^{\top}\mathcal{E}_N\mathcal{E}_N^{\top}U_{N\times K_1}\hat{\Phi}_{K_1\times K_1}^{-1}\right)\\
=&N^{-2}\text{tr}\left(U_{N\times K_1}^{\top}\mathcal{E}_N\mathcal{E}_N^{\top}U_{N\times K_1}\hat{\Phi}_{K_1\times K_1}^{-1}\hat{\Phi}_{K_1\times K_1}^{-1}\right)\\
\leq &[\lambda_{\min}(\hat{\Phi}_{K_1\times K_1})]^{-2}N^{-2}\|U_{N\times K_1}^{\top}\mathcal{E}_N\|^2=O_p\left(\frac{K_1}{N}\right)\ ,
\end{align*}
where the last equality follows that $|\lambda_{\min}(\hat{\Phi}_{K_1\times K_1})|^{-1}=O_p(1)$ and $N^{-2}\|U_{N\times K_1}^{\top}\mathcal{E}_N\|^2=O_p(K_1/N)$ by Markov's inequality. \\
We finally compute the probability order of $A_{3N}$. We define the notation
\begin{align*}
R_N(\gamma^*)=\left(\left\{\mathbb{E}[{\pi}_0(T_1,X_1)Y_1|T_1]-(\gamma^*)^{\top}u_{K_1}(T_1)\right\},...,\left\{\mathbb{E}[{\pi}_0(T_N,X_N)Y_N|T_N]-(\gamma^*)^{\top}u_{K_1}(T_N)\right\}\right)^{\top}\ ,
\end{align*}
then it follows that with probability approaching to 1,
\begin{align*}
\|A_{3N}\|^2=&\left\|\left[\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\sum_{i=1}^Nu_{K_1}(T_i)\left\{\mathbb{E}[{\pi}_0(T_i,\boldsymbol{X}_i)Y_i|T_i]-(\gamma^*)^{\top}u_{K_1}(T_i)\right\}\right]\right\|^2\\
=&N^{-2}\left\|\hat{\Phi}_{K_1\times K_1}^{-1}U_{N\times K_1}^{\top}R_N(\gamma^*)\right\|^2\\
=& N^{-2}\text{tr}\left(\hat{\Phi}_{K_1\times K_1}^{-1}U_{N\times K_1}^{\top}R_N(\gamma^*)R_N(\gamma^*)^{\top}U_{N\times K_1}\hat{\Phi}_{K_1\times K_1}^{-1}\right)\\
=& N^{-2}\text{tr}\left(U_{N\times K_1}^{\top}R_N(\gamma^*)R_N(\gamma^*)^{\top}U_{N\times K_1}\hat{\Phi}_{K_1\times K_1}^{-1}\hat{\Phi}_{K_1\times K_1}^{-1}\right)\\
=& N^{-2}\text{tr}\left(\hat{\Phi}_{K_1\times K_1}^{-1/2}U_{N\times K_1}^{\top}R_N(\gamma^*)R_N(\gamma^*)^{\top}U_{N\times K_1}\hat{\Phi}_{K_1\times K_1}^{-1/2}\hat{\Phi}_{K_1\times K_1}^{-1}\right)\\
\leq &\lambda_{\max}(\hat{\Phi}_{K_1\times K_1}^{-1})N^{-2}\text{tr}\left(\hat{\Phi}_{K_1\times K_1}^{-1/2}U_{N\times K_1}^{\top}R_N(\gamma^*)R_N(\gamma^*)^{\top}U_{N\times K_1}\hat{\Phi}_{K_1\times K_1}^{-1/2}\right)\\
=&\lambda_{\max}(\hat{\Phi}_{K_1\times K_1}^{-1})N^{-1}\text{tr}\left(R_N(\gamma^*)R_N(\gamma^*)^{\top}U_{N\times K_1}(U^{\top}_{N\times K_1}U_{N\times K_1})^{-1}U_{N\times K_1}^{\top}\right)\\
\leq &[\lambda_{\min}(\hat{\Phi}_{K_1\times K_1})]^{-1}N^{-1}\|R_N(\gamma^*)\|^2\\
=&[\lambda_{\min}(\hat{\Phi}_{K_1\times K_1})]^{-1}\cdot \frac{1}{N}\sum_{i=1}^N\left\{\mathbb{E}[{\pi}_0(T_i,\boldsymbol{X}_i)Y_i|T_i]-(\gamma^*)^{\top}u_{K_1}(T_i)\right\}^2=O_p(K_1^{-2\tilde{\alpha}}),
\end{align*}
where the first inequality follows from the fact that $tr(AB)\leq \lambda_{\max}(B)tr(A)$ for any symmetric matrix $B$ and positive semidefinite matrix $A$, the second inequality follows from the same fact and the fact that $U_{N\times K_1}(U^{\top}_{N\times K_1}U_{N\times K_1})^{-1}U_{N\times K_1}^{\top}$ is a projection matrix with maximum eigenvalue 1, and the last equality follows from the fact that $|\lambda_{\min}(\hat{\Phi}_{K_1\times K_1})|^{-1}=O_p(1)$ and the fact that $\frac{1}{N}\sum_{i=1}^N\left\{\mathbb{E}[{\pi}_0(T_i,\boldsymbol{X}_i)Y_i|T_i]-(\gamma^*)^{\top}u_{K_1}(T_i)\right\}^2\leq \sup_{t\in\mathcal{T}}|\mathbb{E}[{\pi}_0(T,X)Y|T]-(\gamma^*)^{\top}u_{K_1}(t)|^2=O(K_1^{-2\tilde{\alpha}})$. Thus we complete the proof of \eqref{claim:gamma_hat-star}.\\
\textbf{(Proof of Asymptotic Normality)}. We have the following decomposition for $\hat{\theta}(t)-\theta(t)$:
\begin{align*}
&\hat{\theta}_K(t)-\theta(t)\\
=&u_{K_1}(t)^{\top}(\hat{\gamma}-\gamma^*)+[(\gamma^*)^{\top}u_{K_1}(t)-\theta(t)]\\
=&u_{K_1}(t)^{\top}\left[\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)\bigg\{\hat{\pi}_K(T_i,\boldsymbol{X}_i)Y_i-\mathbb{E}\left[{\pi}_0(T_i,\boldsymbol{X}_i)Y_i|T_i\right]\bigg\}\right]\\
+&u_{K_1}(t)^{\top}\left[\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)\cdot \bigg\{\mathbb{E}\left[{\pi}_0(T_i,\boldsymbol{X}_i)Y_i|T_i\right]-(\gamma^*)^{\top}u_{K_1}(T_i)\bigg\}\right]\\
&+\bigg[(\gamma^*)^{\top}u_{K_1}(t)-\theta(t)\bigg]\\
=&b_{1N}(t)+b_{2N}(t)+b_{3N}(t)\ ,
\end{align*}
where
\begin{align*}
&b_{1N}(t)=u_{K_1}(t)^{\top}\left[\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)\bigg\{\hat{\pi}_K(T_i,\boldsymbol{X}_i)Y_i-\mathbb{E}\left[{\pi}_0(T_i,\boldsymbol{X}_i)Y_i|T_i\right]\bigg\}\right]\\
& b_{2N}(t)=u_{K_1}(t)^{\top}\left[\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\left[\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)\cdot \bigg\{\mathbb{E}\left[{\pi}_0(T_i,\boldsymbol{X}_i)Y_i|T_i\right]-(\gamma^*)^{\top}u_{K_1}(T_i)\bigg\}\right]\\
&b_{3N}(t)=(\gamma^*)^{\top}u_{K_1}(t)-\theta(t).
\end{align*}
Then we have that
\begin{align*}
&\sqrt{N}{V}_K(t)^{-1/2}\left[\hat{\theta}_K(t)-\theta(t)\right]\\
=&\sqrt{N}{V}_K(t)^{-1/2}b_{1N}(t)+\sqrt{N}{V}_K(t)^{-1/2}b_{2N}(t)+\sqrt{N}{V}_K(t)^{-1/2}b_{3N}(t).
\end{align*}
We shall show that $b_{1N}(t)$ contributes to the asymptotic variance; and $b_{2N}(t)+b_{3N}(t)$ contributes to the asymptotic bias which is asymptotically negligible. Thus to complete the proof of asymptotic normality, it is sufficient to prove the following results:
\begin{itemize}
\item[(i)] $V_K\geq c \|u_{K_1}(t)\|^2$ for some $c>0$;
\item[(ii)] $\sqrt{N}V_K^{-1/2}b_{1N}(t)\xrightarrow{d}N(0,1)$;
\item[(iii)] $\sqrt{N}V_K^{-1/2}b_{2N}(t)=o_p(1)$;
\item[(iv)] $\sqrt{N}V_K^{-1/2}b_{3N}(t)=o_p(1)$.
\end{itemize}
We first prove Result (i). Note that $\lambda_{\min}(\Sigma_{K_1\times K_1})\geq \underline{c}_{\sigma^2}\lambda_{\min}(\Phi_{K_1\times K_1})=\underline{c}_{\sigma^2}\lambda_{\min}(I_{K_1\times K_1})\geq \underline{c}_{\sigma^2}$, we can have
\begin{align*}
V_K=&u_{K_1}^{\top}(t) \Phi_{K_1\times K_1}^{-1}\Sigma_{K_1\times K_1}\Phi_{K_1\times K_1}^{-1}u_{K_1}(t)\\
\geq & \lambda_{\min}(\Sigma_{K_1\times K_1}) u_{K_1}^{\top}(t) \Phi_{K_1\times K_1}^{-1}\Phi_{K_1\times K_1}^{-1}u_{K_1}(t)\\
\geq & \underline{c}_{\sigma^2} \|u_{K_1}(t)\|^2.
\end{align*}
For the claim (ii). Let
\begin{align*}
&\tilde{b}_{1N}(t)=u_{K_1}(t)^{\top}\left[\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\right]^{-1}\\
&\quad\times\Bigg[\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)\bigg\{{\pi}_0(T_i,\boldsymbol{X}_i)Y_i-\mathbb{E}\left[{\pi}_0(T_i,\boldsymbol{X}_i)Y_i|T_i,\boldsymbol{X}_i\right]+\mathbb{E}\left[{\pi}_0(T_i,\boldsymbol{X}_i)Y_i|\boldsymbol{X}_i\right]-\mathbb{E}\left[{\pi}_0(T_i,\boldsymbol{X}_i)Y_i\right]\bigg\}\Bigg]\ .
\end{align*}
Similar to the proof of Lemma \ref{asy:equivalent}, we can have that
\begin{align*}
\sqrt{N}{V}_K(t)^{-1/2}\cdot ({b}_{1N}(t)-\tilde{b}_{1N}(t))=o_p(1)\ .
\end{align*}
Then
\begin{align*}
&\sqrt{N}{V}_K(t)^{-1/2}b_{1N}(t)\\
=&\sqrt{N}{V}_K(t)^{-1/2}\tilde{b}_{1N}(t)+o_p(1)\\
=&\sqrt{N}{V}_K(t)^{-1/2}u_{K_1}(t)^{\top}\hat{\Phi}_{K_1\times K_1}^{-1}N^{-1}U_{N\times K_1}^{\top}\mathcal{E}_N\\
=&\sqrt{N}{V}_K(t)^{-1/2}u_{K_1}(t)^{\top}{\Phi}_{K_1\times K_1}^{-1}N^{-1}U_{N\times K_1}^{\top}\mathcal{E}_N+\sqrt{N}{V}_K(t)^{-1/2}u_{K_1}(t)^{\top}[\hat{\Phi}_{K_1\times K_1}^{-1}-{\Phi}_{K_1\times K_1}^{-1}]N^{-1}U_{N\times K_1}^{\top}\mathcal{E}_N\\
=&B_{1N,1}(t)+B_{1N,2}(t)\ .
\end{align*}
For $B_{1N,1}(t)$, we can simply apply the Liapounov CLT and show that $B_{1N,1}(t)\xrightarrow{d}N(0,1)$. For $B_{1N,2}(t)$, let $\boldsymbol{T}=(T_1,...,T_N)$, we can obtain that
\begin{align*}
&\mathbb{E}\left[B_{1N,2}(t)^2|\boldsymbol{T}\right]\\
=&N^{-1}{V}_K(t)u_{K_1}(t)^{\top}[\hat{\Phi}_{K_1\times K_1}^{-1}-{\Phi}_{K_1\times K_1}^{-1}]\cdot \mathbb{E}\left[U_{N\times K_1}^{\top}\mathcal{E}_N\mathcal{E}_N^{\top}U_{N\times K_1}|\boldsymbol{T}\right]\cdot [\hat{\Phi}_{K_1\times K_1}^{-1}-{\Phi}_{K_1\times K_1}^{-1}]u_{K_1}(t)\\
=&\lambda_{\max}\left(N^{-1} \mathbb{E}\left[U_{N\times K_1}^{\top}\mathcal{E}_N\mathcal{E}_N^{\top}U_{N\times K_1}|\boldsymbol{T}\right]\right){V}_K(t)u_{K_1}(t)^{\top}[\hat{\Phi}_{K_1\times K_1}^{-1}-{\Phi}_{K_1\times K_1}^{-1}]\cdot [\hat{\Phi}_{K_1\times K_1}^{-1}-{\Phi}_{K_1\times K_1}^{-1}]u_{K_1}(t)\\
\leq &\lambda_{\max}\left(N^{-1} \mathbb{E}\left[U_{N\times K_1}^{\top}\mathcal{E}_N\mathcal{E}_N^{\top}U_{N\times K_1}|\boldsymbol{T}\right]\right)\left\{V_{K}^{-1}\|u_{K_1}(t)\|^2\right\}\cdot \|\hat{\Phi}_{K_1\times K_1}^{-1}-{\Phi}_{K_1\times K_1}^{-1}\|^2\\
=&O_p(1)O_p(1)o_p(1)=o_p(1)
\end{align*}
where we use the fact that $N^{-1} \mathbb{E}\left[U_{N\times K_1}^{\top}\mathcal{E}_N\mathcal{E}_N^{\top}U_{N\times K_1}|\boldsymbol{T}\right]=N^{-1}\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\sigma^2(T_i)$ has bounded maximum eigenvalue. Therefore, $B_{1N,1}(t)=o_p(1)$ by the conditional Chebyshev's inequality. Thus (ii) holds. \\
For (iii), by Cauchy-Schwarz's inequality, we can obtain that
\begin{align*}
&\sqrt{N}V_{K}^{-1/2}|b_{2N}(t)|\\
=&N^{-1/2}V_{K}^{-1/2}\left|u_{K_1}(t)^{\top}\hat{\Phi}_{K_1\times K_1}^{-1}U^{\top}_{ N\times K_1}R_N(\gamma^*)\right|\\
\leq &V_{K}^{-1/2}\left\{u_{K_1}(t)^{\top}\hat{\Phi}_{K_1\times K_1}^{-1}\left(N^{-1}U^{\top}_{N\times K_1}U_{N\times K_1}\right)\hat{\Phi}_{K_1\times K_1}^{-1}u_{K_1}(t)\right\}^{\frac{1}{2}} \left\{R_N(\gamma^*)^{\top}R_N(\gamma^*)\right\}^{\frac{1}{2}}\\
\leq &V_{K}^{-1/2}\left\{u_{K_1}(t)^{\top}\hat{\Phi}_{K_1\times K_1}^{-1}u_{K_1}(t)\right\}^{\frac{1}{2}} \left\{R_N(\gamma^*)^{\top}R_N(\gamma^*)\right\}^{\frac{1}{2}}\\
\leq & \{V_{K}^{-1/2}\|u_{K_1}(t)\|\}\cdot |\lambda_{\max}(\Phi_{K_1\times K_1}^{-1})|^{\frac{1}{2}}\cdot O(N^{\frac{1}{2}}\cdot K_1^{-\tilde{\alpha}})\\
=& O(1)\cdot O_p(1)\cdot o_p(1)=o_p(1)\ .
\end{align*}
Similarly, we can show show that $\sqrt{N}V_{K}^{-1/2}|b_{3N}(t)|=o_p(1)$. This completes the proof of the Theorem.
\subsection{Proof Theorem 7.2}
Similar to the proof of Lemma \ref{asy:equivalent}, let $\mu(t,x)=\mathbb{E}[Y|T=t,X=x]$, we decompose $\sqrt{N}(\hat{\psi}_K-\psi)$ as follows:
\begin{align}
&\sqrt{N}(\hat{\psi}_K-\psi)\notag\\
=&\frac{1}{\sqrt{N}}\sum_{i=1}^N\left\{\hat{\pi}_K(T_i,\boldsymbol{X}_i)Y_i-\psi\right\}\notag\\
=&\label{eq:WK_psi} \frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg\{ \left(\hat{\pi}_K(T_i,\boldsymbol{X}_i) - \pi_K^*(T_i,\boldsymbol{X}_i)\right)Y_i -\int_{\mathcal{T}}\int_{\mathcal{X}} \left(\hat{\pi}_K(t,\boldsymbol{x}) - \pi_K^*(t,\boldsymbol{x})\right)\mu(\boldsymbol{x},t) dF_{X,T}(\boldsymbol{x},t)\Bigg\} \\[2mm]
\label{eq:VK_psi} &~ + \frac {1} {\sqrt{N}} \sum_{i=1}^N \Bigg\{\left(\pi_K^*(T_i,\boldsymbol{X}_i) - {\pi_0(T_i,\boldsymbol{X}_i)}\right)Y_i - \int_{\mathcal{T}}\int_{\mathcal{X}} \mu(t,\boldsymbol{x})\left(\pi_K^*(t,\boldsymbol{x}) - {\pi(t,\boldsymbol{x})}\right) dF_{X,T}(\boldsymbol{x},t) \Bigg\} \\
\label{eq:Lemma1_psi}&~ + \sqrt{N}\int_{\mathcal{T}}\int_{\mathcal{X}} \mu(t,\boldsymbol{x})\left(\pi_K^*(t,\boldsymbol{x}) - {\pi_0(t,\boldsymbol{x})}\right) dF_{X,T}(\boldsymbol{x},t) \\[2mm]
&~ \label{eq:tauto_psi}+ {\sqrt{N}}\int_{\mathcal{T}}\int_{\mathcal{X}} \left(\hat{\pi}_K(t,\boldsymbol{x}) - \pi_K^*(t,\boldsymbol{x})\right)\mu(\boldsymbol{x},t)dF_{X,T}(\boldsymbol{x},t) \\
& \qquad - \sqrt{N}\int_{\mathcal{T}}\int_{\mathcal{X}} \mu(t,\boldsymbol{x})\rho''\left(u_{K_1}^{\top}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) m(t;\boldsymbol{\beta}_0) dF_{X,T}(\boldsymbol{x},t) \notag \\[4mm]
& \label{eq:Q_psi} +\sqrt{N} \int_{\mathcal{T}}\int_{\mathcal{X}} \varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\rho''\left(u_{K_1}^{\top}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) dF_{X,T}(\boldsymbol{x},t)\\
&\qquad - \sqrt{N} \int_{\mathcal{T}}\int_{\mathcal{X}} \mu(t,\boldsymbol{x})\rho''\left(u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)A^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) dF_{X,T}(\boldsymbol{x},t)\notag \\[4mm]
& \label{eq:proj_psi}+{\sqrt{N}}\int_{\mathcal{X}} \mu(t,\boldsymbol{x})\rho''\left(u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)u_{K_1}^{\top}(t)A^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) dF_{X,T}(\boldsymbol{x},t)\\
&\qquad +\frac{1}{\sqrt{N}}\sum_{i=1}^N \bigg\{ \pi_0(T_i,\boldsymbol{X}_i)\mu(T_i,\boldsymbol{X}_i)-\mathbb{E}\left[\pi_0(T,\boldsymbol{X})\mu(T,\boldsymbol{X})|\boldsymbol{X}=\boldsymbol{X}_i\right]\notag\\
&\qquad\qquad\qquad\qquad -\mathbb{E}\left[\pi_0(T,\boldsymbol{X})\mu(T,\boldsymbol{X})|T=T_i\right]+\mathbb{E}\left[\pi_0(T,\boldsymbol{X})\mu(T,\boldsymbol{X})\right]\bigg\} \notag \\[4mm]
&~ + \frac {1} {\sqrt{N}} \sum_{i=1}^N \bigg\{ \pi_0(T_i,\boldsymbol{X}_i)Y_i-\pi_0(T_i,\boldsymbol{X}_i)\mu(T_i,\boldsymbol{X}_i)+\mathbb{E}\left[\pi_0(T,\boldsymbol{X})\mu(T,\boldsymbol{X})|T=T_i\right]\label{eq:Normal_psi}\\
& \quad \quad \quad \quad \quad \quad \quad+\mathbb{E}\left[\pi_0(T,\boldsymbol{X})\mu(T,\boldsymbol{X})|\boldsymbol{X}=\boldsymbol{X}_i\right]-\mathbb{E}\left[\pi_0(T,\boldsymbol{X})\mu(T,\boldsymbol{X})\right]-\psi \bigg\}\ .\notag
\end{align}
Using the similar argument for showing that \eqref{eq:WK}-\eqref{eq:proj} are all $o_p(1)$ in the proof of Lemma \ref{asy:equivalent}, we can obtain that the terms \eqref{eq:WK_psi}-\eqref{eq:proj_psi} are all $o_p(1)$, while the term \eqref{eq:Normal_psi} is asymptotically normal and attains the efficiency bound.
\singlespacing
\bibliographystyle{econometrica}
\bibliography{Semiparametric}