EconBase
← Back to paper

A Unified Framework for Efficient Estimation of General Treatment Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

229,226 characters · 20 sections · 37 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Supplemental Material for “A Unified Framework for Efficient Estimation of General Treatment Models ”

\setstretch{1.3}

\listoftables

Complete Simulation Results

In this supplemental material, we present complete simulation results that are partly omitted in the main paper Ai_Motegi_Zhang_cts_treat. See Section (ref) for simulation designs and Section (ref) for results.

Simulation Design

We consider four data generating processes (DGPs):

description{1.0cm} • $T_{i} = X_{1i} + X_{2i} + 0.2 X_{3i} + 0.2 X_{4i} + \xi_{i}$ and $Y_{i} = 1 + X_{1i} + 0.1 X_{2i} + 0.1 X_{3i} + 0.1 X_{4i} + T_{i} + \epsilon_{i}$. • $T_{i} = (X_{1i} + 0.5)^{2} + 0.4 X_{2i} + 0.4 X_{3i} + 0.4 X_{4i} + \xi_{i}$ and $Y_{i} = 1 + X_{1i} + 0.1 X_{2i} + 0.1 X_{3i} + 0.1 X_{4i} + T_{i} + \epsilon_{i}$. • $T_{i} = X_{1i} + X_{2i} + 0.2 X_{3i} + 0.2 X_{4i} + \xi_{i}$ and $Y_{i} = 0.75 X_{1i}^{2} + 0.2 (X_{2i} + 0.5)^{2} + T_{i} + \epsilon_{i}$. • $T_{i} = (X_{1i} + 0.5)^{2} + 0.4 X_{2i} + 0.4 X_{3i} + 0.4 X_{4i} + \xi_{i}$ and $Y_{i} = 0.75 X_{1i}^{2} + 0.2 (X_{2i} + 0.5)^{2} + T_{i} + \epsilon_{i}$.

For each DGP, $\xi_{i} \stackrel{i.i.d.}{\sim} N(0, 9)$, $\epsilon_{i} \stackrel{i.i.d.}{\sim} N(0, 25)$, and $\boldsymbol{X}_{i} = [X_{1i}, X_{2i}, X_{3i}, X_{4i}]^{T} \stackrel{i.i.d.}{\sim} N(0, \boldsymbol{\Sigma})$, where the diagonal elements of $\boldsymbol{\Sigma}$ are all 1 and the off-diagonal elements are all $\rho \in \{ 0.0, 0.2, 0.4 \}$. We draw $J = 1000$ Monte Carlo samples with sample size $N \in \{ 100, 500, 1000 \}$.

Our DGPs have similar structures with Fong_Hazlett_Imai_2018. The characteristics of the four DGPs can be summarized as follows.

description{1.0cm} • $T_{i}$ is linear in $\boldsymbol{X}_{i}$; $Y_{i}$ is linear in $\boldsymbol{X}_{i}$ and $T_{i}$. • $T_{i}$ is nonlinear in $\boldsymbol{X}_{i}$; $Y_{i}$ is linear in $\boldsymbol{X}_{i}$ and $T_{i}$. • $T_{i}$ is linear in $\boldsymbol{X}_{i}$; $Y_{i}$ is nonlinear in $\boldsymbol{X}_{i}$ and linear in $T_{i}$. • $T_{i}$ is nonlinear in $\boldsymbol{X}_{i}$; $Y_{i}$ is nonlinear in $\boldsymbol{X}_{i}$ and linear in $T_{i}$.

In the main paper Ai_Motegi_Zhang_cts_treat, we discuss only DGP-1 and DGP-4 in order to save space. (DGP-4 here is called “DGP-2” in the main paper since DGP-2 and DGP-3 here are skipped.) As shown in the main paper, a parametric version of Fong, Hazlett, and Imai's Fong_Hazlett_Imai_2018 covariate balancing generalized propensity score (CBGPS) estimators produces no bias under DGP-1 and considerable bias under DGP-4.\footnote{ A non-parametric version of the CBGPS estimators is also biased under a data generating process which is similar to DGP-4, and the magnitude of the bias is as large as the parametric version Fong_Hazlett_Imai_2018. } Here we also consider DGP-2 and DGP-3, which serve as intermediate cases where either $T_{i}$ is nonlinear in $\boldsymbol{X}_{i}$ or $Y_{i}$ is nonlinear in $\boldsymbol{X}_{i}$.

Since $Y_{i}$ is linear in $T_{i}$ under all four DGPs, the link function is always of the form $\mathbb{E} [Y^*(t)] = \beta_{1} + \beta_{2} t$. It is straightforward to see that the true coefficients are $(\beta_{1}, \beta_{2}) = (1, 1)$ for each DGP. $\beta_{2}$ is of greater interest than $\beta_{1}$ since $\beta_{2}$ measures the average treatment effect.

To estimate $(\beta_{1}, \beta_{2})$ via our proposed method, we should decide which variables to include in polynomials with respect to $T_{i}$ and $\boldsymbol{X}_{i}$. For each DGP, we use $u_{K_{1}} (t_{i}) = [1, t_{i}, t_{i}^{2}]^{T}$ (i.e. $K_{1} = 3$) and $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}]^{T}$ (i.e. $K_{2} = 9$) or $u_{K_{2}} (\boldsymbol{X}_{i}) = [1, X_{1i}, X_{2i}, X_{3i}, X_{4i}, X_{1i}^{2}, X_{2i}^{2}, X_{3i}^{2}, X_{4i}^{2}, X_{1i} X_{2i}, X_{1i} X_{3i}, X_{1i} X_{4i}, X_{2i} X_{3i}, X_{2i} X_{4i}, X_{3i} X_{4i}]^{T}$ (i.e. $K_{2} = 15$). In the literature of non-parametric estimation of treatment effects, it is common to include the first and second moments of $\boldsymbol{X}_{i}$ with or without cross terms chan2016globally. Hence we consider two cases with and without the cross terms of $\boldsymbol{X}_{i}$.

We compare our estimators with the parametric version of Fong, Hazlett, and Imai's Fong_Hazlett_Imai_2018 CBGPS estimator. The non-parametric version of CBGPS is omitted since the parametric and non-parametric versions exhibit similar performance according to the simulation results of Fong_Hazlett_Imai_2018. Computation of the parametric CBGPS estimator involves two steps. The first step is the estimation of stabilized weights, and the second step is the estimation of average treatment effects. Our covariates are chosen to be $\boldsymbol{X}_{i} = [X_{1i}, X_{2i}, X_{3i}, X_{4i}]^{T}$ for the first step and $\boldsymbol{Z}_{i} = [1, T_{i}, \boldsymbol{X}_{i}^{T}]^{T}$ for the second step.\footnote{ In extra simulations not reported here, we added some irrelevant covariates to the model, as Fong_Hazlett_Imai_2018 did, in order to see whether the finite sample performance is sensitive to redundant covariates. Since the results were not so sensitive to the inclusion of irrelevant covariates, we only report results without the irrelevant covariates here. }

Simulation Results

In this section we present our simulation results. See Tables (ref)-(ref) for DGP-1; Tables (ref)-(ref) for DGP-2; Tables (ref)-(ref) for DGP-3; Tables (ref)-(ref) for DGP-4. For each table, we report the bias, standard deviation, root mean squared error, and coverage probability based on the 95% confidence bands.

As discussed in the main paper Ai_Motegi_Zhang_cts_treat, both of our stabilized-weight (SW) estimator and Fong, Hazlett, and Imai's Fong_Hazlett_Imai_2018 CBGPS estimator are unbiased under DGP-1. The latter often has smaller bias and standard deviation across all sample sizes $N \in \{ 100, 500, 1000 \}$. That is not a surprising result since DGP-1 satisfies the linearity assumption underpinning the CBGPS estimator. The coverage probability, however, is sometimes comparable between the two estimators. See, for example, Table (ref) ($\beta_{2}$) with $\rho = 0.2$ and $N = 500$, where the coverage probability is 0.926 for the SW estimator with $K_{2} = 9$ and 0.890 for the CBGPS estimator. The simulation results are not sensitive to the correlation coefficient $\rho \in \{ 0.0, 0.2, 0.4 \}$. The SW estimators with and without the cross terms of $\boldsymbol{X}_{i}$ lead to roughly similar performance.

Similar implications hold for DGP-2. The CBGPS estimator performs well arguably due to a negligibly small degree of nonlinearity. There is a nonlinear term $(X_{1i} + 0.5)^{2}$ in the DGP of $T_{i}$, but the DGP of $Y_{i}$ is kept to be linear. Hence the CBGPS estimator maintains the sharp performance as in DGP-1.

Under DGP-3, the SW estimator begins to be more comparable with the CBGPS estimator. See, for example, Table (ref) ($\beta_{2}$) with $\rho = 0.4$ and $N = 1000$, where the bias is 0.003 for the SW estimator with $K_{2} = 9$ and 0.004 for the CBGPS estimator; the standard deviation is 0.174 for SW and 0.090 for CBGPS; the coverage probability is 0.932 for SW and 0.892 for CBGPS. Under DGP-3, there are two nonlinear terms $0.75 X_{1i}^{2}$ and $0.2 (X_{2i} + 0.5)^{2}$ in the DGP of $Y_{i}$. It is likely that the CBGPS estimator is negatively affected by those nonlinear features while the SW estimator is more robust against nonlinearity.

As discussed in the main paper Ai_Motegi_Zhang_cts_treat, our SW estimator dominates the CBGPS estimator under DGP-4. See, for example, Table (ref) ($\beta_{1}$) with $\rho = 0.4$ and $N = 1000$, where the bias is 0.016 for the SW estimator with $K_{2} = 9$ and -0.182 for the CBGPS estimator; the standard deviation is 0.605 for SW and 0.194 for CBGPS; the coverage probability is 0.923 for SW and 0.800 for CBGPS. Also see Table (ref) ($\beta_{2}$) with $\rho = 0.4$ and $N = 1000$, where the bias is 0.038 for SW and 0.156 for CBGPS; the standard deviation is 0.170 for SW and 0.080 for CBGPS; the coverage probability is 0.916 for SW and 0.304 for CBGPS.

Under DGP-4, the CBGPS estimator keeps producing negative bias in $\beta_{1}$ and positive bias in $\beta_{2}$. The bias is considerably large, and it does not vanish as sample size $N$ increases. Our estimator, in contrast, produces virtually no bias for any sample size $N \in \{ 100, 500, 1000 \}$. It also achieves accurate coverage probability for larger sample sizes $N \in \{ 500, 1000 \}$. The relative advantage of our estimator stems from the strong degree of nonlinearity contained in DGP-4, where both $T_{i}$ and $Y_{i}$ have nonlinear terms.

In summary, our estimator is by construction robust against the functional form of underlying DGPs, while Fong, Hazlett, and Imai's Fong_Hazlett_Imai_2018 estimator is sensitive to model misspecification. In fact, our approach produces unbiased estimates under all DGPs considered, while Fong, Hazlett, and Imai's Fong_Hazlett_Imai_2018 approach produces biased estimates under DGP-4. In reality, the functional form of an underlying DGP is unknown to the researcher. It is therefore of practical use to employ our estimator in order to accomplish correct inference for dose-response curves.

table[table omitted — 4,239 chars of source]
table[table omitted — 4,230 chars of source]
table[table omitted — 4,263 chars of source]
table[table omitted — 4,248 chars of source]
table[table omitted — 4,226 chars of source]
table[table omitted — 4,222 chars of source]
table[table omitted — 4,249 chars of source]
table[table omitted — 4,233 chars of source]

Assumptions

assumption[Unconfounded Treatment Assignment] For all $t\in \mathcal{T}$, given $\boldsymbol{X}$ , $T$ is independent of $Y^{\ast }(t)$, i.e., $Y^{\ast }(t)\perp T|\boldsymbol{X,}$ for all $t\in \mathcal{T}$.
assumptionThe support $\mathcal{X}$ of $\boldsymbol{X}$ is a compact subset of $\mathbb{R}^{r}$. The support $\mathcal{T}$ of the treatment variable $T$ is a compact subset of $\mathbb{R}$.
assumptionThere exist two positive constants $\eta_1$ and $\eta_2$ such that \begin{equation*} 0 < \eta_1 \leq \pi_0(t,\boldsymbol{x}) \leq \eta_2 <\infty \ ,\ \forall (t, \boldsymbol{x}) \in \mathcal{T}\times\mathcal{X}\ . \end{equation*}
assumptionThere exist $\Lambda _{K_{1}\times K_{2}}\in \mathbb{R} ^{K_{1}\times K_{2}}$ and a positive constant $\alpha >0$ such that \begin{equation*} \sup_{(t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X}}\left\vert (\rho ^{\prime -1}\left( \pi _{0}(t,\boldsymbol{x})\right) -u_{K_{1}}(t)^{\top }\Lambda _{K_{1}\times K_{2}}v_{K_{2}}(\boldsymbol{x})\right\vert =O(K^{-\alpha }). \end{equation*}
assumptionFor every $K_1$ and $K_2$, the smallest eigenvalues of $ \mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^\top\right]$ and $\mathbb{E}\left[ v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^\top\right]$ are bounded away from zero uniformly in $K_1$ and $K_2$.
assumptionThere are two sequences of constants $\zeta _{1}(K_{1})$ and $\zeta _{2}(K_{2})$ satisfying\newline $\sup_{t\in \mathcal{T}}\Vert u_{K_{1}}(t)\Vert \leq \zeta _{1}(K_{1})$ and $ \sup_{\boldsymbol{x}\in \mathcal{X}}\Vert v_{K_{2}}(\boldsymbol{x})\Vert \leq \zeta _{2}(K_{2})$, $K=K_{1}(N)K_{2}(N)$ and $\zeta (K)=\zeta _{1}(K_{1})\zeta _{2}(K_{2})$, such that $\zeta (K)K^{-\alpha }\rightarrow 0$ and $\zeta (K)\sqrt{K/N}\rightarrow 0$ as $N\rightarrow \infty $.
assumptionThe parameter space $\Theta\subset\mathbb{R} ^{p}$ is a compact set and the true parameter $\boldsymbol{\beta}_0$ is in the interior of $\Theta$ , where $p\in\mathbb{N}$.
assumptionThere exists a unique solution $\boldsymbol{\beta }_{0}$ for the optimization problem \begin{equation*} \min_{\boldsymbol{\beta }\in \Theta }\int_{\mathcal{T}}\mathbb{E}\left[ L(Y^{\ast }(t)-g(t;\boldsymbol{\beta }))\right] dF_{T}(t)\ . \end{equation*}
assumption$\mathbb{E}\left[\sup_{\boldsymbol{\beta} \in \Theta }|L\left( Y-g(T;\boldsymbol{\beta} )\right)|^{2}\right]<\infty $.
assumptionThe following conditions hold true: \begin{enumerate} • $g(t;\boldsymbol{\beta })$ is twice continuously differentiable in $ \boldsymbol{\beta }\in \Theta $; • $L(Y-g(T;\boldsymbol{\beta }))$ is differentiable in $\boldsymbol{\beta }$ with probability one, i.e., for any directional vector $\boldsymbol{ \ \eta }\in \mathbb{R}^{p}$, there exists an integrable random variable $ L^{\prime }(Y-g(T;\boldsymbol{\beta }))$ such that{ \ \begin{equation*} \mathbb{P}\left( \lim_{\epsilon \rightarrow 0}\frac{L(Y-g(T;\boldsymbol{\ \beta }+\epsilon \boldsymbol{\eta }))-L(Y-g(T;\boldsymbol{\beta }))}{ \epsilon }=L^{\prime }(Y-g(T;\boldsymbol{\beta }))\cdot \left\langle m(T; \boldsymbol{\beta }),\boldsymbol{\eta }\right\rangle _{\mathbb{R} ^{p}}\right) =1, \end{equation*} } where $\left\langle \cdot ,\cdot \right\rangle _{\mathbb{R}^{p}}$ is the inner product in Euclidean space $\mathbb{R}^{p}$; • $\mathbb{E}\left[ L^{\prime }(Y-g(T;\boldsymbol{\beta }_{0}))^{2} \right] <\infty $. \end{enumerate}
assumptionSuppose that \begin{equation*} \frac{1}{N}\sum_{i=1}^{N}\hat{\pi}_{K}(T_{i},\boldsymbol{X}_{i})L^{\prime }\left( Y_{i}-g(T_{i};\hat{\boldsymbol{\beta }})\right) m(T_{i};\hat{ \boldsymbol{\beta }})=0 \end{equation*} holds with probability approaching one.
assumption$\mathbb{E}\left[ \pi _{0}(T,\boldsymbol{X})L^{\prime }(Y-g(T; \boldsymbol{\beta }))m(T;\boldsymbol{\beta })\right] $ is differentiable with respect to $\boldsymbol{ \beta}$ and $H_{0}:=-\nabla _{\beta }\mathbb{E}\left[ \pi _{0}(T,\boldsymbol{X})L^{\prime }(Y-g(T;\boldsymbol{\beta }))m(T;\boldsymbol{ \beta })\right] \Big|_{\boldsymbol{\beta }=\boldsymbol{\beta }_{0}}$ is nonsingular.
assumption$\varepsilon (t,\boldsymbol{x};{\boldsymbol{ \beta }}_{0}):=\mathbb{E}[L^{\prime }(Y-g(T;\boldsymbol{\beta }_{0}))|T=t, \boldsymbol{X}=\boldsymbol{x}]$ is continuously differentiable in $(t, \boldsymbol{x})$.
assumption\ \begin{enumerate} • $\mathbb{E}\left[ \sup_{\boldsymbol{\beta} \in \Theta }|L^{\prime }(Y-g(T;\boldsymbol{\beta} ))^{2+\delta }\right] <\infty $ for some $\delta >0$; • The function class $\{L^{\prime }(y-g(t;\boldsymbol{\beta} )):\boldsymbol{\beta} \in \Theta \}$ satisfies: \begin{equation*} \mathbb{E}\left[ \sup_{\boldsymbol{\beta} _{1}:\Vert \boldsymbol{\beta}_{1}-\boldsymbol{\beta} \Vert <\delta }\left\vert L^{\prime }(Y-g(T;\boldsymbol{\beta} _{1}))-L^{\prime }(Y-g(T;\boldsymbol{\beta} ))\right\vert ^{2}\right] ^{1/2}\leq a\cdot \delta ^{b} \end{equation*} for any $\forall \boldsymbol{\beta}\in \Theta $ and any small $\delta >0$ and for some finite positive constants $a$ and $b$. \end{enumerate}
assumption$\zeta (K)\sqrt{K^{4}/N}\rightarrow 0$ and $\sqrt{N} K^{-\alpha}\to 0$ as $N\to \infty$.

Efficiency Bound

Proof of Theorem 3.1

Without loss of generality, we only consider the distribution of $(T,\boldsymbol{X},Y)$ to be absolutely continuous with respect to Lebesgue measure, i.e., there exists a density function $f_{T,X,Y}(t,\boldsymbol{x},y)$ such that $dF_{T,X,Y}(t,\boldsymbol{x},y)=f_{T,X,Y}(t,\boldsymbol{x},y)dtd\boldsymbol{x}dy$. For discrete cases, the proof can be established by using a similar argument.\\

We follow the approach of Bickel_Klaassen_Ritov_Wellner_1993 to derive the variance bound of $\boldsymbol{\beta}_0$, see also tchetgen2012semiparametric. Let $\left\{f^{\alpha}_{Y,T,X}(y,t,\boldsymbol{x})\right\}_{\alpha \in\mathbb{R}}$ denote a one dimensional regular parametric submodel with $ f^{\alpha=0}_{Y,T,X}(y,t,\boldsymbol{x})=f_{Y,T,X}(y,t,\boldsymbol{x})$. By definition, $\boldsymbol{\beta}_0$ solves following equation:

align[align omitted — 178 chars of source]

By Assumption (ref), (ref) is equivalent to

align[align omitted — 248 chars of source]

Therefore, the parameter $\boldsymbol{\beta}(\alpha)$ induced by the submodel $f^{\alpha}_{Y,T,X}(y,t,\boldsymbol{x})$ satisfies:

align[align omitted — 314 chars of source]

where $\mathbb{E}^{\alpha}\left[\cdot|T=t,\boldsymbol{X}=\boldsymbol{x}\right]$ denotes taking expectation with respect to the submodel $f^{\alpha}_{Y|T,X}(\cdot|t,\boldsymbol{x})$. \\

Differentiating both sides of (ref) with respect to $\alpha$, evaluating at $\alpha = 0$ and using the condition $Y^*(t)\perp T|\boldsymbol{X}$, we can deduce that

align*[align* omitted — 5,594 chars of source]

Since $H_0=- \nabla_{\beta}\left\{\int_{\mathcal{T}} \mathbb{E}[L'(Y^*(t)-g(t;\boldsymbol{\beta}))]\cdot m(t;\boldsymbol{\beta})f_{T}(t) dt\right\}\Bigg|_{\boldsymbol{\beta}=\boldsymbol{\beta}_0}$ is invertible by Assumption (ref), we get

align[align omitted — 944 chars of source]

The efficient influence function of $\boldsymbol{\beta}_0$, denoted by $S_{eff}(Y,T,\boldsymbol{X};\boldsymbol{\beta}_0)$, is a unique function satisfying the following equation:

align[align omitted — 290 chars of source]

Therefore, to justify our theorem, it suffices to substitute $S_{eff}(Y,T,\boldsymbol{X};\boldsymbol{\beta}_0) = H_0 ^{-1} \psi(Y,T,\boldsymbol{X};\boldsymbol{\beta}_0)$ into (ref) and check the validity. Note that

align[align omitted — 1,005 chars of source]

For the term (ref), we have

align*[align* omitted — 1,364 chars of source]

For the term (ref), we have

align*[align* omitted — 2,386 chars of source]

where the first equality holds in accordance with the definition of $\int_{\mathcal{Y}}L'(y-g(t;\boldsymbol{\beta}_0))f_{Y|X,T}(y|\boldsymbol{x},t)dy=:\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)$. \\

For the term (ref), we have

align*[align* omitted — 2,165 chars of source]

We have proved (ref) holds, hence $S_{eff}$ is the efficient influence function of $\boldsymbol{\beta}_0$.

Particular Case I: Binary Treatment Effects

In this section, we show that when $T\in\{0,1\}$, $g(t;\boldsymbol{\beta})=\beta_0+\beta_1\cdot t$ and $L(v)=v^2$, our general efficiency bound derived in Theorem 3.1 reduces to the well-known efficiency bound for average treatment effects in robins1994estimation and hahn1998role. In accordance with our identification condition, $\beta_0$ and $\beta_1$ are identified by minimizing the following loss function $$\sum_{t\in\{0,1\}}\mathbb{E}[(Y^*(t)-\beta_0-\beta_1\cdot t)^2]\cdot \mathbb{P}(T=t).$$ The solutions are given by $$\beta_0=\mathbb{E}[Y^*(0)], \ \beta_1=\mathbb{E}[Y^*(1)-Y^*(0)].$$ Here $\beta_1$ is the average treatment effects.

corSuppose $T\in\{0,1\}$, $L(v)=v^2$, $g(t;\boldsymbol{\beta})=\beta_0+\beta_1\cdot t$ and the conditions in Theorem 3.1 hold, the efficient influence functions of $\beta_0$ and $\beta_1$ given by Theorem 3.1 reduce to \begin{align*} &S_{eff}(T,\boldsymbol{X},Y;\beta_0)=\phi_2(T,\bold{X},Y;\beta_0), \\ &S_{eff}(T,\boldsymbol{X},Y;\beta_1,\beta_0)=\phi_2(T,\bold{X},Y;\beta_0)-\phi_1(T,\bold{X},Y;\beta_1,\beta_0), \end{align*} where \begin{align*} &\phi_1(T,\boldsymbol{X},Y;\boldsymbol{\beta})=\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot Y^*(1)- \left\{\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}-1\right\}\cdot \mathbb{E}[Y^*(1)|\boldsymbol{X}] -\beta_0-\beta_1, \\ &\phi_2(T,\boldsymbol{X},Y;\boldsymbol{\beta})=\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot Y^*(0)- \left\{\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}-1\right\}\cdot \mathbb{E}[Y^*(0)|\boldsymbol{X}] -\beta_0, \end{align*} and they are the same as the efficient influence functions given in robins1994estimation and hahn1998role.
proofUsing our notation, we have \begin{align*} &\boldsymbol{\beta}_0=(\beta_0,\beta_1)^{\top}\ ,\ g(t;\boldsymbol{\beta}_0)=\beta_0+\beta_1 \cdot t , \quad m(t;\boldsymbol{\beta}_0)=\begin{bmatrix} 1 \\ t \end{bmatrix}\ ,\ H_0=\mathbb{E}\left[m(T;\boldsymbol{\beta}_0)m(T;\boldsymbol{\beta}_0)^{\top}\right]\ ,\\ &\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)=T\cdot\left\{ \mathbb{E}[Y^*(1)-Y^*(0)|\boldsymbol{X}]-\beta_1\right\}+ \mathbb{E}[Y^*(0)|\boldsymbol{X}]-\beta_0, \\ &\pi_0(T,\boldsymbol{X})=\frac{T\cdot p+(1-T)\cdot q}{T\cdot \mathbb{P}(T=1|\boldsymbol{X})+T\cdot \mathbb{P}(T=0|\boldsymbol{X})}=\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p +\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q, \end{align*} where $p=\mathbb{P}(T=1)$ and $q=\mathbb{P}(T=0)$. In accordance with our Theorem 3.1, the efficient influence function of $(\beta_0,\beta_1)$ is \begin{align*} H^{-1}_0\bigg\{\pi_0 (T, \boldsymbol{X})m(T;\boldsymbol{\beta}_0)\left\{ Y-\mathbb{E}[Y| \boldsymbol{X},T]\right\} +\mathbb{E}\left[ \varepsilon (T,\boldsymbol{X}; \boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)| \boldsymbol{X}\right]\bigg\}\ . \end{align*} With some computation, we have \begin{align} H^{-1}_0 =\begin{bmatrix} 1& p \\ p& p \end{bmatrix}^{-1}=\frac{1}{pq}\cdot \begin{bmatrix} p& -p \\ -p& 1 \end{bmatrix}=\begin{bmatrix} \frac{1}{q}& -\frac{1}{q} \\ -\frac{1}{q}& \frac{1}{pq} \end{bmatrix}\ . \end{align} and \begin{align} &\pi_0 (T, \boldsymbol{X})m(T;\boldsymbol{\beta}_0)\left\{ Y-\mathbb{E}[Y| \boldsymbol{X},T]\right\} \notag\\ =&\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p \cdot \begin{bmatrix} 1 \\ T \end{bmatrix}\cdot\bigg\{Y-T\cdot \mathbb{E}[Y^*(1)|\boldsymbol{X}]-(1-T)\cdot \mathbb{E}[Y^*(0)|\boldsymbol{X}]\bigg\} \notag\\ &+\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q\cdot \begin{bmatrix} 1 \\ T \end{bmatrix}\cdot\bigg\{Y-T\cdot \mathbb{E}[Y^*(1)|\boldsymbol{X}]-(1-T)\cdot \mathbb{E}[Y^*(0)|\boldsymbol{X}]\bigg\} \notag\\ =&\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p \cdot \begin{bmatrix} 1 \\ 1 \end{bmatrix}\cdot\bigg\{Y^*(1)- \mathbb{E}[Y^*(1)|\boldsymbol{X}]\bigg\} +\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q \cdot \begin{bmatrix} 1 \\ 0 \end{bmatrix}\cdot\bigg\{Y^*(0)- \mathbb{E}[Y^*(0)|\boldsymbol{X}]\bigg\} \notag\\ =&\begin{bmatrix} \frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot \left\{Y^*(1)- \mathbb{E}[Y^*(1)|\boldsymbol{X}]\right\}\cdot p+\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot \left\{Y^*(0)- \mathbb{E}[Y^*(0)|\boldsymbol{X}]\right\}\cdot q \\[4mm] \frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot \left\{Y^*(1)- \mathbb{E}[Y^*(1)|\boldsymbol{X}]\right\}\cdot p \end{bmatrix} \end{align} and \begin{align} &\mathbb{E}\left[ \varepsilon (T,\boldsymbol{X}; \boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)| \boldsymbol{X}\right] \notag\\ =&\mathbb{E}\Bigg[\bigg(T\cdot\left\{ \mathbb{E}[Y^*(1)-Y^*(0)|\boldsymbol{X}]-\beta_1\right\}+ \mathbb{E}[Y^*(0)|\boldsymbol{X}]-\beta_0\bigg)\cdot \frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p \cdot \begin{bmatrix} 1 \\ T \end{bmatrix} \bigg|\boldsymbol{X}\Bigg] \notag\\ &+\mathbb{E}\Bigg[\bigg(T\cdot\left\{ \mathbb{E}[Y^*(1)-Y^*(0)|\boldsymbol{X}]-\beta_1\right\}+ \mathbb{E}[Y^*(0)|\boldsymbol{X}]-\beta_0\bigg)\cdot \frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q \cdot \begin{bmatrix} 1 \\ T \end{bmatrix} \bigg|\boldsymbol{X} \Bigg] \notag \\ =&\mathbb{E}\Bigg[\bigg( \mathbb{E}[Y^*(1)|\boldsymbol{X}]-\beta_1-\beta_0\bigg)\cdot \frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p \cdot \begin{bmatrix} 1 \\ 1 \end{bmatrix} \bigg|\boldsymbol{X}\Bigg] \notag\\ &+\mathbb{E}\Bigg[\bigg( \mathbb{E}[Y^*(0)|\boldsymbol{X}]-\beta_0\bigg)\cdot \frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q \cdot \begin{bmatrix} 1 \\ 0 \end{bmatrix} \bigg|\boldsymbol{X} \Bigg] \notag\\ =&\begin{bmatrix} \bigg(\mathbb{E}[Y^*(1)|\boldsymbol{X}]-\beta_1-\beta_0\bigg)\cdot p +\bigg(\mathbb{E}[Y^*(0)|\boldsymbol{X}]-\beta_1\bigg)\cdot q \\[4mm] \bigg(\mathbb{E}[Y^*(1)|\boldsymbol{X}]-\beta_1-\beta_0\bigg)\cdot p \end{bmatrix}. \end{align} Therefore, with (ref), (ref), and (ref) we can obtain that \begin{align*} &\pi_0 (T, \boldsymbol{X})m(T;\boldsymbol{\beta}_0)\left\{ Y-\mathbb{E}[Y| \boldsymbol{X},T]\right\}+\mathbb{E}\left[ \varepsilon (T,\boldsymbol{X}; \boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)| \boldsymbol{X}\right]\\ =&\begin{pmatrix} p\cdot \phi_1(T,\boldsymbol{X},Y;\boldsymbol{\beta}_0)+q\cdot \phi_2(T,\boldsymbol{X},Y;\boldsymbol{\beta}_0) \\[2mm] p\cdot \phi_1(T,\boldsymbol{X},Y;\boldsymbol{\beta}_0) \end{pmatrix}, \end{align*} and the efficient influence functions of $\beta_1$ and $\beta_2$ are given by \begin{align*} &\begin{bmatrix} \frac{1}{q}& -\frac{1}{q} \\ -\frac{1}{q}& \frac{1}{pq} \end{bmatrix}\cdot \begin{pmatrix} p\cdot \phi_1(T,\boldsymbol{X},Y;\boldsymbol{\beta})+q\cdot \phi_2(T,\boldsymbol{X},Y;\boldsymbol{\beta}) \\[2mm] p\cdot \phi_1(T,\boldsymbol{X},Y;\boldsymbol{\beta}) \end{pmatrix}=\begin{pmatrix} \phi_2(T,\boldsymbol{X},Y;\boldsymbol{\beta}) \\[2mm] \phi_1(T,\boldsymbol{X},Y;\boldsymbol{\beta})- \phi_2(T,\boldsymbol{X},Y;\boldsymbol{\beta})\ . \end{pmatrix}. \end{align*}

Particular Case II: Multiple Treatment Effects

In this section, we show that when $T\in\{0,1,...,J\}$, $J\in\mathbb{N}$, $g(t;\boldsymbol{\beta})=\sum_{j=0}^J \beta_j \cdot I(t=j)$ and $L(v)=v^2$, our general efficiency bound derived in Theorem 3.1 reduces to the efficiency bound of multi-level treatment effects given in cattaneo2010efficient. In accordance with our proposed identification condition, $\{\beta_j\}_{j=0}^J$ are identified by minimizing the following loss function $$\sum_{j=0}^J\mathbb{E}\left[(Y^*(j)-\beta_j)^2\right]\cdot \mathbb{P}(T=j).$$ The solutions are $\beta_j=\mathbb{E}[Y^*(j)]$ for $j\in\{0,...,J\}$.

corSuppose $T\in\{0,1,...,J\}$, $J\in\mathbb{N}$, $g(t;\boldsymbol{\beta})=\sum_{j=0}^J \beta_j \cdot I(t=j)$, $L(v)=v^2$, and the conditions in Theorem 3.1 hold, the efficient influence functions of $\{\beta_j\}_{j=0}^J$ given by Theorem 3.1 reduce to \begin{align*} S_{eff}(T,\boldsymbol{X},Y;\beta_j)=\frac{I(T=j)}{\mathbb{P}(T=j|\boldsymbol{X})}\cdot \left\{Y^*(j)-\mathbb{E}[Y^*(j)|\boldsymbol{X}]\right\} + \mathbb{E}[Y^*(j)|X]- \beta_j,\ j\in\{0,...,J\}, \end{align*} and they are the same as the efficient influence functions given in cattaneo2010efficient.
proofUsing our notation, we have \begin{align*} \boldsymbol{\beta}_0=(\beta_0,...,\beta_J)^{\top}, \ g(t;\boldsymbol{\beta}_0)=\sum_{j=0}^J \beta_j \cdot I(t=j), \ m(t;\boldsymbol{\beta}_0)=\begin{bmatrix} I(t=0) \\ I(t=1)\\ \vdots \\ I(t=J) \end{bmatrix}, \ H_0=\mathbb{E}\left[m(T;\boldsymbol{\beta}_0)m(T;\boldsymbol{\beta}_0)^\top\right]. \end{align*} Then \begin{align*} \varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)=&\mathbb{E}[Y|T,X]-g(T;\boldsymbol{\beta}_0)\\ =&\sum_{j=0}^J \mathbb{E}[Y^*(j)|X] \cdot I(t=j)-\sum_{j=0}^J \beta_j \cdot I(T=j)\\ =&\sum_{j=0}^J\left( \mathbb{E}[Y^*(j)|X]- \beta_j\right) \cdot I(T=j) \end{align*} and \begin{align*} \pi_0(T,X)=\sum_{j=0}^J\frac{I(T=j)}{\mathbb{P}(T=j|\boldsymbol{X})}\cdot p_j, \ where \ p_j=\mathbb{P}(T=j). \end{align*} Then we have \begin{align*} H_0^{-1}=\mathbb{E}\left[m(T;\boldsymbol{\beta}_0)m(T;\boldsymbol{\beta}_0)^\top\right]^{-1}=\begin{bmatrix} & p_0^{-1} & & & \\ & & p_1^{-1} & & \\ & & \cdots & & \\ & & & & p_J^{-1} \end{bmatrix}, \end{align*} and \begin{align} &\pi_0 (T, \boldsymbol{X})m(T;\boldsymbol{\beta}_0)\left\{ Y-\mathbb{E}[Y| \boldsymbol{X},T]\right\} \notag\\ =&\left\{\sum_{j=0}^J\frac{I(T=j)}{\mathbb{P}(T=j|\boldsymbol{X})}\cdot p_j\right\} \cdot \begin{bmatrix} I(T=0) \\ I(T=1)\\ \vdots \\I(T=J) \end{bmatrix}\cdot\bigg\{Y-\sum_{j=0}^J I(T=j)\cdot \mathbb{E}[Y^*(j)|\boldsymbol{X}]\bigg\} \notag\\ =&\begin{bmatrix} I(T=0) \\ I(T=1)\\ \vdots \\I(T=J) \end{bmatrix} \left\{\sum_{j=0}^J\frac{I(T=j)}{\mathbb{P}(T=j|\boldsymbol{X})}\cdot p_j\cdot Y^*(j)-\sum_{j=0}^J\frac{I(T=j)}{\mathbb{P}(T=j|\boldsymbol{X})}\cdot p_j\cdot \mathbb{E}[Y^*(j)|\boldsymbol{X}] \right\}\notag\\ =& \begin{bmatrix} \frac{I(T=0)}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot p_0\cdot \left\{Y^*(0)-\mathbb{E}[Y^*(0)|\boldsymbol{X}]\right\} \\[2mm] \frac{I(T=1)}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p_1\cdot \left\{Y^*(1)-\mathbb{E}[Y^*(1)|\boldsymbol{X}]\right\}\\ \vdots \\[2mm] \frac{I(T=J)}{\mathbb{P}(T=J|\boldsymbol{X})}\cdot p_J\cdot \left\{Y^*(j)-\mathbb{E}[Y^*(j)|\boldsymbol{X}]\right\} \end{bmatrix} \end{align} and \begin{align*} & \varepsilon (T,\boldsymbol{X}; \boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\\ =&\left\{\sum_{j=0}^J\left( \mathbb{E}[Y^*(j)|X]- \beta_j\right) \cdot I(T=j)\right\}\left\{\sum_{j=0}^J\frac{I(T=j)}{\mathbb{P}(T=j|\boldsymbol{X})}\cdot p_j\right\}\begin{bmatrix} I(T=0) \\ I(T=1)\\ \vdots \\ I(T=J) \end{bmatrix} \\ =&\begin{bmatrix} \frac{I(T=0)}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot p_0\cdot \left\{ \mathbb{E}[Y^*(0)|X]- \beta_0\right\} \\[2mm] \frac{I(T=1)}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p_1\cdot \left\{ \mathbb{E}[Y^*(1)|X]- \beta_1\right\}\\ \vdots \\ \frac{I(T=J)}{\mathbb{P}(T=J|\boldsymbol{X})}\cdot p_J\cdot \left\{ \mathbb{E}[Y^*(j)|X]- \beta_J\right\} \end{bmatrix} \end{align*} and \begin{align} \mathbb{E}\left[ \varepsilon (T,\boldsymbol{X}; \boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)| \boldsymbol{X}\right]= \begin{bmatrix} p_0\cdot \left\{ \mathbb{E}[Y^*(0)|X]- \beta_0\right\} \\[2mm] p_1\cdot \left\{ \mathbb{E}[Y^*(1)|X]- \beta_1\right\}\\ \vdots \\ p_J\cdot \left\{ \mathbb{E}[Y^*(j)|X]- \beta_J\right\} \end{bmatrix}. \end{align} From Theorem 3.1, the efficient influence function of $\boldsymbol{\beta }_0=(\beta_0,...,\beta_J)$ is given by \begin{align*} &H_0^{-1}\left\{\pi_0 (T, \boldsymbol{X})m(T;\boldsymbol{\beta}_0)\left\{ Y-\mathbb{E}[Y| \boldsymbol{X},T]\right\}+ \mathbb{E}\left[ \varepsilon (T,\boldsymbol{X}; \boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)| \boldsymbol{X}\right]\right\}\\ =&\begin{bmatrix} \frac{I(T=0)}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot \left\{Y^*(0)-\mathbb{E}[Y^*(0)|\boldsymbol{X}]\right\} + \mathbb{E}[Y^*(0)|X]- \beta_0\\[2mm] \frac{I(T=1)}{\mathbb{P}(T=1|\boldsymbol{X})} \cdot \left\{Y^*(1)-\mathbb{E}[Y^*(1)|\boldsymbol{X}]\right\} +\mathbb{E}[Y^*(1)|X]- \beta_1\\ \vdots \\[2mm] \frac{I(T=J)}{\mathbb{P}(T=J|\boldsymbol{X})} \cdot \left\{Y^*(j)-\mathbb{E}[Y^*(j)|\boldsymbol{X}]\right\}+ \mathbb{E}[Y^*(j)|X]- \beta_J \end{bmatrix}, \end{align*} which is the same as the efficient influence function developed in Corollary 1 of cattaneo2010efficient.

Particular Case III: Quantile Treatment Effects

In this section, we show that when $T\in\{0,1\}$ is a binary treatment variable, $L(v)=v(\tau-I(v\leq 0))$ is the check function with $\tau\in(0,1)$, and $g(t;\boldsymbol{\beta}_0)=\beta_0\cdot (1-t)+\beta_1\cdot t$, where $\boldsymbol{\beta}_0=(\beta_0,\beta_1)$, our general efficiency bound derived in Theorem 3.1 reduces to the efficiency bound of quantile treatment effects given in Firpo2007Efficient. In accordance with our identification condition, $\beta_0$ and $\beta_1$ are identified by minimizing the following loss function

align*[align* omitted — 136 chars of source]

The solutions are $\beta_0=\inf\{q: \mathbb{P}(Y^*(0)\leq q)\geq \tau\}$ and $\beta_1=\inf\{q: \mathbb{P}(Y^*(1)\leq q)\geq \tau\}$, which are the $\tau^{th}$ quantiles of potential outcomes.

corLet $T\in\{0,1\}$, $f_{Y^*(1)}$ and $f_{Y^*(0)}$ be the probability densities of the potential outcomes $Y^*(1)$ and $Y^*(0)$ respectively, $g(t;\boldsymbol{\beta}_0)=\beta_0\cdot (1-t)+\beta_1\cdot t$, $L(v)=v(\tau-I(v\leq 0))$, and the conditions in Theorem 3.1 hold, then the efficient influence function of $\boldsymbol{\beta}_0$ given by Theorem 3.1 reduces to \begin{align*} S_{eff}(Y,T,\boldsymbol{X};\boldsymbol{\beta}_0) =\begin{bmatrix} \frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot \left\{\frac{\tau-I(Y^*(0)\leq \beta_0) }{f_{Y^*(0)}(\beta_0)}\right\}- \left(\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}-1\right)\cdot \mathbb{E}\left[\frac{\tau-I(Y^*(0)\leq \beta_0)}{f_{Y^*(0)}(\beta_0)} \big|\boldsymbol{X}\right]\\[2mm] \frac{T}{\mathbb{P}(T=1|\boldsymbol{X})} \cdot \left\{\frac{\tau-I(Y^*(1)\leq \beta_1)}{f_{Y^*(1)}(\beta_1)} \right\}- \left(\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}-1\right)\cdot \mathbb{E}\left[\frac{\tau-I(Y^*(1)\leq \beta_1)}{f_{Y^*(1)}(\beta_1)} \big|\boldsymbol{X}\right] \end{bmatrix}, \end{align*} which is the same as the efficient influence function given in Firpo2007Efficient.
proofUsing our notation, we have \begin{align*} &\boldsymbol{\beta}_0=(\beta_0,\beta_1)^{\top},\ g(t;\boldsymbol{\beta}_0)=\beta_0\cdot (1-t)+\beta_1 \cdot t , \ m(t;\boldsymbol{\beta}_0)=\begin{bmatrix} 1-t \\ t \end{bmatrix},\\ &L(v)=v(\tau-I(v\leq 0)), \ L'(v)=\tau-I(v\leq 0)\ a.s., \\ &\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)=T\cdot \mathbb{E}[\tau-I(Y^*(1)\leq \beta_1)|\boldsymbol{X}]+(1-T)\cdot \mathbb{E}[\tau-I(Y^*(0)\leq \beta_0)|\boldsymbol{X}], \\ &\pi_0(T,\boldsymbol{X})=\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p +\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q\ , \ p=\mathbb{P}(T=1), \ q=\mathbb{P}(T=0). \end{align*} Direct computation yields \begin{align*} &\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)L'(Y-g(T;\boldsymbol{\beta}_0))\\ =&\left\{\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p +\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q\right\}\cdot \begin{bmatrix} 1-T \\ T \end{bmatrix}\cdot \bigg\{\tau-I(Y\leq \beta_0\cdot (1-T)+\beta_1\cdot T )\bigg\}\\ =& \begin{bmatrix} \frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q\cdot \left\{\tau-I(Y^*(0)\leq \beta_0) \right\} \\[2mm] \frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p\cdot \left\{\tau-I(Y^*(1)\leq \beta_1) \right\} \end{bmatrix} \end{align*} and \begin{align*} \pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)=\begin{bmatrix} \frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q\cdot \mathbb{E}\left[\tau-I(Y^*(0)\leq \beta_0)|\boldsymbol{X} \right] \\[2mm] \frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p\cdot \mathbb{E}\left[\tau-I(Y^*(1)\leq \beta_1) |\boldsymbol{X}\right] \end{bmatrix} \end{align*} and \begin{align*} \mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)|\boldsymbol{X}\right]=\begin{bmatrix} q\cdot \mathbb{E}\left[\tau-I(Y^*(0)\leq \beta_0)|\boldsymbol{X} \right] \\[2mm] p\cdot \mathbb{E}\left[\tau-I(Y^*(1)\leq \beta_1) |\boldsymbol{X}\right] \end{bmatrix} \end{align*} and \begin{align*} H_0=&\nabla_{\boldsymbol{\beta}}\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta})L'(Y-g(T;\boldsymbol{\beta}))\right] =\begin{bmatrix} -q\cdot f_{Y^*(0)}(\beta_0) & 0 \\ 0 &-p\cdot f_{Y^*(1)}(\beta_1) \end{bmatrix}. \end{align*} Therefore, by Theorem 3.1, the efficient influence function of $\boldsymbol{\beta}_0$ is \begin{align*} &S_{eff}(Y,T,\boldsymbol{X};\boldsymbol{\beta}_0)\\ =&H_0^{-1}\cdot \bigg\{ \pi_0 (T, \boldsymbol{X})m(T;\boldsymbol{\beta }_{0})L'(Y-g(T;\boldsymbol{\beta}_0))-\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta }_{0})\varepsilon (T,\boldsymbol{X}; \boldsymbol{\beta }_{0})\\ & \qquad \qquad+\mathbb{E}\left[ \varepsilon (T,\boldsymbol{X}; \boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta }_{0})| \boldsymbol{X}\right] \bigg\}\\ =&\begin{bmatrix} q^{-1}\cdot \frac{1}{f_{Y^*(0)}(\beta_0)} & 0 \\ 0 &p^{-1}\cdot \frac{1}{f_{Y^*(1)}(\beta_1)} \end{bmatrix}\\ &\times \begin{bmatrix} \frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot q\cdot \left\{\tau-I(Y^*(0)\leq \beta_0) \right\}-q\cdot \left(\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}-1\right)\cdot \mathbb{E}\left[\tau-I(Y^*(0)\leq \beta_0) |\boldsymbol{X}\right]\\[2mm] \frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}\cdot p\cdot \left\{\tau-I(Y^*(1)\leq \beta_1) \right\}-p\cdot \left(\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}-1\right)\cdot \mathbb{E}\left[\tau-I(Y^*(1)\leq \beta_1) |\boldsymbol{X}\right] \end{bmatrix}\\ =&\begin{bmatrix} \frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}\cdot \left\{\frac{\tau-I(Y^*(0)\leq \beta_0) }{f_{Y^*(0)}(\beta_0)}\right\}- \left(\frac{1-T}{\mathbb{P}(T=0|\boldsymbol{X})}-1\right)\cdot \mathbb{E}\left[\frac{\tau-I(Y^*(0)\leq \beta_0)}{f_{Y^*(0)}(\beta_0)} \bigg|\boldsymbol{X}\right]\\[2mm] \frac{T}{\mathbb{P}(T=1|\boldsymbol{X})} \cdot \left\{\frac{\tau-I(Y^*(1)\leq \beta_1)}{f_{Y^*(1)}(\beta_1)} \right\}- \left(\frac{T}{\mathbb{P}(T=1|\boldsymbol{X})}-1\right)\cdot \mathbb{E}\left[\frac{\tau-I(Y^*(1)\leq \beta_1)}{f_{Y^*(1)}(\beta_1)} \bigg|\boldsymbol{X}\right] \end{bmatrix}, \end{align*} which coincides with efficiency bound derived in Firpo2007Efficient.

Convergence Rate of Estimated Stabilized Weights

In this section, we establish the convergence rate of estimated stabilized weights $\hat{\pi}_K(T,\boldsymbol{X})$. Let $G_{K_1\times K_2}^*$, $\Lambda_{K_1\times K_2}^*$ and $\pi_K^*(t,\boldsymbol{x})$ be the theoretical counterparts of $\hat{G}_{K_1\times K_2}$, $\hat{\Lambda}_{K_1\times K_2}$ and $\hat{\pi}_K(t,\boldsymbol{x})$ respectively:

align*[align* omitted — 466 chars of source]

As discussed in Appendix A.3, we assume the sieve basises $u_{K_1}(T)$ and $v_{K_2}(\boldsymbol{X})$ are orthonormalized, i.e.,

align[align omitted — 213 chars of source]

Let

align*[align* omitted — 211 chars of source]

We also recall the following property satisfied by $\pi_0(T,\boldsymbol{X})$: for any integrable functions $u(t)$ and $v(\boldsymbol{X})$,

equation[equation omitted — 161 chars of source]

Lemma $\ref{lemma_pi^*}$

The first lemma states that $\pi_K^*(t,\boldsymbol{x})$ is arbitrarily close to the true stabilized weights $\pi_0(t,\boldsymbol{x})$.

lemmaUnder Assumption (ref)-(ref), we have $$ \sup_{(t,\boldsymbol{x}) \in \mathcal{T} \times \mathcal{X}}|\pi_0(t,\boldsymbol{x}) - \pi^*_K(t,\boldsymbol{x}) | = O\left(K^{-\alpha}\zeta(K)\right),$$ and $$ \mathbb{E}\left[|\pi_0(T,\boldsymbol{X}) - \pi^*_K(T,\boldsymbol{X}) |^2\right] = O\left(K^{-2\alpha}\right),$$ and $$\frac{1}{N}\sum_{i=1}^N|\pi_0(T_i,\boldsymbol{X}_i) - \pi^*_K(T_i,\boldsymbol{X}_i) |^2=O_p\left(K^{-2\alpha}\right).$$
proofBy Assumption (ref), $\pi_0(t,\boldsymbol{x})\in [\eta_1, \eta_2], \ \forall (t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}$ and $(\rho')^{-1}$ is strictly decreasing. Define $$ \overline{\gamma} := \sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times\mathcal{X}} (\rho')^{-1}\left(\pi_0(t,\boldsymbol{x})\right) \leq (\rho')^{-1}(\eta_1) ~~ \text{and} ~~ \underline{\gamma} := \inf_{(t,\boldsymbol{x}) \in \mathcal{T}\times\mathcal{X}} (\rho')^{-1}\left(\pi_0(t,\boldsymbol{x})\right) \geq(\rho')^{-1}(\eta_2),$$ which are two finite constants. By Assumptions (ref), there exist a constant $C> 0$ and a $K_1\times K_2$ matrix $\Lambda_{K_1\times K_2} \in \mathbb{R}^{K_1\times K_2}$ such that \begin{align*} \sup_{(t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X}}\left|(\rho')^{-1}\left(\pi_0(t,\boldsymbol{x})\right) - u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right| < CK^{-\alpha}, \end{align*} which implies \begin{align} u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) \in & \left((\rho')^{-1}\left(\pi_0(t,\boldsymbol{x})\right) - CK^{-\alpha}, (\rho')^{-1}\left(\pi_0(t,\boldsymbol{x})\right) + CK^{-\alpha} \right) \\ \subset & \left[\gamma - CK^{-\alpha}, \overline{\gamma} + CK^{-\alpha}\right], \ \forall (t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}, \notag \end{align} and \begin{align*} &\rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) +CK^{-\alpha}\right) - \rho'(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{X}))\\ < &\pi_0(t,\boldsymbol{x}) - \rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right) \\ < & \rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})-CK^{-\alpha}\right) - \rho'(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})) \ , \forall (t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X}. \end{align*} Let $\Gamma_1:= [\underline{\gamma} -1, \overline{\gamma} + 1]$, by Mean Value Theorem, for large enough $K$, there exist \begin{align*} \xi_1(t,\boldsymbol{x}) \in & \left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}), u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) + CK^{-\alpha}\right)\\ \subset &\left[\gamma -CK^{-\alpha}, \overline{\gamma} + 2CK^{-\alpha}\right]\subset\Gamma_1 \ , \\ \xi_2(t,\boldsymbol{x}) \in& \left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) -CK^{-\alpha}, u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)\\ \subset & \left[\gamma - 2CK^{-\alpha}, \overline{\gamma} + CK^{-\alpha}\right]\subset \Gamma_1, \end{align*} such that \begin{align*} \rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) + CK^{-\alpha}\right) - \rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) \right)= \rho”(\xi_1(t,x))CK^{-\alpha}\geq -a_1CK^{-\alpha} \end{align*} and \begin{align*} \rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) - CK^{-\alpha}\right) - \rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) \right)= -\rho”(\xi_2(t,\boldsymbol{x}))CK^{-\alpha} \leq a_2CK^{-\alpha}, \end{align*} where $ -a_1 :=\inf_{\gamma \in \Gamma_1} \rho''(\gamma)$ and $a_2 := \sup_{\gamma \in \Gamma_1}\left( -\rho''(\gamma)\right)$. Let $a := \max\{a_1, a_2\}$, we have \begin{align} \sup_{(t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X}}\left| {\pi_0(t,\boldsymbol{x})} - \rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) \right)\right| < aCK^{-\alpha}. \end{align} For some fixed $C_2 > 0$ (to be chosen later), define $$ \Upsilon_{K_1\times K_2} := \left\{\Lambda \in \mathbb{R}^{K_1\times K_2}: \|\Lambda - \Lambda_{K_1\times K_2}\| \leq C_2K^{-\alpha}\right\}.$$ For sufficiently large $K_1$ and $K_2$, we have that $\forall \Lambda \in \Upsilon_{K_1\times K_2}$, $\forall (t,\boldsymbol{x}) \in \mathcal{T}\times\mathcal{X}$, \begin{align*} &\left|u_{K_1}(t)^{\top}\Lambda v_{K_2}(\boldsymbol{x}) - u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right|\\ \leq &\|\Lambda - \Lambda_{K_1\times K_2}\| \cdot \sup_{\boldsymbol{x}\in\mathcal{X}}\|v_{K_2}(\boldsymbol{x}) \| \cdot \sup_{t\in\mathcal{T}}\|u_{K_1}(t)\| \leq C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2) . \end{align*} Then in light of (ref) and Assumption (ref), for large enough $K_1$ and $K_2$, $\forall\Lambda \in \Upsilon_{K_1\times K_2}$ and $\forall (t,\boldsymbol{x}) \in \mathcal{T}\times\mathcal{X}$, we can deduce that \begin{align} & u_{K_1}(t)^{\top}\Lambda v_{K_2}(\boldsymbol{x}) \in \left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) - C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2), \right. \\ &\left. \quad \quad \quad \quad \quad \quad \quad \quad \quad u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})+ C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2)\right) \notag \\ & \subset \left[\gamma - CK^{-\alpha} - C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2), \right. \notag\\ & \left. \overline{\gamma} + CK^{-\alpha} + C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2)\right]\subset \Gamma_1.\notag \end{align} By definition \begin{align*} G_{K_1\times K_2}^*\left(\Lambda\right)= \mathbb{E}\left[\rho\left(u_{K_1}(T)^{\top}\Lambda v_{K_2}(\boldsymbol{X})\right)\right] -\mathbb{E}[u_{K_1}(T)]^{\top}\Lambda \mathbb{E}[v_{K_2}(\boldsymbol{X})], \end{align*} is a strictly concave function of $\Lambda$. By (ref), the formula $\tr(AB)=\tr(BA)$ for matrices $A$ and $B$, the facts $\mathbb{E}\left[v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^\top\right]=I_{K_2\times K_2}$ and $\mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^\top\right]=I_{K_1\times K_1}$, we can deduce that{ \begin{align} &\|\nabla {G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\|^2\notag\\ =&\left\|\mathbb{E}\left[\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] -\mathbb{E}[u_{K_1}(T)]\mathbb{E}[v_{K_2}(\boldsymbol{X})]^{\top}\right\|^2 \notag\\ =&\left\|\mathbb{E}\left[\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] -\mathbb{E}[\pi_0(T,\boldsymbol{X})u_{K_1}(T)v_{K_2}(\boldsymbol{X})]^{\top}\right\|^2\quad (by (ref)) \notag\\ =&\left\|\mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] \right\|^2 \notag\\ =&\tr\Bigg\{\mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] \notag \\ &\quad \times \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}v_{K_2}(\boldsymbol{X})u_{K_1}(T)^\top\right]\Bigg\} \notag\\ =&\tr\Bigg\{\mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] \cdot \mathbb{E}\left[u_{K_2}(\boldsymbol{X})u_{K_2}(\boldsymbol{X})^\top\right]\notag \\ &\quad \times \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}v_{K_2}(\boldsymbol{X})u_{K_1}(T)^\top\right]\cdot \mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^\top\right]\Bigg\} \notag\\ =&\mathbb{E}\Bigg[\tr\Bigg\{u_{K_1}(T)^\top\cdot \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] \cdot \mathbb{E}\left[u_{K_2}(\boldsymbol{X})u_{K_2}(\boldsymbol{X})^\top\right]\notag \\ &\qquad \times \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}v_{K_2}(\boldsymbol{X})u_{K_1}(T)^\top\right]\cdot u_{K_1}(T)\Bigg\} \Bigg] \notag\\ =&\mathbb{E}\Bigg[\pi_0(T,\boldsymbol{X}) \cdot u_{K_1}(T)^\top\cdot \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] \cdot u_{K_2}(\boldsymbol{X})\notag \\ &\qquad\times \cdot u_{K_2}(\boldsymbol{X})^\top \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}v_{K_2}(\boldsymbol{X})u_{K_1}(T)^\top\right]\cdot u_{K_1}(T) \Bigg] \quad (by (ref)) \notag\\ =&\mathbb{E}\Bigg[\bigg|{\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}u_{K_1}(T) \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] {\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}v_{K_2}(\boldsymbol{X}) \bigg|^2\Bigg]. \end{align}} Note that the term in the last expression { $${\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}u_{K_1}(T)\cdot \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] {\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}v_{K_2}(\boldsymbol{X})$$} is the $L^2(dF_{T,X})$-projection of $\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}$ on the space spanned by {$\{{\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}u_{K_1}(T)$, ${\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}v_{K_2}(\boldsymbol{X})\}$}, which implies that{ \begin{align} &\mathbb{E}\Bigg[\bigg|{\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}u_{K_1}(T) \mathbb{E}\left[\sqrt{\pi_0(T,\boldsymbol{X})}\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}u_{K_1}(T)v_{K_2}(\boldsymbol{X})^\top\right] {\pi_0(T,\boldsymbol{X})}^{\frac{1}{4}}v_{K_2}(\boldsymbol{X}) \bigg|^2\Bigg]\notag\\ \leq &\mathbb{E}\left[\bigg|\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}\bigg|^2\right]. \end{align}} Now, with (ref), (ref), we can obtain that \begin{align} &\|\nabla {G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| \notag\\ \leq& \mathbb{E}\left[\bigg|\frac{\left\{\rho'\left(u_{K_1}(T)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)-\pi_0(T,\boldsymbol{X})\right\}}{\sqrt{\pi_0(T,\boldsymbol{X})}}\bigg|^2\right]^{\frac{1}{2}}\notag\\ \leq & \frac{1}{\sqrt{\eta_1}} \sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}\left|\rho'\left(u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2} v_{K_2}(\boldsymbol{x})\right)-\pi_0(t,\boldsymbol{x})\right| \qquad (\text{by Assumption (ref)}) \notag \\ \leq & \frac{aC}{\sqrt{\eta_1}}\cdot K^{-\alpha} \qquad (\text{by (ref)}. \end{align} Note that for any $\Lambda \in \partial \Upsilon_{K_1\times K_2}$, i.e. $\|\Lambda - \Lambda_{K_1\times K_2}\|= C_2K^{-\alpha}$, by Mean Value Theorem and the fact $\rho''(y)=-\rho'(y)$, we can deduce that{ \begin{align*} & \quad {G}^*_{K_1\times K_2}(\Lambda) - {G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\\ &= \sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} \frac{\partial}{\partial \lambda_i}{G}^*_{K_1\times K_2}(\lambda_1^K,\ldots,\lambda_{K_2}^K)\\ &\qquad + \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}\frac{1}{2}(\lambda_j - \lambda_j^K)^{\top} \frac{\partial^2}{\partial \lambda_i \partial \lambda_l}{G}^*_{K_1\times K_2}(\bar{\lambda}^K_1,\ldots,\bar{\lambda}^K_{K_2})(\lambda_l - \lambda_l^K) \notag\\ & \leq \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| \\ & +\frac{1}{2} \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} \mathbb{E}\left[\rho\left(u_{K_1}^\top(T)\bar{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right)u_{K_1}(T)u_{K_1}(T)^\top v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right](\lambda_l - \lambda_l^K) \notag\\ & = \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| \\ & \quad -\frac{1}{2} \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} \mathbb{E}\left[\frac{\rho'\left(u_{K_1}^\top(T)\bar{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right)}{\pi_0(T,\boldsymbol{X})}\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^\top v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right](\lambda_l - \lambda_l^K) \notag\\ & \leq \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| \\ &\quad-\frac{a_3}{2\eta_2} \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} \mathbb{E}\left[\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^\top v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right](\lambda_l - \lambda_l^K) \notag \quad \text{(by $a_3=\inf_{y\in\Gamma_1}\{\rho'(y)\}$)}\\ & = \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| \\ &\quad-\frac{a_3}{2\eta_2} \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} \mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^\top\right]\mathbb{E}\left[ v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right](\lambda_l - \lambda_l^K) \quad \text{(by (ref))}\notag\\ & = \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| -\frac{a_3}{2\eta_2} \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} \mathbb{E}\left[ v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right](\lambda_l - \lambda_l^K) \notag\\ & = \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| -\frac{a_3}{2\eta_2} \sum_{j=1}^{K_2}(\lambda_j - \lambda_j^K)^{\top} (\lambda_j - \lambda_j^K) \notag\\ &= \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\|\nabla{G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| - \frac {a_3}{2\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|^2 \notag\\ & = \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\left(\|\nabla {G}^*_{K_1\times K_2}(\Lambda_{K_1\times K_2})\| - \frac {a_3}{2\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\right)\\ &\leq \left\|\Lambda - \Lambda_{K_1\times K_2}\right\|\left(\frac{aC}{\sqrt{\eta_1}}K^{-\alpha}-\frac {a_3}{2\eta_2}\cdot C_2K^{-\alpha}\right), \quad \text{(by (ref))} \end{align*}} where $\bar{\Lambda}_{K_1\times K_2}=(\bar{\lambda}_1^K,...,\bar{\lambda}_{K_2}^K)$ lies on the line joining ${\Lambda}=(\lambda_1,...,\lambda_{K_2})$ and ${\Lambda}_{K_1\times K_2}=(\lambda_1^K,...,\lambda_{K_2}^K)$, which implies $u_{K_1}^\top(t)\bar{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\in \Gamma_1$ by (ref); $a_3=\inf_{y\in\Gamma_1}\{\rho'(y)\}>0$ is a finite positive constant; the fourth and fifth equalities follow from $\mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^\top\right]=I_{K_1\times K_1}$ and $\mathbb{E}\left[v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^\top\right]=I_{K_2\times K_2}$ respectively. Therefore, by choosing $$C_2>\frac{2\eta_2}{a_3}\cdot \frac{aC}{{\sqrt{\eta_1}}},$$ we can obtain the following conclusion: \begin{align} G^*_{K_1\times K_2}(\Lambda_{K_1\times K_2}) > G^*_{K_1\times K_2}(\Lambda) \ , \forall \Lambda \in \partial\Upsilon_{K_1\times K_2}\ . \end{align} Since $G^*_{K_1\times K_2}$ is continuous, (ref) implies that there exists a local maximum of $G^*_{K_1\times K_2}$ in the interior of $\Upsilon_{K_1\times K_2}$. Note that $G^*_{K_1\times K_2}$ is strictly concave with a unique global maximum point $\Lambda^*_{K_1\times K_2}$, therefore we can claim that \begin{align} \Lambda^*_{K_1\times K_2} \in \Upsilon_{K_1\times K_2}^\circ, \ \text{i.e.}\ \|\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\|=O(K^{-\alpha}) \ . \end{align} By Mean Value Theorem, (ref) and (ref), we can deduce that \begin{align} \notag &|\rho'\left(u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right) - \rho'\left(u_{K_1}(t)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)| \\ = &|\rho”(\xi^*(t,\boldsymbol{x}))|\left|u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})-u_{K_1}(t)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right| \notag\\ \leq & -\rho”(\xi^*(t,\boldsymbol{x})) \times \|\Lambda_{K_1\times K_2}-\Lambda_{K_1\times K_2}^* \| \times \sup_{t\in\mathcal{T}}\|u_{K_1}(t)\| \times \sup_{\boldsymbol{x}\in \mathcal{X}}\|v_{K_2}(\boldsymbol{x})\| \notag \\ \leq & a_2C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2) \ ,\notag \end{align} where $ a_2 = \sup_{\gamma \in \Gamma_1} \{ -\rho''(\gamma) \} <\infty$ is a finite positive constant, and $\xi^*(t,\boldsymbol{x})$ lies between the point $u_{K_1}(t)^{\top} \Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})$ and $u_{K_1}(t)^{\top} \Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})$ (note (ref) implies $\xi^*(t,\boldsymbol{x}) \in \Gamma_1 $ for all $(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}$ and large enough $K$). Therefore, using the triangle inequality, and Assumption (ref), we can have \begin{align} &\sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}}\left|{\pi_0(t,\boldsymbol{x})} - \pi^*_K(t,\boldsymbol{x})\right| \notag \\ \leq & \sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}}\left|{\pi_0(t,\boldsymbol{x})} - \rho'\left(u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)\right|\notag\\ &+ \sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}}\left|\rho'\left(u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right) - \rho'\left(u_{K_1}(t)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)\right| \notag \\ \leq & aCK^{-\alpha} + a_2C_2K^{-\alpha}\zeta_1(K_1)\zeta_2(K_2)=O\left(K^{-\alpha}\zeta(K)\right), \notag \end{align} where $\zeta(K)=\zeta_1(K_1)\zeta_2(K_2)$. \\ We next prove $\mathbb{E}\left[\left|{\pi_0(T,\boldsymbol{X})} - \pi^*_K(T,\boldsymbol{X})\right|^2\right]=O\left(K^{-2\alpha}\right)$. By Assumption (ref), we can deduce that { \begin{align*} &\mathbb{E}\left[\left|{\pi_0(T,\boldsymbol{X})} - \pi^*_K(T,\boldsymbol{X})\right|^2\right]\\ \leq &2\cdot \mathbb{E}\left[\left|{\pi_0(T,\boldsymbol{X})} - \rho'\left(u_{K_1}(T)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right)\right|^2\right]+2\cdot \mathbb{E}\left[\left|\rho'\left(u_{K_1}(T)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right) - \rho'\left(u_{K_1}(T)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{X})\right)\right|^2\right]\\ \leq & 2\cdot \sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}}\left|{\pi_0(t,\boldsymbol{x})} - \rho'\left(u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)\right|^2+2\sup_{\gamma\in \Gamma_1}|\rho”(\gamma)|^2\cdot \mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]\\ \leq & O(K^{-2\alpha})+ O(1)\cdot \mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]. \end{align*}} We next compute the order of $\mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]$. Note that $\mathbb{E}[u_{K_1}(T)u_{K_1}(T)^\top]=I_{K_1\times K_1}$, $\mathbb{E}[v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^\top]=I_{K_2\times K_2}$, (ref), (ref) and Assumption (ref), we can deduce that { \begin{align} &\mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]\notag\\ =& \mathbb{E}\left[ u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^{\top}\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top u_{K_1}(T)\right] \notag\\ =& \mathbb{E}\left[\frac{1}{\pi_0(T,\boldsymbol{X})}\pi_0(T,\boldsymbol{X}) u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^{\top}\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top u_{K_1}(T)\right] \notag\\ \leq& \frac{1}{\eta_1}\cdot \mathbb{E}\left[\pi_0(T,\boldsymbol{X}) u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^{\top}\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top u_{K_1}(T)\right] \notag\\ =&\frac{1}{\eta_1}\cdot \int_{\mathcal{T}} u_{K_1}^\top(t)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}\mathbb{E}\left[v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^{\top}\right]\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top u_{K_1}(t)dF_T(t)\quad (\text{by (ref)}) \notag\\ =& \frac{1}{\eta_1}\cdot \int_{\mathcal{T}} u_{K_1}^\top(t)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}\cdot\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top u_{K_1}(t)dF_T(t) \notag\\ =&\frac{1}{\eta_1}\cdot \int_{\mathcal{T}} \tr\Bigg(\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}\cdot\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top u_{K_1}(t)u_{K_1}^\top(t)\Bigg)dF_T(t) \notag\\ =&\frac{1}{\eta_1}\cdot \tr\Bigg(\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}\cdot\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}^\top\Bigg) \notag\\ \leq & \frac{1}{\eta_1}\cdot\|\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\|^2= O(K^{-2\alpha}). \quad (\text{by (ref)}) \end{align}} Therefore, we can obtain \begin{align*} \mathbb{E}\left[\left|{\pi_0(T,\boldsymbol{X})} - \pi^*_K(T,\boldsymbol{X})\right|^2\right]=O\left(K^{-2\alpha}\right). \end{align*} We finally prove $N^{-1}\sum_{i=1}^N\left|{\pi_0(T_i,\boldsymbol{X}_i)} - \pi^*_K(T_i,\boldsymbol{X}_i)\right|^2=O_p\left(K^{-2\alpha}\right)$. Note that by (ref), we can have{ \begin{align*} &\mathbb{E}\left[\left\{\frac{1}{N}\sum_{i=1}^N \left|u_{K_1}^\top(T_i)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}_i) \right|^2-\mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]\right\}^2\right]\\ \leq &\frac{1}{N}\cdot \mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^4\right]\\ \leq& \frac{1}{N}\cdot \mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]\cdot \sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}\left|u_{K_1}^\top(t)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x}) \right|^2\\ \leq & \frac{1}{N}\cdot O(K^{-2\alpha})\cdot \zeta_1(K_1)^2\zeta_2(K_2)^2\cdot \left\|\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\|^2\leq \frac{1}{N}\cdot \zeta(K)^2\cdot O( K^{-4\alpha}), \end{align*}} then in light of Chebyshev's inequality and Assumption (ref), we have \begin{align} &\frac{1}{N}\sum_{i=1}^N \left|u_{K_1}^\top(T_i)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}_i) \right|^2-\mathbb{E}\left[\left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]\notag \\ =&O_p\left(\frac{\zeta(K)}{\sqrt{N}}K^{-2\alpha}\right)=o_p\left(K^{-2\alpha}\right). \end{align} With (ref), (ref), (ref), and Assumption (ref), we can deduce that \begin{align*} &\frac{1}{N}\sum_{i=1}^N\left|{\pi_0(T_i,\boldsymbol{X}_i)} - \pi^*_K(T_i,\boldsymbol{X}_i)\right|^2\\ \leq &\frac{2}{N}\sum_{i=1}^N \left|{\pi_0(T_i,\boldsymbol{X}_i)} - \rho'\left(u_{K_1}(T_i)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)\right|^2\\ &+\frac{2}{N}\sum_{i=1}^N \left|\rho'\left(u_{K_1}(T_i)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right) - \rho'\left(u_{K_1}(T_i)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{X}_i)\right)\right|^2 \\ \leq &2\sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}}\left|{\pi_0(t,\boldsymbol{x})} - \rho'\left(u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)\right|^2\\ &+\sup_{\gamma\in \Gamma_1}|\rho”(\gamma)|^2\cdot \frac{2}{N}\sum_{i=1}^N \left|u_{K_1}^\top(T_i)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}_i) \right|^2 \\ \leq &2\sup_{(t,\boldsymbol{x}) \in \mathcal{T}\times \mathcal{X}}\left|{\pi_0(t,\boldsymbol{x})} - \rho'\left(u_{K_1}(t)\Lambda_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)\right|^2\\ &+2\cdot \sup_{\gamma\in \Gamma_1}|\rho”(\gamma)|^2\cdot \mathbb{E}\left[ \left|u_{K_1}^\top(T)\left\{\Lambda^*_{K_1\times K_2}-\Lambda_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}) \right|^2\right]+o_p\left(K^{-2\alpha}\right) \\ =& O(K^{-2\alpha})+ O(K^{-2\alpha})+o_p\left(K^{-2\alpha}\right) = O_p\left(K^{-2\alpha}\right). \quad (\text{by (ref)}) \end{align*}

Lemma $\ref{lemma_pi^hat}$

lemmaUnder Assumption (ref)-(ref), we have $$ \left\|\hat{\Lambda}_{K_1\times K_2}- \Lambda_{K_1\times K_2}^*\right\| = O_p\left(\sqrt{\frac {K} {N}} \right)\ .$$
proofDefine $$ \hat{S}_N := \frac {1} {N} \sum_{i=1}^N \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \pi_0(T_i,\boldsymbol{X}_i)u_{K_1}(T_i)u_{K_1}(T_i)^{\top}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X}_i)v_{K_2,l}(\boldsymbol{X}_i),$$ where $\lambda_j$ and $\lambda^*_j$ are the $j$-th column of $\Lambda$ and $\Lambda_{K_1\times K_2}^*$ respectively. Since $\hat{S}_N$ is symmetric, using (ref) and the facts that $\mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^{\top }\right]=I_{K_1\times K_1}$ and $\mathbb{E}\left[v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^{\top}\right]=I_{K_2\times K_2}$, we can have \begin{align*} \mathbb{E}\left[\hat{S}_N\right]=&\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \mathbb{E}\left[\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right](\lambda_l-\lambda_l^*)\\ =&\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \mathbb{E}\left[u_{K_1}(T)u_{K_1}(T)^{\top}\right]\mathbb{E}[v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})](\lambda_l-\lambda_l^*)\\ =&\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top}(\lambda_j-\lambda_j^*)=\left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|. \end{align*} Then we can further deduce that \begin{align*} &\mathbb{E}\left[\left|\hat{S}_N - \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\| \right|^2\right]\notag\\ = & \mathbb{E}[\hat{S}_N^2] - 2\mathbb{E}[\hat{S}_N]\left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\| + \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2\notag \\ = &\frac{N}{N^2}\cdot \mathbb{E}\left[\left(\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right)^2\right]\\ &+2\cdot \frac{C_N^2}{N^2}\cdot \mathbb{E}\left[\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right]^2 \notag \\ &-\left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2 \notag \\ = &\frac{1}{N}\mathbb{E}\left[\left(\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right)^2\right]\\ &+\frac{N(N-1)}{N^2}\cdot \mathbb{E}\left[\hat{S}_N\right]^2 -\left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2 \notag \\ =&\frac{1}{N}\mathbb{E}\left[\left(\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right)^2\right]\\ & \quad -\frac{1}{N}\left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2\\ < & \frac{1}{N}\mathbb{E}\left[\left(\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\right)^2\right]. \end{align*} In light of the fact that $$ 0 \leq y^{\top} \left\{\pi_0(t,\boldsymbol{x})u_{K_1}(t)u_{K_1}(t)^{\top} \right\}y \leq \eta_2 \zeta_1(K_1)^2 y^{\top}y \ , \ \forall y\in\mathbb{R}^{K_1},\ \forall (t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X} \ ,$$ we can deduce that \begin{align*} &\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j-\lambda_j^*)^{\top} \left\{ \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}\right\}(\lambda_l-\lambda_l^*)v_{K_2,j}(\boldsymbol{X})v_{K_2,l}(\boldsymbol{X})\\ =&\left[\sum_{j=1}^{K_2}v_{K_2,j}(\boldsymbol{X})(\lambda_j-\lambda_j^*)^{\top}\right] \cdot \left\{ \pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}\right\}\cdot \left[\sum_{l=1}^{K_2}(\lambda_l-\lambda_l^*)v_{K_2,l}(\boldsymbol{X})\right] \\ \leq & \eta_2 \cdot \|u_{K_1}(T)\|^2\cdot \left\|\sum_{i=1}^{K_2}(\lambda_i-\lambda_i^*)^{\top}v_{K_2,i}(\boldsymbol{X}) \right\|^2\\ \leq & \eta_2 \cdot \|u_{K_1}(T)\|^2\cdot \left(\sum_{i=1}^{K_2}\|\lambda_i-\lambda_i^*\|^2\right)\left(\sum_{i=1}^{K_2}v_{K_2,i}(\boldsymbol{X})^2\right)\\ =& \eta_2 \cdot \|u_{K_1}(T)\|^2\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^2\|v_{K_2}(\boldsymbol{X})\|^2 . \end{align*} Therefore, we can obtain that{ \begin{align} &\mathbb{E}\left[\left|\hat{S}_N - \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\| \right|^2\right] \notag\\ \leq & \frac{1}{N} \eta^2_2 \cdot \mathbb{E}\left[ \|u_{K_1}(T)\|^4\cdot \|v_{K_2}(\boldsymbol{X})\|^4 \right]\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4 \notag\\ \leq& \frac{1}{N} \eta^2_2 \cdot \zeta_1(K_1)^2\cdot \zeta_2(K_2)^2\cdot \mathbb{E}\left[ \|u_{K_1}(T)\|^2\cdot \|v_{K_2}(\boldsymbol{X})\|^2 \right]\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4 \notag\\ =& \frac{1}{N} \eta^2_2 \cdot \zeta_1(K_1)^2\cdot \zeta_2(K_2)^2\cdot \mathbb{E}\left[\frac{1}{\pi_0(T,\boldsymbol{X})}\cdot\pi_0(T,\boldsymbol{X}) \|u_{K_1}(T)\|^2\cdot \|v_{K_2}(\boldsymbol{X})\|^2 \right]\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4 \notag\\ \leq& \frac{1}{N} \frac{\eta^2_2}{\eta_1} \cdot \zeta_1(K_1)^2\cdot \zeta_2(K_2)^2\cdot \mathbb{E}\left[\pi_0(T,\boldsymbol{X}) \|u_{K_1}(T)\|^2\cdot \|v_{K_2}(\boldsymbol{X})\|^2 \right]\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4\notag \quad (by Assumption\ (ref))\\ =&\frac{1}{N} \frac{\eta^2_2}{\eta_1} \cdot \zeta_1(K_1)^2\cdot \zeta_2(K_2)^2\cdot \mathbb{E}\left[ \|u_{K_1}(T)\|^2 \right]\cdot\mathbb{E}\left[ \|v_{K_2}(\boldsymbol{X})\|^2 \right]\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4\notag \quad (by (ref))\\ =&\frac{1}{N} \frac{\eta^2_2}{\eta_1} \cdot \zeta_1(K_1)^2\cdot \zeta_2(K_2)^2\cdot K_1\cdot K_2\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4 \notag (since\ \mathbb{E}[ \|u_{K_1}(T)\|^2]=K_1\ and\ \mathbb{E}[ \|v_{K_2}(\boldsymbol{X})\|^2]=K_2)\\ =&\frac{1}{N} \frac{\eta^2_2}{\eta_1} \cdot \zeta(K)^2\cdot K\cdot \left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4\ . \quad (since\ \zeta(K)=\zeta_1(K_1)\zeta_2(K_2)\ and \ K=K_1\cdot K_2) \end{align}} Considering the event set $$E_{N}:= \left\{\hat{S}_N > \frac{1}{2}\left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2,\ \Lambda\neq \Lambda_{K_1\times K_2}^*\right\}\ ,$$ by Chebyshev's inequality, (ref), and Assumption (ref) we can get \begin{align*} &\mathbb{P}\left(\left|\hat{S}_N- \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2\right|\geq \frac{1}{2} \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2,\ \Lambda\neq \Lambda_{K_1\times K_2}^*\right)\\ &\leq \frac{4\mathbb{E}\left[\left|\hat{S}_N - \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\| \right|^2\right]}{\left\|\Lambda-\Lambda^*_{K_1\times K_2}\right\|^4} \\ & \leq\frac{4}{N} \frac{\eta^2_2}{\eta_1}\cdot \zeta(K)^2\cdot K \leq O\left( \frac{\zeta(K)^2K}{N}\right)=o(1), \end{align*} which implies that for any $\epsilon>0$, there exists $N_0(\epsilon)\in\mathbb{N}$ such that $N>N_0(\epsilon)$ large enough \begin{align} \mathbb{P}\left((E_{N})^c\right)<\mathbb{P}\left(\left|\hat{S}_N- \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2\right|\geq \frac{1}{2} \left\|{\Lambda}-{\Lambda}^*_{K_1\times K_2}\right\|^2,\ \Lambda\neq \Lambda_{K_1\times K_2}^*\right) <\frac{\epsilon}{2} \ . \end{align} Note that \begin{align*} &\frac{\partial}{\partial {\lambda_j}}\hat{G}_{K_1\times K_2}(\lambda_1,\ldots,\lambda_{K_2})\\ =&\frac{1}{N}\sum_{i=1}^N\rho'\left(u_{K_1}^{\top}(T_i)\Lambda v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)-\frac{1}{N^2}\sum_{i=1}^N\sum_{l=1}^Nu_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_l)\\ =& \frac{1}{N}\sum_{i=1}^N \left\{\rho'\left(u_{K_1}^{\top}(T_i)\Lambda v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})]u_{K_1}(T_i)\right\} \\ &-\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)\left\{\frac{1}{N}\sum_{l=1}^Nv_{K_2,j}(\boldsymbol{X}_l)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})] \right\} \end{align*} and \begin{align*} &\frac{\partial}{\partial {\lambda_j}}{G}^*_{K_1\times K_2}(\lambda_1,\ldots,\lambda_{K_2})=\mathbb{E}\left[\rho'\left(u_{K_1}^{\top}(T_i)\Lambda v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)\right]-\mathbb{E}[u_{K_1}(T)]\cdot \mathbb{E}[v_{K_2,j}(\boldsymbol{X})] . \end{align*} Since $\Lambda_{K_1\times K_2}^*$ is the unique maximizer of $G^*_{K_1\times K_2}(\cdot)$, then for each $j\in \{1,\ldots,K_2\}$, \begin{align*} &\frac{\partial}{\partial {\lambda_j}}{G}^*_{K_1\times K_2}(\lambda_1^*,\ldots,\lambda^*_{K_2})\\ =&\mathbb{E}\left[\rho'\left(u_{K_1}^{\top}(T)\Lambda_{K_1\times K_2}^* v_{K_2}(\boldsymbol{X})\right)u_{K_1}(T)v_{K_2,j}(\boldsymbol{X})\right]-\mathbb{E}[u_{K_1}(T)]\mathbb{E}[v_{K_2,j}(\boldsymbol{X})]=0. \end{align*} Therefore, for large enough $K$, we can deduce that{ \begin{align} &\mathbb{E}\left[\|\nabla\hat{G}_{K_1\times K_2}(\Lambda^*_{K_1\times K_2})\|^2\right] = \sum_{j=1}^{K_2}\mathbb{E}\left[\left\|\frac{\partial}{\partial {\lambda_j}}\hat{G}_{K_1\times K_2}(\lambda_1^*,\ldots,\lambda^*_{K_2})\right\|^2\right] \\ \leq & 2 \sum_{j=1}^{K_2}\mathbb{E}\left[\left\|\frac{1}{N}\sum_{i=1}^N \left\{\rho'\left(u_{K_1}^{\top}(T_i)\Lambda^*_{K_1\times K_2} v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})]u_{K_1}(T_i)\right\} \right\|^2\right]\notag \\ &+ 2\sum_{j=1}^{K_2}\mathbb{E}\left[\left\|\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)\left\{\frac{1}{N}\sum_{l=1}^Nv_{K_2,j}(\boldsymbol{X}_l)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})] \right\}\right\|^2\right] \notag \\ =& \frac{2}{N^2} \sum_{j=1}^{K_2}\sum_{i=1}^N \mathbb{E}\left[\left\|\rho'\left(u_{K_1}^{\top}(T_i)\Lambda^*_{K_1\times K_2} v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})]u_{K_1}(T_i)\right\|^2\right] \notag \\ &+ 2\sum_{j=1}^{K_2}\mathbb{E}\left[\left\|\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)\left\{\frac{1}{N}\sum_{l=1}^Nv_{K_2,j}(\boldsymbol{X}_l)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})] \right\}\right\|^2\right] \notag \\ \leq & \frac{2}{N^2} \sum_{j=1}^{K_2}\sum_{i=1}^N \mathbb{E}\left[\left\|\rho'\left(u_{K_1}^{\top}(T_i)\Lambda^*_{K_1\times K_2} v_{K_2}(\boldsymbol{X}_i)\right)u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})]u_{K_1}(T_i)\right\|^2\right] \notag \\ &+ 2\sum_{j=1}^{K_2}\mathbb{E}\left[\left|\frac{1}{N}\sum_{l=1}^Nv_{K_2,j}(\boldsymbol{X}_l)-\mathbb{E}[v_{K_2,j}(\boldsymbol{X})]\right|^2\right]\cdot \mathbb{E}\left[\left\|\frac{1}{N}\sum_{i=1}^Nu_{K_1}(T_i)\right\|^2\right] \notag \\ \leq & \frac{4}{N} \sum_{j=1}^{K_2} \left\{ \mathbb{E}\left[\left\|\rho'\left(u_{K_1}^{\top}(T)\Lambda^*_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)u_{K_1}(T)v_{K_2,j}(\boldsymbol{X})\right\|^2\right]\right. +\mathbb{E}[v_{K_2,j}(\boldsymbol{X})^2]\mathbb{E}\left[\|u_{K_1}(T)\|^2\right] \bigg\}\notag \\ &+ \frac{2}{N}\sum_{j=1}^{K_2}\mathbb{E}\left[v_{K_2,j}(\boldsymbol{X})^2\right]\cdot \mathbb{E}\left[\left\|u_{K_1}(T)\right\|^2\right] \notag \\ = & \frac{4}{N} \sum_{j=1}^{K_2} \left\{ \mathbb{E}\left[\frac{\left|\rho'\left(u_{K_1}^{\top}(T)\Lambda^*_{K_1\times K_2} v_{K_2}(\boldsymbol{X})\right)\right|^2}{\pi_0(T,\bold{X})}\cdot \pi_0(T,\bold{X})\cdot \left\|u_{K_1}(T)v_{K_2,j}(\boldsymbol{X})\right\|^2\right]\right. +\mathbb{E}[v_{K_2,j}(\boldsymbol{X})^2]\mathbb{E}\left[\|u_{K_1}(T)\|^2\right] \bigg\}\notag \\ &+ \frac{2}{N}\sum_{j=1}^{K_2}\mathbb{E}\left[v_{K_2,j}(\boldsymbol{X})^2\right]\cdot \mathbb{E}\left[\left\|u_{K_1}(T)\right\|^2\right] \notag \\ \leq & \frac{4}{N} \sum_{j=1}^{K_2} \left\{ \frac{\left(\sup_{\gamma\in\Gamma_1}\rho'(\gamma)\right)^2}{\eta_1}\cdot \mathbb{E}\left[ \pi_0(T,\bold{X})\cdot \left\|u_{K_1}(T)v_{K_2,j}(\boldsymbol{X})\right\|^2\right]\right. +\mathbb{E}[v_{K_2,j}(\boldsymbol{X})^2]\mathbb{E}\left[\|u_{K_1}(T)\|^2\right] \bigg\}\notag \\ &+ \frac{2}{N}\sum_{j=1}^{K_2}\mathbb{E}\left[v_{K_2,j}(\boldsymbol{X})^2\right]\cdot \mathbb{E}\left[\left\|u_{K_1}(T)\right\|^2\right] \notag\\ = & \frac{4}{N} \sum_{j=1}^{K_2} \left\{ \frac{\left(\sup_{\gamma\in\Gamma_1}\rho'(\gamma)\right)^2}{\eta_1}\cdot \mathbb{E}\left[ v_{K_2,j}(\boldsymbol{X})^2\right] \mathbb{E}\left[ \left\|u_{K_1}(T)\right\|^2\right]\right. +\mathbb{E}[v_{K_2,j}(\boldsymbol{X})^2]\mathbb{E}\left[\|u_{K_1}(T)\|^2\right] \bigg\}\notag \\ &+ \frac{2}{N}\sum_{j=1}^{K_2}\mathbb{E}\left[v_{K_2,j}(\boldsymbol{X})^2\right]\cdot \mathbb{E}\left[\left\|u_{K_1}(T)\right\|^2\right] \notag\\ \leq & \frac{1}{N}\left\{ \frac{4}{\eta_1}\left(\sup_{\gamma\in\Gamma_1}\rho'(\gamma)\right)^2+4+2\right\}\cdot \mathbb{E}\left[\left\|u_{K_1}(T)\right\|^2\right]\sum_{j=1}^{K_2}\mathbb{E}\left[v_{K_2,j}(\boldsymbol{X})^2\right] \notag \\ =&\frac{1}{N}\left\{ \frac{4}{\eta_1}\left(\sup_{\gamma\in\Gamma_1}\rho'(\gamma)\right)^2+6\right\} K_1K_2= C_4^2 \frac{K}{N}, \notag \end{align}} where the last inequality follows by Assumption (ref) and $C_4:=\sqrt{ \frac{4}{\eta_1}\left(\sup_{\gamma\in\Gamma_1}\rho'(\gamma)\right)^2+6}$ is a finite universal constant.\\ Let $\epsilon>0$, fix $\dps C_5(\epsilon) > 0$ (to be chosen later) and define \begin{align} \notag \hat{\Upsilon}_{K_1\times K_2}(\epsilon):= \left\{\Lambda \in \mathbb{R}^{K_1\times K_2}: \|\Lambda - \Lambda_{K_1\times K_2}^*\| \leq C_5(\epsilon) C_4 \sqrt{\frac {K} {N}} \right\}. \end{align} For $\forall \Lambda \in \hat{\Upsilon}_{K_1\times K_2}(\epsilon), \forall (t,\boldsymbol{x}) \in\mathcal{T}\times \mathcal{X}$, we can have \begin{align*} & \left|u_{K_1}(t)^{\top}\Lambda v_{K_2}(\boldsymbol{x}) - u_{K_1}(t)^{\top}\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right|\\ \leq& \|\Lambda - \Lambda_{K_1\times K_2}^* \|\sup_{t\in\mathcal{T}}\|u_{K_1}(t)\| \sup_{\boldsymbol{x}\in\mathcal{X}}\|v_{K_2}(\boldsymbol{x})\| \leq C_5(\epsilon)C_4\sqrt{\frac {K} {N}} \zeta_1(K_1)\zeta_2(K_2) , \end{align*} thus for large enough $N$, in accordance with Assumption (ref) and (ref), we have \begin{align} &u_{K_1}(t)^{\top}\Lambda v_{K_2}(\boldsymbol{x}) \in \Bigg[ u_{K_1}(t)^{\top}\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) - C_5(\epsilon)C_4 \zeta_1(K_1)\zeta_2(K_2) \sqrt{\frac {K} {N}}, \notag\\ & \quad \quad \quad \quad \quad \quad \quad \quad \quad u_{K_1}(t)^{\top}\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) + C_5(\epsilon)C_4 \zeta_1(K_1)\zeta_2(K_2) \sqrt{\frac {K} {N}} \Bigg] \notag \\ &\quad \quad \quad \quad \quad \quad \quad \subset \Bigg[ \underline{\gamma} - CK^{-\alpha} - C_5(\epsilon)C_4 \zeta_1(K_1)\zeta_2(K_2) \sqrt{\frac {K} {N}}, \notag \\ & \quad \quad \overline{\gamma} + CK^{-\alpha} + C_5(\epsilon)C_4 \zeta_1(K_1)\zeta_2(K_2) \sqrt{\frac {K} {N}}\Bigg] \subset \Gamma_2(\epsilon)\ , \end{align} where $\Gamma_2(\epsilon):= \left[\underline{\gamma}-1-C_5(\epsilon), \overline{\gamma}+1+C_5(\epsilon)\right]$ is a compact set and independent of $(t,\boldsymbol{x})$. \\ For any $\Lambda \in \partial \hat{\Upsilon}_{K_1\times K_2}(\epsilon)$, there exists $\bar{\Lambda}$ on the line joining $\Lambda$ and $\Lambda_{K_1\times K_2}^*$ such that \begin{align*} \hat{G}_{K_1\times K_2}(\Lambda) =& \hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*) + \sum_{j=1}^{K_2}(\lambda_j - \lambda_j^*)^{\top} \frac{\partial}{\partial \lambda_i}\hat{G}_{K_1\times K_2}(\lambda_1^*,\ldots,\lambda_{K_2}^*)\\ &+ \frac{1}{2}\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^*)^{\top} \frac{\partial^2}{\partial \lambda_i \partial \lambda_l}\hat{G}_{K_1\times K_2}(\bar{\lambda}_1,\ldots,\bar{\lambda}_{K_2})(\lambda_l - \lambda_l^*)\ , \end{align*} where $\bar{\lambda}_j$ denotes the $j$-th column of $\bar{\Lambda}$. For the second order term in above equality, note that $u_{K_1}^{\top}(t)\bar{\Lambda}v_{K_2}(\boldsymbol{x})\in\Gamma_2(\epsilon)$ for all $(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}$, we can further deduce that \begin{align} &\sum_{l=1}^{K_2}\sum_{j=1}^{K_2}(\lambda_j - \lambda_j^*)^{\top} \frac{\partial^2}{\partial \lambda_i \partial \lambda_l}\hat{G}_{K_1\times K_2}(\bar{\lambda}_1,\ldots,\bar{\lambda}_{K_2})(\lambda_l - \lambda_l^*) \\ = &\frac {1} {N} \sum_{i=1}^N\sum_{j=1}^{K_2}\sum_{l=1}^{K_2}(\lambda_j-\lambda^*_{j})^{\top}u_{K_1}(T_i)\rho\left(u_{K_1}^{\top}(T_i)\bar{\Lambda}v_{K_2}(\boldsymbol{X}_i)\right)(\lambda_l - \lambda_l^*)^{\top}u_{K_1}(T_i)v_{K_2,j}(\boldsymbol{X}_i)v_{K_2,l}(\boldsymbol{X}_i) \notag\\ \leq & - \frac {\bar{b}(\epsilon)} {N} \sum_{i=1}^N\sum_{j=1}^{K_2}\sum_{l=1}^{K_2} (\lambda_j-\lambda^*_{j})^{\top}u_{K_1}(T_i) u_{K_1}(T_i)^{\top}(\lambda_l - \lambda_l^*)v_{K_2,j}(\boldsymbol{X}_i)v_{K_2,l}(\boldsymbol{X}_i)\notag\\ = & - \frac {\bar{b}(\epsilon)} {N} \sum_{i=1}^N\sum_{j=1}^{K_2}\sum_{l=1}^{K_2} \frac{1}{\pi_0(T_i,\boldsymbol{X}_i)} (\lambda_j-\lambda^*_{j})^{\top} \pi_0(T_i,\boldsymbol{X}_i)u_{K_1}(T_i) u_{K_1}(T_i)^{\top}(\lambda_l - \lambda_l^*)v_{K_2,j}(\boldsymbol{X}_i)v_{K_2,l}(\boldsymbol{X}_i)\notag\\ \leq & - \frac {\bar{b}(\epsilon)} {N \eta_2} \sum_{i=1}^N\sum_{j=1}^{K_2}\sum_{l=1}^{K_2} (\lambda_j-\lambda^*_{j})^{\top}\pi_0(T_i,\boldsymbol{X}_i) u_{K_1}(T_i) u_{K_1}(T_i)^{\top}(\lambda_l - \lambda_l^*)v_{K_2,j}(\boldsymbol{X}_i)v_{K_2,l}(\boldsymbol{X}_i)\notag\\ =& -\frac{\bar{b}(\epsilon)}{\eta_2}\hat{S}_N\ , \notag \end{align} where $-\bar{b}(\epsilon):= \sup_{\gamma \in \Gamma_2(\epsilon)} \rho''(\gamma)<\infty$ for each fixed $\epsilon$. Therefore, on the event $E_{N}$ and for large enough $N$, we can deduce that for any $\Lambda\in \partial\hat{\Upsilon}_{K_1\times K_2}(\epsilon)$, \begin{align} \text{on the event}\ E_N: &\quad \ \hat{G}_{K_1\times K_2}(\Lambda) - \hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*)\\ &= \sum_{j=1}^{K_2}(\lambda_j - \lambda_j^*)^{\top} \frac{\partial}{\partial \lambda_i}\hat{G}_{K_1\times K_2}(\lambda_1^*,\ldots,\lambda_{K_2}^*) \notag\\ &\quad + \sum_{l=1}^{K_2}\sum_{j=1}^{K_2}\frac{1}{2}(\lambda_j - \lambda_j^*)^{\top} \frac{\partial^2}{\partial \lambda_i \partial \lambda_l}\hat{G}_{K_1\times K_2}(\bar{\lambda}_1,\ldots,\bar{\lambda}_{K_2})(\lambda_l - \lambda_l^*) \notag\\ & \leq \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\|\nabla\hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*)\| -\frac{\bar{b}(\epsilon)}{2\eta_2} \hat{S}_N \ \text{(by (ref))} \notag\\ &\leq \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\|\nabla\hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*)\| - \frac {\bar{b}(\epsilon)}{4\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|^2 \notag\\ & = \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\left(\|\nabla\hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*)\| - \frac {\bar{b}(\epsilon)}{4\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\right) \ , \notag \end{align} where the second inequality follows from definition of the event $E_{N}$. \\ Note that for sufficiently large $N$, by Chebyshev's inequality and (ref) we have \begin{align} &\mathbb{P}\left\{\|\nabla\hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*)\| \geq \frac {\bar{b}(\epsilon)}{4\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\right\}\\ \leq&\frac{16\eta_2^2}{\bar{b}(\epsilon)^2}\cdot \frac{\mathbb{E}\left[\left\|\nabla\hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*)\right\|^2\right]}{ \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|^2} \leq\frac{16\eta_2^2}{\bar{b}(\epsilon)^2C_5^2(\epsilon)} \leq \frac{\epsilon}{2}\ ,\notag \end{align} where the last inequality holds by choosing $$C_5(\epsilon) \geq \sqrt{\frac{32\eta_2^2}{\bar{b}(\epsilon)^2\epsilon}}\ .$$ Therefore, for sufficiently large $N$, by (ref) and (ref) we can derive \begin{align} &\mathbb{P}\left( (E_{N})^c \ \text{or} \ \|\nabla\hat{G}_{K_1\times K_2}(\Lambda^*_{K_1\times K_2})\| \geq \frac {\bar{b}(\epsilon)}{2\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\right) \leq \frac {\epsilon} {2} + \frac {\epsilon} {2} = \epsilon \notag \\ \Rightarrow &\mathbb{P}\left( E_{N} \ \text{and} \ \|\nabla\hat{G}_{K_1\times K_2}(\Lambda^*_{K_1\times K_2})\| < \frac {\bar{b}(\epsilon)}{2\eta_2} \left\|\Lambda - \Lambda_{K_1\times K_2}^*\right\|\right) > 1-\epsilon. \end{align} With (ref) and (ref), we can obtain that $$\mathbb{P}\left\{\hat{G}_{K_1\times K_2}(\Lambda) - \hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*) < 0 ,~~ \forall \Lambda \in \partial\hat{\Upsilon}_{K_1\times K_2}(\epsilon) \right\} \geq 1 - \epsilon\ .$$ Note that the event $\left\{ \hat{G}_{K_1\times K_2}(\Lambda_{K_1\times K_2}^*) > \hat{G}_{K_1\times K_2}(\Lambda) ,~ \forall \Lambda \in \partial\hat{\Upsilon}_{K_1\times K_2}(\epsilon) \right\}$ implies that there exists a local maximizer in the interior of $\hat{\Upsilon}_{K_1\times K_2}(\epsilon)$. Since $\hat{G}_{K_1\times K_2}(\cdot)$ is strictly concave and $\hat{\Lambda}_{K_1\times K_2} $ is the unique global maximizer of $\hat{G}_{K_1\times K_2}$, then \begin{align} \mathbb{P}\left(\hat{\Lambda}_{K_1\times K_2} \in \hat{\Upsilon}_{K_1\times K_2}(\epsilon)\right)>1-\epsilon , \end{align} i.e. $ \left\|\hat{\Lambda}_{K_1\times K_2}- \Lambda_{K_1\times K_2}^*\right\| = O_p\left(\sqrt{\frac {K} {N}} \right)$.

Corollary $\ref{cor:pi^-pi*}$

The next corollary states that $\hat{\pi}_K(t,\boldsymbol{x})$ is arbitrarily close to ${\pi}^*_K(t,\boldsymbol{x})$.

corUnder Assumption (ref)-(ref), we have $$\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\hat{\pi}_K(t,\boldsymbol{x})-{\pi}^*_K(t,\boldsymbol{x})|=O_p\left(\zeta(K)\sqrt{\frac{K}{N}}\right),$$ and $$\int_{\mathcal{T}\times \mathcal{X}}|\hat{\pi}_K(t,\boldsymbol{x})-{\pi}^*_K(t,\boldsymbol{x})|^2dF_{T,X}(t,\boldsymbol{x})=O_p\left(\frac{K}{N}\right),$$ and $$\frac{1}{N}\sum_{i=1}^N|\hat{\pi}_K(T_i,\boldsymbol{X}_i)-{\pi}^*_K(T_i,\boldsymbol{X}_i)|^2=O_p\left(\frac{K}{N}\right).$$
proofFrom the proof of Lemma (ref), we know the facts $\mathbb{P}\left(\hat{\Lambda}_{K_1\times K_2}\in \hat{\Upsilon}_{K_1\times K_2}(\epsilon)\right)>1-\epsilon$ and (ref). Then for any element $\tilde{\Lambda}_{K_1\times K_2}$ lying on the line joining $\hat{\Lambda}_{K_1\times K_2}$ and $\Lambda_{K_1\times K_2}^*$, we can have that $\mathbb{P}(u_{K_1}(t)^{\top}\tilde{\Lambda}_{K_1\times K_2} v_{K_2}(\boldsymbol{x})\in \Gamma_{2}(\epsilon)$ for all $(t,\boldsymbol{x})\in \mathcal{T}\times\mathcal{X})\geq 1-\epsilon$, which implies \begin{align} \sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\rho”(u_{K_1}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x}))|=O_p(1). \end{align} Using Mean Value Theorem, Lemma (ref), and (ref), we can obtain that \begin{align} \notag &\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\hat{\pi}_K(t,\boldsymbol{x})-{\pi}^*_K(t,\boldsymbol{x})|\\ =&\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\rho'\left(u_{K_1}(t)\hat{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right) - \rho'\left(u_{K_1}(t)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right)| \notag \\ \leq &\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\rho”(u_{K_1}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x}))|\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}\left|u_{K_1}(t)\hat{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x})-u_{K_1}(t)\Lambda^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})\right| \notag\\ \leq & O_p(1) \cdot \|\hat{\Lambda}_{K_1\times K_2}-\Lambda_{K_1\times K_2}^* \| \cdot \sup_{t\in\mathcal{T}}\|u_{K_1}(t)\| \cdot\sup_{\boldsymbol{x}\in \mathcal{X}}\|v_{K_2}(\boldsymbol{x})\| \notag \\ \leq & O_p(1)\cdot O_p\left(\sqrt{\frac {K} {N}} \right) \zeta_1(K_1)\cdot \zeta_2(K_2)=O_p\left(\zeta(K)\sqrt{\frac{K}{N}}\right).\notag \end{align} Note that by Mean Value Theorem and (ref), we can deduce that \begin{align*} &\int_{\mathcal{T}\times \mathcal{X}}|\hat{\pi}_K(t,\boldsymbol{x})-{\pi}^*_K(t,\boldsymbol{x})|^2dF_{T,X}(t,\boldsymbol{x})\\ \leq&\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\rho”(u_{K_1}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x}))|^2\int_{\mathcal{T}\times \mathcal{X}}\left|u_{K_1}(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right|^2dF_{T,X}(t,\boldsymbol{x})\\ \leq &O_p(1)\cdot \int_{\mathcal{T}\times \mathcal{X}}\left|u_{K_1}(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right|^2dF_{T,X}(t,\boldsymbol{x}) . \end{align*} We estimate $\int_{\mathcal{T}\times \mathcal{X}}\left|u_{K_1}(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right|^2dF_{T,X}(t,\boldsymbol{x})$. Note that $\mathbb{E}[u_{K_1}(T)[u_{K_1}(T)^\top]=I_{K_1\times K_1}$, $\mathbb{E}[v_{K_2}(\boldsymbol{X})v_{K_2}(\boldsymbol{X})^\top]=I_{K_2\times K_2}$, (ref) and Assumption (ref), we can deduce that{ \begin{align} &\int_{\mathcal{T}\times \mathcal{X}}\left|u_{K_1}(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right|^2dF_{T,X}(t,\boldsymbol{x}) \notag\\ \leq & \int_{\mathcal{T}\times \mathcal{X}}u_{K_1}^\top(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})v_{K_2}(\boldsymbol{x})^\top\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top u_{K_1}(t) dF_{T,X}(t,\boldsymbol{x}) \notag\\ =& \int_{\mathcal{T}\times \mathcal{X}}\frac{1}{\pi_0(t,\boldsymbol{x})}\pi_0(t,\boldsymbol{x}) u_{K_1}^\top(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})v_{K_2}(\boldsymbol{x})^\top\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top u_{K_1}(t) dF_{T,X}(t,\boldsymbol{x}) \notag\\ \leq & \frac{1}{\eta_1} \int_{\mathcal{T}\times \mathcal{X}}\pi_0(t,\boldsymbol{x})\cdot u_{K_1}^\top(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})v_{K_2}(\boldsymbol{x})^\top\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top u_{K_1}(t) dF_{T,X}(t,\boldsymbol{x})\notag \\ = & \frac{1}{\eta_1} \int_{\mathcal{T}} u_{K_1}^\top(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}\left(\int_{\mathcal{X}}v_{K_2}(\boldsymbol{x})v_{K_2}(\boldsymbol{x})^\top dF_X(x)\right)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top u_{K_1}(t) dF_{T}(t) \notag\\ = & \frac{1}{\eta_1} \int_{\mathcal{T}} u_{K_1}^\top(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top u_{K_1}(t) dF_{T}(t) \notag\\ = & \frac{1}{\eta_1} \tr\Bigg( \left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top \int_{\mathcal{T}}u_{K_1}(t)u_{K_1}^\top(t) dF_{T}(t)\Bigg) \notag\\ = & \frac{1}{\eta_1} \tr\Bigg( \left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}^\top \Bigg)\notag \\ =& \frac{1}{\eta_1}\cdot \left\|\hat{\Lambda}_{K_2\times K_2}-\Lambda^*_{K_1\times K_2}\right\|^2 =O_p\left(\frac{K}{N}\right). \end{align}} Then we obtain \begin{align*} \int_{\mathcal{T}\times \mathcal{X}}|\hat{\pi}_K(t,\boldsymbol{x})-{\pi}^*_K(t,\boldsymbol{x})|^2dF_{T,X}(t,\boldsymbol{x})=O_p\left(\frac{K}{N}\right). \end{align*} Similar to (ref), we have { \begin{align} &\frac{1}{N}\sum_{i=1}^N \left|u_{K_1}^\top(T_i)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}_i) \right|^2-\int_{\mathcal{T}\times \mathcal{X}}\left|u_{K_1}(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right|^2dF_{T,X}(t,\boldsymbol{x})\notag \\ =&O_p\left(\frac{\zeta(K)}{\sqrt{N}}\cdot \|\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\|^2\right)=O_p\left(\frac{\zeta(K)}{\sqrt{N}}\cdot \frac{K}{N}\right)=o_p\left(\frac{K}{N}\right). \end{align}} where the last equality holds in light of Assumption (ref). Hence, with (ref) and (ref), we have \begin{align*} &\frac{1}{N}\sum_{i=1}^N|\hat{\pi}_K(T_i,\boldsymbol{X}_i)-{\pi}^*_K(T_i,\boldsymbol{X}_i)|^2 \\ \leq&\sup_{(t,\boldsymbol{x})\in\mathcal{T}\times\mathcal{X}}|\rho”(u_{K_1}(t)\tilde{\Lambda}_{K_1\times K_2}v_{K_2}(\boldsymbol{x}))|^2\cdot \frac{1}{N}\sum_{i=1}^N\left|u_{K_1}(T_i)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{X}_i)\right|^2\\ \leq &O_p(1)\cdot \int_{\mathcal{T}\times \mathcal{X}}\left|u_{K_1}(t)\left\{\hat{\Lambda}_{K_1\times K_2}-\Lambda^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right|^2dF_{T,X}(t,\boldsymbol{x})+o_p\left(\frac{K}{N}\right) \\ \leq &O_p\left(\frac{K}{N}\right)+o_p\left(\frac{K}{N}\right)=O_p\left(\frac{K}{N}\right). \end{align*}

Efficient Estimation

Proof of Theorem 5.17

Since we have proved $\|\hat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0\|\xrightarrow{p}0$ in Theorem 10, we now begin to prove the asymptotic efficiency of $\hat{\boldsymbol{\beta}}$. By Assumption (ref), $\hat{\boldsymbol{\beta}}$ is a unique solution of the following equation:

align[align omitted — 208 chars of source]

with probability approaching to one. Note that $L'(\cdot)$ may be a non-differentiable function, e.g. $L'(v)=\tau-I(v\leq 0)$ in quantile regression, we cannot simply apply Mean Value Theorem on (ref) to obtain the expression for $\sqrt{N}(\hat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0)$. To solve this problem, we define

align*[align* omitted — 139 chars of source]

which is a differentiable function in $\boldsymbol{\beta}$ and by definition $f(\boldsymbol{\beta}_0)=0$. Using Mean Value Theorem, we can obtain that

align*[align* omitted — 191 chars of source]

where $\tilde{\boldsymbol{\beta}}$ lies on the line joining $\hat{\boldsymbol{\beta}}$ and $\boldsymbol{\beta}_0$. Because $\nabla_{\beta}f(\boldsymbol{\beta})$ is continuous in $\boldsymbol{\beta}$ at $\boldsymbol{\beta}_0$, and $\|\hat{\boldsymbol{\beta}}-{\boldsymbol{\beta}}_0\|\xrightarrow{p}0$, then we have

align*[align* omitted — 155 chars of source]

Define the empirical process:

align*[align* omitted — 293 chars of source]

and by Assumption (ref) we can have

align*[align* omitted — 733 chars of source]

By Assumption (ref) and (ref), Theorems 4 and 5 of andrews1994empirical, we can conclude that $\mu_N(\cdot)$ is stochastically equicontinuous, which implies $\mu_N(\hat{\boldsymbol{\beta}}) -\mu_N(\boldsymbol{\beta}_0)\xrightarrow{p}0$. We also note that $\mathbb{E}[\pi_0 (T,\boldsymbol{X})L'(Y -g(T;\boldsymbol{\beta}_0 ))m(T;\boldsymbol{\beta}_0)]=0$, then

align[align omitted — 235 chars of source]

We next claim the following important Lemma (ref), and leave its proof to Section (ref).

lemmaUnder Assumption (ref)-(ref), we have \begin{align} \frac{1}{\sqrt{N}}\sum_{i=1}^N\hat{\pi}_K(T_i,\boldsymbol{X}_i)m(T_i;{ \boldsymbol{\beta}}_0)L'\left\{Y_i-g\left(T_i;{\boldsymbol{\beta}} _0\right)\right\}=\frac{1}{\sqrt{N}}\sum_{i=1}^N\psi(Y_i,T_i,\boldsymbol{X} _i;\boldsymbol{\beta}_0) +o_p(1), \end{align} where \begin{align*} &\psi (Y,T,\boldsymbol{X};\boldsymbol{\beta }_{0}):= \pi_0 (T, \boldsymbol{X})m(T;\boldsymbol{\beta }_{0})L'(Y-g(T;\boldsymbol{\beta}_0))-\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta }_{0})\varepsilon (T,\boldsymbol{X}; \boldsymbol{\beta }_{0}) \\ &\qquad \qquad\qquad+\mathbb{E}\left[ \varepsilon (T,\boldsymbol{X}; \boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta }_{0})| \boldsymbol{X}\right] +\mathbb{E}\left[ \varepsilon (T,\boldsymbol{X}; \boldsymbol{\beta }_{0})\pi_0 (T,\boldsymbol{X})m(T;\boldsymbol{\beta }_{0})| T\right] \ , \end{align*} and $\varepsilon (T,\boldsymbol{X};\boldsymbol{\beta }_{0}):= \mathbb{E} [L'(Y-g(T;\boldsymbol{\beta}_0))|T,\boldsymbol{X}]$.

Lemma (ref) is the most important step for establishing the efficiency of our proposed estimator. A key technique in proving Lemma (ref) is a use of a weighted least square projection of $L'(Y-g(T;\boldsymbol{\beta}_0))$ onto the space linearly spanned by the approximation basis $\{u_{K_{1}}(T),v_{K_{2}}(\boldsymbol{X})\}$.

Combining (ref) and Lemma (ref), we can obtain the asymptotic expression for $\sqrt{N}(\hat{ \boldsymbol{\beta }}-\boldsymbol{\beta }_{0})$:

align*[align* omitted — 304 chars of source]

which leads to our Theorem 5.17.

Proof of Lemma (ref)

Before proving Lemma (ref), we prepare some preliminary notation and results that will be used later. Since $\hat{\Lambda}_{K_1\times K_2}$ is a unique maximizer of the concave function $\hat{G}_{K_1\times K_2}$, then

align*[align* omitted — 259 chars of source]

Using Mean Value Theorem, we can have

align[align omitted — 573 chars of source]

where $\tilde{\Lambda}_{K_1\times K_2}$ lies on the line joining from $\hat{\Lambda}_{K_1\times K_2}$ to ${\Lambda}^*_{K_1\times K_2}$. We define the following notation:

align[align omitted — 229 chars of source]

and

align[align omitted — 408 chars of source]

In light of (ref) we have $$\left\|A^*_{K_1\times K_2}\right\|= O_p\left(\sqrt{\frac{K}{N}}\right). $$ From (ref), $A^*_{K_1\times K_2}$ can also be written as

align[align omitted — 344 chars of source]

We now start to prove Lemma (ref). We decompose $\frac{1}{\sqrt{N}}\sum_{i=1}^N\hat{\pi}_K(T_i,\boldsymbol{X}_i)\left\{Y_i-g\left(T_i;\boldsymbol{\beta}_0\right)\right\}m(T_i;\boldsymbol{\beta}_0)$ as follows:

align[align omitted — 4,048 chars of source]

where $\hat{A}_{K_1\times K_2}$ and $A_{K_1\times K_2}^*$ are defined in (ref) and (ref)}. We show that the terms (ref)-(ref) are all of $o_p(1)$, while the term (ref) is asymptotically normal.\\

For term (ref):

Denoting (ref) by $W_K$ and applying Mean Value Theorem twice, we can obtain

align*[align* omitted — 2,008 chars of source]

where $\tilde{A}_{K_1\times K_2}$ is defined in (ref), and{

align*[align* omitted — 1,352 chars of source]

} and $\xi_3(t,\boldsymbol{x})$ lies between $u_{K_1}(t)\tilde{\Lambda}_{K_1\times K_2}^{\top}v_{K_2}(\boldsymbol{x})$ and $u_{K_1}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x})$. \\

For the term $W_{1K}$, we denote its $k^{th}$ component by $W_{1K,k}$, $k=1,\ldots,p$, and we also let $m_k(T_i;\boldsymbol{\beta}_0)$ be the $k^{th}$ component of $m(T_i;\boldsymbol{\beta}_0)$, i.e.,

align*[align* omitted — 1,295 chars of source]

where

align*[align* omitted — 593 chars of source]

We compute the second moment of $U_{K_2\times K_1}(k)$ to get that

align*[align* omitted — 2,382 chars of source]

where $\dps a_3 := \sup_{\gamma \in \Gamma_1} |\rho''(\gamma)|^2 < +\infty$, the second inequality follows from this definition and the fact that $u_{K_1}^{\top}(t){\Lambda}^*_{K_1\times K_2}v_{K_2}(\boldsymbol{x}) \in \Gamma_1,\ \forall (t,\boldsymbol{x}) \in \mathcal{T}\times\mathcal{X}$ when $K$ is large enough; the third inequality follows from Assumption (ref); the forth inequality follows from Assumption (ref) and the facts

align[align omitted — 297 chars of source]

Then in light of Chebyshev's inequality, Lemma (ref) and Assumption (ref), we have $$|W_{1K,k}|\leq \|U_{K_2\times K_1}\|\|\hat{A}_{K_1\times K_2}\| =O_p(\sqrt{K})O_p\left(\sqrt{\frac{K}{N}}\right)=O_p\left(\sqrt{\frac{K^2}{N}}\right) ,$$ which implies

align[align omitted — 99 chars of source]

For the term $W_{3K}$, since $\xi_3(t,\boldsymbol{x})$ lies between $u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})$ and $u_{K_1}(t)^{\top}\tilde{\Lambda}_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})$, which implies $\xi_3(t,\boldsymbol{x})$ lies between $u_{K_1}(t)^{\top}\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})$ and $u_{K_1}(t)^{\top}\hat{\Lambda}_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})$. Then in light of (ref) and (ref), we have $\mathbb{P}\left(\xi_3(t,\boldsymbol{x})\in \Gamma_2(\epsilon),\ \forall (t,\boldsymbol{x})\in \mathcal{T}\times \mathcal{X}\right)>1-\epsilon$, therefore,

align[align omitted — 142 chars of source]

With (ref), (ref), the fact $\|\tilde{A}_{K_1\times K_2}\|\leq \|\hat{A}_{K_1\times K_2}\|$, Lemma (ref), and Assumption (ref), we can derive that{

align[align omitted — 1,638 chars of source]

}

For the term $W_{2K}$, we can deduce that {

align*[align* omitted — 1,801 chars of source]

} where the fourth inequality follows from the fact that{

align*[align* omitted — 1,212 chars of source]

} Therefore, we can obtain that

align*[align* omitted — 215 chars of source]

Finally, it follows that the term (ref) is of $o_p(1)$ in light of Assumption (ref).\\

For term (ref): Note that

align*[align* omitted — 1,020 chars of source]

where the last equality follows from Lemma (ref). Then by Chebyshev's inequality, we can claim that the term (ref) is of $O_p(\zeta(K)K^{-\alpha})$.\\

\noindentFor term (ref): By Lemma (ref) and Assumption (ref), we can deduce that

align*[align* omitted — 509 chars of source]

For term (ref): By Mean Value Theorem and the definition of $\hat{A}_{K_1\times K_2}$ in (ref), the term (ref) is exactly equal to zero. \\

For term (ref): We can telescope (ref) as follows:

align[align omitted — 1,614 chars of source]

For the term (ref), by Mean Value Theorem, {

align*[align* omitted — 403 chars of source]

} which is $O_p\left(\sqrt{\frac{K^2}{N}}\right)$ from (ref).\\

For the term (ref), we first compute the probability order of $\|A^*_{K_1\times K_2}-\hat{A}_{K_1\times K_2}\|$. Using (ref)}, the fact $\rho''(v)=-\rho'(v)$ and Mean Value Theorem, we have{

align[align omitted — 1,220 chars of source]

} For the term (ref), by (ref) we can write $\hat{A}_{K_1\times K_2}$ as $$ \hat{A}_{K_1\times K_2}=\mathbb{E}_{T,\boldsymbol{X}}\left[\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}\hat{A}_{K_1\times K_2}v_{K_2}(\boldsymbol{X})v^{\top}_{K_2}(\boldsymbol{X})\right], $$ where $\mathbb{E}_{T,\boldsymbol{X}}[\cdot]$ denotes taking expectation with respect to $(T,\boldsymbol{X})$. We telescope (ref) as follows:

align[align omitted — 1,016 chars of source]

For the term (ref), by Lemmas (ref) and (ref), we have that{

align*[align* omitted — 1,818 chars of source]

} For the term (ref), define the linear map $\mathcal{J}(\cdot):\mathbb{R}^{{K_1\times K_2}}\rightarrow \mathbb{R}$ by {

align*[align* omitted — 341 chars of source]

} then $ \eqref{A_diff_1_2}=\mathcal{J}(\hat{A}_{K_1\times K_2})$. For any fixed $M \in \mathbb{R}^{{K_1\times K_2}}$, by (ref) and $M=\mathbb{E}[\pi_0(T,\boldsymbol{X})u_{K_1}(T)u_{K_1}(T)^{\top}M$ $\cdot v_{K_2}(\boldsymbol{X})v^{\top}_{K_2}(\boldsymbol{X})]$, then we have

align*[align* omitted — 1,073 chars of source]

Using Chebyshev's inequality we have $$ |\mathcal{J}(M)| =\|M\| O_p\left(\zeta(K)\sqrt{\frac{K}{N}}\right)\ ,$$ then in light of Lemma (ref), $$\eqref{A_diff_1_2}=\mathcal{J}(\hat{A}_{K_1\times K_2})=\|\hat{A}_{K_1\times K_2}\| O_p\left(\zeta(K)\sqrt{\frac{K}{N}}\right)= O_p\left(\zeta(K)\frac{K}{N}\right)\ .$$ Therefore, $$\eqref{A_difference_1}=\eqref{A_diff_1_1}+\eqref{A_diff_1_2}= O_p\left(N^{-\frac{1}{2}}\zeta(K)\cdot K^{\frac{3}{2}-\alpha}\right)+O_p\left(\zeta(K)\frac{K}{N}\right) \ .$$

For the term (ref), we can deduce that

align*[align* omitted — 1,271 chars of source]

where the fourth inequality follows from (ref) and Lemma (ref). Now, we can obtain

align[align omitted — 392 chars of source]

Using (ref), Assumptions (ref) and (ref), for large enough $N$, we have {

align*[align* omitted — 1,050 chars of source]

} where the second inequality holds since by using the same argument of establishing (ref), we have $$\int_{\mathcal{T}\times\mathcal{X}}\left(u_{K_1}(t)\left\{\hat{A}_{K_1\times K_2}-A^*_{K_1\times K_2}\right\}v_{K_2}(\boldsymbol{x})\right)^2dF_{T,X}(t,\boldsymbol{x})= O( \|\hat{A}_{K_1\times K_2}-A_{K_1\times K_2}^*\|).$$ Therefore, we can obtain that

align*[align* omitted — 304 chars of source]

For term (ref): By the definition of $A^*_{K_1\times K_2}$ in (ref), we have{

align[align omitted — 1,454 chars of source]

} We shall show that both (ref) and (ref) are of $o_p(1)$. Noting $\rho''=-\rho'$, we can telescope (ref) as follows: {

align[align omitted — 1,597 chars of source]

} We shall show that (ref), (ref) and (ref) are all of $o_p(1)$. Note that second moment of (ref) is {

align*[align* omitted — 2,266 chars of source]

} where the third equality holds because $$\int_{\mathcal{T}}\int_{\mathcal{X}}\pi_0(t,\boldsymbol{x})\cdot m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\bigg[\frac{\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)}{\pi_0(t,\boldsymbol{x})}\bigg]u_{K_1}^{\top}(t) \left\{u_{K_1}(T)v^{\top}_{K_2}(\boldsymbol{X})\right\}v_{K_2}(\boldsymbol{x})dF_{X,T}(\boldsymbol{x},t) $$ is the weighted $L^2$-projection of $m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{x};\boldsymbol{\beta}_0)\bigg[\frac{\rho'\left(u_{K_1}^{\top}(t)\Lambda_{K_1\times K_2}^*v_{K_2}(\boldsymbol{x})\right)}{\pi_0(t,\boldsymbol{x})}\bigg]$ on the space linearly spanned by $\{u_{K_1}(t),v_{K_2}(\boldsymbol{x})\}$ with the weighted measure $\pi_0(t,\boldsymbol{x})dF_{T,X}(t,\boldsymbol{x})$. Similarly, we can also show (ref) and (ref) are of $o_p(1)$. Therefore, (ref) is of $o_p(1)$.\\

For the term (ref), since $\rho''(v)=-\rho'(v)$ and the fact $\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)\right]=0$, we telescope it as follows:{

align[align omitted — 2,306 chars of source]

} For the term (ref), since

align*[align* omitted — 475 chars of source]

and by Assumptions (ref), (ref), and (ref), we can deduce that

align*[align* omitted — 194 chars of source]

For the term (ref), noting the fact that $\mathbb{E}\left[\pi_0(T,\boldsymbol{X})m(T;\boldsymbol{\beta}_0)\varepsilon(T,\boldsymbol{X};\boldsymbol{\beta}_0)|\boldsymbol{X}\right]=\int_{\mathcal{T}}m(t;\boldsymbol{\beta}_0)\varepsilon(t,\boldsymbol{X};\boldsymbol{\beta}_0)$ $dF_T(t)$, we can rewrite (ref) as follows:

align[align omitted — 548 chars of source]

By computing the second moment of (ref), we can obtain that{

align*[align* omitted — 1,636 chars of source]

} where $T^*\sim F_T$, $\boldsymbol{X}^*\sim F_X$, and $T^*$ is independent of $\boldsymbol{X}^*$; the first inequality holds by Jensen's inequality; the last equality follows from Lemma (ref) and the fact that

align*[align* omitted — 242 chars of source]

is the $L^2$-projection of $m(T^*;\boldsymbol{\beta}_0)\varepsilon(T^*,\boldsymbol{X}^*;\boldsymbol{\beta}_0)$ on the space spanned by $\{u_{K_1}(T^*),v_{K_2}(\boldsymbol{X}^*)\}$, which implies {

align*[align* omitted — 384 chars of source]

} Thus (ref) is of $o_p(1)$ by Chebyshev's inequality. Similar argument can be applied to show that both (ref) and (ref) are of $o_p(1)$. Therefore, we can have that

align*[align* omitted — 130 chars of source]

Then, we can obtain that

align*[align* omitted — 95 chars of source]

Summing up all orders (ref)-(ref) and using Assumption (ref), we have{

align*[align* omitted — 353 chars of source]

}

Variance Estimation

masry1996multivariate studies the strong consistency of kernel regression estimation. The following conditions are imposed so that Theorem 6 of masry1996multivariate applies. Let $\boldsymbol{u}=(u_1,...,u_{r+1})$ and $K(\boldsymbol{u})=\prod_{j=1}^{r+1}k(u_j)$.\\

\noindentCondition 1. The kernel $K(\boldsymbol{u})\in L_1$ satisfies $\|\boldsymbol{u}\|K(\boldsymbol{u})\in L_{1}$ and $\|\boldsymbol{u}\|^2K(\boldsymbol{u})\in L_{1}$.

\noindentCondition 2. The density function $f_{Y,X,T}(y,\boldsymbol{x},t)$ is uniformly bounded away from zero and above, and also it is uniformly continuous on $\mathbb{R}^{r+1}$.

\noindentCondition 3. (a) The kernel $K(\cdot)$ is bounded with compact support; (b) let $H_j(\boldsymbol{u}):=\boldsymbol{u}^jK(\boldsymbol{u})$, and $|H_j(\boldsymbol{u})-H_j(\boldsymbol{v})|\leq C\|\boldsymbol{u}-\boldsymbol{v}\|$ for all $j$ with $0\leq j\leq 3$.

\noindentCondition 4. The functions $f_{Y,X,T}(y,\boldsymbol{x},t)$, $\mathbb{E}\left[\pi_0(T,\boldsymbol{X})L'(Y-g(T;\boldsymbol{\beta}))|T=t,\boldsymbol{X}=\boldsymbol{x}\right]$, $\mathbb{E}[\pi_0(T,\boldsymbol{X})L'(Y-g(T;\boldsymbol{\beta}))|T=t]$ and $\mathbb{E}\left[\pi_0(T,\boldsymbol{X})L'(Y-g(T;\boldsymbol{\beta}))|\boldsymbol{X}=\boldsymbol{x}\right]$ are twice differentiable, and the derivatives are Lipschitz continuous and uniformly bounded.

\noindentCondition 5. $\mathbb{E}\left[\sup_{\boldsymbol{ \beta}\in \Theta}|L'(Y-g(T;\boldsymbol{\beta}))|^{\sigma}\right]<\infty$ for some $\sigma>2$.

\noindentCondition 6.The bandwidths $h_1\asymp \cdots \asymp h_r \asymp h_Y \asymp h_T\asymp h_N$ go to zero slowly enough such that

align*[align* omitted — 138 chars of source]

Some Extensions

Proof of Theorem 7.1

(Proof of Consistency). Let

align*[align* omitted — 163 chars of source]

then $\hat{\theta}_K(t)=\hat{\gamma}^{\top}u_{K_1}(t)$. By assumption, there exists $\gamma^*\in\mathbb{R}^{K_1}$ such that

align[align omitted — 159 chars of source]

We first claim that

align[align omitted — 156 chars of source]

and the proof will be established later. With the claim (ref), we first show that $\int_{\mathcal{T}}|\hat{\theta}_K(t)-\theta(t)|^2dF_T(t)=O_p\left(\frac{\zeta(K)^2K}{N}+\zeta(K)^2K^{-2\alpha}+K_1^{-2\tilde{\alpha}}\right)$. Note that

align*[align* omitted — 667 chars of source]

With the claim (ref), we next show that $\sup_{t\in\mathcal{T}}|\hat{\theta}_K(t)-\theta(t)|=O_p[\zeta_1(K_1)(\zeta(K)\sqrt{K/N}+\zeta(K)K^{-\alpha}+K_1^{-\alpha})]$. Note that

align*[align* omitted — 789 chars of source]

Finally, we turn back to prove the claim (ref). Note that

align*[align* omitted — 776 chars of source]

where

align*[align* omitted — 604 chars of source]

We first compute the probability order of $A_{1N}$. We use the following notation:

align*[align* omitted — 311 chars of source]

Then we can obtain that

align*[align* omitted — 1,720 chars of source]

where the first inequality follows from the fact that $tr(AB)\leq \lambda_{\max}(B)tr(A)$ for any symmetric matrix $B$ and positive semidefinite matrix $A$, the second inequality follows from the same fact and the fact that $U_{N\times K_1}(U^{\top}_{N\times K_1}U_{N\times K_1})^{-1}U_{N\times K_1}^{\top}$ is a projection matrix with maximum eigenvalue 1, and the fourth inequality follows from the facts that $|\lambda_{\min}(\hat{\Phi}_{K_1\times K_1})|^{-1}=O_p(1)$, $\sup_{(t,x)\in\mathcal{T}\times \mathcal{X}}|\hat{\pi}_K(t,x)-{\pi}_0(t,x)|=O_p\left(\zeta(K)K^{-\alpha}+\zeta(K)\sqrt{K/N}\right)$ and $N^{-1}\sum_{i=1}^NY_i^2=O_p(1)$. \\

Next, we compute the probability order of $A_{2N}$. We can deduce that

align*[align* omitted — 629 chars of source]

where the last equality follows that $|\lambda_{\min}(\hat{\Phi}_{K_1\times K_1})|^{-1}=O_p(1)$ and $N^{-2}\|U_{N\times K_1}^{\top}\mathcal{E}_N\|^2=O_p(K_1/N)$ by Markov's inequality. \\

We finally compute the probability order of $A_{3N}$. We define the notation

align*[align* omitted — 218 chars of source]

then it follows that with probability approaching to 1,

align*[align* omitted — 1,577 chars of source]

where the first inequality follows from the fact that $tr(AB)\leq \lambda_{\max}(B)tr(A)$ for any symmetric matrix $B$ and positive semidefinite matrix $A$, the second inequality follows from the same fact and the fact that $U_{N\times K_1}(U^{\top}_{N\times K_1}U_{N\times K_1})^{-1}U_{N\times K_1}^{\top}$ is a projection matrix with maximum eigenvalue 1, and the last equality follows from the fact that $|\lambda_{\min}(\hat{\Phi}_{K_1\times K_1})|^{-1}=O_p(1)$ and the fact that $\frac{1}{N}\sum_{i=1}^N\left\{\mathbb{E}[{\pi}_0(T_i,\boldsymbol{X}_i)Y_i|T_i]-(\gamma^*)^{\top}u_{K_1}(T_i)\right\}^2\leq \sup_{t\in\mathcal{T}}|\mathbb{E}[{\pi}_0(T,X)Y|T]-(\gamma^*)^{\top}u_{K_1}(t)|^2=O(K_1^{-2\tilde{\alpha}})$. Thus we complete the proof of (ref).\\

(Proof of Asymptotic Normality). We have the following decomposition for $\hat{\theta}(t)-\theta(t)$:

align*[align* omitted — 728 chars of source]

where

align*[align* omitted — 589 chars of source]

Then we have that

align*[align* omitted — 185 chars of source]

We shall show that $b_{1N}(t)$ contributes to the asymptotic variance; and $b_{2N}(t)+b_{3N}(t)$ contributes to the asymptotic bias which is asymptotically negligible. Thus to complete the proof of asymptotic normality, it is sufficient to prove the following results:

itemize$V_K\geq c \|u_{K_1}(t)\|^2$ for some $c>0$; • $\sqrt{N}V_K^{-1/2}b_{1N}(t)\xrightarrow{d}N(0,1)$; • $\sqrt{N}V_K^{-1/2}b_{2N}(t)=o_p(1)$; • $\sqrt{N}V_K^{-1/2}b_{3N}(t)=o_p(1)$.

We first prove Result (i). Note that $\lambda_{\min}(\Sigma_{K_1\times K_1})\geq \underline{c}_{\sigma^2}\lambda_{\min}(\Phi_{K_1\times K_1})=\underline{c}_{\sigma^2}\lambda_{\min}(I_{K_1\times K_1})\geq \underline{c}_{\sigma^2}$, we can have

align*[align* omitted — 303 chars of source]

For the claim (ii). Let

align*[align* omitted — 442 chars of source]

Similar to the proof of Lemma (ref), we can have that

align*[align* omitted — 86 chars of source]

Then

align*[align* omitted — 509 chars of source]

For $B_{1N,1}(t)$, we can simply apply the Liapounov CLT and show that $B_{1N,1}(t)\xrightarrow{d}N(0,1)$. For $B_{1N,2}(t)$, let $\boldsymbol{T}=(T_1,...,T_N)$, we can obtain that

align*[align* omitted — 963 chars of source]

where we use the fact that $N^{-1} \mathbb{E}\left[U_{N\times K_1}^{\top}\mathcal{E}_N\mathcal{E}_N^{\top}U_{N\times K_1}|\boldsymbol{T}\right]=N^{-1}\sum_{i=1}^Nu_{K_1}(T_i)u_{K_1}(T_i)^{\top}\sigma^2(T_i)$ has bounded maximum eigenvalue. Therefore, $B_{1N,1}(t)=o_p(1)$ by the conditional Chebyshev's inequality. Thus (ii) holds. \\

For (iii), by Cauchy-Schwarz's inequality, we can obtain that

align*[align* omitted — 802 chars of source]

Similarly, we can show show that $\sqrt{N}V_{K}^{-1/2}|b_{3N}(t)|=o_p(1)$. This completes the proof of the Theorem.

Proof Theorem 7.2

Similar to the proof of Lemma (ref), let $\mu(t,x)=\mathbb{E}[Y|T=t,X=x]$, we decompose $\sqrt{N}(\hat{\psi}_K-\psi)$ as follows:

align[align omitted — 3,152 chars of source]

Using the similar argument for showing that (ref)-(ref) are all $o_p(1)$ in the proof of Lemma (ref), we can obtain that the terms (ref)-(ref) are all $o_p(1)$, while the term (ref) is asymptotically normal and attains the efficiency bound.

\singlespacing