Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Inference for Nonlinear Endogenous Treatment Effects Accounting for High-Dimensional Covariate Complexity
\affil[a]{Department of Economics, The Chinese University of Hong Kong}
\affil[b]{Department of Statistics, Rutgers University}
abstractNonlinearity and endogeneity are prevalent challenges in causal analysis using observational data. This paper proposes an inference procedure for a nonlinear and endogenous marginal effect function, defined as the derivative of the nonparametric treatment function, with a primary focus on an additive model that includes high-dimensional covariates. Using the control function approach for identification, we implement a regularized nonparametric estimation to obtain an initial estimator of the model. Such an initial estimator suffers from two biases: the bias in estimating the control function and the regularization bias for the high-dimensional outcome model. Our key innovation is to devise the double bias correction procedure that corrects these two biases simultaneously. Building on this debiased estimator, we further provide a confidence band of the marginal effect function. Simulations and an empirical study of air pollution and migration demonstrate the validity of our procedures.
JEL classification: C14, C21, C26, C55
\\
Keywords: Nonlinear causal effects, control function, double bias correction, data-rich environment.
\onehalfspacing
Introduction
Nonlinear treatment effects with endogeneity are prevalent in empirical economic studies. The econometric model could be high-dimensional with the increasing availability of rich datasets that include many potential covariates. For example, in gopalan2021home, the effect of home equity on labor income is found to be nonlinear, while as a measure of home equity, the loan-to-value (LTV) ratio is endogenous. Their study addressed the issue of endogeneity using the synthetic Loan-To-Value (LTV) ratio as an instrumental variable (IV). This ratio is constructed using synthetic loans, the original LTV level, and changes in the house price index. Other control variables might include the original loan amount, purchase price, loan balance, job tenure, and homeowner's age, among many other individual characteristics. Another example, which is used as our later empirical study, is the potential nonlinear effect of environmental pollution on migration chen2022. Air pollution is endogenous due to unobserved common economic factors affecting both migration and pollution. Thermal inversion is a popular IV in related environmental economic studies. We consider high-dimensional controls, including county-level income, expenditure, investment in health and education, etc. A third example is the nature of the job ladder in terms of working time and career advancement. Gicheva2013 found a positive but nonlinear relationship between weekly working hours and hourly wage growth using data\footnote{Data were from the 1979 cohort of National Longitudinal Survey of Youth, U.S. Bureau of Labor Statistics, and GMAT Registrant Survey.} from the high-end labor market. Potential covariates include gender, race, age, the number of children under 18, marital status, the mother's education, the college GPA, the GMAT score, working experience, and whether a person is enrolled in school. The marginal effects seem small at low levels of working hours and increase at high levels like 50 hours or above. However, some missing variables, including parents' involvement in learning and human capital formation, are likely to be correlated with working hours, which causes bias in the results.
Motivated by the real applications above, we propose new estimation and inference procedures for nonlinear treatment functions with high-dimensional covariates. For observations indexed by $i=1,2,\dots,n$ with a scalar outcome $Y_i$, a scalar endogenous treatment $D_i$, high-dimensional covariates $X_{i\cdot}$ and finite-dimensional instruments $Z_{i\cdot}$, we consider the model satisfying the conditional moment restriction $\mathbb{E}(Y_i-g(D_i)-X_{i\cdot}^{\top}\theta|X_{i\cdot},Z_{i\cdot})=0$, where the function $g$ and high-dimensional coefficients $\theta$ are unknown. Granted that our methodology applies to the original nonlinear treatment effect function $g(\cdot)$, we focus on the inference for the marginal effect function, defined as the derivative function $g'(\cdot)$. It is interpreted as the “marginal effect" at given levels of the treatment variable, analogous to the slope coefficient in a linear causal model. The valid inference for the nonlinear marginal effect function, taking into account high-dimensional covariates, is essential for empirical researchers and policy-makers.
This paper uses the control function approach newey1999 for model identification. It is well known in the literature (e.g., neweypowell89,Florens03,aichen2003) that a fully nonparametric IV model encounters an ill-posed inverse problem and the “curse of dimensionality". We thus consider an additive model specified in Section (ref), which is already a challenging problem in the presence of high-dimensional covariates.
Main Results and Contributions
Under the control function setup, our procedure mainly includes the following steps: First, estimate the control function using the LASSO regularization tibshirani1996regression; Second, construct an initial estimator of the marginal effect function using LASSO again; Third, correct the regularization bias in the initial estimator for inference. Our constructed initial estimator suffers from two sources of regularization bias from estimating both the control function in the first step and the nonlinear treatment in the second step. The former causes new obstacles to the inference problem, which is absent from the high-dimensional inference in additive models without endogeneity.
To address this difficulty, we propose a novel debiasing procedure for valid inference of the marginal effect function in the presence of endogeneity, nonlinearity, and high dimensionality. We call it the double bias correction because it corrects two sources of biases. Furthermore, we construct a uniform confidence band for the nonlinear marginal effect in the manner of lu2020kernel, adopting the techniques of multiplier bootstrap chernozhukov2014anti,chernozhukov2014gaussian.
The simulation study supports the validity of our estimator and confidence band of the endogenous marginal effect in the presence of high-dimensional covariates. In our empirical analysis of migration and air pollution, we find new results in addition to those of chen2022, who consider only the linear effect. Specifically, using thermal inversion as IV, the effect of pollution on migration is found to be insignificant when the pollution level is very low or moderately high. The effect is significant when pollution is worse than the “good air quality” level but below the medium level or at high levels. The magnitude of the effect increases at very high levels of pollution. In Section (ref), we provide an explanation of this result based on the reference-dependent principle in behavioral economics.
The main contributions of the paper are summarized as follows.
enumerate• In a high-dimensional and endogenous setting, we propose the double bias correction procedure to correct for regularization biases from both the outcome regression and the nonlinear control function estimation.
• We provide a uniform confidence band for the nonlinear marginal effect function. The development of a uniform confidence band for the nonlinear endogenous effects under high dimensions appears to be novel in the literature.
Literature Review
Our research connects to the literature on nonparametric estimation, causal inference, and high-dimensional models. Endogeneity in general nonparametric models was considered in neweypowell89, Matzkin94, blundell03, Hall05, darolles2011nonparametric. Newey1990 studied the efficiency of a nonparametric reduced form function for the instrumental variables. The control function method for nonparametric IV models was widely studied in newey1999,Horowitz2011,Wooldridge05, Guo16,su2008local,ozabaci2014additive,lee2007endogeneity,aghion2013innovation.
Recently, chen2018optimal, chen2021adaptive developed optimal sup-norm rates and uniform confidence band for the nonlinear treatment function and its derivatives. Breunig2022adaptive considered minimax adaptive estimation of quadratic functionals in nonparametric IV models. babii2020honest developed honest inference for a wide class of ill-posed regression models, including the nonparametric IV model. Distinguished from the above literature focusing on low-dimensional models, we specialize in the valid inference of nonlinear marginal effect that addresses the complexity from high-dimensional covariates. chernozhukov2022locally,chernozhukov2022automatic established automatic debiasing estimators of functionals like average treatment effect for general nonlinear models. Our method differs in that we use the control function approach for identification, and develop a uniform confidence band for the marginal effect function.
A surge of machine learning methods provides solutions to the estimation and inference of linear treatment effects with rich observational features. The popular double machine learning (DML) by chernozhukov2018double corrects the regularization bias from estimating the linear treatment effect with high-dimensional covariates. belloni2014 proposed a post-selection inference procedure for high-dimensional linear treatment effects models. fan2020endogenous considered a model with a mixture of controls and instruments. fan2022testing proposed an overidentification test for high-dimensional linear IV models. Unlike the above studies, we focus on the nonlinear effect.
Another strand of literature focuses on the inference for high-dimensional nonlinear models without endogeneity. The estimation of high-dimensional partially linear models was considered by wang2010estimation,muller2015partial,yu2016minimax. The estimation of high-dimensional additive models has also been frequently discussed meier2009high,huang2010variable,koltchinskii2010sparsity,suzuki2013fast,yuan2016minimax,tan2019doubly. lu2020kernel extended the procedure of javanmard2014confidence to the high-dimensional additive model. kozbur2021inference and gregory2021statistical proposed inferential procedures for additive models through post-selection or debiased estimators. su2019non proposed bootstrapping inference for high-dimensional nonseparable models. guo2022decorrelated used a decorrelated local linear estimator for inference of the first-order derivative of the target function. ning2023 proposed estimation and inference for high-dimensional partially linear models with estimated outcomes. The complexity of endogeneity distinguishes our problem from the abovementioned literature.
Notations. We use “$\stackrel{p}{\to}$” and “$\stackrel{d}{\to}$” to denote convergence in probability and distribution, respectively. The phrase “with probability approaching one as $n\to\infty$" is abbreviated as “w.p.a.1". We use “${\mathbb{E}}(\cdot)$" to denote the expectation and “${\mathbb{E}}_{\mathcal{F}}(\cdot)$" to denote the conditional expectation ${\mathbb{E}}(\cdot|\mathcal{F})$ for any $\sigma$-field $\mathcal{F}$. For any positive sequences $a_n$ and $b_n$, “$a_n\lesssim b_n$” means there exists some constant $C$ such that $a_n\leq Cb_n$, “$a_n\gtrsim b_n$” means $b_n\lesssim a_n$, and “$a_n\asymp b_n$” indicates $a_n\lesssim b_n$ and $b_n\lesssim a_n$. We use $[n]$ for some $n\in\mathbb{N}$ to denote the integer set $\{1,2,\cdots,n\}$. The floor and ceiling functions are $\lfloor\cdot\rfloor$ and $\lceil\cdot\rceil$, respectively. For a $p$-dimensional vector $x=(x_{1,}x_{2},\cdots,x_{p})^{\top}$, the number of nonzero entries is $\|x\|_0$, the $L_{2}$ norm is $\left\Vert x\right\Vert _{2}=\sqrt{\sum_{j=1}^{p}x_{j}^{2}}$, the $L_{1}$ norm is $\left\Vert x\right\Vert_{1}=\sum_{j=1}^{n}\left|x_{j}\right|$, and its maximum norm is $\|x\|_{\infty}=\max_{j\in[p]}|x_{j}|$. For a $p\times r$ matrix $A=(A_{ij})_{i\in[p],j\in[r]}$, we define the maximum norm $\|A\|_\infty = \max_{i,j}|A_{i,j}|$, $L_2$ norm $\|A\|_2 = \sqrt{\lambda_{\max}(A^\top A)}$ and the $L_1$ norm $\|A\|_1 = \max_{j\in[r]}\sum_{i\in[p]}|A_{ij}|$. For a square-integrable function $f(\cdot)$, define its $L_2$ norm as $\|f\|_2 = (\int f^2(x) dx)^{1/2}$. We use $0_p$ to denote the $p\times 1$ null vector. The indicator function is ${\textbf 1}(\cdot)$. For any function $g(\cdot)$, its first-order derivative function is $g^\prime(\cdot)$.
Paper Organization. The remainder of the paper is organized as follows. In Section (ref), we introduce the model and the main methodology. Section (ref) provides the main theoretical justification. In Section (ref), we demonstrate the finite-sample performance of our nonlinear effect estimator and uniform confidence band. Section (ref) provides an empirical example. Section (ref) concludes the paper. Technical proofs and additional simulation results are provided in the Appendix.
Model and Methodology
Model
We consider the following additive model:
align[align omitted — 181 chars of source]
where $Y_i\in\mathbb{R}$ is the outcome variable, $D_i\in\mathbb{R}$ is the univariate treatment variable, $g(\cdot)$ is the unknown treatment function of interest, $X_{i\cdot}\in\mathbb{R}^{p}$ represents the high-dimensional covariates\footnote{Without loss of generality, $X_{i\cdot}$ may include nonlinear terms such as the polynomials of some elementary covariates. The extension of (ref) and (ref) to the nonparametric additive model is possible.}, $Z_{i\cdot}\in\mathbb{R}^{p_z}$ denotes the instrumental variables, and $(u_i,v_i)^\top$ are unmeasured errors. Due to the existence of unmeasured confounders, $u_i$ might be correlated with $v_i$, leading to the endogeneity of $D_i$. In the above model (ref), we consider the additive nonlinear relation between $D_i$ and $Z_i$, where the $\psi_\ell(\cdot)$'s, for $\ell=1,2,\cdots,p_z$, are unknown functions. It accommodates the common linear reduced form with $\psi_\ell(Z_{i\ell})=Z_{i\ell}\psi_\ell$. The additive model in (ref) has been widely considered in the literature belloni2012sparse,fan2018nonparametric,ozabaci2014additive. We assume that the dataset $\mathscr{D}_n := \{Y_i,D_i,X_{i\cdot},Z_{i\cdot}\}_{i\in[n]}$ is independently and identically distributed (i.i.d.). We allow the dimension $p$ of covariates $X_{i\cdot}$ to be larger than $n$ and assume that the high-dimensional nuisance parameters $\theta,\varphi\in{\mathbb{R}}^{p}$ are sparse with the sparsity level $s=\max\{\|\theta\|_0,\|\varphi\|_0\}$. We assume that $p_z$ is a fixed number ($p_z=1$ in many applications).
The main focus of this paper is the statistical inference for the derivative function $g^\prime(\cdot)$.
Formally, we call $g^\prime(\cdot)$ the marginal effect function or marginal effect for short\footnote{The marginal effect function that we define here (see also chen2021adaptive) should be distinguished from the term “marginal treatment effect” as in heckman05. The latter is defined as the average causal effect of $D$ on $Y$ for individuals with some observed $X = x$ and unobserved $U = u$, where $U$ is a continuously distributed random variable satisfying the monotonicity condition of imbens94 with a binary treatment $D$. See the excellent survey by mogstad2018.}. This includes $g(D_i)=D_i\beta$ as a special case, where the derivative $g^\prime(D_i)=\beta$ is the homogeneous treatment effect in linear causal models.
The identification of $g(D_i)$ relies on the control function approach. Specifically, we impose the condition that
equation[equation omitted — 110 chars of source]
which is widely used in the literature florens2008identification,newey1999,imbens2009identification,su2008local,Horowitz2011,Wooldridge05. A simple sufficient condition for ((ref)) is that $(u_i,v_i)$ is independent of the exogeneous covariates $X_{i\cdot}$ and the instruments $Z_{i\cdot}$.
Define $q(v_i) := \mathbb{E}(u_i|v_i)$. The function $q(v_i)$ is called the control function in the literature; see blundell03 for a review. Then we have the following decomposition:
equation[equation omitted — 187 chars of source]
For simplicity, we assume that $\mathbb{E}(g(D_i))$, $\mathbb{E}(q(v_i))$, and $\mathbb{E}(\psi_\ell(Z_{i\ell}))$ for all $\ell\in[p_z]$ are all zero. In practice, we allow nonzero expectations by adding an intercept term to the model and handle the intercept by demeaning.
Estimation
Throughout the paper, we use the B-spline basis functions defined in Section (ref) following chen2018optimal to approximate the nonlinear functions in the model specified by ((ref))-((ref)). We use the uniformly based knots schumaker2007spline for B-spline functions such that the distances between every two adjacent knots are equal. Specifically, we use $B$, with the $(i,j)$-th element $B_{ij}=B_{j}(D_i)$, to denote an $n\times M_D$ matrix of $g(D_i)$'s basis functions; we use $H$, with the $(i,j)$-th element $H_{ij}=H_{j}(v_i)$, to denote an $n\times M_v$ matrix of $q(v_i)$'s basis functions; we use $K_\ell$, with the $(i,j)$-th element $(K_{\ell})_{ij}=K_{j\ell}(Z_{i\ell})$, to denote an $n\times M_\ell$ matrix of $\psi_\ell(Z_{i\ell})$'s basis functions, and $K=(K_1,K_2,\cdots,K_{p_z})$. We define $B(\cdot) = (B_1(\cdot),B_2(\cdot),\dots,B_{M_D}(\cdot))^\top$ as a functional vector that collects the spline functions of $g(\cdot)$, and similarly define $H(\cdot)$ and $K_\ell(\cdot)$. More spline functions result in smaller approximation errors of the nonlinear functions but larger variances in the estimators. The choice of numbers of spline functions $M_D$, $M_v$, and $M_\ell$ are discussed in the third paragraph of Section (ref) for simulation studies.
Let $\mathcal{D}$ denote the support of $D_i$. By Proposition (ref), when the function $g$ belongs to the $\gamma$-th H\"{o}lder Class specified in Definition (ref), there exists a $\beta\in\mathbb{R}^{M_D}$ such that the following approximation error
equation[equation omitted — 44 chars of source]
satisfies
equation[equation omitted — 171 chars of source]
Therefore, the coefficient $\beta$ guarantees accurate spline approximations of both the original treatment function $g(\cdot)$ and its derivative $g^\prime(\cdot)$. Similarly, we use the bases $H(\cdot)$ to approximate $q(\cdot)$, and $K_\ell(\cdot)$ to approximate $\psi_\ell(\cdot)$ for $\ell=1,2,\dots,p_z$. There exist $\eta\in\mathbb{R}^{M_v}$ and $\kappa_\ell\in\mathbb{R}^{M_\ell}$ for $\ell=1,2,\dots,p_z$ such that the approximation errors
equation[equation omitted — 43 chars of source]
and
equation[equation omitted — 97 chars of source]
satisfy similar error bounds as (ref).
We write $B_{i\cdot} = (B_1(D_i),\cdots.B_{M_D}(D_i))^\top$ as the spline functions for the $i$-th individual; $H_{i\cdot}$ and $K_{i\cdot}$ are defined similarly. Furthermore, define $\kappa = (\kappa_1^\top, \kappa_2^\top,\cdots,\kappa_{p_z}^\top)^\top$. With the spline approximation, the original models ((ref))-((ref)) are rewritten as follows:
align[align omitted — 253 chars of source]
where $r_{\psi i} := \sum_{\ell=1}^{p_z}r_{\ell}(Z_{i\ell})$. In order to carry out the regression for the aforementioned models, it is necessary to estimate $H_{i\cdot}$ that includes an unobservable random error $v_i$. To address the high dimensionality in ((ref)), we implement the following partial $L_1$ penalized regression
equation[equation omitted — 246 chars of source]
where the tuning parameter $\lambda_D>0$ is chosen by cross validation. The $L_1$ penalization in (ref) is only applied to $\varphi$ to address the high-dimensionality of $X_{i\cdot}$. We compute the residuals $\widehat{v}_i=D_i- K_{i\cdot}^\top\widehat\kappa - X_{i\cdot}^{\top}\widehat{\varphi}$ and use it to estimate $v_i$.
Recall that $H_j(\cdot)$ for $j=1,2,\dots,M_v$ are basis functions of $q(v_i)$. Define $\widehat{H}_{ij} = H_{j}(\widehat{v}_i)$, $\widehat H_{i\cdot} = (\widehat{H}_{i1},\widehat{H}_{i2},\dots,\widehat{H}_{iM_v})^\top$, and $\widehat{H} = (\widehat H_{1\cdot},\widehat H_{2\cdot},\dots,\widehat H_{n\cdot})^\top$. The model ((ref)) can be rewritten as follows:
equation[equation omitted — 150 chars of source]
where the approximation error $r_i$ admits the following form
equation[equation omitted — 97 chars of source]
The above error $r_i$ includes the spline approximation error $r_g(D_i)+r_q(\widehat{v}_i)$ and the control function estimation error $q(v_i)-q(\widehat{v}_i)$.
To estimate $\beta$, $\eta$, and $\theta$, we use a similar partially $L_1$ penalized regression for the outcome model ((ref)),
equation[equation omitted — 284 chars of source]
where $\lambda_Y>0$ is the tuning parameter chosen by cross-validation. Similar to (ref), the penalty is imposed on the high-dimensional vector $\theta$ only.
Define $B^\prime(d) = (B_j^\prime(d))_{j\in[M_D]}$ as the derivative of spline functions. The plug-in estimator for $g^\prime(d)$ using $\widehat{\beta}$ is
equation[equation omitted — 98 chars of source]
The above plugin estimator admits the following error,
equation[equation omitted — 155 chars of source]
where $r_{g}^\prime(d) = g^\prime(d) - B^\prime(d)^\top\beta $ is a higher-order term from the spline approximation error of $g^\prime(d)$. Even though the penalization in (ref) is not directly applied to the parameter $\beta$, the plug-in estimator $\widehat{g}^\prime(d)$ using $\widehat\beta$ still inherits a certain level of bias from (ref). This is due to the correlation between the treatment variable $D_i$ and the high-dimensional covariates $X_{i\cdot}$. We propose a new debiased estimator for $g^\prime(d)$ in Section (ref), and construct a uniform confidence band for $g^\prime(d)$ based on the bias-corrected estimator in Section (ref).
Double Bias Correction
We aim at a debiased estimator of the marginal effect function $g^\prime(d)$, approximated by $B^\prime(d)^\top \beta$. The primary task is thus to find a debiased estimator of $\beta$.
Earlier literature zhang2014confidence,van2014asymptotically,javanmard2014confidence developed debiased LASSO estimators for inference of high-dimensional models without endogeneity. In this section, we develop a novel debiasing procedure to overcome the unique challenge in our model, which arises from high dimensionality, nonlinearity, and endogeneity.
To fix ideas, let the $p_F$-dimensional vector $\widehat F_{i\cdot}$ collect some “regressors”. The bias-corrected estimator of $\beta$ has the following form:
equation[equation omitted — 155 chars of source]
where $\widehat\varepsilon_i$ is the residual, and $\widehat\Omega_B$ is an estimator of a submatrix composed of the first $M_D$ rows of the inverse Gram matrix of $\widehat F_{i\cdot}$. We will specify $\widehat F_{i\cdot}$ and $\widehat\Omega_B$ in (ref) and (ref), respectively.
In view of ((ref)), we examine the LASSO residual $\widehat\varepsilon_i$ of the model ((ref)) and ((ref)). Define $\widehat{W}_{i\cdot} := (B_{i\cdot},\widehat{H}_{i\cdot})$, $\omega := (\beta^\top,\eta^\top )^\top$, and $\widehat\omega = (\widehat\beta^\top,\widehat\eta^\top )^\top$ as the initial LASSO estimator of $\omega$. The fitted value is defined as $\widehat Y_i = \widehat{W}_{i\cdot}^\top\widehat\omega + X_{i\cdot}^\top\widehat\theta$, and the residual is decomposed as
equation[equation omitted — 345 chars of source]
We see the necessity of double bias correction from the above error decomposition (ref). The first source of bias $\varDelta_i^Y$ comes from the regularization in the outcome regression ((ref)). It can be addressed by existing bias correction methods. The second source of bias $\varDelta_i^v$ comes from the control function approximation using an $L_1$ regularized regression ((ref)). The additional bias $\varDelta_i^v$ brings a unique challenge to our high-dimensional nonparametric endogenous effect model with the control function. Consequently, we need a solution to simultaneously correct $\varDelta_i^Y$ from regularization ((ref)) and $\varDelta_i^v$ from the regularization ((ref)). We call our solution double bias correction to highlight the necessity of removing two biases.
Using Taylor expansion, we have the following approximation of $q(v_i)-q(\widehat v_i),$
\[q(v_i)-q(\widehat v_i) \approx q^\prime(\widehat v_i)( v_i - \widehat v_i) \approx q^\prime(\widehat v_i) X_{i\cdot}^\top (\widehat\varphi - \varphi) + q^\prime(\widehat v_i) K_{i\cdot}^\top (\widehat\kappa - \kappa), \]
where $\widehat\varphi$ and $\widehat\kappa$ are from ((ref)), and the higher-order biases and spline approximation errors are omitted for simplicity of presentation. This approximation is linear in the LASSO estimators $\widehat\varphi$ and $\widehat\kappa$. Thus, the residual (ref) is approximated by
equation[equation omitted — 355 chars of source]
The linearization (ref) is not operational for bias correction since the function $q^\prime(\cdot)$ is infeasible. We thus estimate it by
equation[equation omitted — 122 chars of source]
where $\widehat\eta$ comes from the LASSO regression ((ref)), and $\widehat{H}_{i\cdot}^\prime$ is an $M_v\times 1$ vector with the $j$-th coordinate being the derivative of the spline function $\widehat H_{ij}^\prime = H^\prime_j(\widehat v_i)$. Now the regressors in ((ref)) are approximated by
equation[equation omitted — 234 chars of source]
where $p_F$ is the length of $\widehat F_{i\cdot}$. Notice that the valid instruments $Z_{i\cdot}$ have no direct effects on $Y_i$ in model ((ref)). Nevertheless, the regularization bias $\varDelta_i^v$ from estimating the control function relates to the splines $K_{i\cdot}$ of the instruments $Z_{i\cdot}$. Thus, though the instruments $Z_{i\cdot}$ do not explicitly appear in model ((ref)), their spline functions $K_{i\cdot}$ are included in $\widehat F_{i\cdot}$ and used for bias correction.
We then construct a debiased LASSO estimator based on the residual $\widehat\varepsilon_i$ and the regressors $\widehat F_{i\cdot}$ defined in (ref). Define $\widehat{F} :=(\widehat{F}_{i\cdot}^\top)_{i\in[n]}^\top$. Then $\widehat\Omega_B^\top = (\widehat\Omega_1,\cdots,\widehat\Omega_{M_D})$ is a $p_F\times M_D$ matrix whose $j$-th column is defined as follows:
equation[equation omitted — 307 chars of source]
where $\textbf{i}_j$ is the $j$-th standard basis and $\mu_j$ is a tuning parameter. The constraint $\|\widehat\Sigma_{F} \Omega - \textbf{i}_j\|_{\infty} \leq \mu_j$ is imposed to control the bias term, and $\|n^{-1/2}\widehat F \Omega \|_\infty\leq \mu_j$ bounds the higher-order conditional moments and hence guarantees the asymptotic normality for non-Gaussian errors $\varepsilon_i$. By the bias-corrected estimator $\widetilde\beta$ defined in ((ref)), we finally construct a debiased estimator for the marginal effect function:
equation[equation omitted — 227 chars of source]
where $\widehat m(d)$ = $\widehat\Omega_B^\top B^\prime(d)$. The tuning parameter selection is specified in the last two paragraphs of Section (ref).
We use the debiased estimator ((ref)) for inference of the marginal effect $g^\prime(d)$. Define
equation[equation omitted — 211 chars of source]
For any fixed $d$, a point-wise $100(1-\alpha)\%$ confidence interval is given by
equation[equation omitted — 322 chars of source]
where $z_{1-\alpha/2}$ is the $(1-\alpha/2)$-th quantile of the standard normal distribution. To establish the uniform confidence band of the whole function $g^\prime(d)$ for all $d$ over its domain $\mathcal{D}$, we will use a critical value from multiplier bootstrap in place of the $z_{1-\alpha/2}$ used in ((ref)).
Uniform Confidence Band
We define the following quantity:
equation[equation omitted — 159 chars of source]
where $\widehat{s}(d)$ and $\widehat{\sigma}_\varepsilon$ are defined in ((ref)). Under the regularity conditions stated in Section (ref), we show that $\mathbb{H}_n(d)$ converges in distribution to a standard normal variable for a given treatment level $d$. We proceed to a confidence band for all $d$ over its domain $\mathcal{D}$ by considering the distribution of the empirical process $ \mathbb{H}_n(d) $. Following the techniques of chernozhukov2014anti,chernozhukov2014gaussian, we approximate $\mathbb{H}_n(d)$ by the following Gaussian multiplier process:
equation[equation omitted — 170 chars of source]
where $e_i$ for $1\leq i\leq n$ are i.i.d.\ standard normal variables. Let $\widehat c_n(\alpha)$ be the $(1-\alpha)$-th quantile of $ \sup_{d\in\mathcal{D}} \left|\widehat{\mathbb{H}}_n(d)\right| $. Then, we construct the following confidence band at level $100(1-\alpha)\%$,
equation[equation omitted — 465 chars of source]
The critical value $\widehat c_n(\alpha)$ approximates the $(1-\alpha)$-th quantile of $\sup_{d\in\mathcal{D}}\left| \mathbb{H}_n(d)\right|$, which is the supreme absolute value of an asymptotically Gaussian empirical process. The theoretical justifications of Gaussian approximation by chernozhukov2014anti,chernozhukov2014gaussian suggest that $\widehat c_n(\alpha)$ is a good approximation such that the confidence band ((ref)) is asymptotically honest.
All methods described thus far have utilized the full sample. Unlike the regression models without endogeneity, our regressors $\widehat F_{i\cdot}$ used for bias correction include some initial LASSO estimators that substantially complicate the theoretical analysis for our double bias correction. These estimators include $\widehat\varphi$ and $\widehat\kappa$ that construct the residual $\widehat v_i$ of the LASSO regression ((ref)), and $\widehat\eta$ used to estimate the derivative $q^\prime$ as in ((ref)). Thus, sample-splitting (see details in the following Algorithm (ref)) is required for these initial LASSO estimators in theoretical derivations. We highlight that the sample-splitting is merely a technical restriction. In Section (ref) for numerical studies, we demonstrate that the confidence band using a full sample has a reasonable coverage rate and is more informative than the split-sample confidence band due to a larger effective sample size, and thus we recommend using the full-sample inference for practitioners.
To keep consistent with the theoretical justifications, we summarize the split-sample procedure for a uniform confidence band of $g^\prime(d)$ in Algorithm (ref) below\footnote{The code for implementation is available at \url{https://github.com/ZiweiMEI/HDNPIV}.}. The algorithm for the full-sample inference is summarized in Appendix (ref). We remark that our procedure also works for the original function $g(d)$, with $B^\prime(d)$ in ((ref)) replaced by the original splines $B(d)$ and ((ref))-((ref)) revised accordingly. \\
algorithm[algorithm omitted — 1,616 chars of source]
Asymptotic Theory
Throughout the paper, we consider the number of covariates $p$, the sparsity index $s$, and the number of bases functions $M_D$, $M_v$, and $M_\ell$ for $\ell=1,2,\dots,p_z$ as deterministic functions of the sample size $n$. In asymptotic statements, we only explicitly let $n\to\infty$, thus allowing $p$, $M_D$, $M_v$, and $M_\ell$ all go to infinity, whereas $s$ is either fixed or divergent.
We impose the following theoretical assumptions.
Assumption[Control Function] Assume that the dataset $\mathscr{D}_n := \{Y_i,D_i,X_{i\cdot},Z_{i\cdot}\}_{i\in[n]}$ is independently and identically distributed. Furthermore, suppose that ((ref)) holds, and $\mathbb{E}(v_i|Z_{i\cdot},X_{i\cdot}) = 0$.
We assume that $\sigma_v^2 = \mathbb{E}(v_i^2|Z_{i\cdot},X_{i\cdot})$ and $\sigma_\varepsilon^2 = \mathbb{E}(\varepsilon_i^2|v_i,Z_{i\cdot},X_{i\cdot})$ are positive constants where $\varepsilon_i$ is defined in ((ref)). Furthermore, ${\mathbb{E}}(\varepsilon_i^4|v_i,Z_{i\cdot},X_{i\cdot})$ is bounded by some positive constant independent of $n$ and $p$.
As mentioned earlier in Section (ref), we impose the condition ((ref)) for identification using the control function approach in Assumption (ref). We also impose the conditional homoskedasticity of $v_i$ and $\varepsilon_i$. Even though the consistency of the initial estimators does not require such homoskedasticity assumption, it is unclear how to relax this assumption for establishing asymptotic normality of our double-bias-correction estimators unless we impose additional sparsity conditions on precision matrix of $\widehat F_{i\cdot}$ (the inverse of $\Sigma_{F|\mathcal{L}}$ defined before Proposition (ref)). Since $\widehat F_{i\cdot}$ includes LASSO estimators and spline functions, the sparsity of this precision matrix is difficult to deduce from low-level assumptions. For the sake of simplicity and to maintain focus on our methodological contributions, we will concentrate on homoskedasticity, thereby avoiding any diversions. We further assume the bounded third order moment of the random shock $\varepsilon_i$ to control for the error bounds.
Assumption[Bounded Supports and Densities]
Suppose that $D_i\in\mathcal{D}$, $v_i\in\mathcal{V}$, $X_{ij}\in\mathcal{X}$ $\forall$ $j\in[p]$, and $Z_{i\ell}\in\mathcal{Z}$ for all $\ell\in[p_z]$ are continuous, where $\mathcal{D}=[a_D,b_D]$, $\mathcal{V}=[a_v,b_v]$, $\mathcal{X}=[a_x,b_x]$, and $\mathcal{Z}=[a_z,b_z]$ are compact intervals in $\mathbb{R}$. Additionally, the marginal density functions of $D_i$ and $v_i$, as well as the joint density functions of $(Z_{i1},\cdots,Z_{ip_z})$, are absolutely continuous with density functions bounded below by $c_f>0$ and above by $C_f>0$.
We assume boundedness of $v_i$ and $D_i$ to guarantee small errors of B-spline approximation. This technical assumption has been used in the control function literature newey1999, imbens2009identification,su2008local,ozabaci2014additive. In practice, we can normalize the data to a compact interval. As pointed out by ozabaci2014additive, boundedness is possible to remove with arguments in su2012sieve, inducing additional complications. Finally, we focus on the continuously valued treatment, instruments, and covariates.
We further impose regular conditions on the nonlinear functions $g(\cdot)$, $q(\cdot)$ and $\psi_\ell(\cdot)$ in models (ref) and (ref). We define the functions of the H\"{o}lder Class as follows.
DefinitionThe $\gamma$-th H\"{o}lder Class $\mathcal{H}_{\mathcal{X}}(\gamma,L)$ is the set of $\gamma$-times differentiable functions $f:\mathcal{X}\to\mathbb{R}$ such that its derivative $f^{(\ell)}$ with $\ell=\lfloor\gamma\rfloor$ satisfies the following:
\begin{equation}
|f^{(\ell)}(x)-f^{(\ell)}(y)|\leq L|x-y|^{\gamma-\ell}, for any x,y\in\mathcal{X}.
\end{equation}
When $\mathcal{X}$ is compact in the real line, $f\in\mathcal{H}_{\mathcal{X}}(\gamma,L)$ implies that all $j$-th derivatives of $f$ are bounded in $\mathcal{X}$ for any $j\in\{1,2,\cdots,\ell=\lfloor\gamma\rfloor\}$. To address the error between $\widehat v_i$ and $v_i$, we assume that the domain of $q(\cdot)$ is $\mathcal{V}_q=[a_v-\epsilon_v,b_v+\epsilon_v]$ for some constant $\epsilon_v > 0$, which extends $\mathcal{V}=[a_v,b_v]$, the support of $v_i$, , as specified in Assumption (ref).
Assumption[H\"{o}lder Class] Suppose that
$g(\cdot)\in\mathcal{H}_{\mathcal{D}}(\gamma,L)$, $q(\cdot)\in\mathcal{H}_{\mathcal{V}_q}(\gamma,L)$ and $\psi_\ell(\cdot)\in\mathcal{H}_{\mathcal{Z}}(\gamma,L)$ for some $\gamma\geq 2$ and positive constant $L$. Further assume $\sum_{\ell=1}^{p_z}\psi_\ell^\prime(Z_{i\ell}) \neq 0$ with probability one.
Assumption (ref) about the compact supports and bounded densities for $D_i$ and $v_i$, together with Assumption (ref), guarantees a small spline function approximation error. The condition of nonzero derivative in Assumption (ref) relates to identification, as explained in the following Remark (ref).
Remark[Identifiability]
According to newey1999, the models ((ref))-((ref)) are identifiable if (a) the boundary of the support of $(Z_{i\cdot}^\top, X_{i\cdot}^\top,v_i)$ shares zero probability; (b) all functions are differentiable; and (c) $\sum_{\ell=1}^{p_z}\psi_\ell^\prime(Z_{i\ell})$ is almost surely nonzero. These conditions are reflected in our assumptions. Specifically, condition (a) is implied by the bounded support and density condition by Assumption (ref). Condition (b) is implied by Assumption (ref) that all functions are in the H\"{o}lder class. The last condition (c), imposed by Assumption (ref), means the IVs are relevant to the endogenous variable $D$.
In the following, for any random vector $\zeta = (\zeta_{j})_{j\geq 1}$, we use $\widetilde \zeta = (\widetilde\zeta_{j})_{j\geq 1}$ to denote the standardized vector with $\widetilde\zeta_{ij} :=({\mathbb{E}}(\zeta_{ij}^2))^{-1/2}\zeta_{ij}$.
Assumption[Eigenvalues] Suppose that ${\mathbb{E}}(X_{ij}^2)$ is bounded away from zero and above uniformly for all $j$. Furthermore, define $Q_{i\cdot}:=( K_{i\cdot}^\top, X_{i\cdot}^\top)^\top$ and $U_{i\cdot}=(B_{i\cdot}^\top,H_{i\cdot}^\top,X_{i\cdot}^\top)^\top$. We use $\widetilde Q_{i\cdot}$ and $\widetilde U_{i\cdot}$ to denote their standardized versions. Suppose that the eigenvalues of ${\mathbb{E}}(\widetilde Q_{i\cdot}\widetilde Q_{i\cdot}^\top)$ and ${\mathbb{E}}(\widetilde U_{i\cdot}\widetilde U_{i\cdot}^\top)$ are bounded away from zero and above.
Remark[No Perfect Collinearity]
In Assumption (ref), the vectors $Q_{i\cdot}$ and $U_{i\cdot}$ respectively collect the regressors, or their population version, in the LASSO regressions ((ref)) and ((ref)). Intuitively, Assumption (ref) rules out perfect collinearity of the regressors in LASSO. We need this assumption for the restrictive eigenvalue conditions to derive the LASSO consistency for high-dimensional models bickel2009simultaneous. We impose the assumption on the standardized regressors, since the B-splines $B_{i\cdot}$, $H_{i\cdot}$, and $K_{i\cdot}$ have Gram matrices with eigenvalues convergent to zero as the number of spline bases pass to infinity, and thus are of smaller scales than the covariates $X_{i\cdot}$.
Remark[IV Strength]
The bounded-away-from-zero eigenvalues of ${\mathbb{E}}(\widetilde U_{i\cdot}\widetilde U_{i\cdot}^\top)$ in Assumption (ref) can be considered as a restriction on the IV strength. This condition means that the variables in $U_{i\cdot}$, including (functions of) $D_i$, $v_i$ and $X_{i\cdot}$, are not perfectly correlated. Intuitively, if the nonlinear functions of IVs $\{\psi_\ell(Z_{i\ell})\}_{\ell=1}^{p_z}$ in the model ((ref)) introduce sufficient variations to $D_i$, the variables $D_i$, $v_i$, and $X_{i\cdot}$ in ((ref)) are far away from perfect collinearity, and the eigenvalue condition for ${\mathbb{E}}(\widetilde U_{i\cdot}\widetilde U_{i\cdot}^\top)$ can thus be satisfied. The sufficient variations from $\{\psi_\ell(Z_{i\ell})\}_{\ell=1}^{p_z}$ requires that neither the covariates $X_{i\cdot}$ and the functions of IVs $\psi_\ell(Z_{i\ell})$ for $\ell=1,\dots,p_z$ are perfectly collinear, nor $\sum_{\ell=1}^{p_z}\psi_\ell(Z_{i\ell})$ is close to a constant function. The latter means the treatment $D_i$ and the IVs $Z_{i\cdot}$ are significantly relevant, which holds in many empirical applications.
We further state the restrictions on the asymptotic regime.
AssumptionSuppose the following conditions hold:
\begin{enumerate}[(a)]
• The sparsity index $s = O(n^{1/4-c_0})$ for an arbitrarily small $ c_0 \in (0, \frac{1}{4})$ and $\log p = O(n^{c_\gamma})$ with the constant $c_\gamma$ only depending on $\gamma$ in Assumption (ref).
• The number of spline bases satisfy $M_D\asymp M_v \asymp M_\ell$ for $\ell=1,2,\dots,p_z$. Define $M := \max\left(M_D, M_v, M_1, \dots, M_{p_z} \right)$. Assume $M \asymp n^\nu$ with $\nu\in[\frac{1}{2\gamma+1},\frac{1}{4})$. The LASSO tuning parameters satisfy $\lambda_Y = C_Y\sqrt{\frac{\log (pM)}{n}}$ and $\lambda_D = C_D\sqrt{\frac{\log (pM)}{n}}$ for some sufficiently large constants $C_Y,C_D>0$.
\end{enumerate}
The rates specified in Assumption (ref) of the tuning parameters $\lambda_D$, $\lambda_Y$ are similar to the linear model case, which is applied to the remainder of the paper without further clarifications. These rates are only for technical proofs; in practice, we use cross-validation for data-driven choice of LASSO tuning parameters.
The following theorem establishes the rate of convergence for the initial estimators proposed in (ref). Recall that $\widehat\beta,\widehat\eta$ are the estimated coefficients of spline functions, and $\widehat\theta$ stores the estimated coefficients of covariates.
TheoremSuppose that Assumptions (ref)-(ref) hold. Then the initial estimators $\widehat\beta,\widehat\eta,\widehat\theta$ in ((ref)) and $\widehat g^\prime(\cdot)$ in ((ref)) have the following error bounds w.p.a.1:
\begin{equation}
\|\widehat{\beta}-\beta\|_2^2 + \|\widehat\eta - \eta\|_2^2 \lesssim M^{-2\gamma + 1} + \dfrac{(s+M)M\log (pM)}{n},
\end{equation}
\begin{equation}
\|\widehat{g}^\prime - g^\prime\|_2^2 \lesssim M^{-2(\gamma-1) } + \dfrac{M^2(s+M)\log (pM)}{n},
\end{equation}
\begin{equation}
\|\widehat{\theta}-{\theta}\|_1 \lesssim M^{-2\gamma}\sqrt{\dfrac{n}{\log p}} + (s+M)\sqrt{\dfrac{\log (pM)}{n}}.
\end{equation}
Though we consider estimation of the nonlinear endogenous effect with high-dimensional covariates, the error bounds in Theorem (ref) echo the literature of spline regression without endogeneity. For instance, refer to zhou2000derivative and huang2010variable.
In terms of inference, we need additional theoretical assumptions. We first formalize the independence assumption for split-sample estimators mentioned in the discussions before Algorithm (ref). Define the LASSO estimators $\widehat\varphi^{\text{ind}}$, $\widehat\kappa^{\text{ind}}$, and $\widehat\eta^{\text{ind}}$ as those in Algorithm (ref) using a different dataset independent of $\mathscr{D}_n$ specified in Assumption (ref), where the superscript “ind” means independence.
Assumption[Split-Sample Estimators]We assume that the LASSO estimators $\widehat\varphi^{\text{ind}}$, $\widehat\kappa^{\text{ind}}$, and $\widehat\eta^{\text{ind}}$ are constructed by another i.i.d.\ dataset independent of data $\mathscr{D}_n$ with a sample size $n^\prime \asymp n$.
We define $F_{i\cdot} := (B_{i\cdot}^\top,H_{i\cdot}^\top,{X}_{i\cdot}^\top,q^\prime(v_i) K_{i\cdot}^\top, q^\prime(v_i){X}_{i\cdot}^\top)^\top$ as the population truth of $\widehat F_{i\cdot}$ in ((ref)). We define $F_{ij}^{\rm std} := ({\mathbb{E}}(F_{ij}^2))^{-1/2}F_{ij}$ as the standardized version of $F_{ij}$ for any $j=1,2,\dots,p_F$, and $F_{i\cdot}^{\rm std} = (F_{i1}^{\rm std},F_{i2}^{\rm std},\dots,F_{ip_F}^{\rm std})^\top$.
AssumptionSuppose that $v_i$ is independent of the vector $(X_{i\cdot}^\top,Z_{i\cdot}^\top)^\top$ and $X_{i\cdot}^\top$ has a bounded sub-Gaussian norm. Furthermore, assume that the eigenvalues of ${\mathbb{E}}( F_{i\cdot}^{\rm std} (F_{i\cdot}^{\rm std})^\top)$ are bounded away from zero and above.
The sub-Gaussian norm of a random vector is defined in Definition (ref) in Appendix (ref). We need the bounded sub-Gaussian norm of the whole vector $X_{i\cdot}$ (on top of the compact support in Assumption (ref)) to rule out strong dependence among the high dimensional covariates. Similar to Assumption (ref), Assumption (ref) rules out perfect collinearity among the variables in $F_{i\cdot}$. This implies that $q$ is a nonzero and nonlinear function; otherwise, $q^\prime(v_i)$ is constant, and therefore, $X_{i\cdot}$ and $q^\prime(v_i) X_{i\cdot}$ are perfectly collinear. This assumption ensures the invertibility of $\Sigma_{F|\mathcal{L}}$, and thus, problem (ref) is feasible with high probability. For a linear $q$, a feasible solution is also available. In Appendix (ref), we discuss the feasibility of ((ref)) when $q$ is linear.
We first show the feasibility of (ref). Let $\mathcal{L}$ denote the $\sigma$-field generated by the split-sample LASSO estimators specified in Assumption (ref). We define $\Sigma_{F|\mathcal{L}} := \mathbb{E}_{\mathcal{L}}\left(\widehat F_{i\cdot} \widehat F_{i\cdot}^\top\right)$.
PropositionSuppose that Assumptions (ref)-(ref) hold. Then w.p.a.1,
\begin{align}
\left\|\widehat\Sigma_F \Sigma_{F|\mathcal{L}}^{-1} - I \right\|_{\infty} &\lesssim M\sqrt{\dfrac{\log (pM)}{n}}, \\
\|\widehat F^\top \Sigma_{F|\mathcal{L}}^{-1}\|_\infty &\lesssim M\sqrt{\dfrac{\log (pM)}{n}}.
\end{align}
Proposition (ref) shows that the problem (ref) is feasible w.p.a.1 when
equation[equation omitted — 77 chars of source]
for $j\in[M_D]$ with some constant $C_j$ large enough. Specifically, the first $M_D$ columns of $\Sigma_{F|\mathcal{L}}^{-1}$ fulfill the constraints in (ref) w.p.a.1. Note that the feasible solution is the inverse of the conditional Gram matrix $\Sigma_{F|\mathcal{L}}$ instead of the unconditional one ${\mathbb{E}}(\widehat F_{i\cdot} \widehat F_{i\cdot}^\top)$, since $\{\widehat F_{i\cdot}\}_{i\in[n]}$ are i.i.d.\ conditionally on $\mathcal{L}$, but not unconditionally. The convergence rate of a tuning parameter like ((ref)) is a commonly used technical assumption for proofs javanmard2014confidence,gold2020inference,lu2020kernel. For practitioners, we introduce a data-driven choice of $\mu_j$ in the last paragraph of Section (ref) for simulations.
For asymptotic normality of the debiased estimator, we need stronger restrictions on the sparsity index $s$ and the number of spline bases $M$.
manualassumption{(ref)$^\prime$}
Suppose the following conditions hold:
\begin{enumerate}[(a)]
• The sparsity index $s = O(n^{2/9-c_0})$ for an arbitrarily small $ c_0 \in (0, \frac{2}{9})$ and $\log p = O(n^{c_\gamma})$ with $c_\gamma$ dependent only on $\gamma$ in Assumption (ref).
• The number of spline bases satisfy $M_D\asymp M_v \asymp M_\ell$ for $\ell=1,2,\dots,p_z$. Recall $M = \max\left(M_D, M_v, M_1, \dots, M_{p_z} \right)$ Assume $M \asymp n^\nu$ with $\nu\in(\frac{1}{2\gamma},\frac{1}{4.5})$. The LASSO tuning parameters satisfies $\lambda_Y = C_Y\sqrt{\frac{\log (pM)}{n}}$ and $\lambda_D = C_D\sqrt{\frac{\log (pM)}{n}}$ for some $C_Y,C_D>0$ large enough.
• The tuning parameter $\mu_j$ in (ref) follows (ref).
\end{enumerate}
Assumption (ref) for asymptotic normality of the debiased estimator distinguishes from Assumption (ref) for consistency of initial estimators in the following aspects. First, in Assumption (ref)(a) we require $s = o(n^{2/9})$ that is slightly stronger than the rate $o(n^{1/4})$ in Assumption (ref)(a). Second, a stronger condition $\nu\in (\frac{1}{2\gamma},\frac{1}{4.5})$ is needed in Assumption (ref)(b) to ensure that the error caused by (ref) is small enough to guarantee asymptotic normality. Specifically, the lower bound $\frac{1}{2\gamma}$ ensures that the spline approximation errors are small enough, and the upper bound $\frac{1}{4.5}$ controls for the variance of $\widehat q^\prime(\widehat v_i)$ in (ref) to bound the estimation error of $q^\prime(v_i)$. The range of $\nu$ implies that $\gamma > 2.25$, which imposes slight additional smoothness on the second-order derivatives compared to Assumption (ref). Third, in Assumption (ref), we specify the rate of $\mu_j$ in (ref) for bias correction.
We are now ready to present the main theoretical results for the double bias correction and uniform confidence band. Recall that $\widehat m(d)$ is defined below ((ref)) and $\widehat s(d)$ is defined below ((ref)).
PropositionSuppose that Assumptions (ref)-(ref), (ref), (ref), and (ref) hold. Then for any fixed $d\in\mathcal{D}$, the debiased estimator $\widetilde g^\prime(d)$ in ((ref)) satisfies
\begin{equation}
\sqrt{n}(\widetilde g^\prime(d) - g^\prime(d)) = \mathcal{Z}(d) + \Delta(d),
\end{equation}
where $ \Delta(d) / \widehat s(d) \stackrel{p}{\to} 0$ and $\mathcal{Z}(d) := \dfrac{\widehat m(d)^\top \sum_{i=1}^n \widehat F_{i\cdot} \varepsilon_i}{\sqrt{n}}$ with ${\mathbb{E}}(\mathcal{Z}(d)|\widehat F_{i\cdot}) = 0$ and ${\rm var}[\mathcal{Z}(d)|\widehat F_{i\cdot}] = \left(\sigma_\varepsilon\widehat s(d)\right)^2$. In addition, $\widehat s(d) \asymp M^{1.5}$ uniformly for all $d\in\mathcal{D}$ w.p.a.1.
Proposition (ref) presents a result on the decomposition of the estimation error for the debiased estimator $\widetilde g^\prime(d)$. The asymptotic distribution of $\sqrt{n}(\widehat g^\prime(d) - g^\prime(d))$ is determined by $\mathcal{Z}(d)$ with a zero conditional mean and a conditional variance $(\sigma_\varepsilon\widehat s(d))^2$. It can be shown that $\mathcal{Z}(d)/\sigma_\varepsilon\widehat s(d)$ converges in distribution to a standard normal variable for any fixed $d$. We relegate this intermediate result and highlight the honesty of the uniform confidence band for the whole function displayed in Theorem (ref) below.
The last result in Proposition (ref) implies that the standard error for the doubly bias corrected derivative function estimator is $M^{1.5}/\sqrt{n}$ w.p.a.1. Though we study the nonparametric inference of a derivative function under a high-dimensional model with endogeneity, the standard error of our debiased estimator shares the same order with the nonparametric estimator using spline regressions under the low-dimensional model without endogeneity zhou2000derivative.
Next we present the asymptotic honesty of the confidence band (ref). For simplicity, we only consider the honesty of our confidence band over the H\"{o}lder class of functions $g$ in the following theorem. We treat other nonrandom components in ((ref)) and ((ref)) as fixed, including the nonlinear functions $ \psi_\ell(\cdot)$ for $\ell = 1,2,\dots,p_z$, the coefficients of covariates $\theta$ and $\varphi$, and the distributions of covariates $X_{i\cdot}$, instruments $Z_{i\cdot}$, and unmeasured errors $(u_i,v_i)$.
TheoremUnder the conditions in Proposition (ref),
the $100(1-\alpha)\%$ confidence band $\mathcal{C}_{n,\alpha}^{D}=\{\mathcal{C}_{n,\alpha}(d):d\in\mathcal{D}\}$ defined as (ref) is asymptotically honest in the sense that
\begin{equation}
\mathop{\lim\inf}\limits_{n\to\infty}\inf_{g\in\mathcal{H}_{\mathcal{D}}(\gamma,L)}\Pr\left(g^{\prime}(d)\in\mathcal{C}_{n,\alpha}(d) for all d\in\mathcal{D}\right) \geq 1 - \alpha,
\end{equation}
where the
H\"{o}lder functional class $\mathcal{H}_{\mathcal{D}}(\gamma,L)$ is defined in Definition (ref).
Theorem (ref) formally establishes the asymptotic honesty of our confidence band ((ref)) over the
H\"{o}lder class of functions $g$. The result is shown by the techniques of Gaussian coupling for the suprema of empirical processes by chernozhukov2014anti,chernozhukov2014gaussian.
Simulations
Setup
In the simulation studies, we generate the outcome $Y_i$ and the treatment $D_i$ following models (ref) and (ref). We generate the covariates $X_{i\cdot}$ and IVs $Z_{i\cdot}$ following lu2020kernel. Specifically, let $(U_{ji})_{i\in[n],j\in[p+p_z+1]}$ be i.i.d.\ $U[0,1]$ random variables, and
\[X_{ij} = \dfrac{U_{i,j}+0.3U_{i,p+p_z+1}}{1.3},\ j\in[p],\]
\[Z_{ij} = \dfrac{U_{i,p + j}+0.3U_{i,p+p_z+1}}{1.3},\ j\in[p_z].\]
The error terms are generated from $v_i\sim\text{ i.i.d.\ }\sqrt{12}U[-0.5,0.5]$, $\varepsilon_i \sim\text{ i.i.d.\ }N(0,1)$ and $u_i = q(v_i) + \varepsilon_i$ with $q(v) = v^2 - 1$. We vary $p\in\{150,400,800\}$ for high-dimensional covariates and the sample size $n\in\{500,1000,2000,3000\}$. For the sample-splitting inference, we divide the two samples into two subsets with cardinalities $n_a = \lfloor n/2 \rfloor$ and $n_b = n - n_a$.
We set $\theta = (1_6^\top, 0_{p-6}^\top)^\top$ and $\varphi = (1,-1,1,-1,1,-1,0_{p-6}^\top)^\top$. Additionally, we set $p_z=1$ for a just-identified IV model, which is the leading case in empirical studies. We use $\psi(z)=4(2z-1)^2$ as the nonlinear function measuring the relevance of the IV in (ref).
We consider four cases of the $g$ function in (ref):
\[g_1(d) = 0,\ g_2(d) = d,\ g_3(d)=0.05(d-3)^2,\ g_4(d)=0.02(d-3)^3. \]
The small coefficients in the nonlinear treatment functions balance the relatively large range of $D_i$ for moderate values of marginal functions. We simulate $10^5$ samples following the DGPs (ref) and (ref) and form a compact interval with the lower and upper bounds as the 10th and 90th percentiles of the simulated $D_i$, respectively. We take 1000 grid points in this compact interval and construct the confidence band on these grid points.
We use cubic B-spline functions for estimation and set the number of basis functions as $M_D = M_v = M_z = 5$ following lu2020kernel. As a robustness check, we provide additional simulations in Table (ref) for the finite-sample performance with different numbers of $M_D$, with $M_v$ and $M_z$ fixed to 5. It turns out that a larger $M_D$ brings almost no benefits to the coverage of the confidence bands but produces larger variances due to higher variable dimensions.
The LASSO tuning parameters in (ref) and (ref) are selected based on cross-validation. Following the idea of gold2020inference, we set the tuning parameter $\mu_j$ in (ref) as
$\mu_j = a_0 \cdot \min_\Omega \left\|\widehat\Sigma_F \Omega - \textbf{i}_j\right\|_{\infty}$,
where the factor $a_0$ is chosen to balance the bias control, for which a smaller $a_0$ is desirable, and the confidence band's length, for which a larger $a_0$ is desirable. We set $a_0=1.2$ following gold2020inference.
Numerical Results
Table (ref) shows the simulation results for the linear $g$ functions. The left panel “Full-Sample” shows the results from full-sample inference where we always use all samples for all estimators and bias correction procedures. The right panel “Split-Sample” shows the split-sample results based on the procedures described in Algorithm (ref).
We first focus on the left panel showing the full-sample results. In the “BiasInit” and “BiasDB” columns, we see that compared to the initial plug-in LASSO estimator defined as $\widehat{g}^\prime(\cdot)$ in (ref), the bias-corrected estimator $\widetilde{g}^\prime(\cdot)$ in (ref) significantly reduces the bias from LASSO penalization. The bias of the bias-corrected estimator decreases as the sample size grows. In terms of inference, when $p=150$, the coverage probability of the full-sample confidence bands exceeds the nominal size of 0.95 under all sample sizes. The average confidence band length continues to decrease as the sample size increases. When $p=400$ or $800$, the coverage deviates from the nominal size when the sample is not large enough. As a nonparametric method, our inferential procedure requires a relatively large sample size to handle high-dimensional covariates under finite sample. Among the settings where the coverage is close to the nominal size, the average confidence length decreases as the sample size increases.
We then compare the full-sample and split-sample results. Under the same sample size, the split-sample confidence band, compared to the full-sample one, produces either a much worse coverage, or a far wider confidence band if its coverage is close to the nominal size. These comparative results show that even if sample splitting is required for technical proofs, the full-sample inference dominates the split-sample procedure in practice because of a larger effective sample size. We thus recommend the full-sample inference for practitioners.
Table (ref) shows the results for nonlinear $g$. In general, the performance of our procedure is robust to different settings for the $g$ functions. The simulation results show the validity of our procedure for the uniform confidence band of endogenous marginal effects.
As a robustness check, additional simulation results are provided in Appendix (ref). These additional results show that our procedure is robust to unbounded distributions, although the theory depends on the assumption of boundedness as most high-dimensional nonparametric methods do. In addition, as mentioned in Section (ref), we compare various settings of $M_D$ and show that $M_D=5$ is a reasonable choice since an increase in the number of basis functions does not improve coverage but brings additional variance.
table[table omitted — 5,023 chars of source]
table[table omitted — 5,076 chars of source]
The Nonlinear Effect of Pollution on Migration
In this section, we revisit the study of pollution and migration. chen2022 studied the effects of air pollution on migration in China using changes in the average strength of thermal inversions over five-year periods as a source of exogenous variation for medium-run air pollution levels. Their findings suggested that air pollution is responsible for large changes in inflows and outflows of migration in China. Specifically, they found that a 10% increase in air pollution is capable of reducing the population through outmigration by approximately 2.8% in a given county.
Air pollution might have nonlinear effects on net outmigration. We use the proposed nonlinear treatment effects estimator on the same dataset in chen2022 and enrich it with many other covariates. Our method has two advantages compared to those of the original study. First, we can explore the potential nonlinear effect of pollution on migration. Second, we can improve the estimation accuracy by explicitly controlling for high-dimensional characteristics in the model. Specifically, we consider the following empirical econometric model:
align[align omitted — 170 chars of source]
where for any variable $x_i$, the notation $\ddot{x}_i$ denotes its standardized version with a zero sample mean and a unit standard deviation.
table[table omitted — 2,646 chars of source]
Instead of using the fixed effect to account for unobserved factors of migration, we explicitly use the county-level controls of one cross-sectional data point, which is the five-year average of 2006-2010 that contains 2860 counties. Table (ref) provides the summary statistics of the main variables. The outcome variable, $Y$, denotes the measure of migration in county $i$. Specifically, we use the destination-based net immigration ratio, which is the fraction of people entering a county but with their hukou in their place of origin, net of the outflows. The treatment variable, denoted as $D$, measures the 5-year average concentration of PM2.5. The key identification strategy is the exogenous variation in thermal inversion Ransom1995,Arceo2016. A thermal inversion refers to an abnormal temperature-altitude gradient, where the air becomes hotter instead of cooler with higher altitude and traps pollutants near the ground. Other key covariates include income, expenditure, health, education, grain subsidies, the workforce in agriculture, temperature, precipitation, sunshine, humidity, and wind. To evaluate the performance of our procedure for high-dimensional covariates, we add another 105 geo-economic variables to control for potential omitted variable bias\footnote{These variables include but are not limited to industrial, educational, and environmental variables from the China City Statistical Yearbook and China County Statistical Yearbook. We use a population- or GDP-weighted method to convert some city-level variables to county-level variables.}.
Following the simulation studies, we take 1000 grid points between the 10th and 90th percentiles of $D_i$. The selection of tuning parameters also follows the simulation section. Panel (A) of Figure (ref) illustrates the baseline results. The horizontal axis represents the standardized pollution level $D$, and the vertical axis shows the marginal effect. We see a clear nonlinear pattern of the treatment effect. When the pollution level is below the mean, the marginal effect of pollution on immigration is insignificant, showing that people can tolerate low-level pollution. When pollution approaches the mean value such that the standardized $D$ is approximately zero, air pollution shows a significant negative marginal effect on immigration; thus, pollution starts to drive out the population. As pollution becomes more severe but is at a moderately high level (within one standard deviation from the mean), the effect becomes insignificant. When the pollution level is very high, e.g., exceeding 1.3 standard deviations from the mean, the negative effect becomes significantly negative again and the magnitude continues to increase, indicating that people generally cannot tolerate high-level pollution. Quantitatively, in the significant range of treatment effects, at the low-medium level of pollution, a standard deviation increase in pollution is responsible for an approximately -0.6 standard deviation decrease in destination-based immigration. At the very high level of pollution, a standard deviation increase in pollution is responsible for up to a -1.8 standard deviation decrease in destination-based immigration.
The ineffectiveness of air pollution in the medium range may be explained by the reference dependence in prospect theory Kahneman1979. People in counties with moderately high pollution can tolerate a moderate level of pollution since they treat very high pollution levels as their reference point. Reference points are found to be an important factor in human behavior, such as retirement plans Seibold2021 and labor supply Crawford2011. Another possible explanation is adaptive behavior, such as the use of air purifiers when the pollution level increases.
figure[figure omitted — 252 chars of source]
For a further robustness check, we randomly generate 100 i.i.d.\ standard normal variables and add them to the covariate set. We expect them to be identified as pure noise by the proposed procedure such that they do not severely impact the results. The additional result shown in Panel (B) of Figure (ref) satisfies our expectation, as it presents a very similar pattern as in Panel (A) without these irrelevant noises.
Conclusion
We propose an inferential procedure to construct a uniform confidence band of the endogenous marginal effect function. We include high-dimensional covariates to account for potential omitted variable bias and adopt the control function approach to identify the endogenous effect function. Our main novelty is to design a novel double bias correction procedure to tackle the unique challenge in the nonparametric inference of under endogeneity and the high-dimensional covariate complexity. Our empirical application finds the nonlinear effect of air pollution on migration, which is an important complement to the recent empirical literature. To better facilitate practical studies, it is an exciting topic for future studies to extend our methodology to a model with multiple endogenous variables and heteroskedastic errors, which are also common in empirical applications.
\setcounter{section}{0}
\setcounter{page}{1}
\setcounter{footnote}{0}
{ \bf
center[center omitted — 137 chars of source]
}
{
center[center omitted — 261 chars of source]
}
\newcounter{counter}[section]
\setcounter{table}{0}
\setcounter{equation}{0}
\setcounter{figure}{0}
\setcounter{Theorem}{0}
\setcounter{Assumption}{0}
\setcounter{Lemma}{0}
\setcounter{Remark}{0}
\setcounter{Corollary}{0}
\setcounter{Proposition}{0}
\setcounter{Definition}{0}
\setcounter{subsection}{0}
\setcounter{table}{0}
\setcounter{equation}{0}
\setcounter{figure}{0}
\setcounter{Theorem}{0}
\setcounter{Assumption}{0}
\setcounter{Lemma}{0}
\setcounter{Remark}{0}
\setcounter{Corollary}{0}
\setcounter{Proposition}{0}
\setcounter{algocf}{0}
Algorithm of Full-Sample Inference
algorithm[algorithm omitted — 1,263 chars of source]
The Feasibility of ((ref)) When $q$ Is Linear
The control function $q(\cdot)$ is usually unknown in practice. The methodology proposed in the main text accommodates the special case of a linear $q$ function. To simplify the illustration, we assume $M_D = M_v = M_\ell = M$, and that $\Sigma_{F|\mathcal{L}}$ is observable so that our goal is to find the vector $\xi_j$ such that
equation[equation omitted — 87 chars of source]
for $j\in[M]$, where $\textbf{i}_j$ is the $j$-th standard basis. When $q(\cdot)$ is linear, there exists a constant $q_0$ such that $q^\prime(v) \equiv q_0$ for all $v$. Define $\Psi_i = (B_{i\cdot}^\top, H_{i\cdot}^\top, q_0K_{i\cdot}^\top)^\top \in \mathbb{R}^{M_{\rm all}}$ with $M_{\rm all} := (p_z+2)M$. Recall $X_{i\cdot}\in \mathbb{R}^p$. Thus, we have
\[\Sigma_{F|\mathcal{L}} = \left(
array[array omitted — 597 chars of source]
\right) =: \left(
array[array omitted — 148 chars of source]
\right)\]
where $A_{11} := {\mathbb{E}}_{\mathcal{L}}[\Psi_{i\cdot} \Psi_{i\cdot} ^\top]$, $A_{22} := {\mathbb{E}}_{\mathcal{L}}[X_{i\cdot} X_{i\cdot}^\top]$ and $A_{12}:= {\mathbb{E}}_{\mathcal{L}}[\Psi_{i\cdot} X_{i\cdot}^\top] $. By similar eigenvalue conditions in Assumption (ref), the submatrix
$$\Sigma_{11} := \left(
array[array omitted — 67 chars of source]
\right) \in \mathbb{R}^{(M_{\rm all}+p)\times(M_{\rm all}+p)}$$ is invertible. Define
\[\Omega_O := \left(
array[array omitted — 121 chars of source]
\right)\]
and $$\Sigma_{12}:= (A_{12}^\top,A_{22})^\top.$$ We show the solutions for $\{\xi_j\}_{j\in[M]}$ are the first $M$ columns of $\Omega_O$. Note that
\[
aligned\Sigma_{F|\mathcal{L}} \Omega_O = \left(\begin{array}{cc}
\Sigma_{11} & q_0\Sigma_{12} \\
q_0\Sigma_{12}^\top & q_0^2 A_{22}
\end{array}\right) \cdot \left(\begin{array}{cc}
\Sigma_{11}^{-1} & O_{(M_{\rm all}+p)\times p} \\
O_{p\times (M_{\rm all}+p)} & O_{p\times p}
\end{array}\right) = \left(\begin{array}{cc}
I_{M_{\rm all}+p} & O_{(M_{\rm all}+p)\times p} \\
q_0\Sigma_{12}^\top\Sigma_{11}^{-1} & O_{p\times p}
\end{array}\right).
\]
Observe that $\Sigma_{12}^\top$ is the last $p$ rows of $\Sigma_{11}$. Thus, $\Sigma_{12}^\top\Sigma_{11}^{-1}$ is the last $p$ rows of the identity matrix $I_{M_{\rm all}+p}$. It turns out that the first $M$ columns of $q_0\Sigma_{12}^\top\Sigma_{11}^{-1}$ are all zero. Consequently, the first $M$ columns of $\Sigma_{F|\mathcal{L}} \Omega_O = \left(
array[array omitted — 129 chars of source]
\right)$ are $\{i_j\}_{j\in[M]}$ so that the first $M$ columns of $\Omega_O$ are the solutions to ((ref)).
Proofs
Throughout the proofs, we use $c$ and $C$ (sometimes with subscripts) to denote generic positive constants irrelevant to the sample size, which may vary from place to place. The notations “$\lesssim_p$”, “$\gtrsim_p$”, and $\asymp_p$ indicate “$\lesssim$”, “$\gtrsim$”, and “$\asymp$” hold w.p.a.1 respectively. We assume that function $q$ and its derivatives $q^\prime(\cdot)$, $q^{\prime\prime}(\cdot)$, as well as the B-splines $H(\cdot)$ and the derivatives $H^\prime(\cdot)$, $H^{\prime\prime}(\cdot)$ are continuously extended to the whole real line. Specifically, $q(v) = q(a_v - \epsilon_v)$ for all $v < a_v - \epsilon_v $, and $q(b_v + \epsilon_v)$ for all $v > b_v + \epsilon_v$. Therefore, $\widehat v_i$ will always fall into the support of $v_i$.
Definitions
We first define the B-spline basis following chen2018optimal. We consider a uniform B-spline basis with $m$ interior knots and support $[a,b]$. Let $a = t_{-k} =\cdots = t_0 < t_1 < \cdots < t_{m} = t_{m+1} = \cdots = t_{m+k-1} = b$ denote the extended knot sequence. Let $I_0 = [t_0,t_1),\cdots,I_m = [t_m,t_{m+1}]$. As mentioned in the beginning of Section (ref), we use the uniformly based knots schumaker2007spline for B-spline functions such that the distances between every two adjacent knots are equal. All $I_m$'s thus share the same length.
A basis of degree 0 is constructed by
\[N_{j,0}(x) =
cases1,\ if x \in I_j, \\
0,\ otherwise
\]
for $j=0,\cdots,m$. Bases of degree $k>0$ is defined recursively by
\[N_{j,k}(x) = \dfrac{x-t_j}{t_{j+k}-t_j}N_{j,k-1}(x) + \dfrac{t_{j+k+1}-x}{t_{j+k+1}-t_j}N_{j+1,k-1}(x) \]
where $\frac{1}{0}:=0$ following chen2018optimal. Without loss of generality, we define $N_{j,k}(x)=0$ for $j<-k$ or $j>m$. Furthermore, for any $x\in I_j$, $N_{j^\prime,k}(x)$ is nonzero only if $j^\prime = j,\cdots,j+k$.
Define $B_{j}^{[k^\prime]}(x) := N_{j-k^\prime-1,k^\prime}(x)$ for any $j$ and $k^\prime$. We use $M$ to denote the number of spline functions used in the estimation and inference. By the definition of $N_{j,r}$ above, $M = m + k + 1$. In addition, the B-splines satisfy partition to unity $\sum_{j\in[M]}B_{j}^{[k^\prime]}(x) = 1$ for all $k^\prime \leq k$. Similar definitions and properties also hold for $H(\cdot)$ and $K_\ell(\cdot)$ in the corresponding supports and we will not repeat the statements.
In addition, we provide formal definitions for (conditionally) sub-Gaussian and sub-exponential random variables and vectors vershynin2010introduction.
Definition[Sub-Gaussian norms]
The sub-Gaussian norm of any random variable $x$ (conditional on the sigma-field $\mathcal{F}$) is
\begin{equation}
\|x\|_{\psi_2|\mathcal{F}} := \sup_{q\geq 1}\dfrac{1}{\sqrt{q}}[\mathbb{E}_{\mathcal{F}}|x|^q]^{1/q}.
\end{equation}
For any random vector $X\in\mathbb{R}^p$, we define its (conditional) sub-Gaussian norm as
\begin{equation}
\|X\|_{\psi_2|\mathcal{F}} := \sup_{b\in\mathbb{R}^p:\|b\|_2=1} \|b^\top x\|_{\psi_2|\mathcal{F}}.
\end{equation}
A random variable or vector is called (conditionally) sub-Gaussian if its (conditional) sub-Gaussian norm is uniformly bounded by some absolute constant.
Definition[Sub-Exponential norms]
The sub-exponential norm of any random variable $x$ (conditional on the sigma-field $\mathcal{F}$) is
\begin{equation}
\|x\|_{\psi_1|\mathcal{F}} := \sup_{q\geq 1}\dfrac{1}{q}[\mathbb{E}_{\mathcal{F}}|x|^q]^{1/q}.
\end{equation}
For any random vector $X\in\mathbb{R}^p$, we define its (conditional) sub-exponential norm as
\begin{equation}
\|X\|_{\psi_1|\mathcal{F}} := \sup_{b\in\mathbb{R}^p:\|b\|_2=1} \|b^\top x\|_{\psi_1|\mathcal{F}}.
\end{equation}
A random variable or vector is called (conditionally) sub-exponential if its (conditional) sub-exponential norm is uniformly bounded by some absolute constant.
Preliminary Propositions
Proposition (ref) provides widely used sup-norm bounds of cross-products. Propositions (ref)-(ref) state some properties of spline functions. Proposition (ref) and the consequent corollaries provide the estimation error of $\widehat v_i$ and its functions from the LASSO algorithm ((ref)). Throughout the whole Section (ref), all probabilistic bounds are independent of the function $g$, and thus hold uniformly for any $g\in\mathcal{H}_{\mathcal{D}}(\gamma,L)$.
PropositionUnder Assumptions (ref) and (ref), there exist some absolute constants $c$ and $C$ such that with probability at least $1-c(pM)^{-4}$
\begin{equation}
\left\|\dfrac{X^\top X}{n} - \mathbb{E}\left[\dfrac{X^\top X}{n}\right]\right\|_\infty + \left\|\dfrac{X^\top v}{n} \right\|_\infty + \left\|\dfrac{X^\top \varepsilon}{n} \right\|_\infty \leq C\sqrt{\dfrac{\log (pM) }{n}}.
\end{equation}
for any $t > 0$.
proof[Proof of Proposition (ref)]Since $v$ and $X$ are uniformly bounded, their cross products have bounded sub-Gaussian norm. The first two terms can be thus easily bounded using the Hoeffding-type inequality for centered sub-Gaussian variables vershynin2010introduction and the union bound. The procedures in the proof of Lemma B2 in fan2022testing apply to this proposition.
The bound of the third term applies the moderate deviation inequality belloni2012sparse that
\begin{equation}
\begin{aligned}
\Pr\left(\max_{j\in[p_x]}\dfrac{\left|\sum_{i=1}^n X_{ij}\varepsilon_i\right|}{\sqrt{\sum_{i=1}^n X_{ij}^2\varepsilon_i^2}} > \Phi^{-1}(1-\gamma_0/2p_x)\right) \leq \gamma_0 (1+A/\ell_n^3)
\end{aligned}
\end{equation}
where $A$ is an absolute constant and $\Phi$ is the cumulative distribution function of a standard normal variable, provided that $\ell_n > 0$ and under the i.i.d.\ setting,
\begin{equation}
0 \leq \Phi^{-1}(1-\gamma_0/2p) \leq \dfrac{n^{1/6}}{\ell_n} \min_{j\in[p]} \dfrac{\left( {\mathbb{E}}(X_{ij}^2\varepsilon_i^2)\right)^{1/2}}{\left( {\mathbb{E}}|X_{ij}^3\varepsilon_i^3|\right)^{1/3}} - 1.
\end{equation}
The following arguments verify (ref). By the boundness of $X_{ij}$ in Assumption (ref) and the bounded high-order conditional moment of $\varepsilon_i$ in Assumption (ref),
\[{\mathbb{E}}|X_{ij}^3\varepsilon_i^3| \lesssim {\mathbb{E}}|\varepsilon_i^3| \lesssim 1.\]
By the lower and upper bounds of ${\mathbb{E}}(X_{ij}^2)$ in Assumption (ref),
\[{\mathbb{E}}(X_{ij}^2\varepsilon_i^2) = {\mathbb{E}}(X_{ij}^2{\mathbb{E}}(\varepsilon_i^2|X_{ij})) = \sigma_\varepsilon^2 {\mathbb{E}}(X_{ij}^2) \asymp 1. \]
Thus, if we take $\ell_n=A$, the right hand side of (ref) is lower bounded by $cn^{1/6}$ for some absolute constant $c$. Then when $\gamma_0=M^{-1}$, (ref) follows by the fact that $\Phi^{-1}(1-\gamma_0/2p) \lesssim \sqrt{\log(2p/\gamma_0)} \lesssim \sqrt{\log (pM)}$.
The validity of (ref) implies (ref) when $\gamma_0 = M^{-1}$ and $\ell_n=A$, which also implies that the right hand side of (ref) converges to zero. Thus, w.p.a.1
\[\begin{aligned}
\|n^{-1}X^\top\varepsilon\|_\infty &\leq n^{-1/2}\cdot \Phi^{-1}(1-\gamma_0/2p)\cdot \max_{j\in[p]}\sqrt{n^{-1}\sum_{i=1}^n X_{ij}^2\varepsilon_i^2 } \\
&\lesssim n^{-1/2} \cdot \sqrt{\log (pM)} \cdot \sqrt{n^{-1}\sum_{i=1}^n \varepsilon_i^2 }\\
&\lesssim_p \sqrt{\log (pM) / n},
\end{aligned}\]
where the second row applies the $\Phi^{-1}(1-\gamma_0/2p)\lesssim \sqrt{\log (pM)}$ and the boundness of $X_{ij}$, and the third row applies $n^{-1}\sum_{i=1}^n \varepsilon_i^2 \stackrel{p}{\to} \sigma_\varepsilon^2$ by the i.i.d.\ and the third moment conditions.
PropositionFor all $1\leq i\leq n$ we have \[\max\{\|B_{i\cdot}\|_2, \|H_{i\cdot}\|_2, (\|(K_\ell)_{i\cdot}\|_2)_{\ell\in[p_z]} \}\leq 1,\]
\[\min\{\|B_{i\cdot}\|_2, \|H_{i\cdot}\|_2, (\|(K_\ell)_{i\cdot}\|_2)_{\ell\in[p_z]} \}\geq \dfrac{1}{\sqrt{k}}.\]
proof[Proof of Proposition (ref)]
By schumaker2007spline we know that for each fixed $x$, the cardinality of the set $\{j\in[M]:B_{j}(x) \neq 0\}$ is no greater than $k$. By partition of unity schumaker2007spline that $\sum_{j=1}^{M}B_{j}(D_i)=1$, we deduce that
$$ \|B_{i\cdot}\|_2 \leq \sum_{j=1}^{M}B_{j}(D_i) =1,$$
$$ \|B_{i\cdot}\|_2 \geq \dfrac{1}{\sqrt{k}}\sum_{j=1}^{M}B_{j}(D_i) = \dfrac{1}{\sqrt{k}}.$$
Likewise, we can deduce the same bounds for
$\|H_{i\cdot}\|_2$ and $\|(K_\ell)_{i\cdot}\|_2$.
The following Proposition (ref) is about the first-order derivative of B-Spline functions. Proposition (ref) is a direct result of de1978practical.
PropositionFor any $v\in[a_v-\epsilon_v,b_v+\epsilon_v]$,
$H_j^\prime(v)=M\left(H_{j}^{[k-1]}(v)-H_{j+1}^{[k-1]}(v)\right)$, where $H_{j}^{[k-1]}(v)$ is the $j$-th B-Spline function of degree $k-1$. Similarly, for any $d\in[a_D,b_D]$,
$B_j^\prime(d) = M\left(B_{j}^{[k-1]}(d)-B_{j+1}^{[k-1]}(d)\right)$.
Proposition$\|B^\prime(d))\|_q \asymp M$ and $\|H^\prime(v)\|_q\asymp M$ for $q=1,2$ uniformly for all $d\in[a_D,b_D]$ and $v\in[a_v-\epsilon_v,b_v+\epsilon_v]$. Also, $\|B^{\prime\prime}(d)\|_q,\|H^{\prime\prime}(v)\|_q\lesssim M^2$.
proof[Proof of Proposition (ref)]
We only prove this proposition for $B(d)$ and the arguments for $H(v)$ are exactly the same. We first focus on $q=2$. From Proposition (ref), we have $\|B^\prime(d)\|_2 =\sqrt{\sum_{j=1}^M(B_j^\prime(d))^2} = M\sqrt{\sum_{j=1}^M [B_j^{[k-1]}(d) - B_{j+1}^{[k-1]}(d)]^2}$. Given $B_{M+1}^{[k-1]}(d) = 0$, it is easy to show that
\[\begin{aligned}
\sum_{j=1}^M [B_j^{[k-1]}(d) - B_{j+1}^{[k-1]}(d)]^2 \leq 2\sum_{j=1}^M[B_j^{[k-1]}(d)]^2 + 2\sum_{j=1}^M[B_{j+1}^{[k-1]}(d)]^2
\leq 4\|B^{[k-1]}(d)\|_2^2 \lesssim 4.
\end{aligned}\]
We then show the other side of the inequality. Recall that for any $d$, there exists at most $k$ numbers $j\in[M]$ such that $B_{j}^{[k-1]}(d)\neq 0$. Suppose these $k$ numbers are $j_0,j_0+1,\cdots,j_0+k-1$. Consequently,
\[\sum_{j=1}^M [B_j^{[k-1]}(d) - B_{j+1}^{[k-1]}(d)]^2 \geq \frac{1}{k}\left(\sum_{j=j_0}^{j_0+k-1} \left|B_j^{[k-1]}(d) - B_{j+1}^{[k-1]}(d)\right|\right).\]
It suffices to show that $\sum_{j=j_0}^{j_0+k-1} \left|B_j^{[k-1]}(d) - B_{j+1}^{[k-1]}(d)\right|$ is bounded away from zero. We will show by contradiction that a lower bound is $\frac{1}{k+1}$. Suppose that $\sum_{j=j_0}^{j_0+k-1} \left|B_j^{[k-1]}(d) - B_{j+1}^{[k-1]}(d)\right| < \frac{1}{k+1}$. Then
\[\begin{aligned}
1 = \sum_{j=j_0}^{j_0+k-1} B_j^{[k-1]}(d) = \sum_{j=j_0}^{j_0+k-1} \sum_{i=j}^{j_0+k-1} \left[B_i^{[k-1]}(d) - B_{i+1}^{[k-1]}(d)\right] < \dfrac{k}{k+1}
\end{aligned}\]
which is a contradiction.
Therefore, $\|B^\prime(d)\|_2\asymp M$. As for the $L_1$ norm, note that
\[M \lesssim \|B^\prime(d)\|_2 \leq \|B^\prime(d)\|_1 = M \sum_{j=1}^M |B_j^{[k-1]}(d) - B_{j+1}^{[k-1]}(d)| \leq 2M \]
where the last inequality applies the partition to unity property that $\sum_{j\in[M]} B_j^{[k-1]}(d) = 1$. As for the second-order derivative, note that
\[\begin{aligned}
\|B^{\prime\prime}(d)\|_1 &= \sum_{j\in[M]} |B_j^{\prime\prime}(d)| \leq M\sum_{j\in[M]}|(B_j^{[k-1]})^\prime(d) - (B_{j+1}^{[k-1]})^\prime(d)| \leq M\cdot 2\|(B^{[k-1]})^\prime(d)\|_1 \lesssim M^2.
\end{aligned}\]
This completes the proof of Proposition (ref).
The following proposition about eigenvalues of Gram matrices for the spline bases and their derivatives.
PropositionDefine $B^\prime_{i\cdot}=B^\prime(D_i)$ and $H^\prime_{i\cdot} = H^\prime(v_i)$. Under Assumptions (ref), there exists some absolute constants $c$ and $C$ such that for $M$ large enough,
\begin{equation}
\begin{aligned}
cM^{-1} \leq \lambda_{\min}({\mathbb{E}}(B_{i\cdot} B_{i\cdot}^\top)) \leq \lambda_{\max}({\mathbb{E}}(B_{i\cdot} B_{i\cdot}^\top)) \leq CM^{-1} \\
cM^{-1} \leq \lambda_{\min}({\mathbb{E}}(H_{i\cdot} H_{i\cdot}^\top)) \leq \lambda_{\max}({\mathbb{E}}(H_{i\cdot} H_{i\cdot}^\top)) \leq CM^{-1} \\
cM^{-1} \leq \lambda_{\min}({\mathbb{E}}(K_{i\cdot} K_{i\cdot}^\top)) \leq \lambda_{\max}({\mathbb{E}}(K_{i\cdot} K_{i\cdot}^\top)) \leq CM^{-1}. \\
\end{aligned}
\end{equation}
and
\begin{equation}
\begin{aligned}
\lambda_{\max}({\mathbb{E}}(B^\prime_{i\cdot} (B^\prime_{i\cdot})^\top)) \leq CM \\
\lambda_{\max}({\mathbb{E}}(H^\prime_{i\cdot} (H^\prime_{i\cdot})^\top)) \leq CM.
\end{aligned}
\end{equation}
proof[Proof of Proposition (ref)]
Equation (ref) is a direct result of chen2018optimal. For (ref), we only show the bounds for $H^\prime$ since the same arguments apply to $B^\prime$. Define \[\Delta H_{i\cdot} := \left(H_{j}^{[k-1]}(v_i) - H_{j+1}^{[k-1]}(v_i)\right)_{j\in[M]},\] $H^{[k-1]}_{i\cdot} := \left(H_{j}^{[k-1]}(v_i)\right)_{j\in[M]}$ and $H^{[k-1]}_{+1,{i\cdot}} := \left(H_{j+1}^{[k-1]}(v_i)\right)_{j\in[M]}$. Note that
\[\begin{aligned}
\|{\mathbb{E}}(H_{i\cdot}^\prime H_{i\cdot}^{\prime\top})\|_2 = M^2\left\|{\mathbb{E}}\left[\Delta H_{i\cdot}^{{\otimes2}}\right]\right\|_2 \leq M^2 \left(\left\|{\mathbb{E}}[H^{(k-1){\otimes2}}_{i\cdot}]\right\|_2+\left\|{\mathbb{E}}[H^{(k-1){\otimes2}}_{+1,{i\cdot}}]\right\|_2\right) \lesssim M
\end{aligned}\]
where the first equality applies Proposition (ref) and the last inequality applies (ref).
Below are well-known results about B-spline function approximation that are essential in the theoretical analysis; See de1978practical and zhou2000derivative. The following proposition only shows the spline approximation errors of the function $g(\cdot)$; Similar results hold for other functions $q(\cdot)$ and $\psi_\ell(\cdot)$ and are not listed here for simplicity.
PropositionSuppose $\gamma\geq 2$ and $L>0$ are fixed. Then for any function $g\in\mathcal{H}(\gamma,L)$, there exists a $\beta\in\mathbb{R}^{M_D}$ such that
\begin{equation}
\sup_{d\in{\mathcal{D}}} \left|\sum_{j=1}^{{M}}\beta_jB_{j}(d)-g(d)\right|=O(M^{-\gamma}),
\end{equation}
\begin{equation}
\sup_{d\in{\mathcal{D}}} \left|\sum_{j=1}^{{M}}\beta_jB^\prime_{j}(d)-g^\prime(d)\right|=O(M^{-\gamma + 1}),
\end{equation}
The $O(\cdot)$ holds uniformly for all $g\in\mathcal{H}(\gamma,L)$.
The following Proposition bounds the errors of LASSO estimators in (ref).
PropositionSuppose that Assumptions (ref)-(ref) hold and $\lambda_D = C_D\sqrt{\dfrac{\log (pM)}{n}}$ for some $C_D>0$ large enough. There exist some constants $c$ such that with probability at least $1-c(pM)^{-4}$
\begin{equation}
\|\widehat\kappa - \kappa\|_2 \lesssim \left(\sqrt{\dfrac{n}{\log (pM)}}M^{-\gamma-0.5} + 1\right)M^{-\gamma+0.5} + \sqrt{\dfrac{M^2\log (pM)}{n}} + \sqrt{\dfrac{sM\log (pM)}{n}},
\end{equation}
\begin{equation}
\begin{aligned}
\|\widehat\varphi - \varphi\|_1 & \leq C\left(\sqrt{\dfrac{n}{\log (pM)}}M^{-2\gamma} + (s+M)\sqrt{\dfrac{\log (pM)}{n}} \right)
\end{aligned}
\end{equation}
and
\begin{equation}
\begin{aligned}
\|\widehat\varphi - \varphi\|_2^2 & \leq C\left( M^{-2\gamma} + (s+M)\cdot\dfrac{\log (pM)}{n}\right).
\end{aligned}
\end{equation}
proof[Proof of Proposition (ref)] By the definition of $\widehat\varphi$ and $\widehat\kappa$, we have
\begin{equation}
\frac{1}{n}\|D-K\widehat\kappa-X\widehat{\varphi}\|_2^2+\lambda_D\|\widehat{\varphi}\|_1\leq \frac{1}{n}\|D-K \kappa-X\varphi\|_2^2+\lambda_D\|\varphi\|_1.
\end{equation}
which implies
\begin{equation}
\frac{1}{n}\|K(\kappa-\widehat{\kappa})+X(\varphi-\widehat{\varphi})\|_2^2+\lambda_D\|\widehat{\varphi}\|_1\leq \frac{2}{n}\left( r_\psi+v\right)^\top\left[K(\kappa-\widehat{\kappa})+X(\varphi-\widehat{\varphi})\right] +\lambda_D\|{\varphi}\|_1
\end{equation}
given that $r_{\psi}+v=D-K\kappa-X\varphi$ where $r_{\psi i}:=\sum_{\ell=1}^{p_z}r_{\ell}(Z_{i\ell})$.
In the following, we make use of this basic inequality and further establish the convergence rate of the proposed estimator. The proof consists of two steps.\\
{\bf Step 1: Deduce a more convenient basic inequality.} Note that
\begin{equation}
\frac{1}{n}r_\psi^{\top}\left(K(\kappa-\widehat{\kappa})+X({\varphi}-\widehat{\varphi})\right) \leq R_{D1} \frac{1}{\sqrt{n}}\|K(\kappa-\widehat{\kappa})+X({\varphi}-\widehat{\varphi})\|_2,
\end{equation}
where \begin{equation}
\begin{aligned}
R_{D1}=\sqrt{\frac{1}{n}\sum_{i=1}^{n}r_{\psi i}^2} &\lesssim M^{-\gamma}.
\end{aligned}
\end{equation}
By the inequality $ab\leq a^2+ \frac{b^2}{4}$, we have
\begin{equation}
\begin{aligned}
\frac{1}{n}r_\psi^{\top}\left(K(\kappa-\widehat{\kappa})+X({\varphi}-\widehat{\varphi})\right)\leq R_{D1}^2+ \frac{1}{4{n}}\|K(\kappa-\widehat{\kappa})+X({\varphi}-\widehat{\varphi})\|_n^2
\end{aligned}
\end{equation}
Proposition (ref) implies that $$ \Pr\left(\left\|\frac{1}{n}X^{\top}v\right\|_{\infty}\geq \dfrac{c_0\lambda_D}{2}\right) \leq c(pM)^{-4} $$ for some $0<c_0<1$ and $\lambda_D = C_\lambda\sqrt{M\log p/n}$ for some $C_\lambda$ large enough, and hence we have with probability at least $1 - c(pM)^{-4}$
\begin{equation}
\frac{1}{n}X^{\top}v({\varphi}-\widehat{\varphi})\leq \dfrac{c_0\lambda_D}{2} \|{\varphi}-\widehat{\varphi}\|_1
\end{equation}
By the uniform boundedness of $v_i$, it has a uniformly bounded sub-Gaussian norm (conditional on $Z$). Given that $\{v_i\}_{i\in[n]}$ are still i.i.d. conditional on $Z$, there exist some absolute constants $c$ and $C$ such that for any $t>0$
\[\Pr\left( \left|\sum_{i=1}^n K_{i\ell}v_i\right| > t \Bigg|Z \right)\leq C\exp\left(-\dfrac{ct^2}{\sum_{i=1}^n K_{i\ell}^2}\right).\]
Note that $f(x)=\exp(-1/x)$ is concave when $x>1/2$. By Jensen's inequality,
\[\begin{aligned}
{\mathbb{E}}\left(\exp\left(-\dfrac{ct^2}{\sum_{i=1}^n K_{i\ell}^2}\right)\textbf{1}\left[\sum_{i=1}^n K_{i\ell}^2>ct^2/2\right]\right) &= {\mathbb{E}}\left(\exp\left(-\dfrac{ct^2}{\sum_{i=1}^n K_{i\ell}^2\textbf{1}\left[\sum_{i=1}^n K_{i\ell}^2>ct^2/2\right]}\right)\right) \\
&\leq \exp\left(-\dfrac{ct^2}{{\mathbb{E}}\left(\sum_{i=1}^n K_{i\ell}^2\textbf{1}\left[\sum_{i=1}^n K_{i\ell}^2>ct^2/2\right]\right)}\right) \\
&\lesssim C\exp\left(-\dfrac{Mct^2}{n}\right).
\end{aligned}\]
Then
\[\begin{aligned}
\Pr\left( \left|\sum_{i=1}^n K_{i\ell}v_i\right| > t \right) &\leq \mathbb{E}\left[\Pr\left( \left|\sum_{i=1}^n K_{i\ell}v_i\right| > t \Bigg|Z \right)\textbf{1}\left[\sum_{i=1}^n K_{i\ell}^2>ct^2/2\right]\right] + \Pr\left(\sum_{i=1}^n K_{i\ell}^2>ct^2/2\right) \\
&\leq C\exp\left(-\dfrac{ct^2}{\sum_{i=1}^n \mathbb{E}(K_{i\ell}^2)}\right) + C(pM)^{-4} \leq C\exp\left(-\dfrac{Mct^2}{n}\right) + C(pM)^{-4}.
\end{aligned}\]
Let $t=\sqrt{3n\log (pM)/(Mc)}$. Then by union bound
\[\Pr\left( \max_{\ell\in[p_zM]}n^{-1}\left|\sum_{i=1}^n K_{i\ell}v_i\right| > \sqrt{\dfrac{3}{c}\cdot\dfrac{\log (pM)}{nM}} \right) \leq C(pM)^{-4}.\]
Then with probability at least $1-C(pM)^{-4}$,
\begin{equation}
\|K^{\top}v\|_2^2 \lesssim M \|K^\top v\|_\infty^2 \lesssim \sqrt{\log (pM)}.
\end{equation}
By the inequality $\frac{1}{n}K^\top v(\kappa-\widehat{\kappa})\leq \frac{1}{n}\|K^\top v\|_2\|\kappa -\widehat{\kappa}\|_2$, we further obtain
\begin{equation}
{\mathbb{P}}\left(\left|\frac{1}{n}K^\top v( \kappa -\widehat{\kappa })\right|\geq R_{D2} \|\kappa-\widehat{\kappa}\|_2\right)\leq {\mathbb{P}}\left( \frac{1}{n}\|K^\top v\|_2\geq R_{D2} \right)\lesssim (pM)^{-4},
\end{equation}
where \begin{equation}
R_{D2} \asymp \sqrt{\dfrac{\log (pM)}{n}}.
\end{equation}
By plugging the upper bounds (ref), (ref) and (ref) into the basic inequality (ref), we conclude that with probability at least $1-2(pM)^{-4}$,
\begin{equation}
\begin{aligned}
&\frac{1}{2n}\|K(\kappa - \widehat\kappa)+X({\varphi}-\widehat{\varphi})\|_n^2+(1-c_0){\lambda_D}\|(\widehat{\varphi}-{\varphi})_{\mathcal{S}_\varphi^c}\|_1 \\ \leq&
2R_{D1}^2+2R_{D2}\|\widehat{\kappa}-\kappa\|_2+(1+c_0)\lambda_D\|(\widehat{\varphi} - \varphi)_{\mathcal{S}_\varphi}\|_1,
\end{aligned}
\end{equation}
with $\mathcal{S}_\varphi=\{j\in[p]:\varphi_j\neq 0\}$.
{\bf Step 2: Establish “Restricted Eigenvalue" type concentration.} In the following, we establish concentration bounds for $\frac{1}{2n}\|K(\kappa-\widehat{\kappa})+X({\varphi}-\widehat{\varphi})\|_2^2$.
We first consider the case
\begin{equation}
2R_{D2}\|\kappa-\widehat{\kappa}\|_2+(1+c_0)\lambda_D\|({\varphi}-\widehat{\varphi})_{\mathcal{S}_\varphi}\|_1 \leq c_1 R_{D1}^2.
\end{equation} for some positive constant $c_1>0$. Then by (ref) and (ref)
\begin{equation}
\|\varphi - \widehat\varphi\|_1 = \|(\varphi - \widehat\varphi)_{\mathcal{S}_\varphi}\|_1 + \|(\varphi - \widehat\varphi)_{\mathcal{S}_\varphi^c}\|_1 \leq \left(\dfrac{2+c_1}{(1-c_0)} + \dfrac{1}{1-c_0} \right) \dfrac{R_{D1}^2}{\lambda_D}.
\end{equation}
Besides, (ref) also implies that
\begin{equation}
\|\kappa - \widehat\kappa\|_2 \lesssim \dfrac{c_1R_{D1}^2}{R_{D2}}.
\end{equation}
We further provide the following lemma about the $L_2$ norm of the estimation error of $\widehat\varphi$.
\begin{Lemma}Under the conditions for Proposition (ref), when ((ref)) holds, we have with probability at least $1 - c(pM)^{-4}$ such that
\begin{equation}
\|\widehat\varphi-\varphi\|_2^2 \lesssim R_{D1}^2.
\end{equation}
\end{Lemma}
When (ref) does not hold, then
\begin{equation}
2R_{D2}\|\widehat{\kappa}-\kappa\|_2+(1+c_0)\lambda_D\|(\widehat{\varphi}-{\varphi})_{\mathcal{S}_\varphi}\|_1 > c_1 R_{D1}^2.
\end{equation}
And (ref) implies
\begin{equation}
\begin{aligned}
&\frac{1}{2n}\|K(\kappa-\widehat{\kappa})+X({\varphi}-\widehat{\varphi})\|_n^2+(1-c_0){\lambda_D}\|({\varphi}-\widehat{\varphi})_{\mathcal{S}_\varphi^c}\|_1\\
&\leq \frac{4+2c_1}{c_1}R_{D2}\|\widehat\kappa-\kappa\|_2+\left(1+c_0+\dfrac{2}{c_1}\right)\lambda_D\|(\widehat{\varphi}-{\varphi})_{\mathcal{S}_\varphi}\|_1.
\end{aligned}
\end{equation}
In this case, we consider the restricted parameter space,
\begin{equation}
\mathcal{C}_D=\left\{\delta= \left(\begin{array}{c}
\widehat{\varphi}-\varphi \\
\widehat{\kappa}-\kappa
\end{array}\right): \|(\widehat{\varphi}-{\varphi})_{\mathcal{S}_\varphi^c}\|_1\leq \frac{c_2R_{D2}}{\lambda_D}\|\widehat{\kappa}-\kappa\|_2 + c_3\|(\widehat{\varphi}-{\varphi})_{\mathcal{S}_\varphi}\|_1\right\}
\end{equation}
where $c_2=\frac{4+2c_1}{(1-c_0)c_1}$ and $c_3=\frac{(c_1+c_0c_1+2)}{(1-c_0)c_1}$.
\begin{Lemma} Suppose the conditions in Theorem (ref) hold. Then w.p.a.1
\begin{equation}
\sup_{\widehat{\varphi}-\varphi\in \mathcal{C}_D} \frac{\frac{1}{n}\|K(\kappa-\widehat{\kappa})+X({\varphi}-\widehat{\varphi})\|_2^2}{M^{-1}\|\kappa-\widehat{\kappa}\|_2^2+\|{\varphi}-\widehat{\varphi}\|_2^2} \geq 2c,
\end{equation}
for some universal positive constant $c$.
\end{Lemma}
By Lemma (ref), we have with probability at least $1-c(pM)^{-4}$
\begin{equation}
\frac{1}{2n}\|K(\kappa-\widehat{\kappa})+X({\varphi}-\widehat{\varphi})\|_n^2 \geq {cM^{-1}\|\kappa-\widehat{\kappa}\|_2^2+c\|{\varphi}-\widehat{\varphi}\|_2^2}.
\end{equation}
By adding both sides of (ref) with $(1-c_0)\lambda_D\|(\widehat{\varphi}-{\varphi})_{\mathcal{S}_\varphi}\|_1$, we have
\begin{equation}
\begin{aligned}
\ &\dfrac{1}{2n}\|K(\kappa-\widehat\kappa) + X(\varphi - \widehat \varphi)\|_2^2+2(1-c_0)\lambda_D\|\widehat{\varphi}-{\varphi}\|_1 \\
\leq &2R_{D1}^2 + 2R_{D2}\|\widehat{\kappa}-\kappa\|_2 + 2\lambda_D\|(\widehat{\varphi}-{\varphi})_{\mathcal{S}_\varphi}\|_1.
\end{aligned}
\end{equation}
Since $\frac{1}{2c}a^2+\frac{c }{2}b^2\geq ab$ for any $a,b>0$, we have
\begin{equation}
2R_{D2}\|\widehat{\kappa}-\kappa\|_2\leq \frac{2M}{c}R_{D2}^2+\frac{c}{2M}\|\widehat{\kappa}-\kappa\|_2^2,
\end{equation}
and
\begin{equation}
\begin{aligned}
2\lambda_D\|(\widehat{\varphi}-{\varphi})_{\mathcal{S}_\varphi}\|_1&\leq 2\sqrt{s}\lambda_D \|(\widehat{\varphi}-{\varphi})_{\mathcal{S}_\varphi}\|_2 \\
&\leq \frac{8}{c} {s}\lambda_D^2+\frac{c}{2} \|\widehat{\varphi}-{\varphi}\|_2^2.
\end{aligned}
\end{equation}
By plugging (ref) and (ref) into (ref), we have
\begin{equation}
\dfrac{1}{4n}\|K(\kappa-\widehat\kappa) + X(\varphi - \widehat\varphi)\|_2^2+(1-c_0)\lambda_D\|\widehat{\varphi}-{\varphi}\|_1 - \dfrac{c}{2M}\|\widehat\kappa-\kappa\|_2^2 - \dfrac{c}{2}\|\widehat\varphi - \varphi\|_2^2 \lesssim R_{D1}^2 + MR_{D2}^2 + s\lambda_D^2.
\end{equation}
By (ref) and (ref),
\begin{equation}
\dfrac{c}{2M}\|\kappa - \widehat\kappa\|_2^2 + c\|\widehat\varphi - \varphi\|_2^2 + (1-c_0)\lambda_D\|\widehat\varphi-\varphi\|_1 \lesssim R_{D1}^2 + M R_{D2}^2 + s\lambda_D^2.
\end{equation}
Note that the rates of $R_{D1}$, $R_{D2}$ and $\lambda_D$ are specified in ((ref)), ((ref)) and the statement of Proposition (ref), respectively. Then ((ref)) follows ((ref)) and ((ref)); ((ref)) and ((ref)) follow ((ref)), Lemma (ref) and ((ref)).
CorollaryRecall that $\widehat v_i = D_i - K_{i\cdot}^\top\widehat\kappa - X_{i\cdot}^\top\widehat\varphi$. Suppose that Assumptions (ref)-(ref) hold.
There exist some constants $C$ and $c$ such that with probability at least $1-c(pM)^{-4}$
\begin{equation}
n^{-1} \sum_{i=1}^n \left(\widehat v_i - v_i\right)^2 \leq C\left( \dfrac{(s+M)\log (pM)}{n}+ M^{-2\gamma}\right),
\end{equation}
and
\begin{equation}
\begin{aligned}
\sup_{i\in[n]}\left|\widehat v_i - v_i \right| &\leq C\left(\sqrt{\dfrac{n}{\log (pM)}}M^{-2\gamma} + (s+M)\sqrt{\dfrac{\log (pM)}{n}}+ M^{-\gamma}\right).
\end{aligned}
\end{equation}
In addition, when $\widehat\kappa$ and $\varphi$ are independent of $D_i,K_{i\cdot},X_{i\cdot}$ and $v_i$,
\begin{equation}
{\mathbb{E}}_{\mathcal{L}}\left(\widehat v_i - v_i\right)^2 \leq C\left( \dfrac{(s+M)\log (pM)}{n}+ M^{-2\gamma}\right)
\end{equation}
with probability at least $1-c(pM)^{-4}$.
proof[Proof of Corollary (ref)] Note that
\begin{equation}
\begin{aligned}
n^{-1}\sum_{i=1}^n(\widehat v_i - v_i)^2 &\lesssim n^{-1}\sum_{i=1}^n\left(K_{i\cdot}^\top (\widehat\kappa-\kappa) + X_{i\cdot}^\top(\widehat\varphi - \varphi)\right)^2 + n^{-1}\sum_{i=1}^n r_{\psi i}^2\\
&\lesssim n^{-1}\sum_{i=1}^n\left(K_{i\cdot}^\top (\widehat\kappa-\kappa) + X_{i\cdot}^\top(\widehat\varphi - \varphi)\right)^2 +M^{-2\gamma},\\
\end{aligned}
\end{equation}
and
\begin{equation}
\begin{aligned}
&\ \ \ \ n^{-1}\sum_{i=1}^n\left(K_{i\cdot}^\top (\widehat\kappa-\kappa) + X_{i\cdot}^\top(\widehat\varphi - \varphi)\right)^2 \\
&\lesssim \|\widehat\kappa-\kappa\|_2^2\cdot \left\|\dfrac{\sum_{i=1}^n H_{i\cdot} H_{i\cdot}^\top}{n}-{\mathbb{E}}\left(H_{i\cdot} H_{i\cdot}^\top\right)\right\|_2 + \|\widehat\varphi-\varphi\|_1^2\left\|\dfrac{\sum_{i=1}^n X_{i\cdot} X_{i\cdot}^\top}{n} - {\mathbb{E}}(X_{i\cdot} X_{i\cdot}^\top)\right\|_\infty \\
&\ \ \ \ + \|\widehat\kappa-\kappa\|_2^2\cdot \left\| {\mathbb{E}}\left(H_{i\cdot} H_{i\cdot}^\top\right)\right\|_2 + \|\widehat\varphi - \varphi\|_2^2 \cdot \left\| {\mathbb{E}}(X_{i\cdot} X_{i\cdot}^\top)\right\|_2.
\end{aligned}
\end{equation}
Following similar arguments for ((ref)), we can show that \[\left\|\dfrac{\sum_{i=1}^n H_{i\cdot} H_{i\cdot}^\top}{n}-{\mathbb{E}}\left(H_{i\cdot} H_{i\cdot}^\top\right)\right\|_2 \lesssim \sqrt{(nM)^{-1}\log (pM)}\]
with probability at least $1-C(pM)^{-4}$.
Besides, note that $\|{\mathbb{E}}(H_{i\cdot} H_{i\cdot}^\top)\|_2\lesssim M^{-1}$ and $\|{\mathbb{E}}(X_{i\cdot} X_{i\cdot}^\top)\|_\infty \lesssim 1$. Together with Propositions (ref) and (ref), with probability at least $1-c(pM)^{-4}$,
\begin{equation}
\begin{aligned}
&\ \ \ \ n^{-1}\sum_{i=1}^n\left(K_{i\cdot}^\top (\widehat\kappa-\kappa) + X_{i\cdot}^\top(\widehat\varphi - \varphi)\right)^2 \\
&\lesssim \|\widehat\kappa - \kappa\|_2^2\cdot\left(\sqrt{\dfrac{\log (pM)}{nM}} + \dfrac{1}{M}\right) + \|\widehat\varphi - \varphi\|_1^2 \cdot \sqrt{\dfrac{\log (pM)}{n}} + \|\widehat\varphi - \varphi\|_2^2 \\
&\lesssim \|\widehat\kappa - \kappa\|_2^2\cdot\dfrac{1}{M} + \|\widehat\varphi - \varphi\|_1^2 \cdot \sqrt{\dfrac{\log (pM)}{n}} + \|\widehat\varphi - \varphi\|_2^2 \\
& \lesssim \dfrac{n}{\log (pM)}M^{-4\gamma - 1} + \dfrac{(s+M)\log (pM)}{n} \\ &\lesssim M^{-2\gamma} + \dfrac{(s+M)\log (pM)}{n},
\end{aligned}
\end{equation}
where the last step applies $n = o[M^{2\gamma + 1}]$. Thus by ((ref)),
\[n^{-1}\sum_{i=1}^n (\widehat v_i - v_i)^2 \lesssim M^{-2\gamma} + \dfrac{(s+M)\log (pM)}{n}.\]
Besides,
\[\begin{aligned}
\sup_{i\in[n]} |\widehat v_i - v_i|
&\lesssim \|\widehat\kappa-\kappa\|_2\cdot\sup_{i\in[n]}\|K_{i\cdot}\|_2 + \|\widehat\varphi-\varphi\|_1\cdot \sup_{i\in[n]}\|X_{i\cdot}\|_\infty + \sup_{i\in[n]}|r_{\psi i}|\\
&\lesssim \sqrt{\dfrac{n}{\log (pM)}}M^{-2\gamma} + (s+M)\sqrt{\dfrac{\log (pM)}{n}} + M^{-\gamma}.
\end{aligned}\]
Finally,
\[\begin{aligned}
{\mathbb{E}}_{\mathcal{L}}(\widehat v_i - v_i)^2 &\lesssim \|\widehat\kappa - \kappa\|_2^2\cdot\|{\mathbb{E}}(H_{i\cdot} H_{i\cdot}^\top)\|_2 + \|\widehat\varphi - \varphi\|_2^2\cdot\|{\mathbb{E}}(X_{i\cdot} X_{i\cdot}^\top)\|_2 + {\mathbb{E}}[r_{\psi i}^2] \\
&\lesssim \|\widehat\kappa - \kappa\|_2^2\cdot M^{-1}+ \|\widehat\varphi - \varphi\|_2^2 + M^{-2\gamma},
\end{aligned}\]
where the last inequality applies Proposition (ref) and the bounded eigenvalues of ${\mathbb{E}}(X_{i\cdot} X_{i\cdot}^\top)$. Then ((ref)) holds by
Proposition (ref). This completes the proof of Corollary (ref).
RemarkDefine \begin{equation}
\widehat{\mathcal{V}}_n := \left\{\sup_{i\in[n]}\left|\widehat v_i - v_i \right| \leq C\left(\sqrt{\dfrac{n}{\log (pM)}}M^{-2\gamma} + (s+M)\dfrac{\log (pM)}{n}+ M^{-\gamma}\right)\right\}.
\end{equation}
When $M \asymp n^\nu$ with $\nu\in [\frac{1}{2\gamma+1},\frac{1}{2\gamma})$ and $s^2\log p = o(n)$, ((ref)) implies $\sup_{i\in[n]}|\widehat v_i - v_i| = o(1)$ with probability at least $1-c(pM)^{-4}$. Consequently, $\widehat{\mathcal{V}}_n$ implies $\{\widehat v_i\in [a_v-\epsilon_v,b_v+\epsilon_v]\text{ for all }i\in[n]\}$. By Corollary (ref),
\begin{equation}
\Pr\left[\widehat{\mathcal{V}}_n\right] \geq 1-c(pM)^{-4}.
\end{equation}
By Assumption (ref), it is easy to deduce that, when $\widehat{\mathcal{V}}_n$ holds,
\[|q(v_i) - q(\widehat v_i)| \leq \sup_v|q^\prime(v)|\cdot|v_i - \widehat v_i| \lesssim |v_i - \widehat v_i|,\]
and likewise,
\[|q^\prime(v_i) - q^\prime(\widehat v_i)| \leq \sup_v|q^{\prime\prime}(v)|\cdot|v_i - \widehat v_i| \lesssim |v_i - \widehat v_i|.\]
Thus, when $\widehat{\mathcal{V}}_n$ holds,
\begin{equation}
n^{-1}\sum_{i=1}^n\left(q(v_i) - q(\widehat v_i)\right)^2 + n^{-1}\sum_{i=1}^n\left(q^\prime(v_i) - q^\prime(\widehat v_i)\right)^2 \lesssim \dfrac{(s+M)\log (pM)}{n}+ M^{-2\gamma},
\end{equation}
and
\begin{equation}
\sup_{i\in[n]}\left|q(v_i) - q(\widehat v_i)\right| + \sup_{i\in[n]}\left|q^\prime(v_i) - q^\prime(\widehat v_i)\right| \lesssim \sqrt{\dfrac{n}{\log (pM)}}M^{-2\gamma} + (s+M)\sqrt{\dfrac{\log (pM)}{n}} + M^{-\gamma}.
\end{equation}
CorollarySuppose that Assumptions (ref)-(ref) hold.
There exists some constant $c$ such that with probability at least $1-c(pM)^{-4}$
\begin{equation}
n^{-1} \sum_{i=1}^n \sum_{j=1}^{M} (H_j(v_i)-H_j(\widehat v_i))^2 \lesssim M^2\left( \dfrac{(s+M)\log (pM)}{n}+ M^{-2\gamma}\right),
\end{equation}
and,
\begin{equation}
\begin{aligned}
\sup_{i\in[n]} \sqrt{\sum_{j=1}^M (H_j^\prime (v_i)-H_j^\prime (\widehat v_i))^2} \lesssim M^2\left(\sqrt{\dfrac{n}{\log (pM)}}M^{-2\gamma} + \sqrt{\dfrac{(s+M)^2\log (pM)}{n}} + M^{-\gamma}\right).
\end{aligned}
\end{equation}
Furthermore, when $\widehat\kappa$ and $\widehat\varphi$ are independent of $D_i,K_{i\cdot},X_{i\cdot}$ and $v_i$, with probability at least $1-c(pM)^{-4}$
\begin{equation}
{\mathbb{E}}_{\mathcal{L}}\sum_{j=1}^{M} (H_j(v_i)-H_j(\widehat v_i))^2 \lesssim M^2\left( (s+M)\dfrac{\log (pM)}{n}+ M^{-2\gamma}\right) + \epsilon_n,
\end{equation}
for some $\epsilon_n = O_p[(pM)^{-2}]$.
proof[Proof of Corollary (ref)]
By the mean value theorem, we know that for any $i\in[n]$ and $j\in[M]$, there exists a $v_{ij}^*$ between $v_i$ and $\widehat{v}_i$ such that
\[ |H_j(v_i) - H_j(\widehat{v}_i)| = |H^\prime(v_i^*)|\cdot \left|v_i-\widehat{v}_i\right|.\]
As mentioned in Remark (ref), $\widehat{\mathcal{V}}_n$ implies $\widehat v_i\in [a_v-\epsilon_v,b_v+\epsilon_v]\text{ for all }i\in[n]$ and hence $\{v_i^*\}_{i\in[n]}$ also fall in this compact interval.
Consequently, when $\widehat{\mathcal{V}}_n$ in ((ref)) holds
\begin{equation}\begin{aligned}
\sum_{j=1}^{M}\left(H_j(v_i) - H_j(\widehat{v}_i)\right)^2 &= \sum_{j=1}^{M}\left(v_i-\widehat{v}_i\right)^2\left(H_j^\prime(v_{i}^*)\right)^2 \lesssim M^2(v_i-\widehat v_i)^2,\\
\end{aligned}
\end{equation}
where the second inequality applies Proposition (ref). Thus, $\sum_{i=1}^n\sum_{j=1}^{M}\left(H_j(v_i) - H_j(\widehat{v}_i)\right)^2 \lesssim M^2\sum_{i=1}^n\left(v_i-\widehat{v}_i\right)^2$. Similarly, using Proposition (ref) for second-order derivatives of splines,
\[\begin{aligned}
\sup_{i\in[n]}\sum_{j=1}^{M}\left(H_j^\prime(v_i) - H_j^\prime(\widehat{v}_i)\right)^2 &= \sup_{i\in[n]}\sum_{j=1}^{M}\left(v_i-\widehat{v}_i\right)^2\left(H_j^{\prime\prime}(v_{i}^*)\right)^2 \lesssim M^4\left(v_i-\widehat{v}_i\right)^2.
\end{aligned}\]
Then ((ref)) and ((ref)) hold by ((ref)) and ((ref)) in Corollary (ref). As for ((ref)),
\[\begin{aligned}
{\mathbb{E}}_{\mathcal{L}}\sum_{j=1}^{M} (H_j(v_i)-H_j(\widehat v_i))^2 &= {\mathbb{E}}_{\mathcal{L}}\left[1(\widehat{\mathcal{V}}_n)\sum_{j=1}^{M} (H_j(v_i)-H_j(\widehat v_i))^2\right] + {\mathbb{E}}_{\mathcal{L}}\left[1(\widehat{\mathcal{V}}_n^c)\sum_{j=1}^{M} (H_j(v_i)-H_j(\widehat v_i))^2\right] \\
&\lesssim M^2 {\mathbb{E}}_{\mathcal{L}}(\widehat v_i - v_i)^2 + M{\rm Pr}(\widehat{\mathcal{V}}^c|\mathcal{L}),
\end{aligned}\]
where the first term on the RHS applies the ((ref)) when $\widehat{\mathcal{V}}_n$ holds, and the second term applies the boundness of $H_j$. It suffices to show
\begin{equation}
{\rm Pr}(\widehat{\mathcal{V}}^c|\mathcal{L}) = O_p((pM)^{-4}).
\end{equation}
By Markov inequality, for any $C>0$,
\[\begin{aligned}
\Pr\left[ {\rm Pr}(\widehat{\mathcal{V}}^c|\mathcal{L}) > C(pM)^{-4} \right] \leq \dfrac{{\mathbb{E}}\left[{\rm Pr}(\widehat{\mathcal{V}}^c|\mathcal{L})\right]}{C(pM)^{-4} } = \dfrac{{\rm Pr}(\widehat{\mathcal{V}}^c)}{C(pM)^{-4}} = \dfrac{O(1)}{C},
\end{aligned}\]
and hence $\lim_{C\to\infty}\Pr\left[ {\rm Pr}(\widehat{\mathcal{V}}^c|\mathcal{L}) > C(pM)^{-4} \right] = 0$. Then
\[\begin{aligned}
{\mathbb{E}}_{\mathcal{L}}\sum_{j=1}^{M} (H_j(v_i)-H_j(\widehat v_i))^2
&\lesssim M^2 {\mathbb{E}}_{\mathcal{L}}(\widehat v_i - v_i)^2 + (pM)^{-2},
\end{aligned}\]
and hence ((ref)) holds by ((ref)).
CorollaryFor all $1\leq i\leq n$ we have with probability at least $1-c(pM)^{-4}$,
\[ \dfrac{1}{\sqrt{k}} \leq \|\widehat H_{i\cdot}\|_2 \leq 1.\]
proof[Proof of Corollary (ref)]
The result holds by the proof of Proposition (ref) when $\widehat{\mathcal{V}}_n$ in ((ref)) holds.
Proof of Theorem (ref)
As a preparatory result for honesty of the confidence band, we will show that the high probability error bounds hold uniformly for all $g\in\mathcal{H}_{\mathcal{D}}(\gamma,L)$. We abbreviate $\inf_{g\in\mathcal{H}_{\mathcal{D}}(\gamma,L)}$ and $\sup_{g\in\mathcal{H}_{\mathcal{D}}(\gamma,L)}$ as $\inf_g$ and $\sup_g$ in the remainder of the proofs. Specifically, we will show
equation[equation omitted — 180 chars of source]
equation[equation omitted — 165 chars of source]
Throughout the proof, only $\widehat\omega$, $\widehat\theta$ and the spline approximation error of function $g$, denoted by $r_g$, are dependent on the function $g$, and thus the uniformity over $g$ needs to be taken care of only for the events dependent on these estimators and approximation errors.
Recall that $\omega=(\beta^\top,\eta^\top)^\top$ and $\widehat\omega$ denotes the LASSO estimator. By the definition of LASSO estimators $\widehat{\omega}$ and $\widehat{\theta}$, we have
equation[equation omitted — 192 chars of source]
which implies
equation[equation omitted — 295 chars of source]
given that $r+\varepsilon=y-\widehat{W}\omega-X\theta$ where $r=(r_1,r_2,\cdots,r_n)^\top$ with $r_i$ defined in ((ref)).
In the following, we use this basic inequality and further establish the convergence rate of the proposed estimator. The proof consists of two steps.\\
{\bf Step 1: Deduce a more convenient inequality.} We analyze the terms in (ref) and obtained a more convenient version of (ref). Note that
equation[equation omitted — 213 chars of source]
where
equation[equation omitted — 451 chars of source]
with probability at least $1-c(pM)^{-4}$, where the last inequality applies Assumption (ref), Proposition (ref), and Remark (ref). By the inequality $ab\leq a^2+ \frac{b^2}{4}$, we have
equation[equation omitted — 263 chars of source]
Proposition (ref) implies that $\left\|\frac{1}{n}X^{\top}\varepsilon\right\|_{\infty}\leq \dfrac{c_0\lambda_Y}{2} $ for $0<c_0<1$, and hence we have w.p.a.1
equation[equation omitted — 153 chars of source]
for any $\widehat\theta$. Furthermore,
\[
aligned\begin{aligned}
{\mathbb{E}}\|B^{\top}\varepsilon\|_2^2={\mathbb{E}}\sum_{j=1}^{M}\left(\sum_{i=1}^{n} \varepsilon_i B_{ij}\right)^2&=\sum_{j=1}^{M}\sum_{i=1}^{n} {\mathbb{E}} (\varepsilon_i^2 B_{ij}^2)\\
&=\sum_{j=1}^{M}\sum_{i=1}^{n} {\mathbb{E}} ({\mathbb{E}}(\varepsilon_i^2 B_{ij}^2|X,Z,v)) \\
&=\sum_{j=1}^{M}\sum_{i=1}^{n} {\mathbb{E}} (B_{ij}^2{\mathbb{E}}(\varepsilon_i^2 |X,Z,v)) \\
&= \sigma_\varepsilon^2 \sum_{i=1}^{n} {\mathbb{E}} (\sum_{j=1}^{M}B_{ij}^2) \leq nk\sigma_\varepsilon^2,
\end{aligned}
\]
where the last inequality applies Proposition (ref). In a similar manner we deduce that
\[
aligned\begin{aligned}
{\mathbb{E}}\left[\|\widehat{H}^{\top}\varepsilon\|_2^2\right] \leq n k \sigma_\varepsilon^2.
\end{aligned}
\]
Note that $\|\widehat{W}^{\top}\varepsilon\|_2^2 \leq 2\|B^{\top}\varepsilon\|_2^2 + 2\|\widehat{H}^{\top}\varepsilon\|_2^2$. Then by Markov inequality we deduce that
equation[equation omitted — 293 chars of source]
and thus
equation[equation omitted — 137 chars of source]
where
equation[equation omitted — 117 chars of source]
By ((ref)) and the inequality $\frac{1}{n}\widehat{W}^{\top}\varepsilon(\omega-\widehat{\omega})\leq \frac{1}{n}\|\widehat{W}^{\top}\varepsilon\|_2\|\omega-\widehat{\omega}\|_2$, we further obtain w.p.a.1
equation[equation omitted — 129 chars of source]
for any $\widehat\omega$. By plugging the upper bounds (ref), (ref) and (ref) into the basic inequality (ref), we conclude that w.p.a.1, for any $\widehat\theta$ and $\widehat\omega$
equation[equation omitted — 357 chars of source]
with $\mathcal{S}=\{j\in[p]:\theta_j\neq 0\}$.
{\bf Step 2: Establish restricted eigenvalue-type concentration.} In the following, we establish concentration bounds for $\frac{1}{2n}\|W(\omega-\widehat{\omega})+X({\theta}-\widehat{\theta})\|_2^2$.
We first consider the case
equation[equation omitted — 157 chars of source]
for some positive constant $c_1>0$. It follows from (ref) that
equation[equation omitted — 213 chars of source]
Together with (ref), we have
equation[equation omitted — 255 chars of source]
and
equation[equation omitted — 102 chars of source]
When (ref) does not hold, then
equation[equation omitted — 154 chars of source]
And (ref) implies
equation[equation omitted — 380 chars of source]
In this case, we consider the restricted parameter space,
equation[equation omitted — 338 chars of source]
where $c_2=\frac{4+2c_1}{(1-c_0)c_1}$ and $c_3=\frac{(c_1+c_0c_1+2)}{(1-c_0)c_1}$.
LemmaSuppose the conditions in Theorem (ref) hold. Then w.p.a.1
\begin{equation}
\sup_{ \delta \in \mathcal{C}_0} \frac{\frac{1}{n}\|\widehat{W}(\omega-\widehat{\omega})+X({\theta}-\widehat{\theta})\|_2^2}{M^{-1}\|\omega-\widehat{\omega}\|_2^2+\|{\theta}-\widehat{\theta}\|_2^2} \geq 2c
\end{equation}
for some universal positive constant $c$.
By adding both sides with $(1-c_0)\lambda_Y\|(\widehat{\theta}-{\theta})_\mathcal{S}\|_1$ to (ref), we have
equation[equation omitted — 283 chars of source]
Since $\frac{1}{2(c )}a^2+\frac{c }{2}b^2\geq ab$ for any $a,b>0$, we have
equation[equation omitted — 139 chars of source]
and
equation[equation omitted — 367 chars of source]
By Lemma (ref),
equation[equation omitted — 213 chars of source]
and thus
\[R_2\|\widehat{\omega}-\omega\|_2 + 2\lambda_Y\|(\widehat{\theta}-{\theta})_\mathcal{S}\|_1 \leq \dfrac{M}{2c}R_2^2 + \dfrac{2}{c}s\lambda_Y^2 + \frac{1}{4n}\|\widehat{W}(\omega-\widehat{\omega})+X({\theta}-\widehat{\theta})\|_2^2.\]
Combining the above with (ref), we have
equation[equation omitted — 248 chars of source]
and using ((ref)) again,
equation[equation omitted — 237 chars of source]
All the inequalities above hold w.p.a.1 for any $\widehat\theta$ and $\widehat\omega$. Hence
equation[equation omitted — 146 chars of source]
and
equation[equation omitted — 159 chars of source]
Recall that the bounds of $R_1$ and $R_2$ are given as (ref) and (ref). By combining (ref) and (ref), we establish (ref) and thus (ref); By combining (ref) and (ref), we establish (ref) and thus (ref).
For (ref), note that for any fixed $d$
\[(\widehat g^\prime(d) - g^\prime(d) )^2 = ((\widehat\beta-\beta)^\top B^\prime(d) - r^\prime_g(d))^2 \leq 2[(\widehat\beta-\beta)^\top B^\prime(d)]^2 + 2[r^\prime_g(d)]^2 \]
and $\sup_g [r^\prime_g(d)]^2 \lesssim_p M^{-2\gamma+2}$ by Proposition (ref). Let $f_{\mathcal{D}}(s)$ denote the probability density function of $D_i$. Then (ref) follows by
\[
aligned\sup_g \int_{\mathcal{D}}(\widehat g^\prime - g^\prime )^2 &\lesssim \sup_g (\widehat\beta-\beta)^\top \int_{\mathcal{D}} B^\prime(s)B^\prime(s)^\top {\rm d}s (\widehat\beta-\beta) + M^{-2\gamma+2}\\
&\leq \sup_g c_f^{-1}(\widehat\beta-\beta)^\top \int_{\mathcal{D}} f_{\mathcal{D}}(s) B^\prime(s)B^\prime(s)^\top {\rm d}s (\widehat\beta-\beta) + M^{-2\gamma+2} \\
&= c_f^{-1}\sup_g (\widehat\beta-\beta)^\top {\mathbb{E}}[B^\prime(D_i)B^\prime(D_i)^\top](\widehat\beta-\beta) + M^{-2\gamma+2} \\
&\lesssim M \cdot \sup_g\|\widehat\beta-\beta\|_2^2 + M^{-2\gamma+2} \\
&\lesssim_p \dfrac{M^2(s+M)\log (pM)}{n} + M^{-2\gamma+2}
\]
where the second row applies Assumption (ref), the fourth row applies Proposition (ref), and the last row applies (ref).
Proof of Proposition (ref)
Starting from this proof, we introduce some new definitions and notations to facilitate the proof of the honesty of the uniform confidence band. For any positive sequences $a_n$ and $b_n$, we use $a_n = o_{u.p.}(b_n)$ to represent the fact that for any $\epsilon>0$, $\sup_g \Pr\left\{a_n/b_n > \epsilon \right\} \to 0$ as $n\to\infty$. The abbrevation “u.p.” means the probability converges uniformly for all functions $g$. Also, $a_n \lesssim_{u.p.} b_n$ means there exists some absolute constant $C$ such that $\inf_g \Pr\left\{a_n \leq Cb_n\right\} \to 1$; $a_n \gtrsim_{u.p.} b_n$ and means $b_n \lesssim_{u.p.} a_n$; $a_n \asymp_{u.p.} b_n$ means $a_n \lesssim_{u.p.} b_n$ and $b_n \lesssim_{u.p.} a_n$. We will show stronger results that
equation[equation omitted — 158 chars of source]
and
equation[equation omitted — 149 chars of source]
Define $q_i^\prime := q^\prime(v_i)$, $\widehat{q}^\prime_i:=\widehat{q}^\prime(\widehat v_i)$. Recall that $H$ is the matrix with the $(i,j)$-th element being $H_{ij} = H_j(v_i)$, $\widetilde X_{i\cdot} := q_i^\prime X_{i\cdot}$, $\widetilde K_{i\cdot} := q_i^\prime K_{i\cdot}$ and $F_{i\cdot} := \left(B_{i\cdot}^\top,H_{i\cdot}^\top,\widetilde K_{i\cdot}^\top, X_{i\cdot}^\top, \widetilde X_{i\cdot}^\top\right)^\top$. Define $\mathbb{E}_{\mathcal{L}}(\cdot) := \mathbb{E}\left[\cdot|\mathcal{L}\right]$. The (conditional) covariance matrices are denoted as $\Sigma_F := {\mathbb{E}}(F_{i\cdot} F_{i\cdot}^\top)$ and $\Sigma_{F|\mathcal{L}}:={\mathbb{E}}_{\mathcal{L}}(\widehat F_{i\cdot} \widehat F_{i\cdot}^\top )$.
We first state the following Lemma about the eigenvalues of $\Sigma_{F|\mathcal{L}}$.
LemmaUnder the conditions of Proposition (ref), we have $$1 \lesssim_p \inf_g \lambda_{\min}(\Sigma_{F|\mathcal{L}}^{-1}) \leq \sup_g \lambda_{\max}(\Sigma_{F|\mathcal{L}}^{-1}) \lesssim_p M. $$
Step 1. Show ((ref)).
Recall the definitions of sub-Gaussian and sub-exponential norms given in ((ref)) and ((ref)).
We first bound the sub-Gaussian norm of $\widehat F_{i\cdot}$.
Note that for all $\beta\in\mathbb{R}^M$ such that $\|\beta\|_2=1$,
\[(\beta^\top B_{i\cdot})^2 \leq \|B_{i\cdot}\|_2 \leq \sum_{j=1}^MB_{ij}=1.\]
Thus, $\|B_{i\cdot}\|_{\psi_2|\mathcal{L}}\lesssim 1$ and similarly, $\|\widehat H_{i\cdot}\|_{\psi_2|\mathcal{L}}, \|(K_\ell)_{i\cdot}\|_{\psi_2|\mathcal{L}}\lesssim 1$. Additionally,
\[
aligned\sup_{i\in[n]}\left|\widehat q^\prime(\widehat v_i)\right| &\leq \|\widehat\eta^{ind}-\eta\|_2\sup_{i\in[n]}\|\widehat H_{i\cdot}^\prime\|_2 + \sup_{i\in[n]}|\eta^\top\widehat H_{i\cdot}^\prime - q^\prime(\widehat v_i)| + \sup_{i\in[n]}|q^\prime(\widehat v_i)| \\ &\lesssim \|\widehat\eta^{ind}-\eta\|_2\cdot M + M^{-\gamma+1} + \sup_{v}|q^\prime(v)| \\
&\lesssim \|\widehat\eta^{ind}-\eta\|_2\cdot M + 1.
\]
Thus,
\[
aligned\|\widehat F_{i\cdot}\|_{\psi_2|\mathcal{L}} &\leq \|B_{i\cdot}\|_{\psi_2|\mathcal{L}} + \|\widehat H_{i\cdot}\|_{\psi_2|\mathcal{L}} +
\sup_{i\in[n]} |\widehat q^\prime(\widehat v_i)|\sum_{\ell = 1}^{p_z}\sup_{j\in[M]}\|(K_\ell)_{i\cdot}\|_{\psi_2|\mathcal{L}} \ + \left(1+\sup_{i\in[n]} |\widehat q^\prime(\widehat v_i)|\right)\cdot\|X_{i\cdot}\|_{\psi_2} \\
&\lesssim \|\widehat\eta^{ind}-\eta\|_2\cdot M + 1
\]
and thus $\|\Sigma_{F|\mathcal{L}}^{-1}\widehat F_{i\cdot}\|_{\psi_2|\mathcal{L}} \lesssim \|\Sigma_{F|\mathcal{L}}^{-1}\|_2 \left(\|\widehat\eta^{\text{ind}}-\eta\|_2\cdot M + 1\right) $. In the same manner we can also deduce $|\widehat F_{i\cdot}|\lesssim \|\widehat\eta^{\text{ind}}-\eta\|_2\cdot M + 1$. Thus, any coordinate of $ \widehat F_{i\cdot} \widehat F_{i\cdot}^\top \Sigma_{F|\mathcal{L}}^{-1}$ then has a sub-Gaussian norm norm (conditionally on $\mathcal{L}$) bounded by $${\rm SG}_{\max} := C\left[\|\widehat\eta^{\text{ind}}-\eta\|_2\cdot M + 1\right]^2 \cdot\|\Sigma_{F|\mathcal{L}}^{-1}\|_2$$
where by (ref) and Lemma (ref)
equation[equation omitted — 107 chars of source]
Then by vershynin2010introduction and union bound, for any $t>0$
\[
aligned&\ \ \ \ \sup_g {\rm \Pr}\left(\|\widehat\Sigma_F\Sigma_{F|\mathcal{L}}^{-1} - I_{p_F}\|_\infty > t\Bigg|\mathcal{L}\right) \\
&= \sup_g {\rm \Pr}\left(\max_{j,k\in[p_F]}\left| \left(\dfrac{1}{n}\sum_{i=1}^n \left(\widehat F_{i\cdot} \widehat F_{i\cdot}^\top \Sigma_{F|\mathcal{L}}^{-1} - {\mathbb{E}}_{\mathcal{L}}\left(\widehat F_{i\cdot} \widehat F_{i\cdot}^\top \Sigma_{F|\mathcal{L}}^{-1}\right)\right)\right)_{jk} \right| > t\Bigg|\mathcal{L}\right) \\
&\leq 2p_F^2 \cdot \exp\left(-cn \cdot \dfrac{t^2}{\sup_g {\rm SG}_{\max}^2 } \right) \lesssim_p 2p_F^2 \cdot \exp\left(-cn \cdot \dfrac{t^2}{ M^2 } \right)
\]
where the last inequality applies (ref). Taking $t = \sqrt{3M^2\log p_F / (cn)}$, we have
\[
aligned\sup_g {\rm \Pr}\left(\|\widehat\Sigma_F\Sigma_{F|\mathcal{L}}^{-1} - I_{p_F}\|_\infty > \sqrt{\dfrac{3M^2\log p_F}{cn}}\Bigg|\mathcal{L}\right)
&\leq 2p_F^2\cdot\exp\left(-3\log p_F\right) \to 0.
\]
In other words, the supreme of conditional probability, as a random variable uniformly bounded in $[0,1]$, is $o_p(1)$. Thus by the Bounded Convergence Theorem,
\[
aligned\sup_g {\rm \Pr}\left(\|\widehat\Sigma_F\Sigma_{F|\mathcal{L}}^{-1} - I_{p_F}\|_\infty >\sqrt{\dfrac{3\cdot M^2\log p_F}{cn}}\right)
&\leq {\mathbb{E}}\left[\sup_g {\rm \Pr}\left(\|\widehat\Sigma_F\Sigma_{F|\mathcal{L}}^{-1} - I_{p_F}\|_\infty >\sqrt{\dfrac{3\cdot M^2\log p_F}{cn}}\Bigg|\mathcal{L}\right) \right]\\
&\to 0
\]
and thus $\|\widehat\Sigma_F\Sigma_{F|\mathcal{L}}^{-1} - I\|_\infty \lesssim_{u.p.} M\sqrt{\dfrac{\log p_F}{n}}$.
Step 2. Show ((ref)). Recall that conditionally on $\mathcal{L}$, $\widehat F_{i\cdot}^\top \Sigma_{F|\mathcal{L}}^{-1}$ has sub-Gaussian norm bounded by ${\rm SG}_{\max}^{(1)}:=C\left[\|\widehat\eta^{\text{ind}} - \eta\|_2\cdot M + 1\right]\cdot \|\Sigma_{F|\mathcal{L}}^{-1}\|_2$ with
\[\sup_g {\rm SG}_{\max}^{(1)} \lesssim_p M\]
following the arguments for (ref). Then for any $t>0$
\[
aligned\sup_g {\rm Pr}\left(\|\widehat F\Sigma_{F|\mathcal{L}}^{-1}\|_\infty > t\Bigg|\mathcal{L}\right) &= \sup_g {\rm Pr}\left(\sup_{i\in[n]}\left\|\Sigma_{F|\mathcal{L}}^{-1}\widehat F_{i\cdot} \right\|_\infty > t\Bigg|\mathcal{L}\right) \\
&\leq n\cdot p_F \exp\left(-\dfrac{ct^2}{\sup_g {\rm SG}_{\max}^{(1)}}\right) \lesssim_p n\cdot p_F \exp\left(-\dfrac{ct^2}{M^2}\right).
\]
Taking $t =2 M\sqrt{\log ( n p_F ) / c}$, by similar arguments we deduce
\[
aligned\sup_g {\rm Pr}\left(\|\widehat F\Sigma_{F|\mathcal{L}}^{-1}\|_\infty \lesssim_p CM\sqrt{\log ( np_F)}\Bigg|\mathcal{L}\right)
&\leq n\cdot p_F \exp\left( -2\log (np_F) \right) \to 0.
\]
Using the Bounded Convergence Theorem again, we deduce
\[\sup_g {\rm Pr}\left(\|\widehat F\Sigma_{F|\mathcal{L}}^{-1}\|_\infty > 2M\sqrt{\log (np_F) / c}\right) \to 0,\]
and hence $n^{-1/2}\|\widehat F\Sigma_{F|\mathcal{L}}^{-1}\|_\infty \lesssim_{u.p.} M\sqrt{\dfrac{\log (n p_F )}{n}} \lesssim M\sqrt{\dfrac{\log (pM)}{n}}$.
Proof of Proposition (ref)
To prepare for the honesty, we will show the results with $o_p$ and $\asymp_p$ replaced by $o_{u.p.}$ and $\asymp_{u.p.}$. Recall that $\widehat m(d) := \widehat\Omega_B B^\prime(d) $. Note that
equation[equation omitted — 270 chars of source]
By Proposition (ref),
\[\sup_g \sup_{d\in\mathcal{D}} \left|\sqrt{n}\left( \sum_{i=1}^{M}\beta_jB_j^\prime(d) - g^\prime(d) \right)\right| = O(\sqrt{n}M^{-\gamma+1}).\]
By the definition of the debiased estimator $\widetilde\beta$ in (ref),
equation*[equation* omitted — 405 chars of source]
where \[\mathcal{Z}(d) :=\dfrac{\widehat{m}(d)^\top \sum_{i=1}^n \widehat F_{i\cdot} \varepsilon_i}{\sqrt{n}},\ \varDelta(d) := \dfrac{\widehat{m}(d)^\top \sum_{i=1}^n \widehat F_{i\cdot} (\widehat\varepsilon_i-\varepsilon_i)}{\sqrt{n}} + \sqrt{n}B^\prime(d)^\top(\widehat\beta-\beta).\]
We then need to prove the following results:
enumerate[(S1)]
• The scale of $\widehat s(d)$
\begin{equation}
\widehat{s}(d) := \sqrt{\widehat m(d)^\top\widehat \Sigma_F\widehat m(d)} \asymp_{u.p.} M^{1.5}.
\end{equation}
• $ \sup_{d\in\mathcal{D}}\dfrac{\sqrt{n}M^{-\gamma+1} + |\varDelta(d)|}{\widehat s(d)} \lesssim_{u.p.} \frac{1}{\log n}$ uniformly for all $d$. The additional $1/\log n$ handles the Gaussian approximation for Theorem (ref).
{\bfProof of (S1)}. By Proposition (ref), the vector $j$-th column of $\Sigma_{F|\mathcal{L}}^{-1}$ belongs to the feasible set of the optimization algorithm ((ref)). Define $\textbf{i}_j$ as the $j$-th standard basis with the $j$-th element being one and others being zero. Then,
equation[equation omitted — 396 chars of source]
where the last step applies Proposition (ref), and
equation[equation omitted — 751 chars of source]
where the first step applies the positive semi-definiteness of $\widehat\Omega_B\widehat\Sigma_F\widehat\Omega_B$, the second step follows by the definition of $\widehat\Omega_B$ and the fact that $\Sigma_{F|\mathcal{L}}^{-1}$ is feasible for ((ref)) with uniformly high probability by (ref) and (ref), and the last step applies Lemma (ref).
It then suffices to show that $$\|\Sigma_{F|\mathcal{L}}^{-1}\widehat \Sigma_F \Sigma_{F|\mathcal{L}}^{-1} - \Sigma_{F|\mathcal{L}}^{-1}\|_\infty = o_{u.p.}(1).$$
We have shown $\|\Sigma_{F|\mathcal{L}}^{-1}\widehat F_{i\cdot}\|_{\psi_2|\mathcal{L}} \lesssim \|\Sigma_{F|\mathcal{L}}^{-1}\|_2\cdot (\|\widehat\eta^\text{ind} - \eta\|_2\cdot M+1)$ in the proof of ((ref)) for Proposition (ref).
Any coordinate of $ \Sigma_{F|\mathcal{L}}^{-1}\widehat F_{i\cdot} \widehat F_{i\cdot}^\top \Sigma_{F|\mathcal{L}}^{-1}$ then has a sub-exponential norm (conditionally on $\mathcal{L}$) bounded by
${\rm SE}_{\max} = C\left[\|\widehat\eta^{\text{ind}}-\eta\|_2\cdot M + 1\right]^2 \cdot\|\Sigma_{F|\mathcal{L}}^{-1}\|_2^2$, and by (ref) and Lemma (ref)
$$\sup_g {\rm SE}_{\max} \lesssim_p \left(o(1) + 1\right)^2M^2 \lesssim M. $$
Since ${\mathbb{E}}_{\mathcal{L}}[\Sigma_{F|\mathcal{L}}^{-1}\widehat \Sigma_F \Sigma_{F|\mathcal{L}}^{-1} ] = \Sigma_{F|\mathcal{L}}^{-1}$, by vershynin2010introduction for centered sup-exponential variables and the union bound, for any $t>0$
\[
aligned&\ \ \ \ \sup_g {\rm \Pr}\left(\|\Sigma_{F|\mathcal{L}}^{-1}\widehat\Sigma_F\Sigma_{F|\mathcal{L}}^{-1} - \Sigma_{F|\mathcal{L}}^{-1}\|_\infty > t\Bigg|\mathcal{L}\right) \\
&\leq 2p_F^2 \cdot \exp\left(-cn\cdot\min\left(\dfrac{t^2}{\sup_g {\rm SE}_{\max}^2},\dfrac{t}{\sup_g {\rm SE}_{\max}}\right) \right) \\
&\lesssim_p 2p_F^2 \cdot \exp\left(-cn\cdot\min\left(\dfrac{t^2}{M^4},\dfrac{t}{M^2}\right) \right)
\]
Taking $t = 2\sqrt{M^4\log p_F / cn}$ with $C_t$ large enough, we have
\[\dfrac{t^2}{M^4} \lesssim \dfrac{\log p_F}{n} \to 0\]
with $n$ large enough, implying that $t^2/M^4 \leq t/M^2$.
Thus,
\[
aligned\sup_g {\rm \Pr}\left(\|\Sigma_{F|\mathcal{L}}^{-1}\widehat\Sigma_F\Sigma_{F|\mathcal{L}}^{-1} - \Sigma_{F|\mathcal{L}}^{-1}\|_\infty > C_t\sqrt{\dfrac{M^4\log p_F}{n}}\Bigg|\mathcal{L}\right)
&\lesssim_p 2p_F^2\cdot\exp\left(-4\log p_F\right) \to 0.
\]
In other words, the supreme of the conditional probability is $o_p(1)$. Thus by the Bounded Convergence Theorem,
\[
aligned&\ \ \ \ \sup_g {\rm \Pr}\left(\|\Sigma_{F|\mathcal{L}}^{-1}\widehat\Sigma_F\Sigma_{F|\mathcal{L}}^{-1} - \Sigma_{F|\mathcal{L}}^{-1}\|_\infty > C_t\sqrt{\dfrac{M^4\log p_F}{n}}\right) \\
&\leq {\mathbb{E}}\left[\sup_g {\rm \Pr}\left(\|\Sigma_{F|\mathcal{L}}^{-1}\widehat\Sigma_F\Sigma_{F|\mathcal{L}}^{-1} - \Sigma_{F|\mathcal{L}}^{-1}\|_\infty > C_t\sqrt{\dfrac{M^4\log p_F}{n}}\Bigg|\mathcal{L}\right) \right] \to 0,
\]
and thus $\|\Sigma_{F|\mathcal{L}}^{-1}\widehat\Sigma_F\Sigma_{F|\mathcal{L}}^{-1} - \Sigma_{F|\mathcal{L}}^{-1}\|_\infty \lesssim_{u.p.} M^2\sqrt{\dfrac{\log p_F}{n}} \lesssim M^2\sqrt{\dfrac{\log (pM)}{n}} = o_p(M)$. Then by ((ref)) we have
equation[equation omitted — 146 chars of source]
Together with ((ref)), we deduce
equation[equation omitted — 118 chars of source]
The proof the other side of the inequality (ref) follows javanmard2014confidence. We construct the following estimator
equation[equation omitted — 255 chars of source]
where $\mu = \max_{j\in[M]}\mu_j$ and $I_B$ is the first $M$ rows of the $p_F$-dimensional identity matrix. Note that for any $M\times p_F$ matrix $\widehat\Omega_{B} = (\widehat\Omega_{Bj})^\top_{j\in[M]}$ belonging to the feasible set of (ref), $\widehat\Omega_{B}^\top B^\prime(d)$ belongs to the feasible set of (ref) due to the fact that
equation[equation omitted — 375 chars of source]
where the last inequality follows from the feasibility of $\widehat\Omega_B$ for (ref). Recall that $\widehat m(d) := \widehat\Omega_B B^\prime(d)$. Hence, we have
equation[equation omitted — 145 chars of source]
It is thus sufficient to establish a lower bound for $(m^{*}(d))^\top\widehat\Sigma_F m^{*}(d)$. Due to the feasibility condition of (ref), we have $-B^\prime(d)^{\top} I_B \widehat\Sigma_F m^{*}(d)+\|B^\prime(d)\|_2^2-\mu\|B^\prime(d)\|_1^2\leq 0$ and hence for any $c_0>0$,
equation[equation omitted — 728 chars of source]
where $\widehat\Sigma_B = n^{-1}B^\top B$.
By Proposition (ref), $\|B^\prime(d)\|_2 \asymp \|B^\prime(d)\|_1 \asymp M$. Thus, $\|B^\prime(d)\|_2^2-\mu\|B^\prime(d)\|_1^2 \geq (1-C\mu)\|B^\prime(d)\|_2^2\geq c\|B^\prime(d)\|_2^2$
where the last step applies the fact that $\mu = o(1)$.
Taking maximum of the right hand side of (ref) over all $c_0>0$, we have
equation[equation omitted — 472 chars of source]
It suffices to find an upper bound of $B^\prime(d)^{\top} \widehat\Sigma_B B^\prime(d)$. By ((ref)),
\[\|B^\top B - {\mathbb{E}}(B^\top B)\|_2 \lesssim_p \sqrt{n M^{-1}\log (pM)}\]
and $\|{\mathbb{E}}(B^\top B)\|_2\asymp nM^{-1}$ by Proposition (ref). Thus,
\[
aligned\sup_{d\in\mathcal{D}}\left|\dfrac{B^\prime(d)^\top\left( \widehat\Sigma_B - n^{-1}{\mathbb{E}}(B^\top B)\right)B^\prime(d)}{B^\prime(d)^\top n^{-1}{\mathbb{E}}(B^\top B)B^\prime(d)}\right| &\lesssim_p
\sqrt{\dfrac{\log(pM)}{n}} \to 0
\]
which implies $B^\prime(d)^\top \widehat\Sigma_B B^\prime(d) \asymp_p B^\prime(d)^\top n^{-1}{\mathbb{E}}(BB^\top) B^\prime(d) \asymp M^{-1}\|B^\prime(d)\|_2^2$. Then
equation[equation omitted — 220 chars of source]
Then by (ref), ((ref)), and (ref),
equation[equation omitted — 218 chars of source]
The proof is completed by combining ((ref)) and ((ref)).
{\bfProof of (S2)}.
By ((ref)) and the fact that $M = n^\nu$ with $\nu > \frac{1}{2\gamma+1}$ in Assumption (ref), we have $\sqrt{n}M^{-\gamma+1} = o(M^{\frac{2\gamma+1}{2}-\gamma+1}) = o(M^{1.5}) = o_p(\widehat s(d))$.
We then carefully decompose the residual $\widehat\epsilon_i$. Recall that we define $\varDelta^Y_i = \widehat{W}_{i\cdot}^\top(\omega-\widehat\omega) + X_{i\cdot}(\theta-\widehat\theta)$ and $\varDelta^v_i = q(v_i) - q(\widehat v_i)$ in (ref). Then
equation[equation omitted — 155 chars of source]
Using Taylor expansion,
\[
aligned\varDelta^v_i &= q^\prime(\widehat v_i)(v_i - \widehat v_i) + q^{\prime\prime}(v_i^*)(v_i - \widehat v_i)^2 \\
&= q^\prime(\widehat v_i)(X_{i\cdot}^\top(\widehat\varphi^ind - \varphi) + K_{i\cdot}(\widehat\kappa^ind - \kappa) ) + q^\prime(\widehat v_i)r_{\psi i} + q^{\prime\prime}(v_i^*)(v_i - \widehat v_i)^2 \\
\]
where $r_{\psi i}$ is the spline approximation error in (ref). Define
\[\widehat\varDelta^v_i = \widehat q^\prime(\widehat v_i)(X_{i\cdot}^\top(\widehat\varphi^\text{ind} - \varphi)+K_{i\cdot}^\top(\widehat\kappa^\text{ind}-\kappa))\]
with $\widehat q^\prime(\widehat v_i) = H^\prime(\widehat v_i)^\top\widehat\eta^\text{ind}$.
We deduce that
equation*[equation* omitted — 338 chars of source]
By (ref),
equation[equation omitted — 140 chars of source]
where $\widetilde r_i := (\varDelta^v_i - \widehat\varDelta^v_i) + r_g(D_i) + r_q(\widehat v_i)$ the high order terms in (ref). (ref) yields a decomposition of the error we need to bound, given as
equation[equation omitted — 367 chars of source]
We first bound $\varDelta_1(d)$. Define $\check r_i := r_g(D_i) + r_q(\widehat v_i) + q^\prime(\widehat v_i)r_{\psi i}$ as the spline approximation errors. Then
\[\widetilde r_i = \check r_i + q^{\prime\prime}(v_i^*)(v_i - \widehat v_i)^2 + (q^\prime(\widehat v_i) - \widehat q^\prime(\widehat v_i))\left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]\]
where
\[|\varDelta_1(d)| \leq \left|\dfrac{\widehat{m}(d)^\top \sum_{i=1}^n \widehat F_{i\cdot} \check r_i}{\sqrt{n}}\right| + \left|\dfrac{\widehat{m}(d)^\top \sum_{i=1}^n \widehat F_{i\cdot} (\widetilde r_i - \check r_i)}{\sqrt{n}}\right| =: \varDelta_{11}(d) + \varDelta_{12}(d).\]
When $\nu\in(\frac{1}{2\gamma},\frac{1}{4.5})$ in Assumption (ref) holds, by Cauchy-Schwartz inequality
equation[equation omitted — 265 chars of source]
uniformly for all $d$, where the last equality applies Proposition (ref) and the boundness of the function $q^\prime(\cdot)$, implying that $\sum_{i=1}^n \check r_i^2
= O(n\cdot M^{-2\gamma}) = o(1/\log n)$. For $\varDelta_{12}(d)$, again by Cauchy-Schwartz inequality
\[|\varDelta_{12}(d)| \leq \widehat s(d)\sqrt{ \sum_{i=1}^n (\widetilde r_i - \check r_i)^2}.\]
It thus suffices to show the following lemma.
LemmaUnder the conditions for Theorem (ref), $\sup_g \sqrt{\sum_{i=1}^n (\widetilde r_i - \check r_i)^2}=o_{u.p.}(1/\log n)$.
It remains to bound $\varDelta_2(d)$. Define
\[\pi = (\beta^\top,\eta^\top,\theta^\top,-\kappa^\top,-\varphi^\top)^\top,\]
\[\widehat\pi = (\widehat\beta^\top,\widehat\eta^{\text{ind}\top},\widehat\theta^\top,-\widehat\kappa^{\text{ind}\top},-\widehat\varphi^{\text{ind}\top})^\top,\]
and recall the definition of $\widehat F_{i\cdot}$ in (ref). Thus, \[\varDelta^Y_i+\widehat\varDelta^v_i = \widehat F_{i\cdot}^\top (\pi - \widehat \pi).\]
We use $I_B$ to denote the first $M_D$ rows of the $p_F$ demensional identify matrix. Then uniformly for all $d$,
\[
aligned|\varDelta_2(d)| &= \left|\sqrt{n}B^\prime(d)^\top\widehat\Omega_B \sum_{i=1}^n\widehat F_{i\cdot} \widehat F_{i\cdot}^\top(\pi-\widehat\pi) - B^\prime(d)^\top I_B(\pi-\widehat\pi)\right| = \left|\sqrt{n}B^\prime(d)^\top(\widehat\Omega_B\widehat\Sigma_F - I_B)(\pi-\widehat\pi)\right|. \\
\sup_{d\in\mathcal{D}}|\varDelta_2(d)| &\leq \sup_{d\in\mathcal{D}}\sqrt{n}\cdot \|B^\prime(d)\|_1 \cdot \|\widehat\Omega_B\widehat\Sigma_F - I_B\|_\infty \cdot \|\widehat\pi-\pi\|_1 \\
&\lesssim_{u.p.} \sqrt{n}\cdot M^2 \sqrt{\dfrac{\log (pM)}{n}} \cdot \|\widehat\pi-\pi\|_1\\
\]
where the last step applies the bound of $\|B^\prime(d)\|_1$ by Proposition (ref), and the first restriction of (ref). By (ref) and Proposition (ref), we can deduce that under the conditions for Theorem (ref),
\[\sup_g \|\beta - \widehat\beta\|_2 + \sup_g \|\eta-\widehat\eta\|_2 + \|\kappa - \widehat\kappa^\text{ind}\|_2 = o_p\left(\frac{1}{M\sqrt{\log (pM) }\log n }\right)\]
and by (ref)
\[\sup_g \|\theta - \widehat\theta\|_1 + \|\varphi-\widehat\varphi^\text{ind}\|_1 = o_p\left(\frac{1}{\sqrt{M \log (pM) } \log n}\right).\]
Thus,
\[
aligned\sup_g \|\widehat\pi-\pi\|_1 &\leq \sqrt{M}\sup_g \|\beta-\widehat{\beta}\|_2 + \sqrt{M}\sup_g \|\eta-\widehat{\eta}\|_2 + \sqrt{p_zM}\|\kappa - \widehat\kappa^ind\|_2 + \sup_g \|\theta-\widehat{\theta}\|_1 + \|\varphi - \widehat\varphi^ind\|_1 \\
&=o_p\left(\frac{1}{\sqrt{M \log (pM) } \log n}\right),
\]
which implies
\[ \sup_{d\in\mathcal{D}}|\varDelta_2(d)|\lesssim_{u.p.} \sqrt{n}\cdot M^2 \sqrt{\dfrac{\log (pM)}{n}} \cdot o_{u.p.}\left(\frac{1}{\sqrt{M \log (pM) } \log n}\right) = o_{u.p.}\left(\frac{1}{ \log n}\right) \]
Then, $$\sup_{d\in\mathcal{D}} \frac{\sup_{d}|\varDelta_2(d)|}{\widehat s(d)} = o_p(\frac{M^{1.5}}{\widehat s(d)\log n}) = o_{u.p.}(1/\log n).$$ We complete the proof of Proposition (ref).
Proof of Theorem (ref)
It suffices to prove the following three results.
enumerate[(R1)]
• $\sup_g \left|\widehat\sigma_\varepsilon/\sigma_\varepsilon - 1\right| \lesssim_p \frac{1}{n^{1/4}}$. This result implies $\inf_g \widehat\sigma_\varepsilon \gtrsim_p 1$, and together with (S2) in the proof of Proposition (ref) implies that \[\sup_{d\in\mathcal{D}}\left|\dfrac{\sqrt{n}\left(\widetilde{g}^\prime(d) - g^\prime(d)\right) - \mathcal{Z}(d)}{\widehat\sigma_\varepsilon\cdot \widehat s(d)}\right| \lesssim_{u.p.} \frac{1}{\log n}.\]
• Show $\sup_{d\in\mathcal{D}}|\mathcal{Z}(d)/\widehat s(d)| \lesssim_{u.p.} n^{1/4} / \log n$. This result together with the error bound for $\widehat\sigma_\varepsilon$ in (R1) implies
\begin{equation}
\sup_{d}\left|\mathbb{H}(d) - \widetilde{\mathbb{H}}(d)\right| = o_{u.p.}\left(\dfrac{1}{\sqrt{\log n}}\right)
\end{equation}
where $\mathbb{H}(d) := \dfrac{ \sqrt{n}\left(\widetilde{g}^\prime(d) - g^\prime(d)\right)} {\widehat\sigma_\varepsilon\cdot \widehat s(d)}$ and $\widetilde{\mathbb{H}}(d) := \dfrac{\mathcal{Z}(d)}{ \sigma_\varepsilon\cdot \widehat s(d)}$.
• There exists a version of Gaussian process $\mathbb{H}^{(1)}(d)$ such that
\begin{equation}
\left|\sup_{d}|\widetilde{\mathbb{H}}(d)| - \sup_{d}|\mathbb{H}^{(1)}(d)|\right| \lesssim_{u.p.} \frac{1}{\log n}
\end{equation}
with
\[\mathbb{H}^{(1)}(d) := (\sigma_\varepsilon^2 n)^{-1/2}\sum_{i=1}^n B^\prime(d)^\top \widehat\Omega_B \widehat F_{i\cdot} e^{(1)}_i\]
where $e^{(1)}_i$ are i.i.d.\ $N(0,\sigma_\varepsilon^2)$ variables.
Together with ((ref)) in (R2), we can deduce that
\begin{equation}
\left|\sup_{d}|{\mathbb{H}}(d)| - \sup_{d}|\mathbb{H}^{(1)}(d)|\right| \lesssim_{u.p.} \frac{1}{\log n}.
\end{equation}
We will also show
\begin{equation}
\sup_g {\mathbb{E}}\left[\sup_{d}| \mathbb{H}^{(1)}(d)|\right] \lesssim \sqrt{\log n}
\end{equation}
which by chernozhukov2014anti implies that for any $\epsilon > 0$
\begin{equation}
\sup_g \sup_{x\in\mathbb{R}}\Pr\left[ \left|\sup_{d}|\mathbb{H}^{(1)}(d)| - x \right|\leq \epsilon \right] \lesssim \epsilon\sqrt{\log n}.
\end{equation}
• Recall that $\widehat c_n(\alpha)$ is defined as the $(1-\alpha)$-quantile of $\sup_{d}|\widehat{\mathbb{H}}(d)|$ with $\widehat{\mathbb{H}}(d)$ defined in ((ref)). Let $c_n(\alpha)$ be the $(1-\alpha)$-th quantile of $\sup_{d}|\mathbb{H}^{(1)}(d)|$. Show that for some sequences $\tau_{n}= o(1)$ and $\epsilon_n = o(1/\sqrt{\log (pM)})$,
\begin{equation}
\begin{aligned}
\widehat c_n(\alpha) \geq c(\alpha + \tau_n) - \epsilon_n. \\
\end{aligned}
\end{equation}
Then through (R1)-(R4), we deduce that
\[
aligned&\ \ \ \ \inf_g\Pr\left(g^\prime(d)\in\mathcal{C}_{n,\alpha}(d) for all d\in\mathcal{D}\right) \\
&\geq \inf_g\Pr\left(\sup_{d}|{\mathbb{H}}(d)| \leq \widehat c_n(\alpha)\right) \\
&\geq \inf_g\Pr\left(\sup_{d}|\mathbb{H}^{(1)}(d)| \leq \widehat c_n(\alpha) + o((\log (pM))^{-1/2}) \right) + o(1) \\
&\geq \inf_g\Pr\left(\sup_{d}|\mathbb{H}^{(1)}(d)| \leq c_n(\alpha + \tau_n) - \epsilon_n + o((\log (pM))^{-1/2}) \right) + o(1) \\
&\geq 1 - \alpha - \tau_n -c\left[\epsilon_n + o((\log (pM))^{-1/2})\right]\sqrt{\log (pM)} + o(1) \to 1-\alpha.
\]
The second inequality applies ((ref)). The third inequality applies ((ref)). The last inequality applies ((ref)).
{\bfProof of (R1)}. Note that
\[\left|\dfrac{\widehat\sigma_\varepsilon}{\sigma_\varepsilon}-1\right|=\dfrac{\left|\dfrac{\widehat\sigma_\varepsilon^2}{\sigma_\varepsilon^2}-1\right|}{\left|\dfrac{\widehat\sigma_\varepsilon}{\sigma_\varepsilon}+1\right|} = \sigma_\varepsilon\dfrac{\left| \widehat\sigma_\varepsilon^2 -\sigma_\varepsilon^2\right|}{\left|\widehat\sigma_\varepsilon+\sigma_\varepsilon\right|},\]
It thus suffices to show that $\sup_g |\widehat\sigma_\varepsilon^2 -\sigma_\varepsilon^2|=o_p(1)$. Since
$n^{-1}\sum_{i=1}^n \varepsilon_i^2 - \sigma_\varepsilon^2 = O_p(n^{-1/2})$ by the bounded fourth order moments of $\varepsilon_i$, and
it suffices to show that $\sup_g |\widehat\sigma_\varepsilon^2 - n^{-1} \sum_{i=1}^n\varepsilon_i^2| = \sup_g |n^{-1}\sum_{i=1}^n(\widehat\varepsilon_i^2 - \varepsilon_i^2)| = o_p(1/n^{-1/4})$. As
\[
aligned\left|n^{-1}\sum_{i=1}^n (\widehat\varepsilon_i^2 - \varepsilon_i^2)\right| &\leq n^{-1}\sum_{i=1}^n (\widehat\varepsilon_i - \varepsilon_i)^2 + 2\left|n^{-1}\sum_{i=1}^n (\widehat\varepsilon_i - \varepsilon_i)\varepsilon_i\right| \\
&\leq n^{-1}\sum_{i=1}^n (\widehat\varepsilon_i - \varepsilon_i)^2 + 2\sqrt{n^{-1}\sum_{i=1}^n (\widehat\varepsilon_i - \varepsilon_i)^2 }\cdot \sqrt{n^{-1}\sum_{i=1}^n \varepsilon_i^2},
\]
it remains to show $n^{-1}\sum_{i=1}^n (\widehat\varepsilon_i - \varepsilon_i)^2 = o_p(1/\sqrt{n})$. By ((ref)), $\widehat\varepsilon_i - \varepsilon_i = \widehat W_i^\top(\omega-\widehat\omega) + X_{i\cdot}(\theta - \widehat\theta) + q(v_i) - q(\widehat v_i) + r_g(D_i) + r_q(\widehat v_i)$, then
\[
aligned\dfrac{1}{n}\sum_{i=1}^n(\widehat\varepsilon_i - \varepsilon_i)^2 &\lesssim \dfrac{1}{n}\sum_{i=1}^n\left[\widehat W_i^\top(\omega-\widehat\omega) + X_{i\cdot}(\theta - \widehat\theta)\right]^2 + \dfrac{1}{n}\sum_{i=1}^n [q(v_i) - q(\widehat v_i)]^2 + \dfrac{1}{n}\sum_{i=1}^n[r_g^2(D_i) + r_q^2(\widehat v_i)] \\
&=: \varDelta^\varepsilon_1 + \varDelta^\varepsilon_2 + \varDelta^\varepsilon_3.
\]
(ref), (ref), ((ref)), ((ref)), and $M\gg n^{1/(2\gamma)}$ in Assumption (ref) imply $\sup_g \varDelta^\varepsilon_1 = o_p(n^{-1/2})$.
((ref)) implies $\varDelta^\varepsilon_2 = o_p(n^{-1/2})$. Proposition (ref) implies $\sup_g \varDelta^\varepsilon_3 = o_p(n^{-1/2})$. It completes the proof of $\sup_g |\widehat\sigma_\varepsilon / \sigma_\varepsilon - 1| = o_p(n^{-1/4})$.
{\bf Proof of (R2)}. Note that
equation[equation omitted — 404 chars of source]
where the last inequality applies the bound of $\|B^\prime(d)\|_1$ by Proposition (ref) and the lower bound of $\widehat s(d)$ by (ref).
It suffices to bound $\|\sum_{i=1}^n \widehat\Omega_B \widehat F_{i\cdot} \varepsilon_i \|_\infty$. By the Chebyshev inequality and the union bound,
\[
aligned\Pr\left(\max_{j\in[M_D]}\dfrac{n^{-1/2}\left|\sum_{i=1}^n\widehat\Omega_{Bj}^\top\widehat F_{i\cdot}\varepsilon_i\right|}{\sqrt{ \Omega_{Bj}^\top\widehat\Sigma_F\Omega_{Bj}}\cdot \sigma_\varepsilon } > \tau \Bigg|\mathcal{G} \right) \leq \dfrac{M}{\tau^2}.
\]
Taking $\tau = \sqrt{M\log (pM)}$,
\[
aligned\sup_g \Pr\left(\max_{j\in[M_D]}\dfrac{n^{-1/2}\left|\sum_{i=1}^n\widehat\Omega_{Bj}^\top\widehat F_{i\cdot}\varepsilon_i\right|}{\sqrt{ \Omega_{Bj}^\top\widehat\Sigma_F\Omega_{Bj}}\cdot \sigma_\varepsilon } > \sqrt{M\log (pM)}\Bigg|\mathcal{G} \right) \leq \dfrac{1}{\log (pM)} \to 0.
\]
Using Bounded Convergence Theorem on the supreme of the conditional probability,
\[
aligned&\ \ \ \ \sup_g \Pr\left(\max_{j\in[M_D]}\dfrac{n^{-1/2}\left|\sum_{i=1}^n\widehat\Omega_{Bj}^\top\widehat F_{i\cdot}\varepsilon_i\right|}{\sqrt{ \Omega_{Bj}^\top\widehat\Sigma_F\Omega_{Bj}}\cdot \sigma_\varepsilon } > \sqrt{ M\log (pM)} \right) \\ &\leq {\mathbb{E}}\left(\sup_g \Pr\left(\max_{j\in[M_D]}\dfrac{n^{-1/2}\left|\sum_{i=1}^n\widehat\Omega_{Bj}^\top\widehat F_{i\cdot}\varepsilon_i\right|}{\sqrt{ \Omega_{Bj}^\top\widehat\Sigma_F\Omega_{Bj}}\cdot \sigma_\varepsilon } > \sqrt{ M\log (pM)}\Bigg|\mathcal{G} \right)\right) \to 0
\]
which implies
equation[equation omitted — 341 chars of source]
where the last inequality applies by ((ref)) that $$\max_{j\in[M_D]}\Omega_{Bj}^\top\widehat\Sigma_F\Omega_{Bj} \leq \|\Omega_{B}^\top\widehat\Sigma_F\Omega_{B}^\top \|_\infty \lesssim_{u.p.} M.$$
(ref) and (ref) imply
\[\sup_{d\in\mathcal{D}} \dfrac{|\mathcal{Z}(d)|}{\widehat s(d)} \leq \sup_{d\in\mathcal{D}}\dfrac{\|B^\prime(d)\|_1}{\widehat s(d)}\cdot n^{-1/2}\|\sum_{i=1}^n \widehat\Omega_B \widehat F_{i\cdot} \varepsilon_i \|_\infty \lesssim_{u.p.} \sqrt{M\log (pM)} = o(n^{1/4}/\log n),\]
which completes the proof of (R2).
{\bfProof of (R3)}. Let $\zeta_i(d) := \widehat s(d)^{-1} \widehat m(d)^\top\widehat F_{i\cdot} \varepsilon_i$ and $\zeta_i^*(d) := \widehat s(d)^{-1}\widehat m(d)^\top \widehat F_{i\cdot} e^{(1)}_i$ where $e^{(1)}_i$ are i.i.d.\ $N(0,\sigma_\varepsilon^2)$ variables. Then for any $d_0, d_1, \cdots$, the vector $(\zeta_i^*(d_0),\zeta_i^*(d_1),\cdots)^\top$ is jointly normal conditionally on $\mathcal{G}$ with pairwise covariance for any $(d_0,d_1)$, given as
$$ {\mathbb{E}}_{\mathcal{G}}[\zeta^*_i(d_0)\zeta^*_i(d_1)] = \sigma_\varepsilon^2 \frac{\widehat m(d_0)^\top\widehat F_{i\cdot} \widehat F_{i\cdot} \widehat m(d_1)}{\widehat s(d_0) \widehat s(d_1) } = {\mathbb{E}}_{\mathcal{G}}[\zeta_i(d_0)\zeta_i(d_1)]. $$
Define
\[ \mathbb{G} := \sup_{d\in\mathcal{D}} \sum_{i=1}^n \zeta_i(d),\ \mathbb{G}^\prime := \sup_{d\in\mathcal{D}} \sum_{i=1}^n [- \zeta_i(d)],\]
\[ \mathbb{G}^* := \sup_{d\in\mathcal{D}} \sum_{i=1}^n \zeta_i^*(d),\ \mathbb{G}^{*\prime} := \sup_{d\in\mathcal{D}} \sum_{i=1}^n [-\zeta_i^*(d)].\]
We have the following lemma.
LemmaUnder the conditions for Theorem (ref),
\[ n^{-1/2}|\mathbb{G} - \mathbb{G}^*| \lesssim_{u.p.} \dfrac{1}{\sqrt{\log n}} ,\]
\[ n^{-1/2}|\mathbb{G}^\prime - \mathbb{G}^{*\prime}|
\lesssim_{u.p.} \dfrac{1}{\sqrt{\log n}}.\]
Note that $\sup_{d\in\mathcal{D}}|\sum_{i=1}^n\zeta_i(d)| = \max\{\mathbb{G},\mathbb{G}^\prime \}=\frac{\mathbb{G}+\mathbb{G}^\prime+|\mathbb{G}-\mathbb{G}^\prime|}{2}$ and $\sup_{d\in\mathcal{D}}|\sum_{i=1}^n\zeta_i^*(d)| = \max\{\mathbb{G}^*,\mathbb{G}^{*\prime} \}=\frac{\mathbb{G}^*+\mathbb{G}^{*\prime}+|\mathbb{G}^*-\mathbb{G}^{*\prime}|}{2}$. Thus by Lemma (ref),
equation[equation omitted — 921 chars of source]
Recall that $\widetilde{\mathbb{H}}(d) = \widehat\sigma_\varepsilon n^{-1/2}\sum_{i=1}^n\zeta_i(d)$, $\widehat{\mathbb{H}}(d) = \widehat\sigma_\varepsilon n^{-1/2}\sum_{i=1}^n\zeta_i^*(d)$ and $\mathbb{H}^{(1)}(d) = \sigma_\varepsilon n^{-1/2}\sum_{i=1}^n\zeta_i^*(d)$. By (R1) and ((ref)),
\[\left|\sup_{d}|\widehat{\mathbb{H}}(d)| - \sup_{d}|\widetilde{\mathbb{H}}(d)|\right| \lesssim_{u.p.} \dfrac{1}{\widehat\sigma_\varepsilon} n^{-1/2} \left|\sup_{d\in\mathcal{D}}|\sum_{i=1}^n\zeta_i(d)| - \sup_{d\in\mathcal{D}}|\sum_{i=1}^n\zeta_i^*(d)|\right| \lesssim_{u.p.} \dfrac{1}{\sqrt{\log n}}.\]
If suffices to show
equation[equation omitted — 168 chars of source]
Note that
\[
aligned\left|\sup_{d}|\widehat{\mathbb{H}}(d)| - \sup_{d}|\mathbb{H}^{(1)}(d)|\right| &\leq \left|\dfrac{1}{\widehat\sigma_\varepsilon} - \dfrac{1}{\sigma_\varepsilon}\right| \cdot \sigma_\varepsilon\sup_{d}\left|\mathbb{H}^{(1)}(d)\right| \lesssim_{u.p.} \dfrac{1}{(\log (pM))^{1.5}} \cdot \sup_{d}\left|\mathbb{H}^{(1)}(d)\right|
\]
where the last step applies (R1). If ((ref)) holds, we can use the Markov inequality to deduce
\[\sup_{d}\left|\mathbb{H}^{(1)}(d)\right| = o_{u.p.}((\log n)^{0.75})\]
and thus
\[\left|\sup_{d}|\widehat{\mathbb{H}}(d)| - \sup_{d}|\mathbb{H}^{(1)}(d)|\right| = o_{u.p.}\left(\frac{1}{(\log n)^{0.75}}\right).\]
Then the proof of (R2) will end with the verification of ((ref)). Let $d_j^\varDelta = a_D + j\varDelta_J$ for $j=0,1,\cdots,J$ with some integer $J$ large enough, where $\varDelta_J := \frac{b_D - a_D}{J}$. Then for any $d$, we can find some $d_j^\varDelta$ such that $|d_j^\varDelta - d_j| \leq \varDelta_J$ and hence
equation[equation omitted — 379 chars of source]
Note that $\mathbb{G}^*(d^\varDelta_j)\sim N(0,1)$ for each $j\in[J]$, and thus
\[ {\mathbb{E}}\left[\max_{j\in[J]} |\mathbb{G}^*(d^\varDelta_j)| \right] \lesssim \sqrt{\log J}. \]
For any $d,d^\prime\in\mathcal{D}$,
equation[equation omitted — 483 chars of source]
where the last inequality applies
the definition of $\widehat\Omega_B$ when ((ref)) is feasible\footnote{We define $\widehat\Omega_B = O$ and $\widehat{\mathbb{H}}(d) = \mathbb{H}^{(1)}(d) = \widetilde{\mathbb{H}}(d) = 0$ when ((ref)) is infeasible.} with $\mu_j \asymp M\sqrt{\log (pM) / n}$. Additionally, by ((ref)) and ((ref)), we have almost surely that $$ \inf_g \inf_{d\in\mathcal{D}}\widehat s(d) \gtrsim \inf_{d\in\mathcal{D}}\|B^\prime(d)\|_1^2 / \sqrt{B^\prime(d)^\top\widehat\Sigma_B B^\prime(d)}\gtrsim \inf_{d\in\mathcal{D}}\dfrac{ \|B^\prime(d)\|_1^2 }{\sqrt{\|B^\prime(d)\|_1^2\cdot \|\widehat\Sigma_B\|_\infty}} \gtrsim M$$
with $\widehat\Sigma_B = n^{-1}\sum_{i=1}^n B_{i\cdot} B_{i\cdot}^\top$ and thus $\|\widehat\Sigma_B\|_\infty$ is uniformly bounded. In addition, by the second restriction in (ref),
equation[equation omitted — 198 chars of source]
Also, $\|B^\prime(d)+B^\prime(d^\prime)\|_1 \lesssim M$ by Proposition (ref). Then
\[
aligned&\ \ \ \ \sup_g \|\widehat s(d)^{-1}B^\prime(d) - \widehat s(d^\prime)^{-1}B^\prime(d^\prime)\|_1 \\
&\leq \|B^\prime(d)\|_1\cdot\sup_g \dfrac{\left|\widehat s(d^\prime)^2 - \widehat s(d)^2\right|}{\widehat s(d)\widehat s(d^\prime)[\widehat s(d) + \widehat s(d^\prime)]} + \dfrac{\|B^\prime(d) - B^\prime(d^\prime)\|_1}{\widehat s(d^\prime)} \\
&\lesssim \dfrac{\|B^\prime(d)\|_1\cdot\sup_g |(B^\prime(d) - B^\prime(d^\prime))^\top\widehat\Omega_B \widehat\Sigma_F \widehat\Omega_B^\top (B^\prime(d) + B^\prime(d^\prime))| }{M^3} +\dfrac{\|B^\prime(d) - B^\prime(d^\prime)\|_1}{M}\\
&\lesssim \left|\dfrac{\|B^\prime(d) - B^\prime(d^\prime)\|_1\cdot \sup_g \|\widehat\Omega_B \widehat\Sigma_F \widehat\Omega_B^\top\|_\infty\cdot \|B^\prime(d) + B^\prime(d^\prime)\|_1}{M^{2}}\right| + \dfrac{\|B^\prime(d) - B^\prime(d^\prime)\|_1}{M} \\
&\lesssim \|B^\prime(d) - B^\prime(d^\prime)\|_1\cdot M\log (pM).
\]
By Proposition (ref), there are at most a fixed number (denoted as $K_0$) nonzero elements in the vector $B^\prime(d) - B^\prime(d^\prime)$. Denote this active set as $\mathcal{K}_0$. Also, note that
equation[equation omitted — 360 chars of source]
where the inequality applies Proposition (ref) for the upper bound of second-order derivative $B^{\prime\prime}$, and $d_{2j}$ is between $d$ and $d^\prime$. Thus, $$\sup_g \|\widehat s(d)^{-1}B^\prime(d) - \widehat s(d^\prime)^{-1}B^\prime(d^\prime)\|_1 \lesssim M^3\log (pM) |d - d^\prime|,$$ together with (ref) implying that
\[
alignedn^{-1/2}|\mathbb{G}^*(d) - \mathbb{G}^*(d^\prime)| \lesssim \sqrt{n}M^4(\log (pM))^{1.5}\cdot |d - d^\prime|\cdot \sup_{i\in[n]}|\varepsilon_i|. \\
\]
Thus,
\[
aligned\sup_g {\mathbb{E}}\left(\sup_{|d-d^\prime|\leq\varDelta_J} |n^{-1/2}[\mathbb{G}^*(d) - \mathbb{G}^*(d^\prime)]|\right) &\lesssim \sqrt{n}M^4(\log (pM))^{1.5}\Delta_J {\mathbb{E}}(\sup_{i\in[n]}|\varepsilon_i|) \&\lesssim \sqrt{n}M^4(\log (pM))^2\Delta_J.
\]
Taking $J \asymp n^5$ , we have $\Delta_j \asymp n^{-5}$ and by ((ref))
\[
aligned{\mathbb{E}}\left( \sup_{d\in\mathcal{D}} |\mathbb{H}^{(1)}(d)| \right)
&\leq {\mathbb{E}}\left(\sup_{|d-d^\prime|\leq \varDelta_J} n^{-1/2}|\mathbb{G}^*(d)-\mathbb{G}^*(d^\varDelta_j)|\right) + {\mathbb{E}}\left(\max_{j\in[J]}|\mathbb{G}^*(d^\varDelta_j)|\right) \\
&\lesssim \sqrt{n}M^4(\log (pM))^2\cdot n^{-5} + \sqrt{5\log n} \lesssim \sqrt{\log n}.
\]
{\bf Proof of (R4)}. By ((ref)),
equation*[equation* omitted — 167 chars of source]
with some $\tau_{n}\to 0$. Thus,
\[
aligned\sup_g \Pr\left( \sup_{d}|\mathbb{H}^{(1)}(d)| > \widehat c_n(\alpha ) + \dfrac{C}{(\log n)^{0.75}}\right) &\geq
\sup_g \Pr\left( \sup_{d}|\mathbb{H}^{(1)}(d)| > \widehat c_n(\alpha ) \right) - \tau_{n} \\
&= 1- \alpha - \tau_{n}
\]
which implies $\widehat c_n(\alpha ) < c_n(\alpha + \tau_{n}) - \dfrac{C}{(\log n)^{0.75}}$. The proof completes by taking $\epsilon_n = \dfrac{C}{(\log n)^{0.75}}$.
Proof of Technical Lemmas
proof[Proof of Lemma (ref)]
Following the previous notations, we use $S_x$ to denote ${\rm diag}({\mathbb{E}}(x_{ij}^2))_{j\in[p]}$ for any generic random vector $x_{i\cdot} = (x_{ij})_{j\in[p]}$, and $\widetilde x_{i\cdot} := S_x^{-1/2} x_{i\cdot}$ to denote the standardized version of any random vector $x_{i\cdot}$. Note that
\begin{equation}
\begin{aligned}
\|\widehat\varphi - \varphi\|_2^2 &\lesssim (\widehat\varphi - \varphi)^\top {\mathbb{E}}[X_{i\cdot} X_{i\cdot}^\top] (\widehat\varphi - \varphi) + (\widehat\kappa - \kappa)^\top {\mathbb{E}}[K_{i\cdot} K_{i\cdot}^\top] (\widehat\kappa - \kappa) \\
&= (\widehat\varphi - \varphi)^\top S_X^{1/2}{\mathbb{E}}[\widetilde X_{i\cdot} \widetilde X_{i\cdot}^\top]S_X^{1/2} (\widehat\varphi - \varphi) + (\widehat\kappa - \kappa)^\top S_K^{1/2} {\mathbb{E}}[\widetilde K_{i\cdot} \widetilde K_{i\cdot}^\top]S_K^{1/2} (\widehat\kappa - \kappa) \\
&\lesssim (\widehat\varphi - \varphi)^\top S_X^{1/2} S_X^{1/2} (\widehat\varphi - \varphi) + (\widehat\kappa - \kappa)^\top S_K^{1/2} S_K^{1/2} (\widehat\kappa - \kappa) \\
&\lesssim (\widehat\varphi - \varphi)^\top S_X^{1/2} {\mathbb{E}}[\widetilde X_{i\cdot} \widetilde X_{i\cdot}^\top] S_X^{1/2}(\widehat\varphi - \varphi) + (\widehat\kappa - \kappa)^\top S_K^{1/2} {\mathbb{E}}[\widetilde K_{i\cdot} \widetilde K_{i\cdot}^\top] S_K (\widehat\kappa - \kappa) + \\
&\ \ \ \ 2(\widehat\kappa - \kappa)^\top S_K^{1/2} {\mathbb{E}}[\widetilde K_{i\cdot} \widetilde X_{i\cdot}^\top] S_X^{1/2} (\widehat\varphi - \varphi)\\
&= (\widehat\varphi - \varphi)^\top {\mathbb{E}}[ X_{i\cdot} X_{i\cdot}^\top] (\widehat\varphi - \varphi) + (\widehat\kappa - \kappa)^\top {\mathbb{E}}[ K_{i\cdot} K_{i\cdot}^\top] (\widehat\kappa - \kappa) + \\
&\ \ \ \ 2(\widehat\kappa - \kappa)^\top {\mathbb{E}}[ K_{i\cdot} X_{i\cdot}^\top] (\widehat\varphi - \varphi)\\
\end{aligned}
\end{equation}
where the third and the fourth steps apply Assumption (ref).
By ((ref)) and ((ref)), we deduce that
\begin{equation}
n^{-1}\|K(\widehat\kappa-\kappa) + X(\varphi - \widehat\varphi)\|_2^2 \lesssim R_{D1}^2.
\end{equation}
and also note the following decomposition
\[\begin{aligned}
n^{-1}\|K(\widehat\kappa-\kappa) + X(\varphi - \widehat\varphi)\|_2^2 &= (\widehat\varphi - \varphi)^\top \dfrac{\sum_{i=1}^n X_{i\cdot} X_{i\cdot}^\top}{n} (\widehat\varphi - \varphi) + (\widehat\kappa - \kappa)^\top \dfrac{\sum_{i=1}^n K_{i\cdot} K_{i\cdot}^\top}{n} (\widehat\kappa - \kappa) + \\
&\ \ \ \ 2(\widehat\kappa - \kappa)^\top \dfrac{\sum_{i=1}^n K_{i\cdot} X_{i\cdot}^\top}{n} (\widehat\varphi - \varphi).
\end{aligned} \]
By Proposition (ref), with probability at least $1 - c(pM)^{-4}$
\[\|\dfrac{1}{n}\sum_{i=1}^n\left[K_{i\cdot} X_{i\cdot}^\top - {\mathbb{E}}(K_{i\cdot} X_{i\cdot}^\top) \right]\|_\infty + \|\dfrac{1}{n}\sum_{i=1}^n\left[X_{i\cdot} X_{i\cdot}^\top - {\mathbb{E}}(X_{i\cdot} X_{i\cdot}^\top) \right]\|_\infty \lesssim \sqrt{\dfrac{\log (pM)}{n}},\]
\[n^{-1}\left\|\sum_{i=1}^n [K_{i\cdot} K_{i\cdot}^\top - {\mathbb{E}}(K_{i\cdot} K_{i\cdot}^\top)] \right\|_\infty \lesssim \sqrt{\dfrac{\log (pM)}{n}}.\]
Together with ((ref)) and ((ref)), we deduce that with probability at least $1 - c(pM)^{-4}$
\[\begin{aligned}
\varDelta^D_1 := (\widehat\varphi - \varphi)^\top \dfrac{1}{n}\sum_{i=1}^n\left[X_{i\cdot} X_{i\cdot}^\top - {\mathbb{E}}(X_{i\cdot} X_{i\cdot}^\top) \right] (\widehat\varphi - \varphi) &\leq \|\widehat\varphi - \varphi\|_1^2 \cdot \|\dfrac{1}{n}\sum_{i=1}^n\left[X_{i\cdot} X_{i\cdot}^\top - {\mathbb{E}}(X_{i\cdot} X_{i\cdot}^\top) \right]\|_\infty\\
&\lesssim \dfrac{R_{D1}^4}{\lambda_D^2} \sqrt{\dfrac{\log (pM)}{n}} \lesssim \sqrt{\dfrac{n}{\log (pM)}}R_{D1}^4,
\end{aligned}\]
\[\begin{aligned}
\varDelta^D_2 := (\widehat\kappa - \kappa)^\top \dfrac{1}{n}\sum_{i=1}^n\left[X_{i\cdot} X_{i\cdot}^\top - {\mathbb{E}}(X_{i\cdot} X_{i\cdot}^\top) \right] (\widehat\kappa - \kappa) &\leq \|\widehat\kappa - \kappa\|_1^2 \cdot \|\dfrac{1}{n}\sum_{i=1}^n\left[K_{i\cdot} K_{i\cdot}^\top - {\mathbb{E}}(K_{i\cdot} K_{i\cdot}^\top) \right]\|_\infty \\
&\lesssim \dfrac{MR_{D1}^4}{R_{D2}^2} \sqrt{\dfrac{\log (pM)}{n}} \lesssim \sqrt{\dfrac{nM^2}{\log (pM)}}R_{D1}^4,
\end{aligned}\]
and
\[\begin{aligned}
\varDelta^D_3 &:= 2(\widehat\kappa - \kappa)^\top \dfrac{1}{n}\sum_{i=1}^n\left[K_{i\cdot} X_{i\cdot}^\top - {\mathbb{E}}(K_{i\cdot} X_{i\cdot}^\top) \right] (\widehat\varphi - \varphi) \\
&\leq \sqrt{p_zM} \|\widehat\kappa - \kappa\|_2 \cdot \|\widehat\varphi - \varphi\|_1 \cdot \|\dfrac{1}{n}\sum_{i=1}^n\left[K_{i\cdot} X_{i\cdot}^\top - {\mathbb{E}}(K_{i\cdot} X_{i\cdot}^\top) \right]\|_\infty\\
&\lesssim \dfrac{\sqrt{M} R_{D1}^4}{R_{D2}\lambda_D} \sqrt{\dfrac{\log (pM)}{n}} \lesssim \sqrt{\dfrac{nM}{\log (pM)}}R_{D1}^4.
\end{aligned}\]
Recall that $R_{D1} \asymp M^{-\gamma}$ and thus
\[\sqrt{\dfrac{nM^2}{\log (pM)}}R_{D1}^2 \lesssim o[\sqrt{n}\cdot M^{-2\gamma + 1}] = o[M^{\frac{2\gamma+1}{2}-2\gamma+1}] = o(M^{-\gamma + 1.5}) = o(1)\]
as $\gamma \geq 2$, which implies $\varDelta^D_1 + \varDelta^D_2 + \varDelta^D_3 = o(R_{D1}^2)$. Consequently, by ((ref)) and ((ref)),
\[\begin{aligned}
\|\widehat\varphi - \varphi\|_2^2 &= n^{-1}\|K(\widehat\kappa - \kappa) + X(\widehat\varphi-\varphi)\|_2^2 + \varDelta^D_1 + \varDelta^D_2 + \varDelta^D_3 \\
&\lesssim R_{D1}^2 + o(R_{D1}^2) \lesssim R_{D1}^2 \\
\end{aligned}\]
with probability at least $1-c(pM)^{-4}$.
proof[Proof of Lemma (ref)]
Let $K_\ell$ denote the matrix with the $(i,j)$-th element $K_{ij}=K(v_i)$. Define $Q=(K,X)$ and $\delta = \left(\begin{array}{c}
\widehat{\kappa} -\kappa \\
\widehat\varphi-\varphi
\end{array}\right)$. Then we decompose the target quadratic form as
\begin{equation}
\begin{aligned}
\delta^\top \dfrac{Q^\top Q}{n} \delta
&= \delta^\top \dfrac{{\mathbb{E}}[Q^\top Q]}{n} \delta + \Delta_0
\end{aligned}
\end{equation}
where
\[\Delta_0 = \delta^\top\dfrac{Q^\top Q-{\mathbb{E}}[Q^\top Q]}{n} \delta.\]
The remaining proofs are composed of two steps.
{\bf Step 1: Find a lower bound of $\delta^\top \dfrac{{\mathbb{E}}[Q^\top Q]}{n} \delta$.}
Note that by Assumption (ref) we have
\begin{equation}
\begin{aligned}
\delta^\top \dfrac{{\mathbb{E}}[Q^\top Q]}{n} \delta &= \delta^\top S_Q^{1/2}\dfrac{{\mathbb{E}}[\widetilde Q^\top \widetilde Q]}{n} S_Q^{1/2}\delta \\
&\gtrsim (\widehat{\kappa}-\kappa)^\top S_K^{1/2}\dfrac{{\mathbb{E}}[\widetilde K^\top \widetilde K]}{n} S_K^{1/2}(\widehat{\kappa}-\kappa) +
(\widehat{\varphi}-\varphi)^\top S_X^{1/2}\dfrac{{\mathbb{E}}[\widetilde X^\top \widetilde X]}{n} S_X^{1/2} (\widehat{\varphi}-\varphi) \\
&= (\widehat{\kappa}-\kappa)^\top \dfrac{{\mathbb{E}}[ K^\top K]}{n} (\widehat{\kappa}-\kappa) +
(\widehat{\varphi}-\varphi)^\top \dfrac{{\mathbb{E}}[ X^\top X]}{n} (\widehat{\varphi}-\varphi) \\
&\gtrsim M^{-1}\|\widehat\kappa - \kappa\|_2^2 + \|\widehat\varphi - \varphi\|_2^2.
\end{aligned}
\end{equation}
{\bf Step 2: Bound $\Delta_0$.} Note that each row in the $n\times(p_zM + p)$ matrix $Q$ are independent bounded variables. By Bernstein-type inequality, with probabilty at least $1-c(M\vee p)^{-4}$
\[
\left\|\dfrac{Q^\top Q}{n}-\dfrac{\mathbb{E}(Q^\top Q)}{n}\right\|_\infty \lesssim \sqrt{\dfrac{\log(M\vee p)}{n}}
\]
and hence
\begin{equation*}
|\Delta_0| \lesssim_p \|\delta\|_1^2\sqrt{\dfrac{\log(M\vee p)}{n}}.
\end{equation*}
Note that
\[\begin{aligned}
\|\delta\|_1^2 &= \|\widehat\kappa - \kappa\|_1^2 + \|(\widehat\varphi-\varphi)_\mathcal{S}\|_1^2 + \|(\widehat\varphi-\varphi)_{\mathcal{S}^c}\|_1^2 \\
&\leq M\|\widehat\kappa - \kappa\|_2^2 + \|(\widehat\varphi-\varphi)_\mathcal{S}\|_1^2 + 2\dfrac{c_2^2R_{D2}^2}{\lambda_D^2}\|\widehat\kappa - \kappa\|_2^2 + 2c_3^2\|(\widehat\varphi-\varphi)_\mathcal{S}\|_1^2 \\
&= \left(2M + 2\dfrac{c_2^2R_{D2}^2}{\lambda_D^2}\right)\|\widehat\kappa - \kappa\|_2^2 + (1+2c_3^2)s\|\widehat\varphi-\varphi\|_2^2 \\
&\lesssim M\|\widehat\kappa - \kappa\|_2^2 + s\|\widehat\varphi-\varphi\|_2^2
\end{aligned}
\]
which implies
\begin{equation}
\begin{aligned}
|\Delta_0|
& \lesssim_p M\sqrt{\dfrac{\log (pM)}{n}}\|\widehat\kappa-\kappa\|_2^2 + s\sqrt{\dfrac{\log (pM)}{n}}\|\widehat\varphi-\varphi\|_2^2 \\
&= o_{\rm{a.s.}}\left(M^{-1}\|\widehat\kappa-\kappa\|_2^2 + \|\widehat\varphi-\varphi\|_2^2\right)
.
\end{aligned}
\end{equation}
proof[Proof of Lemma (ref)]Note that the proof of Lemma (ref) depends on Lemma (ref) since the LASSO algorithm (ref) depends on $\widehat v_i$ from (ref).
Let $H$ denote the matrix with the $(i,j)$-th element $H_{ij}=H_j(v_i)$. Define $W=(B,H)$, $\widehat{U}=(\widehat{W},X)$, $U=(W,X)$ and $\delta = \left(\begin{array}{c}
\widehat{\omega}-\omega \\
\widehat{\theta}-\theta
\end{array}\right)$. Then we decompose the target quadratic form as
\begin{equation}
\begin{aligned}
\delta^\top \dfrac{\widehat{U}^\top \widehat{U}}{n} \delta &= \delta^\top \dfrac{{\mathbb{E}}[U^\top U]}{n} \delta + \delta^\top \dfrac{(U-\widehat{U})^\top (U-\widehat{U})}{n} \delta + 2\delta^\top \dfrac{(U-\widehat{U})^\top U }{n} \delta \\
&\ \ \ + \delta^\top\dfrac{U^\top U-{\mathbb{E}}[U^\top U]}{n} \delta \\
&\geq \delta^\top \dfrac{{\mathbb{E}}[U^\top U]}{n} \delta + \Delta_1 + \Delta_2
\end{aligned}
\end{equation}
where
\[\Delta_1 = 2\delta^\top \dfrac{(U-\widehat{U})^\top U }{n} \delta,\ \ \Delta_2 = \delta^\top\dfrac{U^\top U-{\mathbb{E}}[U^\top U]}{n} \delta.\]
The remaining proofs are composed of three steps.
{\bf Step 1: Find a lower bound of $\delta^\top \dfrac{{\mathbb{E}}[U^\top U]}{n} \delta$.}
Note that following ((ref)), we have by Assumption (ref)
\begin{equation}
\begin{aligned}
\delta^\top \dfrac{{\mathbb{E}}[U^\top U]}{n} \delta &\geq
c^*(\widehat{\beta}-\beta)^\top \dfrac{{\mathbb{E}}[B^\top B]}{n}(\widehat{\beta}-\beta) +
c^*(\widehat{\eta}-\eta)^\top \dfrac{{\mathbb{E}}[H^\top H]}{n}(\widehat{\eta}-\eta) \\ &\ \ \ +
c^*(\widehat{\theta}-\theta)^\top \dfrac{{\mathbb{E}}[X^\top X]}{n}(\widehat{\theta}-\theta) \\
&\geq
c^*c_B M^{-1}\|\widehat{\beta}-\beta\|_2^2 +
c^*c_H M^{-1}\|\widehat{\eta}-\eta\|_2^2 +
c^*c_\Sigma\|\widehat{\theta}-\theta\|_2^2 \\
&\geq
c^*(c_B\wedge c_H) M^{-1}\|\widehat{\omega}-\omega\|_2^2 +
c^*c_\Sigma\|\widehat{\theta}-\theta\|_2^2.
\end{aligned}
\end{equation}
{\bf Step 2: Bound $\Delta_1$.} Note that
\begin{equation}
\Delta_1 = \dfrac{2}{n}(\widehat{\omega} - \omega)^\top(W-\widehat{W})^\top W(\widehat{\omega} - \omega) + \dfrac{2}{n}(\widehat{\omega} - \omega)^\top(W-\widehat{W})^\top X(\widehat{\theta}-\theta).
\end{equation}
Define $a_n := M^{-\gamma + 1} + \sqrt{\dfrac{M^2(s+M)\log (pM)}{n}}$.
By Corollary (ref), with probability at least $1-c(pM)^{-4}$
\begin{equation}
\begin{aligned}
\|W-\widehat{W}\|_2 &\leq \sqrt{\sum_{i=1}^n\sum_{j=1}^{M}[H_j(v_i)-H_j(\widehat{v}_i)]^2} \\
&\lesssim \sqrt{n}a_n.
\end{aligned}
\end{equation}
Besides, $ \|W\|_2 \leq \|B\|_2 + \|H\|_2$ and
\begin{equation}
\begin{aligned}
\|B\|_2 = \sqrt{\|B^\top B\|_2}\leq \sqrt{\|B^\top B - {\mathbb{E}}(B^\top B)\|_2} + \sqrt{\|{\mathbb{E}}(B^\top B)\|_2}.\\
\end{aligned}
\end{equation}
Since $\|B_{i\cdot} B_{i\cdot}^\top\|_2\leq \|B_{i\cdot}\|_2^2\leq k_D$ and
\[
\begin{aligned}
\|\sum_{i=1}^n \mathbb{E}[ B_{i\cdot} B_{i\cdot}^\top B_{i\cdot} B_{i\cdot}^\top ] \|_2 &= \| \sum_{i=1}^n \mathbb{E}[ \|B_{i\cdot}\|_2^2 B_{i\cdot} B_{i\cdot}^\top ]\|_2 \\
&= k_D \| \sum_{i=1}^n \mathbb{E}[ B_{i\cdot} B_{i\cdot}^\top ]\|_2 \leq k_D n M^{-1},
\end{aligned}
\]
Then applying the matrix Bernstein inequality tropp2015introduction, we have for any $t>0$
\begin{equation*}
{\mathbb{P}}\left(\left\|\dfrac{B^\top B}{n} - {\mathbb{E}}(\dfrac{B^\top B}{n})\right\|_2 > \dfrac{t}{n} \right) \leq 2M\cdot \exp\left(-\dfrac{t^2}{2k_D nM^{-1}+k_D t}\right).
\end{equation*}
Taking $t=\sqrt{12k_D nM^{-1}\log (M\vee p)}$ we have with probability at least $1-c(M\vee p)^{-4}$
\begin{equation}
\sqrt{\left\|B^\top B - {\mathbb{E}}(B^\top B)\right\|_2} \lesssim (nM^{-1}\log (M\vee p))^{1/4}.
\end{equation}
Together with the fact that $\sqrt{\|{\mathbb{E}}(B^\top B)\|_2}\lesssim \sqrt{nM^{-1}}$ by Proposition (ref), (ref) and (ref) imply $\|B\|_2 \lesssim \sqrt{nM^{-1}\log (pM)}$. Following the same procedures we have $ \|H\|_2 \lesssim \sqrt{nM^{-1}\log (pM)} $ and hence
\begin{equation}
\begin{aligned}
\|W\|_2 \leq \|B\|_2 + \|H\|_2 \lesssim \sqrt{nM^{-1}\log (pM)}.
\end{aligned}
\end{equation}
(ref) and (ref) imply that with probability at least $1-c(M\vee p)^{-4}$
\begin{equation}
\begin{aligned}
\left|\dfrac{2}{n}(\widehat{\omega}-\omega)^\top(W-\widehat{W})^\top W (\widehat{\omega}-\omega)\right| &\leq \dfrac{2}{n}\|\widehat{\omega}-\omega\|_2^2 \|W-\widehat{W}\|_2\|W\|_2 \\
&\lesssim \|\widehat{\omega}-\omega\|_2^2 a_n \sqrt{M^{-1}\log (pM)}.
\end{aligned}
\end{equation}
In addition, (ref) also implies
\begin{equation}
\begin{aligned}
&\ \ \ \ \left| \dfrac{2}{n}(\widehat{\omega}-\omega)^\top(W-\widehat{W})^\top X (\widehat{\theta}-\theta) \right| \\
&\leq 2\sqrt{(\widehat{\omega}-\omega)^\top\dfrac{(W-\widehat{W})^\top(W-\widehat{W})}{n}(\widehat{\omega}-\omega)}\sqrt{(\widehat{\theta}-\theta)^\top\dfrac{X^\top X}{n}(\widehat{\theta}-\theta)} \\
&\leq (\widehat{\omega}-\omega)^\top\dfrac{4C_\Sigma(W-\widehat{W})^\top(W-\widehat{W})}{c^*c_\Sigma n}(\widehat{\omega}-\omega) + (\widehat{\theta}-\theta)^\top\dfrac{c^* c_\Sigma X^\top X}{4C_\Sigma n}(\widehat{\theta}-\theta) \\
&\leq \|\widehat{\omega}-\omega\|^2_2\dfrac{4C_\Sigma\|W-\widehat{W}\|^2_2}{c^*c_\Sigma n} + (\widehat{\theta}-\theta)^\top\dfrac{c^*c_\Sigma X^\top X}{4C_\Sigma n}(\widehat{\theta}-\theta) \\
&\leq C\|\widehat{\omega}-\omega\|^2_2 a_n^2 + (\widehat{\theta}-\theta)^\top\dfrac{c^*c_\Sigma X^\top X}{4C_\Sigma n}(\widehat{\theta}-\theta) \\
\end{aligned}
\end{equation}
for some $C>0$. Besides, by Proposition (ref) we have with probability at least $1-(M\vee p)^{-4}$
\[ \|\dfrac{X^\top X}{n} - \dfrac{{\mathbb{E}}(X^\top X)}{n}\|_\infty \leq C_x \sqrt{\dfrac{\log (pM)}{n}} \]
for some $C_x>0$ large enough. Note that $((\widehat{\omega}-\omega),(\widehat{\theta}-\theta)^\top)^\top \in \mathcal{C}_0$ as defined in (ref),
\begin{equation}
\begin{aligned}
(\widehat{\theta}-\theta)^\top\dfrac{X^\top X}{n}(\widehat{\theta}-\theta) &\leq (\widehat{\theta}-\theta)^\top\dfrac{{\mathbb{E}}(X^\top X)}{n}(\widehat{\theta}-\theta) + \|\widehat{\theta}-\theta\|_1^2\|\dfrac{X^\top X}{n} - \dfrac{{\mathbb{E}}(X^\top X)}{n}\|_\infty \\
&\leq C_\Sigma \|\widehat{\theta}-\theta\|_2^2 + \left(\dfrac{2c_2^2R_2^2}{\lambda_Y^2}\|\widehat{\omega}-\omega\|_2^2 + 2c_3^2s\|\widehat{\theta}-\theta\|_2^2\right) \cdot C_x\sqrt{\dfrac{\log p}{n}} \\
&\leq 2C_\Sigma \|\widehat{\theta}-\theta\|_2^2 + \dfrac{2c_2^2R_2^2}{\lambda_Y^2}C_x\sqrt{\dfrac{\log p}{n}}\|\widehat{\omega}-\omega\|_2^2. \\
\end{aligned}
\end{equation}
Combining (ref), (ref), (ref) and (ref) we deduce that with probability at least $1-c(M\vee p)^{-4}$
\begin{equation}
\begin{aligned}
|\Delta_1| &\lesssim C\|\widehat{\omega}-\omega\|_2^2(a_n^2 + a_n\sqrt{M^{-1}\log (pM)}) + (\widehat{\theta}-\theta)^\top\dfrac{c^*c_\Sigma X^\top X}{4C_\Sigma n}(\widehat{\theta}-\theta) \\
&\leq C\|\widehat{\omega}-\omega\|_2^2(a_n^2 + a_n\sqrt{M^{-1}\log (pM)} + \dfrac{R_2^2}{\lambda_Y^2}\sqrt{\dfrac{\log p}{n}}) + \dfrac{c^*c_\Sigma \|\widehat{\theta}-\theta\|_2^2}{2} \\
&= o_p\left(M^{-1}\|\widehat{\omega}-\omega\|_2^2\right) + \dfrac{c^*c_\Sigma \|\widehat{\theta}-\theta\|_2^2}{2}.
\end{aligned}
\end{equation}
{\bf Step 3: Bound $\Delta_2$.} Note that each row in the $n\times(2M+p)$ matrix $U$ are independent bounded variables. By Bernstein-type inequality, with probabilty at least $1-c(M\vee p)^{-4}$
\[
\left\|\dfrac{U^\top U}{n}-\dfrac{\mathbb{E}(U^\top U)}{n}\right\|_\infty \lesssim \sqrt{\dfrac{\log(M\vee p)}{n}}
\]
and hence
\begin{equation*}
|\Delta_2| \lesssim \|\delta\|_1^2\sqrt{\dfrac{\log(M\vee p)}{n}}.
\end{equation*}
Note that
\[\begin{aligned}
\|\delta\|_1^2 &= \|\widehat{\omega}-\omega\|_1^2 + \|(\widehat{\theta}-\theta)_\mathcal{S}\|_1^2 + \|(\widehat{\theta}-\theta)_\mathcal{N}\|_1^2 \\
&\leq 2M\|\widehat{\omega}-\omega\|_2^2 + \|(\widehat{\theta}-\theta)_\mathcal{S}\|_1^2 + 2\dfrac{c_2^2R_2^2}{\lambda_Y^2}\|\widehat{\omega}-\omega\|_2^2 + 2c_3^2\|(\widehat{\theta}-\theta)_\mathcal{S}\|_1^2 \\
&= \left(2M + 2\dfrac{c_2^2R_2^2}{\lambda_Y^2}\right)\|\widehat{\omega}-\omega\|_2^2 + (1+2c_3^2)s\|\widehat{\theta}-\theta\|_2^2 \\
&\lesssim M \|\widehat{\omega}-\omega\|_2^2 + s\|\widehat{\theta}-\theta\|_2^2
\end{aligned}
\]
which implies
\begin{equation}
\begin{aligned}
|\Delta_2|
&\lesssim M \sqrt{\dfrac{\log (pM)}{n}}\|\widehat\omega-\omega\|_2^2 + s\sqrt{\dfrac{\log (pM)}{n}}\|\widehat\theta-\theta\|_2^2 \\
&= o_{\rm{a.s.}}\left(M^{-1}\|\widehat\omega-\omega\|_2^2 + \|\widehat\theta-\theta\|_2^2\right)
.
\end{aligned}
\end{equation}
proof[Proof of Lemma (ref)]
Define $\Sigma_F = {\mathbb{E}}\left(F_{i\cdot} F_{i\cdot}^\top \right)$, $S_F := \text{diag}(\Sigma_F)$ and recall that $F^{\rm std}_{i\cdot} = S_F^{-1/2}F_{i\cdot}$ in Assumption (ref). The eigenvalues of ${\mathbb{E}}[S_F^{-1/2}\Sigma_{F}S_F^{-1/2}] = {\mathbb{E}}[F^{\rm std}_{i\cdot} (F^{\rm std}_{i\cdot})^\top]$ are bounded away from zero and above by Assumption (ref). It then suffices to show that
\begin{equation}
\sup_g \left\|S_F^{-1/2}\left[\Sigma_F - \Sigma_{F|\mathcal{L}}\right] S_F^{-1/2} \right\|_2 = o_p(1)
\end{equation}
which implies that the eigenvalues of $S_F^{-1/2}\Sigma_{F|\mathcal{L}}S_F^{-1/2}$ are bounded away from zero and above. As $1 \lesssim_p \lambda_{\min}(S_F^{-1/2}) \leq \lambda_{\max}(S_F^{-1/2}) \lesssim_p M$, we deduce that $M^{-1} \lesssim_p \inf_g \lambda_{\min}(\Sigma_{F|\mathcal{L}})\leq \sup_g \lambda_{\max}(\Sigma_{F|\mathcal{L}}) \lesssim_p 1$ by ((ref)).
We then start proving ((ref)). By standard arguments,
\[\begin{aligned}
\|S_F^{-1/2}[\Sigma_F - \Sigma_{F|\mathcal{L}}]S_F^{-1/2} \|_2 &=\left\|S_F^{-1/2}{\mathbb{E}}_{\mathcal{L}}(F_{i\cdot}^{{\otimes2}}-\widehat F_{i\cdot}^{{\otimes2}})S_F^{-1/2}\right\|_2 \\ &\lesssim \left\|S_F^{-1/2}{\mathbb{E}}_{\mathcal{L}}(F_{i\cdot}-\widehat F_{i\cdot})^{{\otimes2}}S_F^{-1/2}\right\|_2 + \left\|{\mathbb{E}}_{\mathcal{L}}S_F^{-1/2}(F_{i\cdot}-\widehat F_{i\cdot})F_{i\cdot}^\top S_F^{-1/2}\right\|_2\\
&=: L_1 + L_2.
\end{aligned}\]
We first bound $L_1$. Given that $\max_{j\in[M_D]}{\mathbb{E}}(H_{ij}^2) \lesssim M$, $\max_{j\in[p_zM]}{\mathbb{E}}[(q^\prime(v_i)K_{ij})^2] \lesssim M$ and $\max_{j\in[p]}{\mathbb{E}}[(q^\prime(v_i)X_{ij})^2] \lesssim 1$, we deduce that
\[\begin{aligned}
L_1 &\lesssim \max_{j\in[M_D]}{\mathbb{E}}(H_{ij}^2)^{-1}\left\|{\mathbb{E}}_{\mathcal{L}}(H_{i\cdot}-\widehat H_{i\cdot})^{{\otimes2}}\right\|_2 + \max_{j\in[p]}{\mathbb{E}}[(q^\prime(v_i)X_{ij})^2]^{-1}\left\|{\mathbb{E}}_{\mathcal{L}}(\widehat q_i^\prime - q_i^\prime)^2X_{i\cdot}^{{\otimes2}}\right\|_2 + \\
&\ \ \ \ \max_{j\in[p_zM]}{\mathbb{E}}[(q^\prime(v_i)K_{ij})^2]^{-1}\left\|{\mathbb{E}}_{\mathcal{L}}(\widehat q_i^\prime - q_i^\prime)^2K_{i\cdot}^{{\otimes2}}\right\|_2 \\
&\lesssim M\left\|{\mathbb{E}}_{\mathcal{L}}(H_{i\cdot}-\widehat H_{i\cdot})^{{\otimes2}}\right\|_2 +\left\|{\mathbb{E}}_{\mathcal{L}}(\widehat q_i^\prime - q_i^\prime)^2X_{i\cdot}^{{\otimes2}}\right\|_2 + M\left\|{\mathbb{E}}_{\mathcal{L}}(\widehat q_i^\prime - q_i^\prime)^2K_{i\cdot}^{{\otimes2}}\right\|_2 \\
&=: L_{11} + L_{12} + L_{13}.
\end{aligned}\]
Bound $L_{11}$. By ((ref)),
\[\begin{aligned}
L_{11} = \left\|n^{-1}\sum_{i=1}^n M{\mathbb{E}}_{\mathcal{L}}(H_{i\cdot}-\widehat H_{i\cdot})^{{\otimes2}}\right\|_2 \leq n^{-1}M\sum_{i=1}^n\sum_{j=1}^{M} {\mathbb{E}}_{\mathcal{L}}(H_{ij}-\widehat H_{ij})^2 = o_p(1).
\end{aligned}\]
Since the $H$ and $\widehat H$ are not dependent on $g$, the $o_p(1)$ holds uniformly for all $g$.
Bound $L_{12}$. Note that
\begin{equation}
\begin{aligned}
(\widehat q_i^\prime - q_i^\prime)^2
&\lesssim \left|q^\prime(\widehat v_i) - q^\prime(v_i)\right|^2 + \left|\eta^\top \widehat H_{i\cdot}^\prime - q^\prime(\widehat v_i)\right|^2 + \left|(\widehat\eta^{ind} - \eta)^\top \widehat H_{i\cdot}^\prime\right|^2 \\
&=: \varDelta^q_{1i} + \varDelta^q_{2i} + \varDelta^q_{3i}.
\end{aligned}
\end{equation}
Define $\varDelta^q_{j}:=\sup_{i\in[n]} \varDelta^q_{ji}$ for $j=1,2,3.$ Eq. ((ref)) shows when $\widehat{\mathcal{V}}_n$ in (ref) holds,
\begin{equation}
\varDelta^q_1 \lesssim c_{1n} := \dfrac{n}{\log (pM)}M^{-4\gamma} + \dfrac{(s+M)^2\log (pM)}{n} + M^{-2\gamma} = o(1).
\end{equation}
When $\widehat{\mathcal{V}}_n$ in ((ref)) does not hold, we claim $\varDelta^q_1$ is uniformly bounded since $q^\prime(\cdot)$ is bounded over a compact interval. Thus
\begin{equation}
\begin{aligned}
\sup_g\left\|{\mathbb{E}}_{\mathcal{L}}\left[ \varDelta^q_1 X_{i\cdot}^{\otimes2}\right]\right\|_2
&\leq \sup_g\left\|{\mathbb{E}}_{\mathcal{L}}\left[ 1(\widehat{\mathcal{V}}_n) \varDelta^q_1 X_{i\cdot}^{\otimes2}\right]\right\|_2 + \sup_g\left\|{\mathbb{E}}_{\mathcal{L}}\left[ 1(\widehat{\mathcal{V}}_n^c) \varDelta^q_1 X_{i\cdot}^{\otimes2}\right]\right\|_2 \\
&\lesssim c_{1n}\cdot \|{\mathbb{E}}(X_{i\cdot} X_{i\cdot}^\top)\|_2 + \Pr(\widehat{\mathcal{V}}_n|\mathcal{L}) \cdot {\mathbb{E}}( \|X_{i\cdot}\|_2^2) \\
&\lesssim c_{1n} + \Pr(\widehat{\mathcal{V}}_n|\mathcal{L})\cdot p = o_p(1)
\end{aligned}
\end{equation}
where the third inequality applies that boundness of $X_{i\cdot}$ and the $o_p(1)$ applies (ref).
Furthermore, by Proposition (ref)
\begin{equation}
\Delta^q_2 = \sup_{v\in[a_v-\epsilon_v,b_v+\epsilon_v]}\left|\sum_{j\in[M]}\eta_j H_j^\prime(v) -q^\prime(v)\right|^2 \lesssim M^{-2\gamma+2}.
\end{equation}
By Proposition (ref) and the functional extensions stated in the beginning of Section (ref),
\begin{equation}
\varDelta^q_3 \leq \sup_{i\in[n]} \|\widehat\eta^ind - \eta\|_2^2\cdot \|\widehat H^\prime_{i\cdot}\|_2^2 \lesssim \|\widehat\eta^ind - \eta\|_2^2\cdot M^2.
\end{equation}
Consequently,
\begin{equation}
\begin{aligned}
\sup_g\left\|{\mathbb{E}}_{\mathcal{L}}\left[ ( \varDelta^q_2 + \varDelta^q_3 ) X_{i\cdot}^{\otimes2}\right]\right\|_2
&\lesssim \left( M^{-2\gamma+2}+\sup_g\|\widehat\eta^\text{ind} - \eta\|_2^2 M^2\right) \cdot \lambda_{\max}({\mathbb{E}}[X_{i\cdot} X_{i\cdot}^\top]) = o_p(1).
\end{aligned}
\end{equation}
The $o_p(1)$ applies the uniform error bound for $\widehat\eta^\text{ind}$ in (ref).
Therefore,
\[ \sup_g L_{12} = \sup_g \left\|{\mathbb{E}}_{\mathcal{L}}(\widehat q_i^\prime - q_i^\prime)^2X_{i\cdot}^{{\otimes2}}\right\|_2 = \sup_g \left\|{\mathbb{E}}_{\mathcal{L}}(\varDelta^q_1 + \varDelta^q_2 + \varDelta^q_3)X_{i\cdot}^{{\otimes2}}\right\|_2 = o_p(1).\]
\underline{Bound $L_{13}$}. Note that $\|M{\mathbb{E}}(K_{i\cdot} K_{i\cdot}^\top)\|_2 \lesssim 1$ by Proposition (ref), and
\begin{equation}
M \|K_{i\cdot}\|_2^2 \leq M\sup_{\{z_\ell\}}\sum_{\ell\in[p_z]}\sum_{j\in[M]} [K_{j\ell}(z_\ell)]^2 \leq M\sup_{\{z_\ell\}}\sum_{\ell\in[p_z]}\left(\sum_{j\in[M]} |K_{j\ell}(z_\ell)|\right)^2 = Mp_z.
\end{equation}
Using these two results for $K_{i\cdot}$, we can deduce
\[\begin{aligned}
\sup_g L_{13} &\lesssim \sup_g M\left\|{\mathbb{E}}_{\mathcal{L}}\left[ (\varDelta^q_1 + \varDelta^q_2 + \varDelta^q_3 ) K_{i\cdot}^{\otimes2}\right]\right\|_2 = o_p(1).
\end{aligned}\]
following similar arguments when bounding $L_{12}$,
\underline{We then bound $L_2$}. It follows that
\[\begin{aligned}
\sup_g L_2 &\lesssim \sup_g\sqrt{\left\|{\mathbb{E}}\left[S_F^{-1/2}(F_{i\cdot}-\widehat F_{i\cdot})^{\otimes2} S_F^{-1/2}\right]\right\|_2} \cdot \sqrt{\|{\mathbb{E}}(F^{\rm std}_{i\cdot}(F^{\rm std}_{i\cdot})^\top)\|_2}\\ &= \sup_g\sqrt{L_1}\cdot O(1) = o_p(1)
\end{aligned}\]
where the second step applies Assumption (ref) where $F^{\rm std}_{i\cdot} = S_F^{-1/2}F_{i\cdot}$. ((ref)) then follows the bounds of $L_1$ and $L_2$.
proof[Proof of Lemma (ref)] Note that
\[\sum_{i=1}^n (\widetilde r_i - \check r_i)^2 \leq \sum_{i=1}^n q^{\prime\prime}(v_i^*)^2(v_i - \widehat v_i)^4 + \sum_{i=1}^n (q^\prime(\widehat v_i) - \widehat q^\prime(\widehat v_i))^2\left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2.\]
By ((ref)), we deduce that
\begin{equation}
\begin{aligned}
\sum_{i=1}^n (v_i - \widehat v_i)^4 \lesssim_p \dfrac{n^3}{(\log (pM))^2}M^{-8\gamma} + \dfrac{(s+M)^4[\log (pM)]^2}{n} + n\cdot M^{-4\gamma} = o(1)
\end{aligned}
\end{equation}
under the conditions for Theorem (ref). Besides, $q^{\prime\prime}$ is uniformly bounded when $\widehat{\mathcal{V}}_n$ holds. Hence,
\[\begin{aligned}
\sum_{i=1}^n q^{\prime\prime}(v_i^*)^2(v_i - \widehat v_i)^4 \lesssim_p \sum_{i=1}^n (v_i - \widehat v_i)^4 = o_p(1).
\end{aligned}\]
It remains to show
\[ \sum_{i=1}^n (q^\prime(\widehat v_i) - \widehat q^\prime(\widehat v_i))^2\left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2 = o_p\left(\frac{1}{(\log n)^2}\right).\]
By ((ref)),
\[\begin{aligned}
(q^\prime(\widehat v_i) - \widehat q^\prime(\widehat v_i))^2 &\lesssim \varDelta^q_{1i} + \varDelta^q_{2i} + \varDelta^q_{3i}
\end{aligned}.\]
It suffices to show that
\[ \sup_g \sum_{i=1}^n (\varDelta^q_{1i} + \varDelta^q_{2i} + \varDelta^q_{3i} )\left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2 = o_p\left(\frac{1}{(\log n)^2}\right).\]
Note that by Proposition (ref),
\[\begin{aligned}
\sup_{i\in[n]} \left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2 &\lesssim \|\widehat\varphi^\text{ind}-\varphi\|_1^2 \sup_{i\in[n]} \|X_{i\cdot}\|_\infty^2 + \|\widehat\kappa^\text{ind} - \kappa\|_2^2 \cdot \sup_{i\in[n]} \|K_{i\cdot}\|_2^2 \\
&\lesssim_p c_{0n} := \dfrac{n}{\log (pM)} M^{-4\gamma} + \dfrac{(s+M)^2 \log (pM)}{n}.
\end{aligned}\]
By Assumption (ref), $c_{0n} = o(n^{-1/2}(\log n)^{-2})$. Furthermore, recall that $\varDelta^q_{j} = \sup_{i\in[n]}\varDelta^q_{ji}$ defined right after ((ref)). Equations ((ref)) and ((ref)) show that
\[ \varDelta^q_1
\lesssim_p \dfrac{n}{\log (pM)} M^{-4\gamma} + \dfrac{(s+M)^2\log (pM)}{n} + M^{-2\gamma} = o(n^{-1/2}) \]
and
\[ \varDelta^q_2
\lesssim M^{-2\gamma + 2} = o(n^{-1/2}). \]
Note that $\varDelta^q_1$ and $\varDelta^q_2$ are not dependent on $g$. Thus, \[\begin{aligned}
&\ \ \ \ \sup_g \sum_{i=1}^n (\varDelta^q_{1i} + \varDelta^q_{2i} )\left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2 \\
&\lesssim n(\varDelta^q_1 + \varDelta^q_2)c_{0n} = n\cdot o_p(n^{-1/2}) \cdot o_p(1/(\log n)^2) = o_p(1/(\log n)^2).
\end{aligned} \]
For $\varDelta_{3i}^q$, by the definition in (ref),
\[\varDelta_{3i}^q \lesssim \|\widehat\eta^\text{ind} - \eta\|_2^2\cdot \sup_{i\in[n]} \|\widehat H^\prime_{i\cdot} - H^\prime_{i\cdot}\|_2^2 + \left|(\widehat\eta^\text{ind} - \eta)^\top H^\prime_{i\cdot} \right|^2 =: \varDelta_{3}^{(1)} + \varDelta_{3i}^{(2)} \]
If suffices to show
\[\sup_g \sum_{i=1}^n (\varDelta_{3}^{(1)} + \varDelta_{3i}^{(2)} )\left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2 = o(1/(\log n)^2).\]
By (ref)
\begin{equation}
\sup_g \|\widehat\eta^ind - \eta\|_2^2 \lesssim_p M^{-2\gamma + 1} + \dfrac{(s+M)^2 \log (pM)}{n} \lesssim \dfrac{(s+M)^2 \log (pM)}{n}
\end{equation}
where the last inequality applies the fact that $M^{2\gamma+1} \gtrsim n$ under Assumption (ref). A similar argument applies to the last inequalities in the following two equations. By (ref),
\[\begin{aligned}
\sup_{i\in[n]} \|\widehat H^\prime_{i\cdot} - H^\prime_{i\cdot}\|_2^2 &\lesssim_p M^4\left(\dfrac{n}{\log (pM)}M^{-4\gamma} + \dfrac{(s+M)^2\log (pM)}{n}+M^{-2\gamma}\right) \\
&\lesssim \dfrac{M^4(s+M)^2\log (pM)}{n}.
\end{aligned}\]
Besides, by ((ref))
\[ n^{-1}\sum_{i=1}^n \left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2 \lesssim_p M^{-2\gamma} + \dfrac{(s+M)\log (pM)}{n} \lesssim \dfrac{(s+M)\log (pM)}{n}.\]
Thus
\[\begin{aligned}
&\ \ \ \ \sup_g \sum_{i=1}^n \varDelta_{3}^{(1)} \left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2\\
&\lesssim_p \dfrac{(s+M)^2\log (pM)}{n}\cdot \dfrac{M^4(s+M)^2\log (pM)}{n} \cdot (s+M)\log (pM) \\
&= \dfrac{M^4(s+M)^5[\log (pM)]^3}{n^2} \lesssim \dfrac{(s^9+M^9)[\log (pM)]^3}{n^2} = o(1/(\log n)^2)
\end{aligned}\]
where the $o(1/(\log n)^2)$ applies Assumption (ref). It remains to show
\[\sup_g \sum_{i=1}^n \left[ \varDelta^{(2)}_{3i} \left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2\right] = o_{p.g.}\left(\frac{1}{(\log n)^2}\right). \]
When we consider $\widehat\eta^\text{ind}$, $\widehat\varphi^\text{ind}$ and $\widehat\kappa^\text{ind}$ that generate the sigma-field $\mathcal{L}$ as given, $\varDelta^{(2)}_{3i}$ is a function of $v_i$, and $X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) $ is a function of $X_{i\cdot}$ and $Z_{i\cdot}$. By independence between $(X_{i\cdot}^\top, Z_{i\cdot}^\top)^\top$ and $v_i$ (conditionally on $\mathcal{L}$), we deduce that
\[n{\mathbb{E}}_{\mathcal{L}}\left( \varDelta^{(2)}_{3i} \left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2\right) = n{\mathbb{E}}_{\mathcal{L}}\left( \varDelta^{(2)}_{3i} \right){\mathbb{E}}_{\mathcal{L}}\left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2.\]
By Proposition (ref) and (ref),
\[n\cdot \sup_g {\mathbb{E}}_{\mathcal{L}}\left( \varDelta^{(2)}_{3i} \right)\lesssim_p n\|\widehat\eta^\text{ind} - \eta\|_2^2 \cdot \lambda_{\max}({\mathbb{E}}(H^\prime_{i\cdot} {H^\prime_{i\cdot}}^\top )) \lesssim_p (s+M)^2\log (pM)\cdot M \]
In addition, by (ref) and (ref)
\[\begin{aligned}
{\mathbb{E}}_{\mathcal{L}}\left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2
&\lesssim_p M^{-2\gamma} + \dfrac{(s+M)\log (pM)}{n} \lesssim \dfrac{(s+M)\log (pM)}{n}
\end{aligned}\]
where the last inequality applies the rate of $M$ in Assumption (ref). Thus,
\[\begin{aligned}
\sup_g n{\mathbb{E}}_{\mathcal{L}}\left( \varDelta^{(2)}_{3i} \left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2\right) \lesssim_p \dfrac{(s+M)^4(\log (pM))^2}{n} = o\left(\dfrac{1}{(\log n)^3}\right)
\end{aligned}\]
where the $o(1)$ applies Assumption (ref). By Markove inequality,
\[\begin{aligned}
&\ \ \ \ \sup_g \Pr\left\{\left| \sum_{i=1}^n \left[ \varDelta^{(2)}_{3i} \left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2\right]\right| > (\log n)^{-2.5}\Bigg|\mathcal{L}\right\} \\
&\leq (\log n)^{2.5}\sup_g n{\mathbb{E}}_{\mathcal{L}}\left( \varDelta^{(2)}_{3i} \left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2\right) = o_p(1).
\end{aligned}\]
Then
\[\begin{aligned}
&\ \ \ \ \sup_g \Pr\left\{\left| \sum_{i=1}^n \left[ \varDelta^{(2)}_{3i} \left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2\right]\right| > (\log n)^{-2.5} \right\} \\
&= \sup_g {\mathbb{E}}\left(\Pr\left\{\left| \sum_{i=1}^n \left[ \varDelta^{(2)}_{3i} \left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2\right]\right| > (\log n)^{-2.5} \Bigg|\mathcal{L}\right\}\right) \\
&\leq {\mathbb{E}}\left(\sup_g \Pr\left\{\left| \sum_{i=1}^n \left[ \varDelta^{(2)}_{3i} \left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2\right]\right| > (\log n)^{-2.5} \Bigg|\mathcal{L}\right\}\right) \\
&\to 0
\end{aligned}\]
where the limit applies the Bounded Convergence Theorem to supreme of the conditional probability as a random variable. Thus, $\sum_{i=1}^n \left[ \varDelta^{(2)}_{3i} \left[X_{i\cdot}^\top (\widehat\varphi^\text{ind}-\varphi ) + K_{i\cdot}^\top (\widehat\kappa^\text{ind} - \kappa) \right]^2\right] = o_{p.g.}(1)$. We complete the proof of Lemma (ref).
proof[Proof of Lemma (ref)]
We only show the Gaussian approximation error for $\mathbb{G}$ and the same arguments apply to $\mathbb{G}^\prime$. Let $\mathcal{G}$ be the sigma-field generated by $\{X_{i\cdot},Z_{i\cdot},v_i\}_{i\in[n]}$ and the LASSO estimators $\widehat\kappa^{\text{ind}}$, $\widehat\varphi^{\text{ind}}$, $\widehat\eta^{\text{ind}}$. Let $d_j^\varDelta = a_D + j\varDelta_J$ for $j=0,1,\cdots,J$ with some integer $J$ large enough, where $\varDelta_J := \frac{b_D - a_D}{J}$. Recall $\mathcal{D}=[a_D,b_D]$ and define $\mathcal{D}^\varDelta:=\{d_j^\varDelta\}_{0\leq j\leq J}$. Define
\[ \mathbb{G} := \sup_{d\in\mathcal{D}} \sum_{i=1}^n \zeta_i(d),\ \mathbb{G}^\varDelta := \sup_{d\in\mathcal{D}^\varDelta} \sum_{i=1}^n \zeta_i(d),\]
\[ \mathbb{G}^* := \sup_{d\in\mathcal{D}} \sum_{i=1}^n \zeta_i^*(d),\ \mathbb{G}^{*\varDelta} := \sup_{d\in\mathcal{D}^\varDelta} \sum_{i=1}^n \zeta_i^*(d)\]
and using triangular inequalities
\begin{equation}
| \mathbb{G} - \mathbb{G}^*| \leq |\mathbb{G} - \mathbb{G}^\varDelta| + |\mathbb{G}^* - \mathbb{G}^{*\varDelta}| + |\mathbb{G}^\varDelta - \mathbb{G}^{*\varDelta}|.
\end{equation}
Bound $|\mathbb{G} - \mathbb{G}^\varDelta| + |\mathbb{G}^* - \mathbb{G}^{*\varDelta}|$. Observe that $|\mathbb{G} - \mathbb{G}^\varDelta| \leq \sup_{|d-d^\prime| \leq \varDelta_J} |\sum_{i=1}^n \left(\xi_i(d) - \xi_i(d^\prime)\right) |$ and we further deduce that
\begin{equation}
\begin{aligned}
|\mathbb{G} - \mathbb{G}^\varDelta| &\leq \sup_{|d-d^\prime| \leq \varDelta_J}\|\widehat s(d)^{-1}B^\prime(d) - \widehat s(d^\prime)^{-1}B^\prime(d^\prime)\|_1 \cdot \left\|\sum_{i=1}^n \widehat\Omega_B \widehat F_{i\cdot} \varepsilon_i \right\|_\infty.
\end{aligned}
\end{equation}
We first bound the supreme of $L_1$ norm. We have $\|B^\prime(d)\|_1$ by Proposition (ref), $\widehat s(d) \asymp_{u.p.} M^{1.5}$ by ((ref)). Then
\[\begin{aligned}
&\ \ \ \ \|\widehat s(d)^{-1}B^\prime(d) - \widehat s(d^\prime)^{-1}B^\prime(d^\prime)\|_1 \\
&\leq \|B^\prime(d)\|_1\cdot\dfrac{\left|\widehat s(d^\prime)^2 - \widehat s(d)^2\right|}{\widehat s(d)\widehat s(d^\prime)[\widehat s(d) + \widehat s(d^\prime)]} + \dfrac{\|B^\prime(d) - B^\prime(d^\prime)\|_1}{\widehat s(d^\prime)} \\
&\lesssim_{u.p.} \dfrac{M\cdot \left|\widehat s(d^\prime)^2 - \widehat s(d)^2\right| }{M^{4.5}} + \dfrac{\|B^\prime(d) - B^\prime(d^\prime)\|_1}{M^{1.5}}\\
&=\dfrac{|(B^\prime(d) - B^\prime(d^\prime))^\top\widehat\Omega_B \widehat\Sigma_F \widehat\Omega_B^\top (B^\prime(d) + B^\prime(d^\prime))|}{M^{3.5}} + \dfrac{\|B^\prime(d) - B^\prime(d^\prime)\|_1}{M^{1.5}} \\
&=\left|\dfrac{\|B^\prime(d) - B^\prime(d^\prime)\|_1\cdot \|\widehat\Omega_B \widehat\Sigma_F \widehat\Omega_B^\top\|_\infty\cdot \|B^\prime(d) + B^\prime(d^\prime)\|_1}{M^{3.5}}\right| + \dfrac{\|B^\prime(d) - B^\prime(d^\prime)\|_1}{M^{1.5}} \\
&\lesssim_{u.p.} \dfrac{\|B^\prime(d) - B^\prime(d^\prime)\|_1}{M^{1.5}}.
\end{aligned}\]
By ((ref)),
\[\|\widehat s(d)^{-1}B^\prime(d) - \widehat s(d^\prime)^{-1}B^\prime(d^\prime)\|_1 \lesssim_{u.p.} M^{0.5}|d - d^\prime|.\]
and thus
\begin{equation}
\begin{aligned}
|\mathbb{G} - \mathbb{G}^\varDelta| &\lesssim_{u.p.} M^{0.5}\sup_{|d-d^\prime|\leq \varDelta_J}|d - d^\prime|\left\|\sum_{i=1}^n \widehat\Omega_B \widehat F_{i\cdot} \varepsilon_i \right\|_\infty \leq M^{0.5}\varDelta_J\left\|\sum_{i=1}^n \widehat\Omega_B \widehat F_{i\cdot} \varepsilon_i \right\|_\infty.
\end{aligned}
\end{equation}
(ref) together with (ref)
\begin{equation}
|\mathbb{G} - \mathbb{G}^\varDelta| \lesssim_{u.p.} \sqrt{n\log (pM)}M^{1.5}\cdot \Delta_J.
\end{equation}
Following exactly the same arguments,
\begin{equation}
|\mathbb{G}^* - \mathbb{G}^{*\varDelta}| \lesssim_{u.p.} \sqrt{n\log (pM)}M^{1.5}\cdot \Delta_J.
\end{equation}
Bound $| \mathbb{G}^\varDelta - \mathbb{G}^{*\varDelta}|$. Using chernozhukov2014gaussian, for any $t>0$
\begin{equation}
\begin{aligned}
\Pr\left\{| \mathbb{G}^\varDelta - \mathbb{G}^{*\varDelta}| > 16 t \Bigg|\mathcal{G}\right]\} \lesssim \dfrac{b_1\log (J\vee n)}{t^2} + \dfrac{(b_2+b_4)[\log (J\vee n)]^2}{t^3} + \dfrac{\log n}{n}
\end{aligned}
\end{equation}
where
\[\begin{aligned}
b_1 &:= {\mathbb{E}}_{\mathcal{G}}\left[\max_{d,d^\prime\in\mathcal{D}^\varDelta}\left|\sum_{i=1}^n(\zeta_{i}(d)\zeta_{i}(d^\prime) - {\mathbb{E}}_{\mathcal{G}}[\zeta_{i}(d)\zeta_{i}(d^\prime)])\right|\right],
\end{aligned}\]
\[\begin{aligned}
b_2 &:= {\mathbb{E}}_{\mathcal{G}}\left[\max_{d\in\mathcal{D}^\varDelta}\sum_{i=1}^n\left|\zeta_i(d)\right|^3\right]
\end{aligned}\]
and
\[\begin{aligned}
b_4 &:= \sum_{i=1}^n{\mathbb{E}}_{\mathcal{G}}\left[\max_{d\in\mathcal{D}^\varDelta}\left|\zeta_i(d)\right|^3\cdot 1\left(\max_{d\in\mathcal{D}^\varDelta}\left|\zeta_i(d)\right|>t/\log (J\vee n)\right)\right]
\end{aligned}\]
Note that
$\widehat F, \widehat\Omega, \widehat s(d)\in\mathcal{G}$ and
\begin{equation}
|\zeta_{i}(d)| = \left| \dfrac{B^\prime(d)^\top\widehat\Omega_B \widehat F_{i\cdot}\varepsilon_i}{\widehat s(d)}\right| \leq \dfrac{\|B^\prime(d)\|_1\cdot \|\widehat F\widehat\Omega_B^\top\|_\infty\cdot |\varepsilon_i|}{\widehat s(d)},
\end{equation}
and
\[\sup_g \sup_{d\in\mathcal{D}} \dfrac{\|B^\prime(d)\|_1\cdot \|\widehat F\widehat\Omega_B^\top\|_\infty }{\widehat s(d)} \lesssim_p \sqrt{M\log (pM) }\]
by the fact that $\sup_{d\in\mathcal{D}}\|B^\prime(d)\|_1\lesssim M$ by Proposition (ref), $\sup_g \|\widehat F\widehat\Omega_B^\top\|_\infty\lesssim M\sqrt{\log (pM) / n}$ by the definition of $\widehat\Omega_B^\top$ in ((ref)), and the lower bound for $\widehat s(d)$ for ((ref)).
Thus,
\[\begin{aligned}
\sup_g b_2 &\leq \sup_g \sup_{d\in\mathcal{D}} \left(\dfrac{\|B^\prime(d)\|_1\cdot \|\widehat F\widehat\Omega_B^\top\|_\infty}{\widehat s(d)}\right)^3\sum_{i=1}^n {\mathbb{E}}_{\mathcal{G}}|\varepsilon_i^3| \lesssim n\cdot (M\log (pM))^{1.5},
\end{aligned}\]
and similarly
\[\sup_g b_4 \lesssim_p n\cdot (M\log (pM))^{1.5}.\]
We then bound $b_1$. Using the Chebyshev's inequality, for any $d,d^\prime\in\mathcal{D}^\varDelta$ and $\tau >0$,
\[\begin{aligned} &\ \ \ \ \sup_g \Pr\left( n^{-1/2}\left|\sum_{i=1}^n(\zeta_{i}(d)\zeta_{i}(d^\prime) - {\mathbb{E}}_{\mathcal{G}}[\zeta_{i}(d)\zeta_{i}(d^\prime)])\right|>\tau \Bigg|\mathcal{G}\right) \\
&\leq \dfrac{1}{n\tau^{2}} \sup_g {\mathbb{E}}_{\mathcal{G}}\left[ \left|\sum_{i=1}^n(\zeta_{i}(d)\zeta_{i}(d^\prime) - {\mathbb{E}}_{\mathcal{G}}[\zeta_{i}(d)\zeta_{i}(d^\prime)])\right|^{2}\right] \\
&= \dfrac{1}{n\tau^{2}}\sup_g \sum_{i=1}^n {\mathbb{E}}_{\mathcal{G}} \left( \zeta_{i}(d)\zeta_{i}(d^\prime) - {\mathbb{E}}_{\mathcal{G}}[\zeta_{i}(d)\zeta_{i}(d^\prime)] \right)^{2} \\
&\leq \dfrac{1}{n\tau^{2}}\sup_g \sum_{i=1}^n\left({\mathbb{E}}_{\mathcal{G}}\left[ \zeta_{i}(d)^2\zeta_{i}(d^\prime)^2\right] \right) \\
&\lesssim \sup_g \sup_{d\in\mathcal{D}} \left(\dfrac{\|B^\prime(d)\|_1\cdot \|\widehat F\widehat\Omega_B^\top\|_\infty}{\widehat s(d)}\right)^2 n^{-1}\sup_g \sum_{i=1}^n {\mathbb{E}}_{\mathcal{G}}(\xi_i(d^\prime)^2\varepsilon_i^2) \\ &\lesssim_p \dfrac{M[\log (pM)]}{\tau^2},\\
\end{aligned}\]
where the fourth row applies ((ref)), and the last step applies \[
\begin{aligned}
\sum_{i=1}^n {\mathbb{E}}_{\mathcal{G}}(\xi_i(d^\prime)^2\varepsilon_i^2) &= \sum_{i=1}^n \dfrac{(B^\prime(d)^\top \widehat\Omega_B\widehat F_{i\cdot} )^2}{\widehat s(d)^2 }{\mathbb{E}}_{\mathcal{G}}(\varepsilon_i^4) \lesssim \sum_{i=1}^n \dfrac{(B^\prime(d)^\top \widehat\Omega_B\widehat F_{i\cdot} )^2}{\widehat s(d)^2 } = n
\end{aligned}\]
where the second step applies the bounded fourth condition moment of $\varepsilon_i$, and the last step applies the definition of $\widehat s(d)$. By union bound, \[\begin{aligned} &\sup_g \Pr\left( n^{-1/2}\max_{d,d^\prime\in\mathcal{D}^\varDelta}\left|\sum_{i=1}^n(\zeta_{i}(d)\zeta_{i}(d^\prime) - {\mathbb{E}}_{\mathcal{G}}[\zeta_{i}(d)\zeta_{i}(d^\prime)])\right|>\tau \Bigg|\mathcal{G}\right) &\lesssim_p \dfrac{J^2 M[\log (pM)]}{\tau^2}\\
\end{aligned}\]
and thus
\[\begin{aligned} &\ \ \ \ \sup_g \Pr\left( (nJ^2M)^{-1/2}(\log (pM))^{-1}\max_{d,d^\prime\in\mathcal{D}^\varDelta}\left|\sum_{i=1}^n(\zeta_{i}(d)\zeta_{i}(d^\prime) - {\mathbb{E}}_{\mathcal{G}}[\zeta_{i}(d)\zeta_{i}(d^\prime)])\right|>\tau \Bigg|\mathcal{G}\right) \\
&\lesssim_p \dfrac{J^2 M[\log (pM)]^2}{ (\sqrt{J^2 M}\log (pM)\tau)^2} = \dfrac{1}{\tau^2}.\\
\end{aligned}\]
By ${\mathbb{E}}(x) = \int_0^\infty \Pr(x > e)de$ for any random variable $x$,
\[\begin{aligned}
&\ \ \ \ (nJ^2M)^{-1/2}(\log (pM))^{-1}\cdot \sup_g b_1 \\
&\leq \int_0^\infty \sup_g \Pr\left( (nJ^2M)^{-1/2}(\log (pM))^{-1}\max_{d,d^\prime\in\mathcal{D}^\varDelta}\left|\sum_{i=1}^n(\zeta_{i}(d)\zeta_{i}(d^\prime) - {\mathbb{E}}_{\mathcal{G}}[\zeta_{i}(d)\zeta_{i}(d^\prime)])\right|>\tau \Bigg|\mathcal{G}\right) d\tau \\
&\lesssim_p 1 + \int_1^\infty \dfrac{1}{\tau^2} d\tau = 2
\end{aligned}\]
and hence $\sup_g b_1\lesssim_p \sqrt{nM}J\log (pM)$. Let $t = \sqrt{n}t_n$. Using (ref) and the bounds for $b_1$, $b_2$, and $b_4$, we have
\[\begin{aligned}
\sup_g \Pr\left\{n^{-1/2}|\mathbb{G}^\varDelta-\mathbb{G}^{*\varDelta}|>16t_n\Bigg|\mathcal{G}\right\} \lesssim_p \dfrac{\sqrt{M}J\log (J\vee n)}{\sqrt{n}t_n^2} + \frac{(M\log (pM))^{1.5}[\log (J\vee n)]^2}{\sqrt{n}t_n^3} + \frac{\log n}{n}.
\end{aligned}\]
Let $J \asymp M^{1.5}\log n$ and $t_n = (\log n)^{-1/2}$, we have
\[\begin{aligned}
\sup_g \Pr\left\{n^{-1/2}|\mathbb{G}^\varDelta-\mathbb{G}^{*\varDelta}|>\frac{16}{\sqrt{\log n}}\Bigg|\mathcal{G}\right\} \lesssim_p \dfrac{M^2(\log n)^2}{\sqrt{n}} + \frac{M^{1.5}[\log (pM)]^5}{\sqrt{n}} + \frac{\log n}{n} = o(1).
\end{aligned}\]
Then
\[\begin{aligned}
\sup_g \Pr\left\{n^{-1/2}|\mathbb{G}^\varDelta-\mathbb{G}^{*\varDelta}|>\frac{16}{\sqrt{\log n}}\right\} &\leq {\mathbb{E}}\left(\sup_g \Pr\left\{n^{-1/2}|\mathbb{G}^\varDelta-\mathbb{G}^{*\varDelta}|>\frac{16}{\sqrt{\log n}}\Bigg|\mathcal{G}\right\} \right)
\\&\to 0
\end{aligned}\]
where the limit uses the Bounded Convergence Theorem for the supreme of conditional probability. Thus
\begin{equation}
n^{-1/2}|\mathbb{G}^\varDelta-\mathbb{G}^{*\varDelta}| \lesssim_{u.p.} \frac{1}{\sqrt{\log n}}.
\end{equation}
Also by ((ref)) and ((ref)) we have \begin{equation}
n^{-1/2}|\mathbb{G} - \mathbb{G}^\varDelta| + n^{-1/2}|\mathbb{G}^* - \mathbb{G}^{*\varDelta}| \lesssim_{u.p.} \dfrac{1}{\sqrt{\log n}}.
\end{equation}
Thus, by ((ref)), ((ref)) and ((ref))
\begin{equation}
n^{-1/2}|\mathbb{G}-\mathbb{G}^*| \lesssim_{u.p.} \frac{1}{\sqrt{\log n}}.
\end{equation}
We complete the proof of Lemma (ref).
Additional Simulation Results
This section includes the robustness simulation results omitted from the main text. The results in this section all use full-sample inference as described in the main text.
As mentioned in Section (ref) of the main text, we first check the performance of our methodology under different choices of $M_D$. As shown in Tables (ref) and (ref), a larger $M_D$ does not benefit the inference in terms of coverage but produces a wider confidence band due to additional variances from a larger dimension of B-Spline functions. It shows that $M_D = 5$ is a reasonable choice.
Next, we check the robustness of our method under different types of distributions that violate the compactness assumption. We set $U_{ji}\sim N(0,12^{-1/2})$ and $v_i\sim N(0,1)$ to maintain the same variances of the data from bounded distributions in the main text. Other settings are unchanged. Table (ref) shows that the performance of the proposed method is robust to violations of compact supports.
table[table omitted — 2,573 chars of source]
table[table omitted — 2,643 chars of source]
table[table omitted — 4,721 chars of source]