Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Estimation and Uniform Inference in Sparse High-Dimensional Additive Models
frontmatter\runtitle{Inference in High-Dimensional Additive Models}
\runauthor{Bach, Klaassen, Kueck, Spindler}
\thankstext{T1}{Version: April, 2024.}
\begin{aug}
,
,\\
\and
\thankstext{T2}{Corresponding author: [email removed].}
\address{University of Hamburg, Heinrich Heine University Düsseldorf}
\end{aug}
\begin{abstract}
We develop a novel method to construct uniformly valid confidence bands for a nonparametric component $f_1$ in the sparse additive model $Y=f_1(X_1)+\ldots + f_p(X_p) + \varepsilon$ in a high-dimensional setting. Our method integrates sieve estimation into a high-dimensional Z-estimation framework, facilitating the construction of uniformly valid confidence bands for the target component $f_1$. To form these confidence bands, we employ a multiplier bootstrap procedure. Additionally, we provide rates for the uniform lasso estimation in high dimensions, which may be of independent interest.
Through simulation studies, we demonstrate that our proposed method delivers reliable results in terms of estimation and coverage, even in small samples.
\end{abstract}
\begin{keyword}[class=MSC]
\kwd[Primary ]{62G08}
\kwd[; Secondary ]{62-07}
\kwd{41A15}
\end{keyword}
\begin{keyword}
\kwd{Additive Models}
\kwd{High-dimensional Setting}
\kwd{Z-estimation}
\kwd{Double Machine Learning}
\kwd{Lasso}
\end{keyword}
Introduction
Nonparametric regression offers a way to estimate the relationship $f$ between a target variable $Y$ and input variables $X=(X_1,\ldots,X_p)^T$ without imposing restrictive functional assumptions. This relationship is expressed as
\[
Y=f(X_1,\ldots,X_p) + \varepsilon,
\]
where $\varepsilon$ denotes the random error term satisfying ${\mathrm{E}}[\varepsilon|X]=0$. However, in settings with a large number of regressors $p$, especially in cases where $p$ exceeds the number of observations $n$, the well-known curse of dimensionality makes it practically impossible to estimate the regression function $f(X_1, \ldots, X_p)$. A popular approach to overcome this limitation of nonparametric estimation in practice is to impose an additive structure of the regression function, leading to the following additive models
align[align omitted — 70 chars of source]
where $f_j(\cdot)$, $j=1,\ldots,p$, are smooth univariate functions. However, even within this framework, challenges can arise when incorporating a large number of regressors or components, or numerous sieve terms to model the component functions $f_j$. While in many studies, the dimension of the problem is fixed, we follow the recent literature and allow the number of parameters to grow asymptotically with and exceed the sample size. This approach provides a reasonable approximation for various empirical settings involving complex data.
The concept of additive models has its origins in the works of friedman1981, stone1985 and hastie1990. Readers seeking an introduction to additive models will find useful textbook treatments in hastie1990 and wood2017. A seminal contribution to the analysis of semi- and nonparametric models, including additive models, was made by Robinson1988. Since then, numerous studies focusing on nonparametric estimation of these models have been published, with contributions from LintonNielsen, Andrews1990, FanJiang, Horowitz2004, HorowitzMammen2006, Mammen2006, Horowitz2012 and many more.
Traditionally, the literature has concentrated on estimation and inference in low-dimensional settings where $p$ is fixed. Recent years, however, have witnessed considerable progress in the understanding and analysis of high-dimensional additive models that allow the number of components to grow with the sample size. For example, the theoretical literature has provided insights into estimation rates in high-dimensional additive models, as in the work of sardy2004, lin2006 and many others ravikumar2009, meier2009, huang2010, koltchinskii2010, kato2012two, petersen2016, lou2016.
The theoretical results in high-dimensional settings rely on a sparsity assumption, which requires that only a small number $s$ of coefficients or components are relevant, meaning they are non-zero. While the number of parameters or covariates exceed the sample size, or the number of parameters can increase along with the sample size, sparsity requires that the number of relevant parameters remains smaller than the sample size. However, it may still increase as the sample size grows, thus introducing additional structure.
While high-dimensional additive models have been studied extensively, there has been limited focus on valid inference within them, particularly in terms of constructing valid hypothesis tests or confidence regions. Approaches to construct confidence bands have been proposed by haerdle1989, sun1994, fan2000, claeskens2003 and zhang2010, but only in the widely studied fixed-dimensional setting. Only recently have new results been derived regarding valid inference in high-dimensional additive models. In the remainder of the introduction section, we will review these recent advancements and highlight our contribution to this emerging literature.\\
The primary aim of our paper is to provide a method for constructing uniformly valid inference and confidence bands in sparse high-dimensional models in the sieve framework. In doing so, we contribute to the growing literature on high-dimensional inference in additive models, especially that on debiased/double machine learning. The double machine learning approach belloni2014, Chernozhukov2018 offers a general framework for uniformly valid inference in high-dimensional settings.
Similar methods, such as those proposed by vandegeer2014 and zhang2014, have also produced valid confidence intervals for low-dimensional parameters in high-dimensional linear models. These studies are based on the so-called debiasing approach, which provides an alternative framework for valid inference. The framework entails a one-step correction of the lasso estimator, resulting in an asymptotically normally distributed estimator of the low-dimensional target parameter. For a survey on post-selection inference in high-dimensional settings and its generalizations, we refer to spindler2015.\\
In research closely related to ours, Kozbur2021 proposes a post-nonparametric double selection approach for a scalar functional of a single component.
We consider the same setting as Kozbur2021, which encompasses a more general additively separable model
\[
Y= f_{1}(X_1)+ f_{-1}(X_2,\dots,X_p) + \varepsilon,
\]
which includes the additive model
\[
Y= f_1(X_1) + \ldots + f_p(X_p) + \varepsilon.
\]
Kozbur2021 focuses on inference on functionals of the form $\theta=a(f_1)$ resulting in pointwise confidence intervals based on a penalized series estimator. Although many of these functionals are of interest, Kozbur2021 approach cannot be easily extended to uniform inference because it considers only differentiable functionals. In contrast, our framework allows us to extend these results and clarify the underlying assumptions. First, building on recent results for inference on high-dimensional target parameters by belloni2018uniformly and Belloni2014b, we are able to construct uniformly valid confidence bands for the whole function $f_1$. Second, the post-nonparametric double selection approach in Kozbur2021 relies on very technical assumptions concerning model selection properties of lasso in Assumption 7 that are difficult to verify (see also the discussion in lu2019 and Example 2 in Kozbur2021). We have much simpler and less technical assumptions since we do not need to rely on model selection properties of lasso. We only make use of the rate of convergence of our lasso estimate to orthogonalize our estimation procedure. In this context, we provide tailored results for uniform lasso estimation that are necessary for establishing valid inference.\\
In a recent study based on the previously mentioned debiasing approach by zhang2014, gregory2016 propose an estimator for the first component $f_1$ in a high-dimensional additive model in which the number of additive components $p$ may increase with the sample size. The estimator is constructed in two steps. The first step entails constructing an undersmoothed estimator based on near-orthogonal projections with a group lasso bias correction. A debiased version of the first step estimator is used to generate pseudo responses $\hat{Y}$. These pseudo responses are then used in the second step, which involves applying a smoothing method to a nonparametric regression problem with $\hat{Y}$ and covariates $X_1$. Under sparsity assumptions concerning the number of nonzero additive components, the so-called oracle property is shown. Accordingly, the proposed estimator in gregory2016 is asymptotically equivalent to the oracle estimator, which is based on the true functions $f_2,\dots,f_p$. The asymptotic properties of the oracle estimator are well understood and carry over to the proposed debiasing estimate, including the methodology for constructing uniformly valid confidence intervals for $f_1$.
Importantly, gregory2016 do not explicitly focus on inference, and their analysis requires much stronger assumptions to obtain the oracle property. For example, these assumptions include normally distributed errors independent of $X$, as well as a bounded support of $X$. Similar to our framework, a large set of basis functions is chosen in their work, such as polynomials or splines, to approximate the components $f_1$ and $f_{-1}$. One distinctive feature of our work compared to gregory2016 is that we allow the degree of approximating functions to grow to infinity with increasing sample size.\\
A procedure explicitly addressing the construction of uniformly valid confidence bands for the components in high-dimensional additive models has been developed by lu2019. The authors emphasize that achieving uniformly valid inference in these models is challenging due to the difficulty of directly generalizing the ideas from the fixed-dimensional case. Whereas confidence bands in the low-dimensional case are mostly built using kernel methods, the estimators for high-dimensional sparse additive models typically rely on sieve estimators based on dictionaries. To derive their results, lu2019 combine both kernel and sieve methods to draw upon the advantages of each, resulting in a kernel-sieve hybrid estimator. This is a two-step estimator with tuning parameters for kernel estimation and sieves estimation, such as the bandwidth and penalization levels, which must be chosen by cross-validation. Because of the local structure of the hybrid estimator, the framework
of lu2019 differs from ours in that they consider an additive local approximation model with sparsity (ATLAS), in which they need to impose a local sparsity structure.
The advantage of our proposed estimator is that we do not have to leave the sieves framework while establishing the uniform validity of the resulting confidence bands. Interestingly, lu2019 conclude that "it is challenging to study the uniform confidence
band through pure sieve-type approaches" (p. 4). However, this is exactly our main contribution: achieving uniform inference within a sieve framework. We accomplish this by framing the problem as one of high-dimensional Z-estimation, utilizing recent results from belloni2018uniformly. Additionally, we offer a theory-driven approach for choosing the penalization level for the lasso estimation steps, eliminating the need for computationally intense cross-validation procedures. Similar to gregory2016, lu2019 assume normally distributed errors that are independent of $X$. In contrast, our model framework allows us to dispense with the normality assumption and only requires the existence of the fourth moment of the error term. Moreover, our main results are also compatible with a heteroskedastic error. Additionally, we can overcome the requirement outlined in lu2019 that the number of relevant regressors is bounded. Instead, our model allows the number of relevant regressors to grow to infinity with increasing sample size.
Organization of the Paper
The paper is organized as follows. Section (ref) introduces and motivates the main regression problem in a high-dimensional additive model. Section (ref) presents the estimation method. In Section (ref), the main result is provided. Section (ref) presents a simulation study, highlighting the small sample properties and implementation of our proposed method. Section (ref) provides the proof of the main theorem. The Appendix includes additional technical material. Appendix (ref) presents a general result for uniform inference on a high-dimensional linear functional. Appendix (ref) provides results in terms of uniform lasso estimation rates in high-dimensions which might be of independent interest. Computational details and additional simulation results are presented in Appendix (ref) and Appendix (ref).
Notation
Throughout the paper, we consider a random element $W$ taken from a common probability space $(\Omega,\mathcal{A},P)$.
We denote by $P\in \mathcal{P}_n$ a probability measure out of a large class of probability measures, which may vary with the sample size (because the model is allowed to change with $n$). By $\mathbb{P}_n$ we denote the empirical probability measure. Additionally, let $\mathbb{E}$ and $\mathbb{E}_n$ be the expectations with regard to $P$, and $\mathbb{P}_n$, respectively. $\mathbb{G}_n(\cdot)$ denotes the empirical process
align*[align* omitted — 108 chars of source]
for a class of suitably measurable functions $\mathcal{F}:\mathcal{W}\to \mathbb{R}$. $\|\cdot\|_{P,q}$ denotes the $L^q(P)$-norm.
Furthermore, $\|v\|_1=\sum_{l=1}^p |v_l|$ denotes the $\ell_1$-norm, $\|v\|_2=\sqrt{v^Tv}$ the $\ell_2$-norm and $\|v\|_0$ equals the number of non-zero components of a vector $v\in\mathbb{R}^p$. We define $v_{-l}:=(v_1,\dots,v_{l-1},v_{l+1}\dots,v_p)^T\in\mathbb{R}^{p-1}$ for any $1\le l\le p$. $\|v\|_\infty=\sup_{l=1,\dots,p }|v_l|$ denotes the $\sup$-norm. Let $c$ and $C$ denote positive constants that are independent of $n$ with values that may change at each appearance. The notation $a_n\lesssim b_n$ means $a_n\le Cb_n$ for all $n$ and some $C$. Furthermore, $a_n=o(1)$ denotes that there exists a sequence $(b_n)_{\ge 1}$ of positive numbers such that $|a_n|\le b_n$ for all $n$, where $b_n$ is independent of $P\in\mathcal{P}_n$ for all $n$, and $b_n$ converges to zero. Finally, $a_n=O_P(b_n)$ means that for any $\epsilon>0$ there exists a $C$ such that $P(a_n>Cb_n)\le\epsilon$ for all $n$.
Setting
Motivation and Illustration
Before introducing the formal setting in Section (ref), we would like to motivate the basic ideas with a simplified example. We consider an additive model with two covariables
align[align omitted — 65 chars of source]
Our goal is to perform valid inference on the object of interest $f_1$. In other words, we would like to provide a uniform confidence band for $f_1$ as illustrated in Figure (ref). Hence, we treat $f_2$ in (ref) as a nuisance function.
Next, we assume that it is possible to represent the two components through an approximately linear representation. For the first component, the representation is given by
$$f_1(X_1) = \theta_0^T g(X_1) + b_1(X_1).$$
Here, $g(X_1)$ is a basis (e.g., a spline basis, sieve terms or a polynomial series) consisting of $d_1$ terms, $\theta_0$ is the corresponding coefficient vector and $b_1(X_1)$ is an approximation error.
For the second component, we assume an analogous representation
align[align omitted — 76 chars of source]
where the basis $h(X_2)$ consists of $d_2$ terms.
We allow the dimensions $d_1$ and $d_2$ to grow with the sample size in order to establish the asymptotic results in high-dimensional settings. Accordingly, the number of components $p$ can also grow with the sample size.\footnote{However, for the ease of notation we will later subsume them in the component $f_{-1}(X_2,\dots,X_p)$.}
To derive a uniformly valid confidence band, estimation and inference of
align[align omitted — 111 chars of source]
is required. In a naive approach to estimate the first component $\theta_{0,1}$ of the vector $\theta_0$, one might employ lasso or other machine learning methods to select the relevant components of the basis expansion terms in $\theta_0$ and $\beta_0$ in the regression model ((ref)). A possible second step could involve estimating a regression of the dependent variable on all components that have been selected in the lasso estimation step. The final estimator for the first component $\theta_{0,1}$ could then be obtained from this regression, and this procedure could be repeated iteratively for all other components in $f_1(x_1)$.
figure[figure omitted — 896 chars of source]
However, this approach is problematic because it generally fails to yield valid results. In other words, using the lasso estimator necessitates accounting for variable selection, leading to the non-standard problem of post-selection inference. Intuitively, the naive procedure based on a single lasso estimation step could result in selection mistakes that invalidate the resulting confidence interval for $\theta_{0,1}$. These critical selection mistakes will not occur for variables with high predictive power for the dependent variable $Y$, but rather refer to variables that are potentially highly correlated with the basis expansion term $g_1()$. This could lead to an omitted variable bias, preventing the estimator from asymptotically converging to a normal distribution. To overcome this limitation of post-selection inference, the statistical literature has developed the double machine learning and the debiasing approach, as mentioned in the introduction.
To address the potential bias introduced by lasso estimation, belloni2014 propose including an auxiliary regression for the corresponding covariate of the target parameter. Here, we consider
align[align omitted — 61 chars of source]
where $\nu$ is an error term and $Z_{-1}$ is defined as
$$Z_{-1}=(g_2(X_1),\dots,g_{d_1}(X_1),h_1(X_{-1}),\dots,h_{d_2}(X_{-1}))^T.$$
Later, we will also allow for an approximation error in this equation. belloni2014 propose including in the final regression not only the covariates selected in the first step of the naive approach but also augmenting this set of variables with Lasso-selected regressors from the auxiliary regression. This procedure is equivalent to constructing a so-called Neyman orthogonal moment function with respect to the nuisance part. This is essential for ensuring valid post-selection inference for the first component of the vector $\theta_0$. In Section (ref), we will provide more details about this property. Heuristically, the additional regression step in Equation (ref) will lead to robustness against moderate selection mistakes.
It can be shown formally, that this procedure implements an orthogonal moment equation
$$\mathbb{E}\left[\psi_1(W,\theta_{0,1},\eta_{0,1})\right]=0,$$
where the first component of $\theta_0$ is our target parameter and all other involved parameters are considered nuisance parameters, denoted $\eta_{0,1}$.
belloni2014 developed an approach for valid inference for one parameter. In high-dimensional additive models, the major technical challenge arises from the need to conduct inference for the potentially high-dimensional vector $\theta_0$. In other words, the number of elements in $\theta_0$ for which we would like to construct a valid confidence region is allowed to grow with the sample size. Each component of $\theta_0$, $\theta_{0,l}$ with $\l=1, \ldots, d_1$, is determined by an orthogonal moment condition and we will demonstrate how uniformly valid confidence bands can be constructed by embedding into a high-dimensional Z-estimation framework. Finally, we illustrate how the estimation of $\theta_0$ can be translated into uniformly valid confidence bands for the target function $f_1\approx \theta_0^T g(\cdot)$ using a multiplier bootstrap procedure.\\
Figure (ref) illustrates the use of our estimation procedure, which is described in Section (ref), by providing a preview of our simulation studies in Section (ref). Our estimation methods yield an unbiased estimate for the target component $f_1$, as well as corresponding $(1-\alpha)$ confidence interval that is valid over a compact interval of $X_1$. In the next section, we consider a more general additively separable model and introduce a general formulation of the underlying problem.
Formal Setting
Consider the following nonparametric additively separable model
$$Y=f(X)+\varepsilon=f_1(X_1)+f_{-1}(X_{-1}) + \varepsilon$$
with $\mathbb{E}[\varepsilon|X]=0$ and $c\le \operatorname{Var}(\varepsilon|X)\le C$. Let the scalar response $Y$ and features $X=(X_1,\dots,X_p)$ take values in $\mathcal{Y}$ and in $\mathcal{X}=\mathcal{X}_1\times\dots\times\mathcal{X}_p$, respectively. We assume the observation of $n$ i.i.d. copies $(W^{(i)})_{i=1}^n= (Y^{(i)},X^{(i)})_{i=1}^n$ of $W=(Y,X)$, with the number of covariates $p$ allowed to grow with the sample size $n$. To ensure identifiability, we assume $\mathbb{E}[f_{1}(X_{1})]=0$. Our aim is to construct uniformly valid confidence regions for the first nonparametric component of the regression function. In other words, we want to find functions $\hat{l}(x)$ and $\hat{u}(x)$ fulfill
align*[align* omitted — 91 chars of source]
Here, $I\subseteq\mathcal{X}_1$ is a bounded interval of interest in which we want to conduct inference. We approximate $f_1$ and $f_{-1}$ using linear combinations of approximating functions $g_1,\dots,g_{d_1}$ and $h_1,\dots,h_{d_2}$, respectively. Define
$$g(x_1):=(g_1(x_1),\dots,g_{d_1}(x_1))^T$$
for $x_1\in\mathbb{R}^{}$ and
$$h(x_{-1}):=(h_1(x_{-1}),\dots,h_{d_2}(x_{-1}))^T$$
for $x_{-1}\in\mathbb{R}^{p-1}$.
It is important to note that we allow the number of approximating functions $d_1$ and $d_2$ to increase with sample size.
Assume that the approximations are given by
align[align omitted — 67 chars of source]
where $\theta_{0,l}\in\Theta_l$ and analogously
align[align omitted — 79 chars of source]
where $b_1$ and $b_2$ denote the error terms. Additionally, it is convenient to define the union of approximation functions
$$z(x):=(g_1(x_1),\dots,g_{d_1}(x_1),h_1(x_{-1}),\dots,h_{d_2}(x_{-1}))^T$$
for $x \in\mathbb{R}^p$, where we abbreviate $Z:=z(X)$.
As a common assumption in high-dimensional statistics, we will impose sparsity in the linear approximation of $f$, meaning that only a relatively small number of the coefficients in $(\theta_0,\beta_0)$, denoted $s_1$, are non-zero which implies that only a subset of approximation functions in $Z$ is needed to approximate $f_1$ and $f_{-1}$ sufficiently well. Of course, the identity of these relevant approximation functions is not known. Sparsity can occur because of two different reasons: Either we have many irrelevant covariables $X_2,\ldots,X_p$ in our model or we have sparse representations of the regression functions.
For each element $g_l$ of $g$, we also consider
align[align omitted — 143 chars of source]
with
$\mathbb{E}[\nu^{(l)}Z_{-l}] = 0$, such that $(\gamma^{(l)}_0)^TZ_{-l}$ is the linear projection of $g_l(X_1)$ onto $Z_{-l}$. Here, $\gamma^{(l,1)}_0$ denotes the sparse part of the coefficient vector $\gamma^{(l)}_0$.
The auxiliary regression ((ref)) is used to construct an orthogonal score function for valid inference in a high-dimensional setting, as described in Section (ref).
Estimating $$f_1(\cdot)\approx \theta_0^Tg(\cdot)$$
can be recast into a general Z-estimation problem of the form
$$\mathbb{E}\left[\psi_l(W,\theta_{0,l},\eta_{0,l})\right]=0,\quad l\in1,\dots, d_1$$
with target parameter $\theta_0$ where the score functions are defined by
align*[align* omitted — 145 chars of source]
Here,
$\eta=(\eta^{(1)},\eta^{(2)},\eta^{(3)})^T$
with
$\eta^{(1)}\in \mathbb{R}^{d_1+d_2-1},\eta^{(2)}\in \mathbb{R}^{d_1+d_2-1}$ and $\eta^{(3)}\in \ell^{\infty}(\mathbb{R}^p)$ are nuisance functions.
The true nuisance parameter $\eta_{0,l}$ is given by
align*[align* omitted — 130 chars of source]
where $\beta_0^{(l)}$ is defined as
$$\beta_0^{(l)}:=(\theta_{0,1},\dots,\theta_{0,l-1},\theta_{0,l+1},\dots\theta_{0,d_1},\beta_{0,1},\dots,\beta_{0,d_2})^T.$$
Essentially, the index $l$ determines which coefficient is not contained in $\beta_0^{(l)}$. Formally, we will assume sparsity of $\beta_0^{(l)}$, i.e. $\|\beta_0^{(l)}\|_0\le s_1$ and we also assume sparsity in the auxiliary regression, i.e. $\|\gamma_0^{(l,1)}\|_0\le s_2$ for all $l=1,\ldots,d_1$. These assumptions are discussed in detail in Comment (ref) and Comment (ref).
The third part of the nuisance functions captures the error made by the approximation of $f_1$ and $f_{-1}$, which is independent from $l$. Therefore, we sometimes omit $l$.
remarkThe score $\psi$ is linear in $\theta$, meaning
\begin{align*}
\psi_l(W,\theta,\eta)=\psi^{a}_l(X,\eta^{(2)})\theta+\psi^b_l(X,\eta)
\end{align*}
with
$$\psi^{a}_l(X,\eta^{(2)})=-g_l(X_1)(g_l(X_1)-(\eta^{(2)})^TZ_{-l})$$
and
$$\psi^b_l(W,\eta)=(Y-(\eta^{(1)})^T Z_{-l}-\eta^{(3)}(X))(g_l(X_1)-(\eta^{(2)})^TZ_{-l})$$
for all $l=1,\dots,d_1$.
remarkThe score function $\psi$ satisfies the moment condition, namely
\begin{align*}
\mathbb{E}\left[\psi_l(W,\theta_{0,l},\eta_{0,l})\right]&=0
\end{align*}
for all $l=1,\dots,d_1$, and, given further conditions mentioned in Section (ref), the near Neyman orthogonality condition
\begin{align*}
D_{l,0}[\eta,\eta_{0,l}]&:=\partial_t\big\{\mathbb{E}[\psi_l(W,\theta_{0,l},\eta_{0,l}+t(\eta-\eta_{0,l}))]\big\}\big|_{t=0}\approx 0,
\end{align*}
where $\partial_t$ denotes the derivative with respect to $t$.
remarkMostly for ease of notation, we consider the nonparametric additively separable model $Y=f(X)+\varepsilon=f_1(X_1)+f_{-1}(X_{-1}) + \varepsilon$. In the statistical literature the most studied special case is the fully additive separable model
$Y=f(X)+\varepsilon=f_1(X_1)+\ldots + f_{p}(X_{p}) + \varepsilon$. This model has more structure that can be exploited, but what matters from a theoretical perspective is the number of parameters $d_2$, respectively the sparsity index $s_1$, for the nuisance parts $f_{-1}(X_{-1})$ or $f_2(X_2)+\ldots + f_{p}(X_{p})$, which can grow with the sample size in our asymptotic framework.
Estimation
In this section, we describe our estimation method and how the uniform valid confidence bands are constructed. First, nuisance functions are estimated using lasso regressions. Subsequently, they are plugged into the moment conditions and solved for the target parameters, which yield an estimate $\hat{f}_1$ for the first component in the additive regression model. The lower and upper curve of the confidence bands are then determined using the estimated covariance matrix and a critical value, which is obtained through a multiplier bootstrap procedure. The technical details for the estimation are given in this section.\\
Estimation and Algorithm
Let
$
g(x)=\left(g_1(x),\dots,g_{d_1}(x)\right)^T\in\mathbb{R}^{d_1\times 1},
$
and
align*[align* omitted — 145 chars of source]
for some vectors
$\theta=(\theta_1,\dots,\theta_{d_1})^T$
and $\eta=(\eta_1,\dots,\eta_{d_1})^T.$
For each $l=1,\dots,d_1$, let $\hat{\eta}_l=\left(\hat{\eta}_l^{(1)},\hat{\eta}_l^{(2)},\hat{\eta}_l^{(3)}\right)$ be an (lasso) estimator of the nuisance function. The estimator $\hat{\theta}_0$ of the target parameter
$\theta_0=(\theta_{0,1},\dots,\theta_{0,d_1})^T$
is defined as the solution of
align[align omitted — 283 chars of source]
where $\epsilon_{n}=O\left(\delta_n\varsigma_n^{-1/2}n^{-1/2}\right)$ is the numerical tolerance. The sequences of positive constants, $\delta_n$ and $\varsigma_n$, are defined in Assumption A.(ref) and Assumption A.(ref). Since the score is linear, the explicit solution is given by
$$\hat{\theta}_{0,l}=-\frac{\mathbb{E}_n[\psi^b_l(W,\hat{\eta}_l)]}{\mathbb{E}_n[\psi^{a}_l(W,\hat{\eta}^{(2)}_l)]}.$$
Therefore, (ref) obviously holds if we use the estimator $\hat{\theta}_{0,l}$. We consider the representation in (ref) because we rely on more general results derived in Appendix (ref) to analyze the theoretical properties of our new estimation procedure.
Lastly, the target function $f_1(\cdot)$ can be estimated by
align[align omitted — 57 chars of source]
Define the Jacobian matrix
align*[align* omitted — 193 chars of source]
with
align*[align* omitted — 103 chars of source]
for all $l=1,\dots,d_1$. Observe that $\Sigma_{\varepsilon\nu}:=\mathbb{E}\big[\psi(W,\theta_0,\eta_0)\psi(W,\theta_0,\eta_0)^T\big]$ is the covariance matrix of $\varepsilon\nu:=(\varepsilon\nu^{(1)},\dots,\varepsilon\nu^{(d_1)})$. Define the approximate covariance matrix
align*[align* omitted — 192 chars of source]
with
align*[align* omitted — 1,095 chars of source]
The approximate covariance matrix can be estimated by replacing every expectation with its empirical analog and plugging in the estimated parameters
align*[align* omitted — 209 chars of source]
This estimated covariance matrix can be used to construct the confidence bands
align*[align* omitted — 188 chars of source]
where $c_{\alpha}$ is a critical value determined by the following standard multiplier bootstrap method introduced in chernozhukov2013gaussian. Define
$$\hat{\psi}_x(\cdot):=(g(x)^T\hat{\Sigma}_n g(x))^{-1/2}g(x)^T\hat{J}^{-1}\psi(\cdot,\hat{\theta}_{0},\hat{\eta}_{0})$$
and let
align*[align* omitted — 177 chars of source]
where $(\xi^{(i)})_{i=1}^n$ are independent standard normal random variables (especially independent from the data $(W^{(i)})_{i=1}^n$). The multiplier bootstrap critical value $c_\alpha$ is given by the $(1-\alpha)$-quantile of the conditional distribution of $\sup_{x\in I}|\hat{\mathcal{G}}_x|$ given $(W^{(i)})_{i=1}^n$. This estimation procedure can be summarized in Algorithm (ref).
algorithm[algorithm omitted — 2,397 chars of source]
Practical Considerations: Estimation of Sparse High-Dimensional Additive Models
In addition to the significance level $\alpha$, the number of bootstrap repetitions and an interval $I$ for inference, it is necessary to specify a dictionary of approximating functions $g_1, \ldots, g_{d_1}$ for $f_1$ and $h_1, \ldots, h_{d_2}$ for $f_{-1}$ as inputs to Algorithm (ref). In our theoretical results and our simulation experiments, we focus on B-splines. From a practical point of view, B-splines are specified by parameters for a polynomial degree and the number of knots. In our simulation study, we use cubic B-splines. The number of knots plays a role for the quality of the approximation for $f_1$ as well as the width of the confidence bands. For example, a higher number of knots leads to a higher number of basis functions and corresponding number of fitted coefficients. In our simulation study, we observed that our results are relatively robust with regard to the number of knots and only change moderately with changes of the specification. We base the choice of the number of knots on preliminary evidence in terms of the approximation quality for $f_1$. We also investigated the use of a cross-validated choice of the number of knots, which would not be backed by our theoretical framework, but could serve as a practical guidance. Our results, which are available on request, indicate that a cross-validated choice of the B-spline specification tends to lead to conservative results, i.e., wide confidence intervals with higher coverage than the nominal level.
Based on the provided input, the algorithm proceeds with various nuisance estimation steps. Due to the large number of coefficients, we base estimation on the lasso or post-lasso, i.e., a lasso estimator that is followed by an ordinary least squares estimation step belloni:2013. The performance of regularized estimators depends on the choice of the corresponding penalty parameters. We base the choice of the lasso penalty on theoretical arguments as provided in Appendix (ref). In our software implementation, we use the rigorous lasso learner as provided by the R package hdm hdm. We also experimented with a cross-validated choice of the lasso penalty as provided by the glmnet package glmnet:2010. However, we find that the cross-validated penalty choice was more costly under computational consideration and also less stable. In our simulation experiments, we observe that the estimation performance of the lasso learner has important implications on the accuracy of our point estimator and, accordingly, on the coverage results of the confidence bands. High-dimensional additive models are characterized by a large number of parameters, which have to be estimated with a possibly small number of observations. In our simulation results in Section (ref), we find that the theory-based lasso learner exhibits a good performance in selecting the sparse parameters even in highly correlated settings.
Once, the lasso estimation has been performed, the corresponding residuals are plugged into the variance-covariance matrix. This, in turn, is used to construct the simultaneous confidence bands via a multiplier bootstrap procedure. The latter is based on a random perturbation of the score function, for example, by an i.i.d. standard normal distributed random variable. This procedure is very appealing from a computational point of view as it does not require resampling and reestimation of the parameters, as for example in classical bootstrap procedures. It is generally recommended to use a large number of bootstrap repetitions, $B\ge 500$.
Main Results
Now, we proceed to specifying the conditions required to create uniformly valid confidence bands using Algorithm (ref). To represent $f_1$ and $f_{-1}$ by their approximations in ((ref)) and ((ref)), we need to choose an appropriate set of approximating functions $g=(g_1,\dots,g_{d_1})$ and $h=(h_1,\dots,h_{d_2})$, respectively. In this context, let $\bar{d}_n:=\max(d_1,d_2,n,e)$ and $C$ be a strictly positive constant independent of $n$ and $l$, where $e$ in $\bar{d}_n$ denotes the Euler's number. $(\tilde{A}_n)_{n \geq 1}$ denotes a sequence of positive constants, possibly growing to infinity with $\tilde{A}_n \geq n$ for all $n$. For more details, we refer to Appendix (ref). The number of non-zero coefficients in Equation ((ref)) and ((ref)), is given by the sparsity index $s=\max(s_1,s_2)$. Additionally, we set $t_1:=\sup_{x\in I} \|g(x)\|_0\le d_1$. The definition of $t_1$ is helpful if the functions $g_l$, $l=1,\dots,d_1$, are local, meaning that for any point $x$ in $I$ there are at most $t_1<<d_1$ non-zero functions. Furthermore, $g(I)$ denotes the image of the approximation functions concerning the interval of interest $I$.\\ \\
The following assumptions are valid uniformly in $n\ge n_0$ and $P\in\mathcal{P}_n$:
assumptiona\ \\
\begin{itemize}
• It holds $\frac{1}{\sqrt{t_1}}\lesssim \inf\limits_{x\in I} \|g(x)\|_2\le C<\infty$, $\quad\sup_{x\in I} \sup_{l=1,\dots,d_1}|g_l(x)|\le C$ and for all $\varepsilon>0$
\begin{align*}
\log N(\varepsilon,g(I),\|\cdot\|_{2})\le Ct_1\log\left(\frac{\tilde{A}_n}{\varepsilon}\right)
\end{align*}
where $N (\epsilon, g(I), || \cdot ||_{2})$ denotes the covering number of $g(I)$ with radius $\epsilon$ with respect to the $|| \cdot ||_{2}$-norm.
•
The approximation errors obey
\begin{align*}
\left(b_1(X_1)+b_2(X_{-1})\right)^2\le Cs_1\log(\bar{d}_n)/n\ (a.s.).
\end{align*}
• It holds $$\mathbb{E}\Big[\nu^{(l)}\big(b_1(X_1)+b_2(X_{-1})\big)\Big]\le C\delta_n \varsigma_n^{-1/2}n^{-1/2}$$
with $\varsigma_n$ defined in Assumption A.(ref) (iv) and $\delta_n=\sqrt{\frac{n^{2/q}s^2\varsigma_n\log^2(\bar{d}_n)}{n}}=o(1)$ where $q$ is defined in Assumption A.(ref) (iii) .
\end{itemize}
Assumption A.(ref)$(i)$ contains regularity conditions on $g$. We assume that the supremem of the $\ell_2$-norm of $g(x)$ is bounded, but the infimum is allowed to decrease with sample size, affecting the growth conditions in A.(ref)$(v)$. Assumption A.(ref)$(ii)$ is a condition on the approximation error and is discussed further in Comment (ref). This assumption is mild because the number of approximating functions may increase with sample size. Lastly, Assumption A.(ref)$(iii)$ ensures that any violation of the exact Neyman Orthogonality due to approximation errors is negligible. It is worth noting that if $b_1(X_{1})$ and $b_2(X_{-1})$ are measurable with respect to $Z_{-l}$ (for example, in a linear approximate sparse setting for the conditional expectation) the exact Neyman Orthogonality holds.
assumptiona\ \\
\begin{itemize}
• For all $l=1,\dots,d_1$, $\Theta_l$ contains a ball of radius $$\log(\log(n))n^{-1/2}\log^{1/2}(d_1\vee e)\varsigma_n\log(n)$$ centered at $\theta_{0,l}$ with
$$\sup\limits_{l=1,\dots,d_1}\sup\limits_{\theta_l\in\Theta_l} |\theta_l|\le C.$$
• It holds
\begin{align*}
\|\beta_0^{(l)}\|_0\le s_1,\quad \|\beta_0^{(l)}\|_2\le C
\end{align*}
for all $l=1,\dots,d_1$ and
\begin{align*}
\max\limits_{l=1,\dots,d_1}\|\gamma_0^{(l)}\|_1\le C
\end{align*}
with
\begin{align*}
\max\limits_{l=1,\dots,d_1}\|\gamma_0^{(l,1)}\|_0\le s_2
\end{align*}
and
\begin{align*}
\max\limits_{l=1,\dots,d_1}\|\gamma_0^{(l,2)}\|_{1}^2&\le C\sqrt{\varsigma_n^{-1/2}s_2^2\log(\bar{d}_n)/n},\\ \max\limits_{l=1,\dots,d_1}\|(\gamma_0^{(l,2)})^TZ_{-l}\|_{P,2}^2&\le C\varsigma_n^{-1}s_2\log(\bar{d}_n)/n.
\end{align*}
• The constructed approximation functions are almost surely bounded,
$$\|Z\|_{\infty}\le C\quad \text{(a.s.)},$$
and
$$\|\varepsilon\|_{P,q}\le C$$
for $q\ge 4$ corresponding to the growth rates in $(v)$.
• It holds
$$c \varsigma_n^{-1} \le \inf\limits_{\|\xi\|_2=1} \mathbb{E}[(\xi^T Z)^2] \le \sup\limits_{\|\xi\|_2=1} \mathbb{E}[(\xi^T Z)^2] \le C \varsigma_n^{-1}$$
and
$$ c \varsigma_n^{-1}\le \inf\limits_{\|\xi\|_2=1} \xi^T\Sigma_\nu\xi\le \sup\limits_{\|\xi\|_2=1} \xi^T\Sigma_\nu\xi \le C \varsigma_n^{-1}.$$
Further,
$$c\varsigma_n^{-2}\le\mathbb{E}\left[(\nu^{(l)})^2 Z_{-l,j}^2\right]\le C\varsigma_n^{-2},$$
uniformly for all $l=1,\dots,d_1$, $j = 1,\dots, (d_1 -1)$ and
\begin{align*}
\max\limits_{l=1,\dots,d_1}\max\limits_{j=1,\dots,d_1 - 1}\|\nu^{(l)} Z_{-l,j}\|_{P,3}&\le C\varsigma_n^{-1}K_{n},\\
\max\limits_{l=1,\dots,d_1}\max\limits_{j=1,\dots,d_1 -1} \mathbb{E}\left[(\nu^{(l)})^4 Z_{-l,j}^4\right]&\le C\varsigma_n^{-4} L_n.
\end{align*}
• There exists a positive number $q\ge 0$ such that the following growth condition is fulfilled:
\begin{itemize}
• $n^{\frac{2}{q}}t_1^3\varsigma_n\log(\tilde{A}_n\varsigma_n^{-1/2})\left(\frac{s^2\log^2(\bar{d}_n) }{n}\vee \frac{t_1^9\varsigma_n^2\log^2(\tilde{A}_n\varsigma_n^{-1/2})}{n}\right)=o(1)$
• $t_1^7\varsigma_n\log(\tilde{A}_n\varsigma_n^{-1/2})\left(\frac{s\log(\bar{d}_n)\log(\tilde{A}_n)}{n}\vee\frac{t_1^{9}\varsigma_n^2\log^6(\tilde{A}_n\varsigma_n^{-1/2})}{n}\right)=o(1)$
\end{itemize}
\end{itemize}
Assumptions A.(ref)$(i)$ and $(ii)$ are regularity and sparsity conditions, permitting the number of nonzero regression coefficients $s_1$ and $s_2$ to grow to infinity with increasing sample size. A detailed comment on the sparsity condition is given in Comment (ref). Assumption A.(ref)$(iii)$ contains tail conditions on the approximating functions
(and therefore on the original variables) as well as for the error term. Assumption A.(ref)$(iv)$ is an eigenvalue condition that restricts the correlation between the basis elements (and therefore between the original variables), as well as the correlation between the error term $\nu_l$ from the auxiliary regressions. The key innovation of A.(ref)$(iv)$ is that we allow the eigenvalues to vanish as $n$ increases. This is essential if we use an increasing number of local approximation functions $g_l$, $l=1,\dots,d_1$, which implies that the support and hence the variance of each $g_l$ is vanishing with increasing sample size. Typically, we have $\varsigma_n\approx d_1$ as we will illustrate in Comment (ref) for B-Splines. This is in line with the Assumption 5 in kozbur2015 which tackles the same problem of local estimators where $\zeta_0(d_1)=d_1^{1/2}$ is the rate to orthogonalize B-Splines. A related assumption on the restricted eigenvalue can also be found in Assumption A3 in lu2019 where the minimum eigenvalue is of order $O(d_1^{-1})$ as well.
Lastly, Assumption A.(ref)$(v)$ provides the growth conditions. These are given in general terms and depend on the choice of the approximation functions. Choosing B-Splines simplifies the growth conditions significantly as we will also outline in Comment (ref).\\
theoremUnder Assumptions A.(ref) and A.(ref), it holds that
\begin{align*}
P\left(\hat{l}(x)\le f_1(x)\le \hat{u}(x), \forall x\in I\right)\to 1-\alpha
\end{align*}
uniformly over $P\in\mathcal{P}_n$ where $c_\alpha$ is a critical value determined through the multiplier bootstrap method.
Theorem (ref) provides uniform confidence bands for the whole target component $f_1$. Particularly, with probability $1-\alpha$, we have
$$\sup_x|\hat{f}_1(x)-f_1(x)|^2=O\left(\frac{sup_x (g(x)^T\hat{\Sigma}_ng(x))c_\alpha^2}{n}\right)=O\left(\frac{\varsigma_n c_\alpha^2}{n}\right)$$
by Assumption A.(ref)$(iv)$. Hence, the rate of convergence/the width of the confidence band strongly depends on the parameter $\varsigma_n$ and therefore on the number of approximating functions $d_1$ where typically $\varsigma_n\approx d_1$ for local approximating functions.
remark[B-Splines]
An appropriate and common choice in series estimation is B-Splines. B-Splines are positive and local in the sense that $g(x)\ge 0$ and $\sup_{x\in I}\|g(x)\|_0\le t_1$ for every $x$, where $t_1$ is the order of the spline. The $l_1$-norm of B-Splines is bounded by 1, meaning
\begin{align*}
\|g(x)\|_1=\sum\limits_{j=1}^{d_1}g_j(x)\le 1
\end{align*}
for every $x$ (due to partition of unity). Hence, Assumption A.(ref)$(i)$ is met with
$$\frac{1}{\sqrt{t_1}}\le\inf\limits_{x\in I} \|g(x)\|_2\le\sup\limits_{x\in I} \|g(x)\|_2\le 1 \quad\text{and}\quad \sup\limits_{x\in I} \sup\limits_{l=1,\dots,d_1}|g_l(x)|\le 1.$$
Now, we go into greater detail regarding the condition on the covering number of the image of $g$. Especially if $t_1<d_1$, the complexity of the approximating functions is reduced significantly. One obtains
\begin{align*}
g(I)\subseteq \bigcup\limits_{j=1}^{\binom{d_1}{t_1}} g^{(j)}(I),
\end{align*}
where each $g^{(j)}(I)$ is only dependent on $t_1$ nonzero components. It is straightforward to see that for each $g^{(j)}(I)$ the covering numbers satisfy
\begin{align*}
N(\varepsilon,g^{(j)}(I),\|\cdot\|_{2})\le \left(\frac{6\sup_{x\in I}\|g(x)\|_2}{\varepsilon}\right)^{t_1}
\end{align*}
(cf. vanweak), implying
\begin{align*}
\log N(\varepsilon,g(I),\|\cdot\|_{2})&\le \log\left(\sum\limits_{j=1}^{\binom{d_1}{t_1}}N(\varepsilon,g^{(j)}(I),\|\cdot\|_{2})\right)\\
&\le \log\left(\left(\frac{e\cdot d_1}{t_1}\right)^{t_1}\left(\frac{6\sup_{x\in I}\|g(x)\|_2}{\varepsilon}\right)^{t_1}\right)\\
&\le t_1\log\left(\left(\frac{6ed_1\sup_{x\in I}\|g(x)\|_2}{t_1}\right)\frac{1}{\varepsilon}\right)\\
&\lesssim t_1\log\left(\frac{d_1}{\varepsilon}\right).
\end{align*}
It is worth noting that the second moment of a univariate B-spline (with a fixed order) is given by
\begin{align*}
\mathbb{E}[g(X)^2]\approx d_1^{-1},
\end{align*}
see e.g. shen1998local or de1978practical. In this case it is reasonable to assume that Assumption
A.(ref)$(iv)$ holds with $\varsigma_n=d_1$. Let us now discuss Assumption A.(ref)$(ii)$. Define the following class of additive, smooth functions:
$$\mathcal{F}_p^m(\tilde{s}):=\{f:X\mapsto\sum_{j\in \tilde{S}}f_j(X_j)\mid|\tilde{S}|\le \tilde{s}, f_j\in C^m[0,1], \mathbb{E}[f_j(X_j)]=0\ \forall j \in \tilde{S}\}$$
and assume $f_{-1}\in \mathcal{F}_p^m(\tilde{s})$. Using B-Splines we are able to approximate $f_{-1}$ sufficiently well. Let $h_1,\dots,h_{\tilde{d}}$ be B-Splines of order $\tilde{t}$ with equidistant knots and $\mathcal{B}(\tilde{d},\tilde{t})$ be the corresponding linear space. Relying on shen1998local, for each $j\in\tilde{S}$, we can find a $f_j^*\in \mathcal{B}(\tilde{d},\tilde{t})$ such that
\begin{align*}
\|f_j-f_j^*\|_\infty = O\left((\tilde{d}-\tilde{t})^{-\tilde{t}}\right)
\end{align*}
for all $\tilde{t}\le m$. This directly implies
\begin{align*}
\sup_{x_{-1}} b_2(x_{-1}) \le C \tilde{s} (\tilde{d}-\tilde{t})^{-\tilde{t}}.
\end{align*}
With the same argument for $f_1\in C^m[0,1]$ and choosing $\tilde{d}=d_1$, we obtain
\begin{align*}
\sup_{x} (b_1(x_1) + b_2(x_{-1})) \le C (\tilde{s}+1) (\tilde{d}-\tilde{t})^{-\tilde{t}}.
\end{align*}
For $\tilde{t} = 2$ and $\tilde{s}= O(1)$, we have to ensure that
\begin{align*}
(sup_{x} (b_1(x_1) + b_2(x_{-1})))^2 \le C \tilde{d}^{-4} \le \frac{s_1\log(\bar{d}_n)}{n}.
\end{align*}
Here, for simplicity, we assume that only a finite number of regressors in our model are relevant ($\tilde{s}=O(1)$) and hence $s_1\approx d_1$ and $\bar{d}_n\approx pd_1$. This implies
\begin{align*}
s_1 = O(n^{1/5}\log^{-1/5}(\bar{d}_n)).
\end{align*}
which can be aligned with our growth rates in Assumption A.(ref)$(v)$ which we discuss next. It is worth noting that in this case
$$ sup_{x} (b_1(x_1) + b_2(x_{-1}))\lesssim d_1^{-2}$$
which is in line with Assumption 3 in kozbur2015 with $K=d_1$ and $\alpha_0=2$.
Choosing the order of B-Splines $t_1 = \log(n)$ and keeping in mind that $\varsigma_n\approx d_1$ and $\tilde{A}_n\approx d_1$, the growth rates in Assumption A.(ref)$(v)$ simplify to
\begin{itemize}
• $n^{\frac{1}{\tilde{q}}}d_1\log(d_1)\left(\frac{s^2\log^2(\bar{d}_n) }{n}\vee \frac{d_1^2\log^2(d_1)}{n}\right)\lesssim n^{\frac{1}{\tilde{q}}}\frac{d_1^3\log^3(\bar{d}_n)}{n} =o(1)$
• $d_1\log(d_1)\left(\frac{d_1\log(\bar{d}_n)\log(d_1)}{n}\vee\frac{d_1^2\log^6(d_1)}{n}\right)\lesssim \frac{d_1^3\log^7(d_1)}{n}=o(1)$
\end{itemize}
for a suitable $\tilde{q}$. This is in line with $s_1\approx d_1\gtrsim n^{1/5}$ which is needed to ensure a sufficient small approximation error in Assumption A.(ref)$(ii)$. We observe that the total number of regressors $p$ and hence the total number of approximation functions $\bar{d}_n\approx d_1p$ can grow at an exponential rate with sample size. This means that the set of approximating functions can be larger than the sample size. This situation is common for lasso-based estimators. Our growth condition is in line with other results in the
literature, e.g., belloni2018uniformly, Belloni2014b and many others.
Obviously, this is only one special case in which the assumptions A.(ref)$(v)$ and A.(ref)$(ii)$ are
satisfied. It is worth noting that the Assumption A.(ref)$(v)$ can still hold true if we allow the number of relevant regressors to increase with sample size.
This discussion points precisely to the following trade-off:
For a given model, such as an additive model with a finite number of relevant regressors, the number of approximating functions $d_1$ and $d_2$ are allowed to grow with sample size to ensure a sufficiently small approximation error in Assumption A.(ref)$(ii)$. On the other hand, more parameters in the linear model increases the estimation error in the nuisance parameters, which is controlled by the growth assumptions in Assumption A.(ref)$(v)$. Therefore, both assumptions need to be balanced. This trade-off can also be observed in Assumption 11.2 and 11.3 in kozbur2015. In contrast to these growth rates, our weaker growth conditions are simultaneously feasible if the target function is twice-differentiable as outlines above. This can be explained by the fact that we are using improvements in series estimators in belloni2018uniformly to relax the assumptions in kozbur2015 who uses results in NEWEY1997147. Particularly, kozbur2015 assumes $n^{-1}d_1^4=o(1)$ in Assumption 4, while we roughly assume $n^{-1}d_1^3=o(1)$ above. Hence, choosing $d_1\approx n^{1/5}$, $\bar{d}_n=d_2$ can be as large as $\exp(o(n^{2/15}))$ which can be larger than $n$ to satisfy the growth condition in a). It is worth noting that the growth condition on $d_2$ in Assumption 11.2 in Kozbur2021 is much more restrictive than in our paper.
The second growth condition in b) guarantees the validity of multiplier bootstrap and allows us to construct uniformly valid confidence regions. It is in line with chernozhukov2013gaussian, but we allow vanishing eigenvalues.
remarkThe sparsity condition in A.(ref)$(ii)$ restricts the number of nonzero regression coefficients $s_1$ and $s_2$ in the Equations ((ref)), ((ref)) and ((ref)). Through this, we especially assume that the regression function $f$ can be adequately approximated using only $s_1$ relevant basis functions. Note that we do not directly control the number of relevant covariables, but rather the total number of approximating functions. This sparsity condition differs from that used by gregory2016 and lu2019 who restrict the number of relevant additive components in the model ((ref)). Our model also includes the approximate sparse setting through the error terms $b_1$ and $b_2$ in ((ref)) and ((ref)), which offers greater flexibility and is more realistic for many applications.\\
The sparsity condition in A.(ref)$(ii)$ which restricts the number of nonzero regression coefficients $s_1$ in the Equations ((ref)) and ((ref)) is standard in high-dimensional regression models. The condition on the number of nonzero regression coefficients $s_2$ in the auxiliary regression is also standard but (ref) needs further discussion.
It assumes that each approximating function itself may be approximated sufficiently
well by $s_2$ or fewer additive components which is comparable to Assumption (B4) in gregory2016 or Assumption 6 in Kozbur2021.
Lemma 1 in gregory2016 shows that (B4) holds under a mixing type condition on the covariates if one uses piecewise polynomials for approximation. Analog to the sparsity assumption in the main regression, sparsity in (ref) ensures fast estimation rates of the nuisance parameters. The auxiliary regression in (ref) is basically used to orthogonalize our estimation procedure. Heuristically, for any M-estimator, the orthogonality property essentially requires the inverse of the Hessian matrix of the population loss function to have sparse columns (cf. Assumption 6 in lu2019). Therefore such a sparsity assumption or related assumption is unavoidable.
Simulation Results
table[table omitted — 601 chars of source]
We assess the empirical performance of our estimator through a series of simulation experiments based on the settings outlined in gregory2016, which in turn build on the work of meier2009. In our first set of simulations, we evaluate the empirical coverage of our simultaneous confidence bands replicating the data generating process from gregory2016 and meier2009, which incorporates a homoskedastic error term. Additionally, because our estimation framework is also compatible with heteroskedastic error terms, we extend the simulation study to include a heteroskedastic version of the aforementioned design.
Furthermore, we investigate the performance of our estimator and confidence bands in scenarios involving interactions among the nuisance functions $f_{-j}(x_{-j})$. These scenarios are challenging in two regards: First, modeling interactions requires a large number of interaction terms, quickly leading to very high-dimensional settings. Second, specifying interactions among various functions in empirical applications is challenging because of the risk that the chosen model will not precisely match the true underlying data generating process. By exploiting a sparse structure, our theoretical framework can tolerate a moderate degree of misspecification in the nuisance functions. We introduce two interacted designs -- one correctly specified and one misspecified--providing a comprehensive assessment of our methodology.
We consider the finite sample performance of our estimator in a high-dimensional additive model of the form
align*[align* omitted — 65 chars of source]
with $i = 1, \ldots, n$ and $j = 1, \ldots, p$. The functions $f_j(x_j)$ are nonlinear for $j=1, 2, 4$ and linear for $j=3$, as can be recognized from their definition in Table (ref). The functions $f_5(x_5), \ldots, f_p(x_p)$ are null components, thus implementing a sparse setting. In all of our considered settings, the regressors $X$ are marginally uniformly distributed on an interval, $I = [-2.5, 2.5]$ with correlation matrix $\Sigma$ with $\Sigma_{k,l}=\rho^{|k-l|}$, $1 \le k,l \le p$. We consider the values $\rho = 0$ and $\rho = 0.5$ because they implement the weakest and strongest correlation structures from the simulations in gregory2016. We generate data sets for scenarios with $n \in \{100, 1000\}$ observations and $p \in \{50, 150\}$ explanatory variables $X_1, \ldots, X_p$.
First, we consider a homoskedastic setting with a standard normally distributed error term $\epsilon \sim N\left(0, 1\right)$, which is identical to the design in gregory2016. Second, we define a comparable heteroskedastic setting by specifying an error term $\epsilon_j\sim N\left(0, \sigma_j(x_j)\right)$ with $\sigma_j(x_j) = \underline{\sigma} \cdot (1 + |x_j|)$ and $\underline{\sigma} = \sqrt{\frac{12}{67}}$. This value of $\underline{\sigma}$ ensures a signal-to-noise ratio that is comparable to the homoskedastic setting. Third, we implement two different interacted scenarios, in which we focus on the estimation of the function $f_1(x_1)$. In the correctly specified interacted scenario, which we will later refer to as Scenario $I$, the specification of the estimated sparse additive model is identical to that of the underlying data generating process. In this model, the function $f_2(x_2)$ is interacted with all other components except for $f_1(x_1)$, making the scenario suitable for our estimation approach. Including the interactions of all $p-1$ nuisance functions leads to a high-dimensional setting. For example, if one specifies cubic B-splines with 3 knots in a setting with $p=150$ covariates, the overall dimension of the estimated model is $d_1+d_2 = 6,228$. In the data generating process of the misspecified interacted Scenario $II$, all non-sparse functions are interacted with each other, except for $f_1(x_1)$. In contrast, the model that is estimated empirically only specifies a basic additive model without any interaction terms. Hence, the interacted Scenario $II$ implements a situation with considerable model misspecification.
In the simulation, we use the estimator and the multiplier bootstrap procedure we proposed in Section (ref) to generate predictions $\hat{f}_j(x_j)$ for the function $f_j(x_j)$ and construct simultaneous confidence bands that are defined in terms of $\hat{l}_j(x_j)$ and $\hat{u}_j(x_j)$. We approximate the functions $f_j(x_j)$ in the additive model by cubic B-splines. Estimation is performed using post-lasso with a theory-based choice of the penalty level as implemented in the R package hdm hdm. Further details related to the implementation in the simulation study can be found in Appendix (ref).\\
Table (ref) presents the empirical coverage achieved by the estimated simultaneous $95\%$-confidence bands in the homoskedastic and heteroskedastic scenario. The results are obtained in $R = 500$ repetitions and the confidence bands are constructed over an interval of values of $x_j$, $I = [-2,2]$. A confidence band is considered to cover the function $f_j(x_j)$ if it entirely contains the true function, or, stated more formally, if for all values of $x_j \in I$ it holds that $\hat{l}_j(x_j)\le f_j(x_j) \le \hat{u}_j(x_j)$.\\
The results can be interpreted as evidence supporting the validity of our inference method in high-dimensional additive models in both homoskedastic and heteroskedastic settings. The coverage of the confidence bands is close to the nominal level across all settings, in particular when the sample size is large ($n=1000$) or if the correlation structure among the covariates is weak ($\rho=0$). We find that a precise approximation of the target function $f_j(x_j)$ through an appropriate specification of the employed B-Splines is essential for the coverage results of the confidence bands.
table[table omitted — 1,872 chars of source]
Figures (ref) to (ref) illustrate the average results for the estimators of $f_1(x_1), \ldots, f_5(x_5)$, along with the corresponding joint confidence bands. The green dashed line and the green shaded areas represent the results for the settings with $\rho = 0$, and the blue curve and shaded areas refer to the setting with $\rho = 0.5$. The red solid line shows the true value of the function $f_j(x_j)$. The corresponding results for the heteroskedastic settings are very similar and, hence, made available in Appendix (ref). The plots depicting the average results provide additional insights into the quality of estimation across different settings. First, they illustrate that the estimation benefits substantially from larger samples. In samples with $n=100$ the confidence bands exhibit noticeable variability and width, whereas in larger samples they become considerably narrower. It is important to note that this observation is expected. For example, in the high-dimensional settings with $n=100$ and $p=150$, cubic B-splines with 3 knots would lead to an overall dimension of $d_1 + d_2 = 900$, which is challenging in terms of estimation. Second, comparing the results across different values of $\rho$ helps elucidate the challenge in terms of approximation quality. For example, Figure (ref) shows the results for $f_4(x_4)$ in the homoskedastic design. Whereas the approximation appears to be very accurate for $\rho = 0$ in the setting with $n=100$, it is considerably worse for $\rho=0.5$. However, in larger samples, the quality of the approximation improves, which is in line with the results of gregory2016. Similar observations can be made for $f_3(x_3)$ and $f_5(x_5)$. As in the previously considered settings, the strength of the correlation of the regressors $X$ plays an important role among the performance of the estimator $\hat{f}_1(x_1)$.
table[table omitted — 730 chars of source]
figure[figure omitted — 1,019 chars of source]
table[table omitted — 728 chars of source]
figure[figure omitted — 1,019 chars of source]
figure[figure omitted — 1,019 chars of source]
figure[figure omitted — 1,019 chars of source]
figure[figure omitted — 1,019 chars of source]
In addition to the homoskedastic and heteroskedastic designs, we investigate the performance of our estimator and confidence bands in interacted simulation settings. The results in Tables (ref) and (ref) and Figures (ref) and (ref) provide simulation evidence in very high-dimensional settings -- both with a correctly (Scenario $I$) and an incorrectly (Scenario $II$) specified additive model. As a consequence of the higher dimensionality and the increased noise in the estimation problem, the confidence intervals are wider than in the basic settings considered previously. In all cases and irrespective of the model misspecification in Scenario $II$, the empirical coverage is close to the nominal level. As in the settings considered before, the strength of the correlation among the regressors $X$ makes estimation of $f_1(x_1)$ more challenging (see for example the top panels in Figure (ref)).
In general, the performance of the estimator and the confidence bands depends on the specification of the cubic B-splines used to approximate the target functions (in terms of the number of knots). In our simulations, we observe that the quality of estimation and width of the confidence bands change only moderately when varying the number of knots. Overall, we conclude that our results are relatively robust to the specifications of the B-splines. We find that the performance of the point estimator for the target component $f_j(x_j)$ relies on variable selection through the lasso estimator. In line with our expectations, the lasso estimator is better at selecting the sparse components of the B-splines in larger samples and settings with weaker correlation structure, i.e., $\rho = 0$. In our simulations, we chose the specification of the B-splines estimator based on preliminary evidence. We also investigated the performance of a cross-validated choice (results available upon request). While the latter lacks theoretical justification, we found it to be relatively conservative, resulting in wide confidence bands. During our simulation experiments, we also investigated the performance of an alternative lasso learner that is based on a cross-validated choice of the penalty term. We found that the cross-validated penalty choice was computationally more expensive and inferior in terms of the estimation performance, as indicated by selecting very few of the sparse components.
Overall, we interpret the simulation results as supportive evidence for our estimation approach. The approximation and coverage results for our estimator and confidence bands across different simulation settings are comparable to those provided in gregory2016. As in gregory2016, the approximation quality of our estimator depends on the underlying correlation structure of the covariates $X$, the parametrization of the B-splines, and the variable selection performance of lasso. In their simulation study, gregory2016 report pointwise coverage results, whereas our results refer to the coverage of the true function $f_j(x_j)$ over all points of $x_j$ in an interval $I$ by a simultaneous confidence band.
figure[figure omitted — 1,121 chars of source]
figure[figure omitted — 1,085 chars of source]
Proofs
proof[Proof of Theorem (ref)]\ \\
We will prove that the Assumptions A.(ref) and A.(ref) imply the Assumptions B.(ref)-B.(ref) stated in Appendix (ref) and then the claim follows by applying Theorem (ref). Without loss of generality, we assume $\min(d_1,n)\ge e$ to simplify notation. \\ \\
Assumption B.(ref)\\
Conditions $(i)$ is directly assumed in A.(ref)$(i)$.
Since the eigenvalues of $\Sigma_{\nu}$ are of order $O(\varsigma_n^{-1})$ due to A.(ref)$(iv)$, it holds
\begin{align*}
c \varsigma_n^{-1}\le | J_{0,l}| \le C \varsigma_n^{-1}
\end{align*}
uniformly over $l=1,\dots,d_1.$
Further, since $c\le \operatorname{Var}(\varepsilon|X)\le C$ a.s., the eigenvalues of $\Sigma_{\varepsilon\nu}$ are also of order $O(\varsigma_n^{-1})$. Hence,
\begin{align*}
\Sigma_n=J_0^{-1}\Sigma_{\varepsilon\nu}(J_0^{-1})^T
\end{align*}
directly implies B.(ref)$(ii)$.\\
Assumption B.(ref)\\
For each $l=1,\dots,d_1$, the moment condition holds
\begin{align*}
\mathbb{E}\left[\psi_l(W,\theta_{0,l},\eta_{0,l})\right]&=\mathbb{E}\big[\varepsilon\nu^{(l)}\big]\\
&=\mathbb{E}\big[\nu^{(l)}\underbrace{\mathbb{E}\left[\varepsilon|X\right]}_{=0}\big]\\
&=0.
\end{align*}
For all $l=1,\dots,d_1$, define the convex set
\begin{align*}
T_l:=\Big\{\eta=(\eta^{(1)},\eta^{(2)},\eta^{(3)})^T:&\eta^{(1)},\eta^{(2)} \in \mathbb{R}^{d_1+d_2-1},\\
&\eta^{(3)}\in\ell^{\infty}(\mathbb{R}^p)\Big\}
\end{align*}
and endow $T_l$ with the norm
\begin{align*}
\|\eta\|_e:=\max\left\{\varsigma_n^{-1/2}\|\eta^{(1)}\|_2,\|\eta^{(2)}\|_2,\|\eta^{(3)}(X)\|_{P,2})^TZ_{-l}\|_{P,2}\right\}.
\end{align*}
Further, let $\tau_n:=\sqrt{\frac{s\log(\bar{d}_n)}{n}}$ with $s=\max(s_1,s_2)$ and define the corresponding nuisance realization set
\begin{align*}
\mathcal{T}_l:=\bigg\{\eta\in T_l:&\eta^{(3)}\equiv 0, \|\eta^{(1)}\|_0\vee\|\eta^{(2)}\|_0\le Cs, \\
&\varsigma_n^{-1/2}\|\eta^{(1)}-\beta_0^{(l)}\|_2\vee\|\eta^{(2)}-\gamma_0^{(l,1)}\|_2\le C\tau_n,\\
&\varsigma_n^{-1/2}\|\eta^{(1)}-\beta_0^{(l)}\|_1\vee\|\eta^{(2)}-\gamma_0^{(l,1)}\|_1\le C\sqrt{s}\tau_n\bigg\}\cup \{\eta_{0,l}\}
\end{align*}
for a sufficiently large constant $C$.
Due to A.(ref)$(ii)$ and $(iii)$, it holds
\begin{align*}
|\nu^{(l)}| &= \left|g_l(X_1) - (\gamma_0^{(l)})^T Z_{-l}\right|\\
&\le |g_l(X_1)| + \|\gamma_0^{(l)}\|_1 \|Z_{-l}\|_{\infty}\\
&\le C
\end{align*}
almost surely.
For $\mathcal{F}:=\{\varepsilon\nu^{(l)}:l=1,\dots,d_1\}$, it holds
\begin{align*}
\mathcal{S}_n:&=\mathbb{E}\left[\sup\limits_{l=1,\dots,d_1}\left|\sqrt{n}\mathbb{E}_n\left[\psi_l(W,\theta_{0,l},\eta_{0,l})\right]\right|\right]\\
&=\mathbb{E}\left[\sup\limits_{f\in\mathcal{F}}\mathbb{G}_n(f)\right]
\end{align*}
and the envelope $\sup_{f\in\mathcal{F}}|f|$ satisfies
\begin{align*}
\|\max_{l=1,\dots,d_1}\varepsilon\nu^{(l)}\|_{P,q}&\le C\|\varepsilon\|_{P,q}\le C
\end{align*}
for $q$ defined in Assumption A.(ref) $(iii)$.
We can apply Lemma P.2 from belloni2018uniformly with $|\mathcal{F}|=d_1$ to obtain
\begin{align*}
\mathcal{S}_n\le C\log^{\frac{1}{2}}(d_1)+ C\log^{\frac{1}{2}}(d_1)\left(n^{\frac{2}{q}}\frac{\log^(d_1)}{n}\right)^{1/2}\lesssim \log^{\frac{1}{2}}(d_1),
\end{align*}
due to A.(ref)$(v)$. Finally, Assumption A.(ref)$(i)$ implies B.(ref)$(i)$. Assumption B.(ref)$(ii)$ holds since for all $l=1,\dots,d_1$, the map $(\theta_l,\eta_l)\mapsto\psi_l(X,\theta_l,\eta_l)$ is twice continuously Gateaux-differentiable on $\Theta_l\times \mathcal{T}_l$, which directly implies the differentiability of the map $(\theta_l,\eta_l)\mapsto\mathbb{E}[\psi_l(X,\theta_l,\eta_l)]$. Additionally, for every $\eta \in\mathcal{T}_l\setminus\{\eta_{0,l}\}$, we have
{\allowdisplaybreaks
\begin{align*}
D_{l,0}[\eta,\eta_{0,l}]&:=\partial_t\big\{\mathbb{E}[\psi_l(W,\theta_{0,l},\eta_{0,l}+t(\eta-\eta_{0,l}))]\big\}\big|_{t=0}\\
&=\mathbb{E}\big[\partial_t\big\{\psi_l(W,\theta_{0,l},\eta_{0,l}+t(\eta-\eta_{0,l}))\big\}\big]\big|_{t=0}\\
&=\mathbb{E}\bigg[\partial_t\bigg\{\Big(Y-\theta_{0,l} g_l(X_1)-\big(\eta^{(1)}_{0,l}+t(\eta^{(1)}-\eta^{(1)}_{0,l})\big)^T Z_{-l}\\
&\quad -\big(\eta^{(3)}_{0,l}(X)+t(\eta^{(3)}(X)-\eta^{(3)}_{0,l}(X))\big)\Big)\\
&\quad\Big(g_l(X_1)-\big(\eta^{(2)}_{0,l}+t(\eta^{(2)}-\eta^{(2)}_{0,l})\big)^T Z_{-l}\Big)\bigg\}\bigg]\bigg|_{t=0}\\
&=\mathbb{E}\left[\varepsilon(\eta^{(2)}_{0,l}-\eta^{(2)})^T Z_{-l}\right]+\mathbb{E}\left[\nu^{(l)}(\eta^{(1)}_{0,l}-\eta^{(1)})^T Z_{-l}\right]\\
&\quad +\mathbb{E}\left[\nu^{(l)}\left(\eta^{(3)}_{0,l}(X)-\eta^{(3)}(X)\right)\right]
\end{align*}
}
with
\begin{align*}
\mathbb{E}\left[\varepsilon(\eta^{(2)}_{0,l}-\eta^{(2)})^T Z_{-l}\right]=\mathbb{E}\left[(\eta^{(2)}_{0,l}-\eta^{(2)})^T Z_{-l}\mathbb{E}[\varepsilon|X]\right]=0,
\end{align*}
\begin{align*}
\mathbb{E}\left[\nu^{(l)}(\eta^{(1)}_{0,l}-\eta^{(1)})^T Z_{-l}\right]&=(\eta^{(1)}_{0,l}-\eta^{(1)})^T\mathbb{E}\left[ Z_{-l}\nu^{(l)}\right]=0,
\end{align*}
and
\begin{align*}
\mathbb{E}\left[\nu^{(l)}\left(\eta^{(3)}_{0,l}(X)-\eta^{(3)}(X)\right)\right]=\ &\mathbb{E}\left[\nu^{(l)}\big(b_1(X_1)+b_2(X_{-1})\big)\right]\le C\delta_n\varsigma_n^{-1/2} n^{-1/2}
\end{align*}
due to Assumption A.(ref)$(iii)$ with $\delta_n=\sqrt{\frac{n^{2/q}s^2\varsigma_n\log^2(\bar{d}_n)}{n}}$. Due to the linearity of the score and the moment condition, it holds
\begin{align*}
\mathbb{E}[\psi_l(W,\theta_l,\eta_{0,l})]=J_{0,l}(\theta_l-\theta_{0,l})
\end{align*}
and by Assumption A.(ref)$(iv)$
$$c\varsigma_n^{-1}\le |J_{0,l}|=\mathbb{E}\left[(\nu^{(l)})^2\right] \le C\varsigma_n^{-1}$$
Assumption B.(ref)$(iv)$ is satisfied.\\
For all $t\in[0,1)$, $l=1,\dots,d_1$, $\theta_l\in\Theta_l$ and $\eta_l\in\mathcal{T}_l\setminus\{\eta_{0,l}\}$, we have
\begin{align*}
&\mathbb{E}\left[\left(\psi_l(W,\theta_l,\eta_l)-\psi_l(W,\theta_{0,l},\eta_{0,l})\right)^2\right]\\
=\ &\mathbb{E}\left[\left(\psi_l(W,\theta_l,\eta_l)-\psi_l(W,\theta_{0,l},\eta_l)+\psi_l(W,\theta_{0,l},\eta_l)-\psi_l(W,\theta_{0,l},\eta_{0,l})\right)^2\right]\\
\le\ &C\bigg(\mathbb{E}\left[\left(\psi_l(W,\theta_l,\eta_l)-\psi_l(W,\theta_{0,l},\eta_l)\right)^2\right]\\
&\quad\vee \mathbb{E}\left[\left(\psi_l(W,\theta_{0,l},\eta_l)-\psi_l(W,\theta_{0,l},\eta_{0,l})\right)^2\right]\bigg)
\end{align*}
with
\begin{align*}
&\mathbb{E}\left[\left(\psi_l(W,\theta_l,\eta_l)-\psi_l(W,\theta_{0,l},\eta_l)\right)^2\right]\\
=\ &|\theta_l-\theta_{0,l}|^2\mathbb{E}\left[\left(g_l(X_1)(g_l(X_1)-(\eta_l^{(2)})^T Z_{-l})\right)^2\right]\\
\le\ &C |\theta_l-\theta_{0,l}|^2 \mathbb{E}\left[\left(\nu^{(l)}+(\eta_{0,l}^{(2)}-\eta_l^{(2)})^T Z_{-l} \right)^2\right]\\
\le\ &C |\theta_l-\theta_{0,l}|^2\left(\|\nu^{(l)}\|^2_{P,2}\vee\|(\eta_{0,l}^{(2)}-\eta_l^{(2)})^T Z_{-l}\|^2_{P,2}\right)\\
\le\ &C |\theta_l-\theta_{0,l}|^2\left(\varsigma_n^{-1}\vee\|\eta_{0,l}^{(2)}-\eta_l^{(2)}\|^2_2\varsigma_n^{-1}\right)\\
\le\ &C \varsigma_n^{-1}|\theta_l-\theta_{0,l}|^2
\end{align*}
due to Assumption A.(ref)$(ii)$, $(iv)$ and the definition of $\mathcal{T}_l$. With similar arguments, we obtain
\begin{align*}
&\mathbb{E}\left[\left(\psi_l(W,\theta_{0,l},\eta_l)-\psi_l(W,\theta_{0,l},\eta_{0,l})\right)^2\right]\\
=\ &\mathbb{E}\Bigg[\bigg(\Big(Y-\theta_{0,l}g_l(X_1)-(\eta_l^{(1)})^T Z_{-l}-\eta_l^{(3)}(X)\Big)\Big(g_l(X_1)-(\eta_l^{(2)})^T Z_{-l}\Big)\\
&\quad - \Big(Y-\theta_{0,l}g_l(X_1)-(\eta_{0,l}^{(1)})^T Z_{-l}-\eta_{0,l}^{(3)}(X)\Big)\Big(g_l(X_1)-(\eta_{0,l}^{(2)})^T Z_{-l}\Big)\bigg)^2\Bigg]\\
=\ &\mathbb{E}\Bigg[\bigg(\Big(Y-\theta_{0,l}g_l(X_1)-(\eta_l^{(1)})^T Z_{-l}-\eta_l^{(3)}(X)\Big)\\
&\quad\cdot\Big((\eta_{0,l}^{(2)}-\eta_l^{(2)})^T Z_{-l}\Big)\\
&\quad + \Big(g_l(X_1)-(\eta_{0,l}^{(2)})^T Z_{-l}\Big)\\
&\quad \cdot\Big((\eta_{0,l}^{(1)}-\eta_l^{(1)})^T Z_{-l}+\eta_{0,l}^{(3)}(X)-\eta_l^{(3)}(X)\Big)\bigg)^2\Bigg]\\
\le\ &C \bigg(\varsigma_n^{-1/2}\|\eta_{0,l}^{(2)}-\eta_l^{(2)}\|_2\vee \varsigma_n^{-1/2}\|\eta_{0,l}^{(1)}-\eta_l^{(1)}\|_2\vee\|\eta^{(3)}_{0,l}(X)\|_{P,2}\bigg)^2\\
\le\ &C\|\eta_{0,l}-\eta_l\|_e^2,
\end{align*}
where we used the definition of $\mathcal{T}_l$, A.(ref)$(ii)$ and
\begin{align*}
\sup_{\|\xi\|_2=1}\mathbb{E}[(\xi^T Z)^2]\le C\varsigma_n^{-1}.
\end{align*}
Therefore, Assumption B.(ref)$(v)(a)$ holds since it is straightforward to show Assumption B.(ref)$(v)(a)$ for $\eta_l = \eta_{0,l}$. It holds
{\allowdisplaybreaks
\begin{align*}
&\bigg|\partial_t\mathbb{E}\Big[\psi_l(W,\theta_l,\eta_{0,l}+t(\eta_l-\eta_{0,l}))\Big]\bigg|\\
=\ &\bigg|\mathbb{E}\bigg[\partial_t\bigg\{\Big(Y-\theta_{0,l} g_l(X_1)-\big(\eta^{(1)}_{0,l}+t(\eta_l^{(1)}-\eta^{(1)}_{0,l})\big)^T Z_{-l}\\
&\quad -\big(\eta^{(3)}_{0,l}(X)+t(\eta_l^{(3)}(X)-\eta^{(3)}_{0,l}(X))\big)\Big)\\
&\quad\cdot\Big(g_l(X_1)-\big(\eta^{(2)}_{0,l}+t(\eta_l^{(2)}-\eta^{(2)}_{0,l})\big)^T Z_{-l}\Big)\bigg\}\bigg]\bigg|\\
=\ &\bigg|\mathbb{E}\bigg[\Big(Y-\theta_{0,l} g_l(X_1)-(\eta^{(1)}_{0,l}+t(\eta_l^{(1)}-\eta^{(1)}_{0,l}))^T Z_{-l}\\
&\quad -(\eta^{(3)}_{0,l}(X)+t(\eta_l^{(3)}(X)-\eta^{(3)}_{0,l}(X)))\Big)\\
&\quad \cdot\Big((\eta^{(2)}_{0,l}-\eta_l^{(2)})^T Z_{-l} \Big)\\
&\quad +\Big(g_l(X_1)-(\eta^{(2)}_{0,l}+t(\eta_l^{(2)}-\eta^{(2)}_{0,l}))^T Z_{-l}\Big)\\
&\quad\cdot\Big((\eta^{(1)}_{0,l}-\eta_l^{(1)})^T Z_{-l}+\eta^{(3)}_{0,l}(X)-\eta_l^{(3)}(X)\Big)\bigg]\bigg|\\
=\ &| I_{1,1} + I_{1,2} + I_{1,3}+ I_{1,4}|
\end{align*}
}
with
\begin{align*}
I_{1,1}&=\mathbb{E}\bigg[\Big(Y-\theta_{0,l} g_l(X_1)-(\eta^{(1)}_{0,l}+t(\eta_l^{(1)}-\eta^{(1)}_{0,l}))^T Z_{-l}\\
&\quad -(\eta^{(3)}_{0,l}(X)+t(\eta_l^{(3)}(X)-\eta^{(3)}_{0,l}(X)))\Big)\Big((\eta^{(2)}_{0,l}-\eta_l^{(2)}))^T Z_{-l}\Big)\bigg]\\
&\le C \varsigma_n^{-1}\|\eta^{(2)}_{0,l}-\eta_l^{(2)}\|_2,\\
I_{1,2}&=\mathbb{E}\bigg[\Big(Y-\theta_{0,l} g_l(X_1)-(\eta^{(1)}_{0,l}+t(\eta_l^{(1)}-\eta^{(1)}_{0,l}))^T Z_{-l}\\
&\quad -(\eta^{(3)}_{0,l}(X)+t(\eta_l^{(3)}(X)-\eta^{(3)}_{0,l}(X)))\Big)\bigg]\\
&=-t\mathbb{E}\bigg[\Big((\eta_l^{(1)}-\eta^{(1)}_{0,l})^T Z_{-l}+(\eta_l^{(3)}(X)-\eta^{(3)}_{0,l}(X))\Big)\bigg]\\
&\le C\varsigma_n^{-1/2},\\
I_{1,3}&=\mathbb{E}\bigg[\Big(g_l(X_1)-(\eta^{(2)}_{0,l}+t(\eta_l^{(2)}-\eta^{(2)}_{0,l}))^T Z_{-l}\Big)\Big((\eta^{(1)}_{0,l}-\eta_l^{(1)})^T Z_{-l}\Big)\bigg]\\
&\le C \varsigma_n^{-1}\|\eta^{(1)}_{0,l}-\eta_l^{(1)}\|_2
\end{align*}
since $\mathbb{E}[\varepsilon|X]=0$ and $\mathbb{E}[\nu^{(l)}Z_{-l}]=0$. Further, it holds
\begin{align*}
I_{1,4}&=\mathbb{E}\bigg[\Big(g_l(X_1)-(\eta^{(2)}_{0,l}+t(\eta_l^{(2)}-\eta^{(2)}_{0,l}))^T Z_{-l}\Big)\Big(\eta^{(3)}_{0,l}(X)\Big)\bigg]\\
&\le C\varsigma_n^{-1/2}\|\eta^{(3)}_{0,l}(X)\|_{P,2}
\end{align*}
since
\begin{align*}
&\quad\mathbb{E}\bigg[\Big(g_l(X_1)-(\eta^{(2)}_{0,l}+t(\eta_l^{(2)}-\eta^{(2)}_{0,l}))^T Z_{-l}\Big)^2\bigg]^{1/2}\\
&\le \mathbb{E}[(\nu^{(l)})^2]^{1/2}+\mathbb{E}\left[\left((\eta_l^{(2)}-\eta^{(2)}_{0,l})^T Z_{-l}\right)^2\right]^{1/2}\\
&\le C\varsigma_n^{-1/2}\left(1+\|\eta_l^{(2)}-\eta^{(2)}_{0,l}\|_2^2\right)^{1/2}.
\end{align*}
This implies Assumption B.(ref)$(v)(b)$. Finally, to obtain Assumption B.(ref)$(v)(c)$, we note that
{\allowdisplaybreaks
\begin{align*}
&\partial_t^2\mathbb{E}\left[\psi_l(W,\theta_{0,l}+t(\theta_l-\theta_{0,l}),\eta_{0,l}+t(\eta_l-\eta_{0,l}))\right]\\
=\ &\partial_t\mathbb{E}\bigg[\Big(Y-\big(\theta_{0,l}+t(\theta_l-\theta_{0,l}) \big)g_l(X_1)-\big(\eta^{(1)}_{0,l}+t(\eta_l^{(1)}-\eta^{(1)}_{0,l})\big)^T Z_{-l}\\
&\quad -\big(\eta^{(3)}_{0,l}(X)+t(\eta_l^{(3)}(X)-\eta^{(3)}_{0,l}(X))\big)\Big)\\
&\quad \cdot\Big((\eta^{(2)}_{0,l}-\eta_l^{(2)}))^T Z_{-l}\Big)\\
&\quad +\Big(g_l(X_1)-\big(\eta^{(2)}_{0,l}+t(\eta_l^{(2)}-\eta^{(2)}_{0,l})\big)^T Z_{-l}\Big)\\
&\quad\cdot\Big((\theta_{0,l}-\theta_{l}) )g_l(X_1)+(\eta^{(1)}_{0,l}-\eta_l^{(1)})^T Z_{-l}+\eta^{(3)}_{0,l}(X)\Big)\bigg]\\
=\ &2\mathbb{E}\bigg[\Big((\theta_{0,l}-\theta_{l}) g_l(X_1)+(\eta_{0,l}^{(1)}-\eta^{(1)}_{l})^T Z_{-l} + \eta_{0,l}^{(3)}(X)\Big)\\
&\quad\cdot\Big((\eta^{(2)}_{0,l}-\eta_l^{(2)}))^T Z_{-l}\Big)\bigg]\\
\le &2\mathbb{E}\bigg[\Big((\theta_{0,l}-\theta_{l}) g_l(X_1)+(\eta_{0,l}^{(1)}-\eta^{(1)}_{l})^T Z_{-l} + \eta_{0,l}^{(3)}(X)\Big)^2\bigg]\\
&\quad\vee 2\mathbb{E}\bigg[\Big((\eta^{(2)}_{0,l}-\eta_l^{(2)}))^T Z_{-l}\Big)^2\bigg]\\
\le &C \left(|\theta_{0,l}-\theta_l|^2\varsigma_n^{-1}\vee \|\eta_{0,l}-\eta_l\|_e^2\right)
\end{align*}
}
using the same arguments as above.\\ \\
Assumption B.(ref)\\
Note that the Assumptions B.(ref)$(ii)$ and $(iii)$ both hold by the construction of $\mathcal{T}_l$ and the Assumptions A.(ref)$(ii)$ and A.(ref)$(ii)$. The main part to verify Assumption B.(ref) is to show that the estimates of the nuisance function are contained in the nuisance realization set with high probability. We will rely on uniform lasso estimation results stated in Appendix (ref). Consider a scaled version of the auxiliary regression
\begin{align*}
\tilde{Z}_l = (\gamma_0^{(l)})^T \tilde{Z}_{-l} + \tilde{\nu}^{(l)}
\end{align*}
with $\tilde{Z}_l:= \varsigma_n^{1/2} Z_l$ and $ \tilde{\nu}^{(l)}:= \varsigma_n^{1/2} \nu^{(l)}$. First, we want to proof that $\hat{\eta}^{(2)}_{l}\in\mathcal{T}_l$ for all $l=1,\dots,d_1$. We have to check the Assumptions C.(ref)$(i)$ to $(iv)$. Due to Assumption A.(ref)$(iii)$, the first part of Assumption C.(ref)$(i)$ is satisfied with $M_n = \varsigma_n^{1/2},$ since we have already shown $|\nu^{(l)}|\le C$ almost surely uniformly over $l = 1,\dots, d_1$. Further,
$$c\le\mathbb{E}\left[(\tilde{\nu}^{(l)})^2 \tilde{Z}_{-l,j}^2\right]\le C,$$
uniformly for all $l=1,\dots,d_1$, $j = 1,\dots, d_1 -1$ and
\begin{align*}
\max\limits_{l=1,\dots,d_1}\max\limits_{j=1,\dots,d_1 - 1}\|\tilde{\nu}^{(l)} \tilde{Z}_{-l,j}\|_{P,3}&\le CK_{n},\\
\max\limits_{j=1,\dots,d_1}\max\limits_{j=1,\dots,d_1 -1} \mathbb{E}\left[(\tilde{\nu}^{(l)})^4 \tilde{Z}_{-l,j}^4\right]&\le C L_n
\end{align*}
due to Assumption A.(ref)$(iv)$. Assumption A.(ref)$(iv)$ also directly implies Assumption C.(ref)$(ii)$. Combining Assumption C.(ref)$(ii)$ with Assumption A.(ref)$(ii)$ implies Assumption C.(ref)$(iii)$, since
\begin{align*}
\max\limits_{l=1,\dots,d_1}\|\gamma_0^{(l,2)}\|_{2}^2\le C\varsigma_n \max\limits_{l=1,\dots,d_1}\|(\gamma_0^{(l,2)})^T Z_{-l}\|_{P,2}^2 \le C s_2\log(\bar{d}_n)/n.
\end{align*}
The growth conditions C.(ref)$(iv)$ are included in Assumption A.(ref)$(v)$.
Therefore,
$$\hat{\eta}^{(2)}_{l}\in \mathcal{T}_l\quad\text{for all } l=1,\dots,d_1$$ with probability $1-o(1)$. To estimate $\eta^{(1)}_{0,l}$, we run a lasso regression of $Y$ on $\tilde{Z}$:
\begin{align*}
Y = (\varsigma_n^{-1/2}\theta_0)^T (\varsigma_n^{1/2} g(X_1)) + (\varsigma_n^{-1/2}\beta_0)^T (\varsigma_n^{1/2} h(X_{-1})) + b_1(X_1) + b_{-1}(X_{-1}) + \varepsilon
\end{align*}
Define the corresponding nuisance parameter
$$\tilde{\beta}_0^{(l)}:=\varsigma_n^{-1/2}\beta_0^{(l)} = \varsigma_n^{-1/2}(\theta_{0,1},\dots,\theta_{0,l-1},\theta_{0,l+1},\dots\theta_{0,d_1},\beta_{0,1},\dots,\beta_{0,d_2})^T.$$
With analogous arguments as in the proof of Theorem (ref), it holds
\begin{align*}
\|\tilde{\beta}_0^{(l)}-\hat{\tilde{\beta}}^{(l)}\|_0&\le Cs_1,\\
\|\tilde{\beta}_0^{(l)}-\hat{\tilde{\beta}}^{(l)}\|_2&\le C\sqrt{\frac{s_1\log(\bar{d}_n)}{n}},\\
\|\tilde{\beta}_0^{(l)}-\hat{\tilde{\beta}}^{(l)}\|_1&\le C\sqrt{\frac{s_1^2\log(\bar{d}_n)}{n}}
\end{align*}
with probability $1-o(1)$ using Assumptions A.(ref)$(ii)$, A.(ref)$(ii)$-$(v)$ and
\begin{align*}
c\le\mathbb{E}\big[\varepsilon^2 \tilde{Z}_{l}^2\big]&= \mathbb{E}\big[\tilde{Z}_{l}^2\underbrace{\mathbb{E}[\varepsilon^2|X]}_{=\operatorname{Var}(\varepsilon|X)} \big]\le C.
\end{align*}
This directly implies that with probability $1-o(1)$ the nuisance realization set $\mathcal{T}_l$ contains $\hat{\eta}^{(1)}_{l}$ for all $l=1,\dots,d_1$.\\
Combining the results above with $\hat{\eta}^{(3)}\equiv 0$, we obtain Assumption B.(ref)$(i)$. Define
\begin{align*}
\mathcal{F}_1:=\big\{\psi_l(\cdot,\theta_l,\eta_l):l=1,\dots,d_1,\theta_l\in\Theta_l,\eta_l\in\mathcal{T}_l\big\}.
\end{align*}
To bound the complexity of $\mathcal{F}_1$, we exclude the true nuisance function (the true nuisance function is the only element of $\mathcal{T}_l$ with a nonzero approximation error):
\begin{align*}
\mathcal{F}_{1,1}:=\big\{\psi_l(\cdot,\theta_l,\eta_l):l=1,\dots,d_1,\theta_l\in\Theta_l,\eta_l\in\mathcal{T}_l\setminus\{\eta_0^{(l)}\}\big\}\subseteq \mathcal{F}_{1,1}^{(1)}\mathcal{F}_{1,1}^{(2)}
\end{align*}
with
\begin{align*}
\mathcal{F}_{1,1}^{(1)}&:=\big\{W\mapsto Y-\theta_l g_l(X_1)-(\eta^{(1)}_l)^T Z_{-l}:l=1,\dots,d_1,\theta_l\in\Theta_l,\eta_l\in\mathcal{T}_l\setminus\{\eta_0^{(l)}\}\big\}\\
\mathcal{F}_{1,1}^{(2)}&:=\big\{W\mapsto g_l(X_1)-(\eta^{(2)}_l)^T Z_{-l}:l=1,\dots,d_1,\theta_l\in\Theta_l,\eta_l\in\mathcal{T}_l\setminus\{\eta_0^{(l)}\}\big\}.
\end{align*}
Note that the envelope $F_{1,1}^{(1)}$ of $\mathcal{F}_{1,1}^{(1)}$ satisfies
\begin{align*}
\|F_{1,1}^{(1)}\|_{P,2q}&\le \bigg\|\sup_{l=1,\dots,d_1}\sup_{\theta_l\in\Theta_l,\|\eta_{0,l}^{(1)}-\eta_l^{(l)}\|_1\le C\sqrt{s}\tau_n}\Big(|\varepsilon| + |\eta_{0}^{(3)}(X)|\\
&\quad + |(\theta_{0,l}-\theta_l)g_l(X_1)|+|(\eta_{0,l}^{(1)}-\eta_l^{(1)})^T Z_{-l}|\Big)\bigg\|_{P,2q}\\
&\lesssim \|\varepsilon\|_{P,2q}+\| \eta_{0}^{(3)}(X)\|_{P,2q}+ \|\sup_{l=1,\dots,d_1}g_l(X_1)\|_{P,2q}\\
&\quad + \varsigma_n^{-1/2}\sqrt{s_1}\tau_n\|\sup_{j=1,\dots,d_1+d_2} Z_j\|_{P,2q}\\
&\lesssim C + \varsigma_n^{-1/2}\sqrt{s_1}\tau_n
\end{align*}
due to A.(ref)$(ii)$, A.(ref)$(iii)$ and analogously
\begin{align*}
\|F_{1,1}^{(2)}\|_{P,2q}\lesssim C + \sqrt{s_2}\tau_n.
\end{align*}
Next, note that due to Lemma 2.6.15 from vanweak the set
\begin{align*}
\mathcal{G}_{1,1}:=\big\{Z\mapsto \xi^T Z: \xi\in \mathbb{R}^{d_1+d_2+1},\|\xi\|_0\le Cs,\|\xi\|_2\le C\big\}
\end{align*}
is a union over $\binom{d_1+d_2+1}{Cs}$ VC-subgraph classes $\mathcal{G}_{1,1,k}$ with VC indices less or equal to $Cs+2$. Therefore, $\mathcal{F}_{1,1}^{(1)}$ and $\mathcal{F}_{1,1}^{(2)}$ are unions over $\binom{d_1+d_2+1}{Cs}$ respectively $\binom{d_1+d_2}{Cs}$ VC-subgraph classes, which combined with Theorem 2.6.7 from vanweak implies
\begin{align*}
\sup_Q\log N(\varepsilon\|F_{1,1}^{(1)}\|_{Q,2},\mathcal{F}_{1,1}^{(1)},\|\cdot\|_{Q,2})\lesssim s_1\log\left(\frac{d_1+d_2}{\varepsilon}\right)
\end{align*}
and
\begin{align*}
\sup_Q\log N(\varepsilon\|F_{1,1}^{(2)}\|_{Q,2},\mathcal{F}_{1,1}^{(2)},\|\cdot\|_{Q,2})\lesssim s_2\log\left(\frac{d_1+d_2}{\varepsilon}\right).
\end{align*}
Using basic calculations, we obtain
\begin{align*}
\sup_Q\log N(\varepsilon\|F_{1,1}\|_{Q,2},\mathcal{F}_{1,1},\|\cdot\|_{Q,2})\lesssim s\log\left(\frac{d_1+d_2}{\varepsilon}\right),
\end{align*}
where $F_{1,1}:=F_{1,1}^{(1)}F_{1,1}^{(2)}$ is an envelope for $\mathcal{F}_{1,1}$ with
\begin{align*}
\|F_{1,1}\|_{P,q}\le \|F_{1,1}^{(1)}\|_{P,2q}\|F_{1,1}^{(2)}\|_{P,2q}\lesssim (C+\varsigma_n^{-1/2}\sqrt{s_1\tau_n})(C+\sqrt{s_2\tau_n})\lesssim C.
\end{align*}
Define
$$\mathcal{F}_{1,2}:=\big\{\psi_l(\cdot,\theta_l,\eta_{0,l}):l=1,\dots,d_1,\theta_l\in\Theta_l\big\}$$
and, with an analogous argument, we obtain
\begin{align*}
\sup_Q\log N(\varepsilon\|F_{1,2}\|_{Q,2},\mathcal{F}_{1,2},\|\cdot\|_{Q,2})\lesssim \log\left(\frac{d_1}{\varepsilon}\right),
\end{align*}
where the envelope $F_{1,2}$ of $\mathcal{F}_{1,2}$ obeys
\begin{align*}
\|F_{1,2}\|_{P,q}\lesssim C.
\end{align*}
Combining the results above, we obtain
\begin{align*}
\sup_Q\log N(\varepsilon\|F_{1}\|_{Q,2},\mathcal{F}_{1},\|\cdot\|_{Q,2})\lesssim s\log\left(\frac{d_1+d_2}{\varepsilon}\right),
\end{align*}
where the envelope $F_1:=F_{1,1}^{(1)}F_{1,1}^{(2)}\vee F_{1,2}$ of $\mathcal{F}_1$ satisfies
\begin{align*}
\|F_{1}\|_{P,q}\lesssim C.
\end{align*}
Therefore, Assumption B.(ref)$(iv)$ holds with $\upsilon_n\lesssim s$ and $a_n = d_1\vee d_2$.\\
At first remark that for each $l=1,\dots,d_1$, we have
\begin{align*}
c&\le \operatorname{Var}(\varepsilon|X)\\
&\le \mathbb{E}\left[\big(Y-\theta_lg_l(X_1)-(\eta_l^{(1)})^T Z_{-l} -\eta^{(3)}(X)\big)^2|X\right]\\
&= \mathbb{E}\left[\big(\varepsilon+ (\theta_{0,l}-\theta_l)g_l(X_1)+(\eta_{0,l}^{(1)} -\eta_l^{(1)})^T Z_{-l} +\eta_{0}^{(3)}(X)\big)^2|X\right]\\
&= \underbrace{\mathbb{E}\left[\varepsilon^2|X\right]}_{=\operatorname{Var}(\varepsilon|X)\le C}+\mathbb{E}\bigg[\big(\underbrace{(\theta_{0,l}-\theta_l)g_l(X_1)}_{\le C}+\underbrace{(\eta_{0,l}^{(1)} -\eta_l^{(1)})^T Z_{-l}}_{\le \|\eta_{0,l}^{(1)} -\eta_l^{(1)}\|_1 C\le C} + \underbrace{\eta_{0}^{(3)}(X)}_{\le C}\big)^2|X\bigg]\\
&\le C,
\end{align*}
due to $c\le Var(\epsilon|X)\le C$ and the growth rates.
For all $f\in\mathcal{F}_1$, we have
\begin{align*}
&\mathbb{E}\left[f^2\right]^{\frac{1}{2}}\\
=&\ \mathbb{E}\left[\big(Y-\theta_lg_l(X_1)-(\eta^{(1)})^T Z_{-l} -\eta^{(3)}(X)\big)^2\big(g_l(X_1)-(\eta^{(2)})^T Z_{-l}\big)^2\right]^{\frac{1}{2}}\\
=&\ \mathbb{E}\Big[\big(g_l(X_1)-(\eta^{(2)})^T Z_{-l}\big)^2\\
&\quad\cdot\mathbb{E}\left[\big(Y-\theta_lg_l(X_1)-(\eta^{(1)})^T Z_{-l} -\eta^{(3)}(X)\big)^2|X\right]\Big]^{\frac{1}{2}}\\
\end{align*}
such that
\begin{align*}
c\varsigma_n^{-\frac{1}{2}}\le \mathbb{E}[f^2]^{\frac{1}{2}}\le C\varsigma_n^{-\frac{1}{2}},
\end{align*}
due to Assumption A.(ref)$(iv)$. This corresponds to Assumption B.(ref)$(v)$. Assumption B.(ref)$(vi)(a)$ holds by the definition of $\tau_n$ and $\upsilon_n\lesssim s$. To verify the next growth condition, we note
\begin{align*}
&(\tau_n+\mathcal{S}_n\log(n)/\sqrt{n})(\upsilon_n\log(\bar{d}_n))^{1/2}+n^{-1/2+1/q}\upsilon_n \log(\bar{d}_n)\\
\lesssim\ &(\tau_n+\log^\frac{1}{2}(d_1)\log(n)/\sqrt{n})(s\log(\bar{d}_n))^{1/2}+n^{-1/2+1/q}s\log(\bar{d}_n)\\
\lesssim\ &\left(n^{\frac{2}{q}}\frac{s^2\log^{2}(\bar{d}_n)}{n}\right)^{\frac{1}{2}}\\
\lesssim\ &\delta_n\varsigma_n^{-1/2}
\end{align*}
for a given $q\ge 4$ with $$\delta_n =\sqrt{\frac{n^{2/q}s^2\varsigma_n\log^2(\bar{d}_n)}{n}}=o(1)$$ due to Assumption A.(ref)$(v)$ and analogously
\begin{align*}n^{1/2}\tau_n^2=\frac{s\log(\bar{d}_n)}{\sqrt{n}}\lesssim\delta_n\varsigma_n^{-1/2}.
\end{align*}
\\ \\
Assumption B.(ref)$(i)-(ii)$ \\
Define
\begin{align*}
\mathcal{F}_0:=\{\psi_x(\cdot):x\in I\},
\end{align*}
where $\psi_x(\cdot):=(g(x)^T\Sigma_n g(x))^{-1/2}g(x)^TJ_0^{-1}\psi(\cdot,\theta_{0},\eta_{0})$. It is easy to verify that B.(ref)$(ii)$ holds with
$$L_n= t_1^{9/2}\varsigma_n^{3/2}.$$
Using the same argument, we can conclude that the envelope $F_0$ of $\mathcal{F}_0$ satisfies
\begin{align*}
\|F_0\|_{P,q}&=\mathbb{E}\left[\sup_{x\in I}\left|(g(x)^T\Sigma_n g(x))^{-1/2}g(x)^TJ_0^{-1}\psi(W,\theta_{0},\eta_{0})\right|^q\right]^{\frac{1}{q}}\\
&\lesssim t_1^{1/2}\varsigma_n^{-1/2} \mathbb{E}\left[\sup_{x\in I}\left|g(x)^TJ_0^{-1}\psi(W,\theta_{0},\eta_{0})\right|^q\right]^{\frac{1}{q}}\\
&= t_1^{1/2}\varsigma_n^{-1/2}\mathbb{E}\left[\sup_{x\in I}\left|\sum_{l=1}^{d_1} g_l(x)J_{0,l}^{-1}\psi_l(W,\theta_{0,l},\eta_{0,l})\right|^q\right]^{\frac{1}{q}}\\
&\lesssim t_1^{3/2}\varsigma_n^{1/2} \mathbb{E}\left[\sup_{l=1,\dots,d_1}\left|\varepsilon \nu^{(l)}\right|^q\right]^{\frac{1}{q}}\\
&\lesssim t_1^{3/2}\varsigma_n^{1/2}
\end{align*}
with $t_1^{3/2}\varsigma_n^{1/2}
\lesssim t_1^{9/2}\varsigma_n^{3/2}=L_n$.
It is worth noting that this can be relaxed to $\|F_0\|_{P,q}\lesssim t_1^{1/2}\varsigma_n^{1/2}$ if B-Splines are used for approximation.
To bound the entropy of $\mathcal{F}_0$, we note that
\begin{align*}
&\quad\big\|\psi_x(W)-\psi_{\tilde{x}}(W)\big\|_{P,2}\\
&=\Big\|(g(x)^T\Sigma_n g(x))^{-1/2}\sum_{l=1}^{d_1} g_l(x) \mathbb{E}[(\nu^{(l)})^2]^{-1}\psi_l(W,\theta_{0,l},\eta_{0,l})\\
&\quad -(g(\tilde{x})^T\Sigma_n g(\tilde{x}))^{-1/2}\sum_{l=1}^{d_1} g_l(\tilde{x}) \mathbb{E}[(\nu^{(l)})^2]^{-1}\psi_l(W,\theta_{0,l},\eta_{0,l})\Big\|_{P,2}\\
&\le | (g(x)^T\Sigma_n g(x))^{-1/2}-(g(\tilde{x})^T\Sigma_n g(\tilde{x}))^{-1/2}|\Big\| g(x)^T J_0^{-1}\psi(W,\theta_{0,l},\eta_{0,l})\Big\|_{P,2}\\
&\quad +(g(\tilde{x})^T\Sigma_n g(\tilde{x}))^{-1/2}\Big\|\big(g(x) -g(\tilde{x})\big)^TJ_0^{-1}\psi(W,\theta_{0,l},\eta_{0,l})\Big\|_{P,2}\\
&\lesssim | (g(x)^T\Sigma_n g(x))^{-1/2}-(g(\tilde{x})^T\Sigma_n g(\tilde{x}))^{-1/2}|\sup_{x\in I}\|g(x)\|_2\varsigma_n^{1/2}\\
&\quad +t_1^{1/2}\|g(x)-g(\tilde{x})\|_2
\end{align*}
due to the eigenvalues of $\Sigma_n$.
Additionally, it holds
\begin{align*}
&|(g(x)^T\Sigma_n g(x))^{-1/2}-(g(\tilde{x})^T\Sigma_n g(\tilde{x}))^{-1/2}|\\
\lesssim\ &\left|\left(\frac{g(\tilde{x})^T\Sigma_n g(\tilde{x})}{g(x)^T\Sigma_n g(x)}\right)^{1/2}-1\right|\varsigma_n^{-1/2} t_1^{1/2}\\
\lesssim\ &|g(\tilde{x})^T\Sigma_n g(\tilde{x})-g(x)^T\Sigma_n g(x)|\varsigma_n^{-3/2} t_1^{3/2}\\
=\ &|(g(x)-g(\tilde{x}))^T\Sigma_n (g(x)+g(\tilde{x}))|\varsigma_n^{-3/2} t_1^{3/2}\\
\lesssim\ &\|g(x)-g(\tilde{x})\|_2\sup_x\|g(x)\|_2\varsigma_n^{-1/2} t_1^{3/2}
\end{align*}
which implies
\begin{align*}
\big\|\psi_x(W)-\psi_{\tilde{x}}(W)\big\|_{P,2}\lesssim \|g(x)-g(\tilde{x})\|_2 t_1^{3/2}.
\end{align*}
Using the same argument as in Theorem 2.7.11 from vanweak, we obtain
\begin{align*}
&\quad\sup_Q\log N(\varepsilon\|F_{0}\|_{Q,2},\mathcal{F}_0,\|\cdot\|_{Q,2})\\
&\lesssim \sup_Q\log N\left(\left(\frac{\varepsilon t_1^{3/2}\varsigma_n^{1/2}}{t_1^{3/2}}\right) t_1^{3/2},\mathcal{F}_0,\|\cdot\|_{Q,2}\right)\\
&\le \log N\left(\left(\varepsilon \varsigma_n^{1/2}\right),g(I),\|\cdot\|_{2}\right)\\
&\lesssim t_1\log\left(\frac{\tilde{A}_n \varsigma_n^{-1/2}}{\varepsilon}\right)
\end{align*}
by Assumption A.(ref)$(i)$. Therefore, Assumption B.(ref)$(i)$ is satisfied with $\varrho_n=t_1$ and $A_n = \tilde{A}_n\varsigma_n^{-1/2}.$\\ \\
Assumption B.(ref)\\
Next, we want to prove that with probability $1-o(1)$ it holds
\begin{align}
\sup_{l=1,\dots,d_1}|\hat{J}_l-J_{0,l}|=\varsigma_n^{-1}\tilde{\epsilon}_n
\end{align}
with $\tilde{\epsilon}_n=o(1)$ where $\hat{J}_l=\mathbb{E}_n[-g_l(X_1)(g_l(X_1)-(\hat{\eta}_l^{(2)})^TZ_{-l})]$. It holds
\begin{align*}
|\hat{J}_l-J_{0,l}|&\le |\hat{J}_l-\mathbb{E}[-g_l(X_1)(g_l(X_1)-(\hat{\eta}_l^{(2)})^TZ_{-l})]|\\
&\quad+|\mathbb{E}[-g_l(X_1)(g_l(X_1)-(\hat{\eta}_l^{(2)})^TZ_{-l})]+J_{0,l}|
\end{align*}
with
\begin{align*}
&|\mathbb{E}[-g_l(X_1)(g_l(X_1)-(\hat{\eta}_l^{(2)})^TZ_{-l})]+J_{0,l}|\\
\le&|\mathbb{E}[g_l(X_1)(\hat{\eta}_l^{(2)}-\eta_{0,l}^{(2)})^T Z_{-l})]|\\
\lesssim &\ \tau_n\varsigma_n^{-1}.
\end{align*}
Let
\begin{align*}
\tilde{\mathcal{G}}_1:=\bigg\{&X\mapsto -g_l(X_1)(g_l(X_1)-(\eta_l^{(2)})^TZ_{-l}):l=1,\dots,d_1,\|\eta_l^{(2)}\|_0\le Cs_2,\\
&\|\eta^{(2)}_l-\eta^{(2)}_{0,l}\|_2\le C\tau_n,\|\eta^{(2)}-\eta^{(2)}_{0,l}\|_1\le C\sqrt{s_2}\tau_n\bigg\}.
\end{align*}
For any $q\ge 2$, the envelope $\tilde{G_1}$ of $\tilde{\mathcal{G}}_1$ satisfies
\begin{align*}
\mathbb{E}[\tilde{G}_1^q]^{\frac{1}{q}}&\le\mathbb{E} \left[\sup_{l=1,\dots,d_1}\sup_{\eta^{(2)}:\|\eta^{(2)}_l-\eta^{(2)}_{0,l}\|_1\le C\sqrt{s}\tau_n} |g_l(X_1)|^q|(g_l(X_1)-(\eta_l^{(2)})^TZ_{-l})|^q\right]^{\frac{1}{q}}\\
&\lesssim \|\sup_{l=1,\dots,d_1}\nu^{(l)}\|_{P,q}\vee
\mathbb{E}\bigg[\sup_{l=1,\dots,d_1}\sup_{\eta^{(2)}:\|\eta^{(2)}_l-\eta^{(2)}_{0,l}\|_1\le C\sqrt{s}\tau_n}(\eta_{0,l}^{(2)}-\eta_l^{(2)})^TZ_{-l})^{q}\bigg]^{\frac{1}{q}}\\
&\lesssim C \vee \sqrt{s_2}\tau_n\\
&\lesssim C
\end{align*}
and, with similar arguments as above, we obtain
\begin{align*}
\sup_{g\in \tilde{\mathcal{G}}_1}\mathbb{E}[g^2]^{\frac{1}{2}}
&\lesssim \varsigma_n^{-1/2} \vee\tau_n\varsigma_n^{-1/2}
\lesssim \varsigma_n^{-1/2}.
\end{align*}
Further, we have
\begin{align*}
\sup_Q\log N(\varepsilon\|\tilde{G}_1\|_{Q,2},\tilde{\mathcal{G}}_1,\|\cdot\|_{Q,2})\lesssim s_2\log\left(\frac{d_1+d_2}{\varepsilon}\right).
\end{align*}
Therefore, by using Lemma P.2 from belloni2018uniformly with $\sigma=\varsigma_n^{-1/2}$, it holds
\begin{align*}
\sup_{l=1,\dots,d_1}|\hat{J}_l-J_{0,l}|&\lesssim \sup_{f\in\tilde{\mathcal{G}}_1}|\mathbb{E}_n[f(X)]-\mathbb{E}[f(X)]|+\tau_n\varsigma_n^{-1}\\
&\lesssim K\left(\varsigma_n^{-1/2}\sqrt{\frac{s_2\log(\bar{d}_n\varsigma_n^{1/2})}{n}}+n^{\frac{1}{q}}\frac{s_2\log^(\bar{d}_n\varsigma_n^{1/2})}{n}\right) +\tau_n\varsigma_n^{-1}\\
&\lesssim\varsigma_n^{-1}\tilde{\epsilon}_n
\end{align*}
with probability $1-o(1)$ and
$$\tau_n\lesssim\tilde{\epsilon}_n\lesssim\left(\sqrt{\frac{s_2\varsigma_n\log(\bar{d}_n\varsigma_n^{1/2})}{n}}+n^{\frac{1}{q}}\frac{s_2\varsigma_n\log^{}(\bar{d}_n\varsigma_n^{1/2})}{n}\right)=o(1).$$
Next, we want to bound the restricted eigenvalues of $\hat{\Sigma}_{\varepsilon\nu}$ with high probability by showing
\begin{align}
\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}|v^T\big(\hat{\Sigma}_{\varepsilon\nu}-\Sigma_{\varepsilon\nu}\big)v|\lesssim u_n
\end{align}
with
\begin{align*}
u_n&\lesssim \left(n^{\frac{3}{q}}\frac{t_1\log(d_1)}{\varsigma_nn}\right)^{1/2}\vee \left(\frac{t_1^2s\log(\bar{d}_n)}{\varsigma_nn}\right)^{1/2}\vee \frac{t_1^2s\log(\bar{d}_n)}{n}\\
& \lesssim \left(n^{\frac{3}{q}}\frac{t_1\log(d_1)}{\varsigma_nn}\right)^{1/2}\vee \left(\frac{t_1^2s\log(\bar{d}_n)}{\varsigma_nn}\right)^{1/2}= o(1)
\end{align*}
for a suitable $q\ge 4$ defined in Assumption A(ref).. Define $\xi_i:=\varepsilon_i\nu_i$, $\hat{\xi}_i:=\hat{\varepsilon}_i\hat{\nu}_i$ and observe that
\begin{align*}
&\hat{\Sigma}_{\varepsilon\nu}-\Sigma_{\varepsilon\nu}\\
=\ &\frac{1}{n}\sum\limits_{i=1}^n \hat{\xi}_i\hat{\xi}_i^T -\mathbb{E}[\xi_i\xi_i^T]\\
=\ &\frac{1}{n}\sum\limits_{i=1}^n \xi_i\xi_i^T -\mathbb{E}[\xi_i\xi_i^T]\\
&\ +\frac{1}{n}\sum\limits_{i=1}^n \xi_i\big(\hat{\xi}_i-\xi_i\big)^T+\frac{1}{n}\sum\limits_{i=1}^n \big(\hat{\xi}_i-\xi_i\big)\xi_i^T+\frac{1}{n}\sum\limits_{i=1}^n \big(\hat{\xi}_i-\xi_i\big)\big(\hat{\xi}_i-\xi_i\big)^T.
\end{align*}
Using the Lemma Q.1 from belloni2018uniformly, we can bound the first part.\\
Due to the tail conditions on $\varepsilon$ and boundedness of $\nu$, we obtain
\begin{align*}
\left(\mathbb{E}\left[\max_{1\le i\le n}\|\varepsilon_i\nu_i\|_\infty^2\right]\right)^{1/2}&\lesssim \mathbb{E}\left[\max_{1\le i\le n}|\varepsilon_i|^2\right]^{1/2}
\\
&\lesssim n^{1/q}
\end{align*}
for $q$ defined in Assumption A.(ref)(iii). Then, Lemma Q.1 implies
\begin{align*}
&\ \mathbb{E}\left[\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\Big|v^T\Big(\frac{1}{n}\sum\limits_{i=1}^n\xi_i\xi_i^T-\mathbb{E}[\xi_i\xi_i^T]\Big)v\Big|\right]\\
=&\ \mathbb{E}\left[\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\Big|\mathbb{E}_n\Big[\big(v^T\xi_i\big)^2-\mathbb{E}\big[\big(v^T\xi_i\big)^2\big]\Big]\Big|\right]\\
\lesssim&\ \tilde{\delta}_n^2+\tilde{\delta}_n \varsigma_n^{-1/2}\lesssim \left(n^{\frac{3}{q}}\frac{t_1\log(d_1)}{\varsigma_nn}\right)^{1/2}
\end{align*}
where we used
\begin{align*}
\tilde{\delta}_n &\lesssim \left(n^{\frac{2}{q}-1}t_1\log^2(t_1)\log(d_1)\log(n)\right)^{\frac{1}{2}}\\
&\lesssim \left(n^{\frac{3}{q}}\frac{t_1\log^(d_1)}{n}\right)^{\frac{1}{2}}\lesssim \varsigma_n^{-1/2}
\end{align*}
by the growth conditions in A.(ref)$(v)$.
Using Markov's inequality, we directly obtain
\begin{align*}
\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\Big|v^T\Big(\frac{1}{n}\sum\limits_{i=1}^n\xi_i\xi_i^T-\mathbb{E}[\xi_i\xi_i^T]\Big)v\Big|\lesssim u_n
\end{align*}
with probability $1-o(1)$. Note that by applying the results on covariance estimation from chen2012masked instead would lead to comparable growth rates.\\
With probability $1-o(1)$, it holds
\begin{align*}
\sup_{l=1,\dots,d_1} |\hat{\theta}_{l}-\theta_{0,l}|\lesssim \varsigma_n^{1/2}\tau_n
\end{align*}
analog to step 1 in the proof of Theorem 2.1 in belloni2018uniformly. Define
\begin{align*}
\tilde{\mathcal{G}}^2_2:=\big\{(\psi_l(\cdot,\theta_l,\eta_l)-\psi_l(\cdot,\theta_{0,l},\eta_{0,l}))^2:\ & l=1,\dots,d_1,|\theta_l-\theta_{0,l}|\le C\varsigma_n^{1/2}\tau_n,\\
& \eta_l\in\mathcal{T}_l\setminus\{\eta_{0,l}\}\big\},
\end{align*}
with
\begin{align*}
\sup_Q\log N(\varepsilon\|\tilde{G}_2^2\|_{Q,2},\tilde{\mathcal{G}}_2^2,\|\cdot\|_{Q,2})\lesssim s\log\left(\frac{d_1+d_2}{\varepsilon}\right).
\end{align*}
Here, $\tilde{G}_2^2$ is a measurable envelope of $\tilde{\mathcal{G}}_2^2$ with
\begin{align*}
\tilde{G}_2^2=\sup_{l=1,\dots,d_1}\sup_{\theta_l:|\theta_l-\theta_{0,l}|\le C\varsigma_n^{1/2}\tau_n,\eta_l\in\mathcal{T}_l}\big(\psi_l(W,\theta_l,\eta_l)-\psi_l(W,\theta_{0,l},\eta_{0,l})\big)^2
\end{align*}
and
\begin{align*}
&\|\tilde{G}_2^2\|_{P,\tilde{q}}\\
\lesssim&\ \Big\|\sup_{l,\theta_l,\eta_l^{(2)}}\Big((\theta_{0,l}-\theta_l)g_l(X_1)\big(g_l(X_1)-(\eta_l^{(2)})^TZ_{-l}\big)\Big)^2\Big\|_{P,\tilde{q}}\\
&\ +\Big\|\sup_{l,\eta_l}\Big(\big(Y-\theta_{0,l}g_l(X_1)-(\eta_l^{(1)})^TZ_{-l}-\eta_l^{(3)}(X)\big)\\
&\quad\quad\big((\eta_{0,l}^{(2)}-\eta_l^{(2)})^TZ_{-l}\big)\Big)^2\Big\|_{P,\tilde{q}}\\
&\ +\Big\|\sup_{l,\eta_l^{(1)},\eta_l^{(3)}}\Big(\big(g_l(X_1)-(\eta_{0,l}^{(2)})^TZ_{-l}\big)\\
&\quad\quad\big((\eta_{0,l}^{(1)}-\eta_l^{(1)})^TZ_{-l}+\eta_{0,l}^{(3)}(X)-\eta_{l}^{(3)}(X)\big)\Big)^2\Big\|_{P,\tilde{q}}\\
&=:T_1+T_2+T_3.
\end{align*}
For any $\tilde{q}\le q/2$ in Assumption A.(ref), it holds
\begin{align*}
T_1&\lesssim \tau_n^2\varsigma_n^\Big\|\sup_{l,\eta_l^{(2)}}\Big(g_l(X_1)\big(g_l(X_1)-(\eta_l^{(2)})^TZ_{-l}\big)\Big)^2\Big\|_{P,\tilde{q}}\\
&\lesssim \tau_n^2\varsigma_n^\Big\|\sup_{l,\eta_l^{(2)}}\Big(g_l(X_1)-(\eta_l^{(2)})^TZ_{-l}\Big)^2\Big\|_{P,\tilde{q}}\\
&\lesssim \tau_n^2\varsigma_n^,
\end{align*}
and
\begin{align*}
T_2&\lesssim s_2\tau_n^2
\end{align*}
where we used that $\|\varepsilon\|_{P,q}\le C$, $\|Z\|_\infty\le C$ (a.s.) and that $\eta^{(2)}_l\in\mathcal{T}_l$. Further, we have
\begin{align*}
T_3&\lesssim \Big\|\sup_{l,\eta_l^{(1)},\eta_l^{(3)}}\Big((\eta_{0,l}^{(1)}-\eta_l^{(1)})^TZ_{-l}+\eta_{0,l}^{(3)}(X)\Big)^2\Big\|_{P,q}\\
&\lesssim s_1\tau_n^2 + \tau_n^2
\end{align*}
due to Assumption A.(ref)$(ii)$. By using an analogous argument as above, we obtain
\begin{align*}
\tilde{\sigma}:&=\sup\limits_{f\in\tilde{\mathcal{G}}_2^2}\mathbb{E}\left[f(X)^2\right]^\frac{1}{2}\\
&=\sup\limits_{l=1,\dots,d_1}\sup\limits_{\theta_l:|\theta_l-\theta_{0,l}|\le C\varsigma^{1/2}\tau_n,\eta_l\in\mathcal{T}_l}\mathbb{E}\left[\left(\psi_l(W,\theta_l,\eta_l)-\psi_l(W,\theta_{0,l},\eta_{0,l})\right)^4\right]^{\frac{1}{2}}\\
&\lesssim \tau_n^2(s\vee\varsigma_n).
\end{align*}
Again, we can apply Lemma P.2 from belloni2018uniformly to obtain
\begin{align*}
\sup\limits_{f\in\tilde{\mathcal{G}}_2^2}|\mathbb{E}_n[f(X)]-\mathbb{E}[f(X)]|\le K&\bigg(\tilde{\sigma}\sqrt{\frac{s\log(\bar{d}_n)}{n}}+n^{\frac{1}{q}}\|\tilde{G}_2^2\|_{P,\tilde{q}}\frac{s\log(\bar{d}_n)}{n}\bigg)\\
&\lesssim (s\vee \varsigma_n)\tau_n^3\vee n^{\frac{1}{\tilde{q}}}(s\vee \varsigma_n)\tau_n^4
\end{align*}
with probability $1-o(1)$. Note that we have already shown Assumption B.(ref)$(v)(a)$ which implies
\begin{align*}
\sup\limits_{f\in\tilde{\mathcal{G}}_2^2} \mathbb{E}[f(X)]&\le C\left(|\theta_l-\theta_{0,l}|^2\varsigma_n^{-1}\vee \|\eta_{0,l}-\eta_l\|_e^2\right)\\
&\lesssim \tau_n^2.
\end{align*}
Combined, this implies
\begin{align}
\sup_{l=1,\dots,d_1}\mathbb{E}_n\left[\left(\hat{\varepsilon}_i\hat{\nu}^{(l)}_i-\varepsilon_i\nu_i^{(l)}\right)^2\right]&\le \sup\limits_{f\in\tilde{\mathcal{G}}_2^2} \mathbb{E}_n[f(X)] \nonumber\\
&\lesssim \tau_n^2+\left((s\vee \varsigma_n)\tau_n^3\vee n^{\frac{1}{\tilde{q}}}(s\vee \varsigma_n)\tau_n^4\right)\lesssim \tau_n^2
\end{align}
by growth the growths assumptions in A.(ref) (v) and, with an analogous argument, we obtain
\begin{align*}
\sup_{l=1,\dots,d_1}\mathbb{E}_n\left[\left(\varepsilon_i\nu_i^{(l)}\right)^2\right]\lesssim \varsigma_n^{-1}.
\end{align*}
Therefore, it holds
\begin{align*}
& \sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}|v^T\frac{1}{n}\sum\limits_{i=1}^n \xi_i\big(\hat{\xi}_i-\xi_i\big)^Tv|\\
=\ &\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}|\mathbb{E}_n\left[v^T \xi_i\big(\hat{\xi}_i-\xi_i\big)^Tv\right]|\\
\le\ & \sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\left|\left(\mathbb{E}_n\left[\left(v^T \xi_i\right)^2\right]\mathbb{E}_n\left[\left(v^T\big(\hat{\xi}_i-\xi_i\big)\right)^2\right]\right)^{\frac{1}{2}}\right|\\
\lesssim\ &\varsigma_n^{-1/2} \sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\left|\left(\mathbb{E}_n\left[\left(v^T\big(\hat{\xi}_i-\xi_i\big)\right)^2\right]\right)^{\frac{1}{2}}\right|\\
=\ & \varsigma_n^{-1/2}\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\left(\sum_{k=1}^{d_1}\sum_{l=1}^{d_1} v_kv_l\mathbb{E}_n\left[(\hat{\varepsilon}_i\hat{\nu}_i^{(k)}-\varepsilon_i\nu_i^{(k)})(\hat{\varepsilon}_i\hat{\nu}_i^{(l)}-\varepsilon_i\nu_i^{(l)})\right]\right)^{\frac{1}{2}}\\
\lesssim\ &\varsigma_n^{-1/2}t_1\sup_{l=1,\dots,d_1}\mathbb{E}_n\left[(\hat{\varepsilon}_i\hat{\nu}_i^{(l)}-\varepsilon_i\nu_i^{(l)})^2\right]^{\frac{1}{2}}\\
\lesssim\ & t_1\varsigma_n^{-1/2}\tau_n
\end{align*}
and
\begin{align*}
\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}|v^T\frac{1}{n}\sum\limits_{i=1}^n \big(\hat{\xi}_i-\xi_i\big)\big(\hat{\xi}_i-\xi_i\big)^Tv|\lesssim t_1^2\tau_n^2
\end{align*}
with probability $1-o(1)$. Combining the steps above, implies ((ref)) if $u_n=o(1)$ which is ensured by the growth conditions. Next, note that for every sparse vector $w\in \mathbb{R}^{d_1}$ ($\|w\|_0\le t_1$) there exists a corresponding matrix $M_w$
$$M_w\in \mathbb{R}^{d_1\times d_1}: (M_w)_{k,l}=\begin{cases}1 \text{ if } w_k\neq 0 \wedge w_l\neq 0\\
0 \text{ else,} \end{cases}$$
such that
\begin{align*}
w^T(\hat{\Sigma}_{\varepsilon\nu}-\Sigma_{\varepsilon\nu})w=w^T\left(M_w \odot(\Sigma_{\varepsilon\nu}-\hat{\Sigma}_{\varepsilon\nu})\right)w.
\end{align*}
Due to ((ref)), it holds
\begin{align*}
\sup_{\|w\|_0\le t_1}\sup_{\|v\|_2=1}\left|v^T\left(M_w \odot(\hat{\Sigma}_{\varepsilon\nu}-\Sigma_{\varepsilon\nu})\right)v\right|&\le \sup_{\|v\|_2=1,\|v\|_0\le t_1}\left|v^T(\hat{\Sigma}_{\varepsilon\nu}-\Sigma_{\varepsilon\nu})v\right|\lesssim u_n,
\end{align*}
which implies
\begin{align*}
\sup_{\|w\|_0\le t_1}\|M_w \odot(\hat{\Sigma}_{\varepsilon\nu}-\Sigma_{\varepsilon\nu})\|_2\lesssim u_n
\end{align*}
and
\begin{align*}
\sup_{\|w\|_0\le t_1}\|M_w \odot\hat{\Sigma}_{\varepsilon\nu}\|_2\lesssim \varsigma_n^{-1}
\end{align*}
due to Assumption A.(ref)$(iv)$. This can be used to show that
\begin{align}
\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}|v^T\big(\hat{\Sigma}_n-\Sigma_n\big)v|\lesssim \varsigma_n\left(\tilde{u}_n\vee\tilde{\epsilon}_n\right)
\end{align}
where $\tilde{u}_n:=\varsigma_n u_n=o(1)$
with probability $1-o(1)$ which can be interpreted as an upper bound for the sparse eigenvalues of $\hat{\Sigma}_n-\Sigma_n$. It holds
\begin{align*}
\hat{\Sigma}_n-\Sigma_n&= \hat{J}^{-1}\hat{\Sigma}_{\varepsilon\nu}(\hat{J}^{-1})^T-J_0^{-1}\Sigma_{\varepsilon\nu}(J_0^{-1})^T\\
&=(\hat{J}^{-2}-J_0^{-2})\hat{\Sigma}_{\varepsilon\nu}+J_0^{-2}\big(\hat{\Sigma}_{\varepsilon\nu}-\Sigma_{\varepsilon\nu}\big).
\end{align*}
Note that
\begin{align*}
&\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}|v^T\big(\hat{J}^{-2}-J_0^{-2})\hat{\Sigma}_{\varepsilon\nu}v|\\
=\ &\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}|v^T\big(\hat{J}^{-2}-J_0^{-2})\left(M_v\odot\hat{\Sigma}_{\varepsilon\nu}\right)v|\\
\le\ &\sup\limits_{\|w\|_0\le t_1}\left\|\left(M_w\odot\hat{\Sigma}_{\varepsilon\nu}\right)\right\|_2\left\|\big(\hat{J}^{-2}-J_0^{-2})\right\|_2\\
\lesssim\ & \varsigma_n^{-1}\left\|\big(\hat{J}^{-2}-J_0^{-2})\right\|_2\\
\lesssim & \varsigma_n\tilde{\epsilon}_n
\end{align*}
due to the sub-multiplicative spectral norm, $|J_{0,l}^{-1}|=O(\varsigma_n)$ and (ref) which implies
\begin{align}
\left\|\hat{J}^{-2}-J_0^{-2}\right\|&\le \left\|\hat{J}^{-1}(\hat{J}^{-1}-J_0^{-1})\right\|+\left\|J_0^{-1}(\hat{J}^{-1}-J_0^{-1})\right\|\nonumber \\
&\lesssim\varsigma_n \left\|\hat{J}^{-1}-J_0^{-1}\right\|\lesssim\varsigma_n^3 \left\|\hat{J}-J_0\right\|=\varsigma_n^2\tilde{\epsilon}_n
\end{align}
with $\left\|\hat{J}-J_0\right\|=\sup_{l=1,\dots,d_1}|\hat{J}_l-J_{0,l}|=\varsigma_n^{-1}\tilde{\epsilon}_n$. The second term can be bounded by
\begin{align*}
&\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}|v^TJ_0^{-2}\big(\hat{\Sigma}_{\varepsilon\nu}-\Sigma_{\varepsilon\nu})v|\lesssim \varsigma_n^2 u_n=\varsigma_n\tilde{u}_n.
\end{align*}
This implies ((ref)). We finally obtain
\begin{align*}
\sup_{x\in I}\left|\frac{(g(x)^T\hat{\Sigma}_n g(x))^{1/2}}{(g(x)^T\Sigma_n g(x))^{1/2}}-1\right|&\le \sup_{x\in I}\left|\frac{(g(x)^T\hat{\Sigma}_n g(x))}{(g(x)^T\Sigma_n g(x))}-1\right|\\
&\lesssim \varsigma_n^{-1}t_1\sup_{x\in I}\left|g(x)^T\big(\hat{\Sigma}_n-\Sigma_n\big) g(x)\right|\\
&\le \varsigma_n^{-1}t_1 \sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}|v^T\big(\hat{\Sigma}_n-\Sigma_n\big)v|\\
&=t_1(\tilde{u}_n\vee\tilde{\epsilon}_n)\lesssim t_1\tilde{u}_n
\end{align*}
with probability $1-o(1)$ which is the first part of Assumption B.(ref) with $\epsilon_n\lesssim t_1 \tilde{u}_n$ and $\epsilon_n t_1 \log(\tilde{A}_n\varsigma_n^{-1/2})=o(1)$. \\ \\
Assumption B.(ref)$(iii)-(iv)$ \\
Define
\begin{align*}
\sigma_x:&=(g(x)^T\Sigma_n g(x))^{1/2},\\
\hat{\sigma}_x:&=(g(x)^T\hat{\Sigma}_n g(x))^{1/2}
\end{align*}
and
$$\hat{\mathcal{F}}_0:=\{\psi_x(\cdot)-\hat{\psi}_x(\cdot):x\in I\}$$
with $\hat{\psi}_x(\cdot):=\hat{\sigma}_x^{-1}g(x)^T\hat{J}_0^{-1}\psi(\cdot,\hat{\theta}_{},\hat{\eta}_{})$. For every $x$ and $\tilde{x}$, it holds
\begin{align*}
&\|\psi_x(W)-\hat{\psi}_x(W)-(\psi_{\tilde{x}}(W)-\hat{\psi}_{\tilde{x}}(W))\|_{\mathbb{P}_n,2}\\
= &\Big\|\sigma_x^{-1}g(x)^TJ_0^{-1}\psi(W,\theta_{0},\eta_{0})-\sigma_{\tilde{x}}^{-1}g({\tilde{x}})^TJ_0^{-1}\psi(W,\theta_{0},\eta_{0}) \\
& -\big(\hat{\sigma}_x^{-1}g(x)^T\hat{J}^{-1}\psi(W,\hat{\theta}_,\hat{\eta}_)-\hat{\sigma}_{\tilde{x}}^{-1}g({\tilde{x}})^T\hat{J}^{-1}\psi(W,\hat{\theta}_,\hat{\eta}_) \big)\Big\|_{\mathbb{P}_n,2}\\
= &\Big\|\sum_{l=1}^{d_1} (\sigma_x^{-1} g_l(x)-\sigma_{\tilde{x}}^{-1}g_l(\tilde{x}))J_{0,l}^{-1}\psi_l(W,\theta_{0,l},\eta_{0,l}) \\
& -\sum_{l=1}^{d_1} (\hat{\sigma}_x^{-1} g_l(x)-\hat{\sigma}_{\tilde{x}}^{-1}g_l(\tilde{x}))\hat{J}_l^{-1}\psi_l(W,\hat{\theta}_{l},\hat{\eta}_{l})\Big\|_{\mathbb{P}_n,2}\\
\le &\Big\|\sum_{l=1}^{d_1} (\sigma_x^{-1} g_l(x)-\sigma_{\tilde{x}}^{-1}g_l(\tilde{x}))\Big(J_{0,l}^{-1}-\hat{J}_{l}^{-1}\Big)\psi_l(W,\theta_{0,l},\eta_{0,l})\Big\|_{\mathbb{P}_n,2} \\
&+\Big\|\sum_{l=1}^{d_1} (\sigma_x^{-1} g_l(x)-\sigma_{\tilde{x}}^{-1}g_l(\tilde{x}))\hat{J}_{l}^{-1}\Big(\psi_l(W,\theta_{0,l},\eta_{0,l})-\psi_l(W,\hat{\theta}_{l},\hat{\eta}_{l})\Big)\Big\|_{\mathbb{P}_n,2}\\
& +\Big\|\sum_{l=1}^{d_1}\Big( (\sigma_x^{-1} g_l(x)-\sigma_{\tilde{x}}^{-1}g_l(\tilde{x}))- (\hat{\sigma}_x^{-1} g_l(x)-\hat{\sigma}_{\tilde{x}}^{-1}g_l(\tilde{x}))\Big)\hat{J}_{l}^{-1}\psi_l(W,\hat{\theta}_{l},\hat{\eta}_{l})\Big\|_{\mathbb{P}_n,2}\\
=:& I_{4,1}+I_{4,2}+I_{4,3}.
\end{align*}
We obtain
\begin{align*}
I_{4,1}&=\Big\|\sum_{l=1}^{d_1} (\sigma_x^{-1} g_l(x)-\sigma_{\tilde{x}}^{-1}g_l(\tilde{x}))\Big(J_{0,l}^{-1}-\hat{J}_{l}^{-1}\Big)\psi_l(W,\theta_{0,l},\eta_{0,l})\Big\|_{\mathbb{P}_n,2} \\
&\le\sigma_x^{-1}\Big\|(g(x)-g(\tilde{x}))^T\Big(J_{0}^{-1}-\hat{J}_^{-1}\Big)\psi(W,\theta_{0},\eta_{0})\Big\|_{\mathbb{P}_n,2}\\
&\quad+|\sigma_x^{-1}-\sigma_{\tilde{x}}^{-1}|\Big\|g(\tilde{x})^T\Big(J_{0}^{-1}-\hat{J}_^{-1}\Big)\psi(W,\theta_{0},\eta_{0})\Big\|_{\mathbb{P}_n,2}\\
&\lesssim \varsigma_n^{-1/2}\sqrt{t_1} \|g(x)-g(\tilde{x})\|_2\sup\limits_{\|v\|_2=1,\|v\|_0\le 2t_1}\Big\|v^T\Big(J_{0}^{-1}-\hat{J}_^{-1}\Big)\psi(W,\theta_{0},\eta_{0})\Big\|_{\mathbb{P}_n,2}\\
&\quad + \varsigma_n^{-1/2}t_1^{3/2}\|g(x)-g(\tilde{x})\|_2\sup\limits_{x\in I}\|g(x)\|_2^{2}\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\Big\|v^T\Big(J_{0}^{-1}-\hat{J}_^{-1}\Big)\psi(W,\theta_{0},\eta_{0})\Big\|_{\mathbb{P}_n,2}\\
&\lesssim \varsigma_n^{-1/2}t_1^{3/2}\|g(x)-g(\tilde{x})\|_2\varsigma_n^{1/2}\tilde{\epsilon}_n\lesssim t_1^{3/2}\tilde{\epsilon}_n\|g(x)-g(\tilde{x})\|_2,
\end{align*}
with $t_1^{3/2}\tilde{\epsilon}_n=o(1)$ where we used that
\begin{align*}
&\ \sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\Big\|v^T\Big(J_{0}^{-1}-\hat{J}_^{-1}\Big)\psi(W,\theta_{0},\eta_{0})\Big\|^2_{\mathbb{P}_n,2}\\
=&\ \sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\Big|v^T\Big(J_{0}^{-1}-\hat{J}_^{-1}\Big)\frac{1}{n}\sum\limits_{i=1}^n \xi_i\xi_i^T \Big(J_{0}^{-1}-\hat{J}_^{-1}\Big)^Tv\Big|\\
\le&\ \left\|J_{0}^{-1}-\hat{J}_^{-1}\right\|^2_2\sup\limits_{\|v\|_0\le t_1}\Big\|M_v\odot\left(\frac{1}{n}\sum\limits_{i=1}^n \xi_i\xi_i^T\right)\Big\|_2\\
\lesssim& (\varsigma_n\tilde{\epsilon}_n)^2\varsigma_n^{-1}\lesssim \varsigma_n\tilde{\epsilon}_n^2
\end{align*}
by (ref). Analogously, we obtain
\begin{align*}
I_{4,2}&=\Big\|\sum_{l=1}^{d_1} (\sigma_x^{-1} g_l(x)-\sigma_{\tilde{x}}^{-1}g_l(\tilde{x}))\hat{J}_{l}^{-1}\Big(\psi_l(W,\theta_{0,l},\eta_{0,l})-\psi_l(W,\hat{\theta}_{l},\hat{\eta}_{l})\Big)\Big\|_{\mathbb{P}_n,2}\\
&\le\sigma_x^{-1}\Big\|(g(x)-g(\tilde{x}))^T\hat{J}_^{-1}\Big(\psi(W,\theta_{0},\eta_{0})-\psi(W,\hat{\theta}_,\hat{\eta}_)\Big)\Big\|_{\mathbb{P}_n,2}\\
&\quad+|\sigma_x^{-1}-\sigma_{\tilde{x}}^{-1}|\Big\|g(\tilde{x})^T\hat{J}_^{-1}\Big(\psi(W,\theta_{0},\eta_{0})-\psi(W,\hat{\theta}_,\hat{\eta}_)\Big)\Big\|_{\mathbb{P}_n,2}\\
&\lesssim \varsigma_n^{-1/2}\sqrt{t_1}\|g(x)-g(\tilde{x})\|_2\sup\limits_{\|v\|_2=1,\|v\|_0\le 2t_1}\Big\|v^T\hat{J}_^{-1}\Big(\psi(W,\theta_{0},\eta_{0})-\psi(W,\hat{\theta}_,\hat{\eta}_)\Big)\Big\|_{\mathbb{P}_n,2}\\
&\quad + \varsigma_n^{-1/2}t_1^{3/2}\|g(x)-g(\tilde{x})\|_2\sup\limits_{x\in I}\|g(x)\|_2^2\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\Big\|v^T\hat{J}_^{-1}\Big(\psi(W,\theta_{0},\eta_{0})-\psi(W,\hat{\theta}_,\hat{\eta}_)\Big)\Big\|_{\mathbb{P}_n,2}\\
&\lesssim \varsigma_n^{-1/2}t_1^{3/2}\|g(x)-g(\tilde{x})\|_2\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\Big\|v^T\hat{J}_^{-1}\Big(\psi(W,\theta_{0},\eta_{0})-\psi(W,\hat{\theta}_,\hat{\eta}_)\Big)\Big\|_{\mathbb{P}_n,2}\\
&\lesssim \varsigma_n^{-1/2}t_1^{3/2} \|g(x)-g(\tilde{x})\|_2 \varsigma_n t_1\tau_n\\
&\lesssim \sqrt{\frac{t_1^5\varsigma_n s\log(\bar{d}_n)}{n}}\|g(x)-g(\tilde{x})\|_2=o(\|g(x)-g(\tilde{x})\|_2),
\end{align*}
by growth conditions where we used that
$$\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\Big\|v^T\hat{J}_{}^{-1}\Big(\psi(W,\theta_{0},\eta_{0})-\psi(W,\hat{\theta}_{},\hat{\eta}_{})\Big)\Big\|_{\mathbb{P}_n,2}\lesssim \varsigma_n t_1\tau_n$$
as shown in (ref). Further, it holds
\begin{align*}
I_{4,3}&= \Big\|\sum_{l=1}^{d_1}\Big( (\sigma_x^{-1} g_l(x)-\sigma_{\tilde{x}}^{-1}g_l(\tilde{x}))- (\hat{\sigma}_x^{-1} g_l(x)-\hat{\sigma}_{\tilde{x}}^{-1}g_l(\tilde{x}))\Big)\hat{J}_{l}^{-1}\psi_l(W,\hat{\theta}_{l},\hat{\eta}_{l})\Big\|_{\mathbb{P}_n,2}\\
&\le \big| \sigma_x^{-1}-\hat{\sigma}_x^{-1} \big|\Big\|(g(x)-g(\tilde{x}))^T\hat{J}_^{-1}\psi(W,\hat{\theta}_,\hat{\eta}_)\Big\|_{\mathbb{P}_n,2}\\
&\quad +\big| (\sigma_x^{-1}-\hat{\sigma}_x^{-1})-(\sigma_{\tilde{x}}^{-1}-\hat{\sigma}_{\tilde{x}}^{-1})\big|\Big\|g(\tilde{x})^T\hat{J}_^{-1}\psi(W,\hat{\theta}_,\hat{\eta}_)\Big\|_{\mathbb{P}_n,2}.
\end{align*}
Note that
\begin{align*}
&\big| (\sigma_x^{-1}-\hat{\sigma}_x^{-1})-(\sigma_{\tilde{x}}^{-1}-\hat{\sigma}_{\tilde{x}}^{-1})\big|\\
=\ &\Big| \frac{1}{\sigma_x \sigma_{\tilde{x}}}(\sigma_{\tilde{x}}-\sigma_x)-\frac{1}{\hat{\sigma}_x \hat{\sigma}_{\tilde{x}}}(\hat{\sigma}_{\tilde{x}}-\hat{\sigma}_{x})\Big|\\
= \ &\frac{1}{\hat{\sigma}_x \hat{\sigma}_{\tilde{x}}}\Big| \frac{\hat{\sigma}_x \hat{\sigma}_{\tilde{x}}}{\sigma_x \sigma_{\tilde{x}}}(\sigma_{\tilde{x}}-\sigma_x)-(\hat{\sigma}_{\tilde{x}}-\hat{\sigma}_{x})\Big|\\
\lesssim\ & t_1\varsigma_n^{-1}\left(\big| (\sigma_{\tilde{x}}-\sigma_x)-(\hat{\sigma}_{\tilde{x}}-\hat{\sigma}_{x})\big|+\Big|\frac{\hat{\sigma}_x \hat{\sigma}_{\tilde{x}}}{\sigma_x \sigma_{\tilde{x}}}-1\Big|\big|\sigma_{\tilde{x}}-\sigma_x\big|\right)
\end{align*}
with
\begin{align*}
\Big|\frac{\hat{\sigma}_x \hat{\sigma}_{\tilde{x}}}{\sigma_x \sigma_{\tilde{x}}}-1\Big|\big|\sigma_{\tilde{x}}-\sigma_x\big|
&\le\Big(\Big|\frac{\hat{\sigma}_x }{\sigma_x }-1\Big|\frac{ \hat{\sigma}_{\tilde{x}}}{ \sigma_{\tilde{x}}}+\Big|\frac{\hat{\sigma}_{\tilde{x} }}{\sigma_{\tilde{x}} }-1\Big|\Big)\big|\sigma_{\tilde{x}}-\sigma_x\big|\\
&\lesssim\epsilon_n\frac{1}{\sigma_x}\big|\sigma^2_{\tilde{x}}-\sigma^2_x\big|\\
&\lesssim\epsilon_n\sqrt{t_1}\varsigma_n^{1/2}\|g(x)-g(\tilde{x})\|_2\sup_x\|g(x)\|_2
\end{align*}
uniformly over $x\in I$ with probability $1-o(1)$ and
\begin{align*}
&\big| (\sigma_{\tilde{x}}-\sigma_x)-(\hat{\sigma}_{\tilde{x}}-\hat{\sigma}_{x})\big|\\
\le\ &\frac{1}{(\hat{\sigma}_{\tilde{x}}+\hat{\sigma}_{x})} \big| (\sigma_{\tilde{x}}^2-\sigma_x^2)-(\hat{\sigma}_{\tilde{x}}^2-\hat{\sigma}_{x}^2)\big|+\Big|\left(\frac{1}{(\sigma_{\tilde{x}}+\sigma_{x})}-\frac{1}{(\hat{\sigma}_{\tilde{x}}+\hat{\sigma}_{x})}\right)(\sigma_{\tilde{x}}^2-\sigma_x^2)\Big|\\
\lesssim\ & \sqrt{t_1}\varsigma_n^{-1/2}\left(\big|(\sigma_{\tilde{x}}^2-\sigma_x^2)-(\hat{\sigma}_{\tilde{x}}^2-\hat{\sigma}_{x}^2)\big|+\Big|\frac{(\hat{\sigma}_{\tilde{x}}+\hat{\sigma}_{x})}{(\sigma_{\tilde{x}}+\sigma_x)}-1\Big| \big|\sigma^2_{\tilde{x}}-\sigma^2_x\big|\right).
\end{align*}
Using an analogous argument as in the verification of Assumption B.(ref), we obtain
\begin{align*}
|(\sigma_x^{2}-\hat{\sigma}_x^{2})-(\sigma_{\tilde{x}}^{2}-\hat{\sigma}_{\tilde{x}}^{2})|&=|(g(x)-g(\tilde{x}))^T(\Sigma_n-\hat{\Sigma}_n) (g(x)+g(\tilde{x}))|\\
&\le \|(\Sigma_n-\hat{\Sigma}_n)(g(x)-g(\tilde{x}))\|_2\sup_{x\in I}\|g(x)\|_2\\
&\lesssim \varsigma_n\tilde{u}_n\|g(x)-g(\tilde{x})\|_2
\end{align*}
with probability $1-o(1)$ where the last inequality holds due the order of the sparse eigenvalues in ((ref)). Additionally,
\begin{align*}
\Big|\frac{(\hat{\sigma}_{\tilde{x}}+\hat{\sigma}_{x})}{(\sigma_{\tilde{x}}+\sigma_x)}-1\Big| \big|\sigma^2_{\tilde{x}}-\sigma^2_x\big|&\le \sup_{x\in I} \Big|\frac{\hat{\sigma}_x }{\sigma_x }-1\Big| \big|\sigma^2_{\tilde{x}}-\sigma^2_x\big|\\
&\lesssim\epsilon_n\varsigma_n\|g(x)-g(\tilde{x})\|_2
\end{align*}
with probability $1-o(1)$. Therefore, we obtain
\begin{align*}
I_{4,3}&\lesssim \sqrt{t_1}\varsigma_n^{-1/2}\epsilon_n\|g(x)-g(\tilde{x})\|_2\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\Big\|v^T\hat{J}_^{-1}\psi(W,\hat{\theta}_,\hat{\eta}_)\Big\|_{\mathbb{P}_n,2}\\
&\quad +t_1^{3/2}\varsigma_n^{-1}\left(\epsilon_n\varsigma_n^{1/2}\vee \varsigma_n^{1/2}\tilde{u}_n\right)\|g(x)-g(\tilde{x})\|_2\sup\limits_{\|v\|_2=1,\|v\|_0\le t_1}\Big\|v^T\hat{J}_^{-1}\psi(W,\hat{\theta}_,\hat{\eta}_)\Big\|_{\mathbb{P}_n,2}\\
&\lesssim t_1^{3/2}\left(\epsilon_n\vee \tilde{u}_n\right) \|g(x)-g(\tilde{x})\|_2.
\end{align*}
Combining the steps above, we obtain
\begin{align*}
\|\psi_x(W)-\hat{\psi}_x(W)-(\psi_{\tilde{x}}(W)-\hat{\psi}_{\tilde{x}}(W))\|_{\mathbb{P}_n,2}\le \|g(x)-g(\tilde{x})\|_2\|\hat{F}_0\|_{\mathbb{P}_n,2}
\end{align*}
with
\begin{align*}
\|\hat{F}_0\|_{\mathbb{P}_n,2}=o(1)
\end{align*}
due to the growth condition in Assumption A.(ref)$(v)(b)$. Using the same argument as Theorem 2.7.11 from vanweak, we obtain with probability $1-o(1)$
\begin{align*}
\log N(\varepsilon,\hat{\mathcal{F}}_0,\|\cdot\|_{\mathbb{P}_n,2})&\le\log N(\varepsilon\|\hat{F}_0\|_{\mathbb{P}_n,2},\hat{\mathcal{F}}_0,\|\cdot\|_{\mathbb{P}_n,2})\\
&\le \log N(\varepsilon,g(I),\|\cdot\|_{2})\\
&\le\bar{\varrho}_n\log\left(\frac{\bar{A}_n}{\varepsilon}\right)
\end{align*}
with $\bar{\varrho}_n=t_1$ and $\bar{A}_n\lesssim \tilde{A}_n$ by Assumption A.(ref)$(i)$. Additionally, it holds
\begin{align*}
&\|\psi_x(W)-\hat{\psi}_x(W)\|_{\mathbb{P}_n,2}\\
=\ &\Big\|\sigma_x^{-1}g(x)^TJ_0^{-1}\psi(W,\theta_{0},\eta_{0})-\hat{\sigma}_x^{-1}g(x)^T\hat{J}^{-1}\psi(W,\hat{\theta}_,\hat{\eta}_)\Big\|_{\mathbb{P}_n,2}\\
\le\ &\sigma_x^{-1}\Big\|g(x)^T\Big(J_0^{-1}-\hat{J}^{-1}\Big)\psi(W,\theta_{0},\eta_{0})\Big\|_{\mathbb{P}_n,2}\\
&+\sigma_x^{-1}\Big\|g(x)^T\hat{J}^{-1}\Big(\psi(W,\theta_{0},\eta_{0})-\psi(W,\hat{\theta}_,\hat{\eta}_)\Big)\Big\|_{\mathbb{P}_n,2}\\
&+|\sigma_x^{-1}-\hat{\sigma}_x^{-1}|\Big\|g(x)^T\hat{J}^{-1}\psi(W,\hat{\theta}_,\hat{\eta}_)\Big\|_{\mathbb{P}_n,2}\\
\lesssim\ &\sqrt{t_1}\tilde{\epsilon}_n+\sqrt{\frac{t_1^3\varsigma_ns\log(\bar{d}_n)}{n}}+\sqrt{t_1}\epsilon_n\\
\lesssim\ &\Bigg(\frac{t_1s_2\varsigma_n\log(\bar{d}_n\varsigma_n^{1/2})}{n}+\left(n^{\frac{1}{q}}\frac{t_1s_2\varsigma_n\log(\bar{d}_n\varsigma_n^{1/2})}{n}\right)^2+\frac{t_1^3\varsigma_ns\log(\bar{d}_n)}{n}\\
&+n^{\frac{3}{q}}\frac{t_1^4\varsigma_n\log(d_1)}{n}+\frac{t_1^5s\varsigma_n\log(\bar{d}_n)}{n}\Bigg)^{1/2}\\
\lesssim & \ \sqrt{\frac{t_1^5s\varsigma_n\log(\bar{d}_n)}{n}}
\end{align*}
for q large enough by (ref), (ref) and using analogous arguments as above. Therefore, B.(ref)$(iii)$ holds with
$$\bar{\delta}_n\lesssim\sqrt{\frac{t_1^5s\varsigma_n\log(\bar{d}_n)}{n}}$$
To complete the proof, we verify all growth conditions from Assumptions B.(ref) and B.(ref). As shown in the verification of B.(ref)$(vi)$, it holds
\begin{align*}
t_1^2\delta_n^2\varrho_n\log(A_n)=t_1^3n^{2/q-1}s^2\varsigma_n\log^2(\bar{d}_n) \log(\tilde{A}_n\varsigma_n^{-1/2})=o(1).
\end{align*}
Additionally,
\begin{align*}
n^{-\frac{1}{7}}L_n^{\frac{2}{7}}\varrho_n\log(A_n)=\left(\frac{t_1^{16}\varsigma_n^3\log^7(\tilde{A}_n\varsigma_n^{-1/2})}{n}\right)^{1/7}=o(1)
\end{align*}
and
\begin{align*}
n^{\frac{2}{3q}-\frac{1}{3}}L_n^{\frac{2}{3}}\varrho_n\log(A_n)=\left(n^{\frac{2}{q}}\frac{t_1^{12}\varsigma_n^3\log^3(\tilde{A}_n\varsigma_n^{-1/2})}{n}\right)^{1/3}=o(1)
\end{align*}
for $q$ large enough due to growth condition in Assumption A.(ref)$(v)(c)$. Note that
\begin{align*}
\varepsilon_n\varrho_n\log(A_n)&=\varepsilon_n t_1\log(\tilde{A}_n\varsigma_n^{-1/2})=t_1^2\tilde{u}_n\log(\tilde{A}_n\varsigma_n^{-1/2})\\
&=t_1^2\sqrt{\frac{t_1^2s\varsigma_n\log(\bar{d}_n)}{n}}\log(\tilde{A}_n\varsigma_n^{-1/2})\\
&=\sqrt{\frac{t_1^6s\varsigma_n\log(\bar{d}_n)\log^2(\tilde{A}_n\varsigma_n^{-1/2})}{n}}=o(1).
\end{align*}
Hence, to conclude we need to show that
\begin{align*}
\bar{\delta}^2_n\bar{\varrho}_n\varrho_n\log(\bar{A}_n)\log(A_n)=\bar{\delta}_n^2 t_1^2\log(\tilde{A}_n)\log(\tilde{A}_n\varsigma_n^{-1/2})=o(1).
\end{align*}
which holds true since
\begin{align*}
\bar{\delta}_n^2 t_1^2\log(\tilde{A}_n\varsigma_n^{-1/2})\log(\tilde{A}_n)
\lesssim \frac{t_1^7s\varsigma_n\log(\bar{d}_n)\log(\tilde{A}_n\varsigma_n^{-1/2})\log(\tilde{A}_n)}{n}
=o(1)
\end{align*}
due to Assumption A.(ref)$(v)(b)$.