EconBase
← Back to paper

Localized Neural Network Modelling of Time Series: A Case Study on US Monetary Policy

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

282,821 characters · 23 sections · 82 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

{7pt} {7pt}

titlepage\begin{center} { \bf Localized Neural Network Modelling of Time Series:\\ A Case Study on US Monetary Policy} { {\sc Jiti Gao$^\dag$, Fei Liu$^\sharp$, Bin Peng$^\dag$ and Yanrong Yang$^*$} $^\dag$Monash University, $^\sharp$Nankai University and $^*$The Australian National University} \today \begin{abstract} In this paper, we investigate a semiparametric regression model under the context of treatment effects via a localized neural network (LNN) approach. Due to a vast number of parameters involved, we reduce the number of effective parameters by (i) exploring the use of identification restrictions; and (ii) adopting a variable selection method based on the group-LASSO technique. Subsequently, we derive the corresponding estimation theory and propose a dependent wild bootstrap procedure to construct valid inferences accounting for the dependence of data. Finally, we validate our theoretical findings through extensive numerical studies. In an empirical study, we revisit the impacts of a tightening monetary policy action on a variety of economic variables, including short-/long-term interest rate, inflation, unemployment rate, industrial price and equity return via the newly proposed framework using a monthly dataset of the US. {\em Keywords}: Dependent Wild Bootstrap; Group-LASSO; Semiparametric Model; Treatment Effects {\em JEL classification}: C14, C22, C45 \end{abstract} \end{center}

Introduction

Neural network (NN) architecture has received increasing attention over the last several decades. On relevant topics, a large number of papers have been published in different journals such as Econometrica, The Annals of Statistics, and Journal of Machine Learning Research, etc. by experts from different disciplines. Apparently, we cannot exhaust the literature, but refer interested readers to BMR2021 and FMZ2021 for extensive reviews from a methodological point of view.

NN usually includes three ingredients: input layer, hidden layer(s), and output layer. We now briefly comment on them one by one. The input layer is possibly the easiest one to understand, as it includes regressors only. The hidden layer(s) involve lots of activation functions and parameters mapping linear combinations of the regressors to a certain range of the real line. Usually, sparsity has to be imposed to ensure a reasonable number of effective parameters and a small set of active activation functions (e.g., SH2020, WL2021). Once the parameters are estimated, one can load the test dataset to evaluate the performance of NN. Finally, the output layer receives the outcome, of which there are two notable types (i.e., quantitative and qualitative). Against this background, a semiparametric regression model under the context of treatment effects (e.g., BCH2014) naturally includes these two types of output in one framework, so it offers a nice structure to start the following semiparametric regression model:

eqnarray[eqnarray omitted — 71 chars of source]

where $\mathbf{x} =(x_1,\ldots, x_d)^\top$ is a $d\times 1$ vector of control variables, $z$ is a treatment/policy variable subject to the influence of $\mathbf{x}$ via the structure $z= I(G(\mathbf{x} )-\eta\ge 0 )$ with $I(\cdot)$ being the indicator function, both $\varepsilon$ and $\eta$ are idiosyncratic error components, and $g$ and $G$, defined on $[-a,a]^d\to \mathbb{R}$ with $a$ being fixed\footnote{We consider the fixed $a$ case in the main text, and explain how to allow them to be defined on $\mathbb{R}^d$ in Appendix (ref).}, are unknown functions.

Model ((ref)) belongs to a class of partially linear models studied extensively in the relevant literature, see, robinson1988, hlg2000, Gao, LR2007, ttg2010, and BCH2014, for example. Existing estimation and inferential methods are mainly based on nonparametric kernel and series methods for the case where the dimensionality of $\mathbf{x}$ is small. In the current big--data environment where the dimensionality of $\mathbf{x}$ is large, there are newly proposed methods, including machine learning based methods.

The main features of model ((ref)), which has been proposed and discussed in BCH2014, are that model ((ref)) involves a binary structure for $z$, and BCH2014 develop a series based approach for the estimation of $\alpha$. By contrast, this paper proposes using a localized NN (LNN) method for the estimation of $\alpha$ and $g(\cdot)$ simultaneously. In addition, this paper also develops an easily implementable dependent wild bootstrap method for the inference of both $\alpha$ and $g(\cdot)$.

Until very recently, the investigation on NN architecture mainly focuses on some fixed design regression models (e.g., Cybenko1989 and many follow-up studies since then), or uses independent and identically distributed (i.i.d.) data (e.g., BK2019,SH2020; and many references therein). There are only limited studies available for us to understand NN with dependent data from theoretical perspective (see, cs1998,chen2007; for example), although NN based methods have been widely used to study time series data in practice (e.g., HOR1996,crs2001,GU2021429,Gu2020; just to name a few). We would like to contribute along this line of research, and thus assume the following time series data are observable:

eqnarray[eqnarray omitted — 74 chars of source]

where, for a positive integer $T$, $[T]$ stands for $\{ 1,\ldots,T\}$.

Meanwhile, it seems that so far the majority of the literature focuses on prediction errors, and barely talks about how to build feasible inferential procedures, such as constructing confidence intervals. A few exceptions known to us are DFLSP2021, FLM2021, CLMZ2022, and hhll24 for example on estimation and testing for the average treatment effect rather than on $g(\cdot)$ that we are also interested in this paper.

Therefore, one important objective and contribution of this paper is that we develop NN based approach to addressing both estimation and inferential issues for both $\alpha_0$ and $g(\cdot)$ using (ref). We also show how $G(\cdot)$ can be recovered via NN practically in Appendix (ref). All things considered, we draw Figure (ref) for the purpose of illustration, in which only the dark area of the hidden layer is activated. A few questions arise naturally:

enumerate[leftmargin=*, noitemsep] • Why are there only a small of number of functions getting activated ? • How does sparsity come to play? If the least absolute shrinkage and selection operator (i.e., LASSO) is employed, how do we define the set of true parameters ? • Provided a set of dependent time series data, can any inference (such as a confidence interval) be established ? and so forth.

{

figure[figure omitted — 1,845 chars of source]

}

Another challenge which arises with the complexity of NN architecture is the transparency of algorithms (see Appendix (ref) for a brief survey of the existing software packages). One main reason is the lack of practical guidelines for establishing a feasible version. Our literature review highlights that social science studies using NN approach rarely provide detailed descriptions of their numerical implementation. While we concur with Athey2019 that machine learning will have a transformative impact on social science, transparent algorithms are crucial to ensure the practical relevance and utility of the findings derived from these approaches.

Having those said, our contributions are as follows.

(i) We establish an approximation procedure that approximates polynomials via NN in a local sense rather than a global sense.

(ii) We then explore the use of identification restrictions and establish the LNN based approach under a set of mild conditions.

(iii) We show that some closed--form expressions can be obtained for the estimators of the parameters of interest.

(iv) Accordingly, asymptotic distributions are derived, and a dependent wide bootstrap procedure is proposed for inferential purposes.

(v) As shown in Theorems 2.1 and 2.2 and their discussions in Section 2 below, the LNN based estimation and inferential methods outperform such results associated with existing methods.

(vi) We validate our theoretical findings through extensive numerical studies.

(vii) In an empirical study, we revisit the impacts of a tightening monetary policy action on a variety of economic variables, including short-/long-term interest rate, inflation, unemployment rate, industrial price and equity return via the newly proposed framework using a monthly dataset of the US.

The rest of the paper is organized as follows. In the main text,

(a) Section (ref) introduces LNN architecture, proposes a group--LASSO based estimation procedure, and then establishes the corresponding asymptotic properties to infer $\alpha_0$ and $g(\cdot)$, respectively;

(b) We provide extensive simulation studies in Section (ref) to examine the finite-sample performance;

(c) Section (ref) presents an empirical study that investigates the average effects of the US monetary policy change on macroeconomic and financial variables;

(d) Section (ref) concludes.

In the online supplementary appendices,

(e) Appendix (ref) includes some discussions on issues associated with practical implementation and also presents a detailed algorithm;

(f) Appendix (ref) discusses the estimation of a fully nonparametric model which is a special case of (ref), explains how to relax the restriction about $a$, and infers $G(\cdot)$ of $z_t$ via LNN;

(g) Appendix (ref) includes additional simulations;

(h) We finally give the proofs in Appendix (ref).

To close this section, we introduce some notation and mathematical symbols. Vectors and matrices are always expressed in bold font. Further, $\| \cdot \|$ denotes the Euclidean norm of a vector or the Frobenius norm of a matrix; $ \mathbf{0}_{a}$ and $ \mathbf{1}_{a}$ are respectively $a\times 1$ vectors of zeros and ones for $a\in \mathbb{N}$ and $\mathbf{I}_a$ denotes an $a\times a$ identity matrix; for a vector of nonnegative integers $\pmb{\mu} = (\mu_1,\ldots, \mu_d)^\top \in \mathbb{N}_0^{d}$ in which $\mathbb{N}_0=0\cup \mathbb{N}$, let $\pmb{\mu}! =\mu_1!\cdots \mu_d!$; $\mathtt{c}$, $\mathtt{C}$ and $O(1)$ always stand for fixed constants, and may be different at each appearance; $\to_P$ and $\to_D$ stand for convergence in probability and convergence in distribution, respectively.

For a function $m\,:\, [-a,a]^d\mapsto \mathbb{R}$, let $\|m \|_{\infty} = \sup_{\mathbf{x}\in [-a,a]^d}|m(\mathbf{x})|$. If the partial derivative of $m(\mathbf{x})$ exists, we write $\frac{\partial^{|\pmb{\mu}|} m(\mathbf{x})}{\partial \mathbf{x}^{\pmb{\mu}}} =\frac{\partial^{|\pmb{\mu}|} m(\mathbf{x})}{\partial x_1^{\mu_1}\cdots \partial x_d^{\mu_d}}$ for short, where $|\pmb{\mu}| =\sum_{j=1}^d \mu_j$. Additionally, let

eqnarray*[eqnarray* omitted — 147 chars of source]

for $\mathbf{a}= (a_0, a_1,\ldots, a_d)^\top \in \mathbb{R}^{d+1}$, $\mathbf{r} = (r_0,r_1,\ldots, r_d)^\top\in \mathbb{N}_0^{d+1}$, and $ \mathbf{x}\in \mathbb{R}^{d}$. Denote that for given $q\geq 1$

eqnarray[eqnarray omitted — 158 chars of source]

where $\mathbf{n}=(n_1,\ldots, n_d)^\top\in \mathbb{N}_0^d$, and the dimension of $\mathscr{P}_q$ is apparently $\text{dim}\mathscr{P}_q =\binom{d+q}{d}\coloneqq d_q$. Accordingly, let

eqnarray[eqnarray omitted — 162 chars of source]

where $m_j(\mathbf{x} \, |\, \mathbf{x}_0)$'s are the basis monomials (centered at $\mathbf{x}_0$) of $ \mathscr{P}_q$. Denote a set

eqnarray[eqnarray omitted — 115 chars of source]

where $x_j$ and $x_{0,j}$ are the $j^{th}$ elements of $\mathbf{x}$ and $\mathbf{x}_0$ respectively, and $h$ is a bandwidth. Finally, let $ \mathbf{H}=\operatorname*{\normalfont\textrm{diag}}\{H_1,\ldots, H_{d_q}\}$ with $H_j =\prod_{k=1}^dh^{-n_{j,k}}$, and let $\mathbf{n}_j =(n_{j,1},\ldots, n_{j,d})^\top$ include the corresponding power terms of $m_j(\mathbf{x} \, |\, \mathbf{x}_0)$ defined in (ref).

Methodology and Asymptotic Theory

In this section, we first introduce LNN in Section (ref) and then infer $\alpha_0$ and $g(\cdot)$ in Section (ref). Notably, several new and useful results are established in Lemmas (ref)-(ref) to show how to approximate $g(\cdot)$ by LNN. These results contribute to the current literature in at least the following two points:

1. In contrast to the current literature that usually allows for all parameters to be estimated from data, we start by presenting some identification conditions, which can help reduce the number of effective parameters significantly.

2. One key idea behind NN architecture is that it can approximate polynomial terms, of which as well understood the linear combination can further approximate unknown functions by standard nonparametric analysis (e.g., BK2019, SH2020).

In this paper, we do the same, but the difference is that we introduce a bandwidth parameter $h$ below. The reason is that approximating polynomials via NN can only be achieved in a local sense rather than a global sense (cf., Lemmas (ref)-(ref) below). This is why we use the terminology LNN. More importantly, $h$ has a direct control on active and non-active activation functions of the hidden layer.

To proceed, we recall the notation defined in Section (ref) and explain how NN architecture works conceptually. In the literature of machine learning, one prefers to target the entire area that the unknown function is defined on, and normally uses a training set to pre-specify a lot of parameters, which do not change with respect to the observations of the test set. By doing so, one just pays some price when calculating the parameters in the first time, and no longer needs to update them when loading the test set. We are now ready to proceed.

LNN Architecture

First, we define a family of sufficiently smooth functions, and formally state the first assumption of the paper.

definition[Continuity] Let $p=q+s$ for some $q\in \mathbb{N}$ and $0<s\le 1$. A function $m\,:\, [-a,a]^d\mapsto \mathbb{R}$ is called $(p, \mathscr{C})$-smooth, if for every $\pmb{\mu} \in \mathbb{N}_0^d$ with $|\pmb{\mu}|=q$ the partial derivative $\frac{\partial^{|\pmb{\mu}|} m(\mathbf{x})}{\partial \mathbf{x}^{\pmb{\mu}}}$ exists and satisfies that for all $\mathbf{x}, \mathbf{z}\in [-a,a]^d$, $ \big|\frac{\partial^{|\pmb{\mu}|} m(\mathbf{x})}{\partial \mathbf{x}^{\pmb{\mu}}} -\frac{\partial^{|\pmb{\mu}|} m(\mathbf{z})}{\partial \mathbf{z}^{\pmb{\mu}}} \big| \le \mathscr{C} \| \mathbf{x} - \mathbf{z}\|^s.$
assumption\begin{enumerate}[leftmargin=*, noitemsep] • Let $g(\mathbf{x})$ be $(p,\mathscr{C})$-smooth, and $\max_{|\pmb{\alpha} |\le q} \| \frac{ \partial^{|\pmb{\alpha}|} g(\mathbf{x}) }{\partial\mathbf{x}^{\pmb{\alpha}} } \|_{\infty}\le \mathtt{c}$ for $0<c<\infty$. • Let Sigmoidal function $\sigma(\cdot)\, : \, \mathbb{R}\to [0,1]$ satisfy that \begin{enumerate}[leftmargin=*, noitemsep] • $\sigma(\cdot)$ is at least $q+1$ times continuously differentiable with bounded derivatives. • A point $u_\sigma\in \mathbb{R}$ exists, where all derivatives up to the order $q$ of $\sigma(\cdot)$ are different from zero. \end{enumerate} \end{enumerate}

Assumption (ref).1 is widely adopted in the literature of nonparametric regression (e.g., LR2007). The main point is that each component in Taylor expansion of $g(\cdot)$ is bounded and also sufficiently smooth. Assumption (ref).2 nests a wide class of activation functions commonly used in the literature as special cases, e.g., Sigmoidal squasher (i.e., $\sigma(x) = \frac{1}{1+\exp(-x)}$), Error function (i.e., $\sigma(x) =\frac{2}{\sqrt{\pi}}\int_0^x \exp(-w^2)dw$), etc. We refer the interested reader to DUBEY202292 for a comprehensive review on different activation functions, and to Appendix (ref) for a detailed example. We acknowledge the growing literature on the rectified linear unit (ReLU) function (e.g., SH2020, FLM2021). As ReLU and Sigmoidal functions require different approximation theories, we therefore focus on the Sigmoidal function in this paper.

lemmaLet Assumption (ref).2 hold. For $\forall x_0\in \mathbb{R}$, there exist $\pmb{\gamma}=(\gamma_1,\ldots, \gamma_{q+1})^\top$ and $\pmb{\beta}=(\beta_1,\ldots, \beta_{q+1})^\top$ with $\gamma_k= \frac{(-1)^{q+k-1} {\mathtt{C}}^q}{\sigma^{(q)} (u_\sigma)} \binom{q}{k-1} $ and $\beta_k= \frac{k-1}{\mathtt{C}}$, for $k\in [q+1]$, such that \begin{eqnarray*} \sup_{|x-x_0|\le h}\left| \sum_{k=1}^{q+1} \gamma_k \sigma (\beta_k \cdot (x-x_0)+ u_\sigma) - (x-x_0)^q \right| =O\left( h^{q+1} \right), \end{eqnarray*} where $\mathtt{C}$ is a constant, and $h$ is a bandwidth.
remark\normalfont\begin{enumerate}[leftmargin=*, noitemsep] • Lemma (ref) yields a recursive relationship, as one can repeatedly invoke Lemma (ref) to replace $(x-x_0)$ inside the activation. It will then yield a NN with multiple hidden layers. However, we do not see any benefit of doing so unless $g(\cdot)$ of (ref) has certain specific structure. • Lemma (ref) is independent of data. Only the order of the polynomial term to be approximated depends on the smoothness of $g(\cdot)$. The quantities $\pmb{\gamma}$, $\pmb{\beta}$ and $u_\sigma$ are fully decided by the activation function and the polynomial term, so they are known prior to regression. The constant $\mathtt{C}$ raises an issue of identifiability, so we simply let $\mathtt{C}=1$ throughout the rest of this paper. $h$ will be decided by the sample size later, and is introduced to control which activation functions will be activated. \end{enumerate}

We then show the feasibility of LNN architecture.

lemma[Feasibility] Suppose that Assumption (ref).2 holds. For $\forall \mathbf{x}_0 \in \mathbb{R}^d$, we define $p(\mathbf{x}\, |\, \mathbf{x}_0, \pmb{\lambda})=\pmb{\lambda}^\top\mathbf{m}(\mathbf{x}\, |\, \mathbf{x}_0),$ where $\pmb{\lambda} =(\lambda_1,\ldots, \lambda_{d_q})^\top$ with $d_q=\binom{d+q}{d}$ and $\|\pmb{\lambda} \|\le \mathtt{c}$. Then there exists a localized neural network of the form: \begin{eqnarray*} s(\mathbf{x}\, |\, \mathbf{x}_0, \widetilde{\pmb{\lambda}}) =(\widetilde{\pmb{\lambda}}\otimes \pmb{\gamma})^\top\pmb{\sigma}(\mathbf{x}\, |\, \mathbf{x}_0 ) \quadwith\quad \pmb{\sigma}(\mathbf{x}\, |\, \mathbf{x}_0 ) =\Big\{\sigma \big(\pi_{j0}+\sum_{k=1}^d\pi_{jk}(x_k-x_{k0}) \big) \Big\}_{d_q(q+1)\times 1} \end{eqnarray*} such that $\sup_{\mathbf{x}\in C_{\mathbf{x}_0,h} }|s(\mathbf{x}\, |\, \mathbf{x}_0, \widetilde{\pmb{\lambda}})-p(\mathbf{x}\, |\, \mathbf{x}_0, \pmb{\lambda})|= O(h^{q+1})$, where $\widetilde{\pmb{\lambda}}= \mathbf{D}^\top \pmb{\lambda}$ with $\mathbf{D}$ being a rotation matrix, and $\pmb{\pi}_j=(\pi_{j0},\cdots,\pi_{jd})^\top$ satisfies \begin{eqnarray*} (\pmb{\pi}_1,\ldots, \pmb{\pi}_{d_q (q+1)})&=&\frac{\operatorname*{\normalfontdiag}\{ h, \mathbf{I}_d\}\mathbf{W} ( \mathbf{I}_{d_q}\otimes \pmb{\beta}^\top)}{d+1} + \bigg[\begin{matrix} u_\sigma\\ \mathbf{0}_{d} \end{matrix}\bigg] \otimes \mathbf{1}_{d_q(q+1)}^\top, \end{eqnarray*} in which $\mathbf{W}$ is a user chosen matrix satisfying that $\max_{j}\|\mathbf{w}_{j}\|\le \sqrt{d+1}$ and $\mathbf{w}_{j_1}\ne \mathbf{w}_{j_2}$ for any given $j_1, j_2\in [d_q]$, and $\mathbf{w}_{j}$ stands for the $j^{th}$ column of $\mathbf{W}$.
remark\normalfont\begin{enumerate}[leftmargin=*, noitemsep] • We use vector form to rewrite each element of $\pmb{\sigma}(\mathbf{x}\, |\, \mathbf{x}_0 ) $, i.e., $\pmb{\sigma}(\mathbf{x}\, |\, \mathbf{x}_0 ) =\Big\{\sigma \big(\pi_{j0}+\sum_{k=1}^d\pi_{jk}(x_k-x_{k0}) \big) \Big\}_{d_q(q+1)\times 1} $ which is exactly what Sigmoidal activation function does (i.e., mapping a liner combination of regressors plus a location parameter to $[0,1]$). By design, $\pmb{\sigma}(\mathbf{x}\, |\, \mathbf{x}_0 )$ naturally and automatically explores different interaction terms of the regressors. • Lemma (ref) infers that LNN architecture rotates the parameters of interest (i.e., $\pmb{\lambda}$) with a pre-determined $d_q\times d_q$ full rank matrix $\mathbf{D}$, which can be easily constructed practically. We present the details in Appendix (ref) for the sake of space. • Without loss of generality, we can let $\mathbf{w}_j= \frac{\sqrt{d+1}}{q} \mathbf{w}_j^*$, where $\{\mathbf{w}_j^*\}$ are the vectors corresponding to the powers of the distinctive terms in the expansion of $(1+x_1+\cdots+x_d)^q$. Thus, $\|\mathbf{w}_j\|\le \frac{\sqrt{d+1}}{q} |\mathbf{w}_j^*|=\sqrt{d+1},$ which ensures Lemma (ref) can be invoked. \end{enumerate}

Using Lemma (ref), the LNN method requires us to focus on a compact set $[-a, a]^d$, as we need to partition it into lots of small cubes. These small cubes may have different names in each appearance. For example, they are referred to as (hyper-) cubes in BK2019 and SH2020, and are referred to as localization in FLM2021. Each cube is corresponding to an effective sample set, which is jointly determined by the point of interest and the bandwidth. Finally, we note that in Appendix (ref), we show that the main results remain valid when $a\rightarrow \infty$ along with the sample size.

That said, for a given integer $M\ge 1$, we subdivide $[-a, a]^d$ into $M^d$ cubes of side length $2h\coloneqq\frac{2a}{M}$, and number these cubes by $C_{\mathbf{x}_{0\mathbf{i}},h} $ with $\mathbf{i}\in [M]^d$:

eqnarray[eqnarray omitted — 138 chars of source]

where $\mathbf{x}_{0\mathbf{i}}$ represents the center of $C_{\mathbf{x}_{0\mathbf{i}},h} $, and $x_j$ and $x_{0\mathbf{i},j}$ are the $j^{th}$ elements of $\mathbf{x}$ and $\mathbf{x}_{0\mathbf{i}}$ respectively.

Let $\widetilde{s}(\mathbf{x} \, |\, \widetilde{\pmb{\Lambda}}) = \sum_{\mathbf{i}\in [M]^d} I_{\mathbf{i},h}(\mathbf{x}) s(\mathbf{x} \,|\, \mathbf{x}_{0\mathbf{i}}, \widetilde{\pmb{\lambda}}_{\mathbf{i}})$, $\widetilde{\pmb{\Lambda}} =\{\widetilde{\pmb{\lambda}}_{\mathbf{i}}\, |\, \mathbf{i}\in [M]^d\}$ and $I_{\mathbf{i},h}(\mathbf{x}) =I(\mathbf{x}\in C_{\mathbf{x}_{0\mathbf{i}},h} )$. We then establish the following important lemma.

lemma[LNN Architecture] Suppose that Assumption (ref) holds. Then we can approximate $g(\mathbf{x})$ by $ \widetilde{s}(\mathbf{x} \, |\, \widetilde{\pmb{\Lambda}})$ such that \begin{eqnarray*} \| g(\mathbf{x}) - \widetilde{s}(\mathbf{x} \, |\, \widetilde{\pmb{\Lambda}}) \|_{\infty}=O(h^p). \end{eqnarray*}
remark\normalfont The use of the indicator function is consistent with the sparsity setting of SH2020, where the author argues that “the network sparsity assumes that there are only few non-zero/active network parameters". In fact, our study clarifies the definition of “the network sparsity" by (i) defining non-active activation functions (such as those in Figure (ref)) which is realized through the use of indicator function $I_{\mathbf{i},h}(\mathbf{x})$; and (ii) pointing out the number of effective parameters. In Section (ref) below, we further explore the second point, as it allows us to define a set of true parameters when using thresholding techniques.

Estimation and Inference for $\alpha_0$ and $g(\cdot)$

We are now ready to work on our estimation method. To accommodate a potentially large number of explanatory variables, we adopt a sparse structure, and impose the following assumption.

assumptionSuppose that there exists an integer $d_0\leq d$ and a function $g_c\,:\, [-a,a]^{d_0}\to \mathbb{R}$ such that $g(\mathbf{x}_0) \equiv g_c(\mathbf{x}_{c,0})$, where $\mathbf{x}_{c,0}$ contains the first $d_0$ elements of $\mathbf{x}_0$.

Here, to be clear, we still require $d_0$ and $d$ to be fixed in theory as in BK2019, SH2020 and FLM2021. As pointed out in fan1996local, when it comes to estimate unknown functions, such settings where $d\geq 4$ and the same size is not big enough can cause the so--called “curse of dimensionality". In our numerical studies, with a relatively large $d_0$, our approach still achieves good finite sample properties.

Under Assumption (ref), it is clear to see that the coefficient vector $\pmb{\lambda}$ of Lemma (ref) inherits the sparsity of $g(\mathbf{x}_t)$. Specifically, we denote the true space as

eqnarray*[eqnarray* omitted — 154 chars of source]

where $\mathbf{n}_0=(n_1,\ldots, n_{d_0})^\top\in \mathbb{N}_0^{d_0}$. We then define $\overline{\mathscr{P}}_q$ as the complement space, such that $\mathscr{P}^0_q \cup \overline{\mathscr{P}}_q=\mathscr{P}_q$ and $\mathscr{P}^0_q \cap \overline{\mathscr{P}}_q=\emptyset$. Simple algebra shows that the dimensions of $\mathscr{P}^0_q$ and $\overline{\mathscr{P}}_q$ are $\text{dim}\mathscr{P}^0_q =\binom{d_0+q}{d_0}\coloneqq d^0_q$ and $\text{dim}\overline{\mathscr{P}}_q = d_q-d^0_q$ respectively. In connection with (ref), we suppose further that $\{m_1(\mathbf{x} \, |\, \mathbf{x}_0),\ldots, m_{d^0_q}(\mathbf{x} \, |\, \mathbf{x}_0)\}$ constitute a basis for $\mathscr{P}^0_q$ without loss of generality\footnote{This implies that we can express $m_j(\mathbf{x} \, |\, \mathbf{x}_0)=m_{c,j}(\mathbf{x}_c \, |\, \mathbf{x}_{c,0})$ for $j=1,\ldots, d_q^0$. However, this requirement is purely for notational simplicity and is not necessary for the validity of our estimation and theoretical development.}. Consequently, the sparsity of $\pmb{\lambda}$ is expressed as $\{\lambda_{j}=0\, |\,j=d_q^0+1,\ldots,d_q\}.$ Moreover, similar arguments to those presented in Lemma (ref) can be applied to show that the true function $g_c(\mathbf{x}_c)$ can be approximated by the oracle NN architecture $\widetilde{s}_c(\mathbf{x}_c \, |\, \widetilde{\pmb{\Lambda}}_c)$: $\| g_c(\mathbf{x}_c) - \widetilde{s}_c(\mathbf{x}_c \, |\, \widetilde{\pmb{\Lambda}}_c) \|_{\infty}=O(h^p)$, where $\widetilde{\pmb{\Lambda}}_c$ is the oracle counterpart of $\widetilde{\pmb{\Lambda}}$.

Drawing upon the sparsity structure of $\pmb{\lambda}$, we propose using the group-LASSO strategy (YL2006) to formulate the following penalized estimation:

eqnarray[eqnarray omitted — 255 chars of source]

where $\{\psi_j\}$ denote the tuning parameters, $\widetilde{Q} (\alpha,\pmb{\Theta})= \sum_{t=1}^T[y_t-z_t\alpha -\widetilde{s}(\mathbf{x}_t \, |\, \pmb{\Theta}) ]^2$, $\mathcal{S} = \{ \widetilde{s}(\mathbf{x} \, |\, \pmb{\Theta})\}$ with $\widetilde{s}(\mathbf{x} \, |\, \cdot)$ being defined in Lemma (ref), $\pmb{\Theta} =\{\pmb{\theta}_{\mathbf{i}}\, |\, \mathbf{i}\in [M]^d\}$ with $\pmb{\theta}_{\mathbf{i}}$'s being $d_q\times 1$ vectors, and $\pmb{\Theta}_{D,j}$ is an $M^d\times 1$ vector with the $j^{th}$ component being $\mathbf{D}^{\top,-1}\pmb{\theta}_{\mathbf{i}}$.

Let $\widetilde{\pmb{\theta}}_{\mathbf{i}}$ and $\widetilde{\pmb{\Theta}}_{D,j}$ be the corresponding estimators of $\pmb{\theta}_{\mathbf{i}}$ and $\pmb{\Theta}_{D,j}$. For the purpose of comparison, we also define the following oracle estimators: $ (\widetilde{\alpha}_c,\widetilde{\pmb{\Theta}}_c)$ of the form:

equation*[equation* omitted — 195 chars of source]

where $\widetilde{Q}_c (\alpha,\pmb{\Theta}_c)= \sum_{t=1}^T[y_t-z_t\alpha -\widetilde{s}_c(\mathbf{x}_{c,t} \, |\, \pmb{\Theta}_c) ]^2$, and $\pmb{\Theta}_c$, $\widetilde{s}_c(\mathbf{x}_{c,t} \, |\, \pmb{\Theta}_c)$ and $\mathcal{S}_{c}$ respectively represent the oracle counterparts of $\pmb{\Theta}$, $\widetilde{s}(\mathbf{x}_{t} \, |\, \pmb{\Theta})$ and $\mathcal{S}$ (i.e., assuming the sparsity is known). Although the oracle estimators $\widetilde{\alpha}_c$ and $\widetilde{\pmb{\Theta}}_c$ are infeasible practically, they can be approximated by the group-LASSO estimators with asymptotically negligible biases. To facilitate the rest of our development, we impose the following time series structure on (ref) in Assumption 3 below.

assumption\begin{enumerate}[leftmargin=*, noitemsep] • $\{(\mathbf{x}_t,\eta_t,\varepsilon_t)\, |\, t\in [T]\}$ are strictly stationary and $\alpha$-mixing with mixing coefficient \begin{eqnarray*} \alpha(t) = \sup_{A\in \mathcal{F}_{-\infty}^0, B\in \mathcal{F}_t^\infty} |P(A)P(B) -P(AB)| \end{eqnarray*} satisfying $\sum_{t= 1}^{\infty} \alpha^{\nu/(2+\nu)}(t) <\infty$ for some $\nu>0$, where $\mathcal{F}_{-\infty}^0$ and $\mathcal{F}_{t}^\infty$ are the $\sigma$-algebras generated by $\{(\mathbf{x}_s,\eta_s,\varepsilon_s): s \leq 0\}$ and $\{(\mathbf{x}_s,\eta_s,\varepsilon_s): s \geq t\}$, respectively. • The probability density function of $\mathbf{x}_1$, say $ f_{\mathbf{x}}(\cdot)$, and the function $G(\mathbf{x})$ are Lipschitz continuous and bounded on $[-a,a]^d$. Additionally, $f_{\mathbf{x}}(\cdot)$ is bounded away from 0 on $[-a,a]^d$. • $E[\varepsilon_1 \,| \, \mathbf{x}_1,\eta_1]=0$, $E[\varepsilon_1^2 \,| \, \mathbf{x}_1, \eta_1]=\sigma_\varepsilon^2$ almost surely (a.s.), and $E[|\varepsilon_1|^{2+\nu} \,| \, \mathbf{x}_1, \eta_1]\le\mathtt{c}$ a.s., where $\nu$ is the same as involved in Assumption (ref).1. \end{enumerate}

Assumption (ref) is standard in the literature of time series analysis (see, Gao, for example). Heterogeneity may also be introduced by, for example, $E[\varepsilon_1^2 \,| \, \mathbf{x}_1,\eta_1]=\psi(\mathbf{x}_1,\eta_1)$, which however makes notation even more complicated than what we need to involve. $\{\mathbf{x}_t\}$ may also have more complex structures such as linear processes, locally stationarity, deterministic trends, etc. Surely, the corresponding development and asymptotic results will need to be modified accordingly, but it will not add too many credits to the original idea of this paper. We therefore focus on the current setting.

To proceed, we introduce some additional notation. Let $\mathbf{H}_c$, $\mathbf{x}_{c,0\mathbf{i}}$, and $\mathbf{m}_c(\mathbf{x}_{c,t}\, |\, \mathbf{x}_{c,0\mathbf{i}})$ be the oracle counterparts of $\mathbf{H}$, $\mathbf{x}_{0\mathbf{i}}$, and $\mathbf{m}(\mathbf{x}_{t}\, |\, \mathbf{x}_{0\mathbf{i}})$. Also, let $\Phi_\eta(x)$ and $f_{\mathbf{x}_c}(\mathbf{x}_{c,\mathbf{i}})$ denote the CDF of $\eta_t$ and the density function of $\mathbf{x}_{c,t}$, respectively. Also define

eqnarray[eqnarray omitted — 768 chars of source]
assumption\begin{enumerate}[leftmargin=*, noitemsep] • Assume that $\inf_{\mathbf{x}\in [-a,a]^d} \lambda_{\min}(\pmb{\Omega}^\ast_0(\mathbf{x}))>0$, where{ \begin{eqnarray*} \pmb{\Omega}^\ast_0(\mathbf{x})=\left( \begin{array}{c c} 2^d\Phi_\eta(G(\mathbf{x})) f_{\mathbf{x}}(\mathbf{x})& \Phi_\eta(G(\mathbf{x})) f_{\mathbf{x}}(\mathbf{x}) \int_{[-1,1]^d} \mathbf{m}(\mathbf{x}\, |\, \mathbf{0})^\top \mathrm{d} \mathbf{x}\\ \Phi_\eta(G(\mathbf{x})) f_{\mathbf{x}}(\mathbf{x}) \int_{[-1,1]^d} \mathbf{m}(\mathbf{x}\, |\, \mathbf{0}) \mathrm{d} \mathbf{x} & f_{\mathbf{x}}(\mathbf{x}) \int_{[-1,1]^d} \mathbf{m}(\mathbf{x}\, |\, \mathbf{0}) \mathbf{m}(\mathbf{x}\, |\, \mathbf{0})^\top \mathrm{d} \mathbf{x} \end{array} \right). \end{eqnarray*}} • Suppose that $\sigma_{c,z}^2=\lim_{T\rightarrow\infty}\frac{1}{T}\sum_{t=1}^T\sum_{s=1}^TE[\widetilde{z}_t\widetilde{z}_s]E[\varepsilon_t\varepsilon_s]$ is a positive constant. • Let { $0<\lim\sup_{T\rightarrow \infty} Th^{d_0 +2p}<\infty$ when $d_0<d$, and $\lim_{T\rightarrow \infty} Th^{d_0 + 2p} =0$ when $d_0=d$, $\lim_{T\rightarrow \infty} \min_{d_q^0+1\leq j\leq d_q}\{\psi_jH_j\}T^{-\frac{1}{2}}=\infty$ and $\lim_{T\rightarrow \infty} \max_{j\in[d^0_q]} \{\psi_jH_j\}T^{-\frac{1}{2}} =0$}. \end{enumerate}

Assumption 4 imposes a set of regularity conditions. Assumptions 4.1 and 4.2 ensure the positive definiteness of the asymptotic covariances, and Assumption 4.3 regulates the bandwidth and the group-LASSO tuning parameters, which can be easily fulfilled. The first part of Assumption 4.3 requires that when the dimensionality of the true covariates satisfies $d_0<d$, there is no need to impose an under--smoothing condition, and it is required to impose the under--smoothing condition when $d_0=d$. For the second part of Assumption 4.3, detailed justifications are available from Appendix A.1 of the online supplementary document for more details.

Given these extra conditions, we establish the following lemma.

lemmaUnder Assumptions (ref)-(ref), we have \begin{enumerate}[leftmargin=*, noitemsep] • $P(\|\widetilde{\pmb{\Theta}}_{D,j}\|=0)\rightarrow1$, for $j=d_q^0+1,\ldots,d_q$; • $|\widetilde{\alpha}- \widetilde{\alpha}_c|=O_P(h^p)+o_P (T^{-\frac{1}{2}})$. \end{enumerate}

Lemma (ref).1 indicates that the group-LASSO method can correctly identify the sparsity of $\pmb{\lambda}$. Lemma (ref).2 shows $\widetilde{\alpha}$ and $ \widetilde{\alpha}_c$ are asymptotically equivalent up to a rate of an order of $O_P(h^p)+o_P (T^{-\frac{1}{2}} )$.

Let $\mathbf{Z}=(z_1,\cdots,z_T)^\top$ and $\widetilde{\mathbf{X}}_{c,\mathbf{i}} = (\widetilde{\mathbf{x}}_{c,\mathbf{i}, 1}, \cdots, \widetilde{\mathbf{x}}_{c,\mathbf{i}, T})^\top$, where $$\widetilde{\mathbf{x}}_{\mathbf{i},c,t} =I_{\mathbf{i},h} (\mathbf{x}_{c,t})(\mathbf{I}_{d^0_q}\otimes \boldsymbol{\gamma}^\top) \pmb{\sigma}_c( \mathbf{x}_{c,t}\, |\,\mathbf{x}_{0,c,\mathbf{i}})$$ and $\pmb{\sigma}_c( \mathbf{x}_{c,t}\, |\,\mathbf{x}_{0,c,\mathbf{i}})$ denotes the oracle counterpart of $\pmb{\sigma}(\mathbf{x}\, |\, \mathbf{x}_0 ) $. Building upon Lemma (ref) of the online supplement, we can now establish the asymptotic distributions of the proposed estimators.

theoremSuppose that Assumptions (ref)-(ref) hold. \begin{enumerate}[leftmargin=*, noitemsep] • As $T\rightarrow\infty$, \begin{equation*} \sqrt{T}\widetilde{M}_{c,z}(\widetilde{\alpha}-\alpha_0+O_P(h^p))\to_D N (0,1), \end{equation*} where $\widetilde{M}_{c,z}=T^{-1}\sigma_{c,z}^{-1}\mathbf{Z}^\top\big(\mathbf{I}_{T}-\sum_{\mathbf{i}\in [M]^{d_0}} \widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigl[\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top \widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigr]^{-1}\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top\big) \mathbf{Z}$. • For $\forall \mathbf{x}_{0}\in [-a,a]^{d}$, let $\widetilde{g}(\mathbf{x}_{0})= \sum_{\mathbf{i}\in [M]^{d}} \widetilde{\mathbf{x}}_{ 0,\mathbf{i}}^\top \widetilde{\pmb{\theta}}_{\mathbf{i}}$, where $\widetilde{\mathbf{x}}_{0, \mathbf{i}}=I_{\mathbf{i},h} (\mathbf{x})(\mathbf{I}_{d_q}\otimes \pmb{\gamma}^\top) \pmb{\sigma}( \mathbf{x}\, |\,\mathbf{x}_{0\mathbf{i}})$. Then, as $T\rightarrow\infty$, \begin{equation*} \sqrt{Th^{d_0}} \widetilde{\sigma}_{\mathbf{x}_{c,0}}^{ -1}(\widetilde{g}(\mathbf{x}_{0}) - g_c(\mathbf{x}_{c,0}) +O_P(h^p)) \to_D N (0, 1 ), \end{equation*} where $\widetilde{\sigma}_{\mathbf{x}_{c,0}}^2 = \sigma_\varepsilon^2\sum_{\mathbf{i}\in [M]^{d_0}} I_{\mathbf{i},h}(\mathbf{x}_{c,0}) \mathbf{m}_c(\mathbf{x}_{c,0}\, |\, \mathbf{x}_{c,0\mathbf{i}})^\top\mathbf{H}_c \pmb{\Sigma}_{c,\mathbf{i}}^{-1}\mathbf{H}_c\mathbf{m}_c(\mathbf{x}_{c,0}\, |\, \mathbf{x}_{c,0\mathbf{i}}).$ \end{enumerate}

Note that the rate of convergence, $\sqrt{T h^{d_0}}$, is faster than that of the conventional kernel estimator of an order of $\sqrt{T h^d}$ due to the fact that $Th^d = Th^{d_0} \, h^{d- d_0} = o\left(T h^{d_0}\right)$ when $d_0<d$. Meanwhile, the order of the bias term is also smaller than that of the conventional kernel estimator due to the construction of the LNN architecture and the definition of $p>2$. This is the main reason that the proposed LNN method outperforms over the conventional kernel method in finite--sample studies particularly when $d_0\geq 4$. Table 1 of Section 3 further supports the finite--sample superiority of the proposed LNN method.

When establishing the second result of Theorem (ref), the terms $E[\varepsilon_1\varepsilon_{1+t}]$ for $t\ge 1$ all vanish in the asymptotic covariance matrix due to the partition of (ref) and the use of the indicator function. As a result, the confidence interval associated with the second result of Theorem (ref) can be easily achieved.

Because $\sigma_{c,z}$ defined in Assumption 4 involves a long--run variance component, however, the serial dependence is not asymptotically negligible when inferring $\widetilde{\alpha}$. Therefore, we propose using the following dependent wild bootstrap method.

enumerate[leftmargin=*, noitemsep] • Based on the estimation residuals $\widehat{\varepsilon}_t=y_t - z_t\widetilde{\alpha} - \widetilde{g}(\mathbf{x}_{t}) $, generate the bootstrap random errors: $\varepsilon^\ast_t =\widehat{\varepsilon}_t\varsigma_t$, where $\varsigma_t$ is an $\ell$-dependent time series that satisfies $E[\varsigma_t]=0$, $E[\varsigma_t^2]=1$, $E[\varsigma_t^4]<\infty$, and $E[\varsigma_t\varsigma_s]=K\left(\frac{t-s}{\ell}\right)$ with $\ell\rightarrow\infty$ as $T\rightarrow\infty$. Here $K(\cdot)$ is a symmetric and Lipschitz continuous kernel function defined on $[-1,1]$, and satisfies $K(0)=1$ and $\int_{-\infty}^\infty K(u)e^{-iux}du\geq 0$ for all $x\in \mathbb{R}$. • Construct the dependent bootstrap variables as $y_t^\ast = z_t\widetilde{\alpha} + \widetilde{g}(\mathbf{x}_{t})+\varepsilon^\ast_t$. With $\{y_t^\ast, z_t, \mathbf{x}_t\}$, we can re-estimate $\alpha$ and $g(\mathbf{x}_{0})$ and obtain the bootstrap group-LASSO estimators $\widetilde{\alpha}^\ast$ and $\widetilde{g}^\ast(\mathbf{x}_0)$. • We repeat above two steps for a sufficiently large number of times and obtain the bootstrap draws.
theoremLet Assumptions (ref)-(ref) hold. Assume further there exists a positive number $\nu^\ast>\nu$ such that $E[|\varepsilon_1|^{2+\nu^\ast} \,| \, \mathbf{x}_1, \eta_1]\le\mathtt{c}$ a.s., and $\ell$ satisfies that $\ell h^{2p}\rightarrow0$, $\frac{\ell}{\sqrt{T}}\rightarrow0$, and $\ell \, T^{\frac{2(\nu-\nu^\ast)}{(2+\nu)(2+\nu^\ast)}}\rightarrow 0$, where $\nu$ is defined in Assumption (ref). Then, as $T\rightarrow\infty$, { \begin{enumerate}[leftmargin=*, noitemsep] • $\sup_{w}\Bigl|{\rm Pr}^\ast\bigl[\sqrt{T}(\widetilde{\alpha}^\ast-\widetilde{\alpha})\leq w \bigr]-{\rm Pr}\bigl[\sqrt{T}(\widetilde{\alpha}-\alpha_0)\leq w \bigr]\Bigr|=o_P(1)$, • $\sup_{w}\Bigl|\text{\normalfont Pr}^* (\sqrt{Th^{d_0}} \widetilde{\sigma}_{\mathbf{x}_{c,0}}^{-1}[\widetilde{g}^*(\mathbf{x}_0) - \widetilde{g}(\mathbf{x}_0)]\le w ) - \Pr (\sqrt{Th^{d_0}} \widetilde{\sigma}_{\mathbf{x}_{c,0}}^{-1}[\widetilde{g}(\mathbf{x}_0) - g_c(\mathbf{x}_{0,c})]\le w )\Bigr|=o_P(1)$, \end{enumerate} where ${\rm Pr}^\ast$} denotes the probability measure conditional on the observed sample.
remark\normalfont \begin{enumerate}[leftmargin=*, noitemsep] • It is a well-known problem with the LASSO-type estimators that simple residual-based bootstrap methods fail to consistently estimate their distributions unless some thresholding techniques are applied to handle zero components Chatterjee2011. However, in the case of group-LASSO estimation, which shares a similar idea to the adaptive-LASSO, it automatically incorporates soft-thresholding penalties. Consequently, there is no need for additional truncation. Similar discussions can be found in Section 4 of Chatterjee2011. • The serial dependence in $\varepsilon_t$ only complicates the inference for $\alpha_0$. For a purely nonparametric model without treatment components, a straightforward wild bootstrap method can be employed to mimic the distribution of $\widetilde{g}(\mathbf{x}_0)$. More detailed discussions are provided in Appendix (ref) of the online supplement. \end{enumerate}

To close this section, we emphasize that as have stated in Section (ref), we provide the detailed study about a fully nonparametric regression in Appendix (ref) (i.e., letting $\alpha_0 =0$). Although it is a special case of (ref), the investigation yields some useful insights and allows us to further clarify some features of LNN compared to the existing literature. Also, we infer $G(\cdot)$ of the binary structure of $z_t$ in Appendix (ref), which is also of great interest in both theory and practice.

Further Discussion

Up to this point, we would like to point out that those questions raised in Section (ref) have all been answered, so Figure (ref) can be understood better. We summarize some key points which may have been discussed previously here and there, and further discuss some remaining issues.

Sparsity --- It is now clear that without sparsity, the total number of activation functions is

eqnarray*[eqnarray* omitted — 79 chars of source]

of which only $d_q(q+1)$ neurons are activated when loading test data. Among the activated activation functions, the number of effective parameters is $d_q$, while the rest of the parameters are predetermined. Provided sparsity, utilizing identification conditions allows us to define the true set of parameters, so we can further reduce the effective parameters via the group-LASSO approach.

Multiple Hidden Layers --- Lemma (ref) yields a recursive relationship, as one can repeatedly invoke Lemma (ref) to replace $(x-x_0)$ inside the activation. It will then yield a LNN architecture with multiple hidden layers. Although having multiple hidden layers is achievable, at this stage it is not clear to us why we should do so. As discussed in Remark (ref), this step is completely independent of data, so we do not see any benefit of doing so unless $g(\cdot)$ (or $G(\cdot)$) has certain specific structure. Under some extra structure on $g(\cdot)$ (or $G(\cdot)$), however, the necessity of developing LNN architecture with multiple hidden layers deserves extra attention in future research.

Dependence & Trending --- LNN automatically eliminates some correlation of observations from different time periods when establishing the asymptotic distribution for the unknown functions. In this paper, we assume that the regressors $\{\mathbf{x}_t\, |\, t\in [T]\}$ are strictly stationary and mixing. In fact, they can have other more complex structures, such as linear processes, locally stationarity, heterogeneity, deterministic trends, etc. As a result, many climate models (such as those in MUDELSEE2019310) may be better captured. For such cases, one may need to revise the assumptions and proofs accordingly depending on detailed research questions.

Data-Splitting --- One may further connect the above results with the data-splitting technique of Chernozhukov2018, and simultaneously explore the sparsity of the binary structure in $z_t$.

Simulation

In this section, we conduct simulations to validate the theoretical findings, focusing specifically on the semiparametric model with sparsity. Additional simulation results are provided in Appendix (ref) of the online supplement to exam additional theoretical results of Appendix (ref), and to demonstrate the newly proposed method works reasonably well even without involving sparsity.

As discussed in Section (ref), many parameters are involved in the LNN architecture. It would be extremely difficult to systematically check every single one in this paper due to the page constraints, so we have to be selective. Also, we are constrained by computing power. That said, the following quantities are pre-fixed without loss of generality.

itemize[leftmargin=*, noitemsep] • Throughout, we use the Sigmoidal squasher, $\sigma(w)= 1/(1+\exp(-w))$, as the activation function. • $\pmb{\pi}_j$'s are generated in exactly the same way as mentioned in Remark (ref). • Let $R=200$ for the bootstrap procedure.

We consider the semiparametric treatment effects model (ref), where $\varepsilon_t = 0.5\varepsilon_{t-1} + N(0, 0.75)$, the $j^{th}$ element of $\mathbf{x}_t$ is independently generated as $x_{t,j}\sim U(-1, 1)$, and $z_t$ is independently generated from Bernoulli ($\cos(|\mathbf{x}_{c,t}^\top \mathbf{1}_{d_0}/d_0|)$) to allow for dependence between $z_t$ and $\mathbf{x}_{0,t}$. We simply let $\alpha_0=1$ and $g(\mathbf{x}_t)=g_c(\mathbf{x}_{c,t}) = 1+\sin (\mathbf{x}_{c,t}^\top \mathbf{1}_{d_0})$, where $\mathbf{x}_{c,t}$ contains the first $d_0$ elements in $\mathbf{x}_t$. The bandwidth $h$ is set as $h=a/M$, where $M$ is the integer closest to $a/h_1$ with $h_1=1.5\cdot T^{-1/(d+2p-0.5)}$. Here, $-0.5$ is to ensure $\sqrt{T} h^{p+d/2}\to 0$ holds. In fact, $h$ is very close to $h_1$, and the current setup is simply to guarantee $M$ is a large positive integer. When designing the simulations, our impression is that the results are not sensitive to the choices of $g(\mathbf{x})$ and the bandwidth. Therefore, in what follows, we only vary the values of $q$, $d$($d_0$), $u_\sigma$. Specifically, we use $d\in\{2, 8, 14\}$, $d_0=d/2$, $q\in\{3,4\}$ and $u_\sigma\in\{-0.5,0.5\}$.

To measure the estimate of $g(\cdot)$, we select a few test points from $[-a,a]^d$. Ideally, we would like to select $L$ points from each dimension, so it gives $L^d$ points to evaluate in total. However, it will create a lot computational overhead, so we select the points $\mathbf{x}_{L,j} \coloneqq (-a+\frac{2aj}{L+1} )\mathbf{1}_d$ for $j\in[L]$. At each test point, we construct the estimate along with the $95\%$ confidence interval using the method of Section (ref). To evaluate the finite sample performance, we compute the root mean squared errors (RMSEs) and coverage rates (CRs) after $n$ replications:

eqnarray*[eqnarray* omitted — 520 chars of source]

where $\widetilde{\alpha}_i$ and $\widetilde{g}_i(\cdot) $ stand for the estimates of $\alpha_0$ and $g(\cdot)$, respectively, at the $i^{th}$ replication. $\text{CR}_{\alpha, ij}$ and $\text{CR}_{g, ij}$ are the 95% confidence intervals of $\widetilde{\alpha}_i^\ast-\widetilde{\alpha}_i$ and $\widetilde{g}_i^*(\mathbf{x}_{L,j})-\widetilde{g}_i(\mathbf{x}_{L,j})$, respectively, based on the bootstrap draws from the $i^{th}$ replication. The number of bootstraps is set to be 200. We select $T\in \{800, 1600,2400 \}$, $L=10$ and $n=1000$ without loss of generality.

RMSEs and coverage rates are reported in Table (ref). As evident in the table, RMSEs of both $\widetilde{\alpha}_i$ and $\widetilde{g}_i(\cdot)$ decrease as $T$ increases from 800 to 2400. Moreover, coverage rates are close to 0.95 indicating that the bootstrap procedure behaves reasonably well. Some additional facts should also be mentioned. The simulation results are not changing significantly with respect to the value of $u_\sigma$, which confirms our argument in Remark (ref). For comparison, we also report the simulated RMSEs and CRs for the LNN estimation without considering the sparsity structure. As evident in Table (ref), the LNN-based group-LASSO estimators generally outperform the LNN estimation, especially in cases with large parameter sets (e.g., $d=14$ and $q=4$).

center[center omitted — 40 chars of source]

To further illustrate the proposed bootstrap method, we present the plots of 95% bootstrap confidence intervals of $g(\cdot)$ for the cases $(d,\mu_\sigma)=(2,0.5)$ and $(d,\mu_\sigma)=(2,-0.5)$ in Figures (ref) and (ref), respectively. For enhanced clarity, we employ a denser grid of evaluation points $\mathbf{x}_{L,ij} = (-a+\frac{2ai}{L+1}, -a+\frac{2aj}{L+1} )$ for $i,j\in[L]$. At each point, we apply the bootstrap procedure and compute the 95% confidence interval. The plots in Figures (ref) and (ref) demonstrate that the true functions are effectively covered across different choices of $T$ and $q$. Moreover, with the increasing sample size, the bootstrap confidence intervals exhibit clear convergence. Furthermore, $q=3$ seems to yield better coverage overall, and the choice of $u_\sigma$ has less impact on the inference.

center[center omitted — 48 chars of source]

Having demonstrated the finite sample performance of the newly proposed framework, we are now ready to examine the impacts of US monetary policy in the following section.

A Case Study

It is widely acknowledged that monetary policy shocks may have significant influence over macroeconomic variables. As noted by Clarida2000, “{\it the difference in policy behaviour could be an important underlying source of the shift in macroeconomic behaviour}". Consequently, over the past few decades, there has been a surge of empirical and theoretical research to study the relationship between monetary policies and various macroeconomic indicators, including interest rate, unemployment rate, economic growth, and asset price. For example, Clarida2000 investigate the role of monetary policy in the macroeconomic stability by establishing connections between the policy interest rate with expected inflation; Bernanke2005 explore how the stock index responds to the unanticipated monetary policy actions; Blanchard2010 study the policy effects on the relationship between inflation and unemployment rate; etc.

The Monetary Policy Effect Model

As widely recognized in the literature, a fundamental question in this field is how to characterize monetary policy shocks or unanticipated policy changes (see Christians1996,Bernanke2005; among others). A commonly adopted approach involves using the disturbance in a regression of monetary policy indicators on lagged observable macroeconomic variables to capture the information conveyed by policy actions that cannot be anticipated by the market. Then, the response of the macroeconomic variables to the policy shock can be measured through another regression between such variables. In this section, we adopt a nonparametric specification for the monetary policy evolution, which can be regarded as a generalization of the linear model employed by Christians1996. Specifically,

equation[equation omitted — 75 chars of source]

where $P_t$ is an indicator variable for the shift to a more tightening monetary policy, $\mathbf{x}_{t}$ contains the macroeconomic predictors available when $P_t$ is determined, $G_P(\cdot)$ is a nonparametricalyy unknown function, and $\eta_t$ is a disturbance term. We then have

eqnarray*[eqnarray* omitted — 77 chars of source]

where $\Phi_\eta(\cdot)$ denotes the CDF of $\eta_t$. Therefore, $\Phi_\eta(G_P(\mathbf{x}_t))$ captures the market anticipation of a tightening policy action by the authority. This concept is in line with the idea of policy propensity scores adopted by Angrist2018. Then, the unanticipated monetary policy shock $S_t$ can be specified as follows:

equation[equation omitted — 70 chars of source]

We can then characterize its relationship with macroeconomic or financial outcome variables ($M_t$) through the following semiparametric model:

equation[equation omitted — 87 chars of source]

where the parameter $\alpha_P$ captures the effects of unanticipated policy shocks and $G_M(\cdot)$ is another nonparametric function that controls the influence from the lagged macroeconomic predictors. By substituting $S_t$ from (ref) into (ref), we obtain

eqnarray[eqnarray omitted — 87 chars of source]

where $g_M(\mathbf{x}_{t})=G_M(\mathbf{x}_{t})-\Phi_\eta(G_P(\mathbf{x}_t))\alpha_P$. The setup of (ref) enables us to estimate the policy effects $\alpha_P$ using the methodology that is proposed in Section (ref).\footnote{In the online supplement, we provide an LNN-based method to estimate the nonparametric binary model (ref). Then, using the estimators $\widehat{G}_P(\mathbf{x}_t)$ and $\widehat{g}_M(\mathbf{x}_{t})$ that are obtained by estimating (ref) and (ref), respectively, we can recover $G_M(\mathbf{x}_{t})$ in (ref) by $\widehat{G}_M(\mathbf{x}_{t}) = \widehat{g}_M(\mathbf{x}_{t})+\Phi_\eta(\widehat{G}_P(\mathbf{x}_t))\widehat{\alpha}_P$.

An alternative approach to recovering the policy effects is to estimate (ref) and (ref) sequentially. However, this procedure involves an essential step where we have to replace the unobservable unanticipated policy shock $S_t$ in (ref) with its estimator $\widehat{S}_t=P_t-\Phi_\eta(\widehat{G}_P(\mathbf{x}_t))$ to construct a feasible estimator for $\alpha_P$. This substitution inevitably introduces additional approximation errors compared to the direct estimation of (ref). } Compared with the traditional parametric models of monetary policy effects, our framework offers several advantages. First, the relationship between variables is not constrained to be linear, allowing for more flexible and realistic modelling of complex interactions. Second, our approach is capable to accommodate a large set of control variables and automatically detect the insignificant predictors.

In what follows, we revisit the impacts of a tightening monetary policy action on a variety of economic variables, including short-/long-term interest rate, inflation, unemployment rate, industrial price and equity return. We accomplish this by employing the proposed semiparametric model with treatment effects on a monthly dataset of the US.

Variables and Data

In the literature Christians1996,Romer2000,Cochrane2002, a commonly used indicator of monetary policy shifts is the change in the federal funds target rate which is announced during Federal Reserve Open Market Committee (FOMC) meetings. Accordingly, we construct an indicator variable to capture the tightening or easing monetary policy whenever there is an increase in the announced target rate. We examine the effects of these policy changes on key macroeconomic and financial outcome variables that have been extensively studied in the literature. The sample period considered in this study is from January, 1989 to September, 2015.

Specifically, we first analyze the policy effects on a variety of interest rates: the effective federal funds rate (FFR) and treasury bond yields quoted on an investment basis at 3-month (Yield3m), 1-year (Yield1y), 2-year (Yield2y), 5-year (Yield5y), and 10-year (Yield10y) maturities. For these variables, we obtain their monthly average data from the FRED, Federal Reserve Bank of St. Louis. Following the approach of Angrist2018, we investigate the effects on inflation, unemployment rate, and industrial output, which are measured by (change in the log value of) the Personal Consumption Expenditures Price Index (PCE), (change in) the Civilian Unemployment Rate (UNRATE), and (change in the log value of) the Industrial Production Index (IP), respectively. The data for these variables are also collected from FRED. In addition to these macroeconomic variables, we follow Bernanke2005 to study the value-weighted return of the S$\&$P500 stocks as a representative of equity returns and the data is sourced from the CRSP index series database at WRDS.

To isolate the effects from the anticipation of the shifts in monetary policy, we incorporate a set of control variables that are sourced from three categories: (i) Monetary policy persistence: we measure the persistence of monetary policy using lagged federal funds target rate and real target rate changes (TRC), along with an indicator variable for FOMC meeting occurrences. (ii) Economic conditions: we include the first and second lags of inflation and unemployment rate to reflect the potential influence of underlying economic conditions on monetary policy changes. (iii) Market expectations: we utilize the federal funds future (FFF) index developed by Angrist2018 to capture market expectations regarding future monetary policy decisions \footnote{To measure market expectations of target rate changes, Angrist2018 develop the FFF variable using federal funds rate derivatives with specific adjustments made for data during FOMC meeting months. For detailed information about the construction of this index, we refer the readers to Angrist2018.}.

center[center omitted — 40 chars of source]

We present a summary of the aforementioned variables along with their descriptive statistics and unit root test results in Table (ref) below. As shown in the table, all variables are stationary at the 5% significance level, except 10-year treasury bond yields which is stationary at the 10% significance level. Therefore, these variables fit our assumption reasonably well.

Estimation Results

Using the methodology outlined in Section (ref), we investigate the response of macroeconomic and financial variables in the subsequent month following the announcement of a federal funds target rate increase. To set up the LNN architecture, we specify $h=T^{-1/(d+2p-0.5)}$, $\mu_\sigma=0.5$ and $q=3$. The estimated policy effects captured by $\widehat{\alpha}_P$ and their bootstrap confidence intervals are reported in Panel A of Table (ref).

Our estimation results reveal that an unanticipated increase in the target rate exerts a significant and positive influence on the federal funds rate and on the treasury bond yield with the maturity of 3 months. However, the effect is statistically insignificant for treasury bond yields with maturities that are more than 1 year, suggesting that tightening monetary policy has a more pronounced impact on short-term interest rates compared to long-term rates. This aligns with previous studies by Cochrane2002 and Angrist2018, who also observe diminishing effects along the yield curve with increasing maturity. Controlling for anticipated changes in the target rate, our estimation indicates no significant effects of monetary policy shocks on changes in inflation, unemployment rate, or industrial price. These findings are consistent with those reported by Angrist2018. Table (ref) presents information about the insignificant predictors identified by the LNN-based group-LASSO method. As shown in Table 3, different numbers of significant predictors, ranging from one to five, are detected for FFR, Yield1y, Yield2y, PCE, UNRATE, and IP. This result demonstrates the empirical significance of the proposed method.

center[center omitted — 47 chars of source]

We then explore the sensitivity of our results by varying the set of control variables, specifically by excluding the FFF or both FFF and lagged values of PCE and UNRATE. The results, which are also presented in Panel A of Table (ref), demonstrate that policy effects appear to be more pronounced with fewer control variables, which is likely due to the inclusion of the influence from anticipated monetary policy changes. Notably, the unemployment rate response becomes significantly negative when FFF is not controlled for. Furthermore, without controlling for FFF and the lagged economic indicators, a noticeable short-term “price puzzle” in the literature Sims1992 emerges --- a temporary increase in price in response to higher interest rate.

To examine the longer-term effects of tightening monetary policy, we compute the effects of target rate changes ($P_t$) on the average changes of future outcome variables: $\overline{M}_{t+L}=\frac{1}{L}\sum_{l=1}^LM_{t+l}$, across different horizons ($L$) up to 24 months. The estimation results are depicted in Figures (ref) and (ref), illustrating an accumulation of positive effects on short-term bond yields within the first 12 months, followed by a gradual decline in subsequent months. As a robustness check, we re-estimate the monetary policy effects using data before August 2008, when the global financial crisis started. The estimation results are reported in Panel B of Table (ref). Notably, the impact of tightening monetary policy in curbing the increase in unemployment rate was significant before the global financial crisis, but it became considerably more muted after the crisis.

center[center omitted — 48 chars of source]

We follow the relevant literature in treatment effects HIR2003 to estimate the average treatment effect (ATE) using the inverse probability weighting approach. Specifically, we calculate the policy propensity scores $\Phi_\eta(\widehat{G}_P(\mathbf{x}_t))$ based on the LNN estimation of the binary model (ref), as discussed in Appendix (ref) of the online supplement, and define the weight as $w_t=\Phi_\eta(\widehat{G}_P(\mathbf{x}_t))^{-P_t}[1-\Phi_\eta(\widehat{G}_P(\mathbf{x}_t))]^{-(1-P_t)}$. Then, we construct the inverse probability weighting estimator for ATE as follows:

equation*[equation* omitted — 116 chars of source]

We omit these results from this study to maintain focus on our main objective.

Conclusion

NN has gained considerable attentions over the past few decades. Yet, many questions, as outlined in Section (ref), have not been satisfactorily addressed in the relevant literature. In this paper, we bring in identification restrictions to the LNN framework from a semiparametric regression perspective, and consider the LNN based estimation and inference for the unknown parameter and function involved in the modelling of time series data. We then integrate the LNN architecture with the group-LASSO technique to achieve consistent estimation and automatic detection of insignificant regressors. The asymptotic distributions are derived accordingly, demonstrating that the sparsity structure can be effectively identified and the estimators exhibit the distributions that are asymptotically equivalent to those for the infeasible oracle estimators.

Additionally, we propose a dependent wild bootstrap procedure to obtain valid inferences in practice. Last but not least, we validate our theoretical findings through extensive numerical studies. In the empirical study, we revisit the impacts of a tightening monetary policy action on a variety of economic variables, including short-/long-term interest rate, inflation, unemployment rate, industrial price and equity return via the newly proposed framework using a monthly dataset of the US.

Several major comments have been made here and there in Section (ref), and some future research directions have been acknowledged along the way. Finally, we hope the current article will shed light on how to produce transparent algorithms to ensure that our research findings are useful for practical implementations and applications.

Acknowledgements

Gao, Peng and Yang acknowledge financial support from the Australian Research Council Discovery Grants Program under Grant Numbers: { DP200102769, DP210100476 and DP230102250}, respectively. Liu's research was financially supported by National Natural Science Foundation of China under Grant Number: 72203114.

table[table omitted — 6,421 chars of source]
figure[figure omitted — 883 chars of source]
figure[figure omitted — 890 chars of source]
table[table omitted — 1,690 chars of source]
table[table omitted — 710 chars of source]
table[table omitted — 4,100 chars of source]
figure[figure omitted — 866 chars of source]
figure[figure omitted — 692 chars of source]

{

center[center omitted — 40 chars of source]

Appendix A includes some discussions on issues associated with practical implementation and also presents a detailed algorithm. Appendix (ref) discusses the estimation of a fully nonparametric model which is a special case of (ref), explains how to relax the restriction about $a$, and infers $G(\cdot)$ of $z_t$ via LNN. Appendix (ref) includes additional simulations. We finally present the preliminary lemmas and the proofs in Appendix (ref).

\setcounter{equation}{0} \setcounter{lemma}{0} \setcounter{section}{0} \setcounter{table}{0} \setcounter{figure}{0} \setcounter{remark}{0} \setcounter{corollary}{0} \setcounter{assumption}{0}

\setcounter{equation}{0} \setcounter{lemma}{0} \setcounter{section}{0} \setcounter{table}{0} \setcounter{figure}{0} \setcounter{remark}{0} \setcounter{corollary}{0} \setcounter{assumption}{0}

Practical Implementation

First, we provide some examples to illustrate our concern about the existing software packages. In the literature, the “neuralnet" R package (see FG2010 for detailed illustration) has been well adopted. When training a NN, a key parameter is called “hidden" that is a vector of integers specifying the number of hidden neurons in each layer by R document, and the package refers to MYA1994 regarding the choice of the number of neurons. However, it is worth pointing out that MYA1994 use a modified AIC criterion to investigate the case with the number of activation functions (as well as the number of parameters) being finite which is reflected in their asymptotic development. As a consequence, the arguments of MYA1994 no longer hold when the number of activation functions is diverging. As we show in this paper, having a diverging number of activation functions is the minimum requirement to achieve asymptotic consistency, and the rate of divergence is also associated with the sample size under a set of minor conditions. Similar issues also apply to “deepnet" package of {\tt R}, “torch.nn.Linear" and “tfl.layers.Linear" of Python, “\textsf{feedforwardnet}" of Matlab, etc.

Next, we comment on the matrix $\mathbf{D}$, the selection of the tuning parameter, a computational algorithm for the LNN based group-LASSO, and some properties of Sigmoidal squasher.

On the Matrix $\mathbf{D}$ --- By the proof of Lemma (ref), we can determine $\mathbf{D}$ through $\mathbf{m}(\mathbf{x}\, |\, \mathbf{x}_0)=\mathbf{D}\mathbf{A}(\mathbf{x} \, | \,\mathbf{x}_0)$, where

eqnarray*[eqnarray* omitted — 205 chars of source]

with $[\pmb{\alpha}_1 ,\ldots, \pmb{\alpha}_{d_q}]= \frac{1}{d+1}\cdot\operatorname*{\normalfont\textrm{diag}}\{h,\mathbf{I}_d \} \mathbf{W}$.

On Tuning Parameters --- An essential consideration for the practical application lies in the selection of tuning parameters: $\psi_1,\cdots,\psi_{d_q}$. We simplify the selection of $d_q$ by utilizing the unpenalized estimators $\widetilde{\pmb{\Theta}}^\ast_j$ along with a common tuning parameter $\psi$:

equation[equation omitted — 119 chars of source]

Using analogous arguments in the proof of Lemma (ref) and Lemma (ref), we can show that $h^{d}H_j^{-2}\|\widetilde{\pmb{\Theta}}^\ast_{j}\|^2$ converges to a positive constant in probability for $j\in[d_q^0]$, while it has the order of $O_P\bigl(\frac{1}{Th^d}\bigr)$ for $j=d_q^0+1,\ldots,d_q$. With (ref), Assumption (ref).3 becomes

eqnarray*[eqnarray* omitted — 172 chars of source]

By the definition of $H_j$, it is evident that $\max_{j\in[d_q^0]}\{H_j\}=O(h^{-q})$ and $\min_{j=d_q^0+1,\ldots,d_q}\{H_j\}=O(1)$. Therefore, a sufficient condition for the tuning parameter will be $\psi_0\sqrt{T}h^{d}\rightarrow\infty$ and $\psi_0T^{-\frac{1}{2}}h^{-q}\rightarrow0$. The remaining task involves the selection of $\psi_0$, which can be achieved by minimising the following information criterion:

eqnarray*[eqnarray* omitted — 147 chars of source]

where $ (\widetilde{\alpha}_\psi,\widetilde{\pmb{\Theta}}_\psi)$ are the group-LASSO estimators using (ref), and ${\rm df}_\psi$ is the number of nonzero coefficients identified by $\widetilde{\pmb{\Theta}}_\psi$. Accordingly, the optimal tuning parameter is obtained by

equation*[equation* omitted — 82 chars of source]

Algorithm for the LNN Based Group-LASSO --- The literature provides well-established computational algorithms for the LASSO estimation FL2001, HL2005. Herein, we adopt the local quadratic approximation procedure.

enumerate[leftmargin=*, noitemsep] • Obtain the unpenalized estimators $(\widetilde{\alpha}^{(0)},\widetilde{\pmb{\Theta}}^{(0)})$ as the initial estimators. • The estimators $(\widetilde{\alpha}^{(m)},\widetilde{\pmb{\Theta}}^{(m)})$ in the $m^{th}$ step are constructed as \begin{eqnarray} \left(\widetilde{\alpha}^{(m)},\widetilde{\pmb{\Theta}}^{(m)}\right) = \operatorname*{\arg\!\min}_{\alpha, \pmb{\Theta}}\left( \widetilde{Q} (\alpha,\pmb{\Theta})+\sum_{j=1}^{d_q}\psi_j\frac{\|\pmb{\Theta}_{D,j}\|^2}{\|\widetilde{\pmb{\Theta}}^{(m-1)}_{D,j}\|}\right), \end{eqnarray} where $\widetilde{\pmb{\Theta}}^{(m-1)}_{D,j}$ denotes the vector containing the $j^{th}$ elements of $\mathbf{D}^{\top,-1}\widetilde{\pmb{\theta}}^{(m-1)}_{\mathbf{i}}$'s in the $(m-1)^{th}$ step. Simple algebra shows that the first-order conditions for $\widetilde{\alpha}^{(m)}$ and $\widetilde{\pmb{\theta}}^{(m)}_{\mathbf{i}}$ are given by \begin{eqnarray} \sum_{t=1}^Tz_t[y_t-z_t\widetilde{\alpha}^{(m)} -\sum_{\mathbf{i}\in [M]^d} \widetilde{\mathbf{x}}_{\mathbf{i}, t}^\top \widetilde{\pmb{\theta}}^{(m)}_{\mathbf{i}}]&=&0, \nonumber\\ -\sum_{t=1}^T\widetilde{\mathbf{x}}_{\mathbf{i}, t}[y_t-z_t\widetilde{\alpha}^{(m)} - \widetilde{\mathbf{x}}_{\mathbf{i}, t}^\top \widetilde{\pmb{\theta}}^{(m)}_{\mathbf{i}}]+\mathbf{D}^{-1}\pmb{\Phi}^{(m-1)}\mathbf{D}^{\top, -1}\widetilde{\pmb{\theta}}^{(m)}_{\mathbf{i}}&=&\mathbf{0}, \end{eqnarray} where $\pmb{\Phi}^{(m-1)}$ is a $d_q\times d_q$ diagonal matrix with its $j^{th}$ diagonal element being $\psi_j/(\|\widetilde{\pmb{\Theta}}^{(m-1)}_{D,j}\|)$. Solving the equations in (ref), we obtain \begin{eqnarray*} \widetilde{\alpha}^{(m)}&=& [\mathbf{Z}^\top \mathbf{M}^{(m-1)}_{x,\phi} \mathbf{Z} ]^{-1}\mathbf{Z}^\top \mathbf{M}^{(m-1)}_{x,\phi} \mathbf{Y}, \nonumber\\ \widetilde{\pmb{\theta}}^{(m)}_{\mathbf{i}}&=& [\widetilde{\mathbf{X}}_{\mathbf{i}}^\top \widetilde{\mathbf{X}}_{\mathbf{i}}+\mathbf{D}^{ -1}\pmb{\Phi}^{(m-1)} \mathbf{D}^{\top, -1}]^{-1}\widetilde{\mathbf{X}}_{\mathbf{i}}^\top [\mathbf{Y}-\mathbf{Z}\widetilde{\alpha}^{(m)} ], \end{eqnarray*} where $\mathbf{Z}=(z_1,\cdots,z_T)^\top$, $\mathbf{Y}=(y_1,\cdots,y_T)^\top$, $\widetilde{\mathbf{X}}_{\mathbf{i}} = (\widetilde{\mathbf{x}}_{\mathbf{i}, 1}, \cdots, \widetilde{\mathbf{x}}_{\mathbf{i}, T})^\top$, and $$\mathbf{M}^{(m)}_{x,\phi}=\mathbf{I}_{T}-\sum_{\mathbf{i}\in [M]^d} \widetilde{\mathbf{X}}_{\mathbf{i}}\left[\widetilde{\mathbf{X}}_{\mathbf{i}}^\top \widetilde{\mathbf{X}}_{\mathbf{i}}+\mathbf{D}^{ -1}\pmb{\Phi}^{(m)}\mathbf{D}^{\top, -1}\right]^{-1}\widetilde{\mathbf{X}}_{\mathbf{i}}^\top.$$ • Repeat Step 2 until numerical convergence.

On Sigmoidal Squasher --- For Sigmoidal squasher, we have

eqnarray*[eqnarray* omitted — 98 chars of source]

in which

eqnarray[eqnarray omitted — 57 chars of source]

is the recursion relation for Stirling numbers of the second kind. We refer interested readers to MINAI1993845 for more details. Sigmoidal squasher is easy to use in the sense that we can arbitrarily choose $u_\sigma$ of Lemma (ref). Without loss of generality, we let $u_\sigma \in \{-0.5, 0.5\}$ in the numerical studies. In Figure (ref), we plot $\sigma^{(k)}(x) $ for $k=0,\ldots, 5$ for the purpose of demonstration.

Extra Theoretical Results

Treatment on $a$

Before we explain how to relax the restriction on $a$, we consider a fully nonparametric model by letting $\alpha_0\equiv 0$ and ignoring the sparsity (i.e., $d\equiv d_0$). The rest settings are identical to Section (ref).

The model to be investigated becomes

eqnarray*[eqnarray* omitted — 53 chars of source]

and the objective function is a simplified version of that involved in (ref):

equation[equation omitted — 119 chars of source]

Accordingly the OLS estimator of $\widetilde{\pmb{\Lambda}}$ is obtained by

eqnarray[eqnarray omitted — 242 chars of source]

and, for $\forall \mathbf{x}_0 \in [-a,a]^d$, the estimator of $g( \mathbf{x}_0)$ is then defined by $\widehat{g}(\mathbf{x}_0) = \widetilde{s}(\mathbf{x}_0 \, |\, \widehat{\pmb{\Theta}} ) $. Equation (ref) admits a closed-form estimator for each $\widehat{\pmb{\theta}}_{\mathbf{i}}$. To see this, we write

eqnarray*[eqnarray* omitted — 235 chars of source]

where $\widetilde{\mathbf{x}}_{\mathbf{i},t} =I_{\mathbf{i},h} (\mathbf{x}_t)(\mathbf{I}_{d_q}\otimes \boldsymbol{\gamma}^\top) \boldsymbol{\sigma}( \mathbf{x}_t\, |\,\mathbf{x}_{0\mathbf{i}})$ and the equality follows from the fact that $I_{\mathbf{i},h} (\mathbf{x}_t)I_{\mathbf{j},h} (\mathbf{x}_t)=0$ for $\mathbf{i}\ne \mathbf{j}$. Thus, for $\forall\mathbf{i}$, the first order condition yields

eqnarray[eqnarray omitted — 238 chars of source]

where the invertibility of $\sum_{t=1}^T\widetilde{\mathbf{x}}_{\mathbf{i},t} \widetilde{\mathbf{x}}_{\mathbf{i},t} ^\top$ is guaranteed asymptotically in view of (ref) and (ref).

After carefully studying (ref) for each $\mathbf{i}$ and repeatedly invoking $I_{\mathbf{i},h} (\mathbf{x}_t)I_{\mathbf{j},h} (\mathbf{x}_t)=0$ for $\mathbf{i}\ne \mathbf{j}$, the following theorem holds.

theoremSuppose that Assumptions (ref) and (ref) hold. For $\forall \mathbf{x}_0\in [-a,a]^d$, \begin{eqnarray*} \sqrt{Th^d} \widehat{\sigma}_{\mathbf{x}_0}^{ -1}(\widehat{g}(\mathbf{x}_0) - g(\mathbf{x}_0) +O_P(h^p)) \to_D N (0, 1 ), \end{eqnarray*} where $\widehat{\sigma}_{\mathbf{x}_0}^2 = \sigma_\varepsilon^2\mathbf{m}(\mathbf{x}_0 \, |\, \mathbf{x}_0)^\top \mathbf{H} \pmb{\Sigma}_{\mathbf{x}_0}^{-1} \mathbf{H} \mathbf{m}(\mathbf{x}_0 \, |\,\mathbf{x}_0)$, and $\pmb{\Sigma}_{\mathbf{x}_0} = f_{\mathbf{x}}( \mathbf{x}_0) \int_{[-1,1]^d} \mathbf{m}(\mathbf{x}\, |\, \mathbf{0}) \mathbf{m}(\mathbf{x}\, |\, \mathbf{0})^\top \mathrm{d} \mathbf{x}$.

Obviously, our LNN based estimation method is simple and easy to implement. Accordingly, we propose the following bootstrap procedure to establish inference in practice.

enumerate[leftmargin=*, noitemsep] • We calculate $\widehat{\varepsilon}_t = y_t -\widehat{g}(\mathbf{x}_t)$ for $t\in [T]$. • Collect i.i.d. draws of $\{\eta_t\, |\, t\in [T]\}$ from $ N(0,1)$, and construct the bootstrap version dependent variables as follows: $ y_t^* = \widehat{g}(\mathbf{x}_t) + \widehat{\varepsilon}_t \eta_t$. We re-estimate $g(\cdot)$ using $\{ (y_t^*, \mathbf{x}_T) \, |\, t\in [T]\}$ as under (ref), and denote the estimate as $ \widehat{g}^*(\cdot)$. • Repeat Step 2 $R$ times, where $R$ is sufficiently large.

For the bootstrap procedure, the following result holds immediately.

theoremLet the conditions of Theorem (ref) hold. Suppose further that $T h^{d+ 2p}\to 0$. For $\forall \mathbf{x}_0\in [-a,a]^d$, we have \begin{eqnarray*} \sup_{w}|\normalfont Pr^* (\sqrt{Th^d} \widehat{\sigma}_{\mathbf{x}_0}^{-1}[\widehat{g}^*(\mathbf{x}_0) - \widehat{g}(\mathbf{x}_0)]\le w ) - \Pr (\sqrt{Th^d} \widehat{\sigma}_{\mathbf{x}_0}^{-1}[\widehat{g}(\mathbf{x}_0) - g(\mathbf{x}_0)]\le w )|=o_P(1). \end{eqnarray*} where $\text{\normalfont Pr}^*$ is the probability measure induced by the bootstrap procedure.

Note that, we require $g(\mathbf{x})$ to be defined on a compact set, but do not impose restriction on the range of $\{\mathbf{x}_t\}$. In fact, for time series data, it may make more sense to assume that $a$ is diverging, which is indeed achievable. We now provide two treatments to relax the restriction on $a$.

Treatment 1: Suppose that $\mathbf{x}_t$ follows a sub-Gaussian distribution, and we can then require $\sqrt{\log (Th^d)} \cdot a\to \infty.$ A similar treatment has also been discussed in LTG2016 for example, so we do not further elaborate it here. However, there is a price that we have to pay, i.e., the slow rate of convergence.

Treatment 2: Alternatively, we can modify the construction of LNN from a nonparametric viewpoint. In the literature of kernel regression, one normally pre-specifies a point of interest (e.g., $\mathbf{x}_0$), and investigates a small area nearby only which is usually decided by some bandwidth(s) converging to 0. As a result, the parameters obtained from the estimation procedure usually vary with respect to the point of interest. In other words, when evaluating different points from the test set, the number of parameters to be estimated will be proportional to the cardinality of the training set (although estimation is always carried on using the same training set). Provided a large test set, it may create lots of overhead from a computational viewpoint, but the advantage is that in theory it allows us to consider $\forall \mathbf{x}_0\in \mathbb{R}^d$.

That said, for $\forall\mathbf{x}_0\in\mathbb{R}^d$, we consider the following objective function:

equation[equation omitted — 151 chars of source]

where the subscript $L$ infers the local version, $\pmb{\theta}$ is $d_q\times 1$ vector satisfying $\| \pmb{\theta}\|<\infty$, we let $I_{0,h}(\mathbf{x}_t) :=I(\mathbf{x}_t \in C_{\mathbf{x}_0,h})$ for short, and $C_{\mathbf{x}_0,h}$ is defined in (ref) already. Then the OLS estimate of $\widetilde{\pmb{\lambda}}$ defined in Lemma (ref) is obtained by

eqnarray*[eqnarray* omitted — 102 chars of source]

and, accordingly, the estimate of $g(\mathbf{x}_0)$ is defined by $\widehat{g}_L(\mathbf{x}_0) = s(\mathbf{x}_0 \, |\, \mathbf{x}_0, \widehat{\pmb{\theta}} ) .$ We can then produce the following corollary.

corollarySuppose that Assumptions (ref) and (ref) hold. As $(1/h, Th^d)\to (\infty, \infty)$, for $\forall \mathbf{x}_0\in \mathbb{R}^d$, \begin{enumerate}[leftmargin=*, noitemsep] • $\sqrt{Th^d} \, \left(\mathbf{H}^{-1} \mathbf{D}^{\top,-1}(\widehat{\pmb{\theta}} -\widetilde{\pmb{\lambda}}) +O_P(h^p))\right) \to_D N\left(\mathbf{0},\sigma_\varepsilon^2\pmb{\Sigma}_{\mathbf{x}_0}^{-1}\right),$$\sqrt{Th^d} \widehat{\sigma}_{\mathbf{x}_0}^{-1}(\widehat{g}_L(\mathbf{x}_0) - g(\mathbf{x}_0) +O_P(h^p)) \to_D N (0, 1 ),$ \end{enumerate} where $\widetilde{\pmb{\lambda}}$ is uniquely determined by $g(\mathbf{x}_0)$.

A Nonparametric Binary Model

So far, we have not explored the binary structure of $z_t$ much, which is also of great interest widely adopted in a wide range of applications (Athey2019). To close our investigation about the model (ref), we use LNN approach to infer $G(\cdot)$.

Recall that

eqnarray*[eqnarray* omitted — 49 chars of source]

For simplicity, we suppose that the respective probability density function (PDF) and the cumulative distribution function (CDF) of $\eta$ are known, and denote the PDF and CDF by $\phi_\eta(\cdot)$ and $\Phi_\eta(\cdot)$ respectively. Here, the information about $\phi_\eta(\cdot)$ and $\Phi_\eta(\cdot)$ is necessary for carrying on likelihood estimation.

Direct calculation shows that

eqnarray*[eqnarray* omitted — 143 chars of source]

which yield $E[z \, |\, \mathbf{x} ] =\Phi_\eta(G(\mathbf{x}))$. Accordingly, the log-likelihood function is defined below:

eqnarray[eqnarray omitted — 197 chars of source]

where the definition of $l_t(\cdot)$ is obvious. To infer $G(\cdot)$, we consider the following objective function:

eqnarray*[eqnarray* omitted — 339 chars of source]

which yields the following maximum likelihood estimator:

eqnarray[eqnarray omitted — 162 chars of source]

Note that our LNN method involves a general unknown function form. As a result, the classical results of likelihood estimation, such as those in NEWEY19942111, no longer hold. Therefore, before establishing an asymptotic distribution using $\widehat{\pmb{\Theta}}$, we state a lemma to show the feasibility of LNN architecture when modelling binary outcomes.

lemmaUnder Assumptions (ref) and (ref), \begin{eqnarray*} \frac{1}{T}\sum_{t=1}^T[\Phi_\eta( g(\mathbf{x}_t))- \Phi_\eta( \widetilde{s}(\mathbf{x}_t\, | \, \widehat{\pmb{\Theta}}) )]^2 =o_P(1) . \end{eqnarray*}

Lemma (ref) provides the consistency, and also bridges the likelihood estimation and the nonlinear least squares approach to some extent. To be precise, Lemma (ref) does not provide any specific consistency for $\forall\, \widehat{\pmb{\theta}}_{\mathbf{i}}$. Instead, it evaluates the overall performance of LNN. More importantly, it says when modelling a binary outcome, the likelihood estimation using the LNN architecture is approximately equivalent to implementing a nonlinear least squares method provided the distribution of $\eta_t$ is correctly specified. In addition, Lemma (ref) further infers that

eqnarray*[eqnarray* omitted — 132 chars of source]

so many remarks made previously can be directly applied. Last but not least, Lemma (ref) facilitates numerical implementation in practice, which will be further discussed in Appendix (ref).

Below, we establish the following asymptotic distribution.

theoremSuppose that Assumptions (ref) and (ref) hold. For $\forall \mathbf{x}_0\in [-a,a]^d$, \begin{eqnarray*} \sqrt{Th^d} \widetilde{\sigma}_{\mathbf{x}_0}^{-1} (\widehat{g}(\mathbf{x}_0) - g(\mathbf{x}_0) +O_P(h^p)) \to_D N (0,1), \end{eqnarray*} where $ \widehat{g}(\mathbf{x}_0) = \widetilde{s}(\mathbf{x}_0 \, |\, \widehat{\pmb{\Theta}} )$, $\widetilde{\sigma}_{\mathbf{x}_0}^2 =\mathbf{m}(\mathbf{x}_0 \, |\, \mathbf{x}_0)^\top \mathbf{H} \widetilde{\pmb{\Sigma}}_{\mathbf{x}_0}^{-1}\mathbf{H} \mathbf{m}(\mathbf{x}_0 \, |\, \mathbf{x}_0)$, and \begin{eqnarray*} \widetilde{\pmb{\Sigma}}_{\mathbf{x}_0}=\frac{f_{\mathbf{x}}(\mathbf{x}_0) \phi_\eta(g(\mathbf{x}_0))^2 }{[1- \Phi_\eta(g(\mathbf{x}_0))] \Phi_\eta(g(\mathbf{x}_0))} \int_{[-1,1]^d} \mathbf{m}(\mathbf{x}\, |\, \mathbf{0}) \mathbf{m}(\mathbf{x}\, |\, \mathbf{0})^\top \mathrm{d} \mathbf{x} . \end{eqnarray*}

In light of PA2012, we propose a score based wild bootstrap approach for inferential purposes as follows.

enumerate[leftmargin=*, noitemsep] • For each bootstrap replication, we collect i.i.d. draws of $\{\eta_t \, |\, t\in [T] \}$ from $N(0,1)$, and calculate \begin{eqnarray} \widehat{\pmb{\theta}}_{\mathbf{i}}^* = \widehat{\pmb{\theta}}_{\mathbf{i}} + \left(\sum_{t=1}^T\frac{\partial^2 \log l_t(\widetilde{s}(\mathbf{x}_t\, | \, \widehat{\pmb{\Theta}})) }{\partial \pmb{\theta}_{\mathbf{i}} \partial \pmb{\theta}_{\mathbf{i}}^\top } \right)^{-1} \sum_{t=1}^T\frac{\partial \log l_t(\widetilde{s}(\mathbf{x}_t\, | \, \widehat{\pmb{\Theta}})) }{\partial \pmb{\theta}_{\mathbf{i}}} \eta_t, \end{eqnarray} where $l_t(\cdot)$ is defined in (ref). • Repeat Step 1 $R$ times, where $R$ is sufficiently large.

It is worth pointing out that the above procedure is computationally efficient in the sense that the right hand side of (ref) enjoys a closed-form expression, which is given in (ref) and (ref) specifically. Practically, the bootstrap procedure may require much less time compared with the estimation of (ref) itself.

The following theorem holds for the above bootstrap procedure.

theoremLet the conditions of Theorem (ref) hold. Suppose further that $T h^{d+ 2p}\to 0$. For $\forall \mathbf{x}_0\in [-a,a]^d$, we have \begin{eqnarray*} \sup_{w}|\normalfont Pr^* (\sqrt{Th^d} \widetilde{\sigma}_{\mathbf{x}_0}^{-1}[\widehat{g}^*(\mathbf{x}_0) - \widehat{g}(\mathbf{x}_0)]\le w ) - \Pr (\sqrt{Th^d} \widetilde{\sigma}_{\mathbf{x}_0}^{-1}[\widehat{g}(\mathbf{x}_0) - g(\mathbf{x}_0)]\le w )|=o_P(1), \end{eqnarray*} where $\text{\normalfont Pr}^*$ is the probability measure induced by the bootstrap procedure, and $\widehat{g}^*(\cdot)$ is yielded by the bootstrap draws in an obvious manner.

Extra Simulation

We now provide extra simulations, which have two focuses: (1). examining the theoretical results in Appendix (ref), and (2). demonstrating the newly proposed method works reasonably well even without involving sparsity.

On the Fully Nonparametric Model --- Consider the following regression model:

equation[equation omitted — 52 chars of source]

where the variables are generated in the same manner as in Section (ref). In what follows, we only vary the values of $(q, d, u_\sigma)$. Specifically, we consider the cases $d=2, 8$. The rest parameters are identical to those in the main text.

To measure the finite sample performance, we select the points as follows:

eqnarray*[eqnarray* omitted — 354 chars of source]

With each dataset, we first estimate all $g(\mathbf{x}_{L,\mathbf{j}})$ using the approach of Appendix (ref), and then construct the corresponding 95% confidence interval using the bootstrap procedure documented in Theorem (ref) for each point. We report $\text{RMSE}_g$ and $\text{CR}_g$ which are defined in the main text already.

First, we draw some plots for the case with $d=2$. In both Figures (ref) and (ref), the first sub-plot is always the true $g(\mathbf{x})$. For the rest of sub-plots, each has three layers. The middle one is the average of estimates over $n$ replications. The top and bottom layers are the averages of the bootstrap draws corresponding to the 97.5% and 2.5% quantiles respectively. A few facts emerge. Overall, the LNN approach can recover the unknown function reasonably well. Also, both figures are very similar, so the results are not sensitive to the choice of $u_\sigma$ as explained in Remark (ref).

More detailed numbers are summarized Table (ref). As expected, when $T$ goes up, RMSE$_g$ converges to 0 and CR$_g$ converges to 0.95. Also, we note that when $d$ increases, RMSE$_g$ increases but still has reasonable performance expect the case with $(d=8,q=4)$. Therefore, it seems that for large $d$, smaller $q$ yields better finite sample performance.

On the Binary Model --- We consider the following data generating process:

eqnarray*[eqnarray* omitted — 128 chars of source]

in which $\eta_t = 0.5\eta_{t-1} + N(0, 0.75)$, the $j^{th}$ element of $\mathbf{x}_t$ is generated as $x_{t,j}\sim U(-a, a)$, and $G(\mathbf{x}) = 1+\sin(\mathbf{x}^\top \mathbf{1}_d/d)$. The rest parameters are identical to the simulation design of the fully nonparametric model.

We first note a computational issue. We reply on “fminunc" function of Matlab to solve

eqnarray*[eqnarray* omitted — 128 chars of source]

which is the same as that in (ref). In order to invoke the minimization process in any statistical software (including R, Matlab, etc.), one needs to provide initial values to the parameters under estimation. As a consequence, the numbers reported below are affected by the initial values more or less. Although it is not our intention to tackle this complicated computational issue in this paper, Lemma (ref) does become useful in this case. Recall that Lemma (ref) bridges the log likelihood estimation and the nonlinear least squares estimation. Therefore, we first conduct an OLS estimation using the approach of Appendix (ref) as the initial value of $\pmb{\Theta}$ for each generated $\{( z_t, \mathbf{x}_t)\, |\, t\in [T]\}$. We then invoke log likelihood estimation as our final estimate of $\pmb{\Theta}$ for each dataset. Even in this case, the computation is rather slow, and the computational time increases dramatically when the number of parameters goes up.

We draw a few plots in Figures (ref) and (ref). The first sub-plot is always the true $G(\mathbf{x})$. For the rest of sub-plots, each has three layers. The middle one is the average of estimates over $n$ replications. The top and bottom layers are the averages of the bootstrap draws corresponding to the 97.5% and 2.5% quantiles respectively. Overall, the LNN approach can recover the unknown function reasonably well. Also, both figures are very similar, so the results are not sensitive to the choice of $u_\sigma$.

We further summarize the detailed numbers in Table (ref). A few facts should be mentioned. First, the coverage rates are reasonably well. As expected, when $T$ goes up, RMSE converges to 0 and CR converges to 0.95. Also, we note that when $d$ increases, RMSE increases but still has reasonable performance. The results are not changing much with respect to the value of $u_\sigma$. Again, it seems that for large $d$, smaller $q$ yields better finite sample performance.

{

Preliminary Lemmas & Proofs

Before proving the theoretical results in Appendices (ref)-(ref), we first present all preliminary lemmas in Appendix (ref).

Preliminary Lemmas

For $\mathbf{a}= (a_0, a_1,\ldots, a_d)^\top \in \mathbb{R}^{d+1}$, $\mathbf{r} = (r_0,r_1,\ldots, r_d)^\top\in \mathbb{N}_0^{d+1}$, and $ \mathbf{x}\in \mathbb{R}^{d}$, we let

eqnarray*[eqnarray* omitted — 177 chars of source]

Obviously, we have $f_{\mathbf{a}}\in \mathscr{P}_q$, where $ \mathscr{P}_q$ is defined in (ref).

lemmaLet $f\,:\, \mathbb{R}^d\to \mathbb{R}$ be a $(p,\mathscr{C})$-smooth function. For $\forall \mathbf{x}_0\in \mathbb{R}^d$, let \begin{eqnarray*} p_q(\mathbf{x}\, |\, \mathbf{x}_0) =\sum_{0\le |\mathbf{J}|\le q} \frac{1}{\mathbf{J}!}\cdot \frac{\partial^{|\mathbf{J}|} f(\mathbf{x}_0) }{\partial \mathbf{x}^{\mathbf{J}}} (\mathbf{x}-\mathbf{x}_0)^{\mathbf{J}}. \end{eqnarray*} Then \begin{eqnarray*} \|f(\mathbf{x})-p_q(\mathbf{x}\, |\, \mathbf{x}_0)\|_{\infty}\le O(1)\|\mathbf{x}-\mathbf{x}_0 \|^p, \end{eqnarray*} where $p=q+s$, $O(1)$ depends on $d$ and $q$ only.
lemmaFor almost all $\mathbf{a}_1,\ldots,\mathbf{a}_{d_q}\in \mathbb{R}^{d+1}$ (with respect to the Lebesgue measure in $\mathbb{R}^{(d+1)\times d_q}$), we have that $\{f_{\mathbf{a}_1} (\mathbf{x}),\ldots,f_{\mathbf{a}_{d_q}} (\mathbf{x})\}$ is a basis of the linear vector space $ \mathscr{P}_q$.
lemmaUnder Assumptions (ref), (ref), and (ref), we have \begin{eqnarray*} \|\widetilde{\alpha}-\alpha_0\|^2+ h^d\sum_{\mathbf{i}\in [M]^d}\|\mathbf{H}^{-1}\mathbf{D}^{\top,-1}(\widetilde{\pmb{\theta}}_{\mathbf{i}}-\widetilde{\pmb{\lambda}}_{\mathbf{i}})\|^2=O_P\left( \frac{1}{Th^{d}}\right). \end{eqnarray*}
lemmaSuppose Assumptions (ref), (ref), and (ref) hold. As $(1/h, Th^{d_0})\to (\infty, \infty)$, \begin{enumerate} • $\sqrt{T}(\widetilde{\alpha}_{c}-\alpha_0+O_P(h^p))\to_D N\left(\mathbf{0}, M_{c,z}^{-2}\sigma_{c,z}^2\right)$; • $\sigma_\varepsilon^{-1}\pmb{\Sigma}_{c,\mathbf{i}}^{1/2}\sqrt{Th^{d_0}}\left(\mathbf{H}_c^{-1} \mathbf{D}_c^{\top,-1}(\widetilde{\pmb{\theta}}_{c,\mathbf{i}} -\widetilde{\pmb{\lambda}}_{c,\mathbf{i}})+O_P(h^p)\right) \to_D N\left(\mathbf{0},\mathbf{I}_{d_q^0}\right)$, for each $\mathbf{i}\in [M]^{d_0}$, \end{enumerate} where $\mathbf{H}_c$, $\mathbf{D}_c$, and $\widetilde{\pmb{\lambda}}_{c,\mathbf{i}}$ are counterpart matrices of $\mathbf{H}$, $\mathbf{D}$, and $\widetilde{\pmb{\lambda}}_{\mathbf{i}}$ for the true model, and \begin{eqnarray*} M_{c,z}&=&\int_{\mathbf{x}\in [-a,a]^d} \Phi_\eta(G(\mathbf{x})) f_{\mathbf{x}}(\mathbf{x})\mathrm{d} \mathbf{x}-\lim_{T\rightarrow\infty}h^{d_0}\sum_{\mathbf{i}\in [M]^{d_0}}\mathbf{M}_{c,\mathbf{i}}^\top \pmb{\Sigma}_{c,\mathbf{i}}^{-1} \mathbf{M}_{c,\mathbf{i}} \nonumber\\ \pmb{\Sigma}_{c,\mathbf{i}}&=&f_{\mathbf{x}_c}(\mathbf{x}_{c,\mathbf{i}}) \int_{[-1,1]^{d_0}} \mathbf{m}_c(\mathbf{x}_c\, |\, \mathbf{0}) \mathbf{m}_c(\mathbf{x}_c\, |\, \mathbf{0})^\top \mathrm{d} \mathbf{x}_c, \end{eqnarray*} with \begin{equation*} \mathbf{M}_{c,\mathbf{i}}=\int_{[-a,a]^{d-d_0}}\Phi_\eta(G(\mathbf{x}_{c,0\mathbf{i}},\mathbf{z})) f_{\mathbf{x}}(\mathbf{x}_{c,0\mathbf{i}},\mathbf{z})\mathrm{d}\mathbf{z} \int_{[-1,1]^{d_0}} \mathbf{m}_c(\mathbf{x}_c\, |\, \mathbf{0}) \mathrm{d} \mathbf{x}_c. \end{equation*}
lemmaSuppose Assumptions (ref) and (ref) hold. As $(1/h, Th^d)\to (\infty, \infty)$, for each $\mathbf{i}\in [M]^d$, \begin{eqnarray*} \sigma_\varepsilon^{-1}\pmb{\Sigma}_{\mathbf{i}}^{1/2}\sqrt{Th^d}\left(\mathbf{H}^{-1} \mathbf{D}^{\top,-1}(\widehat{\pmb{\theta}}_{\mathbf{i}} -\widetilde{\pmb{\lambda}}_{\mathbf{i}})+O_P(h^p)\right) \to_D N\left(\mathbf{0},\mathbf{I}_{d_q}\right), \end{eqnarray*} where $ \pmb{\Sigma}_{\mathbf{i}} =f_{\mathbf{x}}(\mathbf{x}_{\mathbf{i},0}) \int_{[-1,1]^d} \mathbf{m}(\mathbf{x}\, |\, \mathbf{0}) \mathbf{m}(\mathbf{x}\, |\, \mathbf{0})^\top \mathrm{d} \mathbf{x}$.

Before presenting the next lemma, we calculate the partial derivatives of $\log L(\widetilde{s}(\cdot\, | \, \pmb{\Theta}))$ with respect to each $\pmb{\theta}_{\mathbf{i}}$.

eqnarray[eqnarray omitted — 1,261 chars of source]

where the third equality follows from the fact that $\mathbf{x}_t$ can not simultaneous belong to $C_{\mathbf{x}_{0\mathbf{i}},h}$ and $C_{\mathbf{x}_{0\mathbf{j}},h}$ for $\mathbf{i}\ne \mathbf{j}$ by the construction of $\widetilde{s}(\cdot\, | \, \pmb{\Theta})$, and $\widetilde{\mathbf{x}}_{\mathbf{i}, t}$ is the same as that defined in Section (ref).

Based on $\frac{\partial \log L(\widetilde{s}(\cdot\, | \, \pmb{\Theta})) }{\partial \pmb{\theta}_{\mathbf{i}}}$ and some tedious calculation, the second order derivative is

eqnarray[eqnarray omitted — 1,641 chars of source]
lemmaUnder Assumptions (ref) and (ref), for each $\mathbf{i}\in [M]^d$, \begin{enumerate} • $\frac{1}{T}\sum_{t=1}^TI_{\mathbf{i}, h}(\mathbf{x}_t) [g(\mathbf{x}_t) - s(\mathbf{x}_t \, | \, \widetilde{\mathbf{x}}_{\mathbf{i}}, \widehat{\pmb{\theta}}_{\mathbf{i}} )]^2 =o_P(1)$, • $\left\|\frac{1}{T}\mathbf{H}\mathbf{D}\frac{\partial^2 \log L(\widetilde{s}(\cdot\, | \, \widehat{\pmb{\Theta}})) }{\partial \pmb{\theta}_{\mathbf{i}} \partial \pmb{\theta}_{\mathbf{i}}^\top }\mathbf{D}^\top \mathbf{H} -\widetilde{\pmb{\Sigma}}_{\mathbf{i}}\right\| =o_P(1)$, • $\left\|\frac{1}{T}\mathbf{H}\mathbf{D}\frac{\partial^2 \log L(s(\cdot\, | \, \widetilde{\mathbf{x}}_{\mathbf{i}}, \pmb{\theta}_{\mathbf{i}}^*) ) }{\partial \pmb{\theta}_{\mathbf{i}} \partial \pmb{\theta}_{\mathbf{i}}^\top }\mathbf{D}^\top \mathbf{H} -\widetilde{\pmb{\Sigma}}_{\mathbf{i}}\right\| =o_P(1)$, • $\widetilde{\pmb{\Sigma}}_{\mathbf{i}}^{-1/2}\frac{1}{\sqrt{Th^d}} \mathbf{H} \mathbf{D}\frac{\partial \log L(g(\cdot) ) }{\partial \pmb{\theta}_{\mathbf{i}} }\to_D N(\mathbf{0}, \mathbf{I}_{d_q})$, \end{enumerate} where $\pmb{\theta}_{\mathbf{i}}^*$ lies between $\widehat{\pmb{\theta}}_{\mathbf{i}}$ and $ \widetilde{\pmb{\lambda}}_{\mathbf{i}}$, and \begin{eqnarray*} \widetilde{\pmb{\Sigma}}_{\mathbf{i}}=\frac{f_{\mathbf{x}}(\widetilde{\mathbf{x}}_{\mathbf{i}}) \phi_\varepsilon(g(\widetilde{\mathbf{x}}_{\mathbf{i}}))^2 }{[1- \Phi_\varepsilon(g(\widetilde{\mathbf{x}}_{\mathbf{i}}))] \Phi_\varepsilon(g(\widetilde{\mathbf{x}}_{\mathbf{i}}))} \int_{[-1,1]^d} \mathbf{m}(\mathbf{x}\, |\, \mathbf{0}) \mathbf{m}(\mathbf{x}\, |\, \mathbf{0})^\top \mathrm{d} \mathbf{x} . \end{eqnarray*}

Proofs for the LNN Architecture

Proof of Lemma (ref):

This is Lemma 8 of BK2019, so the derivation is omitted. {$\blacksquare$}

Proof of Lemma (ref):

In what follows, let

eqnarray*[eqnarray* omitted — 200 chars of source]

It suffices to show that $f_{\mathbf{a}_1} (\mathbf{x}),\ldots,f_{\mathbf{a}_{d_q}} (\mathbf{x})$ are linearly independent. To do this, let $b_1,\ldots, b_{d_q}\in \mathbb{R}$ be such that

eqnarray[eqnarray omitted — 79 chars of source]

As we explained under (ref), the monomials involved in (ref) are linearly independent. Thus, (ref) implies that

eqnarray*[eqnarray* omitted — 131 chars of source]

Note that $\sharp\widetilde{\mathbf{r}}=d_q$ by design, so we can construct a one-to-one relationship between $k\in [d_q]$ and $\mathbf{r}\in \widetilde{\mathbf{r}}$. Using this relationship, we can construct $p(\mathbf{x})\in \mathscr{P}_q$ as follows:

eqnarray*[eqnarray* omitted — 159 chars of source]

which satisfies

eqnarray[eqnarray omitted — 89 chars of source]

Position 4 in SAUER2006191 implies that (ref) has the only solution $p(\cdot)=0$ in $\mathscr{P}_q$ for Lebesgue almost all $\mathbf{a}_1,\ldots,\mathbf{a}_{d_q}\in \mathbb{R}^{d+1}$, which in turn implies $b_1=\cdots =b_{d_q}=0$. The proof is now completed. {$\blacksquare$}

Proof of Lemma (ref):

Before proceeding further, we would like to point out that in what follows, $\mathtt{C}$ is a constant for the purpose of rescaling only.

By Assumption (ref).2, there is a point $u_\sigma\in \mathbb{R}$ such that none of the derivatives up to the order $q$ is 0 at $u_\sigma$. Thus, we construct the following one-layer NN:

eqnarray[eqnarray omitted — 554 chars of source]

in which the definitions of $\gamma_k$'s and $\beta_k$'s are obvious.

By Assumption (ref).2 again, $\sigma(\cdot)$ is $q+1$ times continuously differentiable. Thus, it can be expanded in a Taylor series with Lagrange remainder around $u_\sigma$ up to order $q$:

eqnarray[eqnarray omitted — 691 chars of source]

where $\xi_k\in [u_\sigma - \frac{k}{\mathtt{C}}\cdot |x-x_0|, u_\sigma + \frac{k}{\mathtt{C}}\cdot |x-x_0|]$ for all $0\le k\le q$.

Note that

eqnarray*[eqnarray* omitted — 345 chars of source]

where $ \Big\{

array[array omitted — 21 chars of source]

\Big\}$ is the Stirling number of the second kind. The Stirling number of the second kind describes the number of options to split a set of $j$ elements into $n$ non-empty subsets, which is equal to 0 for $0\le j<n$, and is equal to 1 for $j=n$. The result holds true for all $j, n\in \mathbb{N}$ (AS1972).

Thus, we can further simplify the right hand side of (ref), and write

eqnarray*[eqnarray* omitted — 376 chars of source]

which in connection with (ref) yields that

eqnarray*[eqnarray* omitted — 696 chars of source]

In view of Assumption (ref).2 and $\mathtt{C}$ being a fixed value, the proof is now completed. {$\blacksquare$}

Proof of Lemma (ref):

By Lemma (ref), we can reconstruct all of $\{m_i(\mathbf{x}\, |\, \mathbf{x}_0)\}$ as follows:

eqnarray*[eqnarray* omitted — 112 chars of source]

where

eqnarray*[eqnarray* omitted — 300 chars of source]

Note that the rotation matrix $\mathbf{D}$ is determined by $\pmb{\alpha}_{j}$'s only, so they are fixed.

Apparently, we have an issue of identification here, because for example we can arbitrarily rescale $\pmb{\alpha}_{j}$'s, and modify $\mathbf{D}$ accordingly without changing $m_i(\mathbf{x}\, |\, \mathbf{x}_0)$ as follows:

eqnarray*[eqnarray* omitted — 151 chars of source]

in which $ \mathbf{B}$ is full rank. Therefore, for the purpose of identification, we regulate $\pmb{\alpha}_j$'s as follows:

eqnarray[eqnarray omitted — 116 chars of source]

in which $\mathbf{I}_h=\operatorname*{\normalfont\textrm{diag}}\{h,\mathbf{I}_d \}$, and $\mathbf{W}$ is defined in the body of this lemma. As a result, for $\forall j\in [d_q]$,

eqnarray*[eqnarray* omitted — 286 chars of source]

so we can invoke Lemma (ref) later on.

Treating $(1,\mathbf{x}^\top-\mathbf{x}_0^\top)\pmb{\alpha}_{j}$ as a whole and using Lemma (ref), we write

eqnarray[eqnarray omitted — 601 chars of source]

where the last line follows from Lemma (ref). Also, $\pmb{\beta}$, $\pmb{\gamma}$ and $u_\sigma$ are known as discussed in Remark (ref). Thus, we can further write

eqnarray[eqnarray omitted — 978 chars of source]

where the last line follows from the facts that $d_q$ is fixed and $\|\pmb{\lambda} \|=O(1)$.

Finally, let $\widetilde{\pmb{\lambda}} =(\widetilde{\lambda}_1,\ldots, \widetilde{\lambda}_{d_q})^\top$ with $\widetilde{\lambda}_j = \pmb{\lambda}^\top \mathbf{d}_j$. Further, in view of the definitions of $\pmb{\sigma}(\mathbf{x}\, |\, \mathbf{x}_0 ) $ and $(\pmb{\pi}_1,\ldots, \pmb{\pi}_{d_q (q+1)})$ in the body of this lemma and (ref), the proof is then completed. {$\blacksquare$}

Proof of Lemma (ref):

Before starting the proof, we introduce a few notations to facilitate the development. First, recall that in the body of this theorem, we have defined

eqnarray*[eqnarray* omitted — 236 chars of source]

where $s (\mathbf{x}\, |\, \widetilde{\mathbf{x}}_{\mathbf{i}}, \widetilde{\pmb{\lambda}}_{\mathbf{i}}) = (\widetilde{\pmb{\lambda}}_{\mathbf{i}}\otimes \pmb{\gamma})^\top\pmb{\sigma} (\mathbf{x}\, |\, \widetilde{\mathbf{x}}_{\mathbf{i}} )$ by the definition of Lemma (ref). Second, note that the leading terms of the $q^{th}$ order Taylor expansion of $g(\mathbf{x}) $ at each $ \widetilde{\mathbf{x}}_{\mathbf{i}}\in (-a,a)^d$ can be written as follows:

eqnarray*[eqnarray* omitted — 438 chars of source]

where the definition of $\pmb{\lambda}_{\mathbf{i}}$ should be obvious in view of the definition of $\mathbf{m}(\mathbf{x}\, |\, \widetilde{\mathbf{x}}_{\mathbf{i}})$ according to (ref).

We are now ready to start the proof, and write

eqnarray*[eqnarray* omitted — 582 chars of source]

Note that by Lemma (ref) we choose $\widetilde{\pmb{\lambda}}_{\mathbf{i}}$ which fulfils the relationship: $\widetilde{\pmb{\lambda}}_{\mathbf{i}} =\mathbf{D}^\top \pmb{\lambda}_{\mathbf{i}}$. It is worth mentioning that although $(\widetilde{\mathbf{x}}_{\mathbf{i}}, \pmb{\lambda}_{\mathbf{i}})$ vary with respect to $\mathbf{i}$, the rotation matrix $\mathbf{D}$ in facts is solely determined by $\mathbf{W}$ of Lemma (ref). Therefore, without loss of generality, we can fix $\mathbf{W}$ over $\mathbf{i}$, as it is user chosen. Then $\mathbf{D}$ remains the same in view of the proof of Lemma (ref).

Next, we write

eqnarray[eqnarray omitted — 678 chars of source]

where the inequality follows from the definition of $I_{\mathbf{i},h}(\mathbf{x})$, and the last step follows from Lemma (ref). Also, we can obtain that

eqnarray[eqnarray omitted — 703 chars of source]

where the inequality follows from the definition of $I_{\mathbf{i},h}(\mathbf{x})$, and the last step follows from Lemma (ref) by letting $\widetilde{\pmb{\lambda}}_{\mathbf{i}} =\mathbf{D}^\top \pmb{\lambda}_{\mathbf{i}}$.

Therefore, based on (ref) and (ref), we obtain

eqnarray*[eqnarray* omitted — 115 chars of source]

where $p=q+s$. The proof is now completed. {$\blacksquare$}

Proofs for the Main Results

Proof of Lemma (ref):

We adopt a similar strategy with that for Lemma A.1 of WX2009 to establish the results in Lemma (ref). For notational simplicity, let $\widetilde{Q}_\psi (\alpha,\pmb{\Theta})=\widetilde{Q} (\alpha,\pmb{\Theta})+\sum_{l=1}^{d_q}\psi_l\|\pmb{\Theta}_{l}\|$, where $\widetilde{Q} (\alpha,\pmb{\Theta})= \sum_{t=1}^T[y_t-z_t\alpha -\widetilde{s}(\mathbf{x}_t \, |\, \pmb{\Theta}) ]^2$. Let $\mathbf{B} =\{ \mathbf{b}_{\mathbf{i}}\, |\, \mathbf{i}\in [M]^d\}$ and $\|\mathbf{B}\|_H^2=\sum_{\mathbf{i}\in [M]^d}\|\mathbf{H}^{-1}\mathbf{D}^{\top,-1}\mathbf{b}_{\mathbf{i}}\|^2$, where $\mathbf{b}_{\mathbf{i}}$ is a $d_q\times 1$ vector of constants. Additionally, denote $d_T=\frac{1}{\sqrt{T}h^d}$.

By FL2001, it suffices to show that for any $\epsilon>0$, there exists a constant $C>0$ such that

eqnarray[eqnarray omitted — 262 chars of source]

We write

eqnarray[eqnarray omitted — 509 chars of source]

where $\pmb{\Lambda}_l$ and $\mathbf{B}_{D,l}$ are vectors that contain the $l$-th elements of $\pmb{\lambda}_{\mathbf{i}}$ and $\mathbf{D}^{\top,-1}\mathbf{b}_{\mathbf{i}}$, respectively.

For $R_1$, directly using the definition of $\widetilde{s}(\mathbf{x} \, |\, \pmb{\Theta})$ and $s(\mathbf{x} \,|\, \mathbf{x}_{0\mathbf{i}}, \pmb{\theta}_{\mathbf{i}})$ gives

eqnarray*[eqnarray* omitted — 621 chars of source]

We can further write

eqnarray*[eqnarray* omitted — 827 chars of source]

We can further write

eqnarray[eqnarray omitted — 642 chars of source]

where $\widetilde{\mathbf{x}}^\ast_{\mathbf{i},t} =I_{\mathbf{i},h} (\mathbf{x}_t)(z_t,\pmb{\sigma}( \mathbf{x}_t\, |\,\mathbf{x}_{0\mathbf{i}})^\top (\mathbf{I}_{d_q}\otimes \pmb{\gamma}^\top)^\top)^\top$ and $\mathbf{b}_{\mathbf{i}}^\ast=(a,\mathbf{b}_{\mathbf{i}}^\top)^\top$. With this notation, we can rewrite $R_{1,1}$:

eqnarray[eqnarray omitted — 855 chars of source]

where $\widetilde{\mathbf{H}}=\text{diag}(1, \mathbf{H}\mathbf{D})$, $\widetilde{\lambda}_{\mathbf{i},\min}$ denotes the smallest eigenvalue of $\frac{1}{Th^d}\sum_{t=1}^T \widetilde{\mathbf{H}} \widetilde{\mathbf{x}}^{\ast}_{\mathbf{i},t} \widetilde{\mathbf{x}}^{\ast\top}_{\mathbf{i},t}\widetilde{\mathbf{H}}^\top$, $\widetilde{\lambda}_{\min}=\min\{\widetilde{\lambda}_{\mathbf{i},\min},|, \mathbf{i}\in [M]^d\}$, and the second equality holds by the fact that $I_{\mathbf{i},h} (\mathbf{x}_t)I_{\mathbf{j},h} (\mathbf{x}_t)=0$ for $\mathbf{i}\ne \mathbf{j}$.

It suffices to explore $\frac{1}{Th^d}\sum_{t=1}^T\mathbf{H}\mathbf{D}\widetilde{\mathbf{x}}_{\mathbf{i},t} \widetilde{\mathbf{x}}^{\top}_{\mathbf{i},t}\mathbf{D}^\top\mathbf{H}$, $\frac{1}{Th^d}\sum_{t=1}^TI_{\mathbf{i},h} (\mathbf{x}_t)z_t^2$, and $\frac{1}{Th^d}\sum_{t=1}^Tz_t\widetilde{\mathbf{x}}^{\top}_{\mathbf{i},t}\mathbf{D}^\top\mathbf{H}$ before we can obtain the probability limit of $\widetilde{\lambda}_{\min}$.

We proceed with $\frac{1}{Th^d}\sum_{t=1}^TI_{\mathbf{i},h} (\mathbf{x}_t)z_t^2$. Note that

eqnarray[eqnarray omitted — 720 chars of source]

where the fourth equality can be proved by directly applying the Lipschitz continuity of $G(\mathbf{x})$ and $f_{\mathbf{x}}(\mathbf{x})$ in Assumption (ref).

Analogously, we can compute the second moment:

eqnarray[eqnarray omitted — 784 chars of source]

For the first term on the right-hand side of (ref), by (ref),

eqnarray[eqnarray omitted — 355 chars of source]

For the second term on the right-hand side of (ref),

eqnarray[eqnarray omitted — 1,008 chars of source]

where $\widetilde{\Phi}_{\eta,s}(\cdot,\cdot)$ denotes the joint CDF of $(\eta_1,\eta_{1+s})$, $f_{\mathbf{x},s}(\mathbf{x}, \mathbf{z})$ is the joint PDF of $(\mathbf{x}_1, \mathbf{x}_{1+s})$, and the last equality holds by the following results which are implied by the $\alpha$-mixing conditions in Assumption (ref),

equation[equation omitted — 310 chars of source]

Analogously, for the third term on the right-hand side of (ref), we have

eqnarray[eqnarray omitted — 306 chars of source]

Combining (ref), (ref), (ref), and (ref), we obtain

equation[equation omitted — 222 chars of source]

Together with (ref), it yields that

equation[equation omitted — 186 chars of source]

Then, we consider $\frac{1}{Th^d}\sum_{t=1}^Tz_t\widetilde{\mathbf{x}}^{\top}_{\mathbf{i},t}\mathbf{D}^\top\mathbf{H}$. Drawing upon (ref) and (ref), it follows from the definition of $\mathbf{H}$ and $\max_j|\mathbf{n}_j|=q$ that

eqnarray[eqnarray omitted — 225 chars of source]

Therefore, we only need to study the convergence of $\frac{1}{Th^d}\sum_{t=1}^TI_{\mathbf{i},h} (\mathbf{x}_t)z_t\mathbf{m}(\mathbf{x}_t\, |\, \mathbf{x}_{0\mathbf{i}})^\top \mathbf{H}$. We have

eqnarray*[eqnarray* omitted — 922 chars of source]

Using analogous arguments to those in (ref), we can show the convergence of its second moment and obtain that

eqnarray[eqnarray omitted — 339 chars of source]

Finally, we consider $\frac{1}{Th^d}\sum_{t=1}^T\mathbf{H}\mathbf{D}\widetilde{\mathbf{x}}_{\mathbf{i},t} \widetilde{\mathbf{x}}^{\top}_{\mathbf{i},t}\mathbf{D}^\top\mathbf{H}$ and write

eqnarray[eqnarray omitted — 1,095 chars of source]

where the first equality follows from Assumption (ref).1, the second equality follows from (ref), and the last step follows from Assumption (ref).2 and the definition of $\mathbf{m}(\mathbf{x}\, |\, \mathbf{0})$ given in (ref).

For the second moment, using the arguments that are analogous to the proof of (ref), we obtain

eqnarray[eqnarray omitted — 316 chars of source]

so we omit the details for now.

Together with (ref) and (ref), it proves that

equation[equation omitted — 123 chars of source]

where $\lambda_{\min,0}=\inf_{\mathbf{x}\in[-a,a]^d}\lambda_{\min}\left\{\pmb{\Omega}^\ast_0(\mathbf{x})\right\}$ with

eqnarray*[eqnarray* omitted — 587 chars of source]

Additionally, Assumption (ref) ensures that $\lambda_{\min,0}>0$.

For $R_{1,2}$, by Lemma (ref), Cauchy–Schwarz inequality, and using the analogous arguments to the derivations of $R_{1,1}$, we can obtain

eqnarray[eqnarray omitted — 523 chars of source]

For $R_{1,3}$, by (ref) and (ref), we can further write

eqnarray[eqnarray omitted — 550 chars of source]

Under Assumption (ref), it is clear to see that $\|R_{1,2}\|=o_P(d_T^2)$ and $\|R_{1,3}\|=o_P(d_T^2)$. Up to now, we have finished the investigation of $R_{1}$ and then proceed with $R_{2}$. Using simple algebra, we obtain

eqnarray[eqnarray omitted — 501 chars of source]

where $\psi_l^\ast = \psi_lH_l$, $\psi^\ast_{\max}=\max\{\psi^\ast_{l},l\in[d^0_q]\}$ and the third inequality holds by the fact $\sum_{l=1}^{d^0_q}H_l^{-1}\|\mathbf{B}_{D,l}\|\leq \sqrt{d^0_q\sum_{l=1}^{d^0_q}H_l^{-2}\|\mathbf{B}_{D,l}\|^2}=\sqrt{d^0_q}\|\mathbf{B}\|_H $. By Assumption (ref), it is straightforward to see that $\|R_2\|=o_P(d_T^2)$.

In summary of the results that are established in (ref), (ref), (ref), (ref), (ref), and (ref), it is implied by Assumption (ref) that the leading order term in $\frac{1}{d_T^2Th^d}\widetilde{Q}_\psi (\alpha_0+d_Tb_0,\widetilde{\pmb{\Lambda}}+d_T\mathbf{B})- \frac{1}{d_T^2Th^d}\widetilde{Q}_\psi (\alpha_0, \widetilde{\pmb{\Lambda}})$ is a quadratic function of $C$ with the coefficient for the quadratic term having the probability limit $\lambda_{\min,0}>0$ and the coefficients for the linear terms being bounded in probability by $O_P(1)$. Consequently, with a sufficiently large $C$, the left-hand side of (ref) is guaranteed to be positive with probability one. This completes the proof of Lemma (ref). {$\blacksquare$}

Proof of Lemma (ref):

(1). With knowledge of the true model ($g(\mathbf{x})=g_c(\mathbf{x}_c)$), we can define the following oracle LNN candidates:

eqnarray*[eqnarray* omitted — 305 chars of source]

where $\mathbf{x}_c$, $\mathbf{x}_{c,0\mathbf{i}}$, and $\pmb{\theta}_{c,\mathbf{i}}$ are oracle counterparts of $\mathbf{x}$, $\mathbf{x}_{0\mathbf{i}}$, and $\pmb{\theta}_{\mathbf{i}}$, respectively, and

eqnarray*[eqnarray* omitted — 377 chars of source]

where $\widetilde{\pmb{\lambda}}_c$, $\pmb{\gamma}_c$, and $\pmb{\pi}_{c,j}$ are the counterparts of $\widetilde{\pmb{\lambda}}$, $\pmb{\gamma}$, and $\pmb{\pi}_{j}$ in the true model. Noteworthily, we assume $\mathbf{m}_c(\mathbf{x}_c|\mathbf{x}_{c,0}) = (m_{c,1}(\mathbf{x}_c|\mathbf{x}_{c,0}),\ldots, m_{c,d_q^0}(\mathbf{x}_c|\mathbf{x}_{c,0}))^\top$ constitute a basis for the true space $\mathscr{P}^0_q$ without loss of generality, and

equation*[equation* omitted — 89 chars of source]

for $j=1,\ldots,d_q^0$.

Using arguments that are analogous to those in the proof of Lemma (ref), we can obtain

eqnarray[eqnarray omitted — 138 chars of source]

where $\widetilde{\pmb{\Lambda}}_c =\{\widetilde{\pmb{\lambda}}_{c,\mathbf{i}}\, |\, \mathbf{i}\in [M]^{d_0}\}$, $\widetilde{\pmb{\lambda}}_{c,\mathbf{i}}$ corresponds to $\pmb{\lambda}_{c,\mathbf{i}}$ up to a rotation matrix $\mathbf{D}_c$, and $\pmb{\lambda}_{c,\mathbf{i}}$ is decided by the Taylor expansion of $g_c(\mathbf{x}_c)$ at the point $\mathbf{x}_{c,0\mathbf{i}}$.

For the oracle estimators $\widetilde{\alpha}_{c}$ and $\widetilde{\pmb{\theta}}_{c,\mathbf{i}}$, it is easy to see that they satisfy the following first-order conditions:

eqnarray[eqnarray omitted — 404 chars of source]

By solving the equations in (ref), we obtain the expressions for $\widetilde{\pmb{\theta}}_{c,\mathbf{i}}$ and $\widetilde{\alpha}_{c}$:

eqnarray[eqnarray omitted — 399 chars of source]

where $\mathbf{Z}=(z_1,\cdots,z_T)^\top$, $\mathbf{Y}=(y_1,\cdots,y_T)^\top$, and $\mathbf{M}_{c,x}=\mathbf{I}_{T}-\sum_{\mathbf{i}\in [M]^{d_0}} \widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigl[\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top \widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigr]^{-1}\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top$ with $\widetilde{\mathbf{X}}_{c,\mathbf{i}} = (\widetilde{\mathbf{x}}_{c,\mathbf{i}, 1}, \cdots, \widetilde{\mathbf{x}}_{c,\mathbf{i}, T})^\top$.

Using (ref) and the expression in (ref), we can further expand $\widetilde{\alpha}_{c}$:

eqnarray*[eqnarray* omitted — 513 chars of source]

where $\pmb{\varepsilon}=(\varepsilon_1,\cdots,\varepsilon_T)^\top$. We then proceed with the convergence of $\frac{1}{T}\mathbf{Z}^\top \mathbf{M}_{c,x} \mathbf{Z}$. We write

eqnarray[eqnarray omitted — 397 chars of source]

For $Q_1$, it is clear to see that

eqnarray[eqnarray omitted — 229 chars of source]

For the second moment,

eqnarray[eqnarray omitted — 1,011 chars of source]

where $\widetilde{\Phi}_{\eta,s}(\cdot,\cdot)$ denotes the joint CDF of $(\eta_1,\eta_{1+s})$, $f_{\mathbf{x},s}(\mathbf{x}, \mathbf{z})$ is the joint PDF of $(\mathbf{x}_1, \mathbf{x}_{1+s})$, and the last equality holds by (ref), which is a standard result under the $\alpha$-mixing conditions in Assumption (ref). Combing (ref) and (ref), we obtain

equation[equation omitted — 177 chars of source]

For $Q_2$, using arguments that are analogous to those for (ref), (ref), and (ref), we can readily obtain

eqnarray[eqnarray omitted — 614 chars of source]

where $\mathbf{H}_c$ and $\mathbf{D}_c$ are counterpart matrices of $\mathbf{H}$ and $\mathbf{D}$ for the true model,

eqnarray[eqnarray omitted — 524 chars of source]

In summary of (ref), (ref), and (ref),

eqnarray[eqnarray omitted — 393 chars of source]

We then proceed with $\mathbf{Z}^\top \mathbf{M}_{c,x}\pmb{\varepsilon}$. Using arguments that are analogous to those in the proof of (ref), we can write

eqnarray[eqnarray omitted — 2,400 chars of source]

where the definitions of $\widetilde{Q}_1$, $\cdots$, $\widetilde{Q}_4$ are obvious.

For notational simplicity, let $\pmb{\xi}_{M,\mathbf{i}}=\frac{1}{Th^{d_0}}\sum_{t=1}^TI_{\mathbf{i},h} (\mathbf{x}_{c,t})z_t\mathbf{H}_c\mathbf{m}_c(\mathbf{x}_{c,t}\, |\, \mathbf{x}_{c,0\mathbf{i}}) -\mathbf{M}_{c,\mathbf{i}}$. For the first term, using the oracle counterpart of (ref), we can readily obtain that

eqnarray*[eqnarray* omitted — 341 chars of source]

To show the convergence of $\widetilde{Q}_1$, it suffices only to study $\widetilde{Q}_1^\ast$. It is clear to see that $E[\widetilde{Q}^\ast_1]=0$ under Assumption (ref) and {

eqnarray[eqnarray omitted — 1,523 chars of source]

}For $\widetilde{Q}_{1,1}^\ast$, using the arguments that are closely related to those in the proof of (ref), we can obtain

eqnarray*[eqnarray* omitted — 244 chars of source]

Therefore, $\widetilde{Q}_{1,1}^\ast=o\bigl(\frac{1}{T} \bigr)$. Analogously, for the second term

eqnarray[eqnarray omitted — 495 chars of source]

where the third inequality holds by Assumption (ref) and the Davydov's inequality for $\alpha$-mixing processes Bosq1996.

It immediately implies that the second and third terms on the right-hand side of (ref) are also asymptotically negligible. In summary of these results, we have $E[\widetilde{Q}_1^{\ast 2}]=o\bigl(\frac{1}{T} \bigr)$ and

equation[equation omitted — 80 chars of source]

Using similar arguments, we can obtain

equation[equation omitted — 137 chars of source]

Then, it suffices only to study $\widetilde{Q}_4$. For notational simplicity, we define

equation*[equation* omitted — 241 chars of source]

In what follows, we first compute the asymptotic covariance of $\frac{1}{\sqrt{T}}\sum_{t=1}^T\widetilde{z}_t\varepsilon_t$ and then employ the small-block and large-block technique for $\alpha$-mixing processes to establish its asymptotic normality.

Write

eqnarray[eqnarray omitted — 630 chars of source]

For the first term, we can use analogous arguments in the proof of (ref) and Assumption (ref) to show that

eqnarray[eqnarray omitted — 172 chars of source]

For $\widetilde{Q}_{4,2}$ and $\widetilde{Q}_{4,3}$, it is clear to see that $\widetilde{Q}_{4,2}=\widetilde{Q}_{4,3}$. It suffices only to study $\widetilde{Q}_{4,2}$. Write

{

eqnarray[eqnarray omitted — 3,072 chars of source]

}where $\sigma_{\varepsilon,s}^2=E[\varepsilon_1\varepsilon_{1+s}]$. Let $Q_0$ represent the limit of the terms on the right-hand side of (ref), as $T\rightarrow\infty$. Together with (ref) and Assumption (ref), it yields that

eqnarray*[eqnarray* omitted — 111 chars of source]

Below, we further use small-block and large-block to prove the normality. To employ the small-block and large-block arguments, we partition the set $\{1,\ldots, T \}$ into $2k_T+1$ subsets with large blocks of size $l_T$ and small blocks of size $s_T$ and the last remaining set of size $T-k_T(l_T+s_T)$, where $l_T$ and $s_T$ are selected such that

eqnarray*[eqnarray* omitted — 187 chars of source]

and $\nu$ is defined in Assumption (ref).1.

For $j=1,\ldots, k_T$, define

eqnarray*[eqnarray* omitted — 279 chars of source]

Note that $\alpha(T) = o(1/T)$ and $k_Ts_T/T\to 0$. By direct calculation, we immediately obtain that

eqnarray*[eqnarray* omitted — 157 chars of source]

Therefore,

eqnarray*[eqnarray* omitted — 139 chars of source]

By Proposition 2.6 of FanYao, we have as $T\to 0$

eqnarray*[eqnarray* omitted — 252 chars of source]

where $i$ is the imaginary unit.

In connection with (ref)-(ref), the Feller condition is fulfilled as follows:

eqnarray*[eqnarray* omitted — 101 chars of source]

Also, we note that

eqnarray[eqnarray omitted — 614 chars of source]

where the first inequality follows from H\"older inequality, the second inequality follows from Chebyshev's inequality, and the second equality follows from Minkowski inequality. Consequently,

eqnarray*[eqnarray* omitted — 188 chars of source]

where the last step follows from the choice of $l_T$ as specified above. Therefore, the Lindberg condition is justified. Using a Cram{\'e}r-Wold device, the CLT follows immediately by the standard argument:

eqnarray[eqnarray omitted — 154 chars of source]

Combing (ref) and (ref) leads to the desired result in Lemma (ref).1.

(2). We now study the asymptotic behaviour of $\widetilde{\pmb{\theta}}_{c,\mathbf{i}}$. Using its expression in (ref), we can further write

eqnarray[eqnarray omitted — 1,048 chars of source]

Using arguments that are analogous to those for (ref) and (ref), we can readily obtain

eqnarray[eqnarray omitted — 386 chars of source]

where $\pmb{\Sigma}_{c,\mathbf{i}}$ and $\mathbf{M}_{c,\mathbf{i}}$ are defined in (ref). Together with Lemma (ref).1 and (ref), these results yield

equation[equation omitted — 332 chars of source]

Drawing upon the $\alpha$-mixing conditions in Assumption (ref), we can use the arguments that are closely related to the small-block and large-block technique employed in the proof of Lemma (ref) to show that

equation*[equation* omitted — 212 chars of source]

Together with (ref), it completes the proof of Lemma (ref).2. {$\blacksquare$}

Proof of Lemma (ref):

(1). For any $l=d_q^0+1,\ldots,d_q$, if $\|\widetilde{\pmb{\Theta}}_{l}\|\neq 0$, we have the following first-order condition for the minimization problem in (ref):

eqnarray[eqnarray omitted — 394 chars of source]

where $\widetilde{Q} (\alpha,\pmb{\Theta})= \sum_{t=1}^T[y_t-z_t\alpha -\widetilde{s}(\mathbf{x}_t \, |\, \pmb{\Theta}) ]^2$.

We proceed with the derivations of $\mathbf{Q}_{l,1}$. In light of (ref), by taking first-order partial derivative of $\widetilde{Q} (\alpha,\pmb{\Theta})$ with respect to $\pmb{\theta}_{\mathbf{i},l}$, we obtain

eqnarray*[eqnarray* omitted — 609 chars of source]

where $\widetilde{x}_{\mathbf{i},tl}$ is the $l$-th element of $\widetilde{\mathbf{x}}_{\mathbf{i},t}$ and $\widetilde{\mathbf{x}}^\ast_{\mathbf{i},t} =I_{\mathbf{i},h} (\mathbf{x}_t)(z_t,\pmb{\sigma}( \mathbf{x}_t\, |\,\mathbf{x}_{0\mathbf{i}})^\top (\mathbf{I}_{d_q}\otimes \pmb{\gamma}^\top)^\top)^\top$, and $\widetilde{\pmb{\theta}}_{\mathbf{i}}^\ast=(\widetilde{a},\widetilde{\pmb{\theta}}_{\mathbf{i}}^ \top)^\top$.

Then, simple algebra gives

eqnarray[eqnarray omitted — 908 chars of source]

where $\widetilde{\pmb{\lambda}}_{\mathbf{i}}^\ast=(a_0,\widetilde{\pmb{\lambda}}_{\mathbf{i}}^ \top)^\top$.

It suffices only to study the first three terms on the right-hand side of (ref) to derive the convergence rate of $\mathbf{Q}_{l,1}$. Using (ref), (ref) and Lemma (ref), we can readily obtain that the first term has the probability order of $O_P(TH_l^{-2})$. For the second term, by Lemma (ref) and (ref), it has the probability order of $O_P(T^2h^{d+2p}H_l^{-2})$. Additionally, we can use analogous arguments in (ref) and show that the third term in (ref) is bounded in probability by $O_P(TH_l^{-2})$. In summary, we can establish the following result for $\mathbf{Q}_{l,1}$:

eqnarray[eqnarray omitted — 69 chars of source]

For $\mathbf{Q}_{l,2}$, it is straightforward to see $\|\mathbf{Q}_{l,2}\|=\psi_l\mathtt{c}$. Under the condition $\min_{d_q^0+1\leq l\leq d_q}\{\psi_lH_l\}T^{-\frac{1}{2}}\rightarrow \infty$, (ref) yields that

equation*[equation* omitted — 75 chars of source]

which leads to a contradictory result to that in (ref). Therefore, we must have $P\left(\|\widetilde{\pmb{\Theta}}_{D,l}\|= 0\right)\rightarrow 1$. It completes the proof of Lemma (ref).1.

(2). Together with Lemma (ref) and the fact that $g(\mathbf{x})=g_c(\mathbf{x}_c)$, (ref) yields

eqnarray[eqnarray omitted — 177 chars of source]

Recall that $\widetilde{Q} (\alpha,\pmb{\Theta})= \sum_{t=1}^T[y_t-z_t\alpha -\widetilde{s}(\mathbf{x}_t \, |\, \pmb{\Theta}) ]^2$. Simple algebra further gives $ \widetilde{s}_{c}(\mathbf{x}_c \, |\, \pmb{\Theta}_c)=\sum_{\mathbf{i}\in [M]^{d_0}} \widetilde{\mathbf{x}}^\top_{c,\mathbf{i},t}\pmb{\theta}_{c,\mathbf{i}}$, where $\widetilde{\mathbf{x}}_{c,\mathbf{i},t} =I_{\mathbf{i},h} (\mathbf{x}_{c,t})(\mathbf{I}_{d^0_q}\otimes \pmb{\gamma}_c^\top) \pmb{\sigma}_c( \mathbf{x}_{c,t}\, |\,\mathbf{x}_{c, 0\mathbf{i}})$.

For the group-LASSO estimators, we can formulate the following first-order conditions:

eqnarray[eqnarray omitted — 473 chars of source]

where $\pmb{\Phi}_c=\mathbf{D}^{-1}\pmb{\Phi}_D\mathbf{D}^{\top,-1}$ with $\pmb{\Phi}_D$ being a $d_q\times d_q$ diagonal matrix with its $l$-th diagonal element being $\psi_l/(2\|\widetilde{\pmb{\Theta}}_{D,l}\|)$, for $l=1,\ldots,d_q$.

Recall that $\widetilde{\pmb{\theta}}_{\mathbf{i},\star}$ contains the first $d_q^0$ elements in $\mathbf{D}^{\top,-1}\widetilde{\pmb{\theta}}_{\mathbf{i}}$. Let $\mathbf{D}_\star$ be the matrix that contains the first $d_q^0$ rows in $\mathbf{D}$ and $\widetilde{\mathbf{x}}_{\mathbf{i}, t,\star} = \mathbf{D}_\star\widetilde{\mathbf{x}}_{\mathbf{i}, t}$. Then, using the sparsity in $g(\mathbf{x})$, Lemma (ref).1, (ref), and (ref), we can rewrite (ref) as

eqnarray[eqnarray omitted — 580 chars of source]

where $\pmb{\Phi}_{D,\star}$ is a $d_q^0\times d_q^0$ diagonal matrix that contains the first $d_q^0$ diagonal elements of $\pmb{\Phi}_D$ and $\mathbf{H}_c$, as defined at an early stage, consists of the first $d_q^0$ elements in $\mathbf{H}$.

Solving (ref), we can establish the following results for $(\widetilde{\alpha}, \widetilde{\pmb{\theta}}_{\mathbf{i},\star})$:

eqnarray[eqnarray omitted — 509 chars of source]

where $\mathbf{M}_{x,\phi,\star}=\mathbf{I}_{T}-\sum_{\mathbf{i}\in [M]^{d_0}} \widetilde{\mathbf{X}}_{\mathbf{i},\star}\left[\widetilde{\mathbf{X}}_{\mathbf{i},\star}^\top \widetilde{\mathbf{X}}_{\mathbf{i},\star}+\pmb{\Phi}_{D,\star}\right]^{-1}\widetilde{\mathbf{X}}_{\mathbf{i},\star}^\top$ with $\widetilde{\mathbf{X}}_{\mathbf{i},\star} = (\widetilde{\mathbf{x}}_{\mathbf{i}, 1,\star}, \cdots, \widetilde{\mathbf{x}}_{\mathbf{i}, T,\star})^\top$.

In light of (ref) and (ref), it suffices only to study the asymptotic behaviour of $\mathbf{M}_{x,\phi,\star}$ before we can establish Lemma (ref).2. Define $\mathbf{M}_{c,x,\phi}=\mathbf{I}_{T}-\sum_{\mathbf{i}\in [M]^{d_0}} \widetilde{\mathbf{X}}_{c,\mathbf{i}}\left[\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top \widetilde{\mathbf{X}}_{c,\mathbf{i}}+\pmb{\Phi}_{D,\star}\right]^{-1}\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top$.

We write

eqnarray*[eqnarray* omitted — 861 chars of source]

Together with (ref), it yields that

eqnarray[eqnarray omitted — 1,199 chars of source]

where the last equality is ensured by the fact that $\{\psi_lH_l\}T^{-\frac{1}{2}}\rightarrow0$ and $h^{d_0}H^{-1}_l\|\widetilde{\pmb{\Theta}}_l\|\geq c$ hold uniformly for $l\in[d^0_q]$ and a positive constant $c$ under Lemma (ref) and the conditions in Assumption (ref).

In light of the definition of $\widetilde{\mathbf{x}}_{\mathbf{i}, t,\star} $ and $\widetilde{\mathbf{x}}_{c,\mathbf{i}, t}$, analogously to (ref) and (ref), we can obtain

eqnarray[eqnarray omitted — 163 chars of source]

Therefore, it is clear to see that $\frac{1}{T}\mathbf{Z}^\top\left(\mathbf{M}_{x,\phi,\star}-\mathbf{M}_{c,x,\phi}\right)\mathbf{Z}=O_P(h^{q+1})$. Together with (ref), it implies that $\frac{1}{T}\mathbf{Z}^\top\left(\mathbf{M}_{x,\phi,\star}-\mathbf{M}_{c,x}\right)\mathbf{Z}$ has the probability order of $o_P\left(\frac{1}{\sqrt{T}}+h^p\right)$. Analogously, we can show that $\frac{1}{T}\mathbf{Z}^\top\left(\mathbf{M}_{x,\phi,\star}-\mathbf{M}_{c,x}\right)\mathbf{Y}$ is also negligible. Therefore, we have

equation[equation omitted — 114 chars of source]

Therefore, the proof of Lemma (ref).2 is complete. {$\blacksquare$}

Proof of Theorem (ref):

(1). Directly applying Lemma (ref) and Lemma (ref), we can establish the desired result in Theorem (ref).1.

(2). Similarly to (ref) and (ref), we have

eqnarray[eqnarray omitted — 220 chars of source]

where $\widetilde{\mathbf{x}}_{c,\mathbf{i},0} =I_{\mathbf{i},h} (\mathbf{x}_{c,0})(\mathbf{I}_{d^0_q}\otimes \pmb{\gamma}_c^\top) \pmb{\sigma}_c( \mathbf{x}_{c,0}\, |\,\mathbf{x}_{c,0\mathbf{i}})$.

Together with Lemma (ref).1, (ref), (ref), and (ref), it gives

eqnarray[eqnarray omitted — 336 chars of source]

Combining (ref) and (ref), we obtain

eqnarray*[eqnarray* omitted — 410 chars of source]

We now study $\widetilde{\pmb{\theta}}_{\mathbf{i},\star}$. With (ref), similarly to (ref), we can further write

eqnarray[eqnarray omitted — 901 chars of source]

Drawing upon (ref) and the expressions for $\widetilde{\pmb{\theta}}_{c,\mathbf{i}}$ and $\widetilde{\pmb{\theta}}_{\mathbf{i},\star}$ that we have derived in (ref) and (ref), respectively,

eqnarray[eqnarray omitted — 1,189 chars of source]

Using the results that are established in (ref) and (ref), we can readily obtain the orders of the first three terms on the right-hand side of (ref) as $o_P\left(\frac{1}{\sqrt{T}}\right)$, $o_P\left(\frac{1}{T}\right)$ and $o_P\left(\frac{1}{\sqrt{T}}\right)$. Therefore, we have

equation[equation omitted — 214 chars of source]

Using the central limit theorem for $\widetilde{\pmb{\theta}}_{c,\mathbf{i}}$ in Lemma (ref), we can establish the asymptotic normality of $\sqrt{Th^{d_0}}(\widetilde{g}(\mathbf{x}_{0})-g_c(\mathbf{x}_{c,0}))$, which has the asymptotic covariance as the limit of

eqnarray*[eqnarray* omitted — 334 chars of source]

where $\pmb{\Sigma}_{c,\mathbf{i}}$ is defined in Lemma (ref). Then, the proof of Theorem (ref).2 is complete. {$\blacksquare$}

Proof of Theorem (ref):

(1) Using arguments that are analogous to those in the proof of Lemma (ref), we can show that the bootstrap Group-LASSO estimators can approximate the bootstrap oracle estimators up to some asymptotically negligible bias terms. Therefore, it suffices only to study the asymptotic behaviour of the bootstrap oracle estimators. Define the bootstrap oracle estimators $ (\widetilde{\alpha}^\ast_c,\widetilde{\pmb{\Theta}}^\ast_c)$ as follows:

equation*[equation* omitted — 210 chars of source]

where $\widetilde{Q}^\ast_c (\alpha,\pmb{\Theta}_c)= \sum_{t=1}^T[y^\ast_t-z_t\alpha -\widetilde{s}_c(\mathbf{x}_{c,t} \, |\, \pmb{\Theta}_c) ]^2$.

By solving the first-order conditions, we obtain the following expressions for $\widetilde{\pmb{\theta}}^\ast_{c,\mathbf{i}}$ and $\widetilde{\alpha}^\ast_{c}$:

eqnarray[eqnarray omitted — 425 chars of source]

where $\mathbf{Z}=(z_1,\cdots,z_T)^\top$, $\mathbf{Y}^\ast=(y^\ast_1,\cdots,y^\ast_T)^\top$, and $\mathbf{M}_{c,x}=\mathbf{I}_{T}-\sum_{\mathbf{i}\in [M]^{d_0}} \widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigl[\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top \widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigr]^{-1}\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top$ with $\widetilde{\mathbf{X}}_{c,\mathbf{i}} = (\widetilde{\mathbf{x}}_{c,\mathbf{i}, 1}, \cdots, \widetilde{\mathbf{x}}_{c,\mathbf{i}, T})^\top$.

Drawing upon the DGP of $y_t^\ast$ and the second expression in (ref), we can further expand $\widetilde{\alpha}^\ast_{c}$ as follows:

eqnarray[eqnarray omitted — 389 chars of source]

where $\widetilde{\mathbf{X}}_{\mathbf{i}} = (\widetilde{\mathbf{x}}_{\mathbf{i}, 1}, \cdots, \widetilde{\mathbf{x}}_{\mathbf{i}, T})^\top$ and $\pmb{\varepsilon}^\ast=(\varepsilon_1^\ast,\cdots,\varepsilon_T^\ast)^\top$.

For the first term, it is clear to see that

eqnarray*[eqnarray* omitted — 413 chars of source]

Together with Lemma (ref).1 and (ref), it immediately yields that the first term in (ref) is asymptotically negligible.

Recall that $\widetilde{z}_t=z_t-\sum_{\mathbf{i}\in [M]^{d_0}} I_{\mathbf{i},h} (\mathbf{x}_{c,t}) \mathbf{M}_{c,\mathbf{i}}^\top \pmb{\Sigma}_{c,\mathbf{i}}^{-1}\mathbf{H}_c\mathbf{m}_c(\mathbf{x}_{c,t}\, |\, \mathbf{x}_{c,0\mathbf{i}})$. For notational simplicity, we further define $\widetilde{z}_t^\ast = z_t-\sum_{\mathbf{i}\in [M]^{d_0}} \mathbf{Z}^\top\widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigl[\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top \widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigr]^{-1}\widetilde{\mathbf{x}}_{c,\mathbf{i},t}$. For the second term on the right-hand side of (ref), we write

eqnarray[eqnarray omitted — 657 chars of source]

Using arguments that are analogous to those in the proofs of (ref) and (ref), we can show that $\mathcal{J}_1=o_P(1)$. For $\mathcal{J}_2$, write

eqnarray[eqnarray omitted — 289 chars of source]

For $\mathcal{J}_{2,1}$, it is straightforward to see that $\frac{1}{\sqrt{T}}\sum_{t=1}^TE[\widetilde{z}_tz_t\varsigma_t]=0$ and

eqnarray*[eqnarray* omitted — 516 chars of source]

It is obvious that the first term has the order $O(1)$ and the second and third terms have the same order. Therefore, it suffices only to study the second term. We have

eqnarray*[eqnarray* omitted — 522 chars of source]

In summary of these results, we have

eqnarray[eqnarray omitted — 125 chars of source]

with implies that $\frac{1}{\sqrt{T}}\sum_{t=1}^T\widetilde{z}_tz_t\varsigma_t=O_P(\sqrt{\ell})$. Together with Theorem (ref).1, it yields that

eqnarray[eqnarray omitted — 188 chars of source]

We now proceed with the derivations of $\mathcal{J}_{2,2}$. By Lemma (ref).1 and (ref), we obtain

eqnarray*[eqnarray* omitted — 328 chars of source]

Similarly to (ref), we can show that

equation*[equation* omitted — 175 chars of source]

Additionally, directly applying Lemma (ref).2, (ref), and Cauchy-Schwarz inequality yields

eqnarray*[eqnarray* omitted — 814 chars of source]

Therefore, we have

equation[equation omitted — 113 chars of source]

Combining (ref) and (ref), we obtain

eqnarray*[eqnarray* omitted — 40 chars of source]

Analogously, we can show that $\mathcal{J}_{3}$ is also asymptotically negligible.

In what follows, we proceed to explore $\mathcal{J}_{4}$ which generates the bootstrap distribution. Let $E^\ast[\cdot]$ and $\text{Var}^\ast(\cdot)$ denote the expectation and variance conditional on the observed sample. We first show that $\text{Var}^\ast( \mathcal{J}_{4})=\sigma_{c,z}^2+o_P(1)$.

It is clear to see that $ E^\ast\bigl[\mathcal{J}_{4}\bigr]=0$. Moreover, we write

eqnarray*[eqnarray* omitted — 694 chars of source]

Using a decomposition that is similar to (ref), we can easily show that $\mathcal{Q}_1=E[\widetilde{z}^2_1 \varepsilon^2_1]+o_P(1)$. Let $s_T$ satisfy that $s_T\rightarrow\infty$ and $s_T^2/\ell\rightarrow0$. For $\mathcal{Q}_2$, we write

eqnarray[eqnarray omitted — 856 chars of source]

For $\mathcal{Q}_{2,1}$, by Davydov's inequality for $\alpha$-mixing processes,

eqnarray[eqnarray omitted — 376 chars of source]

Together with Lipschitz continuity of the kernel function, it yields that

eqnarray[eqnarray omitted — 166 chars of source]

For $\mathcal{Q}_{2,2}$, by (ref), we have

eqnarray[eqnarray omitted — 232 chars of source]

The second equality holds by the fact that $\sum_{t=1}^{T-1} \alpha(t)^{\nu/(2+\nu)}$ and $\sum_{t=1}^{s_T} \alpha(t)^{\nu/(2+\nu)}$ have the same limit as $T\rightarrow\infty$ and $s_T\rightarrow\infty$, which is ensured by the order of $\alpha$-mixing coefficients.

In summary of the results that are established in (ref), (ref), and (ref), we can readily obtain

equation[equation omitted — 151 chars of source]

We then study $\mathcal{Q}_2-E[\mathcal{Q}_2]$. With $\nu$ and $\nu^\ast$ that are defined in Assumption (ref) and Theorem (ref), respectively, we can always define a positive number $r$ through the following equation:

equation*[equation* omitted — 71 chars of source]

Since $\nu^\ast>\nu$, it is clear to see that

eqnarray*[eqnarray* omitted — 82 chars of source]

Thus, we have $r>2$. For notational simplicity, we define a norm $\|\zeta\|_{n}=E\bigl[\|\zeta\|^{n}\bigr]^{1/n}$ for any random variable $\zeta$ and any positive number $n\geq 1$. We have

eqnarray[eqnarray omitted — 834 chars of source]

Additionally, let $\mathcal{F}_t$ and $E_{t}[\cdot]$ be the sigma field generated by $\{\widetilde{z}_s, \varepsilon_s\}_{s=t,t-1,\cdots}$ and the expectation conditional on $\mathcal{F}_t$, respectively. By McLeish's inequality for $\alpha$-mixing processes McLeish1975 and Assumption (ref),

eqnarray*[eqnarray* omitted — 730 chars of source]

for a positive integer $t_0$. Using this result and the Lemma A of Hansen1992, we can readily obtain

eqnarray[eqnarray omitted — 341 chars of source]

where $c_{\nu^\ast}=12\bigl\|\widetilde{z}_1 \bigr\|_{2+\nu^\ast} \bigl\|\varepsilon_1\bigr\|_{2+\nu^\ast}$ and $r^\ast=\min(r,4)$.

Combing (ref) and (ref) gives

eqnarray*[eqnarray* omitted — 166 chars of source]

Under the condition $\ell T^{2/r^\ast-1}\rightarrow 0$, we have

eqnarray[eqnarray omitted — 67 chars of source]

By (ref) and (ref), we can finish the investigation of $\mathcal{Q}_2$ and obtain

eqnarray[eqnarray omitted — 150 chars of source]

A similar argument applies with the index $s$ replacing $t$ for $\mathcal{Q}_3$, so it follows that

eqnarray[eqnarray omitted — 152 chars of source]

Drawing upon the definition of $\sigma_{c,z}^2$ in Assumption (ref) and the convergence of $\mathcal{Q}_1$, the results that are established in (ref) and (ref) can immediately yield

equation[equation omitted — 85 chars of source]

In light of the $\ell$-dependent $\varsigma_t$, we follow the Theorem 3.1 of Shao2010 and adopt the large-block and small-block argument to prove the central limit theorem for $\mathcal{J}_{4}$ conditional on the observed sample. Define $l_T$ and $s_T$ as the lengths for the large and small blocks and $k_T=\lfloor T/(l_T+s_T)\rfloor$ such that $l_T,s_T\rightarrow\infty$ and

equation*[equation* omitted — 118 chars of source]

For $j=1,\ldots, k_T$, define

eqnarray*[eqnarray* omitted — 304 chars of source]

We first show that $\frac{1}{\sqrt{T}}\sum_{j=1}^{k_T} \xi^\ast_{j,2}=o_P(1)$. Since $\frac{\ell}{s_T}\rightarrow 0$, we assume $l_T,s_T>\ell$ without loss of generality. By the definition of $\varsigma_t$, we can observe that $\{\xi^\ast_{1,1},\cdots,\xi^\ast_{k_T,1}\}$ are independent conditional on the observed data, as are $\{\xi^\ast_{1,2},\cdots,\xi^\ast_{k_T,2}\}$. Using (ref) and Assumption (ref), we have

eqnarray*[eqnarray* omitted — 660 chars of source]

Therefore, $\frac{1}{\sqrt{T}}\sum_{j=1}^{k_T} \xi^\ast_{j,2}=o_P(1)$. Analogously, we also have $\frac{1}{\sqrt{T}} \xi^\ast_{0}=o_P(1)$. Next, we establish the asymptotic normality of $\frac{1}{\sqrt{T}}\sum_{j=1}^{k_T} \xi^\ast_{j,1}$ by verifying the Lindeberg condition. Using analogous arguments to those in the proof of (ref), we can first show that $\frac{1}{T}E^\ast\bigl[\bigl(\sum_{j=1}^{k_T} \xi^\ast_{j,1}\bigr)^2 \bigr]=\sigma_{c,z}^2+o_P(1)$. Moreover, for any $\epsilon>0$, we can use the same argument as in (ref) to obtain

eqnarray[eqnarray omitted — 245 chars of source]

For notational simplicity, we define a norm (conditional on the observed sample) $\|\zeta\|^\ast_{n}=E^\ast\bigl[\|\zeta\|^{n}\bigr]^{1/n}$ for any random variable $\zeta$ and any positive number $n\geq 1$. In what follows, we use the Rosenthal inequality to study the order of $\|\xi^\ast_{1,1}\|^\ast_{2+\nu}$, without loss of generality. Noteworthily, Rosenthal inequality is designed for the independent random variables and it is not directly applicable for $\widetilde{z}_t\varepsilon_t\varsigma_t$. Therefore, we further decompose $\xi^\ast_{1,1}$ as $\xi^\ast_{1,1}=\sum_{k=1}^{\ell+1} \xi^\ast_{1,1,k}$, where

equation*[equation* omitted — 167 chars of source]

Then, by the definition of $\varsigma$, it is clear to see that for each $k$, all elements that are involved in the summation in $\xi^\ast_{1,1,k}$ are independent conditional on the observed sample. Using the triangle inequality and Rosenthal inequality sequentially, we obtain

eqnarray[eqnarray omitted — 805 chars of source]

Combining (ref) and (ref) gives

eqnarray*[eqnarray* omitted — 550 chars of source]

We therefore have $\frac{1}{T}\sum_{j=1}^{k_T}E^\ast\bigl[\xi^{\ast 2}_{j,1}\cdot I(|\xi^\ast_{j,1}|\geq \epsilon \sqrt{T}) \bigr]=o_P(1)$, which is the last step of the large-block and small-block technique and it immediately yields the following CLT: $\mathcal{J}_{4}\rightarrow_{D^\ast} N\left(0,\sigma_{c,z}^2\right)$, where $\rightarrow_{D^\ast}$ denotes the convergence in distribution conditional on the observed sample. Recall that we have shown that $\mathcal{J}_{1},\mathcal{J}_{2},$ and $\mathcal{J}_{3}$ are all asymptotically negligible at an earlier stage. Combining these results with (ref), (ref), and Theorem (ref).1 leads to the assertion in Theorem (ref).1.

(2) Let $\Omega_T$ denote the event in which the group-LASSO estimation has correctly identified the sparsity structure of the $g(\mathbf{x})$ function. That is $\|\widetilde{\pmb{\Theta}}_{D,j}\|=0$, for $j=d_q^0+1,\ldots,d_q$. By Lemma (ref), we have $P(\Omega_T)\rightarrow1$, as $T\rightarrow\infty$. Hence, it suffices only to study the asymptotic distribution of $\widetilde{g}^*(\mathbf{x}_0)$ conditional on the observed sample and $\Omega_T$.

On $\Omega_T$ and using analogous arguments to those in the proof of Theorem (ref), we can obtain the following result for the bootstrap estimator $\widetilde{g}^\ast(\mathbf{x}_{0})$:

eqnarray*[eqnarray* omitted — 339 chars of source]

where $\widetilde{\pmb{\theta}}^\ast_{c,\mathbf{i}}$ denotes the oracle bootstrap estimator.

Then, we can use the arguments that are closely related to those in the proof of (ref) to obtain

equation[equation omitted — 356 chars of source]

Then, similar steps to those in the derivation of $\frac{1}{\sqrt{T}}\mathbf{Z}^\top \mathbf{M}_{c,x}\pmb{\varepsilon}^\ast$'s bootstrap distribution in (ref) can be applied here to establish the bootstrap behaviour of the first term on the right-hand side of (ref). Specifically, we can obtain that it converges to $N\left(\mathbf{0},\sigma^2_\varepsilon \pmb{\Sigma}_{c,\mathbf{i}}^{-1}\right)$ conditional on the observed sample, up to some asymptotically negligible terms. In connection with (ref), it leads to the desired result in Theorem (ref).2. {$\blacksquare$}

Proofs for the Fully Nonparametric Model

Proof of Lemma (ref):

First, we expand the expression of $\widehat{\pmb{\theta}}_{\mathbf{i}}$ as follows:

eqnarray*[eqnarray* omitted — 1,510 chars of source]

where the third equality follows from the fact that $I_{\mathbf{i}, h}(\mathbf{x}_t)I_{\mathbf{j}, h}(\mathbf{x}_t)=0$ for $\mathbf{i}\ne \mathbf{j}$, and the definition of $\widetilde{s}(\mathbf{x}_t \, |\, \widetilde{\pmb{\Lambda}})$. Below, we consider the terms on the right-hand side one by one.

We further define

eqnarray[eqnarray omitted — 357 chars of source]

First, we consider $\frac{1}{Th^d} \sum_{t=1}^T\widetilde{\mathbf{x}}_{\mathbf{i},t} \widetilde{\mathbf{x}}_{\mathbf{i},t} ^\top$. By (ref) and (ref), we obtain

eqnarray*[eqnarray* omitted — 363 chars of source]

Also, we note

eqnarray*[eqnarray* omitted — 881 chars of source]

where

eqnarray*[eqnarray* omitted — 410 chars of source]

By Lemma (ref), it is easy to see that

eqnarray*[eqnarray* omitted — 210 chars of source]

Then we can write

eqnarray*[eqnarray* omitted — 1,642 chars of source]

where the first inequality follows from the exercise 5 on page 267 of Magnus, and the second inequality follows from (ref) and (ref).

Based on the above development, we can conclude that

eqnarray*[eqnarray* omitted — 327 chars of source]

Finally, in order to establish the asymptotic distribution, we just need to focus on $\frac{1}{\sqrt{Th^d}}\sum_{t=1}^T\widetilde{\mathbf{m}}(\mathbf{x}_t\, |\, \mathbf{x}_{\mathbf{i},0}) \varepsilon_t$ in view of (ref). Write

eqnarray[eqnarray omitted — 1,240 chars of source]

where the last two terms are the same up to a transpose operation.

For the first term, we can use similar argument to that in (ref) and obtain

eqnarray[eqnarray omitted — 413 chars of source]

Then, we use Assumption (ref) and the Davydov's inequality for $\alpha$-mixing processes (see pages 19-20 in Bosq1996) to show the convergence of the second term on the right-hand side of (ref). Specifically, we have

eqnarray[eqnarray omitted — 198 chars of source]

Moreover, it is clear to see that

eqnarray[eqnarray omitted — 800 chars of source]

By (ref) and (ref),

eqnarray[eqnarray omitted — 341 chars of source]

Analogously, the third term on the right-hand side of (ref) is also negligible.

Thus, we can conclude that

eqnarray[eqnarray omitted — 742 chars of source]

Below, we further use small-block and large-block to prove the normality. To employ the small-block and large-block arguments, we partition the set $\{1,\ldots, T \}$ into $2k_T+1$ subsets with large blocks of size $l_T$ and small blocks of size $s_T$ and the last remaining set of size $T-k_T(l_T+s_T)$, where $l_T$ and $s_T$ are selected such that

eqnarray*[eqnarray* omitted — 192 chars of source]

and $\nu$ is defined in Assumption (ref).1.

For $j=1,\ldots, k_T$, define

eqnarray*[eqnarray* omitted — 508 chars of source]

Note that $\alpha(T) = o(1/T)$ and $k_Ts_T/T\to 0$. By direct calculation, we immediately obtain that

eqnarray*[eqnarray* omitted — 157 chars of source]

Therefore,

eqnarray*[eqnarray* omitted — 195 chars of source]

By Proposition 2.6 of FanYao, we have as $T\to 0$

eqnarray*[eqnarray* omitted — 252 chars of source]

where $i$ is the imaginary unit.

In connection with (ref)-(ref), the Feller condition is fulfilled as follows:

eqnarray*[eqnarray* omitted — 276 chars of source]

Also, we note that

eqnarray*[eqnarray* omitted — 1,056 chars of source]

where the first inequality follows from H\"older inequality, the second inequality follows from Chebyshev's inequality, and the second equality follows from Minkowski inequality. Consequently,

eqnarray*[eqnarray* omitted — 254 chars of source]

where the last step follows from the choice of $l_T$ as specified above. Therefore, the Lindberg condition is justified. Using a Cram{\'e}r-Wold device, the CLT follows immediately by the standard argument. {$\blacksquare$}

Proof of Theorem (ref):

By Lemma (ref), we can write

eqnarray[eqnarray omitted — 1,178 chars of source]

where $\widetilde{\mathbf{x}}_{\mathbf{i},0} =I_{\mathbf{i},h} (\mathbf{x}_0)(\mathbf{I}_{d_q}\otimes \pmb{\gamma}^\top) \pmb{\sigma}( \mathbf{x}_0\, |\,\mathbf{x}_{0\mathbf{i}})$.

By (ref) and (ref), we obtain

eqnarray*[eqnarray* omitted — 270 chars of source]

which immediately yields

eqnarray*[eqnarray* omitted — 194 chars of source]

Together with Lemma (ref) and (ref), it implies that

eqnarray[eqnarray omitted — 512 chars of source]

Therefore, Lemma(ref) implies that $\sqrt{Th^d}(\widehat{g}(\mathbf{x}_0)-g(\mathbf{x}_0))$ is asymptotically normal with the asymptotic covariance being the limit of

eqnarray*[eqnarray* omitted — 298 chars of source]

where $ \pmb{\Sigma}_{\mathbf{i}} =f_{\mathbf{x}}(\mathbf{x}_{0\mathbf{i}}) \int_{[-1,1]^d} \mathbf{m}(\mathbf{x}\, |\, \mathbf{0}) \mathbf{m}(\mathbf{x}\, |\, \mathbf{0})^\top \mathrm{d} \mathbf{x}$.{$\blacksquare$}

Proof of Theorem (ref):

In what follows, we label the quantities associated with the bootstrap procedure by the superscript $^*$, which will not be further explained unless misunderstanding may arise.

By design, we have for $\forall \mathbf{i}$

eqnarray*[eqnarray* omitted — 1,656 chars of source]

where the second equality follows from the definition of $y_t^*$.

Note that the term $\mathbf{A}_{\mathbf{i,2}}$ has been investigated in the proof of Lemma (ref), and is negligible under the condition $\sqrt{T}h^{p+d/2}\to 0$. For the term $\mathbf{A}_{\mathbf{i,3}}$, we can further write

eqnarray*[eqnarray* omitted — 427 chars of source]

where the second equality follows from the fact that $\{\eta_t \}$ are i.i.d. draws from $N(0,1)$ and are independent of the sample.

Therefore, we only need to pay attention to $\mathbf{A}_{\mathbf{i},1}$ below. It suffices to consider $\frac{1}{\sqrt{T}}\sum_{t=1}^T\mathbf{\pmb\xi}_t^*$, where $\mathbf{\pmb\xi}_t^* =\frac{1}{\sqrt{h^d}}\widetilde{\mathbf{x}}_{\mathbf{i},t} \varepsilon_t \eta_t$. As $\{\eta_t \}$ are i.i.d. draws from $N(0,1)$, it is easy to know that

eqnarray*[eqnarray* omitted — 223 chars of source]

in view of the proof of Lemma (ref).

Below, we consider $E^*[\|\pmb{\xi}_1^*\|^2 \cdot I(\|\pmb{\xi}_1^*\| \ge \epsilon \sqrt{T})]$. Write

eqnarray*[eqnarray* omitted — 1,028 chars of source]

where the first inequality follows from H\"older inequality, the second inequality follows from Chebyshev's inequality, the third inequality follows from the definition of $E^*$ and Minkowski inequality, and the last step follows from $\frac{1}{h^d}E \| \widetilde{\mathbf{m}}(\mathbf{x}_1\, |\, \mathbf{x}_{\mathbf{i},0}) \varepsilon_1 \|^{2+\nu}=O(1)$ by the proof of Lemma (ref). Consequently,

eqnarray*[eqnarray* omitted — 166 chars of source]

where the last step follows from $Th^d\to 0$. Therefore, the Lindberg condition is justified. Then the result follows. {$\blacksquare$}

Proof of Corollary (ref):

The proof is a simpler version of that presented for Lemma (ref) and Theorem (ref), therefore it is omitted. {$\blacksquare$}

Proofs for the Binary Model

Proof of Lemma (ref):

(1). First, note that provided $0<x, x_0<1$, we have the following two expressions by the following Taylor expansions:

eqnarray[eqnarray omitted — 214 chars of source]

where both $x^*$ and $x^\dagger$ lie between $x$ and $x_0$.

We are now ready to start our investigation. By (ref) and (ref), write

eqnarray[eqnarray omitted — 1,670 chars of source]

where both $\Phi_t^*$ and $\Phi_t^\dagger$ lie between $\Phi_\eta(\widetilde{s}(\mathbf{x}_t\, | \, \pmb{\Theta}))$ and $\Phi_\eta(g(\mathbf{x}_t))$, and the definitions of $\mathbb{L}_{T,1}$ and $\mathbb{L}_{T,2}$ are obvious.

We then consider $\mathbb{L}_{T,1}$ and $\mathbb{L}_{T,2}$ respectively, and start with $\mathbb{L}_{T,1}$. For notational simplicity, we let

eqnarray*[eqnarray* omitted — 248 chars of source]

Simple algebra shows that

eqnarray[eqnarray omitted — 230 chars of source]

For any given $\Delta \Phi_\eta(g(\cdot))$, we then consider

eqnarray[eqnarray omitted — 626 chars of source]

where the inequality follows from Assumption (ref).1, and Davydov's inequality and (ref). By Lemmas A1 and A2 of WP2003, we immediately obtain that

eqnarray[eqnarray omitted — 92 chars of source]

We next investigate $\mathbb{L}_{T,2}$. Write

eqnarray*[eqnarray* omitted — 605 chars of source]

where the first inequality follows from $\frac{1}{(a+b)^2} \ge \frac{1}{2a^2 + 2b^2}$ because of $(a+b)^2\le 2a^2 + 2b^2$, the second inequality follows from the fact that $\Phi_t^*$ and $\Phi_t^\dagger$ lie between $\Phi_\eta(\widetilde{s}(\mathbf{x}_t\, | \, \pmb{\Theta}))$ and $\Phi_\eta(g(\mathbf{x}_t))$, and the third inequality follows from that $ \frac{1-y_t}{4\cdot 2} + \frac{y_t}{2} \ge \frac{1}{8}$ because of $z_t$ taking the value of 1 or 0 only.

By the fact that $0\ge\frac{1}{T}\log L(g(\cdot)) - \frac{1}{T}\log L(\widetilde{s}(\cdot\, | \, \widehat{\pmb{\Theta}}))$, and (ref) and (ref), we now conclude that

eqnarray*[eqnarray* omitted — 261 chars of source]

which completes the proof of this lemma. {$\blacksquare$}

By Lemma (ref), it is obvious that

eqnarray*[eqnarray* omitted — 650 chars of source]

where the third equality follows from the first result of Lemma (ref) and Lemma (ref).

Proof of Lemma (ref):

(1). Note that by Lemma (ref), we can write

eqnarray*[eqnarray* omitted — 543 chars of source]

where the third equality follows from the fact that $\mathbf{x}_t$ can not simultaneous belong to $C_{\mathbf{x}_{0\mathbf{i}},h}$ and $C_{\mathbf{x}_{0\mathbf{j}},h}$ for $\mathbf{i}\ne \mathbf{j}$. Thus, we must have

eqnarray*[eqnarray* omitted — 190 chars of source]

which completes the proof of the first result.

(2). By (ref), we denote

eqnarray*[eqnarray* omitted — 430 chars of source]

where the definitions of $ \mathbf{L}_j(\cdot)$ for $j=1,2,3$ should be obvious.

First, we consider $\mathbf{L}_2(s(\cdot \, | \, \mathbf{x}_{\mathbf{i},0},\pmb{\theta}_{\mathbf{i}} ))$ and $\mathbf{L}_3(s(\cdot \, | \, \mathbf{x}_{\mathbf{i},0},\pmb{\theta}_{\mathbf{i}} ))$. For $\mathbf{L}_2(s(\cdot \, | \, \mathbf{x}_{\mathbf{i},0},\pmb{\theta}_{\mathbf{i}} ))$, we write

eqnarray[eqnarray omitted — 1,146 chars of source]

where the definition of $ \widetilde{\mathbf{X}}_{\mathbf{i}, t} (\cdot) $ is obvious.

For the first term on the right hand side of (ref), we write

eqnarray*[eqnarray* omitted — 1,335 chars of source]

where the first inequality follows from Cauchy-Schwarz inequality, the second inequality follows from Mean-Value Theorem and the fact that $\phi_\eta(\cdot)$ is uniformly bounded, and the last step follows from the first result of this lemma and (ref).

For the second term on the right hand side of (ref), by some tedious algebra and the first result of this lemma, it is not hard to see that

eqnarray*[eqnarray* omitted — 446 chars of source]

Further, using Assumption (ref) and Billingsley's inequality following a procedure similar (but simplified) as in (ref), and in connection with (ref), we can show that

eqnarray*[eqnarray* omitted — 208 chars of source]

Based on the above development, we are readily to conclude that

eqnarray*[eqnarray* omitted — 155 chars of source]

Similar to the analysis of $ \widetilde{\mathbf{L}}_2(s(\mathbf{x}_t\, | \, \mathbf{x}_{\mathbf{i},0}, \widehat{\pmb{\theta}}_{\mathbf{i}} ))$, we can also obtain that

eqnarray*[eqnarray* omitted — 189 chars of source]

Below, we focus on $\mathbf{L}_1(s(\mathbf{x}_t\, | \, \mathbf{x}_{\mathbf{i},0}, \widehat{\pmb{\theta}}_{\mathbf{i}} ))$, and write

eqnarray*[eqnarray* omitted — 1,229 chars of source]

where the second equality follows from similar steps as those for $\widetilde{\mathbf{L}}_2(s(\mathbf{x}_t\, | \, \mathbf{x}_{\mathbf{i},0}, \widehat{\pmb{\theta}}_{\mathbf{i}} )) $, the third equality follows from a proof similar to those for (ref), and the fourth equality follows from a development similar to (ref).

Thus, we can now conclude that for each $\mathbf{i}$

eqnarray*[eqnarray* omitted — 289 chars of source]

where $\widetilde{\pmb{\Sigma}}_{\mathbf{i}}$ is defined in the body of this lemma.

The proof of the second result is now completed.

(3). In view of the fact that $\pmb{\theta}_{\mathbf{i}}^*$ lies between $\widehat{\pmb{\theta}}_{\mathbf{i}}$ and $ \widetilde{\pmb{\lambda}}_{\mathbf{i}}$, the result follows immediately by going through the same procedure as the second result of this lemma.

(4). Write

eqnarray*[eqnarray* omitted — 806 chars of source]

where $\widetilde{\mathbf{m} }(\mathbf{x}_t\, |\, \mathbf{x}_{\mathbf{i},0})$ is defined in (ref), the second equality follows from (ref), and in the third equality we let

eqnarray*[eqnarray* omitted — 165 chars of source]

for notational simplicity. Moreover, as $g$ is defined on $[-a, a]$, it is easy to know that $ 0<\mathtt{c}\le \Phi_\eta(g(\mathbf{x}_t))\le \mathtt{C}<1$. Thus, we can further write

eqnarray*[eqnarray* omitted — 162 chars of source]

Also, simple algebra shows that

eqnarray*[eqnarray* omitted — 375 chars of source]

We now move on and write

eqnarray*[eqnarray* omitted — 721 chars of source]

Similar to (ref), we have

eqnarray*[eqnarray* omitted — 233 chars of source]

Thus , we can conclude that

eqnarray*[eqnarray* omitted — 909 chars of source]

where the last step follows from a procedure similar to (ref).

Below, we further use small-block and large-block to prove the normality. To employ the small-block and large-block arguments, we partition the set $\{1,\ldots, T \}$ into $2k_T+1$ subsets with large blocks of size $l_T$ and small blocks of size $s_T$ and the last remaining set of size $T-k_T(l_T+s_T)$, where $l_T$ and $s_T$ are selected such that

eqnarray*[eqnarray* omitted — 169 chars of source]

where $\nu$ is defined in Assumption (ref).1

For $j=1,\ldots, k_T$, define

eqnarray*[eqnarray* omitted — 479 chars of source]

Note that $\alpha(T) = o(1/T)$ and $k_Ts_T/T\to 0$. By direct calculation, we immediately obtain that

eqnarray*[eqnarray* omitted — 154 chars of source]

Therefore,

eqnarray*[eqnarray* omitted — 185 chars of source]

By Proposition 2.6 of FanYao, we have as $T\to 0$

eqnarray*[eqnarray* omitted — 254 chars of source]

where $i$ is the imaginary unit. Thus, the Feller condition is fulfilled as follows:

eqnarray*[eqnarray* omitted — 398 chars of source]

Also, we note that

eqnarray*[eqnarray* omitted — 510 chars of source]

where the first inequality follows from H\"older inequality, the second inequality follows from Chebyshev's inequality, and the last step follows from Minkowski inequality. Consequently,

eqnarray*[eqnarray* omitted — 215 chars of source]

which is the Lindberg condition. Using a Cram{\'e}r-Wold device, the CLT follows immediately by the standard argument. {$\blacksquare$}

Proof of Theorem (ref):

By the first order condition, we have

eqnarray*[eqnarray* omitted — 277 chars of source]

where the second equality follows from (ref).

Using Taylor expansion, we have

eqnarray*[eqnarray* omitted — 425 chars of source]

where $\pmb{\theta}_{\mathbf{i}}^*$ lies between $\widehat{\pmb{\theta}}_{\mathbf{i}}$ and $ \widetilde{\pmb{\lambda}}_{\mathbf{i}}$, and the second equality follows from Lemma (ref) and the continuity of $\phi_\eta$ and $\Phi_\eta$.

Thus, by Lemma (ref), the result follows immediately. {$\blacksquare$}

Proof of Theorem (ref):

The proof follows from (ref) and (ref) and a procedure very similar to that given in Theorem (ref). {$\blacksquare$}

figure[figure omitted — 170 chars of source]
figure[figure omitted — 216 chars of source]
figure[figure omitted — 214 chars of source]
table[table omitted — 1,014 chars of source]
figure[figure omitted — 213 chars of source]
figure[figure omitted — 216 chars of source]
table[table omitted — 1,031 chars of source]

}

}