Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
282,821 characters · 23 sections · 82 citation commands
{7pt} {7pt}
Neural network (NN) architecture has received increasing attention over the last several decades. On relevant topics, a large number of papers have been published in different journals such as Econometrica, The Annals of Statistics, and Journal of Machine Learning Research, etc. by experts from different disciplines. Apparently, we cannot exhaust the literature, but refer interested readers to BMR2021 and FMZ2021 for extensive reviews from a methodological point of view.
NN usually includes three ingredients: input layer, hidden layer(s), and output layer. We now briefly comment on them one by one. The input layer is possibly the easiest one to understand, as it includes regressors only. The hidden layer(s) involve lots of activation functions and parameters mapping linear combinations of the regressors to a certain range of the real line. Usually, sparsity has to be imposed to ensure a reasonable number of effective parameters and a small set of active activation functions (e.g., SH2020, WL2021). Once the parameters are estimated, one can load the test dataset to evaluate the performance of NN. Finally, the output layer receives the outcome, of which there are two notable types (i.e., quantitative and qualitative). Against this background, a semiparametric regression model under the context of treatment effects (e.g., BCH2014) naturally includes these two types of output in one framework, so it offers a nice structure to start the following semiparametric regression model:
where $\mathbf{x} =(x_1,\ldots, x_d)^\top$ is a $d\times 1$ vector of control variables, $z$ is a treatment/policy variable subject to the influence of $\mathbf{x}$ via the structure $z= I(G(\mathbf{x} )-\eta\ge 0 )$ with $I(\cdot)$ being the indicator function, both $\varepsilon$ and $\eta$ are idiosyncratic error components, and $g$ and $G$, defined on $[-a,a]^d\to \mathbb{R}$ with $a$ being fixed\footnote{We consider the fixed $a$ case in the main text, and explain how to allow them to be defined on $\mathbb{R}^d$ in Appendix (ref).}, are unknown functions.
Model ((ref)) belongs to a class of partially linear models studied extensively in the relevant literature, see, robinson1988, hlg2000, Gao, LR2007, ttg2010, and BCH2014, for example. Existing estimation and inferential methods are mainly based on nonparametric kernel and series methods for the case where the dimensionality of $\mathbf{x}$ is small. In the current big--data environment where the dimensionality of $\mathbf{x}$ is large, there are newly proposed methods, including machine learning based methods.
The main features of model ((ref)), which has been proposed and discussed in BCH2014, are that model ((ref)) involves a binary structure for $z$, and BCH2014 develop a series based approach for the estimation of $\alpha$. By contrast, this paper proposes using a localized NN (LNN) method for the estimation of $\alpha$ and $g(\cdot)$ simultaneously. In addition, this paper also develops an easily implementable dependent wild bootstrap method for the inference of both $\alpha$ and $g(\cdot)$.
Until very recently, the investigation on NN architecture mainly focuses on some fixed design regression models (e.g., Cybenko1989 and many follow-up studies since then), or uses independent and identically distributed (i.i.d.) data (e.g., BK2019,SH2020; and many references therein). There are only limited studies available for us to understand NN with dependent data from theoretical perspective (see, cs1998,chen2007; for example), although NN based methods have been widely used to study time series data in practice (e.g., HOR1996,crs2001,GU2021429,Gu2020; just to name a few). We would like to contribute along this line of research, and thus assume the following time series data are observable:
where, for a positive integer $T$, $[T]$ stands for $\{ 1,\ldots,T\}$.
Meanwhile, it seems that so far the majority of the literature focuses on prediction errors, and barely talks about how to build feasible inferential procedures, such as constructing confidence intervals. A few exceptions known to us are DFLSP2021, FLM2021, CLMZ2022, and hhll24 for example on estimation and testing for the average treatment effect rather than on $g(\cdot)$ that we are also interested in this paper.
Therefore, one important objective and contribution of this paper is that we develop NN based approach to addressing both estimation and inferential issues for both $\alpha_0$ and $g(\cdot)$ using (ref). We also show how $G(\cdot)$ can be recovered via NN practically in Appendix (ref). All things considered, we draw Figure (ref) for the purpose of illustration, in which only the dark area of the hidden layer is activated. A few questions arise naturally:
{
}
Another challenge which arises with the complexity of NN architecture is the transparency of algorithms (see Appendix (ref) for a brief survey of the existing software packages). One main reason is the lack of practical guidelines for establishing a feasible version. Our literature review highlights that social science studies using NN approach rarely provide detailed descriptions of their numerical implementation. While we concur with Athey2019 that machine learning will have a transformative impact on social science, transparent algorithms are crucial to ensure the practical relevance and utility of the findings derived from these approaches.
Having those said, our contributions are as follows.
(i) We establish an approximation procedure that approximates polynomials via NN in a local sense rather than a global sense.
(ii) We then explore the use of identification restrictions and establish the LNN based approach under a set of mild conditions.
(iii) We show that some closed--form expressions can be obtained for the estimators of the parameters of interest.
(iv) Accordingly, asymptotic distributions are derived, and a dependent wide bootstrap procedure is proposed for inferential purposes.
(v) As shown in Theorems 2.1 and 2.2 and their discussions in Section 2 below, the LNN based estimation and inferential methods outperform such results associated with existing methods.
(vi) We validate our theoretical findings through extensive numerical studies.
(vii) In an empirical study, we revisit the impacts of a tightening monetary policy action on a variety of economic variables, including short-/long-term interest rate, inflation, unemployment rate, industrial price and equity return via the newly proposed framework using a monthly dataset of the US.
The rest of the paper is organized as follows. In the main text,
(a) Section (ref) introduces LNN architecture, proposes a group--LASSO based estimation procedure, and then establishes the corresponding asymptotic properties to infer $\alpha_0$ and $g(\cdot)$, respectively;
(b) We provide extensive simulation studies in Section (ref) to examine the finite-sample performance;
(c) Section (ref) presents an empirical study that investigates the average effects of the US monetary policy change on macroeconomic and financial variables;
(d) Section (ref) concludes.
In the online supplementary appendices,
(e) Appendix (ref) includes some discussions on issues associated with practical implementation and also presents a detailed algorithm;
(f) Appendix (ref) discusses the estimation of a fully nonparametric model which is a special case of (ref), explains how to relax the restriction about $a$, and infers $G(\cdot)$ of $z_t$ via LNN;
(g) Appendix (ref) includes additional simulations;
(h) We finally give the proofs in Appendix (ref).
To close this section, we introduce some notation and mathematical symbols. Vectors and matrices are always expressed in bold font. Further, $\| \cdot \|$ denotes the Euclidean norm of a vector or the Frobenius norm of a matrix; $ \mathbf{0}_{a}$ and $ \mathbf{1}_{a}$ are respectively $a\times 1$ vectors of zeros and ones for $a\in \mathbb{N}$ and $\mathbf{I}_a$ denotes an $a\times a$ identity matrix; for a vector of nonnegative integers $\pmb{\mu} = (\mu_1,\ldots, \mu_d)^\top \in \mathbb{N}_0^{d}$ in which $\mathbb{N}_0=0\cup \mathbb{N}$, let $\pmb{\mu}! =\mu_1!\cdots \mu_d!$; $\mathtt{c}$, $\mathtt{C}$ and $O(1)$ always stand for fixed constants, and may be different at each appearance; $\to_P$ and $\to_D$ stand for convergence in probability and convergence in distribution, respectively.
For a function $m\,:\, [-a,a]^d\mapsto \mathbb{R}$, let $\|m \|_{\infty} = \sup_{\mathbf{x}\in [-a,a]^d}|m(\mathbf{x})|$. If the partial derivative of $m(\mathbf{x})$ exists, we write $\frac{\partial^{|\pmb{\mu}|} m(\mathbf{x})}{\partial \mathbf{x}^{\pmb{\mu}}} =\frac{\partial^{|\pmb{\mu}|} m(\mathbf{x})}{\partial x_1^{\mu_1}\cdots \partial x_d^{\mu_d}}$ for short, where $|\pmb{\mu}| =\sum_{j=1}^d \mu_j$. Additionally, let
for $\mathbf{a}= (a_0, a_1,\ldots, a_d)^\top \in \mathbb{R}^{d+1}$, $\mathbf{r} = (r_0,r_1,\ldots, r_d)^\top\in \mathbb{N}_0^{d+1}$, and $ \mathbf{x}\in \mathbb{R}^{d}$. Denote that for given $q\geq 1$
where $\mathbf{n}=(n_1,\ldots, n_d)^\top\in \mathbb{N}_0^d$, and the dimension of $\mathscr{P}_q$ is apparently $\text{dim}\mathscr{P}_q =\binom{d+q}{d}\coloneqq d_q$. Accordingly, let
where $m_j(\mathbf{x} \, |\, \mathbf{x}_0)$'s are the basis monomials (centered at $\mathbf{x}_0$) of $ \mathscr{P}_q$. Denote a set
where $x_j$ and $x_{0,j}$ are the $j^{th}$ elements of $\mathbf{x}$ and $\mathbf{x}_0$ respectively, and $h$ is a bandwidth. Finally, let $ \mathbf{H}=\operatorname*{\normalfont\textrm{diag}}\{H_1,\ldots, H_{d_q}\}$ with $H_j =\prod_{k=1}^dh^{-n_{j,k}}$, and let $\mathbf{n}_j =(n_{j,1},\ldots, n_{j,d})^\top$ include the corresponding power terms of $m_j(\mathbf{x} \, |\, \mathbf{x}_0)$ defined in (ref).
In this section, we first introduce LNN in Section (ref) and then infer $\alpha_0$ and $g(\cdot)$ in Section (ref). Notably, several new and useful results are established in Lemmas (ref)-(ref) to show how to approximate $g(\cdot)$ by LNN. These results contribute to the current literature in at least the following two points:
1. In contrast to the current literature that usually allows for all parameters to be estimated from data, we start by presenting some identification conditions, which can help reduce the number of effective parameters significantly.
2. One key idea behind NN architecture is that it can approximate polynomial terms, of which as well understood the linear combination can further approximate unknown functions by standard nonparametric analysis (e.g., BK2019, SH2020).
In this paper, we do the same, but the difference is that we introduce a bandwidth parameter $h$ below. The reason is that approximating polynomials via NN can only be achieved in a local sense rather than a global sense (cf., Lemmas (ref)-(ref) below). This is why we use the terminology LNN. More importantly, $h$ has a direct control on active and non-active activation functions of the hidden layer.
To proceed, we recall the notation defined in Section (ref) and explain how NN architecture works conceptually. In the literature of machine learning, one prefers to target the entire area that the unknown function is defined on, and normally uses a training set to pre-specify a lot of parameters, which do not change with respect to the observations of the test set. By doing so, one just pays some price when calculating the parameters in the first time, and no longer needs to update them when loading the test set. We are now ready to proceed.
First, we define a family of sufficiently smooth functions, and formally state the first assumption of the paper.
Assumption (ref).1 is widely adopted in the literature of nonparametric regression (e.g., LR2007). The main point is that each component in Taylor expansion of $g(\cdot)$ is bounded and also sufficiently smooth. Assumption (ref).2 nests a wide class of activation functions commonly used in the literature as special cases, e.g., Sigmoidal squasher (i.e., $\sigma(x) = \frac{1}{1+\exp(-x)}$), Error function (i.e., $\sigma(x) =\frac{2}{\sqrt{\pi}}\int_0^x \exp(-w^2)dw$), etc. We refer the interested reader to DUBEY202292 for a comprehensive review on different activation functions, and to Appendix (ref) for a detailed example. We acknowledge the growing literature on the rectified linear unit (ReLU) function (e.g., SH2020, FLM2021). As ReLU and Sigmoidal functions require different approximation theories, we therefore focus on the Sigmoidal function in this paper.
We then show the feasibility of LNN architecture.
Using Lemma (ref), the LNN method requires us to focus on a compact set $[-a, a]^d$, as we need to partition it into lots of small cubes. These small cubes may have different names in each appearance. For example, they are referred to as (hyper-) cubes in BK2019 and SH2020, and are referred to as localization in FLM2021. Each cube is corresponding to an effective sample set, which is jointly determined by the point of interest and the bandwidth. Finally, we note that in Appendix (ref), we show that the main results remain valid when $a\rightarrow \infty$ along with the sample size.
That said, for a given integer $M\ge 1$, we subdivide $[-a, a]^d$ into $M^d$ cubes of side length $2h\coloneqq\frac{2a}{M}$, and number these cubes by $C_{\mathbf{x}_{0\mathbf{i}},h} $ with $\mathbf{i}\in [M]^d$:
where $\mathbf{x}_{0\mathbf{i}}$ represents the center of $C_{\mathbf{x}_{0\mathbf{i}},h} $, and $x_j$ and $x_{0\mathbf{i},j}$ are the $j^{th}$ elements of $\mathbf{x}$ and $\mathbf{x}_{0\mathbf{i}}$ respectively.
Let $\widetilde{s}(\mathbf{x} \, |\, \widetilde{\pmb{\Lambda}}) = \sum_{\mathbf{i}\in [M]^d} I_{\mathbf{i},h}(\mathbf{x}) s(\mathbf{x} \,|\, \mathbf{x}_{0\mathbf{i}}, \widetilde{\pmb{\lambda}}_{\mathbf{i}})$, $\widetilde{\pmb{\Lambda}} =\{\widetilde{\pmb{\lambda}}_{\mathbf{i}}\, |\, \mathbf{i}\in [M]^d\}$ and $I_{\mathbf{i},h}(\mathbf{x}) =I(\mathbf{x}\in C_{\mathbf{x}_{0\mathbf{i}},h} )$. We then establish the following important lemma.
We are now ready to work on our estimation method. To accommodate a potentially large number of explanatory variables, we adopt a sparse structure, and impose the following assumption.
Here, to be clear, we still require $d_0$ and $d$ to be fixed in theory as in BK2019, SH2020 and FLM2021. As pointed out in fan1996local, when it comes to estimate unknown functions, such settings where $d\geq 4$ and the same size is not big enough can cause the so--called “curse of dimensionality". In our numerical studies, with a relatively large $d_0$, our approach still achieves good finite sample properties.
Under Assumption (ref), it is clear to see that the coefficient vector $\pmb{\lambda}$ of Lemma (ref) inherits the sparsity of $g(\mathbf{x}_t)$. Specifically, we denote the true space as
where $\mathbf{n}_0=(n_1,\ldots, n_{d_0})^\top\in \mathbb{N}_0^{d_0}$. We then define $\overline{\mathscr{P}}_q$ as the complement space, such that $\mathscr{P}^0_q \cup \overline{\mathscr{P}}_q=\mathscr{P}_q$ and $\mathscr{P}^0_q \cap \overline{\mathscr{P}}_q=\emptyset$. Simple algebra shows that the dimensions of $\mathscr{P}^0_q$ and $\overline{\mathscr{P}}_q$ are $\text{dim}\mathscr{P}^0_q =\binom{d_0+q}{d_0}\coloneqq d^0_q$ and $\text{dim}\overline{\mathscr{P}}_q = d_q-d^0_q$ respectively. In connection with (ref), we suppose further that $\{m_1(\mathbf{x} \, |\, \mathbf{x}_0),\ldots, m_{d^0_q}(\mathbf{x} \, |\, \mathbf{x}_0)\}$ constitute a basis for $\mathscr{P}^0_q$ without loss of generality\footnote{This implies that we can express $m_j(\mathbf{x} \, |\, \mathbf{x}_0)=m_{c,j}(\mathbf{x}_c \, |\, \mathbf{x}_{c,0})$ for $j=1,\ldots, d_q^0$. However, this requirement is purely for notational simplicity and is not necessary for the validity of our estimation and theoretical development.}. Consequently, the sparsity of $\pmb{\lambda}$ is expressed as $\{\lambda_{j}=0\, |\,j=d_q^0+1,\ldots,d_q\}.$ Moreover, similar arguments to those presented in Lemma (ref) can be applied to show that the true function $g_c(\mathbf{x}_c)$ can be approximated by the oracle NN architecture $\widetilde{s}_c(\mathbf{x}_c \, |\, \widetilde{\pmb{\Lambda}}_c)$: $\| g_c(\mathbf{x}_c) - \widetilde{s}_c(\mathbf{x}_c \, |\, \widetilde{\pmb{\Lambda}}_c) \|_{\infty}=O(h^p)$, where $\widetilde{\pmb{\Lambda}}_c$ is the oracle counterpart of $\widetilde{\pmb{\Lambda}}$.
Drawing upon the sparsity structure of $\pmb{\lambda}$, we propose using the group-LASSO strategy (YL2006) to formulate the following penalized estimation:
where $\{\psi_j\}$ denote the tuning parameters, $\widetilde{Q} (\alpha,\pmb{\Theta})= \sum_{t=1}^T[y_t-z_t\alpha -\widetilde{s}(\mathbf{x}_t \, |\, \pmb{\Theta}) ]^2$, $\mathcal{S} = \{ \widetilde{s}(\mathbf{x} \, |\, \pmb{\Theta})\}$ with $\widetilde{s}(\mathbf{x} \, |\, \cdot)$ being defined in Lemma (ref), $\pmb{\Theta} =\{\pmb{\theta}_{\mathbf{i}}\, |\, \mathbf{i}\in [M]^d\}$ with $\pmb{\theta}_{\mathbf{i}}$'s being $d_q\times 1$ vectors, and $\pmb{\Theta}_{D,j}$ is an $M^d\times 1$ vector with the $j^{th}$ component being $\mathbf{D}^{\top,-1}\pmb{\theta}_{\mathbf{i}}$.
Let $\widetilde{\pmb{\theta}}_{\mathbf{i}}$ and $\widetilde{\pmb{\Theta}}_{D,j}$ be the corresponding estimators of $\pmb{\theta}_{\mathbf{i}}$ and $\pmb{\Theta}_{D,j}$. For the purpose of comparison, we also define the following oracle estimators: $ (\widetilde{\alpha}_c,\widetilde{\pmb{\Theta}}_c)$ of the form:
where $\widetilde{Q}_c (\alpha,\pmb{\Theta}_c)= \sum_{t=1}^T[y_t-z_t\alpha -\widetilde{s}_c(\mathbf{x}_{c,t} \, |\, \pmb{\Theta}_c) ]^2$, and $\pmb{\Theta}_c$, $\widetilde{s}_c(\mathbf{x}_{c,t} \, |\, \pmb{\Theta}_c)$ and $\mathcal{S}_{c}$ respectively represent the oracle counterparts of $\pmb{\Theta}$, $\widetilde{s}(\mathbf{x}_{t} \, |\, \pmb{\Theta})$ and $\mathcal{S}$ (i.e., assuming the sparsity is known). Although the oracle estimators $\widetilde{\alpha}_c$ and $\widetilde{\pmb{\Theta}}_c$ are infeasible practically, they can be approximated by the group-LASSO estimators with asymptotically negligible biases. To facilitate the rest of our development, we impose the following time series structure on (ref) in Assumption 3 below.
Assumption (ref) is standard in the literature of time series analysis (see, Gao, for example). Heterogeneity may also be introduced by, for example, $E[\varepsilon_1^2 \,| \, \mathbf{x}_1,\eta_1]=\psi(\mathbf{x}_1,\eta_1)$, which however makes notation even more complicated than what we need to involve. $\{\mathbf{x}_t\}$ may also have more complex structures such as linear processes, locally stationarity, deterministic trends, etc. Surely, the corresponding development and asymptotic results will need to be modified accordingly, but it will not add too many credits to the original idea of this paper. We therefore focus on the current setting.
To proceed, we introduce some additional notation. Let $\mathbf{H}_c$, $\mathbf{x}_{c,0\mathbf{i}}$, and $\mathbf{m}_c(\mathbf{x}_{c,t}\, |\, \mathbf{x}_{c,0\mathbf{i}})$ be the oracle counterparts of $\mathbf{H}$, $\mathbf{x}_{0\mathbf{i}}$, and $\mathbf{m}(\mathbf{x}_{t}\, |\, \mathbf{x}_{0\mathbf{i}})$. Also, let $\Phi_\eta(x)$ and $f_{\mathbf{x}_c}(\mathbf{x}_{c,\mathbf{i}})$ denote the CDF of $\eta_t$ and the density function of $\mathbf{x}_{c,t}$, respectively. Also define
Assumption 4 imposes a set of regularity conditions. Assumptions 4.1 and 4.2 ensure the positive definiteness of the asymptotic covariances, and Assumption 4.3 regulates the bandwidth and the group-LASSO tuning parameters, which can be easily fulfilled. The first part of Assumption 4.3 requires that when the dimensionality of the true covariates satisfies $d_0<d$, there is no need to impose an under--smoothing condition, and it is required to impose the under--smoothing condition when $d_0=d$. For the second part of Assumption 4.3, detailed justifications are available from Appendix A.1 of the online supplementary document for more details.
Given these extra conditions, we establish the following lemma.
Lemma (ref).1 indicates that the group-LASSO method can correctly identify the sparsity of $\pmb{\lambda}$. Lemma (ref).2 shows $\widetilde{\alpha}$ and $ \widetilde{\alpha}_c$ are asymptotically equivalent up to a rate of an order of $O_P(h^p)+o_P (T^{-\frac{1}{2}} )$.
Let $\mathbf{Z}=(z_1,\cdots,z_T)^\top$ and $\widetilde{\mathbf{X}}_{c,\mathbf{i}} = (\widetilde{\mathbf{x}}_{c,\mathbf{i}, 1}, \cdots, \widetilde{\mathbf{x}}_{c,\mathbf{i}, T})^\top$, where $$\widetilde{\mathbf{x}}_{\mathbf{i},c,t} =I_{\mathbf{i},h} (\mathbf{x}_{c,t})(\mathbf{I}_{d^0_q}\otimes \boldsymbol{\gamma}^\top) \pmb{\sigma}_c( \mathbf{x}_{c,t}\, |\,\mathbf{x}_{0,c,\mathbf{i}})$$ and $\pmb{\sigma}_c( \mathbf{x}_{c,t}\, |\,\mathbf{x}_{0,c,\mathbf{i}})$ denotes the oracle counterpart of $\pmb{\sigma}(\mathbf{x}\, |\, \mathbf{x}_0 ) $. Building upon Lemma (ref) of the online supplement, we can now establish the asymptotic distributions of the proposed estimators.
Note that the rate of convergence, $\sqrt{T h^{d_0}}$, is faster than that of the conventional kernel estimator of an order of $\sqrt{T h^d}$ due to the fact that $Th^d = Th^{d_0} \, h^{d- d_0} = o\left(T h^{d_0}\right)$ when $d_0<d$. Meanwhile, the order of the bias term is also smaller than that of the conventional kernel estimator due to the construction of the LNN architecture and the definition of $p>2$. This is the main reason that the proposed LNN method outperforms over the conventional kernel method in finite--sample studies particularly when $d_0\geq 4$. Table 1 of Section 3 further supports the finite--sample superiority of the proposed LNN method.
When establishing the second result of Theorem (ref), the terms $E[\varepsilon_1\varepsilon_{1+t}]$ for $t\ge 1$ all vanish in the asymptotic covariance matrix due to the partition of (ref) and the use of the indicator function. As a result, the confidence interval associated with the second result of Theorem (ref) can be easily achieved.
Because $\sigma_{c,z}$ defined in Assumption 4 involves a long--run variance component, however, the serial dependence is not asymptotically negligible when inferring $\widetilde{\alpha}$. Therefore, we propose using the following dependent wild bootstrap method.
To close this section, we emphasize that as have stated in Section (ref), we provide the detailed study about a fully nonparametric regression in Appendix (ref) (i.e., letting $\alpha_0 =0$). Although it is a special case of (ref), the investigation yields some useful insights and allows us to further clarify some features of LNN compared to the existing literature. Also, we infer $G(\cdot)$ of the binary structure of $z_t$ in Appendix (ref), which is also of great interest in both theory and practice.
Up to this point, we would like to point out that those questions raised in Section (ref) have all been answered, so Figure (ref) can be understood better. We summarize some key points which may have been discussed previously here and there, and further discuss some remaining issues.
Sparsity --- It is now clear that without sparsity, the total number of activation functions is
of which only $d_q(q+1)$ neurons are activated when loading test data. Among the activated activation functions, the number of effective parameters is $d_q$, while the rest of the parameters are predetermined. Provided sparsity, utilizing identification conditions allows us to define the true set of parameters, so we can further reduce the effective parameters via the group-LASSO approach.
Multiple Hidden Layers --- Lemma (ref) yields a recursive relationship, as one can repeatedly invoke Lemma (ref) to replace $(x-x_0)$ inside the activation. It will then yield a LNN architecture with multiple hidden layers. Although having multiple hidden layers is achievable, at this stage it is not clear to us why we should do so. As discussed in Remark (ref), this step is completely independent of data, so we do not see any benefit of doing so unless $g(\cdot)$ (or $G(\cdot)$) has certain specific structure. Under some extra structure on $g(\cdot)$ (or $G(\cdot)$), however, the necessity of developing LNN architecture with multiple hidden layers deserves extra attention in future research.
Dependence & Trending --- LNN automatically eliminates some correlation of observations from different time periods when establishing the asymptotic distribution for the unknown functions. In this paper, we assume that the regressors $\{\mathbf{x}_t\, |\, t\in [T]\}$ are strictly stationary and mixing. In fact, they can have other more complex structures, such as linear processes, locally stationarity, heterogeneity, deterministic trends, etc. As a result, many climate models (such as those in MUDELSEE2019310) may be better captured. For such cases, one may need to revise the assumptions and proofs accordingly depending on detailed research questions.
Data-Splitting --- One may further connect the above results with the data-splitting technique of Chernozhukov2018, and simultaneously explore the sparsity of the binary structure in $z_t$.
In this section, we conduct simulations to validate the theoretical findings, focusing specifically on the semiparametric model with sparsity. Additional simulation results are provided in Appendix (ref) of the online supplement to exam additional theoretical results of Appendix (ref), and to demonstrate the newly proposed method works reasonably well even without involving sparsity.
As discussed in Section (ref), many parameters are involved in the LNN architecture. It would be extremely difficult to systematically check every single one in this paper due to the page constraints, so we have to be selective. Also, we are constrained by computing power. That said, the following quantities are pre-fixed without loss of generality.
We consider the semiparametric treatment effects model (ref), where $\varepsilon_t = 0.5\varepsilon_{t-1} + N(0, 0.75)$, the $j^{th}$ element of $\mathbf{x}_t$ is independently generated as $x_{t,j}\sim U(-1, 1)$, and $z_t$ is independently generated from Bernoulli ($\cos(|\mathbf{x}_{c,t}^\top \mathbf{1}_{d_0}/d_0|)$) to allow for dependence between $z_t$ and $\mathbf{x}_{0,t}$. We simply let $\alpha_0=1$ and $g(\mathbf{x}_t)=g_c(\mathbf{x}_{c,t}) = 1+\sin (\mathbf{x}_{c,t}^\top \mathbf{1}_{d_0})$, where $\mathbf{x}_{c,t}$ contains the first $d_0$ elements in $\mathbf{x}_t$. The bandwidth $h$ is set as $h=a/M$, where $M$ is the integer closest to $a/h_1$ with $h_1=1.5\cdot T^{-1/(d+2p-0.5)}$. Here, $-0.5$ is to ensure $\sqrt{T} h^{p+d/2}\to 0$ holds. In fact, $h$ is very close to $h_1$, and the current setup is simply to guarantee $M$ is a large positive integer. When designing the simulations, our impression is that the results are not sensitive to the choices of $g(\mathbf{x})$ and the bandwidth. Therefore, in what follows, we only vary the values of $q$, $d$($d_0$), $u_\sigma$. Specifically, we use $d\in\{2, 8, 14\}$, $d_0=d/2$, $q\in\{3,4\}$ and $u_\sigma\in\{-0.5,0.5\}$.
To measure the estimate of $g(\cdot)$, we select a few test points from $[-a,a]^d$. Ideally, we would like to select $L$ points from each dimension, so it gives $L^d$ points to evaluate in total. However, it will create a lot computational overhead, so we select the points $\mathbf{x}_{L,j} \coloneqq (-a+\frac{2aj}{L+1} )\mathbf{1}_d$ for $j\in[L]$. At each test point, we construct the estimate along with the $95\%$ confidence interval using the method of Section (ref). To evaluate the finite sample performance, we compute the root mean squared errors (RMSEs) and coverage rates (CRs) after $n$ replications:
where $\widetilde{\alpha}_i$ and $\widetilde{g}_i(\cdot) $ stand for the estimates of $\alpha_0$ and $g(\cdot)$, respectively, at the $i^{th}$ replication. $\text{CR}_{\alpha, ij}$ and $\text{CR}_{g, ij}$ are the 95% confidence intervals of $\widetilde{\alpha}_i^\ast-\widetilde{\alpha}_i$ and $\widetilde{g}_i^*(\mathbf{x}_{L,j})-\widetilde{g}_i(\mathbf{x}_{L,j})$, respectively, based on the bootstrap draws from the $i^{th}$ replication. The number of bootstraps is set to be 200. We select $T\in \{800, 1600,2400 \}$, $L=10$ and $n=1000$ without loss of generality.
RMSEs and coverage rates are reported in Table (ref). As evident in the table, RMSEs of both $\widetilde{\alpha}_i$ and $\widetilde{g}_i(\cdot)$ decrease as $T$ increases from 800 to 2400. Moreover, coverage rates are close to 0.95 indicating that the bootstrap procedure behaves reasonably well. Some additional facts should also be mentioned. The simulation results are not changing significantly with respect to the value of $u_\sigma$, which confirms our argument in Remark (ref). For comparison, we also report the simulated RMSEs and CRs for the LNN estimation without considering the sparsity structure. As evident in Table (ref), the LNN-based group-LASSO estimators generally outperform the LNN estimation, especially in cases with large parameter sets (e.g., $d=14$ and $q=4$).
To further illustrate the proposed bootstrap method, we present the plots of 95% bootstrap confidence intervals of $g(\cdot)$ for the cases $(d,\mu_\sigma)=(2,0.5)$ and $(d,\mu_\sigma)=(2,-0.5)$ in Figures (ref) and (ref), respectively. For enhanced clarity, we employ a denser grid of evaluation points $\mathbf{x}_{L,ij} = (-a+\frac{2ai}{L+1}, -a+\frac{2aj}{L+1} )$ for $i,j\in[L]$. At each point, we apply the bootstrap procedure and compute the 95% confidence interval. The plots in Figures (ref) and (ref) demonstrate that the true functions are effectively covered across different choices of $T$ and $q$. Moreover, with the increasing sample size, the bootstrap confidence intervals exhibit clear convergence. Furthermore, $q=3$ seems to yield better coverage overall, and the choice of $u_\sigma$ has less impact on the inference.
Having demonstrated the finite sample performance of the newly proposed framework, we are now ready to examine the impacts of US monetary policy in the following section.
It is widely acknowledged that monetary policy shocks may have significant influence over macroeconomic variables. As noted by Clarida2000, “{\it the difference in policy behaviour could be an important underlying source of the shift in macroeconomic behaviour}". Consequently, over the past few decades, there has been a surge of empirical and theoretical research to study the relationship between monetary policies and various macroeconomic indicators, including interest rate, unemployment rate, economic growth, and asset price. For example, Clarida2000 investigate the role of monetary policy in the macroeconomic stability by establishing connections between the policy interest rate with expected inflation; Bernanke2005 explore how the stock index responds to the unanticipated monetary policy actions; Blanchard2010 study the policy effects on the relationship between inflation and unemployment rate; etc.
As widely recognized in the literature, a fundamental question in this field is how to characterize monetary policy shocks or unanticipated policy changes (see Christians1996,Bernanke2005; among others). A commonly adopted approach involves using the disturbance in a regression of monetary policy indicators on lagged observable macroeconomic variables to capture the information conveyed by policy actions that cannot be anticipated by the market. Then, the response of the macroeconomic variables to the policy shock can be measured through another regression between such variables. In this section, we adopt a nonparametric specification for the monetary policy evolution, which can be regarded as a generalization of the linear model employed by Christians1996. Specifically,
where $P_t$ is an indicator variable for the shift to a more tightening monetary policy, $\mathbf{x}_{t}$ contains the macroeconomic predictors available when $P_t$ is determined, $G_P(\cdot)$ is a nonparametricalyy unknown function, and $\eta_t$ is a disturbance term. We then have
where $\Phi_\eta(\cdot)$ denotes the CDF of $\eta_t$. Therefore, $\Phi_\eta(G_P(\mathbf{x}_t))$ captures the market anticipation of a tightening policy action by the authority. This concept is in line with the idea of policy propensity scores adopted by Angrist2018. Then, the unanticipated monetary policy shock $S_t$ can be specified as follows:
We can then characterize its relationship with macroeconomic or financial outcome variables ($M_t$) through the following semiparametric model:
where the parameter $\alpha_P$ captures the effects of unanticipated policy shocks and $G_M(\cdot)$ is another nonparametric function that controls the influence from the lagged macroeconomic predictors. By substituting $S_t$ from (ref) into (ref), we obtain
where $g_M(\mathbf{x}_{t})=G_M(\mathbf{x}_{t})-\Phi_\eta(G_P(\mathbf{x}_t))\alpha_P$. The setup of (ref) enables us to estimate the policy effects $\alpha_P$ using the methodology that is proposed in Section (ref).\footnote{In the online supplement, we provide an LNN-based method to estimate the nonparametric binary model (ref). Then, using the estimators $\widehat{G}_P(\mathbf{x}_t)$ and $\widehat{g}_M(\mathbf{x}_{t})$ that are obtained by estimating (ref) and (ref), respectively, we can recover $G_M(\mathbf{x}_{t})$ in (ref) by $\widehat{G}_M(\mathbf{x}_{t}) = \widehat{g}_M(\mathbf{x}_{t})+\Phi_\eta(\widehat{G}_P(\mathbf{x}_t))\widehat{\alpha}_P$.
An alternative approach to recovering the policy effects is to estimate (ref) and (ref) sequentially. However, this procedure involves an essential step where we have to replace the unobservable unanticipated policy shock $S_t$ in (ref) with its estimator $\widehat{S}_t=P_t-\Phi_\eta(\widehat{G}_P(\mathbf{x}_t))$ to construct a feasible estimator for $\alpha_P$. This substitution inevitably introduces additional approximation errors compared to the direct estimation of (ref). } Compared with the traditional parametric models of monetary policy effects, our framework offers several advantages. First, the relationship between variables is not constrained to be linear, allowing for more flexible and realistic modelling of complex interactions. Second, our approach is capable to accommodate a large set of control variables and automatically detect the insignificant predictors.
In what follows, we revisit the impacts of a tightening monetary policy action on a variety of economic variables, including short-/long-term interest rate, inflation, unemployment rate, industrial price and equity return. We accomplish this by employing the proposed semiparametric model with treatment effects on a monthly dataset of the US.
In the literature Christians1996,Romer2000,Cochrane2002, a commonly used indicator of monetary policy shifts is the change in the federal funds target rate which is announced during Federal Reserve Open Market Committee (FOMC) meetings. Accordingly, we construct an indicator variable to capture the tightening or easing monetary policy whenever there is an increase in the announced target rate. We examine the effects of these policy changes on key macroeconomic and financial outcome variables that have been extensively studied in the literature. The sample period considered in this study is from January, 1989 to September, 2015.
Specifically, we first analyze the policy effects on a variety of interest rates: the effective federal funds rate (FFR) and treasury bond yields quoted on an investment basis at 3-month (Yield3m), 1-year (Yield1y), 2-year (Yield2y), 5-year (Yield5y), and 10-year (Yield10y) maturities. For these variables, we obtain their monthly average data from the FRED, Federal Reserve Bank of St. Louis. Following the approach of Angrist2018, we investigate the effects on inflation, unemployment rate, and industrial output, which are measured by (change in the log value of) the Personal Consumption Expenditures Price Index (PCE), (change in) the Civilian Unemployment Rate (UNRATE), and (change in the log value of) the Industrial Production Index (IP), respectively. The data for these variables are also collected from FRED. In addition to these macroeconomic variables, we follow Bernanke2005 to study the value-weighted return of the S$\&$P500 stocks as a representative of equity returns and the data is sourced from the CRSP index series database at WRDS.
To isolate the effects from the anticipation of the shifts in monetary policy, we incorporate a set of control variables that are sourced from three categories: (i) Monetary policy persistence: we measure the persistence of monetary policy using lagged federal funds target rate and real target rate changes (TRC), along with an indicator variable for FOMC meeting occurrences. (ii) Economic conditions: we include the first and second lags of inflation and unemployment rate to reflect the potential influence of underlying economic conditions on monetary policy changes. (iii) Market expectations: we utilize the federal funds future (FFF) index developed by Angrist2018 to capture market expectations regarding future monetary policy decisions \footnote{To measure market expectations of target rate changes, Angrist2018 develop the FFF variable using federal funds rate derivatives with specific adjustments made for data during FOMC meeting months. For detailed information about the construction of this index, we refer the readers to Angrist2018.}.
We present a summary of the aforementioned variables along with their descriptive statistics and unit root test results in Table (ref) below. As shown in the table, all variables are stationary at the 5% significance level, except 10-year treasury bond yields which is stationary at the 10% significance level. Therefore, these variables fit our assumption reasonably well.
Using the methodology outlined in Section (ref), we investigate the response of macroeconomic and financial variables in the subsequent month following the announcement of a federal funds target rate increase. To set up the LNN architecture, we specify $h=T^{-1/(d+2p-0.5)}$, $\mu_\sigma=0.5$ and $q=3$. The estimated policy effects captured by $\widehat{\alpha}_P$ and their bootstrap confidence intervals are reported in Panel A of Table (ref).
Our estimation results reveal that an unanticipated increase in the target rate exerts a significant and positive influence on the federal funds rate and on the treasury bond yield with the maturity of 3 months. However, the effect is statistically insignificant for treasury bond yields with maturities that are more than 1 year, suggesting that tightening monetary policy has a more pronounced impact on short-term interest rates compared to long-term rates. This aligns with previous studies by Cochrane2002 and Angrist2018, who also observe diminishing effects along the yield curve with increasing maturity. Controlling for anticipated changes in the target rate, our estimation indicates no significant effects of monetary policy shocks on changes in inflation, unemployment rate, or industrial price. These findings are consistent with those reported by Angrist2018. Table (ref) presents information about the insignificant predictors identified by the LNN-based group-LASSO method. As shown in Table 3, different numbers of significant predictors, ranging from one to five, are detected for FFR, Yield1y, Yield2y, PCE, UNRATE, and IP. This result demonstrates the empirical significance of the proposed method.
We then explore the sensitivity of our results by varying the set of control variables, specifically by excluding the FFF or both FFF and lagged values of PCE and UNRATE. The results, which are also presented in Panel A of Table (ref), demonstrate that policy effects appear to be more pronounced with fewer control variables, which is likely due to the inclusion of the influence from anticipated monetary policy changes. Notably, the unemployment rate response becomes significantly negative when FFF is not controlled for. Furthermore, without controlling for FFF and the lagged economic indicators, a noticeable short-term “price puzzle” in the literature Sims1992 emerges --- a temporary increase in price in response to higher interest rate.
To examine the longer-term effects of tightening monetary policy, we compute the effects of target rate changes ($P_t$) on the average changes of future outcome variables: $\overline{M}_{t+L}=\frac{1}{L}\sum_{l=1}^LM_{t+l}$, across different horizons ($L$) up to 24 months. The estimation results are depicted in Figures (ref) and (ref), illustrating an accumulation of positive effects on short-term bond yields within the first 12 months, followed by a gradual decline in subsequent months. As a robustness check, we re-estimate the monetary policy effects using data before August 2008, when the global financial crisis started. The estimation results are reported in Panel B of Table (ref). Notably, the impact of tightening monetary policy in curbing the increase in unemployment rate was significant before the global financial crisis, but it became considerably more muted after the crisis.
We follow the relevant literature in treatment effects HIR2003 to estimate the average treatment effect (ATE) using the inverse probability weighting approach. Specifically, we calculate the policy propensity scores $\Phi_\eta(\widehat{G}_P(\mathbf{x}_t))$ based on the LNN estimation of the binary model (ref), as discussed in Appendix (ref) of the online supplement, and define the weight as $w_t=\Phi_\eta(\widehat{G}_P(\mathbf{x}_t))^{-P_t}[1-\Phi_\eta(\widehat{G}_P(\mathbf{x}_t))]^{-(1-P_t)}$. Then, we construct the inverse probability weighting estimator for ATE as follows:
We omit these results from this study to maintain focus on our main objective.
NN has gained considerable attentions over the past few decades. Yet, many questions, as outlined in Section (ref), have not been satisfactorily addressed in the relevant literature. In this paper, we bring in identification restrictions to the LNN framework from a semiparametric regression perspective, and consider the LNN based estimation and inference for the unknown parameter and function involved in the modelling of time series data. We then integrate the LNN architecture with the group-LASSO technique to achieve consistent estimation and automatic detection of insignificant regressors. The asymptotic distributions are derived accordingly, demonstrating that the sparsity structure can be effectively identified and the estimators exhibit the distributions that are asymptotically equivalent to those for the infeasible oracle estimators.
Additionally, we propose a dependent wild bootstrap procedure to obtain valid inferences in practice. Last but not least, we validate our theoretical findings through extensive numerical studies. In the empirical study, we revisit the impacts of a tightening monetary policy action on a variety of economic variables, including short-/long-term interest rate, inflation, unemployment rate, industrial price and equity return via the newly proposed framework using a monthly dataset of the US.
Several major comments have been made here and there in Section (ref), and some future research directions have been acknowledged along the way. Finally, we hope the current article will shed light on how to produce transparent algorithms to ensure that our research findings are useful for practical implementations and applications.
Gao, Peng and Yang acknowledge financial support from the Australian Research Council Discovery Grants Program under Grant Numbers: { DP200102769, DP210100476 and DP230102250}, respectively. Liu's research was financially supported by National Natural Science Foundation of China under Grant Number: 72203114.
{
Appendix A includes some discussions on issues associated with practical implementation and also presents a detailed algorithm. Appendix (ref) discusses the estimation of a fully nonparametric model which is a special case of (ref), explains how to relax the restriction about $a$, and infers $G(\cdot)$ of $z_t$ via LNN. Appendix (ref) includes additional simulations. We finally present the preliminary lemmas and the proofs in Appendix (ref).
\setcounter{equation}{0} \setcounter{lemma}{0} \setcounter{section}{0} \setcounter{table}{0} \setcounter{figure}{0} \setcounter{remark}{0} \setcounter{corollary}{0} \setcounter{assumption}{0}
\setcounter{equation}{0} \setcounter{lemma}{0} \setcounter{section}{0} \setcounter{table}{0} \setcounter{figure}{0} \setcounter{remark}{0} \setcounter{corollary}{0} \setcounter{assumption}{0}
First, we provide some examples to illustrate our concern about the existing software packages. In the literature, the “neuralnet" R package (see FG2010 for detailed illustration) has been well adopted. When training a NN, a key parameter is called “hidden" that is a vector of integers specifying the number of hidden neurons in each layer by R document, and the package refers to MYA1994 regarding the choice of the number of neurons. However, it is worth pointing out that MYA1994 use a modified AIC criterion to investigate the case with the number of activation functions (as well as the number of parameters) being finite which is reflected in their asymptotic development. As a consequence, the arguments of MYA1994 no longer hold when the number of activation functions is diverging. As we show in this paper, having a diverging number of activation functions is the minimum requirement to achieve asymptotic consistency, and the rate of divergence is also associated with the sample size under a set of minor conditions. Similar issues also apply to “deepnet" package of {\tt R}, “torch.nn.Linear" and “tfl.layers.Linear" of Python, “\textsf{feedforwardnet}" of Matlab, etc.
Next, we comment on the matrix $\mathbf{D}$, the selection of the tuning parameter, a computational algorithm for the LNN based group-LASSO, and some properties of Sigmoidal squasher.
On the Matrix $\mathbf{D}$ --- By the proof of Lemma (ref), we can determine $\mathbf{D}$ through $\mathbf{m}(\mathbf{x}\, |\, \mathbf{x}_0)=\mathbf{D}\mathbf{A}(\mathbf{x} \, | \,\mathbf{x}_0)$, where
with $[\pmb{\alpha}_1 ,\ldots, \pmb{\alpha}_{d_q}]= \frac{1}{d+1}\cdot\operatorname*{\normalfont\textrm{diag}}\{h,\mathbf{I}_d \} \mathbf{W}$.
On Tuning Parameters --- An essential consideration for the practical application lies in the selection of tuning parameters: $\psi_1,\cdots,\psi_{d_q}$. We simplify the selection of $d_q$ by utilizing the unpenalized estimators $\widetilde{\pmb{\Theta}}^\ast_j$ along with a common tuning parameter $\psi$:
Using analogous arguments in the proof of Lemma (ref) and Lemma (ref), we can show that $h^{d}H_j^{-2}\|\widetilde{\pmb{\Theta}}^\ast_{j}\|^2$ converges to a positive constant in probability for $j\in[d_q^0]$, while it has the order of $O_P\bigl(\frac{1}{Th^d}\bigr)$ for $j=d_q^0+1,\ldots,d_q$. With (ref), Assumption (ref).3 becomes
By the definition of $H_j$, it is evident that $\max_{j\in[d_q^0]}\{H_j\}=O(h^{-q})$ and $\min_{j=d_q^0+1,\ldots,d_q}\{H_j\}=O(1)$. Therefore, a sufficient condition for the tuning parameter will be $\psi_0\sqrt{T}h^{d}\rightarrow\infty$ and $\psi_0T^{-\frac{1}{2}}h^{-q}\rightarrow0$. The remaining task involves the selection of $\psi_0$, which can be achieved by minimising the following information criterion:
where $ (\widetilde{\alpha}_\psi,\widetilde{\pmb{\Theta}}_\psi)$ are the group-LASSO estimators using (ref), and ${\rm df}_\psi$ is the number of nonzero coefficients identified by $\widetilde{\pmb{\Theta}}_\psi$. Accordingly, the optimal tuning parameter is obtained by
Algorithm for the LNN Based Group-LASSO --- The literature provides well-established computational algorithms for the LASSO estimation FL2001, HL2005. Herein, we adopt the local quadratic approximation procedure.
On Sigmoidal Squasher --- For Sigmoidal squasher, we have
in which
is the recursion relation for Stirling numbers of the second kind. We refer interested readers to MINAI1993845 for more details. Sigmoidal squasher is easy to use in the sense that we can arbitrarily choose $u_\sigma$ of Lemma (ref). Without loss of generality, we let $u_\sigma \in \{-0.5, 0.5\}$ in the numerical studies. In Figure (ref), we plot $\sigma^{(k)}(x) $ for $k=0,\ldots, 5$ for the purpose of demonstration.
Before we explain how to relax the restriction on $a$, we consider a fully nonparametric model by letting $\alpha_0\equiv 0$ and ignoring the sparsity (i.e., $d\equiv d_0$). The rest settings are identical to Section (ref).
The model to be investigated becomes
and the objective function is a simplified version of that involved in (ref):
Accordingly the OLS estimator of $\widetilde{\pmb{\Lambda}}$ is obtained by
and, for $\forall \mathbf{x}_0 \in [-a,a]^d$, the estimator of $g( \mathbf{x}_0)$ is then defined by $\widehat{g}(\mathbf{x}_0) = \widetilde{s}(\mathbf{x}_0 \, |\, \widehat{\pmb{\Theta}} ) $. Equation (ref) admits a closed-form estimator for each $\widehat{\pmb{\theta}}_{\mathbf{i}}$. To see this, we write
where $\widetilde{\mathbf{x}}_{\mathbf{i},t} =I_{\mathbf{i},h} (\mathbf{x}_t)(\mathbf{I}_{d_q}\otimes \boldsymbol{\gamma}^\top) \boldsymbol{\sigma}( \mathbf{x}_t\, |\,\mathbf{x}_{0\mathbf{i}})$ and the equality follows from the fact that $I_{\mathbf{i},h} (\mathbf{x}_t)I_{\mathbf{j},h} (\mathbf{x}_t)=0$ for $\mathbf{i}\ne \mathbf{j}$. Thus, for $\forall\mathbf{i}$, the first order condition yields
where the invertibility of $\sum_{t=1}^T\widetilde{\mathbf{x}}_{\mathbf{i},t} \widetilde{\mathbf{x}}_{\mathbf{i},t} ^\top$ is guaranteed asymptotically in view of (ref) and (ref).
After carefully studying (ref) for each $\mathbf{i}$ and repeatedly invoking $I_{\mathbf{i},h} (\mathbf{x}_t)I_{\mathbf{j},h} (\mathbf{x}_t)=0$ for $\mathbf{i}\ne \mathbf{j}$, the following theorem holds.
Obviously, our LNN based estimation method is simple and easy to implement. Accordingly, we propose the following bootstrap procedure to establish inference in practice.
For the bootstrap procedure, the following result holds immediately.
Note that, we require $g(\mathbf{x})$ to be defined on a compact set, but do not impose restriction on the range of $\{\mathbf{x}_t\}$. In fact, for time series data, it may make more sense to assume that $a$ is diverging, which is indeed achievable. We now provide two treatments to relax the restriction on $a$.
Treatment 1: Suppose that $\mathbf{x}_t$ follows a sub-Gaussian distribution, and we can then require $\sqrt{\log (Th^d)} \cdot a\to \infty.$ A similar treatment has also been discussed in LTG2016 for example, so we do not further elaborate it here. However, there is a price that we have to pay, i.e., the slow rate of convergence.
Treatment 2: Alternatively, we can modify the construction of LNN from a nonparametric viewpoint. In the literature of kernel regression, one normally pre-specifies a point of interest (e.g., $\mathbf{x}_0$), and investigates a small area nearby only which is usually decided by some bandwidth(s) converging to 0. As a result, the parameters obtained from the estimation procedure usually vary with respect to the point of interest. In other words, when evaluating different points from the test set, the number of parameters to be estimated will be proportional to the cardinality of the training set (although estimation is always carried on using the same training set). Provided a large test set, it may create lots of overhead from a computational viewpoint, but the advantage is that in theory it allows us to consider $\forall \mathbf{x}_0\in \mathbb{R}^d$.
That said, for $\forall\mathbf{x}_0\in\mathbb{R}^d$, we consider the following objective function:
where the subscript $L$ infers the local version, $\pmb{\theta}$ is $d_q\times 1$ vector satisfying $\| \pmb{\theta}\|<\infty$, we let $I_{0,h}(\mathbf{x}_t) :=I(\mathbf{x}_t \in C_{\mathbf{x}_0,h})$ for short, and $C_{\mathbf{x}_0,h}$ is defined in (ref) already. Then the OLS estimate of $\widetilde{\pmb{\lambda}}$ defined in Lemma (ref) is obtained by
and, accordingly, the estimate of $g(\mathbf{x}_0)$ is defined by $\widehat{g}_L(\mathbf{x}_0) = s(\mathbf{x}_0 \, |\, \mathbf{x}_0, \widehat{\pmb{\theta}} ) .$ We can then produce the following corollary.
So far, we have not explored the binary structure of $z_t$ much, which is also of great interest widely adopted in a wide range of applications (Athey2019). To close our investigation about the model (ref), we use LNN approach to infer $G(\cdot)$.
Recall that
For simplicity, we suppose that the respective probability density function (PDF) and the cumulative distribution function (CDF) of $\eta$ are known, and denote the PDF and CDF by $\phi_\eta(\cdot)$ and $\Phi_\eta(\cdot)$ respectively. Here, the information about $\phi_\eta(\cdot)$ and $\Phi_\eta(\cdot)$ is necessary for carrying on likelihood estimation.
Direct calculation shows that
which yield $E[z \, |\, \mathbf{x} ] =\Phi_\eta(G(\mathbf{x}))$. Accordingly, the log-likelihood function is defined below:
where the definition of $l_t(\cdot)$ is obvious. To infer $G(\cdot)$, we consider the following objective function:
which yields the following maximum likelihood estimator:
Note that our LNN method involves a general unknown function form. As a result, the classical results of likelihood estimation, such as those in NEWEY19942111, no longer hold. Therefore, before establishing an asymptotic distribution using $\widehat{\pmb{\Theta}}$, we state a lemma to show the feasibility of LNN architecture when modelling binary outcomes.
Lemma (ref) provides the consistency, and also bridges the likelihood estimation and the nonlinear least squares approach to some extent. To be precise, Lemma (ref) does not provide any specific consistency for $\forall\, \widehat{\pmb{\theta}}_{\mathbf{i}}$. Instead, it evaluates the overall performance of LNN. More importantly, it says when modelling a binary outcome, the likelihood estimation using the LNN architecture is approximately equivalent to implementing a nonlinear least squares method provided the distribution of $\eta_t$ is correctly specified. In addition, Lemma (ref) further infers that
so many remarks made previously can be directly applied. Last but not least, Lemma (ref) facilitates numerical implementation in practice, which will be further discussed in Appendix (ref).
Below, we establish the following asymptotic distribution.
In light of PA2012, we propose a score based wild bootstrap approach for inferential purposes as follows.
It is worth pointing out that the above procedure is computationally efficient in the sense that the right hand side of (ref) enjoys a closed-form expression, which is given in (ref) and (ref) specifically. Practically, the bootstrap procedure may require much less time compared with the estimation of (ref) itself.
The following theorem holds for the above bootstrap procedure.
We now provide extra simulations, which have two focuses: (1). examining the theoretical results in Appendix (ref), and (2). demonstrating the newly proposed method works reasonably well even without involving sparsity.
On the Fully Nonparametric Model --- Consider the following regression model:
where the variables are generated in the same manner as in Section (ref). In what follows, we only vary the values of $(q, d, u_\sigma)$. Specifically, we consider the cases $d=2, 8$. The rest parameters are identical to those in the main text.
To measure the finite sample performance, we select the points as follows:
With each dataset, we first estimate all $g(\mathbf{x}_{L,\mathbf{j}})$ using the approach of Appendix (ref), and then construct the corresponding 95% confidence interval using the bootstrap procedure documented in Theorem (ref) for each point. We report $\text{RMSE}_g$ and $\text{CR}_g$ which are defined in the main text already.
First, we draw some plots for the case with $d=2$. In both Figures (ref) and (ref), the first sub-plot is always the true $g(\mathbf{x})$. For the rest of sub-plots, each has three layers. The middle one is the average of estimates over $n$ replications. The top and bottom layers are the averages of the bootstrap draws corresponding to the 97.5% and 2.5% quantiles respectively. A few facts emerge. Overall, the LNN approach can recover the unknown function reasonably well. Also, both figures are very similar, so the results are not sensitive to the choice of $u_\sigma$ as explained in Remark (ref).
More detailed numbers are summarized Table (ref). As expected, when $T$ goes up, RMSE$_g$ converges to 0 and CR$_g$ converges to 0.95. Also, we note that when $d$ increases, RMSE$_g$ increases but still has reasonable performance expect the case with $(d=8,q=4)$. Therefore, it seems that for large $d$, smaller $q$ yields better finite sample performance.
On the Binary Model --- We consider the following data generating process:
in which $\eta_t = 0.5\eta_{t-1} + N(0, 0.75)$, the $j^{th}$ element of $\mathbf{x}_t$ is generated as $x_{t,j}\sim U(-a, a)$, and $G(\mathbf{x}) = 1+\sin(\mathbf{x}^\top \mathbf{1}_d/d)$. The rest parameters are identical to the simulation design of the fully nonparametric model.
We first note a computational issue. We reply on “fminunc" function of Matlab to solve
which is the same as that in (ref). In order to invoke the minimization process in any statistical software (including R, Matlab, etc.), one needs to provide initial values to the parameters under estimation. As a consequence, the numbers reported below are affected by the initial values more or less. Although it is not our intention to tackle this complicated computational issue in this paper, Lemma (ref) does become useful in this case. Recall that Lemma (ref) bridges the log likelihood estimation and the nonlinear least squares estimation. Therefore, we first conduct an OLS estimation using the approach of Appendix (ref) as the initial value of $\pmb{\Theta}$ for each generated $\{( z_t, \mathbf{x}_t)\, |\, t\in [T]\}$. We then invoke log likelihood estimation as our final estimate of $\pmb{\Theta}$ for each dataset. Even in this case, the computation is rather slow, and the computational time increases dramatically when the number of parameters goes up.
We draw a few plots in Figures (ref) and (ref). The first sub-plot is always the true $G(\mathbf{x})$. For the rest of sub-plots, each has three layers. The middle one is the average of estimates over $n$ replications. The top and bottom layers are the averages of the bootstrap draws corresponding to the 97.5% and 2.5% quantiles respectively. Overall, the LNN approach can recover the unknown function reasonably well. Also, both figures are very similar, so the results are not sensitive to the choice of $u_\sigma$.
We further summarize the detailed numbers in Table (ref). A few facts should be mentioned. First, the coverage rates are reasonably well. As expected, when $T$ goes up, RMSE converges to 0 and CR converges to 0.95. Also, we note that when $d$ increases, RMSE increases but still has reasonable performance. The results are not changing much with respect to the value of $u_\sigma$. Again, it seems that for large $d$, smaller $q$ yields better finite sample performance.
{
Before proving the theoretical results in Appendices (ref)-(ref), we first present all preliminary lemmas in Appendix (ref).
For $\mathbf{a}= (a_0, a_1,\ldots, a_d)^\top \in \mathbb{R}^{d+1}$, $\mathbf{r} = (r_0,r_1,\ldots, r_d)^\top\in \mathbb{N}_0^{d+1}$, and $ \mathbf{x}\in \mathbb{R}^{d}$, we let
Obviously, we have $f_{\mathbf{a}}\in \mathscr{P}_q$, where $ \mathscr{P}_q$ is defined in (ref).
Before presenting the next lemma, we calculate the partial derivatives of $\log L(\widetilde{s}(\cdot\, | \, \pmb{\Theta}))$ with respect to each $\pmb{\theta}_{\mathbf{i}}$.
where the third equality follows from the fact that $\mathbf{x}_t$ can not simultaneous belong to $C_{\mathbf{x}_{0\mathbf{i}},h}$ and $C_{\mathbf{x}_{0\mathbf{j}},h}$ for $\mathbf{i}\ne \mathbf{j}$ by the construction of $\widetilde{s}(\cdot\, | \, \pmb{\Theta})$, and $\widetilde{\mathbf{x}}_{\mathbf{i}, t}$ is the same as that defined in Section (ref).
Based on $\frac{\partial \log L(\widetilde{s}(\cdot\, | \, \pmb{\Theta})) }{\partial \pmb{\theta}_{\mathbf{i}}}$ and some tedious calculation, the second order derivative is
Proof of Lemma (ref):
This is Lemma 8 of BK2019, so the derivation is omitted. {$\blacksquare$}
Proof of Lemma (ref):
In what follows, let
It suffices to show that $f_{\mathbf{a}_1} (\mathbf{x}),\ldots,f_{\mathbf{a}_{d_q}} (\mathbf{x})$ are linearly independent. To do this, let $b_1,\ldots, b_{d_q}\in \mathbb{R}$ be such that
As we explained under (ref), the monomials involved in (ref) are linearly independent. Thus, (ref) implies that
Note that $\sharp\widetilde{\mathbf{r}}=d_q$ by design, so we can construct a one-to-one relationship between $k\in [d_q]$ and $\mathbf{r}\in \widetilde{\mathbf{r}}$. Using this relationship, we can construct $p(\mathbf{x})\in \mathscr{P}_q$ as follows:
which satisfies
Position 4 in SAUER2006191 implies that (ref) has the only solution $p(\cdot)=0$ in $\mathscr{P}_q$ for Lebesgue almost all $\mathbf{a}_1,\ldots,\mathbf{a}_{d_q}\in \mathbb{R}^{d+1}$, which in turn implies $b_1=\cdots =b_{d_q}=0$. The proof is now completed. {$\blacksquare$}
Proof of Lemma (ref):
Before proceeding further, we would like to point out that in what follows, $\mathtt{C}$ is a constant for the purpose of rescaling only.
By Assumption (ref).2, there is a point $u_\sigma\in \mathbb{R}$ such that none of the derivatives up to the order $q$ is 0 at $u_\sigma$. Thus, we construct the following one-layer NN:
in which the definitions of $\gamma_k$'s and $\beta_k$'s are obvious.
By Assumption (ref).2 again, $\sigma(\cdot)$ is $q+1$ times continuously differentiable. Thus, it can be expanded in a Taylor series with Lagrange remainder around $u_\sigma$ up to order $q$:
where $\xi_k\in [u_\sigma - \frac{k}{\mathtt{C}}\cdot |x-x_0|, u_\sigma + \frac{k}{\mathtt{C}}\cdot |x-x_0|]$ for all $0\le k\le q$.
Note that
where $ \Big\{
\Big\}$ is the Stirling number of the second kind. The Stirling number of the second kind describes the number of options to split a set of $j$ elements into $n$ non-empty subsets, which is equal to 0 for $0\le j<n$, and is equal to 1 for $j=n$. The result holds true for all $j, n\in \mathbb{N}$ (AS1972).
Thus, we can further simplify the right hand side of (ref), and write
which in connection with (ref) yields that
In view of Assumption (ref).2 and $\mathtt{C}$ being a fixed value, the proof is now completed. {$\blacksquare$}
Proof of Lemma (ref):
By Lemma (ref), we can reconstruct all of $\{m_i(\mathbf{x}\, |\, \mathbf{x}_0)\}$ as follows:
where
Note that the rotation matrix $\mathbf{D}$ is determined by $\pmb{\alpha}_{j}$'s only, so they are fixed.
Apparently, we have an issue of identification here, because for example we can arbitrarily rescale $\pmb{\alpha}_{j}$'s, and modify $\mathbf{D}$ accordingly without changing $m_i(\mathbf{x}\, |\, \mathbf{x}_0)$ as follows:
in which $ \mathbf{B}$ is full rank. Therefore, for the purpose of identification, we regulate $\pmb{\alpha}_j$'s as follows:
in which $\mathbf{I}_h=\operatorname*{\normalfont\textrm{diag}}\{h,\mathbf{I}_d \}$, and $\mathbf{W}$ is defined in the body of this lemma. As a result, for $\forall j\in [d_q]$,
so we can invoke Lemma (ref) later on.
Treating $(1,\mathbf{x}^\top-\mathbf{x}_0^\top)\pmb{\alpha}_{j}$ as a whole and using Lemma (ref), we write
where the last line follows from Lemma (ref). Also, $\pmb{\beta}$, $\pmb{\gamma}$ and $u_\sigma$ are known as discussed in Remark (ref). Thus, we can further write
where the last line follows from the facts that $d_q$ is fixed and $\|\pmb{\lambda} \|=O(1)$.
Finally, let $\widetilde{\pmb{\lambda}} =(\widetilde{\lambda}_1,\ldots, \widetilde{\lambda}_{d_q})^\top$ with $\widetilde{\lambda}_j = \pmb{\lambda}^\top \mathbf{d}_j$. Further, in view of the definitions of $\pmb{\sigma}(\mathbf{x}\, |\, \mathbf{x}_0 ) $ and $(\pmb{\pi}_1,\ldots, \pmb{\pi}_{d_q (q+1)})$ in the body of this lemma and (ref), the proof is then completed. {$\blacksquare$}
Proof of Lemma (ref):
Before starting the proof, we introduce a few notations to facilitate the development. First, recall that in the body of this theorem, we have defined
where $s (\mathbf{x}\, |\, \widetilde{\mathbf{x}}_{\mathbf{i}}, \widetilde{\pmb{\lambda}}_{\mathbf{i}}) = (\widetilde{\pmb{\lambda}}_{\mathbf{i}}\otimes \pmb{\gamma})^\top\pmb{\sigma} (\mathbf{x}\, |\, \widetilde{\mathbf{x}}_{\mathbf{i}} )$ by the definition of Lemma (ref). Second, note that the leading terms of the $q^{th}$ order Taylor expansion of $g(\mathbf{x}) $ at each $ \widetilde{\mathbf{x}}_{\mathbf{i}}\in (-a,a)^d$ can be written as follows:
where the definition of $\pmb{\lambda}_{\mathbf{i}}$ should be obvious in view of the definition of $\mathbf{m}(\mathbf{x}\, |\, \widetilde{\mathbf{x}}_{\mathbf{i}})$ according to (ref).
We are now ready to start the proof, and write
Note that by Lemma (ref) we choose $\widetilde{\pmb{\lambda}}_{\mathbf{i}}$ which fulfils the relationship: $\widetilde{\pmb{\lambda}}_{\mathbf{i}} =\mathbf{D}^\top \pmb{\lambda}_{\mathbf{i}}$. It is worth mentioning that although $(\widetilde{\mathbf{x}}_{\mathbf{i}}, \pmb{\lambda}_{\mathbf{i}})$ vary with respect to $\mathbf{i}$, the rotation matrix $\mathbf{D}$ in facts is solely determined by $\mathbf{W}$ of Lemma (ref). Therefore, without loss of generality, we can fix $\mathbf{W}$ over $\mathbf{i}$, as it is user chosen. Then $\mathbf{D}$ remains the same in view of the proof of Lemma (ref).
Next, we write
where the inequality follows from the definition of $I_{\mathbf{i},h}(\mathbf{x})$, and the last step follows from Lemma (ref). Also, we can obtain that
where the inequality follows from the definition of $I_{\mathbf{i},h}(\mathbf{x})$, and the last step follows from Lemma (ref) by letting $\widetilde{\pmb{\lambda}}_{\mathbf{i}} =\mathbf{D}^\top \pmb{\lambda}_{\mathbf{i}}$.
Therefore, based on (ref) and (ref), we obtain
where $p=q+s$. The proof is now completed. {$\blacksquare$}
Proof of Lemma (ref):
We adopt a similar strategy with that for Lemma A.1 of WX2009 to establish the results in Lemma (ref). For notational simplicity, let $\widetilde{Q}_\psi (\alpha,\pmb{\Theta})=\widetilde{Q} (\alpha,\pmb{\Theta})+\sum_{l=1}^{d_q}\psi_l\|\pmb{\Theta}_{l}\|$, where $\widetilde{Q} (\alpha,\pmb{\Theta})= \sum_{t=1}^T[y_t-z_t\alpha -\widetilde{s}(\mathbf{x}_t \, |\, \pmb{\Theta}) ]^2$. Let $\mathbf{B} =\{ \mathbf{b}_{\mathbf{i}}\, |\, \mathbf{i}\in [M]^d\}$ and $\|\mathbf{B}\|_H^2=\sum_{\mathbf{i}\in [M]^d}\|\mathbf{H}^{-1}\mathbf{D}^{\top,-1}\mathbf{b}_{\mathbf{i}}\|^2$, where $\mathbf{b}_{\mathbf{i}}$ is a $d_q\times 1$ vector of constants. Additionally, denote $d_T=\frac{1}{\sqrt{T}h^d}$.
By FL2001, it suffices to show that for any $\epsilon>0$, there exists a constant $C>0$ such that
We write
where $\pmb{\Lambda}_l$ and $\mathbf{B}_{D,l}$ are vectors that contain the $l$-th elements of $\pmb{\lambda}_{\mathbf{i}}$ and $\mathbf{D}^{\top,-1}\mathbf{b}_{\mathbf{i}}$, respectively.
For $R_1$, directly using the definition of $\widetilde{s}(\mathbf{x} \, |\, \pmb{\Theta})$ and $s(\mathbf{x} \,|\, \mathbf{x}_{0\mathbf{i}}, \pmb{\theta}_{\mathbf{i}})$ gives
We can further write
We can further write
where $\widetilde{\mathbf{x}}^\ast_{\mathbf{i},t} =I_{\mathbf{i},h} (\mathbf{x}_t)(z_t,\pmb{\sigma}( \mathbf{x}_t\, |\,\mathbf{x}_{0\mathbf{i}})^\top (\mathbf{I}_{d_q}\otimes \pmb{\gamma}^\top)^\top)^\top$ and $\mathbf{b}_{\mathbf{i}}^\ast=(a,\mathbf{b}_{\mathbf{i}}^\top)^\top$. With this notation, we can rewrite $R_{1,1}$:
where $\widetilde{\mathbf{H}}=\text{diag}(1, \mathbf{H}\mathbf{D})$, $\widetilde{\lambda}_{\mathbf{i},\min}$ denotes the smallest eigenvalue of $\frac{1}{Th^d}\sum_{t=1}^T \widetilde{\mathbf{H}} \widetilde{\mathbf{x}}^{\ast}_{\mathbf{i},t} \widetilde{\mathbf{x}}^{\ast\top}_{\mathbf{i},t}\widetilde{\mathbf{H}}^\top$, $\widetilde{\lambda}_{\min}=\min\{\widetilde{\lambda}_{\mathbf{i},\min},|, \mathbf{i}\in [M]^d\}$, and the second equality holds by the fact that $I_{\mathbf{i},h} (\mathbf{x}_t)I_{\mathbf{j},h} (\mathbf{x}_t)=0$ for $\mathbf{i}\ne \mathbf{j}$.
It suffices to explore $\frac{1}{Th^d}\sum_{t=1}^T\mathbf{H}\mathbf{D}\widetilde{\mathbf{x}}_{\mathbf{i},t} \widetilde{\mathbf{x}}^{\top}_{\mathbf{i},t}\mathbf{D}^\top\mathbf{H}$, $\frac{1}{Th^d}\sum_{t=1}^TI_{\mathbf{i},h} (\mathbf{x}_t)z_t^2$, and $\frac{1}{Th^d}\sum_{t=1}^Tz_t\widetilde{\mathbf{x}}^{\top}_{\mathbf{i},t}\mathbf{D}^\top\mathbf{H}$ before we can obtain the probability limit of $\widetilde{\lambda}_{\min}$.
We proceed with $\frac{1}{Th^d}\sum_{t=1}^TI_{\mathbf{i},h} (\mathbf{x}_t)z_t^2$. Note that
where the fourth equality can be proved by directly applying the Lipschitz continuity of $G(\mathbf{x})$ and $f_{\mathbf{x}}(\mathbf{x})$ in Assumption (ref).
Analogously, we can compute the second moment:
For the first term on the right-hand side of (ref), by (ref),
For the second term on the right-hand side of (ref),
where $\widetilde{\Phi}_{\eta,s}(\cdot,\cdot)$ denotes the joint CDF of $(\eta_1,\eta_{1+s})$, $f_{\mathbf{x},s}(\mathbf{x}, \mathbf{z})$ is the joint PDF of $(\mathbf{x}_1, \mathbf{x}_{1+s})$, and the last equality holds by the following results which are implied by the $\alpha$-mixing conditions in Assumption (ref),
Analogously, for the third term on the right-hand side of (ref), we have
Combining (ref), (ref), (ref), and (ref), we obtain
Together with (ref), it yields that
Then, we consider $\frac{1}{Th^d}\sum_{t=1}^Tz_t\widetilde{\mathbf{x}}^{\top}_{\mathbf{i},t}\mathbf{D}^\top\mathbf{H}$. Drawing upon (ref) and (ref), it follows from the definition of $\mathbf{H}$ and $\max_j|\mathbf{n}_j|=q$ that
Therefore, we only need to study the convergence of $\frac{1}{Th^d}\sum_{t=1}^TI_{\mathbf{i},h} (\mathbf{x}_t)z_t\mathbf{m}(\mathbf{x}_t\, |\, \mathbf{x}_{0\mathbf{i}})^\top \mathbf{H}$. We have
Using analogous arguments to those in (ref), we can show the convergence of its second moment and obtain that
Finally, we consider $\frac{1}{Th^d}\sum_{t=1}^T\mathbf{H}\mathbf{D}\widetilde{\mathbf{x}}_{\mathbf{i},t} \widetilde{\mathbf{x}}^{\top}_{\mathbf{i},t}\mathbf{D}^\top\mathbf{H}$ and write
where the first equality follows from Assumption (ref).1, the second equality follows from (ref), and the last step follows from Assumption (ref).2 and the definition of $\mathbf{m}(\mathbf{x}\, |\, \mathbf{0})$ given in (ref).
For the second moment, using the arguments that are analogous to the proof of (ref), we obtain
so we omit the details for now.
Together with (ref) and (ref), it proves that
where $\lambda_{\min,0}=\inf_{\mathbf{x}\in[-a,a]^d}\lambda_{\min}\left\{\pmb{\Omega}^\ast_0(\mathbf{x})\right\}$ with
Additionally, Assumption (ref) ensures that $\lambda_{\min,0}>0$.
For $R_{1,2}$, by Lemma (ref), Cauchy–Schwarz inequality, and using the analogous arguments to the derivations of $R_{1,1}$, we can obtain
For $R_{1,3}$, by (ref) and (ref), we can further write
Under Assumption (ref), it is clear to see that $\|R_{1,2}\|=o_P(d_T^2)$ and $\|R_{1,3}\|=o_P(d_T^2)$. Up to now, we have finished the investigation of $R_{1}$ and then proceed with $R_{2}$. Using simple algebra, we obtain
where $\psi_l^\ast = \psi_lH_l$, $\psi^\ast_{\max}=\max\{\psi^\ast_{l},l\in[d^0_q]\}$ and the third inequality holds by the fact $\sum_{l=1}^{d^0_q}H_l^{-1}\|\mathbf{B}_{D,l}\|\leq \sqrt{d^0_q\sum_{l=1}^{d^0_q}H_l^{-2}\|\mathbf{B}_{D,l}\|^2}=\sqrt{d^0_q}\|\mathbf{B}\|_H $. By Assumption (ref), it is straightforward to see that $\|R_2\|=o_P(d_T^2)$.
In summary of the results that are established in (ref), (ref), (ref), (ref), (ref), and (ref), it is implied by Assumption (ref) that the leading order term in $\frac{1}{d_T^2Th^d}\widetilde{Q}_\psi (\alpha_0+d_Tb_0,\widetilde{\pmb{\Lambda}}+d_T\mathbf{B})- \frac{1}{d_T^2Th^d}\widetilde{Q}_\psi (\alpha_0, \widetilde{\pmb{\Lambda}})$ is a quadratic function of $C$ with the coefficient for the quadratic term having the probability limit $\lambda_{\min,0}>0$ and the coefficients for the linear terms being bounded in probability by $O_P(1)$. Consequently, with a sufficiently large $C$, the left-hand side of (ref) is guaranteed to be positive with probability one. This completes the proof of Lemma (ref). {$\blacksquare$}
Proof of Lemma (ref):
(1). With knowledge of the true model ($g(\mathbf{x})=g_c(\mathbf{x}_c)$), we can define the following oracle LNN candidates:
where $\mathbf{x}_c$, $\mathbf{x}_{c,0\mathbf{i}}$, and $\pmb{\theta}_{c,\mathbf{i}}$ are oracle counterparts of $\mathbf{x}$, $\mathbf{x}_{0\mathbf{i}}$, and $\pmb{\theta}_{\mathbf{i}}$, respectively, and
where $\widetilde{\pmb{\lambda}}_c$, $\pmb{\gamma}_c$, and $\pmb{\pi}_{c,j}$ are the counterparts of $\widetilde{\pmb{\lambda}}$, $\pmb{\gamma}$, and $\pmb{\pi}_{j}$ in the true model. Noteworthily, we assume $\mathbf{m}_c(\mathbf{x}_c|\mathbf{x}_{c,0}) = (m_{c,1}(\mathbf{x}_c|\mathbf{x}_{c,0}),\ldots, m_{c,d_q^0}(\mathbf{x}_c|\mathbf{x}_{c,0}))^\top$ constitute a basis for the true space $\mathscr{P}^0_q$ without loss of generality, and
for $j=1,\ldots,d_q^0$.
Using arguments that are analogous to those in the proof of Lemma (ref), we can obtain
where $\widetilde{\pmb{\Lambda}}_c =\{\widetilde{\pmb{\lambda}}_{c,\mathbf{i}}\, |\, \mathbf{i}\in [M]^{d_0}\}$, $\widetilde{\pmb{\lambda}}_{c,\mathbf{i}}$ corresponds to $\pmb{\lambda}_{c,\mathbf{i}}$ up to a rotation matrix $\mathbf{D}_c$, and $\pmb{\lambda}_{c,\mathbf{i}}$ is decided by the Taylor expansion of $g_c(\mathbf{x}_c)$ at the point $\mathbf{x}_{c,0\mathbf{i}}$.
For the oracle estimators $\widetilde{\alpha}_{c}$ and $\widetilde{\pmb{\theta}}_{c,\mathbf{i}}$, it is easy to see that they satisfy the following first-order conditions:
By solving the equations in (ref), we obtain the expressions for $\widetilde{\pmb{\theta}}_{c,\mathbf{i}}$ and $\widetilde{\alpha}_{c}$:
where $\mathbf{Z}=(z_1,\cdots,z_T)^\top$, $\mathbf{Y}=(y_1,\cdots,y_T)^\top$, and $\mathbf{M}_{c,x}=\mathbf{I}_{T}-\sum_{\mathbf{i}\in [M]^{d_0}} \widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigl[\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top \widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigr]^{-1}\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top$ with $\widetilde{\mathbf{X}}_{c,\mathbf{i}} = (\widetilde{\mathbf{x}}_{c,\mathbf{i}, 1}, \cdots, \widetilde{\mathbf{x}}_{c,\mathbf{i}, T})^\top$.
Using (ref) and the expression in (ref), we can further expand $\widetilde{\alpha}_{c}$:
where $\pmb{\varepsilon}=(\varepsilon_1,\cdots,\varepsilon_T)^\top$. We then proceed with the convergence of $\frac{1}{T}\mathbf{Z}^\top \mathbf{M}_{c,x} \mathbf{Z}$. We write
For $Q_1$, it is clear to see that
For the second moment,
where $\widetilde{\Phi}_{\eta,s}(\cdot,\cdot)$ denotes the joint CDF of $(\eta_1,\eta_{1+s})$, $f_{\mathbf{x},s}(\mathbf{x}, \mathbf{z})$ is the joint PDF of $(\mathbf{x}_1, \mathbf{x}_{1+s})$, and the last equality holds by (ref), which is a standard result under the $\alpha$-mixing conditions in Assumption (ref). Combing (ref) and (ref), we obtain
For $Q_2$, using arguments that are analogous to those for (ref), (ref), and (ref), we can readily obtain
where $\mathbf{H}_c$ and $\mathbf{D}_c$ are counterpart matrices of $\mathbf{H}$ and $\mathbf{D}$ for the true model,
In summary of (ref), (ref), and (ref),
We then proceed with $\mathbf{Z}^\top \mathbf{M}_{c,x}\pmb{\varepsilon}$. Using arguments that are analogous to those in the proof of (ref), we can write
where the definitions of $\widetilde{Q}_1$, $\cdots$, $\widetilde{Q}_4$ are obvious.
For notational simplicity, let $\pmb{\xi}_{M,\mathbf{i}}=\frac{1}{Th^{d_0}}\sum_{t=1}^TI_{\mathbf{i},h} (\mathbf{x}_{c,t})z_t\mathbf{H}_c\mathbf{m}_c(\mathbf{x}_{c,t}\, |\, \mathbf{x}_{c,0\mathbf{i}}) -\mathbf{M}_{c,\mathbf{i}}$. For the first term, using the oracle counterpart of (ref), we can readily obtain that
To show the convergence of $\widetilde{Q}_1$, it suffices only to study $\widetilde{Q}_1^\ast$. It is clear to see that $E[\widetilde{Q}^\ast_1]=0$ under Assumption (ref) and {
}For $\widetilde{Q}_{1,1}^\ast$, using the arguments that are closely related to those in the proof of (ref), we can obtain
Therefore, $\widetilde{Q}_{1,1}^\ast=o\bigl(\frac{1}{T} \bigr)$. Analogously, for the second term
where the third inequality holds by Assumption (ref) and the Davydov's inequality for $\alpha$-mixing processes Bosq1996.
It immediately implies that the second and third terms on the right-hand side of (ref) are also asymptotically negligible. In summary of these results, we have $E[\widetilde{Q}_1^{\ast 2}]=o\bigl(\frac{1}{T} \bigr)$ and
Using similar arguments, we can obtain
Then, it suffices only to study $\widetilde{Q}_4$. For notational simplicity, we define
In what follows, we first compute the asymptotic covariance of $\frac{1}{\sqrt{T}}\sum_{t=1}^T\widetilde{z}_t\varepsilon_t$ and then employ the small-block and large-block technique for $\alpha$-mixing processes to establish its asymptotic normality.
Write
For the first term, we can use analogous arguments in the proof of (ref) and Assumption (ref) to show that
For $\widetilde{Q}_{4,2}$ and $\widetilde{Q}_{4,3}$, it is clear to see that $\widetilde{Q}_{4,2}=\widetilde{Q}_{4,3}$. It suffices only to study $\widetilde{Q}_{4,2}$. Write
{
}where $\sigma_{\varepsilon,s}^2=E[\varepsilon_1\varepsilon_{1+s}]$. Let $Q_0$ represent the limit of the terms on the right-hand side of (ref), as $T\rightarrow\infty$. Together with (ref) and Assumption (ref), it yields that
Below, we further use small-block and large-block to prove the normality. To employ the small-block and large-block arguments, we partition the set $\{1,\ldots, T \}$ into $2k_T+1$ subsets with large blocks of size $l_T$ and small blocks of size $s_T$ and the last remaining set of size $T-k_T(l_T+s_T)$, where $l_T$ and $s_T$ are selected such that
and $\nu$ is defined in Assumption (ref).1.
For $j=1,\ldots, k_T$, define
Note that $\alpha(T) = o(1/T)$ and $k_Ts_T/T\to 0$. By direct calculation, we immediately obtain that
Therefore,
By Proposition 2.6 of FanYao, we have as $T\to 0$
where $i$ is the imaginary unit.
In connection with (ref)-(ref), the Feller condition is fulfilled as follows:
Also, we note that
where the first inequality follows from H\"older inequality, the second inequality follows from Chebyshev's inequality, and the second equality follows from Minkowski inequality. Consequently,
where the last step follows from the choice of $l_T$ as specified above. Therefore, the Lindberg condition is justified. Using a Cram{\'e}r-Wold device, the CLT follows immediately by the standard argument:
Combing (ref) and (ref) leads to the desired result in Lemma (ref).1.
(2). We now study the asymptotic behaviour of $\widetilde{\pmb{\theta}}_{c,\mathbf{i}}$. Using its expression in (ref), we can further write
Using arguments that are analogous to those for (ref) and (ref), we can readily obtain
where $\pmb{\Sigma}_{c,\mathbf{i}}$ and $\mathbf{M}_{c,\mathbf{i}}$ are defined in (ref). Together with Lemma (ref).1 and (ref), these results yield
Drawing upon the $\alpha$-mixing conditions in Assumption (ref), we can use the arguments that are closely related to the small-block and large-block technique employed in the proof of Lemma (ref) to show that
Together with (ref), it completes the proof of Lemma (ref).2. {$\blacksquare$}
Proof of Lemma (ref):
(1). For any $l=d_q^0+1,\ldots,d_q$, if $\|\widetilde{\pmb{\Theta}}_{l}\|\neq 0$, we have the following first-order condition for the minimization problem in (ref):
where $\widetilde{Q} (\alpha,\pmb{\Theta})= \sum_{t=1}^T[y_t-z_t\alpha -\widetilde{s}(\mathbf{x}_t \, |\, \pmb{\Theta}) ]^2$.
We proceed with the derivations of $\mathbf{Q}_{l,1}$. In light of (ref), by taking first-order partial derivative of $\widetilde{Q} (\alpha,\pmb{\Theta})$ with respect to $\pmb{\theta}_{\mathbf{i},l}$, we obtain
where $\widetilde{x}_{\mathbf{i},tl}$ is the $l$-th element of $\widetilde{\mathbf{x}}_{\mathbf{i},t}$ and $\widetilde{\mathbf{x}}^\ast_{\mathbf{i},t} =I_{\mathbf{i},h} (\mathbf{x}_t)(z_t,\pmb{\sigma}( \mathbf{x}_t\, |\,\mathbf{x}_{0\mathbf{i}})^\top (\mathbf{I}_{d_q}\otimes \pmb{\gamma}^\top)^\top)^\top$, and $\widetilde{\pmb{\theta}}_{\mathbf{i}}^\ast=(\widetilde{a},\widetilde{\pmb{\theta}}_{\mathbf{i}}^ \top)^\top$.
Then, simple algebra gives
where $\widetilde{\pmb{\lambda}}_{\mathbf{i}}^\ast=(a_0,\widetilde{\pmb{\lambda}}_{\mathbf{i}}^ \top)^\top$.
It suffices only to study the first three terms on the right-hand side of (ref) to derive the convergence rate of $\mathbf{Q}_{l,1}$. Using (ref), (ref) and Lemma (ref), we can readily obtain that the first term has the probability order of $O_P(TH_l^{-2})$. For the second term, by Lemma (ref) and (ref), it has the probability order of $O_P(T^2h^{d+2p}H_l^{-2})$. Additionally, we can use analogous arguments in (ref) and show that the third term in (ref) is bounded in probability by $O_P(TH_l^{-2})$. In summary, we can establish the following result for $\mathbf{Q}_{l,1}$:
For $\mathbf{Q}_{l,2}$, it is straightforward to see $\|\mathbf{Q}_{l,2}\|=\psi_l\mathtt{c}$. Under the condition $\min_{d_q^0+1\leq l\leq d_q}\{\psi_lH_l\}T^{-\frac{1}{2}}\rightarrow \infty$, (ref) yields that
which leads to a contradictory result to that in (ref). Therefore, we must have $P\left(\|\widetilde{\pmb{\Theta}}_{D,l}\|= 0\right)\rightarrow 1$. It completes the proof of Lemma (ref).1.
(2). Together with Lemma (ref) and the fact that $g(\mathbf{x})=g_c(\mathbf{x}_c)$, (ref) yields
Recall that $\widetilde{Q} (\alpha,\pmb{\Theta})= \sum_{t=1}^T[y_t-z_t\alpha -\widetilde{s}(\mathbf{x}_t \, |\, \pmb{\Theta}) ]^2$. Simple algebra further gives $ \widetilde{s}_{c}(\mathbf{x}_c \, |\, \pmb{\Theta}_c)=\sum_{\mathbf{i}\in [M]^{d_0}} \widetilde{\mathbf{x}}^\top_{c,\mathbf{i},t}\pmb{\theta}_{c,\mathbf{i}}$, where $\widetilde{\mathbf{x}}_{c,\mathbf{i},t} =I_{\mathbf{i},h} (\mathbf{x}_{c,t})(\mathbf{I}_{d^0_q}\otimes \pmb{\gamma}_c^\top) \pmb{\sigma}_c( \mathbf{x}_{c,t}\, |\,\mathbf{x}_{c, 0\mathbf{i}})$.
For the group-LASSO estimators, we can formulate the following first-order conditions:
where $\pmb{\Phi}_c=\mathbf{D}^{-1}\pmb{\Phi}_D\mathbf{D}^{\top,-1}$ with $\pmb{\Phi}_D$ being a $d_q\times d_q$ diagonal matrix with its $l$-th diagonal element being $\psi_l/(2\|\widetilde{\pmb{\Theta}}_{D,l}\|)$, for $l=1,\ldots,d_q$.
Recall that $\widetilde{\pmb{\theta}}_{\mathbf{i},\star}$ contains the first $d_q^0$ elements in $\mathbf{D}^{\top,-1}\widetilde{\pmb{\theta}}_{\mathbf{i}}$. Let $\mathbf{D}_\star$ be the matrix that contains the first $d_q^0$ rows in $\mathbf{D}$ and $\widetilde{\mathbf{x}}_{\mathbf{i}, t,\star} = \mathbf{D}_\star\widetilde{\mathbf{x}}_{\mathbf{i}, t}$. Then, using the sparsity in $g(\mathbf{x})$, Lemma (ref).1, (ref), and (ref), we can rewrite (ref) as
where $\pmb{\Phi}_{D,\star}$ is a $d_q^0\times d_q^0$ diagonal matrix that contains the first $d_q^0$ diagonal elements of $\pmb{\Phi}_D$ and $\mathbf{H}_c$, as defined at an early stage, consists of the first $d_q^0$ elements in $\mathbf{H}$.
Solving (ref), we can establish the following results for $(\widetilde{\alpha}, \widetilde{\pmb{\theta}}_{\mathbf{i},\star})$:
where $\mathbf{M}_{x,\phi,\star}=\mathbf{I}_{T}-\sum_{\mathbf{i}\in [M]^{d_0}} \widetilde{\mathbf{X}}_{\mathbf{i},\star}\left[\widetilde{\mathbf{X}}_{\mathbf{i},\star}^\top \widetilde{\mathbf{X}}_{\mathbf{i},\star}+\pmb{\Phi}_{D,\star}\right]^{-1}\widetilde{\mathbf{X}}_{\mathbf{i},\star}^\top$ with $\widetilde{\mathbf{X}}_{\mathbf{i},\star} = (\widetilde{\mathbf{x}}_{\mathbf{i}, 1,\star}, \cdots, \widetilde{\mathbf{x}}_{\mathbf{i}, T,\star})^\top$.
In light of (ref) and (ref), it suffices only to study the asymptotic behaviour of $\mathbf{M}_{x,\phi,\star}$ before we can establish Lemma (ref).2. Define $\mathbf{M}_{c,x,\phi}=\mathbf{I}_{T}-\sum_{\mathbf{i}\in [M]^{d_0}} \widetilde{\mathbf{X}}_{c,\mathbf{i}}\left[\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top \widetilde{\mathbf{X}}_{c,\mathbf{i}}+\pmb{\Phi}_{D,\star}\right]^{-1}\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top$.
We write
Together with (ref), it yields that
where the last equality is ensured by the fact that $\{\psi_lH_l\}T^{-\frac{1}{2}}\rightarrow0$ and $h^{d_0}H^{-1}_l\|\widetilde{\pmb{\Theta}}_l\|\geq c$ hold uniformly for $l\in[d^0_q]$ and a positive constant $c$ under Lemma (ref) and the conditions in Assumption (ref).
In light of the definition of $\widetilde{\mathbf{x}}_{\mathbf{i}, t,\star} $ and $\widetilde{\mathbf{x}}_{c,\mathbf{i}, t}$, analogously to (ref) and (ref), we can obtain
Therefore, it is clear to see that $\frac{1}{T}\mathbf{Z}^\top\left(\mathbf{M}_{x,\phi,\star}-\mathbf{M}_{c,x,\phi}\right)\mathbf{Z}=O_P(h^{q+1})$. Together with (ref), it implies that $\frac{1}{T}\mathbf{Z}^\top\left(\mathbf{M}_{x,\phi,\star}-\mathbf{M}_{c,x}\right)\mathbf{Z}$ has the probability order of $o_P\left(\frac{1}{\sqrt{T}}+h^p\right)$. Analogously, we can show that $\frac{1}{T}\mathbf{Z}^\top\left(\mathbf{M}_{x,\phi,\star}-\mathbf{M}_{c,x}\right)\mathbf{Y}$ is also negligible. Therefore, we have
Therefore, the proof of Lemma (ref).2 is complete. {$\blacksquare$}
Proof of Theorem (ref):
(1). Directly applying Lemma (ref) and Lemma (ref), we can establish the desired result in Theorem (ref).1.
(2). Similarly to (ref) and (ref), we have
where $\widetilde{\mathbf{x}}_{c,\mathbf{i},0} =I_{\mathbf{i},h} (\mathbf{x}_{c,0})(\mathbf{I}_{d^0_q}\otimes \pmb{\gamma}_c^\top) \pmb{\sigma}_c( \mathbf{x}_{c,0}\, |\,\mathbf{x}_{c,0\mathbf{i}})$.
Together with Lemma (ref).1, (ref), (ref), and (ref), it gives
Combining (ref) and (ref), we obtain
We now study $\widetilde{\pmb{\theta}}_{\mathbf{i},\star}$. With (ref), similarly to (ref), we can further write
Drawing upon (ref) and the expressions for $\widetilde{\pmb{\theta}}_{c,\mathbf{i}}$ and $\widetilde{\pmb{\theta}}_{\mathbf{i},\star}$ that we have derived in (ref) and (ref), respectively,
Using the results that are established in (ref) and (ref), we can readily obtain the orders of the first three terms on the right-hand side of (ref) as $o_P\left(\frac{1}{\sqrt{T}}\right)$, $o_P\left(\frac{1}{T}\right)$ and $o_P\left(\frac{1}{\sqrt{T}}\right)$. Therefore, we have
Using the central limit theorem for $\widetilde{\pmb{\theta}}_{c,\mathbf{i}}$ in Lemma (ref), we can establish the asymptotic normality of $\sqrt{Th^{d_0}}(\widetilde{g}(\mathbf{x}_{0})-g_c(\mathbf{x}_{c,0}))$, which has the asymptotic covariance as the limit of
where $\pmb{\Sigma}_{c,\mathbf{i}}$ is defined in Lemma (ref). Then, the proof of Theorem (ref).2 is complete. {$\blacksquare$}
Proof of Theorem (ref):
(1) Using arguments that are analogous to those in the proof of Lemma (ref), we can show that the bootstrap Group-LASSO estimators can approximate the bootstrap oracle estimators up to some asymptotically negligible bias terms. Therefore, it suffices only to study the asymptotic behaviour of the bootstrap oracle estimators. Define the bootstrap oracle estimators $ (\widetilde{\alpha}^\ast_c,\widetilde{\pmb{\Theta}}^\ast_c)$ as follows:
where $\widetilde{Q}^\ast_c (\alpha,\pmb{\Theta}_c)= \sum_{t=1}^T[y^\ast_t-z_t\alpha -\widetilde{s}_c(\mathbf{x}_{c,t} \, |\, \pmb{\Theta}_c) ]^2$.
By solving the first-order conditions, we obtain the following expressions for $\widetilde{\pmb{\theta}}^\ast_{c,\mathbf{i}}$ and $\widetilde{\alpha}^\ast_{c}$:
where $\mathbf{Z}=(z_1,\cdots,z_T)^\top$, $\mathbf{Y}^\ast=(y^\ast_1,\cdots,y^\ast_T)^\top$, and $\mathbf{M}_{c,x}=\mathbf{I}_{T}-\sum_{\mathbf{i}\in [M]^{d_0}} \widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigl[\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top \widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigr]^{-1}\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top$ with $\widetilde{\mathbf{X}}_{c,\mathbf{i}} = (\widetilde{\mathbf{x}}_{c,\mathbf{i}, 1}, \cdots, \widetilde{\mathbf{x}}_{c,\mathbf{i}, T})^\top$.
Drawing upon the DGP of $y_t^\ast$ and the second expression in (ref), we can further expand $\widetilde{\alpha}^\ast_{c}$ as follows:
where $\widetilde{\mathbf{X}}_{\mathbf{i}} = (\widetilde{\mathbf{x}}_{\mathbf{i}, 1}, \cdots, \widetilde{\mathbf{x}}_{\mathbf{i}, T})^\top$ and $\pmb{\varepsilon}^\ast=(\varepsilon_1^\ast,\cdots,\varepsilon_T^\ast)^\top$.
For the first term, it is clear to see that
Together with Lemma (ref).1 and (ref), it immediately yields that the first term in (ref) is asymptotically negligible.
Recall that $\widetilde{z}_t=z_t-\sum_{\mathbf{i}\in [M]^{d_0}} I_{\mathbf{i},h} (\mathbf{x}_{c,t}) \mathbf{M}_{c,\mathbf{i}}^\top \pmb{\Sigma}_{c,\mathbf{i}}^{-1}\mathbf{H}_c\mathbf{m}_c(\mathbf{x}_{c,t}\, |\, \mathbf{x}_{c,0\mathbf{i}})$. For notational simplicity, we further define $\widetilde{z}_t^\ast = z_t-\sum_{\mathbf{i}\in [M]^{d_0}} \mathbf{Z}^\top\widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigl[\widetilde{\mathbf{X}}_{c,\mathbf{i}}^\top \widetilde{\mathbf{X}}_{c,\mathbf{i}}\bigr]^{-1}\widetilde{\mathbf{x}}_{c,\mathbf{i},t}$. For the second term on the right-hand side of (ref), we write
Using arguments that are analogous to those in the proofs of (ref) and (ref), we can show that $\mathcal{J}_1=o_P(1)$. For $\mathcal{J}_2$, write
For $\mathcal{J}_{2,1}$, it is straightforward to see that $\frac{1}{\sqrt{T}}\sum_{t=1}^TE[\widetilde{z}_tz_t\varsigma_t]=0$ and
It is obvious that the first term has the order $O(1)$ and the second and third terms have the same order. Therefore, it suffices only to study the second term. We have
In summary of these results, we have
with implies that $\frac{1}{\sqrt{T}}\sum_{t=1}^T\widetilde{z}_tz_t\varsigma_t=O_P(\sqrt{\ell})$. Together with Theorem (ref).1, it yields that
We now proceed with the derivations of $\mathcal{J}_{2,2}$. By Lemma (ref).1 and (ref), we obtain
Similarly to (ref), we can show that
Additionally, directly applying Lemma (ref).2, (ref), and Cauchy-Schwarz inequality yields
Therefore, we have
Combining (ref) and (ref), we obtain
Analogously, we can show that $\mathcal{J}_{3}$ is also asymptotically negligible.
In what follows, we proceed to explore $\mathcal{J}_{4}$ which generates the bootstrap distribution. Let $E^\ast[\cdot]$ and $\text{Var}^\ast(\cdot)$ denote the expectation and variance conditional on the observed sample. We first show that $\text{Var}^\ast( \mathcal{J}_{4})=\sigma_{c,z}^2+o_P(1)$.
It is clear to see that $ E^\ast\bigl[\mathcal{J}_{4}\bigr]=0$. Moreover, we write
Using a decomposition that is similar to (ref), we can easily show that $\mathcal{Q}_1=E[\widetilde{z}^2_1 \varepsilon^2_1]+o_P(1)$. Let $s_T$ satisfy that $s_T\rightarrow\infty$ and $s_T^2/\ell\rightarrow0$. For $\mathcal{Q}_2$, we write
For $\mathcal{Q}_{2,1}$, by Davydov's inequality for $\alpha$-mixing processes,
Together with Lipschitz continuity of the kernel function, it yields that
For $\mathcal{Q}_{2,2}$, by (ref), we have
The second equality holds by the fact that $\sum_{t=1}^{T-1} \alpha(t)^{\nu/(2+\nu)}$ and $\sum_{t=1}^{s_T} \alpha(t)^{\nu/(2+\nu)}$ have the same limit as $T\rightarrow\infty$ and $s_T\rightarrow\infty$, which is ensured by the order of $\alpha$-mixing coefficients.
In summary of the results that are established in (ref), (ref), and (ref), we can readily obtain
We then study $\mathcal{Q}_2-E[\mathcal{Q}_2]$. With $\nu$ and $\nu^\ast$ that are defined in Assumption (ref) and Theorem (ref), respectively, we can always define a positive number $r$ through the following equation:
Since $\nu^\ast>\nu$, it is clear to see that
Thus, we have $r>2$. For notational simplicity, we define a norm $\|\zeta\|_{n}=E\bigl[\|\zeta\|^{n}\bigr]^{1/n}$ for any random variable $\zeta$ and any positive number $n\geq 1$. We have
Additionally, let $\mathcal{F}_t$ and $E_{t}[\cdot]$ be the sigma field generated by $\{\widetilde{z}_s, \varepsilon_s\}_{s=t,t-1,\cdots}$ and the expectation conditional on $\mathcal{F}_t$, respectively. By McLeish's inequality for $\alpha$-mixing processes McLeish1975 and Assumption (ref),
for a positive integer $t_0$. Using this result and the Lemma A of Hansen1992, we can readily obtain
where $c_{\nu^\ast}=12\bigl\|\widetilde{z}_1 \bigr\|_{2+\nu^\ast} \bigl\|\varepsilon_1\bigr\|_{2+\nu^\ast}$ and $r^\ast=\min(r,4)$.
Combing (ref) and (ref) gives
Under the condition $\ell T^{2/r^\ast-1}\rightarrow 0$, we have
By (ref) and (ref), we can finish the investigation of $\mathcal{Q}_2$ and obtain
A similar argument applies with the index $s$ replacing $t$ for $\mathcal{Q}_3$, so it follows that
Drawing upon the definition of $\sigma_{c,z}^2$ in Assumption (ref) and the convergence of $\mathcal{Q}_1$, the results that are established in (ref) and (ref) can immediately yield
In light of the $\ell$-dependent $\varsigma_t$, we follow the Theorem 3.1 of Shao2010 and adopt the large-block and small-block argument to prove the central limit theorem for $\mathcal{J}_{4}$ conditional on the observed sample. Define $l_T$ and $s_T$ as the lengths for the large and small blocks and $k_T=\lfloor T/(l_T+s_T)\rfloor$ such that $l_T,s_T\rightarrow\infty$ and
For $j=1,\ldots, k_T$, define
We first show that $\frac{1}{\sqrt{T}}\sum_{j=1}^{k_T} \xi^\ast_{j,2}=o_P(1)$. Since $\frac{\ell}{s_T}\rightarrow 0$, we assume $l_T,s_T>\ell$ without loss of generality. By the definition of $\varsigma_t$, we can observe that $\{\xi^\ast_{1,1},\cdots,\xi^\ast_{k_T,1}\}$ are independent conditional on the observed data, as are $\{\xi^\ast_{1,2},\cdots,\xi^\ast_{k_T,2}\}$. Using (ref) and Assumption (ref), we have
Therefore, $\frac{1}{\sqrt{T}}\sum_{j=1}^{k_T} \xi^\ast_{j,2}=o_P(1)$. Analogously, we also have $\frac{1}{\sqrt{T}} \xi^\ast_{0}=o_P(1)$. Next, we establish the asymptotic normality of $\frac{1}{\sqrt{T}}\sum_{j=1}^{k_T} \xi^\ast_{j,1}$ by verifying the Lindeberg condition. Using analogous arguments to those in the proof of (ref), we can first show that $\frac{1}{T}E^\ast\bigl[\bigl(\sum_{j=1}^{k_T} \xi^\ast_{j,1}\bigr)^2 \bigr]=\sigma_{c,z}^2+o_P(1)$. Moreover, for any $\epsilon>0$, we can use the same argument as in (ref) to obtain
For notational simplicity, we define a norm (conditional on the observed sample) $\|\zeta\|^\ast_{n}=E^\ast\bigl[\|\zeta\|^{n}\bigr]^{1/n}$ for any random variable $\zeta$ and any positive number $n\geq 1$. In what follows, we use the Rosenthal inequality to study the order of $\|\xi^\ast_{1,1}\|^\ast_{2+\nu}$, without loss of generality. Noteworthily, Rosenthal inequality is designed for the independent random variables and it is not directly applicable for $\widetilde{z}_t\varepsilon_t\varsigma_t$. Therefore, we further decompose $\xi^\ast_{1,1}$ as $\xi^\ast_{1,1}=\sum_{k=1}^{\ell+1} \xi^\ast_{1,1,k}$, where
Then, by the definition of $\varsigma$, it is clear to see that for each $k$, all elements that are involved in the summation in $\xi^\ast_{1,1,k}$ are independent conditional on the observed sample. Using the triangle inequality and Rosenthal inequality sequentially, we obtain
Combining (ref) and (ref) gives
We therefore have $\frac{1}{T}\sum_{j=1}^{k_T}E^\ast\bigl[\xi^{\ast 2}_{j,1}\cdot I(|\xi^\ast_{j,1}|\geq \epsilon \sqrt{T}) \bigr]=o_P(1)$, which is the last step of the large-block and small-block technique and it immediately yields the following CLT: $\mathcal{J}_{4}\rightarrow_{D^\ast} N\left(0,\sigma_{c,z}^2\right)$, where $\rightarrow_{D^\ast}$ denotes the convergence in distribution conditional on the observed sample. Recall that we have shown that $\mathcal{J}_{1},\mathcal{J}_{2},$ and $\mathcal{J}_{3}$ are all asymptotically negligible at an earlier stage. Combining these results with (ref), (ref), and Theorem (ref).1 leads to the assertion in Theorem (ref).1.
(2) Let $\Omega_T$ denote the event in which the group-LASSO estimation has correctly identified the sparsity structure of the $g(\mathbf{x})$ function. That is $\|\widetilde{\pmb{\Theta}}_{D,j}\|=0$, for $j=d_q^0+1,\ldots,d_q$. By Lemma (ref), we have $P(\Omega_T)\rightarrow1$, as $T\rightarrow\infty$. Hence, it suffices only to study the asymptotic distribution of $\widetilde{g}^*(\mathbf{x}_0)$ conditional on the observed sample and $\Omega_T$.
On $\Omega_T$ and using analogous arguments to those in the proof of Theorem (ref), we can obtain the following result for the bootstrap estimator $\widetilde{g}^\ast(\mathbf{x}_{0})$:
where $\widetilde{\pmb{\theta}}^\ast_{c,\mathbf{i}}$ denotes the oracle bootstrap estimator.
Then, we can use the arguments that are closely related to those in the proof of (ref) to obtain
Then, similar steps to those in the derivation of $\frac{1}{\sqrt{T}}\mathbf{Z}^\top \mathbf{M}_{c,x}\pmb{\varepsilon}^\ast$'s bootstrap distribution in (ref) can be applied here to establish the bootstrap behaviour of the first term on the right-hand side of (ref). Specifically, we can obtain that it converges to $N\left(\mathbf{0},\sigma^2_\varepsilon \pmb{\Sigma}_{c,\mathbf{i}}^{-1}\right)$ conditional on the observed sample, up to some asymptotically negligible terms. In connection with (ref), it leads to the desired result in Theorem (ref).2. {$\blacksquare$}
Proof of Lemma (ref):
First, we expand the expression of $\widehat{\pmb{\theta}}_{\mathbf{i}}$ as follows:
where the third equality follows from the fact that $I_{\mathbf{i}, h}(\mathbf{x}_t)I_{\mathbf{j}, h}(\mathbf{x}_t)=0$ for $\mathbf{i}\ne \mathbf{j}$, and the definition of $\widetilde{s}(\mathbf{x}_t \, |\, \widetilde{\pmb{\Lambda}})$. Below, we consider the terms on the right-hand side one by one.
We further define
First, we consider $\frac{1}{Th^d} \sum_{t=1}^T\widetilde{\mathbf{x}}_{\mathbf{i},t} \widetilde{\mathbf{x}}_{\mathbf{i},t} ^\top$. By (ref) and (ref), we obtain
Also, we note
where
By Lemma (ref), it is easy to see that
Then we can write
where the first inequality follows from the exercise 5 on page 267 of Magnus, and the second inequality follows from (ref) and (ref).
Based on the above development, we can conclude that
Finally, in order to establish the asymptotic distribution, we just need to focus on $\frac{1}{\sqrt{Th^d}}\sum_{t=1}^T\widetilde{\mathbf{m}}(\mathbf{x}_t\, |\, \mathbf{x}_{\mathbf{i},0}) \varepsilon_t$ in view of (ref). Write
where the last two terms are the same up to a transpose operation.
For the first term, we can use similar argument to that in (ref) and obtain
Then, we use Assumption (ref) and the Davydov's inequality for $\alpha$-mixing processes (see pages 19-20 in Bosq1996) to show the convergence of the second term on the right-hand side of (ref). Specifically, we have
Moreover, it is clear to see that
By (ref) and (ref),
Analogously, the third term on the right-hand side of (ref) is also negligible.
Thus, we can conclude that
Below, we further use small-block and large-block to prove the normality. To employ the small-block and large-block arguments, we partition the set $\{1,\ldots, T \}$ into $2k_T+1$ subsets with large blocks of size $l_T$ and small blocks of size $s_T$ and the last remaining set of size $T-k_T(l_T+s_T)$, where $l_T$ and $s_T$ are selected such that
and $\nu$ is defined in Assumption (ref).1.
For $j=1,\ldots, k_T$, define
Note that $\alpha(T) = o(1/T)$ and $k_Ts_T/T\to 0$. By direct calculation, we immediately obtain that
Therefore,
By Proposition 2.6 of FanYao, we have as $T\to 0$
where $i$ is the imaginary unit.
In connection with (ref)-(ref), the Feller condition is fulfilled as follows:
Also, we note that
where the first inequality follows from H\"older inequality, the second inequality follows from Chebyshev's inequality, and the second equality follows from Minkowski inequality. Consequently,
where the last step follows from the choice of $l_T$ as specified above. Therefore, the Lindberg condition is justified. Using a Cram{\'e}r-Wold device, the CLT follows immediately by the standard argument. {$\blacksquare$}
Proof of Theorem (ref):
By Lemma (ref), we can write
where $\widetilde{\mathbf{x}}_{\mathbf{i},0} =I_{\mathbf{i},h} (\mathbf{x}_0)(\mathbf{I}_{d_q}\otimes \pmb{\gamma}^\top) \pmb{\sigma}( \mathbf{x}_0\, |\,\mathbf{x}_{0\mathbf{i}})$.
By (ref) and (ref), we obtain
which immediately yields
Together with Lemma (ref) and (ref), it implies that
Therefore, Lemma(ref) implies that $\sqrt{Th^d}(\widehat{g}(\mathbf{x}_0)-g(\mathbf{x}_0))$ is asymptotically normal with the asymptotic covariance being the limit of
where $ \pmb{\Sigma}_{\mathbf{i}} =f_{\mathbf{x}}(\mathbf{x}_{0\mathbf{i}}) \int_{[-1,1]^d} \mathbf{m}(\mathbf{x}\, |\, \mathbf{0}) \mathbf{m}(\mathbf{x}\, |\, \mathbf{0})^\top \mathrm{d} \mathbf{x}$.{$\blacksquare$}
Proof of Theorem (ref):
In what follows, we label the quantities associated with the bootstrap procedure by the superscript $^*$, which will not be further explained unless misunderstanding may arise.
By design, we have for $\forall \mathbf{i}$
where the second equality follows from the definition of $y_t^*$.
Note that the term $\mathbf{A}_{\mathbf{i,2}}$ has been investigated in the proof of Lemma (ref), and is negligible under the condition $\sqrt{T}h^{p+d/2}\to 0$. For the term $\mathbf{A}_{\mathbf{i,3}}$, we can further write
where the second equality follows from the fact that $\{\eta_t \}$ are i.i.d. draws from $N(0,1)$ and are independent of the sample.
Therefore, we only need to pay attention to $\mathbf{A}_{\mathbf{i},1}$ below. It suffices to consider $\frac{1}{\sqrt{T}}\sum_{t=1}^T\mathbf{\pmb\xi}_t^*$, where $\mathbf{\pmb\xi}_t^* =\frac{1}{\sqrt{h^d}}\widetilde{\mathbf{x}}_{\mathbf{i},t} \varepsilon_t \eta_t$. As $\{\eta_t \}$ are i.i.d. draws from $N(0,1)$, it is easy to know that
in view of the proof of Lemma (ref).
Below, we consider $E^*[\|\pmb{\xi}_1^*\|^2 \cdot I(\|\pmb{\xi}_1^*\| \ge \epsilon \sqrt{T})]$. Write
where the first inequality follows from H\"older inequality, the second inequality follows from Chebyshev's inequality, the third inequality follows from the definition of $E^*$ and Minkowski inequality, and the last step follows from $\frac{1}{h^d}E \| \widetilde{\mathbf{m}}(\mathbf{x}_1\, |\, \mathbf{x}_{\mathbf{i},0}) \varepsilon_1 \|^{2+\nu}=O(1)$ by the proof of Lemma (ref). Consequently,
where the last step follows from $Th^d\to 0$. Therefore, the Lindberg condition is justified. Then the result follows. {$\blacksquare$}
Proof of Corollary (ref):
The proof is a simpler version of that presented for Lemma (ref) and Theorem (ref), therefore it is omitted. {$\blacksquare$}
Proof of Lemma (ref):
(1). First, note that provided $0<x, x_0<1$, we have the following two expressions by the following Taylor expansions:
where both $x^*$ and $x^\dagger$ lie between $x$ and $x_0$.
We are now ready to start our investigation. By (ref) and (ref), write
where both $\Phi_t^*$ and $\Phi_t^\dagger$ lie between $\Phi_\eta(\widetilde{s}(\mathbf{x}_t\, | \, \pmb{\Theta}))$ and $\Phi_\eta(g(\mathbf{x}_t))$, and the definitions of $\mathbb{L}_{T,1}$ and $\mathbb{L}_{T,2}$ are obvious.
We then consider $\mathbb{L}_{T,1}$ and $\mathbb{L}_{T,2}$ respectively, and start with $\mathbb{L}_{T,1}$. For notational simplicity, we let
Simple algebra shows that
For any given $\Delta \Phi_\eta(g(\cdot))$, we then consider
where the inequality follows from Assumption (ref).1, and Davydov's inequality and (ref). By Lemmas A1 and A2 of WP2003, we immediately obtain that
We next investigate $\mathbb{L}_{T,2}$. Write
where the first inequality follows from $\frac{1}{(a+b)^2} \ge \frac{1}{2a^2 + 2b^2}$ because of $(a+b)^2\le 2a^2 + 2b^2$, the second inequality follows from the fact that $\Phi_t^*$ and $\Phi_t^\dagger$ lie between $\Phi_\eta(\widetilde{s}(\mathbf{x}_t\, | \, \pmb{\Theta}))$ and $\Phi_\eta(g(\mathbf{x}_t))$, and the third inequality follows from that $ \frac{1-y_t}{4\cdot 2} + \frac{y_t}{2} \ge \frac{1}{8}$ because of $z_t$ taking the value of 1 or 0 only.
By the fact that $0\ge\frac{1}{T}\log L(g(\cdot)) - \frac{1}{T}\log L(\widetilde{s}(\cdot\, | \, \widehat{\pmb{\Theta}}))$, and (ref) and (ref), we now conclude that
which completes the proof of this lemma. {$\blacksquare$}
By Lemma (ref), it is obvious that
where the third equality follows from the first result of Lemma (ref) and Lemma (ref).
Proof of Lemma (ref):
(1). Note that by Lemma (ref), we can write
where the third equality follows from the fact that $\mathbf{x}_t$ can not simultaneous belong to $C_{\mathbf{x}_{0\mathbf{i}},h}$ and $C_{\mathbf{x}_{0\mathbf{j}},h}$ for $\mathbf{i}\ne \mathbf{j}$. Thus, we must have
which completes the proof of the first result.
(2). By (ref), we denote
where the definitions of $ \mathbf{L}_j(\cdot)$ for $j=1,2,3$ should be obvious.
First, we consider $\mathbf{L}_2(s(\cdot \, | \, \mathbf{x}_{\mathbf{i},0},\pmb{\theta}_{\mathbf{i}} ))$ and $\mathbf{L}_3(s(\cdot \, | \, \mathbf{x}_{\mathbf{i},0},\pmb{\theta}_{\mathbf{i}} ))$. For $\mathbf{L}_2(s(\cdot \, | \, \mathbf{x}_{\mathbf{i},0},\pmb{\theta}_{\mathbf{i}} ))$, we write
where the definition of $ \widetilde{\mathbf{X}}_{\mathbf{i}, t} (\cdot) $ is obvious.
For the first term on the right hand side of (ref), we write
where the first inequality follows from Cauchy-Schwarz inequality, the second inequality follows from Mean-Value Theorem and the fact that $\phi_\eta(\cdot)$ is uniformly bounded, and the last step follows from the first result of this lemma and (ref).
For the second term on the right hand side of (ref), by some tedious algebra and the first result of this lemma, it is not hard to see that
Further, using Assumption (ref) and Billingsley's inequality following a procedure similar (but simplified) as in (ref), and in connection with (ref), we can show that
Based on the above development, we are readily to conclude that
Similar to the analysis of $ \widetilde{\mathbf{L}}_2(s(\mathbf{x}_t\, | \, \mathbf{x}_{\mathbf{i},0}, \widehat{\pmb{\theta}}_{\mathbf{i}} ))$, we can also obtain that
Below, we focus on $\mathbf{L}_1(s(\mathbf{x}_t\, | \, \mathbf{x}_{\mathbf{i},0}, \widehat{\pmb{\theta}}_{\mathbf{i}} ))$, and write
where the second equality follows from similar steps as those for $\widetilde{\mathbf{L}}_2(s(\mathbf{x}_t\, | \, \mathbf{x}_{\mathbf{i},0}, \widehat{\pmb{\theta}}_{\mathbf{i}} )) $, the third equality follows from a proof similar to those for (ref), and the fourth equality follows from a development similar to (ref).
Thus, we can now conclude that for each $\mathbf{i}$
where $\widetilde{\pmb{\Sigma}}_{\mathbf{i}}$ is defined in the body of this lemma.
The proof of the second result is now completed.
(3). In view of the fact that $\pmb{\theta}_{\mathbf{i}}^*$ lies between $\widehat{\pmb{\theta}}_{\mathbf{i}}$ and $ \widetilde{\pmb{\lambda}}_{\mathbf{i}}$, the result follows immediately by going through the same procedure as the second result of this lemma.
(4). Write
where $\widetilde{\mathbf{m} }(\mathbf{x}_t\, |\, \mathbf{x}_{\mathbf{i},0})$ is defined in (ref), the second equality follows from (ref), and in the third equality we let
for notational simplicity. Moreover, as $g$ is defined on $[-a, a]$, it is easy to know that $ 0<\mathtt{c}\le \Phi_\eta(g(\mathbf{x}_t))\le \mathtt{C}<1$. Thus, we can further write
Also, simple algebra shows that
We now move on and write
Similar to (ref), we have
Thus , we can conclude that
where the last step follows from a procedure similar to (ref).
Below, we further use small-block and large-block to prove the normality. To employ the small-block and large-block arguments, we partition the set $\{1,\ldots, T \}$ into $2k_T+1$ subsets with large blocks of size $l_T$ and small blocks of size $s_T$ and the last remaining set of size $T-k_T(l_T+s_T)$, where $l_T$ and $s_T$ are selected such that
where $\nu$ is defined in Assumption (ref).1
For $j=1,\ldots, k_T$, define
Note that $\alpha(T) = o(1/T)$ and $k_Ts_T/T\to 0$. By direct calculation, we immediately obtain that
Therefore,
By Proposition 2.6 of FanYao, we have as $T\to 0$
where $i$ is the imaginary unit. Thus, the Feller condition is fulfilled as follows:
Also, we note that
where the first inequality follows from H\"older inequality, the second inequality follows from Chebyshev's inequality, and the last step follows from Minkowski inequality. Consequently,
which is the Lindberg condition. Using a Cram{\'e}r-Wold device, the CLT follows immediately by the standard argument. {$\blacksquare$}
Proof of Theorem (ref):
By the first order condition, we have
where the second equality follows from (ref).
Using Taylor expansion, we have
where $\pmb{\theta}_{\mathbf{i}}^*$ lies between $\widehat{\pmb{\theta}}_{\mathbf{i}}$ and $ \widetilde{\pmb{\lambda}}_{\mathbf{i}}$, and the second equality follows from Lemma (ref) and the continuity of $\phi_\eta$ and $\Phi_\eta$.
Thus, by Lemma (ref), the result follows immediately. {$\blacksquare$}
Proof of Theorem (ref):
The proof follows from (ref) and (ref) and a procedure very similar to that given in Theorem (ref). {$\blacksquare$}
}
}