Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
49,015 characters · 14 sections · 42 citation commands
Nonparametric Cointegrating Regression Functions with Endogeneity and Semi-Long Memory
Regression models that relate two time series have long been of interest in economics, and any number of quantities of interest in such models are nonstationary. If two variables in a linear regression model are integrated (in the time series sense) of the same order but some linear combinations of them are stationary, then the model is said to be a cointegrated regression model. That label is also used to refer to models in which the covariate or regressor process is nonstationary and the response process is related to it through a nonlinear function (see tjostheim2020some for a recent review). Nonlinear cointegrated regression models have been applied, among others, to problems in environmental Kuznets curve, wind turbine data sets, and stock prices (e.g. fasen2013time and references therein). In particular, let
be a nonlinear cointegrating regression model, where $x_k$ is a nonstationary regressor, $u_k$ is an error process, and $f(\cdot)$ is an unknown real function.
Two problems that have been the focus of much research are the estimation of the unknown regression function $f(\cdot)$ nonparametrically and the testing that $f(\cdot)$ has a particular parametric form. Many of the theoretical results related to these problems have assumed strict exogeneity in which the regressor $x_k$ is assumed to be uncorrelated with the regression error $u_k$ (e.g. karlsen2007nonparametric,cai2009functional, WP2009a, wang2014martingale). Often, the regressor $x_k$ is assumed to be a short memory process. Extending this framework by allowing the regressor $x_k$ to be driven by long memory innovations and permitting correlation with $u_k$ (so-called endogeneity) has received less attention in the literature and leads to some technical problems that have not been completely resolved.
Recently, WP2016 established a limit theory for nonparametric estimation of regression models that allow for both endogeneity and long memory in the regressor process. They also developed a test statistic for parametric forms of regression functions of type $f(x)=g(x,\theta_0)$ where $g(\cdot,\theta)$ represents a parametric family of functions with an unknown true parameter value $\theta_0 \in \Theta$. The authors assumed that $\Theta$ is a compact subspace of $\mathbb{R}^m$ for some finite $m$. The test of WP2016 is a modification of a test statistic given by hardle1993comparing for the random sample case. The test was also used in gao2012model for a nonlinear cointegrating model with a martingale error structure and no endogeneity.
A property of the test under long memory and endogeneity is that the limit distribution depends on the value of the fractional differencing parameter $d$ in the nonstationary regressor $x_k$. The motivation for the work reported here was that if one employs the idea of tempering and assumes that there are semi-long memory input shocks to the regressor process $x_k$, the limit distribution of the test statistic given by WP2016 no longer depends on the unknown parameter $d$. We describe this idea in detail and list the consequences which are related to the asymptotic theory for nonparametric cointegrating regression functions.
Here, we summarize the main contributions of our paper as follows. First, we demonstrate the asymptotic properties of the kernel estimator of the unknown regression function $f(x)$ in ((ref)) through semi-long memory and endogeneity and propose its confidence intervals. Second, we develop test statistic for the parametric forms of the regression function under the null hypothesis, where we show that the limit distribution of the test statistic under semi-long memory involves the local time of standard Brownian motion and is free from the fractional differencing parameter $d$. Third, we investigate the properties of the regression function estimator as well as the performance of the test statistic through simulation studies. Finally, we explain how the test statistic can be used in real applications. To the best of our knowledge, these objectives have not been formally studied in the existing literature.
The remainder of the article is organized as follows. The regressor process $x_k$ is described in Section (ref). In Section (ref), we provide some initial results that are suitable for developing the limit theory presented in the following sections. In Section (ref), we introduce nonlinear cointegrating regression models under the assumption that the regressor process involves strongly tempered shocks. We establish the limit distribution of the estimator of the regression function.
In Section (ref), we present a test for parametric forms of the regression function, show that its asymptotic distribution is free of the fractional differencing parameter, and consider its power under local alternatives. Section (ref) contains the results of simulation studies for the estimation of the regression function and the properties of the test statistic. Section (ref) describes the application of the test statistic utilizing an actual data set. Section (ref) contains concluding remarks and some directions for future work. The proofs of all technical results and additional simulations are contained in the Appendix. All the R code implementing the proposed methodology is available at Github repository \url{https://github.com/SepidehMosaferi/Nonparametric-Regression-Functions}.
Throughout the paper, we denote $C, C_1, C_2, ...$ as generic constants which may differ at each appearance. We use $\rightarrow_{P}$ for convergence in probability, $\Rightarrow$ for weak convergence of the associated probability measures, $\rightarrow_{D}$ for convergence in distribution, and $=_{D}$ for equivalence in distribution. For any two functions $f$ and $g$, $f \asymp g $ means $C_1\leq f/g \leq C_2$. We use $||x||=\max_i|x_i|$ for the vector $x=(x_i)$, $\wedge$ for the minimum between two numbers, and a.s. for almost surely. Finally, i.i.d. and f.d.d. mean independent and identically distributed and finite-dimensional distribution, respectively.
The models we consider throughout this article are of the general form ((ref)) in which the regressor process is the sum of input shocks that result in either a long memory process or a semi-long memory process. In the long memory process, we take for $k=1, \ldots, N$
where $\zeta(s)$ is an i.i.d. noise with $\mathbb{E}(\zeta(0))=0$ and $\mathbb{E}(\zeta^2(0))=1$. The coefficient $b_d(j)$ regularly varies at infinity as $j^{d-1}$. In this article, we use
where $\Gamma(d)$ is the gamma function defined as $\int_{0}^{\infty} e^{-x} x^{d-1} dx$, and $ 0<d<1/2$ is the fractional differencing parameter.
In contrast to the long memory process, a semi-long memory process contains strongly tempered shocks as
where $\zeta(s)$ and $b_d(j)$ are the same as for the long memory case. In ((ref)), $\lambda \equiv \lambda_N >0$ is called the tempering parameter and is sample size dependent, which satisfies the following main assumption:
The tempering parameter $\lambda$ in expression ((ref)) allows the range of the fractional differencing parameter $d$ to be extended from $(0,1/2)$ to $(0,\infty)$.
The strongly tempered process $\eqref{SLMx}$ belongs to the general class of stochastic processes called tempered linear processes; see sabzikar2018invariance. These processes have a semi-long memory property in the sense that their autocovariance functions initially resemble that of a long memory process but eventually decay fast at an exponential rate. One special case of such processes is the autoregressive tempered fractionally integrated moving average ARTFIMA$(p,d,\lambda,q)$ process with $p$ as the order of the autoregressive polynomial and $q$ as the order of the moving average polynomial, which has been studied by meerschaert2014tempered and sabzikar2019parameter.
The class of ARTFIMA$(0,d,\lambda,0)$ that does not have autoregressive and moving average components is highly applicable. Some empirical examples include modeling logarithmic returns for AMZN stock prices, modeling geophysical turbulence in water velocity data, and modeling climate data sets given in sabzikar2019parameter. In all cases, the authors have found that by using an ARTFIMA model, we are able to capture aspects of the low-frequency activity of time series better than an ARFIMA model without considering $\lambda$ in the model.
The local time process of a stochastic process $G(x)$ is defined as
where $(t,s) \in \mathbb{R}_{+} \times \mathbb{R}$. Let $d_N:= [\mathbb{E}(x_N^2)]^{1/2}$ and consider $x_{k,N}:=x_k/d_N$ for $1 \leq k \leq N$ to be a triangular array. A function of $x_{k,N}$ that will be used in the sequel takes the form of a sample average of functions of $x_{k,N}$ as
where $g(x)$ is a bounded function such that $\int_{\mathbb{R}} |g(x)| dx< \infty$ and $c_N:= d_N/h$ where $c_N \rightarrow \infty$ and $c_N/N \rightarrow 0$. The bandwidth parameter is denoted by $h$ and satisfies $h \equiv h_N \rightarrow 0$ as $N \rightarrow \infty$. Functions in the form of $S_N$ commonly arise in nonlinear cointegrating regressions, which could be the kernel function $K(.)$ or its squared $K^2(.)$; see karlsen2001nonparametric, karlsen2007nonparametric, and WP2009a.
The limit behavior of $S_N$ will be important in establishing the limit behavior of the kernel estimator of $f(x)$ in the context of nonparametric regression. The limit distribution of $S_N$ is as follows
where $L_{B}(t,s)$ is the local time of standard Brownian motion $B(t)$ at the spatial point $s$; see Proposition (ref) part (i) to follow. When the function $g(.)$ is a kernel density, the integral in expression ((ref)) is unity, and the limit is then the local time of $B(.)$ at the origin; see jeganathan2004convergence and WP2009a for some of the related results. In the case of long memory, the local time given in expression ((ref)) is in the form of $L_{B_{d+1/2}}(t,s)$ and therefore depends on the unknown parameter $d$ (WP2009a). This complicates the limit theory of the kernel estimator of regression function as well as the associated test statistic for long memory processes.
Assume the model ((ref)) where $f(\cdot)$ is an unknown real regression function. To induce endogeneity, let $\eta_k=(\zeta(k),\epsilon(k))'$ be a sequence of random vectors with $\mathbb{E}(\eta_0)=0$ and $\mathbb{E}(\eta_0 \eta_0')=\Sigma$ where
and take $u_k$ in ((ref)) to be $u_k=\sum_{j=0}^{\infty} \psi_j \eta_{k-j}$ for the coefficient vector $\psi_j=(\psi_{j1},\psi_{j2})$. Therefore, $\mathbb{E}(u_0^2)=\sum_{j=0}^{\infty} \psi_j \Sigma \psi_j'$ and $cov(u_k,x_k) \neq 0$ by recalling expression ((ref)). Now, assume $\eta_0$ and $\psi_j$ satisfy the following assumption.
Assumption (ref) allows the error term $u_k$ to be cross-correlated with the regressor $x_k$. Furthermore, assume that the characteristic function $\varphi(t)$ of $\zeta(0)$ satisfies $\int_\mathbb{R} (1+|t|) |\varphi(t)| dt < \infty,$ that ensures smoothness in the corresponding kernel density (see WP2016).
The Nadaraya-Watson regression estimator of $f(x)$ is
where $K_h(s)=h^{-1}K(s/h)$, and $K(\cdot)$ is a non-negative bounded continuous function. In order to establish the asymptotic behavior for the kernel estimate $f(x)$, $\{ x_{k,N} \}_{k \geq 1, N \geq 1}$ needs to be a strong smooth array (see Definition (ref) as well as Proposition (ref) in the Appendix). Now, we impose the following two assumptions on the kernel $K(.)$:
Assumption (ref) is needed for some of the technical proofs and is satisfied for many commonly used kernels that the Fourier transformation of $K(x)$ is integrable. The assumption (ref) can be used in developing analytic forms for the asymptotic bias function in kernel estimation and can be verified for various kernels $K(x)$ and regression functions $f(x)$. The following theorem is the main result on the Nadaraya-Watson kernel estimator of ((ref)) and recall that $\lambda$ is a function of $N$ such that $\lambda \rightarrow 0$ and $N \lambda \rightarrow \infty$ as $N \rightarrow \infty$.
For a fixed value of $x$, $f(x)$ has a continuous $(p+1)$ derivative in a small neighborhood of $x$. For any $h$ satisfying $\sqrt{N} \lambda^d h \rightarrow \infty$ as $h \rightarrow 0$ and $\sqrt{N} \lambda^d h^{2(p+1)+1} \rightarrow 0$, we have $\hat{f}(x) \rightarrow_{P} f(x)$. Therefore, $\hat{f}(x)$ is a consistent estimator of $f(x)$. The consistency of $\hat{f}(x)$ could be followed from Theorem 3.1 of WP2009b.
The term $\mathbb{E}(u^2_0)$ in $d_0^2$ can be estimated by
Expression ((ref)) is appropriate for use in the interval estimation and other inferential procedures. When $\mathbb{E}(u_0^8)< \infty$ and $\int_{\mathbb{R}} K(s) f_1^2(s,x) ds < \infty$ for a given $x$, for any $h$ satisfying $\sqrt{N} \lambda^d h \rightarrow \infty$ and $h \rightarrow 0$, we have $\hat{\sigma}^2_N \rightarrow_{P} \mathbb{E}(u_0^2)$. The consistency of $\hat{\sigma}^2_N$ is stated in Theorem 3.2 in WP2009b, and we refer an interested reader to their manuscript.
When there is no reason to believe that $f(x)$ in ((ref)) follows a particular parametric form, the use of a nonparametric estimator is attractive, but nonparametric estimators generally have a slow convergence rate compared to parametric estimators. It is often possible to determine a plausible parametric regression function and then conduct a test of a hypothesis formulated as,
where $\theta_0$ in ((ref)) is a vector of unknown parameters that belongs to a compact and convex space $\Theta$.
The tests of ((ref)) have been previously considered by hardle1993comparing, horowitz2001adaptive, gao2009specification, wang2012specification, and WP2016 under different assumptions on the data generating mechanism. In fact, gao2009specification and wang2012specification considered a kernel-smoothed U statistic of the form $\sum_{j,k=1, j \neq k}^{N} \hat{u}_k \hat{u}_j K[(x_k-x_j)/h]$ with $\hat{u}_k=y_k-g(x_k,\hat{\theta}_N)$, where $\hat{\theta}_N$ is an estimator of $\theta$. The estimator $\hat{\theta}_N$, can be based on a non-linear least squares method by setting $Q_N(\theta)=\sum_{k=1}^{N}(y_k-g(x_k,\theta))^2$ and minimizing it over $\theta \in \Theta$ to obtain $\hat{\theta}_N = \text{argmin}_{\theta \in \Theta} Q_N(\theta).$
The asymptotic for the U statistic is difficult to extend to the case of endogenous regressors. WP2016 modified a test statistic previously suggested by hardle1993comparing for the consideration of endogenous regressors with long memory. In that case, both the limit theory and the convergence rate of the test statistic depend on the fractional differencing parameter $d$ through the local time of fractional Brownian motion $L_{B_{d+1/2}}(t,s)$, which complicates its use in actual problems. In this section, we consider the statistic under the assumption of semi-long memory input shocks to the regressors $x_k$. Both the limit distribution of WP2016 under long memory, and our limit distribution for use with semi-long memory regressors are based on the following statistic,
The term $\pi(x)$ in ((ref)) is a positive integrable weight function with a compact support. To develop the asymptotic theory of $T_{N}$ in the context of semi-long memory regressors and endogeneity, we provide three assumptions as follows.
As noted in WP2016, Assumption (ref) covers a wide range of functions $g(x,\theta)$ and $\pi(x)$. Typical examples of $g(x,\theta)$ include $\theta e^x/(1+e^x)$, $e^{- \theta |x|}$, $e^{- \theta x^2}$, $e^{\theta |x|}/(1+e^{\theta |x|})$, $\theta \log|x|$, $(x+\theta)^2$, etc. Note that Assumptions (ref)--(ref) can be compared with those imposed by WP2016. In particular, the convergence rate imposed on $\hat{\theta}_N$ by Assumption (ref) has been verified in Section 4 of WP2016 and will be required in our proof of the theorems to follow. We now consider the asymptotic behavior of $T_{N}$ in ((ref)) and its convergence rate in the semi-long memory setting.
To verify that the test has nontrivial power, one can assess the null hypothesis ((ref)) against a local alternative
Here, $\rho_N$ is a sequence of numerical constants that measures the local deviation from the null hypothesis. Also in ((ref)), $m(x)$ is a real function free of $\theta$ without lying in the span of $g(x,\theta)$ and its derivative functions. To make $m(x)$ smooth enough under $H_A$ for the sake of asymptotic power development, we give the following two assumptions.
The conditions on $m(x)$ in Assumption (ref) are week and can be satisfied by a large class of real functions. They are required to ensure the divergence of the test statistic and its consistency under $H_A$.
In this section, we illustrate the properties of the Nadaraya-Watson regression estimator and its confidence intervals. In addition, we study the size and power of the specification test statistic using the same regression functions.
To examine the behavior of regression function estimators, we generate data sets from a model with the form of ((ref)), where we incorporate $\sigma$ as follows:
The regressor process $x_k$ is defined in ((ref)) for the semi-long memory (SLM) setting and as $x_k=x_k^\circ$ where $x_k^\circ$ is given in ((ref)) for the long memory (LM) setting. Let $u_k=\psi u_{k-1}+\epsilon(k)$ with $\psi=0.25$ and $\mathbb{E}(\zeta(k) \epsilon(k)):= \rho_{\zeta,\epsilon}$ (c.f., ((ref))) such that $(\zeta(k), \epsilon(k))$ are i.i.d. $N\Big(0,
\Big)$. Also, let $\rho_{\zeta,\epsilon}=0.5$ and $\sigma=0.2$ in ((ref)).
We consider the following two regression functions:
We set the sample size at $N=\{ 50, 100, 500 \}$, and the number of Monte Carlo replications at $R=2000$.
We use the estimator ((ref)) with an Epanechnikov kernel in the form of $K(u)=0.75(1-u^2)1_{\{|u| \leq 1\}}$. For simplicity, we let $h = \{N^{-1/3}, N^{-1/7}\}$ and $\lambda = N^{-1/5}$ under the assumption of $N \lambda \rightarrow \infty$. We assume $d=\{ 0.1, 0.2, 0.3, 0.4\}$. We study Monte Carlo approximations of absolute bias ($\mathbb{E}\{|\hat{f}(x)-f(x)|\}$), standard deviation (std) ($[\mathbb{E}\{[\hat{f}(x)-E(\hat{f}(x))]^2\}]^{0.5}$) and root of mean squared error (rmse) ($[\mathbb{E}\{[\hat{f}(x)-f(x)]^2\}]^{0.5}$) of $\hat{f}(x)$ for the true regression functions over the interval $[-1,1]$ at $100$ points equally spaced between $-1$ and $1$.
To construct point-wise confidence intervals for the regression function $f(x)$, we use the limit distribution given in ((ref)). An asymptotic $100(1-\alpha)\%$ level confidence interval for $f(x)$ is then given by
where $\hat{\sigma}^2_N$ is given in expression ((ref)). We present the values of absolute bias, std, and rmse graphically for all the values of $x$ and for the two regression functions with $N=500$ and $h=N^{-1/7}$ in Figures (ref) and (ref). For the SLM case, these criteria are stable across values of $d$, while they vary as $d$ changes in the LM case and this is true for both straight-line and quadratic regression functions.
In Figures (ref) and (ref), we illustrate Monte Carlo coverage probabilities of confidence intervals and $(1-\alpha)=0.95$ point-wise intervals with their expected lengths. We observe that the expected lengths and coverage probabilities remain roughly similar for different values of $d$ under the SLM scenario, which is not the case for the LM scenario. Note that coverage was always reduced from the nominal level of $0.95$, but much less so for SLM processes than for LM processes. Coverage also decreased as $d$ increased. Additional results for smaller sample sizes and $h=N^{-1/3}$ are given in the Appendix.
In order to investigate the finite sample performance of test statistic, we first construct the Monte Carlo distributions of test statistic under both long and semi-long memory processes with endogeneity. Then, we use resampling methods to study the size and power of test statistic since their limiting distributions have complex forms and are not practical. In particular, we use the subsampling technique developed by mosaferi2024properties.
Drawing on the application to be presented in Section (ref) we consider two regression functions, a straight line ($y_k=\theta_0+\theta_1 x_k + \sigma u_k$) and a quadratic ($y_k=\theta_0+\theta_1 x_k + \theta_2 x_k^2+ \sigma u_k$), where $(\theta_0,\theta_1,\theta_2)=(0,1,1)$. For simplicity, we let $h=N^{-1/3}$ and $N=\{50, 100, 500 \}$. We assume that the amount of endogeneity is $\rho_{\zeta,\epsilon}=0.5$ and $d=\{0.1, 0.2, 0.3, 0.4, 0.5, 1, 1.5\}$, where $d \geq 0.5$ is only for the case of SLM, again based on experience with the application of Section (ref). Additionally, we let $\lambda=N^{-1/5}$. Other details of the simulation studies conducted to examine the behavior of the tests are the same as those described in Section (ref).
Monte Carlo densities of the test statistic are given in Figures (ref) and (ref). We observe that the densities become further from zero as the value of $d$ increases under the LM case. On the other hand, the densities overlap quite a lot for the SLM case and in particular for $0<d<1/2$.
To study the size of the test, we use the straight line and quadratic regression function assumptions to formulate $H_0$. To study the power of test, we use the $H_A$ given in equation ((ref)). We assume $\rho_N=1/N$ and $m(x_k)=|x_k|$. Therefore, we have $y_k=\theta_0+\theta_1 x_k + \frac{1}{N} |x_k| + \sigma u_k$ and $y_k=\theta_0+\theta_1 x_k + \theta_2 x_k^2+ \frac{1}{N} |x_k| + \sigma u_k$. For the block sizes, we choose the values of $\{[0.5 \sqrt{N}], [\sqrt{N}], [2 \sqrt{N}], [4 \sqrt{N}] \}$. The size and power values were close to $1$ for the LM and SLM cases. These results are not given in the manuscript, but allowed us to identify a bias in the form of the test statistic. We rectify this bias issue and construct a de-biased test statistic using the subsampling procedure and the algorithm described in Section C of mosaferi2024properties.
The values of size and power for the de-biased test statistic are given in Tables (ref) and (ref). Values of size for the straight line model are quite reasonable in both LM and SLM cases for small block sizes but decrease as block size increases. The size for the quadratic model is depressed from nominal, also decreases as block size increases, but is considerably more stable for SLM than LM situations. Power is low across the board, also decreases with block size, and is less stable for LM cases than for SLM ones. Additional results for $d \geq 0.5$ under SLM and smaller sample sizes are given in the Appendix.
In this section, we illustrate the performance of the test statistic given in equation ((ref)) in the analysis of environmental Kuznets curves. In particular, we focus on the relationship of carbon dioxide emissions (CO$_2$) v.s. gross domestic product (GDP) per capita in Spain and sulfur dioxide emission (SO$_2$) v.s. GDP per capita in the United Kingdom. We consider the natural logarithmic transformation of these quantities ($x_k$ and $y_k$) before proceeding. The data set for Spain is from 1950 to 2008 and contains 59 observations. The data set for the United Kingdom is from 1870 to 1999 and contains 130 observations. The data sets are provided on the Github repository of the first author. As mentioned in the previous section, the limit theory of this test does not have a simple form and is not practical, so we again rely on the use of subsampling to compute benchmarks.
Scatter-plots of the relevant variables are shown in Figure (ref). For Spain, the pattern of points suggests a straight-line regression function, while for the United Kingdom the pattern appears to be more quadratic in form. Thus, we test the null hypothesis of $f(x)=\theta_0+\theta_1 x$ for Spain and two null hypotheses of (1) $f(x)=\theta_0+\theta_1 x$ and (2) $f(x)=\theta_0+\theta_1 x + \theta_2 x^2$ for the United Kingdom. We also overlay the Nadaraya-Watson (N-W) regression function estimator $\hat{f}(x)$ from ((ref)) on the plots in Figure (ref). Optimal values for the bandwidth $h$ were selected using cross-validation (e.g. park1990comparison,hardlebook) with the help of N-W regression functions. The optimal bandwidth was $h=0.151$ for Spain and $h=0.107$ for the United Kingdom. The associated least squares values based on $\text{LCV}(h)=\sum_{k=1}^{N}(y_k - \hat{f}_{-k}(x_k))^2$ are given in Table (ref), where $\hat{h}_{\text{optimal}}=\text{argmin}_{h>0}\text{LCV}(h)$. To estimate the values of the fractional differencing parameter $d$ and the tempering parameter $\lambda$, we use their Whittle estimators from artfima package in R (e.g., sabzikar2019parameter). The estimated values of these parameters are also given in Table (ref). Based on the previous section and due to the biased form of test statistic $T_{\lambda,d}$, we consider both the biased and de-biased form of tests for use with subsampling. The resultant p-values are summarized in Table (ref). Furthermore, the p-values for the quadratic null hypothesis for the United Kingdom based on all possible block sizes values are displayed in Figure (ref).
For Spain, we observe that the null hypothesis of a straight line would be rejected when we use the biased form of test statistic, but with the de-biased test statistic this hypothesis would not be rejected. For the United Kingdom, the null hypothesis of a straight line is rejected for both the biased and de-biased forms of the test statistic. A hypothesized quadratic form for the regression function would not be rejected for either the biased or de-biased test statistics. Based on Figure (ref), including a greater number of blocks in the de-biasing process increases the p-values associated with the hypothesis of a quadratic regression function. In particular, for the last plot in Figure (ref), we can see that when 20 blocks are included in de-biasing, the p-value is one for all block sizes.
In this article we have considered the use of tempered or semi-long memory regressor processes within the context of cointegrating regression models for time series analysis. We have developed limit theory that allows the construction of pointwise confidence intervals for a nonlinear regression function when estimation is accomplished through the use of a kernel smoother. We have also demonstrated that the limit distribution of a test statistic for the specification of the regression function does not depend on the differencing parameter $d$ in semi-long memory. Through simulation studies, we have shown that performance of interval estimators of a regression function is more stable and has generally better performance for the semi-long memory case than for long memory. We also developed the limit distribution of the test statistic for their parametric forms and showed that it is free from the unknown parameter of $d$.
Some directions for future work could be related to multivariate nonparametric cointegrating regression with multiple regressors such that they have a semi-long memory property with error terms belonging to the $\alpha$-stable finite-dimensional distributions with infinite variance and infinite lower-order moments. We assume the regression disturbances are serially and cross-correlated with the regressors and permit the associated bandwidth in the regression model to be random and data dependent. One can investigate the uniform convergence rate for the kernel regression estimator and provide model specification test for multivariate cointegrating regressions where the asymptotic distribution of the test statistic is free of any unknown parameters and is proportional to a local time of linear stable motions.
We would like to thank the Editor, Associate Editor, and two anonymous reviewers for their excellent suggestions and comments that have substantially improved the quality of our manuscript.
No potential conflict of interest was reported by the authors.
\setcounter{secnumdepth}{0}