EconBase
← Back to paper

Nonparametric Cointegrating Regression Functions with Endogeneity and Semi-Long Memory

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

49,015 characters · 14 sections · 42 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonparametric Cointegrating Regression Functions with Endogeneity and Semi-Long Memory

abstractThis article develops nonparametric cointegrating regression models with endogeneity and semi-long memory. We assume that semi-long memory is produced in the regressor process by tempering of random shock coefficients. The fundamental properties of long memory processes are thus retained in the regressor process. Nonparametric nonlinear cointegrating regressions with serially dependent errors and endogenous regressors driven by long memory innovations have been considered in WP2016. That work also implemented a statistical specification test for testing whether the regression function follows a parametric form. The limit theory of test statistic involves the local time of fractional Brownian motion. The present paper modifies the test statistic to be suitable for the semi-long memory case. With this modification, the limit theory for the test involves the local time of the standard Brownian motion and is free of the unknown parameter $d$. Through simulation studies, we investigate the properties of nonparametric regression function estimation as well as test statistic. We also demonstrate the use of test statistic through actual data sets. \noindentKeywords: Fractional differencing parameter; test statistic; standard Brownian motion; simulation studies; tempered process.

Introduction

Regression models that relate two time series have long been of interest in economics, and any number of quantities of interest in such models are nonstationary. If two variables in a linear regression model are integrated (in the time series sense) of the same order but some linear combinations of them are stationary, then the model is said to be a cointegrated regression model. That label is also used to refer to models in which the covariate or regressor process is nonstationary and the response process is related to it through a nonlinear function (see tjostheim2020some for a recent review). Nonlinear cointegrated regression models have been applied, among others, to problems in environmental Kuznets curve, wind turbine data sets, and stock prices (e.g. fasen2013time and references therein). In particular, let

equation[equation omitted — 70 chars of source]

be a nonlinear cointegrating regression model, where $x_k$ is a nonstationary regressor, $u_k$ is an error process, and $f(\cdot)$ is an unknown real function.

Two problems that have been the focus of much research are the estimation of the unknown regression function $f(\cdot)$ nonparametrically and the testing that $f(\cdot)$ has a particular parametric form. Many of the theoretical results related to these problems have assumed strict exogeneity in which the regressor $x_k$ is assumed to be uncorrelated with the regression error $u_k$ (e.g. karlsen2007nonparametric,cai2009functional, WP2009a, wang2014martingale). Often, the regressor $x_k$ is assumed to be a short memory process. Extending this framework by allowing the regressor $x_k$ to be driven by long memory innovations and permitting correlation with $u_k$ (so-called endogeneity) has received less attention in the literature and leads to some technical problems that have not been completely resolved.

Recently, WP2016 established a limit theory for nonparametric estimation of regression models that allow for both endogeneity and long memory in the regressor process. They also developed a test statistic for parametric forms of regression functions of type $f(x)=g(x,\theta_0)$ where $g(\cdot,\theta)$ represents a parametric family of functions with an unknown true parameter value $\theta_0 \in \Theta$. The authors assumed that $\Theta$ is a compact subspace of $\mathbb{R}^m$ for some finite $m$. The test of WP2016 is a modification of a test statistic given by hardle1993comparing for the random sample case. The test was also used in gao2012model for a nonlinear cointegrating model with a martingale error structure and no endogeneity.

A property of the test under long memory and endogeneity is that the limit distribution depends on the value of the fractional differencing parameter $d$ in the nonstationary regressor $x_k$. The motivation for the work reported here was that if one employs the idea of tempering and assumes that there are semi-long memory input shocks to the regressor process $x_k$, the limit distribution of the test statistic given by WP2016 no longer depends on the unknown parameter $d$. We describe this idea in detail and list the consequences which are related to the asymptotic theory for nonparametric cointegrating regression functions.

Our contributions

Here, we summarize the main contributions of our paper as follows. First, we demonstrate the asymptotic properties of the kernel estimator of the unknown regression function $f(x)$ in ((ref)) through semi-long memory and endogeneity and propose its confidence intervals. Second, we develop test statistic for the parametric forms of the regression function under the null hypothesis, where we show that the limit distribution of the test statistic under semi-long memory involves the local time of standard Brownian motion and is free from the fractional differencing parameter $d$. Third, we investigate the properties of the regression function estimator as well as the performance of the test statistic through simulation studies. Finally, we explain how the test statistic can be used in real applications. To the best of our knowledge, these objectives have not been formally studied in the existing literature.

Organization of the paper

The remainder of the article is organized as follows. The regressor process $x_k$ is described in Section (ref). In Section (ref), we provide some initial results that are suitable for developing the limit theory presented in the following sections. In Section (ref), we introduce nonlinear cointegrating regression models under the assumption that the regressor process involves strongly tempered shocks. We establish the limit distribution of the estimator of the regression function.

In Section (ref), we present a test for parametric forms of the regression function, show that its asymptotic distribution is free of the fractional differencing parameter, and consider its power under local alternatives. Section (ref) contains the results of simulation studies for the estimation of the regression function and the properties of the test statistic. Section (ref) describes the application of the test statistic utilizing an actual data set. Section (ref) contains concluding remarks and some directions for future work. The proofs of all technical results and additional simulations are contained in the Appendix. All the R code implementing the proposed methodology is available at Github repository \url{https://github.com/SepidehMosaferi/Nonparametric-Regression-Functions}.

Throughout the paper, we denote $C, C_1, C_2, ...$ as generic constants which may differ at each appearance. We use $\rightarrow_{P}$ for convergence in probability, $\Rightarrow$ for weak convergence of the associated probability measures, $\rightarrow_{D}$ for convergence in distribution, and $=_{D}$ for equivalence in distribution. For any two functions $f$ and $g$, $f \asymp g $ means $C_1\leq f/g \leq C_2$. We use $||x||=\max_i|x_i|$ for the vector $x=(x_i)$, $\wedge$ for the minimum between two numbers, and a.s. for almost surely. Finally, i.i.d. and f.d.d. mean independent and identically distributed and finite-dimensional distribution, respectively.

The regressor process

The models we consider throughout this article are of the general form ((ref)) in which the regressor process is the sum of input shocks that result in either a long memory process or a semi-long memory process. In the long memory process, we take for $k=1, \ldots, N$

align[align omitted — 123 chars of source]

where $\zeta(s)$ is an i.i.d. noise with $\mathbb{E}(\zeta(0))=0$ and $\mathbb{E}(\zeta^2(0))=1$. The coefficient $b_d(j)$ regularly varies at infinity as $j^{d-1}$. In this article, we use

equation*[equation* omitted — 120 chars of source]

where $\Gamma(d)$ is the gamma function defined as $\int_{0}^{\infty} e^{-x} x^{d-1} dx$, and $ 0<d<1/2$ is the fractional differencing parameter.

In contrast to the long memory process, a semi-long memory process contains strongly tempered shocks as

align[align omitted — 152 chars of source]

where $\zeta(s)$ and $b_d(j)$ are the same as for the long memory case. In ((ref)), $\lambda \equiv \lambda_N >0$ is called the tempering parameter and is sample size dependent, which satisfies the following main assumption:

itemize[leftmargin=*] • {\bf Semi-Long Memory Assumption}: The tempering parameter $\lambda\to 0$ and $N\lambda\to\infty$ as $N\to\infty$.

The tempering parameter $\lambda$ in expression ((ref)) allows the range of the fractional differencing parameter $d$ to be extended from $(0,1/2)$ to $(0,\infty)$.

The strongly tempered process $\eqref{SLMx}$ belongs to the general class of stochastic processes called tempered linear processes; see sabzikar2018invariance. These processes have a semi-long memory property in the sense that their autocovariance functions initially resemble that of a long memory process but eventually decay fast at an exponential rate. One special case of such processes is the autoregressive tempered fractionally integrated moving average ARTFIMA$(p,d,\lambda,q)$ process with $p$ as the order of the autoregressive polynomial and $q$ as the order of the moving average polynomial, which has been studied by meerschaert2014tempered and sabzikar2019parameter.

The class of ARTFIMA$(0,d,\lambda,0)$ that does not have autoregressive and moving average components is highly applicable. Some empirical examples include modeling logarithmic returns for AMZN stock prices, modeling geophysical turbulence in water velocity data, and modeling climate data sets given in sabzikar2019parameter. In all cases, the authors have found that by using an ARTFIMA model, we are able to capture aspects of the low-frequency activity of time series better than an ARFIMA model without considering $\lambda$ in the model.

Initial results

The local time process of a stochastic process $G(x)$ is defined as

equation*[equation* omitted — 131 chars of source]

where $(t,s) \in \mathbb{R}_{+} \times \mathbb{R}$. Let $d_N:= [\mathbb{E}(x_N^2)]^{1/2}$ and consider $x_{k,N}:=x_k/d_N$ for $1 \leq k \leq N$ to be a triangular array. A function of $x_{k,N}$ that will be used in the sequel takes the form of a sample average of functions of $x_{k,N}$ as

equation*[equation* omitted — 68 chars of source]

where $g(x)$ is a bounded function such that $\int_{\mathbb{R}} |g(x)| dx< \infty$ and $c_N:= d_N/h$ where $c_N \rightarrow \infty$ and $c_N/N \rightarrow 0$. The bandwidth parameter is denoted by $h$ and satisfies $h \equiv h_N \rightarrow 0$ as $N \rightarrow \infty$. Functions in the form of $S_N$ commonly arise in nonlinear cointegrating regressions, which could be the kernel function $K(.)$ or its squared $K^2(.)$; see karlsen2001nonparametric, karlsen2007nonparametric, and WP2009a.

The limit behavior of $S_N$ will be important in establishing the limit behavior of the kernel estimator of $f(x)$ in the context of nonparametric regression. The limit distribution of $S_N$ is as follows

equation[equation omitted — 90 chars of source]

where $L_{B}(t,s)$ is the local time of standard Brownian motion $B(t)$ at the spatial point $s$; see Proposition (ref) part (i) to follow. When the function $g(.)$ is a kernel density, the integral in expression ((ref)) is unity, and the limit is then the local time of $B(.)$ at the origin; see jeganathan2004convergence and WP2009a for some of the related results. In the case of long memory, the local time given in expression ((ref)) is in the form of $L_{B_{d+1/2}}(t,s)$ and therefore depends on the unknown parameter $d$ (WP2009a). This complicates the limit theory of the kernel estimator of regression function as well as the associated test statistic for long memory processes.

Nonparametric regression function estimation

Assume the model ((ref)) where $f(\cdot)$ is an unknown real regression function. To induce endogeneity, let $\eta_k=(\zeta(k),\epsilon(k))'$ be a sequence of random vectors with $\mathbb{E}(\eta_0)=0$ and $\mathbb{E}(\eta_0 \eta_0')=\Sigma$ where

equation[equation omitted — 175 chars of source]

and take $u_k$ in ((ref)) to be $u_k=\sum_{j=0}^{\infty} \psi_j \eta_{k-j}$ for the coefficient vector $\psi_j=(\psi_{j1},\psi_{j2})$. Therefore, $\mathbb{E}(u_0^2)=\sum_{j=0}^{\infty} \psi_j \Sigma \psi_j'$ and $cov(u_k,x_k) \neq 0$ by recalling expression ((ref)). Now, assume $\eta_0$ and $\psi_j$ satisfy the following assumption.

myassumptionLet $\sum_{j=0}^{\infty}\psi_j \neq 0$ and $\sum_{j=0}^{\infty}j^{1/4}(|\psi_{j1}|+|\psi_{j2}|)< \infty$. Also, let $\mathbb{E}||\eta_0||^{\alpha} < \infty$ for $\alpha>2$.

Assumption (ref) allows the error term $u_k$ to be cross-correlated with the regressor $x_k$. Furthermore, assume that the characteristic function $\varphi(t)$ of $\zeta(0)$ satisfies $\int_\mathbb{R} (1+|t|) |\varphi(t)| dt < \infty,$ that ensures smoothness in the corresponding kernel density (see WP2016).

The Nadaraya-Watson regression estimator of $f(x)$ is

equation[equation omitted — 102 chars of source]

where $K_h(s)=h^{-1}K(s/h)$, and $K(\cdot)$ is a non-negative bounded continuous function. In order to establish the asymptotic behavior for the kernel estimate $f(x)$, $\{ x_{k,N} \}_{k \geq 1, N \geq 1}$ needs to be a strong smooth array (see Definition (ref) as well as Proposition (ref) in the Appendix). Now, we impose the following two assumptions on the kernel $K(.)$:

myassumption$K(x)$ is a non-negative bounded continuous function that satisfies $\int_{\mathbb{R}} K(x) dx=1$ and $\int_{\mathbb{R}} |\hat{K}(x)|dx < \infty$, where $\hat{K}(x)=\int_{\mathbb{R}} e^{ixt} K(t) dt$.
myassumptionFor a given $x$, there exists a real positive function $f_1(s,x)$ and $\gamma \in (0,1]$ such that when $h$ is sufficiently small, $|f(hy+x)-f(x)| \leq h^\gamma f_1(y,x) \, \forall y \in \mathbb{R}$ and $\int_{\mathbb{R}}K(s) [f_1(s,x)+f_1^2(s,x)]ds < \infty$ hold.

Assumption (ref) is needed for some of the technical proofs and is satisfied for many commonly used kernels that the Fourier transformation of $K(x)$ is integrable. The assumption (ref) can be used in developing analytic forms for the asymptotic bias function in kernel estimation and can be verified for various kernels $K(x)$ and regression functions $f(x)$. The following theorem is the main result on the Nadaraya-Watson kernel estimator of ((ref)) and recall that $\lambda$ is a function of $N$ such that $\lambda \rightarrow 0$ and $N \lambda \rightarrow \infty$ as $N \rightarrow \infty$.

mytheoremUnder Assumptions (ref)--(ref) and for any $h$ satisfying $\sqrt{N}\lambda^d h \rightarrow \infty$ and $ \sqrt{N} \lambda^d h^{1+2\gamma} \rightarrow 0$ where $\gamma \in (0,1]$, we have \begin{equation} \Big\{\sqrt{N}\lambda^d h\Big\}^{1/2} \Big(\hat{f}(x)-f(x)\Big) \rightarrow_D d_0 \, N(0,1) \, L_{B}^{-1/2}(1,0), \end{equation} where $d_0^2=\mathbb{E}(u^2_{0}) \int_{\mathbb{R}} K^2(s) ds$ and $N(0,1)$ denotes a standard normal which is independent from the local time of Brownian motion $L_{B}(1,0)$. If we normalize the limit form ((ref)), we have: \begin{equation} \Big\{h \sum_{k=1}^{N} K_h(x_k-x) \Big\}^{1/2} \Big(\hat{f}(x)-f(x)\Big) \rightarrow_D N(0,\sigma^2), \end{equation} where $\sigma^2= d_0^2$; see also expression ((ref)) given in the Appendix.

For a fixed value of $x$, $f(x)$ has a continuous $(p+1)$ derivative in a small neighborhood of $x$. For any $h$ satisfying $\sqrt{N} \lambda^d h \rightarrow \infty$ as $h \rightarrow 0$ and $\sqrt{N} \lambda^d h^{2(p+1)+1} \rightarrow 0$, we have $\hat{f}(x) \rightarrow_{P} f(x)$. Therefore, $\hat{f}(x)$ is a consistent estimator of $f(x)$. The consistency of $\hat{f}(x)$ could be followed from Theorem 3.1 of WP2009b.

myremarkThe local time given in Theorem (ref) is the local time of standard Brownian motion, which is independent of the unknown parameter $d$. This is a direct consequence of tempering the time series in the regressor $x_{k}$. In this regard, the result of Theorem (ref) is different from Theorem 2.1 in WP2016, which involves the local time process $L_{B_{d+1/2}}(1,0)$ and is related to a fractional Brownian motion with parameter $d+1/2$.
myremarkBased on Theorem (ref), the bandwidth $h$ has to satisfy certain rate conditions to ensure that the asymptotic distribution on the right side of expression ((ref)) holds. Under semi-long memory and for $d > 0$, we require $\sqrt{N} \lambda^d h \rightarrow \infty$ and $\sqrt{N} \lambda^d h^{1+2 \gamma} \rightarrow 0$.

The term $\mathbb{E}(u^2_0)$ in $d_0^2$ can be estimated by

equation[equation omitted — 127 chars of source]

Expression ((ref)) is appropriate for use in the interval estimation and other inferential procedures. When $\mathbb{E}(u_0^8)< \infty$ and $\int_{\mathbb{R}} K(s) f_1^2(s,x) ds < \infty$ for a given $x$, for any $h$ satisfying $\sqrt{N} \lambda^d h \rightarrow \infty$ and $h \rightarrow 0$, we have $\hat{\sigma}^2_N \rightarrow_{P} \mathbb{E}(u_0^2)$. The consistency of $\hat{\sigma}^2_N$ is stated in Theorem 3.2 in WP2009b, and we refer an interested reader to their manuscript.

Specification test for the regression function

When there is no reason to believe that $f(x)$ in ((ref)) follows a particular parametric form, the use of a nonparametric estimator is attractive, but nonparametric estimators generally have a slow convergence rate compared to parametric estimators. It is often possible to determine a plausible parametric regression function and then conduct a test of a hypothesis formulated as,

equation[equation omitted — 53 chars of source]

where $\theta_0$ in ((ref)) is a vector of unknown parameters that belongs to a compact and convex space $\Theta$.

The tests of ((ref)) have been previously considered by hardle1993comparing, horowitz2001adaptive, gao2009specification, wang2012specification, and WP2016 under different assumptions on the data generating mechanism. In fact, gao2009specification and wang2012specification considered a kernel-smoothed U statistic of the form $\sum_{j,k=1, j \neq k}^{N} \hat{u}_k \hat{u}_j K[(x_k-x_j)/h]$ with $\hat{u}_k=y_k-g(x_k,\hat{\theta}_N)$, where $\hat{\theta}_N$ is an estimator of $\theta$. The estimator $\hat{\theta}_N$, can be based on a non-linear least squares method by setting $Q_N(\theta)=\sum_{k=1}^{N}(y_k-g(x_k,\theta))^2$ and minimizing it over $\theta \in \Theta$ to obtain $\hat{\theta}_N = \text{argmin}_{\theta \in \Theta} Q_N(\theta).$

The asymptotic for the U statistic is difficult to extend to the case of endogenous regressors. WP2016 modified a test statistic previously suggested by hardle1993comparing for the consideration of endogenous regressors with long memory. In that case, both the limit theory and the convergence rate of the test statistic depend on the fractional differencing parameter $d$ through the local time of fractional Brownian motion $L_{B_{d+1/2}}(t,s)$, which complicates its use in actual problems. In this section, we consider the statistic under the assumption of semi-long memory input shocks to the regressors $x_k$. Both the limit distribution of WP2016 under long memory, and our limit distribution for use with semi-long memory regressors are based on the following statistic,

equation[equation omitted — 152 chars of source]

The term $\pi(x)$ in ((ref)) is a positive integrable weight function with a compact support. To develop the asymptotic theory of $T_{N}$ in the context of semi-long memory regressors and endogeneity, we provide three assumptions as follows.

myassumption$K(x) \pi(x)$ has a compact support such that $\int_{\mathbb{R}} K(x)dx=1$ and $|K(x)-K(y)| \leq C |x-y|$ whenever $|x-y|$ is sufficiently small.
myassumptionThere exist $g_1(x)$ and $g_2(x)$ such that for each $\theta, \theta_0 \in \Theta$, $|g(x,\theta)-g(x,\theta_0)| \leq C ||\theta-\theta_0|| g_1(x)$ holds, and for some $0< \beta \leq 1$, $|g_1(x+y)-g_1(x)| \leq C |y|^{\beta} g_2(x)$ holds, whenever $y$ is sufficiently small. Furthermore, $\int_{\mathbb{R}} [1+g_1^2(x)+g^2_2(x)] \pi(x) dx < \infty$ holds.
myassumptionUnder $H_0$, if $\theta_0$ is the true value of $\theta$, $||\hat{\theta}_N-\theta_0||=o_P\Big(\{\sqrt{N} \lambda^d h \}^{-1/2}\Big)$.

As noted in WP2016, Assumption (ref) covers a wide range of functions $g(x,\theta)$ and $\pi(x)$. Typical examples of $g(x,\theta)$ include $\theta e^x/(1+e^x)$, $e^{- \theta |x|}$, $e^{- \theta x^2}$, $e^{\theta |x|}/(1+e^{\theta |x|})$, $\theta \log|x|$, $(x+\theta)^2$, etc. Note that Assumptions (ref)--(ref) can be compared with those imposed by WP2016. In particular, the convergence rate imposed on $\hat{\theta}_N$ by Assumption (ref) has been verified in Section 4 of WP2016 and will be required in our proof of the theorems to follow. We now consider the asymptotic behavior of $T_{N}$ in ((ref)) and its convergence rate in the semi-long memory setting.

mytheoremLet Assumptions (ref) and (ref)--(ref) hold. Then, under $H_0$, we have \begin{equation} T_{\lambda,d}:= \frac{1}{\sqrt{N} \lambda^d h} T_{N} \rightarrow_D d_{(0)}^2 L_{B}(1,0), \end{equation} where $d_{(0)}^2=\mathbb{E}(u^2_0) \int_{\mathbb{R}}K^2(s)ds \int_{\mathbb{R}} \pi(x) dx$, $N^{1/2-\delta_0} \lambda^d h \rightarrow \infty$, and $\delta_0$ is as small as required.
myremark(i) Under semi-long memory settings for $d \in \mathbb{R}_{+}$, the limit distribution of ((ref)) is free of the unknown fractional differencing parameter $d$; however, the convergence rate depends on the unknown parameter $d$ even when $\lambda \rightarrow 0$ and $N \lambda \rightarrow \infty$. (ii) Under the long memory setting for $d \in (0,1/2)$, the test statistic and limit distribution given by WP2016 are $T_{N,d}:= (d_N/Nh) T_N \rightarrow_{D} d^2_{(0)} L_{B_{d+1/2}}(1,0)$, where $d_N \sim N^{d+1/2}$ is $N \rightarrow \infty$, and $T_N$ is as given in ((ref)). The limit distribution of this version of the test statistic relies on $d$ through the fractional Brownian motion process $B_{d+1/2}$. (iii) Under the short memory setting for $d=0$ and $d_N \sim N^{1/2}$, the test statistic and limit distribution are $T_{N,0}=(1/\sqrt{N}h) T_{N} \rightarrow_{D} d^2_{(0)} L_{B}(1,0)$. (iv) Note that the limit distribution of test statistic ((ref)) formulated under the semi-long memory case is similar to the short memory case. In neither case does the limit distribution depend on $d$.

To verify that the test has nontrivial power, one can assess the null hypothesis ((ref)) against a local alternative

equation[equation omitted — 107 chars of source]

Here, $\rho_N$ is a sequence of numerical constants that measures the local deviation from the null hypothesis. Also in ((ref)), $m(x)$ is a real function free of $\theta$ without lying in the span of $g(x,\theta)$ and its derivative functions. To make $m(x)$ smooth enough under $H_A$ for the sake of asymptotic power development, we give the following two assumptions.

myassumption(i) There exist $m_1(x)$ and $\gamma \in (0,1]$ such that for any $y$ sufficiently small, $|m(x+y)-m(x)| \leq C |y|^{\gamma} m_1(x)$ holds. (ii) Let $\int_{\mathbb{R}} [1+m^2(x)+m^2_1(x)] \pi(x) dx < \infty$ and $\int_{\mathbb{R}} m^2(x) \pi(x) dx > 0$ hold. (iii) The function $m(x)$ is not an element of the space spanned by $g(x,\theta)$ and its derivative functions.
myassumptionUnder $H_A$, if $\theta_0$ is the true value of $\theta$, $||\hat{\theta}_N-\theta_0||=o_P\Big(\{\sqrt{N} \lambda^d h \}^{-1/2}\Big)$.

The conditions on $m(x)$ in Assumption (ref) are week and can be satisfied by a large class of real functions. They are required to ensure the divergence of the test statistic and its consistency under $H_A$.

mytheoremLet Assumptions (ref), (ref)--(ref), and (ref)--(ref) hold. Then, under $H_A$, we have \begin{equation*} \lim\limits_{N \rightarrow \infty} P \Big(T_{\lambda,d} \geq T_0 \Big)=1, \end{equation*} where $T_0$ is positive, and $h \rightarrow 0$ satisfies $N^{1/2-\delta_0} \lambda^d h \rightarrow \infty$ for a small enough $\delta_0$ and any $\rho_N$ satisfying $N^{1/2} \lambda^d h \rho_N^2 \rightarrow \infty$.
myremarkBased on Theorem (ref), the test statistic $T_{\lambda,d}$ has nontrivial power against the local alternatives in the form of ((ref)) whenever $\rho_N \rightarrow 0$ at a rate that is slower than $\{\sqrt{N} \lambda^d h\}^{-1/2}$ as $\{\sqrt{N} \lambda^d h\}^{-1} \rightarrow 0$. The proof of Theorem (ref) ensures the divergence of normalized test statistic $T_{\lambda,d}$ and test consistency under $H_A$; see Remarks 3.2 and 3.3 in WP2016 for a related discussion.

Simulation studies

In this section, we illustrate the properties of the Nadaraya-Watson regression estimator and its confidence intervals. In addition, we study the size and power of the specification test statistic using the same regression functions.

Regression function properties

To examine the behavior of regression function estimators, we generate data sets from a model with the form of ((ref)), where we incorporate $\sigma$ as follows:

equation[equation omitted — 54 chars of source]

The regressor process $x_k$ is defined in ((ref)) for the semi-long memory (SLM) setting and as $x_k=x_k^\circ$ where $x_k^\circ$ is given in ((ref)) for the long memory (LM) setting. Let $u_k=\psi u_{k-1}+\epsilon(k)$ with $\psi=0.25$ and $\mathbb{E}(\zeta(k) \epsilon(k)):= \rho_{\zeta,\epsilon}$ (c.f., ((ref))) such that $(\zeta(k), \epsilon(k))$ are i.i.d. $N\Big(0,

psmallmatrix1 & \rho_{\zeta,\epsilon} \\ \rho_{\zeta,\epsilon} & 1

\Big)$. Also, let $\rho_{\zeta,\epsilon}=0.5$ and $\sigma=0.2$ in ((ref)).

We consider the following two regression functions:

itemize• straight line regression function: $f(x)= \theta_0 + \theta_1 x$, where $(\theta_0,\theta_1)=(0,1)$, and • quadratic regression function: $f(x)= \theta_0 + \theta_1 x + \theta_2 x^2$, where $(\theta_0,\theta_1,\theta_2)=(0,1,1)$.

We set the sample size at $N=\{ 50, 100, 500 \}$, and the number of Monte Carlo replications at $R=2000$.

We use the estimator ((ref)) with an Epanechnikov kernel in the form of $K(u)=0.75(1-u^2)1_{\{|u| \leq 1\}}$. For simplicity, we let $h = \{N^{-1/3}, N^{-1/7}\}$ and $\lambda = N^{-1/5}$ under the assumption of $N \lambda \rightarrow \infty$. We assume $d=\{ 0.1, 0.2, 0.3, 0.4\}$. We study Monte Carlo approximations of absolute bias ($\mathbb{E}\{|\hat{f}(x)-f(x)|\}$), standard deviation (std) ($[\mathbb{E}\{[\hat{f}(x)-E(\hat{f}(x))]^2\}]^{0.5}$) and root of mean squared error (rmse) ($[\mathbb{E}\{[\hat{f}(x)-f(x)]^2\}]^{0.5}$) of $\hat{f}(x)$ for the true regression functions over the interval $[-1,1]$ at $100$ points equally spaced between $-1$ and $1$.

To construct point-wise confidence intervals for the regression function $f(x)$, we use the limit distribution given in ((ref)). An asymptotic $100(1-\alpha)\%$ level confidence interval for $f(x)$ is then given by

equation*[equation* omitted — 144 chars of source]

where $\hat{\sigma}^2_N$ is given in expression ((ref)). We present the values of absolute bias, std, and rmse graphically for all the values of $x$ and for the two regression functions with $N=500$ and $h=N^{-1/7}$ in Figures (ref) and (ref). For the SLM case, these criteria are stable across values of $d$, while they vary as $d$ changes in the LM case and this is true for both straight-line and quadratic regression functions.

figure[figure omitted — 646 chars of source]
figure[figure omitted — 709 chars of source]
figure[figure omitted — 648 chars of source]
figure[figure omitted — 711 chars of source]

In Figures (ref) and (ref), we illustrate Monte Carlo coverage probabilities of confidence intervals and $(1-\alpha)=0.95$ point-wise intervals with their expected lengths. We observe that the expected lengths and coverage probabilities remain roughly similar for different values of $d$ under the SLM scenario, which is not the case for the LM scenario. Note that coverage was always reduced from the nominal level of $0.95$, but much less so for SLM processes than for LM processes. Coverage also decreased as $d$ increased. Additional results for smaller sample sizes and $h=N^{-1/3}$ are given in the Appendix.

Test statistic properties

In order to investigate the finite sample performance of test statistic, we first construct the Monte Carlo distributions of test statistic under both long and semi-long memory processes with endogeneity. Then, we use resampling methods to study the size and power of test statistic since their limiting distributions have complex forms and are not practical. In particular, we use the subsampling technique developed by mosaferi2024properties.

Drawing on the application to be presented in Section (ref) we consider two regression functions, a straight line ($y_k=\theta_0+\theta_1 x_k + \sigma u_k$) and a quadratic ($y_k=\theta_0+\theta_1 x_k + \theta_2 x_k^2+ \sigma u_k$), where $(\theta_0,\theta_1,\theta_2)=(0,1,1)$. For simplicity, we let $h=N^{-1/3}$ and $N=\{50, 100, 500 \}$. We assume that the amount of endogeneity is $\rho_{\zeta,\epsilon}=0.5$ and $d=\{0.1, 0.2, 0.3, 0.4, 0.5, 1, 1.5\}$, where $d \geq 0.5$ is only for the case of SLM, again based on experience with the application of Section (ref). Additionally, we let $\lambda=N^{-1/5}$. Other details of the simulation studies conducted to examine the behavior of the tests are the same as those described in Section (ref).

Monte Carlo densities of the test statistic are given in Figures (ref) and (ref). We observe that the densities become further from zero as the value of $d$ increases under the LM case. On the other hand, the densities overlap quite a lot for the SLM case and in particular for $0<d<1/2$.

To study the size of the test, we use the straight line and quadratic regression function assumptions to formulate $H_0$. To study the power of test, we use the $H_A$ given in equation ((ref)). We assume $\rho_N=1/N$ and $m(x_k)=|x_k|$. Therefore, we have $y_k=\theta_0+\theta_1 x_k + \frac{1}{N} |x_k| + \sigma u_k$ and $y_k=\theta_0+\theta_1 x_k + \theta_2 x_k^2+ \frac{1}{N} |x_k| + \sigma u_k$. For the block sizes, we choose the values of $\{[0.5 \sqrt{N}], [\sqrt{N}], [2 \sqrt{N}], [4 \sqrt{N}] \}$. The size and power values were close to $1$ for the LM and SLM cases. These results are not given in the manuscript, but allowed us to identify a bias in the form of the test statistic. We rectify this bias issue and construct a de-biased test statistic using the subsampling procedure and the algorithm described in Section C of mosaferi2024properties.

The values of size and power for the de-biased test statistic are given in Tables (ref) and (ref). Values of size for the straight line model are quite reasonable in both LM and SLM cases for small block sizes but decrease as block size increases. The size for the quadratic model is depressed from nominal, also decreases as block size increases, but is considerably more stable for SLM than LM situations. Power is low across the board, also decreases with block size, and is less stable for LM cases than for SLM ones. Additional results for $d \geq 0.5$ under SLM and smaller sample sizes are given in the Appendix.

figure[figure omitted — 634 chars of source]
figure[figure omitted — 649 chars of source]
table[table omitted — 1,093 chars of source]
table[table omitted — 1,094 chars of source]

Application

In this section, we illustrate the performance of the test statistic given in equation ((ref)) in the analysis of environmental Kuznets curves. In particular, we focus on the relationship of carbon dioxide emissions (CO$_2$) v.s. gross domestic product (GDP) per capita in Spain and sulfur dioxide emission (SO$_2$) v.s. GDP per capita in the United Kingdom. We consider the natural logarithmic transformation of these quantities ($x_k$ and $y_k$) before proceeding. The data set for Spain is from 1950 to 2008 and contains 59 observations. The data set for the United Kingdom is from 1870 to 1999 and contains 130 observations. The data sets are provided on the Github repository of the first author. As mentioned in the previous section, the limit theory of this test does not have a simple form and is not practical, so we again rely on the use of subsampling to compute benchmarks.

figure[figure omitted — 341 chars of source]

Scatter-plots of the relevant variables are shown in Figure (ref). For Spain, the pattern of points suggests a straight-line regression function, while for the United Kingdom the pattern appears to be more quadratic in form. Thus, we test the null hypothesis of $f(x)=\theta_0+\theta_1 x$ for Spain and two null hypotheses of (1) $f(x)=\theta_0+\theta_1 x$ and (2) $f(x)=\theta_0+\theta_1 x + \theta_2 x^2$ for the United Kingdom. We also overlay the Nadaraya-Watson (N-W) regression function estimator $\hat{f}(x)$ from ((ref)) on the plots in Figure (ref). Optimal values for the bandwidth $h$ were selected using cross-validation (e.g. park1990comparison,hardlebook) with the help of N-W regression functions. The optimal bandwidth was $h=0.151$ for Spain and $h=0.107$ for the United Kingdom. The associated least squares values based on $\text{LCV}(h)=\sum_{k=1}^{N}(y_k - \hat{f}_{-k}(x_k))^2$ are given in Table (ref), where $\hat{h}_{\text{optimal}}=\text{argmin}_{h>0}\text{LCV}(h)$. To estimate the values of the fractional differencing parameter $d$ and the tempering parameter $\lambda$, we use their Whittle estimators from artfima package in R (e.g., sabzikar2019parameter). The estimated values of these parameters are also given in Table (ref). Based on the previous section and due to the biased form of test statistic $T_{\lambda,d}$, we consider both the biased and de-biased form of tests for use with subsampling. The resultant p-values are summarized in Table (ref). Furthermore, the p-values for the quadratic null hypothesis for the United Kingdom based on all possible block sizes values are displayed in Figure (ref).

table[table omitted — 352 chars of source]
figure[figure omitted — 568 chars of source]

For Spain, we observe that the null hypothesis of a straight line would be rejected when we use the biased form of test statistic, but with the de-biased test statistic this hypothesis would not be rejected. For the United Kingdom, the null hypothesis of a straight line is rejected for both the biased and de-biased forms of the test statistic. A hypothesized quadratic form for the regression function would not be rejected for either the biased or de-biased test statistics. Based on Figure (ref), including a greater number of blocks in the de-biasing process increases the p-values associated with the hypothesis of a quadratic regression function. In particular, for the last plot in Figure (ref), we can see that when 20 blocks are included in de-biasing, the p-value is one for all block sizes.

table[table omitted — 775 chars of source]

Discussion and future work

In this article we have considered the use of tempered or semi-long memory regressor processes within the context of cointegrating regression models for time series analysis. We have developed limit theory that allows the construction of pointwise confidence intervals for a nonlinear regression function when estimation is accomplished through the use of a kernel smoother. We have also demonstrated that the limit distribution of a test statistic for the specification of the regression function does not depend on the differencing parameter $d$ in semi-long memory. Through simulation studies, we have shown that performance of interval estimators of a regression function is more stable and has generally better performance for the semi-long memory case than for long memory. We also developed the limit distribution of the test statistic for their parametric forms and showed that it is free from the unknown parameter of $d$.

Some directions for future work could be related to multivariate nonparametric cointegrating regression with multiple regressors such that they have a semi-long memory property with error terms belonging to the $\alpha$-stable finite-dimensional distributions with infinite variance and infinite lower-order moments. We assume the regression disturbances are serially and cross-correlated with the regressors and permit the associated bandwidth in the regression model to be random and data dependent. One can investigate the uniform convergence rate for the kernel regression estimator and provide model specification test for multivariate cointegrating regressions where the asymptotic distribution of the test statistic is free of any unknown parameters and is proportional to a local time of linear stable motions.

Acknowledgments

We would like to thank the Editor, Associate Editor, and two anonymous reviewers for their excellent suggestions and comments that have substantially improved the quality of our manuscript.

Disclosure statement

No potential conflict of interest was reported by the authors.

\setcounter{secnumdepth}{0}