EconBase
← Back to paper

Detecting long-range dependence for time-varying linear models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

81,821 characters · 14 sections · 86 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Detecting long-range dependence for time-varying linear models

abstractWe consider the problem of testing for long-range dependence in time-varying coefficient regression models, where the covariates and errors are locally stationary, allowing complex temporal dynamics and heteroscedasticity. We develop KPSS, R/S, V/S, and K/S-type statistics based on the nonparametric residuals. Under the null hypothesis, the local alternatives as well as the fixed alternatives, we derive the limiting distributions of the test statistics. As the four types of test statistics could degenerate when the time-varying mean, variance, long-run variance of errors, covariates, and the intercept lie in certain hyperplanes, we show the bootstrap-assisted tests are consistent under both degenerate and non-degenerate scenarios. In particular, in the presence of covariates the exact local asymptotic power of the bootstrap-assisted tests can enjoy the same order as that of the classical KPSS test of long memory for strictly stationary series. The asymptotic theory is built on a new Gaussian approximation technique for locally stationary long-memory processes with short-memory covariates, which is of independent interest. The effectiveness of our tests is demonstrated by extensive simulation studies and real data analysis.

Keywords: Long-range dependence, Locally stationary process, Spurious long memory, Time-varying models

\footnotetext[1]{E-mail addresses: \href{[email removed]}{[email removed]}(L.Bai), \href{[email removed]}{[email removed]}(W.Wu)}

Introduction

Consider the time-varying coefficient linear model

align[align omitted — 113 chars of source]

where the covariate vector $\mathbf x_{i,n}=(1,x_{i,2,n},...,x_{i,p,n})^\top$ is a $p$-dimensional {\it short-range dependent} (SRD) locally stationary time series and $y_{i,n}$ is the response variable, $t_i = i/n$. At each time point $t_i$, we only observe one realization $(\mathbf x_{i, n},y_{i,n})$ and no repeated measurement is available. The time-varying regression coefficient function $\boldsymbol \beta(\cdot)$ is a $p$-dimensional function with each coordinate a smooth function on $[0,1]$ and the zero mean error process $(e_{i,n})$ is a possibly {\it long-range dependent} (LRD) or {\it long-memory} time series. More precisely, we assume $(e_{i,n})$ is a locally stationary $I(d)$ process i.e., for $1\leq i\leq n$

align[align omitted — 71 chars of source]

where $\mathcal B$ is the lag operator, $d\in [0,1/2)$ is the {\it long-memory parameter} and $u_{i,n}$ is a SRD or {\it short-memory} locally stationary process. The strict definitions of locally stationary and long-memory processes are deferred to (ref). The error model (ref) naturally generalizes classical stationary SRD and LRD processes by allowing their generating mechanism to vary with time. Observe that $(e_{i,n})$ will reduce to the SRD process $(u_{i,n})$ if $d=0$ and will be a LRD process if $d \in (0, 1/2)$. In fact, when $(u_{i,n})$ is stationary, (ref) allows the classical stationary long-memory processes (e.g. FARIMA-GARCH models), which have found extensive application in hydrology (ZHANG2011121, koutsoyiannis2013hydrology), economics and finance (caporale2013long, unemployment2016) and many other fields since first introduced by hurst1951long. Moreover, model (ref) admits heteroscedasticity, i.e., the dependence of $u_{i,n}$ on $(\mathbf x_{r,n})_{r=1}^n$, see (ref) for more details.

The time-varying regression model (ref) with time series errors has attracted enormous attention, see for instance fan2000simultaneous, zhou2010simultaneous and chen2018jbes where the errors are assumed to be SRD, and kulik2012conditional, BeranLongmemory and ferreira2018estimation where LRD errors are considered. The aforementioned research reveals that nonparametric estimators of the time-varying coefficient ${\boldsymbol \beta}(\cdot)$ possess distinct properties under the two scenarios, $d=0$ and $0<d<1/2$. When $d=0$, consider the local linear estimator of the multivariate coefficient function $ {\boldsymbol \beta}(\cdot)$ using the kernel function $K(\cdot)$ and the bandwidth $b_n$, of which the asymptotic behavior rests on the distributions of $(\mathbf x_{i,n})$ and $(e_{i,n})$. In particular, the order of the deviation $|\hat {\boldsymbol \beta}(\cdot)-\boldsymbol \beta(\cdot)|$ is determined by the long-memory parameter $d$. For $d=0$ and $t\in (0,1)$, zhou2010simultaneous shows that under mild conditions,

align[align omitted — 193 chars of source]

where $\mu_2$ and $\phi_0$ are constants determined by $K(\cdot)$ and ${\boldsymbol \Sigma}(t)$ is determined by the moments of the process $(\mathbf x_{i,n}e_{i,n})_{i=1}^n$. Meanwhile, for $d>0$ and $p=1$, Theorem 7.22 in BeranLongmemory shows that for stationary $e_{i,n}$ under regularity conditions,

align[align omitted — 170 chars of source]

where $V(d) = 2c_f \Gamma(1-2d) \sin(\pi d) \int_{-1}^1\int_{-1}^1 K(x)K(y)|x-y|^{2d-1}dxdy$, and $c_f$ is a constant related to the spectral density of errors. Equation (ref) shows that for $d>0$ the convergence rate of $\hat {\boldsymbol \beta}(t)$ is $(nb_n)^{d-1/2}$, which is much slower than the well-known $(nb_n)^{-1/2}$ convergence rate as given by (ref) when $d=0$. Therefore, a crucial problem of the statistical inference of model (ref) is to test

align[align omitted — 74 chars of source]

The testing problem (ref) for (ref) is closely related to the existing tests of `spurious long memory', which refers to the phenomenon that in the presence of regime changes, level shifts or certain deterministic trends, a short memory process could exhibit many properties of a long-memory process, known as the `spurious long-memory' effects, see for example giraitis2001testing, qu2011test, and mccloskey2013memory. These findings motivate the tests for distinguishing genuine and spurious long memory. Among others, qu2011test, preu2013 and SIBBERTSEN201833 consider testing the null hypothesis of stationary long memory against spurious long memory. Meanwhile, several tests have been introduced to test the null hypothesis of spurious long memory, for which a prevailing approach is to assume a specific and parametric form of non-stationarity, see shao2006aos, harris2008testing, and davis2013consistency among others. Recently, there has been growing interest in detecting long memory in the presence of general non-stationarity, see for example Detectinglrdependence which considered locally stationary moving average formulation. In practice, by testing (ref) for model (ref), we are able to identify a new type of `spurious long memory' resulting from the {\it misspecification in the conditional mean}. See our data analysis in (ref) where we apply our method to the Hong Kong circulatory and respiratory data.

The goal of the present work is to test the {\it hypothesis (ref)} under complex and general temporal dynamics, assuming that $(\mathbf x_{i,n})$ and $(u_{i,n})$ belong to the flexible class of locally stationary processes generated by smoothly changing underlying mechanisms. Although some related literature has studied hypothesis (ref) for linear regression models with deterministic covariates, see for example harris2008testing, to the best of the authors' knowledge, this work is the first instance investigating testing (ref) for (ref) in the presence of time series covariates. In the literature, KPSS (Kwiatkowski, Phillips, Schmidt, and Shin, see lee1996power), R/S (range over standard deviation, see hurst1951long ), V/S (rescaled variance, see giraitis2003rescaled), and K/S (which has a limiting distribution of Kolmogoroff-Smirnoff form, see lima2004robustness) tests have been widely used for long memory detection in stationary processes. In this paper, we develop new KPSS, R/S, V/S and K/S-type tests tailored to the non-stationary time series time-varying regression problem (ref). The limiting distributions of the test statistics under the null hypothesis, the local and fixed alternatives are then derived. Our results differ from their stationary counterparts due to the following reasons. (1) Under the null hypothesis, it is well-known that the KPSS, R/S, V/S and K/S tests are built on the partial sum process whose convergence rate is $ n^{-1/2}$. However, the nonparametric estimate $\hat {\boldsymbol \beta}(\cdot)$ induces stochastic errors much larger than $n^{-1/2}$ which will lead to different limiting distributions as well as possible degeneracy. (2) Due to the non-stationary errors and covariates, the partial sum processes cannot be approximated by processes with stationary increments, which makes the test statistics non-pivotal. The major contributions of the paper lie in the following three aspects.

Firstly, our methods are applicable to the locally stationary time series regression, which has found considerable attention in various related fields, see for instance vogt2012, you2019, zhou2010simultaneous and many others. In particular, the flexible locally stationary framework allows the error processes to display conditional and unconditional heteroscedasticity that has been increasingly investigated (see harris2017adaptive and cavaliere2020adaptive) in the context of long-memory models. Both the evolving distributional properties of the locally stationary data and the long-memory properties pose long-standing challenges to the inference of time-varying coefficient linear model (ref) due to the lack of general Gaussian approximation techniques for non-stationary long-memory processes. For stationary long-memory processes, Gaussian approximation has been studied by for example dehling1989empirical. Recently, wu2018 developed Gaussian approximation schemes for a class of locally stationary long-memory linear processes. However, their results cannot accommodate regression problems with time series covariates which requires the analysis of distributional properties of the partial sum process of $(\mathbf x_{i,n} e_{i,n})_{i=1}^n$. In this paper, we address this issue via a further Gaussian approximation theorem, allowing $(\mathbf x_{i,n})$ and $(e_{i,n})$ to be non-stationary SRD and LRD processes, respectively, with flexible dependence between them.

Secondly, we develop effective bootstrap approaches which circumvent the difficult estimation of the non-pivotal limiting distributions of the test statistics under time series non-stationarity. In particular, the test statistics could degenerate when the time-varying mean, variance, long-run variance of errors and covariates, and the intercept lie in certain hyperplanes whose geometry cannot be directly identified from the data. Importantly, regardless of the degeneracy of test statistics, our bootstrap procedures are consistent and possess good finite sample properties. Furthermore, we show that the exact local asymptotic power of the four types of bootstrap-assisted tests can reach the order $O( \log^{-1} n)$ in the presence of time series covariates, no matter whether the test statistics degenerate under the null hypothesis. This rate coincides with shao2007local which studies a similar problem of testing the SRD null hypothesis against LRD local alternatives for strictly stationary series without covariates.

The rest of the paper is organized as follows. Section (ref) introduces necessary notation. Section (ref) formally states the non-stationary LRD model and the related assumptions. Section (ref) provides the test statistics, and establishes the asymptotic results via the new Gaussian approximation theory for the product of non-stationary SRD and LRD processes. Section (ref) discusses the bootstrap algorithms. Section (ref) reports the simulation results and the analysis of Hong Kong circulatory and respiratory data. Section (ref) provides a brief concluding remark. In (ref), we provide the proof of the new Gaussian approximation theory. Detailed proofs, the literature review of KPSS and related tests, the implementation details including the selection of tuning parameters, additional simulation results, data analysis results (including COVID-19 data), and additional algorithms are relegated to the supplement.

Notation

For a matrix $\mathbf A = (a_{ij})_{1\leq i \leq n, 1\leq j \leq m} \in \mathbb{R}^{n\times m}$, let $|\mathbf A| = (\sum_{j=1}^m \sum_{i=1}^n a_{ij}^2)^{1/2}$ and write $\mathbf A \geq 0$ if $\mathbf A$ is semi-positive definite. Let $(\mathbf A)_{(1,1)}$ denote the element of $\mathbf A$ in the first column and first row. Notice that when $m=1$, $\mathbf A$ is a vector. For $\mathbf A\geq 0$ with eigendecomposition $\mathbf A = \mathbf Q\mathbf D \mathbf Q^{\top}$ with orthonormal matrix $\mathbf Q$ and diagonal matrix $\mathbf D$, the root of $\mathbf A$ is defined by $\mathbf A^{1/2} = \mathbf Q \mathbf D^{1/2} \mathbf Q^{\top}$, where $\mathbf D^{1/2}$ is the elementwise root of $\mathbf D$. Let $\mathbf I_p$ denote the $p$-dimensional identity matrix. For a random matrix $\mathbf A$, for $q \geq 1$, let $\|\mathbf A \|_q = (\mathbb{E}|\mathbf A|^q)^{1/q}$ denote the $\mathcal{L}^q$-norm of the random variable $|\mathbf{A}|$ and write $\|\cdot\|=\|\cdot\|_2$ for short. Write $\mathbf A \in \mathcal{L}^q$ if $\|\mathbf A \|_q < \infty$. For a function $f(\cdot)$, write $f \in C[0,1]$ if $f$ is continuous over $[0,1]$, $f \in C^p[0,1]$ if the $p_{th}$ order derivative of $f$ is continuous over $[0,1]$. Write $\mathbf A \in C[0,1]$ and $\mathbf A \in C^p[0,1]$ if each element $A_{ij}(\cdot)$ in $\mathbf A(\cdot)$ is in $C[0,1]$ and $C^p[0,1]$, respectively. For any kernel function $K(\cdot)$, let $K^*(\cdot)$ denote the jackknife equivalent kernel $2 \sqrt{2} K(\sqrt{2}x) - K(x)$. Denote by $\lfloor x\rfloor$ the largest integer smaller or equal to $x$. For any two positive real sequences $a_n$ and $b_n$, write $a_n \asymp b_n$ if $\exists\ 0 <c < C<\infty$ such that $c<\liminf_{n \to \infty} \frac{a_n}{b_n}<\limsup_{n \to \infty} \frac{a_n}{b_n}<C$. Let $t \wedge s$ denote the smaller value in $t$ and $s$. Let `$\Rightarrow$' denote convergence in distribution. Write $\lambda$ as Lebesgue measure on $[0,1]$. Let `$:=$' denote `defined as'.

Model assumptions

We start by introducing the time series time-varying regression model (ref) in detail. Recall model (ref) has the following form

equation[equation omitted — 195 chars of source]

We assume that the process $(u_{i,n})_{i=-\infty}^n$ and the covariate process $(\mathbf x_{i,n})_{i=1}^n$ have the form

align[align omitted — 119 chars of source]

where $\mathcal{F}_i=(\varepsilon_{-\infty},...,\varepsilon_i)$, $(\varepsilon_{i})_{i\in \mathbb Z}$ are $i.i.d.$ random variables, $H$ and $\mathbf W=(W_1,...,W_p)^\top$ are measurable functions such that $H:(-\infty,1]\times \mathbb R^{\mathbb Z}\rightarrow \mathbb R$, $W_s:[0,1]\times \mathbb R^{\mathbb Z}\rightarrow \mathbb R$, $2\leq s\leq p$, while $W_1$ is fixed to be $1$ corresponding to the intercept of the regression. Define $(\varepsilon^{\prime}_i)_{i \in \mathbb{Z}}$ as an $i.i.d.$ copy of $(\varepsilon_i)_{i \in \mathbb{Z}}$, and for $j \geq 0$, let $\mathcal{F}^{*}_j = (\mathcal{F}_{-1},\varepsilon^{\prime}_0, \varepsilon_1,\cdots, \varepsilon_{j-1}, \varepsilon_{j})$. For any (vector) process $\mathbf L(t,\mathcal{F}_i)$, it is said to be $\mathcal L^q$ stochastic Lipschitz continuous in the interval $\mathcal I$ (denoted by $\mathbf L\in \mathrm{Lip}_q(\mathcal I)$) if for $t_1,t_2\in \mathcal I$, there exists a constant $M>0$ such that

align[align omitted — 100 chars of source]

We say the process $\mathbf L(t,\mathcal{F}_i)$ is {\it locally stationary} (LS) on $\mathcal I$ if $\mathbf L(t,\mathcal{F}_i)\in \mathrm{Lip}_q(\mathcal I)$ for some $q\geq 2$. Write $\mathrm{Lip}_q=\mathrm{Lip}_q([0,1])$ for short. The locally stationary process offers a flexible nonparametric device to characterise the complex temporal dynamics of the error and covariate processes in (ref), which is based on Bernoulli shift processes and leads to a general framework of nonlinear processes, see wu2005nonlinear. Other formulations of locally stationary processes include dahlhaus1997 and nason2000wavelet. See dahlhaus2019bej for a comprehensive review. The physical dependence measure of the nonlinear filter $\mathbf L \in \mathcal{L}^q$ ($q>0$) over the interval $\mathcal I$ is defined by

equation[equation omitted — 164 chars of source]

The physical dependence measure $\delta_q(\mathbf L,k, \mathcal I)$ quantifies the influence of the input $\varepsilon_0$ on the output $\mathbf L(t,\mathcal{F}_k)$ over the interval $\mathcal I$. Observe that $\delta_q(\mathbf L,k, \mathcal I)=0$ if $k<0$. Write $\delta_q(\mathbf L,k)=\delta_q(\mathbf L,k, [0,1])$ for short. We proceed to define SRD and LRD non-stationary processes for non-stationary time series.

definitionThe univariate process $( L(t,\mathcal{F}_i))_{i=-\infty}^\infty, t\in \mathcal I$ is said to be SRD if $$ \sum_{j=-\infty}^{\infty} \sup_{t,s \in \mathcal I} \left|\operatorname{Cov}\left( L(t, \mathcal{F}_0), L(s, \mathcal{F}_{j})\right)\right|<\infty, $$ and a LRD process otherwise. The $p$-dimensional vector process $( \mathbf L(t,\mathcal{F}_i))_{i=-\infty}^\infty, t\in \mathcal I$ is SRD if each of its component is SRD, and is otherwise LRD.

Definition (ref) distinguishes the SRD and LRD by the uniform summability of covariance, which naturally extends the traditional definition of long memory in second-order stationary univariate processes, see for example Condition III in Chapter 2 of pipiras2017long. The uniform long-memory definition has been introduced to define LRD and SRD non-stationary time series in wu2018. Various definitions of LRD stationary processes could be found in pipiras2017long and definitions of LRD locally stationary processes are discussed in BERAN2009900, Detectinglrdependence and ferreira2018estimation among others. In this paper, we posit the following assumptions.

assumptionThe zero mean SRD process $(u_{i,n})_{i=-\infty}^{n}$ satisfies \begin{enumerate}[label=(a\arabic*)', itemsep=2pt,topsep=0pt,parsep=0pt] • $H(t,\mathcal{F}_0) \in \mathrm{Lip}_2(-\infty,1]$, $\mathbb{E} H(t,\mathcal{F}_0)=0$, and $ \sup_{ t \in (-\infty,1] }\left\|H\left(t, \mathcal{F}_{0}\right)\right\|_{4}<\infty $. • $ \delta_4(H, k,(-\infty,1]) = O(\chi^k)$ $\text{for some}~ \chi \in (0,1)$. • Define the long-run variance function as $ \sigma^{2}_H(t):=\sum_{k=-\infty}^{\infty} \operatorname{Cov}\left(H\left(t, \mathcal{F}_{0}\right), H\left(t, \mathcal{F}_{k}\right)\right)$, $t \in (-\infty, 1], $ which satisfies that $\underset{t \in (-\infty,1] }{\inf} \sigma_H^2(t)> 0$ and $\underset{t \in (-\infty,1] }{\sup}\sigma_H^2(t) < \infty.$$\sigma_H^2 (\cdot)$ is twice continuously differentiable on $[0,1]$. \end{enumerate}

Condition (ref) imposes the assumptions of local stationarity and finite forth moment on the innovations $(u_{i,n})$. Condition (ref) will be satisfied if $\sup_{t \in (-\infty, 1]} \|\frac{\partial}{\partial t}H(t,\mathcal{F}_0)\| <\infty$. Condition (ref) ensures that the innovations $(u_{i,n})$ are SRD and satisfy geometric measure contraction (GMC). Conditions (ref) and (ref) guarantee that the innovations $(u_{i,n})$ have a finite, non-degenerate and smooth long-run variance. Notice that under the null hypothesis $d=0$, the error process $(e_{i,n})$ reduces to $(u_{i,n})$, which indicates that $(e_{i,n})$ is a SRD process. When $d > 0$, $(e_{i,n})$ is generated by a binomial weighted combination of $u_{i,n}=H(t_i,\mathcal{F}_i)$ starting from the infinite past ($i=-\infty$). To stress that $(e_{i,n})$ is a LRD process under the alternative hypothesis with respect to $d$, in the remaining of this article we write $e_{i,n}$ as $e_{i,n}^{(d)}$ when $d>0$, i.e., $e_{i,n}^{(d)} = (1-B)^{-d} u_{i,n} = \sum_{k=0}^{\infty} \psi_k(d) u_{i-k}$, $\psi_j(d) = \Gamma (j+d) / [\Gamma(d) \Gamma (j+1)]$. We further write $e_{i,n}^{(d)} = H^{(d)}(t_i, \mathcal{F}_{i})$, where $ H^{(d)}(t, \mathcal{F}_{l}) = \sum_{k=0}^{\infty} \psi_k(d) H(t - t_k, \mathcal{F}_{l-k})$. The following Proposition (ref) elaborates that the physical dependence measure of $(e^{(d)}_{i,n})$ relies on $d$.

propositionUnder Assumption (ref), we have uniformly for $l \geq 0$, $0 < d < 1/2$, \begin{align} \delta_p(H^{(d)}, l, (-\infty, 1]) = O\{(1+l)^{d-1}\}. \end{align}

Our formulation of $(e_{i,n})$ in (ref) allows for a wide class of non-stationary SRD and LRD processes under $H_0$ and $H_A$, respectively, including the following examples.

example[Linear locally stationary process] Consider the time-varying FARIMA($p,d,q$) model $(0 < d<1/2)$ (recall that $u_{j,n}=H(t_j,\mathcal{F}_j)$, $-\infty \leq j \leq n$) \begin{align} &(1-\mathcal{B})^{d}e_{i,n}=u_{i,n}, i = 1,2,\cdots, n,with \Phi^{p}(\mathcal{B}, t)H(t,\mathcal{F}_j)= \Theta^{q}(\mathcal{B}, t) \varepsilon_{j},\nonumber \end{align} where $t \in (-\infty,1]$, $j\in \mathbb{Z}$, $\Phi^{p}(z, t)=1+\phi_{1}(t) z+\cdots+\phi_{p}(t) z^{p}$ and $\Theta^{q}(z, t)=1+\theta_{1}(t) z+\cdots+\theta_{q}(t) z^{q}$ are polynomials with degrees $p$ and $q$, and the random variables $(\varepsilon_{i})_{i\in\mathbb Z}$ are $i.i.d.$ with mean $0$ and variance $1$. Assume that for $t \in (-\infty,1]$, $\{\phi_i(t), 1\leq i \leq p\}$ and $\{\theta_j(t), 1 \leq j \leq q\}$ are twice differentiable, $\Phi^{p}(z, t)$ and $\Theta^{q}(z, t)$ do not share the same roots, and $\Phi^{p}(z, t)$ does not have roots in the unit disk $\{|z| \leq 1\}$. Then there exists real-valued differentiable functions $(a_i(t))_{i\geq 0}$ such that $A(z,t) =\Theta^{q}(z, t)/\Phi^{p}(z, t) = \sum_{i=0}^{\infty} a_{i}(t) z^i$, where for $t\in (-\infty ,1]$, $|a_j(t)|$ and $|a_j^{\prime}(t)|$ are summable. Consequently, we have the $\text{MA}(\infty)$ representation \begin{equation} e_{i,n}=(1-\mathcal B)^{-d}u_{i,n}=(1-\mathcal B)^{-d}\sum_{j=0}^\infty a_j(t_i)\mathcal B^j\varepsilon_{i}:=\sum_{j=0}^{\infty} b_{j,i}\varepsilon_{i-j}. \end{equation} When $d=0$, it follows that $b_{j,i} = a_j(t_i)$. Thus, $(e_{i,n})$ is SRD according to Definition (ref). When $d > 0$, by Lemma 3.2 of KOKOSZKA199519, $b_{j,i} =\sum_{l=0}^j \psi_l a_{j-l}(t_{i-l})= L_i(j)j^{d-1}$, where $L_i(\cdot)$ is a slowly varying function for each $i$. Suppose for $1 \leq i \leq n$, $|L_i(j)| \leq L(j)$ for some slowly varying function $L(\cdot)$. Noticing that $\operatorname{Cov}(e_{i,n}, e_{i+k,n}) = \sum_{j=0}^{\infty} b_{j,i}b_{j+k, i+k}$. By Proposition 2.2.9 in pipiras2017long, $ \sup_{1 \leq i \leq n}|\operatorname{Cov}(e_{i,n}, e_{i+k,n})|$ is of order $k^{2d-1}$ for all $i \in \mathbb{Z}$. Hence, (ref) is LRD according to Definition (ref).
example[Nonlinear locally stationary process] Consider the time-varying ARFIMA($p,d,q$)-GARCH($1,1$) process $(1-\mathcal{B})^d e_{i,n} =u_{i,n}, 1\leq i\leq n$, where $u_{j,n}=H(t_j,\mathcal{F}_j)$ and \begin{align} \Phi^{p}(\mathcal{B}, t)H(t, \mathcal{F}_j)= \Theta^{q}(\mathcal{B}, t)v_j(t), v_j(t) =\varepsilon_{j} \sigma_{j}(t), j\in \mathbb Z, t\in (-\infty,1],\nonumber \end{align} where $\sigma^2_{j}(t) = c(t)+\alpha(t)v_{j-1 }^{2}(t) + \beta(t) \sigma^2_{j-1}(t)$, $(\varepsilon_{i})_{i\in\mathbb Z}$ are $i.i.d.$ random variables with mean 0 and variance 1, $c(t), \alpha(t), \beta(t)$ are smooth non-negative functions, and $\Phi^{p}(z, t)=1+\phi_{1}(t) z+\cdots+\phi_{p}(t) z^{p}$ and $\Theta^{q}(z, t)=1+\theta_{1}(t) z+\cdots+\theta_{q}(t) z^{q}$ are polynomials with degrees $p$ and $q$. Assume that for $t \in (-\infty,1]$, $\{\phi_i(t), 1\leq i \leq p\}$, $\{\theta_j(t), 1 \leq j \leq q\}$, $c(t)$, $\alpha(t)$ and $\beta(t)$ are twice differentiable, $\Phi^{p}(z, t)$ and $\Theta^{q}(z, t)$ do not share the same roots, and $\Phi^{p}(z, t)$ does not have roots in the unit disk $\{|z| \leq 1\}$, $(\varepsilon_{i})_{i\in\mathbb Z} \in \mathcal{L}^{8}$, $\sup _{t \in (-\infty,1]} c(t) <\infty$ and $\sup _{t \in (-\infty,1] }\left\|\alpha (t) \varepsilon_{t}^{2}+\beta(t)\right\|_{4}<1$. Then, by Example 2 of wu2011gaussian, \ $(u_{i,n})_{i=-\infty}^{n}$ satisfy (ref) of Assumption (ref). When $d > 0$, Definition (ref) can be verified similarly as given in Example (ref), since $(v_{i}(t))_{i\in \mathbb Z }$ are white noises.

Main results

Since $(e_{i,n})$ is not observable in (ref), we propose to test $H_0$ based on nonparametric residuals. Specifically, we adopt the local linear approach (see for instance fan1993local and fan1996local) to estimate ${\boldsymbol \beta}(t)$ in (ref), i.e.,

equation[equation omitted — 350 chars of source]

where $K(\cdot)$ is a kernel function with finite support $[-1,1]$ and $b_n$ is a bandwidth and $K_{b_n}(\cdot)=K(\cdot/b_n)$. To further eliminate the bias term involving $\boldsymbol \beta^{\prime \prime }(\cdot)$, we use the jackknife bias-corrected estimator in wu2007inference :

equation[equation omitted — 158 chars of source]

Then, we obtain the nonparametric residuals $(\tilde e_{i,n})$, i.e., $ \tilde e_{i,n} = y_{i,n} - \mathbf x_{i,n}^{\top} \tilde{\boldsymbol \beta}(t_i).\nonumber $ For simplicity, define $\tilde S_{r,n} =\sum_{i=\lfloor nb_n \rfloor+1}^r \tilde{e}_{i,n}$, $r=\lfloor nb_n \rfloor+1,\cdots n-\lfloor nb_n\rfloor$. We consider four well-known types of partial sum based test statistics, which are KPSS, R/S, V/S and K/S-type tests built on $(\tilde e_{i,n})$.

enumerate[itemsep=2pt,topsep=0pt,parsep=0pt] • KPSS-type statistic \begin{equation} T_n = \frac{1}{n(n - 2\lfloor nb_n\rfloor)}\sum_{r=\lfloor nb_n \rfloor+1}^{n-\lfloor nb_n\rfloor} \left(\tilde S_{r,n}\right)^2.\end{equation} • R/S-type statistic $ Q_n = \max_{\lfloor nb_n \rfloor + 1 \leq k \leq n - \lfloor nb_n \rfloor } \tilde S_{k,n} - \min_{\lfloor nb_n \rfloor + 1 \leq k \leq n - \lfloor nb_n \rfloor } \tilde S_{k,n}. $ • V/S-type statistic $ M_n = \frac{1}{n(n - 2\lfloor nb_n\rfloor)}\left\{\sum_{k=\lfloor nb_n \rfloor + 1 }^{n-\lfloor nb_n \rfloor } \tilde S_{k,n} ^2 - \frac{1}{n - 2\lfloor nb_n \rfloor}\left(\sum_{k=\lfloor nb_n \rfloor + 1}^{n-\lfloor nb_n \rfloor } \tilde S_{k,n} \right)^2\right\}. $ • K/S-type statistic $ G_n = \max_{\lfloor nb_n \rfloor + 1 \leq k \leq n - \lfloor nb_n \rfloor } \left|\tilde S_{k,n} \right|. $

The above four types of tests have been widely applied to the detection of long memory and many other important problems (e.g. unit root testing) for stationary time series. We refer to Section (ref) of the online supplement for the complete literature review and applications. To the best of our knowledge, all the existing work on the KPSS, R/S, V/S, and K/S tests considers the statistics based on the original series $(e_{i,n})$ or parametric residuals (e.g., the residuals obtained by the removal of the sample mean), and is therefore not applicable to the time-varying coefficient model (ref). Meanwhile, it is well-known that the nonparametric estimators have a slower convergence rate than the corresponding parametric estimators. Therefore, the asymptotic properties of our nonparametric residual-based KPSS and related tests will be very different from their parametric residual-based or original series-based counterparts; in fact we show that the tests can degenerate under certain scenarios (see Theorem (ref)) and thus bootstrap procedures adaptive to the possible degeneracy are proposed for implementation, see Algorithms (ref), also Algorithms \ref*{trend_algorithms} and \ref*{algorithms} of the online supplement. In this article, we use the term `{\it KPSS and related tests}' to represent the four types of tests.

Assumptions

For the sake of brevity, in this section we only discuss the KPSS-type test statistic in detail, and summarize the results of KPSS-related test statistics in Remark (ref). In order to investigate the asymptotic properties of $T_n$ defined by (ref) in the presence of time series covariates, we introduce the following assumptions.

assumptionThe kernel function $K(\cdot)$ is continuous, symmetric and supported on $[-1,1]$.
assumptionEach coordinate of $\boldsymbol \beta(\cdot)$, i.e., $\beta_i(\cdot)$ for $i = 1,\cdots,p$, lies in $C^3[0,1]$.
assumptionLet $\mathbf U(t, \mathcal{F}_i) = H(t, \mathcal{F}_i)\mathbf W(t, \mathcal{F}_i)$, $i = 1, \cdots, n$, s.t. \begin{enumerate}[label=(A\arabic*)] • $\mathbf U(t, \mathcal{F}_i) \in \mathrm{Lip}_2$, $\sup_{t \in [0,1]} \|\mathbf U(t, \mathcal{F}_i) \|_4< \infty.$ • Short-range dependence: $\delta_4(\mathbf U, k) = O(\chi^k)~ \text{for some} ~\chi \in (0,1)$. • The smallest eigenvalue of $ {\boldsymbol \Sigma}(t):=\sum_{j=-\infty}^{\infty} \mathrm{Cov}\left\{\mathbf{U}\left(t, \mathcal{F}_{0}\right), \mathbf{U}\left(t, \mathcal{F}_{j}\right)\right\}$, $ t\in [0,1], $(i.e., the long-run covariance matrix) is bounded away from 0 on $[0,1]$. \end{enumerate}

Assumption (ref) is standard for local linear time series regression, see for instance zhou2010simultaneous. Notice that the first element of $\mathbf U(t, \mathcal{F}_i)$ is $H(t, \mathcal{F}_i)$. When $p=1$ (the time-varying trend model), Assumption (ref) reduces to (ref) in Assumption (ref) with $t \in [0,1]$.

Define $\mathbf M(t) := \mathbb{E} (\mathbf{W}(t,\mathcal{F}_0)\mathbf{W}(t,\mathcal{F}_0)^{\top})$, $t \in [0,1]$. Notice that the first element of $\mathbf W$ is $1$. Write $\boldsymbol{ \mu}_W(\cdot) := (1, \mu_{W,2}(\cdot), \cdots, \mu_{W,p}(\cdot))^{\top} :=\mathbb{E}(\mathbf W(t,\mathcal{F}_0))$. For $p \geq 2$, Let $x_{i,j, n}$ denote the $j$th element in $\mathbf x_{i,n}$, and $\mathbf x_{i,n}^{(-1)} := (x_{i,2,n}, \cdots, x_{i,p,n})^{\top}$. Define $\boldsymbol \mu_W^{(-1)}(\cdot) := (\mu_{W, 2}(\cdot), \cdots, \mu_{W,p}(\cdot))^{\top}$, $\mathbf W^{(-1)}(\cdot, \cdot) := (W_2(\cdot, \cdot),\cdots, W_p(\cdot, \cdot))^\top$.

assumptionThe following conditions hold for the covariates when $p\geq 2$: \begin{enumerate}[label=(B\arabic*),itemsep=2pt,topsep=0pt,parsep=0pt] • The smallest eigenvalue of $\mathbf M(t)$ is bounded away from 0 on $[0, 1]$. • {$\mathbf M(\cdot) \in C^1[0,1]$, $\boldsymbol \mu_W (\cdot) \in C^1[0,1]$}. • $\mathbf{W}\left(t, \mathcal{F}_{i}\right) \in \mathrm{Lip}_{2}, \text { and } \sup _{t \in [0,1]}\left\|\mathbf{W}\left(t, \mathcal{F}_{i}\right)\right\|_{8}<\infty$, $i = 1, \cdots, n$. • $\delta_{8}(\mathbf W^{(-1)},k)=O(\chi^k)$ for some $ \chi \in (0,1)$. • $\mathbb{E}(H(t_j,\mathcal{F}_j)|\mathbf W(t_j,\mathcal{F}_j)) = 0$ for $j = 1,2,\cdots,n,$ a.s.. \end{enumerate}

Condition (ref) ensures that there is no multicolinearity among the explanatory variables. Assumption (ref) guarantees that the $\mathbf M(\cdot)$ and $\boldsymbol \mu_W(\cdot)$ have continuous derivatives. Assumption (ref) requires that the covariates are locally stationary. Condition (ref) imposes that each component of $\mathbf W^{(-1)}(\cdot, \cdot)$ is SRD. Condition (ref) assumes that $\mathbf x_{i,n}$ is a $p$-dimensional random vector uncorrelated with innovations, which is necessary for model identification. Our assumptions are very mild in the sense that we allow nonlinearity and heteroscedasticity for the covariates and errors, as well as the correlation between $\mathbf x_{i,n}$ and $e_{i,n}$. Assumptions (ref) and (ref) can be verified using similar arguments in wu2009quantile. When $p = 1$, we use the convention that $\mathbf x_{i,n}^{(-1)} = \emptyset$, $\mathbf W^{(-1)}(\cdot, \cdot) = \emptyset$, and $\boldsymbol \mu_W^{(-1)}(\cdot) = \emptyset$, $\mathbf M(\cdot)=\boldsymbol \mu_W(\cdot)=1$. Therefore Assumption (ref) always hold in this case.

assumption$\lambda(\boldsymbol \Sigma^{1/2}(u) \mathbf{M}^{-1}(u)\boldsymbol \mu^{\top}_W(u) \neq (\sigma_H(u), 0, \cdots, 0)^{\top}) > 0$.

The test statistics will degenerate under the null hypothesis when the parameters of the regression model (ref) lie in the hyperplane $\{u\in[0,1]: \boldsymbol \Sigma^{1/2}(u) \mathbf{M}^{-1}(u)\boldsymbol \mu^{\top}_W(u)= (\sigma_H(u), 0, \cdots, 0)^{\top}\}$. (ref) excludes such situation.

Asymptotic theory

The following theorem establishes the asymptotic distribution of the KPSS-type statistic $T_n$ (ref) under the null hypothesis.

theoremLet Assumptions (ref), (ref), (ref) and (ref) be satisfied, assuming $nb_n^6 \to 0$, $nb_n^{7/2}/(\log n)^4\to \infty$, we have that under the null hypothesis: (i) If Assumption (ref) holds, \begin{align} T_n\Rightarrow \int_0^1 U^2(t) dt, \end{align} where $U(t)$ is a zero mean continuous Gaussian process with covariance function \begin{align} \mathbb{E}(U( r)U(s)) =: \gamma( r,s) &=\int_0^{ r\wedge s}\sigma_H^2 (u) du - 2 \int_0^{ r\wedge s}\{\boldsymbol \mu^{\top}_W(u)\mathbf{M}^{-1}(u)\boldsymbol \Sigma^{1/2}(u)\}_1\sigma_H (u) du\notag\\ &+ \int_0^{ r\wedge s} \boldsymbol \mu^{\top}_W(u)\mathbf{M}^{-1}(u)\boldsymbol \Sigma(u)\mathbf{M}^{-1}(u)\boldsymbol \mu_W(u) du, \quad r,s \in [0,1], \end{align} where $\{\cdot\}_1$ denotes the first element of a vector. (ii) If Assumption (ref) fails, then $ s_1^{-1}(T_n/b_n-s_2 )\Rightarrow \chi^2_1, $ where $s_1$ and $s_2$ are constants, i.e., $s_1 = 2\sigma^2_H(0) \int_0^1 \left(\int_{v-1}^1 K^*(t) dt\right)^2 dv$, and $s_2 = 2\int_{0}^1 \sigma^2_H(t) dt \int_0^1 \left(\int_{v}^1 K^*(t) dt\right)^2 dv$.

Theorem (ref) reveals that for the time-varying coefficient model with time series covariates, the limiting distribution of $T_n$ depends on the time-varying mean and covariance matrix of the covariates $\mathbf x_{i,n}$ as well as the long-run covariance matrix of $\mathbf x_{i,n} e_{i,n}$. Theorem (ref) is very general since it posits neither the specific form of heteroscedasticity nor the parametric form of the errors. When Assumption (ref) is violated, (ii) shows that $T_n$ degenerates with asymptotic variance $2s_1^2 b_n^2$. After standardization, it converges to $\chi^2_1$ in distribution. The results of R/S, V/S and K/S follow similarly. An important scenario that $T_n$ degenerates under $H_0$ (i.e., $d=0$) is the following time-varying trend model corresponding to $p=1$, i.e.,

equation[equation omitted — 233 chars of source]
remarkTheorem (ref) and other theoretical results in this paper are valid for $p=1$, with $\boldsymbol \beta(t)$ replaced by $\beta_1(t)$, $\boldsymbol \Sigma(t)$ replaced by $\sigma_H^2(t)$, $\boldsymbol \mu_W(t)$ replaced by $1$.

Gaussian approximation

Since the KPSS and related test statistics are constructed based on the partial sum process, to derive the asymptotic properties of the test statistics under the alternative hypothesis, we first study the Gaussian approximation of $\sum_{i=1}^r \mathbf x_{i,n} e_{i,n}^{(d)}$, $1 \leq r \leq n$, which is the partial sum of the product of a SRD and a LRD time series under $H_A$. Though Gaussian approximation theory for stationary processes ( see for instance wu2007strong, dehling1989empirical, wu2006invariance and the reference therein) has been successfully established and widely applied to many fields of statistics, there are only a few results of Gaussian approximation for locally stationary processes. Among them, wu2011gaussian established a flexible Gaussian approximation framework for locally stationary SRD processes, which has served as a fundamental key to the inference of SRD (piecewise) locally stationary processes and functional time series, see for instance chen2015function and wu2018gradient. wu2018 proposed a Gaussian approximation scheme for a class of locally stationary linear LRD processes. However, all the existing Gaussian approximation approaches are not applicable to the partial sum process of the product series $ (\mathbf x_{i,n} e_{i,n}^{(d)})$, which serves as the crucial ingredient for establishing the limiting distribution of $T_n$ under $H_A$. To this end, we shall provide a general Gaussian approximation theorem for the product of LRD and SRD processes. In the remaining of this paper, let $ d_n = c/ \log n$, where $c$ is a positive constant. We substitute $d$ with $d_n$ to differentiate the notation under the fixed alternatives $(d>0)$ and that under the local alternatives $(d=c/\log n)$.

theoremUnder Assumptions (ref) and (ref), on a richer probability space, we have: (i) There exists $\mathbf R_{k, n}=\sum_{j=0}^{\infty} \boldsymbol \mu_W(t_k)\psi_j(d)\sigma_H \left(t_{k-j}\right) v_{k-j}$, where the random variables $(v_i)_{i \in \mathbb Z}$ are $i.i.d.$ $N(0,1)$, s.t. \begin{align} \max_{\lfloor nb_n \rfloor + 1 \leq r \leq n - \lfloor nb_n \rfloor} \left|\sum_{i = \lfloor nb_n \rfloor +1}^r (\mathbf x_{i,n} e_{i,n}^{(d)} - \mathbf{R}_{i,n}) \right| = O_{\mathbb{P}}(\sqrt{n}(\log n)^d + n^{1+\alpha_0(d-1/2)}), \end{align} where $ \alpha_0 \in (1, 4/3) $ and therefore $n^{1+\alpha_0(d-1/2)} = o(n^{d+1/2})$. (ii) Further under Assumption (ref), there exists a sequence of Gaussian processes $$\mathbf{\tilde R}_{i,n} = \sum_{j = 1}^{\infty} \boldsymbol \mu_W(t_i) \psi_{j}(d_n)\sigma_H(t_{i-j}) V_{i - j ,1} + \boldsymbol \Sigma^{1/2}(t_i) \mathbf V_i,$$ where $\mathbf V_i$, $1 \leq i \leq n$, are $i.i.d.$ $N(\mathbf 0, \mathbf I_p)$, s.t. \begin{align} \max_{\lfloor nb_n \rfloor + 1 \leq r \leq n - \lfloor nb_n \rfloor} \left|\sum_{i = \lfloor nb_n \rfloor +1}^r (\mathbf x_{i,n} e_{i,n}^{(d_n)} - \mathbf{\tilde R}_{i,n}) \right| = o_{\mathbb{P}}(n^{1/2}). \end{align}

Since $\|{\mathbf R}_{i,n}\| \asymp n^{d+1/2}$ and $\|\tilde {\mathbf R}_{i,n}\| \asymp n^{1/2}$, the approximation errors of Theorem (ref) (i) and (ii) are asymptotically negligible. It can be also verified that the process $(\mathbf R_{k,n})_{k=1}^n$ is a locally stationary LRD Gaussian process defined by Definition (ref).

remarkThe results of (i) consist of two parts. The rate $\sqrt{n} (\log n)^d$ is due to the approximation of $\max_{1\leq s \leq n}|\sum_{k=1}^s(\boldsymbol x_{k,n} - \boldsymbol\mu_W(t_k))e_{k,n}^{(d)}|$, see (ref). The rate $n^{1+\alpha_0(d-1/2)}$ is due to the approximation of $\max_{1\leq s \leq n}|\sum_{k=1}^s\boldsymbol \mu_W(t_k)e_{k,n}^{(d)}-\mathbf R_{k,n}|$, which extends Theorem 2 in wu2018, see (ref). Specifically, we allow the driving shocks $(u_{i,n})_{i=-\infty}^n$ to be both dependent and heteroscedastic, while wu2018 assumed $(u_{i,n})_{i=-\infty}^n$ to be independent.

It is worth pointing out that (ii) is not an direct consequence of (i). Letting $d=d_n$ in (i), the rate $\sqrt n(\log n)^d$ in (i) will eventually lead to a trivial bound, i.e., $O_{\mathbb{P}}(\sqrt{n})$, which is of the same order as the partial sum $\sum_{i = \lfloor nb_n \rfloor +1}^r \mathbf x_{i,n} e^{(d_n)}_{i,n}$, $\lfloor nb_n \rfloor +1\leq r\leq n-\lfloor nb_n \rfloor $.

Based on the Gaussian approximation result, we proceed to study the limiting distributions of KPSS and related statistics under the fixed and local alternatives for the time-varying coefficient model (ref) with time series covariates. For the sake of brevity, we focus on the KPSS-type statistics. The results for R/S, V/S, and K/S-type statistics can be derived similarly and are summarized in Remark (ref) and \cref*{sub:limits} in the online supplement.

Fixed alternatives

In the following theorem, we establish the asymptotic distribution of the KPSS-type statistic $T_n$ (ref) under the fixed alternatives.

theoremUnder Assumptions (ref), (ref), (ref) and (ref), assuming $nb_n^4/(\log n)^2 \to \infty, n b_n^6 \to 0$, and $\lambda(|\boldsymbol \mu_W^{(-1)}(\cdot))| \neq 0) > 0$, then we have under $H_A$ with long-memory parameter $d$, \begin{align} T_{n}\Gamma^2(d+1) /n^{2d } \Rightarrow \int_0^1 U^2_d(t) dt, \end{align} where $U_d(t)$ is a zero mean continuous Gaussian process with covariance function \begin{align} \mathbb{E}(U_d(r)U_d(s)) := \gamma_d(r,s) = \int_{-\infty}^{r \wedge s}\sigma^2_H(v)\lambda_d(r,v)\lambda_d(s,v) dv , \quad r, s \in [0,1], \end{align} where for $v \leq u \in [0,1]$, $ \lambda_d(u,v) = d\int_{(-v)_+}^{(u-v)_+}t^{d-1}(\check M_W(t+v)-1) dt $, $\check M_W(t) = \boldsymbol \mu_W^{\top}(t) \mathbf{M}^{-1}(t)\boldsymbol \mu_W(t)$, $t \in [0,1]$.

Theorem (ref) proves that the test statistic diverges to infinity at the rate of $n^{2d}$. Furthermore, from Theorem (ref), we observe that the limiting distribution of $T_n$ under the fixed alternatives is independent of ${\boldsymbol \Sigma}(t)$ except $\sigma^2_H(t)$, which is its $(1,1)$ component, while under the null hypothesis the limiting distribution relies on all the components of ${\boldsymbol \Sigma}(t)$, see Theorem (ref). This is because when $d>0$, the stochastic fluctuation of the SRD components $(\mathbf x_{i,n})$ is asymptotic negligible compared to that of $(e^{(d)}_{i,n})$.

Straightforward calculation shows that $\check M_W(t)\geq 1$. The condition $\lambda(|\boldsymbol \mu_W^{(-1)}(\cdot))| \neq 0) > 0$ ensures $\check M_W(t)>1$ in an interval with positive length such that $\lambda_d(u,v)>0$ and excludes the scenario that all the stochastic covariates are zero mean during the whole period (see (ref) for detailed discussion of such scenario) as well as the time-varying trend model,i.e., (ref) with $p=1$. When $\lambda_d(u,v)=0$, we can show that $T_n=o_{\mathbb{P}}(n^{2d})$, i.e., the test statistic is degenerate under $H_A$.

remarkThe error model (ref) is in fact a Type I fractional $I(d)$ process when $d>0$. If a locally stationary Type II fractional $I(d)$ error process is considered, i.e., $(1- \mathcal{B})^d e_{i,n} = u_{i,n}\mathbf 1(i \geq 1)$, the limiting distribution of $T_n$ is almost the same except that the lower bound of the integral in $\gamma_d(r,s)$ is $0$ instead of $-\infty$. We refer to marinucci1999alternative for the definition of Type I and Type II fractional $I(d)$ processes.

Local alternatives

The following theorem presents the asymptotic distribution of the KPSS-type statistic $T_n$ (ref) under the local alternatives.

theoremLet the conditions of (ref) and Assumption (ref) hold. Then under $H_A$ with $d_n=c/\log n$ for a constant $c>0$, we have \begin{align} T_n \Rightarrow \int_0^1 (U^{\circ}(t))^2dt, \end{align} where $U^{\circ}(t)$ is a zero mean continuous Gaussian process with covariance function \begin{align} \mathbb{E}(U^{\circ}(r)U^{\circ}(s))=: \gamma^{\circ} (r,s) &= \check \gamma(r, s) + \gamma(r, s) + 2\tilde \gamma(r, s) , \quad r, s \in [0,1], \end{align} where $\gamma(r, s)$ is defined in (ref), and \begin{align} &\tilde \gamma(r, s) = (e^c - 1)\int_0^{r\wedge s} \sigma_H(t)(\{\boldsymbol \mu_W^{\top}(t) \mathbf M^{-1}(t) \boldsymbol \Sigma^{1/2}(t)\}_1 - \sigma_H(t))(\check M_W(t) - 1) dt, \\ &\check \gamma(r, s) = (e^c -1)^2\int_0^{r\wedge s} \sigma^2_H(t) (\check M_W(t)-1)^2 dt, \end{align} where $\{\cdot\}_1$ and $\check M_W(t)$ are as defined in (ref) and (ref), respectively.

From (ref), we shall see that under the local alternatives $d_n=c/\log n$, the KPSS-type statistic converges to a distribution depending on the mean and covariance matrix of $(\mathbf x_{i,n})$, the long-run covariance matrix of $(\mathbf x_{i,n}e_{i,n})$, as well as the parameter $c$. Careful examination of the proof of Theorem (ref) shows that $T_n$ will converge to the limit in Theorem (ref) if $d_n\log n=o(1)$, indicating that the exact local power of the KPSS-type test is $O(\log^{-1} n)$ for model (ref). For the time-varying trend model (ref), it can be shown that the test statistic is degenerate (since $\lambda(|\boldsymbol \mu_W^{(-1)}(\cdot))| \neq 0) =0$) under local alternatives $d_n = c/\log n$, i.e., $T_n=o_{\mathbb{P}}(1)$. The results for R/S, K/S and V/S-type tests follow similarly.

remarkThe limiting behavior of R/S, V/S and K/S-type statistics defined in (ref) under (ref) can be derived by Theorems (ref), (ref) and (ref) as well as an application of continuous mapping theorem. Their limiting distributions are functions of $U(t)$, $U_d(t)$, $U^{\circ}(t)$ defined therein, see \cref*{sub:limits} in the online supplement for the exact forms.

The bootstrap-assisted procedure

(ref) shows that under the null hypothesis, the limiting distributions of KPSS and related test statistics are functions of the Gaussian process $U(t)$ which involves parameters $\boldsymbol \mu_W(t), \mathbf M(t), \boldsymbol \Sigma(t)$ (or $\sigma^2_H(t)$ when $p=1$). Furthermore, the magnitude and specific form of the limiting distributions depend on whether (ref) holds, which is usually unknown in practice. Therefore, it's impossible to obtain the critical values by directly simulating the Gaussian process $U(t)$. In this section, we provide a consistent bootstrap approach (ref) which mimics the asymptotic behavior of $T_n$ under the null hypothesis no matter whether (ref) is satisfied and yields valid simulated critical values. In addition, the estimation of $\boldsymbol \mu_W(\cdot)$ is not required in our proposed bootstrap tests. More precisely, we employ $\mathbf x_{i,n}^\top \hat{\mathbf M}^{-1}(t)$ instead of $\hat{\boldsymbol \mu}^\top_W(t) \hat{\mathbf M}^{-1}(t)$, where $\hat{\boldsymbol \mu}_W(t)$ stands for any consistent estimator of $\boldsymbol \mu_W(t)$, and the validity of the former construction can be easily verified by noting that in (ref) of (ref), the Gaussian multiplier is independent of $ \mathbf x_{i,n}$, $\hat {\mathbf M}(\cdot)$ and $\hat {\boldsymbol \Sigma}(\cdot)$ where the consistent estimators $\hat{\mathbf M}(t)$ and $\hat {\boldsymbol \Sigma}(t)$ will be discussed later. Thus, the difference between (ref) and the counterpart with $\boldsymbol \mu_W(t)$ can be controlled by the convolution of standard Gaussian multipliers and the partial sum of the zero mean SRD process containing $(\mathbf x_{i,n}-\boldsymbol \mu_W(t_i))$. We only discuss (ref) for the KPSS-type test when $p\geq 2$ in detail in this section. The other algorithms, including \cref*{trend_algorithms} in the online supplement for $p=1$ case corresponding to the time-varying trend model (ref), and \cref*{algorithms} for R/S, V/S, K/S-type tests are moved to the supplement.

algorithm[algorithm omitted — 2,017 chars of source]

To implement (ref), we need to obtain the estimators $\hat {\mathbf M}(t)$ and $\hat {\boldsymbol \Sigma}(t)$. For $\hat {\mathbf M}(t)$ we propose the following estimator that uses directly observed covariates $\mathbf x_{i,n}$:

equation[equation omitted — 143 chars of source]

where $t^* = \max\{\eta_n,\min(t,1-\eta_n)\}$ for some bandwidth $\eta_n \to 0$, $n\eta_n^2 \to \infty$. Under Assumptions (ref) and (ref), after a careful investigation of Lemma 6 of zhou2010simultaneous, we have $\sup_{t \in[\eta_n,1-\eta_n]}|\hat{\mathbf{M}}(t)- \mathbf M(t)|=o_{\mathbb{P}}(1)$, i.e. $\hat{\mathbf{M}}(t)$ is uniformly consistent, see \cref*{lm:scblemma6} in the online supplement for details.

Observing that $\boldsymbol \Sigma(\cdot)$ depends on the unobserved error process $(e_{i,n})$. Therefore, a common approach to estimate $\boldsymbol \Sigma(\cdot)$ is to utilize $\hat e_{i,n}'s$ from the local linear fit (ref), see for instance zhou2010simultaneous and vogt2015detecting. Such plug-in estimate will yield test results that are sensitive to the choices of $b_n$ under $H_0$, and even more sensitive under $H_A$. We also find through our extensive numerical studies (which are not reported in this paper due to limited space) that the power of the test using the plug-in estimator of long-run covariance is often unsatisfactory. Therefore, we adopt a difference-based estimator that does not involves $\hat e_{i,n}'s$.

For $p=1$ we recommend the difference statistic proposed in (4.7) of Section 4.2 of dette2019detecting for $\sigma^2_H(t)$ which is built on the difference of $y_{i,n}$. For $p \geq 2$, it can be shown that a direct extension of the estimator based on the difference of $(\mathbf x_{i,n}y_{i,n}$) is asymptotically biased. Therefore, we adopt the following bias-corrected difference-based estimator proposed by lrvunpublish. Let $\mathbf Q_{k, m}=\sum_{i=k}^{k+m-1} \mathbf x_{i,n}y_{i,n}$, for $t \in [m/n,1-m/n]$,

align[align omitted — 243 chars of source]

where $ \omega(t, i)= K_{\tau_n}\left(t_i-t\right) / \sum_{i=1}^{n} K_{\tau_{n}} \left(t_i-t\right) $ for some bandwidth $\tau_n$ and the kernel function $K(\cdot)$ with support $(-1,1)$. For $t \in [0,m/n)$, set $\acute{{\boldsymbol \Sigma}}(t) = \acute{{\boldsymbol \Sigma}}(m/n)$. For $t \in (1-m/n,1]$, set $\acute{{\boldsymbol \Sigma}}(t) = \acute{{\boldsymbol \Sigma}}(1-m/n)$. The bias-corrected difference-based estimator $\hat {\boldsymbol \Sigma}(t)$ for $t \in [0,1]$ is then defined as:

align[align omitted — 283 chars of source]

where $ \hat{\mathbf A}_{j} = \frac{1}{m}\sum_{i= j-m+1}^{j} (\mathbf x_{i,n} \mathbf x_{i,n}^{\top}\breve {\boldsymbol \beta}(t_i)-\mathbf x_{i+m, n} \mathbf x_{i+m, n}^{\top}\breve {\boldsymbol \beta}(t_{i+m}))$, $ \breve {\boldsymbol \beta}(t) = \boldsymbol \Omega^{-1}(t)\boldsymbol \varpi (t), $ where $\boldsymbol\Omega(t),\boldsymbol \varpi (t)$ are the smoothed versions of $\acute{\boldsymbol \Delta}_{j} := \frac{1}{m}\sum_{i=j-m+1}^{j} \tilde{\mathbf X}_{i,m}\tilde{\mathbf X}_{i,m}^{\top}$ and $\breve {\boldsymbol \Delta}_{j} := \frac{1}{m}\sum_{i=j-m+1}^{j} \tilde{\mathbf X}_{i,m}^{\top}\tilde{\mathbf Y}_{i,m}$, i.e., $ \boldsymbol\Omega(t) = \sum_{j=m}^{n-m} \acute {\boldsymbol \Delta}_{j}\omega(t^*, j)/2,\nonumber$ and $\boldsymbol \varpi (t) = \sum_{j=m}^{n-m} \breve {\boldsymbol \Delta}_{j} \omega(t^*, j)/2. $ To make (ref) operational, it's necessary to select smoothing parameters $\eta_n$, $\tau_n$ and $m$ for $\hat {\mathbf M}(t)$ and $\hat {\boldsymbol \Sigma}(t)$. The selection of smoothing parameters is postponed to \cref*{sel} of the online supplement.

The limiting behavior of the bootstrap tests

In this section, we shall show the asymptotic correctness of the bootstrap test (ref). We shall also prove that the power of (ref) against $d > 0$ can be no less than $(n/m)^{2d}$, and that the exact local power of (ref) can achieve the order $O(\log^{-1} n)$. We start by defining the long-run {\it cross} covariance vector between the locally stationary processes $ (\mathbf U(t,\mathcal{F}_i))$ and $(H(t,\mathcal{F}_j))$.

definitionDefine the long-run cross covariance vector $\mathbf s_{UH}(t) \in \mathbb R^p$ by \begin{align} \mathbf s_{UH}(t) = \sum_{j=-\infty}^{\infty} \mathrm{Cov}(\mathbf U(t, \mathcal{F}_0), H(t, \mathcal{F}_j)), \quad t \in [0, 1].\nonumber \end{align}

When $p=1$, $\mathbf s_{UH}(t)$ degenerates into $\sigma_H^2(t)$. We assume the following conditions for $\hat{\boldsymbol \Sigma}(t)$ used in (ref). Write $\hat{\boldsymbol \Sigma}(t)$ defined in (ref) as $\hat{\boldsymbol \Sigma}_d(t)$ under the fixed alternatives and $\hat{\boldsymbol \Sigma}_{d_n}(t)$ under the local alternatives $d_n$. Let $\mathcal{I} = [\gamma_n, 1-\gamma_n] \subset (0,1)$, where $\gamma_{n}=\tau_{n}+(m+1) / n$.

assumptionThe long-run variance estimator satisfies the following conditions (i) Under the null hypothesis, \begin{equation} \sup _{t \in \mathcal{I}}\left|\hat{{\boldsymbol \Sigma}}(t)-{\boldsymbol \Sigma}(t)\right|= o_{\mathbb{P}}(b_n/\log^2 n).\nonumber \end{equation} (ii) Under the fixed alternatives, \begin{equation} \sup_{t \in \mathcal{I}}\left| m^{-2d}\hat{{\boldsymbol \Sigma}}_d(t) - {\boldsymbol \Sigma}_d(t)\right| = o_{\mathbb{P}}(1),\nonumber \end{equation} where ${\boldsymbol \Sigma}_d(t) = \kappa_2(d)\sigma_H^2(t) \boldsymbol \mu_W(t)\boldsymbol \mu^{\top}_W(t)$, $t \in [0,1]$, and $\kappa_2(d) = \Gamma^{-2}(d+1)\int_{0}^{\infty}(t^d - (t-1)_+^d)(2t^d - (t-1)_+^d - (t+1)^d) dt$. (iii) Under the local alternatives $d_n = c/\log n$, $m = \lfloor n^{\alpha_1} \rfloor$, $\alpha_1 \in (0,1)$, \begin{align} \sup_{t \in \mathcal{I}}\left|\hat{{\boldsymbol \Sigma}}_{d_n}(t) - \check {\boldsymbol \Sigma}(t)\right| = o_{\mathbb{P}}(1),\nonumber \end{align} where $\check {\boldsymbol \Sigma}(t) := {\boldsymbol \Sigma}(t) + (e^{c\alpha_1}-1)^2\sigma_H^2(t) \boldsymbol \mu_W(t)\boldsymbol {\mu}^{\top}_W(t) + (e^{c\alpha_1}-1)(\mathbf s_{UH}(t) \boldsymbol \mu_W^{\top}(t) +\boldsymbol \mu_W(t)\mathbf s^{\top}_{UH}(t))$.

It can be shown that the both the plug-in estimator of zhou2010simultaneous and the bias-corrected estimator (ref) satisfies this condition under suitable bandwidth conditions following Theorem 3.1 of wu2006invariance and the chaining argument of Propostion B.1 of dette2018change. Let $\tilde T_n$ denote the bootstrap statistic (ref) generated in one iteration. Recall the definitions of $U(t)$, $s_1$, $s_2$ in (ref). (ref) gives the limiting distributions of bootstrap statistic $\tilde T_n$.

theorem[Bootstrap under null] Assume the conditions (ref), (ref), (ref), (ref) and (ref) hold, $nb_n^{7/2}/(\log n)^4\to \infty$, $nb_n^6 \to 0$, $\eta_n \to 0$, $n\eta_n^2 \to \infty$. Then, under the null hypothesis, we have (i) if (ref) holds, then $ \tilde T_{n} \Rightarrow \int_0^1 U^2(t) dt$ . (ii) if (ref) doesn't hold, then $ s_1^{-1}(\tilde T_n/b_n-s_2 )\Rightarrow \chi^2_1$.

Combining with (ref), (ref) indicates that the bootstrap test (ref) is asymptotically of level $\alpha$ {\it no matter whether (ref) is satisfied}. We proceed to investigate the behavior of the bootstrap statistic $\tilde T_n$ under fixed and local alternatives. Let $\tilde U_d(t)$, $\check U(t)$ be zero mean continuous Gaussian processes, with the covariance structures defined in the same way as that of $U(t)$ in (ref), where $\boldsymbol \Sigma(t)$ is replaced by $\boldsymbol \Sigma_d(t)$ and $\breve{\boldsymbol \Sigma}(t)$ in (ref), respectively. Let $\sigma_{Hd}^2(t) := (\boldsymbol \Sigma_d(t))_{(1,1)}$, $\check \sigma_{H}^2(t) := (\check{\boldsymbol \Sigma}(t))_{(1,1)}.$

theorem[Bootstrap under alternatives] Under the conditions of (ref), (i) Suppose $\lambda({\boldsymbol \Sigma}^{1/2}_d(t)\mathbf{M}^{-1}(t) \boldsymbol \mu_W(t) \neq (\sigma_{Hd}(t), 0, \cdots, 0)) > 0$ under the fixed alternatives $d>0$. Then, we have $$m^{-2d} \tilde T_{n} \Rightarrow \int_0^1 \tilde U^2_d(t) dt,$$ where $\tilde U_d(t)$ is a zero-mean continuous Gaussian process with covariance function \begin{align} \mathbb{E}(\tilde U_d(r)\tilde U_d(s)) &= \int_0^{r \wedge s }\boldsymbol \mu_W^{\top}(t) \mathbf{M}^{-1}(t) {\boldsymbol \Sigma}_d(t)\mathbf{M}^{-1}(t) \boldsymbol \mu_W(t) dt \nonumber \\ & - 2 \int_{0}^{r \wedge s } \{\boldsymbol \mu_W^{\top}(t) \mathbf M^{-1}(t) {\boldsymbol \Sigma}_d^{1/2}(t)\}_1 \sigma_{Hd}(t) dt + \int_{0}^{r \wedge s} \sigma_{Hd}^2(t) dt, \quad r,s \in [0,1]. \nonumber \end{align} (ii) Suppose $\lambda(\check{\boldsymbol \Sigma}^{1/2}(t)\mathbf{M}^{-1}(t) \boldsymbol \mu_W(t) \neq (\check \sigma_{H}(t), 0, \cdots, 0)) > 0$. For the local alternatives $d_n = c /\log n$ with some positive constant $c$, we have $$ \tilde T_{n} \Rightarrow \int_0^1 \check U^2(t) dt,$$ where $\check U(t)$ is a zero-mean continuous Gaussian process with covariance function \begin{align} \mathbb{E}(\check U(r)\check U(s)) &= \int_0^{r \wedge s } \boldsymbol \mu_W^{\top}(t) \mathbf{M}^{-1}(t) \check {\boldsymbol \Sigma}(t)\mathbf{M}^{-1}(t) \boldsymbol \mu_W(t) dt \\ &- 2 \int_{0}^{r \wedge s } \{\boldsymbol \mu_W^{\top}(t) \mathbf M^{-1}(t) \check {\boldsymbol \Sigma}^{1/2}(t)\}_1 \check \sigma_{H}(t) dt + \int_{0}^{r \wedge s} \check \sigma_{H}^2(t) dt, \quad r,s \in [0,1]. \end{align}

(ref) (i) gives the limiting distribution of $\tilde T_{n}/ m^{2d}$, which means the critical values generated by (ref) for a level $\alpha$ test diverges at the rate of $m^{2d}$. Thus, the bootstrap-assisted test is consistent since (ref) demonstrates that the KPSS-type test statistic $T_n$ diverges at the rate $n^{2d}$ which is much faster than $m^{2d}$. Further together with (ref), (ref) (i) shows that the bootstrap test (ref) is asymptotically correct. Notice that the condition $\lambda({\boldsymbol \Sigma}_d(t)\mathbf{M}^{-1}(t) \boldsymbol \mu_W(t) \neq (\sigma_{Hd}(t), 0, \cdots, 0)) > 0$ prevents the degeneracy of bootstrap statistics. If it is violated, $\tilde T_n= o(m^{2d})$ which will yield higher power than when the condition is fulfilled.

On the other hand, (ref) (ii) and (ref) indicate that the bootstrap test (ref) is able to detect the local alternatives at the rate of $\log^{-1} n$. Observe that under (ref) (iii) as $c\rightarrow 0$, $|\check{\boldsymbol \Sigma}(t)-\boldsymbol \Sigma(t)|\rightarrow 0$ and the covariance structure of $\check{U}(t)$ will converge to the covariance structure of $U(t)$. Therefore, the bootstrap test (ref) has no power when $d_n=o( \log^{-1} n)$, indicating that the proposed test has the exact local power of $O(\log^{-1} n)$ under the condition of (ref). For stationary time series with unknown constant mean, shao2007local has also proved that the KPSS test for long memory has the exact local power $O(\log^{-1} n)$.

An important scenario that $T_n$ degenerates under both null and alternatives is the time-varying trend model (ref). For this case we provide the bootstrap-assisted KPSS test in \cref*{trend_algorithms} of the online supplement, which is the $p=1$ version of (ref) with $\hat {\mathbf M}(t) = 1$, and $\hat{\boldsymbol \Sigma}(t) = \hat \sigma_H^2(t)$ using the difference statistic proposed by dette2019detecting. (ref) (ii) ensures that the level $\alpha$ critical value generated by \cref*{trend_algorithms} converges to the $\alpha_{th}$ quantile of the test statistic $T_n$ under the null hypothesis. Under the alternative hypothesis, the following proposition investigates the power of the test implemented via \cref*{trend_algorithms}.

propositionLet $\tilde T_n$ denote the KPSS-type bootstrap statistic generated from \cref*{trend_algorithms}. Under Assumptions (ref), (ref), (ref), (ref), and the bandwidth conditions $nb_n^4/(\log n)^2 \to \infty, b_n \to 0$, $nb_n/m \to \infty$, we have the following results: (i) Under the fixed alternatives $d > 0$, $ \lim_{n \to \infty} P\left( T_n > \tilde T_n\right) = 1. $ (ii) Suppose $m = \lfloor n^{\alpha_1} \rfloor$, $nb_n = n^{\beta}$ for some $\alpha_1, \beta \in (0,1)$. Then under local alternatives with $d = d_n = c/\log n$ for a sufficiently large constant $c$, $ \lim_{n \to \infty} P\left( T_n > \tilde T_n\right) = 1. $

In addition to the KPSS test, (ref) also provides the bootstrap-assisted V/S, R/S and K/S tests for the time-varying trend model (ref). Similar conclusions as given in (ref) hold for the power of these tests.

remarkFor $p\geq 2$, $T_n$ degenerates under the fixed alternatives if and only if all the stochastic covariates are of mean zero (i.e., $\mu_W^{(-1)}(t) = 0$), see (ref). Under this condition, we can show that $T_n/b_n$ diverges at the rate of $(nb_n)^{2d}$. Combining with (ref) (i), our bootstrap tests (ref) and (ref) are consistent if $b_n (nb_n/m)^{2d}\to \infty$. If $\sum_{i \in \mathbb Z} \mathrm{Cov} (H(t, \mathcal{F}_0), H(t, \mathcal{F}_i)W^{(-1)}(t, \mathcal{F}_i)) = 0$, $T_n$ degenerates under $H_0$ and $H_A$ if and only if $\mu_W^{(-1)}(t) = 0$. In this case, one could implement (ref) and (ref) for R/S, V/S, K/S-type tests via modifying the difference-based long-run covariance estimator: set $\hat{\boldsymbol \Sigma}(t)_{(1,l)} = 0$ and $\hat{\boldsymbol \Sigma}(t)_{(l,1)} = 0$ for $l=2, \cdots, p$. Then, by a further investigation of the proof to (ref), $\tilde T_n/b_n$ is $O_{\mathbb{P}}(m^{2d})$ and thus the tests are consistent when $m/(nb_n) \to 0$. In addition, similar to (ref), we can also show that the local power of the tests with the modified estimator is of order $O(\log^{-1} n)$.

Finite sample performance

In the following simulation studies and data analysis, we examine the size and power performance of the bootstrap-assisted KPSS and related tests (ref), \cref*{trend_algorithms} and \cref*{algorithms} in the online supplement. The number of bootstrap samples is $B = 2000$ and the number of replications is $1000$. The parameters $\mathbf M(t)$, ${\boldsymbol \Sigma}(t)$, $\sigma^2_H(t)$ are estimated by $\hat{\mathbf M}(t)$, $\hat {\boldsymbol \Sigma}(t)$, $\hat \sigma^2_H(t)$ in (ref), with all the smoothing parameters selected by the methods advocated in \cref*{sel} in the online supplement. Let

align[align omitted — 178 chars of source]

where $(\varepsilon_{l})_{l\in \mathbb Z}, (\zeta_{l})_{l\in \mathbb Z}$ are $i.i.d.$ $ N(0,1)$. We consider the following time-varying coefficient model, $$ y_{i, n}=\beta_{1}(t_i)+\beta_{2}(t_i) x_{i, n}+e_{i,n}, \quad i=1, \ldots, n, $$ where $\beta_{1}(t)=4 \sin (\pi t) $, $\beta_{2}(t)=4 \exp \{-2 \left(t-0.5\right)^{2}\} $, $x_{i, n}=W(t_i, \mathcal{F}_{i})$, and $ u_{j,n} = H(t_j,\mathcal{F}_j, \mathcal{G}_j)$. First, we consider the following independent model:

enumerate[label=(\roman*),itemsep=2pt,topsep=0pt,parsep=0pt] • Let $W\left(t, \mathcal{F}_{i}\right)= (0.25 + 0.25\cos(2\pi t))W(t, \mathcal{F}_{i-1})+ 0.25\zeta_{i} + (t-0.5)^2$ , $ H(t, \mathcal{G}_{i})=(0.35 - 0.4(t-0.5)^2) H(t, \mathcal{G}_{i-1})+ 0.8 \varepsilon_{i}$.

Second, we consider the following heteroscedastic model:$$ H(t,\mathcal{F}_{i},\mathcal{G}_{i}) = B\left(t,\mathcal{G}_{i}\right)\sqrt{1+W^2(t, \mathcal{F}_{i})},$$ where $ W\left(t, \mathcal{F}_{i}\right)= (0.1 + 0.1\cos(2\pi t))W(t, \mathcal{F}_{i-1})+ 0.2\zeta_{i} + 0.7(t-0.5)^2, $ and $B\left(t,\mathcal{G}_{i}\right)$ is as considered in the following linear and nonlinear scenarios.

enumerate[label=(ii.\arabic*),itemsep=2pt,topsep=0pt,parsep=0pt] • Linear errors: $ B(t, \mathcal{G}_{i})=(0.3 - 0.4(t-0.5)^2) B(t, \mathcal{G}_{i-1})+ 0.8\varepsilon_{i}. $ • Nonlinear errors: $ B(t, \mathcal{G}_{i})=(0.15 - 0.4(t-0.5)^2) B(t, \mathcal{G}_{i-1})+ 0.8G(t, \mathcal{G}_{i})$,$ G(t, \mathcal{G}_{i}) = \varepsilon_i \sigma_i(t), $ where $\sigma^2_i(t) = 0.9+0.1\cos(\pi/3 + 2\pi t) + (0.1+0.2t)G^2(t, \mathcal{G}_{i-1})+ (0.1 + 0.2t)\sigma^2_{i-1}(t).$

Observe that models (ref) and (ref) are heteroscedastic models with locally stationary AR(1) and locally stationary GARCH(1,1) errors, respectively. (ref) summarizes the performance of our proposed bootstrap-assisted KPSS, R/S, V/S and K/S-type tests for long memory in models (ref) and (ref) with different $b_n's$. We relegate the simulated sizes of model (ref) with different $b_n's$ to \cref*{tb:m0size1} of the online supplement. The empirical sizes of all the four tests are close to their nominal levels and are quite stable when $b_n$ changes within a reasonably wide range. Also, \cref*{tb:Tn} in the online supplement reports the simulated Type I error of the proposed tests with respect to increasing sample sizes. As shown in \cref*{tb:Tn}, our procedures for smoothing parameter selection including GCV and MV selection as well as the difference-based long-run variance estimator work very well in the sense that the simulated sizes of all four tests are quite close to their nominal levels in different sample sizes.

table[table omitted — 1,283 chars of source]
figure[figure omitted — 394 chars of source]

Recall the long-memory error process of (ref), which can be written as $e_{i,n}^{(d)} = (1-\mathcal{B})^{-d}u_{i,n}$ where $\mathcal{B}$ is the lag operator. (ref) displays the power performance of the KPSS and related tests for model (ref) with nominal level $0.1$. The left panel reports simulated rejection rates of KPSS, R/S, V/S, and K/S-type tests as the long memory parameter $d$ increases from $0$ to $0.5$ with sample size $1500$. The power of all four KPSS and related tests increases to 1 as $d$ approaches to $1/2$, while the V/S-type test remains the most powerful among all tests. The right panel depicts the power performance of the four tests as the sample size grows from $200$ to $2500$ when $d = 0.4$. The figure shows that the power of each test increases to $1$ as the sample size grows, and the V/S-type test performs the best among the four tests whenever the sample size is larger than $400$. The power performance of models (ref) and (ref) are shown in \cref*{fig:power1} and \cref*{fig:power3} of the online supplement where the V/S-type test achieves the best power under most alternatives and sample sizes considered.

Analysis of Hong Kong hospital data

Hong Kong circulatory and respiratory data contains daily measurements of pollutants and daily hospital admissions in Hong Kong between January 1st, 1994 and December 31st, 1995. The dataset has been studied by fan2000simultaneous, zhou2010simultaneous, and wu2018gradient among others. They investigated the relationship between the levels of pollutants and the total number of hospital admissions of circulation and respiration,see \cref*{fig:hk_data} in \cref*{hk} of the supplement for the observed series. Under the assumption of $i.i.d.$ observations, fan2000simultaneous claimed that sulphur dioxide ($\text{SO}_2$) is not significant. Using the non-stationary model, zhou2010simultaneous found all three pollutants ($\text{SO}_2$, nitrogen dioxide ($\text{NO}_2$) and dust) are significant. In both cases, they assumed the observations were SRD. We shall examine this assumption via the KPSS and related tests. Consider the following time-varying coefficient model,

equation[equation omitted — 148 chars of source]

where $(y_{i,n})$ is the series of daily total number of hospital admissions of circulation and respiration and $(x_{i, p, n})$, $p = 2, 3, 4$, represent the series of daily levels (in micrograms per cubic meter) of $\text{SO}_2$, $\text{NO}_2$ and dust, respectively. The sample size is $n = 2 \times 365 = 730$. We first investigate whether the covariates $(x_{i, p, n})$, $p = 2, 3, 4$ are SRD series.

We select the smoothing parameters $b_n$, $\eta_n$, $m$ and $\tau_n$ through the methods provided in \cref*{sel} in the online supplement and summarize the selected parameters in \cref*{tb:hk_par} in \cref*{hk} of the online supplement. We test for long memory in the three pollutants series via KPSS and related tests, presenting the $p$-values in (ref). For each pollutant series we fail to reject it is SRD at the significance level $0.05$.

table[table omitted — 515 chars of source]

To test for long memory in the daily hospital admissions, we consider two approaches. The first approach is to model the hospital admissions by the time-varying trend model (ref) and implement the KPSS and related tests, i.e., the \cref*{trend_algorithms} in the online supplement. The KPSS test yields p-values $0.004$, and other tests yield p-values smaller than $2 \times 10^{-4}$. All four tests reject the null hypothesis of short memory at the significance of $0.05$. The second approach is to model the hospital admissions via model (ref) taking into account three pollutants ($\text{SO}_2, \text{NO}_2$ and dust) and apply the KPSS and related tests, i.e., the (ref) and \cref*{algorithms} in the online supplement. The large $p$-values in the right panel of (ref) show that the KPSS and related tests fail to reject that the total number of hospital admissions is SRD at the significance level $0.05$. Although both of the two approaches are asymptotically correct, the misspecification of regression models tends to cause `spurious long memory' in finite samples. Therefore, our results conclude that the SRD assumption for (ref) adopted by fan2000simultaneous, zhou2010simultaneous, wu2018gradient and many others is reasonable.

Conclusion and future work

This paper develops bootstrap-assisted KPSS, R/S, K/S, and V/S-type nonparametric tests to detect long memory in time-varying coefficient linear models where the covariates and errors are allowed to be locally stationary and heteroscedastic. Under the null hypothesis, the fixed and local alternatives, we derive the limiting distributions of those test statistics and bootstrap statistics. In particular, we identify the conditions under which the test statistics degenerate. Although such conditions cannot be directly identified from real data in practice, our proposed bootstrap-assisted KPSS and related tests are always asymptotically correct regardless of such conditions. We also establish the theory of Gaussian approximation to the partial sum process of the product of non-stationary SRD and LRD time series, which is of separate interest and useful for a large class of problems in the analysis of (time-varying) linear models with LRD errors and SRD covariates.

A comprehensive Monte Carlo study supports that our proposed KPSS and related tests have good size and power performance in finite samples and are robust to the choices of smoothing parameters. The tests are applied to Hong Kong circulatory and respiratory data, recognizing `spurious long memory' due to the misspecification in the conditional mean. Therefore, our proposed methods can be employed as regression diagnostics. In \cref*{sec:covid} of the online supplement, we also apply the proposed tests to the COVID-19 data and identify that the time series of log cumulative confirmed cases of Japan and Ireland are LRD, while the time series of log cumulative confirmed deaths of Japan and Ireland are both SRD.

Recent studies have considered LRD models with time-varying long-memory parameter $d(t)$, see for instance Detectinglrdependence and ferreira2018estimation. Extra simulations in \cref*{tvd} of the online supplement evidence that our proposed testing procedures are still consistent against the alternatives with time-varying non-negative $d(t)$, i.e., $d(t)>0$ for $t$ in a sub-interval of $[0,1]$. The derivation of the theoretical behavior of the test statistics and the bootstrap procedure with time-varying $d(t)$ is left for rewarding future work. The extension of the proposed tests to $d \geq 1/2$ or $d<0$ for the null hypothesis (see also wu2006invariance, Ioa2021) is also challenging and meaningful.