Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
81,821 characters · 14 sections · 86 citation commands
Detecting long-range dependence for time-varying linear models
Keywords: Long-range dependence, Locally stationary process, Spurious long memory, Time-varying models
\footnotetext[1]{E-mail addresses: \href{[email removed]}{[email removed]}(L.Bai), \href{[email removed]}{[email removed]}(W.Wu)}
Consider the time-varying coefficient linear model
where the covariate vector $\mathbf x_{i,n}=(1,x_{i,2,n},...,x_{i,p,n})^\top$ is a $p$-dimensional {\it short-range dependent} (SRD) locally stationary time series and $y_{i,n}$ is the response variable, $t_i = i/n$. At each time point $t_i$, we only observe one realization $(\mathbf x_{i, n},y_{i,n})$ and no repeated measurement is available. The time-varying regression coefficient function $\boldsymbol \beta(\cdot)$ is a $p$-dimensional function with each coordinate a smooth function on $[0,1]$ and the zero mean error process $(e_{i,n})$ is a possibly {\it long-range dependent} (LRD) or {\it long-memory} time series. More precisely, we assume $(e_{i,n})$ is a locally stationary $I(d)$ process i.e., for $1\leq i\leq n$
where $\mathcal B$ is the lag operator, $d\in [0,1/2)$ is the {\it long-memory parameter} and $u_{i,n}$ is a SRD or {\it short-memory} locally stationary process. The strict definitions of locally stationary and long-memory processes are deferred to (ref). The error model (ref) naturally generalizes classical stationary SRD and LRD processes by allowing their generating mechanism to vary with time. Observe that $(e_{i,n})$ will reduce to the SRD process $(u_{i,n})$ if $d=0$ and will be a LRD process if $d \in (0, 1/2)$. In fact, when $(u_{i,n})$ is stationary, (ref) allows the classical stationary long-memory processes (e.g. FARIMA-GARCH models), which have found extensive application in hydrology (ZHANG2011121, koutsoyiannis2013hydrology), economics and finance (caporale2013long, unemployment2016) and many other fields since first introduced by hurst1951long. Moreover, model (ref) admits heteroscedasticity, i.e., the dependence of $u_{i,n}$ on $(\mathbf x_{r,n})_{r=1}^n$, see (ref) for more details.
The time-varying regression model (ref) with time series errors has attracted enormous attention, see for instance fan2000simultaneous, zhou2010simultaneous and chen2018jbes where the errors are assumed to be SRD, and kulik2012conditional, BeranLongmemory and ferreira2018estimation where LRD errors are considered. The aforementioned research reveals that nonparametric estimators of the time-varying coefficient ${\boldsymbol \beta}(\cdot)$ possess distinct properties under the two scenarios, $d=0$ and $0<d<1/2$. When $d=0$, consider the local linear estimator of the multivariate coefficient function $ {\boldsymbol \beta}(\cdot)$ using the kernel function $K(\cdot)$ and the bandwidth $b_n$, of which the asymptotic behavior rests on the distributions of $(\mathbf x_{i,n})$ and $(e_{i,n})$. In particular, the order of the deviation $|\hat {\boldsymbol \beta}(\cdot)-\boldsymbol \beta(\cdot)|$ is determined by the long-memory parameter $d$. For $d=0$ and $t\in (0,1)$, zhou2010simultaneous shows that under mild conditions,
where $\mu_2$ and $\phi_0$ are constants determined by $K(\cdot)$ and ${\boldsymbol \Sigma}(t)$ is determined by the moments of the process $(\mathbf x_{i,n}e_{i,n})_{i=1}^n$. Meanwhile, for $d>0$ and $p=1$, Theorem 7.22 in BeranLongmemory shows that for stationary $e_{i,n}$ under regularity conditions,
where $V(d) = 2c_f \Gamma(1-2d) \sin(\pi d) \int_{-1}^1\int_{-1}^1 K(x)K(y)|x-y|^{2d-1}dxdy$, and $c_f$ is a constant related to the spectral density of errors. Equation (ref) shows that for $d>0$ the convergence rate of $\hat {\boldsymbol \beta}(t)$ is $(nb_n)^{d-1/2}$, which is much slower than the well-known $(nb_n)^{-1/2}$ convergence rate as given by (ref) when $d=0$. Therefore, a crucial problem of the statistical inference of model (ref) is to test
The testing problem (ref) for (ref) is closely related to the existing tests of `spurious long memory', which refers to the phenomenon that in the presence of regime changes, level shifts or certain deterministic trends, a short memory process could exhibit many properties of a long-memory process, known as the `spurious long-memory' effects, see for example giraitis2001testing, qu2011test, and mccloskey2013memory. These findings motivate the tests for distinguishing genuine and spurious long memory. Among others, qu2011test, preu2013 and SIBBERTSEN201833 consider testing the null hypothesis of stationary long memory against spurious long memory. Meanwhile, several tests have been introduced to test the null hypothesis of spurious long memory, for which a prevailing approach is to assume a specific and parametric form of non-stationarity, see shao2006aos, harris2008testing, and davis2013consistency among others. Recently, there has been growing interest in detecting long memory in the presence of general non-stationarity, see for example Detectinglrdependence which considered locally stationary moving average formulation. In practice, by testing (ref) for model (ref), we are able to identify a new type of `spurious long memory' resulting from the {\it misspecification in the conditional mean}. See our data analysis in (ref) where we apply our method to the Hong Kong circulatory and respiratory data.
The goal of the present work is to test the {\it hypothesis (ref)} under complex and general temporal dynamics, assuming that $(\mathbf x_{i,n})$ and $(u_{i,n})$ belong to the flexible class of locally stationary processes generated by smoothly changing underlying mechanisms. Although some related literature has studied hypothesis (ref) for linear regression models with deterministic covariates, see for example harris2008testing, to the best of the authors' knowledge, this work is the first instance investigating testing (ref) for (ref) in the presence of time series covariates. In the literature, KPSS (Kwiatkowski, Phillips, Schmidt, and Shin, see lee1996power), R/S (range over standard deviation, see hurst1951long ), V/S (rescaled variance, see giraitis2003rescaled), and K/S (which has a limiting distribution of Kolmogoroff-Smirnoff form, see lima2004robustness) tests have been widely used for long memory detection in stationary processes. In this paper, we develop new KPSS, R/S, V/S and K/S-type tests tailored to the non-stationary time series time-varying regression problem (ref). The limiting distributions of the test statistics under the null hypothesis, the local and fixed alternatives are then derived. Our results differ from their stationary counterparts due to the following reasons. (1) Under the null hypothesis, it is well-known that the KPSS, R/S, V/S and K/S tests are built on the partial sum process whose convergence rate is $ n^{-1/2}$. However, the nonparametric estimate $\hat {\boldsymbol \beta}(\cdot)$ induces stochastic errors much larger than $n^{-1/2}$ which will lead to different limiting distributions as well as possible degeneracy. (2) Due to the non-stationary errors and covariates, the partial sum processes cannot be approximated by processes with stationary increments, which makes the test statistics non-pivotal. The major contributions of the paper lie in the following three aspects.
Firstly, our methods are applicable to the locally stationary time series regression, which has found considerable attention in various related fields, see for instance vogt2012, you2019, zhou2010simultaneous and many others. In particular, the flexible locally stationary framework allows the error processes to display conditional and unconditional heteroscedasticity that has been increasingly investigated (see harris2017adaptive and cavaliere2020adaptive) in the context of long-memory models. Both the evolving distributional properties of the locally stationary data and the long-memory properties pose long-standing challenges to the inference of time-varying coefficient linear model (ref) due to the lack of general Gaussian approximation techniques for non-stationary long-memory processes. For stationary long-memory processes, Gaussian approximation has been studied by for example dehling1989empirical. Recently, wu2018 developed Gaussian approximation schemes for a class of locally stationary long-memory linear processes. However, their results cannot accommodate regression problems with time series covariates which requires the analysis of distributional properties of the partial sum process of $(\mathbf x_{i,n} e_{i,n})_{i=1}^n$. In this paper, we address this issue via a further Gaussian approximation theorem, allowing $(\mathbf x_{i,n})$ and $(e_{i,n})$ to be non-stationary SRD and LRD processes, respectively, with flexible dependence between them.
Secondly, we develop effective bootstrap approaches which circumvent the difficult estimation of the non-pivotal limiting distributions of the test statistics under time series non-stationarity. In particular, the test statistics could degenerate when the time-varying mean, variance, long-run variance of errors and covariates, and the intercept lie in certain hyperplanes whose geometry cannot be directly identified from the data. Importantly, regardless of the degeneracy of test statistics, our bootstrap procedures are consistent and possess good finite sample properties. Furthermore, we show that the exact local asymptotic power of the four types of bootstrap-assisted tests can reach the order $O( \log^{-1} n)$ in the presence of time series covariates, no matter whether the test statistics degenerate under the null hypothesis. This rate coincides with shao2007local which studies a similar problem of testing the SRD null hypothesis against LRD local alternatives for strictly stationary series without covariates.
The rest of the paper is organized as follows. Section (ref) introduces necessary notation. Section (ref) formally states the non-stationary LRD model and the related assumptions. Section (ref) provides the test statistics, and establishes the asymptotic results via the new Gaussian approximation theory for the product of non-stationary SRD and LRD processes. Section (ref) discusses the bootstrap algorithms. Section (ref) reports the simulation results and the analysis of Hong Kong circulatory and respiratory data. Section (ref) provides a brief concluding remark. In (ref), we provide the proof of the new Gaussian approximation theory. Detailed proofs, the literature review of KPSS and related tests, the implementation details including the selection of tuning parameters, additional simulation results, data analysis results (including COVID-19 data), and additional algorithms are relegated to the supplement.
For a matrix $\mathbf A = (a_{ij})_{1\leq i \leq n, 1\leq j \leq m} \in \mathbb{R}^{n\times m}$, let $|\mathbf A| = (\sum_{j=1}^m \sum_{i=1}^n a_{ij}^2)^{1/2}$ and write $\mathbf A \geq 0$ if $\mathbf A$ is semi-positive definite. Let $(\mathbf A)_{(1,1)}$ denote the element of $\mathbf A$ in the first column and first row. Notice that when $m=1$, $\mathbf A$ is a vector. For $\mathbf A\geq 0$ with eigendecomposition $\mathbf A = \mathbf Q\mathbf D \mathbf Q^{\top}$ with orthonormal matrix $\mathbf Q$ and diagonal matrix $\mathbf D$, the root of $\mathbf A$ is defined by $\mathbf A^{1/2} = \mathbf Q \mathbf D^{1/2} \mathbf Q^{\top}$, where $\mathbf D^{1/2}$ is the elementwise root of $\mathbf D$. Let $\mathbf I_p$ denote the $p$-dimensional identity matrix. For a random matrix $\mathbf A$, for $q \geq 1$, let $\|\mathbf A \|_q = (\mathbb{E}|\mathbf A|^q)^{1/q}$ denote the $\mathcal{L}^q$-norm of the random variable $|\mathbf{A}|$ and write $\|\cdot\|=\|\cdot\|_2$ for short. Write $\mathbf A \in \mathcal{L}^q$ if $\|\mathbf A \|_q < \infty$. For a function $f(\cdot)$, write $f \in C[0,1]$ if $f$ is continuous over $[0,1]$, $f \in C^p[0,1]$ if the $p_{th}$ order derivative of $f$ is continuous over $[0,1]$. Write $\mathbf A \in C[0,1]$ and $\mathbf A \in C^p[0,1]$ if each element $A_{ij}(\cdot)$ in $\mathbf A(\cdot)$ is in $C[0,1]$ and $C^p[0,1]$, respectively. For any kernel function $K(\cdot)$, let $K^*(\cdot)$ denote the jackknife equivalent kernel $2 \sqrt{2} K(\sqrt{2}x) - K(x)$. Denote by $\lfloor x\rfloor$ the largest integer smaller or equal to $x$. For any two positive real sequences $a_n$ and $b_n$, write $a_n \asymp b_n$ if $\exists\ 0 <c < C<\infty$ such that $c<\liminf_{n \to \infty} \frac{a_n}{b_n}<\limsup_{n \to \infty} \frac{a_n}{b_n}<C$. Let $t \wedge s$ denote the smaller value in $t$ and $s$. Let `$\Rightarrow$' denote convergence in distribution. Write $\lambda$ as Lebesgue measure on $[0,1]$. Let `$:=$' denote `defined as'.
We start by introducing the time series time-varying regression model (ref) in detail. Recall model (ref) has the following form
We assume that the process $(u_{i,n})_{i=-\infty}^n$ and the covariate process $(\mathbf x_{i,n})_{i=1}^n$ have the form
where $\mathcal{F}_i=(\varepsilon_{-\infty},...,\varepsilon_i)$, $(\varepsilon_{i})_{i\in \mathbb Z}$ are $i.i.d.$ random variables, $H$ and $\mathbf W=(W_1,...,W_p)^\top$ are measurable functions such that $H:(-\infty,1]\times \mathbb R^{\mathbb Z}\rightarrow \mathbb R$, $W_s:[0,1]\times \mathbb R^{\mathbb Z}\rightarrow \mathbb R$, $2\leq s\leq p$, while $W_1$ is fixed to be $1$ corresponding to the intercept of the regression. Define $(\varepsilon^{\prime}_i)_{i \in \mathbb{Z}}$ as an $i.i.d.$ copy of $(\varepsilon_i)_{i \in \mathbb{Z}}$, and for $j \geq 0$, let $\mathcal{F}^{*}_j = (\mathcal{F}_{-1},\varepsilon^{\prime}_0, \varepsilon_1,\cdots, \varepsilon_{j-1}, \varepsilon_{j})$. For any (vector) process $\mathbf L(t,\mathcal{F}_i)$, it is said to be $\mathcal L^q$ stochastic Lipschitz continuous in the interval $\mathcal I$ (denoted by $\mathbf L\in \mathrm{Lip}_q(\mathcal I)$) if for $t_1,t_2\in \mathcal I$, there exists a constant $M>0$ such that
We say the process $\mathbf L(t,\mathcal{F}_i)$ is {\it locally stationary} (LS) on $\mathcal I$ if $\mathbf L(t,\mathcal{F}_i)\in \mathrm{Lip}_q(\mathcal I)$ for some $q\geq 2$. Write $\mathrm{Lip}_q=\mathrm{Lip}_q([0,1])$ for short. The locally stationary process offers a flexible nonparametric device to characterise the complex temporal dynamics of the error and covariate processes in (ref), which is based on Bernoulli shift processes and leads to a general framework of nonlinear processes, see wu2005nonlinear. Other formulations of locally stationary processes include dahlhaus1997 and nason2000wavelet. See dahlhaus2019bej for a comprehensive review. The physical dependence measure of the nonlinear filter $\mathbf L \in \mathcal{L}^q$ ($q>0$) over the interval $\mathcal I$ is defined by
The physical dependence measure $\delta_q(\mathbf L,k, \mathcal I)$ quantifies the influence of the input $\varepsilon_0$ on the output $\mathbf L(t,\mathcal{F}_k)$ over the interval $\mathcal I$. Observe that $\delta_q(\mathbf L,k, \mathcal I)=0$ if $k<0$. Write $\delta_q(\mathbf L,k)=\delta_q(\mathbf L,k, [0,1])$ for short. We proceed to define SRD and LRD non-stationary processes for non-stationary time series.
Definition (ref) distinguishes the SRD and LRD by the uniform summability of covariance, which naturally extends the traditional definition of long memory in second-order stationary univariate processes, see for example Condition III in Chapter 2 of pipiras2017long. The uniform long-memory definition has been introduced to define LRD and SRD non-stationary time series in wu2018. Various definitions of LRD stationary processes could be found in pipiras2017long and definitions of LRD locally stationary processes are discussed in BERAN2009900, Detectinglrdependence and ferreira2018estimation among others. In this paper, we posit the following assumptions.
Condition (ref) imposes the assumptions of local stationarity and finite forth moment on the innovations $(u_{i,n})$. Condition (ref) will be satisfied if $\sup_{t \in (-\infty, 1]} \|\frac{\partial}{\partial t}H(t,\mathcal{F}_0)\| <\infty$. Condition (ref) ensures that the innovations $(u_{i,n})$ are SRD and satisfy geometric measure contraction (GMC). Conditions (ref) and (ref) guarantee that the innovations $(u_{i,n})$ have a finite, non-degenerate and smooth long-run variance. Notice that under the null hypothesis $d=0$, the error process $(e_{i,n})$ reduces to $(u_{i,n})$, which indicates that $(e_{i,n})$ is a SRD process. When $d > 0$, $(e_{i,n})$ is generated by a binomial weighted combination of $u_{i,n}=H(t_i,\mathcal{F}_i)$ starting from the infinite past ($i=-\infty$). To stress that $(e_{i,n})$ is a LRD process under the alternative hypothesis with respect to $d$, in the remaining of this article we write $e_{i,n}$ as $e_{i,n}^{(d)}$ when $d>0$, i.e., $e_{i,n}^{(d)} = (1-B)^{-d} u_{i,n} = \sum_{k=0}^{\infty} \psi_k(d) u_{i-k}$, $\psi_j(d) = \Gamma (j+d) / [\Gamma(d) \Gamma (j+1)]$. We further write $e_{i,n}^{(d)} = H^{(d)}(t_i, \mathcal{F}_{i})$, where $ H^{(d)}(t, \mathcal{F}_{l}) = \sum_{k=0}^{\infty} \psi_k(d) H(t - t_k, \mathcal{F}_{l-k})$. The following Proposition (ref) elaborates that the physical dependence measure of $(e^{(d)}_{i,n})$ relies on $d$.
Our formulation of $(e_{i,n})$ in (ref) allows for a wide class of non-stationary SRD and LRD processes under $H_0$ and $H_A$, respectively, including the following examples.
Since $(e_{i,n})$ is not observable in (ref), we propose to test $H_0$ based on nonparametric residuals. Specifically, we adopt the local linear approach (see for instance fan1993local and fan1996local) to estimate ${\boldsymbol \beta}(t)$ in (ref), i.e.,
where $K(\cdot)$ is a kernel function with finite support $[-1,1]$ and $b_n$ is a bandwidth and $K_{b_n}(\cdot)=K(\cdot/b_n)$. To further eliminate the bias term involving $\boldsymbol \beta^{\prime \prime }(\cdot)$, we use the jackknife bias-corrected estimator in wu2007inference :
Then, we obtain the nonparametric residuals $(\tilde e_{i,n})$, i.e., $ \tilde e_{i,n} = y_{i,n} - \mathbf x_{i,n}^{\top} \tilde{\boldsymbol \beta}(t_i).\nonumber $ For simplicity, define $\tilde S_{r,n} =\sum_{i=\lfloor nb_n \rfloor+1}^r \tilde{e}_{i,n}$, $r=\lfloor nb_n \rfloor+1,\cdots n-\lfloor nb_n\rfloor$. We consider four well-known types of partial sum based test statistics, which are KPSS, R/S, V/S and K/S-type tests built on $(\tilde e_{i,n})$.
The above four types of tests have been widely applied to the detection of long memory and many other important problems (e.g. unit root testing) for stationary time series. We refer to Section (ref) of the online supplement for the complete literature review and applications. To the best of our knowledge, all the existing work on the KPSS, R/S, V/S, and K/S tests considers the statistics based on the original series $(e_{i,n})$ or parametric residuals (e.g., the residuals obtained by the removal of the sample mean), and is therefore not applicable to the time-varying coefficient model (ref). Meanwhile, it is well-known that the nonparametric estimators have a slower convergence rate than the corresponding parametric estimators. Therefore, the asymptotic properties of our nonparametric residual-based KPSS and related tests will be very different from their parametric residual-based or original series-based counterparts; in fact we show that the tests can degenerate under certain scenarios (see Theorem (ref)) and thus bootstrap procedures adaptive to the possible degeneracy are proposed for implementation, see Algorithms (ref), also Algorithms \ref*{trend_algorithms} and \ref*{algorithms} of the online supplement. In this article, we use the term `{\it KPSS and related tests}' to represent the four types of tests.
For the sake of brevity, in this section we only discuss the KPSS-type test statistic in detail, and summarize the results of KPSS-related test statistics in Remark (ref). In order to investigate the asymptotic properties of $T_n$ defined by (ref) in the presence of time series covariates, we introduce the following assumptions.
Assumption (ref) is standard for local linear time series regression, see for instance zhou2010simultaneous. Notice that the first element of $\mathbf U(t, \mathcal{F}_i)$ is $H(t, \mathcal{F}_i)$. When $p=1$ (the time-varying trend model), Assumption (ref) reduces to (ref) in Assumption (ref) with $t \in [0,1]$.
Define $\mathbf M(t) := \mathbb{E} (\mathbf{W}(t,\mathcal{F}_0)\mathbf{W}(t,\mathcal{F}_0)^{\top})$, $t \in [0,1]$. Notice that the first element of $\mathbf W$ is $1$. Write $\boldsymbol{ \mu}_W(\cdot) := (1, \mu_{W,2}(\cdot), \cdots, \mu_{W,p}(\cdot))^{\top} :=\mathbb{E}(\mathbf W(t,\mathcal{F}_0))$. For $p \geq 2$, Let $x_{i,j, n}$ denote the $j$th element in $\mathbf x_{i,n}$, and $\mathbf x_{i,n}^{(-1)} := (x_{i,2,n}, \cdots, x_{i,p,n})^{\top}$. Define $\boldsymbol \mu_W^{(-1)}(\cdot) := (\mu_{W, 2}(\cdot), \cdots, \mu_{W,p}(\cdot))^{\top}$, $\mathbf W^{(-1)}(\cdot, \cdot) := (W_2(\cdot, \cdot),\cdots, W_p(\cdot, \cdot))^\top$.
Condition (ref) ensures that there is no multicolinearity among the explanatory variables. Assumption (ref) guarantees that the $\mathbf M(\cdot)$ and $\boldsymbol \mu_W(\cdot)$ have continuous derivatives. Assumption (ref) requires that the covariates are locally stationary. Condition (ref) imposes that each component of $\mathbf W^{(-1)}(\cdot, \cdot)$ is SRD. Condition (ref) assumes that $\mathbf x_{i,n}$ is a $p$-dimensional random vector uncorrelated with innovations, which is necessary for model identification. Our assumptions are very mild in the sense that we allow nonlinearity and heteroscedasticity for the covariates and errors, as well as the correlation between $\mathbf x_{i,n}$ and $e_{i,n}$. Assumptions (ref) and (ref) can be verified using similar arguments in wu2009quantile. When $p = 1$, we use the convention that $\mathbf x_{i,n}^{(-1)} = \emptyset$, $\mathbf W^{(-1)}(\cdot, \cdot) = \emptyset$, and $\boldsymbol \mu_W^{(-1)}(\cdot) = \emptyset$, $\mathbf M(\cdot)=\boldsymbol \mu_W(\cdot)=1$. Therefore Assumption (ref) always hold in this case.
The test statistics will degenerate under the null hypothesis when the parameters of the regression model (ref) lie in the hyperplane $\{u\in[0,1]: \boldsymbol \Sigma^{1/2}(u) \mathbf{M}^{-1}(u)\boldsymbol \mu^{\top}_W(u)= (\sigma_H(u), 0, \cdots, 0)^{\top}\}$. (ref) excludes such situation.
The following theorem establishes the asymptotic distribution of the KPSS-type statistic $T_n$ (ref) under the null hypothesis.
Theorem (ref) reveals that for the time-varying coefficient model with time series covariates, the limiting distribution of $T_n$ depends on the time-varying mean and covariance matrix of the covariates $\mathbf x_{i,n}$ as well as the long-run covariance matrix of $\mathbf x_{i,n} e_{i,n}$. Theorem (ref) is very general since it posits neither the specific form of heteroscedasticity nor the parametric form of the errors. When Assumption (ref) is violated, (ii) shows that $T_n$ degenerates with asymptotic variance $2s_1^2 b_n^2$. After standardization, it converges to $\chi^2_1$ in distribution. The results of R/S, V/S and K/S follow similarly. An important scenario that $T_n$ degenerates under $H_0$ (i.e., $d=0$) is the following time-varying trend model corresponding to $p=1$, i.e.,
Since the KPSS and related test statistics are constructed based on the partial sum process, to derive the asymptotic properties of the test statistics under the alternative hypothesis, we first study the Gaussian approximation of $\sum_{i=1}^r \mathbf x_{i,n} e_{i,n}^{(d)}$, $1 \leq r \leq n$, which is the partial sum of the product of a SRD and a LRD time series under $H_A$. Though Gaussian approximation theory for stationary processes ( see for instance wu2007strong, dehling1989empirical, wu2006invariance and the reference therein) has been successfully established and widely applied to many fields of statistics, there are only a few results of Gaussian approximation for locally stationary processes. Among them, wu2011gaussian established a flexible Gaussian approximation framework for locally stationary SRD processes, which has served as a fundamental key to the inference of SRD (piecewise) locally stationary processes and functional time series, see for instance chen2015function and wu2018gradient. wu2018 proposed a Gaussian approximation scheme for a class of locally stationary linear LRD processes. However, all the existing Gaussian approximation approaches are not applicable to the partial sum process of the product series $ (\mathbf x_{i,n} e_{i,n}^{(d)})$, which serves as the crucial ingredient for establishing the limiting distribution of $T_n$ under $H_A$. To this end, we shall provide a general Gaussian approximation theorem for the product of LRD and SRD processes. In the remaining of this paper, let $ d_n = c/ \log n$, where $c$ is a positive constant. We substitute $d$ with $d_n$ to differentiate the notation under the fixed alternatives $(d>0)$ and that under the local alternatives $(d=c/\log n)$.
Since $\|{\mathbf R}_{i,n}\| \asymp n^{d+1/2}$ and $\|\tilde {\mathbf R}_{i,n}\| \asymp n^{1/2}$, the approximation errors of Theorem (ref) (i) and (ii) are asymptotically negligible. It can be also verified that the process $(\mathbf R_{k,n})_{k=1}^n$ is a locally stationary LRD Gaussian process defined by Definition (ref).
It is worth pointing out that (ii) is not an direct consequence of (i). Letting $d=d_n$ in (i), the rate $\sqrt n(\log n)^d$ in (i) will eventually lead to a trivial bound, i.e., $O_{\mathbb{P}}(\sqrt{n})$, which is of the same order as the partial sum $\sum_{i = \lfloor nb_n \rfloor +1}^r \mathbf x_{i,n} e^{(d_n)}_{i,n}$, $\lfloor nb_n \rfloor +1\leq r\leq n-\lfloor nb_n \rfloor $.
Based on the Gaussian approximation result, we proceed to study the limiting distributions of KPSS and related statistics under the fixed and local alternatives for the time-varying coefficient model (ref) with time series covariates. For the sake of brevity, we focus on the KPSS-type statistics. The results for R/S, V/S, and K/S-type statistics can be derived similarly and are summarized in Remark (ref) and \cref*{sub:limits} in the online supplement.
In the following theorem, we establish the asymptotic distribution of the KPSS-type statistic $T_n$ (ref) under the fixed alternatives.
Theorem (ref) proves that the test statistic diverges to infinity at the rate of $n^{2d}$. Furthermore, from Theorem (ref), we observe that the limiting distribution of $T_n$ under the fixed alternatives is independent of ${\boldsymbol \Sigma}(t)$ except $\sigma^2_H(t)$, which is its $(1,1)$ component, while under the null hypothesis the limiting distribution relies on all the components of ${\boldsymbol \Sigma}(t)$, see Theorem (ref). This is because when $d>0$, the stochastic fluctuation of the SRD components $(\mathbf x_{i,n})$ is asymptotic negligible compared to that of $(e^{(d)}_{i,n})$.
Straightforward calculation shows that $\check M_W(t)\geq 1$. The condition $\lambda(|\boldsymbol \mu_W^{(-1)}(\cdot))| \neq 0) > 0$ ensures $\check M_W(t)>1$ in an interval with positive length such that $\lambda_d(u,v)>0$ and excludes the scenario that all the stochastic covariates are zero mean during the whole period (see (ref) for detailed discussion of such scenario) as well as the time-varying trend model,i.e., (ref) with $p=1$. When $\lambda_d(u,v)=0$, we can show that $T_n=o_{\mathbb{P}}(n^{2d})$, i.e., the test statistic is degenerate under $H_A$.
The following theorem presents the asymptotic distribution of the KPSS-type statistic $T_n$ (ref) under the local alternatives.
From (ref), we shall see that under the local alternatives $d_n=c/\log n$, the KPSS-type statistic converges to a distribution depending on the mean and covariance matrix of $(\mathbf x_{i,n})$, the long-run covariance matrix of $(\mathbf x_{i,n}e_{i,n})$, as well as the parameter $c$. Careful examination of the proof of Theorem (ref) shows that $T_n$ will converge to the limit in Theorem (ref) if $d_n\log n=o(1)$, indicating that the exact local power of the KPSS-type test is $O(\log^{-1} n)$ for model (ref). For the time-varying trend model (ref), it can be shown that the test statistic is degenerate (since $\lambda(|\boldsymbol \mu_W^{(-1)}(\cdot))| \neq 0) =0$) under local alternatives $d_n = c/\log n$, i.e., $T_n=o_{\mathbb{P}}(1)$. The results for R/S, K/S and V/S-type tests follow similarly.
(ref) shows that under the null hypothesis, the limiting distributions of KPSS and related test statistics are functions of the Gaussian process $U(t)$ which involves parameters $\boldsymbol \mu_W(t), \mathbf M(t), \boldsymbol \Sigma(t)$ (or $\sigma^2_H(t)$ when $p=1$). Furthermore, the magnitude and specific form of the limiting distributions depend on whether (ref) holds, which is usually unknown in practice. Therefore, it's impossible to obtain the critical values by directly simulating the Gaussian process $U(t)$. In this section, we provide a consistent bootstrap approach (ref) which mimics the asymptotic behavior of $T_n$ under the null hypothesis no matter whether (ref) is satisfied and yields valid simulated critical values. In addition, the estimation of $\boldsymbol \mu_W(\cdot)$ is not required in our proposed bootstrap tests. More precisely, we employ $\mathbf x_{i,n}^\top \hat{\mathbf M}^{-1}(t)$ instead of $\hat{\boldsymbol \mu}^\top_W(t) \hat{\mathbf M}^{-1}(t)$, where $\hat{\boldsymbol \mu}_W(t)$ stands for any consistent estimator of $\boldsymbol \mu_W(t)$, and the validity of the former construction can be easily verified by noting that in (ref) of (ref), the Gaussian multiplier is independent of $ \mathbf x_{i,n}$, $\hat {\mathbf M}(\cdot)$ and $\hat {\boldsymbol \Sigma}(\cdot)$ where the consistent estimators $\hat{\mathbf M}(t)$ and $\hat {\boldsymbol \Sigma}(t)$ will be discussed later. Thus, the difference between (ref) and the counterpart with $\boldsymbol \mu_W(t)$ can be controlled by the convolution of standard Gaussian multipliers and the partial sum of the zero mean SRD process containing $(\mathbf x_{i,n}-\boldsymbol \mu_W(t_i))$. We only discuss (ref) for the KPSS-type test when $p\geq 2$ in detail in this section. The other algorithms, including \cref*{trend_algorithms} in the online supplement for $p=1$ case corresponding to the time-varying trend model (ref), and \cref*{algorithms} for R/S, V/S, K/S-type tests are moved to the supplement.
To implement (ref), we need to obtain the estimators $\hat {\mathbf M}(t)$ and $\hat {\boldsymbol \Sigma}(t)$. For $\hat {\mathbf M}(t)$ we propose the following estimator that uses directly observed covariates $\mathbf x_{i,n}$:
where $t^* = \max\{\eta_n,\min(t,1-\eta_n)\}$ for some bandwidth $\eta_n \to 0$, $n\eta_n^2 \to \infty$. Under Assumptions (ref) and (ref), after a careful investigation of Lemma 6 of zhou2010simultaneous, we have $\sup_{t \in[\eta_n,1-\eta_n]}|\hat{\mathbf{M}}(t)- \mathbf M(t)|=o_{\mathbb{P}}(1)$, i.e. $\hat{\mathbf{M}}(t)$ is uniformly consistent, see \cref*{lm:scblemma6} in the online supplement for details.
Observing that $\boldsymbol \Sigma(\cdot)$ depends on the unobserved error process $(e_{i,n})$. Therefore, a common approach to estimate $\boldsymbol \Sigma(\cdot)$ is to utilize $\hat e_{i,n}'s$ from the local linear fit (ref), see for instance zhou2010simultaneous and vogt2015detecting. Such plug-in estimate will yield test results that are sensitive to the choices of $b_n$ under $H_0$, and even more sensitive under $H_A$. We also find through our extensive numerical studies (which are not reported in this paper due to limited space) that the power of the test using the plug-in estimator of long-run covariance is often unsatisfactory. Therefore, we adopt a difference-based estimator that does not involves $\hat e_{i,n}'s$.
For $p=1$ we recommend the difference statistic proposed in (4.7) of Section 4.2 of dette2019detecting for $\sigma^2_H(t)$ which is built on the difference of $y_{i,n}$. For $p \geq 2$, it can be shown that a direct extension of the estimator based on the difference of $(\mathbf x_{i,n}y_{i,n}$) is asymptotically biased. Therefore, we adopt the following bias-corrected difference-based estimator proposed by lrvunpublish. Let $\mathbf Q_{k, m}=\sum_{i=k}^{k+m-1} \mathbf x_{i,n}y_{i,n}$, for $t \in [m/n,1-m/n]$,
where $ \omega(t, i)= K_{\tau_n}\left(t_i-t\right) / \sum_{i=1}^{n} K_{\tau_{n}} \left(t_i-t\right) $ for some bandwidth $\tau_n$ and the kernel function $K(\cdot)$ with support $(-1,1)$. For $t \in [0,m/n)$, set $\acute{{\boldsymbol \Sigma}}(t) = \acute{{\boldsymbol \Sigma}}(m/n)$. For $t \in (1-m/n,1]$, set $\acute{{\boldsymbol \Sigma}}(t) = \acute{{\boldsymbol \Sigma}}(1-m/n)$. The bias-corrected difference-based estimator $\hat {\boldsymbol \Sigma}(t)$ for $t \in [0,1]$ is then defined as:
where $ \hat{\mathbf A}_{j} = \frac{1}{m}\sum_{i= j-m+1}^{j} (\mathbf x_{i,n} \mathbf x_{i,n}^{\top}\breve {\boldsymbol \beta}(t_i)-\mathbf x_{i+m, n} \mathbf x_{i+m, n}^{\top}\breve {\boldsymbol \beta}(t_{i+m}))$, $ \breve {\boldsymbol \beta}(t) = \boldsymbol \Omega^{-1}(t)\boldsymbol \varpi (t), $ where $\boldsymbol\Omega(t),\boldsymbol \varpi (t)$ are the smoothed versions of $\acute{\boldsymbol \Delta}_{j} := \frac{1}{m}\sum_{i=j-m+1}^{j} \tilde{\mathbf X}_{i,m}\tilde{\mathbf X}_{i,m}^{\top}$ and $\breve {\boldsymbol \Delta}_{j} := \frac{1}{m}\sum_{i=j-m+1}^{j} \tilde{\mathbf X}_{i,m}^{\top}\tilde{\mathbf Y}_{i,m}$, i.e., $ \boldsymbol\Omega(t) = \sum_{j=m}^{n-m} \acute {\boldsymbol \Delta}_{j}\omega(t^*, j)/2,\nonumber$ and $\boldsymbol \varpi (t) = \sum_{j=m}^{n-m} \breve {\boldsymbol \Delta}_{j} \omega(t^*, j)/2. $ To make (ref) operational, it's necessary to select smoothing parameters $\eta_n$, $\tau_n$ and $m$ for $\hat {\mathbf M}(t)$ and $\hat {\boldsymbol \Sigma}(t)$. The selection of smoothing parameters is postponed to \cref*{sel} of the online supplement.
In this section, we shall show the asymptotic correctness of the bootstrap test (ref). We shall also prove that the power of (ref) against $d > 0$ can be no less than $(n/m)^{2d}$, and that the exact local power of (ref) can achieve the order $O(\log^{-1} n)$. We start by defining the long-run {\it cross} covariance vector between the locally stationary processes $ (\mathbf U(t,\mathcal{F}_i))$ and $(H(t,\mathcal{F}_j))$.
When $p=1$, $\mathbf s_{UH}(t)$ degenerates into $\sigma_H^2(t)$. We assume the following conditions for $\hat{\boldsymbol \Sigma}(t)$ used in (ref). Write $\hat{\boldsymbol \Sigma}(t)$ defined in (ref) as $\hat{\boldsymbol \Sigma}_d(t)$ under the fixed alternatives and $\hat{\boldsymbol \Sigma}_{d_n}(t)$ under the local alternatives $d_n$. Let $\mathcal{I} = [\gamma_n, 1-\gamma_n] \subset (0,1)$, where $\gamma_{n}=\tau_{n}+(m+1) / n$.
It can be shown that the both the plug-in estimator of zhou2010simultaneous and the bias-corrected estimator (ref) satisfies this condition under suitable bandwidth conditions following Theorem 3.1 of wu2006invariance and the chaining argument of Propostion B.1 of dette2018change. Let $\tilde T_n$ denote the bootstrap statistic (ref) generated in one iteration. Recall the definitions of $U(t)$, $s_1$, $s_2$ in (ref). (ref) gives the limiting distributions of bootstrap statistic $\tilde T_n$.
Combining with (ref), (ref) indicates that the bootstrap test (ref) is asymptotically of level $\alpha$ {\it no matter whether (ref) is satisfied}. We proceed to investigate the behavior of the bootstrap statistic $\tilde T_n$ under fixed and local alternatives. Let $\tilde U_d(t)$, $\check U(t)$ be zero mean continuous Gaussian processes, with the covariance structures defined in the same way as that of $U(t)$ in (ref), where $\boldsymbol \Sigma(t)$ is replaced by $\boldsymbol \Sigma_d(t)$ and $\breve{\boldsymbol \Sigma}(t)$ in (ref), respectively. Let $\sigma_{Hd}^2(t) := (\boldsymbol \Sigma_d(t))_{(1,1)}$, $\check \sigma_{H}^2(t) := (\check{\boldsymbol \Sigma}(t))_{(1,1)}.$
(ref) (i) gives the limiting distribution of $\tilde T_{n}/ m^{2d}$, which means the critical values generated by (ref) for a level $\alpha$ test diverges at the rate of $m^{2d}$. Thus, the bootstrap-assisted test is consistent since (ref) demonstrates that the KPSS-type test statistic $T_n$ diverges at the rate $n^{2d}$ which is much faster than $m^{2d}$. Further together with (ref), (ref) (i) shows that the bootstrap test (ref) is asymptotically correct. Notice that the condition $\lambda({\boldsymbol \Sigma}_d(t)\mathbf{M}^{-1}(t) \boldsymbol \mu_W(t) \neq (\sigma_{Hd}(t), 0, \cdots, 0)) > 0$ prevents the degeneracy of bootstrap statistics. If it is violated, $\tilde T_n= o(m^{2d})$ which will yield higher power than when the condition is fulfilled.
On the other hand, (ref) (ii) and (ref) indicate that the bootstrap test (ref) is able to detect the local alternatives at the rate of $\log^{-1} n$. Observe that under (ref) (iii) as $c\rightarrow 0$, $|\check{\boldsymbol \Sigma}(t)-\boldsymbol \Sigma(t)|\rightarrow 0$ and the covariance structure of $\check{U}(t)$ will converge to the covariance structure of $U(t)$. Therefore, the bootstrap test (ref) has no power when $d_n=o( \log^{-1} n)$, indicating that the proposed test has the exact local power of $O(\log^{-1} n)$ under the condition of (ref). For stationary time series with unknown constant mean, shao2007local has also proved that the KPSS test for long memory has the exact local power $O(\log^{-1} n)$.
An important scenario that $T_n$ degenerates under both null and alternatives is the time-varying trend model (ref). For this case we provide the bootstrap-assisted KPSS test in \cref*{trend_algorithms} of the online supplement, which is the $p=1$ version of (ref) with $\hat {\mathbf M}(t) = 1$, and $\hat{\boldsymbol \Sigma}(t) = \hat \sigma_H^2(t)$ using the difference statistic proposed by dette2019detecting. (ref) (ii) ensures that the level $\alpha$ critical value generated by \cref*{trend_algorithms} converges to the $\alpha_{th}$ quantile of the test statistic $T_n$ under the null hypothesis. Under the alternative hypothesis, the following proposition investigates the power of the test implemented via \cref*{trend_algorithms}.
In addition to the KPSS test, (ref) also provides the bootstrap-assisted V/S, R/S and K/S tests for the time-varying trend model (ref). Similar conclusions as given in (ref) hold for the power of these tests.
In the following simulation studies and data analysis, we examine the size and power performance of the bootstrap-assisted KPSS and related tests (ref), \cref*{trend_algorithms} and \cref*{algorithms} in the online supplement. The number of bootstrap samples is $B = 2000$ and the number of replications is $1000$. The parameters $\mathbf M(t)$, ${\boldsymbol \Sigma}(t)$, $\sigma^2_H(t)$ are estimated by $\hat{\mathbf M}(t)$, $\hat {\boldsymbol \Sigma}(t)$, $\hat \sigma^2_H(t)$ in (ref), with all the smoothing parameters selected by the methods advocated in \cref*{sel} in the online supplement. Let
where $(\varepsilon_{l})_{l\in \mathbb Z}, (\zeta_{l})_{l\in \mathbb Z}$ are $i.i.d.$ $ N(0,1)$. We consider the following time-varying coefficient model, $$ y_{i, n}=\beta_{1}(t_i)+\beta_{2}(t_i) x_{i, n}+e_{i,n}, \quad i=1, \ldots, n, $$ where $\beta_{1}(t)=4 \sin (\pi t) $, $\beta_{2}(t)=4 \exp \{-2 \left(t-0.5\right)^{2}\} $, $x_{i, n}=W(t_i, \mathcal{F}_{i})$, and $ u_{j,n} = H(t_j,\mathcal{F}_j, \mathcal{G}_j)$. First, we consider the following independent model:
Second, we consider the following heteroscedastic model:$$ H(t,\mathcal{F}_{i},\mathcal{G}_{i}) = B\left(t,\mathcal{G}_{i}\right)\sqrt{1+W^2(t, \mathcal{F}_{i})},$$ where $ W\left(t, \mathcal{F}_{i}\right)= (0.1 + 0.1\cos(2\pi t))W(t, \mathcal{F}_{i-1})+ 0.2\zeta_{i} + 0.7(t-0.5)^2, $ and $B\left(t,\mathcal{G}_{i}\right)$ is as considered in the following linear and nonlinear scenarios.
Observe that models (ref) and (ref) are heteroscedastic models with locally stationary AR(1) and locally stationary GARCH(1,1) errors, respectively. (ref) summarizes the performance of our proposed bootstrap-assisted KPSS, R/S, V/S and K/S-type tests for long memory in models (ref) and (ref) with different $b_n's$. We relegate the simulated sizes of model (ref) with different $b_n's$ to \cref*{tb:m0size1} of the online supplement. The empirical sizes of all the four tests are close to their nominal levels and are quite stable when $b_n$ changes within a reasonably wide range. Also, \cref*{tb:Tn} in the online supplement reports the simulated Type I error of the proposed tests with respect to increasing sample sizes. As shown in \cref*{tb:Tn}, our procedures for smoothing parameter selection including GCV and MV selection as well as the difference-based long-run variance estimator work very well in the sense that the simulated sizes of all four tests are quite close to their nominal levels in different sample sizes.
Recall the long-memory error process of (ref), which can be written as $e_{i,n}^{(d)} = (1-\mathcal{B})^{-d}u_{i,n}$ where $\mathcal{B}$ is the lag operator. (ref) displays the power performance of the KPSS and related tests for model (ref) with nominal level $0.1$. The left panel reports simulated rejection rates of KPSS, R/S, V/S, and K/S-type tests as the long memory parameter $d$ increases from $0$ to $0.5$ with sample size $1500$. The power of all four KPSS and related tests increases to 1 as $d$ approaches to $1/2$, while the V/S-type test remains the most powerful among all tests. The right panel depicts the power performance of the four tests as the sample size grows from $200$ to $2500$ when $d = 0.4$. The figure shows that the power of each test increases to $1$ as the sample size grows, and the V/S-type test performs the best among the four tests whenever the sample size is larger than $400$. The power performance of models (ref) and (ref) are shown in \cref*{fig:power1} and \cref*{fig:power3} of the online supplement where the V/S-type test achieves the best power under most alternatives and sample sizes considered.
Hong Kong circulatory and respiratory data contains daily measurements of pollutants and daily hospital admissions in Hong Kong between January 1st, 1994 and December 31st, 1995. The dataset has been studied by fan2000simultaneous, zhou2010simultaneous, and wu2018gradient among others. They investigated the relationship between the levels of pollutants and the total number of hospital admissions of circulation and respiration,see \cref*{fig:hk_data} in \cref*{hk} of the supplement for the observed series. Under the assumption of $i.i.d.$ observations, fan2000simultaneous claimed that sulphur dioxide ($\text{SO}_2$) is not significant. Using the non-stationary model, zhou2010simultaneous found all three pollutants ($\text{SO}_2$, nitrogen dioxide ($\text{NO}_2$) and dust) are significant. In both cases, they assumed the observations were SRD. We shall examine this assumption via the KPSS and related tests. Consider the following time-varying coefficient model,
where $(y_{i,n})$ is the series of daily total number of hospital admissions of circulation and respiration and $(x_{i, p, n})$, $p = 2, 3, 4$, represent the series of daily levels (in micrograms per cubic meter) of $\text{SO}_2$, $\text{NO}_2$ and dust, respectively. The sample size is $n = 2 \times 365 = 730$. We first investigate whether the covariates $(x_{i, p, n})$, $p = 2, 3, 4$ are SRD series.
We select the smoothing parameters $b_n$, $\eta_n$, $m$ and $\tau_n$ through the methods provided in \cref*{sel} in the online supplement and summarize the selected parameters in \cref*{tb:hk_par} in \cref*{hk} of the online supplement. We test for long memory in the three pollutants series via KPSS and related tests, presenting the $p$-values in (ref). For each pollutant series we fail to reject it is SRD at the significance level $0.05$.
To test for long memory in the daily hospital admissions, we consider two approaches. The first approach is to model the hospital admissions by the time-varying trend model (ref) and implement the KPSS and related tests, i.e., the \cref*{trend_algorithms} in the online supplement. The KPSS test yields p-values $0.004$, and other tests yield p-values smaller than $2 \times 10^{-4}$. All four tests reject the null hypothesis of short memory at the significance of $0.05$. The second approach is to model the hospital admissions via model (ref) taking into account three pollutants ($\text{SO}_2, \text{NO}_2$ and dust) and apply the KPSS and related tests, i.e., the (ref) and \cref*{algorithms} in the online supplement. The large $p$-values in the right panel of (ref) show that the KPSS and related tests fail to reject that the total number of hospital admissions is SRD at the significance level $0.05$. Although both of the two approaches are asymptotically correct, the misspecification of regression models tends to cause `spurious long memory' in finite samples. Therefore, our results conclude that the SRD assumption for (ref) adopted by fan2000simultaneous, zhou2010simultaneous, wu2018gradient and many others is reasonable.
This paper develops bootstrap-assisted KPSS, R/S, K/S, and V/S-type nonparametric tests to detect long memory in time-varying coefficient linear models where the covariates and errors are allowed to be locally stationary and heteroscedastic. Under the null hypothesis, the fixed and local alternatives, we derive the limiting distributions of those test statistics and bootstrap statistics. In particular, we identify the conditions under which the test statistics degenerate. Although such conditions cannot be directly identified from real data in practice, our proposed bootstrap-assisted KPSS and related tests are always asymptotically correct regardless of such conditions. We also establish the theory of Gaussian approximation to the partial sum process of the product of non-stationary SRD and LRD time series, which is of separate interest and useful for a large class of problems in the analysis of (time-varying) linear models with LRD errors and SRD covariates.
A comprehensive Monte Carlo study supports that our proposed KPSS and related tests have good size and power performance in finite samples and are robust to the choices of smoothing parameters. The tests are applied to Hong Kong circulatory and respiratory data, recognizing `spurious long memory' due to the misspecification in the conditional mean. Therefore, our proposed methods can be employed as regression diagnostics. In \cref*{sec:covid} of the online supplement, we also apply the proposed tests to the COVID-19 data and identify that the time series of log cumulative confirmed cases of Japan and Ireland are LRD, while the time series of log cumulative confirmed deaths of Japan and Ireland are both SRD.
Recent studies have considered LRD models with time-varying long-memory parameter $d(t)$, see for instance Detectinglrdependence and ferreira2018estimation. Extra simulations in \cref*{tvd} of the online supplement evidence that our proposed testing procedures are still consistent against the alternatives with time-varying non-negative $d(t)$, i.e., $d(t)>0$ for $t$ in a sub-interval of $[0,1]$. The derivation of the theoretical behavior of the test statistics and the bootstrap procedure with time-varying $d(t)$ is left for rewarding future work. The extension of the proposed tests to $d \geq 1/2$ or $d<0$ for the null hypothesis (see also wu2006invariance, Ioa2021) is also challenging and meaningful.