Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,175 characters · 5 sections · 97 citation commands
\centerline{Yeonwoo Rho\footnote{Yeonwoo Rho is Assistant Professor in Statistics at the Department of Mathematical Sciences, Michigan Technological University, Houghton, MI 49931. (Email: [email removed])} and Xiaofeng Shao\footnote{ Xiaofeng Shao is Professor at the Department of Statistics, University of Illinois, at Urbana-Champaign, Champaign, IL 61820 (Email: [email removed]). }} \centerline {\it Michigan Technological University$^1$} \centerline {\it University of Illinois at Urbana-Champaign$^2$} \centerline{\today}
Unit root testing has received a lot of attention in econometrics since the seminal work by Dickey:Fuller:1979, Dickey:Fuller:1981. In their papers, unit root tests were developed under the assumption of independent and identically distributed (i.i.d.) Gaussian errors. Many variants of the Dickey-Fuller test have been proposed, when the error processes are stationary, weakly dependent, and free of the Gaussian assumption. Most of the variants rely on two fundamental approaches to accommodate weak dependence in the error. One approach is the Phillips-Perron test Phillips:1987a, Phillips:Perron:1988, where the longrun variance of the error process is consistently estimated in a nonparametric way, using heteroscedasticity and autocorrelation consistent estimators Newey:West:1987, Andrews:1991. The other approach is the augmented Dickey-Fuller test Said:Dickey:1984, which approximates dependence structure in error processes with an AR($p$) model, where $p$ can grow with respect to the sample size. In addition to these two popular methods and their variants, bootstrap-based tests were also proposed by Paparoditis:Politis:2002, Paparoditis:Politis:2003, Chang:Park:2003, Parker:Paparoditis:Politis:2006, and Cavaliere:Taylor:2009a, among others. For reviews and comparisons of some of these bootstrap-based methods, we refer to Paparoditis:Politis:2005 and Palm:Smeekes:Urbain:2008.
Recently, it has been argued that many macroeconomic and financial series exhibit nonstationary behavior in the error. In particular, heteroscedastic behavior in the error is well known in the unit root testing literature. For instance, the U.S. gross domestic product series was observed to have less variability since the 1980s; see Kim:Nelson:1999, McConnell:Perez-Quiros:2000, Busetti:Taylor:2003, and references therein. Also the majority of macroeconomic data in Stock:Watson:1999 exhibit heteroscedasticity in unconditional variances, as pointed out by Sensier:vanDijk:2004. If there are breaks in the error structure, it is known that traditional unit root tests are biased towards rejecting the unit root assumption Busetti:Taylor:2003. For this reason, a number of unit root tests that are robust to heteroscedasticity have been developed in the literature, such as in Busetti:Taylor:2003, Cavaliere:Taylor:2007, Cavaliere:Taylor:2008a, Cavaliere:Taylor:2008b, Cavaliere:Taylor:2009a, and Smeekes:Urbain:2014. Such tests allow for smooth and abrupt changes in the unconditional or conditional variance in error processes. However, changes in underlying dynamics do not have to be limited to heteroscedastic behavior. For example, in financial econometrics, it has been argued that a slow decay in sample autocorrelation in squared or absolute stock returns might be due to a smooth change in its dynamics rather than a long-memory behavior. Accordingly, Starica:Granger:2005 and Fryzlewicz_etal:2008, among others, proposed to model stock returns series with locally stationary models. A unit root test that is robust to changes in general error dynamics would be useful in identifying a long-memory behavior in such financial series.
In a series of papers by Taylor and coauthors, heteroscedasticity in the error is accommodated by assuming that the error is generated from linear processes with heteroscedastic innovations, i.e.,
Here, $\varepsilon_t$ is assumed to be i.i.d. or to be a martingale difference sequence, and $\omega_t$ is a sequence of deterministic numbers that account for heteroscedasticity. The process ((ref)) can be considered as a generalized linear process. Although this generalization allows for some departures from stationarity, it is still restrictive in the following three aspects. First, this kind of linear process cannot accommodate nonlinearity in the error process and does not include nonlinear models that are popular in time series analysis, such as threshold, bilinear and nonlinear moving average models. Second, this error structure is somewhat special in that temporal dependence and heteroscedasticity can be separated, and it seems that most methods developed to account for heteroscedasticity in the error take advantage of this special error structure. Third, changes in second- or higher-order properties in the error process $u_t$ are not as flexibly accommodated as those in $e_t$. Recently, Smeekes:Urbain:2014 proposed a unit root test for a piecewise modulated stationary process $u_t=\omega_t v_t$, where $v_t$ is a weakly stationary process and $\omega_t$ is a sequence of deterministic numbers that accounts for heteroscedasticity. While a modulated stationary processes is more flexible in handling heteroscedasticity in $u_t$, it is still restrictive in the sense that time dependence and heteroscedasticity can be separated, similarly to ((ref)). For more discussion on the separability of ((ref)) and modulated stationary processes in the context of linear regression models with fixed regressors and nonstationary errors, see Rho:Shao:2015.
In this paper, a general framework of nonstationarity is adapted to capture both smooth and abrupt changes in second- or higher-order properties of the error process. Specifically, the error process is assumed to follow a piecewise locally stationary (PLS) process, which was recently proposed by Zhou:2013, as a generalization of a locally stationary process. Locally stationary processes have received a lot of attention since the seminal works of Priestley:1965 and Dahlhaus:1997. Local stationarity naturally expands the notion of stationarity by allowing a change of the second-order properties of a time series; see Dahlhaus:1997, Mallat:Papanicolaou:Zhang:1998, Giurcanu:Spokoiny:2004, and Zhou:Wu:2009, among others, for more related work. However, locally stationary processes exclude abrupt changes in second- or higher-order properties, which is often observed in real data. To accommodate abrupt changes, PLS processes were proposed to allow for a finite number of breaks in addition to smooth changes. For example, Adak:1998 proposed a PLS model in frequency domain, generalizing the local stationary model due to Dahlhaus:1997. Zhou:2013 proposed another PLS model in time domain as an extension of the framework of Zhou:Wu:2009 and Draghicescu:Guillas:Wu:2009. This PLS process allows for both nonlinearity and piecewise local stationarity and covers a wide range of processes; see Wu:2005, Zhou:Wu:2009, and Zhou:2013 for more details.
Under the general PLS framework for the error, the limiting null distributions of the conventional unit root test statistics are not pivotal; they depend on the local longrun variance of the PLS error and some other nuisance parameters. A direct consistent estimation of the unknown parameters in the limiting null distributions is very involved, unlike the case of stationary errors Phillips:1987a. To overcome this difficulty, we propose to apply the dependent wild bootstrap (DWB) proposed in Shao:2010b to approximate the limiting null distributions. We also provide a rigorous theoretical justification by establishing the functional central limit theorem for the standardized partial sum process of the bootstrapped residuals. This seems to be the first time DWB is justified for PLS processes and in the unit root setting. It suggests the ability of DWB to accommodate both piecewise local stationarity and weak dependence, which can be potentially used for other inference problems related to locally stationary processes.
The rest of the paper is organized as follows. Section (ref) presents the model, along with the test statistics and their limiting distributions under the null and local alternatives. In Section (ref), DWB is described and its consistency is justified. The power behavior under local alternatives is also presented. Section (ref) presents some simulation results. Section (ref) summarizes the paper. Technical details are relegated to the supplementary material.
We now set some standard notation. Throughout the paper, $\stackrel{\cal D}{\longrightarrow}$ is used for convergence in distribution and $\Rightarrow$ signifies weak convergence in $D[0,1]$, the space of functions on [0,1] which are right continuous and have left limits, endowed with the Skorohod metric Billingsley:1968. Let $a_n\asymp c_n$ indicate $a_n/c_n\to1$ as $n\to\infty$. For $a\in\mathbb{R}$, $\lfloor a\rfloor$ denotes the largest integer smaller than or equal to $a$. $B(\cdot)$ denotes a standard Brownian motion, and $N(\mu,\Sigma)$ the (multivariate) normal distribution with mean $\mu$ and covariance matrix $\Sigma$. Set $||X||_p=(E|X|^p)^{1/p}$. Let ${\bf 1}(\mathcal{E})$ be the indicator function, being 1 if the event $\mathcal{E}$ occurs and 0 otherwise.
Following the framework in Phillips:Xiao:1998, consider data $\{y_{1,n},\ldots,y_{n,n}\}$ generated from
and
Here, $\beta$ is a $p\times 1$ vector of coefficients and $z_{t,n}$ is a $p\times 1$ vector of deterministic trend functions, which satisfies the following conditions:
These assumptions include some popular trend functions, such as $(p-1)$st order polynomial trends, and are quite standard in the literature; see Section 2.1 of Phillips:Xiao:1998 and Section 2 of Cavaliere:Taylor:2007. The initial condition, $X_{0,n}=0$ is assumed in order to simplify the argument. This assumption can be relaxed to allow, e.g., $X_{0,n}$ to be bounded in probability, which does not alter our asymptotic results.
Following the framework introduced by Zhou:2013, the error process $\{u_{t,n}\}_{t=1}^n$ is assumed to be mean-zero piecewise locally stationary (PLS) with a finite number of break points. Let the break points be denoted by $b_1,b_2,\ldots,b_\tau$, where $0=b_0<b_1<\ldots<b_\tau<b_{\tau+1}=1$. The process $\{u_{t,n}\}_{t=1}^n$ is considered as a concatenation of $\tau+1$ measurable functions $G_j(s,\mathcal{F}_t):[0,1]\times\mathbb{R}^\infty\to\mathbb{R}$, $j=0,1,\ldots,\tau$, where $$u_{t,n}=G_j(s_t,\mathcal{F}_t)~~~~{\rm if}~b_j\leq s_t< b_{j+1},$$ $s_t=t/n$, $\mathcal{F}_t=(\ldots,\varepsilon_0,\ldots,\varepsilon_{t-1},\varepsilon_t)$, and the $\varepsilon_t$ are i.i.d. random variables with mean 0 and variance 1. The following is further assumed:
If there is no break point and the function $G$ does not depend on its first argument, then the PLS process reduces to a nonlinear causal process $G(\mathcal{F}_t)$, which can accommodate a wide range of stationary processes. A special example is when $G$ is a linear function, in which case $G(\mathcal{F}_t)=\sum_{j=0}^\infty c_j\varepsilon_{t-j}$ is the commonly used linear process. For nonlinear time series models that fall into the above framework of a nonlinear causal process, see Shao:Wu:2007.
By introducing the dependence on the relative location $t/n$, the PLS series naturally extends this stationary causal process $G(\mathcal{F}_t)$ to a locally stationary one; see Zhou:Wu:2009. In particular, assumption (A1) states that if $t/n$ and $t'/n$ are close and there is no break point in between, then $u_{t,n}$ and $u_{t',n}$ are expected to be stochastically close. In other words, the second- or higher-order property of $u_{t,n}$ should be smoothly changing, except for a finite number of break points. This ensures the local stationarity between break points. As stated in Zhou:2013, the goal of extending from locally stationary processes to PLS processes is to allow for abrupt changes in both second- and high-order properties, and to accommodate both nonlinearity and nonstationarity in a broad fashion.
The physical dependence measure in (A3) was introduced in Zhou:Wu:2009 as an extension of its stationary counterpart first introduced by Wu:2005. Assumption (A3) implies that $u_{t,n}$ is locally short-range dependent and that the dependence decays exponentially fast. When the location $s\in[0,1]$ is fixed, the process $\{G_j(s,\mathcal{F}_t)\}_{t\in\mathbb{Z}}$ is stationary for each $j$, and (A4) introduces the time-varying longrun variance parameter $\sigma^2(s)$. Assumptions (A1)--(A4) are similar to those of Zhou:Wu:2009, Wu:Zhou:2011, and Zhou:2013. These assumptions are not the weakest possible for our theoretical results to hold but are satisfied by a wide class of time series models. See the following examples:
Given the observations $\{y_{t,n},z_{t,n}\}_{t=1}^n$, consider testing the unit root hypothesis
The ordinary least squares (OLS) estimator $\widehat{\rho}_n=\left(\sum_{t=1}^n\widehat{X}_{t,n}\widehat{X}_{t-1,n}\right)/\left(\sum_{t=1}^n\widehat{X}_{t-1,n}^2\right)$ of $\rho$ is considered, where $\widehat{X}_{t,n}=y_{t,n}-\widehat{\beta}_n'z_{t,n}$ are the OLS residuals of $y_{t,n}$ regressed on $z_{t,n}$.\footnote{It is worth mentioning that the generalized least squares (GLS) detrending can be used instead of the OLS detrending to make the tests more powerful. This is quite straightforward and well established in the literature Elliott_etal:1996,Muller:Elliott:2003,Smeekes:2013, and will not be pursued in this paper.} We proceed to define two test statistics that are popular in the literature, namely $$\mathbf{T}_n=n(\widehat{\rho}_n-1)~~~~{\rm and}~~~~ \mathbf{t}_n =\frac{\left(\sum_{t=1}^n\widehat{X}_{t-1,n}^2\right)^{1/2}\left(\widehat{\rho}_n-1\right)}{(s_n^2)^{1/2}}, $$ where $s_n^2=(n-2)^{-1}\sum_{t=1}^n\left(\widehat{X}_{t,n}-\widehat{\rho}_n\widehat{X}_{t-1,n}\right)^2$.
In the unit root testing literature, local alternatives, i.e., $\rho_n=1+c/n$, $c<0$, are often considered to examine the behavior of the test when the true $\rho$ is close to the unity. The Ornstein-Uhlenbeck process, $J_c(r)=\int_0^re^{(r-s)c}dB(s)$, is usually involved in the limiting distributions of $\mathbf{T}_n$ and $\mathbf{t}_n$ for this near-integrated case. Under our error assumptions, define a similar process, $J_{c,\sigma}(r):=\int_0^re^{(r-s)c}\sigma(s)dB(s)$. Theorem (ref) below states the limiting distributions of the two test statistics under the null hypothesis, $\rho=1$, and under local alternatives, $\rho_n=1+c/n$, $c<0$.
When $\rho$ is far from the unity, with $c<0$ and large in absolute value, the limiting distributions are far to the left compared to the limiting null distributions. In this case, the unit root null hypothesis would be rejected with high probability. This implies that the unit root tests based on $\mathbf{T}_n$ and $\mathbf{t}_n$ have nontrivial powers under local alternatives. In the special case $\beta\equiv0$ and $\sigma(s)=\sigma$, i.e., when there is no deterministic trend and the error is stationary, the two limiting distributions $\mathcal{L}_{\mathbf{T},c}$ and $\mathcal{L}_{\mathbf{t},c}$ reduce to those found in Theorem 1 of Phillips:1987b.
Notice that when $c=0$, i.e., under the null hypothesis, $J_{c,\sigma|Z}(r)=B_{\sigma|Z}(r)$, which implies $\int_0^1B_{\sigma|Z}(r)\sigma(r) \allowbreak dB(r)=2^{-1}\{B_{\sigma|Z}^2(1)-\int_0^1\sigma^2(r)dr\}$, by It\^{o}'s formula.\footnote{Ito's formula can be written as $dB_\sigma(r)=\sigma(r)dB(r)$. Using It\^{o}'s formula, we derive $B_\sigma^2(r)=B_\sigma^2(0)-\int_0^r2\sigma(s)B_\sigma(s)dB(s)+2^{-1}\int_0^r2\sigma^2(s)ds$, which leads to $\int_0^rB_{\sigma}(s)\sigma(s)dB(s)=2^{-1}\{B_{\sigma}^2(r)-\int_0^r\sigma^2(s)ds\}$.} In this case, the limiting null distributions can be written as
where $B_\sigma(r)=\int_0^r\sigma(s)dB(s)$ and $$B_{\sigma|Z}(r)=B_\sigma(r)-\left\{\int_0^1B_\sigma(s) Z(s)'ds\right\}\left\{\int_0^1 Z(s)Z(s)'ds\right\}^{-1}Z(r)$$ is the Hilbert projection of $B_\sigma(\cdot)$ onto the space orthogonal to $Z(\cdot)$.
Since the limiting null distributions contain a number of unknown parameters, one may try to directly estimate them for inferential purpose. Consistent estimation of the limit of the average marginal variance, $\sigma_u^2$, is not difficult. This can be done by noting that $\sigma_u^2=\int_0^1c(s;0)ds$, with $c(s;h)$ defined in (A4) and in the first paragraph of the technical appendix. As a special case of Lemma A.5, $\sigma_u^2$ can be consistently estimated by $n^{-1}\sum_{t=1}^n \widehat{u}_{t,n}^2$, where $\widehat{u}_{t,n}=\widehat{X}_{t,n}-\widehat{\rho}_n\widehat{X}_{t-1,n}$ is the OLS residual. However, consistent estimation of the local longrun variance $\sigma^2(s)$ is not as simple for a PLS process. In the case of a stationary error process, $c(s;h)$ does not depend on $s$, and $\sigma^2(s)=\sum_{h=-\infty}^\infty {\mbox{cov}}(u_{t},u_{t+h})=\sigma^2$. If it is further assumed that there is no deterministic trend function, i.e., $\beta\equiv0$, then the limiting null distributions $\mathcal{L}_{\mathbf{T}}$ and $\mathcal{L}_{\mathbf{t}}$ reduce to $$\mathcal{L}_{\mathbf{T},\sigma(s)=\sigma}=\frac{\{B(1)^2-\sigma_u^2/\sigma^2\}}{2\int_0^1B(r)^2dr}~~~{\rm and}~~~ \mathcal{L}_{\mathbf{t},\sigma(s)=\sigma}=\frac{\sigma/\sigma_u\{B(1)^2-\sigma_u^2/\sigma^2\}}{2\{\int_0^1B(r)^2dr\}^{1/2}}.$$ These limiting null distributions contain only a couple of unknown parameters and coincide with those in Phillips:1987a. To make inference possible in this stationary error case, as Phillips:1987a and Phillips and Perron (1998) suggested, the longrun variance $\sigma^2$ of the error process may be consistently estimated using heteroscedasticity and autocorrelation consistent (HAC) estimators. The new statistics (see page 287 of Phillips:1987a), adjusted by using consistent estimates of $\sigma$ and $\sigma_u$, have pivotal limiting null distributions. However, in the piecewise locally stationary error case, the usual Phillips-Perron adjustment may not lead to pivotal limiting null distributions. Specifically, the HAC-based estimator of the nuisance parameter $\sigma^2(s)$, $s\in[0,1]$, is not known to be consistent in the PLS framework. The parameter $\sigma^2(s)$ is unknown at infinitely many points, and the integral of $\sigma(s)$ over a Brownian motion needs to be estimated as well as $\sigma^2(s)$ itself. This makes the direct estimation of the unknown parameters in the limiting null distributions difficult.
To implement the (asymptotic) level $\alpha$ test, the $\alpha$-quantiles of the limiting null distributions $\mathcal{L}_{\mathbf{T}}$ and $\mathcal{L}_{\mathbf{t}}$ need to be identified and estimated. However, it is difficult to consistently estimate the unknown parameters $\sigma(s)$ for all $s\in[0,1]$. As an alternative, we shall use a bootstrap method to approximate the limiting null distributions. When the errors are stationary, it is well known that the Phillips-Perron test or the augmented Dickey-Fuller test have size distortions in finite samples, even though they have been proven to work asymptotically. Bootstrap-based methods have been proposed to improve the finite sample performance. Psaradakis:2001, Chang:Park:2003, and Palm:Smeekes:Urbain:2008 used the sieve bootstrap Kreiss:1988 assuming an infinite order AR structure for the error process. Paparoditis:Politis:2003 applied the block bootstrap Kunsch:1989, which randomly samples from overlapping blocks of residuals. Swensen:2003 and Parker:Paparoditis:Politis:2006 extended the stationary bootstrap Politis:Romano:1994 to unit root testing, where not only are overlapping blocks randomly chosen, but also the block size is chosen from a geometric distribution. Cavaliere:Taylor:2009a applied the wild bootstrap Wu:1986 for unit root M tests Perron:Ng:1996, which are modifications of the Phillips-Perron test.
To accommodate both heteroscedasticity and temporal dependence in the error, bootstrap-based methods have been developed by Cavaliere:Taylor:2008b and Smeekes:Taylor:2012. In their papers, the error is assumed to be a linear process with heteroscedastic innovations as in equation ((ref)), and the wild bootstrap Wu:1986 was used along with an AR sieve procedure to filter out the dependence in the error. The combination of an autoregressive sieve (or recolored) filter and wild bootstrap handles heteroscedasticity and serial correlation simultaneously, but its theoretical validity strongly depends upon the linear process assumption on the error. We speculate that this recolored wild bootstrap (RWB) will not work for PLS errors because the naive wild bootstrap can account for heteroscedasticity, but not for weak temporal dependence that is not completely filtered out after applying the AR sieve. The wild bootstrap works in the framework of ((ref)), since temporal dependence is removed by the Phillips-Perron adjustment Cavaliere:Taylor:2008b or the augmented Dickey-Fuller adjustment (i.e., the AR sieve) Cavaliere:Taylor:2009a,Smeekes:Taylor:2012. This type of removal is only possible under the assumption that the error is a heteroscedastic linear process, in which case the longrun variance $\sigma^2(s)$ can be factored into two parts; one part is due to heteroscedasticity in the innovations $e_t$, and the other part is due to temporal dependence (i.e., $\sum_{j=0}^\infty c_j$) in the error, as shown recently by Rho:Shao:2015. As a result, limiting null distributions after the two popular adjustments for temporal dependence depend only on heteroscedasticity, which can be handled by the wild bootstrap. However, for PLS error processes, even after the Phillips-Perron or the augmented Dickey-Fuller adjustment, the limiting null distributions are still affected by temporal dependence in the error. Therefore, RWB is not expected to work in our setting.
To accommodate nonstationarity and temporal dependence in the error, we propose to adopt the so-called dependent wild bootstrap (DWB), which was first introduced by Shao:2010b in the context of stationary time series. It turns out that DWB is capable of mimicking local weak dependence in the error process and provides a consistent approximation of the limiting null distributions of $\mathbf{T}_n$ and $\mathbf{t}_n$. Note that DWB was developed for stationary time series and its applicability was only proved for smooth function models. Smeekes:Urbain:2014 recently proved the validity of several modified wild bootstrap methods, including DWB, for modulated stationary errors in a multivariate setting. However, as discussed in the introduction, the modulated stationary process is somewhat restrictive due to its separable structure of temporal dependence and heteroscedasticity in its longrun variance. Instead, our PLS framework is considerably more general in allowing both abrupt and smooth change in second- and higher-order properties. From a technical viewpoint, our proofs seem more involved than theirs due to the general error framework we adopt.
In the implementation of DWB, pseudo-residuals are generated by perturbing the original (OLS) residuals using a set $\{W_{t,n}\}_{t=1}^n$ of external variables. The difference between DWB and the original wild bootstrap is that $\{W_{t,n}\}_{t=1}^n$ is made to be dependent in DWB, whereas $\{W_{t,n}\}_{t=1}^n$ is assumed to be independent in the usual wild bootstrap. The following assumptions on $\{W_{t,n}\}_{t=1}^n$ are from Shao:2010b:
In practice, $\{W_{t,n}\}_{t=1}^n$ can be sampled from a multivariate normal distribution with mean zero and covariance function ${\rm cov}(W_{t,n},W_{t',n})=a\{(t-t')/l\}$. There are two user-determined parameters: a kernel function $a(\cdot)$ and a bandwidth parameter $l$. The kernel function affects the performance to a lesser degree than the bandwidth parameter $l$, and the choice of $l$ will be discussed in Section (ref). For the kernel function, some commonly used kernels, such as the Bartlett kernel, satisfy (B2).
The DWB algorithm in unit root testing is as follows:
The following theorem provides the core result in the proof of the consistency of DWB in Theorem (ref) and may be of independent interest.
Note that Theorem (ref) holds not only under the null hypothesis $\rho=1$ but also under local alternatives. This property makes the DWB method powerful because the bootstrapped distributions correctly mimic the limiting null distributions under both the null and local alternatives. The DWB method can still correctly approximate the limiting null distribution under local alternatives, mainly because $y_{t,n}^*$ are constructed assuming $\rho=1$ in step 5.
Under the null hypothesis, i.e., when $c=0$, Theorem (ref) establishes the consistency of DWB in approximating the limiting null distributions. Since the bootstrap statistics (asymptotically) replicate the exact null distribution when $\rho=1$, the (asymptotic) size of our unit root test would be exactly the same as the level of the test. On the other hand, if $c$ is negative and far from 0, the probability of rejecting the null, or the asymptotic power of the test, will be close to 1. If $c$ is not 0 but not too far from 0, Theorem (ref) states that the probability of rejecting the null is somewhere between the level of the test and 1. This means that the DWB-based unit root tests have nontrivial power under local alternatives.
In this section, the DWB method is compared with the recolored (sieve) wild bootstrap (RWB) method, which was proposed in Cavaliere:Taylor:2009a. We also propose to combine the AR sieve idea in RWB with DWB and present this method as the recolored dependent wild bootstrap (RDWB). The RDWB statistics are based on the RWB statistics using DWB to determine the critical values of the tests, so RDWB can be considered as a generalization of RWB.
Before introducing our simulation setting and results, we first present some details about (i) the RDWB algorithm and (ii) the size-corrected power calculation similar to the one in Dominguez:Lobato:2001.
First, the RDWB procedure is described below. Rewrite equation ((ref)) as
where $\Delta$ represents the difference operator.
For a fair comparison of power, the following size-corrected power procedure similar to Dominguez:Lobato:2001 is adapted.
The following data generating processes (DGPs) are used for comparison of DWB, RWB, and RDWB in finite samples. For simplicity, set $\beta\equiv 0$ so that $\widehat{X}_{t,n}=X_{t,n}$. Consider ((ref)) and $u_{t,n}$ generated from time-varying moving average (MA) and autoregressive (AR) models with lag 1, $$({\rm MA}_{i,j})~u_{t,n}=e_{j,t,n}+\phi_i(t/n) e_{j,t-1,n},~~~~({\rm AR}_{i,j})~u_{t,n}=e_{j,t,n}+\phi_i(t/n) u_{t-1,n}$$ for $t=1,\ldots,n$, where $e_{j,t,n}=\omega_j(t/n)\varepsilon_t$, $\varepsilon_t\stackrel{i.i.d.}{\sim}N(0,1)$. The MA or AR coefficient $\phi_i(s)$ is possibly time-varying with the following six choices: for $s\in[0,1]$, $$\phi_1(s)=0.8,~\phi_2(s)=-0.8,~\phi_3(s)=0.2+0.6{\bf 1}(s>0.2),$$ $$\phi_4(s)=0.2+0.6{\bf 1}(s>0.8),~\phi_5(s)=0.8-1.6s,~{\rm and}~\phi_6(s)=0.6s-0.8.$$ The function $\omega_j(s)$ governs possible heteroscedastic behavior in $u_{t,n}$ with the following five choices: for $s\in[0,1]$, $$\omega_1(s)=0.5,~\omega_2(s)=0.1+0.5{\bf 1}(s>0.1),~\omega_3(s)=0.1+0.5{\bf 1}(s>0.9),$$ $$\omega_4(s)=0.1+0.5{\bf 1}(0.4<s<0.6),~{\rm and}~\omega_5(s)=0.5s+0.1.$$ Combinations of $\phi_i(s)$ and $\omega_j(s)$ along with the choice of MA or AR lead to 60 DGPs that satisfy the PLS assumption in (A1)--(A4). In particular, if $i=1$ or 2, $\phi_i(s)$ is constant over $s\in[0,1]$. The corresponding $\{u_{t,n}\}$ processes fall into the category of linear processes with heteroscedastic error in ((ref)), making RWB consistent for any choices of $\omega_j(s)$, $j=1,\ldots,5$. These settings are to mirror the setup of the Cavaliere-Taylor papers. For all other settings, the asymptotic consistency of RWB is not guaranteed, whereas DWB and RDWB are expected to work asymptotically. Sudden increases and smooth changes in MA or AR coefficients are presented in the cases with $i=3,4$ and $i=5,6$, respectively. The variance of $e_{j,t,n}$ is a constant ($j=1$), a step function with a sudden increase in the beginning ($j=2$) and end ($j=3$) of the series, a step function with a sudden increase and decrease in the middle ($j=4$), or a smoothly increasing sequence ($j=5$).
The sample sizes $n=100$ and $400$ are considered. The number of Monte-Carlo replications is 2000, and the number of bootstrap replications is $B=1000$ for all bootstrap methods. For local alternatives, $c=0, -5, -10, -15, -20, -25, -30$ are considered. In particular, for DWB and RDWB, in each replication, pseudoseries $(W_{1,n},\ldots,W_{n,n})'$ are generated from i.i.d. $N(\mathbf{0}_n,\Sigma)$, where $\Sigma$ is an $n$ by $n$ matrix with its $(i,j)$th element being $a\{(i-j)/l\}$. Here the Bartlett kernel is used, i.e., $a(s)=(1-|s|){\bf 1}(|s|\leq 1)$. For DWB and RDWB, the bandwidth parameter $l$ is chosen as $l=\lfloor 6(n/100)^{1/4}\rfloor$. That is, $l=6$ if $n=100$, and $l=8$ if $n=400$. In Section B of the supplementary material, (i) full details on the effect of different choices of $l$ for selected DGPs are presented and (ii) a data-driven choice of $l$, the minimum volatility method, is proposed. It seems that the empirical sizes are not overly sensitive to the choice of $l$, as long as $l$ is not too small, and the finite sample size comparison with the MV method in Section B of supplementary material supports the above deterministic choice of $l$.
Tables (ref) and (ref) present the empirical sizes of the three methods when the nominal size is 5%. When the model is stationary with positive coefficient $\phi(s)=\phi_1(s)=0.8$, i.e., (${\rm MA/AR}_{1,j}$) for $j=1,\ldots,5$, all three bootstrap methods produce reasonably accurate sizes, except that the DWB method tends to under-reject for the AR models. This under-rejecting behavior of DWB is observed consistently for most AR models. This might be due to the fact that the DWB method mimics the time dependence in the original data in a manner similar to MA models, so that it does not produce as accurate sizes for AR models as for MA models. The RDWB method nicely compensates this shortcoming by applying an AR-based prewhitening. The prewhitening effect is most noticeable when the model is stationary with negative coefficient $\phi(s)=\phi_2(s)=-0.8$, i.e., (${\rm MA/AR}_{2,j}$) for $j=1,\ldots,5$. For these models with negative autocorrelation, the size-distortion of the DWB method is very large at both sample sizes with slight less distortion for larger sample size. This suggests that although the DWB method should work asymptotically for the negative coefficient case, this convergence could be too slow to be useful in practice. On the other hand, after applying the AR-based prewhitening, similar to RWB, finite sample sizes are brought closer to the nominal level.
A careful examination of RWB shows that it has a fairly accurate size, especially when it is theoretically supported ($i=1,2$). However, for some DGPs with changing MA or AR coefficients ($i=3,4,5,6$), RWB does not seem to be consistent. In particular, in the MA models, the sizes of RWB tend to further deviate from the nominal level as the sample size $n$ increases when there is a sudden increase in the variance in innovations at the latter part of the series ($j=3$) with changing variance (see $({\rm MA}_{3,3})$, $({\rm MA}_{4,3})$, $({\rm MA}_{5,3})$, and $({\rm MA}_{6,3})$) or when both MA coefficient and variance of innovations change smoothly (see $({\rm MA}_{5,5})$). In the AR models, RWB tends to have heavier size distortion as $n$ increases when the AR coefficient changes drastically from negative to positive ($i=5$; see $({\rm AR}_{5,1})$, $({\rm AR}_{5,2})$, $({\rm AR}_{5,3})$, and $({\rm AR}_{5,5})$) or when the AR coefficient is negative and changes smoothly and the variance in innovations suddenly increases at the latter part of the series (see $({\rm AR}_{6,3})$). This size distortion might be an indication that the AR prewhitening (RWB) alone does not work in theory, and the dependence in the error is not completely filtered out. By contrast, RDWB tends to have more accurate sizes for these models, although size distortion due to inaccurate prewhitening is still apparent to a lesser degree. On the other hand, as long as the MA or AR coefficients are nonnegative, DWB without prewhitening is always demonstrated to have more accurate size as $n$ increases. In particular, for MA models with changing MA coefficient ($i=3,4,5$), DWB tends to produce the best size with the most consistent behavior among the three bootstrap methods.
Overall, the size for RDWB seems to be the most reliable among the three bootstrap methods if the underlying DGP is not known. In some unreported simulations, we have observed the following: (i) the large size distortion associated with the DWB method for negative autocorrelation models, (${\rm MA/AR}_{2,1}$), can be reduced to below the nominal $5\%$ level if we use restricted residuals, at the price of power loss; (ii) a comparison with residual block bootstrap in Paparoditis:Politis:2003 shows that the size for the residual block bootstrap can be quite distorted for some DGPs, e.g., (${\rm MA}_{5,5}$). This indicates the inability of residual block bootstrap to consistently approximate the limiting null distribution when the error process is PLS.
Figures (ref) and (ref) present the power curves of $\mathbf{t}_n$ for DWB, RWB, and RDWB for selected DGPs with $n=100$ and $400$, respectively. The size-adjusted power curves in the first panel, (${\rm MA}_{4,1}$), are representative for most of the cases where all three bootstrap methods have reasonably accurate sizes, where (i) DWB tends to have the best power, (ii) RDWB tends to have slightly better power than RWB when $n=100$, and (iii) RWB and RDWB are fairly comparable in terms of sizes and powers in general. Most MA or AR models with positive MA or AR coefficient for at least part of a series ($\phi_i(s)$ with $i=1,3,4,5$) and constant, early break, or smooth change in error variance ($\omega_j(s)$ with $j=1,2,5$) tend to have a similar shape. The second panel, (${\rm MA}_{2,1}$), represents the size-adjusted curves when the MA or AR coefficients are negative at all time points ($i=2,6$) so that the finite sample size of DWB is highly distorted. Even though DWB has the best power, it is not recommended due to its big size distortion in this case. It seems that RWB and RDWB do not have much difference in terms of size-adjusted power.
The last two panels focus on the comparison between RWB and RDWB. The third panel, (${\rm MA}_{1,3}$), is representative when there is a jump in $\omega(s)$ at the end of the series ($j=3$) and when both RWB and RDWB have reasonable sizes. Models with $i=1,2,6$ and $j=3,4$ tend to have a similar pattern if DWB is ignored due to its high size distortion when $i=2,6$. In this case, RWB and RDWB have similar powers, as RWB tends to have slightly better power when $i = 5$ or 6, whereas RDWB tends to have slightly higher power when $i = 1$. The last panel, (${\rm MA}_{6,3}$), is representative for the case when RWB is not consistent. (${\rm AR}_{i,j}$) or (${\rm MA}_{i,j}$) with $i=3,4,5$ and $j=3,4$ fall into this category. In this case, RDWB seems to present the most reasonable size and power. Even though RWB appears to have the best power, RWB does not seem to be consistent due to the considerable increase in its finite sample size as $n$ increases for some models. Complete power curves for all DGPs are presented in the supplementary material as Figures C.1--C.4. It is worth noting that $\mathbf{t}_n$ tends to produce more accurate sizes with little power loss (or slightly better power) than $\mathbf{T}_n$.
In summary, RDWB, the combination of RWB and DWB, appears to work well in finite samples. It tends to produce reasonably high powers and fairly accurate sizes in all models under examination. In the situation when DWB or RWB have a large size distortion, the size accuracy of RDWB is well maintained and its power appears quite reasonable in all cases. One downside associated with RDWB is that it requires two tuning parameters: the truncation lag in the AR sieve and the bandwidth parameter in DWB. In this paper, we choose the number of lags for RWB and RDWB using the MAIC method. As for the bandwidth parameter, it seems that DWB and RDWB are not sensitive to the choice of the bandwidth parameter and the proposed deterministic choice seems to perform reasonably well in finite samples. Given that the DGP is unknown in practice, we shall recommend the use of RDWB.
In this paper, we present a new bootstrap-based unit root testing procedure that is robust to changing second- and higher-order properties in the error process. The error is modeled as a piecewise locally stationary (PLS) process, which is general enough to include time-varying nonlinear processes as well as heteroscedastic linear processes as special cases. In particular, the PLS process does not impose a separable structure on its longrun variance as do heteroscedastic linear processes and modulated stationary processes, which have been adopted in the literature to model heteroscedasticity and weak dependence of the error. Under the PLS framework, the limiting null distributions of two popular test statistics are derived and the dependent wild bootstrap (DWB) method is used to approximate these non-pivotal distributions. The functional central limiting theorem has been established for the standardized partial sum process of the DWB residuals, and bootstrap consistency is justified under local alternatives. The DWB-based unit root test has asymptotically nontrivial local power. The DWB method was originally proposed for stationary time series. By showing its consistency in the PLS setting, we broaden its applicability and its use in the locally stationary context is worth further exploration. For finite sample simulations, we propose a recolored DWB (RDWB), combining the AR sieve idea used in the RWB test with DWB to improve the performance of the DWB-based test. In many cases, the RDWB method tends to provide the most accurate sizes and reasonably good power, compared to the use of DWB or RWB alone. In practice, with little knowledge of the error structure, the RDWB-based test seems preferable due to its robustness for a large class of nonstationary error processes.
\centerline {\bf\sc Acknowledgements} This research was partially supported by NSF grant DMS-1104545. We are grateful to the co-editor and the three referees for their constructive comments and suggestions that led to a substantial improvement of the paper. In particular, we are most grateful to Peter C. B. Phillips, who has gone beyond the call of duty for an editor in carefully correcting our English. We also thank Fabrizio Zanello, Mark Gockenbach, Benjamin Ong, and Meghan Campbell for proofreading. Superior, a high performance computing cluster at Michigan Technological University, was used in obtaining results presented in this publication.