EconBase
← Back to paper

Unit Root Testing with Slowly Varying Trends

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

76,537 characters · 8 sections · 50 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Unit Root Testing with Slowly Varying Trends

\def\spacingset#1{ {#1}} \spacingset{1}

\if11 \fi

\if01 {

center[center omitted — 74 chars of source]

} \fi

abstractA unit root test is proposed for time series with a general nonlinear deterministic trend component. It is shown that asymptotically the pooled OLS estimator of overlapping blocks filters out any trend component that satisfies some Lipschitz condition. Under both fixed-$b$ and small-$b$ block asymptotics, the limiting distribution of the $t$-statistic for the unit root hypothesis is derived. Nuisance parameter corrections provide heteroskedasticity-robust tests, and serial correlation is accounted for by pre-whitening. A Monte Carlo study that considers slowly varying trends yields both good size and improved power results for the proposed tests when compared to conventional unit root tests.

{\it Keywords:} unit root tests, nonlinear trends, heteroskedasticity \\ {\it JEL Classification:} C12, C14, C22 \spacingset{1.5}

Introduction

It is widely debated in the time series literature whether macroeconomic variables such as GDP, inflation, and interest rates are $I(1)$ or $I(0)$ around a deterministic trend. Dickey-Fuller-type unit root tests often fail to reject the null hypothesis for these time series. The trend component of a time series $y_t$ is typically treated as known up to some parameter vector. The most commonly applied unit root tests, such as those developed by dickey1979, said1984, phillips1987, phillipsperron1988, and elliott1996, impose either a constant or a linear trend model. If, however, the deterministic trend component is nonlinear, highly persistent trend-stationary processes can be hardly distinguishable from unit root processes (see, e.g., bierens1997 and becker2006).

It is not only a misspecified trend model that may lead to high power losses, as an overparameterized model can also reduce the power of unit root tests. Therefore, many authors have suggested applying trend models that seem more suitable for macro data. Broken trend models with one-time changes in mean or slope with known breakpoint were first studied by perron1989 and rappoport1989. christiano1992 demonstrated that a broken trend model with an unknown breakpoint is more adequate, and zivot1992, as well as banerjee1992, proposed unit root tests for this framework. Structural changes in innovation variances were studied by hamori1997, kim2002, and cavaliere2005, while cavaliere2011 considered unit root testing under broken trends together with nonstationary volatility. leybourne1998, kapetanios2003, and kilic2011 allowed for exponential smooth transitions from one trend regime to another. bierens1997 approximated a nonlinear mean function with Chebyshev polynomials, and enders2012a proposed a Fourier series approximation of the trend, which are approaches that can be used when the exact form and date of structural changes are unknown. For a comprehensive review on the research on unit root testing see choi2015.

Dickey-Fuller-type tests are based on the $t$-statistic of the first-order autoregressive parameter. In case of a constant trend, the estimator is derived from a regression of $\Delta y_t$ on $(y_{t-1} - \overline y)$, where $\overline y$ is the sample mean. schmidt1992 estimated the constant by the initial observation, which results in a regression of $\Delta y_t$ on $(y_{t-1} - y_1)$. Whereas a constant is often not a good global approximation, in a small block, a smoothly varying trend can be approximated quite closely by a constant. To exploit this fact, we propose a block procedure to filter out the unknown trend component. Blocking was also used in rooch2019 to estimate the fractional integration parameter in a similar situation. We divide the series into $T-B$ overlapping blocks of length $B$. As the blocks can be considered as units of a panel, we follow the panel unit root tests proposed by breitung2000 and levin2002 and consider a pooled regression of $\Delta y_{j+t}$ on $(y_{j+t-1} - y_j)$ for $2 \leq t\leq T$ and $1 \leq j \leq T-B$. The deterministic function is approximated locally by a constant. One could also use higher order local approximations of the trend function, but unreported simulations indicate that these approximations do not work well in samples of usual size. For this reason, we focus on constant local approximations. Under a general class of piecewise continuous trend functions, the resulting pooled estimator is consistent as $B,T \to \infty$. The limiting null distribution of the t-statistic is a functional of a Brownian motion under fixed-$b$ asymptotics. Under small-$b$ asymptotics, a normal distribution is obtained.

The paper is organized as follows: In Section (ref) the autoregressive model with independent and heteroskedastic errors is analyzed together with the asymptotic behavior of the pooled least squares estimator in the presence of a general nonlinear trend component. For both fixed-$b$ and small-$b$ block asymptotics, the limiting distributions are derived under both the unit root hypothesis and under local alternatives. In the presence of heteroskedastic errors, nuisance parameters appear in the limiting distributions, and the estimation of these parameters is discussed. Section (ref) considers pseudo $t$-tests for the unit root hypothesis, and heteroskedasticity-robust test statistics are provided. In Section (ref), a pre-whitening procedure is proposed in order to account for short-run dynamics, while Section (ref) reports on Monte Carlo simulations. The tests are found to have only minor size distortions in small samples and are sized correctly in larger samples. It is shown that in the presence of slowly varying trends, pooled tests tend to yield higher power than conventional unit root tests. Finally, Section (ref) presents the conclusion.

\if01 { While some proofs including those of the main theorems are presented in the Appendix, the more technical proofs are available as online supporting information. }\fi In the following, $W(r)$ denotes a standard Brownian motion and “$\Rightarrow$” stands for weak convergence on the c�dl�g space $D[0,1]$ together with a suitable norm. $\Theta(\cdot)$ denotes the exact order Landau symbol, that is, $a_T = \Theta(b_T)$ if and only if $a_T=O(b_T)$ and $b_T = O(a_T)$, as $T \to \infty$. Moreover, $\lfloor \cdot \rfloor$ is the integer part of its argument, and $\Delta y_t$ stands for the differenced series $y_t - y_{t-1}$. Finally, $\overset{d}{\longrightarrow}$ and $\overset{p}{\longrightarrow}$ denote convergence in distribution and convergence in probability.

The pooled estimator

We are interested in inference concerning the autoregressive parameter $\rho$ in the model

align[align omitted — 101 chars of source]

where $\rho$ is close or equal to one. The deterministic trend component $d_t$ is treated as nonstochastic and fixed in repeated samples, where its functional form is nonparametric and unknown.

assumption[trend component] The trend component is given by $d_t = d(t/T)$, where $d(r)$ is a piecewise Lipschitz continuous function.

Note that any continuously differentiable function is Lipschitz continuous. Lipschitz functions are locally close to a constant value in the sense that there exists some $C < \infty$ such that $|d(r) - d(s)| \leq C|r-s|$ for all $r,s \in \mathbb{R}$. The piecewise Lipschitz condition allows for a partition with a finite number of intervals, such that $d(r)$ is Lipschitz continuous on each interval. This includes both smooth changes as well as abrupt breaks in the trend function. For the initial value, it is assumed that $E[x_0^2] < \infty$. We introduce the pooled estimator and the unit root test statistics under the following assumptions on the error term:

assumption[heteroskedastic errors] The process $\{u_t\}_{t \in \mathbb{N}}$ is independently dis\-trib\-uted with $E[u_t]= 0$, $E[u_t^2] = \sigma_t^2$ and $E[u_t^4] < \infty$, where $\sigma_t = \sigma(t/T)$. The function $\sigma(r)$ is c\`{a}dl\`{a}g, non-stochastic, strictly positive, and bounded.

The principal approach to dealing with a general, slowly varying trend is to approximate the unknown trend locally by a constant. Let $B$ be some blocklength that satisfies $2 \leq B < T$. We divide the time series into $T-B$ overlapping blocks of length $B$ and then block-wise estimate $\rho$ via OLS under a constant trend specification. In the fashion of schmidt1992, as well as breitung1994, the constant trend is estimated by the first observation in each block, which corresponds to the maximum likelihood estimator under the unit root hypothesis $\rho = 1$. Thereafter, by pooling the $T-B$ individual block regressions, we obtain the regression equation

align*[align* omitted — 113 chars of source]

where $\phi = \rho - 1$. The pooled OLS estimator is formulated as

align*[align* omitted — 168 chars of source]

In the following, we derive the asymptotic properties for the numerator and the denominator separately. The numerator and denominator statistics are defined as

align*[align* omitted — 215 chars of source]

such that $\sqrt{BT} (\hat{\rho} - 1) = \mathcal Y_{1,T}/\mathcal Y_{2,T}$. Their counterparts without deterministics are given by

align*[align* omitted — 215 chars of source]

In what follows, we show that, under the block procedure, the deterministic component can be ignored asymptotically. All asymptotic results are jointly derived for $B,T \to \infty$. While the statistics $\mathcal X_{1,T}$ and $\mathcal X_{2,T}$ are infeasible if $d_t$ is unknown, they can be well approximated by $\mathcal Y_{1,T}$ and $\mathcal Y_{2,T}$ in the following sense:

lemmaLet $\rho = 1 - c/\sqrt{BT}$ with $c \geq 0$, let $d_t$ satisfy Assumption (ref), and let $u_t$ satisfy Assumption (ref). Then, as $B,T \to \infty$, $\mathcal Y_{1,T} - \mathcal X_{1,T} = O_P(B^{-1/2})$, and $\mathcal Y_{2,T} - \mathcal X_{2,T} = O_P(T^{-1/2})$.

Accordingly, we obtain $(\mathcal Y_{1,T} - \mathcal X_{1,T}, \mathcal Y_{2,T} - \mathcal X_{2,T}) \overset{p}{\longrightarrow} (0, 0)$ jointly, and the block procedure filters out the trend component in the numerator and the denominator asymptotically. Hence, applying Slutsky's theorem, we can write

align*[align* omitted — 139 chars of source]

This result is valid without any rate restrictions for $B$. In order to obtain the limiting distribution, we formulate some properties for the numerator and denominator statistics.

lemmaLet $\rho = 1 - c/\sqrt{BT}$ with $c \geq 0$, and let $u_t$ satisfy Assumption (ref). Then, as $B,T \to \infty$, the following statements hold true: \begin{itemize} • $\mathcal X_{1,T} = \sum_{j=1}^T q_{j,T} - c \cdot \mathcal W_T$, where $\{q_{j,T}, \ j\leq T, \ T \in \mathbb N\}$ is a martingale difference array with $q_{j,T} = B^{-3/2} T^{-1/2} \sum_{t \in \mathcal{I}_j} \sum_{k=1}^{t-1} u_j u_{j-k}$, $\mathcal{I}_j = \{ t \in \mathbb{N}: \ 1 \leq t \leq B, \ j+B-T \leq t \leq j-1\}$, and $\mathcal W_T = 0.5 \int_0^1 \sigma^2(r) \,\mathrm{d} r + O_P(B^{1/2} T^{-1/2})$. • $Var[\mathcal X_{1,T}] = \Theta(1)$ and $Var[\mathcal X_{2,T}] = \Theta(B T^{-1})$. • If $c=0$ and $\sigma_t^2 = \sigma^2$ for all $t \in \mathbb N$, \begin{align*} v_T^2 := \frac{\sigma^2 Var[\mathcal X_{1,T}]}{E[\mathcal X_{2,T}]} = \frac{(T-B)(2B-1)-2(B-2)}{3B(T-B)}. \end{align*} \end{itemize}

The previous results suggest distinguishing between different rates for $B$, which leads to two fundamentally different types of blocklength asymptotics. The fixed-$b$ approach denotes the case where the relative blocklength $B/T$ converges to some value $b$ with $0 < b < 1$, such that $B$ and $T$ grow at the same rate. In the small-$b$ approach, we consider a relative blocklength that converges to zero, while $B,T \to \infty$.\footnote{Note that the terminology \enquote{fixed-b and small-b asymptotics} was also used in the context of long-run variance estimation. Whereas kiefer2005 used this wording for the asymptotics of the ratio of the truncation point to the sample size, we consider the ratio of the blocklength to the sample size.} As the blocks are overlapping, the error terms in the pooled regression equation are correlated, but, fortunately, the correlation structure is known by construction. Together with the central limit theorem for martingale difference arrays, the following asymptotic result can be established for the small-$b$ case:

theoremLet $\rho = 1 - c/\sqrt{BT}$ with $c \geq 0$, let $d_t$ satisfy Assumption (ref), and let $u_t$ satisfy Assumption (ref). Let $B/T \to 0$ as $B,T \to \infty$. Then, \begin{align*} \mathcal Y_{1,T} \overset{d}{\longrightarrow} \mathcal{N} \bigg(-\frac{c}{2} \int_0^1 \sigma^2(r) \,\mathrm{d} r, \ \frac{1}{3} \int_0^1 \sigma^4(r) \,\mathrm{d} r \bigg), \quad and \ \quad \mathcal Y_{2,T} \overset{p}{\longrightarrow} \frac{1}{2} \int_0^1 \sigma^2(r) \,\mathrm{d} r. \end{align*}

Since $\mathcal Y_{2,T}$ converges in probability to a constant, we have joint convergence of $(\mathcal Y_{1,T}, \mathcal Y_{2,T})$, and the pooled estimator is asymptotically normally distributed under small-$b$ asymptotics. Under the unit root hypothesis $\rho=1$, or, equivalently, if $c = 0$, it follows that

align*[align* omitted — 199 chars of source]

The asymptotic variance of $\hat \rho$ involves integrals of the second- and fourth-order powers of the function $\sigma(r)$, where the factor $\int_0^1 \sigma^4(r) \,\mathrm{d} r/(\int_0^1 \sigma^2(r) \,\mathrm{d} r)^2$ is equal to unity in case of homoskedasticity. This factor also appears in the asymptotic variance matrix of the OLS estimator of the autoregressive coefficient under unconditional heteroskedasticity (see phillips2006).

cavaliere2005 showed that permanent changes in volatility induce a time-shift in the right-hand-side process of the functional central limit theorem. A variance-transformed Brownian process $W_\eta(r)$ appears in the limiting distributions of Dickey-Fuller-type unit root tests. Given the variance profile $\eta$, where $\eta(s) = ( \int_0^1 \sigma^2(r) dr )^{-1} \int_0^s \sigma^2(r) dr$, the transformed process is defined as $W_\eta(r) = W(\eta(r))$, where $W(r)$ is a standard Brownian motion. When imposing fixed-$b$ asymptotics, the numerator and denominator statistics can be represented as a partial sum process of the innovations, which leads to the following limiting result:

theoremLet $\rho = 1 - c/\sqrt{BT}$ with $c \geq 0$, let $d_t$ satisfy Assumption (ref), and let $u_t$ satisfy Assumption (ref). Let $0 < b < 1$, and let $B/T \to b$ as $B,T \to \infty$. Then, \begin{align*} \left(\begin{matrix} \mathcal Y_{1,T} \\ \mathcal Y_{2,T} \end{matrix}\right) \overset{d}{\longrightarrow} \left(\begin{matrix} 0.5 b^{-3/2} \int_0^1 \sigma^2(r) \,\mathrm{d} r \big(\int_0^{1-b} (J_{c,b,\eta}(b+r) - J_{c,b,\eta}(r))^2 - b(1-b) \big) \\ b^{-2} \int_0^1 \sigma^2(r) \,\mathrm{d} r \int_0^{1-b} \int_{r}^{b+r} (J_{c,b,\eta}(s) - J_{c,b,\eta}(r))^2 \,\mathrm{d} s \,\mathrm{d} r \end{matrix}\right), \end{align*} where $J_{c,b,\eta}(r) = \int_0^r e^{-(r-s)c/b} dW_\eta(s)$.

The limiting distributions are represented as functionals of the process $J_{c,b,\eta}$, which is an Ornstein-Uhlenbeck type process that is driven by a variance-transformed Wiener process. Consequently, the pooled estimator is asymptotically represented as a functional of a standard Brownian motion. If $\rho = 1$, the continuous mapping theorem and Theorem (ref) imply that

align*[align* omitted — 269 chars of source]

under fixed-$b$ asymptotics. In comparison to the limiting distribution of the $\rho$-statistic in the Dickey-Fuller framework, the functional includes an additional integral, which results from pooling the block regressions.

In order to estimate the unknown parameters in the limiting distributions, we consider the residuals $\hat{u}_t = y_t - \hat{\rho} y_{t-1}$ for $t=2, \ldots, T$ and their sample mean $\overline{\hat u} = (T-1)^{-1} \sum_{j=2}^T \hat u_j$. Let, for notational convenience, $\hat u_1 = 0$, and let

align*[align* omitted — 774 chars of source]

where $s \in [0,1]$. We obtain the following consistency results:

lemmaLet $\rho = 1-c/\sqrt{BT}$ with $c \geq 0$, let $d_t$ satisfy Assumption (ref), and let $u_t$ satisfy Assumption (ref). \begin{itemize} • $\hat{\sigma}^2 \overset{p}{\longrightarrow} \int_0^1 \sigma^2(r) \,\mathrm{d} r$, as $B,T \to \infty$. • $\sup_{s \in [0,1]} |\hat \eta(s) - \eta(s)| \overset{p}{\longrightarrow} 0$, as $B,T \to \infty$. • $\hat \kappa^2 \overset{p}{\longrightarrow} \int_0^1 \sigma^4(r) \,\mathrm{d} r / \int_0^1 \sigma^2(r) \,\mathrm{d} r$, as $B,T \to \infty$ and $B/T \to 0$. \end{itemize}

Pseudo t-statistics for unit root testing

The principal concept of Dickey-Fuller-type unit root tests is to consider a $t$-test for the null hypothesis $H_0: \rho = 1$. Following this approach in the pooled regression framework, the usual standard error is given by $s_{\hat \rho} = \hat{\sigma} ( \sum_{j=1}^{T-B} \sum_{t=2}^B (y_{t+j-1} - y_j)^2 )^{-1/2} = \hat{\sigma} ( \mathcal Y_{2,T} B^2T )^{-1/2}$ and the conventional $t$-statistic is represented as $(\hat \rho - 1)/s_{\hat \rho} = \sqrt{B} \mathcal Y_{1,T}/\sqrt{\hat \sigma^2 \mathcal Y_{2,T}}$, which diverges in probability under $H_0$. Accordingly, we consider a scaled pseudo $t$-statistic of the form

align[align omitted — 151 chars of source]

which is $O_P(1)$, as $B,T \to \infty$.

In what follows, pseudo $t$-tests are defined for both small-$b$ and fixed-$b$ block asymptotics. In order to get a nuisance-parameter-free limiting distribution under small-$b$ asymptotics, we replace $\hat \sigma$ by $\hat \kappa$ in equation (ref). The small-$b$ pseudo $t$-statistic is given as

align*[align* omitted — 253 chars of source]

The factor $v_T$ is defined in Lemma (ref). Since $v_T \to 2/3$, this term provides a finite-sample correction and scales the asymptotic variance of the $t$-statistic to unity. Under fixed-$b$ asymptotics, a nuisance term appears in the Gaussian process itself. By means of transforming the data with its inverse variance profile, cavaliere2007 showed that the time-transformation in the Gaussian limiting processes can be inverted. The variance profile estimator $\hat{\eta}(s)$ is strictly increasing and admits the unique inverse function $\hat{\eta}^{-1}(s)$. Accordingly, we consider the time-transformed series $\tilde{y}_t = y_{\lfloor \hat{\eta}^{-1}(t/T) T \rfloor}$ for $t=1, \ldots, T$. We replace the original series in the test statistic by $\tilde y_t$ and define

align*[align* omitted — 276 chars of source]

which yields the fixed-$b$ statistic

align*[align* omitted — 306 chars of source]

In practice, the time time-transformed series $\tilde{y}_t$ can have duplicate entries in low volatility periods and therefore may not include all information of the original series in high volatility periods. However, we do not need to discard any observations when transforming the data. We may artificially extend the series. An auxiliary sample size $\widetilde T \geq T$ can be chosen in such a way that $\hat{\eta}^{-1}(t/\widetilde{T}) - \hat{\eta}^{-1}((t-1)/\widetilde{T}) \geq \widetilde T^{-1}$ for all $t = 1, \ldots, \widetilde T$. Then, the grid of width $1/\widetilde T$ is dense enough such that $\tilde{y}_t = y_{\lfloor \hat{\eta}^{-1}(t/\widetilde{T}) \widetilde{T} \rfloor}$, $t=1, \ldots, \widetilde{T}$, includes all sample points of the original series, and the fixed-$b$ statistic may be applied to this auxiliary series. Note that the auxiliary time series is not necessary from a theoretical point of view, but it leads to better test results in small samples.

theoremLet $\rho = 1-c/\sqrt{BT}$ with $c \geq 0$, let $d_t$ satisfy Assumption (ref), and let $u_t$ satisfy Assumption (ref). \begin{itemize} • Let $B/T \to 0$ as $B,T \to \infty$. Then, \begin{align*} \tau-SB \overset{d}{\longrightarrow} \mathcal{N} \bigg( - \frac{c \sqrt 3}{2} \frac{\int_0^1 \sigma^2(r) \,\mathrm{d} r }{\sqrt{ \int_0^1 \sigma^4(r) \,\mathrm{d} r }}, \ 1 \bigg). \end{align*} • Let $0 < b < 1$, and let $B/T \to b$ as $B,T \to \infty$. Then, \begin{align*} \tau-FB \overset{d}{\longrightarrow} \frac{ \int_0^{1-b} \left( J_{c,b}(b+r) - J_{c,b}(r) \right)^2 \,\mathrm{d} r - b(1-b)}{2 \sqrt{ b \int_0^{1-b} \int_r^{b+r} \left( J_{c,b}(s) - J_{c,b}(r) \right)^2 \,\mathrm{d} s \,\mathrm{d} r} }, \end{align*} where $J_{c,b}(r) = \int_0^r e^{-(r-s)c/b} dW(s)$ is a standard Ornstein-Uhlenbeck process. \end{itemize}

The unit root hypothesis is rejected in favor of stationarity if the test statistic is smaller than the $\alpha$-quantile of the limiting distribution for the case $c=0$, where $\alpha$ is the significance level. For $\tau$-SB we can rely on standard normal quantiles as critical values. The limiting distribution of $\tau$-FB is nonstandard. Note that $J_c(r) = W(r)$ if $c = 0$. Table (ref) presents simulated left-tailed quantiles of the null distribution for various relative blocklengths $B/T$ and significance levels.

table[table omitted — 1,458 chars of source]

From the point of view of a practitioner, the $\tau$-SB test has a number of advantages: the distribution is standard normal; thus, there is no need to resort to new tables, and p-values are easy to implement. In fact, the simulations in Section (ref) indicate that the standard normal approximation is quite accurate in small samples if $B = \Theta(T^\gamma)$, where $0.5 \leq \gamma \leq 0.8$. Furthermore, the unit root test is robust to heteroskedasticity without using any data modification method such as those in cavaliere2007 and beare2018 or wild bootstrap implementations (see cavaliere2008bootstrap).

Testing under short-run dynamics

A more realistic scenario for macroeconomic variables is that error terms are serially correlated. We impose the following assumption on the error process:

assumption[serially correlated errors] The process $\{u_t\}_{t \in \mathbb{Z}}$ possesses the moving average representation $u_t = \psi(L) \epsilon_t = \sum_{i=0}^\infty \psi_i \epsilon_{t-i}$ with $\sum_{i=0}^\infty |\psi_i| < \infty$, where $L$ is the usual lag operator. Moreover, all solutions $z$ of the equation $\psi(z) = 0$ satisfy $|z| > 1$. The process $\{\epsilon_t\}_{t \in \mathbb{Z}}$ is independently distributed with $E[\epsilon_t]= 0$, $E[\epsilon_t^2] = \sigma_t^2$ and $E[\epsilon_t^4] < \infty$, where $\sigma_t = \sigma(t/T)$. The function $\sigma(r)$ is c\`{a}dl\`{a}g, non-stochastic, strictly positive, and bounded.

Assumption (ref) implies that the moving average representation of $u_t$ is invertible, and we may write $\theta(L) u_t = u_t - \sum_{i=1}^\infty \theta_i u_{t-i} = \epsilon_t$, where $\theta(z) = 1 - \sum_{i=1}^\infty \theta_i z^i$, and $\sum_{i=1}^\infty |\theta_i| < \infty$. In order to correct for the effect of short-run dynamics, we follow breitung2005, among others, and consider the pre-whitened series $x_t^* = \theta(L) x_t$. By equation (ref), it follows that

align*[align* omitted — 92 chars of source]

where $\epsilon_t$ satisfies the same conditions as $u_t$ under Assumption (ref). Consequently, if the unit root statistics are defined in terms of

align*[align* omitted — 229 chars of source]

instead of $\mathcal X_{1,T}$ and $\mathcal X_{2,T}$, their limiting distributions coincide with those presented in the previous sections.

Since the autoregressive parameters of the error process are unknown, they need to be estimated. In the fashion of said1984 and chang2002, we fix some lag order $p_T$ and consider the AR($p_T$) error representation $u_t = \sum_{i=1}^{p_T} \theta_i u_{t-i} + \epsilon_{p_T,t}$ with $\epsilon_{p_T,t} = \sum_{i=p_T+1}^\infty \theta_i u_{t-i} + \epsilon_t$. Then,

align[align omitted — 126 chars of source]

which is equal to $\sum_{i=1}^{p_T} \theta_i \Delta x_{t-i} + \epsilon_{p_T,T}$ under the unit root hypothesis. The lag order $p_T$ is allowed to grow with the sample size $T$. In what follows, we show that the differenced deterministic terms are asymptotically negligible, as $p_T \to \infty$ with $p_T=o(B^{1/2})$, and we may replace $\Delta x_{t-i}$ by $\Delta y_{t-i}$ for all $i \geq 0$ in the augmented regression equation. Let $(\hat \varphi$, $\hat \theta_1, \ldots, \hat \theta_{p_T})'$ be the least squares coefficient vector from the regression of $\Delta y_t$ on $y_{t-1}, \Delta y_{t-1}, \ldots, \Delta y_{t-{p_T}}$, for $t = p_T+1. \ldots, T$.

lemmaLet $\rho = 1 - c/\sqrt{BT}$ with $c \geq 0$, let $d_t$ satisfy Assumption (ref), and let $u_t$ satisfy Assumption (ref). Then, $\sum_{i=1}^{p_T} (\hat \theta_i - \theta_i) = O_P(p_T B^{-1/2})$, as $p_T,B,T \to \infty$.

The estimated pre-whitened series is defined as $\hat y_{t}^* = y_t - \sum_{i=1}^{p_T} \hat \theta_i y_{t-i}$, and the corresponding numerator and denominator statistics are given by

align*[align* omitted — 268 chars of source]
lemmaLet $\rho = 1 - c/\sqrt{BT}$ with $c \geq 0$, let $d_t$ satisfy Assumption (ref), and let $u_t$ satisfy Assumption (ref). Then, $\hat{\mathcal Y}_{1,T}^* - \mathcal X_{1,T}^* = O_P(p_T B^{-1/2})$, and $\hat{\mathcal Y}_{2,T}^* - \mathcal X_{2,T}^* = O_P(p_T T^{-1/2})$, as $p_T, B,T \to \infty$.

As a direct consequence, $(\hat{\mathcal Y}_{1,T}^* - \mathcal X_{1,T}^*, \hat{\mathcal Y}_{2,T}^* - \mathcal X_{2,T}^*) \overset{p}{\longrightarrow} (0, 0)$ if $p_T=o(B^{1/2})$. Let $\hat \rho^*$ be given by $\sqrt{BT}(\hat \rho^* - 1) = \hat{\mathcal Y}_{1,T}^*/\hat{\mathcal Y}_{2,T}^*$ and let the pre-whitened residuals be defined as $\hat u_t^* = \hat y_t^* - \hat \rho^* \hat y_{t-1}^*$, for $t = p_T+1, \ldots, T$. For notational convenience, let $\hat u_1^* = \ldots = \hat u_{p_T}^* = 0$. The pre-whitened counterparts of the estimators from Lemma (ref) are defined as

align*[align* omitted — 810 chars of source]

Analogously, we consider the time-transformed pre-whitened series $\tilde{y}_t^* = \hat y_{\lfloor \hat{\eta}^{*-1}(t/T) T \rfloor}^*$ for all $t=1, \ldots, T$, where $\hat{\eta}^{*-1}(s)$ is the unique inverse of $\hat \eta^*(s)$, and we define

align*[align* omitted — 290 chars of source]

For any lag order $p_T\geq 0$, the pre-whitened versions of the test statistics are given by

align*[align* omitted — 233 chars of source]

Note that $\tau\text{-SB}_0 = \tau\text{-SB}$ and $\tau\text{-FB}_0 = \tau\text{-FB}$. To summarize, we obtain the following limiting distributions:

theoremLet $\rho = 1- c/\sqrt{BT}$, let $d_t$ satisfy Assumption (ref), and let $u_t$ satisfy Assumption (ref). Furthermore, let $p_T = o(B^{1/2})$. \begin{itemize} • Let $B/T \to 0$ as $B,T \to \infty$. Then, $\hat \kappa^{*2} \overset{p}{\longrightarrow} \int_0^1 \sigma^4(r) \,\mathrm{d} r / \int_0^1 \sigma^2(r) \,\mathrm{d} r$, and \begin{align*} \tau-SB_{p_T} \overset{d}{\longrightarrow} \mathcal{N} \bigg( - \frac{c \sqrt 3}{2} \frac{\int_0^1 \sigma^2(r) \,\mathrm{d} r }{\sqrt{ \int_0^1 \sigma^4(r) \,\mathrm{d} r }}, \ 1 \bigg). \end{align*} • Let $0 < b < 1$, and let $B/T \to b$ as $B,T \to \infty$. Then, $\sup_{r \in [0,1]} |\hat \eta(s) - \eta(s)| \overset{p}{\longrightarrow} 0$, $\hat \sigma^{*2} \overset{p}{\longrightarrow} \int_0^1 \sigma^2(r) \,\mathrm{d} r$, and \begin{align*} \tau-FB_{p_T} \overset{d}{\longrightarrow} \frac{ \int_0^{1-b} \left( J_{c,b}(b+r) - J_{c,b}(r) \right)^2 \,\mathrm{d} r - b(1-b)}{2 \sqrt{ b \int_0^{1-b} \int_r^{b+r} \left( J_{c,b}(s) - J_{c,b}(r) \right)^2 \,\mathrm{d} s \,\mathrm{d} r} }, \end{align*} where $J_{c,b}(r) = \int_0^r e^{-(r-s)c/b} dW(s)$. \end{itemize}

The lag order $p_T$ is typically unknown in practice and can be chosen using conventional lag order selection methods, such as the Bayesian information criterion (BIC) or by the general-to-specific methodology in the fashion of ng1995. The maximum lag order $p_{max}$ can be chosen for instance by the rule of thumb provided by schwert1989. For the special case of a single break in the deterministic component, demetrescu2016 showed that if $p_T$ is determined by a usual information criterion the correct lag length is selected asymptotically.

Simulations

In this section, the finite sample performance of the unit root tests is evaluated by means of Monte Carlo simulations. The analysis includes different specifications for both the deterministic part $d_t$ and the stochastic part $x_t$.

table[table omitted — 1,108 chars of source]
figure[figure omitted — 326 chars of source]

While the zero-trend $d_t = 0$ is the main benchmark, we consider several other trends including sharp breaks and smooth changes of different shapes. The trend specifications are presented in Table (ref) and Figure (ref). The parameter $\lambda$ determines the size of the break. Similar trend functions are also considered in jones2014 in order to evaluate the performance of the unit root test by enders2012a.

The stochastic part $x_t$ is simulated both under the null hypothesis $\rho = 1$ and the alternative hypothesis $\rho = 0.9$. For the errors $u_t$, we consider an independent process as well as the AR(1) process $u_t = 0.5 u_{t-1} + \epsilon_t$ with standard normal innovations. Furthermore, results with heteroskedastic innovations using the variance function $\sigma^2(r) = 1+\lambda \cdot 1_{\{ r\leq 2/3 \}}$ are presented.

The small-$b$ tests are implemented using blocklengths of the form $B=T^\gamma$ with parameters $\gamma \in \{ 0.5, 0.6, 0.7, 0.8 \}$. For the fixed-$b$ versions, we consider $B = b \cdot T$ with relative blocklengths $b \in \{0.2, 0.4, 0.6\}$. For all tests, the lag augmentation order $p_T$ is either fixed or flexibly determined by the BIC with a maximum lag order of $p_{max} = 5$. All empirical size levels are presented for a significance level of 5%, and the models are simulated with 100,000 repetitions for sample sizes of $T=100$ and $T=300$. As noted by muller2003, the power of a unit root test depends on the initial condition, and the initial value is simulated as $x_0 \sim \mathcal{N}(0,\sigma_0^2)$ for $\sigma_0^2 \in \{0, 5, 10\}$.

In order to demonstrate the advantage of the fixed-$b$ and small-$b$ unit root tests, their finite sample results are compared to those obtained by conventional unit root tests. As the main benchmark, we consider the augmented Dickey-Fuller test by said1984 with constant trend specification (ADF henceforth), which is the $t$-test for the hypothesis $\phi = 0$ in the regression $\Delta y_t = \phi y_{t-1} + \beta_0 + \sum_{i=1}^{p_T} \xi_i \Delta y_{t-i} + e_t$.

elliott1996 proposed a feasible point-optimal test with local-to-unity GLS demeaning in the ADF regression. Let the deterministic trend function be given by the vector $z_t$, and let $\alpha^* = 1 - \overline c/T$, where $\overline c \in \mathbb{R}$. Furthermore, let $y_{\overline c, t} = y_t - \alpha^* y_{t-1}$ and $Z_{\overline c, t} = z_t - \alpha^* z_{t-1}$ for $t \geq 2$, and let $y_{\overline c, 1} = y_1$ and $Z_{\overline c, 1} = z_1$. The Dickey-Fuller GLS test is then the $t$-test for the hypothesis $\phi = 0$ in the regression $\Delta y_t^d = \phi y_{t-1}^d + \sum_{i=1}^{p_T} \xi_i \Delta y_{t-i}^d + e_t$, where $y_t^d = y_t - \hat \beta' z_t$ and where $\hat \beta$ is the OLS estimator from a regression of $y_{\overline c, t}$ on $Z_{\overline c, t}$. For the constant trend specification (DF-GLS henceforth), we set $z_t = 1$ and $\overline c = 7$, and, for the linear trend specification (DF-GLS-trend henceforth), $z_t = (1,t)'$ and $\overline c = 13.5$ are considered. Note that the point-optimal test with GLS demeaning is asymptotically equivalent with the Dickey-Fuller test for $d_t = 0$ computed using the series with initial value subtraction (see elliott1996)

An approach that does not assume a precise model for the trend component is that developed by enders2012a (EL henceforth). A flexible Fourier form is used to approximate smooth breaks in the trend function. Structural changes can be captured by the low frequency components of a series. In its simplest form, enders2012a considered the parametric trend model $d(r) = \alpha_0 + \gamma r + \alpha_1 \sin(2 \pi r) + \beta_1 \cos(2 \pi r)$. More frequencies could be included, but doing so could lead to an over-fitting problem. The test works as follows: First, the auxiliary regression $\Delta y_t = \delta_0 + \delta_{1} \Delta \sin(2 \pi t/T) + \delta_{2} \Delta \cos(2 \pi t/T) + v_t$ is considered with OLS estimates $\widehat{\delta}_0$, $\widehat{\delta}_{1}$, and $\widehat{\delta}_{2}$. Let $\widetilde D_t = \hat \delta_0 t + \hat \delta_1 \sin(2 \pi t/T) + \hat \delta_2 \cos (2 \pi t / T)$, which yields the detrended series $\widetilde{S}_t = y_t - \widetilde D_t - (y_1 - \widetilde D_1)$. Finally, the test statistic is given by the $t$-statistic for the null hypothesis $\phi = 0$ in the regression $\Delta y_t = \phi \widetilde{S}_{t-1} + \beta_0 + \beta_{1} \Delta \sin(2 \pi t/T) + \beta_{2} \Delta \cos(2 \pi t/T) + \sum_{i=1}^{p_T} \xi_i \Delta \widetilde{S}_{t-i} + e_t$.

harvey2005, harvey2006 showed that, if $x_0 \sim \mathcal{N}(0, \sigma_\alpha^2/(1-\rho^2))$ for $\rho = 1-c/T$ with $c > 0$ and some $\sigma_\alpha > 0$, the limiting distributions of the ADF and the DF-GLS test depend on the additional nuisance parameter $\sigma_\alpha$. The DF-GLS test is optimal for the zero initial condition $x_0 = 0$, but its power decreases monotonically in $\sigma_\alpha$, while the power of the ADF test increases. Figure (ref) indicates that the pooled tests are less sensitive to this effect across different values of $\sigma_\alpha$. Furthermore, there is no test that outperforms the other tests uniformly across $\sigma_\alpha$ for this situation in terms of size-adjusted power.

figure[figure omitted — 701 chars of source]
table[table omitted — 5,635 chars of source]
table[table omitted — 6,395 chars of source]
table[table omitted — 6,387 chars of source]
table[table omitted — 6,010 chars of source]

Tables (ref)--(ref) present size and actual power results under different model specifications. For smaller sample sizes, the pooled tests have small size distortions, which become larger as the break gets larger. However, for larger sample sizes, the size distortions decline. Overall, the size levels are similar to those obtained from using the conventional unit root tests.

The power of the pooled tests depends on the blocklength. In case of no break, a larger blocklength implies higher power results, which is in line with the theoretical findings that those tests have power in a $1/\sqrt{BT}$ neighborhood of the unit root hypothesis. For blocklengths of $B=T^{0.8}$ in the small-$b$ case and $B = 0.6 T$ in the fixed-$b$ case, the power results are similar to those from the ADF test and the Dickey-Fuller GLS test, where the ordering depends on the initial condition (cf.\ Figure (ref)). Hence, none of the tests dominates the pooled tests uniformly across these small-sample specifications (although, asymptotically, those tests have power in a $1/T$ neighborhood of the unit root hypothesis). Furthermore, smaller blocklengths, such as $T^{0.6}$ in the small-$b$ context and $0.2 T$ in the fixed-$b$ context, still yield reasonably high power. In particular, the EL test performs much worse in all cases. The size and power results obtained under the AR(1) error specification with both fixed and flexible lag augmentation for the pre-whitening scheme are similar to those produced by i.i.d.\ errors.

table[table omitted — 2,930 chars of source]

As the tests are designed to yield higher power in the presence of slowly varying trends and breaks, we compare the size-adjusted powers of the tests under the trend specifications presented in Table (ref) and Figure (ref). For large break sizes $\lambda$, it is shown that the smaller the blocklength, the greater the power results. In most cases, the pooled tests have greater power than the ADF, the DF-GLS, the DF-GLS-trend, and the EL test. Furthermore, the power results of the pooled tests are quite uniform across different trend specifications when compared to those of the conventional tests.

Table (ref) shows that the pooled tests have reasonable size and power properties under the presence of AR(1) errors and different trend specifications. Furthermore, from Table (ref), we can conclude that the tests are sized correctly and have good power properties in the presence of a break in the variance and in the trend function.

The blocklength $B$ is a tuning parameter that needs to be chosen carefully, and any optimality result would depend on the actual trend model. In practice, however, the trend model is unknown, which makes it hard to derive an optimal blocklength. Although theoretical recommendations cannot be formulated based on the current analysis, the small-$b$ tests with $B = T^{0.7}$ and the fixed-$b$ tests with $T = 0.2 B$ yield very promising results for all trend functions studied in this paper and are therefore recommended as the default settings.

Conclusion

We have presented two variants of a unit root test under an unknown trend specification that are robust under both heteroskedasticity and autocorrelation. When applied to finite samples, the tests show good size properties. The fixed-$b$ pooled test statistic converges to a functional of a Brownian motion under the unit root hypothesis, while the small-$b$ variant shows a standard normal distribution in the limit. Autocorrelation-robust versions of the tests were introduced using a pre-whitening scheme. Monte Carlo simulations indicate that, while under the zero-trend specification, the fixed-$b$ and small-$b$ tests perform similar to the conventional tests in terms of size and power, under sharp breaks as well as smooth changes in the trend, their power is much higher. Furthermore, the powers of the tests are less sensitive to the initial value when compared to the augmented Dickey-Fuller test and the Dickey-Fuller GLS test.

Acknowledgements

I would like to thank J�rg Breitung and Matei Demetrescu for their extensive advice and support. My thanks also go to Hans Manner, Markus K�sler, Robinson Kruse-Becher, Dominik Wied, Nazarii Salish, Uwe Hassler, Martin Wagner, the Co-Editor, and two anonymous referees for their helpful comments. The suggestions made by participants attending the 2015 RMSE meeting in Cologne, the SMYE conference 2017 in Halle (Saale), the SNDE conference 2017 in Paris, and the IAAE conference 2017 in Sapporo are also highly appreciated. Furthermore, the usage of the CHEOPS HPC cluster for parallel computing and a conference grant of the International Association for Applied Econometrics are greatfully acknowledged.

Supporting Information

An accompanying R-package for the application of the tests proposed in this article is available online at https://github.com/ottosven/urtrend.

\addcontentsline{toc}{section}{References}