EconBase
← Back to paper

Bootstrapping Nonstationary Autoregressive Processes with Predictive Regression Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

99,074 characters · 25 sections · 118 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

BOOTSTRAPPING NONSTATIONARY AUTOREGRESSIVE PROCESSES WITH PREDICTIVE REGRESSION MODELS BY CHRISTIS KATSOURIS University of Southampton and University of Exeter

\pagenumbering{roman}

abstractWe establish the asymptotic validity of the bootstrap-based IVX estimator proposed by phillipsmagdal2009econometric for the predictive regression model parameter based on a local-to-unity specification of the autoregressive coefficient which covers both nearly nonstationary and nearly stationary processes. A mixed Gaussian limit distribution is obtained for the bootstrap-based IVX estimator. The statistical validity of the theoretical results are illustrated by Monte Carlo experiments for various statistical inference problems.

\setcounter{page}{1} \pagenumbering{arabic}

Introduction.

Consider the first-order autoregressive process $\left\{ X_t \right\}_{t=1}^{+ \infty}$, defined by the following recursive process

align[align omitted — 91 chars of source]

where $\left\{ \varepsilon_t \right\}$ are independent $\mathcal{N} (0,1)$ random sequences. The least squares estimator $\hat{ \vartheta}_n$ of $\vartheta$, based on a sample of $n$ observations $\left\{ X_1,...,X_n \right\}$ is given by the following expression

align[align omitted — 116 chars of source]

A well known asymptotic result is that $\hat{\vartheta}_n$ is a consistent estimator for $\vartheta_n$ for all values of $\vartheta_n \in \left( - \infty, + \infty \right)$. Specifically, the asymptotic distribution of $\hat{\vartheta}_n$ depends on restrictions imposed on the admissible parameter space of $\vartheta$. In particular, in the stable case, which implies that the parameter space of $\vartheta$ takes values within the unit circle, that is, $| \vartheta | < 1$, various seminal studies such as mann1943statistical and anderson1959asymptotic among others, have shown that

align[align omitted — 167 chars of source]

When the true parameter lies on the boundary of the parameter space, such that $| \vartheta | = 1$, it has been shown by white1958limiting and phillips1987towards that the following limiting distribution holds

align[align omitted — 193 chars of source]

where $W(r)$ for some $0 \leq r \leq 1$ is a standard Wiener process within the probability space $\left( \Omega, \mathbb{P}, \mathcal{F}_t \right)$. On the other hand, when the true parameter of the autoregressive process given by expression (ref) is outside the unit circle (explosive parameter region), such that $| \vartheta | > 1$, it has been shown by white1958limiting and anderson1959asymptotic that the following asymptotic result holds

align[align omitted — 172 chars of source]

The particular limiting distribution discontinuity with respect to the parameter space of the autoregressive model of order one has been extensively discussed in the literature (see, bercu2001large, jeganathan1991asymptotic, phillips2007limit, proia2020moderate, benke2021nearly). Without loss of generality, considering for each $n \in \mathbb{N}$, the statistical experiment $\mathcal{E}_n$ corresponding to the observations $\left\{ X_0, ..., X_n \right\}$, then the sequence $\left( \mathcal{E}_n \right)_{ n \in \mathbb{N} }$ is locally asymptotically normal (LAN) when $| \vartheta | < 1$, it is locally asymptotically Brownian functional (LABF) when $| \vartheta | = 1$, and it is locally asymptotically mixed normal (LAMN) when $| \vartheta | > 1$ (see, le1986asymptotic, le2000asymptotics and phillips1989partially). Thus, statistical inference procedures can crucially depend on the true value of the population parameter $\theta_n$ in certain regions of the parameter space.

Motivated by these observations various studies have focus on the development of estimation and inference methodologies for nonstationary time series models (see, rao1978asymptotic, dickey1979distribution, chan1987asymptotic1, chan1988limiting, phillips1987towards, phillips1987time, phillips1988regression, phillips1988weak, phillips1988testing, abadir1993limiting, larsson1995asymptotic, buchmann2007asymptotic and phillips2007limit and horvath2003bootstrap). Specifically, the class of nearly nonstationary time series models as proposed by the seminal studies of phillips1987towards, phillips1987time who developed asymptotics for near-integrated time series provide a unified approach to inference. In practise, expressing the autoregressive coefficient with the local-unit-root specification, such that, $\vartheta_n = \big( 1 + c / n \big)$, where $n$ the sample size and $c$ denotes the nuisance parameter of persistence, allows to encompass such moderate deviations from the unit boundary (e.g., unit root, near-integrated or explosive, see stoykov2019least). Our main research interest is to investigate the finite and large sample properties of bootstrap-based estimators and test statistics (e.g., moreno2012unit) for predictive regression models when regressors are generated as nonstationary autoregressive processes.

Predictive regression models are commonly used for modelling macroeconomic data which exhibit high persistence (see, zhu2014predictive). Furthermore, the nonstandard nature of the inference problem in predictive regressions has been previously discussed in the literature such as by the studies of Phillips1990statistical, cavanagh1995inference and jansson2006optimal, with the main challenge being the robustness of inference methodologies to the nuisance parameter of persistence (see, astill2023bonferroni)). A unified framework for robust estimation and inference regardless of the abstract degree of persistence has been proposed by phillipsmagdal2009econometric and kostakis2015Robust (see, also chen2013uniform, breitung2015instrumental, li2017uniform, liu2019unified, yang2021unified and jayetileke2021predictive). We argue that the implementation of resampling methodologies for nonstationary time series models provides a way for conducting inference, when non-pivotal quantities are involved, which can ensure accurate asymptotic approximations (see more recently demetrescu2022extensions). Thus, establishing the asymptotic validity (in the spirit of cavaliere2020inference) of these bootstrapped approximations is crucial in other testing problems as well such as when considering structural break (see, georgiev2018testing, katsouris2022partial, katsouris2023testing, katsouris2023structural and fei2023ivx).

A commonly used resampling methodology used to study the distributional properties of a statistic of interest is based on the well known Efron's bootstrap method as in efron1986bootstrap. Various studies consider the asymptotic validity of the bootstrap-based estimator in (ref). More specifically, the validity of explosive autoregressive processes is studied by basawa1989bootstrapping, basawa1991bootstrapping while in general asymptotics for the bootstrap in stationary models is predominately concerned with independent observations (see, freedman1981bootstrapping, freedman1984bootstrapping). Although it has been initially documented in the literature that the bootstrap cannot work for dependent processes (see, bose1988edgeworth) it was anticipated that it would work if the dependence is taken care of while resampling (see, politis1994stationary and paparoditis2003residual, paparoditis2005bootstrapping). Therefore, the dependence structure of innovations found in predictive regression models has to be correctly implemented when considering a bootstrap procedure.

Our framework considers that the dependence structure is preserved due to the joint Gaussianity assumption imposed on the error sequences of the predictive regression model along with the nonstationary autoregressive process. In particular, motivated by the limit theory of cointegrating regressions the bootstrap implementation is adjusted to capture the dependence induced by the linear process representation of the innovation sequence (see, li1997bootstrapping, li2001bootstrapping, psaradakis2001bootstrap, psaradakis2001bootstrapURs and inoue2002bootstrapping). Furthermore, the optimization method employed can affect the limiting distribution of related test statistics. The main quantity of interest is the risk of the normalized error, such that, $r_n \left( \hat{\theta}_n - \theta \right)$ of an estimator $\hat{\theta}_n$ with a suitable slow varying rate function $r_n$ where $r_n \to \infty$ as $n \to \infty$. The exact form of the particular rate function depends on the properties of the estimator under consideration; such as $\sqrt{n}-$consistency is only ensured for the OLS estimator while a different rate of convergence holds for the IVX estimator. Consequently, deriving probability bounds for the specific risk quantity using the Berry-Essen approach, as well as for its bootstrapped counterpart, in order to assess the efficiency of finite-sample approximations depend on the chosen estimation methodology\footnote{Related statistical theory is discussed in the books of tanaka2017time and hamilton2020time as well as in the studies of hall2014martingale and billingsley2013convergence. Moreover, bickel1981some presents some asymptotic theory aspects for the bootstrap while mykland1992asymptotic proposes relevant asymptotic expansions.}.

In general, by the classical CLT, the distribution of a pivotal quantity

align[align omitted — 57 chars of source]

tends weakly to $\mathcal{N}(0,1)$. Let $F_n$ be the empirical distribution of $X_1,..., X_n$, putting mass $1/n$ on each $X_t$. When one considers the bootstrap resampling, the $n$ data points $X_1,..., X_n$ are treated as a population, with distribution function $F_n$ and mean $\mu_n$, while $\mu_n^*$ is considered as an estimator of $\mu_n$. In practise, the idea is that the behaviour of the quantity $Q_n^*$ mimics that of $Q_n$. Thus, the distribution of $Q_n^*$ could be computed from the data and used to approximate the unknown sampling distribution of $Q_n$ (see, praestgaard1993exchangeably).

This paper develops limit theory for the asymptotic validity of the bootstrapped based IVX estimator in predictive regression models that satisfies the regularity conditions of valid instrumental based estimators such as relevance and orthogonality while preserving the robust properties to the nuisance parameter of persistence. To the best of our knowledge this is a novel implementation and further establishes the robust properties of the IVX estimator for bootstrap-based inference.

Preliminary Theory.

Concurrent with our work, and independently, georgiev2021extensions also develop asymptotic results for the bootstrap test based on the IVX estimator who focus on predictability tests based on the sub-sample approach. Another key point here is that a comparison of the asymptotic behaviour of various bootstrapping procedures especially for the case of predictive regression models with regressors of abstract degree of persistence will be essential to obtain insights regarding the accuracy of bootstrap approximations. Notice that the IVX estimator is a consistent estimator of the predictive regression model coefficient although it has a different convergence rate in comparison to the OLS estimator. Therefore, the bootstrap which is well known to deliver asymptotic refinements over first-order asymptotic approximations as well as bias corrections can be applied to solve the asymptotic bias problem (see cavaliere2022bootstrap for the counterargument). On the other hand, any additional bias term that appears in the asymptotic expansion reduces the precision of the test.

Consider the partial-sum process below:

align[align omitted — 75 chars of source]

then it holds that

align[align omitted — 141 chars of source]

where $W$ is the standard Brownian motion.

Moreover this convergence suggests that $X_n$ should converge to a solution of the limiting stochastic differential equation. Thus, a key step in the application of the stochastic differential equation approach is to show that the sequence of stochastic integrals in the approximating equation converges to the corresponding stochastic integral in the limit\footnote{A specific example is the case in which one considers the random weighted bootstrap. In this case it is more interesting to study the Marcinkiewicz-Zygmund strong law for randomly weighted sums of dependent random variables under dependent random designs. }. Let $D$ be the space of cadlag functions $f: [0,1] \to \mathbb{R}$ equipped with the Skorokhod topology. As a measurable structure on $D$ we consider the corresponding Borel $\sigma-$algebra $\mathcal{B}(D)$. We write $X_n \overset{d}{\to} X$ whenever $X_n$, X are random variables taking values in $\big( D, \mathcal{B} (D) \big)$ such that $X_n$ converges weakly to $X$ in $D$ as $n \to \infty$, that is, $\mathsf{lim}_{ n \to \infty } \mathbb{E} \big[ f (X_n) \big] \to \mathbb{E} \big[ f(X) \big]$ for all bounded continuous functions $f: D \to \mathbb{R}$.

We have that

align[align omitted — 122 chars of source]

Therefore, it holds that

align[align omitted — 71 chars of source]

We define the Ornstein-Uhlenbeck process $J_c = \big( J_c(t) \big)_{ t \in [0,1] }$ driven by $J$ by setting the following stochastic expression

align[align omitted — 101 chars of source]

The usefulness of the bootstrap resampling method is exactly to provide a way to ensure uniform size that is robust to the presence of effects (such as conditional heteroscedasticity) that can affect sufficient control of the type of errors when conducting hypothesis testing. Moreover, exact calculation of the critical value of the test statistic is approximated based on Monte Carlo simulations. Due to the computational burden of the bootstrap procedure we use relative small sample sizes. However, by taking a sufficiently large number of Monte Carlo samples the approximation error can be made arbitrary small thus ensuring that the bootstrap approximation will work well given regularity conditions hold that we discuss in more details below.

Consider a test statistic $\mathcal{T}_n$ of a parameter estimator such that $\mathcal{T}_n = \mathsf{\eta}(n) \left( \hat{\theta}_n - \theta_0 \right)$. Let $\mathcal{T}^{*}_n$ denote the bootstrap version of $\mathcal{T}_n$ such that $\mathcal{T}^{*}_n = \mathsf{\eta}(n) \left( \hat{\theta}^{*}_n - \theta_0 \right)$. Let $\hat{F}_n (u) := \mathbb{P}^{*} \left( \mathcal{T}^{*}_n \leq u \right), u \in \mathbb{R}$ denote the distribution function conditional on the original pair of data.The bootstrap $\mathsf{p-value}$ is defined as below:

align[align omitted — 123 chars of source]

First-order asymptotic validity of $\hat{p}^{*}_n$ requires that $\hat{p}_n$ converges in distribution to a standard uniform distribution such that $\hat{p}_n \overset{d}{\to} \mathsf{Unif} [0,1]$.

\paragraph{Minimax Optimality} A complementary aim of this paper is to show that the sufficient conditions for the error bounds are indeed necessary in some applications. Let us define a test $\phi$, which is a Borel measurable map, $\phi: \mathcal{X}_n : \mapsto \left\{ 0, 1 \right\}$. Then, for a class of null distributions $\mathcal{P}_0$, we denote the set of all level $\alpha$ tests by

align[align omitted — 163 chars of source]

We focus on the minimax framework to evaluate the performance of test statistics and model estimators which enables us to study the fundamental limits of hypothesis testing while ensuring a (strong) uniform guarantee over a large class of (null and alternative) distributions. The notion of minimax performance is widely used to quantify the difficulty of a statistical problem. More specifically, in the hypothesis testing literature, it is also common to study the performance of tests against fixed or directional alternatives such as cases in which $\beta_n = \beta_0 + \frac{\delta}{\sqrt{n}}$ for some $\delta > 0$. For example, the problem of severe empirical size distortions when $\rho$ is large it has been also documented in the case of structural break tests in dynamic econometric models (see, kuan1994implementing). Therefore, this implies that we are likely to obtain false rejection of constant coefficients in dynamic models when variables have high persistence and endogeneity.

Outline of the paper.

The rest of the paper is organized as follows. In Section (ref), we introduce the predictive regression model and model assumptions. We motivate the implementation of the IV based estimator which is robust to the nuisance parameters of persistence. In Section (ref), we prove the asymptotic validity of the bootstrap IVX estimator and verify that we obtain an identical limit distribution as in the standard case. Section (ref), provides a Monte Carlo Simulation study. In particular, the asymptotic properties of the different bootstrap proposals under the null and under the alternative are investigated and their relative power performance is analyzed. Section (ref), concludes.

Notation.

Throughout the paper, all limits are taken as $n \to \infty$, where $n$ is the sample size. Then, the symbol $"\Rightarrow"$ is used to denote the weak convergence of the associated probability measures as $n \to \infty$ (as defined in billingsley2013convergence) which implies convergence of sequences of random vectors of continuous c\`adl\`ag functions on $\mathcal{D} \left( [0,1] \right)$ equipped with the Skorokhod topology. The symbol $\overset{d}{\to}$ denotes convergence in distribution and $\overset{p}{\to}$ denotes convergence in probability within a suitable probability space $( \Omega, \mathcal{F}_t, \mathbb{P} )$. Finally, $\mathcal{O}(a_n)$ denotes the rate of convergence of estimators, such that if $\theta_n = \mathcal{O}(a_n)$, implies that, $\theta_n / a_n$ is bounded. Also we denote with $\mathbb{P}^{*}$, $\mathbb{E}^{*}$ and $\mathsf{Var}^{*}$ the probability measure, expected value and variance operators induced by the bootstrap resampling procedure conditional on the time series observations of the original sample.

The model, main assumptions and estimators.

In this section we introduce the model under consideration in this paper, the main modelling assumptions as well as two different estimation methodologies proposed in the literature.

Predictive regression model.

Consider the predictive regression model given by

align[align omitted — 182 chars of source]

Then, the regressor $X_t$ of the predictive regression model is assumed to be generated via the following process

align[align omitted — 157 chars of source]

where $\upbeta_n \in \mathbb{R}$ is the (model) parameter of interest while the autoregressive coefficient $\varrho_n \in \mathbb{R}$ is defined such that $\varrho_n = \big( 1 + \frac{ \displaystyle c }{ \displaystyle n^{\gamma} } \big)$ with $c \in ( - \infty , + \infty)$ and $\gamma \in (0,1]$.

Notice that $v_t$ is assumed to be a zero mean, stationary and ergodic process with finite autocovariances given by $\tilde{\sigma} := \mathbb{E} \left( v_t v_{t-j} \right)$ and $\omega^2 = \sum_{j = - \infty}^{ \infty } \tilde{\sigma}$ is finite and nonzero. Furthermore, the nuisance coefficients $c$ and $\gamma$ are considered as tuning parameters which determine the degree of persistence for the regressors and therefore without loss of generality the particular novel feature proposed by phillips1987towards, phillips1987time provides a tractable way to consider the nature of possible nonstatioary regressors when estimating predictive regressions. In terms of distributional assumptions, we assume that the predictive regression model is generated from a bivariate Gaussian distribution, such that $e_t = \left( u_t, v_t \right)^{\prime} \overset{ \textit{i.i.d} }{ \sim } \ \mathcal{N} \big( 0, \boldsymbol{\Sigma} \big)$ with $t \in \left\{ 1,...,n \right\}$. Further related theoretical results which apply to the innovation sequence of the system can be found in phillips1992asymptotics.

assumptionLet $\left\{ \epsilon_t \right\}_{ t \in \mathbb{Z} }$ be a sequence of i.i.d continuous random variables with $\mathbb{E} \left( \epsilon_1 \right) = 0$ and $\mathbb{E} \left( \epsilon^2_1 \right) = 1$, and with a characteristic function $\phi(\tau)$. Then, the regressor $X_t$ of the predictive regression model is assumed to be generated via the following process \begin{align} X_t = \varrho_n X_{t-1} + v_t, \ \ \ X_0 = 0, \ \ \varrho_n = \left( 1 + \frac{ c }{ n^{\gamma} } \right), \ \ \ t \in \left\{ 1,...,n \right\}, \end{align} with $\gamma = 1$, where $\kappa$ is a real constant such that $c \in ( - \infty, 0)$ and $v_t = \sum_{j=0}^{\infty} c_j \epsilon_{t-j}$, with $\sum_{j=0}^{ \infty } c_j \neq 0$ and the summability condition $\sum_{j=0}^{ \infty } j^{1+ \delta} | c_j | < \infty$ holds for some $\delta > 0$.
assumption(i) Let $\left\{ u_t, \mathcal{F}_{t} \right\}_{ t \geq 1 }$, where $\mathcal{F}_{t}$ is a sequence of increasing $\sigma-$fields which is independent of $\epsilon_j$, $j > t $, forms a martingale difference sequence satisfying $\mathbb{E} \big( u^2_{t} | \mathcal{F}_{t-1} \big) \to \sigma^2 > 0$, as $t \to \infty$ and $\underset{ t \geq 1 }{ \mathsf{sup} } \ \mathbb{E} \big( | u_{t}|^4 \big| \mathcal{F}_{t-1} \big) < \infty$. (ii) $X_t$ is adapted to $\mathcal{F}_t$, and there exists a correlated Brownian motion $\big( W, V \big)$ such that the following joint weakly convergence holds \begin{align} \left( \frac{1}{\sqrt{n} } \sum_{j=1}^{[nr]} \epsilon_{j-1} , \frac{1}{\sqrt{n} } \sum_{j=1}^{ [nr] } u_{j} \right) \Rightarrow \big( W(r), V(r) \big) \end{align} on $\mathcal{D}_{ \mathbb{R}^2 } \left( [0,1] \right)^2$ as $n \to \infty$.

Define the partial sum process $V_n(r)$ and $U_n(r)$ as below

align[align omitted — 131 chars of source]
remarkAssumption (ref) and (ref) provide necessary conditions for the development of the asymptotic theory (similarly in Wang2012specification) for the corresponding bootstrap IV based estimator we are aiming to examine its statistical validity. In terms of the estimation procedure for the parameter of the predictive regression model we consider both the least squares estimation as well as the endogenous instrumentation method proposed by phillipsmagdal2009econometric.

Main Results

We are interested to examine the asymptotic behaviour for the $\hat{\upbeta}_n$ estimator, in the case of both least square estimation as well as the two-step endogenous instrumental variable estimation approach followed for constructing the IVX filtration. To derive the asymptotic distribution of both estimators we consider the Ornstein-Uhlenbeck (OU) process

align[align omitted — 145 chars of source]

which satisfies the stochastic differential equation $d J_{c}(r) = c J_{c}(r) dr + dW(r)$. Then, the following asymptotic results hold, based on the local to unity limit law proposed by phillips1987time, phillips1987towards such that $\frac{X_{[nr]}}{\sqrt{n} } \Rightarrow J_{c}(r)$, where $J_{c}(r)$ is the OU process defined in expression (ref).

align[align omitted — 261 chars of source]
remarkThe convergence of the sample moments to their stochastic integral counterparts are proved in the seminal study of phillips1988regression who also derived invariance principles for the multivariate model. In our study we consider the corresponding univariate limit results since we consider the univariate predictive regression model with persistent regressors. Furthermore, we consider the bootstrap invariance principle to obtain the weak convergence of the sample moments for the bootstrapped-based test statistics.

The ordinary least squares estimate of $\upbeta_n$ denoted with $\hat{\upbeta}_n^{ols}$ is given by

align[align omitted — 119 chars of source]

Therefore, since $\sqrt{n} \left( \hat{\upbeta}_n^{ols} - \upbeta \right) = \mathcal{O}_p( \frac{1}{\sqrt{n}} )$ then it converges in distribution which implies

align[align omitted — 292 chars of source]
remarkThe limit distribution given by (ref) shows that the existence of the nuisance parameter $c$ in the underline stochastic process, when the parameter $\upbeta_n$ is obtained based on the least squares estimation weakly convergence to a nonstandard limit. On the other hand this is not the case for the corresponding IVX estimator as we demonstrate in the next section.

IVX estimator

In this section we discuss the main methodology employed for the estimating the nonstationary time series model. Specifically, the IVX instrumentation proposed by phillipsmagdal2009econometric provides important advantages rather than using the conventional differencing approach prior to fitting the model. Although employing the difference operator is a common approach for ensuring covariance stationarity in time series observations due to the fact that it can remove trends and seasonality it can also contribute to loss of information such as the long-memory properties\footnote{In particular, kasparis2015nonparametric consider how the interplay of the long-memory properties of innovations and the persistence properties of regressors in a nonparametric predictive regression affect the limiting distributions of test statistics and corresponding model estimators.} of these observations which are eliminated when differencing is applied. Therefore, the IVX estimator in contrast to its OLS counterpart is found to be robust to the persistence properties of the regressors as captured by the autoregressive specification without altering the persistence properties of regressors. The particular estimation approach can accommodate both nearly nonstationary and nearly stationary processes, that is, generalizing this way the integration order of regressors for the predictive regression.

Consider the instrumental variable $Z_{tn}$ which is based on the IVX methodology of phillipsmagdal2009econometric. Then, the IVX instrumental variable is constructed

align[align omitted — 188 chars of source]

Denote with $\hat{\upbeta}_n^{ivx}$ the corresponding IVX based estimator for the parameter of the predictive regression given by the system of equations in expressions (ref) and (ref). Then, standard IV estimation arguments provide an expression for the IVX estimator

align[align omitted — 127 chars of source]

Thus, we consider the limit distribution of the random variable $\psi_n$ as defined below

align[align omitted — 283 chars of source]
lemmaUnder Assumption (ref) and (ref), the following invariance principle holds \begin{align} \sum_{t=1}^{ [nr] } Z_{t-1} X_{t-1} \Rightarrow - \frac{1}{c_z} \left( r \omega_{xx}^2 + \int_0^r J_{c}(s) dJ_{c}(s) \right), \ \ with \ \ r \in [0,1]. \end{align}
remarkRelated weakly convergence arguments which can be employed for the proof of Lemma (ref) can be found in phillipsmagdal2009econometric. The proof of the particular result consists of a key limit result for deriving the asymptotic distribution of the bootstrap based IVX estimator which is the main goal of this paper.

A simple implementation of the limit given by Lemma (ref) after appropriate normalization shows that the following nuisance-free limiting distribution holds

align[align omitted — 186 chars of source]

where $\psi$ follows a mixed Gaussian distribution. The mixed Gaussian property of the $\psi_n$ random variate, is instrumental in developing inference procedures robust to abstract degree of persistence (see, phillips2013predictive, phillips2016robust).

Therefore, it remains to determine the components of the stochastic covariance matrix $\tilde{ \boldsymbol{\Sigma} }$. More precisely, to determine the covariance matrix of the random variable $\psi$, we consider the martingale difference sequence $\xi_{nt} := \left( \frac{1}{ n^{ ( 1 + \upgamma_z ) / 2 } } Z_{t-1} u_t , \frac{1}{ \sqrt{n} } v_t \right)^{ \prime }$. Then the martingale conditional variance of $\xi_{nt}$ has the following structure

align[align omitted — 510 chars of source]
remarkNote that as shown by phillips2007limit $n^{ - \left( 1 + \upgamma_z / 2 \right)} \sum_{t=1}^n z_{t-1} = o_p(1)$, which implies that the martingale conditional covariance (ref) has off-diagonal elements which converge in probability to zero. Furthermore, note that without loss of generality, we operate under the assumption that $\upgamma_x = 1$ and $c \in (- \infty, 0)$. Relaxing the parameter space of these two quantities allows to consider regressors generated from general nonstationary autoregressive processes, covering both near nonstationary and near stationary processes. Nevertheless, the robust property of the IVX estimator implies that regardless of the persistence properties of regressors asymptotically converges to a pivotal mixed Gaussian distribution. Therefore, the particular large sample property result to conventional statistical inference when considering Wald type statistics after suitable normalization (see, phillips2013predictive, phillips2016robust).

Notice that it has been shown that when the residual process $u_t$ is correlated with the regressor $x_t$, the limiting distribution of the $t-$ratio is no longer normal (see, phillips1986multiple). As a remedy to the particular bias, the fully modified regression proposed by Phillips1990statistical as an alternative way of constructing the test statistic as we present later. Overall a key component of our framework is that it establishes the invariance principle which concerns the weak convergence (within the Shorokhod topology) of the bootstrap partial sum process to Brownian motion. Then, based on the continuous mapping theorem, it can be used to obtain asymptotic distributions of various bootstrapped statistics without making parametric assumptions on the underlying model. In particular this approach is build upon the strong approximation and the Beveridge-Nelson representation of linear processes. The invariance principles for the $\textit{i.i.d}$ innovation and the bootstrapped counterpart are first developed using the strong approximation of the partial sum process by the standard Brownian motion, and subsequently the invariance principles for the general linear process with $\textit{i.i.d}$ innovations and the corresponding bootstrapped process are established by their Beveridge-Nelson representations as in phillips1992asymptotics.

Let the partial sum process of $( \varepsilon_t )$ be defined by

align[align omitted — 90 chars of source]

Then, we have that $W_n \overset{ d }{ \to } W$ in the space of $\mathcal{D} [0,1]$ of cadlag functions, where $W$ is the standard Brownian motion. The space $\mathcal{D} [0,1]$ is equipped with the Skorokhod topology.

In particular for the bootstrap, we first obtain or estimate $\left\{ \varepsilon_t \right\}_{t=1}^n$ from the sample of size $n$ and get $\left\{ \hat{\varepsilon}_t \right\}_{t=1}^n$ . then, we resample from the empirical distribution of $\left\{ \hat{\varepsilon}_t \right\}_{t=1}^n$, that is, the distribution with point probability mass $1 /n$ based on the time series observations of size $n$, to get the bootstrap sample $\left\{ \varepsilon_t^{*} \right\}_{t=1}^n$ as the i.i.d samples from the empirical distribution of $\left\{ \hat{\varepsilon}_t \right\}_{t=1}^n$. Both $( \hat{\varepsilon}_t )$ and $(\varepsilon_t^{*})$ are dependent upon the sample size $n$, and we may more precisely denote them as triangular arrays $( \hat{\varepsilon}_{nt} )$ and $(\varepsilon_{nt}^{*})$.

Furthermore, certain studies have previously suggested the inconsistency of the bootstrap estimator when the parameter is on the boundary of the parameter space (see, andrews2000inconsistency). However, notice within our setting the particular consideration has a different interpretation as such a phenomenon occurs via the local unit root process. In other words, the local-to-unity specification provides a natural bound on these type of processes when modelling dependent time series and therefore the asymptotic behaviour of estimators based on predictive regression models are analytically tractable. Therefore, when considering bootstrapping nonstationary autoregressive processes using predictive regression models, in practise we are mainly concern about deriving an identical weakly convergence argument similar to the non-bootstrapped conventional limit result.

Our main goal is to develop a consistent bootstrap estimator of the distribution of the stochastic quantity $\psi_n := n^{ \frac{ 1 + \upgamma_z }{2} } \left( \hat{\upbeta}_n^{ivx} - \upbeta_n \right)$, regardless of the persistence properties of the regressors of the predictive regression model as captured by the autoregressive process with a LUR coefficient. More precisely, under the assumption that the regressors are generated by the nonstationary process $X_t = \varrho_n X_{t-1} + v_t$ with $\varrho_n = \left( 1 - \frac{ \kappa }{ n^{\upgamma_x} } \right)$ in Assumption (ref), then we aim to investigate whether $\psi_n$ and $\psi_n^{\star} := n^{ \frac{ 1 + \upgamma_z }{2} } \left( \hat{\upbeta}_n^{\star ivx} - \upbeta_n \right)$ converge to the same distribution limit within a suitable probability space. To do this, we consider all features of the distribution of $\psi_n$ such as the covariance matrix in (ref).

The Bootstrap and Hypothesis testing.

A commonly used approach to evaluate the limiting distributions in nonstationary time series models based on asymptotic approximations relies on finite samples bootstrapped samples. The exact type of the bootstrap procedure follwed to replicate the dependence structure of the model can affect both the validity of the results as well as the accuracy of those approximations. In this section, we study the asymptotic behaviour of the bootstrap based IVX estimator for different bootstrap schemes. To do this, we impose additional assumptions for the associated partial sum processes. Let $\mathcal{D}$ be the space of cadlag functions, $f : [0,1] \to \mathbb{R}$ equipped with the Shorokhod topology (e.g., see csorgHo2003donsker).

\paragraph{Notation} Denote with $\mathcal{Z}_{nt}^{*}$ a sequence of bootstrap statistics, then the following modes of convergence hold (see related definitions in goncalves2015Bootstrap)

itemize$\mathcal{Z}_{t}^{*} = o_{ P^{*} }(1)$ in probability, or $\mathcal{Z}_{t}^{*} \overset{ P^{*} }{ \to } 0$ in probability, if for any \begin{align} \epsilon > 0, \delta > 0 \ \ \underset{ n \to \infty }{ \mathsf{lim} } \mathbb{P} \big[ \mathbb{P}^{*} \big( \left| \mathcal{Z}_{n}^{*} \right| > \delta \big) > \epsilon \big] = 0. \end{align} • $\mathcal{Z}_{t}^{*} = \mathcal{O}_{ P^{*} }(1)$ in probability, if for all $\epsilon > 0$ there exists a $M_{\epsilon} < \infty$ such that \begin{align} \underset{ n \to \infty }{ \mathsf{lim} } \mathbb{P} \big[ \mathbb{P}^{*} \big( \left| \mathcal{Z}_{n}^{*} \right| > M_{\epsilon} \big) > \epsilon \big] = 0. \end{align} • $\mathcal{Z}_{n}^{*} \overset{ d^{*} }{ \to } \mathcal{Z}$ in probability if, conditional on the sample, $\mathcal{Z}_{n}^{*}$, weakly converges to $\mathcal{Z}$ under $P^{*}$, for all samples contained in a set with probability converging to one. Specifically, we write $\mathcal{Z}_{n}^{*} \overset{ d^{*} }{ \to } \mathcal{Z}$ in probability if and only if $\mathbb{E}^{*} \big( f \left( \mathcal{Z}_{n}^{*} \right) \big) \to \mathbb{E} \left( f(\mathcal{Z}) \right)$ in probability for any bounded and uniformly continuous function $f$.

Therefore, for the asymptotic validity of the bootstrap algorithm one needs to consider similar arguments as when deriving the limiting distribution of nonstationary time series models. In particular, muller2011efficient discusses various aspects of asymptotic efficiency with respect of the weak convergence of functionals to their brownian motion counterparts.

Bootstrap Procedure.

Similar investigation in the literature with respect to the bootstrap procedure include the study of davidson2008wild who consider a framework for the Wild bootstrap although not in a nonstationary time series regression environment. In this paper we apply the Wild bootstrap algorithm using the OLS as well as the IVX residuals. Another example include the study of phillips1990asymptotic who develop asymptotic theory for residual-based tests of cointegration. Next, we describe the estimation procedure for the bootstrap based IVX estimator, which can help us shed light on the specific asymptotic expressions and properties we need to consider in order to derive its limiting distribution.

Consider again the predictive regression with multiple predictors. Simulate the following DGP

align[align omitted — 146 chars of source]

where the parameters of interest are $\upbeta_n$ and $\varrho_n$. Denote with $u_t = \left( u_{t}, v_{t} \right)^{\prime}$ the martingale difference sequence vector with variance-covariance matrix given by

align[align omitted — 146 chars of source]

where $\boldsymbol{\Sigma}$ is a known positive-definite matrix.

Specifically, we are motivated in establishing the asymptotic validity of the bootstrap for the IVX estimator since its is employed in statistical problems for which the limiting distribution is non-standard such as when testing for structural breaks in predictive regression models and the bootstrap is necessary for deriving critical values. While for instance, we focus in the case of IVX-based predictability tests, the limit results we derive in this paper can be generalized for the supremum IVX-Wald statistic when testing for parameter instability in predictive regression models. The first Bootstrap algorithm we consider is described below.

Bootstrap Algorithm

itemize• Estimation Step \begin{itemize} • Fit the predictive regression to the sample data $\left( Y_t, X_{t-1} \right)^{\prime}$ to obtain the OLS residuals $\hat{u}_t$, for $t \in \left\{ 1,..., n \right\}$. Obtain the estimate of $\hat{\upbeta}_n$. • Fit by OLS the AR(1) model with the LUR coefficient, to the regressor $X_t$, to obtain the OLS residuals $\hat{v}_t$, for $t \in \left\{ 1,..., n \right\}$. Set $\hat{v}_t = 0$. Obtain $\hat{\varrho}$. \end{itemize}

In this section, we assume that the autoregressive coefficient is given by $\rho_n = \left( 1 - \frac{\kappa}{n} \right)$, which allows to examine the limit theory of the estimate of the autoregressive coefficient for moderate deviations from unity. Note that the estimate of $\varrho_n$ is obtained by ordinary least squares estimation which is expressed as

align[align omitted — 122 chars of source]
itemize• Generate the bootstrap sample $\left\{ Y_t^{*}, X_t^{*}, t = 1,...,n \right\}$. \begin{itemize} • Generate the bootstrap innovations $w_t^{*} = \left( u_{t}^{*}, v_{t}^{*\prime} \right)^{\prime} := \left( e_t \otimes \hat{u}_t, e_t \otimes \hat{v}_t \right)$, where the symbol $\otimes$ denotes element-wise vector multiplication, such that $e_t$ for $t \in \left\{ 1,..., n \right\}$, is a scalar sequence independent from the data such that $e_t \overset{ \textit{i.i.d} }{ \sim } \mathcal{N}(0,1)$. • Generate $X_t^{*}$ such that \begin{align} X_{t_b}^{*} = \hat{\varrho} X_{t_b -1}^{*} + u_{t_b}^{*}, \ \ \ \ \ for \ t_b \in \left\{ 1,...,n \right\}, \end{align} with initial conditions $X_0^{*} = 0$ and $b \in \left\{ 1,...,B \right\}$. • Generate the associated bootstrap IVX instrument $z_t^{*}$ as below \begin{align} Z_0^{*} = 0 \ \ \ and \ \ \ Z_t^{*} = \sum_{j=0}^{t_b-1} \left( 1 + \frac{\kappa_z}{n} \right)^j \Delta X_{t_b-j}^{*} , \ \ for \ \ t \in \left\{ 1,...,n \right\}, \end{align} where $\kappa_z$ is the coefficient matrix for the persistence of the bootstrap IVX instruments that contains the same values as the original IVX instruments. • Generate $Y_{ t_b }^{*}$ such that $y_{ t_b }^{*} = \hat{\upbeta} X_{ t_b -1 }^{*} + u_{ t_b }^{*}$. Set $Y_1^{*} = Y_1$. • Repeat Steps 2.1 to 2.4 $B$ times (e.g., $B = 1000$). \end{itemize}
remarkNotice that the asymptotic validity of the bootstrap IVX estimator is an important property of general interest with applications in various inference problems. For example, katsouris2022partial considers a framework for structural break testing in predictive regression models with persistent regressors. In the particular framework the bootstrap validity of the IVX estimator is necessary to ensure the consistent estimation of critical values from the bootstrap limiting distribution of the corresponding structural break test statistics. Furthermore, georgiev2021extensions study the implementation of both the residual wild bootstrap as well as the fixed regressor wild bootstrap for the IVX estimator which is useful when constructing predictability tests (e.g., t-tests or Wald type statistics) based on the predictive regression model.

Then, for the estimation step one can consider estimating $\beta_n$ either via the classical least squares estimation, denoted with $\hat{\beta}_n^{LS}$, or via the IVX instrumentation methodology proposed by phillipsmagdal2009econometric, denoted with $\hat{\beta}_n^{IVX}$. The least squares estimate of $\beta_n$ is

align[align omitted — 116 chars of source]

The corresponding IVX estimate of $\beta_n$ is

align[align omitted — 125 chars of source]

where

align[align omitted — 195 chars of source]

We use for example, $\kappa_z = 1$ and $\gamma_z = 0.95$. \

From Step1A and Step1B, we can obtain the residual sequences, $ \left\{ \hat{u}_t^{IVX}, \hat{v}_t^{LS} \right\}_{t=1}^n$. To simplify the notation, for the remaining of this Section, we simply denote with $\hat{u}_t$ the corresponding residual sequence obtained via the IVX instrumentation procedure and $\hat{v}_t$ the residual sequence obtained via Least squares estimation. That is,

align[align omitted — 74 chars of source]

Next, we define the centered residuals for both sequences, such that

align[align omitted — 135 chars of source]

Denote with $\tilde{F}_n$ the empirical distribution function based on $\left\{ \tilde{e}_t \right\}_{t=1}^n$ where $\tilde{e}_t = \left( \tilde{u}_t, \tilde{v}_t \right)^{\prime}$. Thus, $\tilde{F}_n$ associates mass $n^{-1}$ to each of $\tilde{e}_t, t =1,....,n$. Now, assuming that $\tilde{F}_n$ is the true distribution, draw a random sample $\left\{ e^{*}_t, t = 1,...,n \right\}$ from $\tilde{F}_n$. Therefore, conditionally on $\left( X_1,..., X_{n-1} \right)$, the random variables $\left\{ e^{*}_t, t = 1,...,n \right\}$ are i.i.d with distribution function $\tilde{F}_n$.

Step 2: (Bootstrap Step)

We can now construct the bootstrap sample $\left\{ X_t^{*}, t = 1,...,n \right\}$ recursively via the following expression

align[align omitted — 79 chars of source]

with $X_0 = 0$, where $\hat{\rho}_n$ it has been estimated in Step 1 by Least squares estimation.

Next, we construct the bootstrap sample $\left\{ Y_t^{*}, t = 1,...,n \right\}$ recursively via the following expression

align[align omitted — 80 chars of source]

where $\hat{\beta}_n$ it has been estimated in Step 1 by the IVX instrumentation methodology.

Next, we re-estimate the model parameter of interest, denoted with $\hat{\hat{\beta}}_n$, by applying again the IVX instrumentation on the generated sequence of $\left\{ Y_t^{*}, X_t^{*} \right\}$, from Step 1. That is,

align[align omitted — 144 chars of source]

where

align[align omitted — 207 chars of source]

We use for example, the same values as in Step 1, that is, $\kappa_z = 1$ and $\gamma_z = 0.95$. \

Step 3: (and onwards) \

For the remaining bootstrap steps, that is, $b = \left\{ 1,...,B \right\}$ we generate a sequence of IVX estimates that is $\underline{\beta}^{*IVX} = \left( \beta^{*IVX}_{n1}, ...., \beta^{*IVX}_{nB} \right)$.

The main aim of this Section, is to derive the limit distribution of $\underline{\beta}^{*IVX}$ given $\left( X_1, ... X_n \right)$ and to show that it is the same as the original limit distribution of the IVX estimator. Notice that in practise the bootstrap procedure should be able to capture the time series dependence structure. It is worth mentioning that we need to employ some autoregressive robust covariance matrix estimation methodology when applying our proposed bootstrap algorithm to ensure that the dependence structure of the nonstationary time series model is preserved which safeguards the consistent estimation of the Wald-type statistic (jansson2004error and jansson2012nearly).

Asymptotic distribution result.

We consider a triangular array of paired random variables such that $\left\{ X_{k,n}, Y_{k,n} \right\}$ with $n \geq 1$ and $k \geq 1$ satisfying

align[align omitted — 156 chars of source]

where $e_k = \left( u_k, v_k \right)^{\prime} \sim \ \mathcal{N} \big( 0, \boldsymbol{\Sigma} \big)$. Then, we construct the following triangular array using endogenous information generated by the triangular array $\left\{ X_{k,n} \right\}_{n, k \geq 1}$ given by

align[align omitted — 222 chars of source]

We define the following random sequence

align[align omitted — 127 chars of source]

where $\hat{\upbeta}_n^{ivx}$ is the IVX estimator of $\upbeta_n$. Therefore, the random measure $\mathcal{J}_{n}$ can be shown to have a limiting distribution which is given by the following expression

align[align omitted — 191 chars of source]

where $\mathcal{U} \left( . \right)$ is Brownian motion with variance $\sigma^2_{uu} \tilde{V}$ with $\tilde{V} \equiv - \omega^2_{xx} \big/ 2 \kappa_z$, for $\upgamma_x > \upgamma_z $ since $\upgamma_z \in (0,1)$. Then, the bootstrap estimator $\hat{\upbeta}_n^{*ivx}$ of $\upbeta_n$ is obtained from the corresponding bootstrap sequence $\left\{ Y_t^{*}, X_t^{*} \right\}$ which implies the following random sequence

align[align omitted — 139 chars of source]

where

align[align omitted — 211 chars of source]

Therefore, in practise we need to show that $\mathcal{J}_{n}$ and $\mathcal{J}_{n}^{*}$ have the same limit distribution (we define this step as Condition 1). To do this, we consider the random measure $\mathcal{\zeta}_n$ that corresponds to the paired triangular array (ref)-(ref) and the generated instrument is given in expression (ref). sweeting1989conditional proposed a framework for evaluating conditional weak convergence. Specifically, when considering the asymptotic validity of the bootstrap algorithm it is necessary to consider the conditional weak convergence based on the bootstrapped invariance principles. On the other hand, the dependence structure of the predictive regression model remains fixed therefore the invariance principles based for the probability space that corresponds to the bootstrapped observations is valid due to similar conditions.

Notice also that although the practitioner needs to choose the nuisance parameter of persistence for the instrumental variable. In other words, the statistics of interest such as the model estimator as well as the Wald test are calculated based on the following expression:

align[align omitted — 220 chars of source]

Following a similar approach in the literature (to add references), we need to specify an auxiliary functional of the form $\Psi(.)$ such that

align[align omitted — 101 chars of source]

so that we can derive an associated result that shows a probability limit of the form

align[align omitted — 140 chars of source]

To do that, we can define for example with

align[align omitted — 162 chars of source]

which is taken to be a regular conditional probability distribution function.

remarkThe above statements imply that to show that the bootstrap approximation is valid, then we shall show that along almost all paths of $H_n$ given by expression (ref) converges in distribution to the distribution of $\mathcal{J}$ given by expression (ref) (we define this step as Condition 2). Therefore, to prove the result we are aiming we can try to prove that a corresponding weakly convergence result holds for the quantity $\mathcal{J}_n^{*}$.

Notice that in the literature for bootstrap based inference of time series models (see, basawa1991bootstrapping) the authors prove that the bootstrap estimator of the autoregressive coefficient in an AR(1) model in the explosive case will not have the same limiting distribution as the original estimator, within our framework this might not be affecting the result which we are aiming to prove. In other words, due to the structure of the predictive regression system, in each bootstrap step, to generate the pair $\left\{ X_t^{*}, Y_t^{*} \right\}$ we use the residuals $u_t^{*}$ obtained from the first step regression and then construct the observations $Y_t^{*}$ in the second step regression. Therefore, in practise since we are not interested whether or not the bootstrap estimator of the autoregressive coefficient (under various nonstationary processes) is valid or not, we can still use the particular residual sequences. Therefore, after fitting the predictive regression we can obtain the bootstrap IVX estimator conditional on the nonstationary process. More precisely, the statistical validity of the bootstrap residuals that correspond to the autoregressive process is proved to hold by paparoditis2003residual (see, Appendix (ref)).

Therefore, we can consider the covariance matrix of the random variable $\psi^{*}$, where

align[align omitted — 195 chars of source]

Therefore, we consider the martingale difference sequence $\xi^{*}_{nt} := \left( \frac{1}{ n^{ ( 1 + \upgamma_z ) / 2 } } Z^{*}_{t-1} u^{*}_t , \frac{1}{ \sqrt{n} } v^{*}_t \right)^{ \prime }$. \

Then the martingale conditional variance has the following structure

align[align omitted — 533 chars of source]
remarkIn this paper, in contrary to some related studies in the literature such as georgiev2018testing, cavaliere2020inference and georgiev2021extensions we do not consider the implementation of the fixed regressor bootstrap. The reason for this is simple, since we study estimators for the predictive regression model based on the paired data $\left\{ Y_t, X_t \right\}_{t=1}^n$, then keeping the regressors fixed in each bootstrap iteration asymptotically might not preserve the limiting distribution of the original pair.

Monte Carlo Simulations.

exampleWe consider the following simple data-generating process (DGP): \begin{align} y_t &= \beta_0 x_t + u_t, \ \ \ t = 1, 2,..., n \\ x_t &= \left( 1 - \frac{c}{n} \right) x_{t-1} + v_t, \ \ \ for some \ c > 0. \end{align} where $v_t \sim \mathcal{N} \left( 0, \sigma_v^2 \right)$ and $u_t = \rho u_{t-1} + \epsilon_t$ where $\epsilon_t \sim \mathcal{N} \left( 0, \sigma_{\epsilon}^2 \right)$. In practise, when $\rho = 0$ then $u_t = \epsilon_t$ is the i.i.d residual process.

The particular example can demonstrate our asymptotic results presented in this paper. Then the performance of our bootstrap procedure can be assessed using total variation distances applied to the empirical distribution function and the corresponding empirical distribution function of the residual process (see related limit theory in stroud1972fixed and mykland1992asymptotic). According to horowitz2003bootstrap bootstrap resampling is a method for estimating the distribution of an estimator or test statistic by resampling one's data or a model estimated from the data. Under conditions that hold in a variety of econometric applications, the bootstrap provides approximations to distributions of statistics, coverage probabilities of confidence intervals, and rejection probabilities of tests that are more accurate than the approximations of first-order asymptotic distribution theory.

Conclusion.

In this paper we contribute to the literature of bootstrap methodologies for nonstationary time series models. Specifically, we study the bootstrapping problem of nonstationary autoregressive processes with predictive regression models with near unit root (persistent) regressors. These are regressors which are near-integrated, that is, close to the unit boundary. We establish the asymptotic validity of the bootstrap based estimators and investigate the finite-sample behaviour of these bootstrap procedures by Monte Carlo experiments. The related theory to the bootstrap aspects of the paper is presented in the \hyperref[appn]{Appendix}. Moreover, the IVX instruments are constructed based on endogenous information therefore satisfying the relevance condition for valid instruments. Therefore, the role of the bootstrap is to mimic the dependence structure in the predictive regression model. A discussion regarding the relevance condition of instrumental variables can be found in the paper of hall1996judging.

Therefore, one of the applications of our methodology is to employ the bootstrap for nonstandard and nonpivotal testing problems. Generally a Wald-type statistic is a pivotal function, as occurs for instance when testing linear restrictions on the coefficients of linear regressions. On the other hand, when testing in the predictive regression model we employ the IVX-Wald test which is robust to the nuisance parameter of persistence. We have investigated two estimation methodologies when bootstrapping nonstationary autoregressive processes with predictive regression models for the purpose of obtaining the bootstrapped distribution of test statistics. On the other hand, bootstraps with exchangeable weights are not so wide spread in the econometrics and statistics literature. Thus considering the random weight bootstrap in the context of nonstationary autoregressive processes is a novel contribution to the literature.

As future research we aim to investigate the bootstrap validity of the framework that corresponds to the uniform inference methodology in predictive regression. In particular, the quasi restricted likelihood ratio test (QRLRT) shows to have desirabe properties. More precisely, the restricted likelihood has been found to provide a well-behaved likelihood ratio test in the predictive regression model even when the regressor exhibits almost unit root behaviour. Another interesting further research worth mentioning include to consider the bootstrap validity of additional persistence classes or different functional forms for the predictive regression model as proposed in the framework of duffy2021estimation. Our proposed bootstrap algorithm can be also employed in different settings as well such that within a sequential monitoring framework for structural breaks in Garch $(p,q)$ models (see, berkes2004sequential). Furthermore, although in this paper we consider the construction of Wald-type statistics based on linear restrictions other transformations can be also considered such as non-linear restrictions which are found to have better power performance (see also heimann1996bootstrapping).

small\paragraph{Acknowledgements} I wish to thank Jose Olmo, Tassos Magdalinos and Jean-Yves Pitarakis for helpful discussions during my PhD studies as well as Stathis Paparoditis and Giuseppe Cavaliere. The author of this article acknowledge the use of the IRIDIS High Performance Computing Facility and associated support services at the University of Southampton, in the completion of this work. Financial support from the VC PhD studentship of the University of Southampton is also gratefully acknowledged.
appendix\begin{center} APPENDIX \end{center}

Technical Proofs

(page S.23 in georgiev2021extensions)

We have that

align[align omitted — 244 chars of source]

It then follows by near-integration asymptotics that

align*[align* omitted — 304 chars of source]

Notice that it holds that $\left[ M_v \right]_{\pi} = \left[ J_c \right]_{\pi}$, that is, the two stochastic processes have the same quadratic variation. Furthermore, it holds that

align[align omitted — 120 chars of source]

which holds by the semi-martingale property of $J_c(r)$. Moreover, notice these stochastic processes are monotonically increasing with a continuous quadratic variation function $\left[ M_v \right]( \pi )$ on a compact set $\mathcal{B} (0,1)$, then it is sufficient to show that the asserted weakly convergence of the sample moments to the corresponding stochastic integrals holds point-wise in probability for $\pi \in [0,1]$.

Consider the proof from li2001bootstrapping. In particular, we are interested to show the invariance principles for the partial sums of $\varepsilon_t^{*}$ and $u_t^{*}$. Thus, first we show that

align[align omitted — 111 chars of source]

Let EDF denote the empirical distribution function $d_2(.,.)$ be the Mallows metric defined as $d_p( \nu, \mu) = \mathsf{inf} \ \mathbb{E} \left( \left\lVert V - U \right\rVert^p \right)^{ 1 / p}$. Thus, we denote with

align[align omitted — 259 chars of source]

Thus, we need to show that

align[align omitted — 39 chars of source]

Moreover, we need to establish convergence of finite-dimensional distributions and verify the tightness condition. In particular, for the finite-dimensional convergence, it is sufficient to show that for any finite $k$ and $0 < r_1 < ... < r_k < 1$, if

align[align omitted — 142 chars of source]

and

align[align omitted — 166 chars of source]

then $d_2 \left( V_k^{*}, V_k \right) \to 0$. Thus, it can be shown that $d_2 \left( V_k^{*}, V_k \right)^2 \leq d_2 \left( F_T, \tilde{F}_T \right)^2 \to 0$.

For tightness, it suffices to show that there exists a non-decreasing function $\psi$ such that, for almost all sample paths, for $0 \leq r_1 \leq r \leq r_2 \leq 1$,

align[align omitted — 99 chars of source]

From the paper: Testing for parameter instability in predictive regression models

remarkNotice that considering for example the test statistic $\mathcal{W}^{*}_{ivx}$ we observe that bootstrap invariance principles hold jointly along with stated convergence results. In practise, in terms of convergence in topological spaces, we have joint weak convergence of random measures. The concept is weaker than weak convergence in probability, although it reduces to the latter when the limit distribution is non-random.

Smoothing Local-to-moderate unit root theory.

Consider an autoregressive process with local-to-moderate deviations from unit root of the form $\rho_n = \left( 1 + \frac{c}{K} \right)$, where $K$ converges to infinity with the sample size $n$ and $\rho_n$ approaches unity from the stationary or the explosive side according to the sign of $c$ (see, Phillips2010smoothing).

We consider that such a time series process constitutes of $m$ blocks of $K$ observations with total sample size $n = mK$. Then, partitioning the chronological sequence $\left\{ t = 1,...,n \right\}$ by setting $t = [kj] + k $ for $k \in \left\{ 1,...,K \right\}$ and $j \in \left\{ 0,..., m-1 \right\}$, it is possible to study the asymptotic behaviour of the time series $\left\{ X_t : t = 1,...,n \right\}$ via the asymptotic properties of the time series $\left\{ X_{ Kj + k} : j = 0,..., m-1, k = 1,..., K \right\}$. The latter representation is particularly useful for revealing the transition from non-stationary to stationary autoregression as the number of blocks $m$ increases.

More formally, the process with the above characteristics may be written in the form

align[align omitted — 179 chars of source]

since we have $m$ blocks with $K$ observations in each block, giving us a total of $n = mK$.

Therefore, under this setting when $m=1$, then the usual local to unity model applies, and when $m \to \infty$, then the moderate deviation theory of PM and GP holds.

Let $W$ be a standard Brownian motion and $J_c(t) = \displaystyle \int_{0}^t e^{c(t-s)} dW(s)$ be a corresponding OU process. For each $m$, letting

align[align omitted — 71 chars of source]

we observe that W also a standard Brownian motion and we denote by $\widetilde{J}_c(t) = \int_0^t e^{c(t-s)} d \widetilde{W} (s)$ the associated OU process. Thus, for given $m \geq 1$, we may derive a limit theory for the least squares estimate $\hat{\rho}_{n,m}$ of $\rho_{n,m}$ as $n \to \infty$ using earlier results from standard local to unity asympotics.

Thus, using the identities above we obtain

align[align omitted — 184 chars of source]

the results derived in Phillips imply that, for fixed $m$ and $n \to \infty$, the asymptotic distribution of the least squares estimator takes the following form

align[align omitted — 287 chars of source]

When $c < 0$, sequential limits with $n \to \infty$ followed by $m \to \infty$ lead to the normal asymptotic theory given in PM and GP:

align[align omitted — 610 chars of source]

Bootstrapping Unstable Autoregressive Processes.

Theoretical Background.

In this section, we present related asymptotic theory examples as preliminary theory to the framework proposed in this paper.

Consider the first-order autoregressive process $\left\{ X_t \right\}$, for $t = 1,2,...$,

align[align omitted — 61 chars of source]

where $\left\{ \epsilon_t \right\}$ are independent $\mathcal{N} (0,1)$ random variables.

The least squares estimator $\hat{\beta}_n$ of $\beta$, based on a sample of $n$ observations $\left( X_1,..., X_n \right)$, is given by the following expression

align[align omitted — 129 chars of source]

Note that the asymptotic validity of the bootstrap estimator corresponding to $\hat{\beta}_n$ for the stationary case, viz., $| \beta| < 1$, follows from the work of bose1988edgeworth and the validity for the explosive case, viz., $|\beta | > 1$ has recently established by basawa1989bootstrapping. Both of these papers consider the general case when the distribution of $\left\{ \epsilon_t \right\}$ is not necessarily known. The limit distribution of $\hat{\beta}_n$ in the unstable case $| \beta | = 1$ is known to be nonnormal. It is therefore, of special interest to consider the bootstrap approximation for the distribution of $\hat{\beta}_n$ for the unstable case.

Invalidity of the bootstrap estimator.

\

Let $\xi_n = \left( \sum_{t=1}^n X^2_{t-1} \right)^{1/2} \left( \hat{\beta}_n - \beta \right)$, where $\hat{\beta}_n$ is defined as above. It is then, well known that when $\beta = 1$,

align[align omitted — 182 chars of source]

where $W(t)$ is a standard Wiener process. Then, the corresponding bootstrap sample $\left\{ X_t^{*} \right\}$ is obtained recursively from the following relation

align[align omitted — 83 chars of source]

where $\left\{ \epsilon_t^{*} \right\}$ constitutes a random sample from $\mathcal{N} \left( 0, 1 \right)$. Then, the bootstrap estimator $\hat{\beta}^{*}_n$ of $\beta$ is then defined as in (ref) with $X$'s replaced by $X^{*}$'s.

Let $\xi_n^{*} = \left( \sum_{t=1}^n X^{*2}_{t-1} \right)^{1/2} \left( \hat{\beta}^{*}_n - \beta \right)$ denote the bootstrap version of $\xi_n$. It will be shown that $\xi_n$ and $\xi_n^{*}$ do not have the same limit distribution, thus invalidating the bootstrap.

Therefore, we consider a triangular array $\left\{ X_{k,n}, k \leq 1, n \leq 1 \right\}$ satisfying

align[align omitted — 79 chars of source]

with independent $\epsilon_k \sim \mathcal{N} \left( 0, 1 \right)$ and where $\left\{ b_n \right\}$ is a sequence of numbers such that $n \left( b_n - 1 \right) \to \gamma$. Let

align[align omitted — 220 chars of source]

where $\left\{ W(t) : 0 \leq t \leq 1 \right\}$ is a standard Brownian motion and

align[align omitted — 98 chars of source]

Then, by Theorem 1 of chan1987asymptotic1 we have that

align[align omitted — 117 chars of source]

where

align[align omitted — 160 chars of source]

and where $\mathbb{P}_{b_n}$ indicates the probability distribution induced by the model in (ref). Define

align[align omitted — 129 chars of source]

which is taken to be a regular conditional probability distribution function. Therefore, we define a random measure

align[align omitted — 69 chars of source]

Since $H( \gamma, x )$ given by (ref) is continuous in $\gamma$ for each fixed $x$, we have

align[align omitted — 66 chars of source]
align[align omitted — 141 chars of source]

Therefore, if the bootstrap approximation were valid then along almost all paths $H_n$ given by (ref) would converge in distribution to the distribution of $\xi$ given in (ref). However, we have in fact that

align[align omitted — 89 chars of source]
proofBy almost sure representativeness of convergent laws, it is possible to define $\tilde{\beta}_n, n \leq 1$, and $\tilde{\xi}$ with $\hat{\beta}_n = \tilde{\beta}_n$, $\xi^{\prime} =_{d} \tilde{\xi}$ and \begin{align} n \left( \tilde{\beta}_n - 1 \right) \overset{ d }{ \to } \tilde{ \xi } \ \ almost surely \ as \ n \to \infty. \end{align} Recall that, \begin{align} n \left( \hat{\beta}_n - 1 \right) \overset{ d }{ \to } \xi^{\prime} \ as \ n \to \infty \end{align} Therefore, we have that \begin{align} H_n \left( \tilde{ \beta }_n, x \right) \to H ( \tilde{ \xi }, x ) \ \ almost surely \ as \ n \to \infty. \end{align} Hence for sets of the form \begin{align} A_j = \bigcup_{i=1}^{ m_j } \big( x_{ij}, y_{ij} \big) \end{align} representing a disjoint union of intervals, (ref) implies that \begin{align} \bigg( H_n \left( \tilde{ \beta }_n, A_1 \right),..., H_n \left( \tilde{ \beta }_n, A_k \right) \bigg) \to \bigg( H_n \left( \tilde{ \xi }, A_1 \right),..., H_n \left( \tilde{ \xi }, A_k \right) \bigg) \end{align} almost surely as $n \to \infty$. Since, \begin{align} \bigg( \eta_n ( A_1 ), .... , \eta_n ( A_k ) \bigg) =_{d} \bigg( H_n \left( \tilde{ \beta }_n, A_1 \right),..., H_n \left( \tilde{ \beta }_n, A_k \right) \bigg) \end{align} and \begin{align} \bigg( \eta ( A_1 ), .... , \eta ( A_k ) \bigg) =_{d} \bigg( H_n \left( \tilde{ \xi }, A_1 \right),..., H_n \left( \tilde{ \xi }, A_k \right) \bigg), \end{align} which implies that (ref) follows from (ref).

The above example, demonstrates that the least squares estimate of the autoregressive coefficient in the case of the unstable autoregression model ($\beta = 1$), induces an invalidity of the bootstrap procedure. A similar invalidity of the bootstrap procedure occurs when considering the distribution of $n \left( \hat{\beta}_n - \beta \right)$ as examined by chan1987asymptotic1. In this paper, we aim to demonstrate that in the case of predictive regression model, when we estimate the unknown parameter using the IVX instrumentation procedure, then the limiting distribution of the original estimate is preserved which verifies the validity of the corresponding IVX bootstrap estimate.

Bootstrapping Explosive Autoregressive Processes.

Preliminary results and the bootstrap estimate.

First, we review some basic limit results for the explosive case as in bose1988edgeworth. We assume that $| \beta | > 1$ and define with

align[align omitted — 149 chars of source]

where $\left\{ \epsilon_t \right\}$ are i.i.d random variables with $\mathbf{E} \left( \epsilon_j \right) = 0$ and Var$\left( \epsilon_j \right) = \sigma^2$. Moreover, it can be shown that $U_n$ and $V_n$ are identically distributed for each $n$ and there exist random variables $U$ and $V$ such that

align[align omitted — 96 chars of source]

where $U$ and $V$ are independent and identically distributed. If $\left\{ \epsilon_t \right\}$ are normal, $U$ and $V$ are normal each with mean 0 and variance $\left( 1 - \beta^{-2} \right)^{-1}$. In the general case, $U$ and $V$ can be represented by

align[align omitted — 93 chars of source]

Note that the common characteristic function of $U$ and $V$ is

align[align omitted — 103 chars of source]

where $\phi_{\epsilon}( . )$ is the characteristic function of $\epsilon_t$.

Let $\hat{ \beta}_n$ denote the least squares estimate of $\beta$. The following Theorem summarizes the limit distribution of $\hat{ \beta}_n$ for the case $| \beta | > 1$.

theoremFor $\hat{ \beta}_n$ defined in the autoregressive model, we have that for $| \beta | > 1$, \begin{align} \left( \beta^2 - 1 \right)^{-1} | \beta |^n \left( \hat{ \beta}_n - \beta \right) \to_{d} V \big/ U, \end{align} where $U$ and $V$ are defined in (ref).

We now describe the corresponding bootstrap estimate. Let $\hat{\epsilon}_t = X_t - \hat{ \beta}_n X_{t-1}$ and define $\tilde{\epsilon}_t = \hat{\epsilon}_t- n^{-1} \sum_{t=1}^n \hat{\epsilon}_t$, the centered residuals. Denote by $\tilde{F}_n$ the empirical distribution function based on $\left\{ \tilde{\epsilon}_t, t = 1,...,n \right\}$. Thus, $\tilde{F}_n$ associates mass $n^{-1}$ to each of $\tilde{\epsilon}_t, t =1,....,n$. Now, assuming that $\tilde{F}_n$ is the true distribution, draw a random sample $\left\{ \epsilon^{*}_t, t = 1,...,n \right\}$ from $\tilde{F}_n$. Therefore, conditionally on $\left( X_1,..., X_n \right)$, the random variables $\left\{ \epsilon^{*}_t, t = 1,...,n \right\}$ are i.i.d with distribution function $\tilde{F}_n$. This pseudo-series allow us to construct the bootstrap sample $\left\{ X_t^{*}, t = 1,...,n \right\}$ recursively by the following expression

align[align omitted — 87 chars of source]

with $X_0 = 0$. The bootstrap least squares estimate is then given by

align[align omitted — 126 chars of source]

The main aim of this section (see, also bose1988edgeworth), is to derive the limit distribution of $\hat{\beta}_n^{*}$ given $\left( X_1, X_2,... \right)$ and to show that it is the same as the limit distribution given in Theorem (ref). Before doing that, we present below a useful lemma for the purpose of proving the aforementioned result.

lemmaFor the autoregressive model with $\mathbf{E} \epsilon_t = 0$, Var $\epsilon_t = \sigma^2 < \infty$ and $| \beta | > 1$, we have \begin{align} \frac{1}{n} \sum_{t=1}^n \left( \tilde{ \epsilon}_t - \epsilon_t \right) \overset{ a.s }{ \to } 0, \ \ as \ n \to \infty. \end{align} Moreover, \begin{align} \hat{\beta}_n \overset{ a.s }{ \to } \beta \ \ as \ n \to \infty. \end{align}

In practise, we denote the empirical distribution of the centred residuals $\tilde{\epsilon}_t$ as $\tilde{F}_t$. Notice that the bootstrap innovations $\left\{ \epsilon_t^{*} \right\}$ are conditionally independent with common distribution $\tilde{F}_T$.

Limit distribution of the bootstrap estimate.

We now state the main result of this section.

theoremConditionally on $\left( X_1,..., X_n \right)$ as $n \to \infty$ we have, for $| \beta | > 1$, \begin{align} \left( \hat{ \beta }_n^2 - 1 \right)^{-1} | \hat{ \beta }_n |^{n} \left( \hat{ \beta }_n^{*} - \hat{ \beta }_n \right) \to_{d} V \big/ U \end{align} for almost all sample paths $\left( X_1, X_2, ... \right)$, where $U$ and $V$ are defined in (ref), and $\left( \hat{\beta}_n , \hat{ \beta}_n^{*} \right)$ are defined, above.
proofDefine, \begin{align} U_n^{*} = \sum_{t=1}^n \hat{\beta}^{ - (t - 1 )} \epsilon_t^{*} \ \ and \ \ V_n^{*} = \sum_{t=1}^n \hat{\beta}^{ - (n - t )} \epsilon_t^{*} \end{align} The result in the theorem can be deduced analogously to that of Theorem (ref) provided we show that, conditionally on $\left( X_1, ... X_n \right)$, $\left( U_n^{*}, V_n^{*} \right) \to_{d} \left( U, V \right)$ for almost all sample paths, where $U_n^{*}$, $V_n^{*}$, $U$ and $V$ are as defined above. This is due to the fact that the limit distributions of $\hat{ \beta}_n^{*}$ and $\hat{\beta}_n$ are determined respectively by those of $\left( U_n^{*}, V_n^{*} \right)$ and $\left( U, V \right)$. Thus, for the remaining derivations we shall show that the conditionally characteristic function of $\left( U_n^{*}, V_n^{*} \right)$ given $\left( X_1, ... X_n \right)$ converges to the characteristic function of $\left( U, V \right)$, for almost all sample paths $\left( X_1, X_2, ... \right)$. Let $\mathbf{E}^{*} (.)$ denote the expectation with respect to the distribution $\tilde{F}_n$ (of $\epsilon_t^{*}$ ) conditional on $\left( X_1, ... X_n \right)$. Let $\phi_{ U_n^{*} (\tau)} = \mathbf{E}^{*} \left( \text{exp} \ i\tau U_n^{*} \right)$. Thus, it is shown that \begin{align} \phi_{ U_n^{*} (\tau)} = \prod_{t=1}^n \mathbf{E}^{*} \left( exp \ i \tau \hat{\beta}_n^{-(t-1)} \epsilon_t^{*} \right) \end{align} converges a.s to $\phi_{ U ( \tau )}$ for each $\tau \in \mathbb{R}$. To accomplish this it suffices to show that \begin{align} \underset{ m \to \infty }{ lim } \ \underset{ n \leq m }{ sup } \sum_{t = m}^n \bigg| \mathbf{E}^{*} \left( exp \ i \tau \hat{\beta}_n^{-(t-1)} \epsilon_t^{*} \right) - 1 \bigg| = 0 \ \ \ a.s, \end{align} where the convergence in (ref) is uniform on compact sets in $\mathbb{R}$ and it holds that \begin{align} \underset{ n \to \infty }{ \text{lim} } \bigg| \mathbf{E}^{*} \left( \text{exp} \ i \tau \hat{\beta}_n^{-(t-1)} \epsilon_t^{*} \right) - \phi_{ \epsilon_1 } \left( \tau \beta^{-(t-1)} \right) \bigg| = 0 \ \ \ \text{a.s}. \end{align} for each $t \geq 1$ and uniformly in $\tau$ in bounded intervals. First, note that \begin{align} \bigg| \mathbf{E}^{*} \left( \text{exp} \ i \tau \hat{\beta}_n^{-(t-1)} \epsilon_t^{*} \right) - 1 \bigg| &\leq \mathbf{E}^{*} \left( |\tau \hat{ \beta}_n^{(t-1)} \epsilon_t^{*} | \right) \nonumber \\ &= | \tau | | \hat{ \beta}_n |^{-(t-1)} \frac{1}{n} \sum_{j=1}^n | \tilde{\epsilon}_j^{*} |. \end{align} Since by Lemma (ref) $\hat{\beta}_n \overset{ \text{a.s} }{ \to } \beta$, $| \beta | > 1$ and $( 1/ n ) \sum_{j=1}^n | \tilde{\epsilon}_t | \overset{ \text{a.s} }{ \to } \mathbb{E} | \epsilon_1 | < \infty$, it follows that almost for almost each sample path there exist positive integers $C$ and $m$ such that $( 1/ n ) \sum_{j=1}^n | \tilde{\epsilon}_j | \leq C$ and $| \hat{\beta}_n | \geq \delta > 1$ for all $n \geq m$. Thus, for almost each sample path \begin{align} \bigg| \mathbf{E}^{*} \left( \text{exp} \ i \tau \hat{\beta}_n^{-(t-1)} \epsilon_t^{*} \right) - 1 \bigg| \leq | \tau | \delta^{- (t-1)} C \end{align} for all $n \geq m$, and is established. Next, we have that \begin{align} \bigg| \mathbf{E}^{*} &\left( \text{exp} \ i \tau \hat{\beta}_n^{-(t-1)} \epsilon_t^{*} \right) - \phi_{ \epsilon_1 } \left( \tau \beta^{-(t-1)} \right) \bigg| \nonumber \\ &\leq \bigg| \frac{1}{n} \sum_{j=1}^n \text{exp} \ i \tau \hat{\beta}_n^{-(t-1)} \tilde{\epsilon}_j - \frac{1}{n} \sum_{j=1}^n \text{exp} \ i \tau \hat{\beta}_n^{-(t-1)} \epsilon_j \bigg| \nonumber \\ &+ \bigg| \frac{1}{n} \sum_{j=1}^n \text{exp} \ i \tau \hat{\beta}_n^{-(t-1)} \epsilon_j - \frac{1}{n} \sum_{j=1}^n \text{exp} \ i\tau \beta_n^{-(t-1)} \epsilon_j \bigg| \nonumber \\ &+ \bigg| \frac{1}{n} \sum_{j=1}^n \text{exp} \ i\tau \beta_n^{-(t-1)} \epsilon_j - \mathbf{E} \left( \text{exp} \ i\tau \beta^{-(t-1)} \epsilon_1 \right) \bigg|. \end{align} Therefore, for the first term in (ref) we obtain, \begin{align*} \bigg| \frac{1}{n} &\sum_{j=1}^n \text{exp} \ i\tau \hat{\beta}_n^{-(t-1)} \tilde{\epsilon}_j - \frac{1}{n} \sum_{j=1}^n \text{exp} \ i\tau \hat{\beta}_n^{-(t-1)} \epsilon_j \bigg| \\ &\leq \frac{1}{n} \sum_{j=1}^n \big| \text{exp} \ i\tau \hat{\beta}_n^{-(t-1)} \epsilon_j \big| \big| \text{exp} \ i \tau \hat{\beta}_n^{-(t-1)} \left( \tilde{\epsilon}_j - \epsilon_j \right) - 1 \big| \\ &\leq \frac{| \tau |}{n} \sum_{j=1}^n \big| \hat{\beta}_n^{-(t-1)} \big| \big| \tilde{\epsilon}_j - \epsilon_j \big| \\ &\leq \left( |\tau| \hat{\beta}_n^{-2(t-1)} \right)^{1 / 2} \left( \frac{|\tau| }{n} \sum_{j=1}^n \left( \tilde{\epsilon}_j - \epsilon_j \right)^2 \right)^{1 / 2} \to 0 \ \ \ \text{a.s}. \end{align*} uniformly on $\tau$ on compact subsets by Lemma (ref). Similarly, we obtain that \begin{align} \big| \hat{\beta}_n^{-(t-1)} - \beta^{-(t-1)} \big| \frac{| \tau |}{n} \sum_{j=1}^n | \epsilon_j | \to 0 \ \ \ \text{a.s}. \end{align} uniformly in $\tau$ on compact subsets, which implies that the second term in (ref) goes to zero almost surely. By the Glivenko-Cantelli theorem and the convergence theorem for characteristic functions, the third term in (ref) goes to 0 a.s. and uniformly for $\tau$ in a compact set. Hence, (ref) is established, and it follows that \begin{align} \phi_{U^{*}_n} ( \tau ) = \prod_{t=1}^n \phi_{\epsilon^{*}_1} \left( \tau \hat{\beta}_n^{-(t-1)} \right) \overset{ \text{a.s} }{ \to } \prod_{t=1}^{\infty } \phi_{\epsilon_1 } \left( \tau \beta^{-(t-1)} \right) = \phi_U ( \tau ). \end{align} Similarly, it can be shown that \begin{align} \phi_{U^{*}_n, V^{*}_n} ( \tau, s ) \overset{ \text{a.s} }{ \to } \prod_{t=1}^{\infty } \phi_{\epsilon^{*}_1} \left( \tau \hat{\beta}^{-(t-1)} \right) \prod_{t=1}^{\infty } \phi_{\epsilon^{*}_1} \left( s \hat{\beta}^{-(t-1)} \right) = \phi_U ( \tau ) \phi_V ( s ). \end{align}
corollaryIf $\left\{ \epsilon_t \right\}$ are assumed normal, we have, for $| \beta | > 1$, conditionally on $\left( X_1,..., X_n \right)$, \begin{align} T_n = \hat{ \sigma}_n^{-1} \left( \sum_{t=1}^n X^{*2}_{t-1} \right)^{1 / 2} \left( \hat{\beta}_n^{*} - \hat{\beta}_n \right) \to_d \mathcal{N} \left( 0,1 \right), \end{align} for almost all sample paths, where \begin{align} \hat{ \sigma}_n^2 = \sum_{t=1}^n \left( X_t - \hat{\beta}_n X_{t-1} \right)^2. \end{align}

Residual based Block Bootstrap for unit root testing.

Estimation framework.

Consider the autoregressive model given by

align[align omitted — 58 chars of source]

where $\rho_n = 1 + c / n$, $c < 0$ and the stationary process $\left\{ U_t \right\}$ satisfies: $E \left( U_t \right) = 0$, $E \left( | U_t |^{\nu} \right) < \infty$ for some $\nu > 2$, $f_U(0) > 0$ and $\sum_{k=1}^{\infty} \alpha (k)^{1-2 / \nu } < \infty$, where $\alpha(.)$ denotes the strong mixing coefficient of $\left\{ U_t \right\}$. In this Section, we summarize the RBB testing Algorithm as proposed by paparoditis2003residual. More specifically, the algorithm is carried out conditionally on the original data $\left\{ X_1,....,X_n \right\}$ and implicitly defines a bootstrap probability mechanism denoted by $P^{*}$ that is capable of generating bootstrap pseudo-series of the type $\left\{ X_t^{*}, t = 1,2,... \right\}$.

RBB Testing Algorithm.

The RBB procedure could be indeed an interesting application to examine, but we first consider the validity of the bootstrap IVX estimate constructed via the simple bootstrap procedure for obtaining the residual sequence. Then after establishing this asymptotic equivalence, we could examine the implementation of the RBB algorithm for obtaining the IVX bootstrap estimate.

enumerate• First calculate the centered residuals given by \begin{align} \hat{v}_t = \left( X_t - \tilde{\rho}_n X_{t-1} \right) - \frac{1}{n-1} \sum_{j = 2}^n \big( X_j - \tilde{\rho}_n X_{j-1} \big) \end{align} for $t = 2,3,...,n $ where $\tilde{\rho}_n = \tilde{\rho}_n \big( X_1, X_2, ..., X_n \big)$ is a consistent estimator of $\rho_n$ based on the observations $\left\{ X_1,..., X_n \right\}$. • Choose a positive integer $b$, where $b < n$ and let $i_0,...,i_{k-1}$ be drawn i.i.d with distribution uniform on the set $\left\{1,2,..., n - b \right\}$, such that for example, $k = \left[ \frac{(n - 1)}{b} \right]$. Then, the procedure constructs a bootstrap pseudo-series $\left\{ X_1^{*},...., X_l^{*} \right\}$ where $l = kb + 1$, as follows \begin{equation} X_t^{*} = \begin{cases} X_1 & , for \ t=1, \\ \hat{\mu} + X^{*}_{t-1} + \hat{v}_{i_{m} + s} & , for \ t = 2,3,...,l, \end{cases} \end{equation} where $m = \left[ \frac{(t-2)}{b} \right]$, $s = t - mb - 1$, and $\hat{\mu}$ is a drift parameter that is either equal to zero or represents a consistent estimator of $\mu$. • Let $\hat{\rho}_n$ be the estimator used to perform the unit root test. Compute the pseudo-statistic $\rho^{*}$ which is the test statistic $\hat{\rho}_l$ based on the pseudo-data $\left\{ X_1^{*},...., X_l^{*} \right\}$.

Taking into account the asymptotic theory of the regression statistic $n \left( \hat{\rho}_{LS} - 1 \right)$ for near integrated processes, the following theorem about the asymptotic local power behavior of the RBB based test can be established.

theoremLet \begin{align} \mathcal{B}_{RBB,n} \left( \rho_n ; \alpha \right) \to \mathbb{P} \left( J \leq \mathcal{C}_{\alpha} - c \right) \end{align} convergence in probability, where $\mathcal{C}_{\alpha}$ is the $\alpha$ quantile of the distribution of \begin{align} \bigg( W^2(1) - \sigma_U^2 / \sigma \bigg) \bigg( 2 \int_0^1 W^2 (r) dr \bigg)^{-1} \end{align} where J is a random variable which has the following distribution \begin{align} \bigg( \int_0^1 J_c (r) dW(r) + \left( 1 - \frac{ \sigma^2_U }{ \sigma^2 } \right) \bigg)\bigg( \int_0^1 J_c^2 (r) dr \bigg)^{-1} \end{align} and $J_c(r) = \int_0^1 e^{(r-s)c} dW(s)$ is the OU process generated by the stochastic differential equation $dJ_c(r) = cJ_c(r) dr + dW(r)$ with initial condition $J_c(0) = 0$.

FCLT for the Bootstrap Partial Sum process.

The asymptotic properties of the RBB testing procedure are based on the stochastic behaviour of the standardized partial sum process $\left\{ S_{m}^{*}(r), 0 \leq r \leq 1 \right\}$, defined by

align[align omitted — 188 chars of source]

and also

align[align omitted — 186 chars of source]

where for example, we define that $v_1^{*} \equiv X_1$, $v_t^{*} = X_t^{*} - \hat{\beta} - X_{t-1}^{*}$, for $t = 2,3,...,m$, and $\sigma^{*2} = \text{var} \left( m^{- 1 / 2} \sum_{ j = 1}^m v_j^{*} \right)$.

Note that, the partial sum process $S_{m}^{*}(r)$ is considered to be a random element in the function space $\mathcal{D}[0,1]$, which represent the space of all real valued functions on the interval $[0,1]$ that are right continuous at each point and have finite left limits.

The following theorem, given by paparoditis2003residual, shows that under a general set of assumptions on the process $\left\{ X_t \right\}$, and conditionally on the observed series $X_1,X_2,..., X_n$, the bootstrap partial sum process defined by (ref) and (ref) converges weakly to the standard Wiener process on $[0,1]$. Moreover, we denote with $T_n^{*} = T_n^{*} \left( X_1^{*}, X_2^{*},...., X_n^{*} \right)$ is a random sequence based on the bootstrap sample $X_1^{*}, X_2^{*},...., X_n^{*}$ and $G$ is a random measure, then the we use $T_n^{*} \Rightarrow G$ to denote the convergence in probability, which means that the distance between the law of $T_n^{*}$ and the law of $G$ tends to zero in probability for any distance metrizing weak convergence.

theoremLet $\left\{ X_t \right\}$ be a stochastic process, and assume that the process $\left\{ v_t \right\}$ defined by $v_t = X_t - \rho X_{t-1}$ satisfies the above regulatory conditions, and let $\tilde{\rho}_n$ be an estimator of $\rho$. Moreover, if $b \to \infty$ such that $b \big/ \sqrt{n} \to 0$ as $n \to \infty$, then \begin{align} S_{m}^{*} \Rightarrow W \ \ \ in probability. \end{align}

Therefore, the above result along with a bootstrap version of the continuous mapping theorem enables us to apply the block bootstrap proposal of this paper in order to approximate the null distribution of a variety of different test statistics.

theoremAssume that the process $\left\{ X_t \right\}$ satisfies the above conditions with $\beta = 0$. If $b \to \infty$ but $b \big/ \sqrt{n}$ as $n \to \infty$, then we have that \begin{align} \underset{ x \in \mathbb{R} }{ sup } \ \bigg| \mathbb{P}^{*} \left( l \left( \hat{ \rho }^{*LS}_n - 1 \right) \leq x \big| X_1,...,X_n \right) - \mathbb{P}_0 \left( \left( \hat{\rho}^{LS} - 1 \right) \leq x \right) \bigg| \to 0 \end{align} in probability, and \begin{align} \underset{ x \in \mathbb{R} }{ sup } \ \bigg| \mathbb{P}^{*} \left( l \left( \hat{ \rho }^{*LS}_{c,n} - 1 \right) \leq x \big| X_1,...,X_n \right) - \mathbb{P}_0 \left( \left( \hat{\rho}_{c,n}^{LS} - 1 \right) \leq x \right) \bigg| \to 0 \end{align} in probability, where $\mathbb{P}_{0}$ denotes the probability measure corresponding to the case where the statistics $\hat{ \rho }^{*LS}_{n}$ and $\hat{ \rho }^{*LS}_{c,n}$ are computed from a stretch of size $n$ from the unit root process obtained by integrating $\left\{ U_t \right\}$.