EconBase
← Back to paper

Partial Sum Processes of Residual-Based and Wald-type Break-Point Statistics in Time Series Regression Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

80,093 characters · 16 sections · 73 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Partial Sum Processes of Residual-Based and Wald-type Break-Point Statistics in Time Series Regression Models

I would like to thank Jose Olmo, Tassos Magdalinos and Jean-Yves Pitarakis for guidance, support and continuous encouragement throughout the PhD programme. Financial support from the VC PhD studentship of the University of Southampton is gratefully acknowledged. The author declares no conflicts of interest. All remaining errors are mine.} } }

center[center omitted — 31 chars of source]
abstractWe revisit classical asymptotics when testing for a structural break in linear regression models by obtaining the limit theory of residual-based and Wald-type processes. First, we establish the Brownian bridge limiting distribution of these test statistics. Second, we study the asymptotic behaviour of the partial-sum processes in nonstationary (linear) time series regression models. Although, the particular comparisons of these two different modelling environments is done from the perspective of the partial-sum processes, it emphasizes that the presence of nuisance parameters can change the asymptotic behaviour of the functionals under consideration. Simulation experiments verify size distortions when testing for a break in nonstationary time series regressions which indicates that the normalized Brownian bridge limit cannot provide a suitable asymptotic approximation in this case. Further research is required to establish the cause of size distortions under the null hypothesis of parameter stability. \\ Keywords: Structural break tests, Wald-type statistic, OLS-CUSUM statistic, Brownian bridge. \\

INTRODUCTION

Severe time series fluctuations manifesting as market exuberance are considered by econometricians as early warning signs of upcoming economic recessions which are not usually explained by common boom and bust cycles (see, greenwood2020predictable, baron2021banking and katsouris2021sequential). Furthermore, the economic aspects of prolonged economic policy uncertainty as well as as more recently pandemics can impact the robustness of parameter estimates due to increased model uncertainty. More precisely, such economic phenomena can appear in specification functions as structural breaks to parameter coefficients. Therefore, the identification and estimation of the true break-points can improve forecasts and reduce model uncertainty. Specifically, in the literature there is a plethora of statistical methodologies for structural break detection under different modelling environments (see, bai1997estimating, bai1998estimating and pitarakis2004least among others). In this paper we focus on the properties of the partial sum processes when constructing test statistics for testing the null hypothesis of no parameter instability in time series regression models under the assumption of stationary regressors vis-a-vis nonstationary regressors.

Although we do not introduce any new testing methodologies we review some important asymptotic theory results related to testing for structural breaks in time series regression models. More precisely, we study the asymptotic behaviour of partial-sum processes when these are employed to construct residual-based and Wald-type statistics in time series regression models. Our motivation for revisiting some of these aspects is to add on the discussion regarding the adequacy of finite-distribution approximations to large-sample theory of statistics. In particular, a common misconception when deriving asymptotic theory of test statistics and estimators is the fact that the large sample theory certainly provides good approximations to finite-sample results (see, chernoff1958asymptotic). Lastly, we consider two different modelling frameworks that is a stationary time series regression model and a nonstationary time series regression.

In terms of the structural break testing framework, we consider various examples which cover both the case of a known break-point as well as an unknown break-point, under parametric assumptions regarding the distribution of the innovations. Specifically, throughout this paper we assume that the innovations driving the error terms of both the stationary and nonstationary time series models are $\textit{i.i.d}$, with a known distribution, that is, are Gaussian distributed which implies that we are within the realm of parametric methods. Moreover, relaxing the particular assumption by assuming that distribution function of these innovations is considered as a nuisance parameter, requires to consider a semiparametric framework for estimation and inference purposes which is beyond the scope of this paper. Therefore, we have that the sequence of innovations, $\{ \epsilon_t \}_{t=1}^{\infty}$ to be $\epsilon_t \sim_{ \textit{i.i.d} } \mathcal{N}(0, \sigma^2)$ which implies independence and homoscedasticity\footnote{Notice that these two assumptions can indeed be quite strong. For example weakly dependent and heteroscedastic innovations can reflect more accurately the properties of aggregate time series and therefore such assumptions can be also included via appropriate modifications.}. Furthermore, the assumption of stationary innovation sequences is also imposed to facilitate the limit theory and this property holds regardless of the time series properties of the regressors and regressand.

In particular, we are interested in obtaining limit results of the following form

align[align omitted — 155 chars of source]

where $\mathcal{C}_T(s)$ is the test statistic and $W(s)$ is the standard Wiener process for some $s \in [0,1]$. However, the emphasis in this paper is the investigation of the main asymptotic theory aspects based of the regression model under consideration in terms of the properties of regressors. Therefore, the specific comparison allow us to focus on the implementation of the partial-sum processes when deriving asymptotic theory results for test statistics for these two classes of time series regression models, that is, the classical linear regression versus nonstationary time series regression. The investigation of the properties of partial-sum processes for the residuals of non-linear regression models such as ARMA, ARMAX, ARCH or GARCH models can be indeed quite fruitful, but we leave the particular considerations for future research. Related studies include the papers of kulperger2005high, aue2006strong and gronneberg2019partial.

All limits are taken as $T \to \infty$ where $T$ is the sample size. The symbol $"\Rightarrow"$ is used to denote the weak convergence of the associated probability measures as $T \to \infty$ (as defined in billingsley2013convergence) which implies convergence of sequences of random vectors of continuous cadlag functions on $[0,1]$ within a Skorokhod topology. The symbol $\overset{d}{\to}$ denotes convergence in distribution and $\overset{p}{\to}$ denotes convergence in probability. The remainder of the paper is organized as follows. Section (ref) discusses some examples related to the use of partial-sum processes for residual-based statistics in linear regression models. Section (ref) discusses Wald-type statistics in linear regression models. Section (ref) discusses some aspects related to the asymptotic theory for the structural break testing framework in nonstationary time series regressions. Section (ref) concludes and discusses further aspects for research.

RESIDUAL-BASED STATISTICS

The use of partial-sum processes for model residuals when testing the null hypothesis of no structural break using residual-based statistics appears in various applications in the literature katsouris2021sequential. Firstly, brown1975techniques proposed the OLS-CUSUM test constructed based on cumulated sums of recursive residuals for testing for the presence of a single structural break in coefficients of the linear regression model (see, also kramer1988testing and ploberger1992cusum). Other residual-based statistics found in the literature include the OLS-MOSUM test proposed by chu1995mosum which is constructed as sums of a fixed number of residuals that move across the whole sample. Therefore, in this case the statistic can be more sensitive in detecting parameter changes in comparison to the cumulated sums of recursive residuals. Furthermore, the literature evolved towards the construction of structural break statistics within an online monitoring framework which includes a window of fixed size (historical period) and an out-of-sample estimation window (monitoring period). In particular, chu1995moving and kuan1994implementing proposed the Moving Estimates and Recursive Estimates statistics for parameter stability respectively. A unified framework has been then proposed by the seminal study of chu1996monitoring and also further examined by leisch2000monitoring for the case of the generalized fluctuation test.

Some additional considerations include the study of the characteristics of structural change which includes the frequency of breaks in time series such as single vis-a-vis multiple break points (e.g., bai1997estimating and bai1998estimating) as well as the nature of structural change, which implies detecting structural breaks in the conditional mean vis-a-vis the conditional variance or higher moments of partial sum processes (e.g., horvath2001empirical, kulperger2005high, pitarakis2004least, andreou2006monitoring, andreou2009structural). Furthermore, an alternative asymptotic analysis of residual-based statistics is proposed by andreou2012alternative. The current literature extensively studies structural changes in the mean and variance of regression coefficients, however less attention is given to the study of structural break testing due to smooth changes in the persistence properties of regressors as it is defined within the framework of local-to-unity for autoregressive models. Specifically, smooth transitions of stochastic processes from $I(0)$ to $I(1)$ and other non-stationarities can be employed for the detection of bubbles in financial markets (see, horvath2020sequential). In this paper, our aim is to provide a discussion of the use of partial sum processes of residual-based and Wald-type statistics for these two different modelling environments.

OLS-CUSUM test

The OLS-CUSUM test statistic (see, kramer1988testing) belongs to the class of residual based statistics (see, e.g., stock1994unit) and it involves the partial sum processes of regression residuals based on the model under consideration. Following the literature we define a general class of regression residuals based on the partial sum process as proposed by kulperger2005high given by Definition (ref) and Theorem (ref) below.

definitionThe partial sum process of residuals is given by \begin{align*} \widehat{ \mathcal{U} }_T (s) = \sum_{t=1}^{ \floor{ Ts} } \widehat{u}_t , \ \ for some \ s \in [0,1]. \end{align*} with the partial sum process of the corresponding innovations is similarly defined as \begin{align*} \mathcal{U}_T (s) = \sum_{t=1}^{ \floor{ Ts} } u_t, \ \ for some \ s \in [0,1]. \end{align*}
theoremSuppose that $\sqrt{T} \big| \hat{ \theta}_T - \theta \big| = \mathcal{O}_p(1)$ where $\hat{\theta}_T$ is the set of estimated model parameters and $\theta$ is in the interior of $\Theta$. If $\mathbb{E} \left( | u_0| \right) < \infty$, then \begin{align} \underset{ s \in [0,1] }{ \mathsf{sup} } \ \frac{1}{ \sqrt{T} } \ \bigg| \bigg( \widehat{ \mathcal{U} }_T(s) - s \widehat{ \mathcal{U} }_T(s) \bigg) - \bigg( \mathcal{U}_T(s) - s \mathcal{U}_T(s) \bigg) \bigg| = o_p(1). \end{align} Then, the invariance principle for partial sums for an i.i.d sequence $\big\{ u_t \big\}$ such that \begin{align} \left\{ M_T (s) := \frac{ \mathcal{U}_T(s) - s \mathcal{U}_T(s) }{ \sigma_u \sqrt{T} }, \ s \in [0,1] \right\} \ \ and \ \ \left\{ \widehat{M}_T (s) := \frac{ \widehat{ \mathcal{U} }_T(s) - s \widehat{ \mathcal{U} }_T(s) }{ \sigma_u \sqrt{T} }, \ s \in [0,1] \right\} \end{align} where $\sigma_u^2 = \mathbb{E}\left( u^2 - \mu_u \right)^2$ and $\mu_u = \mathbb{E}\left( u_0 \right)$, implies that the functionals $M_T (s)$ and $\widehat{M}_T (s)$ both converge weakly in the Skorokhod topology $\mathcal{D}[0,1]$, to a Brownian bridge $\bigg\{ \mathcal{BB}(s) = W(s) - s W(1), \ s \in [0,1] \bigg\}$.
remarkThe functional $M_T (s)$ represents a partial-sum process and $W(s)$ is the corresponding standard Wiener process on the interval $[0,1]$ which preserves the continuous mapping theorem and convergence in a suitable functional space for any continuous function $g(.)$, such that $g\left( M_T (s) \right) \overset{\mathcal{D}}{\to} g \left( W(s) \right)$. Furthermore, both Definition (ref) and Theorem (ref) corresponds to partial-sum process of the residual sequence of a time series regression model with a general specification form, under the assumption of stationarity and ergodicity. Similarly, we can generalize these limit results to second order partial-sum processes for square residual sequences. More precisely, the first and second order partial-sum processes represent functionals of the OLS-CUSUM and OLS-CUSUM squared (see, deng2008limit).

The proof of Theorem (ref) is omitted which demonstrates a weak invariance principle; a stronger version of Donsker's classical functional central limit theorem (see, kulperger2005high and csorgHo2003donsker). In particular, the weakly convergence of the asymptotic distribution of the OLS-CUSUM statistic for the classical regression model is studied by aue2013structural. Related limit results can be found in the book of csorgo1997limit. The development of the asymptotic theory for the residual-based and Wald-type statistics when testing for structural breaks is based on the validity of Theorem (ref). For the remaining of this section we consider some standard examples from the literature to demonstrate the use of the residual-based statistics and their corresponding limit results when conducting statistical inference.

exampleConsider the location model formulated as below \begin{align} y_t = \mu_1 \mathbf{1} \{ t \leq k \} + \mu_2 \mathbf{1} \{ t > k \} + \epsilon_t, \ t =1,...,T \end{align} where $\mu_1$ and $\mu_2$ are both deterministic and the break point $k = \floor{Ts}$ for some $s \in [0,1]$ is an unknown fixed fraction of the full sample. Suppose a parametric assumption on the error term holds, such that $\epsilon_t \sim_{ \textit{i.i.d} } \mathcal{N} (0,1)$ for $t = 1,...,T$ and the following moment conditions apply \begin{align} (i). \ \ \mathbb{E} \left( \epsilon_t | \mathcal{F}_{t-1} \right) = 0 \ \ \ \ \ (ii). \ \ \mathbb{E} \left( \epsilon^2_t | \mathcal{F}_{t-1} \right) = \sigma^2_{\epsilon} \ \ \ \ \ (iii). \ \ \mathbb{E} \left( | \epsilon_t | \right)^{2 + d} < \infty, \ \ d > 0. \end{align} Then, the OLS estimator $\widehat{\mu} = T^{-1} \sum_{t=1}^T y_t$ of $\mu$ is a $\sqrt{T}-$consistent estimator and asymptotically normal. Furthermore, based on the location model formulation given by (ref) the objective is to conduct a two-sided test of the testing hypothesis $\mathbb{H}_0: \mu_1 = \mu_2$. Under the null hypothesis we obtain a consistent estimator $\hat{\sigma}^2_{\epsilon}$ of $\sigma^2_{\epsilon} < \infty$ such that $\hat{\sigma}^2_{\epsilon} \overset{ p }{ \to } \sigma^2_{\epsilon}$ as $T \to \infty$ as well as the OLS-residuals defined as $\widehat{\epsilon}_t = \big( y_t - \widehat{\mu} \big)$. Thus, we have that \begin{align} \widehat{\epsilon}_t \equiv \big( y_t - \bar{y}_T \big) = \mu + \epsilon_t - \frac{1}{T} \sum_{t=1}^T \left( \mu + \epsilon_t \right) \end{align}

Based on Example (ref), we consider that the following functional central limit theorem (FCLT) holds

align[align omitted — 114 chars of source]

which applies to the unobservable innovation terms of the above time series model that includes only a model intercept, where $W(.)$ is the standard Brownian motion such that $W(s) \sim N(0,s)$. For notational convenience we also denote with $\sigma_{\epsilon} W(s) \equiv B(s)$ for some $s \in [0,1]$.

Consider the standardized innovations defined as $\epsilon_t^o = \displaystyle \left( \epsilon_t - \frac{1}{T} \sum_{t=1}^T \epsilon_t \right) \equiv \big( \epsilon_t - \bar{\epsilon} \big)$ and suppose that $\epsilon_t$ is an $\textit{i.i.d}$ sequence of innovations. Then the centered partial sum process is defined as below

align[align omitted — 119 chars of source]

and the corresponding residual centered partial sum process is defined as below

align[align omitted — 149 chars of source]

Moreover, let $\widehat{\sigma}^2_m = \widehat{M}^2_T (1) / T$ to be the sample variance estimator of the above partial sum process. Then,

align[align omitted — 190 chars of source]

which shows that the self-normalized centered partial sum process $\big\{ \widehat{M}_T (s) / \widehat{\sigma}^2_m, \ 0 \leq s \leq 1 \big\}$ behaves as the residuals $\left\{ \widehat{\epsilon}_t \right\}_{t=1}^T$ are asymptotically the same as the unobservable innovations $\left\{ \epsilon_t \right\}_{t=1}^T$.

Therefore, the OLS-CUSUM statistic based on the standardized residuals is expressed as below

align[align omitted — 509 chars of source]

Then, it can be shown that $\mathcal{C}_T(k) \Rightarrow \big[ W(s) - s W(1) \big]$ which is Brownian bridge. In particular, the weakly convergence of the OLS-CUSUM statistic is based on $\underset{ 0 \leq k \leq T }{ \mathsf{max} } \mathcal{C}_T(k) \Rightarrow \underset{ 0 \leq s \leq 1 }{ \mathsf{sup} } \big[ W(s) - s W(1) \big]$.

remarkNotice that Example (ref) considers a linear time series model with no covariates, thus the structural break statistic can be considered as a difference in means for the two sub-samples, under the null hypothesis. The partial sum process for the innovation sequences of the model follows standard Brownian motion limit results. Furthermore, in Example (ref) we consider the OLS-CUSUM statistic for the classical regression model as proposed by the seminal study of ploberger1992cusum.
exampleConsider the following time series model \begin{align} y_t = x_t^{\prime} \beta_1 \mathbf{1} \{ t \leq k \} + x_t^{\prime} \beta_2 \mathbf{1} \{ t > k \} + \epsilon_t, \ t =1,...,T \end{align} where $y_t$ is the dependent variable, $x_t = \big[ 1, x_{2,t},...,x_{p,t} \big]^{\prime} = \big[ 1, \tilde{x}_t^{\prime} \big]^{\prime}$ is a $(p+1)-$dimensional vector and the error term $\epsilon_t$ are assumed to be i.i.d $(0, \sigma_\epsilon^2)$. We define $x_{1,t} \equiv x_t^{\prime} \mathbf{1} \{ t \leq k \}$ and $x_{2,t} \equiv x_t^{\prime} \mathbf{1} \{ t > k \}$ for $k = [Ts]$ with $s \in [0,1]$. Under the null hypothesis of no structural break $\mathbb{H}_0: \beta_1 = \beta_2$ with $\widehat{\beta}_T$ the $\sqrt{T}$-consistent estimator for $\beta$ such that $\sqrt{T} \left( \hat{\beta}_T - \beta \right) = \mathcal{O}_p(1)$ is bounded in probability.

Based on Example (ref) the OLS-CUSUM statistic, $\mathcal{C}_T(k)$, is obtained using the OLS residuals under the null hypothesis, defined as $\widehat{\epsilon}_t = y_t - \widehat{\beta}_T x_t = \epsilon_t - x_t^{\prime} \left( \widehat{\beta}_T - \beta \right)$. Then, we obtain that

align[align omitted — 234 chars of source]

In particular, we aim to show that $\mathcal{C}_T(k) {\Rightarrow} \big[ W(s) - s W(1) \big]$, which shows weakly convergence of the statistic to the Brownian bridge process. To do this, we consider the OLS residuals which can be expressed as below

align[align omitted — 269 chars of source]

Furthermore, the following result holds

align[align omitted — 181 chars of source]

A short proof on the asymptotic result above is provided here. We can express the left side of ((ref)) as an inner product since our framework allows such representation

align[align omitted — 264 chars of source]

Notice that it holds that

align[align omitted — 327 chars of source]

Then, the second term with a matrix decomposition for $Q = \left( \frac{1}{T} \sum_{t=1}^T x_t x_t^{\prime} \right)$ where $x_t= \big[ 1, \tilde{x}_t^{\prime} \big]^{\prime}$ can be expressed as below

align[align omitted — 452 chars of source]

since $\underset{ T \to \infty }{ \text{lim}} \frac{1}{T} \sum_{t=1}^{ \floor{Ts} } \tilde{x}_t \tilde{x}_t^{ \prime} = \tilde{\boldsymbol{Q}}$. Also, note that $

bmatrix[bmatrix omitted — 79 chars of source]

^{-1} =

bmatrix[bmatrix omitted — 84 chars of source]

$.

Therefore, we obtain

align[align omitted — 471 chars of source]

where $\boldsymbol{0}$ is $(p-1)$ dimensional column vectors of zeros. Then, using the asymptotic result given by ((ref)) as well as the expression for the OLS residuals as in ((ref)) the OLS-CUSUM statistic can be expressed as

align*[align* omitted — 666 chars of source]

Notice that the second term above converges to zero in probability such that

align[align omitted — 98 chars of source]

then the limit result follows as given below

align[align omitted — 304 chars of source]

Thus, the OLS-CUSUM statistic weakly converges to the Brownian bridge limit uniformly for $s \in [0,1]$.

exampleSimilarly to Examples (ref) and (ref) we can derive limit theory results for the OLS-CUSUM squared test (see detailed proofs in deng2008limit). Following Definition (ref) and Theorem (ref) the corresponding test statistic is defined as below \begin{align} \mathcal{C}^{(2)}_T(k) = \underset{ 0 \leq k \leq T }{ \mathsf{max} } \left( \frac{ \sum_{t=1}^{k} \widehat{\epsilon}^2_t - \frac{k}{T} \sum_{t=1}^T \widehat{\epsilon}^2_t }{ \widehat{\sigma}_{\epsilon} \sqrt{T} } \right) = \underset{ 0 \leq s \leq 1 }{ \mathsf{sup} } \left\{ \frac{1}{ \widehat{\sigma}_{\epsilon} } \left( \frac{1}{\sqrt{T}} \sum_{t=1}^{ \floor{Ts} } \widehat{\epsilon}^2_t - \frac{s}{\sqrt{T}} \sum_{t=1}^T \hat{\epsilon}^2_t \right) \right\} \end{align} Furthermore, the following two sufficient conditions hold \begin{align} \underset{ 1 \leq j \leq n }{ \mathsf{max} } \left| \frac{1}{ \sqrt{T} } \sum_{t=1}^j \epsilon_t x_t^{\prime} \left( \widehat{\beta}_T - \beta \right) \right| \overset{ p }{ \to } 0 \\ \underset{ 1 \leq j \leq n }{ \mathsf{max} } \ \frac{1}{ \sqrt{T} } \left( \widehat{\beta}_T - \beta \right)^{\prime} \sum_{t=1}^j x_t x_t^{\prime} \left( \widehat{\beta}_T - \beta \right) \overset{ p }{ \to } 0 \end{align} Therefore, it can be proved that the weakly convergence result to a Brownian bridge follows \begin{align} \mathcal{C}^{(2)}_T(k) \Rightarrow \underset{ 0 \leq s \leq T }{ \mathsf{sup} } \ \big[ W(s) - s W(1) \big]. \end{align}

WALD TYPE STATISTICS

In this section we examine the implementation of Wald-type statistics when constructing an equivalent structural break test for the conditional mean of a regression model. The particular methodology has been advanced by the seminal work of hawkins1987test and andrews1993tests.

definitionWe define the following $p-$dimensional limiting distribution for some $\pi \in (0,1)$ such that \begin{align} \mathcal{Q}_p( \pi ) := \frac{ \bigg[ W_p(\pi) - \pi W_p(1) \bigg]^{\prime} \bigg[ W_p(\pi) - \pi W_p(1) \bigg] }{ \pi ( 1 - \pi ) } \end{align} where $W_p(.)$ is $p \times 1$ vector of independent standard Brownian motions for some $p \geq 1$, where $p$ represents the number of parameters that the null hypothesis is testing for stability.
remarkThe limit process $\mathcal{Q}_p( \pi )$ is referred to as the square of a standardized tied-down Bessel process of order $p$ (see, andrews1993tests). In particular, the limit process given by Definition (ref) a is employed when deriving the asymptotic distribution of sup-Wald statistics, such that $\underset{ \pi \in (0,1) }{ \mathsf{sup} } W_T( \pi ) \overset{ d }{ \to } \underset{ \pi \in (0,1) }{ \mathsf{sup} } \mathcal{Q}_p( \pi )$.

We mainly consider univariate regression models but the framework can be also generalized to time series regressions with multiple regressors. Furthermore, when applying the supremum functional to Wald-type statistics we assume that the unknown break-point $\pi \in (0,1)$ lies in a symmetric subset of the particular unit set, to ensure that inference is not at the boundary of the parameter space (see, andrews2001testing).

Unconditional Mean test

Consider the following time series regression model

align[align omitted — 116 chars of source]

where $\theta_1$ and $\theta_2$ are both deterministic and the break point $k = \floor{Ts}$ for some $s \in (0,1)$ is an unknown fixed fraction of the full sample period. The following assumption hold:

assumptionLet $\epsilon_t$ be a sequence of random variables. Then, the following moment conditions hold. \begin{enumerate} • $\mathbb{E} \left( \epsilon_t | \mathcal{F}_{t-1} \right) = 0$ with an asymptotic variance as $T \to \infty$ given as below \begin{align} \mathsf{var}\left( T^{-1/2} \sum_{t=1}^{ \floor{Ts} } \epsilon_t \right) = \frac{ \floor{Ts} }{T} \sigma^2_{\epsilon} \overset{ p }{ \to } s \sigma^2_{\epsilon}, \ \ \ \sigma^2_{\epsilon} = \underset{ T \to \infty }{ \mathsf{lim} } \frac{1}{T} \mathbb{E} \left[ \left( \sum_{t=1}^T \epsilon_t \right)^2 \right] \end{align} • sup$_t$ $\mathbb{E} \left( | \epsilon_t | \right)^{2 + d} < \infty$ for some $d > 2$. • $\{ \epsilon_t \}_{t=1}^T$ satisfies the following Functional Central Limit Theorem (FCLT), \begin{align} \frac{1}{\sqrt{T}} \sum_{t=1}^{ \floor{Ts} } \epsilon_t \overset{\mathcal{D}}{\to} \sigma_{\epsilon} W(s) \equiv B(s) \end{align} \end{enumerate}
remarkAssumption A1 gives the moment conditions of the classical regression model with only intercept. In particular, $\mathsf{Var} \left( T^{-1/2} \sum_{t=1}^{ \floor{Ts} } \epsilon_t \right)$ gives the asymptotic variance of the partial-sum of innovations. Since the limiting variance is bounded then convergence in distribution follows.

The null and alternative hypothesis are as below

align[align omitted — 101 chars of source]

and we consider the following test statistic for detecting structural change in the unconditional mean of the simple linear regression with only intercept.

align[align omitted — 680 chars of source]

We can show that the asymptotic variance of the normalized statistical distance measure $\sqrt{T} \left( \bar{y_{1}} - \bar{y_{2}} \right)$ is given by $\mathsf{Avar} \left[\sqrt{T} \left( \bar{y_{1}} - \bar{y_{2}} \right) \right] = \frac{ \sigma^2_{\epsilon} }{ s (1-s) } \equiv \sigma^2_{z}$ by noting that $\underset{ T \to \infty }{ \text{lim} } \frac{ T }{ \floor{Ts} } = \frac{1}{s}$ and $\underset{ T \to \infty }{ \text{lim} } \frac{ T }{ T - \floor{Ts}} = \frac{1}{s(1-s)}$. According to aue2013structural testing the null hypothesis of equal means across a $p$-dimensional multivariate time series is equivalent to constructing a $p-$dimensional CUSUM process and testing for a structural break at an unknown break point $k$. Due to the fact that the partial-sum process representing the CUSUM statistic has a weak convergence to a $p-$dimensional Brownian Motion process, then the quadratic form of the partial-sum process weakly convergence to the sum of squared independent Brownian bridges. The specific property permits to establish an equivalence between a supremum Wald-type statistic and a CUSUM-type statistic. In particular, we show that the asymptotic distribution of the OLS-CUSUM statistic can be deduced from the asymptotic distribution of $\mathcal{Z}_T$ and vice-versa.

theoremLet $\mathcal{Z}_T(s)$ be the test statistic for testing the null hypothesis $\mathbb{H}_0: \theta_1 = \theta_2$ of no structural change in the unconditional mean as defined by ((ref)) for some $0 \leq k \leq T$. Assume the conditions A1 to A3 given by Assumption (ref) hold and let $\mathcal{Z}^o_T(s)$ be an OLS-CUSUM type statistic. Then, both statistics weakly converge in the Skorokhod space $\mathcal{D}[0,1]$ to functionals of a Brownian bridge such as $\{ \mathcal{BB}(s): s \in [0,1] \}$ where $\mathcal{BB}(s) = \big[ W(s) - s W(1) \big]$. \begin{align} \mathcal{Z}_T (r) &= \underset{ s \in [\nu , 1 - \nu] }{ \mathsf{sup} } \left\{ T \frac{ \left( \bar{y_{1}} - \bar{y_{2}} \right)^2 }{ \widehat{\sigma}^2_{ z } } \right\} \Rightarrow \underset{ s \in [\nu , 1 - \nu] }{ \mathsf{sup} } \frac{ \big[ W(s) - s W(1) \big]^2 }{s(1-s)} \\ \mathcal{Z}^o_T(s) &= \underset{ s \in [\nu , 1 - \nu] }{ \mathsf{sup} } \left\{ \frac{1}{ \sqrt{T} } \frac{1}{\sqrt{ \widehat{\sigma}_{\epsilon} } } \left( \sum_{t=1}^{ \floor{Ts} } y_t - s \sum_{t=1}^T y_t \right) \right\} \Rightarrow \underset{ s \in [\nu , 1 - \nu] }{ \mathsf{sup} } \big[ W(s) - s W(1) \big] \end{align}
remarkTheorem 2 implies that the $\mathcal{Z}_T(s)$ statistic of equality of means in a SLR model with only intercept across the unknown break point has an equivalent weak convergence to the CUSUM statistic of the time series under consideration using a suitable normalization constant.

Next, we consider the corresponding supremum Wald-type statistic for testing for structural change in the unconditional mean of the classical regression model with only intercept. The test is formulated as below

align[align omitted — 349 chars of source]

where $\widehat{\sigma}_{\epsilon}^2 = \frac{1}{T} \sum_{t=1}^T \hat{\epsilon}_t^2(k) $ the residual variance under the null hypothesis and is a consistent estimator of $\sigma_{\epsilon}$ such that $\widehat{\sigma}_{\epsilon}^2 \overset{ p }{ \to } \sigma_{\epsilon}^2$. Thus, by substituting the restriction matrix $\mathcal{\boldsymbol{R}} = \big[ \boldsymbol{I} - \boldsymbol{I} \big]$, $\boldsymbol{Z} = \big[ \boldsymbol{X}_1 \ \boldsymbol{X}_2 \big]^{\prime}$ and the estimator $\widehat{\boldsymbol{\Theta} } = \big[ \widehat{\theta}_1 \ \widehat{\theta}_2 \big]^{\prime}$ of $\boldsymbol{\Theta}$ into the above formulation of the Wald statistic we obtain the expression

align[align omitted — 291 chars of source]

We observe that for the unconditional mean model it holds that

align*[align* omitted — 187 chars of source]

Since in this section we consider mean shifts, $X_1$ stacks the elements of $x_t$ for which $t \leq k$, that is, $\mathbf{1} \{ t \leq k \}$ and $X_2$ stacks the elements of $x_t$ for which $t > k$, that is, $\mathbf{1} \{ t \leq k \}$. Therefore, using the corresponding matrix notation the OLS estimators can be written as below

align*[align* omitted — 234 chars of source]

Thus, the Wald statistic can be expressed as

align*[align* omitted — 263 chars of source]

Under the null hypothesis $\mathbb{H}_0: \theta_1 = \theta_2$, we have that since $\left( \bar{y_{1}} - \bar{y_{2}} \right) \equiv \left( \widehat{\theta}_1 - \widehat{\theta}_2 \right)$ we obtain the expression

align*[align* omitted — 316 chars of source]

Therefore, the Wald test is equivalently written as below

align*[align* omitted — 471 chars of source]

A simple application of the WLLN for the variance of the OLS estimator implies that $\widehat{\sigma}^2_{\epsilon} \overset{ p }{ \to } \sigma_{\epsilon}$ as $T \to \infty$. Moreover since it holds that $\frac{k}{T} \frac{T-k}{T} \overset{ p }{ \to } s(1-s)$, we obtain the weak convergence for the sup-Wald statistic

align[align omitted — 260 chars of source]

The above asymptotic result indeed verifies the weakly convergence of the sup-Wald statistic into the supremum of a normalized squared Brownian bridge, specifically for the unconditional mean specification of the regression model. Then, sstatistical inference is conducted based on the null hypothesis, $\mathbb{H}_0: \theta_1 = \theta_2$, which is rejected for large values of the sup-Wald statistic with a significance level $\alpha \in (0,1)$. Thus, the exact form of the limiting distribution of the statistic is employed to obtain associated critical values, denoted with $c_{\alpha}$ such that $\mathbb{P} \left( \mathcal{W}_T^{*} (s) > c_{\alpha} \right) > 0$ with $\underset{ T \to \infty }{ \mathsf{lim} } \mathbb{P} \left( \mathcal{W}_T^{*} (s) > c_{\alpha} \right) = 1$.

Conditional Mean test

Consider the following model

align[align omitted — 167 chars of source]

where $y_t$ is the dependent variable, $x_t$ is a $p \times 1$ vector of regressors, $\epsilon_t$ is an unobservable disturbance term with $\mathbb{E}\left( \epsilon_t | x_t \right) = 0$ almost surely and $\theta_1$ and $\theta_2$ the regression coefficients formulated under the null hypothesis of no structural break. Define with $x_{1,t} \equiv x_t^{\prime} \mathbf{1} \{ t \leq k \}$ and $x_{2,t} \equiv x_t^{\prime} \mathbf{1} \{ t > k \}$ where $k = \floor{Ts}$ for some $s \in (0,1)$. Furthermore, notice that the regressor vector $x_t$ can contain exogenous regressors and lagged dependent variables with unknown integration order. Moreover, under suitable regularity assumptions one can consider a static, a dynamic time series regression model as well as models with integrated regressors or cointegrated regressors.

The null hypothesis of interest is

align[align omitted — 112 chars of source]

The alternative hypothesis $\mathbb{H}_A$ is that $\mathbb{H}_0$ is false, which implies that the regression coefficient has a structural break at some unknown break point in the sample. Under $\mathbb{H}_0$, the unknown constant parameter vector $\theta$ can be consistently estimated using the OLS estimator such that

align[align omitted — 159 chars of source]

Under the alternative $\mathbb{H}_A$, $\theta_0 \equiv \theta_t$ can be considered as a time-varying parameter vector, which implies, $\theta_t \equiv \theta_1$ for $1 \leq t \leq k$ and $\theta_t \equiv \theta_2$ for $k + 1 \leq t \leq T$. To facilitate the estimation and inference based on the above regression model we impose the following regularity conditions.

assumptionLet $\epsilon_t$ be a sequence of random variables. Then the following moment conditions and FCLT hold, which allow $x_t$ and $\epsilon_t$ to be weakly dependent. \begin{enumerate} • $\{\epsilon_t\}_{t=1}^{T}$ is a homoskedastic martingale difference sequence (m.d.s) such that $\mathbb{E} \left( \epsilon_t | \mathcal{F}_{t-1} \right) = 0$ and $\mathbb{E}\left( \epsilon_t^2 \right) =\sigma^2$, where $\mathcal{F}_{t-1} = \left\{ x^{\prime}_t, x^{\prime}_{t-1},..., \epsilon_{t-1}, \epsilon_{t-2},... \right\}$. • $\{x_t\}_{t=1}^{T}$ has at least a finite second moment, that is, sup$_t$ $\sum_{t=1}^{T} ||x_t||^{2 + d} < \infty$ where $d > 0$ and the following weak law of large numbers (WLLN) holds \begin{align} \frac{1}{T} \sum_{t=1}^{T} x_t x_t^{\prime} \overset{p}{\to} \boldsymbol{Q}_T \ \ as \ \ T \to \infty \ such that \ \underset{ s \in [0,1] }{ \mathsf{sup} } \bigg\rvert \frac{1}{T} \sum_{t=1}^{ \floor{Ts} } x_t x_t^{\prime} - s\boldsymbol{Q}_T \bigg\rvert = o_p(1) \end{align} where $\boldsymbol{Q}_T$ is a $(p \times p)$ non-stochastic finite and positive definite matrix. • $\{x_t \epsilon_t \}_{t=1}^{T}$ satisfies the following Functional Central Limit Theorem\footnote{The multivariate FCLT is examined in the studies of Phillips1986multiple and wooldridge1988some.}(FCLT), \begin{align} \mathcal{S}_T(s) = \left( T^{-1/2} \boldsymbol{\Omega}_T^{-1/2} \sum_{t=1}^{ \floor{Ts} } x_t \epsilon_t \right) \Rightarrow \boldsymbol{W}_p(s), \ s \in (0,1) \ and \ \boldsymbol{\Omega}_T = \mathbb{E} \left( x_t x_t^{\prime} \epsilon^2_t \right) >0. \end{align} \end{enumerate}

In other words, the general moment conditions given by Assumption (ref) permits to consider both residual-based statistics as well as Wald-type statistics as suitable detectors when testing for a break-point within the full sample. In particular, chen2012testing consider the implementation of generalized Hausman-type tests using nonparametric estimation and inference techniques.

Asymptotic Convergence of sup-Wald test

Next, we derive the asymptotic convergence of the Wald test when detecting a single structural break in the regression for the classical regression model with a conditional mean specification form, under Assumptions B1 to B3. Furthermore, under the null hypothesis the break-point is unidentified, thus to facilitate statistical inference we consider the corresponding supremum Wald-type statistic based on the OLS estimator which is expressed as $\mathcal{W}^{*}_T(s) := \underset{ s \in [ \nu, 1 - \nu] }{ \text{sup} } \mathcal{W}_T(s)$ for some $\nu \in (0,1)$.

Using the null hypothesis $\mathbb{H}_0: \theta_1 = \theta_2$ of no structural break we obtain an equivalent representation using the linear restriction matrix $\mathcal{ \boldsymbol{R} }$ of rank $q$, which implies that $\mathbb{H}_0: \mathcal{\boldsymbol{R} } \boldsymbol{\Theta} = \boldsymbol{0}$. We define $\boldsymbol{\Theta} = \big[ \theta_1 \ \theta_2 \big]^{\prime}$, $\boldsymbol{Z} = \big[ \boldsymbol{X}_1 \ \boldsymbol{X}_2 \big]^{\prime}$ and $\mathcal{\boldsymbol{R}} = \big[ \boldsymbol{I} - \boldsymbol{I} \big]$ and prove that the asymptotic distribution of the corresponding sup-Wald statistic weakly convergences to the supremum of a normalized squared Brownian bridge.

theoremUnder the null hypothesis $\mathbb{H}_0: \theta_1 = \theta_2$ with $0 < \nu < 1$ we define $\mathcal{W}^{*}_T(s)$ \begin{align} \mathcal{W}^{*}_T(s) := \underset{ s \in [ \nu, 1 - \nu] }{ sup } \mathcal{W}_T(s) \Rightarrow \underset{ s \in [\nu , 1 - \nu] }{ \mathsf{sup} } \frac{ \big[ \boldsymbol{W}_p(s) - s \boldsymbol{W}_p(1) \big]^{\prime} \boldsymbol{Q}_T^{-1} \big[ \boldsymbol{W}_p(s) - s \boldsymbol{W}_p(1) \big] }{s(1-s)} . \end{align} where $\mathcal{W}_T(s)$ denotes the Wald statistic for testing the null hypothesis $\mathbb{H}_0: \theta_1 = \theta_2$ and is expressed as \begin{align} \mathcal{W}_T = \frac{1}{ \widehat{\sigma}_{\epsilon}^2 } \left( \mathcal{ \boldsymbol{R} } \boldsymbol{\Theta} \right)^{ \prime } \left[ \mathcal{ \boldsymbol{R} } \left( \boldsymbol{Z}^{\prime} \boldsymbol{Z} \right)^{-1} \mathcal{\boldsymbol{R} }^{\prime} \right]^{-1} \left( \mathcal{ \boldsymbol{R} } \boldsymbol{\Theta} \right) \end{align} with $\widehat{\sigma}^2_{\epsilon} = \displaystyle \frac{1}{T} \sum_{t=1}^T \left( y_t - x_t^{\prime} \widehat{\theta}_T \right)$ a consistent estimate of $\sigma^2_{\epsilon}$ the OLS variance under the null hypothesis.
remarkThe above asymptotic theory analysis demonstrates that CUSUM-type statistics for detecting a structural break in the classical regression model with a single unknown break-point typically weakly converge to Brownian bridge functionals, such that $\mathcal{BB}(s) := \big[ W_p(s) - sW_p(1) \big]$ while the corresponding Wald-type tests weakly converge to a normalized version of the CUSUM test with a normalized constant given by $k_T := \left\{ \frac{k}{T} \left( 1 - \frac{k}{T} \right) \right\}^{1/2}$ where $k = \floor{Ts}$ for some $s \in [0,1]$.

Monte Carlo simulation experiments can be used to compare the asymptotic validity and performance of OLS-CUSUM type statistics and Wald-type tests by obtaining associated empirical size and power results. An important criticism of CUSUM-type statistics is that these tests are based on residuals under the null hypothesis which implies that the test is not designed with a specific alternative under consideration. Therefore, although the use of these tests can lead to a monotonically increasing power, in practise it can be stochastically dominated by the power of a Wald type test (see, also andreou2008restoring). On the other hand the advantage of using a Wald-type statistic when testing for a structural break, is that the construction of the test allows to incorporate residuals obtained either under the null or under the alternative hypothesis, providing this way superior power performance. Intuitively, using the residual variance of the unrestricted model leads to better finite-sample power since the Wald-type statistic contains information from the alternative model (alternative hypothesis).

STRUCTURAL BREAK TESTING UNDER NONSTATIONARITY

Although the purpose of this paper is to explain the main challenges one faces when testing for a structural break in nonstandard econometric problems we discuss the main intuition when developing associated asymptotic theory using some examples. More specifically, in order to accommodate the nonstationary aspect in time series models\footnote{Standard regularity conditions for estimation and inference in nonstationarity time series models, is the asymptotic theory developed for nearly unstable autoregressive processes. The limit theory of time series models such as the asymptotic inference of AR(1) processes first examined by mann1943statistical has been extended to the non-stationary asymptotic case as it captured via the local to unity framework which allows for the autoregressive coefficient of a univariate AR(1) to be expressed in the form $\rho = 1 + c/n$, where $n$ the sample size and $c$ the unknown degree of persistence of the data generating process.} we need to consider a suitable probability space in which weakly convergence arguments holds. Due to the fact that the asymptotic terms of sample moments of nonstationary time series models often involve non-standard limit results such as convergence to stochastic integrals, care is needed when applying related convergence arguments. In particular, the stochastic integrals of the form $\int_0^1 W dW (s)$ can be shown to converge weakly, under the null hypothesis, to the associated stochastic integral with the limiting Brownian motions, under the assumption of the existence of cadlag functionals in the unit interval equipped with the Skorokhod topology (see, muller2011efficient).

Linear Restrictions Testing

We begin our analysis by discussing the main limit theory which is employed in nonstationary time series models, by providing two examples: (i) a time series regression with integrated regressors and (ii) a predictive regression with persistent regressors. Although, such asymptotics are commonly used when considering the asymptotic behaviour of t-tests and Wald-type tests based on linear restrictions on the parameter coefficients, these are applicable when constructing the corresponding structural break statistics.

Time Series Regression with Integrated Regressors

A class of nonstationary time series models include the linear regressions with integrated regressors as proposed by the studies of cavanagh1995inference and jansson2006optimal among others.

example(Cointegrating Regression, banerjee1993co) Consider the following bivariate system of co-integrated variables $\{ y_t \}_{t=1}^{\infty}$ and $\{ x_t \}_{t=1}^{\infty}$ such that \begin{align} y_t &= \beta x_t + u_t \\ x_t &= x_{t-1} + \epsilon_t \end{align} with $u_t \sim N(0, \sigma^2_u )$, \ $\epsilon_t \sim N(0, \sigma^2_{\epsilon} )$ and $\mathbb{E}( u_t \epsilon_s ) = \sigma_{u \epsilon } \ \forall \ t \neq s$.

The OLS estimator of $\beta$ is given by $ \widehat{ \beta } = \left( \sum_{t=1}^T x_t^2 \right)^{-1} \left( \sum_{t=1}^T y_t x_t \right)$ which implies

align[align omitted — 148 chars of source]

Thus, for this regression model due to the presence of the integrated regressor the limit is expressed as

align[align omitted — 128 chars of source]

Next, in order to derive the limiting distribution of the model parameter for the case of integrated regressors, we first consider the limiting distribution of the sample moment $\left( T^{-1} \sum_{t=1}^T x_t u_t \right)$. To do this, we assume the existence of a conditional distribution for $u_t$ given $\epsilon_t$, which implies the conditional mean form

align[align omitted — 205 chars of source]

Furthermore, we define with $W_{\epsilon} (r)$ and $W_{v} (r)$ to be two independent Wiener processes on $\mathcal{C}[0,1]$. Therefore,

align[align omitted — 222 chars of source]

Substituting $x_t = x_{t-1} + \epsilon_t$ into the above expression we obtain

align*[align* omitted — 438 chars of source]

Most importantly, within this setting the following asymptotic results hold

align[align omitted — 163 chars of source]

and it has been also proved in the seminal study of Phillips1987time (see, also Phillips1986multiple)

align[align omitted — 298 chars of source]

Putting the above together we obtain that

align[align omitted — 268 chars of source]

Furthermore, phillips1988asymptotic proved the following limit result

align[align omitted — 120 chars of source]

Therefore, under the null hypothesis $\mathbb{H}_0: \beta = 0$, it follows that

align[align omitted — 297 chars of source]

Thus, the $t-$statistic, denoted as $\mathcal{T}_{\beta = 0}$ for testing the null hypothesis, $\mathbb{H}_0: \beta = 0$, is written as

align[align omitted — 277 chars of source]

has the following asymptotic distribution

align[align omitted — 580 chars of source]
remarkNotice that the derived limiting distribution of the Student-t statistic indicates that the $t-$ratio of $\widehat{\beta}$ does not follow a standard normal distribution unless $\phi = 0$; in which case the structure of the model implies that the regressor $x_t$ is exogenous for the estimation of the model parameter $\beta$ which is the main parameter of interest for inference purposes. In particular, when $\phi \neq 0$ then the first term of the above limiting distribution gives rise to second-order or endogeneity bias, which although asymptotically negligible in estimating $\beta$ due to super consistency, can appear in finite-samples causing size distortions when obtaining the empirical size of the test.

Predictive Regression Model

Predictive regression models are extensively used in time series econometrics and the empirical finance literature for examining the stock return predictability puzzle as proposed by campbell2006efficient. A standard predictive regression has the following econometric specification (see, kostakis2015robust)

align[align omitted — 98 chars of source]

The innovation sequence $(\epsilon_t,u_t)$ is generated such that $(\epsilon_t,u_t) \sim_{ \textit{i.i.d} } \mathcal{N} (0,\Sigma)$ where $\Sigma =

bmatrix[bmatrix omitted — 97 chars of source]

$.

Similarly to the previous example, the null hypothesis of interest using linear restrictions on the model parameter $\beta$, is formulated such that $\mathbb{H}_0: \beta = 0$. The main econometric challenges when conducting statistical inference using the predictive regression model includes the problem of embedded endogeneity due to the innovation structure of the system as well as the nuisance parameter of persistence, $c$, when the autocorrelation coefficient of the model is expressed with the local-to-unity specification. As a result, depending on the value of the autocorrelation coefficient, the asymptotics for the parameter of the predictive regression model take a different form which makes statistical inference challenging. Specifically, when $|\rho| < 1$, then $x_t$ is known to be stationary, when $\rho = 1$ then $x_t$ is unit root or integrated and when $c<0$ is assumed to follow a local-to-unity or nearly integrated process. The literature has proposed various methodologies for conducting statistical inference robust to the nuisance parameter of persistence. For instance, Phillips2007limit study the limit theory of time series models which includes regressors that are close to the unit root boundary\footnote{The authors consider the limit distribution theory in both the near-stationary $(c < 0)$ and the near-explosive cases $(c > 0)$.}.

In terms of the asymptotic theory that correspond to the predictive regression model, we consider the partial-sum process for the integrated regressor. The particular aspect is important especially in comparison to when constructing residual-based statistics in which case the main quantity of interest is the partial-sum process of the residuals corresponding to stationary innovations. Denote with $\mathcal{F}_{T,t-1}$ the $\sigma-$algebra generated by the random variables and with $X_{ \floor{Tr} }$ the partial-sum process of interest. Then, under the assumption that the $x_t$ is generated as a local-unit-root process the weakly convergence of the partial-sum functional corresponds to a uniform convergence to an Ornstein-Uhlenbeck (OU) process\footnote{The continuous time OU diffusion process given by $dy_t = \theta y_t dt + \sigma dw_t, \ y_0 = b, t > 0 $, where $\theta$ and $\sigma >0$ are unknown parameters and $w_t$ is the standard Wiener process, has a unique solution to $\{ y_t \}$ which is expressed as $ y_t = \text{exp} \left( \theta t \right) b + \sigma \int_0^t \text{exp} \left[ \theta( t - s) \right] dw_s \equiv \text{exp} \left( \theta t \right) + \sigma J_{\theta} \left( t \right)$ (see, e.g., perron1991continuous).} rather to the standard Wiener process (see, Phillips1987time and related limit theory in durrett1978functional) as

align[align omitted — 118 chars of source]

In other words, the assumptions we impose regarding the parametrization of the autocorrelation coefficient $\rho$ can change the asymptotic behaviour of the stochastic difference equation. Generally, statistical inference is nonstandard in the sense that when $\rho_T = \left( 1 + \frac{c}{T} \right)$ for some nuisance parameter $c$, then the testing problem concerning the parameter $\beta$ exhibit nonstandard large-sample properties under local-to-unity asymptotics.

Predictive tests

To provide some further clarity regarding the effect of expressing the autocorrelation coefficient in terms of moderate deviations from unity, to the validity of conventional inference methods, we consider as an example the stationary autoregressive model AR(1), $y_t = \rho_T y_{t-1} + u_t$ where $u_t \overset{ i.i.d }{ \sim }(0, \sigma^2)$ and $| \rho | < 1$. In this case, it is a well-known fact that the limit distribution of the t-test for testing the null hypothesis $\mathbb{H}_0: \rho_T = 0$ with $\mathcal{T}_T ( \rho_T) = \displaystyle \frac{ \widehat{\rho}_T - \rho }{ \widehat{\sigma} } \Rightarrow \mathcal{N}(0,1)$ converges to a standard normal distribution. In addition, the asymptotic distribution of $\mathcal{T}_T ( \rho_T)$ is invariant even under the assumption of conditional heteroscedasticity which implies $\mathbb{E}\left( u_t^2 | \mathcal{F}_{t-1} \right) = \sigma^2_t$ and $\underset{ t \in \mathbb{Z} }{ \text{sup} } | \hat{\sigma}^2_t - \sigma^2_t | = o_p(1)$, a condition for consistent estimation.

On the other hand, the limiting distribution of the model parameter $\beta$ of the predictive regression model as well as the associated t-test for testing the null hypothesis, $\mathbb{H}_0: \beta = 0$, appears to be challenging due to the fact that it is found to be nonstandard and the corresponding t-test is non-pivotal since it depends on the nuisance parameter $c$ (see, cavanagh1995inference and campbell2006efficient). Consequently, given the focus of our study to the asymptotic behaviour of partial-sum processes when constructing test statistics in nonstationary time series models, we illustrate the related asymptotic theory with some examples.

It is worth mentioning that the partial-sum processes, $X_{ \floor{Tr}} (r)$, are considered to be (maximally) invariant with respect to the presence of the model intercept $\mu$. Therefore under the null hypothesis, $\mathbb{H}_0: \beta = 0$, joint weak convergence of observation processes to their Brownian motion counterparts holds. Specifically, an application of the invariance principle proposed by Phillips1987time such that $\frac{x_{ \floor{Tr} }}{\sqrt{T} } \Rightarrow J_c(r)$, where $J_c(r) = \int_{0}^r e^{ (r-s)c} dB_c(s)$ is a standard Ornstein-Uhlenbeck process, implies that

align[align omitted — 206 chars of source]

Since $T \left( \widehat{\beta}_T - \beta \right) = \mathcal{O}_p(1)$ is bounded in probability, then we can establish the usual mode of converges in distribution in the same probability space such that

align[align omitted — 293 chars of source]

Consider the stationary case such that $| \rho | < 1$, then by partitioning the covariance matrix $\Sigma$ similar to the regression model with the integrated regressor, we use the decomposition $\epsilon_{1.2 t} = \epsilon_t - \frac{ \sigma_{\epsilon u} }{ \sigma_{\epsilon} } u_t$ which implies

align[align omitted — 499 chars of source]

The first term of expression ((ref)) since in includes a conditional error term then the corresponding limit distribution converges to a mixed normal limit. Moreover, we consider the joint convergence of the martingale sequences $\left\{ \sum_{t=1}^n x_{t-1} \epsilon_t \right\}$ and $\left\{ \sum_{t=1}^n u_t \right\}$ are defined on the same probability space

align[align omitted — 105 chars of source]

Thus, the conditional covariance matrix of the martingale vector $\xi_{Tt}$ is given by

align[align omitted — 440 chars of source]

Using the partition matrix identity, $\Sigma_{1.2} = \Sigma_{11} - \Sigma_{12} \Sigma_{22}^{-1} \Sigma_{21}$, to the predictive regression model we obtain the relation $\frac{ \displaystyle \sigma_{1.2} }{ \displaystyle \sigma^2_{\epsilon} } = 1 - \frac{ \displaystyle \sigma^2_{ u \epsilon } }{ \displaystyle \sigma^2_{\epsilon} \sigma^2_{u} }$. Therefore, the following mixed normal limit convergence holds

align[align omitted — 253 chars of source]

Hence, using expressions ((ref)) and ((ref)) and the limit result below

align[align omitted — 202 chars of source]

for some $r \in (0,1)$, we obtain an analytical expression for the asymptotic distribution of the t-statistic $\mathcal{T}_T ( \beta_T ) = \frac{ \widehat{\beta}_T - \beta }{ \widehat{ \sigma}_{\beta} } \Rightarrow \phi \widehat{\mathcal{M}}(c) + (1 - \phi^2)^{1/2} \mathcal{Z}$, where $\mathcal{Z} \sim \mathcal{N}(0,1)$ is independent of the random quantity $\widehat{\mathcal{M}}(c)$ and $c$ denotes the nuisance parameter of persistence.

More precisely, it holds that

align*[align* omitted — 471 chars of source]

where $\phi = \frac{ \displaystyle \sigma_{ \epsilon u } }{ \displaystyle \sigma_{\epsilon} \sigma_{u}}$, and $\widehat{\mathcal{M}}(c) = \frac{ \displaystyle \int_{0}^1 J_c(s) dB_{x}(s) }{ \displaystyle \sigma^2_{u} \int_{0}^1 J^2_c(s) ds}$.

In summary, the t-statistic for the predictive regression coefficient has been proved to have a non-standard limiting distribution which implies that normal or chi-square based inference is not available in practise, due to the endogeneity problem as well as the existence of persistence regressors. Therefore, the particular non-standard testing problem makes it difficult to conduct inference without prior knowledge regarding the exact value of the coefficient of persistence and in practise cannot be consistently estimated. Suggested solutions to overcome this problem include the Bonferroni confidence interval proposed by cavanagh1995inference and elliott1996efficient, the conditional likelihood approach that uses sufficient statistics proposed by jansson2006optimal and the control function approach proposed by elliott2011control.

Structural Break Testing

Setting against the background described in details in Section 4.1 we now discuss the main challenges for the development of the structural break testing framework.

Asymptotic Distributions of Test Statistics

In this section we consider the implementation of OLS-CUSUM and Wald-type statistics within the local-to-unity framework when detecting instabilities in the parameters of a predictive regression model with persistent regressors. Although in this paper we consider an in-sample monitoring scheme, under suitable modifications an on-line (sequential) monitoring scheme as proposed by chu1996monitoring can provide an early warning mechanism for risk management purposes based on macroeconomic and financial conditions. The implementation of such a framework can be interpreted as a dynamic methodology for testing for parameter instability under the assumption of time-varying persistence properties.

Therefore, we are interested in proposing suitable testing methodologies for detecting structural change in the vector of regression coefficients $\theta$ of the following time series regression model

align[align omitted — 155 chars of source]

where $x_t = R_T x_{t-1} + u_t$ with $R_T = \left( 1 - \frac{C}{T} \right)$ and $C = \mathsf{diag} \left\{ c_1,..., c_p \right\}$ (see, kostakis2015robust). Under the null hypothesis of no structural break $\mathbb{H}_0: \beta_1 = \beta_2$. Our proposition aims to incorporate neglected non-linearities such as structural breaks to the current LUR framework. For example, certain non-linear functions\footnote{For example, wang2012specification consider a specification test for nonlinear nonstationary models within the LUR framework of cointegrating regression system.} of $I(1)$ processes can wrongly behave like stationary long memory processes (see, e.g., kasparis2014nonlinearity). A recent approach which considers structural breaks under such conditions is presented by berenguer2020cumulated. In this paper, we consider the weak dependence assumption.

OLS-CUSUM test statistic

Next, we focus on the limit theory of the residual-based statistic, CUSUM test, constructed using the OLS residuals of the predictive regression under the null hypothesis, $\mathbb{H}_0: \beta_1 = \beta_2$. The OLS residuals of the predictive regression are given by $\widehat{ \epsilon }_t^{ ols } = y_t - x_t^{\prime} \widehat{\beta}^{ ols }$ and the corresponding OLS-CUSUM statistic is expressed in the usual way as below for $s \in [0,1]$

align[align omitted — 530 chars of source]

The OLS residuals of the predictive regression can be expressed as $\widehat{ \epsilon }_t^{ ols } = y_t - x_t^{\prime} \widehat{\beta}^{ ols } \equiv \epsilon_t - x_t^{\prime} \left( \widehat{\beta}_T - \beta \right) $

align[align omitted — 482 chars of source]

Therefore, using expression ((ref)) the OLS-CUSUM statistic within the LUR framework becomes

align[align omitted — 660 chars of source]

As we can clearly observe from the second term of the above expression that corresponds to the limiting distribution of the OLS-CUSUM statistic in a predictive regression model with persistent regressors the dependence on the nuisance parameter of persistence, $c$, makes the limit result non-standard and non-pivotal. In other words, the implementation of a residual-based statistic in a predictive regression model using OLS residuals is considered to be problematic in the derivation of the asymptotic distribution due to its dependence on the nuisance degree of persistence of the autoregressive specification of the model. Furthermore, for $k = \floor{Tr}$ we define the following term for notation simplicity

align[align omitted — 164 chars of source]

Then, the weakly convergence result can be written as below

align[align omitted — 150 chars of source]

In summary, the weakly convergence of the in-sample OLS-CUSUM statistic includes the component $\widetilde{J}_{\infty} ( c, r)$ which depends on the nuisance parameter $c$ and thus can affect the true size of the test under the null hypothesis of no parameter instability given the fact that we cannot consistently estimate the coefficient of persistence. However, for example the term $\int_0^r J_c(s) ds - r \int_0^1 J_c(s) ds$ it is likely to be quite small and therefore can be considered not to be contributing to huge size distortions. An extensive Monte Carlo study can shed light on the particular aspect of the proposed test for detecting structural change in predictive regression models with persistent regressors. Thus, we have demonstrated that when testing for a structural break in linear time series regression models, the partial-sum processes of conventional test such as those of residual-based and Wald-type statistics have different properties when information regarding the integration order of regressors is available in the form of the LUR specification form.

CONCLUSION

In this paper we establish the Brownin Bridge limiting distributions in a fairly standard settings, that is linear regression models under the assumption of stationarity and ergodicity when constructing residual-based and Wald-type statistics for testing the null hypothesis of no parameter instability. In particular, in all those cases we have demonstrated that the normalized Brownian bridge limit holds for both test statistics. Additionally, we investigate whether this property also holds for nonstationary time series regression models with integrated or persistent regressors. Our asymptotic theory analysis has demonstrated that while in the classical linear regression model with stationary regressors the convergence of the test statistics to brownian bridge limit results hold, in the case of the nonstationary time series model it appears to be the case that the limiting distribution is non-standard and non-pivotal due to the dependence of the distribution to the nuisance parameter of persistence.

Based on the general assumption that the underline stochastic processes are mean-reverting, then we can establish adequate approximations to finite-sample moments for the model under consideration regardless of the econometric environment operates under the assumption of stationarity or we consider the settings of a nonstationary time series model. On the other hand, since the main feature of nonstationary time series models is the parametrization of the autocorrelation coefficient with respect to the nuisance parameter of persistence, this implies that limiting distributions are non-standard and non-pivotal which makes inference difficult. In particular, despite the large availability of macroeconomic and financial variables which can be included as regressors in predictive regressions (such as financial ratios, diffusion indices, fundamentals), practitioners have no prior knowledge regarding the persistence properties of predictors so conventional estimation and inference methods for model parameters, such as predictability tests, forecast evaluation tests as well as structural break testing require to handle the nuisance parameter of persistent.

Therefore, further research is needed to propose suitable statistical methodologies that take into consideration these challenges, especially when testing for the presence of a structural break in nonstationary time series models. Practically, the extension of structural break tests in nonstationary time series models, such as predictive regressions, which are particularly useful when information regarding the time series properties of regressors is not lost by taking the first difference for instance, is crucial for both theoretical and empirical studies. In a subsequent paper, we propose a formal econometric framework and develop the associated asymptotic theory for Wald-type statistics under regressors nonstationarity.