Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
114,633 characters · 15 sections · 127 citation commands
Break-Point Date Estimation for Nonstationary Autoregressive and Predictive Regression Models
\pagenumbering{roman}
Lecturer in Economics, Department of Economics, University of Exeter Business School, Exeter EX4 4PU, United Kingdom. E-mail Address: \textcolor{blue}{[email removed]}.\\ } } }
\setcounter{page}{1} \pagenumbering{arabic}
Structural break inference is an important task when the robustness of econometric estimation is concerned. Although various existing methodologies in the time series econometrics literature focus on testing for the presence of parameter instability in regression coefficients, the estimation of the exact break-points and their statistical properties is more challenging, especially under the assumption of regressors nonstationarity. We develop point estimation and asymptotic theory for break-point estimators in nonstationary autoregressive and predictive regression model with nonstationary regressors. We first present the structural break test for the model coefficient of the predictive regression model based on the usual least squares estimate of the coefficient of the first-order autoregressive model as well as the predictive regression model. Second we using the IVX instrumentation which is found to be robust to the abstract degree of persistence (see, kostakis2015robust), we construct IVX-based structural break tests\footnote{Relevant asymptotic theory analysis for the Wald-type statistics we employ in this article are presented by katsouris2023predictability, katsouris2023testing.} (see, katsouris2023predictability,katsouris2023structural, katsouris2023testing) and corresponding break-point estimators. Overall, the accurate dating of a structural break is of paramount important especially since predictive regression models are commonly used to detect the so-called "pockets of predictability". Thus, correctly identifying periods of predictability regardless of the presence of a structural break is crucial for asset pricing and risk management purposes.
Determining the asymptotic distribution of the break-point estimator in nonstationary regression models is a challenging task. Specifically, for the case of a shift with fixed magnitude it can be shown that the limiting distribution of the change-point estimator depends on the underlying distribution of the innovation in a complicated manner (see, hinkley1970inference). The least squares estimation of the change-point in mean and variance in a linear time series regression model is obtained by pitarakis2004least, although the nonstationary properties of regressors are not considered. Furthermore, pitarakis2012jointly, pitarakis2014joint develops suitable econometric frameworks for jointly testing the null hypothesis of no structural change in nonstationary times series regression models. More recently dalla2020asymptotic and stark2022testing consider relevant aspects for structural break testing and dating in regression models under dependence. The commonly used assumption is that the break occurs at an unknown time location within the full sample, such that, $t = \floor{ \tau T }$, where $\tau \in (0,1)$ and these limits are also used when obtaining moment functionals.
Specifically, predictive regressions with regressors generated via the local-to-unity parametrization has recently seen a growing attention in the literature (see, phillipsmagdal2009econometric, gonzalo2012regime, phillips2014confidence, kostakis2015robust, kasparis2015nonparametric, demetrescu2020testing and duffy2021estimation). The main feature of these time series regression models, is that an $I(0)$ integrated dependent variable (such as stock returns) is regressed against a persistent predictor (such as the divided-price ratio) and this allows to construct predictability tests (see, also zhu2014predictive). In statistical terms, the process $\left( Y_t, \boldsymbol{X}_{t-1} \right)_{ t \in \mathbb{Z} }$ with a martingale difference sequence $\xi_t = \left( u_t, v_t \right)^{\prime}$ such that $\left( Y_t, \boldsymbol{X}_{t-1}, \xi_t \right)_{ t \in \mathbb{Z} }$ is generated by a predictive regression model and $\boldsymbol{X}_{t}$ is an autoregressive process expressed using the local-to-unity parametrization, it allows us to examine the persistence properties of regressors across various regimes. On the other hand, these frameworks operate under the assumption of a parameter constancy in the full sample.
In this paper, we study the break-point estimators for structural break tests in predictive regression models with possible nonstationary regressors. The stability of the autoregressive processes is determined by the local-to-unity parametrization. Specifically, we focus on autoregressive processes which are close to the unit boundary but have different order of convergence, namely high persistent regressors which are $o_P \left( n^{-1 / 2} \right)$ and mildly integrated regressors which are $o_P \left( n^{- \upgamma / 2} \right)$ for some $\gamma \in (0,1)$ and a positive persistence coefficient $c > 0$. These features do complicate the asymptotic theory analysis in some extend but their properties are useful for break-point estimation and dating in the aforementioned settings. For instance, conventional structural break tests for the parameters of linear regression models employ the widely used sup-Wald test proposed by andrews1993tests. However, the distributional theory of the Andrews's test crucially depends on the strict stationarity assumption of regressors. In contrast to the literature that focuses on structural break testing in linear regressions, the predictive regression model is usually fitted to economic datasets which contain time series that are highly persistent. Therefore, within such econometric environment the traditional law of large numbers and central limit theorems can invalidate the standard econometric assumptions of linear regression models, which affect the large sample approximations. As a result distorted inferences can occur when testing for parameter instability in predictive regressions when these features are not accommodated in the asymptotic theory of the tests.
Our first objective is to theoretically demonstrate the impact of the presence of the nuisance parameter of persistence to the limiting distribution of test statistics and break-point estimators. Specifically, the limit result obtained by andrews1993tests implies the use of the supremum functional on the Brownian Bridge process defined as $\mathsf{sup}_{ s \in [0,1] } \big[ W_n(s) - s W_n (1) \big]$. Under weak convergence, we have process convergence where $B$ is a Brownian motion, that is, a pivot process, hence enabling the practitioner to use standard tabulated critical values. Conveniently, katsouris2023predictability, katsouris2023testing show that the OLS based sup-Wald statistic in the presence of nearly integrated regressors weakly converges into the standard NBB limit, however in the case of regressors with high persistence the same limit is no longer valid. As a result, this complicates the asymptotic theory analysis of break-point estimators especially since usually these nonstationary features are not a prior known. However, before proceeding to the limit behaviour of break-point estimators we discuss an alternative estimation procedure which is found to be robust to the unknown persistence properties can be applied to a structural-break testing while producing similar asymptotic behaviour. In this direction, one can employ the IVX instrumental variable estimation approach proposed by phillipsmagdal2009econometric and extensively examined by kostakis2015robust in the context of predictability tests and by gonzalo2012regime in the context of predictability tests for threshold predictive regression models\footnote{More specifically, the limit theory of kostakis2015robust provides a unified framework for robust inference and testing regardless the persistence properties of regressors. A simple example is the application of the IVX-Wald test for inferring the individual statistical significance of predictors under abstract degree of persistence. Further scenarios such as predictors of mixed integration order see phillips2013predictive and phillips2016robust in which cases the mixed normality assumption still holds.}.
Thus, we study the statistical inference problem of structural break testing at an unknown break-point therefore the supremum functional is implemented and two test statistics are considered based on two different parameter estimation methods. The first estimation method considers the OLS estimator, while the second method considers an instrumental variable based estimator, namely the IVX estimator proposed by phillipsmagdal2009econometric. These two estimators of the predictive regression model have different finite-sample and asymptotic properties, which allows us to compare the limiting distributions of the proposed tests for the two different types of persistence of the regressors. For both test statistics we assume that the regressors included in the model are permitted to be only one of the two persistence types which simplifies the asymptotic theory of the tests, however the presence of nuisance parameters under the null of parameter constancy, that is, the unknown break-fraction and the coefficient of persistence, requires careful examination of the asymptotic theory. An additional caveat is the inclusion of an intercept in the predictive regression model which induces different limiting distributions when is assumed to be stable vis-a-vis the case in which is permitted to shift. In summary, we are interested to examine the consistency and convergence rates of the break-point estimators that correspond to these two estimation methodologies for the coefficients of the predictive regression model.
The asymptotic theory of the present paper hold due to the invariance principle of the partial sum process of $x_t$, where $x_t = \left( 1 - \frac{c}{n^{\upgamma}} \right) x_{t - 1}$, as proposed by phillips1987time and is considered to be the building block for related limit results when considering time series models. We denote with $\hat{ \mathcal{U} }_n(s) := \frac{1}{ \sqrt{n} } \sum_{ t = 1 }^{ \floor{ ns } } x_{t}$, for some $s \in [0,1]$ and with $\hat{ \mathcal{U} }_{ n^{\upgamma} }(s) := \frac{1}{ T^{ \upgamma / 2 } } \sum_{ t = 1 }^{ \floor{ n^{ \upgamma } s } } x_{t}$, for some $s \in [0,1]$ and $0 < \upgamma < 1$ for the invariance principle of the partial sum process of $x_t$ in the case of mildly integrated processes, for the corresponding limit as proposed by phillips2005limit. Motivated from the aforementioned seminal work, as well as the framework of phillipsmagdal2009econometric, similarly in this paper we use these invariance principles of the partial sums processes that correspond to the instrumental variable, IVX, proposed by phillipsmagdal2009econometric, specifically within a structural break testing framework. These results allow us to formally obtain the limiting distributions of the proposed tests with respect to the nuisance parameter of persistence along with the unknown break-point location, and observe in which cases we obtain nuisance-free inference that can simplify significantly the hypothesis testing procedure.
All random elements are defined with a probability space, denoted by $\left( \Omega, \mathcal{F}, \mathbb{P} \right)$. All limits are taken as $n \to \infty$, where $n$ is the sample size. The symbol $"\Rightarrow"$ is used to denote the weak convergence of the associated probability measures as $n \to \infty$. The symbol $\overset{d}{\to}$ denotes convergence in distribution and $\overset{\text{plim}}{\to}$ denotes convergence in probability, within the probability space. Let $\left\{ Y_t, \boldsymbol{X}_t \right\}_{t=1}^n$, where $\boldsymbol{X}_t = \left( X_{1t},..., X_{pt} \right)$ denote the corresponding random variables of the underline joint distribution function. The rest of the paper is organized as follows. Section (ref) presents the asymptotic theory for break-point detection in AR(1) models. Section (ref) considers the asymptotic theory for break-point detection predictive regressions.
Our study contributes to the time series econometrics literature in several ways. Firstly, we propose a set of Wald type statistics for detecting parameter instability in predictive regression models robust to the persistence properties of regressors. We derive analytical forms of the limit distributions of these test statistics and show under which conditions these limit results correspond to the conventional NBB result providing this way a clear estimation and inference strategy to practitioners interested to implement a structural break test in predictive regression models with possibly nonstationary regressors. Therefore the significance of the proposed framework in the broader econometric literature includes the provision of detailed asymptotic results for these structural break tests as well as necessary data transformations and functional forms which are can be implemented so that statistical inference can be simplified. Therefore, the construction of structural break tests in predictive regression models would not have been possible without one first consider the asymptotic behaviour of estimators of the predictive regression model.
Secondly, in a similar spirit as in saikkonen2006break, we aim to investigate the properties of estimators of the time period where a shift has taken place. In particular, under the assumption of a possible single structural break, we point identify the shift in the model coefficients of the first order autoregressive and predictive regression model. Moreover, two alternative estimators for the break date are considered, and their asymptotic properties are derived under various assumptions regarding the local alternatives in which case the size of the shift is considered. These results have various further applications as they can then be used to explore the implications of inference in predictive regression models after estimation of breaks. Lastly, we aim to perform a more detailed and more insightful investigation of the small-sample properties of the break date estimators and the resulting structural break tests by extending the simulation design and empirical findings of katsouris2023predictability. Notice that the break-point date estimator denoted with $k = \uptau n$, where $\uptau \in (0,1)$, needs to be estimated using a criterion function. However, the estimation approach of the econometric model will also consequently affect the statistical properties of the corresponding break-point estimator. Rearranging the break-point estimators gives the following expression
A relevant asymptotic theory question of interest is under which conditions the above convergence in probability holds. Roughly speaking, the main idea of the framework here is that asymptotically the break date can then be located at the true break date or within a neighbourhood of the true break date. Therefore, one will need to carefully consider the required assumptions that provide identification conditions for the break date. In addition, we will need to consider possible aspects of the structural break testing environment under regressors nonstationarity that can potentially make the break date estimator $\hat{\uptau}$ inconsistent, resulting to an incorrect estimation of the break date for the modelling environment under consideration.
In terms of the stochastic integral approximations we consider in this paper, a relevant literature include phillips1987time, phillips1987towards, phillips1988regression as well as kurtz1991weak and hansen1992convergence who present various examples of weak convergence of stochastic integrals and stochastic differential equations upon which the limit theory of this paper is based on. The related limit theory proves that the continuous time OU diffusion process given by
where $c$ and $\sigma >0$ are unknown parameters and $W_t$ is the standard Wiener process, has a unique solution to $\{ J(t) \}_{ t \in [0,1]}$, such as $J(t) = \displaystyle \mathsf{exp} \left( c t \right) b + \sigma \int_0^t \mathsf{exp} \left\{ c( t - s) \right\} dW (s) \equiv \text{exp} \left(c t \right) + \sigma J_{c} \left( t \right)$ (see, perron1991continuous). This representation provides a way of determining the asymptotic convergence of each of the component which the expression of the Wald type statistic can be decomposed to. \color{black}
When the break-point is known one can apply a Chow-type test statistic. In particular, sun2022asymptotically investigate the asymptotic distribution of a Chow-based test in the presence of heteroscedasticity and autocorrelation. However, when the break-point is unknown, the estimation procedure for break-point date estimation is usually based on the optimization of a criterion function. Moreover, the presence of a possible single against multiple-breaks will require modifications to the formulation of the relevant hypotheses, the criterion function as well as the asymptotic theory analysis. Specifically, in the literature of structural break testing for linear regression models, bai1997estimating propose a statistical procedure for detecting multiple breaks sequentially (one-by-one) rather than using a simultaneous estimation approach of multiple breaks. In summary while bai1997estimating proposed a sample-splitting method to estimate the breaks one at a time by minimizing the residual sum of squares, bai1998estimating proposed to estimate the breaks simultaneously by minimizing the residual sum of squares. The advantages of the former method lies in its computational savings and its robustness to misspecification in the number of breaks. A number of issues arise in the presence of multiple breaks. More precisely, the determination of the number of breaks, the estimation of the break points given the number as well as the statistical analysis of the resulting estimators. Simultaneous and sequential methods are fundamentally different methodologies that yield different break-point estimators. However, one of the drawbacks of sequential break-point algorithms is the complexity in deriving convergence rates of the estimated break-point. In particular, bai1997estimating demonstrate that sequentially obtaining estimated break points are $n-$consistent, which corresponds to the same rate as in the case of simultaneous estimation\footnote{ In particular, for the simultaneous estimators of break, their asymptotic distributions in stationary models are symmetric, but the computational burden is heavy. The least-squares operations are of order $O( n^2 )$ even under the most efficient algorithm. In contrast, for the sequential estimators of breaks, the computational burden is light (the least-squares operations are of order $O( n )$, but the asymptotic distributions of the estimators, are asymmetric due to the misspecification of the models. Hence, additional efforts, such as repartitioning the sample, are needed in order to obtain symmetrical asymptotic distributions of the estimators.}.
Although these features have not been investigated in the case of predictive regression models with possible nonstatationary regressors or regressors with mixed integration order. In this article, we aim to provide some insights on some relevant cases (not all of the nonstationary regimes). Within our framework, the first stage of the procedure implies testing for the possible presence of a structural break under high persistence, that is, the regressors of the autoregressive and the predictive regression models are parametrized using the local-to-unity parameter which induces a nearly-integrated process. Furthermore, the accuracy of the estimator depends on whether the break-point estimator is bounded. Thus, the robustness of the break-point estimator can be verified by its ability to be as close as possible to the true break-point within the full sample. \color{black}
From the unit root and structural break literature perspective, relevant testing methodologies include saikkonen2006break as well as the single-break homoscedasticity-based persistence persistence change tests proposed by harvey2006modified. Testing for multiple break-points are discussed in the studies of lumsdaine1997multiple and carrion2009gls. Moreover, kejriwal2008limit study estimation and inference in cointegrated regression models with multiple structural changes allowing both stationary and integrated regressors; by deriving the constistency, rate of convergence and the limit distribution of the estimated break fractions. If the coefficients of the integrated regressors are allowed to change, the estimated break fractions are asymptotically dependent so that confidence itervals need to be constructed jointly. Recently, kejriwal2020bootstrap propose bootstrap procedures for detecting multiple persistence shifts in heteroscedastic time series such as cointegrating regressions. The bootstrap procedure proposed by kejriwal2020bootstrap for detecting multiple breaks as an example here is as below:
The asymptotic validity of the two-step procedure follows from (i) the test in the first step is asymptotically pivotal under the null and consistent against alternatives involving a change in at least one parameter and (ii) the break fraction is consistently estimated as long as any of the parameters are subject to a break. In particular, the second fact ensures that the F-test in the second step converges to a chi-square distribution under the null hypothesis of no structural change in the subvector of interest. More precisely, this result follows since the estimate of the break fraction is fast enough to ensure that the limiting distribution of the parameter estimate is the same that would prevail if the break date was known. In summary, the particular framework reveals that existing partial break sup-Wald tests diverge with $n$ when the coefficients are not being tested are subject to change. Thus, kejriwal2020bootstrap propose a simple two-step procedure which first tests for joint parameter stability and subsequently conducts a standard chi-squared stability test on the coefficients of interest allowing the other coefficients to change at the breakpoints estimated by minimizing the sum of squared residuals in the pure structural change model. The procedure proposed by kejriwal2020bootstrap estimates the number of breaks using a sequential test of the null hypothesis of $( \ell \geq 1 )$ breaks against the alternative of $( \ell + 1 )$ breaks. The particular approach is useful especially in cases when the null of no break could be rejected against the alternative hypothesis of at least one break. Lastly, casini2022generalized propose a generalized laplace inference in multiple change-points models framework, although not suitable for detecting multiple-breaks in cointegrating and predictive regression models.
Generally speaking, estimation and inference procedures of econometric models that do not account for such data-driven selection of nuisance parameters such as an unknown structural break or an unknown threshold variable can perform poorly when used for empirical studies andrews2021inference (see, also gonzalo2012regime,gonzalo2017inferring). In this direction, zhu2022testing propose a framework which considers the possible presence of both a structural break and a threshold effect in predictive regression models, robust to the degree of persistence. In the self-normalization literature two relevant studies include choi2022subsample and zhang2018unsupervised. Then, to correct for serial dependence, it often requires a consistent estimator of the long-run variance, namely, the spectral density at zero frequency. The particular quantity involves autocovariances of all orders, and a data-driven bandwidth is usually needed for its estimator to be adaptive to the underlying dependence. Specifically, shao2010testing uses the self-normalization approach for single change-point testing in time series. Note that, the self-normalization approach implies that instead of restoring to a consistent estimator of the long-run variance, one relies on a sequence of recursive estimators to form the normalizer and in turn pivotalize the asymptotic distribution of the test statistic. A framework for single structural break testing based on the self-normalization approach is given by ling2007testing. Lastly, andrews2021inference study methodologies for inference after estimation of structural breaks (see, also fiteni2002robust and busetti2003variance).
We present some relevant illustrative examples to the modelling environment under consideration.
Consider testing the joint null hypothesis\footnote{Note that we can also test predictability in the pre-break and post-break subsamples, using $H_0: \boldsymbol{\beta}_1 = 0$ and $H_0: \boldsymbol{\beta}_1 = 0$, respectively.} $H_0: \boldsymbol{\beta}_1 = \boldsymbol{\beta}_2 = \boldsymbol{0}$. Under the null hypothesis, the predictive regression model reduces to a change in mean model as below (see, katsouris2022partial):
A suitable stopping rule used for the change-point detection rely either on thresholding or on the optimization of a model selection criterion. Various methods for multiple change-point detection exist in the literature which include the dynamic programming method, to detect multiple change points in the exponential family of distributions. According to casini2022generalized, for multiple break-points such that $m$ change-points, the inference framework can be constructed by denoting with $j \in \left\{ 1,..., m+1 \right\}$, where by convention $n_0^0 = 0$ and $n_{m+1}^0 = n$. In the multiple break-point setting, such that $\big( n_1^0,..., n_m^0 \big)$, there are $( m + 1)$ regimes, each corresponding to a distinct parameter value $\delta_j^0$, which needs to be estimated. Therefore, when the full sample has $n$ available observations, the aim is to simultaneously estimate the unknown regression coefficients together with the break points. Moreover, the asymptotic theory analysis for the multiple break-points case follow directly from the single break case. However, especially for nonstationary time series regressions the existence of multiple break points might be problematic due to possible changes in the peristence properties of regressors from one regime to the other. Neverthless, the procedure and estimation criterion proposed by casini2022generalized can be employed to identify these multiple structural breaks in predictive regression models as well (see, barigozzi2018simultaneous).
Therefore, the class of estimators for inference in multiple change-points regressions relies on a certain criterion function (see, casini2022generalized)
As a result, in order to establish the large-sample properties of break-point estimators and test statistics, in a similar spirit as in casini2022generalized, we can consider the shrinkage theoretical framework of bai1998estimating and qu2007estimating (see, also lavielle2000least and nkurunziza2020inference). On the other hand, necessary modifications of the asymptotic theory analysis is required in order to incorporate the features of the local-to-unity parametrization as well as the use of the proposed sup-Wald type statistics based on the OLS and the IVX estimation.
Break point estimation is an important component in change-point detection problems. In order to obtain some useful insights we begin by considering the statistical properties that correspond to structural break tests and break-point estimators for an AR(1) autoregressive model, where the local-to-unity parametrization is omitted rendering this way a stationary AR$(1)$ model under suitable parameter space restrictions. We follow the study of pang2021estimating.
\paragraph{Motivation} katsouris2023testing, katsouris2023predictability established the asymptotic behaviour of structural break tests in predictive regression models for the sup-Wald OLS and sup-Wald IVX statistic.
We also proved that under the assumption of a known break-point the limiting distributions converge to a nuisance parameter free distribution $\chi^2$ regardless of the persistence properties. However, the studies of katsouris2023testing, katsouris2023predictability didn't consider the asymptotic theory of break-point date estimation, which we aim to present in this article.
Therefore, we are interested to estimate the structural parameters $\beta_1$ and $\beta_2$ and the time (or location) of change $\uptau_0$. Furthermore, we estimate the following model:
We could also denote with $\mathcal{S}_n = \mathcal{S}_{1n} ( \beta_1, \uptau ) + \mathcal{S}_{2n} ( \beta_2, \uptau )$ where $\mathcal{S}_n \equiv \text{RSS}_{n}( \uptau )$ such that
Therefore, in practice the criterion function which will need to be employed to estimate the break-point date estimators requires an iterative procedure. In other words, for each $\tau \in \Pi$, we obtain the regression parameter estimators pre$-\floor{ n \tau }$ and post$-\floor{ n \tau }$ such that
for $j \in \left\{ 1,2 \right\}$ respectively. Furthermore, in practice the shift point will be estimated as the sample partition that minimizes the objective function concentrated in $\tau$ such that
Moreover, we have that $\hat{\vartheta}_n := \big( \hat{\beta}_{1n}, \hat{\beta}_{2n}, \hat{\uptau}_n \big)^{\prime}$, where $\hat{\beta}_{jn} = \hat{\beta}_{jn} \left( \hat{\uptau}_n \right)$ are the coefficient estimators for $j \in \left\{ 1, 2 \right\}$. Moreover, the size of the jump will be estimated and denoted with $\hat{\delta}_n = \left( \hat{\beta}_{1n} - \hat{\beta}_{2n} \right)$. Notice that $\hat{\beta}_1$ has a scaled Dickey-Fuller distribution. The term $\hat{\beta}_2$ can be described as being asymptotically normally distributed with a random variance, which occurs due to the asymptotic limit given by the following expression:
The numerator follows a CLT and the denominator converges to a random variable. Moreover, the maintained hypothesis is that the shift exists, which implies that $\delta \neq 0$. Consider for example the set of these indicator functions such that $\boldsymbol{1} \left\{ t \leq \uptau n \right\}$ and $\boldsymbol{1} \left\{ t \leq \hat{ \uptau } n \right\}$ where $\hat{ \uptau }$ is an estimator of the unknown break fraction $\uptau$. Then, showing that for the OLS estimator of the predictive regression model it holds that $n \left( \hat{\uptau} - \uptau \right) = \mathcal{O}_p(1)$ implies also that $ \floor{ \hat{\uptau} n } = \floor{ ( \hat{\uptau} - \uptau )n + \uptau n } = \floor{ O_p(1) + \uptau n } = \floor{ \uptau n } + o_p(1)$.
Relevant frameworks which consider a structural change type estimation and inference of a first-order autoregressive model includes the framework proposed by kurozumi2023fluctuation. Although the particular framework corresponds to a fluctuation type monitoring test\footnote{Notice that the monitoring testing approach (see, chu1996monitoring, leisch2000monitoring, aue2009delay and horvath2020sequential among others) corresponds to a different implementation and estimation procedure of structural change in time series regressions. One of the main difference is the use of a historical and a monitoring period during which model estimates and residuals are constructed with the purpose of detecting structural breaks. We leave these considerations as future research. Another aspect worth emphasizing is that in this article the sup-Wald type statistics correspond ton an iterative estimation step in order to construct a sequence of test statistics based on fitting the regression model within the full-sample and comparing the model estimates across two subsamples that correspond to the pre-break and post-break part. } for detecting the presence of explosive behaviour in time series data (see, arvanitis2018mildly and skrobotov2023testing). On the other hand, in this article we consider the statistical properties of break-point estimators in predictive regression models within the full-sample, so these two aspects are considered to be the main contributions of our study. Furthermore, our work can be useful in relevant applications from the financial economics literature since knowing the exact distributional properties of the break-point estimators when the econometrician employs a predictive regression model either for detecting slope instabilities or testing for predictability robust against parameter instability.
Therefore, we can see from the limit result given on Theorem 1 that $\hat{\beta}_1 ( \hat{\uptau}_n )$ and $\hat{\beta}_2 ( \hat{\uptau}_n )$ are both asymptotically normally distributed with variance depending on $\beta_1$, $\beta_2$ and $\uptau_0$.
Notice that in the model the change point $\uptau_0$ is unknown and has to be estimated by $\hat{\uptau}_n$. Thus, to show the consistency of $\hat{\uptau}_n$, the common practice in the structural break literature is to show that $( 1 / n ) RSS_n( \uptau )$ converges uniformly to a nonstochastic function that has a unique minimum at $\uptau = \uptau_0$. Therefore, we focus on the asymptotic behaviour of the quantity $( 1 / n ) RSS_n( \uptau )$. The following Lemma is useful in deriving the limiting behaviour of the criterion function $( 1 / n ) RSS_n( \uptau )$ and in proving Theorem 1, which follows.
Consider the criterion function $\left( 1 / n \right) RSS ( \uptau)$ when $\beta_2 = 1$ behaves very differently. Then, under Assumptions (A1)-(A3) we have that
Following the econometric framework of chong2001structural and pang2021estimating, further limit results which will need to be established in the case one replaces the assumption of a stationary autoregressive coefficient with the local-to-unity parametrization includes:
Notice that there is an asymptotic gap between $( 1 / n) RSS_n( \uptau_0 )$ and $( 1 / n ) RSS ( \uptau )$. Thus to examine the consistency of $\hat{\beta}_1$, we have to investigate the transitional behaviour of $( 1 / n ) RSS_n ( \uptau )$. Note that for any constant $c > 0$,
where
When $\alpha < 1 / 2$, it holds that $\theta_n \left( \alpha, c \right) \overset{ p }{ \to } 1, \ \ \ \hat{\beta}_1 \left( \uptau_0 + c n^{ \alpha - 1 } \right) \overset{ p }{ \to } \beta_1$, and $\frac{1}{n} RSS_n \left( \uptau_0 + c n^{ \alpha - 1 } \right) \overset{ p }{ \to } \sigma^2$. Moreover, for $\frac{1}{2} < \alpha < 1$ we have that
The above results simply imply that if the convergence rate of $\hat{\uptau}_n$ is faster that $n^{1/2}$ then $\hat{\beta}_1$ will be consistent, otherwise it will be inconsistent. Next, we derive the asymptotic behaviour of the criterion function $( 1 / n ) RSS_n ( \uptau )$. Another useful quantity is as below:
Clearly, the challenging task here is that when we impose that assumption that the autoregressive coefficient of the first order autoregression model is expressed via the local-to-unity parametrization, then we expect that the break-point estimators will depend on the nuisance parameter of persistence. In other words, the fact that the asymptotic distribution of these break-point estimators will not be nuisance parameter-free can be challenging when critical values are needed (e.g., case of the monitoring scheme). We leave these considerations for future research.
Following the existing literature to find the limiting distribution of $\hat{\beta}_1 ( \hat{\uptau}_n )$, in the case of the stationary AR$(1)$ model, notice that $\left( \hat{\uptau}_n - \uptau_0 \right) = \mathcal{O}_p \left( n^{-1} \right)$ and it holds that
From the above derivations we can see that $\hat{\beta}_1 ( \hat{\uptau}_n )$ and $\hat{\beta}_1 ( \uptau_0 )$ have the same asymptotic distribution, since $\left\{ y_{t-1} \epsilon_t, \mathcal{F}_t \right\}_{t=1}^{\floor{ n \uptau_0 } }$ is a martingale difference sequence, with $\mathbb{E} \left[ y_{t-1} \epsilon_t | \mathcal{F}_{t-1} \right] = 0$ and $\sum _{t=1}^{ \floor{ n \uptau_0 } } \mathbb{E} \left[ \left( y_{t-1} \epsilon_t \right)^2| \mathcal{F}_{t-1} \right] \overset{ p }{ \to } \frac{ \sigma^4 }{ 1 - \beta_1^2 } < \infty$. Applying the central limit theorem for martingale difference sequences and the fact that
Next, to find the limiting distribution of $\hat{\beta}_2 ( \hat{\uptau}_n )$, notice that $\left( \hat{\uptau}_n - \uptau_0 \right) = \mathcal{O}_p ( \frac{1}{n} )$ and
It can be prove that both $\hat{\beta}_1 ( \hat{\uptau}_n )$ and $\hat{\beta}_2 ( \uptau_0 )$ have the same asymptotic distribution, which implies
By the central limit theorem for martingale difference sequences and by independence of the two martingale differences given previously, we have that
The first term above weakly converges to $\left[ \sigma^2 B^2 ( \uptau_0 ) \right] / \left( 1 - \beta_2^2 \right)$. Furthermore, because $\left| y_{k_0} / \sqrt{n} \right| = \mathcal{O}_p(1)$ and $\left( 1 / \sqrt{n} \right) \text{sup}_{ t > k_0 } \left| \epsilon_t \right| = o_p(1)$, then we can show that the second term is bounded by
Although in this paper we assume that the innovation sequence of the model has a linear process representation, other studies in the literature imposes a NED condition when developing structural break tests in stationary time series regression models (see, ling2007testing, kim2020mean and lee2014functional). Notice that the IVX estimator has the property that decorrelates the system and therefore conventional invariance principles and weak convergence results hold without requiring to consider a topological convergence in a different space. Furthermore, another important feature of the predictive regression model is that one can incorporate serial correlation in the error term and therefore using the IVX instrumentation mixed Gaussianity distributional converges still holds.
At this point, we begin our asymptotic theory analysis by investigating the corresponding expressions of the criterion function when the functional form of the linear predictive regression model with a conditional mean function is employed. We write the residual sum of squares as below:
which after expanding out, the RSS expression can be written as below:
Furthermore, we need to determine the validity of the following expansion in the case of the predictive regression model.
Furthermore, we consider the asymptotic behaviour of the break-estimators when we are testing the stability of the model parameters of the predictive regression model based on the IVX estimator of the predictive regression model. In particular, within our framework we implement the IVX filter proposed by PM (2009) which implies the use of a mildly integrated instrumental variable that has the following form
for some $c_z > 0$ and $0 < \delta < 1$, where $\delta$ is the exponent rate of the persistence coefficient which corresponds to the instrumental variable. Notice that the above filtering method transforms the autoregressive process $x_t$, which can be either stable or unstable, into a mildly integrated process which is less persistent a Nearly Integrated array such as the case of $x_t$.
The case study we illustrate in the previous section is based on the specific approach proposed by pang2021estimating. In this article since the estimation methodology of model coefficients is taken into consideration, then we need to establish the asymptotic theory separately for the break-point estimator which based on the OLS optimization versus the break-point IVX optimization, in order to evaluate the consistency and statistical properties of the estimated break-point. Our objective in is to evaluate the properties of $k_1$ which is an estimator of the location of the break-point $k_1^0$ in the slope parameters and intercept of the predictive regression model (see, kostakis2015robust).
In particular, the estimator of the break-point is obtained by minimizing the concentrated sum of squares errors function as
where $\hat{\beta}_1 (k)$ and $\hat{\beta}_2 (k)$ denote the least squares estimators of the slope parameters within each regime for given $k$. Alternatively, we can also reformulate $\hat{k}_1$ as below
where $S_n$ denotes the full sample sum of squared errors. Thus, we are interested to establish the weak consistency of $\hat{\uptau} = \hat{k}_1/n$ under a certain set of assumptions. This formulation allows us to establish the weak consistency of the break fraction $\hat{\uptau} = \hat{k}_1 / n$ under certain set of assumptions. Thus, to have more meaningful power comparisons we study in more details the asymptotic behaviour of the break estimators under the alternative hypothesis based on the two estimators. To do this, we follow the methodology proposed with the framework of pang2021estimating, but in our setting we focus in the case of a single unknown break-point. We consider separately the break-point estimator based on the OLS versus the IVX estimators to evaluate the consistency of the estimated break-point and obtain corresponding convergence rates\footnote{The finite sample properties of $\hat{\uptau}_1$ is important not only because of the direct economic implications that the accurate dating of a structural break in the mean may entail but also for the subsequent analysis which could involve the search for further breaks in the variance, typically based on the residual sequence (see, pitarakis2004least).}.
Step 1: For any given $0 < \tau < 1$, denote with
Then the change-point estimator of $\uptau_2^0$ is defined as below
where
Once we obtain $\hat{\uptau}_{2,n}$, the least squares estimator of $\beta_3$ is represented by $\hat{\beta}_3^{OLS} \left( \hat{\uptau}_{2,n} \right)$ and the OLS of $k_2^{0}$ is denoted by $\hat{k}_2 = [ \hat{\uptau}_{2,n} n ]$.
Step 2: For any given $0 < \tau < \hat{\tau}_{2,n}$, the OLS estimators of the parameters $\beta_1$ and $\beta_2$ are given by
respectively. Then the change-point estimator of $\uptau_1^0$ is defined as below
where
Once we obtain $\hat{\uptau}_{1,n}$, the least squares estimator of $\beta_1$ and $\beta_2$ are represented by $\hat{\beta}_1^{OLS} \left( \hat{\uptau}_{1,n} \right)$ and $\hat{\beta}_2^{OLS} \left( \hat{\uptau}_{1,n} \right)$ respectively, and the OLS of $k_1^{0}$ is denoted by $\hat{k}_1 = [ \hat{\uptau}_{1,n} n ]$.
Step 1: For any given $0 < \tau < 1$, denote with
Then the change-point estimator of $\uptau_2^0$ is defined as below
where
Once we obtain $\hat{\uptau}_{2,n}$, the IVX estimator of $\beta_3$ is represented by $\hat{\beta}_3^{IVX} \left( \hat{\uptau}_{2,n} \right)$ and the IVX of $k_2^{0}$ is denoted by $\hat{k}_2^{IVX} = [ \hat{\uptau}_{2,n} n ]$.
Step 2: For any given $0 < \tau < \hat{\tau}_{2,n}$, the IVX estimators of the parameters $\beta_1$ and $\beta_2$ are given by
respectively. Then the change-point estimator of $\uptau_1^0$ is defined as below
Once we obtain $\hat{\uptau}_{1,n}$, IVX estimator of $\beta_1$ and $\beta_2$ are represented by $\hat{\beta}_1^{IVX} \left( \hat{\uptau}_{1,n} \right)$ and $\hat{\beta}_2^{IVX} \left( \hat{\uptau}_{1,n} \right)$ respectively, and the IVX of $k_1^{0}$ is denoted by $\hat{k}_1^{IVX} = [ \hat{\uptau}_{1,n} n ]$.
Therefore, notice that each of the covariates of the predictive regression are modelled using the autoregressive process $x_t = \rho_n x_{t-1} + v_t$, $x_0 = 0$. Specifically, when $\rho_n = \rho$ such that $| \rho | < 1$ then, $x_t$ is a stationary weakly dependent process. However, the present study focuses on cases where $x_t$ is nonstationary. In particular, in these cases $\rho_n = \left( 1 + \frac{c}{n} \right)$. More precisely, if the autoregressive parameter is fixed with $\rho_n = \rho$ and $| \rho | < 1$, then $x_t$ is asymptotically stationary and weakly dependent. Therefore, this framework helps to unify structural break testing in predictive regression models in cases where the properties of the predictor is not known thus offering robustness to integration order duffy2021estimation.
where $Y_{t} \in \mathbb{R}$ is an 1$-$dimensional vector and $\boldsymbol{x}_t \in \mathbb{R}^{p \times n}$ is a $p-$dimensional vector of local unit root regressors, with an initial condition $\boldsymbol{x}_0 = 0$. Moreover, $\boldsymbol{C} = \mathsf{diag} \{ c_1,...,c_p \}$ is a $p \times p$ diagonal matrix which determines the degree of persistence of the regressors by the unknown persistence coefficients $c_i$'s which are assumed to be positive constants. Define with $\boldsymbol{\eta}_t = \left( u_{t}, \boldsymbol{v}_{t}^{\prime} \right)^{\prime}$. Then, the partial sum process constructed from $\boldsymbol{\eta}_t $ satisfies a multivariate invariance principle. That is, for $r \in [0,1]$ and as $n \to \infty$ we have (where $\Rightarrow$ denotes weak convergence in distribution),
where B(r) is a $p-$dimensional Brownian motion with covariance matrix
Now $\boldsymbol{\Omega}$ and $\boldsymbol{B}(r)$ are partitioned as below
with partitions given by
Furthermore, note that it can be proved that
Another important aspect is to consider the case of weakly dependent errors which applies removing the independence assumption of the error sequence $\left\{ \epsilon_t \right\}_{ t = 1}^n$. Specifically, the case of weakly dependent errors assumes the following linear process representation
Furthermore, to ensure that $\epsilon_t$'s are weakly dependent error terms we assume that $a(1) := \sum_{ j = 0}^{ \infty } a_j \neq 0$, $\sum_{ j = 0}^{ \infty } j^{3 / 2} | a_j | < \infty$, and $\left\{ e_t \right\}_{ t = 1}^n$ is a sequence of i.i.d random variables with mean zero and variance $0 < \sigma_e^2 < \infty$. Moreover, the consistency of three estimators that correspond to the three regimes, when the estimated break is incorporated, needs to be established such that (Theorem 4.1 in pang2021estimating)
where $a(2) = \sum_{ j = 0 }^{ \infty } a^2_j$, and $\xi$ and $\zeta$ are two independent standard Cauchy variates.
Local to Unity
Consider the following predictive regression model
where $\beta_1 = \left( 1 - \frac{\gamma}{n} \right)$ with $\gamma \in \mathbb{R}$, $\beta_2 = \beta_{2n} = \left( 1 + \frac{c_1}{k_n} \right)$ with $c_1 > 0$, $\beta_3 = \beta_{3n} = \left( 1 - \frac{c_2}{h_n} \right)$ with $c_2 > 0$ and $u_t = cn^{- \eta} + \epsilon_t$ with $c \in \mathbb{R}$ and $\eta > 1 /2$ when $t \leq k_1^0$, and $u_t = \epsilon_t$ when $t > k_1^0$.
Then $\hat{\beta}_1 \left( \hat{\tau}_{1,n} \right)$, $\hat{\beta}_2 \left( \hat{\tau}_{1,n} \right)$ and $\hat{\beta}_3 \left( \hat{\tau}_{2,n} \right)$ are all consistent, and their asymptotic distributions are respectively given by
where $\mathcal{X}$ and $\mathcal{Z}$ are two random variables obeying $\mathcal{N} \left( 0, \frac{1}{2c_1} \right)$ and $\mathcal{N} \left( 0, \frac{1}{2c_2} \right)$ respectively. Moreover, the random variables $\mathcal{X}$, $\mathcal{Z}$ and $\left\{ W(s), 0 \leq s \leq \tau_1^0 \right\}$ are mutually independent.
Additional Hypothesis Testing as in pang2021estimating:
We discuss a hypothesis test problem which is important for testing for the equality of the persistence in the two shifting regimes. Suppose that $k_n = n^{ \alpha_1 }$ with $0 < \alpha_1 < 1$ and $h_n = n^{ \alpha_2 }$ with $0 < \alpha_2 < 1$. Therefore, the null hypothesis is given by $H_0: c_1 = c_2 = 0$. Under the null, this implies that the underlying stochastic process has an autoregressive unit root throughout the sample. The alternative hypothesis can be formulated as below $H_0: c_1 \neq 0$ and $c_2 \neq 0$. Therefore, under the alternative hypothesis the underlying process is generated by the AR(1) model with $k_n = n^{ \alpha_1 }$ and $h_n = n^{ \alpha_2 }$. Denote with $\hat{ \beta } = \frac{ \sum_{t=1}^n y_t y_{t-1} }{ \sum_{t=1}^n y^2_{t-1} }$, and the associated test statistic is given by
Furthermore, it is well known that under $H_0$, we have that
which implies that the the t-ratio given by the random variable $t_n$ is bounded in probability under $H_0$. On the other hand, since we can prove that $| t_n |$ will go to infinity in probability under $H_1$, it implies that the t-statistic has the ability to disciminate between the data generated from a unit root model vis-a-vis the date generated from the model specification above.
Our main research objective is to investigate the statistical properties and asymptotic behaviour of break-point estimators when a single change-point occurs in predictive regression models.
In this article we consider testing for structural break in univariate time series regressions models, an active research literature considers statistical methodologies for estimation and inference in change-point models for multivariate time series (see, preuss2015detection among others). We leave such considerations and extensions of our framework as future research. Other extensions include consider a sequence of break-point estimators in a high-dimensional predictive regression model using shrinkage type estimators to obtain the break-point date estimators (see, the excellent framework proposed recently by tu2023penetrating\footnote{Although the term "sporadic predictability", roughly speaking reflects predictability adapted to the changing environment based on the myriads of global and local macroeconomic shocks not to mention that the first word of the particular paper title is mildly not appropriate for an economics paper.}, see also the discussion on the shrinkage inference approach in katsouris2023quantile and the proposed shrinkage methodology presented in chen2017estimation, nkurunziza2010shrinkage as well as nkurunziza2021inference).