EconBase
← Back to paper

Predictability Tests Robust against Parameter Instability

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

117,776 characters · 22 sections · 105 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Predictability Tests Robust against Parameter Instability

\pagenumbering{roman}

center[center omitted — 31 chars of source]

We consider Wald type statistics designed for joint predictability and structural break testing based on the instrumentation method of phillipsmagdal2009econometric. We show that under the assumption of nonstationary predictors: (i) the tests based on the OLS estimators converge to a nonstandard limiting distribution which depends on the nuisance coefficient of persistence; and (ii) the tests based on the IVX estimators can filter out the persistence under certain parameter restrictions due to the supremum functional. These results contribute to the literature of joint predictability and parameter instability testing by providing analytical tractable asymptotic theory when taking into account nonstationary regressors. We compare the finite-sample size and power performance of the Wald tests under both estimators via extensive Monte Carlo experiments. Critical values are computed using standard bootstrap inference methodologies. We illustrate the usefulness of the proposed framework to test for predictability under the presence of parameter instability by examining the stock market predictability puzzle for the US equity premium. \\

JEL classification: C12, C22, C26, C32, C53, C58.

Keywords: Stock Return Predictability, Parameter Instability, Predictive Regression, IVX filtration, Local-to-unity asymptotic theory, Ornstein-Uhlenbeck process, bootstrap.

\setcounter{page}{1} \pagenumbering{arabic}

Introduction

The stock predictability puzzle is an important research theme appeared in the financial econometrics literature which has seen growing attention in recent years\footnote{In empirical finance the predictive regression model provides a suitable testing framework for aspects such as conditional asset pricing and investment strategy performance evaluation. An extensive study about the latter is presented by pesaran1995predictability. The author examines the predictability of stock returns from the perspective of optimal portfolio decision. Moreover, the author introduces the concept of episodic predictability, arguing that predictability can be explained in accordance to certain macroeconomic events. Moreover, related applied macroeconometric applications include the forecasting of macroeconomic fundamentals as well as the examination of short-horizon vis-a-vis long-horizon predictability.}. Various studies have demonstrated that forecast performance via predictive regressions with macroeconomic variables as predictors, is countercyclical, that is, there is a co-movement with the business cycle. Further studies focused on identifying periods of episodic predictability; a terminology adapted to describe the phenomenon of predictors "switching on" during certain periods, while appear to have no predictive ability during other periods (e.g., see gonzalo2012regime, gonzalo2017inferring, chinco2019sparse, demetrescu2020testing). These periods of unstable forecasting ability appear in time series models in the form of parameter instability (see, rossi2012out, inoue2017rolling, pitarakis2017simple and georgiev2018testing). Considering this vibrant discussion, we aim to study how the degree of persistence affects the asymptotic theory of related to the above aspects. Our goal is to tackle the following econometric question: "How does the time series properties of predictors affect predictability testing under the presence of parameter instability?"

The predictive regression model operates under the strong assumption of parameter stability, which could be violated in certain regions of the sample\footnote{A related study to parameter instability in prediction models is presented by paye2006instability. In this paper, the authors propose a framework for identifying and estimating multiple breaks.}. Moreover, the majority of the structural break literature propose testing methodologies under the assumption of stationary time series. Therefore, the development of a robust joint test against both predictability and structural break is an important aspect to tackle. Related issues from the predictability literature, include the high persistence of predictors as well as the econometric implications of testing for structural break given the existence of nonstationary regressors as captured via the local-to-unit root specification of the predictive regression. In this paper we propose an econometric framework for jointly testing against both predictability and structural break. The tests are constructed in a similar manner as the Wald type statistics proposed by gonzalo2012regime in a threshold predictive regression model\footnote{The particular framework and in the subsequent paper of gonzalo2017inferring the authors propose tests which capture the effects of linearity and the presence of a threshold effect in predictive regressions with persistent predictors.}. Our contributions are threefold: (i) We propose a test statistic which jointly tests against the alternative hypothesis of both predictability and structural break and show that the test is robust to the persistence properties of predictors in the model;

(ii) We show that the test statistic has a nuisance parameter free limiting distribution under the assumption of stationary and mildly stationary regressors, while it converges to a nonpivotal asymptotic distribution under the assumption of local-to-unit root (LUR) or integrated regressors; (iii) We provide an extensive Monte Carlo simulation study where we compare the empirical size and power of the sup IVX-Wald test. We also present simulations to examine the finite-sample properties of the corresponding sup OLS-Wald statistic. Lastly, we employ the proposed testing framework to investigate the predictability puzzle based on the equity premium of US stock returns.

\paragraph{Related Literature}

The ideas in this paper are related to research done in two different fields. From finance, it is related to the stock return predictability literature and from the econometrics and statistics perspective it is related to the literature of parameter instability and structural break testing methodologies. We aim to bridge the gap in the literature by proposing a testing framework which incorporates both aspects.

Firstly, we provide a general overview of the predictability literature, starting with the implementation of simple t-tests to detect statistical significance in stable relations of predictants such as equity index returns, on the lagged time series of predictors. From extensive empirical applications, this practise has been proved to cause distorted inference due to the presence of high persistent predictors, since nonstandard terms appear in the limiting distribution of the t-test. In particular, stambaugh1999predictive observed this finite sample bias\footnote{Notice that the Stambaugh bias correction is based on the studies of marriott1954bias and kendall1954note who proposed suitable bias correction in autocorrelations.} which occurs when the classical least squares estimator is employed for statistical inference. Furthermore, amihud2004predictive consider a second-order bias-correction and propose a reduced-bias OLS based estimator. Both aforementioned methods are considered to provide a post-estimation bias correction in finite samples.

Secondly, the predictability literature was extended to nonstandard inference to account for the presence of nonstationary predictors; some of the most notable contributions include the Bonferroni-type approach as in cavanagh1995inference and campbell2006efficient, the conditional likelihood method based on sufficient statistics\footnote{The idea of sufficient statistics considers optimal tests invariant under transformation based on the curved exponential family. This framework allows to use the conditional restrictions testing problem in the presence of nuisance parameters and is particularly appealing in the case of near integrated regressors.} proposed by jansson2006optimal and the control function method proposed by elliott2011control. The drawback of these testing methodologies is that the asymptotic theory has some undesirable properties such as the uncorrectable bias due to the presence of nuisance parameter in the limit distribution (see, phillips2013predictive). Moreover, the particular approaches are computational complex especially in multivariate predictive regression settings. For instance, kasparis2015nonparametric propose a set of nonparametric predictive tests which allows for nonlinearities and offers robustness to integration order.

Thirdly, a novel approach that recently has been attracting much attention is the instrumental variable-based test proposed by kostakis2015Robust (KMS, hereafter), which is build upon the theoretical framework developed by phillipsmagdal2009econometric. This methodology, referred to as IVX-Wald test, provides a robust framework for predictive regression models which is valid for predictors with general persistence properties. Specifically, the asymptotic theory shows that the IVX estimator which converges to a mixed Gaussian distribution, successfully removes the long-run endogeneity, that appears due to the innovation structure of the model, and provides a pivotal statistic robust under different degrees of persistence or even regressors of mixed integration (e.g., see phillips2016robust). Hence, a self-normalized Wald statistic can be constructed that converges to a nuisance parameter free $\chi^2$ limiting distribution. More importantly, the IVX filtration implies a direct inference procedure via the various moment approximations (e.g., long-run covariance matrices) and can be easily extended to the multivariate predictive regression model under certain regulatory conditions.

The aforementioned literature operates under the assumption of parameter constancy\footnote{Note that, the literature of structural change goes back to 1940s and 1950s with the pioneering work of Wald on sequential hypothesis testing as well as the seminal work of page1954continuous which proposes methodologies for detecting anomalies in control charting, an idea further developed by chu1995moving and chu1996monitoring who consider testing for structural break in the sense of contaminated and non-contaminated periods in linear regression models. Further seminal studies for structural break tests include chow1960tests, hawkins1987test and ploberger1990local among others.} which implies a stable predictive relationship over the sample period. However, due to the nature of economic conditions, shifting between periods of market tranquillity and periods of market exuberance, the phenomenon of episodic predictability\footnote{Note that, formal econometric methods to test for episodic predictability have been recently proposed by demetrescu2020testing, who consider related subsampling techniques.} has been proposed to capture these "pockets" of predictability across business cycles (see, farmer2019pockets). This implies the existence of time-varying predictability which can be examined within an econometric framework which accommodates time-varying parameters. In this paper, we approach this aspect by proposing to test for the existence of joint predictability and structural break, in the form of a single parameter shift at an unknown break-point. Thus, our approach is closer to the framework of andrews1993tests who proposed Wald-type statistics for testing for a single structural break at an unknown break-point.

The presence of structural break in predictive regression models implies variation in predictability across time. As a result, standard testing methodologies for break detection, such as Andrew's Wald statistics for linear regression models as well as the family of tests proposed by bai1998estimating can no longer be valid for nonstationary\footnote{The frameworks proposed by phillipsmagdal2009econometric and magdalinos2009limit explicitly examine the various forms of nonstationarity attributed due to the structure of the predictive and cointegrated system, as captured by specific characteristics of the system. This is a major distinction in the literature which previously thought that the presence of nonstationarity in the regressors as a charateristic occurring due to the autoregressive specification of the model predictors solely.} predictors

as captured by the autoregressive specification of the predictive regression model. This invalidity is demonstrated by the simulation studies of paye2006instability. More specifically, the authors show that in the case of highly correlated innovations and (unfiltered) persistent predictors, an implementation of a sup-F test and UDMax test can cause severe size distortions when testing for structural breaks in time series. Moreover, the optimal test of Elliot and Muller has good empirical size performance, but distortions appear in the asymptotic properties of the power function of the test.

The literature of structural break tests has recently focused in the construction of suitable testing frameworks for predictive regression models. cai2015testing propose modelling smooth structural breaks using an $L_2-$ type statistic, in predictive regressions with nonstationary predictors. A different approach include the study of pitarakis2017simple who consider a CUSUM-type statistic. Additionally, due to the occurrence of nonstandard limiting distributions in standard structural break test implementations, georgiev2018testing and georgiev2019bootstrap propose the use of a fixed regressor wild bootstrap procedure\footnote{Notice that the particular fixed regressor wild bootstrap procedure is an extension to the fixed regressor bootstrap approach employed by hansen2000testing to deal with the appearance of a nonstandard limiting distribution due to nonstationary regressors.} to approximate critical values. Another related study to the bootstrapping approach of structural break tests is presented by boldea2019bootstrapping. However, the proposed framework is restricted to the case of exogenous regressors.

Despite the recent developments of the structural break literature for predictive regressions, these methodologies consider the detection of parameter instability in predictive regression models without simultaneously testing whether the tests are robust to the presence of predictability around the break-point location. The only existing study that jointly tests for the existence of both effects is the paper of demetrescu2020testing, who propose to combine testing for predictability using appropriate subsampling techniques\footnote{The subsampling technique is similarly employed by davidson2010tests in the context of detecting structural break due to break in cointegration of the relation under investigation.}(see also similar techniques employed by hansen2000sample). The authors propose to use bootstrap-based inference due to the existence of a non-pivotal limiting distribution, which is a robust approach to provide statistical validity of the tests.

Further aspects related to the IV based approach of KMS that we follow in this paper, can be also examined within the above parameter instability testing procedures. Recent applications include magdalinos2020least who consider predictability tests with GARCH-type effects (see also, gungor2020small ) as well as yang2020testing who consider a modification of the KMS test that accounts for serial correlation in the error term of the linear predictive regression. Moreover, pang2020estimating propose testing methodologies for multiple structural breaks under the presence of nonstationary predictors with an application to financial bubble detection. All these features demonstrate additional refinements one can consider as future research related to our proposed framework.

\paragraph{Outline} The paper is organized as follows. Section (ref), presents the predictive regression model along with the background assumptions. This Section also includes a review of the IVX instrumentation procedure of KMS, which is employed for the construction of the proposed test statistic. Section (ref), presents the asymptotic theory of the test statistic under different degrees of regressors persistence. Section (ref) presents an extensive Monte Carlo simulation study. Section (ref) illustrates an application to the US stock returns which provides evidence of the empirical relevance of the proposed tests for jointly testing predictability and structural break. Section (ref) concludes. All the mathematical proofs and related results are included in the Appendix of the paper.

Econometric Model and main assumptions

In this section we present the theoretical background within which we operate. We denote with $\left( \Omega, \mathcal{F}, \mathbb{P} \right)$ a suitable probability space on which all of the random elements are defined\footnote{Notice that this is an important assumption, since our framework can be extended to Hilbert spaces within which we can denote the expectation operator and higher moments as inner products. In that scenario, structural break has a slightly different interpretation. We leave this as a future research endeavour. }. Moreover, throughout the paper, all limits are taken as $T \to \infty$, where $T$ is the sample size. The symbol $"\Rightarrow"$ is used to denote the weak convergence of the associated probability measures as $T \to \infty$. The symbol $\overset{d}{\to}$ denotes convergence in distribution and $\overset{\text{plim}}{\to}$ denotes convergence in probability, within the probability space. Let $\left\{ Y_t, X_t \right\}_{t=1}^T$ denote the corresponding random variables of the underline distributions.

Predictive Regression Model

Consider the predictive regression model with a possible single structural break

align[align omitted — 172 chars of source]

where $y_{t} \in \mathbb{R}$ is an one$-$dimensional vector and $x_t \in \mathbb{R}^{p \times 1}$ is a $p-$dimensional vector of predictors, with an initial condition $x_0 = \mathcal{O}_p(1)$ which we assume that does not affect the limit theory. Moreover, the autocorrelation matrix is expressed as below

align[align omitted — 112 chars of source]

where $\mathbf{C} = diag \{ c_1,...,c_p \}$ is a $p \times p$ matrix. The degree of persistence in the regressors is determined by the unknown persistence coefficients $c_i$'s which are assumed to be positive constants and the exponent rate $\gamma_x \in \mathbb{R}$, that is, $\gamma_x < 1$, $\gamma_x \in (0,1)$ or $\gamma_x > 1$.

The predictive regression model given by (ref)-(ref) accommodates the existence of a single unknown structural break at location $k$. Therefore, the model parameters take the following form which indicate a time-varying parameter vector (i.e., parameter instability)

align[align omitted — 234 chars of source]

where $\beta_1 \in \mathbb{R}^{p \times 1}$ and $\beta_2 \in \mathbb{R}^{p \times 1}$ while for the univariate model $\beta_1 \in \mathbb{R}$ and $\beta_2 \in \mathbb{R}$.

To develop asymptotic theory suitable for robust inference in the predictive regression model under the presence of parameter instability and nonstationarity we impose the following assumptions and regulatory conditions as described below.

assumptionLet $e_t = \left( u_{t}, v_{t}^{\prime} \right)^{\prime}$ be a $(p+1)-$dimensional vector. The innovation sequence $e_t$ is a conditionally homoscedastic martingale difference sequence ( m.d.s ) such that the following two moment conditions hold: \begin{enumerate} • $\mathbb{E} \left[ e_{t} | \mathcal{F}_{t-1} \right] = 0$, where $\mathcal{F}_t = \sigma \left( u_t, u_{t-1},... \right)$ is the natural filtration. • $\mathbb{E} \left[ e_{t} e_{t}^{\prime} | \mathcal{F}_{t-1} \right] = \mathbf{\Sigma}_{ee}$, where $\mathbf{\Sigma}_{ee} \in \mathbb{R}^{(p+1) \times (p+1)}$ is a positive-definite covariance matrix, which has the following form: \begin{align*} \mathbf{\Sigma}_{ee} = \begin{bmatrix} \sigma_u^2 & \sigma^{\prime}_{uv} \\ \sigma_{vu} & \mathbf{\Sigma}_{vv} \end{bmatrix} > 0. \end{align*} with $\sigma_u^2 \in \mathbb{R}$, $\sigma_{uv} \in \mathbb{R}^{p \times 1}$ and $\mathbf{\Sigma}_{vv} \in \mathbb{R}^{p \times p}$, where $p$ is the number of predictors. \end{enumerate}

Assumption A1 indicates that the innovation vector is a martingale difference sequence while A2 indicates that the innovation sequence is conditionally homoscedastic. Under conditions A1 and A2, the following Functional Central Limit Theorem (FCLT) applies

align*[align* omitted — 311 chars of source]

where $\mathbf{\Sigma}_{ee}^{1/2} W(r)$ is a $(p+1)-$dimensional Brownian motion with covariance matrix $\mathbf{\Sigma}_{ee}$.

The above error structure provides a realistic interpretation of macroeconomic shocks. In other words, the shocks to $y_{t}$ and $x_{t-1}$, that is, $u_{t}$ and $v_{t}$ respectively, appear to be contemporaneously correlated, a commonly used assumption in predictive regression models. Related definitions can be found in phillips1986multiple, phillips1987time and phillips1988regression. Moreover, the implementation of a FCLT\footnote{A standard FCLT for linear regression models is introduced by Theorem 7.17 in white2001asymptotic.} allows to derive the limiting distribution of Wald type statistics for jointly testing predictability and structural break within the framework of this paper.

assumptionThe innovation to $x_t$ is a linear process with the following representation \begin{align*} v_t := \Phi(L) \epsilon_t \equiv \sum_{j=0}^{\infty} \Phi_j \epsilon_{t-j}, \ \ \ \epsilon_t \sim^{i.i.d} \left( 0, \mathbf{\Sigma}_{ee} \right) \end{align*} where $\left\{ \Phi \right\}_{j=0}^{\infty}$ is a sequence of absolute summable constant matrices such that $\sum_{j=0}^{ \infty} \Phi_j$ has full rank and $\Phi_0 = I_p$ with $\Phi (1) \neq 0$, allowing for the presence of serial correlation.

To allow for the detection of possible structural breaks via the predictive regression model and avoid the presence of nuisance parameters at the boundary of the parameter space we also impose the following assumption.

assumptionLet $k = \pi T$, where $0 < \pi < 1$. Then, the fraction of structural break defined as $\pi_0 = k_0 / T$ is within the interior of $(0,1)$ for some fixed $\pi_0$ parameter\footnote{Notice that this mathematical statement does not imply that we consider only known break-points.}.
assumptionThe estimator $\widehat{\pi}$, is $\mathcal{F}_t-$measurable and $[0,1]-$valued. Then, consistent estimation of the break-point implies that $\exists$ some $m_0$ such that $\widehat{\pi} - \pi = \mathcal{O}_p \left( T^{-m_0} \right)$.
remarkAssumption (ref) gives the linear process representation of innovations proposed by phillips1992asymptotics and can accommodate the case of serial correlation or even conditional heteroscedasticity with suitable specifications. Assumption (ref) excludes structural breaks at the boundaries of the sample. Assumption (ref) ensures the existence of a consistent estimator of the unknown break-point which imply that $\underset{ T \to \infty }{ \text{lim} } \mathbb{E} \big[ | \hat{\pi} - \pi | \big] = 0$. For instance, assuming that we are testing for an unknown break-point, to avoid unidentified parameters under the null hypothesis the supremum functional is employed for the construction of the Wald type statistics we propose in the next section.

When the time-varying parameters $\alpha_t$ and $\beta_t$ are both constant over the full sample, that is, $\alpha_1 = \alpha_2 = \alpha$ and $\beta_1 = \beta_2 = \beta$, then the regression specification given by (ref)-(ref) reduces to the standard predictive regression model. In that case, the hypothesis testing of interest is finding statistical evidence against the null hypothesis of no predictability\footnote{For instance, the nonparametric predictive test of kasparis2015nonparametric is constructed by comparing the nonparametric functional form estimator with the parametric counterpart estimator. In our setup, the estimation procedure is simpler to that aspect, however it still requires to carefully examine the stochastic properties of the Wald type statistics under nonstationarity and the presence of a single structural break.}. Furthermore, it has been shown in the literature when the ordinary least squares estimation is employed, then statistical inference is nonstandard due to the fact that the limiting distribution of the tests depends on the unknown persistence parameter, $c_i$, and results to uncorrectable asymptotic bias. Since in our setup we are interested in detecting a single structural break an appropriate bias correction which occurs due to endogeneity should take into consideration the presence of these two regimes\footnote{Notice that the underline regimes of our framework are motivated by the hypothesis of parameter instability in the predictive regression model rather than the presence of an independent threshold variable which induces the presence of the two regimes, as in gonzalo2012regime, gonzalo2017inferring.}.

Robust Inference with the IVX filtration

To deal with the aforementioned challenges for inference in predictive regression models, a robust procedure has been introduced by phillipsmagdal2009econometric and extended by kostakis2015Robust to detect the presence of predictability via the IVX-Wald test. The particular methodology, called IVX filtation, allows for the endogenous formulation of instruments based on the information contained in the regressors of the predictive regression. As a result, the degree of persistence of the instrumental variable has degree of persistence explicitly controlled so that the process is mildly integrated.

Specifically, the IVX filtration considers the first order difference of the corresponding autocorrelation regression such as it can be expressed as below

align[align omitted — 129 chars of source]

The particular first difference is not an innovation process unless the regressor belongs to the persistence class of integrated processes. However, it behaves asymptotically as an innovation after linear filtering by a matrix consisting of near-stationary roots\footnote{Note this assumption is a key idea in the development of the asymptotic theory in cointegrated systems for regressors with various types of nonstationarity (see, phillipsmagdal2009econometric).}.

Therefore, the procedure requires to choose an artificial coefficient matrix of the form

align[align omitted — 131 chars of source]

where $\mathbf{C}_z = diag \{ c_{z,1},...,c_{z,p} \}, c_z > 0$ for all $i \in \{1,...,p \}$. Then, the instrumental regressor matrix $\widetilde{ z }_t \in \mathbb{R}^{T \times p}$ can be constructed as below

align[align omitted — 205 chars of source]

To see this, notice that using expression (ref) we obtain the following decomposition

align[align omitted — 253 chars of source]

which can be written via the following expression

align[align omitted — 200 chars of source]

Therefore, the IVX filtration proposes to use the constructed $\widetilde{z }_t$ instruments for the regressors $x_t$ which are considered to behave asymptotically as mildly-integrated processes.

More explicitly, by replacing $x_t$ with the instrument $z_t$ which has a controllable degree of persistence, result to a robust inference procedure which accounts for the effects of nonstationarity. Then, the tuning parameters, which are the exponent rate $\delta_z$ and the diagonal matrix $C_z$ are selected to ensure that $z_t$ is mildly integrated; less persistence than a unit root or a regressor assumed to be generated via a local-unit-root process.

Thus, the IV based approach, implies that the IVX estimator is expressed as below

align[align omitted — 210 chars of source]

where $\bar{x}_{T-1} = \frac{1}{T} \sum_{t=1}^{T} x_{t-1}$ and $\bar{y}_{T} = \frac{1}{T} \sum_{t=1}^{T} y_t$, are the corresponding sample means.

As shown by Theorem A in the Appendix of KMS, the IVX estimator converges to a mixed Gaussian\footnote{The mixed Gaussianity property of the IVX estimator is also extensively examined in the papers of phillips2013predictive and phillips2016robust under various integration orders.} limiting distribution, which holds regardless of the degree of persistence of the regressors in the model. In turn, this property allows to construct a self-normalized Wald-type statistic which is shown to converge to a standard $\chi^2-$distribution.

The classical testing hypothesis implies that, $\mathbb{H}_0: \mathcal{R} \beta = 0$ vs $\mathbb{H}_1: \mathcal{R} \beta \neq 0$, where $\mathcal{R}$ the full rank $q \times p$ restriction matrix with rank $q$. Then, the IVX-Wald statistic can be used to test the null hypothesis of no predictability in predictive regression models.

The IVX-Wald statistic is expressed as below

align[align omitted — 122 chars of source]

where $\mathbf{Q}_{\mathcal{R}}$ is a consistent estimator of the asymptotic variance-covariance matrix of $\widetilde{\beta}^{ IVX}$ that accommodates both long-run endogeneity caused by the correlation between the error terms of the system, that is, $u_{t}$ and $v_{t}$, and the finite-sample distortion which results from removing the model intercept. The covariance matrix $\mathbf{Q}_{\mathcal{R}}$, is derived via the following fully modified (FM) estimation of the system covariance terms

align[align omitted — 580 chars of source]

where $\widetilde{z}_{T-1} = \frac{1}{T} \sum_{t=1}^{T} \tilde{z}_{t-1}$. Notice that for the univariate predictive regression then $\widehat{ \mathbf{\Omega} }_{FM} \equiv \sigma^2_{FM}$ and $\widehat{\mathbf{\Sigma} }_{uu} \equiv \widehat{ \sigma }^2_u$, where $\widehat{ \sigma }^2_u$ is a consistent estimator of $\sigma^2_u$. In that case, we can use the bias correction of KMS for the univariate model which is given by

align[align omitted — 301 chars of source]

A key aspect for the computation of the IVX-Wald statistic is the estimation procedure for the matrices $\widehat{ \mathbf{ \Omega} }_{uv}$ and $\widehat{ \mathbf{ \Omega} }_{uu}$, which represent the estimated long-run covariance between $u_{t}$ and $v_{t}$ and the long-run variance of $u_{t}$ respectively. More specifically, these covariance matrices can be constructed using nonparametric kernels with preselected bandwidth parameters, such as the Newey-West type estimators (see, KMS for detailed description and related studies such as newey1987simple and andrews1991heteroskedasticity).

Furthermore, it can be proved that under the null hypothesis which imposes a set of linear restrictions in the parameter vector $\beta$, we obtain that $\mathcal{W}^{IVX}_T \implies \chi^2(q)$ as $T \to \infty$ (see, Theorem 1 in KMS). This important limit theory result provides a unified framework for robust inference and testing regardless the persistence properties of regressors. A simple example is the application of the IVX-Wald test for inferring the individual statistical significance of predictors under abstract degree of persistence. For additional scenarios such as predictors of mixed integration order see phillips2013predictive and phillips2016robust. In these cases the powerful property of the IVX-Wald statistic still holds.

Comparing OLS against IVX based estimators

In this section, we present some preliminary comparisons between the OLS-Wald and the IVX-Wald statistics, under the null hypothesis of no structural break, to motivate further our research. These comparisons aim to shed some light on the ability of the OLS estimator to detect parameter instability under the assumption of nonstationary predictors. Furthermore, we examine the performance of the IVX estimator for robustifying inference in predictive regression models with persistent\footnote{Notice that the assumption of persistent or integrated regressors is commonly used in the time series econometrics literature due to the nature of economic and financial variables.} predictors under the presence of parameter instability which has been recently a stylized fact in various economic data.

We consider that the pair $\left\{ \left( y_t, x_t \right): 1 \leq t \leq T \right\}$ is generated by the following bivariate predictive regression system drawn from an i.i.d normally distributed error sequence

align[align omitted — 146 chars of source]

where $\left( u_t, v_t \right) \overset{ i.i.d }{ \sim } \mathcal{N}

pmatrix[pmatrix omitted — 91 chars of source]

$ with $c_i > 0$ or $c_i < 0$ where $y_t \in \mathbb{R}$ and $x_t \in \mathbb{R}$.

We are particularly interested to design Wald statistics robust to nonstationarity which allows to test jointly for the presence of predictability and parameter instability in the predictive regression model given by (ref) and (ref). To demonstrate that the estimator of the model parameter can affect the limiting distribution of the test and thus statistical inference we consider via a Monte Carlo experiment testing for a structural break. Thus, we can apply the sup-Wald statistics using each estimator separately by partitioning the sample, estimating the test statistic for each partition and then taking the supremum over the generated sequence of test statistics\footnote{Formal assumptions and asymptotic theory regarding the implementation of the sup-Wald test can be found in andrews1993tests. In this section, we demonstrate simple comparisons of the finite sample performance of the two Wald-type statistics based on the OLS and IVX estimators respectively.}.

Initially, we consider that the covariance matrix $\Sigma$ which captures the dependence between the predictive regression and the persistent regressor has a unit variance. Notice that in this paper we consider a single structural change as in the seminal paper of andrews1993tests. When the break-point location, denoted with $\pi$ is known, then to construct the Wald statistic we estimate model parameters $\beta_1 ( \pi )$ for $t = \left\{ 1,..., T \pi \right\}$ and $\beta_2 ( \pi )$ for $t = \left\{ T \pi + 1,..., T \right\}$ based on the observations of the corresponding sub-samples. When the break-point location is unknown then we consider the closed interval $\pi \in [\pi_1, \pi_2]$.

Under the null hypothesis of no structural break, $\mathbb{H}_0: \beta_1 = \beta_2$, which implies parameter constancy across the full sample. We report rejection frequencies from 5,000 Monte Carlo replications, using the predictive regression model with a fixed parameter $\beta = 0.25$. For the purpose of comparability we use the critical values that correspond to the NNB asymptotic critical values given by Table 1 of andrews1993tests. The empirical size results for the sup OLS-Wald statistic are presented on Table (ref) and (ref), while additional results with 10,000 replications can be found on Table (ref) and (ref).

As we can see from these rejection frequencies, under the assumption of persistent predictors (low values of $c$) when we use the critical values that correspond to the standard NBB limiting distribution we obtain size distortions which considerably deteriorate for larger correlation values between the error sequences $u_t$ and $v_t$. Thus, using the OLS estimator especially when modelling nonstationarity can produce biased inference. The econometric intuition is clear, constructing structural break statistics based on the OLS estimator when the predictor follows a local-to-unit root process causes not correctable size distortions of the tests due to the fact that the degree of persistence coefficient $c_i$ is unknown a prior and cannot be consistently estimated (see, Proposition (ref)). Furthermore, when we implement the sup IVX-Wald statistic as explained in the previous section then we will need to carefully examine the stochastic properties of the proposed tests.

Firstly, we briefly examine in this setting the asymptotic theory of the two estimators to understand which components of the limiting distribution result into nonstandard sampling results especially due to the introduction of parameter instability.

Notice that the appearance of nonstandard limiting distribution occurs in the case we estimate the sup OLS-Wald statistic, since when we estimate the model parameter $\beta$ by OLS its limiting distribution depends on the following two components

align[align omitted — 269 chars of source]

The first component converges to the integral of a squared demeaned OU process, while the second component weakly converges to a stochastic integral of that demeaned OU process. Therefore, under the assumption of nonstationary regressors and due to the appearance of the corresponding stochastic integral approximations, the second moment of the asymptotic distribution of the model estimator have a nonconstant variance and a nonstandard limiting distribution. For instance, georgiev2018testing derive the limiting distribution of the sup-Wald test statistic under the local-to-unit-root generating process, which results to a non-pivotal asymptotic theory. In order to deal with this problem, the authors propose to use the fixed regressor bootstrap which is asymptotically valid even under the nonstationarity assumption.

Secondly, in the case we estimate the model parameter using the IVX instrumentation, then the following matrix moments appear

align[align omitted — 233 chars of source]

The asymptotic theory of the above moment matrices is examined extensively by KMS. A summary of the main results is presented by Lemma (ref) in Appendix (ref) of this paper.

Both OLS and IVX estimators weakly converge to a mixed Gaussian distribution, thus to better understand the differences between the OLS and the IVX estimator\footnote{The asymptotic distribution of the IVX estimator in the univariate predictive regression model is proved to be mixed Gaussian. A proof of this asymptotic result can be found in phillips2016robust.} and how these manifest in the asymptotic theory of the tests, we examine the stochastic quantity

align[align omitted — 162 chars of source]

The $\mathcal{P}_c \left( \pi \right)$ quantity which appears in the limiting distribution of the Wald statistics when an IVX estimator is employed, is a random variable. Thus, considering the analytical form of this random quantity will require to derive an analytical approximation with stochastic terms of higher order, which is a challenging task.

Similarly, the corresponding expression for the OLS estimator is is a function of the stochastic quantity

align[align omitted — 159 chars of source]

To obtain some insights regarding the asymptotic behaviour of the stochastic quantities $\mathcal{P}^{IVX}_c \left( \pi \right)$ and $\mathcal{P}^{OLS}_c \left( \pi \right)$ we use integral approximation simulation techniques with 10,000 replications for a fixed $\pi \in [ 0.15, 0.85 ]$. Specifically, to approximate the underline OU process which drives these stochastic quantities we use a $1000-$step random walk.

Moreover, to make these heuristic arguments more apparent, we plot the mean and $95\%$ confidence intervals of the particular empirical fluctuation processes for values of the persistence parameter $c \in \left\{ 1, 5, 100 \right\}$. As we can see from Figure 1, for low persistence levels the asymptotic variance of these processes exhibit a linear growth rate for both the OLS and the IVX estimators since $\mathcal{P}_c \left( \pi \right)$ has approximate linear increments with $\pi$. In other words, since the predictive regression model does not assume stationarity for the persistence properties of regressors then by allowing for breaks if we use the OLS based estimator then the accumulation of information as seen by the the second moments has a faster convergence rate especially in the case of high persistence. On the other hand, when we use the IVX estimator this accumulation of information is smoother for different values of $\pi$ due to the properties of the particular estimator in filtering out abstract degrees of persistence, even under time-varying settings.

FIGURE 1. \color{black}

\color{black}

Joint Predictability and Structural Break Testing

We focus on the linear predictive regression model and in particular the univariate case. Our proposed testing framework and related asymptotic theory can be extended to the multivariate predictive regression with suitable notation modification. Testing for a single structural break at an unknown break point location within the full sample requires that the multiple predictive regression is expressed as below

align[align omitted — 210 chars of source]

where $k = \floor{ T \pi}$ for some $\pi \in (0,1)$, with $\floor{.}$ denoting the integer part operator.

To derive the asymptotic theory of the tests we employ the local-unit-root specification proposed by phillips1987time. Specifically, since the set of regressors are assumed to be generated via the LUR process $x_t = \left( I_p - \frac{C}{T} \right)$ $x_{t-1} + v_t$, we consider the following $p-$dimensional Gaussian process

align[align omitted — 86 chars of source]

which satisfies the Black-Scholes differential equation $d K_c(r) \equiv c K_c(r) + d B_v(r)$, with $K_c(r) =0$, implying also that $K_c(r) \equiv \sigma_v J_c(r)$, where $\displaystyle J_c(r) = \int_0^r e^{(r-s)C} d W_v(s)$ and $K_c(r)$ the Ornstein-Uhlenbeck, (OU) process,\footnote{The OU is a stationary Gaussian process with an autocorrelation function that decays exponentially over time. Moreover, the continuous time OU diffusion process has a unique solution.} which encompasses the unit root case such that $J_c(r) \equiv B_v(r)$, for $C = 0$.

Classical Least Squares Estimation

In this Section, we derive the asymptotic theory result that corresponds to a standard sup-Wald test, based on the OLS estimator, when testing for a single structural break in predictive regressions with multiple predictors.

We denote with $\mathbf{X}_1 := \big[ \mathbf{1} \{ t \leq k \} \ \ x_{t} \mathbf{1} \{ t \leq k \} \big]$, $\mathbf{X}_2 := \big[ \mathbf{1} \{ t > k \} \ \ x_{t} \mathbf{1} \{ t > k \} \big]$ and define $\mathbf{X} = [ \mathbf{X}_1 \ \mathbf{X}_2 ] \in \mathbb{R}^{T \times 2 (p+1)}$ the corresponding partitioned matrix and $\mathbf{\mathcal{R}} = \left[ \mathbf{I}_{p+1} \ - \mathbf{I}_{p+1} \right]$ with $\mathbf{I}_{p+1}$ denoting an identity matrix. Moreover, we denote with $\beta := ( \beta_1 , \beta_2 )^{\prime}$, the parameter vector, then the predictive regression is $y = X \beta + u$. Thus, the OLS Wald statistic for testing the null hypothesis of no structural break, that is, $\mathbb{H}_0: \theta_1 = \theta_2$ against $\mathbb{H}_1: \theta_1 \neq \theta_2$, where $\theta_j = (\alpha_j, \beta_j )$ for $j=1,2$, is given by the following expression

align[align omitted — 323 chars of source]

Statistical inference under the null hypothesis of no structural break in the predictive regression are conducted using the supremum functional. The sup OLS-Wald statistic is

align[align omitted — 177 chars of source]

where $ 0 < \pi_1 < 1$ and $ 0 < \pi_2 < 1$ with $\pi_2 = 1 - \pi_1$.

Under the null hypothesis the breakpoint $k$ is unidentified and so the supremum functional selects the maximum Wald statistic corresponding to a sequence of Wald statistics evaluated at values within the interval $[\pi_1, \pi_2]$. Specifically, in order to construct the corresponding supremum Wald tests we split the full sample $\left\{ y_j, x_{j-1} \right\}_{j=1}^T$ into two sub-samples; the first sub-sample corresponds to the time period before time $t$, $\left\{ y_j, x_{j-1} \right\}_{j=1}^t$ and the second sub-sample corresponds to the time period after time $t$, $\left\{ y_j, x_{j-1} \right\}_{j=t+1}^T$. Due to the fact that we operate within the Skorokhod topology the sample moments that correspond to these sub-samples weakly converge to the corresponding asymptotic result which are based on the $\mathcal{D}[0,1]$ topology. For instance, when deriving the sample moments we can use the notation $t \in \left[ t_1, t_2 \right]$, where $t_1$ and $t_2$ are the lower and upper bounds for the possible break-point location.

theoremIf Assumptions (ref) and (ref) hold and $\pi$ denotes the unknown break-point, then the sup OLS-Wald statistic under the null hypothesis $\mathbb{H}_0: \theta_1 = \theta_2$ against $\mathbb{H}_1: \theta_1 \neq \theta_2$, where $\theta_j = (\alpha_j, \beta_j )^{\prime}$ with $j=1,2$ based on the predictive regression model (ref)-(ref) with $\gamma_x = 1$\footnote{Note that this parameter restriction for the exponent rate implies that $x_t$ follows a LUR process.} weakly converges to the following limiting distribution \begin{align} \widetilde{\mathcal{W}}^{OLS}( \pi ) \equiv \underset{ \pi \in [ \pi_1 , \pi_2 ] }{ sup } \ \mathcal{W}_T^{OLS}( \pi ) \Rightarrow \underset{ \pi \in [ \pi_1 , \pi_2 ] }{ sup } \ \bigg\{ \widetilde{ \mathbf{N} } ^{\prime}_c( \pi ) \widetilde{ \mathbf{M} }_c( \pi )^{-1} \widetilde{ \mathbf{N} }_c( \pi ) \bigg\} \end{align} where \begin{align} \widetilde{ \mathbf{M} }_c( \pi )^{-1} = \widetilde{ \mathbf{G} }_c(\pi) - \widetilde{\mathbf{G} }_c(\pi) \widetilde{\mathbf{G} }_c(1)^{-1} \widetilde{\mathbf{G} }_c(\pi) \end{align} and \begin{align} \widetilde{ \mathbf{N} }_c( \pi ) = \left\{ \widetilde{\mathbf{G} }_c(\pi)^{-1} \widetilde{H} _c(\pi) - \bigg[ \widetilde{\mathbf{G}}_c(1) - \widetilde{\mathbf{G}}_c(\pi) \bigg]^{-1} \bigg[ \widetilde{H}_c(1) -\widetilde{H}_c(\pi) \bigg] \right\}^{\prime} \end{align} such that \begin{align} \widetilde{\mathbf{G} }_c(\pi) := \int_0^{\pi} \widetilde{ K }_c( r ) \widetilde{ K }^{\prime}_c( r ) dr \ \ and \ \ \widetilde{H}_c( \pi ) := \int_0^{\pi} \widetilde{ K }_c( r ) d B_u(r) \end{align}

with $\widetilde{\mathbf{G} }_c(\pi) \in \mathbb{R}^{ ( p +1 ) \times ( p +1 )}$ a positive-definite stochastic matrix and $\widetilde{H}_c(\pi) \in \mathbb{R}^{ ( p +1 ) \times 1}$.

Theorem (ref) provides a novel result in the literature and our first theoretical contribution, demonstrating that the sup OLS-Wald statistic as a structural break test on an unknown location based on the predictive regression model with persistent predictors, does not converge to the NBB result as proposed by andrews1993tests in linear regressions. Thus, we show that the asymptotic theory of the test depends on the nuisance parameter of persistence $c_i$. However, when the break-point is known a prior, such as $\pi \equiv \pi_0$, it can be easily proved that the limit theory of the OLS-Wald test converges to a NBB even in the case of persistent predictors. The proof of Theorem (ref) is shown in Appendix (ref).

IVX based Estimation

In this section, we examine the IVX-Wald based tests for jointly testing for the presence of predictability and parameter instability based on the predictive regression model. The null hypothesis remains the same as in the IVX-Wald test, but the alternative hypothesis is constructed so that it captures both effects. We examine separately the case of stable model intercept and the case of unstable model intercept. The model estimates of the two sub-samples are used to construct the proposed test statistics. We denote with $\widetilde{\beta}_1^{IVX} \left( t \right)$ and $\widetilde{\beta}_2^{IVX} \left( t \right)$ the IVX estimates of the two sub-samples; and with $\widetilde{Q}_1 \left( t \right)$ and $\widetilde{Q}_2 \left( t \right)$ the corresponding covariance matrices. The sample estimates are given by

align[align omitted — 409 chars of source]

Joint Wald Tests under stable model intercept

We begin our asymptotic theory analysis by considering the case in which the model intercept $\alpha$ is assumed to be stable, that is, under the null hypothesis there is no structural break in the model intercept. Then, the testing hypothesis of interest is given by

align[align omitted — 83 chars of source]

For evaluating our hypothesis we use the IVX-Wald test which has the following form\footnote{Note that the covariance matrices are computed based on the long-covariance matrices and corresponding FM corrections given by KMS, see also kasparis2015nonparametric. These definitions are employed since we do not rule out the strong assumption of a weakly covariance dependence in the model.}

align[align omitted — 294 chars of source]

We denote with $\widetilde{\mathcal{W}}^{IVX}_{\beta}(t) \equiv \underset{ t \in [ t_1, t_2 ] }{ \text{sup}} \mathcal{W}_{\beta}^{IVX}(t)$, the corresponding sup IVX-Wald statistic, since we consider the case of the unknown break-point. The two IVX estimates of $\beta$ and their corresponding asymptotic variances are computed, using the data from each sub-sample separately. Thus, the $\mathcal{W}_{\beta}^{IVX}(t)$ statistic can be thought as a Chow-type statistic for detecting break at time $t$, such that $t_1 \leq t \leq t_2$. Then, due to the unknown nature of the break-point the sup Wald test is simply a sequence of Chow-type statistics within the same probability space. Theorem (ref) presents the limiting distribution of $\widetilde{\mathcal{W}}^{IVX}_{\beta}(t)$, under the null hypothesis of no parameter instability in the predictive regression model.

theoremIf Assumptions (ref) and (ref) hold and $\pi$ denotes the unknown break-point, then the sup IVX-Wald statistic under the null hypothesis (with $\alpha$ known to be stable a priori) $\mathbb{H}_0 : \alpha_1 = \alpha_2$ and $\beta_1 = \beta_2$, based on the predictive regression model (ref)-(ref) and no restriction on the exponent rate $\gamma_x$, has the following asymptotic behaviour \begin{align} \widetilde{\mathcal{W}}^{IVX}_{\beta}(t) \Rightarrow \underset{ \pi \in [ \pi_1, \pi_2 ]}{sup} \bigg\{ \mathbf{N}( \pi )^{\prime} \mathbf{M}( \pi )^{-1} \mathbf{N}( \pi ) \bigg\} \end{align} with $t \in [t_1, t_2]$\footnote{Note that, since we operate within the Skorokhod topology $\mathcal{D} \left( 0, 1 \right)$ we can apply standard weakly convergence arguments.} and $\pi \in [\pi_1, \pi_2]$, where \begin{align} \mathbf{N}( \pi ) &= B_p( \pi ) - \mathbf{R}( \pi ) B_p( 1 ) \\ \mathbf{M}( \pi ) &= \pi \big( \mathbf{I}_p - \mathbf{R}( \pi ) \big) \big( \mathbf{I}_p - \mathbf{R}( \pi ) \big)^{\prime} + (1 - \pi ) \mathbf{R}( \pi ) \mathbf{R}( \pi )^{\prime} \end{align} such that \begin{equation} \mathbf{R} ( \pi ) = \begin{cases} \left( \pi \mathbf{\Omega}_{xx} + \int_0^{\pi} \mathbf{B} dB^{\prime} \right) \left( \mathbf{\Omega}_{xx} + \int_0^{1} \mathbf{B} dB^{\prime} \right)^{-1} & ,if \ \gamma_x > 1 \\ \\ \left( \pi \mathbf{\Omega}_{xx} + \int_0^{\pi } \mathbf{J}_C dJ_C^{\prime} \right) \left( \mathbf{\Omega}_{xx} + \int_0^{1} \mathbf{J}_C dJ_C^{\prime} \right)^{-1} & ,\text{if} \ \gamma_x = 1 \\ \\ \pi \mathbf{I}_p & , \text{if} \ \gamma_x < 1 \end{cases} \end{equation} where $B(.)$ is a $p-$dimensional standard Brownian motion, $J_C (\pi) = \int_0^{\pi} e^{C (\pi - s)} dB(r)$ is an \textit{Ornstein-Uhkenbeck} (OU) process and we denote with $\underline{J}_C (\pi) = J_C (\pi) - \int_0^1 J_C(s) ds$ and $\underline{B} (\pi ) = B(\pi) - \int_0^1 B(s) ds$ the demeaned processes of $J(\pi)$ and $B(\pi)$ respectively.

Theorem (ref) demonstrates that the supremum functional of the IVX-Wald test weakly convergence to a stochastic quadratic functional and does not have the same asymptotic behaviour as the corresponding IVX-Wald statistic which is proved to follow a $\chi^2_p$ distribution (see, KMS, phillips2013predictive and phillips2016robust).

An important implication of Theorem (ref), which consists our second contribution to the limit theory of structural break tests in predictive regression models with nonstationary predictors, is that we show that the dependence of the limiting distribution to the unknown parameter of persistence takes different forms by restricting further the parameter space of the exponent rate. Specifically, Theorem (ref) demonstrates that when testing for a structural break in predictive regression models, under the assumption that the predictors are integrated or nearly integrated, then the limiting distribution of the test, weakly converge to a nonstandard process, which verifies the corresponding result already mentioned by hansen2000testing). On the other hand, when the predictors are stationary or mildly stationary, then the test statistic weakly converges to the familiar squared tied-down Bessel process as proved by andrews1993tests in the case of linear regression models. The latter implies that when we consider separately the special case for which $\gamma_x < 1$, which covers both the cases of mildly integrated regressors, i.e., $\gamma_x \in (0,1)$, and stationary regressors, i.e., $\gamma_x = 0$, then it can be proved that the limiting distribution of the sup IVX-Wald test converges to the standard NBB result\footnote{Note that in the case of a known break-point we can easily prove a convergence to a $\chi^2_p$ distribution which is free of nuisance parameters and so conventional inference methods apply}. Corollary (ref) below is a direct implication of Theorem (ref) and summarizes this finding.

corollaryUnder the assumptions and definitions given by Theorem (ref), when $\gamma_x < 1$ then the following asymptotic distribution holds \begin{align} \widetilde{\mathcal{W}}^{IVX}_{\beta}(t) \Rightarrow \underset{ \pi \in [ \pi_1, \pi_2 ]}{sup} \frac{ \mathcal{BB}_p ( \pi )^{\prime} \mathcal{BB}_p ( \pi ) }{ \pi (1 - \pi) }, \end{align} where $\mathcal{BB}_p ( . )$ is a $p-$dimensional standard Brownian bridge.
remarkThe weakly convergence of the sup IVX-Wald test into a normalized squared Brownian bridge, when the exponent rate is restricted such that $\gamma_x < 1$, implies standard statistical inference due to the known distribution. One rejects $\mathbb{H}_0$ for large values of the sup IVX-Wald test based on a significance level $\alpha$ such that $0 \leq \alpha \leq 1$ and thus the limit distribution can be used to derive associated critical values, denoted with $c_{\alpha}$ such that $\mathbb{P} \big( \widetilde{\mathcal{W}}_{\beta}^{IVX} ( \pi ) > c_{\alpha} \big) > 0$ with $\underset{ T \to \infty }{ \text{lim} } \mathbb{P} \big( \widetilde{\mathcal{W}}_{\beta}^{IVX} ( \pi ) > c_{\alpha} \big) = 1$.

Our findings presented by Theorem (ref) verify the conjuncture of hansen2000testing who argues that when testing for a structural break based on a sup-functional induces an asymptotic non-pivotal distribution under the assumption of nonstationarity. However, this seminal study has not examined in details certain forms of nonstationarity which can occur and how these are manifested in the limit theory of the tests. In particular, within our framework we demonstrate using local-to-unit root asymptotic arguments that the nonstandard limiting distribution occurs when $x_t$ is properly modelled as a nonstationary stochastic process, and specifically the NBB result no longer holds when $x_t$ is either a nearly integrated or an integrated process.

The proofs of both Theorem (ref) and Corollary (ref) can be found in Appendix (ref). Next, we focus on designing test statistics for jointly testing against predictability and structural break in predictive regression models. We define a joint Wald test based on the IVX estimator under the assumption of an unknown break-point as below

align[align omitted — 240 chars of source]

where the supremum functional applies only on the second component of the test above. Then, the Corollary (ref) below gives the related limit theory result.

corollary(i) If the conditions of Theorem (ref) hold, then under the null hypothesis, $\mathbb{H}_0: \alpha_1 = \alpha_2$ and $\beta_1 = \beta_2 = 0$ and no restriction on the exponent rate $\gamma_x$, the large sample theory of the test statistic specified by (ref) has the following form \begin{align} \widetilde{ \mathcal{W} }_{\beta}^{joint} \Rightarrow B(1)^{\prime} B(1) + \underset{ \pi \in [ \pi_1, \pi_2 ]}{ sup } \bigg\{ \mathbf{N}( \pi )^{\prime} \mathbf{M}( \pi )^{-1} \mathbf{N}( \pi ) \bigg\}, \end{align} where $B(.)$, $\mathbf{N}( \pi)$ and $\mathbf{N}( \pi )$ are defined in Theorem (ref). \\ (ii) As a special case, when $\gamma_x < 1$, it follows that \begin{align} \widetilde{\mathcal{W}}_{\beta}^{joint} \Rightarrow \chi^2_p + \underset{ \pi \in [ \pi_1, \pi_2 ]}{ sup } \frac{ \mathcal{BB}_p ( \pi )^{\prime} \mathcal{BB}_p ( \pi ) }{ \pi (1 - \pi ) }, \end{align} where $\mathcal{BB}_p(.)$ is a $p-$dimensional standard Brownian bridge and $\chi^2_p$ denotes the $\chi^2$ random variable with $p$ degrees of freedom. Furthermore, the two stochastic quantities of the limiting distribution are assumed to be independent.
remarkNotice that Corollary (ref) demonstrates that when $\gamma_x < 1$, the large sample theory of the test statistic $\mathcal{W}_{ \beta }$ is pivotal. Therefore, in this special case by having a limiting distribution being free of nuisance parameters, asymptotic critical values for testing the null hypothesis can be easily obtained. The particular test statistic provides a methodology for testing for both predictability and structural break which is robust to the persistence properties of regressors after replacing the OLS with an IVX estimator.

The main arguments we use for the proof of Corollary (ref) is to consider the joint testing hypothesis as a composite hypothesis based on the mutually exclusive parameter space, and thus we construct the limit theory of this test based on the joint formulation of the two separate testing hypotheses. By employing the asymptotic matrix moments based on the IVX instrumentation we prove that the limiting distribution of the joint test can be decomposed into two components. These two components are considered to be independent random variables and therefore practically we could use the critical values of the limiting distributions that corresponds to each of these stochastic quantities.

Joint Wald Tests under unstable model intercept

Notice that when a model intercept is included then under the null hypothesis, when $\alpha_1 = \alpha_2$, for example if we are testing whether $H_0: \beta_1 = \beta_2$, then the particular hypothesis would be equivalent to testing the null hypothesis of no structural break. However, in the case of an unstable model intercept, that is, $\alpha_1 \neq \alpha_2$, then testing the null hypothesis $H_0: \beta_1 = \beta_2$ can result to a different asymptotic theory due to the break in the model intercept. In this section, we consider the case that a potential structural break occurs in both the intercept and the slope coefficient $\beta$ of the predictive regression model. The two proposed tests aim to test: (i) the null hypothesis of no break on model intercepts while the slope coefficients remain at a fixed level; (ii) the null hypothesis of no break on model intercepts while the slope coefficients are both equal to zero.

The proposed test statistic is expressed as below

align[align omitted — 262 chars of source]

where the covariance estimators are computed as below

align[align omitted — 379 chars of source]

Then, the joint Wald test is expressed using the following form

align[align omitted — 201 chars of source]

To distinguish between the standard Wald statistics in the case of known break-point and the Wald type statistics that correspond to the unknown break-point we denote with $\widetilde{\mathcal{W}}^{IVX}_{\alpha}(t) = \underset{ t \in [ t_1, t_2 ]}{\text{sup}} \mathcal{W}^{IVX}_{\alpha}(t)$ and $\widetilde{\mathcal{W}}^{joint}_{\alpha \beta }(t) = \underset{ t \in [ t_1, t_2 ]}{\text{sup}} \mathcal{W}^{joint}_{\alpha \beta}(t)$. Both of these test statistics require to split the sample into subsamples of window size $[ t_1, t_2 ] \subset (1,T)$ and estimate the joint Wald test based on the observations from each subsample. For instance, the last two components of the joint test requires to estimate the Wald IVX for the full sample plus the supremum statistics for the model intercept and the model slope separately. Notice that the model intercept has different asymptotic properties, in particular, convergence rate to the true population parameter when the IVX estimator is used for the slope parameter. Therefore, we use the covariance estimators given by (ref) and (ref) which is constructed based on the residuals from the fitted predictive regression of each subsample while we add a bias correction based on the covariance estimator of the slope coefficients since the estimator of the intercept is conditional on the IVX estimator in the case we switch from the OLS to the IVX estimation procedure.

propositionConsider the predictive regression model given by expressions (ref)-(ref). If Assumption (ref)-(ref) hold and $\alpha$ is known to be unstable a priori, then under the null hypothesis $\mathbb{H}_0: \alpha_1 = \alpha_2 \ \ \text{and} \ \ \beta_1 = \beta_2 = \beta$, as T $\to \infty$ the following limit result holds \begin{align} \widetilde{\mathcal{W}}^{IVX}_{\alpha}(t) \Rightarrow \underset{ \pi \in [ \pi_1, \pi_2 ]}{ sup } \frac{ \mathcal{BB}_1 ( \pi )^{\prime} \mathcal{BB}_1 ( \pi ) }{ \pi (1 - \pi) } \end{align} where $\mathcal{BB}_1 ( . )$ is a one-dimensional standard Brownian bridge.
remarkNotice that Proposition (ref) provides an asymptotic result for a composite hypothesis since we consider jointly testing for a structural break in the model intercept and the slope coefficients while we test, that under the null hypothesis the slope coefficient has a fixed parameter value $\beta$. Additionally, we can investigate the limiting distribution of the joint Wald test when we assume that under the null hypothesis there is no predictability.
propositionConsider the predictive regression model given by expressions (ref)-(ref). If Assumption (ref)-(ref) hold and $\alpha$ is known to be unstable a priori, then under the null hypothesis $\mathbb{H}_0: \alpha_1 = \alpha_2$ and $\beta_1 = \beta_2 = 0$, as $T \to \infty$ the following limit results hold: \ (i) \begin{align} \widetilde{\mathcal{W}}^{joint}_{\alpha \beta }(t) \Rightarrow B(1)^{\prime} B(1) + \underset{ \pi \in [ \pi_1, \pi_2 ]}{ sup } \bigg\{ \widetilde{ \mathbf{N} }( \pi )^{\prime} \widetilde{ \mathbf{M} } ( \pi )^{-1} \widetilde{ \mathbf{N} }( \pi ) \bigg\} \end{align} where $\widetilde{ \mathbf{N} }( \pi ) = \big( \mathcal{BB}_1( \pi ), \mathbf{N}( \pi ) \big)^{\prime}$ and $\widetilde{ \mathbf{M} }( \pi ) = \begin{pmatrix} \pi(1- \pi) & 0 \\ 0 & \mathbf{M} (\pi) \end{pmatrix}$. The terms, $\mathbf{N}(\pi)$ and $\mathbf{M} (\pi)$ are defined in Theorem (ref). \\ (ii) As a special case, when $\gamma_x \in (0,1)$, it holds that \begin{align} \widetilde{\mathcal{W}}^{joint}_{\alpha \beta }(t) \Rightarrow \chi^2_p + \underset{ \pi \in [\pi_1, \pi_2 ]}{ sup } \frac{ \mathcal{BB}_{p+1} ( \pi )^{\prime} \mathcal{BB}_{p+1} ( \pi ) }{ \pi (1 - \pi) } \end{align} where $\mathcal{BB}_{p+1} ( . )$ is a $(p+1)-$dimensional standard Brownian bridge, and $\chi^2_p$ is a random variable following a $\chi^2$ distribution with $p$ degrees of freedom.
remarkNotice that, Proposition (ref) shows that the limiting distribution of the joint test for both predictability and structural break, when we consider simultaneously testing whether there is a structural break to the model intercept and no predictability using the set of regressors of the model, has an asymptotic distribution which takes a different form when we consider different values of the parameter space of the exponent rate.

As we have seen by our extensive asymptotic theory analysis provided in this Section, the asymptotic distribution of the Joint IVX-Wald tests can be affected by various scenarios, such as the inclusion of model intercept in the predictive regression as well as the parameter space of the exponent rate.

Monte-Carlo Simulation Study

In this section, we present a Monte Carlo simulation study in order to examine the finite size properties of the proposed Wald-type statistics in terms of their empirical size and power performance, under the null hypothesis of no joint parameter instability and predictability. In practise, the degree of persistence in the time series of the regressors is unknown. That is, both the coefficient of persistence $c_i$ as well as the exponent rate $\gamma_x$ are both unknown parameters to the researcher. Moreover, we have proved that the limiting distribution of the Wald-type statistics for detecting structural break in predictive regression models depend on these unknown properties of the regressors for certain parameter value restrictions on the coefficient of persistence. In particular when the exponent rate of persistence has an exponential rate , then for both cases i.e., testing for structural break or jointly testing for predictability and structural break we have a nonstandard limiting distribution which depends on the corresponding stochastic integrals which are functions of the unknown break-point. Now, in the case we known a prior that the exponent rate equals to one, which means we are within the realm of near integrated or LUR regressors then the tests converge to a standard NBB distribution which is easy to tabulate critical values. Nevertheless, in any case in empirical applications critical values can be obtained by either simulations or more advanced bootstrap methodologies.

The MC simulation study aims to shed light on the main theoretical results of the paper. Therefore, to demonstrate the above theoretical result, under the null hypothesis of no structural break, we can generate a DGP with no breaks in the coefficients of the predictive regression. Then, constructing the sup-Wald test and using the Andrews' critical values we can observe whether size distortions indeed occur in this scenario\footnote{Recall that under the assumption of persistent predictors, we prove via Proposition (ref) of the paper that when testing for parameter instability in the predictive regression the limiting distribution of the sup-Wald statistic no longer follows a normalized squared brownian bridge, (NBB) for an unknown structural break. In particular, when contacting inference or assessing the statistical validity of the sup-Wald statistic via a MC experiment, using the corresponding critical values of the sup-Wald test proposed by andrews1993tests we can observe that leads to size distortions due to the non-NBB limiting distribution.}. Secondly, via Proposition (ref) of the paper we propose an alternative approach to overcome this problem. In particular, using an IV based sup-Wald test, which is constructed using the IVX instrumentation, we prove that the limiting distribution of the statistic indeed weakly converges to a NBB, which allow us to use the Andrews' critical values, avoiding this way to simulate critical values which can be computational complex. Below, we present the DGP, the test statistics as well as the size and power comparisons for the proposed tests of our econometric framework.

Experiment Design

We use the following data generating process (DGP) where $y_t$ is a scalar and $x_t$ is a vector of LUR predictors (with the property of being highly persistent).

align[align omitted — 295 chars of source]

with $t \in \{1,...,T \}$ and $T = \{ 100, 250, 500, 1000 \}$ for $B = 5,000$ replications.

Furthermore, we consider the effect of different localizing coefficients of persistence across the predictors. We use $c_i \in \{ 1, 5, 10, 20 \}$ for $i = 1,2$, which cover various cases of LUR regressors, with smaller values implying that we impose the assumption of higher persistence and lower values implying that existence of mild persistence in the predictor.

The covariance matrix of the innovations $\underline{e}_t = \left( u_t, \underline{v}_t \right)^{\prime} \sim \mathcal{N} \left( \underline{0}_{3 \times 1}, \Sigma_{ee} \right)$ we assume that is parametrised with the following structure

align[align omitted — 228 chars of source]

Clearly, we can see that the covariance matrix given by expression (ref) allows to consider various scenarios regarding the contemporaneous correlation of the regressors and the dependent variables which has related economic interpretation. We use the following predetermined covariance matrices

align*[align* omitted — 240 chars of source]

Under both the null and the alternative hypothesis, the predictive regression is generated using the following econometric specification (using (ref) and (ref))

align[align omitted — 218 chars of source]

where $k = \floor{ T \pi}$ for some $\pi \in (0,1)$.

Additionally under the null hypothesis, of no parameter instability, we simulate the DGP using the same vector of regression parameters before and after the breakpoint, such as $\underline{\beta}_1 = \underline{\beta}_2$. Furthermore, under the alternative hypothesis of parameter instability, we can generate a sequence of local alternatives by using a different vector of regression parameters before and after the breakpoint. Note that, since our framework consider a single unknown structural break, the DGP is re-constructed for each window $[\pi_1, \pi_2]$ in order to apply the supremum functional.

Test statistics

We assess the statistical validity in finite and large samples of the Wald based statistics within our proposed framework, that is, the sup Wald-OLS test and the sup-Wald IVX test. In particular, we are interested to verify any size distortions under the null hypothesis of no structural change, when using the sup Wald-OLS test and we expect to observe improvements in the empirical size of the sup Wald-IVX test. Moreover, we examine the size and power performance of the two statistics across different degree of persistence as well as the rate at which we allow the IVX instrumentation procedure to create a more mildly integrated regressor (as given by the parameters $c_z$ and $\delta \equiv \delta_0$).

For the large sample properties of the test statistics a standard convergence result apply, that is, $\mathcal{W}_T ( \pi ; \delta ) \to \mathcal{W} ( \pi ; \delta )$ as $T \to \infty$. Furthermore, the testing hypothesis of interest is a two-sided type hypothesis which is expressed as below

align[align omitted — 153 chars of source]

The test statistics are computed via expressions (ref) and (ref), while standard regularity and invariance principles holds (e.g, uniform convergence for the compact space $\underline{\theta} \in \Theta$ such that $( \underline{\beta}_1, \underline{\beta}_2)^{\prime} \subset \underline{\theta}$ and $\Theta \in \mathbb{R}^q$).

Wald-OLS statistic

align[align omitted — 356 chars of source]

Wald-IVX statistic

align[align omitted — 344 chars of source]

where $\mathcal{Q}_{ \mathcal{R} }$ is defined as below

align[align omitted — 292 chars of source]

In the case of a known break-point we can use the critical values proposed by andrews1993tests with an appopriate significance size such as $\alpha = 5\%$ for both test statistics. Since we consider an unknown break point, then the sup functional gives the statistic with the maximum value after estimating a sequence of statistics over the interval $\pi \in [ \pi_1, 1 - \pi_2]$, and thus critical values need to be used due to the fact that the limiting distribution in this case is not a standard $\chi^2-$distribution. Furthermore, the data that we generate from the DGP given by (ref) and (ref) only differ over the values of $c_1$ and $c_2$ and are the same for different values of the sample size $T$. Thus, as the sample size increases the degree of persistence of the endogenous regressors remain the same one (for both predictors). Moreover, the increase in the sample size aims to reflect the properties of finite versus large sample asymptotics for the Wald type statistics in testing for a single unknown break-point in predictive regression models with persistent regressors. Furthermore, in order to control for the existence of empirical size distortions when testing for structural break via the Wald-IVX statistic, we use the bias corrected IVX estimators as proposed by phillipsmagdal2009econometric. The particular bias correction, implies that we account for an overestimation effect when projecting the instrumental variable towards the direction of the dependent variable (due to endogeneity).

The bias corrected IVX estimator has the following form

align[align omitted — 167 chars of source]

where the estimator of $\Delta$ is given by

align[align omitted — 145 chars of source]

The above non-parametric estimator is known as Newey-West type estimator (see, newey1987simple) and allows to estimate the bias correction without imposing additional parametric assumptions. Optimal bandwidth choices can be of the form $m = \eta T^{1/5}$, where $\eta$, is a positive constant. The bias correction is applied to both $\tilde{ \underline{\beta}_1 }^{\text{IVX}}$ and $\tilde{ \underline{\beta}_2 }^{\text{IVX}}$.

Size Comparison

To conduct a size comparison of the test statistics, we examine the empirical rejection rates of the sup-Wald OLS and sup-Wald IVX tests for detecting single structural change with persistent predictors, under the null hypothesis of no structural change, that is, $\mathbb{H}_0: \underline{\beta}_1 = \underline{\beta}_2$ using a significance level $\alpha = 5\%$ and compare using different critical value approximations. In particular, we repeat the empirical-size experiment using the asymptotic critical values from (i) Table 4 of gonzalo2012regime, (ii) Table 1 of andrews1993tests), (iii) the bootstrap generated critical values within the MC step. More specifically, to do this, we generate 5,000 datasets from DGP (ref) and (ref) for various values of $\underline{\beta}^0$ and compute the frequency of rejecting the null hypothesis.

align*[align* omitted — 239 chars of source]

Table (ref) (for sup Wald-OLS) and Table (ref) and (ref) (for sup Wald-IVX), present the probabilities of rejection of the two Wald-type statistics at the $5\%$ nominal rate, under the null hypothesis (for Design 1). In particular, we consider different values for the exponent rate of the IVX parameter, such as $\delta \in \{ 0.75, 0.95 \}$ in order to investigate the varying effect of the degree of persistent of the instrumental variable for detecting structural change in predictive regressions with persistent regressors, as well as different localizing coefficient of persistence, $c_i \in \{ 1, 5, 10, 20 \}$. Comparing the empirical size results for Table (ref) versus Table (ref) and (ref), we can see that the sup Wald IVX produces values for the empirical size quite close to the nominal size for critical value $c_{\alpha} = 13.42$, $\delta = 0.95$ and $ -0.5 \leq \rho \leq 0.5$. The particular critical value is a closer representation to the $\alpha-$quantile from the corresponding limiting distribution.

In Table (ref) and Table (ref) we present the empirical size under the null hypothesis for the model with a single predictor and no-intercept. We can observe that size distortions appear for larger values of correlation between the $u_t$ and $v_t$ and this is more severe for high persistence regressors (i.e., low values of the coefficient of persistence). Moreover, we also consider the case of explosive regressors, that is, $c_i < 0$. Clearly, the empirical size in this scenario has higher values since we use the critical value of Andrews (8.85) which is not the corresponding critical value of the asymptotic distribution of the test statistic for detecting structural break in the predictive regression model under the assumption of explosive regressors. Notice also that in the case we have $p > 1$ and we use the sup OLS-Wald statistic for testing for joint predictability and structural break then the size distortions are much higher than in the case of $p =1$ and this is due to the nonstandard limiting distribution of the test statistic under the assumption of nonstationary regressors.

All main conclusions are in line with similar findings in the literature of predictive regression models, such as larger size distortions appear as the correlation between u(t) and v(t) increases, and this effect is more apparent in the case of persistent predictors (i.e., lower values of the persistence coefficient c). Moreover, we see that the Sup OLS-Wald statistic with the standard asymptotics (Andrews) is clearly immune to persistence when there is no correlation between u(t) and v(t). Furthermore, size distortions appear under the existence of nonzero correlation between the error sequences $u_t$ and $v_t$.

\color{black}

In other words, when Cov $\left( u_t, v_t \right) \neq 0$, then under the assumption of nonstationary predictors which is captured via the Local to unit root specification, then the limiting distribution no longer follows the standard NBB result of Andrews. In fact in this case it depends on the degree of persistence $c_i$.

We observe that even though the simulated asymptotic values of the sup IVX-Wald test for detecting a single structural break in the predictive regression model with no model intercept is not quite close to the 8.85 cv that corresponds to the NBB result of Andrews, we can verify that the asymptotic result does not depend on the nuisance parameter $c$.

Moreover, from the empirical simulations across different values of $\rho$ we can see that the simulated asymptotic critical values are quite close when observing at a specific value of $c$. This is not surprising since the fully modified covariance estimator incorporated in the construction of the covariance matrix for the IVX takes into account the dependence structure of the regressors by applying a common long-run covariance structure. This holds across different values of $c_i$. \color{black}

Illustrative Examples

We consider for instance the case of bivariate predictor persistence and examine the finite-sample performance of the proposed tests. In this case, we use the $c_1 = 1$ and $c_2 = 5$ considering this way two predictors with different degree of persistence, however which belongs to the same persistence class as defined in KMS. In particular, the quantile results we obtain since are based on finite-sample approximations then they have slightly lower value of the true asymptotic quantile that corresponds to the asymptotic distribution of the test.

\color{black}

Power Comparison

To conduct a power comparison we compare the rejection rates under the alternative hypothesis of structural change, $\mathbb{H}_1: \underline{\beta}_1 \neq \underline{\beta}_2$. Specifically, under the alternative hypothesis, we generate data using the predictive regression given by expression (ref) and then find statistical evidence against the null hypothesis by computing the two test statistics. In particular, comparing the same value of the localizing coefficient across different sample sizes, the power function is indeed monotonically increasing.

We consider a sequence of local alternatives of the form $\widetilde{\underline{\beta}} = \left( \underline{\beta}^0 + b / T\right)$ and implement the power function for breaks in both the model intercept and the coefficients of the predictors. We introduce parameter instability via the following:

align*[align* omitted — 250 chars of source]

Table 4 and 5 presents the probabilities of rejection of the two Wald-type statistics at the $5\%$ nominal rate, under the alternative hypothesis. The proposed methodology is recommended in cases in which the practitioner has some prior information regarding the persistent properties of predictors included in the predictive regression. Moreover, for a given value of $c_z$ (e.g., $c_z \equiv 1$), then for a larger exponent rate $( \delta = \delta_0 )$ in the IVX instrument can lead to higher power in exprense of lower values for the empirical size, while a smaller $\delta_0$ can achieve better size corrections. The number of predictors can be another source of distorted inferences and thus in high dimensional settings further corrections might be needed to correct the size and power of the proposed tests.

The Monte Carlo experiments verify that our proposed IV based Wald test provides a solution to the problem of possible size and power distortions when testing for parameter instability in predictive regressions with persistent predictors. Implementation is therefore straightforward via the use of standard statistical tables. In particular, once the magnitude of the sup-Wald IVX statistic has been computed in the case of a known break-point it then suffices to obtain the relevant critical values from Table 1 in andrews1993tests (see e.g., pitarakis2008comment).

remarkNote that simulating critical values as the results given by Table 1 of andrews1993tests, can be done with other numerical approximation methods. Further studies among others include the papers of estrella2003critical and anatolyev2012another. Moreover, hansen1997approximate proposes a methodology to obtain simulated p-values for stability tests in linear regression models, which can be used for the empirical size of the Wald-OLS statistic in our framework.
table[table omitted — 2,815 chars of source]
smallTable (ref) presents finite-sample empirical sizes for the sup Wald-OLS test, with nominal level $\alpha = 5\%$ for $B = 5,000$ replications. The predictive regression model under the null hypothesis, $H_0: \beta_1 = \beta_2$, is given by, $y_{t} = 0.25 x_{t-1} + u_t$, $x_t = (1- \frac{c}{T}) x_{t-1} + v_t$, with $\Sigma_{ee} = \begin{bmatrix} 1 & \rho \\ \rho & 1 \end{bmatrix}$.
table[table omitted — 2,818 chars of source]
smallTable (ref) presents finite-sample empirical sizes for the sup Wald-OLS test, with nominal level $\alpha = 5\%$ for $B = 5,000$ replications. The predictive regression model under the null hypothesis, $H_0: \beta_1 = \beta_2$, is given by, $y_{t} = 0.25 x_{t-1} + u_t$, $x_t = (1- \frac{c}{T}) x_{t-1} + v_t$, with $\Sigma_{ee} = \begin{bmatrix} 1 & \rho \\ \rho & 1 \end{bmatrix}$.

\color{red} OLD TABLE (to update with the ones after bootstrapping) \color{black}

table[table omitted — 7,168 chars of source]
smallTable (ref) presents finite-sample sizes for the sup Wald-IVX test, with nominal size $\alpha = 5\%$. The predictive regression model under the null hypothesis is given by $y_{t} = 0.25 + 0.5 x_{t-1} + u_t$, $x_t = (1- \frac{c_1}{T}) x_{t-1} + v_t$, with $\Sigma_{ee} = \begin{bmatrix} 0.25 & \sigma_{uv} \\ \sigma_{uv} & 0.75 \end{bmatrix}$, $\rho = \displaystyle \frac{ \sigma_{uv} }{ \sigma_u \sigma_v }$, as given above. Furthermore, for the IVX estimation step, we use an IVX persistence parameter $\delta \in \left\{ 0.75, 0.95 \right\}$ and the localizing coefficient is set to $c_{z} = 1$. The number of replications is $B = 5,000$.

Empirical Application

In this section we present an empirical application aiming to shed light on the literature of stock return predictability. Identifying periods of predictability has important implications in various aspects of finance. Related modern reviews of these aspects are presented by kostakis2018taking and chinco2019sparse, among others. However, despite the extensive research of the field, the findings are still rather mixed kasparis2015nonparametric (see, welch2008comprehensive for a full discussion). For instance, aspects such as the chosen sample period or the selected predictors can give different conclusions. Additionally, parameter instability due to certain economic events can also affect the reliability of predictability tests. Therefore, it is of paramount importance to develop robust testing methodologies for inferring predictability under conditions such as parameter instability or the presence of nonstationary regressors. Using the predictive regression model studied in this paper\footnote{Alternative model specifications can be considered; for instance a model which considers expected returns in relation to macroeconomic conditions and forecasting uncertainty. A first move towards this direction is presented by atanasov2020consumption who examine consumption fluctuations and expected returns with respect to the predictability literature. More specifically, the authors indeed find statistical evidence of predictability at the one-quarter horizon using the IVX testing approach of KMS.} our primary focus is to examine the robustness of the proposed tests.

Data Description

We focus on examining the presence of predictability for monthly US stock market excess returns over the period 1990-2019 using the set of variables considered in welch2008comprehensive\footnote{The dataset can be retrieved from Amit Goyal's website at \url{http://www.hec.unil.ch/agoyal/}. Detailed descriptions of variables can be found in the Online Appendix of welch2008comprehensive. } which capture economic and financial conditions for the US economy.

\paragraph{Predictant} The dependent variable is the monthly equity premium (excess return) of the US stock market based on the S$\&$P500 index. We construct the excess return as in kasparis2015nonparametric, that is, the difference between the total rate of return and the risk-free rate for the same sample period. As a proxy of the US stock market return, we use the value-weighted S$\&$P500 total stock market return including dividends. The risk-free rate is the 3-month T-bill rate obtained from the database of FRED\footnote{Time series of macroeconomic variables, such as the US inflation rate and the T-bill rate can be found at \url{ https://fred.stlouisfed.org/}. Notice also that the proxy of equity premium and the other financial variables we consider in this paper, are commonly used in the predictability literature, see gonzalo2012regime, kasparis2015nonparametric,kostakis2015Robust and kostakis2018taking.}.

\paragraph{Predictors} The predictor variables we consider include: dividend-payout ratio (d/e), long-term yield (lty), dividend yield spread (dy), dividend-price-ratio (d/p), T-bill rate (tbl), earnings-price-ratio (e/p), book-to-market ratio (b/m), default yield spread (dfs), net equity expansion (ntis), term spread (tms) and inflation rate (inf).

Predictability Tests

To begin with, we apply the simple predictability test on the full sample for the period 1946-2019 using as predictant the S$\&$P500 Equity Premium. However, in this paper we take a slightly different approach than the literature. In particular, we focus on the subsample spanning the period 1990Q1 to 2019Q4, and consider monthly sampling frequency. Our first goal is to examine the stock return predictability of this subsample, that is, to identify the financial variables which are individually statistical significant as well as to identify for evidence of joint statistical significance. Furthermore, our second goal is to repeat the same exercise using the proposed joint test of predictability and parameter instability and compare the results we obtain. The chosen subsample includes the period of the 2008 financial crisis, so it is natural to assume that certain predictors might exhibit structural break around that economic event. Therefore, this is a suitable sample to assess the statistical validity of the proposed methodology for the case of a single structural break. Certain limitations of our approach are on sight, however these do not invalidate our findings. In particular, we do not consider the existence of multiple structural breaks neither we consider sample splitting techniques which can affect the power of the tests especially when using an out-of-sample forecasting scheme.

Firstly, summarizing our findings is Table (ref) which presents predictability tests based on both the classical least squares estimator and the IVX estimator. Notice that for these set of tests we consider the regressors on their original form in order to preserve the degree of persistence and have comparability between the two estimators under examination. Table (ref) presents structural break tests for the regressors. In this case the traditional structural break tests are implemented under the assumption of stationary time series, by taking the first difference of the regressors before fitting the AR(1) models.

Secondly, we examine the short-horizon predictability via the proposed joint predictability and structural break Wald statistics.

Conclusion

In this paper, we have extensively examined the asymptotic theory of tests for Joint predictability and parameter instability under the assumption of nonstationary regressors. We compare our results with previous seminal work in the fields of both structural break testing in linear regressions as well as predictability testing in predictive regression models. We find some interesting results not previously presented in the literature. Firstly, using the OLS estimator for the parameters of the predictive regression model we show that the limiting distribution of the Wald statistic for testing for a single structural break has a nonstandard limiting distribution which depends on the unknown coefficient of persistence. Secondly, by employing the IVX estimator proposed in the literature as a robust estimator which filters out the abstract degree of persistence in regressors, we have proved that the limiting distribution of the Joint tests takes different forms which weakly converges to a functional of a Brownian bridge in some instances while converges to a nonstandard limiting distribution which depends on the coefficient of persistence when regressors are assume to be highly persistent.

Conducting inference, such as structural break testing, on the regression coefficient of the predictive regression model with multiple highly persistent regressors can lead to a nonstandard limiting distribution. The proposed testing methodology ensures that the limiting distribution of the structural break tests is free of any nuisance parameters, such as the unknown localizing coefficient of persistence. Moreover, we consider "pure" structural change as it is defined by andrews1993tests, in the sense that the entire parameter vector is subject to structural change under the alternative hypothesis. The asymptotic distribution of the IV based Wald test is found to be given by the supremum of the normalized squared Brownian Bridge (NBB). Thus, the exact limiting distribution of the IV based Wald test allows to determine critical values similar to the case of the linear regression model, without further simulations and additional computational cost. This holds in the case of mildly integrated or integrated regressors, while in the case of nonstationary regressors further investigation is needed to determine exact critical values since the limiting distribution has a dependence on the nuisance coefficient of persistence.

The developed asymptotic theory for the sup-Wald IVX tests under various degrees of persistence we consider in this paper, indicate that the asymptotic behaviour of the tests when the supremum functional is included, is different from the corresponding limit theory in the case of only linear restrictions to the parameters of the predictive regression. Nevertheless, the robustness of the IVX instrumentation provides a way to determine an analytic form of the asymptotic distribution for different levels of persistence, a feature often seen in time series data. This feature appears in various empirical finance applications in which the available information set for current or future economic conditions with many times persistent properties and existence of parameter instability.