EconBase
← Back to paper

Sequential monitoring for cointegrating regressions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

83,716 characters · 15 sections · 98 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Sequential monitoring for cointegrating regressions

frontmatter\runtitle{Monitoring cointegration} \begin{aug} \hskip .2cm \runauthor{L. Trapani and E. Whitehouse} \address{$^{\dag}$University of Nottingham\\ \printead{e1}\\ } \address{$^{*}$Newcastle University\\ \printead{e2}\\ } \end{aug} \begin{abstract} We develop monitoring procedures for cointegrating regressions, testing the null of no breaks against the alternatives that there is either a change in the slope, or a change to non-cointegration. After observing the regression for a calibration sample $m$, we study a CUSUM-type statistic to detect the presence of change during a monitoring horizon $m+1,...,T$. Our procedures use a class of boundary functions which depend on a parameter, $0 \leq \eta \leq \frac{1}{2}$, whose value affects the delay in detecting the possible break. Technically, these procedures are based on almost sure limiting theorems whose derivation is not straightforward. We therefore define a monitoring function which - at every point in time - diverges to infinity under the null, and drifts to zero under alternatives. We cast this sequence in a randomised procedure to construct an i.i.d. sequence, which we then employ to define the detector function. Our monitoring procedure rejects the null of no break (when correct) with a small probability, whilst it rejects with probability one over the monitoring horizon in the presence of breaks. \end{abstract} \begin{keyword}[class=MSC] \kwd{62F03} \kwd{62L10} \kwd{62M10} \end{keyword} \begin{keyword } \kwd{cointegration} \kwd{structural change} \kwd{sequential monitoring} \kwd{randomized tests} \end{keyword}

Introduction

In this paper, we study the following cointegrating regression

equation[equation omitted — 96 chars of source]

where $\left( y_{i},x_{i}^{\prime }\right) ^{\prime }$ is a $(p+1)\times 1$, $I\left( 1\right) $ vector and $\epsilon _{i}\ $is a stationary innovation. In particular, we investigate the issue of monitoring ((ref)), after a calibration period of length $m$, during which our maintained assumptions are that (i) ((ref)) is a cointegrating relationship and (ii) the slope $\beta $ is constant. From $i=m+1$ onwards, we check whether the relationship in ((ref)) remains constant, or whether either the slope $\beta $ changes, or ((ref)) becomes a non-cointegrating regression (or both).

The timely detection of structural change is arguably of great importance in the context of any regression model: whilst there is an extensive literature on the general topic of on-line detection of changes (see e.g. csorgo1997 for a survey), in the econometrics literature this issue has received some limited attention since the contribution by chu. Recent articles that study this topic have focused on linear regression models (lajos04, aue2006, lajos07, kap-monitor), large factor models (bt1), and also cointegrating regressions (steland, wied, wagner). In particular, wied consider, essentially, the same problem as in our paper; namely, they propose several statistics for the on-line detection of structural breaks in a model like ((ref)), considering both the possibility of a change in the slope $\beta$ and a change to a non-cointegrating regression.

From a methodological viewpoint, we use a residual-based detector to test for the null hypothesis of no change over the monitoring horizon $m+1\leq i\leq T$. Note that this corresponds to a closed-end procedure (aue12), since monitoring - as can be expected to happen in practice - stops after $T$. The family of detectors which we propose here are based on the sum of squared residuals. Simulations show that our monitoring scheme has excellent finite sample properties, with low occurrence of false detections and very good power versus both alternatives under consideration. Other detectors are also possible (see, for example, the various statistics considered in homm, albeit in a different context).

From a technical point of view, as pointed out by lajos04 and lajos07, the detectors employed in monitoring procedures depend upon a parameter, henceforth denoted as $\eta$, which can vary in the interval $\left[ 0,\frac{ 1}{2}\right] $. Constructing test statistics when $\eta =0$ (see e.g. chu) requires, as a technical tool, weak convergence, and therefore one can employ a huge variety of results which are well-known in the literature (see e.g. the book by billingsley). On the other hand, the choice $\eta =0$ is known to often yield inferior results, in particular resulting in a longer delay in detection of a break (aue2004). In order to overcome this issue, it is usually recommended to choose $\eta >0$ (lajos07). However, from a technical viewpoint, using $\eta >0$ requires having stronger forms of convergence than weak convergence, with fewer results available (we refer to the textbook by csorgo1997 for an excellent treatment of the subject). For example, to the best of our knowledge we are not aware of strong approximations like the ones derived in kmt1 and kmt2 for convergence to stochastic integrals, where usually \textquotedblleft weak\textquotedblright\ results are used instead (see chanwei; and phillipsweak). In light of this, we only rely on (almost sure) rates, and we develop a family of statistics - computed at each $m+1\leq i\leq T$ - which diverge to positive infinity under the null of no break, whilst they drift to zero in the presence of breaks. We then randomize such statistics at each point in time $i $: the outcome of our randomisation is a sequence of random variables which, under the null of no break, are i.i.d. with finite moments up to any order, whilst they diverge to infinity in the presence of a break. Finally, we employ the newly generated sequence in order to construct the same detectors as in lajos04 and lajos07, being able to rely on the theory spelt out in those papers. Using randomisation is helpful when the properties of a certain statistic are not known, or depend on nuisance parameters: in this respect, it might be envisaged that randomisation serves a similar purpose to the bootstrap or to self-normalisations (see dette2019likelihood for an example of self-normalisation in the context of monitoring). In the econometric literature, randomisation has been employed in a wide variety of contexts, including testing for forecasting superiority (corradi2006), stationarity ( bandi2014), finiteness of moments (trapani16), boundary problems (HT16) and determining the number of common factors in a large factor models (trapani17). In our context, however, we do not employ randomisation to produce a randomised test, but to construct a \textquotedblleft well-behaved\textquotedblright\ sequence which, in turn, can be employed to define an easy-to-study test statistic. In this respect, our contribution uses the same approach as bt1, who study monitoring for structural change in the context of a large, stationary factor models. By relying solely on rates, we require quite mild assumptions; all the theory can be based on using a standard OLS\ estimator, with no need for more specialised estimators like, say, the FM-OLS estimator (phillips-hansen) or a Dynamic OLS estimator (saikkonen); and, finally, we do not need to rely on the accuracy of the long-run variance estimator.

The remainder of the paper is organised as follows. In Section (ref), we provide the relevant assumptions, and then report theoretical results on estimation and the monitoring procedure. Extensions to e.g. the case of deterministics are in Section (ref). In Section (ref) we demonstrate the performance of our monitoring procedure through both a Monte Carlo simulation exercise (Section (ref)) and an empirical application to US housing market data (Section (ref)). Section (ref) concludes. Proofs of the main results are in Section (ref). All technical lemmas and some proofs are relegated to the Supplement.

Throughout the paper we use the notation $c_{0}$, $c_{1}$,... to denote positive and finite constants, that do not depend on the sample size; their value is allowed to change from line to line. We use the expression \textquotedblleft a.s.\textquotedblright\ as short-hand for \textquotedblleft almost surely\textquotedblright ; the ordinary limit is denoted as \textquotedblleft $\rightarrow $\textquotedblright . Finally, for a vector $a $ and a matrix $A$, $\left\Vert a\right\Vert $ and $\left\Vert A\right\Vert $ represent the Euclidean norm. Other notation is introduced later on in the paper.

Theory

We begin with introducing some notation and the main assumptions that should hold under the null of no break (Section (ref)); we then move to discuss the two alternative hypotheses which we consider, namely a change in the slope and/or a change to a non-cointegrating equation (Section (ref)). Finally, in Section (ref), we discuss the relevant CUSUM process, and the randomisation algorithm.

Main assumptions

Recall ((ref))

equation*[equation* omitted — 59 chars of source]

which we assume to be valid during the calibration period $1 \leq i \leq m$, with

equation[equation omitted — 53 chars of source]

We also define the long run variances of $u_{i}$ and $\epsilon _{i}$ as

align[align omitted — 326 chars of source]

We consider the following assumption.

assumptionIt holds that: (i) $\epsilon _{i}$ and $u_{i}$ have mean zero with (a) $E\left\vert \epsilon _{i}\right\vert ^{2}<\infty $ for $1\leq i\leq T$, and $0<\sigma _{\epsilon }^{2}<\infty $; and (b) $\Sigma _{u}$ is positive definite with $\left\Vert \Sigma _{u}\right\Vert $; (ii) $ E\left\Vert x_{0}\right\Vert ^{2}<\infty $ and \begin{equation} \sup_{1\leq i\leq t}\left\Vert x_{i}-W_{x}\left( i\right) \right\Vert =O_{a.s.}\left( t^{1/2-\delta ^{\prime }}\right) , \end{equation} for some $0<\delta ^{\prime }<\frac{1}{2}$, where $W_{x}\left( i\right) $ is a $p$-dimensional Wiener process with increments of variance $\Sigma _{u}$; (iii) $E\left\Vert \sum_{i=1}^{t}x_{i}\epsilon _{i}\right\Vert ^{2}\leq c_{0}t^{2}$, for all $1\leq t\leq T$; (iv) $E\left\Vert \sum_{i=1}^{t}x_{i}x_{i}^{\prime }\right\Vert \leq c_{0}t^{2}$, for all $ 1\leq t\leq T$.

Assumption (ref)(i) is a standard second moment condition which is required to hold under the null of no change, and also when the slope $\beta $ changes but ((ref)) remains a cointegrating relationship. Note that, by part (i)(b), we rule out cointegration among the regressors. Part (ii) of the assumption, in essence, states that a strong approximation exists for the partial sums process $x_{i}$. This is a high-level assumption, which could be replaced by more primitive requirements on the existence of moments for the innovation $u_{i}$, and some form of weak dependence. It can be envisaged, as far as moments are concerned, that at least $E\left\Vert u_{i}\right\Vert ^{2+\delta }<\infty $ is required for some $\delta >0$; thence, ((ref)) would follow immediately if $u_{i}$ is i.i.d. (see kmt1 and kmt2 for the univariate case, and gotze for the multidimensional one), and also under fairly general forms of weak dependence such as the case of stationary causal processes including linear models, Volterra series and models with conditional heteroskedasticity (see wu2005, and berkesliuwu). Interestingly, in the literature it is relatively common to assume a weak Invariance Principle to hold in lieu of assuming weak dependence (see e.g. Assumption 2 in wied). Part (ii) of Assumption (ref) serves exactly the same purpose, except for the fact that in our paper we need almost sure rates. Parts (iii) and \textit{(iv)} could also be derived under more primitive conditions on moments, serial dependence, and possible correlation between $u_{i}$ and $\epsilon _{i}$. For example, the results could be shown by standard arguments in the case of $u_{i}$ and $\epsilon _{i}$ being \textit{i.i.d.} and independent of each other; in this case, existence of second moments would suffice. Part \textit{(iv)} can be shown under more general forms of dependence, e.g. in the case of linear processes by exploiting the results in solo. Also, in part \textit{(iii)}, the requirement of independence between $u_{i}$ and $\epsilon _{i}$ is not necessary: again under the assumption of linear processes, for example, it could be shown (see, \textit{inter alia}, durlauf, and phillips-hansen) that this part of the assumption can hold also in the presence of endogeneity.

Hypotheses of interest and the construction of the monitoring procedure

We base our on-line monitoring on the theory developed in lajos04 and lajos07. We assume that the data are collected for an initial calibration period of size $m$ where no break occurs; this can be viewed as the historic sample available to the researcher. We then define the (length of) the monitoring horizon $T_{m}$ as $T_{m}=T-m$. Thus, if $T$ represents the total period considered, $m$ is the amount of time elapsed until the beginning of the monitoring period. In essence, $m$ is going to be the sample size used by the researcher for estimation. Choosing $T_{m}$ - that is, choosing where to stop the monitoring - is an important issue in sequential analysis, since it can be argued that monitoring comes at a cost (see the original paper by wald); in this paper, we allow for $ T_{m}\rightarrow \infty $, under the assumption that monitoring is costless - this assumption can be realistic when analysing economic series, although not in other contexts (see e.g. the comments in chu).

Alternative hypotheses of interest

Under the null hypothesis of our monitoring scheme, ((ref)) is a cointegrating relationship for the whole monitoring horizon, and the slope $ \beta $ does not change; rewriting ((ref)) as

equation*[equation* omitted — 63 chars of source]

we have

equation[equation omitted — 166 chars of source]

Conversely, when the null does not hold, there could be at least two interesting, non mutually exclusive alternatives. In the first case, there could be a structural change whereby, after $i=m$, $\beta $ changes:

equation[equation omitted — 100 chars of source]

In ((ref)), $m\leq k^{\ast }<T$ is the potential breakdate. In addition to this (or as an alternative), ((ref)) may switch to being a non-cointegrating relationship at some point in time, viz.

equation[equation omitted — 125 chars of source]

In both cases, the case of no break is represented by having $k^{\ast }=T$.

As a general comment to our hypothesis testing framework, we point out that the set-up in ((ref))-((ref)) mirrors the analysis in wied very closely. In particular, the null hypothesis is the intersection of two (very different) requirements: (a) the fact that there is no time variation in the structural parameter $\beta $ in ((ref)) over the monitoring horizon, under the implicitly maintained hypothesis that ((ref)) is always a cointegrating regression; and (b) the fact that ((ref)) is indeed a cointegrating relationship during the monitoring horizon. This could be the the set-up of interest in various applications (see e.g. Section (ref)); furthermore, an \textquotedblleft omnibus\textquotedblright\ procedure which is powerful versus a global alternative could be viewed as advantageous in order to avoid having to test under a maintained hypothesis whose validity may not always be assumed. On the other hand, the monitoring procedure proposed in this paper (and in wied) can be argued to be \textquotedblleft non-constructive\textquotedblright : after rejecting the null and finding evidence of a change in the nature of ((ref)), it is not clear which of the two alternatives the change can be ascribed to. In the literature, there are tests for more focussed alternatives which could, in principle, be extended into monitoring procedures. For example, under the maintained assumption that $\epsilon _{i}$ is stationary over the whole monitoring horizon, one could think of extending the test for breaks in cointegrating regressions proposed by kejriwal. Similarly, under the maintained assumption that $\beta $ is constant for the whole interval $ 1\leq i\leq T$, a monitoring procedure could, in principle, be constructed using the residuals $\widehat{\epsilon }_{i}$, e.g. by extending the test for a change in persistence proposed by busetti. Indeed, under the same maintained hypotheses mentioned above, our procedure can also be employed to test, separately, versus the two alternatives mentioned above. In this respect, the monitoring scheme proposed in this paper could be viewed as a preliminary step: upon finding evidence that a change occurred, the researcher may decide to use a more specialised procedure to disentangle the nature of the change in ((ref)).

In order to analyse the case of ((ref)), we need the following assumption which characterizes the behaviour of $\epsilon _{i}$ under $ H_{A,2}$.

assumptionUnder $H_{A,2}$, it holds that (i) \begin{equation} \sup_{k^{\ast }+1\leq i\leq t}\left\vert \epsilon _{i}-W_{\epsilon }\left( i\right) \right\vert =O_{a.s.}\left( t^{1/2-\delta ^{\prime \prime }}\right) , \end{equation} for all $k^{\ast }+1\leq t\leq T$ and some $0<\delta ^{\prime \prime }<\frac{ 1}{2}$, where $W_{\epsilon }\left( i\right) $ is a Wiener process with increments of positive variance equal to the long-run variance of $ u_{i}^{\epsilon }$; (ii) \begin{equation*} E\left\vert \sum_{i=k^{\ast }+1}^{t}\epsilon _{i}^{2}\right\vert \leq c_{0}t^{2}, \end{equation*} for all $k^{\ast }+1\leq t\leq T$.

Assumption (ref) supersedes parts (i) and (iii) of Assumption (ref) in order to accommodate for the presence of a switch to a non-cointegrating regression. According to the assumption, in essence, after the breakdate $k^{\ast }$ the innovation $\epsilon _{i}$ becomes a unit root process.

The monitoring function

Our monitoring scheme is based on a non-recursive estimator of $\beta $: estimation is carried out using the sample $1\leq i\leq m$ once and for all, without updating the estimate as $i$ elapses. We focus only on this merely for the sake of a concise discussion: this choice is not the only possible one. lajos04, inter alia, propose a recursive monitoring procedure (as well as a non-recursive one), where $\beta $ is estimated at each $i$ using an expanding sample. It seems reasonable to conjecture that, even in our context, the non-recursive scheme is probably likely to be less affected by outliers, thus ensuring a better size control, whilst the recursive procedure should be, by design, more sensitive to breaks.

Let

equation[equation omitted — 130 chars of source]

where dependence on the sample size $m$ will be omitted whenever possible, and define the residuals

equation[equation omitted — 171 chars of source]

for $m+1\leq i\leq T$ onwards. At each $k$, we define the cumulative process

equation[equation omitted — 162 chars of source]

for $1\leq k\leq T_{m}$.

commentWe begin by providing some heuristic arguments to motivate our choice. Our approach is based on the cumulative sums of some transformation of the residuals, as is natural in these contexts (see e.g. wied, lajos04) - although, somewhat differently from wied, we found it easier to study cumulative processes involving squares of residuals. The main idea, however, is that, in the presence of a break, such cumulative sums would grow much larger, over time, than their \textquotedblleft natural\textquotedblright\ growth rate in the absence of breaks. Based on well-known properties of unit root processes (see e.g. durlauf), it can be expected that the fluctuations of the monitoring functions should grow linearly over time if there is no break; if there is a break, some \textquotedblleft non-centrality\textquotedblright\ term should arise which makes the fluctuations grow at a rate faster than linear. Upon expanding the summands $x_{i}\widehat{\epsilon }_{i}$ and $ \widehat{\epsilon }_{i}^{2}$, it can be seen that the \textquotedblleft non-centrality\textquotedblright\ of $Q_{1}\left( m;k\right) $ is given by $ \Delta _{\beta }x_{i}^{2}$ under $H_{A,1}$: the cumulative sums of this term should diverge quadratically over time, and such divergence should be strong, in that it can be expected that \begin{equation} \lim \inf_{t\rightarrow \infty }\frac{\ln \ln t}{t^{2}} \sum_{i=1}^{t}x_{i}^{2}>0 a.s. \end{equation} This would ensure the ability to detect a break in $\beta $. On the other hand, under $H_{A,2}$, the \textquotedblleft non-centrality\textquotedblright\ of $Q_{1}\left( m;k\right) $ is given by $ x_{i}\epsilon _{i}$. Whilst it is true that the partial sums of $ x_{i}\epsilon _{i}$ grow at a quadratic rate, there is no guarantee that this growth is as strong as in ((ref)). Indeed, one can expect \footnote{ This can be understood even better by considering the zero-mean, continuous process $\Gamma \left( u\right) =W_{x}\left( u\right) W_{\epsilon }\left( u\right) $ for which $\Gamma \left( 0\right) =0$. Let \begin{equation*} \Delta \left( r\right) =\int_{0}^{r}\Gamma \left( u\right) du. \end{equation*} If $\Delta \left( r\right) $ has no zeros on the interval $\left( 0,a\right] $, then $E\left[ \Delta \left( r\right) \right] \neq 0$ on $\left( 0,a\right] $, which contradicts the fact that $E\left[ \Delta \left( r\right) \right] =0 $ due to $E\Gamma \left( u\right) =0$ for all $u$.} \begin{equation} \lim \inf_{t\rightarrow \infty }\sum_{i=1}^{t}x_{i}\epsilon _{i}=0 a.s. \end{equation} We conjecture that the following Chung-type LIL (chung) should hold \begin{equation} \lim \inf_{t\rightarrow \infty }\frac{\ln \ln t}{t^{2}}\sup_{1\leq j\leq t}\left\vert \sum_{i=1}^{j}x_{i}\epsilon _{i}\right\vert >0 a.s., \end{equation} but a result like ((ref)) is obviously much weaker, as far as the divergence of the monitoring function is concerned, than that of ((ref)). In light of these considerations, a procedure based on $ Q_{1}\left( m;k\right) $ should have power versus $H_{A,1}$, and may have some, limited power, under $H_{A,2}$. Conversely, the monitoring function $Q_{2}\left( m;k\right) $ is designed to ensure power versus $H_{A,2}$: its \textquotedblleft non-centrality\textquotedblright\ is driven by the cumulative sums of $ \epsilon _{i}^{2}$ for which, under $H_{A,2}$, a result like ((ref) ) holds. Under $H_{A,1}$, however, it can be seen that the \textquotedblleft non-centrality\textquotedblright\ term is driven by

Estimation of $\protect\sigma _{\protect\epsilon }^{2}$

In ((ref)), $\widehat{\sigma }_{\epsilon }^{2}$ is an estimator of $ \sigma _{\epsilon }^{2}$. In our paper, we use a weighted-sum-of-covariance estimator. In order to apply our theory, we need to show the almost sure convergence of $\widehat{\sigma }_{\epsilon }^{2}$ to a positive limit; thus, this section of our paper can be compared to berkesbartlett.

Let $\rho _{l}^{\left( \epsilon \right) }$ denote the $l$-th order autocovariance of $\epsilon _{i}$, i.e. $\rho _{l}^{\left( \epsilon \right) }=E\left( \epsilon _{i}\epsilon _{i-l}\right) $. This can be estimated as

equation[equation omitted — 155 chars of source]

Based on ((ref)), we define

equation[equation omitted — 209 chars of source]
commentIn ((ref)), we have used the Bartlett kernel with bandwidth $H$, although other choices are also possible as long as standard assumptions are satisfied - see e.g. andrews1991.

Let $y_{i,l}^{\left( \epsilon \right) }=\epsilon _{i}\epsilon _{i-l}-\rho _{l}^{\left( \epsilon \right) }$. We need the following regularity conditions

assumptionIt holds that: (i) $\epsilon _{i}$ is covariance stationary with $ E\left\vert \epsilon _{i}\right\vert ^{4}<\infty $ for all $i$; (ii) $ \sum_{l=0}^{\infty }l\left\vert \rho _{l}^{\left( \epsilon \right) }\right\vert <\infty $; (iii) $E\left\vert \sum_{i=l+1}^{m}y_{i,l}^{\left( \epsilon \right) }\right\vert ^{2}\leq c_{0}m $.

It holds that

propositionWe assume that Assumptions (ref)-(ref) are satisfied. As $\min \left( m,H\right) \rightarrow \infty $ \begin{equation} \widehat{\sigma }_{\epsilon }^{2}=\sigma _{\epsilon }^{2}+o_{a.s.}\left( \frac{H}{m^{1/2}}\left( \ln m\right) ^{3+\varepsilon }\left( \ln \ln m\right) \left( \ln H\right) ^{2+\varepsilon }\right) +O\left( \frac{1}{H} \right) , \end{equation} for every $\varepsilon >0$.

In Proposition (ref), a crucial role is played by the bandwidth $H$ . In order to ensure consistency, ((ref)) requires that $H\rightarrow \infty$ and

equation*[equation* omitted — 146 chars of source]

as $m\rightarrow \infty $.

commentso that the speed of convergence is maximised by choosing \begin{equation} H=H\left( m\right) =c_{0}\frac{m^{1/4}}{\left( \ln m\right) ^{5/2+\varepsilon }\left( \ln \ln m\right) ^{1/2}}. \end{equation}

The monitoring scheme

The main idea underpinning ((ref)) is that, by construction, $Q\left( m;k\right) $ should pick up the presence of a break, which would introduce a drift in its fluctuations. In order to check whether $Q\left( m;k\right) $ is growing \textquotedblleft naturally\textquotedblright , i.e. without breaks, or not, we introduce the function

equation[equation omitted — 128 chars of source]

where the choice of $\gamma $ depends on the length of the monitoring horizon. Heuristically, the function $g\left( m;k\right) $ should control the growth rate of $Q\left( m;k\right) $: this is driven by a term proportional to the cumulative sum of $\epsilon _{i}^{2}$ - which is controlled by $m+k$ in ((ref)) - and one proportional to the cumulative sum of $x_{i}^{2}$, multiplied by the (square of the) estimation error $ \beta -\widehat{\beta }_{m}$ - which is controlled by the term $\left( \frac{ m+k}{m}\right) ^{2}$ in ((ref)).

assumptionIt holds that: (i) $T_{m}=c_{0}m^{\theta }$ for some $\theta >1$ and $0 < c_{0} < \infty$; (ii) if $k^{\ast } < T$, $k^{\ast }=O\left( m^{\theta ^{\prime }}\right) $ with $0\leq \theta ^{\prime }<\theta $; (iii) $\lim \inf_{m\rightarrow \infty }\frac{T_{m}}{m}>0$.

Assumption (ref) states that the monitoring horizon should go on for a sufficiently long time (part (i)), and obviously include the breakdate if there is a break (part (ii)). In particular, part (i), with its implications, is very similar to equation (1.12) in lajos07, who also consider the case where monitoring goes on for an infinite time (unless a change is detected).

In practice, $\theta $ is also a given parameter, which is calculated from Assumption (ref)(i), once $m$ and $T_{m}$ have been set. Hence, $\gamma $ is calculated according to the rule

equation[equation omitted — 67 chars of source]

where $\delta $ is chosen as $0<\delta <1$. In principle, any value of $ \delta $ will ensure the validity of the theory below. In essence, $\gamma$ is chosen as a fraction of $\frac{1}{\theta-1}$; clearly, choosing $\delta$ close to $1$ yields a small $\gamma$, which in turn makes the divergence of $g\left( m;k\right) $ as $m \rightarrow \infty$ slower than in the case of a $\delta$ closer to zero. We discuss the practical impact of the choice of $\delta $ (and $\gamma $) on the ability of the monitoring procedure to detect breaks in Section (ref). \newline The function $g\left( m;k\right) $ has been chosen so as to distinguish the growth rate that $Q\left( m;k\right) $ should have if there were no break, from the rate at which it would diverge if there were a break. Heuristically, in absence of breaks, $Q\left( m;k\right) $ should grow, but slower than $g\left( m;k\right) $; on the other hand, if there is a break, its presence in the residuals $\widehat{\epsilon }_{i}$ should make $Q\left( m;k\right) $ grow at a faster pace, and faster than $g\left( m;k\right) $ itself. We point out that the term $\left( m+k\right) $ in ((ref)) is a rather coarse estimate, and in principle it could be refined; however, this term is anyway dominated by the second component of $g\left( m;k\right) $, and ((ref)) yields very good results in simulations.

Define

equation[equation omitted — 89 chars of source]

Based on the above, we expect that $\psi _{m,k}$ drifts to zero as $m$ and $ T_{m}$ diverge if there is no break, whereas it should explode if there is a break; note that we only consider rates. Indeed, in order to separate such rates even better, we use the transformation

equation[equation omitted — 98 chars of source]

By construction, $\widetilde{\psi }_{m,k}$ has the opposite behaviour as $ \psi _{m,k}$: it can be expected that $\widetilde{\psi }_{m,k}$ drifts to zero in the presence of a break (that is, under the alternative); conversely, it should diverge to positive infinity if there is no break (that is, under the null). Indeed, in the Appendix, we prove that, as $ m\rightarrow \infty $

align*[align* omitted — 176 chars of source]

Given that the test statistic $\widetilde{\psi }_{m,k}$ does not converge to a non-degenerate random variable under the null (or the alternative), we propose to use a randomised version of $\widetilde{\psi }_{m,k}$. We present this as an algorithm, whose output will be a sequence of i.i.d. random variables, with a known distribution (at least asymptotically) under $ H_{0}$, and which diverge under $H_{A,1}$ and $H_{A,2}$.

description• For each $k$, generate an i.i.d. $N\left( 0,1\right) $\ sequence $\left\{ \xi _{j}^{\left( k\right) },1\leq j\leq R\right\} $. • Generate the Bernoulli sequence $\zeta _{j}^{\left( k\right) }\left( u\right) =I\left( \left\vert \widetilde{\psi } _{m,k}\right\vert ^{1/2}\xi _{j}^{\left( k\right) }\leq u\right) $. • Compute \begin{equation} \vartheta _{m,R}^{\left( k\right) }\left( u\right) =\frac{2}{R^{1/2}} \sum_{j=1}^{R}\left( \zeta _{j}^{\left( k\right) }\left( u\right) -\frac{1}{2 }\right) . \end{equation} • Define \begin{equation} \Theta _{m,R}^{\left( k\right) }=\int_{-\infty }^{+\infty }\left\vert \vartheta _{m,R}^{\left( k\right) }\left( u\right) \right\vert ^{2}dF\left( u\right), \end{equation} where $F\left( u\right) $ is a distribution.

Some comments on the sequence $\left\{ \Theta _{m,R}^{\left( k\right) },1\leq k\leq T_{m}\right\} $ are in order. Consider first the case of the null of no break. The Bernoulli random variable $\zeta _{j}^{\left( k\right) }\left( u\right) $ should - asymptotically - be equal to $1$ or $0$ with probability $\frac{1}{2}$, and thus have mean $\frac{1}{2}$. In this case, when constructing $\vartheta _{m,R}^{\left( k\right) }\left( u\right) $, a Central Limit Theorem holds and therefore we expect $\Theta _{m,R}^{\left( k\right) }$ to have a chi-square distribution. On the other hand, under the alternative of a break, $\zeta _{j}^{\left( k\right) }\left( u\right) $ should be (heuristically) $0$ or $1$ with probability $0$ or $1$ (depending on the sign of $u$) - thus its mean is not $\frac{1}{2}$, and a Law of Large Numbers should hold. Note finally that, by construction, conditionally on the sample, the sequence $\{\Theta _{m,R}^{\left( k\right) }\}_{k=1}^{T_{m}}$ is independent across $k$; also, by integrating out $u$ in Step 4, the statistic $\Theta _{m,R}^{\left( k\right) }$ becomes invariant to the choice of this specification.

The following regularity conditions are needed:

assumptionIt holds that: $F\left( u\right)$ is a non-degenerate continuous distribution with (i) $\int_{-\infty }^{+\infty }u^{2}dF\left( u\right) <\infty $; (ii) the sequences $\left\{ \xi _{j}^{\left( k\right) },1\leq j\leq R\right\} $ are independent across $k$.

Let now $P^{\ast }$ represent the conditional probability with respect to $ \left\{ u_{i},\epsilon _{i},1\leq i\leq T\right\} $; we use the notation \textquotedblleft $\overset{D^{\ast }}{\rightarrow }$\textquotedblright\ and \textquotedblleft $\overset{P^{\ast }}{\rightarrow }$\textquotedblright\ to define, respectively, conditional convergence in distribution and in probability according to $P^{\ast }$. It holds that

theoremWe assume that Assumptions (ref)-(ref) hold. As $\min \left( m,R\right) \rightarrow \infty $ with \begin{equation} R\exp \left( -m^{\gamma }\right) \rightarrow 0, \end{equation} under $H_{0}$\ it holds that, for each $1\leq k\leq T_{m}$ \begin{equation*} \Theta _{m,R}^{\left( k\right) }\overset{D^{\ast }}{\rightarrow }\chi _{1}^{2}, \end{equation*} for almost all realisations of $\left\{ u_{i},\epsilon _{i},1\leq i\leq T\right\} $.
theoremWe assume that Assumptions (ref)-(ref) hold. As $\min \left( m,R\right) \rightarrow \infty $, under $H_{A,1}\cup H_{A,2}$ \ it holds that, for each $k\geq \left\lfloor m^{\max \left\{ 1,\theta ^{\prime }\right\} \left( 1+\varepsilon \right) }\right\rfloor $ for all $ \varepsilon >0$ \begin{equation*} \frac{1}{R}\Theta _{m,R}^{\left( k\right) }\overset{P^{\ast }}{\rightarrow } 1, \end{equation*} for almost all realisations of $\left\{ u_{i},\epsilon _{i},1\leq i\leq T\right\} $.

Theorems (ref) and (ref) are intermediate results. Theorem (ref) stipulates that under the null $\Theta _{m,R}^{\left( k\right) }$ has, asymptotically, a $\chi _{1}^{2}$ distribution; this result is of independent interest, and we will make use of it to show that $\Theta _{m,R}^{\left( k\right) }$ has finite moments of order $2+\varepsilon $ with $\varepsilon >0$. Note that the only thing that is required is the fact that $\Theta _{m,R}^{\left( k\right) }$ is an i.i.d. sequence, with finite moments of order $2+\varepsilon $: this is the building block on which we can construct a detector whose properties can be studied analytically. In this respect, any other transformation of $\vartheta _{m,R}^{\left( k\right) }\left( u\right) $ (e.g., the absolute value, or a power thereof) will also work, giving exactly the same results as in Theorem (ref) below; the only advantage of defining $\Theta _{m,R}^{\left( k\right) }$ as in Step 4 above is that its asymptotics has already been studied (see e.g. HT16).

The theorems contain a restriction on the relative rate of divergence of the pre-monitoring sample size $m$ and the artificial sample size $R$. Heuristically, note that our monitoring procedure is based on having a bounded sequence with finite moments under the null. As the proof of Theorem (ref) shows, under the null the statistic $\Theta _{m,R}^{\left( k\right) }$ has a non-centrality term which vanishes as long as ((ref)) is satisfied. Conversely, under the alternative it is required that $\Theta _{m,R}^{\left( k\right) }$ should pass to infinity: Theorem (ref) ensures that this occurs at a rate equal to $R$. Thus, Theorem (ref) and ((ref)) provide a family of selection rules for $R$. Given that we only need convergence and divergence, the role played by $R$ can be expected to be rather marginal, which is also confirmed by our simulations (see Section (ref)). However, we note that the choice $R=m$ satisfies ((ref)). Theorem (ref), conversely, states that, under the alternative where there is a break at $k^{\ast }$, $\Theta _{m,R}^{\left( k\right) }$ diverges to positive infinity after $k^{\ast }$.

commentRecall that $\Theta _{m,R}^{\left( k\right) }$, by construction, is also independent across $k$, conditional on the sample.

In light of these results, we build a monitoring function, based on the use of the cumulative sums process. Define the detectors

equation[equation omitted — 172 chars of source]

As can be noted, $d\left( m;k\right) $ is the CUSUM\ process of $\left\{ \Theta _{m,R}^{\left( k\right) },1\leq k\leq T_{m}\right\} $, after centering and standardizing. \\

Similarly to the literature on structural breaks (see e.g. csorgo1997), we now need to define a family of threshold functions such that if the CUSUM process exceeds the threshold, a change is detected. A standard choice (see chu) is

align[align omitted — 186 chars of source]

Based on this choice, the FLCT yields that, for every $x$

align[align omitted — 224 chars of source]

where $B$ is a standard Brownian motion. The limiting law of this expression involves a Brownian motion; intuitively, this being a heteroskedastic process, this procedure may not be the most powerful one; this is further corroborated by aue2004, who show that the delay in detecting a changepoint increases as $\eta \rightarrow 0$. A possibility would be to re-scale the monitoring function as suggested in lajos04 and lajos07, viz. using

align[align omitted — 218 chars of source]

with $\eta \in \left[ 0,\frac{1}{2}\right] $, and $c_{\alpha ,m}$\ a critical value. Intuitively, the difference with ((ref)) is that the monitoring function is now smaller than before, which should ensure higher power. From a technical point of view, however, choosing $\eta >0$ entails having to use a different asymptotics, based on almost sure as opposed to weak convergence. The fact that the building blocks of $d\left( m;k\right) $ are the $\Theta _{m,R}^{\left( k\right) }$s - which are, conditional on the sample, i.i.d. and with finite moments - entails that that we can use an array of almost sure results (see the book by csorgo1997), which in turn makes it possible to carry out the monitoring using $\eta >0$ in ((ref) ).\\ We point out that, despite the considerations above, the choice of threshold functions is by no means unique, and, as chu put it, \textquotedblleft often dictated by mathematical convenience rather than optimality\textquotedblright ; in our case, we have defined $\nu ^{\ast }\left( m;k\right) $ as per ((ref)) give that the calculation of crossing probabilities (made according to ((ref))) is tractable - see lajos04 and lajos07. We then define the stopping rule as

equation[equation omitted — 149 chars of source]

setting $\widehat{k}_{m}=T_{m}$ when ((ref)) does not hold for $ 1\leq k\leq T_{m}$.

The critical value $c_{\alpha ,m}$ is defined, for a given level $\alpha $, as

align[align omitted — 318 chars of source]

in ((ref)), $\left\{ B\left( t\right) ,-\infty <t<\infty \right\} $ is a standard Brownian motion, whereas in ((ref)) we have defined

equation[equation omitted — 143 chars of source]

critical values for ((ref)) - which do not depend on $m$ - can be found in Table 1 in lajos04.

We need the following assumption, which restricts ((ref)) and strengthens Assumption (ref).

assumptionIt holds that: (i) \begin{equation*} m^{1/2+\tau }R\exp \left( -m^{\gamma }\right) \rightarrow 0, \end{equation*} as $\min \left( m,R\right) \rightarrow \infty $, for $\tau >0$; (ii) $ \int_{-\infty }^{+\infty }u^{4+\tau }dF\left( u\right) <\infty $, for some $ \tau >0$.
commentAssumption (ref) restricts ((ref)), by virtue of the presence of $m^{1/2+\tau }$, and therefore it constraints the choice of the size of the artificial sample, $R$.

It holds that

theoremWe assume that Assumptions (ref)-(ref) are satisfied. As $\min \left( m,R\right) \rightarrow \infty $ with ((ref)), under $H_{0}$ it holds that \begin{align} & P^{\ast }\left[ \max_{1\leq k\leq T_{m}}\frac{d\left( m;k\right) }{\nu ^{\ast }\left( m;k\right) }\leq x\right] \rightarrow P\left[ \sup_{0\leq t\leq 1}\frac{\left\vert B\left( t\right) \right\vert }{t^{\eta }}\leq x \right] for \eta <\frac{1}{2}, \\ & P^{\ast }\left[ \max_{1\leq k\leq T_{m}}\frac{d\left( m;k\right) }{\nu ^{\ast }\left( m;k\right) }\leq \frac{x+D_{m}}{A_{m}}\right] \rightarrow \exp \left( -\exp \left( -x\right) \right) for \eta =\frac{1}{2}, \end{align} for $-\infty <x<\infty $ and almost all realisations of $\left\{ u_{i},\epsilon _{i},1\leq i\leq T\right\} $. As $\min \left( m,R\right) \rightarrow \infty $, under $H_{A,1}\cup H_{A,2}$ it holds that \begin{equation} \max_{1\leq k\leq T_{m}}\frac{d\left( m;k\right) }{\nu ^{\ast }\left( m;k\right) }\overset{P^{\ast }}{\rightarrow }\infty , for any \eta \in \left[ 0,\frac{1}{2}\right] , \end{equation} for almost all realisations of $\left\{ u_{i},\epsilon _{i},1\leq i\leq T\right\} $.

Theorem (ref) implies the following

corollaryUnder the assumptions of Theorem (ref) it holds that: \begin{align} & \lim_{\min (m,R)\rightarrow \infty }P^{\ast }\left( \widehat{k} _{m}<T_{m}\right) \leq \alpha , \ \ under H_{0}, \\ & \lim_{\min (m,R)\rightarrow \infty }P^{\ast }\left( \widehat{k} _{m}<T_{m}\right) =1, \ \ under H_{A,1}\cup H_{A,2}, \end{align} for almost all realisations of $\left\{ u_{i},\epsilon _{i},1\leq i\leq T\right\} $.

Discussion and extensions

In this section, we investigate two aspects of the monitoring procedure proposed above. Firstly, we examine the impact of various test specifications on the power of our procedure (Section (ref)); secondly, we consider the presence of deterministics in ((ref)) (Section (ref)).

Power versus shrinking alternatives and the impact of $g\left( m;k\right)$

Our proposed methodology depends on several specifications in the construction of the monitoring function, and in the algorithm to compute the test statistic. In this section, we discuss the impact of such specifications on the power of the monitoring procedure. In particular, in this section we discuss the impact of $\gamma $ in the threshold function $ g\left( m;k\right) $ on power versus shrinking alternatives. In Section (ref), we also comment on the choices of $u$ and its distribution.

In order to discuss the impact of $\gamma $, we focus on a simple set-up where there are no deterministics, viz. on model ((ref))

equation*[equation* omitted — 59 chars of source]

and we consider the presence of power versus the local-to-null set-ups

eqnarray[eqnarray omitted — 289 chars of source]

In ((ref)), we assume

equation*[equation* omitted — 65 chars of source]

with $u_{i}^{\epsilon }$ independent of $u_{i}^{\upsilon }$. Specifically, in ((ref)), we consider a shrinking break where $\Delta _{\beta }\left( m\right) \rightarrow 0$, whereas in ((ref)), inspired by busetti, we model the local-to-null case as having $\sigma _{\upsilon }\left( m\right) \rightarrow 0$. Note that we consider, in both equations, the break as shrinking with $m$, since this can be viewed as the sample size on which estimation is based.

Heuristically, as the proof of Theorem (ref) shows, in order for the monitoring procedure to detect changes, it is necessary that $\psi _{m,k}\rightarrow \infty $ a.s.; thus, intuitively, $\Delta _{\beta }\left( m\right) $ and $\sigma _{\upsilon }\left( m\right) $, as they drift to zero, must be \textquotedblleft slow enough\textquotedblright\ to ensure that $ \psi _{m,k}$ diverges. We formalise this in the following theorem

theoremWe assume that Assumptions (ref)-(ref) are satisfied. Then, under $H^{\ast}_{A,1}$, equation ((ref)) holds as $ m\rightarrow \infty $ as long as \begin{eqnarray} m^{\delta -\varepsilon }\Delta _{\beta }\left( m\right) &\rightarrow &\infty , when \theta \leq 2, \\ m^{\theta \frac{\theta -2+\delta }{2\left( \theta -1\right) }}\Delta _{\beta }\left( m\right) &\rightarrow &\infty , when \theta >2, \end{eqnarray} for some $\varepsilon >0$. Under $H^{\ast}_{A,2}$, equation ((ref)) holds as $ m\rightarrow \infty $ as long as \begin{eqnarray} m^{\delta -\varepsilon }\sigma _{\upsilon }\left( m\right) &\rightarrow &\infty , when \theta \leq 2, \\ m^{\theta \frac{\theta -2+\delta }{2\left( \theta -1\right) }}\sigma _{\upsilon }\left( m\right) &\rightarrow &\infty , when \theta >2. \end{eqnarray}

Theorem (ref), together with ((ref)), illustrates what happens to the power of the monitoring procedure depending on the value of $\gamma $. As can be expected in light of the definition of $g\left( m;k\right) $, choosing a \textquotedblleft small\textquotedblright\ $\gamma$ (which corresponds to choosing $\delta$ close to $1$) enhances the power of the procedure, which, conversely, declines as $\gamma$ increases. This can be understood by noting that the noncentrality of $Q\left( m;k\right) $ is divided by $g\left( m;k\right) $ too. When $\theta \leq 2$, the procedure could potentially (depending on $\delta$) be able to detect breaks as small as $O\left( \frac{1}{m^{1-\epsilon}}\right)$, with $\epsilon > 0$ arbitrarily small. When $\theta > 2$ - that is, when monitoring goes on for a very long time - the ability to detect a small break increases. \\ Note that an alternative could have been to express the break as shrinking with the whole (calibration plus monitoring) sample size, $T$, as done in wied.

The monitoring procedure in the presence of deterministics

In this section, we consider the following extension of ((ref))

equation[equation omitted — 93 chars of source]

Equation ((ref)) contains, with respect to the previous model, a constant and a deterministic trend; other extensions could of course be possible. Our hypothesis testing framework is the same as in the previous section, namely we test for

equation*[equation* omitted — 153 chars of source]

versus the two alternatives

eqnarray*[eqnarray* omitted — 189 chars of source]

Note that, for the sake of a concise discussion, we do not consider changes in $\mu _{0}$ or $\mu _{1}$, although again this would be perfectly possible in principle.

Our monitoring scheme can be adapted as follows. As is typical in this case, we propose to demean and detrend both $y_{i}$ and $x_{i}$, by estimating

eqnarray*[eqnarray* omitted — 85 chars of source]

using OLS, and then computing

equation[equation omitted — 201 chars of source]

where $\widehat{u}_{i}^{x}$ and $\widehat{u}_{i}^{y}$ are the OLS residuals from the regressions above. After defining

equation*[equation* omitted — 86 chars of source]

for $m+1\leq i\leq T$, we use the recursively detrended residuals

equation[equation omitted — 150 chars of source]

where

equation*[equation* omitted — 320 chars of source]

Note that other detrending schemes could be proposed also; for example, in the construction of $\widehat{\epsilon }_{i}^{d}$, one could use non-recursive estimates $\widehat{\mu }_{0}$ and $\widehat{\mu }_{1}$ (computed once and for all using the sample $1\leq i\leq m$). \newline We then define, as in ((ref)), the cumulative process

equation[equation omitted — 192 chars of source]

where the long-run variance estimator $\widetilde{\sigma }_{\epsilon }^{2}$ is computed exactly as in ((ref)), using $\widehat{\epsilon } _{i}^{d}$. We need the following assumption, which complements Assumption (ref)(i).

assumptionIt holds that: (i) $E\left\Vert \sum_{i=1}^{t}i\epsilon _{i}\right\Vert ^{2}\leq c_{0}t^{3}$, for all $1\leq t\leq T$, and (ii) $ E\left\Vert \sum_{i=1}^{t}ix_{i}\right\Vert ^{2}\leq c_{0}t^{5}$, for all $ 1\leq t\leq T$.

The next theorem shows that, when using $Q^{d}\left( m;k\right) $ instead of $Q\left( m;k\right) $ in constructing the monitoring procedure, the same results hold.

theoremWe assume that Assumptions (ref)-(ref) are satisfied. Then, when constructing $d\left( m;k\right) $ using $Q^{d}\left( m;k\right) $, ((ref))-((ref)) hold.

Numerical and empirical evidence

In this section, we illustrate the properties of our procedure through a Monte Carlo exercise (Section (ref)), and through an application to US housing market data (Section (ref)).

Simulations

We consider the DGP in ((ref)), with the addition of a constant term, and with $p=1$, namely

equation*[equation* omitted — 99 chars of source]

to evaluate the finite sample performance of our proposed procedure. As discussed in section (ref), we can allow for a constant term in the DGP through recursive demeaning of the data. Incorporating a constant in this fashion allows us to directly compare our new procedure to the equivalent constant-only version of that proposed by wied. Noting that our demeaned procedure is mean-invariant, we set $\mu_{0}=0$. We set $\beta =1$ for $1\leq i\leq m$, although unreported experiments show that, as can be expected, this value has no impact on the results.

Innovations $\left\{ \epsilon _{i},u_{i}\right\} $ have been generated as follows

align[align omitted — 449 chars of source]

In ((ref)), we allow for $AR\left( 1\right) $ dynamics in $u_{i}$, setting $\rho ^{\left( x\right) }\in \left\{ 0,0.5\right\} $. We have also experimented with other values, noting that results hardly change. In order to control for the signal-to-noise ratio, we have generated the idiosyncratic innovation $v_{i}^{u}$ as i.i.d. $N\left( 0,\sigma _{u}^{2}\right) $; by ((ref)). This entails that the signal-to-noise ratio is exactly equal to $\sigma _{u}^{2}$, and we have used $\sigma _{u}^{2}=2$ in our experiments. In unreported simulations, we considered $ \sigma_{u}^{2}=1$, with qualitatively similar results. As far as ((ref) ) is concerned, we have generated $v_{i}^{e}$ as i.i.d. $N\left( 0,1\right) $. Serial dependence in the error term $\epsilon _{i}$ is explicitly allowed for through $\rho ^{\left( \epsilon \right) }$; note that when $\rho ^{\left( \epsilon \right) }=1$, this corresponds to $H_{A,2}$, that is, ((ref)) becomes a non-cointegrating regression. We report results for $\rho ^{\left( \epsilon \right) }\in \left\{ 0,0.5,0.9\right\} $. In ((ref)), we also consider the possible presence of endogeneity through the coefficient $\rho ^{\left( x\epsilon \right) }$, using $\rho ^{\left( x\epsilon \right) }\in \left\{ 0,0.5\right\} $. The long-run variance of $ \epsilon _{i}$\ is estimated as in ((ref)), setting $ H=\left\lfloor m^{1/6}\right\rfloor $, where $\left\lfloor \cdot \right\rfloor $ denotes the integer part.

As far as the other specifications of the experiment are concerned, we report results for $T\in \left\{ 100,200,400\right\} $ and $m\in \left\{ \frac{T}{4},\frac{T}{2}\right\} $. When considering the presence of a break, we have set the changepoint $k^{\ast }=m+\frac{T}{4}$. Experimenting with other breakdates does not change results in any remarkable way. Under $ H_{A,1}$, we have set $\beta _{i}=\beta +\Delta _{\beta }I\left[ i > k^{\ast }\right] $, with $\Delta _{\beta }\in \left\{ 0.5,1\right\} $. In addition to reporting empirical rejection frequencies under the alternative, we also report the delay in changepoint detection, defined as

equation[equation omitted — 81 chars of source]

We now turn to describing the implementation of the test and of the randomisation algorithm. As far as the former is concerned, we have computed $g\left( m;k\right) $ setting, according to ((ref)), $\gamma =0.45$. Results are similar, especially for large $m$, when using $\gamma =0.4$ and $ \gamma =0.5$.

commentThe choice of $\gamma $ directly affects the rate of convergence/divergence of $\widetilde{\psi }_{m,k}$, and it does have an important impact on the empirical rejection frequencies. Indeed, as can be expected, decreasing $ \gamma $ results in better power (but worse null rejection frequencies, with a mild tendency to over-reject in small samples - $T=100$ - when using $ \gamma =0.4$), and vice versa when $\gamma $ is increased to $\gamma =0.5$.

In the randomisation algorithm, based on ((ref)), we set $R=m$; altering this specification (which we have tried in some unreported experiments) is virtually inconsequential on the empirical rejection frequencies under both $H_{0}$ and $H_{A,1}\cup H_{A,2}$. Finally, we discuss the choice of $u$. Extracting $u$ from a distribution - as we make explicit in Step 4 of our algorithm - has the advantage that $u$ gets integrated out in the construction of $\Theta _{m,R}^{\left( k\right) }$, making this invariant to the support of $u$ itself. In this respect, choosing $F\left( u\right) $ as the standard normal distribution is a possibility, which is very easy to implement. Indeed, in order to construct $\Theta _{m,R}^{\left( k\right) }$ practically, we can use a Gauss-Hermite quadrature to approximate the integral that defines it, viz.

equation[equation omitted — 204 chars of source]

where the $z_{s}$s, $1\leq s\leq n_{S}$, are the zeros of the Hermite polynomial $H_{n_{S}}\left( z\right) $ and the weights $w_{s}$ are defined as

equation[equation omitted — 157 chars of source]

Thus, when constructing $\theta _{m,R}^{\left( k\right) }\left( u\right) $, we construct $n_{S}$ of these statistics, each with $u=\sqrt{2}z_{s}$; the values of the roots $z_{s}$, and of the corresponding weights $w_{s}$, are tabulated e.g. in salzer. in our case, we have used $n_{S}=2$, so that $u=\pm 1$ with equal weight $\frac{1}{2}$; we note that in unreported experiments we tried $n_{S}=4$ with the corresponding weights, but there were no changes up to the $4$-th decimal in the empirical rejection frequencies.

We report results for $\eta=\{0,0.45,0.49,0.50\}$ for the threshold function in ((ref)). We offer a direct comparison of our procedure to that of wied. We focus our attention on the IM-OLS version of their test in our simulations as the authors state a preference for IM-OLS, relative to the FM-OLS and D-OLS approaches that they also consider, on the basis of its finite sample performance. We denote this procedure $WW$--$IM$ in what follows. The nominal significance level $\alpha$, in the computation of critical values defined in ((ref)) and ((ref)), has been set as $\alpha =0.05$. Finally, all experiments have been carried out using $1,000$ replications.

Empirical rejection frequencies under $H_{0}$ are reported in Table (ref). We point out that, in our context, the notion of size (control) differs from the standard Neyman-Pearson testing paradigm. In the latter, empirical rejection frequencies are expected to be close to their nominal level. In the context of a sequential testing procedure like ours, as pointed out by lajos07 (see also the comments in Ch. 9 in Sen), the primary purpose is to keep the false detection rate below the chosen level $\alpha$. Indeed, the proportion of false discoveries should go to zero, since the monitoring can continue for an infinite amount of time. In this respect, whilst this is the case for all four values of $\eta$ considered in our procedure, setting $\eta=0$ yields the best results, with empirical rejection frequencies approaching zero across many settings of $m$ and $T$. Examining Panel A of the table, the case of no serial dependence in the error terms, it is clear that for $T=100$, $m=25$ all test procedures exhibit empirical rejection frequencies above their nominal significance levels, with the degree of distortion higher for our procedures than for the $WW$--$IM$ procedure. However, for $T=100$, $m=50$, whilst $WW$--$IM$ offer slightly inflated empirical rejection frequencies (between 0.052 and 0.080 depending on the value of $\rho ^{\left( x \epsilon \right) }$ and $\rho ^{\left( x \right) }$), our test procedure offers empirical rejection frequencies lower than the nominal significance level. For $T=200$ and $T=400$ we observe a similar pattern of results to that under $T=100$, $m=50$; note that the $WW$--$IM$ procedure exhibits a higher degree of upwards distortion (with frequencies up to 0.080 observed).

Allowing for AR(1) dynamics in $u_{i}$ has a small upwards effect on the empirical rejection frequencies of all tests for $T=100$ and $m=25$. For all other combinations of $m$ and $T$, little effect is observed for our test procedures, with no effect at all in the case of $T=400$, whereas the $WW$--$IM$ test is observed to be a little more sensitive to these AR(1) dynamics. Allowing for endogeneity results in modest increases in the empirical rejection frequencies for our procedures, for most settings of $m$ and $T$, whereas it has the opposite effect for the the $WW$--$IM$ test, resulting in modest decreases in size relative to the no endogeneity case.

Considering Panel B, serial dependence in $\epsilon _{i}$ of $\rho ^{\left( \epsilon \right) }=0.5$ results in upwards size distortion for $WW$--$IM$ for all settings of $m$, $T$, $\rho ^{\left( x \epsilon \right) }$ and $\rho ^{\left( x \right) }$. In contrast, with the exception of $T=100$, $m=25$, where all procedures exhibit empirical rejection frequencies somewhat higher than the nominal significance level, the size of our procedures are robust to this degree of serial correlation, with only very small differences observed in the empirical rejection frequencies from the no serial dependence case (no larger than 0.006 for the settings considered here). This is a pleasing result, given that serial dependence is likely in practice. Finally, turning our attention to Panel C, the case of high serial dependence, $\rho ^{\left( \epsilon \right) }=0.9$, we find that the $WW$--$IM$ procedure exhibits substantial over-sizing for all combinations of $m$, $T$, $\rho ^{\left( x \epsilon \right) }$ and $\rho ^{\left( x \right) }$. As before, with the exception of $T=100$, $m=25$, our procedures display less size distortion than those of $WW$--$IM$. When examining the performance of our procedures, we notice here the role that $m$ plays, with smaller empirical rejection frequencies observed for $m=\frac{T}{2}$ than for $m=\frac{T}{4}$, for a given $T$; and with empirical rejection frequencies decreasing as $T$ increases for a given setting of $m$.

Empirical rejection frequencies and the associated detection delays under $ H_{A,1}$ (i.e. under a change in $\beta$ in the cointegrating regression) are reported in Tables (ref) and (ref) respectively. Considering first the rejection frequencies in Table (ref), in the case of no serial dependence in the errors (Panel A), it is clear that our monitoring procedures offer excellent power, with $\eta=\{0.45,0.49,0.5\}$ outperforming $WW$--$IM$ over most combinations of $\Delta _{\beta }$, $T$, $m$, $\rho ^{\left( x \right) }$, and $\rho ^{\left( x \epsilon \right) }$. There are only 8 instances out of the 48 different combinations of settings in Panel A where $WW$--$IM$ achieves higher power, all cases where $m=\frac{T}{2}$. This result is somewhat anticipated, given that our test allows for $T_{m} \rightarrow \infty$, assumes that the monitoring horizon should go on for a sufficiently long time, and is thus expected to perform better for small $m$ relative to $T_{m}$, whereas wied choose $m$ to be large relative to $T$, as discussed in section (ref). It is pleasing however, that even in the case of $m=\frac{T}{2}$, our procedures outperform $WW$--$IM$ in terms of power in the majority of cases. Despite its small empirical rejection frequencies under the null, our procedure with $\eta=0$ also performs very well in terms of power, with rejection frequencies under $H_{A,1}$ very similar to those of $\eta=\{0.45,0.49,0.5\}$ in most cases. In addition to the effects of $\rho ^{\left( x \right) }$, and $\rho ^{\left( x \epsilon \right) }$, our procedures appear to be robust to serial dependence in the errors, with high levels of power maintained under different settings of $\rho ^{\left( \epsilon \right) }$. Comparing the four values of $\eta$ that we consider here, no setting uniformly outperforms the others, with little difference in rejection frequencies observed between these values.

Turning our attention to the detection delays reported in Table (ref), we observe that increasing $m$, increasing $T$, and increasing $\Delta _{\beta }$ all contribute towards reducing the detection delay, as we might expect. Contrary to the empirical rejection frequencies, when considering detection delays a ranking does emerge amongst the different values of $\eta$ for our procedure, with $\eta=0$ resulting in a longer detection delay relative to the other settings. This detection delay can be seen as a trade-off for the very small null rejection frequencies exhibited in Table (ref). Setting $ \eta=\{0.45,0.49,0.5\}$ produces the shortest detection delay across the various settings of $m$, $T$ and $\Delta _{\beta }$ considered here, with very little to distinguish between these settings. Our procedure is capable of detecting a break in the parameters of the cointegrating regression shortly after the break occurs, as little as 2.7 observations on average after the break for the case of $T=400$, $m=200$ and $\Delta _{\beta }=1$, where $\rho ^{\left(\epsilon \right)}=0$, $\rho ^{\left(x \epsilon \right)}=0$, $\rho ^{\left(x\right)}=0.5$ and $\eta=0.49$. The $WW$--$IM$ procedure incurs a longer detection delay than our test procedure for every setting considered here.

Finally, empirical rejection frequencies and detection delays under $H_{A,2}$ (i.e. under a switch from a cointegrating to a non-cointegrating regression) are given in Tables (ref) and (ref) respectively. Considering first the rejection frequencies in Table (ref), we note that our procedure is able to offer good levels of power against this alternative hypothesis for most settings. Relative to our results for $H_{A,1}$ in Table (ref), increasing $m$ has a more severe effect on the empirical rejection frequencies, particularly for smaller values of $T$. Examining Panel A, our procedure outperforms $WW$--$IM$ in the majority of cases. Exceptions occur in some instances where $m=\frac{T}{2}$, $\rho ^{\left(x \epsilon \right)}=0.5$.

Comparing Panels A with Panels B and C, it is clear that serial correlation in the errors has the effect of reducing the empirical rejection frequencies for all tests, with a higher degree of serial correlation corresponding to a lower rejection frequency. Of course, this result is to be anticipated given the nature of the alternative hypothesis. In general, with serial correlation of $\rho ^{\left(\epsilon \right)}\{=0.5,0.9\}$, our procedure performs better for $m=\frac{T}{4}$ and $WW$--$IM$ performs better for $m=\frac{T}{2}$, although we note that the empirical rejection frequencies reported here are not size-adjusted, and given the degree of over-sizing exhibited by especially $WW$--$IM$ in Table (ref), it is hard to directly compare the tests' performance.

Considering the detection delays under $H_{A,2}$ in Table (ref), we again observe that the delay decreases as $m$ and $T$ increase. We note that delay detections are generally longer under $H_{A,2}$ than for equivalent settings under $H_{A,1}$. As with the detection delays under $H_{A,1}$, when considering our procedure, setting $\eta=0$ provides the longest delay in detection, with $\eta=\{0.45,0.49,0.5\}$ providing the quickest detection of a break. $WW$--$IM$ exhibits a longer detection delay than our procedure with $\eta=\{0.45,0.49,0.5\}$ across all settings, except in one instance\footnote{$T=400$, $m=200$, $\rho ^{\left(\epsilon \right)}=0$, $\rho ^{\left(x \epsilon \right)}=0$ and $\rho ^{\left(x \right)}=0.5$ } where it is still outperformed by $\eta=\{0.45,0.5\}$.

When considering detection delays, it is possible that a test detects a break prematurely, which would lead to a negative delay for that replication according to (4.4), which in turn could result in a misleadingly low reported average delay in Tables (ref) and (ref). To further examine the estimated break dates found by these procedures, and to verify whether premature detection is of concern here, in Figures (ref) and (ref) we consider histograms of the estimated break dates found by our procedure using $\eta=0$ and $\eta=0.45$, as well as the $WW$--$IM$ procedure. For simplicity, we consider the case of $\rho ^{\left(x \right)}=0$, $\rho ^{\left(\epsilon \right)}=0$ and $\rho ^{\left(x \epsilon \right)}=0$. Figure (ref) displays estimated break dates under $H_{A,1}$, with (ref) considering $T=200$, $m=\frac{T}{4}$ and $\Delta_{\beta}=1$. It is clear that our procedure using either $\eta=0$ or $\eta=0.45$ provides more accurate break date estimation than $WW$--$IM$, with only a very small difference between these settings of $\eta$. A premature break date is found in only a handful of replications, in the case of $\eta=0.45$, suggesting that early detection is not a significant problem for our test. In Figure (ref) we set $T=400$, $m=\frac{T}{2}$ and $\Delta_{\beta}=0.5$, a more challenging circumstance for our procedure as it is designed for small $m$ relative to $T_{m}$. Nevertheless, our procedure displays more accuracy than that of $WW$-$IM$ here.

Figure (ref) displays estimated break dates under $H_{A,2}$, with (ref) and (ref) considering the same settings of $T$ and $m$ as in (ref) and (ref) respectively. Again, we are able to note the accuracy of our procedure relative to $WW$--$IM$, and the relatively small numbers of replications where a break is detected before the true break date, $k^{*}$.

Although $\eta=0$ provides the lowest null empirical rejection frequencies, we argue that a sequential monitoring test based on $\eta=0.45$ provides the best overall performance given that it maintains a null empirical rejection frequency below the nominal significance level across most settings of $m$ and $T$, as well as providing the shortest detection delays under both $ H_{A,1}$ and $H_{A,2}$ (although we note that there is very little difference in performance between $\eta=\{0.45,0.49,0.5\}$).

Empirical application

To demonstrate the practical relevance of the procedure developed in section (ref), and inspired by the empirical work of anundsen and wied, we investigate the possibility that the US housing market experienced a structural break in cointegration. Based on the life-cycle model of housing under the assumption of no arbitrage for the housing market, anundsen analyses two fundamentals-driven cointegrating relationships. The first approach, known as the price-to-rent model, relies on the user cost of a property being equal to the cost of renting a property of similar quality in equilibrium, and is given by:

equation[equation omitted — 87 chars of source]

where $ph_{t}$ is the logarithm of real housing prices at period $t$, $r_{t}$ is the logarithm of real rents, and $UC_{t}$ is the real direct user cost of housing, computed as

equation*[equation* omitted — 86 chars of source]

where $\tau^{y}_{t}$ is the marginal personal income tax rate (measured here at twice the median income), $\tau^{p}_{t}$ is the marginal tax rate on personal property, $i_{t}$ is the nominal interest rate, $\pi_{t}$ is overall price inflation, and $\delta_{t}$ is the housing depreciation rate.

The second approach, known as the inverted demand model, assumes that imputed rent is a function of income and housing stock, and is given by the below equation:

equation[equation omitted — 152 chars of source]

where $y_{t}$ is the logarithm of real per capita disposable income and $ h_{t}$ is the logarithm of the per capita housing stock.

Assuming that the variables in ((ref)) and ((ref)) are $ I(1)$, economic theory predicts that $u_{t}$ and $\widetilde{u}_{t}$ are both $I(0)$. That is, two cointegrating relationships exist between housing prices and their fundamentals. A breakdown of these cointegrating regressions therefore indicates that housing prices are no longer being driven by these fundamentals. Following the definition of stiglitz, inter alia, that an asset bubble exists when its price no longer appears to be justified by the value of its fundamental components, anundsen interprets a breakdown in these cointegrating relationships as evidence of a bubble in housing prices. Indeed, following anundsen, we also allow for an intercept and linear trend term in each model, viz.

equation[equation omitted — 103 chars of source]

and

equation[equation omitted — 194 chars of source]

anundsen applies models ((ref))-((ref)) to quarterly US housing market data over the sample period 1976:Q1 - 2010:Q4. Specifically, he estimates vector autoregression models and undertakes Johansen cointegration testing for both the price-to-rent and inverted demand equations using expanding sub-samples of the data, starting with an initial sub-sample from 1976:Q1 - 1995:Q4 and then subsequently adding four new observations until the full sample is used. This analysis finds evidence of a cointegrating relationship in the housing market up until 2002 in the price-to-rent model; evidence in favour of cointegration disappears when 2002:Q4 is included in the sample, but there is evidence of a return to a cointegrating relationship towards the end of the sample. As far as the inverted demand model is concerned, a similar pattern is found, with cointegration breaking down in 2001:Q4. These results imply the emergence of bubble behaviour in the housing market beginning in 2001-2002; however, wied highlight that the analysis suffers from the problem of multiple testing, leading to uncontrolled size. Considering the same dataset, they apply their real time monitoring procedure to models ((ref))-((ref)), with $WW$--$IM$ detecting a breakdown in cointegration at 2006:Q4 for the price-to-rent model and 2004:Q2 for the inverted demand model (with the FM-OLS version of the procedure finding a slightly earlier break of 2003:Q2 for this model). This delay in detection, relative to the results of anundsen, can be viewed as a trade-off for asymptotic validity, and therefore controlled size.

We apply our sequential monitoring procedure to the dataset discussed above, containing information on US house prices from 1976:Q1 - 2010:Q4.\footnote{Detailed information on the dataset sources and construction are contained in anundsen. The dataset has been downloaded from the Journal of Applied Econometrics data archive.} In line with the previous two studies, our calibration sample runs from 1976:Q1 - 1995:Q4, such that $m=80$; effectively, this means that \textquotedblleft future data\textquotedblright (and our monitoring) starts in 1996:Q1 (this being the earliest possible break date that we can detect). We set $R=m$ and $n_{s}=2$ as before. We allow for a constant and linear trend in both models through the recursive demeaning and detrending method discussed in section (ref). In view of our simulation results in the previous section, we have used $\eta=0.45$. \footnote{ Although, in unreported results, we note that setting $\eta=\{0,0.49,0.50\}$ provides the same break date estimate as $\eta=0.45$ for both models.} Similarly, based on the Monte Carlo evidence, we set $\gamma=0.4$, which should ensure size control, while decreasing the detection delay.

Figure (ref) displays the residuals obtained from our estimation of the price to rent and inverted demand models, using recursive demeaning and detrending. From visual inspection, it is clear that the residuals of both models undergo a period of mean-reverting behaviour in the earlier part of the sample, whereas more persistent behaviour in the residuals is observed from the early 2000s, lending support to the hypothesis that a structural break in cointegration occurs during the sample period. Considering first the price-to-rent model in Figure (ref), our sequential monitoring procedure finds evidence of a break in cointegration in 2005:Q3, 5 quarters earlier than the $WW$--$IM$ test, although still somewhat later than the detection date in the initial experiment of anundsen. Examining the inverted demand model, in Figure (ref), evidence is found of a break in cointegration in 2004:Q1, somewhat earlier than for the price-to-rent model, in line with the results of wied and anundsen. Thus, our results support the claim of a breakdown in fundamentals-driven cointegrating relationships in the US housing markets during the housing bubble of the 2000s.

Conclusions

In this paper we have investigated the issue of monitoring a cointegrating regression. Having stability as the null hypothesis, we develop a procedure to detect changes in the regression coefficients and/or from cointegration to non-cointegration. Our procedure is based on using the cumulative sums of squared residuals; at each point in the monitoring horizon, we randomise the cumulative sum process, thereby obtaining an i.i.d. sequence with finite moments of arbitrarily high order. We then use the results in lajos04 and lajos07 to construct a family of procedures which may be viewed as a complement to the results in wied.

We point out that, as well as deriving the aforementioned statistics, in this paper we have proposed a general methodology to construct monitoring schemes in the context of a cointegrating regression. The approach we propose can be readily generalised to use other statistics (e.g., upon calculating the relevant rates, even the KPSS type statistic employed in wied could be randomised and used in our algorithm), or to other hypothesis testing frameworks. As a leading example, wagner consider the very interesting case where ((ref)) is, to begin with, a non-cointegrating regression with $\epsilon _{i}\sim I\left( 1\right) $, and the purpose of monitoring is to verify whether ((ref)) becomes a cointegrating regression, with $\epsilon _{i}\sim I\left( 0\right) $. Although we leave this interesting research question for future study, we point out that a monitoring scheme for this case could be readily developed. Indeed, one could use exactly the same approach as we do, using

equation*[equation* omitted — 73 chars of source]

instead of ((ref)). This and others issues are under investigation by the authors.

comment\section{Data availability statement} The dataset employed in Section (ref) may be found online in the supporting information tab for this article (file name: Anundsendata.xslx). \section{Supporting information} Additional supporting information (technical lemmas and the proof of Proposition 1) is contained in the Supplement (TrapaniWhitehouseSupplement.pdf), which can be found online in the supporting information tab for this article.

{ {{.2cm} }}