The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
83,682 characters
Sequential monitoring for cointegrating regressions
\begin{frontmatter}
\title{Sequential monitoring for cointegrating regressions}
\runtitle{Monitoring cointegration}
\begin{aug}
\author{\fnms{Lorenzo} \snm{Trapani}$^{\dag}$\ead[label=e1]{[email removed]}}\hskip .2cm
\author{\fnms{Emily} \snm{Whitehouse}$^{*}$\ead[label=e2]{[email removed]}}
\runauthor{L. Trapani and E. Whitehouse}
\affiliation{University of Nottingham\thanksmark{m1} and Newcastle University \thanksmark{m2}}
\address{$^{\dag}$University of Nottingham\\
\printead{e1}\\
}
\address{$^{*}$Newcastle University\\
\printead{e2}\\
}
\end{aug}
\begin{abstract}
We develop monitoring procedures for cointegrating regressions, testing the null of no breaks against the alternatives that there is either a change in the slope, or a change to non-cointegration. After observing the regression for a calibration sample $m$, we study a CUSUM-type statistic to detect the presence of change during a monitoring horizon $m+1,...,T$. Our procedures use a class of boundary functions which depend on a parameter, $0 \leq \eta \leq \frac{1}{2}$, whose value affects the delay in detecting the possible break. Technically, these procedures are based on almost sure limiting theorems whose derivation is not straightforward. We therefore define a monitoring function which - at every point in time - diverges to infinity under the null, and drifts to zero under alternatives. We cast this sequence in a randomised procedure to construct an \textit{i.i.d.} sequence, which we then employ to define the detector function. Our monitoring procedure rejects the null of no break (when correct) with a small probability, whilst it rejects with probability one over the monitoring horizon in the presence of breaks.
\end{abstract}
\begin{keyword}[class=MSC]
\kwd{62F03}
\kwd{62L10}
\kwd{62M10}
\end{keyword}
\begin{keyword
}
\kwd{cointegration}
\kwd{structural change}
\kwd{sequential monitoring}
\kwd{randomized tests}
\end{keyword}
\end{frontmatter}
\newpage
\section{Introduction\label{introduction}}
In this paper, we study the following cointegrating regression
\begin{equation}
y_{i}=\beta ^{\prime }x_{i}+\epsilon _{i}\text{, }1\leq i\leq T,
\label{model-1}
\end{equation}
where $\left( y_{i},x_{i}^{\prime }\right) ^{\prime }$ is a $(p+1)\times 1$,
$I\left( 1\right) $ vector and $\epsilon _{i}\ $is a stationary innovation.
In particular, we investigate the issue of monitoring (\ref{model-1}), after
a calibration period of length $m$, during which our maintained assumptions are
that \textit{(i)} (\ref{model-1}) is a cointegrating relationship and
\textit{(ii)} the slope $\beta $ is constant. From $i=m+1$ onwards, we check
whether the relationship in (\ref{model-1}) remains constant, or whether
either the slope $\beta $ changes, or (\ref{model-1}) becomes a
non-cointegrating regression (or both).
The timely detection of structural change is arguably of great importance in
the context of any regression model: whilst there is an extensive literature
on the general topic of on-line detection of changes (see e.g.
\citealp{csorgo1997} for a survey), in the econometrics literature this
issue has received some limited attention since the contribution by
\citet{chu}. Recent articles that study this topic have focused on linear
regression models (\citealp{lajos04}, \citealp{aue2006}, \citealp{lajos07},
\citealp{kap-monitor}), large factor models (\citealp{bt1}), and also
cointegrating regressions (\citealp{steland}, \citealp{wied},
\citealp{wagner}). In particular, \citet{wied} consider, essentially, the
same problem as in our paper; namely, they propose several statistics for
the on-line detection of structural breaks in a model like (\ref{model-1}),
considering both the possibility of a change in the slope $\beta$ and a
change to a non-cointegrating regression.
From a methodological viewpoint, we use a residual-based detector to test
for the null hypothesis of no change over the monitoring horizon $m+1\leq
i\leq T$. Note that this corresponds to a closed-end procedure (\citealp{aue12}), since monitoring - as can be expected to happen in practice - stops after $T$.
The family of detectors
which we propose here are based on the sum of squared residuals. Simulations show that our monitoring scheme
has excellent finite sample properties, with low occurrence of false detections and very good
power versus both alternatives under consideration. Other detectors are also
possible (see, for example, the various statistics considered in
\citealp{homm}, albeit in a different context).
From a technical point of view, as pointed out by \citet{lajos04} and \citet{lajos07}, the detectors employed in monitoring procedures depend upon a parameter,
henceforth denoted as $\eta$, which can vary in the interval $\left[ 0,\frac{
1}{2}\right] $. Constructing test statistics when $\eta =0$ (see e.g. \citealp{chu}) requires, as a
technical tool, weak convergence, and therefore one can employ a huge
variety of results which are well-known in the literature (see e.g. the book
by \citealp{billingsley}). On the other hand, the choice $\eta =0$ is known
to often yield inferior results, in particular resulting in a longer delay
in detection of a break (\citealp{aue2004}). In order to overcome this issue, it is usually
recommended to choose $\eta >0$ (\citealp{lajos07}). However, from a
technical viewpoint, using $\eta >0$ requires having stronger forms of
convergence than weak convergence, with fewer results available (we refer to
the textbook by \citealp{csorgo1997} for an excellent treatment of the
subject). For example, to the best of our knowledge we
are not aware of strong approximations like the ones derived in \citet{kmt1}
and \citet{kmt2} for convergence to stochastic integrals, where usually
\textquotedblleft weak\textquotedblright\ results are used instead (see
\citealp{chanwei}; and \citealp{phillipsweak}). In light of this, we only
rely on (almost sure) rates, and we develop a family of statistics -
computed at each $m+1\leq i\leq T$ - which diverge to positive infinity
under the null of no break, whilst they drift to zero in the presence of
breaks. We then randomize such statistics at each point in time $i $: the
outcome of our randomisation is a sequence of random variables which, under
the null of no break, are \textit{i.i.d.} with finite moments up to any
order, whilst they diverge to infinity in the presence of a break. Finally,
we employ the newly generated sequence in order to construct the same
detectors as in \citet{lajos04} and \citet{lajos07}, being able to rely on
the theory spelt out in those papers. Using randomisation is helpful when
the properties of a certain statistic are not known, or depend on nuisance
parameters: in this respect, it might be envisaged that randomisation serves
a similar purpose to the bootstrap or to self-normalisations (see \citealp{dette2019likelihood} for an example of self-normalisation in the context of monitoring). In the econometric literature,
randomisation has been employed in a wide variety of contexts, including
testing for forecasting superiority (\citealp{corradi2006}), stationarity (
\citealp{bandi2014}), finiteness of moments (\citealp{trapani16}), boundary
problems (\citealp{HT16}) and determining the number of common factors in a
large factor models (\citealp{trapani17}). In our context, however, we do
not employ randomisation to produce a randomised test, but to construct a
\textquotedblleft well-behaved\textquotedblright\ sequence which, in turn,
can be employed to define an easy-to-study test statistic. In this respect,
our contribution uses the same approach as \citet{bt1}, who study monitoring
for structural change in the context of a large, stationary factor models.
By relying solely on rates, we require quite mild assumptions; all the theory can be based on using a
standard OLS\ estimator, with no need for more specialised estimators like,
say, the FM-OLS estimator (\citealp{phillips-hansen}) or a Dynamic OLS
estimator (\citealp{saikkonen}); and, finally, we do not need to rely on the accuracy of the long-run
variance estimator.
The remainder of the paper is organised as follows. In Section \ref{theory},
we provide the relevant assumptions, and then report theoretical results on
estimation and the monitoring procedure. Extensions to e.g. the case of deterministics are in Section \ref{extensions}. In Section \ref{montecarlo} we
demonstrate the performance of our monitoring procedure through both a Monte
Carlo simulation exercise (Section \ref{simulations}) and an empirical
application to US housing market data (Section \ref{empirics}). Section \ref
{conclusions} concludes. Proofs of the main results are in Section \ref{proofs}. All technical lemmas and some proofs are relegated to
the Supplement.
Throughout the paper we use the notation $c_{0}$, $c_{1}$,... to denote
positive and finite constants, that do not depend on the sample size; their
value is allowed to change from line to line. We use the expression
\textquotedblleft a.s.\textquotedblright\ as short-hand for
\textquotedblleft almost surely\textquotedblright ; the ordinary limit is
denoted as \textquotedblleft $\rightarrow $\textquotedblright . Finally, for
a vector $a $ and a matrix $A$, $\left\Vert a\right\Vert $ and $\left\Vert
A\right\Vert $ represent the Euclidean norm. Other notation is introduced
later on in the paper.
\section{Theory\label{theory}}
We begin with introducing some notation and the main assumptions that should
hold under the null of no break (Section \ref{a-h0}); we then move to
discuss the two alternative hypotheses which we consider, namely a change in
the slope and/or a change to a non-cointegrating equation (Section \ref
{monitor}). Finally, in Section \ref{sequential}, we discuss the relevant
CUSUM process, and the randomisation algorithm.
\subsection{Main assumptions\label{a-h0}}
Recall (\ref{model-1})
\begin{equation*}
y_{i}=\beta ^{\prime }x_{i}+\epsilon _{i},
\end{equation*}
which we assume to be valid during the calibration period $1 \leq i \leq m$, with
\begin{equation}
x_{i}=x_{i-1}+u_{i}. \label{model-x}
\end{equation}
We also define the long run variances of $u_{i}$ and $\epsilon _{i}$ as
\begin{align}
\Sigma _{u}& =\lim_{m\rightarrow \infty }E\left( \frac{1}{\sqrt{m}}
\sum_{i=1}^{m}u_{i}\right)\left( \frac{1}{\sqrt{m}}\sum_{i=1}^{m}u_{i}
\right)^{\prime }, \label{sigma-u} \\
\sigma _{\epsilon }^{2}& =\lim_{m\rightarrow \infty }Var\left( \frac{1}{
\sqrt{m}}\sum_{i=1}^{m}\epsilon _{i}\right) . \label{sigma-e}
\end{align}
We consider the following assumption.
\begin{assumption}
\label{as-2}It holds that: (i) $\epsilon _{i}$ and $u_{i}$ have mean zero
with (a) $E\left\vert \epsilon _{i}\right\vert ^{2}<\infty $ for $1\leq
i\leq T$, and $0<\sigma _{\epsilon }^{2}<\infty $; and (b) $\Sigma _{u}$ is
positive definite with $\left\Vert \Sigma _{u}\right\Vert $; (ii) $
E\left\Vert x_{0}\right\Vert ^{2}<\infty $ and
\begin{equation}
\sup_{1\leq i\leq t}\left\Vert x_{i}-W_{x}\left( i\right) \right\Vert
=O_{a.s.}\left( t^{1/2-\delta ^{\prime }}\right) , \label{sip-x}
\end{equation}
for some $0<\delta ^{\prime }<\frac{1}{2}$, where $W_{x}\left( i\right) $ is
a $p$-dimensional Wiener process with increments of variance $\Sigma _{u}$;
(iii) $E\left\Vert \sum_{i=1}^{t}x_{i}\epsilon _{i}\right\Vert ^{2}\leq
c_{0}t^{2}$, for all $1\leq t\leq T$; (iv) $E\left\Vert
\sum_{i=1}^{t}x_{i}x_{i}^{\prime }\right\Vert \leq c_{0}t^{2}$, for all $
1\leq t\leq T$.
\end{assumption}
Assumption \ref{as-2}\textit{(i)} is a standard second moment condition
which is required to hold under the null of no change, and also when the
slope $\beta $ changes but (\ref{model-1}) remains a cointegrating
relationship. Note that, by part \textit{(i)}(b), we rule out cointegration among the regressors. Part \textit{(ii)} of the assumption, in essence, states that
a strong approximation exists for the partial sums process $x_{i}$. This is
a high-level assumption, which could be replaced by more primitive
requirements on the existence of moments for the innovation $u_{i}$, and
some form of weak dependence. It can be envisaged, as far as moments are
concerned, that at least $E\left\Vert u_{i}\right\Vert ^{2+\delta }<\infty $
is required for some $\delta >0$; thence, (\ref{sip-x}) would follow
immediately if $u_{i}$ is \textit{i.i.d.} (see \citealp{kmt1} and
\citet{kmt2} for the univariate case, and \citealp{gotze} for the
multidimensional one), and also under fairly general forms of weak
dependence such as the case of stationary causal processes including linear
models, Volterra series and models with conditional heteroskedasticity (see
\citealp{wu2005}, and \citealp{berkesliuwu}). Interestingly, in the
literature it is relatively common to assume a weak Invariance Principle to
hold in lieu of assuming weak dependence (see e.g. Assumption 2 in
\citealp{wied}). Part \textit{(ii)} of Assumption \ref{as-2} serves exactly
the same purpose, except for the fact that in our paper we need almost sure
rates. Parts \textit{(iii)} and \textit{(iv)} could also be derived under
more primitive conditions on moments, serial dependence, and possible
correlation between $u_{i}$ and $\epsilon _{i}$. For example, the results
could be shown by standard arguments in the case of $u_{i}$ and $\epsilon
_{i}$ being \textit{i.i.d.} and independent of each other; in this case,
existence of second moments would suffice. Part \textit{(iv)} can be shown
under more general forms of dependence, e.g. in the case of linear processes
by exploiting the results in \citet{solo}. Also, in part \textit{(iii)}, the
requirement of independence between $u_{i}$ and $\epsilon _{i}$ is not
necessary: again under the assumption of linear processes, for example, it
could be shown (see, \textit{inter alia}, \citealp{durlauf}, and
\citealp{phillips-hansen}) that this part of the assumption can hold also in
the presence of endogeneity.
\subsection{Hypotheses of interest and the construction of the monitoring
procedure\label{monitor}}
We base our on-line monitoring on the theory developed in \citet{lajos04}
and \citet{lajos07}. We assume that the data are collected for an initial
calibration period of size $m$ where no break occurs; this can be viewed as
the historic sample available to the researcher. We then define the (length
of) the monitoring horizon $T_{m}$ as $T_{m}=T-m$. Thus, if $T$ represents
the total period considered, $m$ is the amount of time elapsed until the
beginning of the monitoring period. In essence, $m$ is going to be the
sample size used by the researcher for estimation. Choosing $T_{m}$ - that
is, choosing where to stop the monitoring - is an important issue in
sequential analysis, since it can be argued that monitoring comes at a cost
(see the original paper by \citealp{wald}); in this paper, we allow for $
T_{m}\rightarrow \infty $, under the assumption that monitoring is costless
- this assumption can be realistic when analysing economic series, although
not in other contexts (see e.g. the comments in \citealp{chu}).
\subsubsection{Alternative hypotheses of interest\label{subsub}}
Under the null hypothesis of our monitoring scheme, (\ref{model-1}) is a
cointegrating relationship for the whole monitoring horizon, and the slope $
\beta $ does not change; rewriting (\ref{model-1}) as
\begin{equation*}
y_{i}=\beta_{i} ^{\prime }x_{i}+\epsilon _{i},
\end{equation*}
we have
\begin{equation}
H_{0}:\left\{
\begin{array}{l}
\beta _{i}=\beta \\
\epsilon _{i}\text{ is stationary}
\end{array}
\right. \text{ for }1\leq i\leq T_{m}. \label{null}
\end{equation}
Conversely, when the null does not hold, there could be at least two
interesting, non mutually exclusive alternatives. In the first case, there
could be a structural change whereby, after $i=m$, $\beta $ changes:
\begin{equation}
H_{A,1}:\beta _{i}=\beta +\Delta _{\beta }I\left[ i>k^{\ast }\right] .
\label{break}
\end{equation}
In (\ref{break}), $m\leq k^{\ast }<T$ is the potential breakdate. In
addition to this (or as an alternative), (\ref{model-1}) may switch to being
a non-cointegrating relationship at some point in time, viz.
\begin{equation}
H_{A,2}:\epsilon _{i}=\epsilon _{i-1}+u_{i}^{\epsilon }\text{ for }k^{\ast
}+1\leq i\leq T. \label{spurious}
\end{equation}
In both cases, the case of no break is represented by having $k^{\ast }=T$.
As a general comment to our hypothesis testing framework, we point out that
the set-up in (\ref{null})-(\ref{spurious}) mirrors the analysis in
\citet{wied} very closely. In particular, the null hypothesis is the
intersection of two (very different) requirements: \textit{(a)} the fact
that there is no time variation in the structural parameter $\beta $ in (\ref
{model-1}) over the monitoring horizon, under the implicitly maintained
hypothesis that (\ref{model-1}) is always a cointegrating regression; and
\textit{(b)} the fact that (\ref{model-1}) is indeed a cointegrating
relationship during the monitoring horizon. This could be the the set-up of
interest in various applications (see e.g. Section \ref{empirics});
furthermore, an \textquotedblleft omnibus\textquotedblright\ procedure which
is powerful versus a global alternative could be viewed as advantageous in
order to avoid having to test under a maintained hypothesis whose validity
may not always be assumed. On the other hand, the monitoring procedure
proposed in this paper (and in \citealp{wied}) can be argued to be
\textquotedblleft non-constructive\textquotedblright : after rejecting the
null and finding evidence of a change in the nature of (\ref{model-1}), it
is not clear which of the two alternatives the change can be ascribed to. In
the literature, there are tests for more focussed alternatives which could,
in principle, be extended into monitoring procedures. For example, under the
maintained assumption that $\epsilon _{i}$ is stationary over the whole
monitoring horizon, one could think of extending the test for breaks in
cointegrating regressions proposed by \citet{kejriwal}. Similarly, under the
maintained assumption that $\beta $ is constant for the whole interval $
1\leq i\leq T$, a monitoring procedure could, in principle, be constructed
using the residuals $\widehat{\epsilon }_{i}$, e.g. by extending the test
for a change in persistence proposed by \citet{busetti}. Indeed, under the
same maintained hypotheses mentioned above, our procedure can also be
employed to test, separately, versus the two alternatives mentioned above.
In this respect, the monitoring scheme proposed in this paper could be
viewed as a preliminary step: upon finding evidence that a change occurred,
the researcher may decide to use a more specialised procedure to disentangle
the nature of the change in (\ref{model-1}).
\bigskip
In order to analyse the case of (\ref{spurious}), we need the following
assumption which characterizes the behaviour of $\epsilon _{i}$ under $
H_{A,2}$.
\begin{assumption}
\label{ha-2}Under $H_{A,2}$, it holds that (i)
\begin{equation}
\sup_{k^{\ast }+1\leq i\leq t}\left\vert \epsilon _{i}-W_{\epsilon }\left(
i\right) \right\vert =O_{a.s.}\left( t^{1/2-\delta ^{\prime \prime }}\right)
, \label{sip-2}
\end{equation}
for all $k^{\ast }+1\leq t\leq T$ and some $0<\delta ^{\prime \prime }<\frac{
1}{2}$, where $W_{\epsilon }\left( i\right) $ is a Wiener process with
increments of positive variance equal to the long-run variance of $
u_{i}^{\epsilon }$; (ii)
\begin{equation*}
E\left\vert \sum_{i=k^{\ast }+1}^{t}\epsilon _{i}^{2}\right\vert \leq
c_{0}t^{2},
\end{equation*}
for all $k^{\ast }+1\leq t\leq T$.
\end{assumption}
Assumption \ref{ha-2} supersedes parts \textit{(i)} and \textit{(iii)} of
Assumption \ref{as-2} in order to accommodate for the presence of a switch
to a non-cointegrating regression. According to the assumption, in essence,
after the breakdate $k^{\ast }$ the innovation $\epsilon _{i}$ becomes a
unit root process.
\subsubsection{The monitoring function\label{monitor-function}}
Our monitoring scheme is based on a non-recursive estimator of $\beta $:
estimation is carried out using the sample $1\leq i\leq m$ once and for all,
without updating the estimate as $i$ elapses. We focus only on this merely
for the sake of a concise discussion: this choice is not the only possible
one. \citet{lajos04}, \textit{inter alia}, propose a recursive monitoring
procedure (as well as a non-recursive one), where $\beta $ is estimated at
each $i$ using an expanding sample. It seems reasonable to conjecture that,
even in our context, the non-recursive scheme is probably likely to be less
affected by outliers, thus ensuring a better size control, whilst the
recursive procedure should be, by design, more sensitive to breaks.
Let
\begin{equation}
\widehat{\beta }_{m}=\left( \sum_{i=1}^{m}x_{i}x_{i}^{\prime }\right)
^{-1}\sum_{i=1}^{m}x_{i}y_{i}, \label{beta}
\end{equation}
where dependence on the sample size $m$ will be omitted whenever possible,
and define the residuals
\begin{equation}
\widehat{\epsilon }_{i}=y_{i}-\widehat{\beta }_{m}^{\prime }x_{i}=\epsilon
_{i}+\left( \beta -\widehat{\beta }_{m}\right) ^{\prime }x_{i},
\label{residual}
\end{equation}
for $m+1\leq i\leq T$ onwards. At each $k$, we define the cumulative process
\begin{equation}
Q\left( m;k\right) =\left\vert \frac{1}{\widehat{\sigma }_{\epsilon }^{2}}
\sum_{i=m+1}^{m+k}\widehat{\epsilon }_{i}^{2}\right\vert , \label{q-2}
\end{equation}
for $1\leq k\leq T_{m}$.
\begin{comment}
We begin by providing some heuristic arguments to motivate our choice. Our
approach is based on the cumulative sums of some transformation of the
residuals, as is natural in these contexts (see e.g. \citealp{wied},
\citealp{lajos04}) - although, somewhat differently from \citet{wied}, we
found it easier to study cumulative processes involving squares of
residuals. The main idea, however, is that, in the presence of a break, such
cumulative sums would grow much larger, over time, than their
\textquotedblleft natural\textquotedblright\ growth rate in the absence of
breaks. Based on well-known properties of unit root processes (see e.g.
\citealp{durlauf}), it can be expected that the fluctuations of the
monitoring functions should grow linearly over time if there is no break; if
there is a break, some \textquotedblleft non-centrality\textquotedblright\
term should arise which makes the fluctuations grow at a rate faster than
linear. Upon expanding the summands $x_{i}\widehat{\epsilon }_{i}$ and $
\widehat{\epsilon }_{i}^{2}$, it can be seen that the \textquotedblleft
non-centrality\textquotedblright\ of $Q_{1}\left( m;k\right) $ is given by $
\Delta _{\beta }x_{i}^{2}$ under $H_{A,1}$: the cumulative sums of this term
should diverge quadratically over time, and such divergence should be
strong, in that it can be expected that
\begin{equation}
\lim \inf_{t\rightarrow \infty }\frac{\ln \ln t}{t^{2}}
\sum_{i=1}^{t}x_{i}^{2}>0\text{ a.s.} \label{lim-inf-1}
\end{equation}
This would ensure the ability to detect a break in $\beta $. On the other
hand, under $H_{A,2}$, the \textquotedblleft
non-centrality\textquotedblright\ of $Q_{1}\left( m;k\right) $ is given by $
x_{i}\epsilon _{i}$. Whilst it is true that the partial sums of $
x_{i}\epsilon _{i}$ grow at a quadratic rate, there is no guarantee that
this growth is as strong as in (\ref{lim-inf-1}). Indeed, one can expect
\footnote{
This can be understood even better by considering the zero-mean, continuous
process $\Gamma \left( u\right) =W_{x}\left( u\right) W_{\epsilon }\left(
u\right) $ for which $\Gamma \left( 0\right) =0$. Let
\begin{equation*}
\Delta \left( r\right) =\int_{0}^{r}\Gamma \left( u\right) du.
\end{equation*}
If $\Delta \left( r\right) $ has no zeros on the interval $\left( 0,a\right]
$, then $E\left[ \Delta \left( r\right) \right] \neq 0$ on $\left( 0,a\right]
$, which contradicts the fact that $E\left[ \Delta \left( r\right) \right]
=0 $ due to $E\Gamma \left( u\right) =0$ for all $u$.}
\begin{equation}
\lim \inf_{t\rightarrow \infty }\sum_{i=1}^{t}x_{i}\epsilon _{i}=0\text{ a.s.
} \label{lim-inf-2}
\end{equation}
We conjecture that the following Chung-type LIL (\citealp{chung}) should hold
\begin{equation}
\lim \inf_{t\rightarrow \infty }\frac{\ln \ln t}{t^{2}}\sup_{1\leq j\leq
t}\left\vert \sum_{i=1}^{j}x_{i}\epsilon _{i}\right\vert >0\text{ a.s.,}
\label{lim-inf-3}
\end{equation}
but a result like (\ref{lim-inf-3}) is obviously much weaker, as far as the
divergence of the monitoring function is concerned, than that of (\ref
{lim-inf-1}). In light of these considerations, a procedure based on $
Q_{1}\left( m;k\right) $ should have power versus $H_{A,1}$, and may have
some, limited power, under $H_{A,2}$.
Conversely, the monitoring function $Q_{2}\left( m;k\right) $ is designed to
ensure power versus $H_{A,2}$: its \textquotedblleft
non-centrality\textquotedblright\ is driven by the cumulative sums of $
\epsilon _{i}^{2}$ for which, under $H_{A,2}$, a result like (\ref{lim-inf-1}
) holds. Under $H_{A,1}$, however, it can be seen that the \textquotedblleft
non-centrality\textquotedblright\ term is driven by
\end{comment}
\subsubsection{Estimation of $\protect\sigma _{\protect\epsilon }^{2}$\label
{lr}}
In (\ref{q-2}), $\widehat{\sigma }_{\epsilon }^{2}$ is an estimator of $
\sigma _{\epsilon }^{2}$. In our paper, we use a weighted-sum-of-covariance
estimator. In order to apply our theory, we need to show the almost sure
convergence of $\widehat{\sigma }_{\epsilon }^{2}$ to a positive limit;
thus, this section of our paper can be compared to \citet{berkesbartlett}.
Let $\rho _{l}^{\left( \epsilon \right) }$ denote the $l$-th order
autocovariance of $\epsilon _{i}$, i.e. $\rho _{l}^{\left( \epsilon \right)
}=E\left( \epsilon _{i}\epsilon _{i-l}\right) $. This can be estimated as
\begin{equation}
\widehat{\rho }_{l}^{\left( \epsilon \right) }=\frac{1}{m}\sum_{i=l+1}^{m}
\widehat{\epsilon }_{i}\widehat{\epsilon }_{i-l}. \label{rho-e}
\end{equation}
Based on (\ref{rho-e}), we define
\begin{equation}
\widehat{\sigma }_{\epsilon }^{2}=\widehat{\rho }_{0}^{\left( \epsilon
\right) }+2\sum_{l=1}^{H}\left( 1-\frac{l}{H+1}\right) \widehat{\rho }
_{l}^{\left( \epsilon \right) }. \label{sig-e-hat}
\end{equation}
\begin{comment}
In (\ref{sig-e-hat}), we have used the Bartlett kernel with bandwidth $H$,
although other choices are also possible as long as standard assumptions are
satisfied - see e.g. \citet{andrews1991}.
\end{comment}
Let $y_{i,l}^{\left( \epsilon \right) }=\epsilon _{i}\epsilon _{i-l}-\rho
_{l}^{\left( \epsilon \right) }$. We need the following regularity conditions
\begin{assumption}
\label{lrv}It holds that: (i) $\epsilon _{i}$ is covariance stationary with $
E\left\vert \epsilon _{i}\right\vert ^{4}<\infty $ for all $i$; (ii) $
\sum_{l=0}^{\infty }l\left\vert \rho _{l}^{\left( \epsilon \right)
}\right\vert <\infty $; (iii) $E\left\vert \sum_{i=l+1}^{m}y_{i,l}^{\left(
\epsilon \right) }\right\vert ^{2}\leq c_{0}m $.
\end{assumption}
It holds that
\begin{proposition}
\label{long-run}We assume that Assumptions \ref{as-2}-\ref{lrv} are
satisfied. As $\min \left( m,H\right) \rightarrow \infty $
\begin{equation}
\widehat{\sigma }_{\epsilon }^{2}=\sigma _{\epsilon }^{2}+o_{a.s.}\left(
\frac{H}{m^{1/2}}\left( \ln m\right) ^{3+\varepsilon }\left( \ln \ln
m\right) \left( \ln H\right) ^{2+\varepsilon }\right) +O\left( \frac{1}{H}
\right) , \label{sig-e}
\end{equation}
for every $\varepsilon >0$.
\end{proposition}
In Proposition \ref{long-run}, a crucial role is played by the bandwidth $H$
. In order to ensure consistency, (\ref{sig-e}) requires that $H\rightarrow \infty$ and
\begin{equation*}
\frac{H}{m^{1/2}}\left( \ln m\right)
^{3+\varepsilon }\left( \ln \ln m\right) \left( \ln H\right) ^{2+\varepsilon
}\rightarrow 0,
\end{equation*}
as $m\rightarrow \infty $.
\begin{comment}
so that the speed of convergence is maximised by
choosing
\begin{equation}
H=H\left( m\right) =c_{0}\frac{m^{1/4}}{\left( \ln m\right)
^{5/2+\varepsilon }\left( \ln \ln m\right) ^{1/2}}. \label{bwidth}
\end{equation}
\end{comment}
\subsection{The monitoring scheme\label{sequential}}
The main idea underpinning (\ref{q-2}) is that, by construction, $Q\left(
m;k\right) $ should pick up the presence of a break, which would introduce a
drift in its fluctuations. In order to check whether $Q\left( m;k\right) $
is growing \textquotedblleft naturally\textquotedblright , i.e. without
breaks, or not, we introduce the function
\begin{equation}
g\left( m;k\right) =\left[ \left( m+k\right) +\left( \frac{m+k}{m}\right)
^{2}\right] ^{1+\gamma }, \label{g-1}
\end{equation}
where the choice of $\gamma $ depends on the length of the monitoring
horizon. Heuristically, the function $g\left( m;k\right) $ should control
the growth rate of $Q\left( m;k\right) $: this is driven by a term
proportional to the cumulative sum of $\epsilon _{i}^{2}$ - which is
controlled by $m+k$ in (\ref{g-1}) - and one proportional to the cumulative
sum of $x_{i}^{2}$, multiplied by the (square of the) estimation error $
\beta -\widehat{\beta }_{m}$ - which is controlled by the term $\left( \frac{
m+k}{m}\right) ^{2}$ in (\ref{g-1}).
\begin{assumption}
\label{horizon}It holds that: (i) $T_{m}=c_{0}m^{\theta }$ for
some $\theta >1$ and $0 < c_{0} < \infty$; (ii) if $k^{\ast } < T$, $k^{\ast }=O\left( m^{\theta
^{\prime }}\right) $ with $0\leq \theta ^{\prime }<\theta $; (iii) $\lim
\inf_{m\rightarrow \infty }\frac{T_{m}}{m}>0$.
\end{assumption}
Assumption \ref{horizon} states that the monitoring horizon should go on for
a sufficiently long time (part \textit{(i)}), and obviously include the
breakdate if there is a break (part \textit{(ii)}). In particular, part
\textit{(i)}, with its implications, is very similar to equation (1.12) in \citet{lajos07}, who also consider the case where monitoring goes on for an infinite time (unless a change is detected).
In practice, $\theta $ is also a given parameter, which is calculated from
Assumption \ref{horizon}\textit{(i)}, once $m$ and $T_{m}$ have been set.
Hence, $\gamma $ is calculated according to the rule
\begin{equation}
\gamma =\frac{1-\delta }{\theta -1}, \label{gamma}
\end{equation}
where $\delta $ is chosen as $0<\delta <1$. In principle, any value of $
\delta $ will ensure the validity of the theory below. In essence, $\gamma$ is chosen as a fraction of $\frac{1}{\theta-1}$; clearly, choosing $\delta$ close to $1$ yields a small $\gamma$, which in turn makes the divergence of $g\left( m;k\right) $ as $m \rightarrow \infty$ slower than in the case of a $\delta$ closer to zero. We discuss the practical impact of the choice of $\delta $ (and $\gamma $) on the ability of the
monitoring procedure to detect breaks in Section \ref{local-alt}. \newline
The function $g\left( m;k\right) $ has been chosen so as to distinguish the
growth rate that $Q\left( m;k\right) $ should have if there were no break,
from the rate at which it would diverge if there were a break.
Heuristically, in absence of breaks, $Q\left( m;k\right) $ should grow, but
slower than $g\left( m;k\right) $; on the other hand, if there is a break,
its presence in the residuals $\widehat{\epsilon }_{i}$ should make $Q\left(
m;k\right) $ grow at a faster pace, and faster than $g\left( m;k\right) $
itself. We point out that the term $\left( m+k\right) $ in (\ref{g-1}) is a
rather coarse estimate, and in principle it could be refined; however, this
term is anyway dominated by the second component of $g\left( m;k\right) $,
and (\ref{g-1}) yields very good results in simulations.
Define
\begin{equation}
\psi _{m,k}=\frac{Q\left( m;k\right) }{g\left( m;k\right) }. \label{psi}
\end{equation}
Based on the above, we expect that $\psi _{m,k}$ drifts to zero as $m$ and $
T_{m}$ diverge if there is no break, whereas it should explode if there is a
break; note that we only consider rates. Indeed, in order to separate such
rates even better, we use the transformation
\begin{equation}
\widetilde{\psi }_{m,k}=\exp \left( \frac{1}{\psi _{m,k}}\right) -1.
\label{psi-1}
\end{equation}
By construction, $\widetilde{\psi }_{m,k}$ has the opposite behaviour as $
\psi _{m,k}$: it can be expected that $\widetilde{\psi }_{m,k}$ drifts to
zero in the presence of a break (that is, under the alternative);
conversely, it should diverge to positive infinity if there is no break
(that is, under the null). Indeed, in the Appendix, we prove that, as $
m\rightarrow \infty $
\begin{align*}
& P\left\{ \omega :\widetilde{\psi }_{m,k}=\infty \right\} =1\text{ under }
H_{0}, \\
& P\left\{ \omega :\widetilde{\psi }_{m,k}=0\right\} = 1\text{ under }H_{A}.
\end{align*}
Given that the test statistic $\widetilde{\psi }_{m,k}$ does not converge to
a non-degenerate random variable under the null (or the alternative), we
propose to use a randomised version of $\widetilde{\psi }_{m,k}$. We present
this as an algorithm, whose output will be a sequence of \textit{i.i.d.}
random variables, with a known distribution (at least asymptotically) under $
H_{0}$, and which diverge under $H_{A,1}$ and $H_{A,2}$.
\begin{description}
\item[\textbf{Step 1}] For each $k$, generate an \textit{i.i.d. }$N\left(
0,1\right) $\textit{\ }sequence $\left\{ \xi _{j}^{\left( k\right) },1\leq
j\leq R\right\} $.
\item[\textbf{Step 2}] Generate the Bernoulli sequence $\zeta _{j}^{\left(
k\right) }\left( u\right) =I\left( \left\vert \widetilde{\psi }
_{m,k}\right\vert ^{1/2}\xi _{j}^{\left( k\right) }\leq u\right) $.
\item[\textbf{Step 3}] Compute
\begin{equation}
\vartheta _{m,R}^{\left( k\right) }\left( u\right) =\frac{2}{R^{1/2}}
\sum_{j=1}^{R}\left( \zeta _{j}^{\left( k\right) }\left( u\right) -\frac{1}{2
}\right) . \label{theta-minor}
\end{equation}
\item[\textbf{Step 4}] Define
\begin{equation}
\Theta _{m,R}^{\left( k\right) }=\int_{-\infty }^{+\infty }\left\vert
\vartheta _{m,R}^{\left( k\right) }\left( u\right) \right\vert ^{2}dF\left(
u\right), \label{theta-maior}
\end{equation}
where $F\left( u\right) $ is a distribution.
\end{description}
Some comments on the sequence $\left\{ \Theta _{m,R}^{\left( k\right)
},1\leq k\leq T_{m}\right\} $ are in order. Consider first the case of the
null of no break. The Bernoulli random variable $\zeta _{j}^{\left( k\right)
}\left( u\right) $ should - asymptotically - be equal to $1$ or $0$ with
probability $\frac{1}{2}$, and thus have mean $\frac{1}{2}$. In this case,
when constructing $\vartheta _{m,R}^{\left( k\right) }\left( u\right) $, a
Central Limit Theorem holds and therefore we expect $\Theta _{m,R}^{\left(
k\right) }$ to have a chi-square distribution. On the other hand, under the
alternative of a break, $\zeta _{j}^{\left( k\right) }\left( u\right) $
should be (heuristically) $0$ or $1$ with probability $0$ or $1$ (depending
on the sign of $u$) - thus its mean is not $\frac{1}{2}$, and a Law of Large
Numbers should hold. Note finally that, by construction, conditionally on
the sample, the sequence $\{\Theta _{m,R}^{\left( k\right) }\}_{k=1}^{T_{m}}$
is independent across $k$; also, by integrating out $u$ in Step 4, the
statistic $\Theta _{m,R}^{\left( k\right) }$ becomes invariant to the choice
of this specification.
The following regularity conditions are needed:
\begin{assumption}
\label{regularity}It holds that: $F\left( u\right)$ is a non-degenerate continuous distribution with (i) $\int_{-\infty }^{+\infty
}u^{2}dF\left( u\right) <\infty $; (ii) the sequences $\left\{ \xi
_{j}^{\left( k\right) },1\leq j\leq R\right\} $ are independent across $k$.
\end{assumption}
Let now $P^{\ast }$ represent the conditional probability with respect to $
\left\{ u_{i},\epsilon _{i},1\leq i\leq T\right\} $; we use the notation
\textquotedblleft $\overset{D^{\ast }}{\rightarrow }$\textquotedblright\ and
\textquotedblleft $\overset{P^{\ast }}{\rightarrow }$\textquotedblright\ to
define, respectively, conditional convergence in distribution and in
probability according to $P^{\ast }$. It holds that
\begin{theorem}
\label{theta-1}We assume that Assumptions \ref{as-2}-\ref{regularity} hold.
As $\min \left( m,R\right) \rightarrow \infty $ with
\begin{equation}
R\exp \left( -m^{\gamma }\right) \rightarrow 0, \label{restriction}
\end{equation}
under $H_{0}$\ it holds that, for each $1\leq k\leq T_{m}$
\begin{equation*}
\Theta _{m,R}^{\left( k\right) }\overset{D^{\ast }}{\rightarrow }\chi
_{1}^{2},
\end{equation*}
for almost all realisations of $\left\{ u_{i},\epsilon _{i},1\leq i\leq
T\right\} $.
\end{theorem}
\begin{theorem}
\label{theta-2}We assume that Assumptions \ref{as-2}-\ref{regularity} hold.
As $\min \left( m,R\right) \rightarrow \infty $, under $H_{A,1}\cup H_{A,2}$
\ it holds that, for each $k\geq \left\lfloor m^{\max \left\{ 1,\theta
^{\prime }\right\} \left( 1+\varepsilon \right) }\right\rfloor $ for all $
\varepsilon >0$
\begin{equation*}
\frac{1}{R}\Theta _{m,R}^{\left( k\right) }\overset{P^{\ast }}{\rightarrow }
1,
\end{equation*}
for almost all realisations of $\left\{ u_{i},\epsilon _{i},1\leq i\leq
T\right\} $.
\end{theorem}
Theorems \ref{theta-1} and \ref{theta-2} are intermediate results. Theorem
\ref{theta-1} stipulates that under the null $\Theta _{m,R}^{\left( k\right)
}$ has, asymptotically, a $\chi _{1}^{2}$ distribution; this result is of
independent interest, and we will make use of it to show that $\Theta
_{m,R}^{\left( k\right) }$ has finite moments of order $2+\varepsilon $ with
$\varepsilon >0$. Note that the only thing that is required is the fact that
$\Theta _{m,R}^{\left( k\right) }$ is an \textit{i.i.d.} sequence, with
finite moments of order $2+\varepsilon $: this is the building block on
which we can construct a detector whose properties can be studied
analytically. In this respect, any other transformation of $\vartheta
_{m,R}^{\left( k\right) }\left( u\right) $ (e.g., the absolute value, or a
power thereof) will also work, giving exactly the same results as in Theorem
\ref{monitoring} below; the only advantage of defining $\Theta
_{m,R}^{\left( k\right) }$ as in Step 4 above is that its asymptotics has
already been studied (see e.g. \citealp{HT16}).
The theorems contain a restriction on the relative rate of divergence of the
pre-monitoring sample size $m$ and the artificial sample size $R$. Heuristically, note that our monitoring procedure is based on having a bounded sequence with finite moments under the null. As the proof of Theorem \ref{theta-1} shows, under the null the statistic $\Theta _{m,R}^{\left( k\right)
}$ has a non-centrality term which vanishes as long as (\ref{restriction}) is satisfied. Conversely, under the alternative it is required that $\Theta _{m,R}^{\left( k\right)
}$ should pass to infinity: Theorem \ref{theta-2} ensures that this occurs at a rate equal to $R$. Thus, Theorem \ref{theta-2} and (\ref{restriction}) provide a family of selection rules for $R$. Given that we only need convergence and divergence, the role played by $R$ can be expected to be rather marginal, which is also confirmed by our simulations (see
Section \ref{montecarlo}). However, we note that the choice $R=m$ satisfies (\ref
{restriction}). Theorem \ref{theta-2}, conversely, states that, under the
alternative where there is a break at $k^{\ast }$, $\Theta _{m,R}^{\left(
k\right) }$ diverges to positive infinity after $k^{\ast }$.
\begin{comment}
Recall that $\Theta
_{m,R}^{\left( k\right) }$, by construction, is also independent across $k$,
conditional on the sample.
\end{comment}
In light of these results, we build a monitoring function, based on the use
of the cumulative sums process. Define the detectors
\begin{equation}
d\left( m;k\right) =\left\vert \sum_{i=m+1}^{m+k}\frac{\Theta _{m,R}^{\left(
i\right) }-1}{\sqrt{2}}\right\vert ,\text{ }1\leq k\leq T_{m}.
\label{detector}
\end{equation}
As can be noted, $d\left( m;k\right) $ is the CUSUM\ process of $\left\{
\Theta _{m,R}^{\left( k\right) },1\leq k\leq T_{m}\right\} $, after
centering and standardizing. \\
Similarly to the literature on structural breaks (see e.g. \citealp{csorgo1997}), we now need to define a family of threshold functions such that if the CUSUM process exceeds the threshold, a change is detected. A standard choice (see \citealp{chu}) is
\begin{align}
& \nu \left( m;k\right) =c_{\alpha ,m}\nu ^{\ast }\left( m;k\right) ,
\label{nu-1-1} \\
& \nu ^{\ast }\left( m;k\right) =m^{1/2}\left( 1+\frac{k}{m}\right) . \label{nu-2-1}
\end{align}
Based on this choice, the FLCT yields that, for every $x$
\begin{align}
& P^{\ast }\left[ \max_{1\leq k\leq T_{m}}\frac{d\left( m;k\right) }{\nu
^{\ast }\left( m;k\right) }\leq x\right] \rightarrow P\left[ \sup_{0\leq
t\leq 1} \left\vert B\left( t\right) \right\vert \leq x
\right] ,
\end{align}
where $B$ is a standard Brownian motion. The limiting law of this expression involves a Brownian motion; intuitively, this being a heteroskedastic process, this procedure may not be the most powerful one; this is further corroborated by \citet{aue2004}, who show that the delay in detecting a changepoint increases as $\eta \rightarrow 0$. A possibility would be to re-scale the monitoring function as suggested in \citet{lajos04} and
\citet{lajos07}, viz. using
\begin{align}
& \nu \left( m;k\right) =c_{\alpha ,m}\nu ^{\ast }\left( m;k\right) ,
\label{nu-1} \\
& \nu ^{\ast }\left( m;k\right) =m^{1/2}\left( 1+\frac{k}{m}\right) \left(
\frac{k}{m+k}\right) ^{\eta }, \label{nu-2}
\end{align}
with $\eta \in \left[ 0,\frac{1}{2}\right] $, and $c_{\alpha ,m}$\ a
critical value. Intuitively, the difference with (\ref{nu-1-1}) is that the monitoring function is now smaller than before, which should ensure higher power. From a technical point of view,
however, choosing $\eta >0$ entails having to use a different asymptotics,
based on \textit{almost sure} as opposed to \textit{weak} convergence. The
fact that the building blocks of $d\left( m;k\right) $ are the $\Theta
_{m,R}^{\left( k\right) }$s - which are, conditional on the sample, \textit{
i.i.d.} and with finite moments - entails that that we can use an array of
almost sure results (see the book by \citealp{csorgo1997}), which in turn
makes it possible to carry out the monitoring using $\eta >0$ in (\ref{nu-2}
).\\
We point out that, despite the considerations above, the choice of threshold functions is by no
means unique, and, as \citet{chu} put it, \textquotedblleft often dictated by
mathematical convenience rather than optimality\textquotedblright ; in our
case, we have defined $\nu ^{\ast }\left( m;k\right) $ as per (\ref{nu-2})
give that the calculation of crossing probabilities (made according to (\ref
{eta-1})) is tractable - see \citet{lajos04} and \citet{lajos07}. We then define the
stopping rule as
\begin{equation}
\widehat{k}_{m}=\inf \left\{ 1\leq k\leq T_{m}\text{ s.t. }d\left(
m;k\right) \geq \nu \left( m;k\right) \right\} , \label{stopping}
\end{equation}
setting $\widehat{k}_{m}=T_{m}$ when (\ref{stopping}) does not hold for $
1\leq k\leq T_{m}$.
The critical value $c_{\alpha ,m}$ is defined, for a given level $\alpha $,
as
\begin{align}
& P\left[ \sup_{0\leq t\leq 1}\frac{\left\vert B\left( t\right) \right\vert
}{t^{\eta }}\leq c_{\alpha ,m}\right] =1-\alpha ,\text{ for }\eta <\frac{1}{2
}, \label{eta-1} \\
& c_{\alpha ,m}=\frac{D_{m}-\ln \left( -\ln \left( 1-\alpha \right) \right)
}{A_{m}},\text{ for }\eta =\frac{1}{2}; \label{eta-2}
\end{align}
in (\ref{eta-1}), $\left\{ B\left( t\right) ,-\infty <t<\infty \right\} $ is
a standard Brownian motion, whereas in (\ref{eta-2}) we have defined
\begin{equation}
A_{m}=\left( 2\ln \ln m\right) ^{1/2}\text{ and }D_{m}=2\ln \ln m+\frac{1}{2}
\ln \ln \ln m-\frac{1}{2}\ln \pi ; \label{am-dm}
\end{equation}
critical values for (\ref{eta-1}) - which do not depend on $m$ - can be found in Table 1 in \citet{lajos04}.
We need the following assumption, which restricts (\ref{restriction}) and
strengthens Assumption \ref{regularity}.
\begin{assumption}
\label{restrict-2}It holds that: (i)
\begin{equation*}
m^{1/2+\tau }R\exp \left( -m^{\gamma }\right) \rightarrow 0,
\end{equation*}
as $\min \left( m,R\right) \rightarrow \infty $, for $\tau >0$; (ii) $
\int_{-\infty }^{+\infty }u^{4+\tau }dF\left( u\right) <\infty $, for some $
\tau >0$.
\end{assumption}
\begin{comment}
Assumption \ref{restrict-2} restricts (\ref{restriction}), by virtue of the
presence of $m^{1/2+\tau }$, and therefore it constraints the choice of the
size of the artificial sample, $R$.
\end{comment}
It holds that
\begin{theorem}
\label{monitoring}We assume that Assumptions \ref{as-2}-\ref{restrict-2} are
satisfied.
As $\min \left( m,R\right) \rightarrow \infty $ with (\ref{restriction}),
under $H_{0}$ it holds that
\begin{align}
& P^{\ast }\left[ \max_{1\leq k\leq T_{m}}\frac{d\left( m;k\right) }{\nu
^{\ast }\left( m;k\right) }\leq x\right] \rightarrow P\left[ \sup_{0\leq
t\leq 1}\frac{\left\vert B\left( t\right) \right\vert }{t^{\eta }}\leq x
\right] \text{ for }\eta <\frac{1}{2}, \label{null-asy-1} \\
& P^{\ast }\left[ \max_{1\leq k\leq T_{m}}\frac{d\left( m;k\right) }{\nu
^{\ast }\left( m;k\right) }\leq \frac{x+D_{m}}{A_{m}}\right] \rightarrow
\exp \left( -\exp \left( -x\right) \right) \text{ for }\eta =\frac{1}{2},
\label{null-asy-2}
\end{align}
for $-\infty <x<\infty $ and almost all realisations of $\left\{
u_{i},\epsilon _{i},1\leq i\leq T\right\} $.
As $\min \left( m,R\right) \rightarrow \infty $, under $H_{A,1}\cup H_{A,2}$
it holds that
\begin{equation}
\max_{1\leq k\leq T_{m}}\frac{d\left( m;k\right) }{\nu ^{\ast }\left(
m;k\right) }\overset{P^{\ast }}{\rightarrow }\infty ,\text{ for any }\eta
\in \left[ 0,\frac{1}{2}\right] , \label{power}
\end{equation}
for almost all realisations of $\left\{ u_{i},\epsilon _{i},1\leq i\leq
T\right\} $.
\end{theorem}
Theorem \ref{monitoring} implies the following
\begin{corollary}
\label{corollary}Under the assumptions of Theorem \ref{monitoring} it holds
that:
\begin{align}
& \lim_{\min (m,R)\rightarrow \infty }P^{\ast }\left( \widehat{k}
_{m}<T_{m}\right) \leq \alpha ,\text{ \ \ under }H_{0}, \label{size} \\
& \lim_{\min (m,R)\rightarrow \infty }P^{\ast }\left( \widehat{k}
_{m}<T_{m}\right) =1,\text{ \ \ under }H_{A,1}\cup H_{A,2}, \label{pow}
\end{align}
for almost all realisations of $\left\{ u_{i},\epsilon _{i},1\leq i\leq
T\right\} $.
\end{corollary}
\section{Discussion and extensions\label{extensions}}
In this section, we investigate two aspects of the monitoring procedure
proposed above. Firstly, we examine the impact
of various test specifications on the power of our procedure (Section \ref
{local-alt}); secondly, we consider the presence of deterministics in (\ref
{model-1}) (Section \ref{deterministics}).
\subsection{Power versus shrinking alternatives and the impact of $g\left( m;k\right)$\label{local-alt}}
Our proposed methodology depends on several specifications in the
construction of the monitoring function, and in the algorithm to compute the
test statistic. In this section, we discuss the impact of such
specifications on the power of the monitoring procedure. In particular, in
this section we discuss the impact of $\gamma $ in the threshold function $
g\left( m;k\right) $ on power versus shrinking alternatives. In Section \ref
{montecarlo}, we also comment on the choices of $u$ and its distribution.
In order to discuss the impact of $\gamma $, we focus on a simple set-up
where there are no deterministics, viz. on model (\ref{model-1})
\begin{equation*}
y_{i}=\beta ^{\prime }x_{i}+\epsilon _{i},
\end{equation*}
and we consider the presence of power versus the local-to-null set-ups
\begin{eqnarray}
H^{\ast}_{A,1} &:&\beta _{i}=\beta +\Delta _{\beta }\left( m\right) I\left[
i>k^{\ast }\right] , \label{ha-1-local} \\
H^{\ast}_{A,2} &:&\epsilon _{i}=\sigma _{\upsilon }\left( m\right) \upsilon
_{i}+u_{i}^{\epsilon }\text{ for }k^{\ast }+1\leq i\leq T.
\label{ha-2-local}
\end{eqnarray}
In (\ref{ha-2-local}), we assume
\begin{equation*}
\upsilon _{i}=\upsilon _{i-1}+u_{i}^{\upsilon },
\end{equation*}
with $u_{i}^{\epsilon }$ independent of $u_{i}^{\upsilon }$. Specifically, in (\ref{ha-1-local}), we consider a
shrinking break where $\Delta _{\beta }\left( m\right) \rightarrow 0$,
whereas in (\ref{ha-2-local}), inspired by \citet{busetti}, we model the
local-to-null case as having $\sigma _{\upsilon }\left( m\right) \rightarrow
0$. Note that we consider, in both equations, the break as shrinking with $m$, since this can be viewed as the sample size on which estimation is based.
Heuristically, as the proof of Theorem \ref{monitoring} shows, in order for
the monitoring procedure to detect changes, it is necessary that $\psi
_{m,k}\rightarrow \infty $ a.s.; thus, intuitively, $\Delta _{\beta }\left(
m\right) $ and $\sigma _{\upsilon }\left( m\right) $, as they drift to zero,
must be \textquotedblleft slow enough\textquotedblright\ to ensure that $
\psi _{m,k}$ diverges. We formalise this in the following theorem
\begin{theorem}
\label{drift}We assume that Assumptions \ref{as-2}-\ref{restrict-2} are
satisfied. Then, under $H^{\ast}_{A,1}$, equation (\ref{power}) holds as $
m\rightarrow \infty $ as long as
\begin{eqnarray}
m^{\delta -\varepsilon }\Delta _{\beta }\left( m\right) &\rightarrow &\infty
,\text{ when }\theta \leq 2, \label{delta-drift-1} \\
m^{\theta \frac{\theta -2+\delta }{2\left( \theta -1\right) }}\Delta _{\beta
}\left( m\right) &\rightarrow &\infty ,\text{ when }\theta >2,
\label{delta-drift-2}
\end{eqnarray}
for some $\varepsilon >0$. Under $H^{\ast}_{A,2}$, equation (\ref{power}) holds as $
m\rightarrow \infty $ as long as
\begin{eqnarray}
m^{\delta -\varepsilon }\sigma _{\upsilon }\left( m\right) &\rightarrow
&\infty ,\text{ when }\theta \leq 2, \label{sigma-drift-1} \\
m^{\theta \frac{\theta -2+\delta }{2\left( \theta -1\right) }}\sigma
_{\upsilon }\left( m\right) &\rightarrow &\infty ,\text{
when }\theta >2. \label{sigma-drift-2}
\end{eqnarray}
\end{theorem}
Theorem \ref{drift}, together with (\ref{gamma}), illustrates what happens
to the power of the monitoring procedure depending on the value of $\gamma $. As can be expected in light of the definition of $g\left( m;k\right) $, choosing a \textquotedblleft small\textquotedblright\ $\gamma$ (which corresponds to choosing $\delta$ close to $1$) enhances the power of the procedure, which, conversely, declines as $\gamma$ increases. This can be understood by noting that the noncentrality of $Q\left( m;k\right) $ is divided by $g\left( m;k\right) $ too. When $\theta \leq 2$, the procedure could potentially (depending on $\delta$) be able to detect breaks as small as $O\left( \frac{1}{m^{1-\epsilon}}\right)$, with $\epsilon > 0$ arbitrarily small. When $\theta > 2$ - that is, when monitoring goes on for a very long time - the ability to detect a small break increases. \\
Note that an alternative could have been to express the break as shrinking with the whole (calibration plus monitoring) sample size, $T$, as done in \citet{wied}.
\subsection{The monitoring procedure in the presence of deterministics \label{deterministics}}
In this section, we consider the following extension of (\ref{model-1})
\begin{equation}
y_{i}=\mu _{0}+\mu _{1}i+\beta ^{\prime }x_{i}+\epsilon _{i}.
\label{model-2}
\end{equation}
Equation (\ref{model-2}) contains, with respect to the previous model, a
constant and a deterministic trend; other extensions could of course be
possible. Our hypothesis testing framework is the same as in the previous
section, namely we test for
\begin{equation*}
H_{0}:\left\{
\begin{array}{l}
\beta _{i}=\beta \\
\epsilon _{i}\text{ is stationary}
\end{array}
\right. \text{ for }1\leq i\leq T_{m},
\end{equation*}
versus the two alternatives
\begin{eqnarray*}
H_{A,1} &:&\beta _{i}=\beta +\Delta _{\beta }I\left[ i>k^{\ast }\right] , \\
H_{A,2} &:&\epsilon _{i}=\epsilon _{i-1}+u_{i}^{\epsilon }\text{ for }
k^{\ast }+1\leq i\leq T.
\end{eqnarray*}
Note that, for the sake of a concise discussion, we do not consider changes
in $\mu _{0}$ or $\mu _{1}$, although again this would be perfectly possible in
principle.
Our monitoring scheme can be adapted as follows. As is typical in this case,
we propose to demean and detrend both $y_{i}$ and $x_{i}$, by estimating
\begin{eqnarray*}
y_{i} &=&a_{0}+a_{1}i+u_{i}^{y}, \\
x_{i} &=&b_{0}+b_{1}i+u_{i}^{x},
\end{eqnarray*}
using OLS, and then computing
\begin{equation}
\widehat{\beta }_{m}^{d}=\left( \sum_{i=1}^{m}\widehat{u}_{i}^{x}\widehat{u}
_{i}^{x\prime }\right) ^{-1}\sum_{i=1}^{m}\widehat{u}_{i}^{x}\widehat{u}
_{i}^{y}, \label{beta-hat-detrend}
\end{equation}
where $\widehat{u}_{i}^{x}$ and $\widehat{u}_{i}^{y}$ are the OLS residuals from the regressions above. After defining
\begin{equation*}
\widetilde{\epsilon }_{i}=y_{i}-\widehat{\beta }_{m}^{d\prime }x_{i},
\end{equation*}
for $m+1\leq i\leq T$, we use the recursively detrended residuals
\begin{equation}
\widehat{\epsilon }_{i}^{d}=\widetilde{\epsilon }_{i}-\left( \widehat{\mu }
_{0,i}+\widehat{\mu }_{1,i}i\right) , \label{res-detrend}
\end{equation}
where
\begin{equation*}
\left(
\begin{array}{c}
\widehat{\mu }_{0,i} \\
\widehat{\mu }_{1,i}
\end{array}
\right) =\left[ \sum_{j=1}^{i}\left(
\begin{array}{cc}
1 & j \\
j & j^{2}
\end{array}
\right) \right] ^{-1}\sum_{j=1}^{i}\left(
\begin{array}{c}
\widetilde{\epsilon }_{j} \\
j\widetilde{\epsilon }_{j}
\end{array}
\right) .
\end{equation*}
Note that other detrending schemes could be proposed also; for example, in
the construction of $\widehat{\epsilon }_{i}^{d}$, one could use
non-recursive estimates $\widehat{\mu }_{0}$ and $\widehat{\mu }_{1}$
(computed once and for all using the sample $1\leq i\leq m$).
\newline
We then define, as in (\ref{q-2}), the cumulative process
\begin{equation}
Q^{d}\left( m;k\right) =\left\vert \frac{1}{\widetilde{\sigma }_{\epsilon
}^{2}}\sum_{i=m+1}^{m+k}\left( \widehat{\epsilon }_{i}^{d}\right)
^{2}\right\vert , \label{q-detrend}
\end{equation}
where the long-run variance estimator $\widetilde{\sigma }_{\epsilon }^{2}$
is computed exactly as in (\ref{sig-e-hat}), using $\widehat{\epsilon }
_{i}^{d}$. We need the following assumption, which complements Assumption
\ref{as-2}\textit{(i)}.
\begin{assumption}
\label{as-detrend}It holds that: (i) $E\left\Vert \sum_{i=1}^{t}i\epsilon
_{i}\right\Vert ^{2}\leq c_{0}t^{3}$, for all $1\leq t\leq T$, and (ii) $
E\left\Vert \sum_{i=1}^{t}ix_{i}\right\Vert ^{2}\leq c_{0}t^{5}$, for all $
1\leq t\leq T$.
\end{assumption}
The next theorem shows that, when using $Q^{d}\left( m;k\right) $ instead of
$Q\left( m;k\right) $ in constructing the monitoring procedure, the same
results hold.
\begin{theorem}
\label{detrending}We assume that Assumptions \ref{as-2}-\ref{as-detrend} are
satisfied. Then, when constructing $d\left( m;k\right) $ using $Q^{d}\left(
m;k\right) $, (\ref{null-asy-1})-(\ref{power}) hold.
\end{theorem}
\section{Numerical and empirical evidence\label{montecarlo}}
In this section, we illustrate the properties of our procedure through a
Monte Carlo exercise (Section \ref{simulations}), and through an application
to US housing market data (Section \ref{empirics}).
\subsection{Simulations\label{simulations}}
We consider the DGP in (\ref{model-1}), with the addition of a constant term, and with $p=1$, namely
\begin{equation*}
y_{i}= \mu_{0} + \beta x_{i}+\epsilon _{i},\text{ with }x_{i}=\sum_{j=1}^{i}u_{j};
\end{equation*}
to evaluate the finite sample performance of our proposed procedure. As discussed in section \ref{deterministics}, we can allow for a constant term in the DGP through recursive demeaning of the data. Incorporating a constant in this fashion allows us to directly compare our new procedure to the equivalent constant-only version of that proposed by \cite{wied}. Noting that our demeaned procedure is mean-invariant, we set $\mu_{0}=0$. We set $\beta =1$ for $1\leq i\leq m$, although unreported experiments show that, as can be expected, this value has no impact on the results.
Innovations $\left\{ \epsilon _{i},u_{i}\right\} $ have been generated as
follows
\begin{align}
u_{i}& =\rho ^{\left( x\right) }u_{i-1}+v_{i}^{u}, \label{ar-u} \\
\epsilon _{i}& =\left( \frac{1+\left( \rho ^{\left( x\epsilon \right)
}\right) ^{2}Var\left( v_{i}^{u}\right) }{1-\left( \rho ^{\left( \epsilon
\right) }\right) ^{2}}\right) ^{-1/2}\epsilon _{i}^{\ast }, \label{s-t-n} \\
\epsilon _{i}^{\ast }& =\rho ^{\left( \epsilon \right) }\epsilon
_{i-1}^{\ast }+v_{i}^{e}+\rho ^{\left( x\epsilon \right) }v_{i}^{u},
\label{ar-e}
\end{align}
In (\ref{ar-u}), we allow for $AR\left( 1\right) $ dynamics in $u_{i}$,
setting $\rho ^{\left( x\right) }\in \left\{ 0,0.5\right\} $. We have also
experimented with other values, noting that results hardly change. In order
to control for the signal-to-noise ratio, we have generated the
idiosyncratic innovation $v_{i}^{u}$ as \textit{i.i.d.} $N\left( 0,\sigma
_{u}^{2}\right) $; by (\ref{s-t-n}). This entails that the signal-to-noise
ratio is exactly equal to $\sigma _{u}^{2}$, and we have used $\sigma
_{u}^{2}=2$ in our experiments. In unreported simulations, we considered $
\sigma_{u}^{2}=1$, with qualitatively similar results. As far as (\ref{ar-e}
) is concerned, we have generated $v_{i}^{e}$ as \textit{i.i.d.} $N\left(
0,1\right) $. Serial dependence in the error term $\epsilon _{i}$ is
explicitly allowed for through $\rho ^{\left( \epsilon \right) }$; note that
when $\rho ^{\left( \epsilon \right) }=1$, this corresponds to $H_{A,2}$,
that is, (\ref{model-1}) becomes a non-cointegrating regression. We report
results for $\rho ^{\left( \epsilon \right) }\in \left\{ 0,0.5,0.9\right\} $. In
(\ref{ar-e}), we also consider the possible presence of endogeneity through
the coefficient $\rho ^{\left( x\epsilon \right) }$, using $\rho ^{\left(
x\epsilon \right) }\in \left\{ 0,0.5\right\} $. The long-run variance of $
\epsilon _{i}$\ is estimated as in (\ref{sig-e-hat}), setting $
H=\left\lfloor m^{1/6}\right\rfloor $, where $\left\lfloor \cdot
\right\rfloor $ denotes the integer part.
As far as the other specifications of the experiment are concerned, we
report results for $T\in \left\{ 100,200,400\right\} $ and $m\in \left\{
\frac{T}{4},\frac{T}{2}\right\} $. When considering the presence of a break,
we have set the changepoint $k^{\ast }=m+\frac{T}{4}$. Experimenting with
other breakdates does not change results in any remarkable way. Under $
H_{A,1}$, we have set $\beta _{i}=\beta +\Delta _{\beta }I\left[ i > k^{\ast
}\right] $, with $\Delta _{\beta }\in \left\{ 0.5,1\right\} $. In addition
to reporting empirical rejection frequencies under the alternative, we also
report the delay in changepoint detection, defined as
\begin{equation}
delay=\frac{\widehat{k}_{m}-k^{\ast }}{k^{\ast }}. \label{delay}
\end{equation}
We now turn to describing the implementation of the test and of the
randomisation algorithm. As far as the former is concerned, we have computed
$g\left( m;k\right) $ setting, according to (\ref{g-1}), $\gamma =0.45$.
Results are similar, especially for large $m$, when using $\gamma =0.4$ and $
\gamma =0.5$.
\begin{comment}
The choice of $\gamma $ directly affects the rate of convergence/divergence
of $\widetilde{\psi }_{m,k}$, and it does have an important impact on the
empirical rejection frequencies. Indeed, as can be expected, decreasing $
\gamma $ results in better power (but worse null rejection frequencies, with
a mild tendency to over-reject in small samples - $T=100$ - when using $
\gamma =0.4$), and vice versa when $\gamma $ is increased to $\gamma =0.5$.
\end{comment}
In the randomisation algorithm, based on (\ref{restriction}), we set $R=m$;
altering this specification (which we have tried in some unreported
experiments) is virtually inconsequential on the empirical rejection
frequencies under both $H_{0}$ and $H_{A,1}\cup H_{A,2}$. Finally, we discuss the choice of $u$. Extracting $u$ from a
distribution - as we make explicit in Step 4 of our algorithm - has the
advantage that $u$ gets integrated out in the construction of $\Theta
_{m,R}^{\left( k\right) }$, making this invariant to the support of $u$
itself. In this respect, choosing $F\left( u\right) $ as the standard normal distribution is a possibility, which is very easy to implement. Indeed, in order to construct $\Theta _{m,R}^{\left( k\right) }$ practically, we can use a Gauss-Hermite quadrature to approximate the integral that defines it,
viz.
\begin{equation}
\Theta _{m,R}^{\left( k\right) }=\frac{1}{\sqrt{\pi }}
\sum_{s=1}^{n_{S}}w_{s}\left( \vartheta _{m,R}^{\left( k\right) }\left( \sqrt{2}
z_{s}\right) \right) ^{2}, \label{upsilon-feasible}
\end{equation}
where the $z_{s}$s, $1\leq s\leq n_{S}$, are the zeros of the Hermite
polynomial $H_{n_{S}}\left( z\right) $ and the weights $w_{s}$ are defined
as
\begin{equation}
w_{s}=\frac{\sqrt{\pi }2^{n_{S}-1}\left( n_{S}-1\right) !}{n_{S}\left[
H_{n_{S}-1}\left( z_{s}\right) \right] ^{2}}. \label{hermite-weights}
\end{equation}
Thus, when constructing $\theta _{m,R}^{\left( k\right) }\left( u\right) $,
we construct $n_{S}$ of these statistics, each with $u=\sqrt{2}z_{s}$; the
values of the roots $z_{s}$, and of the corresponding weights $w_{s}$, are
tabulated e.g. in \citet{salzer}. in our case, we have used $n_{S}=2$, so
that $u=\pm 1$ with equal weight $\frac{1}{2}$; we note that in unreported
experiments we tried $n_{S}=4$ with the corresponding weights, but there
were no changes up to the $4$-th decimal in the empirical rejection
frequencies.
We report results for $\eta=\{0,0.45,0.49,0.50\}$ for the
threshold function in (\ref{nu-2}). We offer a direct comparison of our procedure to that of \cite{wied}. We focus our attention on the IM-OLS version of their test in our simulations as the authors state a preference for IM-OLS, relative to the FM-OLS and D-OLS approaches that they also consider, on the basis of its finite sample performance. We denote this procedure $WW$--$IM$ in what follows. The nominal significance level $\alpha$, in the computation of critical values defined in (\ref{eta-1}) and (\ref{eta-2}), has been set as $\alpha =0.05$. Finally, all experiments have
been carried out using $1,000$ replications.
Empirical rejection frequencies under $H_{0}$ are reported in Table
\ref{tab:Table1}.
We point out that, in our context, the notion of size (control) differs from the standard
Neyman-Pearson testing paradigm. In the latter, empirical rejection
frequencies are expected to be close to their nominal level. In the context
of a sequential testing procedure like ours, as pointed out by
\citet{lajos07} (see also the comments in Ch. 9 in \citealp{Sen}), the
primary purpose is to keep the false detection rate \textit{below} the
chosen level $\alpha$. Indeed, the proportion of false discoveries should go
to zero, since the monitoring can continue for an infinite amount of time.
In this respect, whilst this is the case for all four values of $\eta$
considered in our procedure, setting $\eta=0$ yields the best results, with empirical
rejection frequencies approaching zero across many settings of $m$ and $T$. Examining Panel A of the table, the case of no serial dependence in the error terms, it is clear that for $T=100$, $m=25$ all test procedures exhibit empirical rejection frequencies above their nominal significance levels, with the degree of distortion higher for our procedures than for the $WW$--$IM$ procedure. However, for $T=100$, $m=50$, whilst $WW$--$IM$ offer slightly inflated empirical rejection frequencies (between 0.052 and 0.080 depending on the value of $\rho ^{\left( x \epsilon \right) }$ and $\rho ^{\left( x \right) }$), our test procedure offers empirical rejection frequencies lower than the nominal significance level. For $T=200$ and $T=400$ we observe a similar pattern of results to that under $T=100$, $m=50$; note that the $WW$--$IM$ procedure exhibits a higher degree of upwards distortion (with frequencies up to 0.080 observed).
Allowing for \textit{AR(1)} dynamics in $u_{i}$ has a small upwards effect on the empirical rejection frequencies of all tests for $T=100$ and $m=25$. For all other combinations of $m$ and $T$, little effect is observed for our test procedures, with no effect at all in the case of $T=400$, whereas the $WW$--$IM$ test is observed to be a little more sensitive to these \textit{AR(1)} dynamics. Allowing for endogeneity results in modest increases in the empirical rejection frequencies for our procedures, for most settings of $m$ and $T$, whereas it has the opposite effect for the the $WW$--$IM$ test, resulting in modest decreases in size relative to the no endogeneity case.
Considering Panel B, serial dependence in $\epsilon _{i}$ of $\rho ^{\left( \epsilon \right) }=0.5$ results in upwards size distortion for $WW$--$IM$ for all settings of $m$, $T$, $\rho ^{\left( x \epsilon \right) }$ and $\rho ^{\left( x \right) }$. In contrast, with the exception of $T=100$, $m=25$, where all procedures exhibit empirical rejection frequencies somewhat higher than the nominal significance level, the size of our procedures are robust to this degree of serial correlation, with only very small differences observed in the empirical rejection frequencies from the no serial dependence case (no larger than 0.006 for the settings considered here). This is a pleasing result, given that serial dependence is likely in practice. Finally, turning our attention to Panel C, the case of high serial dependence, $\rho ^{\left( \epsilon \right) }=0.9$, we find that the $WW$--$IM$ procedure exhibits substantial over-sizing for all combinations of $m$, $T$, $\rho ^{\left( x \epsilon \right) }$ and $\rho ^{\left( x \right) }$. As before, with the exception of $T=100$, $m=25$, our procedures display less size distortion than those of $WW$--$IM$. When examining the performance of our procedures, we notice here the role that $m$ plays, with smaller empirical rejection frequencies observed for $m=\frac{T}{2}$ than for $m=\frac{T}{4}$, for a given $T$; and with empirical rejection frequencies decreasing as $T$ increases for a given setting of $m$.
Empirical rejection frequencies and the associated detection delays under $
H_{A,1}$ (i.e. under a change in $\beta$ in the cointegrating regression)
are reported in Tables \ref{tab:Table2a} and \ref{tab:Table2b} respectively.
Considering first the rejection frequencies in Table \ref{tab:Table2a}, in the case of no serial dependence in the errors (Panel A), it is clear that our monitoring procedures offer excellent power, with $\eta=\{0.45,0.49,0.5\}$ outperforming $WW$--$IM$ over most combinations of $\Delta _{\beta }$, $T$, $m$, $\rho ^{\left( x \right) }$, and $\rho ^{\left( x \epsilon \right) }$. There are only 8 instances out of the 48 different combinations of settings in Panel A where $WW$--$IM$ achieves higher power, all cases where $m=\frac{T}{2}$. This result is somewhat anticipated, given that our test allows for $T_{m} \rightarrow \infty$, assumes that the monitoring horizon should go on for a sufficiently long time, and is thus expected to perform better for small $m$ relative to $T_{m}$, whereas \citet{wied} choose $m$ to be large relative to $T$, as discussed in section \ref{sequential}. It is pleasing however, that even in the case of $m=\frac{T}{2}$, our procedures outperform $WW$--$IM$ in terms of power in the majority of cases. Despite its small empirical rejection frequencies under the null, our procedure with $\eta=0$ also performs very well in terms of power, with rejection frequencies under $H_{A,1}$ very similar to those of $\eta=\{0.45,0.49,0.5\}$ in most cases. In addition to the effects of $\rho ^{\left( x \right) }$, and $\rho ^{\left( x \epsilon \right) }$, our procedures appear to be robust to serial dependence in the errors, with high levels of
power maintained under different settings of $\rho ^{\left( \epsilon \right)
}$. Comparing the four values of $\eta$ that we consider here, no setting uniformly outperforms the others,
with little difference in rejection frequencies observed between these values.
Turning our attention to the detection delays reported in Table \ref
{tab:Table2b}, we observe that increasing $m$, increasing $T$, and
increasing $\Delta _{\beta }$ all contribute towards reducing the detection
delay, as we might expect. Contrary to the empirical rejection frequencies,
when considering detection delays a ranking does emerge amongst the
different values of $\eta$ for our procedure, with $\eta=0$ resulting in a longer detection
delay relative to the other settings. This detection delay can be seen as a
trade-off for the very small null rejection frequencies exhibited in Table \ref{tab:Table1}. Setting $
\eta=\{0.45,0.49,0.5\}$ produces the shortest detection delay across
the various settings of $m$, $T$ and $\Delta _{\beta }$ considered here, with very little to distinguish between these settings. Our procedure is capable of detecting a break in the parameters of the
cointegrating regression shortly after the break occurs, as little as 2.7
observations on average after the break for the case of $T=400$, $m=200$ and
$\Delta _{\beta }=1$, where $\rho ^{\left(\epsilon \right)}=0$, $\rho ^{\left(x \epsilon \right)}=0$, $\rho ^{\left(x\right)}=0.5$ and $\eta=0.49$. The $WW$--$IM$ procedure incurs a longer detection delay than our test procedure for every setting considered here.
Finally, empirical rejection frequencies and detection delays under $H_{A,2}$
(i.e. under a switch from a cointegrating to a non-cointegrating regression)
are given in Tables \ref{tab:Table3a} and \ref{tab:Table3b} respectively.
Considering first the rejection frequencies in Table \ref{tab:Table3a}, we note that our procedure is able to offer good levels of power against this alternative hypothesis for most settings. Relative to our results for $H_{A,1}$ in Table \ref{tab:Table2a}, increasing $m$ has a more severe effect on the empirical rejection frequencies, particularly for smaller values of $T$. Examining Panel A, our procedure outperforms $WW$--$IM$ in the majority of cases. Exceptions occur in some instances where $m=\frac{T}{2}$, $\rho ^{\left(x \epsilon \right)}=0.5$.
Comparing Panels A with Panels B and C, it is clear that serial correlation in the errors has the effect of reducing the empirical rejection frequencies for all tests, with a higher degree of serial correlation corresponding to a lower rejection frequency. Of course, this result is to be anticipated given the nature of the alternative hypothesis. In general, with serial correlation of $\rho ^{\left(\epsilon \right)}\{=0.5,0.9\}$, our procedure performs better for $m=\frac{T}{4}$ and $WW$--$IM$ performs better for $m=\frac{T}{2}$, although we note that the empirical rejection frequencies reported here are not size-adjusted, and given the degree of over-sizing exhibited by especially $WW$--$IM$ in Table \ref{tab:Table1}, it is hard to directly compare the tests' performance.
Considering the detection delays under $H_{A,2}$ in Table \ref{tab:Table3b}, we again observe that the delay decreases as $m$ and $T$ increase. We note that delay detections are generally longer under $H_{A,2}$ than for equivalent settings under $H_{A,1}$. As with the detection delays under $H_{A,1}$, when considering our procedure, setting $\eta=0$ provides the longest delay in detection, with $\eta=\{0.45,0.49,0.5\}$ providing the quickest detection of a break. $WW$--$IM$ exhibits a longer detection delay than our procedure with $\eta=\{0.45,0.49,0.5\}$ across all settings, except in one instance\footnote{$T=400$, $m=200$, $\rho ^{\left(\epsilon \right)}=0$, $\rho ^{\left(x \epsilon \right)}=0$ and $\rho ^{\left(x \right)}=0.5$ } where it is still outperformed by $\eta=\{0.45,0.5\}$.
When considering detection delays, it is possible that a test detects a break prematurely, which would lead to a negative delay for that replication according to (4.4), which in turn could result in a misleadingly low reported average delay in Tables \ref{tab:Table2b} and \ref{tab:Table3b}. To further examine the estimated break dates found by these procedures, and to verify whether premature detection is of concern here, in Figures \ref{histHA1} and \ref{histHA2} we consider histograms of the estimated break dates found by our procedure using $\eta=0$ and $\eta=0.45$, as well as the $WW$--$IM$ procedure. For simplicity, we consider the case of $\rho ^{\left(x \right)}=0$, $\rho ^{\left(\epsilon \right)}=0$ and $\rho ^{\left(x \epsilon \right)}=0$. Figure \ref{histHA1} displays estimated break dates under $H_{A,1}$, with \ref{HA1_200} considering $T=200$, $m=\frac{T}{4}$ and $\Delta_{\beta}=1$. It is clear that our procedure using either $\eta=0$ or $\eta=0.45$ provides more accurate break date estimation than $WW$--$IM$, with only a very small difference between these settings of $\eta$. A premature break date is found in only a handful of replications, in the case of $\eta=0.45$, suggesting that early detection is not a significant problem for our test. In Figure \ref{histHA2} we set $T=400$, $m=\frac{T}{2}$ and $\Delta_{\beta}=0.5$, a more challenging circumstance for our procedure as it is designed for small $m$ relative to $T_{m}$. Nevertheless, our procedure displays more accuracy than that of $WW$-$IM$ here.
Figure \ref{histHA2} displays estimated break dates under $H_{A,2}$, with \ref{HA2_200} and \ref{HA2_400} considering the same settings of $T$ and $m$ as in \ref{HA1_200} and \ref{HA1_400} respectively. Again, we are able to note the accuracy of our procedure relative to $WW$--$IM$, and the relatively small numbers of replications where a break is detected before the true break date, $k^{*}$.
Although $\eta=0$ provides the lowest null empirical rejection frequencies, we argue that a sequential monitoring test based on $\eta=0.45$ provides the best overall performance given that it maintains a null empirical rejection frequency below the nominal significance level across most settings of $m$ and $T$, as well as providing the shortest detection delays under both $
H_{A,1}$ and $H_{A,2}$ (although we note that there is very little difference in performance between $\eta=\{0.45,0.49,0.5\}$).
\subsection{Empirical application\label{empirics}}
To demonstrate the practical relevance of the procedure developed in section
\ref{theory}, and inspired by the empirical work of \citet{anundsen} and
\citet{wied}, we investigate the possibility that the US housing market
experienced a structural break in cointegration. Based on the life-cycle
model of housing under the assumption of no arbitrage for the housing
market, \citet{anundsen} analyses two fundamentals-driven cointegrating
relationships. The first approach, known as the \textit{price-to-rent}
model, relies on the user cost of a property being equal to the cost of
renting a property of similar quality in equilibrium, and is given by:
\begin{equation} \label{pricerent}
ph_{t} = \gamma_{r}r_{t} + \gamma_{UC}UC_{t} + u_{t}
\end{equation}
where $ph_{t}$ is the logarithm of real housing prices at period $t$, $r_{t}$
is the logarithm of real rents, and $UC_{t}$ is the real direct user cost of
housing, computed as
\begin{equation*}
UC_{t} = (1-\tau^{y}_{t})(i_{t}+\tau^{p}_{t}) - \pi_{t} + \delta_{t},
\end{equation*}
where $\tau^{y}_{t}$ is the marginal personal income tax rate (measured here
at twice the median income), $\tau^{p}_{t}$ is the marginal tax rate on
personal property, $i_{t}$ is the nominal interest rate, $\pi_{t}$ is
overall price inflation, and $\delta_{t}$ is the housing depreciation rate.
The second approach, known as the \textit{inverted demand} model, assumes
that imputed rent is a function of income and housing stock, and is given by
the below equation:
\begin{equation}
ph_{t}=\widetilde{\gamma }_{y}y_{t}+\widetilde{\gamma }_{h}h_{t}+\widetilde{
\gamma }_{UC}UC_{t} + \widetilde{u}_{t}, \label{invdemand}
\end{equation}
where $y_{t}$ is the logarithm of real per capita disposable income and $
h_{t}$ is the logarithm of the per capita housing stock.
Assuming that the variables in (\ref{pricerent}) and (\ref{invdemand}) are $
I(1)$, economic theory predicts that $u_{t}$ and $\widetilde{u}_{t}$ are
both $I(0)$. That is, two cointegrating relationships exist between housing
prices and their fundamentals. A breakdown of these cointegrating
regressions therefore indicates that housing prices are no longer being
driven by these fundamentals. Following the definition of \cite{stiglitz},
\textit{inter alia}, that an asset bubble exists when its price no longer
appears to be justified by the value of its fundamental components, \cite
{anundsen} interprets a breakdown in these cointegrating relationships as
evidence of a bubble in housing prices. Indeed, following \citet{anundsen},
we also allow for an intercept and linear trend term in each model, viz.
\begin{equation}
ph_{t}=\theta _{1}+\theta _{2}t+\theta _{3}r_{t}+\theta _{4}UC_{t}+u_{t},
\label{ptr-2}
\end{equation}
and
\begin{equation}
ph_{t}=\widetilde{\theta }_{1}+\widetilde{\theta }_{2}t+\widetilde{\theta }
_{3}y_{t}+\widetilde{\theta }_{4}h_{t}+\widetilde{\theta }_{4}UC_{t}+
\widetilde{u}_{t}. \label{id-2}
\end{equation}
\citet{anundsen} applies models (\ref{ptr-2})-(\ref{id-2}) to quarterly US
housing market data over the sample period 1976:Q1 - 2010:Q4. Specifically,
he estimates vector autoregression models and undertakes Johansen
cointegration testing for both the price-to-rent and inverted demand
equations using expanding sub-samples of the data, starting with an initial
sub-sample from 1976:Q1 - 1995:Q4 and then subsequently adding four new
observations until the full sample is used. This analysis finds evidence of
a cointegrating relationship in the housing market up until 2002 in the
price-to-rent model; evidence in favour of cointegration disappears when
2002:Q4 is included in the sample, but there is evidence of a return to a
cointegrating relationship towards the end of the sample. As far as the
inverted demand model is concerned, a similar pattern is found, with
cointegration breaking down in 2001:Q4. These results imply the emergence of
bubble behaviour in the housing market beginning in 2001-2002; however,
\citet{wied} highlight that the analysis suffers from the problem of
multiple testing, leading to uncontrolled size. Considering the same
dataset, they apply their real time monitoring procedure to models (\ref
{ptr-2})-(\ref{id-2}), with $WW$--$IM$ detecting a breakdown in cointegration at 2006:Q4 for the price-to-rent model and 2004:Q2 for the inverted demand model (with the FM-OLS version of the procedure finding a slightly earlier break of 2003:Q2 for this model). This delay in detection, relative to the results of \citet{anundsen}, can be viewed as a trade-off for asymptotic validity, and therefore controlled size.
We apply our sequential monitoring procedure to the dataset discussed above, containing information on US house prices from 1976:Q1 - 2010:Q4.\footnote{Detailed information on the dataset sources and construction are contained
in \cite{anundsen}. The dataset has been downloaded from the \textit{Journal
of Applied Econometrics} data archive.} In line with the previous two
studies, our calibration sample runs from 1976:Q1 -
1995:Q4, such that $m=80$; effectively, this means that \textquotedblleft future data\textquotedblright (and our monitoring) starts in 1996:Q1 (this being the earliest possible break date that we can detect). We set $R=m$ and $n_{s}=2$ as before. We allow for a constant and linear trend in both models through the recursive demeaning and detrending method discussed in section \ref{deterministics}. In view of our
simulation results in the previous section, we have used $\eta=0.45$.
\footnote{
Although, in unreported results, we note that setting $\eta=\{0,0.49,0.50\}$
provides the same break date estimate as $\eta=0.45$ for both models.} Similarly, based on the Monte Carlo evidence, we set $\gamma=0.4$, which should ensure size control, while decreasing the detection delay.
Figure \ref{empexgraphs} displays the residuals obtained from our estimation of the price to rent and inverted demand models, using recursive demeaning and detrending. From visual inspection, it is
clear that the residuals of both models undergo a period of mean-reverting
behaviour in the earlier part of the sample, whereas more persistent
behaviour in the residuals is observed from the early 2000s, lending support
to the hypothesis that a structural break in cointegration occurs during the
sample period. Considering first the price-to-rent model in Figure \ref{ptr}, our sequential monitoring procedure finds evidence of a break in
cointegration in 2005:Q3, 5 quarters earlier than the $WW$--$IM$ test, although
still somewhat later than the detection date in the initial experiment of
\cite{anundsen}. Examining the inverted demand model, in Figure \ref{id},
evidence is found of a break in cointegration in 2004:Q1, somewhat earlier than for the price-to-rent model, in line with the results of \cite{wied} and \cite{anundsen}. Thus, our results support the claim of a breakdown in fundamentals-driven cointegrating relationships in the US housing markets during the housing bubble of the 2000s.
\section{Conclusions\label{conclusions}}
In this paper we have investigated the issue of monitoring a cointegrating
regression. Having stability as the null hypothesis, we develop a procedure
to detect changes in the regression coefficients and/or from
cointegration to non-cointegration. Our procedure is based
on using the cumulative sums of squared residuals; at each point in the
monitoring horizon, we randomise the cumulative sum process, thereby
obtaining an \textit{i.i.d.} sequence with finite moments of arbitrarily
high order. We then use the results in \citet{lajos04} and \citet{lajos07}
to construct a family of procedures which may be viewed as a complement to
the results in \citet{wied}.
We point out that, as well as deriving the aforementioned statistics, in
this paper we have proposed a general methodology to construct monitoring
schemes in the context of a cointegrating regression. The approach we
propose can be readily generalised to use other statistics (e.g., upon
calculating the relevant rates, even the KPSS type statistic employed in
\citet{wied} could be randomised and used in our algorithm), or to other
hypothesis testing frameworks. As a leading example, \citet{wagner} consider
the very interesting case where (\ref{model-1}) is, to begin with, a
non-cointegrating regression with $\epsilon _{i}\sim I\left( 1\right) $, and
the purpose of monitoring is to verify whether (\ref{model-1}) becomes a
cointegrating regression, with $\epsilon _{i}\sim I\left( 0\right) $.
Although we leave this interesting research question for future study, we
point out that a monitoring scheme for this case could be readily developed.
Indeed, one could use exactly the same approach as we do, using
\begin{equation*}
\widehat{\psi }_{m,k}=\exp \left( \psi _{m,k}\right) -1,
\end{equation*}
instead of (\ref{psi-1}). This and others issues are under investigation by
the authors.
\begin{comment}
\section{Data availability statement}
The dataset employed in Section \ref{empirics} may be found online in the
supporting information tab for this article (file name: Anundsendata.xslx).
\section{Supporting information}
Additional supporting information (technical lemmas and the proof of Proposition 1) is contained in the Supplement (TrapaniWhitehouseSupplement.pdf), which can be found online in the supporting information tab for this article.
\end{comment}
{\small {\setlength{\bibsep}{.2cm}
\bibliographystyle{chicago}
\bibliography{biblio}
}}
\clearpage