EconBase
← Back to paper

Sequential monitoring for explosive volatility regimes

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

157,977 characters · 22 sections · 96 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Sequential monitoring for explosive volatility regimes

\address{Lajos Horv\'ath, Department of Mathematics, University of Utah, US} \address{Lorenzo Trapani, University of Leicester, UK, and Universita' di Pavia, Italy} \address{Shixuan Wang, Department of Economics, University of Reading, UK} \subjclass{Primary 62M10, 91B84; secondary 60G10, 62F12}

abstractIn this paper, we develop two families of sequential monitoring procedure to (timely) detect changes in a GARCH(1,1) model. Whilst our methodologies can be applied for the general analysis of changepoints in GARCH(1,1) sequences, they are in particular designed to detect changes from stationarity to explosivity or vice versa, thus allowing to check for \textquotedblleft volatility bubbles\textquotedblright . Our statistics can be applied irrespective of whether the historical sample is stationary or not, and indeed without prior knowledge of the regime of the observations before and after the break. In particular, we construct our detectors as the CUSUM process of the quasi-Fisher scores of the log likelihood function. In order to ensure timely detection, we then construct our boundary function (exceeding which would indicate a break) by including a weighting sequence which is designed to shorten the detection delay in the presence of a changepoint. We consider two types of weights: a lighter set of weights, which ensures timely detection in the presence of changes occurring \textquotedblleft early, but not too early\textquotedblright\ after the end of the historical sample; and a heavier set of weights, called \textquotedblleft R\'{e}nyi weights\textquotedblright \thinspace\ which is designed to ensure timely detection in the presence of changepoints occurring very early in the monitoring horizon. In both cases, we derive the limiting distribution of the detection delays, indicating the expected delay for each set of weights. Our theoretical results are validated via a comprehensive set of simulations, and an empirical application to daily returns of individual stocks.

\doublespacing

Introduction

In recent years, developing tools for the (ex-ante or ex-post) detection of the onset or the collapse of a bubble in financial markets has been one of the most active research areas in financial econometrics; we refer the reader, in particular, to the seminal articles on ex-post detection by phillips2011, and phillips2015testing2, and also to skrobotov2023testing for a review. As far as ex-ante detection - that is, finding the onset or collapse of a bubble in real time, as new data come in - is concerned, this important issue has also been studied in numerous recent contributions; although a comprehensive literature review goes beyond the scope of this paper, we refer to the articles by homm2012testing, phillips2015testing, and whitehouse2023real, inter alia; in particular, the paper by whitehouse2023real also contains a comprehensive literature review of in-sample and online bubble detection methods. A common trait to the vast majority of the existing literature is its reliance on a linear specification, usually an AutoRegressive (AR) model, to capture regime changes in the dynamics of log prices. Whilst such a modelling choice can be justified from the theoretical point of view (see e.g. phillips2011dating), and whilst its analytical tractability offers obvious advantages from a mathematical standpoint (see e.g. phillips2007limit, and aue2007limit), a major issue is that using an AR model is fraught with difficulties when monitoring for changes from an explosive towards a stationary regime. Several promising solutions have been proposed such as the reverse regression approach by phillips2018financial. horvath2023real propose a different model, based on a Random Coefficient Autoregressive (RCA) specification, where inference is always standard normal irrespective of stationarity or the lack thereof (aue2011), and develop a family of sequential monitoring procedures based on the weighted CUSUM process, to check whether the deterministic part of the autoregressive root changes over time. Such a testing set-up also encompasses both the case of a switch from a stationary to an explosive regime (thus indicating the start of a bubble phenomenon), and a change from an explosive to a stationary regime (thus signalling the collapse of a bubble).

The theory developed in horvath2023real paves the way to a more general research question, namely developing sequential monitoring techniques which are robust to both the initial regime (i.e., which can be employed irrespective of whether the observations in the training/historical sample are stationary or explosive), and to the type of change which occurs after a changepoint (i.e., which can detect changes from stationarity to another stationary regime, or to a explosive regime; or from an explosive regime to another explosive regime or a stationary one). From a technical viewpoint, this question is nontrivial for at least three\ reasons. First, proposing a changepoint detection methodology whose asymptotics is the same irrespective of stationarity or explosivity is not easy per se, because the partial sum processes which constitute the building blocks of e.g. CUSUM-based statistics require completely different approximations depending on whether the observations are stationary or not. Second, in order to ensure timely changepoint detection, weighted versions of the CUSUM process need to be considered, with different sets of weights ensuring optimal detection delays for different changepoint locations within the monitoring horizon. Third, on account of the previous point, it is important to offer, to the applied user, a battery of results on the limiting distribution of the detection delays, so as to gauge the expected detection delay depending on the location of the changepoint and the weighing scheme employed.

Motivated by the questions above, in this paper we investigate, with an emphasis on completeness, the issue of sequential detection for regime changes in a GARCH(1,1) sequence

equation[equation omitted — 179 chars of source]

In particular, we develop two families of detectors based on the weighted CUSUM process of the quasi-Fisher scores associated with ((ref)): one with \textquotedblleft mild\textquotedblright\ weights, designed to detect changes that may occur \textquotedblleft not too early\textquotedblright\ after the start of the monitoring period; and one with heavy weights, designed instead to detect changes occurring \textquotedblleft very early\textquotedblright\ after the start of the monitoring period. The latter is based on applying to the CUSUM process a set of (heavy) weights, resulting in a family of test statistics known as R\'{e}nyi statistics (see horvath2021 and horvathmiller, for in-sample tests, and ghezzi2024fast, for sequential monitoring). Our methodologies can, in principle, be applied to detect any type of change in the vector $\left( \omega ,\alpha ,\beta \right) $. However, seeing as: (1) our main interest is in detecting changes between regimes (e.g., from stationarity to nonstationarity, or vice versa); the stationarity or lack thereof of ((ref)) is determined solely by $\alpha $ and $\beta $; and (3) in the presence of nonstationarity, $\omega $ is not identified ( jensen2004asymptotic), we focus on monitoring for changes only in the sub-vector $\left( \alpha ,\beta \right) $. Detecting shifts in the behaviour of the (conditional) volatility process $\sigma _{i}$ is important in general; as hillebrand2005neglecting notes, when neglecting a break inference is biased in finite samples, and the sum of the estimated autoregressive parameters $\alpha $ and $\beta $ converges to one. Furthermore, changes (and, occasionally, explosions) in the volatility of time series are often observed in practice (see e.g. bloom2007uncertainty, and jurado2015measuring). Hence, finding the start or the end of an explosive regime in $\sigma _{i}$ is of practical relevance because, as richter2023testing put it, one \textquotedblleft often sees sudden, integrated, or mildly explosive behaviour in the second moment of the process which bounces back after a while\textquotedblright\ (p. 468). Changes between stationarity and explosivity in ((ref)) can be interpreted as volatility bubbles , i.e. events in which the second moment of the data (rather than the data, e.g. prices, themselves) experiences periods of exhuberance. The link between a volatility phenomenon and a \textquotedblleft proper\textquotedblright\ bubble has not been fully explored yet (see jurado2015measuring), and, empirically, explosive regimes in volatility can be ascribed to various sources in addition to bubbles (see sornette2018can, and richter2023testing). Nevertheless, as jarrow2023explosion put it, \textquotedblleft price bubbles result from excess speculative trading decoupled from the asset's fundamentals (dividends and liquidation value), which increases the asset's price volatility to extreme levels\textquotedblright\ (p. 478). Hence, an analysis of the regimes of the volatility of financial variables is bound to contribute to a better understanding of bubble phenomena.

Based on the discussion above, in this paper we propose a battery of tests for the sequential monitoring of the volatility of financial variables, which complements the existing tests for bubbles based on the conditional mean. Specifically, we make at least three contributions. First, we study online detection of changes between stationarity and explosivity and vice versa, which - to the best of our knowledge - is a novel result in the literature, and which complements the ex-post detection statistics studied in richter2023testing. Second, we develop the full-blown theory for R \'{e}nyi statistics in the context of sequential monitoring of a GARCH(1,1) model. Third, we derive the limiting distribution of detection delays for all our monitoring schemes, including those based on R\'{e}nyi statistics; this is an entirely novel result, which complements the results in horvath2020sequential.\newline

The remainder of the paper is organised as follows. We discuss our workhorse model and the main assumptions, as well as the test statistics, in Section (ref). The theory is reported in Section (ref): in particular, the asymptotics under the null is in Section (ref), and the full-blown asymptotics of the detection delay in the presence of a changepoint is in Section (ref). We validate our theory through a comprehensive set of simulations (Section (ref)), and an empirical application to daily returns of individual stocks (Section (ref)). Section (ref) concludes.

NOTATION. We define the Euclidean norm of a vector $a$ as $\Vert a\Vert $. We denote the integer part of a number as $\left\lfloor \cdot \right\rfloor $ . We use: \textquotedblleft a.s\textquotedblright\ for \textquotedblleft almost sure(ly)\textquotedblright ; \textquotedblleft $\rightarrow $ \textquotedblright\ for the ordinary limit; \textquotedblleft $\overset{ \mathcal{D}}{\rightarrow }$\textquotedblright\ for convergence in distribution; \textquotedblleft $\overset{\mathcal{P}}{\rightarrow }$ \textquotedblright\ for convergence in probability; \textquotedblleft $ \overset{{\mathcal{D}}}{=}$\textquotedblright\ for equality in distribution. Positive, finite constants are denoted as $c_{0}$, $c_{1}$, ... and their value may change from line to line. Other, relevant notation is introduced later on in the paper.

Model, assumptions, and hypothesis testing

Model, assumptions and hypotheses of interest

The time dependent GARCH (1,1) sequence is defined by the recursion

equation[equation omitted — 192 chars of source]

where $\sigma _{0}^{2}$, $y_{0}^{2}$ are initial values, and $\alpha _{i}$, $ \beta _{i}$ and $\omega _{i}$ are positive parameters.

Whilst the hypothesis testing framework is spelt out below, our monitoring schemes are all based on the maintained assumption that we have $m$ observations which form a stable period (this is also known as the non-contamination assumption in CSW96), viz.

equation[equation omitted — 150 chars of source]

We denote the value of the common parameter in (ref) as $ \mbox{\boldmath${\theta}$}_{0}=(\alpha _{0},\beta _{0},\omega _{0})^{\top }$ . Prior to spelling out the main assumptions, we review the conditions for the stationarity of $y_{i}$. As nelson1991conditional shows (see also bougerol1992strict, and francq2012strict)

enumerate• if $E\log \left\vert \alpha _{0}\epsilon _{0}^{2}+\beta _{0}\right\vert <\infty $, then $\sigma _{i}$ converges exponentially fast to a unique, strictly stationary and ergodic solution $\left\{ \overline{ \sigma }_{i},-\infty <i<\infty \right\} $ for all initial values $\epsilon _{0}$ and $\sigma _{0}$; • if $E\log \left\vert \alpha _{0}\epsilon _{0}^{2}+\beta _{0}\right\vert >\infty $, then $\sigma _{i}$ is nonstationary with $\sigma _{i}\overset{a.s.}{\rightarrow }\infty $ exponentially fast ( nelson1991conditional); • if $E\log \left\vert \alpha _{0}\epsilon _{0}^{2}+\beta _{0}\right\vert =\infty $, then $\sigma _{i}$ is nonstationary, but this is a much more delicate case; indeed, kluppelberg2004continuous show that $\sigma _{i}\overset{\mathcal{P}}{\rightarrow }\infty $, but a.s. divergence to infinity cannot be established. \newline

We will develop several monitoring schemes for the null hypothesis that the parameter $\mbox{\boldmath${\theta}$}_{0}$ undergoes no changes after the historical training period $1\leq i\leq m$, i.e.

equation[equation omitted — 168 chars of source]

Under the alternative, we assume that there is a change at time $m+k^{\ast }$ ; whilst this would correspond to having $\left( \omega _{m+k^{\ast }-j},\alpha _{m+k^{\ast }-j},\beta _{m+k^{\ast }-j}\right) $ $\neq $ $\left( \omega _{m+k^{\ast }+j},\alpha _{m+k^{\ast }+j},\beta _{m+k^{\ast }+j}\right) $ for all $j\geq 0$, it is well know that the $\omega _{i}$'s cannot be identified in explosive, nonstationary regimes ( francq2012strict). Hence, we will test for

align[align omitted — 309 chars of source]

i.e., for the possible presence of changes in $\alpha _{i}$ and $\beta _{i}$ only. Note that these are anyway the parameters of interest, since the stationarity (or lack thereof) of $y_{i}$ is not affected by $\omega _{i}$.

We require the following assumptions on $\mbox{\boldmath${\theta}$}_{0}$, and on the innovations $\{\epsilon _{i},-\infty <i<\infty \}$.

assumptionIt holds that: $\alpha _{0}>0$, $\beta _{0}>0$ and $\omega _{0}>0$.
assumptionIt holds that: (i) $\{\epsilon _{i},-\infty <i<\infty \} $ are independent and identically distributed random variables; (ii) $\epsilon _{0}^{2}$ is nondegenerate; (iii) $E\epsilon _{0}=0$ , $E\epsilon _{i}^{2}=1$, $0<\mbox{\rm var}(\epsilon _{0}^{2})$, and $ E|\epsilon _{0}|^{\kappa }<\infty $ with some $\kappa >4$.

Assumptions (ref) and (ref) are standard. In particular, it is worth noting that, in Assumptions (ref), there is no requirement as to the stationarity properties of $\left\{ y_{i},1\leq i\leq m\right\} $: the observations in the training sample can belong to a stationary or explosive volatility regime.

Estimation and monitoring schemes

As stated in the Introduction, the main purpose of our analysis is to offer a detection scheme which finds changes in the parameters of a GARCH(1,1) model as soon as possible after the training period. In this section, we propose several detectors, all based on the CUSUM process of the quasi-Fisher scores.

As is typical, we estimate the parameter $\mbox{\boldmath${\theta}$}_{0}$, using the training sample, by Quasi Maximum Likelihood (QML). The $\log $ likelihood function is given by

equation*[equation* omitted — 207 chars of source]

where $\mbox{\boldmath${\theta}$}=(\omega ,\alpha ,\beta )^{\top }$, $\omega >0,\alpha >0$ and $\beta >0$, and the random functions $\bar{\sigma}_{i}^{2}( \mbox{\boldmath${\theta}$})$ given by the recursion

equation*[equation* omitted — 175 chars of source]

the recursion starts from the initial values $y_{0}$ and $\bar{\sigma} _{0}^{2}$. The QML estimator computed from the historical sample is denoted as $\hat{\mbox{\boldmath${\theta}$}}_{m}$, with

equation*[equation* omitted — 205 chars of source]

where the (compact) space $\mbox{\boldmath${\theta}$}$ is defined as

equation*[equation* omitted — 252 chars of source]

$0<\underline{\omega },\overline{\omega },\underline{\alpha },\overline{ \alpha },\underline{\beta },\overline{\beta }<\infty $.

Starting with the initial values ${y}_{m}$ and ${\sigma }_{m}^{2}$, we define the random functions $\bar{\sigma}_{m+k}^{2}(\mbox{\boldmath${ \theta}$})$ based on the observations after the historical sample by the recursion

equation*[equation* omitted — 175 chars of source]

with the log likelihood function given by

equation*[equation* omitted — 209 chars of source]

Hence, we define the CUSUM\ process of the quasi-Fisher scores as

equation[equation omitted — 298 chars of source]

Heuristically, under the null of no change, the scores have zero mean; hence, the partial sum process ${\mathcal{r}}_{m,k}(\mbox{\boldmath${ \theta}$})$ calculated at $\hat{\mbox{\boldmath${\theta}$}}_{m}$\ should also fluctuate around zero with increasing variance. Conversely, in the presence of a break (at, say, $k^{\ast }$), $\hat{\mbox{\boldmath${\theta}$}} _{m}$ is a biased estimator for the \textquotedblleft new\textquotedblright\ parameter $\mbox{\boldmath${\theta}$}_{m+k^{\ast }}$; thus, ${\mathcal{r}} _{m,k}(\mbox{\boldmath${\theta}$})$, calculated at $\hat{\mbox{\boldmath${ \theta}$}}_{m}$, should have a drift term. In the light of these heuristic considerations, we propose the following detector

equation[equation omitted — 203 chars of source]

where

equation*[equation* omitted — 450 chars of source]

Based on ((ref)), a break is flagged as soon as the detector ${ \mathcal{D}}_{m}(k)$ exceeds a threshold. We call such a threshold the boundary function. Similarly to CSW96, lajos04, horvath2020sequential, homm2012testing, we use the boundary function, designed for a closed-ended procedure - i.e., for a procedure which terminates at a certain time, say ${\mathcal{n}}$, if there is no change

equation[equation omitted — 137 chars of source]

On account of ((ref)) and ((ref)), a changepoint is found at a stopping time $\tau _{m}$ defined as

equation[equation omitted — 324 chars of source]

The (user-chosen) parameter $\eta $ in ((ref)) determines the timeliness of changepoint detection of our sequential monitoring procedure. aue2004delay and aue2008monitoring show that, as $\eta $ approaches $1$, changepoints are detected with a smaller and smaller delay depending on their location; on the other hand, different values of $\eta $ work better for different changepoint locations, as also pointed out in a recent contribution by kirch2022sequential. In particular, values of $0\leq \eta <1$ are able to offer short detection delays for breaks that do not occur \textquotedblleft too early\textquotedblright\ after $m$.

In order to detect earlier changes, ghezzi2024fast (building on previous contributions by horvath2021, and horvathmiller , developed for in-sample changepoint detection) suggest using R\'{e}nyi type statistics, with stopping rule

equation[equation omitted — 347 chars of source]

where $r$ is a trimming sequence specified in Assumption (ref), and

equation[equation omitted — 118 chars of source]

We note that ((ref)) and ((ref)) exclude the case $\eta =1$. Indeed, aue2004delay show that, as far as stopping times under the alternative are concerned, using $\eta =1$ would produce the shortest detection time. The case $\eta =1$ requires to be studied separately. Following lajos07, we modify the boundary function. Let

equation*[equation* omitted — 94 chars of source]

We use the boundary functions

equation[equation omitted — 146 chars of source]

and

equation[equation omitted — 160 chars of source]

The corresponding stopping times - denoted as $\tau _{m}^{\ast }$ and $\bar{ \tau}_{m}^{\ast }$ - are defined exactly in the same way as $\tau _{m}$ and $ \bar{\tau}_{m}$ in ((ref)) and ((ref)), respectively, using now the boundaries $g_{m}^{\ast }(k)$ and $\bar{g}_{m}^{\ast }(k)$.\newline

As mentioned above, we consider a closed-ended scheme, which terminates ${ \mathcal{n}}$ periods after $m$. The following assumptions characterise the length of the monitoring horizon and of the trimming sequence $r$ defined in ((ref)); in particular, Assumption (ref) is designed in order to consider only early changepoint detection.

assumptionIt holds that ${\mathcal{n}}\rightarrow \infty $, and ${ \mathcal{n}}/m\rightarrow 0$.
assumptionIt holds that, in equation ((ref)), $r\rightarrow \infty $ and $r/{\mathcal{n}}\rightarrow 0$.

Asymptotics

In this section, we investigate the asymptotic behaviour of our monitoring schemes under the null and under the alternative hypotheses.

Asymptotics under the null

Let $\mathbf{W}(t)=(W_{1}(t),W_{2}(t)),t\geq 0$ be a two dimensional standard Wiener process - i.e.,\ $\{W_{1}(t),t\geq 0\}$ and $ \{W_{2}(t),t\geq 0\}$ are two independent Gaussian processes with $ EW_{1}(t)=EW_{2}(t)=0$, and covariance kernel $EW_{1}(t)W_{1}(s)$ $=$ $ EW_{2}(t)W_{2}(s)$ $=\min (t,s)$.

theoremWe assume that $H_{0}$ of ((ref)) and Assumptions (ref)--(ref) hold, and that $E\log (\alpha _{0}\epsilon _{0}^{2}+\beta _{0})\neq 0$.\newline (i) If $0\leq \eta <1$, then we have \begin{equation*} \lim_{m\rightarrow \infty }P\{\tau _{m}={\mathcal{n}}\}=P\left\{ \sup_{0<t\leq 1}\frac{1}{t^{\eta }}\left\Vert \mathbf{W}(t)\right\Vert ^{2}\leq {\mathcal{c}}\right\} . \end{equation*} (ii) If in addition, Assumption (ref) also holds and $\eta >1$, then we have \begin{equation*} \lim_{m\rightarrow \infty }P\{\bar{\tau}_{m}={\mathcal{n}}\}=P\left\{ \sup_{1\leq t<\infty }\frac{1}{t^{\eta }}\left\Vert \mathbf{W}(t)\right\Vert ^{2}\leq {\mathcal{c}}\right\} . \end{equation*}

Theorem (ref) offers a rule to calculate the asymptotic critical values; for a given nominal level $\alpha $, the critical value ${\mathcal{c}} _{\alpha }$ is defined as

equation*[equation* omitted — 150 chars of source]

for all $0\leq \eta <1$, and

equation*[equation* omitted — 152 chars of source]

for $\eta >1$. Using the scale transformation of the Wiener process, it immediately follows that $\{\mathbf{W}(t),t>0\}$ $\overset{{\mathcal{D}}}{=}$ $\{t\mathbf{W}(1/t),t>0\}$; hence, for all $\eta >1$

equation*[equation* omitted — 182 chars of source]

Theorem (ref) rules out the boundary case $E\log (\alpha _{0}\epsilon _{0}^{2}+\beta _{0})=0$; this is because we would need an exact (and large enough) rate of divergence for $\left\vert \sigma _{i}\right\vert $ as $ i\rightarrow \infty $, but this result is not available in the case $E\log (\alpha _{0}\epsilon _{0}^{2}+\beta _{0})=0$ (see francq2012strict , p. 823; and also Theorem 4 in HT2019, where a similar problem is encountered, and the discussion thereafter).\newline

Theorem (ref) does not consider the case $\eta =1$, which corresponds to the stopping times $\tau _{m}^{\ast }$ and $\bar{\tau}_{m}^{\ast }$ based on the boundaries defined in ((ref)) and ((ref)) respectively. The case $\eta =1$ is studied separately in the following theorem.

theoremWe assume that $H_{0}$ of ((ref)) and Assumptions (ref)--(ref) hold, and that $E\log (\alpha _{0}\epsilon _{0}^{2}+\beta _{0})\neq 0$. Then, for all $-\infty <{\mathcal{c}}<\infty $, it holds that\newline (i) $\lim_{m\rightarrow \infty }P\{\tau _{m}^{\ast }={\mathcal{n}}\}=\exp \left( -e^{-{\mathcal{c}}}\right) $. (ii) If in addition Assumption (ref) also holds, then we have $ \lim_{m\rightarrow \infty }P\{\bar{\tau}_{m}^{\ast }={\mathcal{n}}\}=\exp \left( -e^{-{\mathcal{c}}}\right) .$

Theorem (ref) stipulates that the asymptotic critical values, for a given nominal level $\alpha $, can be calculated as

equation[equation omitted — 149 chars of source]

using ${\mathcal{c}}_{\alpha }$\ or $\overline{{\mathcal{c}}}_{\alpha }$\ according as ((ref)) or ((ref)) is employed. Although the theorem offers an explicit formula to compute asymptotic critical values, these are bound to be inaccurate due to the slow convergence to the Extreme Value distribution. In particular, simulations show that, in finite samples, asymptotic critical values overstate the true values thus leading to low power.

Asymptotics under the alternative

We now study the behaviour of our monitoring schemes under the alternative, focussing, in particular, on the limiting distribution of the detection delay. We report the limiting distribution of the detection delay when using $0\leq \eta <1$ in Section (ref); in Section (ref), we report the limiting distribution of the detection delay when using $\eta >1$.

In both cases, under the alternative $H_{A}$, of ((ref)), the parameter $ \mbox{\boldmath${\theta}$}_{0}=(\alpha _{0},\beta _{0},\omega _{0})^{\top }$ changes to $\mbox{\boldmath${\theta}$}_{A}=(\alpha _{A},\beta _{A},\omega _{A})^{\top }$ satisfying

assumption\;$\alpha_A>0, \beta_A>0, \omega_0>0$, $\bar{ \mbox{\boldmath${\theta}$}}_0 \neq \bar{\mbox{\boldmath${\theta}$}}_A$, $ \bar{\mbox{\boldmath${\theta}$}}_0=(\alpha_0,\beta_0)^\top$ and $\bar{ \mbox{\boldmath${\theta}$}}_A=(\alpha_A, \beta_A)^\top$.

Detection delays with $0\leq \protect\eta <1$

We begin by investigating the asymptotic behaviour of the stopping time $ \tau _{m}$ defined in ((ref)). Whilst the result in Theorem (ref) below is valid for all cases, we need to introduce some preliminary notation, separately, for the two cases: (1) when the sequence is stationary after the change and (2) when the sequence is explosive after the change.\newline

\qquad Preliminary notation\newline

We begin by introducing some preliminary notation for the former case, i.e.

equation[equation omitted — 81 chars of source]

Under ((ref)), after the change the observations are exponentially close to $\{\hat{x}_{i},-\infty <i<\infty \}$, a stationary GARCH (1,1) sequence given by

equation[equation omitted — 169 chars of source]

We also define the log likelihood function

equation*[equation* omitted — 179 chars of source]

where $\hat{h}_{i}^{2}(\mbox{\boldmath${\theta}$})=\omega +\alpha \hat{x} _{i-1}^{2}+\beta \hat{h}_{i-1}^{2}(\mbox{\boldmath${\theta}$})$. Let

equation*[equation* omitted — 227 chars of source]

and define the size of the change as

equation[equation omitted — 129 chars of source]

We define the covariance matrix

equation[equation omitted — 475 chars of source]

and

equation[equation omitted — 165 chars of source]

where $A_{m}=\mbox{\boldmath${\Delta}$}^{\top }\hat{\mathbf{D}}_{m}^{-1} \mbox{\boldmath${\Delta}$}$. Finally (as far as preliminary notation is concerned), we define $u_{\mathcal{n}}>0$ as the unique solution of the equation

equation[equation omitted — 187 chars of source]

and $u^{\ast }>0$ as the solution of

equation[equation omitted — 86 chars of source]

It is easy to see that $u_{\mathcal{n}}\rightarrow u^{\ast }$, and that $ u^{\ast }=1$, if $t^{\ast }=0$. We are now ready to introduce the main notation to spell out the properties of the stopping time $\tau _{m}$ when the observations change into a stationary sequence.\newline

We now introduce the preliminary notation for the case when the observations turn into an explosive sequence after the time of change, i.e.

equation[equation omitted — 84 chars of source]

jensen2004asymptotic proved that

equation*[equation* omitted — 302 chars of source]

exists. Similarly to $\mbox{\boldmath${\Delta}$}$ in equation ((ref) ), we define the size of the change under $H_{A}$ as $\mbox{\boldmath${ \Upsilon}$}={\mathcal{r}}^{(2)}(\mbox{\boldmath${\theta}$}_{0})\neq \mathbf{0 }$. Similarly to ((ref)), we define

equation*[equation* omitted — 159 chars of source]

where $B_{m}=\mbox{\boldmath${\Upsilon}$}^{\top }\hat{\mathbf{D}}_{m}^{-1} \mbox{\boldmath${\Upsilon}$}$. Finally, similarly to $u_{\mathcal{n}}$ and $ u^{\ast }$ we define $\tilde{u}_{\mathcal{n}}$ and $\tilde{u}^{\ast }$ as the solutions of the equations

equation*[equation* omitted — 275 chars of source]

After the change in the parameters, the gradient of the likelihood function is approximated with the sequences

equation*[equation* omitted — 214 chars of source]

and

equation*[equation* omitted — 177 chars of source]

Analogously to $\mbox{\boldmath${\Sigma}$}_{1}$ in ((ref)), we finally introduce

equation[equation omitted — 349 chars of source]

\qquad Main notation\newline

After spelling out the preliminary notation for the two cases of the observations begin stationary and nonstationary, we now introduce the main notation. As we will see in Theorem (ref) below, in several cases the delay $\tau _{m}-k^{\ast }$ converges to a standard normal random variable after being centered and rescaled; the centering and rescaling for $\tau _{m}-k^{\ast }$ depend on whether $t^{\ast }<\infty $ or $t^{\ast }=\infty $ ($\tilde{t}^{\ast }<\infty $ or $\tilde{t}^{\ast }=\infty $, equivalently). In particular

enumerate• when $t^{\ast }<\infty $ ($\tilde{t}^{\ast }<\infty $, respectively), $ \tau _{m}-k^{\ast }$ will be centered around \begin{equation*} v_{1,{\mathcal{n}}}=\left\{ \begin{array}{ll} u_{\mathcal{n}}\left( {\mathcal{n}}^{1-\eta }\frac{ {\mathcal{c}}}{ A_{m}}\right) ^{1/(2-\eta )},\;\;\;if \;\;\;E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})<0, & \\ \tilde{u}_{\mathcal{n}}\left( {\mathcal{n}}^{1-\eta }\frac{ {\mathcal{c}}}{ B_{m}}\right) ^{1/(2-\eta )},\;\;\; if\;\;\;E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})>0, & \end{array} \right. \end{equation*} and rescaled by \begin{equation*} v_{2,{\mathcal{n}}}=\left\{ \begin{array}{ll} \left( t^{\ast }\boldmath${\Delta}$^{\top }\mathbf{D}^{-1} \boldmath${\Delta}$+\left( 1-\frac{ t^{\ast }}{ t^{\ast }+u^{\ast }}\right) \boldmath${\Delta}$ ^{\top }\mathbf{D}^{-1}\boldmath${\Sigma}$_{1}\mathbf{D}^{-1} \mbox{\boldmath${\Delta}$}\right) ^{1/2}(t^{\ast }+u^{\ast })^{3/2} & \\ \;\;\;\;\times \left( {\mathcal{n}}^{1-\eta } \frac{{\mathcal{c}}}{ A_{m}}\right) ^{1/(4-2\eta )},\;\;\;\mbox{if}\;\;\;E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})<0, & \\ \left( \tilde{t}^{\ast }\mbox{\boldmath${\Upsilon}$}^{\top } \mathbf{D}^{-1}\mbox{\boldmath${\Upsilon}$}+\left( 1-\frac{ \tilde{t}^{\ast }}{\tilde{t}^{\ast }+\tilde{u} ^{\ast }}\right) \mbox{\boldmath${\Upsilon}$}^{\top }\mathbf{D}^{-1} \mbox{\boldmath${\Sigma}$}_{2}\mathbf{D}^{-1}\mbox{\boldmath${\Upsilon}$} \right) ^{1/2}(\tilde{t}^{\ast }+\tilde{u}^{\ast })^{3/2} & \\ \;\;\;\;\times \left( { \mathcal{n}}^{1-\eta }\frac{{\mathcal{c}}}{ B_{m}} \right) ^{1/(4-2\eta )},\;\;\;\mbox{if}\;\;\;E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})>0. & \end{array} \right. \end{equation*} • when $t^{\ast }=\infty $ ($\tilde{t}^{\ast }<\infty $, respectively), $ \tau _{m}-k^{\ast }$ will be centered around \begin{equation*} v_{3,{\mathcal{n}}}=\left\{ \begin{array}{ll} \left( \frac{{\mathcal{c}}}{ A_{m}} {\mathcal{n}}^{1-\eta }(k^{\ast })^{\eta }\right) ^{1/2},\;\;\; \mbox{if}\;\;\;E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})<0, & \\ \left( \frac{{\mathcal{c}}}{ B_{m}}{\mathcal{n}}^{1-\eta }(k^{\ast })^{\eta }\right) ^{1/2},\;\;\;\mbox{if}\;\;\;E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})>0, & \end{array} \right. \end{equation*} and rescaled by \begin{equation*} v_{4,{\mathcal{n}}}=\left\{ \begin{array}{ll} \frac{(\mbox{\boldmath${\Delta}$}^{\top }\mathbf{D} ^{-1}\mbox{\boldmath${\Delta}$})^{1/2}}{ A_{m}}(k^{\ast })^{1/2},\;\;\;\mbox{if}\;\;\;E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})<0, & \\ & \\ \frac{(\mbox{\boldmath${\Upsilon}$}^{\top }\mathbf{ D}^{-1}\mbox{\boldmath${\Upsilon}$})^{1/2}}{ B_{m}}(k^{\ast })^{1/2},\;\;\;\mbox{if}\;\;\;E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})>0. & \end{array} \right. \end{equation*}

In order to present the limiting distribution for both cases, we define: the Gaussian process $\{\mbox{\boldmath${\Gamma}$}(t),t\geq 0\}$, with $E \mbox{\boldmath${\Gamma}$}(t)=\mathbf{0}$ and $E\mbox{\boldmath${\Gamma}$}(t) \mbox{\boldmath${\Gamma}$}^{\top }(s)=\min (t,s)\mathbf{D}$; the random variables

equation[equation omitted — 156 chars of source]
equation[equation omitted — 610 chars of source]

and, finally, the asymptotic variances

equation[equation omitted — 929 chars of source]
equation[equation omitted — 414 chars of source]

\newline

Let ${\mathcal{N}}$ denote a standard normal random variable, and let

equation*[equation* omitted — 84 chars of source]
theoremWe assume that $H_{A}$ of ((ref)) and Assumptions (ref)--(ref) and (ref) hold, and that $E\log (\alpha _{0}\epsilon _{0}^{2}+\beta _{0})\neq 0$. Then, for all $0\leq \eta <1$ \newline (i) If \begin{equation} \lim_{m\rightarrow \infty }\frac{k^{\ast }}{{\mathcal{n}}^{(1-\eta )/(2-\eta )}}<\infty , \end{equation} then \begin{equation*} \frac{\tau _{m}-k^{\ast }-v_{1,{\mathcal{n}}}}{v_{2,{\mathcal{n}}}}\overset{{ \mathcal{D}}}{\rightarrow }\bar{s}_{1}{\mathcal{N}}(0,1). \end{equation*} (ii) If \begin{equation} \lim_{m\rightarrow \infty }\frac{k^{\ast }}{{\mathcal{n}}^{(1-\eta )/(2-\eta )}}=\infty , \end{equation} and $\bar{u}=0$, then \begin{equation*} \frac{\tau _{m}-k^{\ast }-v_{3,{\mathcal{n}}}}{v_{4,{\mathcal{n}}}}\overset{{ \mathcal{D}}}{\rightarrow }\bar{s}_{2}{\mathcal{N}}(0,1). \end{equation*} (iii) If $0<\bar{u}<1$, then \begin{equation*} \lim_{m\rightarrow \infty }P\left\{ \frac{\tau _{m}-k^{\ast }}{(k^{\ast })^{1/2}}>x\right\} =P\left\{ \bar{u}^{1-\eta }\max \left( {\mathcal{A}}_{1}, {\mathcal{A}}_{2}(x)\right) \leq {\mathcal{c}}\right\} . \end{equation*}

In order to understand the practical implications of Theorem (ref), note that (up to some positive and finite constant)

equation*[equation* omitted — 197 chars of source]
equation*[equation* omitted — 175 chars of source]

The case ((ref)) corresponds to a \textquotedblleft very early\textquotedblright\ break; in this case, Theorem (ref) states that the expected delay is approximately $v_{1,{\mathcal{n}}}$, i.e. that it is approximately equal to $\left( m^{1-\eta }\right) ^{1/(2-\eta )}$. Clearly, as $\eta $ increases, $v_{1,{\mathcal{n}}}$ decreases; the dispersion around the expected delay, measured by $v_{2,{\mathcal{n}}}$, also decreases, indicating that the choice of $\eta $ plays a role in determining the delay in detecting (very early) changepoints, and that larger values of $\eta $ reduce such a delay. In the presence of an \textquotedblleft early, but not so early\textquotedblright\ break - corresponding to case (ii) of the theorem, where recall that $ k^{\ast }=o\left( {\mathcal{n}}\right) $, the expected delay $v_{3,{\mathcal{ n}}}$ still decreases as $\eta $ increases, as long as $k^{\ast }={\mathcal{n }}^{\gamma }$, for $\gamma >1/(2-\eta )$, but the dispersion around the expected delay - given by the standardization $v_{4,{\mathcal{n}}}$ - does not depend on $\eta $. Finally, the case of a late(r) change is studied in part (iii) of the theorem: in such a case, $\eta $ - and therefore the weight function in the definition of the detector - does not play any role.

Detection delays when $\protect\eta >1$

We now investigate the asymptotic behaviour of the stopping time $\tilde{\tau }_{m}$ defined in ((ref)) - that is, when the detector is a R\'{e} nyi type statistic with $\eta >1$. In such a case, the asymptotic behaviour of the detection delay uses the same notation irrrespective of whether $ E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})<0$ or $>0$.

Let

equation*[equation* omitted — 100 chars of source]

We define two independent normal random vectors $\mathbf{N}_{1}$ and $ \mathbf{N}_{2}$ such that $E\mathbf{N}_{1}=\mathbf{0},E\mathbf{N}_{2}= \mathbf{0},E\mathbf{N}_{1}\mathbf{N}_{1}^{\top }=\mathbf{D}$ and

equation*[equation* omitted — 332 chars of source]

Similarly to ${\mathcal{A}}_{1}$ and ${\mathcal{A}}_{2}(x)$ in ((ref)) and ((ref)), we define

equation[equation omitted — 171 chars of source]
equation[equation omitted — 905 chars of source]

and the centering sequence

equation*[equation* omitted — 599 chars of source]

and the rescaling sequence $v_{6,{\mathcal{n}}}=(k^{\ast })^{1/2}$.

theoremWe assume that $H_{A}$ of ((ref)) and Assumptions (ref)--(ref) and (ref) hold, and that $E\log (\alpha _{0}\epsilon _{0}^{2}+\beta _{0})\neq 0$. Then, for all $1<\eta <2$.\newline (i) If $k^{\ast }\leq r$ and ${\mathcal{a}}=0$ hold, then \begin{equation*} \lim_{m\rightarrow \infty }P\{\bar{\tau}_{m}=r\}=1. \end{equation*} (ii) If $k^{\ast }\leq r$ and ${\mathcal{a}}>0$ hold, then \begin{equation*} \lim_{m\rightarrow \infty }P\{\bar{\tau}_{m}=r\}=P\left\{ ({\mathcal{a}} ^{1/2}\mathbf{N}_{1}+{\mathcal{a}}\boldmath${\Delta}$+\mathbf{N} _{2})^{\top }\mathbf{D}^{-1}({\mathcal{a}}^{1/2}\mathbf{N}_{1}+{\mathcal{a}} \boldmath${\Delta}$+\mathbf{N}_{2})>{\mathcal{c}}\right\} . \end{equation*} (iii) If $k^{\ast }>r$ and ${\mathcal{a}}<\infty $, then \begin{equation*} \lim_{m\rightarrow \infty }P\left\{ \bar{\tau}_{m}>k^{\ast }+xr^{1/2}\right\} =P\left\{ \max \left( {\mathcal{B}}_{1},{\mathcal{B}} _{2}(x)\right) \leq {\mathcal{c}}\right\} . \end{equation*} (iv) If $k^{\ast }>r,\bar{u}<1$ and ${\mathcal{a}}=\infty $ hold, then \begin{equation*} \frac{\bar{\tau}_{m}-k^{\ast }-v_{5,{\mathcal{n}}}}{v_{6,{\mathcal{n}}}} \overset{{\mathcal{D}}}{\rightarrow }\bar{s}_{2}{\mathcal{N}}. \end{equation*}

Similarly to Theorem (ref), Theorem (ref) describes the detection delay when using R\'{e}nyi type statistics depending on the location of the break; to the best of our knowledge, this is the first time such a result has ever been derived. Part (i) of the theorem is also derived in ghezzi2024fast, and, in essence, it states that if the break occurs prior to the trimming sequence $r$ in the R\'{e}nyi type statistics, then it is identified straight at $r$ - that is, as soon as the R \'{e}nyi type statistics starts the monitoring. Parts (ii) and (iii) of the theorem refine and extend the results in ghezzi2024fast. Finally, part (iv) states that, in the case of a break occurring late - or, better, much later than $r$ - a large value of $\eta $ could even be detrimental because the centering sequence $v_{5,{ \mathcal{n}}}$ diverges with $k^{\ast }$, at a faster rate as $\tau $ increases. This confirms the common wisdom (see kirch2022asymptotic and kirch2022sequential), and the findings in ghezzi2024fast, that R\'{e}nyi type statistics are designed for the fast detection of very early occurring breaks, whereas they may yield suboptimal results for later breaks.

Simulations

In this section, we assess the finite sample performance of our monitoring procedures via Monte Carlo simulations. According to the theory in Section (ref), we can have two classes of monitoring schemes, based on $ \eta \neq 1$ (covered by Theorem (ref)) and $\eta=1$ (covered by Theorem (ref)). For the sake of brevity, here we only focus on the case $\eta \neq 1$. We consider several data generating processes (DGP). We use three lengths of the historical training sample $m=500,1000,5000$ and two lengths of the monitoring $\mathcal{n}=250,500$. The sequential procedure is performed $5,000$ times with independently generated samples, and the percentage of simulations for which the detector crosses the boundary functions is reported for several values of $\eta $. For the R\'{e}nyi type statistic based on Theorem (ref)(ii), we follow horvath2021 and set $r=\sqrt{\mathcal{n}}$. Guidelines on implementation are provided in Section (ref) of the Supplement. To obtain critical values, we simulate two independent standard Wiener processes $W_{1}(t)$ and $W_{2}(t)$ on a grid of $100,000$ equally spaced points in the unit interval $\left[ 0,1\right] $ and compute $\sup_{0<t\leq 1}\frac{1}{t^{\eta }}\left( W_{1}^{2}(t)+W_{2}^{2}(t)\right) $ and $ \sup_{0<t\leq 1}\frac{1}{t^{1-\eta }}\left( W_{1}^{2}(t)+W_{2}^{2}(t)\right) $. We repeat this by $100,000$ times and obtain the empirical $90\%$, $95\%$ , and $99\%$ percentiles of the above two quantities, corresponding to the critical values at $10\%$, $5\%$, and $1\%$ levels based on Theorem (ref) (i) and (ii). Critical values are in Table (ref).

table[table omitted — 797 chars of source]

The boundary functions in Section (ref) are designed for the case $ m\rightarrow \infty $. However, preliminary simulations show that the empirical sizes based on those boundary functions tend to be larger than the nominal levels in finite samples, in particular for DGPs with the Student's $ t$ errors. To make our monitoring schemes more practical under small finite samples, we suggest to \textquotedblleft tune\textquotedblright\ the boundary functions as

align[align omitted — 447 chars of source]

for ((ref)) and ((ref)) respectively. The intuition underpinning the term $\left( 1+1/\log (m)\right) ^{2}$ is to boost the boundary function in small samples. The term $\left( 1+k/m\right) ^{2}$ as is typically employed when the monitoring horizon is \textquotedblleft long\textquotedblright\ (see horvath2020sequential, horvath2021monitoring, and zenhya). Although this term is inconsequential for the asymptotic theory in our set-up, we find that it can further improve the empirical size. Both tuning terms are asymptotically negligible. and only play a role in finite samples to achieve better size control at no expense for power. The proposed tuning is tailored to DGPs with Student's $t$ errors, rather than Gaussian errors; indeed, heavy tails are a well-known stylised fact of financial returns.

Empirical size under the null

Under the null hypothesis, the realisation of GARCH(1,1) in the historical training period ($1\leqslant i\leqslant m$) and in the monitoring period ($ m+1\leqslant i\leqslant m+\mathcal{n}$) is

equation*[equation* omitted — 182 chars of source]

where $\epsilon _{i}$ follows a standard normal distribution or the Student's $t$ distribution with $7$ degrees of freedom. Since our monitoring procedure does not require the historical sample to be stationary or not, we choose the following two set of GARCH(1,1) parameters, taken from francq2012strict: (i) $(\omega _{0},\alpha _{0},\beta _{0})=(0.10,0.18,0.80)$, which represents the stationary case; (ii) $(\omega _{0},\alpha _{0},\beta _{0})=(0.10,0.30,0.80)$, corresponding to the nonstationary case since $E\log (\alpha \epsilon _{0}^{2}+\beta )>0$ under errors following either the standard normal or the Student's $t$ distributions. \newline

Table (ref) reports the empirical sizes at $5\%$ significance level for the monitoring scheme based on Theorem (ref)( i) for different values of $\eta $. A noticeable feature is that a larger $\eta $ results in a higher rejection rates, and a smaller $\eta $ is more conservative in rejection. Under the Student's $t$ errors, $\eta =0.3$ is a good choice because the monitoring procedure has reasonably good empirical sizes when $m=1,000$, and the empirical sizes for $m=5,000$ are closer to the theoretical level of $5\%$. Under Gaussian errors, the monitoring procedure is slightly under-sized, which is mainly due to the additionally tuning we imposed in (ref). For practical use, the tendency to under-reject with Gaussian errors may not necessarily be a concern, because the empirical power does not seem to be affected, as shown in Section (ref). Lastly, the simulation results show that our monitoring schemed works reasonably well for both stationary and nonstationary GARCH(1,1) models.\newline

table[table omitted — 2,075 chars of source]

Table (ref) in the Supplement contains the empirical sizes for R\'{e}nyi type statistics. The rejection rates are slightly higher than the $5\%$ nominal level. We note that, in principle, it would be possible to design a different tuning for R\'{e}nyi type statistics.

Empirical power under $H_{A}$

We now turn to the analysis of the empirical power. Under the alternative, the data is generated by

equation*[equation* omitted — 48 chars of source]

and

equation*[equation* omitted — 284 chars of source]

where the parameter $\boldmath{\theta }_{0}$$=(\alpha _{0},\beta _{0},\omega _{0})^{\top }$ changes to $\boldmath{\theta }_{A}$$=(\alpha _{A},\beta _{A},\omega _{A})^{\top }$ at time $m+k^{\ast }$. We consider two scenarios for the time of change: (a) $k^{\ast }=\lfloor \sqrt{\mathcal{n}} \rfloor $ corresponds to a change occurring \textquotedblleft early, but not too early\textquotedblright\ after the historical sample; (b) $ k^{\ast }=0.5{\mathcal{n}}$ indicates a change happening much later than $r$ . \newline

There are many possible ways of changes under the alternative. To keep our results clean, we set $\omega _{0}=\omega _{A}=0.1$ and $\alpha _{0}=\alpha _{A}=0.18$, and concentrate on a change in $\beta $ under the following four representative alternatives:

itemize$\beta_0=0.8, \beta_A=0.6$, i.e. a change from a stationary to another stationary regime, • $\beta_0=0.8, \beta_A=0.9$, i.e. a change from a stationary to an explosive regime, • $\beta_0=0.9, \beta_A=0.8$, i.e. a change from an explosive to a stationary regime, • $\beta_0=0.9, \beta_A=1.0$, i.e. a change from an explosive to another explosive regime.\newline

Tables (ref) and (ref) show the empirical power of the monitoring scheme based on Theorem (ref)( i) at $5\%$ significance level when $\mathcal{n}=500$ for a change at $k^{\ast }=\lfloor \sqrt{\mathcal{n}}\rfloor $ and $k^{\ast }=\lfloor 0.5 \mathcal{n}\rfloor $, respectively.\footnote{ The empirical power of $\mathcal{n}=250$ (not reported) is marginally lower than the empirical power of $\mathcal{n}=500$.} There are five major observations. First, our monitoring scheme is highly effective in detecting changes under $H_{A,2}$ and $H_{A,3}$ for both early and late changes. These alternatives result in a change between a stationary regime and an explosive regime, which is relatively easy to detect. Second, the monitoring scheme exhibits high power in detecting early changes under $H_{A,1}$ and $H_{A,4}$ . These alternatives represent a change within either a stationary or an explosive regime. Third, there is a deterioration in power when detecting late changes under $H_{A,1}$ and $H_{A,4}$, although satisfactory levels can be achieved by using a large(r) training sample size of $m=5,000$. Fourth, the power is relatively lower when using the Student's $t$ distribution errors compared to normal errors. Lastly, there is only a marginal decline observed in the power with a larger value of $\eta $.\newline

table[table omitted — 3,058 chars of source]
table[table omitted — 2,984 chars of source]

Tables (ref) and (ref) in the Supplement provide the empirical power for the R\'{e}nyi type statistics based on Theorem (ref)(ii) under the same setting. When detecting early changes at $k^{\ast }=\lfloor \sqrt{\mathcal{n}}\rfloor $, similar observations as above apply; the monitoring schemes with $\eta =1.3$ and $1.5$ proves to be effective. However, one noticeable difference is that a larger value of $\eta $ is detrimental in the power. In particular, $\eta =1.7$ and $2$ suffer a remarkable loss of power under $H_{A,1}$. As far as late changes ($k^{\ast }=\lfloor 0.5\mathcal{n}\rfloor $) are concerned, the R\'{e}nyi type statistics become much less effective, as predicted by the theory. This is because R\'{e}nyi type statistics are devised for the fast detection of very early changes, whilst being suboptimal for late changes. \newline

It is also worthwhile to examine the stopping time $\tau _{m}$ and $\bar{\tau }_{m}$ in order to investigate the detection delays of our monitoring procedures. Figure (ref) shows the boxplot of the detection delays of $\tau _{m}$ and $\bar{\tau}_{m}$ for a change at $k^{\ast }=\lfloor \sqrt{\mathcal{n}}\rfloor $ under $H_{A,2}$ when $m=500$, $ \mathcal{n}=500$. For the monitoring procedure based on Theorem (ref)( i), it is consistent with our theory that larger values of $\eta $ reduce the detection delay. Considering the R\'{e}nyi type statistics based on Theorem (ref)(ii), there is only a marginal difference in using various values of $\eta $. Comparing the detection delay between the monitoring procedures based on Theorem (ref)(i) and (ii ), we can clearly see the merit of the R\'{e}nyi type statistics for the fast detection of early changes, as evidenced by shorter detection delays.

figure[figure omitted — 453 chars of source]

Empirical illustration

We illustrate our monitoring procedures using daily returns of individual stocks. We focus on four stocks: Apple Inc. (ticker: AAPL, Permno: 14593), Middlefield Banc Corp. (ticker: MBCN, Permno: 14932), Genetic Technologies Ltd (ticker: GENE, Permno: 90899), and NTS Realty Holdings LP (ticker: NLP, Permno: 90508). We download daily returns (without dividend) from the CRSP database.\footnote{ We choose to use the daily returns without dividend, rather than log difference of prices, to avoid the complication due to stock splits.} We consider two periods in order to showcase the detection for four types of changes (in the sense of the four different alternatives in our simulation). Depending on the specific purpose of the researcher, one can choose between the monitoring procedures based on Theorem (ref)(i) and ( ii). Based on our simulations,if the aim is to quickly detect very early changes, we suggest using the R\'{e}nyi type statistics based on Theorem (ref)(ii), with some tolerance for the compromise in size and power; conversely, if the purpose is to have good size control and high power, it is recommended to use the monitoring procedure based on Theorem (ref)(i). In this application, our preference is to have a good balance of size and power, and the procedure based on Theorem (ref)(i) (with the choice of $\eta =0.3$) delivers a good performance with sample sizes similar to the dataset used in this section. \footnote{ We relegate the results using R\'{e}nyi weights in Section (ref) of the Supplement.} Before applying our monitoring procedure, it is necessary to ensure there is no change during the historical training sample. To this end, we use the test developed by horvath2024detecting (horvath2024detecting, labeled as HW(2024) hereinafter). Their test is to detect changes in GARCH(1,1) processes without assuming stationarity, which can accommodate either stationary or nonstationary historical sample. A rejection of their test indicates there is no change of $(\alpha ,\beta )$ in the GARCH(1,1) during the historical sample. At the same time, we are keen to understand which the type of four changes may occur. Thus, we firstly examine whether our historical sample is stationary or not by employing the nonstationarity test developed by francq2012strict (francq2012strict, labeled as FZ(2012) hereinafter). At the end of our monitoring horizon, we use the FZ(2012) test again to check the stationarity of the samples after the change (if there is one).

Change from a stationary regime

To illustrate changepoint detection from a stationary regime, we choose the training period of 2016--2019 (1007 trading days) and the monitoring period of 2020--2021 (507 trading days). The training period is before the outbreak of COVID-19, while the monitoring period is in the pandemic. We apply our monitoring procedure for the stocks of AAPL and MBCN during this period. Table (ref) (Columns 1 and 2) reports the results of the sequential monitoring procedure, as well as other information, including HW(2024) test, FZ(2012) test, and parameter estimates. HW(2024) test indicates that there is no parameter change during the training sample for AAPL and MBCN. The nonstationarity test of FZ(2012) indicates that they are both stationary during the training sample. Our sequential monitoring detects a change of AAPL on July 31st, 2020 and a change of MBCN on May 8th, 2020. Based on the FZ(2012) test for the sample after the change, we can conclude that AAPL experienced a change from a stationary to another stationary regime, while MBCN shifted from a stationary regime to a nonstationary one. Figure (ref) contains returns series (upper panel) during the monitoring period and the detector versus the boundary function (lower panel).

table[table omitted — 1,805 chars of source]
figure[figure omitted — 373 chars of source]

Change from a nonstationary regime

We now consider detection from a nonstationary regime, and use 2007--2010 (1011 trading days) as the training period and 2011-2022 (505 trading days) as the monitoring period. The training period covers the global financial crisis (GFC), while the monitoring period follows the GFC but includes the European debt crisis. In this period, we monitor GENE and NLP. The results of the sequential monitoring procedure, alongside other supporting information, are displayed in Columns 3 and 4 of Table (ref). Based on the HW(2024) test, we cannot reject that the return series of GENE and NLP have change in the training period. As evidenced by the nonstationarity test of FZ(2012), both stocks are in the nonstationary regime during the training period. Our sequential monitoring procedure reveals a change of GENE on April 27th, 2011 and a change of NLP on September 4th, 2012. After applying FZ(2012) nonstationarity test on the sample after the change, it is found that the change of GENE is from a nonstationary regime to a stationary regime, whilst the change of NLP is from a nonstationary to another nonstationary regime. It is also interesting to note that GENE after the change is in a strict stationary regime, but not in a second-order stationary regime. Figure (ref) shows their returns series (upper panel) during the monitoring period and the detector versus the boundary function (lower panel).

figure[figure omitted — 371 chars of source]

Conclusions and discussions

In this paper, we complement the existing literature on (ex-ante) testing for bubble phenomena by proposing a family of weighted, CUSUM-based statistics to detect changes in the parameters of a GARCH(1,1) process. Our monitoring procedure can be applied irrespective of whether, in the training sample, the observations are stationary or explosive, and it is able to detect all types of changes: (a) from a stationary to another stationary regime (which is helpful in order to avoid the issues concerning the consistent estimation of a GARCH(1,1) process spelt out in hillebrand2005neglecting); (b) from a stationary to an explosive regime (which contains information on the possible inception of a bubble); (c) from an explosive to a stationary regime (which, in the light of the previous point, could shed light on the cooling off the turbulence associated with a bubble on a financial market); and (d) from an explosive to another explosive regime (which, depending on the direction of the change - towards a more or a less explosive regime - could indicate whether exuberant volatility is heating up or cooling down). Technically, we propose two families of statistics, both based on weighted versions of the CUSUM process of the quasi-Fisher scores: one family uses lighter weights, and it is designed to detect, optimally, changes occurring not immediately after the start of the monitoring horizon; the other family uses heavier, R\'{e} nyi-type weights, which make it more sensitive to changepoints occurring immediately after the end of the training period. For both cases, we study the limiting distribution of the detection delays; to the best of our knowledge, no such results exist for the case of a GARCH(1,1) models, and no results in general exist for the case of R\'{e}nyi statistics. Given the interest in the detection of bubble phenomena, and the scant amount of contributions in the context of detection of changes in the volatility, we believe that our paper should be a useful addition to the toolbox of the financial econometrician.

{ {\ } }

\setcounter{section}{0} \setcounter{subsection}{-1} \setcounter{subsubsection}{-1} \setcounter{equation}{0} \setcounter{lemma}{0} \setcounter{theorem}{0}

Implementation guidelines, further Monte Carlo evidence, and further empirical results

Practical guidance on the implementation

In this section, we provide a detailed step-by-step guidance on implementing our sequential monitoring procedures (based on Theorem (ref)), which could be useful for researchers who want to apply it without studying the underlying theory. The steps are as follows:

enumerate• Use QMLE to estimate the parameters $\boldsymbol{\theta }$ using the training sample \begin{equation*} \hat{\boldsymbol{\theta }}_{m}=\rm argmax\left\{ \sum_{i=1}^{m}\bar{ \ell}_{i}(\boldsymbol{\theta }):\;\boldsymbol{\theta }\in \boldsymbol{\Theta }\right\} , \end{equation*} where \begin{equation*} \bar{\ell}_{i}(\boldsymbol{\theta })=\log \bar{\sigma}_{i}^{2}+\dfrac{ y_{i}^{2}}{\bar{\sigma}_{i}^{2}}, \end{equation*} and \begin{equation*} \bar{\sigma}_{i}^{2}(\boldsymbol{\theta })=\omega +\alpha y_{i-1}+\beta \bar{ \sigma}_{i-1}^{2}(\boldsymbol{\theta }),\;\;\;\;\;1\leq i\leq m. \end{equation*} • Calculate the quasi-Fisher scores of the log likelihood function \begin{equation*} \frac{\partial \bar{\ell}_{i}(\hat{\boldsymbol{\theta }}_{m})}{\partial \alpha }=\frac{\partial \bar{\ell}_{i}(\hat{\boldsymbol{\theta }}_{m})}{ \partial \bar{\sigma}_{i}^{2}(\hat{\boldsymbol{\theta }}_{m})}\frac{\partial \bar{\sigma}_{i}^{2}(\hat{\boldsymbol{\theta }}_{m})}{\partial \alpha }= \frac{1}{\bar{\sigma}_{i}^{2}(\hat{\boldsymbol{\theta }}_{m})}\left( 1-\frac{ y_{i}^{2}}{\bar{\sigma}_{i}^{2}(\hat{\boldsymbol{\theta }}_{m})}\right) \frac{\partial \bar{\sigma}_{i}^{2}(\hat{\boldsymbol{\theta }}_{m})}{ \partial \alpha },\;\;\;\;\;1\leq i\leq m+\mathcal{n}, \end{equation*} \begin{equation*} \frac{\partial \bar{\ell}_{i}(\hat{\boldsymbol{\theta }}_{m})}{\partial \beta }=\frac{\partial \bar{\ell}_{i}(\hat{\boldsymbol{\theta }}_{m})}{ \partial \bar{\sigma}_{i}^{2}(\hat{\boldsymbol{\theta }}_{m})}\frac{\partial \bar{\sigma}_{i}^{2}(\hat{\boldsymbol{\theta }}_{m})}{\partial \beta }=\frac{ 1}{\bar{\sigma}_{i}^{2}(\hat{\boldsymbol{\theta }}_{m})}\left( 1-\frac{ y_{i}^{2}}{\bar{\sigma}_{i}^{2}(\hat{\boldsymbol{\theta }}_{m})}\right) \frac{\partial \bar{\sigma}_{i}^{2}(\hat{\boldsymbol{\theta }}_{m})}{ \partial \beta },\;\;\;\;\;1\leq i\leq m+\mathcal{n}. \end{equation*} To implement the above calculation, we need the following recursive equations: \begin{equation*} \frac{\partial \bar{\sigma}_{i}^{2}(\hat{\boldsymbol{\theta }}_{m})}{ \partial \alpha }=y_{i-1}^{2}+\beta \frac{\partial \bar{\sigma}_{i-1}^{2}( \hat{\boldsymbol{\theta }}_{m})}{\partial \alpha }, \end{equation*} \begin{equation*} \frac{\partial \bar{\sigma}_{i}^{2}(\hat{\boldsymbol{\theta }}_{m})}{ \partial \beta }=\bar{\sigma}_{i-1}^{2}(\hat{\boldsymbol{\theta }} _{m})+\beta \frac{\partial \bar{\sigma}_{i-1}^{2}(\hat{\boldsymbol{\theta }} _{m})}{\partial \beta }, \end{equation*} with initial values set to \begin{equation*} \frac{\partial \bar{\sigma}_{0}^{2}(\hat{\boldsymbol{\theta }}_{m})}{ \partial \alpha }=\frac{\partial \bar{\sigma}_{0}^{2}(\hat{\boldsymbol{ \theta }}_{m})}{\partial \beta }=0, \end{equation*} which follows fiorentini1996analytic and pelagatti2009variance. • Estimate the covariance matrix of the derivatives in the training sample \begin{equation*} \hat{\mathbf{D}}_{m}=\dfrac{1}{m}\sum_{i=i}^{m}\left( \frac{\partial \bar{ \ell}_{i}(\hat{\boldsymbol{\theta }}_{m})}{\partial \alpha },\frac{\partial \bar{\ell}_{i}(\hat{\boldsymbol{\theta }}_{m})}{\partial \beta }\right) ^{\top }\left( \frac{\partial \bar{\ell}_{i}(\hat{\boldsymbol{\theta }}_{m}) }{\partial \alpha },\frac{\partial \bar{\ell}_{i}(\hat{\boldsymbol{\theta }} _{m})}{\partial \beta }\right) . \end{equation*} • Obtain the detector \begin{equation*} \mathcal{D}_{m}(k)=\mathcal{r}_{m,k}^{\top }(\hat{\boldsymbol{\theta }}_{m}) \hat{\mathbf{D}}_{m}^{-1}\mathcal{r}_{m,k}(\hat{\boldsymbol{\theta }} _{m}),\;\;\;\;\;m\leq k<m+\mathcal{n}, \end{equation*} where \begin{equation*} \mathcal{r}_{m,k}(\boldsymbol{\theta })=\sum_{i=m+1}^{m+k}\left( \frac{ \partial \bar{\ell}_{i}(\boldsymbol{\theta })}{\partial \alpha },\frac{ \partial \bar{\ell}_{i}(\boldsymbol{\theta })}{\partial \beta }\right) ^{\top }. \end{equation*} • Based on Theorem (ref)(i), a break is flagged as soon as the detector $\mathcal{D}_{m}(k)$ exceeds the tuned boundary function \begin{equation*} \mathfrak{g}_{m}(k)=\mathcal{c}\mathcal{n}\left( 1+\dfrac{1}{\log (m)} \right) ^{2}\left( 1+\dfrac{k}{m}\right) ^{2}\left( \dfrac{k}{\mathcal{n}} \right) ^{\eta },\quad \;\;where\;\;1\leq k\leqslant \mathcal{n}-1; \end{equation*} based on Theorem (ref)(ii), a break is indicated as soon as the detector $\mathcal{D}_{m}(k)$ goes above the tuned boundary function \begin{equation*} \bar{\mathfrak{g}}_{m}(k)=\mathcal{c}r\left( 1+\dfrac{1}{\log (m)}\right) ^{2}\left( 1+\dfrac{k}{m}\right) ^{2}\left( \frac{k}{r}\right) ^{\eta },\quad \;\;where\;\;r\leq k\leqslant \mathcal{n}-1; \end{equation*} otherwise, no break is reported.

Additional simulation results

table[table omitted — 2,066 chars of source]
table[table omitted — 3,032 chars of source]
table[table omitted — 2,977 chars of source]

Further results - empirical application using R\'{e}nyi weights

Figure (ref) presents the detector $\mathcal{D}_m(k)$ (with the choice of $\eta=1.5$) versus the boundary function $\bar{\mathfrak{g}} _m(k)$ based on Theorem 3.1 (ii) for the four stocks during the same periods analysed in Section 5. The monitoring procedure based on the R\'enyi type statistics detects a change of AAPL on March 2nd, 2020, a change of MBCN on March 13th, 2020, a change of GENE on April 27th, 2011, and a change of NLP on September 4th, 2012. Such results illustrate the merit of R\'enyi type statistics for the fast detection of very early changes, in particular for AAPL (nearly 5 months earlier) and MBCN (about 2 months earlier), compared to the procedure based on Theorem 3.1(i).

figure[figure omitted — 397 chars of source]

\setcounter{subsection}{-1} \setcounter{subsubsection}{-1} \setcounter{equation}{0} \setcounter{lemma}{0} \setcounter{theorem}{0}

Technical lemmas

Henceforth, we use $\nabla $ and $\nabla ^{2}$ for the gradient vector and the Hessian matrix. We begin with some facts and notation which we will use throughout the proofs.\newline

If (ref) holds, then there is a unique, stationary, non anticipative sequence $\{x_{i},-\infty <i<\infty \}$ satisfying

equation*[equation* omitted — 146 chars of source]

(cf. berkes2003garch, and francq2019garch). Next we define

equation*[equation* omitted — 140 chars of source]

$\mbox{\boldmath${\theta}$}=(\omega ,\alpha ,\beta )^{\top }$. The stationary version of the likelihood function is

equation*[equation* omitted — 169 chars of source]

and we define

equation*[equation* omitted — 120 chars of source]

Similarly, the stationary version of ${\mathcal{r}}_{m,k}( \mbox{\boldmath${\theta}$})$ is

equation[equation omitted — 275 chars of source]

berkes2003garch and francq2019garch showed that the matrix

equation*[equation* omitted — 86 chars of source]

exists and it is non singular.

After the change the observations are given by

equation[equation omitted — 180 chars of source]
lemmaIf Assumptions (ref) and (ref) hold, then we have \begin{equation*} m\left( \hat{\boldmath${\theta}$}_{m}-\boldmath${\theta}$ _{0}\right) =-{\mathbf{H}}^{-1}\nabla {{\mathcal{f}}}_{m}( \boldmath${\theta}$_{0})+O_{P}\left( \frac{1}{m}\right) . \end{equation*} \begin{proof} The lemma is taken from berkes2003garch and francq2019garch (cf.\ also Lemma 5.3 in horvath2024detecting). \end{proof}
lemmaIf Assumptions (ref)--(ref) hold, then we have \begin{equation} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{k^{\zeta }}\sup_{ \boldmath${\theta}$\in \boldmath${\theta}$}\left\Vert { \mathcal{r}}_{m,k}(\boldmath${\theta}$)-\bar{{\mathcal{r}}}_{m,k}( \boldmath${\theta}$)\right\Vert =O_{P}(1), \end{equation} with some $\zeta <1/2$ and \begin{equation} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}^{1/2}(k/{\mathcal{n}} )^{\gamma }}\sup_{\boldmath${\theta}$\in \boldmath${\theta}$ }\left\Vert {\mathcal{r}}_{m,k}(\mbox{\boldmath${\theta}$})-\bar{{\mathcal{r} }}_{m,k}(\mbox{\boldmath${\theta}$})\right\Vert =o_{P}(1), \end{equation} for all $0\leq \gamma <1/2$. \begin{proof} It is shown in berkes2003garch, francq2004maximum, and francq2019garch that there is a constant $0<\rho <1$ such that \begin{equation*} \sup_{\mbox{\boldmath${\theta}$}\in \mbox{\boldmath${\theta}$}}\left\vert \bar{\sigma}_{m+i}^{2}(\mbox{\boldmath${\theta}$})-\bar{h}_{m+i}^{2}( \mbox{\boldmath${\theta}$})\right\vert =O\left( \rho ^{i}\right) \quad \quad a.s. \end{equation*} \begin{equation*} \sup_{\mbox{\boldmath${\theta}$}\in \mbox{\boldmath${\theta}$}}\left\Vert \nabla \bar{\sigma}_{m+i}^{2}(\mbox{\boldmath${\theta}$})-\nabla \bar{h} _{m+i}^{2}(\mbox{\boldmath${\theta}$})\right\Vert =O\left( \rho ^{i}\right) \quad \quad a.s. \end{equation*} and \begin{equation*} \left\vert y_{m+i}^{2}-x_{m+i}^{2}\right\vert =O\left( \rho ^{i}\right) \quad \quad a.s. \end{equation*} Since $\underline{\omega }\leq \bar{\sigma}_{h}^{2}(\mbox{\boldmath${ \theta}$})$ and $\underline{\omega }\leq \bar{h}_{j}^{2}(\mbox{\boldmath${ \theta}$})$ for all $j\geq 1$, elementary calculations yield (ref). Now (ref) implies \begin{equation*} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}^{1/2}(k/{\mathcal{n}} )^{\gamma }}\sup_{\mbox{\boldmath${\theta}$}\in \mbox{\boldmath${\theta}$} }\left\Vert {\mathcal{r}}_{m,k}(\mbox{\boldmath${\theta}$})-\bar{{\mathcal{r} }}_{m,k}(\mbox{\boldmath${\theta}$})\right\Vert =O_{P}(1){\mathcal{n}} ^{-1/2+\gamma }\max_{1\leq k\leq {\mathcal{n}}}k^{-\gamma +\zeta }=o_{P}(1). \end{equation*} \end{proof}
lemmaIf Assumptions (ref)--(ref) hold, then we have \begin{equation*} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}^{1/2}(k/{\mathcal{n}} )^{\gamma }}\left\Vert {\mathcal{r}}_{m,k}(\hat{\boldmath${\theta}$} _{m})-\bar{{\mathcal{r}}}_{m,k}(\boldmath${\theta}$_{0})\right\Vert =o_{P}(1), \end{equation*} for all $0\leq \gamma <1/2$. \begin{proof} Lemma (ref) yields \begin{equation*} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}^{1/2}(k/{\mathcal{n}} )^{\gamma }}\sup_{\boldmath${\theta}$\in \boldmath${\theta}$ }\left\Vert {\mathcal{r}}_{m,k}(\hat{\boldmath${\theta}$}_{m})-\bar{{ \mathcal{r}}}_{m,k}(\hat{\boldmath${\theta}$}_{m})\right\Vert =o_{P}(1). \end{equation*} It is proven in berkes2003garch that there is a neighbourhood of $ \mbox{\boldmath${\theta}$}_{0}$, say $\mbox{\boldmath${\theta}$}_{0}$, such that \begin{equation*} \max_{1\leq k\leq {\mathcal{n}}}\sup_{\mbox{\boldmath${\theta}$}\in \mbox{\boldmath${\theta}$}_{0}}\frac{1}{k}\left\Vert \nabla ^{3}{\mathcal{f}} _{m,k}(\mbox{\boldmath${\theta}$})\right\Vert =O_{P}(1). \end{equation*} Using a two term Taylor expansion coordinate-wise, we conclude that \begin{equation} \nabla {\mathcal{f}}_{m,k}(\hat{\mbox{\boldmath${\theta}$}}_{m})-\nabla { \mathcal{f}}_{m,k}(\mbox{\boldmath${\theta}$}_{0})=\nabla ^{2}{\mathcal{f}} _{m,k}(\mbox{\boldmath${\theta}$}_{0})(\hat{\mbox{\boldmath${\theta}$}}_{m}- \mbox{\boldmath${\theta}$}_{0})+\mathbf{R}_{m,k}, \end{equation} and \begin{equation*} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}^{1/2}(k/{\mathcal{n}} )^{\gamma }}\Vert \mathbf{R}_{m,k}\Vert =O_{P}({\mathcal{n}}^{1/2})\Vert \hat{\mbox{\boldmath${\theta}$}}_{m}-\mbox{\boldmath${\theta}$}_{0}\Vert ^{2}. \end{equation*} The central limit for decomposable Bernoulli shifts (cf. horvath2024detecting) yields that \begin{equation} \Vert \nabla {\mathcal{f}}_{m}(\mbox{\boldmath${\theta}$}_{0})\Vert =O_{P}\left( m^{1/2}\right) , \end{equation} and therefore Lemma (ref) and Assumption (ref) give \begin{equation} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}^{1/2}(k/{\mathcal{n}} )^{\gamma }}\Vert \mathbf{R}_{m,k}\Vert =o_{P}(1). \end{equation} According to the ergodic theorem (cf. breiman1968) \begin{equation} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{k}\left\Vert \nabla ^{2}{\mathcal{f} }_{m,k}(\mbox{\boldmath${\theta}$}_{0})\right\Vert =O_{P}(1), \end{equation} and therefore Lemma (ref) implies \begin{align*} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}^{1/2}(k/{\mathcal{n}} )^{\gamma }}\left\Vert \nabla ^{2}{\mathcal{f}}_{m}(\mbox{\boldmath${ \theta}$}_{0})(\hat{\mbox{\boldmath${\theta}$}}_{m}-\mbox{\boldmath${ \theta}$}_{0})\right\Vert & =O_{P}\left( m^{-1/2}\right) \max_{1\leq k\leq { \mathcal{n}}}\frac{k}{{\mathcal{n}}^{1/2}(k/{\mathcal{n}})^{\gamma }} \\ & =O_{P}\left( \left( \frac{{\mathcal{n}}}{m}\right) ^{1/2}\right) =o_{P}(1). \end{align*} Given that $\bar{{\mathcal{r}}}_{m,k}(\mbox{\boldmath${\theta}$})$ is made of the first two coordinates of $\nabla {\mathcal{f}}_{m,k}( \mbox{\boldmath${\theta}$})$, the proof is complete. \end{proof}
lemmaIf Assumptions (ref)--(ref) hold, then, for any $m$ , we can a define Gaussian process $\{\mbox{\boldmath${\Gamma}$}_{m}(t)\in R^{2},t\geq 0\}$ with $E\mbox{\boldmath${\Gamma}$}_{m}(t)=\mathbf{0}$ and $E \mbox{\boldmath${\Gamma}$}_{m}(t)\mbox{\boldmath${\Gamma}$}_{m}(s)=\min (t,s) \mathbf{D}$, \begin{equation} \mathbf{D}=E\left[ \left( \frac{\partial \bar{f}_{1}(\boldmath${ \theta}$_{0})}{\partial \alpha },\frac{\partial \bar{f}_{1}( \boldmath${\theta}$_{0})}{\partial \beta }\right) ^{\top }\left( \frac{\partial \bar{f}_{1}(\boldmath${\theta}$_{0})}{\partial \alpha } ,\frac{\partial \bar{f}_{1}(\boldmath${\theta}$_{0})}{\partial \beta } \right) \right] , \end{equation} such that for any $0\leq \gamma <1/2$ \begin{equation*} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}^{1/2}(k/{\mathcal{n}} )^{\gamma }}\left\Vert \bar{{\mathcal{r}}}_{m,k}(\boldmath${\theta}$ _{0})-\boldmath${\Gamma}$_{m}(k)\right\Vert =o_{P}(1), \end{equation*} \begin{proof} It is shown in horvath2024detecting that $\{\nabla {\mathcal{f}} _{m,k}(\mbox{\boldmath${\theta}$}_{0}),k\geq 1\}$ is a decomposable Bernoulli shift. Hence $\{\bar{{\mathcal{r}}}_{m,k}(\mbox{\boldmath${ \theta}$}_{0}),k\geq 1\}$ is also a decomposable Bernoulli shift, and therefore by Lemma S2.1 of aue2014dependent we can define Wiener processes $\mbox{\boldmath${\Gamma}$}_{m}(t)\in R^{2}$ with $E \mbox{\boldmath${\Gamma}$}_{m}(t)=\mathbf{0}$ and $E\mbox{\boldmath${ \Gamma}$}_{m}(t)\mbox{\boldmath${\Gamma}$}_{m}(s)=\min (t,s)\mathbf{D}$ such that \begin{equation} \sup_{1\leq k<\infty }\frac{1}{k^{\zeta }}\left\Vert \bar{{\mathcal{r}}} _{m,k}(\mbox{\boldmath${\theta}$}_{0})-\mbox{\boldmath${\Gamma}$} _{m}(k)\right\Vert =O_{P}(1), \end{equation} with some $\zeta <1/2$. Since \begin{equation*} \sup_{1\leq k\leq {\mathcal{n}}}\frac{k^{\zeta }}{{\mathcal{n}}^{1/2}(k/{ \mathcal{n}})^{\gamma }}=o(1), \end{equation*} the result is proven. \end{proof}
lemmaIf Assumptions (ref)--(ref) hold, then we have \begin{equation} \max_{r\leq k\leq {\mathcal{n}}}\frac{1}{r^{1/2}(k/r)^{\gamma }}\sup_{ \boldmath${\theta}$\in \boldmath${\theta}$}\left\Vert { \mathcal{r}}_{m,k}(\boldmath${\theta}$)-\bar{{\mathcal{r}}}_{m,k}( \boldmath${\theta}$)\right\Vert =o_{P}(1), \end{equation} \begin{equation} \max_{r\leq k\leq {\mathcal{n}}}\frac{1}{r^{1/2}(k/r)^{\gamma }}\left\Vert \bar{{\mathcal{r}}}_{m,k}(\hat{\boldmath${\theta}$}_{m})-\bar{{ \mathcal{r}}}_{m,k}(\boldmath${\theta}$_{0})\right\Vert =o_{P}(1), \end{equation} and \begin{equation} \max_{r\leq k\leq {\mathcal{n}}}\frac{1}{r^{1/2}(k/r)^{\gamma }}\left\Vert \bar{{\mathcal{r}}}_{m,k}(\mbox{\boldmath${\theta}$}_{0})- \mbox{\boldmath${\Gamma}$}_{m}(t)\right\Vert =o_{P}(1), \end{equation} for all $\gamma >1/2$, where $\{\mbox{\boldmath${\Gamma}$}_{m}(t),t\geq 0\}$ is defined in Lemma (ref). \begin{proof} It follows from (ref) that \begin{align*} \max_{r\leq k\leq {\mathcal{n}}}& \frac{1}{r^{1/2}(k/r)^{\gamma }}\sup_{ \mbox{\boldmath${\theta}$}\in \mbox{\boldmath${\theta}$}}\left\Vert { \mathcal{r}}_{m,k}(\mbox{\boldmath${\theta}$})-\bar{{\mathcal{r}}}_{m,k}( \mbox{\boldmath${\theta}$})\right\Vert \\ & =\max_{r\leq k\leq {\mathcal{n}}}\frac{k^{\zeta }}{r^{1/2}(k/r)^{\gamma }} \max_{1\leq k\leq {\mathcal{n}}}\sup_{\mbox{\boldmath${\theta}$}\in \mbox{\boldmath${\theta}$}}\frac{1}{k^{\zeta }}\left\Vert {\mathcal{r}} _{m,k}(\mbox{\boldmath${\theta}$})-\bar{{\mathcal{r}}}_{m,k}( \mbox{\boldmath${\theta}$})\right\Vert \\ & =O_{P}\left( r^{\zeta -1/2}\right) =o_{P}(1), \end{align*} since $r\rightarrow \infty $; hence (ref) is proven. Turning to (ref), the expansion in (ref) implies \begin{align*} \max_{r\leq k\leq {\mathcal{n}}}\frac{1}{r^{1/2}(k/r)^{\gamma }}\Vert \mathbf{R}_{m,k}\Vert & =\left\{ \begin{array}{ll} O_{P}\left( ({\mathcal{n}}/m)(r/{\mathcal{n}})^{\gamma }r^{-1/2}\right) ,\;\;\;\mbox{if}\;\;\;1/2<\gamma <1, & \\ O_{P}\left( r^{1/2}/m\right) ,\;\;\;\mbox{if}\;\;\;\gamma \geq 1, & \end{array} \right. \\ & =o_{P}(1). \end{align*} Now Lemma (ref), (ref) and (ref) yield \begin{align*} \max_{r\leq k\leq {\mathcal{n}}}\frac{1}{r^{1/2}(k/r)^{\gamma }}\Vert \nabla \bar{{\mathcal{r}}}_{m,k}(\mbox{\boldmath${\theta}$}_{0})(\hat{ \mbox{\boldmath${\theta}$}}_{m}-\mbox{\boldmath${\theta}$}_{0})\Vert & =\left\{ \begin{array}{ll} O_{P}\left( ({\mathcal{n}}/m)(r/{\mathcal{n}})^{\gamma -1/2}\right) ,\;\;\; \mbox{if}\;\;\;1/2<\gamma <1 & \\ O_{P}\left( (r/{\mathcal{n}})^{\gamma -1/2}({\mathcal{n}}/m)^{1/2}\right) ,\;\;\;\mbox{if}\;\;\;\gamma \geq 1 & \end{array} \right. \\ & =o_{P}(1), \end{align*} completing the proof of (ref). Finally, the approximation in (ref) is an immediate consequence of Lemma (ref), since \begin{align*} \max_{r\leq k\leq {\mathcal{n}}}\frac{1}{r^{1/2}(k/r)^{\gamma }}\left\Vert \bar{{\mathcal{r}}}_{m,k}(\mbox{\boldmath${\theta}$}_{0})- \mbox{\boldmath${\Gamma}$}_{m}(t)\right\Vert & \leq \max_{r\leq k\leq { \mathcal{n}}}\frac{k^{\zeta }}{r^{1/2}(k/r)^{\gamma }}\max_{1\leq k\leq { \mathcal{n}}}\frac{1}{k^{\zeta }}\left\Vert \bar{{\mathcal{r}}}_{m,k}( \mbox{\boldmath${\theta}$}_{0})-\mbox{\boldmath${\Gamma}$}_{m}(t)\right\Vert \\ & =o_{P}(1). \end{align*} \end{proof}
lemmaIf Assumptions (ref) and (ref) hold, then we have for all $1<\overline{\omega }$ there is $0<\rho <1$ such that \begin{equation*} \sup_{1\leq i<\infty }\sup_{1/\overline{\omega }\leq \omega \leq \overline{ \omega }}\frac{1}{\rho ^{i}}|\bar{p}_{m+i,1}(\bar{\boldmath${\theta}$} _{0},\omega )-{\mathcal{z}}_{m+i,1}|=O_{P}(1) \end{equation*} and \begin{equation*} \sup_{1\leq i<\infty }\sup_{1/\overline{\omega }\leq \omega \leq \overline{ \omega }}\frac{1}{\rho ^{i}}|\bar{p}_{m+i,2}(\bar{\boldmath${\theta}$} _{0},\omega )-{\mathcal{z}}_{m+i,2}|=O_{P}(1). \end{equation*} \begin{proof} Using the proofs from jensen2004asymptotic, horvath2024detecting established the result as their Lemma 5.7. \end{proof}
lemmaWe assume that Assumptions (ref)--(ref) hold. \newline (i) If $0\leq \gamma <1/2$, then we have \begin{equation*} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}^{1/2}(k/{\mathcal{n}} )^{\gamma }}\left\Vert {\mathcal{r}}_{m,k}(\hat{\boldmath${\theta}$} _{m})-\mathbf{q}_{m+k}\right\Vert =o_{P}(1). \end{equation*} (ii) If in addition, Assumption (ref) also holds, and $\gamma >1/2$, then we have \begin{equation*} \max_{r\leq k\leq {\mathcal{n}}}\frac{1}{r^{1/2}(k/r)^{\gamma }}\left\Vert { \mathcal{r}}_{m,k}(\hat{\boldmath${\theta}$}_{m})-\mathbf{q} _{m+k}\right\Vert =o_{P}(1). \end{equation*} \begin{proof} Let $\bar{\mbox{\boldmath${\theta}$}}_{0}=(\alpha _{0},\beta _{0})$. jensen2004asymptotic showed there is a neighbourhood of $\bar{ \mbox{\boldmath${\theta}$}}_{0}$, say $\bar{\mbox{\boldmath${\theta}$}}_{0}$ , such that \begin{equation} \sup_{\omega \leq \omega \leq \overline{\omega }}\sup_{(\alpha ,\beta )\in \bar{\boldmath${\theta}$}_{0}}\frac{1}{k}\left\Vert \nabla {\mathcal{r}}_{m,k}(\alpha ,\beta ,\omega )\right\Vert =O_{P}(1). \end{equation} Let $\bar{\mbox{\boldmath${\theta}$}}_{m}$ denote the vector of the first two coordinates of $\hat{\mbox{\boldmath${\theta}$}}_{m}$. jensen2004asymptotic proved that \begin{equation} \sup_{\omega \leq \omega \leq \overline{\omega }}\Vert \bar{ \boldmath${\theta}$}_{m}-\bar{\mbox{\boldmath${\theta}$}}_{0}\Vert =O_{P}\left( m^{-1/2}\right) \end{equation} (note that the estimator $\bar{\mbox{\boldmath${\theta}$}}_{m}$ might depend on $\omega )$. Using the mean value theorem coordinate--wise, we get from Lemma (ref) that \begin{equation*} \Vert {\mathcal{r}}_{m,k}(\hat{\mbox{\boldmath${\theta}$}}_{m})-\mathbf{q} _{m+k}\Vert \leq \sup_{\underline{\omega }\leq \omega \leq \overline{\omega } }\sup_{(\alpha ,\beta )\in \bar{\mbox{\boldmath${\theta}$}}_{0}}\left\Vert \nabla {\mathcal{r}}_{m,k}(\alpha ,\beta ,\omega )\right\Vert \Vert \bar{ \mbox{\boldmath${\theta}$}}_{m}-\bar{\mbox{\boldmath${\theta}$}}_{0}\Vert ^{1/2} \end{equation*} with probability tending to 1. Thus we get from (ref) and (ref) that \begin{align*} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}^{1/2}(k/{\mathcal{n}} )^{\gamma }}\left\Vert {\mathcal{r}}_{m,k}(\hat{\mbox{\boldmath${\theta}$}} _{m})-\mathbf{q}_{m+k}\right\Vert & =O_{P}\left( m^{-1/2}{\mathcal{n}} ^{-1/2+\gamma }\max_{1\leq k\leq {\mathcal{n}}}k^{1-\gamma }\right) \\ & =O_{P}\left( \left( \frac{{\mathcal{n}}}{m}\right) ^{1/2}\right) =o_{P}(1). \end{align*} Similarly, \begin{align*} \max_{r\leq k\leq {\mathcal{n}}}& \frac{1}{r^{1/2}(k/r)^{\gamma }}\left\Vert {\mathcal{r}}_{m,k}(\hat{\mbox{\boldmath${\theta}$}}_{m})-\mathbf{q} _{m+k}\right\Vert \\ & =O_{P}\left( m^{-1/2}r^{-1/2+\gamma }\max_{r\leq k\leq {\mathcal{n}} }k^{1-\gamma }\right) \\ & =\left\{ \begin{array}{ll} O_{P}\left( ({\mathcal{n}}/m)^{1/2}(r/{\mathcal{n}})^{\gamma _{1}/2}\right) ,\quad \;\mbox{if}\;\;\;1/2<\gamma \leq 1 & \\ O_{P}\left( ({\mathcal{n}}/m)^{1/2}\right) ,\quad \;\mbox{if}\;\;\;\gamma >1 & \end{array} \right. \\ & =o_{P}(1). \end{align*} \end{proof}
lemmaIf Assumptions (ref)--(ref) hold, then for any $m$ we can define a Gaussian process $\{\mbox{\boldmath${\Gamma}$}_{m}(t),t\geq 0\}$, $E\mbox{\boldmath${\Gamma}$}_{m}(t)=\mathbf{0},E\mbox{\boldmath${ \Gamma}$}_{m}(t)\mbox{\boldmath${\Gamma}$}_{m}(s)=\min (t,s)\mathbf{D}$ with \begin{equation} \mathbf{D}=E[1-\epsilon _{m}^{2}]^{2}E\left[ \left( {\mathcal{z}}_{m,1},{ \mathcal{z}}_{m,2}\right) ^{\top }\left( {\mathcal{z}}_{m,1},{\mathcal{z}} _{m,2}\right) \right] . \end{equation} such that \newline (i) if $0\leq \gamma <1/2$, then \begin{equation} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}^{1/2}(k/{\mathcal{n}} )^{\gamma }}\left\Vert \mathbf{q}_{m+k}-\boldmath${\Gamma}$ _{m}(k)\right\Vert =o_{P}(1), \end{equation} (ii) if in addition, Assumption (ref) also holds and $\gamma >1/2$, then \begin{equation} \max_{r\leq k\leq {\mathcal{n}}}\frac{1}{r^{1/2}(k/r)^{\gamma }}\left\Vert \mathbf{q}_{m+k}-\boldmath${\Gamma}$_{m}(k)\right\Vert =o_{P}(1). \end{equation} \begin{proof} horvath2024detecting showed that $\{\mathbf{q}_{m+k},k\geq 1\}$ is a decomposable Bernoulli shift and therefore Lemma S2.1 of aue2014dependent implies that for any $m$ there are Gaussian processes $\{\mbox{\boldmath${\Gamma}$}_{m}(k),k\geq 1\}$ such that \begin{equation} \sup_{1\leq k<\infty }\frac{1}{k^{\zeta }}\left\Vert \mathbf{q}_{m+k}- \boldmath${\Gamma}$_{m}(k)\right\Vert =O_{P}(1) \end{equation} with some $\zeta <1/2$, and $E\mbox{\boldmath${\Gamma}$}(t)=\mathbf{0},\;E \mbox{\boldmath${\Gamma}$}(t)\mbox{\boldmath${\Gamma}$}^{\top }(s)=\min (t,s) \mathbf{D}$ and $\mathbf{D}$ is defined in (ref). The approximation in ( (ref)) implies both (ref) and (ref) in the same way as (ref) implies Lemma (ref) and (ref). \end{proof}
lemmaWe assume that $H_{A}$, Assumptions (ref)--(ref) and (ref) are satisfied.\newline (i) If $E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})<0$, then \begin{equation*} \max_{1\leq k\leq k^{\ast }}\frac{{\mathcal{D}}_{m,k}}{g_{m}(k)}=O_{P}\left( \left( \frac{k^{\ast }}{{\mathcal{n}}}\right) ^{1-\eta }\right) . \end{equation*} (ii) If $E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})<0$, then \begin{equation*} \max_{r\leq k\leq \max (r,k^{\ast })}\frac{{\mathcal{D}}_{m,k}}{\bar{g} _{m}(k)}=O_{P}\left( 1\right) . \end{equation*} \begin{proof} The lemma follows from the proof of Theorem (ref). \end{proof}

Next we assume that (ref) holds. The observations are generated by (ref) but in contrast to the previous case, the sequence is explosive and cannot be approximated with a stationary sequence. However, we can approximate the gradient of the log likelihood function with a stationary sequence, we need only minor modifications of the arguments used before. We use the following sequences to approximate the gradient vector of the log likelihood function when (ref) holds for an explosive sequence:

equation*[equation* omitted — 214 chars of source]

and

equation*[equation* omitted — 177 chars of source]

Let

equation*[equation* omitted — 190 chars of source]
lemmaIf $H_{A}$, Assumptions (ref)--(ref), (ref) are satisfied and $E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})>0$, then we have \begin{equation*} \max_{k^{\ast }<i\leq {\mathcal{n}}}\frac{1}{(i-k^{\ast })^{\zeta }} \left\Vert [{\mathcal{r}}_{m,m+i}-{\mathcal{r}}_{m,m+k^{\ast }}]-(i-k^{\ast })\boldmath${\Upsilon}$-\tilde{\mathbf{q}}_{m,k^{\ast },i}\right\Vert =O_{P}(1), \end{equation*} with some $\zeta <1/2$. \begin{proof} After the changepoint the observations change into an explosive sequence with parameter $\mbox{\boldmath${\theta}$}_{A}$, hence the proofs of Lemmas (ref) and (ref) can be repeated. \end{proof}
lemmaIf $H_{A}$, Assumptions (ref)--(ref), (ref) are satisfied and $E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})>0$, then we can define Gaussian processes $\left\{ \tilde{\mbox{\boldmath${\Gamma}$}} _{m}(t),t\geq 0\right\} $ with $E\tilde{\mbox{\boldmath${\Gamma}$}}_{m}(t)= \mathbf{0},E\tilde{\mbox{\boldmath${\Gamma}$}}_{m}(t)\tilde{ \mbox{\boldmath${\Gamma}$}}_{m}^{\top }(s)=\min (t,s)\mbox{\boldmath${ \Sigma}$}_{2}$, where $\mbox{\boldmath${\Sigma}$}_{2}$ is defined in (ref), such that \begin{equation*} \max_{k^{\ast }<k\leq {\mathcal{n}}}\frac{1}{(k-k^{\ast })^{\zeta }} \left\Vert \tilde{\mathbf{q}}_{m,k^{\ast },i}-(i-k^{\ast }) \boldmath${\Upsilon}$-\tilde{\boldmath${\Gamma}$} _{m}(i-k^{\ast })\right\Vert =O_{P}(1). \end{equation*} \begin{proof} The proof is the same as of that of Lemma (ref). \end{proof}

\setcounter{subsection}{-1} \setcounter{subsubsection}{-1} \setcounter{equation}{0} \setcounter{lemma}{0} \setcounter{theorem}{0}

Proofs

proof[Proof of Theorem (ref)] We begin by showing the theorem under stationarity. It follows from Lemmas (ref)--(ref) that \begin{equation*} \sup_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}(k/{\mathcal{n}} )^{\eta }}\left\Vert {\mathcal{r}}_{m,k}^{\top }(\hat{\boldmath${ \theta}$}_{m})\mathbf{D}^{-1}{\mathcal{r}}_{m,k}(\hat{\boldmath${ \theta}$}_{m})-\boldmath${\Gamma}$_{m}^{\top }(k)\mathbf{D}^{-1} \boldmath${\Gamma}$_{m}(k)\right\Vert =o_{P}(1). \end{equation*} The covariance structure of $\mbox{\boldmath${\Gamma}$}_{m}(k)$ implies that \begin{equation*} \sup_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}(k/{\mathcal{n}} )^{\eta }}\left\vert \boldmath${\Gamma}$_{m}^{\top }(k)\mathbf{D}^{-1} \boldmath${\Gamma}$_{m}(k)\right\vert \overset{{\mathcal{D}}}{=} \sup_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}(k/{\mathcal{n}} )^{\eta }}\left\Vert \mathbf{W}(k)\right\Vert ^{2}, \end{equation*} where $\{\mathbf{W}(t),t\geq 0\}$ has independent coordinates, identically distributed Wiener processes. Using the the scale transformation and continuity of the Wiener process we conclude \begin{equation*} \sup_{1\leq k\leq {\mathcal{n}}}\frac{1}{{\mathcal{n}}(k/{\mathcal{n}} )^{\eta }}\left\Vert \mathbf{W}(k)\right\Vert ^{2}\overset{{\mathcal{D}}}{=} \sup_{1/{\mathcal{n}}\leq k/{\mathcal{n}}\leq }\frac{1}{(k/{\mathcal{n}} )^{\eta }}\left\Vert \mathbf{W}(k/{\mathcal{n}})\right\Vert ^{2}\overset{{ \mathcal{D}}}{\rightarrow }\sup_{0<t\leq }\frac{1}{t^{\eta }}\left\Vert \mathbf{W}(t)\right\Vert ^{2}. \end{equation*} The proof of part \textit{(i)} of the theorem is complete when (ref) holds since \begin{equation} \left\Vert \hat{\mathbf{D}}_{m}-\mathbf{D}\right\Vert =o_{P}\left( 1\right) , \end{equation} with $\mathbf{D}$ defined in (ref). We now turn to proving the second part of the Theorem. We note \begin{align*} \max_{r\leq k\leq {\mathcal{n}}}\frac{1}{r(k/r)^{\eta }}\mbox{\boldmath${ \Gamma}$}_{m}^{\top }(k)\mathbf{D}^{-1}\mbox{\boldmath${\Gamma}$}_{m}(k)& \overset{{\mathcal{D}}}{=}\max_{r\leq k\leq {\mathcal{n}}}\frac{1}{ r(k/r)^{\eta }}\left\Vert \mathbf{W}(k)\right\Vert ^{2} \\ & \overset{{\mathcal{D}}}{=}\max_{1\leq k/r\leq {\mathcal{n}}/r}\frac{1}{ (k/r)^{\eta }}\left\Vert \mathbf{W}(k/r)\right\Vert ^{2} \\ & \overset{{\mathcal{D}}}{\rightarrow }\max_{1\leq t<\infty }\frac{1}{ t^{\eta }}\left\Vert \mathbf{W}(t)\right\Vert ^{2} \end{align*} due to the almost sure continuity of the Wiener process, and the Law of the Iterated Logarithm for Wiener processes. The result in part \textit{(ii)} of the theorem now follows from Lemma (ref) and (ref). Under nonstationarity, the proof is exactly the same as above, only Lemmas (ref)--(ref) are replaced with Lemmas (ref)--(ref). We also note that according to jensen2004asymptotic (cf.\ also francq2019garch), it holds that $\left\Vert \hat{\mathbf{D}}_{m}- \mathbf{D}\right\Vert =o_{P}(1)$, where $\mathbf{D}$ is defined in (ref) , thus completing the proof.
proof[Proof of Theorem (ref)] We note that under our conditions, it holds that $\Vert \hat{\mathbf{D}}_{m}- \mathbf{D}\Vert =O_{P}(m^{-\delta })$, for some $\delta >0$. Hence in the proofs we can replace $\hat{\mathbf{D}}_{m}^{-1}$ with $\mathbf{D}^{-1}$ in the definition of the detector, without loss of generality. We wish to point out that the definition of $\mathbf{D}$ depends on the sign of $E\log (\alpha _{0}\epsilon _{0}^{2}+\beta _{0})$.\newline First we assume that $E\log (\alpha _{A}\epsilon _{0}^{2}+\beta _{A})<0$ holds. To prove part (i) of the theorem, we note that according to Lemma (ref) \begin{equation} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{k^{1/2}}\sup_{\boldmath${ \theta}$\in \boldmath${\theta}$}\Vert {\mathcal{r}}_{m,k}( \boldmath${\theta}$)-\bar{{\mathcal{r}}}_{m,k}(\boldmath${ \theta}$)\Vert =O_{P}(1), \end{equation} and \begin{equation} \max_{a\leq k\leq {\mathcal{n}}}\frac{1}{k^{1/2}}\sup_{\boldmath${ \theta}$\in \mbox{\boldmath${\theta}$}}\Vert {\mathcal{r}}_{m,k}( \mbox{\boldmath${\theta}$})-\bar{{\mathcal{r}}}_{m,k}(\mbox{\boldmath${ \theta}$})\Vert =O_{P}\left( a^{-1/2+\zeta }\right) , \end{equation} where $\zeta <1/2$ and $a=(\log {\mathcal{n}})^{\phi }$ with $\phi >4$. Following the proof of Lemma (ref) we get \begin{equation*} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{k^{1/2}}\Vert \mathbf{R}_{m,k}\Vert =O_{P}({\mathcal{n}}^{1/2})\Vert \hat{\mbox{\boldmath${\theta}$}}_{m}- \mbox{\boldmath${\theta}$}_{0}\Vert ^{2}=O_{P}\left( \frac{{\mathcal{n}} ^{1/2}}{m}\right) , \end{equation*} and by (ref) \begin{equation*} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{k^{1/2}}\left\Vert \nabla ^{2}{ \mathcal{f}}_{m,k}(\mbox{\boldmath${\theta}$}_{0})(\hat{\mbox{\boldmath${ \theta}$}}_{m}-\mbox{\boldmath${\theta}$})\right\Vert =O_{P}\left( \frac{{ \mathcal{n}}^{1/2}}{m^{1/2}}\right) , \end{equation*} \begin{equation*} \max_{a\leq k\leq {\mathcal{n}}}\frac{1}{k^{1/2}}\left\Vert \nabla ^{2}{ \mathcal{f}}_{m,k}(\mbox{\boldmath${\theta}$}_{0})(\hat{\mbox{\boldmath${ \theta}$}}_{m}-\mbox{\boldmath${\theta}$})\right\Vert =O_{P}\left( \frac{ a^{1/2}}{m^{1/2}}\right) . \end{equation*} Using the approximation in (ref) we conclude \begin{equation} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{k^{1/2}}\left\Vert \bar{{\mathcal{r} }}_{m,k}(\mbox{\boldmath${\theta}$}_{0})-\mbox{\boldmath${\Gamma}$} _{m}(k)\right\Vert =O_{P}\left( 1\right) \end{equation} and \begin{equation} \max_{1\leq k\leq a}\frac{1}{k^{1/2}}\left\Vert \bar{{\mathcal{r}}}_{m,k}( \mbox{\boldmath${\theta}$}_{0})-\mbox{\boldmath${\Gamma}$}_{m}(k)\right\Vert =O_{P}\left( a^{-1/2+\zeta }\right) . \end{equation} The Darling--Erd\H{o}s (cf. csorgo1997, pp.\ 363--365) yields \begin{equation*} \frac{1}{2\log \log {\mathcal{n}}}\max_{1\leq k\leq {\mathcal{n}}}\frac{1}{k} \mbox{\boldmath${\Gamma}$}_{m}^{\top }(k)\mathbf{D}^{-1}\mbox{\boldmath${ \Gamma}$}_{m}(k)\overset{P}{\rightarrow }1, \end{equation*} and \begin{equation*} \max_{a\leq k\leq {\mathcal{n}}}\frac{1}{k}\mbox{\boldmath${\Gamma}$} _{m}^{\top }(k)\mathbf{D}^{-1}\mbox{\boldmath${\Gamma}$}_{m}(k)=O_{P}\left( \log \log \log {\mathcal{n}}\right) . \end{equation*} Thus we obtain that \begin{equation*} \lim_{m\rightarrow \infty }P\left\{ \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{ k}{\mathcal{r}}_{m,k}^{\top }(\hat{\mbox{\boldmath${\theta}$}}_{m})\mathbf{D} ^{-1}{\mathcal{r}}_{m,k}(\hat{\mbox{\boldmath${\theta}$}}_{m})=\max_{a\leq k\leq {\mathcal{n}}}\frac{1}{k}{\mathcal{r}}_{m,k}^{\top }(\hat{ \mbox{\boldmath${\theta}$}}_{m})\mathbf{D}^{-1}{\mathcal{r}}_{m,k}(\hat{ \mbox{\boldmath${\theta}$}}_{m})\right\} =1. \end{equation*} Putting together our estimates on the set $a\leq k\leq {\mathcal{n}}$, we get \begin{equation*} \max_{a\leq k\leq {\mathcal{n}}}\frac{1}{k}\left\vert {\mathcal{r}} _{m,k}^{\top }(\hat{\mbox{\boldmath${\theta}$}}_{m})\mathbf{D}^{-1}{\mathcal{ r}}_{m,k}(\hat{\mbox{\boldmath${\theta}$}}_{m})-\mbox{\boldmath${\Gamma}$} _{m}^{\top }(k)\mathbf{D}^{-1}\mbox{\boldmath${\Gamma}$}_{m}(k)\right\vert =O_{P}\left( (\log {\mathcal{n}})^{-\bar{\phi}}\right) , \end{equation*} with some $\bar{\phi}>0$. Since \begin{equation*} \max_{a\leq k\leq {\mathcal{n}}}\mbox{\boldmath${\Gamma}$}_{m}^{\top }(k) \mathbf{D}^{-1}\mbox{\boldmath${\Gamma}$}_{m}(k)\overset{{\mathcal{D}}}{=} \max_{a\leq k\leq {\mathcal{n}}}\Vert \mathbf{W}(k)\Vert ^{2}, \end{equation*} where $\{\mathbf{W}(t),t\geq 0\}$ is defined in Theorem (ref). Using Appendix 3 in csorgo1997, we obtain that \begin{equation*} \lim_{{\mathcal{n}}\rightarrow \infty }P\left\{ a(\log {\mathcal{n}} )\max_{a\leq k\leq {\mathcal{n}}}\Vert \mathbf{W}(k)\Vert \leq x+b_{2}(\log { \mathcal{n}})\right\} =\exp (-e^{-x}) \end{equation*} for all $x$. Hence the proof Theorem (ref)(i) is complete when $E\log (\alpha _{0}\epsilon _{0}^{2}+\beta _{0})<0$.\newline Lemma (ref), (ref) and (ref) imply that (ref) and (ref) can be replaced with \begin{equation*} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{k^{1/2}}\left\Vert {\mathcal{r}} _{m,k}(\hat{\mbox{\boldmath${\theta}$}}_{m})-\mathbf{q}_{m+k}\right\Vert =O_{P}(1), \end{equation*} and \begin{equation*} \max_{a\leq k\leq {\mathcal{n}}}\frac{1}{k^{1/2}}\left\Vert {\mathcal{r}} _{m,k}(\hat{\mbox{\boldmath${\theta}$}}_{m})-\mathbf{q}_{m+k}\right\Vert =O_{P}(a^{-1/2+\zeta }), \end{equation*} with some $\zeta <1/2$, when $E\log (\alpha _{0}\epsilon _{0}^{2}+\beta _{0})>0$. The approximation in (ref) yields \begin{equation} \max_{1\leq k\leq {\mathcal{n}}}\frac{1}{k^{1/2}}\left\Vert \mathbf{q}_{m+k}- \mbox{\boldmath${\Gamma}$}_{m}(k)\right\Vert =O_{P}(1), \end{equation} and \begin{equation} \max_{a\leq k\leq {\mathcal{n}}}\frac{1}{k^{1/2}}\left\Vert \mathbf{q}_{m+k}- \mbox{\boldmath${\Gamma}$}_{m}(k)\right\Vert =O_{P}(a^{-1/2+\zeta }). \end{equation} The approximations in (ref) and (ref) imply part \textit{(i) } of the theorem, in the same way as (ref) and (ref) yielded (ref) under condition $E\log (\alpha _{0}\epsilon _{0}^{2}+\beta _{0})<0$.\newline To prove part \textit{(ii)} of the theorem, we use again the approximations in (ref) and (ref). By the scale transformation of Wiener processes we have \begin{equation} \sup_{r\leq t\leq {\mathcal{n}}}\frac{1}{t^{1/2}}\Vert \mathbf{W}(t)\Vert \overset{{\mathcal{D}}}{=}\sup_{1\leq t\leq {\mathcal{n}}/r}\frac{1}{t^{1/2}} \Vert \mathbf{W}(t)\Vert \end{equation} and therefore the Darling--Erd\H{o}s limit result (cf. csorgo1997, pp.\ 363--368) with (ref) and (ref) yields \begin{equation*} \frac{1}{2\log \log ({\mathcal{n}}/r)}\max_{r\leq k\leq {\mathcal{n}}}\frac{1 }{k}{\mathcal{r}}_{m,k}^{\top }(\hat{\mbox{\boldmath${\theta}$}}_{m})\mathbf{ D}^{-1}{\mathcal{r}}_{m,k}\overset{P}{\rightarrow }1 \end{equation*} and \begin{equation*} \max_{r\leq k\leq rg}\frac{1}{k}{\mathcal{r}}_{m,k}^{\top }(\hat{ \mbox{\boldmath${\theta}$}}_{m})\mathbf{D}^{-1}{\mathcal{r}} _{m,k}=O_{P}\left( \log \log \log ({\mathcal{n}}/r)\right) \end{equation*} with \begin{equation*} g=(\log ({\mathcal{n}}/r))^{\bar{\phi}},\;\;\;\bar{\phi}>4. \end{equation*} Now we get \begin{equation*} \lim_{m\rightarrow \infty }P\left\{ \max_{r\leq k\leq {\mathcal{n}}}\frac{1}{ k}{\mathcal{r}}_{m,k}^{\top }(\hat{\mbox{\boldmath${\theta}$}}_{m})\mathbf{D} ^{-1}{\mathcal{r}}_{m,k}=\max_{rg\leq k\leq {\mathcal{n}}}\frac{1}{k}{ \mathcal{r}}_{m,k}^{\top }(\hat{\mbox{\boldmath${\theta}$}}_{m})\mathbf{D} ^{-1}{\mathcal{r}}_{m,k}\right\} =1. \end{equation*} Using the approximation of ${\mathcal{r}}_{m,k}(\hat{\mbox{\boldmath${ \theta}$}}_{m})$ with Gaussian processes we get \begin{equation*} \max_{rg\leq k\leq {\mathcal{n}}}\frac{1}{k}\left\vert {\mathcal{r}} _{m,k}^{\top }(\hat{\mbox{\boldmath${\theta}$}}_{m})\mathbf{D}^{-1}{\mathcal{ r}}_{m,k}-\mbox{\boldmath${\Gamma}$}_{m}^{\top }(k)\mathbf{D}^{-1} \mbox{\boldmath${\Gamma}$}_{m}(k)\right\vert =O_{P}\left( (\log ({\mathcal{n} }/r))^{-\bar{\phi}}\right) \end{equation*} with some $\bar{\phi}>0$. Now using the results on pp. 363--365 in csorgo1997 with (ref) yields \begin{equation*} \lim_{m\rightarrow \infty }P\left\{ a(\log ({\mathcal{n}}/r))\max_{rg\leq k\leq {\mathcal{n}}}\frac{1}{k^{1/2}}\left( \mbox{\boldmath${\Gamma}$} _{m}^{\top }(k)\mathbf{D}^{-1}\mbox{\boldmath${\Gamma}$}_{m}(k)\right) ^{1/2}\leq x+b_{2}(\log ({\mathcal{n}}/r))\right\} =\exp (-e^{-x}), \end{equation*} completing the proof of part \textit{(ii)} of the theorem when $E\log (\alpha _{0}\epsilon _{0}^{2}+\beta _{0})<0$. The same arguments can be used when $E\log (\alpha _{0}\epsilon _{0}^{2}+\beta _{0})>0$.
proof[Proof of Theorem (ref)] First we assume that (ref) holds. After $m+k^{\ast }$, the sequence can turn into (asymptotically) stationary sequence. Since (ref) holds, we have \begin{equation*} \max_{k^{\ast }+1\leq j<\infty }\sup_{\boldmath${\theta}$\in \boldmath${\theta}$}\frac{1}{j-k^{\ast }}\Vert {\mathcal{r}}_{m,j}( \boldmath${\theta}$)-{\mathcal{r}}_{m,k^{\ast }}(\boldmath${ \theta}$)-(j-k^{\ast }){\mathcal{r}}^{(1)}(\boldmath${\theta}$)\Vert =o_{P}(1), \end{equation*} where \begin{equation*} {\mathcal{r}}^{(1)}(\boldmath${\theta}$)=E\bar{{\mathcal{r}}} _{m,k^{\ast }+1}^{(1)}(\mbox{\boldmath${\theta}$}), \end{equation*} \begin{equation*} \bar{{\mathcal{r}}}_{m,k^{\ast }+j}^{(1)}(\mbox{\boldmath${\theta}$} )=\sum_{i=k^{\ast }+1}^{k^{\ast }+j}\left( \frac{\partial \hat{\ell}_{i}( \mbox{\boldmath${\theta}$})}{\partial \alpha },\frac{\partial \hat{\ell}_{i}( \mbox{\boldmath${\theta}$})}{\partial \beta }\right) ^{\top }, \end{equation*} Under $H_{A}$ we have that $\mbox{\boldmath${\Delta}$}={\mathcal{r}}^{(1)}( \mbox{\boldmath${\theta}$}_{0})\neq \mathbf{0}.$ Next we write for $ k>k^{\ast }$ \begin{equation} {\mathcal{r}}_{m,k}(\hat{\mbox{\boldmath${\theta}$}}_{m})=(k-k^{\ast }) \mbox{\boldmath${\Delta}$}+\mathbf{m}_{m,k} \end{equation} with \begin{equation} \mathbf{m}_{m,k}={\mathcal{r}}_{m,k^{\ast }}(\hat{\mbox{\boldmath${\theta}$}} _{m})+[{\mathcal{r}}_{m,k}(\hat{\mbox{\boldmath${\theta}$}}_{m})-{\mathcal{r} }_{m,k^{\ast }}(\hat{\mbox{\boldmath${\theta}$}}_{m})-(k-k^{\ast }) \mbox{\boldmath${\Delta}$}] \end{equation} We use the decomposition \begin{equation} {\mathcal{D}}_{m,k}=(k-k^{\ast })^{2}A_{m}+2(k-k^{\ast })\mbox{\boldmath${ \Delta}$}^{\top }\hat{\mathbf{D}}_{m}^{-1}\mathbf{m}_{m,k}+\mathbf{m} _{m,k}^{\top }\mathbf{D}_{m}^{-1}\mathbf{m}_{m,k}, \end{equation} where $A_{m}$ is defined in (ref). We note that, for any ${\mathcal{ M}}$ \begin{equation*} \max_{1\leq k\leq {\mathcal{M}}}\frac{{\mathcal{D}}_{m,k}}{g_{m}(k)}=\max \left( \max_{1\leq k\leq k^{\ast }}\frac{{\mathcal{D}}_{m,k}}{g_{m}(k)} ,\max_{k^{\ast }+1\leq k\leq {\mathcal{M}}}\frac{{\mathcal{D}}_{m,k}}{ g_{m}(k)}\right) . \end{equation*} First we consider the case when (ref) holds. The observations $y_{i}$ can be approximated with a stationary sequence defined in (ref). Recalling the definitions of $t^{\ast },\bar{u},u_{{\mathcal{n}}}$ and $ u^{\ast }$ in (ref)--(ref), we define \begin{equation*} {\mathcal{M}}=k^{\ast }+u_{{\mathcal{n}}}\left( {\mathcal{n}}^{1-\eta }\frac{ {\mathcal{c}}}{A_{m}}\right) ^{1/(2-\eta )}+R, \end{equation*} with \begin{equation*} R=x{(u^{\ast }+t^{\ast })^{3/2}}\left( \left( {\mathcal{n}}^{1-\eta }\frac{{ \mathcal{c}}}{A_{m}}\right) ^{1/(2-\tau )}\right) ^{1/2}. \end{equation*} Now according to Theorem (ref) \begin{equation*} \left( \max_{1\leq k\leq k^{\ast }}\frac{{\mathcal{D}}_{m,k}}{k^{\eta }}-{ \mathcal{c}}{\mathcal{n}}^{1-\eta }\right) \left( {\mathcal{n}}^{(1-\eta )/(2-\eta )}\right) ^{-3/2+\eta }\overset{\mathcal{P}}{\rightarrow }-\infty . \end{equation*} It follows from the proof of Theorem (ref) \begin{equation*} \max_{k^{\ast }+1\leq k\leq {\mathcal{M}}}\frac{\mathbf{m}_{m,k}^{\top } \mathbf{D}_{m}^{-1}\mathbf{m}_{m,k}}{g_{m}(k)}=O_{P}(1), \end{equation*} and for any $\bar{\eta}<1/2$ \begin{equation*} \max_{k^{\ast }+1\leq k\leq {\mathcal{M}}}\frac{1}{k-k^{\ast }}\frac{\Vert \mathbf{m}_{m,k}\Vert }{{\mathcal{n}}^{1/2}(k/{\mathcal{n}})^{\bar{\eta}}} =O_{P}(1). \end{equation*} Let $0<\delta <1$ and write \begin{equation*} {\mathcal{n}}^{1-\eta }\max_{k^{\ast }+1\leq k\leq (1-\delta ){\mathcal{M}}} \frac{{\mathcal{D}}_{m,k}}{g_{m}(k)}=(1-\delta )^{2-\eta }{\mathcal{M}} ^{2-\eta }O_{P}(1), \end{equation*} where $O_{P}(1)$ does not depend on $\delta $. Thus we conclude \begin{equation*} \left( {\mathcal{n}}^{1-\eta }\max_{k^{\ast }+1\leq k\leq (1-\delta ){ \mathcal{M}}}\frac{{\mathcal{D}}_{m,k}}{g_{m}(k)}-{\mathcal{n}}^{1-\eta }\right) ({\mathcal{n}}^{1-\eta })^{(-3/2+\eta )/(2-\eta )}\overset{\mathcal{ P}}{\rightarrow }-\infty . \end{equation*} Our calculations show that we need to consider the asymptotic properties of \begin{align*} & \left( \max_{(1-\delta ){\mathcal{M}}\leq k\leq {\mathcal{M}}} \frac{1}{k^{\eta }}\left( (k-k^{\ast })^{2}A_{m}+2(k-k^{\ast }) \mbox{\boldmath${\Delta}$}^{\top }\hat{\mathbf{D}}_{m}^{-1}\mathbf{m} _{m,k}\right) \right. \\ & \left. -\left[ {\mathcal{c}}{\mathcal{n}}^{1-\eta }-\frac{1}{{\mathcal{M}} ^{\eta }}({\mathcal{M}}-k^{\ast })^{2}A_{m}\right] \right) \left( {\mathcal{n }}^{(1-\eta )/(2-\eta )}\right) ^{-3/2+\eta }. \end{align*} Since the model is the same until $y_{m+k^{\ast }}$, the $k^{\ast }$th observation following the training sample, we have that \begin{align} \mathbf{D}=\lim_{k^{\ast }\rightarrow \infty }& \frac{1}{m+k^{\ast }} E\left\{ \left( \frac{\partial \bar{{\mathcal{r}}}_{m,m+k^{\ast }}( \mbox{\boldmath${\theta}$}_{0})}{\partial \alpha },\frac{\partial \bar{{ \mathcal{r}}}_{m,m+k^{\ast }}(\mbox{\boldmath${\theta}$}_{0})}{\partial \beta }\right) ^{\top }\right. \\ & \left. \times \left( \frac{\partial \bar{{\mathcal{r}}} _{m,m+k^{\ast }}(\mbox{\boldmath${\theta}$}_{0})}{\partial \alpha },\frac{ \partial \bar{{\mathcal{r}}}_{m,m+k^{\ast }}(\mbox{\boldmath${\theta}$}_{0}) }{\partial \beta }\right) \right\} . \notag \end{align} (Note that the formula for $\mathbf{D}$ depends on whether the training sample is stationary or non stationary, cf. berkes2003garch and jensen2004asymptotic). Along the lines of the proof of Theorem (ref) one can show that \begin{equation} {\mathcal{M}}^{-3/2}\left( \lfloor {\mathcal{M}}t\rfloor -\lfloor k^{\ast }/{ \mathcal{M}}\rfloor \right) \mbox{\boldmath${\Delta}$}^{\top }\hat{\mathbf{D} }_{m}^{-1}\mathbf{m}_{m,{\mathcal{M}}t}\overset{{\mathcal{D}}[1-\delta ,1]}{ \longrightarrow }(t-t^{\ast }/(u^{\ast }+t^{\ast }))Z(t), \end{equation} where $\{Z(t),t\geq t^{\ast }/(t^{\ast }+u^{\ast })\}$ is a Gaussian process with $EZ(t)=0$ and \begin{equation*} EZ(t)Z(s)=t^{\ast }\mbox{\boldmath${\Delta}$}^{\top }\mathbf{D}^{-1} \mbox{\boldmath${\Delta}$}+[\min (t,s)-t^{\ast }/(t^{\ast }+u^{\ast })] \mbox{\boldmath${\Delta}$}^{\top }\mathbf{D}^{-1}\mbox{\boldmath${\Sigma}$} _{1}\mathbf{D}^{-1}\mbox{\boldmath${\Delta}$}, \end{equation*} where $\mbox{\boldmath${\Sigma}$}_{1}$ is defined in (ref). Hence \begin{equation*} {\mathcal{M}}^{-1/2+\eta }\max_{(1-\delta ){\mathcal{M}}\leq k\leq {\mathcal{ M}}}\frac{1}{k^{\tau }}(k-k^{\ast })\mbox{\boldmath${\Delta}$}^{\top }\hat{ \mathbf{D}}_{m}^{-1}\mathbf{m}_{m,k}\overset{{\mathcal{D}}}{\rightarrow } \sup_{1-\delta \leq t\leq 1}\frac{1}{t^{\eta }}(t-t^{\ast })Z(t). \end{equation*} Also, for any $x>0$ \begin{align*} \lim_{\delta \rightarrow 0}\limsup_{m\rightarrow \infty }& P\left\{ {\mathcal{M}}^{-1/2+\eta }\max_{(1-\delta ){\mathcal{M}}\leq k\leq { \mathcal{M}}}\frac{1}{k^{\eta }}\left\vert (k-k^{\ast })\mbox{\boldmath${ \Delta}$}^{\top }\hat{\mathbf{D}}_{m}^{-1}\mathbf{m}_{m,k}-k^{\ast } \mbox{\boldmath${\Delta}$}^{\top }\hat{\mathbf{D}}_{m}^{-1}\mathbf{m}_{m,{ \mathcal{M}}}\right\vert >x\right\} \\ & =0. \end{align*} Since $(k-k^{\ast })^{2}/t^{\eta }$ is increasing on $[(1-\delta ){\mathcal{M }},{\mathcal{M}}]$ we get that \begin{align*} & \lim_{\delta \rightarrow 0}\limsup_{m\rightarrow \infty }P\left\{ \sup_{(1-\delta ){\mathcal{M}}\leq k\leq {\mathcal{M}}}\frac{1}{ k^{\eta }}\left( (k-k^{\ast })^{2}A_{m}+2(k-k^{\ast })\mbox{\boldmath${ \Delta}$}^{\top }\hat{\mathbf{D}}_{m}^{-1}\mathbf{m}_{m,k}\right) \right. \\ & \left. =\frac{1}{{\mathcal{M}}^{\eta }}\left[ ({\mathcal{M}}-k^{\ast })^{2}A_{m}+2({\mathcal{M}}-k^{\ast })\mbox{\boldmath${\Delta}$}^{\top }\hat{ \mathbf{D}}_{m}^{-1}\mathbf{m}_{m,{\mathcal{M}}}\right] \right\} =1. \end{align*} Next we write \begin{align*} P& \left\{ \frac{1}{{\mathcal{M}}^{\eta }}\left[ ({\mathcal{M}}-k^{\ast })^{2}A_{m}+2({\mathcal{M}}-k^{\ast })\mbox{\boldmath${\Delta}$}^{\top }\hat{ \mathbf{D}}_{m}^{-1}\mathbf{m}_{m,{\mathcal{M}}}\right] \leq {\mathcal{c}}{ \mathcal{n}}^{1-\eta }\right\} \\ & =P\left\{ 2({\mathcal{M}}-k^{\ast })[\mbox{\boldmath${\Delta}$}^{\top } \hat{\mathbf{D}}_{m}^{-1}\mathbf{m}_{m,{\mathcal{M}}}]{\mathcal{M}} ^{-3/2}\leq \left[ {\mathcal{c}}{\mathcal{n}}^{1-\eta }-\frac{1}{{\mathcal{M} }^{\eta }}({\mathcal{M}}-k^{\ast })^{2}A_{m}\right] {\mathcal{M}}^{-3/2+\eta }\right\} . \end{align*} Using (ref) we get \begin{equation*} {\mathcal{M}}^{-1/2}\mbox{\boldmath${\Delta}$}^{\top }\hat{\mathbf{D}} _{m}^{-1}\mathbf{m}_{m,{\mathcal{M}}}\overset{{\mathcal{D}}}{\rightarrow } \bar{s}_{1}{\mathcal{N}}, \end{equation*} where ${\mathcal{N}}$ denotes a standard normal random variable, $\bar{s} _{1} $ is defined in (ref) and $\mathbf{D}$ is the covariance matrix in (ref). Also, \begin{equation*} \left[ {\mathcal{n}}^{1-\eta }-\frac{1}{{\mathcal{M}}^{\eta }}({\mathcal{M}} -k^{\ast })^{2}A_{m}\right] {\mathcal{M}}^{-3/2+\eta }\rightarrow -2x, \end{equation*} and thus we conclude \begin{align*} \lim_{m\rightarrow \infty }& P\left\{ 2({\mathcal{M}}-k^{\ast })[ \mbox{\boldmath${\Delta}$}^{\top }\hat{\mathbf{D}}_{m}^{-1}\mathbf{m}_{m,{ \mathcal{M}}}]{\mathcal{M}}^{-3/2}\leq \left[ {\mathcal{c}}{\mathcal{n}} ^{1-\eta }-\frac{1}{{\mathcal{M}}^{\eta }}({\mathcal{M}}-k^{\ast })^{2}A_{m} \right] {\mathcal{M}}^{-3/2+\eta }\right\} \\ & =P\left\{ \bar{s}_{1}{\mathcal{N}}\leq -x\right\} =1-\Phi (x/\bar{s}_{1}). \end{align*} We assume now that $t^{\ast }=\infty $ and $\bar{u}=0$ also holds. We modify the definition of ${\mathcal{M}}$ as \begin{equation*} {\mathcal{M}}=k^{\ast }+\left( \frac{{\mathcal{c}}}{A_{m}}{\mathcal{n}} ^{1-\eta }(k^{\ast })^{\eta }\right) ^{1/2}+R\quad \;\mbox{with}\quad \;R= \frac{x}{A_{m}}(k^{\ast })^{1/2}. \end{equation*} Proceeding as in the previous case we get that for any $0<\delta <1$ \begin{align} \lim_{m\rightarrow \infty }& P\left\{ \tau _{m}>{\mathcal{M}}\right\} \\ & =\lim_{m\rightarrow \infty }\left\{ \max_{(1-\delta ){\mathcal{M}}\leq k\leq {\mathcal{M}}}\frac{1}{k^{\eta }}\left( (k-k^{\ast })^{2}A_{m}+2(k-k^{\ast })\mbox{\boldmath${\Delta}$}^{\top }\hat{\mathbf{D}} _{m}^{-1}\mathbf{m}_{m.k}\right) \leq {\mathcal{c}}{\mathcal{n}}^{1-\eta }\right\} \notag \\ & =\lim_{m\rightarrow \infty }P\left\{ 2({\mathcal{M}}-k^{\ast }) \mbox{\boldmath${\Delta}$}^{\top }\hat{\mathbf{D}}_{m}^{-1}\mathbf{m}_{m,{ \mathcal{M}}}\leq \left( {\mathcal{c}}{\mathcal{n}}^{1-\eta }-\frac{1}{{ \mathcal{M}}^{\eta }}({\mathcal{M}}-k^{\ast })^{2}A_{m}\right) {\mathcal{M}} ^{\eta }\right\} \notag \\ & =\lim_{m\rightarrow \infty }P\left\{ 2(k^{\ast })^{-1/2} \mbox{\boldmath${\Delta}$}^{\top }\hat{\mathbf{D}}_{m}^{-1}\mathbf{m}_{m,{ \mathcal{M}}}\leq \left( {\mathcal{c}}{\mathcal{n}}^{1-\eta }-\frac{1}{{ \mathcal{M}}^{\tau }}({\mathcal{M}}-k^{\ast })^{2}A_{m}\right) \frac{{ \mathcal{M}}^{\eta }}{({\mathcal{M}}-k^{\ast })(k^{\ast })^{1/2}}\right\} . \notag \end{align} Using the definitions of ${\mathcal{M}}$ and $R$ we have \begin{equation*} \left( {\mathcal{c}}{\mathcal{n}}^{1-\eta }-\frac{1}{{\mathcal{M}}^{\eta }}({ \mathcal{M}}-k^{\ast })^{2}A_{m}\right) \frac{{\mathcal{M}}^{\eta }}{({ \mathcal{M}}-k^{\ast })(k^{\ast })^{1/2}}\rightarrow -2x. \end{equation*} As in the proof of (ref) we have \begin{equation*} (k^{\ast })^{-1/2}\mbox{\boldmath${\Delta}$}^{\top }\hat{\mathbf{D}}_{m}^{-1} \mathbf{m}_{m,{\mathcal{M}}}\overset{{\mathcal{D}}}{\rightarrow }\bar{s}_{2}{ \mathcal{N}}, \end{equation*} where ${\mathcal{N}}$ is a standard normal random variable and $\bar{s}_{2}$ is defined in (ref). Now we conclude \begin{equation*} \lim_{m\rightarrow \infty }P\{\tau _{m}>{\mathcal{M}}\}=1-\Phi (x/\bar{s} _{2}), \end{equation*} where $\Phi $ is the standard normal distribution function. \newline In the final case we write \begin{equation*} k^{\ast }=u_{m}{\mathcal{n}}\;\;\;\;\mbox{with}\;\;\;u_{m}\rightarrow \bar{u} \in (0,1). \end{equation*} In this case, we define \begin{equation*} {\mathcal{M}}=k^{\ast }+x(k^{\ast })^{1/2}. \end{equation*} We write \begin{equation*} \max_{1\leq k\leq {\mathcal{M}}}\frac{{\mathcal{D}}_{m,k}}{{\mathcal{n}}(k/{ \mathcal{n}})^{\eta }}=\max \left( U_{1,m},U_{2,m}\right) , \end{equation*} \begin{equation*} U_{1,m}=\max_{1\leq k\leq k^{\ast }}\frac{{\mathcal{D}}_{m,k}}{{\mathcal{n}} (k/{\mathcal{n}})^{\eta }} \end{equation*} and \begin{equation*} U_{2,m}=\max_{k^{\ast }+1\leq k\leq {\mathcal{M}}}\frac{{\mathcal{D}}_{m,k}}{ {\mathcal{n}}(k/{\mathcal{n}})^{\eta }}. \end{equation*} We use the decomposition in (ref) and we note \begin{equation*} \left( \frac{1}{k^{\ast }}\right) ^{1/2}\max_{k^{\ast }+1\leq k\leq { \mathcal{M}}}\Vert {\mathcal{r}}_{m,k}(\hat{\mbox{\boldmath${\theta}$}}_{m})- {\mathcal{r}}_{m,k^{\ast }}(\hat{\mbox{\boldmath${\theta}$}}_{m})\Vert =O_{P}(1). \end{equation*} It follows from the proof of Theorem (ref) that \begin{equation*} \left( \max_{1\leq k\leq k^{\ast }}\frac{{\mathcal{D}}_{m,k}}{k^{\ast }(k/k^{\ast })^{\eta }},\left( \frac{1}{k^{\ast }}\right) ^{1/2}{\mathcal{r}} _{m,k^{\ast }}(\hat{\mbox{\boldmath${\theta}$}}_{m})\right) \overset{{ \mathcal{D}}}{\rightarrow }\left( \sup_{0\leq s\leq 1}\frac{ \mbox{\boldmath${\Gamma}$}(s)^{\top }\mathbf{D}^{-1}\mbox{\boldmath${ \Gamma}$}(s)}{s^{\eta }},\mbox{\boldmath${\Gamma}$}(1)\right) , \end{equation*} where $\{\mbox{\boldmath${\Gamma}$}(s),s\geq 0\}$ is a Gaussian process with $E\mbox{\boldmath${\Gamma}$}(s)=\mathbf{0}$ and $E\mbox{\boldmath${\Gamma}$} (s)\mbox{\boldmath${\Gamma}$}^{\top }(t)=\min (t,s)\mathbf{D}.$ We write \begin{equation*} U_{2,m}=\max_{1\leq j\leq x(k^{\ast })^{1/2}}\left( \frac{k^{\ast }+j}{m} \right) ^{1-\eta }\left( {\mathcal{r}}_{m,k^{\ast }}+j\mbox{\boldmath${ \Delta}$}+[{\mathcal{r}}_{m,k^{\ast }}-{\mathcal{r}}_{m,k}]\right) ^{\top } \hat{\mathbf{D}}_{m}^{-1}\left( {\mathcal{r}}_{m,k^{\ast }}+j \mbox{\boldmath${\Delta}$}+[{\mathcal{r}}_{m,k^{\ast }}-{\mathcal{r}} _{m,k}]\right) ^{\top }. \end{equation*} Thus we get \begin{equation*} \left( U_{m,1},U_{m,2}\right) \overset{{\mathcal{D}}}{\rightarrow }\left( \sup_{0<t\leq 1}\frac{\bar{u}^{1-\eta }}{t^{\eta }}\mbox{\boldmath${\Gamma}$} ^{\top }\mathbf{D}^{-1}\mbox{\boldmath${\Gamma}$}(t),\bar{u}^{1-\eta }\sup_{0\leq t\leq x}(t\mbox{\boldmath${\Delta}$}+\mbox{\boldmath${\Gamma}$} (1))^{\top }\mathbf{D}^{-1}(t\mbox{\boldmath${\Delta}$}+\mbox{\boldmath${ \Gamma}$}(1))\right) . \end{equation*} Next we assume that (ref) holds. We replace (ref) and (ref) with \begin{equation*} {\mathcal{r}}_{m,k}(\hat{\mbox{\boldmath${\theta}$}}_{m})=(k-k^{\ast }) \mbox{\boldmath${\Upsilon}$}+\tilde{\mathbf{m}}_{m,k}, \end{equation*} \begin{equation*} \tilde{\mathbf{m}}_{m,k}={\mathcal{r}}_{m,k^{\ast }}(\hat{ \mbox{\boldmath${\theta}$}}_{m})+[{\mathcal{r}}_{m,k}(\hat{ \mbox{\boldmath${\theta}$}}_{m})-{\mathcal{r}}_{m,k^{\ast }}(\hat{ \mbox{\boldmath${\theta}$}}_{m})-(k-k^{\ast })\mbox{\boldmath${\Upsilon}$}], \end{equation*} \begin{equation*} {\mathcal{D}}_{m,k}=(k-k^{\ast })^{2}B_{m}+2(k-k^{\ast })\mbox{\boldmath${ \Upsilon}$}^{\top }\hat{\mathbf{D}}_{m}^{-1}\tilde{\mathbf{m}}_{m,k}+\tilde{ \mathbf{m}}_{m,k}^{\top }\hat{\mathbf{D}}_{m}^{-1}\tilde{\mathbf{m}}_{m,k}. \end{equation*} Following the same arguments as in the case of (ref), we get that if $\bar{u}=0$, then \begin{align*} & \lim_{m\rightarrow \infty }P\{\tau _{m}>\tilde{{\mathcal{M}}}\} \\ & =\lim_{m\rightarrow \infty }P\left\{ 2(\tilde{{\mathcal{M}}}-k^{\ast })[ \mbox{\boldmath${\Upsilon}$}^{\top }\hat{\mathbf{D}}_{m}^{-1}\tilde{\mathbf{m }}_{m,\tilde{{\mathcal{M}}}}]\tilde{{\mathcal{M}}}^{-3/2}\leq \left[ { \mathcal{c}}{\mathcal{n}}^{1-\eta }-\frac{1}{\tilde{{\mathcal{M}}}^{\eta }} \left( \tilde{{\mathcal{M}}}-k^{\ast }\right) ^{2}B_{m}\right] \tilde{{ \mathcal{M}}}^{-3/2+\eta }\right\} , \end{align*} and the definition of $\tilde{{\mathcal{M}}}$ depends on if $\tilde{t}^{\ast }<\infty $ is finite or not. If $\tilde{t^{\ast }}<\infty $ we use \begin{equation*} \tilde{{\mathcal{M}}}=k^{\ast }+\tilde{u}_{{\mathcal{n}}}\left( {\mathcal{n}} ^{1-\eta }\frac{{\mathcal{c}}}{B_{m}}\right) ^{1/(2-\eta )}+\tilde{R}, \end{equation*} with \begin{equation*} \tilde{R}=x{(\tilde{u}^{\ast }+\tilde{t}^{\ast })^{3/2}}\left( \left( { \mathcal{n}}^{1-\eta }\frac{{\mathcal{c}}}{B_{m}}\right) ^{1/(2-\eta )}\right) ^{1/2}. \end{equation*} Using the definition of $\bar{{\mathcal{M}}}$ we get \begin{equation} \lim_{m\rightarrow \infty }\left[ {\mathcal{c}}{\mathcal{n}}^{1-\tau }-\frac{ 1}{\tilde{{\mathcal{M}}}^{\tau }}\left( \tilde{{\mathcal{M}}}-k^{\ast }\right) ^{2}B_{m}\right] \tilde{{\mathcal{M}}}^{-3/2+\tau }=-2x. \end{equation} Applying (ref) and Lemma (ref) with the independence of the approximating Gaussian processes we conclude \begin{equation} (\tilde{{\mathcal{M}}}-k^{\ast })[\mbox{\boldmath${\Upsilon}$}^{\top }\hat{ \mathbf{D}}_{m}^{-1}\tilde{\mathbf{m}}_{m,\tilde{{\mathcal{M}}}}]\tilde{{ \mathcal{M}}}^{-3/2}\overset{{\mathcal{D}}}{\rightarrow }\bar{s}_{1}{ \mathcal{N}}, \end{equation} where ${\mathcal{N}}$ is a standard normal random variable and $\bar{s}_{1}$ is defined in (ref). The covariance matrix $\mathbf{D}$ is defined in (ref). If $\tilde{t}^{\ast }=\infty $ and $\bar{u}=0$, then we use \begin{equation*} {\mathcal{M}}=k^{\ast }+\left( \frac{{\mathcal{c}}}{B_{m}}{\mathcal{n}} ^{1-\eta }(k^{\ast })^{\eta }\right) ^{1/2}+\tilde{R},\quad \;\mbox{with} \quad \;\tilde{R}=\frac{x}{B_{m}}(k^{\ast })^{1/2}. \end{equation*} We still have (ref) but (ref) is replaced with \begin{equation*} (\tilde{{\mathcal{M}}}-k^{\ast })[\mbox{\boldmath${\Upsilon}$}^{\top }\hat{ \mathbf{D}}_{m}^{-1}\tilde{\mathbf{m}}_{m,\tilde{{\mathcal{M}}}}]\tilde{{ \mathcal{M}}}^{-3/2}\overset{{\mathcal{D}}}{\rightarrow }\bar{s}_{2}{ \mathcal{N}}, \end{equation*} where $\bar{s}_{2}$ is given in (ref). As in the previous case, if $0< \bar{u}<1$, then we use again ${\mathcal{M}}=k^{\ast }+x(k^{\ast })^{1/2}$. Since the sum of gradients (after removing the mean) of the log likelihood function of the observations after the change is negligible, we get that \begin{equation*} \max_{1\leq k\leq {\mathcal{M}}}\frac{{\mathcal{D}}_{m}(k)}{g_{m}(k)}\overset {{\mathcal{D}}}{\rightarrow }\bar{u}^{1-\eta }\max \left( \sup_{0<t\leq 1} \frac{1}{t^{\eta }}\mbox{\boldmath${\Gamma}$}^{\top }(t)\mathbf{D}^{-1} \mbox{\boldmath${\Gamma}$}(t),\sup_{0\leq t\leq x}(t\mbox{\boldmath${ \Upsilon}$}+\mbox{\boldmath${\Gamma}$}(1)^{\top }\mathbf{D}^{-1}(t \mbox{\boldmath${\Upsilon}$}+\mbox{\boldmath${\Gamma}$}(t))\right) , \end{equation*} where recall that $\{\mbox{\boldmath${\Gamma}$}(t),t\geq 0\}$ is a Gaussian process $E\mbox{\boldmath${\Gamma}$}(t)=\mathbf{0}$ and $E \mbox{\boldmath${\Gamma}$}(t)\mbox{\boldmath${\Gamma}$}^{\top }(s)=\min (t,s) \mathbf{D}$. Next we consider the R\'{e}nyi type detector. First we investigate the case when after the change we have an (asymptotically) stationary sequence. Assume that $k^{\ast }\leq r$, i.e.\ the change occurred before the monitoring started. By definition, \begin{equation*} P\{\bar{\tau}_{m}=r\}=P\left\{ \frac{{\mathcal{D}}_{m}(r)}{\bar{g}_{m}(r)} \geq 1\right\} . \end{equation*} The proof is based again on (ref) and (ref). It follows from Lemma (ref) that \begin{equation*} \frac{1}{(k^{\ast })^{1/2}}\Vert {\mathcal{r}}_{m,k^{\ast }}\Vert =O_{P}(1) \end{equation*} and if $k^{\ast }\rightarrow \infty $ and $r-k^{\ast }\rightarrow \infty $, then \begin{equation*} \left( \frac{1}{(k^{\ast })^{1/2}}{\mathcal{r}}_{m,k^{\ast }},\frac{1}{ (r-k^{\ast })^{1/2}}\left[ {\mathcal{r}}_{m,r}(\hat{\mbox{\boldmath${ \theta}$}}_{m})-{\mathcal{r}}_{m,k^{\ast }}(\hat{\mbox{\boldmath${\theta}$}} _{m})-(r-k^{\ast })\mbox{\boldmath${\Delta}$}\right] \right) \overset{{ \mathcal{D}}}{\rightarrow }\left( \mathbf{N}_{1},\mathbf{N}_{2}\right) \end{equation*} where $\mathbf{N}_{1},\mathbf{N}_{2}$ are independent standard normal random variables with covariance matrices $\mbox{\boldmath${\Sigma}$}_{1}$ and $ \mbox{\boldmath${\Sigma}$}_{2}$ defined in (ref) and (ref), respectively. Thus we get \begin{equation*} \frac{1}{r^{1/2}}\left( (r-k^{\ast })\mbox{\boldmath${\Delta}$}+\mathbf{m} _{m,r}\right) \overset{{\mathcal{D}}}{\rightarrow }{\mathcal{a}}^{1/2} \mathbf{N}_{1}+{\mathcal{b}}+\mathbf{N}_{2}. \end{equation*} If ${\mathcal{a}}<\infty $, then we use \begin{equation*} {\mathcal{M}}=k^{\ast }+xr^{1/2}. \end{equation*} It follows from the proof of Theorem (ref) that for all $K>0$ \begin{align*} & \left\{ \frac{1}{r^{1/2}}{\mathcal{r}}_{m,rt}(\hat{\mbox{\boldmath${ \theta}$}}_{m}),\frac{1}{r^{1/2}}\left( {\mathcal{r}}_{m,k^{\ast }}+[{ \mathcal{r}}_{m,xr^{1/2}}-{\mathcal{r}}_{m,k^{\ast }}-xr^{1/2}]\right) ,t\geq 0,x\geq 0\right\} \\ & \overset{{\mathcal{D}}[0,{\mathcal{a}}]\times \lbrack 0,K]}{ \longrightarrow }\left\{ \mbox{\boldmath${\Gamma}$}_{1}(t), \mbox{\boldmath${\Gamma}$}_{1}({\mathcal{a}})+x\mbox{\boldmath${\Delta}$} \right\} \end{align*} where $\{\mbox{\boldmath${\Gamma}$}_{1}(t),t\geq 0\}$ is Gaussian process with $E\mbox{\boldmath${\Gamma}$}_{1}(t)=\mathbf{0}$, $E\mbox{\boldmath${ \Gamma}$}_{1}(t)\mbox{\boldmath${\Gamma}$}_{1}^{\top }(s)=\min (t,s) \mbox{\boldmath${\Sigma}$}_{1}$ with $\mbox{\boldmath${\Sigma}$}_{1}$ defined in (ref). Thus we conclude \begin{align*} \sup_{r\leq k\leq {\mathcal{M}}}& \frac{{\mathcal{D}}_{m}(k)}{g_{m}(k)}= \frac{1}{{\mathcal{c}}}\max \left( \max_{r\leq k\leq k^{\ast }}\frac{{ \mathcal{D}}_{m}(k)}{(k/r)^{\eta }r},\;\max_{k^{\ast }<k\leq {\mathcal{M}}} \frac{{\mathcal{D}}_{m}(k)}{(k/r)^{\eta }r}\right) \\ & \overset{{\mathcal{D}}}{\rightarrow }\frac{1}{{\mathcal{c}}{\mathcal{a}} ^{\eta }}\max \left( \sup_{1\leq t\leq {\mathcal{a}}}\mbox{\boldmath${ \Gamma}$}_{1}^{\top }(t)\mathbf{D}^{-1}\mbox{\boldmath${\Gamma}$} _{1}(t),\;\max_{0\leq s\leq x}(\mbox{\boldmath${\Gamma}$}_{1}({\mathcal{a}} )+s\mbox{\boldmath${\Delta}$})^{\top }\mathbf{D}^{-1}(\mbox{\boldmath${ \Gamma}$}_{1}({\mathcal{a}})+s\mbox{\boldmath${\Delta}$}).\right) . \end{align*} For the final case ${\mathcal{a}}=\infty $ we use \begin{equation*} {\mathcal{M}}=k^{\ast }+\left( \frac{{\mathcal{c}}}{A_{m}}\frac{(k^{\ast })^{\eta }}{r^{\eta -1}}\right) ^{1/2}+x(k^{\ast })^{1/2}. \end{equation*} We observe that \begin{equation} \frac{k^{\ast }}{((k^{\ast })^{\eta }/r^{\eta -1})^{1/2}}=\left( \left( \frac{k^{\ast }}{r}\right) ^{2-\eta }r\right) ^{1/2}\rightarrow \infty \end{equation} and \begin{equation} \frac{(k^{\ast })^{\eta }/r^{\eta -1}}{k^{\ast }}=\left( \frac{k^{\ast }}{r} \right) ^{\eta -1}\rightarrow \infty . \end{equation} As in the previous cases, \begin{align*} \lim_{m\rightarrow \infty }& P\{\bar{\tau}_{m}>{\mathcal{M}}\} \\ & =\lim_{m\rightarrow \infty }P\left\{ 2({\mathcal{M}}-k^{\ast }) \mbox{\boldmath${\Delta}$}^{\top }\hat{\mathbf{D}}_{m}^{-1}\mathbf{m}_{m,{ \mathcal{M}}}<\left[ {\mathcal{c}}r^{1-\eta }-\frac{({\mathcal{M}}-k^{\ast })^{2}A_{m}}{{\mathcal{M}}^{\eta }}\right] {\mathcal{M}}^{\eta }\right\} \\ & =\lim_{m\rightarrow \infty }P\left\{ 2{\mathcal{M}}^{-1/2} \mbox{\boldmath${\Delta}$}^{\top }\hat{\mathbf{D}}_{m}^{-1}\mathbf{m}_{m,{ \mathcal{M}}}<\left[ {\mathcal{c}}r^{1-\eta }-\frac{({\mathcal{M}}-k^{\ast })^{2}A_{m}}{{\mathcal{M}}^{\eta }}\right] \frac{1}{{\mathcal{M}}-k^{\ast }}{ \mathcal{M}}^{-1/2+\eta }\right\} . \end{align*} We obtain from (ref) and (ref) that \begin{equation*} \left[ {\mathcal{c}}r^{1-\eta }-\frac{({\mathcal{M}}-k^{\ast })^{2}A_{m}}{{ \mathcal{M}}^{\eta }}\right] \frac{1}{{\mathcal{M}}-k^{\ast }}{\mathcal{M}} ^{-1/2+\eta }\rightarrow -2x. \end{equation*} We get from (ref) that \begin{equation*} {\mathcal{M}}^{-1/2}\mbox{\boldmath${\Delta}$}^{\top }\hat{\mathbf{D}} _{m}^{-1}\mathbf{m}_{m,{\mathcal{M}}}\overset{{\mathcal{D}}}{\rightarrow } \bar{s}_{2}{\mathcal{N}}, \end{equation*} where ${\mathcal{N}}$ is a standard normal random variable and $\bar{s}_{2}$ is defined in (ref). We now conclude \begin{equation*} \lim_{m\rightarrow \infty }P\{\bar{\tau}_{m}>{\mathcal{M}}\}=1-\Phi (x/\bar{s }_{2}). \end{equation*} The proof when (ref) holds is essentially the same so the details are omitted.