Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
35,616 characters · 7 sections · 27 citation commands
Improving the accuracy of bubble date estimators under time-varying volatility
\baselineskip= 6mm
Non-stationary volatility is sometimes observed in time series (in particular, financial data) but discussion of the break dates estimators under non-stationary volatility has limited attention in the literature. One of the exceptions is harris2020level, in which the estimation of level shift was improved by correcting the original time series by non-parametrically estimated time varying variance. While the explosive bubble model was proposed by PWY2011 and extended by phillips2015a,phillips2015b and harvey2017improving, in which the time series is generated by a unit root process followed by an explosive regime that is again followed by a unit root regime (or with a possible stationary correction market in a recovery regime), the importance of non-stationary volatility accommodation in bubble detection methods was discussed by HLST2016 and phillips2020real, the latter of which proposed a modification of the wild bootstrap recursive algorithm (based on the expanding sample) of HLST2016 for obtaining the dates of the bubble(s) and also addressed the multiplicity testing problem. harvey2020sign considered the minimization of the sign based statistic for obtaining the dates of the bubble but did not provide any finite sample performance. On the other hand, as discussed in harvey2017improving and pang2021estimating (PDC hereafter), the break dates estimators based on the minimization of the sum of the squared residuals are more accurate than the recursive method of phillips2015a,phillips2015b under the assumption of homoskedasticity. Nevertheless, as far as we know, there are no studies which accommodate the non-stationary volatility behaviour into the estimation of the bubble dates based on the minimization of the sum of the squared residuals.
Recently, PDC and kurozumi2022asymptotic investigated the asymptotic behaviour of the bubble date estimators. In particular, they obtained the consistency of the collapsing date estimator by minimizing the sum of the squared residuals using the two-regime model (even though the true model has four regimes), allowing non-stationary volatility. Due to the consistency, one could split the whole sample at the estimated break date and consider the estimation of the date of the origination of the bubble using the sample before the estimated collapsing date and the date of the market recovery using the sample after the estimated collapsing date. This sample splitting approach closely resembles that of harvey2017improving by minimizing the full SSR based on the four regimes model, but computationally less involved and, as PDC demonstrated, performs better in terms of estimation accuracy of the break dates. On the contrary to the collapsing date of the bubble, the consistency of the dates of the origination of the bubble and the market recovery depend on the extent of the explosive regime and collapsing regime. In other words, if the explosive speed is not sufficiently fast, then PDC and kurozumi2022asymptotic obtained only the consistency of the estimators of the break fractions, not the break date.
In this paper, we propose a two-step algorithm for estimating the emerging date, the collapasing date, and the recovering date of a bubble under non-stationary volatility. First, due to the consistency of the break dates (fractions) estimators regardless of heteroskedasticity, we estimate these break dates as proposed by PDC and kurozumi2022asymptotic and collect the residuals of the fitted four-regime model. Second, we estimate non-parametrically the time-varying error variance from these residuals and perform the GLS-based sample splitting approach, which minimizes the weighted SSRs. Monte-Carlo simulations demonstrate the performance of our correction method for a model with a one time break in volatility, especially when this break occurs at the beginning or the end of the sample. The empirical application consists of different time series of cryptocurrencies for which the two methods of identifying the bubble dates are performed: One without volatility correction and another with volatility correction.
The remainder of this paper is organized as follows. Section 2 formulates the model and assumptions. In Section 3, we define the main GLS-based procedure under a general type of weights. The choice of the specific weights are discussed in Section 4 and the new two-step algorithm is proposed. The finite sample performance of the estimated break dates is demonstrated in Section 5, and the empirical example is given in Section 6. Section 7 concludes the paper.
Let us consider the following bubble's emerging and collapsing model for $t=1,2,\ldots,T$:
where $y_0=o_p(T^{1/2})$, $c_0\geq 0$, $\eta_0 > 1/2$, $\phi_a > 1$, $\phi_b < 1$, $c_1\geq 0$, and $\eta_1 > 1/2$. We assume that the market is normal in the first and last regimes in the sense that the time series $y_t$ is a unit root process (a random walk) with possibly positive drift shrinking to 0. The process starts exploding at $t=k_e+1$ at a rate of $\phi_a$, which is typically only slightly greater than one and thus sometimes characterized as a mildly explosive specification. The explosive behavior stops at $t=k_c$ and $y_t$ is collapsing at a rate of $\phi_b < 1$ in the next regime, followed by the normal market regime. This model can be seen as a structural change model with the break points being given by $k_e$, $k_c$, and $k_r$. The corresponding break fractions are defined as $\tau_e\coloneqq k_e/T$, $\tau_c\coloneqq k_c/T$, and $\tau_r\coloneqq k_r/T$, respectively. We would like to estimate these break dates as accurately as possible.
For model (ref), we make the following assumption.
By Assumption (ref), the break fractions are distinct and not too close each other. Assumption (ref) allows for various kinds of nonstationary unconditional volatility in the shocks, such as a volatility shift (possibly multiple times) and linear and non-linear transitions. Under Assumption (ref), it is well known that the functional central limit theorem (FCLT) holds for the partial sum process of $\{\varepsilon_t\}$ normalized by $\sqrt{T}$, which weakly converges to a variance transformed Brownian motion as shown by CavaliereTaylor2007a,CavaliereTaylor2007b.
Following PDC and kurozumi2022asymptotic, we estimate the break dates one at a time. As model (ref) can be expressed as
PDC and kurozumi2022asymptotic proposed to fit the one-time structural change model without a constant and to estimate the break point by minimizing the sum of the squared residuals. It is shown that the estimated break date, $\hat{k}_c$, is consistent for $k_c$. We then split the whole sample into the two subsamples, and from the fist subsample before $\hat{k}_c$, the emerging date of the explosive behavior is estimated by fitting a one-time structural change model again, while $k_r$ is estimated from the second subsample after $\hat{k}_c$. These estimated break fractions, $\hat{\tau}_e\coloneqq\hat{k}_e/T$ and $\hat{\tau}_r\coloneqq\hat{k}_r/T$, are shown to be consistent and further, $\hat{k}_e$ ($\hat{k}_r$) is consistent for $k_e$ ($k_r$) if, roughly speaking, $\phi_a$ deviates from 1 sufficiently ($\phi_a-1 > 1-\phi_b$). See PDC and kurozumi2022asymptotic for details.
Although the above estimated break dates (fractions) are consistent under nonstationary volatility in Assumption (ref), the efficiency gain would be expected by estimating the break dates based on the weighted sum of the squared residuals (SSR). To be more precise, let $\delta_t$ be a generic series of weights $\delta_t$ and then the weighted SSR based on a one-time structural change model is given by
where $\sum_{t=\ell}^m$ is abbreviated just as $\sum_{\ell}^m$. As $SSR(k,\delta_t,\phi_a,\phi_b)$ is minimized at \[ \hat{\phi}_a(k,\delta_t)\coloneqq\frac{\sum_1^k y_{t-1}y_t{\delta}_t^{-2}}{\sum_1^k y_{t-1}^2{\delta}_t^{-2}} \quad\mbox{and}\quad \hat{\phi}_b(k,\delta_t)\coloneqq\frac{\sum_{k+1}^T y_{t-1}y_t{\delta}_t^{-2}}{\sum_{k+1}^T y_{t-1}^2{\delta}_t^{-2}} \] for given $k$ and $\delta_t$, the estimator of $k_c$ is given by \[ \hat{k}_c(\delta_t)\coloneqq \arg\min_{\underline{\tau}_c \leq k/T \leq \overline{\tau}_c} SSR(k,\delta_t), \] where $0<\underline{\tau}_c<\tau_c<\overline{\tau}_c < 1$ and $SSR(k,\delta_t)\coloneqq SSR(k,\delta_t,\hat{\phi}_a (k,\delta_t),\hat{\phi}_b(k,\delta_t))$. The corresponding break fraction estimator is defined as $\hat{\tau}_c(\delta_t)\coloneqq \hat{k}_c(\delta_t)/T$.
Once we obtained the estimator of $k_c$, we can move on to the estimation of $k_e$ and $k_r$. For $k_e$, the estimation is based on the minimization of the weighted sum of the squared residuals using the first sub-sample, and the estimator is defined as \[ \hat{k}_e(\delta_t)\coloneqq \arg\min_{\underline{\tau}_e\leq k/T\leq \overline{\tau}_e} SSR_1(k,\delta_t) \] where $0 <\underline{\tau}_e < \tau_e < \overline{\tau}_e <{\hat{\tau}}_c$ and \[ SSR_1(k,\delta_t) \coloneqq \sum_{1}^k{\delta}_t^{-2}\left(y_t-\hat{\phi}_c(k,\delta_t)y_{t-1}\right)^2+\sum_{k+1}^{\hat{k}_c(\delta_t)}{\delta}_t^{-2}\left(y_t-\hat{\phi}_d(k,\delta_t)y_{t-1}\right)^2 \] \[ \mbox{with}\quad \hat{\phi}_c(k,\delta_t)\coloneqq\frac{\sum_{1}^k y_{t-1}y_t{\delta}_t^{-2}}{\sum_{1}^k y_{t-1}^2{\delta}_t^{-2}}\quad\mbox{and}\quad \hat{\phi}_d(k,\delta_t)\coloneqq\frac{\sum_{k+1}^{\hat{k}_c(\delta_t)} y_{t-1}y_t{\delta}_t^{-2}}{\sum_{k+1}^{\hat{k}_c(\delta_t)} y_{t-1}^2{\delta}_t^{-2}}. \] The corresponding break fraction estimator is defined as $\hat{\tau}_e(\delta_t)\coloneqq \hat{k}_e(\delta_t)/T$. For notational convenience, we suppressed the dependence of $\hat{k}_e(\delta_t)$, $\hat{\tau}_e(\delta_t)$, and $SSR_1(k,\delta_t)$ on $\hat{k}_c(\delta_t)$.
On the other hand, for the estimation of $k_r$, we minimize the weighted sum of the squared residuals using the second sub-sample, and the estimator is defined as \[ \hat{k}_r(\delta_t)\coloneqq \arg\min_{\underline{\tau}_r\leq k/T\leq \overline{\tau}_r} SSR_2(k,\delta_t) \] where $\hat{\tau}_c <\underline{\tau}_r <\tau_r < \overline{\tau}_r <1$ and \[ SSR_2(k,\delta_t) \coloneqq \sum_{\hat{k}_c(\delta_t)+1}^k{\delta}_t^{-2}\left(y_t-\hat{\phi}_e(k,\delta_t)y_{t-1}\right)^2+\sum_{k+1}^T{\delta}_t^{-2}\left(y_t-\hat{\phi}_f(k,\delta_t)y_{t-1}\right)^2 \] \[ \mbox{with}\quad \hat{\phi}_e(k,\delta_t)=\frac{\sum_{\hat{k}_c(\delta_t)+1}^k y_{t-1}y_t{\delta}_t^{-2}}{\sum_{\hat{k}_c(\delta_t)+1}^k y_{t-1}^2{\delta}_t^{-2}}\quad\mbox{and}\quad \hat{\phi}_f(k,\delta_t)=\frac{\sum_{k+1}^T y_{t-1}y_t{\delta}_t^{-2}}{\sum_{k+1}^T y_{t-1}^2{\delta}_t^{-2}}. \] The corresponding break fraction estimator is defined by $\hat{\tau}_r(\delta_t)\coloneqq \hat{k}_r(\delta_t)/T$.
We call the above method the sample splitting approach based on the weighted least least squares (WLS) method. Note that the special case where $\delta_t=1$ for all $t$, called the OLS method in this paper, corresponds to the estimation method by PDC with the sample ranging from 1 to $k_r$ and that by kurozumi2022asymptotic.
To implement the sample splitting approach based on the WLS method in pracitce, we need to choose a weight function $\delta_t$ appropriately. In our model, it is natural to choose the volatility function $\sigma_t$ as the weight $\delta_t$ to obtain the efficiency gain but such a WLS estimation is infeasible because the volatility function is unknown. In this article, we follow xuphillips2008adaptive and estimate $\sigma_t$ by a kernel-based method. More precisely, we first estimate $\tau_e$, $\tau_c$, and $\tau_r$ by the sample splitting approach based on the OLS method ($\delta_t=1$ for all $t$) as proposed by PDC and kurozumi2022asymptotic. Then, using the estimated break dates denoted as $\hat{\tau}_e(1)$, $\hat{\tau}_c(1)$, and $\hat{\tau}_r(1)$, we estimate
by the least squares method and obtain the residuals $\hat{e}_t$, where $D_t(a,b)=\mathbb I(\lfloor aT\rfloor<t\leq\lfloor bT\rfloor)$ with $\mathbb{I}(\cdot)$ being the indicator function. Next, $\hat{\sigma}_t^2$ is calculated as
$K(\cdot)$ is a bounded nonnegative continuous kernel function defined on the real line with $\int_{-\infty}^{\infty}K(s)ds=1$, and $b$ is a bandwidth parameter. Finally, by plugging $\hat{\sigma}_t^2$ into $\delta_t^2$ in the sample splitting approach, we obtain the estimators $\hat{\tau}_e(\hat{\sigma}_t)$, $\hat{\tau}_c(\hat{\sigma}_t)$, and $\hat{\tau}_c(\hat{\sigma}_t)$. xuphillips2008adaptive showed that the estimation accuracy of the coefficient in a stable autoregressive model improves by the adaptive (WLS) estimation and we investigate if it works for the estimation of the bubble's dates in the next section.
In this section, we examine the performance of the estimates of the bubble regimes dates in finite samples if the error variance is subject to changes in volatility.
The Monte-Carlo simulations reported in this section are based on the series generated by (ref) with $y_0=1500$ and $\{\varepsilon_t\}\sim IIDN(0,1)$. Data are generated from this DGP for samples of $T=(400,800)$ with $50,000$ replications.\footnote{All simulations were programmed in R with rnorm random number generator.} We set the drift terms in the first and fourth regimes to $c_0T^{-\eta_0}=1/800$ and $c_1T^{-\eta_0}=1/800$, respectively, following PDC. In this experiment, we focus on local to unit root behaviour characterized by $\phi_a=1+c_a/T$ and $\phi_b=1-c_b/T$ where $c_a$ takes values among $\{4,5,6\}$ whereas $c_b$ is fixed at $6$.
For the dates of bubble regimes, we use $(\tau_e,\tau_c,\tau_r)$ to be equal to (0.4,0.6,0.7). This setting seems to be empirically relevant considering Japanese stock price, its logarithm, US house price index, and cryptocurrencies. We consider the case with a one-time break in volatility at date $\tau $, so that the volatility function $\sigma_t$ has the following form: \[\sigma_t^2=\sigma_0^2+\delta (\sigma_1^2-\sigma_0^2)\mathbb I(t>\lfloor \tau T\rfloor)\] with $\sigma_1/\sigma_0$ takes values among $\{1/5, 5\}$ and $\tau$ takes values among $\{0.2, 0.8\}$.
As in kurozumi2022asymptotic, in the minimization of $SSR(k/T)$, $SSR_1(k/T)$, and $SSR_2(k/T)$, we excluded the first and last 5% observations from the permissible break date $k$. For example, when estimating $k_r$ based on $SSR_2(k/T)$, the permissible break date $k$ ranges from $\hat{k}_c+0.05T+1$ to $0.95T$. If the break date estimate $\hat{k}_c$ exceeds $0.95T$, then we cannot estimate $k_r$; we do not include such a case in any bins of the histogram and thus the sum of the heights of the bins does not necessarily equal one for $\hat{k}_r$ in some cases. Similarly, we cannot estimate $\hat{k}_e$ when $\hat{k}_c < 0.05T$. To save space, we pick up several selected cases in the following and the other cases are provided in the online appendix.
Figure 1 presents the histograms of $\hat{k}_c$ when $\tau=0.8$, $s_0/s_1=1/5$, and $T=400$. The left column shows the results based on the OLS based method, while the right column corresponds to the WLS based method. In this case, the process becomes more volatile at the end of the sample and thus it would be difficult to distinguish between the explosive and collapsing behavior (from $\tau=0.4$ to 0.7) and a random walk with high volatility (from $\tau=0.8$ to 1). As expected, the OLS method tends to incorrectly choose the end of the sample as the collapsing date when $c_a=4$ as shown in Figure 1(a), although the local peak is observed at around the true break fraction ($\tau_c=0.6$). As the size of the bubble ($c_a$) gets larger, the local peak becomes higher as is observed in Figures 1(c) and (e) (note that the vertical axis is different depending on the value of $c_a$). On the contrary, we can observe from Figures 1(b), (d), and (f) that the WLS method can estimate the collapsing date more accurately than the OLS method; the finite sample distribution has a mode at the true break fraction and the frequency of correctly estimating the true date by WLS is about twice of that by OLS.
Figure 2 shows the histograms of $\hat{k}_c$ when $\tau=0.2$, $s_0/s_1=5$, and $T=400$. In this case, there exists a unit root regime with high volatility at the beginning of the sample and thus it is expected that the histograms would have positive frequencies before $\tau=0.2$. In fact, this is the case as is observed in Figure 2, although the accuracy is much better than the case in Figure 1. Overall, the WLS based method can detect the true collapsing date more often than the OLS based method. For example, when $c_a=0.4$ and $c_b=0.6$, the relative frequency of correct detection of the true collapsing date rises from 0.25 to 0.35 by introducing the adaptive procedure. We can also observe that the WLS method incorrectly detect the collapsing date at the beginning of the sample less frequently than the OLS method.
Figure 3 presents the histograms of $\hat{k}_e$ when $\tau=0.2$, $s_0/s_1=5$, and $T=400$, which is the same case as in Figure 2. Overall, when the size of the bubble is small with $c_a=4$, it is difficult to estimate the emerging date ($\tau_e=0.4$) accurately, but for large values of $c_a$, the accuracy of $\hat{k}_e$ improves and the histograms has a peak at $0.4$ as in Figures 3(c)--(f). Again, in this case, the performance of the estimator based on the WLS method is better than that based on the OLS method.
The other results are briefly summarized in the online appendix. Overall, Monte-Carlo simulations demonstrate that the accuracy of the estimators of the break dates improves significantly in some cases, while in other cases we cannot find any difference between the distribution of the estimator based on the OLS method and that based on the WLS method. Because our volatility correction does not deteriorate the finite sample performance of the break dates estimators, we recommend using the sample splitting approach with the WLS based method in any cases.
In this section, we demonstrate the application of the two sample splitting approaches to the top largest cryptocurrencies by capitalization (btc, eth, xrp, xlm, bch, ltc, eos, bnb, ada, xtz, etc, xmr) for daily observations. In all cases, the closing price in US dollars at 00:00 GMT on the corresponding day is used. Recently, kurozumi2022time investigated the explosive behaviour of these time series and detected the explosiveness as well as non-stationary volatility behavior. We implemented the two estimation methods for each calendar year (365 observations from January 1 to December 31) from 2014 to 2019, if the data of the corresponding currencies are available in that year. We report only the cases where the two methods return the different estimates of the break dates, because the purpose of this section is to demonstrate how effective the WLS method is for identifying the dates of the explosive behavior. Therefore, we omit the cases where the break dates are the same in both the methods.
We found eight cases where at least one of the estimated break dates is different. The results are presented in Figures (ref)-(ref). In each figure, the black line shows the sample path of the corresponding cryptocurrency, the three red doted lines the estimated dates of the emergence, collapse, and recovery based on the OLS method, and the three blue dashed lines those estimated by the WLS method.
For xrp in 2014 in Figure (ref), the series is collapsing from the beginning of the sample and it seems to be explosive, at least by visual inspection, at the end of the sample. Clearly, our model (ref) with one explosive regime is not valid in the corresponding year. In such a case, both methods cannot identify the correct break dates. This example demonstrates that we should carefully choose the sample periods in which only one set of the four regimes should be included in the same order as in (ref).
Figure (ref) shows xrm in 2015. We can observe that the same collapsing and recovering dates of the explosive behavior are obtained by the two methods, whereas the estimated emerging date by the WLS is about one month earlier than that by the OLS, if we take nonstationary volatility into account.
Figure (ref) shows eth in 2016, which may becomes explosive twice by visual inspection. It seems that the WLS method successfully detect the explosive behavior of eth in early 2016, whereas the OLS method erroneously assigns the second peak of the process as the recovering date.
The currency xlm in 2017 is given in Figure (ref), in which there exists the small explosive behavior at the middle of the sample and the much large explosiveness is observed at the end of the sample, which is not compatible with our model (ref). Nevertheless, the WLS method seems to detect the first explosiveness well, whereas $\hat{k}_c$ estimated by the OLS method is no longer the collapsing date.
Figure (ref) for etc in 2017 and Figure (ref) for xmr in 2017 are similar to Figure (ref) in that the time series has two explosiveness in the sample. Again, for etc in 2017, the first exuberance is well identified by the WLS method whereas the OLS method seems to fail to accurately estimate the recovering date. On the other hand, it seems to be difficult to identify the break dates by the both methods for xmr in 2017.
Figure (ref) shows the sample path of xlm 2018, which has several small humps in this sample period. Although the three break dates estimated by the WLS may be interpreted as the emerging, collapsing, and recovering dates, they may not correspond to the one specific explosiveness but to some of the several humps. On the other hand, the three estimated dates by the OLS method cannot be interpreted as designated by theory.
Figure (ref) shows bnb in 2018, in which large explosiveness is observed at the beginning of the sample and there seems to exists a mild explosive and collapsing behavior in most part of the sample. It seems that the WLS method captures this second behavior, although the collapsing regime is relatively short by taking volatility shift into account. It seems that the estimated collapsing date by the OLS method seems to be incorrect and it might be either the recovering date or the emerging date.
As a whole, the volatility correction by the WLS method seems to work well, except for several cases where the explosive behavior is observed more than twice. We also observed that the WLS method can be robust to the short explosiveness either at the beginning or end of the sample if there exists another exuberance in the middle of the sample, although it is desirable to set up the sample periods in which only one set of the exuberance is included. For that purpose, the procedure proposed by phillips2015a,phillips2015b may be useful.
We proposed the algorithm for volatility correction in estimation of the dates of the bubble in four-regime model. The method consists of the following steps: Estimation of the break dates without volatility correction; non-parametric estimation of the volatility function by replacing the true break dates with the estimated ones; WLS-based estimation of the dates of the bubble. The Monte-Carlo results show that the estimated break dates are at least as accurately as those under the homoskesasticity assumption and better in some cases. The empirical illustration using the cryptocurrencies demonstrates the different performance of the two methods, with and without volatility correction and that the WLS method returns the adequate break dates more often than the OLS method.
\setcounter{page}{1}