Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
222,518 characters · 6 sections · 34 citation commands
Testing for a Forecast Accuracy Breakdown under Long Memory
\def\spacingset#1{ {#1}} \spacingset{1}
\newtheorem{Theorem}{Theorem} \newtheorem{Proposition}{Proposition} \newtheorem{Proof}{Proof} \newtheorem{Remark}{Remark} \newtheorem{Assumption}{Assumption} \newtheorem{lemma}{Lemma} \newtheorem{Definition}{Definition}
\DeclarePairedDelimiter\ceil{\lceil}{\rceil} \DeclarePairedDelimiter\floor{\lfloor}{\rfloor}
\if00 \fi
\if10 {
} \fi
{\it Keywords:} Electricity prices; Forecast failure; Local power; Out-of-sample forecast; Structural change
\spacingset{1.8}
The presence of low frequency contaminations, such as structural breaks, poses a challenge in distinguishing true long memory from spurious long memory. This is due to the fact that both concepts share similar characteristics, such as a singularity in the periodogram near the zero Fourier frequency and significant autocorrelations at large lags (diebold2001long, granger2004occasional, mikosch2004nonstationarities). Tests for structural breaks are closely related to tests for forecast breakdowns, which test for structural breaks in a forecast error loss function. perron2006dealing reviews the extensive literature on structural breaks. Perr21 extend classical tests for structural breaks to the context of forecast failure. giacomini2009detecting propose a test to retrospectively assess whether a given forecast model provides stable forecasts by comparing in-sample and out-of-sample averages of a forecast error loss series. A forecast breakdown is defined as a decline in the out-of-sample performance of a forecasting model relative to its in-sample performance. To measure the forecast performance, the popular squared error loss function is used in this paper. Unlike structural break tests applied to the forecasting model, forecast accuracy breakdown tests allow for misspecified models and they can detect variance changes when a quadratic loss function is used. Furthermore, they assist the applied econometrician in determining the accuracy of a given forecast model and whether it requires modification. It is important to note, however, that the detection of a forecast breakdown does not necessarily indicate that the forecast model needs to be altered. For instance, christoffersen1997optimal demonstrate that the optimality of a point forecast does not depend on the variance of the errors when a symmetric loss function is used. Nevertheless, a change in the variance can cause a forecast breakdown, since tests for forecast failures cannot distinguish between symmetric and asymmetric loss functions. \\ In practical applications, it is natural to distinguish between in-sample and out-of-sample periods. In retrospective testing, however, an artificial separation between these periods is required, since the test period serves as a pseudo-out-of-sample period. The forecasting model is estimated in the in-sample period, and to assess whether the forecast accuracy changes, it should provide stable forecasts. To obtain forecasts, Perr21 recommend using a fixed forecast scheme, since it is the only forecast scheme that leads to a monotonic power function, mainly because it ensures the maximum difference between the in-sample and out-of-sample means of the loss series. Rolling and recursive schemes can induce power losses due to contamination problems. Under a fixed scheme, the test proposed by giacomini2009detecting is a Wald test for a single mean change in the total loss series at a predetermined date. The total loss series combines in-sample and out-of-sample losses. The test suffers from power losses if the break date is not specified precisely. To address this issue, rossi2012out propose a sup-Wald test that maximizes the giacomini2009detecting test over a predefined range of potential break dates. When multiple changes occur, the test still suffers from non-monotonic power functions. For this reason, Perr21 propose a double sup-Wald (DSW) test for a structural break in the mean of the out-of-sample loss series, estimating the breakpoint with the bai1998estimating estimator. The proposed method involves performing SW tests for each choice of the in-sample period. Ultimately, the DSW test statistic value is the largest value among all in-sample period choices. In this paper, we extend the DSW test to the long memory framework by using a memory and autocorrelation consistent (MAC) estimator of the long-run variance proposed by robinson1995gaussian and by incorporating a long memory robust estimator of the breakpoint. A related topic is discussed by paza2023optimal. They obtain optimal forecasts in a long memory time series setting under discrete structural breaks. To the best of our knowledge, our study is the first to connect long memory time series with forecast failures.\\ Krus18 derive theoretical evidence for a long memory transfer from a time series to forecasts and to forecast differentials, especially in the presence of biased forecasts. We show that this transfer also applies to the squared forecast error, which is a component of a forecast differential. We consider biased and unbiased forecasts, as well as the presence and absence of a fractional cointegration relationship between a long memory time series and its forecast.\\ The remainder of the paper is organized as follows. Section (ref) contains the theoretical derivation of the long memory transfer from the time series to the forecasts and to the squared forecast error. This justifies the need for a long memory robust forecast accuracy breakdown test, which is presented in Section (ref). To derive the finite sample size and power properties of the test, we perform a Monte Carlo study in Section (ref). Section (ref) applies the proposed test to European and U.S. electricity prices, while Section (ref) contains concluding remarks. All proofs are consolidated in the Appendix.
Before presenting the test to detect a change in forecast accuracy, we analyze the transfer of long memory from a time series and its forecast to the loss function under specific settings. We distinguish between biased and unbiased forecasts. Our analysis follows Krus18, who derive long memory transfer properties for forecast error loss differentials. The proofs of all propositions are provided in the Appendix. To create a framework, we restrict the time series and its forecast to be stationary long memory processes.
Assumption (ref) is convenient since dittmann2002properties derive long memory properties for squares and cross products of Gaussian long memory processes. We now distinguish between the presence of fractional cointegration between the time series and its forecast and the absence of this property. A fractional cointegration relationship refers to the existence of a linear combination of long memory time series that is integrated of reduced order. The memory transfer results are based on the asymptotic behavior of autocovariance functions of products and squares of long memory time series. In addition, Proposition 3 of chambers1998long can be used, as it states that the memory order of a linear combination of fractionally integrated series is equal to the maximum order of the individual components. leschinski2017memory generalizes this result for the broader class of long memory processes and discusses the memory of products of long memory time series. We derive the long memory transfer for the squared error loss function, which we also use for our test in Section (ref),
First, consider the case of the absence of fractional cointegration.
Under Assumption (ref), the long memory transfer is summarized in the following proposition.
Proof. See the Appendix.\\ The result is obtained by replacing $y_t$ and $\hat{y}_t$ by their centered series $a_t^* = a_t - \mu_a$ for $a_t \in \{y_t,\hat{y}_t\}$. As shown in the Appendix, the squared forecast error in Eq. ((ref)) can then be rearranged as
Proposition (ref) shows that if the forecast is biased, the memory order of the original terms dominates. If the forecast is unbiased, the first part of Eq. ((ref)) drops and the product and squared series remain. As shown in the Appendix, the memory orders of the squared series dominate the memory of the product series and the result follows.\\ We now relax Assumption (ref) and allow for fractional cointegration between the series and the forecast.
$d_x - \delta$ refers to the memory order of the linear combination of $y_t$ and $\hat{y}_t$, and they are fractionally cointegrated for $\delta > 0$. As can be seen from Assumption (ref), $x_t$ is the common factor that drives the long memory properties of $y_t$ and $\hat{y}_t$. We restrict the fractional cointegration to a form where the two series can be represented as linear functions of their common factor. In fact, our previous Assumption (ref) of the absence of fractional cointegration may be inappropriate in a forecasting setup. It follows directly from Assumption (ref) that
There are three cases to consider under fractional cointegration.
The memory transfer results for the three cases are obtained by substituting the relations from Assumption (ref) into the squared forecast error in Eq. ((ref)). By rearranging the resulting equation, we obtain an expression similar to that in Eq. ((ref)). We obtain the following result.
Proof. See the Appendix.\\ In case of a biased forecast with $\kappa_y \neq \kappa_{\hat{y}}$, the memory order of $x_t$ dominates. In the other two cases, the memory is reduced. Under a biased forecast and $\kappa_y = \kappa_{\hat{y}}$, the resulting memory parameter of the squared forecast error is $\tilde{d}$. The exact order of $\tilde{d}$ cannot be determined because squares and products of the errors $\epsilon_y$ and $\epsilon_{\hat{y}}$ are involved, for which the memory orders are unknown. In the remaining case of an unbiased forecast with $\kappa_y = \kappa_{\hat{y}}$, the forecast is an accurate prediction of the underlying series. Here, the memory of $x_t^{*^2}$ dominates, but this memory is also smaller than that of $x_t^*$, as shown in the Appendix.\\ In summary, our main results in Propositions (ref) and (ref) show that long memory can be transferred from the underlying time series to the forecast and finally, to the squared forecast error. The memory transfer depends crucially on the (un)biasedness of the forecast and the presence of fractional cointegration.
After deriving the long memory transfer from the time series to the forecast and finally to the squared forecast error, we now turn to the detection of a forecast failure. The accuracy of the point forecasts is evaluated in the out-of-sample period. We use the popular out-of-sample squared error loss function
where $L^o (m) = L_{m+\tau}^o,\ldots,L_{T}^o$ is the out-of-sample loss sequence of size $n = T-m-\tau+1$. Under the null hypothesis, \[ H_0: E[L_t^o] = \mu_0, \qquad \forall t = \tau + 1, \ldots, T-\tau +1, \] there is no structural break in the mean of the out-of-sample loss series. The alternative hypothesis is given by: \[ H_1: E[L_t^o] \neq E[L_{t+1}^o] \] for at least one $t = \tau + 1, \ldots, T-\tau$. Thus, our test is capable of detecting a single break in the forecast performance, but it does not rule out the existence of multiple breaks. We do not take into account the presence of multiple breaks because there is no long memory robust breakpoint estimator yet.\\ The underlying test for a break in the predictive accuracy is a Wald-type test proposed by Perr21. To compute the test statistic, the out-of-sample loss series is calculated for each in-sample period of size $m$ in an interval $[m_0,m_1]$, where the range between $m_0$ and $m_1$ is usually a fraction of the size of the largest out-of-sample period. The SW test is then applied to each out-of-sample loss series, and we take the maximum over all SW test statistics for all choices of the in-sample period, resulting in a DSW test. The test statistic is given by
where
$\varepsilon$ is a trimming parameter that is set to $0.1$ throughout the simulations and application. For a given choice of $m$, the test maximizes the difference between the demeaned restricted sum of squared residuals of the out-of-sample loss series $\left(SSR_{L^o(m)}\right)$ and the demeaned unrestricted SSR of the out-of-sample loss series $\left(SSR\left(T_b(m)\right)_{L^o(m)}\right)$. $T_b(m)$ denotes the breakpoint, which results from a long memory robust CUSUM test proposed by horvath1997effect for Gaussian processes and extended by wang2008change to general linear processes. Both papers derive the limit distributions that converge to the supremum of a fractional Brownian bridge. The difference between the SSR is standardized by the MAC long-run variance estimate of the sample mean proposed by robinson2005robust and abadir2009two, which is applied to the demeaned out-of-sample loss series with a break at time $T_b(m)$. The MAC estimator is defined as
where
and
is the periodogram, $\lambda_j = 2\pi j/n$ are the Fourier frequencies and the bandwidth $w$ converges to infinity and satisfies $w = o(n/(log\ n)^2)$. Finally,
The MAC estimator requires an estimate of the long memory parameter of the out-of-sample loss series under the alternative. For this purpose, we use the local Whittle estimator of kunsch1987statistical and robinson1995gaussian,
For the bandwidth choices, we follow the simulation results of abadir2009two and choose $w=\floor{n^{0.8}}$ for the bandwidth of the MAC estimator and $w_{\scriptscriptstyle LW}=\floor{n^{0.65}}$ for the bandwidth of the local Whittle estimator.\\ Since we are trying to detect a change in forecast accuracy, for each $m \in [m_0,m_1]$ the selected in-sample period should be stable in order to provide stable forecasts. To obtain stable forecasts, a fixed forecasting scheme should be chosen where the in-sample period consists of observations $1,\ldots,m$, because a rolling and a recursive scheme ensure that the mean of the in-sample losses approaches the mean of the out-of-sample losses rather quickly. This prevents forecast accuracy breakdown tests from generating power under a rolling and a recursive scheme. A fixed scheme may not optimize a given loss function, but that is not the goal of forecast breakdown tests.\\ We require the following assumption and additionally that $T, m\ \text{and}\ n$ converge to infinity at the same rate.
The limit distribution of the test is derived in the Appendix and summarized in the following proposition.
Proof. See the Appendix.
To examine the finite sample size and power properties of our proposed test, we conduct a Monte Carlo simulation study.
The critical values in Table (ref) are obtained under the null hypothesis with $5{,}000$ repetitions of $T=1{,}000$ observations of fractional white noise (FWN) processes with a fractional parameter $d \in \{0.1,0.2,0.3,0.4\}$, Gaussian innovations and a significance level of $10\%$, $5\%$, and $1\%$. In addition, we account for short-run dynamics by adding an autoregressive parameter $\phi \in \{0.1,0.2,0.3,0.4\}$ to the FWN processes. The corresponding critical values can be found in Table (ref) in the Appendix. The critical values, the size and the power results are obtained with $m_0 = \floor{0.2T}$, so that $20\%$ of the sample data is used for the smallest choice of the in-sample period, $\Bar{\mu}=0.3$, which implies that $30\%$ of the largest out-of-sample period is used for the window defining $m_1$, the size of the largest in-sample period. We consider a forecast horizon of $\tau=1$ and a trimming value of $\varepsilon=0.1$. Changing these values slightly does not significantly change the results. We only report results for a fixed forecast scheme because, similar to Perr21, we find that this is the only forecast scheme that ensures a monotonic power function. We elaborate on the loss of power under a recursive and a rolling forecast scheme in the supplementary material.\\ We conduct a local power analysis comparing the Perr21 test with the heteroscedasticity and autocorrelation consistent estimator of andrews1991heteroskedasticity for the long-run variance correction factor in Eq. ((ref)) and the bai1998estimating breakpoint estimator (hereafter abbreviated as HAC test) with our proposed test (hereafter abbreviated as MAC test). For the size and the power we use $2{,}000$ Monte Carlo repetitions and $T=500$ observations. Moreover, we let $d \in \{0.1,0.2,0.3,0.4\}$ and $\phi \in \{0,0.1,0.2,0.3,0.4\}$ and we consider mean shifts in the simulated autoregressive fractionally integrated moving average (ARFIMA) processes starting from the three hundredth observation of increasing size $k T^{d-1/2}$ for $k = 0,1,...,25$.\\
Table (ref) shows the empirical size for nominal significance levels of $10\%$, $5\%$, and $1\%$. On average, the empirical size of our MAC test is closer to the nominal size than the empirical size of the HAC test. At the same time, our test is slightly undersized on average, while the HAC test is oversized for all parameter combinations. The liberalness of the HAC test increases on average with increasing fractional parameter value. For increasing short-run dynamics, the liberalness of the HAC test seems to decrease, at least for $d=0.4$.\\
Figure (ref) shows the local power functions for FWN processes with a nominal significance level of 5%. Overall, the MAC test is superior in terms of power for all values of the fractional parameter. The power differences increase with increasing fractional parameter value and the HAC test with $d=0.4$ even loses power.\\
Figure (ref) introduces short-run dynamics of increasing size to the simulated FWN processes. Our test is still superior in terms of power, but the power of the HAC test increases overall. On average, both tests lose power when the short-run dynamics increase.\\ We also considered variance changes in the simulated ARFIMA processes, as giacomini2009detecting state that forecast accuracy breakdown tests with a squared error loss function can detect variance changes as well. However, as we have shown in Section (ref), a memory reduction to zero is possible when the forecast is unbiased. Both tests are still consistent under a variance break, but they perform similarly in terms of power, so we do not report the results.
The global energy crisis that commenced in 2021 resulted in a significant surge in electricity prices. In this analysis, we apply our proposed test to European and U.S. daily wholesale electricity prices with the objective of identifying differences in the severity and timing of the changes in forecasting performance in response to the crisis. The European electricity price series include data from Denmark, France, Germany, the Netherlands, and Norway. They are obtained from the German Federal Network Agency. The U.S. electricity price hubs considered are Mass Hub, Mid-C, Palo Verde, PJM West, and SP-15. The data set was obtained from the U.S. Energy Information System. The regions the price hubs cover can be found on the organization's website. It should be noted that additional U.S. price hubs are not considered in this analysis due to incomplete data. In order to facilitate comparison with the U.S. data, weekends have been removed from the European data set. In view of the sharp rise in inflation following the global pandemic, all price series were adjusted for inflation using an electricity-specific price index. The initial year of our analysis, 2015, has been designated as the base year. The electricity price indices for the European countries are obtained from the corresponding national statistical authorities\footnote{Statistics Denmark, The National Institute of Statistics and Economic Studies (France), Statistisches Bundesamt (Germany), Centraal Bureau voor de Statistiek (The Netherlands), Statistics Norway.}, while the corresponding index for the U.S. has been obtained from the U.S. Bureau of Labor Statistics.
Figure (ref) shows the inflation-adjusted electricity price series for Germany and the price hub PJM West, spanning from 01/01/2015 to 05/31/2024. It also illustrates the autocorrelation function (ACF) and the periodogram of the corresponding series. The remaining price series, ACFs and periodograms are shown in Figures (ref), (ref), (ref) and (ref) in the Appendix. The series are highly persistent with a pole at frequency zero, indicating that the series possess long memory properties. The characteristic weekly seasonality of electricity prices, which arises from the change in electricity demand and generation during the week as opposed to the weekend, is not evident in the data set, as weekends are not included. The seasonal component of electricity prices has already been analyzed by haldrup2006regime in a long memory framework. \\ For the application of our test, we select the smallest in-sample period ending on 01/01/2018 and the largest in-sample period ending on 12/31/2019. The critical values for each time series are obtained by simulating $5{,}000$ ARFIMA paths with $T=1{,}000$ observations, $m_0 = \floor{0.2T}$, $\Bar{\mu}=0.3$, $\tau=1$ and $\varepsilon=0.1$.
Table (ref) shows the critical values, the test statistic and the break date if the test statistic is significant at any significance level. The European price series are subject to a break in forecast accuracy that is between the end of July and the end of November 2021. Thus, the breaks are before the Russian invasion of Ukraine on 02/24/2022 and the corresponding European Union sanctions against Russia. The increase in global electricity prices in 2021 is due to increased demand for energy following the pandemic, which could not be met by supply. Another factor that may have had an impact on electricity prices in the European Union is the start of Phase IV of the EU ETS in 2021. Phase IV involves a higher, linearly increasing reduction factor for the EU-wide cap on emission allowances. As a result of the cap, the price per ton of CO$_2$ and the price of gas have risen sharply, while high gas prices drive up electricity prices through the merit order effect. For the U.S. price series, a break is identified in 2021 for Mass Hub and PJM West. The break in August 2020 for SP-15 is statistically significant at the $10 \%$ level.
In this paper, we propose a DSW test for detecting a forecast accuracy breakdown under a long memory time series setting. To account for the long memory properties, we standardize the test statistic with a MAC long-run variance estimate of the sample mean. The corresponding breakpoint is estimated with a long memory robust CUSUM test. The finite sample size and power properties of the test are derived in a Monte Carlo simulation. We show that our approach is superior in terms of size and power to a HAC estimate for the long-run variance if the underlying process exhibits long memory properties. The practical relevance of the method is demonstrated by applying the test to European and U.S. electricity prices. We find that European countries were more exposed to the global energy crisis that started in 2021.\\ There are ample opportunities to extend our flexible approach. The quadratic loss function may be replaced by any other loss function that the applied econometrician deems appropriate. Future research should also allow for multiple breaks, which may also occur in the in-sample period.