Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,266 characters · 8 sections · 0 citation commands
Testing for equal predictive accuracy with strong dependence
\affil[1]{University of York} \affil[2]{Universit\`{a} degli Studi di Milano}
JEL classification codes: C12; C32; C53. Keywords: strong autocorrelation, forecast evaluation, Diebold and Mariano test. Total Words: 5,744
Accurate forecasts are extremely important for forward-looking decision making. Weather forecasts often have a dedicated section even on the daily news, and predictions of the diffusion of the COVID-19 pandemic have had a critical impact on the life of most people. In economics, decisions over individual savings, firm-level investments, government fiscal policies and central bank monetary policies rely on forecasts of, among others, future economic activity and price levels.
To discriminate between good and bad forecasts, \citeasnoun{diebold1995comparing} [DM hereafter] suggested comparing alternative forecasts using a test for equal predictive ability. The DM test is based on a loss function associated with the forecast errors of each forecast, and allows to test the null hypothesis of zero expected loss differential between two competing forecasts. This approach takes forecast errors as model-free, and the test is valid also when the forecasts are produced from unknown models, as for example from forecast survey data. In addition, if the objective is to compare forecasting methods as opposed to forecasting models, then \citeasnoun{giacomini2006tests} showed that in an environment with asymptotically non-vanishing estimation uncertainty the DM test can still be applied.
The DM test allows us to test for equal predictive accuracy using any loss function, and the test statistic is asymptotically valid for contemporaneously correlated, serially correlated, and non-normal forecast errors. The test relies on the assumption that the loss differential is weakly dependent. The rationale for this assumption is that, under mean squared error (MSE) loss, optimal $q$-step ahead forecasts should generate at most MA($q-1$) errors. Thus if the considered forecast approximates the optimal forecast, its forecast errors should not be too correlated over time, although dependence beyond the MA($q-1$) boundary may take place. In practice, forecasts with errors that are fairly correlated can occur not only when the considered forecast fails to approximate the optimal forecast under MSE loss, but also when the forecast is optimal under an alternative loss function \citeaffixed{patton2007properties}{see} or it is evaluated on a relatively short sample.
Still, we can encounter situations in which the loss differential is highly autocorrelated even in the presence of a prediction with weakly dependent forecast errors and in samples of moderate size. This can happen when the DM test is used to compare the predictive ability of a selected forecast against a naive benchmark. This is a common practice, as naive benchmarks are cost-effective and readily available at any time, so they provide a standard reference for comparisons. Using simple benchmarks allows us to understand the added value of a specific forecasting technique, as it is desirable that predictions from sophisticated forecasting methods (for example complex models or expensive surveys) are more accurate than naive benchmarks. However, naive forecasts may in some cases generate relevant autocorrelation in the loss differential.
In this paper, we study the performance of the DM test when the assumption of weak autocorrelation of the loss differential does not hold. We characterise strong dependence as local to unity as in \citeasnoun{Phillips1987LocUnity} and \citeasnoun{PhillipsMagd2007MildlyIntegrated}. This definition is at odds with the more popular characterisation in the literature that treats strong autocorrelation and long memory as synonyms. Local to unity, however, seems well suited to derive reliable guidance when the sample is not very large, as it is the case in many applications in economics. With this definition the strength of the dependence is determined also by the sample size: a stationary AR(1) process with root close to unity may be treated as weakly dependent in a very large sample, but standard asymptotics may be a poor guidance for cases with smaller samples, and local to unity asymptotics may be more informative. We show that the power of the DM test decreases as the dependence increases, making it more difficult to obtain statistically significant evidence of superior predictive ability against less accurate benchmarks. We also find that after a certain threshold the test has no power and the correct null hypothesis is spuriously rejected. \Copy{R2c7t}{These results caution us to seriously consider the dependence properties of the loss differential before the application of the DM test, especially when naive benchmarks are considered. In this respect, a unit root test could be a valuable diagnostic for the preliminary detection of critical situations.}
To illustrate the problems associated with the DM test when there is dependence in the loss differential, we consider the case in which an AR(1) forecast for inflation in the Euro Area is compared to two naive benchmarks: a constant 2% prediction (that represents the inflation target in the Euro Area) and a rolling average prediction. These benchmark predictions have highly dependent forecast errors. As a consequence, the loss differential is dependent and the DM test fails to reject the null of equal predictive accuracy, even if the benchmarks are less accurate than the AR(1) forecast for short forecasting horizons.
In the literature, there has been some attention to the issue of forecast evaluation in presence of persistence. \citeasnoun{corradi2001predictive} examined the DM statistic in the presence of cointegration, whereas \citeasnoun{rossi2005testing} examined the effect of high persistence on the loss differential. \citeasnoun{mccracken2020diverging} provided an example to show that using a fixed and finite estimation window can result in loss differentials that depend on the first observations, so that the time series of the loss differential in the DM test is not ergodic for the mean. These works considered a framework with parameter estimation error; instead \citeasnoun{clark1999finite}, \citeasnoun{Khalaf2017} and \citeasnoun{kruse2019comparing} took forecast errors as primitives. \citeasnoun{clark1999finite} and \citeasnoun{Khalaf2017}, in particular, provided convincing simulation evidence that the DM test is incorrectly sized in the presence of dependence. \citeasnoun{giacomini2006tests} and \citeasnoun{coroneo2020comparing} showed that under benign forms of weak dependence, the correct size may be restored using bootstrap or fixed smoothing asymptotics, respectively, but results in \citeasnoun{Khalaf2017} suggest that even bootstrap is only a useful and reliable guidance when the depedence is not too strong, relative to the sample size. On the other hand \citeasnoun{kruse2019comparing} derived the properties of the DM test in the presence of long memory using standard asymptotics, and memory and autocorrelation consistent standardisation. \citeasnoun{kruse2019comparing} had a sample of 4883 observations, so long memory seems a reasonable modelling strategy in their case. However, this approach cannot be applied to moderate sample size, such as the ones typically encountered in macroeconomic forecasting.
Of course, the DM test can be seen as a particular application of the standard $t$-test on the mean in presence of dependence. A certain level of persistence can be accommodated for the $t$-test using bootstrap or alternative asymptotics, see for example \citeasnoun{gonccalves2011block} for the former, and \citeasnoun{kiefer2005new}, \citeasnoun{sun2014let}, \citeasnoun{sun2014fixed}, \citeasnoun{lazarus2018har} for the latter. Following \citeasnoun{muller2014hac} or \citeasnoun{GiraitisPhillips2012Local}, it is clear that the even the limit distribution in \citeasnoun{kiefer2005new} and related works does not provide a useful guidance when the persistence is too strong in relation to the sample size. \citeasnoun{muller2014hac} does provide an algorithm to perform a reliable test in that case, although the sample size required is larger than the samples that are usually available in forecast evaluation exercises. As a promising alternative option, \citeasnoun{henzi2021valid} and \citeasnoun{choe2023comparing} recently proposed procedures based on sequential testing that do not require assumptions on the dependence of the process. These new procedures may therefore be robust even in situations of dependence, however the assumption of bounded loss differentials may limit their applicability.
As it is clear from this review of the literature, the core of the statistical results discussed here should not come as a surprise. However, their implication in the context of testing for equal predictive ability remains relevant, in particular when good forecasts are compared against poor benchmarks, with relevant autocorrelations, as our results suggest that in these cases the blind application of the DM test leads to incorrect conclusions. Even more importantly, it may be more difficult to reject the null hypothesis when good forecasts are compared to poor competitors, than when the same good forecasts are compared to competitors that are nearly as good. This perverse and undesirable feature of the DM test should be kept in mind in forecast evaluation exercises.
The paper is organised as follows. We formally introduce the DM test in Section (ref), and derive the limit properties of the DM statistics in presence of dependence in Section (ref). We investigate the practical implication of our theoretical findings in a Monte Carlo exercise (Section (ref)) and in the empirical application (Section (ref)). Details on the assumptions of the DGP and formal derivations are in the Appendix.
The DM test was introduced to compare two forecasts of a time series, according to a user chosen loss metric. For $t=\{1, \hdots, T\}$, denoting the forecast errors as $e_{1,t}$ and $e_{2,t}$, respectively, and the loss function $L(.)$, \citeasnoun{diebold1995comparing} consider the loss differential
and test the null hypothesis of equal predictive ability $H_0: E(d_t)=0$ against the alternative $H_1: E(d_t) \neq 0$. The key assumptions by \citeasnoun{diebold1995comparing} and \citeasnoun{diebold2015comparing} are that $d_t$ is stationary and weakly dependent, and that the average loss, $\overline{d}=\frac{1}{T} \sum_{t=1}^{T}{d_t}$, follows a Central Limit Theorem. In particular, denoting $\mu=E(d_t)$, it is assumed that $\sqrt{T}{(\overline{d}-\mu)} \rightarrow_d {N(0,{\sigma^2})}$ as $T \rightarrow \infty$, where $0< \sigma^2<\infty $ is the long run variance of $d_t$.
Thus, inference on $E(d_t)$ can be based on the normalised limit
Denoting $\widehat{\sigma}^2$ an estimate of $\sigma^2$, $\widehat{\sigma}=\sqrt{\widehat{\sigma}^2}$, then the classical DM test uses the statistic
where the null hypothesis of equal predictive ability is rejected at 5% significance level against a two-sided alternative if the realization of $|DM|$ is above the 1.96 threshold.
The original DM test exploits the consistency of $\widehat{\sigma}^2$ to justify the standard normal as the limit distribution under the null. This may generate rather poor size performance in finite sample, see \citeasnoun{diebold1995comparing} and also \citeasnoun{clark1999finite}. With fixed smoothing asymptotics, the limit for $\widehat{\sigma}^2$ is derived under alternative asymptotics, accounting for the distribution of $\widehat{\sigma}^2$ and, as a consequence, the DM statistic does not have limit standard normal distribution, but the alternative limit provides a better approximation of the distribution of the DM statistic in finite samples. As the alternative distribution depends on the way $\sigma^2$ is estimated, we consider two cases: the weighted autocovariance estimate using the Bartlett kernel and the weighted periodogram estimate using the Daniell kernel.
Denoting by $\widehat{\gamma}_l$ the sample autocovariance of lag $l$, the weighted autocovariance estimate of the long run variance using the Bartlett kernel is
where $M$ is a user-chosen bandwidth parameter, and
Under $H_0$,
where the distribution of $\Phi_A(b)$ depends on $b$; this is characterised in \citeasnoun{kiefer2005new}, where relevant quantiles are also provided.
For the weighted periodogram estimate, denoting $w(\lambda)=\frac{1}{\sqrt{2 \pi T}} \sum_{t=1} ^ {T} d_t e^{i \lambda t}$ the Fourier transform of $d_t$ at frequency $\lambda$, and $I(\lambda)=|w(\lambda)|^2$ as the periodogram, the Daniell weighted periodogram estimate is
where, for integer $j$, $\lambda_j=\frac{2 \pi j}{T}$ are the Fourier frequencies and $m$ is a user-chosen bandwidth parameter. The test statistic is then given by
Under $H_0$,
where $t_{2m}$ is the Student's $t$-distribution with $2m$ degrees of freedom, see \citeasnoun{coroneo2020comparing} for more details.
The key assumption in the construction of the DM test is that $d_t$ is weakly dependent. This assumption seems reasonable in the context of forecasts, as it is well known that, under MSE loss, optimal $q-$step ahead forecasts should be at most $MA(q-1)$. For this very reason, \citeasnoun{diebold1995comparing} considered estimating $\sigma^2$ using only the first $q-1$ autocovariances of $d_t$, and verified that this assumption was met in the data in the empirical application that they presented.
However, in practice, it is not uncommon to have strong autocorrelation in the loss differential. This can happen in presence of forecasts that are optimal under alternative loss functions, or when the forecast evaluation sample $T$ is short. In addition, it is common practice to apply the DM test to test for equal predictive accuracy of a selected forecast against a naive benchmark, resulting in dependent forecast errors and loss differential, possibly even strongly autocorrelated. Section (ref) contains an example illustrating how stochastically trending behavior may appear in the loss differentials.
Denoting $\mu=E(d_t)$ and $y_t=d_t-\mu$, so that
we assume that
where $u_t$ is a zero mean, weakly dependent process with long run variance $\omega$. We consider two different models for $\rho_T$: in subsection (ref) we discuss the local-to-unity AR(1) approximation as in \citeasnoun{Phillips1987LocUnity} (alongside with the standard unit root model); in subsection (ref) we consider the moderate deviations from a unit root as in \citeasnoun{PhillipsMagd2007MildlyIntegrated}. These models are a convenient representation of dependence for $y_t$ when the dimension in $T$ is relatively short, as it is indeed the case in many empirical studies. In both cases, we refer to Appendix (ref) for a detailed presentation and discussion of the assumptions, and for the derivation of the results.
In this case, we assume for $\rho_t$ in (ref)
When $c$ is in the neighbourhood of $0$, $\rho_T$ is approximated as $1+{c/T}$, i.e. $ \rho_T \sim 1+{c/T}$ as $ c \rightarrow 0.$ When $c=0$ then the process $y_t$ has a unit root, and it is initialised setting the initial condition $y_0=O_p(1)$.
Our model is completed by formalising the assumptions on $u_t$.
Denoting $g(\lambda)$ as the spectral density of $u_t$, Assumption (ref) implies that $g(0)>0$ and that $g(\lambda)<\infty$ uniformly in $\lambda$. Assumption (ref) is sufficient to establish the functional central limit theorem (FCLT) for a stationary, weakly dependent linear process as in \citeasnoun{PhillipsSolo1992}, see remark 3.5 of \citeasnoun{PhillipsSolo1992}, and notice that condition $\sum_{j=0}^{\infty} j^{1/2}|\psi_j|<\infty$ implies (16) of \citeasnoun{PhillipsSolo1992}, as discussed on page 973.
Define
where $W(r)$ is a standard Brownian motion. The process $J_c(r)$ is a Ornstein-Uhlenbeck process: for given $r$ it is normally distributed (when $c=0$ the Ornstein-Uhlenbeck process is the standard Brownian motion). We refer to \citeasnoun{Phillips1987LocUnity} for a detailed discussion, but we state some important results from Lemma 1 of \citeasnoun{Phillips1987LocUnity}:
where $J_c(r)=W(r)$ when $c=0$. The limit in ((ref)) follows proceeding as in \citeasnoun{Phillips1987LocUnity} but using the FCLT for linear processes instead of for mixing sequencies (also see \citeasnoun{Chan_and_Wei}); ((ref)) and ((ref)) are then due to the continuous mapping theorem. The result for $\overline{y}$ in particular means that the sample average is not consistent in the neighbourhood of a unit root.
Denoting
we can now establish the limit properties of the $DM_A$ and $DM_P$ statistics.
We refer to Appendix (ref) for a more detailed derivation of this and other results in this section. In view of (ref) or (ref), as $M/T \rightarrow 0$ or $m \rightarrow \infty$ the DM test statistic diverges even when the null hypothesis is correct, thus giving spurious evidence of superior predictive ability. As we interpret the local to unit root as an approximation of an AR(1) in finite sample, this result suggests that the DM test diverges in presence of a root that is stationary but close to 1.
Next, we present the limit of the $DM_A$ and $DM_P$ statistics using fixed smoothing asymptotic. \\ Denoting
When $c=0$, so that the process $y_t$ is characterised by a unit root, the limit distribution $\widetilde{J}_c(r)$ is replaced by $\widetilde{W}(r)={W}(r)-r {W}(1)$ where $W(r)$ is a standard Brownian motion, and $\overline{J_c}$ is replaced by $\overline{W}=\int_0 ^1 W(r)dr$.
The limits in (ref) and (ref) exhibit the self-normalization property as in \citeasnoun{Shao2015Selfnormalization}. It is interesting to compare the limits in Theorem (ref) and in Theorem (ref) with results in the literature. It is well known that the standardised mean is diverging in presence of strongly autocorrelated series when $m \rightarrow \infty$ is assumed (and an analogue result would hold when $M/T \rightarrow 0$), but not when $m$ is assumed fixed, see for example \citeasnoun{McElroyPolitis2012ET} and \citeasnoun{hualde2017fixed}. In the context of local to unity process, this result was established, for example, in \citeasnoun{Sun2014EssayPCB}. The advantages of series variance estimators of the long run variance in presence of autocorrelation is also discussed in \citeasnoun{MULLER2007LRVestimates}. Our interest in the limits in Theorem (ref) and in Theorem (ref) is, however, slightly different, as we do not see these as alternative limits under different asymptotics, but rather as guidance that show properties of the DM test for relatively small and large $m$ (large $M/T$ and small $M/T$). The limits in Theorem (ref) and in Theorem (ref) are only relevant in the sense that we can establish whether the test statistic is divergent or not.
As $c$ in (ref) varies between $-\infty$ and $0$, it is possible to use the theory from subsection (ref) for any AR(1) model with positive autocorrelation. However, the limits in Theorem (ref) and in Theorem (ref) may not provide a valuable guideline when $\rho_T$ is not in fact in the very close neighbourhood of unity. For this situation, \citeasnoun{PhillipsMagd2007MildlyIntegrated} generalise $\rho_T$ to moderate deviations from the unit root. We simplify the model slightly and consider
Moderate deviations from the unit root following (ref) are also discussed in \citeasnoun{PM2007Chapterinbook}. \citeasnoun{GiraitisPhillips2012Local} provide a generalisation of some results under a weaker condition, similar to $(1-\rho_T)T\rightarrow \infty$. We strengthen Assumption (ref) slightly, as
Under (ref), $\overline{d}$ is a consistent estimate of $\mu$ only when $\alpha \in (0,1/2)$, but the CLT in (ref) still holds for any $\alpha \in (0,1)$, see Theorem 2.1 and the discussion on page 168 of \citeasnoun{GiraitisPhillips2012Local}. Recalling that, for any $T$, $\sigma^2=(1-\rho_T)^{-2} \omega^2$ and noticing that this is proportional to $T^{2 \alpha}$ in large samples, the rate of convergence of the CLT is reduced to $\sqrt{T^{1-2\alpha}}$ (the theory does not cover the $\alpha=1$ case, but notice that $\sqrt{T^{1-2\alpha}} \rightarrow T^{-1/2}$ as $\alpha \rightarrow 1$ and this is the rate in (ref), suggesting a proximity of the two representations; the extension of (ref) under (ref) is explored more in \citeasnoun{PM2007Chapterinbook}).
In this section, we investigate the properties of the DM statistic in the neighbourhood of unity in a Monte Carlo exercise. We consider the DGP
where
and $u_{t}$ independent from $\varepsilon_s$ for all $s,t$.
Notice that we have assumed $E(x_t)=0$, this is without loss of generality because in (ref) we could use deviations from the mean, $y_t=(\alpha+\beta E(x_{t-1}))+\beta (x_{t-1}-E(x_{t-1}))+u_{t}$. Finally, we also set $\beta=1$, again without loss of generality.
We consider two forecasting strategies:
so $\widehat{\beta}=1$ (in general it would be $\beta$ ) and $\widetilde{y}=\alpha$, and
and
so, for
then $E(d_t)=0$ when $\delta=0$. Notice that as $|\phi|<1$ then both $x_t$, $y_t$ and $d_t$ are mixing with sufficient rate, $E(d_s^2)<\infty$ (in view of the Gaussianity) and the long run variance exists. \\ Remark From Lemma 1 of \citeasnoun{dittmann2002properties}, $x_{t-1}^2$ is AR(1) with coefficient $\phi^2$; $\alpha u_t$ is an independent process and $x_{t-1} u_t$ is Martingale difference. Thus $d_t$ is like AR(1) plus noise.
As in any realistic situation the sample size $T$ is given, whether the standard normal limit or Theorem (ref) is a better approximation depends on the relative interplay between $\phi$ and $T$ (by the same token, the Lemma in \citeasnoun{dittmann2002properties} should also be seen as an approximation, when $\phi$ is close to 1, relative to $T$). For values of $\phi$ close to 1 (relative to $T$) the limit (ref) would be a better approximation for the sample average, and Theorem (ref) is a more appropriate guideline; conversely, (ref) and the standard normal limit for the DM statistic should be a more reliable guideline when $\phi$ is not close to 1, relative to $T$. Thus, the same value of $\phi$ could generate spurious rejections or not depending on the sample size. In this Monte Carlo study, we thus consider a range of values for $\phi$ and $T$ to assess the interplay between these two key elements.
We consider two sample sizes, $T=50$ and $T=100$, and a range of bandwidths spanning $m=1$ to $m=\left\lfloor T^{2/3} \right\rfloor$ when the average periodogram is used, and a range of bandwidths spanning from $M=\left\lfloor T^{1/4} \right\rfloor$ to $M=T$ when the weighted autocovariance estimator is used. For each experiment, we run $10,000$ repetitions, and we compute the empirical frequencies of rejections of the two-sided version of the test, i.e. we compare the $|DM|$ statistic against the appropriate $5 \%$ critical value from the $t_{2m}$ or the $\Phi_A(b)$ distribution, respectively, the latter as in \citeasnoun{kiefer2005new}. We always use these critical values since they yield better size properties, as discussed, for example, in \citeasnoun{lazarus2018har} or in \citeasnoun{coroneo2020comparing}.
We consider $\sigma_{\varepsilon}^2 = 1$, $\sigma_u^2 = 1$, and a range of values for $\phi$; $\alpha$ is as in (ref) for two values of $\delta$. We set $\delta=0$ to observe the effects on the empirical size; to observe the effects on power we set $\delta= -\sqrt{\sigma^2_\varepsilon/(1-\phi^2)}+\sqrt{\sigma^2_\varepsilon/(1-\phi^2)-1)}$ as this yields $E(d_t)=-1$: with this choice we can observe how the power changes as the persistence increases, for the same deviation from the null hypothesis.
We report in Tables (ref)-(ref) the simulation results using the weighted autocovariance and the weighted periodogram estimates of the long-run variance, respectively. Results confirm the findings in Section (ref). In particular:
These results support the key conclusions that we derived in Section (ref). In particular, in the size exercise, the distortion increases with $\phi$ and with the bandwidth $m$ (and decreases with $M$). In the power exercise, the power drops as $\phi$ increases from 0 to 0.85, but notice that for large values of $\phi$ this power is in fact fictitious, in the sense that it rather reflects the spurious size distortion that we observed under the null. \Copy{R2.5t}{Automatic procedures to select the bandwidth, as in \citeasnoun{delgado1996optimal}, \citeasnoun{robinson1983review} or in \citeasnoun{newey1994automatic} would not solve these problems, although the fact that smaller $m$s (larger $M$s) are automatically selected as the autocorrelation in the loss function increases, would at least avoid the most adverse effects. }
To illustrate the problems associated with the DM test when there is dependence in the loss differential, in this section we present the case in which a forecast for inflation with weakly dependent forecast errors is compared to two strongly dependent naive benchmarks. In particular, we consider quarterly predictions for the inflation rate in the Euro Area from a standard AR(1) model, as in \citeasnoun{forni2003financial} and \citeasnoun{marcellino2003macroeconomic}. As for the benchmarks, we consider a constant 2% prediction (that represents the inflation target in the Euro Area) and a rolling average (RA) prediction.
We use data on the Harmonized Index of Consumer Prices from the FRED database, and we compute quarterly year-on-year inflation rates from 1996.Q1 to 2020.Q4. \Copy{R1.4t}{Given that our objective is to compare forecasting methods as opposed to forecasting models, we consider the case of non-vanishing estimation uncertainty, as in \citeasnoun{giacomini2006tests}, and estimate all coefficients and rolling averages using a rolling window of 10 years.} We compute predictions for horizons from 1 quarter to 8 quarters-ahead, and we evaluate them on the period from 2010.Q1 to 2020.Q4 (44 observations) using a quadratic loss function.
The series of inflation and the forecasts for selected forecast horizons for the AR(1) model are shown in Figure (ref), along with the $2\%$ and the rolling average benchmarks. A visual inspection of the plots immediately suggests that the forecast from the AR(1) model is clearly superior to the $2\%$ benchmark for the 2-quarter horizon, but, not surprisingly, the superior performance of the forecast from the AR(1) model is eroded as longer forecasting horizons are considered. Additional plots for the realised forecast errors, the realised losses and the realised loss differentials are reported in Appendix (ref).
In Table (ref), we report summary statistics for the forecast errors (defined as the realised value minus the prediction) for the AR(1) and the two benchmark predictions. The forecast errors are all negative on average, implying that in this period inflation in the Euro Area has been lower than predicted by the AR(1) and the benchmarks. This result is not generated by a few large negative errors, as all the median forecast errors are also negative.
The average and median forecast errors for the AR(1) increase (in absolute value) with the forecast horizon, but they remain lower than the ones of the two benchmarks for all the forecasting horizons. The standard deviations of the AR(1) forecast errors also increase with the forecasting horizon, and they are smaller than the ones of the two benchmarks for forecasting horizons up to 4 quarters. Finally, we also present the autocorrelation structure and the ADF tests for the forecast errors (we estimated the model with the intercept, with lags selected by means of the BIC). Especially at the lowest horizons, the autocorrelations of the errors from the AR(1) forecasts declines fairly quickly, in comparison with the autocorrelations of the benchmarks. Despite this fact, the ADF test fails to reject the null hypothesis in all the cases, except for the one period horizon. We interpret this as a situation of low power of the ADF test, and therefore as evidence of persistence, but possibly not a unit root. We verify this interpretation by analysing the properties of the realised forecast losses reported in Table (ref). The average realised losses associated to the AR(1) forecast are lower, at least for forecasts up to six quarters, and less dispersed than the ones of the two benchmarks, so they are, in this sense, more precise. Moreover, the losses from the AR(1) predictions are not very correlated for short forecasting horizons. As we increase the forecasting horizon the dependence increases, but the autocorrelations still decay reasonably quickly. On the other hand, the two benchmarks display large and persistent autocorrelations in their realised forecast losses at all forecasting horizons. We further investigate the dependence in the realised losses using the ADF test: the difference in the persistence that we observed in the sample autocorrelations of the realised losses is confirmed by the outcome of the ADF test, where the unit root hypothesis is rejected only for the forecasts from the AR(1) model (and only for short horizons).
Overall, these results suggest that the AR(1) model should be more precise for short-term forecasts, but this superiority could be masked empirically by the excessive dependence in the benchmarks. On the other hand, the AR(1) does not seem to produce more precise forecasts than the benchmarks at longer horizons, and the outcomes of the unit roots tests should be interpreted as a warning that any potential statistical significant difference may be spurious.
Summary statistics of the loss differentials, computed as the loss of the benchmark minus the loss of the AR(1), reported in Table (ref), show that at short horizons loss differentials are positive, indicating that AR(1) predictions may be more accurate than the benchmarks. As the forecasting horizon increases, the average loss differential decreases. In particular, for the RA benchmark, it becomes negative, so that at 7 and 8 quarters ahead RA predictions might be more accurate than AR(1) predictions.
However, the table also shows that the loss differentials are characterised by relevant autocorrelations, even at short forecasting horizons: the properties of the loss differentials are therefore heavily affected by the benchmark considered, and even with a forecast with weakly dependent loss it is possible to have strong autocorrelation of the loss differential. In the last column of the table, we report the augmented Dickey–Fuller test statistic, which clearly indicates that even at short forecasting horizons the null of unit root of the loss differential cannot be rejected. With these levels of dependence, the DM test statistic is going to be subject to the drawbacks described in Section (ref).
We report the outcome of the DM tests for the null of equal predictive ability of AR(1) predictions with respect to a rolling average and a constant 2% benchmarks in Tables (ref) and(ref).
We consider tests in which the DM statistics are computed estimating the long-run variances as in (ref) or in (ref). For $\widehat{\sigma}_A$ we used bandwidths $M$ as $\lfloor T^{2/9} \rfloor$, $\lfloor T^{1/3} \rfloor$, $\lfloor T^{1/2} \rfloor$ and $T$, taking values $2$, $3$, $6$ and $44$ (the case $\lfloor T^{1/4} \rfloor$ is not present as this is again $2$); for $\widehat{\sigma}_P$ the bandwidth values $m$ of 1, $\left\lfloor T^{1/4} \right\rfloor $, $\left\lfloor T^{1/3} \right\rfloor $, $\left\lfloor T^{1/2} \right\rfloor $ and $\left\lfloor T^{2/3} \right\rfloor$, that for a sample of 44 observations are respectively 2, 3, 6 and 12. In all cases we use critical values from the corresponding fixed smoothing ($\Phi_A(b)$ or $t_{2m}$) distribution.
Results in Tables (ref) and (ref) are in line with our theory, as the outcome of the DM test reflects the autocorrelation documented in Table (ref). In view of the high autocorrelation of $d_t$, the test may be affected by size distortion, especially with the larger bandwidths $m$, $m=\lfloor T^{1/2} \rfloor$ and $m=\lfloor T^{2/3} \rfloor$, or for short $M$. We therefore discard results for these bandwidths. Even results with $m=\lfloor T^{1/3} \rfloor$ or $M=\lfloor T^{1/2} \rfloor$ should be considered with caution in this case, especially when the accuracy of the forecasts is compared against the $2 \%$ benchmark, as the autocorrelation seems to be particularly strong in that case. Remarkably, this is also the only case in which the equal predictive accuracy null between the AR(1) and the 2$\%$ forecast is significant at $5 \%$ level using $\widehat{\sigma}_P$ and the $m=\lfloor T^{1/3} \rfloor$ bandwidth; with shorter bandwidths, on the other hand, the null hypothesis of equal predictive accuracy is never rejected at $5\%$ level. This seems to be a disappointing outcome, given the apparent superior performance of forecasts from the AR(1) model at short horizons (as documented in Figure (ref) and in Table (ref)), and we suspect that it is due to the lack of power of the DM test in the presence of autocorrelation. The results when $\widehat{\sigma}_A$ and $M=T$ are used are perhaps slightly more convincing, at least when the one period ahead forecasts are evaluated. Overall, these results highlight how applying the DM test when the loss differential is not weakly dependent may generate unrealiable results.
In this paper, we have verified that the DM test may be seriously misleading in presence of strong autocorrelation in the loss differential. \citeasnoun{diebold2015comparing} mentions that “[o]f course forecasters may not achieve optimality, resulting in serially correlated, and indeed forecastable, forecast errors. But I(1) nonstationarity of forecast errors takes serial correlation to the extreme”. This is certainly true. However, the DM test is often used against naive benchmarks, for which an I(1) forecast error (or with root close enough to 1, given the sample size) may not be impossible. While this may be seen as an “abuse” of the DM test, it seems desirable that a test is robust to such abuse. Our results warn that this is not the case, and that the DM test may perform poorly, generating size distortion or low power, also in the presence weakly dependent processes with autocorrelation close to unity.
For comparing forecasts, the DM test is “the only game in town”, as noted in \citeasnoun{diebold2015comparing}. However, one should be aware that the game has its rules. In the empirical application, we used the DM test to compare AR(1) inflation forecasts to two naive benchmarks. Results indicate that, using a quadratic loss, the test fails exactly because the benchmark forecasts are not optimal under MSE loss, which is not a nice feature. This does not mean that one should not use the DM test. Rather, our work suggests that one should take the recommendation in \citeasnoun{diebold2015comparing} to use diagnostic procedures to assess the validity of the assumption of weak dependence of the loss differential very seriously.
\addcontentsline{toc}{section}{References}