EconBase
← Back to paper

Testing Quantile Forecast Optimality

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

77,769 characters · 14 sections · 61 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing Quantile Forecast Optimality

abstractQuantile forecasts made across multiple horizons have become an important output of many financial institutions, central banks and international organisations. This paper proposes misspecification tests for such quantile forecasts that assess optimality over a set of multiple forecast horizons and/or quantiles. The tests build on multiple Mincer-Zarnowitz quantile regressions cast in a moment equality framework. Our main test is for the null hypothesis of autocalibration, a concept which assesses optimality with respect to the information contained in the forecasts themselves. We provide an extension that allows to test for optimality with respect to larger information sets and a multivariate extension. Importantly, our tests do not just inform about general violations of optimality, but may also provide useful insights into specific forms of sub-optimality. A simulation study investigates the finite sample performance of our tests, and two empirical applications to financial returns and U.S. macroeconomic series illustrate that our tests can yield interesting insights into quantile forecast sub-optimality and its causes. \newline \noindentJEL Classification: C01, C12, C22, C52, C53 \newline \noindentKeywords: Forecast evaluation, forecast rationality, multiple horizons, quantile regression.

\onehalfspacing

Introduction

Economic and financial forecasters have become increasingly interested in making quantile predictions, often across different quantile levels and at multiple horizons into the future. In financial markets, for instance, such multi-step quantile predictions are produced due to the 10-day value-at-risk (VaR) requirements of the Basel Committee on Banking Supervision.\footnote{See for instance: https://www.bis.org/publ/bcbs148.pdf [Last Accessed: 18/12/20]} In the growth-at-risk (GaR) literature on the other hand, ABG19 propose quantile models to predict downside risks to real gross domestic product (GDP) growth at horizons ranging from one quarter ahead to one year ahead. These methods are now widely implemented in academic research PRRH20,BS20 and in international institutions like the IMF PEJLAFW2019, and are typically applied across various quantile levels. This trend for multi-horizon quantile forecasts has also developed into a growing literature in nowcasting GaR that typically uses several intra-period nowcast horizons ADP18,FMS21,CCM20. Finally, it is common for central banks, such as the Bank of England, to produce fan charts of key economic variables such as GDP growth, unemployment or the Consumer Price Index (CPI) inflation rate across several quantile levels and horizons.

However, despite the expansion in empirical and methodological research, there is currently very little statistical guidance for assessing whether a set of multi-step ahead, multi-quantile forecasts are consistent with respect to the outcomes observed. This consistency is often referred to as `optimality', `rationality', or `calibration' in the literature, with `full optimality' referring to optimality relative to the information set known to the forecaster, while a weaker form of optimality known as `autocalibration' is defined with respect to the information contained in the forecasts themselves gneiting2013,Tsyplakov2013. This paper aims to fill this gap in the literature by proposing various (out-of-sample) optimality tests for quantile forecasts that can accommodate predictions either derived from known econometric forecasting models, or from external sources like institutional or professional forecasters. Specifically, we develop tests that assess optimality of quantile forecasts over multiple forecast horizons and multiple quantiles simultaneously.

The main test of this paper is a joint test of autocalibration for quantile forecasts obtained across different horizons and quantile levels. The test is based on a series of quantile Mincer-Zarnowitz (MZ) regressions GLLS11 across all quantile levels and horizons, which are in turn used to construct a test statistic for the null hypothesis of autocalibration across horizons and quantiles using a set of moment equalities RS2010,AS2010. We suggest a block bootstrap procedure to obtain critical values for the test. The bootstrap is simple to implement and avoids the need to estimate a large variance-covariance matrix that would be required in a more standard Wald-type test. We establish the first-order asymptotic validity of these bootstrap critical values.

The test of autocalibration based on MZ regressions can provide valuable information to forecasters. In particular, failure to reject the null hypothesis of autocalibration suggests that the forecaster may proceed to use the forecasts as they are without the need to `re-calibrate' them. On the other hand, if the null hypothesis is rejected, the test hints at directions for improvement of the forecasts. That is, it informs the forecaster about the horizons, quantiles or horizon-quantile combinations that contributed strongest to the rejection of the null, and thus an improvement of the forecasts is warranted. In addition, the estimated MZ regression can be used to infer about the nature of the deviations from autocalibration, when the forecasts are plotted alongside the realisations. The estimated MZ coefficients may also be used to perform a re-calibration or the original forecasts, as has been suggested in the case of mean forecasts in the recent work of C22.

We provide two extensions of this test for autocalibration. The first extension allows for additional predictors in the MZ regressions, which we call the augmented quantile Mincer-Zarnowitz test. This test is operationally similar to that of the first test, but may provide richer information to the forecaster. It tests a stronger form of optimality relative to a larger information set than autocalibration. If autocalibration is not rejected, but the null hypothesis of the augmented test is rejected, it indicates that the additional variables used in the MZ regression carry additional informational content which should be used in making the forecasts themselves. The second extension allows to test optimality for multiple time series variables and not just for a single variable. Testing multiple time series variables simultaneously may be useful in cases where we are interested in testing whether one type of model delivers optimal forecasts for multiple macroeconomic variables, or across different financial asset returns, for instance different companies from the same sector.

Finally, as a separate contribution, we also outline in the appendix a test for monotonically non-decreasing expected quantile loss as the forecast horizon increases. This extends the result of PT2012 to the quantile case whereas they focussed on the mean squared forecast error (MSFE) case for optimal multi-horizon mean forecasts. The test makes use of empirical moment inequalities using the Generalised Moment Selection (GMS) procedure of AS2010. This test can also be seen as complementary to monotonicity tests used in the nowcasting literature for the MSFE of mean nowcasts FG17.

We provide two empirical applications of our methodology. The first application applies the basic MZ test to classical VaR forecasts for S&P 500 returns constructed from a GARCH(1,1) model via the GARCH bootstrap pascual2006. We test jointly over the quantile levels 0.01, 0.025 and 0.05 and horizons from 1 to 10 trading days. Autocalibration is rejected overall and the miscalibration of the forecasts gets stronger for larger forecast horizons and more extreme quantiles. Furthermore, a clear pattern emerges over all quantiles and horizons regarding the conditional quantile bias: the VaR forecasts tend to underestimate risk in calmer times, but overestimate it in more stressful periods.

The second application applies the test in the spirit of the emerging GaR literature, where we focus on the extensions of our test using the augmented MZ test and the test with multiple time series. We expand on the work of ABG19 to formally investigate the performance of simple quantile regression models using financial conditions indicators in predicting a range of U.S. macroeconomic series. Interestingly, we find that the forecasts across four different series and a range of quantile levels and horizon are sub-optimal in that they are not autocalibrated. However, further analysis of the results shows that this sub-optimality is present only in inflation-type series and not in real series like industrial production and employment growth. We also find poorer calibration at the most extreme quantile under consideration.

In relation to the existing literature, this paper extends the work on quantile forecast optimality or, in other words, absolute evaluation of quantile forecasts. The focus of this literature has been on single-horizon prediction at a single quantile, which mainly stems from the extensive body of research on backtesting VaR, such as christoffersen1998, engle2004caviar, EO10,EO11, GLLS11 and nolde2017. Our work also complements the literature on testing the relative forecast performance of conditional quantile models such as GK05, M15 or more recently CFG2022. Finally, as our focus lies on testing for optimality across horizons, the paper also relates to Q17, who emphasized the importance of multi-horizon forecast evaluation to avoid multiple testing issues in the context of relative evaluation of mean forecasts. The only work on multi-horizon optimality testing we are aware of is PT2012, who consider the case of mean forecasts as well and discuss several implications of optimality specific to the multi-horizon context and how to construct tests for them, most notably the monotonicity of expected loss over horizons, which we extend to the quantile case in the appendix.

The rest of the paper is organised as follows. Section (ref) lays out the notion of quantile forecast optimality that will provide the foundation of our tests. Section (ref) then introduces the test for autocalibration via MZ regression, along with the bootstrap methodology and theory. Section (ref) extends the test to the augmented MZ test and the test for multiple variables, while Section (ref) gives the two empirical applications of our methods. Finally, Section (ref) concludes the paper.

The appendix contains results from a Monte Carlo study (Section (ref)), where we assess the finite sample properties of the MZ and augmented MZ tests across various sample sizes and bootstrap block lengths. The appendix also contains the proofs for the theoretical results (Sections (ref) and (ref)) along with the monotonicity test (Section (ref)). Sections (ref) and (ref) provide additional empirical results and graphs for the VaR and the GaR application, respectively. Finally, all tests of the paper are provided as R functions in the R package quantoptimR available at https://github.com/MarcPohle/quantoptimR.

Quantile Forecast Optimality

Consider a multivariate stochastic process $\{\mathbf{V}_{t}\}_{t \in \mathbb{Z}}$, where $\mathbf{V}_{t}$ is a random vector which contains a response variable of interest $y_{t}$ and other observable predictors. We denote the forecaster's information set at time $t$ by $\mathcal{F}_{t}=\sigma(\mathbf{V}_{s}; s\leq t)$, where $\sigma(.)$ denotes the $\sigma$-algebra generated by a set of random variables. Assuming a continuous outcome $y_{t}$ for the rest of the paper, our target functional is the conditional $\tau$-quantile of $y_{t}$ given $\mathcal{F}_{t-h}$:

equation*[equation* omitted — 100 chars of source]

where $F_{y_{t}|\mathcal{F}_{t-h}} \left(\cdot \right)$ is the cumulative distribution function of $y_{t}$ conditional on $\mathcal{F}_{t-h}$. We denote an $h$-step ahead forecast at time $t-h$ for this $\tau$-quantile $q_{t} \left(\tau |\mathcal{F}_{t-h} \right)$ by $\widehat{y}_{\tau,t,h}$, and assume that we observe these forecasts $\widehat{y}_{\tau,t,h}$ for each target period $t$ at multiple horizons, $h\in\mathcal{H}=\{1,\ldots,H\}$, and multiple quantile levels, $ \tau \in \mathcal{T}=\{\tau_{1},\ldots,\tau_{K}\}\subset [0+\varepsilon,1-\varepsilon]$ with $\varepsilon>0$, for some finite integers $H$ and $K$, respectively. That is, at each time point $t$ we have a matrix of forecasts, $\left(\widehat{y}_{\tau,t,h}\right)_{\tau=\tau_1,...,\tau_K,h=1,...,H}$. In addition, throughout the paper, we will assume strict stationarity of $\{\mathbf{V}_{t}\}_{t \in \mathbb{Z}}$ and finite first moments of the forecasts $\widehat{y}_{\tau,t,h}$ and $y_t$ itself, see Assumptions A1 and A2 in Section (ref).

Since our focus lies on the evaluation of quantile forecasts, the loss function used for evaluation in this context is the `tick' or `check' loss which is well-known from quantile regression. This is written as $L_{\tau} \left(y_{t+h} - \widehat{y}_{\tau,t,h} \right) = \rho_{\tau} \left( y_{t+h} - \widehat{y}_{\tau,t,h} \right)$, where $\rho_{\tau}(u) = u \left(\tau - 1\{u<0\} \right)$ and where $1\{.\}$ denotes the indicator function giving a value of one when the expression is true and zero otherwise.

While relative forecast evaluation deals with comparing different forecasting methods or models, mainly by ranking them via their expected loss, the subject of this paper is absolute forecast evaluation across different quantile levels and/or forecasting horizons, in other words the assessment of a particular forecasting model or method in terms of absolute evaluation criteria for multiple quantile levels and horizons. These evaluation criteria are usually different forms of optimality (or `rationality'/`calibration'). We start by defining and discussing various forms of quantile forecast optimality before showing how to operationalise the latter for testing.

definition[Optimality] An $h$-step ahead forecast $\widehat{y}^{\ast}_{\tau,t,h|\mathcal{I}_{t-h}}$ for the $\tau$-quantile is optimal relative to an information set $\mathcal{I}_{t-h} \subset \mathcal{F}_{t-h}$ if: \begin{equation*} \widehat{y}^{\ast}_{\tau,t,h|\mathcal{I}_{t-h}} = \arg \min_{\widehat{y}_{\tau,t,h}} \mathrm{E} \left[L_{\tau} \left( y_{t} - \widehat{y}_{\tau,t,h} \right) | \mathcal{I}_{t-h} \right]. \end{equation*} We simply call it optimal and denote it by $\widehat{y}^{\ast}_{\tau,t,h}$ if $\mathcal{I}_{t-h}=\mathcal{F}_{t-h}$, i.e.\ if it is optimal relative to the full information set: $\widehat{y}^{\ast}_{\tau,t,h}\equiv \widehat{y}^{\ast}_{\tau,t,h|\mathcal{F}_{t-h}}$.

Analogous to the case of mean forecasts granger1969, an optimal quantile forecast relative to an information set can alternatively be characterised as being equal to the respective conditional quantile provided the information set is sufficiently large and includes the forecasts themselves. Specifically, since `tick' loss is a strictly consistent scoring function for the corresponding quantile (see Definition 1 and Proposition 1 in gneiting2011), it holds that an $h$-step ahead forecast $\widehat{y}_{\tau,t,h}$ for the $\tau$-quantile is optimal relative to any information (sub-)set $\mathcal{I}_{t-h}$ satisfying $\sigma \left(\widehat{y}_{\tau,t,h} \right) \subset \mathcal{I}_{t-h} \subset \mathcal{F}_{t-h}$, where $\sigma \left(\widehat{y}_{\tau,t,h} \right) $ denotes the sigma algebra spanned by the forecast itself, if and only if:

equation[equation omitted — 422 chars of source]

While interest often lies in testing the null hypothesis of (full) optimality relative to the information set $\mathcal{F}_{t-h}$, which amounts to testing if the forecast, $\widehat{y}_{\tau,t,h}$, is equal to its target, $q_{t}(\tau |\mathcal{F}_{t-h})$, the possibly large and generally unknown information set $\mathcal{F}_{t-h}$ usually makes direct tests of this hypothesis difficult in practice. Thus, we next discuss weaker forms of optimality that will form the basis of our test(s) in Sections (ref) and (ref) below. In fact, in Section (ref) of the appendix (see Lemma (ref)) we show formally that these weaker forms of optimality may always be viewed as a direct implication of optimality with respect to the `full' information set $\mathcal{F}_{t-h}$. That is, any $h$-step ahead forecast optimal with respect to the full information set $\mathcal{F}_{t-h}$, is also optimal relative to any `smaller' information (sub-)set $\mathcal{I}_{t} \subset \mathcal{F}_t$.

A special case of this `weaker' form of optimality is optimality with respect to the information contained in the forecast itself, $\sigma \left(\widehat{y}_{\tau,t,h} \right)$, or autocalibration, a term first coined by Tsyplakov2013 and gneiting2013 in the context of probabilistic forecasts.

definition[Autocalibration] An $h$-step ahead forecast $\widehat{y}_{\tau,t,h}$ for the $\tau$-quantile is autocalibrated if it holds that: $\widehat{y}_{\tau,t,h} = q_{t} \left(\tau |\sigma(\widehat{y}_{\tau,t,h}) \right)$.

On the one hand, autocalibration may be regarded as a direct implication of full optimality that is particularly suitable for testing as it only relies upon the forecasts themselves and does not require any assumptions on the information set $\mathcal{F}_{t-h}$ or a selection of variables from it. On the other hand, however, autocalibration may also be viewed as a criterion for absolute forecast evaluation in its own right for several reasons. Firstly, the concept has a clear interpretation since a forecast user provided with autocalibrated forecasts should use them as they are and not transform or `recalibrate' them. Secondly, only involving forecasts and observations and no information set that depends on other quantities, it comes closest to the idea of forecast calibration as a concept of consistency between forecasts and observations gneiting2007. Thirdly, autocalibration might often be a more reasonable criterion to demand from forecasts than full optimality, which is a often hard to fulfill in practice. Finally, the Murphy decomposition of expected loss P20 shows that autocalibration is a fundamental property of forecasts in that expected loss is driven by only two forces: deviations from autocalibration and the information content of the forecasts.

The next section will outline how to test autocalibration, viewed either as an implication of full optimality or as a forecast property its own right, across multiple quantile levels and horizons simultaneously.

Quantile Mincer-Zarnowitz Test

Null Hypothesis and Quantile Mincer-Zarnowitz Regressions

While autocalibration testing has a long tradition in econometrics through the use Mincer-Zarnowitz regressions for mean forecasts mincer1969, the latter may also be used directly for the case of quantiles GLLS11. Definition (ref) in fact suggests that a natural test for autocalibration of an $h$-step ahead forecast for the $\tau$-level quantile may be based on checking whether, for a given sample of outcomes and quantile forecasts at level $\tau_{k}$ and horizon $h$, it holds that: \[ q_{t} \left(\tau |\widehat{y}_{\tau,t,h} \right) = \alpha_{h}^{\dag}(\tau_{k})+ \widehat{y}_{\tau_{k},t,h}\beta_{h}^{\dag}(\tau_{k} )=\widehat{y}_{\tau_{k},t,h} \] almost surely. More generally, since our goal is to test for autocalibration over multiple forecast horizons and quantile levels jointly, we specify such a linear quantile regression model for every horizon $h\in\mathcal{H}$ and $\tau_{k}\in\mathcal{T}$ as follows:

equation[equation omitted — 260 chars of source]

where $\mathbf{X}_{\tau_{k},t,h}=(1,\widehat{y}_{\tau_{k},t,h})^{\prime}$ and $\boldsymbol{\beta}_{h}^{\dag} (\tau_{k} )=(\alpha_{h}^{\dag}(\tau_{k}),\beta_{h}^{\dag}(\tau_{k} ))^{\prime}$. Here, the population coefficient vector $\boldsymbol{\beta}^{\dag}_{h} (\tau_{k} )=(\alpha^{\dag}_{h}(\tau_{k}),\beta^{\dag}_{h}(\tau_{k} ) )^{\prime}$ of this linear quantile regression model is defined as:

equation[equation omitted — 230 chars of source]

where $\mathcal{B}$ denotes the parameter space satisfying conditions set out in Assumption A3 below. The composite null hypothesis is given by:

equation[equation omitted — 124 chars of source]

for all $h\in\mathcal{H}$ and $\tau_{k}\in\mathcal{T}$ versus $H^{\text{MZ}}_{1}: \{\alpha_{h}^{\dag}(\tau_{k})\neq 0 \} \quad \text{and/or}\quad \{ \beta_{h}^{\dag}(\tau_{k} )\neq 1\}$ for at least some $h\in\mathcal{H}$ and $\tau_{k}\in\mathcal{T}$. Testing the null hypothesis in ((ref)) not only yields a multi-horizon, multi-quantile test of autocalibration, but also provides us with an idea about possible deviations from the null. In particular, examining the contributions of single horizons, quantiles or horizon-quantile combinations to the overall test statistic, which will be introduced in Subsection (ref) below, may also be informative about deviations from autocalibration. Moreover, the empirical counterpart of $q_{t} \left(\tau |\widehat{y}_{\tau,t,h} \right) = \alpha_{h}^{\dag}(\tau_{k})+ \widehat{y}_{\tau_{k},t,h}\beta_{h}^{\dag}(\tau_{k} )$ may be interpreted as autocalibrated forecasts such that, for a specific value of the forecast $\widehat{y}_{\tau_{k},t,h}$, the deviations between the regression line (or recalibrated forecast) and the forecasts themselves, $\widehat{y}_{\tau_{k},t,h} - q_{t} \left(\tau |\widehat{y}_{\tau,t,h} \right)$ can be interpreted as the quantile version of a conditional bias. The direction and size of this conditional quantile bias informs us about deficiencies of forecasts in certain situations, a point that we will come back to and illustrate in the applications in Section (ref).

Test Statistic and Bootstrap

In what follows, assume that we observe an evaluation sample of size $P$, in other words a scalar-valued time series of observations starting at some point in time $R+1 \in \mathbb{Z}$, $\{y_{t}\}_{t=R+1}^{T}$, and a matrix-valued time series of forecasts, $\left\{ \left(\widehat{y}_{\tau,t,h}\right)_{\tau=\tau_1,...,\tau_K,h=1,...,H} \right\}_{t=R+1}^{T}$. Moreover, we may also observe a vector of additional variables $\mathbf{Z}_{t-h}$ from the forecaster's information set $\mathcal{F}_{t-h}$. We will write this additional sample of vector-valued time series as $\{\mathbf{Z}_{t}\}_{t=R+1-H}^{T-1}$.

Using the taxonomy of GR2010, forecasts $\widehat{y}_{\tau_{k},t,h}$ may stem either from `forecasting methods' or from `forecasting models'. In the former case, we are typically without knowledge about the underlying model such as with forecasts from the Survey of Professional Forecasters (SPF), or may use forecasts that depend on parameters estimated in-sample using so-called limited-memory estimators based on a finite rolling estimation window. In the case of `forecasting models' on the other hand, we need to account for the contribution of estimation uncertainty to the asymptotic distribution of the statistic. However, since the focus of this paper lies on detecting systematic forecasting bias rather than dealing with specific forms of estimation error, we consider the latter only under the expanding scheme with a `large' in-sample estimation window. Specifically, for forecasting models, we assume that we also observe $R$ additional observations of $y_{t}$ prior to $R+1$ that may be used as estimation window (note that the $R$ in-sample observations also comprise $H$ observations that are used to produce the initial out-of-sample forecast for period $R+1$), and that $P/R \rightarrow 0$ as $P,R\rightarrow \infty$. This allows us to abstract from estimation error in the analysis and to focus on systematic features of the forecasts.\footnote{In addition, it also allows us to resample forecasts directly in the bootstrap.}

The parametric models we consider in this paper take the form $m(\mathbf{W}_{t-h};\boldsymbol{\theta}^{\dag}_{\tau_{k},h})$, where $\mathbf{W}_{t-h}$ denotes a vector of predictor variable(s) and we assume for simplicity that $\mathbf{W}_{t-h}$ is a subset of $\mathbf{V}_{t-h}$. Moreover, $\boldsymbol{\theta}^{\dag}_{\tau_{k},h}$ is a population parameter vector that needs to be estimated in a first step, while the function $m(\mathbf{W}_{t-h};\cdot)$ on the other hand is assumed to be a `smooth' function of the parameter vector in the sense of Assumption A6 below. For instance, $m(\mathbf{W}_{t-h};\boldsymbol{\theta}^{\dag}_{\tau_{k},h})$ could itself take the form of a linear quantile regression model: \[ q_{t,h}\left(\tau_{k}|\mathbf{W}_{t-h}\right)=m(\mathbf{W}_{t-h};\boldsymbol{\theta}^{\dag}_{\tau_{k},h})=\mathbf{W}_{t-h}^{\prime}\boldsymbol{\theta}^{\dag}_{\tau_{k},h}. \] Alternatively, the model could also take the form of a nonlinear location scale model: \[ q_{t,h}\left(\tau_{k}|\mathbf{W}_{t-h}\right)=m(\mathbf{W}_{t-h};\boldsymbol{\theta}^{\dag}_{\tau_{k},h})=m_{\mu}\left(\mathbf{W}_{t-h};\boldsymbol{\theta}^{\dag}_{h,\mu}\right)+\sigma\left(\mathbf{W}_{t-h};\boldsymbol{\theta}^{\dag}_{h,\sigma}\right)q_{t,h,\epsilon}\left(\tau_{k}\right), \] where $\boldsymbol{\theta}^{\dag}_{\tau_{k},h}=(\boldsymbol{\theta}^{\dag\prime}_{h,\mu},\boldsymbol{\theta}^{\dag\prime}_{h,\sigma},q_{t,h,\epsilon}\left(\tau_{k}\right))^{\prime}$ and $q_{t,h,\epsilon}\left(\tau_{k}\right)$ denotes the unconditional $\tau_{k}$ quantile of the error term of the location scale model.

To accommodate both `forecasting methods' and `forecasting models', we adopt a more generic notation in what follows and let $\mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta}_{\tau_{k},h}^{\dag})$ stand either for the vector of stemming from a forecasting method or for the population vector of Mincer-Zarnowitz regressors stemming from a corresponding forecasting model. On the contrary, when forecasts have been generated from a model that has been estimated through the recursive (i.e., expanding) scheme using the first $t-h$ observations of the sample, with $t=R+1, R+2,\ldots$, we write $\mathbf{X}_{\tau_{k},t,h}(\widehat{\boldsymbol{\theta}}_{\tau_{k},t,h})$ to denote the dependence on the estimated parameter vector $\widehat{\boldsymbol{\theta}}_{\tau_{k},t,h}$. Note that even though we focus on the recursive scheme in light of the applications, the rolling estimation scheme whereby a window of the last $R$ observations is used (running from $t-h-R+1$ to $t-h$) or the fixed scheme where the parameter vector is estimated only once, i.e. $\widehat{\boldsymbol{\theta}}_{\tau_{k},t,h}=\widehat{\boldsymbol{\theta}}_{\tau_{k},R+1-h,h}$, are equally compatible with our set-up and the assumptions below could be adapted in a straightforward manner. Finally, `forecasting methods' can be accommodated by assuming $\widehat{\boldsymbol{\theta}}_{\tau_{k},t,h}=\boldsymbol{\theta}_{\tau_{k},h}^{\dag}$ almost surely.

Thus, to implement the test for the null hypothesis in ((ref)) versus the alternative hypothesis, we first estimate the coefficient vector as:

equation[equation omitted — 377 chars of source]

for each $h$ and $\tau_{k}$. With these estimates at hand, different possibilities to construct a suitable test statistic exist. More specifically, since the number of elements in $\mathcal{H}$ and $\mathcal{T}$ is finite, one option is to construct a Wald-type test based on the estimates in ((ref)) together with a suitable estimator of the variance-covariance matrix. However, when interest lies in testing ((ref)) against its complement for a larger number of quantile levels and horizons, constructing a Wald test involves estimating a large variance-covariance matrix, which can be difficult in practice and may lead to a poor finite sample performance. On the other hand, as we argue below, using a moment equality based test in combination with the nonparametric bootstrap does not suffer from this drawback. In fact, the moment equality framework can be extended easily to other set-ups which give rise to an even larger number of equalities (see Section (ref)).

To see the possibility of a moment equality based test, note that under $H^{\text{MZ}}_{0}$ and the Assumptions A1 to A7 outlined below it holds that:

align*[align* omitted — 506 chars of source]

pointwise in $h$ and $\tau_{k}$, where the matrix $\mathbf{J}_{h}(\tau_{k} )$ is given by:

equation[equation omitted — 368 chars of source]

and $f_{t,h}(\cdot)$ is defined in Assumption A4. In fact, under $H^{\text{MZ}}_{0}$, in Section (ref) of the appendix we establish the linear Bahadur representation:

align*[align* omitted — 517 chars of source]

The above representation motivates the use of a moment equality type statistic for a test of autocalibration. Thus, define the set $ \mathcal{C}^{\text{MZ}}=\left\{(h,\tau_{k},j):\ h\in \mathcal{H},\ \tau_{k}\in\mathcal{T},\ j\in\{1,2\}\right\}$, and let $\vert \mathcal{C}^{\text{MZ}}\vert=\kappa$ denote the cardinality of $\mathcal{C}^{\text{MZ}}$, while $s=1,\ldots,\kappa$ is a generic element from $\mathcal{C}^{\text{MZ}}$. For the test statistic, define $\widehat{m}_{s}$ either as $\widehat{\alpha}_{h}(\tau_{k})$ or as $(\widehat{\beta}_{h}(\tau_{k} )-1)$ for a specific $\tau_{k}$ and $h$. A test statistic for the null hypothesis in ((ref)) is then given by:

equation[equation omitted — 115 chars of source]

Note that the non-studentised statistic in ((ref)) above does not require estimation of the asymptotic variance, and consequently will be non-pivotal as its asymptotic distribution does depend on the full variance-covariance matrix. Heuristically, under $H^{\text{MZ}}_{0}$, and the conditions imposed in Theorem 1 below:

equation[equation omitted — 349 chars of source]

where $\boldsymbol{\Sigma}$ is the asymptotic variance-covariance matrix of the empirical moment equalities scaled by $\sqrt{P}$, which is unknown in practice and depends on features of the data generating process (DGP). Of course, this nuisance parameter problem can be taken into account by using a suitable bootstrap procedure HLN11. In particular, we will generate bootstrap critical values using the moving block bootstrap (MBB) of K1989, whose formal validity for quantile regression with time series observations has been established by GLN2018. In doing so, we will resample the forecasts directly as the limiting distribution will be derived under the condition that $P/R\rightarrow \pi=0$, implying that forecast estimation error does not feature into the asymptotic distribution of the test statistic W96. We do so to abstract from the dependence on a particular estimator, which would require further details about the underlying forecasting models.

We generate bootstrap samples of length $P$ consisting of $K_{b}$ blocks of length $l$ such that $P=K_{b}l$. We draw the starting index $I_{j}$ of each block $1,\ldots,K_{b}$, $\{I_{j},I_{j+1},\ldots,I_{j+l}\}$, from a discrete random uniform distribution on $[R+1,T-l]$. These indices are used to resample from $ \left\{y_{t},\widehat{y}_{\tau,t,h}\right\}_{t=R+1}^{T}$ jointly for each $\tau=\tau_1,...,\tau_K$ and $h=1,...,H$. This way we generate $B$ bootstrap samples, each with $ \left\{y_{t}^b,\widehat{y}^b_{\tau,t,h}\right\}_{t=R+1}^{T}$ for all $\tau=\tau_1,...,\tau_K$ and $h=1,...,H$. For each bootstrap sample, we construct bootstrap equivalents of ((ref)) and then the corresponding bootstrap statistic:

equation[equation omitted — 140 chars of source]

The critical value is then given by the $(1-\alpha)$ quantile of the empirical bootstrap distribution of $\widehat{U}^{b}_{\text{MZ}}$ over $B$ draws, say $c_{B,P,(1-\alpha)}$.

Assumptions and Asymptotic Validity

For the asymptotic validity of this procedure, we make the following assumptions:

A1: The outcome variable $y_{t}$ is strictly stationary and $\beta$-mixing with the mixing coefficient satisfying $\beta(k)=O\left(k^{-\frac{\epsilon}{\epsilon-1}}\right)<\infty$ for $\epsilon>1$.

A2: For all $h\in \mathcal{H}$, $\tau_{k}\in \mathcal{T}$, and $\boldsymbol{\theta}\in \boldsymbol{\Theta}$, $\mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta})$ is strictly stationary, and satisfies the mixing condition from A1 as well as $\mathrm{E}\left[\left\Vert \mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta})\right\Vert^{2\epsilon+2}\right]<\infty$, where $\Vert\cdot\Vert$ denotes the Euclidean norm and $\boldsymbol{\Theta}$ is defined in A6 below. The distribution of $\mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta})$ is absolutely continuous with Lebesgue density.

A3: For every $\tau_{k}\in\mathcal{T}$ and $h\in\mathcal{H}$, assume that the parameter space of $\boldsymbol{\beta}_{h} (\tau_{k} )$, $\mathcal{B}$, is a compact and convex set. For every $\tau_{k}\in\mathcal{T}$ and $h\in\mathcal{H}$, the coefficient vector $\boldsymbol{\beta}_{h}^{\dag} (\tau_{k} )$ from ((ref)) lies in the interior of $\mathcal{B}$.

A4: For all $h\in\mathcal{H}$ and $\tau_{k}\in \mathcal{T}$, the conditional distribution function of $y_{t}$ (given $\mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta}^{\dag}_{\tau_{k},h})$), $F(\cdot|\mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta}^{\dag}_{\tau_{k},h}))\equiv F_{t,h}(\cdot) $, admits a continuous Lebesgue density, $ f(\cdot|\mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta}))\equiv f_{t,h}(\cdot)$, which is bounded away from zero and infinity for all $u$ in $\mathcal{U}=\{u:0<F_{t,h}(u)<1\}$ almost surely. For all $h$, the density $f_{t,h}(\cdot)$ is integrable uniformly over $\mathcal{U}$.

A5: For every $h\in\mathcal{H}$ and $\tau_{k}\in\mathcal{T}$, assume that the matrix $\boldsymbol{J}_{h}(\tau_{k})$ defined in ((ref)) is positive definite and that: \[ \mathrm{E}\left[\mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta}_{\tau_{k},h}^{\dag}) \left( 1\left\{ y_{t}\leq \mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta}_{\tau_{k},h}^{\dag})^{\prime }\boldsymbol{\beta}^{\dag}_{\tau_{k},h}\left( \tau_{k} \right) \right\} -\tau_{k} \right)\right]=\mathbf{0}. \]

A6: Assume that $\boldsymbol{\Theta}$ is compact and that, for each $\tau_{k}\in\mathcal{T}$ and $h\in\mathcal{H}$, $\boldsymbol{\theta}_{\tau_{k},h}^{\dag}$ lies in its interior. For all $\boldsymbol{\theta}_{1},\boldsymbol{\theta}_{2}\in\boldsymbol{\Theta}$ and $t$, it holds that: \[ \Vert\mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta}_{1})-\mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta}_{2}) \Vert \leq B(\mathbf{X}_{\tau_{k},t,h})\Vert \boldsymbol{\theta}_{1}-\boldsymbol{\theta}_{2}\Vert \] for some non-negative, real-valued function $B(\mathbf{X}_{\tau_{k},t,h})$ satisfying $\mathrm{E}[ B(\mathbf{X}_{\tau_{k},t,h})^2 ]<\infty$. In addition, assume that for every $h\in\mathcal{H}$ and $\tau_{k}\in\mathcal{T}$, the estimator $\widehat{\boldsymbol{\theta}}_{\tau_{k},t,h}$ satisfies: \[ \sup_{t\geq R+1}\Vert \widehat{\boldsymbol{\theta}}_{\tau_{k},t,h}-\boldsymbol{\theta}_{\tau_{k},h}^{\dag} \Vert =O_{\mathrm{Pr}}\left(\frac{1}{\sqrt{R}}\right) . \]

A7: Assume that $R,P,l\rightarrow \infty$ as $T\rightarrow \infty$ with $P/R\rightarrow \pi= 0$ and $l/P\rightarrow 0$.

Assumption A1 imposes some mild restrictions on the time dependence of the data that are in turn linked to the existence and finiteness of corresponding moments in A2. On the other hand, the continuity of $\mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta})$ for any given $\tau_{k}$, $h$, and $\boldsymbol{\theta}$ only serves the purpose to simplify some of the arguments in the proofs of Section (ref) in the appendix, and could be relaxed at the expense of more cumbersome notation. Assumptions A3-A5 are required to derive the limiting distribution of the linear quantile regression estimator KX06,GMO2011. In fact, A3 and A4 represent standard assumptions on the parameter space and the smoothness of the (conditional) distribution of $y_{t}$. A5 ensures asymptotic normality of the quantile regression estimator, with the moment condition representing the quantile equivalent of the well known orthogonality condition from ordinary least squares KW2003. Assumption A6 on the other hand is only needed when the focus lies on forecasting models. Specifically, it places restrictions on the underlying parametric models, but is in fact compatible with commonly used location scale or linear quantile regression models that satisfy the Lipschitz condition in A6 and that can be estimated at rate $\sqrt{R}$. In particular, as KX2009 propose a two-step estimation procedure for linear GARCH models based on linear quantile regression, our set-up also comprises the latter type of models that are frequently used in finance applications. Finally, Assumption A7 governs the rates at which $P$ and $R$ as well as the block length $l$ may grow to infinity. In particular, and in analogy to W96, we require $\pi=0$ for estimation error to be ignorable asymptotically and to be able to resample directly from the forecasts (rather than to resample from the realised predictors). In turn, this allows us to focus on miscalibration as a structural feature of the models.

We are now ready to derive the asymptotic properties of the statistic under the null hypothesis:

theoremAssume that A1 to A7 hold, and that $\boldsymbol{\Sigma}\in\mathbb{R}^{\kappa\times\kappa}$ is positive definite. Then under $H^{\text{MZ}}_{0}$: \[ \lim_{T,B\rightarrow \infty}\Pr\left(\widehat{U}_{\text{MZ}}>c_{B,P,(1-\alpha)}\right)=\alpha. \]

Theorem (ref) establishes the asymptotic size control of the moment equality test. It is easy to implement using the moving block bootstrap and performs very well in terms of finite sample size and power, which we assess through several simulations. These can be found in Section (ref) of the appendix. There we provide two contrasting simulation set-ups in subsections (ref) and (ref) to match both the macroeconomic and financial applications below, which differ in sample sizes, target quantile levels and time series characteristics (we choose AR and GARCH processes as DGPs). In terms of block length choice, we find that values like $l=10$ work well in financial applications with thousands of daily observations and a lower length like $l=4$ is adequate in macroeconomic situations with only hundreds of monthly observations. However, the results do not change significantly when tweaking the block length. We also provide additional simulations in Subsection (ref) for augmented quantile Mincer-Zarnowitz test that we discuss next.

Extensions

In the following two subsections, we will describe two extensions of the test from Section (ref), firstly to accommodate additional predictors $\mathbf{Z}_{t-h}$ in the Mincer-Zarnowitz regression to test different forms of optimality, and secondly to test for autocalibration of quantile forecasts for multiple time series simultaneously.

Augmented Quantile Mincer-Zarnowitz Test

Recalling the characterisation of optimality relative to any information set $\mathcal{I}_{t-h}\subset \mathcal{F}_{t-h}$ from (ref), it becomes clear that the Mincer-Zarnowitz set-up from the previous section may also be used to test stronger forms of optimality with respect to larger information sets than $\sigma(\widehat{y}_{\tau_{k},t,h})$. More precisely, while $H_{0}^{MZ}$ vs. $H_{1}^{MZ}$ is a test of autocalibration, it does not check if all available valuable information from $\mathcal{F}_{t-h}$ was incorporated into the forecasting model or taken into account by the forecaster. We therefore suggest the idea of augmented quantile Mincer-Zarnowitz regressions, where a vector of additional regressors $\mathbf{Z}_{t-h} \in \mathcal{F}_{t-h}$ is added to the regression model in ((ref)) to test for optimality relative to $\sigma(\widehat{y}_{\tau_{k},t,h},\mathbf{Z}_{t-h})$, see also ET2016 for a discussion of augmented Mincer-Zarnowitz regressions in the context of mean forecasts.

That is, in analogy to the previous section, we again specify a linear quantile regression model for every horizon $h\in\mathcal{H}$ and $\tau_{k}\in\mathcal{T}$ as follows:

equation[equation omitted — 225 chars of source]

where with slight abuse of notation we use the same symbols as in the previous section for the first two regression coefficients as well as the error term, and suppress the possible dependence of the forecasts on some parameter vector $\boldsymbol{\theta}^{\dag}$. If the coefficients of $\mathbf{Z}_{t-h}$ in the population augmented Mincer-Zarnowitz regression are non-zero, i.e. $\boldsymbol{\gamma}_{h}^{\dag} (\tau_{k})\neq \boldsymbol{0}$, there is valuable information in $\mathbf{Z}_{t-h}$ that has not been incorporated into the forecasts yet. As a result, those variables or a subset thereof should be included into the model to improve forecast accuracy.

Formally, the null hypothesis we test is given by:

equation[equation omitted — 198 chars of source]

for all $h\in\mathcal{H}$ and $\tau_{k}\in\mathcal{T}$ versus $H^{\text{AMZ}}_{1}: \{\alpha_{h}^{\dag}(\tau_{k})\neq 0 \}$ and/or $\{ \beta_{h}^{\dag}(\tau_{k} )\neq 1\}$ and/or $\{\boldsymbol{\gamma}_{h}^{\dag} (\tau_{k} ) \neq \boldsymbol{0} \}$ for at least some $h\in\mathcal{H}$ and $\tau_{k}\in\mathcal{T}$.

In contrast to a standard Mincer-Zarnowitz quantile regression, the augmented version requires a choice of variables from $\mathcal{F}_{t-h}$. These variables included in $\mathbf{Z}_{t-h}$ have to be chosen a priori and may, in some situations, suggest themselves naturally as illustrated in our macroeconomic application. In other situations, however, it might be hard to pick those variables from a potentially very large information set. In those cases, regularised or factor-augmented Mincer-Zarnowitz regressions could be used instead of ((ref)). We leave this extension to future research and state the asymptotic size result of a moment equality based test of $H^{\text{AMZ}}_{0}$ vs. $H^{\text{AMZ}}_{1}$. Specifically, let $\widehat{\widetilde{m}}_{s}$ denote the quantile estimators $\widehat{\alpha}_{h}(\tau_{k})$, $(\widehat{\beta}_{h}(\tau_{k} )-1)$, or $\widehat{\boldsymbol{\gamma}}_{h}(\tau_{k})$ for a specific quantile level $\tau_{k}$ and horizon $h$. Similarly, in analogy to before, let $\widetilde{\kappa}$ denote the total (finite) number of moment equalities to be tested. The corresponding test statistic, $\widehat{U}_{\text{AMZ}}$, is given by:

equation[equation omitted — 129 chars of source]

Finally, define the `augmented' covariate and coefficient vectors: \[ \widetilde{\mathbf{X}}_{\tau_{k},t,h}(\boldsymbol{\theta}_{\tau_{k},h}^{\dag} )=(\mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta}_{\tau_{k},h}^{\dag} )^{\prime},\mathbf{Z}_{t-h}^{\prime })^{\prime} \] and $\widetilde{\boldsymbol{\beta}}_{h}^{\dag}(\tau_{k})=(\alpha^{\dag}_{h}(\tau_{k}),\beta^{\dag}_{h}(\tau_{k}),\boldsymbol{\gamma}^{\dag}_{h}(\tau_{k})^{\prime})^{\prime}$, respectively, the latter with parameter space $\widetilde{\mathcal{B}}$. We obtain the following result for the test defined by ((ref)):

corollaryAssume that A1 to A7 hold with $\mathbf{X}_{\tau_{k},t,h}(\cdot)$, $\boldsymbol{\beta}_{h}^{\dag}(\tau_{k})$, and $\mathcal{B}$ replaced by $\widetilde{\mathbf{X}}_{\tau_{k},t,h}(\cdot)$, $\widetilde{\boldsymbol{\beta}}_{h}^{\dag}(\tau_{k})$, and $\widetilde{\mathcal{B}}$, respectively. Moreover, assume that: \[ \operatorname*{plim}_{P\rightarrow \infty} \mathrm{Var}\left[\sqrt{P}\left(\begin{array}{c} \widehat{\widetilde{m}}_{1}\\\vdots \\\widehat{\widetilde{m}}_{\overline{\kappa}}\end{array}\right)\right]\equiv \widetilde{\mathbf{\Sigma}}\in\mathbb{R}^{\widetilde{\kappa}\times\widetilde{\kappa}} \] is positive definite, where $\mathrm{Var}\left[\cdot\right]$ denotes the variance operator. Then under $H^{\text{AMZ}}_{0}$: \[ \lim_{T,B\rightarrow \infty}\Pr\left(\widehat{U}_{\text{AMZ}}>c_{B,P,(1-\alpha)}\right)=\alpha. \]

Multivariate Quantile Mincer-Zarnowitz Test

While Section (ref) establishes a test for autocalibration of quantile forecasts of a single time series $y_{t}$, researchers may sometimes be interested in testing this property across several time series $\mathbf{y}_{t}=(y_{1,t},\ldots,y_{G,t})^{\prime}$ using an array of $h$ step ahead $\tau_{k}$-level forecasts $\left(\widehat{\mathbf{y}}_{\tau,t,h}\right)_{\tau=\tau_1,...,\tau_K,h=1,...,H}$, where $\widehat{\mathbf{y}}_{\tau_{k},t,h}=(\widehat{y}_{1,\tau_{k},t,h},\ldots,\widehat{y}_{G,\tau_{k},t,h})^{\prime}$, where $h\in\mathcal{H}$, $\tau_{k}\in\mathcal{T}$, and finite $G \in \mathbb{N}$. For instance, we may be interested in testing for forecast autocalibration jointly across different industries or sectors; components of GDP growth like consumption or export growth; or different macro series like in our application in Section (ref).

We now sketch the extension of the autocalibration test to such a multivariate set-up. To this end, consider again the following linear quantile regression model:

equation[equation omitted — 235 chars of source]

Here, for each group $i\in\{1,\ldots,G\}$, the coefficient vector $\boldsymbol{\beta}^{\dag}_{i,h} (\tau_{k} )=(\alpha^{\dag}_{i,h}(\tau_{k}),\beta^{\dag}_{i,h}(\tau_{k} ) )^{\prime}$ is defined as:

equation[equation omitted — 304 chars of source]

with $\mathbf{X}_{i,\tau_{k},t,h}(\boldsymbol{\theta}^{\dag}_{i,\tau_{k},h})^{\prime}=(1, \widehat{y}_{i,\tau_{k},t,h})^{\prime}$.\footnote{Note that we use $\boldsymbol{\theta}_{i}^{\dag}$ to denote the possible dependence of the forecast for series $i$ on a forecasting model.} Therefore, the set-up in ((ref)) allows for miscalibration also at the individual time series level (e.g., industry or sector) if for some $i$,$h$, and $\tau_{k}$ it holds that $\alpha_{i,h}(\tau_{k})\neq 0$ and/or $\beta_{i,h}(\tau_{k})\neq 1$, respectively.

The sample analogue of ((ref)), using the evaluation sample of observations $\{y_{t}\}_{t=R+1}^{T}$ and array-valued forecasts $\widehat{y}_{i,\tau,t,h}$ for each $i$ is given by:

equation[equation omitted — 310 chars of source]

As in Section (ref), we are interested in testing the composite null hypothesis:

equation[equation omitted — 116 chars of source]

for all $h\in\mathcal{H}$, $\tau_{k}\in\mathcal{T}$, $i=1,\ldots,G$, versus $H^{\text{MMZ}}_{1}: \{\alpha_{i,h}(\tau_{k})\neq 0 \}$ and/or $\{ \beta_{i,h}(\tau_{k} )\neq 1\}$ for at least some $h$, $\tau_{k}$, and $i$. Note that, since $G$ is finite, for a given horizon $h$ and quantile level $\tau_{k}$ (with A1-A7 adapted to hold for $y_{i,t}$ and $\mathbf{X}_{i,\tau_{k},t,h}(\boldsymbol{\theta}_{i,\tau_{k},h}^{\dag})$, $i=1,\ldots,G$), the estimator in ((ref)) is consistent for $\boldsymbol{\beta}^{\dag}_{i,h} (\tau_{k} )$ for each $i$. The test statistic for the null hypothesis in ((ref)) versus its complement is therefore given by $\widehat{U}_{\text{MMZ}}=\sum_{s=1}^{\overline{\kappa}} \left(\sqrt{P}\widehat{\overline{m}}_{s}\right)^{2}$, where either $\widehat{\overline{m}}_{s}=\widehat{\alpha}_{i,h}(\tau_{k})$ or $\widehat{\overline{m}}_{s}=\widehat{\beta}_{i,h}(\tau_{k})-1$, respectively, and in a similar way as above, $\overline{\kappa}$ denotes the total number of moment conditions which in this case accounts for the $G$ different series being used in the test.

In order to construct a suitable bootstrap statistic in analogue to Subsection (ref), we construct bootstrap analogues $\widehat{\boldsymbol{\beta}}^{b}_{i,h}(\tau_{k})$ of ((ref)) from bootstrap samples of length $P=K_{b}l$ from $K_{b}$ blocks of length $l$ by resampling again from the series of forecast-observation pairs, where the forecasts in this case are array-valued. The bootstrap procedure does not just take the horizon, but also the group structure as given, which ensures that the dependence of the original data across horizons $h$ as well as across series $i=1,\ldots,G$ is maintained. More precisely, we again draw the starting index $I_{j}$ of each block of forecasts and observations $1,\ldots,K_{b}$ from a discrete random uniform distribution on $[R+1,T-l]$. These indices are used to resample from $ \left\{y_{t},\left(\widehat{\mathbf{y}}_{\tau,t,h}\right)_{\tau=\tau_1,...,\tau_K,h=1,...,H}\right\}_{t=R+1}^{T}$. This way we generate $B$ bootstrap samples, each with $ \left\{y_{t}^b,\left(\widehat{\mathbf{y}}^b_{\tau,t,h}\right)_{\tau=\tau_1,...,\tau_K,h=1,...,H}\right\}_{t=R+1}^{T}$.

For each $i=1,\ldots,G$, we then construct a corresponding bootstrap estimator given by: \[ \widehat{\boldsymbol{\beta} }^{b}_{i,h}(\tau_{k} )=\arg \min_{\boldsymbol{b}_{i}\in \mathcal{B}}\frac{1}{P} \sum_{s=R+1}^{T}\left(\rho _{\tau }\left( y_{i,s}^{b}-\mathbf{X}_{i,\tau_{k},s,h}(\widehat{\boldsymbol{\theta}}_{i,\tau_{k},s,h}^{b})^{\prime}\boldsymbol{b}_{i}\right)\right). \] The final bootstrap statistic becomes $\widehat{U}^{b}_{\text{MMZ}}=\sum_{s=1}^{\overline{\kappa}} \left(\sqrt{T}(\widehat{\overline{m}}^{b}_{s}-\widehat{\overline{m}}_{s})\right)^{2}$, where $\widehat{\overline{m}}^{b}_{s}$ is equal to $\widehat{\alpha}^{b}_{i,h}(\tau_{k})$ or $\widehat{\beta}^{b}_{i,h}(\tau_{k})-1$, respectively. Constructing critical values on the basis of $\widehat{U}^{b}_{\text{MMZ}}$, $b=1,\ldots,B$, as in Subsection (ref), the following corollary holds:

corollaryAssume that Assumptions A1 to A7 hold with $y_{t}$, $\mathbf{X}_{\tau_{k},t,h}(\boldsymbol{\theta}_{\tau_{k},h}^{\dag})$, and $\widehat{\boldsymbol{\theta}}_{\tau_{k},t,h}$ replaced by $y_{i,t}$, $\mathbf{X}_{i,\tau_{k},t,h}(\boldsymbol{\theta}_{i,\tau_{k},h}^{\dag})$, and $\widehat{\boldsymbol{\theta}}_{i,\tau_{k},t,h}$, respectively, for every $i=1,\ldots,G$. Moreover, assume that: \[ \operatorname*{plim}_{P\rightarrow \infty} \mathrm{Var}\left[\sqrt{P}\left(\begin{array}{c} \widehat{\overline{m}}_{1}\\\vdots \\\widehat{\overline{m}}_{\overline{\kappa}}\end{array}\right)\right]\equiv \overline{\mathbf{\Sigma}}\in\mathbb{R}^{\overline{\kappa}\times\overline{\kappa}}, \] is positive definite, where $\mathrm{Var}\left[\cdot\right]$ denotes the variance operator. Then, under $H^{\text{MMZ}}_{0}$: \[ \lim_{T,B\rightarrow \infty}\Pr\left(\widehat{U}_{\text{MMZ}}>c_{B,P,(1-\alpha)}\right)=\alpha. \]

Empirical Applications

In this section, we provide two empirical applications. The first is a finance application applying the MZ test to test the optimality of GARCH predictions of the tail quantiles of financial returns. The second application uses the MZ test extensions to assess the optimality of GaR forecasts made across a range of U.S. macroeconomic variables.

Empirical Application 1: Financial Returns

Forecasts of lower tail quantiles of returns distributions (or upper tail quantiles of loss distributions) play an important role in financial risk management as the most prominent risk measures are either themselves tail quantiles or defined in terms of tail quantiles (see he2022 for a recent overview). In particular, the VaR at level $\tau$ is just the $\tau$-quantile of the returns $y_{t}$, $VaR(\tau)=q(\tau)$, where usually $\tau$ is chosen to be either 0.05 or 0.01. Expected shortfall as well as median shortfall are also defined in terms of quantiles of the return distribution and our multi-quantile evaluation framework can be useful in the evaluation of those risk measures as well (see Subsection (ref) of the appendix for a discussion). Producing and backtesting VaR forecasts is thus a central task in financial risk management. Therefore, as discussed in the introduction, the majority of contributions to quantile forecast optimality testing was motivated by this problem. However, the literature focused on single-quantile forecasts, while it may be of interest to check VaR at both the 0.01 and the 0.05 level, or even a grid of several VaR levels to approximate the whole tail distribution. In addition, note that risk management requires forecasts of risk measures over multiple horizons, e.g.\ for one day ahead and cumulative losses over the next ten trading days. Nevertheless, extant evaluation methods focus on a single horizon (with the noteworthy exception of barendse2023 who also consider multi-horizon evaluation of VaR and expected shortfall) and consequently one-day-ahead forecasts are usually evaluated. Our tests solve those problems as they enable joint evaluation of VaR forecasts over multiple levels and horizons.

To illustrate the use of the Mincer-Zarnowitz test for the evaluation of forecasts of financial risk measures, we apply it to multi-horizon, multi-quantile forecasts for daily S&P 500 returns. We consider horizons from $h=1$ through $h=10$ in terms of trading days and three quantile levels $\tau\in\{0.01,0.025,0.05\}$. The classic model for return volatility and VaR forecasting is the GARCH(1,1) model bollerslev1986. As no closed-form formula for multi-period-ahead GARCH quantile forecasts is available, except for the case of Gaussian innovations, we use the GARCH bootstrap of pascual2006. It draws standardised residuals from the estimated one-period-ahead model to simulate draws multiple periods in the future, from which quantiles can be obtained.\footnote{We use the implementation of the GARCH bootstrap from the rugarch package in R ghalanos2014.} We choose student-$t$ errors for the estimation of the model.

Our sample consists of daily S&P 500 returns from January 3rd 2000 to June 27th 2022, amounting to 5634 observations.\footnote{Data taken from the Oxford-Man Realized Library: https://realized.oxford-man.ox.ac.uk/data/download [Last accessed: 05/07/22]} We use recursive pseudo-out-of-sample forecasting with an initial estimation window of size 3000 for $h=10$ or, in other words, $R=3009$, leading to an evaluation sample of size $P=2625$. Figure (ref) in the appendix displays the one-day ahead forecasts for the three quantiles and the realisations. The forecasts for the other horizons look very similar, but are expectedly a bit wider.

We first use our Mincer-Zarnowitz test over the three quantiles, $\mathcal{T}=\{0.01,0.025,0.05\}$, and ten horizons, $\mathcal{H}=\{1,...,10\}$. We use $B=1000$ bootstrap draws and a block length of $l=10$. Table (ref) presents the results. With a $p$-value of 0.01 there is clear evidence against the null of autocalibration. Alternative block length choices of $l=5$ and $l=20$ lead to very similar $p$-values (0.011 and 0.015) which is promising in that the results are insensitive to block length.

table[table omitted — 323 chars of source]

As the test statistic from (ref) can directly be interpreted as an empirical distance from the null, consisting of scaled (by $\sqrt{P}$), squared deviations of all the Mincer-Zarnowitz regression coefficients from their values under the null, we can also look at the individual contributions to this statistic from single quantiles and single horizons or single quantile-horizon combinations, displayed in Table (ref). From this table a clear picture emerges. The outer quantiles and the longer forecast horizons contribute more to the test statistic, and thus show stronger evidence for miscalibration. Since risk management is typically concerned about the performance of a certain risk model such as the GARCH(1,1) across a range of quantile levels or horizons, this also demonstrates that a common practice to evaluate those models only for a specific choice of the latter may lead to incorrect conclusions about the overall performance of the prediction model.\footnote{In fact, Table (ref) in Section (ref) of the appendix, which contains the $p$-values for individual autocalibration tests at given values $\tau$ and $h$, illustrates that such `telescoping' practice may indeed be misleading as there is no strong evidence against autocalibration from some individual level quantile horizon combinations.}

table[table omitted — 929 chars of source]

The tests may also convey information about how models could be improved by a closer look at the Mincer-Zarnowitz regression lines themselves. For instance, using $h=1$ and $\tau=0.01$ as an illustrative example since all estimated intercepts are negative and all slopes of the regression lines less than one (see Subsection (ref) of the appendix), Figure (ref) shows the scatter plot of forecast-observation pairs alongside the estimated Mincer-Zarnowitz regression line and the diagonal. The latter represents the population regression line under $H_{0}^{MZ}$, in other words when $\alpha_{1}^{\dag}(0.01)=0$ and $\beta_{1}^{\dag}(0.01)=1$, respectively. The discrepancy between the Mincer-Zarnowitz regression line and the diagonal thus suggests that the forecasts are in fact mis-calibrated. Contrasting the two, it becomes clear that in calmer times (when the forecasts and realizations are less extreme, i.e. closer to 0) the GARCH(1,1) forecasts tend to under-predict the actual risk, in other words the quantile forecasts are not extreme enough, while in more volatile times the forecasts tend to overestimate risk.

figure[figure omitted — 238 chars of source]

Finally, in Subsection (ref) of the appendix, we examine the robustness of our results against various specification changes and examine other popular calibration tests designed for single horizon and quantile pairs. Specifically, we find that the above results do not change qualitatively when we alter the specification to a different estimation scheme (rolling) and to smaller estimation window sizes. Similarly, results remained also unchanged when omitting the COVID-19 period, or when experimenting with the GJR-GARCH model glosten1993. Finally, when looking at the autocalibration tests of engle2004caviar with the respective quantile forecast as regressor and the test of christoffersen1998, we find that the former provides $p$-values that are very close to the individual level $p$-values of the Mincer-Zarnowitz test, while this is not the case for the test of christoffersen1998 which tests different implications of full optimality.

Empirical Application 2: U.S. Macro Series

In this section we use our tests to explore the optimality of model-based forecasts of various U.S. macroeconomic series. The analysis of quantile forecasts for macroeconomic series has become widespread since studies like M15. More recently, the GaR literature has emerged to provide a tool to monitor downside risk to economic growth using quantile predictions. This approach typically analyses quarterly real GDP growth using financial conditions indicators (see ABG19), and has been subsequently applied to other quarterly macro series like employment and inflation by AABG21.

However, in spite of the increasing interest in quantile forecasting in macroeconomics, none of these papers subject their models to the type of forecast optimality test we develop in this paper. We aim to fill this gap in the empirical literature, applying our tests to shed light on the optimality of commonly-used models in predicting various macro series.

Instead of using quarterly data we propose the use of monthly variables (also used recently in similar contexts by CFG2022) and we will focus on the same four target variables analysed in M15. These series, all transformed to stationarity using the growth rate, are the Consumer Price Index for All Urban Consumers (CPIAUCSL), Industrial Production: Total Index (INDPRO), All Employees, Total Nonfarm (PAYEMS) and Personal Consumption Expenditures Excluding Food and Energy (Chain-Type Price Index) (PCEPILFE).\footnote{All series in the study are taken from the Federal Reserve Economic Data (FRED). Url: https://fred.stlouisfed.org/ [Last accessed: 08/03/22]} These series are very close in nature to the quarterly series analysed in AABG21 and will be regressed on an autoregressive term and the Chicago Fed National Financial Conditions Index (NFCI) as in ABG19.

Specifically, we use the direct forecasting scheme to generate quantile forecasts at quantile levels $\tau_k$ for $k=1,...,K$ and horizons $h=1,...,H$ as follows:

equation[equation omitted — 170 chars of source]

where $y_{t-h}$ is the autoregressive term corresponding to one of the four target variables mentioned above and $x_{t-h}$ is the NFCI. The parameter estimates are obtained by the standard quantile regression estimator and are indexed both by $\tau_k$ and $h$ to denote that a separate quantile regression is run at each quantile and horizon as in the direct scheme, as well as by $t$ as the forecasts are generate in a pseudo out-of-sample fashion as mentioned below. In essence, ((ref)) boils down to a forecast made by a quantile autoregressive distributed lag (QADL) model galvao2013 using the direct forecasting scheme.

The data series span the period 1984M1 to 2019M12, giving a total number of $T=432$ monthly observations. We use the recursive out-of-sample scheme and split the sample into equal portions for the initial estimation sample and the evaluation sample, $R=P=216$. This gives an evaluation sample size, $P$, around the middle of the range of Monte Carlo simulations in Section (ref) of the appendix. In making forecasts using ((ref)) we will use horizons $h=1,...,12$ and quantile levels $\tau_k\in\{0.1,0.25,0.5\}$. The use of these quantile levels allows us to focus on the left part of the distribution, as is common in GaR studies such as AABG21, but also includes the median as an important case of predicting the centre of the distribution. For the bootstrap implementation we use $B=1000$ bootstrap draws and employ a block length of $l=4$ as this is seen to work well in the simulation study in the appendix.

The results in Table (ref) display the results of the Mincer-Zarnowitz tests for autocalibration, with further graphical insight into the behaviour of the out-of-sample predictions given in Subsection (ref) of the appendix. We first analyse the joint Mincer-Zarnowitz test (`Joint') which works on multiple time series, as described above, where in this context we have $G=4$ target variables and we jointly test for autocalibration across all series to avoid the multiple testing problem. The results in the first row of Table (ref) show that there is indeed some evidence against autocalibration when looking across all four macro series. The $p$-value of 0.07 indicates that there is evidence at the 10% significance level that the QADL-type model does not produce well-calibrated forecasts jointly across these four series, for forecast horizons $h=1,...,12$ and quantile levels $\tau_k\in\{0.1,0.25,0.5\}$.

table[table omitted — 556 chars of source]

With this in mind, it is useful to dig further into the individual series to see which of them are likely to be causing the rejection of the joint null of autocalibration. The remainder of Table (ref) displays the Mincer-Zarnowitz test when performed individually for each series. For the two real series, industrial production and employment, we see little evidence against the null. On the other hand, for the two price-type series we see somewhat different results with a clear rejection in the case of PCEPILFE and a $p$-value just under 10% for CPIAUCSL. This suggests that the QADL-type approach suggested by ABG19 does indeed appear appropriate for real macroeconomic series but less-so for price series.\footnote{In Section (ref) of the appendix, we also try to isolate the specific horizons and/or quantile levels that contribute the most to the rejection. Our findings suggest that for PCEPILFE and CPIAUCSL the smallest contribution to the statistic comes from quantile level $\tau_k=0.5$, while the $\tau_k=0.1$ quantile level contributes more substantially, even though no systematic conclusion can be drawn beyond $h=1$. This seems to indicate that further study should consider investigating the types of series which might deliver better-calibrated predictions in the far-left tail of the distribution of inflation-type series. }

One final exercise we perform is to apply the augmented MZ test where we use additional predictors in the MZ regression. Table (ref) displays the results for each of the four series above, where in each case the remaining three variables were used as the augmenting regressors. This serves as a simple check to see if any of these other variables would have been able to improve the forecasts if they were added to the forecasting model, especially for the real variables for which the weaker null of autocalibration was not rejected. However, the results in Table (ref) are similar to the non-augmented version of the test. As expected, the stronger null is rejected as well for the inflation type series, which already showed rejections for the weaker null of autocalibration. More interestingly, for the real variables, we still get no rejections. This suggests that we are not able to improve these forecasts by the addition of inflation type variables to the forecasting model.

table[table omitted — 492 chars of source]

Conclusion

This paper deals with the absolute evaluation of quantile forecasts in situations where predictions are made over multiple horizons and possibly multiple quantile levels. We propose multi-horizon, multi-quantile tests for optimality by employing quantile Mincer-Zarnowitz regressions and a moment equality framework with a bootstrap methodology which avoids the estimation of a large covariance matrix. The main quantile Mincer-Zarnowitz test is of the null hypothesis of autocalibration, which is a fundamental property of forecast consistency. We also provide two extensions. The first extension tests a stronger null hypothesis, which allows us to add further important variables to the information set with respect to which optimality is tested. This augmented quantile Mincer-Zarnowitz test thus makes it possible to examine if the information contained in those variables was used optimally by the forecaster. The second extension is a multivariate quantile Mincer-Zarnowitz test and allows us to check autocalibration of forecasts for multiple time series at possibly multiple horizons and quantiles.

Our tests allow for an overall decision about the quality of a forecasting approach, whether it is a single model used over multiple horizons and quantiles or a mix of different models and expert judgement employed by an institution. Crucially, it avoids the multiple testing problem inherent to most practical situations, where many forecasts are made over horizons, quantiles or multiple variables. Importantly, our testing framework is constructive in that it does not only provide a formal procedure to reach this overall decision, but may also provide valuable feedback about possible weaknesses of the forecasting approach under consideration and how it could be improved.

There are many possible future avenues arising from our work, for instance the evaluation of distributional or probabilistic forecasts gneiting2014. Since these distributional forecasts are considered quantile calibrated when the corresponding quantile forecasts for all quantiles are autocalibrated, one future extension of our work may look into optimality testing across many quantiles.

Appendix