EconBase
← Back to paper

Conditional Method Confidence Set

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

80,528 characters · 15 sections · 61 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Conditional Method Confidence Set

titlepage\thispagestyle{empty} \begin{abstract} This paper proposes a Conditional Method Confidence Set (CMCS) which allows to select the best subset of forecasting methods with equal predictive ability conditional on a specific economic regime. The test resembles the Model Confidence Set by Hansen.2011 and is adapted for conditional forecast evaluation. We show the asymptotic validity of the proposed test and illustrate its properties in a simulation study. The proposed testing procedure is particularly suitable for stress-testing of financial risk models required by the regulators. We showcase the empirical relevance of the CMCS using the stress-testing scenario of Expected Shortfall. The empirical evidence suggests that the proposed CMCS procedure can be used as a robust tool for forecast evaluation of market risk models for different economic regimes. \end{abstract} Keywords: forecast evaluation, conditional loss function, Value-at-Risk, Expected Shortfall\\

Introduction

Forecasting performance of econometric methods has been a critical focus in financial and macroeconomic research, largely driven by the strict demands of institutional regulators who aim at mitigating systemic risk Ellis.2022. In the aftermath of the 2008 Financial Crisis and the recent COVID-19 pandemic, the need for robust, reliable risk assessment models has become more pressing. Regulators, such as the Basel Committee on Banking Supervision, have responded with comprehensive guidelines, including stress-testing requirements, to ensure that the econometric tools employed by banks and financial institutions are sufficiently robust during periods of market turbulence. From an econometric standpoint, these regulatory requirements highlight the necessity for advanced statistical procedures capable of selecting and evaluating models under various market regimes. A prominent example of such regulations is the requirement to “stress-test” methods employed by financial institutions to forecast Expected Shortfall, or the worst expected loss associated with the market portfolio. The periods of stress are defined by a variety of risk factors provided by the regulator, associated with different liquidity horizons, where a forecasting method is expected to be robust with respect to high values of risk factors across all liquidity horizons, which can be interpreted as economic regimes.

This paper contributes to the literature by proposing a Conditional Method Confidence Set (CMCS), a robust tool for evaluating the performance of forecasting methods under various financial stress scenarios. The CMCS builds on the Model Confidence Set (MCS) test by Hansen.2011 and extends it to allow for state-dependent model comparisons, where forecast accuracy is assessed conditional on a specific economic regime. Our empirical application focuses on stress-testing Expected Shortfall forecasts across different liquidity horizons as required by BaselCommittee.2019, offering new insights into the robustness and reliability of downside risk forecasts in financial markets.

Existing literature extensively debates the role of model misspecification, estimation errors, and the choice of loss functions in evaluating forecast accuracy. The impact of these factors on forecast comparison is critical when models are misspecified or rely on non-nested information sets. Patton.2020 demonstrates that the rankings of models, within the class of consistent loss functions, can change based on the specific choice of the loss function. In regulated environments, where financial downside risks such as Value-at-Risk and Expected Shortfall need to be forecasted, robustness of forecast evaluation is crucial. Gourieroux.2021 address this challenge by introducing robust forecast intervals that account for model misspecification, which is particularly important for stress-testing. Zhu.2022 propose a framework where conditioning variables or instruments are used to detect the best performing forecast. This approach builds on the earlier methods by introducing the role of conditioning variables to improve forecast quality, thereby shifting the focus from unconditional to conditional forecast evaluation.

The challenge of evaluating forecast accuracy becomes more complex in the context of misspecified models. Giacomini.2006 introduced the first formal approach to comparing potentially misspecified forecasts, conditional on an information set, by developing the Conditional Predictive Ability (CPA) test. Their framework allows for a pairwise comparison of forecasts based on conditional expected loss. Building on this, Li.2022 with Li.2020 extended the CPA test to allow for uniform inference on conditional loss differentials, proposing tests that can assess the hypothesis that the loss differential between models remains zero across all conditioning variables. Giacomini.2010 and Richter.2020 further contributed to the growing body of literature by proposing a dynamic framework to test for equal predictive ability, which adjusts model rankings as forecast performance shifts over time due to structural changes in the data-generating process.

A more recent contribution by Borup.2017 and Borup.2024 introduced a dynamic forecast combination (DFC) framework. This method enables the comparison of models in a multivariate setting, extending the pairwise comparison by Giacomini.2006 to identify the best set of models based on their conditional performance. Hansen.2011 provided a crucial advancement in the model selection with the concept of the Model Confidence Set (MCS), which allows for the identification of a subset of models that exhibit equal predictive ability. This set-based approach is particularly valuable in contexts where multiple models need to be considered, and their predictive accuracy needs to be assessed simultaneously. Recent work by Arnold.2024 extends this concept by incorporating sequential testing methods, allowing the MCS to be applied in a real-time framework with continuous updates as new data becomes available.

Our paper complements the unconditional approach of the MCS by Hansen.2011: even if one cannot distinguish between the predictive ability of a set of models on average, the models may display very different forecast accuracy conditional on the state of the environment. In this case, the CMCS refines the MCS as, for each state, it delivers a set of models with indistinguishable predictive ability that may differ strongly from the unconditional one.

The methods of Li.2020 and Li.2022 apply to a setting that is conceptually different from ours. They consider a conditional expected loss that is a continuous function, which rules out conditioning on indicator variables - yet, many observable states of the economy are discrete, and only a few states may be of interest. Moreover, the method of Li.2022 yields a set of models that are weakly superior over all values of the conditioning set. Consequently, as pointed out by the authors, this set is potentially empty, as uniform weak superiority may be too strong an assumption under misspecification. In contrast, our CMCS is less restrictive, as it also applies if the relative conditional predictive ability between two models varies from state to state.

While the testing procedures by \textcites{Giacomini.2006}{Borup.2024} indicate if there are differences in the conditional predictive ability, they do not directly identify the state (variable) that underlies these differences. They address this issue by proposing an selection rule based on the predicted forecasting loss, which, however, may lead to eliminations based on weak evidence. This procedure is conceptually analogous to the regression F-test, which Hansen.2011 discuss. However, our CMCS procedure ensures that eliminations are based on sufficient evidence by adopting statewise the coherency requirement that Hansen.2011 lay out.

The remainder of the paper is organized as follows. In Section (ref) we propose the Conditional Method Confidence Set and discuss its theoretical properties. Section (ref) illustrates the theoretical properties of the proposed test and compares it to the existing Wald-type tests for conditional predictive ability. Section (ref) provides empirical evidence of the performance of the proposed test in the context of stress-testing the Expected Shortfall forecasts. Section (ref) summarizes the main findings and gives an outlook on future research.

Conditional MCS

This section develops a state-wise multiple testing procedure that delivers conditional method confidence sets. We extend the Model Confidence Set (MCS) procedure by Hansen.2011 by applying it to statewise losses, i.e., subsamples of losses that are selected based on the state of the world at the forecast origin. We show the asymptotic validity of our approach.

Motivating example

To illustrate the theoretical results presented below and introduce some notation consider an example of the risk manager making a choice between two forecasting methods: a GARCH(1,1) model with Gaussian innovations and a GARCH(1,1) model with innovations following a standardized Student-t distribution \( sst(\nu) \) with \( \nu \) degrees of freedom. Both models can be used to e.g. forecast market volatility and then assume a location-scale parametrization to forecast the Expected Shortfall and the Value-at-Risk.

The risk manager estimates the parameters of each model \( i \), \( i \in \{1, 2\} \), using an estimation window of length \( r_{i} \), i.e., she uses a maximum of \( r= \max_{i} \{r_{1}, r_{2} \} \) observations. Using estimated parameters, each model makes a forecast of VaR and ES at time \( t+k \), where, for simplicity, we consider the case \( k=1 \), i.e., a one-day-ahead forecast. Following the realization of the financial return, she uses a statistical loss function \( L \) to assess the accuracy of the forecasts, which yields a scalar loss corresponding to each forecast. Thus, for each time \( t \) such that \( t+1\) is included in the out-of-sample window, she obtains the loss differential \( d_{12, t} = l_{1, t} - l_{2, t} \) to compare the forecasts of the two GARCHes.

Next, assume that the DGP of the financial return at time \( t+1 \) follows a GARCH model with innovations from the random variable \( \Psi_{t+1} \) that has a state-dependent distribution: conditional on a bivariate state variable \( S_{t} \), the innovations are distributed as

equation[equation omitted — 152 chars of source]

The state variable \( S_{t} \) takes two values, non-crisis and crisis, which we index and refer to by \( l=1, 2 \). From the DGP, it is clear that both models are misspecified: they assume the wrong distribution of the innovations with (unconditional) probability \( P(S_{t} = 2) \) and \( P(S_{t} = 1) \), respectively. In the following, we use the term method to stress that the models are misspecified and that the forecasts depend on estimated parameters.

Assume that the regulator observes the state of the market, \( S_{t} \), with some error, and reports its observations. The error is such that, when the regulator reports non-crisis, the GARCH with Gaussian innovations has a smaller expected loss than the GARCH with \( sst(\nu) \) innovations, and vice versa when the regulator reports a crisis. Consequently, exploiting the information provided by the regulator can lead to improved forecasting accuracy if one chooses the superior method.

Thus, the risk manager needs to evaluate the methods' forecasting performance in both states \( l \in \{1, 2\}\), for which she makes \(n^{1} \) and \( n^{2} \) observations each of the loss differential \( d_{12, t} \). Therefore, she tries to make inference about the expected value of \( d^{1}_{12, \tau}, \tau \in {1, 2, \ldots, n^{1}} \) and \( d^{2}_{12, \tau}, \tau \in {1, 2, \ldots, n^{2}} \), which are the loss differentials that correspond to the state of non-crisis and crisis, respectively.

The CMCS approach performs two separate tests about the mean of the conditional loss differentials \( d^{1}_{12}, \tau \) and \( d^{2}_{12, \tau} \) each, based on the test statistics

align[align omitted — 188 chars of source]

where \( \widehat{var}(\bar{d}^{l}_{12}), \; l \in \{1, 2\} \) is a consistent estimator of the variance of \( \bar{d}^{l}_{12} \). For \(n^{1} \) and \( n^{2} \) large enough, the tests reveal for each state the method that is superior.

If she decides to use the test by \textcites{Giacomini.2006} (which is nested by the multivariate extension of Borup.2024) to select one method for each state, she applies a two-step-procedure. First, she performs a Wald type test that uses the vector of test functions \( h_{t} = (1, \mathbbm{I}_{\{s_{t}=1\}})^{\prime} \in R^{2 \times 1}\) to obtain the instrumented loss differential \( z_{t} = h_{t} d_{t} = (d_{t}, d_{t}\mathbbm{I}_{\{s_{t}=1\}} )^{\prime} \). The test statistic is then

align[align omitted — 97 chars of source]

where \( \hat{\Sigma} \) is a consistent estimator of the variance-covariance matrix of \( z_{t} \).

If the test rejects that \( \mathbb{E}[z_{t}] = (0,0)^{\prime} \), she concludes that there are differences in the conditional predictive ability, but she does not have evidence in which state they exist.

Second, she thus applies the proposed decision rule. With two methods and two disjoint states, the decision rule is based on the sign of \( \bar{d}^{1}_{12} \) and \( \bar{d}^{2}_{12} \), i.e., she chooses the method with the smaller conditional average loss.

Compared to this combination of conditional test and decision rule by \textcites{Giacomini.2006}{Borup.2024}, the statewise testing of the CMCS ensures that choosing one model over the other is based on sufficiently strong evidence in finite samples, which we exemplify in Section (ref).

The unconditional MCS approach reduces to a Diebold-Mariano test (Diebold.1995) that uses all observations to perform a test about the unconditional mean of \( d_{12, t} \) based on

align[align omitted — 112 chars of source]

where \( \widehat{var}(\bar{d}_{12}) \) is a consistent estimator of the variance of \( \bar{d}_{12} \). Thus, for \( n \) large enough, the MCS reveals the method that is better on average, but will make inferior forecasts in one of the states. The CMCS, however, yields for each state the method that is superior, and therefore enables the risk manager to make more accurate forecasts.

For ease of exposition, this illustration considers only two forecasting methods, which reduces the comparison problem to a pairwise one. For \( m \geq 3 \), the CMCS approach can be implemented using \( T^{l}_{max}, \; l \in \{1, 2\} \), which takes the maximum over the individual conditional t statistics, while the MCS test can be implemented analoguously based on the unconditional \( T_{max} \) test statistic. The multivariate test by Borup.2024 is based on the Kronecker product of \( h_{t} \) and the \(m-1\) vector of loss differences between a baseline method and the remaining ones.

Description of the environment

We consider a stochastic process \( \mathbf{W} \equiv {W_{t}:\Omega \to R^{s+1}, s \in N, t=1, 2, \ldots} \) on a complete probability space \( ( \Omega, \mathcal{F}, P) \), where \( W_{t} \equiv (Y_{t}, X_{t}^{'})^{'}\) is observable. \( Y_{t}: \Omega \to R \) is the variable of interest, while \( X_{t}: \Omega \to R^{s} \) are the predictor variables, and \( \mathcal{F}_{t}=\sigma(W^{'}_{1}, \ldots, W^{'}_{t})^{'} \) (cf. \textcites{White.1994}{Giacomini.2006}{Borup.2024}).

Assume there is a finite set of \( m \) competing forecasting methods that make univariate forecasts of some functional \( F \) of the variable \( Y \), e.g., the conditional mean, median or quantile. At each time \( t \), method \( i \) makes a forecast of \( F( Y_{t+k}) \), i.e., \( k\)-steps-ahead. We denote the forecast as \( \hat{f}^{i}_{t, k, r^{i}} = f^{i}(W_{t}, W_{t-1}, \ldots, W_{t-r^{i}+1}; \hat{\beta}^{i}_{t, r^{i}} ) \) for \( i=1, \ldots, m\), where \( f^{i} \) is an \( \mathcal{F}_{t}\)-measurable function.

The subscript \(r^{i} \) on \( \hat{f} \) indicates that the forecast is generated using \(r^{i}\) observations prior to time \( t \). Moreover, \( \hat{\beta}^{i}_{t, r^{i}} \) denotes the estimates that the \( i^{th} \) forecasting method uses to generate the forecast. These estimates can be parametric, semi-parametric, or nonparametric.

Let \( r = \max \{r^{1}, \ldots, r^{m} \} \), i.e., \( r \) is the maximum length of an estimation window over the set of competing forecasting methods. Along the lines of \textcites{Giacomini.2006}{Borup.2024}, we require that \( r < \infty \) is finite. This rules out an expanding estimation window, but includes the rolling or fixed window estimation scheme with both fixed and time-varying \( r^{i} \). Consequently, we may compare nested models, as the estimation error does not vanish, which would lead to degenerate limiting distributions (Giacomini.2006).

For each method \( i \), we obtain a total of \( n \) pairs of the target variable \(Y_{t+k} \) and the \(k\)-steps-ahead forecasts made at time \( t \). We evaluate the forecasts using a real-valued, scalar loss function \(L(Y_{t+k}, \hat{f}^{i}_{t, k, r_{i}}) \). While we focus on statistical loss functions that are strictly consistent in the sense of Gneiting.2007, economic measures such as utility or a monetary criterion are also admissible. The shorthand notation \( L_{i, t} \) denotes the loss associated with the forecast that method \( i \) makes at time \( t \).

We aim to test for differences in the predictive ability of the competing forecasting methods when we observe a specific economic condition, e.g., an oil price shock or a financial crisis. We condition on \( d \) disjoint states\footnote{ While we restrict the admissible conditioning variables as compared to Giacomini.2006 and Borup.2024, this allows us to obtain interpretable MCSs. For further discussion, see Section (ref).}, i.e., we may think of the state variable \( S_{t} \) as a one-dimensional categorical random variable. We can thus represent realizations of \( S_{t} \) as a \( d-1 \) vector of indicator variables \( \tilde{s}_{t} \), which is \( \mathcal{F}_{t}\)-measurable. This yields the specific test functions \( h_{t} = (1, \tilde{s}_{t}^{\prime})^{\prime} \) as used in the tests by Giacomini.2006 and Borup.2024.

Let superscript \( l \) indicate that a random variable or observation is conditional on state \( l =1, \ldots, d \). If helpful, \( l=0 \) denotes the unconditional case. We denote the index set associated with all \( n \) out-of-sample losses \( I = \{ t_{0}, t_{0}+1, \ldots, t_{0} + n-1 \} \), where \( t_{0} \) is the earliest forecast origin. Moreover, we write \( I^{l} = \{ t: S_{t} = s^{l}, \: t \in I \} \), i.e., the observations for which we observe state \( l \) at the forecast origin, with \( n^{l} = | I^{l}| \). In the following, for each \( l \), we use \( \tau \in \{1, \ldots, n^{l} \} \) to index the observations of \(L_{i,t} | S_{t} = s^{l} \). Thus, \( \{ L^{l}_{\tau} \} \) denotes the subsequence of \( \{ L_{t}\} \) such that \(t \in I^{l}\), i.e., \( \{L_{t} | S_{t} = s^{l}\} \).

Conditional hypotheses

For a given state or economic regime $l$, e.g., a stress period as defined by regulator (BaselCommittee.2019), we want to formally test for differences in the predictive ability of \( m \) competing forecasting methods. Thus, for the state $l$ we define the conditional relative performance variables as

equation[equation omitted — 156 chars of source]

where \( \mathcal{M}^{\cdot, 0} \) is the initial set of all \( m \) methods. The competing methods are ranked in terms of their conditional expected losses.

definitionThe set of conditionally superior objects is defined by \\ \( \mathcal{M}^{l,*} \equiv \{ i \in \mathcal{M}^{\cdot, 0}: \mu^{l}_{ij} \leq 0 \text{ for all } j \in \mathcal{M}^{\cdot, 0} , \; l \in \{0, 1, \cdots, d\} \}\).

Definition (ref) is the conditional equivalent to Hansen.2011. In this paper we impose an assumption that \( \mu^{l}_{ij} \equiv \mathbb{E}[d^{l}_{ij, n^{l}\tau}] < \infty \), and that \( \operatorname{sgn} ( \mathbb{E}[d^{l}_{ij, n^{l}\tau}] ) \) does not depend on \( \tau \) for all \( i, j \in \mathcal{M}^{\cdot, 0} \), which implies that the conditional ranking of the forecasting methods is stable over time.

The hypotheses which need to be tested to find the set $\mathcal{M}^{l,*}$ take the form \[ H^{l}_{0, \mathcal{M} }: \mu^{l}_{ij} = 0 \text{ for all } i,j \in \mathcal{M}, \; l \in \{0, 1, \cdots, d\} \] and \[ H^{l}_{A, \mathcal{M} }: \mu^{l}_{ij} \neq 0 \text{ for some } i,j \in \mathcal{M}, \; l \in \{0, 1, \cdots, d\}, \] where \( \mathcal{M}^{l} \subset \mathcal{M}^{\cdot, 0}. \)

Equivalently, we can express the hypotheses in terms of \( \mu^{l}_{i \cdot} \equiv \mathbb{E}[d^{l}_{i \cdot, n^{l}\tau}] \), where \( d^{l}_{i \cdot, n^{l}\tau} = L^{l}_{i, \tau} - \dfrac{1}{m} \displaystyle \sum_{j=1}^{m} L^{l}_{j, \tau} \). We define the conditional method confidence set as any subset of \( \mathcal{M}^{\cdot, 0} \) that contains \( \mathcal{M}^{l, *} \).

CMCS Testing procedure

Operationally, the CMCS procedure is very similar to the unconditional MCS procedure by Hansen.2011. The difference of the CMCS compared to the unconditional testing is that the CMCS is performed statewise on the time series of losses, which correspond to a state $l$.

Generally, a main concern in testing multiple hypotheses is the control of the familywise error rate (FWER), which is defined as making at least one false rejection. In the context of testing predictive ability, this means to eliminate at least one method with the smallest expected loss. One principle to control the FWER while avoiding pairwise comparisons is the so-called closure method by Marcus.1976. To reject a hypothesis, every intersection hypothesis, i.e., a hypothesis that nests the individual hypothesis, must be rejected (see Lehmann.2022, ch. 9.2). This ensures that the rejection of a hypothesis is based on sufficiently strong evidence.

In the MCS testing procedure, Hansen.2011 impose a similar requirement that they call “coherency” between test and elimination rule, and which needs to be fulfilled to devise a valid multiple testing procedure that avoids pairwise comparison.

We impose analogous assumptions about the statewise testing procedures that we provide in Section (ref). We adopt the coherency requirement statewise, which implies the we control the FWER for each state. As we consider disjoint states, CMCS-based decisions are robust against false rejections.

Moreover, we impose assumptions such that the conditional loss differentials have finite moments and display finite temporal dependence, thus central limit theorems for mixing random variables apply.

assumptionFor some \( r>2\) and \( \gamma >0 \), it holds that \( \{ W_{t} \equiv (Y_{t}, X_{t}^{'})^{'}\} \) and test functions \( \{ h_{t} \} \) are \( \alpha \)-mixing of order \( -r/(r-2) \).

Define \( z_{ij, t} = h_{t} d_{ij,t} \). With Corollary (ref), \( \{ z_{ij, t}\}_{i,j \in \mathcal{M}^{0}} \) is \( \alpha \)-mixing of order \( -r/(r-2) \), and thus \( \{ d^{l}_{ij, n^{l}\tau}\}_{i,j \in \mathcal{M}^{0}} \) is \( \alpha \)-mixing of at most order \( -r/(r-2) \).

assumptionFor some \( r>2\) and \( \gamma >0 \), it holds that \( \mathbb{E} | d^{l}_{ij, n^{l}\tau}|^{r+\gamma} < \infty \) for all \( i, j, l, n^{l}, \tau \).
assumptionMoreover, assume about \( \{ d^{l}_{ij, n^{l}\tau}\}_{i,j \in \mathcal{M}^{0}} \) for all \( l, n^{l}, \tau \) that (i) \( \operatorname{sgn} \mathbb{E}(d^{l}_{ij, n^{l}\tau}) \) is constant across $n^{l}\tau$, (ii) \( var(d^{l}_{ij, n^{l}\tau}) > 0 \), and (iii) \( var((n^{l})^{-1/2} \sum^{n^{l}}_{\tau=1}d^{l}_{ij, n^{l}\tau}) > \delta > 0 \) for all \( n^{l} \) sufficiently large.
assumptionWe assume that \( P(S_{t}=s^{l}) > \varepsilon > 0 \) for each \( t, l \), with \( \varepsilon >0\) such that \( n^{l} \to \infty \text{ as } n \to \infty \).

Assumptions (ref) to Assumptions (ref) are essentially analogous to those of \textcites{Giacomini.2006}{Hansen.2011}{Borup.2024}; the data may exhibit both considerable heterogeneity and temporal dependence. Assumption (ref) (i) is an additional assumption to \textcites{Giacomini.2006}{Borup.2024}. It implies that we permit arbitrary structural breaks, except for changes in the sign of the expected loss differentials, i.e., as long as the structural breaks do not alter the conditional ranking of the forecasting methods. Especially, our method imposes restrictions on the state variables, in the sense that it requires careful conditioning by the researcher or regulator. This is less restrictive than the strict stationarity that Hansen.2011 impose on \( d_{ij, t} \) in the unconditional case.

In contrast to \textcites{Giacomini.2006}{Borup.2024}, assuming that the conditional ranking of the forecasting methods is stable rules out arbitrary shifts in the means of the conditional loss differentials also under non-stationarity. This implies that the null and the alternative are exhaustive, i.e., if \( \mathbb{E}[\{ d^{l}_{ij, n^{l}\tau}\}] = 0 \) then it also holds that \( \mathbb{E}[\{ d^{l}_{ij, n^{l *}\tau}\}] = 0 \) for any sequence \( \{n^{l*}\} \), and analogously in the case of an inequality.

Finally, Assumption (ref) ensures that the statewise sample size \(n^{l} \) goes to infinity as \( n \) goes to infinity for each state \(l\).

\paragraph The specific multiple testing procedure that we consider to test \( H^{l}_{0, \mathcal{M} } \) uses \( t^{l}_{i\cdot} = \bar{d}^{l}_{i\cdot, n^{l}} / \sqrt{\widehat{var}(\bar{d}^{l}_{i\cdot, n^{l}})} \). In brief, if \( T^{l}_{max, \mathcal{M}} = \max_{i} t^{l}_{i\cdot} \) exceeds a critical value \( c \), apply the elimination rule \(e^{l}_{max, \mathcal{M}} \) to decide which method to eliminate, where \( e^{l}_{max, \mathcal{M}}=\operatorname*{arg\,max\,}_{i \in \mathcal{M}^{l}} t^{l}_{i\cdot} \) removes the method with the largest conditional standardized excess loss. Theorem (ref) below provides the asymptotic properties of this testing procedure.

Let \( \rho^{l}_{n^{l}} \) denote the conditional \( m \times m \) correlation matrix that is implied by the covariance matrix \( \Omega^{l}_{n^{l}} \) of Lemma (ref). Further, given the vector of random variables \( \xi^{l}_{n^{l}} \sim N_{m}(0, \rho^{l}_{n^{l}})\), we let \( F^{l}_{\rho_{n^{l}}} \) denote the distribution of \( \max_{i} \xi^{l}_{i, n^{l}} \). We define \( \bar{V}^{l}_{n^{l}} = (\bar{d}^{l}_{i\cdot, n^{l}}, \ldots, \bar{d}^{l}_{m\cdot, n^{l}})^{'}. \), \( l =1, \ldots, d \).\\

theoremLet Assumptions (ref), (ref) and (ref) hold and suppose that \( \widehat{(\omega^{l}_{i, n^{l}})}^{2} \equiv \widehat{var}((n^{l})^{1/2}\bar{d}^{l}_{i \cdot, n^{l}}) = n^{l} \widehat{var}(\bar{d}^{l}_{i\cdot, n^{l}}) \overset{p}{\to}(\omega^{l}_{i, n^{l}})^{2}, \) where \( (\omega^{l}_{i, n^{l}})^{2}, i =1, \ldots, m\), are the diagonal elements of \( \Omega^{l}_{n^{l}}\). Under \( H^{l}_{0, \mathcal{M}} \), we have \( T^{l}_{max, \mathcal{M}} \overset{d}{\to} F^{l}_{\rho_{n^{l}}} \), and under the alternative hypothesis \( H^{l}_{A, \mathcal{M}} \), we have \( T^{l}_{max, \mathcal{M}} \to \infty \) in probability. Moreover, under the alternative hypothesis, we have \( T^{l}_{max, \mathcal{M}} = t^{l}_{j \cdot} \), where \( j=e^{l}_{max, \mathcal{M}} \notin \mathcal{M}^{l,*} \) for \( n^{l} \) sufficiently large.
proofLet \( \Lambda^{l}_{n^{l}} \equiv \operatorname{diag}((\omega^{l}_{1, n^{l}})^{2}, \ldots, (\omega^{l}_{m, n^{l}})^{2}) \) and \( \hat{\Lambda}^{l}_{n^{l}} \equiv \operatorname{diag}(\widehat{(\omega^{l}_{1, n^{l}})}^{2}, \ldots, \widehat{(\omega^{l}_{m, n^{l}})}^{2}) \). From Lemma (ref) it follows that \( \xi^{l}_{n^{l}} = ( \xi^{l}_{1,n^{l}}, \ldots, \xi^{l}_{m,n^{l}})^{'} \equiv (\Lambda^{l}_{n^{l}})^{-1/2} (n^{l})^{-1/2} \bar{V^{l}_{n^{l}}} \overset{d}{\to}N_{n^{l}}(0, \rho^{l}_{n^{l}})\), since \( \rho^{l}_{n^{l}} = (\Lambda^{l}_{n^{l}})^{-1/2} \Omega^{l}_{n^{l}} (\Lambda^{l}_{n^{l}})^{-1/2} \).\\ From \( t^{l}_{i\cdot} = \dfrac{\bar{d}^{l}_{i\cdot, n^{l}}}{\sqrt{\widehat{var}(\bar{d}^{l}_{i\cdot, n^{l}})}} = (n^{l})^{1/2} \bar{d}^{l}_{i\cdot, n^{l}} / \hat{\omega}^{l}_{i, n^{l}} = \xi^{l}_{i,n^{l}} \frac{\omega^{l}_{i, n^{l}}}{\hat{\omega}^{l}_{i, n^{l}}}\), it now follows that \( T^{l}_{max, \mathcal{M}} = \max_{i} t^{l}_{i\cdot} = \max^{l}_{i}((\widehat{\Lambda}^{l}_{n^{l}})^{-1/2} (n^{l})^{1/2}\bar{V^{l}_{n^{l}}})_{i} \overset{d}{\to}F^{l}_{\rho_{n^{l}}}.\) \\ Under the alternative, \( \bar{d}^{l}_{j\cdot, n^{l}} \overset{d}{\to} \mathbb{E}(\bar{d}^{l}_{j\cdot, n^{l}}) > 0 \) for any \( j \notin \mathcal{M}^{*}\), so that both \( t_{j \cdot, n} \) and \(T^{l}_{max, \mathcal{M}} \) converge to infinity at rate \((n^{l})^{1/2} \) in probability. Moreover, it follows that \( j=e_{max, \mathcal{M}} \notin \mathcal{M}^{l,*} \) for \( n \) sufficiently large.

Theorem (ref) shows that \( T^{l}_{max, \mathcal{M}} \) converges to the maximum of a normal distribution under the null, while it detects inferior methods under the alternative.

Bootstrap implementation

We implement the CMCS testing procedure by performing the bootstrapped MCS testing procedure as put forth by Hansen.2011, but use the statewise losses separately for each state $l$. Consequently, we bootstrap a functional of the mean of weakly dependent time series, i.e., the conditional losses. As discussed by Hansen.2011, the bootstrap implicitly accounts for the correlations among the (conditional) loss differentials, which influence \( F^{l}_{\rho_{n^{l}}} \). Goncalves.2002 show that the block bootstrap variance estimator for the sample mean is consistent under the type of dependence that we consider. The CMCS bootstrap procedure is outlined in Appendix (ref).

\endinput

Simulation Study

This section illustrates the finite sample properties of the proposed CMCS by means of Monte Carlo simulation. In Section (ref) we illustrate the power properties of the conditional tests compared to the unconditional ones. In Section (ref) we compare the properties of the t-test, which the CMCS is built upon, to the Wald test underlying the DFC by Borup.2024. Finally, Section (ref) sheds light on the reasons behind the differences between t-test and Wald type test in the context of conditional forecast evaluation.

CMCS power properties

Hansen.2011 define the power of the unconditional MCS procedure as the average number of elements in the unconditional MCS \( \mathcal{ \widehat{M}}_{1-\alpha}^{*} \).\footnote{The power of a multiple testing procedure can be defined in several ways, see, e.g., Romano.2005, for a discussion thereof.} Analoguously, we define the power of the CMCS procedure for each state as the average number of elements in the statewise CMCS \( \mathcal{ \widehat{M}}_{1-\alpha}^{l, *} \).

We consider a setting in which all competing methods have the same unconditional predictive ability, while their conditional predictive ability varies according to the current state. The setup of these simulations is similar to the one in \textcites{Giacomini.2006}{Borup.2017}. Accordingly, we define a state variable \( S_{t} \), where \(P(S_{t} = 1)=p \), and \(P(S_{t} = 2)=1-p \). Similarly to the simulation studies in the literature, we set \( p=0.5 \). For \( n \in \{150, 500 , 1000\} \) observations, we generate \(5000 \) sequences of losses according to \( \mathbf{L}_{t+1} = \boldsymbol{\mu}_{t+1} + \boldsymbol{\epsilon}_{t+1} \), where for each method \( i \in \{1, \ldots, m \} \)

equation[equation omitted — 182 chars of source]

with \( c_{i} = \frac{2(i-1)}{m-1}\), and error terms \(\boldsymbol{\epsilon}_{t+1} \), that follow a multivariate normal \( \mathcal{N} (\mathbf{0}, \Phi) \) with \( \Phi = \mathbb{I}_{m_{0}} \). The choice of the \( c_{i} \) is such that the vector of expected conditional losses is equally spaced, and consequently \( \boldsymbol{\mu}_{t} \) is such that the largest difference in the conditional predictive ability is \( 2 \mu \), and that the conditional predictive ability is symmetric in the two states: in state 1 (\(S_{t} = 1\) ), model 1 is conditionally the best and model \(m \) the worst. In state 2, (\(S_{t} = 2\) ), the ranking is reversed. Moreover, it holds that the absolute value of the conditional loss differential between two models \( i, j \) is the same in both states for all \( i, j \in \mathcal{M}^{\cdot, 0}\).

We consider \( \mu \in [0.1, 0.5 ] \), which covers the same range of the expected loss differential between the two models as in Giacomini.2006. We perform both the unconditional MCS testing procedure, and the statewise CMCS for \(m=10\) models, and set the level of the test \( \alpha=0.05 \).

Figure (ref) presents the power - the average number of models in the final set of models - of both the MCS and the CMCS for \( n \in \{150, 500, 1000\} \) observations in total, i.e., an expected number of conditional observations of \( n/2 \). The DGP implies that, in expectation, the power is the same in state 1 and state 2, though the CMCS will contain different models according to the state of the economy: it holds that \(\mathbb{E}[d^{1}_{ij}] = -\mathbb{E}[d^{2}_{ij}] \) for all \( i,j, \in \{1, \ldots, m \}\). Varying the conditional expected loss differentials and the state probability \( p \) would imply that the power of the CMCS is different in the two states.

figure[figure omitted — 965 chars of source]

Both testing procedures display the desired behaviour. On the one hand, the unconditional MCS rarely eliminates any methods, as the null of unconditional equal predictive ability holds. On the other hand, the CMCS clearly has power, as the number of methods in the confidence set decreases with $\mu$. The proposed CMCS considerably narrows down the set of models for which we cannot reject equal conditional predictive ability. Comparing the different panels on Figure (ref) indicates that the power of the CMCS test increases with the sample size $n$, as the number of methods in the confidence set, for a given distance form the null $\mu$, is lower, the larger the $n$.

Two models, two states: Wald vs. t-test

The simulation study below illustrates the difference between the proposed CMCS approach and the two-step-approach of \textcites{Giacomini.2006}{Borup.2024}, i.e., the combination of a Wald type test and a decision rule. For clarity and ease of interpretation we consider a simple case of two forecasting methods and two disjoint states, where the forecast evaluation is based on the sign of the loss differential for a given state: if the Wald test provides evidence to reject the null of equal predictive ability, the procedure defines the MCS as the method that has the smaller average loss conditional on the current state.

As above, we define a state variable \( S_{t} \), where \(P(S_{t} = 1)=p \), and \(P(S_{t} = 2)=1-p \), with \( p\in \{0.2, 0.5\} \). For \( n=500 \) observations, we generate \(10,000 \) time series of losses according to \( \mathbf{L}_{t+1} = \boldsymbol{\mu}_{t+1} + \boldsymbol{\epsilon}_{t+1} \), where

equation[equation omitted — 216 chars of source]

where \( v \in [0, 1] \) and the error terms \(\boldsymbol{\epsilon}_{t+1} \) follow a multivariate normal \( \mathcal{N} (\mathbf{0}, \Phi) \) with \( \Phi = \mathbb{I}_{m_{0}} \).

We consider \( \mu \in [0.05, 0.3 ] \) and first focus on the case \(p=0.5\), when both states occur with equal probability. In state 1, method 1 is superior, and the difference in the conditional predictive ability is \( \Delta_{1} =- 2 \mu \). In state 2, if \( v =1\), the difference in the conditional predictive ability is \( \Delta_{2} = 2 \mu \), i.e. method 2 is superior, whereas the unconditional predictive ability is the same. For \( v \in (0, 1) \), method 2 is superior in state 2, while method 1 is unconditionally superior. For \(v=0 \), both methods have the same conditional equal predictive ability in state 2, and method 1 is uniformly weakly superior.

We use \( c^{\star} \) as shorthand notation for \( c( \mathcal{D}^{\star}, 1-\alpha) \), the critical value of a distribution \( \mathcal{D}^{\star} \) when we perform a test at the \( 1-\alpha \) confidence level. \( P(|T^{\star}| > c^{\star}) \) denotes the simulated probability with which the absolute value of test statistic \(T^{\star} \) exceeds the critical value \(c^{\star}\) of its distribution under the null hypothesis, which is the rejection rate of the test. \( T^{1}\) and \( T^{2}\) denote the test statistics for the statewise t-tests, while \( T^{h} \) is the Wald-type test statistic used by Giacomini.2006. The test statistics are defined in Section (ref), Equations (ref) and (ref). We report results for the level of the test \( \alpha=0.05\).

Table (ref) presents the results for \(v=0 \), where in state 1, we would expect the method confidence set to contain the model 1. The further is the distance from the null, captured by increasing $\Delta_{1}$ in table rows, the greater is the power of the t-test for state 1. The t-test for state 2, as expected, produces nominal rejection rates of about 5%, as both methods have the same CPA for this state. The null of equal predictive ability holds in state 2, and conditional on a rejection of the Wald test, we always eliminate one method, which implies a type 1 error of 100% for this testing procedure.

Table (ref) presents the results for \( v \in (0, 1]\). Firstly, for small values of \( v \), \( v < 0.5 \), we observe that using the Wald type test has less power than the t test in state 1. Compared to the statewise t-test, the critical value increases with the degrees of freedom of the \( \chi^{2}_{df} \)-distribution, and state 2 fails to provide sufficient evidence to exceed this value.

Secondly, small values of \( v \) are associated with frequently estimating the wrong sign of the expected loss differential in state 2. If \( \mu \) is large enough such that the differences in predictive ability in state 1 drive the rejections of the Wald test, this then leads to large “type III” or directional errors in state 2: the decision rule eliminates the superior or equally accurate method 2 far more often than the nominal level of the test.

Thirdly, if \( v \) is large, the distance from the conditional null in state 2 is larger and the sign of the predicted loss differential is usually estimated correctly. In this case, the Wald type testing procedure based on insufficient evidence has more power than the statewise t tests, without committing large type III errors.

These simulation results illustrate the need for additional testing as formalized in the closed testing procedure (Marcus.1976), and also highlight the loss of power of the Wald type test if the difference in predictive ability in one state is very small.

table[table omitted — 1,497 chars of source]
table[table omitted — 3,607 chars of source]

For $p = 0.2$, Tables (ref) and (ref) present the results for \(v=0 \) and \( v \in (0, 1]\), respectively. In addition to the above mentioned mechanisms, these tables show the loss of power of the Wald test if there are large differences in predictive ability in the state that occurs with smaller probability (here \( p=0.2\) ), while the more frequent state displays only small differences in predictive ability.

table[table omitted — 1,537 chars of source]
table[table omitted — 3,581 chars of source]

\FloatBarrier

To provide a theoretical explanation for the smaller power of the Wald test for small values of \( v \), i.e, \( \Delta_{2} \), we provide the following lemma.

lemmaAssume that the DGP follows Equation (ref), where \( \Delta_{1} <0 \) and \( \Delta_{2} \), denote the conditional expected losses in state 1 and state 2, respectively, while \( \sigma^{2} \) denotes the common variance of the conditional losses. Moreover, assume that the covariance matrix \( \Sigma \) of \( z_{t} = h_{t} d_{t} = (d_{t}, d_{t}\mathbbm{I}_{\{s_{t}=1\}} )^{\prime} \) is known, and that \(n_{1} = p n \), \( n_{2} = (1-p) n \) are fixed. Then, it holds that \begin{align*} T^{h} & \equiv n \left[ D + E + F \right] \\ &=n \bigl [ p^{2} (\overline{d^{1}})^{2} \dfrac{ \sigma^{2} + p \Delta_{2}^{2} }{ p \sigma^2 \left( (1 - p) \Delta_1^2 + p \Delta_2^2 + \sigma^2 \right)} +2 p(1-p) \bar{d^{1}} \bar{d^{2}} \dfrac{\Delta_{1} \Delta_{2}}{\sigma^2 \left( (1 - p) \Delta_1^2 + p \Delta_2^2 + \sigma^2 \right)} \\ &+ (1-p)^{2} (\overline{d^{2}})^{2} \dfrac{ \sigma^{2} + (1-p) \Delta_{1}^{2} }{(1 - p) \sigma^2 \left( (1 - p) \Delta_1^2 + p \Delta_2^2 + \sigma^2 \right)} \bigr ] \\ &< n_{1} \dfrac{(\overline{d^{1}})^{2}}{\sigma^{2}} + n [E +F], \end{align*}
proofSee the Algebra in Appendix (ref).
remarkWe see that \( n_{1} \dfrac{(\overline{d^{1}})^{2}}{\sigma^{2}} \) is proportional to \( T^{1} \) (see Equation (ref)). In the special case \( \Delta_{2} = 0 \), it holds that \( E=F=0 \), i.e., \( T^{h} = n D < n_{1} \dfrac{(\overline{d^{1}})^{2}}{\sigma^{2}} \). If the critical value \( c(\chi^{2}_{1}, 1-\alpha) < n_{1} \dfrac{(\overline{d^{1}})^{2}}{\sigma^{2}} \), but \( n D < c(\chi^{2}_{2}, 1-\alpha) \), a rejection depends on the evidence that comes from state 2 in terms \( E \) and \( F \). Additionally, \( E \) and \( F \) need to compensate for the inflated denominator in term \( A \). For \( \Delta_{2} \) small enough, the Wald type test thus has less power than the statewise test in state 1.

\FloatBarrier

Rejection regions

To further examine the behaviour of the Wald test and statewise t-tests, we plot the rejection regions as a function of the average conditional out-of-sample loss for the same DGP as in Section (ref). We fix \( n=500 \), and vary \( p \), i.e., the probability of being in state 1, and the expected conditional loss differentials \( \mathbf{\Delta}=(\Delta_{1}, \Delta_{2})^{\prime} \), to convey how state probabilities and distance from the null impact the power of the tests.

Moreover, we use the true value of the covariance matrix of the loss differentials obtained from the simulations to construct the test statistics, and set the number of observations per state to their expected values \( np \) and \( n(1-p)\), respectively. We derive the expression for the Wald type test statistic in Appendix (ref).

On each panel of Figure (ref), the black square depicts the expected value of the conditional loss differentials \( \Delta \). To illustrate the distance from the null of conditional equal predictive ability, the outer black bars around the black squares correspond to the 2.5% and 97.5% quantiles of the sample mean of the loss differentials, and the inner bars to the 25% and 75% quantiles respectively. The areas where the null of equal conditional predictive ability is not rejected of the statewise t-tests are shaded in grey, while the non-rejection area of the Wald test lies within the black ellipsis.

The two upper panels on Figure (ref) show the case when there are large conditional differences in the predictive ability, while the two bottom panels illustrate a setting with a large expected loss difference in state 1, and a much smaller one in state 2, i.e., when the predictive ability difference is close to the null in state 2. In the first case, as illustrated by the upper panels, the Wald test has sufficient power to reject the null of CPA, and the decision rule yields reliable results as the sign of the expected conditional loss differential is usually estimated correctly.

The lower left panel demonstrates one pitfall of combining the Wald type test with the elimination rule as in Giacomini.2006: while the rejection is driven by the large conditional loss differential in state 1, the superior model is eliminated very frequently in state 2 if the sign of the loss differential was estimated incorrectly.

The lower right panel illustrates the loss of power if there are large differences in predictive ability in the state that occurs with smaller probability, \( p=0.3\) in this case, while the more frequent state displays only small differences in the predictive ability. This setting relates to the states of crisis / non-crisis in financial markets: the forecasting methods' performances differ strongly in times of crisis, for which relatively few observations are available, while their performance is very similar during calm periods. Relative to the Wald type test, the loss of power of the statewise t-test for state 1 is much smaller.

figure[figure omitted — 712 chars of source]

Empirical Evidence

This section demonstrates how CMCS can be used in the context of forecasting and stress testing Expected Shortfall (ES).

Downside measures of market risk

Value-at-Risk (VaR) has been used by the Basel Committee on Banking Supervision (BCBS) to assess market risk since 1996 and it measures the worst possible loss of a portfolio, which can happen with small probability. Empirically VaR is estimated as a 1% quantile of the profit and loss distribution of a portfolio. Let $r_t = \ln(P_t) - \ln(P_{t-1})$ be the daily log return process, where $P_t$ is the closing price on day $t$, $t= 1, \ldots T$. We assume that daily returns $r_t$ follow a specific conditional distribution with a cumulative distribution function $D_t$: $r_t|\mathcal{F}_{t-1} \sim d_t$, where $d_t$ is the probability density function corresponding to $D_t$ and $\mathcal{F}_{t-1}$ denotes past filtration. A one-day ahead forecast of the return quantile at a level $p$, known as the VaR, is defined as $\text{VaR}_{t+1}(p) = D_{t+1}^{-1}(p)$. ES has been introduced in 2016 as a more “prudent” risk measure, which estimates the expected value of the loss, conditional on VaR threshold being crossed, usually at the $p = 2.5\%$ level. The ES is then defined as $\text{ES}_{t+1}(p) = \operatorname*{E} \left[ \left. r_{t+1} \right| r_{t+1}<\text{VaR}_{t+1}(p),\mathcal{F}_{t} \right]$. BaselCommittee.2019 requires the banks to report one day ahead VaR forecasts at $p =1\%$ and 10 days ahead ES forecasts at $p = 2.5\%$. The forecasts of VaR are checked to fit the definition of a quantile with various tests, e.g. out of 100 reported forecasts the nominal level of 1% implies 1 day where the loss is lower than the provided VaR. The quality of VaR forecasts determines the scaling, or the penalty factor for the banks' capital requirements, which in turn are measured based on ES. Additionally ES forecasts should be stress-tested based on the risk factors of different liquidity horizons. BCBS specifies several liquidity horizons and associated risk factors for the ES stress-testing, which are summarized in Table (ref).

table[table omitted — 715 chars of source]

The period of “stress” is defined as the most severe time period of 252 days, where the risk factor was at its highest, and the considered sample of the risk factor should include the time period starting from January 2007. The coefficient to be reported, $ES_{BCBS}$ is defined in (ref):

equation[equation omitted — 149 chars of source]

where $ES(j = k), \; k = 2,...,5$ denotes the ES forecast for the “stress” period of the risk factor at liquidity horizon $LH_k$, $T$ denotes the forecasting horizon, e.g. 1 or 10 days. The scaling factors $LH_j$ are defied in Table (ref). Table (ref) in the Appendix reports risk factors implemented in the empirical study below.

figure[figure omitted — 344 chars of source]

Figure (ref) provides a graphical illustration of the stress period associated with different liquidity horizons. Notably, the more “stressful” periods for risk factors such as volatility or small cap prices are quite different from the high risk periods associated with interest or foreign exchange rates. Given the heterogeneity of the risk factors we expect the conditional forecast combination decomposition vary significantly accords $LH$.

In this paper we consider daily return data of 30 large cap stock returns. The details about the stocks and some selective descriptive statistics are presented in Appendix (ref). The testing period, which is used to construct the forecast combinations and to stress test the methods, is from January 2000 until December 2014 and the performance of forecast combinations is evaluated during the period from January 2015 until December 2019 ($H = 1250$). We consider 20 models for the forecast combination, which are estimated on a 2 or 4 years of data prior to 2014 in a rolling window fashion, i.e. model parameters are updated with every shift of the estimation and testing windows.

table[table omitted — 819 chars of source]

The model choice is motivated by the empirical properties of the daily return data, which exhibit strong overkurtosis and non-zero skewness. We also include historical simulation (HS) and filtered historical simulation (FHS) models, which are widespread in the professional world BaroneAdesi.1999. The details on model implementation are presented in Table (ref).

Forecasting and stress-testing ES

To construct a forecast combination for different stress periods we compare the proposed CMCS test with the dynamic forecast combination (DFC) by Borup.2024. Both testing procedures are implemented based on the consistent loss function for the ES, which is elicitable jointly with VaR Fissler.2016:

align[align omitted — 253 chars of source]

where $p$ denotes the probability level and equals to 2.5% as specified in BaselCommittee.2019, $H_t = \mbox{$\mathrm{1l}$\,}(r_t\leq VaR_t) $ denotes the hit at time $t$, $G_1(x) = x$, $G_2(x) = \exp(x)/(1+exp(x))$, $\xi_2(x) = \ln(1+ \exp(x))$, $a = \ln(2)$ Taylor.2020. For both CMCS and Wald test we use the significance level $\alpha = 0.05$. The CMCS is implemented with a bootstrap of $B = 100$ iterations. The forecast combination for both testing procedures is constructed as an average of the ES forecasts which remain in the Model Confidence Set.

Appendix Tables (ref) and (ref) report average out-of-sample conditional ES forecasts for all liquidity horizons and across all stocks for the forecast combination based on the proposed CMCS for 1 and 10 days ahead respectively. The last column of the table reports the coefficient $ES_{BCBS}$ defined in (ref) and indicates the overall riskiness of the investment across the liquidity horizons. The variability in ES forecasts across liquidity horizons highlights that assets do not exhibit uniform risk exposure. Some assets might exhibit increased risk exposure to shorter-term factors (e.g. interest rates, equity prices), while others might be more susceptible to longer-term factors (e.g. commodity prices). For instance, the ES forecast for BAC (Bank of America) is more sensitive to stress periods in foreign exchange rates and sovereign bond interest rates, rather than commodity prices. In contrast, a stock like GE exhibits relatively small changes across liquidity horizons, suggesting it is less sensitive to the associated risk factors.

figure[figure omitted — 338 chars of source]
figure[figure omitted — 328 chars of source]

The time series of the CMCS ES forecast combination for the BAC and GE stocks are presented in Figures (ref) and (ref). Lines of different colours correspond to liquidity horizons.

The plots illustrate that the CMCS dynamically adopts the composition of the forecast: a large expected loss for the BAC stock at the beginning of the year is associated with the $LH = 20$ depicted as a blue line, whereas at the end of the year all of the risk factors, including the unconditional forecast, are of a similar level. The forecast of expected loss for the GE on the contrary does not highlight a specific risk factor to which the stock could be particularly sensitive to.

figure[figure omitted — 631 chars of source]
figure[figure omitted — 639 chars of source]

Figures (ref) and (ref) provide insights into the forecast composition of the CMCS. The heatmaps depict the out-of-sample average frequency of a given model (in y-axis) to be selected for a forecast combination for a given stock (in x-axis) for 1 and 10 days ahead forecasts respectively. The heatmaps illustrate that the model selection considerably differs across (i) forecasting horizons (Figures (ref) and (ref)); (ii) risk factors (panels on each figure); and (iii) across stocks (patterns of each heatmap). The warmer the colour, the more frequently the method was included in the CMSC. Notably, the conservative Extreme Value Theory (EVT) models are excluded from the CMCS for all stocks and liquidity horizons and ES models which are based on historical simulation are in general selected fewer times. Furthermore, the forecasting horizon plays a role, e.g. an EGARCH with Student-t innovations estimated on a large window of 1000 observations is rarely picked up for 10 day forecasting horizon, compared to the 1 day forecasting horizon. The selection of the same model, but estimated on a different estimation window, differs as well. Figure (ref) demonstrates that for the 10 days forecasting horizon the EGARCH-t model estimated on a shorter window of 500 observations is selected in the CMCS rather frequently, compared to the same model estimated on 1000 observations. Furthermore, the conditional model selection drastically differs from the unconditional one, the latter being more sparse. The conditional method selection varies largely across stocks, as well as liquidity horizons: comparison of the second and the last heatmaps on Figure (ref) demonstrate that for different risk factors a different set of models is selected for the same stock. These results indicate that the model performance in downside risk measurement is very heterogeneous and call for data-driven methods in the risk management.

Next, we consider an alternative MCS-based forecast combination - the DFC by Borup.2024. The test statistics of DFC crucially depends on the covariance estimator of the loss differentials, see Borup.2017 for the details. Given the daily frequency of the ES forecasts, we implement the HAC estimator with a truncated kernel with a truncation lag of a quarter of the sample and report the results based on sample covariance estimator and larger truncation lag in Appendix (ref). The results for the sample covariance estimator, as expected, differ substantially from the HAC estimator, however the performance of the testing procedure for a different truncation lag is comparable to the results reported below.

figure[figure omitted — 667 chars of source]

Appendix Tables (ref) - (ref) report the average ES forecasts of the DFC method for each stock. The reported forecasts for the expected loss are more negative compared to the results for the proposed CMCS in Tables (ref) and (ref). This difference can be explained by the model selection patterns in Figure (ref): the more conservative ES forecasts of EVT and RiskMetrics are often included in the DFC forecast. Overall, the testing procedure seem to lack in power to discriminate the performance of candidate models during the stress periods of different risk factors. The regulatory requirements of the ES stress testing imply that the considered sample in which the forecasts are stress-tested is considerably large, however the “conditioning” period of stress is relatively short, making it difficult for the test of DFC, which is based on the whole sample, to differentiate between the models.

figure[figure omitted — 367 chars of source]

As a consequence, the resulting forecast combination includes under-performing models, which results in over-estimation of downside risk. Providing overly conservative foreacsts implies that the financial institutions would have to increase their reserves drastically. Furthermore, as depicted in Figure (ref), the forecast combination is unstable over time, as the test does not exclude volatile forecasts. The performace of the Wald test-bsed DFC for different forecasting horizons and truncation lags are presented in Appedix (ref). These results imply that (i) the testing procedure is rather sensitive to the truncation lag; (ii) for the HAC estimator the test lacks power to differentiate between the methods, risk factors and forecasting horizon. For the down-side risk stress testing and forecasting we therefore suggest to rely on the proposed state-wise CMCS, which offers a robust tool for data driven conditional method selection.

Conclusions

This paper proposes a new methodology for evaluating the forecasting performance of econometric models under varying financial conditions, introducing the Conditional Method Confidence Set (CMCS). The CMCS extends the traditional Model Confidence Set (MCS) of Hansen.2011 by incorporating conditional forecast evaluation, allowing for model comparison based on specific economic regimes. This approach is particularly relevant in the context of stress-testing financial downside risk measures, such as Value-at-Risk (VaR) and Expected Shortfall (ES), which are critical for regulatory compliance under the Basel Accords.

Our theoretical contributions include the development of a state-dependent testing procedure that is asymptotically valid. By allowing forecast accuracy to be evaluated conditionally on specific market states, the CMCS approach offers a more robust and adaptable method for selecting models in volatile financial environments.

Empirically, we demonstrate the efficacy of the CMCS in stress-testing ES forecasts across different liquidity horizons implied by a variety of risk factors and economic conditions. The results show that different assets react differently to stress scenarios for different risk factors, and that the proposed CMCS method provides a consistent framework for identifying the best-performing models under these conditions. Importantly, the conditional testing approach captures the heterogeneous nature of risk exposure across assets and forecasting horizons, offering more insights into forecast performance compared to traditional unconditional methods.

Future work may examine different means of accounting for false discoveries, such as controlling the false discovery rate or the mixed directional FWER. This may improve the power of the CMCS and its ability to lead to more accurate forecasts, but is beyond the scope of this paper. Moreover, exploring other test statistics to construct the CMCS could result in improved finite sample properties. Finally, investigating additional empirical applications may help assess the usefulness of our proposed method.

\addcontentsline{toc}{section}{References} \printbibliography

appendix