EconBase
← Back to paper

Encompassing Tests for Value at Risk and Expected Shortfall Multi-Step Forecasts based on Inference on the Boundary

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

100,942 characters · 17 sections · 153 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Encompassing Tests for Value at Risk and Expected Shortfall Multi-Step Forecasts based on Inference on the Boundary

\def\spacingset#1{ {#1}} \spacingset{1}

abstractWe propose forecast encompassing tests for the Expected Shortfall (ES) jointly with the Value at Risk (VaR) based on flexible link (or combination) functions. Our setup allows testing encompassing for convex forecast combinations and for link functions which preclude crossings of the combined VaR and ES forecasts. As the tests based on these link functions involve parameters which are on the boundary of the parameter space under the null hypothesis, we derive and base our tests on nonstandard asymptotic theory on the boundary. Our simulation study shows that the encompassing tests based on our new link functions outperform tests based on unrestricted linear link functions for one-step and multi-step forecasts. We further illustrate the potential of the proposed tests in a real data analysis for forecasting VaR and ES of the S&P\,500 index.

{\it Keywords}: asymptotic theory on the boundary, joint elicitability, multi-step ahead and aggregate forecasts, forecast evaluation and combinations \\[0.5em] {\it JEL}: C12, C52, C58

\onehalfspacing

Introduction

For nearly two decades, financial institutions and regulators have advocated Value at Risk (VaR) as the main tool for risk management and capital allocation. Owing to a number of weaknesses, including the failure of capturing (extreme) tail risks and hence discouraging risk diversification Artzner1999, Acerbi2002, Tasche2002, the Basel Committee on Banking Supervision (BCBS) has recently adopted Expected Shortfall (ES), complementing and in parts substituting VaR as the fundamental measure for market risk Basel2013, Basel2016, Basel2017, Basel2019.

The ES at level $\alpha \in (0,1)$ is defined as the expected return beyond the $\alpha$-quantile and it is widely used as a coherent measure of tail risks Artzner1999, Tasche2002. Nonetheless, its inherent deficiency is that the ES is not elicitable on its own, meaning that the ES cannot be obtained as the unique minimizer of the expectation of a loss (scoring) function, see e.g., Gneiting2011. However, Fissler2016 show that the VaR and ES are jointly elicitable (or 2-elicitable). This joint elicitability property directly hints towards evaluating the ES jointly with the VaR in a unified framework FisslerZiegelGneiting2016, as in the present study concerning forecast encompassing tests.

Forecast encompassing of two competing forecasts tests whether one forecast alone performs not worse than any forecast combination, stemming from some parametric combination formula, also denoted by link functions in this article. If this holds, the rival forecast contains no additionally useful information relative to the first forecast HendryRichard1982, MizonRichard1986. This makes forecast encompassing tests an attractive tool for the empirical comparison of competing forecasts, especially when focusing on efficiency gains stemming from forecast combinations.\footnote{ For recent empirical applications of forecast encompassing tests, see e.g., Taylor2005, Busetti2013, FuertesOlmo2013, Costantini2017, Liu2017, Zhao2017, Tsiotas2018, Clements2020, YouLiu2020 among others.} As meaningful measures of forecast performance are based on strictly consistent loss functions Gneiting2011, this forcefully illustrates the importance of the existence of such loss functions for testing forecast encompassing. Hence, we build our encompassing tests on joint loss functions for the VaR and ES Fissler2016, and on recently developed joint semiparametric VaR and ES models Patton2019, DimiBayer2019, Taylor2019, Barendse2020.

As the main methodological contribution of this paper, we introduce encompassing tests for the ES jointly with the VaR based on flexible link functions or combination formulas, which allow for several important specifications that go beyond those of existing encompassing tests of e.g.\ GiacominiKomunjer2005 and DimiSchnaitmann2020. While linear forecast combination methods with unrestricted parameters are the most prominent class of link functions used for encompassing tests,\footnote{see e.g., HendryRichard1982,MizonRichard1986, Diebold1989, GiacominiKomunjer2005, ClementsHarvey2009, ClementsHarvey2010, DimiSchnaitmann2020.} more flexible approaches are especially important for joint tests of the VaR and ES: First, unrestricted linear link functions regularly result in VaR and ES crossings, i.e.\ days where the optimally combined ES forecast is larger than the VaR forecast, which immediately contradicts their definitions Taylor2020. To this end, we propose the no-crossing link functions which impede such crossings. Second, convex forecast combinations present an attractive alternative as their structure can stabilize the forecast performance and reduce the estimation noise Timmermann2006, Hansen2008, Bayer2018, which is particularly important for the case of semiparametric models for the VaR and ES with extreme probability levels DimiBayer2019.

The link function specifications considered in this article imply that certain model parameters are on the boundary of the parameter space under the null hypothesis. This boundary issue is exemplified by encompassing tests for convex forecast combinations, which entails testing whether the convex combination parameter is one (or zero). Under the null hypothesis, this parameter lies on the boundary of the admissible parameter space, i.e., the unit interval. Hence, we derive novel and nonstandard asymptotic theory for the model parameters and the resulting Wald test statistics for semiparametric models for the VaR and ES which allows some (or all) of the true model parameters to be on the boundary of the parameter space. For this, we follow the approach of Andrews1999 and Andrews2001, where the proofs use empirical process methods of Andrews1994 and Doukhan1995. To render our tests practically feasible, we draw critical values from the resulting nonstandard asymptotic distributions of the Wald test statistics obtained from simulations involving the solution of quadratic programming problems.

The proposed encompassing tests allow for testing one-step ahead, multi-step ahead and multi-step aggregate forecasts, where the consideration of multi-step forecasts requires the application of a VaR and ES specific adaption of the HAC (Heteroskedasticity and Autocorrelation Consistent) estimator of NeweyWest1987 and Andrews1991. The examination of multi-step (aggregate) forecasts is particularly relevant for the risk measures VaR and ES due to the explicit calls for 10-day aggregate VaR and ES forecasts of the Basel2016, Basel2017, Basel2019, Basel2020. Furthermore, this goes beyond many recent papers concerning forecast evaluation procedures for the VaR and ES, which mainly focus on one-step ahead forecasts.\footnote{see e.g.\ Kratz2017, CostanzinoCurran2018, BayerDimi2019, CouperierLeymarie2019, Patton2019, DimiSchnaitmann2020.}

Our simulations show that the encompassing tests for the VaR and ES based on our new link functions and on inference on the boundary exhibit accurate empirical sizes and good power properties. In particular, we find that these test specifications outperform classical tests based on unrestricted linear link functions for the VaR and ES of DimiSchnaitmann2020 throughout all considered simulation designs and for all, one-step ahead, multi-step ahead, and multi-step aggregate forecasts. However, we find that long forecast horizons (e.g., $10$ days) paired with short evaluation periods of less than $1000$ trading days result in unreliable test decisions. This simulation result sheds a critical light on the recent evaluation methods based on relatively short evaluation periods proposed by the Basel2016, Basel2017, Basel2019, Basel2020.

We empirically illustrate the usefulness of the VaR and ES encompassing tests based on convex link functions by comparing one-day ahead, and 10-day ahead and aggregate VaR and ES forecasts for daily S&P\,500 index returns. We estimate eleven risk models in a rolling forecast scheme, including several GARCH specifications and the ES-specific semiparametric models of Taylor2019 and Patton2019. For the evaluation period from July 2008 to June 2020, we find that the new tests assign much higher optimal weights to models specified with asymmetric volatilities, asymmetric and fat-tailed residual distributions and dynamic higher moments, especially for multi-step forecasts.

Our paper is closely related to the recent work of DimiSchnaitmann2020, but differs in the following ways. First, our proposed encompassing tests extend the ones of DimiSchnaitmann2020 with more flexible link functions by allowing the true parameters under the null to be on the boundary of the parameter space. Second, while the theoretical contribution of DimiSchnaitmann2020 focuses on the inclusion of misspecified models for encompassing tests, we establish inference on the boundary of the parameter space. Third, while our considered link functions allow for theoretically appealing specifications, they also exhibit clearly superior empirical properties in a wide range of simulations. Fourth, following the recent regulations of the Basel2019 (Basel2019, Basel2020), we consider multi-step ahead and aggregate VaR and ES forecasts in the simulations and the empirical application of this article.

The remainder of the paper is organized as follows. Section (ref) proposes the joint encompassing tests based on flexible link functions and develops asymptotic theory for the joint VaR and ES models for parameters on the boundary of the parameter space. Section (ref) presents simulations for the encompassing tests and Section (ref) applies the tests to VaR and ES forecasts for the S&P\,500 index. Section (ref) concludes. The proofs are given in Appendix (ref), and a supplementary material document contains additional material for the paper, where references starting with \textcolor{blue}{S.} refer to the supplement.

Theory

Setup and Notation

We follow the general setup of GiacominiKomunjer2005 and DimiSchnaitmann2020 while further allowing for multi-step forecasts. For this, we consider a stationary stochastic process $Z = \left\{ Z_t: \Omega \to \mathbb{R}^{\tilde l+1}, \, t= 1,\dots,R, \,\tilde l \in \mathbb{N}, \,R\in \mathbb{N} \right\}$, which is defined on some common and complete probability space $(\Omega, \mathcal{F}, \mathbb{P})$, where $\mathcal{F} = \left\{ \mathcal{F}_t, t = 1, \dots, R \right\}$ and $\mathcal{F}_t = \sigma \left\{ Z_s, s\le t \right\}$. We partition the stochastic process as $Z_t = (Y_t,X_t)$, where $Y_t: \Omega \to \mathbb{R}$ is an absolutely continuous random variable of interest and $X_t: \Omega \to \mathbb{R}^{\tilde l}$ is a vector of explanatory variables. For some fixed forecast horizon $h \in \mathbb{N}$, we denote the conditional distribution of $Y_{t+h}$ given the information set $\mathcal{F}_t$ by $F_t$. Accordingly, $\mathbb{E}_t$, $\operatorname{Var}_t$ and $h_t$ denote the expectation, variance and density corresponding to $F_t$. The conditional VaR of $Y_{t+h}$ given $\mathcal{F}_t$ at probability level $\alpha \in (0,1)$ is formally defined as

align[align omitted — 126 chars of source]

and given that $F_t$ is continuous at its $\alpha$-quantile, the conditional ES of $Y_{t+h}$ given $\mathcal{F}_t$ at level $\alpha \in (0,1)$ is defined by

align[align omitted — 149 chars of source]

In order to allow for forecast evaluation of multi-step (ahead and aggregate) forecasts with horizon $h \in \mathbb{N}$ in an out-of-sample fashion, we split $R = S+T+h-1$, where $S \in \mathbb{N}$ denotes the length of the in-sample and $T \in \mathbb{N}$ of the out-of sample window. In detail, for all $t \in \mathbb{N}$, such that $S \le t \le S+T-1$, we generate $h$-step ahead VaR and ES forecasts for the random variables $Y_{t+h}$ (i.e.\ for the sequence $(Y_{S+h}, \dots, Y_{S+T+h-1})$) based on the previous $S$ data points. For convenience of the notation, we define the set $\mathfrak{T} := \{ t \in \mathbb{N}: S \le t \le S+T-1 \}$ corresponding to the time points the forecasts are issued for the out-of-sample period.

We further denote the competing, $\mathcal{F}_t$-measurable, $h$-step forecasts for the VaR and ES by $\hat q_{j,t}$ and $\hat e_{j,t}$, for $j=1,2$. Following GiacominiKomunjer2005, we assume that these are generated through a function $f \big( \gamma_{t,S}, Z_t, Z_{t-1},\dots \big)$, which is fixed over time. For this, $\gamma_{t,S}$ denotes the (estimated or fixed) model parameters at time $t$, or alternatively the semi- or non-parametric estimator used in the construction of the forecasts, (possibly) estimated by data from the in-sample period of length $S$. This construction allows for forecasting schemes with fixed (or no) parameters, forecasting schemes with model parameters $\gamma_{t,S}$ that are estimated only once, and rolling window forecasting schemes where the parameters $\gamma_{t,S}$ are re-estimated in each step GiacominiKomunjer2005. In our testing approach, we focus on evaluation of the entire forecasting method as e.g.\ in GiacominiKomunjer2005 and GiacominiWhite2006, instead of on a forecasting model, as e.g.\ in West1996, West2001. The stacked forecasts are denoted by $\hat{\boldsymbol{q}}_t = (\hat q_{1,t}, \hat q_{2,t})$ for the VaR, and by $\hat{\boldsymbol{e}}_t = (\hat e_{1,t}, \hat e_{2,t})$ for the ES.\footnote{A generalization of our framework to test encompassing for multiple competing forecasts ($K\ge2$) in the sense of HarveyNewbold2000 is readily available by generalizing the notation as $\hat{\boldsymbol{q}}_t = (\hat q_{1,t}, \dots, \hat q_{K,t})$ and $\hat{\boldsymbol{e}}_t = (\hat e_{1,t}, \dots, \hat e_{K,t})$ and by further using suitable specifications for the link functions and the null hypotheses in the subsequent derivations.} In our notation of the forecasts, we stress the dependence on $t$, the time-point they are issued, while suppressing the dependence on the forecast horizon $h$ as it is treated as fixed.

Let $r_t$ denote financial log-returns for day $t$. Then, our theoretical setup allows for the treatment of classical multi-step ($h$-step) ahead forecasts, but also for $h$-step aggregate forecasts in the sense of an aggregated return over $h$ days, such as the 10-day aggregate VaR and ES forecasts explicitly stated in the regulatory framework of the Basel2019, Basel2020. For classical $h$-step ahead forecasts, we use $Y_{t+h} = r_{t+h}$, while for $h$-step aggregate forecasts we choose $Y_{t+h} = \sum_{s=1}^{h} r_{t+s}$.

In the following exposition, all vectors refer to column vectors. For splitting of subvectors, we often abuse notation and write $\theta = (\theta_1, \theta_2)$ instead of $\theta = (\theta_1^\top, \theta_2^\top)^\top$. The operator $\nabla$ denotes the derivative with respect to $\theta$. All limits below are taken “as $T \to \infty$” unless stated otherwise and $\overset{P}{\longrightarrow}$ and $\overset{d}{\longrightarrow}$ denote convergence in probability and distribution respectively. Let $:=$ denote an equality “by definition”. Furthermore, let $\mathbb{R}_+$ and $\mathbb{R}_-$ denote the non-negative and non-positive real half-lines respectively and we define $\mathbb{R}_C = \{ z \in \mathbb{R}: |z| \le C \}$ to be a sufficiently large compact subset of the real numbers (for some $C \in \mathbb{R}_+$ large enough).

Joint Encompassing Tests for VaR and ES Forecasts

For the introduction of the joint encompassing tests for VaR and ES forecasts, we follow DimiSchnaitmann2020 and define the flexible link (or combination) functions

alignat{3} &g^q: \mathfrak{Q} \times \mathfrak{E} \times \Theta \to \mathbb{R}&, \qquad &(\hat{\boldsymbol{q}}_t, \hat{\boldsymbol{e}}_t, \theta) \mapsto g^q(\hat{\boldsymbol{q}}_t, \hat{\boldsymbol{e}}_t, \theta), \\ &g^e: \mathfrak{Q} \times \mathfrak{E} \times \Theta \to \mathbb{R}&, \qquad &(\hat{\boldsymbol{q}}_t, \hat{\boldsymbol{e}}_t, \theta) \mapsto g^e(\hat{\boldsymbol{q}}_t, \hat{\boldsymbol{e}}_t, \theta),

based on some (compact) parameter space $\Theta \subset \mathbb{R}^k$, where $\mathfrak{Q}$ and $\mathfrak{E}$ denote the random spaces of the VaR and ES forecasts. These link functions represent the parametric, functional forms of the forecast combinations we consider.\footnote{The link functions can alternatively be interpreted as (semi-) parametric models for the conditional quantile (VaR) and ES of $F_t$ as in Patton2019.} E.g., in the classical case of testing forecast encompassing, these link functions are linear with (essentially) unrestricted parameter spaces. For convenience of notation, we henceforth use the short forms

align[align omitted — 243 chars of source]

We further assume that there exists a unique test parameter value $\theta_\ast \in \Theta$ such that $g^q(\hat{\boldsymbol{q}}_t, \hat{\boldsymbol{e}}_t, \theta_\ast) = \hat q_{1,t}$ and $g^e(\hat{\boldsymbol{q}}_t, \hat{\boldsymbol{e}}_t, \theta_\ast) = \hat e_{1,t}$ almost surely. This assumption ensures that the parametric link function allows for the trivial forecast combination of only choosing the first forecast.\footnote{As the encompassing tests in this article are always formulated as forecast (pair) one encompasses forecast (pair) two, we only assume the existence of $\theta_\ast$ corresponding to the first (pair) of forecasts. Testing the inverted encompassing hypothesis that the second pair of forecasts encompasses the first forecast pair can be carried out by interchanging the forecast pairs. Alternatively, one could assume that a value $\tilde \theta_\ast$ exists such that $g^q(\hat{\boldsymbol{q}}_t, \hat{\boldsymbol{e}}_t, \tilde \theta_\ast) = \hat q_{2,t}$ and $g^e(\hat{\boldsymbol{q}}_t, \hat{\boldsymbol{e}}_t, \tilde \theta_\ast) = \hat e_{2,t}$ holds almost surely.} In the classical case of unrestricted linear link functions, $\theta_\ast$ often corresponds to $(1,0)$ or $(0,1,0)$, depending on whether an intercept is included in the model. We refer to Section (ref) for details and examples of these link functions.

Gneiting2011 shows that the ES stand-alone is not elicitable, i.e.\ there do not exist suitable (strictly consistent) loss functions, which are a central ingredient for encompassing tests GiacominiKomunjer2005, DimiSchnaitmann2020. Fissler2016 overcome this deficiency and show that there exist joint loss functions for the VaR and ES and further characterize this class subject to mild regularity conditions by

align[align omitted — 340 chars of source]

where the function $\mathfrak{g}$ is twice continuously differentiable and increasing, $\phi$ is three times continuously differentiable, strictly increasing and strictly convex, and $a$ and $\mathfrak{g}$ are $Y_{t+h}$-integrable functions. The most prominent candidate of this class is the zero-homogeneous loss function NoldeZiegel2017AAS, sometimes called the FZ0 loss,

align[align omitted — 225 chars of source]

which is obtained by choosing $\mathfrak{g}(z) = 0$, $a(z) = 0$ and $\phi(z) = -\log(-z)$ in ((ref)). We henceforth often use the short notations $\rho_t(\theta) := \rho \big( Y_{t+h}, g^q_t(\theta) , g^e_t(\theta) \big)$ and $\rho^{\text{FZ0}}_t(\theta) := \rho^{\text{FZ0}} \big( Y_{t+h}, g^q_t(\theta) , g^e_t(\theta) \big)$.

Using the general class of loss functions in ((ref)), we define the true regression (or combination) parameter $\theta^0 \in \Theta$ by

align[align omitted — 201 chars of source]

which is independent of $t$ as we assume stationarity of the process $Z$.\footnote{See e.g.\ Patton2019, DimiBayer2019, BayerDimi2019, DimiSchnaitmann2020 and Barendse2020 for details on joint (semi-) parametric models for the VaR and ES.} The strict consistency result of the loss function from Fissler2016 together with further weak regularity conditions on the link functions implies that

align[align omitted — 155 chars of source]

almost surely, which justifies the notion of the true regression parameter.

We now define joint forecast encompassing for the VaR and ES following GiacominiKomunjer2005 and DimiSchnaitmann2020.

definition[Joint VaR and ES Forecast Encompassing] We say that the pair $\big( \hat q_{1,t}, \hat e_{1,t}\big)$ jointly encompasses $\big( \hat q_{2,t}, \hat e_{2,t}\big)$ at time $t$ with respect to the link functions $g^q$ and $g^e$ if and only if \begin{align} \mathbb{E} \left[ \rho \big( Y_{t+h}, \hat q_{1,t}, \hat e_{1,t} \big) \right] = \mathbb{E} \left[ \rho \big( Y_{t+1}, g^q(\hat{\boldsymbol{q}}_t, \hat{\boldsymbol{e}}_t, \theta^0), g^e(\hat{\boldsymbol{q}}_t, \hat{\boldsymbol{e}}_t, \theta^0) \big) \right], \end{align} where the loss function $\rho$ is given in ((ref)).

This holds if and only if $\theta^0 = \theta_\ast$ as we impose uniqueness of the parameter $\theta_\ast$. The intuition behind the specification in ((ref)) is that the forecasts $(\hat q_{1,t}, \hat e_{1,t})$ generate the same expected loss as an optimal forecast combination $\big( g^q(\hat{\boldsymbol{q}}_t, \theta^0), g^e(\hat{\boldsymbol{e}}_t, \theta^0) \big)$ based on the optimal combination parameter defined in ((ref)). Hence, using the first pair of forecasts $(\hat q_{1,t}, \hat e_{1,t})$ is the optimal, but trivial forecast combination. From a different point of view, this implies that the second pair of forecasts $(\hat q_{2,t}, \hat e_{2,t})$ does not add any useful information which is not already contained in $(\hat q_{1,t}, \hat e_{1,t})$.

If the interest is mainly placed on the performance of the competing ES forecasts, one can consider the auxiliary ES encompassing test in the spirit of DimiSchnaitmann2020.\footnote{Application of the strict encompassing test of DimiSchnaitmann2020 in the setting of the present article further requires combining the asymptotic theory under misspecification of DimiSchnaitmann2020 with the theory of estimation and testing at the boundary of the present article.}

definition[Auxiliary ES Forecast Encompassing] We say that the forecast $\hat e_{1,t}$ auxiliarily encompasses its rival $\hat e_{2,t}$ at time $t$ with respect to the link functions $g^q$ and $g^e$ if and only if \begin{align} \mathbb{E} \left[ \rho \big( Y_{t+h}, g^q(\hat{\boldsymbol{q}}_t, \hat{\boldsymbol{e}}_t, \theta^0), \hat e_{1,t} \big) \right] = \mathbb{E} \left[ \rho \big( Y_{t+1}, g^q(\hat{\boldsymbol{q}}_t, \hat{\boldsymbol{e}}_t, \theta^0), g^e(\hat{\boldsymbol{q}}_t, \hat{\boldsymbol{e}}_t, \theta^0) \big) \right], \end{align} where the loss function $\rho$ is given in ((ref)).

Finding testable conditions for the auxiliary test, corresponding to the condition $\theta^0 = \theta_\ast$ for the joint test, has to be done on a case-by-case basis for the link functions under consideration, see Section (ref) for further details.

Given a sample of competing forecasts and corresponding realizations, we can test whether the sequence of joint VaR and ES forecasts $( \hat q_{1,t}, \hat e_{1,t} )$ encompasses the sequence $( \hat q_{2,t}, \hat e_{2,t} )$ for all $t \in \mathfrak{T}$ (in the out-of-sample period) by estimating the parameters of the semiparametric models

align[align omitted — 199 chars of source]

where $Q_\alpha(u_{t+h}^q \mid \mathcal{F}_t) = 0$ and $\operatorname{ES}_\alpha(u_{t+h}^e \mid \mathcal{F}_t) = 0$ almost surely for all $t \in \mathfrak{T}$ by using the M-estimator introduced in Patton2019 and DimiBayer2019, and by testing whether $\theta_\ast = \theta^0$ using a Wald test.

Differently from DimiSchnaitmann2020 and the remaining literature on testing forecast encompassing, we allow the true, optimal parameter $\theta^0$ to be on the boundary of $\Theta$ under the null hypothesis. This facilitates the consideration of several important link function specifications. E.g., this enables to test encompassing for link specifications which theoretically prevent crossings of the combined VaR and ES forecasts in the sense that $g_t^e(\theta) \le g_t^q(\theta)$ almost surely for all $t \in \mathfrak{T}$ Taylor2020. Furthermore, we can test forecast encompassing based on convex forecast combinations, which stabilizes the parameter estimation. While the subsequent section focuses mainly on these two examples, our approach is by no means limited to these link functions.

The Link Function Specifications

In this section, we introduce three link function specifications which are of interest for this article, where other link functions can be treated along the lines of this section by employing an equivalent split of the parameter vector and by formulating the null hypotheses accordingly. The treatment of asymptotic theory on the boundary in the sense of Andrews2001, detailed in Section (ref) of the present article, requires splitting the parameter vector $\theta$ into the following structurally different subvectors,

align[align omitted — 102 chars of source]

where $\beta_1 \in \mathcal{B}_1 \subseteq \mathbb{R}^{p_1}$, $\beta_2 \in \mathcal{B}_2 \subseteq \mathbb{R}^{p_2}$, $\delta \in \Delta \subseteq \mathbb{R}^{q}$ and $\psi \in \Psi \subseteq \mathbb{R}^{s}$, where $p_1 + p_2 + q + s = k$, $p := p_1 + p_2$ and $\Theta = \mathcal{B}_1 \times \mathcal{B}_2 \times \Delta \times \Psi$. The intuition behind this decomposition is the following: (1) the null hypothesis we test for is based on $\beta_1$ only, and $\beta_1$ may or may not be on the boundary of the parameter space; (2) $\beta_2$ may or may not be on the boundary, but it is not tested for; (3) $\delta$ is not on the boundary, and it is not tested for; (4) $\psi$ is not tested for, it may or may not be on the boundary, and the off-diagonal elements of the matrix $\mathcal{T}$, defined later in ((ref)), corresponding to interactions of $\psi$ and $(\beta_1, \beta_2, \delta)$ are zero.

Most importantly, the null hypothesis is based on $\beta_1$ only, while the remaining parameters can be thought of as nuisance parameters, required for the estimation of the model. The distinction between $\psi$ and the remaining parameter subvectors (in particular $\beta_2$) is that the imposed nullity of certain off-diagonal elements of $\mathcal{T}$ implies that the asymptotic distribution of $\beta_1$ is not affected by whether $\psi$ is on the boundary or not.

Using the subvector decomposition in ((ref)), we can formally introduce the link functions and the corresponding null hypotheses of interest for the joint and auxiliary encompassing tests. The subsequent orderings of the parameters $\theta$ follows the ordering in the decomposition in ((ref)). All following encompassing null hypotheses are formulated for the test that the forecast pair $(\hat q_{1,t}, \hat e_{1,t})$ encompasses $(\hat q_{2,t}, \hat e_{2,t})$, whereas the reverse tests can be defined by simply interchanging the forecast pairs.

enumerate[label=(\arabic*)] • (Unrestricted) Linear: The unrestricted linear link functions are given by \begin{align} g_t^q(\theta) &= \theta_6 + \theta_3 \hat q_{1,t} + \theta_4 \hat q_{2,t}, \qquad and \\ g_t^e(\theta) &= \theta_5 + \theta_1 \hat e_{1,t} + \theta_2 \hat e_{2,t}, \end{align} where the parameter space $\Theta := \mathbb{R}_C^6$ is essentially unrestricted, as the constant $C$ can be chosen sufficiently large. We henceforth denote these link functions as linear. We then test (a) $\mathbb{H}_0^{\text{Joint}}: (\theta_1, \theta_2, \theta_3, \theta_4) = (1,0,1,0)$, and (b) $\mathbb{H}_0^{\text{Aux}}: (\theta_1, \theta_2) = (1,0)$.\footnote{In terms of the subvectors decomposition in ((ref)), we can assign $\beta_1 := (\theta_1, \theta_2, \theta_3, \theta_4)$ and $\delta := (\theta_5, \theta_6)$ for the joint test and $\beta_1 := (\theta_1, \theta_2)$ and $\delta := (\theta_3, \theta_4, \theta_5, \theta_6)$ for the auxiliary test. As the parameter subvector $\beta_1$ is in the interior of the parameter space under the null for both tests, classical asymptotic theory is sufficient for this unrestricted linear link function specification. } This corresponds to the standard case of forecast encompassing tests FairShiller1989, ClementsHarvey2009, which is already considered by DimiSchnaitmann2020 for the case of the VaR and ES. As none of the parameters are on the boundary under the null, standard asymptotic theory is sufficient here and we use this specification as the benchmark in this paper. • Convex Combinations: We consider the link functions \begin{align} g_t^q(\theta) &= \theta_4 + \theta_2 \hat q_{1,t} + (1-\theta_2) \hat q_{2,t}, \qquad and \\ g_t^e(\theta) &= \theta_3 + \theta_1 \hat e_{1,t} + (1-\theta_1) \hat e_{2,t}, \end{align} where $\Theta := [0,1]^2 \times\mathbb{R}_C^2$. We then test the following null hypothesis: \begin{enumerate} • $\mathbb{H}_0^{\text{Joint}}: (\theta_1, \theta_2) = (1,1)$, and we assign $\beta_1 := (\theta_1, \theta_2) \in \mathcal{B}_1 := [0,1]^2$, and $\delta := (\theta_3, \theta_4) \in \Delta := \mathbb{R}_C^2$. • $\mathbb{H}_0^{\text{Aux}}: \theta_1 = 1$, and we assign $\beta_1 := \theta_1 \in \mathcal{B}_1 := [0,1]$, $\beta_2 := \theta_2 \in \mathcal{B}_2 := [0,1]$ and $\delta := (\theta_3, \theta_4) \in \Delta := \mathbb{R}_C^2$. \end{enumerate} In comparison with the linear link functions, the convex forecast combinations require estimation of less parameters and therefore stabilizes the parameter estimation, especially for highly correlated forecasts.\footnote{Notice that for the estimation of joint VaR and ES models, especially for extreme probabilities such as $\alpha = 2.5\%$, adding additional parameters is costly in terms of both, computation times and estimation noise, see e.g.\ the simulations of DimiBayer2019 for details.} For both hypotheses formulated above, $\theta_1$ and $\theta_2$ are on the boundary under the null, while $\theta_3$ and $\theta_4$ are not. The latter parameters are assigned to $\delta$ instead of $\psi$ as the matrix $\mathcal{T}$, given in ((ref)), does not have null entries at the respective points. As the tested parameters are on the boundary of the parameter space under the null hypotheses of both tests, their corresponding Wald test statistics are subject to a non-standard asymptotic distribution Andrews1999, Andrews2001. • No VaR and ES Crossing: We consider the link functions \begin{align} g_t^q(\theta) &= \theta_3 + \theta_1 \hat e_{1,t} + (1-\theta_1) \hat e_{2,t} + \theta_2 \big(\hat q_{1,t} -\hat e_{1,t} \big) + (1-\theta_2) \big(\hat q_{2,t} -\hat e_{2,t} \big), \; \text{ and } \\ g_t^e(\theta) &= \theta_3 + \theta_1 \hat e_{1,t} + (1-\theta_1) \hat e_{2,t}, \end{align} where $\Theta := [0,1]^2 \times \mathbb{R}_C^2$. These link functions imply that $g_t^q(\theta) \ge g_t^e(\theta)$ holds almost surely for all $t \in \mathfrak{T}$, which can be interpreted as a necessary condition for sensible (combinations of) VaR and ES forecasts, which is closely related to the issue of \textit{quantile crossings} in quantile regression Koenker2005book. We then test \begin{enumerate} • $\mathbb{H}_0^{\text{Joint}}: (\theta_1, \theta_2) = (1,1)$, and we assign $\beta_1 := (\theta_1, \theta_2) \in \mathcal{B}_1 := [0,1]^2$, and $\delta := \theta_3 \in \Delta := \mathbb{R}_C$. • $\mathbb{H}_0^{\text{Aux}}: \theta_1 = 1$, and we assign $\beta_1 := \theta_1 \in \mathcal{B}_1 := [0,1]$, $\beta_2 := \theta_2 \in \mathcal{B}_2 := [0,1]$ and $\delta := \theta_3 \in \Delta := \mathbb{R}_C$. \end{enumerate} As in the convex setup, the tested parameters are on the boundary under the null and non-standard asymptotic theory is required.

While we focus on these examples of link functions in this article, the asymptotic theory presented in the subsequent section is valid for a many other interesting link functions, such as link functions without intercepts, nonlinear functions, and further specifications which prevent a crossing of the VaR and ES forecasts.

Asymptotic Theory on the Boundary of the Parameter Space

In this section, we derive the asymptotic theory for the M-estimator\footnote{In order to ensure global convergence of the M-estimator by avoiding local minima, we utilize the implementation of the R package esreg BayerDimi2019esreg based on the Iterated Local Search (ILS) meta-heuristic of Lourenco2003. See Section 3 of DimiBayer2019 for further details.} $\hat \theta_T$, given by

align[align omitted — 247 chars of source]

Classical asymptotic theory for the M-estimator $\hat \theta_T$, as given in Patton2019, states that given certain regularity conditions,

align[align omitted — 194 chars of source]

where

align[align omitted — 538 chars of source]

with

align[align omitted — 444 chars of source]

The function $\psi_t(\theta)$ corresponds to the gradient of the loss function $\rho_t(\theta)$ almost surely, i.e.\ on the set $\{ \theta \in \Theta: \; Y_{t+h} \not= g^q_t(\theta) \}$, which has probability one as the distribution $F_t$ is assumed to be absolutely continuous.

The asymptotic normality result in ((ref)) crucially depends on the regularity condition that the true parameter $\theta^0$ is in the interior of the parameter space, $\operatorname{int}(\Theta)$. This condition is violated under the null hypothesis for many interesting specifications of the link functions for the considered encompassing tests, as further outlined in Section (ref). Andrews1999 derives the non-standard asymptotic distribution of the parameter estimates in a general setup, which allows for parameters to be on the boundary and Andrews2001 extends this result to the asymptotic distribution of the resulting Wald test statistics.

Intuitively, the condition $\theta^0 \in \operatorname{int}(\Theta)$ implies that parameters to all sides (in a neighborhood) of $\theta^0$ are contained in $\Theta$ such that the estimator $\hat \theta_T$ is allowed to vary to all sides of $\theta^0$. The asymptotic normality result in ((ref)) formalizes this intuition by quantifying this variation as a limiting normal distribution. In contrast, if $\theta^0$ is on the boundary of $\Theta$, the estimator $\hat \theta_T$ cannot attain values to all sides of $\theta^0$, as values in some directions are excluded through the boundary. Consequently, in these cases the asymptotic distribution is more complicated and non-standard, which we formalize through deriving asymptotic theory on the boundary in the following. For this, we make the following assumptions. \newcounter{AssumptionCounter}

assumption$ $ \begin{enumerate}[label=(\Alph*)] • The parameter space is given as the product space $\Theta = \mathcal{B}_1 \times \mathcal{B}_2 \times \Delta \times \Psi$, where each of these four spaces is compact and restricted by individual inequality constraints: \begin{itemize} • $\mathcal{B}_1 = \big\{ \beta_1 \in \mathbb{R}^{p_1}: \Gamma_{\beta_1} \beta_1 \le r_{\beta_1} \big\}$, where $\Gamma_{\beta_1}$ is a $l_{\beta_1} \times p_1$ matrix and $r_{\beta_1}$ a $l_{\beta_1}$-dimensional vector, • $\mathcal{B}_2 = \big\{ \beta_2 \in \mathbb{R}^{p_2}: \Gamma_{\beta_2} \beta_2 \le r_{\beta_2} \big\}$, where $\Gamma_{\beta_2}$ is a $l_{\beta_2} \times p_2$ matrix and $r_{\beta_2}$ a $l_{\beta_2}$-dimensional vector, • $\Delta = \big\{ \delta \in \mathbb{R}^{q}: \Gamma_{\delta} \delta \le r_{\delta} \big\}$, where $\Gamma_{\delta}$ is a $l_{\delta} \times q$ matrix and $r_{\delta}$ a $l_{\delta}$-dimensional vector, • $\Psi = \big\{ \psi \in \mathbb{R}^{s}: \Gamma_{\psi} \psi \le r_{\psi} \big\}$, where $\Gamma_{\psi}$ is a $l_{\psi} \times s$ matrix and $r_{\psi}$ a $l_{\psi}$-dimensional vector. \end{itemize} • The process $Z_t$ is stationary and $\beta$-mixing of size $- r/( r-1)$ for some $ r>1$. • It holds that $\mathbb{E}\big[ \sup_{\theta \in \Theta} |\rho_t(\theta)|^{2r} \big] < \infty$ and $\mathbb{E} \left[ \sup_{\theta \in \Theta} ||\psi_t(\theta)||^{2r} \right] < \infty$ for all $\theta \in \Theta$ and some $\delta > 0$, where $ r>1$ is given in condition (ref).\footnote{We state these conditions as high-level moment conditions depending on $\rho_t(\theta)$ and $\psi_t(\theta)$. The derivations for primitive moment conditions for the semiparametric models for the VaR and ES for specific choices of the functions $\mathfrak{g}(\cdot)$ and $\phi(\cdot)$ are straight-forward, but the resulting conditions are often rather convoluted, see e.g.\ Appendix A of DimiBayer2019 and Assumption 2 (C) and (D) of Patton2019.} • The distribution of $Y_{t+h}$ given $\mathcal{F}_t$, denoted by $F_t$, is absolutely continuous with continuous and strictly positive density $h_t$, which is bounded from above almost surely on the whole support of $F_t$ and Lipschitz continuous. • The link functions $g_t^q(\theta)$ and $g_t^e(\theta)$ are $\mathcal{F}_t$-measurable, twice continuously differentiable in $\theta$ on $\operatorname{int}(\Theta)$ almost surely, and directionally differentiable on the boundary of $\Theta$. Moreover, if for some $\theta_1, \theta_2 \in \Theta$, $\mathbb{P} \big( g_t^q(\theta_1) = g_t^q(\theta_2) \cap g_t^e(\theta_1) = g_t^e(\theta_2) \big) = 1$, then $\theta_1 = \theta_2$. • The matrices $\mathcal{I}$ and $\mathcal{T}$ have full rank. • The matrix-elements of $\mathcal{T}$ governing the dependence of $(\beta_1, \beta_2, \delta)$ and of $\psi$ are zero. \setcounter{AssumptionCounter}{\value{enumi}} \end{enumerate}

Apart from the conditions (ref) and (ref), these assumptions are similar to the ones of Patton2019 and DimiSchnaitmann2020. However, as we base our proofs on stochastic equicontinuity and empirical process theory Andrews1994, instead of on the approach of Weiss1991, some of the conditions differ slightly. One main difference is that we assume the slightly stronger dependence condition of $\beta$-mixing (instead of $\alpha$-mixing) in order to show stochastic equicontinuity of the empirical process based on the theory of Doukhan1995. Notice that the parameter space in condition (ref) can conveniently be expressed through $l$ inequality constraints using an $l \times k$ matrix $ \Gamma_\theta$ and an $l$-dimensional vector $r_\theta$ as\footnote{In fact, $r_\theta = (r_{\beta_1}, r_{\beta_2}, r_{\delta}, r_{\psi})$ and by expressing $\Gamma_\theta$ as a $4 \times 4$ block matrix, the individual blocks $\Gamma_{\beta_1}$, $\Gamma_{\beta_2}$, $\Gamma_{\delta}$ and $\Gamma_{\psi}$ appear on its diagonal with rectangular zero-blocks everywhere else.}

align[align omitted — 123 chars of source]

This general formulation allows for flexible product spaces of closed real intervals.

theoremSuppose Assumption (ref) holds. Then \begin{align} \sqrt{T} \big( \hat \theta_T - \theta^0 \big) \overset{d}{\longrightarrow} \hat \lambda, \qquad where \qquad \hat \lambda = \underset{\lambda \in \Lambda}{\arg \inf}\; (\lambda - Z)^\top \mathcal{T} (\lambda - Z), \end{align} with $Z= \mathcal{T}^{-1} G$, $G \sim \mathcal{N}(0, \mathcal{I})$ and $\Lambda = \big\{ \lambda \in \mathbb{R}^k: \Gamma_{\theta}^{(b)} \lambda \le 0 \big\}$, where $\Gamma_{\theta}^{(b)}$ denotes the submatrix of $\Gamma_\theta$ from (ref), which consists of the rows of $\Gamma_\theta$ for which all inequalities $\Gamma_{\theta}^{(b)} \theta^0 \le r_\theta$ hold as an equality.

The proof of Theorem (ref) verifies the necessary assumptions in Andrews1999 and Andrews2001.\footnote{ Notice that the notation in Andrews2001 includes the nuisance parameter $\pi \in \Pi$ which we do not require. Thus, following the comment on p.692 of Andrews2001, we simply employ a parameter space $\Pi = \{\pi_0\}$ consisting of a single point $\pi_0$, e.g.\ $\pi_0 = 0$, and suppress the dependency on $\pi$ in the notation. } If $\theta^0 \in \operatorname{int}(\Theta)$, none of the inequalities in ((ref)) is binding and $\Lambda = \mathbb{R}^k$. This implies that $\hat \lambda = Z$ almost surely in ((ref)), which results in the classical asymptotic normality result given in ((ref)). In contrast, if $\theta^0$ is on the boundary of $\Theta$, the $\arg \inf$ in ((ref)) results in a non-standard asymptotic distribution of the stabilizing transformation $\sqrt{T} \big( \hat \theta_T - \theta^0 \big)$.

Subvector Inference

In the notation of the subvector decomposition of $\theta$ in ((ref)), we only test parametric restrictions for the subvector $\beta_1$, which might be substantially smaller than $\theta$. Thus, the formulation of the arg\,inf in ((ref)) might be unnecessarily complex in these situations. To address this issue, we derive inference for the subvector $\beta = (\beta_1, \beta_2)$ of $\theta$ by following the general approach of Andrews1999, Andrews2001. In some instances, this considerably simplifies the solution of the arg\,inf in ((ref)).

For this, we define the subvector $\gamma := (\beta, \delta) = (\beta_1, \beta_2, \delta)$, which contains all parameters in $\theta$ but $\psi$, with the intuition that $\psi$ does not have any influence on the asymptotic distribution of $\gamma$ through the nullity restrictions on $\mathcal{T}$ imposed in condition (ref) in Assumption (ref). We define the following quantities for the subvectors $\beta$ and $\gamma$,

align[align omitted — 154 chars of source]

where $ {\mathcal{T}_\gamma}$ denotes the upper-left $(p+q) \times (p+q)$ submatrix of $\mathcal{T}$ and $G_\gamma$ the upper $(p+q)$-dimensional subvector of $G$. The following theorem states the asymptotic distribution of the subvector $\beta$.

theoremGiven Assumption (ref), it holds that \begin{align} \sqrt{T} \big( \hat \beta_T - \beta^0 \big) \overset{d}{\longrightarrow} \hat \lambda_\beta, \end{align} where \begin{align} \hat \lambda_\beta = \underset{\lambda_\beta \in \Lambda_\beta}{\arg \inf}\; (\lambda_\beta - Z_\beta)^\top \big( H {\mathcal{T}_\gamma}^{-1} H^\top \big)^{-1} (\lambda_\beta - Z_\beta), \end{align} and $\Lambda_\beta = \big\{ \lambda_\beta \in \mathbb{R}^p: \Gamma_{\beta}^{(b)} \lambda _\beta\le 0 \big\}$. The matrix $\Gamma_{\beta}^{(b)}$ denotes the sub-matrix of $\Gamma_{\beta}$, which consists of the rows of $\Gamma_{\beta}$ for which the inequality $\Gamma_{\beta} \beta^0 \le r_\beta$ holds as an equality.

Theorem (ref) shows that the asymptotic distribution of $\beta$ is entirely unaffected by the parameter $\psi$. In contrast, the subvector $\delta$ (which is contained in $\gamma$) influences the asymptotic distribution of $\beta$ through the weighting matrix in the quadratic programming problem in ((ref)), even though $\delta$ itself is not on the boundary of the parameter space.

While closed-form representations for the distribution of $\hat \lambda_\beta$ (and of $\hat \lambda$) are only available in special cases Andrews1999, we can conveniently simulate from its distribution in a straight-forward fashion by solving a quadratic programming problem. For this, notice that the minimization problem in ((ref)) is equivalent to solving

align[align omitted — 349 chars of source]

where $\Gamma_\beta^{(b)}$ is given as in Theorem (ref) and specifies the binding inequality restrictions of $\Lambda_\beta$. Consequently, we can draw samples from the Gaussian random variable $G_\gamma$, and for each sampled value, we solve the quadratic programming problem given in ((ref)). The respective solutions then form a sample of the random variable $\hat \lambda_\beta$, whose distribution is asymptotically equivalent to the one of $\sqrt{T} \big( \hat \beta_T - \beta^0 \big)$.

The Wald Test Statistic

We now consider a Wald test for the null hypothesis $\mathbb{H}_0: \beta_1 = \beta_{1 \ast}$ for some $\beta_{1 \ast} \in \mathcal{B}_1$, which may or may not be on the boundary of $\mathcal{B}_1$. We define the Wald test statistic for the null hypothesis $\mathbb{H}_0: \beta_1 = \beta_{1\ast}$ as

align[align omitted — 162 chars of source]

with weighting matrix $\hat V_T^{-1}$, given by

align[align omitted — 173 chars of source]

where $H_1 := [I_p, \mathbf{0}_{p \times q}]$, and where $\hat{\mathcal{T}}_{T \gamma}$ and $\hat{\mathcal{I}}_{T \gamma}$ are the upper left $(p+q) \times (p+q)$ submatrices of $\hat{\mathcal{T}}_{T}$ and $\hat{\mathcal{I}}_{T}$, respectively, which are consistent estimators for the matrices $\mathcal{T}$ and $\mathcal{I}$. For the matrix $\mathcal{T}$, we use the estimator

align[align omitted — 512 chars of source]

where the bandwidth $c_T$ satisfies $c_T = o(1)$ and $c_T^{-1} = o(T^{1/2})$. In the specification of $\hat{\mathcal{T}}_{T}$, the term $\mathds{1}_{\{| Y_{t+h} - g_t^q(\hat \theta_T))| \le c_T \}}/(2 c_T)$ is a nonparametric estimator of the conditional density $h_t(g^q_t(\theta^0))$, which is also employed in Engle2004 and Patton2019.

As we allow for multi-step ahead (aggregate) forecasts in this treatment, we employ a HAC estimator NeweyWest1987, Andrews1991 for the matrix $\mathcal{I}$,

align[align omitted — 330 chars of source]

based on some weight functions $z(j,m) \to 1$ and the bound (or bandwidth) $m_T = o(T^{1/4})$. Furthermore, $\psi_t(\theta)$ is given in ((ref)) and we define $\mathfrak{T}_j := \{ t \in \mathbb{N}: S+j \le t \le S+T-1 \}$ for all $j \ge 0$. As the functions $\psi_t( \theta)$ are not continuous in $\theta$, we generalize the consistency proofs of the HAC estimator in NeweyWest1987 to nonsmooth objective functions in Lemma (ref) in the supplementary material. For the asymptotic distribution of the Wald test statistic, we impose the following assumptions.

assumption$ $ \begin{enumerate}[label=(\Alph*)] \setcounter{enumi}{\value{AssumptionCounter}} • $m_T \to \infty$ such that $m_T = o(T^{1/4})$ and $z(j,m) \to 1$ as $m \to \infty$. • $c_T = o(1)$ and $c_T^{-1} = o(T^{1/2})$. • The functions $g_t^q(\theta)$ and $g_t^e(\theta)$ are three times continuously differentiable (in $\theta$) and the following moments are finite, $\mathbb{E} \left[ \sup_{\tilde{\theta} \in U(\theta, \delta) } \left|\left| \nabla_\theta \tilde A_t(\theta) \right| \right|^{2 r} \right]$, $\mathbb{E} \left[ \sup_{\tilde{\theta} \in U(\theta, \delta) } \left|\left| \nabla_\theta \tilde B_t(\theta) \right| \right|^{2 r} \right]$, \\ $\mathbb{E} \left[ \sup_{\tilde{\theta} \in U(\theta, \delta) } \left| \tilde A_t(\tilde \theta) \right|^{2r} \times \sup_{\tilde{\theta} \in U(\theta, \delta) } \left|\left| \nabla_\theta g_t^q(\tilde \theta) h_t(g_t^q(\tilde \theta)) \right| \right|^{2r} \right]$, \\ and $\mathbb{E} \left[ \sup_{\theta \in \Theta} \left| \left| \psi_t( \theta) \right| \right|^{2(r+\delta)} \right]$, for some $\delta > 0$, where $\tilde A_t(\theta)$ and $\tilde B_t(\theta)$ are given in ((ref)) and ((ref)) in the supplementary material. \end{enumerate}

Conditions (ref) and (ref) are standard in the literature on HAC estimators and estimating the conditional density, see e.g., NeweyWest1987, Engle2004 and Patton2019. The strengthened moment conditions (ref) are required to establish stochastic equicontinuity of the discontinuous function $\frac{1}{T} \sum_{t \in \mathfrak{T}_j} \psi_t(\theta) \psi_{t-j}^\top (\theta)$ for consistency of the HAC estimator.

theoremSuppose Assumption (ref) and Assumption (ref) hold. Then \begin{align} W_T \overset{d}{\longrightarrow} W := \hat \lambda_{\beta_1}^\top V^{-1} \hat \lambda_{\beta_1}, \end{align} where $V$ denotes the probability limit of $\hat V_T$ and $\hat \lambda_{\beta_1}$ is the upper $p_1$-dimensional subvector of $ \hat \lambda_{\beta}$, given in Theorem (ref).

Using the simulation procedure for the distribution of $\hat \lambda_{\beta}$ described after Theorem (ref), we can easily simulate draws from $\hat \lambda_{\beta_1}$ and consequently from the distribution of $W$ by using the formula in ((ref)). Hence, we obtain simulated, asymptotic critical values for the Wald test statistic.

We further use a variant of the HAC estimator NeweyWest1987, Andrews1991, which is specifically designed for the semiparametric VaR and ES models. For most classical HAC estimators, estimation of the contemporaneous variance $\mathbb{E} \big[ \psi_t(\theta^0) \psi_{t}^\top(\theta^0) \big]$ is straight-forward by employing a sample counterpart. The major challenge in consistently estimating the matrix $\mathcal{I}$ in ((ref)) is then the inclusion of the (sample) autocovariances $\mathbb{E} \big[ \psi_t( \theta^0) \psi_{t-j}^\top( \theta^0) \big]$ such that the resulting estimator is positive definite.

However, for the VaR and ES, and especially for extreme quantile levels, estimation of the contemporaneous variance $\mathbb{E} \big[ \psi_t( \theta^0) \psi_{t}^\top( \theta^0) \big]$ is cumbersome in itself as it depends on the conditional truncated variance $\operatorname{Var}_t( Y_{t+h} | Y_{t+h} \le g_t^q(\theta^0))$, see e.g.\ DimiBayer2019. For this, we employ the scl-sp estimator of DimiBayer2019, which is based on the regularizing assumption that the quantile residuals $u_{t+h}^q = Y_{t+h} - g_t^q(\theta^0)$ follow a location-scale model, conditional on the employed covariates. Imposing a location-scale model might cause some misspecification in the estimation, but it allows to use all observations to estimate a conditional variance, and then obtain the conditional truncated variance through a transformation formula for location-scale models. We obtain this estimator by replacing the outer product estimator of the contemporaneous variance by the scl-sp estimator,

align[align omitted — 193 chars of source]

where $\widetilde{\Omega}_{T,0}$ denotes the scl-sp estimator of DimiBayer2019.

Even though the parametric link functions in ((ref)) depend explicitly on the forecasts $\hat{\boldsymbol{q}}_t$ and $\hat{\boldsymbol{e}}_t$, it is important to note that the asymptotic theory of this section also holds for general semiparametric models for the VaR and ES in the sense of Patton2019. Consequently, the asymptotic theory and the proposed Wald test can further be employed for testing (the nullity) of coefficients in the dynamic models of Taylor2019 and Patton2019, which are on the boundary of the parameter space under the null hypothesis. Furthermore, the strict ES encompassing test of DimiSchnaitmann2020 allows for testing encompassing of ES forecasts without their accompanying VaR forecasts, which potentially introduces model misspecification in the parametric models. The asymptotic theory for the M-estimator presented here can easily be adapted to the misspecified case by replacing the matrices $\mathcal{T}$ and $\mathcal{I}$ with their misspecification-robust counterparts of DimiSchnaitmann2020, and by replacing the respective steps in the proof of Theorem (ref).

Simulations

In this section, we evaluate the empirical properties of the encompassing tests based on the three different link functions specified in Section (ref), and on the asymptotic theory of Section (ref). Section (ref) numerically illustrates the effect testing on the boundary has on the asymptotic distribution of the parameters. Subsequently, we analyze the size and power properties of the encompassing tests in Section (ref) for one-step ahead forecasts and in (ref) for multi-step ahead and aggregate forecasts.

The Asymptotic Distribution on the Boundary

We illustrate how true parameters on the boundary of the parameter space affect the asymptotic distribution of the M-estimator through simulations. For this, we simulate data according to the standard GARCH model with Gaussian innovations described in ((ref)) in Section (ref) with an out-of-sample window length of $T=2500$. We estimate the parameters of the three considered link functions for the joint encompassing test that tests whether forecasts stemming from the (true) GARCH model encompass forecasts from the GJR-GARCH model given in ((ref)).

Figure (ref) illustrates the distribution of the parameter estimates by plotting histograms over 10000 simulation replications for the intercept and slope parameters of the respective ES link functions $g_t^e$, whose true values equal zero and one respectively throughout all link functions. For the (unrestricted) linear link function, all true parameters are in the interior of the parameter space and we find that the histograms for both parameters closely approximate the asymptotic normal distribution, derived and employed by Patton2019 and DimiSchnaitmann2020. In contrast, for the convex and no-crossing link functions, the slope parameter is bounded between zero and one, i.e.\ its true value of one is on the boundary of the parameter space. This results in the non-standard distributions illustrated by the histograms for the slope parameters in the second and third plot in the lower row of Figure (ref). The histograms approximate the asymptotic distribution consisting of a mixture of a point mass at one and a half-normal distribution, which is considerably different from asymptotic normality. This behavior directly carries over to the resulting asymptotic distributions of the Wald test statistics which substantiates the necessity of the non-standard asymptotic theory on the boundary presented in Section (ref).

figure[figure omitted — 338 chars of source]

While this behavior is not unexpected for the parameters on the boundary, the asymptotic distribution of the intercept parameters, which themselves are in the interior of the parameter space, is also affected due to the joint estimation. For instance, we observe a slight skewness in the distribution of the intercept parameter of the convex link function contrasting the Gaussian distribution of the linear intercept parameter.

One-Step Ahead Forecasts

In this section, we investigate the empirical performance of our new encompassing tests for one-step ahead forecasts. For this, we consider encompassing of VaR and ES forecasts stemming from a standard GARCH and a GJR-GARCH model Bollerslev1986, Glosten1993, which are given by $r_{j,t+1} = \sigma_{j,t+1} u_{t+1}$, for $j=1,2$, where the two distinct volatility specifications are given by

align[align omitted — 275 chars of source]

Furthermore, we employ two different residual distributions,

align[align omitted — 169 chars of source]

where the latter denotes a skewed $t$-distribution, parameterized as in fernandez1998bayesian and giot2003value, with zero mean, unit variance, a skewness parameter of $0.8$ and $5$ degrees of freedom. For the two GARCH models paired with the two residual distributions, optimal one-step ahead VaR and ES forecasts are given by $\hat q_{j,t} = z_\alpha \sigma_{j,t+1}$ and $\hat e_{j,t} = \xi_\alpha \sigma_{j,t+1}$ for $j=1,2$, where $z_\alpha$ and $\xi_\alpha$ are the $\alpha$-quantile and $\alpha$-ES of the standard normal and the skewed $t$-distribution, respectively.\footnote{Regarding the time index, notice that $\hat q_{j,t}$ and $\hat e_{j,t}$ represent $\mathcal{F}_t$-measurable forecasts for the return $r_{j,t+1}$, while $\sigma_{j,t+1}$ is equivalently based on time $t$ information and corresponds to the conditional volatility of $r_{j,t+1}$.} For both distributions, we simulate $Y_{t+1} = r_{t+1} = \big( (1-\pi) \sigma_{1,t+1} + \pi \sigma_{2,t+1} \big) u_{t+1}$ for 11 equally spaced values of $\pi \in [0,1]$, where $u_{t+1}$ is given as in ($a$) and ($b$) in (ref).

We consider encompassing tests comparing the respective GARCH and GJR-GARCH volatility specifications, where we analyze the models based on Gaussian and $t$-distributed residuals in separate simulation setups. For each forecast pair, we test two null hypotheses: the first tests whether the first forecast encompasses the second, indicated by $\mathbb{H}_0^{(1)}$, whereas the second tests the reverse, i.e.\ that forecast two encompasses forecast one, indicated by $\mathbb{H}_0^{(2)}$. These two null hypotheses correspond to the cases $\pi = 0$ and $\pi=1$ in the simulation design above. For all intermediate values of $\pi \in (0,1)$, the returns are generated as linear combinations of the models, and both null hypotheses should be rejected. For both encompassing tests, we employ the scl-sp covariance estimator of DimiBayer2019 described in Section (ref).\footnote{Section (ref) in the supplemental material shows that the results for employing a HAC estimator are qualitatively equivalent for one-step ahead forecasts.} All following results are based on 2000 Monte Carlo replications.

table[table omitted — 3,777 chars of source]

Table (ref) reports the empirical test sizes of the joint VaR and ES and the auxiliary ES encompassing tests based on the three link functions described in Section (ref) for a nominal size of $5\%$. For this, we consider the two GARCH specifications described in (ref) and (ref) for various out-of-sample sizes ranging from $T=250$ to $T=5000$. We find that the tests based on the convex and no-crossing link functions outperform the ones build on the linear link function, especially for smaller out-of-sample sizes: the tests based on the linear link function are in some instances severely oversized, while the other two link functions exhibit empirical sizes generally below $10\%$, even for the smallest of the considered sample sizes. Note for this that a sample size of $T=250$ is considered to be very small for VaR and ES forecasts at a probability level of $\alpha = 2.5\%$, as this corresponds to only six VaR violations on average. This result can be explained by the reduced number of estimated parameters for both, the convex and no-crossing link functions, and by the theoretically appealing property of excluding VaR and ES crossings for the no-crossing specification.

We further find that the auxiliary ES encompassing test generally exhibits more accurate (smaller) sizes than the joint VaR and ES test throughout almost all considered designs. This behavior is particularly evident for the process with skewed-$t$ innovations. As the joint test includes testing of the quantile parameters, the asymptotic covariance matrix additionally contains the density quantile function $h_t( g_t^q(\theta^0))$ in (ref), which is particularly challenging to estimate for small probability levels (see e.g.\ Koenker1978, Koenker2005book, DimiBayer2019).

figure[figure omitted — 866 chars of source]

Figure (ref) shows size-adjusted power curves for the joint VaR and ES and the auxiliary ES tests based on the three link functions for a nominal significance level of $5\%$ and for the various settings described above.\footnote{Figure (ref) in the supplementary material shows the corresponding raw power of the tests.} For computing the size-adjusted power, we follow the approach of DavidsonMacKinnon1998. For an increasing degree of misspecification through $\pi$, we find increasing power throughout all considered tests and processes. Both, the convex and no-crossing link function specifications exhibit better (size-adjusted) power than the linear link function throughout all considered processes, sample sizes and values of $\pi$. While the convex link function exhibits a slightly superior performance for the joint test, the no-crossing link function performs slightly better for the auxiliary ES test. We provide simulation results for two additional processes outside the location-scale family in Section (ref) in the supplementary material, where the results for these forecasts are comparable to those obtained here.

Multi-Step Ahead and Aggregate Forecasts

In this section, we consider multi-step ahead and multi-step aggregate forecasts for the VaR and ES. For any $h > 1$, we set $Y_{j,t+h} = r_{j,t+h}$ for multi-step ahead forecasts, and $Y_{j,t+h} = \sum_{s=1}^h r_{j,t+s}$ for multi-step aggregate forecasts, where the returns $r_{j,t+h}$ are simulated from the respective GARCH specifications in (ref) - (ref) for $j=1,2$. In order to simulate returns which follow a (probabilistic) convex combination of these two processes, we simulate Bernoulli draws $\pi_{t+h} \sim \operatorname{Bern}(\pi)$ for 11 equally spaced values of $\pi \in [0,1]$, and let $Y_{t+h} = (1-\pi_{t+h}) Y_{1,t+h} + \pi_{t+h} Y_{2,t+h}$.

WongSo2003 and Loennbark2016 among others illustrate that even though the conditional variance of multi-step ahead (aggregate) forecasts for (quadratic) GARCH models is easily tractable, the entire conditional distribution is not. This implies that multi-step ahead (aggregate) VaR and ES forecasts cannot be obtained equivalently to one-step ahead forecasts by simply multiplying their conditional multi-step ahead (aggregate) volatilities with the quantile or ES of the residual distribution. Consequently, we employ a simulation method proposed by WongSo2003 which yields very accurate approximations of the true VaR and ES forecasts: for all out-of-sample time points $t \in \mathfrak{T}$, we simulate $R=10000$ sample paths from the respective GARCH model for $h$ days into the future and in order to obtain multi-period ahead (aggregate) VaR and ES forecasts, we (point-wisely) take the empirical quantile and ES over the $R$ sample paths of the simulated $h$-period ahead (aggregated) returns.

Here, we restrict attention to the DGP based on Gaussian residuals, the convex link function and on the joint VaR and ES encompassing test as the $t$-distributed residuals and the auxiliary tests perform comparably in the previous section. However, we consider $h$-step ahead and $h$-step aggregate VaR and ES forecasts with forecasting horizons of $h=1,2,5$ and 10 days. This allows to investigate the properties of the test for increasing forecast horizons $h$. We employ a HAC estimator with the embedded scl-sp estimator of DimiBayer2019 for the contemporaneous variance as described in Section (ref), as in particular the multi-period aggregate forecasts exhibit a correlated behavior due to their inherently overlapping nature. In Section (ref) in the supplementary material, we discuss four different covariance estimators and show that the HAC estimator augmented with the scl-sp estimator performs best.

table[table omitted — 2,113 chars of source]

Table (ref) reports the tests sizes and Figure (ref) presents size-adjusted power\footnote{The size-adjusted power plots for $h=10$ and $T \in \{250,500,1000\}$ in Figure (ref) exhibit test sizes under the null hypotheses slightly above $5\%$. These are an artifact stemming from the fact that slightly more than $5\%$ of the simulated $p$-values are exactly zero, rendering an exact size-adjustment in the sense of DavidsonMacKinnon1998 infeasible.} plots of the joint VaR and ES encompassing test for multi-step ahead and multi-step aggregate forecasts for a nominal significance level of $5\%$.\footnote{Figure (ref) in the supplementary material shows the corresponding raw power. Table (ref), Figure (ref) and Figure (ref) in the supplementary material show test results for the auxiliary ES encompassing test.} The encompassing tests for $h$-step ahead forecasts are well-sized, especially for larger sample sizes and for small horizons $h$. The empirical sizes deteriorate slightly with an increasing forecast horizon $h$. While the general behavior is similar for $h$-step aggregate forecasts, these tests suffer considerably more from an increase of the forecasting horizon $h$. The inferior performance of multi-period aggregate forecasts is not surprising given that the moment conditions of the aggregate forecasts are heavily correlated due to the overlapping definition of the aggregate forecasts.

figure[figure omitted — 1,005 chars of source]

Concerning the size-adjusted power, depicted in Figure (ref), we observe similar patterns. For $h=1,2$, the size-adjusted power increases substantially for an increasing degree of misspecification for all considered settings. For longer forecast horizons $h=5,10$, the test power is generally lower for both forecast types. As before, the encompassing tests for $h$-step ahead forecasts exhibit better properties than for $h$-step aggregate forecasts. This can again be explained by the inherent correlation in $h$-step aggregate forecasts which necessitates the use of a sufficiently large amount of observations for a consistent estimation of the asymptotic covariance matrix together with its nuisance quantities and autocorrelation structure. Hence, with an increasing forecast horizon $h$, a larger out-of-sample period is required to obtain encompassing tests with reliable test decisions. Furthermore, small sample sizes paired with large forecast horizons (e.g., $T=250$ and $h=10$) yield almost flat (size-adjusted) power curves which implies that the test becomes unreliable and practical applications should be interpreted very carefully in these scenarios.\footnote{Along these lines, Harvey2017 notice similar small-sample issues for forecast encompassing tests and tests for equal predictive ability DieboldMariano1995 for multi-step ahead forecasts.} This negative result is remarkable concerning the planned evaluation of 10-day ahead aggregate ES forecasts Basel2019.

Empirical Application

This section empirically illustrates the usefulness of the proposed encompassing tests by comparing alternative VaR and ES forecasts for daily S&P\,500 returns from August 4, 2000 to June 19, 2020 including a total of 5000 daily observations. We conduct a rolling window forecasting scheme with $S=2000$ estimation observations, and $T=3000$ evaluation points starting on July 22, 2008. We follow the Basel Accords Basel2017, Basel2019 and employ $\alpha = 2.5\%$. Based on the simulation results of Section (ref), we restrict our attention to the tests based on the convex link functions with intercepts in the empirical application.

Particularly, we consider a total of eleven competing risk models for forecasting VaR and ES, including: (i) a rolling Historical Simulation using the window length of 250 days, (ii) the RiskMetrics model, (iii) the GARCH(1,1) model with normal innovations (GARCH-N), and the GJR-GARCH(1,1) model of Glosten1993 with skewed Student-$t$ distributed innovations (GJR-ST), (iv) the GARCH and GJR-GARCH models with asymmetric Laplace innovations (GARCH-AL and GJR-AL) and the same models with a time varying shape parameter (GARCH-AL-TVP and GJR-AL-TVP) of Chen2012, (v) the symmetric absolute value (SAV-) and asymmetric slope (AS-) CAViaR-ES models of Taylor2019, and (vi) the one factor GAS model (GAS-1F) of Patton2019. Details for the risk models of Chen2012, Taylor2019 and Patton2019 are given in Section (ref) in the supplementary material and an additional absolute evaluation in the form of backtests for these models is given in Section (ref) in the supplementary material.

One-Step Ahead Forecasts

In this subsection, we analyze pairwise encompassing for one-step ahead VaR and ES forecasts using the encompassing tests based on the convex link functions. For each model pair, we estimate the combination weights and test both null hypotheses, i.e.\ that model one encompasses model two and vice versa. We obtain simulated critical values for the test through Theorem (ref) and by employing the scl-sp estimator of DimiBayer2019. Due to the simulation results of Section (ref) in the supplementary material, we do not consider estimation of HAC-terms in the covariance for one-step ahead forecasts.

table[table omitted — 2,428 chars of source]

We report the summarized results for the joint VaR and ES and the auxiliary ES encompassing tests with a significance level of $5\%$ for all pairwise combinations of the eleven risk models in Table (ref), where the models (in the table rows) are sorted according to their encompassing performance. Out of the ten model combinations each individual model is subject to, we report the instances how often both null hypotheses are rejected (denoted by "Combination" or "Comb"), not rejected ("Inconclusive" or "Incon"), only the first one is rejected ("Encompassed" or "E'ed"), and only the second one is rejected ("Encompassing" or "E'ing"). Notice that the "Combination" column is based on rejecting both null hypotheses, which constitutes a multiple testing problem and the results have to be interpreted at a Bonferroni corrected significance level of $10\%$, while each individual tests are based on a nominal significance level of $5\%$.\footnote{Table (ref) in the supplementary material reports the correlations of the VaR and ES forecasts and Table (ref) additionally reports the estimated (convex) combination weights together with the test decisions for the combinations of the six bestperforming models, chosen by the absolute evaluation in Table (ref).}

We find that the GJR-GARCH models with Skew-t and asymmetric Laplace innovations achieve the best forecasting performance among the competing models. Interestingly, the CAViaR-ES models of Taylor2019 and the GAS-1F model of Patton2019, which are specifically developed for jointly forecasting VaR and ES, generally do not perform as good as the GARCH specifications. As expected, the RiskMetrics and Historical Simulation models perform worst. Furthermore, we find many instances of rejections of both encompassing hypotheses, implying that a forecast combination via the estimated encompassing weights is superior to both individual models. This result justifies the usefulness of the proposed encompassing tests, and is in line with the arguments for forecast combinations of GiacominiKomunjer2005, Timmermann2006, Taylor2020 and DimiSchnaitmann2020.

10-Step Ahead and Aggregate Forecasts

In this subsection, we apply the proposed encompassing tests to 10-day ahead and aggregate VaR and ES forecasts. Note that 10-day aggregate VaR and ES forecasts are required by the Basel Accords for minimal capital requirement and risk weighted assets Basel2019, Basel2020. As it is unclear how to obtain multi-step ahead forecasts from the CAViaR-ES models of Taylor2019 and the GAS-1F model of Patton2019, we reduce the set of evaluation models to the seven members of the GARCH family. For these models, we obtain multi-step ahead and multi-step aggregate VaR and ES forecasts through the simulation method of WongSo2003, further described in Section (ref). Such a simulation-based forecasting is necessary as the conditional distribution of multi-step returns generally differs from the imposed innovation distribution of the model, and thus, VaR and ES forecasts cannot be obtained through classical location-scale formulas as for one-step ahead forecasts. Based on the results of Section (ref) in the supplementary material, we use a HAC covariance estimator NeweyWest1987, augmented with the scl-sp estimator of DimiBayer2019 for the contemporaneous variance component to perform the encompassing tests for multi-step ahead and aggregate forecasts.

table[table omitted — 3,204 chars of source]

Table (ref) reports the summarized encompassing test results for 10-step ahead and aggregate forecasts.\footnote{Table (ref) and Table (ref) in the supplementary material report the detailed test results. Table (ref) additionally reports correlations for the 10-day ahead and aggregate VaR and ES forecasts.} The test results show that for both, 10-step ahead and aggregate forecasts, the best performing model is the GJR-GARCH model with asymmetric Laplace innovations and a time-varying shape parameter. We find almost no cases of double rejections, i.e.\ forecast combinations are not (significantly) preferred over the stand-alone models. This can be an artifact from the lower power for multi-step forecasts as illustrated in Section (ref) or from the high(er) correlations of the forecasts, reported in Table (ref) in the supplementary material.

Overall, the empirical results show that a model specified with an asymmetric volatility process and a skewed error distribution, such as the GJR-ST model, outperforms the competing models considered in this paper for one-step ahead VaR and ES forecasts. Moreover, models based on an asymmetric innovation distribution with time-varying parameters, such as, GJR-AL-TVP and GARCH-AL-TVP models, perform better than the other competing models. Note that the time-varying scale parameter of the asymmetric Laplace distribution produces both time-varying skewness and kurtosis for the innovation distribution. We find that specifying time-varying higher moments for a risk model substantially improves the model forecasting performance in both multi-step ahead and aggregate risk forecasts, much more than in one-step ahead forecasts.

Conclusion

This article proposes joint encompassing tests which compare one-step and multi-step VaR and ES forecasts based on general semiparametric forecast combination methods (link functions) for the VaR and ES. While unrestricted linear methods are often employed in encompassing tests for functionals like the mean and quantiles (the VaR) as e.g.\ in HendryRichard1982, GiacominiKomunjer2005, different combination methods are of particular interest for the ES. E.g., our no-crossing link specification theoretically circumvents crossings of the predicted VaR and ES, which is conceptually desirable but not straight-forward to achieve Taylor2020.

Our employed link functions imply that some of the tested parameters are on the boundary of the parameter space under the null hypothesis, which necessitates non-standard asymptotic theory. Based on the general framework of Andrews1999, Andrews2001, we provide such novel asymptotic theory for the proposed encompassing tests and for the accompanying Wald test statistics, which allows for inference and testing on the boundary. Our simulations show that the proposed VaR and ES forecast encompassing tests based on the convex and no-crossing link functions exhibit superior size and power properties than those based on unrestricted linear link functions. By employing the proposed encompassing tests in a real data analysis, we find that building risk models on specifications including time-varying higher moments substantially improves the model forecasting performance, especially for multi-step ahead and aggregate VaR and ES forecasts.

Our framework allows for several straight-forward extensions. Encompassing tests for multiple VaR and ES forecasts in the sense of HarveyNewbold2000 can directly be implemented through our asymptotic theory by adapting the link functions. Furthermore, the asymptotic theory allows for encompassing tests based on any strictly consistent loss function for the VaR and ES. Incorporating estimation risk into these tests can be obtained by combining our theory with the work of EscancianoOlmo2010, Du2017, BarendseKole2019, and incorporating model misspecification through combining our theory with the one of DimiSchnaitmann2020. Encompassing tests for different functionals such as e.g., the mean, quantiles, expectiles or probability densities based on link functions which require testing on the boundary (e.g., using convex link functions) can be implemented through adapting our asymptotic theory to semiparametric models for the functional under consideration. Eventually, our asymptotic theory can be used to test (e.g., for nullity of) model parameters on the boundary of the parameter space for the semiparametric VaR and ES models of Patton2019, Taylor2019 or Gerlach2019, along the lines of Francq2009.

Acknowledgments

Our work has been supported by the University of Hohenheim, the Klaus Tschira Foundation and the University of Konstanz. A previous version of this paper circulated with the title "A Regression-based Joint Encompassing Test for Value-at-Risk and Expected Shortfall Forecasts".

\onehalfspacing {5pt}