Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
75,746 characters · 12 sections · 66 citation commands
How to Compare Copula Forecasts?
\baselineskip18pt \setcounter{totalnumber}{50} \setcounter{topnumber}{50} \setcounter{bottomnumber}{50} \abovedisplayskip1.5ex plus1ex minus1ex \belowdisplayskip1.5ex plus1ex minus1ex \abovedisplayshortskip1.5ex plus1ex minus1ex \belowdisplayshortskip1.5ex plus1ex minus1ex
Modeling the dependence between several outcomes is of interest in forecasting economic, financial or meteorological variables, among others. Since the seminal work of Skl59, it is known that copulas provide a complete description of the dependence structure of a random vector with continuous marginal distributions. Therefore, applications of copula modeling are by now widespread in forecasting contexts. For instance, in finance, copulas have been used for forecasting the Value-at-Risk GHS09, for portfolio construction Pat04, for predictive modeling of exchange rates CKL13, and for systemic risk prediction BC19. We refer to the review articles by MR12 and Pat13 for many more examples, and to MLT13 for a meteorological forecasting application.
A key step in copula modeling is the choice of a suitable parametric family, such as the Gaussian copula, the $t$-copula, an Archimedean copula, or any mixture thereof. If dependence varies over time, an additional choice involves the driving process for the copula parameters, such as the dynamics proposed by Pat06 or the more recent generalized autoregressive score (GAS) dynamics of CKL13. The sheer number of choices that a practitioner is faced with in copula modeling calls for statistically sound tools for model selection and model comparison.
The main purpose of this paper is to develop such tools. Model comparison commonly hinges on strictly consistent scoring functions (often also called loss functions). Taking as arguments the forecast and the observation, these scoring functions are strictly minimized in expectation by the true model choice. If such a strictly consistent score exists for a certain functional of interest, such as the mean or a quantile, this functional is called elicitable. E.g., the mean functional is elicitable, with the square loss as a strictly consistent scoring function. But not all functionals are elicitable, e.g., variance and Expected Shortfall fail to be elicitable Gne11. The reason is that these functionals do not have convex level sets (CxLS), which is a necessary condition for elicitability Osband1985, Gne11. Our first main result, Proposition (ref), shows that copulas generally fail to be elicitable, due to a lack of the CxLS property.
A notable exception to our non-elicitability result is on Fr\'echet classes, that is, on classes $\mathcal{F}$ of multivariate distributions where the distributions only differ in terms of their dependence structure, i.e., their copulas, but their marginal distribution structure is the same, i.e., for any $F, G\in\mathcal{F}$, it holds that $F_i = G_i$ for all marginal cumulative distribution functions $F_i$ and $G_i$ ($i=1, \ldots, d$). In practical terms, this means that the correct marginal model is known, and the forecaster is only concerned with model choice for the copula. We provide possible strictly consistent scores for this important situation. In the previous literature on copula choice, researchers have often relied on the Kullback--Leibler information criterion (KLIC) to decide between different copula models with fixed marginals CF06,DPV10. Our positive result not only formally justifies this practice by grounding it in a sound decision-theoretic framework, but it also points towards a wealth of other scores that can be used (see Proposition (ref)).
To mitigate the problem of (in general) non-elicitable copulas, it is possible to adopt a strategy pioneered by Osband1985. He suggests to turn a non-elicitable functional into an elicitable one by pairing it with an appropriate auxiliary functional; the pairs (variance, mean) and (Value-at-Risk, Expected Shortfall) being prime examples of this strategy FZ16a. Applied to our situation, we can pair the copula with the marginal distributions. This amounts to specifying the full joint distribution, which is elicitable with proper scores readily available, e.g., the log-score or the continuous ranked probability score GR07. However, when comparing the full joint distribution with proper scores, we cannot determine if differences in the ranking are due to superior marginal forecasts or due to superior copula forecasts.
We address this attribution problem by drawing on the recently proposed notion of multi-objective elicitability FH24. This concept works with bivariate (instead of real-valued) scores, which are uniquely minimized in expectation by the true forecast with respect to the lexicographic order. Such scores are called strictly consistent multi-objective scores. Our second main contribution is to show that strictly consistent multi-objective scores exist for the tuple consisting of the copula and all marginal distributions (see Theorem (ref)).
The third main contribution of this article is to show how the multi-objective scores from Theorem (ref) can be used in two-step tests for predictive ability. In the first step, the marginals are tested for equal predictive accuracy. The second step then tests for equal or superior predictive ability in the copula forecasts. This procedure allows a rejection of the null to be clearly attributed to the accuracy of the marginal or of the copula forecast (see Section (ref)). At each step, our tests are akin to standard Diebold--Mariano tests DM95. However, critical values are computed in such a way that size is controlled across the two steps.
Monte Carlo simulations demonstrate that our two-step tests hold size even in relatively small samples (Section (ref)). Also the power of the tests in discriminating between copula forecasts of different quality is high. Our study also reveals that the tests' power in the copula component is higher when the marginal forecasts are less seriously misspecified. Moreover, we find that attribution also works well in the sense that differences in predictive ability of the marginals (copula) is detected almost exclusively in the first step (second step).
We illustrate the use of our two-step tests in an application to copula predictions for returns from international stock market indices (Section (ref)). The forecasting models we consider present the current state-of-the-art, including DCC--GARCH models of Eng02 and $t$--GAS models inspired by CKL13. Using our proposed methods, one of our findings is that $t$--GAS models and DCC--GARCH models produce equally accurate forecasts for the marginal densities, yet copula forecasts of the latter model are to be preferred.
Once a copula model has been chosen using the methodology proposed in this paper, a natural question is whether the adopted copula provides a sufficiently accurate description of the dependence structure. This is important because scores only provide information on the relative quality of forecasts, therefore only delivering the “least bad” choice. For checking the absolute quality of the copula forecasts, the methods of ZiegelGneiting2014 may be used. In this sense, our results nicely complement the aforementioned ones.
The remainder of the paper proceeds as follows. In Section (ref), we introduce some basic notation and definitions. Section (ref) states our main theoretical results and Section (ref) introduces our two-step tests. Section (ref) contains the Monte Carlo simulations and Section (ref) presents our empirical application. The final Section (ref) concludes with a discussion of our results. To promote flow in the main text, the proofs of all technical results are relegated to the Appendix.
We assume that all random variables in the following are defined on a non-atomic probability space $(\Omega,\mathcal{A},\ensuremath{\mathrm{P}})$. Suppose that $\bm Y=(Y_1,\ldots,Y_d)^\prime$ is the random vector, whose dependence structure is to be modeled. We denote its joint cumulative distribution function (cdf) by $F$ and the marginal cdfs by $F_i$ ($i=1,\ldots,d$). Then, by Skl59's Skl59 theorem MFE15 there exists a cdf $C\colon[0,1]^d\rightarrow[0,1]$ with uniform marginals, called a copula, such that
From this representation it is clear that the copula fully describes the dependence structure of $\bm Y$, with marginal information absorbed in the $F_i$. Skl59 also shows that the copula is unique if all marginal cdfs are continuous. Vice versa, for any copula $C$ and any marginal distributions $F_1,\ldots,F_d$, the right-hand side of (ref) defines a cdf on $\mathbb{R}^d$. Following almost all of the copula literature, we shall assume continuous marginals in this paper, implying that the copula $C$ in (ref) is unique.
In the sequel, we denote the set of all $d$-dimensional copulas by $\mathcal{C}$, and the set of all $d$-dimensional cdfs with continuous marginals by $\mathcal{F}(\mathbb{R}^d)$. In particular, $\mathcal{F}(\mathbb{R})$ denotes the set of all continuous cdfs on $\mathbb{R}$. The subset $\mathcal{F}([0,1]^d)\subset\mathcal{F}(\mathbb{R}^d)$ contains all cdfs with support on $[0,1]^d$. The set $\mathcal{F}$ will always refer to a subset of $\mathcal{F}(\mathbb{R}^d)$. Using this notation, Sklar's theorem implies the existence of the mapping \[ T_{\operatorname{cop}}\colon \mathcal{F}(\mathbb{R}^d) \to \mathcal{C}, \] which maps a joint cdf $F\in\mathcal{F}(\mathbb{R}^d)$ with continuous marginals to its unique copula $T_{\operatorname{cop}}(F)\in\mathcal{C}$, satisfying (ref). Finally, a Fr\'echet class is a subclass of $\mathcal{F}(\mathbb{R}^d)$ with fixed and given marginal distribution structure $\{F_i\}_{i=1,\ldots, d}$. That means, any two members of a Fr\'echet class only differ in terms of their copulas.
Consider some functional $T\colon\mathcal{F}\to\mathsf{A}$, mapping any cdf in $\mathcal{F}$ to some action domain $\mathsf{A}$. In the decision-theoretic framework of Gne11, forecasts for $T$ are evaluated and ranked in terms of scoring functions $S\colon \mathsf{A}\times\mathbb{R}^d\to\mathbb{R}$. They assign a forecast $x\in\mathsf{A}$ the (negatively oriented) score $S(x,\bm y)\in\mathbb{R}$ if the observation $\bm y\in\mathbb{R}^d$ materializes. Scoring functions are also often called loss functions in the econometric literature and are then denoted by an $L$ GW06,LRV12,CI20. The score $S$ is called strictly $\mathcal{F}$-consistent for $T$ if the expected score is uniquely minimized by the correctly specified functional, that is, if \[ \operatorname{E}\big[S(T(F),\bm Y)\big]<\operatorname{E}\big[S(x,\bm Y)\big]\qquad\text{for all }F\in\mathcal{F}, \ \bm Y\sim F, \text{ and for all } x\in\mathsf{A}, \ x\neq T(F). \] Consequently, strictly consistent scoring functions incentivize truthful forecasting. To see this, suppose that the goal is to elicit the true belief of the forecaster for the functional $T$. If the forecaster is “rewarded” with the negative score $-S(x,\bm Y)$ for a forecast $x$ and a verifying observation $\bm Y\sim F$ that materializes, then her expected reward $\operatorname{E}[-S(x,\bm Y)]$ is maximized by the true report $x=T(F)$. In this sense, the scoring function $S$ serves as a “truth serum” that elicits the true belief. This example also explains why we say that $T$ is elicitable on $\mathcal{F}$, if a strictly $\mathcal{F}$-consistent scoring function exists for a functional $T$.
Strict consistency of a scoring function tells us that lower expected scores are to be preferred, because the expected score is minimized by the true report. In practice, by the law of large numbers, we can approximate the expected score by the sample average of the scores. Then, the forecast with the lowest average score is the preferred one. The following example makes this concrete.
If the functional $T$ maps to some finite-dimensional action domain $\mathsf{A}\subseteq \mathbb{R}^k$, we speak of a point forecast; e.g., when $T$ is the mean functional or a quantile. For $d=k=1$, the prime examples for strictly consistent scores are the square loss, $S(x,y) = (x-y)^2$, for the mean, and the absolute loss, $S(x,y)=|x-y|$, for the median. On the other hand, probabilistic forecasts amount to specifying the full predictive distribution. In this case, $T$ is the identity map on $\mathcal{F}$ and the action domain $\mathsf{A}$ corresponds to $\mathcal{F}$ itself. Traditionally, strictly consistent scoring functions for probabilistic forecasts are called strictly proper scores or strictly proper scoring rules GR07. Examples include the log-score $S_{\log}(F,\bm y)=-\log f(\bm y)$ (with $f(\cdot)$ denoting the density of $F$) and the continuous ranked probability score (CRPS) $S_{\operatorname{CRPS}}(F,\bm y)=\operatorname{E}\Vert\bm X-\bm y\Vert-(1/2)\operatorname{E}\Vert\bm X-\tilde{\bm X}\Vert$ for independent $\bm X\sim F$ and $\tilde{\bm X}\sim F$ GR07.
As a matter of fact, the copula functional $T_{\operatorname{cop}}\colon \mathcal{F}(\mathbb{R}^d)\to\mathcal{C}$ we study in this paper bridges the realms of point forecasts and probabilistic forecasts; it summarizes the distribution (like point forecasts), yet with an infinite-dimensional action domain (like probabilistic forecasts).
To establish the non-elicitability of the copula functional $T_{\operatorname{cop}}$, we recall that the convex level sets property is a necessary condition for elicitability Osband1985,Gne11.
Even though the original proof of Lemma (ref) is provided for point forecasts, that is, $\mathsf{A}\subseteq \mathbb{R}^k$, the proof works exactly the same for any kind of action domain, in particular also for $\mathsf{A}=\mathcal{C}$. We provide the proof in the Appendix for the sake of completeness.
At first sight, it may appear that the focus on distributions with bounded support in Proposition (ref) is restrictive. However, the opposite is the case. Even if one considers a very small class of distributions that only contains distributions with bounded support (including their translations and one mixture) do copulas fail to be elicitable on that class. This then trivially implies that copulas are not elicitable on larger classes of distributions (such as $\mathcal{F}(\mathbb{R}^d)$).
Copulas share the property of non-elicitability with many other functionals, such as the variance, the Expected Shortfall and the mode Gne11,Hei14. The practical implication of the non-elicitability is that copula forecasts cannot be sensibly compared on their own. This is because no score exists for ranking the copula forecasts that satisfies even the minimal quality criterion of being uniquely minimized in expectation by the true report.
In contrast to Proposition (ref), the copula functional has convex level sets on any Fr\'echet class $\mathcal{F}$. Recall that a Fr\'echet class is defined by the property that any members $F,G\in\mathcal{F}$ may only differ in their dependence structure, while their marginal distributions coincide, i.e., $F_i=G_i$ for all $i=1,\ldots,d$. In particular, assuming that the copulas of two distributions in $\mathcal{F}$ coincide therefore means the two distributions already coincide. Hence, also any convex combination coincides. This is a necessary condition for the following elicitability result.
The score in (ref) directly draws on Sklar's Theorem (ref) and exploits the strict propriety of $S$ on $\mathcal{F}$. On the other hand, the score in (ref) exploits the fact that if $\bm Y$ has continuous marginals $F_1,\ldots, F_d$, the vector of probability transforms $\big(F_1(Y_1), \ldots, F_d(Y_d)\big)'$ is distributed like the copula of $\bm Y$. Hence, the first argument of the score can directly use a copula. Interestingly, the score in (ref) stays strictly consistent even without the assumption of continuous marginals. And on larger classes with varying marginal forecasts, it constitutes a strictly proper scoring rule; see Proposition (ref). In contrast, (ref) hinges on the assumption of continuous marginals and cannot be extended to a strictly proper score on its own when varying the marginal forecasts. However, it will be a building block for our multi-objective scores of Theorem (ref).
Proposition (ref) demonstrates that copula forecasts can be compared via strictly consistent scoring functions on Fr\'echet classes, i.e., if the forecasts of the marginal distributions structure is fixed and correctly specified. When this requirement is not met, it is natural to ask---in the spirit of Osband1985---whether the copula is jointly elicitable with some other functional, the marginals being the natural candidates. Indeed, it is immediate from (ref) that specifying the copula and the marginal distributions amounts to specifying the complete predictive distribution. Therefore, the copula and the marginals are jointly elicitable, as we formally state next.
Similarly as the pairs (variance, mean) and (Value-at-Risk, Expected Shortfall), the pair (copula, marginals) is elicitable. Yet, comparing the copula and the marginals jointly as suggested in the above proposition, amounts to a test of the full predictive distribution. This does not allow us to pinpoint differences in average scores to either the copula or the marginals, completely defeating the purpose of the exercise.
To tackle this issue, we draw on the concept of multi-objective scoring functions introduced by FH24. In particular, instead of scalar scores (taking values in $\mathbb{R}$), we use two-dimensional scores taking values in $\mathbb{R}^2$. The use of scalar scores (such as the squared loss) pervades the forecasting literature. This is due to historical reasons and the fact that scalars can easily be compared using the canonical order $\leq$ on $\mathbb{R}$. When considering (as we do) scores in $\mathbb{R}^2$, we need to equip $\mathbb{R}^2$ with an order that ultimately allows us to compare forecasts.
A suitable order on $\mathbb{R}^2$ in that context is given by the lexicographic order, $\preceq_{\mathrm{lex}}$ FH24. Recall that on $\mathbb{R}^2$ it holds that $(x_1,x_2)^\prime\preceq_{\mathrm{lex}} (y_1,y_2)^\prime$ if $x_1<y_1$ or if ($x_1=y_1$ and $x_2\leq y_2$). Obviously, the lexicographic order is a total order on $\mathbb{R}^2$, meaning that we can compare any two points of $\mathbb{R}^2$. This is important because our end goal is to compare predictions for which it is crucial that a conclusive ranking emerges. For example, while $(3,5)^\prime\preceq_{\mathrm{lex}} (5,3)^\prime$, it neither holds that $(3,5)^\prime\geq_{\operatorname{comp}} (5,3)^\prime$ nor $(3,5)^\prime\leq_{\operatorname{comp}} (5,3)^\prime$ with respect to the componentwise order $\leq_{\operatorname{comp}}$ on $\mathbb{R}^2$ (which is not a total order), such that the points cannot be compared. We shall write $(x_1,x_2)^\prime\prec_{\mathrm{lex}} (y_1,y_2)^\prime$ if $(x_1,x_2)^\prime\preceq_{\mathrm{lex}} (y_1,y_2)^\prime$ and $(x_1,x_2)^\prime\neq (y_1,y_2)^\prime$.
The next result is one of our main contributions and establishes the multi-objective elicitability of the copula functional together with the marginal distributions with respect to the lexicographic order. This means that the action domain corresponds to $\mathsf{A}=\mathcal{C}\times \big(\mathcal{F}(\mathbb{R})\big)^d$.
Similarly to univariate strictly consistent scores, our bivariate strictly consistent multi-objective scores also play the role of a “truth serum”: Suppose that, in a first step, the forecaster receives a “reward” of $-S_{\operatorname{marg}}\big(\{F_i\}_{i=1,\ldots, d}, \bm Y\big)$ if she issues the forecasts $\{F_i\}_{i=1,\ldots, d}$ and $\bm Y$ materializes and, in a second step, gets $-S_{\operatorname{cop}}\big(C,\{F_i\}_{i=1,\ldots, d}, \bm Y\big)$ for the copula forecast $C$. Then, in expectation, the step-wise reward-maximizing strategy is to report the true marginals of $\bm Y$ in the first step and the true copula in the second.
Interestingly, the second component in (ref), $S_{\operatorname{cop}}$, coincides with the score in (ref) of Proposition (ref). The reason is that Theorem (ref) nests the situation of Proposition (ref). The latter treats the situation when the marginal distribution structure is given and fixed. Hence, in a comparison with the bivariate scores from (ref), the first component $S_{\operatorname{marg}}$ is the same, such that the comparison solely hinges on the second component, $S_{\operatorname{cop}}$.
We now demonstrate how to use the theoretical results of the previous section for formal hypothesis testing. The $\mathbb{R}^{d}$-valued verifying observations for the copula forecasts are denoted by $\bm Y_1,\ldots,\bm Y_n$ and the information set relevant to the forecaster at time $t$ is $\mathcal{F}_{t}$. Examples include the trivial sigma field $\mathcal{F}_{t}=\{\emptyset,\Omega\}$ (i.e., unconditional “forecasting”) or the information set generated by past observations $\mathcal{F}_{t}=\sigma(\bm Y_{t},\bm Y_{t-1},\ldots,)$. However, the precise content of $\mathcal{F}_{t}$ does not matter in the following. Let us denote by $F_{1t}$ and $F_{2t}$ the competing $\mathcal{F}_{t-1}$-measurable distributional forecasts for the cdf of $\bm Y_t\mid\mathcal{F}_{t-1}$ ($t=1,\ldots,n$). From these, the marginal forecasts $F_{1it}$ and $F_{2it}$ ($i=1,\ldots,d$), and the copula forecasts $C_{1t}$ and $C_{2t}$ can be obtained.
Since the score is the metric by which to judge forecast quality (also for multi-objective scores), tests comparing predictive accuracy are based on the bivariate score differences
When the marginal forecasts are identical, then $d_{m,t}=0$ such that there can only be differences in the second component $d_{c,t}$. In this case, one can use the standard DM95 procedure to test equal predictive ability (i.e., $\operatorname{E}[d_{c,t}]=0$) or superior predictive ability (i.e., $\operatorname{E}[d_{c,t}]\leq0$) in the copula component; see Example (ref). Recall that, as a lower score is preferred, the null $\operatorname{E}[d_{c,t}]\leq0$ corresponds to the situation where the first set of forecasts is superior (or, more precisely, non-inferior) to the second set of forecasts.
Therefore, the case different from the standard testing framework arises when the marginal forecasts are different. The two null hypotheses we consider in this case are
The two-sided hypothesis $H_0^{=}$ simply tests equal predictive accuracy, while the “one and a half-sided” hypothesis $H_0^{\preceq_{\mathrm{lex}}}$ tests whether the marginals are equally accurate and the first set of copula forecasts is non-inferior to the second set.
We stress that, despite the fact that we are interested only in the copula forecasts in this paper, it is essential to also test for equally accurate marginals (i.e., $\operatorname{E}[d_{m,t}]=0$) in both $H_0^{=}$ and $H_0^{\preceq_{\mathrm{lex}}}$. To see this, recall that due to the non-elicitability of copulas (Proposition (ref)), copula forecasts cannot generally be compared on their own. Only when the marginals are equally accurate (a special case of this being identical marginal forecasts), is it possible to sensibly compare copula forecasts. This is because only when $\operatorname{E}[d_{m,t}]=0$ (such that $\operatorname{E}[S_{\operatorname{marg}}(\{F_{1it}\}_{i=1,\ldots,d}, \bm Y_t)] = \operatorname{E}[S_{\operatorname{marg}}(\{F_{2it}\}_{i=1,\ldots,d}, \bm Y_t)]$) does the lexicographic order consider the copula component, such that the ordering of $\operatorname{E}[S_{\operatorname{cop}}(C_{1t},\{F_{1it}\}_{i=1,\ldots,d}, \bm Y_t) ]$ and $\operatorname{E}[S_{\operatorname{cop}}(C_{2t},\{F_{2it}\}_{i=1,\ldots,d}, \bm Y_t) ]$ becomes decisive in preferring one forecast over the other. Therefore, it is essential in $H_0^{=}$ and $H_0^{\preceq_{\mathrm{lex}}}$ to also test $\operatorname{E}[d_{m,t}]=0$.
We test both $H_0^{=}$ and $H_0^{\preceq_{\mathrm{lex}}}$ via the average score differences $\overline{\bm d}_n:=(\overline{d}_{m,n}, \overline{d}_{c,n})^\prime:=\frac{1}{n}\sum_{t=1}^{n}\bm d_t$, which form the sample counterpart to $\operatorname{E}[\bm d_t]$. Let $c_{1n}$ and $c_{2n}$ denote two critical values, whose precise choice is specified in Theorem (ref) below. We are inclined to reject $H_0^{=}$ if in a first step $\sqrt{n}|\overline{d}_{m,n}|>c_{1n}$ or---if the first step did not lead to a rejection---it holds in a second step that $\sqrt{n}|\overline{d}_{c,n}|>c_{2n}$. Our test of $H_0^{\preceq_{\mathrm{lex}}}$ proceeds similarly except that the second step rejects if $\sqrt{n}\overline{d}_{c,n}>c_{2n}$.
Since $\overline{d}_{m,n}$ and $\overline{d}_{c,n}$ are generally correlated, a crucial quantity in computing $c_{1n}$ and $c_{2n}$ is an estimate of the long-run variance $\bm \varOmega_n=\operatorname{Var}(\sqrt{n}\overline{\bm d}_n)$ of the average score differences. We estimate $\bm \varOmega_n$ via
where $m_n$ is an integer sequence of cutoffs and $w_{n,h}$ is a triangular array of scalar weights. Both $m_n$ and $w_{n,h}$ are specified in more detail in the following assumption.
The point of Assumption (ref), which is similar to the conditions in Theorem 4 of GW06, is merely to ensure that a multivariate central limit theorem holds for $\bm d_t$ and that $\bm \varOmega$ can be estimated consistently. Hence, it can be replaced by any other assumption implying these two properties. We refer to GW06 for some extensive discussion of the wide range of (marginal and copula) estimation schemes covered by the mixing assumption B2.
Theorem (ref) amounts to a step-wise application of the standard Diebold--Mariano test. Yet, the critical values are chosen differently, such that the overall (asymptotic) size of our two-step test is preserved. Figure (ref) depicts the non-rejection regions of the above two-step tests for $H_0^{=}$ in panel (a) and for $H_0^{\preceq_{\mathrm{lex}}}$ in panel (b). These non-rejection regions depend on the critical values $c_{1n}$ and $c_{2n}$, which are implicitly determined by (ref) and (ref). In our numerical examples, we choose $c_{1n}$ and $c_{2n}$ such that both probabilities on the left-hand sides of (ref) and (ref) are equal to $\alpha/2$, respectively. Note that the computation of rectangular probabilities for multivariate normal distributions (as required for the computation of $c_{1n}$ and $c_{2n}$) used to be a computational challenge Gen04. Nowadays, this no longer poses a problem using, e.g., the R package mvtnorm BG09.
We mention that while we have adopted the standard DM95 testing framework here, it would also be possible to follow, e.g., the fixed-smoothing approach to predictive ability testing by CI20. We omit details for brevity.
As a data-generating process (DGP) we consider a $d$-variate CCC--GARCH model Bol90. Such models have successfully replicated the (co-)movements of returns on speculative assets. We consider more recent models in this tradition in the empirical application of Section (ref). However, for the present purpose, we prefer a simpler model. Let $\mathcal{F}_{t-1}=\sigma(\bm Y_{t-1},\bm Y_{t-2},\ldots)$ denote the information set consisting of past observables. These are generated by the standard CCC--GARCH specification \[ \bm Y_t=(Y_{1t},\ldots,Y_{dt})^\prime=\bm \varSigma_t(\bm \theta^\circ)\bm \varepsilon_t,\qquad t=1,\ldots,n. \] Here, $\bm \varSigma_t(\bm \theta^\circ)=\operatorname{diag}(\bm \sigma_t)$ with $\bm \sigma_t=(\sigma_{1t},\ldots,\sigma_{dt})^\prime$ is the $\mathcal{F}_{t-1}$-measurable diagonal matrix containing the volatilities of $\bm Y_t$ (given $\mathcal{\mathcal{F}}_{t-1}$) with true parameter vector $\bm \theta^\circ$. The multiplicative $\mathbb{R}^d$-valued Gaussian noise $\bm \varepsilon_t$ is independent of $\mathcal{F}_{t-1}$ and independent, identically distributed (i.i.d.) with mean zero, unit variance and correlation matrix $\bm R$ (written $\bm \varepsilon_t\overset{\text{i.i.d.}}{\sim}\mathcal{N}(\bm 0,\bm R)$).
The individual volatilities follow GARCH(1,1) dynamics with identical parameters across $i$, such that $\sigma_{it}^2=\sigma_{it}^2(\bm \theta^\circ)=\omega^\circ+\alpha^\circ Y_{i,t-1}^2+\beta^\circ \sigma_{i,t-1}^2$ for $i=1,\ldots,d$ and $\bm \theta^\circ=(\omega^\circ, \alpha^\circ, \beta^\circ)^\prime$. For the parameters, we choose $\omega^\circ=0.001$, $\alpha^\circ=0.1$, $\beta^\circ=0.5$, and we let $\rho=0.5$ be the equicorrelation parameter of the correlation matrix $\bm R$. Hence, we assume that $\bm R$ has ones on the main diagonal and $\rho$ everywhere else. For the dimension, we set $d=5$. The sample sizes we consider are $n\in\{150,\ 300\}$. Our choices of $d$ and $n$ are inspired by the empirical application in the next section.
By construction $\bm Y_t\mid\mathcal{F}_{t-1}\sim\mathcal{N}\big(\bm 0,\bm \varSigma_t(\bm \theta^\circ)\bm R\bm \varSigma_t(\bm \theta^\circ)\big)$, and we denote the cdf of this normal distribution by $\Phi(\bm \theta^\circ,\rho;\mathcal{F}_{t-1})$. The notation emphasizes the dependence of $\Phi$ on $\mathcal{F}_{t-1}$ here to make clear that this conditional cdf depends on past observables $\bm Y_{t-1},\bm Y_{t-2},\ldots$. From the above, the copula of $\bm Y_t\mid\mathcal{F}_{t-1}$ is Gaussian with equicorrelation parameter $\rho$, and the conditional marginal distributions of $\bm Y_t$ are given by $Y_{it}\mid\mathcal{F}_{t-1}\sim\mathcal{N}\big(0,\sigma_{it}^2(\bm \theta^\circ)\big)$ for $i=1,\ldots,d$. Of course, the facts that $\bm Y_t\mid\mathcal{F}_{t-1}$ has a Gaussian copula with equicorrelation parameter $\rho$ and that $Y_{it}\mid\mathcal{F}_{t-1}\sim\mathcal{N}\big(0,\sigma_{it}^2(\bm \theta^\circ)\big)$ ($i=1,\ldots,d$) fully determine the joint distribution of $\bm Y_t\mid\mathcal{F}_{t-1}$. Crucially, our construction allows us to separately vary the quality of the marginal forecasts (via $\bm \theta^\circ$) and the copula forecasts (via $\rho$).
Specifically, we consider the two distributional “forecasts”
where $\delta^{x}_{it}$ ($i\in\{1,2\}$, $x\in\{\operatorname{marg}, \operatorname{cop}\}$) are mutually independent i.i.d. draws from a uniform $\mathcal{U}[1-\Delta_i^x,1+\Delta_i^x]$-distribution. Therefore, the positive $\delta^{x}_{it}$'s may be interpreted as multiplicative noises with $\operatorname{E}[\delta^x_{it}]=1$, which contaminate the true parameters at each point in time. More precisely, $\delta^{\operatorname{marg}}_{it}$ serves to disturb the forecasts of the marginals and $\delta^{\operatorname{cop}}_{it}$ those of the copula. The multiplicative disturbances may be interpreted as mimicking estimation error of the true parameters. However, use of $\delta^{\operatorname{marg}}_{it}$ and $\delta^{\operatorname{cop}}_{it}$ (instead of some actual parameter estimator) allows us to vary the quality of the marginal and copula forecasts independently of each other. Of course, this comes at the cost of imposing the unrealistic assumption that the true parameters are known. We discuss this feature of our setup below. Note that, as in a realistic forecasting setting, the forecasts $F_{it}$ for time $t$ only depend on past information captured by $\mathcal{F}_{t-1}$.
We consider the following five settings for the two competing forecasts $F_{1t}$ and $F_{2t}$, each characterized by a specific choice of $\Delta_1^{\operatorname{marg}}$, $\Delta_2^{\operatorname{marg}}$, $\Delta_1^{\operatorname{cop}}$ and $\Delta_2^{\operatorname{cop}}$:
For the size results in (i), both forecasts are equally contaminated (i.e., $\Delta_1^{\operatorname{marg}}=\Delta_2^{\operatorname{marg}}$ and $\Delta_1^{\operatorname{cop}}=\Delta_2^{\operatorname{cop}}$), implying that the null hypotheses $H_{0}^{=}$ and $H_{0}^{\preceq_{\mathrm{lex}}}$ hold. In the settings (ii)--(v), the second set of forecasts $F_{2t}$ is always less (or equally) contaminated than the first (i.e., $\Delta_{2}^x\leq\Delta_1^x$ for $x\in\{\operatorname{marg},\operatorname{cop}\}$). Therefore, (ii)--(v) all correspond to cases where the alternative to both $H_{0}^{=}$ and $H_{0}^{\preceq_{\mathrm{lex}}}$ is true. The setup in (iii) is analogous to that in (ii), except that both marginals are more seriously misspecified (but equally so).
We mention that it would be desirable for a more realistic setting to actually estimate the parameters of our CCC--GARCH model using, e.g., the QMLE FZ10. However, with a realistic parameter estimator, it would be very difficult to construct forecasts whose quality can be varied for the marginals and the copula separately. Therefore, we have opted for the above method in (ref) of “contaminating” the true parameters.
In applying our two-step tests from Theorem (ref) we use the log-score for $S_i$ ($i=1,\ldots,d$) and $S$ to compute our multi-objective score from Theorem (ref). Furthermore, we choose the critical values such that the probabilities on the left-hand side of (ref) and (ref) are each equal to $\alpha/2$ for $\alpha=0.05$. In computing $\widehat{\bm \varOmega}_n$ we follow the recommendation of DM95 and use $w_{n,h}=0$, such that $\widehat{\bm \varOmega}_n$ equals the sample variance of the score differences.
We draw the following conclusions from the results in Table (ref), which are based on 10,000 replications of the tests of $H_0^{=}$ and $H_{0}^{\preceq_{\mathrm{lex}}}$ for each of (i)--(v). The overall rejection frequencies in the column “Joint” indicate that size is close to the nominal level of $\alpha=5\%$ (row (i)) and power is high (rows (ii)--(v)). The columns “Marginal” and “Copula” decompose the total number of rejections (in column “Joint”) into rejections that occurred in the first step (for the marginal component) and the second step (for the copula component), respectively. By construction of our two-step test, these rejection frequencies should each equal $\alpha/2=2.5\%$ under the null. Table (ref) shows that this is roughly the case; see rows labeled (i).
Furthermore, we see that if the difference in the forecasts is only in the marginals (setup (iv)), then the null is rejected mainly in the first step. Similarly, if differences in predictive ability exist exclusively in the copula component, then our test identifies these mainly in the second step, as it should be; see the results for setups (ii)--(iii). We conclude that attribution using our tests works well.
It is also noteworthy that for setups (ii) and (iii), where the marginals are equally misspecified, the rejection probability in the first step roughly equals the expected $2.5\%$. Setting (iii), where the copula forecasts are as different as under (ii), shows that it becomes increasingly more difficult to distinguish between the copula forecasts, when the marginals are more seriously misspecified. The lesson to be learned is that, while the marginals are not formally tested in the second step, they do have an influence on the power of the copula test. This is because the copula score differences based on $S_{\operatorname{cop}}$ do implicitly depend on the marginal forecasts; see the definition of $S_{\operatorname{cop}}$ in Theorem (ref) and the discussion of Remark (ref).
We generally observe the highest power when there are differences in predictive ability of both marginals and copulas for forecasts (v). Moreover, we also see the expected increase in power for larger sample sizes $n$. Finally, the tests of $H_0^{\preceq_{\mathrm{lex}}}$ have higher power than those for $H_0^{=}$, since the deviation from the null is in the direction of the alternative of $H_0^{\preceq_{\mathrm{lex}}}$. This is also as expected for one (and a half)-sided tests.
We illustrate our two-step tests for a broad set of international stock market indices. Specifically, we consider copula forecasts for the joint daily log-returns $\bm Y_t=(Y_{1t},\ldots,Y_{dt})^\prime$ on the S&P 500, DAX, CAC 40, Hang Seng Index and Nikkei 225 (such that $d=5$). Data on the index returns were downloaded from \href{https://www.wsj.com/market-data/quotes}{www.wsj.com/market-data/quotes}. The dependence structure of these indices are of importance, e.g., in building internationally diversified portfolios. Our sample contains 7 years of daily data from 2016 to 2022. For simplicity, we only keep those observations where returns for all indices are available.
We compare the copula forecasts of seven models. The first two models are benchmark DCC--GARCH processes, one with normally distributed innovations ($\mathcal{N}$--DCC) and one with multivariate Student's $t$-distributed innovations ($t$--DCC). These models have been found to provide very accurate variance-covariance matrix forecasts LRV12,LRV13. The next two models we consider are CCC--GARCH processes, as already used in the simulations in Section (ref). Again, we use a normal distribution for the innovations ($\mathcal{N}$--CCC) and a $t$-distribution ($t$--CCC). As our fifth model, we employ the GO--GARCH of Van02, which has Gaussian innovations. The final two competitors are not from the class of multivariate GARCH models, but from the class of generalized autoregressive score (GAS) models CKL13,Har13. In one case, we assume returns to have a conditional normal distribution ($\mathcal{N}$--GAS), where the marginal scales vary over time (using a GAS update), yet the marginal locations and correlations are time-invariant.\footnote{We have also tried to fit this model with time-varying correlations, yet we only obtained nonsensical parameter estimates.} In the other case, we assume returns whose conditional distribution is multivariate $t$ ($t$--GAS). Here, the evolution of the univariate scales and the correlation of the multivariate $t$-distribution follow GAS dynamics with identity scaling matrix. The degrees of freedom parameter, in contrast, is left constant over time for parsimony. GAS models have become increasingly popular in time series modeling due to their flexibility and their strong empirical performance in forecasting applications. The webpage \href{https://www.gasmodel.com} {www.gasmodel.com}, in particular, provides some testimony to the increasing popularity of GAS models. We refer to Bea19a for a more complete description of all the above models as well as their implementation in R.
We use a fixed window of 6 years to estimate the models, for which we use the rmgarch package of rmgarch and the GAS package of ABC19. We keep the remaining 1 year (corresponding to $\bm Y_1,\ldots,\bm Y_n$ with $n=223$) for comparing the copula forecasts. Throughout, we denote the distributional forecasts by $F_{1t}$ and $F_{2t}$. From these, the copula forecasts $C_{1t}$ and $C_{2t}$ are derived as well as the respective copula densities $c_{1t}$ and $c_{2t}$. Likewise, we obtain the forecasts of the marginals $F_{1it}$ and $F_{2it}$ ($i=1,\ldots,d$) and the appertaining density forecasts $f_{1it}$ and $f_{2it}$, respectively.
To compare forecasts, we use the average score differences as defined in (ref). In doing so, we use the multi-objective score $\bm S$ from Theorem (ref) (Equation (ref)) with the log-scores for $S_i$ and $S$ (see Example (ref)), such that
with similar formulae for the second set of forecasts. To give an impression of the quality of the forecasts issued by the seven models, we display the average scores for the marginals $\overline{S}_{\operatorname{marg},n}=\frac{1}{n}\sum_{t=1}^{n}S_{\operatorname{marg}}(\{F_{1it}\}_{i=1,\ldots,d},\bm Y_{t})$ and the copula $\overline{S}_{\operatorname{cop},n}=\frac{1}{n}\sum_{t=1}^{n}S_{\operatorname{cop}}(C_{1t},\{F_{1it}\}_{i=1,\ldots,d},\bm Y_{t})$ in Table (ref). Recall that, everything else equal, a lower average score is preferable. The magnitude of the average marginal scores seems to be comparable for all models, except for the $\mathcal{N}$--GAS model that has a much higher score. For the copula component, the GARCH-type models all perform equally well, with the GAS models lagging equally far behind.
More formally, we now provide the results of our two-step tests, as detailed in Section (ref). We first carry out tests of $H_0^{=}$, i.e., the joint hypothesis of equal predictive ability in the marginals and the copula. Table (ref) shows the results for all possible pairwise comparisons. In that table, an “M” indicates a rejection of our two-step test already in the first step that compares the marginal forecasts (such that $\sqrt{n}|\overline{d}_{m,n}|>c_{1n}$ holds in the notation of Theorem (ref)). On the other hand, a “C” signals a rejection only in the second step analyzing the copula forecasts (i.e., $\sqrt{n}|\overline{d}_{c,n}|>c_{2n}$ but $\sqrt{n}|\overline{d}_{m,n}|\leq c_{1n}$).
Comparing the GAS with the GARCH models in Table (ref), we find significant differences in forecast quality already in the marginals for the $\mathcal{N}$--GAS model, but also in the second step test of the copula forecasts for the $t$--GAS. Between the CCC and the DCC models, we find non-rejections of the null and rejections of it either due to the marginals or the copulas. Despite the different predictive ability of the various CCC and DCC models, the differences in forecast quality between the GO--GARCH model and the other GARCH models are always insignificant.
Strictly speaking, a rejection of $H_0^{=}$ only allows us to conclude that forecasts are of different quality. Therefore, to assess which model should actually be preferred, we also carry out tests of $H_0^{\preceq_{\mathrm{lex}}}$. We do so for all pairwise comparisons in Table (ref). Specifically, in Table (ref) we test $H_0^{\preceq_{\mathrm{lex}}}$ with the models in the first column issuing the first set of forecasts $F_{1t}$ and the models in the first row corresponding to the second set of forecasts $F_{2t}$. In other words, the non-inferiority of the column-listed models is tested with respect to the row-listed models.
While this does not mechanically have to be the case, we find that the entries in the lower triangular part of Table (ref) are identical to those in Table (ref). However, Table (ref) provides a statistically sound indication of which copula forecasts are to be preferred (at least in those cases, where the marginal forecasts perform equally well). For instance, while we cannot infer anything about the copula forecasts of the $\mathcal{N}$--GAS relative to the other models (because the marginals are of significantly different predictive accuracy), we reject the non-inferiority of the copula forecasts of the $t$--GAS model compared with all CCC and DCC models. Therefore, the dependence structure predictions of the CCC/DCC model class significantly outperform those of the $t$--GAS.
Turning to the upper triangular part of Table (ref), we see that the entries with an “M” in the lower triangular part necessarily reappear (in transposed form) in the upper triangular part. This is because the first-step test of the marginals is two-sided. However, “C” entries may be different for the upper and lower triangular parts, because of the one-sided nature of the copula forecast accuracy test. For instance, while the $t$--GAS model is inferior in terms of its copula forecasts to the $\mathcal{N}$--DCC model (lower left entry of the table), the $\mathcal{N}$--DCC model is non-inferior to the $t$--GAS model (upper right entry of the table).
Finally, we zoom in on one specific comparison of two models, viz. the $t$--DCC and the $t$--GAS as two benchmarks in their respective model classes. Here, Table (ref) reveals that there is little difference for the marginals, while the differences in the copula forecasts lead to a rejection of $H_0^{=}$. In contrast, the hypothesis $H_0^{\preceq_{\mathrm{lex}}}$ (with $t$--DCC providing the first set of forecasts, and $t$--GAS the second set) cannot be rejected since the copula forecasts of the DCC--GARCH model are non-inferior to those of the GAS model. (For the marginals, it is clear that we again do not reject by construction of our two-step test.)
Figure (ref) illustrates the comparison of the $t$--DCC and the $t$--GAS by plotting the average score differences over time (i.e., $\frac{1}{T}\sum_{t=1}^{T}\bm d_t$ for $T=1,\ldots,n$). For the marginal scores in panel (a), there is evidently little difference. However, the copula scores in panel (b) are significantly negative---a sign of the superiority of the DCC--GARCH model when it comes to capturing the dependence dynamics over time.
We stress that such an attribution of the rejection to the marginals or the copula would not have been possible by inspecting only the differences in the standard log-scores of the predictive distributions $F_{1t}$ and $F_{2t}$ (panel (c)). Note that this log-score equals (see Equation (ref))
Therefore, the log-score is the sum of the marginal score $S_{\operatorname{marg}}(\{F_i\}_{i=1,\ldots,d},\bm y)$ and the copula score $S_{\operatorname{cop}}(C,\{F_i\}_{i=1,\ldots,d},\bm y)$.\footnote{Note that if we had chosen a different scoring rule as building blocks for $S_{\operatorname{marg}}$ and $S_{\operatorname{cop}}$, we would not get such an additive decomposition.} In particular, adding up the marginal and copula average score differences in panels (a) and (b) yields the average density score differences in panel (c). While a simple Diebold--Mariano test based on the above log-score in (ref) likewise reveals that $H_0^{=}$ can be rejected, the reason for the rejection remains unclear. This contrasts with our two-step test, which clearly attributes the rejection to the quality of the copula forecasts.
While copulas are widely used in predictive applications in economics and finance, surprisingly little is known about how to carry out a statistically sound forecast comparison. In this paper, we fill this gap by establishing the following three results. First, unless the marginal structure is given and fixed, copula forecasts cannot be compared on their own. Second, when the marginal structure is given and fixed, that is, on Fr\'echet classes, any strictly proper score can validly be used for comparing copula forecasts. Third, exploiting the existence of multi-objective scores for the copula and the marginals, we demonstrate how to compare these joint forecasts while still being able to attribute differences in the forecast accuracy to the quality of the marginal forecasts or the copula forecasts.
Our motivation to introduce multi-objective scores is quite different from that of FH24. Being interested in the comparison of systemic risk forecasts, FH24 need to overcome the non-elicitability of AB16's AB16 conditional Value-at-Risk (CoVaR)---a very popular systemic risk measure. Even the strategy of combining CoVaR with VaR as an auxiliary functional does not yield elicitability. This motivates FH24 to introduce multi-objective elicitability---a property that turns out to hold for the pair (CoVaR, VaR). In our case, the situation is quite different. The functional of interest, i.e., the copula, is jointly elicitable with the marginals (unlike the pair (CoVaR, VaR)). However, this property does not allow us to pinpoint performance differences specifically to the copula models. Therefore, the main advantage of multi-objective elicitability (vis-\`{a}-vis elicitability) in our context is that it allows for attributing differences in performance to either the marginals or the copula.
The non-elicitability result of copulas presented in this paper and the multi-objective elicitability together with the marginal distributions also has consequences for estimation. It demonstrates why it is generally necessary to first estimate the marginals $\hat F_1, \ldots, \hat F_d$ and then to infer the copula from the pseudo-sample from the copula, $\big\{\big(\hat F_1(Y_{1t}), \ldots, \hat F_d(Y_{dt})\big)'\big\}_{t=1, \ldots, n}$. The second step can be done with empirical score minimization of the score in (ref); see e.g., Chapter 7.5 in MFE15. This two-step procedure amounts to empirical score minimization of the multi-objective scores from Theorem (ref) with respect to the lexicographic order.
In light of the results of this paper and those in FH24, we expect there to be many more areas where the concept of multi-objective elicitability may be fruitfully applied. For instance, in a risk-management context, the pair consisting of VaR and Expected Shortfall (ES) is elicitable and therefore also multi-objective elicitable. While forecasts for the pair (VaR, ES) may be compared drawing on its elicitability, it is only the pair's multi-objective elicitability that allows for performance attribution to either the VaR or the ES component. Explorations such as these are left for future research.
\onehalfspacing