EconBase
← Back to paper

The Efficiency Gap

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

215,812 characters · 24 sections · 180 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The Efficiency Gap

abstractAbstract. Parameter estimation via M- and Z-estimation is equally powerful in semiparametric models for one-dimensional functionals due to a one-to-one relation between corresponding loss and identification functions via integration and differentiation. For multivariate functionals such as multiple moments, quantiles, or the pair (Value at Risk, Expected Shortfall), this one-to-one relation fails and not every identification function possesses an antiderivative. The most important implication is an efficiency gap: The most efficient Z-estimator often outperforms the most efficient M-estimator. We theoretically establish this phenomenon for multiple quantiles at different levels and for the pair (Value at Risk, Expected Shortfall), and illustrate the gap numerically. Our results further give guidance for pseudo-efficient M-estimation for semiparametric models of the Value at Risk and Expected Shortfall.

Keywords: Efficient semiparametric estimation; Expected Shortfall; M-estimation; Quantiles; Loss functions \\ JEL Codes: C14, C22, C32, C51, C58, G32

Introduction

\onehalfspacing

Given some real-valued response variable $Y_t$ and some $p$-dimensional vector of covariates $X_t$, one is often interested in modelling the effect of the covariates on the response variable through regression models. E.g., one might be interested in the average effect of economic and financial conditions as e.g.\ inflation on GDP growth. The classical mean regression technique captures the average effect by modelling the expectation of the conditional distribution of $Y_t$ given $X_t$, denoted by $F_{t}$. However, researchers are often interested in different properties of this conditional distribution, e.g., in low quantiles if attention is focused on downside risks of GDP growth as in Adrian2019. This can be facilitated through quantile regression Koenker1978, where one parametrically models the quantile of the conditional distribution $F_{t}$.

More generally, one is interested in a certain statistical functional $\Gamma$ of the conditional distribution $F_{t}$, where the functional maps a (conditional) distribution to a real-valued outcome. The functional of interest varies among disciplines: E.g., quantitative risk managers are specifically interested in models for risk measures such as conditional variances (volatility), quantiles (Value at Risk, VaR), expectiles and Expected Shortfall (ES) Bollerslev1986, Engle2004, Efron1991, Patton2019. Epidemiological forecasts, of particular importance due to the COVID-19 pandemic, often focus on prediction intervals, which commonly consist of two quantiles Bracher2021NatComm, Cramer2022.

It is common practice to model the functional as some parametric model $\Gamma(F_{t}) = m(X_t,\theta_0)$ for some unique parameter $\theta_0 \in \Theta \subseteq \mathbb{R}^q$. This specification is commonly referred to as semiparametric: Even though the model $m$ itself is parametric, it does not specify the full conditional distribution $F_{t}$, but only a functional $\Gamma$ thereof Newey1990, BKRW1998book.

While standard approaches often model every functional of interest separately, joint semiparametric models for multivariate (or vector-valued) functionals have desirable advantages in many instances: A joint treatment of two quantile levels is e.g.\ beneficial for prediction intervals Shrestha2006, Bracher2021NatComm, it can impede quantile crossings GourierouxJasiak2008, WhiteKimManganelli2015, Catania2019, and it generally improves efficiency. More fundamentally, there are cases where M- or Z-estimation of univariate models is infeasible such as for the variance, ES and Range Value at Risk (RVaR, also called “interquantile expectation”, which nests the trimmed mean), since suitable loss or identification functions for these functionals do not exist; see Osband1985, Weber2006, WangWei2020, DFZ_CharMest. However, such objective functions exist for an appropriate multivariate functional; see FisslerZiegel2016 for the pair (VaR, ES), Osband1985 for the pair (mean, variance), and FisslerZiegel2019 for the triplet of the RVaR with two quantiles. These examples motivate our consideration of joint estimation of multivariate models.

Estimation of the parameter $\theta_0$ in semiparametric models is regularly carried out by either minimum (M-) or zero (Z-) estimation NeweyMcFadden1994. Given these estimators are consistent and asymptotically normal, one favors an efficient estimator with an associated covariance matrix which is as small as possible. Besides more accurate estimates, this allows for more powerful inference through tests and confidence intervals.

In this article, we investigate the efficiency of M-estimators, based on some loss functions, in particular in relation to Z-estimators, which are based on identification functions (or moment conditions). We show the existence of an “efficiency gap” for multivariate functionals in the sense that the semiparametric M-estimator cannot attain the Z-estimation or semiparametric efficiency bound in the sense of Stein1956. For this, we make use of a recent result of DFZ_CharMest that fully characterizes the class of consistent, semiparametric M-estimators for general functionals through the classes of strictly consistent loss functions from the literature on forecast evaluation Gneiting2011, FisslerZiegel2016. For vector-valued functionals, theses latter classes are considerably smaller than the corresponding classes of identification functions used in Z-estimation. This is in stark contrast to the univariate case, where these classes are almost equivalent and M- and Z-estimation can be equally efficient. As a stepping stone, we derive the novel result that the “optimal instrument matrix" of Chamberlain1987 and Newey1993 is not only a sufficient, but also a necessary condition for efficient Z-estimation.

Throughout the article, we recurrently make use of the running example of a double quantile model---i.e., a semiparametric model for two quantiles at different levels---to illustrate and exemplify our general theoretical results. In particular, we derive conditions for the occurrence of the efficiency gap and illustrate these in simulations. Our results directly generalize to finitely many quantiles. This model class arises naturally in the following fields of applications: In quantitative risk management, one is interested in quantiles (VaR) of financial returns at two small probability levels, say 1% and 2.5%, which directly motivates the joint modelling of two quantiles Engle2004, WhiteKimManganelli2015, Catania2019. Furthermore, prediction intervals can naturally be defined as the interval spanned by two (conditional) quantiles with levels of e.g., $5\%$ and $95\%$ BrehmerGneiting2020, FisslerHlavinovaRudloff2019Theory, Bracher2021NatComm. Eventually, the entire conditional distribution can conveniently be approximated by multiple conditional quantiles; see e.g., Buchinsky1994, Angrist2006, Chernozhukov2010 for microeconomic and Adrian2019 for macroeconomic applications. While models for individual quantile levels could be estimated separately, an important methodological demand on reasonable models is to impede quantile crossings Koenker2005, which can be achieved through joint models as in GourierouxJasiak2008, WhiteKimManganelli2015 and Catania2019. Moreover, joint estimation generally improves efficiency.

We further illustrate that the efficiency gap arises for the popular and recently proposed joint models for the VaR and ES Patton2019. While the as yet common choices of loss functions used for their M-estimation are rather ad hoc, we provide two novel pseudo-efficient loss functions, that is, choices which result in efficient M-estimation at least in specific (but realistic) situations. We illustrate their superiority in simulations, especially for small probability levels that are of particular importance in risk management. The first pseudo-efficient choice is “surprisingly feasible” in the sense that it requires very little pre-estimates compared to classical semiparametric models for the mean or quantiles. This finding suggests an improved and practically relevant M-estimator for semiparametric VaR and ES models. We anticipate that the efficiency gap generalizes to joint models for various other vector-valued functionals like multiple expectiles or the RVaR, jointly with corresponding quantiles.

The paper is organized as follows. Section (ref) formally introduces M- and Z-estimation and relates these to the literature on forecast evaluation that we quickly review in Section (ref). Section (ref) considers efficient M- and Z-estimation of general semiparametric models and attainability of the semiparametric efficiency bound. In Section (ref), we establish the efficiency gap for double quantile models and models for (VaR, ES), which is illustrated in simulations in Section (ref). The Supplementary Material contains all proofs in Section (ref), analyzes efficient estimation of the pair (mean, variance) in Section (ref), discusses the impact of the gap on equivariant estimation in Section (ref), and contains further technical details in its subsequent sections.

M- and Z-estimation

We consider a time series $Z_t = (Y_t,X_t)$, $t\in\mathbb N$, where $Y_t$ are real-valued response variables and $X_t$ are $\mathbb{R}^p$-valued regressors, that can potentially contain lagged values of $Y_t$, allowing for autoregressive models. Let $\mathcal{F}_\mathcal{Z}$ be a class of possible joint distributions of $Z_t$ that formalizes the uncertainty about the distribution of our time series. $\mathcal{F}_\mathcal{Z}$ induces a class $\mathcal{F}_\mathcal{X}$ of marginal distributions of $X_t$ and a class $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ of conditional distributions, $F_t$, of $Y_t$ given $X_t$. Whenever they exist, we denote the conditional density by $f_t$, the conditional expectation by $\mathbb{E}_t[\cdot] = \mathbb{E}[\cdot \,|\, X_t]$ and the conditional variance by $\operatorname{Var}_t(\cdot) = \operatorname{Var}(\cdot \,|\, X_t)$. Equalities of random variables are meant to hold almost surely if not stated otherwise.

Let $\Gamma \colon\mathcal{F}_{\mathcal{Y}|\mathcal{X}}\to\Xi\subseteq\mathbb{R}^k$ be some $k$-dimensional and measurable functional of the conditional distributions $F_t$. Standard examples for univariate functionals are the mean or quantiles. Later on, we consider a pair of two quantiles and the pair consisting of the VaR and ES as examples for multivariate functionals. Let $\Theta\subseteq \mathbb{R}^q$ be a parameter space with non-empty interior, $\operatorname{int}(\Theta)$, and $m\colon \mathbb{R}^p\times \Theta\to \Xi$ a parametric and (in $\theta$) differentiable model for the functional $\Gamma$. We denote the gradient of its $j$-th component by the column vector $\nabla_{\theta} m_j(X_t,\theta)\in\mathbb{R}^q$, $j= 1,\dots,k$. We work under the following assumption of a correctly specified model with a unique parameter.

assFor all distributions $F_{Z_t} \in \mathcal{F}_\mathcal{Z}$ of $Z_t=(Y_t, X_t)$, there is a unique and time-independent parameter $\theta_0 = \theta_0(F_{Z_t}) \in \operatorname{int}(\Theta)$ such that $m(X_t,\theta_0) = \Gamma(F_t)$ for all $t\in\mathbb N$.

We dispense with a strong stationarity assumption on the time series $Z_t$, however, Assumption (ref) imposes a semiparametric stationarity assumption in that the parameter $\theta_0$, and hence the functional $\Gamma(F_t)$ is time-independent, allowing e.g., for heteroskedasticity.

Following Huber1967 and NeweyMcFadden1994, we consider M-estimators for $\theta_0$

equation[equation omitted — 164 chars of source]

based on possibly time-varying loss functions $\rho_t$, which are the key ingredient of the M-estimator. The core condition on $\rho_t$ for the consistency of $\widehat \theta_{M,T}$ is that

align[align omitted — 230 chars of source]

which we call strict $\mathcal{F}_{\mathcal{Z}}$-model-consistency of $\rho_t$ for $m$ as in DFZ_CharMest; also see Gourieroux1987.

A standard alternative to M-estimation are zero (Z-) or method of moments (MM-) estimators Hansen1982, NeweyMcFadden1994, given by

align[align omitted — 168 chars of source]

The name arises since the minimization in (ref) essentially sets the average of the $q$-dimensional, possibly time varying functions $\psi_t$ to zero. Hence, consistency of the Z-estimator crucially relies on the strict unconditional $\mathcal{F}_{\mathcal{Z}}$-identification condition

align[align omitted — 219 chars of source]

The functions $\psi_t$ in (ref) are often the gradients of the losses $\rho_t$ in (ref). We do not consider the standard extension to generalized method of moments (GMM) estimation, where $\psi_t$ can be of larger dimension than $\theta$, as the exactly identified case in (ref) suffices for efficient estimation; see Theorem (ref) and Remark (ref) for details.

For semiparametric estimation, there exist a multitude of choices for the functions $\rho_t$ and $\psi_t$ that satisfy the conditions (ref) and (ref) respectively GourierouxMonfortTrognon1984, Komunjer2005. This opens up the possibilities to optimally choose $\rho_t$ and $\psi_t$, e.g., for efficient estimation Newey1993. To characterize such bounds, it is essential to characterize the entire classes of functions $\rho_t$ and $\psi_t$ such that (ref) and (ref) hold. For this, DFZ_CharMest formally connect these conditions to the notions of strictly consistent loss and strict identification functions from the literature on forecast evaluation Gneiting2011, which we shortly review in the following.

Strictly consistent loss and strict identification functions

Throughout this section, let $Y \sim F \in \mathcal{F}$ be a real-valued random variable, where $\mathcal{F}$ is a generic class of probability distributions on $\mathbb{R}$. We consider the single-valued functional $\Gamma \colon \mathcal{F} \to\Xi$ that attains values in the $k$-dimensional action domain $\Xi\subseteq \mathbb{R}^k$. A map $a\colon\mathbb{R}\times \Xi \to\mathbb{R}^\ell$, $\ell \in \mathbb{N}$, is called $\mathcal{F}$-integrable if $\mathbb{E}[a(Y, \xi)]$ exists and is finite for all $Y \sim F\in\mathcal{F}$ and for all $\xi\in\Xi$.

definition[Consistency and elicitability] An $\mathcal{F}$-integrable map $\rho\colon \mathbb{R}\times \Xi\to\mathbb{R}$ is an $\mathcal{F}$-consistent loss function for a functional $\Gamma\colon \mathcal{F}\to\Xi$ if \begin{equation} \mathbb{E} \big[ \rho \big(Y,\Gamma(F)\big)\big] \le \mathbb{E} \big[ \rho (Y, \xi) \big] \qquad for all Y \sim F\in\mathcal{F}, \ for all \xi\in\Xi\,. \end{equation} If equality in (ref) implies that $\xi = \Gamma(F)$, then $\rho$ is called strictly $\mathcal{F}$-consistent for $\Gamma$. A functional $\Gamma$ is elicitable on $\mathcal{F}$ if there is a strictly $\mathcal{F}$-consistent loss function for it.

The crucial difference to unconditional model consistency in (ref) is that in (ref), the expectation is only taken with respect to $Y$. The whole classes of (strictly) consistent losses are characterized for many functionals Gneiting2011, FisslerZiegel2016. E.g., under richness conditions on the class $\mathcal{F}$, one can show that $\rho$ is (strictly) $\mathcal{F}$-consistent for the mean functional if and only if it is a Bregman loss $\rho(y,\xi) = \phi(y)- \phi(\xi) +\phi'(\xi)(\xi-y) + \kappa(y)$ where $\phi$ is a (strictly) convex function on $\mathbb{R}$ with subgradient $\phi'$, and the function $\kappa\colon\mathbb{R}\to\mathbb{R}$ is such that $\mathcal{F}$-integrability holds Savage1971, Gneiting2011. This class nests the omnipresent squared loss $\rho(y,\xi) = (y-\xi)^2$. Likewise, under similar richness conditions and if $\mathcal{F}$ contains only distributions with a unique $\alpha$-quantile, a loss is strictly $\mathcal{F}$-consistent for the $\alpha$-quantile with $\alpha \in (0,1)$, if and only if $\rho$ is a generalized piecewise linear loss functions $\rho(y,\xi) = (\mathds{1}_{\{y\le \xi\}} - \alpha )(g(\xi)-g(y)) + \kappa(y)$, where $g$ is (strictly) increasing, and $\kappa$ a function ensuring $\mathcal{F}$-integrability Gneiting2011b. This class nests the well known pinball loss $\rho(y,\xi) = (\mathds{1}_{\{y\le \xi\}} - \alpha)(\xi - y)$. The following running example is used illustratively throughout the paper.

RunningExmpConsider the double quantile $\Gamma(\cdot) = \big(Q_\alpha(\cdot), Q_\beta(\cdot)\big)$ at probability levels $0<\alpha<\beta<1$ and a class $\mathcal{F}$ of strictly increasing distribution functions fulfilling the richness Assumption (ref) in Appendix (ref). FisslerZiegel2016 characterizes the class of (strictly) $\mathcal{F}$-consistent losses $\rho\colon\mathbb{R}\times\Xi\to\mathbb{R}$, $\Xi\subseteq \mathbb{R}^2$, for $\Gamma$ as \begin{align} \begin{aligned} \rho(y, \xi_1, \xi_2) = &\big( \mathds{1}_{\{y \le \xi_1\}} -\alpha \big) \big(g_{1}(\xi_1) - g_{1}(y) \big) \\ + \, &\big(\mathds{1}_{\{y \le \xi_2\}} - \beta \big) \big(g_{2}(\xi_2) - g_{2}(y) \big) + \kappa(y), \end{aligned} \end{align} where $g_{1}, g_2: \mathbb{R} \to \mathbb{R}$ are (strictly) increasing, and $\kappa\colon\mathbb{R}\to\mathbb{R}$ is such that $\rho$ is $\mathcal{F}$-integrable. Strikingly, this means that the whole class of (strictly) consistent losses for the double quantile coincides with the sum of (strictly) consistent losses for the individual quantiles.

In forecast evaluation, identification functions are used to check (conditional) calibration of forecasts NoldeZiegel2017, DimiPattonSchmidt2019, akin to goodness-of-fit tests.

definition[Identification function and identifiability] An $\mathcal{F}$-integrable map $\varphi\colon \mathbb{R} \times \Xi\to\mathbb{R}^k$ is an $\mathcal{F}$-identification function for a functional $\Gamma\colon \mathcal{F}\to\Xi \subseteq \mathbb{R}^k$ if \( \mathbb{E}\big[ \varphi\big(Y,\Gamma(F)\big) \big] =0 \) for all $Y \sim F \in \mathcal{F}$. If additionally $ \mathbb{E} \big[ \varphi(Y,\xi) \big] =0$ implies that $\xi = \Gamma(F) $ for all $F\in\mathcal{F}$ and for all $\xi\in\Xi$, it is a strict $\mathcal{F}$-identification function for $\Gamma$. A functional $\Gamma$ is called identifiable on $\mathcal{F}$ if there is a strict $\mathcal{F}$-identification function for it.

Given a strict $\mathcal{F}$-identification function $\varphi\colon\mathbb{R}\times\Xi\to\mathbb{R}^k$ for a functional $\Gamma\colon\mathcal{F}\to\Xi\subseteq \mathbb{R}^k$, DFZ_OsbandID shows that under some regularity conditions and richness conditions on $\mathcal{F}$, the full class of strict $\mathcal{F}$-identification functions for $\Gamma$ is given by

align[align omitted — 182 chars of source]

This characterization result is valid for any identifiable functional. In contrast, there is no such general characterization result available for the class of strictly consistent loss functions for a given elicitable functional. They need to be established on a case-by-case basis.

RunningExmpLet $\mathcal{F}$ be the class of continuous and strictly increasing distribution functions. The double quantile functional possesses a strict $\mathcal{F}$-identification function $\varphi(y,\xi_1,\xi_2) = \big( \mathds{1}_{\{y \le \xi_1 \}} - \alpha, \; \mathds{1}_{\{y \le \xi_2 \}} - \beta \big)^\intercal$. Equation (ref) provides a rich family of further strict $\mathcal{F}$-identification functions, e.g., choosing $h(\xi_1,\xi_2)= \big(\begin{smallmatrix} 1 & 1\\ 0 & 1 \end{smallmatrix}\big)$ leads to $\varphi'(y,\xi_1,\xi_2) = h(\xi_1,\xi_2) \varphi(y,\xi_1,\xi_2) = \big( \mathds{1}_{\{y \le \xi_1 \}} - \alpha + \mathds{1}_{\{y \le \xi_2 \}} - \beta, \; \mathds{1}_{\{y \le \xi_2 \}} - \beta \big)^\intercal$.

There is an intimate relationship between (strictly) consistent loss functions and strict identification functions for $\Gamma$ via differentiation and integration. For one-dimensional functionals $\Gamma$, these two classes are essentially equivalent: On the one hand, under sufficient smoothness and regularity conditions, first-order conditions yield that the derivative of any (strictly) consistent loss for $\Gamma$ is an identification function, whose strictness however requires some additional care. On the other hand, Osband's principle Osband1985, Gneiting2011 implies that---under sufficient regularity conditions---if $\varphi$ is an oriented identification function for $\Gamma$, then for any consistent loss $\rho$ there is a real-valued function $h$ such that

align[align omitted — 211 chars of source]

The relation between loss and identification functions is more involved for multivariate functionals $\Gamma$, and it turns out that there are considerably more identification functions than consistent losses. This disparity proves to be consequential for efficient estimation of semiparametric models for vector-valued functionals, as discussed in the subsequent sections of this article.

In more detail, the gradient of any (strictly) consistent loss is still a (multivariate) identification function for $\Gamma$. For the reverse direction, (ref) holds equivalently with $h$ being $(k\times k)$-matrix valued. However, $h(\xi) \mathbb{E} [ \varphi(Y,\xi) ]$ can only have an antiderivative if the Hessian $\nabla_\xi^2 \mathbb{E} [ \rho(Y,\xi) ]$ is symmetric; see FisslerZiegel2016 for a rigorous statement. This result imposes strong conditions on $h$ as illustrated with the following running example.

RunningExmpFisslerZiegel2016 yields that the derivative of any expected (strictly) $\mathcal{F}$-consistent loss function for the double quantile takes the form $h(\xi_1,\xi_2) \mathbb{E} \big[ \varphi(Y,\xi_1,\xi_2) \big]$ where $h(\xi_1,\xi_2) = \text{diag}\big(w_1(\xi_1), w_2 (\xi_2)\big)$ and $w_1$, $w_2$ are non-negative, subject to the richness Assumption (ref). This constitutes the argument for the characterization of all (strictly) consistent loss functions in (ref) where clearly $w_j = g_j'$, $j=1,2$. On the other hand, there is evidently a considerably larger class of $\mathbb{R}^{2\times 2}$-valued functions $h$ such that $\det(h(\xi_1, \xi_2))\neq 0$ for all $(\xi_1, \xi_2)\in\Xi$. E.g., $\varphi'$ in Running Example (ref) cannot arise as the derivative of a strictly consistent loss for the double quantile functional as the corresponding $h=\big(\begin{smallmatrix} 1 & 1\\ 0 & 1 \end{smallmatrix}\big)$ is not diagonal.

We refer to Supplement Section (ref) for further remarks and technical details on the connection between loss and identification functions.

Efficient Semiparametric Estimation

Recall the M-estimator at (ref), where the losses $\rho_t$ have to satisfy (ref), which closely resembles the notion of strict consistency in Section (ref). DFZ_CharMest shows that these two conditions are equivalent under Assumptions (ref) and (ref), i.e., a semiparametric M-estimator is consistent if and only if a (strictly) consistent loss function is used.

RunningExmpLet $m(X_t,\theta) = \big( q_\alpha(X_t,\theta), q_\beta(X_t,\theta) \big)^\intercal$ be some semiparametric model for the double quantile functional where $\theta \in \Theta \subseteq \mathbb{R}^q$. Then, DFZ_CharMest yields that under our Assumptions (ref) and (ref), a loss $\rho\colon \mathbb{R}\times\Xi\to\mathbb{R}$, $\Xi\subseteq \mathbb{R}^2$, is $\mathcal{F}_\mathcal{Z}$-model-consistent for $m$ if and only if $\rho$ is of the form given in (ref). This implies that the M-estimator for the double quantile model can only be consistent if $\rho$ is of the form given in (ref).

Such characterization results for the full class of consistent M-estimators allow to determine an asymptotically most efficient M-estimator. To this end, it is helpful to relate the asymptotic distributions of M- and Z-estimators, which coincide if the identification functions $\psi_t$ of the latter match the derivative with respect to $\theta$ of the loss $\rho_t$ of the former; see e.g., Theorems 3.1, 3.2, and the discussion on p.\ 2145 in NeweyMcFadden1994 for details. For non-differentiable losses, this rationale holds on the level of the differentiable conditional expectations NeweyMcFadden1994. Consequently, in the sequel we say that an M-estimator has an equivalent Z-estimator if the derivative of (the conditional expectation of) the loss function with respect to $\theta$ equals (the conditional expectation of) the identification function almost surely. Also notice that the asymptotic covariance of M-estimators is invariant to rescaling by constants $c$ and additions of terms $\kappa_t(Y_t)$ with the consequence that we dispense with a discussion of these terms in the sequel.

Following Chamberlain1987, Gourieroux1987, Newey1990 among many others, we consider functions $\psi_t$ in ((ref)) based on conditional moment conditions of the form

align[align omitted — 163 chars of source]

where $\varphi$ is a strict identification function for the functional $\Gamma$ and the $q \times k$ matrices $A_t(X_t, \theta)$ are often called instrument matrices. We denote their sequence by $\mathbb{A} = (A_t)_{t\in\mathbb N}$ and the resulting Z-estimator at (ref) by $\widehat \theta_{Z,T,\mathbb{A}}$. This restriction is justified by three reasons: First, moment conditions of the form (ref) generally suffice to reach the semiparametric efficiency bound Chamberlain1987. Second, the derivatives of strictly consistent loss functions take that form, where $A_t(X_t, \theta)$ matches the model gradient. Third, despite the convenient result (ref), a characterization of all consistent Z-estimators in terms of their functions $\psi_t$ is not available; see e.g., Roehrig1988, Komunjer2012 and the supplement of DFZ_CharMest.

Henceforth, we assume that the considered M- and Z-estimators are consistent and asymptotically normal. Primitive conditions for this are widely available, see e.g., Huber1967, Weiss1991, NeweyMcFadden1994, Andrews1994, davidson1994stochastic. These conditions include classical moment and dependence conditions on the process $(Y_t, X_t)_{t\in\mathbb N}$ together with smoothness assumptions on the conditional expectations of the employed loss and identification functions, and crucially, an identification condition for the model parameters. For M-estimators, this identification condition is conveniently fulfilled through DFZ_CharMest by employing strictly consistent loss functions. However, the analogue condition for the Z-estimator that $\psi_t$ are strict $\mathcal{F}_\mathcal{Z}$-identification functions for $\theta_0$ is more difficult to establish and generally has to be verified on a case-by-case basis; see e.g., Section (ref) for specific results for our running example of the double quantile models.

Henceforth, we impose the following assumption that ensures the specific form of the matrix $\Sigma_{T,\mathbb{A}}$ given in ((ref)), in particular the absence of “HAC” terms NeweyWest1987.

assSuppose that the sequence $\big(\psi_{t} (Y_t, X_t, \theta_0) \big)_{t\in\mathbb N}$ is uncorrelated.

Under the above conditions on the Z-estimator $\widehat \theta_{Z,T, \mathbb{A}}$, it holds that

align[align omitted — 224 chars of source]

where the asymptotic covariance is governed by the terms

align[align omitted — 426 chars of source]

where, for any $\theta\in\Theta$,

align[align omitted — 379 chars of source]

We say that an asymptotically normal estimator is efficient if there is no other asymptotically normal estimator with a smaller covariance matrix in the Loewner order $\succcurlyeq$. For two positive semi-definite matrices $A$ and $B$, we say that $A \succcurlyeq B$ if and only if $A - B$ is positive semi-definite. Motivated by the discussion in Newey1990, we deliberately omit an analysis of “superefficient” estimators. The following theorem establishes necessary and sufficient conditions for efficient Z-estimation by extending the theory of Hansen1985, Chamberlain1987 and Newey1993. Notice that the theorem also holds in the case $Y_t \in \mathbb{R}^d$, $d > 1$.

theoremUnder Assumptions (ref) and (ref), let $\varphi$ be a strict $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-identification function for $\Gamma$. Let $\widehat{\theta}_{Z,T, \mathbb{A}^{\ast}}$ be the Z-estimator at (ref) that is asymptotically normal and based on the strict unconditional $\mathcal{F}_{\mathcal{Z}}$-identification function at (ref) with instrument matrices $A_{t,C}^\ast(X_t,\theta)$ such that \begin{align} A_{t,C}^\ast(X_t,\theta_0) = C D_t(X_t,\theta_0)^\intercal S_t(X_t,\theta_0)^{-1} \qquad for all t \in \mathbb{N}, \end{align} where $S_t(X_t,\theta_0)$ and $D_t(X_t,\theta_0)$ are given at (ref) and (ref), assuming that $S_t(X_t,\theta_0)$ is invertible, and $C$ is any deterministic and invertible $q\times q$ matrix. Then: \begin{enumerate}[label=\normalfont (\roman*)] • The asymptotic covariance matrix of the Z-estimator $\widehat{\theta}_{Z,T, \mathbb{A}^{\ast}}$ is the limit (for $T \to \infty$) of \begin{align} \Lambda_T^{-1} := \left( \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ D_t(X_t,\theta_0)^\intercal S_t(X_t,\theta_0)^{-1} D_t(X_t,\theta_0) \right] \right)^{-1}. \end{align} • For any sequence of instrument matrices $\mathbb{A} = (A_t)_{t\in\mathbb N}$, and $\Delta_{T,\mathbb{A}}$, $\Sigma_{T,\mathbb{A}}$ as given at (ref) and (ref), it holds that $\Delta_{T,\mathbb{A}}^{-1} \Sigma_{T,\mathbb{A}} \Delta_{T,\mathbb{A}}^{-1} \succcurlyeq \Lambda_T^{-1}$ for all $T\ge1$. • If for some $t\in\{1, \ldots, T\}$ and for any non-singular and deterministic matrix $C$ it holds that $\mathbb P\big(A_t(X_t,\theta_0) \not= A_{t,C}^\ast(X_t,\theta_0) \big)>0$, then $\Delta_{T,\mathbb{A}}^{-1} \Sigma_{T,\mathbb{A}} \Delta_{T,\mathbb{A}}^{-1} \succcurlyeq \Lambda_T^{-1}$ and $\Delta_{T,\mathbb{A}}^{-1} \Sigma_{T,\mathbb{A}} \Delta_{T,\mathbb{A}}^{-1} \not= \Lambda_T^{-1}$. \end{enumerate}

Parts (ref) and (ref) of Theorem (ref) are direct time series generalizations of the efficiency result of Hansen1985, Chamberlain1987, and Newey1993. Together, they state that $\Lambda_T^{-1}$ is an asymptotic efficiency bound for the general Z-estimator for semiparametric models and that the Z-estimator based on the choice $A_{t,C}^\ast(X_t,\theta)$ for all $t\in\mathbb N$ which fulfills ((ref)) attains this efficiency bound, and is consequently an efficient Z-estimator. Thus, parts (ref) and (ref) of Theorem (ref) can be understood as a sufficient condition for efficient semiparametric Z-estimation.

Conversely, part (ref) can be interpreted as a necessary condition for efficient estimation and is novel to the literature. It states that efficient semiparametric estimation can only be carried out by choosing instrument matrices satisfying ((ref)) almost surely. Otherwise, there is some $v\in\mathbb{R}^q$ such that the asymptotic variance of the linear combination $v^\intercal \widehat{\theta}_{Z,T, \mathbb{A}}$ is larger than the asymptotic variance of $v^\intercal \widehat{\theta}_{Z,T, \mathbb{A}^{\ast}}$. This necessary condition for efficient estimation is crucial for the following sections where we show that for certain functionals, the M-estimator of semiparametric models cannot attain the Z-estimation efficiency bound and consequently neither the semiparametric efficiency bound in the sense of Stein1956, which is further discussed in Section (ref)

RunningExmpFor the double quantile model based on the identification function $\varphi$ from Running Example (ref), Theorem (ref) implies that efficient Z-estimation is based on the efficient instrument matrix $A^\ast_{t,C}(X_t,\theta_0) = C D_t(X_t,\theta_0)^\intercal S_t(X_t,\theta_0)^{-1}$, where $C$ is some deterministic and nonsingular matrix and where \begin{equation} S_t(X_t,\theta_0) = \begin{pmatrix} \alpha (1-\alpha) & \alpha (1-\beta) \\ \alpha (1-\beta) & \beta (1-\beta) \end{pmatrix}, \qquad D_t(X_t,\theta_0) = \begin{pmatrix} f_t(q_\alpha(X_t, \theta_0)) \nabla_{\theta} q_\alpha(X_t,\theta_0)^\intercal \\ f_t(q_\beta(X_t, \theta_0)) \nabla_{\theta} q_\beta(X_t,\theta_0)^\intercal \end{pmatrix}. \end{equation} The asymmetric roles of $\alpha$ and $\beta$ in $S_t(X_t,\theta_0) $ stem from the convention that w.l.o.g.\ $\alpha<\beta$.
remThe efficient instrument matrix $A_{t,C}^\ast(X_t,\theta_0)$ in Theorem (ref) depends on the specific choice of an identification function $\varphi$ However, invoking the characterization result (ref), if we used a different identification function $\varphi'\big(Y,m(X_t, \theta)\big) = h\big(m(X_t, \theta)\big)\varphi\big(Y,m(X_t, \theta)\big)$ in (ref), where $h$ has full rank, the resulting matrices $S_t(X_t,\theta)$ and $D_t(X_t,\theta)$ in (ref), (ref) would change, but the induced conditional moment conditions (ref) would remain unchanged. Hence, the efficiency bound $\Lambda_T^{-1}$ is invariant to the choice of $\varphi$, and Theorem (ref) can be interpreted as global, $\varphi$-independent, necessary and sufficient conditions for efficiency.
remWhile overidentified GMM-estimation \begin{align} \widehat \theta_{GMM,T} = \operatorname*{arg\,min}_{\theta \in\Theta} \left(\frac{1}{T} \sum_{t=1}^T \psi_t (Y_t , X_t,\theta) \right)^\intercal W_T \left(\frac{1}{T} \sum_{t=1}^T \psi_t (Y_t , X_t,\theta) \right), \end{align} with $s$-dimensional ($s > q$) functions $\psi_t$ and a positive definite weighting matrix $W_T$ can generally improve efficiency compared to Z-estimation Hansen1982, HallBook2005, when employing the efficient instrument choice in (ref), there is no additional efficiency gain through using overidentifying moment restrictions. For details, see e.g., Newey1993, and notice that the proofs of Theorem (ref), (ref) and (ref) work identically when including overidentifying moment restrictions together with a weighting matrix $W_T$ as in (ref). Consequently, we restrict attention to efficient instrument Z-estimation in Theorem (ref).

Semiparametric Models for Vector-Valued Functionals

Semiparametric Double Quantile Models

Consider the double quantile model $m(X_t,\theta) = \big( q_\alpha(X_t,\theta), \, q_\beta(X_t,\theta) \big)^\intercal$ at fixed levels $0 < \alpha < \beta < 1$ from our Running Examples (ref)--(ref), whose importance is motivated in the Introduction. The results of this section hold equivalently for multiple quantiles at different levels. Let $\widehat \theta_{Z,T,\mathbb{A}}$ be the Z-estimator defined via (ref) and (ref) based on some sequence of instrument matrices $\mathbb{A}$ and the strict $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-identification function $\varphi(y,\xi_1,\xi_2) = \big( \mathds{1}_{\{y \le \xi_1 \}} - \alpha, \; \mathds{1}_{\{y \le \xi_2 \}} - \beta \big)^\intercal$ assuming that all distributions in $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ are differentiable at their $\alpha$- and $\beta$-quantiles with strictly positive derivatives. Recall from Remark (ref) that the initial choice of $\varphi$ is irrelevant. The exact form (for this choice of $\varphi$) of the efficient instrument matrix $A^\ast_{t,C}(X_t,\theta_0) = C D_t(X_t,\theta_0)^\intercal S_t(X_t,\theta_0)^{-1}$ is given in Running Example (ref).

Under Assumptions (ref) and (ref) in Appendix (ref), DFZ_CharMest yields that the full class of consistent M-estimators at (ref) is given by the class of (strictly) $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent loss functions for two quantiles in (ref). For any sequence $G= (g_{1,t},g_{2,t})_{t\in\mathbb N}$ of such functions, we denote the corresponding M-estimators defined via (ref) by $\widehat \theta_{M,T,G}$.

We assume that the M- and Z-estimators are consistent and asymptotically normal. Primitive conditions for this are discussed before Assumption (ref). Strict unconditional $\mathcal{F}_\mathcal{Z}$-model consistency of $\rho_t$ for the M-estimator is guaranteed by Theorem DFZ_CharMest for the strictly consistent losses at (ref). For the strict unconditional identification of the Z-estimator, we refer to Proposition (ref) in Section (ref) which shows strict identification for the efficient Z-estimator in linear models. While generalizations of these conditions are desirable, their derivation is known to be “quite difficult” NeweyMcFadden1994.

The following theorem establishes that, under certain conditions, the M-estimator of the double quantile model is subject to the efficiency gap, i.e., it cannot attain the Z-estimation efficiency bound, and consequently neither the semiparametric efficiency bound.

theoremSuppose that Assumptions (ref), (ref) together with Assumptions (ref) and (ref) in Appendix (ref) hold for the double quantile model at levels $0<\alpha<\beta<1$, $\widehat \theta_{M,T,G}$ is asymptotically normal and the following further regularity conditions hold: \begin{enumerate}[label=\normalfont {(DQ\arabic*)}, leftmargin=*] • The parameters of the individual models are separated, $m(X_t,\theta) = \big( q_\alpha(X_t,\theta^\alpha), q_\beta(X_t,\theta^\beta) \big)^\intercal$, where $\theta = \big( \theta^\alpha, \theta^\beta \big) \in \operatorname{int}{\Theta} \subseteq \mathbb{R}^q$, with $\theta^\alpha \in \mathbb{R}^{q_1}$ and $\theta^\beta \in \mathbb{R}^{q_2}$ and $q_1 + q_2 = q$. • For all $t\in\mathbb N$, and for all $A\in \mathcal A$ with $\mathbb P(A)=1$ there are $q_1+1$ mutually different $v_1, \ldots, v_{q_1+1} \in \big\{\nabla_{\theta^\alpha} q_\alpha(X_t(\omega),\theta_0^\alpha) \in \mathbb{R}^{q_1} \colon \omega\in A\big\}\subseteq \mathbb{R}^{q_1}$, such that any subset of cardinality $q_1$ of $\{ v_1,\dots,v_{q_1+1} \}$ is linearly independent. The analogue assertion holds for the gradient $\nabla_{\theta^\beta} q_\beta(X_t,\theta_0^\beta)$, replacing $q_1$ by $q_2$. • For all $t\in\mathbb N$, $F_t$ is differentiable at $q_\alpha(X_t,\theta_0^\alpha)$ and $q_\beta(X_t,\theta_0^\beta)$ and the derivatives satisfy $f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) > 0$ and $f_t\big(q_\beta(X_t,\theta_0^\beta)\big) >0$, and $g'_{1,t}(\xi_1)>0,$ $g'_{2,t}(\xi_2)>0$ for all $\xi_1,\xi_2$. \end{enumerate} Then, the following statements hold: \begin{enumerate}[label = \normalfont {(\Alph*)}] • Let $\nabla_{\theta^\alpha} q_\alpha(X_t,\theta_0^\alpha) = \nabla_{\theta^\beta} q_\beta(X_t,\theta_0^\beta)$ for all $t\in\mathbb N$. The M-estimator $\widehat \theta_{M,T,G}$ attains the Z-estimation efficiency bound in (ref) if and only if the following three conditions hold: \begin{align} \exists c_1>0 \ \forall t\in\mathbb N: \quad f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) &= c_1 f_t\big(q_\beta(X_t,\theta_0^\beta)\big) \quad a.s., \\ \exists c_2>0 \ \forall t\in\mathbb N: \quad g_{1,t}'\big(q_\alpha(X_t,\theta_0^\alpha)\big) &= c_2 f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) \quad a.s., \\ \exists c_3>0 \ \forall t\in\mathbb N: \quad g_{2,t}'\big(q_\beta(X_t,\theta_0^\beta)\big) &= c_3 f_t\big(q_\beta(X_t,\theta_0^\beta)\big) \quad a.s. \end{align} • Furthermore, if (ref) or (ref) is violated, then $\widehat \theta_{M,T,G}$ does not attain the Z-estimation efficiency bound in (ref). \end{enumerate}

A discussion of the conditions of Theorem (ref) is in order. Assumptions (ref) and (ref) are required to characterize the class of M-estimators; see the previous Running Examples and DFZ_CharMest and FisslerZiegel2016 for a discussion. The separated parameter condition (ref) contains a large class of possible models. E.g., it nests classically used individual quantile models for separate probability levels $\alpha$ and $\beta$. These parameters can also be restricted through inequality relations, e.g., to impede quantile crossings. While models with joint parameters would also be interesting, completely different methods of proof are required to generalize Theorem (ref) along these lines. Our simulation results in Section (ref) indicate that the efficiency gap carries over to joint parameter models, and is numerically even more severe.

Condition (ref) concerns the variability of the model gradient and is slightly stronger than the classical assumption on univariate models $m$ that the matrix $\mathbb{E} \big[ \nabla_\theta m(X_t,\theta_0) \nabla_\theta m(X_t,\theta_0)^\intercal \big]$ is of full rank for all $t \in \mathbb{N}$. E.g., consider a linear model with explanatory variable $X_t = (1,V_t)^\intercal$, where $V_t$ attains only $0$ and $1$ with positive probability. Then, $\mathbb{E} \big[ \nabla_\theta m(X_t,\theta_0) \nabla_\theta m(X_t,\theta_0)^\intercal \big] = \mathbb{E} \big[ X_t X_t^\intercal \big]$ is positive definite whereas condition (ref) is not fulfilled. However, if $V_t$ attains at least three different values with positive probability (or if its distribution is absolutely continuous), (ref) holds. Condition (ref) is standard for asymptotic normality in quantile regressions.

The gradient condition $\nabla_{\theta^\alpha} q_\alpha(X_t,\theta_0^\alpha) = \nabla_{\theta^\beta} q_\beta(X_t,\theta_0^\beta)$ in (ref) is mainly motivated through models that are linear in the parameters, where these gradients are simply $X_t$. In contrast, statement (ref) holds independent of this gradient condition for general semiparametric models with separated parameters, but does not provide sufficient conditions for efficient M-estimation.

Section (ref) shows that the efficiency gap indeed affects the important diagonal entries of the covariance matrix, which is not immediate from Theorem (ref).

For the remainder of this subsection, we assume for simplicity that the gradient condition $\nabla_{\theta^\alpha} q_\alpha(X_t,\theta_0^\alpha) = \nabla_{\theta^\beta} q_\beta(X_t,\theta_0^\beta)$ holds, putting us in the situation of (ref). Then, the core condition of this theorem on the underlying process is ((ref)). Given that ((ref)) holds, the remaining conditions ((ref)) and ((ref)) are fulfilled by using the obvious choices

align[align omitted — 200 chars of source]

These conditions coincide with classical efficient semiparametric quantile estimation (for one quantile only) in Komunjer2010a, Komunjer2010b. We refer to ((ref)) as the pseudo-efficient choices as they attain the Z-estimation efficiency bound only in certain situations.

We now analyze the validity of the core condition ((ref)) for double quantile models of the form

align[align omitted — 141 chars of source]

where the two quantile-innovations $(u_t^\alpha)_{t\in\mathbb N}$ and $(u_t^\beta)_{t\in\mathbb N}$ satisfy the quantile-stationarity conditions $Q_\alpha(u_t^\alpha \,|\, X_t ) = 0$ and $Q_\beta(u_t^\beta \,|\, X_t) = 0$, such that Assumption (ref) is satisfied. Apart from this assumption, these innovations can be heterogeneously distributed. Clearly, $u_t^\alpha$ and $u_t^\beta$ are generally dependent.

Such correctly specified double quantile models can for instance be generated through a process

align[align omitted — 89 chars of source]

for functions $\zeta\colon \mathbb{R}^p\to\mathbb{R}$, $\eta\colon \mathbb{R}^p\to (0,\infty)$, where the innovations $(\varepsilon_t)_{t\in\mathbb N}$ are themselves independent, independent of $(X_t)_{t\in\mathbb N}$, and where $z_\alpha = F_{\varepsilon_t}^{-1}(\alpha)$ and $z_\beta = F_{\varepsilon_t}^{-1}(\beta)$ are time-independent. Then, the conditional quantiles at level $\alpha \in (0,1)$ (and equivalently for $\beta$) are given by $q_\alpha(X_t, \theta_0^\alpha) = Q_\alpha(Y_t|X_t) = \zeta(X_t) + \eta(X_t) z_\alpha$. E.g., if $\zeta(X_t)$ and $\eta(X_t)$ are linear in $X_t$, as in the simulation setup in Section (ref), we also get linear conditional quantile models $q_\alpha(X_t, \theta_0^\alpha)$ and $q_\beta(X_t, \theta_0^\beta)$. While the process in (ref) resembles the ubiquitous class of location-scale processes, the quantities $\zeta(X_t)$ and $\eta(X_t)$ possibly lose their interpretation as location and scale for sufficiently heterogeneously distributed innovations $\varepsilon_t$.

For a process in (ref), the density transformation formula yields that ((ref)) is equivalent to

align[align omitted — 277 chars of source]

This implies that for processes of the form ((ref)), the M-estimator $\widehat \theta_{M,T,G}$ of the double quantile model is able to attain the efficiency bound (based on the choices in ((ref))), if and only if the density ratio in ((ref)) is constant in $t$. Consequently, for any i.i.d.\ innovations $(\varepsilon_t)_{t\in\mathbb N}$, the M-estimator based on the choices ((ref)) attains the Z-estimation efficiency bound.

However, one can easily construct examples where condition ((ref)) is violated, e.g., by considering Student's $t$-distributed innovations $\varepsilon_t \sim t_{\nu_t}(\mu_t,\sigma_t^2)$ with time-varying degrees of freedom $\nu_t$, and where the time-varying means and standard deviations are given by

align[align omitted — 230 chars of source]

where $t_\nu = t_\nu(0,1)$. These choices ensure that for $\alpha, \beta \in (0,1)$, $\alpha < \beta$, we have $Q_\alpha \left( t_{\nu_t} \big( \mu_{t}, \sigma_{t}^2 \big) \right) = z_\alpha$ and $Q_\beta \left( t_{\nu_t} \big( \mu_{t}, \sigma_{t}^2 \big) \right) = z_\beta$ for all $t \in \mathbb{N}$, and hence, the quantile-stationarity condition is satisfied while simultaneously condition ((ref)) is violated for all quantile levels such that $\alpha \not= 1-\beta$.

For centered or equal-tailed prediction intervals with $\alpha= 1-\beta < 0.5$, we can choose skewed normally distributed innovations Azzalini1985 $\varepsilon_t \sim \mathcal{SN} (\mu_t,\sigma_t^2, \gamma_t)$ with time-varying skewness $\gamma_t$, where the means $\mu_t$ and the standard deviations $\sigma_t$ are given by

align[align omitted — 403 chars of source]

where $\mathcal{SN} (\gamma_1) := \mathcal{SN} (0,1,\gamma_1)$. Then, $Q_\alpha \left( \mathcal{SN} (\mu_t, \sigma_t^2, \gamma_t) \right) = z_\alpha$ and $Q_\beta \left( \mathcal{SN} (\mu_t, \sigma_t^2, \gamma_t) \right) = z_\beta$ for all $t \in \mathbb{N}$ and for all $\alpha, \beta \in (0,1)$, $\alpha < \beta$. We employ these models in the simulations in Section (ref), where we numerically confirm the theoretical results of this section.

Constructing further processes where the M-estimator cannot attain the Z-estimation efficiency bound can be carried out along these lines, where the crucial condition is that the data generating mechanism must go beyond the class of simple location-scale processes with i.i.d.\ residuals. Interesting candidates are GAS models of Creal2013, and specifically for quantiles, the CAViaR specifications of Engle2004 and WhiteKimManganelli2015.

In summary, there exists an efficiency gap for the double quantile model. Its presence mainly depends on the underlying process through the key condition in (ref). The elementary reason for this efficiency gap is the relatively narrow class of strictly consistent loss functions for quantiles at different levels in ((ref)), which only consists of the sum of strictly consistent losses for the individual quantiles. In particular, this class is much smaller than the corresponding class of strict identification functions; see the Running Example (ref) for details.

Semiparametric Joint Quantile and ES Models

Consider a joint model for the quantile (or VaR) and ES at level $\alpha \in (0,1)$, given by $m(X_t,\theta) = \big( q_\alpha(X_t,\theta), e_\alpha(X_t,\theta) \big)^\intercal$, where $q_\alpha(X_t,\theta)$ is a model for the $\alpha$-quantile and $e_\alpha(X_t, \theta)$ denotes a model for the $\operatorname{ES}_\alpha$ at level $\alpha$. For a random variable $Z$ with quantiles $Q_u(Z)$, the $\operatorname{ES}_\alpha(Z)$ is defined as $ \frac{1}{\alpha} \int_0^\alpha Q_u(Z) \mathrm{d}u$ that simplifies to $\operatorname{ES}_\alpha(Z) = \mathbb{E} \left[ Z \; | \; Z \le Q_\alpha(Z) \right]$ if $\mathbb{P}\big(Z \le Q_\alpha(Z)\big)=\alpha$.

As shown by Gneiting2011 and Weber2006, ES is generally neither elicitable nor identifiable and thus, Theorem 1 (ii) and (iv) DFZ_CharMest and Propositions S1 and S3 in its supplementary material provide formal evidence that both M- and Z-estimation of semiparametric models for the conditional ES stand-alone are infeasible. However, FisslerZiegel2016 show that under mild conditions, the pair $(Q_\alpha, \operatorname{ES}_\alpha)$ is jointly elicitable and identifiable, and further characterize the class of strictly consistent loss functions. Due to the recent introduction of ES into the Basel framework as the standard risk measure in banking regulation Basel2016, there is a fast-growing interest in semiparametric models for ES (jointly with the quantile) and Patton2019, DimiBayer2019, Taylor2019, DimiSchnaitmann2019, GuillenETAL2021, among many others, utilize these losses for M-estimation of joint semiparametric models.

Suppose that $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ contains only continuous and strictly increasing distribution functions with an integrable lower tail. Consider the strict $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-identification function

align[align omitted — 225 chars of source]

and define the Z-estimator $\widehat \theta_{Z,T,\mathbb{A}}$ via (ref) and (ref) based on some sequence of instrument matrices $\mathbb{A}$. From Theorem (ref), we get that the efficient estimator has to fulfil the condition $A^\ast_{t,C}(X_t,\theta_0) = C D_t(X_t,\theta_0)^\intercal S_t(X_t,\theta_0)^{-1}$ for some deterministic and nonsingular matrix $C$, where

align[align omitted — 741 chars of source]

Under Assumptions (ref) and (ref), DFZ_CharMest shows that M-estimation can be carried out if and only if a (strictly) $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent loss functions for the pair $(Q_\alpha, \operatorname{ES}_\alpha)$ is used. FisslerZiegel2016 show that under Assumption (ref), this whole class is given by

align[align omitted — 376 chars of source]

where $\xi_1\mapsto g_t(\xi_1) + \xi_1\phi_t'(\xi_2)/\alpha$ is (strictly) increasing for each $\xi_2$, $\phi_t$ is (strictly) convex and $\rho_t$ is $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-integrable. For sequences $G = (g_t)_{t\in\mathbb N}$ and $\Phi = (\phi_t)_{t\in\mathbb N}$ of such functions, we denote the M-estimator defined via (ref) and (ref) by $\widehat \theta_{M,T,G, \Phi}$.

The following theorem establishes that, under certain conditions, the M-estimator of the joint quantile and ES regression model is subject to the efficiency gap.

theoremSuppose that Assumptions (ref), (ref) together with Assumptions (ref) and (ref) in Appendix (ref) hold for the joint quantile and ES model at level $\alpha\in(0,1)$, $\widehat \theta_{M,T,G, \Phi}$ is asymptotically normal and the following further regularity conditions hold: \begin{enumerate}[label=\normalfont {(QES\arabic*)}, leftmargin=*] • The parameters of the individual models are separated, $m(X_t,\theta) = \big(q_\alpha(X_t,\theta^q), e_\alpha(X_t,\theta^e) \big)^\intercal$, where $\theta = \big( \theta^q, \theta^e \big) \in \Theta \subseteq \mathbb{R}^q$, with $\theta^q \in \mathbb{R}^{q_1}$ and $\theta^e \in \mathbb{R}^{q_2}$ and $q_1 + q_2 = q$. • For all $t\in\mathbb N$, and for all $A\in \mathcal A$ with $\mathbb P(A)=1$ there are $q_1+1$ mutually different $v_1, \ldots, v_{q_1+1} \in \big\{\nabla_{\theta^q} q_\alpha(X_t(\omega),\theta_0^q) \in \mathbb{R}^{q_1} \colon \omega\in A\big\}\subseteq \mathbb{R}^{q_1}$, such that any subset of cardinality $q_1$ of $\{ v_1,\dots,v_{q_1+1} \}$ is linearly independent. The analogue assertion holds for the gradient $\nabla_{\theta^e} e_\alpha(X_t,\theta_0^e)$, replacing $q_1$ by $q_2$. • For all $t\in\mathbb N$, $F_t$ is differentiable at $q_\alpha(X_t,\theta_0^q)$ with $f_t\big(q_\alpha(X_t,\theta_0^q)\big) > 0$ and $ g_t'(\xi_1) + \phi_t'(\xi_2)/\alpha >0$ and $\phi_t''(\xi_2) >0$ for all $\xi_1, \xi_2$. \end{enumerate} Then, the following statements hold: \begin{enumerate}[label = \normalfont {(\Alph*)}] • Let $\nabla_{\theta^q} q_\alpha(X_t,\theta_0^q) = \nabla_{\theta^e} e_\alpha(X_t,\theta_0^e)$ for all $t \in \mathbb{N}$. The M-estimator $\widehat \theta_{M,T,G, \Phi}$ attains the Z-estimation efficiency bound in (ref) if and only if the following five conditions hold: \begin{align} \exists c_1>0 \ \forall t\in\mathbb N :\quad \operatorname{Var}_t \big( Y_t \big| Y_t \le q_\alpha(X_t,\theta^q_0) \big) = c_1 \big( q_\alpha(X_t,\theta^q_0) - e_\alpha(X_t,\theta^e_0) \big)^2\quad a.s., \\ \exists c_2>0 \ \forall t\in\mathbb N :\quad f_t\big(q_\alpha(X_t,\theta^q_0)\big) = \frac{c_2}{q_\alpha(X_t,\theta^q_0) - e_\alpha(X_t,\theta^e_0)} \quad a.s.,\\ \exists c_3>0 \ \forall t\in\mathbb N :\quad \phi_t”\big(e_\alpha(X_t,\theta^e_0)\big) =\frac{c_3}{\operatorname{Var}_t \big( Y_t \big| Y_t \le q_\alpha(X_t,\theta^q_0) \big)} \quad a.s., \\ \exists c_4 \in\mathbb{R} \ \forall t\in\mathbb N \ \exists c_{5,t} \in \mathbb{R}:\quad g_t'\big(q_\alpha(X_t,\theta^q_0)\big) = c_4 f_t\big(q_\alpha(X_t,\theta^q_0)\big) + c_{5,t}\quad a.s., \\ \forall t\in\mathbb N :\quad \phi_t'\big(e_\alpha(X_t,\theta^e_0)\big) = \frac{c_3}{c_1 c_2} f_t\big(q_\alpha(X_t,\theta^q_0)\big) - \alpha c_{5,t} \quad a.s. \end{align} • Furthermore, if (ref), or (ref), or \begin{align} \exists c_6>0 \ \forall t\in\mathbb N:\quad g_t'\big(q_\alpha(X_t,\theta^q_0)\big) + \phi_t'\big(e_\alpha(X_t,\theta^e_0)\big)/\alpha = c_6 f_t\big(q_\alpha(X_t,\theta^q_0)\big) \quada.s. \end{align} is violated, then $\widehat \theta_{M,T,G, \Phi}$ does not attain the Z-estimation efficiency bound in (ref). \end{enumerate}

The general structure of Theorem (ref) is similar to Theorem (ref): Statement (ref) provides necessary and sufficient conditions as to when the M-estimation and Z-estimation efficiency bounds coincide, using the additional assumption on the model gradients $\nabla_{\theta^q} q_\alpha(X_t,\theta_0^q) = \nabla_{\theta^e} e_\alpha(X_t,\theta_0^e)$. Dispensing with the latter condition, (ref) provides necessary conditions only. Also, the conditions (ref)\,--\,(ref) resemble the conditions (ref)\,--\,(ref) and are satisfed for a large class of processes and estimators, see the discussion after Theorem (ref) and in Patton2019.

For the remainder for this section, we assume that the gradient condition $\nabla_{\theta^q} q_\alpha(X_t,\theta_0^q) = \nabla_{\theta^e} e_\alpha(X_t,\theta_0^e)$ holds, putting us in the situation of (ref). Then, the core conditions for efficiency of the joint quantile and ES models are given in ((ref))\,--\,((ref)), where the conditions ((ref)), ((ref)) only depend on the underlying process and do not involve $g_t$ and $\phi_t$, resembling condition ((ref)). These two conditions result from the rather restrictive shape of the class of (strictly) consistent loss functions in ((ref)), see FisslerZiegel2016 for details. Section (ref) further analyzes the validity of ((ref)), ((ref)) with results resembling the ones for double quantile models from the previous section. Given that these conditions hold, efficient M-estimation can be performed by employing suitable choices of $g_t$ and $\phi_t$ satisfying (ref)\,--\,(ref), which are further discussed in Section (ref) and which resemble conditions ((ref)) and ((ref)).

Conditions ((ref))\,--\,((ref)) and (ref) illustrate the concordance with mean and quantile regression models. Condition (ref) (which can be split into (ref) and (ref) under the equality of the model gradients) is closely related to the efficient choice for semiparametric quantile models, see Komunjer2010a, Komunjer2010b, and Section (ref) of this article. However, in contrast to classical quantile regression, it is important to notice that given ((ref)) and ((ref)) hold, the choice $g_t(z) = 0$ (resulting from $c_4=0$ and $c_{5,t}=0$) facilitates efficient estimation through a suitable choice of the function $\phi_t$. Moreover, condition ((ref)) resembles the classical condition of efficient least squares estimation of GourierouxMonfortTrognon1984, where the second derivative of $\phi_t$ is proportional to the reciprocal of the conditional variance. As ES is a tail expectation, one also needs to consider the tail variance in ((ref)).

Barendse2020 considers two-step estimation and a related two-step efficiency bound for semiparametric quantile and ES models that we discuss and relate to our results in Section (ref).

Efficient Estimation of Joint Semiparametric Quantile and ES Models

Here, we discuss feasible choices for $g_t$ and $\phi_t$ satisfying ((ref))\,--\,((ref)) and (ref) to facilitate efficient M-estimation for semiparametric joint quantile and ES models based on Theorem (ref). To this end, we assume that ((ref)) and ((ref)) hold for the underlying process and defer a discussion of these conditions to Section (ref). An obvious solution satisfying ((ref))\,--\,((ref)) is

alignat{3} \begin{aligned} g_t^eff1(\xi_1) &= d_1 F_t(\xi_1), \qquad & for \xi_1 > e_\alpha(X_t, \theta^e_0), \\ \phi_t^eff1(\xi_2) &= - d_2\log \big( q_\alpha(X_t, \theta^q_0) - \xi_2 \big) \qquad & for \xi_2 < q_\alpha(X_t, \theta^q_0), \end{aligned}

for all $t\in \mathbb{N}$ and for some constants $d_1 \ge 0$ and $d_2 > 0$, which we refer to as the first pseudo-efficient choices. Motivated by the condition

align[align omitted — 253 chars of source]

for some $c >0$, given in ((ref)) in the proof of Theorem (ref) and in the two-step efficiency bound of Barendse2020, a second pseudo-efficient choice, satisfying ((ref))\,--\,((ref)), is given by

align[align omitted — 449 chars of source]

for constants $d_3 > 0$, $d_4\ge0$, where $v_t = \operatorname{Var}_t \big( Y_t \big| Y_t \le q_\alpha(X_t,\theta^q_0) \big)$ and $q_t = q_\alpha(X_t, \theta^q_0)$. Then, \[ {\phi_t'}^\text{eff2}(\xi_2) = -\frac{d_3}{\sqrt{(1-\alpha)v_t}} \arctan \left( \frac{\sqrt{1-\alpha}(q_t - \xi_2) }{ \sqrt{v_t}} \right) +\frac{\pi d_3(1+d_4)}{2\sqrt{(1-\alpha)v_t}}>0, \] and ${\phi_t''}^\text{eff2}(\xi_2) = {d_3} \big(v_t + (1-\alpha)(q_t - \xi_2)^2\big)^{-1}>0$, for all $\xi_2 < q_t$.

This illustrates that, given that ((ref)) and ((ref)) hold, there exist different efficient M-estimators. Furthermore, if ((ref)) and ((ref)) do not hold jointly, Theorem (ref) (ref) cannot be employed for a statement on efficiency of different M-estimators and it is generally unclear which choices of $g_t$ and $\phi_t$ result in the most efficient estimator. We analyze this numerically for location-scale process with heteroskedastic innovations in the simulation study in Section (ref). The results there also suggest that there is an efficiency gap in models with joint parameters.

As it is common for efficient semiparametric estimation (cf.\ GourierouxMonfortTrognon1984, Komunjer2010a, Komunjer2010b), the efficient choice depends on the knowledge of the true parameter vector $\theta_0$ and further unknown quantities such as the conditional density $f_t$ evaluated at $q_\alpha(X_t,\theta^q_0)$ or the quantile-truncated variance $\operatorname{Var}_t \big( Y_t \big| Y_t \le q_\alpha(X_t,\theta^q_0) \big)$. In practice, one usually applies a two-step estimation approach where the unknown quantities in the efficient choices are substituted by consistent estimates. Notably, the pseudo-efficient M-estimators based on the first choices $g_t(\xi_1)=0$ and $\phi_t(\xi_2)$ in ((ref)) are remarkably feasible in the sense that they only require a first-step estimate of the quantile-specific parameters. This is considerably easier than the required nonparametric first-step estimators of the conditional variance or the conditional distribution function in efficient M-estimation of mean and quantile regressions.

A further interesting fact arises from a comparison of ((ref)) to the predominantly used loss functions with homogeneous loss differences of degree zero NoldeZiegel2017, given by

align[align omitted — 149 chars of source]

Patton2019 build their M-estimation approach on these choices and DimiBayer2019 numerically show that such M-estimators are relatively efficient.

Comparing the choice $g_t(\xi_1) = 0$ to the efficient choice in ((ref)) illustrates the elegance of the parsimonious choice $d_1=0$. By further comparing the choices of $\phi_t$ in ((ref)) and ((ref)), we see that the zero-homogeneous loss function only deviates from the pseudo-efficient choice in ((ref)) through the translation by $q_\alpha(X_t, \theta^q_0)$. This justifies the choice of Patton2019 ex post and theoretically explains the good numerical performance observed by BayerDimi2019. While the zero-homogeneous choice requires strictly negative values for the conditional ES, employing the closely related efficient choice in ((ref)) makes this condition redundant and instead, we only have to impose the natural condition that the conditional ES is smaller than the conditional quantile. Interestingly, when $d_1=0$, (ref) also constitutes a strictly consistent loss with zero-homogeneous loss differences, when allowing the (itself 1-homogenous) quantile as an input parameter. This does not contradict NoldeZiegel2017 as they naturally do not allow the true quantile as an input parameter.

Processes Generating an Efficiency Gap in joint Quantile and ES Models

In this section, we discuss attainability of the process conditions ((ref)) and ((ref)), which are necessary for the M-estimator to match the Z-estimation efficiency bound under the gradient condition $\nabla_{\theta^q} q_\alpha(X_t,\theta_0^q) = \nabla_{\theta^e} e_\alpha(X_t,\theta_0^e)$. We consider joint quantile and ES models of the form

align[align omitted — 118 chars of source]

where the innovations $(u_t^q)_{t\in\mathbb N}$ and $(u_t^e)_{t\in\mathbb N}$ satisfy the semiparametric stationarity conditions $Q_\alpha(u_t^q \,|\, X_t ) = 0$ and $\operatorname{ES}_\alpha(u_t^e \,|\, X_t) = 0$, such that Assumption (ref) is satisfied.

Such correctly specified models can for instance be generated through the process in (ref), where we---slightly differently from the residual assumption in (ref)---impose that $z_\alpha = F_{\varepsilon_t}^{-1}(\alpha)$ and $s_\alpha = \operatorname{ES}_\alpha(\varepsilon_t)$ are time-independent, such that Assumption (ref) holds. Apart from that, the innovations may be heterogeneously distributed. We then get that $Q_\alpha(Y_t\,|\,X_t) = \zeta(X_t)+ \eta(X_t) z_\alpha$ and $\operatorname{ES}_\alpha(Y_t\,|\,X_t) = \zeta(X_t) + \eta(X_t) s_\alpha$. E.g., if $\zeta(X_t)$ and $\eta(X_t)$ are linear in $X_t$, as in the simulation setup in Section (ref), we also get linear models for $q_\alpha(X_t, \theta_0^q)$ and $e_\alpha(X_t, \theta_0^e)$.

It further holds that $\operatorname{Var}_t \big( Y_t \,|\, Y_t \le q_\alpha(X_t,\theta^q_0) \big) = \eta(X_t)^2 \operatorname{Var}_t \big( \varepsilon_t \,|\, \varepsilon_t \le z_\alpha \big)$, and $f_t \big( q_\alpha(X_t,\theta_0^q) \big) = f_{\varepsilon_t}(z_\alpha) / \eta(X_t) = f_{\varepsilon_t}(z_\alpha)(z_\alpha - s_\alpha)/\big(q_\alpha(X_t,\theta^q_0) - e_\alpha(X_t,\theta^e_0) \big)$. Thus, for stationary innovations $(\varepsilon_t)_{t\in\mathbb N}$, the quantities $\operatorname{Var}_t \big( \varepsilon_t \,|\, \varepsilon_t \le z_\alpha \big)$ and $f_{\varepsilon_t}(z_\alpha)$ are constant, which implies that the conditions ((ref)) and ((ref)) are satisfied, and hence, any M-estimator based on choices for $g_t$ and $\phi_t$ satisfying ((ref))\,--\,((ref)) attains the Z-estimation efficiency bound.

Similarly to Section (ref), we can easily construct processes which generate an efficiency gap by considering time-varying innovation distributions. E.g., we consider independent and Student's $t$-distributed innovations $\varepsilon_t \sim t_{\nu_t}(\mu_t,\sigma_t^2)$ with time-varying degrees of freedom $\nu_t$ and

align[align omitted — 263 chars of source]

The conditions in (ref) are such that the quantile-ES stationarity condition is satisfied. For this process, it still holds that $\operatorname{Var}_t \big( Y_t\,|\,Y_t \le q_\alpha(X_t,\theta_0) \big) = \eta(X_t)^2 \operatorname{Var} \big( \varepsilon_t\,|\,\varepsilon_t \le z_\alpha \big)$, as $\varepsilon_t$ is independent of $X_t$. However, the quantity $\operatorname{Var} \big( \varepsilon_t\,|\,\varepsilon_t \le z_\alpha\big)$ is generally time-varying, and consequently, this violates ((ref)) and hence generates an efficiency gap.

Numerical Illustration of the Efficiency Gap

In this section, we numerically illustrate the efficiency gap for double quantile and joint quantile and ES models by approximating the expectations (over the covariates) in ((ref))\,--\,((ref)) in simulations. We use 1000 simulation replications each consisting of a sample size of $T=2000$.

Double Quantile Models

For the double quantile models, we simulate according to the process in ((ref)), where $X_t \stackrel{\textrm{iid}}{\sim} 3 \operatorname{Beta}(3, 1.5)$, $\zeta(X_t) = 10 + 0.5 X_t$, and $\eta(X_t) = 0.5 + 0.5 X_t$. For the model innovations $\varepsilon_t$, we choose the following three different specifications: (a) $\varepsilon_t \stackrel{iid}{\sim} \mathcal{N}(0,1)$; (b) $\varepsilon_t \sim t_{\nu_t}(\mu_t,\sigma_t^2)$ with time-varying degrees of freedom, $\nu_t = 3 \times \mathds{1}_{\{t \le T/2\}} + 100 \, \mathds{1}_{\{t > T/2\}}$, where $\mu_t$ and $\sigma_t$ are given in ((ref)); and (c) $\varepsilon_t \sim \mathcal{SN} (\mu_t,\sigma_t^2, \gamma_t)$ follows a skewed normal distribution with time-varying skewness, $\gamma_t = 0.9 \, \mathds{1}_{\{t > T/2\}}$, where $\mu_t$ and $\sigma_t$ are given in ((ref)).

These choices are motivated through the theoretical considerations of Section (ref) that for models of the form (ref) with i.i.d.\ residuals, the M-estimator is able to attain the Z-estimation efficiency bound, while conversely, it cannot do so for heterogeneously distributed innovations. The heterogeneously skewed process in (c) is motivated by symmetric prediction intervals where $\alpha = 1- \beta$. Empirically, scenario (b) (and similarly for (c)) can be motivated by a breakpoint model for the degree of heavy-tailedness of the innovations: A period of stress (first part of the sample) exhibiting heavy tails is followed by a relatively calm period (second part of the sample), which is resembled by an innovation-distribution with considerably lighter tails.

For the considered processes, it holds that $Q_\alpha(Y_t|X_t)= (10 + 0.5 z_\alpha) + (0.5 + 0.5z_\alpha) X_t$, and $Q_\beta(Y_t|X_t)= (10 + 0.5 z_\beta) + (0.5 + 0.5z_\beta) X_t$. We estimate linear models with separated parameters,

align[align omitted — 188 chars of source]

In order to consider models with joint parameters, we use a slightly modified parametrization of the process by using $\eta(X_t) = 0.5 X_t$, which implies that $Q_\alpha(Y_t|X_t)= 10 + (0.5 + 0.5z_\alpha) X_t$, and $Q_\beta(Y_t|X_t)= 10 + (0.5 + 0.5z_\beta ) X_t$. Hence, we use the (correctly specified) joint intercept models

align[align omitted — 190 chars of source]

Table (ref) reports the relative standard deviations of the estimated parameters normalized by the corresponding efficiency bound. The row denoted “Efficiency Bound” reports the raw standard deviation. We consider the probability levels $(\alpha, \beta) \in \big\{ (1\%, 2.5\%), (5\%, 95\%), (1\%, 90\%) \big\}$, where the first choice is important for VaR modeling in risk management, while the remaining two consider estimation of a symmetric and an asymmetric prediction interval. Panels A-C consider the separated parameter models in ((ref)) while Panels D-F consider models with joint parameters in ((ref)). We show results for the joint M-estimator using the general loss function in ((ref)) paired with the choices $g_{t}(\xi) = g_{1,t}(\xi) = g_{2,t}(\xi)$ given in the first column of Table (ref) together with the Z-estimation efficiency bound. $F_{\text{Log}}$ denotes the distribution function of a standard logistic distribution. Tables (ref) and (ref) show results for additional probability levels.

table[table omitted — 7,018 chars of source]

The numerical results generally confirm the conclusions of Section (ref): the pseudo-efficient M-estimator with $g_{t}(\xi) = F_t(\xi)$ attains the efficiency bound for homoskedastic innovation distributions, while it cannot attain the efficiency bound for both heteroskedastic processes. Furthermore, as discussed in Section (ref), for symmetric quantile levels as in Panel B, the symmetrically heteroskedastic process in (b) is not sufficient for generating an efficiency gap, whereas the asymmetric process in (c) is sufficient. The latter claim can be seen by the slightly larger standard deviation of $\theta_3$ in Panel B (c), which is not a numerical artifact as it is supported by our theory in Section (ref). Remarkably, even for models with separated parameters, where the pseudo-efficient choices are efficient estimators for the individual quantile models Komunjer2010a, Komunjer2010b, the corresponding joint M-estimator does not attain the (joint) efficiency bound for the processes with heteroskedastic innovations, see Panel A and Section (ref).

We observe that the gap becomes numerically larger for quantile levels in the tails of the conditional distributions (Panel A) and for quantile levels which are close together. This makes it particularly relevant in the VaR literature, where VaR is often reported for multiple, extreme quantile levels. The first observation can be explained by condition ((ref)). For a numerically large efficiency gap, one requires heterogeneity of the conditional densities at the respective quantiles $f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big)$ and $f_t\big(q_\beta(X_t,\theta_0^\beta)\big)$, which is more common in the tails of the conditional distributions than in their central regions. The second observation can be explained by noting that the efficiency gap is driven by the non-zero term $\alpha (1-\beta)$ in the off-diagonal entries of $S_t(X_t,\theta_0)$ in ((ref)). It is particularly large for $\alpha \approx \beta \approx 1/2$ and particularly small for $\alpha <\hspace{-0.3em}< \beta$.

Panels D--F in Table (ref) present results for the models with a joint intercept parameter. We find that the general Z-estimation efficiency bound is still valid, which substantiates the statement of Theorem (ref). Differently from models with separated parameters, the pseudo-efficient choices $g_{t}(\xi) = F_t(\xi)$ generally cannot attain the efficiency bound, even in the homoskedastic residual case, which indicates that the efficiency gap applies to an even wider class of processes in joint parameter models. For both heteroskedastic innovation distributions, the efficiency gap exists and is larger in magnitude. Furthermore, the efficiency gap becomes substantially larger, especially in the example of Panel F, while the pseudo-efficient choices still result in the most efficient estimator among the considered choices of M-estimators. These results show that the efficiency gap is present for a large class of double quantile models and data generating processes, which goes beyond the theoretically considered models of Theorem (ref).

Joint Quantile and ES Models

For joint quantile and ES models with separated model parameters, we use the process in ((ref)) and utilize parametric choices which result in strictly negative ES values, $X_t \stackrel{\textrm{iid}}{\sim} 3 \times \operatorname{Beta}(3, 1.5)$, $\zeta(X_t) = -1 -0.5 X_t$, and $\eta(X_t) = 0.5 + 0.5 X_t$. For the model innovations $\varepsilon_t$, we choose the following two specifications: (a) $\varepsilon_t \stackrel{iid}{\sim} \mathcal{N}(0,1)$; and (b) $\varepsilon_t \sim t_{\nu_t}(\mu_t,\sigma_t^2)$ with time-varying degrees of freedom, $\nu_t = 3 \, \mathds{1}_{\{t \le T/2\}} + 100 \, \mathds{1}_{\{t > T/2\}}$, where $\mu_t$ and $\sigma_t$ are given in ((ref)). These choices are motivated through the theoretical considerations of Section (ref) that for location-scale models with i.i.d.\ residuals, the M-estimator is able to attain the Z-estimation efficiency bound, while conversely, it cannot do so for heterogeneously distributed innovations. Also recall the empirical motivation of breakpoint models from Section (ref)

For the considered process, it holds that $Q_\alpha(Y_t|X_t)= (-1 + 0.5z_\alpha) + (0.5z_\alpha -0.5) X_t$ and $\operatorname{ES}_\alpha(Y_t|X_t) = (-1 + 0.5 s_\alpha) + (0.5s_\alpha -0.5) X_t$. We estimate the following linear models with separated parameters,

align[align omitted — 190 chars of source]

which satisfy the conditions of Theorem (ref). We further consider linear models with joint model parameters where the conditions of Theorem (ref) do not hold in order to assess efficient estimation of quantile--ES models beyond the model classes considered in Theorem (ref). For this, we use a slightly modified parameterisation of the process by using $\eta(X_t) = 0.5 X_t$, which implies that $Q_\alpha(Y_t|X_t)= -1 + (0.5z_\alpha -0.5) X_t$ and $\operatorname{ES}_\alpha(Y_t|X_t) = -1 + (0.5s_\alpha -0.5) X_t$. We use the (correctly specified) joint intercept models

align[align omitted — 190 chars of source]

We consider the quantile and ES at joint probability levels $\alpha \in \{1\%, 2.5\%, 10\%\}$. For $g_t$, we use the two (pseudo-efficient) choices $g_t(\xi_1)=0$ and $g_t(\xi_1)=F_t(\xi_1)$ coupled with several choices of $\phi_t$, see the first two columns of Table (ref) for a detailed list. The first two choices of $\phi_t$ correspond to sub-optimal choices as already noticed by DimiBayer2019, whereas the next choice $\phi_t(\xi_2) = -\log(-\xi_2)$ coincides with the ubiquitous zero-homogeneous loss. The latter two choices $\phi_t^{\text{eff1}}$ and $\phi_t^{\text{eff2}}$ are the pseudo-efficient choices given in ((ref)) and ((ref)).

table[table omitted — 4,118 chars of source]

Panel A of Table (ref) presents the approximated parameter standard deviations for $\alpha = 2.5\%$ and for the separated parameter models in ((ref)). The results confirm the theoretical considerations of Section (ref): the M-estimator based on either of the pseudo-efficient choices, $\phi_t^\text{eff1}$ and $\phi_t^\text{eff2}$, attains the Z-estimation efficiency bound for location-scale models with homoskedastic innovations, while there is an efficiency gap for heteroskedastic innovation distributions with a magnitude of up to $15\%$. Table (ref) in Section (ref) reports additional results for $\alpha = 1\%$ and $\alpha = 10\%$, which shows that the efficiency gap is more pronounced for small(er) probability levels, corresponding to the most important cases for the risk measures VaR and ES.

As indicated by ((ref)), our simulation results confirm that, though counterintuitive at first sight, efficient M-estimation in the homoskedastic case can be accomplished by both, the traditional efficient choice of quantile regression, $g_t(\xi_1) = F_t(\xi_1)$, and by the zero-function $g_t(\xi_1) = 0$. Furthermore, both pseudo-efficient choices $\phi^\text{eff1}$ and $\phi^\text{eff2}$ are able to attain the efficiency bound in the homoskedastic setting for separated parameter models. However, their performance differs for heteroskedastic models, where the choice $\phi^\text{eff2}$ delivers more efficient ES estimates but at the same time slightly less efficient quantile estimates. The function $\phi_t(\xi_2) = - \log(-\xi_2)$ performs almost as well as the pseudo-efficient choices throughout all considered designs, which is not surprising given its similar form to $\phi^\text{eff1}$.

Panel B of Table (ref) presents results for the models with joint parameters given in ((ref)). While the Z-estimation efficiency bound is still valid, it cannot be attained by any of the M-estimators utilized in the simulation study, even in the homoskedastic case, for any of the chosen pseudo-efficient choices. This implies that the efficiency gap for joint quantile and ES models goes beyond the model class considered in Theorem (ref). This holds similarly for the heteroskedastic case, where the efficiency gap becomes quantitatively much larger: the standard deviations of the pseudo-efficient choices are between $34\%$ and $77\%$ larger than the efficiency bound.

As in the heteroskedastic case of Panel A, the second pseudo-efficient choice slightly outperforms the first one also for this example of joint parameter models. Finally, among the considered M-estimators, the ubiquitous zero-homogeneous choice $g_t(\xi_1)=0$ and $\phi_t(\xi_2) = - \log(-\xi_2)$ performs relatively well and even outperforms both pseudo-efficient choices for the heteroskedastic innovation distributions and the joint parameter models of Panel B. This is especially remarkable given that, in contrast to the pseudo-efficient choices, it does not require any pre-estimates in practice.

Conclusion

The results of this paper have important consequences. On the theoretical side, they motivate the consideration of semiparametric M-estimation efficiency bounds, which will generally not coincide with the semiparametric efficiency bound of Stein1956 for multivariate functionals. On the practical side, they suggest the use of a new pseudo-efficient and feasible loss function for M-estimation of semiparametric VaR and ES models Patton2019, Taylor2019, which recently attain a lot of attraction. We anticipate that similar results can be derived for multiple expectiles or the interquantile expectation (RVaR) Barendse2020.

If interest is particularly on efficient estimation for general functionals, our findings suggest the following practical recommendations: If the M-estimator attains the Z-estimation efficiency bound, it seems advisable to use the former due to its often superior numerical stability, and as DFZ_CharMest guarantees its consistency, which can be problematic for (efficient) Z- and GMM-estimation due to missing global identification. On the other hand, in the presence of an efficiency gap, the efficient Z-estimator is more attractive as its efficiency dominates all M-estimators. Its possibly lacking global identification could be remedied by using a consistent pre-estimate---e.g., from (pseudo-efficient) M-estimation---and restricting the numerical optimization algorithm to a local search around the pre-estimate, and by relying on local identification as e.g., suggested by NeweyMcFadden1994.

appendix\setcounter{ass}{2} \section{Additional Assumptions} This section restates assumptions from other papers we use in our theory: Assumption (ref) combines Assumptions 2 and 3 of DFZ_CharMest. Assumptions (ref) and (ref) provide sufficient conditions for the respective characterization results of loss functions for the double quantile and joint quantile and ES models, respectively FisslerZiegel2016. \begin{ass} \begin{enumerate} • For all random variables $Z=(Y,X) \sim F_Z \in \mathcal{F}_{\mathcal{Z}}$, assume that the map $m(X,\cdot)\colon\Theta\to\Xi$ is surjective almost surely. Moreover, the conditional expectation $\mathbb{E}\big[\rho\big(Y,m(X,\theta)\big)\big|X\big]$ is continuous in $\theta$ almost surely. • For all random variables $Z=(Y,X) \sim F_Z \in \mathcal{F}_{\mathcal{Z}}$ and for any event $A\in \sigma(X)$ with positive probability $\mathbb P(A)>0$ the conditional distribution $F_{Z|A}$ is also in $\mathcal{F}_{\mathcal{Z}}$. \end{enumerate} \end{ass} \begin{ass} Let $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ contain all continuously differentiable distribution functions with positive derivatives (densities) such that the double quantile maps surjectively on $\Xi$, which is assumed to be an open and path connected subset of $\mathbb{R}^2$. \end{ass} \begin{ass} Let $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ contain all continuously differentiable distribution functions with positive derivatives (densities) and integrable lower tail such that $(Q_\alpha, \operatorname{ES}_\alpha)$ maps surjectively on $\Xi\subseteq\mathbb{R}^2$, which is assumed to be an open and path connected subset of $\mathbb{R}^2$. \end{ass}

\singlespacing {2pt plus 0.3ex}

thebibliography\bibitem[Ackerberg et al., 2014]{Ackerberg2014} Ackerberg, D., Chen, X., Hahn, J., and Liao, Z. (2014). \newblock {Asymptotic Efficiency of Semiparametric Two-step GMM}. \newblock {\em Review of Economic Studies}, 81(3):919--943. \bibitem[Adrian et al., 2019]{Adrian2019} Adrian, T., Boyarchenko, N., and Giannone, D. (2019). \newblock Vulnerable growth. \newblock {\em American Economic Review}, 109(4):1263--89. \bibitem[Andrews, 1994]{Andrews1994} Andrews, D. (1994). \newblock {E}mpirical {P}rocess {M}ethods in {E}conometrics. \newblock In Engle, R. F. and McFadden, D., editors, {\em Handbook of Econometrics}, volume 4, chapter 37, pages 2247--2294. Elsevier. \bibitem[Angrist et al., 2006]{Angrist2006} Angrist, J., Chernozhukov, V., and Fernandez-Val, I. (2006). \newblock Quantile regression under misspecification, with an application to the u.s. wage structure. \newblock {\em Econometrica}, 74(2):539--563. \bibitem[Azzalini, 1985]{Azzalini1985} Azzalini, A. (1985). \newblock A class of distributions which includes the normal ones. \newblock {\em Scandinavian Journal of Statistics}, 12(2):171--178. \bibitem[Barendse, 2022]{Barendse2020} Barendse, S. (2022). \newblock {Efficiently Weighted Estimation of Tail and Interquantile Expectations}. \newblock {\em Preprint}. \newblock \href{https://dx.doi.org/10.2139/ssrn.2937665}{https://dx.doi.org/10.2139/ssrn.2937665}. \bibitem[Bartalotti, 2013]{Bartalotti2013} Bartalotti, O. (2013). \newblock {GMM Efficiency and IPW Estimation for Nonsmooth Functions}. \newblock {\em Preprint}. \newblock \href{http://repec.tulane.edu/RePEc/pdf/tul1301.pdf}{http://repec.tulane.edu/RePEc/pdf/tul1301.pdf}. \bibitem[{Basel Committee}, 2016]{Basel2016} {Basel Committee} (2016). \newblock {M}inimum capital requirements for {M}arket {R}isk. \newblock Technical report, Bank for International Settlements. \newblock \href{http://www.bis.org/bcbs/publ/d352.pdf}{http://www.bis.org/bcbs/publ/d352.pdf}. \bibitem[Bayer and Dimitriadis, 2022]{BayerDimi2019} Bayer, S. and Dimitriadis, T. (2022). \newblock Regression based expected shortfall backtesting. \newblock {\em Journal of Financial Econometrics}, 20(3). \bibitem[Bellini and Bignozzi, 2015]{BelliniBignozzi2015} Bellini, F. and Bignozzi, V. (2015). \newblock On elicitable risk measures. \newblock {\em Quantitative Finance}, 15(5):725--733. \bibitem[Bickel et al., 1998]{BKRW1998book} Bickel, P. J., Klaassen, C. A. J., Ritov, Y., and Wellner, J. A. (1998). \newblock {\em Efficient and Adaptive Estimation for Semiparametric Models}. \newblock Johns Hopkins series in the mathematical sciences. Springer New York. \bibitem[Bollerslev, 1986]{Bollerslev1986} Bollerslev, T. (1986). \newblock Generalized autoregressive conditional heteroskedasticity. \newblock {\em Journal of Econometrics}, 31(3):307--327. \bibitem[Bracher et al., 2021]{Bracher2021NatComm} Bracher, J., Wolffram, D., Deuschel, J., G{\"o}rgen, K., Ketterer, J. L., Ullrich, A., Abbott, S., Barbarossa, M. V., Bertsimas, D., Bhatia, S., et al. (2021). \newblock A pre-registered short-term forecasting study of covid-19 in germany and poland during the second wave. \newblock {\em Nature Communications}, 12(1):1--16. \bibitem[Brehmer and Gneiting, 2021]{BrehmerGneiting2020} Brehmer, J. and Gneiting, T. (2021). \newblock {Scoring interval forecasts: Equal-tailed, shortest, and modal interval}. \newblock {\em Bernoulli}, 27(3):1993--2010. \bibitem[Buchinsky, 1994]{Buchinsky1994} Buchinsky, M. (1994). \newblock {Changes in the U.S. Wage Structure 1963-1987: Application of Quantile Regression}. \newblock {\em Econometrica}, 62(2):405--458. \bibitem[Catania and Luati, 2019]{Catania2019} Catania, L. and Luati, A. (2019). \newblock Semiparametric modeling of multiple quantiles. \newblock {\em Preprint}. \newblock \href{http://dx.doi.org/10.2139/ssrn.3494995}{http://dx.doi.org/10.2139/ssrn.3494995}. \bibitem[Chamberlain, 1987]{Chamberlain1987} Chamberlain, G. (1987). \newblock {Asymptotic efficiency in estimation with conditional moment restrictions}. \newblock {\em Journal of Econometrics}, 34(3):305--334. \bibitem[Chernozhukov et al., 2010]{Chernozhukov2010} Chernozhukov, V., Fern{\'a}ndez-Val, I., and Galichon, A. (2010). \newblock Quantile and probability curves without crossing. \newblock {\em Econometrica}, 78(3):1093--1125. \bibitem[Cramer et al., 2022]{Cramer2022} Cramer, E. Y., Ray, E. L., Lopez, V. K., Bracher, J., Brennen, A., Castro Rivadeneira, A. J., Gerding, A., Gneiting, T., House, K. H., Huang, Y., et al. (2022). \newblock Evaluation of individual and ensemble probabilistic forecasts of covid-19 mortality in the united states. \newblock {\em Proceedings of the National Academy of Sciences}, 119(15):e2113561119. \bibitem[Creal et al., 2013]{Creal2013} Creal, D., Koopman, S. J., and Lucas, A. (2013). \newblock Generalized autoregressive score models with applications. \newblock {\em Journal of Applied Econometrics}, 28(5):777--795. \bibitem[Davidson, 1994]{davidson1994stochastic} Davidson, J. (1994). \newblock {\em Stochastic Limit Theory: An Introduction for Econometricians}. \newblock Advanced Texts in Econometrics. OUP Oxford. \bibitem[Dimitriadis and Bayer, 2019]{DimiBayer2019} Dimitriadis, T. and Bayer, S. (2019). \newblock A joint quantile and expected shortfall regression framework. \newblock {\em Electronic Journal of Statistics}, 13(1):1823--1871. \bibitem[Dimitriadis et al., 2022a]{DFZ_CharMest} Dimitriadis, T., Fissler, T., and Ziegel, J. (2022a). \newblock Characterizing {M}-estimators. \newblock {\em Preprint}. \newblock \href{https://arxiv.org/abs/2208.08108}{https://arxiv.org/abs/2208.08108}. \bibitem[Dimitriadis et al., 2022b]{DFZ_OsbandID} Dimitriadis, T., Fissler, T., and Ziegel, J. (2022b). \newblock Osband's principle for identification functions. \newblock {\em Preprint}. \newblock \href{https://arxiv.org/abs/2208.07685}{https://arxiv.org/abs/2208.07685}. \bibitem[Dimitriadis et al., 2019]{DimiPattonSchmidt2019} Dimitriadis, T., Patton, A. J., and Schmidt, P. (2019). \newblock {Testing Forecast Rationality for Measures of Central Tendency}. \newblock {\em Preprint}. \newblock \href{https://arxiv.org/abs/1910.12545}{https://arxiv.org/abs/1910.12545}. \bibitem[Dimitriadis and Schnaitmann, 2021]{DimiSchnaitmann2019} Dimitriadis, T. and Schnaitmann, J. (2021). \newblock Forecast encompassing tests for the expected shortfall. \newblock {\em International Journal of Forecasting}, 37(2):604--621. \bibitem[Efron, 1991]{Efron1991} Efron, B. (1991). \newblock Regression percentiles using asymmetric squared error loss. \newblock {\em Statistica Sinica}, 1(1):93--125. \bibitem[Engle and Manganelli, 2004]{Engle2004} Engle, R. F. and Manganelli, S. (2004). \newblock {CAV}ia{R}: {C}onditional {A}utoregressive {V}alue at {R}isk by {R}egression {Q}uantiles. \newblock {\em Journal of Business & Economic Statistics}, 22(4):367--381. \bibitem[Fissler et al., 2021]{FisslerHlavinovaRudloff2019Theory} Fissler, T., Frongillo, R., Hlavinov\'a, J., and Rudloff, B. (2021). \newblock Forecast evaluation of quantiles, prediction intervals, and other set-valued functionals. \newblock {\em Electronic Journal of Statistics}, 15(1):1034--1084. \bibitem[Fissler and Ziegel, 2016]{FisslerZiegel2016} Fissler, T. and Ziegel, J. F. (2016). \newblock {Higher order elicitability and Osband's principle}. \newblock {\em Annals of Statistics}, 44(4):1680--1707. \bibitem[Fissler and Ziegel, 2019]{FisslerZiegel2019} Fissler, T. and Ziegel, J. F. (2019). \newblock Order-sensitivity and equivariance of scoring functions. \newblock {\em Electronic Journal of Statistics}, 13(1):1166--1211. \bibitem[Gneiting, 2011a]{Gneiting2011} Gneiting, T. (2011a). \newblock {Making and Evaluating Point Forecasts}. \newblock {\em Journal of the American Statistical Association}, 106:746--762. \bibitem[Gneiting, 2011b]{Gneiting2011b} Gneiting, T. (2011b). \newblock Quantiles as optimal point forecasts. \newblock {\em International Journal of Forecasting}, 27(2):197--207. \bibitem[Gourieroux and Jasiak, 2008]{GourierouxJasiak2008} Gourieroux, C. and Jasiak, J. (2008). \newblock {Dynamic quantile models}. \newblock {\em Journal of Econometrics}, 147(1):198--205. \bibitem[Gourieroux et al., 1987]{Gourieroux1987} Gourieroux, C., Monfort, A., and Renault, E. (1987). \newblock Consistent {M-estimators} in a semi-parametric model. \newblock {\em CEPREMAP Working Paper 8720}. \newblock \href{http://www.cepremap.fr/depot/couv_orange/co8720.pdf}{http://www.cepremap.fr/depot/couv_orange/co8720.pdf}. \bibitem[Gourieroux et al., 1984]{GourierouxMonfortTrognon1984} Gourieroux, C., Monfort, A., and Trognon, A. (1984). \newblock Pseudo maximum likelihood methods: Theory. \newblock {\em Econometrica}, 52(3):681--700. \bibitem[Guillen et al., 2021]{GuillenETAL2021} Guillen, M., Berm{\'u}dez, L., and Pitarque, A. (2021). \newblock Joint generalized quantile and conditional tail expectation regression for insurance risk analysis. \newblock {\em Insurance: Mathematics and Economics}, 99:1--8. \bibitem[Hall, 2005]{HallBook2005} Hall, A. (2005). \newblock {\em Generalized Method of Moments}. \newblock Advanced texts in econometrics. Oxford University Press. \bibitem[Hansen, 1982]{Hansen1982} Hansen, L. P. (1982). \newblock Large sample properties of generalized method of moments estimators. \newblock {\em Econometrica}, 50(4):1029--54. \bibitem[Hansen, 1985]{Hansen1985} Hansen, L. P. (1985). \newblock A method for calculating bounds on the asymptotic covariance matrices of generalized method of moments estimators. \newblock {\em Journal of Econometrics}, 30(1):203--238. \bibitem[Hristache and Patilea, 2016]{Hristache2016} Hristache, M. and Patilea, V. (2016). \newblock Semiparametric efficiency bounds for conditional moment restriction models with different conditioning variables. \newblock {\em Econometric Theory}, 32(4):917--946. \bibitem[Huber, 1967]{Huber1967} Huber, P. J. (1967). \newblock {T}he behavior of maximum likelihood estimates under nonstandard conditions. \newblock In {\em Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability}, pages 221--233. Berkeley: University of California Press. \bibitem[Jankov{\'a} and van de Geer, 2018]{Jankova2018} Jankov{\'a}, J. and van de Geer, S. (2018). \newblock Semiparametric efficiency bounds for high-dimensional models. \newblock {\em Annals of Statistics}, 46(5):2336--2359. \bibitem[Koenker, 2005]{Koenker2005} Koenker, R. (2005). \newblock {\em {Quantile Regression}}. \newblock Cambridge University Press, Cambridge. \bibitem[Koenker and Bassett, 1978]{Koenker1978} Koenker, R. and Bassett, G. (1978). \newblock {R}egression quantiles. \newblock {\em Econometrica}, pages 33--50. \bibitem[Komunjer, 2005]{Komunjer2005} Komunjer, I. (2005). \newblock Quasi-maximum likelihood estimation for conditional quantiles. \newblock {\em Journal of Econometrics}, 128(1):137--164. \bibitem[Komunjer, 2012]{Komunjer2012} Komunjer, I. (2012). \newblock Global identification in nonlinear models with moment restrictions. \newblock {\em Econometric Theory}, 28(4):719--729. \bibitem[Komunjer and Vuong, 2010a]{Komunjer2010b} Komunjer, I. and Vuong, Q. (2010a). \newblock Efficient estimation in dynamic conditional quantile models. \newblock {\em Journal of Econometrics}, 157(2):272--285. \bibitem[Komunjer and Vuong, 2010b]{Komunjer2010a} Komunjer, I. and Vuong, Q. (2010b). \newblock Semiparametric efficiency bound in time-series models for conditional quantiles. \newblock {\em Econometric Theory}, 26(02):383--405. \bibitem[Lambert, 2019]{Lambert2013} Lambert, N. (2019). \newblock {Elicitation and Evaluation of Statistical Forecasts}. \newblock {\em Preprint}. \newblock \href{https://ai.stanford.edu/ nlambert/papers/elicitation_statistical_forecasts.pdf}{https://ai.stanford.edu/nlambert/papers/elicitation_statistical_forecasts.pdf}. \bibitem[Nau, 1985]{Nau1985} Nau, R. F. (1985). \newblock {Should Scoring Rules Be `Effective'?} \newblock {\em Management Science}, 31(5):527--535. \bibitem[Newey and West, 1987]{NeweyWest1987} Newey, W. and West, K. (1987). \newblock A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix. \newblock {\em Econometrica}, 55(3):703--08. \bibitem[Newey, 1990]{Newey1990} Newey, W. K. (1990). \newblock Semiparametric efficiency bounds. \newblock {\em Journal of Applied Econometrics}, 5(2):99--135. \bibitem[Newey, 1993]{Newey1993} Newey, W. K. (1993). \newblock {Efficient Estimation of Models with Conditional Moment Restrictions}. \newblock In Maddala, G., Rao, C., and Vinod, H., editors, {\em {Handbook of Statistics, Volume 11: Econometrics}}. \bibitem[Newey and McFadden, 1994]{NeweyMcFadden1994} Newey, W. K. and McFadden, D. (1994). \newblock {L}arge sample estimation and hypothesis testing. \newblock In Engle, R. F. and McFadden, D., editors, {\em {Handbook of Econometrics}}, volume 4, chapter 36, pages 2111--2245. Elsevier. \bibitem[Nolde and Ziegel, 2017]{NoldeZiegel2017} Nolde, N. and Ziegel, J. F. (2017). \newblock {Elicitability and backtesting: Perspectives for banking regulation}. \newblock {\em Annals of Applied Statistics}, 11(4):1833--1874. \bibitem[Osband, 1985]{Osband1985} Osband, K. H. (1985). \newblock {\em {Providing Incentives for Better Cost Forecasting}}. \newblock PhD thesis, University of California, Berkeley. \newblock \href{https://doi.org/10.5281/zenodo.4355667}{https://doi.org/10.5281/zenodo.4355667}. \bibitem[Patton, 2011]{Patton2011} Patton, A. (2011). \newblock Volatility forecast comparison using imperfect volatility proxies. \newblock {\em Journal of Econometrics}, 160(1):246--256. \bibitem[Patton et al., 2019]{Patton2019} Patton, A. J., Ziegel, J. F., and Chen, R. (2019). \newblock Dynamic semiparametric models for expected shortfall (and value-at-risk). \newblock {\em Journal of Econometrics}, 211(2):388 -- 413. \bibitem[Prokhorov and Schmidt, 2009]{ProkhorovSchmidt2009} Prokhorov, A. and Schmidt, P. (2009). \newblock {GMM} redundancy results for general missing data problems. \newblock {\em Journal of Econometrics}, 151(1):47--55. \bibitem[Roehrig, 1988]{Roehrig1988} Roehrig, C. S. (1988). \newblock Conditions for identification in nonparametric and parametric models. \newblock {\em Econometrica}, 56(2):433--447. \bibitem[Savage, 1971]{Savage1971} Savage, L. J. (1971). \newblock {Elicitation of Personal Probabilities and Expectations}. \newblock {\em Journal of the American Statistical Association}, 66:783--801. \bibitem[Shrestha and Solomatine, 2006]{Shrestha2006} Shrestha, D. L. and Solomatine, D. P. (2006). \newblock Machine learning approaches for estimation of prediction interval for the model output. \newblock {\em Neural Networks}, 19(2):225--235. \bibitem[Spady and Stouli, 2018]{Spady2018} Spady, R. and Stouli, S. (2018). \newblock Simultaneous mean-variance regression. \newblock {\em Preprint}. \newblock \href{https://arxiv.org/abs/1804.01631}{https://arxiv.org/abs/1804.01631}. \bibitem[Stein, 1956]{Stein1956} Stein, C. (1956). \newblock Efficient nonparametric testing and estimation. \newblock In {\em {Proceedings of the Third Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Contributions to the Theory of Statistics}}, pages 187--195, Berkeley, Calif. University of California Press. \bibitem[Steinwart et al., 2014]{SteinwartPasinETAL2014} Steinwart, I., Pasin, C., Williamson, R., and Zhang, S. (2014). \newblock {Elicitation and Identification of Properties}. \newblock {\em JMLR Workshop Conf. Proc.}, 35:1--45. \bibitem[Taylor, 2019]{Taylor2019} Taylor, J. W. (2019). \newblock Forecasting value at risk and expected shortfall using a semiparametric approach based on the asymmetric laplace distribution. \newblock {\em Journal of Business & Economic Statistics}, 37(1):121--133. \bibitem[Wang and Wei, 2020]{WangWei2020} Wang, R. and Wei, Y. (2020). \newblock Risk functionals with convex level sets. \newblock {\em Mathematical Finance}, 30(4):1337--1367. \bibitem[Weber, 2006]{Weber2006} Weber, S. (2006). \newblock {Distribution-Invariant Risk Measures, Information, and Dynamic Consistency}. \newblock {\em Mathematical Finance}, 16:419--441. \bibitem[Weiss, 1991]{Weiss1991} Weiss, A. A. (1991). \newblock Estimating nonlinear dynamic models using least absolute error estimation. \newblock {\em Econometric Theory}, 7(01):46--68. \bibitem[White et al., 2015]{WhiteKimManganelli2015} White, H., Kim, T.-H., and Manganelli, S. (2015). \newblock {VAR for VaR: Measuring tail dependence using multivariate regression quantiles}. \newblock {\em Journal of Econometrics}, 187(1):169--188.

\onehalfspacing \setcounter{page}{1} \setcounter{footnote}{0}

center[center omitted — 209 chars of source]

\setcounter{section}{0} \setcounter{table}{0} \setcounter{figure}{0}

\setcounter{ass}{5}

Proofs for the Results of the Main Paper

proof[Proof of Theorem (ref)] For $A_{t,C}^\ast(X_t,\theta_0) = C D_t(X_t,\theta_0)^\intercal S_t(X_t,\theta_0)^{-1}$, one obtains that $\Delta_{T,\mathbb{A}^\ast} = \frac{1}{T} \sum_{t=1}^T C \, \mathbb{E} \left[ D_t(X_t,\theta_0)^\intercal S_t(X_t,\theta_0)^{-1} D_t(X_t,\theta_0) \right]$ and $\Sigma_{T,\mathbb{A}^\ast} = \Delta_{T,\mathbb{A}^\ast}C^\intercal$. Thus, the asymptotic covariance of the Z-estimator based on the choice $A_{t,C}^\ast(X,\theta_0)$ is the limit of \[ \Delta_{T,\mathbb{A}^\ast}^{-1} \Sigma_{T,\mathbb{A}^\ast}\left(\Delta_{T,\mathbb{A}^\ast}^{-1}\right)^\intercal = \Lambda_T^{-1} = \left( \frac{1}{T} \sum_{t=1}^T \mathbb{E} \big[ D_t(X_t,\theta_0)^\intercal S_t(X_t,\theta_0)^{-1} D_t(X_t,\theta_0) \big] \right)^{-1} \] for all deterministic and non-singular choices of $C$, which shows part (ref) of Theorem (ref). As the asymptotic covariance is independent of the choice of $C$, without loss of generality we continue with $C = I_q$ for the proof of part (ref) and henceforth use the notation $A_{t}^\ast = A_{t, I_q}^\ast$. We define the random vector $\chi_{t,T} = \big( \Delta_{T,\mathbb{A}}^{-1} A_t(X_t,\theta_0) - \Lambda_{T}^{-1} A_{t}^\ast(X_t,\theta_0) \big) \varphi\big(Y_t,m(X_t,\theta_0) \big)$ for all $t, 1\le t \le T, T \ge 1$. Straight-forward calculations yield that \begin{align*} &\frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \chi_{t,T} \chi_{t,T}^\intercal \right] = \\ &\Delta_{T,\mathbb{A}}^{-1} \left( \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ A_t(X_t,\theta_0) \varphi\big(Y_t, m(X_t,\theta_0) \big) \varphi\big(Y_t,m(X_t,\theta_0)\big)^\intercal A_t(X_t,\theta_0)^\intercal \right] \right) (\Delta_{T,\mathbb{A}}^\intercal)^{-1} \\ + \, &\Lambda_T^{-1} \left( \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ A_t^\ast(X_t,\theta_0) \varphi\big(Y_t, m(X_t,\theta_0) \big) \varphi\big(Y_t, m(X_t,\theta_0) \big)^\intercal A_t^\ast(X_t,\theta_0)^\intercal \right] \right) (\Lambda_T^{-1})^\intercal \\ - \, &\Delta_{T,\mathbb{A}}^{-1} \left( \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ A_t(X_t,\theta_0) \varphi\big(Y_t, m(X_t,\theta_0) \big) \varphi\big(Y_t, m(X_t,\theta_0) \big)^\intercal A_t^\ast(X_t,\theta_0)^\intercal \right] \right) (\Lambda_T^{-1})^\intercal \\ - \, &\Lambda_T^{-1} \left( \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ A_t^\ast(X_t,\theta_0) \varphi\big(Y_t, m(X_t,\theta_0) \big) \varphi\big(Y_t, m(X_t,\theta_0) \big)^\intercal A_t(X_t,\theta_0)^\intercal \right] \right) (\Delta_{T,\mathbb{A}}^\intercal)^{-1} \\ = \, &\Delta_{T,\mathbb{A}}^{-1} \left( \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ A_t(X_t,\theta_0) S_t(X_t,\theta_0) A_t(X_t,\theta_0)^\intercal \right] \right) (\Delta_{T,\mathbb{A}}^\intercal)^{-1} \\ + \, &\Lambda_T^{-1} \left( \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ D_t(X_t,\theta_0)^\intercal S_t(X_t,\theta_0)^{-1} D_t(X_t,\theta_0) \right] \right) \Lambda_T^{-1} \\ - \, &\Delta_{T,\mathbb{A}}^{-1} \left( \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ A_t(X_t,\theta_0) D_t(X_t,\theta_0) \right] \right) \Lambda_T^{-1} \\ - \, & \Lambda_T \left( \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ D_t(X_t,\theta_0)^\intercal A_t(X_t,\theta_0)^\intercal \right] \right) (\Delta_{T,\mathbb{A}}^\intercal)^{-1} \\ = \, &\Delta_{T,\mathbb{A}}^{-1} \Sigma_{T,\mathbb{A}} \Delta_{T,\mathbb{A}}^{-1} - \Lambda_T^{-1}, \end{align*} which is positive semi-definite for all $T \ge 1$ as the sum of outer products, which concludes the proof of part (ref). For the proof of part (ref), assume that for some $t = 1,\dots,T$, the matrix $A_t(X_t,\theta)$ is such that $A_t(X_t,\theta_0) \not= A_{t,C}^\ast(X_t,\theta_0)$ for any non-singular and deterministic matrix $C$ with positive probability. Then, for some $t = 1,\dots,T$, the matrix \( M_{T,\mathbb{A}} (X_t,\theta_0) := \Delta_{T,\mathbb{A}}^{-1} A_t(X_t,\theta_0) - \Lambda_{T}^{-1} A_{t,C}^\ast(X_t,\theta_0) \) is nonzero with positive probability, as otherwise $A_t(X_t,\theta_0) = A_{t, \tilde C}^\ast(X_t,\theta_0)$ almost surely with $\tilde C = \Delta_{T,\mathbb{A}} \Lambda_{T}^{-1} C$. This implies that the matrix $M_{T,\mathbb{A}} (X_t,\theta_0)$ has positive rank with positive probability. Furthermore, the matrix $S_t(X_t,\theta_0)$ defined in (ref) is positive definite with probability one for all $t = 1,\dots,T$ by assumption. Consequently, we can apply the Cholesky decomposition and get that there exists a lower triangular matrix $G_t(X_t,\theta_0)$ with strictly positive diagonal entries such that $S_t(X_t,\theta_0) = G_t(X_t,\theta_0) G_t(X_t,\theta_0)^\intercal$ almost surely, i.e., the matrix $G_t(X_t,\theta_0)$ has full rank almost surely. Thus, the matrix \( B_{T,\mathbb{A},t}(X_t,\theta_0) := M_{T,\mathbb{A}}(X_t,\theta_0) G_t(X_t,\theta_0) \) has positive rank for some $t = 1,\dots,T$ with positive probability by Sylvester's rank inequality as it is the product of matrices with strictly positive rank (with positive probability) and full rank (almost surely), respectively. Consequently, there exists a $j \in \{ 1,\dots,k\}$ such that \begin{align} \mathbb{P} \big( B_{T,\mathbb{A},t}(X_t,\theta_0)^\intercal e_j \not= 0 \big) > 0, \qquad for some t = 1,\dots,T, \end{align} where $e_j$ is the $j$-th standard basis vector of $\mathbb{R}^k$. Thus, \begin{align*} &e_j^\intercal \left( \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \chi_{t,T} \chi_{t,T}^\intercal \right] \right) e_j = \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[e_j^\intercal M_{T,\mathbb{A}}(X_t,\theta_0) S_t(X_t,\theta_0) M_{T,\mathbb{A}}(X_t,\theta_0)^\intercal e_j \right] \\ &= \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[e_j^\intercal B_{T,\mathbb{A},t}(X_t,\theta_0) B_{T,\mathbb{A},t}(X_t,\theta_0)^\intercal e_j \right] = \frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \left\| B_{T,\mathbb{A},t}(X_t,\theta_0)^\intercal e_j \right\|^2 \right] > 0, \end{align*} for all $T \ge 1$, since all summands are non-negative and, invoking (ref), at least one summand must be strictly positive, which shows that the matrix \( \Delta_{T,\mathbb{A}}^{-1} \Sigma_{T,\mathbb{A}} \Delta_{T,\mathbb{A}}^{-1} - \Lambda_T^{-1} \) has at least one strictly positive eigenvalue, which concludes the proof of the theorem.
proof[Proof of Theorem (ref)] Under Assumptions (ref), (ref) and (ref), DFZ_CharMest and FisslerZiegel2016 yield that any consistent M-estimator of semiparametric double quantile models is based on classical (strictly) consistent loss functions for the pair of two quantiles, given in ((ref)). Furthermore, the M- and Z-estimator have identical asymptotic covariance if and only if the moment conditions of the Z- and derivative of the loss of the M-estimator coincide, or, respectively, their conditional expectations coincide, see the discussion after (ref). Thus, in the following we compare whether the derivatives of any strictly consistent loss function given in ((ref)) can attain the efficient moment conditions of the Z-estimator almost surely. We get that all identification functions which correspond to an M-estimator (in the form of a derivative of the conditional expectation almost surely) are given by \begin{align*} \psi_{g_{1,t},g_{2,t}}(Y_t,X_t,\theta) = \begin{pmatrix} \nabla_{\theta^\alpha} q_\alpha(X_t,\theta^\alpha) g_{1,t}'\big(q_\alpha(X_t,\theta^\alpha)\big) \big( \mathds{1}_{\{ Y_t \le q_\alpha(X_t,\theta^\alpha)\}} - \alpha \big) \\ \nabla_{\theta^\beta} q_\beta(X_t,\theta^\beta) g_{2,t}'\big(q_\beta(X_t,\theta^\beta)\big) \big( \mathds{1}_{\{ Y_t \le q_\beta(X_t,\theta^\beta)\}} - \beta \big) \end{pmatrix}, \end{align*} which can be written as $\psi_{g_{1,t},g_{2,t}}(Y_t,X_t,\theta) = A_{g_{1,t},g_{2,t}} (X_t,\theta) \varphi\big(Y_t, m(X_t,\theta) \big)$, where \begin{align} A_{g_{1,t},g_{2,t}} (X_t,\theta) = \begin{pmatrix} \nabla_{\theta^\alpha} q_\alpha(X_t,\theta^\alpha) g_{1,t}'\big(q_\alpha(X_t,\theta^\alpha)\big) & 0 \\ 0 & \nabla_{\theta^\beta} q_\beta(X_t,\theta^\beta) g_{2,t}'\big(q_\beta(X_t,\theta^\beta)\big) \end{pmatrix}. \end{align} We start by showing statement (ref), assuming that the Z-estimation efficiency bound is attained by the M-estimator. From Theorem (ref), part (ref) and (ref), we get that the efficient instrument choice is given by $A^\ast_{t,C}(X_t,\theta_0) = C D_t(X_t,\theta_0)^\intercal S_t(X_t,\theta_0)^{-1}$, at the true parameter $\theta_0$, where $C$ is some deterministic and nonsingular matrix and where $D_t(X_t,\theta_0)$ and $S_t(X_t,\theta_0)$ are given in (ref). Furthermore, Theorem (ref) part (ref) shows that any choice of $A_t(X_t,\theta_0)$ which deviates from $A^\ast_{t,C}(X_t,\theta_0)$ (at the true parameter $\theta_0$) with positive probability for some $t \in \mathbb{N}$, cannot attain the efficiency bound. Thus, in the following we show by contradiction that the general instrument matrix of the M-estimator, $A_{g_{1,t},g_{2,t}}(X_t,\theta_0)$, given in ((ref)), cannot attain the necessary form $A^\ast_{t,C}(X_t,\theta_0)$ at the true parameter $\theta_0$ with probability one for any deterministic matrix $C$. For this, we assume that there exists a deterministic and non-singular $q \times q$ matrix $C$ and functions $g_{1,t}$ and $g_{2,t}$ such that $A^\ast_{t,C}(X_t,\theta_0) = A_{g_{1,t},g_{2,t}}(X_t,\theta_0)$ almost surely for all $t \in \mathbb{N}$. We split $C = \left( \begin{smallmatrix} C_{11} & C_{12} \\ C_{21} & C_{22} \end{smallmatrix} \right)$ in its respective parts, where $C_{11} \in \mathbb{R}^{q_1 \times q_1}$, $C_{22} \in \mathbb{R}^{q_2 \times q_2}$, and $C_{12}, C_{21}^\intercal \in \mathbb{R}^{q_1 \times q_2}$. Then, the equation $A^\ast_{t,C}(X_t,\theta_0) = A_{g_{1,t},g_{2,t}}(X_t,\theta_0)$ is equivalent to \begin{align} \begin{aligned} &\begin{pmatrix} \alpha (1-\alpha) g_{1,t}'\big(q_\alpha(X_t,\theta_0^\alpha)\big) \nabla_{\theta^\alpha} q_\alpha(X_t,\theta_0^\alpha) & \alpha (1-\beta) g_{1,t}'\big(q_\alpha(X_t,\theta_0^\alpha)\big) \nabla_{\theta^\alpha} q_\alpha(X_t,\theta_0^\alpha) \\ \alpha (1-\beta) g_{2,t}'\big(q_\beta(X_t,\theta_0^\beta)\big) \nabla_{\theta^\beta} q_\beta(X_t,\theta_0^\beta) & \beta (1-\beta) g_{2,t}'\big(q_\beta(X_t,\theta_0^\beta)\big) \nabla_{\theta^\beta} q_\beta(X_t,\theta_0^\beta) \end{pmatrix} \\ & = \begin{pmatrix} f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) C_{11} \nabla_{\theta^\alpha} q_\alpha(X_t,\theta_0^\alpha) & f_t\big(q_\beta(X_t,\theta_0^\beta)\big) C_{12} \nabla_{\theta^\beta} q_\beta(X_t,\theta_0^\beta) \\ f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) C_{21} \nabla_{\theta^\alpha} q_\alpha(X_t,\theta_0^\alpha) & f_t\big(q_\beta(X_t,\theta_0^\beta)\big) C_{22} \nabla_{\theta^\beta} q_\beta(X_t,\theta_0^\beta) \end{pmatrix}, \end{aligned} \end{align} which must hold element-wise for all four sub-components. Equality of the upper left component yields that there is some $A\in \mathcal A$ with $\mathbb P(A)=1$ such that \begin{align} \xi_t(\omega) \cdot \nabla_{\theta^\alpha} q_\alpha\big(X_t(\omega),\theta_0^\alpha\big) = C_{11} \cdot \nabla_{\theta^\alpha} q_\alpha\big(X_t(\omega),\theta_0^\alpha\big), \quad \forall \omega\in A \end{align} for the scalar random variable $\xi_t := \alpha (1-\alpha) g_{1,t}'\big(q_\alpha(X_t,\theta_0^\alpha)\big)/f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big)$. Equation ((ref)) is an eigenvalue problem for the deterministic matrix $C_{11}$ with stochastic eigenvalues $\xi_t(\omega)$ and eigenvectors $ \nabla_{\theta^\alpha} q_\alpha(X_t(\omega),\theta_0^\alpha)$, $\omega\in A$. We now show that this equation only holds if $\xi_t$ is constant on $A$. By Assumption (ref), there are $\omega_1, \ldots, \omega_{q_1+1}\in A$ such that for $v_\ell:= \nabla_{\theta^\alpha} q_\alpha\big(X_t(\omega_\ell),\theta_0^\alpha\big)$, $\ell\in\{1, \ldots, q_1+1\}$, any subset of cardinality $q_1$ of $\{v_1, \ldots, v_{q_1+1}\}$ is linearly independent. As $C_{11}$ is a deterministic $q_1 \times q_1$ matrix, it can have at most $q_1$ different eigenvalues. Let $\lambda_1, \dots, \lambda_{q_1}$ be the eigenvalues of $C_{11}$ (not necessarily different, thus counted multiple times for higher algebraic multiplicities) ordered such that $v_i$ is an eigenvector for eigenvalue $\lambda_i$ for all $i = 1,\dots,q_1$. Invoking that $v_1,\dots,v_{q_1}$ are linearly independent, it holds that $ \sum_{ \lambda \in \{ \lambda_1,\dots,\lambda_{q_1} \} } \operatorname{dim}(E_\lambda) = q_1, $ where the summation ignores repetitions in the set $\{ \lambda_1,\dots,\lambda_{q_1} \}$ and where $E_\lambda$ denotes the eigenspace corresponding to eigenvalue $\lambda$. The eigenvector $v_{q_1+1}$ must be contained in $E_{\lambda_i}$ for some $i=1,\dots,q_1$ as otherwise, the sum of the geometric multiplicities would exceed $q_1$. If $\operatorname{dim}(E_{\lambda_i}) = l < q_1$, then $E_{\lambda_i}$ is spanned by $l$ elements of $\{ v_1,\dots,v_{q_1}\}$, and as $v_{q_1+1}$ is contained in $E_{\lambda_i}$, these $l$ elements of $\{ v_1,\dots,v_{q_1}\}$ then must be linearly dependent together with $v_{q_1+1}$. This contradicts Assumption (ref). Thus, $\operatorname{dim}(E_{\lambda_i}) = q_1$ and consequently, the geometric multiplicity of $\lambda_i$ is $q_1$, which then must equal the algebraic multiplicity. Hence, all eigenvalues of $C_{11}$ are equal, $\lambda_1 = \cdots = \lambda_{q_1}$, and consequently, $\xi_t$ is constant on $A$, implying that it is constant almost surely. This implies that $g_{1,t}'\big(q_\alpha(X_t,\theta_0^\alpha) \big) = c_2 f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big)$ almost surely for some constant $c_2 > 0$ and for all $t \in \mathbb{N}$, i.e., ((ref)). An analogous proof for the lower right entry of ((ref)) shows ((ref)), which concludes the proof of (ref). For (ref) we start with the “only if” direction assuming that the M-estimator attains the efficiency bound. From part (ref), we already obtain that ((ref)) and ((ref)) must hold. Exploiting $\nabla_{\theta^\alpha} q_\alpha(X_t,\theta_0^\alpha) = \nabla_{\theta^\beta} q_\beta(X_t,\theta_0^\beta)$ and $g_{1,t}'\big(q_\alpha(X_t,\theta_0^\alpha)\big) = c_2 f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big)$, the upper right component of ((ref)) implies that \begin{align} \frac{ \alpha (1-\beta) c_2 f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big)}{f_t\big(q_\beta(X_t,\theta_0^\beta)\big)} \cdot \nabla_{\theta^\alpha} q_\alpha(X_t,\theta_0^\alpha) = C_{12} \cdot \nabla_{\theta^\alpha} q_\alpha(X_t,\theta_0^\alpha), \end{align} almost surely. Applying the same eigenvalue argument to ((ref)) (recalling that $\nabla_{\theta^\alpha} q_\alpha(X_t,\theta_0^\alpha) = \nabla_{\theta^\beta} q_\beta(X_t,\theta_0^\beta)$ implies that $q_1=q_2$ such that $C_{12}$ is quadratic) yields ((ref)). For the “if” implication in (ref), we assume that ((ref)), ((ref)) and ((ref)) hold. We choose $C_{11} = \alpha (1-\alpha )c_2 I_{q_1 \times q_1}$, $C_{12} = \alpha (1-\beta) c_1 c_2 I_{q_1 \times q_2}$, $C_{21} = \alpha (1-\beta) c_3/c_1 I_{q_2 \times q_1}$ and $C_{22} = \beta (1-\beta) c_3 I_{q_2 \times q_2}$, where $\operatorname{det}(C) \not= 0$ follows from $0<\alpha<\beta<1$. Thus, straightforward calculations yield that $A_{g_{1,t},g_{2,t}} (X_t,\theta_0) = A^\ast_{t,C}(X_t,\theta_0)$ holds almost surely for all $t \in \mathbb{N}$. Applying Theorem (ref) yields the claim.
proof[Proof of Theorem (ref)] This proof follows the general ideas of the proof of Theorem (ref). Under Assumptions (ref), (ref) and (ref), DFZ_CharMest and FisslerZiegel2016 yield that any consistent M-estimator of joint quantile and ES models is based on classical (strictly) consistent loss functions given in ((ref)). Thus, in the following we analyze whether the derivatives of any consistent loss function are able to match the efficient moment conditions of Theorem (ref) almost surely. We get that all identification functions which correspond to an M-estimator (in the form of a derivative almost surely) are given by $\psi_{g_t,\phi_t} \big( Y_t, X_t, \theta \big)$ equalling \[ \begin{pmatrix} \nabla_{\theta^q} q_\alpha(X_t,\theta^q) \Big( g_t'\big(q_\alpha(X_t,\theta^q)\big) + \phi_t'\big(e_\alpha(X_t,\theta^e)\big)/\alpha \Big) \left( \mathds{1}_{\{Y_t \le q_\alpha(X_t,\theta^q)\}} - \alpha \right) \\[0.3em] \nabla_{\theta^e} e_\alpha(X_t,\theta^e) \phi_t''\big(e_\alpha(X_t,\theta^e)\big) \left( e_\alpha(X_t,\theta^e) - q_\alpha(X_t,\theta^q) + \frac{1}{\alpha}(q_\alpha(X_t,\theta^q) - Y_t) \mathds{1}_{\{Y_t \le q_\alpha(X_t,\theta^q)\}} \right) \end{pmatrix}. \] This implies that the moment conditions corresponding to an M-estimator can be written as $\psi_{g_t,\phi_t} \big( Y_t, X_t, \theta \big) = A_{g_t,\phi_t}(X_t,\theta) \varphi \big( Y_t, m(X_t,\theta) \big)$, where $\varphi\big( Y_t, m(X_t,\theta) \big)$ is given in ((ref)), and $A_{g_t,\phi_t}(X_t,\theta)$ is a diagonal matrix with entries $\Big( g_t'\big(q_\alpha(X_t,\theta^q)\big) + \phi_t'\big(e_\alpha(X_t,\theta^e)\big)/\alpha \Big) \times \nabla_{\theta^q} q_\alpha(X_t,\theta^q)$ and $ \phi_t''\big(e_\alpha(X_t,\theta^e)\big) \nabla_{\theta^e} e_\alpha(X_t,\theta^e)$. To show (ref), we assume that the Z-estimation efficiency bound is attained by the M-estimator. From Theorem (ref), we get that the efficient estimator has to fulfill the condition $A^\ast_{t,C}(X_t,\theta_0) = C D_t(X_t,\theta_0)^\intercal S_t(X_t,\theta_0)^{-1}$ for some deterministic and nonsingular matrix $C$, where $D_t(X_t,\theta_0)$ and $S_t(X_t,\theta_0)$ are given in (ref) and (ref). Thus, we verify whether there exists a deterministic and non-singular $q \times q$ matrix $C$ (and appropriate functions $g_t$ and $\phi_t$) such that $A^\ast_{t,C}(X_t,\theta_0) = A_{g_t,\phi_t}(X_t,\theta_0)$ almost surely, i.e., whether $C D_t(X_t,\theta_0)^\intercal = A_{g_t,\phi_t}(X_t,\theta_0) S_t(X_t,\theta_0)$ holds almost surely. By splitting the matrix $C = \begin{pmatrix} C_{11} & C_{12} \\ C_{21} & C_{22} \end{pmatrix}$ in its respective parts, where $C_{11} \in \mathbb{R}^{q_1 \times q_1}$, $C_{22} \in \mathbb{R}^{q_2 \times q_2}$, and $C_{12}, C_{21}^\intercal \in \mathbb{R}^{q_1 \times q_2}$, this simplifies to the following four equalities, \begin{align} C_{11} \nabla_{\theta^q} q_\alpha(X_t,\theta^q_0) = (1-\alpha) \frac{\alpha g_t'\big(q_\alpha(X_t,\theta^q_0)\big) + \phi_t'\big(e_\alpha(X_t,\theta^e_0)\big)}{ f_t\big(q_\alpha(X_t,\theta^q_0)\big)} \nabla_{\theta^q} q_\alpha(X_t,\theta^q_0), \\ \begin{aligned} C_{12} \nabla_{\theta^e} e_\alpha(X_t,\theta^e_0) &= (1-\alpha) \big( q_\alpha(X_t,\theta^q_0) - e_\alpha(X_t,\theta^e_0) \big) \\ &\qquad \times \left( g_t'\big(q_\alpha(X_t,\theta^q_0)\big) + \phi_t'\big(e_\alpha(X_t,\theta^e_0)\big)/\alpha \right) \nabla_{\theta^q} q_\alpha(X_t,\theta^q_0), \end{aligned} \\ C_{21} \nabla_{\theta^q} q_\alpha(X_t,\theta^q_0) = \frac{(1-\alpha) \big( q_\alpha(X_t,\theta^q_0) - e_\alpha(X_t,\theta^e_0) \big) \phi_t”\big(e_\alpha(X_t,\theta^e_0)\big) }{f_t\big(q_\alpha(X,\theta^q_0)\big) } \nabla_{\theta^e} e_\alpha(X_t,\theta^e_0), \\ \begin{aligned} &C_{22} \nabla_{\theta^e} e_\alpha(X_t,\theta^e_0) = \phi_t”\big(e_\alpha(X_t,\theta^e_0)\big) \nabla_{\theta^e} e_\alpha(X_t,\theta^e_0) \\ &\quad \times \left( \frac{1}{\alpha} \operatorname{Var}_t \big( Y_t \big| Y_t \le q_\alpha(X_t,\theta^q_0) \big) + \frac{1-\alpha}{\alpha} \big( e_\alpha(X_t,\theta^e_0) - q_\alpha(X_t,\theta^q_0) \big)^2\right), \end{aligned} \end{align} which have to hold almost surely. Using the same eigenvalue argument as in the proof of Theorem (ref), equation ((ref)) implies that \begin{align} (1 - \alpha) \big( \alpha g_t'\big(q_\alpha(X_t,\theta^q_0)\big) + \phi_t'\big(e_\alpha(X_t,\theta^e_0)\big) \big) = \tilde c_1 f_t\big(q_\alpha(X_t,\theta^q_0)\big) \end{align} almost surely for some constant $\tilde c_1 > 0$. Equation ((ref)) follows by setting $c_6 = \tilde c_1/(\alpha(1-\alpha))$. Similarly, ((ref)) implies that \begin{align} \frac{\tilde c_2}{\phi_t”\big(e_\alpha(X_t,\theta^e_0)\big)} = \frac{1}{\alpha} \operatorname{Var}_t \big( Y_t \big| Y_t \le q_\alpha(X_t,\theta^q_0) \big) + \frac{1-\alpha}{\alpha} \big( e_\alpha(X_t,\theta^e_0) - q_\alpha(X_t,\theta^q_0) \big)^2 \end{align} almost surely for some constant $\tilde c_2 >0$. Furthermore, combining ((ref)) and ((ref)) implies \begin{align*} &C_{12} C_{21} \nabla q_\alpha(X_t,\theta^q_0) = \nabla q_\alpha(X_t,\theta^q_0) (1-\alpha)^2 \big/ \alpha \\ &\times \frac{\big( q_\alpha(X_t,\theta^q_0) - e_\alpha(X_t,\theta^e_0) \big)^2 \phi_t”\big(e_\alpha(X_t,\theta^e_0)\big) \left( \alpha g_t'\big(q_\alpha(X_t,\theta^q_0)\big) + \phi_t'\big(e_\alpha(X_t,\theta^e_0)\big) \right)}{ f_t\big(q_\alpha(X_t,\theta^q_0)\big)} \end{align*} almost surely and employing the same eigenvalue argument again yields that \begin{align} \begin{aligned} &\big( q_\alpha(X_t,\theta^q_0) - e_\alpha(X_t,\theta^e_0) \big)^2 \phi_t”\big(e_\alpha(X_t,\theta^e_0)\big) \left( \alpha g_t'\big(q_\alpha(X_t,\theta^q_0)\big) + \phi_t'\big(e_\alpha(X_t,\theta^e_0)\big) \right) \\ &= \frac{\tilde c_3 \alpha}{(1-\alpha)^2} f_t\big(q_\alpha(X_t,\theta^q_0)\big) \end{aligned} \end{align} almost surely for some constant $\tilde c_3 > 0$. Substituting ((ref)) and ((ref)) into ((ref)) finally yields \begin{align} \operatorname{Var}_t \big( Y_t \big| Y_t \le q_\alpha(X_t,\theta^q_0) \big) = (1-\alpha) \left( \frac{ \tilde c_1 \tilde c_2}{ \tilde c_3} - 1 \right) \big( q_\alpha(X_t,\theta^q_0) - e_\alpha(X_t,\theta^e_0) \big)^2. \end{align} almost surely. By defining the constant $c_1 := (1-\alpha) \left( \frac{ \tilde c_1 \tilde c_2}{ \tilde c_3} - 1 \right)$, we obtain ((ref)), where the positivity of $c_1$ follows from that fact that both sides of (ref) are positive. Substituting ((ref)) into ((ref)) yields ((ref)), which concludes the proof of statement (ref). For (ref) we start with the “only if” direction, assuming that the M-estimator attains the efficiency bound. From part (ref) we obtain that (ref), (ref) and (ref) must hold. Employing the same eigenvalue argument as before, we obtain from ((ref)) that $ \tilde c_4 f_t\big(q_\alpha(X_t,\theta^q_0)\big)/(1-\alpha) = \big(q_\alpha(X_t,\theta^q_0) - e_\alpha(X_t,\theta^e_0) \big)\phi_t''\big(e_\alpha(X_t,\theta^e_0)\big) $ for some constant $\tilde c_4>0$, where we additionally exploited that $\nabla_{\theta^q} q_\alpha(X_t,\theta_0^q) = \nabla_{\theta^e} e_\alpha(X_t,\theta_0^e)$ almost surely. Combining (ref) and (ref) yields $\phi_t''\big(e_\alpha(X_t,\theta^e_0)\big) = c_3/c_1 \big( q_\alpha(X_t,\theta^q_0) - e_\alpha(X_t,\theta^e_0) \big)^{-2}$, which in turn leads us to \begin{align} f_t\big(q_\alpha(X_t,\theta^q_0)\big) = \frac{c_2}{q_\alpha(X_t,\theta^q_0) - e_\alpha(X_t,\theta^e_0)}, \end{align} almost surely for where $c_2 =( 1-\alpha) c_3/(c_1 \tilde c_4) > 0$, establishing (ref). Using again that $\phi_t''\big(e_\alpha(X_t,\theta^e_0)\big) = c_3/c_1 \big( q_\alpha(X_t,\theta^q_0) - e_\alpha(X_t,\theta^e_0) \big)^{-2}$ and since the support of $e_\alpha(X_t,\theta^e_0)$ is a non-degenerate interval by assumption, it must hold that \begin{align} \phi_t'\big(e_\alpha(X_t,\theta^e_0)\big) = \frac{c_3/c_1}{ q_\alpha(X_t,\theta^q_0) - e_\alpha(X_t,\theta^e_0)} + \tilde c_{5,t}, \end{align} almost surely for some deterministic, but possibly time-varying constant $\tilde c_{5,t} \in \mathbb{R}$ for all $t \in \mathbb{N}$. Combining ((ref)), ((ref)) and ((ref)) yields that \begin{align} g_t'\big(q_\alpha(X_t,\theta^q_0)\big) = c_4 f_t\big(q_\alpha(X_t,\theta^q_0)\big) + c_{5,t}, \end{align} \color{black} where $c_4 := \left( \frac{\tilde c_1}{\alpha (1 - \alpha) } - \frac{c_3}{\alpha c_1 c_2} \right) \in \mathbb{R}$ and $c_{5,t} := - \tilde c_{5,t}/\alpha$, which establishes (ref). Eventually, employing ((ref)) and ((ref)) yields that $ \phi_t'\big(e_\alpha(X_t,\theta^e_0)\big)/\alpha = \tilde c_1f_t\big(q_\alpha(X_t,\theta^q_0)\big)/(\alpha (1-\alpha)) - c_4 f_t\big(q_\alpha(X_t,\theta^q_0)\big) - c_{5,t}, $ and hence $\phi_t'\big(e_\alpha(X_t,\theta^e_0)\big) = c_3 f_t\big(q_\alpha(X_t,\theta^q_0)\big)/(c_1 c_2) - \alpha c_{5,t}, $ which shows (ref) and concludes this direction. For the “if” implication in statement (ref), we assume the conditions ((ref))\,--\,((ref)). Choosing $C_{11} = \alpha (1-\alpha) \left( c_4 + \frac{c_3}{\alpha c_1 c_2} \right) I_{q_1 \times q_1}$, $C_{12} = \frac{(1-\alpha)}{\alpha c_1 c_3^2} \big( \alpha c_1 c_2 c_4 + c_3 \big) I_{q_1 \times q_2}$, $C_{21} = \frac{(1-\alpha) c_3}{c_1 c_2} I_{q_2 \times q_1}$ and $C_{22} = \frac{c_1 + 1 - \alpha}{\alpha c_1 c_3} I_{q_2 \times q_2}$, automatically yields $\operatorname{det}(C) \not= 0$, and straight-forward calculations yield that ((ref))\,--\,((ref)) are satisfied and thus, the identity $A_{g_t, \phi_t} (X_t,\theta_0) = A^\ast_{t,C}(X_t,\theta_0)$ holds almost surely for all $t \in \mathbb{N}$. Applying Theorem (ref) yields the claim.

Details on the connection between loss and identification functions

Multivariate functionals

Section (ref) in the main paper provided the arguments why, roughly speaking, there are “more” strict identification functions than strictly consistent loss functions for multivariate functionals. Interestingly, in extreme cases, e.g., for prediction intervals symmetric around the mean or the median, multivariate functionals may admit strict identification functions, but may even fail to be elicitable at all, due to the said integrability conditions FisslerHlavinovaRudloff2019Theory. By virtue of the arguments DFZ_CharMest, this gap in turn induces a gap between the classes of consistent M- and Z-estimators, which is illustrated by the following Remark (ref). Note that in order to use Osband's principle in Remark (ref), and in line with the discussion in NeweyMcFadden1994, it is sufficient to assume that the conditional expectation $\mathbb{E}\big[\rho\big(Y_t,m(X_t, \theta)\big)\big|X_t\big]$ is differentiable in $\theta$ almost surely. This allows us to treat also losses that are per se not differentiable, such as the pinball loss.

remLet Assumption (ref) hold for some $k$-dimensional functional $\Gamma\colon \mathcal{F}_{\mathcal{Y}|\mathcal{X}}\to\Xi$ with a strict $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-identification function $\varphi\colon \mathbb{R}\times\Xi\to\mathbb{R}^k$ and a strictly $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent loss function $\rho \colon \mathbb{R}\times \Xi\to\mathbb{R}$. Suppose that $\mathbb{E}\big[\rho\big(Y_t,m(X_t, \theta)\big)\big|X_t\big]$ is differentiable in $\theta$ almost surely. Under the richness conditions on $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ of FisslerZiegel2016, we then have \begin{equation} \nabla_\theta \mathbb{E}\big[\rho\big(Y_t,m(X_t, \theta)\big)\big|X_t\big] = \nabla_\theta m(X_t, \theta)^\intercal \cdot h\big(m(X_t,\theta)\big)\cdot\mathbb{E}\big[\varphi\big(Y_t,m(X_t, \theta)\big)\big|X_t\big] \,, \end{equation} where $h$ takes values in $\mathbb{R}^{k\times k}$ and the gradient $\nabla_\theta m(X_t,\theta)$ is in $\mathbb{R}^{k\times q}$. Comparing (ref) with (ref), one obtains the identity $A(X_t,\theta) = \nabla_\theta m(X_t, \theta)^\intercal \cdot h\big(m(X_t,\theta)\big)$ for the instrument matrix. The presence of $h$ and the limitations of the choice of $h$ discussed above yield that there are considerably fewer (strict) model-consistent losses than (strict) moment functions.

Univariate functionals

If $\Gamma$ is univariate, its mixture-continuity implies that every strictly consistent loss $\rho$ is (strictly) order-sensitive meaning that $\xi\mapsto \bar \rho(F,\xi)$ is (strictly) decreasing (increasing) for $\xi\le \Gamma(F)$ (for $\xi\ge \Gamma(F)$); see Nau1985, Lambert2013 and BelliniBignozzi2015. The functional $\Gamma$ is mixture-continuous if for all $F_0, F_1\in\mathcal{F}$ such that $(1-\lambda)F_0 + \lambda F_1\in\mathcal{F}$ for all $\lambda \in[0,1]$ the function $[0,1]\ni \lambda\mapsto \Gamma((1-\lambda)F + \lambda G)$ is continuous. Therefore, the derivative of $\rho$ is an oriented identification function in the sense that $\nabla_\xi \bar \rho(F,\xi)\le 0$ ($\ge0$) if $\xi\le \Gamma(F)$ ($\ge\Gamma(F)$). Intuitively, this excludes the existence of additional local minima of the expected loss, while possible saddle points still remain an issue. Moreover, Osband's principle in dimension one (ref). Even if $\rho$ is strictly consistent, its derivative is not necessarily a strict identification function due to possible saddle points of the expected loss. That means even if $\varphi$ is strict, $h$ might vanish at some points, see SteinwartPasinETAL2014 and NeweyMcFadden1994 for further details and examples. If $\varphi$ is oriented and strict, then $h$ is non-negative. This means, on the other hand, that we can also start with an oriented strict identification function, multiply it with any positive $h$, and integrate it. This results in a strictly order-sensitive, and therefore, strictly consistent loss SteinwartPasinETAL2014. If $\varphi$ is not strict and $h$ simply non-negative, the resulting loss is merely consistent. This leads to the fact that there is a one-to-one relation between consistent losses and oriented identification functions for $\Gamma$.

remIf the identification function fails to be oriented, it can still be integrated, but does not yield a consistent score. E.g., $\varphi(y,\xi) = (\mathds{1}\{\xi\ge0\} - \mathds{1}\{\xi<0\})(\xi-y)$ is a strict identification function for the mean which fails to be oriented. It is easy to check that the integral $\rho(y,\xi) = \int_0^\xi \varphi(y,z)\,\mathrm{d} z $ is not a strictly consistent loss function for the mean. This identification function constitutes a counterexample to SteinwartPasinETAL2014.

Semiparametric Models for Multiple Moments

We consider joint semiparametric models for the first and second moments, denoted by $\Gamma_{\textrm{mom}}$, and closely related, joint models for mean and variance, $\Gamma_{(\mathbb{E}, \operatorname{Var})}$. Since mean and variance are considered as the most important functionals in classical statistics, the related class of ARMA-GARCH models Bollerslev1986 is omnipresent in the econometric literature, and is often estimated through M- or Z-estimation. See also Spady2018 for joint mean--variance regression models.

assLet $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$ contain all square integrable distributions such that $(\mathbb{E}, \operatorname{Var})$ maps surjectively on $\Xi\subseteq\mathbb{R}^2$, which is assumed to be an open and path connected subset of $\mathbb{R}^2$.

Assumption (ref) is required to characterize the classes of consistent loss and identification functions. Recalling that $\Gamma_{\textrm{mom}}$ and $\Gamma_{(\mathbb{E}, \operatorname{Var})}$ are in bijection, we can invoke the revelation principle Osband1985, Gneiting2011 to relate the corresponding strict $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-identification and strictly $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent loss functions. Strict identification functions are given by

align[align omitted — 293 chars of source]

Given Assumptions (ref) and (ref), DFZ_CharMest yields that the full class of consistent M-estimators at (ref) is determined by the full class of (strictly) $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent loss functions. Using the revelation principle and following FisslerZiegel2016, under Assumption (ref), the class of all differentiable (strictly) $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-consistent loss functions is given by

align[align omitted — 597 chars of source]

where $\phi_t: \{(\xi_1, \xi_2)\in\mathbb{R}^2\,|\, \xi_1^2 \le \xi_2\} \to \mathbb{R}$ are (strictly) convex and twice differentiable functions with gradient $\nabla \phi_t$, and $\kappa_t$ is an $\mathcal{F}_{\mathcal{Y}|\mathcal{X}}$-integrable function. For any sequence $\Phi = (\phi_t)_{t\in\mathbb N}$ of such functions, we denote the corresponding M-estimators defined via (ref) by $\widehat \theta^{\textrm{mom}}_{M,T,\Phi}$ and $\widehat \theta^{(\mathbb{E}, \operatorname{Var})}_{M,T,\Phi}$.

propositionUnder Assumptions (ref), (ref) together with Assumptions (ref) and (ref) in Appendix (ref), suppose that the M-estimators $\widehat \theta^{\textrm{mom}}_{M,T,\Phi}$ and $\widehat \theta^{(\mathbb{E}, \operatorname{Var})}_{M,T,\Phi}$ for the first two moments and for $(\mathbb{E}, \operatorname{Var})$ are asymptotically normal. If almost surely \begin{align} \phi_t(z) = \frac{1}{2} z^\intercal \boldsymbol{\operatorname{Var}}_t \left( \big( Y_t, Y_t^2 \big) \right)^{-1} z \qquad \forall t\in\mathbb N, \end{align} then these M-estimators attain the corresponding Z-estimation efficiency bounds in (ref).
proof[Proof of Proposition (ref)] We first consider the case of the double moment functional. Straight-forward calculations yield that the class of identification functions corresponding to the M-estimators based on loss functions given in ((ref)) is given by \begin{align} \psi_{\phi_t}(Y_t,X_t,\theta) = A_{\phi_t}(X_t,\theta) \cdot \varphi_{mom}\big(Y_t, m(X_t,\theta) \big), \end{align} where $\varphi$ is given in ((ref)) and $ A_{\phi_t}(X_t,\theta) = \begin{pmatrix} \nabla_{\theta} m_1(X_t,\theta) & \nabla_{\theta} m_2(X_t,\theta) \end{pmatrix} \cdot \nabla^2 \phi_t\big( m(X_t,\theta) \big).$ Applying Theorem (ref) yields that the efficiency bound can be attained by a Z-estimator (and for equivalent M-estimators) if and only if $A^\ast_{t,C}(X_t,\theta) = C D_t(X_t,\theta)^\intercal S_t(X_t,\theta)^{-1}$ almost surely, where $C$ is some deterministic and non-singular matrix, and where \begin{align*} S_t(X_t,\theta_0) = \boldsymbol{\operatorname{Var}}_t \big( (Y_t, Y_t^2 ) \big), \qquad and \qquad D_t(X_t,\theta_0) = \begin{pmatrix} \nabla_{\theta} m_1(X_t,\theta_0)^\intercal \\ \nabla_{\theta} m_2(X_t,\theta_0)^\intercal \end{pmatrix}. \end{align*} By choosing $C = I_q$ and the strictly convex quadratic form $\phi_t(z) = \frac{1}{2} z^\intercal \boldsymbol{\operatorname{Var}}_t \big( (Y_t, Y_t^2 ) \big)^{-1} z$ for all $t \in \mathbb{N}$ and for all $z \in \mathbb{R}^2$, this yields that $\nabla^2 \phi_t \big( m(X_t,\theta_0) \big) = \boldsymbol{\operatorname{Var}}_t \big( (Y_t, Y_t^2 ) \big)^{-1}$ almost surely. Consequently, the M-estimator for the double moment regression is able to attain the efficient instrument matrix $A^\ast_{t,C}(X_t,\theta_0)$ (at $\theta_0$) and consequently the Z-estimation efficiency bound. For the situation of mean and variance, (ref) takes the form \( \psi_{\phi}(Y_t,X_t,\theta) = \tilde A_{t,\phi}(X_t,\theta) \cdot \varphi_{(\mathbb{E}, \operatorname{Var})}\big(Y_t, m(X_t,\theta) \big), \) where $\varphi$ is given in ((ref)) and where \begin{align*} \tilde A_{t,\phi}(X_t,\theta) = \begin{pmatrix} \nabla_\theta m_1(X_t,\theta)^\intercal \\ \nabla_\theta v(X_t,\theta)^\intercal + 2 m_1(X_t,\theta) \nabla_\theta m_1(X_t,\theta)^\intercal \end{pmatrix}^\intercal \cdot \nabla^2 \phi_t \begin{pmatrix} m_1(X_t,\theta) \\ v(X_t,\theta) + m_1^2(X_t,\theta) \end{pmatrix}. \end{align*} Straight-forward calculations yield that $S_t(X_t,\theta_0) = \boldsymbol{\operatorname{Var}}_t \big( (Y_t, Y_t^2 ) \big)$ and \[ D_t(X_t,\theta_0) = \begin{pmatrix} \nabla_{\theta} m_1(X_t,\theta_0)^\intercal \\ \nabla_{\theta} v (X_t,\theta_0)^\intercal + 2 m_1(X_t,\theta_0) \nabla_{\theta} m_1(X_t,\theta_0)^\intercal \end{pmatrix}. \] Thus, upon using $\mathbb{R}^2\ni z \mapsto \phi_t(z) = \frac{1}{2} z^\intercal \boldsymbol{\operatorname{Var}}_t \big( (Y_t, Y_t^2 ) \big)^{-1} z$, the efficient choice can be attained.

This result is in line with the classical univariate mean regression, where both, M- and Z-estimators are able to attain the Z-estimation efficiency bound and the most efficient Bregman loss is given by the squared loss, weighted with the inverse of the conditional variance. Intuitively, this attainability can be explained by the fact that the classes of strictly consistent joint loss functions given in (ref) are relatively large due to the presence of the general convex function $\phi_t$, being a function in two arguments.

For the first two moments, this can be illustrated by comparing it to a minimal subclass in this context, namely the class only consisting of the sum of (strictly) consistent loss functions for the individual components, the first and second moment. This arises from (ref) when $\phi_t$ takes the additive form $\phi_{\textrm{add}, t}(\xi_1,\xi_2) = \phi_{1,t}(\xi_1)+\phi_{2,t}(\xi_2)$, where $\phi_{i,t}$ are both (strictly) convex. Since the Hessian $\nabla^2\phi_{\textrm{add}, t}$ is diagonal, $\phi_{\textrm{add}, t}$ can only take the form in (ref) for the special situation when $Y_t$ and $Y_t^2$ are conditionally uncorrelated. Since the class of convex functions on $\mathbb{R}^2$ is far larger than the sum of two convex functions in the individual components, the efficiency bound can be attained. For the pair of mean and variance, note that one cannot decompose the loss into a sum of strictly consistent losses for each component, due to the variance failing to be elicitable in general. In particular, this also shows the importance of modelling the variance jointly with the mean. However, an additive decomposition of $\phi_t$ as discussed above is also possible for mean and variance.

These results are in stark contrast to the double quantile (DQ) regression framework considered Section (ref), where the gap arises since the class of strictly consistent losses is relatively small, coinciding with the described minimal additive class.

Further implications of the gap: Equivariance properties

Patton2011 and NoldeZiegel2017 provide arguments for the usage of homogeneous loss functions for forecast comparison and ranking. More generally, FisslerZiegel2019 advocate for loss functions that respect equivariance properties of the functional of interest. Besides homogeneity, a major equivariance property of interest is translation equivariance, or---more generally speaking---linear equivariance; see FisslerZiegel2019. Again, we focus on two interesting pairs of functionals, (mean, variance) and $(Q_\alpha, \operatorname{ES}_\alpha)$, $\alpha\in(0,1)$. For any random variable $Y$ with finite second moment and any scalar $c\in\mathbb{R}$, the following identities hold

align[align omitted — 290 chars of source]

Suppose one is to model the functional $(Q_\alpha, \operatorname{ES}_\alpha)$ with a parametric model (possibly with joint model parameters) of the form $m(X,\theta)=\big(q_\alpha(X,\theta), e_\alpha(X,\theta)\big)$, where $\theta= \big(\theta^{(1)}, \ldots, \theta^{(q)}\big)\in\Theta\subseteq\mathbb{R}^q$, with intercept parameters, say \[

pmatrix[pmatrix omitted — 57 chars of source]

=

pmatrix[pmatrix omitted — 164 chars of source]

. \] Then, under Assumption (ref), the correctly specified parameter $\theta_0$ has the following equivariance property for $(Y,X)\in\mathcal{Z}$ and $c\in\mathbb{R}$ such that $(Y+c,X)\in\mathcal{Z}$:

align[align omitted — 211 chars of source]

Similar results apply to the pair (mean, variance), where, of course, the intercept transformation only appears in the mean-component.

Similarly, given data $(\boldsymbol{Y}, \boldsymbol{X}) = (Y_t, X_t)_{t=1, \ldots, T}$, it would be desirable to find a similar translation equivariance property for an estimator $\widehat \theta_T = \widehat \theta_T (\boldsymbol{Y}, \boldsymbol{X} )$:

align[align omitted — 295 chars of source]

Under Assumption (ref) of a correctly specified model, (ref) holds for the probability limit of any consistent estimator. However, in finite samples or under model misspecification, it may well fail unless there is some additional structure in the estimator. For example, the OLS-estimator clearly satisfies (ref) and (ref), relying on the fact that the squared loss $\rho(y,\xi) = \frac12(y-\xi)^2$ is translation invariant. Also, the corresponding Z-estimator is translation equivariant, since the standard identification function $\varphi(y,\xi) = y-\xi$ is translation invariant and the instrument matrix $A(X,\theta) = X$ is independent of $\theta$.

It turns out that both two-dimensional functionals in (ref) possess strict identification functions that respect the respective equivariance properties described there, namely

align*[align* omitted — 354 chars of source]

Using instrument matrices which are independent of $\theta$, they induce Z-estimators which obey the translation equivariance in their intercept components. However, Propositions 4.9 and 4.10 in FisslerZiegel2019 ascertain that for both functional pairs, there are no strictly consistent loss function with these equivariance properties---at least under general and realistic assumptions. This rules out the existence of corresponding M-estimators with this property---another manifestation of the gap between these two classes of estimators.

Details on the Efficiency in Double Quantile Models

Theorem (ref) part (ref), which is based on Theorem (ref) part (ref), merely implies that the difference of the asymptotic covariances between any M-estimator and the joint efficient Z-estimator is positive semi-definite with at least one positive eigenvalue. One could plausibly suspect that this is purely caused by differing off-diagonal “covariance” terms, and that the diagonal entries---i.e., the estimation “variances” of the parameters---coincide (at least for the pseudo-efficient M-estimator). The following example illustrates that the efficiency gap also arises for the diagonal entries.

To simplify the exposition, we consider a stationary process $(Y_t,X_t)_{t\in\mathbb N}$ with a univariate $X_t \in \mathbb{R}$, and two linear “slope only” models $q_\alpha(X_t,\theta^\alpha) = X_t\theta^\alpha$ and $q_\beta(X_t,\theta^\beta) = X_t\theta^\beta$. Using the shorthands $f_\alpha:= f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big)$ and $f_\beta:= f_t\big(q_\beta(X_t,\theta_0^\beta)\big)$, the asymptotic variance of the individual efficient Z-estimator for the $\theta^\alpha$ component is $\alpha(1-\alpha)/\mathbb{E}[f_\alpha^2 X_t^2]$. On the other hand, the asymptotic variance of the joint efficient Z-estimator for the $\theta^\alpha$ component is \[ \frac{\alpha(1-\alpha)}{\mathbb{E}[f_\alpha^2X_t^2]} \; \times \; \frac{\alpha(1-\alpha)\beta(1-\beta) - \alpha^2(1-\beta)^2}{\alpha(1-\alpha)\beta(1-\beta) - \alpha^2(1-\beta)^2 \mathbb{E}[f_\alpha f_\beta X_t^2]^2 /\big(\mathbb{E}[f_\alpha^2 X_t^2]\mathbb{E}[ f_\beta^2 X_t^2]\big)}\,, \] which is generally smaller than \( \alpha(1-\alpha)/\mathbb{E}[f_\alpha^2X_t^2], \) since the Cauchy--Schwartz inequality implies that $\mathbb{E}[f_\alpha f_\beta X_t^2]^2 /\big(\mathbb{E}[f_\alpha^2 X_t^2]\mathbb{E}[ f_\beta^2 X_t^2]\big)\le 1$. The latter holds with equality if and only if $f_\alpha X_t$ and $f_\beta X_t$ are colinear almost surely, once again stressing the importance of condition (ref). This effect can numerically be observed in our simulations in panels A of Table (ref) by comparing the efficiency bound to the pseudo-efficient choices based on the choices $F_t(\xi)$.

The Two-Step Estimation Efficiency Bound

In related work, Barendse2020 considers efficiency among the class of two-step estimators of semiparametric models for the quantile and ES with separated parameters. These two-step estimators utilize a quantile regression to estimate the quantile parameters in the first step and a restricted and weighted least squares estimator in the second step for the model parameters of the conditional ES. The author considers efficiency among the possible estimation weights from the second step weighted least squares estimator, see Barendse2020 for details. This procedure amounts to efficiency of the ES parameters in isolation, which generally results in more restrictive efficiency bounds than efficiency of the joint model parameters considered in this article.

In our notation, the class of two-step estimators can be characterized by the general form (ref), the identification functions in (ref) and the class of instrument matrices

align[align omitted — 261 chars of source]

For these estimators, the theory of ProkhorovSchmidt2009, Bartalotti2013 can be used to establish that the asymptotic distribution of the joint Z- and the two-step estimators coincide. Consequently, the family of two-step estimators of Barendse2020 form a subclass of the general class of Z-estimators we consider in this article. Hence, it follows that the resulting two-step estimation efficiency bound is no smaller than the general Z-estimation efficiency bound of Theorem (ref). While these two bounds can coincide in special situations, they generally do not as illustrated in the following.

For the special case of location-scale models with stationary innovations $(\varepsilon_t)_{t \in \mathbb{N}}$ discussed in Section (ref), the efficient weights of Barendse2020 coincide with the choice of $\phi_t''$ implied by a combination of ((ref)) and ((ref)) in Theorem (ref). This illustrates that for this special case, and in terms of the ES parameters, $\theta^e$, considered in isolation, the efficient two-step and the efficient M-estimator are equally efficient; see Barendse2020. However, if efficiency is considered for the full parameter vector, $\theta$, the two-step estimator using the instrument matrix ((ref)) is generally less efficient, which is caused by the inefficient choice of the first-step standard quantile regression. In this special case, joint efficiency could be guaranteed by employing an efficient quantile regression estimator in the first step, see e.g., Komunjer2010a, Komunjer2010b. We refer to the simulation results of Sections (ref) and (ref) and in particular to Tables (ref) and (ref) for a numerical illustration.

More generally, Barendse2020 illustrates that, taken in isolation, the ES specific asymptotic sub-covariance matrix of the M-estimator $\widehat \theta^e$ is subject to his two-step efficiency bound. However, this does not hold if one considers the entire covariance matrix of the joint model parameters for the quantile and ES. This can be observed by comparing ((ref)) with the efficient instrument matrix $A_t^\ast$ given in ((ref)) and ((ref)): while $A_{t,C}^\ast$ generally requires non-zero off-diagonal blocks, the matrix $A_t^\dagger$ is restricted to a block diagonal matrix with zero off-diagonal blocks.

Recall that under the gradient condition that $\nabla_{\theta^q} q_\alpha(X_t,\theta_0^q) = \nabla_{\theta^e} e_\alpha(X_t,\theta_0^e)$ for all $t \in \mathbb{N}$ almost surely, part (ref) of Theorem (ref) implies that if conditions ((ref)) or ((ref)) fail to hold, the M-estimator cannot attain the Z-estimation efficiency bound. As Barendse2020 informally shows that the two-step estimators are equivalent to the class of M-estimators in terms of the efficiency of the ES parameters, this illustrates that the two-step estimators also cannot attain the Z-estimation efficiency bound in this setting. (Formally, relating $A_t^\dagger(X_t, \theta_0)$ in ((ref)) to the efficient choice $A_t^\ast(X_t, \theta_0)$ and employing Theorem (ref) as in the proof of Theorem (ref) yields the desired result.) Besides supporting our claim of an existing efficiency gap for the joint quantile and ES models, this illustrates that the two-step efficiency bound of Barendse2020 does generally not coincide with the general Z-estimation efficiency bound of Hansen1985, Chamberlain1987, and Newey1993.

We illustrate the theoretical considerations of this section numerically through the simulation setup of Sections (ref) and (ref). In Panel A of Table (ref) and in Panels A and B of Table (ref), we additionally report the two-step efficiency bound in the line denoted “Barendse Bound". For the homoskedastic innovations and for the ES specific parameters, the two-step efficiency bound coincides with the Z-estimation efficiency bound, while it does not for the quantile parameters. This is primarily caused by the inefficient first-step quantile estimation---using an efficient quantile estimator (based on $g_t(\xi_1) = F_t(\xi_1)$) would equate both efficiency bounds in the homoskedastic case. In contrast, in the heteroskedastic case, the two-step efficiency bound is considerably larger than the Z-estimation efficiency bound for all four considered parameters. Interestingly, the choice of $\phi^\text{eff2}$ motivated by this two-step estimation efficiency bound exhibits equally efficient ES parameters while the quantile parameters show larger standard deviations.

Connections to the semiparametric efficiency bound

The main focus of this article lies on the Z-estimation efficiency bound for conditional moment restrictions of Hansen1985, Chamberlain1987 and Newey1993. In the context of i.i.d.\ processes and differentiable moment conditions, Chamberlain1987 shows that this bound coincides with the general semiparametric efficiency bound in the sense of a least favorable submodel of Stein1956; c.f.\ Newey1990 and BKRW1998book for surveys on this matter and Ackerberg2014, Jankova2018, Hristache2016, Komunjer2010a for some recent progress.

The definition of the semiparametric efficiency bound builds on the idea that the data stems from a parametric submodel, i.e., a parametric model which completely specifies the full distribution, contains the correctly specified model, and satisfies the semiparametric model assumption. E.g., if we consider a semiparametric model for the conditional mean, we do not make any assumptions about the exact conditional distribution beyond the mean assumption. Any model which parametrizes the full conditional distribution (e.g., a normal distribution with parameterized variance) is such a parametric submodel. Estimation of any parametric submodel is subject to the classical Cram\'{e}r--Rao efficiency bound, which can be attained, e.g., by maximum likelihood estimation using the true parametric distribution, dispensing with a discussion of superefficient estimators. For any parametric submodel, a consistent and asymptotically normal semiparametric estimator is contained in the class of estimators for this parametric submodel and thus, it is subject to the parametric Cram\'{e}r--Rao efficiency bound. Consequently, any semiparametric estimator has an asymptotic variance which is no smaller than the Cram\'{e}r--Rao bound for any parametric submodel. Hence, the semiparametric efficiency bound is defined as the supremum of the Cram\'{e}r--Rao bounds of all parametric submodels.

The results of this paper concerning efficient estimation are derived with respect to the Z-estimation efficiency bound of Hansen1985, Chamberlain1987, and Newey1993. In applications to smooth objective functions and i.i.d.\ processes, the result of Chamberlain1987 can be used to equate these two bounds. However, as we are not aware of a general relation of these bounds for non-i.i.d.\ processes, we cannot preclude that the semiparametric efficiency bound is strictly smaller (in the Loewner order) than the Z-estimation efficiency bound in certain situations. Consequently, all following assertions are stated in relation to the Z-estimation efficiency bound. This does not affect our main conclusion in terms of efficient estimation: When the M-estimator cannot attain the Z-estimation efficiency bound, it also cannot attain the semiparametric efficiency bound, irrespectively of whether these quantities coincide.

Identification of the Efficient Z-estimator for double quantile models

Following DFZ_CharMest, strict model consistency can directly be obtained by employing strictly consistent loss functions and a no-perfect collinearity condition of the model gradient. In contrast, this is more involved for the Z-estimator. Thus, the following proposition shows strict model identification for an efficient Z-estimator and for a large class of models.

Note that Theorem (ref) asserts that the Z-estimator is efficient based on any choice $A^*_{t,C}(X_t,\theta)$ of instrument matrix such that $A^*_{t,C}(X_t,\theta_0) = CD_t(X_t,\theta_0)^\intercal S_t(X_t,\theta_0)^{-1}$, see (ref). This means we only have a condition on $A^*_{t,C}(X_t,\theta)$ for $\theta= \theta_0$. To come up with such a matrix, there are two straight forward ways how to guarantee this. First, we might set $A^*_{t,C}(X_t,\theta) = CD_t(X_t,\theta)^\intercal S_t(X_t,\theta)^{-1}$, and second, we might choose $A^*_{t,C}(X_t,\theta)$ to be constant in $\theta$ and equal to $CD_t(X_t,\theta_0)^\intercal S_t(X_t,\theta_0)^{-1}$. For practical purposes, the latter situation is often hard or infeasible to implement, since it usually requires knowledge of the unknown true parameter $\theta_0$ (and additional quantities of the conditional distribution $F_t$).

For the particular situation of the double quantile model, using the canonical identification function $\varphi$ given in Running Example (ref), $S_t(X_t,\theta_0)$ takes the form (ref), which means it is entirely independent of any knowledge on the underlying DGP whatsoever. This makes the latter choice attractive and reasonably feasible.

propositionWe assume that (a) the double quantile model is linear with separated parameters, i.e., $Q_\alpha(Y_t|X_t) = q_\alpha(X_t,\theta_0^\alpha) = X_t^\intercal \theta_0^\alpha$ and $Q_\beta(Y_t|X_t) = q_\alpha(X_t,\theta_0^\beta)= X_t^\intercal \theta_0^\beta$, such that $\theta_0 = (\theta_0^\alpha, \theta_0^\beta)\in\operatorname{int}(\Theta)$, (b) for all $t\in\mathbb N$, $F_t$ is differentiable with a strictly positive derivative $f_t$, and (c), there exists a possibly time-dependent deterministic constant $c_t > 0$, such that $f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) = c_t f_t\big(q_\beta(X_t,\theta_0^\beta)\big)$ almost surely. Then, the moment function of the efficient Z-estimator of the DQR model is a strict $\mathcal{F}_\mathcal{Z}$-identification function for $\theta_0$, i.e., it holds that \begin{align*} \mathbb{E} \left[ A_t^\ast(X_t, \theta_0) \varphi \big( Y_t, m(X_t, \theta) \big) \right] = 0 \quad \Longleftrightarrow \quad \theta = \theta_0, \end{align*} where $A_t^\ast(X_t, \theta_0)$ is given in (ref).

At the cost of some more tedious notation, Proposition (ref) can be generalized to the situation of linear models with not necessarily separated parameters, so long as there is at least one component that is used for modelling one quantile only, respectively. E.g., in a simple linear regression model, the two quantile models might have the same slope, but a different intercept, or vice versa, they might have the same intercept, but a different slope. Generalising the assertion much beyond linear models seems to be difficult due to the application of the mean value theorem in the proof.

proof[Proof of Proposition (ref)] It holds that \( \mathbb{E} \left[ A_t^\ast(X_t, \theta_0) \varphi \big( Y_t, m(X_t, \theta_0) \big) \right] = 0 \) since we have that $\mathbb{E}\left[\varphi \big( Y_t, m(X_t, \theta_0) \big)\big|X_t\right]=0$. The reverse direction is a little more involved. For this, straight-forward calculations yield that for any $\theta\in\Theta$ \begin{align*} \mathbb{E} \left[ A_t^\ast(X_t, \theta_0) \varphi \big( Y_t, m(X_t, \theta) \big) \right] = \mathbb{E} \left[ U_1 \nabla_\theta q_\alpha(X_t,\theta_0^\alpha) + U_2 \nabla_\theta q_\beta(X_t,\theta_0^\beta) \right], \end{align*} where the scalar and $\sigma(X_t)$-measurable random variables $U_1$ and $U_2$ are given by \begin{align*} U_1 &= \frac{ f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) }{\alpha (1-\alpha) \beta - \alpha^2 (1-\beta)} \big( \beta a - \alpha b \big) \qquad and \\ U_2 &= \frac{ f_t\big(q_\beta(X_t,\theta_0^\beta)\big) }{\beta (1-\alpha) (1-\beta) - \alpha (1-\beta)^2} \big( -(1-\beta) a + (1- \alpha) b \big), \end{align*} with $a = F_t\big(q_\alpha(X_t,\theta^\alpha)\big) - \alpha$ and $b = F_t\big(q_\beta(X_t,\theta^\beta)\big) - \beta$. As $\nabla_\theta q_\alpha(X_t,\theta_0^\alpha) = \begin{pmatrix} X_t \\0 \end{pmatrix}$ and $\nabla_\theta q_\beta(X_t,\theta_0^\beta) = \begin{pmatrix} 0\\mathcal{X}_t \end{pmatrix}$, it holds that $\mathbb{E} \left[ A_t^\ast(X_t, \theta_0) \varphi \big( Y_t, m(X_t, \theta) \big) \right] = 0$ if and only if \begin{align} \mathbb{E} \left[ f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) \big( \beta a - \alpha b \big) X_t \right] = 0 \ and \ \mathbb{E} \left[ f_t\big(q_\beta(X_t,\theta_0^\beta)\big) \big( (1-\beta) a - (1- \alpha) b \big) X_t \right] = 0. \end{align} As $f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) = c_t f_t\big(q_\beta(X_t,\theta_0^\beta)\big)$ almost surely by assumption (where $c_t$ is deterministic), this implies that \begin{align*} \beta \mathbb{E} \big[ f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) a X_t \big] - \alpha \mathbb{E} \big[ f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) b X_t \big] &= 0 \qquad and \\ c_t (1-\beta) \mathbb{E} \big[ f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) a X_t \big] - c_t (1-\alpha) \mathbb{E} \big[ f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) b X_t \big] &= 0. \end{align*} Solving this system of equations, where we exploit that $c_t\neq0$, and combining it with (ref) and the fact that $\alpha\neq \beta$, we arrive at \begin{align} \mathbb{E} \left[ f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) a X_t \right] = 0 \quad and \quad \mathbb{E} \left[ f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) b X_t \right] = 0. \end{align} We now proceed by a proof through contradiction with an argument similar as in DimiBayer2019. For this, assume that $\theta \not= \theta_0$. Using the zero-condition in ((ref)), we get \begin{align*} 0 &= \mathbb{E} \left[ f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) a X_t^\intercal \right] \big( \theta^\alpha - \theta_0^\alpha \big) \\ &= \mathbb{E} \left[ f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) X_t^\intercal \big( \theta^\alpha - \theta_0^\alpha \big) \big( F_t\big(q_\alpha(X_t,\theta^\alpha)\big) - F_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) \big) \right] \\ &= \mathbb{E} \left[ f_t\big(q_\alpha(X_t,\theta_0^\alpha)\big) f_t\big(q_\alpha(X_t,\tilde \theta^\alpha)\big) \Big(X_t^\intercal \big( \theta^\alpha - \theta_0^\alpha \big) \Big)^2 \right], \end{align*} where we have used the mean value theorem and the linearity of the model to obtain the last identity and where $\tilde \theta^\alpha = (1-\lambda) \theta_0^\alpha + \lambda \theta^\alpha$ for some $\lambda\in[0,1]$. By assumption, the density is strictly positive such that we can conclude that $\mathbb P\big(X_t^\intercal ( \theta^\alpha - \theta_0^\alpha ) =0\big)=1$. Then, due to Assumption (ref), it must hold that $\theta^\alpha = \theta_0^\alpha$. Employing a similar argument to $\theta^\beta$ yields that $\theta^\beta= \theta_0^\beta$, which concludes this proof.

Additional Simulation Results

In this section, we report simulation results for additional probability levels for the double quantile, and the joint quantile and ES models discussed in Sections (ref) and (ref). Specifically, Tables (ref) and (ref) present results for the double quantile model and Table (ref) for the joint quantile and ES model. The format of these tables follows Tables (ref) and (ref) from the main document.

table[table omitted — 6,208 chars of source]
table[table omitted — 5,323 chars of source]
table[table omitted — 6,695 chars of source]

\FloatBarrier