The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
66,723 characters
\bibliographystyle{joe}
\citationstyle{dcu}
\lineskip=1ex \baselineskip 4ex
\setlength{\arraycolsep}{0.02in}
\pagestyle{empty}
\begin{center}
\noindent{\Large\bf Estimation risk in conditional expectiles}\vskip 3em
\end{center}
\textbf{Marcelo Fernandes}\\
Sao Paulo School of Economics, FGV\\\\
\textbf{Víctor Henriques}\\
Argus Media\\\\
\textbf{Eduardo Fonseca Mendes}\\
Sao Paulo School of Business Administation, FGV\vskip 3em
\noindent\textbf{Abstract:}~~We establish the consistency and asymptotic normality of a two-step estimator of conditional expectiles in the context of conditional scale models. We first estimate the conditional variance parameters by quasi-maximum likelihood and then compute the unconditional expectile of the innovations using the empirical distribution of the standardized residuals. We show how replacing true innovations with standardized residuals affects the asymptotic variances of both conditional and unconditional expectile estimators. Finally, our empirical analysis reveals that conditional expectiles assess tail risk in cryptomarkets in a more robust manner than traditional quantile-based risk measures, such as value at risk and expected shortfall.\\
\textbf{Keywords:} Asymmetric least squares, quantiles, tail risk.\vfill
{\small\noindent\textbf{Acknowledgments:}~~We are indebted to Silvia Gonçalves, Giuseppe Cavaliere, Abdelaati Daouia, Alan De Genaro, and Edu Horta, as well as to seminar participants at the UBRI Connect (Zurich, 2024), High-Voltage Econometrics Workshop (Lecce, 2024), Research in Options (Rio de Janeiro, 2024), São Paulo School of Advanced Sciences on High-Dimensional Modeling (São Paulo, 2025), BCN Workshop in Financial Econometrics (Barcelona, 2025), Italian Congress of Econometrics and Empirical Economics (Palermo, 2025), SoFiE Annual Conference (Paris, 2025), TSE Financial Econometrics Conference (Hammamet, 2025), Brazilian School of Time Series and Econometrics (Campinas, 2025), Meetings of the Brazilian Finance Society (São Paulo, 2025), Meetings of the Brazilian Econometric Society (São Paulo, 2025), Workshop on Econometrics and Complex Networks (Buenos Aires, 2026), Virtual Time Series Seminar, University of Illinois at Urbana-Champaign, and University of São Paulo. Fernandes and Mendes acknowledge financial support not only from FAPESP (2023/01728-0) and CNPq (302278/2018-4 and 302153/2022-5), but also from the Silicon Valley Community Foundation through the University Blockchain Research Initiative (UBRI, 2022/199610). The usual disclaimer applies.}
\newpage\pagestyle{plain}
\lineskip=1ex \baselineskip 5ex
\section{Introduction}
Quantile-based measures, such as value at risk (VaR) and expected shortfall (ES), dominate risk assessments. The Basel II regulatory framework calculates minimum capital requirements using a VaR approach (\url{www.bis.org/publ/bcbs24.pdf}). However, value at risk is not a satisfactory risk measure for two reasons. First, it does not adequately account for diversification gains because it violates the subadditivity property that characterizes coherent risk measures \cite{artzner1999coherent,follmer2002convex}. Second, since the value at risk corresponds to a quantile of the distribution, it defines what constitutes a tail realization, but without assessing expected severity. Accordingly, the Basel III regulatory framework suggests using expected shortfall (\url{www.bis.org/publ/bcbs265.pdf}). By measuring the expected loss in excess of the value at risk, ES effectively accounts for not only diversification gains, but also the severity of the tail realization.
The quest for better risk measures is not over yet. \cn{gneiting2011} demonstrates that ES is not individually elicitable, so there is no natural backtest procedure to assess the performance of ES forecasts. In addition, quantile-based risk measure estimates consider only the relative frequency of observations above or below their corresponding predictions \cite{daouia2018estimation}, depending heavily on the tail of the loss distribution \cite{kuan2009assessing,daouia2019extremiles}. More specifically, ES is too conservative because it restricts attention to a given quantile, whereas VaR is too lenient because it ignores severity. As such, risk measures based on a given quantile arguably either underestimate or overestimate risk exposures.
The alternative class of risk measures based on \pc{newey1987asymmetric} asymmetric least squares (ALS) has been gaining traction in the literature. In particular, expectiles (XP) are a least-squares analog of quantiles, offering the only coherent (law-invariant) risk measure that simultaneously accounts for diversification gains and allows straightforward backtesting \cite{ziegel2016coherence,daouia2023expectile}. Moreover, it depends on both the probability and severity of the tail event, making them particularly suitable for actuarial and portfolio-allocation problems. This paper establishes the consistency and asymptotic normality of the two-step estimators of the conditional and unconditional expectiles for conditional scale models. As in \cn{francq2015risk}, we first estimate the scale parameters by Gaussian quasi-maximum likelihood (QML) and then estimate the innovation expectile using the empirical distribution of the standardized residuals. Because innovations are unobservable, inference must account for the first-step estimation error. We derive a Bahadur representation that explicitly shows how replacing innovations with standardized residuals affects the second-step estimation. We also obtain a closed-form asymptotic variance, consistent feasible variance estimators, and a conditional Gaussian approximation for the conditional expectile given past realizations.
We corroborate our asymptotic results with Monte Carlo experiments based on designs that reproduce the main stylized facts of asset returns. The simulations assess the finite-sample bias and dispersion of the two-step estimators of the conditional and unconditional expectiles, as well as the quality of the Gaussian approximation based on the asymptotic variance estimator. Both bias and root mean squared error decrease sharply with sample size, whereas standard errors are on average very close to the standard deviation over Monte Carlo replications. In addition, the actual coverage rates of the Wald confidence intervals for both conditional and unconditional expectiles are close to nominal values.
Our approach aligns with the literature. \cn{gao2008estimation} establish consistency and asymptotic normality of the value-at-risk and expected shortfall based on GARCH standardized residuals. \cn{francq2015risk} extend the asymptotic theory to cover several GARCH-type specifications: e.g., exponential GARCH, asymmetric power ARCH, and GJR-GARCH models \cite{nelson1991conditional,ding1993long,glosten1993relation}. We are the first to assess how the estimation of conditional scale parameters affects the inference on conditional and unconditional expectiles. \cn{holzmann2016expectile} and \cn{kratschmer2017statistical} develop the asymptotic theory for the estimation of unconditional expectiles at a fixed level $\tau$ using independent and identically distributed (iid) data. \cnm{daouia2018estimation} (2018, \cy{daouia2020tail}) derive the asymptotic distribution of a weighted ALS estimator for extreme expectiles at level $\tau_n$, with $\tau_n\to 1$ as the sample size $n$ grows. \cn{girard2021extreme} extend the analysis to consider the estimation of conditional extreme expectiles in heavy-tailed heteroskedastic regressions using a two-step approach. In particular, they show that substituting standardized residuals for true innovations does not affect the asymptotic distribution of extreme quantile estimators because the latter converges at a slower rate $\sqrt{n(1-\tau_n)}$ than the first-step estimation error shrinks to zero. Unfortunately, this is not the case here. The estimation of expectiles at a fixed level $\tau$ converges at the same $\sqrt{n}$-rate as the estimation of the conditional mean and variance parameters, and hence the latter affects the asymptotic distribution of the former.
Finally, we empirically assess the performance of conditional expectiles relative to quantile-based risk measures in cryptocurrency markets \ca{MS22}{for an overview, see}. Our motivation is twofold. First, crypto assets appeal to investors mainly due to their low correlation with traditional asset classes \ca{BB22}{see, among others,}. This suggests that we should not assess tail risk by looking only at the value at risk, as we would miss out any diversification benefit. Second, extreme tail events are relatively common in crypto markets \cite{gkillas2018application,STT18,borri2019,nguyen2020investigating}. This casts doubt on the suitability of the expected-shortfall measure, given that the latter is very hard to estimate precisely under heavy tails. As such, crypto markets make fertile ground for the application of conditional expectiles in the stipulation of minimum capital requirements.
We find that the one-step-ahead prediction intervals of the conditional expectiles at the 1\% level are more precise and tighter for every currency, yielding reasonable capital requirements. In particular, the prediction intervals of the value at risk at the 1\% level are between 11.4\$ and 50.9\% wider, whereas those of the expected shortfall at the 1\% level are between 122.1\% and 183.9\% wider. In addition, we explore the one-to-one mapping between quantiles and expectiles to assess how cryptocurrency risks evolve over time through the lens of the expected gain-loss ratio. Our empirical analyses contribute to a better understanding of cryptocurrency markets, complementing previous studies in the literature. For instance, \cn{zhang2021downside} examine how downside risk affects the cross-section of cryptocurrency returns, whereas \cnm{MS19} (2019, \cy{MS20}) investigate price discovery and arbitrage opportunities in cryptomarkets, respectively.
The remainder of this paper proceeds as follows. Section \ref{sec:ALS} discusses the main aspects of expectile-based risk measures. Section \ref{sec:model} introduces the conditional scale model and the two-step estimators of conditional and unconditional expectiles. Section \ref{sec:estimation} derives the asymptotic theory, while Section \ref{sec:mc} reports Monte Carlo simulations that assess the finite-sample behavior of the two-step estimators and their standard errors. Section \ref{sec:crypto_application_clt} examines tail risk in cryptomarkets. Section \ref{sec:conclusion} offers some concluding remarks. Appendix A collects technical Assumptions~and proofs.
\section{Expectiles}\label{sec:ALS}
Let the loss $L\in\mathbb{R}$ be a square integrable random variable with loss distribution function $F_L$. \cn{newey1987asymmetric} define the $\tau$-th expectile ${\textrm{XP}}_{\!\tau}$ as
\begin{equation}\label{eq:expectiles_optimization_problem}
{\textrm{XP}}_{\!\tau}=\underset{\xi\in\mathbb{R}}{\arg\min}\int_{\mathbb{R}}|\tau-\mathbf{1}(\ell<\xi)|(\ell-\xi)^2\textrm{d}F_L(\ell)=\underset{\theta\in\mathbb{R}}{\arg\min}\,\E\left[\rho_\tau(L-\xi)\right]
\end{equation}
where $\rho_\tau(u)=|\tau-\mathbf{1}(u\le 0)|u^2$ is the expectile check function and $\mathbf{1}(A)$ denotes the indicator function that takes value one if $A$ is true, zero otherwise. Apart from coinciding with the mean for $\tau=1/2$, the expectile at any $\tau\in(0,1)$ is well defined and unique for any integrable random variable \cite{newey1987asymmetric,abdous1995relating}.
Replacing the absolute deviation in the quantile check function by a quadratic deviation facilitates optimization in a substantial manner. The quantile objective function is not continuously differentiable at zero, whereby numerical implementation occasionally leads to quantile-crossing functions. In addition, quantile estimators consider only the relative frequency of observations above or below their corresponding predictions \cite{daouia2018estimation,daouia2019extremiles}, with asymptotic distributions that strongly depend on the density function and smoothing parameters \cite{cheng1997unified}. In contrast, the expectile objective function is continuously differentiable almost everywhere. Implementation is straightforward by iterative reweighted least squares \cite{daouia2018estimation}, whereas asymptotic normality requires only finite second moments \cite{holzmann2016expectile}.
The first-order condition of \eqref{eq:expectiles_optimization_problem} also indicates that expectiles depend both on the probability and magnitude of tail realizations:
\begin{displaymath}
\begin{aligned}
0&=\left.\frac{\textrm{d}}{\textrm{d}\xi}\int_{\mathbb{R}}|\mathbf{1}(\ell\le\xi)-\tau|(\ell-\xi)^2\,\textrm{d} F_L(\ell)\right|_{\xi={\textrm{XP}}_{\!\tau}}\\
&=\left.\frac{\textrm{d}}{\textrm{d}\xi}\int_{-\infty}^\xi(1-\tau)(\ell-\xi)^2\,\textrm{d}F_L(\ell) \right|_{\xi={\textrm{XP}}_{\!\tau}}+\left.\frac{\textrm{d}}{\textrm{d}\xi}\int_\xi^\infty\tau(\ell-\xi)^2\,\textrm{d}F_L(\ell)\right|_{\xi={\textrm{XP}}_{\!\tau}}\\
&=2(1-\tau)\int_{-\infty}^{{\textrm{XP}}_{\!\tau}}({\textrm{XP}}_{\!\tau}-\ell)\,\textrm{d}F_L(\ell) +2\tau\int_{{\textrm{XP}}_{\!\tau}}^\infty({\textrm{XP}}_{\!\tau}-\ell)\,\textrm{d}F_L(\ell).
\end{aligned}
\end{displaymath}
It then follows for $u^+=\max\{u,0\}$ and $u^-=\min\{u,0\}$ that
\begin{displaymath}
\tau=\frac{\E(L-{\textrm{XP}}_{\!\tau})^-}{\E(L-{\textrm{XP}}_{\!\tau})}=\frac{\int_{-\infty} ^{{\textrm{XP}}_{\!\tau}}({\textrm{XP}}_{\!\tau}-\ell) \,\textrm{d}F_L(\ell)}{\int_\mathbb{R}({\textrm{XP}}_{\!\tau}-\ell)\,\textrm{d}F_L(\ell)},
\end{displaymath}
implying that the expectile corresponds to the ratio of the average deviation of $Y$ below ${\textrm{XP}}_{\!\tau}$ to the overall average deviation \cite{kuan2009assessing}. As such, evaluating \pc{keating2002universal} omega ratio at the expectile yields $\Omega_Y({\textrm{XP}}_{\!\tau})=(1-\tau)/\tau$. See \cn{remillard2013statistical}, \cn{bellini2017risk} and \cn{bellini2018expectiles} for more links between expectiles and expected gain-loss ratios.
Expectiles also relate to quantiles in several aspects. First, although they both characterize the entire distribution, only the expectile function is continuous and monotonically increasing on $\tau$ for any distribution. Second, there is a straightforward link between expectiles and quantiles:
\begin{equation}\label{eq:tau equivalence}
\tau(\alpha)={\textrm{XP}}_{\!\tau}^{-1}(q_\alpha)=\frac{\int_{-\infty}^{q_\alpha}|\ell-q_\alpha| \,\mathrm{d}F_L(\ell)}{\int_\mathbb{R}|\ell-q_\alpha|\,\mathrm{d}F_L(\ell)},
\end{equation}
so that $\textrm{XP}_{\tau(\alpha)}=q_\alpha$ \cite{yao1996asymmetric}. Third, it is possible to estimate quantiles using expectiles, given that both are contained within the convex hull of the distribution's support \cite{waltrup2015expectile,daouia2019quantiles}. In fact, quantiles are a strict subset of the corresponding expectiles. Conveniently, this remains true in a regression context if the distribution belongs to the location-scale family \cite{yao1996asymmetric}.
After a slow start \cite{breckling1988m,efron1991regression,jones1994expectiles,abdous1995relating,yao1996asymmetric}, the literature on expectiles has recently gained traction \cite{martin2014expectiles,bellini2017risk,kratschmer2017statistical,daouia2018estimation,daouia2020tail,girard2021extreme}. The primary reason is that the expectile is the only coherent risk measure that allows straightforward backtesting due to elicitability \cite{bellini2015elicitable,ziegel2016coherence,daouia2023expectile}.
Coherent risk measures are desirable because they satisfy the axioms of monotonicity, translation invariance, subadditivity, and positive homogeneity. Monotonicity reflects that, if one position is less risky than another, the risk measure should assign a lower risk value. The translation invariance axiom states that adding a constant amount to the position's payout should increase riskiness by exactly that amount. The subadditivity axiom ensures that there are diversification gains in pooling risks, whereas positive homogeneity dictates that multiplying the position by a positive constant should scale the risk measure by the same constant. See \cn{mcneil2015quantitative} for more details. Mathematical properties aside, we must always keep in mind that, in practice, we have to assess risk measures empirically through estimates and/or forecasts. This is exactly the idea of elicitability, which requires the feasibility of backtesting risk measures through minimization of expected scores \cite{gneiting2011,bellini2015elicitable,ziegel2016coherence}.
\begin{comment}
Recall that a score is a function $S:\mathbb{R}\times\mathbb{R}\to[0,\infty)$ such that $S(u,v)\ge 0$ with $S(u,v)=0$ if and only if $u=v$; $S(u,v)$ is increasing for $u>v$ and decreasing for $u<v$; and $S(u,v)$ is continuous in $u$. \cn{gneiting2011} shows that strictly consistent, homogeneous scoring functions for expectiles are of the form $S(r,L)=\mathbf{1}\{L>r\}(1-2\tau)\big(\bar\phi(r)-\bar\phi(L)-\bar\phi'(r)(r-L)\big)-(1-\tau)\big(\bar\phi(r)-\bar\phi'(r)(r-L)\big)$, with $\bar\phi$ denoting a strictly convex and integrable function with subgradient $\bar\phi'$. The natural choice is the 2-homogeneous scoring function given by $\phi(r)=r^2$ as in \cn{newey1987asymmetric}, though \cn{nolde2017elicitability} also entertains a 0-homogeneous alternative with $\phi(r)=\ln r$:
\begin{eqnarray}
S_{\textrm{XP,2}}(r,L)&=&-\mathbf{1}\{L>r\}(1-2\tau)(L-r)^2+(1-\tau)r(r-2L)\label{eq:xp_score_sq}\\
S_{\textrm{XP,0}}(r,L)&=&-\mathbf{1}\{L>r\}(1-2\tau)\left(\ln\frac{L}{r}+1-\frac{L}{r}\right)-(1-\tau)\left(\ln r-1+\frac{L}{r}\right),~\mbox{for}~r>0\label{eq:xp_score_log}
\end{eqnarray}
respectively. The corresponding scoring functions for the value-at-risk and expected-shortfall measures are respectively
\begin{eqnarray}
S_{\textrm{VaR,1}}(r,L)&=&\big(1-\alpha-\mathbf{1}\{L>r\}\big)r+\mathbf{1}\{L>r\}L\label{eq:var_score_linear}\\
S_{\textrm{VaR,0}}(r,L)&=&\big(1-\alpha-\mathbf{1}\{L>r\}\big)\ln r+\mathbf{1}\{L>r\}\ln L,~\mbox{for}~r>0\label{eq:var_score_log}\\
S_{\textrm{VaR,ES,1/2}}(r_1,r_2,L)&=&\mathbf{1}\{x>r_1\}\,\frac{x-r_1}{2\sqrt{r_2}}+(1-v)\frac{r_1+r_2}{2\sqrt{r_2}}\label{eq:ES_score_half}\\
S_{\textrm{VaR,ES,0}}(r_1,r_2,L)&=&\mathbf{1}\{x>r_1\}\,\frac{x-r_1}{r_2}+(1-v)\left(\frac{r_1}{r_2}-1+\ln r_2\right),~\mbox{for}~r_2>0.\label{eq:ES_score_log}
\end{eqnarray}
The individual scores for the value at risk are 1- and 0-homogeneous functions, whereas the joint scores for value at risk and expected shortfall are $1/2$- and 0-homogeneous functions. See, for more details, \cn{THOMSON1979360}, \cn{saerens2000building}, \cn{acerbi2014back}, \cn{fissler2016higher}, and \cn{nolde2017elicitability}.
\end{comment}
\subsection{Comparison between expectiles and quantile-based risk measures}
In this section, we compare expectiles with traditional quantile-based risk measures. We start with the value-at-risk measure $\operatorname{VaR}_\alpha$ at level $\alpha$, which reads
\begin{displaymath}
\operatorname{VaR}_\alpha=q_\alpha=F_L^{-1}(\alpha)=\inf\{\ell\in\mathbb{R}:\,F_L(\ell)\ge\alpha\},
\end{displaymath}
where $q_\alpha$ is the quantile function at level $\alpha\in(0,1)$. By definition, quantiles automatically satisfy the axioms of monotonicity, translation invariance, and positive homogeneity. In addition, they are elicitable for strictly increasing distribution functions \cite{THOMSON1979360,saerens2000building}. Unfortunately, value at risk does not account for the expected loss and diversification gains \cite{danielsson2001academic}.
\cn{artzner1999coherent}, \cn{acerbi2001expected}, \cn{rockafellar2002conditional} argue that the expected shortfall (ES), as defined by the expected loss given that it exceeds the value at risk, addresses both weaknesses of the VaR measure. For any integrable loss $L$ with distribution $F_L$, the ES at level $\alpha\in(0,1)$ is
\begin{displaymath}
\mathrm{ES}_\alpha=\E\left[L\mid L>\operatorname{VaR}_\alpha(L)\right].
\end{displaymath}
In particular, ES is a severity-based risk measure that belongs to the class of spectral risk measures \cite{acerbi2002spectral}. However, it fails in two accounts. First, ES is not elicitable by itself, so that backtesting is far from straightforward \cite{gneiting2011}. Second, it yields very imprecise estimates in finite samples for large values of $\alpha$ \cite{hull2014shortfalls}, especially for heavy-tailed loss distributions \cite{yamai2002comparative}.
Both value at risk and expected shortfall depend heavily on the shape of the tail, though. While the ES is too conservative because it is conditional only on tail events, the VaR is too lenient because it does not account for their severity. As such, quantile-based risk measures either underestimate or overestimate the risk exposure of a position \cite{kuan2009assessing,daouia2019extremiles}.
\cn{bellini2017risk} interpret the expectile risk measure as the amount of capital that should be added to a position to produce a sufficiently high expected gain–loss ratio. As the only risk measure that meets the conditions for coherence, law invariance and elicitability \cite{ziegel2016coherence}, expectiles are extremely convenient for modeling, forecasting and backtesting purposes. In particular, the key advantage of expectiles over VaR and ES is the fact that ALS estimation uses the available data more efficiently, by exploiting both the severity and probability of tail events \cite{daouia2018estimation}. Both VaR and ES estimates disregard the shape of the loss distribution to the left of the corresponding quantile. In addition, the precision of the ES estimates depends heavily on $\alpha$, sample size, and tail thickness. This obviously poses a problem for establishing capital requirements in the case of highly volatile and heavy-tailed asset returns \cite{yamai2002comparative,hull2014shortfalls}.
For $\tau=\alpha>1/2$, expectiles are always below both value-at-risk and expected-shortfall measures for every sample size, regardless of tail thickness. However, this ignores the expectile-quantile mapping in \eqref{eq:tau equivalence}. For instance, the first percentile of the standard Gaussian distribution is close to its expectile at $\tau=0.00145$ \cite{bellini2017risk,nolde2017elicitability}. \cn{chen2018exactitude} shows that, if $\tau(\alpha)$ is such that $\textrm{XP}_{\tau(\alpha)}=\textrm{VaR}_\alpha$, then
\begin{equation}
\label{eq:expectile as expected shortfall}
\tau(\alpha)=\frac{\alpha\,(\textrm{ES}_\alpha-\textrm{VaR}_\alpha)}{\textrm{VaR}_\alpha+2\alpha\,(\textrm{ES}_\alpha-\textrm{VaR}_\alpha)}
\end{equation}
provided that the loss distribution has mean zero. Straightforward manipulations then yield
\begin{equation}
\label{eq:expectile-based expected shortfall}
\textrm{ES}_{\alpha}=\textrm{VaR}_{\alpha}\left(1+\frac{1}{\alpha(\Omega_{\alpha}-1)}\right).
\end{equation}
This makes the connection with the omega ratio \cite{taylor2022forecasting}. Solving for $\Omega_\alpha$ gives way to
\begin{equation}
\label{eq:omega value-at-risk}
\Omega_\alpha=1+\frac{\textrm{VaR}_{\alpha}}{\alpha(\textrm{ES}_{\alpha}-\textrm{VaR}_{\alpha})},
\end{equation}
allowing us to compute the expected gain-loss ratio as a function of the quantile for some fixed level $\alpha$. In addition, it makes clear that, for a fixed $\alpha$, we should expect the gain-loss ratio to increase (or decrease) as the gap between expected shortfall and value at risk shrinks (or enlarges, respectively). As such, we can infer by means of the omega ratio whether a given value-at-risk model is conservative (or permissive).
In summary, expectiles offer a different perspective of the loss distribution than quantile-based risk measures. However, the literature on ALS estimation is mostly under random sampling. Only a few studies address the estimation of conditional expectiles in a time-series context \cite{taylor2008estimating,kuan2009assessing,bellini2017risk,girard2021extreme}. In the next section, we discuss a two-step estimator of the conditional expectile in a similar context to \cn{francq2015risk}. In particular, due to translation invariance and positive homogeneity, it is straightforward to compute conditional expectiles in conditional scale models by first estimating the conditional volatility of asset returns and then computing the unconditional expectiles of their standardized residuals.
\subsection{Conditional expectiles}\label{sec:model}
As expectiles are coherent risk measures, they are stable under affine transformations. This means that, in a location-scale model, the conditional expectile at a fixed level $\tau$ depends exclusively on the conditional mean and volatility, and of the $\tau$-th expectile of the innovation. As it turns out, asset returns typically exhibit a mean close to zero, but a highly persistent time-varying volatility with leverage effects \cite{bollerslev1994arch} that leads to skewness and heavy tails in asset returns \cite{fama1965behavior}. Accordingly, we entertain a GARCH-type approach to model continuously compounded returns $\{y_t\}$:
\begin{equation}\label{eq:conditional volatility model}
y_{t+1}=\sigma_t\,\eta_{t+1},\qquad\qquad\text{with}~~\sigma_t(\bs{\theta}_0)=\sigma(y_t,y_{t-1},\ldots;\,\bs{\theta}_0)~~\text{for}~~t\in\mathbb{Z},
\end{equation}
where $\{\eta_t\}$ is a sequence of iid random variables with zero mean and unit variance, independent of past returns (i.e., $y_s\perp\!\!\!\perp\eta_t$ for $s<t$), and $\bs{\theta}_0\in\mathbb{R}^m$ is a vector of model parameters. The sequence $\{y_t\}$ is a strictly stationary and ergodic solution to \eqref{eq:conditional volatility model}. This setting is very general, nesting the most popular GARCH-type models in the literature. For example, \pc{bollerslev1986generalized} GARCH(1,1) model is such that $\sigma_t^2=\omega_0+\alpha_0\,y_t^2+\beta_0\,\sigma_{t-1}^2$, where $\bs{\theta}_0=(\omega_0,\alpha_0,\beta_0)'$.
Given our interest in expectiles, we henceforth assume that log-returns $y_{t+1}$ are integrable and follow a strictly stationary process adapted to its natural filtration $\{\mathcal{F}_t\}$. The conditional expectile at level $\tau$ then reads
\begin{equation}\label{eq:conditional expectile}
{\textrm{XP}}_{\!\tau}(y_{t+1}|\mathcal{F}_t)=\sigma_t(\bs{\theta}_0)\,{\textrm{XP}}_{\!\tau}^\eta,
\end{equation}
where ${\textrm{XP}}_{\!\tau}^\eta$ is the innovation expectile for $\tau\in(0,1)$.
\section{Estimation of conditional expectiles}\label{sec:estimation}
If we could observe the true innovations, their empirical expectile at the level $\tau$ would solve
\begin{displaymath}
\widehat\xi_n^{\,0}=\underset{\xi\in\mathbb R}{\arg\min}\,\frac{1}{2n}\sum_{t=1}^n\big|\tau-\mathbf 1(\eta_t<\xi)\big|\,(\eta_t-\xi)^2.
\end{displaymath}
Equivalently, $\widehat\xi_n^{\,0}$ is the unique solution of $n^{-1}\sum_{t=1}^n\phi_\tau(\eta_t-\xi)=0$, where $\phi_\tau(u)=u\,w_\tau(u)$ with $w_\tau(u)=\big|\mathbf 1(u<0)-\tau\big|$ for $u\in\mathbb R$. This estimator is strongly consistent as long as the first moment is finite \cite{holzmann2016expectile,kratschmer2017statistical}. Unfortunately, this estimator is infeasible in our case as we do not observe the true innovations. We therefore proceed in two steps as in \cn{francq2015risk}. We first estimate $\bs\theta_0$ in the conditional variance specification by quasi-maximum likelihood and then estimate the expectile using the standardized residuals
\begin{displaymath}
\widehat\eta_t=\frac{y_t}{\widetilde\sigma_t(\widehat{\bs\theta}_n)},\qquad t=1,\ldots,n,
\end{displaymath}
where $\widetilde\sigma_t$ considers arbitrary fixed initial values.
Let $h$ denote the instrumental density in the QML estimation, $g(y,s)=\log\big(s^{-1}h(y/s)\big)$, and
\begin{displaymath}
\widetilde G_n(\bs\theta)=\frac{1}{n}\sum_{t=1}^n g\big(y_t,\widetilde\sigma_t(\bs\theta)\big).
\end{displaymath}
The first-step QML estimator is then $\widehat{\bs\theta}_n=\underset{\bs\theta\in\Theta}{\arg\max}\,\widetilde G_n(\bs\theta)$. In addition, for generic $(\bs\theta,\xi)$, let $\widetilde Q_n(\xi,\bs\theta)=\frac1n\sum_{t=1}^n\widetilde\psi_{t,\tau}(\bs\theta,\xi)$, with $\widetilde\psi_{t,\tau}(\bs\theta,\xi)=\phi_\tau\Big(\frac{y_t}{\widetilde\sigma_t(\bs\theta)}-\xi\Big)$. For every $\bs\theta\in\Theta$, let ${\widehat{\textrm{XP}}_{\!\tau,\bs\theta}}$ denote the unique zero of $\widetilde Q_n(\cdot,\bs\theta)$, which we estimate by $\widehat\xi_n:={\widehat{\textrm{XP}}_{\!\tau,\widehat{\bs\theta}_n}}$ using the standardized residuals.
\begin{theorem}[Consistency] \label{thm:consistency_expectiles}
Let Assumptions~\ref{ass:innovation_process} to~\ref{ass:differentiability_instrumental} hold with $r=2$ in Assumption~\ref{ass:moments}. It then follows that $\widehat{\bs\theta}_n\asconv\bs\theta_0$ and $\widehat\xi_n\asconv\xi_0={\textrm{XP}}_{\!\tau}^\eta$.
\end{theorem}
Consistency follows under standard conditions. Assumption~\ref{ass:innovation_process} requires that innovations are independent and identically distributed (iid) with mean zero and unit variance, and that the innovation distribution is continuous at the expectile $\xi_0$. Assumptions~\ref{ass:stationarity_mixing} and~\ref{ass:compactness} dictate respectively that $y_t$ is a stationary and ergodic process, whereas $\bs\theta_0$ lies in the interior of a compact parameter space. Assumption~\ref{ass:volatility_process} constrains the volatility process to secure identification and differentiability. Assumption~\ref{ass:approximation} ensures that we can approximate the volatility process using a finite history. Assumption~\ref{ass:moments} requires moment conditions of order $r\ge2$. Finally, the identification and smoothness conditions in Assumptions~\ref{ass:identification_instrumental} and~\ref{ass:differentiability_instrumental} automatically hold for the Gaussian quasi-likelihood. See Appendix~\ref{sec:proof} for more details.
Next, we quantify how replacing the innovations by standardized residuals affects the estimation of the unconditional expectile in the second step. At the true parameter values, let $\bs{D}_t=\partial_{\bs\theta}\log\sigma_t(\bs\theta_0)$, $\bs{J}_0=\E(\bs{D}_t)$, $\psi_t=\phi_\tau(\eta_t-\xi_0)$, and $\Psi_0=\tau\{1-F_\eta(\xi_0)\}+(1-\tau)F_\eta(\xi_0)$. Define also $a(u)=\left.\frac{\partial}{\partial s}\,g(u,s)\right|_{s=1}$, $\bs{s}_t=a(\eta_t)\bs{D}_t$, $\bs{H}_0=-\E\left[\partial_{\bs\theta\bs\theta'}^2g\big(y_t,\sigma_t(\bs\theta_0)\big)\right]$, $\bs\Sigma_s=\E(\bs{s}_t\bs{s}_t')$, $\bs\sigma_{s\psi}=\E(\bs{s}_t\psi_t)$, $\bs\Sigma_\theta=\bs{H}_0^{-1}\bs\Sigma_s\bs{H}_0^{-1}$, $\bs\sigma_{\psi\theta}=\bs{H}_0^{-1}\bs\sigma_{s\psi}$, and $\sigma_\psi^2=\E(\psi_t^2)$. Asymptotic normality ensues if we strengthen the moment conditions: namely, by setting $r=4$ in Assumption~\ref{ass:moments} and by imposing Assumption~\ref{ass:qml_moments} on the moments of the QML score and curvature.
\begin{theorem}[Asymptotic normality] \label{thm:asymptotic_expectiles}
Let Assumptions~\ref{ass:innovation_process} to~\ref{ass:qml_moments} hold with $r=4$ in Assumption~\ref{ass:moments}. It then follows that $\sqrt{n}(\widehat\xi_n-\xi_0)\dconv\mathcal N(0,V_\xi)$, where
\begin{displaymath}
V_\xi=\frac{\sigma_\psi^2}{\Psi_0^2}+\xi_0^2\bs{J}_0'\bs\Sigma_\theta\bs{J}_0-2\frac{\xi_0}{\Psi_0}\,\bs{J}_0'\bs\sigma_{\psi\theta}.
\end{displaymath}
\end{theorem}
The first term in $V_\xi$ is the variance that would arise if we could observe the true innovations. The remaining terms capture the effects of estimating the parameters in the conditional variance. The latter involves not only the sampling error in the estimation of $\bs\theta_0$, but also how it correlates with the sampling error in the estimation of the unconditional expectile $\xi_0$.
To simplify notation, denote by $c_{n+1}=\sigma_{n+1}(\bs\theta_0)\xi_0$ the conditional expectile ${\textrm{XP}}_{\!\tau}(y_{t+1}|\mathcal{F}_t)$ at time $n+1$ given the current state at time $n$. To conduct inference on $\widehat{c}_{n+1}=\widetilde\sigma_{n+1}(\widehat{\bs\theta}_n)\widehat\xi_n$, let $D_{n+1}=\partial_{\bs\theta}\log\sigma_{n+1}(\bs\theta_0)$ and, for the finite-memory approximation in Assumption~\ref{ass:forecast}, let $D_{n+1}^{[m]}=\partial_{\bs\theta}\log\sigma_{n+1}^{[m]}(\bs\theta_0)$. Choose a deterministic sequence $m_n$ such that $m_n\to\infty$, $\frac{m_n}{n}\to0$, and define $\mathcal H_n=\sigma(\eta_{n-m_n+1}, \ldots,\eta_n)$. Let $d_{BL}$ denote the bounded-Lipschitz distance between probability laws on $\mathbb R$.
\begin{theorem}[Gaussian prediction intervals]\label{thm:conditional_expectiles}
Let Assumptions~\ref{ass:innovation_process} to~\ref{ass:qml_moments} and~\ref{ass:forecast} hold with $r=4$ in Assumption~\ref{ass:moments}, and define
\begin{eqnarray*}
V_{n+1}&=&\sigma_{n+1}^2(\bs\theta_0)\left[\Psi_0^{-2}\sigma_\psi^2+\xi_0^2(\bs{D}_{n+1}-\bs{J}_0)'\bs\Sigma_\theta (\bs{D}_{n+1}-\bs{J}_0)+2\xi_0\Psi_0^{-1}(\bs{D}_{n+1}-\bs{J}_0)'\bs\Sigma_{\psi\theta}\right]\\
V_{n+1}^{[m_n]}&=&\big(\sigma_{n+1}^{[m_n]}(\bs\theta_0)\big)^2\Big[\Psi_0^{-2}\sigma_\psi^2 +\xi_0^2(\bs{D}_{n+1}^{[m_n]}-\bs{J}_0)'\bs\Sigma_\theta(\bs{D}_{n+1}^{[m_n]}-\bs{J}_0)+2\xi_0\Psi_0^{-1} (\bs{D}_{n+1}^{[m_n]}-\bs{J}_0)'\bs\Sigma_{\psi\theta}\Big].
\end{eqnarray*}
It then follows that $d_{BL}\Big(\mathcal{L}\left\{\sqrt{n}(\widehat{c}_{n+1}-c_{n+1})\,\big|\,\mathcal{H}_n\right\},\mathcal{N}(0,V_{n+1}^{[m_n]})\Big)\pconv 0$. Moreover, it also holds that $V_{n+1}^{[m_n]}-V_{n+1}\pconv 0$, so that $d_{BL}\Big(\mathcal{L}\left\{\sqrt{n}(\widehat{c}_{n+1}-c_{n+1})\,\big|\, \mathcal{H}_n\right\},\mathcal{N}(0,V_{n+1})\Big)\pconv 0$.
\end{theorem}
Theorem~\ref{thm:conditional_expectiles} strengthens the finite-history approximation in Assumption~\ref{ass:approximation} by imposing the uniform contraction in Assumption~\ref{ass:forecast} to obtain asymptotically-valid Wald confidence intervals for one-step-ahead forecasts of the conditional expectile. Although they adequately account for the first-step estimation of the scale parameters, they are infeasible. Proposition~\ref{prop:variance_estimation} in Appendix~\ref{sec:proof} provides consistent estimators for each component of the asymptotic variance, which we denote by adding a hat. We next show that the resulting Wald statistic weakly converges to a standard Gaussian distribution, enabling us to compute feasible confidence intervals for the conditional expectiles. To do so, we introduce only one additional regularity condition, which targets the plug-in estimators of the asymptotic variance components (see Assumption~\ref{ass:plugin} in Appendix~\ref{sec:proof}).
\begin{corollary}[Feasible prediction interval for conditional expectiles]\label{cor:conditional_expectile_interval}
Suppose that Assumptions~\ref{ass:innovation_process} to~\ref{ass:forecast} hold with $r=4$ in Assumption~\ref{ass:moments}, and assume that $\bs\Sigma_Y$ is positive definite. Define $\widehat {V}_{n+1}=\widetilde\sigma_{n+1}^2(\widehat{\bs\theta}_n)\Big[ \widehat\Psi_n^{-2}\widehat\sigma_\psi^2+\widehat\xi_n^2(\widehat{\bs D}_{n+1}-\widehat{\bs J}_n)'\widehat{\bs\Sigma}_\theta(\widehat{\bs D}_{n+1}-\widehat{\bs J}_n)+2\widehat\xi_n\widehat\Psi_n^{-1}(\widehat{\bs D}_{n+1}-\widehat{\bs J}_n)'\widehat{\bs\sigma}_{\psi\theta}\Big]$, where $\widehat{\bs D}_{n+1}=\partial_{\bs\theta}\log\widetilde\sigma_{n+1}(\widehat{\bs\theta}_n)$. It then follows that $\widehat {V}_{n+1}-{V}_{n+1}\pconv0$ and that
\[
d_{BL}\left(\mathcal{L}\left\{\frac{\sqrt{n}(\widehat{c}_{n+1}-c_{n+1})}{\widehat{V}_{n+1}^{1/2}} \middle|\mathcal{\bs H}_n\right\},~\mathcal{N}(0,1)\right)\pconv0.
\]
This means that $\Pr\left\{ c_{n+1}\in\mathcal{I}_{n+1}(1-\alpha)\,\middle|\,\mathcal{H}_n\right\}\pconv 1-\alpha$, where
\[
\mathcal I_{n+1}(1-\alpha)=\left[\widehat{c}_{n+1}-z_{1-\alpha/2}\sqrt{\widehat{V}_{n+1}/n},~\widehat{c}_{n+1}+z_{1-\alpha/2} \sqrt{\widehat{V}_{n+1}/n}\right]
\]
with $z_{1-\alpha/2}$ denoting the $(1-\alpha/2)$-quantile of the standard Gaussian distribution.
\end{corollary}
In the next section, we evaluate the finite-sample distribution of the two-step estimator, paying close attention to whether the asymptotic variance estimator accurately quantifies the uncertainty in the expectile estimates.
\section{Monte Carlo study}\label{sec:mc}
We next assess how informative our asymptotic results are for the finite-sample behavior of the two-step estimators of both conditional and unconditional expectiles. In particular, we study the effects of sample size, volatility persistence, tail thickness, and leverage effects. The experiments distinguish inference for the unconditional expectile of the innovations from that for the more economically relevant one-step-ahead conditional expectile.
We generate returns according to
\begin{equation}
y_t=\sigma_t(\bs\theta_0)\eta_t,
\end{equation}
where $\{\eta_t\}$ is an independent and identically distributed sequence with zero mean and unit variance. The volatility process follows a GJR-GARCH(1,1) specification:
\begin{equation}
\sigma_t^2(\bs\theta_0)=\omega_0+\alpha_0 y_{t-1}^2+\gamma_0 y_{t-1}^2\mathbf{1}(y_{t-1}<0)+\beta_0\sigma_{t-1}^2(\bs\theta_0),
\label{eq:mc-gjr}
\end{equation}
where we calibrate $\omega_0$ to match an unconditional volatility of 20\% per year, assuming 252 trading days; $\alpha_0=0.05$; and $p_0=\alpha_0+\beta_0+\gamma_0/2$ dictates the amount of persistence under symmetric innovations. We entertain low and high levels of persistence by setting $p_0\in\{0.90,0.98\}$. In particular, we consider both symmetric and asymmetric responses to past returns: $\gamma_0=0$ and $\beta_0\in\{0.85,0.93\}$ for the standard GARCH, whereas $\gamma_0=0.08$ and $\beta_0\in\{0.81,0.89\}$ for the asymmetric GJR-GARCH specication. The latter is convenient because it isolates the incremental effect of leverage while fixing the overall persistence in volatility.
Finally, we draw innovations from a standard Gaussian distribution or from $t_\nu$ distributions with $\nu\in\{4,8\}$ degrees of freedom. We standardize the latter by $\sqrt{(\nu-2)/\nu}$ to ensure unit variance. The $t_8$ distribution provides a heavy-tailed benchmark design, whereas the $t_4$ distribution has finite variance but no finite fourth moment. It serves as a stress test that violates the sufficient conditions for QML inference.
For each volatility model, innovation distribution, and persistence level, we consider sample sizes of $n\in\{500,1000,2500,5000\}$ time-series data after a burn-in period of 500 observations. In each of the 10,000 replications, we estimate the correctly specified volatility model by Gaussian QML, imposing not only nonnegativity, but also $\alpha+\beta+\frac{\gamma}{2}\le0.999$, with $\gamma=0$ in the symmetric model. We then calculate the standardized residuals $\widehat\eta_t=y_t/\widetilde\sigma_t(\widehat{\bs\theta}_n)$ and their sample expectile $\widehat\xi_n\equiv{\widehat{\textrm{XP}}_{\!\tau,\widehat{\bs\theta}_n}}$ at $\tau=0.05$. In addition, we also estimate the conditional expectile $c_{n+1}=\sigma_{n+1}(\bs\theta_0)\xi_0$ at time $n+1$ by $\widehat{c}_{n+1}=\widetilde\sigma_{n+1}(\widehat{\bs\theta}_n)\widehat\xi_n$. Finally, we compute the population expectiles of the innovations from the truncated-moment equations of the Gaussian and $t$ distributions.
We assess the bias and root mean squared error (RMSE) of the conditional and unconditional expectile estimators, as well as the average length and coverage rate of their nominal 95\% confidence intervals. Table~\ref{tab:mc-xi-wald-portrait} reveals that bias and RMSE drop significantly as the sample size increases from 500 to 1,000 time-series observations, with a bias becoming negligible for samples of at least 2,500 observations. This pattern holds for both volatility specifications, as well as for every persistence level and distribution, even for the $t_4$ innovations that do not meet the regularity conditions in the asymptotic theory. In addition, the coverage rates of the 95\% Wald confidence intervals range from 0.907 to 0.951, improving monotonically with sample size. Interestingly, they are slightly better for the asymmetric GARCH specification. In contrast, increasing persistence or tail thickness slightly worsens coverage, especially for smaller sample sizes. Finally, the standard deviation of the Wald statistic $Z_{\xi,n}=\sqrt{n}(\widehat\xi_n-\xi_0)/\widehat{V}_\xi^{1/2}$ is very close to one across replications.
\begin{table}[!ht]\vspace*{1.5em}
\caption{Performance of the two-step estimator of the unconditional expectile at $\tau=0.05$\\ {\footnotesize The target is the innovation expectile $\xi_0\equiv{\textrm{XP}}_{\!\tau}^\eta$ at $\tau=0.05$, which we estimate by $\widehat\xi_n\equiv{\widehat{\textrm{XP}}_{\!\tau,\widehat{\bs\theta}_n}}$. For each of the 10,000 Monte Carlo replications, we report the bias and root mean squared error (RMSE) of the expectile estimator, as well as the average length and empirical coverage of the 95\% confidence interval, for each experimental design. The latter considers both standard GARCH and GJR-GARCH specifications, with different sample sizes and persistent levels (low and high). The innovations come either from a Gaussian distribution or from a $t$-student distribution (with 4 or 8 degrees of freedom).}}
\label{tab:mc-xi-wald-portrait}
\begin{adjustbox}{width=1\textwidth}
\begin{tabular}{llr*{9}{c}}
\toprule
&&&\multicolumn{4}{c}{GARCH}&&\multicolumn{4}{c}{GJR-GARCH}\\
\cline{4-7}\cline{9-12}
distribution & persistence & sample size~~ & bias & RMSE & SD($Z_{\xi,n}$) & coverage && bias & RMSE & SD($Z_{\xi,n}$) & coverage\\
\midrule
normal & low & 500~~ & -0.0057 & 0.0800 & 1.067 & 0.936 && -0.0048 & 0.0682 & 1.034 & 0.941\\
& & 1,000~~ & -0.0016 & 0.0378 & 1.017 & 0.945 && -0.0015 & 0.0373 & 1.010 & 0.945\\
& & 2,500~~ & -0.0005 & 0.0234 & 0.995 & 0.949 && -0.0006 & 0.0234 & 0.995 & 0.949\\
& & 5,000~~ & -0.0002 & 0.0166 & 0.998 & 0.949 && -0.0003 & 0.0166 & 0.998 & 0.949\\
\addlinespace[1pt]
& high & 500~~ & -0.0242 & 0.2802 & 1.089 & 0.921 && -0.0183 & 0.1211 & 1.053 & 0.927\\
& & 1,000~~ & -0.0092 & 0.0398 & 1.031 & 0.935 && -0.0086 & 0.0393 & 1.018 & 0.937\\
& & 2,500~~ & -0.0040 & 0.0240 & 1.000 & 0.944 && -0.0036 & 0.0240 & 0.999 & 0.945\\
& & 5,000~~ & -0.0020 & 0.0168 & 1.000 & 0.947 && -0.0018 & 0.0168 & 1.000 & 0.948\\
\addlinespace[3pt]
$t_8$ & low & 500~~ & -0.0075 & 0.1051 & 1.076 & 0.933 && -0.0088 & 0.1371 & 1.045 & 0.937\\
& & 1,000~~ & -0.0032 & 0.0468 & 1.027 & 0.942 && -0.0032 & 0.0459 & 1.017 & 0.944\\
& & 2,500~~ & -0.0011 & 0.0292 & 1.011 & 0.943 && -0.0013 & 0.0292 & 1.011 & 0.944\\
& & 5,000~~ & -0.0004 & 0.0204 & 0.994 & 0.951 && -0.0005 & 0.0204 & 0.994 & 0.950\\
\addlinespace[1pt]
& high & 500~~ & -0.0350 & 0.4175 & 1.092 & 0.917 && -0.0230 & 0.1471 & 1.064 & 0.922\\
& & 1,000~~ & -0.0111 & 0.0489 & 1.035 & 0.934 && -0.0105 & 0.0484 & 1.026 & 0.939\\
& & 2,500~~ & -0.0046 & 0.0298 & 1.016 & 0.940 && -0.0043 & 0.0298 & 1.015 & 0.940\\
& & 5,000~~ & -0.0022 & 0.0207 & 0.997 & 0.950 && -0.0020 & 0.0207 & 0.997 & 0.949\\
\addlinespace[3pt]
$t_4$ & low & 500~~ & -0.0432 & 1.1820 & 1.125 & 0.918 && -0.0317 & 1.0495 & 1.064 & 0.924\\
& & 1,000~~ & -0.0117 & 0.0671 & 1.065 & 0.929 && -0.0135 & 0.0653 & 1.051 & 0.930\\
& & 2,500~~ & -0.0064 & 0.0445 & 1.045 & 0.928 && -0.0076 & 0.0438 & 1.043 & 0.927\\
& & 5,000~~ & -0.0034 & 0.0325 & 1.037 & 0.935 && -0.0042 & 0.0319 & 1.035 & 0.936\\
\addlinespace[1pt]
& high & 500~~ & -0.0820 & 2.0999 & 1.118 & 0.907 && -0.0512 & 1.4201 & 1.082 & 0.908\\
& & 1,000~~ & -0.0204 & 0.0701 & 1.076 & 0.917 && -0.0218 & 0.0690 & 1.061 & 0.918\\
& & 2,500~~ & -0.0101 & 0.0450 & 1.047 & 0.922 && -0.0111 & 0.0444 & 1.042 & 0.922\\
& & 5,000~~ & -0.0053 & 0.0328 & 1.041 & 0.928 && -0.0060 & 0.0320 & 1.035 & 0.929\\
\bottomrule\\
\end{tabular}\end{adjustbox}\end{table}
Table~\ref{tab:mc-ce-wald-portrait} documents the performance of the conditional expectile estimator (bias and RMSE) across Monte Carlo replications for both GARCH-type specifications. It also reports the average width and empirical coverage of the 95\% Wald confidence interval. As before, bias and RMSE shrink substantially as the sample size increases from 500 to 1,000 time-series observations. This does not depend on the experimental design, remaining true for every persistence level, distribution, and volatility specification. A similar pattern arises for the 95\% Wald confidence intervals, which show remarkably little variation in coverage (between 0.927 and 0.952). Lastly, the standard deviation of $Z_{c,n+1}$ is slightly closer to unit than the standard deviation of the Wald statistic for the unconditional quantile.
\begin{table}[!ht]\vspace*{2em}
\caption{Performance of the two-step estimator of the conditional expectile at $\tau=0.05$\\ {\footnotesize The target is the conditional expectile $c_{n+1}\equiv\sigma_{n+1}\xi_0$, which we estimate by $\widehat{c}_{n+1}\equiv\widetilde\sigma_{n+1}(\widehat{\bs\theta}_n)\widehat{\xi}_n$. For each of the 10,000 Monte Carlo replications, we report the bias and root mean squared error (RMSE) of the conditional expectile estimator, as well as the average length and empirical coverage of the 95\% prediction interval, for each experimental design. The latter considers both standard GARCH and GJR-GARCH specifications, with different sample sizes and persistent levels (low and high). The innovations come either from a Gaussian distribution or from a $t$-student distribution (with 4 or 8 degrees of freedom).}}
\label{tab:mc-ce-wald-portrait}
\begin{adjustbox}{width=1\textwidth}
\begin{tabular}{llr*{9}{c}}
\toprule
&&&\multicolumn{4}{c}{GARCH}&&\multicolumn{4}{c}{GJR-GARCH}\\
\cline{4-7}\cline{9-12}
distribution & persistence & sample size & bias & RMSE & SD($Z_{c,n+1}$) & coverage && bias & RMSE & SD($Z_{c,n+1}$) & coverage\\
\midrule
normal & low & 500 & -0.0042 & 0.1161 & 1.046 & 0.938 && -0.0037 & 0.1273 & 1.045 & 0.940\\
& & 1,000 & 0.0003 & 0.0748 & 1.043 & 0.942 && -0.0001 & 0.0848 & 1.048 & 0.941\\
& & 2,500 & -0.0008 & 0.0477 & 1.018 & 0.951 && -0.0009 & 0.0537 & 1.008 & 0.949\\
& & 5,000 & -0.0002 & 0.0329 & 0.999 & 0.952 && -0.0004 & 0.0373 & 1.000 & 0.948\\
& high & 500 & -0.0248 & 0.2199 & 1.103 & 0.928 && -0.0234 & 0.1557 & 1.064 & 0.938\\
& & 1,000 & -0.0105 & 0.0779 & 1.043 & 0.936 && -0.0114 & 0.0881 & 1.032 & 0.941\\
& & 2,500 & -0.0052 & 0.0478 & 1.007 & 0.949 && -0.0055 & 0.0551 & 1.013 & 0.947\\
& & 5,000 & -0.0026 & 0.0335 & 1.002 & 0.950 && -0.0026 & 0.0374 & 1.000 & 0.951\\
$t_8$ & low & 500 & -0.0056 & 0.1560 & 1.071 & 0.935 && -0.0075 & 0.1929 & 1.069 & 0.938\\
& & 1,000 & -0.0024 & 0.1008 & 1.059 & 0.938 && -0.0019 & 0.1174 & 1.066 & 0.938\\
& & 2,500 & 0.0003 & 0.0625 & 1.031 & 0.940 && -0.0002 & 0.0722 & 1.028 & 0.943\\
& & 5,000 & 0.0002 & 0.0440 & 1.014 & 0.947 && -0.0001 & 0.0497 & 1.000 & 0.949\\
& high & 500 & -0.0352 & 0.3557 & 1.137 & 0.928 && -0.0284 & 0.1904 & 1.077 & 0.939\\
& & 1,000 & -0.0119 & 0.1066 & 1.058 & 0.938 && -0.0142 & 0.1292 & 1.049 & 0.942\\
& & 2,500 & -0.0048 & 0.0625 & 1.023 & 0.946 && -0.0053 & 0.0731 & 1.018 & 0.947\\
& & 5,000 & -0.0022 & 0.0448 & 1.016 & 0.946 && -0.0025 & 0.0494 & 1.002 & 0.949\\
$t_4$ & low & 500 & -0.0150 & 0.2437 & 1.107 & 0.927 && -0.0153 & 0.2566 & 1.102 & 0.931\\
& & 1,000 & -0.0058 & 0.1620 & 1.095 & 0.932 && -0.0075 & 0.1895 & 1.094 & 0.936\\
& & 2,500 & -0.0044 & 0.1104 & 1.051 & 0.942 && -0.0059 & 0.1283 & 1.038 & 0.945\\
& & 5,000 & -0.0015 & 0.0831 & 1.018 & 0.945 && -0.0033 & 0.1024 & 1.019 & 0.944\\
& high & 500 & -0.0440 & 0.6450 & 1.147 & 0.927 && -0.0293 & 0.2531 & 1.097 & 0.933\\
& & 1,000 & -0.0168 & 0.1764 & 1.094 & 0.932 && -0.0196 & 0.2046 & 1.074 & 0.938\\
& & 2,500 & -0.0065 & 0.1196 & 1.020 & 0.949 && -0.0073 & 0.1258 & 1.012 & 0.949\\
& & 5,000 & -0.0039 & 0.1045 & 1.013 & 0.946 && -0.0057 & 0.1309 & 1.008 & 0.948\\
\bottomrule
\end{tabular}\end{adjustbox}\end{table}
This improvement relative to the unconditional expectile estimator is consistent with the large-sample theory for two-step estimators of conditional risk measures. The asymptotic variance accounts for the uncertainty both in the estimation of the innovation expectile and in the one-step-ahead volatility forecast. The influence function of the conditional expectile attenuates part of the first-step estimation error that remains visible in the unconditional expectile estimator, bringing the dispersion of the Wald statistic closer to one and coverage closer to nominal levels.
We complement the analysis in two ways. First, we compute the skewness and kurtosis of the Wald statistics $Z_{\xi,n}=\sqrt{n}(\widehat\xi_n-\xi_0)/\widehat{V}_\xi^{1/2}$ and $Z_{c,n+1}=\sqrt{n}(\widehat{c}_{n+1}-c_{n+1})/\widehat{V}_{n+1}^{1/2}$ across replications. We find that their skewness and excess kurtosis are close to zero, especially for samples of at least 1,000 observations, confirming that the Gaussian approximation performs very well in finite samples. Second, we also look at the widths of the 95\% prediction intervals of the conditional expectiles at $\tau=0.05$ relative to those of the value at risk and expected shortfall at $\alpha=0.05$, which we compute using \pc{gao2008estimation} standard errors. Figure~\ref{fig:mc_interval_widths_level} reveals that they are relatively tighter for the conditional expectile at $\tau=0.05$ than for the value at risk and, especially, expected shortfall at $\alpha=0.01$. This is across the board, regardless of the sample size, innovation distribution, persistence level, and volatility specification. Moreover, further simulations show that fixing both quantile and expectile levels to 1\% increases the expectile advantage over the expected shortfall, especially for $t$ innovations.
Although imposing $\alpha=\tau$ is appropriate for statistical comparisons, it implies completely different capital requirements. To make them comparable, we can fix the value at risk and then compute $\tau(\alpha)$ such that the expectile and value at risk coincide, as in~\ref{eq:tau equivalence}. The regulatory guidelines of the Bank for International Settlements suggest that the expected shortfall at the 2.5\% quantile level imposes a similar capital requirement to the value at risk at 1\% (\cnm{BIS2013}, 2013, \cy{BIS2014}). We thus compare their 95\% prediction intervals to those of the conditional expectile of $\tau(0.01)$, which typically takes a more extreme value between 0.33 and 0.66 across Monte Carlo experiments. Figure~\ref{fig:mc_interval_widths_capital} shows that the average lengths of the prediction intervals are not so different for sample sizes of 500 observations, despite the large differences in expectile and quantile levels (e.g., almost tenfold between XP and ES). For larger samples, the average length of the XP prediction intervals become between 20\% and 25\% larger than those of the value at risk and expected shortfall.
\begin{figure}[!ht]
\centering\caption{Relative lengths of the prediction intervals of the tail risk measures with $\alpha=\tau=0.05$\\
{\footnotesize The plots summarize the replication-level ratio between the lengths of the nominal 95\% prediction intervals of the conditional expectile, value at risk and expected shortfall with $\alpha=\tau=0.05$. Top row refers to the ratios with respect to the value-at-risk interval length, whereas bottom rows to the expected shortfall.}}
\includegraphics[height=.375\textheight,trim= 0 .25cm 0 .5cm,clip]{figures/mc_interval_length_ratios_0p05.pdf}
\label{fig:mc_interval_widths_level}
\end{figure}
\begin{figure}[!ht]\vspace*{.375cm}
\centering\caption{Relative lengths of the prediction intervals given comparable capital requirements\\
{\footnotesize The plots summarize the replication-level ratio between the lengths of the nominal 95\% prediction intervals of the value at risk at $\alpha=0.01$, expected shortfall at $\alpha=0.025$, and conditional expectile at $\tau(0.01)$. Top row refers to the ratios with respect to the value-at-risk interval length, whereas bottom rows to the expected shortfall.}}
\includegraphics[height=.375\textheight,trim = 1.7cm 3.25cm 1.8cm 3cm,clip]{figures/mc_interval_length_ratios_capital.pdf}
\label{fig:mc_interval_widths_capital}
\end{figure}
\section{Tail risk in cryptomarkets}\label{sec:crypto_application_clt}
We collect data from Yahoo Finance on the daily values of the cryptocurrencies with the highest market values at the end of 2023: namely, Bitcoin (BTC), Ethereum (ETH), and Binance Coin (BNB).\footnote{~~Yahoo Finance sources their crypto data from CoinMarketCap. Quantitative results remain virtually the same using Coin Codex data, alleviating concerns with crypto data quality \cite{SSY26}.} For the sake of comparison, we also gather the daily exchange rate between the Euro and US dollar (EUR). The sample ranges from January 2016 to December 2023 for EUR and BTC, but only from November 2017 to December 2023 for BNB and ETH.
Figure~\ref{fig:drawdown} documents this difference in a more direct manner by plotting the distribution of their historical drawdowns. We calculate the drawdown at time $t$ as the difference between the maximum cumulative return up to time $t$ and the cumulative return at time $t$, so that it measures the percentage drop from the all-time high. Crypto drawdowns are much larger (and more persistent) than EUR drawdowns, which remain comparatively modest throughout the sample period.
To deal with nonstationarity in the exchange rates, we compute their daily percentage changes (in logs): i.e., $y_t=100\ln(P_t/P_{t-1})$, where $P_t$ is the value of the currency in US dollars at time $t$. Table~\ref{tab:summary} documents that sample sizes are larger for crypto assets because they trade every day of the week, even on bank holidays. However, the most striking feature is the range in which daily cryptocurrency returns dwell. Their minimum values reflect drops in value of about 50\%, with highest returns varying from 22.5\% to more than 53\%, in just one day! This variation leads to very high levels of daily volatility, from 3.75\% to 5.46\%.
In line with the risk-return tradeoff, typical daily returns are also very high, with average and median values that vary between 9\% and 23\% and between 8\% and 15\%, respectively. In comparison, log-changes in the EUR range from -2.81\% to 1.82\%, with zero mean and a daily volatility of 0.47\%. The excess kurtosis in every exchange rate is consistent with heavy tails and conditional heteroskedasticity. In turn, skewness is particularly strong for cryptocurrencies, which might indicate that volatility responds asymmetrically to positive and negative returns. These features obviously invalidate any risk assessment based exclusively on volatility, instead calling for tail risk measures.\newpage
\begin{figure}[!ht]
\centering\caption{Daily exchange rates vis-à-vis the US dollar from January 2016 to December 2023\\
{\footnotesize We plot the value of Binance Coin (BNB), Bitcoin (BTC), Ethereum (ETH), and Euro (EUR) in US dollar.}}
\includegraphics[width=\linewidth, trim = 1cm 4cm 1cm 3cm, clip]{figures/figure_1_exchange_rates_bw.pdf}
\label{fig:prices}
\end{figure}
\begin{figure}[!ht]
\centering\caption{Box plots of the historical drawdowns from 2018 and 2023\\
{\footnotesize We plot the distributions of the historical drawdowns of the BNB, BTC, ETH and EUR exchange rates. We start as from January 2018 to have at least one year of observations for each currency in the initial sample.}}
\includegraphics[height=8.75cm,trim = 1cm 2.75cm 1cm 3cm,clip]{figures/figure_2_drawdowns_bw.pdf}
\label{fig:drawdown}
\end{figure}
\begin{table}[!ht]
\vspace*{1em}\caption{Descriptive statistics for the daily log-changes in exchange rates\\ {\footnotesize We report the number of observations (\#obs) for each time series, as well as their main summary statistics. In particular, we display not only their minimum (min), average (mean) and maximum (max) values, but also their standard deviation (std dev), skewness (skew), kurtosis (kurt), and quartiles ($q_{0.25}$, median and $q_{0.75}$). The sample ranges from January 2016 to December 2023 for EUR and BTC, and from November 2017 to December 2023 for BNB and ETH.}}
\begin{adjustbox}{width=1\textwidth}
\begin{tabular}{p{1.85cm}crrrrrrrrr}
\hline
currency & \#obs & min & $q_{0.25}$ & median & $q_{0.75}$ & max & mean & std dev & skew & kurt\\
\hline
BNB & 2,243 & -54.31 & -1.88 & 0.10 & 2.35 & 52.92 & 0.23 & 5.46 & 0.40 & 20.03\\
BTC & 2,921 & -46.47 & -1.27 & 0.15 & 1.69 & 22.51 & 0.16 & 3.75 & -0.72 & 14.82\\
ETH & 2,243 & -55.07 & -1.90 & 0.08 & 2.38 & 23.47 & 0.09 & 4.80 & -0.93 & 13.84\\
EUR & 2,082 & -2.81 & -0.28 & 0.00 & 0.28 & 1.82 & 0.00 & 0.47 & -0.05 & 4.90\\
\hline
\end{tabular}\end{adjustbox}
\label{tab:summary}
\end{table}
Further analysis shows little autocorrelation, but strong persistence in conditional volatility. Given the skewness in the crypto assets, we estimate GJR-GARCH(1,1) models for each currency using Gaussian QML over rolling windows of 1,000 observations. The asymmetric volatility specification matters because ignoring leverage effects could well distort tail forecasts and their prediction intervals. For each window, we first compute standardized residuals and then estimate not only their unconditional expectile, but also the corresponding conditional expectiles by plugging in one-step-ahead volatility forecasts. Finally, it is straightforward to compute their 90\% prediction intervals using the plug-in variance estimators.
Figure~\ref{fig:empirical confidence interval} depicts the one-step-ahead forecasts of the value at risk, expected shortfall and expectile at level $\alpha=\tau=0.01$ from January 2, 2020 to December 30, 2023. It is apparent that the expected shortfall spikes more sharply than the other tail risk measures, whereas the conditional expectile has the smoothest variation. Moreover, the latter prediction intervals are much tighter than those of the quantile-based risk measures: namely, between 66.3\% and 89.8\% of the length for the value at risk and from 35.2\% to 45.0\% of that for the expected shortfall. We assess the one-step ahead forecasts of the value at risk and expectile at $\alpha=\tau=0.01$ using statistical tests and backtest procedures. In particular, we check the value-at-risk performance using coverage and duration tests \cite{kupiec1995techniques,christoffersen1998evaluating,christoffersen2004backtesting}, whereas we conduct \pc{mcneil2000estimation} bootstrap-based residual tests to evaluate conditional expectiles.
\begin{figure}[!ht]
\caption{Conditional tail risk measures at level $\alpha=\tau=0.01$\\
{\footnotesize The plots in the first column display the daily percentage changes in the value of Binance Coin (BNB), Bitcoin (BTC), Ethereum (ETH), and Euro (EUR) in US dollars from January 2020 to December 2023, as well as the one-day-ahead forecasts of the value at risk, expected shortfall and expectile at level $\alpha=\tau=0.01$. In the second column, we depict how the expectile level $\tau(\alpha)$ should vary over time to make expectile and value at risk coincide.}}
\includegraphics[width=\textwidth,trim = 1cm 0 0 0,clip]{figures/figure_3_conditional_tail_risk_color.pdf}\vspace*{-1.5em}
\label{fig:empirical confidence interval}
\end{figure}\clearpage
\begin{table}[!th]\centering\caption{Frequency of value at risk and expectile forecast exceedances at $\alpha=\tau=0.01$\\ {\footnotesize We report the expected and realized number of exceedances for the one-step-ahead forecasts of the value at risk and expectile based on rolling windows of 1,000 observations. Exceedances occur if realized loss exceeds the tail risk measure. Apart from the corresponding relative frequency (i.e., number of realized exceedances over number of out-of-sample forecasts), we also document the p-values of the coverage tests for the value at risk and the bootstrap-based p-value of the residual test for the expectile.}}
\label{tab:app_backtest}
\begin{tabular}{lccrccrccrcc}
\toprule
& \multicolumn{2}{c}{BNB} && \multicolumn{2}{c}{BTC} && \multicolumn{2}{c}{ETH} && \multicolumn{2}{c}{EUR}\\
\cline{2-3}\cline{5-6}\cline{8-9}\cline{11-12}
daily exceedances & VaR & XP && VaR & XP && VaR & XP && VaR & XP\\
\midrule
expected & 12.43 & 12.43 && 19.21 & 19.21 && 12.43 & 12.43 && 10.82 & 10.82\\
realized & 10 & 16 && 20 & 36 && 8 & 19 && 14 & 36\\
relative frequency & 0.80\% & 1.29\% && 1.04\% & 1.87\% && 0.64\% & 1.53\% && 1.29\% & 3.33\%\\
coverage tests\\
~~~~unconditional & 0.47 & && 0.86 & && 0.18 & && 0.35 &\\
~~~~conditional & 0.71 & && 0.80 & && 0.38 & && 0.54 &\\
~~~~duration & 0.03 & && 0.30 & && 0.02 & && 0.81 &\\
residual test & & 0.48 && & 0.79 && & 0.06 && & 0.98\\
\bottomrule
\end{tabular}\end{table}
Table~\ref{tab:app_backtest} shows that number of VaR exceedances is close to the nominal one-percent target: 1.05\% for BNB, 1.04\% for BTC, 0.72\% for ETH, and 1.20\% for EUR. As expected, there are more exceedances for the one-step-ahead forecasts of the expectiles at level $\tau=1\%$ in view that setting $\tau=\alpha$ does not impose equivalent capital requirements. Finally, the coverage and residual tests cannot reject the congruency of the one-step-ahead forecasts of the tail risk measures at the 5\% significance level.
\begin{table}[!htb]\vspace*{1em}\centering\caption{Prediction intervals of the tail risk measures with comparable capital requirements\\ {\footnotesize We consider the value at risk at the 1\% quantile level, the expected shortfall at the 2.5\% quantile level, and the expectile at the level $\tau(\alpha)=\tau(0.01)$ such that it coincides with the value at risk at level $\alpha=0.01$. We report the number of rolling windows (\#windows) we deploy to produce out-of-sample forecasts of the tail risk measures, as well as the average value of $\tau(\alpha)$. We also display not only the average difference between capital requirements and realized losses (capital buffer), but also the average width of the 95\% prediction intervals of the one-step-ahead forecasts of each tail risk measure (interval width).}}
\label{tab:comparison_risk_measures}
\begin{tabular}{lccrcccrccc}
\toprule
&&&& \multicolumn{3}{c}{capital buffer} && \multicolumn{3}{c}{interval width}\\
\cline{5-7}\cline{9-11}
currency~~~~ & \#windows & $\bar\tau(\alpha)=\bar\tau(0.01)$ && VaR & ES & XP && VaR & ES & XP\\
\midrule
BNB & 1,243 & 0.47\% && 13.97 & 14.70 & 13.97 && 7.86 & 6.33 & 7.21\\
BTC & 1,921 & 0.34\% && 10.69 & 11.24 & 10.69 && 4.50 & 5.79 & 7.64\\
ETH & 1,243 & 0.32\% && 14.02 & 14.04 & 14.02 && 6.86 & 6.88 & 8.38\\
EUR & 1,082 & 0.38\% && 1.17 & 1.26 & 1.17 && 0.35 & 0.41 & 0.54\\
\bottomrule
\end{tabular}
\end{table}
Although useful for calibration, the above exceedance analysis does not consider comparable capital requirements given that we use the same quantile/expectile levels for each tail risk measure \cite{bellini2017risk}. Figure~\ref{fig:empirical confidence interval} indeed shows how the expectile levels of each currency must change over time to match the VaR capital requirements. We then fix the same capital requirement for the value at risk and expectile by setting $\alpha=0.01$ and $\tau(\alpha)=\tau(0.01)$ in \eqref{eq:tau equivalence}. In addition, we also compute the expected shortfall at the 2.5\% quantile level following the regulatory guidelines in \cnm{BIS2013} (2013, \cy{BIS2014}). Table~\ref{tab:comparison_risk_measures} shows not only that the average values of $\tau(\alpha)=\tau(0.01)$ vary from 0.32\% to 0.47\% across exchange rates, but also that the average differences between the ES$_{0.025}$ and VaR$_{0.01}$ capital requirements are quite small in percentage losses: 73 bps for BNB, 55 bps for BTC, 2 bps for ETH, and 9 bps for EUR. Lastly, as expected from the Monte Carlo results, the prediction intervals of the conditional expectiles are relatively wider on average than those of the value at risk and expected shortfall because $\tau(\alpha)$ is much smaller than $\alpha=0.01$.
\section{Conclusion}\label{sec:conclusion}
In this paper, we consider a two-step approach for the estimation of conditional expectiles. In particular, we first estimate the parameters of the conditional variance by quasi-maximum likelihood and then employ standardized residuals to compute the unconditional expectiles of the innovations. We quantify the impact of first-step estimation error on the asymptotic distribution of the expectile estimator, showing how to compute consistent standard errors for the unconditional expectile and prediction intervals for the conditional expectiles. Empirically, we assess how the conditional expectile performs in comparison with traditional quantile-based risk measures in the analysis of daily cryptocurrency returns.
There are at least two directions worth exploring in future research. The first is to extend our asymptotic theory to deal with nonzero expected returns \cite{francq2018estimation}. The second is to study the high-dimensional case of large portfolios \cite{francq2020virtual}. As it happens, expectiles are particularly suitable for handling several assets because they are coherent measures even under nonelliptical distributions \cite{artzner1999coherent}.
\lineskip=.5ex \baselineskip 3.75ex
\bibliography{ALS}
\newpage