EconBase
← Back to paper

Extremal Quantile Regression: An Overview

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

80,425 characters · 26 sections · 86 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Extremal Quantile Regression: An Overview

abstractExtremal quantile regression, i.e. quantile regression applied to the tails of the conditional distribution, counts with an increasing number of economic and financial applications such as value-at-risk, production frontiers, determinants of low infant birth weights, and auction models. This chapter provides an overview of recent developments in the theory and empirics of extremal quantile regression. The advances in the theory have relied on the use of extreme value approximations to the law of the Koenker and Bassett (1978) quantile regression estimator. Extreme value laws not only have been shown to provide more accurate approximations than Gaussian laws at the tails, but also have served as the basis to develop bias corrected estimators and inference methods using simulation and suitable variations of bootstrap and subsampling. The applicability of these methods is illustrated with two empirical examples on conditional value-at-risk and financial contagion.

Introduction

In 1895, the Italian econometrician Vilfredo Pareto discovered that the power law describes well the tails of income and wealth data. This simple observation stimulated further applications of the power law to economic data including zipf:1949, mandelbrot:1963, fama:1965, praetz:1972, sen:1973, and longin:1996, among many others. It also opened up a theory to analyze the properties of the tails of the distributions so-called Extreme Value (EV) theory, which was developed by gnedenko:1943 and dehaan:1970. jansen:1991 applied this theory to analyze the tail properties of US financial returns and concluded that the 1987 market crash was not an outlier; rather, it was a rare event whose magnitude could have been predicted by prior data. This work stimulated numerous other studies that rigorously documented the tail properties of economic data embrechts:1997.

victor:annals extended the EV theory to develop extreme quantile regression models in the tails, and analyze the properties of the koenker:1978 quantile regression estimator, called extremal quantile regression. This work builds especially upon feigin:1994 and knight:linear, which studied the most extreme, frontier regression case in the location model. Related results for the frontier case -- the regression frontier estimators -- were developed by smith:1994, chernozhukov:1998, jureckova:1999, and portnoy:jur. portnoy-koenker:1989 and gjkp:1993 implicitly contained some results on extending the normal approximations to intermediate order regression quantiles (moderately extreme quantiles) in location models. The complete theory for intermediate order regression quantiles was developed in victor:annals. jureckova:2016 recently characterized properties of averaged extreme regression quantiles.

In this chapter we review the theory of extremal quantile regression. We start by introducing the general setup that will be used throughout the chapter. Let $Y$ be a continuous response variable of interest with distribution function $F_Y(y) = {\mathrm{P}}(Y \leq y)$. The marginal $\tau$-quantile of $Y$ is the left-inverse of $y \mapsto F_Y(y)$ at $\tau$, that is $Q_Y(\tau) := \inf\{y: F_{Y}(y) \geq \tau\}$ for some $\tau \in (0,1).$ Let $X$ be a $d_x$-dimensional vector of covariates related to $Y$, $F_X$ be the distribution function of $X$, and $F_{Y}(y | x)= {\mathrm{P}}(Y \leq y | X=x)$ be the conditional distribution function of $Y$ given $X=x$. The conditional $\tau$-quantile of $Y$ given $X=x$ is the left-inverse of $y \mapsto F_{Y}(y | x)$ at $\tau$, that is $Q_Y(\tau | x) := \inf\{y: F_{Y}(y | x) \geq \tau\}$ for some $\tau \in (0,1).$ We refer to $x \mapsto Q_{Y}(\tau|x)$ as the $\tau$-quantile regression function. This function measures the effect of $X$ on $Y$, both at the center and at the tails of the outcome distribution. A marginal or conditional $\tau$-quantile is extremal whenever the probability index $\tau$ is either close to zero or close to one. Without loss of generality, we focus the discussion on $\tau$ close to zero.

The analysis of the properties of the estimators of extremal quantiles relies on EV theory. This theory uses sequences of quantile indexes $\{\tau_T\}_{T=1}^{\infty}$ that change with the sample size $T$. Let $\tau_T T$ be the order of the $\tau_T$-quantile. A sequence of quantile index and sample size pairs $\{\tau_T, T\}_{T=1}^\infty$ is said to be an {\it extreme order} sequence if $\tau_T \searrow 0$ and $ \tau_T T \rightarrow k \in (0, \infty)$ as $T \to \infty$; an {\it intermediate order} sequence if $\tau_T \searrow 0$ and $ \tau_T T \rightarrow \infty$ as $T \to \infty$; and a {\it central order} sequence if $\tau_T$ is fixed as $T \rightarrow \infty$. In this chapter we show that each of these sequences produce different asymptotic approximations to the distribution of the quantile regression estimators. The extreme order sequence leads to an EV law in large samples, whereas the intermediate and central sequences lead to normal laws. The EV law provides a better approximation to the extremal quantile regression estimators.

We conclude this introductory section with a review of some applications of extremal quantile regression to economics and finance.

example[Conditional Value-at-Risk] The Value-at-Risk (VaR) analysis seeks to forecast or explain low quantiles of future portfolio returns of an institution, $Y$, using current information, $X$ vl,engle:2004. Typically, the extremal $\tau$-quantile regression functions $x \mapsto Q_Y(\tau | x)$ with $\tau=0.01$ and $\tau=0.05$ are of interest. The VaR is a risk measure commonly used in real-life financial management, insurance, and actuarial science embrechts:1997. We provide an empirical example of VaR in Section (ref).
example[Determinants of Birthweights] In health economics, we may be interested in how smoking, absence of prenatal care, and other maternal behavior during pregnancy, $X$, affect infant birthweights, $Y$ abrevaya. Very low birthweights are connected with subsequent health problems and therefore extremal quantile regression can help identify factors to improve adult health outcomes. chernozhukov:2011 provide an empirical study of the determinants of extreme birthweights.
example[Probabilistic Production Frontiers] An important application to industrial organization is the determination of efficiency or production frontiers, pioneered by aigner:1968, timmer, and aigner:1976. Given the cost of production and possibly other factors, $X$, we are interested in the highest production levels, $Y$, that only a small fraction of firms, the most efficient firms, can attain. These (nearly) efficient production levels can be formally described by the extremal $\tau$-quantile regression function $x \mapsto Q_Y(\tau | x)$ for $\tau \in [1-\varepsilon, 1)$ and $\varepsilon > 0$; so that only a $\varepsilon$-fraction of firms produce $Q_Y(\tau | X)$ or more.
example[Approximate Reservation Rules] In labor economics, flinn:1982 proposed a job search model with approximate reservation rules. The reservation rule measures the wage level, $Y$, below which a worker with characteristics, $X$, accepts a job with small probability $\varepsilon$, and can by described by the extremal $\tau$-quantile regression $x \mapsto Q_Y(\tau | x)$ for $\tau \in (0,\varepsilon]$.
example[Approximate $(S,s)$-Rules] The $(S,s)$-adjustment models arise as an optimal policy in many economic models arrow:ss. For example, the capital stock, $Y$, of a firm with characteristics $X$ is adjusted sharply up to the level $S(X)$ once it has depreciated below some low level $s(X)$ with probability close to one, $1-\varepsilon$. This conditional $(S,s)$-rule can be described by the extremal $\tau$-quantile regression functions $x \mapsto Q_Y(\tau|x)$ and $x \mapsto Q_Y(1-\tau|x)$ for $\tau \in (0,\varepsilon]$.
example[Structural Auction Models] Consider a first-price procurement auction where bidders hold independent valuations. dp:superconsistent modelled the winning bid, $Y$, as $Y = c(Z) \beta(N) + \varepsilon$, where $c(Z)$ is the efficient cost function that depends on the bid characteristics, $Z$, $\beta(N) \geq 1$ is a mark-up that approaches 1 as the number of bidders $N$ approaches infinity, and the disturbance $\varepsilon$ captures small bidding mistakes independent of $Z$ and $N$. By construction, the structural function $(z,n) \mapsto c(z) \beta(n)$ corresponds to the extremal quantile regression function $x \mapsto Q_Y(\tau | x)$ for $x = (z,n)$ and $\tau \in (0 , {\mathrm{P}}(\varepsilon \leq 0)]$.
example[Other Recent Applications] Following the pioneering work of powell:1984, altonji:2012 applied extremal quantile regression to estimate extensive margins of demand functions with corner solutions. d'haultfoeuille:2015 used extremal quantile regression to deal with endogenous sample selection under the assumption that there is no selection at very high quantiles of the response variable $Y$ conditional on covariates $X$. They applied this approach to the estimation of the black wage gap for young males in the US. zhang:2015 employed extremal quantile regression methods to estimate tail quantile treatment effects under a selection on observables assumption.

Notation: The symbol $\to_d$ denotes convergence in law. For two real numbers $a,b$, $a \ll b$ means that $a$ is much less than $b$. More notation will be introduced when it is first used.

Outline: The rest of the chapter is organized as follows. Section (ref) reviews models for marginal and conditional extreme quantiles. Section (ref) describes estimation and inference methods for extreme quantile models. Section (ref) presents two empirical applications of extremal quantile regression to conditional value-at-risk and financial contagion.

Extreme Quantile Models

This section reviews typical modeling assumptions in extremal quantile regression. They embody Pareto conditions on the tails of the distribution of the response variable $Y$ and linear specifications for the $\tau$-quantile regression function $x \mapsto Q_Y(\tau|x)$.

Pareto-Type and Regularly Varying Tails

The theory for extremal quantiles often assumes that the tails of the distribution of $Y$ have Pareto-type behavior, meaning that the tails decay approximately as a power function, or more formally, a regularly varying function. The Pareto-type tails encompass a rich variety of tail behaviors, from thick to thin tailed distributions, and from bounded to unbounded support distributions.

: Define the variable $U$ by $U := Y$ if the lower end-point of the support of $Y$ is $-\infty$ and by $U := Y-Q_Y(0)$ if the lower end-point of the support of $Y$ is finite. In words, $U$ is a shifted copy of $Y$ whose support ends at either $-\infty$ or $0$. The assumption that the random variable $U$ exhibits a {\em Pareto-type tail} is stated by the following two equivalent conditions: \footnote{$a \sim b$ means that $a/b \to 1$ with an appropriate notion of limit.}

alignat{4} Q_U(\tau) &\sim L(\tau) \cdot \tau^{-\xi} &\qquad as \qquad && \tau &\searrow 0, \\ F_U(u) &\sim \bar{L}(u) \cdot u^{-1/\xi} &\qquad as \qquad && u &\searrow Q_U(0),

for some $\xi \neq 0$, where $\tau \mapsto L(\tau)$ is a non-parametric slowly-varying function at $0$, and $u \mapsto \bar{L}(u)$ is a non-parametric slowly-varying function at $Q_U(0)$. \footnote{A function $z \mapsto f(z)$ is said to be {\em slowly-varying} at $z_0$ if $\lim_{z \searrow z_0} f(z)/f(mz) = 1$ for every $m > 0$} The leading examples of slowly-varying functions are the constant function and the logarithmic function. The number $\xi$ as defined in ((ref)) or ((ref)) is called the {\em extreme value (EV) index} or the {\em tail index}.

The absolute value of $\xi$ measures the heavy-tailedness of the distribution. The support of a Pareto-type tailed distribution necessarily has a finite lower bound if $\xi < 0$ and an infinite lower bound if $\xi > 0$. Distributions with $\xi>0$ include stable, Pareto, Student's $t$, and many others. For example, the $t$-distribution with $\nu$ degrees of freedom has $\xi=1/\nu$ and exhibits a wide range of tail behaviors. In particular, setting $\nu =1$ yields the Cauchy distribution which has heavy tails with $\xi=1$, while setting $\nu =30$ gives a distribution that has light tails with $\xi =1/30$ and is very close to the normal distribution. On the other hand, distributions with $\xi < 0$ include the uniform, exponential, Weibull, and many others.

The assumption of Pareto-type tails can be equivalently cast in terms of a regular variation assumption, as is commonly done in the EV theory. A distribution function $u \mapsto F_U(u)$ is said to be {\em regularly varying} at $u = Q_U(0)$ with index of regular variation $-1/\xi$ if $$\lim_{y\searrow Q_U(0)} F_U(ym) / F_U(y) = {m}^{-1/\xi} \qquad \text{for every} \qquad m > 0.$$ This condition is equivalent to the regular variation of the quantile function $\tau \mapsto Q_U(\tau)$ at $\tau = 0$ with index $-\xi$, $$\lim_{\tau \searrow 0} Q_U(\tau m) / Q_U(\tau) = {m}^{-\xi} \qquad \text{for every} \qquad m > 0.$$

The case of $\xi=0$ corresponds to the class of rapidly varying distribution functions. Such distribution functions have exponentially light tails, with the normal and exponential distributions being the chief examples. For the sake of simplicity, we omit this case from our discussion. Note, however, that since the limit distribution of the main statistics is continuous in $\xi$, the inference theory for $\xi = 0$ is included by taking $\xi \to 0$.

Extremal Quantile Regression Models

The most common model for the quantile regression (QR) function is the linear in parameters specification

equation[equation omitted — 132 chars of source]

and for every $x \in \mathbf{X}$, the support of $X$. This linear functional form not only provides computational convenience but also has good approximation properties. Thus, the set $B(x)$ can include transformations of $x$ such as polynomials, splines, indicators or interactions such that $x \mapsto B(x)' \beta(\tau)$ is close to $x \mapsto Q_Y(\tau | x)$. In what follows, without loss of generality we lighten the notation by using $x$ instead of $B(x)$. We also assume that the $d_x$-dimensional vector $x$ contains a constant as the first element, has a compact support $\mathbf{X}$, and satisfies the regularity conditions stated in Assumption 3 of chernozhukov:2011. Compactness is needed to ensure the continuity and robustness of the mapping from extreme events in $Y$ to the extremal QR statistics. Even if $\mathbf{X}$ is not compact, we can select the data for which $X$ belongs to a compact region.

The main additional assumption for extremal quantile regression is that $Y$, transformed by some auxiliary regression line $X'\beta_e$, has Pareto-type tails. More precisely, together with ((ref)), it assumes that there exists an auxiliary regression parameter $\beta_e \in \mathbb{R}^{d_x}$ such that the disturbance $V := Y - X'\beta_e$ has lower end point $s = 0$ or $s = -\infty$ a.s., and its conditional quantile function $Q_V(\tau | x)$ satisfies the tail equivalence relationship:

equation[equation omitted — 179 chars of source]

for some quantile function $Q_U(\tau)$ that exhibits a Pareto-type tail ((ref)) with EV index $\xi$, and some vector parameter $\gamma$ such that $E[X]' \gamma = 1$ and $X' \gamma > 0$ a.s.

Condition ((ref)) imposes a location-scale shift model. This model is more general than the standard location shift model that replaces $x' \gamma$ by a constant, because it permits conditional heteroskedasticity that is common in economic applications. Moreover, condition ((ref)) only affects the far tails, and therefore allows covariates to affect extremal and central quantiles very differently. Even at the tails, the local effect of the covariates is approximately given by $\beta(\tau) \approx \beta_e + \gamma Q_U(\tau)$, which can be heterogenous across extremal quantiles.

Existence and Pareto-type behavior of the conditional quantile density function is also often imposed as a technical assumption that facilitates the derivation of inference results. Accordingly, we will assume that the conditional quantile density function $\partial Q_V(\tau | x) / \partial \tau$ exists and satisfies the tail equivalence relationship $$\partial Q_V(\tau | x) / \partial \tau \sim x' \gamma \cdot \partial Q_U(\tau) / \partial \tau \qquad \text{as} \qquad \tau \searrow 0$$ uniformly in $x \in \mathbf{X},$ where $\partial Q_U(\tau) / \partial \tau$ exhibits Pareto-type tails as $\tau \searrow 0$ with EV index $\xi+1$.

Estimation and Inference Methods

This section reviews estimation and inference methods for extremal quantile regression. Estimation is based on the koenker:1978 quantile regression estimator. We consider both analytical and resampling methods. These methods are introduced in the univariate case of marginal quantiles and then extended to the multivariate or regression case of conditional quantiles. We start by imposing some general sampling conditions.

Sampling Conditions

We assume that we have a sample of $(Y,X)$ of size $T$ that is either independent and identically distributed (i.i.d.) or stationary and weakly-dependent, with extreme events satisfying a non-clustering condition. In particular, the sequence $\{(Y_t, X_t) \}_{t=1}^T$ is assumed to form a stationary, strongly mixing process with geometric mixing rate, that satisfies the condition that curbs clustering of extreme events chernozhukov:2011. The assumption of mixing dependence is standard in econometrics white:2001. The non-clustering condition is of the meyer:1973 type and states that the probability of two extreme events co-occurring at nearby dates is much lower than the probability of just one extreme event. For example, it assumes that a large market crash is not likely to be immediately followed by another large crash. This assumption is convenient because it leads to limit distributions of extremal quantile regression estimators as if independent sampling had taken place. The plausibility of the non-clustering assumption is an empirical matter.

Univariave Case: Marginal Quantiles

The analog estimator of the marginal $\tau$-quantile is the sample $\tau$-quantile, $$ \hat{Q}_Y(\tau) = Y_{(\lfloor \tau T \rfloor)}, $$ where $Y_{(s)}$ is the $s^{th}$ order statistic of $(Y_1, \ldots, Y_T)$, and $\lfloor z \rfloor$ denotes the integer part of $z$. The sample $\tau$-quantile can also be computed as a solution to an optimization program,

equation*[equation* omitted — 132 chars of source]

where $\rho_\tau(u) := (\tau - 1\{u < 0\}) u$ is the asymmetric absolute deviation function of fox:rubin.

We review the asymptotic behavior of the sample quantiles under extreme and intermediate order sequences, and describe inference methods for extremal marginal quantiles.

Extreme Order Approximation

Recall the following classical result from EV theory on the limit distribution of extremal order statistics: as $T \to \infty$, for any integer $k\geq 1$ such that $k_T := \tau_T T \to k$,

equation[equation omitted — 143 chars of source]

where

equation[equation omitted — 107 chars of source]

The variable $U$ is defined in Section (ref) and $\{\mathcal{E}_1, \mathcal{E}_2, \cdots\}$ is an i.i.d.\ sequence of standard exponential variables. We call $\hat{Z}_T(k_T)$ the {\em canonically-normalized quantile (CN-Q) statistic} because it depends on the canonical scaling constant $A_T$. The result ((ref)) was obtained by gnedenko:1943 for i.i.d. sequences of random variables. It continues to hold for stationary weakly-dependent series, provided the probability of extreme events occurring in clusters is negligible relative to the probability of a single extreme event meyer:1973. leadbetter extended the result to more general time series processes.

The limit in ((ref)) provides an EV distribution as an approximation to the finite-sample distribution of $\hat Q_Y(\tau_T)$, given the scaling constant $A_T$. The limit distribution is characterized by the EV index $\xi$ and the variables $\Gamma_k$. The index $\xi$ is usually unknown but can be estimated by one of the methods described below. The variables $\Gamma_k$ are {\it gamma} random variables. The limit variable $\hat{Z}_\infty(k)$ is therefore a transformation of gamma variables, which has finite mean if $\xi < 1$ and has finite moments of up to order $1 / \xi$ if $\xi > 0$. Moreover, the limit EV distribution is not symmetric, predicting significant median asymptotic bias in $\hat Q_Y(\tau_T)$ with respect to $Q_Y(\tau_T)$ that motivates the use of the median-bias correction techniques discussed below.

The classical result ((ref)) is often not feasible for inference on $Q_Y(\tau_T)$ because the constant $A_T$ is unknown and generally cannot be estimated consistently bertail:2004. One way to deal with this problem is to add strong parametric assumptions on the non-parametric, slowly varying function $\tau \mapsto L(\tau)$ in Equation ((ref)) in order to estimate $A_T$ consistently. This approach is discussed in Section (ref). An alternative is to consider the {\em self-normalized quantile (SN-Q) statistic}

equation[equation omitted — 167 chars of source]

for $m> 1$ such that $mk$ is an integer. For example, $m = p/k_T + 1 = p/k + 1 + o(1)$ for some spacing parameter $p \geq 1$ (e.g.\ $p = 5$). The scaling factor $\mathcal{A}_T$ is completely a function of data and therefore feasible, avoiding the need for the consistent estimation of $A_T$. chernozhukov:2011 shows that as $T \to \infty$, for any integer $k\geq 1$ such that $k_T := \tau_T T \to k$,

equation[equation omitted — 160 chars of source]

The limit distribution in ((ref)) only depends on the EV index $\xi$, and its quantiles can be easily obtained by the resampling methods described in Section (ref).

Intermediate Order Approximation

Dekkers:dehaan show that as $\tau \searrow 0$ and $k_T \to \infty$, and under further regularity conditions,

equation[equation omitted — 178 chars of source]

where $\mathcal{A}_T$ is defined as in ((ref)). This result yields a normal approximation to the finite-sample distribution of $\hat{Q}_Y(\tau)$. Note this normal approximation holds only when $k_T \to \infty$, while the EV approximation ((ref)) holds not only when $k_T \to k$ but also when $k_T \to \infty$ because the EV distribution converges to the normal distribution as $k \to \infty$. In finite samples, we may interpret the condition $k_T \to \infty$ as requiring that $k_T \geq 30$.

Figure (ref), taken from chernozhukov:2011, shows that the EV distribution provides a better approximation to the distribution of extremal sample quantiles than the normal distribution when $k_T < 30$. It plots the quantiles of these distributions against the quantiles of the finite-sample distribution of the sample $\tau$-quantile for $T=200$ and $\tau \in \{ .025, .2, .3\}$. If either the EV or the normal distributions were to coincide with the exact distribution, then their quantiles would fall on the 45 degree line shown by the solid line. When the order $\tau T$ is $5$ or $40$, the quantiles of the EV distribution are very close to the 45 degree line, and in fact are much closer to this line than the quantiles of the normal distribution. Only for the case when the order $\tau T$ becomes $60$, do the quantiles of the EV and normal distributions become comparably close to the 45 degree line.

figure[figure omitted — 586 chars of source]

Estimation of $\xi$

Some inference methods for extremal marginal quantiles require a consistent estimator of the EV index $\xi$. We describe two well-known estimators of $\xi$. The first estimator, due to pickands:1975, relies on the ratio of sample quantile spacings:

equation[equation omitted — 207 chars of source]

Under further regularity conditions, for $\tilde{\tau}_T \searrow 0$ and $\tilde{\tau}_T T \to \infty$ as $T \to \infty$, \[ \sqrt{\tilde{\tau}_T T} (\hat{\xi}_P - \xi) \ \to_d \ \mathcal{N} \biggl( 0, \, \frac{\xi^2(2^{2\xi+1}+1)}{[2 (2^{\xi}-1) \ln 2]^2} \biggr) \] The second estimator, developed by hill, is the moment estimator:

equation[equation omitted — 206 chars of source]

which is applicable when $\xi>0$ and $\hat{Q}_Y(\tilde{\tau}_T) < 0$. This estimator is motivated by the maximum likelihood method that fits an exact power law to the tail data. Under further regularity conditions, for $\tilde{\tau}_T \searrow 0$ and $\tilde{\tau}_T T \to \infty$ as $T \to \infty$, \[ \sqrt{\tilde{\tau}_T T} (\hat{\xi}_H - \xi) \ \to_d \ \mathcal{N}(0, \xi^2). \] The previous limit results can be used to construct the confidence intervals and median-bias corrections for $\xi$. We give an example of these confidence intervals and corrections in Section (ref).

embrechts:1997 provide methods for choosing $\tilde{\tau}_T$. There is a variance-bias trade-off, as the variance of the estimator decreases, but the bias increases, as $\tilde{\tau}_T$ increases. Another view on the choice of $\tilde{\tau}_T$ is that the statistical models are approximations, but not literal descriptions of the data. In practice, the dependence of $\hat{\xi}$ on the threshold $\tilde{\tau}_T$ reflects that power laws with different values of $\xi$ would fit some tail regions better than others. Therefore, if the interest lies in making the inference on $Q_Y(\tau_T)$ for a particular $\tau_T$, it seems reasonable to use $\hat{\xi}$ constructed using $\tilde{\tau}_T = \tau_T$ or the closest $\tilde{\tau}_T$ to $\tau_T$ subject to $\tilde{\tau}_T T \geq 30$. This condition ensures to have a sufficient number of observations to estimate $\xi$.

Estimation of $A_T$

To use the CN-Q statistic for inference, we need to estimate the scaling constant $A_T$ defined in (ref). This requires additional strong restrictions on the underlying model. For instance, assume that the slowly varying function $\tau \mapsto L(\tau)$ is just a constant $L$, i.e., as $\tau \searrow 0$,

equation[equation omitted — 171 chars of source]

Then, we can estimate $L$ by

equation[equation omitted — 175 chars of source]

where $\hat{\xi}$ is either the Pickands or Hill estimator given in (ref) and (ref), and $\tilde{\tau}_T$ can be chosen using the same methods as in the estimation of $\xi$. Then, the estimator of $A_T$ is

equation[equation omitted — 90 chars of source]

Computing Quantiles of the Limit EV Distributions

The inference and bias corrections for extremal quantiles are based on the EV approximations given in (ref) and (ref), with an estimator in place of the EV index if needed. In practice, it is convenient to compute the quantiles of the EV distributions using simulation or resampling methods, instead of an analytical method. Here we illustrate two of such methods: extremal bootstrap and extremal subsampling. The bootstrap method relies on simulation, whereas the subsampling method relies on drawing subsamples from the original sample. Subsampling has the advantages that it does not require the estimation of $\xi$ and is consistent under general conditions (e.g.\ subsampling does not require i.i.d.\ data). Nevertheless, bootstrap is more accurate than subsampling when a stronger set of assumptions holds. It should be noted here that the empirical or nonparametric bootstrap is not consistent for extremal quantiles bickel.

The extremal bootstrap is based on simulating samples from a random variable with the same tail behavior as $Y$. Consider the random variable,\footnote{The variable $Y^*$ follows the generalized extreme value distribution, which nests the Frechet, Weibull, and Gumbell distributions. There are other possibilities, for example, the standard exponential distribution can be replaced by the standard uniform distribution, in which case $Y$ would follow the generalized Pareto distribution.}

equation[equation omitted — 132 chars of source]

This variable has the quantile function

equation[equation omitted — 90 chars of source]

which satisfies Condition ((ref)) because $Q_{Y^*}(\tau) - 1/\xi \sim \tau^{-\xi}/\xi$. The extremal bootstrap estimates the distribution of $Z_T(k_T) = \mathcal{A}_T (\hat{Q}_Y(\tau_T) - Q_Y(\tau_T))$ by the distribution of $Z^*_T(k_T) = \mathcal{A}_T (\hat{Q}_{Y^*}(\tau_T) - Q_{Y^*}(\tau_T))$ obtained by simulation. This approximation reproduces both the EV limit ((ref)) under extreme order sequences and the normal limit ((ref)) under intermediate order sequences. Algorithm (ref) describes the implementation of this method.

algorithm[algorithm omitted — 911 chars of source]

Extremal bootstrap can be also applied to estimate the distribution of other statistics, including the estimators of the EV index in (ref) or (ref) and the extrapolation estimators of Section (ref).

We now describe an extremal subsampling method to estimate the distributions of $Z_T(k_T)$ and $\hat{Z}_T(k_T)$ developed by chernozhukov:2011. It is based on drawing subsamples of size $b < T$ from $(Y_1, \ldots, Y_T)$ such that $b \to \infty$ and $b/T \to 0$ as $T \to \infty$, and computing the subsampling version of the SN-statistic $Z_T(k_T)$ as

equation[equation omitted — 211 chars of source]

where $\hat{Q}_Y^{b}(\tau)$ is the sample $\tau$-quantile in the subsample of size $b$, and $\tau_b := (\tau_T T)/b$. Similarly, the subsampling version of the CN-Q statistic $\hat{Z}_T(k_T)$ is

equation[equation omitted — 110 chars of source]

where $\hat{A}_b$ is a consistent estimator of $A_b$. For example, under the parametric restrictions specified in ((ref)), we can set $\hat{A}_b = \hat{L} b^{-\hat{\xi}}$ for $\hat{L}$ as in ((ref)) and $\hat{\xi}$ one of the estimators in (ref) or (ref).

The distributions of $Z^*_{b,T}(k_T)$ and $\hat Z^*_{b,T}(k_T)$ over all the subsamples estimate the distributions of $Z_T(k_T)$ and $\hat Z_T(k_T)$, respectively. The number of possible subsamples depends on the structure of serial dependence in the data, but it can be very large. In practice, the distributions over all subsamples are approximated by the distributions over a smaller number $S$ of randomly chosen subsamples, such that $S \to \infty$ as $T \to \infty$ politis:1999 . politis:1999 and bertail:2004 provide methods for choosing the subsample size $b$. Algorithm (ref) describes the implementation of the extremal subsampling.

algorithm[algorithm omitted — 1,101 chars of source]

Extremal subsampling differs from conventional subsampling, which is inconsistent for extremal quantiles. This difference can be more clearly appreciated in the case of $\hat{Z}_T(k_T)$. Here, conventional subsampling would recenter the subsampling version of the statistic by the estimator in the full sample $\hat{Q}_Y(\tau_T)$. Recentering in this way requires $A_b/A_T \to 0$ for consistency (see Theorem 2.2.1 in politis:1999), but $A_b/A_T \to \infty$ when $\xi > 0$. Thus, when $\xi > 0$ the extremal sample quantiles $\hat{Q}_Y(\tau_T)$ diverge rendering conventional subsampling to be inconsistent. In contrast, extremal subsampling uses $\hat{Q}_Y(\tau_b)$ for recentering. This sample quantile may diverge, but because it is an intermediate order quantile if $b/T \to 0$, the speed of its divergence is strictly slower than that of $A_T$. Hence, extremal subsampling exploits the special structure of the order statistics to do the recentering.

Median Bias Correction and Confidence Intervals

chernozhukov:2011 construct asymptotically median-unbiased estimators and $(1-\alpha)$-confidence intervals (CI) for $Q_Y(\tau_T)$ based on the SN-Q statistic as \[ \hat{Q}_Y(\tau) - \frac{\hat{c}_{1/2}}{\mathcal{A}_T} \qquad \text{and} \qquad \biggl[ \hat{Q}_Y(\tau) - \frac{\hat{c}_{1-\alpha/2}}{\mathcal{A}_T}, \ \hat{Q}_Y(\tau) - \frac{\hat{c}_{\alpha/2}}{\mathcal{A}_T} \biggr], \] where $\hat{c}_p$ is a consistent estimator of the $p$-quantile of $Z_T(k_T)$ that can be obtained using Algorithms (ref) or (ref). chernozhukov:2011 also construct asymptotically median-unbiased estimators and $(1-\alpha)$-CIs for $Q_Y(\tau_T)$ based on the CN-Q statistic as \[ \hat{Q}_Y(\tau) - \frac{\hat{c}'_{1/2}}{\hat{A}_T} \qquad \text{and} \qquad \biggl[ \hat{Q}_Y(\tau) - \frac{\hat{c}'_{1-\alpha/2}}{\hat{A}_T}, \ \hat{Q}_Y(\tau) - \frac{\hat{c}'_{\alpha/2}}{\hat{A}_T} \biggr], \] where $\hat{c}'_p$ is a consistent estimator of the $p$-quantile of $\hat Z_T(k_T)$ that can be obtained using Algorithm (ref), and $\hat{A}_T$ is a consistent estimator of $A_T$ such as (ref).

Extrapolation Estimator for Very Extremes

Sample $\tau$-quantiles can be very inaccurate estimators of marginal $\tau$-quantiles when $\tau T$ is very small, say $\tau T < 1$. For such very extremal cases we can construct more precise estimators using the assumptions on the behavior of the tails. In particular, we can estimate less extreme quantiles reliably, and extrapolate them to the quantile of interest using the tail assumptions.

Dekkers:dehaan develop the following extrapolation estimator:

equation[equation omitted — 221 chars of source]

where $\tau_T \ll \tilde \tau_T$ and $\hat \xi$ is a consistent estimator of $\xi$ such as (ref) or (ref). Then, for $\tilde \tau_T T \to \tilde k$ and $\tau_T T \to k$ with $\tilde{k} > k$, \[ \frac{\tilde{Q}_Y(\tau_T) - Q_Y(\tau_T)}{\hat{Q}_Y(\tilde \tau_T) - \hat{Q}_Y(2\tilde \tau_T)} \ \to_d \ \frac{(\tilde k/k)^\xi - 2^{-\xi}}{1-2^{-\xi}} + \frac{1-(\Gamma_{\tilde{k}}/k)^\xi}{e^{\xi \mathcal{E}_{\tilde{k}}}-1}, \] where $\mathcal{E}_{\tilde{k}}$ and $\Gamma_{\tilde{k}}$ are independent, $\Gamma_{\tilde{k}}$ has a standard gamma distribution with shape parameter $(2\tilde{k}+1)$, and $\mathcal{E}_{\tilde{k}} \sim \sum_{j=\tilde{k}+1}^{2\tilde{k}} Z_j/j$ with $Z_1, Z_2, \dots$ i.i.d. standard exponential. he:2016 proposed the closely related estimator

equation[equation omitted — 217 chars of source]

Under some regularity conditions, they show that for $\tau_T/\tilde \tau_T \to 0$ as $T \to \infty,$ this estimator converges to a normal distribution jointly with the EV index estimator $\hat \xi$.

The estimators in ((ref)) and ((ref)) have good properties provided that the quantities on the right-hand side are well estimated, which in turn requires that $\tilde \tau_T T$ be large, and that the Pareto-type tail model be a good approximation.

Multivariate Case: Conditional Quantiles

The $\tau$-quantile regression ($\tau$-QR) estimator of the conditional $\tau$-quantile $Q_Y(\tau | x) = x'\beta(\tau)$ is:

equation[equation omitted — 197 chars of source]

This estimator was introduced by laplace:1818 for the median case, and extended by koenker:1978 to include other quantiles and regressors.

In this section, we review the asymptotic behavior of the QR estimator under extreme and intermediate order sequences, and describe inference methods for extremal quantile regression. The analysis for the multivariate case parallels the analysis for the univariate case in Section (ref).

Extreme Order Approximation

Consider the {\em canonically-normalized quantile regression (CN-QR) statistic}

equation[equation omitted — 131 chars of source]

and the {\em self-normalized quantile regression (SN-QR) statistic}

equation[equation omitted — 208 chars of source]

where $\bar{X}_T := T^{-1} \sum_{t=1}^T X_t$ and $m$ is a real number such that $k (m-1) > d_x$ for $k_T = \tau_T T \to k$. For example, $m = (d_x+p)/k_T + 1 = (d_x+p)/k + 1 + o(1)$ where $p \geq 1$ is a spacing parameter (e.g.\ $p = 5$). The CN-QR statistic is generally infeasible for inference because it depends on the unknown canonical normalization constant $A_T$. This constant can only be estimated consistently under strong parametric assumptions, which will be discussed in Section (ref). The SN-QR statistic is always feasible because it uses a normalization that only depends on the data.

chernozhukov:2011 show that for $k_T \to k > 0$ as $T \to \infty$,

equation[equation omitted — 75 chars of source]

where for $\chi = 1$ if $\xi < 0$ and $\chi = -1$ if $\xi > 0$,

equation[equation omitted — 241 chars of source]

where $\{\mathcal{X}_1, \mathcal{X}_2, \dots\}$ is an i.i.d.\ sequence with distribution $F_X$; $\{\Gamma_1, \Gamma_2, \dots\} := \{\mathcal{E}_1, \mathcal{E}_1+\mathcal{E}_2, \dots\}$; $\{\mathcal{E}_1, \mathcal{E}_2, \dots\}$ is an i.i.d.\ sequence of standard exponential variables that is independent of $\{\mathcal{X}_1, \mathcal{X}_2, \dots\}$; and $\{y\}_+ := \max(0,y)$. Furthermore,

equation[equation omitted — 186 chars of source]

The limit EV distributions are more complicated than in the univariate case, but they share some common features. First, they depend crucially on the gamma variables $\Gamma_t$, are not necessarily centered at zero, and can have a significant first-order asymptotic median bias. Second, as mentioned above, the limit distribution of the CN-QR statistic in Equation ((ref)) is generally infeasible for inference due to the difficulty in consistently estimating the scaling constant $A_T$.

remark[Very Extreme Order Quantiles] feigin:1994, smith:1994, chernozhukov:1998, portnoy:jur, and knight:linear derived related results for canonically normalized linear programing or frontier regression estimators under very extreme order sequences where $\tau_T T \searrow 0$ as $T \to \infty$.

Intermediate Order Approximation

victor:annals shows that for $\tau_T \searrow 0$ and $k_T \to \infty$ as $T \to \infty$,

equation[equation omitted — 187 chars of source]

where $\mathcal{A}_T$ is defined as in ((ref)). As in the univariate case, this normal approximation provides a less accurate approximation to the distribution of the extremal quantile regression than the EV approximation when $k_T \not\to \infty$. The condition $k_T \to \infty$ can be interpreted in finite samples as requiring that $k_T/d_x \geq 30$, where $k_T/d_x$ is a dimension-adjusted order of the quantile explained in Section (ref).

Estimation of $\xi$ and $\gamma$

Some inference methods for extremal quantile regression require consistent estimators of the EV index $\xi$ and the scale parameter $\gamma$. The regression analog of the Pickands estimator is

equation[equation omitted — 250 chars of source]

This estimator is consistent if $\tilde{\tau}_T T \to \infty$ and $\tilde{\tau}_T \searrow 0$ as $T \to \infty$. Under additional regularity conditions, for $\tilde{\tau}_T \searrow 0$ and $\tilde{\tau}_T T \to \infty$ as $T \to \infty$,

equation[equation omitted — 190 chars of source]

The regression analog of the Hill estimator is

equation[equation omitted — 228 chars of source]

which is applicable when $\xi>0$ and $X_t\hat{\beta}(\tilde{\tau}_T)<0$. Under further regularity conditions, for $\tilde{\tau}_T \searrow 0$ and $\tilde{\tau}_T T \to \infty$,

equation[equation omitted — 119 chars of source]

These limit results can be used to construct confidence intervals for $\xi$. The scale parameter $\gamma$ can be estimated by

equation[equation omitted — 196 chars of source]

which is consistent if $\tilde{\tau}_T T \to \infty$ and $\tilde{\tau}_T \searrow 0$ as $T \to \infty$.

The choice of $\tilde{\tau}_T$ is similar to the univariate case in Section (ref). This time, however, one needs to take into account the multivariate nature of the problem. For example, if the interest lies in making the inference on $\beta(\tau_T)$ for a particular $\tau_T$, it is reasonable to set $\tilde{\tau}_T$ equal to the closest value to $\tau_T$ such that $\tilde{\tau}_T T/d_x \geq 30$. We refer again the reader to Section (ref) for a discussion on the difference in the choice of $\tilde{\tau}_T$ between the univariate and multivariate cases.

Estimation of $A_T$

To use the CN-QR statistic for inference, we need to estimate the scaling constant $A_T$ defined in (ref). This requires strong restrictions and an additional estimation procedure. For example, assume that the non-parametric slowly varying component $L(\tau)$ of $A_T$ is replaced by a constant $L$, i.e., as $\tau \searrow 0$

equation[equation omitted — 171 chars of source]

Then we can estimate the constant $L$ by

equation[equation omitted — 191 chars of source]

where $\hat{\xi}$ is either the Pickands or Hill estimator given in (ref) or (ref). Thus, the scaling constant $A_T$ is estimated by $$\hat{A}_T := \hat{L} T^{-\hat{\xi}}.$$

Computing Quantiles of the Limit EV Distributions

We consider inference and asymptotically median unbiased estimation for linear functions of the coefficient vector $\beta(\tau)$, $\psi' \beta(\tau),$ for some nonzero vector $\psi \in \mathbb{R}^{d_x}$, based on the EV approximations $\psi'\hat Z_{\infty}(k)$ from (ref) and $\psi' Z_{\infty}(k)$ from (ref). We describe three methods to compute critical values of the limit EV distributions: analytical computation, extremal bootstrap, and extremal subsampling. The analytical and bootstrap methods require estimation of the EV index $\xi$ and the scale parameter $\gamma$. Subsampling applies under more general conditions than the other methods, and hence we would recommend the use of it. However, the analytical and bootstrap methods can be more accurate than subsampling if the data satisfy a stronger set of assumptions.

The analytical computation method is based directly on the limit distributions (ref) and (ref) replacing $\xi$ and $\gamma$ by consistent estimators. Define the $d_x$-dimensional random vector:

equation[equation omitted — 281 chars of source]

where $\hat{\chi} = 1$ if $\hat{\xi} < 0$ and $\hat{\chi} = -1$ if $\hat{\xi} > 0$, $\hat{\xi}$ is an estimator of $\xi$ such as (ref) or (ref), $\hat{\gamma}$ is an estimator of $\gamma$ such as (ref), $\{\Gamma_1, \Gamma_2, \dots\} = \{\mathcal{E}_1, \mathcal{E}_1 + \mathcal{E}_2, \dots\}$, $\{\mathcal{E}_1, \mathcal{E}_2, \dots\}$ is an i.i.d.\ sequence of standard exponential variables, and $\{\mathcal{X}_1, \mathcal{X}_2, \dots\}$ is an i.i.d.\ sequence independent of $\{\mathcal{E}_1, \mathcal{E}_2, \dots\}$ with distribution function $\hat{F}_X$, where $\hat{F}_X$ is any smooth consistent estimator of $F_X$, e.g., a smoothed empirical distribution function of the sample $(X_1, \ldots, X_T)$. Also, let \[ Z^\ast_\infty(k) = \frac{\sqrt{k} \hat{Z}^\ast_\infty(k)}{\bar{X}_T'(\hat{Z}^\ast_\infty(mk)-\hat{Z}^\ast_\infty(k))+\hat{\chi}(m^{-\hat{\xi}}-1)k^{-\hat{\xi}}}. \] The quantiles of $\psi'\hat Z_{\infty}(k)$ and $\psi' Z_{\infty}(k)$ are estimated by the corresponding quantiles of the $\psi' \hat{Z}^\ast_\infty(k)$ and $\psi' Z^\ast_\infty(k)$, respectively. In practice, these quantiles can only be evaluated numerically via the following algorithm.

algorithm[algorithm omitted — 856 chars of source]

The extremal bootstrap is computationally less demanding than the analytical methods. It is based on simulating samples from a random variable with the same tail behavior as $(Y_1, \ldots, Y_T)$. Consider the bootstrap sample $\{(Y_1^*, X_1), \dots, (Y_T,^* X_T)\}$, where

equation[equation omitted — 171 chars of source]

and $\{X_1, \dots, X_T\}$ is a fixed set of observed regressors from the data. The variable $Y_t^*$ has the conditional quantile function

equation[equation omitted — 142 chars of source]

The extremal bootstrap approximates the distribution of $Z_T(k_T) = \mathcal{A}_T (\hat{\beta}(\tau)-\beta(\tau))$ by the distribution of $Z^*_T(k_T) = \mathcal{A}_T (\hat{\beta}^*(\tau)- \beta^*(\tau))$ where $\hat{\beta}^*(\tau)$ is the $\tau$-QR estimator in the bootstrap sample. This approximation reproduces both the EV limit (ref) under extreme value sequences, and the normal limit (ref) under intermediate order sequences. The distribution of $Z^*_T(k_T)$ can be obtained by simulation using the algorithm:

algorithm[algorithm omitted — 1,020 chars of source]

chernozhukov:2011 developed an extremal subsampling method to estimate the distributions of $\hat Z_T(k_T)$ and $Z_T(k_T)$. It is based on drawing subsamples of size $b < T$ from $\{(X_t,Y_t)\}_{t=1}^T$ such that $b \to \infty$ and $b/T \to 0$ as $T \to \infty$, and computing the subsampling version of the SN-QR statistic as

equation[equation omitted — 234 chars of source]

where $m=(d_x + p)/(\tau_T T)$ for some {\em spacing parameter} $p \geq 1$ (e.g.\ $p=5$), $\hat{\beta}_b(\tau)$ is the $\tau$-QR estimator in the subsample of size $b$, $\bar{X}_{b,T}$ is the sample mean of the regressors in the subsample, and $\tau_b := (\tau_T T)/b$.\footnote{In practice, it is reasonable to use the following finite-sample adjustment to $\tau_b$: $\tau_b = \min\{(\tau_T T)/b, 0.2\}$ if $\tau_T < 0.2$, and $\tau_b = \tau_T$ if $\tau_T \geq 0.2$. The idea is that $\tau_T$ is adjusted to be non-extremal if $\tau_T > 0.2$, and the subsampling procedure reverts to central order inference. The truncation of $\tau_b$ by $0.2$ is a finite-sample adjustment that restricts the key statistics $Z^*_{b,T}(k_T)$ to be extremal in subsamples. These finite-sample adjustments do not affect the asymptotic arguments.} Similarly, the subsampling version of the CN-QR statistic $\hat{Z}_T(k_T)$ is

equation[equation omitted — 113 chars of source]

where $\hat{A}_b$ is a consistent estimator for $A_b$. For example, $\hat{A}_b = \hat{L} b^{-\hat{\xi}},$ for $\hat{L}$ given by ((ref)) and $\hat{\xi}$ is the estimator of $\xi$ given in (ref) or (ref).

As in the univariate case, the distributions of $Z^*_{b,T}(k_T)$ and $\hat Z^*_{b,T}(k_T)$ over all the possible subsamples estimate the distributions of $Z_T(k_T)$ and $\hat Z_T(k_T)$, respectively. These distributions can be obtained by simulation using the algorithm:

algorithm[algorithm omitted — 997 chars of source]

The comments of Section (ref) on the choice of subsample size, number of simulations, and differences with conventional subsampling also apply to the regression case.

Median Bias Correction and Confidence Intervals

chernozhukov:2011 construct asymptotically median-unbiased estimators and $(1-\alpha)$-CIs for $\psi'\beta(\tau)$ based on the SN-QR statistic as \[ \psi' \hat{\beta}(\tau) - \frac{\hat{c}_{1/2}}{\mathcal{A}_T} \qquad \text{and} \qquad \biggl[ \psi' \hat{\beta}(\tau) - \frac{\hat{c}_{1-\alpha/2}}{\mathcal{A}_T}, \ \psi' \hat{\beta}(\tau) - \frac{\hat{c}_{\alpha/2}}{\mathcal{A}_T} \biggr], \] where $\hat{c}_p$ is a consistent estimator of the $p$-quantile $c_\alpha$ of $Z_T(k_T)$ that can be obtained using Algorithms (ref), (ref), or (ref). chernozhukov:2011 also construct asymptotically median-unbiased estimators and $(1-\alpha)$-CIs for $\psi'\beta(\tau)$ based on the CN-QR statistic as \[ \psi' \hat{\beta}(\tau) - \frac{\hat{c}_{1/2}'}{A_T} \qquad \text{and} \qquad \biggl[ \psi' \hat{\beta}(\tau) - \frac{\hat{c}_{1-\alpha/2}'}{A_T}, \ \psi' \hat{\beta}(\tau) - \frac{\hat{c}_{\alpha/2}'}{A_T} \biggr], \] where $\hat{c}_p'$ is a consistent estimator of the $p$-quantile of $\hat{Z}_T(k_T)$ that can be obtained using Algorithms (ref) or (ref).

Extrapolation Estimator for Very Extremes

The $\tau$-QR estimators can be very inaccurate when $\tau T/d_x$ is very small, say $\tau T/d_x < 1$. We can construct extrapolation estimators for these cases that use the assumptions on the behavior of the tails. By analogy with the univariate case,

equation[equation omitted — 234 chars of source]

or

equation[equation omitted — 232 chars of source]

where $\tau_T \ll \tilde \tau_T$, and $\hat \xi$ is the Pickands or Hill estimator of $\xi$ in (ref) or (ref). he:2016 derived the joint asymptotic distribution of $(\breve{\beta}(\tau_T),\hat{\xi}_P)$. wang:2012 developed other extrapolation estimators for heavy-tailed distributions with $\xi > 0$.

The estimators in ((ref)) and ((ref)) have good properties provided that the quantities on the right-hand side are well estimated, which in turn requires that $\tilde \tau_T T/d_x$ be large, and that the Pareto-type tail model be a good approximation. To construct the confidence interval for $\beta(\tau_T)$ based on extrapolation, we can apply the extremal subsampling to the statistic $$\widetilde{\mathcal{A}}_T [\tilde{\beta}(\tau_T) - \beta(\tau_T)], \quad \widetilde{\mathcal{A}}_T = \frac{\sqrt{\tilde \tau_T T}}{\bar{X}_T'(\hat{\beta}(m\tilde \tau_T)-\hat{\beta}(\tilde \tau_T))}.$$ For the estimator ((ref)), we can also use analytical methods based on the asymptotic distribution given in Corollary 3.4 of he:2016.

EV Versus Normal Inference

chernozhukov:2011 provided a simple rule of thumb for the application of EV inference. Recall that the order of a sample $\tau_T$-quantile from a sample of size $T$ is $\tau_T T$ (rounded to the next integer). This order plays a crucial role in determining the quality of the EV or normal approximations. Indeed, the former requires $\tau_T T \to k$, whereas the latter requires $\tau_T T \to \infty$. In the regression case, in addition to the order of the quantile, we need to take into account $d_x$, the dimension of $X$. As an example, consider the case where all $d_x$ covariates are indicators that divide equally the sample into subsamples of size $T/d_x$. Then, each of the components of the $\tau_T$-QR estimator will correspond to a sample quantile of order $\tau_T T/d_x$. We may therefore think of $\tau_T T/d_x$ as a dimension-adjusted order for quantile regression.

A common simple rule for the application of the normal is that the sample size is greater than 30. This suggests that we should use extremal inference whenever $\tau_T T/d_x \lesssim 30$. This simple rule may or may not be conservative. For example, when regressors are continuous, the computational experiments in chernozhukov:2011 show that the normal inference performs as well as the EV inference provided that $\tau_T T/d_x \gtrsim 15$ to $20$, which suggests using EV inference when $\tau_T T/d_x \lesssim 15$ to $20$ for this case. On the other hand, if we have an indicator in $X$ equal to one only for 2% of the sample, then the coefficient of this indicator behaves as a sample quantile of order $.02 \tau_T T = \tau_T T/ 50$, which would motivate using EV inference when $\tau_T T/50 \lesssim 15$ to $20$ in this case. This rule is far more conservative than the original simple rule when $d_x \ll 50$. Overall, it seems prudent to use both EV and normal inference methods in most cases, with the idea that the discrepancies between the two can indicate extreme situations.

Empirical Applications

We consider two applications of extremal quantile regression to conditional value-at-risk and financial contagion. We implement the empirical analysis in \verb"R" language with Koenker (koenker:2016) \verb"quantreg" package and the code from chernozhukov-du:2008 and chernozhukov:2011. The data are obtained from Yahoo! Finance.\footnote{The dataset and the code are available online at Fern\'andez-Val's website: http://sites.bu.edu/ivanf/research/.}

Value-at-Risk Prediction

We revisit the problem of forecasting the conditional value-at-risk of a financial institution posed by vl with more recent methodology. The response variable $Y_t$ is the daily return of the Citigroup stock, and the covariates $X_{1t}$, $X_{2t}$, and $X_{3t}$ are the lagged daily returns of the Citigroup stock (C), the Dow Jones Industrial Index (DJI), and the Dow Jones US Financial Index (DJUSFN), respectively. The lagged own return captures dynamics, DJI is a measure of overall market return, and the DJUSFN is a measure of market return in the financial sector. We estimate quantiles of $Y_t$ conditional on $X_t = (1, X_{1t}^+, X_{1t}^-, X_{2t}^+, X_{2t}^-, X_{3t}^+, X_{3t}^-)$ with $x^+ = \max\{x, 0\}$ and $x^- = -\min\{x, 0\}$. There are 1,738 daily observations in the sample covering the period from January 1, 2009 to November 30, 2015.

Figure (ref) plots the QR estimates $\hat{\beta}(\tau)$ along with 90% pointwise CIs. The solid lines represent the extremal CIs and the dashed lines the normal CIs. The extremal CIs are computed by the extremal subsampling method described in Algorithm (ref) with the subsample size $b = \lfloor 50 + \sqrt{1,738} \rfloor = 91$ and the number of simulations $S = 500$. We use the SN-QR statistic with spacing parameter $p=5$. The normal CIs are based on the normal approximation with the standard errors computed with the method proposed by powell:1991.\footnote{We used the command summary.rq with the option ker in the quantreg package to compute the standard errors.} Figure (ref) plots the median bias-corrected QR estimates along with 90% pointwise CIs for the lower tail (note that due to the median bias-correction, the coefficient estimates are slightly different from Figure (ref)). The bias correction is also implemented using extremal subsampling with the same specifications.

figure[figure omitted — 323 chars of source]
figure[figure omitted — 356 chars of source]

We focus the discussion on the impact of downward movements of the explanatory variables (the C lag $X_{1t}^-$, the DJI lag $X_{2t}^-$, and the DJUSFN lag $X_{3t}^-$) on the extreme risk, that is, on the low conditional quantiles of the Citigroup stock return. To interpret the results, it is helpful to keep in mind that if the covariates were completely irrelevant (i.e. independent from the response), then their coefficients would be equal to zero uniformly over $\tau$, except for the constant term. The intercept would coincide with the unconditional quantile of Citigroup daily return. Another general remark is that we would expect the estimates and CIs to be more volatile at the tails than at the center due to data sparsity. Figures (ref) and (ref) show that most of the coefficients are insignificant throughout the distribution, what confirms the expected unpredictability of the stock returns. However, we do find that the coefficient on the Citigroup's lagged return $X_{1t}^-$ is significantly different from zero in the extreme low quantiles (see the upper right figure in Figure (ref)). This suggests that from 2009 to 2015, a past drop in the stock price of Citigroup has significantly pushed down the extreme low quantiles of the current stock price. Informally speaking, the negative return on the stock price induced the risk of a further negative outcome in the near future.

Comparing the CIs produced by the extremal inference and the normal inference, Figure (ref) shows that they closely match in the central region, while Figures (ref) reveals that the normal CIs are often narrower than the extremal CIs in the tails, especially for $\tau < 0.05$. As briefly mentioned in Section (ref), the extremal CIs coincide with the normal CIs when the situation is non-extremal. Therefore, this discrepancy indicates that the normal CIs on the tails substantially underestimates the sampling variation and hence it might lead to a substantial undercoverages in the CIs.

We next characterize the tail properties of the model. Table (ref) reports the estimates of the EV index $\xi$ obtained by the Hill estimator in ((ref)), together with bias corrected estimates and 90% CIs based on ((ref)), which were obtained using the QR extremal bootstrap of Algorithm (ref) with $S=500$ applied to the Hill estimator. The bias-corrected estimates of $\xi$ are relatively stable even at the extreme tails. They are greater than zero, confirming that the distribution of stock returns has a much thicker lower tail than the normal distribution. It is noteworthy that none of these estimates were used to produce the fig. (ref) and (ref) because they were obtained from extremal subsampling method applied to the SN-QR statistic.

table[table omitted — 420 chars of source]

Having characterized the EV index, we can now estimate the very extreme quantiles using extrapolation methods. We set $\hat{\xi}$ to be the estimate with $\tau = 0.05$, and compute the extrapolation estimator ((ref)) for $\tau = 0.005$, $0.001$, and $0.0001$ in Table (ref). For comparison purposes, the first column reports the $\tau$-QR estimates for $\tau = 0.005$ obtained from ((ref)). This estimator cannot be calculated for the other quantile indexes considered. We find some discrepancies between the two estimators especially for the coefficients of the negative lags at $\tau = 0.005$. Figure (ref) plots the predicted values for the conditional $0.005$-quantiles in the second half of 2015 obtained from the QR and extrapolation estimators. The standard QR fit uses sample data that contains few observations on the extreme events, while the extrapolated fit uses the tail model and a reliably estimated conditional $0.05$-quantile coefficients to predict the magnitude of such events. The quality of this prediction clearly depends on whether the tails model is accurate.

table[table omitted — 1,065 chars of source]
figure[figure omitted — 237 chars of source]

Application 2: Contagion of Financial Risk

We consider an application to contagion of financial risk between commercial banks. The response variable $Y_t$ is the daily return of the Citigroup stock (C), and the covariates $X_{1t}$, $X_{2t}$, and $X_{3t}$ are the contemporaneous daily returns of the stocks of other banks, namely, Bank of America (BAC), JPMorgan Chase & Co.\ (JPM), and Wells Fargo & Co.\ (WFC). As in the previous section, we estimate the quantiles of $Y_t$ conditional on $X_t = (1, X_{1t}^+, X_{1t}^-, X_{2t}^+, X_{2t}^-, X_{3t}^+, X_{3t}^-)$ using 1,738 daily observations covering the period from January 1, 2009 to November 30, 2015.

Figure (ref) plots the QR estimates $\hat \beta(\tau)$ along with 90% pointwise CIs. The solid lines represent the extremal CIs and the dashed lines the normal CIs. The extremal CIs are computed by the extremal subsampling method described in Algorithm (ref) with the subsample size $b = \lfloor 50 + \sqrt{1,738} \rfloor = 91$ and the number of simulations $S = 500$. We use the SN-QR statistic with spacing parameter $p=5$. The normal CIs are based on the normal approximation with the standard errors computed with the method proposed by powell:1991. Figure (ref) plots the median bias-corrected QR estimates along with 90% pointwise CIs for the lower tail. The bias correction is also implemented using extremal subsampling with the same specifications.

We find a significant effect of Bank of America's risk on Citigroup's risk. Observe that the coefficient of BAC ($+$) is positive and that of BAC ($-$) is negative across most of the quantiles. This tells that BAC and C hold similar portfolios and that there might be a direct contagion of BAC's risk to C's risk (negative return of BAC is likely to cause negative return of C). Similar observation holds for JPM's risk onto C's risk. However, there are no such contagion effect of WFC's risk onto C's. In fig. (ref) we see that the negative return of Bank of America's stock has a large effect on the extreme low quantile of C's return, while its positive return has no significant effect. This indicates that Bank of America's risk has an asymmetric and large impact on its competitor. As in the value at risk application, we find that the normal and extremal CIs are similar in the central region, while the normal CIs are narrower than the extremal CI in the tails, especially for $\tau < 0.05$.

figure[figure omitted — 329 chars of source]
figure[figure omitted — 362 chars of source]

Table (ref) reports the estimates of the EV index $\xi$ obtained by the Hill estimator ((ref)), together with bias corrected estimates and 90% CIs based on ((ref)), which were obtained using the QR extremal bootstrap of Algorithm (ref) with $S=500$ applied to the Hill estimator. Again we find estimates significantly greater than zero confirming that stock returns have thick lower tails relative to the normal distribution. Table (ref) shows the estimates of the QR coefficients for very low quantiles obtained from QR and the extrapolation estimator ((ref)) with $\hat{\xi} = 0.263,$ the estimate from table (ref) for $\tau = 0.05$. The largest difference between the regression and extrapolation estimates occur for the WFC's return. Here we find a large negative coefficient for the return (-) that indicates that there might be contagion of financial risk from WFC to C at very low quantiles. Figure (ref) contrasts the predicted values for the conditional $0.005$-quantiles in the second half of 2015 obtained from the QR and extrapolation estimators. Overall, the two methods produce similar estimates, although the extrapolated estimator predicts deeper troughs in the quantiles.

table[table omitted — 426 chars of source]
table[table omitted — 1,027 chars of source]
figure[figure omitted — 243 chars of source]