The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
67,461 characters
Estimation and Inference about Tail Features with Tail Censored Data
\title{
\vspace{-6ex}
Estimation and Inference about Tail Features with Tail Censored Data\thanks{
We thank Ulrich M\"{u}ller for very inspiring discussions. We also thank Tim
Christensen, Bo Honor\'{e}, Jia Li, Andrew Patton, Alexis Toda, Mark Watson,
and participants at numerous seminar/conference presentations for very
helpful comments.}}
\author{\textsc{Yulong Wang}\thanks{\textit{Address}: Department of
Economics, Syracuse University, 127 Eggers Hall, Syracuse, NY 13244. \textit{
E-mail}: \texttt{[email removed]}} \\
Syracuse University \and \textsc{Zhijie Xiao}\thanks{\textit{Address}:
Department of Economics, Boston College, 380 Maloney Hall, Chestnut Hill, MA
02467. \textit{E-mail}: \texttt{[email removed]}} \\
Boston College}
\date{First version: May 2019\\
This version: February 2020}
\maketitle
\begin{abstract}
This paper considers estimation and inference about tail features when the
observations beyond some threshold are censored. We first show that ignoring
such tail censoring could lead to substantial bias and size distortion, even
if the censored probability is tiny. Second, we propose a new maximum
likelihood estimator (MLE) based on the Pareto tail approximation and derive
its asymptotic properties. Third, we provide a small sample modification to
the MLE by resorting to Extreme Value theory. The MLE with this modification
delivers excellent small sample performance, as shown by Monte Carlo
simulations. We illustrate its empirical relevance by estimating (i) the
tail index and the extreme quantiles of the US individual earnings with the
Current Population Survey dataset and (ii) the tail index of the
distribution of macroeconomic disasters and the coefficient of risk aversion
using the dataset collected by \citeasnoun{Barro08}. Our new empirical findings
are substantially different from the existing literature.
\addtocontents{toc}{nprotectnenlargethispage*{1000pt}} \flushleft\textbf{
Keywords:} Extreme Value theory, power law, extreme quantile, tail index
\addtocontents{toc}{nprotectnpagebreak}
\end{abstract}
\setcounter{page}{1} \small \normalsize
\section{Introduction}
Tail risk and extreme events are important research topics in economics and
finance. In many applications, the features of interest are tail properties
such as tail index and extreme quantiles. Existing literature has
extensively studied the case with fully observed datasets. In comparison,
this article explores the case with tail censoring. We argue that it is
important to take into account the censoring if the research interest is in
the tail, even when the censoring fraction is small. We provide a new method
to construct estimators and confidence intervals for tail features.
Suppose one has a random sample from some underlying distribution $F$, where
the observations larger than some threshold $T$ are replaced with $T$ or
simply unobserved. In principle, tail features cannot be even identified if
they entirely depend on the right tail part of $F$ that is beyond $T$.
However, we can back out the tail-related features by extrapolation under
two assumptions. They are that (i) the tail of $F$ can be well approximated
by some suitably chosen parametric distribution, and (ii) $T$ is
sufficiently large so that only a small fraction of samples are censored.
The first assumption has been thoroughly studied in the statistic literature
and is satisfied by many commonly used distributions. The second assumption
is also satisfied in many interesting empirical applications, which motivate
this article.
Our first motivating example is the Current Population Survey (CPS) dataset,
which has been the primary data source used for investigating the
distributions of individual earnings and household income in the US.
Featured studies using CPS data include \citeasnoun{Armour13} and \citeasnoun{Eika19},
among many others. In CPS, the individual earnings larger than some
threshold $T$ are typically censored (also called topcoded) and replaced
with $T$ for confidential reasons.\footnote{
The topcoding has constantly been changing. Description of the topcoding
mechanism is available at https://cps.ipums.org/cps/topcodes\_tables.} In
2019, the censoring threshold is 310000 USD, leading to an approximately
0.58\% censoring fraction in the full sample of individuals between 18 and
70 years old. This quantity is also substantially different across the
subsamples defined by race and gender but remains small, as seen in Table
\ref{tbl intro}. Using this dataset, we aim to estimate and construct
confidence intervals for the tail index that measures the tail heaviness of
the income distribution and the extreme quantiles.
\begin{table}[H]
\begin{center}
\caption{Fractions and Numbers of Censored Observations in the 2019 CPS
Dataset}\label{tbl intro}
\vspace{+2ex}
\begin{tabular}{ccccccc}
\hline\hline
& $n$ & {\small cen\%} & {\small cen\#} & $n$ & {\small cen\%} & {\small
cen\#} \\ \hline
\multicolumn{1}{l}{\small full sample} & \multicolumn{1}{r}{\small 115424} &
{\small 0.582} & {\small 672} & & & \\
{\small race\TEXTsymbol{\backslash}gender} & \multicolumn{3}{c}{\small male}
& \multicolumn{3}{c}{\small female} \\ \hline
\multicolumn{1}{l}{\small all} & \multicolumn{1}{r}{\small 55553} & {\small
0.884} & {\small 491} & \multicolumn{1}{r}{\small 59871} & {\small 0.302} &
{\small 181} \\
\multicolumn{1}{l}{\small white} & \multicolumn{1}{r}{\small 43371} &
{\small 0.966} & {\small 419} & \multicolumn{1}{r}{\small 45424} & {\small
0.310} & {\small 141} \\
\multicolumn{1}{l}{\small Asian} & \multicolumn{1}{r}{\small 3676} & {\small
1.360} & {\small 50} & \multicolumn{1}{r}{\small 4099} & {\small 0.537} &
{\small 22} \\
\multicolumn{1}{l}{\small Hispanic} & \multicolumn{1}{r}{\small 44420} &
{\small 1.002} & {\small 445} & \multicolumn{1}{r}{\small 48192} & {\small
0.322} & {\small 155} \\
\multicolumn{1}{l}{\small black} & \multicolumn{1}{r}{\small 6144} & {\small
0.195} & {\small 12} & \multicolumn{1}{r}{\small 7827} & {\small 0.115} &
{\small 9} \\ \hline
\end{tabular}
\vspace{-4ex}
\end{center}
\begin{singlespacing}
\begin{footnotesize}
Note: Entries are the sample sizes ($n$), the fractions in percentage
(cen\%) and the numbers (cen\#) of censored observations in individual
earnings from the March CPS variable ERN\_VAL. Data are available at
https://usa.ipums.org/usa/.
\end{footnotesize}
\end{singlespacing}
\end{table}
In our second application, we examine the size distribution of macroeconomic
disasters, which is investigated by \citeasnoun{Barro08} and \citeasnoun{Barro11}. In
particular, \citeasnoun{Barro11} define a macroeconomic disaster if the annual
Gross Domestic Product (GDP) (or consumption) declines by more than 10\%.
The authors collect the data in 36 countries from 1870 to 2005 and construct
a sample of approximately 5000 observations. However, data are missing for
four countries during WWII due to government collapse or fighting wars.
These observations correspond to the end-of-world case, and hence \citeasnoun
{Barro11} concern that they are the largest observations but censored. In
this situation, the tail censoring fraction is about 0.1\%, and the
parameter of interest is the tail index and the coefficient of risk
aversion. Note that the censoring threshold $T$ is unknown here.
One would think that a tiny censoring fraction makes nearly no effect if we
ignore it. This is true if the object of interest lies in the mid-sample,
such as the median. However, such ignorance could lead to substantial bias
and size distortion if some tail features are of interest. To examine this,
we conduct an extensive Monte Carlo study and find that even the 0.1\% tail
censoring could lead to poor finite sample performance in some commonly used
methods, including, for example, the classical Hill (1975)'s estimator.
\nocite{Hill75}
To accommodate the censoring, existing studies typically rely on some
parametric assumption of the whole distribution (e.g., \citeasnoun{Aban06}, \citeasnoun
{Jenkins10}, and \citeasnoun{Burkhauser10}). Then, tail features can be expressed
as functions of the unknown parameters and estimated by the maximum
likelihood estimator (MLE). However, the parametric assumption on the whole
distribution may lead to a substantial misspecification error when the
object of interest is in the tail. This is because tail features such as
very large quantiles are typically on a large scale, and hence small
misspecification can be considerably amplified. For example, the standard
normal distribution and the Student-t distribution with 20 degrees of
freedom share almost the same shape in the mid-sample but exhibit
substantially different extreme quantiles. Such misspecification is
documented by \citeasnoun{Brzezinski2013} in a large-scale simulation study.
Instead of modeling the whole underlying distribution $F$, we focus on the
\textit{tail} part only and approximate it with the generalized Pareto
distribution (GPD). Such a Pareto-tail approximation holds for many commonly
used distributions, including, for example, Student-t, F, Beta, and Gaussian
distributions. See Chapter 1 of \citeasnoun{deHaan07} for an overview. Given this
approximation, we first pick some tail cutoff $u$, such as some large
empirical quantile, and treat the observations larger than $u$ (but still
less than $T$) as stemming from the censored Pareto tail. Let $k$ denote the
number of these effective \textit{tail} observations. Then, we can fit them
into the classical Tobit model and conduct the MLE of the unknown
parameters. Under some mild regularity conditions, we show that the MLE is
consistent and asymptotically normal as $k$ diverges, enabling the
construction of confidence intervals. Then we can estimate extreme quantiles
by expressing them as functions of the Pareto parameters. This is formally
studied in Section \ref{sec inck}.
The proposed maximum likelihood method can be quite good in finite samples
for some applications, but it is also easy to find examples where the
asymptotic distributions provide poor approximations. A fundamental
limitation of our MLE and many other existing approaches in studying tail
features is that they require restrictive conditions for the choice of $k$
(and equivalently $u$). On the one hand, $k$ has to be sufficiently large to
support enough observations stemming from the approximately Pareto tail for
the consistency and the asymptotic Gaussianity. On the other hand, $k$ has
to be sufficiently small relative to the whole sample size $n$ so that the
tail Pareto approximation incurs a negligible bias. Such a delicate balance
is technically reflected in the conditions that $k\rightarrow \infty $ and $
k/n\rightarrow 0$. As such, for some combinations of $n$ and $F$, it is hard
to find the $k$ that leads to satisfactory inference. This situation is
close in spirit to the bias-variance trade-off in choosing the bandwidth in
the standard kernel regressions.
To alleviate the above issue, we propose a small sample modification of the
MLE when $k$ is only moderate, say 100. In particular, we follow \citeasnoun
{MuellerWang17} to consider $k$ as a fixed number and study the fixed-$k$
asymptotic embedding by resorting to Extreme Value (EV) theory. Instead of
treating the tail observations as \textit{independent} copies from the GPD,
EV theory treats them as \textit{dependent} random variables with a joint EV
distribution. Such dependence is negligible when $k$ is large but plays an
important role when $k$ is only moderate. Using the EV approximation, we
propose new confidence intervals for the tail index and extreme quantiles,
which have excellent coverage probabilities. This is studied in detail in
Section \ref{sec fk}.
In summary, the method that we propose is a hybrid approach. We suggest
using the MLE for estimation and inference when $k$ (and $n$) is large
enough and switching to the fixed-$k$ intervals otherwise. We choose the
switching cutoff to be $k\gtrless 250$ based on our Monte Carlo experiments
in Section \ref{sec mc}.
Returning to the CPS application, $n$ is large enough in the full sample to
support a large $k$, and hence the MLE is expected to perform well. However,
in the Asian male subsample, $n$ is only 3676. Then choosing $k$ as a small
fraction, say 5\%, of $n$, leads to only 180 tail observations and triggers
the switching. By using the new approach, we make several interesting
empirical findings. First, the tail index is substantially different across
genders and races, while the existing literature commonly focuses on the
full sample and finds the tail index to be approximately 0.5. See \citeasnoun
{TodaWang2019} and references therein. Second, extreme quantiles also
considerably vary across genders and races. In particular, the 99.9\%
quantile of all males can be twice larger than that of the black male group.
Third, the tail features are also substantially different across ages. The
middle-aged groups exhibit heavier tails and larger extreme quantiles than
the groups below 30 or above 60 years old.
In the macroeconomic disaster application, we use the new method to
construct confidence intervals for the tail index and the coefficient of
risk aversion. We find substantially different results from those in \citeasnoun
{Barro11}. In particular, we obtain a significantly heavier tail in the
disaster distribution. This further results in a smaller coefficient of
relative risk aversion, approximately 0.75 instead of 3. Our Monte Carlo
simulation statistically justifies such a vast difference.
The rest of the paper is organized as follows. Section \ref{sec inck}
develops the MLE, and Section \ref{sec fk} provides the small sample
modification. Section \ref{sec mc} reports Monte Carlo simulations, and
Section \ref{sec emp} applies the new approach to the CPS and the
macroeconomic disaster examples. Section \ref{sec conclusion} concludes with
some remarks. All proofs and computational details are collected in the
Appendix.
\section{The Maximum Likelihood Estimator\label{sec inck}}
Consider a random sample $\{Y_{i}\}_{i=1}^{n}$ generated from some
cumulative distribution function (CDF) $F$. Due to censoring, the
econometrician observes the pair $\left( Y_{i}^{0},D_{i}\right) ^{\intercal
} $ such that
\begin{eqnarray}
Y_{i}^{0} &=&D_{i}T+\left( 1-D_{i}\right) Y_{i} \label{dgp} \\
D_{i} &=&\mathbf{1}\left[ Y_{i}>T\right] , \notag
\end{eqnarray}
where $T$ denotes some constant censoring threshold and $\mathbf{1}\left[
\cdot \right] $ the indicator function. Without loss of generality, we focus
on the right tail. Define $m=\dsum_{i=1}^{n}D_{i}$ as the number of censored
observations. We assume the density of $Y_{i}$, denoted as $f(\cdot )$, is
continuous and positive so that $\mathbb{P}\left( Y_{i}=T\right) =0$.
The model (\ref{dgp}) has spawned a vast literature about estimation and
inference about the mid-sample features, such as median, non-extreme
quantiles, and regression coefficients. See, for example, \citeasnoun{Powell1986},
\citeasnoun{Portnoy2003}, and \citeasnoun{HongTamer2003}. These mid-sample features are
typically estimated at the root-$n$ rate. In contrast, the tail features are
estimated at a much slower rate since only the largest observations are
informative about the right tail.
Define
\begin{equation*}
F_{u}\left( y\right) =\frac{F\left( u+y\right) -F\left( u\right) }{1-F\left(
u\right) }
\end{equation*}
as the conditional CDF given that $Y_{i}$ is larger than some pre-specified
tail cutoff $u$. We aim to approximate $F_{u}\left( y\right) $ by the
generalized Pareto distribution, which is given by
\begin{equation}
G\left( y;\xi ,\sigma \right) =\left\{
\begin{array}{lcc}
1-\left( 1+\frac{\xi y}{\sigma }\right) ^{-1/\xi } & & \xi \neq 0 \\
1-\exp \left( -y/\sigma \right) & & \xi =0
\end{array}
\right. \text{ } \label{GPD}
\end{equation}
with $y\in
\mathbb{R}
^{+\text{ }}$if $\xi \geq 0$ and $y\in \left( 0,-\sigma /\xi \right) $
otherwise. Denote $y_{0}$ as the right end-point of the support of $Y_{i}$.
It is well established in the statistic literature (e.g., \citeasnoun
{BalkemadeHaan1974} and \citeasnoun{Pickands1975}) that the GPD is a good
approximation of $F$ in the tail, in the sense that
\begin{equation}
\lim_{u\rightarrow y_{0}}\sup_{0<y<y_{0}-u}\left\vert F_{u}\left( y\right)
-G\left( y;\xi ,\sigma \right) \right\vert =0 \label{GPD approx}
\end{equation}
for some scale $\sigma $ implicitly depending on $u$, if and only if $F$ is
in the domain of attraction of one of the three limit laws. The parameter $
\xi $ is referred to as the tail index, which is uniquely determined by $F$
and characterizes its tail heaviness. See Chapter 1 of \citeasnoun{deHaan07} for
an overview.
The tail approximation (\ref{GPD approx}) is a mild assumption as it is
satisfied by many commonly used distributions. In particular, the positive $
\xi $ case covers distributions with a Pareto-type tail such as Pareto,
Student-t, and F distributions.\footnote{
In the standard Pareto distribution with the CDF $\mathbb{P}(Y>y)\propto
y^{-\alpha }$, the tail index $\xi $ equals $1/\alpha $. We focus on $\xi $
instead of $\alpha $ for notational ease.} The case with $\xi =0$ covers the
distributions with finite moments of any order. Leading examples are normal
and log-normal distributions. The situation with a negative $\xi $ includes
the distributions with a finite right end-point. For expositional
simplicity, we focus our discussion on the case with $\xi >0$, which covers
the empirical applications with heavy tails.
In practice, we usually choose $u$ as some large order statistic of $Y_{i}$,
say the 95\% empirical quantile. We let $T=T_{n}$ and $u=u_{n}$ depend on
the sample size $n$ and assume $T_{n}>u_{n}$ (otherwise there is no
observation). Also, denote $k$ as the number of the observations between $
u_{n}$ and $T_{n}$ and $\{Y_{(1)}\geq Y_{(2)}\geq ,\ldots ,\geq Y_{(n)}\}$
the order statistics\footnote{
This is different from the conventional notation for order statistics, that
is, $\{Y_{n:n}\geq Y_{n:n-1}\geq ,\ldots ,\geq Y_{n:1}\}$. We think this
alternative is more intuitive in our setup, especially in Section \ref{sec
fk}.} by descending sorting. Then effectively the available observations are
the censored largest $m+k$ order statistics
\begin{equation}
\left( \underset{m}{\underbrace{T_{n},\ldots ,T_{n}}},Y_{(m+1)},\ldots
,Y_{(m+k)}\right) ^{\intercal }, \label{Y}
\end{equation}
where the largest $m$ order statistics, $\{Y_{(1)},\ldots ,Y_{(m)}\}$ are
censored. Using (\ref{GPD approx}), we can write the conditional
log-likelihood of the tail observations as
\begin{align*}
\mathcal{L}_{n}\left( \xi ,\sigma \right) & =\dsum_{i=1}^{m+k}\left\{
D_{i}\log \left( 1-F_{u}\left( T_{n}-u_{n}\right) \right) +\left(
1-D_{i}\right) \log \frac{f\left( Y_{(i)}-u_{n}\right) }{1-F\left(
u_{n}\right) }\right\} \\
& \approx \dsum_{i=1}^{m+k}\left\{ D_{i}\log \left( 1-G\left(
T_{n}-u_{n}\right) \right) +\left( 1-D_{i}\right) \log g\left(
Y_{(i)}-u_{n};\xi ,\sigma \right) \right\} \\
& =\dsum_{i=1}^{m+k}\left\{ -\frac{D_{i}}{\xi }\log \left( 1+\frac{\xi
\left( T_{n}-u_{n}\right) }{\sigma }\right) -\left( 1-D_{i}\right) \log
\sigma \right. \\
& \text{ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ }\left. -\left( 1-D_{i}\right)
\left( 1+\frac{1}{\xi }\right) \log \left( 1+\frac{\xi \left(
Y_{(i)}-u_{n}\right) }{\sigma }\right) \right\} ,
\end{align*}
where $g\left( y;\xi ,\sigma \right) =\partial G\left( y;\xi ,\sigma \right)
/\partial y$. Then the MLE of $\xi $ and $\sigma $ are constructed as
\begin{equation}
\left( \hat{\xi},\hat{\sigma}\right) ^{\intercal }=\arg \max_{\left(
0,\infty \right) ^{2}}\mathcal{L}_{n}\left( \xi ,\sigma \right) .
\label{MLE}
\end{equation}
To derive the asymptotic properties of the MLE, we make the following
assumptions. To simplify notations, we write $\alpha =1/\xi $ when
convenient and define $L\left( y\right) =y^{\alpha }(1-F\left( y\right) )$.
\begin{condition}
\label{cond iid}$\left( Y_{i}^{0},D_{i}\right) ^{\intercal }$ is
independently and identically generated from (\ref{dgp}).
\end{condition}
\begin{condition}
\label{cond Hall82}$F(\cdot )$ is continuously differentiable with $0\left.
<\right. f(\cdot )\left. <\right. \bar{f}$ for some constant $\bar{f}<\infty
$ and satisfies $L\left( y\right) =C(1+\delta y^{-\beta }+o\left( y^{-\beta
}\right) )$ for some constants $\beta >0$, $C\neq 0$ and $\delta \in
\mathbb{R}
$.
\end{condition}
\begin{condition}
\label{cond top}$T_{n}\rightarrow \infty $ and $T_{n}/u_{n}\rightarrow
\kappa \in \left( 1,\infty \right) .$
\end{condition}
\begin{condition}
\label{cond kn}$k\rightarrow \infty $ and $k=o\left( n^{2\beta /\left(
\alpha +2\beta \right) }\right) .$
\end{condition}
Condition \ref{cond iid} assumes a random sample generated with censoring.
Condition \ref{cond Hall82} is imposed by \citeasnoun{Hall82}, which states that
the Pareto tail approximation involves a second-order bias of the order $
y^{-\beta }$. This is imposed to avoid technical complexity and can be
relaxed with other weaker conditions (cf.\ \citeasnoun{Goldie87}). Consider the
Student-t distribution with zero mean, unit variance, and $v$ degrees of
freedom, for example. The CDF\ is given by
\begin{equation*}
1-F_{t(v)}(y)=Cy^{-v}(1+\delta y^{-2}+O(y^{-4}))\text{ as }y\rightarrow
\infty \text{.}
\end{equation*}
Then Condition \ref{cond Hall82} holds with $\xi =1/v$ and $\beta =2$.
Condition \ref{cond top} assumes the censoring threshold is larger than the
tail cutoff. Specifically, the censoring is asymptotically negligible if $
\kappa =\infty $, and leads to no tail observation if $\kappa =1$. Condition
\ref{cond kn} specifies the choice of the tail cutoff and equivalently the
number of tail observations $k$. The very last assumption that $k=o\left(
n^{2\beta /\left( \alpha +2\beta \right) }\right) $ satisfies $
k/n\rightarrow 0$ and is imposed for expositional simplicity. We can relax
it into $kn^{-2\beta /\left( \alpha +2\beta \right) }\rightarrow \mu $ for
some $\mu \in
\mathbb{R}
$ as in \citeasnoun{Smith87} and Chapter 4 of \citeasnoun{deHaan07}. Doing so leads to a
non-zero mean in the asymptotic normal distribution, which further depends
on $\mu $, $\beta $, and $\delta $. These nuisance parameters are hardly
estimable in practice, and hence researchers typically choose a sufficiently
small $k$ to retain the asymptotic zero mean. This is similar in spirit to
the undersmoothing in standard kernel regressions.
Under these conditions, the following proposition establishes the asymptotic
normality of the MLE.
\begin{proposition}
\label{prop GPDindex}Suppose Conditions \ref{cond iid}-\ref{cond kn} hold.
Then
\begin{equation*}
k^{1/2}\binom{\hat{\xi}-\xi }{\frac{\hat{\sigma}}{\sigma }-1}\overset{d}{
\rightarrow }\mathcal{N}\left( 0,M^{-1}\right) ,
\end{equation*}
where the elements of $M$ are given by
\begin{eqnarray*}
M_{11} &=&\frac{2}{\left( 1+\xi \right) (1+2\xi )}+\frac{\kappa ^{-2-1/\xi }
}{(1+\xi )(1+2\xi )\xi ^{2}}\times \\
&&\left\{ -1-\xi +\kappa (2+4\xi )-\kappa ^{2}(1+\xi )(1+2\xi )\right\} \\
M_{22} &=&\frac{1}{\left( 1+2\xi \right) }-\frac{\kappa ^{-2-1/\xi }}{\left(
1+2\xi \right) } \\
M_{12} &=&\frac{1}{\left( 1+\xi \right) (1+2\xi )}+\frac{\kappa ^{-2-1/\xi }
}{\left( 1+\xi \right) (1+2\xi )\xi ^{2}}\times \\
&&\left\{ -\left( 1+\xi \right) ^{2}+\left( 1-2\kappa \right) \left( 1+\xi
\right) (1+2\xi )+\kappa \left( 2+\xi \right) \left( 1+2\xi \right) \right\}
.
\end{eqnarray*}
\end{proposition}
The tail censoring complicates the asymptotic variance substantially as
compared with the no censoring case (cf.\ \citeasnoun{Smith87}). In particular,
when the censoring is asymptotically negligible ($\kappa =\infty $), the
information matrix reduces to
\begin{equation*}
M=\left[
\begin{array}{cc}
\frac{2}{\left( 1+\xi \right) (1+2\xi )} & \frac{1}{\left( 1+\xi \right)
(1+2\xi )} \\
\frac{1}{\left( 1+\xi \right) (1+2\xi )} & \frac{1}{\left( 1+2\xi \right) }
\end{array}
\right] .
\end{equation*}
Given the MLE of $\xi $ and $\sigma $, we can further estimate the extreme
quantile $Q\left( 1-p\right) \equiv \inf \{y:1-p\leq F\left( y\right) \}$.
We set $p=p_{n}\rightarrow 0$ to capture the extremeness. The estimator can
be constructed by inverting (\ref{GPD}), that is,
\begin{equation*}
\hat{Q}\left( 1-p_{n}\right) =u_{n}+\frac{\hat{\sigma}}{\hat{\xi}}\left(
d_{n}^{\hat{\xi}}-1\right) ,
\end{equation*}
where $d_{n}=\left( m+k\right) /\left( p_{n}n\right) $. To derive a
non-trivial asymptotic result, we let $d_{0}\equiv \lim_{n\rightarrow \infty
}d_{n}>0$ so that the target quantile is of the same or larger magnitude of $
u_{n}$ (otherwise it can be estimated by the corresponding empirical
quantile). The following proposition derives the asymptotic distribution of $
\hat{Q}\left( 1-p_{n}\right) $.
\begin{proposition}
\label{prop GPDquan}Suppose Conditions \ref{cond iid}-\ref{cond kn} hold. If
$d_{0}>0$, then
\begin{equation*}
k^{1/2}\frac{\hat{Q}\left( 1-p_{n}\right) -Q\left( 1-p_{n}\right) }{\sigma
q_{\xi }\left( d_{n}\right) }\overset{d}{\rightarrow }\mathcal{N}\left(
0,\Sigma \right)
\end{equation*}
where $q_{\xi }\left( t\right) =\xi ^{-1}t^{\xi }\log t$ and
\begin{equation*}
\Sigma =\left( 1,\frac{d_{0}^{\xi }-1}{\xi q_{\xi }\left( d_{0}\right) }
\right) M^{-1}\left( 1,\frac{d_{0}^{\xi }-1}{\xi q_{\xi }\left( d_{0}\right)
}\right) ^{\intercal }+q_{\xi }\left( d_{0}\right) ^{-2}.
\end{equation*}
\end{proposition}
Proposition \ref{prop GPDquan} establishes the asymptotic normality of the
extreme quantile estimator. Then the confidence intervals for $\xi $ and $
Q\left( 1-p_{n}\right) $ can be constructed by plugging in the estimators
for the asymptotic variance.
\section{Small Sample Modification under the Fixed-$k$ Asymptotics\label{sec
fk}}
The results in Section 2 suggest that the asymptotic normal approximation
can be used for inference about the tail features as $k$ goes to $\infty $.
In practice, however, the choice of the tail sample size $k$ is widely
accepted as a challenging question even without censoring. This is because a
good selection of $k$ has to balance the tail approximation bias and the
variance delicately. Ultimately, the underlying distribution has to be
reasonably close to the Pareto distribution in the tail to guarantee a
satisfactory finite sample performance.
The asymptotic approximation can be quite accurate for some cases, but it is
also easy to find examples where the limiting normal distribution provides a
poor approximation. Consider the example that $F$ is a mixture of the
standard normal distribution\ with probability 0.8 and some Pareto
distribution with probability 0.2. Such a mixture structure implies that
only the very few largest observations are informative about the true\ tail.
In this case, choosing a large $k$ means including too many contaminating
observations from the mid-sample, while choosing a small $k$ invalidates the
asymptotic Gaussianity. In\ principle, there is no such a procedure that
consistently justifies whether a given $k$ is appropriate when $F$ is
entirely unknown. See Theorem 5.1 in M\"{u}ller and Wang (2017) for a
discussion on the non-censored case.
Therefore, in this article, we do not focus on the choice of $k$ but instead
treat it as given.\ In some cases, $k$ is determined by some economic theory
or empirical guidance. For example, in the macroeconomic disaster
application, the economic definition of disasters for more than 10\% of GDP
decline yields the choice of $k$. In other cases, we may employ some
data-driven algorithms that balance the Pareto approximation bias and the
variance. See, for example, \citeasnoun{Hall82}, \citeasnoun{Drees01}, and \citeasnoun
{Clauset09}.
When $k$ and $n/k$ are both sufficiently large, we would expect the MLE in (
\ref{MLE}) based on the increasing-$k$ asymptotics to work well.
Nevertheless, $k$ is only moderate in some situations, including our
macroeconomic disaster application and the Asian male subsample in CPS. This
causes a small sample issue that the asymptotic Gaussianity is questionable.
To find a better alternative, we resort to the asymptotic embedding that
requires $n$ diverges, but $k$ remains a fixed constant. Under this fixed-$k$
asymptotic framework, the consistent estimation of the tail index and the
extreme quantiles are out of the question since the tail sample size is
fixed. Fortunately, inference about these tail features is still
implementable, as we discuss in this section.
We first study the tail index $\xi $. EV theory (the
Fisher--Tippett--Gnedenko theorem) suggests that when the underlying
distribution is within the maximum domain of attraction (e.g., Chapter 1 of
\citeasnoun{deHaan07}), the sample maximum is asymptotically distributed as the EV
distribution, which is parametric and entirely characterized by $\xi $.
Specifically, our Condition \ref{cond Hall82} is sufficient for the maximum
domain of attraction assumption. Then, EV theory implies that there exist
sequences of constants $a_{n}$ and $b_{n}$ such that, up to some location
and scale normalization,
\begin{equation}
\frac{Y_{(1)}-b_{n}}{a_{n}}\overset{d}{\rightarrow }X_{1}, \label{evt_1}
\end{equation}
where the CDF of $X_{1}$ is given by
\begin{equation}
V_{\xi }(x)=\left\{
\begin{array}{lcl}
1-\exp \left( -\left( 1+\xi x/\sigma \right) ^{-1/\xi }\right) & & \xi \neq
0 \\
1-\exp \left( -\exp \left( -x/\sigma \right) \right) & & \xi =0.
\end{array}
\right. \label{evt dist}
\end{equation}
We can subsume $\sigma $ into $a_{n}$ so that $\sigma $ is always $1$ in (
\ref{evt dist}). In additional to the sample maximum, EV theory also extends
to the first $m+k$ order statistics such that if (\ref{evt_1}) holds, then
for any fixed $m$ and $k$,
\begin{equation}
\left( \frac{Y_{(1)}-b_{n}}{a_{n}},...,\frac{Y_{(m+k)}-b_{n}}{a_{n}}\right)
^{\intercal }\overset{d}{\rightarrow }\left( X_{1},...,X_{m+k}\right)
^{\intercal }. \label{evt_conv}
\end{equation}
The joint density of $\mathbf{(}X_{1},\ldots ,X_{m+k}\mathbf{)}^{\intercal }$
is given by $V_{\xi }(x_{m+k})\prod_{i=1}^{m+k}v_{\xi }(x_{i})/V_{\xi
}(x_{i})$ on $x_{m+k}\leq x_{m+k-1}\leq \ldots \leq x_{1}$, where $v_{\xi
}(x)=dV_{\xi }(x)/dx$.
Since the first $m$ elements are censored, the effective tail observations
asymptotically reduce to
\begin{equation*}
\mathbf{X}=\left( X_{m+1},...,X_{m+k}\right) ^{\intercal },
\end{equation*}
whose density is derived in the following proposition.
\begin{proposition}
\label{prop evt}Suppose Conditions \ref{cond iid} and \ref{cond Hall82}
hold. Then for any fixed positive integers $m$ and $k$, there exist
sequences of constants $a_{n}$ and $b_{n}$ such that
\begin{equation*}
\frac{\left( Y_{(m+1)},\ldots ,Y_{(m+k)}\right) ^{\intercal }-b_{n}\iota _{k}
}{a_{n}}\text{ }\overset{d}{\rightarrow }\mathbf{X,}
\end{equation*}
where $\iota _{k}\ $denotes the $k\times 1$ vector of ones, and the joint
density of $\mathbf{X}$ is given by
\begin{eqnarray}
f_{\mathbf{X|}\xi }\left( x_{m+1},...,x_{m+k}\right) &=&\frac{1}{m!}\left(
-\log V_{\xi }\left( x_{m+1}\right) \right) ^{m}V_{\xi }\left(
x_{m+k}\right) \tprod_{i=m+1}^{m+k}v_{\xi }\left( x_{i}\right) /V_{\xi
}\left( x_{i}\right) \label{fx} \\
&=&\frac{1}{m!}\exp \left(
\begin{array}{c}
-\frac{m}{\xi }\log \left( 1+\xi x_{m+1}\right) -\left( 1+\xi x_{m+k}\right)
^{-1/\xi } \\
-\left( 1+\frac{1}{\xi }\right) \dsum_{i=1}^{k}\log \left( 1+\xi
x_{m+i}\right)
\end{array}
\right) . \notag
\end{eqnarray}
\end{proposition}
It is clear that the elements of $\mathbf{X}$ are dependent as captured by
the term $\left( -\log V_{\xi }\left( x_{m+1}\right) \right) ^{m}V_{\xi
}\left( x_{m+k}\right) $, which is negligible if $k$ is large but plays a
vital role when $k$ is only moderate. This is the fundamental difference
between the increasing-$k$ and the fixed-$k$ asymptotic embeddings, which
are respectively used by the MLE and its small sample modification. From now
on, we use bold letters to denote vectors.
If the constants $a_{n}$ and $b_{n}$ were known, the vector
\begin{equation*}
\mathbf{Y}=\left( Y_{(m+1)},\ldots ,Y_{(m+k)}\right) ^{\intercal }
\end{equation*}
is then approximately distributed as $\mathbf{X}$, and the limiting problem
is reduced to the small sample parametric one: constructing a confidence
interval based on one draw $\mathbf{X}$ whose density $f_{\mathbf{X}|\xi }$
is known up to $\xi $. However, $a_{n}$ and $b_{n}$ respectively correspond
to the scale $\sigma $ and the tail location $u$. Therefore, they ultimately
depend on $F$ and are challenging to estimate. Consider the standard Pareto
distribution, for example. The Pareto exponent $\alpha $ is simply $1/\xi $.
Then the fact that $a_{n}=n^{\xi }$ implies that a small estimation bias in $
\xi $ could be amplified by the $n$-power and lead to a poor inference.
To avoid the knowledge (and estimation) of $a_{n}$ and $b_{n}$, we consider
the following self-normalized statistics:
\begin{eqnarray}
\mathbf{Y}^{\ast } &=&\frac{\mathbf{Y}-Y_{(m+k)}\iota _{k}}{
Y_{(m+1)}-Y_{(m+k)}} \label{max_inv} \\
&=&\left( 1,\frac{Y_{(m+2)}-Y_{(m+k)}}{Y_{(m+1)}-Y_{(m+k)}},...,\frac{
Y_{(m+k-1)}-Y_{(m+k)}}{Y_{(m+1)}-Y_{(m+k)}},0\right) ^{\intercal }. \notag
\end{eqnarray}
It is easy to establish that $\mathbf{Y}^{\mathbf{\ast }}$ is maximal
invariant with respect to the group of location and scale transformations
(cf. Chapter 6 of \citeasnoun{Lehmann05}). In words, the statistic constructed as
a function of $\mathbf{Y}^{\ast }$ remains unchanged if data are shifted and
multiplied by any non-zero constant. This invariance is also intuitive since
the tail shape should preserve no matter how data are linearly transformed.
The continuous mapping theorem and Proposition \ref{prop evt} yield that
\begin{equation}
\mathbf{Y}^{\ast }\overset{d}{\rightarrow }\mathbf{X}^{\ast }=\left( 1,\frac{
X_{m+2}-X_{m+k}}{X_{m+1}-X_{m+k}},...,\frac{X_{m+k-1}-X_{m+k}}{
X_{m+1}-X_{m+k}},0\right) ^{\intercal }, \label{X_sm}
\end{equation}
which is again invariant to location and scale transformation. By change of
variables, the density of $\mathbf{X}^{\ast }$ is given by
\begin{equation}
f_{\mathbf{X}^{\mathbf{\ast }}|\xi }\left( \mathbf{x}^{\mathbf{\ast }
}\right) =\frac{\Gamma \left( k+m\right) }{m!}\int_{0}^{\infty }s^{k-2}\exp
\left( -\frac{m}{\xi }\log \left( 1+\xi s\right) \right) e\left( \mathbf{x}^{
\mathbf{\ast }},s\right) ds, \label{fxstar}
\end{equation}
where $e\left( \mathbf{x}^{\mathbf{\ast }},s\right) =\exp \left( -(1+1/\xi
)\sum_{i=1}^{k}\log (1+\xi x_{i}^{\mathbf{\ast }}s)\right) $ and $x_{i}^{
\mathbf{\ast }}$ denotes the $i$th component of $\mathbf{x}^{\mathbf{\ast }}$
. See Appendix A.1 for more details.
The density (\ref{fxstar}) can be used to conduct inference about $\xi $. In
particular, consider the hypothesis testing problem
\begin{equation*}
H_{0}:\xi =\xi _{0}\text{ against }H_{1}:\xi \in \Xi \backslash \{\xi _{0}\},
\end{equation*}
where $\Xi $ denotes the parameter space of $\xi $. We set $\Xi =(0,1)$ in
the Monte Carlo simulations to cover the distributions with an unbounded
support and a finite mean, which can be easily extended. Since the
alternative hypothesis is composite, we follow \citeasnoun{Andrews94} and \citeasnoun
{EMW15} to consider the weighted average alternative
\begin{equation*}
\int_{\Xi }f_{\mathbf{X}^{\ast }|\xi }\left( \mathbf{x}^{\ast }\right)
dW(\xi ),
\end{equation*}
where the weighting measure $W$ reflects the importance a researcher
attaches to different alternative values of $\xi $. In practice, we set $
W(\cdot )$ to be the CDF of the standard uniform distribution for simplicity.
Given the density of $\mathbf{X}^{\ast }$, we construct the likelihood-ratio
test as
\begin{equation}
\varphi \left( \mathbf{x}^{\ast }\right) =\mathbf{1}\left[ \frac{\int_{\Xi
}f_{\mathbf{X}^{\ast }|\xi }\left( \mathbf{x}^{\ast }\right) dW(\xi )}{f_{
\mathbf{X}^{\ast }|\xi _{0}}\left( \mathbf{x}^{\ast }\right) }>\mathrm{cv}
\left( \xi _{0},k,m\right) \right] , \label{LRtest}
\end{equation}
where $\mathrm{cv}\left( \xi _{0},k,m\right) $ denotes the critical value
that depends on the significance level, the null value $\xi _{0}$, the tail
sample size $k$, and the number of the censored observations $m$. The
critical value is obtained by simulation, and the test is implemented by
replacing $\mathbf{X}^{\ast }$ with $\mathbf{Y}^{\mathbf{\ast }}$ in finite
samples. The continuous mapping theorem and Proposition \ref{prop evt} yield
that $\mathbb{E}\left[ \varphi \left( \mathbf{Y}^{\mathbf{\ast }}\right)
\right] \rightarrow \mathbb{E}\left[ \varphi \left( \mathbf{X}^{\ast
}\right) \right] $, which equals the nominal level under the null
hypothesis. The confidence interval is constructed by inverting the test.
Note that this method does not require the knowledge of the censoring
threshold $T$, which applies to cases such as the macroeconomic disaster.
Following \citeasnoun{MuellerWang17}, we can also construct the confidence
intervals of the extreme quantiles under the fixed-$k$ asymptotics. To this
end, we focus on the $Q(1-p_{n})$ quantile with $p_{n}=O(n^{-1})$, which
captures the fact that the object of interest is of the same order of
magnitude as the sample maximum. In particular, we consider $p_{n}=h/n$ for
some fixed $h>0$. Then EV theory implies that $\left( Q\left( 1-h/n\right)
-b_{n}\right) /a_{n}$ converges to the $e^{-h}$ quantile of $X_{1}$, denoted
$q(\xi ,h)=\left( h^{-\xi }-1\right) /\xi $. Again, the research problem
would become inference about $q(\xi ,h)$ based on the $k\times 1$ vector of
observations $\mathbf{X}$ if $a_{n}$ and $b_{n}$ were known. Without loss of
generality, we construct a confidence set $S(\mathbf{Y})\subset
\mathbb{R}
$ such that $\mathbb{P}\left( Q\left( 1-p_{n}\right) \left. \in \right. S(
\mathbf{Y})\right) \geq 1-\mathrm{lv}$, at least as $n\rightarrow \infty $,
where $\mathrm{lv}$ denotes the significance level.
To eliminate $a_{n}$ and $b_{n}$, we use the self-normalized vector $\mathbf{
Y}^{\ast }$ as in (\ref{fxstar}). Besides, we also impose location and scale
equivariance on our confidence interval $S$. Specifically, we impose that
for any constants $a>0$ and $b$, our interval $S$ satisfies that $S(a\mathbf{
Y}+b)=aS(\mathbf{Y})+b$, where $aS(\mathbf{Y})+b=\{y:(y-b)/a\in S(\mathbf{Y}
)\}$. Under this equivariance constraint, we can write
\begin{eqnarray*}
\mathbb{P}(Q\left( 1-p_{n}\right) \left. \in \right. S(\mathbf{Y})) &=&
\mathbb{P}\left( \frac{Q\left( 1-p_{n}\right) -b_{n}}{a_{n}}\in S\left(
\frac{\mathbf{Y}-b_{n}\iota _{k}}{a_{n}}\right) \right) \\
&=&\mathbb{P}\left( \frac{Q\left( 1-p_{n}\right) -Y_{(m+k)}}{
Y_{(m+1)}-Y_{(m+k)}}\in S\left( \mathbf{Y}^{\ast }\right) \right) \\
&\rightarrow &\mathbb{P}_{\xi }\left( \frac{q(\xi ,h)-X_{m+k}}{
X_{m+1}-X_{m+k}}\in S(\mathbf{X}^{\ast })\right) ,
\end{eqnarray*}
where the notation $\mathbb{P}_{\xi }$ (and $\mathbb{E}_{\xi }$ below)
indicates that the randomness is entirely characterized by $\xi $
asymptotically. The asymptotic problem then is the construction of a
location and scale equivariant $S$ that satisfies
\begin{equation}
\mathbb{P}_{\xi }\left( \frac{q(\xi ,h)-X_{m+k}}{X_{m+1}-X_{m+k}}\in S(
\mathbf{X}^{\ast })\right) \geq 1-\mathrm{lv}\text{ for all }\xi \in \Xi
\label{asy_cov}
\end{equation}
since any $S$ that satisfies Proposition \ref{prop evt} and the equivariance
constraint also satisfies
\begin{equation*}
\lim \inf_{n\rightarrow \infty }\mathbb{P}(Q\left( 1-p_{n}\right) \in S(
\mathbf{Y}))\geq 1-\mathrm{lv}.
\end{equation*}
This problem involves a single observation $\mathbf{X}\in
\mathbb{R}
^{k}$ from a parametric distribution indexed only by the scalar parameter $
\xi \in \Xi $.
In principle, there could still be many solutions that satisfy the
asymptotic size constraint. To obtain the optimal one, we consider the
weighted average expected length criterion
\begin{equation}
\int \mathbb{E}_{\xi }[\func{lgth}(S(\mathbf{X}))]dW(\xi )\text{,}
\label{asy_length}
\end{equation}
where $W$ again denotes some weighting measure on $\Xi $, and $\func{lgth}
(A)=\int \mathbf{1}[y\in A]dy$ for any Borel set $A\subset
\mathbb{R}
$.
To solve the program of minimizing (\ref{asy_length}) subject to (\ref
{asy_cov}) among all equivariant set estimators $S$, we introduce
\begin{equation*}
Y^{\ast }(\xi )=\frac{q(\xi ,h)-X_{m+k}}{X_{m+1}-X_{m+k}},
\end{equation*}
and write $\mathbb{E}_{\xi }[\func{lgth}(S(\mathbf{X}))]\left. =\right.
\mathbb{E}_{\xi }[(X_{m+1}-X_{m+k})\func{lgth}(S(\mathbf{X}^{\ast }))]\left.
=\right. \mathbb{E}_{\xi }[\kappa _{\xi }(\mathbf{X}^{\ast })\func{lgth}(S(
\mathbf{X}^{\ast }))]$ with $\kappa _{\xi }(\mathbf{X}^{\ast })=\mathbb{E}
_{\xi }[X_{m+1}-X_{m+k}|\mathbf{X}^{\ast }]$. Thus, our problem becomes
\begin{equation}
\begin{tabular}{l}
$\min_{S(\cdot )}\int_{\Xi }\mathbb{E}_{\xi }[\kappa _{\xi }(\mathbf{X}
^{\ast })\func{lgth}(S(\mathbf{X}^{\ast }))]dW(\xi )$ \\
\multicolumn{1}{r}{$\text{s.t. }\mathbb{P}_{\xi }\left( Y^{\ast }(\xi )\in S(
\mathbf{X}^{\ast })\right) \geq 1-\mathrm{lv}\text{ for all }\xi \in \Xi .$}
\end{tabular}
\label{program}
\end{equation}
There are two advantages to translate the asymptotic problem into (\ref
{program}). First, (\ref{program}) does not require the knowledge of the
censoring threshold $T$ but only the number of censored observations $m$.
Second, (\ref{program}) only involves $S$ evaluated at $\mathbf{X}^{\ast }$
and hence $\mathbf{Y}^{\ast }$ in practice. This means the knowledge of $
a_{n}$ and $b_{n}$ is asymptotically unnecessary as long as $n$ is
sufficiently larger than $m+k$. Note that any solution to (\ref{program})
also provides the form of $S$, that is, $S(\mathbf{X})=(X_{m+1}-X_{m+k})S(
\mathbf{X}^{\ast })+X_{m+k}$. So once $S(\cdot )$ is determined, the
confidence interval can be constructed in practice by plugging in
\begin{equation*}
(Y_{(m+1)}-Y_{(m+k)})S(\mathbf{Y}^{\ast })+Y_{(m+k)}.
\end{equation*}
To make further progress in solving (\ref{program}), we write the problem in
the following Lagrangian form:
\begin{equation*}
\min_{S(\cdot )}\int_{\Xi }\mathbb{E}_{\xi }[\kappa _{\xi }(\mathbf{X}^{\ast
})\func{lgth}(S(\mathbf{X}^{\ast }))]dW(\xi )+\int_{\Xi }\mathbb{P}_{\xi
}\left( Y^{\ast }(\xi )\in S(\mathbf{X}^{\ast })\right) d\Lambda (\xi ),
\end{equation*}
where the non-negative measure $\Lambda $ denotes the Lagrangian weights
that guarantee the asymptotic size constraint. By writing the expectations
above as integrals over the densities $f_{\mathbf{X}^{\ast }\mathbf{|}\xi }$
of $\mathbf{X}^{\ast }$ and $f_{Y^{\ast }(\xi ),\mathbf{X}^{\ast }|\xi }$ of
$(Y^{\ast }(\xi ),\mathbf{X}^{\ast })$, the solution of the above problem is
given by
\begin{equation}
S(\mathbf{x}^{\ast })=\left\{ y:\int_{\Xi }\kappa _{\xi }(\mathbf{x}^{\ast
})f_{\mathbf{X}^{\ast }|\xi }(\mathbf{x}^{\ast })dW(\xi )<\int_{\Xi
}f_{Y^{\ast }(\xi ),\mathbf{X}^{\ast }|\xi }(y,\mathbf{x}^{\ast })d\Lambda
(\xi )\right\} . \label{S_Lambda}
\end{equation}
The integrals can be numerically calculated by Gaussian quadrature, and then
the only remaining challenge is to find some suitable Lagrangian weights $
\Lambda $. We solve this challenge by the numerical approach developed in
\citeasnoun{EMW15}. The MATLAB program and the weights $\Lambda $ are available at
the author's website. Note that $\Lambda $ only needs to be computed once by
the author instead of empirical researchers. Then the most time-consuming
part in solving the program (\ref{program}) is the numerical integration,
which costs only a few seconds in a modern PC. Further details are provided
in Appendix A.1.
\section{Monte Carlo Simulations \label{sec mc}}
This section examines the finite sample performance of the proposed method
and compares it with several popular existing methods. We generate random
samples from four commonly used distributions: the generalized Pareto
distribution with $\xi =$ $0.5$ and $\sigma =1$ (GPD), the absolute value of
the Student-t distribution with 2 degrees of freedom (\TEXTsymbol{\vert}t(2)
\TEXTsymbol{\vert}), the F distribution with parameters 4 and 4 (F(4,4)),
and the double Pareto-lognormal distribution (dPlN), that is,
\begin{equation*}
Y=\exp \left( c_{1}+c_{2}Z_{1}+\xi Z_{2}-c_{3}Z_{3}\right) ,
\end{equation*}
where $Z_{1}$, $Z_{2}$, $Z_{3}$ are independent and $Z_{1}\sim N(0,1)$, and $
Z_{2},Z_{3}\sim Exp(1)$. For parameter values, we set $c_{1}=0$, $c_{2}=0.5$
, $\xi =0.5$, and $c_{3}=1$, which are typical values for income data as
documented in \citeasnoun{Toda12}. In particular, the dPlN distribution is the
product of independent double Pareto and lognormal variables. It has been
documented to fit well to size distributions of economic variables including
income (\citeasnoun{Reed03}), city size (\citeasnoun{Giesen10}), and consumption (\citeasnoun
{Toda17}). In all four DGP's, the true value of the tail index is 0.5.
Regarding the tail censoring, we set the censoring threshold $T$ as the 99\%
and 99.9\% quantiles of the underlying distributions, implying that the
censored probability (cen\_p) is either 1\% or 0.1\%.
We first consider some widely used estimators in empirical studies. Due to
space limitations, we only report the results of Hill (1975)'s estimator
\nocite{Hill75} and the bias-corrected estimator proposed by \citeasnoun{Gabaix11}
(denoted GI). The confidence intervals are based on their asymptotic
normality and the plug-in estimators of their asymptotic variances. The
sample size $n$ is 1000, 2000, and 5000, and $k$ is set as $\left[ 0.05n
\right] $ for both methods, where $[A]$ denotes the closest integer of $A$.
All results are based on 1000 simulations.
Table \ref{tbl index Hill} depicts the mean biases and the coverage
probabilities of these two methods. Several key findings can be summarized
as follows. First, both the Hill and the GI estimators suffer from severe
biases, and the confidence intervals based on them exhibit substantial
undercoverage. This holds even if the censoring probability is only 0.1\%.
Second, ignoring the upper tail censoring tends to underestimate the tail
index, which implies a misleadingly thin tail. This is seen in Section \ref
{sec disaster} when we study the macroeconomic disasters. Finally,
unreported results show that other methods reviewed in Chapter 3 of \citeasnoun
{deHaan07} also suffer from substantial undercoverage. Therefore, it is
crucial to take the censoring into account, even if the censoring
probability is tiny.
\begin{table}[H]
\begin{center}
\caption{Small Sample Properties of Estimation and Inference about Tail
Index, Igorning Tail Censoring}\label{tbl index Hill}
\vspace{+2ex}
\begin{tabular}{lcccccccccc}
\hline
cen\_p & & \multicolumn{4}{c}{$1\%$} & & \multicolumn{4}{c}{$0.1\%$} \\
\cline{3-6}\cline{8-11}
& & \multicolumn{2}{c}{Bias} & \multicolumn{2}{c}{Cov} & &
\multicolumn{2}{c}{Bias} & \multicolumn{2}{c}{Cov} \\ \hline
n=1000 & & Hill & GI & Hill & GI & & Hill & GI & Hill & GI \\
GPD & & -0.18 & -0.24 & 0.03 & 0.00 & & -0.04 & -0.07 & 0.88 & 0.89 \\
t(2) & & -0.16 & -0.23 & 0.10 & 0.00 & & -0.02 & -0.06 & 0.93 & 0.93 \\
F(4,4) & & -0.12 & -0.20 & 0.39 & 0.05 & & 0.03 & -0.03 & 0.98 & 0.98 \\
dPlN & & -0.17 & -0.24 & 0.05 & 0.00 & & -0.04 & -0.04 & 0.89 & 0.90 \\
\hline
n=2000 & & Hill & GI & Hill & GI & & Hill & GI & Hill & GI \\
GPD & & -0.18 & -0.24 & 0.00 & 0.00 & & -0.04 & -0.07 & 0.85 & 0.81 \\
t(2) & & -0.16 & -0.23 & 0.00 & 0.00 & & -0.02 & -0.06 & 0.93 & 0.86 \\
F(4,4) & & -0.12 & -0.20 & 0.12 & 0.00 & & 0.03 & -0.03 & 0.97 & 0.97 \\
dPlN & & -0.17 & -0.24 & 0.00 & 0.00 & & -0.04 & -0.04 & 0.88 & 0.82 \\
\hline
n=5000 & & Hill & GI & Hill & GI & & Hill & GI & Hill & GI \\
GPD & & -0.18 & -0.24 & 0.00 & 0.00 & & -0.04 & -0.08 & 0.73 & 0.45 \\
t(2) & & -0.16 & -0.23 & 0.00 & 0.00 & & -0.02 & -0.07 & 0.90 & 0.63 \\
F(4,4) & & -0.12 & -0.20 & 0.00 & 0.00 & & 0.03 & -0.03 & 0.93 & 0.95 \\
dPlN & & -0.17 & -0.24 & 0.00 & 0.00 & & -0.03 & -0.08 & 0.77 & 0.47 \\
\hline
\end{tabular}
\vspace{-4ex}
\end{center}
\begin{singlespacing}
\begin{footnotesize}
Note: Entries are the biases and coverage probabilities (Cov) of the 95\%
confidence intervals based on Hill's estimator (Hill) and Gabaix and
Ibragimov (2010)'s estimator (GI). Data are generated from the Pareto(0.5),
the absolute value of Student-t(2), the F(4,4), and the dPIN distributions
with the censored probability (cen\_p) being 1\% or 0.01\%. The results are
based on 1000 simulations.
\end{footnotesize}
\end{singlespacing}
\end{table}
Now we implement the new method proposed in Sections \ref{sec inck} and \ref
{sec fk}. Table \ref{tbl index} depicts the coverage and length of the 95\%
maximum likelihood confidence intervals (denoted ml) based on Proposition
\ref{prop GPDindex} and those of the fixed-$k$ intervals (denoted fk) by
inverting (\ref{LRtest}). Several interesting findings can be made as
follows. First, the maximum likelihood confidence intervals are
substantially longer than the fixed-$k$ ones when the sample size is not
large. Besides, the coverage probability is smaller than the nominal level
when the censoring is at the 99.9\% quantile. This is because the asymptotic
normality cannot perform well when $k$ is not large. Second, in comparison,
the fixed-$k$ ones always deliver the nominal size with shorter length,
especially when the sample size is not large. Finally, when $n$ reaches 5000
(and $k$ reaches 250), the maximum likelihood intervals are comparable with
the fixed-$k$ ones. Hence a simple rule-of-thumb choice of the switching
cutoff is $k\lessgtr 250$, provided $n$ is sufficiently large.
\begin{table}[H]
\begin{center}
\caption{Small Sample Properties of Inference about Tail Index}\label{tbl
index}
\vspace{+2ex}
\begin{tabular}{lcccccccccc}
\hline
cen\_p & & \multicolumn{4}{c}{$1\%$} & & \multicolumn{4}{c}{$0.1\%$} \\
\cline{3-6}\cline{8-11}
& & \multicolumn{2}{c}{Cov} & \multicolumn{2}{c}{Lgth} & &
\multicolumn{2}{c}{Cov} & \multicolumn{2}{c}{Lgth} \\ \hline
n=1000 & & ml & fk & ml & fk & & ml & fk & ml & fk \\
GPD & & 0.98 & 0.93 & 1.39 & 0.73 & & 0.91 & 0.95 & 0.88 & 0.70 \\
t(2) & & 0.98 & 0.94 & 1.40 & 0.73 & & 0.90 & 0.95 & 0.87 & 0.70 \\
F(4,4) & & 0.99 & 0.93 & 1.39 & 0.73 & & 0.90 & 0.95 & 0.87 & 0.70 \\
dPlN & & 0.99 & 0.93 & 1.30 & 0.73 & & 0.90 & 0.94 & 0.87 & 0.70 \\ \hline
n=2000 & & ml & fk & ml & fk & & ml & fk & ml & fk \\
GPD & & 0.96 & 0.94 & 0.99 & 0.69 & & 0.93 & 0.94 & 0.63 & 0.58 \\
t(2) & & 0.96 & 0.93 & 0.99 & 0.69 & & 0.92 & 0.93 & 0.62 & 0.58 \\
F(4,4) & & 0.96 & 0.94 & 0.99 & 0.69 & & 0.93 & 0.94 & 0.63 & 0.58 \\
dPlN & & 0.96 & 0.94 & 0.99 & 0.69 & & 0.92 & 0.93 & 0.62 & 0.58 \\ \hline
n=5000 & & ml & fk & ml & fk & & ml & fk & ml & fk \\
GPD & & 0.97 & 0.94 & 0.63 & 0.54 & & 0.95 & 0.94 & 0.40 & 0.39 \\
t(2) & & 0.96 & 0.94 & 0.62 & 0.54 & & 0.93 & 0.92 & 0.40 & 0.38 \\
F(4,4) & & 0.97 & 0.95 & 0.63 & 0.54 & & 0.94 & 0.93 & 0.40 & 0.38 \\
dPlN & & 0.96 & 0.93 & 0.62 & 0.54 & & 0.94 & 0.94 & 0.40 & 0.38 \\ \hline
\end{tabular}
\vspace{-4ex}
\end{center}
\begin{singlespacing}
\begin{footnotesize}
Note: Entries are the coverage probabilities (Cov) and the averaged length
(Lgth) of the maximum likelihood intervals (ml) and the fixed-$k$ intervals
(fk) for the tail index. Data are generated from the Pareto(0.5), the
absolute value of Student-t(2), the F(4,4), and the dPIN distributions with
the censored probability (cen\_p) being 1\% or 0.01\%. The results are based
on 1000 simulations. The level of significance is 5\%.
\end{footnotesize}
\end{singlespacing}
\end{table}
Tables \ref{tbl quan99} depicts the coverage probabilities and lengths of
the confidence intervals of the 99\% quantile, using either the maximum
likelihood method as in Proposition \ref{prop GPDquan}\ or the fixed-$k$
method (\ref{S_Lambda}). Both methods deliver satisfactory size and length
properties, although the maximum likelihood intervals suffer from slight
undercoverage. However, as we target the more extreme 99.9\% quantile as in
Table \ref{tbl quan999}, such undercoverage is substantial when $k$ is less
than 250. In contrast, the fixed-$k$ ones always perform excellently. These
results reinforce our switching cutoff at $k=250$.
\begin{table}[H]
\begin{center}
\caption{Small Sample Properties of Inference about the $0.99$ Quantile}
\label{tbl quan99}
\vspace{+2ex}
\begin{tabular}{lcccccccccc}
\hline
cen\_p & & \multicolumn{4}{c}{$1\%$} & & \multicolumn{4}{c}{$0.1\%$} \\
\cline{3-6}\cline{8-11}
& & \multicolumn{2}{c}{Cov} & \multicolumn{2}{c}{Lgth} & &
\multicolumn{2}{c}{Cov} & \multicolumn{2}{c}{Lgth} \\ \hline
n=1000 & & ml & fk & ml & fk & & ml & fk & ml & fk \\
GPD & & 0.95 & 0.94 & 6.71 & 7.09 & & 0.91 & 0.96 & 4.83 & 5.79 \\
t(2) & & 0.95 & 0.94 & 6.73 & 7.13 & & 0.91 & 0.94 & 4.87 & 5.89 \\
F(4,4) & & 0.94 & 0.94 & 11.67 & 12.17 & & 0.91 & 0.95 & 8.32 & 9.91 \\
dPlN & & 0.95 & 0.94 & 4.92 & 5.25 & & 0.91 & 0.95 & 3.54 & 4.37 \\ \hline
n=2000 & & ml & fk & ml & fk & & ml & fk & ml & fk \\
GPD & & 0.96 & 0.94 & 4.60 & 4.86 & & 0.92 & 0.96 & 3.38 & 3.88 \\
t(2) & & 0.96 & 0.94 & 4.60 & 4.76 & & 0.93 & 0.96 & 3.43 & 3.89 \\
F(4,4) & & 0.96 & 0.95 & 8.04 & 8.09 & & 0.92 & 0.95 & 5.85 & 6.72 \\
dPlN & & 0.96 & 0.94 & 3.44 & 3.50 & & 0.93 & 0.96 & 2.52 & 2.90 \\ \hline
n=5000 & & ml & fk & ml & fk & & ml & fk & ml & fk \\
GPD & & 0.97 & 0.96 & 2.85 & 2.93 & & 0.93 & 0.94 & 2.12 & 2.38 \\
t(2) & & 0.97 & 0.95 & 2.84 & 2.91 & & 0.94 & 0.96 & 2.15 & 2.40 \\
F(4,4) & & 0.97 & 0.95 & 4.93 & 5.05 & & 0.93 & 0.94 & 3.68 & 4.10 \\
dPlN & & 0.97 & 0.95 & 2.08 & 2.12 & & 0.93 & 0.96 & 1.57 & 1.77 \\ \hline
\end{tabular}
\vspace{-4ex}
\end{center}
\begin{singlespacing}
\begin{footnotesize}
Note: Entries are the coverage probabilities (Cov) and the averaged length
(Lgth) of the maximum likelihood intervals (ml) and the fixed-$k$ intervals
(fk) for the 99\% quantiles. Data are generated from the Pareto(0.5), the
absolute value of Student-t(2), the F(4,4), and the dPIN distributions with
the censored probability (cen\_p) being 1\% or 0.01\%. The results are based
on 1000 simulations. The level of significance is 5\%.
\end{footnotesize}
\end{singlespacing}
\end{table}
\begin{table}[H]
\begin{center}
\caption{Small Sample Properties of Inference about the $0.999$ Quantile}
\label{tbl quan999}
\vspace{+2ex}
\begin{tabular}{lcccccccccc}
\hline
cen\_p & & \multicolumn{4}{c}{$1\%$} & & \multicolumn{4}{c}{$0.1\%$} \\
\cline{3-6}\cline{8-11}
& & \multicolumn{2}{c}{Cov} & \multicolumn{2}{c}{Lgth} & &
\multicolumn{2}{c}{Cov} & \multicolumn{2}{c}{Lgth} \\ \hline
n=1000 & & ml & fk & ml & fk & & ml & fk & ml & fk \\
GPD & & 0.86 & 0.92 & 128.9 & 102.6 & & 0.83 & 0.95 & 61.68 & 75.08 \\
t(2) & & 0.85 & 0.93 & 123.0 & 103.1 & & 0.83 & 0.94 & 60.87 & 72.62 \\
F(4,4) & & 0.85 & 0.91 & 226.8 & 178.1 & & 0.83 & 0.95 & 106.3 & 130.8 \\
dPlN & & 0.85 & 0.93 & 93.16 & 78.21 & & 0.82 & 0.95 & 45.16 & 55.13 \\
\hline
n=2000 & & ml & fk & ml & fk & & ml & fk & ml & fk \\
GPD & & 0.89 & 0.92 & 79.20 & 72.25 & & 0.87 & 0.94 & 39.71 & 44.91 \\
t(2) & & 0.88 & 0.94 & 75.60 & 71.65 & & 0.88 & 0.95 & 39.32 & 45.96 \\
F(4,4) & & 0.89 & 0.92 & 136.6 & 126.5 & & 0.88 & 0.93 & 69.35 & 80.28 \\
dPlN & & 0.88 & 0.92 & 56.17 & 51.48 & & 0.87 & 0.93 & 28.99 & 32.38 \\
\hline
n=5000 & & ml & fk & ml & fk & & ml & fk & ml & fk \\
GPD & & 0.94 & 0.95 & 43.06 & 40.41 & & 0.92 & 0.94 & 23.96 & 24.15 \\
t(2) & & 0.92 & 0.95 & 40.75 & 39.25 & & 0.92 & 0.93 & 23.39 & 24.15 \\
F(4,4) & & 0.92 & 0.95 & 73.55 & 70.90 & & 0.91 & 0.94 & 41.11 & 41.62 \\
dPlN & & 0.92 & 0.92 & 30.65 & 28.32 & & 0.91 & 0.93 & 17.31 & 18.13 \\
\hline
\end{tabular}
\vspace{-4ex}
\end{center}
\begin{singlespacing}
\begin{footnotesize}
Note: Entries are the coverage probabilities (Cov) and the averaged length
(Lgth) of the maximum likelihood intervals (ml) and the fixed-k intervals
(fk) for the 99.9\% quantiles. Data are generated from the Pareto(0.5), the
absolute value of Student-t(2), the F(4,4), and the dPIN distributions with
the censored probability (cen\_p) being 1\% or 0.01\%. The results are based
on 1000 simulations. The level of significance is 5\%.
\end{footnotesize}
\end{singlespacing}
\end{table}
\section{Empirical Applications\label{sec emp}}
Many applications in economics and finance involve estimation and inference
of tail features with censored data. In this section, we apply the proposed
method to the two datasets we discussed earlier in this paper. Our empirical
analysis highlights the potential of our approach.
\subsection{US Individual Earnings\label{sec income}}
Our first application is about the tail features of the individual earnings
distribution. Following the convention, we use the variable ERN\_VAL in the
March CPS dataset and drop the individuals that are younger than 18 or older
than 70 years old. This yields 115,424 observations in the 2019 sample. The
censoring threshold is 310000 USD, which leads to a 0.58\% censoring
fraction in the full sample and various censoring fractions in different
subsamples. The first several columns in Table \ref{tbl cps} present the
sample sizes ($n$) and the numbers (cen\#) and the fractions (cen\%) of the
censored observations, respectively. We use the previously introduced method
to construct the 95\% confidence intervals of the tail index and the 99\%
and 99.9\% quantiles. Specifically, we follow the simulation study to use
the maximum likelihood confidence intervals developed in Section \ref{sec
inck} when $k$ is larger than 250 and switch to the fixed-$k$ confidence
intervals (\ref{S_Lambda}) otherwise. The last six columns in Table \ref{tbl
cps} present the results with $k=\left[ 0.05n\right] $. The results based on
other choices are similar and reported in Appendix A.3.
Several interesting findings can be summarized as follows. First, in Panel
A, the tail index is around 0.5 in the full sample, as commonly found in the
existing literature. But it is substantially different across subsamples.
Second, the tail also exhibits substantial heterogeneity across genders. In
particular, the male sample has significantly higher quantiles than the
female at both the 99\% and 99.9\% levels. Third, this difference also
exists across races. In particular, the 99.9\% quantile of all males is at
least twice larger than that of the black males. All such heterogeneity
provides new evidence for potential racial and gender discrimination.
Finally, Panel B depicts the heterogeneity across ages, with substantially
heavier tails showing up in the middle-aged groups.
\begin{table}[H]
\begin{center}
\caption{Empirical Results in 2019 March CPS Data}\label{tbl cps}
\begin{tabular}{lccccccccc}
\hline\hline
\multicolumn{10}{c}{Panel A: 95\% confidence intervals in
race-and-gender-based subsamples} \\
race-gender & $n$ & cen\# & cen\% & \multicolumn{2}{c}{tail index} &
\multicolumn{2}{c}{Q(0.99)} & \multicolumn{2}{c}{Q(0.999)} \\ \hline
full sample & \multicolumn{1}{r}{115424} & 672 & 0.58 & (0.41 & 0.52) &
(24.24 & 25.32) & (62.63 & 75.61) \\
all males & \multicolumn{1}{r}{55553} & 491 & 0.88 & (0.35 & 0.53) & (28.64
& 30.68) & (67.71 & 92.58) \\
all females & \multicolumn{1}{r}{59871} & 181 & 0.30 & (0.42 & 0.55) & (18.02
& 19.03) & (46.34 & 58.01) \\
white males & \multicolumn{1}{r}{43371} & 419 & 0.97 & (0.88 & 1.00) & (29.83
& 33.56) & (145.2 & 279.0) \\
white females & \multicolumn{1}{r}{45424} & 141 & 0.31 & (0.42 & 0.57) &
(18.09 & 19.28) & (46.56 & 60.66) \\
Asian males & \multicolumn{1}{r}{3676} & 50 & 1.36 & (0.00 & 0.45) & (30.73
& 37.20) & (48.03 & 86.76) \\
Asian females & \multicolumn{1}{r}{4099} & 22 & 0.54 & (0.35 & 0.94) & (20.77
& 27.16) & (45.12 & 145.0) \\
Hispanic males & \multicolumn{1}{r}{44420} & 445 & 1.00 & (0.71 & 0.95) &
(31.10 & 34.75) & (119.3 & 214.0) \\
Hispanic females & \multicolumn{1}{r}{48192} & 155 & 0.32 & (0.47 & 0.62) &
(18.70 & 19.90) & (50.17 & 66.13) \\
black males & \multicolumn{1}{r}{6144} & 12 & 0.20 & (0.22 & 0.58) & (15.92
& 18.22) & (28.73 & 49.39) \\
black females & \multicolumn{1}{r}{7827} & 9 & 0.16 & (0.16 & 0.44) & (13.90
& 15.64) & (25.25 & 37.18) \\ \hline
\multicolumn{10}{c}{Panel B: 95\% confidence intervals in age-based
subsamples} \\
age & $n$ & cen\# & cen\% & \multicolumn{2}{c}{tail index} &
\multicolumn{2}{c}{Q(0.99)} & \multicolumn{2}{c}{Q(0.999)} \\ \hline
18-30 & 27829 & 35 & 0.13 & (0.33 & 0.49) & (12.69 & 13.60) & (28.16 & 36.16)
\\
30-40 & 25213 & 158 & 0.63 & (0.28 & 0.50) & (24.12 & 26.36) & (52.24 &
75.73) \\
40-50 & 23419 & 213 & 0.91 & (0.83 & 1.00) & (28.89 & 33.82) & (119.7 &
297.2) \\
50-60 & 21767 & 196 & 0.90 & (0.52 & 0.83) & (28.19 & 32.21) & (78.05 &
154.2) \\
60-65 & 17196 & 70 & 0.41 & (0.17 & 0.41) & (20.95 & 23.13) & (41.86 & 59.31)
\\ \hline
\end{tabular}
\vspace{-4ex}
\end{center}
\begin{singlespacing}
\begin{footnotesize}
Note: Entries are the sample size ($n$), the number of censored observations
(cen\#), the censored fraction in percentage points (cen\%),\ 95\%
confidence intervals of the tail index and those of the 99\% and 99.9\%
quantiles measured in 10$^{4}$\ USD. The results are based on the variable
ERN\_VAL in the CPS dataset and equivalently the variable inclongj from the
IPUMS dataset. Data are available at https://cps.ipums.org/cps.
\end{footnotesize}
\end{singlespacing}
\end{table}
\subsection{Macroeconomic Disasters\label{sec disaster}}
This section studies the size distribution of macroeconomic disasters, which
is an important research topic in macroeconomics. \citeasnoun{Barro08} and \citeasnoun
{Barro11} construct and analyze the dataset that consists of annual GDP (and
consumption) growth rates in 36 countries from 1870 to 2005. The authors
sort these observations and define a macroeconomic disaster if the GDP
declines by more than 10\%. This leads to $k=157$ tail observations. Then
the authors fit these data to the (double) Pareto distribution to estimate
the Pareto exponent, which is the reciprocal of the\ tail index, and back
out the coefficient of the relative risk aversion by a theoretical model
(eq.2 in \citeasnoun{Barro11}).
However, the largest disasters tend to be missing because some governments
collapsed or were fighting wars (p.1581 in \citeasnoun{Barro11}). Ignoring these
missing data in the upper tail could lead to substantial bias, as we show in
the\ Monte Carlo simulations. We revisit this problem by applying our fixed-$
k$ method since $k$ is only moderate. Specifically, the most recent data
missing happens in four countries, which are Greece, Malaysia, the
Philippines, and Singapore during WWII. Therefore, we set $m=4$ and apply
the fixed-$k$ method to construct the 95\% confidence intervals for the tail
index $\xi $ and those for the coefficient of relative risk aversion by
solving eq.2 in \citeasnoun{Barro11}. For comparison, we also construct the
intervals based on Hill (1975)'s estimator and the bias-reduced estimator
(GI) proposed by \citeasnoun{Gabaix11}. Table \ref{empirical_index} presents the
result.
As shown in the table, the fixed-$k$ intervals contain substantially larger
values of the tail index than the other two methods that ignore the tail
censoring. This is coherent with our simulation results in Table \ref{tbl
index Hill}. By taking the reciprocal, the Pareto exponent is estimated to
be approximately 7 in \citeasnoun{Barro11} but less than 1 by the new method.
Therefore, taking the tail censoring into account leads to a substantially
heavier tail in the disaster size. Accordingly, the coefficient of risk
aversion is found to be around 0.75, which is significantly lower than 3 in
\citeasnoun{Barro11}. These results undermine their conclusion that "the (Hill)
estimate of the upper-tail exponent is likely to have only a small upward
bias due to missing extreme observations, which have to be few in number."
\begin{table}[H]
\begin{center}
\caption{Empirical Results in Macroeconomic Disasters}\label{empirical_index}
\vspace{+2ex}
\begin{tabular}{lcccccc}
\hline\hline
Method & \multicolumn{2}{c}{Hill} & \multicolumn{2}{c}{GI} &
\multicolumn{2}{c}{New} \\ \hline
& \multicolumn{6}{c}{Tail Index} \\
95\% CIs & (0.12 & 0.17) & (0.16 & 0.26) & (0.57 & 1.00) \\ \hline
& \multicolumn{6}{c}{Coefficient of Risk Aversion} \\
95\% CIs & (3.73 & 5.10) & (2.44 & 3.88) & (0.58 & 1.04) \\ \hline\hline
\end{tabular}
\vspace{-4ex}
\end{center}
\begin{singlespacing}
\begin{footnotesize}
Note: Entries are 95\% confidence intervals (CIs) of the tail index of the
disaster size distribution and the coefficient of risk aversion, based on
the Hill estimator (Hill), the bias-reduced estimator proposed by \citeasnoun
{Gabaix11} (GI), and the fixed-$k$ method by inverting (\ref{LRtest}). Data
are available at https://scholar.harvard.edu/barro/data\_sets.
\end{footnotesize}
\end{singlespacing}
\end{table}
\section{Concluding Remarks\label{sec conclusion}}
This paper develops a new approach to estimate and conduct inference about
tail features for censored data. The method can be viewed as a hybrid
approach that uses the maximum likelihood estimation when the tail sample
size is large and switches to a small sample modification otherwise. As
shown in Monte Carlo simulations, the new method has excellent small sample
performance.
This new approach is empirically relevant in broad areas studying tail
features (e.g., tail index and extreme quantiles). We illustrate this with
the March CPS data and the macroeconomic disaster data and find considerably
different results from the existing literature.
There are theoretical extensions and empirical applications of our method,
which we suppress in the current paper due to space limitations. We list a
few here. First, our method naturally applies to the no censoring case by
setting $\kappa =\infty $ in the MLE and $m=0$ in the fixed-$k$ method.
Besides, we can follow \citeasnoun{MuellerWang19} to construct the (quantile)
unbiased estimation of the tail features, which could perform better in
terms of mean absolute deviation and mean squared error, especially when $k$
is not large.
Second, many other tail features can be learned by our new method as long as
they can be expressed as functions of the tail index. For example, the
conditional tail expectation is another important risk measure in finance,
which is defined as the expectation conditional on being larger than some
high quantile, that is, $\mathbb{E}\left[ Y_{i}|Y_{i}>Q(1-p)\right] $. By
reparametrizing $p=h/n$ for some $h>0$ and using EV theory, we can obtain
that that $(\mathbb{E}\left[ Y_{i}|Y_{i}>Q(1-h/n)\right] -b_{n})/a_{n}\left.
\rightarrow \right. h^{-\xi }/(\xi (1-\xi ))-1/\xi $, which again entirely
depends on $\xi $ and $h$ (p.1336 in \citeasnoun{MuellerWang19}). Then we can
construct the fixed-$k$ intervals for this quantity in an analogous fashion
to (\ref{program}).
Finally, our method also allows from weak dependent data if some additional
regularity condition is satisfied. In particular, EV theory holds under weak
dependence, such as $\alpha $-mixing, as long as the largest order
statistics do not show up in a cluster. This is referred to as the
non-cluster condition. See, for example, \citeasnoun{Leadbetter83}, \citeasnoun{Obrien87}
, \citeasnoun{Mikosch00}, \citeasnoun{Chernozhukov05}, and \citeasnoun{Chernozhukov11}.