Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,473 characters · 8 sections · 0 citation commands
Estimation and Inference about Tail Features with Tail Censored Data
\setcounter{page}{1}
Tail risk and extreme events are important research topics in economics and finance. In many applications, the features of interest are tail properties such as tail index and extreme quantiles. Existing literature has extensively studied the case with fully observed datasets. In comparison, this article explores the case with tail censoring. We argue that it is important to take into account the censoring if the research interest is in the tail, even when the censoring fraction is small. We provide a new method to construct estimators and confidence intervals for tail features.
Suppose one has a random sample from some underlying distribution $F$, where the observations larger than some threshold $T$ are replaced with $T$ or simply unobserved. In principle, tail features cannot be even identified if they entirely depend on the right tail part of $F$ that is beyond $T$. However, we can back out the tail-related features by extrapolation under two assumptions. They are that (i) the tail of $F$ can be well approximated by some suitably chosen parametric distribution, and (ii) $T$ is sufficiently large so that only a small fraction of samples are censored. The first assumption has been thoroughly studied in the statistic literature and is satisfied by many commonly used distributions. The second assumption is also satisfied in many interesting empirical applications, which motivate this article.
Our first motivating example is the Current Population Survey (CPS) dataset, which has been the primary data source used for investigating the distributions of individual earnings and household income in the US. Featured studies using CPS data include \citeasnoun{Armour13} and \citeasnoun{Eika19}, among many others. In CPS, the individual earnings larger than some threshold $T$ are typically censored (also called topcoded) and replaced with $T$ for confidential reasons.\footnote{ The topcoding has constantly been changing. Description of the topcoding mechanism is available at https://cps.ipums.org/cps/topcodes_tables.} In 2019, the censoring threshold is 310000 USD, leading to an approximately 0.58% censoring fraction in the full sample of individuals between 18 and 70 years old. This quantity is also substantially different across the subsamples defined by race and gender but remains small, as seen in Table (ref). Using this dataset, we aim to estimate and construct confidence intervals for the tail index that measures the tail heaviness of the income distribution and the extreme quantiles.
In our second application, we examine the size distribution of macroeconomic disasters, which is investigated by \citeasnoun{Barro08} and \citeasnoun{Barro11}. In particular, \citeasnoun{Barro11} define a macroeconomic disaster if the annual Gross Domestic Product (GDP) (or consumption) declines by more than 10%. The authors collect the data in 36 countries from 1870 to 2005 and construct a sample of approximately 5000 observations. However, data are missing for four countries during WWII due to government collapse or fighting wars. These observations correspond to the end-of-world case, and hence \citeasnoun {Barro11} concern that they are the largest observations but censored. In this situation, the tail censoring fraction is about 0.1%, and the parameter of interest is the tail index and the coefficient of risk aversion. Note that the censoring threshold $T$ is unknown here.
One would think that a tiny censoring fraction makes nearly no effect if we ignore it. This is true if the object of interest lies in the mid-sample, such as the median. However, such ignorance could lead to substantial bias and size distortion if some tail features are of interest. To examine this, we conduct an extensive Monte Carlo study and find that even the 0.1% tail censoring could lead to poor finite sample performance in some commonly used methods, including, for example, the classical Hill (1975)'s estimator. \nocite{Hill75}
To accommodate the censoring, existing studies typically rely on some parametric assumption of the whole distribution (e.g., \citeasnoun{Aban06}, \citeasnoun {Jenkins10}, and \citeasnoun{Burkhauser10}). Then, tail features can be expressed as functions of the unknown parameters and estimated by the maximum likelihood estimator (MLE). However, the parametric assumption on the whole distribution may lead to a substantial misspecification error when the object of interest is in the tail. This is because tail features such as very large quantiles are typically on a large scale, and hence small misspecification can be considerably amplified. For example, the standard normal distribution and the Student-t distribution with 20 degrees of freedom share almost the same shape in the mid-sample but exhibit substantially different extreme quantiles. Such misspecification is documented by \citeasnoun{Brzezinski2013} in a large-scale simulation study.
Instead of modeling the whole underlying distribution $F$, we focus on the tail part only and approximate it with the generalized Pareto distribution (GPD). Such a Pareto-tail approximation holds for many commonly used distributions, including, for example, Student-t, F, Beta, and Gaussian distributions. See Chapter 1 of \citeasnoun{deHaan07} for an overview. Given this approximation, we first pick some tail cutoff $u$, such as some large empirical quantile, and treat the observations larger than $u$ (but still less than $T$) as stemming from the censored Pareto tail. Let $k$ denote the number of these effective tail observations. Then, we can fit them into the classical Tobit model and conduct the MLE of the unknown parameters. Under some mild regularity conditions, we show that the MLE is consistent and asymptotically normal as $k$ diverges, enabling the construction of confidence intervals. Then we can estimate extreme quantiles by expressing them as functions of the Pareto parameters. This is formally studied in Section (ref).
The proposed maximum likelihood method can be quite good in finite samples for some applications, but it is also easy to find examples where the asymptotic distributions provide poor approximations. A fundamental limitation of our MLE and many other existing approaches in studying tail features is that they require restrictive conditions for the choice of $k$ (and equivalently $u$). On the one hand, $k$ has to be sufficiently large to support enough observations stemming from the approximately Pareto tail for the consistency and the asymptotic Gaussianity. On the other hand, $k$ has to be sufficiently small relative to the whole sample size $n$ so that the tail Pareto approximation incurs a negligible bias. Such a delicate balance is technically reflected in the conditions that $k\rightarrow \infty $ and $ k/n\rightarrow 0$. As such, for some combinations of $n$ and $F$, it is hard to find the $k$ that leads to satisfactory inference. This situation is close in spirit to the bias-variance trade-off in choosing the bandwidth in the standard kernel regressions.
To alleviate the above issue, we propose a small sample modification of the MLE when $k$ is only moderate, say 100. In particular, we follow \citeasnoun {MuellerWang17} to consider $k$ as a fixed number and study the fixed-$k$ asymptotic embedding by resorting to Extreme Value (EV) theory. Instead of treating the tail observations as independent copies from the GPD, EV theory treats them as dependent random variables with a joint EV distribution. Such dependence is negligible when $k$ is large but plays an important role when $k$ is only moderate. Using the EV approximation, we propose new confidence intervals for the tail index and extreme quantiles, which have excellent coverage probabilities. This is studied in detail in Section (ref).
In summary, the method that we propose is a hybrid approach. We suggest using the MLE for estimation and inference when $k$ (and $n$) is large enough and switching to the fixed-$k$ intervals otherwise. We choose the switching cutoff to be $k\gtrless 250$ based on our Monte Carlo experiments in Section (ref).
Returning to the CPS application, $n$ is large enough in the full sample to support a large $k$, and hence the MLE is expected to perform well. However, in the Asian male subsample, $n$ is only 3676. Then choosing $k$ as a small fraction, say 5%, of $n$, leads to only 180 tail observations and triggers the switching. By using the new approach, we make several interesting empirical findings. First, the tail index is substantially different across genders and races, while the existing literature commonly focuses on the full sample and finds the tail index to be approximately 0.5. See \citeasnoun {TodaWang2019} and references therein. Second, extreme quantiles also considerably vary across genders and races. In particular, the 99.9% quantile of all males can be twice larger than that of the black male group. Third, the tail features are also substantially different across ages. The middle-aged groups exhibit heavier tails and larger extreme quantiles than the groups below 30 or above 60 years old.
In the macroeconomic disaster application, we use the new method to construct confidence intervals for the tail index and the coefficient of risk aversion. We find substantially different results from those in \citeasnoun {Barro11}. In particular, we obtain a significantly heavier tail in the disaster distribution. This further results in a smaller coefficient of relative risk aversion, approximately 0.75 instead of 3. Our Monte Carlo simulation statistically justifies such a vast difference.
The rest of the paper is organized as follows. Section (ref) develops the MLE, and Section (ref) provides the small sample modification. Section (ref) reports Monte Carlo simulations, and Section (ref) applies the new approach to the CPS and the macroeconomic disaster examples. Section (ref) concludes with some remarks. All proofs and computational details are collected in the Appendix.
Consider a random sample $\{Y_{i}\}_{i=1}^{n}$ generated from some cumulative distribution function (CDF) $F$. Due to censoring, the econometrician observes the pair $\left( Y_{i}^{0},D_{i}\right) ^{\intercal } $ such that
where $T$ denotes some constant censoring threshold and $\mathbf{1}\left[ \cdot \right] $ the indicator function. Without loss of generality, we focus on the right tail. Define $m=\dsum_{i=1}^{n}D_{i}$ as the number of censored observations. We assume the density of $Y_{i}$, denoted as $f(\cdot )$, is continuous and positive so that $\mathbb{P}\left( Y_{i}=T\right) =0$.
The model ((ref)) has spawned a vast literature about estimation and inference about the mid-sample features, such as median, non-extreme quantiles, and regression coefficients. See, for example, \citeasnoun{Powell1986}, \citeasnoun{Portnoy2003}, and \citeasnoun{HongTamer2003}. These mid-sample features are typically estimated at the root-$n$ rate. In contrast, the tail features are estimated at a much slower rate since only the largest observations are informative about the right tail.
Define
as the conditional CDF given that $Y_{i}$ is larger than some pre-specified tail cutoff $u$. We aim to approximate $F_{u}\left( y\right) $ by the generalized Pareto distribution, which is given by
with $y\in \mathbb{R} ^{+\text{ }}$if $\xi \geq 0$ and $y\in \left( 0,-\sigma /\xi \right) $ otherwise. Denote $y_{0}$ as the right end-point of the support of $Y_{i}$. It is well established in the statistic literature (e.g., \citeasnoun {BalkemadeHaan1974} and \citeasnoun{Pickands1975}) that the GPD is a good approximation of $F$ in the tail, in the sense that
for some scale $\sigma $ implicitly depending on $u$, if and only if $F$ is in the domain of attraction of one of the three limit laws. The parameter $ \xi $ is referred to as the tail index, which is uniquely determined by $F$ and characterizes its tail heaviness. See Chapter 1 of \citeasnoun{deHaan07} for an overview.
The tail approximation ((ref)) is a mild assumption as it is satisfied by many commonly used distributions. In particular, the positive $ \xi $ case covers distributions with a Pareto-type tail such as Pareto, Student-t, and F distributions.\footnote{ In the standard Pareto distribution with the CDF $\mathbb{P}(Y>y)\propto y^{-\alpha }$, the tail index $\xi $ equals $1/\alpha $. We focus on $\xi $ instead of $\alpha $ for notational ease.} The case with $\xi =0$ covers the distributions with finite moments of any order. Leading examples are normal and log-normal distributions. The situation with a negative $\xi $ includes the distributions with a finite right end-point. For expositional simplicity, we focus our discussion on the case with $\xi >0$, which covers the empirical applications with heavy tails.
In practice, we usually choose $u$ as some large order statistic of $Y_{i}$, say the 95% empirical quantile. We let $T=T_{n}$ and $u=u_{n}$ depend on the sample size $n$ and assume $T_{n}>u_{n}$ (otherwise there is no observation). Also, denote $k$ as the number of the observations between $ u_{n}$ and $T_{n}$ and $\{Y_{(1)}\geq Y_{(2)}\geq ,\ldots ,\geq Y_{(n)}\}$ the order statistics\footnote{ This is different from the conventional notation for order statistics, that is, $\{Y_{n:n}\geq Y_{n:n-1}\geq ,\ldots ,\geq Y_{n:1}\}$. We think this alternative is more intuitive in our setup, especially in Section (ref).} by descending sorting. Then effectively the available observations are the censored largest $m+k$ order statistics
where the largest $m$ order statistics, $\{Y_{(1)},\ldots ,Y_{(m)}\}$ are censored. Using ((ref)), we can write the conditional log-likelihood of the tail observations as
where $g\left( y;\xi ,\sigma \right) =\partial G\left( y;\xi ,\sigma \right) /\partial y$. Then the MLE of $\xi $ and $\sigma $ are constructed as
To derive the asymptotic properties of the MLE, we make the following assumptions. To simplify notations, we write $\alpha =1/\xi $ when convenient and define $L\left( y\right) =y^{\alpha }(1-F\left( y\right) )$.
Condition (ref) assumes a random sample generated with censoring. Condition (ref) is imposed by \citeasnoun{Hall82}, which states that the Pareto tail approximation involves a second-order bias of the order $ y^{-\beta }$. This is imposed to avoid technical complexity and can be relaxed with other weaker conditions (cf.\ \citeasnoun{Goldie87}). Consider the Student-t distribution with zero mean, unit variance, and $v$ degrees of freedom, for example. The CDF\ is given by
Then Condition (ref) holds with $\xi =1/v$ and $\beta =2$.
Condition (ref) assumes the censoring threshold is larger than the tail cutoff. Specifically, the censoring is asymptotically negligible if $ \kappa =\infty $, and leads to no tail observation if $\kappa =1$. Condition (ref) specifies the choice of the tail cutoff and equivalently the number of tail observations $k$. The very last assumption that $k=o\left( n^{2\beta /\left( \alpha +2\beta \right) }\right) $ satisfies $ k/n\rightarrow 0$ and is imposed for expositional simplicity. We can relax it into $kn^{-2\beta /\left( \alpha +2\beta \right) }\rightarrow \mu $ for some $\mu \in \mathbb{R} $ as in \citeasnoun{Smith87} and Chapter 4 of \citeasnoun{deHaan07}. Doing so leads to a non-zero mean in the asymptotic normal distribution, which further depends on $\mu $, $\beta $, and $\delta $. These nuisance parameters are hardly estimable in practice, and hence researchers typically choose a sufficiently small $k$ to retain the asymptotic zero mean. This is similar in spirit to the undersmoothing in standard kernel regressions.
Under these conditions, the following proposition establishes the asymptotic normality of the MLE.
The tail censoring complicates the asymptotic variance substantially as compared with the no censoring case (cf.\ \citeasnoun{Smith87}). In particular, when the censoring is asymptotically negligible ($\kappa =\infty $), the information matrix reduces to
Given the MLE of $\xi $ and $\sigma $, we can further estimate the extreme quantile $Q\left( 1-p\right) \equiv \inf \{y:1-p\leq F\left( y\right) \}$. We set $p=p_{n}\rightarrow 0$ to capture the extremeness. The estimator can be constructed by inverting ((ref)), that is,
where $d_{n}=\left( m+k\right) /\left( p_{n}n\right) $. To derive a non-trivial asymptotic result, we let $d_{0}\equiv \lim_{n\rightarrow \infty }d_{n}>0$ so that the target quantile is of the same or larger magnitude of $ u_{n}$ (otherwise it can be estimated by the corresponding empirical quantile). The following proposition derives the asymptotic distribution of $ \hat{Q}\left( 1-p_{n}\right) $.
Proposition (ref) establishes the asymptotic normality of the extreme quantile estimator. Then the confidence intervals for $\xi $ and $ Q\left( 1-p_{n}\right) $ can be constructed by plugging in the estimators for the asymptotic variance.
The results in Section 2 suggest that the asymptotic normal approximation can be used for inference about the tail features as $k$ goes to $\infty $. In practice, however, the choice of the tail sample size $k$ is widely accepted as a challenging question even without censoring. This is because a good selection of $k$ has to balance the tail approximation bias and the variance delicately. Ultimately, the underlying distribution has to be reasonably close to the Pareto distribution in the tail to guarantee a satisfactory finite sample performance.
The asymptotic approximation can be quite accurate for some cases, but it is also easy to find examples where the limiting normal distribution provides a poor approximation. Consider the example that $F$ is a mixture of the standard normal distribution\ with probability 0.8 and some Pareto distribution with probability 0.2. Such a mixture structure implies that only the very few largest observations are informative about the true\ tail. In this case, choosing a large $k$ means including too many contaminating observations from the mid-sample, while choosing a small $k$ invalidates the asymptotic Gaussianity. In\ principle, there is no such a procedure that consistently justifies whether a given $k$ is appropriate when $F$ is entirely unknown. See Theorem 5.1 in M\"{u}ller and Wang (2017) for a discussion on the non-censored case.
Therefore, in this article, we do not focus on the choice of $k$ but instead treat it as given.\ In some cases, $k$ is determined by some economic theory or empirical guidance. For example, in the macroeconomic disaster application, the economic definition of disasters for more than 10% of GDP decline yields the choice of $k$. In other cases, we may employ some data-driven algorithms that balance the Pareto approximation bias and the variance. See, for example, \citeasnoun{Hall82}, \citeasnoun{Drees01}, and \citeasnoun {Clauset09}.
When $k$ and $n/k$ are both sufficiently large, we would expect the MLE in ( (ref)) based on the increasing-$k$ asymptotics to work well. Nevertheless, $k$ is only moderate in some situations, including our macroeconomic disaster application and the Asian male subsample in CPS. This causes a small sample issue that the asymptotic Gaussianity is questionable. To find a better alternative, we resort to the asymptotic embedding that requires $n$ diverges, but $k$ remains a fixed constant. Under this fixed-$k$ asymptotic framework, the consistent estimation of the tail index and the extreme quantiles are out of the question since the tail sample size is fixed. Fortunately, inference about these tail features is still implementable, as we discuss in this section.
We first study the tail index $\xi $. EV theory (the Fisher--Tippett--Gnedenko theorem) suggests that when the underlying distribution is within the maximum domain of attraction (e.g., Chapter 1 of \citeasnoun{deHaan07}), the sample maximum is asymptotically distributed as the EV distribution, which is parametric and entirely characterized by $\xi $. Specifically, our Condition (ref) is sufficient for the maximum domain of attraction assumption. Then, EV theory implies that there exist sequences of constants $a_{n}$ and $b_{n}$ such that, up to some location and scale normalization,
where the CDF of $X_{1}$ is given by
We can subsume $\sigma $ into $a_{n}$ so that $\sigma $ is always $1$ in ( (ref)). In additional to the sample maximum, EV theory also extends to the first $m+k$ order statistics such that if ((ref)) holds, then for any fixed $m$ and $k$,
The joint density of $\mathbf{(}X_{1},\ldots ,X_{m+k}\mathbf{)}^{\intercal }$ is given by $V_{\xi }(x_{m+k})\prod_{i=1}^{m+k}v_{\xi }(x_{i})/V_{\xi }(x_{i})$ on $x_{m+k}\leq x_{m+k-1}\leq \ldots \leq x_{1}$, where $v_{\xi }(x)=dV_{\xi }(x)/dx$.
Since the first $m$ elements are censored, the effective tail observations asymptotically reduce to
whose density is derived in the following proposition.
It is clear that the elements of $\mathbf{X}$ are dependent as captured by the term $\left( -\log V_{\xi }\left( x_{m+1}\right) \right) ^{m}V_{\xi }\left( x_{m+k}\right) $, which is negligible if $k$ is large but plays a vital role when $k$ is only moderate. This is the fundamental difference between the increasing-$k$ and the fixed-$k$ asymptotic embeddings, which are respectively used by the MLE and its small sample modification. From now on, we use bold letters to denote vectors.
If the constants $a_{n}$ and $b_{n}$ were known, the vector
is then approximately distributed as $\mathbf{X}$, and the limiting problem is reduced to the small sample parametric one: constructing a confidence interval based on one draw $\mathbf{X}$ whose density $f_{\mathbf{X}|\xi }$ is known up to $\xi $. However, $a_{n}$ and $b_{n}$ respectively correspond to the scale $\sigma $ and the tail location $u$. Therefore, they ultimately depend on $F$ and are challenging to estimate. Consider the standard Pareto distribution, for example. The Pareto exponent $\alpha $ is simply $1/\xi $. Then the fact that $a_{n}=n^{\xi }$ implies that a small estimation bias in $ \xi $ could be amplified by the $n$-power and lead to a poor inference.
To avoid the knowledge (and estimation) of $a_{n}$ and $b_{n}$, we consider the following self-normalized statistics:
It is easy to establish that $\mathbf{Y}^{\mathbf{\ast }}$ is maximal invariant with respect to the group of location and scale transformations (cf. Chapter 6 of \citeasnoun{Lehmann05}). In words, the statistic constructed as a function of $\mathbf{Y}^{\ast }$ remains unchanged if data are shifted and multiplied by any non-zero constant. This invariance is also intuitive since the tail shape should preserve no matter how data are linearly transformed.
The continuous mapping theorem and Proposition (ref) yield that
which is again invariant to location and scale transformation. By change of variables, the density of $\mathbf{X}^{\ast }$ is given by
where $e\left( \mathbf{x}^{\mathbf{\ast }},s\right) =\exp \left( -(1+1/\xi )\sum_{i=1}^{k}\log (1+\xi x_{i}^{\mathbf{\ast }}s)\right) $ and $x_{i}^{ \mathbf{\ast }}$ denotes the $i$th component of $\mathbf{x}^{\mathbf{\ast }}$ . See Appendix A.1 for more details.
The density ((ref)) can be used to conduct inference about $\xi $. In particular, consider the hypothesis testing problem
where $\Xi $ denotes the parameter space of $\xi $. We set $\Xi =(0,1)$ in the Monte Carlo simulations to cover the distributions with an unbounded support and a finite mean, which can be easily extended. Since the alternative hypothesis is composite, we follow \citeasnoun{Andrews94} and \citeasnoun {EMW15} to consider the weighted average alternative
where the weighting measure $W$ reflects the importance a researcher attaches to different alternative values of $\xi $. In practice, we set $ W(\cdot )$ to be the CDF of the standard uniform distribution for simplicity.
Given the density of $\mathbf{X}^{\ast }$, we construct the likelihood-ratio test as
where $\mathrm{cv}\left( \xi _{0},k,m\right) $ denotes the critical value that depends on the significance level, the null value $\xi _{0}$, the tail sample size $k$, and the number of the censored observations $m$. The critical value is obtained by simulation, and the test is implemented by replacing $\mathbf{X}^{\ast }$ with $\mathbf{Y}^{\mathbf{\ast }}$ in finite samples. The continuous mapping theorem and Proposition (ref) yield that $\mathbb{E}\left[ \varphi \left( \mathbf{Y}^{\mathbf{\ast }}\right) \right] \rightarrow \mathbb{E}\left[ \varphi \left( \mathbf{X}^{\ast }\right) \right] $, which equals the nominal level under the null hypothesis. The confidence interval is constructed by inverting the test. Note that this method does not require the knowledge of the censoring threshold $T$, which applies to cases such as the macroeconomic disaster.
Following \citeasnoun{MuellerWang17}, we can also construct the confidence intervals of the extreme quantiles under the fixed-$k$ asymptotics. To this end, we focus on the $Q(1-p_{n})$ quantile with $p_{n}=O(n^{-1})$, which captures the fact that the object of interest is of the same order of magnitude as the sample maximum. In particular, we consider $p_{n}=h/n$ for some fixed $h>0$. Then EV theory implies that $\left( Q\left( 1-h/n\right) -b_{n}\right) /a_{n}$ converges to the $e^{-h}$ quantile of $X_{1}$, denoted $q(\xi ,h)=\left( h^{-\xi }-1\right) /\xi $. Again, the research problem would become inference about $q(\xi ,h)$ based on the $k\times 1$ vector of observations $\mathbf{X}$ if $a_{n}$ and $b_{n}$ were known. Without loss of generality, we construct a confidence set $S(\mathbf{Y})\subset \mathbb{R} $ such that $\mathbb{P}\left( Q\left( 1-p_{n}\right) \left. \in \right. S( \mathbf{Y})\right) \geq 1-\mathrm{lv}$, at least as $n\rightarrow \infty $, where $\mathrm{lv}$ denotes the significance level.
To eliminate $a_{n}$ and $b_{n}$, we use the self-normalized vector $\mathbf{ Y}^{\ast }$ as in ((ref)). Besides, we also impose location and scale equivariance on our confidence interval $S$. Specifically, we impose that for any constants $a>0$ and $b$, our interval $S$ satisfies that $S(a\mathbf{ Y}+b)=aS(\mathbf{Y})+b$, where $aS(\mathbf{Y})+b=\{y:(y-b)/a\in S(\mathbf{Y} )\}$. Under this equivariance constraint, we can write
where the notation $\mathbb{P}_{\xi }$ (and $\mathbb{E}_{\xi }$ below) indicates that the randomness is entirely characterized by $\xi $ asymptotically. The asymptotic problem then is the construction of a location and scale equivariant $S$ that satisfies
since any $S$ that satisfies Proposition (ref) and the equivariance constraint also satisfies
This problem involves a single observation $\mathbf{X}\in \mathbb{R} ^{k}$ from a parametric distribution indexed only by the scalar parameter $ \xi \in \Xi $.
In principle, there could still be many solutions that satisfy the asymptotic size constraint. To obtain the optimal one, we consider the weighted average expected length criterion
where $W$ again denotes some weighting measure on $\Xi $, and $\func{lgth} (A)=\int \mathbf{1}[y\in A]dy$ for any Borel set $A\subset \mathbb{R} $.
To solve the program of minimizing ((ref)) subject to ((ref)) among all equivariant set estimators $S$, we introduce
and write $\mathbb{E}_{\xi }[\func{lgth}(S(\mathbf{X}))]\left. =\right. \mathbb{E}_{\xi }[(X_{m+1}-X_{m+k})\func{lgth}(S(\mathbf{X}^{\ast }))]\left. =\right. \mathbb{E}_{\xi }[\kappa _{\xi }(\mathbf{X}^{\ast })\func{lgth}(S( \mathbf{X}^{\ast }))]$ with $\kappa _{\xi }(\mathbf{X}^{\ast })=\mathbb{E} _{\xi }[X_{m+1}-X_{m+k}|\mathbf{X}^{\ast }]$. Thus, our problem becomes
There are two advantages to translate the asymptotic problem into ((ref)). First, ((ref)) does not require the knowledge of the censoring threshold $T$ but only the number of censored observations $m$. Second, ((ref)) only involves $S$ evaluated at $\mathbf{X}^{\ast }$ and hence $\mathbf{Y}^{\ast }$ in practice. This means the knowledge of $ a_{n}$ and $b_{n}$ is asymptotically unnecessary as long as $n$ is sufficiently larger than $m+k$. Note that any solution to ((ref)) also provides the form of $S$, that is, $S(\mathbf{X})=(X_{m+1}-X_{m+k})S( \mathbf{X}^{\ast })+X_{m+k}$. So once $S(\cdot )$ is determined, the confidence interval can be constructed in practice by plugging in
To make further progress in solving ((ref)), we write the problem in the following Lagrangian form:
where the non-negative measure $\Lambda $ denotes the Lagrangian weights that guarantee the asymptotic size constraint. By writing the expectations above as integrals over the densities $f_{\mathbf{X}^{\ast }\mathbf{|}\xi }$ of $\mathbf{X}^{\ast }$ and $f_{Y^{\ast }(\xi ),\mathbf{X}^{\ast }|\xi }$ of $(Y^{\ast }(\xi ),\mathbf{X}^{\ast })$, the solution of the above problem is given by
The integrals can be numerically calculated by Gaussian quadrature, and then the only remaining challenge is to find some suitable Lagrangian weights $ \Lambda $. We solve this challenge by the numerical approach developed in \citeasnoun{EMW15}. The MATLAB program and the weights $\Lambda $ are available at the author's website. Note that $\Lambda $ only needs to be computed once by the author instead of empirical researchers. Then the most time-consuming part in solving the program ((ref)) is the numerical integration, which costs only a few seconds in a modern PC. Further details are provided in Appendix A.1.
This section examines the finite sample performance of the proposed method and compares it with several popular existing methods. We generate random samples from four commonly used distributions: the generalized Pareto distribution with $\xi =$ $0.5$ and $\sigma =1$ (GPD), the absolute value of the Student-t distribution with 2 degrees of freedom (\TEXTsymbol{\vert}t(2) \TEXTsymbol{\vert}), the F distribution with parameters 4 and 4 (F(4,4)), and the double Pareto-lognormal distribution (dPlN), that is,
where $Z_{1}$, $Z_{2}$, $Z_{3}$ are independent and $Z_{1}\sim N(0,1)$, and $ Z_{2},Z_{3}\sim Exp(1)$. For parameter values, we set $c_{1}=0$, $c_{2}=0.5$ , $\xi =0.5$, and $c_{3}=1$, which are typical values for income data as documented in \citeasnoun{Toda12}. In particular, the dPlN distribution is the product of independent double Pareto and lognormal variables. It has been documented to fit well to size distributions of economic variables including income (\citeasnoun{Reed03}), city size (\citeasnoun{Giesen10}), and consumption (\citeasnoun {Toda17}). In all four DGP's, the true value of the tail index is 0.5. Regarding the tail censoring, we set the censoring threshold $T$ as the 99% and 99.9% quantiles of the underlying distributions, implying that the censored probability (cen_p) is either 1% or 0.1%.
We first consider some widely used estimators in empirical studies. Due to space limitations, we only report the results of Hill (1975)'s estimator \nocite{Hill75} and the bias-corrected estimator proposed by \citeasnoun{Gabaix11} (denoted GI). The confidence intervals are based on their asymptotic normality and the plug-in estimators of their asymptotic variances. The sample size $n$ is 1000, 2000, and 5000, and $k$ is set as $\left[ 0.05n \right] $ for both methods, where $[A]$ denotes the closest integer of $A$. All results are based on 1000 simulations.
Table (ref) depicts the mean biases and the coverage probabilities of these two methods. Several key findings can be summarized as follows. First, both the Hill and the GI estimators suffer from severe biases, and the confidence intervals based on them exhibit substantial undercoverage. This holds even if the censoring probability is only 0.1%. Second, ignoring the upper tail censoring tends to underestimate the tail index, which implies a misleadingly thin tail. This is seen in Section (ref) when we study the macroeconomic disasters. Finally, unreported results show that other methods reviewed in Chapter 3 of \citeasnoun {deHaan07} also suffer from substantial undercoverage. Therefore, it is crucial to take the censoring into account, even if the censoring probability is tiny.
Now we implement the new method proposed in Sections (ref) and (ref). Table (ref) depicts the coverage and length of the 95% maximum likelihood confidence intervals (denoted ml) based on Proposition (ref) and those of the fixed-$k$ intervals (denoted fk) by inverting ((ref)). Several interesting findings can be made as follows. First, the maximum likelihood confidence intervals are substantially longer than the fixed-$k$ ones when the sample size is not large. Besides, the coverage probability is smaller than the nominal level when the censoring is at the 99.9% quantile. This is because the asymptotic normality cannot perform well when $k$ is not large. Second, in comparison, the fixed-$k$ ones always deliver the nominal size with shorter length, especially when the sample size is not large. Finally, when $n$ reaches 5000 (and $k$ reaches 250), the maximum likelihood intervals are comparable with the fixed-$k$ ones. Hence a simple rule-of-thumb choice of the switching cutoff is $k\lessgtr 250$, provided $n$ is sufficiently large.
Tables (ref) depicts the coverage probabilities and lengths of the confidence intervals of the 99% quantile, using either the maximum likelihood method as in Proposition (ref)\ or the fixed-$k$ method ((ref)). Both methods deliver satisfactory size and length properties, although the maximum likelihood intervals suffer from slight undercoverage. However, as we target the more extreme 99.9% quantile as in Table (ref), such undercoverage is substantial when $k$ is less than 250. In contrast, the fixed-$k$ ones always perform excellently. These results reinforce our switching cutoff at $k=250$.
Many applications in economics and finance involve estimation and inference of tail features with censored data. In this section, we apply the proposed method to the two datasets we discussed earlier in this paper. Our empirical analysis highlights the potential of our approach.
Our first application is about the tail features of the individual earnings distribution. Following the convention, we use the variable ERN_VAL in the March CPS dataset and drop the individuals that are younger than 18 or older than 70 years old. This yields 115,424 observations in the 2019 sample. The censoring threshold is 310000 USD, which leads to a 0.58% censoring fraction in the full sample and various censoring fractions in different subsamples. The first several columns in Table (ref) present the sample sizes ($n$) and the numbers (cen\#) and the fractions (cen%) of the censored observations, respectively. We use the previously introduced method to construct the 95% confidence intervals of the tail index and the 99% and 99.9% quantiles. Specifically, we follow the simulation study to use the maximum likelihood confidence intervals developed in Section (ref) when $k$ is larger than 250 and switch to the fixed-$k$ confidence intervals ((ref)) otherwise. The last six columns in Table (ref) present the results with $k=\left[ 0.05n\right] $. The results based on other choices are similar and reported in Appendix A.3.
Several interesting findings can be summarized as follows. First, in Panel A, the tail index is around 0.5 in the full sample, as commonly found in the existing literature. But it is substantially different across subsamples. Second, the tail also exhibits substantial heterogeneity across genders. In particular, the male sample has significantly higher quantiles than the female at both the 99% and 99.9% levels. Third, this difference also exists across races. In particular, the 99.9% quantile of all males is at least twice larger than that of the black males. All such heterogeneity provides new evidence for potential racial and gender discrimination. Finally, Panel B depicts the heterogeneity across ages, with substantially heavier tails showing up in the middle-aged groups.
This section studies the size distribution of macroeconomic disasters, which is an important research topic in macroeconomics. \citeasnoun{Barro08} and \citeasnoun {Barro11} construct and analyze the dataset that consists of annual GDP (and consumption) growth rates in 36 countries from 1870 to 2005. The authors sort these observations and define a macroeconomic disaster if the GDP declines by more than 10%. This leads to $k=157$ tail observations. Then the authors fit these data to the (double) Pareto distribution to estimate the Pareto exponent, which is the reciprocal of the\ tail index, and back out the coefficient of the relative risk aversion by a theoretical model (eq.2 in \citeasnoun{Barro11}).
However, the largest disasters tend to be missing because some governments collapsed or were fighting wars (p.1581 in \citeasnoun{Barro11}). Ignoring these missing data in the upper tail could lead to substantial bias, as we show in the\ Monte Carlo simulations. We revisit this problem by applying our fixed-$ k$ method since $k$ is only moderate. Specifically, the most recent data missing happens in four countries, which are Greece, Malaysia, the Philippines, and Singapore during WWII. Therefore, we set $m=4$ and apply the fixed-$k$ method to construct the 95% confidence intervals for the tail index $\xi $ and those for the coefficient of relative risk aversion by solving eq.2 in \citeasnoun{Barro11}. For comparison, we also construct the intervals based on Hill (1975)'s estimator and the bias-reduced estimator (GI) proposed by \citeasnoun{Gabaix11}. Table (ref) presents the result.
As shown in the table, the fixed-$k$ intervals contain substantially larger values of the tail index than the other two methods that ignore the tail censoring. This is coherent with our simulation results in Table (ref). By taking the reciprocal, the Pareto exponent is estimated to be approximately 7 in \citeasnoun{Barro11} but less than 1 by the new method. Therefore, taking the tail censoring into account leads to a substantially heavier tail in the disaster size. Accordingly, the coefficient of risk aversion is found to be around 0.75, which is significantly lower than 3 in \citeasnoun{Barro11}. These results undermine their conclusion that "the (Hill) estimate of the upper-tail exponent is likely to have only a small upward bias due to missing extreme observations, which have to be few in number."
This paper develops a new approach to estimate and conduct inference about tail features for censored data. The method can be viewed as a hybrid approach that uses the maximum likelihood estimation when the tail sample size is large and switches to a small sample modification otherwise. As shown in Monte Carlo simulations, the new method has excellent small sample performance.
This new approach is empirically relevant in broad areas studying tail features (e.g., tail index and extreme quantiles). We illustrate this with the March CPS data and the macroeconomic disaster data and find considerably different results from the existing literature.
There are theoretical extensions and empirical applications of our method, which we suppress in the current paper due to space limitations. We list a few here. First, our method naturally applies to the no censoring case by setting $\kappa =\infty $ in the MLE and $m=0$ in the fixed-$k$ method. Besides, we can follow \citeasnoun{MuellerWang19} to construct the (quantile) unbiased estimation of the tail features, which could perform better in terms of mean absolute deviation and mean squared error, especially when $k$ is not large.
Second, many other tail features can be learned by our new method as long as they can be expressed as functions of the tail index. For example, the conditional tail expectation is another important risk measure in finance, which is defined as the expectation conditional on being larger than some high quantile, that is, $\mathbb{E}\left[ Y_{i}|Y_{i}>Q(1-p)\right] $. By reparametrizing $p=h/n$ for some $h>0$ and using EV theory, we can obtain that that $(\mathbb{E}\left[ Y_{i}|Y_{i}>Q(1-h/n)\right] -b_{n})/a_{n}\left. \rightarrow \right. h^{-\xi }/(\xi (1-\xi ))-1/\xi $, which again entirely depends on $\xi $ and $h$ (p.1336 in \citeasnoun{MuellerWang19}). Then we can construct the fixed-$k$ intervals for this quantity in an analogous fashion to ((ref)).
Finally, our method also allows from weak dependent data if some additional regularity condition is satisfied. In particular, EV theory holds under weak dependence, such as $\alpha $-mixing, as long as the largest order statistics do not show up in a cluster. This is referred to as the non-cluster condition. See, for example, \citeasnoun{Leadbetter83}, \citeasnoun{Obrien87} , \citeasnoun{Mikosch00}, \citeasnoun{Chernozhukov05}, and \citeasnoun{Chernozhukov11}.