Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
80,832 characters · 14 sections · 0 citation commands
Fixed-$k$ Inference for Conditional Extremal Quantiles
\setcounter{page}{1}
Tail risks and extreme events are important research topics in economics. In many applications with multivariate analysis, features of interest are conditional tail properties such as conditional extremal quantiles. This article provides a new method to construct confidence intervals for conditional extremal quantiles from a fixed number $k$ of nearest-neighbor tail observations. Advantages of the proposed method are three-fold: first, it is robust against flexible distributional assumptions unlike parametric methods; second, the procedure yields asymptotically valid confidence intervals for any fixed tuning parameter $k$ unlike existing kernel methods that rely on sequences of moving tuning parameters for asymptotically valid inference; and third, our confidence intervals enjoy a uniform coverage property over a set of data generating processes involving a set of values of the tail index. In the existing literature, methods of inference about conditional quantiles concern about middle quantiles, e.g., \citeasnoun{Qu2015} -- also see \citeasnoun{Qu2019} -- based on local quantile estimators of \citeasnoun{Fan1994} and \citeasnoun{Yu1998}. We aim to complement this existing literature by proposing a method of inference about conditional extremal quantiles.
Compared with unconditional tail features, the conditional tail counterparts are more difficult to study. This is because conditional tails depend on both marginal distributions and their joint behavior. Although marginal distributions can be generally assumed to be approximately Pareto near the tails,\footnote{ This statement follows from the Pickands-Balkema-de Haan Theorem (\citeasnoun {BalkemadeHaan1974} and \citeasnoun{Pickands1975}). See \citeasnoun{deHaan07} for an overview.} joint distributions cannot be generally assumed to be approximated by a fully parametric joint distribution and thus are harder to study given very limited tail observations. To model a covariate-dependent yet tractable tails, the seminal paper by \citeasnoun{Chernozhukov05} extends the quantile regression (QR) estimator of \citeasnoun{Koenker78} to tails, and proposes a method called the extremal quantile regression (EQR). \citeasnoun{Chernozhukov11} further investigate the EQR to construct confidence intervals (CIs) based on subsampling.
The EQR approach is based on the assumption that the conditional extremal quantile can be well approximated by a parametric location-scale shift model:
for $\tau \rightarrow 1$, where $\mu \left( x\right) $ and $\sigma \left(x\right) $ are parametric functions that capture the location and scale, respectively. The element $(1-\tau )^{-\xi }$ can be treated as the quantile function of a standard Pareto distribution, that is, $\mathbb{P}\left( Y>y\right) \sim y^{-1/\xi }$ where $1/\xi $ is the Pareto exponent and $\xi $ is the tail index. This single parameter captures the tail shape in the way that a larger $\xi $ implies a heavier tail. The assumption of model ((ref)) simplifies the conditional tail distribution so that the covariate $X$ only affects the location and scale, but not the shape.\footnote{\citeasnoun{WangLi13} formally establish that the location-shift model assumption is equivalent to assuming $\xi $ remains constant across $x$.} This is satisfied if $X$ and $Y$ are jointly normal but violated by many other joint distributions. Unlike mid-sample features, misspecification bias could be substantial in studying tail ones.\footnote{With this said, we remark that the existing literature suggests a couple of ways in which one can rationalize a possibly misspecified quantile regression. \citeasnoun{Angrist06} show that the parametric linear quantile regression function minimizes a weighted distance to the true nonparametric quantile regression function. \citeasnoun{Kato17} show that the linear quantile regression parameter is a weighted average of the slopes of the true nonparametric quantile regression function.} In this paper, we consider a wider class of flexible joint distribution models using a repeated cross-sectional or panel data structure.
There are a number of reasons for which we want to study conditional tail features, such as conditional extremal quantiles, under flexible joint distribution models. First, conditional value-at-risk (VaR) is a risk measure commonly used in financial management, insurance, and actuarial science. Estimation and inference are studied by \citeasnoun{Chernozhukov01} and \citeasnoun{Engle04}, among others. \citeasnoun{Adrian16} propose a new measure for systemic risk, $\Delta $-CoVar, defined as the difference between two conditional VaRs. The tail shape governs the third-and higher-order moments of the portfolio return, which typically depend on other economic factors, e.g., business cycles. As this is excluded by the location-scale model ((ref)), it is preferred to accommodate a larger class of joint distributions. Second, \citeasnoun{Kelly14} find that extreme event risk affects asset pricing in the U.S. stock market. The shape parameter measures tail risk and varies with other stock characteristics such as stock size. Third, macroeconomists are interested in analyzing lower tails of the conditional distributions of GDP growth rate given financial conditions in the recent growth-at-risk literature -- see \citeasnoun{Adrian19} for example. Fourth, top wealth inequality is an active research question in macro-finance literature (see, for example, \citeasnoun{Piketty03}, \citeasnoun{Gabaix16}, and \citeasnoun{Jones18}). The tail of the wealth distribution is well documented to follow Pareto, and the exponent is in general a function of fundamentals in general equilibrium models. For example, \citeasnoun{Beare17} derive a formula for the Pareto exponent and comparative statics results, and \citeasnoun{Toda19} applies that formula in a general equilibrium context. Finally, investigating factors of infants' birth weights, such as mother's demographic characteristics and maternal behaviors, is an important question in health economics (e.g., \citeasnoun{Abrevaya01} and \citeasnoun{Koenker01}). The lower tails of the conditional distribution are especially of interest for their critical health consequences -- see \citeasnoun{Chernozhukov11}. Other economic issues about conditional tail features can be found in the comprehensive review by \citeasnoun{Chernozhukov17}.
The existing literature suggests alternative approaches besides those based on the parametric location-scale specification ((ref)). To our best knowledge, they all focus on estimation, as opposed to inference, and can be roughly categorized into two classes. The first class maintains some parametric form but relaxes the location-shift model to allow for some nonlinearity. \citeasnoun{WangTsai09} assume that $\xi (x)$ equals to $ \exp(x^{\intercal }\theta_{0}) $ for some unknown parameter $\theta _{0}$. \citeasnoun{WangLi13} assume that the Box-Cox transformed $Y$ has linear conditional quantiles in $X$. The second class is fully nonparametric and constructs some local smooth estimators, including, for example, \citeasnoun {Beirlant04}, \citeasnoun{Gardes10}, \citeasnoun{Gardes12}, \citeasnoun{Daouia13}, and \citeasnoun {Martins-Filho18}.
In this article, we focus on statistical inference rather than estimation, and provide confidence intervals (CIs) of a conditional extremal quantile that have preferred coverage and length properties. Our proposed method applies to both repeated cross-sectional data and panel data. The main idea is very intuitive. Consider the case of using panel data of $(Y,X)$ to fix ideas, and suppose that one is interested in the conditional extremal quantile of $Y $ given $X=x_{0}$, denoted by $Q_{Y|X=x_{0}}(\tau)$. If, for every individual, there exists some time period in which $X$ takes the value $x_{0}$, then we can simply collect the associated $Y$'s and form a cross-sectional sample from $F_{Y|X=x_{0}}$. Since this is infeasible especially when $X$ is continuous, we instead collect from each individual's time series the induced $Y$ associated with $X$ that is the nearest neighbor (NN) of $x_{0}$ . These induced $Y$'s are now approximately stemming from $ F_{Y|X=x_{0}}$, and the large (respectively, small) order statistics from them can be used for inference about the upper (respectively, lower) conditional extremal quantile $Q_{Y|X=x_{0}}(\tau)$. For multi-dimensional covariates, this is done by defining the NN measured by a certain choice of metric, such as the one induced by the Euclidean norm. If a linear regression model is appropriate, then the NN can also be defined using the linear index.
The above approximation approach is formalized by establishing a new extreme value (EV) theory. The theory is based on the large-$n$ and large-$T$ asymptotics, where $n$ and $T$ denote the sample sizes in cross-sectional and time-series dimensions, respectively. A large $T$ guarantees that the NN is close enough to the query point $x_{0}$, while a large $n$ provides enough observations from a more accurate tail sample. Given the new EV theory, we apply it to construct new confidence intervals for the conditional extremal quantiles.
Our proposed approach only requires some smoothness condition on the joint distribution and hence enjoys more robustness against functional form specification than existing methods. A natural question is how much efficiency we lose by using only one out of $T$ observations in each time series. It turns out that if the tail shape depends on the covariate highly nonlinearly,\footnote{ See Section (ref) for concrete numerical settings which this qualitative phrase stands for.} then our proposed NN method dominates existing methods in both coverage and length when $T$ is only moderately large, say 50. When $T$ is very large, say 500, the new CIs also deliver comparable lengths to the kernel regression method with the optimal bandwidth -- see the Monte Carlo results ahead in Section (ref) for more details.
As a by-product of our main result, we also develop CIs for extremal quantiles of the coefficients in a random coefficient regression model. In particular, suppose that $Y_{it}$ and $X_{it}$ are generated from the model $Y_{it}=\alpha _{i}+X_{it}^{\intercal }\beta _{i}+u_{it}$, where $(\alpha_{i},\beta _{i}^{\intercal })^{\intercal}$ is a random vector drawn from some unknown distribution. We first construct the least squares estimators of $\alpha_{i} $ and $\beta _{i}$ using the $i$-th time series for all $i$ and collect the largest (smallest) order statistics from these estimates. We then show that the estimation error is negligible under the large $n$ and large $T$ framework, and hence the largest (smallest) order statistics among these estimates again satisfy the desired EV theory, which further supports the application of the fixed-$k$ CIs for extremal quantiles of $\alpha _{i}$ and $\beta _{i}$. This complements the existing literature focusing on the mid-sample properties of heterogeneous effects (e.g., \citeasnoun{Hsiao04} and \citeasnoun{Wooldridge05}).
Applying the proposed methods, we study the tail risk of extremely low birth weight conditional on mothers' behavioral and demographics characteristics. We find that signs of major effects are the same as those found in preceding studies based on parametric models. On the other hand, we find that some effects exhibit different magnitudes from those reported in the previous studies based on parametric models.
The rest of the paper is organized as follows. Section (ref) presents the main results of this paper. Section (ref) presents Monte Carlo simulation studies. Section (ref) presents an empirical application. Section (ref) concludes the paper. All mathematical proofs and additional details are found in the appendix.
Notation Let $\overset{p}{\rightarrow }$ denote convergence in probability and $\overset{d}{\rightarrow }$ denote convergence in distribution as $n,T\rightarrow \infty $. Let $\mathbf{1}[A]$ denote the indicator function of a generic event $A$. Let $\left\vert \left\vert B\right\vert \right\vert $ denote the Euclidean norm of a vector or matrix $ B $, and let $C$ denote a generic constant whose value may change across lines. Let $B_{\delta }(x)$ denote a generic open ball centered at $x$ with radius $\delta $. When $X$ denotes a column vector and $c$ a scalar, the notation $X-c$ is understood as the vector $X-(c,\ldots ,c)^{\intercal }$.
We present the main result of this paper in this section. Let $X$ denote a $ \dim (X)\times 1$ vector of continuous random variables with uniformly positive joint PDF.\footnote{ While we focus on continuous random variables in our presentation, the method can also accommodate discrete random variables. Suppose that the covariate vector is written as $(X^{\prime },W^{\prime })^{\prime }$ where the subvector $X$ consists of continuous random variables and the subvector $ W$ consists of discrete random variables. Suppose that one is interested in the conditional extremal quantiles given $(X^{\prime },W^{\prime })^{\prime }= (x_0^{\prime },w_0^{\prime })^{\prime }$. We can then extract the subsample with $W=w_0$, and then apply our proposed method for the subsample. } The main object of interest is the conditional extremal quantile $ Q_{Y|X=x_{0}}\left( \tau \right) $ of $Y$ given $X = x_{0}$ for a pre-specified $x_{0}\in \mathbb{R} ^{\dim \left( X\right) }$ and $\tau \rightarrow 1$. For ease of exposition, we consider a balanced\footnote{ This is only for notational ease. The new approach is valid as long as $T$ is large for all $i$.} repeated cross-sectional or panel data set $\{Y_{it},X_{it}\}_{i=1:n,t=1:T}$ that is i.i.d. across $i$ and strictly stationary and weakly dependent across $t$ . Section (ref) presents an informal overview of our proposed method, and Section (ref) gives a formal theoretical justification. Finally, Section (ref) presents an extension of the main theoretical results to linear random coefficient models.
Our method consists of the following three steps. In the first step, we make use of the repeated cross-sectional or panel data structure by selecting a subsample induced by the distances of the covariates $\{X_{it}\}$ to the query point $x_{0}$. The subsample will be a $k\times 1$ random vector denoted as $\mathbf{Y}$. In the second step, by appealing to the extreme value theory, we show that after some normalization, $\mathbf{Y}$ converges in distribution to a well-defined limiting random variable, $\mathbf{V}$, whose distribution $f_{ \mathbf{V}}$ is parametric and uniquely determined by the tail features of $ F_{Y|X=x_{0}}$. In particular, $f_{\mathbf{V}}$ will be uniquely characterized by a scalar parameter $\xi $ that fully captures the tail heaviness of $F_{Y|X=x_{0}}$. Note that $\xi $ depends on the query point $ x_{0}$, which will be suppressed in our notations for simplicity when there is no confusion. Since $Q_{Y|X=x_{0}}\left( \tau \right) $ can also be uniquely expressed as a function of $\xi $ after suitably normalizing $\tau $ , the asymptotic problem becomes conceptually straightforward: constructing inference for a function of $\xi (x_{0})$ given a random draw $\mathbf{V}$. This type of problems is studied by \citeasnoun{EMW15}, who provide a generic argument to construct optimal inference when there exists a nuisance parameter under a null hypothesis. In the third and final step, we tailor their arguments to inference about $Q_{Y|X=x_{0}}\left( \tau \right) $ with $ \xi $ being the nuisance parameter.
The next three subsubsections introduce details of these three steps in order. Section (ref) then follows up by presenting regularity conditions and the main theoretical result of the paper that guarantees that our confidence interval constructed in the three-step procedure controls coverage asymptotically and uniformly over a set of data generating processes.
First, we select our subsample $\mathbf{Y}$ as follows.
The key idea for such a selection is heuristically illustrated by the following derivation - a formal argument is presented as a proof of Theorem (ref) in Appendix (ref). For each $i$, denote the NN among $ \{X_{it}\}_{t=1}^{T} $ to $x_{0}$ as $X_{i,(x_{0})}$. Then for any $y\in \mathbb{R} $,
where $\dot{x}_{i}$ lies between $X_{i,\left( x_{0}\right)}$ and $x_{0}$. The first equality is by the definition of conditional expectation. The second one follows under the strict stationarity. The third equality is valid if the conditional CDF is smooth. The last convergence holds if the NN converges to its query point $x_{0}$ and if the CDF is smooth with bounded derivatives.
The above derivation states that the collection of the induced order statistics $Y$ associated with the NN of $x_{0}$ can be treated as approximately stemming from the true conditional CDF $F_{Y|X=x_{0}}$ asymptotically. Thus the largest (cross-sectional) order statistics $\mathbf{ Y}$ can be treated as draws from the tail of $F_{Y|X=x_{0}}$.
To proceed with the second step, we need some regularity conditions about $ F_{Y|X=x_{0}}$. For readability, we introduce only one of the conditions here with the remaining of them discussed in Section (ref). Specifically, we assume that $F_{Y|X=x_{0}}$ is within the domain of attraction (DOA) of the extreme value distribution (denoted by $ F_{Y|X=x_{0}}\in \mathcal{D}(G_{\xi })$), in the sense that there exist sequences of constants $a_{n}$ and $b_{n}$ such that for every $v$,
where
This DOA condition is extensively studied in the statistics literature and is satisfied by many commonly used distributions, including, for example, Pareto, Student-t, F, Gaussian, and even uniform distributions. See Chapter 1 in \citeasnoun{deHaan07} for a complete review.
Under the DOA assumption and the cross-sectional i.i.d. assumption, we show that for any fixed $k$,
where the joint probability density function (PDF) of $\mathbf{V}$ is given by
for $v_{k}\leq v_{k-1}\leq \ldots \leq v_{1}$ with $g_{\xi }(v)=\partial G_{\xi }(v)/\partial v$, and zero otherwise.
Note that the constants $a_{n}$ and $b_{n}$ depend on $\xi $, and their estimates thus tend to exhibit large magnitudes of sensitivity to tail observations. For example, $a_{n}$ is $n^{\xi }$ if $F_{Y}$ is standard Pareto. Since a small estimation error in $\xi $ is amplified by the $n$ -power, inference relying on a good estimate of $\xi $ and the scale usually requires a large $k$ and a even larger sample size $n$. Besides, the $G_{\xi }(v_{k})$ term in ((ref)) suggests that the largest $k$ order statistics are not asymptotically independent, given any fixed $k$. \footnote{We also derived the estimation and inference method based on the increasing-$k$ asymptotics in a previous version of this article. Their performance is dominated by the fixed-$k$ approach, especially in case with only moderate sample sizes. Therefore, we present the fixed-$k$ result exclusively in this version for illustrational simplicity.}
We aim for a $(1-\alpha)$ CI for the conditional extremal quantile $ Q_{Y|X=x_{0}}(\tau )$ for $\tau $ close to 1. Specifically, we rewrite $\tau$ as $1-h/n$ for some $h>0$ following \citeasnoun{Chernozhukov05} and \citeasnoun {Chernozhukov11}. This setup means that the extremal quantile is of the same order of the sample maximum from $n$ random draws from the true conditional CDF $F_{Y|X=x_{0}}$.
Our objective is to construct a confidence set $S(\mathbf{Y})\subset \mathbb{ R} $ such that $\mathbb{P}(Q_{Y|X=x_{0}}(\tau )\in S(\mathbf{Y}))\geq 1-\alpha + o(1)$, as $n\rightarrow \infty $ and $T\rightarrow \infty $. Under the DOA assumption, calculations show that
Note that $q(\xi ,h)$ is the $\exp (h)$ quantile of $V_{1}$. Since it is shared by both $\mathbf{Y}$ and $Q_{Y|X=x_{0}}(1-h/n)$, we can impose location and scale equivariance on the CI to cancel them out. Specifically, we impose that for any constants $a>0$ and $b$, $S(a\mathbf{Y}+b)=aS(\mathbf{ Y})+b$, where $aS(\mathbf{Y})+b=\{y:(y-b)/a\in S(\mathbf{Y})\}$. Under this equivariance constraint, we can write
where we introduce the self-normalized statistics
and highlight with the subscript $\xi $ that the densities of $V^{q}$ and $ \mathbf{V}^{\ast }$ now depend solely on $\xi $. These can be computed by using ((ref)), ((ref)), and a change of variables.
Since $\xi $ is unknown, we impose the size constraint uniformly for all the values of $\xi $ that are empirically relevant. In this sense the fixed-$k$ approach is more robust against misspecification, especially when the sample size is not large enough to support a precise estimation of $\xi$. Let $\Xi \subset\mathbb{R}$ be the set of tail indices for which we impose the asymptotically correct coverage.\footnote{ We use\ $\Xi =[-1/2,1/2]$ for inference about conditional extremal quantiles in later applications, which covers all the distributions with finite variance. This range can be easily extended.} The asymptotic problem then is to construct a location and scale equivariant $S$ that satisfies
since any $S$ that satisfies ((ref)) also satisfies $\lim \inf_{n\rightarrow \infty ,T\rightarrow \infty }\mathbb{P} (Q_{Y|X=x_{0}}(1-h/n)\in S(\mathbf{Y}))\left. \geq \right. 1\left. -\right.\alpha $ by the continuous mapping theorem. Among all the solutions to this problem, we choose the optimal one that minimizes the weighted average expected length criterion
where $W$ is a positive measure with support on $\Xi $,\footnote{ We use the uniform weight in later sections.} and $\text{lgth}(A)=\int \mathbf{1}[y\in A]dy$ for any Borel set $A\subset \mathbb{R}$. The equivariance of $S$ further implies $\mathbb{E}_{\xi }[\text{lgth}(S(\mathbf{ V}))]\left. =\right. \mathbb{E}_{\xi }[(V_{1}-V_{k})\text{lgth}(S(\mathbf{V} ^{\ast }))]$. Thus the program of minimizing ((ref)) subject to ( (ref)) among all equivariant sets $S$ asymptotically becomes
where we abuse the notation of $\mathbb{E}_{\xi }$ and $\mathbb{P}_{\xi }$ to emphasize that the distributions of $V^{q}$ and $\mathbf{V}^{\ast }$ depend on $\xi $ (and further on $x_{0}$). Note that any solution to ((ref)) also provides the form of $S$, that is, $S(\mathbf{V} )=(V_{1}-V_{k})S(\mathbf{V}^{\ast })+V_{k}$. Once $S(\cdot )$ is determined, therefore, the confidence interval can be constructed in practice by plugging in
In solving ((ref)), we write the problem in the following Lagrangian form:
where the non-negative measure $\Lambda $ denotes the Lagrangian weights that guarantee the asymptotic coverage constraint. By defining $\kappa (\mathbf{V }^{\ast };\xi )=\mathbb{E}_{\xi }[V_{1}-V_{k}|\mathbf{V}^{\ast }]$ and writing the expectations above as integrals over the densities $f_{\mathbf{V} ^{\ast }}$ and $f_{V^{q},\mathbf{V}^{\ast }}$ of $\mathbf{V}^{\ast }$ and $ (V^{q},\mathbf{V}^{\ast })$, respectively, the solution of the above problem is given by
The integrals can be numerically computed by Gaussian quadrature. To find suitable Lagrangian weights $\Lambda $, we appeal to the generic algorithm in \citeasnoun{EMW15}, who provide a numerical method to construct $\Lambda $. We tailor their arguments to our conditional extreme tail inference problem and provide the corresponding MATLAB program on the author's website. The computation cost is only several seconds using a modern PC. Note that $ \Lambda $ only needs to constructed once by the author but not the empirical users. See Section (ref) for more details.
Other tail-related quantities, such as the conditional tail expectations, are also covered by our proposed method as long as they can be expressed as functions of the conditional tail index. We discuss such an extension in Section (ref). In the following subsection, we formally introduce all the regularity conditions and formalize the uniform coverage property of the confidence interval ((ref)).
Our asymptotic theory requires the following four conditions.
Condition 1.1 requires the data to be independent across $i$ and weakly dependent across $t$, which is plausibly satisfied by the Natality Vital Statistics that we use for our empirical application. In addition, this condition also requires the density of $X$ to be positive in an open neighborhood around the query point $x_{0}$. This condition is sufficient to establish that the NN converges to the query point $x_{0}$ almost surely at some power rate. To the best of our knowledge, this is the first result about the (almost sure and L$^{2}$) convergence rate of the NN under weak dependence, whose proof is non-trivial. We formalize this result as Lemma (ref) in Appendix A.1, which might be of independent research interest.\footnote{ The $\beta $-mixing condition allows the application of Berbee's lemma (\citeasnoun {Berbee87}) in establishing Lemma (ref). This is assumed to avoid technical complexity and can be relaxed to other forms of weak dependence.} Note that we intentionally choose only one NN to allow for weak dependence across $t$. If data are independent across both $i$ and $t$, then more than one NNs can be chosen to enlarge the effective sample. We leave this for future research.
This condition requires that the underlying conditional distribution is in the domain of attraction of the generalized EV distribution. This is a mild condition as it is satisfied by many commonly used joint distributions. In particular, it generalizes the conditional location-scale shift model ((ref)) by allowing $\mu (x),\sigma (x),$ and $\xi (x)$ to be all unknown (but smooth) functions of $x$. The case of negative $\xi (x_{0})$ is included only for comprehensiveness, since the $Y$ in most applications involving tail features has an unbounded support that entails a non-negative $\xi(x_{0})$. To illustrate the mildness of this condition, we discuss the following three examples. Our Condition 1.2 is satisfied in all three of them, but the location-scale model assumption ((ref)) is not.
Let $y_{0}$ denote the end-point of the conditional CDF, that is, $ y_{0}=Q_{Y|X=x_{0}}\left( 1\right) \leq \infty $. The next condition is a high level regularity assumption on the smoothness of the conditional tail.
Condition 1.3 requires that the derivatives of the conditional CDF and PDF are smooth and decay quickly. This is a mild condition again, which is satisfied by the above examples by straightforward calculation. For readability, we provide low-level primitive assumptions as sufficient conditions for Condition 1.3 and discuss them in Appendix (ref).
Condition 1.4 requires both $n$ and $T$ to be large. A large $n$ guarantees that the error due to the EV approximation is negligible, and a large $T$ controls the distance between the NN and the query point. The parameter $ \lambda $ can be any constant in the open unit interval, and hence $T$ can be much smaller than $n$.
Under the above conditions, we establish the asymptotically correct uniform coverage by confidence interval ((ref)) in the following theorem, which is the main result of this paper.
We conclude this subsection with a discussion of main properties, advantages and disadvantages of the our proposed fixed-$k$ confidence interval. First, since the confidence interval is based on a fixed number $k$ of tail observations, its length does not decrease in $n$. On the other hand, the length decreases in $k$. Second, unlike kernel regression approaches that require a sequence of moving tuning parameter (i.e., the bandwidth parameter tending to zero), our method only relies on a `fixed' tuning parameter which is $k$. While common data-driven choice rules for bandwidths are not theoretically compatible with inference for their failure to undersmooth estimates, our method based on any fixed $k$ of a researcher's choice guarantees asymptotically valid inference. Third, our fixed-$k$ approach allows for the confidence interval to have a uniform size control property over a set of data generating processes involving a set of values of the tail index, while the existing methods have not been shown to share this uniformity property. This property of our fixed-$k$ approach is useful because $\xi$ is practically unknown to researchers and thus the size control should be uniform for all the values of $\xi$ that are empirically relevant.
The proof strategy for our main result, namely Theorem (ref), and thus our proposed method of constructing confidence intervals apply to other contexts. Among others, inference for extremal quantiles of random coefficients in linear regression models is also possible with our proposed strategy. In this section, we study this class of models which have been widely used in empirical studies in economics.
Consider the model
where $(\alpha _{i},\beta _{i}^{\intercal })^{\intercal }$ denotes random coefficients and $u_{it}$ denotes an error term. This setup has been studied by numerous papers in the literature, and covers the classic panel linear regression model with fixed effects in which $\beta _{i}=\beta_{0}$ for all $ i$. As long as Conditions 1.1-1.4 are satisfied, the previously introduced methods naturally apply here for inference on the conditional extremal quantiles of $Y$. In addition, the model ((ref)) allows us to conduct inference on the unconditional tail features of the random coefficients, $\alpha _{i}$ and $\beta _{i}$. The remainder of this subsection illustrates a procedure to this end.
Let $(\hat{\alpha}_{i},\hat{\beta}_{i}^{\intercal })^{\intercal }$ be the OLS estimator by regressing $Y_{it}$ on $(1,X_{it}^{\intercal })^{\intercal}$ using the time series associated with the $i$-th individual. Collect $\{( \hat{\alpha}_{i},\hat{\beta}_{i}^{\intercal })^{\intercal }\}_{i=1}^{n}$ and sort each series of estimates in the descending order. We then define
that is, the largest $k$ order statistics of $\{\hat{\alpha}_{i}\}$, and
that is, the largest $k$ order statistics of the $j$-th coordinate of $\{ \hat{\beta}_{i}\}_{i=1}^{n}$, for each $j$. Without loss of generality, we focus on the first coordinate of $\beta _{i}$, and suppress the subscript $j$ from our notations for simplicity.
Now, we substitute $\mathbf{A}$ or $\mathbf{B}$ in $S(\cdot )$ as in ((ref)) to construct the confidence interval for extremal quantiles of $ \alpha _{i}$ or $\beta _{i}$. The following conditions are imposed for a theoretical guarantee of correct asymptotic coverage.
Condition 2.1 is similar to Condition 1.1. Since the objects of interest are the unconditional extremal quantiles of $\alpha _{i}$ and $\beta _{i}$, we do not need the NN condition on covariates. The dependence structure is left unspecified as long as it is sufficient for Condition 2.3. Condition 2.2 assumes that the distributions of $\alpha _{i}$ and $\beta _{i}$ are in the domains of attraction of $G_{\xi _{\alpha }}$ and $G_{\xi _{\beta }}$, respectively. Condition 2.3 requires that the estimator $ (\hat{\alpha}_{i},\hat{\beta}_{i}^\intercal)^\intercal$ is consistent for all $i$ and the moments of sample averages of $u_{it}$ and $X_{it}$ across $t$ are bounded. If the tail index is non-positive, then these bounds need to be stronger to accommodate the fact that $ a_{n}\rightarrow 0$.\footnote{ A straightforward calculation yields that normal distribution satisfies Condition 2.3, if $\sup_{i}\left\vert \left\vert (\hat{\alpha}_i,\hat{\beta}_{i}^\intercal)^\intercal-(\alpha_i,\beta _{i}^\intercal)^\intercal\right\vert \right\vert\left. =\right. O_{p}(T^{-\varepsilon })$ and $\sup_{i}\left\vert \bar{u}_{i}\right\vert \left. =\right. O_{p}(T^{-\varepsilon })$ for some $\varepsilon >0$ and if $ n/T\left. \rightarrow \right. \lambda $ for some $\lambda \left. \in \right.(0,\infty )$. This can be seen by $1/f_{\alpha }\left( Q_{\alpha }\left( 1\left.-\right. 1/n\right) \right) \left. \leq \right. O(\log (n))$ when $f_{\alpha}$ and $Q_{\alpha }$ are standard normal density and quantile functions, respectively (cf. Example 1.1.7 in \citeasnoun{deHaan07}).}
Under these conditions, the following corollary establishes the asymptotic coverage.
We conduct Monte Carlo experiments to examine the small sample performance of the new approach. In Section (ref), we first consider the simple panel data $\{Y_{it},X_{it}\}$ without any fixed effect. In Section (ref), we compare the efficiency of the new approach with the kernel estimator, which essentially uses more than one NNs. In Section (ref), we consider the linear random coefficient regression setup ((ref)).
We continue to consider the three examples in Section (ref) as the data generating processes (DGPs). In all experiments, generated data are i.i.d. across $i$, but are dependent across $t$. The dependence structure across $t$ is specified as follows.
We construct CIs for $Q_{Y|X=x_{0}}\left( 1-h/n\right) $ with $x_{0}=0$ and $ 1.65$ (the 50% and 95% quantiles of $X$, respectively) and $h=1$ and $5$. The sample sizes $n$ and $T$ are either 200 or 500, with smaller combinations exercised in later experiments.
We compare results across three approaches: (i) the fixed-$k$ approach (fixed-$k$) introduced in this paper, (ii) quantile regression (QR), and (iii) bootstrapping the empirical quantile (Boot). We produce the fixed-$k$ CI using $k=20$ in most cases if not otherwise noted. The space of $\xi $ is restricted to be $[-1/2,1/2]$. For the QR approach, we run a quantile regression of $ Y_{it}$ on $X_{it}$ and a constant at the $\tau $ quantile for each $i\,$. The conditional quantile is estimated at $\hat{\beta}_{0i}+x_{0}\hat{\beta} _{1i}$ where $\hat{\beta}_{0i}$ and $\hat{\beta}_{1i}$ are the coefficient estimates using the $i$-th individual's observations. The CI is defined the 2.5% and 97.5% quantiles of these $n$ estimates. The bootstrap CI is based on bootstrapping the empirical $\tau $ quantile in $\{Y_{i,[x_{0}]} \}_{i=1}^{n}$. The bootstrap size is 200.
Tables (ref)-(ref) depict the coverage probabilities (Cov) and the average lengths (Lgth) of the above three methods based on 500 simulation draws. The fixed-$k$ approach performs well in terms of both the coverage and length across all the specifications. Regarding the QR method, recall that the conditional quantile is a linear function of $X$ in the first DGP but not in the other two. Therefore, not surprisingly, the CIs based on QR perform well in the first DGP but deliver substantial undercoverage and longer length in the other two due to misspecification. The bootstrap approach is robust to misspecification but requires the asymptotic normal approximation, which performs well only in the mid-sample and does not near the tails. As such, the bootstrap intervals exhibit more undercoverage for $h=1$ than for $h=5$.
We conclude the current subsection with a remark about the choice of $k$. A larger $k$ leads to more tail observations and hence shorter confidence intervals, but is subject to a larger approximation bias due to including too many mid-sample data. This indicates that the choice of $k$ is difficult, especially when $n$ is only moderate. It is actually impossible to choose a uniformly best $k$ allowing the underlying CDF to be flexible (see Theorem 1 of \citeasnoun{MuellerWang17}). The CDFs in our Monte Carlo designs are all well behaved so that such a value of $k$ as large as 40% of the sample size performs well. This is seen in Table (ref), which reports the numbers for $k=20$ and $50$.
Our new approach takes only one NN in each time series, which raises the question of efficiency loss. We answer this by comparing our fixed-$k$ approach with the kernel smoothing method proposed by \citeasnoun{Gardes10}. In particular, we first pool the panel data into a cross-sectional sample. Suppose the object of interest is still $Q_{Y|X=x_{0}}(\tau )$. We follow \citeasnoun{Gardes10} to pick the bin $B_{b_{nT}}(x_{0})$ centered at $x_{0}$ with a bandwidth $b_{nT}$. Since there is no theoretical justification for the optimal choice of $b_{nT}$, we take the rule-of-thumb choice $c(nT)^{-1/5}$ with different values of the constant $c$. Now a certain choice of $b_{nT}$ leads to a certain collection of $Y^{\prime }s$ whose paired $X^{\prime }s$ are in the bin $B_{b_{nT}}(x_{0})$. Sort these induced $Y^{\prime }s$ in the descending order into $\{Y_{(1)}\geq Y_{(2)}\geq ...\geq Y_{(m)}\}$ where $m$ denotes the local sample size determined by the bandwidth. Such local sample size is approximately $nTb_{n}$ in the kernel smoothing (as opposed to $n$ in our new approach).
Given the induced $Y$'s, the conditional quantile is estimated as $\hat{Q} _{Y|X=x_{0}}(\tau ) {\footnotesize =} Y_{(\left\lfloor (1-\tau )m\right\rfloor )}$, that is, the $\left\lfloor (1-\tau )m\right\rfloor $-th largest order statistics in the induced $Y$'s where $\left\lfloor (1-\tau )m\right\rfloor $ denotes the integer part of $\left( 1-\tau \right) m$. \citeasnoun{Gardes10} show that, under $m(1-\tau )\rightarrow \infty $ and some other regularity conditions,
Then, the CI of $Q_{{\footnotesize Y|X=}x_{0}}(\tau )$ is constructed by the delta method and plugging in some consistent estimator of $\xi _{0}$. One choice which they propose is the Hill-type estimator
for some choice of $k<m$.
For comparisons, we implement our fixed-$k$ approach by using the panel data and the above kernel approach by pooling the data. In particular, we implement the conditional Pareto DGP in the previous experiment with $n=200$ and $T$ ranging from 50 to 500. For the fixed-$k$ CI, we set $k=50$. For the kernel method, we implement $c\in \{0.1,0.25,0.5,1,2\}$ and set $k$ (in the Hill-type index estimator ((ref))) as the largest integer less than or equal to $m/4$.
Table (ref) presents the coverages and the lengths of the fixed-$k$ and the kernel CIs. Several interesting observations can be made. First, the kernel approach is sensitive to the choice of the bandwidth. In particular, a correct coverage relies on a narrow window of the bandwidth choice. A larger choice can lead to a substantial undercoverage since the smoothing bias dominates quickly in the tail. Second, when $T$ is only moderately large (say 25 and 50), the fixed-$k$ CIs are much shorter than the kernel one and both of them have good coverage properties. This is because the fully nonparametric method ignores the domain-of-attraction information, which is utilized by our fixed-$k$ method. Third, when $T$ is very large, say 500, choosing only one NN does incur an efficiency loss as we compare the lengths between our fixed-$k$ and the kernel CIs. But such a loss is approximately in a factor of two or three instead of $T^{1/2}$. This means that a general covariate-dependent tail is very difficult to estimate in a fully nonparametric way.
In Table (ref), we consider a two-dimensional standard normal $X$ and generate $Y_{it}$ by $Y_{it}|X_{it}=x\sim \pm $Pa$(\xi (x))$ with $\xi (x_{1},x_{2})=x_{1}+x_{2}+0.5$. The kernel method is illustrated with $c\in \{0.5,1,2,4\}$. All the other parameter choices for both methods remain unchanged as those used for Table (ref). The results clearly suggest that our fixed-$k$ method together with the NN choice dominates the kernel method in terms of both the coverage probabilities and lengths. In particular, the kernel method suffers from the curse of dimensionality as the dimension of $X$ increases.
As a final remark of this subsection, we also implement the standard kernel weighted quantile regression method designed for the mid-sample quantiles (cf.\ Chapter 10 of \citeasnoun{LiRacine07}). Given a large $T$, the target $1-1/n$ conditional quantile is relatively in the mid-sample after pooling the panel data into a cross-sectional one, and hence the confidence interval based on asymptotic normality might work. However, unreported Monte Carlo simulations show that this method works only if $T$ is substantially larger than (e.g, five times as much as) $n$. In our experiments, it is strictly dominated by the method proposed by \citeasnoun{Gardes10}.
In this section, we consider the linear random coefficient model $ Y_{it}=\alpha _{i}+X_{it}\beta _{0}+u_{it}$, where the generated observations including the random coefficients are i.i.d. across $i$. For the time series dependence, we set $\alpha_{i}=T^{-1}\sum_{t=1}^{T}X_{it}$ and $X_{it}=\rho X_{it-1}+e_{it}$ with $e_{it}\sim ^{iid}\mathcal{N}\left( 0,(1-\rho ^{2})\right) $ and $X_{i0}\sim \mathcal{N}\left( 0,1\right) $. The conditional distributions of $u_{it}$ given $X_{it}=x$ are specified as follows.
We use the same set of the three approaches as in the Section (ref) to construct CIs for the conditional extremal quantile $ Q_{Y_{it}|X_{it}=x_{0}}\left( \tau \right) =Q_{\varepsilon _{it}|X_{it}=x_{0}}\left( \tau \right) +x_{0}\beta _{0}$, where $\varepsilon _{it}$ denotes $\alpha _{i}+u_{it}$. Specifically, our fixed-$k$ approach is conducted in two ways: with or without using the standard within least squares estimator of $\beta _{0}$. For the former (fixed-$k$ w.\thinspace LS), we first estimate $\beta _{0}$ using the standard within estimator $ \hat{\beta}$ and back out $\hat{\varepsilon}_{it}=Y_{it}-X_{it}\hat{\beta}$. We then implement the steps in Section (ref) to construct the CIs for the conditional quantiles of $\varepsilon _{it}$. The CIs for $ Q_{Y_{it}|X_{it}=x_{0}}\left( \tau \right) $ are obtained by adding back $ x_{0}\hat{\beta}$. For the one ignoring the linear regression structure (fixed-$k$ w/o LS), we directly use $(Y_{it},X_{it})^\intercal$, and apply Steps 1-3 in Section (ref).
Table (ref) presents the results for $n\in \{100,200\}$ and $ T\in \{25,50,200,500\}$. Several interesting observations can be made. First, the errors in the conditional t and conditional Pareto models do not have finite variances when $x_{0}$ is 0, and hence the LS estimator of $ \beta _{0}$ behaves poorly. This leads to a poor performance of the fixed-$k$ approach if the linear regression model is utilized. This problem can be solved by using the least absolute deviation (LAD) estimator as shown in unreported results. In comparison, the fixed-$k$ CIs without using the linear regression model always perform well given a large enough sample size. Second, the QR approach still suffers from undercoverage in all three specifications since the normal and the Student-t DGPs have nonlinear heteroskedasticity and the conditional Pareto DGP violates the constant tail shape condition. Finally, the bootstrap method performs poorly if the extremal quantiles under investigation are too far in the tail.
In Table (ref), we study the CIs for high quantiles of $ \alpha_{i}$ and $\beta _{i}$ with data generated from $Y_{it}= \alpha_{i}+X_{it}\beta _{i}+u_{it}$, where $\left( \alpha _{i},\beta _{i},X_{it},u_{it}\right) ^{\intercal }\sim ^{iid}\mathcal{N} \left(0,I_{4}\right) $. The i.i.d. condition is across both $i$ and $t$ in this setting. We first estimate $\alpha _{i}$ and $\beta _{i}$ by regressing $Y_{it}$ on $(1,X_{it})^{\intercal }$ with $T$ observations from individual $ i$. We then collect the estimates, $\hat{\alpha}_{i}$ and $\hat{\beta}_{i}$, for all $i$ and sort them in the descending order to apply each of the fixed- $k$, QR, and bootstrap methods. The QR estimator is simply the empirical quantile among the estimators for all $i$, whose asymptotic variance is estimated by the standard kernel density estimator with the rule-of-thumb bandwidth. The results suggest that the fixed-$k$ approach with NN dominates the other two in both coverage and length, especially when the sample size is only moderate.
In this section, we reconsider the extremely low birth weights and their relationships with mother's demographic characteristics and maternal behaviors, which addresses an important question in health economics. We use the detailed natality data published by the National Center for Health Statistics, which has been used by \citeasnoun{Abrevaya01}, \citeasnoun{Koenker01}, and \citeasnoun{Chernozhukov11} among many others. We follow these preceding studies, but our analysis is different from theirs in two aspects. First, these preceding studies use the cross-sectional data in one time period, while we collect the repeated cross-sectional samples from January 1989 to December 2002.\footnote{ We chose this specific period for two reasons. First, this period contains the time period of the cross-sectional data used by \citeasnoun{Abrevaya01}, \citeasnoun {Koenker01}, and \citeasnoun{Chernozhukov11}. Second, these periods maintain the identical variable definitions.} Second, the previous studies all made some parametric model assumptions, including either the linear projection model or the (extremal) quantile regression model. In contrast, our fixed-$k$ method is nonparametric, allowing for nonparametric joint distributions. Accordingly, some of our findings are different from those in the previous studies.
Details of our implementation are as follows. First, we follow the previously mentioned literature -- \citeasnoun{Abrevaya01} in particular -- to choose included covariates. Our dependent variable is the infant birth weight measured in kilograms, and the continuous covariates include mother's age and net weight gain (wtgain) during pregnancy. All the remaining covariates are discrete, and hence we consider the subsamples constructed from various combinations of the categorical variables. For comparison, we set a benchmark subsample in which the infant is a boy, the mother is white and married, has levels of education less than a high school degree, had her first prenatal visit in the first trimester (natal1), and did not smoke during pregnancy. Second, since the samples are repeated cross-sectional, it is more natural to switch the labeling of the indices $i$ and $t$, and first take the NN within each month. The query point is set at age equal to 27 and wtgain equal to 30, corresponding to their respective median values. The NN is then measured by the Euclidean norm after standardizing each of the two variables with mean zero and unit variance. Using the same notation as in Section (ref), we have $n=168$ and $T$ is at least 100 in every subsample. Thus, our fixed-$k$ asymptotic framework with a large $n$ and a large $T$ is suitable with this data. Third, we set $k=30$ based on our simulation results in the previous section and construct the 95% fixed-$k$ confidence intervals for the conditional $p$ -quantiles with $p$ ranging from 1% to 10%. Figure (ref) depicts these confidence intervals in the benchmark subsample and six alternative subsamples corresponding to one and only one of following scenarios: the mother has at least high school diploma; the infant is a girl; the mother is unmarried; the mother is black; the mother does not have prenatal visit during pregnancy; and the mother smokes 10 cigarettes per day on average.\footnote{ For the majority of the subsamples, the number of cigarettes as recorded in data takes only a few discrete values, including 0, 5, 10, and 20. Therefore, we treat it as a discrete random variable in our study.}
We can make the following observations in Figure (ref). First, the effects of changing the covariates are found to have a similar pattern as in the previous studies. In particular, compared with the benchmark subsample, the conditional quantile of infant birth weight decreases substantially if the mother is black, did not have a prenatal visit, and/or smoked during pregnancy. These results reconfirm the signs of the effects reported by the previous studies. Second, on the other hand, the magnitudes of these effects are larger than those documented in the previous studies. Specifically, \citeasnoun{Abrevaya01} finds that smoking ten cigarettes leads to approximately 200 fewer grams at the 10th percentile of the infant birth weight, compared with smoking no cigarettes. \citeasnoun{Chernozhukov11} finds the quantile regression coefficient associated with the number of cigarettes is nearly zero at the 1st percentile (their Figure 8). On the other hand, the last sub-figure in Figure (ref) suggests that the difference can be over 1000 grams at the 1st percentile if we compare the mid-value between the upper and lower bounds of the confidence intervals between these two subsamples. Finally, the effects on extremal birth weight quantiles induced by the demographic characteristics vary across levels of quantiles, instead of remaining fixed.
This paper develops a new nonparametric method of inference for conditional extremal quantiles using a fixed number $k$ of nearest-neighbor tail observations in repeated cross-sectional or panel data. There are three advantages of our proposed method. First, it is robust against flexible distributional assumptions unlike parametric methods. Second, the procedure yields asymptotically valid confidence intervals for any fixed tuning parameter $k$, unlike existing kernel methods that rely on a sequence of moving tuning parameters for asymptotically valid inference. Third, our confidence intervals enjoy the uniform coverage property over a set of data generating processes involving a set of values of the tail index.
The key insight is that the induced order statistics in each time series can be treated as approximately stemming from the true conditional distribution, and the large order statistics among these induced values can then be used to make inference on extremal quantiles. By focusing on the induced order statistics, we effectively reduce the conditional tail problem into an unconditional one. Monte Carlo simulations show that the new method delivers preferred small sample performance in terms of coverage probability and length.
The new method is more flexible than the extremal quantile regression because the latter assumes that the conditional extremal quantile is a parametric location-shift model. If a linear regression model is imposed, then our proposed method can be easily combined with any existing consistent estimator of structural parameters and applies to inference on extremal quantiles of the random coefficients.
Applying the proposed method to Natality Vital Statistics, we reexamine factors of extremely low birth weights that have been analyzed by preceding studies. We find that signs of major effects are the same as those found in preceding studies based on parametric models. On the other hand, we find that some effects exhibit different magnitudes from those reported in the previous studies based on parametric models.