Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
40,733 characters · 8 sections · 1 citation commands
Rearranging Edgeworth-Cornish-Fisher Expansions
Approximations to the distribution of sample statistics of higher order than the order $n^{-1/2}$ provided by the central limit theorem are of central interest in the theory of asymptotic statistics. See, e.g., \citeasnoun{bhattacharya_rao}, \citeasnoun{Rothenberg:handbook}, \citeasnoun{hall_bootstrap_book}, \citeasnoun{blinnikov_moessner}, \citeasnoun{vaart:text}, and \citeasnoun{cramer_book}. An important tool for performing these refinements is provided by the Edgeworth expansion (\citeasnoun{Edgeworth_1905}, \citeasnoun{Edgeworth_1907}), which approximates the distribution of the statistics of interest around the limit distribution (often the normal distribution) using a combination of Hermite polynomials with coefficients defined in terms of population moments. Inverting the expansion yields a related higher order approximation, the Cornish-Fisher expansion (\citeasnoun{Cornish_Fisher_1938}, \citeasnoun{Fisher_Cornish_1960}), to the quantiles of the statistic around the quantiles of the limiting distribution.
One important shortcoming of either the Edgeworth or Cornish-Fisher expansions is that the resulting approximations to the distribution and quantile functions are not necessarily increasing, which violates an obvious monotonicity requirement. This comes from the fact that the polynomials involved in the expansion are not monotone. Here we propose to use a procedure, called the rearrangement, to restore the monotonicity of the approximations and, perhaps more importantly, to improve the estimation properties of these approximations. The resulting improvement is due to the fact that the rearrangement necessarily brings the non-monotone approximations closer to the true monotone target function.
The main findings of the paper can be illustrated through a single picture given in Figure 1, where we plot the true distribution function of a standardized sample mean $X$ based on a small sample, a third order Edgeworth approximation to that distribution, and the rearrangement of the third order approximation. We see that the Edgeworth approximation is sharply non-monotone and provides a rather poor approximation to the distribution function. The rearrangement merely sorts the value of the approximate distribution function in increasing order. One can see that the rearranged approximation, in addition to being monotonic, is a much better approximation to the true function than the original approximation.
We organize the rest of the paper as follows. In Section (ref), we describe the rearrangement and qualify the approximation property it provides for monotonic functions. In Section (ref), we introduce the rearranged Edgeworth-Cornish-Fisher expansions and explain how they produce better approximations to distributions and quantiles of sample statistics. In Section (ref), we illustrate the procedure with several additional examples.
In what follows, let $\mathcal{X}$ be a compact interval. We first consider an interval of the form $\mathcal{X}=[0,1]$. Let $f(x)$ be a measurable function mapping $\mathcal{X}$ to $K$, a bounded subset of $\Bbb{R}$. Let $F_{f}(y) := \int_{\mathcal{X}} 1\{ f(u) \leq y \} du$ denote the distribution function of $f(X)$ when $X$ follows the uniform distribution on $[0,1]$. Let $$f^*(x) = Q_f(x) := \inf \left \{ y \in \Bbb{R}: F_{f}(y) \geq x \right \}$$ be the quantile function of $F_f(y)$. Thus, $$ f^*(x): = \inf \left \{ y \in \Bbb{R}: \left [ \int_{\mathcal{X}} 1\{ f(u) \leq y \} du \right] \geq x \right \}. $$ This function $f^*$ is called the increasing rearrangement of the function $f$. The rearrangement is a tool extensively used in functional analysis and optimal transportation (see, e.g., \citeasnoun{HLP52} and \citeasnoun{villani}.) It originates in the work of Chebyshev, who used it to prove a set of inequalities (\citeasnoun{bronshtein_handbook_mathematics}, p. 31). Here, we employ this tool to improve approximations of monotone functions, such as the Edgeworth-Cornish-Fisher approximations to the distribution and quantile functions of sample statistics.
The rearrangement operator simply transforms a function $f$ to its quantile function $f^*$. That is, $x\mapsto f^*(x)$ is the quantile function of the random variable $f(X)$ when $X\sim U(0,1)$. Another convenient way to think of the rearrangement is as a sorting operation: Given values of the function $f(x)$ evaluated at $x$ in a fine enough mesh of equidistant points, we simply sort the values in increasing order. The function created in this way is the rearrangement of $f$.
Finally, if $\mathcal{X}$ is of the form $[a,b]$ with $a < b$, let $\bar x(x) = (x-a)/(b-a) \in [0,1]$ for $x \in \mathcal{X}$, $x(\bar{x}) = a + (b-a)\bar{x} \in [a,b]$ for $\bar{x} \in [0,1]$, and $\bar f^*$ be the rearrangement of the function $\bar f(\bar x) = f(x(\bar x))$ defined on $\bar{ \mathcal{X}}=[0,1]$. Then, the rearrangement of $f$ is defined as $$ f^*(x) := \bar f^*(\bar x(x)). $$
The following result establishes that the rearrangement always improves the quality of the approximation to a monotone target function.
The first part of Proposition (ref) states the weak inequality ((ref)), and the second part states the strict inequality ((ref)). As an implication, Corollary (ref) states that the inequality is strict for $p \in (1, \infty)$ if the initial approximation $\widehat f(x)$ is decreasing on a subset of $\mathcal{X}$ having positive measure, while the target function $f_0(x)$ is increasing on $\mathcal{X}$ (where by increasing, we mean strictly increasing throughout). Proposition (ref) establishes that the rearranged approximation $\widehat f^*$ has a smaller estimation error in the $L_p$ norm than the initial approximation whenever the latter is not monotone. This is a very useful and generally applicable property that is independent of the way the initial approximation of $f_0$ is obtained.
Proof of Proposition (ref). We consider the case where $\mathcal{X} =[0,1]$ only, as the more general intervals can be dealt similarly. The first part establishes the weak inequality, following, in part, the strategy in Lorentz's proof. The proof focuses directly on obtaining the result stated in the proposition. The second part establishes the strong inequality.
Proof of Part 1. We assume first that the functions $\hat f$ and $f_0$ are step functions, constant on intervals $((s-1)/r, s/r]$, $s=1,\ldots,r$. For each step function $f$ with $r$ steps we associate an $r$-vector $f$ whose $s$-th element, denoted $f_{s}$, equals to the value of function $f$ on the $s$-th interval, and vice versa. Let us define the sorting operator $S$ acting on vectors (and functions) $f$ as follows. Let $k$ be an integer in $1,\ldots,r$ such that $f_k > f_m$ for some $m>k$. If $k$ does not exist, set $Sf = f$. If $k$ exists, set $Sf$ to be a $r$-vector with the $k$-th element equal to $f_m$, the $m$-th element equal to $f_k$, and all other elements equal to the corresponding elements of $f$. Finally, given a vector $Sf$ there is a step function $Sf$ associated to it, as stated above.
For any submodular function $L: \Bbb{R}^2 \to \Bbb{R}_+$, by $f_k \geq f_m$, $f_{0m} \geq f_{0k}$ and the definition of the submodularity, $L(f_m, f_{0k}) + L (f_k, f_{0m}) \leq $ $L(f_k, f_{0k}) + L(f_m, f_{0m}).$ Thus conclude that $ \int_{\mathcal{X}} L\{S\hat f(x), f_0(x)\} dx \leq \int_{\mathcal{X}} L\{ \hat f(x), f_0(x) \}dx,$ using that we integrate step functions. Applying the sorting operator a sufficient finite number of times to $\hat f$, we obtain a completely sorted, that is, rearranged, vector $\hat f^*$. Thus, we can express $\hat f^*$ as $\hat f^*= S \ldots S \hat f$, where the operator $S$ is applied finitely many times. By repeating the argument above, each application weakly reduces the estimation error. Therefore,
Next we extend this result to general measurable functions $\hat f$ and $f_0$ mapping $[0,1]$ to $K$, where $f_0$ is a quantile function. Take a subsequence of bounded step functions $\hat f^{(q)}$ and $f_0^{(q)}$, with $f_0^{(q)}$ being quantile functions, converging to $\hat f$ and $f_0$ almost everywhere as index $q \to \infty$ along an increasing sequence of integers. The almost everywhere convergence of $\hat f^{(q)}$ to $\hat f$ implies the almost everywhere convergence of its quantile function $\hat f^{*(q)}$ to the quantile function of the limit, $\hat f^*$ (\citeasnoun{vaart:text}, p. 305). Since ((ref)) holds for each $q$ along the sequence, the dominated convergence theorem implies that ((ref)) also holds for the general case.
It remains to show the existence of the subsequence in the preceding paragraph. Using series expansion in the Haar basis, any function in $L^2[0,1]$ can be approximated in $L^2$ norm by a sequence of $r$-step functions, where $r=2^j-1$ and $j =1,\ldots,\infty$ pollard:measure. Hence there is a sequence of step functions $ \hat f^{(r)}$ and $f^{(r)}_0$ converging to $\hat f$ and $f_0$ in $L^2$ norm; the functions in the sequence necessarily take values in $K$; by \citeasnoun{pollard:measure}, p. 38, we can extract a further subsequence $\hat f^{(q)}$ and $f^{(q)}_0$, with $q$ running over an increasing sequence of integers, converging to $\hat f$ and $f_0$ almost everywhere. Finally, replace $f^{(q)}_0$ by their quantile functions, i.e., rearrangements, which retain the almost everywhere convergence property to $f_0$ by \citeasnoun{vaart:text}, p. 305. \qed
Proof of Part 2. Consider the step functions, as defined in the proof of Part 1. By setting $r$ sufficiently large, we can take them to satisfy the following hypotheses: there exist regions $\mathcal{X}_0$ and $ \mathcal{X}'_0$, each of measure greater than $\delta>0$, such that for all $x \in \mathcal{X}_0$ and $x' \in \mathcal{X}_0'$, we have that (i) $x' > x$, (ii) $\hat f(x) > \hat f(x') + \epsilon$, and (iii) $f_0(x') > f_0(x) + \epsilon$, for $\epsilon>0$ specified in the proposition. For any strictly submodular function $L: \Bbb{R}^2 \to \Bbb{R}_+$ we have that $\eta = \inf \{ L(v',t) + L(v,t') - L(v,t) - L(v',t') \} >0, $ where the infimum is taken over all $v, v', t, t'$ in the set $K$ such that $v' \geq v + \epsilon$ and $t' \geq t + \epsilon$. We can begin sorting by exchanging an element $\hat f(x)$, $x \in \mathcal{X}_0$, of $r$-vector $\hat f$ with an element $\hat f(x')$, $x' \in \mathcal{X}_0'$, of $r$-vector $\hat f$. This induces a sorting gain of at least $\eta$ times $1/r$. The total mass of points that can be sorted in this way is at least $\delta$. We then proceed to sort all of these points in this way, and then continue with the sorting of other points. After the sorting is completed, the total gain from sorting is at least $\delta \eta$. That is, $ \int_{\mathcal{X}} L\{\hat f^*(x), f_0(x)\}dx \leq \int_{\mathcal{X}} L\{\hat f(x), f_0(x)\} dx - \delta \eta.$
We then extend this inequality to the general measurable functions exactly as in the proof of Part 1. \qed
In the next section, we apply rearrangements to improve the Edgeworth-Cornish-Fisher and related approximations to distribution and quantile functions.
We first consider the quantile case. Let $Q_n$ be the quantile function of a statistic $X_n$, i.e., $$ Q_n(u) = \inf\{ x \in \Bbb{R}: Pr[X_n \leq x] \geq u \}, $$ which we assume to be strictly increasing. Let $\widehat Q_n$ be an approximation to the quantile function $Q_n$ satisfying the following relation:
where $a_n$ is some sequence of positive numbers going to zero as $n \rightarrow \infty$, $\mathcal{U}_n=[\varepsilon_n, 1- \varepsilon_n] \subseteq [0,1]$, and $\varepsilon_n< 1$ is some sequence of positive numbers possibly going to zero as $n \rightarrow \infty$. For example, ((ref)) holds with $\varepsilon_n = n^{-c}$, $c>0$, under certain conditions on the moments (e.g., \citeasnoun{hall_bootstrap_book}).
The leading example of such an approximation is the inverse Edgeworth, or Cornish-Fisher, expansion of the quantile function of a sample mean. If $X_n$ is the standardized sample mean, $X_n = n^{-1/2} \sum_{i=1}^n (Y_i - E[Y_i])/\sqrt{Var(Y_i)}$, based on a random sample $(Y_1,..., Y_n)$ of $Y$, then we have the following $J$-th order expansion
provided that a set of regularity conditions, specified, e.g., in \citeasnoun{zolotarev_pobabilistic}, hold. Here $\Phi$ and $\Phi^{-1}$ denote the distribution function and quantile function of a standard normal random variable. The first three terms of the expansion are given by the polynomials,
where $\lambda$ is the skewness and $\kappa$ is the kurtosis of the random variable $Y$. The Cornish-Fisher expansion is one of the central approximations of the asymptotic statistics. Unfortunately, an inspection of the expressions for the polynomials in ((ref)) reveals that this expansion does not generally deliver a monotone approximation of the quantile function. This shortcoming has been pointed and discussed in detail for example by \citeasnoun{hall_bootstrap_book}. The nature of the polynomials is such that there always exists a large enough range $\mathcal{U}_n$ over which the Cornish-Fisher approximation is not monotone, cf., \citeasnoun{hall_bootstrap_book}. As an example, in the case of the second order approximation ($J=2$) we have that, for $\lambda <0$
that is, the Cornish-Fisher “quantile" function $\widehat Q_n$ is decreasing far enough in the tails. This example merely suggests a potential problem that may or may not apply to practically relevant ranges of probability indices $u$. Indeed, specific numerical examples given below show that in small samples the non-monotonicity can occur in practically relevant ranges. Of course, in sufficiently large samples, the regions of non-monotonicity are squeezed quite far into the tails.
Let $\widehat Q_n^*$ be the rearrangement of $\widehat Q_n$. Then we have that for any $p \in [1,\infty]$, the rearranged quantile function reduces the approximation error of the original approximation:
with the first inequality holding strictly for $p \in (1,\infty)$ whenever $\widehat Q_n$ is decreasing on a region of $\mathcal{U}_n$ of positive measure. We can give the following probabilistic interpretation to this result. Under condition ((ref)), there exists a variable $U= F_n(X_n)$, where $F_{n}$ is the distribution function of $X_{n}$, such that both the stochastic expansion
and the expansion
hold.\footnote{$\widehat Q_n^*(U)$ is defined only on $\mathcal{U}_n$, so we can set $\widehat Q_n^*(U)=Q_n(U)$ outside $\mathcal{U}_n$, if needed. Of course, $U \not \in \mathcal{U}_n$ with probability going to zero if $\varepsilon_{n} \searrow 0$ as $n \rightarrow \infty$.} However, the variable $\widehat Q_n^*(U)$ in ((ref)) is a better coupling to the statistic $X_n$ than $\widehat Q_n(U)$ in ((ref)), in the following sense: For each $p \in [1, \infty]$,
where $1_n =1\{ U \in \mathcal{U}_n\}$. Indeed, property ((ref)) immediately follows from ((ref)).
The above improvements apply in the context of the sample mean $X_n$. In this case, the probabilistic interpretation above is directly connected to the higher order central limit theorem of \citeasnoun{zolotarev_pobabilistic}, which states that under ((ref)) we have the following higher-order probabilistic central limit theorem,
The term $\widehat Q_n(U)$ is Zolotarev's higher-order refinement over the first order normal term $\Phi^{-1}(U)$. \citeasnoun{sun_loader_mccormick} employ an analogous higher-order probabilistic central limit theorem to improve the construction of confidence intervals.
The application of the rearrangement to Zolotarev's term actually delivers a clear improvement in the sense that it also leads to a probabilistic higher order central limit theorem
where the leading term $\widehat Q^*_n(U)$ is closer to $X_n$ than Zolotarev's term $Q_n(U)$, in the sense of ((ref)).
We summarize the above discussion into a formal proposition.
We next consider distribution functions. Let $F_n(x)$ be the distribution function of a statistic $X_n$, assumed to be strictly increasing, and $\widehat F_n(x)$ be an approximation to this distribution such that the following relation holds:
where $a_n$ is some sequence of positive numbers going to zero as $n \rightarrow \infty$, and $\mathcal{X}_n =[- b_n, c_n]$ is an interval in $\Bbb{R}$ for some sequences of positive scalars $b_n$ and $c_n$ possibly growing to infinity. The choice of $b_n$ and $c_n$ is unrestricted under certain conditions on the moments (e.g., \citeasnoun{hall_bootstrap_book}).
The leading example of such an approximation is the Edgeworth expansion of the distribution function of the sample mean. If $X_n$ is the standardized sample mean, $X_n = n^{-1/2} \sum_{i=1}^n (Y_i - E[Y_i])/\sqrt{Var(Y_i)}$, based on a random sample $(Y_1,..., Y_n)$ of $Y$, then we have the following $J$-th order expansion
for some $C >0$, provided that a set of regularity conditions, specified, e.g., in \citeasnoun{hall_bootstrap_book}, hold. The first three terms of the approximation are given by
where $\Phi$ and $\phi$ denote the distribution function and density function of a standard normal random variable, and $\lambda$ and $\kappa$ are the skewness and kurtosis of the random variable $Y$, respectively. Here too, the Edgeworth expansion is one of the central approximations of the asymptotic statistics. Unfortunately, like the Cornish-Fisher expansion, it generally does not provide a monotone approximation of the distribution function. This shortcoming has been pointed and discussed in detail by \citeasnoun{barton_dennis}, \citeasnoun{draper_tierney}, \citeasnoun{sargan_edgeworth}, and \citeasnoun{balitskaya_zolotukhina}, among others.
Let $\widehat F_n^*$ be the rearrangement of $\widehat F_n$. Then, we have that for any $p \in [1,\infty]$, the rearranged Edgeworth approximation reduces the approximation error of the original Edgeworth approximation:
with the first inequality holding strictly for $p \in (1,\infty)$ whenever $\widehat F_n$ is decreasing on a region of $\mathcal{X}_n$ of positive measure.
In some cases, it can be worthwhile to weigh different areas of the support differently than the Lebesgue (flat) weighting prescribes. For example, it may be desirable to rearrange $\widehat F_n$ using $F_n$ as a weighting measure. Indeed, by using $F_n$ as a weight, we obtain a better matching with the P-value: $ P=F_n(X_n)$ (in this quantity, $X_n$ is drawn according to the true $F_n$). Using such a weight will provide a probabilistic interpretation for the rearranged Edgeworth expansion, analogous to the probabilistic interpretation for the rearranged Cornish-Fisher expansion. Although the weight (true $F_n$) is not available, we can use the standard normal measure $\Phi$ as the weighting measure instead. We may also construct an initial rearrangement with the Lebesgue weight, and use it as weight itself for a further weighted rearrangement (and even continue to iterate in this fashion). Using non-Lebesgue weights may also be desirable when we want the improved approximations to weigh the tails more heavily. Whatever the reason may be for further non-Lebesgue weighting, they have the following properties, which follow immediately in light of Remark 3.
Let $\Lambda$ be a distribution function that admits a positive density with respect to the Lebesgue measure on the region $\mathcal{U}_n=[\varepsilon_n, 1- \varepsilon_n]$ for the quantile case and on the region $\mathcal{X}_n=[-b_n, c_n]$ for the distribution case. Then, if ((ref)) holds, the $\Lambda$-weighted rearrangement $\widehat Q^*_{n, \Lambda}$ of the function $\widehat Q_n$ satisfies
where the first equality holds strictly when $\widehat Q$ is decreasing on a subset of positive $\Lambda$-measure. Furthermore, if ((ref)) holds, then the $\Lambda$-weighted rearrangement $\widehat F^*_{n, \Lambda}$ of the function $\widehat F_n$ satisfies
In addition to the Log-normal example given in the introduction, we use the Gamma distribution to illustrate the improvements that the rearrangement provides. Let $(Y_1,...,Y_n)$ be an i.i.d. sequence of Gamma(1/16,16) random variables. The statistic of interest is the standardized sample mean $X_n = n^{-1/2} \sum_{i=1}^n (Y_i - E[Y_i])/\sqrt{Var(Y_i)}$. We consider samples of sizes $n=4, 8, 16$, and $32$. In this example, the distribution function $F_n$ and quantile function $Q_n$ of the statistic $X_n$ are available in a closed form, making it easy to compare them to the Edgeworth approximation $\widehat F_n$ and the Cornish-Fisher approximation $\widehat Q_n$, as well as to the rearranged Edgeworth approximation $\widehat F_n^*$ and the the rearranged Cornish-Fisher approximation $\widehat Q_n^*$. For the Edgeworth and Cornish-Fisher approximations, as defined in the previous section, we consider third order expansions, that is we set $J=3$.
Figure (ref) compares the true distribution function $F_n$, the Edgeworth approximation $\widehat F_n$, and the rearranged Edgeworth approximation $\widehat F_n^*$. The standard normal first order approximation is also included as a benchmark of comparison. We see that the rearranged Edgeworth approximation not only solves the monotonicity problem, but also consistently does a better job at approximating the true distribution than the Edgeworth approximation. Table (ref) further supports this point by presenting the numerical results for the $L_p$ approximation errors, calculated according to the formulas given in the previous section. We see that the rearrangement reduces the approximation error quite substantially in most cases.
Figure (ref) compares the true quantile function $Q_n$, normal first order approximation, Cornish-Fisher approximation $\widehat Q_n$, and rearranged Cornish-Fisher approximation $\widehat Q_n^*$. Here too we see that the rearrangement not only solves the non-monotonicity problem, but also brings the approximation closer to the truth. Table (ref) further supports this point numerically, showing that the rearrangement reduces the $L_p$ approximation error quite substantially in most cases.
In this paper, we have applied the rearrangement procedure to monotonize Edgeworth and Cornish-Fisher expansions and any other related expansions of distribution and quantile functions. The benefits of doing so are twofold. First, we have obtained approximations to the distribution and quantile curves of the statistics of interest which satisfy the logical monotonicity restriction, unlike those directly given by the truncation of the series expansions. Second, we have shown that doing so results in better approximation properties.