Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
161,589 characters · 30 sections · 88 citation commands
Modified Wilcoxon-Mann-Whitney tests of stochastic dominance
\hypersetup{ colorlinks=true, linkcolor=customblue, citecolor=customblue, urlcolor=customblue, filecolor=customblue, allcolors=customblue }
Tests of stochastic dominance occupy a central place in empirical economics. They are routinely employed to compare income distributions, portfolio returns, treatment effects, and many other outcomes of interest. The econometric literature on the subject begins with M89; other foundational contributions include A96, DD00, BD03 and LMW05. A major advance of central relevance to the present article was the development of a powerful bootstrap test of stochastic dominance in LSW10. The test, called the Linton-Song-Whang test or simply the LSW test in what follows, uses a novel bootstrap procedure incorporating a preliminary estimate of a contact set. The effect of contact set estimation is to broaden the set of null configurations at which the limiting rejection rate is equal to the nominal level, thereby improving power against nearby alternatives. Subsequent literature exploring variations upon the LSW test includes DH16, which uses a selective recentering method in place of contact set estimation; LT21, which proposes to estimate the contact set by the method of empirical likelihood; and ZWC24, which uses contact set estimation in conjunction with a variance-weighted Kolmogorov-Smirnov statistic.
The LSW test served as a leading example in the general theory of bootstrap inference for directionally differentiable functionals developed in FS19. The theory provides high-level conditions that guarantee the asymptotic validity of bootstrap procedures based on the estimation of a directional derivative, commonly one characterized by a contact set. The LSW test is shown to satisfy these conditions. By placing the LSW test within a more general framework of bootstrap inference based on contact set estimation, FS19 facilitated the adaptation of the central methodological insight in LSW10 to other hypothesis testing problems of interest, some only tangentially related to stochastic dominance. Examples include tests of stochastic monotonicity S18, density ratio ordering BS19, matrix rank CF19, Lorenz dominance SB21, instrument validity S23, inverse stochastic dominance JSH24 and almost stochastic dominance SS26.
The present article contributes to this literature by developing a new rank-based test of first-order stochastic dominance. Our test statistic is the one-sided Wilcoxon-Mann-Whitney (WMW) statistic studied in ST96. The one-sided WMW statistic can be computed from the procentile-procentile (P-P) plot for two samples drawn from two populations. The P-P plot is simply the empirical cumulative distribution function (cdf) for the first sample composed with the empirical quantile function for the second sample. When the P-P plot rises above the 45-degree line, this may be taken as evidence that the first population does not first-order stochastically dominate the second population. (ref) displays a P-P plot for two samples of size $10$. The one-sided WMW statistic is equal to the area shaded in red multiplied by a number depending on the two sample sizes, in this case $\sqrt{5}$. Ignoring the small white triangles above the 45-degree line---which may be understood to result from the approximation of an integral by a sum and are negligible for large sample sizes---the one-sided WMW statistic is proportional to the area below the P-P plot and above the 45-degree line, and may thus be understood to constitute an area-based measurement of the evidence against first-order stochastic dominance. It depends on the sample observations only through their pooled ranks because the P-P plot depends only on those ranks. For this reason the one-sided WMW statistic is said to be a rank-based statistic.
The asymptotic distribution of the one-sided WMW statistic when the two population distributions are equal was obtained in ST96 for cases where the two samples are drawn independently from two populations. It is the distribution of the area that lies beneath a Brownian bridge and above zero. Because this distribution is free of nuisance parameters, it is simple to implement an asymptotically valid test of first-order stochastic dominance using the one-sided WMW statistic; one need merely compare the statistic to tabulated critical values. However, in many economic applications samples are not drawn independently of one another, but rather are drawn as matched pairs from a single bivariate population distribution. In (ref) we show that the asymptotic distribution of the one-sided WMW statistic under matched pairs sampling is more complicated in form than under independent sampling, and can be represented as a functional of a tied-down Brownian sheet whose covariance kernel is determined by the copula linking paired observations. The dependence of the asymptotic distribution on the unknown copula complicates inference and invalidates the use of critical values tabulated for cases where samples are independent.
To overcome this obstacle we provide, in (ref), two bootstrap procedures for computing a critical value for the one-sided WMW statistic. The first is a standard application of the bootstrap while the second incorporates an estimate of the contact set; i.e., the set on which the population P-P curve is equal to the 45-degree line. The contact set estimator requires a user-specified tuning parameter and applies a variance-weighted exclusion rule which, in the case of matched pairs sampling, uses the empirical copula to account for dependence within pairs. (ref) show that both bootstrap procedures result in tests whose limiting rejection rates are no greater than the nominal level at all null configurations and are equal to one at all alternative configurations. The latter result further shows that the bootstrap critical value incorporating contact set estimation delivers a limiting rejection rate equal to the nominal level everywhere on the boundary of the null, by which we mean the set of all null configurations with positive measure contact set. We comment on the very close relationship between our procedure involving contact set estimation and the LSW test in (ref), drawing attention to some differences connected to the construction of the contact set and to the rank-based nature of the one-sided WMW statistic.
The outcome of numerical simulations pertaining to the small sample performance of our bootstrap critical values, both with independent samples and with matched pairs, is reported in (ref). For both sampling frameworks the simulations show that contact set estimation can produce a large improvement in power, as is the case for the LSW test, while maintaining control of type I error for reasonable choices of the tuning parameter. The results are particularly encouraging in cases where there is strong dependence within matched pairs. (ref) contains a brief empirical illustration of our procedures involving Canadian family income data. We offer some concluding thoughts in (ref). Mathematical proofs are provided in (ref).
Let $F_1:\mathbb{R}\rightarrow[0,1]$ and $F_2:\mathbb{R}\rightarrow[0,1]$ be cdfs. Define the quantile function $Q_2:(0,1)\to\mathbb R$ by $Q_2(u)=\inf\left\{x\in\mathbb R:F_2(x)\ge u \right\}$, and define the P-P curve $R:[0,1]\to[0,1]$ by
The P-P curve is also commonly called an ordinal dominance curve or receiver operating characteristic curve. The definition at $R(u)$ at the endpoints $u=0$ and $u=1$ guarantees that $R$ is continuous at those endpoints. Discontinuities at intermediate values of $u$ are excluded by the following assumption.
Absolute continuity of $R$ is a stronger property than continuity, and implies the existence of a density $r:[0,1]\to\mathbb R$ for $R$, uniquely determined up to null sets. To be concrete we choose $r$ to be equal to the derivative of $R$ at all differentiability points of $R$, and equal to one on the zero measure set where $R$ is not differentiable. The density $r$ is nonnegative and integrable because $R$ is nondecreasing. Absolute continuity of $R$ does not require $F_1$ or $F_2$ to be continuous, but does require that any discontinuities of $F_1$ fall outside of the interior of the convex support of $F_2$.
Absolute continuity of $R$ does not imply that $r$ is bounded. We wish not to exclude cases where $r$ is unbounded because they arise naturally in simple examples. For instance, if $F_1$ and $F_2$ are Gaussian cdfs with equal variances but different means $\mu_1$ and $\mu_2$ then $R$ is absolutely continuous and $r(u)$ diverges to infinity either as $u\downarrow0$ if $\mu_1<\mu_2$ or as $u\uparrow1$ if $\mu_1>\mu_2$. It has nevertheless been quite common in prior research involving P-P curves to make assumptions implying that $r$ is bounded. Examples include BM15 and BS19, where it is assumed that $R$ is continuously differentiable.
We consider two sampling frameworks: independent samples and matched pairs. Our simultaneous treatment of the two sampling frameworks follows past econometric literature including BDB14 and SB21. In both frameworks it should be understood that $\{X_i^1\}_{i=1}^\infty$ and $\{X_i^2\}_{i=1}^\infty$ are sequences of independent and identically distributed (iid) random variables drawn from $F_1$ and $F_2$ respectively. The observed samples are $\{X_i^1\}_{i=1}^{n_1}$ and $\{X_i^2\}_{i=1}^{n_2}$. The sample sizes $n_1$ and $n_2$ should be understood to implicitly depend monotonically on an underlying index $n\in\mathbb N$, which will tend to infinity in subsequent asymptotic arguments. In both sampling frameworks we simply let $n_2=n_2(n)=n$. In the matched pairs sampling framework we also let $n_1=n_1(n)=n$ and the sample of pairs $\{(X_i^1,X_i^2)\}_{i=1}^n$ is assumed to be iid. In the independent sampling framework the iid samples $\{X_i^1\}_{i=1}^{n_1}$ and $\{X_i^2\}_{i=1}^{n_2}$ are assumed to be independent of one another, and to model the relative growth of samples sizes we assume that
Note that (ref) is automatically satisfied in the matched pairs sampling framework, with $\lambda=1/2$. Note also that the convergence to $\lambda\in(0,1)$ in (ref) implies that $n_1n_2/(n_1+n_2)\to\infty$ if and only if $n_1\to\infty$.
The copula linking the marginal cdfs $F_1$ and $F_2$ will play an important role in the matched pairs sampling framework. In this framework we let $C:[0,1]^2\to[0,1]$ be a copula for each pair $(X_i^1,X_i^2)$. Sklar's theorem establishes the existence of $C$, and shows that $C$ is uniquely defined on the product of the closed ranges of $F_1$ and $F_2$. In the independent sampling framework we simply let $C:[0,1]^2\to[0,1]$ be the product copula $C(u,v)=uv$.
We seek to test the null hypothesis that $F_1$ first-order stochastically dominates $F_2$ (weakly). When $R$ is continuous a simple argument shows that $F_1$ first-order stochastically dominates $F_2$ if and only if $R$ does not exceed the 45-degree line. The hypotheses we seek to discriminate between are thus
Define the empirical cdfs $\hat{F}_1:\mathbb R\to[0,1]$ and $\hat{F}_2:\mathbb R\to[0,1]$ by
Further define the empirical quantile function $\hat{Q}_2:(0,1)\to\mathbb R$ by $\hat{Q}_2(u)=\inf\{x\in\mathbb R:\hat{F}_2(x)\geq u\}$, and define the P-P plot $\hat{R}:[0,1]\to[0,1]$ by
Throughout this article, all notation decorated with a circumflex (or “hat”) or with a tilde is implicitly indexed by $n$ and refers to something estimated from data.
The test statistic we consider is the one-sided WMW statistic studied in ST96. It provides an estimate of the area below the P-P curve and above the 45-degree line. Let $L^1[0,1]$ be the usual normed space of Lebesgue integrable functions from $[0,1]$ into $\mathbb R$, and define the functional $\mathcal H:L^1[0,1]\to\mathbb R$ by
The area below $R$ and above the 45-degree line is then $\mathcal H(R)$. We have $\mathcal H(R)=0$ under $\text{H}_0$, and $\mathcal H(R)>0$ under $\text{H}_1$. Let $X^2_{(1)},\dots,X^2_{(n_2)}$ be the sample observations $X^2_1,\dots,X^2_{n_2}$ ranked from smallest to largest. The one-sided WMW statistic $\hat{S}$ is defined by
Using the fact that $\hat{R}(u)=\hat{F}_1(X^2_{(i)})$ when $(i-1)/n_2<u\leq i/n_2$, a simple argument shows that
We may therefore regard $\hat{S}$ to be a convenient approximation to $\displaystyle{T_n^{1/2}\mathcal H(\hat{R})}$, with the error bound $\displaystyle{T_n^{1/2}/(2n_2)}$ vanishing asymptotically under (ref). The approximation error may be understood to correspond to the small white triangles above the 45-degree line in (ref). In general there are at most $n_2$ such triangles and each has area $1/(2n_2^2)$.
We will first discuss the asymptotic behavior of $\displaystyle{T_n^{1/2}(\hat{R}-R)}$, and then explain how the asymptotic behavior of $\hat{S}$ may be deduced from that of $\displaystyle{T_n^{1/2}(\hat{R}-R)}$ by applying the delta-method.
We call the normalized P-P plot $\displaystyle{T_n^{1/2}(\hat{R}-R)}$ the P-P process. A central ingredient to our study of the asymptotic behavior of the one-sided WMW statistic is the fact that, as $n\to\infty$, the P-P process converges in distribution in $L^1[0,1]$. Our definition of convergence in distribution is the one stated in V98 for sequences of “random elements” of a metric space, in this case the separable space $L^1[0,1]$. The notation $\rightsquigarrow$ will be used to signify convergence in distribution in a metric space. If that metric space is $\mathbb R$ then we will instead write $ \mathrel{\mathop{\rightarrow}\limits^{ \vbox to0.4ex{\kern0.2\ex@ \hbox{$\mathrm{\scriptscriptstyle{D}\,}$}\vss}}}$.
The limit in distribution of the P-P process may be expressed in terms of a centered Gaussian process $\mathcal B:[0,1]^2\to\mathbb R$ with covariance kernel
In GS87 the process $\mathcal B$ is referred to as a tied-down Brownian sheet with intensity measure $C$. The marginal processes $\mathcal B_1:[0,1]\to\mathbb R$ and $\mathcal B_2:[0,1]\to\mathbb R$ defined by $\mathcal B_1(u)=\mathcal B(u,1)$ and $\mathcal B_2(u)=\mathcal B(1,u)$ are Brownian bridges. The two Brownian bridges are independent if and only if $C$ is the product copula.
From $\mathcal B$ we construct a centered Gaussian process $\mathcal R:[0,1]\to\mathbb R$ by setting
Here $\lambda$ is the limit appearing in (ref), which is equal to $1/2$ in the matched pairs sampling framework or may take any value in $(0,1)$ in the independent sampling framework. As discussed in BK26, in the matched pairs sampling framework the distribution of $\mathcal R$ depends on $C$ only through the values taken by $C$ on the product of the closed ranges of $F_1$ and $F_2$, these values being uniquely determined by Sklar's theorem.
(ref) may be compared to, for instance, Theorem 3.1 in ACH87 for independent sampling, or Lemma 1.1 in WT21 for matched pairs sampling. Those results place stronger regularity conditions on $F_1$ and $F_2$ and establish convergence in distribution to $\mathcal R$ with respect to a uniform metric. Such convergence cannot be established under the assumptions of (ref) because these assumptions---the only relevant part being the requirement that $R$ is absolutely continuous---do not imply that $\mathcal R$ has bounded sample paths. This will not create difficulties because, as we will see, the convergence in distribution in $L^1[0,1]$ established by (ref) suffices to suitably control the asymptotic behavior of the one-sided WMW statistic.
In view of (ref), the difference between the test statistic $\hat{S}$ and $\displaystyle{T_n^{1/2}\mathcal H(\hat{R})}$ is asymptotically negligible. It will be more convenient for us to study the behavior of the latter quantity. The null hypothesis of first-order stochastic dominance is satisfied if and only if $\mathcal H(R)=0$. In this case we have
In view of the convergence in distribution of $\displaystyle{T_n^{1/2}(\hat R-R)}$ established in (ref), we can obtain the limit distribution of $\displaystyle{T_n^{1/2}(\mathcal H(\hat R)-\mathcal H(R))}$ in (ref) by applying the delta-method. Define the functional $\mathcal H'_R:L^1[0,1]\to\mathbb R$ by
where $B_+$ and $B_0$ are the sets
Adopting the terminology introduced in LSW10, we refer to the set $B_0$ as the contact set. It is established in (ref) that $\mathcal H'_R$ is the Hadamard directional derivative of $\mathcal H$ at $R$. See FS19 for the definition of Hadamard directional differentiability and a discussion of the delta-method oriented toward applications in econometrics. By applying the delta-method with the Hadamard directionally differentiable map $\mathcal H$ we arrive at the following consequence of (ref).
It is observed in ST96 that, under independent sampling and at null configurations such that $F_1=F_2$ (i.e., the least favorable case), the test statistic $\hat S$ converges in distribution to $\smallint_0^1\max\{\mathcal B(u),0\}\,\mathrm{d}u$, where $\mathcal B$ is a Brownian bridge. This follows from (ref) by noting that, under independent sampling, if $F_1=F_2$ then $\mathcal R$ is a Brownian bridge. We may therefore construct a test of $\text{H}_0$ with limiting rejection frequency no greater than $\alpha$ at all null configurations, and equal to $\alpha$ at null configurations with $F_1=F_2$, by rejecting $\text{H}_0$ when $\hat S$ exceeds the $(1-\alpha)$-quantile of $\smallint_0^1\max\{\mathcal B(u),0\}\,\mathrm{d}u$. The $0.9$, $0.95$ and $0.99$ quantiles are reported in ST96 to be $0.39$, $0.48$ and $0.68$, respectively.
In the matched pairs sampling framework $\mathcal R$ is no longer a Brownian bridge at all null configurations such that $F_1=F_2$, and depends on the unknown copula $C$. The critical values reported in ST96 therefore no longer apply. In the following section we propose a bootstrap scheme to produce critical values for $\hat S$ which apply under both independent sampling and matched pairs. Further, we show how power may be improved by incorporating an implicit estimate of the contact set into our scheme, similar to what is done in LSW10.
We consider two methods for constructing critical values. The first may be regarded as a standard implementation of the bootstrap. The second is a modification of the first based on an implicit estimate of the contact set.
Our baseline procedure to obtain a bootstrap critical value for the test statistic $\hat S$ is as follows. We first construct bootstrap cdfs $\hat F^\ast_1:\mathbb R\to[0,1]$ and $\hat F^\ast_2:\mathbb R\to[0,1]$ by setting
where $W^1_{n_1}=(W^1_{1,n_1},\ldots,W^1_{n_1,n_1})$ and $W^2_{n_2}=(W^2_{1,n_2},\ldots,W^2_{n_2,n_2})$ are random weights generated independently of the data. The way in which the weights are generated depends on the sampling framework. With independent samples we draw $W^1_{n_1}$ and $W^2_{n_2}$ independently of one another from the multinomial distribution with equal probabilities over the categories $1,\ldots,n_1$ and $1,\ldots,n_2$ respectively. With matched pairs we draw $W^1_n$ from the multinomial distribution with equal probabilities over the categories $1,\ldots,n$, and then set $W^2_n=W^1_n$. In either sampling framework, we then construct the bootstrap quantile function $\hat{Q}^\ast_2:(0,1)\to\mathbb R$ by setting $\hat Q^\ast_2(u)=\inf\{x\in\mathbb R:\hat F^\ast_2(x)\geq u\}$, and the bootstrap P-P plot $\hat R^\ast:[0,1]\to[0,1]$ by setting
We then compute the bootstrap test statistic
To obtain a test with nominal level $\alpha$ we independently generate a large number $N$ of bootstrap test statistics and choose as our critical value the $\lceil N(1-\alpha)\rceil$-th smallest of these, where $\lceil\cdot\rceil$ rounds up to the nearest integer.
Our second bootstrap procedure involves modifying the bootstrap statistic defined in (ref) so as to incorporate an implicit estimate of $B_0$, the contact set. The first step in this procedure is to compute, for $i\in\{1,\dots,n_2\}$, an estimate $\hat{V}_i$ of the variance of $\displaystyle{T_n^{1/2}\hat R(i/n_2)}$. This is discussed in more detail below. Next, to generate a single bootstrap test statistic, we generate $\hat{R}^\ast$ in the same way as in the standard bootstrap procedure described in (ref), and then compute the modified bootstrap test statistic
where $\tau_n\in(0,\infty)$ is a tuning parameter. Note that if we set $\tau_n=\infty$ in (ref) then, ignoring the possibility that $\hat{V}_i=0$ for some $i$, we recover the standard bootstrap test statistic defined in (ref). To obtain a test with nominal level $\alpha$ we independently generate a large number $N$ of modified bootstrap test statistics and choose as our critical value the $\lceil N(1-\alpha)\rceil$-th smallest of these. We reject $\text{H}_0$ when $\hat S$ exceeds this critical value.
The role of the indicator function in (ref) is to exclude summands for which $\hat R(i/n_2)$ falls below $i/n_2$ by at least $\tau_n$ estimated standard deviations. It provides an implicit estimate of the contact set. This will be made more clear in (ref). In the development of asymptotics to follow we will assume that $\tau_n$ diverges to infinity at a controlled rate as $n\to\infty$. See (ref) below. The numerical simulations reported in (ref) may be used to guide the choice of $\tau_n$ in practice.
The estimators $\hat{V}_i$ may be chosen to approximate pointwise variances of $\mathcal R$. It suffices for our purposes to focus on estimators that work well on the contact set $B_0$. If we set $R(u)=u$ and $r(u)=1$ in (ref) then, by working with the covariance kernel in (ref), we find that $\mathrm{Var}(\mathcal R(u))=u-C(u,u)$. We therefore propose setting
where with matched pairs we define $\hat{C}:[0,1]^2\to[0,1]$ to be the empirical copula
To study the asymptotic behavior of our bootstrap procedures we require a bootstrap analogue to the convergence in distribution of the P-P process established in (ref). See V98 for the definition of convergence in distribution conditional on the data in probability.
(ref) may be roughly understood to mean that, for large $n$, the distribution of $\displaystyle{T_{n}^{1/2}}(\hat{R}^\ast-\hat{R})$ conditional on the data is, with high probability, close to the distribution of $\mathcal R$. This is useful to know because the distribution of $\displaystyle{T_{n}^{1/2}(\hat{R}^\ast-\hat{R})}$ conditional on the data is precisely what we simulate by bootstrapping.
Let $I:[0,1]\to[0,1]$ be the identity map, and recall the definition of $\mathcal H'_R$ given in (ref). The standard bootstrap test statistic $\hat{S}^\ast$ defined in (ref) satisfies
If $\mathrm{H}_0$ is true then $\mathcal H'_R\leq\mathcal H'_{I}$. Thus
In view of (ref) we may expect that the distribution of the quantity on the right-hand side of the last inequality conditional on the data is close to the distribution of $\mathcal H'_R(\mathcal R)$ for large $n$. We know from (ref) that $\hat{S}$ converges in distribution to $\mathcal H'_R(\mathcal R)$ under $\mathrm{H}_0$. This suggests that if $\mathrm{H}_0$ is true then a reasonable upper bound on the quantiles of $\hat{S}$ may be provided by the corresponding quantiles of $\hat{S}^\ast$ conditional on the data. On the other hand, if $\mathrm{H}_1$ is true then we know from (ref) that $\hat{S}$ diverges in probability to infinity as $n$ grows, whereas (ref) and (ref) together suggest that the quantiles of $\hat{S}^\ast$ conditional on the data ought to become close to the corresponding quantiles of $\mathcal H'_{I}(\mathcal R)$.
The loose reasoning provided in the previous paragraph provides a heuristic justification for rejecting $\mathrm{H}_0$ when $\hat{S}$ exceeds the standard bootstrap critical value computed as described in (ref). To provide a rigorous justification we will need to exclude certain degenerate cases. The next assumption serves this purpose.
The assumption on $R$ excludes cases where $F_1$ assigns all mass either entirely to the right or entirely to the left of the convex support of $F_2$. The assumption on $C$ excludes cases where there is extreme positive dependence within matched pairs. Note that the Fr\'{e}chet-Hoeffding upper bound for $C(u,u)$ is $u$.
To study our modified bootstrap procedure we adopt an asymptotic framework in which the tuning parameter $\tau_n$ is assumed to diverge to infinity at a controlled rate as $n\to\infty$.
We claimed in (ref) that the indicator function in (ref) has the effect of providing an implicit estimate of the contact set $B_0$. The implicit estimate we were referring to is
though we note that under $\mathrm{H}_1$ it is more natural to regard $\hat{B}_0$ as an estimate of $B_0\cup B_+$. Of course $B_+$ is empty under $\mathrm{H}_0$. Using $\hat{B}_0$ we define the data-dependent functional $\hat{\mathcal H}':L^1[0,1]\to\mathbb R$ by
The functional $\hat{\mathcal H}'$ may be viewed as an implicit estimate of the directional derivative $\mathcal H'_R$ defined in (ref), again noting that $B_+$ is empty under $\mathrm{H}_0$. The modified bootstrap test statistic $\tilde S^\ast$ defined in (ref) may be rewritten as
revealing the connection between $\tilde S^\ast$ and the implicitly estimated contact set.
The representation of $\tilde{S}^\ast$ given in (ref) is useful because it facilitates the application of results in FS19 providing conditions sufficient for the validity of modified bootstrap procedures. By showing that $\hat{\mathcal H}'$ suitably approximates $\mathcal H'_R$ under $\mathrm{H}_0$, we are able to use (ref) above and Theorem 3.2 in FS19 to show that if $\mathrm{H}_0$ is true then, for large $n$, the distribution of $\displaystyle{\hat{\mathcal H}'(T_n^{1/2}(\hat R^\ast-\hat R))}$ conditional on the data is, with high probability, close to the distribution of $\mathcal H'_R(\mathcal R)$. See (ref) for a precise statement. We use (ref) to establish the following result.
The role of the constant $\eta$ appearing in the statement of (ref) is to control the limiting rejection frequency at null configurations with zero measure contact set, i.e.\ case (ii). At such configurations both the test statistic $\hat S$ and modified bootstrap critical value $\tilde c_{1-\alpha}$ converge in probability to zero, making it difficult to characterize the rejection frequency within our first-order asymptotic framework. See DH16, BS19 and SB21 for further discussion of this issue. Also see LSW10, where the regularity condition introduced in Definition 3 plays a similar role to $\eta$ by excluding null configurations at which the critical value converges in probability to zero. In DH16 it is recommended to set $\eta$ equal to a very small value such as $10^{-6}$ in practice, and to use $\max\{\tilde c_{1-\alpha},\eta\}$ as a critical value rather than $\tilde c_{1-\alpha}$; however it is also reported that in numerical simulations there is no difference between setting $\eta=10^{-6}$ and $\eta=0$. We have found the same in numerical simulations and on this basis recommend setting $\eta=0$.
Our procedure for modifying bootstrap critical values closely resembles the implementation of the bootstrap in the Linton-Song-Whang (LSW) test of stochastic dominance, proposed in LSW10. The discussion therein focuses on the matched pairs sampling framework, and our discussion in this section will do the same, though we note that the LSW test may also be applied, with obvious modifications, in a setting with independent differently-sized samples. The LSW test statistic for the null hypothesis of first-order stochastic dominance is
where $p=2$ and where $w$ is a user-specified weight function. To test higher-order stochastic dominance the empirical cdfs $\hat{F}_1$ and $\hat{F}_2$ are replaced with cumulative integrals thereof, and $w$ should be chosen to ensure that $\hat{L}$ is finite. To test first-order stochastic dominance one may simply set $w\equiv 1$; see W19. However if one were to replace $w(x)\,\mathrm{d}x$ with $\mathrm{d}\hat{F}_2(x)$ in (ref) then the resulting statistic with $p=2$ is a one-sided Cram\'{e}r-von Mises statistic. Moreover, if we instead set $p=1$ and if there are no ties in the observations $X_1^2,\dots,X_n^2$ then
see (ref). Thus, in the absence of tied observations, the one-sided WMW statistic (scaled by $\sqrt{2}$) is obtained by modifying the definition of the LSW test statistic in (ref) so that $p=1$ and so that $w(x)\,\mathrm{d}x$ is replaced by $\mathrm{d}\hat{F}_2(x)$.
The contact set relevant for the LSW test (of first-order stochastic dominance) is the set $D_0\subseteq\mathbb R$ comprised of all $x$ such that $F_1(x)=F_2(x)$. The LSW test relies on an estimate of $D_0$ given by
where $\kappa_n$ is a tuning parameter chosen such that $\kappa_n\to0$ and $\sqrt{n}\,\kappa_n\to\infty$ as $n\to\infty$. The critical value for the LSW test is computed from the simulated distribution, conditional on the data, of the bootstrap statistic $\tilde{L}^\ast$ defined by
A comparable expression for the modified bootstrap statistic $\tilde{S}^\ast$ (again scaled by $\sqrt{2}$) is
see (ref). There is thus a close connection between the modified WMW test and the LSW test.
Let $G$ be the set of all monotone bijections $g:\mathbb R\to\mathbb R$. The property of first-order stochastic dominance is invariant under $G$ in the following sense: if $X^1$ and $X^2$ are random variables such that $X^1$ first-order stochastically dominates $X^2$, then $g(X^1)$ first-order stochastically dominates $g(X^2)$ for every $g\in G$. Consider the effect on the LSW test and on the modified WMW test of applying some $g\in G$ to all of the sample observations $X_i^1$ and $X_i^2$. The outcome of the modified WMW test is unaffected by the application of $g$ because $\hat{R}$, $\hat{R}^\ast$ and $\hat{C}$ (the last of which is used to construct the variance estimates $\hat{V}_i$) all depend on the sample observations only through their pooled ranks. This property is not shared by the LSW test due to the use of the fixed measure $w(x)\,\mathrm{d}x$ rather than the empirical measure $\mathrm{d}\hat{F}_2(x)$ in (ref) and (ref). Applying a transformation $g\in G$ to all observations can change the outcome of the LSW test from rejection to non-rejection, or vice-versa. We illustrate this phenomenon in the empirical application reported in (ref). The non-invariance of the LSW test of first-order stochastic dominance to transformations of the data in $G$ means that the test does not satisfy Lehmann's principle of invariance---see LR22---and so may be susceptible to manipulation by unscrupulous practitioners. This concern is relevant only when testing first-order stochastic dominance, as higher orders of stochastic dominance are not invariant under $G$.
The estimated contact set $\hat{D}_0$ used to implement the LSW test is comprised of those points $x\in\mathbb R$ such that $\hat{F}_1(x)-\hat{F}_2(x)$ is between the uniform thresholds $\pm\,\kappa_n$; see (ref). The corresponding set $\hat{B}_0$ defined in (ref) is comprised of those points $u\in(0,1)$ such that $T_n^{1/2}(\hat{R}(u)-u)$ is above the variance-weighted threshold $-\tau_n\hat{V}^{1/2}_{\lceil n_2 u\rceil}$. An alternative definition of the modified bootstrap statistic $\tilde{S}^\ast$ that uses two-sided uniform thresholding as in the LSW test is
which should be compared to our preferred definition of $\tilde{S}^\ast$ in (ref). Close inspection of the proof of (ref) in (ref) shows that this result remains valid under the alternative definition of $\tilde{S}^\ast$, with only minor adjustments needed in the proof of (ref). We do not have a good theoretical justification for preferring (ref) to (ref), but recommend using (ref) on the basis of unreported numerical simulations. It may be surprising that the indicator function in (ref) is not chosen to exclude summands for which $T_n^{1/2}(\hat{R}(i/n_2)-i/n_2)\geq\tau_n \hat{V}_{i}^{1/2}$, as doing so would seem to mechanically improve power. However, if such summands are excluded then for each fixed $n$ a larger value of $\tau_n$ is needed to suitably control the rate of type I error, leaving the overall effect on power ambiguous.
To investigate the small sample properties of the modified WMW test we ran a number of Monte Carlo simulations. In each simulation we used $10^5$ Monte Carlo repetitions to compute rejection frequencies. In each of these repetitions we randomly generated an iid sample of pairs $\{(X_i^1,X_i^2)\}_{i=1}^{n}$, with $n$ varying from $25$ to $1000$ as described below. The samples were generated with $F_1$ normalized to be the uniform distribution on $[0,1]$, with $F_2$ chosen to obtain a desired P-P curve $R$ as described below, and with $C$ chosen to be the Gaussian copula with correlation parameter $\rho$ equal to $0$, $.25$, $.5$ or $.75$. Setting $\rho=0$ places us in the independent sampling framework with equally-sized samples, while setting $\rho>0$ places us in the matched pairs sampling framework with positive dependence within pairs. Bootstrap critical values were computed using $10^3$ bootstrap samples.
To provide a point of comparison we also report rejection frequencies obtained using the test of first-order stochastic dominance proposed in DH16. See W19 for a succinct treatment. The Donald-Hsu (DH) test has asymptotic properties comparable to those established for the modified WMW test in (ref). However the DH test is based on the one-sided Kolmogorov-Smirnov statistic rather than the one-sided WMW statistic, and uses a selective recentering method to modify bootstrap critical values rather than a contact set estimator. We use the DH test in our simulations rather than the LSW test because the non-invariance of the latter test to strictly increasing transformations of the data makes it easy to manipulate the design of simulations so that the test appears more or less powerful. Like the modified WMW test, the DH test of first-order stochastic dominance depends on the sample observations only through their pooled ranks.
In (ref) we report rejection frequencies at the least favorable case $R(u)=u$, i.e.\ $F_1=F_2$, for the modified WMW test and for the DH test with independent equally-sized samples. Rejection frequencies are reported at the nominal levels $\alpha=.05,.01$ and for sample sizes $n=25,50,100,200,500,1000$. The tuning parameter $\tau_n$ for the modified WMW test was set equal to the values $.5,.75,1,1.25,1.5,\infty$, with the value $\infty$ corresponding to the standard bootstrap critical value described in (ref). The tuning parameter for the DH test was set equal to the values $-.025,-.05,-.1,-.15,-.2,-\infty$, with the value $-\infty$ corresponding to one of the standard bootstrap procedures described in BD03. Note that the simulations reported in DH16 use a tuning parameter value of $-.1\sqrt{\log\log(n_1+n_2)}$, which decreases from $-.117$ to $-.142$ as the two equal sample sizes increase from $25$ to $1000$, thus falling within the range of tuning parameter values considered here.
(ref) shows that the modified WMW test delivers a rejection frequency that is generally close to, but slightly less than, the nominal level. Some over-rejection is observed with the smallest sample sizes and tuning parameters. On the other hand, the DH test delivers a rejection frequency that is generally close to, but slightly greater than, the nominal level. As expected, the rejection frequencies rise as the tuning parameters decrease in magnitude.
In (ref) we report rejection frequencies obtained with independent samples of size $n=500$ and nominal level $\alpha=.05$ for two parametric families of P-P curves satisfying the null hypothesis of first-order stochastic dominance. The top-left and bottom-left panels of (ref) display the two families of P-P curves. The family in the top-left has the parametrization $R_\gamma(u)=u^{1+\gamma}$ with $\gamma\geq0$, and the family in the bottom-left has the parametrization
with $\gamma\geq0$, where $\Phi$ is the standard normal cdf and $\Phi^{-1}$ the corresponding quantile function. Note that for the bottom-left family the contact set is always $[.5,1)$ when $\gamma>0$, so that we are on the boundary of the null, whereas for the top-left family the contact set is empty when $\gamma>0$, so that we are in the interior of the null. In both families we obtain the least favorable case when $\gamma=0$.
In the top-center and bottom-center panels of (ref) we plot the rejection frequencies for the modified WMW test as a function of the parameter $\gamma$ for the P-P curves in the top-left and bottom-left panels. Separate curves are plotted for each of the tuning parameter values $\tau_n=.5,.75,1,1.25,1.5,\infty$, with the curves for smaller tuning parameter values lying above those for larger values. We superimpose triangles on the curve for $\tau_n=\infty$ to improve visibility. An interesting pattern is apparent wherein the rejection frequencies with $\tau_n<\infty$ initially decrease as we raise $\gamma$ above zero -- essentially decreasing to zero in the top-center panel -- before rising back toward the nominal level of .05 as $\gamma$ becomes larger. The rejection frequencies with $\tau_n=\infty$ decrease smoothly to zero in both the top-center and bottom-center panels. The results in the top-center panel are particularly encouraging because these null configurations do not belong to the boundary of the null and therefore, as discussed at the end of (ref), the modified bootstrap critical values are not guaranteed by (ref) to produce limiting rejection frequencies no greater than the nominal level unless they are constrained not to fall below a small positive constant $\eta>0$. Similar to DH16, we have found that placing a small positive lower bound such as $\eta=10^{-6}$ on the modified bootstrap critical values does not meaningfully affect rejection probabilities, so we simply set $\eta=0$.
In the top-right and bottom-right panels of (ref) we plot the rejection frequencies obtained using the modified WMW test with $\tau_n=.75$ and the DH test with tuning parameter $-.139$ (the latter superimposed with asterisks) as a function of the parameter $\gamma$ for the P-P curves in the top-left and bottom-left panels. In the top-right panel we see that the rejection frequencies for the DH test do not share the interesting non-monotone behavior exhibited by the modified WMW test, and instead drop quickly to zero as we raise $\gamma$ above zero. On the other hand, we see in the bottom-right panel that the DH test is more successful than the WMW test in maintaining a rejection frequency close to the nominal level on the boundary of the null, at least for the family of P-P curves considered here.
In (ref) we report rejection frequencies at the least favorable case using the matched pairs sampling framework. These results may be compared directly to those reported in (ref) for the independent sampling framework. The only difference between the simulation designs is that with matched pairs a Gaussian copula with parameter $\rho=.25,.5,.75$ is used to generate dependence within pairs and, as described in (ref), the empirical copula is used to estimate the contact set. The results in (ref) are broadly similar to those in (ref), with the rejection frequencies using the modified WMW test tending to fall below the nominal level, and with the rejection frequencies using the DH test tending to fall modestly above the nominal level. The degree to which the rejection frequencies for the modified WMW test fall below the nominal level increases as the dependence between matched pairs increases.
In (ref) we report rejection frequencies with matched pairs ($n=500$, $\alpha=.05)$ corresponding to the two parametric families of P-P curves used to produce the results displayed in (ref). The left, center and right columns of panels in (ref) correspond to the correlation parameter values $.25$, $.5$ and $.75$, respectively. The top (bottom) row of panels in (ref) corresponds to the family of P-P curves displayed in the top-left (bottom-left) panel in (ref). Each panel of (ref) displays curves plotting the rejection frequency for three tests: the modified WMW test with $\tau_n=.75$ (superimposed with squares) and with $\tau_n=\infty$ (superimposed with triangles) and the DH test with tuning parameter $-.139$ (superimposed with asterisks). The results are very similar overall to those reported in (ref) for the independent sampling framework. The curves for the modified WMW test with $\tau_n=\infty$ and for the DH test are difficult to distinguish in the top row of panels.
In (ref) we report rejection frequencies obtained with independent samples of size $n=500$ and nominal level $\alpha=.05$ for two parametric families of P-P curves not satisfying the null hypothesis of first-order stochastic dominance in general. The two families are displayed in the top-left and bottom-left panels of (ref). The family in the top-left has parametrization $R_\gamma(u)=u^{1-\gamma}$ with $\gamma\in[0,1)$, and the family in the bottom-left has parametrization $R_\gamma(u)=\Phi(\mathrm{e}^\gamma\Phi^{-1}(u))$ with $\gamma\in\mathbb R$. In both families we obtain the least favorable case when $\gamma=0$, while for other values of $\gamma$ the null hypothesis is not satisfied. The crucial difference between the two families is that, when the null hypothesis is not satisfied, the graph of $R_\gamma$ is everywhere above the 45-degree line (except at the endpoints zero and one) with the top-left family but is partially below the 45-degree line with the bottom-left family.
The top-center and bottom-center panels in (ref) display the rejection frequencies for the modified WMW test with tuning parameter values $\tau_n=.5,.75,1.1.25,1.5,\infty$. The curve for $\tau_n=\infty$, which corresponds to the standard bootstrap critical value, has triangles superimposed. In both panels we see the rejection frequencies increase from approximately $\alpha=.05$ to one as $\gamma$ moves away from zero, reflecting the consistency of the tests. In the top-center panel the curves plotted for different tuning parameter values are indistinguishable. In the bottom-center panel there is a clear separation between the curves, with the rejection frequencies using the standard bootstrap critical value well below the rejection frequencies using the modified bootstrap critical value. The very different behavior displayed in the two panels can be understood by observing that the modified bootstrap statistic $\tilde{S}^\ast$ differs from the standard bootstrap statistic $\hat{S}^\ast$ only when the P-P plot $\hat{R}$ falls below the 45-degree line by more than some threshold depending on the tuning parameter. This happens very infrequently in the simulations generating the rejection frequencies in the top-center panel because here we have $R_\gamma(u)>u$ for all $u\in(0,1)$ when $\gamma>0$; but frequently in those generating the rejection frequencies in the bottom-center panel because here we have $R_\gamma(u)<u$ for all $u\in(.5,1)$ when $\gamma<0$, and $R_\gamma(u)<u$ for all $u\in(0,.5)$ when $\gamma>0$.
In the top-right and bottom-right panels of (ref) we plot the rejection frequencies for the modified WMW test with $\tau_n=.75$ and the DH test with tuning parameter $-.139$ against one another. Rejection frequencies for the DH test are superimposed with asterisks. We see that the rejection frequencies for the two tests are similar. A slight power advantage for the modified WMW test is observed in the top-right panel, and a slight power advantage for the DH test in the bottom-right panel.
(ref) shows how the rejection frequencies plotted in (ref) are affected when there is positive dependence between paired observations. The left, center and right columns of panels in (ref) correspond to $\rho=.25,.5,.75$ respectively, while the top (bottom) row of panels corresponds to the family of P-P curves displayed in the top-left (bottom-left) panel in (ref). Rejection frequencies are plotted for the modified WMW test with $\tau_n=.75$ (superimposed with squares) and with $\tau_n=\infty$ (superimposed with triangles), and for the DH test with tuning parameter $-.139$ (superimposed with asterisks). In both rows of panels we see that the power for all three tests improves as $\rho$ increases. In the top row of panels we see that the rejection rates for the modified WMW test with $\tau_n=.75$ and with $\tau_n=\infty$ are indistinguishable, as they were with independent samples in (ref). The slight power advantage of these tests over the DH test widens as $\rho$ increases, becoming quite substantial with $\rho=.75$. In the bottom row of panels we see that the slight power advantage of the DH test observed with independent samples erodes as the correlation parameter increases. The rejection frequencies for the modified WMW test with $\tau_n=.75$ and the DH test are nearly indistinguishable when $\rho=.5$ or $\rho=.75$.
We illustrate the modified WMW test of first-order stochastic dominance with an application to Canadian family income distributions. Our dataset is the same as the one used in BD03 and DH16, so our results may be compared directly to those reported there. The data consist of before-tax and after-tax family incomes in the Canadian Family Expenditure Survey for the years 1978 and 1986. There are 8526 observations in the former year and 9470 in the latter. We test two null hypotheses: that the income distribution in 1986 first-order stochastically dominates the income distribution in 1978 (abbreviated as $1986\gtrsim1978$), and that the income distribution in 1978 first-order stochastically dominates the income distribution in 1986 ($1978\gtrsim1986$), reporting results for before-tax and after-tax income separately.
In (ref) we display the P-P plots for before-tax and after-tax incomes. The two plots are similar. We see that they lie neither entirely above nor entirely below the 45-degree line, indicating the possibility that both null hypotheses may be rejected. In particular, the P-P plots show that the estimated 30th percentile of the income distribution, either before-tax or after-tax, declined slightly between 1978 and 1986. The population P-P curves must be everywhere no greater than the 45-degree line if $1986\gtrsim1978$, or everywhere no less than the 45-degree line if $1978\gtrsim1986$.
In (ref) we report the p-values of tests of first-order stochastic dominance. The first two rows of p-values, for the Barrett-Donald (BD) and DH tests, are taken directly from Tables 6 and 7 in DH16, and confirmed to be correct (up to random variation intrinsic to the bootstrap) in independent calculation. The DH test uses a tuning parameter value of $-.15\approx-.1\sqrt{\log\log(n_1+n_2)}$. To obtain a tuning parameter value for the modified WMW test, we observe that setting $\tau_n=.5$ was effective in controlling the false rejection rate at the least favorable case in the simulations reported in (ref) for the largest sample size of 1000 observations per sample. Since $.5\approx.35\sqrt{\log\log(n_1+n_2)}$ when $n_1=n_2=1000$, and $.53\approx.35\sqrt{\log\log(n_1+n_2)}$ when $n_1=8526$ and $n_2=9470$, we set $\tau_n=.53$ in this illustration. Similar to what is done in DH16, to reduce computational burden we computed one-sided WMW statistics using a discrete approximation based on 2000 grid points spread evenly over the domain of the P-P plots. We calculated p-values using $10^5$ randomly generated bootstrap samples.
We see in (ref) that our modification to bootstrap critical values leads to a meaningful reduction in the p-values obtained using the one-sided WMW statistic. With the before-tax and after-tax incomes, the null hypothesis $1986\gtrsim1978$ is not rejected at the 5% nominal level using the standard bootstrap critical value, but is rejected using the modified bootstrap critical value. The same is true for the null hypothesis $1978\gtrsim1986$ using the after-tax incomes, and in this case the modified bootstrap critical value also leads to rejection at the 1% (or even 0.1%) nominal level. Both the standard and modified bootstrap critical values lead to rejection of the null hypothesis $1978\gtrsim1986$ at the 1% nominal level using the before-tax incomes.
The p-values obtained using the one-sided WMW statistic with the standard bootstrap critical value are larger than those obtained using the one-sided Kolmogorov-Smirnov statistic with the standard bootstrap critical value (i.e., the BD test), and the p-values obtained using the one-sided WMW statistic with the modified bootstrap critical value are larger than those obtained using the one-sided Kolmogorov-Smirnov statistic with the modified bootstrap critical value (i.e., the DH test). This therefore appears to be a case where the particular way in which first-order stochastic dominance is violated is more easily detected with the one-sided Kolmogorov-Smirnov statistic than with the one-sided WMW statistic.
The final three rows of (ref) report p-values for the LSW test with $\kappa_n=.07\approx3(2T_n)^{-1/2}\log\log(2T_n)$. The first of the three rows shows the p-values obtained using the original income data, while the second and third show those obtained when the square or square-root function is applied to incomes. As discussed in (ref), the LSW test of first-order stochastic dominance is not invariant to strictly increasing transformations of the data, so applying the square or square-root function to incomes can affect the p-value obtained. Indeed, the final three rows of (ref) differ substantially, and the outcome of the LSW test of $1986\gtrsim1978$ switches from rejection to non-rejection at the $5\%$ nominal level after applying the square function to after-tax incomes. Transforming incomes using the square or square-root function does not affect the p-values reported in the first four rows of (ref) because the corresponding tests are invariant to strictly increasing transformations of the data.
We have shown in this article that the bootstrap can be used to implement tests of first-order stochastic dominance based on the one-sided WMW statistic in settings involving either independent sampling or matched pair sampling. Moreover, we have shown that modifying the bootstrap so that it incorporates an estimate of the contact set can lead to a large improvement in power while maintaining control of false rejection rates. Our procedure can be applied alongside existing tests using other statistics, such as the LSW test and the DH test, in routine applications of stochastic dominance testing. As tests based on different statistics may differ in their propensity to detect different violations of stochastic dominance, it can be informative to compare the outcome of multiple tests.
Versions of the LSW and DH tests suitable for testing second-order stochastic dominance are provided in LSW10 and DH16. It would be useful to extend the methods proposed in this article so as to likewise obtain a test of second-order stochastic dominance. A complicating factor is the central role played by the P-P process and bootstrap P-P process in our asymptotic arguments, and our reliance on the representation of the one-sided WMW statistic as a functional of the P-P plot. The P-P plot is not a suitable tool for assessing second-order stochastic dominance because such dominance is not invariant under strictly increasing transformations and so cannot be expressed as a property of the P-P curve. It has recently been shown in LL25 that second-order stochastic dominance can instead be expressed as a property of the Lorenz P-P curve. The Lorenz P-P curve is a nondecreasing function from $[0,1]$ into $[0,1]$ obtained by composing the inverse of one unscaled Lorenz curve with another unscaled Lorenz curve, and is everywhere no greater than the 45-degree line precisely when second-order stochastic dominance is satisfied. Bootstrap tests of second-order stochastic dominance proposed in LL25 are similar in spirit to those proposed in BD03 and do not make use of a contact set estimator. It seems likely that the power of the Lando-Legramanti tests could be improved by incorporating contact set estimation, similar to what has been done here and in LSW10. We leave this as a topic for future research.
(ref) is proved in (ref). (ref) are proved in (ref). (ref) are proved in BK26. (ref) establish auxiliary lemmas respectively concerning the consistency of the modified bootstrap and the regularity of asymptotic distributions.
With (ref) in hand, the proof of (ref) is a straightforward application of the delta-method for Hadamard directionally differentiable maps. The latter property is weaker than Hadamard differentiability because the derivative is not required to be linear. See Definition 2.1 and Theorem 2.1 in FS19 for the definition of Hadamard directional differentiability and a corresponding statement of the delta-method. The idea originates in S90,S91 and D93. See also BM15 and K16.
Note that the directional derivative $\mathcal H'_\theta$ is not linear unless $\theta(u)\neq u$ for a.e.\ $u\in[0,1]$. Thus $\mathcal H$ is not Hadamard differentiable everywhere on $L^1[0,1]$. It is Hadamard directionally differentiable everywhere on $L^1[0,1]$.
The following lemma establishing good behavior of the modified bootstrap test statistic $\tilde{S}^\ast$ will be used in the proof of (ref).
We will prove (ref) by applying Theorem 3.2 in FS19, which provides general conditions under which bootstrap procedures based on an estimated contact set are well-behaved. The proof will also make use of the following two lemmas.
A technical obstacle to proving (ref) is the need to demonstrate that the quantile function for the asymptotic distribution of the bootstrap test statistic is continuous and strictly increasing at $1-\alpha$, where $\alpha\in(0,1/2)$. (ref) supplies this ingredient. To prove (ref) we use a theorem on convex functionals of Gaussian processes in DLS98 to show that the only possible mass point for the asymptotic distribution is zero, with any remaining mass spread smoothly over the entire nonnegative halfline. A separate argument, given in the proof of (ref), shows that the mass at zero can be no greater than $1/2$. Closely related econometric applications of the relevant theorem in DLS98 appear in, for instance, A02, CF05, ACF06, CH06 and CHT07.
The next lemma is used to prove (ref).
We conclude by proving (ref). The proofs of both results make use of the following lemma, which may be compared to Lemma 10.11(i) in K08.