Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
93,932 characters · 24 sections · 64 citation commands
Tests for almost stochastic dominance
In this section we present the main ideas of this work, as well as the general context in which they are included and the motivation of our proposals.
Stochastic orders, also known as stochastic dominance (SD) rules in the economic literature, are partial order relations in the set of probability measures; see Shaked-Shanthikumar-2006 and Levy-2016. As SD rules allow comparing and ranking random elements (according to some specific criterion), they are useful in many different disciplines. Economics is one of the areas in which stochastic orders are especially relevant and appear more frequently. There is a large body of literature on this topic in fields such as inequality analysis, actuarial science, portfolio insurance, operations management, inventory or financial optimization.
An advantage of (global) stochastic comparisons between two variables over (partial) comparisons through summary statistics is that, if two distributions are stochastically ordered---with respect to a certain relation---we can derive many important consequences regarding the underlying distributions. Typically, SD implies an ordering among many features or summary measures of the underlying distributions simultaneously. For instance, if the Lorenz ordering holds, then several inequality measures (such as the Gini index) of the involved variables are ordered in the same way; see Arnold-Sarabia-2018. Moreover, many dominance rules are in fact integral stochastic orders (see Muller-1997), that is, they are defined by comparing expected values of functions (in a certain class) of the variables. In this way, SD is linked to expected utility theory in the context of decision theory under uncertainty; see Fishburn-1984.
Given the practical importance of checking, based on observed data, whether it is (statistically) reasonable to assume that two random variables, \(X_1\) and \(X_2\), are ordered, in the literature there are various hypothesis tests of the type
where `$\preceq$' denotes the stochastic order of interest; see, e.g., Anderson-1996, Barrett-Donald-2003, Zheng-2002, Berrendero-Carcamo-2011, Barrett-Donald-Bhattacharya-2014 and Sun-Beare-2021.
SD rules are partial orders and, hence, not every pair of distributions can be ranked. To overcome this important drawback, Leshno-Levy-2002 introduced the notion of almost stochastic dominance (ASD), which refers to the situation where strict SD does not hold but the region where the dominance rule is violated is relatively small. In risk analysis, ASD describes the situation where “most” decision makers would prefer one uncertain prospect over another one, which is equivalent to eliminating some extreme or pathological utility functions from the ordering condition. One advantage of relaxing the strict dominance criterion by its almost counterpart is that the latter is a more flexible concept and allows ordering more pairs of distributions. A second advantage of ASD is that its definition (see Section (ref)) intrinsically depends on the degree to which the strict stochastic dominance, say $X_1 \preceq X_2$, is unfulfilled. This quantity, usually called violation ratio and denoted by the parameter $\epsilon$ (with $0\leq\epsilon\leq 1$), has an intuitive interpretation: small values of $\epsilon$ mean that the variables are very close to being ordered; Levy-2012. Conversely, the larger the \(\epsilon\), the more distant of strict dominance is the relationship between the variables.
In what follows of this introduction we state the questions of interest at the core of this work using only the above intuitive definition of ASD. For this purpose, let us denote by `$\preceq^\epsilon$' the ASD rule with parameter $\epsilon >0$. We will see in Section (ref) that, if $X_1 \preceq^{\epsilon_1} X_2$ holds for a certain $\epsilon_1>0$, then $X_1 \preceq^{\epsilon_2} X_2$, for all $\epsilon_2>\epsilon_1$. Therefore, in this context it is important to determine the minimum violation ratio (MVR), that is, the smallest possible value of $\epsilon$ for which the two variables are almost ordered, given by
In the literature there is yet no clear distinction between the concept of the MVR $\epsilon_0$ and that of a single $\epsilon$ for which ASD holds. The value of $\epsilon$ can be preselected by the practitioner as the acceptable degree of non-fulfillment in the SD condition. For instance, using solely the sample information, we might be interested in checking whether the ASD holds with $\epsilon=0.05$ when actually the (unknown) population MVR is $\epsilon_0=0.01$.
In practice, after observing samples from two variables, the statement that two distributions are ordered cannot be proved statistically, as one would desire. That is, interchanging the roles of the null and alternative hypotheses in (ref) results in an ill-posed problem. Indeed, given a pair of ordered distributions, we can generally find pairs of non-ordered distributions arbitrarily close to the initial ones and hence the null and alternative hypotheses are indistinguishable; see Ermakov-2017. However, ASD provides a more flexible setting, since the collection of pairs of almost ordered distributions is much larger than the set of pairs of strictly ordered variables. Thus, we can implement statistical tests of the form
where the ASD assumption is placed in the alternative hypothesis. Observe that rejection of the null hypothesis in (ref) means that there is statistical evidence that the two variables are almost ordered, while, in the usual tests of type (ref), non-rejection of \(H_0\) only means there is not enough evidence against it. This difference between the conclusions of (ref) and (ref) makes the ASD test relevant in many applications where the aim is to establish the (almost) ordering between the variables; see Huang-Kan-Tzeng-Wang-2021.
As noted by Huang-Kan-Tzeng-Wang-2021, the MVR \(\epsilon_0\) in (ref) is a key point in the analysis of ASD relations. Further, the definition of the MVR (and its distinction from the parameter $\epsilon$) allows us to write the problems that we consider here in a clear and concise way. For instance, given a fixed value of $\epsilon>0$, test (ref) is equivalent to $H_0: \epsilon_0>\epsilon$ versus $H_1:\epsilon_0\le \epsilon$. In this work we are interested in the following hypothesis tests:
Observe that test (a) in (ref) essentially coincides with the ASD test in (ref) with a standard and better-posed formulation of null and alternative. Test (b) corresponds to $H_0:\ X_1 \preceq^\epsilon X_2$ versus $H_1:\ X_1 \npreceq^\epsilon X_2$, a weaker version of the usual SD test in (ref). The last test in (c) amounts to obtaining a confidence interval for the MVR.
In this work we present a general framework to treat, in a unified way, the problems that we consider for several stochastic orders simultaneously, as well as their “almost” versions. Our main contributions are the following:
In Section (ref) we define the notions of $s$-SD and $s$-ASD. In Section (ref) we introduce the 2DSD index and enumerate its main properties, in particular how it characterizes both strict SD and ASD. Section (ref) introduces the plug-in estimators of the 2DSD index and the MVR. Under suitable conditions, we prove the strong consistency and determine the asymptotic distributions of the normalized estimators. We also show the (a.s.) consistency of the corresponding bootstrap estimators. In particular, we provide a consistent resampling scheme to approximate the asymptotic distribution of any statistic defined through the $\mathrm{L}^1$-norm. In Section (ref) we propose bootstrap rejection regions, based on the estimated MVR, for the ASD tests in (ref). The procedure is intuitive and fast to compute for the usual orders even for moderately large data sets. The performance of the tests is checked through simulations. In Section (ref) we analyse two real data sets with the proposed methodology. The detailed proofs of the results are collected in the online Appendix.
A stochastic order relationship between two random elements $X_1$ and $X_2$ is typically established by the pointwise comparison of a real function $s:D_s\to\ensuremath{\mathbb{R}}$ in the two populations, in the sense that we can define the comparison rule:
where $s_1$ and $s_2$ denote the function $s$ in populations \(X_1\) and \(X_2\), respectively, and \(D_s\) is a suitable set, usually the domain of $s$ or the support of the variables. If (ref) holds, we say that $X_2$ stochastically dominates $X_1$ with respect to $s$, and denote it by $X_2$ $s$-SD $X_1$.
We refer to $s$ as the target function of the stochastic ordering as it represents the characteristic that we want to compare in both populations: survival time, distribution of wealth, salaries, accumulated risk, total size, variability, concentration, dependency, etc. The function \(s\) corresponding to the population \(X\) usually depends on the probability distribution of \(X\) in a way determined by the specific stochastic order under consideration. The framework of $s$-SD encompasses many well-known SD rules.
The definition of SD in terms of a function \(s\) becomes more flexible when the condition `\(s_1(t)\ge s_2(t)\)' in ((ref)) is allowed to fail in a “small” subset of \(t\)'s. From now on we assume that any pair \((s_1,s_2)\) fulfils that the difference \(s_1-s_2\in \mathrm{L}^1\), the space of functions $s:D_s \to \ensuremath{\mathbb{R}}$ endowed with the usual $\mathrm{L}^1$-norm, \(\Vert s\Vert =\int_{D_s} |s(t)|\, \text{\rm d} t\) for \(s\in \mathrm{L}^1\). Generalizing previous definitions of ASD (Leshno-Levy-2002, Tsetlin-et-al-2015, and Zheng2018), for $\epsilon \ge 0$, we consider the following ASD rule:
If (ref) holds, we say that $X_2$ $\epsilon$-almost $s$-stochastically dominates $X_1$, written \(X_2\) $s$-ASD($\epsilon$) \(X_1\). Leshno-Levy-2002 defined ASD for the first and second-degree SD. Subsequently, Tsetlin-et-al-2015 extended this idea to the $N$th-degree SD (\(N>2\)) and Zheng2018 considered the Lorenz ordering.
The degree of noncompliance of the stochastic ordering is quantified by the parameter $\epsilon$, called the {\em relative area violation} by Levy-2012 and the {\em violation ratio} of the ordering condition by Huang-Kan-Tzeng-Wang-2021. Observe that when the functions $s_1$ and $s_2$ are smooth enough (right-continuous, for instance), a value \(\epsilon=0\) is equivalent to proper SD as in ((ref)). This is fulfilled for the common choices of \(s\) in the most frequently used stochastic orders; see Examples (ref)--(ref). Note that it always holds that $X_1 \preceq_s^\epsilon X_2$ for $\epsilon\ge 1$. Thus, only values $\epsilon<1$ make sense. Further, it is always satisfied that $X_1\preceq_s^{0.5} X_2$ or $X_2\preceq_s^{0.5} X_1$.
To carry out the tests in (ref), we propose to use a bidimensional index to characterize both strict SD and ASD defined in ((ref)) and ((ref)), respectively. There are several reasons behind the need for a two-dimensional characterization of SD and ASD instead of using a simpler one-dimensional statistic. First, it is convenient that any unidimensional index intending to measure relative dominance (of one variable with respect to another one) takes values on a reference (fixed) bounded symmetric interval. Assume that some normalization has allowed this interval to be $[-1,1]$, in such a way that the extremes of the interval correspond to the strict SD cases, e.g., the index is 1 when \(X_1\preceq_s X_2\) and $-1$ when \(X_2\preceq_s X_1\). Secondly, the index should be continuous in the sense that if there are two random sequences \((X_{1,n_1})\), \( (X_{2,n_2})\) converging to $X_1$ and $X_2$, as \(n_1,n_2\to\infty\), respectively, in some suitable mode of stochastic convergence, then the values of the dominance index between \(X_{1,n_1}\) and \(X_{2,n_2}\) should converge to the index between \(X_1\) and \(X_2\). This continuity property is essential in practice, when using random samples from the populations to construct an empirical counterpart of the index. However, these two requirements are incompatible for a one-dimensional index. For example, we can think of a sequence \(X_{1,n}\) converging to \(X\) as \(n\to\infty\) such that \(X_{1,n}\preceq_s X\) for all \(n\). The sequence of indices corresponding to \((X_{1,n},X)\) would be \(1\) for all \(n\), but the index of \((X,X)\) would be naturally \(0\).
However, it is possible to construct an easy-to-handle, two-dimensional index characterizing both the strict SD (ref), as well as the $\epsilon$-ASD ((ref)). Indeed, our following proposal satisfies, among other desirable properties, the two requirements described before: symmetry and continuity. We define the {\em 2DSD (2-dimensional Stochastic Dominance) index}, associated to the ordering $\preccurlyeq_s$ in (ref), as
This definition is inspired by the inequality index introduced in Baillo-Carcamo-Mora-2022. The first component of the index (ref) provides a signed value on the fulfillment of the $s$-SD between the two variables. The second component is a measure of the “distance” between the two distributions in terms of the function $s$. In Sections (ref) and (ref) we use the new index (ref) to estimate the MVR in (ref) and to define suitable critical regions for the ASD tests (ref).
For the examples of target functions \(s\) considered in Section (ref), we derive the expression of the 2DSD index.
{\bf Example (ref) {\rm (First stochastic dominance index)}.} Let $F_1$ and $F_2$ be the distribution functions of two integrable variables $X_1$ and $X_2$, respectively. The 2DSD index for the first stochastic order is given by
where \(\mu_1\) and \(\mu_2\) are the expectations of \(X_1\) and \(X_2\), respectively, and
is the celebrated $\mathrm{L}^1$-Wasserstein distance; see Villani-2009.
{\bf Example (ref) {\rm (Second stochastic dominance index)}.} Assuming that $\text{\rm E} X_1=\text{\rm E} X_2$ and that $\text{\rm E}(X_1^2)$ and $\text{\rm E}(X_2^2)$ are finite, we have
where $\operatorname*{Var}(X)$ stands for the variance of $X$ and $d_2(X_1,X_2)$ is the Zolotarev (ideal) metric of order $2$; see Rachev-2013. Under the same assumptions, it turns out that the 2DSD index for the stop-loss order (Example (ref)) has the same expression as (ref). Likewise, according to Remark (ref), assumptions $\text{\rm E} X_1=\text{\rm E} X_2$ and $\text{\rm E}(X_1^2) , \text{\rm E}(X_2^2) < \infty$ can be dispensed with if the variables $X_1$ and $X_2$ have compact support.
{\bf Example (ref) {\rm (Lorenz dominance index)}.} Let $X_1$ and $X_2$ be positive and integrable random variables with Lorenz curves $\ell_1$ and $\ell_2$, respectively. The 2DSD index associated to the Lorenz order is given by
where \( G(\ell)= 1-2 \| \ell \| \) is the Gini index of the variable $X$ with Lorenz curve $\ell$. This was essentially the index introduced by Baillo-Carcamo-Mora-2022.
The following result collects the main properties of the 2DSD index $\mathcal{I}_s$ defined in (ref).
Figure (ref) provides a graphical outline of the previous properties. Parts (ref) and (ref) in Proposition (ref) show that the 2DSD index \(\mathcal{I}_s (X_1,X_2)\) characterizes both proper $s$-SD and $s$-ASD($\epsilon$). Equation (ref) is crucial to derive a suitable test statistic for the ASD hypothesis tests (ref). Property (ref) provides the statistical consistency of the potential estimators of the index whenever the underlying empirical functions (playing the role of $s_{1,n_1}$ and $s_{2,n_2}$) converge in $\mathrm{L}^1$ to their population versions. This $\mathrm{L}^1$ convergence has to be established for the particular choice of the target function $s$; see Section (ref).
In this section we carry out a deep analysis of the asymptotic properties of the natural estimators of the 2DSD index in (ref) and the associated MVR in (ref) for the $s$-ASD. We also provide conditions that guarantee the consistency of the bootstrap estimators of these quantities. In the proofs (see the Supplemental Material) we use empirical process theory, results on differentiability of functionals and the extended functional delta method.
First, we propose a natural estimator for the 2DSD index given in (ref). For $j=1,2$, we denote by $\{X_{j,i}\}_{i=1}^{n_j}$, $n_j \in\mathbb{N}$, a random sample from $X_j$. For simplicity, we assume that both samples are mutually independent. However, convergence results similar to those in this section can be obtained when we observe “matched pairs", $ \{ (X_{1,i},X_{2,i}) \}_{i=1}^n $, drawn from a bivariate distribution $(X_1,X_2)$ with copula $C$ satisfying that its maximal correlation is strictly less than one; see Beare-2010. As pointed out in Barrett-Donald-Bhattacharya-2014 and Sun-Beare-2021, this second setting is more reasonable when we have one sample of individuals and two measures of the variable of interest.
The estimation of the 2DSD index $\mathcal{I}_s$ depends on the selected target function $s$. In the usual practical situations, as in Examples (ref)--(ref), $s$ is expressed in terms of the distribution function, $F$, or the quantile function, \(F^{-1}\). In such cases, we propose a plug-in estimator of $\mathcal{I}_s$ based on the empirical measures of the samples. For $j=1,2$, we consider the empirical distribution functions of the samples
where $1_A$ stands for the indicator function of the set $A$. The corresponding empirical quantile functions are $\hat{F}^{-1}_j(x)=\inf\{ y \ge 0 : \hat F_j (y)\ge x \}$ ($0<x<1$). The plug-in estimator \(\hat s\) of \(s\) is obtained by replacing \(F\) (resp., \(F^{-1}\)) with its empirical counterpart $\hat{F}$ (resp., \(\hat{F}^{-1}\)). Then, the plug-in estimator of the index $\ensuremath{\mathcal{I}}_s$ is the empirical 2DSD index given by
For instance, for the Lorenz order (Example (ref)) the empirical Lorenz curves are
where $\hat \mu_j = \sum_{i=1}^{n_j} X_{j,i}/n_j$ are the sample means. Then, the empirical 2DSD index is \(\hat \ensuremath{\mathcal{I}}_\ell (X_1,X_2) = (\Vert \hat\ell_1\Vert-\Vert \hat\ell_2\Vert,\Vert \hat\ell_1-\hat\ell_2\Vert)\).
The next proposition shows that the strong consistency of the empirical index $\hat \ensuremath{\mathcal{I}}_s (s_1,s_2)$ in (ref) follows from the $\mathrm{L}^1$ almost sure convergence of the estimators $\hat s_1$ and $\hat s_2$ to their population counterparts.
The next result provides, for some usual target functions, simple conditions under which the \(\hat s_j\) are $\mathrm{L}^1$ a.s.\ consistent and the assumptions of Proposition (ref) are fulfilled. These results usually follow from Glivenko–Cantelli and the dominated convergence theorems. Corollary (ref) considers 2DSD indices associated with the first, second and Lorenz dominance, but similar results can be derived for higher-degree stochastic orders.
The next theorem provides necessary conditions so that the (normalized) plug-in estimator of the 2DSD index, $\hat\ensuremath{\mathcal{I}}_s$ in (ref), converges in distribution. This convergence follows from the weak convergence of the process $\hat{s}_1-\hat{s}_2$ in the space $\mathrm{L}^1$ together with an extended version of the functional delta method; see Shapiro-1991. We use the arrow `$\rightsquigarrow$' to denote the weak convergence of probability measures in the underlying space.
In all the examples in this work the normalizing sequence in (ref) and (ref) is $r_{n_1,n_2} = \sqrt{{n_1n_2}/{(n_1+n_2)}}$. However, this rate might be different in other cases than those considered here. The following corollary provides necessary and sufficient conditions for the limit distribution in (ref) to be bivariate normal.
Observe that $\{s_1=s_2\}$ is the contact set between $s_1$ and $s_2$. The fact that this set has measure zero means that the two target functions are essentially different.
To apply Theorem (ref) or the subsequent Corollary (ref), we first need to check the convergence in (ref). This has to be established for each particular choice of $s$. In the following proposition we give conditions under which this weak convergence is fulfilled for the examples considered above. For the sake of clarity, we enumerate below all the conditions needed, although not all of them are required simultaneously in each example.
Assumption (ref) means that the sampling scheme is (weakly) balanced. Assumption (ref) amounts to saying that the variables $X_j$ belong to the Lorentz space $\mathcal{L}^{2,1}$ (see Grafakos) and it is equivalent to the convergence of the empirical process in $\mathrm{L}^1$ (see del Barrio). Condition $\Lambda_{2,1}(X_j)<\infty$ is slightly stronger than $\text{\rm E} X_j^2<\infty$; it holds when $\text{\rm E} X_j^{2+\epsilon}<\infty$, for some $\epsilon>0$. Assumption (ref) says that $X_j$ is in the Lorentz space $\mathcal{L}^{4,2}$, which is equivalent to the convergence of the integrated empirical process in $\mathrm{L}^1$; see Carcamo. Again, this condition is a bit stronger than $\text{\rm E} X_j^4<\infty$. These integrability conditions are optimal to obtain the convergence of the underlying processes. Other works on this subject assume that variables have compact support (see Huang-Kan-Tzeng-Wang-2021), a much more demanding requirement. Finally, Assumption (ref) is necessary to derive the convergence of the quantile process in $\mathrm{L}^1$ through the Hadamard differentiability of the inverse map plus the convergence of the empirical process in $\mathrm{L}^1$; see for example Sun-Beare-2021 and Kaji-2018. This is used to derive the weak $\mathrm{L}^1$ convergence of the empirical Lorenz processes.
In the following proposition we show the convergence in $\mathrm{L}^1$ of the processes associated to the usual target functions. From now on, for $j=1,2$, $\mathbb{B}_{F_j}$ denote two independent \(F_j\)-Brownian bridges.
Under the assumptions of Lemma 2.1 in Sun-Beare-2021, the result in Proposition (ref) (ref) can be established in the space of continuous functions on $[0,1]$ with the supremum norm, which entails the convergence in $\mathrm{L}^1$. Note that all the limiting process in Proposition (ref) are centered Gaussian processes. Therefore, Corollary (ref) provides necessary and sufficient conditions so that the (normalized) estimator of the 2DSD index in the usual examples is asymptotically normal (with zero mean). Specifically, in Examples (ref), (ref) and (ref) consider $D_s$ as the union of the supports of the two distributions and assume that $\lambda\in(0,1)$ in Assumption (ref). In Example (ref), let $D_s=[0,1]$. In this situation, the limit in (ref) has a centered bivariate normal distribution if and only if the contact set $\{s_1=s_2\}$ has zero Lebesgue measure.
The literature on estimation of the MVR \(\epsilon_0\) in (ref) is scarce. For the usual order, Bali-2009 and Levy-2012 outline an estimator based on the definition of ASD (ref) and the plug-in principle. Huang-Kan-Tzeng-Wang-2021 propose an estimator for \(\epsilon_0\) in a very specific context of insurance deductibles, for the second-degree SD and for compactly supported variables.
Here, we propose a consistent estimator of $\epsilon_0>0$ based on the empirical 2DSD index. Observe that, by equation (ref) in Proposition (ref) (ref), we can express the MVR as \(\epsilon_0 = g(\ensuremath{\mathcal{I}}(s_1,s_2))\) with \(g(x,y)=(1-x/y)/2\). Thus, a natural estimator of the MVR $\epsilon_0$ is
When \(\Vert s_1-s_2\Vert>0 \), then $g$ is differentiable at $\ensuremath{\mathcal{I}}(s_1, s_2)$, so the strong consistency of (ref) follows from Proposition (ref) and Corollary (ref). Moreover, whenever \(\epsilon_0>0\), the asymptotic distribution of (a standardized version of) $\hat{\epsilon}$ can be derived from Theorem (ref) by application of the (first-order) delta method. Specifically, under the assumptions of Theorem (ref) and when $\epsilon_0>0$, some simple computations show that
From the theoretical point of view, the previous consistency and asymptotic results ensure that the estimators (ref) and (ref) are statistically reasonable and provide the precise rates of convergence. However, the corresponding limits are not distribution-free: they depend on the underlying (and usually unknown) populations $F_1$ and $F_2$. Moreover, even in the case where the limiting variable is bivariate normal (see Corollary (ref)), simulations show that convergence to the limit is relatively slow. Therefore, to make inferences on the 2DSD index \(\ensuremath{\mathcal{I}}_s\) in (ref) and the MVR \(\epsilon_0\) in (ref), we resort to a bootstrap scheme; see Section (ref).
Here, we give general conditions that guarantee the strong consistency of the bootstrap estimators of the 2DSD index and MVR. From Theorem (ref) and Corollary (ref), we see that the asymptotic distribution of the 2DSD index changes radically depending on whether the contact set $\{s_1=s_2\}$ has measure zero or not. This phenomenon has been observed previously. For tests of stochastic dominance as in (ref), the limiting distribution of the test statistic in Linton-Song-Whang-2010 is discontinuous and hence not locally uniformly regular; see Bickel-1993. For this reason we have divided this section into two parts depending on whether the contact set $\{s_1=s_2\}$ has zero or positive measure.
The set of crossing points between two target functions usually consists of a finite (or empty) collection of points. Therefore, in the usual practical situations, the contact set has zero Lebesgue measure and the results of this part are particulary useful. For each $j=1,2$, the bootstrap sample $\{X_{j,i}^*\}_{i=1}^{n_j}$ is obtained by redrawing independently from the empirical distribution $\hat F_j$ in (ref). The estimator of \(s_j\) derived from this bootstrap sample is denoted by \(\hat s_j^*\).
The proof of Theorem (ref) follows from the full Hadamard differentiability of the involved mappings and Fang-Santos. As in the previous results, the strong $\mathrm{L}^1$-consistency of the bootstrap process in ((ref)) has to be established for the selected \(s\). The following proposition shows the strong consistency (in $\mathrm{L}^1$) of the bootstrap processes underlying the estimated indices in Examples (ref)--(ref). The proof of this result can be derived from Gine-Zinn, where general conditions for the consistency of the bootstrap process in a Banach space are established. However, in the online Appendix we include a simpler alternative proof of this result for the considered examples. In particular, together with Theorem (ref), we conclude that the corresponding bootstrapped 2DSD indices and MVR's are consistent a.s. The precise definitions of the involved processes can be found in the online Appendix.
Theorem (ref) shows that the standard bootstrap is no longer valid to estimate the asymptotic distribution of the 2DSD index whenever $\{s_1=s_2\}$ has positive measure. In such a case, the $\mathrm{L}^1$-norm (second component of the 2DSD index in (ref)) is non-differentiable at $s_1-s_2$; see Lemma 1 in the online Appendix or Carcamo. In this situation, Fang-Santos provide different consistent modifications of the bootstrap scheme. The idea is that, although the $\mathrm{L}^1$-norm is not fully differentiable, it is Hadamard directionally differentiable and the directional derivative can be appropriately estimated to obtain a consistent alternative.
As suggested in Fang-Santos, it is convenient to employ an estimator of the directional derivative based on its analytical expression to construct a consistent resampling scheme. In the following result we use the explicit expression of the derivative of the $\mathrm{L}^1$-norm (given in the second component of the limit in (ref)) together with the ideas in Linton-Song-Whang-2010 about enlargements of the contact set. As a result, we obtain a consistent resampling method for statistics defined through the $\mathrm{L}^1$-norm. As this result might be of interest in other situations, we state the next proposition in general terms: $\theta_0\in \mathrm{L}^1$, $\hat{\theta}_n$ is an estimator of $\theta_0$, and $\hat{\theta}_n^*$ its bootstrapped version.
The quantity $a_n$ in Theorem (ref) plays the role of a step size or tuning parameter in the approximation of the derivative. Usually, $r_n = \sqrt{n}$, and in this case typical choices of $a_n$ are $a_n = c/\log n$ or $c/\log\log n$, with $c>0$ a constant.
Theorem (ref) provides a consistent resampling scheme for all the tests in (ref). For instance, following the notation in Fang-Santos, from (ref), the test (ref)(a) can be restated as
where \(\theta_0=s_1-s_2\) and
As the first summand of $\phi$ in (ref) is linear, we can use Theorem (ref) to obtain a suitable bootstrap estimator of $\phi(\theta_0)$. Specifically, for \(\hat\theta_n=\hat s_1-\hat s_2\) and \(r_n=\sqrt{n_1n_2/(n_1+n_2)}\), we consider
with $(a_n)$ being any sequence of real numbers such that $a_n\downarrow 0$ and $a_n r_n\uparrow\infty$. By Theorem (ref), the distribution of \(\hat\phi_n'(r_n(\hat\theta_n^*-\hat\theta_n))\), conditional on the sample, is a consistent estimator of \(r_n(\phi(\hat\theta_n)-\phi(\theta_0))\).
After obtaining the asymptotic distribution of the empirical index (Theorem (ref) and Corollary (ref)), it seems natural to use these results to derive a rejection region for the ASD tests in (ref). However, simulations show that the distribution of the standardized estimated 2DSD index is only approximately Gaussian for very large sample sizes. Then, for moderate sample sizes, asymptotic distributions could differ significantly from the corresponding finite sample distributions. Theorems (ref) and (ref) together with Proposition (ref) guarantee that suitable bootstrap rejection regions for the tests in (ref) are consistent. We have to distinguish between the two cases indicated in Section (ref).
Assuming that the set \(\{s_1=s_2\}\) has zero measure, we propose a standard bootstrap algorithm to construct rejection regions for the tests (ref), at a significance level \(\alpha\), where \(\epsilon\) is a pre-specified violation ratio of interest:
Step 1. Extract \(B\) bootstrap samples from each of the original samples.
Step 2. For each \(b=1,\ldots,B\) and the corresponding pair of bootstrap samples (ref), compute the bootstrapped empirical inequality index \(\hat{\mathcal I}^{*b}\) and MVR \(\hat\epsilon_0^{*b}\).
Step 3. Denote by \(\hat\epsilon_0^{*(\alpha)}\) the \(\alpha\)-quantile of the bootstrapped MVR's \(\hat\epsilon_0^{*1},\ldots,\hat\epsilon_0^{*B}\). We reject the null hypothesis in the test given in
In the general case when \(\{s_1=s_2\}\) might not have zero measure, we can use the modified bootstrap scheme obtain in Section (ref) based on the estimator of the directional derivative in (ref). This implementation is slightly more complicated than that of Case 1, because the testing procedure in Case 2 depends on fixing in advance the value of \(\epsilon\). For example, we reject the null hypothesis in (ref) (which is equivalent to ((ref)a)) when the test statistic \(r_n\phi(\hat\theta_n)\) exceeds \(c_{1-\alpha}\), the \((1-\alpha)\)-quantile of the corresponding asymptotic distribution when \(\phi(\theta_0)=0\). Then, the critical value \(c_{1-\alpha}\) can be estimated by \(\hat c_{1-\alpha}\), the \((1-\alpha)\)-quantile of \(\hat\phi_n'(r_n(\hat\theta_n^*-\hat\theta_n))\), conditional on the sample, where $\hat\phi_n'$ is in (ref).
We finally observe that the proof of Theorems (ref) and (ref) (see the online appendix) ensure that Fang-Santos can be applied and this implies the pointwise size control of the testing procedures described in this section.
We present the results of a Monte Carlo study for the test ((ref)a). In Table (ref) we indicate the four scenarios considered, along with the MVR \(\epsilon_0\) and the stochastic orders with respect to which the test is carried out (graphical representations of the distribution functions or the Lorenz curves appear in the online appendix). In Scenario 1, the distribution functions \(F_1\) and \(F_2\) intersect only once in the right tail and fulfill \(X_2 \preceq_{s}^\epsilon X_1\) for \(\epsilon\geq\epsilon_0\). In Scenario 2, \(F_1\) and \(F_2\) cross once too, but near the origin, with a smaller MVR. We also perform the ASD test with respect to the Lorenz order (Scenario 3), simulating samples from two generalized beta distributions of the second kind (GB2). The scale parameter of the GB2 distribution is chosen so that \(\text{\rm E}(X_j)=1\), for \(j=1,2\). Finally, Scenario 4 is designed to illustrate Case 2, as the distribution functions \(F_1\) and \(F_2\) coincide over the interval $[0,1]$; see the specific expression of \(F_2\) in the appendix.
The technical details of the simulation study are as follows. The significance level is \( \alpha=0.05 \). In every scenario we simulate \(N=500\) Monte Carlo samples (of the same size \(n_1=n_2=n\)) from each of the two models (populations 1 and 2). To examine the effect of the sample size, in Case 1 (resp. Case 2) we selected \(n\) as 1000, 5000, 10000, 20000 and 50000 (resp., 1000, 5000 and 10000). For each Monte Carlo run we drew \(B=2000\) bootstrap samples and computed \( \hat\epsilon_0^{*(0.95)} \) in Case 1. With respect to Case 2, the sequence \(a_n\) used in (ref) is \(a_n = c \, \log(n_1n_2/(n_1+n_2))/r_n\) (similar to the choice in Linton-Song-Whang-2010), with \(r_n=\sqrt{n_1n_2/(n_1+n_2)}\). In Case 2 the power of the testing procedure was evaluated for a grid of values of \(\epsilon\) in (0,1) and for different values of \(c\), namely 0.1, 0.01, 0.002 and 0.001. Sampling from the composite distribution in Scenario 4 was carried out with the R library mistr (Sablica-Hornik). For Scenarios 1 and 4, Figure (ref) displays, for the values of \(\epsilon\) in the horizontal axis, the power of the test ((ref)a), in Scenario 1 (Case 1) for the different values of the sample size \(n\) and in Scenario 4 (Case 2) for \(n=5000\) and the values of \(c\). The power curves corresponding to Scenarios 2 and 3 are analogous to those of Scenario 1 and appear in the online appendix, as well as the power curves for Scenario 4 when \(n=1000\) and \(n=10000\).
\color{black}
The simulations confirm the consistency of the bootstrap testing procedure as previously observed: the asymptotic rejection probability is zero in the interior of the null and the asymptotic power is equal to 1 under the alternative hypothesis. Figure (ref)(a) shows that, as the sample size \(n\) increases, the power approaches 1 when the alternative hypothesis is true and remains bounded by 5% when the null is true. In the power functions for Scenarios 2 and 3 (online appendix), the significance level is very close to or below the 5% nominal level for all values of the sample size \(n\). Scenario 2 (resp., 3) corresponds to a crossing point of the distribution (resp., Lorenz) functions near the origin. Since the crossing point in these two scenarios lies in a high density zone, there is more sample information available to better estimate that point and the MVR \(\epsilon_0\) when strict domination fails. This is the opposite of Scenario 1, in which the two distribution functions \(F_1\) and \(F_2\) cross in a low density zone, and the estimated MVR depends on the integral of the difference \(F_1-F_2\) from the crossing point to infinity. Hence, it takes a larger \(n\) for the ecdf of Figure (ref) (a) to approach the 5% significance level under \(H_0\). Taking \(B>2000\) bootstrap samples does not significantly improve the simulation results. Regarding Scenario 4 (where the contact set has positive measure), we see that the bootstrap ASD testing procedure devised for Case 2 has a good performance and it is fairly robust with respect to the value of \(c\). Observe that a “tightly” estimated contact set (\(c=0.01\) for \(n=1000\), \(c=0.01\), 0.002 or 0.002 for \(n=5000\) and \(c=0.002\) for \(n=10000\)) leads to the power function adjusting better to $0.05$ in \(\epsilon=\epsilon_0\) and to 1 for \(\epsilon>\epsilon_0\).
Here the bootstrap ASD testing procedures introduced in Section (ref) are illustrated with real data. We consider two stochastic orders, the usual one (see Example (ref)) and the Lorenz order (Example (ref)). The results of this section, together with the simulations of Section (ref), support the use of this general estimating and testing procedures for the minimum violation ratio \(\epsilon_0\) in practical, real data settings.
The gender pay gap, the difference in wages between women and men, has been analysed in various ways, but it is generally accepted that females systematically earn less than males. It is natural to compare the salaries between both groups via their empirical distribution functions. For instance, Maasoumi-Wang-2019 carry out first---and second---order SD tests as (ref) of female vs. male wage distributions in United States for all the years in the period 1976--2013. In 12 of these years the authors were unable to establish any of the two types of dominance and they suggest that, in this case, the claim that women are worse off than men is not robust. However, a graphical representation of the pairs (women and men) of empirical distribution functions for some years gives insight into the relative position of these functions when strict dominance cannot be statistically accepted: there is a crossing of the distribution functions at the left tail (utmost lowest wages), while the cdf of male wages generally lies under the female one elsewhere.
In this context the ASD test ((ref)a) is very useful. The MVR \(\epsilon_0\) in (ref) provides a quantitative measure of the ASD degree and, consequently, of the gender gap. A MVR close to \(0\) in the test of almost dominance of male wages over female ones (\(H_1\)) is an objective answer supporting the conclusion of gender inequality. To illustrate the procedure we use microdata from the 2018 Wage Structure Survey in Spain, available at the web page of the Spanish INE (Instituto Nacional de Estad\'{\i}stica). The variable under consideration is the salary per working hour; alternatively, we also examined the annual gross salary. There are many factors that can be taken into account: full- and part-time contracts, age cohorts,\dots For the majority of these groups, the empirical distribution function of men's salaries is clearly below or equal that of women, and generally there is a huge difference between both functions except in the lowest percentiles. Figure (ref) is an example of this situation: we display the distribution functions of female (\(n_1=94168\)) and male (\(n_2=122558\)) hourly salaries. The men's cdf is less than the women's for all the sample values. The empirical distributions are ordered according to the first order and, thus, the associated empirical 2DSD index lies on the diagonal (see Suppl. Figure 3a): \( \hat{\mathcal{I}}_F (X_1,X_2) =(2.4218,2.4218)\) and the estimated MVR \(\hat\epsilon_0\) is 0. For a significance level $\alpha=0.05$, the bootstrap ASD test (Case 1), with \(B=2000\) bootstrap samples, yields \( \hat\epsilon_0^{*(0.95)}=4\cdot 10^{-6}\). Thus, there is a strong statistical evidence supporting that men's wages almost dominate women's in the usual order and, given the extremely small value of \( \hat\epsilon_0^{*(0.95)}\), that the range of salaries where strict domination fails is negligible in comparison with the rest.
Considering only full-time workers between 20 and 29 years old (\(n_1=6100\) and \(n_2=9249\)), we obtain an interesting situation where \(\hat F_2\) is generally below \(\hat F_1\), even though they cross in two occasions, in the left tail and near the 90th percentile, and are fairly close; see Figure (ref). As expected, the empirical 2DSD index for the usual stochastic order is very close to the diagonal: \( \hat{\mathcal{I}}_F (X_1,X_2) =(0.9850,0.9864)\) . For a level \(\alpha=0.05\) and \(B=2000\) bootstrap samples (Suppl. Figure 3b), the test ((ref)a) with the alternative hypothesis that men's wages ASD women's in the first-SD yields \(\hat\epsilon_0^{*(0.95)}=0.047\) (Case 1). Hence the area between the cdf's where the strict dominance fails is, with 95% confidence, less than 5% of the overall area between them, which clearly indicates a high degree of first-order ASD of men's earnings over women's.
To illustrate the performance of the ASD test described in Case 2 of Section (ref) we selected the hourly salaries of individuals employed in the public sector, whose empirical distribution functions for the two sex apparently coincide on the interval of lowest wages (Supplementary Figure 4). There are \(n_1=19719\) women and \(n_2=15834\) men. For a significance level \( \alpha=0.05 \), \(B=2000\) bootstrap samples, \(a_n = c \, \log(n_1n_2/(n_1+n_2))/r_n\), with \(r_n=\sqrt{n_1n_2/(n_1+n_2)}\), we carried out the test (ref)(a) on a grid of \(\epsilon\)'s and for different choices of \(c\). For each \(c\), we determined the minimum \(\epsilon\) for which \(H_0\) is rejected (Supplementary Table 1). These values of \(\epsilon\) are all less than 0.0016, which again supports the conclusion that, even in the public sector, stochastic dominance of male wages fails to hold in a very small interval of salaries. \color{black}
When the variable \(X\) is the income of a household, the distribution is frequently described by the Lorenz curve. The MVR \(\epsilon_0\) and the 2DSD index \(\mathcal{I}_\ell (X_1,X_2)\) associated to the Lorenz order measure the relative inequality between two income distributions, \(X_1\) and \(X_2\). We consider the ASD test ((ref)a), with respect to the Lorenz order, to check whether income distribution in population 2 is almost more unequal than in population 1 (alternative hypothesis). Using income microdata from Spain, available at the INE database, we determine the equivalised disposable income (the variable \(X\)).
We observe that the ASD test can be used to perform a cross-temporal comparison of a random variable holding fixed the rest of conditions, for instance, the location (Spain in this case). Population 1 is then the variable of interest (Spanish income) in a reference year (here 2008). Population 2 is the same variable in a posterior year. The test ((ref)a) checks the almost Lorenz dominance of the income \(X_2\) over the income \(X_1\) in 2008. The year 2008 was chosen as reference because it marked the start of a deep financial and economic crisis in Spain and several other countries. Although the recession that took place in Spain ended in 2014, income distribution in later years is still more unequal than in 2008. The estimated 2DSD index and the quantile \(\hat\epsilon_0^{*(0.95)}\) derived from the test reflect the changes in income inequality between the two years; see also Baillo-Carcamo-Mora-2022.
Depending on the year posterior to 2008, the result of the ASD test is actually very different. For instance, when \(X_2\) is the income in 2016 (see Figure (ref)(a)), the Lorenz curve of 2008 is above or equals that of 2016 for every percentile except an interval contained in $(0.99,1)$. The situation is a clear candidate for Lorenz ASD. Indeed, the estimated 2DSD index (Figure (ref)(b)) is \(\hat\ensuremath{\mathcal{I}}_\ell = (0.006773,0.006777)\) and the ASD test carried out with \(B=2000\) bootstrap samples gives \(\hat\epsilon_0^{*(0.95)}=0.082\), with $\alpha=0.05$. In other words, in 2016 the income distribution in Spain was almost more unequal than in 2008. However, in 2018 the economic and financial situation in Spain had substantially improved. The Lorenz curve of 2018 is above that of 2008 in the interval (0.68,1). The value of \(\hat\epsilon_0^{*(0.95)}=0.73\) (see Supplementary Figure 5) indicates there is no almost dominance in this case. In Figure (ref) we display the estimated MVR and the minimum \(\epsilon\) for which there is evidence that a year almost Lorenz dominates 2008. The (almost) increase in income inequality with respect to 2008 is evident for 2011, 2012 and from 2014 to 2017. Note that, even if \(\hat\epsilon_0\) is very small in 2010, the Lorenz curves of 2008 and 2010 are so near each other that the estimated index \(\hat\ensuremath{\mathcal{I}}_\ell\) is very close to the origin and the bootstrapped indices are distributed all around the origin in the grey region of Figure (ref) (see also Supplementary Figure 5). This explains why the \(\hat\epsilon_0^{*(0.95)}\) corresponding to 2010 is so large compared to the estimated MVR.
We are very grateful to the Editor, Associate Editor and two reviewers for all the comments on the first version of the paper. We especially appreciate that we were encouraged to consider the case where the contact set has positive measure, which resulted in Theorem (ref).
Supported by the Spanish Agencia Estatal de Investigaci\'{o}n through projects PID2019-109387GB-I00 (A.B. and J.C.), PID2021-124195NB-C32 (C.M-C.) and the Severo Ochoa Programme CEX2019-000904-S (C.M.-C.). C.M-C.\ has also been supported by the ERC Advanced Grant 834728.