EconBase
← Back to paper

Testing Conditional Stochastic Dominance at Target Points

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

99,830 characters · 19 sections · 49 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing Conditional Stochastic Dominance at Target Points

spacing{1.1} \begin{abstract} This paper introduces a novel test for conditional stochastic dominance at specific values of the conditioning covariates, referred to as target points. The test is relevant for analyzing treatment effects, discrimination, and income inequality. We propose a Kolmogorov--Smirnov-type test statistic based on induced order statistics from independent samples. A key feature is its data-independent critical value, which eliminates the need for resampling techniques such as the bootstrap, as well as kernel smoothing or parametric assumptions. Instead, the test relies on a tuning parameter to select relevant observations. We establish that the induced order statistics converge to independent draws from the true conditional distributions and that the test is asymptotically of level $\alpha$ under weak regularity conditions. Monte Carlo simulations confirm the strong finite-sample performance of our method. \end{abstract}

KEYWORDS: stochastic dominance, regression discontinuity design, induced order statistics, rank tests, permutation tests.

JEL classification codes: C12, C14.

\thispagestyle{empty}

Introduction

The concept of stochastic dominance has long been central to numerous areas of applied research. This paper examines a specific aspect of stochastic dominance: testing conditional stochastic dominance (CSD) at specific values of the conditioning covariate, referred to as target points. Such conditional comparisons are crucial in many contexts, including evaluating treatment effects in social programs within a regression discontinuity design, analyzing economic disparities across demographic groups, and investigating potential discrimination in decision-making processes.

Unconditional stochastic dominance methods, which analyze entire distributions, have been extensively studied and widely applied in the literature, with foundational contributions dating back to hodges:1958 and mcfadden:1989 and more recent developments in abadie:2002, barrett/donald:2003, Linton/maasoumi/whang:2005, and linton/song/whang:2010, among others. However, in many empirical settings, the primary interest lies in dominance conditional on a subset of the population defined by specific characteristics or values of a conditioning variable. For instance, in regression discontinuity designs, the nature of the methodology necessitates comparing outcome distributions conditional on the cutoff of the running variable (donald/hsu/barrett:2012, shen/zhang:16, goldman/kaplan:2018, qu/yoon:2019). Likewise, in wage discrimination studies, researchers may seek to compare wage distributions across demographic groups while controlling for observed skill levels (becker:1957, canay/etal:2024, bharadwaj/deb/renou:24).

The primary goal of this paper is to test whether the conditional cumulative distribution function (CDF) of one variable stochastically dominates that of the other at specific values of a conditioning variable. Formally, we consider the null hypothesis \[ H_0: F_Y(t | z) \leq F_X(t | z) \quad \text{for all } (t, z) \in \mathbf{R} \times \mathcal{Z}~, \] against the alternative that there exists some \((t, z)\) for which the reversed inequality holds strictly. Here, \( F_Y(\cdot | z) \) and \( F_X(\cdot | z) \) represent the conditional CDFs of the random variables \( Y \) and \( X \), respectively, given \( Z = z \). Importantly, we focus on situations where the set of target values, $\mathcal{Z}$, is not the entire support of $Z$, but rather a finite collection of points (including the case of a singleton). To address this testing problem, we propose a novel procedure that leverages induced order statistics based on independent samples from \( Y \) and \( X \). Our test statistic, a Kolmogorov--Smirnov-type measure, captures the maximal deviation between the empirical CDFs of the two samples, conditional on observations near the target points. Crucially, the critical value we propose is derived in a deterministic, non-data-dependent manner once a tuning parameter is accounted for, ensuring computational simplicity.

Our contributions in this paper are both methodological and theoretical. First, we introduce a novel test for CSD at target points, particularly suited for settings where researchers seek to compare distributions conditional on covariates at specific values of the conditioning covariates. The proposed test exploits induced order statistics, leveraging observations closest to the target conditioning point to construct empirical CDFs that form the basis of our test statistic. Unlike traditional methods, our approach neither relies on kernel smoothing nor imposes parametric assumptions on conditional distributions. However, it requires a tuning parameter, which serves a role analogous to bandwidth selection in nonparametric estimation.

Second, we establish the asymptotic validity of our test under two alternative asymptotic frameworks. In the first framework, the numbers of effective “local” observations, $q_y$ and $q_x$, are held fixed as the sample size $n \to \infty$. In this setting, we show that the test statistic converges to a limiting experiment in which the induced order statistics behave as independent draws from the true conditional CDFs at the target point. This convergence yields a feasible critical value that remains valid without relying on resampling methods such as the bootstrap. Our regularity conditions in this fixed-$q$ framework are mild: the conditional CDFs at the target point may contain finitely many discontinuities; both $Y$ and $X$ may be continuous, discrete, or mixed; and we require only that the conditional CDFs $F_Y(t \mid z)$ and $F_X(t \mid z)$ be continuous in $z$ at the target point, uniformly in $t$. This is a weaker condition than the smoothness assumptions, such as twice differentiability in $z$, typically imposed in nonparametric methods. In the second framework, the numbers of effective observations $q_y$ and $q_x$ are allowed to diverge as $n \to \infty$. Building on recent results on convergence rates for induced order statistics bugni/canay/kim:26b, we derive an explicit rate at which the rejection probability of our test approaches its nominal level as a function of $q_y$, $q_x$, and $n$. This second framework complements the fixed-$q$ results by demonstrating that the test continues to control size when the number of effective observations increases, and it provides practical guidance for the data-dependent selection of the tuning parameters $q_y$ and $q_x$.

Third, we show that the proposed critical value aligns with the one obtained from a permutation-based approach when the random variables $Y$ and $X$ are both continuous, thus establishing a natural connection between our method and the broader literature on permutation-based inference. To the best of our knowledge, this result provides the first formal justification for the validity of permutation-based inference in testing unconditional stochastic dominance. We demonstrate that the critical value of our test cannot be improved when both $Y$ and $X$ are continuous. However, this result does not extend to the case when either $Y$ or $X$ is discretely distributed. For this latter case, we introduce a refined critical value, which is typically smaller than the default one we propose and only a function of the support points for $Y$ and $X$. This refinement enhances power relative to the default critical value, though it comes at the cost of increased computational complexity.

Finally, we examine the finite-sample performance of our test through Monte Carlo simulations and present an empirical application to illustrate its implementation. Both exercises suggest that the data-dependent rule we propose for selecting the key tuning parameters performs well in practice.

Our work contributes to the extensive literature on stochastic dominance testing, building on seminal contributions such as anderson:1996, davidson/duclos:2000, abadie:2002, barrett/donald:2003, Linton/maasoumi/whang:2005, and linton/song/whang:2010. These studies examine the null hypothesis of unconditional stochastic dominance and predominantly rely on asymptotic arguments and resampling techniques. Our approach to testing CSD differs in that, in the limit experiment, the conditional testing problem simplifies to a finite-sample unconditional testing problem. In the context of CSD testing, prior research---including delgado/escaciano:2013, gonzalo/olmo:2014, chang-lee-whang-2015-EJ, and andrews-shi-2017-JOE---evaluate stochastic dominance over a range or the entire support of a continuous conditioning variable. Our work diverges from this literature by testing CSD at specific target points. A distinct line of research, including donald/hsu/barrett:2012, shen/zhang:16, goldman/kaplan:2018, and qu/yoon:2019, studies CSD within regression discontinuity designs, where dominance is defined conditional on cutoffs. All of these methods assume continuity of conditional distributions, limiting their applicability in empirical settings featuring discrete mass points. Our method relaxes this constraint and accommodates distributions with finitely many discontinuities. The resulting test is novel, computationally simple, and valid across a broader class of distributions.

Our work closely aligns with the well-established literature on testing the equality of two continuous distributions. Foundational contributions by gnedenko/korolyuk:1951, korolyuk:1955, and blackman:1956 established that the finite-sample distribution of the two-sample one-sided Kolmogorov--Smirnov test statistic is pivotal under the null, deriving closed-form expressions under various simplifying assumptions. Later research by hodges:1958, hajek/sidak:1967, and durbin:1973 developed methods to approximate this finite-sample distribution. Although our null hypothesis differs, our critical value coincides with the corresponding quantile of this distribution. Notably, hodges:1958 and mcfadden:1989 proposed that the two-sample one-sided Kolmogorov--Smirnov test, originally designed for testing equality of two distributions, could be adapted to test stochastic dominance under continuity assumptions, though without formal proof. We provide a rigorous justification for this claim.

The remainder of the paper is organized as follows. Section (ref) defines the testing problem and introduces notation. Section (ref) presents the induced order statistics that form the basis of our procedure and introduces the proposed test. Section (ref) establishes its asymptotic validity under two alternative asymptotic frameworks. Section (ref) discusses several refinements and extensions, including results on rates of convergence, a data-dependent rule for selecting the tuning parameters, and a refinement of the test that increases power when $Y$ and $X$ are discrete. Importantly, Section (ref) also provides formal results on the validity of permutation tests for stochastic dominance with continuously distributed random variables, offering novel arguments for the validity of permutation-based inference in this setting. Section (ref) evaluates the finite-sample performance of our test through Monte Carlo simulations. Finally, Section (ref) concludes. All proofs are contained in the Appendix.

Testing problem

Let $(Y,Z)$ and $(X,Z)$ denote random pairs, each supported on $\mathbf{R}\times\mathbf{R}$. Let $P_Y$ and $P_X$ denote the joint distributions of $(Y,Z)$ and $(X,Z)$, respectively. For $t\in\mathbf{R}$ and $z$ in the support of $Z$, define the conditional distribution functions \[ F_Y(t | z) := P_Y\{Y \le t \mid Z=z\} \qquad\text{and}\qquad F_X(t | z) := P_X\{X \le t \mid Z=z\}~. \]

We are interested in testing the null hypothesis:

equation[equation omitted — 115 chars of source]

versus the alternative hypothesis:

equation*[equation* omitted — 117 chars of source]

where $\mathcal{Z} = \{z_1, \dots, z_L\}$ is a finite set of target values of the conditioning variable. The case where $L = 1$ and $Z$ is a continuously distributed random variable is both simpler and particularly relevant in empirical applications. To minimize notational clutter, we focus on this case for most of the remainder of the paper, with Section (ref) addressing the case where $L > 1$.

To test this hypothesis, we assume the analyst observes two independent samples:

equation[equation omitted — 115 chars of source]

We refer to the first sample as the $Y$-sample, which consists of $n_y$ i.i.d.\ draws from $P_Y$, and to the second sample as the $X$-sample, which consists of $n_x$ i.i.d.\ draws from $P_X$. Independence of the samples implies that the data-generating distribution is the product measure \[ P := P_Y \otimes P_X~. \] All probability and expectation statements in what follows are taken with respect to $P$ unless otherwise noted. We denote by $\mathbf P$ the class of data-generating distributions $P$ satisfying the regularity conditions imposed later, and let

equation[equation omitted — 97 chars of source]

denote the subset of distributions $P\in \mathbf P$ satisfying the null hypothesis in (ref).

Test based on induced ordered statistics

Let the observed data be those given in (ref). Let $(q_y,q_x)$ be two small positive integers (relative to $(n_y,n_x)$) and consider the point $z_{0}\in \mathcal Z$. The test we propose is based on the following two samples:

itemize• The $q_y$ values of $\{Y_i:1\le i\le n_y\}$ associated with the $q_y$ values of $\{Z_i:1\le i\le n_y\}$ closest to $z_{0}$, and • The $q_x$ values of $\{X_j:1\le j\le n_x\}$ associated with the $q_x$ values of $\{Z_j:1\le j\le n_x\}$ closest to $z_{0}$.

To define these samples formally, we introduce $g$-order statistics for the conditioning variable $Z$, where $g(Z) := |Z - z_{0}|$; see reiss:89 and kaufmann/reiss:92. For any two values \( z, z' \in \mathcal{Z} \), we define the ordering \(\le_{g}\) as follows: \[ z \le_{g} z' \quad \text{if and only if} \quad g(z) \le g(z')~. \] This defines a $g$-ordering on the set $\mathcal{Z}$. The $g$-order statistics \( Z_{g,(i)} \) are then the values of \( Z \) ordered according to this criterion: \[ Z_{g,(1)} \le_{g} Z_{g,(2)} \le_{g} \cdots \le_{g} Z_{g,(n)}~. \] If there are ties in the $g$-ordering, they can be resolved arbitrarily—for example, by relying on the original sample order.

We then take the values of $\{Y_i:1\le i\le n_y\}$ associated with the $q_y$ smallest $g$-ordered statistics of $Z$ in the $Y$-sample, denoted by

equation[equation omitted — 80 chars of source]

That is, $Y_{n_y,[j]}=Y_k$ if $Z_{g,(j)}=Z_k$ for $k=1,\dots,q_y$. Similarly, we take the values of $\{X_i:1\le i\le n_x\}$ associated with the $q_y$ smallest $g$-ordered statistics of $Z$ in the $X$-sample, denoted by

equation[equation omitted — 80 chars of source]

The random variables in (ref) and (ref) are referred to as induced order statistics or concomitants of order statistics; see david/galambos:74,bhattacharya:74,canay/kamat:18. Intuitively, we view these samples as independent samples of \(Y\) and \(X\), conditional on \(Z\) being `close' to \(z_{0}\). A key feature of our test is that it relies solely on these induced order statistics. Letting \(n := \min\{n_y,n_x\}\) and \(q := q_y + q_x\), the effective (pooled) sample is

equation[equation omitted — 138 chars of source]

It is also important to note that the first \(q_y\) elements of \(S_n\) are associated with the \(Y\)-sample, while the remaining \(q_x\) elements come from the \(X\)-sample.

Having defined the induced order statistics, we now define our test statistic as

equation[equation omitted — 210 chars of source]

where the empirical CDFs are

equation*[equation* omitted — 186 chars of source]

The test statistic in (ref) is a two-sample one-sided Kolmogorov--Smirnov (KS) test statistic, see hajek/sidak/sen:99, and the test we propose rejects the null hypothesis in (ref) when $T(S_n)$ exceeds a critical value, defined next.

To introduce the critical value, let $\alpha\in(0,1)$ be given and $\{U_j: 1\le j\le q\}$ be a sequence of uniform random variables i.i.d., that is, $U_j\sim U[0,1]$. Define $\Delta(u)$ as

equation[equation omitted — 149 chars of source]

and

equation[equation omitted — 179 chars of source]

The test we propose for the null hypothesis in (ref) is

equation[equation omitted — 81 chars of source]

We reiterate that (ref) corresponds to our test with a single target point ($L = 1$). Section (ref) extends this framework to the general case with multiple target points ($L > 1$). Furthermore, in addition to our default critical value in (ref), we provide a refined critical value tailored for discrete $Y$ and $X$ with finite support in Section (ref).

The choice of the two-sample one-sided KS test statistic in (ref) is crucial for accommodating cases where \( Y \) and \( X \) are discrete or mixed. While one could construct analogues of our test in (ref) using alternative statistics commonly employed in the stochastic dominance literature---such as one-sided versions of the Cram\'er--von Mises or Anderson--Darling statistics---our asymptotic validity result does not generally extend to these alternatives. See the discussion following Theorem (ref).

remarkFrom a computational perspective, a notable feature of the test $\phi(S_n)$ is that the supremum over $\mathbf{R}$ in (ref) can be replaced by a maximum over the set $\{1, \dots, q_y\}$. This simplification arises because the KS test statistic increases only when evaluated at points corresponding to the $Y$-sample, which are the first $q_y$ elements in the vector $S_n$. Consequently, the supremum in $c_{\alpha}(q_y, q)$ can also be replaced with a maximum over $\{1, \dots, q_y\}$. This allows us to compute the critical value $c_{\alpha}(q_y, q)$ with arbitrary accuracy through simulation.
remarkThe values of $(q_y, q_x)$ are tuning parameters chosen by the researcher. Section (ref) shows that the test is asymptotically valid when $(q_y, q_x)$ are either held fixed or allowed to diverge slowly with the sample size. In Section (ref), we develop data-dependent guidelines for selecting $(q_y, q_x)$ that are compatible with these asymptotic requirements, and we assess their practical performance through simulations in Section (ref) and in the empirical application in Section (ref).
remarkA natural application of our test is in the context of a sharp regression discontinuity design, where an outcome $\tilde{Y}$ depends on a running variable $\tilde{Z}$, and the point of interest is the cutoff $z_0$. The observed sample in this case is $\{(\tilde{Y}_i, \tilde{Z}_i): 1 \le i \le n\}$, and the two samples needed for implementing our test are the following: \begin{align*} &\{(Y_i, Z_i): 1 \le i \le n_y\} := \{(\tilde{Y}_i, \tilde{Z}_i): 1 \le i \le n_y such that Z_i \le z_0\} and \\ &\{(X_i, Z_i): 1 \le i \le n_x\} := \{(\tilde{Y}_i, \tilde{Z}_i): 1 \le i \le n_x such that Z_i > z_0\} . \end{align*} Importantly, this formulation shows that the point $z_0$ can be either an interior or a boundary point in its support. We apply our test to this setting in our Monte Carlo simulations in Section (ref) and our empirical application in Section (ref).

Asymptotic framework and formal results

In this section, we derive the asymptotic properties of the test in (ref) using two alternative asymptotic frameworks. The first one requires $q := q_y+q_x$ to be fixed as $n\to \infty$ and represents a finite sample situation where the effective number of observations used by the test is too small to credibly invoke approximations that require a “large” value of $q$. The second framework requires $q\to \infty$ slowly as $n\to \infty$, and represents a finite sample situation where the effective number of observations used by the test is large enough to invoke approximations for “large” $q$.

Asymptotic results for fixed q

In this section, we examine the asymptotic properties of the test in (ref) within a framework where $q := q_y+q_x$ is fixed and $n:=\min\{n_y,n_x\}\to \infty$. We first derive the asymptotic properties of induced order statistics in (ref), and then present our main theorem.

We start by deriving a result on the induced order statistics collected in the vector $S_n$ in (ref). To do so, we make the following assumptions.

assumptionFor any $\varepsilon >0$ and $z\in \mathcal{Z}$, $P\{Z\in (z-\varepsilon ,z+\varepsilon )\}>0$.
assumptionFor any $z\in \mathcal{Z}$ and sequence $ z_{k}\to z$, $\sup_{t\in \mathbf{R} }|F_{Y}(t|z_{k})-F_{Y}(t|z)|\to 0$ and $\sup_{t\in \mathbf{R} }|F_{X}(t|z_{k})-F_{X}(t|z)|\to 0$.

Assumption (ref) requires that the distribution of $Z$ is locally dense at each of the points in $\mathcal{Z}$. This includes the case where $Z$ has a mass point at $z\in \mathcal Z$. Assumption (ref) is a smoothness assumption required to guarantee that conditioning on observations close to $z$ is informative about the distribution conditional on $Z=z$.

theoremLet Assumptions (ref) and (ref) hold. Then, \begin{equation} S_n \stackrel{d}{\to} S=(S_1,\dots,S_{q}) , \end{equation} where for any $s:=(s_1,\dots,s_{q})\in \mathbf{R}^{q}$, the random vector $S$ satisfies \begin{equation*} P\{S\le s\} = \prod_{j=1}^{q_y} F_Y(s_j|z_{0}) \cdot \prod_{j=q_y+1}^{q} F_X(s_j|z_{0}) . \end{equation*}

Theorem (ref) is a special case of Theorem (ref) in the appendix with $L=1$, which generalizes canay/kamat:18 to accommodate multiple conditioning values. It establishes that the limiting distribution of the induced order statistics in the vector $S_n$ is such that the elements of the vector, denoted by $S$, are mutually independent. Specifically, the first $q_y$ elements of this vector follow the distribution $F_Y(\cdot|z_{0})$, while the remaining $q_x$ elements follow $F_X(\cdot|z_{0})$. The proof leverages the fact that the induced order statistics $S_n$ in (ref) are conditionally independent given $(Z_1,\dots,Z_n)$, with conditional CDFs $$F_Y(\cdot|{Z_{n_y,(1)}}),\dots,F_Y(\cdot|{Z_{n_y,(q_y)}}),F_X(\cdot|{Z_{n_x,(1)}}),\dots,F_X(\cdot|{Z_{n_x,(q_x)}})~.$$ The result then follows by showing that $Z_{n_y,(j)}\overset{p}{\to}z_0$ and $Z_{n_x,(j)}\overset{p}{\to}z_0$ for all $j\in\{1,\dots,q\}$, and invoking standard properties of weak convergence. Theorem (ref) plays a crucial role in our asymptotic validity result presented in Theorem (ref).

In addition to Assumptions (ref) and (ref), we also require that, conditional on $Z=z$, the random variables $Y$ and $X$ have distributions with a finite number of discontinuity points. To state this assumption formally, let $\mathcal{D}_{Y}(z)$ and $\mathcal{D}_{X}(z)$ denote the sets of discontinuity points of the CDFs of ${ Y|Z=z }$ and ${ X|Z=z }$, respectively.

assumptionFor any $z \in \mathcal{Z}$, $|\mathcal{D}_{Y}(z)|$ and $|\mathcal{D}_{X}(z)|$ are finite.

It is important to note that Assumption (ref) allows both $Y$ and $X$ to be continuous, discrete, or mixed random variables. However, it excludes cases where these variables have countably many discontinuities conditional on $z\in \mathcal Z$. We also point out that Theorem (ref) does not require Assumption (ref).

We now formalize our main result in Theorem (ref), which shows that the test defined in (ref) is asymptotically level $\alpha$ under the assumptions we just introduced. Below, we denote by $E_P[\cdot]$ the expected value with respect to the distribution $P\in \mathbf P$.

theoremLet $\mathbf P$ the space of distributions that satisfy Assumptions (ref), (ref), and (ref) hold, and let $\mathbf P_{0}$ be as in (ref). Let $\alpha\in(0,1)$ be given, $T$ be the KS test statistic in (ref), and $\phi(\cdot)$ be the test in (ref). Then, \begin{equation} \limsup_{n \to \infty} E_{P}[\phi(S_n)]\le \alpha \qquadfor all P \in \mathbf P_{0} . \end{equation}

Theorem (ref) establishes the asymptotic validity of the test in (ref). There are three main reasons why the inequality in (ref) may be strict, resulting in the limiting rejection probability strictly below $\alpha$. First, for distributions $P$ in the interior of $\mathbf{P}_0$, where the inequality in (ref) holds strictly for some $t \in \mathbf{R}$, the test is expected to reject with probability less than $\alpha$, with a magnitude depending on the `distance' of $P$ from the boundary of $\mathbf{P}_0$. Second, in cases where $Y$ or $X$ are not continuously distributed, the critical value defined in (ref) serves as an upper bound for the desired quantile, as discussed further in Section (ref). Finally, the test statistic $\Delta(u)$ in (ref) is discretely distributed, taking only a limited number of distinct values. Consequently, the achieved significance level

equation[equation omitted — 133 chars of source]

satisfies $\bar{\alpha} \leq \alpha$ by definition, but may be strictly less than $\alpha$. Whether $\bar{\alpha} = \alpha$ occurs or not depends on whether the critical value $c_\alpha(q_y, q)$ aligns exactly with one of the discrete jumps in the CDF of $\sup_{u \in (0,1)} \Delta(u)$, which depends on $\alpha$, $q_y$, and $q$.

We derive Theorem (ref) by linking the weak convergence of the induced order statistics \( S_n \) to the limit variable \( S \) in (ref) with the finite-sample validity of the test \( \phi(S) \) in the limit experiment. This connection becomes nontrivial when the data are not continuously distributed. The KS statistic plays a central role in addressing these challenges. First, our proof leverages the fact that the rank of the induced order statistics is preserved as the sample size grows. The KS statistic, being rank-based, ensures that the rejection rate of \( \phi(S_n) \) converges to that of \( \phi(S) \) in the limit experiment. Second, the structure of the KS statistic implies that our test controls size in the limit experiment over our class of null distributions---including those that are discrete or mixed---thereby establishing asymptotic validity. By contrast, as shown in Section (ref) in the appendix, analogous results generally fail when using the one-sided versions of the Cramér--von Mises or Anderson--Darling statistics, which are commonly used in the stochastic dominance testing. Nonetheless, in the special case where the data are continuously distributed, we show in Theorem (ref) that the Cramér--von Mises-based test remains valid.

remarkgoldman/kaplan:2018 develop an inference framework for the two-sample one-sided hypothesis test in (ref), interpreting it as a multiple-testing procedure and with applications to the regression discontinuity design. As in this section, they adopt an asymptotic framework with fixed $q$ and construct critical values by simulating i.i.d.\ $U(0,1)$ random variables given a test statistic. Our approach to stochastic dominance testing, however, differs from theirs in several important respects. First, we focus on the KS statistic, whereas they advocate for a different statistic based on the so-called Dirichlet approach, which they argue may offer power advantages. This distinction allows us to accommodate discrete or mixed distributions, while their analysis is restricted to continuously distributed data; see Appendix (ref) for details.\footnote{Although goldman/kaplan:2018 state that their test remains asymptotically valid—albeit conservative—for discrete data, they do not provide a formal proof. By contrast, Theorem (ref) establishes this property for the KS statistic. Interestingly, asymptotic validity does not extend to other test statistics: Appendix (ref) shows that it fails for tests based on the Cramér--von Mises and Anderson--Darling statistics.} Second, we also derive results under an asymptotic framework where $q \to \infty$ as $n \to \infty$. This extension allows us to obtain data-dependent rules for choosing the tuning parameters that satisfy the required rate-of-convergence conditions.

Asymptotic results for large q

In this section, we examine the asymptotic properties of the test in (ref) within a framework where $q:= q_y+q_x \to \infty $ as $n\to \infty$. Our results here follow from the recently derived rates of convergence for induced order statistics in bugni/canay/kim:26b, and so we keep the discussion brief.

The result in Theorem (ref) establishes that the limiting distribution of the induced order statistics in the vector $S_n$ is such that the elements of the limit vector, denoted by $S$, are mutually independent. Thus, $S_n$ serves as an “approximation’’ to $S$, and while the infeasible test $\phi(S)$ would be ideal, in practice we must work with $\phi(S_n)$. Since $\phi \le 1$, the variational characterization of total variation implies

equation[equation omitted — 132 chars of source]

where $\mathcal L(\cdot)$ denotes the distribution of a random vector and $\mathrm{TV}$ denotes total variation distance. Hence, a bound on the rate of convergence of $\mathrm{TV}\big(\mathcal L(S_n),\mathcal L(S)\big)$ would immediately yield a rate at which the rejection probability $E_P[\phi(S_n)]$ is asymptotically bounded by $\alpha$, since in our setting $ E_P[\phi(S)]\le \alpha$. Unfortunately, Theorem (ref) is silent about the rate for the convergence of $S_n$ to $S$.

Theorem (ref) in Appendix (ref) leverages the results in bugni/canay/kim:26b to show that, under assumptions slightly stronger than those stated in Section (ref), it follows that

equation[equation omitted — 162 chars of source]

and so

equation[equation omitted — 133 chars of source]

Hence, the test defined in (ref) is asymptotically level $\alpha$ under those same assumptions, provided that $q_y = o(n_y^{2/3})$ and $q_x = o(n_x^{2/3})$. We rely on these rates of convergence in the next section to derive data-dependent rules of thumb for selecting the two tuning parameters.

remarkTheorem (ref) in Appendix (ref) also shows that $\phi(S_n)$ in (ref) is consistent for any fixed alternative $P\in \mathbf P_1$ when $q\to \infty$ at the specified rates.
remarkIn a framework where $q \to \infty$, one could alternatively define a test using a critical value based on the $1-\alpha$ quantile of the limiting distribution of the (scaled) KS statistic in (ref). Denote this limiting critical value by $c_{\alpha}(\infty)$. Numerical evaluations show that our scaled critical value is smaller than the limiting critical value: $\sqrt{q_yq_x/q}\,c_{\alpha}(q_y,q) \le c_{\alpha}(\infty)$ for $\alpha \in \{0.1, 0.05, 0.01\}$ and $(q_y,q_x) \in \{10,20,\ldots,500\}^2$, with strict inequality in most cases. For example, when $\alpha=5\%$ and $(q_y,q_x) = (70,70)$---a value consistent with our simulations---we obtain $\sqrt{q_yq_x/q}\,c_{\alpha}(q_y,q) = 1.1832 < 1.2239 = c_{\alpha}(\infty)$. The smaller critical value associated with our test translates into an improvement in power of roughly $4\%$ under a local-power approximation for location-shift alternatives. We omit the details for brevity, but they are available upon request.

Discussion and extensions

Data-dependent choice of tuning parameters

We now discuss the practical considerations for implementing our test. We propose a data-dependent method for two tuning parameters $q_y$ and $q_x$, drawing on arguments from armstrong/kolesar:18, similar to the approach used by bugni/canay:21. Importantly, the analysis from these arguments is consistent with the rates of convergence implied by (ref). Concretely, this method leverages a bias-variance trade-off inherent in the estimation of the conditional CDFs used in the test statistic for $\phi(S_n)$, within an asymptotic framework where $q$ can grow slowly with $n$. Our goal is to provide practical guidance for choosing these tuning parameters based on the data, rather than claiming optimality or even validity of any sort. We examine the performance of this rule via Monte Carlo simulations in Section (ref) and use it in the empirical application in Section (ref).

We propose choosing $q_x$ and $q_y$ using the following data-dependent rules:

equation[equation omitted — 227 chars of source]

and

equation[equation omitted — 229 chars of source]

where $\mu_Z := E[Z]$, $\sigma_Z^2 := \text{Var}[Z]$, $\phi_{\mu_Z,\sigma_Z}(\cdot)$ denotes the probability density function of a normal distribution with mean $\mu_Z$ and variance $\sigma_Z^2$, and $\rho_Y$ and $\rho_X$ are the correlation coefficients between $Y$ and $Z$, and $X$ and $Z$, respectively.

To provide some intuition as to why this rule of thumb may be reasonable, assume that the random variable $Z$ is continuous with a density function $f_Z(\cdot)$ satisfying

equation*[equation* omitted — 105 chars of source]

and any values $z_1,z_2\in \mathcal Z$. In addition, suppose that the conditional CDF of $Y$ satisfies

equation*[equation* omitted — 121 chars of source]

and that the conditional CDF of $X$ satisfies the same condition with a constant $C_X$. It can be shown that the standardized bias $B_{n_y,q_y}$ associated with the estimator of the conditional CDF $F_Y(\cdot|z_{0})$ satisfies

equation[equation omitted — 111 chars of source]

Let $t^{\ast}$ denote the right-hand side of (ref). Solving for $q_y$, we obtain

equation*[equation* omitted — 103 chars of source]

The proposed data-dependent rule in (ref) and (ref) can be viewed as under-smoothed approximations of these values, where the unknown Lipschitz constants are approximated by the working model $Z \sim N(\mu_Z, \sigma_Z)$. This guarantees that $q_y = o(n_y^{2/3})$, which is the condition discussed in Section (ref) to obtain (ref). The constant multiplying $n_y^{1/2}$ in (ref) is intuitive for two reasons. First, it reflects that a steeper density at $z_{0}$, or a steeper derivative of the conditional CDFs at $z_{0}$ calls for smaller $q_y$. In such cases, nearby observations provide a poorer approximation of the quantities at $z_{0}$. Since the maximum slope is determined by the constants $C_Z$ and $C_Y$, the rule is inversely proportional to these constants. Second, the rule accounts for low density at $z_{0}$. When $f_Z(z_{0})$ is small, the $q_y$ nearest observations are likely to be farther from $z_{0}$, again requiring smaller $q_y$. While one could replace the normality assumption with a nonparametric estimator of $f_Z(\cdot)$, it is unfortunately impossible to adaptively choose $C_Z$ and $C_Y$ for testing (ref) armstrong/kolesar:18. Since any data-dependent rule for $q$ must reference $C_Z$ and $C_Y$, we prioritize simplicity and use normality for both $f_Z(\cdot)$ and the associated constants.

Refined critical value for discrete data

While our default critical value in (ref) is valid as long as the distributions of random variables $Y$ and $X$ have finitely many discontinuities, it is possible to construct a smaller refined critical value when both variables are discretely distributed with a limited number of support points. Our test with the refined critical value still maintains the asymptotic validity, though it comes at the cost of additional computational complexity.

To motivate the refined critical value, let $\mathbf Y$ and $\mathbf X$ denote the support of $Y$ and $X$, respectively. Our proof for the asymptotic validity (specifically, Theorem (ref) in the appendix) relies on the following inequalities:

equation*[equation* omitted — 202 chars of source]

where $\mathcal{U} := \cup_{t\in \mathbf Y}\{u = F_X(t|z_{0}) \}$ is the set of values that $F_X(t|z_{0})$ takes as $t$ varies over $\mathbf Y$ and $\Delta(u) $ is given in (ref). Once we replace the set $ \mathcal{U}$ with the interval $(0,1)$, our default critical value $c_{\alpha}(q_y,q)$, defined as a quantile of $\sup_{u\in(0,1)}\Delta(u)$ in (ref), ensures the probability is bounded below $\alpha.$ However, replacing the set $ \mathcal{U}$ with $(0,1)$ could be unnecessarily conservative when $ \mathcal{U}$ contains only a few points---either because $\mathbf Y$ contains only a few points or because $F_X(t|z_{0})$ takes few distinct values as $t$ varies. In fact, the cardinality of the set $ \mathcal{U}$ is determined by the smaller support size of $Y$ and $X$.

To define our refined critical value, let $r$ denote the smaller support size of $Y$ and $X$,

equation*[equation* omitted — 60 chars of source]

and let $\mathbf{U}_r$ denote the collection of all ordered $r$-tuples of distinct points in (0,1),

equation*[equation* omitted — 100 chars of source]

We denote an arbitrary element of $\mathbf{U}_r$ by $\mathcal{U}_r$. Our refined critical value is defined as

equation[equation omitted — 228 chars of source]

with $\Delta(u)$ as in (ref), and the refined test for the null hypothesis in (ref) when either $Y$ or $X$ is discretely distributed with a limited number of support points is thus

equation[equation omitted — 91 chars of source]

Section (ref) presents the general version of this test for the case $L>1$.

The power gains of using $c_{\alpha}^{r}(q_y, q)$ over $c_{\alpha}(q_y, q)$ are most pronounced when $r$ is small. Our numerical analysis shows the largest gains for $r \le 10$. Thus, this refinement is most effective for discrete data with a limited number of support points, rather than all discrete settings. Moreover, the computational cost of $c_{\alpha}^{r}(q_y, q)$ increases with $r$, and so as $r$ grows, the gains diminish while the cost rises.

We propose to compute $c_{\alpha}^{r}(q_y, q)$ numerically, by solving the following optimization problem:

equation[equation omitted — 272 chars of source]

where $\mathbf{T}$ is the support of $\Delta(u)$ in (ref). Here, $c_{\rm ub} = c_{\alpha}(q_y, q)$ and $c_{\rm lb}$ is given by

equation*[equation* omitted — 213 chars of source]

The fact that $c_{\alpha}(q_y, q)$ provides a valid upper bound is unsurprising given the preceding discussion. On the other hand, $c_{\rm lb}$ serves as a valid lower bound because $\left\{\frac{1}{1+r},\frac{2}{1+r},\dots,\frac{r}{1+r} \right\}$ is a specific element in $\mathbf U_{r}$. In our numerical evaluations, we often found that $c_{\alpha}^{r}(q_y, q) = c_{\rm lb}$, but not always. This indicates that the additional optimization in (ref) cannot be generally avoided. For modest values of $r$ and $q$, however, this optimization step is computationally straightforward, primarily due to the relatively small number of points typically found in $[c_{\rm lb}, c_{\rm ub}] \cap \mathbf{T}$.

remarkThe support $\mathbf{T}$ of $\Delta(u)$ is a discrete subset of $[-1,1]$ and can be easily enumerated for modest values of $q$. Specifically, the support has cardinality bounded by $(q_y+1)(q_x+1)$, and it is independent of the realizations of the random variables as well as the specific value that $ u \in (0,1)$ takes.

Properties of our test in the limit experiment

In this section, we study the properties of the test $\phi(\cdot)$ in (ref) in the limit experiment associated with the asymptotic framework in Section (ref). By Theorem (ref), this test is equivalent to $\phi(S)$, where

equation[equation omitted — 177 chars of source]

In words, in the limit experiment, we observe one random sample of size $q_y$ from the distribution $F_Y(\cdot|z_{0})$ and the other independent random sample of size $q_x$ from the distribution $F_X(\cdot|z_{0})$. The KS test statistic in (ref) is a function of $S$, and the critical value $c_{\alpha}(q_y, q)$ remains unchanged. We begin our discussion by focusing on the case where $S$ is continuously distributed.

The finite-sample properties of the two-sample one-sided KS statistic have been extensively studied in the literature of testing equality of two continuous (unconditional) distributions. Early works established that the test statistic's finite-sample distribution is pivotal under the null and developed algorithms for its computation (gnedenko/korolyuk:1951,korolyuk:1955,blackman:1956,hodges:1958,hajek/sidak:1967, and durbin:1973). Our critical value is obtained from this pivotal finite-sample distribution, despite the different null hypothesis.

This connection to the literature of testing equality of two distributions arises from the observation that the distribution \(F_Y(\cdot|z_{0}) = F_X(\cdot|z_{0})\) is the least favorable within the set of null distributions \(\mathbf{P}_0\) in (ref) satisfying stochastic dominance, a point first made by lehmann:1951 and later reiterated by hodges:1958,mcfadden:1989, and goldman/kaplan:2018. We define the subset of continuous distributions in \(\mathbf{P}_0\) that satisfy \(F_Y(\cdot|z_{0}) = F_X(\cdot|z_{0})\) as \(\mathbf{P}^*_0\), and denote a generic element in \(\mathbf{P}^*_0\) by \(P^*\). Although these papers studied the distribution of KS statistic under \(P^*\), they did not provide a formal proof that \(P^*\) determines the size of the test under the null hypothesis in (ref). For completeness, we formally state this result in Lemma (ref) below, and provide its proof.

lemmaLet \(\mathbf{P}^*_0 \subset \mathbf{P}_0\) be the subset of distributions \(P^*\) of the random variable \(S\) in (ref), such that for a continuous CDF $F$, \(S_j \sim F\) for all \(j = 1, \ldots, q\). Let \(\phi(\cdot)\) be the test defined in (ref). Then, for any $P^*\in \mathbf P^*_0$, we have \[ \sup_{P \in \mathbf{P}_0} E_P[\phi(S)] = E_{P^*}[\phi(S)] = \bar{\alpha}~, \] where $\bar{\alpha}\le \alpha $ is defined in (ref).

Lemma (ref) shows that when $S \sim P^*\in \mathbf P^*_0$, our test is `exact' in that it achieves the closest possible rejection rate to $\alpha$, defined as $\bar{\alpha}$ in (ref). In this case, we have $$ T(S) \stackrel{d}{=} T(U) \quad \text{where}\quad \{U_j \sim U[0,1]: 1 \leq j \leq q\} \; \text{are i.i.d.}$$ This demonstrates that the analytical (finite-sample) critical value $c_{\alpha}(q_y, q)$ for the one-sided KS test can be accurately approximated by simulating uniform random variables. However, when $S \sim P\in \mathbf P_0$ is such that $P\{S_i = S_{j}: i\not=j\} > 0$, the connection $T(S) \stackrel{d}{=} T(U)$ breaks down. In this case, $c_{\alpha}(q_y, q)$ is no longer the finite-sample analytical quantile of $T(S)$, but rather a valid upper bound.

When \(S\) is continuously distributed, the proposed test \(\phi(S)\) is equivalent to a non-randomized permutation test. This connection establishes the validity of permutation tests for testing stochastic dominance. While hodges:1958 and mcfadden:1989 suggested that a permutation test could be used for this purpose, they did not provide a formal justification. To the best of our knowledge, our proof of this result is novel.

To formally define a permutation test, we introduce the following notation. Let \(\mathbf{G}\) denote the set of all permutations \(\pi = (\pi(1), \dots, \pi(q))\) of \(\{1, \dots, q\}\). The permuted values of \(S\) are given by \[ S^{\pi} = (S_{\pi(1)}, \dots, S_{\pi(q)})~. \] The (non-randomized) permutation test is then defined as follows:

align[align omitted — 276 chars of source]

It follows from standard arguments (see, e.g., lehmann/romano:05) that when $S$ is invariant to permutations, i.e., $S \stackrel{d}{=} S^{\pi}$, the randomized version of the test $\phi^{\rm p}(S)$ is exact in finite samples. However, under the null hypothesis in (ref), we have that $S \stackrel{d}{\not =} S^{\pi}$ for some $P \in \mathbf{P}_0$, and so invariance (or the so-called randomization hypothesis) fails. Therefore, the traditional finite-sample arguments for validity no longer apply. Alternative arguments that claim validity of permutation tests when invariance does not hold typically require $q \to \infty$; see chung/romano:13,canay/romano/shaikh:17, and bugni/canay/shaikh:18, among others. In our current setting, where $q$ is fixed, such arguments do not apply.

We contribute to this literature by demonstrating that, when \(S\) is continuously distributed, our test is equivalent to a non-randomized permutation test. The formal statement of this result follows.

lemmaFor any random variable $\tilde{S}\in \mathbf{R}^{q}$ with $P\{\tilde{S}_{i}\neq \tilde{S}_{j} : i\not=j\}=1$, we have \begin{equation} P\{ c_{\alpha }^{\rm p}(\tilde{S}) =c_{\alpha }( q_{y},q)\} =1 . \end{equation} Moreover, (ref) no longer holds if $\tilde{S}$ is such that $P\{\tilde{S}_{i}=\tilde{S}_{j}: i\not=j\}>0$.

Lemma (ref) shows that our data-independent critical value in (ref) is equivalent to the critical value of a permutation test when the random variable $\tilde{S}$ has no ties. Lemma (ref) immediately implies when $S_n$ in (ref) is continuously distributed, we have

equation[equation omitted — 96 chars of source]

It follows from (ref) and Theorem (ref) that, when $S$ is continuously distributed, a non-randomized permutation test controls the limiting rejection probability under the null hypothesis in (ref). Importantly, this result holds even though invariance does not hold for all $P \in \mathbf{P}_0$.

remarkLemmas (ref) and (ref) establish the validity of permutation tests for testing stochastic dominance in finite-sample settings. This result illustrates an instance where permutation tests can provide finite-sample valid inference even in settings where the randomization hypothesis does not hold. The only similar result we are aware of is that of caughey/etal:2023, who consider a design-based framework with the null hypothesis $\tau_i \le 0$, where $\tau_i$ represents a unit-level treatment effect. Like our work, their result is valid under the condition that the random variables are continuously distributed or that a random tie-breaking rule is applied to handle ties.

The results in Lemmas (ref) and (ref) reveal interesting and novel connections between our test and classical arguments involving finite-sample critical values and permutation tests. However, these results critically depend on the random variable $S$ being continuously distributed and do not extend to cases where $S$ is discretely distributed.

When \( S \) is discrete and ties occur with positive probability, i.e., \( P\{S_i = S_{j}\} > 0: i\not=j \), the finite-sample distribution of the KS test statistic \( T(S) \) depends on the number and location of these ties. A natural approach to handle ties is to redefine the test statistic to randomly break them, effectively making the test a randomized one. While this would allow us to establish an analog of Lemma (ref) for discrete data, we do not pursue such an extension, as randomized tests are rarely used in practice. Despite our best efforts, we were unable to demonstrate that a permutation test could control size under the null hypothesis of stochastic dominance in (ref) when $S$ is discrete, without relying on random tie-breaking rules.

Simulations

In this section, we evaluate the finite-sample performance of the test in (ref) for $L=1$ or the test proposed in Section (ref) for $L>1$ through a simulation study. We present a variety of data-generating processes to illustrate both the strengths and potential limitations of our test.

We consider seven distinct designs with four cases, (a) to (d), in each design as follows:

itemize• {\bf Case (a)}: the null hypothesis holds with equality and $L=1$ • {\bf Case (b)}: the null hypothesis holds with strict inequality and $L=1$ • {\bf Case (c)}: the null hypothesis holds with equality and $L=2$ • {\bf Case (d)}: the null hypothesis is violated and $L=1$

The first three designs are based on the following location–scale model:

equation[equation omitted — 118 chars of source]

where $U$ and $V$ are random variables with specified distributions, and the conditioning variable $Z$ follows a non-negative Beta$(2,2)$ distribution. Under this location–scale model, whenever $\sigma_Y(z_{\ell}) = \sigma_X(z_{\ell})$, the null hypothesis in (ref) holds as long as

equation[equation omitted — 75 chars of source]

Design 1 satisfies (ref) for all $z\in(0,1)$. Design 2 satisfies (ref) at $z_{\ell} = 0.5$, but violates the inequality for $z > 0.5$, which may affect the performance of our test in finite samples due to its reliance on induced order statistics. Design 3 is such that $U$ and $V$ are $U[0,1]$, which guarantees that (ref) holds even when $\sigma_Y(z_{\ell}) < \sigma_X(z_{\ell})$, provided $\mu_Y(z_{\ell}) - \mu_X(z_{\ell}) \ge \sigma_X(z_{\ell}) - \sigma_Y(z_{\ell})$.

Design 4 is a slight variation of the location-scale model to accommodate a regression discontinuity design (RDD):

equation[equation omitted — 169 chars of source]

Following shen/zhang:16, we consider the case where $U$ and $V$ are independent $N(0,1)$ random variables, $Z \sim 2\mathrm{Beta}(2,2) - 1$, and both $\mu_Y(Z)$ and $\mu_X(Z)$ are defined in terms of the function

equation[equation omitted — 82 chars of source]

A key feature shared by Designs 1 through 4 is that both $X$ and $Y$ are continuously distributed; a feature not shared by the next three designs.

Design 5 is such that $U$ and $V$ follow log-normal distributions with the bottom 20% censored. This setup reflects features of wage distributions, which often exhibit a point mass at the minimum wage. Design 6 defines conditional probabilities \( P\{X = k | Z\} \) and \( P\{Y = k | Z\} \) as $$ P\{X = k | Z\} = \frac{e^{\theta^x_k(\frac{3}{2}-Z)}}{\sum_{j=1}^3 e^{\theta^x_j(\frac{3}{2}-Z)}} ~~\text{and}~~ P\{Y = k | Z\} = \frac{e^{\theta_k^y(\frac{3}{2}-Z)}}{\sum_{j=1}^3 e^{\theta^y_j(\frac{3}{2}-Z)}} ~\text{for }k=1,2,3, $$ where, for $\mu_Y(z) \in [-1,1]$,

equation*[equation* omitted — 144 chars of source]

The parameter $ \theta^x_k = (-0.5, -1.5, -2)$ controls the baseline log-odds of each category for $X$, while the factor $(3/2 - Z)$ introduces a monotonic dependence on $Z$. Finally, Design 7 considers $$ X|Z \sim B\Big([25Z], \frac{1}{2}\Big)\quad \text{and}\quad Y|Z \sim B\Big([25Z] + \mu_Y(Z), \frac{1}{2}\Big)~,$$ where $B(\cdot, \cdot)$ denotes a binomial distribution and $[x]$ represents the nearest integer to $x$.

figure[figure omitted — 348 chars of source]
figure[figure omitted — 348 chars of source]

The parameter values for all Designs are reported in Appendix (ref), where we also report the mean values of \(q_y^*\) and \(q_x^*\) across simulations.

We report results for sample size $n = 1,000$ and nominal level $\alpha = 10\%$, based on $10,000$ Monte Carlo simulations to test the null hypothesis in (ref). To implement the test $\phi(\cdot)$ in (ref), denoted by `KS' in the figures, we select the tuning parameters $(q_y, q_x)$ using the data-dependent rules in (ref) and (ref). For Designs 1-3 with $L=1$, we compare our test with the one proposed in goldman/kaplan:2018. For the RDD Design 7 with $L=1$, we compare our test with the method in shen/zhang:16.\footnote{SZ is implemented using the authors’ rule of thumb, whereas for GK we use our own rule of thumb, as the original paper does not provide guidance on how to select their tuning parameter. Note that neither of these tests is defined when $L>1$.} For the mixed and discrete designs, we present results for our test in (ref) as well as for its refined version in (ref).

Figure (ref) reports rejection probabilities under the null hypothesis for the continuously distributed designs. When the data-generating process satisfies the null in (ref) with equality (case (a)), the KS test’s rejection probabilities closely align with the nominal level and consistently outperform the GK and SZ tests. When the null holds with strict inequality (case (b)), rejection rates fall below \(\alpha\), consistent with our critical value serving as a valid upper bound for the true quantile. Case (c), where \(L = 2\), shows behavior similar to case (a), where \(L = 1\). Overall, the KS test exhibits excellent size control.

figure[figure omitted — 329 chars of source]

Figure (ref) reports rejection probabilities under the null hypothesis for the mixed and discretely distributed designs. When the data-generating process satisfies the null in (ref) with equality, the KS test we propose in (ref) works well when the data is mixed, but it is conservative when the data is discrete with few support points. Designs 6 and 7 illustrate that the refined critical value \(c^{\rm r}_{\alpha}(q, q_y)\) in (ref) offers a more accurate approximation of the true quantile of the KS test statistic, though it may still be somewhat conservative.

Figure (ref) reports rejection probabilities under the alternative hypothesis for all designs. The power of the KS test is similar to that of GK and could be above, below, or roughly the same. The power of the KS test is much higher than that of SZ for this design. This is worth noting given that the KS test uses about $90$ observations (for both $q_y$ and $q_x$) while the SZ test uses $276$ effective observations on each side of the threshold. The power results also demonstrate that the refined critical value \(c^{\rm r}_{\alpha}(q, q_y)\) in (ref) enhances power in discrete cases.

Concluding remarks

This paper introduces a novel test for conditional stochastic dominance (CSD) at target points, offering a flexible, nonparametric approach that avoids kernel smoothing while ensuring computational efficiency. By leveraging induced order statistics, our method constructs empirical CDFs using observations closest to the target conditioning point. We establish the asymptotic properties of our test, demonstrating its validity under weak regularity conditions, and derive a critical value that eliminates the need for resampling techniques such as the bootstrap. Additionally, we extend our framework to better handle discrete data, proposing a refined critical value that enhances the power of the test with minimal additional information. Monte Carlo simulations align with our theoretical results and suggest that our test performs well in finite samples, making our test readily applicable to empirical research in economics, finance, and public policy.

An important feature of our test is its simplicity. Once the key tuning parameters are computed, the test only requires a standard test statistic with a deterministic critical value, without the need for kernels, local polynomials, bias correction, or bandwidth selection. Furthermore, our test admits a clear interpretation in the limit experiment, which allows us to connect it with classical analytical critical values and permutation-based tests. In this sense, our findings contribute to the broader literature on stochastic dominance testing by refining conditional inference methods and establishing new links between permutation-based and rank-based approaches. One open question we did not address in this paper concerns the validity of permutation-based tests for the hypothesis of stochastic dominance when both random variables, $Y$ and $X$, are discrete. Despite attempts to formalize this result, we were unable to prove or disprove it. Extensive Monte Carlo simulations (not reported here) suggest that the test may be valid, and this is an area we plan to explore further.

\appendices \setcounter{equation}{0}

Proof of the main results

Proof of Theorem (ref)

By Theorem (ref),

equation*[equation* omitted — 62 chars of source]

where the elements of $S$ are independent, and $S_{j}\sim F_{Y}(\cdot |z_{0})$ for $j=1,\cdots ,q_{y}$ and $S_{j}\sim F_{X}(\cdot |z_{0})$ for $ j=q_{y}+1,\cdots ,q$. By the almost-sure representation theorem, we have a sequence of random vectors $\{\Tilde{S}_{n}: 1\le i\le \infty \}$ and a random vector $\Tilde{S}$ defined on a common probability space $(\Omega ,\mathcal{A },\Tilde{P})$ such that

equation*[equation* omitted — 133 chars of source]

Let $R(s)$ denote the rank of $s$, which maps $s$ to a permutation of $\{1,2,\dots ,q\}$. Define the event that the rank of the two vectors coincides as follows,

equation*[equation* omitted — 62 chars of source]

To reach the conclusion, it suffices to show that

equation[equation omitted — 66 chars of source]

To see this, consider the following argument,

align*[align* omitted — 265 chars of source]

where (1) holds by $\Tilde{S}_{n}\overset{d}{=}S_{n}$, (2) by the fact that $T$ is invariant to rank-preserving transformations, (3) by (ref) and $\phi(\cdot)\in \{0,1\}$, and (4) by $P \in {\bf P}_0$ and Theorem (ref).

We devote the remainder of the proof to establishing (ref). Define \(\mathcal{D} = \mathcal{D}_{X}(z_{0}) \cup \mathcal{D}_{Y}(z_{0})\), where \(\mathcal{D}_{X}(z_{0})\) and \(\mathcal{D}_{Y}(z_{0})\) denote the sets of discontinuity points as specified in Assumption (ref). For any \(\varepsilon > 0\), let

align*[align* omitted — 444 chars of source]

Observe that

align[align omitted — 116 chars of source]

To establish this, consider the following argument. For any \(i,j = 1, \dots, q\), there are three possible cases: (i) \(\tilde{S}_{i} < \tilde{S}_{j}\), (ii) \(\tilde{S}_{i} > \tilde{S}_{j}\), or (iii) \(\tilde{S}_{i} = \tilde{S}_{j}\). First, consider case (i), where \(\tilde{S}_{i} < \tilde{S}_{j}\). Under \(E_{1}(\varepsilon)\), this implies \(\tilde{S}_{i} < \tilde{S}_{j} - \varepsilon\). Under \(E_{n,2}(\varepsilon)\), we have \(\tilde{S}_{n,i} - \varepsilon/2 < \tilde{S}_{i}\) and \(\tilde{S}_{j} < \tilde{S}_{n,j} + \varepsilon/2\). Combining these inequalities yields \(\tilde{S}_{n,i} < \tilde{S}_{n,j}\), as required. Case (ii) follows identically by reversing the roles of \(i\) and \(j\). Finally, consider case (iii), where \(\tilde{S}_{i} = \tilde{S}_{j}\). Under \(E_{1}(\varepsilon)\), this implies \(\tilde{S}_{i} = \tilde{S}_{j} \in \mathcal{D}\). By \(E_{n,3}\), it follows that \(\tilde{S}_{n,i} = \tilde{S}_{n,j} \in \mathcal{D}\). Since this argument holds for all \(i, j = 1, \dots, q\), we conclude that \(F_{n}\) follows, as desired.

By (ref), (ref) follows that there exits $\varepsilon >0$ such that

equation[equation omitted — 118 chars of source]

For arbitrary $\delta >0$, (ref) follows from finding $ \varepsilon =\varepsilon (\delta )>0$ and $N(\delta) $ such that $\tilde{P}\{ E_{1}(\varepsilon )\cap E_{n,2}(\varepsilon )\cap E_{n,3}\} \geq 1-\delta $ for all $n\geq N(\delta) $. Let $\varepsilon_{1}=\inf \{ \Vert \tilde{d}-d\Vert /2:d<\tilde{ d}\in \mathcal{D}\} >0$. By Lemma (ref), $\exists \varepsilon _{2}>0$ such that, for $i\not=j=1,\ldots ,q$,

align*[align* omitted — 334 chars of source]

Finally, set $\varepsilon =\min \{ \varepsilon _{1},\varepsilon _{2}\} >0$ for the remainder of the proof. By elementary arguments, it suffices to show that: (i) $\tilde{P}\{ E_{1}(\varepsilon) ^{c}\} \leq \delta /3$, (ii) $\exists N_{2}(\delta) \in \mathbf{N}$ s.t. $ \tilde{P}\{ E_{n,2}(\varepsilon) ^{c}\} \leq \delta /3$ for all $n\geq N_{2}(\delta) $, and (iii)\ $\exists N_{3}(\delta) \in \mathbf{N}$ s.t. $\tilde{P}\{ E_{n,3}(\varepsilon) ^{c}\} \leq \delta /3$ for all $ n\geq N_{3}( \delta ) $. We divide the rest of the proof into three results.

First, we show that $\tilde{P}\{ E_{1}(\varepsilon) ^{c}\} \leq \delta /3$. To this end, pick $ i\not=j=1,\ldots ,q$ arbitrarily. Note that

equation*[equation* omitted — 186 chars of source]

where (1) holds by $i\not=j=1,\ldots,q$, ${\tilde{S}}_{i}$ and ${ \tilde{S}}_{j}$ being identically distributed, and $\varepsilon =\min \{ \varepsilon _{1},\varepsilon _{2}\} $. From here, we conclude that

equation*[equation* omitted — 220 chars of source]

as desired. Second, we show that $\exists N_{2}(\delta) \in \mathbf{N}$ such that $\tilde{P}\{ E_{n,2}(\varepsilon) ^{c}\} \leq \delta /3$ for all $n\geq N_{2}(\delta) $. To see this, note that

equation[equation omitted — 128 chars of source]

By $\Tilde{S}_{n}\overset{a.s.}{\to }\Tilde{S}$, $\exists N_{2}(\delta) $ such that the right-hand side is less than $ \delta /3$, as desired. Finally, we show that $\exists N_{3}(\delta) \in \mathbf{N}$ such that $\tilde{P}\{ E_{n,3}(\varepsilon) ^{c}\} \leq \delta /3$ for all $n\geq N_{3}(\delta) $. To see this, note that

align*[align* omitted — 300 chars of source]

By Lemma (ref), $\exists N_{3}(\delta) $ such that the right-hand side is less than $\delta /3$, as desired. This completes the proof of (ref) and the theorem.

Proof of Lemma (ref)

Note that

equation*[equation* omitted — 138 chars of source]

where (1) holds by $P^{\ast }\in \mathbf{P}_{0}$ and (2) by Theorem (ref). To complete the proof, it suffices to show that $E_{P^{\ast }}[\phi (S)]=\bar{\alpha}$. To this end, consider the following argument:

align[align omitted — 736 chars of source]

where (1) holds by (ref), (2) holds by pollard:02 and the same arguments used in the proof of Theorem (ref), (3) follows from the continuity of $F$ guaranteeing that

equation*[equation* omitted — 94 chars of source]

for $ \Delta(u) := \frac{1}{q_y}\sum_{j=1}^{q_y} I\{U_{j} \le u\}-\frac{1}{q_x}\sum_{j=q_y+1}^{q} I\{U_{j} \le u\}$ as defined in (ref), and (4) by definition of $\bar{\alpha}$ in (ref).

Proof of Lemma (ref)

Let $Q:=\{Q_{1},\dots ,Q_{q}\}=\{1,2,\dots ,q\}$ and denote by $Q^{\pi }:=\{Q_{\pi (1)},Q_{\pi (2)}\dots ,Q_{\pi (q)}\}$ the permutation $\pi =(\pi (1),\pi (2),...,\pi \left( q\right) )$ of $Q$. Let $R(s)$ denote the rank of $s$, which maps $s$ to a permutation of $Q$. Since the KS statistic $T(\cdot )$ in (ref) is a rank statistic, it follows that for any $s$

equation[equation omitted — 71 chars of source]

where $T^{\ast }$ is a known function; see hajek/sidak/sen:99. That is, the KS test statistic depends on $S$ only through $R(S)$. Define

equation[equation omitted — 234 chars of source]

We divide the rest of the argument into four steps.

Step 1. For any $s\in \mathbf{R}^{q}$ with $s_{i}\not=s_{j} $ for $i\not=j$, and $c_{\alpha }^{\mathrm{p}}(s)$ as in (ref),

equation*[equation* omitted — 73 chars of source]

To establish this, consider the following derivation,

align[align omitted — 897 chars of source]

Here, (1) follows by (ref), (2) follows since $ s_{i}\not=s_{j}$ for $i\not=j$ implies that $R(s^{\pi })=(R(s))^{\pi }$ and $ R(s)=Q^{\bar{\pi}(s)}$ for some $\bar{\pi}\left( s\right) \in \mathbf{G}$, (3) by $\mathbf{G}=\tilde{\mathbf{G}}:=\{\pi \circ \bar{\pi}(s):\pi \in \mathbf{G}\}$, which follows from the fact that $\mathbf{G}$ is a group, and (4) by (ref).

Step 2. $c_{\alpha }^{\mathrm{p}}=c_{\alpha }\left( q_{y},q\right) $, where $c_{\alpha }\left( q_{y},q\right)$ is defined in (ref).

Let $\{U_i : 1\le i\le q\}$ be i.i.d.\ with $U_{i}\sim U\left( 0,1\right) $ and let $\hat{\pi}$ be a uniformly chosen permutation from $ \mathbf{G}$, independent of $U$. Note that $c_{\alpha }\left( q_{y},q\right) $ is the $\left( 1-\alpha \right) $-quantile of $T(U)$ and, by (ref), $ c_{\alpha }^{\mathrm{p}}$ is the $\left( 1-\alpha \right) $-quantile of the CDF $\frac{1}{\left\vert \mathbf{G}\right\vert }\sum_{\pi \in \mathbf{G} }I\left\{ T^{\ast }(Q^{\pi })\leq x\right\} $. The desired result then follows from noting that $\frac{1}{\left\vert \mathbf{G}\right\vert }\sum_{\pi \in \mathbf{G}}I\left\{ T^{\ast }(Q^{\pi })\leq x\right\} $ is the CDF of $T(U)$, as we show next.

Let $E:=\left\{ U_{i}\not=U_{j}\text{ for }i\not=j\right\} $. For any $x\in \mathbf{R}$, our desired result follows from this derivation:

align*[align* omitted — 637 chars of source]

Here, (1) holds by $U\overset{d}{=}U^{\hat{\pi}}$, (2) and (5) by $P\{E\}=1$, (3) by $\hat{\pi}\perp U$ and $\hat{\pi}$ uniformly chosen in $\mathbf{G}$, and (4) by repeating the arguments used to derive (ref).

Step 3. By Step 1, $\{S_{i}\not=S_{j}$ for $ i\not=j\}\subseteq \left\{ c_{\alpha }^{\mathrm{p}}(S)=c_{\alpha }^{\mathrm{p}}\right\} $, and so $ P\left\{ c_{\alpha }^{\mathrm{p}}(S)=c_{\alpha }^{\mathrm{p}}\right\} =1$ holds by our assumption. By Step 2, $c_{\alpha }^{\mathrm{p}}=c_{\alpha }\left( q_{y},q\right) $. The desired result follows from combining these points.

Step 4. To show the last statement, consider $S=\{1 :1\le j \le q\}$. Then, $T(S^{\pi })=T(S)=0$ for all $\pi \in \mathbf{G}$, and so $c_{\alpha }^{\mathrm{p}}(S)=0$. On the other hand, Step 2 implies $c_{\alpha }^{\mathrm{p}}=c_{\alpha }\left( q_{y},q\right) $, which are positive for typical values of $\left( q_{y},q,\alpha \right) $. For example, $q_{y}=1$, $q=2$, and $\alpha =0.1$ yield $c_{\alpha }^{\mathrm{p}}=c_{\alpha }\left( q_{y},q\right) =0.5.$

Supporting technical results

For any $\varepsilon>0$, we use $o_{\varepsilon}(1)$ to denote an expression that converges to zero as $\varepsilon \to 0$. Analogously, for any $n \in \mathbf{N}$, we use $o_{n}(1)$ to denote an expression that converges to zero as $n \to \infty$.

Auxiliary theorems

theoremLet Assumptions (ref) and (ref) hold with $\mathcal{Z}=\{z_{1},z_{2},\dots ,z_{L}\}$, and let $S^{\ell}_n$ be defined as in (ref). Then, \begin{equation} ({S_{n}^{1}}',\ldots ,{S_{n}^{L}}')\overset{d}{\rightarrow }({S^{1}}',\ldots ,{S^{L}}') , \end{equation} where, for any $(s_{1}^{\ell},\dots,s_{q_y^\ell}^{\ell},s_{q_y^\ell + 1}^{\ell},\dots,s_{q^\ell}^{\ell})\in \mathbf{R} ^{q^\ell}$ for each $\ell =1,\dots,L$ with $q^\ell = q_y^\ell+q_x^\ell$, $({S^{1}}',\ldots ,{S^{L}}')$ has the following distribution: \begin{equation*} P\left\{ \bigcap_{\ell =1}^{L}\bigcap_{i=1}^{q^{\ell }}\left\{ S_{i}^{\ell }\leq s_{i}^{\ell }\right\} \right\} = \prod_{\ell =1}^{L}\prod_{i=1}^{q_{y}^{\ell }}F_{Y}(s_{i}^{\ell }|z_{\ell })\prod_{j=1}^{q_{x}^{\ell }}F_{X}(s_{j+q_{y}^{\ell }}^{\ell }|z_{\ell }) , \end{equation*}
proofFor each $\ell =1,\dots ,L$, let $M_y^{\ell}$ denote the subset of the indices $i=1,\dots,n_y$ corresponding to the $q_y^{\ell}$ first $g$-order statistics $(Z_{\ell,(1)},\dots,Z_{\ell,(q_y^{\ell})})$, and let $M_x^{\ell}$ denote the subset of the indices $j=1,\dots,n_x$ corresponding to the $q_x^{\ell}$ first $g$-order statistics $(Z_{\ell,(1)},\dots,Z_{\ell,(q_x^{\ell})})$. Let $E_n$ denote the following event: \begin{equation*} E_{n} =E_{n_y,y} \cap E_{n_x,x} where E_{n_y,y}:=\left\{\bigcap_{\ell=1}^{L}M_y^{\ell} =\emptyset \right\} and E_{n_x,x}:=\left\{\bigcap_{\ell=1}^{L}M_x^{\ell} =\emptyset \right\} . \end{equation*} In words, $E_n$ means that the subsets of the data used in each of the $L$ tests have no observations in common. We begin by showing that \begin{equation} P\{E_{n}\} \rightarrow 1 . \end{equation} Since the two sample are independent, $P\{E_{n}\} = P\{E_{n_y,y}\}P\{E_{n_x,x}\}$ and so we only prove $P\{E_{n_y,y}\}\to 1$ as the other case is analogous. Let $\varepsilon :=\frac{1}{2}\min \left\{ \left\vert z_{\ell }-z_{\ell'}\right\vert :\ell \not=\ell',~\ell ,\ell'=1,\ldots ,L\right\}>0$. For each $\ell =1,\ldots ,L $, let $B_{\ell,y}=\sum_{i=1}^{n_y}I\{|Z_{i}-z_{\ell}|\le \varepsilon \}$, and note that \begin{equation*} \bigcap_{\ell =1}^{L}\left\{ B_{\ell,y}\geq q^{\ell }_y \right\} \subseteq E_{n} . \end{equation*} From here, we have that (ref) follows if we show that \begin{equation} P\{B_{\ell,y}<q^{\ell}_y\}\rightarrow 0 for all \ell =1,\ldots ,L . \end{equation} By Assumption (ref) and $\max_{\ell=1,\ldots ,L}q^{\ell}$ being bounded, $\exists N(\varepsilon )$ s.t.\ for all $n_y\geq N\left( \varepsilon \right) $, \begin{equation} 0<\frac{1}{2} P\{ |Z_{i}-z_{\ell}| \le \varepsilon \}\leq P\{ |Z_{i}-z_{\ell}|\le \varepsilon \} -\max_{\ell=1,\ldots ,L}\frac{q^{\ell}_y}{n_y} . \end{equation} Then, for all $n_y\geq N( \varepsilon ) $, we have \begin{align*} P\{ B_{\ell,y}<q^{\ell}_y\} &=P\left\{ {B_{\ell,y}}/{n_y} -P\{ |Z_{i}-z_{\ell}|\le \varepsilon \} <{q^{\ell}_y}/{n_y} - P\{|Z_{i}-z_{\ell}|\le \varepsilon \} \right\} \\ &\overset{(1)}{\geq }P\{ \vert {B_{\ell,y}}/{n_y}-P\{|Z_{i}-z_{\ell}| \le \varepsilon\} \vert > \frac{1}{2} P\{ |Z_{i}-z_{\ell}| \le \varepsilon \}\} \overset{(2)}{\rightarrow }0 , \end{align*} as desired, where (1) holds by (ref) and (2) holds by the LLN as $n_y \to \infty$, as $B_{\ell,y}=\sum_{i=1}^{n_y}I\{|Z_{i}-z_{\ell}|\le \varepsilon \} \sim Bi(n_y,P\{ |Z_{i}-z_{\ell}|\le \varepsilon \} ) $. We are now ready to prove the desired result. For any $(s_{1}^{\ell},\dots,s_{q_y^\ell}^{\ell},s_{q_y^\ell + 1}^{\ell},\dots,s_{q^\ell}^{\ell})\in \mathbf{R} ^{q^\ell}$ for each $\ell =1,\dots,L$ with $q^\ell = q_y^\ell+q_x^\ell$, we have \begin{align} &P\left\{ \cap_{\ell =1}^{L} \right. \left. \cap_{i=1}^{q^{\ell }} \left\{ S_{n,i}^{\ell }\leq s_{i}^{\ell }\right\} \right\} \notag\\ &\overset{(1)}{=}E\left[ \left. P\left\{ \bigcap_{\ell =1}^{L}\left\{ \bigcap_{i=1}^{q_{y}^{\ell }}\left\{ Y_{n,\left[ i\right] }^{\ell }\leq s_{i}^{\ell }\right\} \bigcap \bigcap_{j=1}^{q_{x}^{\ell }}\left\{ X_{n,\left[ j\right] }^{\ell }\leq s_{j+q_{y}^{\ell }}^{\ell }\right\} \right\} \right\vert \mathcal{A}\right\} \right] \notag \\ &\overset{(2)}{=}E\left[ \prod_{\ell =1}^{L}\left. P\left\{ \left\{ \bigcap_{i=1}^{q_{y}^{\ell }}\left\{ Y_{n,\left[ i\right] }^{\ell }\leq s_{i}^{\ell }\right\} \bigcap \bigcap_{j=1}^{q_{x}^{\ell }}\left\{ X_{n,\left[ j\right] }^{\ell }\leq s_{j+q_{y}^{\ell }}^{\ell }\right\} \right\} \right\vert \mathcal{A}\right\} I\{E_{n}\} \right] +o_{n}\left( 1\right) \notag \\ &\overset{(3)}{=}E\left[ \prod_{\ell =1}^{L}\left. P\left\{ \left\{ \bigcap_{i=1}^{q_{y}^{\ell }}\left\{ Y_{n,\left[ i\right] }^{\ell }\leq s_{i}^{\ell }\right\} \bigcap \bigcap_{j=1}^{q_{x}^{\ell }}\left\{ X_{n,\left[ j\right] }^{\ell }\leq s_{j+q_{y}^{\ell }}^{\ell }\right\} \right\} \right\vert \mathcal{A}\right\} \right] +o_{n}\left( 1\right) \notag \\ &\overset{(4)}{=}E\left[ \prod_{\ell =1}^{L}\prod_{i=1}^{q_{y}^{\ell }}F_{Y}(s_{i}^{\ell }|Z_{n_{y},\left( i\right) }^{\ell })\prod_{j=1}^{q_{x}^{\ell }}F_{X}(s_{j+q_{y}^{\ell }}^{\ell }|Z_{n_{x},\left( j\right) }^{\ell })\right] , \end{align} where (1) holds by the LIE with $\mathcal{A}$ equal to the sigma-algebra generated by the $Z$ observations from both samples, (2) by (ref) and the fact $E_{n}$ implies that the subsets of the data used in each of the $L$ tests have no observations in common, so they are independent conditional on $\mathcal{A}$, (3) by (ref), and (4) by repeating the arguments in the proof of Theorem (ref). Next, we show that \begin{align} Z_{n_{y},\left( i\right) }^{\ell }\overset{p}{\rightarrow }z_\ell for all i=1,\dots,q_{y} and \ell =1,\ldots ,L ,\\ Z_{n_{x},\left( j\right) }^{\ell }\overset{p}{\rightarrow }z_\ell for all j=1,\dots,q_{x}\text{ and } \ell =1,\ldots ,L . \end{align} We only show (ref), as (ref) can be shown analogously. To this end, fix $\ell =1,\ldots ,L$ arbitrarily. We prove the result by complete induction on $i=1,\dots,q_{y}$. Take $i=1$ and fix $\epsilon>0$ arbitrarily. Then, \begin{align} P\{|Z_{n_{y},(1)}^{\ell }-z_{\ell }| <\varepsilon \} &=P\{\text{at least }1 \text{ of }\{Z_{i}:1\leq i\leq n_{y}\}\text{ is s.t }\{|Z_{i}-z_{\ell }|<\varepsilon \}\} \nonumber \\ &\overset{(1)}{=}\sum_{u=1}^{n_{y}}\binom{n_{y}}{u}P\{|Z-z_{\ell }| <\varepsilon \}^{u}[1-P\{|Z-z_{\ell }|<\varepsilon \}]^{n_{y}-u} \nonumber\\ & \overset{(2)}{=}1-P\{|Z-z_{\ell }|\geq \varepsilon \}^{n_{y}} \overset{(3)}{\rightarrow }1 , \end{align} as desired, where (1) holds by the fact that $\{Z_{i}:1\leq i\leq n_{y}\}$ are identically distributed, (2) by the Binomial Theorem, and (3) by Assumption (ref). For the inductive step, we assume $Z_{n_{y},(j)}^{\ell }-z_{\ell }=o_{p}(1)$ for $j\in \{1,\ldots ,q_{y}-1\}$, and prove that $ Z_{n_{y},(j+1)}^{\ell }-z_{\ell }=o_{p}(1)$. For this, consider the following derivation, \begin{align*} &P\{|Z_{n_{y},(j)}^{\ell }-z_{\ell }| <\varepsilon \}\\ &=P\{\text{at least }j \text{ of the }\{Z_{i}:1\leq i\leq n_{y}\}\text{ are s.t. }\{|Z_{i}-z_{\ell }|<\varepsilon \}\} \\ &\overset{(1)}{=}\sum_{u=j}^{n_{y}}\binom{n_{y}}{u}P\{|Z-z_{\ell }| <\varepsilon \}^{u}P\{|Z-z_{\ell }|\geq \varepsilon \}^{n_{y}-u} \\ & \overset{(2)}{=}P\{|Z_{n_{y},(j+1)}^{\ell }-z_{\ell }|<\varepsilon \}+ \binom{n_{y}}{j}P\{|Z-z_{\ell }|< \varepsilon \}^{j}P\{|Z-z_{\ell }|\geq \varepsilon \}^{n_{y}-j} , \end{align*} where (1) holds by the fact that $\{Z_{i}:1\leq i\leq n_{y}\}$ are identically distributed and (2) by the Binomial Theorem. The desired result then follows from assumption that $Z_{n_{y},(j)}^{\ell }-z_{\ell }=o_{p}(1)$ and the following derivation \begin{align*} &\binom{n_{y}}{j}P\{|Z-z_{\ell }|<\varepsilon \}^{j}[1-P\{|Z-z_{\ell }|<\varepsilon \}]^{n_{y}-j}\\ &\leq n_{y}^{j}P\{|Z-z_{\ell }|\geq \varepsilon \}^{n_{y}-j}=\left[ exp^{\frac{j\ln n_{y}}{n_{y}-j}}P\{|Z-z_{\ell }|\geq \varepsilon \}\right] ^{n_{y}-j}\overset{(1)}{\rightarrow} 0 , \end{align*} where (1) follows from Assumption (ref) and noticing that $\exists N$ s.t.\ $exp^{\frac{ j\ln n_{y}}{n_{y}-j}}P\{|Z-z_{\ell }|\geq \varepsilon \}\leq P\{|Z-z_{\ell }|\geq \varepsilon \}<1$ for all $n_{y}>N$ and any $j\in \{1,\dots ,q_{y}-1\}$. By (ref), (ref), and Assumption (ref), we have \begin{equation} \lim_{n\rightarrow \infty }P\left\{ \bigcap_{\ell =1}^{L}\bigcap_{i=1}^{q_{\ell }}\left\{ S_{n,i}^{\ell }\leq s_{i}^{\ell }\right\} \right\} = \prod_{\ell =1}^{L}\prod_{i=1}^{q_{y}^{\ell }}F_{Y}(s_{i}^{\ell }|z_{\ell })\prod_{j=1}^{q_{x}^{\ell }}F_{X}(s_{j+q_{y}^{\ell }}^{\ell }|z_{\ell }) . \end{equation} By definition of convergence in distribution, (ref) implies (ref), as desired.
theoremLet $S$ be the random variable in Theorem (ref). Then, for any $P\in \mathbf P_0$ and $\alpha\in(0,1)$, we obtain \[ E_P[ \phi (S)] \leq \bar{\alpha}\le \alpha ~, \] where \begin{equation} \bar{\alpha} := P\left\lbrace \sup_{u\in\mathbf (0,1)} \left(\frac{1}{q_y}\sum_{j=1}^{q_y} I\{U_{j} \le u\}-\frac{1}{q_x}\sum_{j=q_y+1}^{q} I\{U_{j} \le u\} \right) > c_{\alpha}(q_y,q) \right\rbrace . \end{equation} Moreover, the first inequality becomes an equality under $P$ such that $F_Y(t|z_{0})=F_X(t|z_{0})$ for all $t\in \mathbf R$ and these are continuous functions of $t\in \mathbf R$.
proofRecall that $S = (S_1, \dots,S_{q_y},S_{q_y+1},\dots,S_{q})$ are independent, and such that $S_j \sim F_Y(t|z_{0})$ for all $j =1,\dots, q_y$ and $S_{q_y+j} \sim F_X(t|z_{0})$ for all $j=1,\dots, q_x$. Denote by $Q_Y(\cdot|z_{0})$ and $Q_X(\cdot|z_{0})$ the quantile functions of $F_Y(t|z_{0})$ and $ F_X(t|z_{0})$. Consider the following argument: { \begin{align*} &E_P[ \phi (S)] \\ &= P\left\lbrace \max_{k \in \{1,\dots,q_y\}} \left(\frac{1}{q_y}\sum_{j=1}^{q_y} I\{S_j \le S_k\}-\frac{1}{q_x}\sum_{j=q_y+1}^{q} I\{S_j \le S_k\} \right) > c_{\alpha}(q_y,q) \right\rbrace \\ &\overset{(1)}{=} P\left\lbrace \sup_{t\in\mathbf Y} \left(\frac{1}{q_y}\sum_{j=1}^{q_y} I\{Q_Y(U_j|z_{0}) \le t\}-\frac{1}{q_x}\sum_{j=q_y+1}^{q} I\{Q_X(U_j|z_{0}) \le t\} \right) > c_{\alpha}(q_y,q) \right\rbrace \\ &\overset{(2)}{=} P\left\lbrace \sup_{t\in\mathbf Y} \left(\frac{1}{q_y}\sum_{j=1}^{q_y} I\{U_{j} \le F_Y(t|z_{0})\}-\frac{1}{q_x}\sum_{j=q_y+1}^{q} I\{U_{j} \le F_X(t|z_{0})\} \right) > c_{\alpha}(q_y,q) \right\rbrace \\ &\overset{(3)}{\leq} P\left\lbrace \sup_{t\in\mathbf Y} \left(\frac{1}{q_y}\sum_{j=1}^{q_y} I\{U_{j} \le F_X(t|z_{0})\}-\frac{1}{q_x}\sum_{j=q_y+1}^{q} I\{U_{j} \le F_X(t|z_{0})\} \right) > c_{\alpha}(q_y,q) \right\rbrace \\ &\overset{(4)}{=} P\left\lbrace \sup_{u\in \mathcal{U}} \left(\frac{1}{q_y}\sum_{j=1}^{q_y} I\{U_{j} \le u\}-\frac{1}{q_x}\sum_{j=q_y+1}^{q} I\{U_{j} \le u\} \right) > c_{\alpha}(q_y,q) \right\rbrace \\ &\overset{(5)}{\leq} P\left\lbrace \sup_{u\in\mathbf (0,1)} \left(\frac{1}{q_y}\sum_{j=1}^{q_y} I\{U_{j} \le u\}-\frac{1}{q_x}\sum_{j=q_y+1}^{q} I\{U_{j} \le u\} \right) > c_{\alpha}(q_y,q) \right\rbrace \\ &\overset{(6)}{=}\bar{\alpha} , \end{align*}} where (1) follows from the quantile transformation and the fact that replacing $\max_{k \in \{1,\dots,q_y\}}$ with $\sup_{t\in \mathbf Y}$ does not affect the magnitude of the test statistic, (2) follows from pollard:02, (3) from $P\in \mathbf P_0$, (4) from a simple change of variables and $\mathcal{U} := \cup_{t\in \mathbf Y}\{u = F_X(t|z_{0}) \}$, (5) from $ \mathcal{U}\subseteq (0,1)$, and (6) from the definition of $c_{\alpha}(q_y,q)$ and $\bar{\alpha}$, and the fact that $\{U_{j}:j=1,\dots,q\}$ are i.i.d.\ distributed as $U( 0,1) $. To conclude the proof, it suffices to show that: (i) inequality (3) holds as an equality under the condition \( F_Y(t \mid z_{0}) = F_X(t \mid z_{0}) \) for all \( t \in \mathbf{R} \), and (ii) inequality (5) holds as an equality when these functions are continuous in \( t \). The first claim is immediate. For the second, it follows from the fact that the continuity of the CDFs implies \( \mathcal{U} = (0,1) \). By elementary properties of CDFs, $ \lim_{t \to -\infty} F_X(t \mid z_{0}) = 0$ and $\lim_{t \to \infty} F_X(t \mid z_{0}) = 1$. By the intermediate value theorem, for any \( u \in (0,1) \), there exists \( t \in \mathbf{R} \) such that \( u = F_X(t \mid z_{0}) \), implying \( u \in \mathcal{U} \), as desired.

Auxiliary lemmas

lemmaSuppose that $S_n \sim P$ with $P$ satisfying Assumptions (ref), (ref), and (ref) hold. Consider a sequence of random vectors $\{\tilde{S}_{n}:1\le n< \infty \}$ and a random vector $\tilde{S}$ defined on a common probability space $ (\Omega ,\mathcal{A},\tilde{P})$ such that \begin{equation} \Tilde{S}_{n}\overset{d}{=}S_{n},\quad \Tilde{S}\overset{d}{=}S,\quad and \quad \Tilde{S}_{n}\overset{a.s.}{\to }\Tilde{S} . \end{equation} Then, for any $j\in \{1,\dots ,q\}$, \begin{equation*} \tilde{P}\{\tilde{S}_{n,j}\neq \tilde{S}_{j}, \tilde{S}_{j}\in \mathcal{D} \}=o_{n}(1) . \end{equation*}
proofWe focus on an arbitrary \( j \in \{1, \dots, q_{y}\} \). The argument for \( j \in \{q_{y} + 1, \dots, q\} \) follows analogously by replacing \( Y \) with \( X \). By Lemma (ref), it suffices to show that \begin{equation} \sup_{t \in \mathbb{R}} \big| F_{\tilde{S}_{n,j}}(t) - F_{\tilde{S}_{j}}(t) \big| = o_{n}(1) , \end{equation} where \( F_{\tilde{S}_{n,j}} \) and \( F_{\tilde{S}_{j}} \) denote the distribution functions of \( \tilde{S}_{n,j} \) and \( \tilde{S}_{j} \) in \( (\Omega, \mathcal{A}, \tilde{P}) \). By the proof in Theorem (ref) (with $L=1$ and $\mathcal{Z}=\{z_{0}\}$), we have $ F_{S_{j}}(t)=F_{Y}(t|z_{0})$ and $ F_{S_{n,j}}(t)=E[F_{Y}(t|Z_{n_{y},(j)})]$ where $ Z_{n_{y},(j)}\overset{p}{\to }z_{0}$. By these and (ref), (ref) follows from \begin{equation} \sup_{t\in \mathbf{R}}\vert E[F_{Y}(t|Z_{n_{y},(j)})]-F_{Y}(t|z_{0})\vert =o_{n}(1) . \end{equation} For any $\varepsilon >0$, \begin{align*} E[F_{Y}(t|Z_{n_{y},(j)})] &=\int_{\vert z-z_{0}\vert \leq \varepsilon }F_{Y}(t|z)dP_{Z_{n_{y},(j)}}( z) +\int_{\vert z-z_{0}\vert >\varepsilon }F_{Y}(t|z)dP_{Z_{n_{y},(j)}}( z) \\ &\leq \int_{\vert z-z_{0}\vert \leq \varepsilon }F_{Y}(t|z)dP_{Z_{n_{y},(j)}}( z) +P( \vert Z_{n_{y},(j)}-z_{0}\vert >\varepsilon ) \\ &\overset{(1)}{=} \int_{\vert z-z_{0}\vert \leq \varepsilon }F_{Y}(t|z)dP_{Z_{n_{y},(j)}}( z) +o_{n}(1) , \end{align*} where (1) holds by $Z_{n_{y},(j)}\overset{p}{\to }z_{0}$. Then, \begin{equation*} \sup_{t\in \mathbf{R}}\vert E[F_{Y}(t|Z_{n_{y},(j)})]-F_{Y}(t|z_{0})\vert \leq \sup_{t\in \mathbf{R} }\sup_{\vert z-z_{0}\vert \leq \varepsilon }\vert F_{Y}(t|z)-F_{Y}(t|z_{0})\vert +o_{n}(1) . \end{equation*} Fix $\delta >0$ arbitrarily. By Assumption (ref), $\exists \varepsilon >0$ such that $ \sup_{t\in \mathbf{R}}\sup_{\vert z-z_{0}\vert \leq \varepsilon }\vert F_{Y}(t|z)-F_{Y}(t|z_{0})\vert <\delta /2$. For all large enough $n_y$, the right-hand side is bounded above by $\delta $. Since the choice of $\delta $ was arbitrary, (ref) follows.
lemmaLet $V_{1}$ and $V_{2}$ be independent random variables that are discontinuous at a finite set of points $\mathcal{D}_{1}$ and $\mathcal{D}_{2}$, respectively. Then, for any $\delta >0$, $\exists \varepsilon >0$ small enough s.t. \begin{align*} P\{ \cup _{d\in \mathcal{D}_{1}}\{ |V_{1}-d|<\varepsilon \} \cap \{V_{1}\in \mathcal{D}_{1}^{c}\}\} &<\delta , \\ P\{ \{ |V_{1}-V_{2}|<\varepsilon \} \cap \{V_{1}\in \mathcal{D}_{1}^{c}\}\cap \{ V_{2}\in \mathcal{D}_{2}^{c}\} \} &<\delta . \end{align*}
proofSet $\bar{\varepsilon}:=\{ \min \vert d-\tilde{d}\vert :d,\tilde{d}\in \mathcal{D}_{1}\cap d\not=\tilde{d}\} >0$. For any $\varepsilon \in (0,\bar{\varepsilon}/2),$ \begin{align*} &P\{ \cup _{d\in \mathcal{D}_{1}}\{ |V_{1}-d|<\varepsilon \} \cap \{V_{1}\in \mathcal{D}_{1}^{c}\}\} \\ &\leq \sum_{d\in \mathcal{D}_{1}}P\{ \{|V_{1}-d|<\varepsilon \}\cap \{V_{1}\in \mathcal{D}_{1}^{c}\}\} \\ &\leq _{(1)}\sum_{d\in \mathcal{D}_{1}}P\{ V_{1}\in (d-\varepsilon ,d)\cup (d,d+\varepsilon )\} \overset{(2)}{=} o_{\varepsilon }(1) , \end{align*} where (1) holds by $\varepsilon \in (0,\bar{\varepsilon}/2)$ and (2) by the fact that $(d-\varepsilon ,d)\cap (d,d+\varepsilon )$ has no discontinuities in the CDF. The right-hand side is less than $\delta $ by making $\varepsilon $ arbitrarily small, as desired. Also, for any $\varepsilon \in (0,\bar{\varepsilon}/2)$, \begin{align*} &P\{ \{ |V_{1}-V_{2}|<\varepsilon \} \cap \{V_{1}\in \mathcal{D}_{1}^{c}\}\cap \{ V_{2}\in \mathcal{D}_{2}^{c}\} \} \\ &\overset{(1)}{=} \int_{\mathcal{D}_{2}^{c}}P\{ \{ |V_{1}-v_{2}|<\varepsilon \} \cap \{V_{1}\in \mathcal{D} _{1}^{c}\}\} dP_{V_{2}}( v_{2}) \overset{(2)}{=} o_{\varepsilon }(1) , \end{align*} where (1) holds by $V_{1}\perp V_{2}$, and (2) holds by Lemma (ref) and the dominated convergence theorem. The right-hand side is less than $\delta $ by making $\varepsilon $ arbitrarily small, as desired.
lemmaConsider a random variable $V$ whose CDF is discontinuous at a finite set of points $\mathcal{D}$. Then, for any $v\in \mathbf{R}$, \begin{equation*} P\{ \{ |V-v|<\varepsilon \} \cap \{V\in \mathcal{D}^{c}\}\} =o_{\varepsilon}(1) . \end{equation*}
proofSet $\bar{\varepsilon}:=\{ \min \vert d-\tilde{d}\vert :d,\tilde{d}\in \mathcal{D}_{1}\cap d\not=\tilde{d}\} >0$. Fix $\varepsilon <\bar{\varepsilon}/2$. There are two possibilities: $v\in \mathcal{D}$ or $v\not\in \mathcal{D}$. First, consider $v\in \mathcal{D}$. In this case, \begin{equation*} P\{ |V-v|<\varepsilon \} \cap \{V\in \mathcal{D}^{c}\} \leq P\{ V\in ( v-\varepsilon ,v) \cap ( v,v+\varepsilon ) \} \overset{(1)}{=} o_{\varepsilon}(1) , \end{equation*} where (1) holds because $(d-\varepsilon ,d)\cap (d,d+\varepsilon )$ has no mass points. Second, consider $ v\not\in \mathcal{D}$. Then, \begin{equation*} P\{ \{ |V-v|<\varepsilon \} \cap \{V\in \mathcal{D} ^{c}\}\} \leq F_{V}(v+\varepsilon )-F_{V}(v-\varepsilon )\overset{(2)}{=} o_{\varepsilon }(1) , \end{equation*} where (1) by the fact that $( v-\varepsilon ,v+\varepsilon ] $ has no mass points, so $F_{V}$ is continuous on that interval.
lemmaLet $\{V_{n}:n\ge 1\}$ be a sequence of random variables that satisfy $V_{n}\overset{p}{\to }V$, where $V$ is a random variable whose CDF is discontinuous at a finite set of points $\mathcal{D}$. Furthermore, assume $\sup_{t\in \mathbf{R}}|F_{V_{n}}(t)-F_{V}(t)|\to 0$. Then, \begin{equation*} P\{\{V\in \mathcal{D}\} \cap \{V_{n}\neq V\}\}\to 0 . \end{equation*}
proofFix $\delta >0$ arbitrarily. It suffices to find $N=N(\delta) $ such that $P\{V\in \mathcal{D},V_{n}\neq V\}<\delta $ for all $n>N$. Set $\bar{\varepsilon}:=\{ \min \vert d-\tilde{d}\vert :d, \tilde{d}\in \mathcal{D}\cap d\not=\tilde{d}\} >0$. For any $ \varepsilon \in (0,\bar{\varepsilon}/2),$ consider the following argument. \begin{align*} &P\{\{V\in \mathcal{D}\} \cap \{V_{n}\neq V\}\} \\ &=\sum_{d\in \mathcal{D}}P\{\{V=d\} \cap \{V_{n}\neq d\} \} \\ & \overset{(1)}{=} \sum_{d\in \mathcal{D}}P\{\{V=d\} \cap \{V_{n}\neq d\} \cap \{|V_{n}-d|<\varepsilon\}\}+o_{n}(1) \\ & \leq \sum_{d\in \mathcal{D}}P\{V_{n}\in (d-\varepsilon ,d)\cap (d,d+\varepsilon )\}+o_{n}(1) \\ & \overset{(2)}{=} \sum_{d\in \mathcal{D}}P\{V \in (d-\varepsilon ,d)\cap (d,d+\varepsilon )\}+o_{n}(1) \\ & \overset{(3)}{=} o_{\varepsilon }(1)+o_{n}(1) , \end{align*} where (1) holds by $V_{n}\overset{p}{\to }V$, (2) by $\sup_{t\in \mathbf{R}}|F_{V_{n}}(t)-F_{V}(t)|=o_{n}(1) $, and (3) by the fact that $(d-\varepsilon ,d)\cap (d,d+\varepsilon )$ has no discontinuities in the CDF. For all large enough $n$ and small enough $\varepsilon $, the right-hand side is bounded by $\delta $, as desired.