EconBase
← Back to paper

Probability of worthwhile effect of monotone-response treatments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

93,909 characters · 12 sections · 43 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Probability of worthwhile effect of monotone-response treatments

abstractExperiments may, by design, prevent one from observing on a single subject both the response to a treatment and to its absence. Because of this, marginal distributions for both cases may be observable but not their joint distribution, thus obscuring the distribution of the treatment effect. We examine the case where we impose that the treatment effect is nonnegative, also called monotone treatment response, a common assumption relevant to many practical applications. We solve the problems of best- and worst-case probabilities that the treatment effect exceeds a given value, using an explicit construction for the dependence scheme in each case. Such problems can equivalently be described, in different contexts, as risk aggregation under dependence uncertainty and an order constraint, and as optimal transport with a particular cost function. Keywords: clinically relevant benefit, nonnegative treatment effect, counterfactual causality, dependence uncertainty, optimal transport.

Introduction

In a simple experiment, we study the effect of a treatment. The responses of subjects to the treatment, $Y$, and to the absence of treatment, $X$, are monitored. Responses $X$ and $Y$ are random variables; their distributions, respectively $\mu$ and $\nu$, may either be interpreted as that of the (deterministic) treatment response within a population or of the aleatoric response of any single subject. The treatment effect is $Y-X$.

Most considerations on the treatment effect are answered by elementary statistics if we are able to observe both responses for every subject. However, this is often not possible for experimental studies outside of controlled environments or conducted over a long time frame, or sometimes due to ethical considerations. Consider, for instance, a study on the effects of malnutrition on children's health: a researcher cannot require subjects to be malnourished for the length of the study. The researcher must rather rely on observational data, meaning that data on responses to treatment and non-treatment are collected separately. This leads to two major impediments for statistical purposes:

enumerate• The observed marginal distributions, from each group's data, may not be the true marginal distributions of $X$ and $Y$, as some biases may arise, e.g., from selection. • The separate collection of data obscures the coupling between $X$ and $Y$, and hence also the distribution of the treatment effect $Y-X$.

Let us provide a simple example to illustrate this second impediment.

exampleFor an experiment, we have the following samples respectively without and with treatment: $\mathbf{x}=\{1,2,3,4\}$ and $\mathbf{y}=\{2,3,4,5\}$. One may think that the treatment has had an effect of 1 for every subject, thus supposing the coupling $ \{(1,2), (2,3), (3,4), (4,5)\}$; another may rather think that the treatment was only effective on one subject, with an effect of 4, thus supposing the coupling $ \{(1,5), (2,2), (3,3), (4,4)\}$. Without additional information on the relation between $\mathbf{x}$ and $\mathbf{y}$, we cannot distinguish either coupling from the data.

Experimental designs based on observational data are prominent in many fields, such as sociology, statistics and econometrics; and their study led to the counterfactual conception of causal effects. One may consult BP94, WM99, D00, S09, and D15 for an introduction and some (philosophical, statistical) considerations on counterfactual causality.

As underscored in all of these works, studies involving counterfactual causality will almost always examine the average treatment effect, $\mathbb{E}[Y-X]$. This has the marked advantage of dispelling impediment (S2) because $\mathbb{E}[Y-X]$ is unaffected by the dependence between $Y$ and $X$. Most seminal statistical works on counterfactual causality thus exclusively focused on mitigating (S1); see notably the series of works M94,M95, M97, M07, MP00, MP09.

Average treatment effect, however, may be ill-suited to some situations. In Example (ref), both couplings yield the same average effect, but may lead to different decisions. If, for example, treatments are antidepressants and responses are their side effect on blood pressure, the first coupling, with a constant 1 effect, would probably lead to the drug being prescribed by a physician, as they may judge that the side effect is tolerable. The second coupling, however, would probably incite physicians to avoid prescribing the drug, as they deem too risky to potentially see a dramatic increase of blood pressure. See F99 and YG24 for further discussions on the limitations of average treatment effect in medical trials. Let us present other examples illustrating this rhetoric.

exampleA tutoring service believes that parents will renew their subscription if their children perform at least 10 points better on their tests. The tutoring service wants to obtain bounds on the number of renewals, given the tutored students' grades compared to the non-tutored students.
exampleEx-smokers will continue to abstain from smoking only if they see an increase in their overall health. Public health services want to assess for the worst case of smoking relapse given health data.

We remark that the analytics under study in these examples do not depend on the actual responses $X$ or $Y$ but solely on the treatment effect $Y-X$. However, the average treatment effect is unable to answer the threshold-related considerations.

For such considerations, one may rather consider the probability of a worthwhile effect; see, e.g., F99. Let $k\geq 0$ be the level above which a treatment effect is considered clinically relevant. The probability of worthwhile effect for a given $k$ is the value $\mathbb{P}(Y-X > k)$. F99 discussed probabilities of worthwhile effect as an alternative to p-values in the context of medical treatments. An issue for p-values is that statistical significance of the treatment effect is unrelated to whether this effect is clinically worthwhile. F99 argued that a medical researcher should not heuristically determine what constitutes a worthwhile effect based on the collected data; probabilities of worthwhile effect effectively require the researcher to specify beforehand the threshold for worthwhile effect $k$. For instance, in our earlier example, the physician would establish which threshold $k$ of increase in blood pressure is dangerous and then inquire about the probability of worthwhile effect for $k$. In Examples (ref) and (ref), thresholds for worthwhile effects are specified as $k=10$ and $k=0$ respectively.

When examining probabilities of worthwhile effect instead of the average treatment effect, impediment (S2) comes back: the value $\mathbb{P}(Y-X > k)$ is impacted by the dependence relation between $X$ and $Y$. Since there is already a long string of literature on how to mitigate (S1), as we specified above, we will focus on (S2) and assume that the treatment has fully identified marginals, to remove (S1).

Because of (S2), probabilities of worthwhile effect cannot be computed unless we specify the dependence relation between $X$ and $Y$. One common assumption is to suppose that $X$ and $Y$ are comonotonic.\footnote{Two random variables $Z,W$ are comonotonic if $(Z(\omega) - Z(\omega^{\prime}))(W(\omega)- W(\omega^{\prime}))\geq 0$ for all $\omega, \omega^{\prime}\in\Omega$.} If $\nu$ is a translation of $\mu$, then this is called the Treatment-Unit Additivity assumption D15, which effectively means that $Y-X$ is degenerate. Another common assumption is to suppose that $X$ and $Y$ are independent D00. Letting $k=0$, the probability of worthwhile effect in that case corresponds, up to a small adjustment for ties, to the Mann-Whitney parameter, used in the Wilcoxon-Mann-Whitney test for causal effects W45,MW47 and later introduced by D16 under the name D-value. Not only are independence and comonotonicity very different dependence schemes, they also mean very strong assumptions on the treatment's effect. A researcher may well not want to commit to either. In particular, the independence assumption is heavily criticized by GFBSFGR20, and Hand's paradox H92 is an example of how misleading it can be. As GFBSFGR20 note, to rely on such assumptions for causal inferences marks an obliviousness to impediment (S2), perhaps forgotten because the average treatment effect bypasses it. Consistent with such criticism, our approach will rather consist of deriving worst- and best-case probabilities of worthwhile effect under the full uncertainty stemming from (S2).

In this paper, we will study probabilities of worthwhile effect of treatments with the additional assumption of a monotone response. Monotone treatment response, first studied by M97, stipulates that responses weakly increase with the strength of treatment. For a binary treatment, this means $Y\geq X$.\footnote{Such relations for random variables are to be interpreted as holding almost surely.} Put equivalently: the treatment effect $Y-X$ is nonnegative almost surely, meaning that the treatment cannot decrease the subject's response.

This assumption may seem strong, but is far from unreasonable in many contexts, for instance in a study of the effects of private tutoring on academic success H14. Other examples include the effects of schooling on elders' cognition ABFFFKKS25, of a smaller health-insurance deductible on the number of doctor visits GS06, of moral suasion on tax evasion BCST20, of food assistance programs on food security KPR16, GKP17, C22, and of food security on children's health GK09, GKP12. In Example (ref), both couplings satisfy the assumption of monotone treatment response $ y_i \geq x_i$ for every element $i$. Thus, it is clear that, while limiting the number of possible couplings, the monotone-treatment-response assumption alone in general does not suffice to fully identify the joint distribution of responses.

We examine worst- and best-case probabilities of worthwhile effect respectively defined as

align[align omitted — 384 chars of source]

where $X\buildrel \mathrm{d} \over \sim \mu$ means that the distribution of $X$ is $\mu$.

We do not make further assumptions than full identification of marginals and monotone treatment response; in particular, we do not impose any shape or specific features for the marginal distributions, as long as they allow for monotone treatment response to be possible. The set of possible joint distributions given this constraint is studied in AMZ20.

In the remainder of this section, we examine problems (ref)--(ref) from different perspectives: of risk aggregation and of optimal transport; we also introduce some notation. In Section (ref), we solve the problems for binary treatments. In Section (ref), we solve the problems for non-binary treatments, where intermediate degrees of treatment constitute a source of partial identification of the joint distribution for the full treatment, from the monotone-treatment-response assumption. Section (ref) concludes.

Formulation in optimal transport

In optimal transport theory, the primal Kantorovich problem consists of minimizing an expected cost over all joint distributions with given marginals. See RR06, V09 and F24 for an introduction to optimal transport. We rewrite problems (ref)--(ref) in terms of optimal transport:

align[align omitted — 484 chars of source]

Note how the second terms of the cost functions $c$ and $c^*$ enforce the ordering constraint from the monotone-treatment-response assumption. For atomic distributions (see Section (ref)), one may replace $\infty$ by a finite number that is large enough. Optimal transport with this ordering constraint is studied by NW22 under the name directional optimal transport. They solve problems similar to (ref)--(ref), but with the first term of $c$ or $c^*$ being submodular or supermodular.\footnote{A function $(x,y)\mapsto f(x,y)$ is submodular if $f(x_1,y_1) + f(x_2,y_2) \leq f(x_1\wedge x_2, y_1 \wedge y_2) + f(x_1\vee x_2, y_1 \vee y_2)$ for all $(x_1,y_1),(x_2,y_2)$ in its domain. It is supermodular if the reverse inequality holds for all $(x_1,y_1),(x_2,y_2)$ in its domain.} Here, the functions $(x,y)\mapsto \mathds{1}_{\{y-x > k\}}$ and $(x,y)\mapsto \mathds{1}_{\{y-x \leq k\}}$ are neither submodular nor supermodular. Although it will not be our approach to solving them, strong duality holds for problems (ref)--(ref); see Appendix (ref).

Formulation in risk aggregation under dependence uncertainty

In the risk-management literature, the framework obtained by dispelling (S1) but retaining (S2) is called dependence uncertainty: we suppose that the marginal distributions of $(X,Y)$ are known, but the dependence relation between its components is unknown. See EP10 for considerations and motivations for such a framework in the context of risk management. CLW22 proposed and solved the following problems:

align[align omitted — 366 chars of source]

Note that the only difference between their problems and ours, beyond the weak inequality in (ref), is that $X+Y$ in (ref)--(ref) takes the place of $Y-X$ in (ref)--(ref). This difference has a significant impact on the optimal couplings, making the problems (ref)--(ref) completely different from (ref)--(ref); in particular, the results of CLW22 cannot be applied to our problems. We discuss their connections in Appendix (ref).

Without the constraint of monotone treatment response, the literature on risk aggregation under dependence uncertainty is rich and extends beyond the case of two variables: most notably, R82 gives the solution to the worst-case $\mathbb{P}(X+Y>k)$ and best-case $\mathbb{P}(X+Y\geq k)$. BJW14 and MW15 examine the set of all possible distributions of the sum of several random variables with given marginal distributions; bounds on best- and worst-case values of the quantile of the sum are given in EWW15 and algorithms to numerically compute these values are developed in EPR13 and BLLW24.

Preliminaries and notation

All random variables live on an atomless probability space $(\Omega, \mathcal{F}, \mathbb{P})$. All order terms like “increasing" are in the weak sense. Write $\mathbb{R}_+=[0,\infty)$. Let $\mathcal{M}$ be the set of all finite Borel measures on $\mathbb{R}$, equipped with the usual order $\le$. Note that $(\mathcal{M},\leq)$ is a lattice; $\wedge$ and $\vee$ respectively mean the infimum and supremum operators on that lattice. The probability measure $\delta_x$ is the point-mass at $x\in \mathbb{R}$. The survival function of a measure $\mu\in\mathcal{M}$ is the function $S_{\mu} : \mathbb{R} \to \mathbb{R}_+ $ given by $S_\mu(x)= \mu((x,\infty))$. We write $\mu \le_{\rm st} \nu$ if $S_{\mu}\le S_{\nu}$. An equivalent property to $\mu \le_{\rm st} \nu$ is

equation*[equation* omitted — 134 chars of source]

Note that $\mu\le \nu$ implies $\mu\le_{\rm st } \nu$, and both relations are transitive. The quantile function of a measure $\mu\in\mathcal{M}$ is the function $F^{-1}_\mu : (0,\mu(\mathbb{R})) \to \mathbb{R}$ defined as $ F^{-1}_{\mu}(u) = \inf\{x\in\mathbb{R}: \mu(\mathbb{R})-S_{\mu}(x) \geq u\},$ $u\in(0,1).$ For a random variable $Z$, its distribution is the probability measure $\mu\in \mathcal M$ given by $\mu(A)=\mathbb{P}(Z\in A)$ for Borel $A\subseteq \mathbb{R}$, and we denote this by $Z\buildrel \mathrm{d} \over \sim \mu$. While the exposition on treatment effects in the Introduction considered only probability measures (instead of general measures), we do not necessarily make this assumption henceforth and allow for non-probability measures as well in most of our results.

Let $\mathbb H=\{(x,y)\in \mathbb{R}^2: y \ge x\}$. Let $\Pi$ be the set of all Borel measures on $\mathbb{R}^2$. For $\pi\in \Pi$, we let $P_1 (\pi)$ denote the first marginal of $\pi$ and $P_2( \pi)$ the second marginal of $\pi$. For $\mu,\nu \in \mathcal M$, define the set

equation[equation omitted — 174 chars of source]

which is a set of semi-couplings between $\mu$ and $\nu$ in the sense of HS13. For simplicity, we refer to any element in $ \mathcal{H}(\mu,\nu)$ as a coupling between $\mu$ and $\nu$. For $ \mathcal{H}(\mu,\nu)$ to be nonempty, it is necessary and sufficient that $\mu\le_{\rm st}\nu$; see MS02. Note that $\pi(\mathbb{R}^2) = \mu(\mathbb{R})\le \nu(\mathbb{R})$ for all $\pi\in \mathcal H(\mu,\nu)$ when $\mu \le_{\rm st} \nu$. AMZ20 and NW22 analyzed $ \mathcal{H}(\mu,\nu)$ when $\mu$ and $\nu$ are probability measures.

Probabilities of worthwhile effect for binary treatments

For measures $\mu$ and $\nu$ with $\mu\le_{\rm st} \nu$ that are not necessarily probability measures, we generalize problems (ref)--(ref) in the Introduction by rewriting them as

align[align omitted — 351 chars of source]

Our solutions to problems (ref)--(ref) are predicated on coupling algorithms. In this section, we present these algorithms, solving the problems for specific discrete distributions, and discuss the adequacy of approximating the solutions for the general discrete case and the continuous case this way. Although problems (ref) and (ref) look very similar, the solution to one does not follow directly from the other's by symmetry.

For $\mu$ and $\nu$ that are probability measures with finite means and $\mu\leq_{\rm st}\nu$, crude bounds for $Q^{\inf}_k(\mu,\nu)$ and $Q^{\sup}_k(\mu,\nu)$ are,

equation*[equation* omitted — 185 chars of source]

where the last inequality follows from Markov's inequality. The algorithms presented below leverage the identification of $\mu$ and $\nu$ to produce improved bounds, tailored to the specified marginal distributions.

Solutions for atomic measures with specific atom size

We first study a special case of discrete measures $\mu$ and $\nu.$ A measure $\mu$ is atomic with size $a>0$ if

equation[equation omitted — 80 chars of source]

for some $x_1, \ldots, x_m \in\mathbb{R}$. When $\mu $ has total mass $1$ (that is, $a=1/m$), $\mu$ is simply an empirical distribution. We call $x_1, \ldots, x_m$ the locations of atoms, $m$ the number of atoms, and $a$ the size of an atom; note that multiple atoms may be at identical locations. For atomic measures $\mu$ and $\nu$ with the same size (meaning that they share the same $a>0$), if $\mu\le_{\rm st}\nu$ holds, then $m\le n$, where $m$ is the number of atoms in $\mu$ and $n$ is the number of atoms in $\nu$.

As our first main result, the minimization in (ref) is attained by Algorithm (ref), and the maximization in (ref) is attained by Algorithm (ref). We explain the intuition behind these algorithms first. Throughout, for ease of exposition, let $x$ denote a generic atom location in $\mu$ and $y$ denote a generic atom location in $\nu$.

For the construction of the pairs $(x,y)$ in the optimal $\pi$ for (ref), we couple atoms in $\mu$ iteratively, starting from largest location to smallest location. For each atom, when its turn to be coupled comes, we check whether there is an uncoupled atom in $\nu$ allowing to satisfy $y-x\leq k$. If so, it couples to the one at the largest location among them. If it is unable to satisfy $y-x\leq k$, it sacrifices itself and takes down the atom in $\nu$ at the largest location; this effectively helps the next-to-be-coupled atom in $\mu$, by not being a nuisance, since that probability atom in $\nu$ would not allow any atom at smaller locations in $\mu$ to satisfy $y-x\leq k$ anyway. By always transporting to atoms in $\nu$ at the largest allowed locations, either when helping toward the objective $y-x\leq k$ or not, we make sure that probability mass at smaller locations in $\mu$ has the best chance of satisfying $y-x\leq k$ when its turn to be coupled comes. What could occur, however, would be an atom in $\mu$ satisfying $y-x\leq k$ preventing an atom at a smaller location to do so, but this situation results in effectively the same value for $\int \mathds{1}_{\{(x,y)\in\mathbb{R}^2: y-x>k\}}\mathrm{d}\pi$, $\pi \in \mathcal{H}(\mu,\nu)$, as switching would necessarily cause the first atom to not satisfy it anymore and both atoms have the same size. This construction is made precise in Algorithm (ref).

algofloat\caption \fbox{ \begin{minipage}{0.9\textwidth} \onehalfspacing For $\mu, \nu \in \mathcal{M}$ that are atomic with the same size $a>0$ and $\mu\le_{\rm st}\nu$, let $x_1 \leq \cdots \leq x_m$ and $y_1 \leq \cdots \leq y_n$ be the ordered locations of the atoms in $\mu$ and $\nu$. \begin{itemize} • Set $\mathcal{I}_m: = \{1,\ldots,n \}$. • For $i=m, m-1, \ldots, 1$, in descending order, do: \begin{itemize} • Check whether $\mathcal{J}_i := \{j\in \mathcal{I}_i: x_i\leq y_j\leq x_i+k\}$ is empty. • If $\mathcal{J}_i$ is non-empty, set $d_i := \sup \mathcal{J}_i$;\\ if $\mathcal{J}_i$ is empty, set $d_i: = \sup \mathcal{I}_i$. • Define $\mathcal{I}_{i-1}: = \mathcal{I}_i\backslash \{ d_i\}$. \end{itemize} • Return $Q^{\rm A}_{k}(\mu,\nu): = a (\sum_{i=1}^m \mathds{1}_{\{y_{d_i} - x_i > k\}}) $. \end{itemize} \end{minipage} }

The idea for constructing the optimal $\pi$ for (ref) is similar to the one for (ref). Iteratively for each $x$ in descending order, we couple it to the yet-uncoupled atom in $\nu$ at the smallest location $y$ such that $y-x > k$, and, if not possible, to the one at the smallest location such that $y \geq x$. This construction maximizes the total portion of mass in $\nu$ transported such that $y- x> k$. The rest of the transport instructions is predicated on the principle of least nuisance: If the mass portion is capable of helping toward the objective $y-x> k$, it should do so by employing the smallest location $y$. If it cannot help, it should be the least troublesome and take with itself the smallest possible values of $y$ under the constraint $y\geq x$: this leaves the greatest flexibility for subsequent to-be-coupled atoms to satisfy $y-x> k$. By going through atoms in $\mu$ from the largest to the smallest location, we make sure that unhelpful atoms are so by obligation. It also ensures that at least some uncoupled atom always permits $y\geq x$, given stochastic dominance. This construction is made precise in Algorithm (ref).

algofloat\caption \fbox{ \begin{minipage}{0.9\textwidth} \onehalfspacing For $\mu, \nu \in \mathcal{M}$ that are atomic with the same size $a>0$, let $x_1 \leq \cdots \leq x_m$ and $y_1 \leq \cdots \leq y_n$ be the ordered locations of the atoms in $\mu$ and $\nu$. \begin{itemize} • Set $\mathcal{I}_m := \{1,\ldots,n \}$. • For $i=m, m-1, \ldots, 1$, in descending order, do: \begin{itemize} • Check whether $\mathcal{J}_i := \{j\in \mathcal{I}_i: y_j > x_i+k\}$ is empty. • If $\mathcal{J}_i$ is non-empty, set $d_i := \inf \mathcal{J}_i$;\\ if $\mathcal{J}_i$ is empty, set $d_i := \inf \{j\in \mathcal{I}_i : y_j\geq x_i\}$. • Define $\mathcal{I}_{i-1} := \mathcal{I}_i\backslash \{ d_i\}$. \end{itemize} • Return $Q^{\rm B}_k(\mu,\nu) := a (\sum_{i=1}^m \mathds{1}_{\{y_{d_i} - x_i > k\}}) $. \end{itemize} \end{minipage} }

The optimality claimed for Algorithms (ref)--(ref) is formally stated in Theorem (ref).

theoremLet $\mu, \nu \in \mathcal{M}$ be atomic with the same size $a>0$, $\mu\leq_{\rm st}\nu$, and $k\geq 0$. It holds that \begin{enumerate}[label={\rm (\roman*)}, ref={\rm (\roman*)}] • $ Q^{\inf}_k(\mu,\nu)=Q^{\rm A}_k(\mu,\nu) $ with $Q^{\rm A}_k(\mu,\nu)$ specified in Algorithm (ref); • $ Q^{\sup}_k(\mu,\nu)=Q^{\rm B}_k(\mu,\nu) $ with $Q^{\rm B}_k(\mu,\nu)$ specified in Algorithm (ref). \end{enumerate} Moreover, $ Q^{\inf}_k(\mu,\nu)$ and $ Q^{\sup}_k(\mu,\nu) $ take values in the set $\{0,a,2a,\dots, m a\}$, where $m$ is in (ref).

The proof of Theorem (ref) is presented in Appendix (ref). Despite the similarity in the two problems (ref)--(ref) and the two algorithms, items (ref) and (ref) require different proof arguments.

Algorithms (ref)--(ref) not only provide the optimal values for (ref)--(ref), they also construct corresponding optimizers (not necessarily unique), which are described below. For any atomic coupling $\pi$ with size $a$, define its coupling multiset as the multiset $\Gamma$ of atom locations satisfying

equation*[equation* omitted — 67 chars of source]

Note that, for a given size $a$, a coupling multiset entirely defines a coupling. We denote by $\Gamma^{\rm A}_{k,\mu,\nu}$ the coupling multiset made of pairs $(x_{i},y_{d_i}) $ produced by Algorithm (ref). The corresponding coupling is denoted by $\gamma^{\rm A}_{k,\mu,\nu} $, that is, $$\gamma^{\rm A}_{k,\mu,\nu} = a\sum_{(x,y)\in\Gamma^{\rm A}_{k,\mu,\nu}}\delta_{(x,y)},$$ where $a$ is the size of one atom. In terms of optimal transport, $\gamma^{\rm A}_{k,\mu,\nu}$ is an unbalanced transport plan (“unbalanced" reflects the possibility of $\mu(\mathbb{R})< \nu(\mathbb{R})$), and it is optimal by Theorem (ref). Similarly, let $\Gamma^{\rm B}_{k,\mu,\nu}$ be the coupling multiset produced by Algorithm (ref), and $\gamma^{\rm B}_{k,\mu,\nu}$ be the corresponding transport plan. We provide examples of $\Gamma^{\rm A}_{k,\mu,\nu}$ and $\Gamma^{\rm B}_{k,\mu,\nu}$ for specific $\mu$ and $\nu$ below.

exampleConsider probability measures $\mu$ and $\nu$, atomic with the same size, with atoms at locations $\mathbf{x}=\{0,0.5,2,4,4.5\}$ and $\mathbf{y} = \{3.5,4,5.5,6,7\}$. Note that $m=n=5$. Let $k=3$. Using Algorithm (ref), we obtain $Q^{\rm A}_3(\mu,\nu) = 1/5$, which, by Theorem (ref), is equal to $Q_{3}^{\inf}(\mu,\nu)$. The coupling multiset $\Gamma^{\rm A}_{3,\mu,\nu}$ produced in the process is \begin{equation*} \Gamma_{3,\mu,\nu}^{\rm A} = \{(4.5,7),(4,6),(2,4),(0.5,3.5),(0,5.5)\}. \end{equation*} This coupling multiset is depicted in Figure (ref) where in green are the pairs such that $y-x > 3$ and in red are the ones such that $y-x \leq 3$. Note how ${\Gamma}_{k,\mu,\nu}^{\rm A}$ yields a smaller probability of worthwhile effect compared to comonotonic and DL couplings (see Appendix (ref)): comonotonic coupling would yield $\mathbb{P}(Y-X>3) = 3/5$ and DL-coupling, $\mathbb{P}(Y-X > 3) = 2/5$, for $X\buildrel \mathrm{d} \over \sim \mu$, $Y\buildrel \mathrm{d} \over \sim \nu$.
figure[figure omitted — 1,408 chars of source]
exampleConsider probability measures $\mu$ and $\nu$, atomic with same size, with atoms at locations $\mathbf{x}=\{0,0.5,2, 3,4,5\}$ and $\mathbf{y} = \{3,3.5,4,5,6,7\}$, and let $k=2$. From Algorithm (ref), we obtain $Q^{\rm B}_2(\mu,\nu) = 2/3$. By Theorem (ref), we thus have $Q^{\sup}_2(\mu,\nu) = 2/3$. The coupling multiset $\Gamma_{2,\mu,\nu}^{\rm B}$ produced in the process is \begin{equation*} \Gamma_{2,\mu,\nu}^{\rm B} = \{(5,5),(4,7),(3,6),(2,3),(0.5,3.5), (0,4)\}. \end{equation*} This coupling multiset is depicted in Figure (ref) where in green are the pairs such that $y-x > 2$ and in red are the ones such that $y-x \leq 2$. The coupling multiset ${\Gamma}_{k,\mu,\nu}^{\rm B}$ yields a larger probability of worthwhile effect compared to comonotonic and DL couplings: both would produce $\mathbb{P}(Y-X>2) = 1/3$, for $X\buildrel \mathrm{d} \over \sim \mu$, $Y\buildrel \mathrm{d} \over \sim \nu$.
figure[figure omitted — 1,540 chars of source]

It is clear from Figure (ref) that $\gamma^{\rm A}_{k, \mu, \nu}$ is not the unique transport plan solving (ref), as one could have rather coupled, in Example (ref), $x_4$ to $y_5$ and $x_5$ to $y_4$, still respecting the constraints, and it would have no incidence on the probability of worthwhile effect. Similarly, an alternative to $\gamma^{B}_{k,\mu,\nu}$ in the case of Example (ref) would be the transport plan obtained by rather coupling $x_1$ to $y_2$ and $x_2$ to $y_3$.

The next result highlights some stability from the algorithms, in the sense that the coupling multiset obtained may be separated in two parts, depending on whether atom pairs meet the objective or not, and re-running the algorithm on either part will not alter the coupling multiset.

propositionThe coupling multiset $\Gamma^{\rm A}_{k,\mu,\nu}$ is separable in the following sense: Let $\Gamma^{>} = \{(x,y) \in \Gamma^{\rm A}_{k,\mu,\nu}: y-x > k\}$ and $\Gamma^{\leq} = \Gamma^{\rm A}_{k,\mu,\nu}\backslash \Gamma^{>}$. Then, we have $\Gamma^{>} = \Gamma^{A}_{k,\mu\vert_{\Gamma^{>}} ,\nu\vert_{\Gamma^{>}}}$ and $\Gamma^{\leq} = \Gamma^{A}_{k,\mu\vert_{\Gamma^{\leq}} ,\nu\vert_{\Gamma^{\leq}}}$, where $\mu|_{\Gamma}$ (resp. $\nu|_{\Gamma}$) means restriction to the first (resp. second) component of multiset $\Gamma$'s elements. The coupling multiset $\Gamma^{\rm B}_{k,\mu,\nu}$ is separable similarly.

The separability of the coupling multiset will prove useful below in showing the stability of the couplings produced by the algorithms for limit cases of their applications.

Some results on the coupling solutions

This subsection gathers some results on the couplings solving (ref)--(ref); the aim is to provide further insight on the structure of such couplings and how small variations in the parameters of the problems may affect the solution. Although we previously only discussed solving (ref)--(ref) for atomic measures with same size, the results in this subsection hold for any types of measures. Theorems (ref) and (ref) below will be required in solving the problems for general discrete measures and continuous measures, in the next subsections.

First, we remark that values $Q_k^{\inf}(\mu,\nu)$ and $Q_{k}^{\sup}(\mu,\nu)$ not only bound the possible probabilities of worthwhile treatment effect for $\mu$ and $\nu$, but mark the interval of all of their possible values; we have

equation[equation omitted — 260 chars of source]

This is because the probability $\int \mathds{1}_{\{y-x>k\}}\mathrm{d} \pi$ is linear in the coupling $\pi$ and $\mathcal{H}(\mu,\nu)$ is convex. The attainability of the lower bound is claimed in Theorem (ref) just below; the upper bound may be not attained, see Appendix (ref).

theoremIn (ref), the infimum is attained, and a solution $\pi^*$ can be chosen such that $P_2(\pi^*)\leq_{\rm st}P_2(\pi)$ for all $\pi\in\mathcal{H}(\mu,\nu)$.

The selection result of Theorem (ref) is noteworthy in particular since $\nu\leq_{\rm st} \nu^{\prime}$ does not imply $Q_{k}^{\inf}(\mu,\nu) \leq Q_{k}^{\inf}(\mu,\nu^{\prime})$, as in the following example.

exampleConsider $\mu,\nu,\nu^{\prime}$, atomic probability measures with same size, with atoms respectively at locations $\mathbf{x} = \{0,5\}$, $\mathbf{y}=\{3,8\}$ and $\mathbf{y}^{\prime} = \{5,8\}$. Then, $\nu\leq_{\rm st}\nu^{\prime}$ but $Q_{1}^{\inf}(\mu,\nu) = 1$ and $Q_{1}^{\inf}(\mu,\nu^{\prime}) = 0.5$. Next, if we consider $\nu^{\dagger}$ a measure atomic with same size as $\mu$, with atom locations at $\mathbf{y}^{\dagger} = \{3,5,8\} $, we have $Q_1^{\inf}(\mu,\nu^{\dagger}) = 0.5$ and, by Theorem (ref), the coupling solution $\pi^*$ may effectively be taken so that $P_2(\pi^*) =: \nu^* \leq_{\rm st} P_2(\pi)$ for every $\pi\in\mathcal{H}(\mu,\nu^{\dagger})$. Namely, $\pi^*$ is described by the coupling multiset $\Gamma^* = \{(0,3), (5,5)\}$. Let us remark that $\nu,\nu^{\prime},\nu^* \leq \nu^{\dagger}$, and that $\nu^*\leq_{\rm st} \nu$ and $\nu^{*}\leq_{\rm st}\nu^{\prime}$.

In the case where $\mu$ and $\nu$ are probability measures and $k=0$, the upper bound in (ref) is attained and the solution is given by the comonotonic coupling, as indicated in the following proposition. Apart from the case $k=0$, however, the optimizers do not have a simple dependence structure like comonotonicity or countermonotonicity.

propositionLet $\mu,\nu \in \mathcal{M}$ be probability measures, and $\mu\leq_{\rm st} \nu$. Then, $$Q^{\sup}_0(\mu,\nu) = 1- \int_0^1 \mathds{1}_{\{F_{\mu}^{-1}(u) = F_{\nu}^{-1}(u)\}}(u)\mathrm{d}u = \mathbb{P}(Y - X >0 ),$$ where $X \buildrel \mathrm{d} \over \sim \mu$ and $Y\buildrel \mathrm{d} \over \sim \nu$ are comonotonic.

For the remainder of Section (ref) (that is, also for upcoming Sections (ref) and (ref)), we will discuss only the problem in (ref). While results for (ref) would be similar in essence, some adjustments to the mathematical settings are required, which we believe would overburden the discussion; rather, see Appendix (ref). The result in the following theorem constrains how much $Q_k^{\inf}$ varies when adding or removing mass in the coupling pool.

theoremConsider $\mu,\nu,\mu^{\prime},\nu^{\prime}\in\mathcal{M}$ such that $\mu\leq_{\rm st}\nu$, $\mu^{\prime}\leq_{\rm st}\nu^{\prime}$, $\mu\geq \mu^{\prime}$ and $\nu\leq \nu^{\prime}$. For every $k\geq 0$, it holds that \begin{equation*} 0 \leq Q^{\inf}_k(\mu,\nu) - Q^{\inf}_k(\mu^{\prime},\nu^{\prime}) \leq \mu(\mathbb{R}) - \mu^{\prime}(\mathbb{R}) + \nu^{\prime}(\mathbb{R}) - \nu(\mathbb{R}). \end{equation*}

Theorem (ref) indicates that adding or removing small portions of mass in $\mu$ or $\nu$ cannot snowball into large variations for $Q_k^{\inf}(\mu,\nu)$; the disruption is at most the amount of mass added. Let us remark that the 0 lower bound is non-trivial for $\mu \neq \mu^{\prime}$. In particular, it indicates that a portion of mass in $\mu$ never leads to major compromises in satisfying the objective $y-x\leq k$ because it has to be coupled. In Appendix (ref), we provide an example where this would be the case for $Q_k^{\sup}(\mu,\nu)$, and this predicates the need to tweak the setting accordingly for approximations regarding problem (ref).

Discrete measures

We next extend our solutions to (ref)--(ref) for $\mu$ and $\nu$ that are discrete measures but not necessarily atomic with the same size. These solutions are obtained as limit cases to those for measures that are atomic with the same size.

Let $\mu$ and $\nu$ be two discrete measures such that $|\mathrm{supp}(\nu)|<\infty$ and $\mu\leq_{\rm st}\nu$. Define $\widetilde{\mu}_r$ and $\widetilde{\nu}_r$, $r\in\mathbb{N}$, as atomized versions of $\mu$ and $\nu$, described by

align[align omitted — 342 chars of source]

and this construction ensures that the ordering constraint $\widetilde{\mu}_r\le_{\rm st} \widetilde{\nu}_r$ holds for every atomization level $r$. It also highlights the importance of allowing for $\mu(\mathbb{R})<\nu(\mathbb{R})$ in Theorem (ref) because we will apply it to $ \widetilde{\mu}_r $ and $ \widetilde{\nu}_r$, which do not have the same total mass in general. We require the support of $\nu$ to be finite, otherwise $\widetilde{\nu}_r$ would have infinite mass for any $r\in\mathbb{N}$.

propositionLet $\mu, \nu\in\mathcal{M}$ be discrete measures such that $|\mathrm{supp}(\nu)|<\infty$ and $\mu\leq_{\rm st}\nu$. For any $k\geq 0$, it holds that $Q_k^{\rm A}(\widetilde{\mu}_r, \widetilde{\nu}_r) \uparrow Q_k^{\inf}(\mu, \nu)$ as $r\to\infty$.

Executing the algorithm on atomized versions of $\mu$ and $\nu$ will not only yield an approximate value of $Q^{\inf}_k(\mu,\nu)$, it most notably produces a coupling that is distributionally close to the optimal coupling, thus giving valuable information on the dependence scheme solving (ref).

theoremLet $\mu, \nu\in\mathcal{M}$ be discrete measures such that $|\mathrm{supp}(\nu)|<\infty$ and $\mu\leq_{\rm st}\nu$. For any $k\geq 0$, there exists $\gamma \in \mathcal{H}(\mu,\nu)$ such that $\gamma_{k,\widetilde{\mu}_r,\widetilde{\nu}_r}^{\rm A} \to \gamma$ as $r\to \infty$.

For discrete $\mu$ and $\nu$, define $\gamma^{A}_{k,\mu,\nu} = \lim_{r\to\infty} \gamma^{A}_{k,\widetilde{\mu}_r,\widetilde{\nu}_r} $ as per Theorem (ref). Although Proposition (ref) and Theorem (ref) require $|\mathrm{supp}(\nu)|<\infty$, it is possible to handle cases where the measures have infinite support. For example, in cases where $\mu$ and $\nu$ have a support that is unbounded to the right, one does not have a natural starting point to construct $\gamma^{\mathrm{A}}_{k,\mu,\nu}$ and the atomization of $\nu$ would produce unbounded measures, but a possible workaround is by finding $t\in\mathbb{R}$ such that $\mu \vert_{[t,\infty)} \leq \nu^{[-k]}\vert_{[t+k,\infty)} $, where $\nu^{[-k]}(A)=\nu(A+k)$ for every Borel set $A$ on $\mathbb{R}$. Then, it is certain that all atom locations in $[t,\infty)$ are able to be coupled so that they satisfy the objective $y-x \leq k$, specifically by coupling every $x$ to $y=x+k$. We continue by constructing the coupling leftward from $t$. We illustrate this in the following example.

exampleConsider $ \mu =$ Poisson$(1)$ and $\nu =$ Poisson$(2)$. Obviously, $\mu\leq_{\rm st} \nu$; we also note that $\mu\vert_{[3,\infty)} \leq \nu^{[-1]}\vert_{[4,\infty)}$. Suppose $k=1$. The transport plan $\gamma^{\rm A}_{1,\mu,\nu}$ solving (ref) is depicted in Figure (ref), and we remark that for every $x\in\{3,4,\ldots\}$, mass at that location in $\mu$ is coupled to location $y=x+1$ in $\nu$. We obtain $Q_1^{\inf}(\mu,\nu) = 0.0626$.
figure[figure omitted — 2,102 chars of source]

Absolutely continuous measures

We now let $\mu$ and $\nu$ be absolutely continuous measures, with survival functions $S_{\mu}$ and $S_{\nu}$, whose derivatives are $-f_{\mu}$ and $-f_{\nu}$; we call $f_{\mu}$ and $f_{\nu}$ the densities of $\mu$ and $\nu$. We obtain the solution to (ref) as a limit case to solutions for the discrete case.

Consider $\overline{\mu}_{s}$ and $\overline{\nu}_{s}$, for $s\in\mathbb{N}$, discretized versions of $\mu$ and $\nu$ defined by

align[align omitted — 467 chars of source]

For a discretization level $s$, mass of interval $[x,x+{2}^{-s})$ in $\mu$, rounded down according to the smallest value of $f_{\mu}$ on that interval, is attributed to the middle point of the interval in $\overline{\mu}_s$; same for $\overline{\nu}_s$, but the amount of mass is rounded up according to the largest value of $f_{\nu}$ on the interval. Note how the intervals are designed such they divide in two going from discretization level $s$ to $s+1$, and thus no two discretization levels share points for their support. If $\mu\leq_{\rm st}\nu$, then $\overline{\mu}_{s} \leq_{st} \overline{\nu}_{s}$ holds true for every $s\in\mathbb{N}$; indeed, for every $z \in\mathrm{supp}(\overline{\mu}_{s})\cup \mathrm{supp}(\overline{\nu}_{s})$,

equation*[equation* omitted — 257 chars of source]
propositionLet $\mu, \nu \in\mathcal{M}$ have respective densities $f_{\mu}$ and $f_{\nu}$, and assume that $\mu\leq_{\rm st}\nu$ and $\mathrm{supp}(\nu)$ is compact. If $f_{\mu}$ and $f_{\nu}$ are piecewise Lipschitz continuous, then, as $s\to \infty$, for any $k\geq 0$, $Q_k^{\inf}(\overline{\mu}_{s},\overline{\nu}_{s}) \to Q_k^{\inf}(\mu,\nu)$.
figure[figure omitted — 2,314 chars of source]

Note that we require $f_{\mu}$ and $f_{\nu}$ to be piecewise Lipschitz continuous to have some regularity in approximating $\mu$ and $\nu$ by $\overline{\mu}_{s}$ and $\overline{\nu}_{s}$ as we let $s$ become large; piecewise Lipschitz continuity, combined with compactness of $\mathrm{supp}(\nu)$, also ensures that $\overline{\nu}_{s}$ has finite total mass for every discretization level $s\in\mathbb{N}$. The construction in (ref)--(ref) not only permits to approximate the value of $Q_k^{\inf}(\mu,\nu)$ with measures on which Algorithm (ref) is applicable, by Propositions (ref)--(ref), applications of Algorithm (ref) will moreover yield a consistent coupling as $s$ tends to infinity.

theoremLet $\mu,\nu\in\mathcal{M}$ with $\mu\leq_{\rm st}\nu$. Suppose that densities $f_{\mu}$ and $f_{\nu}$ are piecewise Lipschitz continuous and that $\mathrm{supp}(\nu)$ is compact. For any $k\geq 0$, there exists $\gamma \in \mathcal{H}(\mu,\nu)$ such that $\gamma_{k,\overline{\mu}_{s},\overline{\nu}_{s} }^{\rm A} \to \gamma$ weakly as $s\to \infty$.

Define $\gamma_{k,\mu,\nu}^{\rm A} = \lim_{s\to\infty} \gamma_{k,\overline{\mu}_{s}, \overline{\nu}_{s}}^{\rm A}$ as per Theorem (ref). We illustrate the discretization approach in the following example.

exampleLet $\mu =\mathrm{Uniform}(0,2)$ and $\nu = \mathrm{Uniform}(0,4)$, and thus $\mu\leq_{\rm st}\nu$. Suppose $k=1$; we examine $\gamma_{1,\mu,\nu}^{\rm A}$. By taking small discretization steps, and after atomization, we notice that, for each atom in $\mu$ at a given location $1<x\leq 2$, Algorithm (ref) couples it to the two largest atoms in $\nu$, at locations $y_1,y_2$, such that $y_i-x\leq 1$, $i\in\{1,2\}$. Then, for atoms in $\mu$ located between 0 and 1, only half of the atoms at a given location can be coupled in $\nu$ such that $x=y$, and Algorithm (ref) couples the other half to the largest uncoupled location in $\nu$. As we let the discretization level $s\to\infty$, we obtain that $\gamma^{\rm A}_{1, \mu,\nu}$ is the transport plan illustrated in Figure (ref). In green are the transport portions such that $y-x > 1$; in red are the ones such that $y-x \leq 1$. By Proposition (ref), we have $Q_1^{\rm inf}(\mu,\nu) = 0.25$.

The transport plan from Example (ref) shows that in general $\gamma^{A}_{k,\mu,\nu}$ is not Monge-type. A transport plan is Monge-type if the corresponding transport map is described by a deterministic kernel. As illustrated by the double arrows in Figure (ref), the kernel transporting from $\mu$ to $\nu$ in the example is two-atomic for $x\in [0,1]$, namely $x\mapsto (\delta_x + \delta_{3+x})/2$ on that portion.

Non-binary treatments

Hitherto, we only discussed binary treatments, that is, treatments that may be applied either fully or not, without gradation. One may inquire about the case where a participant may receive partial treatment, producing new marginal response distributions to account for. While still interested in the probability of worthwhile effect of the full treatment, we may impose the coupling to be monotone with respect to the partial-treatment responses as well. Example (ref) below shows, in a simple case, how this constrains the admissible couplings for $X$ and $Y$.

exampleSuppose that, for a given treatment, participants' responses to the treatment are $\mathbf{y} = \{2,4\}$ and responses to its absence are $\mathbf{x} = \{0,2\}$. The coupling multiset maximizing $\mathbb{P}(Y-X > 3)$ is $\{(2,2),(0,4)\}$: no effect half of the time, and an effect of 4 the other half. Now, suppose that we also observed responses to a partial treatment, $\mathbf{z}=\{1,3\}$. This additional information, paired with monotonicity of the treatment, dispels the possibility that the full treatment may have no effect, and thus leaves $\{(0,2),(2,4)\}$ as the only feasible coupling multiset for full treatment. This is illustrated in Figure (ref).
figure[figure omitted — 2,099 chars of source]

In Example (ref), we only considered one degree of partial treatment, but one could have more, leading to more layers of marginal distributions for which monotonicity has to be satisfied in the coupling. Let $Z_1,\ldots,Z_h$ be the responses to each of $h\in\mathbb{N}$ degrees of partial treatment, in increasing order of intensity, and their distributions are denoted by $\eta_1,\ldots, \eta_h$. Aligned with the assumptions regarding full treatment, we suppose that $\eta_1,\ldots,\eta_h$ are fully identified. Monotone treatment response translates to $X\leq Z_1\leq \cdots \leq Z_h \leq Y$. For this to be possible, it must hold that $\mu \leq_{\rm st} \eta_1 \leq_{\rm st} \cdots \leq_{\rm st} \eta_h \leq_{\rm st} \nu$. In the general setting, we allow $\eta_1,\ldots,\eta_h\in\mathcal{M}$ to be non-probability measures, and this stochastic ordering must also hold.

While we do not directly examine the effect of partial treatment, it gives valuable information by restricting the set of admissible measures in $\mathcal{H}(\mu,\nu)$. For any $h\in\mathbb{N}$, define

equation*[equation* omitted — 134 chars of source]

Let $\Theta_{h}$ be the set of finite measures on $\mathbb{R}^{h+2}$ supported on $\mathbb{H}_h$. Define

equation*[equation* omitted — 196 chars of source]

For $\theta\in \Theta_h$, we let $P_{XY}(\theta)$ denote the pairwise marginal of $\theta$ for the first and last dimensions. The set of admissible measures satisfying the additional partial-treatment constraints is given by

equation*[equation* omitted — 175 chars of source]

The worst-case and best-case probabilities of worthwhile effect become

align[align omitted — 475 chars of source]

Because $\mathcal{H}(\mu,\nu \mid \eta_1,\ldots,\eta_h) \subseteq \mathcal{H}(\mu,\nu)$, we have the following elementary bounds for the solutions of (ref)--(ref):

equation*[equation* omitted — 174 chars of source]

Modifying Algorithms (ref) and (ref) to consider the partial-treatment responses, we hereby present Algorithms (ref) and (ref). Our second main result is that Algorithm (ref) solves the minimization problem in (ref) and that Algorithm (ref) solves the maximization problem in (ref). We first discuss the intuition behind each.

Algorithm (ref) employs the same heuristics as Algorithm (ref) in identifying optimal atom locations in $\nu$ to which each atom in $\mu$ should be coupled. Additional steps during each of these couplings assess whether the partial-treatment constraints may be satisfied when this optimal coupling yields $y-x\leq k$. If it is not possible, the algorithm will recommend to take the atoms in $ \eta_1,\ldots,\eta_h$ and $\nu$ at the largest locations. If it is possible, that atom in $\nu$ at this location is selected along the rightmost atoms in $\eta_1,\ldots,\eta_h$ that are to the left of the upper-degree treatment. Indeed, the directionality of monotone-treatment-response constraints makes atoms in $\eta_1,\ldots,\eta_h$ at the larger locations the more restrictive. The idea is therefore that committing the larger-location atoms in $\eta_1,\ldots,\eta_h$ to the coupling of a given atom in $\mu$ is the lesser nuisance for the atoms that are yet to be coupled. The main difficulty is to prove that, by this choice, an atom in $\mu$ satisfying the objective cannot prevent more than one atom from satisfying it because of these committed atoms in $\eta_1,\ldots,\eta_h$. This is the main part of the argument in the proof of Theorem (ref) in Appendix (ref). If the algorithm recommends the largest atom in $\nu$, there is no need to verify whether the ordering constraint may be satisfied: it is guaranteed to be possible from stochastic ordering. We present an application of Algorithm (ref) below.

Algorithms (ref) and (ref) differ more than just from the heuristic difference between Algorithms (ref) and (ref); they also differ in how they verify that the partial-treatment constraint is satisfied. As stated above, the larger the location of an atom in any of $\eta_1,\ldots,\eta_h$, the more restrictive it is in coupling $\mu$ to $\nu$. Algorithm (ref) always tries to select the largest location for atoms in $\nu$, among either set of the atoms helping toward the objective or the others. Therefore, this selection mechanism flows in the same direction as the restrictiveness of the atoms in $\eta_1,\ldots,\eta_h$. Hence, we only need to check whether one atom in $\nu$, in the set of those meeting the objective, allows to satisfy the partial-treatment constraints; if it does not, the others in that set, at smaller locations, do not either. Algorithm (ref), by contrast, always aims to select atoms at smallest locations, whether it is among the set of those meeting the objective $y-x > k$ or the others, meeting $y\geq x$. Hence, its selection process is opposite to the partial-treatment restrictiveness direction. As a result, if the favored atom in the set meeting the objective $y-x > k$ cannot satisfy the partial treatment constraint, the algorithm must check every other atom in that set, as they might satisfy it, before opting for the other set, for which the same process must be repeated. Algorithm (ref) has therefore a higher degree of complexity than Algorithm (ref), from this opposite directionality.

algofloat[t] \caption \fbox{ \begin{minipage}{0.9\textwidth} \onehalfspacing For measures $\mu$, $\nu$ and $\eta_1, \ldots, \eta_h$, atomic with the same size $a>0$ and $\mu \leq_{\rm st} \eta_1 \leq_{\rm st} \dots \leq_{\rm st} \eta_h \leq_{\rm st} \nu$, let $x_1 \leq \cdots \leq x_m$ and $y_1 \leq \cdots \leq y_n$ be the ordered locations of atoms in $\mu$ and $\nu$ and, for each $\ell\in\{1, \ldots,h\}$, let $z_{\ell,1}\leq \cdots \leq z_{\ell, s_{\ell}}$ be the ordered locations of atoms in $\eta_{\ell}$. \begin{itemize} • Set $\mathcal{I}_m := \{1,\ldots,n \}$ and, for each $\ell \in \{1,\ldots,h\}$, set $\mathcal{S}_{\ell, m} := \{1,\ldots,s_{\ell} \}$. • For $i=m, m-1, \ldots, 1$, in descending order, do: \begin{itemize} • Check whether $\mathcal{J}_i := \{j\in \mathcal{I}_i: x_i \leq y_j \leq x_i+k\}$ is empty. If it is non-empty, do: \begin{itemize} • Set $d^*_i := \sup \mathcal{J}_i$. • Set $c_{i,h} := \sup\{r\in\mathcal{S}_{h,i} : z_{h,r} \leq y_{d^*_i}\}$ (skip if the set is empty). • For $\ell = h-1$, $h-2$, $\ldots$, 1, in descending order, do (skip if $h=1$): \begin{itemize} • Set $c_{i,\ell} := \sup\{r\in \mathcal{S}_{\ell, i} : z_{\ell, r} \leq z_{\ell+1, c_{i,\ell+1}}\}$ (skip if the set is empty). \end{itemize} • If $c_{i,1}$ was defined and, in which case, $z_{1,c_{i,1}}\geq x_i$, do: \begin{itemize} • Set $d_i := d_i^*$. • Define $\mathcal{S}_{\ell, i-1} := \mathcal{S}_{\ell, i}\backslash \{c_{i,\ell} \}$ for each $\ell\in\{1,\ldots,h\}$. \end{itemize} \end{itemize} • If $d_i$ is still undefined at this step, do: \begin{itemize} • Set $d_i := \sup \mathcal{I}_i$. • Define $\mathcal{S}_{\ell, i-1} := \mathcal{S}_{\ell, i}\backslash \{\sup \mathcal{S}_{\ell, i}\}$ for each $\ell\in\{1,\ldots,h\}$. \end{itemize} • Define $\mathcal{I}_{i-1} := \mathcal{I}_i\backslash\{d_i\}$. \end{itemize} • Return $Q^{\mathrm{A}^{\eta}}_k(\mu,\nu | \eta_1,\ldots,\eta_h) := a(\sum_{i=1}^m \mathds{1}_{\{y_{d_i} - x_i > k\}})$. \end{itemize} \end{minipage} }
algofloat[t] \caption \fbox{ \begin{minipage}{0.9\textwidth} \onehalfspacing For measures $\mu$, $\nu$ and $\eta_1, \ldots, \eta_h$, atomic with the same size $a>0$ and $\mu \leq_{\rm st} \eta_1 \leq_{\rm st} \dots \leq_{\rm st} \eta_h \leq_{\rm st} \nu$, let $x_1 \leq \cdots \leq x_m$ and $y_1 \leq \cdots \leq y_n$ be the ordered locations of atoms in $\mu$ and $\nu$ and, for each $\ell\in\{1, \ldots,h\}$, let $z_{\ell,1}\leq \cdots \leq z_{\ell, s_{\ell}}$ be the ordered locations of atoms in $\eta_{\ell}$. \begin{itemize} • Set $\mathcal{I}_m := \{1,\ldots,n \}$ and, for each $\ell \in \{1,\ldots,h\}$, set $\mathcal{S}_{\ell, m} := \{1,\ldots,s_{\ell} \}$. • For $i=m, m-1, \ldots, 1$, in descending order, do: \begin{itemize} • For each element $j$ of $\mathcal{J}_i := \{j\in \mathcal{I}_i: y_j \geq x_i\}$, do: \begin{itemize} • Set $c_{i,j,h} := \sup\{r\in\mathcal{S}_{h,i} : z_{h,r} \leq y_j\}$ (skip if the set is empty). • For $\ell = h-1$, $h-2$, $\ldots$, 1, in descending order, do (skip if $h=1$): \begin{itemize} • Set $c_{i,j,\ell} := \sup\{r\in \mathcal{S}_{\ell, i} : z_{\ell, r} \leq z_{\ell+1, c_{i,j,\ell+1}}\}$ (skip if the set is empty). \end{itemize} • If $c_{i,j,1}$ was defined and, in which case, $z_{1,c_{i,j,1}}\geq x_i$, do: \begin{itemize} • If $y_j> x_i + k$, include $j$ in a set $\mathcal{J}_i^*$. • Otherwise, include $j$ in a set $\mathcal{J}_i^{\prime}$. \end{itemize} \end{itemize} • Check whether $\mathcal{J}_i^{*}$ is empty. • If $\mathcal{J}_i^*$ is non-empty, set $d_i := \inf \mathcal{J}_i^*$.\\ If $\mathcal{J}_i^*$ is empty, set $d_i := \inf \mathcal{J}_i^{\prime}$. • Define $\mathcal{S}_{\ell, i-1} := \mathcal{S}_{\ell, i}\backslash \{c_{i,d_i, \ell}\}$ for each $\ell \in\{1,\ldots, h\}$. • Define $\mathcal{I}_{i-1} := \mathcal{I}_i\backslash\{d_i\}$. \end{itemize} • Return $Q^{\mathrm{B}^{\eta}}_k(\mu,\nu | \eta_1,\ldots,\eta_h) := a(\sum_{i=1}^m \mathds{1}_{\{y_{d_i} - x_i > k\}})$. \end{itemize} \end{minipage} }
theoremLet $\mu,\eta_1,\ldots,\eta_h,\nu \in \mathcal{M}$ be atomic with the same size $a>0$, $\mu \leq_{\rm st} \eta_1 \leq_{\rm st} \dots \leq_{\rm st} \eta_h \leq_{\rm st} \nu$, and $k\geq 0$. It holds that \begin{enumerate}[label={\rm (\roman*)}, ref={\rm (\roman*)}] • $ Q^{\inf}_k(\mu,\nu\mid \eta_1,\ldots,\eta_h)=Q^{\mathrm{A}^{\eta}}_k(\mu,\nu\mid \eta_1,\ldots,\eta_h) $, the latter specified in Algorithm (ref); • $ Q^{\sup}_k(\mu,\nu\mid \eta_1,\ldots,\eta_h)=Q^{\mathrm{B}^{\eta}}_k(\mu,\nu\mid \eta_1,\ldots,\eta_h) $, the latter specified in Algorithm (ref). \end{enumerate} Moreover, $Q^{\inf}_k(\mu,\nu \mid \eta_1,\ldots,\eta_h)$ and $Q^{\sup}_k(\mu,\nu \mid \eta_1,\ldots,\eta_h)$ take values in the set $\{0,a,2a,\dots, m a\}$, where $m$ is in (ref).
figure[figure omitted — 2,116 chars of source]
exampleConsider probability measures $\mu$ for no-treatment responses, $\eta_1$ and $\eta_2$ for two degrees of partial-treatment responses, and $\nu$ for full-treatment responses. Each consists of equiprobable atoms at respective locations $\mathbf{x} = \{0,1,4,6 \}$, $\mathbf{z}_1 = \{1,2,4,7\}$, $\mathbf{z}_2 = \{2,4,5,7\}$ and $\mathbf{y} = \{2,4,6,9\}$. Let $k=2$. Algorithm (ref) starts the coupling process with the atom at $x_4=6$. According to Algorithm (ref), the preferred atom in $\nu$ should be $y_3 = 6$; however, Algorithm (ref) needs to verify whether it can satisfy the partial-treatment constraint. Looking down, selecting atoms in $\eta_1$ and $\eta_2$ at largest admissible locations, we obtain $z_{1,c_{4,1}} = 4 < 6$, so it is not possible to satisfy the constraint. Therefore, instead, Algorithm (ref) couples atom at $x_4 = 6$ to atom at $y_4 = 9$, the largest location, and commits atoms at $z_{1,4} = 7$ and $z_{2,4} = 7$ to this couple, effectively creating tuple (6,7,7,9) for the joint distribution $\pi\in\mathcal{H}(\mu,\nu\mid\eta_1,\eta_2)$. Then it is the turn of the atom at $x_3 = 4$ to be coupled. The favored atom in $\nu$ according to Algorithm (ref) would be $y_3 = 6$, and, after verification, it satisfies the partial-treatment constraint, by committing atoms at $z_{1,3} = 4$ and $z_{2,3} = 5$, so the algorithm selects this atom. The coupling multiset thus constructed is depicted in Figure (ref). Algorithm (ref) yields $Q_2^{\mathrm{A}^{\eta}}(\mu,\nu\mid\eta_1,\eta_2) = 0.5$; we remark that it is larger than $Q_2^{\rm A}(\mu,\nu)$, without the partial-treatment constraint. By Theorem (ref), we have $Q^{\inf}_2(\mu,\nu\mid\eta_1,\eta_2) = 0.5$.
figure[figure omitted — 1,635 chars of source]
exampleConsider probability measures $\mu$ for no-treatment responses, $\eta$ for partial-treatment responses, and $\nu$ for full-treatment responses. Each comprises four atoms of the same size at respective locations $\mathbf{x}=\{0,2,4,6\}$, $\mathbf{z}=\{2,4,5,7\}$, and $\mathbf{y}=\{3,5,6,8\}$. Assume $k=2$. Algorithm (ref) returns $Q_2^{\mathrm{B}^{\eta}}(\mu,\nu\mid\eta) = 0.5$; therefore, $Q_2^{\sup}(\mu,\nu\mid\eta) = 0.5$ by Theorem (ref). In the beginning of the coupling process, when coupling the atom at $x_4=6$, Algorithm (ref) selects atom at $y_4=8$ instead of the one at $y_3=6$, as would be recommended by Algorithm (ref). The verification indeed yields $z_{c_{4,3}} = 4 < 6$, indicating that Algorithm (ref)'s choice would not allow to satisfy the partial-treatment constraints. The complete coupling multiset produced by Algorithm (ref) is depicted in Figure (ref). Without the partial treatment, choosing to couple atoms at $x_4 = 6$ and $y_3 = 6$ would have allowed to couple atoms at $x_3 = 4$ and $y_4 = 8$, and atoms at $x_2 = 2$ and $y_2 = 5$, letting one more couple satisfy the objective $y-x> 2$; effectively, $Q_2^{\rm B}(\mu,\nu) = 0.75$.

By using the same atomization and discretization methods, we can also obtain results similar to those in Sections (ref)--(ref) under partial-treatment constraints.

The approach from Algorithms (ref) and (ref) does not extend to continuous treatments, where we would have a continuum of measures $\{\eta_u: u\in (0,1)\}$ to represent the degrees of partial treatment. Let us remark, however, that continuity of treatments conflicts with the assumption of fully identified marginal distributions (and no further structural assumptions) in practical applications: for continuous treatments, almost surely no two subjects will be administered the same level of treatment and thus only one point may be observed for each marginal distribution.

Conclusion

The probability of worthwhile effect for a treatment requires, to be computable, identification of the joint distribution of treatment responses, unlike the average treatment effect, which is commonly employed but ill-suited to some contexts. We examined the range of possible probabilities of worthwhile effect of a monotone-response treatment with the assumption that only marginal distributions are fully identifiable. With no further identification assumptions, Algorithms (ref) and (ref) yield the extremal probabilities of worthwhile effect for atomic distributions. Results for discrete and continuous distributions follow by convergence after atomization and discretization. Partial identification may come from responses to intermediate degrees of treatment, in which case Algorithms (ref) and (ref) produce the extremal probabilities of worthwhile effect. While we mainly grounded this paper in the statistical considerations of treatment effects, these results also transpose to other frameworks; they notably solve extant problems in optimal transport and risk aggregation under model uncertainty.

Acknowledgments

Ruodu Wang is supported by the Natural Sciences and Engineering Research Council of Canada (RGPIN-2018-03823, RGPAS-2018-522590) and Canada Research Chairs (CRC-2022-00141).

thebibliography{10} \bibitem[\citeauthoryear{Amin et al.}{Amin et al.}{2025}]{ABFFFKKS25} Amin, V., Behrman, J. R., Fletcher, J. M., Flores, C. A., Flores-Lagunes, A., Kohler, I., Kohler, H.-P. and Stites, S. D. (2025). Causal effects of schooling on memory at older ages in six low-and middle-income countries: Nonparametric evidence with harmonized datasets. Journals of Gerontology, Series B: Psychological Sciences and Social Sciences, 80(6), gbaf057. \bibitem[\citeauthoryear{Arnold et al.}{Arnold et al.}{2020}]{AMZ20} Arnold, S., Molchanov, I. and Ziegel, J. F. (2020). Bivariate distributions with ordered marginals. Journal of Multivariate Analysis, 177, 104585. \bibitem[\citeauthoryear{Balke and Pearl}{Balke and Pearl}{1994}]{BP94} Balke, A. and Pearl, J. (1994). Counterfactual probabilities: computational methods, bounds and applications. Uncertainty in Artificial Intelligence: Proceedings of the 10th Conference, 46--54. \bibitem[\citeauthoryear{Bernard et al.}{Bernard et al.}{2014}]{BJW14} Bernard, C., Jiang, X. and Wang, R. (2014). Risk aggregation with dependence uncertainty. Insurance: Mathematics and Economics, \textbf{54}, 93--108. \bibitem[\citeauthoryear{Blanchet et al.}{Blanchet et al.}{2025}]{BLLW24} Blanchet, J., Lam, H., Liu, Y. and Wang, R. (2025). Convolution bounds on quantile aggregation. \textit{Operations Research}, \textbf{73}(5), 2761--2781. \bibitem[\citeauthoryear{Bott et al.}{Bott et al.}{2020}]{BCST20} Bott, K. M., Cappelen, A. W., S\o rensen, E. \O. and Tungodden, B. (2020). You’ve got mail: A randomized field experiment on tax evasion. \textit{Management Science}, \textbf{66}(7), 2801--2819. \bibitem[\citeauthoryear{Chen et al.}{Chen et al.}{2022}]{CLW22} Chen, Y., Lin, L. and Wang, R. (2022). Risk aggregation under dependence uncertainty and an order constraint. \textit{Insurance: Mathematics and Economics}, \textbf{102}, 169--187. \bibitem[\citeauthoryear{Cho}{Cho}{2022}]{C22} Cho, S. J. (2022). The effect of aging out of the Women, Infants, and Children (WIC) program on food insecurity. \textit{Health Economics}, \textbf{31}(4), 664--685. \bibitem[\citeauthoryear{C\^ot\'e and Wang}{C\^ot\'e and Wang}{2026}]{CW26} C\^ot\'e, B. and Wang, R. (2026). On convex order and supermodular order without finite mean. \textit{Insurance: Mathematics and Economics}, \textbf{128}, 103234. \bibitem[\citeauthoryear{Dawid}{Dawid}{2000}]{D00} Dawid, A. P. (2000). Causal inference without counterfactuals. \textit{Journal of the American Statistical Association}, \textbf{95}(450), 407--424. \bibitem[\citeauthoryear{Dawid}{Dawid}{2015}]{D15} Dawid, A. P. (2015). Statistical causality from a decision theoretic perspective. \textit{Annual Review of Statistics and its Application}, \textbf{2}, 272--303. \bibitem[\citeauthoryear{Demidenko}{Demidenko}{2016}]{D16} Demidenko, E. (2016). The p-value you can’t buy. \textit{The American Statistician}, \textbf{70}(1), 33--38. \bibitem[\citeauthoryear{Embrechts and Puccetti}{Embrechts and Puccetti}{2010}]{EP10} Embrechts, P. and Puccetti, G. (2010). Risk aggregation. \textit{In Copula Theory and Its Applications: Proceedings of the Workshop Held in Warsaw, 25-26 September 2009}, 111--126. \bibitem[\citeauthoryear{Embrechts et al.}{Embrechts et al.}{2013}]{EPR13} Embrechts, P., Puccetti, G. and R\"uschendorf, L. (2013). Model uncertainty and VaR aggregation. \textit{Journal of Banking and Finance}, \textbf{37}(8), 2750--2764. \bibitem[\citeauthoryear{Embrechts et al.}{Embrechts et al.}{2015}]{EWW15} Embrechts, P., Wang, B. and Wang, R. (2015). Aggregation-robustness and model uncertainty of regulatory risk measures. \textit{Finance and Stochastics}, \textbf{19}, 763--790. \bibitem[\citeauthoryear{Friesecke}{Friesecke}{2024}]{F24} Friesecke, G. (2024). \textit{Optimal Transport: A Comprehensive Introduction to Modeling, Analysis, Simulation, Applications.} Society for Industrial and Applied Mathematics, Philadelphia. \bibitem[\citeauthoryear{Froehlich}{Froehlich}{1999}]{F99} Froehlich, G. W. (1999). What is the chance that this study is clinically significant? A proposal for Q values. \textit{Effective Clinical Practice}, \textbf{2}, 234--239. \bibitem[\citeauthoryear{Gerfin and Schellhorn}{Gerfin and Schellhorn}{2006}]{GS06} Gerfin, M. and Schellhorn, M. (2006). Nonparametric bounds on the effect of deductibles in health care insurance on doctor visits---Swiss evidence. \textit{Health Economics}, \textbf{15}(9), 1011--1020. \bibitem[\citeauthoryear{Greenland et al.}{Greenland et al.}{2020}]{GFBSFGR20} Greenland, S., Fay, M. P., Brittain, E. H., Shih, J. H., Follmann, D. A., Gabriel, E. E. and Robins, J. M. (2020). On causal inferences for personalized medicine: How hidden causal assumptions led to erroneous causal claims about the D-value. \textit{The American Statistician}, \textbf{74}(3), 243--248. \bibitem[\citeauthoryear{Gundersen and Kreider}{Gundersen and Kreider}{2009}]{GK09} Gundersen, C. and Kreider, B. (2009). Bounding the effects of food insecurity on children’s health outcomes. \textit{Journal of Health Economics}, \textbf{28}(5), 971--983. \bibitem[\citeauthoryear{Gundersen et al.}{Gundersen et al.}{2012}]{GKP12} Gundersen, C., Kreider, B. and Pepper, J. (2012). The impact of the National School Lunch Program on child health: a nonparametric bounds analysis. \textit{Journal of Econometrics}, \textbf{166}(1), 79--91. \bibitem[\citeauthoryear{Gundersen et al.}{Gundersen et al.}{2017}]{GKP17} Gundersen, C., Kreider, B. and Pepper, J. V. (2017). Partial identification methods for evaluating food assistance programs: a case study of the causal impact of SNAP on food insecurity. \textit{American Journal of Agricultural Economics}, \textbf{99}(4), 875--893. \bibitem[\citeauthoryear{Hand}{Hand}{1992}]{H92} Hand, D. J. (1992). On comparing two treatments. \textit{The American Statistician}, \textbf{46}(3), 190--192. \bibitem[\citeauthoryear{Hof}{Hof}{2014}]{H14} Hof, S. (2014). Does private tutoring work? The effectiveness of private tutoring: a nonparametric bounds analysis. \textit{Education Economics}, \textbf{22}(4), 347--366. \bibitem[\citeauthoryear{Huesmann and Sturm}{Huesmann and Sturm}{2013}]{HS13} Huesmann, M. and Sturm, K.-T. (2013). Optimal transport from Lebesgue to Poisson. \textit{Annals of Probability}, \textbf{41}(4), 2426--2478. \bibitem[\citeauthoryear{Jaffe and Raban}{Jaffe and Raban}{2025}]{JR25} Jaffe, A. Q., and Raban, D. (2025). Coupling theory, optimal transport, and Strassen's theorem beyond regular orders. \textit{arXiv:2509.21616}. \bibitem[\citeauthoryear{Ji et al.}{Ji et al.}{2023}]{JLS23} Ji, W., Lei, L. and Spector, A. (2023). Model-agnostic covariate-assisted inference on partially identified causal effects. \textit{arXiv:2310.08115}. \bibitem[\citeauthoryear{Kim}{Kim}{2014}]{K14} Kim, J. H. (2014). Identifying the distribution of treatment effects under support restrictions. \textit{arXiv:1410.5885}. \bibitem[\citeauthoryear{Kreider et al.}{Kreider et al.}{2016}]{KPR16} Kreider, B., Pepper, J. V. and Roy, M. (2016). Identifying the effects of WIC on food insecurity among infants and children. \textit{Southern Economic Journal}, \textbf{82}(4), 1106--1122. \bibitem[\citeauthoryear{Liu and Wang}{Liu and Wang}{2021}]{LW21} Liu, F. and Wang, R. (2021). A theory for measures of tail risk. \textit{Mathematics of Operations Research}, \textbf{46}(3), 1109--1128. \bibitem[\citeauthoryear{Mann and Whitney}{Mann and Whitney}{1947}]{MW47} Mann, H. B., and Whitney, D. R. (1947). On a test of whether one of two random variables is stochastically larger than the other. \textit{Annals of Mathematical Statistics}, \textbf{18}(1), 50--60. \bibitem[\citeauthoryear{Manski}{Manski}{1994}]{M94} Manski, C. F. (1994). The selection problem. \textit{Advances in Econometrics}, \textbf{1}, 147--170. Cambridge University Press, Cambridge. \bibitem[\citeauthoryear{Manski}{Manski}{1995}]{M95} Manski, C. F. (1995). \textit{Identification Problems in the Social Sciences}. Harvard University Press, Cambridge. \bibitem[\citeauthoryear{Manski}{Manski}{1997}]{M97} Manski, C. F. (1997). Monotone treatment response. \textit{Econometrica}, \textbf{65}(6), 1311--1334. \bibitem[\citeauthoryear{Manski}{Manski}{2007}]{M07} Manski, C. F. (2007). \textit{Identification for Prediction and Decision}. Harvard University Press, Cambridge. \bibitem[\citeauthoryear{Manski and Pepper}{Manski and Pepper}{2000}]{MP00} Manski, C. F. and Pepper, J. V. (2000). Monotone instrumental variables, with an application to the returns to schooling. \textit{Econometrica}, \textbf{68}(4), 997--1010. \bibitem[\citeauthoryear{Manski and Pepper}{Manski and Pepper}{2009}]{MP09} Manski, C. F. and Pepper, J. V. (2009). More on monotone instrumental variables. \textit{Econometrics Journal}, \textbf{12}(1), S200--S216. \bibitem[\citeauthoryear{Mao and Wang}{Mao and Wang}{2015}]{MW15} Mao, T. and Wang, R. (2015). On aggregation sets and lower-convex sets. \textit{Journal of Multivariate Analysis}, \textbf{138}, 170--181. \bibitem[\citeauthoryear{M\"uller and Stoyan}{M\"uller and Stoyan}{2002}]{MS02} M\"uller, A. and Stoyan, D. (2002). \textit{Comparison Methods for Stochastic Models and Risks.} Wiley, Hoboken. \bibitem[\citeauthoryear{Nutz and Wang}{Nutz and Wang}{2022}]{NW22} Nutz, M. and Wang, R. (2022). The directional optimal transport. \textit{Annals of Applied Probability}, \textbf{32}(2), 1400--1420. \bibitem[\citeauthoryear{Rachev and R{\"u}schendorf}{Rachev and R{\"u}schendorf}{2006}]{RR06} Rachev, S. T. and R{\"u}schendorf, L. (2006). \textit{Mass Transportation Problems, Volume 1: Theory.} Springer, New York. \bibitem[\citeauthoryear{R\"uschendorf}{R\"uschendorf}{1982}]{R82} R\"uschendorf, L. (1982). Random variables with maximum sums. \textit{Advances in Applied Probability}, \textbf{14}(3), 623--632. \bibitem[\citeauthoryear{R{\"u}schendorf}{R{\"u}schendorf}{2013}]{R13} R{\"u}schendorf, L. (2013). {\em Mathematical Risk Analysis. Dependence, Risk Bounds, Optimal Allocations and Portfolios}. Springer, Heidelberg. \bibitem[\citeauthoryear{Senn}{Senn}{2009}]{S09} Senn, S. (2009). Three things that every medical writer should know about statistics. \textit{Journal of the European Medical Writers Association}, \textbf{18}(3), 159--162. \bibitem[\citeauthoryear{Shaked and Shanthikumar}{Shaked and Shanthikumar}{2007}]{SS07} Shaked, M. and Shanthikumar, J. G. (2007). \textit{Stochastic Orders.} Springer, New York. \bibitem[\citeauthoryear{Villani}{Villani}{2009}]{V09} Villani, C. (2009). \textit{Optimal Transport: Old and New}. Springer, Heidelberg. \bibitem[\citeauthoryear{Wang and Wu}{Wang and Wu}{2025}]{WW25} Wang, R. and Wu, Q. (2025). The reference interval in higher-order stochastic dominance. \textit{Economic Theory Bulletin}, \textbf{13}, 263--277. \bibitem[\citeauthoryear{Wang and Zitikis}{Wang and Zitikis}{2021}]{WZ21} Wang, R. and Zitikis, R. (2021). An axiomatic foundation for the Expected Shortfall. \textit{Management Science}, \textbf{67}(3), 1413--1429. \bibitem[\citeauthoryear{Wilcoxon}{Wilcoxon}{1945}]{W45} Wilcoxon, F. (1945). Individual comparisons by ranking methods. \textit{Biometrics Bulletin}, \textbf{1}(6), 80-83. \bibitem[\citeauthoryear{Winship and Morgan}{Winship and Morgan}{1999}]{WM99} Winship, C. and Morgan, S. L. (1999). The estimation of causal effects from observational data. \textit{Annual Review of Sociology}, \textbf{25}(1), 659--706. \bibitem[\citeauthoryear{Yarnell and Goligher}{Yarnell and Goligher}{2024}]{YG24} Yarnell, C. J. and Goligher, E. C. (2024). Interpreting posterior probabilities in Bayesian analyses of clinical trials. \textit{Lancet Respiratory Medicine}, \textbf{12}(3), 188--190.