Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
93,909 characters · 12 sections · 43 citation commands
Probability of worthwhile effect of monotone-response treatments
In a simple experiment, we study the effect of a treatment. The responses of subjects to the treatment, $Y$, and to the absence of treatment, $X$, are monitored. Responses $X$ and $Y$ are random variables; their distributions, respectively $\mu$ and $\nu$, may either be interpreted as that of the (deterministic) treatment response within a population or of the aleatoric response of any single subject. The treatment effect is $Y-X$.
Most considerations on the treatment effect are answered by elementary statistics if we are able to observe both responses for every subject. However, this is often not possible for experimental studies outside of controlled environments or conducted over a long time frame, or sometimes due to ethical considerations. Consider, for instance, a study on the effects of malnutrition on children's health: a researcher cannot require subjects to be malnourished for the length of the study. The researcher must rather rely on observational data, meaning that data on responses to treatment and non-treatment are collected separately. This leads to two major impediments for statistical purposes:
Let us provide a simple example to illustrate this second impediment.
Experimental designs based on observational data are prominent in many fields, such as sociology, statistics and econometrics; and their study led to the counterfactual conception of causal effects. One may consult BP94, WM99, D00, S09, and D15 for an introduction and some (philosophical, statistical) considerations on counterfactual causality.
As underscored in all of these works, studies involving counterfactual causality will almost always examine the average treatment effect, $\mathbb{E}[Y-X]$. This has the marked advantage of dispelling impediment (S2) because $\mathbb{E}[Y-X]$ is unaffected by the dependence between $Y$ and $X$. Most seminal statistical works on counterfactual causality thus exclusively focused on mitigating (S1); see notably the series of works M94,M95, M97, M07, MP00, MP09.
Average treatment effect, however, may be ill-suited to some situations. In Example (ref), both couplings yield the same average effect, but may lead to different decisions. If, for example, treatments are antidepressants and responses are their side effect on blood pressure, the first coupling, with a constant 1 effect, would probably lead to the drug being prescribed by a physician, as they may judge that the side effect is tolerable. The second coupling, however, would probably incite physicians to avoid prescribing the drug, as they deem too risky to potentially see a dramatic increase of blood pressure. See F99 and YG24 for further discussions on the limitations of average treatment effect in medical trials. Let us present other examples illustrating this rhetoric.
We remark that the analytics under study in these examples do not depend on the actual responses $X$ or $Y$ but solely on the treatment effect $Y-X$. However, the average treatment effect is unable to answer the threshold-related considerations.
For such considerations, one may rather consider the probability of a worthwhile effect; see, e.g., F99. Let $k\geq 0$ be the level above which a treatment effect is considered clinically relevant. The probability of worthwhile effect for a given $k$ is the value $\mathbb{P}(Y-X > k)$. F99 discussed probabilities of worthwhile effect as an alternative to p-values in the context of medical treatments. An issue for p-values is that statistical significance of the treatment effect is unrelated to whether this effect is clinically worthwhile. F99 argued that a medical researcher should not heuristically determine what constitutes a worthwhile effect based on the collected data; probabilities of worthwhile effect effectively require the researcher to specify beforehand the threshold for worthwhile effect $k$. For instance, in our earlier example, the physician would establish which threshold $k$ of increase in blood pressure is dangerous and then inquire about the probability of worthwhile effect for $k$. In Examples (ref) and (ref), thresholds for worthwhile effects are specified as $k=10$ and $k=0$ respectively.
When examining probabilities of worthwhile effect instead of the average treatment effect, impediment (S2) comes back: the value $\mathbb{P}(Y-X > k)$ is impacted by the dependence relation between $X$ and $Y$. Since there is already a long string of literature on how to mitigate (S1), as we specified above, we will focus on (S2) and assume that the treatment has fully identified marginals, to remove (S1).
Because of (S2), probabilities of worthwhile effect cannot be computed unless we specify the dependence relation between $X$ and $Y$. One common assumption is to suppose that $X$ and $Y$ are comonotonic.\footnote{Two random variables $Z,W$ are comonotonic if $(Z(\omega) - Z(\omega^{\prime}))(W(\omega)- W(\omega^{\prime}))\geq 0$ for all $\omega, \omega^{\prime}\in\Omega$.} If $\nu$ is a translation of $\mu$, then this is called the Treatment-Unit Additivity assumption D15, which effectively means that $Y-X$ is degenerate. Another common assumption is to suppose that $X$ and $Y$ are independent D00. Letting $k=0$, the probability of worthwhile effect in that case corresponds, up to a small adjustment for ties, to the Mann-Whitney parameter, used in the Wilcoxon-Mann-Whitney test for causal effects W45,MW47 and later introduced by D16 under the name D-value. Not only are independence and comonotonicity very different dependence schemes, they also mean very strong assumptions on the treatment's effect. A researcher may well not want to commit to either. In particular, the independence assumption is heavily criticized by GFBSFGR20, and Hand's paradox H92 is an example of how misleading it can be. As GFBSFGR20 note, to rely on such assumptions for causal inferences marks an obliviousness to impediment (S2), perhaps forgotten because the average treatment effect bypasses it. Consistent with such criticism, our approach will rather consist of deriving worst- and best-case probabilities of worthwhile effect under the full uncertainty stemming from (S2).
In this paper, we will study probabilities of worthwhile effect of treatments with the additional assumption of a monotone response. Monotone treatment response, first studied by M97, stipulates that responses weakly increase with the strength of treatment. For a binary treatment, this means $Y\geq X$.\footnote{Such relations for random variables are to be interpreted as holding almost surely.} Put equivalently: the treatment effect $Y-X$ is nonnegative almost surely, meaning that the treatment cannot decrease the subject's response.
This assumption may seem strong, but is far from unreasonable in many contexts, for instance in a study of the effects of private tutoring on academic success H14. Other examples include the effects of schooling on elders' cognition ABFFFKKS25, of a smaller health-insurance deductible on the number of doctor visits GS06, of moral suasion on tax evasion BCST20, of food assistance programs on food security KPR16, GKP17, C22, and of food security on children's health GK09, GKP12. In Example (ref), both couplings satisfy the assumption of monotone treatment response $ y_i \geq x_i$ for every element $i$. Thus, it is clear that, while limiting the number of possible couplings, the monotone-treatment-response assumption alone in general does not suffice to fully identify the joint distribution of responses.
We examine worst- and best-case probabilities of worthwhile effect respectively defined as
where $X\buildrel \mathrm{d} \over \sim \mu$ means that the distribution of $X$ is $\mu$.
We do not make further assumptions than full identification of marginals and monotone treatment response; in particular, we do not impose any shape or specific features for the marginal distributions, as long as they allow for monotone treatment response to be possible. The set of possible joint distributions given this constraint is studied in AMZ20.
In the remainder of this section, we examine problems (ref)--(ref) from different perspectives: of risk aggregation and of optimal transport; we also introduce some notation. In Section (ref), we solve the problems for binary treatments. In Section (ref), we solve the problems for non-binary treatments, where intermediate degrees of treatment constitute a source of partial identification of the joint distribution for the full treatment, from the monotone-treatment-response assumption. Section (ref) concludes.
In optimal transport theory, the primal Kantorovich problem consists of minimizing an expected cost over all joint distributions with given marginals. See RR06, V09 and F24 for an introduction to optimal transport. We rewrite problems (ref)--(ref) in terms of optimal transport:
Note how the second terms of the cost functions $c$ and $c^*$ enforce the ordering constraint from the monotone-treatment-response assumption. For atomic distributions (see Section (ref)), one may replace $\infty$ by a finite number that is large enough. Optimal transport with this ordering constraint is studied by NW22 under the name directional optimal transport. They solve problems similar to (ref)--(ref), but with the first term of $c$ or $c^*$ being submodular or supermodular.\footnote{A function $(x,y)\mapsto f(x,y)$ is submodular if $f(x_1,y_1) + f(x_2,y_2) \leq f(x_1\wedge x_2, y_1 \wedge y_2) + f(x_1\vee x_2, y_1 \vee y_2)$ for all $(x_1,y_1),(x_2,y_2)$ in its domain. It is supermodular if the reverse inequality holds for all $(x_1,y_1),(x_2,y_2)$ in its domain.} Here, the functions $(x,y)\mapsto \mathds{1}_{\{y-x > k\}}$ and $(x,y)\mapsto \mathds{1}_{\{y-x \leq k\}}$ are neither submodular nor supermodular. Although it will not be our approach to solving them, strong duality holds for problems (ref)--(ref); see Appendix (ref).
In the risk-management literature, the framework obtained by dispelling (S1) but retaining (S2) is called dependence uncertainty: we suppose that the marginal distributions of $(X,Y)$ are known, but the dependence relation between its components is unknown. See EP10 for considerations and motivations for such a framework in the context of risk management. CLW22 proposed and solved the following problems:
Note that the only difference between their problems and ours, beyond the weak inequality in (ref), is that $X+Y$ in (ref)--(ref) takes the place of $Y-X$ in (ref)--(ref). This difference has a significant impact on the optimal couplings, making the problems (ref)--(ref) completely different from (ref)--(ref); in particular, the results of CLW22 cannot be applied to our problems. We discuss their connections in Appendix (ref).
Without the constraint of monotone treatment response, the literature on risk aggregation under dependence uncertainty is rich and extends beyond the case of two variables: most notably, R82 gives the solution to the worst-case $\mathbb{P}(X+Y>k)$ and best-case $\mathbb{P}(X+Y\geq k)$. BJW14 and MW15 examine the set of all possible distributions of the sum of several random variables with given marginal distributions; bounds on best- and worst-case values of the quantile of the sum are given in EWW15 and algorithms to numerically compute these values are developed in EPR13 and BLLW24.
All random variables live on an atomless probability space $(\Omega, \mathcal{F}, \mathbb{P})$. All order terms like “increasing" are in the weak sense. Write $\mathbb{R}_+=[0,\infty)$. Let $\mathcal{M}$ be the set of all finite Borel measures on $\mathbb{R}$, equipped with the usual order $\le$. Note that $(\mathcal{M},\leq)$ is a lattice; $\wedge$ and $\vee$ respectively mean the infimum and supremum operators on that lattice. The probability measure $\delta_x$ is the point-mass at $x\in \mathbb{R}$. The survival function of a measure $\mu\in\mathcal{M}$ is the function $S_{\mu} : \mathbb{R} \to \mathbb{R}_+ $ given by $S_\mu(x)= \mu((x,\infty))$. We write $\mu \le_{\rm st} \nu$ if $S_{\mu}\le S_{\nu}$. An equivalent property to $\mu \le_{\rm st} \nu$ is
Note that $\mu\le \nu$ implies $\mu\le_{\rm st } \nu$, and both relations are transitive. The quantile function of a measure $\mu\in\mathcal{M}$ is the function $F^{-1}_\mu : (0,\mu(\mathbb{R})) \to \mathbb{R}$ defined as $ F^{-1}_{\mu}(u) = \inf\{x\in\mathbb{R}: \mu(\mathbb{R})-S_{\mu}(x) \geq u\},$ $u\in(0,1).$ For a random variable $Z$, its distribution is the probability measure $\mu\in \mathcal M$ given by $\mu(A)=\mathbb{P}(Z\in A)$ for Borel $A\subseteq \mathbb{R}$, and we denote this by $Z\buildrel \mathrm{d} \over \sim \mu$. While the exposition on treatment effects in the Introduction considered only probability measures (instead of general measures), we do not necessarily make this assumption henceforth and allow for non-probability measures as well in most of our results.
Let $\mathbb H=\{(x,y)\in \mathbb{R}^2: y \ge x\}$. Let $\Pi$ be the set of all Borel measures on $\mathbb{R}^2$. For $\pi\in \Pi$, we let $P_1 (\pi)$ denote the first marginal of $\pi$ and $P_2( \pi)$ the second marginal of $\pi$. For $\mu,\nu \in \mathcal M$, define the set
which is a set of semi-couplings between $\mu$ and $\nu$ in the sense of HS13. For simplicity, we refer to any element in $ \mathcal{H}(\mu,\nu)$ as a coupling between $\mu$ and $\nu$. For $ \mathcal{H}(\mu,\nu)$ to be nonempty, it is necessary and sufficient that $\mu\le_{\rm st}\nu$; see MS02. Note that $\pi(\mathbb{R}^2) = \mu(\mathbb{R})\le \nu(\mathbb{R})$ for all $\pi\in \mathcal H(\mu,\nu)$ when $\mu \le_{\rm st} \nu$. AMZ20 and NW22 analyzed $ \mathcal{H}(\mu,\nu)$ when $\mu$ and $\nu$ are probability measures.
For measures $\mu$ and $\nu$ with $\mu\le_{\rm st} \nu$ that are not necessarily probability measures, we generalize problems (ref)--(ref) in the Introduction by rewriting them as
Our solutions to problems (ref)--(ref) are predicated on coupling algorithms. In this section, we present these algorithms, solving the problems for specific discrete distributions, and discuss the adequacy of approximating the solutions for the general discrete case and the continuous case this way. Although problems (ref) and (ref) look very similar, the solution to one does not follow directly from the other's by symmetry.
For $\mu$ and $\nu$ that are probability measures with finite means and $\mu\leq_{\rm st}\nu$, crude bounds for $Q^{\inf}_k(\mu,\nu)$ and $Q^{\sup}_k(\mu,\nu)$ are,
where the last inequality follows from Markov's inequality. The algorithms presented below leverage the identification of $\mu$ and $\nu$ to produce improved bounds, tailored to the specified marginal distributions.
We first study a special case of discrete measures $\mu$ and $\nu.$ A measure $\mu$ is atomic with size $a>0$ if
for some $x_1, \ldots, x_m \in\mathbb{R}$. When $\mu $ has total mass $1$ (that is, $a=1/m$), $\mu$ is simply an empirical distribution. We call $x_1, \ldots, x_m$ the locations of atoms, $m$ the number of atoms, and $a$ the size of an atom; note that multiple atoms may be at identical locations. For atomic measures $\mu$ and $\nu$ with the same size (meaning that they share the same $a>0$), if $\mu\le_{\rm st}\nu$ holds, then $m\le n$, where $m$ is the number of atoms in $\mu$ and $n$ is the number of atoms in $\nu$.
As our first main result, the minimization in (ref) is attained by Algorithm (ref), and the maximization in (ref) is attained by Algorithm (ref). We explain the intuition behind these algorithms first. Throughout, for ease of exposition, let $x$ denote a generic atom location in $\mu$ and $y$ denote a generic atom location in $\nu$.
For the construction of the pairs $(x,y)$ in the optimal $\pi$ for (ref), we couple atoms in $\mu$ iteratively, starting from largest location to smallest location. For each atom, when its turn to be coupled comes, we check whether there is an uncoupled atom in $\nu$ allowing to satisfy $y-x\leq k$. If so, it couples to the one at the largest location among them. If it is unable to satisfy $y-x\leq k$, it sacrifices itself and takes down the atom in $\nu$ at the largest location; this effectively helps the next-to-be-coupled atom in $\mu$, by not being a nuisance, since that probability atom in $\nu$ would not allow any atom at smaller locations in $\mu$ to satisfy $y-x\leq k$ anyway. By always transporting to atoms in $\nu$ at the largest allowed locations, either when helping toward the objective $y-x\leq k$ or not, we make sure that probability mass at smaller locations in $\mu$ has the best chance of satisfying $y-x\leq k$ when its turn to be coupled comes. What could occur, however, would be an atom in $\mu$ satisfying $y-x\leq k$ preventing an atom at a smaller location to do so, but this situation results in effectively the same value for $\int \mathds{1}_{\{(x,y)\in\mathbb{R}^2: y-x>k\}}\mathrm{d}\pi$, $\pi \in \mathcal{H}(\mu,\nu)$, as switching would necessarily cause the first atom to not satisfy it anymore and both atoms have the same size. This construction is made precise in Algorithm (ref).
The idea for constructing the optimal $\pi$ for (ref) is similar to the one for (ref). Iteratively for each $x$ in descending order, we couple it to the yet-uncoupled atom in $\nu$ at the smallest location $y$ such that $y-x > k$, and, if not possible, to the one at the smallest location such that $y \geq x$. This construction maximizes the total portion of mass in $\nu$ transported such that $y- x> k$. The rest of the transport instructions is predicated on the principle of least nuisance: If the mass portion is capable of helping toward the objective $y-x> k$, it should do so by employing the smallest location $y$. If it cannot help, it should be the least troublesome and take with itself the smallest possible values of $y$ under the constraint $y\geq x$: this leaves the greatest flexibility for subsequent to-be-coupled atoms to satisfy $y-x> k$. By going through atoms in $\mu$ from the largest to the smallest location, we make sure that unhelpful atoms are so by obligation. It also ensures that at least some uncoupled atom always permits $y\geq x$, given stochastic dominance. This construction is made precise in Algorithm (ref).
The optimality claimed for Algorithms (ref)--(ref) is formally stated in Theorem (ref).
The proof of Theorem (ref) is presented in Appendix (ref). Despite the similarity in the two problems (ref)--(ref) and the two algorithms, items (ref) and (ref) require different proof arguments.
Algorithms (ref)--(ref) not only provide the optimal values for (ref)--(ref), they also construct corresponding optimizers (not necessarily unique), which are described below. For any atomic coupling $\pi$ with size $a$, define its coupling multiset as the multiset $\Gamma$ of atom locations satisfying
Note that, for a given size $a$, a coupling multiset entirely defines a coupling. We denote by $\Gamma^{\rm A}_{k,\mu,\nu}$ the coupling multiset made of pairs $(x_{i},y_{d_i}) $ produced by Algorithm (ref). The corresponding coupling is denoted by $\gamma^{\rm A}_{k,\mu,\nu} $, that is, $$\gamma^{\rm A}_{k,\mu,\nu} = a\sum_{(x,y)\in\Gamma^{\rm A}_{k,\mu,\nu}}\delta_{(x,y)},$$ where $a$ is the size of one atom. In terms of optimal transport, $\gamma^{\rm A}_{k,\mu,\nu}$ is an unbalanced transport plan (“unbalanced" reflects the possibility of $\mu(\mathbb{R})< \nu(\mathbb{R})$), and it is optimal by Theorem (ref). Similarly, let $\Gamma^{\rm B}_{k,\mu,\nu}$ be the coupling multiset produced by Algorithm (ref), and $\gamma^{\rm B}_{k,\mu,\nu}$ be the corresponding transport plan. We provide examples of $\Gamma^{\rm A}_{k,\mu,\nu}$ and $\Gamma^{\rm B}_{k,\mu,\nu}$ for specific $\mu$ and $\nu$ below.
It is clear from Figure (ref) that $\gamma^{\rm A}_{k, \mu, \nu}$ is not the unique transport plan solving (ref), as one could have rather coupled, in Example (ref), $x_4$ to $y_5$ and $x_5$ to $y_4$, still respecting the constraints, and it would have no incidence on the probability of worthwhile effect. Similarly, an alternative to $\gamma^{B}_{k,\mu,\nu}$ in the case of Example (ref) would be the transport plan obtained by rather coupling $x_1$ to $y_2$ and $x_2$ to $y_3$.
The next result highlights some stability from the algorithms, in the sense that the coupling multiset obtained may be separated in two parts, depending on whether atom pairs meet the objective or not, and re-running the algorithm on either part will not alter the coupling multiset.
The separability of the coupling multiset will prove useful below in showing the stability of the couplings produced by the algorithms for limit cases of their applications.
This subsection gathers some results on the couplings solving (ref)--(ref); the aim is to provide further insight on the structure of such couplings and how small variations in the parameters of the problems may affect the solution. Although we previously only discussed solving (ref)--(ref) for atomic measures with same size, the results in this subsection hold for any types of measures. Theorems (ref) and (ref) below will be required in solving the problems for general discrete measures and continuous measures, in the next subsections.
First, we remark that values $Q_k^{\inf}(\mu,\nu)$ and $Q_{k}^{\sup}(\mu,\nu)$ not only bound the possible probabilities of worthwhile treatment effect for $\mu$ and $\nu$, but mark the interval of all of their possible values; we have
This is because the probability $\int \mathds{1}_{\{y-x>k\}}\mathrm{d} \pi$ is linear in the coupling $\pi$ and $\mathcal{H}(\mu,\nu)$ is convex. The attainability of the lower bound is claimed in Theorem (ref) just below; the upper bound may be not attained, see Appendix (ref).
The selection result of Theorem (ref) is noteworthy in particular since $\nu\leq_{\rm st} \nu^{\prime}$ does not imply $Q_{k}^{\inf}(\mu,\nu) \leq Q_{k}^{\inf}(\mu,\nu^{\prime})$, as in the following example.
In the case where $\mu$ and $\nu$ are probability measures and $k=0$, the upper bound in (ref) is attained and the solution is given by the comonotonic coupling, as indicated in the following proposition. Apart from the case $k=0$, however, the optimizers do not have a simple dependence structure like comonotonicity or countermonotonicity.
For the remainder of Section (ref) (that is, also for upcoming Sections (ref) and (ref)), we will discuss only the problem in (ref). While results for (ref) would be similar in essence, some adjustments to the mathematical settings are required, which we believe would overburden the discussion; rather, see Appendix (ref). The result in the following theorem constrains how much $Q_k^{\inf}$ varies when adding or removing mass in the coupling pool.
Theorem (ref) indicates that adding or removing small portions of mass in $\mu$ or $\nu$ cannot snowball into large variations for $Q_k^{\inf}(\mu,\nu)$; the disruption is at most the amount of mass added. Let us remark that the 0 lower bound is non-trivial for $\mu \neq \mu^{\prime}$. In particular, it indicates that a portion of mass in $\mu$ never leads to major compromises in satisfying the objective $y-x\leq k$ because it has to be coupled. In Appendix (ref), we provide an example where this would be the case for $Q_k^{\sup}(\mu,\nu)$, and this predicates the need to tweak the setting accordingly for approximations regarding problem (ref).
We next extend our solutions to (ref)--(ref) for $\mu$ and $\nu$ that are discrete measures but not necessarily atomic with the same size. These solutions are obtained as limit cases to those for measures that are atomic with the same size.
Let $\mu$ and $\nu$ be two discrete measures such that $|\mathrm{supp}(\nu)|<\infty$ and $\mu\leq_{\rm st}\nu$. Define $\widetilde{\mu}_r$ and $\widetilde{\nu}_r$, $r\in\mathbb{N}$, as atomized versions of $\mu$ and $\nu$, described by
and this construction ensures that the ordering constraint $\widetilde{\mu}_r\le_{\rm st} \widetilde{\nu}_r$ holds for every atomization level $r$. It also highlights the importance of allowing for $\mu(\mathbb{R})<\nu(\mathbb{R})$ in Theorem (ref) because we will apply it to $ \widetilde{\mu}_r $ and $ \widetilde{\nu}_r$, which do not have the same total mass in general. We require the support of $\nu$ to be finite, otherwise $\widetilde{\nu}_r$ would have infinite mass for any $r\in\mathbb{N}$.
Executing the algorithm on atomized versions of $\mu$ and $\nu$ will not only yield an approximate value of $Q^{\inf}_k(\mu,\nu)$, it most notably produces a coupling that is distributionally close to the optimal coupling, thus giving valuable information on the dependence scheme solving (ref).
For discrete $\mu$ and $\nu$, define $\gamma^{A}_{k,\mu,\nu} = \lim_{r\to\infty} \gamma^{A}_{k,\widetilde{\mu}_r,\widetilde{\nu}_r} $ as per Theorem (ref). Although Proposition (ref) and Theorem (ref) require $|\mathrm{supp}(\nu)|<\infty$, it is possible to handle cases where the measures have infinite support. For example, in cases where $\mu$ and $\nu$ have a support that is unbounded to the right, one does not have a natural starting point to construct $\gamma^{\mathrm{A}}_{k,\mu,\nu}$ and the atomization of $\nu$ would produce unbounded measures, but a possible workaround is by finding $t\in\mathbb{R}$ such that $\mu \vert_{[t,\infty)} \leq \nu^{[-k]}\vert_{[t+k,\infty)} $, where $\nu^{[-k]}(A)=\nu(A+k)$ for every Borel set $A$ on $\mathbb{R}$. Then, it is certain that all atom locations in $[t,\infty)$ are able to be coupled so that they satisfy the objective $y-x \leq k$, specifically by coupling every $x$ to $y=x+k$. We continue by constructing the coupling leftward from $t$. We illustrate this in the following example.
We now let $\mu$ and $\nu$ be absolutely continuous measures, with survival functions $S_{\mu}$ and $S_{\nu}$, whose derivatives are $-f_{\mu}$ and $-f_{\nu}$; we call $f_{\mu}$ and $f_{\nu}$ the densities of $\mu$ and $\nu$. We obtain the solution to (ref) as a limit case to solutions for the discrete case.
Consider $\overline{\mu}_{s}$ and $\overline{\nu}_{s}$, for $s\in\mathbb{N}$, discretized versions of $\mu$ and $\nu$ defined by
For a discretization level $s$, mass of interval $[x,x+{2}^{-s})$ in $\mu$, rounded down according to the smallest value of $f_{\mu}$ on that interval, is attributed to the middle point of the interval in $\overline{\mu}_s$; same for $\overline{\nu}_s$, but the amount of mass is rounded up according to the largest value of $f_{\nu}$ on the interval. Note how the intervals are designed such they divide in two going from discretization level $s$ to $s+1$, and thus no two discretization levels share points for their support. If $\mu\leq_{\rm st}\nu$, then $\overline{\mu}_{s} \leq_{st} \overline{\nu}_{s}$ holds true for every $s\in\mathbb{N}$; indeed, for every $z \in\mathrm{supp}(\overline{\mu}_{s})\cup \mathrm{supp}(\overline{\nu}_{s})$,
Note that we require $f_{\mu}$ and $f_{\nu}$ to be piecewise Lipschitz continuous to have some regularity in approximating $\mu$ and $\nu$ by $\overline{\mu}_{s}$ and $\overline{\nu}_{s}$ as we let $s$ become large; piecewise Lipschitz continuity, combined with compactness of $\mathrm{supp}(\nu)$, also ensures that $\overline{\nu}_{s}$ has finite total mass for every discretization level $s\in\mathbb{N}$. The construction in (ref)--(ref) not only permits to approximate the value of $Q_k^{\inf}(\mu,\nu)$ with measures on which Algorithm (ref) is applicable, by Propositions (ref)--(ref), applications of Algorithm (ref) will moreover yield a consistent coupling as $s$ tends to infinity.
Define $\gamma_{k,\mu,\nu}^{\rm A} = \lim_{s\to\infty} \gamma_{k,\overline{\mu}_{s}, \overline{\nu}_{s}}^{\rm A}$ as per Theorem (ref). We illustrate the discretization approach in the following example.
The transport plan from Example (ref) shows that in general $\gamma^{A}_{k,\mu,\nu}$ is not Monge-type. A transport plan is Monge-type if the corresponding transport map is described by a deterministic kernel. As illustrated by the double arrows in Figure (ref), the kernel transporting from $\mu$ to $\nu$ in the example is two-atomic for $x\in [0,1]$, namely $x\mapsto (\delta_x + \delta_{3+x})/2$ on that portion.
Hitherto, we only discussed binary treatments, that is, treatments that may be applied either fully or not, without gradation. One may inquire about the case where a participant may receive partial treatment, producing new marginal response distributions to account for. While still interested in the probability of worthwhile effect of the full treatment, we may impose the coupling to be monotone with respect to the partial-treatment responses as well. Example (ref) below shows, in a simple case, how this constrains the admissible couplings for $X$ and $Y$.
In Example (ref), we only considered one degree of partial treatment, but one could have more, leading to more layers of marginal distributions for which monotonicity has to be satisfied in the coupling. Let $Z_1,\ldots,Z_h$ be the responses to each of $h\in\mathbb{N}$ degrees of partial treatment, in increasing order of intensity, and their distributions are denoted by $\eta_1,\ldots, \eta_h$. Aligned with the assumptions regarding full treatment, we suppose that $\eta_1,\ldots,\eta_h$ are fully identified. Monotone treatment response translates to $X\leq Z_1\leq \cdots \leq Z_h \leq Y$. For this to be possible, it must hold that $\mu \leq_{\rm st} \eta_1 \leq_{\rm st} \cdots \leq_{\rm st} \eta_h \leq_{\rm st} \nu$. In the general setting, we allow $\eta_1,\ldots,\eta_h\in\mathcal{M}$ to be non-probability measures, and this stochastic ordering must also hold.
While we do not directly examine the effect of partial treatment, it gives valuable information by restricting the set of admissible measures in $\mathcal{H}(\mu,\nu)$. For any $h\in\mathbb{N}$, define
Let $\Theta_{h}$ be the set of finite measures on $\mathbb{R}^{h+2}$ supported on $\mathbb{H}_h$. Define
For $\theta\in \Theta_h$, we let $P_{XY}(\theta)$ denote the pairwise marginal of $\theta$ for the first and last dimensions. The set of admissible measures satisfying the additional partial-treatment constraints is given by
The worst-case and best-case probabilities of worthwhile effect become
Because $\mathcal{H}(\mu,\nu \mid \eta_1,\ldots,\eta_h) \subseteq \mathcal{H}(\mu,\nu)$, we have the following elementary bounds for the solutions of (ref)--(ref):
Modifying Algorithms (ref) and (ref) to consider the partial-treatment responses, we hereby present Algorithms (ref) and (ref). Our second main result is that Algorithm (ref) solves the minimization problem in (ref) and that Algorithm (ref) solves the maximization problem in (ref). We first discuss the intuition behind each.
Algorithm (ref) employs the same heuristics as Algorithm (ref) in identifying optimal atom locations in $\nu$ to which each atom in $\mu$ should be coupled. Additional steps during each of these couplings assess whether the partial-treatment constraints may be satisfied when this optimal coupling yields $y-x\leq k$. If it is not possible, the algorithm will recommend to take the atoms in $ \eta_1,\ldots,\eta_h$ and $\nu$ at the largest locations. If it is possible, that atom in $\nu$ at this location is selected along the rightmost atoms in $\eta_1,\ldots,\eta_h$ that are to the left of the upper-degree treatment. Indeed, the directionality of monotone-treatment-response constraints makes atoms in $\eta_1,\ldots,\eta_h$ at the larger locations the more restrictive. The idea is therefore that committing the larger-location atoms in $\eta_1,\ldots,\eta_h$ to the coupling of a given atom in $\mu$ is the lesser nuisance for the atoms that are yet to be coupled. The main difficulty is to prove that, by this choice, an atom in $\mu$ satisfying the objective cannot prevent more than one atom from satisfying it because of these committed atoms in $\eta_1,\ldots,\eta_h$. This is the main part of the argument in the proof of Theorem (ref) in Appendix (ref). If the algorithm recommends the largest atom in $\nu$, there is no need to verify whether the ordering constraint may be satisfied: it is guaranteed to be possible from stochastic ordering. We present an application of Algorithm (ref) below.
Algorithms (ref) and (ref) differ more than just from the heuristic difference between Algorithms (ref) and (ref); they also differ in how they verify that the partial-treatment constraint is satisfied. As stated above, the larger the location of an atom in any of $\eta_1,\ldots,\eta_h$, the more restrictive it is in coupling $\mu$ to $\nu$. Algorithm (ref) always tries to select the largest location for atoms in $\nu$, among either set of the atoms helping toward the objective or the others. Therefore, this selection mechanism flows in the same direction as the restrictiveness of the atoms in $\eta_1,\ldots,\eta_h$. Hence, we only need to check whether one atom in $\nu$, in the set of those meeting the objective, allows to satisfy the partial-treatment constraints; if it does not, the others in that set, at smaller locations, do not either. Algorithm (ref), by contrast, always aims to select atoms at smallest locations, whether it is among the set of those meeting the objective $y-x > k$ or the others, meeting $y\geq x$. Hence, its selection process is opposite to the partial-treatment restrictiveness direction. As a result, if the favored atom in the set meeting the objective $y-x > k$ cannot satisfy the partial treatment constraint, the algorithm must check every other atom in that set, as they might satisfy it, before opting for the other set, for which the same process must be repeated. Algorithm (ref) has therefore a higher degree of complexity than Algorithm (ref), from this opposite directionality.
By using the same atomization and discretization methods, we can also obtain results similar to those in Sections (ref)--(ref) under partial-treatment constraints.
The approach from Algorithms (ref) and (ref) does not extend to continuous treatments, where we would have a continuum of measures $\{\eta_u: u\in (0,1)\}$ to represent the degrees of partial treatment. Let us remark, however, that continuity of treatments conflicts with the assumption of fully identified marginal distributions (and no further structural assumptions) in practical applications: for continuous treatments, almost surely no two subjects will be administered the same level of treatment and thus only one point may be observed for each marginal distribution.
The probability of worthwhile effect for a treatment requires, to be computable, identification of the joint distribution of treatment responses, unlike the average treatment effect, which is commonly employed but ill-suited to some contexts. We examined the range of possible probabilities of worthwhile effect of a monotone-response treatment with the assumption that only marginal distributions are fully identifiable. With no further identification assumptions, Algorithms (ref) and (ref) yield the extremal probabilities of worthwhile effect for atomic distributions. Results for discrete and continuous distributions follow by convergence after atomization and discretization. Partial identification may come from responses to intermediate degrees of treatment, in which case Algorithms (ref) and (ref) produce the extremal probabilities of worthwhile effect. While we mainly grounded this paper in the statistical considerations of treatment effects, these results also transpose to other frameworks; they notably solve extant problems in optimal transport and risk aggregation under model uncertainty.
Ruodu Wang is supported by the Natural Sciences and Engineering Research Council of Canada (RGPIN-2018-03823, RGPAS-2018-522590) and Canada Research Chairs (CRC-2022-00141).