EconBase
← Back to paper

Quantifying Distributional Model Risk in Marginal Problems via Optimal Transport

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

107,282 characters · 33 sections · 119 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Quantifying Distributional Model Risk in Marginal Problems via Optimal Transport

abstractThis paper studies distributional model risk in marginal problems, where each marginal measure is assumed to lie in a Wasserstein ball centered at a fixed reference measure with a given radius. Theoretically, we establish several fundamental results including strong duality, finiteness of the proposed Wasserstein distributional model risk, and the existence of an optimizer at each radius. In addition, we show continuity of the Wasserstein distributional model risk as a function of the radius. Using strong duality, we extend the well-known Makarov bounds for the distribution function of the sum of two random variables with given marginals to Wasserstein distributionally robust Markarov bounds. Practically, we illustrate our results on four distinct applications when the sample information comes from multiple data sources and only some marginal reference measures are identified. They are: partial identification of treatment effects; externally valid treatment choice via robust welfare functions; Wasserstein distributionally robust estimation under data combination; and evaluation of the worst aggregate risk measures.

Introduction

Distributionally robust optimization (DRO) has emerged as a powerful tool for hedging against model misspecification and distributional shifts. It minimizes distributional model risk (DMR) defined as the worst risk over a class of distributions lying in a distributional uncertainty set, see blanchet2019quantifying. Among many different choices of uncertainty sets, Wasserstein DRO (W-DRO) with distributional uncertainty sets based on optimal transport costs has gained much popularity, see Kuhn_2019 and Blanchet2021 for recent reviews. W-DRO has found successful applications in robust decision making in all disciplines including economics, finance, machine learning, and operations research. Its success is largely credited to the strong duality and other nice properties of the Wasserstein DMR (W-DMR). The objective of this paper is to propose and study W-DMR in marginal problems where only some marginal measures of a reference measure are given, see e.g., Kellerer:1984to,rachev1998mass,villani2009optimal, villani2021topics, and ruschendorf1991bounds.

In practice, marginal problems arise from either the lack of complete data or an incomplete model. In insurance and risk management, computing model-free measures of aggregate risks such as Value-at-Risk and Expected Short-Fall is of utmost importance and routinely done. When the exact dependence structure between individual risks is lacking, researchers and policy makers rely on the worst risk measures defined as the maximum value of aggregate risk measures over all joint measures of the individual risks with some fixed marginal measures, see Embrechts_2010 and Embrechts_2013; In causal inference, distributional treatment effects such as the variance and the proportion of participants who benefit from the treatment depend on the joint distribution of the potential outcomes. Even with ideal randomized experiments such as double-blind clinical trials, the joint distribution of potential outcomes is not identified and as a result, only the lower and upper bounds on distributional treatment effects are identified from the sample information, see FAN_2009, Fan_2010_SharpBounds,Fan_2012_CI_quantiles, Fan2017, Ridder_2007, Firpo_2019; In algorithmic fairness when the sensitive group variable is not observed in the main data set, assessment of unfairness measures must be done using multiple data sets, see Kallus2022. Abstracting away from estimation, all these problems involve optimizing the expected value of a functional of multiple random variables with fixed marginals and thus belong to the class of marginal problems for which optimal transport related tools are important.\footnote{When the marginals are univariate, optimal transport problem can be conveniently expressed in terms of copulas. Fan_2010_SharpBounds, Fan_2012_CI_quantiles, FAN_2009, Fan2017, Ridder_2007, and Firpo_2019 explicitly use copula tools. }

The marginal measures in the afore-mentioned applications and general marginal problems are typically empirical measures computed from multiple data sets such as in the evaluation of worst aggregate risk measures or identified under specific assumptions such as randomization or strong ignorability in causal inference. Developing a unified framework for hedging against model misspecification and/or distributional shifts in marginal measures motivates the current paper.

Theoretically, this paper makes several contributions to the literature on distributional robustness and the literature on marginal problems. First, it introduces Wasserstein distributional model risk in marginal problems (W-DMR-MP), where each marginal measure is assumed to lie in a Wasserstein ball centered at a fixed reference measure with a given radius. We focus on the important case with two marginals and consider both non-overlapping and overlapping marginals. For non-overlapping marginal measures, when the radius is zero, the W-DMR-MP reduces to the marginal problems or optimal transport problems studied in Kellerer:1984to,rachev1998mass,villani2009optimal, villani2021topics. For overlapping marginals, when the radius is zero, the W-DMR-MP reduces to the overlapping marginals problem studied in ruschendorf1991bounds; Second, we establish strong duality for our W-DMR with both non-overlapping and overlapping marginals under similar conditions to those for W-DMR, see zhang2022simple, blanchet2019quantifying, and gao2022WDRO. As a first application of our strong duality result for non-overlapping marginals, we extend the well-known Marakov bounds for the distribution function of the sum of two random variables to Wasserstein distributionally robust Makarov bounds; Third, we prove finiteness of the W-DMR-MP and existence of an optimizer at each radius. Based on both results, we show that the identified set of the expected value of a smooth functional of random variables with fixed marginals is a closed interval; Fourth, we show continuity of the W-DMR in marginal problems as a function of the radius. Together these results extend those for W-DMR in blanchet2019quantifying, zhang2022simple, and yue2022linear; Lastly, we extend our formulations and theory to W-DMR with multi-marginals. On a technical note, our proofs build on existing work on W-DMR such as blanchet2019quantifying, zhang2022simple, and yue2022linear. However, an additional challenge due to the presence of multiple marginal measures in our Wasserstein uncertain sets is the verification of the existence of a joint measure with overlapping marginals. We make use of existing results for a given consistent product marginal system in vorob1962consistent, kellerer1964verteilungsfunktionen, and shortt1983combinatorial to address this issue.

Practically, we demonstrate the flexibility and broad applicability of our W-DMR-MP via four distinct applications when the sample information comes from multiple data sources. First, we consider partial identification of treatment effects when the marginal measures of the potential outcomes lie in their respective Wasserstein balls centered at the measures identified under strong ignorability. The validity of strong ignorability is often questionable when unobservable confounders may be present. We apply our W-DMR-MP to establishing the identified sets of treatment effects which can be used to conducting sensitivity analysis to the selection-on-observables assumption. For average treatment effects, we show that when the cost functions are separable, incorporating covariate information does not help shrink the identified set; on the other hand, for non-separable cost functions such as the Mahalanobis distance, incorporating covariate information may help shrink the identified set; Second, in causal inference when the optimal treatment choice is to be applied to a target population different from the training population, adjaho2022 introduces robust welfare functions defined by W-DMR to study externally valid treatment choice. The W-DMR-MP we propose allows us to dispense with the assumption of a known dependence structure for the reference measure in adjaho2022. When shifts in the covariate distribution are allowed, we show that our robust welfare function is upper bounded by the worst robust welfare function of adjaho2022; Third, one important application of W-DMR is in distributionally robust estimation and classification. However as awasthi2022distributionally points out,\footnote{See Graham2016 and Chen_2008 for general data combination problems.} some sensitive variables may not be observed in the same data set as the response variable rendering W-DRO inapplicable. We apply W-DMR-MP to distributionally robust estimation under data combination;\footnote{(ref) provides a detailed comparison of our set up and awasthi2022distributionally.} Fourth, applying our W-DMR-MP to the evaluation of the worst aggregate risk measures allows us to dispense with the known marginals assumption in Embrechts_2010 and Embrechts_2013.

The rest of this paper is organized as follows. (ref) reviews the W-DMR and strong duality, introduces our W-DMR-MP, and then presents four motivating examples. (ref) establishes strong duality and Wasserstein distributionally robust Marakov bounds. (ref) studies finiteness of W-DMR-MP and existence of optimal solutions. Moreover, we show that the identified set of the expected value of a smooth functional of random variables with fixed marginals is a closed interval. (ref) establishes continuity of W-DMR-MP as a function of the radius. (ref) revisits the motivating examples in (ref). (ref) extends our W-DMR-MP to more than two marginals. The last section offers some concluding remarks. Technical proofs are relegated to a series of appendices.

We close this section by introducing the notation used in the rest of this paper. For two sets $A$ and $B$, the relative complement is denoted by $A \setminus B$. Let $\overline{\mathbb{R}} = \mathbb{R} \cup \left\{- \infty, \infty \right \}$, $[d]=\{1,2,...,d\}$, $\mathbb{R}^d_{+} = \left \{ x \in \mathbb{R}^d : x_i \geq 0, \ \forall i \in [d] \right\}$, and $\mathbb{R}^d_{++} = \left\{ x \in \mathbb{R}^d : x_i > 0, \ \forall i \in [d] \right\}$. For any real numbers $x, y \in \mathbb{R}$, we define $x \wedge y := \min\{x, y\}$ and $x \vee y : = \max\{x, y\}$. The Euclidean inner product of $x$ and $y$ in $\mathbb{R}^d$ is denoted by $\langle x, y \rangle$. For any real matrix $W \in \mathbb{R}^{m \times n}$, let $A^\top$ denote the transpose of $W$. For an extended real function $f$ on $\mathcal{X}$, the positive part $f^+$ and the negative part $f^-$ are defined as $f^{+}(x) = \max \left\{ f(x), 0 \right\}$ and $f^{-}(x) = \max \left\{ -f(x), 0 \right\}$, respectively.

For any Polish space $\mathcal{S}$, let $\mathcal{B}_\mathcal{S}$ be the associated Borel $\sigma$-algebra and $\mathcal{P}(\mathcal{S})$ be the collection of probability measures on $\mathcal{S}$. Given a Polish probability space $(\mathcal{S}, \mathcal{B}_{\mathcal{S}}, \nu)$, let $\mathcal{B}_{\mathcal{S}}^\nu$ denote the $\nu$-completion of $\mathcal{B}_{\mathcal{S}}$. Given a probability space $\left(\Omega, \mathcal{F}, \mathbb{P} \right)$ and a map $T: \Omega \rightarrow \mathcal{S}$, let $T\# \mu$ denote the push forward of $\mathbb{P}$ by $T$, i.e., $(T\# \mathbb{P})(A) = \mathbb{P} \left( T^{-1} (A) \right)$ for all $A \in \mathcal{B}_{\mathcal{S}}$, where $T^{-1}(A) = \left\{ \omega\in\Omega: T(\omega) \in A \right\}$. The law of a random variable $S: \Omega \rightarrow \mathbb{R}$ is denoted by $\mathrm{Law}(S)$ which is the same as $S\# \mathbb{P}$. For any $\mu, \nu \in \mathcal{P}(\mathcal{S})$, let $\Pi(\mu, \nu)$ denote the set of all couplings (or joint measures) with marginals $\mu$ and $\nu$.

For any $\mathcal{B}_{\mathcal{S}}^{\nu}$-measurable function $f$, let $\int_{\mathcal{S}} f d \nu$ denote the integral of $f$ in the completion of $(\mathcal{S}, \mathcal{B}_{\mathcal{S}}, \nu)$. For a random element $S: \Omega \rightarrow \mathcal{S}$ with $\mathrm{Law}(S) = \nu$, we write $\mathbb{E}_\nu \left[f(S)\right] = \int_{\mathcal{S}} f d \nu$. Given $p \in (0, \infty)$ and a Borel measure $\nu$ on $\mathcal{S}$, let $L^p(\nu) := L^p(\mathcal{S},\mathcal{B}_{\mathcal{S} } , \nu)$ denote the set of all the $\mathcal{B}_{\mathcal{S}}^{\nu}$-measurable functions $f: \mathcal{S}\rightarrow \mathbb{R}$ such that $\| f\|_{L^p(\nu)} := \left( \int_{\mathcal{S}} |f|^p d\nu \right)^{1/p}< \infty$.

W-DMR and Motivating Examples

In this section, we first review W-DMR and then introduce W-DMR in marginal problems. Lastly, we present four motivating examples of marginal problems which will be used to illustrate our results in the rest of this paper.

A Review of W-DMR and Strong Duality

W-DMR is defined as the worst model risk over a class of distributions lying in a Wasserstein uncertainty set composed of all probability measures that are a fixed Wasserstein distance away from a given reference measure, see blanchet2019quantifying.

Before presenting W-DMR, we review some basic definitions. Let $\mathcal{X}$ be a Polish (metric) space with a metric $\boldsymbol{d}$.

definition[Optimal transport cost] Let $\mu, \nu \in \mathcal{P}(\mathcal{X})$ be given probability measures. The {\it optimal transport cost} between $\mu$ and $\nu$ associated with a cost function $c: \mathcal{X}\times \mathcal{X} \rightarrow \mathbb{R}_+ \cup \{\infty\}$ is defined as \[ \boldsymbol{K}_c(\mu, \nu) = \inf_{\pi \in \Pi(\mu, \nu)} \int_{\mathcal{X} \times \mathcal{X}} c \, d\pi. \]

When the cost function $c$ is lower-semicontinuous, there exists an optimal coupling corresponding to $\boldsymbol{K}_c(\mu, \nu)$. In other words, there exists $\pi^* \in \Pi(\mu, \nu)$ such that $\boldsymbol{K}_c(\mu, \nu)=\int_{\mathcal{X} \times \mathcal{X}} c \, d\pi^*$ villani2009optimal.

definition[Wasserstein distance] Let $p \in [1, \infty)$. The {\it Wasserstein distance} of order $p$ between any two measures $\mu$ and $\nu$ on Polish metric space $(\mathcal{X}, \boldsymbol{d})$ is defined by \begin{align*} \boldsymbol{W}_{ p}(\mu, \nu) &= \left[ \inf_{\pi \in \Pi(\mu, \nu)} \int_{\mathcal{X} \times \mathcal{X} } \boldsymbol{d}^p \, d \pi \right]^{1/p}. \end{align*}

Throughout this paper, we make the following assumption on the cost function $c$.

assumptionLet $(\mathcal{X}, \mathcal{B}_{\mathcal{X}})$ be a Borel space associated to $\mathcal{X}$. The cost function $c: \mathcal{X} \times \mathcal{X} \to \mathbb{R}_{+} \cup \{\infty\}$ is measurable and satisfies $c(x, y) = 0$ if and only if $x = y$.

(ref) implies that for $\mu, \nu \in \mathcal{P}(\mathcal{X})$, $\mu = \nu$ if and only if $\boldsymbol{K}_c(\mu, \nu) = 0$. When $c$ is the metric $\boldsymbol{d}$ on $\mathcal{X}$, $ \boldsymbol{K}_c(\mu, \nu)$ coincides with the Wasserstein distance of order 1 (Kantorovich-Rubinstein distance) between $\mu$ and $\nu$ defined in (ref).

For a given function $f: \mathcal{X} \to \mathbb{R}$, blanchet2019quantifying define W-DMR as

align*[align* omitted — 160 chars of source]

where $\Sigma_{\mathrm{DMR}}(\delta)$ is the Wasserstein uncertainty set\footnote{By convention, we call all uncertainty sets based on optimal transport costs as Wasserstein uncertainty sets.} centered at a reference measure $\mu \in \mathcal{P}(\mathcal{X})$ with radius $\delta\geq 0$, i.e.,

align*[align* omitted — 140 chars of source]

(ref) allows the cost function $c$ to be asymmetric and take value $\infty$, where the latter corresponds to the case that there is no distributional shift in some marginal measure of $\mu$.

remarkUnder (ref), $\Sigma_{\mathrm{DMR}}(0)=\{\mu\}$ and \[ \mathcal{I}_{\mathrm{DMR}}(0) = \int_{\mathcal{X}} f \, d \mu. \]

It is well-known that under mild conditions, strong duality holds for $\mathcal{I}_{\mathrm{DMR}}(\delta)$ when $\delta > 0$ (c.f., blanchet2019quantifying,gao2022WDRO,zhang2022simple). To be self-contained, we restate the strong duality result in zhang2022simple for Polish space below.\footnote{The strong duality result in zhang2022simple allows for general space $\mathcal{X}$.}

theorem[{zhang2022simple}] Let $(\mathcal{X}, \mathcal{B}_{\mathcal{X}}, \mu)$ be a probability space. Let $\delta \in (0, \infty)$ and $f: \mathcal{X} \to \mathbb{R}$ be a measurable function such that $\int_{\mathcal{X}} f \, d \mu > -\infty$. Suppose the cost function satisfies (ref). Then, for any $\delta > 0$, \begin{align} \mathcal{I}_{\mathrm{DMR}}(\delta) = \inf_{\lambda \in \mathbb{R}_{+}} \left\{ \lambda \delta + \int_{\mathcal{X}} \sup_{x' \in \mathcal{X}} [ f(x') - \lambda c(x, x') ] \, d \mu(x) \right\}, \end{align} where $\lambda c(x, x')$ is defined to be $\infty$ when $\lambda = 0$ and $c(x, x') = \infty$.

In the rest of this paper, we keep the convention that for any cost function $c$, $\lambda c(x, y) = \infty$ when $\lambda = 0$ and $c(x, y) = \infty$.

W-DMR in Marginal Problems

Non-overlapping Marginals

Let $\mathcal{V} := \mathcal{S}_1 \times \mathcal{S}_2$ be the product space of two Polish spaces $\mathcal{S}_1$ and $\mathcal{S}_2$. Let $\mu_1$ and $\mu_2$ be Borel probability measures on $\mathcal{S}_1$ and $\mathcal{S}_2$ respectively. Following ruschendorf1991bounds (see also Embrechts_2010), we call the Fr\'{e}chet class of all probability measures on $\mathcal{V}$ having marginals $\mu_1$ and $\mu_2$ the Fr\'{e}chet class with non-overlapping marginals denoted as $\mathcal{F}(\mathcal{V}; \mu_{1}, \mu_{2}) := \mathcal{F}(\mu_{1}, \mu_{2})$. Note that $\mathcal{F}(\mu_{1}, \mu_{2})=\Pi(\mu_1, \mu_2)$.

Let $g:\mathcal{V}\rightarrow \mathbb{R}$ be a measurable function satisfying the following assumption.

assumptionThe function $g:\mathcal{V}\rightarrow \mathbb{R}$ is measurable such that $\int_{\mathcal{V}} g d\gamma_0 > -\infty$ for some $\gamma_0 \in \Pi(\mu_1, \mu_2) \subset \mathcal{P}(\mathcal{V})$.

The marginal problem associated with $\mu_1$ and $\mu_2$ is defined as

align*[align* omitted — 125 chars of source]

It is essentially an optimal transport problem, where the $\sup$ operation is replaced with the $\inf$ operation, see Kellerer:1984to,rachev1998mass,villani2009optimal, villani2021topics) or (ref) for a review of strong duality for $\mathcal{I}_{\mathrm{M} }(\mu_1,\mu_2)$.

The W-DMR with non-overlapping marginals we propose extends the marginal problem by allowing each marginal measure of $\gamma$ to lie in a fixed Wasserstein distance away from a reference measure. Specifically, for any $\gamma \in \mathcal{P}(\mathcal{V} )$, let $\gamma_{1}$ and $\gamma_{2}$ denote the projection of $\gamma$ on $\mathcal{S}_1$ and $\mathcal{S}_2$, respectively. The W-DMR with non-overlapping marginals is defined as

equation[equation omitted — 190 chars of source]

where $\Sigma_{\mathrm{D} }(\delta)$ is the uncertainty set given by \[ \Sigma_{\mathrm{D} }(\delta) := \Sigma_{\mathrm{D} }(\mu_1, \mu_2, \delta) = \left\{ \gamma \in \mathcal{P}(\mathcal{V}): \ \boldsymbol{K}_1( \mu_1, \gamma_1) \leq \delta_1, \ \boldsymbol{K}_2( \mu_2, \gamma_2) \leq \delta_2 \right\}, \] in which $\boldsymbol{K}_1$ and $\boldsymbol{K}_2$ are optimal transport costs associated with cost functions $c_1$ and $c_2$, respectively, and $\delta := (\delta_{1}, \delta_{2}) \in \mathbb{R}_{+}^2$ is the radius of the uncertainty set. Obviously $\Sigma_{\mathrm{D} }(\delta)$ is non-empty for all $\delta \in \mathbb{R}_{+}^2$.

remark(i) Under (ref) and (ref), it holds that $\mathcal{I}_{\mathrm{D} }(\delta) > - \infty$ for all $\delta \in \mathbb{R}_{+}^2$, see (ref); (ii) Under (ref), the uncertainty set $\Sigma_{\mathrm{D} }(0)=\Pi(\mu_{1}, \mu_{2})$ and thus $\mathcal{I}_{\mathrm{D} }(0) =\mathcal{I}_{\mathrm{M} }(\mu_1,\mu_2)$.

Overlapping Marginals

Let $\mathcal{S} := \mathcal{Y}_1 \times \mathcal{Y}_2 \times \mathcal{X}$ be the product space of three Polish spaces $\mathcal{Y}_1$, $\mathcal{Y}_2$, and $\mathcal{X}$. Let $\mathcal{S}_1 := \mathcal{Y}_1 \times \mathcal{X}$ and $\mathcal{S}_2 := \mathcal{Y}_2 \times \mathcal{X}$. Let $\mu_{13} \in \mathcal{P}(\mathcal{S}_1)$ and $\mu_{23} \in \mathcal{P}(\mathcal{S}_2)$ be such that the projection of $\mu_{13}$ and the projection of $\mu_{23}$ on $\mathcal{X}$ are the same. Following ruschendorf1991bounds (see also Embrechts_2010), we call the Fr\'{e}chet class of all probability measures on $\mathcal{S}$ having marginals $\mu_{13}$ and $\mu_{23}$ the Fr\'{e}chet class with overlapping marginals and denote it as $\mathcal{F}(\mathcal{S}; \mu_{13}, \mu_{23}) := \mathcal{F}(\mu_{13}, \mu_{23}).$ Unlike the non-overlapping case, $\mathcal{F}(\mu_{13}, \mu_{23})$ is different from the class of couplings $\Pi(\mu_{13}, \mu_{23})$.

Let $f:\mathcal{S} \rightarrow \mathbb{R}$ be a measurable function satisfying the following assumption.

assumptionThe function $f:\mathcal{S} \rightarrow \mathbb{R}$ is measurable such that $\int_{\mathcal{S}} f d\nu_0 > - \infty$ for some $\nu_0 \in \mathcal{F}(\mu_{13}, \mu_{23}) \subset \mathcal{P}(\mathcal{S})$.

ruschendorf1991bounds studies the following marginal problem with overlapping marginals:

align*[align* omitted — 148 chars of source]

As shown in ruschendorf1991bounds, the marginal problem with overlapping marginals can be computed via the marginal problem with non-overlapping marginals through the following relation:

align*[align* omitted — 204 chars of source]

where for each fixed $x\in\mathcal{X}$, $\mu_{\ell|3}( \cdot| x)$ denote the conditional measure of $Y_\ell$ given $X=x$, and the inner optimization problem is a marginal problem with non-overlapping marginals.

For any $\gamma \in \mathcal{P}(\mathcal{S} )$, let $\gamma_{13}$ and $\gamma_{23}$ denote the projections of $\gamma$ on $\mathcal{Y}_1 \times \mathcal{X}$ and $\mathcal{Y}_2 \times \mathcal{X}$, respectively. The W-DMR with overlapping marginals is defined as

equation[equation omitted — 165 chars of source]

where $\Sigma(\delta)$ is the uncertainty set given by \[ \Sigma(\delta):=\Sigma(\mu_{13}, \mu_{23}, \delta) = \left\{ \gamma \in \mathcal{P}(\mathcal{S} ): \ \boldsymbol{K}_1( \mu_{13}, \gamma_{13}) \leq \delta_1, \ \boldsymbol{K}_2( \mu_{23}, \gamma_{23}) \leq \delta_2 \right\} \] in which $\delta := (\delta_{1}, \delta_{2}) \in \mathbb{R}_{+}^2$ is the radius of the uncertainty set, and $\boldsymbol{K}_1$ and $\boldsymbol{K}_2$ are optimal transport costs associated with $c_1$ and $c_2$. We note that $\Sigma(\delta)$ is non-empty for all $\delta \in \mathbb{R}_{+}^2$.

remark(i) (ref) imply that $\mathcal{I}(\delta)>-\infty$ for all $\delta\ge 0$, see (ref); (ii) When $\delta=0$, the uncertainty set $\Sigma(0)=\mathcal{F}(\mu_{13}, \mu_{23})$ and $ \mathcal{I}(0)= \mathcal{I}_{\mathrm{M} }(\mu_{13},\mu_{23})$.

Motivating Examples

In this section, we present four distinct examples to demonstrate the wide applicability of the W-DMR in marginal problems. The first example is concerned with partial identification of treatment effect parameters when commonly used assumptions in the literature for point identification fail; the second example is concerned with distributionally robust optimal treatment choice; the third one is an application of W-DMR-MP in distributionally robust estimation under data combination; and the last one concerns measures of aggregate risk.

For the first two examples, we adopt the potential outcomes framework for a binary treatment. Let $D \in \{0,1\}$ represent an individual's treatment status, and $Y_1\in \mathcal{Y}_1\subset \mathbb{R}$ and $Y_2\in \mathcal{Y}_2\subset \mathbb{R}$ denote the potential outcomes under treatments $D = 0$ and $D = 1$, respectively. Let the observed outcome be \[ Y = D Y_2 + (1-D) Y_1. \]

To focus on introducing the main ideas, we adopt the selection-on-observables framework stated in (ref) below.

assumption\begin{enumerate} • Conditional Independence: The potential outcomes are independent of treatment assignment conditional on covariate $X\in\mathcal{X}\subset \mathbb{R}^q$ for $q \geq 1$, i.e., \[ (Y_1, Y_2) \ \rotatebox[origin=c]{90}{$\models$} \ D \;|\; X; \] • Common Support: For all $x\in \mathcal{X}$, $0<p(x)<1$, where $p(x):=\mathbb{P}(D=1|X=x)$. \end{enumerate}

Suppose a random sample on $ (Y,X,D)$ is available. Then under (ref), the marginal conditional distribution functions of $Y_1, Y_2$ given $X=x$ are point identified: \[ F_{Y_1|X}(y|x)=\mathbb{P}(Y_1\leq y|X=x)=\mathbb{P}(Y\leq y|X=x,D=0)\] and \[ F_{Y_2|X}(y|x)=\mathbb{P}(Y_2\leq y|X=x)=\mathbb{P}(Y\leq y|X=x,D=1). \] As a result, the probability measures $\mu_{13}$ of $(Y_{1},X)$ and $\mu_{23}$ of $(Y_{2},X)$ are identified as well.

Partial Identification of Treatment Effects

(ref) is commonly used to identify treatment effect parameters and optimal treatment choice. However the validity of (ref) may be questionable when there are unobserved confounders. W-DMR-MP presents a viable approach to studying sensitivity of causal inference to deviations from (ref) by varying the marginal measures of a joint measure of $(Y_1, Y_2, X)$ in Wasserstein uncertainty sets centered at reference measures consistent with (ref). Specifically, let $f$ be a measurable function of $Y_1, Y_2$. Consider treatment effects of the form: $\theta_o:=\mathbb{E}_o[f(Y_1,Y_2)]$, where $\mathbb{E}_o$ denotes expectation with respect to the true measure. It includes the average treatment effect (ATE) for which $f(Y_1, Y_2)=Y_2 - Y_1$ and the distributional treatment effect such as $\mathbb{P}_o(Y_2 - Y_1\geq 0)$, where $\mathbb{P}_o$ denotes the probability computed under the true measure.

Consider the identified set for $\theta_o$ defined as \[ \Theta(\delta) := \left\{ \int_{\mathcal{S} } f(y_1,y_2) \, d \gamma(y_1, y_2, x) : \gamma \in \Sigma( \delta) \right\}, \] where

align*[align* omitted — 196 chars of source]

in which $\mu_{13}$ and $\mu_{23}$ are the identified measures of $(Y_1, X)$ and $(Y_2, X)$ under (ref). Under mild conditions, we show in (ref) that the identified set $\Theta(\delta)$ is a closed interval given by \[ \Theta(\delta)= \left[ \min_{\gamma \in \Sigma(\delta)} \int_\mathcal{S} f(y_1,y_2) \, d \gamma(s), \max_{\gamma \in \Sigma(\delta)} \int_\mathcal{S} f(y_1,y_2) \, d \gamma(s) \right], \] where the lower and upper limits of the interval are characterized by the W-DMR-MP.\footnote{Since $\inf_{\gamma \in \Sigma(\delta)} \int_{\mathcal{S}} f(y_1, y_2) d \gamma(s)$ can be rewritten as $ - \sup_{\gamma \in \Sigma(\delta)} \int_{\mathcal{S}} [- f(y_1, y_2)] d \gamma(s)$, we also refer to the lower limit as W-DMR-MP. } When $\delta=0$, Fan2017 establish a characterization of $\Theta(0)$ via marginal problems with overlapping marginals.

The identified set $\Theta(\delta)$ can be used to conduct sensitivity analysis to deviations from (ref). We note that sensitivity analysis to other commonly used assumptions such as the threshold-crossing model can be done by taking the reference measures as the measures identified under these alternative assumptions, see FAN_2009.

Robust Welfare Function

In empirical welfare maximization (EWM), an optimal choice/policy is chosen to maximize the expected welfare estimated from a training data set and then applied to a target population, see Kitagawa_2018. EWM assumes that the target population and the training data set come from the same underlying probability measure. This may not be valid in important applications. Motivated by designing externally valid treatment policy, adjaho2022 introduces a robust welfare function which allows the target population to differ from the training population. In this paper, we revisit adjaho2022's robust welfare function and propose a new one based on W-DMR with overlapping marginals.

adjaho2022 adopts the following definition of a robust welfare function:

align*[align* omitted — 119 chars of source]

where $d : \mathcal{X} \rightarrow \{0, 1\}$ is a measurable policy function, i.e., $d(X)$ is $0$ or $1$ depending on $X$ and $\Sigma_0(\delta_0)$ is the Wasserstein uncertainty set centered at a joint measure $\mu$ for $(Y_1, Y_2, X)$ consistent with (ref), i.e.,

align*[align* omitted — 139 chars of source]

where $\boldsymbol{K}_c(\mu, \gamma)$ is the optimal transport cost with cost function $c: \mathcal{S} \times \mathcal{S} \rightarrow \mathbb{R}_{+}\cup \{\infty\}$.

Noting that (ref) only identifies the marginal measures $\mu_{13},\mu_{23}$ of the reference measure $\mu$ in $\Sigma_0(\delta_0)$, we define a new robust welfare function as

equation*[equation* omitted — 118 chars of source]

where $\Sigma(\delta) = \Sigma(\mu_{13}, \mu_{23}, \delta)$ is the uncertainty set for W-DMR with overlapping marginals.

W-DRO Under Data Combination

An important application of W-DMR is W-DRO. Let $f: \mathcal{Y}_1 \times \mathcal{Y}_2 \times \mathcal{X}\times \Theta \to \mathbb{R}$ be a loss function with an unknown parameter $\theta \in \Theta\subset \mathbb{R}^q$. W-DRO under data combination is defined as

align[align omitted — 142 chars of source]

where $\Sigma(\delta)$ is the uncertainty set for the overlapping case. For each $\theta \in \Theta$, the inner optimization is a W-DMR with overlapping marginals. In practice, we need to choose the reference measures $\mu_{13}$ and $\mu_{23}$ based on the sample information. Focusing on logit model, where $\mathcal{Y}_1 = \{+1, -1\}$ is the space for the dependent variable, and $\mathcal{Y}_2$ and $\mathcal{X}$ are feature spaces/covariate space, and

align*[align* omitted — 96 chars of source]

awasthi2022distributionally proposes a method dubbed `Robust Data Join' in which the empirical measures constructed from the two data sets are used as reference measures. Specifically, let $\widehat{\mu}_{13}$ and $\widehat{\mu}_{23}$ denote empirical measures based on two separate data sets. The uncertainty set in awasthi2022distributionally takes the following form: \[ \Sigma_{\mathrm{RDJ} } (\delta):= \left\{ \gamma \in \mathcal{P}(\mathcal{S} ): \ \boldsymbol{K}_1( \widehat{\mu}_{13}, \gamma_{13}) \leq \delta_1, \ \boldsymbol{K}_2(\widehat{\mu}_{23}, \gamma_{23}) \leq \delta_2 \right\}, \] where \[ c_1((y_1, x), (y_1', x')) = \| x - x' \|_{p} + \kappa_1 | y_1 - y_1'| \quad \text{and} \] \[ c_2((y_2, x), (y_2, x')) = \| x - x' \|_{p} + \kappa_2 \| y_2 - y_2' \|_{p'} \] with $\kappa_1\geq 1$, $\kappa_2\geq 1$, $p\geq 1$, and $p'\geq 1$.

Note that awasthi2022distributionally's `Robust Data Join' is different from our W-DMR with non-overlapping marginals because the measure of interest $\gamma \in \mathcal{P}(\mathcal{S} )$ has overlapping marginals. It is also different from our W-DMR with overlapping marginals because the reference measures $\widehat{\mu}_{13}$ and $\widehat{\mu}_{23}$ may not have overlapping marginals. Unlike the uncertainty set for W-DMR, $\Sigma_{RDJ}(\delta)$ is empty when $\delta=0$.

Risk aggregation

Let $S_1, S_2$ be random variables representing individual risks defined on Polish spaces $\mathcal{S}_1, \mathcal{S}_{2}$, respectively. Let $\mu_1, \mu_2$ be probability measures of $S_1, S_2$. Let $\mathcal{V} = \mathcal{S}_1 \times \mathcal{S}_2$ and $g: \mathcal{V} \to \mathbb{R}$ be a risk aggregating function. Applying W-DMR with non-overlapping marginals to the risk aggregation function $g$, we can compute the worst aggregate risk when the joint measure of the individual risks varies in the uncertainty set $\Sigma_{\mathrm{D} }(\delta)$. This is different from the set-up in Eckstein2020, where the following robust risk aggregation problem is studied:

align*[align* omitted — 116 chars of source]

where

align*[align* omitted — 133 chars of source]

in which $\boldsymbol{K}_c$ is the optimal transport cost associated with a cost function $c: \mathcal{V} \times \mathcal{V} \to \mathbb{R}_{+}$. Since $\gamma \in \Sigma_{\Pi}(\delta_0)$ is a coupling of $(\mu_1, \mu_2)$, we have that $\Sigma_{\Pi}(\delta_0) \subset \Sigma_{\mathrm{D} }(0)$ and thus $\mathcal{I}_\Pi(\delta_0) \le \mathcal{I}_{\mathrm{D} }(0)$.

Strong Duality and Distributionally Robust Makarov Bounds

In this section, we establish strong duality for our W-DMR-MP and apply it to develop Wasserstein distributionally robust Makarov bounds.

Non-overlapping Marginals

For a measurable function $g: \mathcal{V} \rightarrow \mathbb{R}$ and $\lambda := (\lambda_1, \lambda_2) \in \mathbb{R}^2_{+}$, we define the function $g_\lambda: \mathcal{V}\rightarrow \mathbb{R} \cup \{\infty \}$ as \[\tag{2.1} g_{\lambda}(v) := \sup_{v^\prime \in \mathcal{V}} \varphi_\lambda(v, v^\prime), \] where $\varphi_\lambda: \mathcal{V} \times \mathcal{V} \rightarrow \mathbb{R} \cup \{ -\infty\}$ is given by \[ \varphi_\lambda(v, v^\prime) = g\left(s_1^\prime, s_2^\prime\right) - \lambda_1 c_1 \left(s_1, s_1^\prime\right) -\lambda_2 c_2 \left(s_2, s_2^\prime\right), \] with $v:= (s_1, s_2)$ and $v^\prime := (s_1^\prime, s^\prime_2)$. Similarly, define $g_{\lambda_1, 1}: \mathcal{V} \rightarrow \mathbb{R} \cup \{+\infty\}$ and $g_{\lambda_2, 2}: \mathcal{V} \rightarrow \mathbb{R} \cup \{+\infty\}$ as

align*[align* omitted — 261 chars of source]

The dual problem $\mathcal{J}_{\mathrm{D}}(\delta)$ corresponding to the primal problem $\mathcal{I}_{\mathrm{D}}(\delta)$ is defined as follows:

align[align omitted — 785 chars of source]
theoremSuppose that (ref) hold. Then, $\mathcal{I}_{\mathrm{D}}(\delta) = \mathcal{J}_{\mathrm{D}}(\delta)$ for all $\delta \in \mathbb{R}_{+}^2 \setminus \{0\}$.

Unlike the dual for W-DMR, the dual for W-DMR with non-overlapping marginals in (ref) involves a marginal problem with non-overlapping marginals $\mu_1, \mu_2$ due to the lack of knowledge on the dependence of the joint measure $\mu$. Computational algorithms developed for optimal transport can be used to solve the marginal problem, see Peyre2018. For empirical measures $\mu_1, \mu_2$, the marginal problem is a discrete optimal transport problem and there are efficient algorithms to compute it, see Peyre2018. For general measures $\mu_1, \mu_2$, strong duality may be employed in the numerical computation of the marginal problem. For instance, consider the case when $\delta > 0$. When $g_\lambda (v)$ is Borel measurable, several strong duality results are available, see e.g., villani2009optimal, villani2021topics. For a general function $g$ and cost functions $c_1,c_2$, $g_\lambda (v)$ is not guaranteed to be Borel measurable. However, for Polish spaces, the set $\{v \in \mathcal{V}:g_{\lambda}(v) \ge u\}$ is an analytic set for all $u \in \overline{\mathbb{R}}$ (and $g_{\lambda}$ is universally measurable), since $g$, $c_1$ and $c_2$ are Borel measurable (see blanchet2019quantifying and Bertsekas1978). This allows us to apply strong duality for the marginal problem in Kellerer:1984to restated in (ref) to the marginal problem involving $g_\lambda (v)$, see (ref) in (ref).

Without additional assumptions on the function $g$ and the cost functions, the dual $\mathcal{J}_{\mathrm{D}}(\delta)$ in (ref) for interior points $\delta \in \mathbb{R}_{++}^2$ and the dual for boundary points may not be the same. To illustrate, plugging in $\delta_2 = 0$ in the dual form for interior points in (ref), we obtain

align*[align* omitted — 214 chars of source]

It is different from the dual $\mathcal{J}_{\mathrm{D}}(\delta_1,0)$ for $\delta_1>0$, since

align*[align* omitted — 248 chars of source]

When the function $g$ and the cost functions satisfy assumptions in (ref), the dual $\mathcal{J}_{\mathrm{D}}(\delta)$ in (ref) for interior points $\delta \in \mathbb{R}_{++}^2$ and the dual for boundary points are the same so that

align*[align* omitted — 225 chars of source]

for all $\delta \in \mathbb{R}_{+}^2$.

remarkFor Polish spaces, (ref) generalizes the strong duality in zhang2022simple restated in Theorem 2.1. Our proof is based on that in zhang2022simple. However, due to the presence of two marginal measures in the uncertainty set $\Sigma_D(\delta)$, we need to verify the existence of a joint measure when some of its overlapping marginal measures are fixed, and we rely on existing results for a given consistent product marginal system studied in vorob1962consistent, kellerer1964verteilungsfunktionen, and shortt1983combinatorial, see (ref) for a detailed review.
remarkSimilar to Sinha2017 for W-DMR in marginal problems, we can define an alternative W-DMR through linear penalty terms, i.e., \begin{align*} \sup_{\gamma \in \mathcal{P}(\mathcal{V})} \left\{\int_{\mathcal{V}} g d \gamma - \lambda_1 \boldsymbol{K}_1(\mu_1, \gamma_1) - \lambda_2 \boldsymbol{K}_2(\mu_2, \gamma_2): \boldsymbol{K}_\ell(\mu_\ell, \gamma_\ell) < \infty for \ell = 1, 2\right\} \end{align*} with $\lambda_1, \lambda_2 \in \mathbb{R}_{++}$. The proof of (ref) implies that the dual form of this problem is $\sup_{\varpi \in \Pi(\mu_1, \mu_2)} \int g_\lambda d \varpi$ under the condition in (ref).

Overlapping Marginals

Let $\phi_\lambda : \mathcal{V} \times \mathcal{S}\rightarrow \mathbb{R} \cup \{- \infty \}$ be \[

aligned\phi_\lambda(v, s^{\prime}) & := f(s^{\prime})-\lambda_1 c_1 (s_1, s_1^{\prime})-\lambda_2 c_2(s_2, s_2^{\prime}),

\] where $v = (s_1, s_2)$, $s^\prime = (y^\prime_0,y^\prime_1,x^\prime)$, $s^\prime_\ell = (y^\prime_\ell, x^\prime)$ and $s_\ell=(y_\ell, x_\ell)$. Define the function $f_\lambda : \mathcal{V} \rightarrow \overline{\mathbb{R}}$ associated with $f$ as \[ f_\lambda(v) := \sup_{s^\prime \in \mathcal{S} } \phi_\lambda (v, s^{\prime}). \] Similarly, we define $f_{\lambda, 1}: \mathcal{V} \to \overline{\mathbb{R}}$ and $f_{\lambda, 2}: \mathcal{V} \to \overline{\mathbb{R}}$ as follows:

align*[align* omitted — 314 chars of source]

in which $s_1 = (y_1, x_1)$ and $s_2 = (y_2, x_2)$. The dual problem $\mathcal{J}(\delta)$ corresponding to the primal problem $\mathcal{I}(\delta)$ is defined as follows:

align[align omitted — 783 chars of source]
theoremSuppose that (ref) hold. Then, $\mathcal{I}(\delta) = \mathcal{J}(\delta)$ for all $\delta \in \mathbb{R}_{+}^2 \setminus \{0\}$.

An interesting feature of the dual for overlapping marginals is that it involves marginal problems with non-overlapping marginals, i.e., $\sup_{\varpi \in \Pi(\mu_{13}, \mu_{23})} \int_{\mathcal{V} }f_\lambda (v) d \varpi (v)$, although the uncertainty set in the primal problem involves overlapping marginals. Compared with the non-overlapping marginals case, overlapping marginals in the uncertainty set make the relevant consistent product marginal system in the verification of the existence of a joint measure more complicated, see the proof of (ref). Nonetheless, the non-overlapping marginals in the dual allow us to apply (ref) to the marginal problem involving $f_\lambda$, $f_{\lambda, 1}$ and $f_{\lambda, 2}$, see (ref) in (ref).

Under the assumptions in (ref), we have

align*[align* omitted — 217 chars of source]

for all $\delta \in \mathbb{R}_{+}^2$.

remarkSimilar to the non-overlapping case, we can define an alternative W-DMR with overlapping marginals through linear penalty terms, i.e., \begin{align*} \sup_{\gamma \in \mathcal{P}(\mathcal{S})} \left\{\int_{\mathcal{S}} g \, d \gamma - \lambda_1 \boldsymbol{K}_1(\mu_{13}, \gamma_{13}) - \lambda_2 \boldsymbol{K}_2(\mu_{23}, \gamma_{23}): \boldsymbol{K}_\ell(\mu_{\ell3}, \gamma_{\ell3}) < \infty for \ell = 1, 2\right\}, \end{align*} with $\lambda_1, \lambda_2 \in \mathbb{R}_{++}$. The proof of (ref) implies that the dual form of this problem is $\sup_{\varpi \in \Pi(\mu_{13}, \mu_{23})} \int_{\mathcal{V}} f_\lambda \, d \varpi$ under the conditions in (ref).

Wasserstein Distributionally Robust Makarov Bounds

Let $\mathcal{S}_1=\mathbb{R}$, $\mathcal{S}_2=\mathbb{R}$, $\mu_1\in \mathcal{P}(\mathcal{S}_1)$, and $\mu_2\in \mathcal{P}(\mathcal{S}_2)$. Further, let $Z=S_1+S_2$, where $S_1,S_2$ are random variables whose probability measures are $\mu_1, \mu_2$ respectively. For a given $z\in \mathbb{R}$, let $F_{Z}(z)=\mathbb{E}_o[g(S_1, S_2)]$, where $g(s_1, s_2)= \mathds{1}\left\{s_1 + s_2 \le z\right\}$.

Sharp bounds on the quantile function $F^{-1}_{Z}\left( \cdot\right) $ are established in Makarov1982) and referred to as the Makarov bounds. Inverting the Makarov bounds lead to sharp bounds on the distribution function $F _{Z}\left( z\right)$, see Ruschendorf_1982_RV_Maximum and Frank_1987_best_bounds. They are given by

align*[align* omitted — 311 chars of source]

Since the quantile bounds first established in Makarov1982) and the above distribution bounds are equivalent, we also refer to the latter as Makarov bounds. Makarov bounds have been successfully applied in distinct areas. For example, the upper bound on the quantile of $Z$ is known as the worst VaR of $Z$, see Embrechts2003UsingCopulaeBound, Embrechts_2005_WorstVaR; Makarov bounds are also used to study partial identification of distributional treatment effects when the treatment assignment mechanism identifies the marginal measures of the potential outcomes such as in (ref), see Fan_2009_PartialIdentification,Fan_2010_SharpBounds,Fan_2012_CI_quantiles, FAN_2009, Fan2017, Ridder_2007, and Firpo_2019.

Applying (ref), we extend Makarov bounds to allow for possible misspecification of the marginal measures and call the resulting bounds Wasserstein distributionally robust Makarov bounds.

corollary[Wasserstein distributionally robust Makarov bounds] Suppose that $g(s_1, s_2) = \mathds{1}(s_1 + s_2 \le z)$ and $c_{\ell}(s_\ell, s_\ell') = |s_\ell - s_\ell'|^2$ for $\ell = 1, 2$. For all $\delta\in\mathbb{R}_{+}^2$, \begin{align*} &\sup_{\gamma\in \Sigma_{\mathrm{D}}(\delta)} \mathbb{E}_{\gamma}[g(S_1, S_2)]\\ = &\inf_{\lambda\in\mathbb{R}_{+}^2} \Bigg(\langle\lambda, \delta\rangle +\sup_{ \varpi\in \Pi(\mu_1,\mu_2) } \Bigg[\int_{\{s_1 + s_2 > z\}} \left[1 - \frac{ \lambda_1\lambda_2 (s_1 + s_2 - z)^2}{\lambda_1 + \lambda_2} \right]^+ d\varpi(s_1, s_2)\\ & \quad + \mathbb{E}_{\varpi} \Big[\mathds{1}\left\{S_1 + S_2 \le z\right\} \Big]\Bigg]; \end{align*} \begin{align*} &\inf _{\gamma \in \Sigma_{\mathrm{D}}(\delta)} \mathbb{E}_\gamma\left[g\left(S_1, S_2\right)\right] \\ = & \sup_{\lambda\in\mathbb{R}_{+}^2} \Bigg[- \langle\lambda, \delta\rangle +\inf_{ \varpi\in \Pi(\mu_1,\mu_2) } \Bigg\{- \int_{\{s_1 + s_2 \leq z\} } \left[1 - \frac{ \lambda_1\lambda_2 (s_1 + s_2 - z)^2}{\lambda_1 + \lambda_2} \right]^+ d\varpi(s_1, s_2)\\ & \quad \qquad \quad + \mathbb{E}_{\varpi} \Big[\mathds{1}\left\{S_1 + S_2 \le z\right\} \Big] \Bigg]. \end{align*}

We note that $g_{\lambda}(v)$ is bounded and continuous in $v$, and convex in $\lambda$, and $\Pi(\mu_1, \mu_2)$ is compact. Applying Fan_1953_MinimaxThms's minimax theorem, we can interchange the order of $\inf$ and $\sup$ in the dual in the above corollary and get

align*[align* omitted — 448 chars of source]

This expression is very insightful, where the inner infimum term characterizes possible deviations of the true marginal measures from the reference measures.

Finiteness of the W-DMR-MP and Existence of Optimizers

In this section, we assume that all the reference measures belong to appropriate Wasserstein spaces and prove finitness of the W-DMR-MP and existence of an optimizer.

definition[Wasserstein space] The Wasserstein space of order $p \geq 1$ on a Polish space $\mathcal{X}$ with metric $ \boldsymbol{d}$ is defined as \[ \mathcal{P}_p( \mathcal{X} ) = \left\{ \mu \in \mathcal{P}( \mathcal{X} ): \int_{\mathcal{X} } \boldsymbol{d}(x_0, x)^p d \mu(x) < \infty \right\}, \] where $x_0 \in \mathcal{X}$ is arbitrary.
assumption\leavevmode \begin{assumpenum} • In the non-overlapping case, we assume that $\mu_{1} \in \mathcal{P}_{p_1}(\mathcal{S}_1)$ and $\mu_{2} \in \mathcal{P}_{p_2}(\mathcal{S}_2)$ for some $p_1 \ge 1$ and $p_2 \ge 1$; • In the overlapping case, we assume that $\mu_{13} \in \mathcal{P}_{p_1}(\mathcal{S}_1)$ and $\mu_{23} \in \mathcal{P}_{p_2}(\mathcal{S}_2)$ for some $p_1 \ge 1$ and $p_2 \ge 1$. \end{assumpenum}
assumptionThe cost function $c_\ell: \mathcal{S}_\ell \times \mathcal{S}_\ell \to \mathbb{R} \cup \{\infty\}$ is of the form $c_\ell(s_\ell, s_\ell^\prime ) = \boldsymbol{d}_{\mathcal{S}_\ell } (s_\ell, s_\ell^\prime )^{p_\ell}$, where $(\mathcal{S}_\ell, \boldsymbol{d}_{\mathcal{S}_\ell }) $ is a Polish space and $p_\ell \geq 1$ for $\ell=1, 2$.

Finiteness of the W-DMR-MP

For non-overlapping case, we establish the following result.

theoremSuppose that (ref) hold. Then for all $\delta\in \mathbb{R}_{++}^2$, $\mathcal{I}_{\mathrm{D} }(\delta) < \infty$ if and only if there exist $v^\star := (s_1^\star, s_2^\star) \in \mathcal{V}$ and a constant $M > 0$ such that for all $(s_1, s_2)\in \mathcal{V}$, \begin{equation} g(s_1, s_2) \leq M \left[ 1 + \boldsymbol{d}_{\mathcal{S}_1} (s_1^\star,s_1 )^{p_1}+ \boldsymbol{d}_{ \mathcal{S}_2 }(s_2^\star,s_2)^{p_2} \right], \end{equation} where $p_1$ and $p_2$ are defined in (ref).

The inequality in (ref) is a growth condition on the function $g$. It extends the growth condition in yue2022linear for W-DMR to our W-DMR with non-overlapping narginals.

For the overlapping case, the following result holds.

theoremSuppose that (ref) hold. Then for all $\delta\in \mathbb{R}_{++}^2$, $\mathcal{I}(\delta) < \infty$ if and only if there exist $ (s_1^\star, s_2^\star) \in \mathcal{S}_1 \times \mathcal{S}_2$ and a constant $M > 0$ such that \begin{equation} f(s) \leq M \left[ 1 +\boldsymbol{d}_{\mathcal{S}_1} (s_1^\star, s_1)^{p_1}+\boldsymbol{d}_{ \mathcal{S}_2 }(s_2^\star, s_2)^{p_2} \right], \end{equation} for all $ s \in \mathcal{S}$, where $s := (y_1, y_2, x), s_\ell := (y_\ell, x)$ and $s_\ell^\star := (y_\ell^\star, x^\star)$ for $\ell = 1, 2$, and $p_1$ and $p_2$ are defined in (ref).

The growth condition (ref) on the function $f$ extends the growth condition in yue2022linear for W-DMR. When \[ \boldsymbol{d}_{\mathcal{S}_\ell} ( (y_\ell, x), (y^{\prime}_\ell, x^\prime) ) = \boldsymbol{d}_{\mathcal{Y}_\ell} ( y_\ell,y^{\prime}_\ell ) + \boldsymbol{d}_{\mathcal{X}} ( x, x^{\prime} ), \] condition (ref) is satisfied if and only if there exist $s^\star := (y_1^\star, y_2^\star, x^\star)$ and a constant $M > 0$ such that \[ f(s) \leq M \left[ 1 + \boldsymbol{d}_{\mathcal{Y}_1} (y_1, y_1^\star )^{p_1}+ \boldsymbol{d}_{\mathcal{Y}_2} (y_2, y_2^\star )^{p_2}+ \boldsymbol{d}_{ \mathcal{X} }(x, x^\star )^{p_1 \wedge p_2} \right], \] for all $s = (y_1,y_2,x) \in \mathcal{S}$.

remarkThe conditions in (ref) are sufficient conditions for $\mathcal{I}_{\mathrm{D}}(\delta)$ and $\mathcal{I}(\delta)$ to be finite for all $\delta \in \mathbb{R}_{+}^2$ including boundary points because $\mathcal{I}_{\mathrm{D}}(\delta)$ and $\mathcal{I}(\delta)$ are non-decreasing.

Existence of Optimizers

definitionA metric space $(\mathcal{X}, \boldsymbol{d})$ is said to be proper if for any $r>0$ and $x_0 \in \mathcal{X}$, the closed ball $\overline{B}(x_0, r) := \{ x \in \mathcal{X}: \boldsymbol{d}(x,x_0) \leq r \}$ is compact.

Examples of proper metric spaces include finite dimensional Banach spaces and complete Riemannian manifolds, see yue2022linear.

assumption$(\mathcal{S}_1, \boldsymbol{d}_{\mathcal{S}_1})$ and $(\mathcal{S}_2, \boldsymbol{d}_{\mathcal{S}_2})$ are proper.

(ref) imply that $\Sigma_{\mathrm{D}}(\delta)$ and $\Sigma(\delta)$ are weakly compact, see (ref) in Appendix C. Given weak compactness of the uncertainty sets $\Sigma_{\mathrm{D}}(\delta)$ and $\Sigma(\delta)$, it is sufficient to show that the mapping: $\gamma \to \int g d\gamma$ is upper semi-continuous over $\gamma \in \Sigma_{\mathrm{D} }(\delta)$ for the non-overlapping case, and the mapping: $\gamma \to \int f d\gamma$ is upper semi-continuous over $\gamma \in \Sigma(\delta)$ for the overlapping case. In (ref) below, we provide conditions for $g$ and $f$ ensuring upper semi-continuity of each map and thus the existence of optimal solutions for $\mathcal{I}_{\mathrm{D}}(\delta)$ and $\mathcal{I}(\delta)$.

theoremSuppose that (ref) hold. Further, assume that $g$ is upper-semicontinuous, and there exist a constant $M>0$, $v^\star := (s_1^\star, s_2^\star) \in \mathcal{V}$ and $p_\ell^\prime \in (0, p_\ell)$ for $\ell=1,2$, such that \begin{align} g(v) \leq M\left[1+ \boldsymbol{d}_{\mathcal{S}_1} (s^\star_1, s_1)^{p_1^{\prime}} + \boldsymbol{d}_{\mathcal{S}_2} (s^\star_2,s_2)^{p_2^{\prime}} \right], \end{align} for all $v := (s_1, s_2) \in \mathcal{V}$. Then an optimal solution of (ref) exists for all $\delta \in \mathbb{R}_{+}^2$.
theoremSuppose that (ref) hold. Further, assume that $f$ is upper-semicontinuous, and there exist $ (s_1^\star, s_2^\star) \in \mathcal{S}_1 \times \mathcal{S}_2$, a constant $M>0$, $p_\ell^{\prime} \in (0, p_\ell)$ for $\ell = 1, 2$, such that \begin{align} f(s) \leq M \left[ 1 +\boldsymbol{d}_{\mathcal{S}_1} (s_1^\star, s_1)^{p_1^{\prime}}+\boldsymbol{d}_{ \mathcal{S}_2 }(s_2^\star,s_2 )^{p_2^{\prime}} \right], \end{align} for all $ s \in \mathcal{S}$ where $s := (y_1, y_2, x), s_\ell := (y_\ell, x)$ and $s_\ell^\star := (y_\ell^\star, x_\ell^\star)$ for $\ell = 1, 2$. Then an optimal solution of (ref) exists for all $\delta \in \mathbb{R}_{+}^2$.

Characterization of Identified Sets

In some applications, such as the partial identification of treatment effects introduced in (ref), the identified sets of $\theta_{\mathrm{D}o} := \mathbb{E}_{o}[g(S_1, S_2)]$ and $\theta_o := \mathbb{E}_{o}[f(S)]$ are of interest, where $S$ is a random variable whose probability measure belongs to $\Sigma(\delta)$, and $S_1$ and $S_2$ are random variables whose joint probability measure belongs to $\Sigma_{D}(\delta)$. They are:

align*[align* omitted — 297 chars of source]

By applying finiteness and existence results, we show below that under mild conditions, the identified sets $\Theta_{\mathrm{D}}(\delta)$ and $\Theta(\delta)$ are both closed intervals.

proposition\leavevmode \begin{enumerate}[label=(\roman*)] • Suppose (ref) hold. In addition, $g$ is continuous, and $|g|$ satisfies Condition (ref). Then, for $\delta \in \mathbb{R}_{+}^2$, we have \[ \Theta_{\mathrm{D}}(\delta)=\left[\min _{\gamma \in \Sigma_{\mathrm{D}}(\delta)} \int_{\mathcal{S}_1 \times \mathcal{S}_2} g \, d \gamma, \max _{\gamma \in \Sigma_{\mathrm{D}}(\delta)} \int_{\mathcal{S}_1 \times \mathcal{S}_2} g \, d \gamma \right], \] where both the lower and upper bounds are finite. • Suppose (ref) hold. In addition, $f$ is continuous and $|f|$ satisfies Condition (ref). Then for $\delta \in \mathbb{R}_{+}^2$, we have \[ \Theta(\delta)=\left[\min _{\gamma \in \Sigma(\delta)} \int_{\mathcal{S}} f \, d \gamma, \max_{\gamma \in \Sigma(\delta)} \int_{\mathcal{S}} f \, d \gamma \right], \] where both the lower and upper bounds are finite. \end{enumerate}

The strong duality in Section 3 can be used to evaluate the lower and upper bounds.

Continuity of the DMR-MP Functions

In this section, we establish continuity of the W-DMR-MP functions $\mathcal{I}_{\mathrm{D} }(\delta)$ and $\mathcal{I}(\delta)$ for all $\delta\in\mathbb{R}_+^2$ under similar conditions to those in zhang2022simple. Compared with zhang2022simple, our analysis is more involved, because the boundary in our case includes not only the origin $(0,0)$ but also $(\delta_1,0)$ and $(0,\delta_2)$ for all $\delta_1 > 0$ and $\delta_2 > 0$.

Non-overlapping Marginals

(ref) implies that under (ref), $\mathcal{I}_{\mathrm{D} }(\delta)$ is a concave function for $\delta \in \mathbb{R}_{+}^2$ and hence is continuous on $\mathbb{R}_{++}^2$. We provide the main assumption for the continuity of $\mathcal{I}_{\mathrm{D}}(\delta)$ on $\mathbb{R}_{+}^2$ in this subsection.

assumptionLet $\Psi: \mathbb{R}_+^2 \rightarrow \mathbb{R}_+$ be a continuous, non-decreasing, and concave function with $\Psi(0,0) =0$. Suppose the function $g: \mathcal{V} \rightarrow \mathbb{R}$ satisfies \begin{align} g(v) - g(v^\prime) \leq \Psi \left( c_1(s_1, s_1^\prime), c_2(s_2, s_2^\prime) \right), \end{align} for all $v=(s_1, s_2) \in \mathcal{V}$ and $v^\prime= (s^\prime_1, s^\prime_2) \in \mathcal{V}$.

The function $\Psi$ in (ref) plays the role of the modulus of continuity of $g$. To illustrate, consider the following example.

exampleSuppose (ref) holds, i.e., $c_\ell(s_\ell, s_\ell^\prime) = \boldsymbol{d}_{\mathcal{S}_\ell}(s_\ell, s_\ell^\prime)^{p_\ell}$ for some $p_\ell \geq 1$, $\ell=1,2.$ \begin{enumerate}[label=(\roman*)] • Define a product metric $\boldsymbol{d}_{\mathcal{V}}$ on $\mathcal{V} = \mathcal{S}_1 \times \mathcal{S}_2$ as \[ \boldsymbol{d}_{\mathcal{V}} ((s_1,s_2), (s_1^\prime, s_2^\prime)) = \boldsymbol{d}_{\mathcal{S}_1} (s_1,s_1^\prime) + \boldsymbol{d}_{\mathcal{S}_2} (s_2, s_2^\prime). \] Let $\Psi (x,y) = x^{1/p_1} + y^{1/p_2}$. Then, $\boldsymbol{d}_{\mathcal{V}} ((s_1,s_2), (s_1^\prime, s_2^\prime)) = \Psi \left( c_1 (s_{1}, s_{1}^{\prime}) , c_2(s_{2}, s_{2}^{\prime}) \right)$. On the metric space $(\mathcal{V}, \boldsymbol{d}_{\mathcal{V} })$, the function $g$ is continuous and has $\omega: x\mapsto x$ as modulus of continuity. Moreover, (ref) implies the growth condition in (ref). • Suppose $p_1=p_2$. Define a product metric $\boldsymbol{d}_{\mathcal{V}}$ on $\mathcal{V} = \mathcal{S}_1 \times \mathcal{S}_2$ as \[ \boldsymbol{d}_{\mathcal{V}} ((s_1,s_2), (s_1^\prime, s_2^\prime)) = \left[ \boldsymbol{d}_{\mathcal{S}_1} (s_1,s_1^\prime)^p + \boldsymbol{d}_{\mathcal{S}_2} (s_2, s_2^\prime)^p \right]^{1/p}. \] Let $\Psi (x,y) = (x+y)^{1/p}$. Then, $\boldsymbol{d}_{\mathcal{V}} ( (s_1,s_2), (s_1^\prime, s_2^\prime) ) = \Psi \left( c_1 (s_{1}, s_{1}^{\prime}) , c_2(s_{2}, s_{2}^{\prime}) \right)$. On the metric space $(\mathcal{V}, \boldsymbol{d}_{\mathcal{V} })$, the function $g$ is continuous and has $\omega: x\mapsto x$ as modulus of continuity. (ref) also implies the growth condition in (ref). • Suppose $p_1 \neq p_2$. Define a product metric $\boldsymbol{d}_{\mathcal{V}}$ on $\mathcal{V} = \mathcal{S}_1 \times \mathcal{S}_2$ as \[ \boldsymbol{d}_{\mathcal{V}} ((s_1,s_2), (s_1^\prime, s_2^\prime)) = \boldsymbol{d}_{\mathcal{S}_1} (s_1,s_1^\prime) \vee \boldsymbol{d}_{\mathcal{S}_2} (s_2, s_2^\prime). \] Then, (ref) implies \[ g(v)-g (v^{\prime}) \leq \Psi\left( \boldsymbol{d}_{\mathcal{V}} ( v, v^\prime) , \boldsymbol{d}_{\mathcal{V}} (v, v^\prime) \right) = \omega( \boldsymbol{d}_{\mathcal{V}} (v, v^\prime) ) . \] where $\omega: x \mapsto \Psi(x,x)$ is a concave function. On the metric space $(\mathcal{V},\boldsymbol{d}_{\mathcal{V}})$, the function $g$ is continuous and has $\omega: x \mapsto \Psi(x,x)$ as modulus of continuity. \end{enumerate}
theoremSuppose (ref) hold and $\mathcal{I}_{\mathrm{D}}(\delta) < \infty$ for some $\delta > 0$. Then, the function $\mathcal{I}_{\mathrm{D} }(\delta)$ is continuous on $\mathbb{R}_{+}^2$.

Two implications follow. First, under (ref) and (ref), \[\mathcal{I}_{\mathrm{D} }(0) = \sup_{\gamma \in \Pi(\mu_{1}, \mu_{2})} \int_{\mathcal{V} } g\, d\gamma. \] Continuity facilitates sensitivity analysis as $\delta$ approaches zero; Second, under the assumptions in (ref), we have

align*[align* omitted — 225 chars of source]

for all $\delta \in \mathbb{R}_{+}^2$. As a result, the dual $\mathcal{J}_{\mathrm{D} }(\delta)$ in (ref) is continuous for all $\delta\in\mathbb{R}_+^2$.

Overlapping Marginals

(ref) implies that under (ref), $\mathcal{I}(\delta)$ is a concave function for $\delta \in \mathbb{R}_{+}^2$ and hence is continuous on $\mathbb{R}_{++}^2$. We provide the main assumption for the continuity of $\mathcal{I}(\delta)$ on $\mathbb{R}_{+}^2$ below.

To simplify the technical analysis, we maintain (ref) in this section. Since the metrics in $\mathcal{Y}_1$ and $\mathcal{Y}_2$ are not specified, we introduce an auxiliary function $\rho_\ell$ from $\mathcal{Y}_\ell \times \mathcal{Y}_\ell$ to $\mathbb{R}_+$ induced by the cost function $c_\ell$, $\ell = 1, 2$.

assumptionFor $\ell=1,2$, there exists a function $\rho_\ell$ from $\mathcal{Y}_\ell \times \mathcal{Y}_\ell$ to $\mathbb{R}_+$ such that \begin{assumpenum} • $\rho_\ell$ is symmetric, i.e., $\rho_\ell(y_\ell, y_\ell^\prime) =\rho_\ell(y_\ell^\prime, y_\ell)$ for all $y_\ell, y_\ell^\prime \in \mathcal{Y}_\ell$; • there is $q_\ell\in [1, p_\ell]$ such that $\rho_\ell(y_\ell, y_\ell^\prime) \leq \boldsymbol{d}_{\mathcal{S}_\ell } (s_\ell, s_\ell^\prime)^{q_\ell}$ for all $s_\ell \equiv (y_\ell, x) \in \mathcal{S}_\ell$ and $s^\prime_\ell \equiv (y^\prime_\ell, x^\prime) \in \mathcal{S}_\ell$; • there is a constant $N > 0$ such that $\rho_\ell(y_\ell, y_\ell^\prime) \leq N \left[ \rho_\ell(y_\ell, y_\ell^\star ) + \rho_\ell( y_\ell^\star , y_\ell^\prime) \right]$ for all $y_\ell, y_\ell^\prime, y_\ell^\star \in \mathcal{Y}_\ell$. \end{assumpenum}

We now introduce the main assumption on $f$.

assumptionFor $\ell = 1,2$, let $\Psi_{\ell} : \mathbb{R}^2_+ \rightarrow \mathbb{R}_+$ be continuous, non-decreasing, and concave satisfying $\Psi_\ell(0,0) = 0$. Suppose for all $s = (y_1, y_2, x)$ and $s^\prime = (y^\prime_1, y^\prime_2, x^\prime)$, it holds that \[ f(y_1, y_2, x) - f(y_1^\prime, y_2^\prime, x^\prime) \leq \Psi_1 \left ( c_1(s_1, s_1^\prime), \rho_2(y_2, y_2^\prime) \right), \] and \[ f(y_1, y_2, x) - f(y_1^\prime, y_2^\prime, x^\prime) \leq \Psi_2 \left ( \rho_1(y_1, y_1^\prime), c_2(s_2, s_2^\prime) \right). \]

Like (ref), (ref) depends on the cost functions $c_1,c_2$. It also depends on the auxiliary functions $\rho_1,\rho_2$. The functions $\Psi_1, \Psi_2$ play the role of the modulus of continuity.

example[$p_j$-product metric] Let $(\mathcal{Y}_1,\boldsymbol{d}_{\mathcal{Y}_1} ), (\mathcal{Y}_2, \boldsymbol{d}_{\mathcal{Y}_2})$, and $(\mathcal{X}, \boldsymbol{d}_{\mathcal{X} })$ be Polish (metric) spaces. For $p_\ell \geq 1$, define the $p_\ell$-product metric on $\mathcal{S}_\ell$ as \[\boldsymbol{d}_{\mathcal{S}_\ell} (s_\ell, s_\ell^\prime) = \left[ \boldsymbol{d}_{\mathcal{Y}_\ell}(y_\ell , y_\ell^\prime)^{p_\ell} + \boldsymbol{d}_{\mathcal{X} }(x, x^\prime)^{p_\ell} \right]^{1/p_\ell}. \] Let \[ \rho_\ell( y_\ell , y_\ell^\prime ):= \inf_{ x_\ell, x_\ell^\prime \in \mathcal{X} } \boldsymbol{d}_{\mathcal{S}_\ell} \left( (y_\ell, x_\ell) , (y_\ell^{\prime}, x_\ell^{\prime}) \right)^{p_\ell}. \] It is easy to show that $ \rho_\ell( y_\ell , y_\ell^\prime ) =\boldsymbol{d}_{\mathcal{Y}_\ell}(y_\ell, y_\ell^\prime)^{p_\ell}$ and (ref) is satisfied with $N=2^{p_\ell}$. Moreover, (ref) reduces to \[ \begin{aligned} f(y_1, y_2, x) - f(y_1^\prime, y_2^\prime, x^\prime) &\leq \Psi_1 \left( \boldsymbol{d}_{\mathcal{S}_1}\left(s_1, s_1^{\prime}\right)^{p_1}, \boldsymbol{d}_{\mathcal{Y}_2}\left(y_2, y_2^{\prime}\right)^{p_2} \right) \quad \text{and } \\ f(y_1, y_2, x) - f(y_1^\prime, y_2^\prime, x^\prime) & \leq \Psi_2 \left(\boldsymbol{d}_{\mathcal{Y}_1} \left(y_1, y_1^{\prime} \right)^{p_1}, \boldsymbol{d}_{\mathcal{S}_2}\left(s_2, s_2^{\prime}\right)^{p_2} \right). \end{aligned} \] When $p_1 = p_2 = p$, (ref) may be reduced to a simpler form. To see this, define two functions $\psi_1$ and $\psi_2$ from $\mathbb{R}^3$ to $\mathbb{R}^2$ as $\psi_1: (z_1,z_2, z) \mapsto ( z_1 + z , z_2)$ and $\psi_2: (z_1,z_2, z) \mapsto ( z_1, z_2 + z )$. We can see that \[ \Psi_1 \left( \boldsymbol{d}_{\mathcal{S}_1}(s_1, s_1^\prime)^p, \rho_2(y_1, y_1^\prime)^{p} \right) = \Psi_1 \circ \psi_1 \left( \boldsymbol{d}_{\mathcal{Y}_1}(y_1, y_1^\prime)^{p} , \boldsymbol{d}_{\mathcal{Y}_2} (y_2, y_2^\prime )^p , \boldsymbol{d}_{\mathcal{X}}(x, x^\prime)^p \right), \] \[ \Psi_2 \left( \rho_1(y_1, y_1^\prime)^{p} , \boldsymbol{d}_{\mathcal{S}_2}(s_2, s_2^\prime)^p \right) = \Psi_2 \circ \psi_2 \left( \boldsymbol{d}_{\mathcal{Y}_1}(y_1, y_1^\prime)^{p} , \boldsymbol{d}_{\mathcal{Y}_2}(y_2, y_2^\prime )^p , \boldsymbol{d}_{\mathcal{X}}(x, x^\prime)^p \right). \] Since $\psi_j$ is linear, $\Phi_j = \Psi_j \circ \psi_j$ is still continuous, non-decreasing and concave. (ref) is reduced to the following condition: \[ f(y_1, y_2, x) - f(y_1^\prime, y_2^\prime, x^\prime) \leq \Phi_j \left( \boldsymbol{d}_{\mathcal{Y}_1}(y_1, y_1^\prime)^{p} , \boldsymbol{d}_{\mathcal{Y}_2} (y_2, y_2^\prime )^p , \boldsymbol{d}_{\mathcal{X}}(x, x^\prime)^p \right) \] for all $(y_1,y_2, x) \in \mathcal{S}$ and $(y_1^\prime, y_2^\prime, x^\prime) \in \mathcal{S}$.
theoremSuppose (ref) hold, and $\mathcal{I}(\delta) < \infty$ for some $\delta > 0$. Then the function $\mathcal{I}(\delta)$ is continuous on $\mathbb{R}_{+}^2$.

Like the non-overlapping case, two implications follow. First, under (ref) and (ref), \[ \mathcal{I}(0)= \sup_{\gamma \in \mathcal{F}(\mu_{13}, \mu_{23}) } \int_\mathcal{S} f\, d \gamma. \] Continuity facilitates sensitivity analysis as $\delta$ approaches zero; Second, under the assumptions in (ref), we have

align*[align* omitted — 217 chars of source]

for all $\delta \in \mathbb{R}_{+}^2$. As a result, the dual $\mathcal{J}(\delta)$ in (ref) is continuous for all $\delta\in\mathbb{R}_+^2$.

Motivating Examples Revisited

In this section, we apply the results in Sections 3-5 to the examples introduced in Section 2.

Partial Identification of Treatment Effects

In addition to characterizing $\Theta(\delta)$ introduced in Section 2, we also study the identified set for $\theta_{Do}= \mathbb{E}_o[f(Y_1,Y_2)]$ without using the covariate information: \[ \Theta_{\mathrm{D}}(\delta) := \left\{ \int_{\mathcal{Y}_1 \times \mathcal{Y}_2 } f(y_1, y_2) \, d \gamma(y_1 ,y_2) : \gamma \in \Sigma_{\mathrm{D} }(\delta) \right\}, \] where

align*[align* omitted — 244 chars of source]

in which $\boldsymbol{K}_{Y_1}$ and $\boldsymbol{K}_{Y_2}$ are the optimal transport costs associated with cost functions $c_{Y_1}$ and $c_{Y_2}$, respectively.

Characterization of the Identified Sets

When $f$ is continuous and conditions in (ref) are satisfied, the identified sets $\Theta_{\mathrm{D}}(\delta)$ and $\Theta(\delta)$ are both closed intervals with upper limits given by W-DMR for non-overlapping and overlapping marginals respectively. This allows us to apply our duality results in Section 3 to evaluate and compare $\Theta_{\mathrm{D}}(\delta)$ and $\Theta(\delta)$.

Let $\mathcal{I}_{\mathrm{D} }(\delta)$ and $\mathcal{I}(\delta)$ denote the upper bounds of $\Theta_{\mathrm{D}}(\delta)$ and $\Theta(\delta)$, respectively, where \[ \mathcal{I}_{\mathrm{D} }(\delta) = \sup_{\gamma \in \Sigma_{\mathrm{D} }(\delta)} \int_{\mathcal{Y}_1 \times \mathcal{Y}_2} f(y_1, y_2) \, d \gamma(y_1, y_2) \text{ and } \mathcal{I}(\delta) = \sup_{\gamma \in \Sigma(\delta)} \int_{\mathcal{S}} f(y_1, y_2) \, d \gamma(y_1, y_2, x). \] (ref) establishes robust versions of existing results on the identified sets of treatment effects under (ref), see Fan2017. Sensitivity to deviations from (ref) can be examined via $\Theta_{\mathrm{D}}(\delta)$ and $\Theta(\delta)$ by varying $\delta$. For example, when $f$ satisfies assumptions in (ref), $\mathcal{I}(\delta) $ and $\mathcal{I}_{\mathrm{D} }(\delta)$ are continuous on $\mathbb{R}_+^2$. As a result, \[\lim_{\delta\rightarrow 0 }\mathcal{I}(\delta)=\mathcal{I}(0) \quad \text{and} \quad \lim_{\delta\rightarrow 0 }\mathcal{I}_{\mathrm{D} }(\delta)=\mathcal{I}_{\mathrm{D} }(0). \]

For a general function $f$, the lower and upper limits of the identified sets $\Theta_{\mathrm{D}}(\delta)$ and $\Theta(\delta)$ need to be computed numerically. When $f$ is additively separable, we show that duality results in Section 3 simplify the evaluation of $\Theta_D(\delta)$ and $\Theta(\delta)$. Since the lower bounds of $\Theta_D(\delta)$ and $\Theta(\delta)$ can be computed in a similar way by applying duality to $-f(y_1, y_2)$, we omit details for the lower bounds.

assumptionLet $f: (y_1,y_2, x) \mapsto f_1(y_1) + f_2(y_2)$ from $\mathcal{S}$ to $\mathbb{R}$, where $f_\ell \in L^1(\mu_{\ell 3})$ for $\ell=1,2$.

To avoid tedious notation, we also treat $f$ as a function from $\mathcal{Y}_1 \times \mathcal{Y}_2$ to $\mathbb{R}$. Under (ref), it is easy to show that

align*[align* omitted — 509 chars of source]

where $(f_\ell)_{\lambda_\ell}: \mathcal{Y}_\ell \rightarrow \mathbb{R}$ is given by \[ (f_\ell)_{\lambda_\ell}(y_\ell) = \sup_{ y_\ell^\prime \in \mathcal{Y}_\ell } \left\{ f_\ell(y_\ell^{\prime})-\lambda_\ell c_{Y_\ell}(y_{\ell}, y_{\ell}')\right\}. \] That is, when $f$ is an additively separable function, the W-DMR for non-overlapping marginals is the sum of two W-DMRs associated with the marginals regardless of the cost functions.

Depending on the cost functions, the W-DMR for overlapping marginals may be different from the sum of two W-DMRs associated with the marginals.

definition[Ref. {Chen_2022}] We say that a function $f: \mathcal{X} \times \mathcal{Y} \to \mathbb{R}$ is separable if each $x$ and $y$ can be optimized regardless of the other variable. In other words, \begin{align*} \operatorname{\mathop{\mathrm{argmin}}}_{x, y} f(x, y) = \left (\operatorname{\mathop{\mathrm{argmin}}}_{x \in \mathcal{X}} f(x, y'), \operatorname{\mathop{\mathrm{argmin}}}_{y\in \mathcal{Y}} f(x', y) \right ) \end{align*} for any $x' \in \mathcal{X}$ and $y' \in \mathcal{Y}$.
assumptionFor $\ell = 1, 2$, the cost function $c_\ell((y_\ell, x_\ell), (y_\ell', x_\ell'))$ is separable with respect to $(y_\ell, y_\ell')$ and $(x_\ell, x_\ell')$.
exampleLet $a_\ell: \mathcal{Y}_{\ell} \times \mathcal{Y}_{\ell} \to \mathbb{R}_{+} \cup \{\infty\}$ and $b_\ell: \mathcal{X} \times \mathcal{Y} \to \mathbb{R}_{+} \cup \{\infty\}$ satisfy (ref). Let $s = (y, x)$ and $s' = (y', x')$. Then $c(s, s') = a(y, y') + b(x, x')$ is separable with respect to $(x, x')$ and $(y, y')$. Also, both $c(s, s')=(a(y, y') +1)(b(x, x')+1)-1$ and $c(s, s') = \left[a(y, y')^p + b(x, x')^p\right]^{1/p}$ for $p \ge 1$ are separable with respect to $(x, x')$ and $(y, y')$ even though they are not additively separable.
propositionFor $\ell=1,2$, let $c_\ell: (\mathcal{Y}_\ell \times \mathcal{X}) \times (\mathcal{Y}_\ell \times \mathcal{X}) \rightarrow \mathbb{R}_+$ denote the cost function for $\Theta(\delta)$. Suppose that $c_\ell$ satisfies (ref) and the marginal measure of $\mu_{\ell 3}$ on $\mathcal{Y}_\ell$ coincides with $\mu_\ell$, i.e., $\mu_{\ell,3} = \mathrm{Law}(Y_\ell, X)$ with $\mu_\ell = \mathrm{Law}(Y_\ell)$. Under (ref), one has $\mathcal{I}(\delta) = \mathcal{I}_{\mathrm{D} }(\delta)$, where $\mathcal{I}_{\mathrm{D} }(\delta)$ is based on the cost function $c_{Y_\ell}$ on $\mathcal{Y}_\ell \times \mathcal{Y}_\ell$ given by \[ c_{Y_{\ell}}\left(y_{\ell}, y_{\ell}^{\prime}\right)=\inf _{x_{\ell}, x_{\ell}^{\prime} \in \mathcal{X}} c_{\ell}\left(\left(y_{\ell}, x_{\ell}\right),\left(y_{\ell}^{\prime}, x_{\ell}^{\prime}\right)\right). \]

It is easy to verify that $c_{Y_\ell}(y_\ell, y_\ell^\prime) = 0$ if and only if $y_\ell = y_\ell^\prime$.

This proposition implies that for separable cost functions, the W-DMR for overlapping marginals equals the W-DMR for non-overlapping marginals with cost function $c_{Y_\ell}(y_\ell, y_\ell^\prime)$. As a result, the covariate information does not help shrink the identified set.

Average Treatment Effect

Suppose $f(y_1, y_2) = y_2 - y_1$ and $c_{\ell}((y, x), (y_{\ell}, x_{\ell})) = |y - y'|^{2} + \| x_\ell - x_\ell' \|^2$ for $\ell = 1, 2$. Let $\tau_{ATE}=\mathbb{E}[Y_2 - Y_1]$. Then (ref) implies that the upper bound on $\tau_{ATE}$ is given by

align*[align* omitted — 153 chars of source]

In the rest of this section, we demonstrate that when (ref) is violated, the W-DMR for overlapping marginals may be smaller than the W-DMR for non-overlapping marginals and, as a result, $\Theta(\delta)$ is a proper subset of $\Theta_D(\delta)$.

Consider the squared Mahalanobis distance with respect to a positive definite matrix. That is,

align*[align* omitted — 100 chars of source]

where $V_\ell =

pmatrix[pmatrix omitted — 81 chars of source]

$ is a positive definite matrix. It is easy to show that

align*[align* omitted — 187 chars of source]

where $s_\ell = (y_\ell, x_\ell)$ and $s_{\ell}' = (y_\ell', x_\ell')$.

propositionLet $\mathcal{I}$ be the primal of the overlapping W-DMR problem under \begin{align*} c_{\ell} (s_\ell, s_\ell') = (s_\ell - s_\ell')^{\top} V_\ell^{-1} (s_\ell - s_\ell'). \end{align*} Let $\mathcal{I}_{\mathrm{D}}$ be the primal of the non-overlapping W-DMR problem under $ c_{Y_\ell}(y_\ell ,y_\ell') $. Assume that $\mathbb{E} \|X\|_2^2 < \infty$, $\mathbb{E}|Y_1|^2 < \infty$, and $\mathbb{E}|Y_2|^2 < \infty$. Then, $ \mathcal{I}(\delta) \le \mathcal{I}_{\mathrm{D}}(\delta)$ for all $\delta > 0$.
propositionSuppose that all the conditions in (ref) hold. Then, \begin{propenum} • for all $\delta \in \mathbb{R}^2_{+}$, \begin{align*} \mathcal{I}_{\mathrm{D}}(\delta) & = \mathbb{E}[Y_2] - \mathbb{E}[Y_1] + V_{1, YY}^{1/2} \; \delta_1^{1/2} + V_{2, YY}^{1/2}\; \delta_2^{1/2}, \\ \mathcal{I}(\delta) & = \mathbb{E}[Y_2] - \mathbb{E}[Y_1] + \inf_{\lambda \in \mathbb{R}_{++}^2} \Bigg\{ \lambda_1 \delta_1 + \lambda_2 \delta_2 + \frac{1}{4\lambda_{1}} \left(V_{1} / V_{1, XX} \right) + \frac{1}{4\lambda_{2}} \left(V_{2} / V_{2, XX} \right) \\ & \qquad \qquad \qquad \qquad \qquad \qquad \quad + \frac{1}{4} V_o^{\top} \left(\lambda_1 V_{1, XX}^{-1} + \lambda_2 V_{2, XX}^{-1}\right)^{-1} V_o \Bigg\}, \end{align*} where $V_{\ell} / V_{\ell, XX} : = V_{\ell, YY} - V_{\ell, YX} V_{\ell, XX}^{-1} V_{\ell, XY}$ is the Schur complement of $V_{\ell, XX}$ in $V_{\ell}$ for $\ell = 1, 2$, and $V_o = V_{2, XX}^{-1} V_{2, XY} - V_{1, XX}^{-1} V_{0, XY}$; • $\mathcal{I}_{\mathrm{D}}(\delta) = \mathcal{I}(\delta)$ for all $\delta \in \mathbb{R}_{+}^2$ if and only if $V_{1, XY} = V_{2, XY} = 0$; • $\mathcal{I}_{\mathrm{D}}(\delta)$ and $\mathcal{I}(\delta)$ are continuous on $\mathbb{R}^2_+$. \end{propenum}

(ref) and (ref) imply that for non-separable Mahalanobis cost functions, the information in covariates may help shrink the identified set since $\mathcal{I}_{\mathrm{D}}(\delta) < \mathcal{I}(\delta)$ for some $\delta$ under mild conditions. (ref) also implies that (i) $\mathcal{I}(0) = \mathcal{I}_{\mathrm{D}}(0) = \mathbb{E}[Y_2] - \mathbb{E}[Y_1]$; (ii) $\mathcal{I}(\delta_1, 0) = \mathcal{I}_{\mathrm{D}}(\delta_1, 0)$ and $\mathcal{I}(0, \delta_2) = \mathcal{I}_{\mathrm{D}}(0, \delta_2)$ for all $\delta_1 \ge 0$ and $\delta_2 \ge 0$.

Comparison of Robust Welfare Functions

Recall that

align*[align* omitted — 225 chars of source]

where

align*[align* omitted — 317 chars of source]

Consider the following cost function $c_{\ell}$ for $\ell = 1, 2$:

align*[align* omitted — 82 chars of source]

where $s_\ell = (y_\ell, x_\ell)$, $s_\ell' = (y_\ell', x_\ell')$, and $ c_{Y_1}(y_1, y_1')$ and $c_{Y_2}(y_2, y_2')$ are cost functions for $Y_1$ and $Y_2$, respectively, and $b(x, x')$ is some function on the space $\mathcal{X}$ satisfying (ref). When $b(x, x') = \infty \mathds{1}\{x \ne x'\}$, $\mathbb{P}(X = X') = 1$ for any probability measure in uncertainty set.

adjaho2022 establishes strong duality for $\mathrm{RW}_0(d)$ under several cost functions. For comparison purposes, we restate the following Proposition in adjaho2022 which allows distributional shifts in covariate $X$.

proposition(Proposition 4.1 in adjaho2022) Suppose $Y_1$ and $Y_2$ are unbounded and $\mathbb{E}\left\| X \right\|_2^2$ is finite. Let the cost function $c: \mathcal{S} \times \mathcal{S} \rightarrow \mathbb{R}_{+}$ be given by \begin{align*} c(s, s') = |y_1 - y_1'| + |y_2 - y_2'| + \| x' - x \|_2, \end{align*} for $s = (y_1, y_2, x)$ and $s' = (y_1', y_2', x')$. Then \begin{align*} \mathrm{RW}_0(d) = \sup_{\eta \ge 1} \left\{ \mathbb{E}_{\mu}\left[\max\{Y_2 + \eta h_1(X), Y_1 + \eta h_0(X) \}\right] - \eta \delta_0 \right\}, \quad where \end{align*} \begin{align*} h_0(x) = \inf_{u \in \mathcal{X}: d(u) = 0} \| x - u \|_2 and h_1(x) = \inf_{u \in \mathcal{X}: d(u) = 1} \| x - u \|_2. \end{align*}

This proposition implies that $\mathrm{RW}_0(d)$ depends on the choice of the reference measure $\mu$. Since only the marginals $\mu_{13}$ and $\mu_{23}$ are identified under (ref), adjaho2022 suggest three possible choices for $\mu$ by imposing specific dependence structures on $\mu$:

itemize$Y_1$ and $Y_2$ are perfectly positively dependent conditional on $X = x$; • $Y_1$ and $Y_2$ are conditionally independent given $X = x$; • $Y_1$ and $Y_2$ are perfectly negatively dependent conditional on $X = x$.

Section 4.3.1 in adjaho2022 shows that their robust welfare function $\mathrm{RW}_0(d)$ is minimized when $Y_1$ and $Y_2$ are perfectly negatively dependent conditional on $X = x$.

The following proposition evaluates $\mathrm{RW}(d)$ via the duality result in Section 3 and compares it with $\mathrm{RW}_0(d)$.

propositionConsider \begin{align*} c_{\ell}(s_\ell, s_\ell') & = | y_\ell - y_\ell' | + \| x_\ell - x_\ell' \|_2. \end{align*} Assume that $Y$ is unbounded and $\mathbb{E}|Y_1|$, $\mathbb{E}|Y_2|$, and $\mathbb{E}\| X \|_2^2$ are finite. Then, \begin{propenum} • the robust welfare function $\mathrm{RW}(d)$ based on $\Sigma(\delta)$ has the following dual reformulation: \begin{align*} \mathrm{RW}(d) = \sup_{\lambda \ge 1} \left[ \inf_{\pi \in \Pi(\mu_{13}, \mu_{23})} \int_{\mathcal{V}} \min\{ y_2 + \varphi_{\lambda, 1} (x_1, x_2), y_1 + \varphi_{\lambda, 0} (x_1, x_2) \} d \pi(v) - \langle \lambda, \delta \rangle \right], \end{align*} where $v = (y_1, x_1, y_2, x_2)$, and \begin{align*} \varphi_{\lambda, 0} (x_1, x_2) & = \min_{x': d(x') = 0} \bigg( \lambda_1 \| x_1 - x' \|_2 + \lambda_2 \| x_2 - x' \|_2 \bigg), \\ \varphi_{\lambda, 1} (x_1, x_2) & = \min_{x': d(x') = 1} \bigg( \lambda_1 \| x_1 - x' \|_2 + \lambda_2 \| x_2 - x' \|_2 \bigg); \end{align*} • When $\delta_0 = \delta_1 = \delta_2$, $\mathrm{RW}(d)\le \mathrm{RW}_0^*(d)$, where $\mathrm{RW}_0^*(d)$ is the robust welfare function $\mathrm{RW}_0(d)$ based on the reference measure $\pi^*=\int\max\{\mu_{1|3}+\mu_{2|3}-1,0 \}d\mu_3$. \end{propenum}

Part (ii) of the above proposition implies that $\mathrm{RW}(d) \le \mathrm{RW}_0(d)$ for any reference measure $\mu \in \mathcal{F}(\mu_{13}, \mu_{23})$.

W-DRO for Logit Model Under Data Combination

We revisit the logit model in (ref) and make the following assumption.

assumption(i) Let $(Y_1, Y_2, X)$ follow some unknown measure $\mu$. Let $D$ denote a binary random variable independent of $(Y_1, Y_2, X)$ such that we observe $(Y_1, X)$ when $D = 0$, and $(Y_2, X)$ when $D = 1$; (ii) Let $\{Y_{1i}, X_{1i}\}_{i=1}^{n_1}$ be the data set from $(Y_1, X)$, and $\{Y_{2i}, X_{2i}\}_{i=1}^{n_2}$ be the data set from $(Y_2, X)$.

Under this assumption, $X|D=1$ has the same distribution as $X|D=0$ and the empirical distributions of the two data sets are consistent estimators of the population reference measures for $(Y_1, X)$ and $(Y_2, X)$.

Suppose (ref) hold. Then (ref) implies that for all $\delta > 0$,

align*[align* omitted — 229 chars of source]

where

align*[align* omitted — 194 chars of source]

with $v = (y_1, x_1, y_2, x_2)$.

Let $\widehat{\mu}_{13}$ and $\widehat{\mu}_{23}$ denote the empirical measures based on the two data sets. The dual form of $\mathcal{I}(\delta)$ can be estimated by

align*[align* omitted — 283 chars of source]

A direct consequence of Kellerer:1984to is that

align*[align* omitted — 483 chars of source]

where the last expression reduces to the dual in awasthi2022distributionally for the cost functions \[ c_1((y_1, x), (y_1', x')) = \| x - x' \|_{p} + \kappa_1 | y_1 - y_1'| \quad \text{and} \] \[ c_2((y_2, x), (y_2, x')) = \| x - x' \|_{p} + \kappa_2 \| y_2 - y_2' \|_{p'}. \]

W-DMR with Multi-marginals

Sections 2-6 present a detailed study of W-DMR with two marginals. In this section, we briefly introduce W-DMR with more than two marginals or multi-marginals and discuss strong duality for non-overlapping and overlapping marginals.\footnote{For multi-marginals, the collection of given marginals can be more complicated than the non-overlapping and overlapping marginals (see ruschendorf1991bounds, Embrechts_2010 and Doan2015), we leave a complete treatment of the W-DMR with multi-marginals in future work.} Applications include extension of risk aggregation in (ref) to any finite number of individual risks and robust treatment choice in (ref) to multi-valued treatment.

Non-overlapping Marginals

Let $\mathcal{V}:=\prod_{\ell \in [L]} \mathcal{S}_\ell$ for Polish spaces $\mathcal{S}_\ell$ for $\ell \in [L]$, and $\mu_\ell$ be a probability measure on $(\mathcal{S}_{\ell}, \mathcal{B}_{\mathcal{S}_\ell})$. Let $\Pi(\mu_1, \dotsc, \mu_L)$ be the set of all possible couplings of $\mu_1, \dotsc, \mu_L$. Further, let $g:\mathcal{V}\rightarrow \mathbb{R}$ be a measurable function satisfying the following assumption.

assumptionThe function $g:\mathcal{V}\rightarrow \mathbb{R}$ is a measurable function such that $\int_{\mathcal{V}} g d\gamma_0> -\infty$ for some $\gamma_0 \in \Pi(\mu_1, \dotsc, \mu_L) \subset \mathcal{P}(\mathcal{V})$.

For any $\gamma \in \mathcal{P}(\mathcal{V})$, let $\gamma_{\ell}$ denote the projection of $\gamma$ on $\mathcal{S}_{\ell}$ for $\ell \in [L]$. The W-DMR with non-overlapping multi-marginals is formulated as

align*[align* omitted — 128 chars of source]

where $\Sigma_{\mathrm{D} }(\delta)$ is the uncertainty set defined as

align*[align* omitted — 170 chars of source]

in which $\delta = (\delta_1, \dotsc, \delta_L) \in \mathbb{R}_{+}^L$ is the radius of the uncertainty set.

For a generic vector $v \in \mathbb{R}^L$ and $A \subset [L]$, we write $v_A = (v_{A, 1}, \dotsc, v_{A, L}) \in \mathbb{R}^{L}$ as follows:

align*[align* omitted — 156 chars of source]

We also define $\tilde{c}_\ell: \mathcal{S}_{\ell} \times \mathcal{S}_{\ell} \to \mathbb{R}_{+} \cup \{\infty\}$ as:

align*[align* omitted — 285 chars of source]

For a function $g: \mathcal{V} \rightarrow \mathbb{R}$ and $\lambda := (\lambda_1, \dotsc, \lambda_L) \in \mathbb{R}^L_{+}$, we define the function $g_{\lambda, A}: \mathcal{V} \rightarrow \mathbb{R} \cup \{\infty \}$ as

align*[align* omitted — 194 chars of source]

with $v:= (s_1, \dotsc, s_L)$ and $v^\prime := (s_1^\prime, \dotsc, s_L^\prime)$.

theorem[Non-overlapping case] Suppose that (ref) hold. Then, for any $\delta \in \mathbb{R}_{++}^L$ and $A \subset [L]$, we have \begin{align*} \mathcal{I}_{\mathrm{D} }(\delta_{A} ) = \inf_{\lambda \in \mathbb{R}_{+}^L} \left[ \langle \lambda, \delta_{A} \rangle + \sup_{\pi \in \Pi(\mu_1, \dotsc, \mu_L)} \int_{\mathcal{V}} g_{\lambda, A} \, d \pi \right]. \end{align*}

In practice, the dual in (ref) involves the computation of the multi-marginal problem, $\sup_{\pi \in \Pi(\mu_1, \dotsc, \mu_L)} \int_{\mathcal{V}} g_{\lambda} \, d \pi$, see Pass_2010,Pass_2012,Pass_2015,von_Lindheim_2022,Nenna_2022,mehta2023efficient for detailed studies of properties and computation of multi-marginal problems for specific functions $g_{\lambda}$. For general possibly non Borel-measurable $g_{\lambda}$, the strong duality in Kellerer:1984to could be applied. The established result is stated in (ref) in (ref).

Overlapping Marginals

Let $\mathcal{S} := \left(\prod_{\ell \in [L]} \mathcal{Y}_\ell \right) \times \mathcal{X}$, where $\mathcal{Y}_\ell$ for $\ell \in [L]$ and $\mathcal{X}$ are Polish spaces. Let $\mathcal{S}_\ell := \mathcal{Y}_\ell \times \mathcal{X}$ for $\ell \in [L]$. Let $\mu_{\ell,L+1} \in \mathcal{P}(\mathcal{S}_\ell)$ for $\ell \in [L]$ be such that the projections of $\mu_{\ell, L+1}$ on $\mathcal{X}$ are the same for $\ell \in [L]$. We call the Fr\'{e}chet class of all probability measures on $\mathcal{S}$ having marginals $\left(\mu_{1,L+1} \right)_{\ell \in [L]}$ the Fr\'{e}chet class with overlapping marginals and denote it as $\mathcal{F}\left(\mathcal{S}; (\mu_{\ell,L+1})_{\ell \in [L]} \right) := \mathcal{F}\left( \left(\mu_{\ell,L+1} \right)_{\ell \in [L]} \right)$. This class is the star-like system of marginals in ruschendorf1991bounds and Embrechts_2010, see also Doan2015.

Moreover, let $f:\mathcal{S} \rightarrow \mathbb{R}$ be a measurable function satisfying the following assumption.

assumptionThe function $f:\mathcal{S} \rightarrow \mathbb{R}$ is a measurable function such that $\int_{\mathcal{S}} f \, d\nu_0 > - \infty$ for some $\nu_0 \in \Pi(\mu_{1,L+1},..., \mu_{L,L+1}) \subset \mathcal{P}(\mathcal{S})$.

For any $\gamma \in \mathcal{P}(\mathcal{S})$, let $\gamma_{\ell, L+1}$ denote the projection of $\gamma$ on $\mathcal{Y}_{\ell} \times \mathcal{X}$ for $\ell \in [L]$. Similar to the two marginals case, the W-DMR with overlapping multi-marginals is defined as

align*[align* omitted — 104 chars of source]

where $\Sigma(\delta)$ is the uncertainty set defined as

align*[align* omitted — 173 chars of source]

in which $\delta = (\delta_1, \dotsc, \delta_L) \in \mathbb{R}_{+}^L$ is the radius of the uncertainty set.

For a function $f: \mathcal{V} \rightarrow \mathbb{R}$, $\lambda := (\lambda_1, \dotsc, \lambda_L) \in \mathbb{R}^L_{+}$, and $A \subset [L]$, we define the function $f_{\lambda, A}: \mathcal{V} \to \overline{\mathbb{R}}$ as follows:

align*[align* omitted — 192 chars of source]

where $v = (s_1, \dotsc, s_L)$, $s^\prime = (y^\prime_1, \dotsc, y^\prime_L, x^\prime)$, $s^\prime_\ell = (y^\prime_\ell, x^\prime)$ and $s_\ell=(y_\ell, x_\ell)$, and

align*[align* omitted — 309 chars of source]
theorem[Overlapping case] Suppose that (ref) hold. Then, for any $\delta \in \mathbb{R}^L_{++}$ and $A \subset [L]$, we have \begin{align*} \mathcal{I}(\delta_{A} ) = \inf_{\lambda \in \mathbb{R}_{+}^L} \left[ \langle \lambda, \delta_{A} \rangle + \sup_{\pi \in \Pi(\mu_{1, L+1},\dotsc, \mu_{L,L+1})} \int_{\mathcal{V}} f_{\lambda, A} \, d \pi \right]. \end{align*}

Similar to the nonoverlapping case, strong duality holds for the inner multi-marginal problem under additional conditions. The result is stated in (ref) of (ref).

Treatment Choice for Multi-valued Treatment

We apply strong duality to multi-valued treatment in Kido2022. Let $d: \mathcal{X} \rightarrow [L]$ be a policy function or treatment rule on $\mathcal{X}$ and $Y_\ell \in \mathbb{R}$ denote the potential outcome under the treatment $\ell$ for $\ell \in [L]$. Consider the policy function defined as

align*[align* omitted — 81 chars of source]

Kido2022 introduces the following robust welfare function.

align*[align* omitted — 175 chars of source]

where the uncertainty set $\Sigma_{\mathrm{M}}(\delta_0)$ is based on the conditional distribution of $(Y_\ell)_{\ell \in [L]}$ given $X$:

align*[align* omitted — 238 chars of source]

in which the cost function $c$ associated with $\boldsymbol{K}$ is

align*[align* omitted — 100 chars of source]

Note that the uncertainty set $\Sigma_{\mathrm{M}}(\delta_0)$ does not allow any potential shift\footnote{Kido2022 mentions the possibility of allowing for covariate shift by incorporating uncertainty sets in Mo2020,Zhao2019 for the distribution of the covariate in future work. } in $X$. When $Y_1, \dotsc, Y_L$ are unbounded, Kido2022 shows that

align*[align* omitted — 279 chars of source]

We apply W-DMR for overlapping marginals with the following cost function:

align*[align* omitted — 95 chars of source]

and define a robust welfare function as

align*[align* omitted — 139 chars of source]
propositionFor $\ell \in [L]$, let \begin{align*} c_{\ell}(s_\ell, s_\ell') = | y_\ell - y_\ell'| + \| x_{\ell} - x_\ell' \|_{2}. \end{align*} Assume that $Y_\ell$ is unbounded, $\mathbb{E}[\|X\|_2^2] < \infty$ and $\mathbb{E}[|Y_\ell|] < \infty$. Then \begin{align*} \mathrm{RW}(d) & = \sup_{\lambda \ge 1} \left\{ \inf_{\pi \in \Pi(\mu_{1,L+1}, \dotsc, \mu_{L, L+1})} \int_{\mathcal{V}} \min_{\ell \in [L]} \{y_\ell + \phi_{\lambda, \ell}(x_1, \dotsc, x_L) \} d \pi(s) - \langle \lambda, \delta \rangle \right\}, \end{align*} where \begin{align*} \varphi_{\lambda, \ell}(x_1, \dotsc, x_L) = \min_{x', d(x') = \ell} \sum_{\ell=1}^{L} \lambda_\ell \| x_\ell - x' \|_{2}. \end{align*}

(ref) is an extension of (ref).

Concluding Remarks

In this paper, we have introduced W-DMR in marginal problems for both non-overlapping and overlapping marginals and established fundamental results including strong duality, finiteness of the proposed W-DMR, and existence of an optimizer at each radius. We have also shown continuity of the W-DMR-MP as a function of the radius. Applicability of the proposed W-DMR in marginal problems and established properties is demonstrated via distinct applications when the sample information comes from multiple data sources and only some marginal reference measures are identified. To the best of the authors' knowledge, this paper is the first systematic study of W-DMR in marginal problems. Many open questions remain including the structure of optimizers of W-DMR for both non-overlapping and overlapping marginals, efficient numerical algorithms, and estimation and inference in each motivating example. Another useful extension is to consider objective functions that are nonlinear in the joint probability measure such as the Value-at-Risk of a linear portfolio of risks in Puccetti_2012 and robust spectral measures of risk in Ghossoub_2023,Ennaji2022.

\printbibliography