Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
107,282 characters · 33 sections · 119 citation commands
Quantifying Distributional Model Risk in Marginal Problems via Optimal Transport
Distributionally robust optimization (DRO) has emerged as a powerful tool for hedging against model misspecification and distributional shifts. It minimizes distributional model risk (DMR) defined as the worst risk over a class of distributions lying in a distributional uncertainty set, see blanchet2019quantifying. Among many different choices of uncertainty sets, Wasserstein DRO (W-DRO) with distributional uncertainty sets based on optimal transport costs has gained much popularity, see Kuhn_2019 and Blanchet2021 for recent reviews. W-DRO has found successful applications in robust decision making in all disciplines including economics, finance, machine learning, and operations research. Its success is largely credited to the strong duality and other nice properties of the Wasserstein DMR (W-DMR). The objective of this paper is to propose and study W-DMR in marginal problems where only some marginal measures of a reference measure are given, see e.g., Kellerer:1984to,rachev1998mass,villani2009optimal, villani2021topics, and ruschendorf1991bounds.
In practice, marginal problems arise from either the lack of complete data or an incomplete model. In insurance and risk management, computing model-free measures of aggregate risks such as Value-at-Risk and Expected Short-Fall is of utmost importance and routinely done. When the exact dependence structure between individual risks is lacking, researchers and policy makers rely on the worst risk measures defined as the maximum value of aggregate risk measures over all joint measures of the individual risks with some fixed marginal measures, see Embrechts_2010 and Embrechts_2013; In causal inference, distributional treatment effects such as the variance and the proportion of participants who benefit from the treatment depend on the joint distribution of the potential outcomes. Even with ideal randomized experiments such as double-blind clinical trials, the joint distribution of potential outcomes is not identified and as a result, only the lower and upper bounds on distributional treatment effects are identified from the sample information, see FAN_2009, Fan_2010_SharpBounds,Fan_2012_CI_quantiles, Fan2017, Ridder_2007, Firpo_2019; In algorithmic fairness when the sensitive group variable is not observed in the main data set, assessment of unfairness measures must be done using multiple data sets, see Kallus2022. Abstracting away from estimation, all these problems involve optimizing the expected value of a functional of multiple random variables with fixed marginals and thus belong to the class of marginal problems for which optimal transport related tools are important.\footnote{When the marginals are univariate, optimal transport problem can be conveniently expressed in terms of copulas. Fan_2010_SharpBounds, Fan_2012_CI_quantiles, FAN_2009, Fan2017, Ridder_2007, and Firpo_2019 explicitly use copula tools. }
The marginal measures in the afore-mentioned applications and general marginal problems are typically empirical measures computed from multiple data sets such as in the evaluation of worst aggregate risk measures or identified under specific assumptions such as randomization or strong ignorability in causal inference. Developing a unified framework for hedging against model misspecification and/or distributional shifts in marginal measures motivates the current paper.
Theoretically, this paper makes several contributions to the literature on distributional robustness and the literature on marginal problems. First, it introduces Wasserstein distributional model risk in marginal problems (W-DMR-MP), where each marginal measure is assumed to lie in a Wasserstein ball centered at a fixed reference measure with a given radius. We focus on the important case with two marginals and consider both non-overlapping and overlapping marginals. For non-overlapping marginal measures, when the radius is zero, the W-DMR-MP reduces to the marginal problems or optimal transport problems studied in Kellerer:1984to,rachev1998mass,villani2009optimal, villani2021topics. For overlapping marginals, when the radius is zero, the W-DMR-MP reduces to the overlapping marginals problem studied in ruschendorf1991bounds; Second, we establish strong duality for our W-DMR with both non-overlapping and overlapping marginals under similar conditions to those for W-DMR, see zhang2022simple, blanchet2019quantifying, and gao2022WDRO. As a first application of our strong duality result for non-overlapping marginals, we extend the well-known Marakov bounds for the distribution function of the sum of two random variables to Wasserstein distributionally robust Makarov bounds; Third, we prove finiteness of the W-DMR-MP and existence of an optimizer at each radius. Based on both results, we show that the identified set of the expected value of a smooth functional of random variables with fixed marginals is a closed interval; Fourth, we show continuity of the W-DMR in marginal problems as a function of the radius. Together these results extend those for W-DMR in blanchet2019quantifying, zhang2022simple, and yue2022linear; Lastly, we extend our formulations and theory to W-DMR with multi-marginals. On a technical note, our proofs build on existing work on W-DMR such as blanchet2019quantifying, zhang2022simple, and yue2022linear. However, an additional challenge due to the presence of multiple marginal measures in our Wasserstein uncertain sets is the verification of the existence of a joint measure with overlapping marginals. We make use of existing results for a given consistent product marginal system in vorob1962consistent, kellerer1964verteilungsfunktionen, and shortt1983combinatorial to address this issue.
Practically, we demonstrate the flexibility and broad applicability of our W-DMR-MP via four distinct applications when the sample information comes from multiple data sources. First, we consider partial identification of treatment effects when the marginal measures of the potential outcomes lie in their respective Wasserstein balls centered at the measures identified under strong ignorability. The validity of strong ignorability is often questionable when unobservable confounders may be present. We apply our W-DMR-MP to establishing the identified sets of treatment effects which can be used to conducting sensitivity analysis to the selection-on-observables assumption. For average treatment effects, we show that when the cost functions are separable, incorporating covariate information does not help shrink the identified set; on the other hand, for non-separable cost functions such as the Mahalanobis distance, incorporating covariate information may help shrink the identified set; Second, in causal inference when the optimal treatment choice is to be applied to a target population different from the training population, adjaho2022 introduces robust welfare functions defined by W-DMR to study externally valid treatment choice. The W-DMR-MP we propose allows us to dispense with the assumption of a known dependence structure for the reference measure in adjaho2022. When shifts in the covariate distribution are allowed, we show that our robust welfare function is upper bounded by the worst robust welfare function of adjaho2022; Third, one important application of W-DMR is in distributionally robust estimation and classification. However as awasthi2022distributionally points out,\footnote{See Graham2016 and Chen_2008 for general data combination problems.} some sensitive variables may not be observed in the same data set as the response variable rendering W-DRO inapplicable. We apply W-DMR-MP to distributionally robust estimation under data combination;\footnote{(ref) provides a detailed comparison of our set up and awasthi2022distributionally.} Fourth, applying our W-DMR-MP to the evaluation of the worst aggregate risk measures allows us to dispense with the known marginals assumption in Embrechts_2010 and Embrechts_2013.
The rest of this paper is organized as follows. (ref) reviews the W-DMR and strong duality, introduces our W-DMR-MP, and then presents four motivating examples. (ref) establishes strong duality and Wasserstein distributionally robust Marakov bounds. (ref) studies finiteness of W-DMR-MP and existence of optimal solutions. Moreover, we show that the identified set of the expected value of a smooth functional of random variables with fixed marginals is a closed interval. (ref) establishes continuity of W-DMR-MP as a function of the radius. (ref) revisits the motivating examples in (ref). (ref) extends our W-DMR-MP to more than two marginals. The last section offers some concluding remarks. Technical proofs are relegated to a series of appendices.
We close this section by introducing the notation used in the rest of this paper. For two sets $A$ and $B$, the relative complement is denoted by $A \setminus B$. Let $\overline{\mathbb{R}} = \mathbb{R} \cup \left\{- \infty, \infty \right \}$, $[d]=\{1,2,...,d\}$, $\mathbb{R}^d_{+} = \left \{ x \in \mathbb{R}^d : x_i \geq 0, \ \forall i \in [d] \right\}$, and $\mathbb{R}^d_{++} = \left\{ x \in \mathbb{R}^d : x_i > 0, \ \forall i \in [d] \right\}$. For any real numbers $x, y \in \mathbb{R}$, we define $x \wedge y := \min\{x, y\}$ and $x \vee y : = \max\{x, y\}$. The Euclidean inner product of $x$ and $y$ in $\mathbb{R}^d$ is denoted by $\langle x, y \rangle$. For any real matrix $W \in \mathbb{R}^{m \times n}$, let $A^\top$ denote the transpose of $W$. For an extended real function $f$ on $\mathcal{X}$, the positive part $f^+$ and the negative part $f^-$ are defined as $f^{+}(x) = \max \left\{ f(x), 0 \right\}$ and $f^{-}(x) = \max \left\{ -f(x), 0 \right\}$, respectively.
For any Polish space $\mathcal{S}$, let $\mathcal{B}_\mathcal{S}$ be the associated Borel $\sigma$-algebra and $\mathcal{P}(\mathcal{S})$ be the collection of probability measures on $\mathcal{S}$. Given a Polish probability space $(\mathcal{S}, \mathcal{B}_{\mathcal{S}}, \nu)$, let $\mathcal{B}_{\mathcal{S}}^\nu$ denote the $\nu$-completion of $\mathcal{B}_{\mathcal{S}}$. Given a probability space $\left(\Omega, \mathcal{F}, \mathbb{P} \right)$ and a map $T: \Omega \rightarrow \mathcal{S}$, let $T\# \mu$ denote the push forward of $\mathbb{P}$ by $T$, i.e., $(T\# \mathbb{P})(A) = \mathbb{P} \left( T^{-1} (A) \right)$ for all $A \in \mathcal{B}_{\mathcal{S}}$, where $T^{-1}(A) = \left\{ \omega\in\Omega: T(\omega) \in A \right\}$. The law of a random variable $S: \Omega \rightarrow \mathbb{R}$ is denoted by $\mathrm{Law}(S)$ which is the same as $S\# \mathbb{P}$. For any $\mu, \nu \in \mathcal{P}(\mathcal{S})$, let $\Pi(\mu, \nu)$ denote the set of all couplings (or joint measures) with marginals $\mu$ and $\nu$.
For any $\mathcal{B}_{\mathcal{S}}^{\nu}$-measurable function $f$, let $\int_{\mathcal{S}} f d \nu$ denote the integral of $f$ in the completion of $(\mathcal{S}, \mathcal{B}_{\mathcal{S}}, \nu)$. For a random element $S: \Omega \rightarrow \mathcal{S}$ with $\mathrm{Law}(S) = \nu$, we write $\mathbb{E}_\nu \left[f(S)\right] = \int_{\mathcal{S}} f d \nu$. Given $p \in (0, \infty)$ and a Borel measure $\nu$ on $\mathcal{S}$, let $L^p(\nu) := L^p(\mathcal{S},\mathcal{B}_{\mathcal{S} } , \nu)$ denote the set of all the $\mathcal{B}_{\mathcal{S}}^{\nu}$-measurable functions $f: \mathcal{S}\rightarrow \mathbb{R}$ such that $\| f\|_{L^p(\nu)} := \left( \int_{\mathcal{S}} |f|^p d\nu \right)^{1/p}< \infty$.
In this section, we first review W-DMR and then introduce W-DMR in marginal problems. Lastly, we present four motivating examples of marginal problems which will be used to illustrate our results in the rest of this paper.
W-DMR is defined as the worst model risk over a class of distributions lying in a Wasserstein uncertainty set composed of all probability measures that are a fixed Wasserstein distance away from a given reference measure, see blanchet2019quantifying.
Before presenting W-DMR, we review some basic definitions. Let $\mathcal{X}$ be a Polish (metric) space with a metric $\boldsymbol{d}$.
When the cost function $c$ is lower-semicontinuous, there exists an optimal coupling corresponding to $\boldsymbol{K}_c(\mu, \nu)$. In other words, there exists $\pi^* \in \Pi(\mu, \nu)$ such that $\boldsymbol{K}_c(\mu, \nu)=\int_{\mathcal{X} \times \mathcal{X}} c \, d\pi^*$ villani2009optimal.
Throughout this paper, we make the following assumption on the cost function $c$.
(ref) implies that for $\mu, \nu \in \mathcal{P}(\mathcal{X})$, $\mu = \nu$ if and only if $\boldsymbol{K}_c(\mu, \nu) = 0$. When $c$ is the metric $\boldsymbol{d}$ on $\mathcal{X}$, $ \boldsymbol{K}_c(\mu, \nu)$ coincides with the Wasserstein distance of order 1 (Kantorovich-Rubinstein distance) between $\mu$ and $\nu$ defined in (ref).
For a given function $f: \mathcal{X} \to \mathbb{R}$, blanchet2019quantifying define W-DMR as
where $\Sigma_{\mathrm{DMR}}(\delta)$ is the Wasserstein uncertainty set\footnote{By convention, we call all uncertainty sets based on optimal transport costs as Wasserstein uncertainty sets.} centered at a reference measure $\mu \in \mathcal{P}(\mathcal{X})$ with radius $\delta\geq 0$, i.e.,
(ref) allows the cost function $c$ to be asymmetric and take value $\infty$, where the latter corresponds to the case that there is no distributional shift in some marginal measure of $\mu$.
It is well-known that under mild conditions, strong duality holds for $\mathcal{I}_{\mathrm{DMR}}(\delta)$ when $\delta > 0$ (c.f., blanchet2019quantifying,gao2022WDRO,zhang2022simple). To be self-contained, we restate the strong duality result in zhang2022simple for Polish space below.\footnote{The strong duality result in zhang2022simple allows for general space $\mathcal{X}$.}
In the rest of this paper, we keep the convention that for any cost function $c$, $\lambda c(x, y) = \infty$ when $\lambda = 0$ and $c(x, y) = \infty$.
Let $\mathcal{V} := \mathcal{S}_1 \times \mathcal{S}_2$ be the product space of two Polish spaces $\mathcal{S}_1$ and $\mathcal{S}_2$. Let $\mu_1$ and $\mu_2$ be Borel probability measures on $\mathcal{S}_1$ and $\mathcal{S}_2$ respectively. Following ruschendorf1991bounds (see also Embrechts_2010), we call the Fr\'{e}chet class of all probability measures on $\mathcal{V}$ having marginals $\mu_1$ and $\mu_2$ the Fr\'{e}chet class with non-overlapping marginals denoted as $\mathcal{F}(\mathcal{V}; \mu_{1}, \mu_{2}) := \mathcal{F}(\mu_{1}, \mu_{2})$. Note that $\mathcal{F}(\mu_{1}, \mu_{2})=\Pi(\mu_1, \mu_2)$.
Let $g:\mathcal{V}\rightarrow \mathbb{R}$ be a measurable function satisfying the following assumption.
The marginal problem associated with $\mu_1$ and $\mu_2$ is defined as
It is essentially an optimal transport problem, where the $\sup$ operation is replaced with the $\inf$ operation, see Kellerer:1984to,rachev1998mass,villani2009optimal, villani2021topics) or (ref) for a review of strong duality for $\mathcal{I}_{\mathrm{M} }(\mu_1,\mu_2)$.
The W-DMR with non-overlapping marginals we propose extends the marginal problem by allowing each marginal measure of $\gamma$ to lie in a fixed Wasserstein distance away from a reference measure. Specifically, for any $\gamma \in \mathcal{P}(\mathcal{V} )$, let $\gamma_{1}$ and $\gamma_{2}$ denote the projection of $\gamma$ on $\mathcal{S}_1$ and $\mathcal{S}_2$, respectively. The W-DMR with non-overlapping marginals is defined as
where $\Sigma_{\mathrm{D} }(\delta)$ is the uncertainty set given by \[ \Sigma_{\mathrm{D} }(\delta) := \Sigma_{\mathrm{D} }(\mu_1, \mu_2, \delta) = \left\{ \gamma \in \mathcal{P}(\mathcal{V}): \ \boldsymbol{K}_1( \mu_1, \gamma_1) \leq \delta_1, \ \boldsymbol{K}_2( \mu_2, \gamma_2) \leq \delta_2 \right\}, \] in which $\boldsymbol{K}_1$ and $\boldsymbol{K}_2$ are optimal transport costs associated with cost functions $c_1$ and $c_2$, respectively, and $\delta := (\delta_{1}, \delta_{2}) \in \mathbb{R}_{+}^2$ is the radius of the uncertainty set. Obviously $\Sigma_{\mathrm{D} }(\delta)$ is non-empty for all $\delta \in \mathbb{R}_{+}^2$.
Let $\mathcal{S} := \mathcal{Y}_1 \times \mathcal{Y}_2 \times \mathcal{X}$ be the product space of three Polish spaces $\mathcal{Y}_1$, $\mathcal{Y}_2$, and $\mathcal{X}$. Let $\mathcal{S}_1 := \mathcal{Y}_1 \times \mathcal{X}$ and $\mathcal{S}_2 := \mathcal{Y}_2 \times \mathcal{X}$. Let $\mu_{13} \in \mathcal{P}(\mathcal{S}_1)$ and $\mu_{23} \in \mathcal{P}(\mathcal{S}_2)$ be such that the projection of $\mu_{13}$ and the projection of $\mu_{23}$ on $\mathcal{X}$ are the same. Following ruschendorf1991bounds (see also Embrechts_2010), we call the Fr\'{e}chet class of all probability measures on $\mathcal{S}$ having marginals $\mu_{13}$ and $\mu_{23}$ the Fr\'{e}chet class with overlapping marginals and denote it as $\mathcal{F}(\mathcal{S}; \mu_{13}, \mu_{23}) := \mathcal{F}(\mu_{13}, \mu_{23}).$ Unlike the non-overlapping case, $\mathcal{F}(\mu_{13}, \mu_{23})$ is different from the class of couplings $\Pi(\mu_{13}, \mu_{23})$.
Let $f:\mathcal{S} \rightarrow \mathbb{R}$ be a measurable function satisfying the following assumption.
ruschendorf1991bounds studies the following marginal problem with overlapping marginals:
As shown in ruschendorf1991bounds, the marginal problem with overlapping marginals can be computed via the marginal problem with non-overlapping marginals through the following relation:
where for each fixed $x\in\mathcal{X}$, $\mu_{\ell|3}( \cdot| x)$ denote the conditional measure of $Y_\ell$ given $X=x$, and the inner optimization problem is a marginal problem with non-overlapping marginals.
For any $\gamma \in \mathcal{P}(\mathcal{S} )$, let $\gamma_{13}$ and $\gamma_{23}$ denote the projections of $\gamma$ on $\mathcal{Y}_1 \times \mathcal{X}$ and $\mathcal{Y}_2 \times \mathcal{X}$, respectively. The W-DMR with overlapping marginals is defined as
where $\Sigma(\delta)$ is the uncertainty set given by \[ \Sigma(\delta):=\Sigma(\mu_{13}, \mu_{23}, \delta) = \left\{ \gamma \in \mathcal{P}(\mathcal{S} ): \ \boldsymbol{K}_1( \mu_{13}, \gamma_{13}) \leq \delta_1, \ \boldsymbol{K}_2( \mu_{23}, \gamma_{23}) \leq \delta_2 \right\} \] in which $\delta := (\delta_{1}, \delta_{2}) \in \mathbb{R}_{+}^2$ is the radius of the uncertainty set, and $\boldsymbol{K}_1$ and $\boldsymbol{K}_2$ are optimal transport costs associated with $c_1$ and $c_2$. We note that $\Sigma(\delta)$ is non-empty for all $\delta \in \mathbb{R}_{+}^2$.
In this section, we present four distinct examples to demonstrate the wide applicability of the W-DMR in marginal problems. The first example is concerned with partial identification of treatment effect parameters when commonly used assumptions in the literature for point identification fail; the second example is concerned with distributionally robust optimal treatment choice; the third one is an application of W-DMR-MP in distributionally robust estimation under data combination; and the last one concerns measures of aggregate risk.
For the first two examples, we adopt the potential outcomes framework for a binary treatment. Let $D \in \{0,1\}$ represent an individual's treatment status, and $Y_1\in \mathcal{Y}_1\subset \mathbb{R}$ and $Y_2\in \mathcal{Y}_2\subset \mathbb{R}$ denote the potential outcomes under treatments $D = 0$ and $D = 1$, respectively. Let the observed outcome be \[ Y = D Y_2 + (1-D) Y_1. \]
To focus on introducing the main ideas, we adopt the selection-on-observables framework stated in (ref) below.
Suppose a random sample on $ (Y,X,D)$ is available. Then under (ref), the marginal conditional distribution functions of $Y_1, Y_2$ given $X=x$ are point identified: \[ F_{Y_1|X}(y|x)=\mathbb{P}(Y_1\leq y|X=x)=\mathbb{P}(Y\leq y|X=x,D=0)\] and \[ F_{Y_2|X}(y|x)=\mathbb{P}(Y_2\leq y|X=x)=\mathbb{P}(Y\leq y|X=x,D=1). \] As a result, the probability measures $\mu_{13}$ of $(Y_{1},X)$ and $\mu_{23}$ of $(Y_{2},X)$ are identified as well.
(ref) is commonly used to identify treatment effect parameters and optimal treatment choice. However the validity of (ref) may be questionable when there are unobserved confounders. W-DMR-MP presents a viable approach to studying sensitivity of causal inference to deviations from (ref) by varying the marginal measures of a joint measure of $(Y_1, Y_2, X)$ in Wasserstein uncertainty sets centered at reference measures consistent with (ref). Specifically, let $f$ be a measurable function of $Y_1, Y_2$. Consider treatment effects of the form: $\theta_o:=\mathbb{E}_o[f(Y_1,Y_2)]$, where $\mathbb{E}_o$ denotes expectation with respect to the true measure. It includes the average treatment effect (ATE) for which $f(Y_1, Y_2)=Y_2 - Y_1$ and the distributional treatment effect such as $\mathbb{P}_o(Y_2 - Y_1\geq 0)$, where $\mathbb{P}_o$ denotes the probability computed under the true measure.
Consider the identified set for $\theta_o$ defined as \[ \Theta(\delta) := \left\{ \int_{\mathcal{S} } f(y_1,y_2) \, d \gamma(y_1, y_2, x) : \gamma \in \Sigma( \delta) \right\}, \] where
in which $\mu_{13}$ and $\mu_{23}$ are the identified measures of $(Y_1, X)$ and $(Y_2, X)$ under (ref). Under mild conditions, we show in (ref) that the identified set $\Theta(\delta)$ is a closed interval given by \[ \Theta(\delta)= \left[ \min_{\gamma \in \Sigma(\delta)} \int_\mathcal{S} f(y_1,y_2) \, d \gamma(s), \max_{\gamma \in \Sigma(\delta)} \int_\mathcal{S} f(y_1,y_2) \, d \gamma(s) \right], \] where the lower and upper limits of the interval are characterized by the W-DMR-MP.\footnote{Since $\inf_{\gamma \in \Sigma(\delta)} \int_{\mathcal{S}} f(y_1, y_2) d \gamma(s)$ can be rewritten as $ - \sup_{\gamma \in \Sigma(\delta)} \int_{\mathcal{S}} [- f(y_1, y_2)] d \gamma(s)$, we also refer to the lower limit as W-DMR-MP. } When $\delta=0$, Fan2017 establish a characterization of $\Theta(0)$ via marginal problems with overlapping marginals.
The identified set $\Theta(\delta)$ can be used to conduct sensitivity analysis to deviations from (ref). We note that sensitivity analysis to other commonly used assumptions such as the threshold-crossing model can be done by taking the reference measures as the measures identified under these alternative assumptions, see FAN_2009.
In empirical welfare maximization (EWM), an optimal choice/policy is chosen to maximize the expected welfare estimated from a training data set and then applied to a target population, see Kitagawa_2018. EWM assumes that the target population and the training data set come from the same underlying probability measure. This may not be valid in important applications. Motivated by designing externally valid treatment policy, adjaho2022 introduces a robust welfare function which allows the target population to differ from the training population. In this paper, we revisit adjaho2022's robust welfare function and propose a new one based on W-DMR with overlapping marginals.
adjaho2022 adopts the following definition of a robust welfare function:
where $d : \mathcal{X} \rightarrow \{0, 1\}$ is a measurable policy function, i.e., $d(X)$ is $0$ or $1$ depending on $X$ and $\Sigma_0(\delta_0)$ is the Wasserstein uncertainty set centered at a joint measure $\mu$ for $(Y_1, Y_2, X)$ consistent with (ref), i.e.,
where $\boldsymbol{K}_c(\mu, \gamma)$ is the optimal transport cost with cost function $c: \mathcal{S} \times \mathcal{S} \rightarrow \mathbb{R}_{+}\cup \{\infty\}$.
Noting that (ref) only identifies the marginal measures $\mu_{13},\mu_{23}$ of the reference measure $\mu$ in $\Sigma_0(\delta_0)$, we define a new robust welfare function as
where $\Sigma(\delta) = \Sigma(\mu_{13}, \mu_{23}, \delta)$ is the uncertainty set for W-DMR with overlapping marginals.
An important application of W-DMR is W-DRO. Let $f: \mathcal{Y}_1 \times \mathcal{Y}_2 \times \mathcal{X}\times \Theta \to \mathbb{R}$ be a loss function with an unknown parameter $\theta \in \Theta\subset \mathbb{R}^q$. W-DRO under data combination is defined as
where $\Sigma(\delta)$ is the uncertainty set for the overlapping case. For each $\theta \in \Theta$, the inner optimization is a W-DMR with overlapping marginals. In practice, we need to choose the reference measures $\mu_{13}$ and $\mu_{23}$ based on the sample information. Focusing on logit model, where $\mathcal{Y}_1 = \{+1, -1\}$ is the space for the dependent variable, and $\mathcal{Y}_2$ and $\mathcal{X}$ are feature spaces/covariate space, and
awasthi2022distributionally proposes a method dubbed `Robust Data Join' in which the empirical measures constructed from the two data sets are used as reference measures. Specifically, let $\widehat{\mu}_{13}$ and $\widehat{\mu}_{23}$ denote empirical measures based on two separate data sets. The uncertainty set in awasthi2022distributionally takes the following form: \[ \Sigma_{\mathrm{RDJ} } (\delta):= \left\{ \gamma \in \mathcal{P}(\mathcal{S} ): \ \boldsymbol{K}_1( \widehat{\mu}_{13}, \gamma_{13}) \leq \delta_1, \ \boldsymbol{K}_2(\widehat{\mu}_{23}, \gamma_{23}) \leq \delta_2 \right\}, \] where \[ c_1((y_1, x), (y_1', x')) = \| x - x' \|_{p} + \kappa_1 | y_1 - y_1'| \quad \text{and} \] \[ c_2((y_2, x), (y_2, x')) = \| x - x' \|_{p} + \kappa_2 \| y_2 - y_2' \|_{p'} \] with $\kappa_1\geq 1$, $\kappa_2\geq 1$, $p\geq 1$, and $p'\geq 1$.
Note that awasthi2022distributionally's `Robust Data Join' is different from our W-DMR with non-overlapping marginals because the measure of interest $\gamma \in \mathcal{P}(\mathcal{S} )$ has overlapping marginals. It is also different from our W-DMR with overlapping marginals because the reference measures $\widehat{\mu}_{13}$ and $\widehat{\mu}_{23}$ may not have overlapping marginals. Unlike the uncertainty set for W-DMR, $\Sigma_{RDJ}(\delta)$ is empty when $\delta=0$.
Let $S_1, S_2$ be random variables representing individual risks defined on Polish spaces $\mathcal{S}_1, \mathcal{S}_{2}$, respectively. Let $\mu_1, \mu_2$ be probability measures of $S_1, S_2$. Let $\mathcal{V} = \mathcal{S}_1 \times \mathcal{S}_2$ and $g: \mathcal{V} \to \mathbb{R}$ be a risk aggregating function. Applying W-DMR with non-overlapping marginals to the risk aggregation function $g$, we can compute the worst aggregate risk when the joint measure of the individual risks varies in the uncertainty set $\Sigma_{\mathrm{D} }(\delta)$. This is different from the set-up in Eckstein2020, where the following robust risk aggregation problem is studied:
where
in which $\boldsymbol{K}_c$ is the optimal transport cost associated with a cost function $c: \mathcal{V} \times \mathcal{V} \to \mathbb{R}_{+}$. Since $\gamma \in \Sigma_{\Pi}(\delta_0)$ is a coupling of $(\mu_1, \mu_2)$, we have that $\Sigma_{\Pi}(\delta_0) \subset \Sigma_{\mathrm{D} }(0)$ and thus $\mathcal{I}_\Pi(\delta_0) \le \mathcal{I}_{\mathrm{D} }(0)$.
In this section, we establish strong duality for our W-DMR-MP and apply it to develop Wasserstein distributionally robust Makarov bounds.
For a measurable function $g: \mathcal{V} \rightarrow \mathbb{R}$ and $\lambda := (\lambda_1, \lambda_2) \in \mathbb{R}^2_{+}$, we define the function $g_\lambda: \mathcal{V}\rightarrow \mathbb{R} \cup \{\infty \}$ as \[\tag{2.1} g_{\lambda}(v) := \sup_{v^\prime \in \mathcal{V}} \varphi_\lambda(v, v^\prime), \] where $\varphi_\lambda: \mathcal{V} \times \mathcal{V} \rightarrow \mathbb{R} \cup \{ -\infty\}$ is given by \[ \varphi_\lambda(v, v^\prime) = g\left(s_1^\prime, s_2^\prime\right) - \lambda_1 c_1 \left(s_1, s_1^\prime\right) -\lambda_2 c_2 \left(s_2, s_2^\prime\right), \] with $v:= (s_1, s_2)$ and $v^\prime := (s_1^\prime, s^\prime_2)$. Similarly, define $g_{\lambda_1, 1}: \mathcal{V} \rightarrow \mathbb{R} \cup \{+\infty\}$ and $g_{\lambda_2, 2}: \mathcal{V} \rightarrow \mathbb{R} \cup \{+\infty\}$ as
The dual problem $\mathcal{J}_{\mathrm{D}}(\delta)$ corresponding to the primal problem $\mathcal{I}_{\mathrm{D}}(\delta)$ is defined as follows:
Unlike the dual for W-DMR, the dual for W-DMR with non-overlapping marginals in (ref) involves a marginal problem with non-overlapping marginals $\mu_1, \mu_2$ due to the lack of knowledge on the dependence of the joint measure $\mu$. Computational algorithms developed for optimal transport can be used to solve the marginal problem, see Peyre2018. For empirical measures $\mu_1, \mu_2$, the marginal problem is a discrete optimal transport problem and there are efficient algorithms to compute it, see Peyre2018. For general measures $\mu_1, \mu_2$, strong duality may be employed in the numerical computation of the marginal problem. For instance, consider the case when $\delta > 0$. When $g_\lambda (v)$ is Borel measurable, several strong duality results are available, see e.g., villani2009optimal, villani2021topics. For a general function $g$ and cost functions $c_1,c_2$, $g_\lambda (v)$ is not guaranteed to be Borel measurable. However, for Polish spaces, the set $\{v \in \mathcal{V}:g_{\lambda}(v) \ge u\}$ is an analytic set for all $u \in \overline{\mathbb{R}}$ (and $g_{\lambda}$ is universally measurable), since $g$, $c_1$ and $c_2$ are Borel measurable (see blanchet2019quantifying and Bertsekas1978). This allows us to apply strong duality for the marginal problem in Kellerer:1984to restated in (ref) to the marginal problem involving $g_\lambda (v)$, see (ref) in (ref).
Without additional assumptions on the function $g$ and the cost functions, the dual $\mathcal{J}_{\mathrm{D}}(\delta)$ in (ref) for interior points $\delta \in \mathbb{R}_{++}^2$ and the dual for boundary points may not be the same. To illustrate, plugging in $\delta_2 = 0$ in the dual form for interior points in (ref), we obtain
It is different from the dual $\mathcal{J}_{\mathrm{D}}(\delta_1,0)$ for $\delta_1>0$, since
When the function $g$ and the cost functions satisfy assumptions in (ref), the dual $\mathcal{J}_{\mathrm{D}}(\delta)$ in (ref) for interior points $\delta \in \mathbb{R}_{++}^2$ and the dual for boundary points are the same so that
for all $\delta \in \mathbb{R}_{+}^2$.
Let $\phi_\lambda : \mathcal{V} \times \mathcal{S}\rightarrow \mathbb{R} \cup \{- \infty \}$ be \[
\] where $v = (s_1, s_2)$, $s^\prime = (y^\prime_0,y^\prime_1,x^\prime)$, $s^\prime_\ell = (y^\prime_\ell, x^\prime)$ and $s_\ell=(y_\ell, x_\ell)$. Define the function $f_\lambda : \mathcal{V} \rightarrow \overline{\mathbb{R}}$ associated with $f$ as \[ f_\lambda(v) := \sup_{s^\prime \in \mathcal{S} } \phi_\lambda (v, s^{\prime}). \] Similarly, we define $f_{\lambda, 1}: \mathcal{V} \to \overline{\mathbb{R}}$ and $f_{\lambda, 2}: \mathcal{V} \to \overline{\mathbb{R}}$ as follows:
in which $s_1 = (y_1, x_1)$ and $s_2 = (y_2, x_2)$. The dual problem $\mathcal{J}(\delta)$ corresponding to the primal problem $\mathcal{I}(\delta)$ is defined as follows:
An interesting feature of the dual for overlapping marginals is that it involves marginal problems with non-overlapping marginals, i.e., $\sup_{\varpi \in \Pi(\mu_{13}, \mu_{23})} \int_{\mathcal{V} }f_\lambda (v) d \varpi (v)$, although the uncertainty set in the primal problem involves overlapping marginals. Compared with the non-overlapping marginals case, overlapping marginals in the uncertainty set make the relevant consistent product marginal system in the verification of the existence of a joint measure more complicated, see the proof of (ref). Nonetheless, the non-overlapping marginals in the dual allow us to apply (ref) to the marginal problem involving $f_\lambda$, $f_{\lambda, 1}$ and $f_{\lambda, 2}$, see (ref) in (ref).
Under the assumptions in (ref), we have
for all $\delta \in \mathbb{R}_{+}^2$.
Let $\mathcal{S}_1=\mathbb{R}$, $\mathcal{S}_2=\mathbb{R}$, $\mu_1\in \mathcal{P}(\mathcal{S}_1)$, and $\mu_2\in \mathcal{P}(\mathcal{S}_2)$. Further, let $Z=S_1+S_2$, where $S_1,S_2$ are random variables whose probability measures are $\mu_1, \mu_2$ respectively. For a given $z\in \mathbb{R}$, let $F_{Z}(z)=\mathbb{E}_o[g(S_1, S_2)]$, where $g(s_1, s_2)= \mathds{1}\left\{s_1 + s_2 \le z\right\}$.
Sharp bounds on the quantile function $F^{-1}_{Z}\left( \cdot\right) $ are established in Makarov1982) and referred to as the Makarov bounds. Inverting the Makarov bounds lead to sharp bounds on the distribution function $F _{Z}\left( z\right)$, see Ruschendorf_1982_RV_Maximum and Frank_1987_best_bounds. They are given by
Since the quantile bounds first established in Makarov1982) and the above distribution bounds are equivalent, we also refer to the latter as Makarov bounds. Makarov bounds have been successfully applied in distinct areas. For example, the upper bound on the quantile of $Z$ is known as the worst VaR of $Z$, see Embrechts2003UsingCopulaeBound, Embrechts_2005_WorstVaR; Makarov bounds are also used to study partial identification of distributional treatment effects when the treatment assignment mechanism identifies the marginal measures of the potential outcomes such as in (ref), see Fan_2009_PartialIdentification,Fan_2010_SharpBounds,Fan_2012_CI_quantiles, FAN_2009, Fan2017, Ridder_2007, and Firpo_2019.
Applying (ref), we extend Makarov bounds to allow for possible misspecification of the marginal measures and call the resulting bounds Wasserstein distributionally robust Makarov bounds.
We note that $g_{\lambda}(v)$ is bounded and continuous in $v$, and convex in $\lambda$, and $\Pi(\mu_1, \mu_2)$ is compact. Applying Fan_1953_MinimaxThms's minimax theorem, we can interchange the order of $\inf$ and $\sup$ in the dual in the above corollary and get
This expression is very insightful, where the inner infimum term characterizes possible deviations of the true marginal measures from the reference measures.
In this section, we assume that all the reference measures belong to appropriate Wasserstein spaces and prove finitness of the W-DMR-MP and existence of an optimizer.
For non-overlapping case, we establish the following result.
The inequality in (ref) is a growth condition on the function $g$. It extends the growth condition in yue2022linear for W-DMR to our W-DMR with non-overlapping narginals.
For the overlapping case, the following result holds.
The growth condition (ref) on the function $f$ extends the growth condition in yue2022linear for W-DMR. When \[ \boldsymbol{d}_{\mathcal{S}_\ell} ( (y_\ell, x), (y^{\prime}_\ell, x^\prime) ) = \boldsymbol{d}_{\mathcal{Y}_\ell} ( y_\ell,y^{\prime}_\ell ) + \boldsymbol{d}_{\mathcal{X}} ( x, x^{\prime} ), \] condition (ref) is satisfied if and only if there exist $s^\star := (y_1^\star, y_2^\star, x^\star)$ and a constant $M > 0$ such that \[ f(s) \leq M \left[ 1 + \boldsymbol{d}_{\mathcal{Y}_1} (y_1, y_1^\star )^{p_1}+ \boldsymbol{d}_{\mathcal{Y}_2} (y_2, y_2^\star )^{p_2}+ \boldsymbol{d}_{ \mathcal{X} }(x, x^\star )^{p_1 \wedge p_2} \right], \] for all $s = (y_1,y_2,x) \in \mathcal{S}$.
Examples of proper metric spaces include finite dimensional Banach spaces and complete Riemannian manifolds, see yue2022linear.
(ref) imply that $\Sigma_{\mathrm{D}}(\delta)$ and $\Sigma(\delta)$ are weakly compact, see (ref) in Appendix C. Given weak compactness of the uncertainty sets $\Sigma_{\mathrm{D}}(\delta)$ and $\Sigma(\delta)$, it is sufficient to show that the mapping: $\gamma \to \int g d\gamma$ is upper semi-continuous over $\gamma \in \Sigma_{\mathrm{D} }(\delta)$ for the non-overlapping case, and the mapping: $\gamma \to \int f d\gamma$ is upper semi-continuous over $\gamma \in \Sigma(\delta)$ for the overlapping case. In (ref) below, we provide conditions for $g$ and $f$ ensuring upper semi-continuity of each map and thus the existence of optimal solutions for $\mathcal{I}_{\mathrm{D}}(\delta)$ and $\mathcal{I}(\delta)$.
In some applications, such as the partial identification of treatment effects introduced in (ref), the identified sets of $\theta_{\mathrm{D}o} := \mathbb{E}_{o}[g(S_1, S_2)]$ and $\theta_o := \mathbb{E}_{o}[f(S)]$ are of interest, where $S$ is a random variable whose probability measure belongs to $\Sigma(\delta)$, and $S_1$ and $S_2$ are random variables whose joint probability measure belongs to $\Sigma_{D}(\delta)$. They are:
By applying finiteness and existence results, we show below that under mild conditions, the identified sets $\Theta_{\mathrm{D}}(\delta)$ and $\Theta(\delta)$ are both closed intervals.
The strong duality in Section 3 can be used to evaluate the lower and upper bounds.
In this section, we establish continuity of the W-DMR-MP functions $\mathcal{I}_{\mathrm{D} }(\delta)$ and $\mathcal{I}(\delta)$ for all $\delta\in\mathbb{R}_+^2$ under similar conditions to those in zhang2022simple. Compared with zhang2022simple, our analysis is more involved, because the boundary in our case includes not only the origin $(0,0)$ but also $(\delta_1,0)$ and $(0,\delta_2)$ for all $\delta_1 > 0$ and $\delta_2 > 0$.
(ref) implies that under (ref), $\mathcal{I}_{\mathrm{D} }(\delta)$ is a concave function for $\delta \in \mathbb{R}_{+}^2$ and hence is continuous on $\mathbb{R}_{++}^2$. We provide the main assumption for the continuity of $\mathcal{I}_{\mathrm{D}}(\delta)$ on $\mathbb{R}_{+}^2$ in this subsection.
The function $\Psi$ in (ref) plays the role of the modulus of continuity of $g$. To illustrate, consider the following example.
Two implications follow. First, under (ref) and (ref), \[\mathcal{I}_{\mathrm{D} }(0) = \sup_{\gamma \in \Pi(\mu_{1}, \mu_{2})} \int_{\mathcal{V} } g\, d\gamma. \] Continuity facilitates sensitivity analysis as $\delta$ approaches zero; Second, under the assumptions in (ref), we have
for all $\delta \in \mathbb{R}_{+}^2$. As a result, the dual $\mathcal{J}_{\mathrm{D} }(\delta)$ in (ref) is continuous for all $\delta\in\mathbb{R}_+^2$.
(ref) implies that under (ref), $\mathcal{I}(\delta)$ is a concave function for $\delta \in \mathbb{R}_{+}^2$ and hence is continuous on $\mathbb{R}_{++}^2$. We provide the main assumption for the continuity of $\mathcal{I}(\delta)$ on $\mathbb{R}_{+}^2$ below.
To simplify the technical analysis, we maintain (ref) in this section. Since the metrics in $\mathcal{Y}_1$ and $\mathcal{Y}_2$ are not specified, we introduce an auxiliary function $\rho_\ell$ from $\mathcal{Y}_\ell \times \mathcal{Y}_\ell$ to $\mathbb{R}_+$ induced by the cost function $c_\ell$, $\ell = 1, 2$.
We now introduce the main assumption on $f$.
Like (ref), (ref) depends on the cost functions $c_1,c_2$. It also depends on the auxiliary functions $\rho_1,\rho_2$. The functions $\Psi_1, \Psi_2$ play the role of the modulus of continuity.
Like the non-overlapping case, two implications follow. First, under (ref) and (ref), \[ \mathcal{I}(0)= \sup_{\gamma \in \mathcal{F}(\mu_{13}, \mu_{23}) } \int_\mathcal{S} f\, d \gamma. \] Continuity facilitates sensitivity analysis as $\delta$ approaches zero; Second, under the assumptions in (ref), we have
for all $\delta \in \mathbb{R}_{+}^2$. As a result, the dual $\mathcal{J}(\delta)$ in (ref) is continuous for all $\delta\in\mathbb{R}_+^2$.
In this section, we apply the results in Sections 3-5 to the examples introduced in Section 2.
In addition to characterizing $\Theta(\delta)$ introduced in Section 2, we also study the identified set for $\theta_{Do}= \mathbb{E}_o[f(Y_1,Y_2)]$ without using the covariate information: \[ \Theta_{\mathrm{D}}(\delta) := \left\{ \int_{\mathcal{Y}_1 \times \mathcal{Y}_2 } f(y_1, y_2) \, d \gamma(y_1 ,y_2) : \gamma \in \Sigma_{\mathrm{D} }(\delta) \right\}, \] where
in which $\boldsymbol{K}_{Y_1}$ and $\boldsymbol{K}_{Y_2}$ are the optimal transport costs associated with cost functions $c_{Y_1}$ and $c_{Y_2}$, respectively.
When $f$ is continuous and conditions in (ref) are satisfied, the identified sets $\Theta_{\mathrm{D}}(\delta)$ and $\Theta(\delta)$ are both closed intervals with upper limits given by W-DMR for non-overlapping and overlapping marginals respectively. This allows us to apply our duality results in Section 3 to evaluate and compare $\Theta_{\mathrm{D}}(\delta)$ and $\Theta(\delta)$.
Let $\mathcal{I}_{\mathrm{D} }(\delta)$ and $\mathcal{I}(\delta)$ denote the upper bounds of $\Theta_{\mathrm{D}}(\delta)$ and $\Theta(\delta)$, respectively, where \[ \mathcal{I}_{\mathrm{D} }(\delta) = \sup_{\gamma \in \Sigma_{\mathrm{D} }(\delta)} \int_{\mathcal{Y}_1 \times \mathcal{Y}_2} f(y_1, y_2) \, d \gamma(y_1, y_2) \text{ and } \mathcal{I}(\delta) = \sup_{\gamma \in \Sigma(\delta)} \int_{\mathcal{S}} f(y_1, y_2) \, d \gamma(y_1, y_2, x). \] (ref) establishes robust versions of existing results on the identified sets of treatment effects under (ref), see Fan2017. Sensitivity to deviations from (ref) can be examined via $\Theta_{\mathrm{D}}(\delta)$ and $\Theta(\delta)$ by varying $\delta$. For example, when $f$ satisfies assumptions in (ref), $\mathcal{I}(\delta) $ and $\mathcal{I}_{\mathrm{D} }(\delta)$ are continuous on $\mathbb{R}_+^2$. As a result, \[\lim_{\delta\rightarrow 0 }\mathcal{I}(\delta)=\mathcal{I}(0) \quad \text{and} \quad \lim_{\delta\rightarrow 0 }\mathcal{I}_{\mathrm{D} }(\delta)=\mathcal{I}_{\mathrm{D} }(0). \]
For a general function $f$, the lower and upper limits of the identified sets $\Theta_{\mathrm{D}}(\delta)$ and $\Theta(\delta)$ need to be computed numerically. When $f$ is additively separable, we show that duality results in Section 3 simplify the evaluation of $\Theta_D(\delta)$ and $\Theta(\delta)$. Since the lower bounds of $\Theta_D(\delta)$ and $\Theta(\delta)$ can be computed in a similar way by applying duality to $-f(y_1, y_2)$, we omit details for the lower bounds.
To avoid tedious notation, we also treat $f$ as a function from $\mathcal{Y}_1 \times \mathcal{Y}_2$ to $\mathbb{R}$. Under (ref), it is easy to show that
where $(f_\ell)_{\lambda_\ell}: \mathcal{Y}_\ell \rightarrow \mathbb{R}$ is given by \[ (f_\ell)_{\lambda_\ell}(y_\ell) = \sup_{ y_\ell^\prime \in \mathcal{Y}_\ell } \left\{ f_\ell(y_\ell^{\prime})-\lambda_\ell c_{Y_\ell}(y_{\ell}, y_{\ell}')\right\}. \] That is, when $f$ is an additively separable function, the W-DMR for non-overlapping marginals is the sum of two W-DMRs associated with the marginals regardless of the cost functions.
Depending on the cost functions, the W-DMR for overlapping marginals may be different from the sum of two W-DMRs associated with the marginals.
It is easy to verify that $c_{Y_\ell}(y_\ell, y_\ell^\prime) = 0$ if and only if $y_\ell = y_\ell^\prime$.
This proposition implies that for separable cost functions, the W-DMR for overlapping marginals equals the W-DMR for non-overlapping marginals with cost function $c_{Y_\ell}(y_\ell, y_\ell^\prime)$. As a result, the covariate information does not help shrink the identified set.
Suppose $f(y_1, y_2) = y_2 - y_1$ and $c_{\ell}((y, x), (y_{\ell}, x_{\ell})) = |y - y'|^{2} + \| x_\ell - x_\ell' \|^2$ for $\ell = 1, 2$. Let $\tau_{ATE}=\mathbb{E}[Y_2 - Y_1]$. Then (ref) implies that the upper bound on $\tau_{ATE}$ is given by
In the rest of this section, we demonstrate that when (ref) is violated, the W-DMR for overlapping marginals may be smaller than the W-DMR for non-overlapping marginals and, as a result, $\Theta(\delta)$ is a proper subset of $\Theta_D(\delta)$.
Consider the squared Mahalanobis distance with respect to a positive definite matrix. That is,
where $V_\ell =
$ is a positive definite matrix. It is easy to show that
where $s_\ell = (y_\ell, x_\ell)$ and $s_{\ell}' = (y_\ell', x_\ell')$.
(ref) and (ref) imply that for non-separable Mahalanobis cost functions, the information in covariates may help shrink the identified set since $\mathcal{I}_{\mathrm{D}}(\delta) < \mathcal{I}(\delta)$ for some $\delta$ under mild conditions. (ref) also implies that (i) $\mathcal{I}(0) = \mathcal{I}_{\mathrm{D}}(0) = \mathbb{E}[Y_2] - \mathbb{E}[Y_1]$; (ii) $\mathcal{I}(\delta_1, 0) = \mathcal{I}_{\mathrm{D}}(\delta_1, 0)$ and $\mathcal{I}(0, \delta_2) = \mathcal{I}_{\mathrm{D}}(0, \delta_2)$ for all $\delta_1 \ge 0$ and $\delta_2 \ge 0$.
Recall that
where
Consider the following cost function $c_{\ell}$ for $\ell = 1, 2$:
where $s_\ell = (y_\ell, x_\ell)$, $s_\ell' = (y_\ell', x_\ell')$, and $ c_{Y_1}(y_1, y_1')$ and $c_{Y_2}(y_2, y_2')$ are cost functions for $Y_1$ and $Y_2$, respectively, and $b(x, x')$ is some function on the space $\mathcal{X}$ satisfying (ref). When $b(x, x') = \infty \mathds{1}\{x \ne x'\}$, $\mathbb{P}(X = X') = 1$ for any probability measure in uncertainty set.
adjaho2022 establishes strong duality for $\mathrm{RW}_0(d)$ under several cost functions. For comparison purposes, we restate the following Proposition in adjaho2022 which allows distributional shifts in covariate $X$.
This proposition implies that $\mathrm{RW}_0(d)$ depends on the choice of the reference measure $\mu$. Since only the marginals $\mu_{13}$ and $\mu_{23}$ are identified under (ref), adjaho2022 suggest three possible choices for $\mu$ by imposing specific dependence structures on $\mu$:
Section 4.3.1 in adjaho2022 shows that their robust welfare function $\mathrm{RW}_0(d)$ is minimized when $Y_1$ and $Y_2$ are perfectly negatively dependent conditional on $X = x$.
The following proposition evaluates $\mathrm{RW}(d)$ via the duality result in Section 3 and compares it with $\mathrm{RW}_0(d)$.
Part (ii) of the above proposition implies that $\mathrm{RW}(d) \le \mathrm{RW}_0(d)$ for any reference measure $\mu \in \mathcal{F}(\mu_{13}, \mu_{23})$.
We revisit the logit model in (ref) and make the following assumption.
Under this assumption, $X|D=1$ has the same distribution as $X|D=0$ and the empirical distributions of the two data sets are consistent estimators of the population reference measures for $(Y_1, X)$ and $(Y_2, X)$.
Suppose (ref) hold. Then (ref) implies that for all $\delta > 0$,
where
with $v = (y_1, x_1, y_2, x_2)$.
Let $\widehat{\mu}_{13}$ and $\widehat{\mu}_{23}$ denote the empirical measures based on the two data sets. The dual form of $\mathcal{I}(\delta)$ can be estimated by
A direct consequence of Kellerer:1984to is that
where the last expression reduces to the dual in awasthi2022distributionally for the cost functions \[ c_1((y_1, x), (y_1', x')) = \| x - x' \|_{p} + \kappa_1 | y_1 - y_1'| \quad \text{and} \] \[ c_2((y_2, x), (y_2, x')) = \| x - x' \|_{p} + \kappa_2 \| y_2 - y_2' \|_{p'}. \]
Sections 2-6 present a detailed study of W-DMR with two marginals. In this section, we briefly introduce W-DMR with more than two marginals or multi-marginals and discuss strong duality for non-overlapping and overlapping marginals.\footnote{For multi-marginals, the collection of given marginals can be more complicated than the non-overlapping and overlapping marginals (see ruschendorf1991bounds, Embrechts_2010 and Doan2015), we leave a complete treatment of the W-DMR with multi-marginals in future work.} Applications include extension of risk aggregation in (ref) to any finite number of individual risks and robust treatment choice in (ref) to multi-valued treatment.
Let $\mathcal{V}:=\prod_{\ell \in [L]} \mathcal{S}_\ell$ for Polish spaces $\mathcal{S}_\ell$ for $\ell \in [L]$, and $\mu_\ell$ be a probability measure on $(\mathcal{S}_{\ell}, \mathcal{B}_{\mathcal{S}_\ell})$. Let $\Pi(\mu_1, \dotsc, \mu_L)$ be the set of all possible couplings of $\mu_1, \dotsc, \mu_L$. Further, let $g:\mathcal{V}\rightarrow \mathbb{R}$ be a measurable function satisfying the following assumption.
For any $\gamma \in \mathcal{P}(\mathcal{V})$, let $\gamma_{\ell}$ denote the projection of $\gamma$ on $\mathcal{S}_{\ell}$ for $\ell \in [L]$. The W-DMR with non-overlapping multi-marginals is formulated as
where $\Sigma_{\mathrm{D} }(\delta)$ is the uncertainty set defined as
in which $\delta = (\delta_1, \dotsc, \delta_L) \in \mathbb{R}_{+}^L$ is the radius of the uncertainty set.
For a generic vector $v \in \mathbb{R}^L$ and $A \subset [L]$, we write $v_A = (v_{A, 1}, \dotsc, v_{A, L}) \in \mathbb{R}^{L}$ as follows:
We also define $\tilde{c}_\ell: \mathcal{S}_{\ell} \times \mathcal{S}_{\ell} \to \mathbb{R}_{+} \cup \{\infty\}$ as:
For a function $g: \mathcal{V} \rightarrow \mathbb{R}$ and $\lambda := (\lambda_1, \dotsc, \lambda_L) \in \mathbb{R}^L_{+}$, we define the function $g_{\lambda, A}: \mathcal{V} \rightarrow \mathbb{R} \cup \{\infty \}$ as
with $v:= (s_1, \dotsc, s_L)$ and $v^\prime := (s_1^\prime, \dotsc, s_L^\prime)$.
In practice, the dual in (ref) involves the computation of the multi-marginal problem, $\sup_{\pi \in \Pi(\mu_1, \dotsc, \mu_L)} \int_{\mathcal{V}} g_{\lambda} \, d \pi$, see Pass_2010,Pass_2012,Pass_2015,von_Lindheim_2022,Nenna_2022,mehta2023efficient for detailed studies of properties and computation of multi-marginal problems for specific functions $g_{\lambda}$. For general possibly non Borel-measurable $g_{\lambda}$, the strong duality in Kellerer:1984to could be applied. The established result is stated in (ref) in (ref).
Let $\mathcal{S} := \left(\prod_{\ell \in [L]} \mathcal{Y}_\ell \right) \times \mathcal{X}$, where $\mathcal{Y}_\ell$ for $\ell \in [L]$ and $\mathcal{X}$ are Polish spaces. Let $\mathcal{S}_\ell := \mathcal{Y}_\ell \times \mathcal{X}$ for $\ell \in [L]$. Let $\mu_{\ell,L+1} \in \mathcal{P}(\mathcal{S}_\ell)$ for $\ell \in [L]$ be such that the projections of $\mu_{\ell, L+1}$ on $\mathcal{X}$ are the same for $\ell \in [L]$. We call the Fr\'{e}chet class of all probability measures on $\mathcal{S}$ having marginals $\left(\mu_{1,L+1} \right)_{\ell \in [L]}$ the Fr\'{e}chet class with overlapping marginals and denote it as $\mathcal{F}\left(\mathcal{S}; (\mu_{\ell,L+1})_{\ell \in [L]} \right) := \mathcal{F}\left( \left(\mu_{\ell,L+1} \right)_{\ell \in [L]} \right)$. This class is the star-like system of marginals in ruschendorf1991bounds and Embrechts_2010, see also Doan2015.
Moreover, let $f:\mathcal{S} \rightarrow \mathbb{R}$ be a measurable function satisfying the following assumption.
For any $\gamma \in \mathcal{P}(\mathcal{S})$, let $\gamma_{\ell, L+1}$ denote the projection of $\gamma$ on $\mathcal{Y}_{\ell} \times \mathcal{X}$ for $\ell \in [L]$. Similar to the two marginals case, the W-DMR with overlapping multi-marginals is defined as
where $\Sigma(\delta)$ is the uncertainty set defined as
in which $\delta = (\delta_1, \dotsc, \delta_L) \in \mathbb{R}_{+}^L$ is the radius of the uncertainty set.
For a function $f: \mathcal{V} \rightarrow \mathbb{R}$, $\lambda := (\lambda_1, \dotsc, \lambda_L) \in \mathbb{R}^L_{+}$, and $A \subset [L]$, we define the function $f_{\lambda, A}: \mathcal{V} \to \overline{\mathbb{R}}$ as follows:
where $v = (s_1, \dotsc, s_L)$, $s^\prime = (y^\prime_1, \dotsc, y^\prime_L, x^\prime)$, $s^\prime_\ell = (y^\prime_\ell, x^\prime)$ and $s_\ell=(y_\ell, x_\ell)$, and
Similar to the nonoverlapping case, strong duality holds for the inner multi-marginal problem under additional conditions. The result is stated in (ref) of (ref).
We apply strong duality to multi-valued treatment in Kido2022. Let $d: \mathcal{X} \rightarrow [L]$ be a policy function or treatment rule on $\mathcal{X}$ and $Y_\ell \in \mathbb{R}$ denote the potential outcome under the treatment $\ell$ for $\ell \in [L]$. Consider the policy function defined as
Kido2022 introduces the following robust welfare function.
where the uncertainty set $\Sigma_{\mathrm{M}}(\delta_0)$ is based on the conditional distribution of $(Y_\ell)_{\ell \in [L]}$ given $X$:
in which the cost function $c$ associated with $\boldsymbol{K}$ is
Note that the uncertainty set $\Sigma_{\mathrm{M}}(\delta_0)$ does not allow any potential shift\footnote{Kido2022 mentions the possibility of allowing for covariate shift by incorporating uncertainty sets in Mo2020,Zhao2019 for the distribution of the covariate in future work. } in $X$. When $Y_1, \dotsc, Y_L$ are unbounded, Kido2022 shows that
We apply W-DMR for overlapping marginals with the following cost function:
and define a robust welfare function as
(ref) is an extension of (ref).
In this paper, we have introduced W-DMR in marginal problems for both non-overlapping and overlapping marginals and established fundamental results including strong duality, finiteness of the proposed W-DMR, and existence of an optimizer at each radius. We have also shown continuity of the W-DMR-MP as a function of the radius. Applicability of the proposed W-DMR in marginal problems and established properties is demonstrated via distinct applications when the sample information comes from multiple data sources and only some marginal reference measures are identified. To the best of the authors' knowledge, this paper is the first systematic study of W-DMR in marginal problems. Many open questions remain including the structure of optimizers of W-DMR for both non-overlapping and overlapping marginals, efficient numerical algorithms, and estimation and inference in each motivating example. Another useful extension is to consider objective functions that are nonlinear in the joint probability measure such as the Value-at-Risk of a linear portfolio of risks in Puccetti_2012 and robust spectral measures of risk in Ghossoub_2023,Ennaji2022.
\printbibliography