Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
95,715 characters · 20 sections · 123 citation commands
Partial Identification in Moment Models with Incomplete Data---A Conditional Optimal Transport Approach
Many parameters of interest in economic models satisfy a finite number of moment conditions. The seminal paper of hansen1982large developed a systematic approach to identification, estimation, and inference when all economic variables in the moment condition model are observable in one data set. Since then, the generalized method of moments has become an influential tool in empirical research in economics. However, researchers often do not have access to a single data set that contains observations on all variables and must rely on multiple data sources that cannot be matched; see the discussion in Section (ref). This paper develops the first systematic method to study the identification of parameters in moment condition models using multiple data sets.
Specifically, consider the following moment model,
where $Y_{1}\in\mathcal{Y}_{1}\subseteq\mathbb{R}^{d_{1}},Y_{0}\in\mathcal{Y}_{0}\subseteq\mathbb{R}^{d_{0}}$, and $X\in\mathcal{X}\subseteq\mathbb{R}^{d_{x}}$ denote three distinct random variables; $\theta^{\ast}\in\Theta\subseteq\mathbb{R}^{d_{\theta}}$ is the true unknown parameter; $m$ is a known $k$ dimensional moment function; and $\mathbb{E}_{o}$ denotes the expectation with respect to the true but unknown joint distribution $\mu_{o}$ of $\left(Y_{1},Y_{0},X\right)$.
This paper studies identification of $\theta^{\ast}$ (and known functions of $\theta^{\ast}$)\footnote{When $\delta^{\ast}=G\left(\theta^{\ast}\right)$ is the parameter of interest for some function $G\left(\cdot\right)$, we say $G\left(\cdot\right)$ is known if either its functional form is known or (point) identifiable.} in a prevalent case of incomplete data in empirical research, where the relevant information on $\left(Y_{1},Y_{0},X\right)$ is contained in two datasets, and the units of observations in different datasets cannot be matched: the first dataset identifies the distribution of $\left(Y_{1},X\right)$ and the second dataset identifies the distribution of $\left(Y_{0},X\right)$. We develop the first general unified methodology for studying the sharp identified set (identified set, hereafter) of $\theta^{\ast}$ (and known functions of $\theta^{\ast}$) in the general moment model ((ref)) with incomplete data. Our method allows for (non-additively separable) moment function of any finite dimension, parameter $\theta^{*}$ of any finite dimension, and two arbitrary random variables $Y_{1}$ and $Y_{0}$, each of any finite dimension.
\paragraph*{General Contributions}
The available sample information in the incomplete data scenario that we study in this paper allows identification of the distribution functions of $\left(Y_{1},X\right)$ and $\left(Y_{0},X\right)$, but does not identify the distribution of $\left(Y_{1},Y_{0},X\right)$ without additional assumptions on the joint distribution of $\left(Y_{1},Y_{0},X\right)$. For the general moment model ((ref)), we establish a characterization of the identified set of $\theta^{\ast}$ in terms of a continuum of inequalities involving conditional optimal transport (COT hereafter).\footnote{Conditional optimal transport is an optimal transport between two conditional measures; see carlier2016vector and hosseini2025conditional for rigorous accounts of COT. For expositional simplicity, we often use optimal transport even though it is a conditional optimal transport in this paper.}
For an important class of affine moment models, where the moment function $m\left(\cdot;\theta\right)$ can be written as an affine transformation of $\theta$ with the expectation of the linear transformation matrix being identifiable, but not the translation vector, we show that by applying our general characterization, the identified set for $\theta^{\ast}$ in this class of models is characterized by a continuum of linear inequalities and is thus convex. Furthermore, we derive the expression of the support function of the convex identified set and show that the value of the support function in each direction can be obtained by solving a COT problem. As we demonstrate using several running examples, our novel characterization through COT allows us to make use of theoretical and computational tools for optimal transport (OT hereafter; e.g., rachev2006mass, villani2008optimal, santambrogio2015optimal, and peyre2019computational) to further simplify the characterization and develop computationally more tractable identified sets for specific models than the existing ones in the literature.\footnote{ Other works that have employed optimal transport in partial identification analysis in different set-ups from ours include galichon2011set, galichon2016optimal, and d2025partially. We refer interested readers to molinari2020microeconometrics for a recent survey.} The well-known Fr\'echet problem reviewed in Section (ref) is a simple example of an affine moment model to which our characterization of the identified set applies. We show in Section (ref) that bounding causal effect in meango2025combining is an example of the Fr\'echet problem and the COTs in our characterization provide closed-form expressions for the dual optimizations in Corollary 1 in meango2025combining.\footnote{Mathematically, the bounds problem in meango2025combining is similar to that in fan2014identifying, fan2016estimation. All are variations of the short and long regression in cross2002regressions; see also molinari2006generalization. }
To take full advantage of our results for affine moment models, we propose a two-step procedure to obtain the identified set of the parameter of interest. In the first step of this procedure, we construct an auxiliary parameter that satisfies an affine moment model and establish its identified set. The auxiliary parameter is chosen so that the parameter of interest can be written as a known transformation of the auxiliary parameter. In the second step, we compute the identified set of the parameter of interest from the known transformation of the convex identified set obtained in the first step. The two-step procedure substantially broadens the applicability of our results, since it applies to parameters of interest that may not even satisfy the moment model ((ref)) as we illustrate via the algorithmic measures in Examples (ref) and (ref).
\paragraph*{Results for Running Examples}
To demonstrate the versatility and advantages of our general method and the proposed two-step procedure, we study three examples of partially identified parameters with incomplete data and establish their identified sets using our approach. We show that our method leads to novel and significantly improved results in all examples compared with those in the literature.
Example (ref) is a linear projection model. We consider two different types of sample information, where the first extends that in d2024linear and hwang2023bounding; the second data structure extends that in kitagawa2023linear. The moment function is, in general, not affine, and the identified set of the parameter of interest may not be convex. We employ the proposed two-step procedure, where like hwang2023bounding, we express the parameter of interest as a known function of an unknown parameter $\theta^{\ast}$ satisfying an affine moment model. In the first step, we show that the identified set of $\theta^{\ast}$ is convex and its support function is given by an integral of COT with quadratic costs, the most studied OT problem; see brenier1991polar, santambrogio2015optimal, and peyre2019computational. In the second step, we obtain the identified set of the parameter of interest from that of $\theta^{*}$. Using Proposition 2.17 in santambrogio2015optimal, we show that when specialized to the model in d2024linear, our identified set reduces to that in d2024linear. When specialized to hwang2023bounding, our result leads to the identified set that is a proper subset of the superset proposed in hwang2023bounding. Based on our characterization of the identified set, we also propose a computationally simpler outer set that is also a subset of the superset in hwang2023bounding.
Examples (ref) and (ref) are two algorithmic measures first studied in kallus2022assessing (KMZ, hereafter) in the data combination setup. The measures themselves do not satisfy the moment model ((ref)). To apply our results, we first express them as known transformations of some auxiliary parameter $\theta^{*}$, such that $\theta^{*}$ satisfies model ((ref)) with an affine moment function. Consequently, the two-step procedure applies, and the results we establish for the identified set of affine models apply to $\theta^{*}$ in both examples. Specifically, for Example (ref) on the demographic disparity (DD) measure, we provide a closed-form expression for the support function of the identified set for any collection of DD measures and\emph\textit{\emph{any}} number of protected classes. Furthermore, we show that the identified set is a polytope with a closed-form expression for its vertices. For the specific collection of DD measures studied in KMZ, our closed-form expression solves the infinite dimensional linear programming formulation in KMZ. We analyze the theoretical time complexity of both our support function and vertex-based procedures and show that they are significantly faster than KMZ's method. Using the empirical example in KMZ, we demonstrate that the improvement can be over 15,000-fold.
For the true-positive rate disparity (TPRD) measure in Example (ref), KMZ propose to construct the convex hull of the identified set. They characterize the support function of the convex hull through a computationally costly, NP-hard non-convex optimization problem. Choosing the auxiliary parameter vector $\theta^{\ast}$ cleverly, we express the identified set for any finite number of TPRD measures as a continuous map of the identified set for $\theta^{\ast}$, where the identified set for $\theta^{\ast}$ is convex. Consequently, the identified set for any finite number of TPRD measures is connected, and any single TPRD measure is a closed interval regardless of the number of protected classes. For a single TPRD measure, we provide closed-form expressions for the lower and upper bounds of the closed interval extending the result in KMZ developed for two protected classes. For multiple TPRD measures, we demonstrate that the support function of the identified set for the auxiliary parameter $\theta^{\ast}$ can be computed through a conditional partial optimal transport problem. Technically, we show that the conditional partial OT problem always admits a solution with monotone support. Exploring the monotone support of the solution, we develop a novel algorithm, called Dual Rank Equilibration AlgorithM (DREAM), for computing the identified set for $\theta^{\ast}$ and for any collection of TPRD measures. We establish its theoretical time complexity and demonstrate that DREAM is faster than linear programming using the empirical example in KMZ.
Data combination problems are pervasive in applied work. Examples include studies of the effect of age at school entry on the years of schooling by combining data from the US censuses in 1960 and 1980 (angrist1992effect); differentiated product demand using both micro and macro survey data (berry2004differentiated); studies of long-run returns to college attendance using PSID/NLSY and Addhealth data (fan2014identifying,fan2016estimation); repeated cross-sectional data in difference-in-differences designs (fanManzanares2017partial); evaluations of long-term treatment effects combining experimental and observational data (athey2019surrogate); algorithmic fairness analysis where decision outcomes and protected attributes are recorded in separate datasets (kallus2022assessing); study of neighborhood effects, where longitudinal residence information and information on individual heterogeneity are contained in separate datasets (hwang2023bounding); inter-generational income mobility, where linked income data across generations are often unavailable (santavirta2024name); consumption research, where income and consumption are often measured in two different datasets (crossley2022regression); and counterfactual analyses of actual choice using stated and revealed preference data (meango2025combining).
Existing works in the incomplete data case fall into two broad categories. The first category considers the case where the moment function is additively separable in $Y_{1}$ and $Y_{0}$ and establishes point identification of $\theta^{\ast}$ under weak conditions. See, e.g., angrist1992effect, hahn1998role, chen2008semiparametric, graham2016efficient, and athey2019surrogate.
The second category develops partial identification results for specific models. Except for the running examples we describe below, most existing work studying partial identification under data combination can be reformulated as examples of a special case of model ((ref)), where the moment function is
for a known function $h$ of dimension $k=1$. Equivalently, $\theta^{\ast}=\mathbb{E}_{o}\left[h\left(Y_{1},Y_{0}\right)\right]$. For univariate $Y_{1}$ and $Y_{0}$ with given marginals, computing lower and upper bounds for $\theta^{*}$ for a known function $h$ has a long history in probability literature, and is referred to as the (general) Fr\'{e}chet problem; see ridder2007econometrics and Section 3 in fan2014copulas for a detailed discussion of this problem, including early references and its relation to copulas. These bounds have been used to establish identified sets for $\theta^{*}$ in a broad range of applications, including distributional treatment effects in program evaluation and bivariate option pricing in finance; see e.g., fan2010sharp,fan2012confidence, fan2010partial, fan2017partial, and firpo2019partial. More recently, fan2023partial study identification in model ((ref)) with general moment function $m$ and univariate $Y_{1}$ and $Y_{0}$ via the copula approach and propose a valid inference procedure.\footnote{fan2023partial include a point-identified nuisance parameter $\gamma^{\ast}$ in their moment function to separate the parameter of interest $\theta^{\ast}$ from $\gamma^{\ast}$. Model ((ref)) can be accommodated to having the nuisance parameter by appending $\gamma^{\ast}$ to $\theta^{\ast}$.} However, the copula approach is limited to univariate $Y_{1}$ and $Y_{0}$. For multivariate $Y_{1}$ and $Y_{0}$, the lower and upper bounds on $\theta^{*}$ are defined by COTs with ground cost function $h$ or $-h$. lin2025estimation propose consistent estimation using the primal formulation of COT, while ji2023model study estimation and inference using their dual formulation.
\paragraph*{Organization} Section (ref) introduces our sample information and three motivating examples. In Section (ref), we first characterize the identified set for $\theta^{\ast}$ in our model ((ref)) in terms of a continuum of inequalities, and then study the properties of the identified set for the special class of affine moment models. Finally, we apply our characterization to the Fr\'{e}chet problem in ((ref)) and revisit meango2025combining. Sections (ref) to (ref) apply our general results to each of the motivating examples. The last section offers some concluding remarks. Technical proofs are relegated to Appendix (ref). Appendix (ref) contains additional results on each motivating example.
\paragraph*{Notation} We let $\boldsymbol{0}$ denote the zero vector and $I_{d}$ denote the identity matrix of dimension $d\times d$. We use $\mathds{1}\left\{ \cdot\right\} $ to denote the indicator function. For any vector $v\in\mathbb{R}^{d}$, we use $\left\Vert v\right\Vert $ to denote its Euclidean norm. We denote $\mathbb{S}^{d}$ as the unit sphere of dimension $d$. For any function $h$ and measure $\mu$, we denote $\int \vert h\vert d\mu < \infty$ by $h \in L^{1}(\mu)$. For any cumulative distribution function $F$ defined on $\mathbb{R}$, we let $F^{-1}\left(t\right)\equiv\inf\left\{ x:F\left(x\right)\geq t\right\} $ denote the quantile function. For any two random variables $W$ and $V$, we use $F_{W\mid v}^{-1}\left(\cdot\right)$ to denote the quantile function derived from the distribution of $W$ conditioning on $V=v$.
Our model is defined by ((ref)) and Assumption (ref) below, where $\mu_{1X}$ and $\mu_{0X}$ denote probability distributions of $\left(Y_{1},X\right)$ and $\left(Y_{0},X\right)$, respectively.
We maintain Assumption (ref) throughout the paper. To avoid the trivial case, we assume that $d_{1}\geq1$ and $d_{0}\geq1$, but $d_{x}\geq0$. If $d_{x}=0$, then there is no common $X$ in both datasets. Assumption (ref) (iii) can be interpreted as a comparability assumption. Consider, for example, the first dataset that contains only observations on $\left(Y_{1},X\right)$. Since $Y_{0}$ is missing, the distribution of $\left(Y_{0},X\right)$ is not identifiable. Assumption (ref) (iii) requires the underlying probability distribution of $\left(Y_{0},X\right)$ in the first dataset to be the same as the one of $\left(Y_{0},X\right)$ in the second dataset. The latter probability distribution $\mu_{0X}$ is identifiable because the second dataset contains observations on $\left(Y_{0},X\right)$. Furthermore, it implies that the distribution of $X$ in both datasets is the same.
In the rest of this section, we present three motivating examples and use them to illustrate our main results throughout the paper.
Let $\Theta_{I}\subseteq\Theta$ denote the identified set for $\theta^{\ast}$. It is defined as
where $\mathbb{E}_{\mu}$ denotes the expectation taken with respect to some probability measure $\mu\in\mathcal{M}\left(\mu_{1X},\mu_{0X}\right)$, and $\mathcal{M}\left(\mu_{1X},\mu_{0X}\right)$ is the class of probability measures whose projections on $\left(Y_{1},X\right)$ and $\left(Y_{0},X\right)$ are $\mu_{1X}$ and $\mu_{0X}$.
Equivalently, the identified set $\Theta_{I}$ can be defined via conditional distributions. Let $\mu_{1\mid x}$ and $\mu_{0\mid x}$ be the conditional distributions of $Y_{1}$ on $X=x$ and $Y_{0}$ on $X=x$, respectively. Denote $\mu_{X}$ as the probability distribution of $X$. Recall that $\mathcal{Y}_{1}$, $\mathcal{Y}_{0}$, and $\mathcal{X}$ are the supports of random variables $Y_{1}$, $Y_{0}$, and $X$. For any $x\in\mathcal{X}$, let $\mathcal{M}\left(\mu_{1\mid x},\mu_{0\mid x}\right)$ denote the class of probability distributions whose projections on $Y_{1}$ and $Y_{0}$ are $\mu_{1\mid x}$ and $\mu_{0\mid x}$.\footnote{All the statement with “for any $x\in\mathcal{X}$” can be relaxed to “for almost any $x\in\mathcal{X}$ with respect to $\mu_{X}$ measure”. We ignore such mathematical subtlety to simplify the discussion.} The identified set $\Theta_{I}$ defined in ((ref)) can be equivalently expressed as \[\Theta_{I}=\left\{ \theta\in\Theta:
\right\} . \] To simplify the notation, from now on, we omit the integration variables when there is no confusion.
This section provides a COT approach to characterize the identified set for $\theta^{*}$ in model ((ref)). Consider the following set defined through a continuum of inequalities:
For any given $x\in\mathcal{X}$ and $\theta\in\Theta$, $\mathcal{KT}_{t^{\top}m}\left(\mu_{1\mid x},\mu_{0\mid x};x,\theta\right)$ is the COT cost between conditional measures $\mu_{1\mid x}$ and $\mu_{0\mid x}$ with (ground) cost function $t^{\top}m\left(y_{1},y_{0},x;\theta\right)$. Here we state the explicit dependence of $\mathcal{KT}_{t^{\top}m}\left(\mu_{1\mid x},\mu_{0\mid x};x,\theta\right)$ on $x$ and $\theta$ because the cost function $t^{\top}m\left(y_{1},y_{0},x;\theta\right)$ is allowed to depend on $x$ and $\theta$. With the following assumption, we show that $\Theta_{I}=\Theta_{o}$.
Assumption (ref) (i) holds automatically when the supports of $\mu_{0\mid x}$ and $\mu_{1\mid x}$ are finite for every $x\in\mathcal{X}$, i.e., $\mu_{0\mid x}$ and $\mu_{1\mid x}$ are discrete measures, because any function on a finite set is continuous. Assumption (ref) (ii) is satisfied if $m$ is bounded. It also holds if $m$ does not depend on $x$, both $\mathcal{Y}_{1}$ and $\mathcal{Y}_{0}$ are bounded, and Assumption (ref) (i) holds. As a result, Examples (ref) and (ref) satisfy Assumption (ref). For Example (ref), the moment function satisfies Assumption (ref) (ii) if $\mathbb{E}\left[\left\Vert Y_{1}\right\Vert ^{2}\right]<\infty$ and $\mathbb{E}\left[\left\Vert Y_{0}\right\Vert ^{2}\right]<\infty$.
Proofs of theorems and propositions in this paper are provided in the appendix. The proof of Theorem (ref) relies on the minimax theorem derived in vianney2015minmax and the result on the weak compactness of $\mathcal{M}\left(\mu_{1\mid x},\mu_{0\mid x}\right)$ in the proof of Proposition 2.1 in villani2003topics.
Theorem (ref) characterizes the identified set for $\theta^{\ast}$ via a continuum of inequality constraints on the integral of the COT cost $\mathcal{KT}_{t^{\top}m}\left(\mu_{1\mid x},\mu_{0\mid x};x,\theta\right)$ with respect to $\mu_{X}$. This novel characterization allows us to explore both theoretical and computational tools, including duality theory developed in the OT literature, to study the identified set $\Theta_{I}$; see e.g., rachev2006mass, villani2008optimal, santambrogio2015optimal, and peyre2019computational.
\paragraph*{Covariate-Assisted Identified Set}
The conditioning variable $X$ provides information on the joint distribution of $\left(Y_{1},Y_{0}\right)$ by restricting the conditional distributions of $Y_{1}$ on $X=x$ and $Y_{0}$ on $X=x$ to be $\mu_{1\mid x}$ and $\mu_{0\mid x}$, which are identifiable from the datasets. The identified set obtained using only subvectors of $X$ can be potentially larger than the covariate-assisted identified set, which explores the information of the full vector $X$.
Partition $X\equiv\left(X_{p},X_{np}\right)$, where $X_{p}\in\mathbb{R}^{d_{x_{p}}}$ and $X_{np}\in\mathbb{R}^{d_{x_{np}}}$. Suppose that the moment function only depends on $X_{p}=x_{p}$ and denote it as $m\left(y_{1},y_{0},x_{p};\theta\right)$. This includes Fr\'{e}chet problem ((ref)) and Examples (ref)-(ref) for $d_{x_{p}}=0$. Denote by $\mu_{1\mid x_{p}}$ and $\mu_{0\mid x_{p}}$ the conditional distributions of $Y_{1}$ and $Y_{0}$ on $X_{p}=x_{p}$, respectively. Let $\mu_{X_{p}}$ be the probability distribution of $X_{p}$. Define the following set constructed using only $X_{p}$:
The following proposition shows that $\Theta_{I}$ is at most as large as $\Theta^{O}$.
Proposition (ref) extends Theorem 3.3 in fan2017partial established for Fr\'{e}chet problem ((ref)) with univariate $Y_{1}$ and $Y_{0}$. As long as $X_{np}$ is not independent of both $Y_{1}$ and $Y_{0}$, we expect the covariate-assisted identified set $\Theta_{I}$ to be smaller than $\Theta^{O}$. In general, the stronger $X$ and $Y_{1}$ (or $Y_{0}$) are dependent on each other, the smaller is the covariate-assisted identified set for $\theta^{\ast}$. The second part of the proposition shows that if the conditional measure $\mu_{0\mid x}$ is Dirac at $g_{0}(x)$, then $\Theta_{I}$ reduces to the identified set for the moment model with complete data $(Y_{1},X)$. Therefore, when one of $Y_{1}$ and $Y_{0}$ is perfectly dependent on $X$, $\theta^{\ast}$ is point identified under the standard rank condition. Without the common $X$ in both datasets, this would only be possible in the extreme case that one of $Y_{1}$ and $Y_{0}$ is constant almost surely.
In this section, we show that if the moment function $m$ is affine in $\theta$, then $\Theta_{I}$ is convex with a simple expression for its support function. Throughout the discussion, we assume that Assumption (ref) holds so that Theorem (ref) applies.
For each running example, we formulate the moment functions and $\theta$ in a way that ensures Assumption (ref) holds. Without loss of generality, we focus on the first decomposition in the following discussion. Define
Note that $\mathcal{KT}_{t^{\top}m_{b}}\left(\mu_{1\mid x},\mu_{0\mid x};x\right)$ does not depend on $\theta$, because the cost function $t^{\top}m_{b}\left(y_{1},y_{0},x\right)$ does not depend on $\theta$.
Under Assumption (ref), the convexity of $\Theta_{I}$ follows immediately from Theorem (ref) (i), because the constraints in the expression of $\Theta_{I}$ in Theorem (ref) (i) are affine in $\theta$.
For a convex set $\Psi\subseteq\mathbb{R}^{d_{\theta}}$, its support function $h_{\Psi}\left(\cdot\right):\mathbb{S}^{d_{\theta}}\rightarrow\mathbb{R}$ is defined pointwise by $h_{\Psi}\left(q\right)=\sup_{\psi\in\Psi}q^{\top}\psi$. Our novel characterization of the identified set in Theorem (ref) provides a straightforward way to obtain its support function. Define \[ \Theta_{R}\equiv\left\{ \theta\in\mathbb{R}^{d_{\theta}}:t^{\top}\mathbb{E}\left[m_{a}\left(Y_{0},X\right)\right]\theta\leq-\int\mathcal{KT}_{t^{\top}m_{b}}\left(\mu_{1\mid x},\mu_{0\mid x};x\right)d\mu_{X}\text{ for all }t\in\mathbb{S}^{k}\right\} , \] where $\mathcal{KT}_{t^{\top}m_{b}}\left(\mu_{1\mid x},\mu_{0\mid x};x\right)$ is defined in ((ref)). We make the following two additional assumptions. The first assumption is on the parameter space $\Theta$.
By definition, we have that $\Theta_{R}\cap\Theta=\Theta_{I}$. Assumption (ref) requires that the parameter space $\Theta$ does not provide additional information on the identified set $\Theta_{I}$ given $\Theta_{R}$. It implies that the boundary of $\Theta_{I}$ is completely determined by the continuum of inequalities in Theorem (ref) (i). For all the running examples, we can easily choose $\Theta$ to make Assumption (ref) hold. The second assumption is on the moment function.
Assumption (ref) requires that the number of moments is the same as the number of parameters, and that the coefficient matrix is of full rank. It implies that if the joint distribution $\mu_{o}$ of $\left(Y_{1},Y_{0},X\right)$ is identifiable from the data, then $\theta^{\ast}$ is just identified but not overly-identified. For the examples in the paper, Assumption (ref) is satisfied either automatically because $\mathbb{E}\left[m_{a}\left(Y_{0},X\right)\right]$ is a diagonal matrix with positive diagonal entries or under very mild conditions.
Under Assumption (ref), the expression of $\Theta_{I}$ in Theorem (ref) (i) is equivalent to
As a result, we have that $q^{\top}\theta\leq s\left(q\right)$ for all $\theta\in\Theta_{I}$. Moreover, applying the result in rockafellar1997ConvexAnalysis, we establish that $s(q)$ is indeed the support function of $\Theta_{I}$, as it is a positively homogeneous, proper, closed, and convex function satisfying $s(0) = 0$.
Proposition (ref) shows that we can obtain the value of the support function for any direction $q$ by computing the OT cost $\mathcal{KT}_{q^{\top}\mathbb{E}\left[m_{a}\left(Y_{0},X\right)\right]^{-1}m_{b}}\left(\mu_{1\mid x},\mu_{0\mid x};x\right)$ for each $x$ and then integrating with respect to $\mu_{X}$. Moreover, if the parameter of interest is expressed as a linear map from $\theta^{\ast}$ through the linear operator $L$, then based on Proposition (ref), the support function of the identified set for the parameter of interest can be simply calculated as $s\left(L^{\star}q\right)$, where $L^{\star}$ is the adjoint operator with respect to the inner product.
In this and the next two sections, we construct identified sets for the parameters in Examples (ref)-(ref) by applying our general results, and compare them with the existing results for each example. For Example (ref), we maintain the assumption that $\mathbb{E}\left[\left\Vert Y_{1}\right\Vert ^{2}\right]<\infty$ and $\mathbb{E}\left[\left\Vert Y_{0}\right\Vert ^{2}\right]<\infty$, an assumption that is also imposed in both d2024linear and hwang2023bounding.
Given the identified set $\Theta_{I}$ for $\theta^{\ast}$, the identified set for $\delta^{*}$ can be obtained by mapping each element in $\Theta_{I}$ through ((ref)), ((ref)), or ((ref)), depending on the model. We focus our discussion on constructing the identified set $\Theta_{I}$ in this section. We provide a detailed analysis for the first data type and summarize the identified set for the second data type in Remark (ref).
The moment function for $\theta^{*}$ is affine by its functional form:
Theorem (ref) implies that the identified set for $\theta^{\ast}$ is
Partitioning $t$ as $\left(t_{s}^{\top},t_{r,1}^{\top},\ldots,t_{r,\left(d_{1}-1\right)}^{\top}\right)^{\top}$ with $t_{s}\in\mathbb{R}^{d_{0}}$ and $t_{r,j}\in\mathbb{R}^{d_{0}}$ for $j=1,\ldots,d_{1}-1$, we can express $\mathcal{KT}_{t^{\top}m_{b}}\left(\mu_{1\mid x},\mu_{0\mid x};x\right)$ as a COT with a product cost function:
The first term on the right-hand side of Equation ((ref)) is half of the Wasserstein distance between the conditional distribution of $\left(t_{s}^{\top}Y_{0},t_{r,1}^{\top}Y_{0},\cdots,t_{r,\left(d_{1}-1\right)}^{\top}Y_{0}\right)$ and that of $Y_{1}$ given $X=x$. Wasserstein distance is the value function of the most studied and best understood OT problem. Recent breakthroughs in computational OT have contributed to the surge in its applications in diverse disciplines; see peyre2019computational for a comprehensive account of computational methods and applications. In addition, when both marginals are Gaussian, closed-form expressions for Wasserstein distance are known; see the discussion in santambrogio2015optimal and galichon2016optimal.
One important feature of the OT problem in ((ref)) is that the `effective' dimension of the marginal measures is equal to the dimension of $Y_{1}$ regardless of the dimension of $Y_{0}$, which can be high. Consequently, the computational burden only depends on $d_{1}$ and is independent of $d_{0}$. Moreover, in Appendix (ref), we show that by reordering elements of $\theta^{\ast}$, we can obtain a similar but alternative characterization of $\Theta_{I}$ such that the COT consists of a Wasserstein distance between the conditional distribution of $\left(t_{1}^{\top}Y_{1},\cdots,t_{d_{0}}^{\top}Y_{1}\right)$ and that of $Y_{0}$ given $X=x$ with $t_{j}\in\mathbb{R}^{d_{1}}$ for $j=1,\ldots,d_{0}$. For such a characterization, the computational burden would only depend on $d_{0}$.
When $d_{1}>1$ as in hwang2023bounding, the value function of Equation ((ref)) does not have a closed-from solution. However, a celebrated theorem of brenier1991polar implies that under mild regularity condition on $\mu_{1|x}$, the solution to ((ref)) concentrates on the gradient of a convex function, and the value function can be written as $\int-\nabla\phi_{t}\left(y_{1},x\right)y_{1}d\mu_{1\mid x},$ where $\nabla\phi_{t}\left(Y_{1},x\right)$ denotes the gradient of $\phi_{t}\left(\cdot,x\right)$, a convex function for each $t$ and $x$.
Due to the functional form of $G\left(\cdot\right)$, the resulting set $\Delta_{I}$ may not be convex. On the other hand, because $\Theta_{I}$ is convex and $G\left(\cdot\right)$ is continuous, the set $\Delta_{I}$ is connected and the projection of $\Delta_{I}$ onto any one-dimensional subspace is a connected interval.
In d2024linear, $d_{1}=1$. Proposition 2.17 in santambrogio2015optimal implies that Equation ((ref)) has a closed-form solution and
Consequently, we derive the identified set $\Theta_{I}$ from ((ref)) as
Following d2024linear, we let $\Theta=\mathbb{R}^{d_{0}}$. Assumptions (ref) and (ref) are both satisfied. For any $q_{0}\in\mathbb{S}^{d_{0}}$, Proposition (ref) provides the support function of $\Theta_{I}$:
Assume that the inverse of the second moment matrix $\mathbb{E}\left[\left(Y_{0}^{\top},X_{p}^{\top}\right)^{\top}\left(Y_{0}^{\top},X_{p}^{\top}\right)\right]$ exists, and denote its block partition by $\left(A_{ij}\right)_{i,j\in\left\{ 0,p\right\} }$. The same assumption is imposed in d2024linear. It follows from Equation ((ref)) that $\delta^{\ast}=\left(A_{00}^{\top},A_{p0}^{\top}\right)^{\top}\theta^{\ast}+\left(A_{0p}^{\top},A_{pp}^{\top}\right)^{\top}\mathbb{E}\left[X_{p}Y_{1}\right]$, which is an affine map. The standard result on the support function provides the following proposition on the support function of $\Delta_{I}$, which coincides with Theorems 1 and 2 of d2024linear.
The moment function $m$ is affine with
The identified set $\Theta_{I}$ of $\theta^{\ast}$ follows from Theorem (ref) (i), where \[ \mathcal{KT}_{t^{\top}m_{b}}\left(\mu_{1\mid x},\mu_{0\mid x};x\right)=\inf_{\mu_{10\mid x}\in\mathcal{M}\left(\mu_{1\mid x},\mu_{0\mid x}\right)}\int\int\sum_{j=1}^{J}-t_{j}\mathds{1}\left\{ y_{1}=1,y_{0}=a_{j}\right\} d\mu_{10\mid x}. \] Let $\mathtt{\boldsymbol{d}}\left(y_{1}\right)=\mathds{1}\left\{ y_{1}=1\right\} $ and $\mathtt{\boldsymbol{d}}_{t}\left(y_{0}\right)=\sum_{j=1}^{J}t_{j}\mathds{1}\left\{ y_{0}=a_{j}\right\} $. Then $\mathcal{KT}_{t^{\top}m_{b}}\left(\cdot\right)$ equals to \[ \inf_{\mu_{10\mid x}\in\mathcal{M}\left(\mu_{1\mid x},\mu_{0\mid x}\right)}\int\int-\mathtt{\mathtt{\boldsymbol{d}}}\left(y_{1}\right)\times\mathtt{\mathtt{\boldsymbol{d}}}_{t}\left(y_{0}\right)d\mu_{10\mid x}\left(y_{1},y_{0}\right)=-\int_{\Pr\left(Y_{1}=0\mid x\right)}^{1}F_{D_{t}\mid x}^{-1}\left(u\right)du \] with $D_{t}\equiv \mathtt{\boldsymbol{d}}_{t}\left(Y_{0}\right)$. The last equality follows from the monotone rearrangement inequality.
Given $p$, $D_{E^{\top}p}$ is a discrete random variable. As a result, we can easily obtain its quantile function. The closed-form expression of the support function of $\Delta_{DD}$ in ((ref)) provides an easy way to obtain the identified set for any single DD measure $\delta_{DD}^{\ast}\left(j,j^{\dagger}\right)$ for $J\geq2$; see Appendix (ref) for details. Furthermore, the analytical expression of the support function in ((ref)) allows us to obtain the vertex representation of the identified set $\Delta_{DD}$, which substantially simplifies its computation.
In the following, we prove that both $\Theta_{I}$ and $\Delta_{DD}$ are polytopes and provide their vertex representations. We first derive the vertices for $\Theta_{I}$ whose support function is
The ones for $\Delta_{DD}$ can be obtained by premultiplying $E$ to the vertices for $\Theta_{I}$.
We need to introduce some notation. For the set $M\equiv\left\{ 1,\ldots,J\right\} $, define $\mathfrak{S}_{M}$ as its symmetric group, which is the group of all permutations of $M$. Let $\sigma$ denote a generic element of $\mathfrak{S}_{M}$, i.e., a permutation of $M$. The order of $\mathfrak{S}_{M}$, denoted as $\left|\mathfrak{S}_{M}\right|$, equals $J!$. We can enumerate each element of $\mathfrak{S}_{M}$ as $\sigma_{s}$ for $s=1,\ldots,\left|\mathfrak{S}_{M}\right|$. The values of $\Pr\left(Y_{0}=a_{j}\right)$ for $j\in M$ are identifiable from the sample information. For any $\sigma_{s}\in\mathfrak{S}_{M}$, we define the following polyhedral region of $q$: \[ R_{s}\equiv\left\{ q\in\mathbb{R}^{J}:q_{\sigma_{s}\left(1\right)}\Pr\left(Y_{0}=a_{\sigma_{s}\left(1\right)}\right)^{-1}\leq\ldots\leq q_{\sigma_{s}\left(J\right)}\Pr\left(Y_{0}=a_{\sigma_{s}\left(J\right)}\right)^{-1}\right\} . \] Given $q\in R_{s}$, for any $s=1,\ldots,\left|\mathfrak{S}_{M}\right|$, the quantile function $F_{D_{q}\mid x}^{-1}\left(u\right)$ is a step function taking the value $\sum_{j=1}^{J}\frac{q_{\sigma_{s}\left(j\right)}}{\Pr\left(Y_{0}=a_{\sigma_{s}\left(j\right)}\right)}U\left(u,j,x\right)$, where \[ U\left(u,j,x\right)\equiv\mathds{1}\left\{ \sum_{j=1}^{j-1}\Pr\left(Y_{0}=a_{\sigma_{s}\left(j\right)}\mid x\right)<u<\sum_{j=1}^{j}\Pr\left(Y_{0}=a_{\sigma_{s}\left(j\right)}\mid x\right)\right\} \] and $\Pr\left(Y_{0}=a_{j}\mid x\right)$ denotes the conditional probability of $Y_{0}=a_{j}$ given $X=x$. This is a linear function of $q$. Since the integrations with respect to $u$ and $\mu_{X}\left(x\right)$ in Equation ((ref)) do not affect linearity, it holds that $h_{\Theta_{I}}\left(q\right)$ is linear in $q$ for $q\in R_{s}$. Because each $R_{s}$ is a convex cone and $\cup_{s=1}^{\left|\mathfrak{S}_{M}\right|}R_{s}=\mathbb{R}^{J}$, $h_{\Theta_{I}}\left(q\right)$ is a piecewise linear function in $q\in\mathbb{R}^{J}$ on a finite number of convex cones. Theorem 2.3.4 in hug2010course shows that $\Theta_{I}$ is a polytope. Furthermore, for each $s=1,\ldots,\left|\mathfrak{S}_{M}\right|$, because $h_{\Theta_{I}}\left(q\right)$ is linear in $q$ for $q\in R_{s}$, there is a vector $v_{s}$ such that $h_{\Theta_{I}}\left(q\right)=q^{\top}v_{s}$ for $q\in R_{s}$. Let $v_{s}\equiv\left(v_{s,1},\ldots,v_{s,J}\right)$. Then each element $v_{s,\sigma_{s}\left(j\right)}$ for $j=1,\ldots,J$ can be computed by \[ \int\int_{\Pr\left(Y_{1}=0\mid x\right)}^{1}\frac{1}{\Pr\left(Y_{0}=a_{\sigma_{s}\left(j\right)}\right)}U\left(u,j,x\right)dud\mu_{X}. \] There are $\left|\mathfrak{S}_{M}\right|$ number of vectors $v_{s}$ (not necessarily unique) such that $h_{\Theta_{I}}\left(q\right)=q^{\top}v_{s}$ for $q\in R_{s}$. The set $\left\{ v_{1},\ldots,v_{\left|\mathfrak{S}_{M}\right|}\right\} $ contains all the vertices of the identified set $\Theta_{I}$.\footnote{Not all $v_{s}$ for $s=1,\ldots,\left|\mathfrak{S}_{M}\right|$ are vertices. Some may lie inside the convex hull of others.} This proves the following proposition.
KMZ study a specific collection of DD measures: $E^{\star}\theta^{\ast}=\left[\delta_{DD}^{\ast}\left(1,J\right),\ldots,\delta_{DD}^{\ast}\left(J-1,J\right)\right]$. Denote $\Delta_{DD}^{\star}$ as the identified set for $E^{\star}\theta^{\ast}$. They propose to construct $\Delta_{DD}^{\star}$ using its support function as \[ \Delta_{DD}^{\star}\equiv\left\{ \delta :p^{\top}\delta\leq\int\varPhi_{K}\left(p,x\right)d\mu_{X}\left(x\right)\textrm{ for all }p\in\mathbb{S}^{J-1}\right\} , \] where $\int\varPhi_{K}\left(p,x\right)d\mu_{X}\left(x\right)$ is the value of the support function evaluated at $p$. Based on their Proposition 10, KMZ propose a linear programming approach to computing $\varPhi_{K}\left(p,x\right)$. For any given $p$ and $x$, there are $2J$ variables and about $5J$ (equality and inequality) constraints in their linear program. In Appendix (ref), we discuss the time complexity of the approach in KMZ and our methods, and show that both our support function approach and vertex representation approach are significantly more computationally efficient than the KMZ method.
For the numerical comparison, we use the empirical application in KMZ, which studies the identified set for demographic disparity in mortgage credit decisions using HMDA (Home Mortgage Disclosure Act) dataset. Following KMZ, we focus on three racial groups: White, Black, and Asian and Pacific Islander (API). We consider three sets of proxy variables for race: geolocation (county) only, annual income only, and both geolocation and annual income. To construct $\Delta_{DD}$, we apply the vertex representation $\textrm{conv}\left\{ Ev_{1},\ldots,Ev_{\left|\mathfrak{S}_{M}\right|}\right\} $ as discussed in Proposition (ref). In practice, this representation requires only the conditional probabilities $\Pr\left(Y_{1}=0\mid x_{i}\right)$ and $\Pr\left(Y_{0}=a_{j}\mid x_{i}\right)$ for $j=1,\ldots,J$, where $x_{i}$ for $i=1,\ldots,n$ are observations of $X$. Since these values are provided in the KMZ code, we use them directly instead of recomputing them from the raw data.\footnote{The computed conditional probabilities in Sections (ref) and (ref) can be downloaded from https://github.com/CausalML/FairnessWithUnobservedProtectedClass.} For the KMZ method, we rely on their published code and set the number of direction vectors to 100. As both methods require the same conditional probabilities, the comparison of computational time considers only the steps after these probabilities are obtained. Additional details on this empirical analysis are provided in Appendix (ref).
Figure (ref) shows that our method and the approach in KMZ yield the same identified set. However, our method is substantially faster. The vertex representation approach computes the identified set for each proxy variable in under 70 milliseconds, whereas the KMZ method requires approximately 20 minutes. This corresponds to a speed improvement of over 15,000-fold relative to the KMZ method.
The moment function is affine with $m_{a}\left(y_{0},x\right)$ being the $2J\times2J$ identity matrix and
Consequently, the identified set for $\theta^{*}$ is \[ \Theta_{I}=\left\{ \theta\in\Theta:\sum_{j=1}^{2J}t_{j}\theta_{j}\leq-\int\mathcal{KT}_{t^{\top}m_{b}}\left(\mu_{1\mid x},\mu_{0\mid x};x\right)d\mu_{X}\text{ for all }t\in\mathbb{S}^{2J}\right\} \] with support function for any $q\in\mathbb{S}^{2J}$ given by
Let $\Delta_{TPRD}$ denote the identified set for $K$ different TPRD measures represented by $G\left(\theta^{\ast}\right)$. We can compute $\Delta_{TPRD}$ as a non-linear map from $\Theta_{I}$: \[ \Delta_{TPRD}=\left\{ \delta=G\left(\theta\right):\theta\in\Theta_{I}\right\} =\left\{ \delta=G\left(\theta\right):q^{\top}\theta\leq h_{\Theta_{I}}\left(q\right)\textrm{ for all }q\in\mathbb{S}^{2J}\right\} , \] where $h_{\Theta_{I}}\left(\cdot\right)$ is defined in ((ref)). Because $\Theta_{I}$ is convex and the map $G$ is continuous, we know that $\Delta_{TPRD}$ is connected, which implies that the identified set for any single TPRD measure is a closed interval. In Appendix (ref), we derive the closed-form expressions for the lower and upper endpoints of the interval, extending the result in KMZ established for $J=2$.
For multiple TPRD measures, we need to compute $\mathcal{KT}_{q^{\top}m_{b}}\left(\cdot\right)$, for which we propose a novel computationally efficient algorithm in Section (ref). Our algorithm solves the reformulation of $\mathcal{KT}_{q^{\top}m_{b}}\left(\cdot\right)$ as the value function of a linear programming problem in ((ref)).
Given $q\equiv\left(q_{1},\ldots,q_{2J}\right)$, define $c\left(i,j\right)\equiv-q_{\left(1-i\right)J+j}$ for $i=0,1$ and $j=1,\ldots,J$. Because $Y_{1}$ and $Y_{0}$ are discrete random variables, we define, with a slight abuse of notation,
The COT cost $\mathcal{KT}_{t^{\top}m_{b}}\left(\cdot\right)$ can then be rewritten as the following optimization problem with the argument $\mu_{10\mid x}\left(i,1,j\right)$:
Note that the second constraint is an inequality rather than an equality because the COT cost $\mathcal{KT}_{q^{\top}m_{b}}\left(\cdot\right)$ does not involve $\mu_{10\mid x}\left(y_{1},y_{0}\right)$ for $y_{1}=\left(y_{1s},0\right)$. Consequently, ((ref)) is actually a conditional partial optimal transport, where $\mu_{0\mid x}\left(j\right)$ and $\mu_{1\mid x}\left(i,1\right)$ are identified from the sample information. Furthermore, since $i$ in the optimization problem ((ref)) takes two values only, we can always make $c\left(i,j\right)$ submodular by relabeling the index $j$. Without loss of generality, we will assume from now on that the cost function $c\left(i,j\right)$ in ((ref)) is submodular.
The following lemma shows that this reformulation of the partial optimal transport problem admits a solution whose support is monotone.
Exploiting the structure of the solution, we propose Algorithm (ref) to compute $\mathcal{KT}_{q^{\top}m_{b}}\left(\cdot\right)$ and call it Dual Rank Equilibration AlgorithM (DREAM). DREAM is built on an equivalent characterization of ((ref)) which allows it to use the ranks of $c\left(i,j\right)$ to equilibrate between minimizing the partial transport cost and respecting the constraints in ((ref)). The equilibration is done twice by first separately adjusting $\mu_{10\mid x}^{\ast}\left(0,1,j\right)$ and $\mu_{10\mid x}^{\ast}\left(1,1,j\right)$ across index $j$ and then jointly adjusting $\mu_{10\mid x}^{\ast}\left(0,1,j\right)$ and $\mu_{10\mid x}^{\ast}\left(1,1,j\right)$ for a fixed $j$---Dual Rank Equilibration AlgorithM. Appendix (ref) provides a detailed description of DREAM. DREAM involves only basic arithmetic, comparison, and logical operations, making it straightforward to implement directly in code. In fact, it can be executed by hand. In Appendix (ref), we provide theoretical time complexity analysis of DREAM and show that it is faster than directly solving the linear programming ((ref)), particularly when $J$ is large. \RestyleAlgo{ruled} \SetKwComment{Comment}{/* }{ */}
\RestyleAlgo{ruled} \SetKwComment{Comment}{/* }{ */}
Different from our approach, KMZ study the convex hull of the identified set for a collection of TPRD measures instead of the identified set itself. They focus on $\left[\delta_{TPRD}^{\ast}\left(1,J\right),\ldots,\delta_{TPRD}^{\ast}\left(J-1,J\right)\right]$, which is one example of our TPRD measures with $G^{\star}\left(\theta^{\ast}\right)\equiv\left[g_{1,J}\left(\theta^{\ast}\right),\ldots,g_{J-1,J}\left(\theta^{\ast}\right)\right]$. In their Proposition 11, KMZ propose to compute the support function of the convex hull at each direction from a complicated double optimization problem, where the inner maximization problem is a linear program and the outer maximization is a non-convex optimization problem, which is known to be NP-hard.
We applied DREAM and gurobi, a general-purpose linear programming solver, to the empirical application studied in KMZ.\footnote{We tried and failed to replicate the figures in KMZ using their approach due to the computational challenge of the NP-hard non-convex optimization involved.} The application investigates TPRD in Warfarin dosing with the ClinPGx/PharmGKB dataset used in IWPC_2009.\footnote{The dataset can be downloaded from https://www.clinpgx.org/downloads.} The analysis focuses on three protected attributes (White, Black, and Asian) and three specifications of proxy variables: genetic factors alone, current medications alone, and both genetic factors and current medications jointly. Because values of $\mu_{0\mid x}\left(\cdot\right)$ and $\mu_{1\mid x}\left(\cdot\right)$ in $(\ref{eq:Partial transport problem})$ for each observed $X$ have been computed in KMZ, we use them directly. During implementation, we first sample vectors $q_{1},\ldots,q_{N_{q}}$ uniformly from the $2J$-dimensional unit sphere, along with vectors $\theta_{1},\ldots,\theta_{N_{\theta}}$ from $\prod_{j=1}^{2J}\left[\theta_{j}^{L},\theta_{j}^{U}\right]$, where $\theta_{j}^{L}$ and $\theta_{j}^{U}$ are lower and upper bounds for each $\theta_{j}$ that are defined in Appendix Corollary (ref). Then, we construct the set $\widehat{\Theta}_{I}=\left\{ \theta\in\left\{ \theta_{1},\ldots,\theta_{N_{\theta}}\right\} :q^{\top}\theta\leq h_{\Theta_{I}}\left(q\right)\textrm{ for all }q=q_{1},\ldots,q_{N_{q}}\right\}$. Finally, we apply the mapping to each $\theta\in\widehat{\Theta}_{I}$ to obtain an approximation of $\Delta_{TPRD}$: $\widehat{\Delta}_{TPRD}=\left\{ \delta=G\left(\theta\right):\theta\in\widehat{\Theta}_{I}\right\}$.\footnote{Based on DREAM, the solution to ((ref)) depends only on the ranks of linear combinations of $c\left(i,j\right)$ for $i=0,1$ and $j=1,\ldots,J$ (or equivalently elements of the direction vector $q$). As a result, one can show that $\Theta_{I}$ is a polytope, and obtain its vertex representation by following the similar analysis in Section (ref). However, because the representation is complex and we still need to sample $\theta$ to construct $\widehat{\Delta}_{TPRD}$, we recommend using $\widehat{\Theta}_{I}$.} Appendix (ref) provides more detail.
Figure (ref) plots the identified sets obtained from DREAM, which are identical to the sets produced by the gurobi solver up to a negligible numerical error. The runtime comparison shows that our approach is more than twice as fast as the commercial linear programming solver.
In this paper, we have developed the first unified approach to study the identified set for a finite-dimensional parameter in a general moment equality model with incomplete data.\footnote{ The methods developed in this paper could, in principle, be extended to the case of more than two datasets using results in pass2015multi. We leave extensions for future work.} Through several examples, we have demonstrated the advantages and simplicity of our approach. By exploring recent developments in optimal transport, our method often leads to equivalent, yet computationally more tractable, identified sets for specific models than the existing ones in the literature.
This paper provides the first step towards developing a complete set of econometric techniques for moment models under data combination. Building on the identification analysis in this paper, the next step is to construct valid estimation and inference in the general moment model ((ref)) with incomplete data. For Example (ref) without common covariate $X$, d2024linear develops estimation and inference for $\theta^{*}$. For the special case that $k=1$ and $\theta^{*}=\mathbb{E}\left[h\left(Y_{1},Y_{0},X)\right)\right]$, ji2023model and lin2025estimation construct estimation and inference using primal and dual formulation, respectively. We leave the general case for future work.
We thank St\'{e}phane Bonhomme, Xiaohong Chen, Marc Henry, Ruixuan Liu, and participants of the BIRS Workshop on Optimal Transport and Distributional Robustness in March 2024, 2025 IAER Econometrics Workshop, and 2025 China Annual Conference of the Chinese Economists Association; seminar participants in the statistics department at the University of Washington, economics departments at Johns Hopkins University and Yale University, as well as Amazon in October 2024 for useful feedback. Brendan Pass is pleased to acknowledge the support of the Natural Sciences and Engineering Research Council of Canada Discovery Grant numbers 04658-2018 and 04864-2024.