EconBase
← Back to paper

Lee Bounds for Random Objects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

59,836 characters · 12 sections · 50 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Lee bounds for random objects

\address[D. Kurisu]{Center for Spatial Information Science, The University of Tokyo, 5-1-5, Kashiwanoha, Kashiwa-shi, Chiba 277-8568, Japan.} \email{[email removed]}

\address[Y. Okamoto]{Graduate School of Economics, Kyoto University, Yoshida Honmachi, Sakyo, Kyoto 606-8501, Japan. } \email{[email removed]}

\address[T. Otsu]{Department of Economics, London School of Economics, Houghton Street, London, WC2A 2AE, UK. } \email{[email removed]}

abstractIn applied research, Lee (2009) bounds are widely applied to bound the average treatment effect in the presence of selection bias. This paper extends the methodology of Lee bounds to accommodate outcomes in a general metric space, such as compositional and distributional data. By exploiting a representation of the Fr\'echet mean of the potential outcome via embedding in an Euclidean or Hilbert space, we present a feasible characterization of the identified set of the causal effect of interest, and then propose its analog estimator and bootstrap confidence region. The proposed method is illustrated by numerical examples on compositional and distributional data.

Introduction

Randomized controlled trials are among the most credible approaches for answering causal questions in economics, medical science, and many other disciplines. However, their validity can be undermined by endogenous sample attrition or non-response, which are ubiquitous in empirical studies. A common approach to address this issue, developed in a seminal work by Lee:2009, is to bound the average treatment effect of units whose outcomes are observed regardless of the treatment status. Recent empirical studies employing the Lee bounds include Alfonsi_etal:2020, Atal_etal:2024, and Dobbie_Fryer:2015, among many others.

The methodology of the Lee bounds has been extended in several directions to deal with noncompliance issues Chen_Flores:2015, continuous treatments Lee_liu:2025, multi-layered selection Kroft_etal:2025, and relaxed monotonicity assumptions as well as multiple outcomes Semenova:2025, for example. However, almost all existing works focus on outcomes that lie in a Euclidean space, even though many economically important outcomes do not live in Euclidean spaces.

One prominent class is compositional data, where components are nonnegative and sum to one. For example, labor and macro economists study intra-household time allocation (e.g., Cardia_Gomme:2018, Lise_Yamada:2019), while public economists analyze the composition of government expenditures (e.g., Brender_Drazen:2013, Shelton:2007). Other relevant examples include the composition of household consumption, energy mix across sources, agricultural land use, and portfolio shares in finance.

Another important class consists of distribution-valued outcomes. Economists have studied income and wage distributions (e.g., Battistin_etal:2009, DiNardo_etal:1996) as well as price dispersion (e.g., Cavallo:2017, Pennerstorfer_etal:2020). Related examples include the distributions of daily step counts, daily television viewing time, or hourly electricity consumption, which may be of interest in health, education, and public economics.

Beyond these classes, researchers also encounter network outcomes, such as cross-country trade and travel flows, as well as transaction networks among firms. Interval-valued data are likewise ubiquitous in econometric analysis---for example, when outcomes are reported in brackets for confidentiality, elicited as ranges to improve response rates, or observed with rounding error. Finally, correlation matrices are central objects for studying connectivity, including applications such as product recommendation and brain functional connectivity. We refer to Dubey:2024 and references therein for a general review of statistical analysis of such random objects, along with a range of applications.

Despite their prevalence, standard implementations of Lee bounds are not directly applicable to these cases, since such outcomes are not situated in Euclidean spaces. To overcome this limitation, we extend the scope of partial identification analysis via Lee bounds by adapting the methodology of metric statistics for random objects. In particular, we define the expected potential outcomes for random objects by introducing the notion of the Fr\'echet mean, a direct generalization of the conventional mean toward a general metric space,\footnote{Perhaps a popular example of the Fr\'echet mean in economics is the Aumann mean for random sets (e.g., BeMo08), which is shown to be a special case of the Fr\'echet mean (kuri:25), and our methodology also applies to random sets (see, Example (ref) below).} and then establish partial identification results on the expected potential outcome by exploiting a representation in an embedded Euclidean or Hilbert space, where the bounding argument in Lee:2009 may be adapted. Intuitively, our identified set is obtained by pulling back the Lee-type bounds on the embedded Euclidean or Hilbertian variables, and this approach enables analysis that takes into account the geometric structure of the metric space in which the outcome takes values. Based on our partial identification strategy, the treatment effect can be characterized by a difference in the embedded space, a difference in the projections, or geodesics in the metric space. Our estimator for the identified set is obtained by taking the sample analogs, and we present a valid bootstrap procedure for the confidence region of our identified set.

Causal inference on random object outcomes has been increasingly popular in recent econometrics and statistics literature, e.g., Gunsilius:2023, kurisu_etal:24, KuZhOtMu25a, KuZhOtMu25b, and ZhKuOtMu25. The present paper contributes to this literature by extending the method of Lee bounds to random objects. In contrast to the above papers, to the best of our knowledge, this is the first paper that conducts partial identification analysis on the Fr\'echet mean in a general metric space. In this sense, this paper opens a new avenue for applications of the methodologies of partial identification and metric statistics.

This paper is organized as follows. After closing this section with a recap of the conventional Lee bounds, Section (ref) showcases our partial identification analysis by focusing on compositional data. Then Section (ref) presents our general methodology for partial identification. In Section (ref), we discuss estimation and bootstrap inference on the proposed identified set. Section (ref) provides an empirical illustration. Finally, Appendix (ref) discusses additional examples of random objects, and Appendix (ref) contains the proofs of the theoretical results.

Recap: Lee bounds

For the reader's convenience, we start with the standard Lee bounds for a scalar outcome. We also discuss potential issues that arise and are resolved in the subsequent sections. Consider a binary treatment $D\in \{0,1\}$ (one for the treatment and zero otherwise) with potential outcomes $Y_0$ and $Y_1$ that take values in $\mathcal{Y}\subset\mathbb{R}$. Let $S_{0}, S_{1}\in\{0,1\}$ denote potential selection indicators (one for non-missing and zero otherwise). As in Lee:2009, we assume that $D$ is independent from $(Y_1, Y_0, S_1, S_0)$ (random assignment) and $S_1 \ge S_0$ (monotonicity). Instead of the latent variables $(Y_1, Y_0, S_1, S_0)$, we observe $D$ and $(S, Y)$ that are generated from the general sample selection model:

equation[equation omitted — 219 chars of source]

When the outcome variable $Y$ is scalar-valued, we typically partially identify the average treatment effect for a certain subpopulation Lee:2009:

align*[align* omitted — 141 chars of source]

where we refer to the group of individuals with $S_{1} = S_{0} = 1$ as the always-observed group. The second term $\E{Y_{0} \mid S_{1} = S_{0} = 1}$ is point identified as $\E{Y \mid S = 1, D=0}$, but the first term $\E{Y_{1} \mid S_{1} = S_{0} = 1}$ is not. To partially identify the first term, observe that

align*[align* omitted — 217 chars of source]

where $p=\P{S=1 \mid D=0}/\P{S=1 \mid D=1}$. Then, by considering the worst (resp. best) case scenario where the smallest (resp. largest) $100\times p\%$ values of observed $Y$ (conditional on $S=1, D=1$) are entirely attributed to $Y_{1}$ of the always-observed group, we can obtain the lower (resp. upper) bound on the first term, leading to the following bounds:

align*[align* omitted — 114 chars of source]

where $Q_{Y}(\cdot)$ denotes the conditional quantile function of $Y$ given $D = 1, S = 1$. Based on these bounds, we can obtain the Lee bounds on the average treatment effect $\E{Y_{1} - Y_{0} \mid S_{1} = S_{0} = 1}$ for the always-observed group.

In the subsequent sections, we extend this partial identification strategy to outcomes that take values in more general metric spaces (i.e., random objects rather than Euclidean variables). One of the key issues is how to define and interpret a “mean” in such settings. For Euclidean outcomes, the usual expectation admits a geometric interpretation as the “center” of the distribution. This interpretation, however, is inherently metric-dependent, and as a result, when the outcome variable lies in a non-Euclidean space, the conventional Euclidean mean may ignore the geometric property and perhaps lead to a misleading notion of centrality.

To illustrate this issue, consider a three-part compositional outcome. Let $c_0=(1/3,1/3,1/3)$ denote the “barycenter” of the space of compositional data, and consider an interior point $c_1 = (2/3, 1/6, 1/6)$ and a boundary point $c_2 = (0, 1/2, 1/2)$. Under the Euclidean metric, as illustrated in Figure (ref), the distances from the barycenter to these two points are identical: $d(c_0, c_1) = d(c_0, c_2)$.

figure[figure omitted — 568 chars of source]

From a compositional perspective, however, these two displacements are qualitatively different even if they have the same Euclidean magnitude. The move from $c_0$ to $c_1$ remains within the interior of the simplex and corresponds to a finite reallocation of mass among components. In contrast, the move from $c_0$ to $c_2$ pushes the first component to zero, i.e., to the boundary (or an extreme point) of the simplex. This latter behavior can be considered more extreme. Accordingly, a metric that respects this kind of compositional geometry, such that $d(c_0, c_1) < d(c_0, c_2)$ as in Figure (ref) based on the Aitchison metric, is more natural.

The same consideration applies in a more general setting considered in Section (ref). Therefore, to extend Lee's (2009) bounding procedure to random object outcomes, we begin by introducing an appropriate metric that allows statistical analysis to reflect the geometry of the space in which outcomes take values. We then define an extended notion of the mean (i.e., Fr\'echet mean) with respect to the metric and develop the Lee-type identified region for the Fr\'echet mean of random object outcomes. Another key issue is how to operationalize the bounding approach. We address this issue by exploiting a well-behaved embedding that reduces our partial identification problem in a metric space to the one in a familiar Euclidean or Hilbert space.

Benchmark example: Compositional data

Compositional data play a central role in economic analysis, as many key economic objects are inherently defined as allocations across mutually exclusive categories that sum to a whole. For example, household time allocation among labor, leisure, and housework is a fundamental concept for understanding labor supply decisions and intrahousehold bargaining. Analogously, the allocation of total expenditure across goods underpins classical consumer choice theory. In evaluating cash transfer programs, researchers likewise examine how households allocate consumption across categories; for example, the share devoted to health-related versus non-health food expenditures.

To accommodate these important scenarios, we now extend the scope of the Lee bounds to a compositional data scenario. Recall the treatment indicator ($D$) and selection indicators ($S_1, S_0, S$) defined as before. We retain the random assignment and monotonicity assumptions. Suppose that the potential outcomes $Y_0$ and $Y_1$ take values in the space for compositional data with positive components:

align*[align* omitted — 136 chars of source]

We impose the strict positivity condition $y_j > 0$ only to simplify the discussion that follows. Cases with zero components are accommodated within our general framework in Section (ref).

To introduce a notion of the expectation of the potential outcome $Y_t \in \mathcal{Y}$ for $t\in\{0,1\}$, let us consider the Aitchison metric $d$ on $\mathcal{Y}$ defined as

align*[align* omitted — 45 chars of source]

where $\|\cdot\|$ is the Euclidean metric on $\mathbb{R}^k$ and

align[align omitted — 181 chars of source]

The Aitchison metric is commonly applied for compositional data analysis (Aitchison:1982). As illustrated in Figure (ref), with the Aitchison metric, the resulting distances reflect the underlying geometric structure. Based on the metric space $(\mathcal{Y},d)$, we introduce the Fr\'echet mean of $Y_t\in\mathcal{Y}$, which is defined as

align*[align* omitted — 104 chars of source]

The Fr\'echet mean is a direct generalization of the conventional mean for the Euclidean space toward a general metric space, and its statistical analysis has been increasingly popular in recent literature (see, e.g., Dubey:2024). Analogously, the conditional Fr\'echet mean is defined as $\mathbb{E}_\oplus[Y_t \,| \,\cdot\,]=\mathrm{arg\,min}_{y \in \mathcal{Y}}\E{d^2(y, Y_t)\,|\,\cdot\,}$, and the average treatment effect for the always-observed group is characterized by some contrast of $\mu_{1\oplus}$ and $\mu_{0\oplus}$, where $\mu_{t\oplus}=\mathbb{E}_\oplus[Y_t\mid S_1=S_0=1]$.

As in the scalar outcome case, the assumptions of random assignment and monotonicity imply that $\mu_{0\oplus}$ is point identified as

equation[equation omitted — 127 chars of source]

In contrast, $\mu_{1\oplus}$ is not point identified as in the scalar outcome case. Thus, our main task is to characterize the sharp identified set for $\mu_{1\oplus}$. To this end, we note that (i) the map $\Psi:\mathcal{Y} \to \mathbb{R}^k$ is an isometric and injective map from $\mathcal{Y}$ to $\mathbb{R}^k$, and (ii) the image space $\Psi(\mathcal{Y}) = \{x \in \mathbb{R}^k: \sum_{j=1}^k x_j =0\}$ is a closed convex subset of $\mathbb{R}^k$. Due to these properties, $\mu_{1\oplus}$ can be written as

align*[align* omitted — 84 chars of source]

Note that the expectation $\mathbb{E}[\cdot]$ on the right-hand side is the one for the conventional Euclidean vector, and this representation motivates the following two-step procedure to construct an identified set for $\mu_{1\oplus}$:

align*[align* omitted — 299 chars of source]

For Step 1, a similar reasoning to the scalar-valued case applies, and the sharp identified set of $\E{\Psi(Y_1)\mid S_1=S_0=1}$ can be written as

align[align omitted — 138 chars of source]

where $\mathscr{A}$ is the collection of all subsets $\mathcal{A}$ such that $\P{\Psi(Y)\in \mathcal{A} \mid D = 1, S = 1} = p$ (the set in ((ref)) coincides with the set $\mathcal{I}_1$ in our general result in Section (ref)). Furthermore, a standard result in convex analysis implies that the identified set in ((ref)) admits an alternative tractable characterization $\mathcal{S}_1$. For Step 2, we can show that $\Psi^{-1}(\mathcal{S}_1)$ is indeed the sharp identified set of $\mu_{1\oplus}$, so the main result of this section is summarized as follows.

propositionConsider the setup of this section (in particular, assume random assignment and monotonicity) with some regularity conditions (Assumption (ref) below). Then, $\mu_{0\oplus}$ is point identified as in ((ref)). Moreover, $\mu_{1\oplus}$ is partially identified and its sharp identified set is $\Psi^{-1}(\mathcal{S}_1)$.

This proposition is obtained as a special case of the general result in the next section. Before turning to our general result, we illustrate the proposed strategy and its appeal using a numerical example.

Numerical illustration

This subsection illustrates our proposed procedure. We consider a hypothetical randomized controlled trial setting with sample attrition using data from the 2024 American Time Use Survey. We restrict attention to respondents aged 25-55 who are employed and report positive time spent on leisure, labor-market work, and housework.\footnote{We define leisure as the sum of first-tier codes 1 (Personal Care), 11 (Eating and Drinking), 12 (Socializing, Relaxing, and Leisure), and 13 (Sports, Exercise, and Recreation). Labor-market work corresponds to first-tier code 5 (Work & Work-Related Activities). We define housework as the sum of first-tier codes 2 (Household Activities), 3 (Caring for & Helping Household Members), 4 (Caring for & Helping Nonhousehold Members), and 7 (Consumer Purchases). See \url{https://www.bls.gov/tus/lexicons/lexiconwex2024.pdf} for the list of detailed activities within each category.} The resulting sample size is $1,397$. We then randomly assign individuals to two groups, treating the split as if it were a treatment-control assignment ($711$ in the control group and $686$ in the treatment group). To mimic attrition, we randomly drop observations according to group-specific selection probabilities. Specifically, we set $\P{S=1\mid D=1}=0.90$ and vary the control-group retention rate over $\P{S=1\mid D=0}\in\{0.85,0.70\}$. Since $\P{S=1\mid D=1}$ is held fixed, a smaller $\P{S=1\mid D=0}$ implies a smaller $p$, which in turn increases the degree of “contamination” $1-p$ and limits identifying power. Because there is, in fact, no treatment, we expect no treatment effect and hence assess whether the estimated bounds $\widehat{\mathcal{S}}_1$ contain the point-identified Fr\'echet mean for the controlled group $\widehat{\mu}_{0\oplus}$.

figure[figure omitted — 951 chars of source]

Figure (ref) presents the estimated bounds $\widehat{\mathcal{S}}_1$ (inner, purple region) and point-identified Fr\'echet mean $\widehat{\mu}_{0\oplus}$ (square dot), as well as the 95% bootstrap confidence region for ${\mathcal{S}}_1$ (outer, gray region), whose validity is shown in Section (ref). We can see that the bounds and confidence sets contain $\widehat{\mu}_{0\oplus}$, suggesting the validity of our proposed bounds. In the low-attrition case (Figure (ref)), the identified set is reasonably tight and informative. When the attrition rate is higher (Figure (ref)), the identified set expands, but it still provides non-trivial identifying power.

table[table omitted — 1,438 chars of source]

Table (ref) shows the projections of the estimated bounds and confidence region, as well as the naively applied standard Lee bounds. In Panel A, we can find that the componentwise Lee bounds do not contain the Fr\'echet mean for the control group (which coincides with the counterfactual Fr\'echet mean as well in our scenario) in any dimensions. This pattern suggests that the standard Lee bounds may fail to deliver a valid identified set for the center of the counterfactual outcome distribution (as measured by the Fr\'echet mean). By contrast, the componentwise projections of our proposed bounds do cover $\widehat{\mu}_{0\oplus}$. We also note that these confidence intervals are simultaneously valid, as they are obtained by projecting the confidence region for $\Psi^{-1}(\mathcal{S}_1)$.

In Panel B, although the naive Lee bounds become wider, the bounds for the leisure component still fail to cover $\widehat{\mu}_{0\oplus}$. In contrast, the componentwise projections of our proposed bounds remain valid, although they are less informative than in the low-attrition case shown in Panel A.

Generalization: Lee bounds for random objects

Setup and examples

This section extends the benchmark example in the last section and presents a general framework to conduct partial identification analysis for the sample selection model in ((ref)), where the outcome variable $Y$ resides in a metric space $(\mathcal{Y},d)$.\footnote{Although we focus on the sample selection model, our analysis also applies to the contaminated data model studied by Horowitz_Manski:1995; see Remark (ref) below.} To begin with, we redefine the notation and formally introduce our setup.

assumptionLet $Y_0, Y_1 \in \mathcal{Y}$ be potential outcomes and $S_0, S_1 \in \{0,1\}$ be potential selection indicators for a treatment $D \in \{0,1\}$. $\mathcal{Y}$ is equipped with a metric $d$. Assume that $D$ is independent from $(Y_1, Y_0, S_1, S_0)$ (random assignment) and $S_1 \geq S_0$ with probability one (monotonicity). We observe $D$ and $(S, Y)$ generated from the selection model in ((ref)).

This setup is identical to the one in Lee:2009 except that the support $\mathcal{Y}$ can be non-Euclidean, and we are concerned with causal inference by a randomized experiment in the presence of sample selection. The monotonicity assumption $S_1 \ge S_0$, which rules out “defiers” ($S_1=0, S_0=1$) from the population, requires that the assignment $D$ can affect the sample selection only in one direction. This assumption is standard in many econometric models and aligns well with the threshold-crossing model Vytlacil:2002.

In this setup, the average treatment effect of interest is characterized by some contrast of the Fr\'echet means \[ \mu_{t\oplus}=\mathbb{E}_\oplus[Y_t\mid S_1=S_0=1]=\mathop{\rm arg~min}\limits_{y \in \mathcal{Y}}\E{d^2(y,Y_{t})\mid S_{1}=S_{0}=1}, \] for $t \in \{0,1\}$. Again, $\mu_{0\oplus}$ is point identified as in ((ref)) under Assumption (ref). Thus, we focus on partial identification of $\mu_{1\oplus}$, and impose the following assumptions on the metric space $(\mathcal{Y},d)$.

assumption\quad \begin{description} • There exists an injective map $\Psi$ from $\mathcal{Y}$ to a separable Hilbert space $\mathcal{H}$ with the metric $d_\mathcal{H}$ such that $d(x,y)=d_\mathcal{H}(\Psi(x),\Psi(y))$. Furthermore, $\Psi(\mathcal{Y})$ is a closed convex subset of $\mathcal{H}$. • $\E{\| \Psi(Y) \|^2} < \infty$, and the inner product $\langle u, \Psi(Y)\rangle$ induced by $d_\mathcal{H}$ is continuously distributed conditional on $D = 1$ and $S = 1$ for each $u\in\mathcal{H}$ with $\|u\| = 1$. \end{description}

Assumption (ref) (ii) is a regularity condition requiring the existence of the second moment and continuity of the distribution. The key assumption in our analysis is hence Assumption (ref) (i). Assumption (ref) (i) enables us to characterize the (conditional) Fr\'echet mean of the random variable $Y$ using the map $\Psi$. In fact, the (conditional) Fr\'echet mean of $Y$ is obtained by pulling back, via $\Psi^{-1}$, the (conditional) expectation of the embedded variable $\Psi(Y)$ in the Hilbert space $\mathcal{H}$, which is uniquely determined by the Riesz representation theorem. We refer to kuri:25 for further details on this point. Many of the common metric spaces studied in the literature of metric statistics are known to satisfy Assumption (ref) (i). Here, we list some examples of metric spaces that are commonly used in statistical analysis of metric space-valued data. Additional examples are contained in Appendix (ref).

example[Interval data] Set-valued data have a wide range of applications, for example, in census and survey data analysis and in biostatistics BeMo08,adusumilli2017empirical,li2021local. For $d\geq 1$, let $K_{kc}(\mathbb{R}^d)$ denote the collection of nonempty compact and convex subsets of $\mathbb{R}^d$ and let $K_{kc}^B(\mathbb{R}^d)=\{F \in K_{kc}(\mathbb{R}^d): \sup_{f \in F}\|f\|\leq B\}$ where $\|\cdot\|$ is the standard Euclidean metric and $B$ is a positive constant. We can introduce a metric $d_{kc}$ on $K_{kc}(\mathbb{R}^d)$ defined as $d_{kc}(F,G)=\sqrt{\int_{\mathbb{S}^{d-1}}\{s_F(t)-s_G(t)\}^2\,dt}$, where $s_F(t)=\sup_{f \in F}t'f$ is the support function of $F$ and $\mathbb{S}^{d-1}=\{x \in \mathbb{R}^d:\|x\|=1\}$ is the $(d-1)$-dimensional unit sphere. In this case, the mapping $\Psi(F) = s_F(\cdot)$ is an isometry from $K_{kc}^B(\mathbb{R}^d)$ to $L^2(\mathbb{S}^{d-1})$, the space of functions such that $\int_{\mathbb{S}^{d-1}}f^2(x)dx<\infty$, and the image space $\Psi(K_{kc}^B(\mathbb{R}^d))$ is a closed convex subset of $L^2(\mathbb{S}^{d-1})$ kuri:25. In particular, for $d=1$, $K_{kc}(\mathbb{R})=\{[a,b]: -\infty<a\leq b<\infty\}$ is the space for interval data, and its image space $\Psi(K_{kc}(\mathbb{R}))$ can be identified with $\{(x,y) \in \mathbb{R}^2: y+x\geq 0\}$, which is a closed convex subset of $\mathbb{R}^2$.
example[Functional data] Data consisting of functions are referred to as functional data. Such data are prevalent in longitudinal studies, environmental monitoring, and biomedical imaging ramsay2005functional, hsing2015theoretical, wang2016functional. Those functions are typically assumed to be situated in the space $L^2(\mathcal{I})$ of real-valued square-integrable functions on a compact interval $\mathcal{I}$. This space carries the inner product $\langle f_1, f_2 \rangle_{L^2}=\int_{\mathcal{I}} f_1(x)f_2(x)\,dx$ and the corresponding $L^2$ metric, given by $\|f_1-f_2\|_{L^2} = \sqrt{\int_\mathcal{I}\{f_1(x) - f_2(x)\}^2\,dx}$. The space $L^2(\mathcal{I})$ is closed and convex with respect to the $L^2$ metric. In this example, we can set $\mathcal{H} = L^2(\mathcal{I})$ and take $\Psi$ as the identity map $\Psi(f) = f$.
example[One-dimensional probability distributions] Distributional data arise when each data point is regarded as a probability distribution, and they have been applied in many fields, including economics and multi-cohort studies petersen2022modeling,Gunsilius:2023,zhou2024wasserstein. Let $\mathcal{W}$ denote the space of Borel probability measures on a real line, with finite second moments. This space becomes a metric space when equipped with the 2-Wasserstein distance $d_\mathcal{W}(\mu_1, \mu_2) = \sqrt{\int_0^1 \{F_{\mu_1}^{-1}(u) - F_{\mu_2}^{-1}(u)\}^2\,du}$. Here, $F_{\mu}^{-1}$ denotes the quantile function of a probability distribution $\mu$. The metric space $(\mathcal{W}, d_\mathcal{W})$ is called the Wasserstein space panaretos2020invitation, and distributional data are typically assumed to reside within this space. The mapping $\Psi(\mu) = F_{\mu}^{-1}$ is clearly an isometry from $\mathcal{W}$ to the Hilbert space $L^2([0, 1])$. Moreover, the image $\Psi(\mathcal{W})$ is characterized as the set of all square-integrable, almost everywhere increasing functions on $[0, 1]$, and it is closed and convex in $L^2([0, 1])$ bigot2017geodesic.
example[Networks] Consider a simple, undirected, and weighted network with a set of nodes $\{v_1, \ldots, v_m\}$ and a set of bounded edge weights $ \{w_{pq}: p, q=1, \ldots, m\}$, where $0 \le w_{pq} \le W$. Such a network can be uniquely represented by its graph Laplacian matrix $L = (l_{pq}) \in \mathbb{R}^{m^2}$, defined as \[ l_{pq} = \begin{cases} -w_{pq} & \text{if}\,\, p \neq q, \\ \sum_{r \neq p}w_{pr} & \text{if}\,\, p = q, \end{cases} \quad p, q = 1, \ldots, m. \] The space of graph Laplacians is given by $\mathcal{L}_m = \{L = (l_{pq}): L = L', L1_m = 0_m, -W \le l_{pq} \le 0 \,\, \text{for}\,\, p \neq q\}$ where $1_m$ and $0_m$ are the $m$-vectors of ones and zeros, respectively. This space provides a natural framework for characterizing network structures such as social connections, transportation links, or gene regulator networks kola:14, zhou2022network,severn2022manifold. Equipped with the Frobenius metric $d_F$ defined as $d_F(L_1,L_2) = [\text{tr}\{(L_1-L_2)'(L_1-L_2)\}]^{1/2}$, the space of graph Laplacians $\mathcal{L}_m$ forms a closed and convex subset of the Euclidean space $\mathbb{R}^{m^2}$ zhou2022network.

Main results

Now, we are in a position to derive the sharp identified set of $\mu_{1\oplus}=\mathbb{E}_\oplus[Y_1\mid S_1=S_0=1]$ (see Definition (ref) for the formal definition of the sharp identified set). Under Assumption (ref), the potential Fr\'echet mean can be expressed as $\mu_{1\oplus}=\Psi^{-1}(\E{\Psi(Y_{1})|S_{1}=S_{0}=1})$. Thus, as in the last section, we begin by characterizing an identified set for $\E{\Psi(Y_{1})|S_{1}=S_{0}=1}$. Let $p= \P{S=1 | D=0}/\P{S=1 | D=1}$ and observe that \[ \P{\Psi(Y)\in\mathcal{A}|D = 1, S = 1} = p\P{\Psi(Y_{1})\in\mathcal{A}|S_{1} = S_{0} = 1} + (1-p)\P{\Psi(Y_{1})\in\mathcal{A}|S_{1} = 1, S_{0} = 0}, \] where $\mathcal{A}$ is an arbitrary measurable subset of $\mathcal{H}$. Hence, a $100\times p\%$ of the observed subpopulation with $D = 1, S = 1$ corresponds to the always-observed group. Then, by letting $\mu(\mathcal{A})\coloneqq \P{\Psi(Y)\in\mathcal{A}|D = 1, S = 1}$, the sharp identified set for $\E{\Psi(Y_1) \mid S_1=S_0=1}$ is obtained as

align[align omitted — 273 chars of source]

where $\int_{\mathcal{H}} \cdot \,d\nu$ means the Bochner integral, and the sharpness of $\mathcal{I}_1$ is shown in Appendix (ref). Our partial identification results are presented as follows.

theoremSuppose Assumptions (ref) and (ref) hold true. Then \begin{description} • The convex closure of $\mathcal{I}_1$ satisfies \begin{equation} \overline{\mathrm{conv}}\,\mathcal{I}_1 = \bigcap_{\substack{u\in\mathcal{H}\\ \|u\|_\mathcal{H}=1}}\left\{v\in\mathcal{H} : \langle u, v \rangle \leq \sigma_{\mathcal{I}_1}(u)\right\} \eqqcolon \mathcal{S}_1, \end{equation} where \[ \sigma_{\mathcal{I}_1}(u)\coloneqq \E{\langle u, \Psi(Y)\rangle\mid D = 1, S = 1, \langle u, \Psi(Y)\rangle \geq Q_{\langle u, \Psi(Y)\rangle}(1-p)}, \] and $Q_{\langle u, \Psi(Y)\rangle}(1-p)$ denotes the $(1-p)$-th conditional quantile of $\langle u, \Psi(Y)\rangle$ given $D = 1$ and $S = 1$. In other words, $\mathcal{S}_1$ is a valid identified set of $\E{\Psi(Y_{1})|S_1=S_0=1}$. • $\Psi^{-1}(\mathcal{S}_1)$ is a valid identified set of $\mu_{1\oplus}$. • When $\Psi(Y)$ is finite dimensional, it holds $\mathcal{I}_1=\mathcal{S}_1$, i.e., $\mathcal{S}_1$ is the sharp identified set of $\E{\Psi(Y_{1})|S_1=S_0=1}$. Furthermore, $\Psi^{-1}(\mathcal{S}_1)$ is the sharp identified set of $\mu_{1\oplus}$. \end{description}

Theorem (ref) establishes a Lee-type identified set for the conditional Fr\'echet mean $\mu_{1\oplus}$ of general random objects. Theorem (ref) (iii) shows that, when $\Psi(Y)$ is finite-dimensional, the proposed procedure delivers a sharp identified set for $\mu_{1\oplus}$. Representing a closed convex set as in (ref) is standard in econometrics (e.g., BeMo08), and this representation is computationally feasible in practice. Further details are discussed in Section (ref). We note that incorporating covariates is straightforward (see Remark (ref) below).

When $\Psi(Y)$ is infinite-dimensional, evaluating (ref) may be infeasible, and $\mathcal{S}_1$ need not coincide with the sharp identified set for $\mu_{1\oplus}$. Nevertheless, one may still be able to obtain a sharp identified set over a finite grid of points by applying the same procedure to a finite-dimensional projection of $\Psi(Y)$, which is useful to approximate $\mathcal{S}_1$ or to obtain a feasible, valid identified set. See Remark (ref) below for the case of distributional outcomes.

remark[Covariates] Suppose that a baseline covariates vector $X$ with support $\mathcal{X}\subset\mathbb{R}^k$ is available. The previous argument is valid conditional on $X=x$ so that we can obtain the identified set $\mathcal{S}_1(x)$ for each $x\in\mathcal{X}$. Then the identified set of $\E{\Psi(Y_{1})|S_1=S_0=1}$ is given by $\mathcal{S}_1^X= \{\E{s(X)\mid S=1,D=0}: s(\cdot)\text{ is measurable and }s(x) \in \mathcal{S}_1(x),\,\forall x\in\mathcal{X}\}$, or equivalently, \begin{align*} \mathcal{S}_1^X = \bigcap_{\substack{u\in\mathcal{H}\\ \|u\|_\mathcal{H}=1}}\left\{v\in\mathcal{H} : \langle u, v \rangle \leq \E{\sigma_X(u) \mid S=1,D=0}\right\}, \end{align*} where $\sigma_x(u)$ is the conditional analog of $\sigma_{\mathcal{I}_1}(u)$. The identified set of $\mu_{1\oplus}$ can be obtained by $\Psi^{-1}(\mathcal{S}_1^X)$. When the covariates are discrete, or when continuous covariates are discretized as in Lee:2009, the identified set is simply given by \begin{align*} \mathcal{S}_1^X = \bigcap_{\substack{u\in\mathcal{H}\\ \|u\|_\mathcal{H}=1}}\left\{v\in\mathcal{H} : \langle u, v \rangle \leq \sum_{x\in\mathcal{X}} w_x\sigma_{x}(u)\right\}, \end{align*} where $w_x = \P{X=x \mid S=1, D=0}$. $\square$
remark[Distributional outcome] Recall the one-dimensional distributional outcome case discussed in Example (ref) (i.e., $\mathcal{Y} = \mathcal{W}$). Then $\Psi(\mathcal{Y})$ can be regarded as the space of quantile functions. Take a finite number of $0<q_1<\cdots<q_k<1$ “evaluation points” and consider a map $\Pi : \mathcal{Y} \ni Y \mapsto (F_Y^{-1}(q_1),\ldots, F_Y^{-1}(q_k)) \in \mathbb{R}^k$. Note that the image space $\Pi(\mathcal{Y})=\{(x_1,\dots,x_k) \in \mathbb{R}^k: -\infty<x_1\leq \dots \leq x_k<\infty\}$ is a closed convex subset of $\mathbb{R}^k$. As a result, by applying Theorem (ref) (iii), we can derive the sharp bounds on $\E{\Pi(Y_{1})|S_1=S_0=1}$ by \begin{align*} \mathcal{S}_{1,\Pi}= \bigcap_{\substack{u\in\mathbb{R}^k\\ \|u\|=1}}\left\{v\in\mathbb{R}^k : \langle u, v \rangle \leq \sigma(u)\right\}, \end{align*} where $\sigma(u)= \mathbb{E}[\langle u, \Pi(Y)\rangle\mid D = 1, S = 1, \langle u, \Pi(Y)\rangle \geq Q_{\langle u, \Pi(Y)\rangle}(1-p)]$. Then, we can construct the following valid identified set of $\E{\Psi(Y_{1})|S_1=S_0=1}$: \begin{align*} \mathcal{I}_1 \subset \left\{f\in\Psi(\mathcal{W}) : \left(f(q_1),\ldots,f(q_k)\right) \in \mathcal{S}_{1,\Pi}\right\}. \end{align*} In practice, although the right-hand side is still infeasible, we can approximate and visualize this identified set by taking a large $k$, picking a fine grid of points in $\mathcal{S}_{1,\Pi}$, and then drawing the linearly interpolated values $f(q_1),\ldots,f(q_k)$. Alternatively, one can rely on a feasible superset $[L(\cdot), U(\cdot)]$, where $L(q)\coloneqq s_j^\ell$ if $q\in[q_j,q_{j+1})$ and $U(q)\coloneqq s_{j+1}^u$ if $q\in(q_j,q_{j+1}]$, and $[s_j^\ell, s_j^u]$ denotes the projection of $\mathcal{S}_{1,\Pi}$ onto the $j$-th element.$\square$

We conclude our identification analysis with a remark on the contaminated and corrupted data models in the sense of Horowitz_Manski:1995.

remark[Contaminated and corrupted data model] We begin with the standard contaminated data model. Suppose we are interested in the mean of $Y_1 \in \mathbb{R}$, but only observe the contaminated sample $Y = Y_1 Z + Y_0 (1-Z) \in \mathbb{R}$, where $Z\in\{0,1\}$ is the (non-)contamination indicator taking zero if the realization is contaminated. Suppose the contamination $Z$ is independent of our interest $Y_1$. Horowitz_Manski:1995 provides the sharp bounds on $\E{Y_1}$ under the assumption that $\P{Z=0}\leq \lambda <1$, where $\lambda$ is a known or identified constant. Our analysis in this section, for the case where $Y,Y_1,Y_0\in\mathcal{Y}$, immediately applies to this context. In particular, we have that $\mathbb{E}_{\oplus}[Y_1] \in \Psi^{-1}(\mathcal{S})$, where \begin{equation*} \mathcal{S} = \bigcap_{\substack{u\in\mathcal{H}\\ \|u\|_\mathcal{H}=1}}\left\{v\in\mathcal{H} : \langle u, v \rangle \leq \sigma(u)\right\}, \end{equation*} with $\sigma(u)= \E{\langle u, \Psi(Y)\rangle\mid \langle u, \Psi(Y)\rangle \geq Q_{\langle u, \Psi(Y)\rangle}(\lambda)}$, and $Q_{\langle u, \Psi(Y)\rangle}(q)$ denotes the $q$-th quantile of $\langle u, \Psi(Y)\rangle$. Suppose that $Z$ may be correlated with $Y_1$ (i.e., corrupted data case) but we know that $Y_1\in\mathcal{Y}_1 \subset \mathcal{Y}$ with known $\mathcal{Y}_1$ such that $\Psi(\mathcal{Y}_1)$ is closed and convex. The previous argument suggests that $\E{\Psi(Y_1)\mid Z=1} \in \mathcal{S}$, and hence we obtain the identified set \begin{align*} \E{\Psi(Y_1)} \in (1-\P{Z=0}) \mathcal{S} \oplus \P{Z=0} \Psi(\mathcal{Y}_1) \subset (1-\lambda) \mathcal{S} \oplus \lambda \Psi(\mathcal{Y}_1), \end{align*} where the last inclusion uses the convexity of $\Psi(\mathcal{Y}_1)$ and $\mathcal{S}\subset\Psi(\mathcal{Y}_1)$. $\square$

Characterizing treatment effects

Once the sharp identified set for $\mu_{1\oplus}$ is obtained, the remaining question is how to summarize the treatment effect. Unlike the Euclidean-outcome case, when outcomes take values in a general metric space $\mathcal{Y}$, one cannot in general define treatment effects by algebraic differencing of potential means, since $\mathcal{Y}$ need not admit a linear structure. Nevertheless, researchers can define treatment effects in ways that are tailored to their substantive objectives. Indeed, there can be several approaches to defining causal effects, which we list below.

Difference in the embedded space. The first possibility is to define the causal effect through differences in an embedded space. For a general metric space that can be embedded into a Hilbert space $\mathcal{H}$, we can define the contrast $\Psi(\mu_{1\oplus})-\Psi(\mu_{0\oplus})$ as the causal effect of interest, provided that the embedded object $\Psi(Y)$ has a reasonable interpretation.

A leading example is the case in which $\mathcal{Y}$ is the $2$-Wasserstein space for distributional outcomes. In this setting, it is common to work with quantile-function representations and to define a quantile-treatment-effect-type parameter as the difference in the embedded space of quantile functions $\Psi(\mathcal{Y})$ (e.g., LiKoWa23 and va:25):

align*[align* omitted — 128 chars of source]

In this case, the identified set for the embedded contrast can be written as the Minkowski difference

align*[align* omitted — 139 chars of source]

Difference in the projections. A second option is to focus on a real-valued contrast obtained by projecting the mean objects $\mu_{0\oplus}$ and $\mu_{1\oplus}$ onto $\mathbb{R}$. For example, suppose $Y$ is a compositional outcome and the researcher's interest is in the percentage-point change in the first component. Let $\Pi:\mathcal{Y}\to\mathbb{R}$ denote the projection onto the first coordinate, i.e., $\Pi(y)=y_1$ for $y=(y_1,\ldots,y_k)\in\mathcal{Y}$. Then a natural causal estimand is $\Pi(\mu_{1\oplus})-\Pi(\mu_{0\oplus})$. More generally, the same construction applies whenever $\Pi:\mathcal{Y}\to\mathbb{R}$ has a meaningful interpretation in the application at hand. In this case, an identified set for the projected contrast is given by

align*[align* omitted — 205 chars of source]

Geodesics. As a third option, one may employ the geodesic average treatment effect (GATE) proposed by kurisu_etal:24. The GATE is defined as the entire geodesic path from $\mu_{0\oplus}$ to $\mu_{1\oplus}$:

align*[align* omitted — 229 chars of source]

Intuitively, for $\alpha, \beta\in\mathcal{Y}$, the geodesic $\gamma_{\alpha,\beta}$ is defined as the shortest path from $\alpha$ to $\beta$. For a precise definition of geodesics in a general metric space and for the relationship between GATE and the conventional ATE defined for Euclidean outcomes, see kurisu_etal:24. The full geodesic includes both the direction and the magnitude of the change induced by the treatment. This provides insight into how the underlying object evolves along the shortest intrinsic path. Based on the identified set $\Psi^{-1}(\mathcal{S}_1)$ for $\mu_{1\oplus}$, the identified set of the GATE $\gamma_{\mu_{0\oplus},\mu_{1\oplus}}$ is given by the following set of geodesics: \[ \left\{ \gamma_{\mu_{0\oplus},\mu} : \mu\in\Psi^{-1}(\mathcal{S}_1) \right\}. \] Provided that $\mathcal{S}_1$ is sharp, the induced identified sets for these causal estimands are sharp as well.

Implementation

This section provides detailed estimation and inference procedures for finite-dimensional outcomes. In the infinite-dimensional case, the same approach applies after projecting the object of interest onto a finite-dimensional parameter.

Estimation

We begin with the estimation procedure. Recalling that

align*[align* omitted — 117 chars of source]

the Fr\'echet mean $\mu_{0\oplus}$ can be estimated by its sample analog, or $\Psi^{-1}(\widehat{\mathbb{E}}[\Psi(Y) \mid S=1, D=0])$, where $\widehat{\mathbb{E}}[\Psi(Y) \mid S=1, D=0])$ is the standard sample conditional mean in the Euclidean space. The sharp identified set of $\mu_{1\oplus}$ can be estimated by

align[align omitted — 189 chars of source]

where $\mathbb{S}^{d-1}=\{x \in \mathbb{R}^d:\|x\|=1\}$ is the $(d-1)$-dimensional unit sphere,

align*[align* omitted — 372 chars of source]

$\widehat{p} = \{(\sum_{i=1}^n S_i(1-D_i))/(\sum_{i=1}^n (1-D_i))\}/\{(\sum_{i=1}^n S_i D_i)/(\sum_{i=1}^n D_i)\}$, and $\widehat{Q}_{\langle u, \Psi(Y)\rangle}(q)$ denotes the $q$-th sample conditional quantile of $\langle u, \Psi(Y)\rangle$ given $D = 1, S = 1$. In practice, the intersection “$\bigcap_{u\in \mathbb{S}^{d-1}}$” in (ref) can be approximated arbitrarily well by evaluating (ref) on a sufficiently fine grid of points on the unit sphere $\mathbb{S}^{d-1}$. When $d$ is small, one may use a grid with (approximately) equal angular spacing. For example, see BeMoMo10 for the case $d=2$; for $d=3$, one may employ a Fibonacci lattice. In higher-dimensional cases, an approximation based on random number generation is often more convenient. For instance, if $Z\sim \mathcal{N}(\mathbf{0}, \mathbf{I}_d)$, then $Z/\|Z\|$ is uniformly distributed on the $(d-1)$-dimensional unit sphere Devroye:1986, and thus can be used to form an approximation of the intersection.

Inference

This subsection discusses a method for uncertainty quantification. For $\mu_{0\oplus}$, the quantity $\E{\Psi(Y)\mid S=1,D=0}$ is point identified and corresponds to a standard Euclidean conditional mean. Accordingly, conventional procedures (e.g., the bootstrap) apply, and one can construct a confidence region for $\E{\Psi(Y)\mid S=1,D=0}$, denoted by $\widehat{\mathcal{R}}_0$, in a standard manner. The corresponding confidence region for $\mu_{0\oplus}$ is then given by $\Psi^{-1}(\widehat{\mathcal{R}}_0)$.

We therefore focus on uncertainty quantification for the sharp identified set $\mathcal{S}_1$. Let $\alpha\in(0,1)$ be a researcher-specified significance level. For each $b\in\{1,\ldots, B\}$, compute

align[align omitted — 194 chars of source]

where $\widetilde{\sigma}_{\mathcal{I}_1}^{(b)}(u)$ is the bootstrap counterpart of $\widehat{\sigma}_{\mathcal{I}_1}(u)$, and $\widehat{V}_{\sigma}(u)$ is the estimator of the asymptotic variance $V_{\sigma}(u)$ of $\widehat{\sigma}_{\mathcal{I}_1}(u)$, which is given by Lee:2009. In practice, this supremum is approximated by the maximum over some fine grid points. Then, we compute the critical value

align*[align* omitted — 116 chars of source]

and obtain the confidence region of $\mathcal{S}_1$ as

align*[align* omitted — 227 chars of source]

where $\mathbb{S}^{d-1}$ is the $(d-1)$-dimensional unit sphere. The following proposition establishes the validity of this procedure.

assumption\quad \begin{description} • $\Psi(Y)\mid (S=1,D=1)$ has a bounded, closed, and convex support $K\subset\mathbb{R}^d$. Moreover, $\Psi(Y)$ admits a conditional density $f_{\Psi(Y)}$ given $(S=1,D=1)$ that is continuous on $K$ and satisfies $0<\underline{f} \leq f_{\Psi(Y)}(x) \leq \overline{f} <\infty$ for all $x\in K$. • $0<\E{SD}$, $0<\E{D}<1$, $\E{S\mid D=0}<\E{S\mid D=1}<1$, and $\inf_{u\in\mathbb{S}^{d-1}} V_{\sigma}(u)>0$. • $(Y_1, Y_0, S_1, S_0, D)$ is i.i.d. across individuals. \end{description}
propositionUnder Assumptions (ref) and (ref), it holds $\lim_{n\to\infty}\mathbb{P}[\mathcal{S}_1 \subset \widehat{\mathcal{R}}_1] = 1-\alpha$.

In view of the proof of Proposition (ref), the joint confidence region for $(\Psi(\mu_{0\oplus}), \mathcal{S}_1)$ is also readily available. Instead of (ref), define

align*[align* omitted — 343 chars of source]

where $\widehat{\xi}(j)$ and $\widetilde{\xi}^{(b)}(j)$ are the sample- and bootstrap- estimators of $\xi(j)\coloneqq$ the $j$-th element of $\E{\Psi(Y)\mid S=1,D=0}$, and ${V}_{0}(j)$ is the asymptotic variance of $\widehat{\xi}(j)$ with a consistent estimator $\widehat{V}_{0}(j)$. Then, using the critical value $\widehat{\text{j.c.v.}} =$ the $(1-\alpha)$-th quantile of $\{T_{\mathtt{joint}}(b) : b=1,\ldots,B\}$, one can construct a joint confidence region such that

align[align omitted — 216 chars of source]

where

align*[align* omitted — 303 chars of source]

and $\widehat{\mathcal{R}}_{1,\mathtt{joint}}$ is constructed as before but replacing $\widehat{\text{c.v.}}$ with $\widehat{\text{j.c.v.}}$.

corollaryIn addition to the assumptions made in Proposition (ref), suppose that $\E{S(1-D)}>0$, $\Psi(Y)\mid (S=1,D=0)$ has a bounded support, and $\min_{j\in\{1,\ldots,d\}} V_0(j)>0$. Then, (ref) holds.

This general construction of the joint confidence region will be useful for conducting a hypothesis test of the treatment effects (see Section (ref)). When the treatment effect is defined as the difference in the embedded space $\Psi(\mathcal{Y})$, the $1-\alpha$ confidence region for the treatment effect is given by $\widehat{\mathcal{R}}_{1,\mathtt{joint}} \ominus \widehat{\mathcal{R}}_{0,\mathtt{joint}}$. If the treatment effect is measured by the projection difference, $\Pi(\Psi^{-1}(\widehat{\mathcal{R}}_{1,\mathtt{joint}})) \ominus \Pi(\Psi^{-1}(\widehat{\mathcal{R}}_{0,\mathtt{joint}}))$ provides a valid confidence interval. When the geodesic treatment effect is employed, $\{ \gamma_{\nu,\mu} : \nu \in \Psi^{-1}(\widehat{\mathcal{R}}_{0,\mathtt{joint}})\,\text{and}\, \mu\in\Psi^{-1}(\widehat{\mathcal{R}}_{1,\mathtt{joint}})\}$ has a correct asymptotic coverage.

Empirical illustration

This section illustrates our proposed method using a distributional outcome: sleep duration (hours slept per night). In particular, we study how individuals' sleep responds to financial incentives. Sleep is a central component of time allocation and a potentially important input into economic performance (e.g., Rao_etal:2021, Streatfeild_etal:2021), and recent work in economics has examined whether incentives can shift sleep behavior (e.g., Avery_etal:2025, Bessone_etal:2021). We use a subset of the data from Bessone_etal:2021 and reanalyze the experiment with our method, focusing on distributional changes in sleep hours.

In Bessone_etal:2021, low-income adults in India were employed in a paid data-entry job and wore an actigraphy-based wearable device throughout the study. After eight baseline days, participants were randomly assigned to one of three arms (control, sleep-aid devices$+$incentives, and sleep-aid devices$+$encouragement) and were further cross-randomized to a workplace nap offer. In the incentives arm, participants received sleep-improving devices and were offered monetary incentives tied to increases in actigraph-measured sleep relative to baseline. In this illustration, we focus on the incentives-versus-control comparison and pool over the nap assignment. As a result, the sample sizes are 152 in the control group and 150 in the treatment (incentives) group.

As the distributional sleep outcome, we use the nightly sleeping hours from days 9-28 and treat the resulting set of hours for each individual as a distributional outcome. For illustration, we classify an individual as a non-respondent if sleep data are missing for at least three nights out of these 20 nights. An assumption here is that, when the number of missing nights is below the threshold, missingness is at random and therefore does not distort the distribution; otherwise, the missingness is potentially endogenous. Under this setup, the attrition rate is 12.5% in the control group and 10.7% in the treatment group.

figure[figure omitted — 727 chars of source]

We estimate the identified set following the procedure described in Remark (ref) and Section (ref). We choose $k=17$ evaluation points, with $q_1=0.10, q_2=0.15, \ldots, q_{17}=0.90$, and estimate $\mathcal{S}_{1,\Pi}$. We then randomly draw $200{,}000$ points $\{(f(q_1),\ldots,f(q_{17}))_j\}_{j=1}^{200{,}000}$ from $\widehat{\mathcal{S}}_{1,\Pi}$, and construct $200{,}000$ linearly interpolated curves connecting $f(q_1),\ldots,f(q_{17})$. The 95% joint confidence region is computed following Corollary (ref) and visualized in the same manner as $\widehat{\mathcal{S}}_{1,\Pi}$.

The estimation results are shown in Figure (ref). We observe several notable features. First, the estimated identified set (inner, purple band) is tight and informative. This is because $\widehat{p}=(1-0.125)/(1-0.107)\approx 0.98$ is close to one, implying that the observed distribution is only mildly “contaminated” and the identified set is therefore close to point identification. Second, the identified set lies above the control mean (red curve), and the joint confidence region shows no overlap between the treatment and control groups. These findings suggest that providing incentives shifts the distribution of sleep upward. Moreover, the observations are consistent with $\mu_{1\oplus}$ first-order stochastically dominating $\mu_{0\oplus}$, which is a richer implication than a comparison of average sleep duration alone.